Cheapest LLM API

Cheapest LLM API Calculator

Find the lowest estimated monthly model cost for your workload across major LLM providers. Change request volume, tokens, cached input, and buffer assumptions to see which model is cheapest for that scenario.

Find the cheapest LLM API

Choose a workload, enter monthly requests and token size, then compare provider cost before committing.

  1. 1Pick workloadUse the closest API usage scenario.
  2. 2Enter tokensAdd request volume, input, output, and cache.
  3. 3Compare modelsCheck monthly cost and effective rate.
  4. 4Share or exportCopy the link or download the CSV.
Lowest estimated monthly cost$0.00
Lowest-cost model-
Budget with buffer$0.00
Monthly tokens0
Models compared0

Next: copy the share link or export the CSV, then request a worksheet or send a pricing correction if the assumptions need review.

ProviderModelInput / 1MCached / 1MOutput / 1MDiscountMonthly costCost / request
Cheapest is workload-specific: the table sorts by estimated monthly cost for the exact inputs above. Treat the result as a budget shortlist, not a final model recommendation.
Input-heavy

Large prompts change the winner

Long system prompts, examples, and retrieved context make input pricing and cache discounts matter.

Output-heavy

Long answers can dominate cost

Report generation, coding, and reasoning tasks can shift spend toward output token pricing.

Production

Add real-world buffer

Retries, prompt experiments, evaluation jobs, and traffic spikes can move real spend above a clean estimate.

How to find the cheapest LLM API for your app

Start with one common task in your product, not your whole application. Estimate monthly requests, average input tokens, average output tokens, and whether repeated context can use cached input pricing. The cheapest model for that task appears at the top of the comparison table.

If you are budgeting a full product instead of one model call, use the AI Cost Calculator. If you want a broader provider comparison, use the LLM API Pricing Comparison.

Workload patternCost driverWhat to test next
Simple classification or routingRequest count and short outputsTry low-cost small models first.
Chatbot with long instructionsRepeated input contextCheck cached input savings.
Report or code generationOutput tokensCompare output pricing and answer length.
RAG search answersRetrieval context plus generationEstimate RAG cost separately.

FAQ

What is the cheapest LLM API?

The cheapest LLM API changes with input tokens, output tokens, cached input, and request volume. A small model can be cheapest for simple tasks, while quality requirements may justify a more expensive model.

Should I always pick the lowest-cost model?

No. Use the cheapest model that reliably completes the task. Compare quality, latency, context window, safety behavior, tool support, and data controls before production use.

Sources

Prices in this site were last checked on 2026-07-04. Pricing can change without notice.