Cheapest LLM API
Cheapest LLM API Calculator
Find the lowest estimated monthly model cost for your workload across major LLM providers. Change request volume, tokens, cached input, and buffer assumptions to see which model is cheapest for that scenario.
Find the cheapest LLM API
Choose a workload, enter monthly requests and token size, then compare provider cost before committing.
- 1Pick workloadUse the closest API usage scenario.
- 2Enter tokensAdd request volume, input, output, and cache.
- 3Compare modelsCheck monthly cost and effective rate.
- 4Share or exportCopy the link or download the CSV.
Next: copy the share link or export the CSV, then request a worksheet or send a pricing correction if the assumptions need review.
| Provider | Model | Input / 1M | Cached / 1M | Output / 1M | Discount | Monthly cost | Cost / request |
|---|
Large prompts change the winner
Long system prompts, examples, and retrieved context make input pricing and cache discounts matter.
Long answers can dominate cost
Report generation, coding, and reasoning tasks can shift spend toward output token pricing.
Add real-world buffer
Retries, prompt experiments, evaluation jobs, and traffic spikes can move real spend above a clean estimate.
How to find the cheapest LLM API for your app
Start with one common task in your product, not your whole application. Estimate monthly requests, average input tokens, average output tokens, and whether repeated context can use cached input pricing. The cheapest model for that task appears at the top of the comparison table.
If you are budgeting a full product instead of one model call, use the AI Cost Calculator. If you want a broader provider comparison, use the LLM API Pricing Comparison.
| Workload pattern | Cost driver | What to test next |
|---|---|---|
| Simple classification or routing | Request count and short outputs | Try low-cost small models first. |
| Chatbot with long instructions | Repeated input context | Check cached input savings. |
| Report or code generation | Output tokens | Compare output pricing and answer length. |
| RAG search answers | Retrieval context plus generation | Estimate RAG cost separately. |
FAQ
What is the cheapest LLM API?
The cheapest LLM API changes with input tokens, output tokens, cached input, and request volume. A small model can be cheapest for simple tasks, while quality requirements may justify a more expensive model.
Should I always pick the lowest-cost model?
No. Use the cheapest model that reliably completes the task. Compare quality, latency, context window, safety behavior, tool support, and data controls before production use.
Sources
Prices in this site were last checked on 2026-07-04. Pricing can change without notice.