The lowest LLM API price per million tokens is not necessarily the lowest cost for a production workload. Input and output volume, cached context, context length, reasoning, retries, batch processing, and task completion can all change the final bill.
This LLM API pricing comparison covers representative models from OpenAI API pricing, Anthropic Claude pricing, Google Gemini API pricing, and DeepSeek API pricing. It also shows how to calculate monthly API costs, compare models by workload, evaluate cost per useful result, and decide when API access or self-hosting makes more economic sense.
Pricing note — Prices in this article were checked on October 8, 2026. Figures use standard paid API rates in USD per 1 million tokens unless stated otherwise. Provider pricing can change after publication, so the official pricing documentation should be checked before making a purchasing decision.
Methodology — Official provider pricing is used as the primary source. Rates are compared on a per-million-token basis, with separate conditions noted where pricing changes by context length, cache status, processing mode, or time period. The comparison focuses on representative production models rather than every available API model.
LLM API Pricing Comparison
The table focuses on representative current models rather than attempting to list every available model. A larger model list does not necessarily produce a more useful comparison because the economic result depends on workload, context, caching, output volume, and processing requirements.

For OpenAI, the table uses the standard short-context production rate. OpenAI lists separate long-context rates, so models with large context windows should be evaluated against the applicable context tier rather than treated as having one universal token price.
| Model | Provider | Input | Cached Input | Output | Context | Typical Use Case |
|---|---|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0.10 | $0.01 | $0.60 | 1.05M | Cost-sensitive, high-volume tasks |
| GPT-5.6 Sol | OpenAI | $2.00 | $0.20 | $10.00 | 1.05M | Complex production workloads |
| GPT-6 Astra | OpenAI | $5.00 | $0.50 | $25.00 | 1.05M | Advanced reasoning and coding |
| Claude Haiku 5.5 | Anthropic | $0.10* | $0.01* | $0.50* | 1M | Classification and extraction |
| Claude Sonnet 5.5 | Anthropic | $2.00 | $0.10 | $10.00 | 1M | General production and agents |
| Claude Opus 5.5 | Anthropic | $4.00 | $0.20 | $20.00 | 1M | Advanced coding and knowledge work |
| Claude Fable 5.1 | Anthropic | $10.00 | $0.25 | $50.00 | 1M | Demanding reasoning and agents |
| Gemini 3.8 Flash | $0.75† | $0.075† | $3.75† | Large | High-volume coding and agents | |
| Gemini 3.1 Pro | $2.00‡ | $0.20‡ | $12.00‡ | Large | Multimodal and complex reasoning | |
| DeepSeek V4.1 Flash | DeepSeek | $0.15§ | $0.003§ | $0.60§ | 1M | Cost-sensitive high-volume workloads |
* Claude Haiku 5.5 uses the listed rate for prompts up to 100,000 input tokens. Higher prompt lengths use higher input, output, and cache rates.
† Gemini 3.8 Flash uses its introductory standard rate through December 31, 2026. The listed price increases to $1.50 input and $7.50 output on January 1, 2027.
‡ Gemini 3.1 Pro uses $2 input and $12 output for prompts up to 200,000 tokens. Above that threshold, the rates increase to $4 input and $18 output.
§ DeepSeek V4.1 Flash rates shown are off-peak rates for cache-miss input and output. Peak rates are $0.30 input and $1.20 output.
OpenAI's current pricing table lists GPT-5.6 Luna at $0.10 per million input tokens and $0.60 per million output tokens under its short-context rate, GPT-5.6 Sol at $2 and $10, and GPT-6 Astra at $5 and $25. The same pricing table provides separate long-context rates.
Anthropic currently lists Claude Haiku 5.5 from $0.10 per million input tokens and $0.50 per million output tokens, Claude Sonnet 5.5 at $2 and $10, Claude Opus 5.5 at $4 and $20, and Claude Fable 5.1 at $10 and $50. Anthropic also publishes separate cache and Batch pricing.
Google currently lists Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Gemini 3.1 Pro uses a separate pricing tier above 200,000 input tokens.
DeepSeek currently uses deepseek-flash for V4.1 Flash. Its official pricing distinguishes cache hits, cache misses, peak periods, and off-peak periods.
How LLM API Pricing Works
Most LLM APIs charge according to the number and type of tokens processed. The main cost categories are input tokens, output tokens, cached input, and in some cases additional processing or context tiers.
Input Tokens
Input tokens include the information sent to the model, such as system instructions, user prompts, conversation history, retrieved documents, tool results, and other context.
Low input pricing is particularly relevant when an application sends large amounts of context but generates relatively short responses.
Output Tokens
Output tokens are generated by the model and are often priced higher than input tokens.
This makes response length a major cost variable for chatbots, coding agents, content generation, and other output-heavy workloads.
For example, GPT-5.6 Luna uses a $0.10 input rate and a $0.60 output rate under the standard short-context pricing tier. GPT-6 Astra uses $5 input and $25 output under the same tier.
Two applications using the same model can therefore have very different bills if one generates substantially more output.
Cached Input
Cached input is designed for repeated context.
Typical examples include long system instructions, recurring documents, stable tool definitions, and repeated conversation context.
Anthropic currently charges Claude Sonnet 5.5 $0.10 per million cache-read tokens compared with $2 for standard input. Claude Fable 5.1 has a cache-read rate of $0.25 compared with $10 for standard input.
OpenAI also publishes separate cached input rates for its current models.
The economic benefit depends on how much of the application's context is repeated and how often that context produces cache hits.
What Actually Determines Your LLM API Cost
The rate card tells you the unit price. Your workload determines how many billable tokens and calls you generate.
Input and Output Ratio
Consider two applications with the same request volume.
The first sends 1,000 input tokens and generates 100 output tokens. The second sends 1,000 input tokens and generates 2,000 output tokens.
The second application is much more sensitive to output pricing.
This makes input-only comparisons unreliable for workloads that generate long responses.
Request Volume
A single request price does not tell you what a production application will spend each month.
For example, 100,000 requests with an average of 1,000 input tokens and 500 output tokens produce:
100 million input tokens
50 million output tokens
At that scale, even a small difference in per-million-token pricing can become significant.
Context Length
Long-context applications can consume large amounts of input tokens even when request volume is relatively low.
This matters for:
Document analysis
RAG applications
Coding agents
Legal research
Large conversation histories
Multi-document workflows
Context window size also does not tell the whole cost story.
Google's Gemini 3.1 Pro currently charges $2 per million input tokens for prompts up to 200,000 tokens and $4 above that threshold. Output pricing increases from $12 to $18 per million tokens beyond the same threshold.
OpenAI similarly separates short-context and long-context rates. For GPT-5.6 Luna, the current pricing table lists $0.10 input and $0.60 output for short context, compared with $0.20 input and $0.90 output for long context.
A larger context window can therefore expand what the model can process without making large-context requests automatically economical.
Reasoning Usage
Reasoning models can generate additional thinking tokens during processing.
Billing treatment varies by provider and model. Google, for example, includes thinking tokens in output pricing for supported Gemini models.
For reasoning-heavy applications, visible answer length may therefore underestimate the generation that affects the bill.
Retries and Additional Calls
One user request does not necessarily equal one model call.
An agent may call a model to interpret the task, select a tool, process its result, verify an answer, and recover from an unsuccessful step.
Each additional call can add token usage.
The relevant unit for cost analysis is therefore often the workflow, not the individual API request.
LLM API Cost per Successful Task
The cheapest model per token is not necessarily the cheapest model per successful outcome.
Consider this illustrative example.
| Model | Cost per Attempt | Success Rate | Expected Attempts |
|---|---|---|---|
| Model A | $1.00 | 80% | 1.25 |
| Model B | $1.80 | 98% | 1.02 |
Model A has the lower price per attempt, but repeated attempts reduce its cost advantage.
A simple way to think about this is
Effective cost = API spend ÷ successful outcomes
This is an analytical framework rather than a universal accounting standard.
There is also a difference between provider failure and task failure.
A provider failure can mean a timeout, rejected request, or service error. A task failure means the API returned an answer, but the result did not meet the application's requirements.
Task failures can trigger retries, fallback models, human review, or additional processing, all of which affect total cost.
This distinction is increasingly reflected in provider guidance. Anthropic explicitly recommends comparing models by cost per completed task because a more capable model can sometimes finish a task with fewer turns, searches, context rereads, and backtracking.
Cost per Useful Output
Some workloads cannot be evaluated with a simple success or failure metric.
Content generation, coding, research, summarization, and agent workflows can produce technically valid responses that still require substantial human editing.
For these applications, a more useful question is
How much does it cost to produce an output that is actually usable?
That cost can include API usage, human review, engineering effort, moderation, retries, and downstream processing.
A higher-priced model can therefore be economical when its additional capability reduces the amount of corrective work or additional model calls required.
How to Calculate Your Monthly LLM API Cost
The basic calculation is
Monthly input cost = input tokens ÷ 1,000,000 × input price
Monthly output cost = output tokens ÷ 1,000,000 × output price
Monthly API cost = monthly input cost + monthly output cost
A realistic estimate should then account for cached input, Batch processing, long-context pricing, reasoning usage, retries, and additional model calls.
Calculate Your LLM API Cost
Start with four measurements.
| Input | What to Measure |
|---|---|
| Monthly input tokens | Total tokens sent to the model |
| Monthly output tokens | Total tokens generated |
| Input price | Provider rate per 1M tokens |
| Output price | Provider rate per 1M tokens |
Then separate cached and uncached input where the provider uses different rates.
For production planning, also estimate the percentage of requests that require retries, fallback models, or multiple agent calls. A token-only estimate can understate the cost of multi-step workflows.
Blended LLM API Cost
A blended cost combines input and output pricing using a defined workload ratio.
For example, a workload with 3 million input tokens and 1 million output tokens has a different effective token cost from a workload with 1 million input tokens and 3 million output tokens, even when both use the same model.
A simple blended rate can be calculated as
Blended cost per 1M tokens = total API cost ÷ total tokens processed
The ratio must always be stated because changing the input-to-output mix changes the result.
This is also why two pricing comparison sites can rank the same models differently. One may assume an input-heavy workload, while another may use a generation-heavy workload.
Blended cost is therefore useful for comparison, but it should not replace a workload-specific monthly estimate.
Monthly Cost Example
Assume a workload generates 100 million input tokens and 50 million output tokens each month.
The following estimates use the listed standard rates. They assume no caching, Batch discounts, retries, or extra tool charges. The Gemini 3.1 Pro estimate assumes prompts remain at or below 200,000 input tokens, while the DeepSeek figures show both off-peak and peak rates.
| Model | Input Cost | Output Cost | Estimated Monthly Cost |
|---|---|---|---|
| GPT-5.6 Luna | $10 | $30 | $40 |
| GPT-5.6 Sol | $200 | $500 | $700 |
| GPT-6 Astra | $500 | $1,250 | $1,750 |
| Claude Haiku 5.5 | $10 | $25 | $35 |
| Claude Sonnet 5.5 | $200 | $500 | $700 |
| Claude Opus 5.5 | $400 | $1,000 | $1,400 |
| Claude Fable 5.1 | $1,000 | $2,500 | $3,500 |
| Gemini 3.8 Flash | $75 | $187.50 | $262.50 |
| Gemini 3.1 Pro | $200 | $600 | $800 |
| DeepSeek V4.1 Flash, off-peak | $15 | $30 | $45 |
| DeepSeek V4.1 Flash, peak | $30 | $60 | $90 |
These calculations illustrate rate-card differences rather than forecast actual production bills. Real costs can change with cache hit rates, token distributions, context tiers, retries, tool calls, and model routing.
The OpenAI figures use the standard short-context pricing tier. Long-context requests can use higher rates, so the GPT-5.6 Luna example should not be applied to every 100M-token workload without checking the context distribution.
The DeepSeek figures demonstrate why time-based pricing matters. Off-peak rates are half of peak rates, so the same token volume costs twice as much during the listed peak periods.
LLM API Pricing by Workload
There is no single cheapest LLM API for every application.
| Workload | Cost Factor to Prioritize |
|---|---|
| Classification | Input price and throughput |
| Data extraction | Input price and output length |
| Chat assistants | Input and output ratio |
| Long-document analysis | Context and input pricing |
| Coding agents | Output, reasoning, retries, and tool calls |
| RAG systems | Context reuse and caching |
| Batch processing | Batch pricing and latency |
| Multi-turn agents | Repeated context and output |
| High-volume automation | Unit cost and throughput |
High-Volume Classification
Classification often uses relatively small prompts and short outputs.
A lower-cost model can therefore be attractive when evaluation shows that it meets the required quality level without excessive retries or escalation.
Long-Document Analysis
Long-document workflows are more sensitive to input volume and context pricing.
A large context window can reduce the need to split documents, but the same capability can increase the number of input tokens billed per request.
The useful comparison is context capacity combined with the price at the context size your workload actually requires.
Coding and Agent Workloads
Coding agents can generate large amounts of output and make multiple calls during one task.
Output pricing, reasoning behavior, tool usage, retries, and cache reuse can therefore matter more than the headline input rate.
Batch Processing
Batch workloads introduce a latency tradeoff.
If users do not need immediate results, a discounted asynchronous processing tier can reduce the cost of large jobs. If the application requires interactive responses, the lower batch rate may not be usable.
Prompt Caching and Batch Processing
Caching and Batch processing affect API economics in different ways.
Prompt Caching
Prompt caching works best when an application repeatedly sends the same context.
Examples include
Long system instructions
Repeated documents
Stable tool definitions
Multi-turn conversation context
Shared application instructions
The key variable is the cache hit rate. If most requests contain unique context, the potential saving is limited. If a large portion of every request is reused, cached input can materially reduce the effective input cost.
Anthropic currently charges cache reads at a fraction of the standard input rate. Claude Sonnet 5.5 uses $0.10 per million cache-read tokens compared with $2 for standard input, while Claude Fable 5.1 uses $0.25 compared with $10.
OpenAI also provides separate cached input rates.
Batch Processing
Batch APIs trade response time for lower processing costs.
They can be useful for
Offline classification
Document processing
Dataset generation
Evaluation jobs
Large-scale extraction
Scheduled content processing
Anthropic currently offers a 50% discount on input and output tokens through its Batch API. Google also offers Batch pricing at 50% of its standard Gemini 3.8 Flash rate through the current introductory pricing period.
Batch pricing is useful only when the workload can tolerate asynchronous processing.
Which LLM API Is Cheapest?
“Cheapest” changes depending on what is being measured.
| Comparison | Current Example | Important Condition |
|---|---|---|
| Lowest input price | GPT-5.6 Luna, Claude Haiku 5.5 | Context and provider pricing tier matter |
| Lowest output price | GPT-5.6 Luna, Claude Haiku 5.5 | Quality and task completion still matter |
| High-volume workloads | Lower-cost models | Throughput, retries, and concurrency affect total cost |
| Advanced workloads | Depends on task | Quality, reasoning, context, and tool support matter |
| Flexible processing | DeepSeek V4.1 Flash can benefit from off-peak pricing | Workload must tolerate scheduling |
GPT-5.6 Sol and Claude Sonnet 5.5 both use a $2 input and $10 output standard rate. Gemini 3.1 Pro uses $2 input and $12 output for prompts up to 200,000 tokens.
At this level, token price alone becomes a weak selection criterion. The relevant comparison is the cost and quality of completing the required task.
API Pricing vs Self-Hosting
Using an API and hosting a model yourself represent different cost structures.
| Factor | API | Self-Hosting |
|---|---|---|
| Infrastructure | Provider managed | Your responsibility |
| Upfront cost | Usually lower | Usually higher |
| Scaling | Provider managed | You manage capacity |
| Model optimization | Provider managed | Your responsibility |
| Engineering | Lower operational burden | Higher operational burden |
| Maintenance | Provider managed | Your responsibility |
| Low-volume workloads | Often attractive | May be difficult to justify |
| High stable volume | Can become expensive | May become more economical |
Self-hosting becomes more economically attractive when infrastructure utilization is high enough to spread fixed costs across a large, predictable workload.
The break-even point depends on hardware utilization, model size, throughput, uptime, engineering cost, and workload stability.
A self-hosted model therefore does not have a simple equivalent to a provider's price per million tokens. Infrastructure and engineering costs need to be converted into a comparable cost per useful unit of work.
How to Reduce LLM API Costs
Use Smaller Models for Simpler Tasks
Classification, extraction, formatting, simple summarization, and straightforward transformations may not require a premium model.
The decision should come from task evaluation. If a cheaper model produces acceptable results without substantially higher retry or review rates, the lower token price can translate into lower effective cost.
Control Output Length
Output tokens can be considerably more expensive than input tokens.
If an application generates unnecessary explanations or repeated content, reducing output length can lower cost without changing the model.
Response limits and structured output can help when long responses provide little additional value.
Reuse Stable Context
Prompt caching can reduce the cost of repeated system prompts and other stable context.
The strongest candidates are applications with long instructions, repeated documents, or multi-turn interactions where a substantial portion of the input remains unchanged.
Use Batch Processing
Move workloads that do not require real-time responses to Batch processing when the provider offers a meaningful discount.
The relevant calculation is not simply the percentage discount. It is the saving relative to the operational value of faster results.
Route Requests by Complexity
A single model does not need to handle every task.
A multi-model architecture can use lower-cost models for routine requests and reserve more expensive models for tasks that benefit from additional capability.
This can reduce blended cost while preserving access to stronger models when they are actually needed.
Test the Expensive Tail of Your Workload
Average requests can hide the tasks that consume the most money.
Difficult requests may require more output, additional tool calls, retries, fallback models, or human review.
When evaluating a model, test both typical requests and the difficult portion of the workload. A model that looks inexpensive on average may have a different cost profile on the tasks most likely to fail.
Anthropic's current cost guidance explicitly recommends pricing the tail of the workload rather than relying only on median requests. Its examples show that model rankings can change when cost is measured per completed task instead of per token.
Monitor Cost by Feature
A monthly API bill does not tell you which part of an application is expensive.
Track usage by
Model
Feature
Request type
Input tokens
Output tokens
Cache usage
Retry rate
Cost per completed task
This makes it easier to identify expensive workflows and determine whether a cheaper model, shorter output, caching, or routing would actually reduce total cost.
How to Choose an LLM API
A practical selection process can be kept focused.
1. Estimate Monthly Usage
Start with requests, input tokens, and output tokens.
2. Identify the Workload
Determine whether the application is mainly classification, generation, retrieval, coding, reasoning, or agentic execution.
3. Check Context Requirements
Estimate both typical and unusually large context sizes.
4. Compare the Full Pricing Structure
Look beyond standard input and output rates.
Check caching, Batch processing, long-context pricing, rate limits, tool charges, and any time-based pricing.
5. Test Models on Your Own Data
Public benchmarks provide useful context, but your workload can behave differently.
A focused evaluation can compare
Accuracy
Task completion
Output quality
Latency
Token consumption
Retry frequency
Test the difficult tail of the workload as well as the average case.
6. Calculate Effective Cost
Compare API spending against the quality and successful task rate you actually need.
For generation-heavy or agentic workflows, also consider human review and downstream processing.
7. Recheck Pricing Regularly
LLM API pricing changes quickly.
New models, price reductions, deprecations, processing tiers, caching changes, and routing changes can alter the economics of an existing architecture.
DeepSeek is one example of how a model transition can affect API economics. Its current documentation states that the legacy V4 Flash names are routed to V4.1 Flash and billed at the Flash rate, while peak and off-peak pricing are tied to specific UTC windows.
FAQ
Which LLM API is the cheapest?
There is no single cheapest API for every workload. Some models have very low input rates, while others become more competitive when caching, Batch processing, context tiers, or peak and off-peak pricing are considered.
How much does 1 million LLM tokens cost?
The answer depends on the model and whether the tokens are input, output, cached, or subject to another pricing tier. Current mainstream API rates range from well below one dollar to tens of dollars per million tokens.
Is OpenAI cheaper than Claude?
It depends on the models being compared. GPT-5.6 Sol and Claude Sonnet 5.5 currently have the same standard input and output rates at $2 and $10 per million tokens, while other models have substantially different prices.
Is Gemini cheaper than GPT?
Some Gemini models have lower standard rates than comparable OpenAI models, while others are similarly priced or more expensive. Gemini also offers Batch, Flex, Priority, caching, and context-dependent pricing, so the answer depends on the workload.
Are free LLM APIs really free?
Some providers offer free tiers, but free access generally comes with restrictions such as model availability, rate limits, or usage conditions. Production applications should be evaluated using the provider's paid pricing and applicable limits.
How do I calculate monthly LLM API costs?
Multiply monthly input tokens by the input price per million tokens, then multiply monthly output tokens by the output price per million tokens. Add the two amounts and adjust for caching, Batch processing, reasoning usage, retries, and additional model calls.
Is self-hosting cheaper than using an LLM API?
It can become economical at sufficiently high and stable utilization, but there is no universal break-even point. Infrastructure, hardware utilization, engineering, scaling, optimization, and maintenance all need to be included in the comparison.
Final Takeaway
LLM API pricing should be compared at the level of the work being completed, not only at the token level.
A useful evaluation follows
Token price → usage pattern → blended cost → effective cost → useful output → operating cost
A low-cost model can be the right choice when it reliably handles a task at scale. A higher-priced model can also be economical if it completes difficult tasks with fewer retries, shorter outputs, fewer tool calls, or less human intervention.
Caching, Batch processing, model routing, context-aware pricing, and self-hosting can further change the result.
If you are also evaluating AI-generated content as part of an AI workflow, Unfox AI can help you check your own text and compare results across different samples. AI detection results are best treated as one signal rather than conclusive proof, since results can vary with the model, text length, language, writing style, and editing.




