Back to Blog
AI Models & Tokens

LLM API Pricing Comparison for 2026

Compare OpenAI, Claude, Gemini, and DeepSeek API pricing, token costs, monthly estimates, caching, and batch rates.

Unfox AI

Unfox AI

Content Team

26 min read

The lowest LLM API price per million tokens is not necessarily the lowest cost for a production workload. Input and output volume, cached context, context length, reasoning, retries, batch processing, and task completion can all change the final bill.

This LLM API pricing comparison covers representative models from OpenAI API pricing, Anthropic Claude pricing, Google Gemini API pricing, and DeepSeek API pricing. It also shows how to calculate monthly API costs, compare models by workload, evaluate cost per useful result, and decide when API access or self-hosting makes more economic sense.

Pricing note — Prices in this article were checked on October 8, 2026. Figures use standard paid API rates in USD per 1 million tokens unless stated otherwise. Provider pricing can change after publication, so the official pricing documentation should be checked before making a purchasing decision.

Methodology — Official provider pricing is used as the primary source. Rates are compared on a per-million-token basis, with separate conditions noted where pricing changes by context length, cache status, processing mode, or time period. The comparison focuses on representative production models rather than every available API model.

The lowest LLM API price per million tokens is not necessarily the lowest cost for a production workload. Input and output volume, cached context, context length, reasoning, retries, batch processing, and task completion can all change the final bill.

LLM API Pricing Comparison

The table focuses on representative current models rather than attempting to list every available model. A larger model list does not necessarily produce a more useful comparison because the economic result depends on workload, context, caching, output volume, and processing requirements.

LLM API pricing comparison by input, cached input, output, and context

For OpenAI, the table uses the standard short-context production rate. OpenAI lists separate long-context rates, so models with large context windows should be evaluated against the applicable context tier rather than treated as having one universal token price.

ModelProviderInputCached InputOutputContextTypical Use Case
GPT-5.6 LunaOpenAI$0.10$0.01$0.601.05MCost-sensitive, high-volume tasks
GPT-5.6 SolOpenAI$2.00$0.20$10.001.05MComplex production workloads
GPT-6 AstraOpenAI$5.00$0.50$25.001.05MAdvanced reasoning and coding
Claude Haiku 5.5Anthropic$0.10*$0.01*$0.50*1MClassification and extraction
Claude Sonnet 5.5Anthropic$2.00$0.10$10.001MGeneral production and agents
Claude Opus 5.5Anthropic$4.00$0.20$20.001MAdvanced coding and knowledge work
Claude Fable 5.1Anthropic$10.00$0.25$50.001MDemanding reasoning and agents
Gemini 3.8 FlashGoogle$0.75†$0.075†$3.75†LargeHigh-volume coding and agents
Gemini 3.1 ProGoogle$2.00‡$0.20‡$12.00‡LargeMultimodal and complex reasoning
DeepSeek V4.1 FlashDeepSeek$0.15§$0.003§$0.60§1MCost-sensitive high-volume workloads

* Claude Haiku 5.5 uses the listed rate for prompts up to 100,000 input tokens. Higher prompt lengths use higher input, output, and cache rates.

† Gemini 3.8 Flash uses its introductory standard rate through December 31, 2026. The listed price increases to $1.50 input and $7.50 output on January 1, 2027.

‡ Gemini 3.1 Pro uses $2 input and $12 output for prompts up to 200,000 tokens. Above that threshold, the rates increase to $4 input and $18 output.

§ DeepSeek V4.1 Flash rates shown are off-peak rates for cache-miss input and output. Peak rates are $0.30 input and $1.20 output.

OpenAI's current pricing table lists GPT-5.6 Luna at $0.10 per million input tokens and $0.60 per million output tokens under its short-context rate, GPT-5.6 Sol at $2 and $10, and GPT-6 Astra at $5 and $25. The same pricing table provides separate long-context rates.

Anthropic currently lists Claude Haiku 5.5 from $0.10 per million input tokens and $0.50 per million output tokens, Claude Sonnet 5.5 at $2 and $10, Claude Opus 5.5 at $4 and $20, and Claude Fable 5.1 at $10 and $50. Anthropic also publishes separate cache and Batch pricing.

Google currently lists Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Gemini 3.1 Pro uses a separate pricing tier above 200,000 input tokens.

DeepSeek currently uses deepseek-flash for V4.1 Flash. Its official pricing distinguishes cache hits, cache misses, peak periods, and off-peak periods.

How LLM API Pricing Works

Most LLM APIs charge according to the number and type of tokens processed. The main cost categories are input tokens, output tokens, cached input, and in some cases additional processing or context tiers.

Input Tokens

Input tokens include the information sent to the model, such as system instructions, user prompts, conversation history, retrieved documents, tool results, and other context.

Low input pricing is particularly relevant when an application sends large amounts of context but generates relatively short responses.

Output Tokens

Output tokens are generated by the model and are often priced higher than input tokens.

This makes response length a major cost variable for chatbots, coding agents, content generation, and other output-heavy workloads.

For example, GPT-5.6 Luna uses a $0.10 input rate and a $0.60 output rate under the standard short-context pricing tier. GPT-6 Astra uses $5 input and $25 output under the same tier.

Two applications using the same model can therefore have very different bills if one generates substantially more output.

Cached Input

Cached input is designed for repeated context.

Typical examples include long system instructions, recurring documents, stable tool definitions, and repeated conversation context.

Anthropic currently charges Claude Sonnet 5.5 $0.10 per million cache-read tokens compared with $2 for standard input. Claude Fable 5.1 has a cache-read rate of $0.25 compared with $10 for standard input.

OpenAI also publishes separate cached input rates for its current models.

The economic benefit depends on how much of the application's context is repeated and how often that context produces cache hits.

What Actually Determines Your LLM API Cost

The rate card tells you the unit price. Your workload determines how many billable tokens and calls you generate.

Input and Output Ratio

Consider two applications with the same request volume.

The first sends 1,000 input tokens and generates 100 output tokens. The second sends 1,000 input tokens and generates 2,000 output tokens.

The second application is much more sensitive to output pricing.

This makes input-only comparisons unreliable for workloads that generate long responses.

Request Volume

A single request price does not tell you what a production application will spend each month.

For example, 100,000 requests with an average of 1,000 input tokens and 500 output tokens produce:

  • 100 million input tokens

  • 50 million output tokens

At that scale, even a small difference in per-million-token pricing can become significant.

Context Length

Long-context applications can consume large amounts of input tokens even when request volume is relatively low.

This matters for:

  • Document analysis

  • RAG applications

  • Coding agents

  • Legal research

  • Large conversation histories

  • Multi-document workflows

Context window size also does not tell the whole cost story.

Google's Gemini 3.1 Pro currently charges $2 per million input tokens for prompts up to 200,000 tokens and $4 above that threshold. Output pricing increases from $12 to $18 per million tokens beyond the same threshold.

OpenAI similarly separates short-context and long-context rates. For GPT-5.6 Luna, the current pricing table lists $0.10 input and $0.60 output for short context, compared with $0.20 input and $0.90 output for long context.

A larger context window can therefore expand what the model can process without making large-context requests automatically economical.

Reasoning Usage

Reasoning models can generate additional thinking tokens during processing.

Billing treatment varies by provider and model. Google, for example, includes thinking tokens in output pricing for supported Gemini models.

For reasoning-heavy applications, visible answer length may therefore underestimate the generation that affects the bill.

Retries and Additional Calls

One user request does not necessarily equal one model call.

An agent may call a model to interpret the task, select a tool, process its result, verify an answer, and recover from an unsuccessful step.

Each additional call can add token usage.

The relevant unit for cost analysis is therefore often the workflow, not the individual API request.

LLM API Cost per Successful Task

The cheapest model per token is not necessarily the cheapest model per successful outcome.

Consider this illustrative example.

ModelCost per AttemptSuccess RateExpected Attempts
Model A$1.0080%1.25
Model B$1.8098%1.02

Model A has the lower price per attempt, but repeated attempts reduce its cost advantage.

A simple way to think about this is

Effective cost = API spend ÷ successful outcomes

This is an analytical framework rather than a universal accounting standard.

There is also a difference between provider failure and task failure.

A provider failure can mean a timeout, rejected request, or service error. A task failure means the API returned an answer, but the result did not meet the application's requirements.

Task failures can trigger retries, fallback models, human review, or additional processing, all of which affect total cost.

This distinction is increasingly reflected in provider guidance. Anthropic explicitly recommends comparing models by cost per completed task because a more capable model can sometimes finish a task with fewer turns, searches, context rereads, and backtracking.

Cost per Useful Output

Some workloads cannot be evaluated with a simple success or failure metric.

Content generation, coding, research, summarization, and agent workflows can produce technically valid responses that still require substantial human editing.

For these applications, a more useful question is

How much does it cost to produce an output that is actually usable?

That cost can include API usage, human review, engineering effort, moderation, retries, and downstream processing.

A higher-priced model can therefore be economical when its additional capability reduces the amount of corrective work or additional model calls required.

How to Calculate Your Monthly LLM API Cost

The basic calculation is

Monthly input cost = input tokens ÷ 1,000,000 × input price

Monthly output cost = output tokens ÷ 1,000,000 × output price

Monthly API cost = monthly input cost + monthly output cost

A realistic estimate should then account for cached input, Batch processing, long-context pricing, reasoning usage, retries, and additional model calls.

Calculate Your LLM API Cost

Start with four measurements.

InputWhat to Measure
Monthly input tokensTotal tokens sent to the model
Monthly output tokensTotal tokens generated
Input priceProvider rate per 1M tokens
Output priceProvider rate per 1M tokens

Then separate cached and uncached input where the provider uses different rates.

For production planning, also estimate the percentage of requests that require retries, fallback models, or multiple agent calls. A token-only estimate can understate the cost of multi-step workflows.

Blended LLM API Cost

A blended cost combines input and output pricing using a defined workload ratio.

For example, a workload with 3 million input tokens and 1 million output tokens has a different effective token cost from a workload with 1 million input tokens and 3 million output tokens, even when both use the same model.

A simple blended rate can be calculated as

Blended cost per 1M tokens = total API cost ÷ total tokens processed

The ratio must always be stated because changing the input-to-output mix changes the result.

This is also why two pricing comparison sites can rank the same models differently. One may assume an input-heavy workload, while another may use a generation-heavy workload.

Blended cost is therefore useful for comparison, but it should not replace a workload-specific monthly estimate.

Monthly Cost Example

Assume a workload generates 100 million input tokens and 50 million output tokens each month.

The following estimates use the listed standard rates. They assume no caching, Batch discounts, retries, or extra tool charges. The Gemini 3.1 Pro estimate assumes prompts remain at or below 200,000 input tokens, while the DeepSeek figures show both off-peak and peak rates.

ModelInput CostOutput CostEstimated Monthly Cost
GPT-5.6 Luna$10$30$40
GPT-5.6 Sol$200$500$700
GPT-6 Astra$500$1,250$1,750
Claude Haiku 5.5$10$25$35
Claude Sonnet 5.5$200$500$700
Claude Opus 5.5$400$1,000$1,400
Claude Fable 5.1$1,000$2,500$3,500
Gemini 3.8 Flash$75$187.50$262.50
Gemini 3.1 Pro$200$600$800
DeepSeek V4.1 Flash, off-peak$15$30$45
DeepSeek V4.1 Flash, peak$30$60$90

These calculations illustrate rate-card differences rather than forecast actual production bills. Real costs can change with cache hit rates, token distributions, context tiers, retries, tool calls, and model routing.

The OpenAI figures use the standard short-context pricing tier. Long-context requests can use higher rates, so the GPT-5.6 Luna example should not be applied to every 100M-token workload without checking the context distribution.

The DeepSeek figures demonstrate why time-based pricing matters. Off-peak rates are half of peak rates, so the same token volume costs twice as much during the listed peak periods.

LLM API Pricing by Workload

There is no single cheapest LLM API for every application.

WorkloadCost Factor to Prioritize
ClassificationInput price and throughput
Data extractionInput price and output length
Chat assistantsInput and output ratio
Long-document analysisContext and input pricing
Coding agentsOutput, reasoning, retries, and tool calls
RAG systemsContext reuse and caching
Batch processingBatch pricing and latency
Multi-turn agentsRepeated context and output
High-volume automationUnit cost and throughput

High-Volume Classification

Classification often uses relatively small prompts and short outputs.

A lower-cost model can therefore be attractive when evaluation shows that it meets the required quality level without excessive retries or escalation.

Long-Document Analysis

Long-document workflows are more sensitive to input volume and context pricing.

A large context window can reduce the need to split documents, but the same capability can increase the number of input tokens billed per request.

The useful comparison is context capacity combined with the price at the context size your workload actually requires.

Coding and Agent Workloads

Coding agents can generate large amounts of output and make multiple calls during one task.

Output pricing, reasoning behavior, tool usage, retries, and cache reuse can therefore matter more than the headline input rate.

Batch Processing

Batch workloads introduce a latency tradeoff.

If users do not need immediate results, a discounted asynchronous processing tier can reduce the cost of large jobs. If the application requires interactive responses, the lower batch rate may not be usable.

Prompt Caching and Batch Processing

Caching and Batch processing affect API economics in different ways.

Prompt Caching

Prompt caching works best when an application repeatedly sends the same context.

Examples include

  • Long system instructions

  • Repeated documents

  • Stable tool definitions

  • Multi-turn conversation context

  • Shared application instructions

The key variable is the cache hit rate. If most requests contain unique context, the potential saving is limited. If a large portion of every request is reused, cached input can materially reduce the effective input cost.

Anthropic currently charges cache reads at a fraction of the standard input rate. Claude Sonnet 5.5 uses $0.10 per million cache-read tokens compared with $2 for standard input, while Claude Fable 5.1 uses $0.25 compared with $10.

OpenAI also provides separate cached input rates.

Batch Processing

Batch APIs trade response time for lower processing costs.

They can be useful for

  • Offline classification

  • Document processing

  • Dataset generation

  • Evaluation jobs

  • Large-scale extraction

  • Scheduled content processing

Anthropic currently offers a 50% discount on input and output tokens through its Batch API. Google also offers Batch pricing at 50% of its standard Gemini 3.8 Flash rate through the current introductory pricing period.

Batch pricing is useful only when the workload can tolerate asynchronous processing.

Which LLM API Is Cheapest?

“Cheapest” changes depending on what is being measured.

ComparisonCurrent ExampleImportant Condition
Lowest input priceGPT-5.6 Luna, Claude Haiku 5.5Context and provider pricing tier matter
Lowest output priceGPT-5.6 Luna, Claude Haiku 5.5Quality and task completion still matter
High-volume workloadsLower-cost modelsThroughput, retries, and concurrency affect total cost
Advanced workloadsDepends on taskQuality, reasoning, context, and tool support matter
Flexible processingDeepSeek V4.1 Flash can benefit from off-peak pricingWorkload must tolerate scheduling

GPT-5.6 Sol and Claude Sonnet 5.5 both use a $2 input and $10 output standard rate. Gemini 3.1 Pro uses $2 input and $12 output for prompts up to 200,000 tokens.

At this level, token price alone becomes a weak selection criterion. The relevant comparison is the cost and quality of completing the required task.

API Pricing vs Self-Hosting

Using an API and hosting a model yourself represent different cost structures.

FactorAPISelf-Hosting
InfrastructureProvider managedYour responsibility
Upfront costUsually lowerUsually higher
ScalingProvider managedYou manage capacity
Model optimizationProvider managedYour responsibility
EngineeringLower operational burdenHigher operational burden
MaintenanceProvider managedYour responsibility
Low-volume workloadsOften attractiveMay be difficult to justify
High stable volumeCan become expensiveMay become more economical

Self-hosting becomes more economically attractive when infrastructure utilization is high enough to spread fixed costs across a large, predictable workload.

The break-even point depends on hardware utilization, model size, throughput, uptime, engineering cost, and workload stability.

A self-hosted model therefore does not have a simple equivalent to a provider's price per million tokens. Infrastructure and engineering costs need to be converted into a comparable cost per useful unit of work.

How to Reduce LLM API Costs

Use Smaller Models for Simpler Tasks

Classification, extraction, formatting, simple summarization, and straightforward transformations may not require a premium model.

The decision should come from task evaluation. If a cheaper model produces acceptable results without substantially higher retry or review rates, the lower token price can translate into lower effective cost.

Control Output Length

Output tokens can be considerably more expensive than input tokens.

If an application generates unnecessary explanations or repeated content, reducing output length can lower cost without changing the model.

Response limits and structured output can help when long responses provide little additional value.

Reuse Stable Context

Prompt caching can reduce the cost of repeated system prompts and other stable context.

The strongest candidates are applications with long instructions, repeated documents, or multi-turn interactions where a substantial portion of the input remains unchanged.

Use Batch Processing

Move workloads that do not require real-time responses to Batch processing when the provider offers a meaningful discount.

The relevant calculation is not simply the percentage discount. It is the saving relative to the operational value of faster results.

Route Requests by Complexity

A single model does not need to handle every task.

A multi-model architecture can use lower-cost models for routine requests and reserve more expensive models for tasks that benefit from additional capability.

This can reduce blended cost while preserving access to stronger models when they are actually needed.

Test the Expensive Tail of Your Workload

Average requests can hide the tasks that consume the most money.

Difficult requests may require more output, additional tool calls, retries, fallback models, or human review.

When evaluating a model, test both typical requests and the difficult portion of the workload. A model that looks inexpensive on average may have a different cost profile on the tasks most likely to fail.

Anthropic's current cost guidance explicitly recommends pricing the tail of the workload rather than relying only on median requests. Its examples show that model rankings can change when cost is measured per completed task instead of per token.

Monitor Cost by Feature

A monthly API bill does not tell you which part of an application is expensive.

Track usage by

  • Model

  • Feature

  • Request type

  • Input tokens

  • Output tokens

  • Cache usage

  • Retry rate

  • Cost per completed task

This makes it easier to identify expensive workflows and determine whether a cheaper model, shorter output, caching, or routing would actually reduce total cost.

How to Choose an LLM API

A practical selection process can be kept focused.

1. Estimate Monthly Usage

Start with requests, input tokens, and output tokens.

2. Identify the Workload

Determine whether the application is mainly classification, generation, retrieval, coding, reasoning, or agentic execution.

3. Check Context Requirements

Estimate both typical and unusually large context sizes.

4. Compare the Full Pricing Structure

Look beyond standard input and output rates.

Check caching, Batch processing, long-context pricing, rate limits, tool charges, and any time-based pricing.

5. Test Models on Your Own Data

Public benchmarks provide useful context, but your workload can behave differently.

A focused evaluation can compare

  • Accuracy

  • Task completion

  • Output quality

  • Latency

  • Token consumption

  • Retry frequency

Test the difficult tail of the workload as well as the average case.

6. Calculate Effective Cost

Compare API spending against the quality and successful task rate you actually need.

For generation-heavy or agentic workflows, also consider human review and downstream processing.

7. Recheck Pricing Regularly

LLM API pricing changes quickly.

New models, price reductions, deprecations, processing tiers, caching changes, and routing changes can alter the economics of an existing architecture.

DeepSeek is one example of how a model transition can affect API economics. Its current documentation states that the legacy V4 Flash names are routed to V4.1 Flash and billed at the Flash rate, while peak and off-peak pricing are tied to specific UTC windows.

FAQ

Which LLM API is the cheapest?

There is no single cheapest API for every workload. Some models have very low input rates, while others become more competitive when caching, Batch processing, context tiers, or peak and off-peak pricing are considered.

How much does 1 million LLM tokens cost?

The answer depends on the model and whether the tokens are input, output, cached, or subject to another pricing tier. Current mainstream API rates range from well below one dollar to tens of dollars per million tokens.

Is OpenAI cheaper than Claude?

It depends on the models being compared. GPT-5.6 Sol and Claude Sonnet 5.5 currently have the same standard input and output rates at $2 and $10 per million tokens, while other models have substantially different prices.

Is Gemini cheaper than GPT?

Some Gemini models have lower standard rates than comparable OpenAI models, while others are similarly priced or more expensive. Gemini also offers Batch, Flex, Priority, caching, and context-dependent pricing, so the answer depends on the workload.

Are free LLM APIs really free?

Some providers offer free tiers, but free access generally comes with restrictions such as model availability, rate limits, or usage conditions. Production applications should be evaluated using the provider's paid pricing and applicable limits.

How do I calculate monthly LLM API costs?

Multiply monthly input tokens by the input price per million tokens, then multiply monthly output tokens by the output price per million tokens. Add the two amounts and adjust for caching, Batch processing, reasoning usage, retries, and additional model calls.

Is self-hosting cheaper than using an LLM API?

It can become economical at sufficiently high and stable utilization, but there is no universal break-even point. Infrastructure, hardware utilization, engineering, scaling, optimization, and maintenance all need to be included in the comparison.

Final Takeaway

LLM API pricing should be compared at the level of the work being completed, not only at the token level.

A useful evaluation follows

Token price → usage pattern → blended cost → effective cost → useful output → operating cost

A low-cost model can be the right choice when it reliably handles a task at scale. A higher-priced model can also be economical if it completes difficult tasks with fewer retries, shorter outputs, fewer tool calls, or less human intervention.

Caching, Batch processing, model routing, context-aware pricing, and self-hosting can further change the result.

If you are also evaluating AI-generated content as part of an AI workflow, Unfox AI can help you check your own text and compare results across different samples. AI detection results are best treated as one signal rather than conclusive proof, since results can vary with the model, text length, language, writing style, and editing.

AI detection results are best treated as one signal rather than conclusive proof, since results can vary with the model, text length, language, writing style, and editing.

Unfox AI

Written by Unfox AI

Content Team

Passionate about creating exceptional content and sharing knowledge with the community.

Related Articles

How AI Content Detectors Work
1 min read

How AI Content Detectors Work

Learn how AI content detectors analyze writing patterns, calculate scores, and why their results are not definitive.

Ready to start your next project?

Join thousands of developers who are already building amazing applications with our platform.