Back to Blog
AI Models & Tokens

How to Calculate OpenAI API Token Usage & Cost

Learn how to calculate OpenAI API token usage, cost per request, cached tokens, and monthly API spending.

Unfox AI

Unfox AI

Content Team

14 min read

To calculate OpenAI API cost, multiply each token category by its applicable rate, then add the results.

Input cost = input tokens ÷ 1,000,000 × input price per 1M tokens

Output cost = output tokens ÷ 1,000,000 × output price per 1M tokens

For a basic text request, adding these amounts gives you an estimated cost. Cached input, reasoning tokens, long-context pricing, processing modes, or paid tools can require additional calculations.

A reliable workflow is to estimate tokens before a request, measure actual usage after it, then verify aggregate usage as your application scales.

ai token cta1

What Counts as a Token in the OpenAI API?

A token is a unit that a language model uses to process text. It can represent a complete word, part of a word, punctuation, or a character.

A tokenizer might keep a common word as one token while splitting a less common word into several pieces. Language, punctuation, formatting, and model encoding can therefore change the token count for text with the same number of words.

For common English text, OpenAI gives a rough estimate of about four characters or three-quarters of a word per token. That makes 100 tokens roughly 75 words, but this relationship should be used for planning rather than exact API calculations.

Use the level of measurement appropriate to the task.

  • Word or character count for a rough early estimate

  • A tokenizer for a better estimate before sending a request

  • API usage data for actual request-level usage

For cost calculations, use token counts whenever the actual prompt is available rather than assuming a fixed word-to-token conversion.

Types of OpenAI API Tokens and How They Are Billed

The text typed by a user is not necessarily the entire billable input. A complete request can contain several token categories with different rates.

OpenAI API input output cached and reasoning token billing flow

Input Tokens

Input tokens can include

  • System or developer instructions

  • User prompts

  • Previous conversation messages

  • Retrieved documents or RAG context

  • Few-shot examples

  • Tool definitions

  • Structured output schemas

A 20-token user question can therefore produce a request containing hundreds or thousands of input tokens if the application also sends conversation history, instructions, or retrieved context.

Output Tokens

Output tokens are generated by the model.

Input and output should be calculated separately because their rates can differ. A request using 1,000 input tokens and 1,000 output tokens should not be treated as 2,000 identically priced tokens.

This difference becomes more important in high-volume applications or workloads that generate long responses.

Cached Input and Reasoning Tokens

OpenAI Prompt Caching can change the rate applied to eligible input tokens when prompt prefixes are reused.

On models with explicit cache-write pricing, input tokens may be billed as uncached input, cache writes, or cache reads. A cache write is not an additional input charge applied to the same tokens.

Caching is not automatically cheaper. A cache write can cost more than ordinary input on supported models, while later cache reads can cost substantially less. The total benefit depends on whether the cached prefix receives enough reuse.

Reasoning models can generate reasoning tokens that are not visible in the final response. These tokens count as output tokens for billing and context-window accounting, so a short visible answer does not necessarily mean low output usage.

How to Estimate OpenAI Token Usage Before an API Call

The appropriate estimation method depends on whether you already have the text that will be sent to the API.

Use Text Length for a Rough Estimate

If the final prompt does not exist yet, word or character count can support early budgeting.

Using the approximate relationship of 100 tokens to 75 common English words, 750 words might be around 1,000 tokens. Code, numbers, punctuation, non-English languages, formatting, unusual vocabulary, and model encoding can move the actual count away from that estimate.

Use a Tokenizer for a Better Estimate

Once the actual prompt exists, count its tokens rather than estimating from words.

Developers can use tools such as Tiktoken for programmatic plain-text tokenization. OpenAI also provides input-token counting for complete Responses API inputs, which can account for request structure beyond plain text.

The Unfox AI Token Calculator can also estimate token usage from prompts, documents, and other text before an API request.

Plain-text counting does not necessarily capture the complete request. Messages, tools, schemas, images, files, and other structured inputs can affect actual API usage.

How to Calculate OpenAI API Cost

Once you have token counts, calculate each billing category separately using the rate for the exact model your application uses.

Step 1 — Find the Current Model Price

Check the official OpenAI API pricing for the exact model or model ID.

Depending on the model and workload, pricing can distinguish between

  • Input

  • Cached input

  • Cache writes

  • Output

  • Short and long context

  • Processing modes

Current production budgets should use current official rates rather than prices copied from older articles.

Step 2 — Calculate Input Cost

Use

Input cost = input tokens ÷ 1,000,000 × input rate

If input tokens fall into multiple billing categories, calculate each group at its applicable rate. Do not charge the same tokens once as ordinary input and again as cache writes or reads.

Step 3 — Calculate Output Cost

Use

Output cost = output tokens ÷ 1,000,000 × output rate

For reasoning models, output usage can include reasoning tokens that do not appear in the visible response.

Step 4 — Add Other Applicable Usage

For a basic request

Estimated cost = input cost + output cost

If the workload uses paid tools or other separately billed capabilities, calculate those charges separately and add them to the token cost.

OpenAI API Cost Calculation Example

The following rates are hypothetical and used only to demonstrate the calculation. Check current OpenAI pricing before applying the method to a real workload.

Suppose a text model costs $2 per million input tokens and $10 per million output tokens.

ComponentToken UsageHypothetical RateEstimated Cost
Input4,000$2 per 1M$0.008
Output800$10 per 1M$0.008
Total4,800—$0.016

At 5,000 similar requests per day

Daily cost = $0.016 × 5,000 = $80

At 30 days per month

Monthly cost = $80 × 30 = $2,400

The per-request figure is more useful for forecasting than the headline price per million tokens because it incorporates the input-output mix of the actual workload.

How to Check Actual Token Usage After an API Call

Pre-request token counting is an estimate. After the request, use the usage data returned by the API to measure what was actually processed.

For the Responses API, usage can include

  • input_tokens

  • output_tokens

  • total_tokens

Chat Completions uses

  • prompt_tokens

  • completion_tokens

  • total_tokens

In Python, you can inspect the response usage object.

print(response.usage)

Depending on the endpoint and model, detailed usage data can also expose cached or reasoning token information.

Use three levels of measurement for cost tracking.

  1. Estimate before the request with a tokenizer or token calculator.

  2. Measure individual requests with API usage data.

  3. Verify aggregate consumption with the OpenAI Usage Dashboard or Usage API.

A gap between these levels can reveal request components or production behavior that the original estimate did not capture.

Why Your Actual OpenAI API Cost May Be Higher Than Your Estimate

Correct arithmetic can still produce a poor forecast if the assumed request does not represent the production workload.

Cost DriverWhy It Changes Usage
Conversation historyEarlier messages may become input again
Long system promptsRepeated instructions add input tokens
RAG contextRetrieved documents increase request size
Reasoning usageInternal reasoning can add output usage
RetriesRepeated requests consume additional tokens
Multi-step workflowsOne user action can trigger several API calls
Model changesDifferent models use different rates
ToolsSome capabilities introduce separate billing
Long contextSome models apply different rates above context thresholds

Conversation history illustrates the problem clearly.

Turn 1

System instructions + User 1 → Output 1

Turn 2

System instructions + User 1 + Output 1 + User 2 → Output 2

Turn 3

Previous context + User 3 → Output 3

Even if each new user message remains short, the input can grow as previous context is sent again.

RAG creates a similar effect. A one-sentence question can become a much larger API request when the application adds system instructions and retrieved document passages.

Production forecasts should therefore use complete representative requests rather than user-message length alone.

How to Estimate Monthly OpenAI API Cost

Once representative requests have been measured, monthly cost can be estimated with

Monthly cost = average cost per request × monthly requests

For a user-based application

Monthly cost = average cost per request × requests per user × active users × active days

The main source of uncertainty is usually the average request cost, not the multiplication.

Sample the workloads your application actually handles. Include meaningful variations such as short and long prompts, different conversation depths, and common RAG request sizes.

Measure actual API usage across those samples, calculate an average or useful range, then apply realistic request volume.

After deployment, compare the forecast with aggregate usage. If they diverge, check whether prompt length, context size, model selection, retry frequency, or workflow structure changed before revising the budget assumption.

Practical Ways to Reduce OpenAI API Token Costs

Cost optimization should target token usage that does not improve the required result.

Choose a Model Based on Cost per Completed Task

Price per million tokens does not show the full cost of a workflow.

A lower-priced model can lose its cost advantage if it needs additional retries, longer prompts, or multiple calls to complete a task that another model handles in one request. Compare representative end-to-end task costs rather than token rates alone.

Remove Context That Does Not Improve the Result

Old conversation messages, irrelevant retrieved passages, excessive examples, and repeated instructions increase input usage.

For RAG applications, evaluate whether each retrieved passage improves the answer rather than maximizing the amount of context sent to the model.

Control Unnecessary Output

Match the output allowance to the task.

A classification label or compact JSON response generally does not need the same output budget as a long-form generation task. Appropriate output limits reduce unnecessary generated tokens.

Use Prompt Caching Where It Fits

Prompt caching is most useful when eligible prompt prefixes are reused across requests.

Keep stable instructions, examples, tool definitions, and other repeated content in reusable prefixes when the workload supports caching, then monitor actual cache usage rather than assuming repetition will produce cache hits.

Consider the Batch API for Non-Urgent Work

The OpenAI Batch API currently offers a 50 percent cost discount compared with synchronous APIs and completes batches within a 24-hour turnaround window.

It can fit workloads such as large-scale classification, evaluations, embeddings, and other processing that does not require an immediate response.

A Simple OpenAI API Cost Checklist

Before treating an API estimate as a production budget, confirm

  • The exact API model or model ID

  • Current official pricing

  • Typical complete input-token usage

  • Typical output-token usage

  • Cache writes or reads

  • Reasoning-token usage

  • Conversation-history growth

  • API calls triggered by one user action

  • Separately billed tools or capabilities

  • Expected daily or monthly request volume

  • Request-level actual usage

  • Aggregate usage compared with the forecast

Missing several of these variables means the estimate is still based on incomplete workload assumptions.

Final Thoughts

A useful OpenAI API cost estimate requires both pre-request calculation and post-request measurement.

Estimate token usage from the real prompt, apply the current rates for each relevant billing category, then measure request-level usage through the API. Use those measured requests to forecast volume and verify the result against aggregate usage as the application scales.

The Unfox AI Token Calculator can help estimate token usage from prompts or documents before you apply current OpenAI API pricing.

ai token cta2

FAQ

How much does the OpenAI API cost per token?

There is no single OpenAI API price per token. Rates depend on the model, token category, context conditions, and processing mode.

OpenAI typically publishes text pricing per one million tokens, so use the current rate for the exact API model and billing category in your calculation.

How do I calculate OpenAI API token usage?

Use a tokenizer or token-counting tool to estimate usage before a request. After the request, use the API response's usage data to measure actual input, output, and total token counts.

How many OpenAI tokens are in 1,000 words?

Using OpenAI's rough estimate for common English text, 1,000 words is approximately 1,333 tokens.

Actual counts vary with language, punctuation, formatting, vocabulary, and model encoding, so use a tokenizer when the text is available.

Do input and output tokens cost the same?

Not necessarily. OpenAI models can apply different rates to input and output tokens.

Calculate each category separately using the current pricing for the model you use.

Are cached OpenAI tokens cheaper?

Cache reads can be cheaper than ordinary input on supported models, but caching is not automatically cheaper overall.

On models with explicit cache-write pricing, the initial write can cost more than uncached input. The total cost depends on how often the cached prefix is subsequently reused.

Does ChatGPT Plus include OpenAI API usage?

No. ChatGPT subscriptions and OpenAI API billing are separate.

API activity is billed separately according to the applicable API pricing.

Unfox AI

Written by Unfox AI

Content Team

Passionate about creating exceptional content and sharing knowledge with the community.

Related Articles

How ChatGPT Tokens Work
1 min read

How ChatGPT Tokens Work

Learn how ChatGPT tokens work, from tokenization and token IDs to context limits, output generation, and API usage.

Ready to start your next project?

Join thousands of developers who are already building amazing applications with our platform.