LLM API Cost Compared

Published per-million-token prices are only half the equation. Tokenizer differences, caching, batching and reasoning tokens change the real number.

Comparing LLM API prices looks easy: vendors publish a rate per million input tokens and per million output tokens, so you put them in a spreadsheet and pick the cheapest. In practice the published rate is only half the equation, and teams that optimise on it alone routinely end up with a bill they cannot explain.

This guide covers the rates, the multipliers that actually determine what you pay, and a costing method that survives contact with a real invoice.

On pricing accuracy. Model lineups and prices change frequently — new model generations ship every few months and prices move with them. The figures below were cross-checked across several public pricing trackers at the time of writing. Verify against the provider's official pricing page before committing to a budget or a volume contract, and treat every number here as a starting point rather than a quote.

The published rates

Prices below are US dollars per one million tokens, standard (non-batch) tier, for models whose pricing is consistent across multiple independent pricing trackers. Models whose published prices conflicted between sources have been deliberately excluded — better to omit than to mislead.

ModelVendorInput / 1MOutput / 1MOutput ÷ input
GPT-4.1 nanoOpenAI$0.10$0.404.0×
Gemini 2.5 Flash-LiteGoogle$0.10$0.404.0×
GPT-4o miniOpenAI$0.15$0.604.0×
Gemini 2.5 FlashGoogle$0.30$2.508.3×
GPT-4.1 miniOpenAI$0.40$1.604.0×
Claude Haiku 3.5Anthropic$0.80$4.005.0×
Claude Haiku 4.5Anthropic$1.00$5.005.0×
Gemini 2.5 ProGoogle$1.25$10.008.0×
o4-miniOpenAI$1.10$4.404.0×
GPT-4.1OpenAI$2.00$8.004.0×
o3OpenAI$2.00$8.004.0×
GPT-4oOpenAI$2.50$10.004.0×
Claude Sonnet 4.6Anthropic$3.00$15.005.0×
Claude Opus 4.xAnthropic$5.00$25.005.0×

Two structural observations jump out. First, the spread between the cheapest and most expensive output rate is roughly 60× — model choice dominates every other decision you will make. Second, the output-to-input ratio varies by vendor far more than by tier: Google's ratios are markedly higher than OpenAI's or Anthropic's. That matters enormously, because it means the cheapest vendor depends on your workload shape.

Four multipliers the price list hides

1. Tokenizer differences

Identical text produces different token counts on different providers, because each trains its own tokenizer. A vendor with a 20% lower headline rate but a tokenizer that emits 25% more tokens for your prompts is more expensive, not less. This is the most commonly missed factor in vendor comparisons.

The fix is empirical and takes an afternoon: take 1,000 representative prompts from your production traffic, run each through every candidate provider's counting endpoint, and compute cost per prompt — not cost per token. The ranking frequently differs from the published price ranking.

2. Long-context tiers

Some providers charge a premium above a context threshold — commonly 200K tokens. If your workload involves very large documents, the headline input rate may not apply to the portion above the threshold. Others charge a flat rate at any context length. For long-document workloads this single line item can dominate the total.

3. Reasoning tokens

Reasoning models generate internal reasoning tokens before producing a visible answer, and those are billed as output tokens. A request whose visible answer is 300 tokens may bill for several thousand. If you are evaluating a reasoning model on price, benchmark total output tokens from the usage object, not the length of the answer you see.

4. Retries and failures

Failed requests still consume input tokens in most implementations, and retry storms during an incident can produce a bill spike that dwarfs normal traffic. Set retry budgets, use exponential backoff with jitter, and alert on token spend per hour — not just on monthly spend.

Three real workloads, costed

Assume 10 million requests per month. The same volume produces wildly different bills depending on shape.

Workload A — Classification (input-heavy)

500 input tokens, 10 output tokens per request. Routing support tickets.

Input:  10M × 500 = 5.0B tokens
Output: 10M × 10  = 0.1B tokens

GPT-4o mini:   (5.0B/1M × $0.15) + (0.1B/1M × $0.60) = $750 + $60   = $810
GPT-4o:        (5.0B/1M × $2.50) + (0.1B/1M × $10.00) = $12,500 + $1,000 = $13,500
Claude Opus:   (5.0B/1M × $5.00) + (0.1B/1M × $25.00) = $25,000 + $2,500 = $27,500

17× difference between the cheapest and most expensive option, for a task where a small model is usually equally accurate. This is where routing pays for itself.

Workload B — Code generation (output-heavy)

800 input tokens, 3,000 output tokens per request.

Input:  10M × 800  = 8.0B tokens
Output: 10M × 3000 = 30B tokens

Gemini 2.5 Pro:  (8.0B/1M × $1.25) + (30B/1M × $10.00) = $10,000 + $300,000 = $310,000
GPT-4.1:         (8.0B/1M × $2.00) + (30B/1M × $8.00)  = $16,000 + $240,000 = $256,000
Claude Sonnet:   (8.0B/1M × $3.00) + (30B/1M × $15.00) = $24,000 + $450,000 = $474,000

Here the ranking inverts relative to the input-only case, because output dominates and output ratios differ by vendor. Note also that at this scale, quality matters more than price: if the cheaper model produces code that needs a retry, the retry costs more than the saving.

Workload C — RAG over long documents (very input-heavy)

40,000 input tokens, 400 output tokens per request.

Input:  10M × 40000 = 400B tokens
Output: 10M × 400   = 4B tokens

GPT-4.1:        (400B/1M × $2.00) + (4B/1M × $8.00)  = $800,000 + $32,000 = $832,000
Gemini 2.5 Pro: (400B/1M × $1.25) + (4B/1M × $10.00) = $500,000 + $40,000 = $540,000
GPT-4.1 nano:   (400B/1M × $0.10) + (4B/1M × $0.40)  = $40,000  + $1,600  = $41,600

At this scale the answer is not "pick a cheaper model" — it is "send fewer tokens". Better chunking, reranking, and prompt caching will each move this number by more than any vendor switch. Cutting retrieved context from 40,000 to 10,000 tokens saves more than switching from the most expensive to the cheapest model.

Model routing: the biggest lever

Nothing else on this page comes close. Most production systems send every request to one model, but request difficulty is heavily skewed: the large majority of traffic is routine, and only a small fraction genuinely needs frontier capability.

A cascade — try the cheap model, escalate on low confidence or failed validation — typically captures 60–80% of the cost of a uniform frontier-model deployment with no measurable quality loss, provided you have a way to detect failure. That detection is the hard part and it is task-specific: schema validation for extraction, a unit test for code, a second small model as judge for free-form text.

Routing also gives you something vendor-level decisions cannot: per-request granularity. You can route by user tier, by task, by time of day, and you can change it without a migration.

Caching and batching

Prompt caching discounts repeated prefixes. If your system prompt and tool schemas are constant — and they are, in almost every application — this is close to free money. The discount on cache reads is steep, often an order of magnitude off base input. The cache hit requires an exact prefix match, so the shared portion must come first in the request and be byte-identical across calls. Ordering your prompt with the stable content at the very top is usually the only change needed.

Batch APIs typically halve the price for work that does not need an interactive response: nightly enrichment, bulk classification, embedding generation, report production. If your workload tolerates latency measured in hours, batch is a guaranteed 50% discount.

Both are orthogonal to model choice, which means they stack with routing. A system that routes, caches and batches can plausibly run at a small fraction of the cost of the naive version of itself.

A cost model you can defend

Do not estimate from average token counts. Build the model from measured distributions:

  1. Sample real traffic. Take 1,000 representative requests. Not synthetic examples — real ones, including the long tail, because the tail dominates cost.
  2. Log usage on every call. Store input tokens, output tokens, model, and a task identifier. This is ground truth and costs you one line of code.
  3. Model the distribution, not the mean. Cost follows a heavy tail. The 99th percentile request can cost 50× the median, and your monthly bill is driven disproportionately by that tail.
  4. Multiply by retry rate. Measure it. It is rarely 1.0.
  5. Add a growth term. Usage grows; a model that is accurate at today's volume is useless at ten times the volume.
  6. Re-validate quarterly. Prices move and model lineups change.

Our AI Token Counter lets you run the per-request arithmetic on your real prompts before you commit to an architecture.

Five budgeting mistakes

Optimising on input price when you are output-bound

Code generation, long-form writing and agentic loops are output-heavy. Comparing vendors on input rate alone will point you the wrong way.

Ignoring the tokenizer

A 20% cheaper rate with 25% more tokens is a loss. Benchmark with your own prompts.

Forgetting the system prompt

A 2,000-token system prompt on 10M requests is 20 billion input tokens. That is a budget line, and system prompts accrete — audit them periodically.

No per-request cost attribution

If you cannot attribute spend to a feature, a customer or a team, you cannot optimise it. Tag every call and aggregate by tag. This is also what lets you charge fairly if you resell.

Treating the estimate as a ceiling

Development traffic is not production traffic. Real users write longer inputs, ask follow-ups, and hit your retry paths. Whatever your pilot cost, production will differ — usually upward.

Frequently asked questions

It depends entirely on your workload shape. Input-heavy workloads favour low input rates; output-heavy workloads are dominated by output rates, whose ratios to input differ substantially by vendor. Benchmark your own prompts rather than picking from a price table.

Generation is sequential and compute-bound — each token requires a full forward pass and cannot be produced ahead of the previous one. Ingestion is parallelisable across the whole prompt. The price difference reflects real compute asymmetry.

It depends on what fraction of your input is a repeated prefix. For applications with a large, constant system prompt or tool schema, cache-read discounts are steep enough to change the economics materially. Put stable content first in the request so the prefix matches exactly.

They were cross-checked across public pricing trackers at the time of writing, but model lineups and prices change frequently. Always verify on the provider's official pricing page before committing to a budget.

No. On narrow, well-specified tasks — classification, extraction, routing, formatting — small models frequently match frontier models at a fraction of the cost. The gap widens on tasks requiring multi-step reasoning, long context synthesis, or nuanced writing.