How AI Token Counting Works
A token is not a word, and the difference between the two is where most LLM cost estimates go wrong. Here is how tokenizers actually split your text.
If you have ever estimated a prompt as "about 500 words, so maybe 700 tokens" and then watched the invoice come in 40% higher, you have run into the gap between how humans segment language and how models do. Closing that gap is not academic — token count is the unit you are billed on, the unit your context window is measured in, and the unit that determines whether your request succeeds at all.
This guide explains what a token is, how the algorithm that produces them works, why the same text yields different counts on different providers, and how to get a number you can take to a budget meeting.
What a token actually is
A token is a sub-word unit drawn from a fixed vocabulary that the model was trained against. The vocabulary is decided before training and never changes afterward. GPT-family models use vocabularies in the range of 100,000–200,000 entries; the model's embedding layer has exactly one row per token, and every piece of text you send must be expressed as a sequence of those rows.
The key property is that tokens are sub-word. Common short words map to a single token. Rarer or longer words are decomposed into several. Consider how a tokenizer handles a family of related English words:
"token" -> ["token"] 1 token
"tokens" -> ["token", "s"] 2 tokens
"tokenize" -> ["token", "ize"] 2 tokens
"tokenization" -> ["token", "iz", "ation"] 3 tokens
"tokenizer" -> ["token", "izer"] 2 tokens
"untokenizable" -> ["un", "token", "iz", "able"] 4 tokens
Notice that the morphological structure falls out of the algorithm — un-,
-ize, -ation, -able recur as units because they
recur in the training corpus. This is not designed in; it is learned from statistics.
Byte-pair encoding, explained
Most modern LLM tokenizers use byte-pair encoding (BPE) or a close variant. The idea is almost embarrassingly simple, and it is worth understanding because it explains every counterintuitive behaviour that follows.
Start with a corpus and split everything into individual characters (or bytes). Count which adjacent pair occurs most frequently. Merge that pair into a single new symbol. Repeat — tens of thousands of times — until the vocabulary reaches its target size. What you end up with is a vocabulary of frequently co-occurring character clusters, from single characters up to common words and word fragments.
Iteration 0: l o w e r
Iteration 1: l o w er ("er" is most frequent pair)
Iteration 2: l ow er ("ow" merges)
Iteration 3: lo w er ("lo" merges)
Iteration 4: low er ("low" merges)
At inference time the encoder does not run merges from scratch. It applies a pre-tokenisation regex that splits text into candidate chunks — words, numbers, punctuation runs — and then applies the learned merge rules to each chunk independently. That pre-tokenisation step is why token boundaries never cross certain boundaries: a word and the space before it may merge into one token, but the trailing punctuation usually will not.
Two consequences follow directly, and they explain most surprises:
- Frequency determines efficiency. Common English text tokenises efficiently because common sequences have dedicated vocabulary entries. Rare technical jargon, unusual names, and random strings do not, and cost more tokens per character.
- Numbers are chunked in groups of up to three digits.
2026is usually two tokens (20,26) rather than one, because the vocabulary was learned from text where leading digits recur far more often than complete multi-digit numbers.
Why models don't just use words
If tokens are awkward, why not use words? Two hard constraints force the answer.
Vocabulary explosion. English has hundreds of thousands of words and morphological variants, and new ones — product names, technical terms, slang — appear constantly. A word-level vocabulary would need millions of entries, each requiring training data to learn a good embedding. Rare words would have terrible representations. Sub-word units keep the vocabulary to a manageable size while still composing any word from known pieces.
Unknown words. At inference time you will inevitably encounter a
string never seen in training. A word-level model has no representation for it and must
fall back to a generic <UNK> token, destroying the meaning. A
byte-level BPE model can always fall back to individual bytes and produce
something reasonable. This is why modern tokenizers are byte-level: the
fallback is guaranteed to exist.
There is a genuine downside, and it explains a famous class of model failures. Because the model sees text as character clusters rather than characters, tasks that require character-level reasoning are hard for it. Asking a model to count the letters in a word, or to reverse a string, or to tell you which character is at position 17, produces unreliable answers — not because the model is bad at reasoning, but because the input representation it operates on does not align with the task. If your application needs character-level manipulation, do it in code and give the model the result.
Why token boundaries surprise you
Leading whitespace is part of the token
In most GPT-family tokenizers, a word preceded by a space is a different
token from the same word without one. " dog" and "dog" have
different IDs. This is why prompts that differ only in whitespace can produce
measurably different token counts and, occasionally, subtly different outputs.
Capitalisation creates new tokens
"Token" and "token" are separate entries. So is
"TOKEN". Text in all caps, or with inconsistent capitalisation, tokenises
less efficiently than cleanly-cased text. This is a real effect: uppercasing a document
can increase its token count by 20–30%.
Non-Latin scripts cost more per character
BPE vocabularies are learned from corpora that are overwhelmingly English. Text in other scripts has far fewer dedicated entries and must be decomposed further. A Chinese character, for instance, typically consumes one to two tokens where an English character averages well under half a token. Cyrillic, Arabic, Devanagari and Korean behave similarly. If you are building a multilingual product, budget for this — the non-English half of your user base can cost two to three times more per character than the English half.
Code and JSON are expensive
Source code tokenises poorly relative to prose. Indentation consumes tokens (whitespace runs have their own entries), identifiers are novel and therefore split into fragments, and repeated punctuation is never merged. A JSON payload is worse still, because every key name is quoted and quoted strings are tokenised character-class by character-class. Empirically, code runs roughly 1.5–2× more tokens per character than natural English prose.
Emoji are very expensive
A single emoji is frequently two to four tokens, because emoji sequences in UTF-8 are multi-byte and the vocabulary has sparse coverage. A "cheap" reaction emoji can cost as much as a short sentence.
The same text, different counts
Each provider trains its own tokenizer on its own corpus with its own vocabulary size and its own merge rules. There is no standard, and the results genuinely differ.
This has a direct commercial consequence that catches teams out. Suppose you are choosing between two providers and one publishes a per-million-token price 20% lower than the other. If its tokenizer emits 25% more tokens for your actual prompts, the cheaper headline price is more expensive in practice. Always benchmark on your own real prompts — the published rate is only half the equation.
The same issue appears mid-migration. When a provider releases a new model with a new tokenizer, your existing prompts may produce noticeably different token counts without you changing a single character. Anthropic's newer model families, for example, moved to a tokenizer that can produce substantially more tokens for identical input than earlier generations. Teams that migrated and kept their old cost model were surprised.
How to get an exact count
Estimation rules of thumb are fine for planning and useless for billing. Here is how to get the real number for each major provider.
OpenAI — tiktoken
import tiktoken
enc = tiktoken.get_encoding("o200k_base") # GPT-4o family
n = len(enc.encode("Your prompt here"))
# Or, for a chat request, count what you will actually send:
enc = tiktoken.encoding_for_model("gpt-4o")
total = sum(len(enc.encode(m["content"])) for m in messages)
Note that encoding_for_model is the safer call, because the correct
encoding varies by model. Encoding with the wrong one silently gives you a plausible
but wrong number.
Anthropic — count_tokens endpoint
import anthropic
client = anthropic.Anthropic()
res = client.messages.count_tokens(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "Your prompt here"}],
)
print(res.input_tokens)
Google Gemini — countTokens
response = model.count_tokens("Your prompt here")
print(response.total_tokens)
The most reliable method of all
Every provider returns a usage object on every response. That number is
what you are billed on, so log it. Aggregating usage across your traffic
gives you ground truth for both your cost model and your capacity planning:
# OpenAI response
response.usage.prompt_tokens # input
response.usage.completion_tokens # output
# Anthropic response
response.usage.input_tokens
response.usage.output_tokens
The usage object also catches tokens you did not write: injected system
preambles, tool schemas your framework serialises for you, and reasoning tokens on
reasoning models. A hand count of your prompt text will never include those, which is
why logged usage is almost always higher than a naive estimate.
From tokens to money
Providers publish separate rates for input and output per million tokens, and output is consistently priced higher — typically 3× to 5× — because generation is sequential and compute-bound while ingestion is parallelisable.
cost = (input_tokens / 1_000_000) * input_price
+ (output_tokens / 1_000_000) * output_price
The ratio between input and output volumes tells you which lever matters. A RAG system that retrieves 40,000 tokens of context and returns a 200-word answer is overwhelmingly input cost — optimise retrieval, chunking and caching. A code generator that takes a 500-token instruction and emits 4,000 tokens is overwhelmingly output cost — optimise output length constraints and model choice. These two systems need opposite optimisations, and knowing which one you have is the first step.
Use our AI Token Counter & Cost Estimator to model this with your own text before you commit to an architecture.
Six ways to reduce token usage
1. Cache repeated prefixes
If your system prompt and tool schemas are identical across requests — and they almost always are — prompt caching turns repeated input tokens into a deeply discounted read. Cache-read discounts are frequently an order of magnitude off the base rate. This is usually the single largest saving available to a production system and often requires only a one-line change.
2. Stop resending the whole conversation
Naive chat implementations resend the entire history every turn. Turn 30 resends turns 1 through 29, so cost grows quadratically with conversation length while the marginal value of the oldest turns approaches zero. Summarise or truncate older turns past a threshold.
3. Prune retrieved context
In RAG, more retrieved chunks is not monotonically better. Beyond a point, additional context dilutes attention and adds cost for no accuracy gain. Rerank and take the top k, tuned empirically on your own evaluation set — often k between 3 and 8 beats 20.
4. Constrain output length
Output tokens cost the most, so an unbounded generation is an unbounded bill. Cap
max_tokens at something just above your realistic need, and ask for concise
output explicitly. "Answer in under 100 words" is a genuine cost control.
5. Choose the model per task
Routing is the highest-leverage architectural decision available. Classification, extraction and routing tasks rarely need a frontier model. Route the easy 80% of traffic to a small model and escalate only what fails — the cost difference between tiers is frequently 20× or more for equivalent quality on simple tasks.
6. Trim your prompt
Prompt text is input tokens paid for on every single request. A 2,000-token system prompt across ten million requests a month is a real line item. Audit it periodically; system prompts accrete instructions over time and are rarely pruned.
Frequently asked questions
There is no fixed ratio. In typical English prose one token averages about 0.75 words, or roughly 4 characters — but that average hides enormous variance. Code, JSON, non-Latin scripts, numbers and emoji all cost substantially more per unit of meaning.
Different providers use different tokenizers with different vocabularies and merge rules, so identical text yields different counts. When comparing vendors, benchmark your real prompts rather than comparing published per-million prices alone.
Yes. Leading whitespace is typically merged into the following word as part of the same token, and punctuation runs have their own tokens. Indented code therefore costs measurably more than the same code without indentation.
The request is rejected with an error, or silently truncated depending on the provider and your configuration. Silent truncation is the dangerous case — you lose part of your input and the model answers confidently about what remains. Count tokens before sending.
Yes. Internal reasoning tokens are billed as output tokens even though they do not appear in the final answer. On reasoning models, a request returning 300 visible tokens may bill for several thousand.
Related guides
LLM API Cost Compared
Published per-million-token prices are only half the equation. Tokenizer differences, caching, batching and reasoning tokens change the real number.
JWT Explained
A JWT is three Base64 segments and a signature. Understanding what the signature does — and does not — prevent is the whole game.
Prompt Engineering for Developers
Prompt engineering is mostly specification writing. Here is a structure that produces reliable output, and the failure patterns that break it.