AI Token Counter & Cost Estimator
Count tokens and estimate API cost before you ship a prompt.
Why token count — not word count — determines what you pay
Every commercial LLM API bills by the token, and a token is not a word. It is a
sub-word unit produced by a tokenizer: a byte-pair-encoding (BPE) vocabulary of a few
hundred thousand entries that the model was trained against. The word
unbelievable is one word but typically three tokens
(un, believ, able). A single emoji can be two
tokens. A two-character CJK phrase can be two or three. That is why the popular
"divide characters by four" rule of thumb drifts by 20–30% on real text, and why an
estimate that is good enough for a budget meeting is not good enough for a monthly
invoice you have to defend.
This counter runs a BPE-style pre-tokenisation pass — the same word / number / punctuation splitting rules that GPT-family tokenizers apply — then estimates sub-word splits per chunk. On typical English prose it lands within about 10% of the exact count. Treat it as a planning number, not an invoice number.
How to get an exact count
There is only one authoritative source for the number you will be billed on, and it is always the provider:
- OpenAI — use the
tiktokenlibrary with the encoding matching your model, or readresponse.usage.prompt_tokensandcompletion_tokensfrom any API response. - Anthropic — call the
/v1/messages/count_tokensendpoint before sending, or read theusageobject on the response. - Google Gemini — use
countTokensfrom the SDK, or theusageMetadatafield on responses.
import tiktoken
enc = tiktoken.get_encoding("o200k_base") # GPT-4o family
n = len(enc.encode("Your prompt goes here"))
print(f"{n} tokens")
The usage object on a real response is always preferable, because it
also captures tokens you did not write yourself: tool schemas, system-injected
preambles, and chat history that your framework appends silently.
The cost formula
Providers publish separate input and output rates per million tokens, and output tokens are consistently priced higher — usually 3× to 5× — because generation costs more compute than ingestion. Total cost is:
cost = (input_tokens / 1_000_000) * input_price
+ (output_tokens / 1_000_000) * output_price
This matters more than most teams expect. A RAG application that stuffs 40,000 tokens of retrieved context into every request but returns a 200-token answer is almost entirely input cost. A code-generation endpoint that receives a 500-token prompt and emits 4,000 tokens of code is almost entirely output cost. The two workloads have opposite optimisation strategies, and knowing which one you have tells you where to spend engineering effort.
Four things that quietly inflate your bill
1. Tokenizer differences between providers
The same text produces different token counts on different providers. Anthropic's newer models use a tokenizer that can produce meaningfully more tokens for identical input than the GPT-family encodings. If you are comparing two vendors on published per-million prices alone, you can be misled: a 20% cheaper headline rate can be erased by a tokenizer that emits 25% more tokens. Always benchmark with your own real prompts.
2. Chat history re-sent on every turn
Most chat applications resend the entire conversation on each request. A 30-turn conversation does not cost 30× one turn — it costs roughly the sum of an arithmetic series, because turn 30 resends turns 1 through 29. Truncation, summarisation of older turns, or provider-side prompt caching are the standard fixes.
3. Prompt caching
Every major provider now discounts repeated prefixes. If your system prompt and tool schemas are identical across requests and you are not using caching, you are paying several times over for the same tokens. Cache-read discounts are steep — frequently an order of magnitude off the base input rate.
4. Reasoning models
Reasoning models emit internal reasoning tokens before answering. Those are billed as output tokens and are invisible in the final answer. A request that returns 300 visible tokens may bill for several thousand.
Privacy note
Everything on this page runs in JavaScript in your browser tab. Your text is never sent to a server, never logged, and never used to train anything. You can verify this by opening your browser's network inspector before typing — or by loading this page and then disconnecting from the internet entirely. The tool keeps working.
Frequently asked questions
On ordinary English prose the estimate is usually within about 10% of the exact count. It is deliberately conservative on code, JSON and non-Latin scripts, which all tokenise less efficiently than prose. For a number you will put in front of finance, use tiktoken or the provider's count_tokens endpoint.
Different providers use different tokenizers and vocabularies. The same string genuinely produces different counts. If you are comparing vendors, compare on your own prompts, not on published per-million prices alone.
No. The counter is pure client-side JavaScript. Nothing you type leaves your browser. There is no backend call, no analytics on input content, and no server-side logging.
Yes, on essentially every provider — typically 3–5× more, because generation is compute-bound while ingestion is parallelisable. Workloads that produce long outputs should optimise for output length first.
It is the maximum number of tokens — input plus output — a model can consider in one request. Exceeding it produces an error or silent truncation. Counting tokens before you send is how you avoid discovering the limit at runtime.