LLM Token Counter: See exactly how GPT models split your text.
Runs in your browser — nothing you paste leaves this page. How we prove that
LLM Token Counter playground
Other vendors (Claude, Gemini, Llama) use different tokenizers — counts here are approximate for them; use the vendor's count endpoint for billing.
Results update as you type — press Enter to run now.
Tokens
o200k_base — GPT-5, GPT-4.1, GPT-4o, o1 / o3 / o4-mini
Paste a prompt, a document or a JSON payload, pick the o200k_base or cl100k_base encoding, and get the exact token count with every token boundary highlighted — plus characters, words, bytes and the cost at your own price per million tokens, all computed in your browser, no signup required.
The Gap
Models read tokens. You write characters.
A language model never sees your text as letters. A tokenizer first cuts it into tokens — whole common words, pieces of rarer ones, punctuation, runs of spaces — and the model works on their IDs. In English a token averages about four characters, but that is only an average: code, numbers, URLs and non-Latin scripts break into many more tokens per character.
The count matters three times over. The context window is a token budget, so it decides whether a document fits at all. Billing is per token, usually quoted per million with input and output priced separately. Rate limits are often set in tokens per minute, so a long prompt can throttle you before the request count does. Estimating from characters is how a prompt that “should fit” gets truncated.
As of October 2026, OpenAI’s current models — GPT-5, GPT-4.1, GPT-4o and the o-series — use o200k_base, a vocabulary of about 200,000 tokens. GPT-4, GPT-3.5 Turbo and the text-embedding-3 models use the older cl100k_base, about 100,000. The larger vocabulary holds more whole words outside English, which is why the same Hindi sentence in the reference below costs 26 tokens on one and 66 on the other.
Claude, Gemini and Llama use their own tokenizers, so these counts are an approximation for them — use the vendor’s token-count endpoint when you need the exact number.
The Pipeline
How it works.
Four deterministic steps run on every change — all inside your browser tab, with the same byte-pair encoding tables OpenAI publishes for tiktoken.
-
Load the vocabulary.
The chosen encoding’s merge table is fetched once from this site and cached by the browser; switching to the other encoding loads its own table the same way.
-
Split, then merge.
A regular expression splits the text into words, numbers, punctuation and whitespace runs; byte-pair encoding then merges the UTF-8 bytes of each piece by rank into token IDs.
-
Map tokens back to text.
The IDs are decoded in order and grouped so a character that needs several tokens — an emoji, a CJK glyph — is shown as one span instead of broken bytes.
-
Count and price.
Tokens, characters, words, UTF-8 bytes and characters per token, plus tokens ÷ 1,000,000 × the price you enter. Nothing leaves the tab.
Token Reference
Two encodings, one price formula.
Real counts from the same tokenizer the panel runs, and the arithmetic that turns a count into a bill — so you can check both by hand.
The same sentence, counted twice
In plain English the two encodings agree. The further a script is from English, the more the larger o200k_base vocabulary saves.
Same sentence, two encodings (gpt-tokenizer 4.0.0)
English "Kubernetes schedules pods onto nodes that have enough free CPU and memory."
74 chars · 74 bytes o200k_base 14 cl100k_base 14
German "Kubernetes plant Pods auf Knoten mit genug freier CPU und genügend Speicher ein."
80 chars · 81 bytes o200k_base 17 cl100k_base 21
Japanese "Kubernetes は十分な CPU とメモリがあるノードに Pod を配置します。"
43 chars · 87 bytes o200k_base 20 cl100k_base 26
Hindi "Kubernetes पॉड को उन नोड्स पर शेड्यूल करता है जिनमें पर्याप्त CPU और मेमोरी हो।"
79 chars · 183 bytes o200k_base 26 cl100k_base 66 From tokens to cost
Divide by a million, multiply by your rate, and price input and output separately. The rates below are round illustrative numbers, not any vendor’s price list.
cost = tokens ÷ 1,000,000 × price_per_1M_tokens
prompt (input) 1,840 tokens
reply (output) 420 tokens
rates $1.00 / 1M input · $4.00 / 1M output ← illustrative; use your own
input 1,840 ÷ 1,000,000 × 1.00 = $0.00184
output 420 ÷ 1,000,000 × 4.00 = $0.00168
per call = $0.00352
× 50,000 calls a day = $176.00 a day Next Step
Running the model yourself? Size the GPU next.
A long context costs memory as well as money. The LLM VRAM Calculator sizes the weights, the KV cache for your context length and the runtime overhead of a local model, and names the smallest card that fits.
"Hello, world!" 13 chars → 4 tokens
Hello | , | world | !
o200k_base = 4
cl100k_base = 4 FAQ
Questions, answered.
Tap a question to expand the answer.
What is a token?
A token is the unit a language model reads and is billed by: a common word, a piece of a longer word, a punctuation mark or a run of spaces. The tokenizer splits text into tokens from a fixed vocabulary learned during training, so frequent English words are usually one token each while rare words break into several. As a rough guide, English prose averages around four characters per token, which is why the counter shows characters per token next to the total.
What is the difference between o200k_base and cl100k_base, and which models use them?
They are the two OpenAI byte-pair encodings, named after their vocabulary sizes of roughly 200,000 and 100,000 tokens. As of October 2026, o200k_base is used by GPT-5, GPT-4.1, GPT-4o and the o-series reasoning models, and cl100k_base by GPT-4, GPT-3.5 Turbo and the text-embedding-3 models. The larger vocabulary usually needs fewer tokens for the same text, especially outside English, so pick the encoding that matches the model you are sending to.
Why does the count differ for Claude, Gemini or Llama?
Each model family ships its own tokenizer with its own vocabulary, so the same text splits into a different number of tokens. This counter implements only the OpenAI encodings, which makes it an approximation for any other model — useful for sizing a prompt, not for billing. For an exact figure, use the vendor's own token-counting endpoint or the usage numbers its API returns.
Why do emoji and Chinese, Japanese or Korean text cost more tokens?
Byte-pair encodings work on UTF-8 bytes, and the vocabulary is dominated by text that was common in training data. An emoji or a CJK character takes three or four bytes, and when that byte sequence is not a single vocabulary entry it is split into several tokens. That is why a short line of emoji can cost more than a full English sentence, and why o200k_base, with its larger multilingual vocabulary, often counts CJK text lower than cl100k_base.
How does the "$ per 1M tokens" field work?
Type the price your provider charges per million tokens and the counter multiplies it by the token count to show an estimated cost. No prices are built in, because they change often and differ by model, tier and region. Input and output tokens are usually priced differently, so enter the input price to cost a prompt and the output price to cost a reply of the same length.
Does code tokenize differently from prose?
Yes. Code is full of punctuation, operators, brackets and indentation, and each of those tends to split into its own token or a short run, so code usually yields fewer characters per token than English prose. Runs of spaces are often merged into a single token, which helps indented code, but unusual identifiers and long string literals still break into several pieces. Paste a real file to see the split rather than estimating from its length.
Why does this count differ from the usage the chat API reports?
The counter tokenizes exactly the text you paste. A chat completion request wraps every message in role and separator tokens, which adds a few tokens per message plus a few for the reply, and tools or a system prompt add their own. Expect the API's prompt usage to be slightly higher than this count for a single message, and higher still for a long conversation.
Does my text ever leave my browser?
No. The token counter runs 100% client-side: the tokenizer is loaded from this site into the page and your text is tokenized in your browser tab. Nothing is uploaded to a server, and there is no account or signup, so you can safely paste internal prompts or documents.
More free, private DevOps tools.
The LLM Token Counter is one tool in OpsCanopy — a growing canopy of browser-based validators, converters and testers that never touch a server.
Related tools
New to AI & local LLMs? Read the AI & local LLMs guide →
42 free tools, every one offline-capable — opscanopy.com works with no signup and nothing uploaded.
More for AI workloads: the LLM VRAM Calculator and the Data Size Converter, or browse the full tools directory.
Counts are exact for OpenAI models that use these encodings; chat formatting adds a few tokens per message, and other vendors’ tokenizers differ, so confirm against your provider’s usage report before you budget. OpsCanopy is free and open.