LLM Tokenization

Large language models don't read characters or words; they read tokens. A tokenizer splits text into chunks from a fixed vocabulary, usually tens of thousands to a couple hundred thousand entries, and maps each chunk to an integer ID. The model only ever sees and predicts those IDs. Common words are often a single token, rarer words are split into pieces, and every character of every language can still be represented by falling back to bytes.

Tokenization sounds like plumbing, but it shapes everything practical about working with LLMs: what you pay (APIs bill per token), how much fits in a context window, how fast responses stream, and a long list of odd behaviors, from miscounting letters in a word to worse performance in some languages.

TL;DR

Quick Example

Inspecting tokens with an open tokenizer (tiktoken, used by OpenAI models):

And counting tokens for a Claude request before sending it:

Notice that leading spaces belong to tokens (' isn', ' but') and that "Tokenization" becomes two pieces.

Core Concepts

Why Subwords?

Byte-Pair Encoding (BPE)

BPE builds a vocabulary by starting from single bytes and repeatedly merging the most frequent adjacent pair in a training corpus into a new token, until the vocabulary reaches its target size. Encoding new text applies the learned merges. Byte-level BPE starts from the 256 possible bytes, so any UTF-8 text, emoji, or binary-looking string is representable, and nothing is ever "unknown".

WordPiece and SentencePiece / Unigram

Vocabulary Size

Vocabularies have grown from ~32k (early Llama) to ~100k–260k in current models. Larger vocabularies encode text in fewer tokens, especially multilingual text and code, which speeds generation and stretches the context window, at the cost of a larger embedding table.

Special Tokens and Chat Templates

Tokenizers reserve special tokens for structure: beginning and end of sequence, message and role boundaries, tool calls, and padding. Chat models are trained on conversations formatted with a specific chat template. Hosted APIs apply it for you. With open models, use the tokenizer's apply_chat_template rather than hand-formatting prompts, since a wrong template noticeably degrades output.

Tokens in Practice

Cost and Limits

API pricing is per million input and output tokens (output is typically several times more expensive), and context windows and max-output limits are measured in tokens. Rough English heuristics:

Code, JSON, non-Latin scripts, and heavy whitespace use more tokens per character. Images, audio, and PDFs are also converted into tokens by multimodal models.

Counting Tokens

Quirks Tokenization Explains

Best Practices

Budget in Tokens, Not Characters

Set limits (maximum document size, chat history length, retrieved chunks for RAG) in tokens measured with the target model's tokenizer. Character-based truncation can cut too much or overflow the context.

Chunk Documents on Semantic Boundaries

When splitting for embeddings or retrieval, target a token size (for example 300–800 tokens) but split on headings, paragraphs, or sentences, not mid-word, and add small overlaps. See embeddings.

Keep Prompts Lean

Verbose system prompts, repeated instructions, and pretty-printed JSON all cost tokens on every request. Compact formats (minified JSON, concise instructions) and prompt caching cut costs without changing behavior.

Use Structured Outputs Instead of Token Tricks

Rather than relying on logit biases or stop sequences keyed to specific tokens, use structured outputs or tool calling for reliable formats. They're robust to tokenizer differences between models.

Common Mistakes

Estimating With the Wrong Tokenizer

Truncating Strings by Characters

Cutting user input at 10,000 characters might be 2,500 tokens of English or 6,000 tokens of code, and can split a multi-byte character. Truncate by tokens using the tokenizer.

Forgetting Output Tokens

Context limits apply to input plus output. A prompt that fills the window leaves no room for the answer, and a low max_tokens truncates responses mid-sentence (check the stop reason).

FAQ

How many tokens is a word?

In English, about 1.3 tokens per word on average, or roughly 4 characters per token. Common words are usually one token, long or rare words several. Other languages and code vary widely, so measure with the model's tokenizer when it matters.

Why do different models count the same text differently?

Each model family trains its own tokenizer with its own vocabulary and merges. A larger or more multilingual vocabulary encodes the same text in fewer tokens. Even versions within one family can change tokenizers.

Do images and audio use tokens?

Yes. Multimodal models convert images, audio, and documents into token sequences: images typically cost hundreds to a few thousand tokens depending on resolution, and PDFs cost text tokens plus image tokens per page. Providers document their formulas and include them in usage reporting.

Will tokenization go away?

Research on byte-level and "tokenizer-free" models continues, and some architectures learn dynamic chunking of bytes. For now, production LLMs rely on subword tokenizers because shorter sequences make training and inference much cheaper.

Related Topics

References