LLM Sampling & Decoding

At every step, a large language model outputs a probability distribution over its entire vocabulary: how likely each possible next token is. Decoding is the procedure that picks one. Always taking the most likely token (greedy decoding) gives predictable but often repetitive text. Sampling picks randomly according to the probabilities, shaped by parameters like temperature, top-p, and top-k, trading predictability for variety.

These settings are among the few knobs you control directly in an API call. They matter less than good prompts and context, but the wrong values cause real problems: creative writing that reads like a form letter, or data extraction that invents a different answer on every run.

TL;DR

Quick Example

The same API with different settings for different jobs:

Some models, especially reasoning models with extended thinking enabled, restrict or ignore sampling parameters. Check the model's documentation.

Core Concepts

Logits, Softmax, and the Distribution

The final layer of the transformer produces one logit per vocabulary token. Softmax converts them into probabilities:

where T is the temperature. Decoding then selects a token, appends it, and repeats until a stop condition.

Greedy and Beam Search

Temperature

Temperature doesn't make a model smarter or more factual. A low temperature makes it consistent, including consistently wrong if the most likely answer is wrong.

Truncation Samplers

These remove unlikely tokens before sampling, cutting off the long tail of nonsense:

Penalties and Biases

Stop Conditions

Generation ends when the model emits its end-of-turn token, a stop sequence you specified appears, or max tokens is reached. Always inspect the stop reason: max_tokens means the output was cut off, which is a common source of truncated JSON and half-finished answers.

Determinism and Reproducibility

Even at temperature 0, hosted models can return slightly different outputs across calls. Floating-point non-associativity in batched GPU computation, different hardware, and backend updates all introduce tiny differences that can flip a close decision between tokens. Some APIs offer a seed parameter for best-effort reproducibility. Design systems to tolerate variation:

Constrained Decoding

Structured outputs constrain decoding so only tokens consistent with a JSON schema or grammar can be chosen. The output is then guaranteed to parse and match required fields, regardless of sampling settings. It's more reliable than prompting for JSON and hoping. See structured outputs. Open-model runtimes (llama.cpp grammars, vLLM guided decoding, Outlines) provide the same capability locally.

Best Practices

Start From Defaults, Change One Thing

Provider defaults are tuned for general use. Lower the temperature for deterministic tasks, raise it modestly for creative ones, and evaluate the effect. Don't tune temperature and top-p together unless you're measuring carefully; their effects compound.

Set max_tokens Deliberately

Too low truncates answers; too high allows runaway outputs and cost. Size it to the task, and handle max_tokens stop reasons explicitly: continue, retry with a higher limit, or fail gracefully.

Generate Multiple Samples When Quality Matters

For hard problems, sampling several candidates at moderate temperature and choosing the best by self-consistency voting, a verifier, tests (for code), or an LLM judge often beats a single greedy answer.

Keep Sampling Out of Your Correctness Story

If correctness depends on getting lucky with sampling, fix the prompt, context, or task design instead. Sampling controls style and diversity; accuracy comes from the model, the context, and verification.

Common Mistakes

Using High Temperature for Extraction

Treating Temperature 0 as a Guarantee

Tests asserting exact string equality on LLM output at temperature 0 will eventually flake. Assert on parsed structure, key facts, or evaluator scores instead.

Ignoring Truncation

Parsing a response without checking stop_reason leads to JSONDecodeErrors and silently incomplete answers when outputs hit max_tokens.

FAQ

What temperature should I use?

For extraction, classification, coding, and factual Q&A, 0–0.3. For general assistant chat, the provider default (often around 1.0, with top-p truncation) or somewhat lower. For creative writing and brainstorming, about 0.8–1.0. Evaluate on your own tasks; the best value is task-specific.

What's the difference between top-p and top-k?

Top-k always keeps a fixed number of candidate tokens. Top-p keeps however many tokens are needed to reach a cumulative probability, so it adapts: few candidates when the model is confident, more when it's unsure. Top-p is generally the better default of the two.

Does lower temperature reduce hallucinations?

Somewhat for obscure details, since it avoids low-probability guesses, but not fundamentally. If the model's most likely answer is wrong, low temperature returns that wrong answer consistently. Grounding with retrieval, citations, and verification address hallucinations far more effectively.

Why can't I set temperature on some models?

Reasoning models and extended-thinking modes manage their own decoding for the reasoning process, and some providers fix or restrict sampling parameters for them. Steer those models through instructions, reasoning effort settings, and output format rather than temperature.

Related Topics

References