LLM Sampling & Decoding
At every step, a large language model outputs a probability distribution over its entire vocabulary: how likely each possible next token is. Decoding is the procedure that picks one. Always taking the most likely token (greedy decoding) gives predictable but often repetitive text. Sampling picks randomly according to the probabilities, shaped by parameters like temperature, top-p, and top-k, trading predictability for variety.
These settings are among the few knobs you control directly in an API call. They matter less than good prompts and context, but the wrong values cause real problems: creative writing that reads like a form letter, or data extraction that invents a different answer on every run.
TL;DR
- The model produces logits (raw scores) for every token; softmax turns them into probabilities.
- Temperature rescales them: low (0–0.3) → focused and repeatable; ~1 → the model's natural distribution; high → more random.
- Top-k samples only from the k most likely tokens; top-p (nucleus) from the smallest set whose probabilities sum to p; min-p drops tokens far less likely than the top one.
- Stop sequences and max tokens end generation; check the stop reason to detect truncation.
- Temperature 0 isn't perfectly deterministic on hosted APIs; design for small variations.
- Rule of thumb: adjust **temperature or top-p, not both**, and leave defaults unless you have a reason.
Quick Example
The same API with different settings for different jobs:
Some models, especially reasoning models with extended thinking enabled, restrict or ignore sampling parameters. Check the model's documentation.
Core Concepts
Logits, Softmax, and the Distribution
The final layer of the transformer produces one logit per vocabulary token. Softmax converts them into probabilities:
where T is the temperature. Decoding then selects a token, appends it, and repeats until a stop condition.
Greedy and Beam Search
- Greedy decoding always picks the argmax token. It's deterministic in principle and fine for short, factual outputs, but it can loop or produce bland text in long generations.
- Beam search keeps several candidate sequences and chooses the best overall. It was common in machine translation but is rarely used for chat LLMs, where it tends to produce generic, repetitive text.
Temperature
Temperature doesn't make a model smarter or more factual. A low temperature makes it consistent, including consistently wrong if the most likely answer is wrong.
Truncation Samplers
These remove unlikely tokens before sampling, cutting off the long tail of nonsense:
- Top-k: keep only the k highest-probability tokens (for example k = 40).
- Top-p (nucleus): keep the smallest set of tokens whose cumulative probability ≥ p (for example 0.9). It adapts to the distribution: when the model is confident, few tokens qualify; when it's uncertain, more do.
- Min-p: keep tokens whose probability is at least
min_p × p(top token). It's popular in open-model runtimes for staying coherent at higher temperatures.
Penalties and Biases
- Frequency and presence penalties (in some APIs) reduce the probability of tokens already used, discouraging repetition.
- Repetition penalty (in open-model runtimes) is a similar idea.
- Logit bias raises or lowers specific token IDs. It's brittle because it depends on the tokenizer; prefer structured outputs for format control.
Stop Conditions
Generation ends when the model emits its end-of-turn token, a stop sequence you specified appears, or max tokens is reached. Always inspect the stop reason: max_tokens means the output was cut off, which is a common source of truncated JSON and half-finished answers.
Determinism and Reproducibility
Even at temperature 0, hosted models can return slightly different outputs across calls. Floating-point non-associativity in batched GPU computation, different hardware, and backend updates all introduce tiny differences that can flip a close decision between tokens. Some APIs offer a seed parameter for best-effort reproducibility. Design systems to tolerate variation:
- Validate outputs against schemas, and retry on failure.
- Evaluate on distributions of outputs, not single runs. See LLM evaluation.
- Cache responses where exact repeatability matters.
Constrained Decoding
Structured outputs constrain decoding so only tokens consistent with a JSON schema or grammar can be chosen. The output is then guaranteed to parse and match required fields, regardless of sampling settings. It's more reliable than prompting for JSON and hoping. See structured outputs. Open-model runtimes (llama.cpp grammars, vLLM guided decoding, Outlines) provide the same capability locally.
Best Practices
Start From Defaults, Change One Thing
Provider defaults are tuned for general use. Lower the temperature for deterministic tasks, raise it modestly for creative ones, and evaluate the effect. Don't tune temperature and top-p together unless you're measuring carefully; their effects compound.
Set max_tokens Deliberately
Too low truncates answers; too high allows runaway outputs and cost. Size it to the task, and handle max_tokens stop reasons explicitly: continue, retry with a higher limit, or fail gracefully.
Generate Multiple Samples When Quality Matters
For hard problems, sampling several candidates at moderate temperature and choosing the best by self-consistency voting, a verifier, tests (for code), or an LLM judge often beats a single greedy answer.
Keep Sampling Out of Your Correctness Story
If correctness depends on getting lucky with sampling, fix the prompt, context, or task design instead. Sampling controls style and diversity; accuracy comes from the model, the context, and verification.
Common Mistakes
Using High Temperature for Extraction
Treating Temperature 0 as a Guarantee
Tests asserting exact string equality on LLM output at temperature 0 will eventually flake. Assert on parsed structure, key facts, or evaluator scores instead.
Ignoring Truncation
Parsing a response without checking stop_reason leads to JSONDecodeErrors and silently incomplete answers when outputs hit max_tokens.
FAQ
What temperature should I use?
For extraction, classification, coding, and factual Q&A, 0–0.3. For general assistant chat, the provider default (often around 1.0, with top-p truncation) or somewhat lower. For creative writing and brainstorming, about 0.8–1.0. Evaluate on your own tasks; the best value is task-specific.
What's the difference between top-p and top-k?
Top-k always keeps a fixed number of candidate tokens. Top-p keeps however many tokens are needed to reach a cumulative probability, so it adapts: few candidates when the model is confident, more when it's unsure. Top-p is generally the better default of the two.
Does lower temperature reduce hallucinations?
Somewhat for obscure details, since it avoids low-probability guesses, but not fundamentally. If the model's most likely answer is wrong, low temperature returns that wrong answer consistently. Grounding with retrieval, citations, and verification address hallucinations far more effectively.
Why can't I set temperature on some models?
Reasoning models and extended-thinking modes manage their own decoding for the reasoning process, and some providers fix or restrict sampling parameters for them. Steer those models through instructions, reasoning effort settings, and output format rather than temperature.
Related Topics
- Large Language Models — How LLMs work end to end
- Transformer Architecture — Where the logits come from
- Tokenization — What exactly is being sampled
- Structured Outputs — Constrained decoding for reliable formats
- Prompt Engineering — The bigger lever for output quality
- LLM Hallucinations — Why sampling alone doesn't fix accuracy