LLM Hallucinations
A hallucination is output from a large language model that is fluent and confident but false or unsupported: a citation to a paper that doesn't exist, an API method the library never had, a wrong date, a summary that attributes claims to a document that never made them. Hallucinations are the most common reason LLM features fail in production, and they're especially dangerous because they look exactly like correct answers.
They can't be eliminated entirely, since generating plausible text is what these models do, but they can be made rare and low-impact. Good systems ground the model in real sources, let it say "I don't know", verify important claims, and design interfaces so users can check what matters.
TL;DR
- LLMs generate plausible continuations, not verified facts. Knowledge lives implicitly in weights, and gaps get filled with likely-sounding text.
- Two broad types: factuality errors (wrong about the world) and faithfulness errors (not supported by the provided context).
- The strongest mitigation is grounding: give the model the right sources (RAG, tools, documents) and require citations to them.
- Explicitly permit uncertainty ("say you don't know if the sources don't answer it"), which cuts fabrication substantially.
- Verify high-stakes outputs with checks: quote matching, tests for code, schema validation, second-pass review, or humans.
- Measure hallucination rates with evaluation sets; don't rely on spot checks.
Quick Example
A grounded Q&A prompt that requires quotes and allows abstention:
Core Concepts
Why Models Hallucinate
- Next-token prediction: training rewards producing likely text. When the model lacks a fact, the most likely continuation is still a fluent, specific-sounding answer.
- Compressed, imperfect knowledge: facts are stored diffusely in weights (see transformer architecture). Rare facts, long-tail entities, exact numbers, and recent events are stored weakly or not at all.
- Training and evaluation incentives: benchmarks and feedback that reward confident answers over "I don't know" teach models to guess.
- Knowledge cutoff: models don't know about events after their training data ends unless you provide them.
- Context problems: missing, contradictory, or overly long context, or instructions the model loses track of (see context windows).
- Pressure from the prompt: leading questions ("Which paper proved X?" when none did) invite fabrication.
Types of Hallucination
Fabricated package names in AI-generated code are a real security risk: attackers register those names ("slopsquatting"). Always verify that dependencies exist and are the ones you intended.
Mitigation Strategies
1. Ground the Model in Sources
Provide the facts instead of hoping the model knows them: retrieve relevant documents (RAG, agentic RAG), give it tools (search, database queries, calculators, code execution), and pass authoritative data directly. Grounding converts "recall from memory" into "read and report", which models do far more reliably.
2. Require Citations and Quotes
Ask for answers tied to specific sources: document IDs, quotes, or line references. Citations make answers checkable by users and machines, and asking for quotes first improves faithfulness.
3. Allow and Encourage "I Don't Know"
Tell the model explicitly that abstaining is acceptable and preferred to guessing. Define what to do when information is missing or ambiguous. Without this permission, models tend to answer anyway.
4. Constrain the Task
Narrow, well-specified tasks hallucinate less than open-ended ones. Use structured outputs with enums for fields that must come from a fixed set, and extract before you generate (pull facts into a structure, then write from that structure).
5. Verify
- Deterministic checks: quotes exist in the source, citations point to real documents, numbers match the database, code compiles and passes tests, URLs resolve.
- Model-based checks: a second pass (or a different model) reviews the answer against the sources and flags unsupported claims. This is an LLM-as-judge setup.
- Self-consistency: sample several answers; disagreement signals low confidence.
- Human review for high-stakes outputs (legal, medical, financial, customer-facing commitments).
6. Use Stronger Models and Reasoning Where It Matters
Larger and newer models hallucinate less on many benchmarks, and reasoning models catch more of their own errors. Route high-stakes or complex queries to stronger configurations.
Measuring Hallucinations
- Build an evaluation set of real questions with known answers and sources, including questions the sources can't answer (to test abstention).
- Score faithfulness (claims supported by context), correctness (matches ground truth), and abstention quality (declines when it should, answers when it can).
- Track these metrics per model, prompt, and retrieval change, in CI and production sampling. See LLM evaluation and LLMOps.
Best Practices
Design the UX for Verification
Show sources inline, link citations to the exact passage, distinguish generated summaries from quoted text, and make it easy to report errors. Users catch mistakes when the interface helps them.
Match Autonomy to Stakes
Let the model act freely where errors are cheap and reversible (drafts, suggestions); require confirmation or human approval where they aren't (sending emails, changing records, legal and medical advice). Guardrails enforce these boundaries.
Keep Knowledge Fresh Outside the Model
Don't fine-tune to teach facts that change. Retrieval over a maintained knowledge base keeps answers current and makes corrections instant. Fine-tuning is better for style, format, and behavior.
Common Mistakes
Trusting Confident Tone
Fluency and certainty in the output carry no information about correctness. Evaluate claims, not tone.
Asking for Sources Without Providing Them
Assuming RAG Solves It
Retrieval reduces hallucination only if the right documents are retrieved and the model is instructed to stay within them. Poor chunking, irrelevant results, or permission to "use general knowledge" bring hallucinations right back. Evaluate retrieval quality separately from generation.
FAQ
Can hallucinations be eliminated completely?
Not with current LLMs; producing plausible text is fundamental to how they work. They can be reduced dramatically through grounding, citations, abstention, verification, and good task design, and their impact can be contained with review and UX safeguards.
Does lowering the temperature stop hallucinations?
Only marginally. Low temperature avoids unlikely tokens, but if the model's most likely answer is wrong, it'll give that wrong answer consistently. See LLM sampling. Grounding and verification matter far more.
Why do models invent citations and URLs?
Citations have a very regular format that's easy to imitate: authors, year, title, venue. Without access to real sources, the model generates something with the right shape. Provide sources, use search tools, and verify every link and reference programmatically.
How do I detect hallucinations automatically?
Combine checks: verify quoted text against sources, validate citations and links, compare extracted facts with authoritative data, run code and tests, and use a judge model that assesses whether each claim is supported by the provided context. No single method catches everything, so layer them according to stakes.
Related Topics
- Large Language Models — How LLMs work end to end
- RAG — Grounding answers in retrieved sources
- LLM Evaluation — Measuring faithfulness and correctness
- AI Guardrails — Limiting harm from bad outputs
- LLM Sampling — What decoding settings can and can't fix
- Context Windows — Context problems that cause errors