LLM Hallucinations

A hallucination is output from a large language model that is fluent and confident but false or unsupported: a citation to a paper that doesn't exist, an API method the library never had, a wrong date, a summary that attributes claims to a document that never made them. Hallucinations are the most common reason LLM features fail in production, and they're especially dangerous because they look exactly like correct answers.

They can't be eliminated entirely, since generating plausible text is what these models do, but they can be made rare and low-impact. Good systems ground the model in real sources, let it say "I don't know", verify important claims, and design interfaces so users can check what matters.

TL;DR

Quick Example

A grounded Q&A prompt that requires quotes and allows abstention:

Core Concepts

Why Models Hallucinate

Types of Hallucination

Fabricated package names in AI-generated code are a real security risk: attackers register those names ("slopsquatting"). Always verify that dependencies exist and are the ones you intended.

Mitigation Strategies

1. Ground the Model in Sources

Provide the facts instead of hoping the model knows them: retrieve relevant documents (RAG, agentic RAG), give it tools (search, database queries, calculators, code execution), and pass authoritative data directly. Grounding converts "recall from memory" into "read and report", which models do far more reliably.

2. Require Citations and Quotes

Ask for answers tied to specific sources: document IDs, quotes, or line references. Citations make answers checkable by users and machines, and asking for quotes first improves faithfulness.

3. Allow and Encourage "I Don't Know"

Tell the model explicitly that abstaining is acceptable and preferred to guessing. Define what to do when information is missing or ambiguous. Without this permission, models tend to answer anyway.

4. Constrain the Task

Narrow, well-specified tasks hallucinate less than open-ended ones. Use structured outputs with enums for fields that must come from a fixed set, and extract before you generate (pull facts into a structure, then write from that structure).

5. Verify

6. Use Stronger Models and Reasoning Where It Matters

Larger and newer models hallucinate less on many benchmarks, and reasoning models catch more of their own errors. Route high-stakes or complex queries to stronger configurations.

Measuring Hallucinations

Best Practices

Design the UX for Verification

Show sources inline, link citations to the exact passage, distinguish generated summaries from quoted text, and make it easy to report errors. Users catch mistakes when the interface helps them.

Match Autonomy to Stakes

Let the model act freely where errors are cheap and reversible (drafts, suggestions); require confirmation or human approval where they aren't (sending emails, changing records, legal and medical advice). Guardrails enforce these boundaries.

Keep Knowledge Fresh Outside the Model

Don't fine-tune to teach facts that change. Retrieval over a maintained knowledge base keeps answers current and makes corrections instant. Fine-tuning is better for style, format, and behavior.

Common Mistakes

Trusting Confident Tone

Fluency and certainty in the output carry no information about correctness. Evaluate claims, not tone.

Asking for Sources Without Providing Them

Assuming RAG Solves It

Retrieval reduces hallucination only if the right documents are retrieved and the model is instructed to stay within them. Poor chunking, irrelevant results, or permission to "use general knowledge" bring hallucinations right back. Evaluate retrieval quality separately from generation.

FAQ

Can hallucinations be eliminated completely?

Not with current LLMs; producing plausible text is fundamental to how they work. They can be reduced dramatically through grounding, citations, abstention, verification, and good task design, and their impact can be contained with review and UX safeguards.

Does lowering the temperature stop hallucinations?

Only marginally. Low temperature avoids unlikely tokens, but if the model's most likely answer is wrong, it'll give that wrong answer consistently. See LLM sampling. Grounding and verification matter far more.

Why do models invent citations and URLs?

Citations have a very regular format that's easy to imitate: authors, year, title, venue. Without access to real sources, the model generates something with the right shape. Provide sources, use search tools, and verify every link and reference programmatically.

How do I detect hallucinations automatically?

Combine checks: verify quoted text against sources, validate citations and links, compare extracted facts with authoritative data, run code and tests, and use a judge model that assesses whether each claim is supported by the provided context. No single method catches everything, so layer them according to stakes.

Related Topics

References