RAG Evaluation

A RAG system has many moving parts (parsing, chunking, embeddings, hybrid search, reranking, prompting, and the model), and a change to any of them can quietly improve some answers while breaking others. Evaluation makes those trade-offs visible. Without it, teams tune by anecdote ("this answer looks better now") and ship regressions.

Effective RAG evaluation separates retrieval quality (did we find the right passages?) from generation quality (did the model answer correctly and faithfully from them?), uses a golden dataset of realistic questions, combines deterministic metrics with calibrated LLM-as-judge scoring, and keeps running in CI and production.

TL;DR

Quick Example

A minimal evaluation loop: retrieval recall plus a judge for faithfulness and correctness.

Core Concepts

The Golden Dataset

A good evaluation set is the most valuable asset in a RAG project:

Synthetic question generation (an LLM writing questions from chunks) helps bootstrap coverage, but it skews toward easy, lexically similar questions. Mix it with real data.

Retrieval Metrics

Generation Metrics

LLM-as-Judge

Model graders scale evaluation to thousands of examples. Make them reliable:

See LLM evaluation.

Tooling

Frameworks and platforms that implement these metrics and workflows include Ragas, DeepEval, TruLens, Promptfoo, Arize Phoenix, LangSmith, Braintrust, and cloud evaluation services. They add dataset management, experiment comparison, and tracing. The metrics matter more than the tool; start with simple scripts if that gets evaluation running today.

Evaluation in the Development Loop

  1. Baseline: measure the current pipeline on the golden set.
  2. Change one thing: chunk size, embedding model, top-k, reranker, prompt, or model.
  3. Compare: retrieval and generation metrics side by side, plus cost and latency.
  4. Inspect failures: read the examples that got worse, not just the averages.
  5. Gate in CI: fail builds when key metrics regress beyond a threshold.
  6. Monitor production: sample real traffic for judge scoring, track user feedback (thumbs, corrections, escalations), and add new failures to the golden set.

See LLMOps for the operational side.

Best Practices

Diagnose by Stage

Low recall means fix retrieval (chunking, hybrid search, query rewriting, metadata filters). High recall but unfaithful answers means fix prompting (grounding instructions, citations, allowing abstention) or the model. Stage-level metrics tell you where to look.

Include Unanswerable Questions

A system that always answers looks great until users ask something outside the corpus. Measure how often it correctly abstains, and how often it wrongly refuses.

Track Cost and Latency Alongside Quality

A change that adds 2% correctness but doubles latency or token cost may not be worth it. Report tokens per answer, p95 latency, and quality together.

Version Everything

Record dataset version, pipeline configuration, prompts, model versions, and judge version for every run, so results are reproducible and comparable over time.

Common Mistakes

Judging Only the Final Answer

End-to-end scores hide why quality changed. Always log the retrieved chunks and score retrieval independently.

Evaluating on Synthetic Questions Alone

LLM-generated questions often reuse the source passage's wording, making retrieval look far better than it is for real users' phrasing. Include real queries.

Trusting an Unvalidated Judge

A judge that disagrees with humans 30% of the time produces confident but misleading metrics. Spot-check judge verdicts regularly, and recalibrate prompts when you change judge models.

FAQ

What are the most important RAG metrics?

Retrieval recall@k (is the needed information in the context?) and faithfulness (is the answer supported by that context?), plus correctness against reference answers. Add answer relevance, citation accuracy, and abstention quality as your system matures.

How many examples do I need in an evaluation set?

Start with 50–100 high-quality, diverse examples. That's enough to catch major regressions. Grow toward several hundred as you discover failure modes, and keep subsets per category (product area, question type) to spot localized regressions.

Is LLM-as-judge reliable?

Reasonably, when designed well: specific rubrics, reference answers and context provided, binary criteria, reasoning before verdicts, and validation against human labels. It's not a replacement for periodic human review, especially in high-stakes domains.

How do I evaluate RAG in production without labels?

Use reference-free metrics (faithfulness to retrieved context, answer relevance, context relevance) judged by an LLM on sampled traffic, plus behavioral signals: user feedback, follow-up rephrasing, escalations to humans, and click-through on citations. Feed discovered failures back into the golden set.

Related Topics

References