New batches starting this week Β· Limited seats

RAG Evaluation Metrics Explained: Faithfulness, Relevance, Context Precision and Recall

A metric-by-metric guide to RAG evaluation: what each retrieval and generation metric measures, the inputs it needs, what a low score means and which component to fix.

RAG evaluation metrics: context precision, context recall, faithfulness, answer relevance and citation accuracy
Last updated Β· 14 min read Β· 3,189 words

RAG evaluation metrics tell you which part of a retrieval-augmented generation system is failing: retrieval (the wrong passages came back), generation (the right passages, but a wrong or unsupported answer), or the behaviour around them (bad citations, answering when it should decline). The core set is context precision and context recall for retrieval, faithfulness, answer relevance and answer correctness for generation, plus citation accuracy and refusal correctness, each computed by an LLM judge, against a reference, or deterministically in code. This guide goes metric by metric: what each measures, what it needs, what a low score means and what to fix.

For the broader process (test sets, judge biases, CI gates, monitoring), see our LLM evaluation guide; for the basics, what RAG is and how it works.

The four inputs and three ways to compute a metric

Almost every RAG metric uses some of four things you should log per test case: the question; the retrieved contexts (chunks in rank order, with IDs, as actually placed in the prompt); the answer, including citations; and the reference, meaning an expert-approved answer, a list of relevant chunk or document IDs, or both.

  • Deterministic: plain code over IDs and strings. Hit rate, recall@k, MRR, nDCG (given relevance labels) and "every cited ID was retrieved" need no judge.
  • Reference-based: compares output with expert ground truth. Context recall and answer correctness cannot be computed without a reference.
  • LLM-judge: a model makes a judgement, usually claim by claim. Faithfulness and answer relevance are usually reference-free, so they also run on production samples.
 question -> retriever -> contexts -> generator -> answer
              |           |            |          |
       hit@k, MRR   ctx precision faithfulness relevance
       nDCG         ctx recall    citations    correctness
                                               refusal

Retrieval metrics

Evaluate retrieval first: if the evidence never reached the prompt, nothing downstream can fix it.

Context precision

Measures whether retrieved chunks are relevant and whether relevant ones are ranked above irrelevant ones. Computed by marking each retrieved chunk relevant or not (by a judge against the reference, by human labels or by known relevant IDs), then averaging precision at each rank where a relevant chunk appears, much like average precision in search. Reference-free variants judge against the generated answer instead and inherit its errors. Inputs: question, ordered contexts, ideally a reference.

Low score: noise in the context window. Fix first-stage retrieval (hybrid search, metadata filters on document status and region), add or tune a reranker, or send fewer chunks. Pitfall: precision says nothing about what is missing; five relevant chunks can still omit the exception clause that changes the answer.

Context recall

Measures whether retrieval found everything needed for the correct answer. Computed by splitting the reference answer into statements and having a judge check whether each can be attributed to the retrieved contexts; the score is the share supported. With labelled relevant IDs you can compute an ID-based recall in code instead. Inputs: contexts and a reference; there is no meaningful reference-free version.

Low score: the evidence was not in the prompt. Usual causes: missing documents, chunking that split a rule from its exception, a filter excluding the right version, weak embeddings for your vocabulary, or top-k too small. Fix ingestion and retrieval, not the prompt. Pitfall: a vague reference inflates recall, and recall rewards retrieving more, so watch precision alongside it.

Hit rate and recall@k

Hit rate@k is the share of questions with at least one relevant chunk in the top k. Recall@k is the share of a question's relevant chunks that appear in the top k, averaged over questions. Both are deterministic, so they suit fast iteration on chunking, embeddings and search.

Low score: the right material is not surfacing within the window you send. Compare k values: healthy recall@20 but poor recall@5 means a reranker will help; poor recall@20 means the problem is earlier. Pitfalls: chunk IDs change when you re-chunk, so label relevance at document-and-section level or re-map IDs. Hit rate is generous: one of three needed chunks counts as a hit.

MRR (mean reciprocal rank)

For each question take 1 divided by the rank of the first relevant chunk (1 at rank 1, one half at rank 2, zero if none found), then average. Deterministic. Low score: relevant material is found but ranked low; improve reranking or hybrid fusion. Pitfall: MRR ignores everything after the first hit, so it can look fine while a second required passage is missing.

nDCG, plainly

Normalised discounted cumulative gain scores the whole ranking and allows graded relevance (say 2 for "directly answers", 1 for "useful background", 0 for irrelevant). Each grade is divided by a rank discount (conventionally log2 of rank plus one) and summed, then divided by the same sum for the ideal ordering of known relevant items, giving 0 to 1. Low score: poor ordering or missing key chunks. Pitfall: graded labels take expert time, and labellers disagree.

Generation metrics

Faithfulness (groundedness)

Measures whether every claim in the answer is supported by the retrieved context: the closest single metric to "is it hallucinating relative to its sources?" Computed by a judge that splits the answer into atomic claims and checks each against the contexts; score is supported claims over total claims. Inputs: answer and contexts, no reference.

Low score: the generator is adding knowledge from training, over-generalising or blending passages. Tighten the grounding instruction ("answer only from the context; say when it is silent"), lower temperature, try a model that follows grounding better, or cut noisy context. Pitfalls: faithful is not correct; an answer can be perfectly faithful to a superseded policy. Claim extraction is itself a judge step, and hedges like "please consult your manager" may count as unsupported claims unless your rubric says otherwise.

Answer relevance

Measures whether the answer addresses the question asked, completely and without padding. It checks focus, not truth. Computed in one common approach by having a judge generate questions the answer would suit and measuring embedding similarity to the original question; others ask a judge directly whether each part of the question is addressed. Inputs: question and answer.

Low score: evasive or generic answers, or a two-part question answered only in part. Fix the prompt, answer format or query decomposition. Pitfalls: a confident wrong answer can score high, and correct refusals often score low, so exclude should-decline cases and measure them with refusal correctness.

Answer correctness

Measures whether the answer matches the expert reference in substance. Computed by a judge comparing claims in answer and reference (present in both, only in the answer, only in the reference) into an F1-style score, sometimes blended with embedding similarity. For short facts such as a date, number or yes/no, normalised match in code beats any judge. Inputs: question, answer, reference.

Low score: read it with the others. With low recall it is a retrieval problem; with high recall and high faithfulness the reference may be stale or the model chose the wrong passage among several; with low faithfulness it is generation drift. Pitfall: references go stale; version them with their sources.

Citation accuracy and refusal correctness

Citation accuracy

Two layers. Deterministic: every cited ID must exist in the retrieved set for that request; anything else is a fabricated citation and fails outright. Judge-based: does the cited chunk support the sentence it is attached to (citation precision), and does each factual sentence that needs a source have one (citation recall)? Inputs: answer with citation markers, contexts with IDs.

Low score: fabricated IDs point to the output format; enforce structured citations and validate in code. Real IDs on the wrong sentence often mean chunks are too large or too similar. Pitfall: checking only that citations exist; a real chunk cited for a claim it does not support is still wrong.

Refusal correctness

Measures whether the system says "I don't know" when, and only when, it should: the answer is not in the knowledge base, the user may not see the source, or the question is out of scope. Computed by labelling every test case answerable or should-decline, detecting refusals deterministically (a structured answered: false field beats phrase matching) or with a small classifier, then reporting two rates: should-decline cases correctly declined, and answerable cases wrongly declined. Inputs: question, answer, label and, for permission cases, the test user's role.

Low score: answering when it should decline points to a weak grounding instruction, a retriever that always returns something loosely related, or missing permission filters. Over-declining points to poor recall or an over-cautious prompt. Pitfall: tuning hard against hallucination can quietly create an over-refusing assistant; track both rates.

To compute these on your own pipeline with pgvector, rerankers, Ragas and LangSmith, Cloudsoft's AI, GenAI and Agentic AI course builds evaluation into every RAG lab.

Worked example with toy numbers

All numbers here are illustrative toy values chosen to show the arithmetic, not results from any real system.

Consider a bank building an internal assistant for branch operations staff on KYC procedures. Test question: "When must a dormant savings account holder re-submit KYC documents before reactivation?" Experts labelled three relevant chunks: R1 (the reactivation rule), R2 (acceptable documents) and R3 (an exception for senior citizens). The retriever returns five chunks: irrelevant, R1, irrelevant, R2, irrelevant. R3 is missing.

MetricToy calculationToy result
Hit rate@5At least one relevant chunk in top 51 (hit)
Recall@52 of 3 relevant chunks retrieved0.67
Reciprocal rankFirst relevant at rank 20.50
Context precisionAverage of precision at rank 2 (1/2) and rank 4 (2/4)0.50
nDCG@5 (binary)DCG = 1/log2(3) + 1/log2(5) = 1.06; ideal (R1-R3 at ranks 1-3) = 1 + 0.63 + 0.50 = 2.130.50
Context recall2 of 3 reference statements supported by context0.67
Faithfulness3 of the answer's 4 claims supported0.75

Retrieval missed the exception and ranked noise first; the generator filled the gap with an unsupported claim. Fix retrieval first (the exception was likely chunked away from its rule), then re-measure faithfulness. Across three questions with first relevant hits at ranks 2, 1 and none, MRR would be (0.5 + 1 + 0) / 3 = 0.5.

Diagnosis table: symptom, metric, fix

SymptomMetric that confirms itLikely fix
Answers miss part of a rule or its exceptionLow context recall, low recall@kStructure-aware or parent-child chunking; raise top-k; check ingestion coverage
Document never found by code or nameLow hit rate on exact-term questionsHybrid keyword plus vector search; identifiers in metadata
Relevant chunk found but buriedLow MRR or nDCG, decent recall@20Add or tune a reranker; adjust fusion
Facts not in any sourceLow faithfulness, recall fineStricter grounding prompt; lower temperature; another model
Correct per an old policyHigh faithfulness, low correctnessFilter superseded documents; refresh references
Citations to chunks never retrievedDeterministic citation check failsStructured citation output validated in code
Confident answers to out-of-scope questionsLow correct-decline rate"Not covered" instruction; retrieval relevance cut-off; permission filters
Refuses legitimate questionsHigh false-decline rateImprove recall; soften over-cautious prompt

Building a test set with references

Each test case should carry the question as real users phrase it (typos and mixed English and regional-language terms included), an expert-approved reference answer, the relevant documents and sections (graded if you will use nDCG), an answerable or should-decline label, the test user's role for permission cases, category tags and the document versions the reference depends on.

Seed it from support tickets and chat logs. Synthetic generation (Ragas and similar tools can produce question, context and reference sets) helps coverage, but have an expert review every synthetic case; generated questions tend to copy the document's wording, which makes retrieval look easier than it is. Include unanswerable, multi-document and superseded-version questions, and re-map relevant IDs after re-chunking.

Calibrating judge metrics against human labels

Judge-based faithfulness, relevance, context precision and citation support are a model's opinion until checked against people:

  1. Sample cases across categories, including known failures.
  2. Have domain experts label the same unit the judge scores: claim supported, chunk relevant, question addressed.
  3. Run the judge on the same sample and measure agreement per metric, as raw agreement and a chance-corrected measure such as Cohen's kappa.
  4. Read disagreements; they usually expose a vague rubric or experts who disagree with each other.
  5. Fix the rubric, pin the judge model and prompt, and repeat until agreement is acceptable to the risk owners. Re-calibrate when the judge, rubric or domain changes.

Prefer a judge model different from the generator, and binary decisions over 1-to-10 scales.

Where Ragas fits

Ragas is an open-source Python library for evaluating RAG and other LLM applications. It implements several of the metrics above, including faithfulness, answer (response) relevance, context precision, context recall and answer correctness, mostly via LLM judges and embeddings, plus synthetic test set generation. Metric names, variants and APIs have changed between releases, so pin your version and check its documentation. Calibrate its scores like any judge; retrieval ranking, citation and refusal checks are often simpler in your own code.

Setting targets from a baseline

There is no universal "good" score; any copied threshold came from someone else's data, judge and rubric. Instead:

  1. Freeze a test set version and a pinned judge, and run the current system for a baseline per metric and category.
  2. Repeat the run unchanged to see how much scores move; that noise band is your minimum meaningful difference.
  3. Review baseline failures with the business owner and agree acceptable levels per category by risk. A wrong dress-code answer and a wrong KYC answer do not cost the same.
  4. Make high-risk cases (fabricated citations, answering a should-decline permission case) must-pass rather than averaged.
  5. Gate releases on no metric dropping beyond the agreed tolerance, and raise the bar as the system improves.

CI wiring and production sampling are covered in the broader evaluation process. For a full build, see the RAG knowledge assistant project; when RAG becomes a tool inside an agent, add the trajectory checks from how to evaluate AI agents. Running this discipline inside a customer's environment, with their data and risk owners, is a large part of what engineers practise in Cloudsoft's FDE PRO program.

Frequently asked questions

What are the main RAG evaluation metrics?

For retrieval: context precision, context recall, hit rate, recall@k, MRR and nDCG. For generation: faithfulness, answer relevance and answer correctness. Enterprise systems also need citation accuracy and refusal correctness, which check that sources are real and that the system declines when the answer is not available.

What is the faithfulness metric in RAG?

Faithfulness measures whether the claims in a generated answer are supported by the retrieved context. A judge model splits the answer into claims and checks each against the context, and the score is the share of supported claims. It needs no reference answer, but a faithful answer can still be wrong if the source is wrong or outdated.

What is the difference between context precision and context recall?

Context precision asks whether retrieved chunks are relevant and ranked near the top, so it penalises noise. Context recall asks whether retrieval found all the information needed for the reference answer, so it penalises missing evidence. Low precision usually calls for reranking or filtering; low recall calls for better ingestion, chunking or search.

What are Ragas metrics?

Ragas is an open-source Python library that implements several RAG evaluation metrics, including faithfulness, response relevance, context precision, context recall and answer correctness, mostly using LLM judges and embeddings. Names and APIs vary between versions, so pin the version you use and calibrate its scores against human labels on your own data.

What is a good faithfulness or context recall score?

There is no universal number. Measure a baseline on your own versioned test set with a pinned judge, check how much scores vary between identical runs, and agree acceptable levels per category with the business owner based on risk. Then gate releases on regressions from that baseline.

How do you check that a RAG system says I don't know correctly?

Label each test case answerable or should-decline, detect refusals with a structured output field or a simple classifier, and report two rates: should-decline cases correctly declined and answerable cases wrongly declined. Tracking only one hides either hallucination or over-refusal.

Which RAG metric should I fix first?

Start with retrieval. If context recall or recall@k is low, the evidence never reached the model and no prompt change will fix the answer. Once retrieval is healthy, work on faithfulness, citation accuracy and answer relevance.

To build retrieval pipelines and practise reading these scores on real projects, explore Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad, in our Ameerpet classroom beside Ameerpet Metro or live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us