New batches starting this week Β· Limited seats

LLM Evaluation Interview Questions and Answers 2026 (50 Questions)

50 LLM evaluation interview questions with model answers, from golden sets and metrics to LLM-as-a-judge calibration, RAG and agent evaluation, CI gates, fairness, cost and production scenarios.

LLM evaluation interview questions 2026: 50 questions on golden sets, metrics, LLM-as-judge, RAG and agent evals and CI gates
Last updated Β· 36 min read Β· 7,996 words

LLM evaluation interview questions test whether you can prove that an AI system works, not just build one: how you assemble an eval set, which metrics you trust, when an LLM judge is reliable, and how you stop a prompt or model change from quietly breaking production. This guide collects 50 high-value questions with model answers, covering golden sets, RAG and agent evaluation, LLM-as-a-judge, online signals, CI regression gates, contamination, fairness and cost, followed by ten scenarios of the kind senior interviewers like to use.

How to use this guide

Evaluation now comes up in almost every AI engineering loop, whether the role says GenAI engineer, AI engineer, MLOps or Forward Deployed Engineer. If you want the conceptual background first, read our LLM evaluation guide. What interviewers typically probe at each level:

  • Freshers and junior engineers: why LLM output cannot be tested with simple assertions, what exact match, F1 and ROUGE measure, and what a golden set is.
  • Mid-level engineers: faithfulness, context precision and recall, judge rubrics, synthetic data, and wiring evaluations into a pipeline.
  • Senior and architect roles: calibrating judges against humans, agent trajectory evaluation, online experiments, contamination, fairness slices, evaluation cost, and diagnosing a regression nobody saw coming.

Answer each question aloud before reading the model answer. Where you can, attach an example from something you have built; interviewers remember a concrete eval story far longer than a list of metric names.

Why evaluation is different for LLMs

1. Why can't you test an LLM application the way you test normal software?

Answer: Traditional software is deterministic: the same input gives the same output, so you assert equality. An LLM application is probabilistic and open-ended. The same prompt can produce different wording on each run, many different answers can be correct, and a fluent answer can be wrong. Quality is also multi-dimensional: an answer can be accurate but unhelpful, helpful but unsafe, or correct but in the wrong format. So instead of pass/fail unit tests on exact strings, you evaluate distributions of behaviour over a representative dataset, using a mix of deterministic checks, reference-based metrics, model-based judges and human review.

Interview tip: Mention that you still keep deterministic tests where they fit: JSON schema validation, required fields, tool-call names, banned phrases. Not everything needs a judge.

2. What exactly are you evaluating: the model or the system?

Answer: Usually the system. A production LLM application is a pipeline: prompt templates, retrieval, re-ranking, tool calls, guardrails, post-processing and the model itself. Model evaluation (public benchmarks, capability tests) tells you whether a model is a reasonable candidate. System or task evaluation tells you whether your application, with your data, prompts and tools, does the job your users need. Swapping the model with everything else fixed is one experiment; changing the chunking strategy with the model fixed is another. Good evaluation lets you attribute a change in quality to the component you changed.

3. What are the main dimensions of LLM output quality?

Answer: A practical set is: correctness (is it factually right for the task), groundedness or faithfulness (is it supported by the provided context), relevance (does it address the question), completeness, format compliance (schema, length, language, tone), safety (toxicity, harmful content, data leakage), refusal behaviour (refuses what it should, answers what it should), and operational qualities (latency, token cost). Not every application needs all of them. A ticket classifier mostly needs accuracy and format compliance; a customer-facing assistant needs groundedness, safety and tone as well.

Interview tip: Tie the dimensions to business risk. "For a bank, an ungrounded answer about charges is a complaint and possibly a compliance issue, so faithfulness gets a hard threshold."

4. What is the difference between offline and online evaluation?

Answer: Offline evaluation runs your system against a fixed dataset before release: it is repeatable, cheap to compare across versions and safe, but only as good as the dataset. Online evaluation measures behaviour on real traffic after release: user feedback, implicit signals, sampled judge scores and A/B tests. It reflects reality but is noisy, slower and carries risk. You need both. Offline evaluation is the gate that stops obvious regressions; online evaluation finds what your dataset missed, and those findings flow back into the offline set.

  production traffic --> sample + label --> golden set
        ^                                     |
        |                                     v
  online signals <-- release <-- CI gate <-- offline eval

5. Why is "it looked good in the demo" not evaluation?

Answer: A demo is a handful of hand-picked inputs, usually the happy path, judged by the person who built the system. It has no coverage of edge cases, no measurement of variance across runs, no baseline to compare against and no record that someone else can reproduce. Evaluation means a versioned dataset that represents real usage, defined metrics with thresholds, results stored per version, and a way to rerun everything after any change. This gap is one of the main reasons AI demos fail in enterprise production.

Building evaluation sets

6. What is a golden set, and how do you build one?

Answer: A golden set is a curated, versioned collection of inputs with expected outputs or grading criteria, reviewed by people who know the domain. To build one: collect real questions (support tickets, search logs, user interviews), cover the main intents and the risky edge cases, write reference answers or acceptance criteria with subject-matter experts, record metadata (intent, difficulty, language, source document), and store it in version control or an evaluation platform. Start small and high quality rather than large and noisy; a few hundred carefully reviewed items usually tell you more than thousands of unchecked ones.

Real-world example: Consider an insurer building a claims-policy assistant. The golden set includes common questions ("Is physiotherapy covered?"), questions whose answer differs by policy variant, questions the assistant must decline (legal advice), and questions with no answer in the documents, each tagged by product line so results can be sliced.

7. When is synthetic evaluation data useful, and what are its risks?

Answer: Synthetic data, generated by an LLM from your documents or from seed examples, is useful when you have little real traffic, need coverage of rare cases, want adversarial or multilingual variants, or must avoid using sensitive production data. The risks: generated questions tend to be easier and more "textbook" than real ones, they inherit the generator's style and blind spots, they can leak the answer phrasing from the source chunk (making retrieval look better than it is), and if the same model family generates and judges, scores can be inflated. Mitigate by having humans review a sample, mixing synthetic with real questions, tagging synthetic items so you can report them separately, and generating harder variants deliberately. Our guide on synthetic data for AI testing covers generation techniques in more depth.

8. How do you use production data to improve an eval set?

Answer: Sample production traces regularly, with stratification so you see each intent and not only the most common ones. Prioritise traces with negative signals: thumbs-down, escalations to a human, retries, long sessions, low judge scores. Have reviewers label them, write the expected behaviour, and add them to the golden set. Before storing anything, apply your data-protection rules: mask or remove personal data and follow consent and retention policies (in India, the DPDP Act applies to personal data you process). Over time the golden set drifts toward the real distribution of hard cases, which is what you want.

Interview tip: Say "every production incident becomes a regression test". It shows you close the loop.

9. How big should an evaluation set be?

Answer: Big enough that the difference you care about is larger than the noise. If you want to detect a small change in pass rate, you need more items than if you only want to catch large regressions; with a few dozen items, a swing of several items between runs can be pure randomness. Practical approach: start with a few hundred items for the core set, run each configuration more than once if outputs are non-deterministic, report confidence intervals or at least run-to-run variance, and keep separate smaller sets for specific risks (safety, prompt injection, refusals, format). Coverage of intents and edge cases matters more than raw count.

10. How do you keep an eval set from going stale or leaking into training?

Answer: Version it, record which system version was scored on which dataset version, and refresh it from production on a schedule. Retire items whose source documents changed, otherwise you penalise correct new answers. Keep a held-out portion that prompt engineers do not look at while tuning, because iterating prompts against the same visible examples overfits just like training does. If you fine-tune, make sure eval items never enter the training data, and check for near-duplicates, not only exact copies.

Metrics

11. When are exact match and F1 appropriate?

Answer: Exact match (after normalising case, whitespace and punctuation) works when there is one canonical answer: a classification label, an extracted invoice number, a yes/no decision, an SQL result set. Token-level F1, the harmonic mean of token precision and recall against a reference, gives partial credit for short extractive answers where wording varies slightly. Both break down for long, free-form answers because a correct paraphrase scores poorly. Use them for structured tasks and extraction, and for the structured parts of a larger output (for example, the category field in a JSON response).

12. What are the limits of BLEU and ROUGE for LLM evaluation?

Answer: BLEU measures n-gram precision against references and came from machine translation; ROUGE measures n-gram recall and came from summarisation. Both reward surface word overlap, not meaning. A correct answer phrased differently scores low, a wrong answer that copies key phrases ("the claim is not covered" versus "the claim is covered") scores high, and they say nothing about faithfulness, safety or helpfulness. They also need reference texts, which are expensive for open-ended tasks. They remain useful as cheap trend signals for tightly constrained tasks such as templated summaries, but you should not gate a release on them alone.

Interview tip: Give the negation example. It shows you understand why overlap metrics are blind to meaning.

13. How does semantic similarity scoring work, and where does it fail?

Answer: You embed the generated answer and the reference answer with an embedding model and compute cosine similarity, or use a cross-encoder that scores the pair directly. It tolerates paraphrase far better than BLEU or ROUGE. It fails in three ways: embeddings capture topic more than truth, so "the fee is 500 rupees" and "the fee is 5,000 rupees" look very similar; thresholds are arbitrary and model-specific; and long answers with one wrong sentence still score high. Use semantic similarity as one signal, combined with claim-level checks for facts, numbers and negations.

14. What is faithfulness (groundedness), and how is it measured?

Answer: Faithfulness measures whether every claim in the answer is supported by the context the system was given, regardless of whether the claim is true in the world. A common method: an LLM breaks the answer into atomic claims, then checks each claim against the retrieved context, and the score is the fraction of supported claims. Natural language inference models can do the entailment check instead of an LLM. Faithfulness is the main metric for catching hallucination in RAG. Note that a faithful answer can still be wrong if the retrieved document is outdated, which is why you also measure correctness against references. Our explainer on LLM hallucinations covers the causes.

15. What do answer relevance, context precision and context recall measure?

Answer: Answer (response) relevance measures whether the answer addresses the question asked, penalising evasive, incomplete or padded responses; one approach generates questions from the answer and compares them with the original. Context precision measures whether the retrieved chunks are relevant and whether the relevant ones are ranked near the top. Context recall measures whether the retrieved context contains all the information needed to produce the reference answer, so it typically needs a reference. Together they separate retrieval problems (low context recall) from generation problems (good context, low faithfulness).

MetricQuestion it answersNeeds reference?
FaithfulnessIs the answer supported by the context?No
Answer relevanceDoes it address the question?No
Context precisionIs retrieved context relevant and well ranked?Depends on variant
Context recallDid retrieval find everything needed?Usually yes
Answer correctnessDoes it match the expected answer?Yes

16. How do you evaluate toxicity and safety?

Answer: Use a dedicated safety set: harmful requests, jailbreak attempts, prompt injection inside documents, requests for personal data, and borderline but legitimate requests. Score outputs with safety classifiers or a provider's content-safety service, plus a judge with a clear policy rubric, and send a sample to human reviewers. Measure both directions: the rate of unsafe completions and the rate of wrongly blocked safe requests. Safety evaluation is continuous, not a one-off, because attacks evolve; it overlaps with AI red teaming, which hunts for new failure modes, while evaluation checks known ones on every release.

17. Why do refusal rates matter, and how do you measure over-refusal?

Answer: A system that refuses too little is unsafe; a system that refuses too much is useless and users route around it. Build two labelled sets: requests that must be refused or escalated, and requests that look sensitive but are legitimate ("how do I report a fraudulent transaction?"). Measure the refusal rate on each. A judge or classifier labels each output as answered, refused or partially answered. Track both numbers over releases, because tightening guardrails or switching models often improves one and silently worsens the other. Refusals should also be graceful: explain, and offer a next step or human handoff.

18. Which deterministic checks should every LLM eval suite include?

Answer: Cheap, unambiguous checks that run on every item before any judge: output parses as valid JSON or matches the schema, required fields present, enum values valid, length within limits, response in the requested language, no leaked system prompt or internal identifiers, no personal data patterns (phone, Aadhaar, PAN formats) where they should not appear, correct citation format, and tool calls with valid names and argument types. They are fast, free, reproducible and catch a surprising share of regressions, especially after a model upgrade. See function calling and structured outputs for how to enforce schemas at generation time.

LLM-as-a-judge

19. What is LLM-as-a-judge, and when should you use it?

Answer: LLM-as-a-judge means using a language model, with a grading prompt and rubric, to score or compare outputs on qualities that are hard to compute: helpfulness, faithfulness, tone, policy compliance, reasoning quality. Use it when outputs are open-ended, the criteria can be written down clearly, and human review at full scale is too slow or expensive. Do not use it where a deterministic check works, and do not trust it until you have measured its agreement with human labels on your own data. A judge is a measurement instrument, and instruments need calibration.

20. How do you design a good judge rubric?

Answer: Make each criterion specific and observable, score one dimension per judge call rather than "overall quality", define every score level with concrete descriptions, include short examples of each level, ask for a brief justification before the score so the reasoning is inspectable, and give the judge exactly the inputs it needs (question, context, reference, answer). Binary or three-level scales are usually more consistent than ten-point scales. Return structured output so scores parse reliably.

Criterion: Faithfulness
Inputs: question, retrieved_context, answer
PASS: every factual claim is supported by context
PARTIAL: minor unsupported detail, core answer OK
FAIL: any claim contradicts or is absent from context
Output JSON: {"reasoning": "...", "label": "PASS"}

Interview tip: Say you version rubrics like code. Changing a rubric changes the metric, so scores before and after are not comparable.

21. What biases do LLM judges have?

Answer: Well-documented ones include position bias (preferring the first or second answer in a pairwise comparison), verbosity bias (preferring longer answers), self-preference (favouring outputs from the same model family), style over substance (rewarding confident, well-formatted text), sensitivity to prompt wording, and weak domain knowledge in specialised fields such as medicine or tax. Mitigations: swap positions and only count consistent verdicts, instruct and test for length neutrality, use a judge from a different model family than the system under test, use references where possible, and calibrate against domain experts.

22. How do you calibrate an LLM judge against human labels?

Answer: Have domain experts label a representative sample independently, ideally with two reviewers per item so you also know human-to-human agreement. Run the judge on the same items. Measure agreement with accuracy plus a chance-corrected statistic such as Cohen's kappa, and look at the confusion matrix, because a judge that misses failures (false passes) is more dangerous than one that is slightly harsh. Read the disagreements, refine the rubric or examples, and repeat on fresh items to avoid overfitting the rubric to the calibration set. Recalibrate whenever you change the judge model, the rubric or the domain. Human-to-human agreement is the practical ceiling; do not expect the judge to beat it.

23. Pairwise or pointwise judging: when do you use each?

Answer: Pointwise judging scores one output against a rubric (pass/fail or a scale). It is good for absolute thresholds, release gates and monitoring because each output gets its own score. Pairwise judging shows two outputs for the same input and asks which is better. It is more sensitive for comparing two prompts or models, because relative judgments are easier than absolute ones, but it does not tell you if either output is acceptable and needs position swapping to control bias. A common pattern: pairwise during development to pick the stronger variant, pointwise in CI and production to enforce minimum quality.

24. How do you make judge scores reliable enough to gate a release?

Answer: Pin the judge model version and set low or zero temperature, version the rubric, use structured output, and measure judge self-consistency by scoring the same items several times. Gate on metrics with proven human agreement only, and set thresholds with a margin that exceeds run-to-run noise. Combine judge metrics with deterministic checks so the gate does not rest on one model's opinion. Keep a small human-reviewed audit sample in each release so drift in the judge itself gets noticed.

RAG evaluation

25. How do you evaluate a RAG system end to end?

Answer: Evaluate retrieval and generation separately, then together. Retrieval: for each question, label which documents or chunks are relevant, then measure recall at k, precision or context precision, and ranking metrics such as MRR or nDCG. Generation: given the retrieved context, measure faithfulness, answer relevance and correctness against references. End to end: answer correctness, citation accuracy (do cited sources actually support the sentence), and behaviour on unanswerable questions. Slice results by document type, intent and language. Our article on RAG evaluation metrics goes metric by metric, and the RAG interview questions guide covers the architecture side.

 question -> retriever -> chunks -> generator -> answer
               |                       |          |
      recall@k, ranking     faithfulness   correctness,
      context precision     relevance      citations

26. How do you evaluate behaviour on questions that have no answer in the knowledge base?

Answer: Include deliberately unanswerable questions in the golden set: topics outside the corpus, questions about products that do not exist, and questions whose answer was removed in a document update. The expected behaviour is an honest "I could not find this in the available documents" plus a next step, not a plausible invention. Measure the rate of correct abstention and the rate of false abstention on answerable questions. This catches the most damaging RAG failure: confident answers built from loosely related chunks.

27. How do you tell whether a RAG regression comes from retrieval or generation?

Answer: Look at the per-stage metrics on the failing items. If context recall dropped, the needed information never reached the model: check chunking, embeddings, filters, index freshness and re-ranker changes. If context recall is fine but faithfulness or correctness dropped, the model had the information and misused it: check prompt changes, context ordering, context length and model version. A useful diagnostic is to rerun the generator with gold context injected; if answers recover, the problem is retrieval.

Real-world example: Consider a GCC IT team in Hyderabad whose helpdesk assistant started giving wrong VPN instructions. Context recall had dropped only for questions about one tool, and the cause was a re-indexing job that silently skipped scanned PDFs. Generation was fine.

Agent evaluation

28. How is evaluating an agent different from evaluating a single LLM call?

Answer: An agent takes multiple steps: it plans, calls tools, reads results and decides what to do next, often across several turns. So you evaluate at more than one level: the final outcome (was the task completed correctly), the trajectory (were the steps sensible, efficient and safe), and individual steps (was each tool selected and called correctly). Agents also have side effects, so evaluation must run against sandboxed or mocked tools, and you need metrics for cost and step count, because an agent that succeeds after many redundant tool calls is still a problem. The AI agent evaluation guide covers this in depth.

29. Trajectory evaluation versus outcome evaluation: what is the difference?

Answer: Outcome evaluation checks the end state: was the refund issued for the right amount, is the ticket in the right queue, does the database contain the expected row. It is the closest to business value and tolerant of different valid paths. Trajectory evaluation inspects the sequence of actions: did the agent verify identity before changing an account, did it avoid calling a write tool before reading, did it loop or repeat calls. It catches dangerous or wasteful behaviour that happened to end well. Trajectories can be compared to a reference sequence (exact, in-order or any-order match) or graded by a judge against rules. Use outcome for success rate, trajectory for safety and efficiency.

30. How do you measure tool-call correctness?

Answer: Break it into tool selection (right tool for the step, or correctly no tool), argument correctness (valid schema, correct values extracted from the conversation, no hallucinated IDs), call ordering and dependencies, and handling of tool errors (retry, ask the user, or escalate instead of inventing a result). Deterministic checks validate names and schemas; reference trajectories check expected arguments; a judge assesses whether the choice made sense given the context. For write actions, also check that approvals were requested where policy requires them.

Interview tip: Mention that you evaluate with mocked tools that return realistic success, empty and error responses, because an agent that only ever sees happy tool results has not really been tested.

31. How do you evaluate a multi-turn conversation or session?

Answer: Score at session level for goal completion and at turn level for correctness, relevance and instruction following. To test at scale, use a simulated user, an LLM given a persona and a goal, that converses with the agent until the goal is met or a turn limit is hit, then judge the transcript. Check that the agent keeps context across turns, does not repeat questions already answered and handles mid-conversation changes ("actually, make it the other account"). Managed services now support this pattern; for example, Amazon Bedrock AgentCore Evaluations offers built-in evaluators at session level (such as goal success rate), trace level (such as correctness, faithfulness and helpfulness) and tool level (tool selection and parameter accuracy), scored from OpenTelemetry-based traces. Check current documentation for the evaluator list. The Bedrock AgentCore interview questions guide covers the service itself.

Online evaluation

32. How do you run an A/B test for an LLM feature?

Answer: Define the primary metric before starting (task completion, resolution without escalation, conversion), plus guardrail metrics (safety incidents, latency, cost, complaint rate). Randomise at user or session level, not per request, so a user does not bounce between variants within one conversation. Decide the sample size and duration in advance to avoid stopping when the numbers look good. Watch for novelty effects, because users interact differently with something new. Combine the experiment with sampled judge scores on both arms, so you understand why one variant wins.

33. What explicit and implicit user signals are useful, and what are their pitfalls?

Answer: Explicit signals: thumbs up/down, ratings, written feedback, "report a problem". Implicit signals: copying the answer, follow-up rephrasing of the same question, abandoning the session, escalating to a human, editing a generated draft heavily before sending, accepting or rejecting a suggestion. Pitfalls: explicit feedback is sparse and skewed toward strong opinions, implicit signals are ambiguous (a short session can mean success or frustration), and both are biased by UI design. Treat them as pointers to traces worth reviewing, not as ground truth.

34. How do you monitor quality continuously in production?

Answer: Trace every request with inputs, retrieved context, tool calls, outputs, latency and cost, using OpenTelemetry or a tracing platform. Run cheap deterministic checks on all traffic, and run judge evaluators on a sample, stratified by intent and weighted toward risky flows. Put scores on dashboards with alerts on drops, and route low-scoring traces to a review queue that feeds the golden set. Track drift in input topics too: a new product launch can change what users ask. The broader setup is covered in our AI observability guide.

Regression gates in CI

35. How do you build an evaluation gate in a CI/CD pipeline?

Answer: Treat prompts, retrieval configuration, model IDs and rubrics as versioned artefacts. On each pull request that touches them, run a fast suite: deterministic checks plus a small, representative judge-scored subset. On merge or before release, run the full suite. Compare against the current production baseline, not an absolute number alone, and fail the build if any blocking metric drops beyond a set tolerance or any critical test (safety, policy) fails. Post a report with per-slice scores and the specific failing examples so reviewers can act. Our article on CI/CD for AI applications shows the pipeline structure.

 PR opened
   |
   v
 lint + schema tests --> fast eval subset --> report
   |                                          |
 merge --> full eval vs baseline --> pass? --> deploy
                                       |
                                       no --> block

36. How do you handle non-determinism and flaky evaluation results in CI?

Answer: Reduce variance where possible: fixed temperature, pinned model and judge versions, cached judge results for unchanged outputs. Measure the remaining variance by running the baseline several times, and set tolerances wider than that noise. Separate hard gates (schema, safety, critical policy items, which must pass every time) from soft metrics (judge averages, which must not drop beyond a tolerance). For borderline results, rerun the affected slice or escalate to human review instead of letting engineers retry until green.

Interview tip: Interviewers like hearing that you would rather have a smaller, stable gate that people trust than a large, flaky one that everyone learns to ignore.

Benchmarks, contamination, fairness and cost

37. What is the difference between public benchmarks and task evals, and what is contamination?

Answer: Public benchmarks measure general capabilities (reasoning, coding, knowledge) on standard datasets and help shortlist models. Task evals measure your application on your data and decide whether something ships. Benchmark scores often fail to predict task performance because your domain, language mix, documents and format rules differ. Contamination means benchmark items, or close paraphrases, appeared in a model's training data, so a high score may reflect memorisation rather than ability. Signs include suspiciously high scores on older benchmarks and sharp drops on freshly written equivalent items. The defence for your own evals: keep private, held-out, regularly refreshed sets that were never published. The LLM interview questions guide covers benchmark design from the model side.

38. How do you evaluate an LLM application for fairness?

Answer: Define which groups and attributes matter for the use case and the law (gender, region, language, age, disability, caste or religion where relevant to Indian deployments). Build counterfactual test pairs that change only the attribute, such as the same loan query with different names or cities, and compare outputs, scores, tone and refusal rates. Slice all main metrics by language and user segment, because quality often drops for regional languages or code-mixed Hinglish. Use judges for stereotyping and tone, calibrated against diverse human reviewers. Document findings and fixes. Our guide on bias and fairness testing for AI walks through methods in detail.

Real-world example: Consider a retailer whose product assistant answers well in English but gives shorter, less accurate answers to Telugu and Hindi queries. An aggregate score would hide this; a per-language slice exposes it immediately.

39. How do you control the cost of evaluation?

Answer: Evaluation cost comes from running the system under test, running judges, and human review time. Controls: run deterministic checks first and only judge items that pass them; use a smaller judge model for simple criteria after verifying its agreement with humans; cache outputs and judge scores for unchanged items; tier your suites (small on every PR, full nightly or pre-release); sample production traffic instead of scoring all of it; and use batch inference where the provider offers it. Spend human review time on calibration and disagreement cases, not on re-checking easy passes. Track evaluation spend per release like any other cloud cost.

Evaluation tooling

40. What kinds of evaluation tools exist, and how would you choose between them?

Answer: The tools fall into a few categories, and many teams combine more than one:

  • Metric libraries: Ragas (RAG and agent metrics such as faithfulness, response relevancy, context precision and context recall) and DeepEval (pytest-style LLM test cases with built-in and custom metrics, including G-Eval-style rubric judges).
  • Tracing and evaluation platforms: LangSmith (datasets, experiments, online evaluators, tracing for LangChain and other code) and Langfuse (open-source tracing, datasets, LLM-as-a-judge evaluators and human annotation, self-hostable).
  • Config-driven test runners: promptfoo (open-source CLI where prompts, providers and assertions are declared in config, with red-teaming features; now owned by OpenAI and still maintained as open source).
  • Provider evaluation services: Amazon Bedrock Evaluations for models and knowledge bases, Amazon Bedrock AgentCore Evaluations for agents (online, on-demand and batch evaluation with built-in and custom evaluators), evaluators in Microsoft Foundry (for example groundedness and relevance), and the Gen AI evaluation service on Google Cloud's Vertex AI (Google has been rebranding parts of Vertex AI, so check current naming).

Choose based on where your traces already live, whether data must stay in your environment (self-hosting, region), which metrics you need, CI integration, and team familiarity. The metric definitions matter more than the tool: always read how a library computes a metric before trusting its number, because the same metric name can mean different things in different tools.

Interview tip: Avoid tool loyalty in interviews. Say what each category does and how you would validate any tool's metrics on your own data.

If you want to practise this end to end, from building golden sets and judges to wiring evaluation into CI on cloud infrastructure, the APEX AI, ML, Cloud and Cyber Security program covers evaluation alongside model development, cloud deployment and AI security.

Real-world scenario questions

41. Your LLM judge disagrees with human reviewers on a large share of items. What do you do?

Answer: Stop using the judge for gating until agreement is understood, then treat it as a measurement problem: find out whether the judge, the rubric or the human labels are at fault.

What I would check:

  1. Human-to-human agreement on the same items. If reviewers disagree with each other, the rubric is ambiguous and no judge will fix it.
  2. The confusion matrix: is the judge too lenient (false passes) or too strict, and on which slices?
  3. A sample of disagreements read side by side, looking for patterns such as length bias, missing domain knowledge or a criterion the humans apply that the rubric never states.
  4. Whether the judge receives the same inputs as humans (context, reference, policy document).
  5. Whether the judge model or rubric changed recently without recalibration.

Production consideration: Rewrite the rubric with explicit level descriptions and examples drawn from the disagreements, consider a stronger or different-family judge, split a compound criterion into separate judges, and re-measure on fresh items. Keep human review in the loop for the slice where the judge stays unreliable.

42. Offline eval scores went up after a release, but user complaints increased. Why?

Answer: The eval set or metrics no longer represent what users value, or the release changed something the eval does not measure.

What I would check:

  1. Whether the golden set matches current traffic: compare intent distribution in production with the eval set.
  2. Unmeasured dimensions: latency, answer length, tone, refusal rate, formatting on mobile, language.
  3. Judge bias: a new prompt producing longer, more confident answers can raise judge scores while annoying users.
  4. Overfitting: prompts tuned repeatedly against the visible eval set.
  5. Complaint traces themselves, read in detail and grouped by failure type.

Production consideration: Add the complaint cases to the golden set, add the missing metrics, run a held-out set, and use online A/B testing for changes that affect user experience. Metrics are proxies; when proxy and reality diverge, reality wins.

43. A new model version breaks your output format and downstream parsing fails. How do you handle it?

Answer: Roll back or pin the previous model version immediately, then find out why the format check did not catch it before release.

What I would check:

  1. Whether the model was pinned to a specific version or an alias that moved automatically.
  2. Which format failures appear: extra prose around JSON, markdown fences, renamed fields, different date formats, changed tool-call structure.
  3. Whether structured output or schema-constrained decoding was available and unused.
  4. Whether the eval suite had schema-validation tests on enough real inputs, including long and multilingual ones.
  5. How the parser behaves on failure: does it retry, repair or fail loudly?

Production consideration: Pin model versions, treat every model change as a release that must pass the full eval gate, use native structured outputs where supported, validate every response against a schema at runtime with a safe fallback, and keep a model-migration checklist. Prompts that worked on one model often need adjustment for the next.

44. Faithfulness scores are high, but subject-matter experts say many answers are wrong. What is happening?

Answer: The answers are faithful to the retrieved context, but the context itself is wrong, outdated or incomplete, or the faithfulness judge is too lenient.

What I would check:

  1. Source documents behind the wrong answers: superseded policy versions, duplicates with conflicting content.
  2. Context recall and correctness against references, not faithfulness alone.
  3. Whether the judge checks every claim or just the gist.
  4. Metadata filters for effective dates and document status.

Production consideration: Faithfulness measures consistency with context, not truth. Pair it with answer correctness on an expert-reviewed set and fix content governance, so retired documents leave the index.

45. Your agent's task success rate is good in testing but it is too expensive and slow in production. How do you evaluate the fix?

Answer: Add efficiency to the evaluation, then optimise against quality and cost together.

What I would check:

  1. Step counts and tool calls per task from traces, looking for loops, repeated reads and unnecessary planning steps.
  2. Token usage per step and how much context is resent each turn.
  3. Whether a smaller model could handle routing or simple steps without lowering outcome scores.
  4. Production task mix compared with the test set; real tasks may be longer.

Production consideration: Report success rate, average and tail latency, and cost per successful task side by side for each variant. A change ships only if success stays within tolerance while cost or latency improves. Our LLM latency optimisation guide lists the usual levers.

46. A stakeholder asks, "Is the assistant ready for launch?" How do you answer with evidence?

Answer: Translate evaluation results into risk-based readiness criteria that were agreed with the business before testing began.

What I would check:

  1. Results against agreed thresholds for each dimension: correctness, faithfulness, safety, refusals, latency, cost.
  2. Per-slice results, especially for high-risk intents and less common languages.
  3. Known failure modes and their mitigations (human handoff, guardrails, scope limits).
  4. Monitoring and rollback readiness for launch day.

Production consideration: Recommend a staged rollout, such as internal users first and then a limited customer segment, with online evaluation running and clear criteria for widening. Readiness is a documented decision, not a feeling.

47. Consider a hospital deploying a discharge-summary assistant. How would you design its evaluation?

Answer: Clinical text is high risk, so evaluation is built around clinician review and strict critical-error definitions, with automated metrics as support.

What I would check:

  1. A golden set of de-identified records reviewed by clinicians, covering common and complex cases.
  2. Claim-level faithfulness to the source notes, with zero tolerance for invented medications, doses or diagnoses.
  3. Omission checks for critical items such as allergies and follow-up instructions.
  4. Judge calibration against clinicians, with ongoing clinician sign-off on a sample.
  5. Privacy controls on evaluation data and logs.

Production consideration: The assistant drafts and a clinician approves. Evaluation continues after launch by tracking how much clinicians edit each draft, which is a strong implicit quality signal.

48. Two prompt versions have nearly identical average scores. How do you decide?

Answer: Averages hide differences, so look at the distribution, slices and cost before declaring a tie.

What I would check:

  1. Whether the difference is within run-to-run noise; if so, it is a tie on quality.
  2. Per-slice results: one version may be stronger on high-risk intents.
  3. Failure severity, not just count: one critical failure outweighs several minor ones.
  4. A pairwise judge comparison with position swapping, plus human review of differing outputs.
  5. Token cost, latency and maintainability of each prompt.

Production consideration: If quality is truly equal, pick the cheaper, simpler prompt, or run an online test if user experience might differ.

49. Your RAG assistant performs worse on Hindi and Telugu questions. How do you investigate and fix it?

Answer: Treat each language as its own evaluation slice and test the pipeline stage by stage.

What I would check:

  1. Whether the golden set has enough real, native-written questions per language, including code-mixed and transliterated text.
  2. Retrieval recall per language: does the embedding model handle these languages well, and are the documents in the same language as the queries?
  3. Generation quality and fluency per language, judged by native speakers.
  4. Whether the judge itself is reliable in that language.

Production consideration: Fixes may include a multilingual embedding model, query translation before retrieval, language-aware re-ranking, or translated documents. Then make per-language thresholds part of the CI gate so regressions in one language cannot hide in the average.

50. You join a team with no evaluation at all for a live assistant. What do you set up in your first few weeks?

Answer: Build the smallest loop that produces trustworthy signal, then grow it.

What I would check:

  1. Tracing on all requests, if missing, so you can see inputs, context, outputs and cost.
  2. A first golden set from real traces and stakeholder interviews, reviewed with domain experts.
  3. Deterministic checks and two or three judge metrics tied to the main business risks, calibrated on a labelled sample.
  4. A baseline score for the current production version.
  5. A CI gate on prompt and model changes, and a weekly review of low-scoring production traces.

Production consideration: Start with metrics the business understands and trusts. A small, credible evaluation system that blocks one bad release earns the support you need to expand it. This is the kind of work Forward Deployed Engineers do inside customer organisations: take an AI system from demo to measurable enterprise outcome. The FDE engineer interview questions guide covers that role.

Key takeaways

  • LLM evaluation measures behaviour across a representative dataset, combining deterministic checks, reference metrics, calibrated judges and human review.
  • A versioned golden set built from real usage and expert review is the foundation; synthetic data adds coverage but must be reviewed and reported separately.
  • Overlap metrics such as BLEU and ROUGE are weak for open-ended answers; faithfulness, relevance and context precision and recall diagnose RAG stage by stage.
  • An LLM judge is an instrument: write specific rubrics, control for bias and calibrate against human agreement before gating on it.
  • Agents need outcome, trajectory and tool-level evaluation, run against sandboxed tools, with cost and step count measured alongside success.
  • Offline gates in CI stop known regressions; online signals and A/B tests find what the dataset missed, and those cases flow back into it.
  • Slice every metric by intent, language and user group; averages hide the failures that matter.

Interview preparation checklist

  • Build a small golden set (fifty to a hundred items) for a RAG app and score it with deterministic checks plus two judge metrics.
  • Label a sample yourself, measure judge agreement with your labels, and be ready to discuss the disagreements.
  • Run one evaluation library such as Ragas or DeepEval and explain exactly how one of its metrics is computed.
  • Wire an eval run into a GitHub Actions workflow that fails on a regression against a stored baseline.
  • Evaluate one agent task with mocked tools that return success, empty and error responses.
  • Prepare one story about a metric that misled you and what you changed.
  • Be able to sketch the offline-to-online feedback loop on a whiteboard.
  • Review the cost of your eval runs and how you would reduce it.
  • Practise the scenario questions aloud, using the check-then-fix structure.

FAQ

What skills are needed for LLM evaluation roles?

You need Python, statistics basics such as sampling, variance and agreement measures, an understanding of RAG and agent architectures, experience writing judge rubrics, and familiarity with tracing and CI pipelines. Domain sense matters too, because good evaluation criteria come from understanding what users and the business actually need.

Is LLM evaluation a separate job or part of AI engineering?

Both exist. Some teams have dedicated evaluation or AI quality engineers, but in most teams evaluation is a core responsibility of AI engineers, GenAI engineers, MLOps engineers and Forward Deployed Engineers. Interviewers increasingly treat evaluation skill as a sign of production experience.

How should a fresher prepare for LLM evaluation interview questions?

Learn the metric definitions in this guide, then build one small RAG project with a golden set, a few judge metrics and a CI check. Being able to show your eval results, explain a failure you found and describe how you fixed it is more convincing than memorising metric names.

Do I need to know statistics for AI evaluation interviews?

You need practical statistics rather than theory: sample size and noise, confidence intervals, comparing two variants, and agreement measures such as Cohen's kappa. Senior roles may also ask about experiment design for A/B tests and avoiding early stopping.

Which tools should I learn for LLM evaluation?

Learn one metric library such as Ragas or DeepEval, one tracing and evaluation platform such as LangSmith or Langfuse, and understand how a config-driven runner such as promptfoo works. Awareness of cloud provider evaluation services helps for platform-specific roles. Concepts transfer between tools, so depth in one is more valuable than shallow knowledge of all.

Are RAG evaluation interview questions different from general LLM evaluation questions?

They overlap, but RAG questions focus on separating retrieval quality from generation quality using metrics such as context recall, context precision and faithfulness, plus behaviour on unanswerable questions and citation accuracy. General LLM evaluation also covers judges, safety, agents and online testing.

How are agent evaluation interview questions usually framed?

They usually ask how you measure task success, how you evaluate the sequence of tool calls, how you test safely when tools have side effects, and how you handle multi-turn sessions. Expect at least one scenario about an agent that succeeds but is too slow, too costly or takes unsafe shortcuts.

Is AI evaluation a good career direction for Indian engineers?

Evaluation skills are increasingly expected across AI roles at GCCs, product companies and services firms in Hyderabad, Bengaluru and elsewhere, because enterprises need evidence before they ship AI to customers. The skills also combine well with testing, SDET, data and MLOps backgrounds.

Evaluation is where AI projects earn trust. If you want to build that skill with hands-on projects, cloud deployment and AI security, explore the APEX program at Cloudsoft. If your goal is to take AI systems into customer environments, including RAG assistants and agents evaluated with Ragas, LangSmith and Langfuse, look at the AI Forward Deployed Engineer FDE PRO program, which includes five enterprise projects, a capstone and placement support until you're placed. Both run as classroom sessions in Ameerpet or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us