New batches starting this week Β· Limited seats

LLM Evaluation: How to Test LLM Apps Before and After Launch

LLM evaluation turns 'it looked fine when I tried it' into evidence. This guide covers test sets, evaluation methods, RAG and agent metrics, CI regression gates and production monitoring.

LLM evaluation workflow: test set, automated checks, LLM-as-judge, human review and monitoring in production
Last updated Β· 15 min read Β· 3,287 words

LLM evaluation is how you find out, with evidence rather than impressions, whether a large language model application gives correct, grounded, safe answers, and whether it still does after you change a prompt, model, index or tool. In practice it means a versioned test set, a mix of deterministic checks, reference comparisons, LLM-as-judge scoring and human review, run before every release as a regression gate and continued after launch through tracing, user feedback and sampled production review. This guide covers how to build that system for chat assistants, RAG applications and AI agents.

Why LLM evaluation matters

Traditional software is tested with assertions: the same input gives the same output, and a unit test either passes or fails. LLM applications break that assumption in three ways.

  • Non-determinism. The same question can produce differently worded answers on each run, and sometimes a different conclusion. One good-looking response tells you very little about the next hundred.
  • Silent regressions. A prompt tweak that fixes one complaint can break five cases nobody re-checked. A model upgrade from your provider, a new chunking strategy or a re-embedded index can shift behaviour across the board without any error being raised.
  • Open-ended outputs. "Correct" is often a judgement: is the answer complete, grounded in the right source, in the right tone, free of anything it should not reveal? You need ways to score qualities, not just compare strings.

Most of the production failures described in why AI demos fail in enterprise production trace back to one root cause: the team never measured quality systematically, so they could not see it degrade. Evaluation turns "it looked fine when I tried it" into a number you can track and protect in CI.

Building an evaluation test set

Your evaluation is only as good as the questions you test with. A test set (also called an eval set, golden set or dataset) is a versioned collection of inputs, each with whatever you need to judge the output: a reference answer, the source documents that should be used, the tool calls that should happen, or simply a rubric.

What to include

  • Golden questions. The core tasks the application exists for, with reference answers and source references written or approved by domain experts.
  • Edge cases. Ambiguous questions, questions spanning two documents, questions about superseded policy versions, regional language mixed with English and questions whose correct answer is "this is not covered; contact the HR team".
  • Adversarial cases. Prompt injection attempts ("ignore your instructions and..."), attempts to extract other users' data or the system prompt, jailbreak phrasing, malicious content planted inside retrieved documents, and requests that should be refused.
  • Cases from real logs. Once you have pilot or production traffic, the best new test cases come from it: questions users actually asked, especially ones that got thumbs-down feedback, escalations or support tickets. Redact personal data before they enter the dataset.

How to manage it

Store the test set in version control or in your evaluation platform with explicit versions, so every score is tied to a dataset version. Tag each case by category so you can see where quality moved. Start with a few dozen high-quality cases rather than hundreds of unchecked ones; synthetic cases help coverage but need human review. Keep a separate held-out slice you do not tune against, so you can detect when the team has overfitted prompts to the visible set.

Types of evaluation

No single method covers everything. Mature teams layer four kinds of checks, cheapest and most reliable first.

MethodWhat it checksStrengthsLimitations
Deterministic checksFormat, schema, required fields, banned content, citation IDs, latency, token costFast, cheap, fully repeatableCannot judge meaning or quality
Reference-basedSimilarity or match against a known correct answerObjective when a clear answer existsPenalises correct answers worded differently; needs reference answers
LLM-as-judgeFaithfulness, relevance, completeness, tone, policy compliance via a rubricScales to open-ended outputsBiased and inconsistent unless calibrated against humans
Human reviewDomain correctness, usefulness, subtle riskThe ground truth for judgement callsSlow, costly, needs clear guidelines

Deterministic checks

Anything you can assert in code, assert in code. Does the output parse as the expected JSON? Is the classification label one of the allowed values? Does every citation point to a chunk that was actually retrieved? Does the response avoid account numbers, Aadhaar or PAN patterns? Did the request stay under your latency and cost budget? These checks are cheap, repeatable and catch many real bugs.

Reference-based evaluation

When there is a known right answer, compare against it. For extraction and classification, exact match, precision, recall and F1 work well. For free text, string metrics are weak because correct answers vary in wording; embedding similarity or an LLM judge asked "does this answer contain the same facts as the reference?" works better.

LLM-as-judge

An LLM-as-judge uses a model to grade outputs against a rubric: "Given the question, the retrieved context and the answer, is every claim in the answer supported by the context? Answer PASS or FAIL and explain why." It scales to open-ended quality, and it is also the easiest method to fool yourself with. Known pitfalls:

  • Position bias: in pairwise comparisons, judges tend to favour whichever answer appears first (or second). Swap the order and run both.
  • Verbosity bias: longer, more confident answers often score higher even when they are no more correct.
  • Self-preference: a model may rate outputs from its own family more favourably. Where possible, judge with a different model from the one that generated the answer.
  • Vague rubrics: "rate quality from 1 to 10" produces noisy scores. Binary or small categorical judgements with explicit criteria are far more stable.
  • Drift: if the judge model changes, your scores change. Pin the judge model and prompt, and version them like code.

Calibrate before you trust it. Have domain experts label a sample of outputs, run the judge on the same sample and measure agreement (simple agreement rate or Cohen's kappa). Disagreements usually reveal an unclear rubric. Refine and repeat until agreement is acceptable to the people who own the risk, then re-check periodically. An uncalibrated judge is an opinion with a decimal point.

Human review

Humans remain the final authority for high-stakes judgement: whether a claims explanation is accurate under policy, whether a clinical summary omits something important. Give reviewers a written guideline with examples, a simple label set and a queue prioritised by risk. Their labels feed back into the test set and into judge calibration.

RAG evaluation metrics explained

A retrieval-augmented generation (RAG) system can fail in retrieval, in generation, or both, so you evaluate the two halves separately. Four metrics, popularised by the Ragas framework and now common vocabulary across tools, do most of the work.

MetricQuestion it answersNeeds a reference?A low score suggests
FaithfulnessAre the claims in the answer supported by the retrieved context?NoThe model is adding information not in the sources
Answer relevanceDoes the answer actually address the question asked?NoEvasive, off-topic or padded answers
Context precisionAre the retrieved chunks relevant, and are relevant ones ranked near the top?UsuallyNoisy retrieval or weak reranking
Context recallDid retrieval find all the information needed to produce the reference answer?YesMissing documents, poor chunking, wrong filters

Faithfulness is typically computed by breaking the answer into individual claims and checking each one against the context, so the score reflects the share of supported claims. Note what it does not measure: an answer can be perfectly faithful to the wrong document. Answer relevance checks focus, not truth. Context precision and context recall tell you whether to fix the retriever before touching the prompt; if recall is poor, no prompt engineering will help because the answer was never in the context.

Read them together. High recall with low faithfulness means retrieval is fine and the generator is drifting. Low recall with high faithfulness means the model is honestly answering from incomplete evidence. Add answer correctness against a reference where you have one, and a deterministic citation check. Do not copy a "good score" from a blog post: the acceptable level depends on your domain and risk, and should be agreed with the business owner using your own baseline. The same approach applies whether you choose RAG or fine-tuning.

Evaluating AI agents

AI agents plan, call tools and take actions over several steps, so judging only the final message misses most of what can go wrong. Evaluate the outcome and the path.

  • Task success. Did the agent achieve the goal? Define it checkably: the ticket was created with the right category and priority, the refund was calculated correctly, the report contains the required fields. Where possible, verify against the state of the system (in a sandbox), not the agent's claim that it succeeded.
  • Tool-call accuracy. Did it choose the right tool, with valid and correct arguments, in a sensible order? Compare against expected tool calls for each test case and check arguments in code.
  • Unsafe-action rate. How often does the agent attempt something it should not: an unauthorised write, an action outside its scope, skipping a required human approval, acting on instructions injected through a document or email? For high-risk actions the target is "never", enforced by permissions and approval gates, with evaluation proving the guardrails hold.
  • Efficiency. Steps, tokens and time per task.
  • Trajectory review. Have reviewers read full traces of the agent's reasoning, tool calls and observations for a sample of tasks, especially failures and near-misses.

Consider an IT team in a Hyderabad global capability centre (GCC) piloting an agent that triages service desk tickets and resets passwords through a ticketing system's API. The final replies in testing look polite and correct. Trajectory review shows something else: for some requests the agent calls the reset tool before verifying the requester's identity, because the verification tool's description is ambiguous. A tool-order check and an unsafe-action test case for "reset without verification" now block any release where it happens. Our agentic AI interview questions cover these agent design trade-offs in more depth.

If you want to practise this hands-on, building RAG pipelines and agents and then measuring them with Ragas and LangSmith, Cloudsoft's AI, GenAI and Agentic AI course covers evaluation as part of the build, not as an afterthought.

Online evaluation and monitoring after launch

Offline evaluation tells you a release is safe to ship. Online evaluation tells you how it behaves with real users, real documents and real load, which always differ from the test set.

Tracing

Record every request as a trace: the input, the prompt as sent, retrieved chunks with scores, each tool call and result, model and prompt versions, token counts, latency and the final output. LangSmith and Langfuse provide LLM-specific tracing with evaluation built on top, and OpenTelemetry gives a vendor-neutral way to emit traces into the observability stack your platform team already runs. Redact sensitive fields before storage.

Feedback and sampled scoring

  • Explicit feedback: thumbs up/down with an optional reason, attached to the trace.
  • Implicit signals: rephrased questions, escalations to a human, tickets reopened after an agent closed them.
  • Sampled automated scoring: run reference-free metrics such as faithfulness and answer relevance, plus your calibrated judges, on a sample of production traces.
  • Human review queues: route low-scoring, negatively rated and high-risk traces to reviewers.

Alert on trends, not single bad answers: a drop in faithfulness for one document category, rising refusals after an index rebuild, a cost-per-request spike after a prompt change. Confirmed failures become new test cases, so the same bug cannot ship twice. If you come from operations, this is service reliability applied to AI; Cloudsoft's SRE training covers the observability, SLO and alerting foundations it builds on.

Evaluation in CI: regression gates

Treat prompts, model choices, retrieval settings, tool definitions and judge rubrics as code. Any change to them triggers the evaluation suite in your pipeline, for example in GitHub Actions, just like unit tests.

  • Tier the suite. Fast deterministic checks and a small smoke set on every pull request; the full test set with LLM judges on merge or before release.
  • Compare against a baseline, not an absolute number. The gate fails if a metric drops beyond a tolerance the team agreed, or if any must-pass case fails (adversarial, safety and compliance cases should be must-pass).
  • Handle non-determinism. Run noisy cases several times and check per-category results, not just the average.
  • Re-run on external change. Schedule the suite when the provider updates a model or when source documents are re-ingested, even if your code did not change.

Here is the workflow end to end:

 real logs + experts + adversarial cases
        |
        v
 versioned test set (tagged by category)
        |
 change: prompt / model / index / tool
        |
        v
 CI: deterministic checks -> smoke set
        |
        v
 full suite: reference + calibrated judges
        |
   gate vs baseline --fail--> fix, re-run
        | pass
        v
 release (canary) -> tracing + feedback
        |
        v
 sampled scoring + human review queue
        |
        +--> new failures added to test set

This loop is the Observe and Improve part of moving AI from demo to enterprise outcome. Our playbook on taking AI from POC to production shows where these evaluation gates sit in each project stage.

LLM evaluation tools

The tooling changes quickly; a few widely used options:

ToolWhat it isTypical use
RagasOpen-source Python library for evaluating RAG and LLM applicationsFaithfulness, answer relevance, context precision and recall; test set generation
DeepEvalOpen-source evaluation framework with a pytest-style interfaceWriting LLM test cases with built-in metrics that run in CI
promptfooOpen-source CLI and library driven by config filesComparing prompts and models side by side; red-teaming and security testing
LangSmithLangChain's platform for tracing, datasets and evaluation (works with or without LangChain)Tracing, dataset management, experiments, online evaluators, annotation queues
LangfuseOpen-source LLM engineering platform, self-hostable or managedTracing, prompt management, evaluation scores and human annotation
Arize PhoenixOpen-source observability and evaluation tool built on OpenTelemetryTrace inspection, LLM-judge evaluations, debugging retrieval
OpenTelemetryVendor-neutral standard and SDKs for traces, metrics and logsEmitting AI traces into existing observability backends

A common combination is Ragas or DeepEval for offline metrics in CI plus LangSmith or Langfuse for tracing and online scoring. Keep test sets and rubrics exportable to avoid lock-in.

Common LLM evaluation mistakes

  • Vibe checks instead of a test set. Trying a few questions by hand before each release is not evaluation and cannot catch regressions.
  • Engineers writing the ground truth alone. Without domain experts, you measure agreement with the engineer, not correctness.
  • Trusting an uncalibrated judge. Judge scores without human agreement checks can be confidently wrong.
  • Only looking at averages. A stable average can hide a collapse in one category.
  • Evaluating only the final answer of an agent. Unsafe or wasteful trajectories that happen to end well are still failures.
  • Stopping at launch. Users, documents and models change; without online evaluation, quality decays unseen.

Designing this evaluation discipline inside a customer's environment, with their data, experts and risk owners, is a core part of what Forward Deployed Engineers do; Cloudsoft's FDE PRO program practises it across its enterprise projects and the simulated GlobalBank engagement.

Frequently asked questions

What is LLM evaluation?

LLM evaluation is the practice of systematically measuring the quality, safety and reliability of a large language model application using a test set and a combination of deterministic checks, reference comparisons, LLM-as-judge scoring and human review, both before release and in production.

How do you evaluate an LLM application?

Build a versioned test set of golden questions, edge cases, adversarial cases and real user questions. Run deterministic checks first, then reference-based metrics and calibrated LLM judges, with human review for high-risk cases. Gate releases in CI against a baseline and keep evaluating in production through tracing, feedback and sampled scoring.

What are the main RAG evaluation metrics?

The four most common are faithfulness (is the answer supported by the retrieved context), answer relevance (does it address the question), context precision (are retrieved chunks relevant and well ranked) and context recall (did retrieval find everything needed for the reference answer).

Is LLM-as-judge reliable?

It can be, once calibrated. Judges show position, verbosity and self-preference biases and are sensitive to vague rubrics. Use explicit binary or categorical criteria, pin the judge model and prompt, and measure agreement against human expert labels before relying on the scores.

How do you evaluate AI agents?

Measure task success against the real end state, tool-call accuracy including arguments and order, unsafe-action rate, and efficiency in steps, tokens and time. Review full trajectories for a sample of tasks, because an agent can reach a correct final answer through an unsafe path.

What is the difference between offline and online evaluation?

Offline evaluation runs a fixed test set before release to catch regressions. Online evaluation monitors live traffic after launch using tracing, user feedback, sampled automated scoring and human review, and it supplies new test cases for the offline set.

Which tools are used for LLM evaluation?

Common choices include Ragas and DeepEval for metrics, promptfoo for prompt comparison and red-teaming, LangSmith, Langfuse and Arize Phoenix for tracing and evaluation, and OpenTelemetry for vendor-neutral tracing into existing observability systems.

Evaluation is the skill that separates engineers who can build an AI demo from engineers who can keep an AI system trustworthy in production. Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad takes you from LLM fundamentals through RAG, agents, evaluation and observability with hands-on labs, in our Ameerpet classroom beside Ameerpet Metro or live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us