Synthetic data for LLM evaluation means using a language model (plus templates and code) to generate test questions, conversations, attacks and database records, so you can test an AI system before real user traffic exists or where real data cannot be used. Done well, it gives you fast coverage of topics, edge cases and adversarial inputs; done carelessly, it gives you a large, easy, self-confirming test set that makes a weak system look strong. This guide covers how to generate each kind of synthetic test data, the quality controls that make it trustworthy, and how to mix it with real logged cases over time.
If you are new to evaluation itself, start with our LLM evaluation guide (test sets, judges, CI gates) and the RAG evaluation metrics breakdown. Both mention synthetic cases briefly; this article is the deep dive on producing them.
Why generate synthetic test data at all
Three problems push teams towards synthetic data:
- Cold start. A new RAG assistant or agent has no production logs. Subject-matter experts can write a few dozen golden questions, but they rarely have time to write hundreds. A generator drafts candidates from your documents; experts review instead of author, which is much faster.
- Coverage of edge cases. Real traffic is dominated by common questions. The rare ones (a policy exception, a two-document comparison, a question the system must decline) are where failures hide, and generation lets you fill those categories deliberately.
- Privacy. In banking, healthcare and insurance, copying real customer conversations into a test repository is often not allowed. Synthetic customers, accounts and chats let developers and CI work without personal data, with caveats covered below.
Synthetic data cannot tell you how real users phrase things or how often each case occurs. It is a bootstrap and coverage tool, not a replacement for real cases.
The end-to-end generation workflow
Every type of synthetic test data follows the same loop.
[Coverage plan: topics x types x risk]
|
v
[Seed inputs: docs, schemas, personas]
|
v
[Generator LLM + templates] --> raw cases
|
v
[Auto filters: dedup, trivial, leakage,
answerable-without-retrieval]
|
v
[Human review: sample + all high-risk]
|
v
[Versioned test set v1.x, tagged]
|
v
[Run evals] --> failures --> new seeds
^ |
+---- real logs ----+
The coverage plan is the step most teams skip. Before generating anything, write a small matrix: the topics your system must handle, the question types (single fact, multi-hop, comparison, numeric, should-decline, adversarial) and a risk level for each cell. Generation then fills cells instead of producing whatever the model finds easiest.
Generating RAG question-answer pairs from documents
The standard approach for a RAG system is to sample chunks from the corpus, ask a generator model to write a question that the chunk answers, plus a reference answer and the IDs of the supporting chunks. Tools such as Ragas include generators for this, including multi-hop variants; pin the version you use.
Practical rules that make these pairs useful:
- Force the generator away from the source wording. Generated questions tend to reuse exact phrases from the chunk, which makes keyword and vector retrieval look far better than it will be with real users. Instruct the generator to write as a customer who has not read the document.
- Generate by type, not just by chunk. Ask explicitly for comparisons across two documents, questions needing a number from a table, questions whose answer changed between policy versions, and questions the corpus cannot answer.
- Keep the provenance. Store the source chunk IDs and document version with each case. When a document changes, you can regenerate stale cases.
- Write the reference answer as claims. A short list of required facts is easier to review and score than a polished paragraph.
Generating adversarial and safety cases
Given categories and a few examples, a generator produces attack variations well. Useful categories:
- Prompt injection in the user message ("ignore your instructions and...") and, more importantly, indirect injection hidden inside a retrieved document, email or ticket.
- Requests for data the user should not see: another customer's account, internal-only documents, system prompts.
- Out-of-scope requests it must decline, and borderline ones it should not over-refuse.
- Requests that push an agent towards an unsafe tool action, such as closing an account or approving a refund without confirmation.
Hosted models may refuse some attack content, so keep a seed list written by your security team, with the model only paraphrasing. Always pair each attack set with benign look-alikes: a guardrail that blocks every message mentioning "password" passes the attack set and fails real users. See AI guardrails for the controls these cases test.
Persona-based simulated users for conversations and agents
Single questions do not test multi-turn assistants or agents. For those, teams use a user simulator: a second LLM given a persona and a goal, which converses with your system until the goal is met, it gives up, or a turn limit is reached.
A persona specifies the goal, what the user knows, their style (terse, rambling, frustrated) and hidden facts revealed only when asked, for example "your card was blocked after a foreign transaction, but you only mention the country if asked". Hidden facts test whether the agent asks clarifying questions instead of guessing.
Score these runs on the end state and the trajectory: did the right tool calls happen with the right arguments, did the agent confirm before acting, did it hand over to a human when it should? Our AI agent evaluation guide covers trajectory metrics. Watch for over-cooperative simulators; real users contradict themselves and change goals mid-conversation, so add personas who do that.
Paraphrase variation and multilingual or code-mixed variants
A robust system should give the same answer to the same intent phrased differently. Generate several paraphrases of each reviewed seed: formal, casual, with typos, very short, buried in a complaint. Measure answer consistency across the group, not just accuracy on the original.
For Indian deployments this matters even more. Staff and customers of a Hyderabad bank or a GCC helpdesk write in English, Hindi, Telugu and often code-mixed text such as Hinglish or Telugu written in Roman script ("na debit card block ayyindi, ela unblock cheyali?"). Generators can produce these, but quality varies by language and script, so a native speaker must review a sample.
Structured-data synthesis for text-to-SQL tests
Text-to-SQL systems need two kinds of synthetic data: the questions and the database they run against. Generate a realistic schema-compatible dataset (customers, accounts, transactions) with code and a seeded random generator, so every run is reproducible, and plant the awkward cases deliberately: NULLs, duplicate names, time zones, cancelled orders, accounts with no transactions, values that look numeric but are stored as text.
Then generate natural-language questions with a gold SQL query for each, and compare result sets rather than SQL strings, because many different queries are correct. Validate every gold query by executing it; a gold query that errors or returns an empty set by accident is a broken test. Include refusal cases (writes, restricted tables) and ambiguous questions that need clarification. Our text-to-SQL agent project shows where these tests fit in the build.
Quality control: the part that decides whether the set is any good
Raw generated cases are drafts. These filters, roughly in order, turn them into a test set:
| Check | What it catches | How |
|---|---|---|
| Deduplication | Near-identical questions that inflate counts and over-weight one topic | Embedding similarity with a threshold, then keep one per cluster |
| Trivial-question filter | Questions that copy a sentence from the chunk and swap one word | Lexical overlap between question and source chunk; flag high overlap |
| Answerable without retrieval | Questions the model answers from general knowledge, so they do not test RAG | Ask a model with no context; if it gets the reference right, drop or rewrite |
| Groundedness of reference | Reference answers the generator invented | Check every reference claim against the cited chunks |
| Category balance | Sets dominated by easy single-fact lookups | Count cases per cell of the coverage plan |
| Human review | Wrong, ambiguous or unrealistic cases the filters miss | Expert review of a random sample plus every high-risk case |
Track the reviewer rejection rate per batch; if it is high, fix the generator prompt and regenerate rather than hand-patching. For high-risk categories (regulatory answers, money movement, medical guidance), review every case.
Avoid leakage between generator and judge
The subtle failure is circularity. If the same model family writes the questions, writes the reference answers, powers the application and acts as the judge, they share blind spots. Scores look excellent and mean little.
Mitigations: use a different model (ideally a different provider) for generation than for the application under test; use a third configuration for judging, calibrated against human labels; keep a human-written slice no model touched; and keep a held-out set you never tune against.
Mixing synthetic data with real logged cases
The synthetic share of your test set should shrink as real traffic arrives:
- Pre-launch: expert-written golden cases plus reviewed synthetic cases for coverage and adversarial testing.
- Pilot: add real questions from pilot users (with consent and redaction), especially the ones that failed or got thumbs-down.
- Production: sample traced conversations regularly, label failures, and add them. Use real failures as seeds for new synthetic variants, so one real bug becomes a family of regression cases.
Always report metrics separately by source (synthetic, expert-written, real logged). A system that scores much higher on synthetic than real cases tells you the generator is too easy.
Keeping synthetic sets versioned
A test set that changes silently makes every comparison meaningless. Store test sets as files in Git (JSONL works well) or in your evaluation platform's dataset feature, with a version number and a changelog. Each case should carry an ID, source (synthetic, expert, real), generator model and prompt version, source document IDs and versions, category tags, risk level, review status and reviewer.
When documents change, regenerate affected cases and bump the version; compare releases only on the same version. LangSmith and Langfuse both support datasets and experiment comparison. This fits the CI pattern in CI/CD for AI applications.
Privacy: synthetic is not automatically anonymous
A common and risky assumption is that anything a model generated is safe to share. If you prompt a generator with real customer records, real chat transcripts or real tickets and ask for "similar" examples, the output can reproduce names, account numbers, addresses or rare combinations of attributes that identify a person.
Safer practice:
- Generate from schemas, policies and descriptions, not from real records, wherever possible.
- If you must seed from real data, redact or tokenise personal fields first, and run a PII scan over the output.
- Check generated records for near-matches against the source data, especially rare combinations (a small town plus an unusual occupation plus an age).
- Keep derived synthetic sets under the source data's access rules until checked.
- Sending real records to a hosted generator is itself a data transfer that compliance must approve.
Illustrative example: a bank policy assistant test set
Consider a bank building an internal assistant that answers staff questions about retail banking policies: account opening, KYC, card disputes, fee waivers and loan prepayment. There are no logs, and real customer conversations cannot be used in development.
The team starts with a coverage plan: each policy area crossed with question types (single fact, multi-document, numeric from a fee table, changed-in-latest-circular, should-decline, adversarial) and a risk level. Fee and KYC rules are marked high risk.
Generation runs in four passes: Q&A pairs from policy chunks using a different provider from the assistant, with claim-list references and chunk IDs; paraphrases including branch shorthand, Hinglish and Telugu-English; adversarial cases from a security-team seed list (a named customer's details, injected text in a pasted email, pressure to invent an exception); and persona conversations such as a new employee unsure of terminology or a manager who keeps asking for exceptions.
Filters drop duplicates and generic "what is KYC" questions answerable without retrieval. Compliance officers review every fee and KYC case and a sample of the rest; the first batch copies the circulars too closely, so the prompt is changed and the batch regenerated. The approved set is tagged v1.0 with its document versions.
When a new circular changes a fee, the cases citing the old document version are regenerated and the set becomes v1.1. After the pilot, real staff questions are added and reported separately. The same pattern underpins our banking AI assistant project.
If you want hands-on practice building RAG systems, agents and the evaluation pipelines around them, Cloudsoft's AI, GenAI and Agentic AI course covers generation, evaluation with Ragas and tracing with LangSmith and Langfuse in labs.
Common mistakes with synthetic test data
- Measuring volume instead of coverage. Many unreviewed lookups are worth less than fewer reviewed cases across the plan.
- Letting the generator copy the source. High lexical overlap makes retrieval look solved.
- Using one model for everything. Generator, application and judge sharing biases produce flattering, circular scores.
- No should-decline cases. A set where every question is answerable never tests hallucination on missing information.
- Never retiring synthetic cases. As real logged cases accumulate, stale or unrealistic synthetic ones should be removed or down-weighted.
QA engineers will recognise much of this from test data management in conventional testing; our sibling guide on AI in software testing for QA and SDET covers the wider testing picture. Taking these practices into a customer's real environment, with their documents, data rules and approvers, is a core part of what Forward Deployed Engineers do.
Frequently asked questions
What is synthetic data for LLM evaluation?
It is test data, such as questions with reference answers, simulated conversations, adversarial prompts or database records, generated by a language model and code rather than collected from real users. It helps before real traffic exists, for rare edge cases and where personal data cannot be used.
Can I evaluate a RAG system using only synthetic questions?
Only as a starting point. Synthetic questions help before launch, but they tend to be easier and closer to the document wording than real questions. Add expert-written and real logged cases soon and report metrics by source.
How do I generate test questions with an LLM?
Write a coverage plan of topics, question types and risk levels, then prompt a generator model with sampled document chunks or schemas and ask for a question, a reference answer as a list of claims and the supporting source IDs. Ask explicitly for harder types, then filter and review.
How many synthetic test cases do I need?
There is no fixed number. Aim for enough reviewed cases in each cell of your coverage plan to see a regression in that category, prioritising high-risk areas. A smaller, reviewed, balanced set is more useful than a large unreviewed one.
Should the same model generate the test data and judge the answers?
Preferably not. When the generator, the application and the judge share a model family, they share blind spots and preferences, which inflates scores. Use different models and calibrate the judge against human labels.
Is synthetic data automatically safe from a privacy point of view?
No. If it was generated from real records or transcripts, it can reproduce personal details or identifying combinations of attributes. Generate from schemas where possible, redact inputs and scan outputs.
How do I know whether a synthetic question is too easy?
Check its lexical overlap with the source chunk and ask a model to answer it without any retrieved context. If overlap is high or the model answers correctly without context, the question does not meaningfully test retrieval and should be rewritten or dropped.
How should synthetic test sets be versioned?
Store them as files or platform datasets with a version number and changelog, and record for each case its source, generator model and prompt version, source document versions, category tags and review status.
To build evaluation skills like these with guided labs in our Ameerpet classroom or live online, explore Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad, or call +91 96660 19191 to book a free demo.



