AI system design interview questions test something different from classic system design. You are still expected to reason about load balancers, queues and databases, but the interviewer mainly wants to see whether you can turn a vague GenAI idea into a system with clear requirements, a token and cost budget, a retrieval or agent architecture, an evaluation plan, security boundaries and a believable answer to "what happens when the model is wrong?" This guide covers 50 high-value questions: 15 on the design method itself, then 35 full designs, from enterprise document Q&A and text-to-SQL to an LLM gateway, an agent with approvals and an offline batch-enrichment pipeline.
How to use this guide
- Junior and mid-level loops (AI engineer, GenAI developer) usually pick one design, such as a RAG assistant, and check that you can clarify requirements, sketch components and name an evaluation metric.
- Senior loops (senior AI engineer, AI architect, Forward Deployed Engineer) expect estimation, model routing, failure modes, tenant and permission boundaries, and cost levers, all inside 45 to 60 minutes.
- Customer-facing loops add a discovery layer: the interviewer plays a stakeholder and rewards you for asking about the business outcome before drawing boxes.
Read Q1 to Q15 first. They give you the method that every later design reuses. For each design, practise saying the clarifying questions aloud, drawing the diagram from memory and defending one trade-off in depth. All numbers in this guide are illustrative placeholders for practising estimation, not benchmarks or vendor prices.
Contents
- The design method: requirements, estimation and models (Q1βQ10)
- Evaluation, failure modes and security boundaries (Q11βQ15)
- Knowledge, search and document designs (Q16βQ23)
- Agent and automation designs (Q24βQ31)
- Developer and data designs (Q32βQ36)
- Conversational and multimodal designs (Q37βQ42)
- Platform and risk designs (Q43βQ50)
- Key takeaways
- Interview preparation checklist
- FAQ
The design method: requirements, estimation and models
1. How should you structure a 45β60 minute AI system design interview?
Answer: Use a fixed sequence so you never forget the parts interviewers score: clarify the problem and success metric, estimate volume, tokens, latency and cost, draw a baseline architecture, deep-dive the riskiest component, then cover evaluation, failure modes, security and cost optimisation. Keep the first five to eight minutes for questions. Candidates who start drawing a vector database in minute one usually design the wrong system.
clarify (users, task, data, risk, metric) | estimate (QPS, tokens, latency, cost) | baseline design (simplest thing that works) | deep dive (retrieval / agent / data path) | evaluate -> failure modes -> security -> cost | rollout plan + what I would do next
Interview tip: Say the sequence out loud at the start. It shows structure and lets the interviewer redirect you early if they care more about one part.
2. Which clarifying questions matter most for an LLM-based system?
Answer: Group them so you cover the ground quickly:
- Users and task: who uses it, what decision or action it supports, answer versus action.
- Data: sources, formats, freshness, volume, who owns them, and which permissions apply.
- Risk: what a wrong answer costs, whether a human reviews output, regulatory constraints such as data residency or India's DPDP Act.
- Scale and experience: users, peak concurrency, languages, channels, latency expectations.
- Success: the business metric (handle time, deflection, analyst hours saved) and the budget per request or per month.
The single most revealing question is "what happens today without AI?" It gives you the baseline to beat and usually exposes the real data sources.
3. How do you decide whether a problem needs an LLM at all?
Answer: Use an LLM when the input is unstructured language, the output needs generation or flexible interpretation, and some error is tolerable or reviewable. If the logic is deterministic (tax rules, eligibility checks), use code. If you have plenty of labelled data and a fixed label set, a classical classifier is cheaper and more predictable. If users just need to find a document, good search may be enough. Senior answers often combine them: rules for hard constraints, a classifier for routing, an LLM only for the language-heavy step. Saying "this part should not be an LLM" is a strong signal in an interview, not a weak one.
4. How do you choose between prompting, RAG, fine-tuning and an agent?
Answer: Choose by what is missing: knowledge, behaviour or action.
| Approach | Use when | Main cost or risk |
|---|---|---|
| Prompting only | The model already knows enough; task is format or style | Prompt fragility; no private knowledge |
| RAG | Answers depend on private, changing documents; citations needed | Retrieval quality, permissions, ingestion pipeline |
| Fine-tuning | Consistent format, tone or domain behaviour at high volume; smaller model needed | Training data, re-training on model changes, no fresh facts |
| Agent with tools | The task needs multiple steps or actions in other systems | Loops, unsafe actions, harder evaluation |
These combine: many production systems are a RAG workflow with one or two tool calls, not a free-roaming agent.
5. How do you estimate tokens per request?
Answer: Add up each part of the prompt and the expected output, then multiply by the number of model calls per user request. An illustrative RAG request:
- System prompt and instructions: about 800 tokens
- Retrieved context: 5 chunks of about 400 tokens = 2,000 tokens
- Conversation history (summarised): about 600 tokens
- User question: about 50 tokens
- Output: about 300 tokens
That is roughly 3,450 input and 300 output tokens for one call. If the design also uses a query-rewrite call and a guardrail check, count those too; agents may make five to fifteen calls per task, each carrying the growing transcript. Remember that token counts differ by tokenizer and language: Indian-language text often uses more tokens per word than English, which matters for a multilingual design. See tokens and context windows explained.
6. How do you estimate monthly cost for an LLM feature?
Answer: Requests per month multiplied by tokens per request, priced separately for input and output, plus retrieval, hosting and observability. An illustrative internal assistant:
users 20,000 employees usage 4 questions/day x 22 days requests/month 20,000 x 4 x 22 = 1.76M input tokens 1.76M x 3,450 = ~6.1B output tokens 1.76M x 300 = ~0.53B model cost (6.1B/1M x Pin) + (0.53B/1M x Pout)
Pin and Pout are the provider's current per-million-token prices, which you look up rather than quote from memory. Then state the levers: shorter system prompts, fewer and smaller chunks, prompt caching for the static prefix, routing simple questions to a smaller model, a semantic cache for repeated questions and batch APIs for offline work. Interviewers care far more about the formula and the levers than about the final rupee figure.
7. How do you build a latency budget for a GenAI request?
Answer: Split end-to-end latency into stages and give each a target, then decide what the user perceives. An illustrative chat budget: authentication and policy checks around 20β50 ms, query rewrite 200β400 ms, hybrid retrieval 100β300 ms, reranking 100β200 ms, then time to first token from the model and the token generation rate. With streaming, the user perceives time to first token, so that becomes the key target, while total generation time matters for tool-calling agents and APIs that need the full response. Levers include running retrieval branches in parallel, skipping the rewrite for simple queries, smaller models for intermediate steps, prompt caching and keeping prompts short. For a full treatment, read LLM latency optimisation.
Interview tip: Ask whether the experience is interactive chat, voice (much tighter) or a background job (latency barely matters, cost does).
8. How do you choose models in a system design?
Answer: Choose per step, not per product. Classification, routing and extraction often run well on small, fast models; complex reasoning or long synthesis may need a larger or reasoning model; embeddings and rerankers are separate choices. Selection criteria: quality on your own evaluation set, latency, cost, context window, structured output and tool-calling reliability, language coverage, data residency and contractual terms, and availability in the cloud the customer already uses (Amazon Bedrock, Azure OpenAI, Gemini on Google Cloud, or self-hosted open-weight models). Put models behind a gateway or abstraction so you can swap them, because model versions get deprecated. Do not hardcode a model name as the answer; describe the evaluation that picks it. See how to choose an LLM for the enterprise.
9. Which non-functional requirements are specific to AI systems?
Answer: On top of availability, latency and security, AI systems add:
- Quality targets: groundedness, answer correctness, refusal behaviour, measured on an agreed evaluation set.
- Cost per task and token budgets per tenant or user.
- Reproducibility: prompt, model and index versions logged so an answer can be explained later.
- Data residency and retention for prompts, outputs and logs.
- Provider dependency: quotas, rate limits, deprecation of model versions and a fallback plan.
- Auditability: who asked what, which sources were used, which tool actions were taken and who approved them.
10. How do you design around provider rate limits and quotas?
Answer: Treat model capacity like any scarce dependency. Know the tokens-per-minute and requests-per-minute limits for each model and region, and size peak load against them. Then add a gateway that enforces per-tenant quotas, retries with exponential backoff and jitter on throttling errors, and routes to a secondary region or an approved alternative model when a limit is hit. Separate interactive traffic from batch traffic with different queues and priorities so a nightly job cannot starve live users. For predictable high volume, evaluate reserved or provisioned capacity options from your provider and compare them with on-demand pricing; check current documentation for what each cloud offers.
Evaluation, failure modes and security boundaries
11. What does an evaluation plan look like in a design interview?
Answer: Four layers. First, an offline golden dataset of real or realistic questions with expected sources and reference answers, built with subject-matter experts. Second, component metrics: retrieval recall and precision, faithfulness, answer relevance, tool-call correctness, format validity. Third, LLM-as-judge scoring calibrated against human labels on a sample, so you know when the judge disagrees with people. Fourth, online signals after launch: user feedback, escalation or override rates, task completion and cost per task. Tie it together with release gates: no prompt, model or index change ships unless the regression suite passes. Tools such as Ragas, LangSmith and Langfuse help run and track these. Go deeper with LLM evaluation and the LLM evaluation interview questions.
12. What failure modes should every AI design address?
Answer: Name them and give each a detection and a mitigation.
| Failure | Detection | Mitigation |
|---|---|---|
| Hallucinated or unsupported answer | Faithfulness checks, citation validation | Grounded prompts, "I don't know" path, human review for high risk |
| Retrieval miss | Retrieval recall on eval set, low-score alerts | Hybrid search, reranking, better chunking |
| Prompt injection via documents or tools | Red-team tests, output policy checks | Treat content as data, least-privilege tools, approvals |
| Agent loops or runaway cost | Step and token counters | Step limits, budgets, timeouts |
| Provider outage or throttling | Error-rate alerts | Fallback model or region, degraded mode |
| Stale or wrong index | Freshness metrics per source | Incremental ingestion, source-of-truth checks |
| Quality drift after a change | Regression suite, online metrics | Versioned prompts, canary rollout, rollback |
13. Where do you draw security boundaries in an AI architecture?
Answer: At every place where trust changes. The model itself is not a security boundary: anything in its context can influence its output, so enforcement must happen in code around it.
user --(SSO token)--> app / orchestrator
| B1: authn, authz, tenant
+------------+-------------+
| |
retrieval (ACL filter) tools (scoped creds)
| B2: content is | B3: actions
| untrusted data | need policy
+------------+-------------+
|
gateway (redact, log)
| B4: data leaves org?
model provider
B1 ties every request to a real identity and tenant. B2 treats retrieved documents and tool results as untrusted input that may contain injected instructions. B3 checks each tool call against policy and the user's own permissions, with approvals for risky actions. B4 controls what data leaves the organisation and how the provider may retain it. See AI guardrails for the control catalogue.
14. How do you design observability for an LLM application?
Answer: Trace every request end to end, with a span for each retrieval, rerank, model call, guardrail and tool call. Record the prompt template version, model identifier, token counts, latency, cost, retrieved document IDs and the final outcome. OpenTelemetry gives you a standard way to propagate traces across services, and LLM-focused tools such as Langfuse or LangSmith add prompt and evaluation views. Two design points matter in interviews: decide what you must not log in clear text (PII, secrets), with redaction before storage and retention rules; and connect traces to evaluation, so a poor production answer can become a new test case in one click.
15. How do you close an AI design answer and propose a rollout?
Answer: Summarise in three sentences: the requirement you optimised for, the architecture, and the main risk with its mitigation. Then propose a phased rollout: offline evaluation with a golden set, a shadow or internal pilot with a small user group, a limited release with human review and tight budgets, then wider rollout once metrics hold. Finish with what you would do with more time, such as a fine-tuned small model for a high-volume step, or a richer permission model.
Knowledge, search and document designs
Every design below follows the same shape: clarifying questions, an architecture sketch, key components, trade-offs, scaling, evaluation, and security and cost. In the interview, you will not have time for all of it at full depth; pick the one or two parts where the real risk lies.
16. Design an enterprise document Q&A assistant for 20,000 employees.
Answer: A permission-aware RAG system over SharePoint, Confluence and policy PDFs, available in Teams and a web app, answering with citations and refusing when the answer is not in the sources.
Clarifying questions: Which sources and how many documents? How often do they change? Do document permissions vary by group? Which languages? Is a wrong answer an inconvenience or a compliance problem (HR, legal, safety)?
SharePoint / Confluence / PDFs
| connectors + ACL sync
parse -> chunk -> embed -> index (vectors + ACLs)
|
user (SSO) -> API -> rewrite -> hybrid search
| filter: user groups
rerank -> top k
|
LLM (grounded, cited)
|
answer + citations + feedback
Key components: source connectors that sync content and ACLs, a parser that keeps headings and tables, a hybrid (keyword plus vector) index such as PostgreSQL with pgvector or a managed search service, a reranker, a grounded prompt with a mandatory citation format, and a feedback loop.
Trade-offs: pre-filtering by ACL at query time is safer than post-filtering; larger chunks give context but dilute retrieval; a query-rewrite step improves follow-up questions but adds latency.
Scaling: incremental ingestion driven by change events, not nightly full re-indexing; cache embeddings by content hash; scale the API statelessly.
Evaluation: golden questions per department, retrieval recall, faithfulness, citation correctness and a "must refuse" set of questions a given role cannot see.
Security and cost: ACL filters enforced in the retrieval query, PII redaction in logs, prompt caching for the static prefix. The full build is in our RAG knowledge assistant project, and retrieval-depth questions are in the RAG interview questions guide.
17. Design a multi-tenant RAG SaaS product.
Answer: A platform where many customer organisations upload their own documents and get an assistant, with strict data isolation, per-tenant configuration and usage-based billing.
Clarifying questions: How many tenants and how unequal are they? Do enterprise tenants demand dedicated infrastructure or their own keys? Can tenants bring their own model contracts? What residency regions are required?
tenant users -> edge (auth, tenant_id from token)
|
control plane: tenants, plans,
quotas, model config, keys
|
data plane (shared pool)
ingest workers query service
| |
index: tenant_id partition / per-tenant
| |
LLM gateway (per-tenant quota, cost tags)
Key components: a control plane for tenant metadata and plans, a data plane that derives tenant_id only from the verified token, tenant-partitioned storage, per-tenant encryption keys for premium tiers, and a gateway that meters tokens per tenant.
Trade-offs: pooled indexes with tenant filters are cheap but need rigorous testing; a dedicated index or database per tenant is easier to prove isolated but costs more. A common answer is pooled for small tenants, siloed for large or regulated ones.
Scaling: noisy-neighbour protection through per-tenant rate limits and separate ingestion queues; shard by tenant.
Evaluation: per-tenant quality dashboards plus cross-tenant leakage tests in CI (tenant A's queries must never retrieve tenant B's chunks).
Security and cost: row-level security or equivalent as defence in depth, per-tenant cost attribution for billing. More patterns in multi-tenant AI SaaS architecture.
18. Design a document ingestion pipeline that processes millions of files.
Answer: An event-driven, idempotent pipeline that discovers files, parses them into structured text, chunks, embeds and indexes them, and can reprocess everything when the parser or embedding model changes.
Clarifying questions: File types (scanned PDFs, spreadsheets, emails)? Initial backfill size versus daily change rate? Freshness target? Is OCR needed? Which metadata and permissions must travel with each chunk?
sources -> change events -> queue
|
workers: fetch -> detect type
|
parse (text / layout / OCR / tables)
|
chunk -> enrich metadata -> embed
|
index upsert (doc_id, version, ACL)
|
dead-letter queue + ingestion metrics
Key components: a queue with retries and a dead-letter queue, type-specific parsers, content hashing to skip unchanged files, a version field per document so deletes and updates propagate, and embedding calls batched for throughput.
Trade-offs: layout-aware parsing and OCR are slower and costlier but essential for tables and scans; a separate backfill lane keeps the initial load from delaying fresh updates.
Scaling: horizontal workers sized against embedding API limits; reprocessing via a new index version and an alias swap, so a re-embed never takes search offline.
Evaluation: parse-quality sampling, chunk statistics, retrieval tests after each pipeline change.
Security and cost: malware scanning on upload, ACLs copied at ingestion, deletions honoured quickly for data-protection reasons. Parsing details are in document parsing for RAG.
19. Design an AI contract review assistant for a legal team.
Answer: A system that extracts key clauses, compares them with the company's playbook, flags deviations with risk levels and suggests fallback language, while a lawyer makes every decision.
Clarifying questions: Contract types (NDA, MSA, vendor)? Is there a written playbook with approved positions? Which output does legal want: a redline, a risk memo or a checklist? Confidentiality constraints on sending contracts to a model provider?
contract upload -> parse (clauses, numbering)
|
clause classifier -> clause records
|
playbook retrieval (approved + fallback)
|
LLM compare -> {clause, deviation, risk, cite}
|
schema validation -> reviewer UI
|
lawyer accepts / edits -> audit log
Key components: clause segmentation that preserves numbering, a playbook stored as structured positions rather than prose, structured outputs with a JSON schema, and a reviewer interface showing the source clause beside each finding.
Trade-offs: clause-by-clause comparison is more accurate and auditable than one whole-contract prompt but costs more calls; a fine-tuned clause classifier can cut cost at volume.
Scaling: asynchronous processing per contract with progress updates; contracts are long, so parallelise per clause.
Evaluation: lawyer-labelled contracts measuring clause extraction recall and deviation precision; missed high-risk clauses are the key metric.
Security and cost: matter-level access control, no retention by the provider, full audit of edits. See the contract review AI project.
20. Design semantic search for an e-commerce catalogue.
Answer: A search that understands queries such as "cotton kurta for office under 1500" by combining keyword relevance, vector similarity and structured filters, with the LLM used mainly offline and for query understanding.
Clarifying questions: Catalogue size and update rate? Languages and transliterated queries (Hinglish)? Latency target for search (often well under a second)? Is the goal conversion, zero-result reduction or both?
query -> query understanding (small model / rules)
{terms, category, price, attributes}
|
+--------------+--------------+
keyword (BM25) vector (ANN)
+--------------+--------------+
| filters: price, stock
fusion -> learning-to-rank
|
results (+ cached facets)
Key components: a fast query parser that extracts filters, hybrid retrieval with reciprocal rank fusion, product embeddings generated from titles, attributes and enriched descriptions, and a ranking stage using behavioural signals.
Trade-offs: calling a large LLM per search is usually too slow and costly; use it offline to enrich products and for rare, long queries. Vector search alone misses exact SKUs and brand names, which is why hybrid wins.
Scaling: cache popular queries, precompute embeddings, scale the ANN index with replicas.
Evaluation: offline relevance judgements (NDCG), zero-result rate, then A/B tests on click-through and conversion.
Security and cost: few PII concerns, but guard against prompt injection in seller-supplied text used for enrichment.
21. Design a recommendation system that explains its recommendations with an LLM.
Answer: Keep the recommender a conventional ranking model; use the LLM only to turn the real reasons behind each recommendation into a short, accurate explanation.
Clarifying questions: Domain (products, courses, funds)? Must explanations be generated in real time? Are there regulatory rules on advice, as in financial products?
user events -> feature store
|
candidate gen -> ranking model
|
top N + reason codes (e.g. "bought X",
"similar users", "price drop")
|
LLM: explain ONLY from reason codes
|
template fallback if LLM fails / slow
Key components: reason codes emitted by the ranker, a constrained prompt that may only verbalise those codes, a template fallback, and caching of explanations per item and reason combination.
Trade-offs: free-form explanations read well but may invent reasons, which damages trust; constrained explanations are safer. Pre-generating explanations offline is cheaper than per-request generation.
Scaling: because combinations repeat, a cache absorbs most traffic.
Evaluation: faithfulness of explanations to reason codes, user trust surveys, and whether explanations change click-through.
Security and cost: avoid exposing sensitive inferences ("because you searched for a medical condition"); filter reason codes by sensitivity.
22. Design a regulatory change tracker for a bank's compliance team.
Answer: Consider a bank that must review every new regulator circular and map it to internal policies. The system ingests circulars, summarises obligations, finds affected internal policies and creates review tasks for compliance owners.
Clarifying questions: Which regulators and formats? How quickly must a circular be triaged? Who owns each policy? Does compliance want obligation extraction, a gap analysis or both?
regulator feeds -> fetch + parse circulars
|
LLM: extract obligations (schema, cited)
|
retrieve related internal policies (RAG)
|
LLM: impact note per policy (cited both)
|
task in GRC tool -> owner review -> sign-off
Key components: an obligation schema (who, what, deadline, source paragraph), a policy index with owners as metadata, and integration with the governance or ticketing tool for tasks.
Trade-offs: high recall matters more than precision; it is better to send an owner an irrelevant task than to miss an obligation. A knowledge graph of policies, controls and obligations helps at scale but takes effort to maintain.
Scaling: volume is low; the design priority is reliability and traceability, not throughput.
Evaluation: compliance-labelled circulars measuring obligation recall and mapping accuracy.
Security and cost: every note cites both the circular and the policy paragraph; a human signs off. Domain controls are covered in generative AI in banking.
23. Design an insurance claims intake system that extracts data from documents.
Answer: A pipeline that receives claim forms, hospital bills and photos, extracts structured fields, validates them against the policy and routes the claim to straight-through processing or a human adjuster.
Clarifying questions: Document types and quality (phone photos, handwriting)? Which fields drive the decision? Is the AI allowed to approve anything, or only prepare the file? Turnaround target?
claim docs -> classify doc type
|
OCR / multimodal extract -> field JSON
|
validate: schema, policy DB, totals, dates
|
confidence + rules -> route
| |
straight-through adjuster queue
(low value, all valid) (with highlights)
Key components: document classification, a multimodal model or OCR plus LLM for extraction, deterministic validation against the policy system, per-field confidence and an adjuster UI that highlights the source region for each value.
Trade-offs: routing thresholds trade automation rate against error rate; start conservative. Rules, not the model, decide eligibility.
Scaling: queue-based, with spikes after natural disasters; autoscale workers and protect model quotas.
Evaluation: field-level accuracy on a labelled set, routing precision, adjuster correction rate.
Security and cost: health data needs strict access control and retention rules; fraud signals feed a separate model.
Agent and automation designs
24. Design a customer support agent for a telecom or retail company.
Answer: An assistant that answers from the knowledge base, looks up the customer's orders or plan through tools, performs a small set of safe actions and hands off to a human with full context when it should.
Clarifying questions: Channels (web, app, WhatsApp, voice)? Which actions are allowed (track order, raise refund)? Refund limits? Languages? Current escalation flow and target deflection?
customer -> channel adapter -> session + auth
|
intent + risk classifier
| | |
FAQ (RAG) tool calls handoff
(orders, CRM) (agent desk +
| summary)
policy check: action limits
|
response -> CSAT + trace
Key components: customer authentication before account tools, a small toolset with typed schemas, policy checks on actions (refund caps), and a handoff that passes a summary and transcript to the human agent desk.
Trade-offs: a workflow graph with bounded steps is more predictable than an open agent; more autonomy raises deflection but also risk.
Scaling: stateless workers with session state in a store; peak-hour autoscaling.
Evaluation: resolution rate, correct handoff rate, policy-violation rate in simulated conversations, CSAT.
Security and cost: no account data before authentication; refunds above a limit need approval. Full walkthrough: customer support agent project.
25. Design an IT operations incident agent.
Answer: An agent that, when an alert fires, gathers logs, metrics, recent deployments and similar past incidents, proposes a probable cause and remediation, and runs only approved runbook actions.
Clarifying questions: Which monitoring and ticketing tools (ServiceNow, Jira, PagerDuty)? Read-only first or remediation too? Which actions are safe to automate? On-call workflow?
alert -> incident ticket (ServiceNow)
|
triage agent (LangGraph)
| | | |
logs metrics deploys past incidents
(via MCP servers / APIs, read-only)
|
hypothesis + evidence -> ticket note
|
runbook action? -> approval (on-call) -> run
|
verify recovery -> close + postmortem
Key components: read-only diagnostic tools exposed through MCP servers or APIs, a similar-incident index, runbooks as allow-listed parameterised actions, and an approval step in chat or the ticket.
Trade-offs: automatic remediation shortens recovery but risks making incidents worse; start with diagnosis plus one-click approved actions.
Scaling: alert storms need deduplication and correlation before the agent runs, or cost explodes. If you expose tools over MCP, note that the 2026-07-28 specification revision is stateless (no initialize handshake or protocol sessions), which makes MCP servers easier to run behind an ordinary load balancer; many deployed clients still use earlier revisions, so check modelcontextprotocol.io.
Evaluation: replay past incidents and score root-cause accuracy and time to useful hypothesis.
Security and cost: scoped service credentials, no shell access for the model, full audit. See the ServiceNow AI agent project.
26. Design an agent that requires human approval for risky actions.
Answer: Separate proposing from executing. The agent produces a structured proposal; a policy engine decides whether it can run automatically, needs approval or is forbidden; the executor runs only approved, unchanged proposals.
Clarifying questions: Which actions and their blast radius? Who can approve, and within what time? What happens if nobody approves? Audit and segregation-of-duties rules?
agent -> proposal {tool, args, reason, risk}
|
policy engine (code, not LLM)
| | |
auto-run needs approval deny
|
durable wait (checkpoint state)
|
approver sees diff + evidence -> yes / no
|
executor: re-check policy, hash args, run
|
audit: who proposed / approved / ran
Key components: durable workflow state (for example LangGraph checkpoints) so a wait of hours survives restarts, an approval UI showing exactly what will change, argument hashing so the approved action cannot be altered, and expiry for stale approvals.
Trade-offs: too many approvals cause rubber-stamping; tier actions by risk and approve batches where safe.
Scaling: approvals are asynchronous; the agent must resume from saved state, not re-plan.
Evaluation: policy tests that try to sneak risky actions past the engine, approval latency, override rates.
Security and cost: the approver must not be the requester for sensitive actions. Patterns are in human-in-the-loop AI.
27. Design an employee onboarding agent that works across HR, IT and identity systems.
Answer: A workflow agent that, when HR confirms a new joiner, creates accounts, requests a laptop, assigns training and answers the joiner's questions, with each system integrated through a dedicated tool server.
Clarifying questions: Systems involved (HRMS, Microsoft Entra ID, ServiceNow, LMS)? Which steps are fixed versus judgement-based? Who approves access to sensitive groups?
HRMS "joiner confirmed" event
|
onboarding workflow (fixed steps)
| | |
identity IT assets training
MCP server MCP server MCP server
(Entra ID) (ServiceNow) (LMS)
|
LLM only for: role-based access suggestions,
joiner Q&A, exception summaries
|
manager approves access -> provision -> log
Key components: a deterministic workflow for the known steps, one MCP server per system with narrow tools, the LLM used where judgement adds value, and manager approval before group membership is granted.
Trade-offs: letting the LLM orchestrate everything looks impressive but is harder to test; fixed steps plus LLM-assisted exceptions is more robust.
Scaling: low volume, bursty around joining dates; idempotent steps so retries never create duplicate accounts.
Evaluation: end-to-end tests in a sandbox tenant; access-suggestion accuracy reviewed by IT.
Security and cost: each MCP server holds only the credentials it needs; tool descriptions from third-party servers are reviewed, since they enter the model's context. For protocol depth, see the MCP interview questions.
28. Design an AI assistant for SOC alert triage.
Answer: An assistant that enriches each security alert with asset, identity and threat-intelligence context, summarises what happened, suggests a severity and next steps, and leaves containment decisions to analysts.
Clarifying questions: SIEM and EDR tools in use? Alert volume and false-positive rate today? Can the system isolate hosts or disable users, or only recommend?
SIEM alert -> dedupe / correlate
|
enrich: asset DB, identity, intel, history
|
LLM: timeline + severity + next steps (cited)
|
analyst console (accept / adjust)
|
containment action? -> analyst approval -> SOAR
Key components: deterministic enrichment before the LLM, a structured triage output, links to raw evidence for every claim, and execution through existing SOAR playbooks.
Trade-offs: logs are attacker-controlled input, so prompt injection risk is real; keep the model's output advisory.
Scaling: high volume; summarise only correlated incidents, not every raw alert.
Evaluation: agreement with analyst severity, missed true positives, analyst time per incident.
Security and cost: the assistant's own access is privileged and must be monitored. See the SOC AI assistant project.
29. Design an accounts-payable invoice matching agent.
Answer: Consider a retailer receiving thousands of supplier invoices. The system extracts invoice data, performs three-way matching against purchase orders and goods receipts, explains mismatches and routes exceptions to finance staff.
Clarifying questions: ERP system? Invoice formats (e-invoice, PDF, scans)? Tolerances for price and quantity differences? Can matched invoices be posted automatically?
invoice inbox -> extract (fields + line items)
|
ERP lookup: PO, goods receipt (API)
|
matching rules in code (tolerances)
| |
matched -> post mismatch
(within limits) |
LLM: explain mismatch + suggest
vendor email draft
|
finance reviewer decides
Key components: extraction to a strict schema, deterministic matching (never ask the LLM to do arithmetic you can do in code), exception explanations and draft vendor communications.
Trade-offs: auto-posting saves effort but needs strict limits and audit; duplicate-invoice detection must be deterministic.
Scaling: month-end peaks; batch extraction through queues.
Evaluation: field accuracy, match accuracy versus finance decisions, exception handling time.
Security and cost: bank-detail changes on invoices are a classic fraud signal and must always go to a human.
30. Design a personal assistant with long-term memory.
Answer: An assistant that remembers user preferences, ongoing projects and past decisions across sessions, retrieves only relevant memories per turn and lets the user see, edit and delete what it remembers.
Clarifying questions: What should be remembered (preferences, facts, tasks)? How long? Which tools (calendar, email)? Must users control memory explicitly?
turn -> context builder
| short-term: recent turns
| long-term: memory search (user_id)
v
LLM (+ calendar / email tools)
|
memory writer (async): extract candidate
facts -> dedupe / update / expire -> store
|
memory UI: view, edit, delete
Key components: separate short-term and long-term memory, an asynchronous memory writer that extracts durable facts, conflict resolution (a new fact replaces an old one), time-based decay and a user-facing memory view.
Trade-offs: remembering everything bloats context and creates privacy risk; remembering too little feels forgetful. Explicit "remember this" plus conservative automatic extraction is a reasonable middle.
Scaling: memory is per user and partitions naturally.
Evaluation: recall of relevant memories on scripted multi-session tests, wrong-memory rate.
Security and cost: memories are personal data: encryption, deletion on request and no cross-user retrieval. See AI agent memory.
31. Design a research agent that writes an analyst briefing from many sources.
Answer: A planner breaks the question into sub-questions, workers search internal and approved external sources in parallel, and a writer produces a cited briefing that a human analyst reviews.
Clarifying questions: Which sources are allowed? Expected length and turnaround (minutes, not seconds)? Is the output internal or client-facing?
question -> planner (sub-questions, budget)
|
+-------------+-------------+
worker 1 worker 2 worker 3
(internal) (filings) (news, allowed)
+-------------+-------------+
| notes + source IDs
writer -> cited draft
|
checker: every claim has a source?
|
analyst review -> publish
Key components: an explicit budget (steps, tokens, time), parallel workers with isolated context, a notes store, and a citation checker.
Trade-offs: multi-agent designs parallelise well but multiply cost; a single agent with good tools may be enough for narrow questions.
Scaling: run as a background job with progress updates; cap concurrent jobs per user.
Evaluation: analyst-rated usefulness, citation accuracy, coverage of a known answer set.
Security and cost: external web content may carry injected instructions; workers have no write tools. More agent design depth is in the AI agent developer interview questions.
Developer and data designs
32. Design an AI code review assistant for pull requests.
Answer: A service triggered by pull request events that reviews the diff with relevant repository context, posts a small number of high-confidence comments and never blocks merges on its own.
Clarifying questions: Languages and repository size? What should it catch (bugs, security, style, missing tests)? Self-hosted Git or GitHub? Can code leave the network?
PR opened / updated (webhook)
|
fetch diff + changed files + related code
(symbol index, tests, CODEOWNERS, rules)
|
static analysis results (linters, SAST)
|
LLM review -> candidate comments
|
filter: confidence, dedupe, max N
|
post inline comments -> thumbs up / down
Key components: diff-aware context retrieval (callers, callees, tests), existing static analysis fed in as evidence, a comment filter to avoid noise, and feedback captured per comment.
Trade-offs: more comments catch more issues but developers stop reading them; precision matters more than recall here. Whole-repository context is expensive; retrieve only related symbols.
Scaling: queue reviews, cancel stale runs when a new commit arrives, cache repository indexes.
Evaluation: seeded-bug benchmarks from past incidents, comment acceptance rate.
Security and cost: treat PR text and code comments as untrusted (injection), use a read-only token, consider a self-hosted model for sensitive code. See the GitHub AI agent project.
33. Design a text-to-SQL analytics assistant for business users.
Answer: Users ask questions in plain language; the system retrieves the relevant schema and business definitions, generates SQL, validates it, runs it read-only with the user's own data permissions and explains the result.
Clarifying questions: Warehouse and number of tables? Is there a semantic layer or metric definitions? Who may see which rows? Must answers match finance-reported figures?
question -> retrieve: tables, columns, metric
definitions, example queries
|
LLM -> SQL (dialect-specific)
|
validate: parse, allow-list tables, LIMIT,
no DDL / DML, cost estimate
|
execute as user (read-only, row security)
|
result + SQL shown + chart -> feedback
Key components: schema retrieval rather than dumping every table into the prompt, a library of verified example queries, a SQL validator, execution under the user's identity and showing the SQL for transparency.
Trade-offs: generating against a semantic layer (defined metrics) is more accurate than raw tables but needs that layer to exist; a self-correction loop on SQL errors helps but adds latency.
Scaling: query cost limits and timeouts on the warehouse; cache common questions.
Evaluation: execution accuracy on a gold question set (compare result sets, not SQL strings).
Security and cost: never run as a privileged service account. Full build: text-to-SQL agent project.
34. Design a system that generates test cases from requirements.
Answer: A tool that reads user stories and acceptance criteria from Jira, generates structured test cases and, optionally, automation skeletons, which QA engineers review before they enter the test management system.
Clarifying questions: Manual test cases, automated scripts or both? Which frameworks? How consistent are the requirements? Traceability needs for audits?
Jira story + acceptance criteria
|
retrieve: related stories, existing tests,
domain glossary
|
LLM -> test cases (schema: steps, data,
expected, requirement_id)
|
coverage check: every criterion has a test
|
QA review -> test management tool
Key components: a strict test-case schema with traceability to requirement IDs, duplicate detection against existing tests, and a coverage check.
Trade-offs: generated tests are only as good as the requirements; flag ambiguous criteria instead of guessing.
Scaling: sprint-level batches; low real-time pressure.
Evaluation: QA acceptance rate, defects found by generated tests, coverage of acceptance criteria.
Security and cost: no production data in generated test data; use synthetic data.
35. Design an offline batch-enrichment pipeline for a large catalogue.
Answer: Consider a marketplace with millions of product listings that need clean attributes, categories and short descriptions. Run it as a batch pipeline optimised for cost and correctness, not latency.
Clarifying questions: How many items, and how many change daily? Which fields? Acceptable turnaround (hours or days)? How are errors corrected?
catalogue snapshot -> select changed items
|
build requests (template + item JSON)
|
provider batch API / own GPU workers
|
validate: schema, allowed values, length
| |
valid -> stage invalid -> retry
| (other prompt /
sample QA review model) -> DLQ
|
publish to catalogue (versioned)
Key components: change detection so only new or edited items are processed, deterministic request IDs for idempotency, schema and enumeration validation, a sampling-based human QA step, and versioned output so a bad run can be rolled back.
Trade-offs: many providers offer batch APIs at a lower price than real-time calls, in exchange for longer completion windows; check current pricing and limits. A smaller or fine-tuned model may be good enough for attribute extraction at this volume.
Scaling: shard by category, checkpoint progress, respect batch size limits.
Evaluation: field accuracy on a labelled sample per category before full runs.
Security and cost: estimate total tokens before launching; a prompt bug multiplied by millions of items is expensive, so run a pilot slice first.
36. Design a high-volume ticket classification and routing service.
Answer: Classify incoming support or IT tickets into categories, priority and team. At high volume, the strongest design is often a small fine-tuned model or classifier with an LLM fallback for low-confidence cases.
Clarifying questions: Ticket volume and number of categories? Is labelled history available? How costly is misrouting? Latency needs?
ticket -> PII redaction
|
small model / classifier -> label + score
| |
score high -> route score low -> LLM
| (few-shot,
| label defs)
still unsure -> human
|
corrections -> training data -> retrain
Key components: a confidence-based cascade, label definitions maintained by the service desk, and a feedback loop turning reassignments into training data.
Trade-offs: an LLM for every ticket is simple but costly at volume; a cascade needs more MLOps work. Category definitions change, so plan retraining.
Scaling: a small model serves cheaply on CPU or a modest GPU.
Evaluation: per-category precision and recall, reassignment rate.
Security and cost: redact before any external call.
If you want to practise designs like these on realistic enterprise systems (ServiceNow, Entra ID, Bedrock, Kubernetes) with a trainer reviewing your trade-offs, the AI Forward Deployed Engineer course (FDE PRO) builds five enterprise projects and a simulated "GlobalBank" customer engagement over 12 weeks.
Conversational and multimodal designs
37. Design a multilingual voice bot for a bank or utility call centre.
Answer: A real-time pipeline of speech-to-text, a dialogue engine with tools, and text-to-speech, supporting English and Indian languages, with authentication, interruption handling and a warm transfer to human agents.
Clarifying questions: Which languages and code-mixed speech? Telephony platform? Which intents are in scope? How is the caller authenticated? Latency tolerance?
caller -> telephony (SIP) -> media stream
|
streaming STT (language ID) -> text
|
dialogue: intent -> tools (balance, outage)
small fast LLM, short replies
|
streaming TTS (same language) -> caller
|
barge-in detection | transfer + summary
Key components: streaming STT and TTS, voice activity detection for interruptions, a fast model for turn handling, short responses suited to speech, and OTP or voice-based authentication handled outside the LLM.
Trade-offs: a cascaded pipeline (STT, LLM, TTS) is easier to control and audit; speech-to-speech models can feel more natural but are harder to constrain. Latency targets are much tighter than for chat.
Scaling: concurrent call capacity sizes the system, not requests per second.
Evaluation: word error rate per language, task completion, transfer rate, response latency per turn.
Security and cost: no sensitive data read aloud without authentication; call recording consent. See voice AI agents.
38. Design a WhatsApp assistant for an Indian business.
Answer: An assistant on the WhatsApp Business Platform that answers questions, shares order status, books appointments and hands over to staff, while respecting WhatsApp's messaging rules and opt-in requirements.
Clarifying questions: Which business flows? Languages (English, Hindi, Telugu)? Volume? Is there a backend for orders and bookings? Who handles handoffs and when?
WhatsApp user -> Business Platform webhook
|
verify signature -> dedupe by message ID
|
queue -> conversation worker
|
state store (per phone, consent, language)
|
intent -> RAG FAQ | booking tool | handoff
|
reply (text / buttons / template msg)
Key components: webhook signature verification, idempotent processing (webhooks can be redelivered), conversation state per user, interactive buttons for structured choices, and pre-approved templates for messages outside the customer service window.
Trade-offs: buttons and lists reduce ambiguity and cost compared with free text; full LLM conversation is flexible but less predictable.
Scaling: acknowledge webhooks quickly and process asynchronously.
Evaluation: resolution rate, handoff quality, language accuracy.
Security and cost: phone numbers are personal data under the DPDP Act; minimise and retain carefully. Check Meta's current policies and pricing. Full build: WhatsApp AI assistant project.
39. Design a meeting summariser for an enterprise.
Answer: A service that takes meeting recordings or transcripts, produces a summary, decisions and action items with owners, and pushes actions to the task tool after the organiser confirms.
Clarifying questions: Platform (Teams, Zoom)? Is a transcript already available? Meeting length? Who receives the summary? Confidential meetings excluded?
meeting ends -> transcript (+ speaker labels)
|
chunk by time / topic (long meetings)
|
map: per-chunk notes -> reduce: summary,
decisions, actions {owner, due, quote}
|
organiser review -> send + create tasks
Key components: diarisation, map-reduce summarisation for long meetings, action items linked to transcript timestamps, and an organiser review step.
Trade-offs: long-context models can take a full transcript, but map-reduce gives better traceability and handles very long meetings; owner attribution is the hardest part.
Scaling: asynchronous jobs with morning peaks.
Evaluation: human-rated faithfulness, action-item recall versus a manual list.
Security and cost: respect meeting permissions and sensitivity labels; summaries inherit the meeting's access list.
40. Design a clinical documentation assistant for a hospital.
Answer: Consider a hospital where doctors spend significant time writing notes. The assistant drafts a structured note from the consultation transcript and the patient record, and the doctor reviews, edits and signs it.
Clarifying questions: Note format and specialties? Integration with the hospital information system? Consent for recording? Data residency? Is any clinical suggestion allowed, or only documentation?
consult audio (consent) -> STT (medical vocab)
|
patient context: allergies, meds (HIS API)
|
LLM -> draft note (sections, sources)
|
flags: unsupported statements, missing items
|
doctor edits + signs -> HIS (versioned)
Key components: a medical-vocabulary STT, a note template per specialty, statement-to-transcript linking and mandatory clinician sign-off.
Trade-offs: staying strictly in documentation keeps the risk profile lower than diagnostic suggestions.
Scaling: outpatient peaks; near-real-time drafts at the end of a consult.
Evaluation: clinician edit distance, omission and fabrication rates on reviewed samples.
Security and cost: health data is sensitive; in-region processing, strict retention and access logging are table stakes.
41. Design an AI tutor for an edtech platform.
Answer: A tutor that explains concepts from the course material, gives hints rather than answers on graded work, adapts to the learner's level and supports Indian languages.
Clarifying questions: Age group (minors change safety needs)? Subjects? Is homework graded? Languages? Concurrent learners during exam season?
learner -> tutor API (learner profile)
|
retrieve: lesson content, past mistakes
|
policy: hint mode for graded items
|
LLM (pedagogy prompt) -> answer / hint
|
safety filter -> response -> mastery update
Key components: RAG grounded in the official syllabus, a learner model tracking mastery, a hint-only policy for assessments and content safety filters.
Trade-offs: a helpful tutor that simply gives answers hurts learning; a strict one frustrates learners.
Scaling: heavy evening and exam-season peaks; small models for routine explanations.
Evaluation: correctness on syllabus questions, teacher review, learning outcome comparisons.
Security and cost: stronger safety and data rules for minors; per-learner usage caps.
42. Design a CRM sales copilot that drafts follow-ups and updates opportunities.
Answer: A copilot inside the CRM that summarises account history, drafts follow-up emails after calls and proposes field updates that the salesperson confirms.
Clarifying questions: CRM platform? Data sources (emails, call notes)? Can it send emails or only draft? Which fields can it update?
salesperson opens account
|
gather: CRM record, recent emails, call notes
|
LLM -> summary + next steps + email draft
|
proposed field updates (diff view)
|
user confirms -> CRM API (user's own token)
Key components: retrieval scoped to the user's CRM permissions, drafts never sent automatically, and field updates shown as a diff.
Trade-offs: automatic updates reduce data-entry work but corrupt pipeline data if wrong; confirmation keeps quality high.
Scaling: per-seat usage; cache account summaries and refresh on change.
Evaluation: draft acceptance rate, edit distance, field-update accuracy.
Security and cost: no cross-account leakage.
Platform and risk designs
43. Design an LLM gateway for a large organisation.
Answer: A central service that every application uses to reach models. It handles authentication, routing, quotas, redaction, logging, cost attribution, caching and failover, so individual teams do not each rebuild them.
Clarifying questions: How many teams and providers? Streaming needed? Latency overhead allowed? Which policies are central (PII, allowed models)?
apps (API key / workload identity)
|
LLM gateway
authn -> policy (allowed models, PII)
-> quota / rate limit (team, app)
-> cache lookup
-> router (cost, health, region)
| | |
Bedrock Azure OpenAI self-hosted
|
logs, traces, cost tags -> FinOps dashboards
Key components: a unified API, virtual keys per application, policy enforcement, provider adapters, streaming pass-through, retries and fallback, and usage records tagged by team and use case.
Trade-offs: a gateway is a single point of failure and adds a few milliseconds; deploy it redundantly and keep it thin. Normalising every provider feature into one API loses provider-specific capabilities; allow pass-through where needed.
Scaling: stateless, horizontally scaled; quotas in a fast shared store.
Evaluation: gateway overhead, failover drills, policy test suites.
Security and cost: provider keys live only in the gateway. See LLM gateway explained and enterprise AI architecture.
44. Design an evaluation platform for many AI teams.
Answer: A shared service where teams register datasets, define metrics, run evaluations against any version of a prompt, model or pipeline, and gate releases on the results.
Clarifying questions: How many teams and app types (RAG, agents, extraction)? CI integration? Human labelling needed? Data sensitivity of datasets?
datasets (versioned) metrics (code, judges)
\ /
eval run (app version X)
|
runner: call app -> collect traces
|
scorers: exact, schema, retrieval, LLM judge
|
results store -> compare with baseline
|
CI gate pass / fail | human review queue
Key components: versioned datasets, pluggable scorers, judge prompts calibrated against human labels, run comparison views, and a CI integration that blocks a regression.
Trade-offs: LLM judges scale cheaply but carry bias and drift; human review is accurate but slow. Use both, with agreement tracked.
Scaling: parallel runners with provider-quota awareness; cache unchanged results.
Evaluation: of the platform itself, judge-human agreement.
Security and cost: datasets built from production traces need PII scrubbing and access control.
45. Design a content moderation pipeline for user-generated content.
Answer: A layered pipeline: fast classifiers for clear cases, an LLM for nuanced policy judgement on borderline content, and human moderators for appeals and the hardest decisions.
Clarifying questions: Content types (text, images, video)? Languages? Policy categories? Volume and latency (pre-publish or post-publish)? Legal obligations for takedown timelines?
new content -> hash match (known bad)
|
fast classifiers (per category, per lang)
| | |
clear ok borderline clear violation
| | |
publish LLM + policy text block + log
|
low confidence -> human queue
|
decisions + appeals -> retraining data
Key components: hash matching, classifier tiers, policy documents injected into the LLM prompt, a moderator console and an appeals workflow.
Trade-offs: pre-publish moderation is safer but delays posting; thresholds trade over-blocking against under-blocking per category.
Scaling: the cheap tiers absorb most volume, so the LLM sees only a fraction.
Evaluation: per-category precision and recall, per-language parity, appeal overturn rate.
Security and cost: moderator wellbeing and data handling; adversarial users probe the system, so red-team regularly.
46. Design a shared guardrails service.
Answer: A service that applications call before and after model calls for PII detection and redaction, prompt-injection detection, topic restrictions and output policy checks, so controls are consistent across teams.
Clarifying questions: Inline in the gateway or a library? Latency budget? Which policies vary per app? Languages?
request -> input checks
(PII redact, injection score, topic)
|
model call
|
output checks
(PII, policy, schema, grounding, toxicity)
| |
pass block / rewrite / escalate
|
decision log (policy version, scores)
Key components: policy-as-configuration per application, fast detectors (regex plus models) and a log of every decision with the policy version. Managed options such as Amazon Bedrock Guardrails exist; check current features.
Trade-offs: more checks add latency and false positives; run independent checks in parallel and tier them by risk. Guardrails reduce risk but do not replace least privilege.
Scaling: stateless and horizontally scaled, often co-located with the gateway.
Evaluation: attack and benign test sets per policy, false-positive rate. See AI red teaming.
Security and cost: the redaction mapping itself is sensitive data; protect it.
47. Design a prompt management and experimentation system.
Answer: A registry where prompts are versioned like code, linked to evaluation results, deployed through environments and rolled out gradually with A/B tests.
Clarifying questions: Who edits prompts (engineers, domain experts)? How often? Review requirements? Must every production answer be traceable to a prompt version?
prompt edit -> registry (version, owner)
|
eval suite run -> scores vs baseline
|
review + approve -> staging
|
canary: small share of traffic -> metrics
|
promote / roll back (label "prod")
Key components: versioned templates with variables, environment labels, links to eval runs, traffic splitting and the prompt version stamped on every trace.
Trade-offs: storing prompts in Git gives review and history; a registry allows non-engineers to edit and change prompts without redeploys. Many teams use Git as the source of truth with a sync.
Scaling: prompts are cached in the application; a fetch failure must fall back to the last known version.
Evaluation: no promotion without passing the regression set.
Security and cost: prompt changes are production changes; restrict who can promote. See the prompt engineering interview questions.
48. Design a self-hosted LLM inference platform on Kubernetes.
Answer: A GPU serving platform running open-weight models behind an internal API, chosen when data cannot leave the network, volume is high and steady, or a fine-tuned model must be served.
Clarifying questions: Why self-host (residency, cost, customisation)? Model sizes? Traffic pattern? GPU availability in the region? Team capacity to operate it?
internal apps -> LLM gateway
|
K8s (EKS / GKE / AKS) inference service
model servers (e.g. vLLM) on GPU nodes
continuous batching, KV cache
|
autoscaler on queue depth / GPU use
|
model registry -> weights (object store)
|
metrics: tokens/s, TTFT, queue, GPU memory
Key components: an inference engine with continuous batching, GPU node pools, model weights pulled from a registry, autoscaling on queue depth and warm capacity because cold starts for large models are slow.
Trade-offs: self-hosting gives control and can be cheaper at steady high utilisation, but idle GPUs are expensive and operations are hard; quantisation saves memory but needs quality evaluation.
Scaling: scale replicas per model; route by model and size.
Evaluation: quality parity tests against a hosted baseline, load tests.
Security and cost: network isolation, licence review of model weights, GPU cost dashboards. See self-hosting LLMs.
49. Design a semantic cache for LLM responses.
Answer: A cache that returns a stored answer when a new question is semantically equivalent to a previous one, used only where answers do not depend on the user or on fast-changing data.
Clarifying questions: How repetitive is traffic? Are answers personalised or permission-dependent? How fresh must they be?
question -> normalise -> exact cache?
| miss
embed -> nearest cached question
|
similarity >= threshold AND same scope
(tenant, role, source version)?
| yes | no
cached answer full pipeline
|
store with TTL + scope
Key components: an exact-match layer, an embedding-based layer, cache keys scoped by tenant, permission set and index version, TTLs, and invalidation when sources change.
Trade-offs: a loose threshold returns wrong answers for subtly different questions ("cancel order" versus "cancel subscription"); a strict one rarely hits. Provider prompt caching (reusing a static prefix) is a separate, safer saving.
Scaling: vector index plus a key-value store; small compared with the main system.
Evaluation: sampled cache hits reviewed for equivalence, hit rate.
Security and cost: a cache shared across permission scopes is a data-leak path; scope it. For vector index choices, see the vector database interview questions.
50. Design a fraud alert narrative generator for a bank's investigations team.
Answer: When the fraud detection model raises an alert, the system assembles transaction facts and model reason codes, then drafts an investigator-ready narrative and, where applicable, a draft suspicious activity report that a human finalises.
Clarifying questions: Alert volume? Report format required by the regulator? Which data can the model see? Does the narrative influence the decision or only document it?
fraud model alert + reason codes
|
facts builder (SQL, deterministic):
transactions, devices, counterparties
|
LLM: narrative ONLY from facts JSON
(every number referenced by ID)
|
validator: numbers / dates match facts
|
investigator edits + decides -> case system
Key components: deterministic fact assembly, a prompt forbidding any figure not in the facts, a validator that checks every amount and date against the source, and full case audit.
Trade-offs: narrative fluency versus strict fidelity; prefer fidelity, with templates for fixed sections.
Scaling: asynchronous per alert; peaks follow detection model runs.
Evaluation: factual consistency (automated), investigator edit rate, report acceptance by the compliance reviewer.
Security and cost: the LLM never decides whether fraud occurred; data stays in-region with tight access.
Key takeaways
- Spend the first minutes on clarifying questions: users, data, permissions, risk of a wrong answer, scale and success metric.
- Estimate tokens, cost and latency with a formula and illustrative numbers, then name the levers that move them.
- Choose models per step through evaluation, and hide them behind a gateway so they can change.
- Keep deterministic work (arithmetic, eligibility, permissions, matching) in code; use the LLM for language.
- Draw trust boundaries explicitly: identity, retrieved content as untrusted, tool policy and data leaving the organisation.
- Every design needs an evaluation plan with a golden set, component metrics and release gates.
- Risky actions follow propose, check policy, approve, execute and audit.
Interview preparation checklist
- Practise the Q1 sequence until you can run it without notes in under an hour.
- Prepare a reusable estimation sheet: requests per month, tokens per request, model calls per task, cost formula.
- Draw five core designs from memory: document Q&A, support agent, text-to-SQL, LLM gateway and an agent with approvals.
- For each, prepare one deep trade-off you can defend for five minutes.
- Know one evaluation tool (Ragas, LangSmith or Langfuse) and one tracing approach (OpenTelemetry) well enough to describe a real setup.
- Rehearse failure-mode answers: provider outage, prompt injection, stale index, agent loop, cost spike.
- Build at least one design end to end and be ready to explain what broke and what you changed.
- Do a mock interview where someone plays a sceptical stakeholder asking about risk and cost.
FAQ
What is an AI system design interview?
It is an interview round where you design a complete AI-powered system, such as a RAG assistant or an agent, covering requirements, architecture, models, evaluation, failure handling, security and cost, rather than only writing code or explaining concepts.
How is AI system design different from traditional system design?
Traditional design focuses on throughput, storage and consistency. AI system design adds non-deterministic outputs, token-based cost and latency, retrieval quality, evaluation, prompt injection and human review, while still requiring the classic distributed systems basics.
Which roles have AI system design rounds?
Senior AI engineers, GenAI developers, AI architects, ML platform engineers and Forward Deployed Engineers commonly face them. Junior roles may get a lighter version focused on a single RAG or chatbot design.
How should I prepare for a GenAI system design interview?
Learn a repeatable method, practise estimation with illustrative numbers, memorise a few core architectures, build at least one system end to end and do timed mock interviews where you explain trade-offs aloud.
Do I need to know specific cloud services for AI system design?
It helps to know one cloud well, such as AWS with Amazon Bedrock or Azure with Azure OpenAI, but interviewers mostly score concepts: gateways, retrieval, identity, queues, evaluation and observability. Name services only where you are confident they fit.
Should I give exact token prices in the interview?
No. Prices change often. Show the cost formula, use clearly illustrative numbers and explain the levers, such as smaller models, caching, shorter context and batch processing.
How much coding is involved in an AI system design round?
Usually little or none in the design round itself, but you may be asked to sketch an API, a tool schema or a prompt structure. Other rounds in the same loop often test Python coding.
What is the most common mistake in LLM system design interviews?
Jumping to a vector database and an agent framework before clarifying users, data, permissions and success metrics. The second most common mistake is having no evaluation plan.
Can a DevOps or cloud engineer do well in AI system design interviews?
Yes. Infrastructure, identity, reliability and cost skills transfer directly. Add RAG, agents, evaluation and guardrails by building a few real projects, and you will have strengths many model-focused candidates lack.
Designing these systems in an interview is easier once you have shipped one. Cloudsoft's FDE PRO program covers enterprise RAG, agents, MCP, security, Kubernetes deployment and evaluation over 12 weeks with 120+ hours of live sessions and 60+ labs, plus placement support until you're placed, in Ameerpet (beside Ameerpet Metro) or live online. If you need broader AI/ML, cloud and security foundations first, look at the APEX AI, ML, Cloud and Security program. Call +91 96660 19191 for a free demo.



