New batches starting this week Β· Limited seats

AI Engineer Interview Questions and Answers 2026 (100 Questions)

A broad hub of 100 AI engineer interview questions for 2026, from LLM fundamentals and RAG to agents, evaluation, LLMOps, security, system design and 20 production scenarios.

AI engineer interview questions 2026: 100 questions across Python, LLMs, RAG, agents, cloud AI, LLMOps and system design
Last updated Β· 52 min read Β· 11,434 words

These 100 AI engineer interview questions cover the whole AI application engineering job: LLM fundamentals, context engineering, RAG, agents and tools, evaluation, data pipelines, cloud AI services, LLMOps, security, cost, system design and real production scenarios. In 2026, AI engineer interviews care less about whether you can call a model API and more about whether you can ship an AI feature that stays correct, secure, observable and affordable once real users depend on it. This guide covers the breadth. For depth on any one area, follow the links to our specialist question banks.

How to use this guide

  • Freshers and early-career: interviewers check fundamentals. Can you explain tokens, embeddings, RAG and tool calling correctly, and write clean Python around a model call? Focus on questions 1–50.
  • Mid-level developers moving into AI: expect evaluation, data pipelines, deployment, cost and security. Interviewers want to see that you treat an LLM as an unreliable dependency. Focus on questions 28–77.
  • Senior and lead roles: system design and scenario rounds carry the most weight. You will be asked to justify trade-offs, describe how things fail, and say how you would know a system is working. Focus on questions 72–100.
  • Answer out loud before you read each model answer. Then rewrite it using a project you have actually built. Interviewers spot memorised answers quickly.

Contents

Role and fundamentals

1. What does an AI engineer do, in one paragraph?

Answer: An AI engineer builds software products and internal systems on top of AI models, mostly pre-trained foundation models reached through an API or self-hosted. The job covers designing the context a model sees (prompts, retrieved documents, tool results), connecting the model to data and enterprise systems, building evaluation so you know whether it works, and running it in production with security, monitoring and cost controls. The model is one component. The AI engineer owns the system around it.

Interview tip: End with one sentence about something you shipped, for example "for my last project that meant a retrieval pipeline, an eval set of 150 questions and a cost dashboard." That tells the interviewer you have done the work, not just read about it.

2. How is an AI engineer different from an ML engineer and a data scientist?

Answer: Each role starts from a different place and gets measured differently.

RoleCore questionTypical artefacts
Data scientistWhat does the data tell us, and which model or analysis answers the business question?Analyses, experiments, notebooks, trained models, reports
ML engineerHow do we train, serve and keep a model reliable at scale?Feature pipelines, training jobs, model registry, serving infrastructure
AI engineerHow do we turn foundation models into a working, safe product feature?RAG pipelines, agents, tool integrations, eval suites, APIs, guardrails

In practice the roles overlap. An AI engineer rarely trains models from scratch, but must understand enough ML to decide when fine-tuning, a small classifier or a traditional model beats an LLM.

3. How does an AI engineer compare with a Forward Deployed Engineer?

Answer: Both build production AI. What differs is where they sit and what they are accountable for. An AI engineer usually works on a product or platform team and owns a feature end to end. A Forward Deployed Engineer is embedded with a specific customer and owns the outcome inside that customer's environment, including discovery, messy integrations, security reviews and stakeholder communication. The FDE role was popularised by Palantir. AI labs and enterprise software companies now hire FDEs too. We compare the two in detail in FDE vs AI engineer.

4. When should you NOT use an LLM?

Answer: Skip the LLM when a deterministic rule, a SQL query, a search index or a small trained classifier solves the problem more cheaply, more predictably or more explainably. Examples: validating a PAN format, calculating EMI, routing tickets by a fixed product code, or fraud scoring where regulators expect an explainable model. LLMs earn their place on unstructured language: summarising, extracting from messy documents, answering questions over text, and choosing among tools.

Interview tip: Interviewers ask this to check judgement. A candidate who wants an LLM for everything is a red flag.

5. What is the difference between AI, generative AI and agentic AI?

Answer: AI is the broad field of systems that perform tasks that need intelligence. Generative AI is the subset that produces new content (text, code, images, audio), usually with foundation models. Agentic AI uses generative models in a loop: the model plans, calls tools, observes results and acts toward a goal with some autonomy. Each step up adds capability and also adds risk, because the system moves from producing text to taking actions.

6. What are the main building blocks of an LLM application?

Answer: A typical production LLM application has these layers:

  1. A client or channel: web, Teams, WhatsApp or an API.
  2. An application backend that handles auth, sessions and business rules.
  3. Context assembly: the system prompt, conversation history, retrieved documents and tool definitions.
  4. A model access layer, often behind a gateway.
  5. Tools and integrations.
  6. Guardrails on input and output.
  7. Observability: traces, logs, cost.
  8. Evaluation: offline test sets and online feedback.

Weak candidates list only the model and a vector database. Strong candidates name evaluation and observability without being asked.

7. What does "from AI demo to enterprise outcome" mean to you?

Answer: A demo proves a model can do a task once, on friendly inputs. An enterprise outcome means the task gets done reliably for real users, on real data, under real constraints: permissions, audit, latency targets, budget and change management. The gap between the two is where most AI engineering effort goes, and it is where most projects fail. Name the gaps specifically: data quality, access control, evaluation, edge cases, integration, monitoring and user adoption. Our article on why AI demos fail in enterprise production lists these failure modes.

Python and software engineering for AI

8. Which Python skills matter most for an AI engineer?

Answer: These matter most:

  • async I/O, because model calls and tool calls are network-bound;
  • type hints and Pydantic models for structured inputs and outputs;
  • a web framework such as FastAPI;
  • HTTP clients with timeouts and retries;
  • streaming responses;
  • packaging and dependency pinning;
  • testing with pytest, including mocking model calls;
  • basic data handling with pandas or polars.

Notebooks are fine for exploration, but production code needs modules, tests and configuration that lives outside the code. Our Python for AI engineers guide covers the toolkit.

9. Why is async important when calling LLMs, and what are the pitfalls?

Answer: A model call can take seconds. With synchronous code a single worker sits idle during that time. Async lets one process handle many concurrent requests and lets you fan out independent calls, such as several retrievals or tool calls, in parallel. The pitfalls:

  • A blocking library inside an async route, for example a synchronous database driver, stalls the whole event loop.
  • Unbounded concurrency can hit provider rate limits. Use semaphores.
  • Requests without timeouts can hang forever.
  • Exceptions inside gathered tasks get lost if nobody handles them.

Interview tip: Mention asyncio.gather with a concurrency limit and per-call timeouts. That shows you have actually run this in production.

10. How do you get reliable structured output from an LLM?

Answer: Use the provider's native structured output or function-calling mode with a JSON schema rather than asking for JSON in free text. Define the schema as a Pydantic model and validate every response. If validation fails, retry once with the error message included, and after that fall back to a safe default or a human. Keep schemas small and flat, use enums for categorical fields, and make fields optional only when the model genuinely may not know the value. Schema-valid output can still be wrong, so you still need business validation, such as checking that a claimed amount does not exceed the policy limit. More in function calling and structured outputs.

11. How do you write unit tests for code that calls an LLM?

Answer: Split the code into deterministic and non-deterministic parts. Unit-test the deterministic parts normally: prompt assembly, parsing, validation, tool routing and error handling, with the model client mocked to return fixed responses, including malformed ones. Test model behaviour separately with an evaluation suite that runs against real models and scores outputs. In CI, run unit tests on every commit and run a smaller eval subset on prompt or model changes. Asserting exact string equality on live model output is fragile, so don't.

12. How do you handle retries, timeouts and rate limits for model APIs?

Answer: Set an explicit timeout on every call. Retry only errors that can succeed on a second attempt (throttling, 5xx, connection resets), using exponential backoff with jitter and a retry cap. Respect retry-after headers. Never blindly retry calls that have side effects, such as a tool that creates a ticket. Use idempotency keys there. Put a circuit breaker in front of the provider so a provider outage degrades gracefully, either to a fallback model or a "try again later" message, instead of queueing thousands of stuck requests. At scale, a gateway can centralise rate limiting and failover. See LLM gateway explained.

13. How would you structure the codebase of an LLM service?

Answer: Separate these concerns:

  • API layer: routes, auth, request validation.
  • Orchestration: the workflow or agent graph.
  • Prompts: versioned templates kept as files, not as strings buried in code.
  • Retrieval: indexing and search.
  • Tools: one module per integration with typed interfaces.
  • Model clients: an adapter so you can swap providers.
  • Configuration: model IDs, temperature, limits.
  • Evaluation: datasets, scorers, runners.

Keep model IDs and prompts in configuration so changing them is a reviewed, testable change rather than a hot edit.

14. What does good logging look like in an AI application?

Answer: Use structured logs with a request or trace ID on every line. For each model call, record the model ID, the prompt version, token counts, latency, cost estimate, any tool calls with their arguments and results, guardrail decisions and the final status. Store full prompts and responses only where policy allows, and redact PII before logging. Traces such as OpenTelemetry spans are more useful than flat logs for multi-step agents, because they show which step was slow or failed.

LLM fundamentals

15. Explain how an LLM generates text, without hand-waving.

Answer: An LLM is a transformer trained to predict the next token given the previous tokens. At inference time, it encodes the input tokens, computes a probability distribution over the vocabulary for the next token, samples or selects one token, appends it, and repeats until it hits a stop condition or a token limit. The model has no lookup into a fact database. Anything it "knows" is encoded in its weights from training, plus whatever you put in the context. That is why grounding and evaluation matter. Our separate LLM interview questions go deeper on model internals.

16. What are tokens and context windows, and why do they matter to an engineer?

Answer: Tokens are the sub-word units the model reads and writes. Pricing, latency and limits are all counted in tokens. The context window is the maximum number of tokens (input plus output) the model can attend to in one call. Engineering consequences:

  • Long prompts cost more and slow down time to first token.
  • Very long contexts can bury relevant facts in the middle.
  • Non-English text, including Indian languages, often uses more tokens per word.
  • You must budget space for the output.

17. What do temperature and top-p control, and how do you set them?

Answer: Both control randomness in sampling. Temperature scales the probability distribution: lower values make the model more deterministic, higher values make it more varied. Top-p (nucleus sampling) limits choices to the smallest set of tokens whose combined probability reaches p. For extraction, classification and tool calling, use low temperature. For brainstorming or creative copy, use higher values. Tune one of the two, not both together. Some reasoning-focused models restrict or ignore these parameters, so check the provider documentation for each model.

18. Why do LLMs hallucinate, and how do you reduce it?

Answer: The model produces plausible continuations, not verified facts. When the context lacks the answer, or when the question invites a confident tone, it fills the gap with fluent guesses. To reduce hallucination:

  • Ground answers in retrieved sources and require citations.
  • Instruct the model to say "I don't know" and actually test that it does.
  • Lower the temperature.
  • Constrain outputs with schemas.
  • Verify claims against tools or databases.
  • Measure faithfulness in evaluation.

You can reduce hallucination but you cannot eliminate it, so design the UX and approvals around that.

19. What are embeddings, and how are they different from the LLM itself?

Answer: An embedding model maps text (or images) to a fixed-length vector where semantically similar inputs land close together. It is usually a separate, smaller model from the generator. Embeddings power semantic search, clustering, deduplication and classification. The embedding model you choose sets the quality ceiling for retrieval: if relevant chunks never get retrieved, no generator can fix the answer. If you change the embedding model, you have to re-embed the whole corpus.

20. What are reasoning models, and when are they worth the cost?

Answer: Reasoning models spend extra inference-time compute on intermediate thinking before they answer. That improves performance on multi-step problems such as maths, planning, complex code and tricky policy interpretation. They cost more and respond more slowly. Use them for hard, low-volume decisions or as a planner. Use faster general models for high-volume, simple steps such as classification, extraction and formatting. Many good systems combine both.

21. How do you choose between a large hosted model, a small model and a self-hosted open-weight model?

Answer: Start from requirements: task difficulty, latency, volume, data residency, cost and how much operations work you can take on. Large hosted models are strongest on hard tasks and need no infrastructure. Small models are cheaper and faster for narrow tasks, especially after fine-tuning. Self-hosted open-weight models give you control over data and versions, but you own GPUs, scaling, patching and safety. Run the same evaluation set across candidates and decide with data.

Prompting and context engineering

22. What is the difference between prompt engineering and context engineering?

Answer: Prompt engineering is writing the instructions: role, task, constraints, examples, output format. Context engineering is the broader job of deciding everything that goes into the model's context window on each call: system instructions, selected history, retrieved documents, tool definitions, tool results, memory and user metadata. It also covers what to leave out. In production systems, most quality problems turn out to be context problems (wrong documents, stale history, too many tools) rather than wording problems. See the context engineering guide.

23. What goes into a good system prompt for a production assistant?

Answer: A good system prompt covers:

  • A clear role and scope: what the assistant does and does not do.
  • Grounding rules: answer only from the provided sources, cite them, say when the answer isn't there.
  • Tone and audience.
  • The output format.
  • Escalation behaviour.
  • Explicit handling of common edge cases you found in testing.

Keep it versioned and reviewed like code. Don't put secrets or authorisation logic in it. A system prompt is not a security boundary, and users can often get the model to reveal it.

24. When do few-shot examples help, and when do they hurt?

Answer: Few-shot examples help when the output format or style is hard to describe in words, for example a specific extraction layout or a classification with subtle boundaries. They hurt when the model copies surface details from them, when they push out more useful context, or when they bias the model toward the example labels. Choose examples that cover edge cases rather than easy cases, and for large label sets consider retrieving examples dynamically per query. Measure with and without examples on your eval set, rather than assuming they help.

25. How do you manage conversation history in a long chat?

Answer: Don't resend the full transcript forever. Common strategies:

  • A sliding window of recent turns.
  • A running summary of older turns.
  • Extraction of durable facts (the user's account type, the open ticket ID) into structured state.
  • Retrieval over past turns when they become relevant.

Keep tool results compact: store the full result, but pass the model only the fields it needs. Test long-conversation behaviour explicitly, because quality often degrades quietly after many turns.

26. How do you version and test prompts?

Answer: Store prompts as templates in the repository (or in a prompt registry), each with a version ID. Treat every change as a pull request that runs the eval suite and compares scores with the previous version. Log the prompt version with every production call so you can link quality regressions to a specific change. For risky changes, roll out to a fraction of traffic first. Tools such as LangSmith or Langfuse can manage prompt versions and evaluation runs together.

27. What is prompt chaining, and when is it better than one big prompt?

Answer: Prompt chaining splits a task into sequential calls, for example classify, then retrieve, then draft, then check. Each step has a narrow job and can be tested on its own. Chaining beats one big prompt when the steps need different models or temperatures, when you need to validate intermediate results, or when one prompt keeps getting part of the task wrong. The costs are more latency, more calls and more places to fail. If a single well-structured prompt passes your evals, keep it simple.

RAG

28. Explain RAG and why enterprises use it.

Answer: Retrieval-augmented generation retrieves relevant passages from your own data at query time and puts them in the model's context, so the answer is grounded in current, private sources. Enterprises use it because:

  • knowledge changes faster than you can retrain a model;
  • answers need citations for trust and audit;
  • access control can be enforced at retrieval time;
  • it is usually cheaper and more controllable than fine-tuning for knowledge.

Start with what is RAG, and use our dedicated RAG interview questions for retrieval-level depth.

29. Walk through the indexing and query pipelines of a RAG system.

Answer:

INDEXING
sources -> parse -> clean -> chunk -> embed
        -> store (vectors + text + metadata + ACLs)

QUERY
question -> rewrite -> hybrid search (filter by ACL)
         -> rerank -> top-k chunks -> prompt
         -> LLM -> answer + citations -> checks

Every stage can introduce errors: parsing loses tables, chunking splits a clause, the embedding misses an acronym, filters drop the right document, reranking reorders it. When debugging, find the stage that failed before changing the prompt.

30. How do you choose a chunking strategy?

Answer: Let the document structure and the question types decide. Structure-aware chunking (by headings, sections or clauses) usually beats fixed-size splitting for policies, contracts and manuals. Fixed-size chunks with overlap are a reasonable baseline for uniform prose. Keep tables whole or convert them to text with their headers. Attach the section title and document metadata to each chunk. Then test several chunk sizes against your eval questions and pick on retrieval metrics.

31. What is hybrid search and why add a reranker?

Answer: Hybrid search combines keyword search (BM25), which catches exact terms such as product codes, error IDs and names, with vector search, which catches paraphrases. The two result lists are fused, often with reciprocal rank fusion. A reranker (usually a cross-encoder) then rescores the top candidates against the query far more precisely than embedding similarity can, so the few chunks you send to the model are the right ones. This combination is a common fix when "the answer is in the index but RAG doesn't find it."

32. How do you enforce document permissions in RAG?

Answer: Enforce them at retrieval time, not in the prompt. Store access metadata (groups, roles, departments, sensitivity labels) with each chunk at indexing time, synced from the source system. At query time, resolve the user's identity and groups from the identity provider (for example Microsoft Entra ID) and apply them as a mandatory filter in the search query. Telling the model "do not reveal HR documents" is not access control. Also handle permission changes: deleted documents and revoked access must reach the index quickly, and you should test that they do.

33. RAG or fine-tuning: how do you decide?

Answer: RAG is for knowledge: facts that change, need citations or need permissions. Fine-tuning is for behaviour: a consistent format, a domain style, a narrow classification, or getting a smaller model to match a larger one on a specific task. Use both when the model needs domain behaviour and also current facts. Try prompting and RAG first, because they are cheaper to change. See RAG vs fine-tuning.

34. What are agentic RAG and GraphRAG, and when are they justified?

Answer: In agentic RAG, an agent decides whether to retrieve, which source to search, how to reformulate the query, and whether the results are good enough or it should search again. It helps with multi-hop questions and multiple sources, at the cost of latency and unpredictability. GraphRAG builds a knowledge graph of entities and relationships from the corpus and uses it for questions that depend on connections ("which suppliers are linked to delayed shipments across regions"). Both are justified only once a well-tuned plain RAG baseline fails on a measurable class of questions.

Agents and tools

35. What is an AI agent, precisely?

Answer: An agent is a system where a model decides, in a loop, which actions to take to reach a goal. It reads the goal and context, picks a tool, observes the result, and continues until it finishes or hits a limit. The defining feature is that the model controls the flow. In a fixed workflow, the code controls the flow and the model fills in steps. Many production "agents" are actually workflows with one or two model-driven decisions, and that is often the right design. For depth, see our agentic AI interview questions.

36. How does tool (function) calling work under the hood?

Answer: You send the model a list of tool definitions: name, description and a JSON schema for the arguments. When the model decides a tool is needed, it returns a structured tool-call request instead of text. Your code (not the model) validates the arguments, checks permissions, executes the tool and sends the result back as a new message. The model then continues. The model never executes anything itself. That is why the execution layer is where you enforce authorisation, validation, rate limits and audit.

37. Workflow or autonomous agent: how do you decide?

Answer: Use a deterministic workflow when the steps are known and the order matters, as with approvals, compliance processes and data pipelines. Use an agent when the path depends on what the system discovers along the way, such as troubleshooting or research across sources. A common middle ground is a workflow with bounded agentic steps: a graph where some nodes let the model choose among a few tools, with clear exits and human checkpoints. Frameworks such as LangGraph model this explicitly.

38. What is MCP, and how does it change tool integration?

Answer: The Model Context Protocol is an open protocol, introduced by Anthropic in late 2024, that standardises how AI applications connect to tools, data (resources) and prompt templates through MCP servers. Instead of writing a custom integration for every app and model combination, you build one MCP server for a system such as Jira or a database, and any MCP-capable client can use it. It does not remove the need for authentication, authorisation or input validation. Those stay your job. If you build servers with the official Python SDK, note that in v2 the high-level FastMCP class was renamed MCPServer. Check the SDK's migration guide. Depth: MCP interview questions.

39. How do you stop an agent from looping or running away?

Answer: Put hard limits outside the model: maximum steps, maximum tool calls per tool, maximum tokens and cost per run, and a wall-clock timeout. Detect repetition, such as the same tool called with the same arguments. Give the model a clear "finish" or "escalate" action. Make tool errors informative so the model can recover instead of retrying blindly. Log every step so you can see where loops start. For long-running agents, checkpoint state so a failure resumes cleanly instead of restarting.

40. What kinds of memory can an agent have?

Answer: Working memory is the current context and state of the run. Short-term or session memory is the conversation and intermediate results for this session. Long-term memory is facts, preferences or past outcomes stored outside the model (in a database or vector store) and retrieved when relevant. Episodic memory records past runs that the agent can learn from. Long-term memory raises privacy and correctness questions: what gets stored, who can see it, how it is corrected or deleted.

41. When are multi-agent systems a good idea, and when are they overkill?

Answer: Multi-agent designs help when sub-tasks need different tools, permissions or prompts, as in an IT-ops system with separate diagnosis, change and communication agents, or when the work runs in parallel. They are overkill when one agent with good tools can do the job. Every extra agent adds coordination overhead, latency, cost and new failure modes such as agents misunderstanding each other. Start with one agent, split only when evals or security boundaries require it, and keep a supervisor or explicit graph so control flow stays understandable.

Evaluation

42. Why is evaluation the core skill of an AI engineer?

Answer: Without evaluation, every change to a prompt, model, chunking setting or tool is a guess, and you cannot tell an improvement from a regression. With a decent eval suite you can upgrade models safely, compare vendors, catch regressions in CI and give stakeholders evidence rather than impressions. In enterprise projects, "how will we know it works?" is often the question that decides whether a pilot gets approved for production. See LLM evaluation for the full method.

43. How do you build an evaluation dataset from scratch?

Answer: Start with real or realistic inputs. Take questions from support tickets, search logs or subject-matter experts. Write expected answers or acceptance criteria, plus source references for RAG. Cover common cases, edge cases, questions that must be refused or escalated, adversarial inputs, and questions whose answer isn't in the data. Start small, around a hundred well-chosen items, and grow the set from production failures. Version the dataset, and keep a held-out portion so you don't tune prompts to the test.

Real-world example: Consider an insurer's claims assistant. The most valuable eval items usually come from claims handlers listing questions where a confidently wrong answer would cause a mis-payment. Those items deserve the strictest scoring.

44. Which metrics would you use to evaluate a RAG system?

Answer: Measure retrieval and generation separately.

  • Retrieval: context recall (were the needed passages retrieved?), context precision (how much of what was retrieved was relevant?), and hit rate or MRR against labelled source documents.
  • Generation: faithfulness (is the answer supported by the context?), answer relevance, correctness against a reference, and citation accuracy.
  • Operational: refusal behaviour when the answer is missing, latency and cost per answer.

Libraries such as Ragas implement several of these. What matters more than the tool is calibrating the metrics against human judgement on your own data.

45. What is LLM-as-a-judge, and what are its weaknesses?

Answer: LLM-as-a-judge uses a model to score outputs against a rubric, for example "is every claim supported by the sources? yes/no with a reason." It scales evaluation beyond what humans can review. Its weaknesses:

  • Position and verbosity bias: judges often favour longer answers or the first option.
  • Self-preference toward outputs from the same model family.
  • Inconsistency across runs.
  • Blind spots on domain correctness.

To mitigate: use narrow binary or small-scale rubrics, ask for the reason before the score, randomise order in pairwise comparisons, and measure how well the judge agrees with expert labels before you trust it.

46. How do you evaluate an agent rather than a single response?

Answer: Evaluate the trajectory as well as the final answer. Did the agent complete the task? Did it pick the right tools with valid arguments? Did it avoid forbidden actions, ask for approval when required, and finish within step and cost budgets? Build scenario tests with mocked or sandboxed tools and check the end state, for example "ticket created with correct priority, no email sent." Track task success rate, tool-call accuracy, steps per task and escalation rate.

47. How do offline evals and online monitoring fit together?

Answer: Offline evals run before release on a fixed dataset and act as the gate for prompt, model and code changes. Online monitoring runs on live traffic: user feedback, escalation and retry rates, sampled LLM-judge scoring of real conversations, guardrail trigger rates, latency and cost. Online signals reveal new failure types. You turn those into new offline test cases, so the eval set improves as the product does. Shipping without both is how silent quality decay goes unnoticed.

If you want hands-on practice with evaluation, RAG and production deployment alongside cloud and security, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program covers these as connected skills rather than isolated topics.

Data pipelines

48. What does a production ingestion pipeline for AI look like?

Answer: A production ingestion pipeline does the following:

  • Pulls from source connectors (SharePoint, Confluence, S3, databases, ticketing tools), incrementally where possible.
  • Parses documents into text and structure.
  • Cleans and deduplicates.
  • Extracts metadata and permissions.
  • Chunks, embeds and upserts into the index.
  • Records lineage (which source version produced which chunk).

It needs scheduling, retries, dead-letter handling for failed documents, and metrics on documents processed, failed and skipped. Treat it as a real data pipeline with an owner and alerting, not a one-off script.

49. How do you handle PDFs, scanned documents and tables?

Answer: Pick the parser by document type. Text PDFs can be extracted directly, but check reading order and multi-column layouts. Scans need OCR or a document-intelligence service. Tables need layout-aware parsing that keeps headers with cells, often converted to Markdown or row-wise sentences. Images and charts may need a multimodal model to describe them. Sample parsed output by hand before indexing anything, because parsing errors are invisible downstream and show up as "the model is wrong".

50. How do you keep an index fresh and consistent with the source?

Answer: Use change detection (webhooks, change feeds or modified timestamps) for incremental updates. Use stable document IDs so updates replace old chunks rather than duplicating them. Propagate deletes. Run a periodic full reconciliation to catch drift. Store the source version and indexing time on each chunk so answers can show how fresh they are, and alert when the time since the last successful sync exceeds what the business accepts. Stale permissions are as dangerous as stale content.

51. How do you handle PII in data pipelines for AI?

Answer: Classify data at ingestion. Decide per field whether to exclude it, mask it, tokenise it or keep it with stricter access control. Detect PII (Aadhaar, PAN, phone numbers, account numbers) with pattern rules plus a named-entity detector, and redact before indexing or logging when the use case does not need it. Keep processing within the required region. Document purpose and retention to meet obligations such as India's DPDP Act.

Cloud AI services

52. Give a high-level map of AI services on AWS.

Answer: Amazon Bedrock provides managed access to foundation models from several providers, along with Knowledge Bases for managed RAG, Guardrails, and evaluation features. For agents, AWS now points new projects to Amazon Bedrock AgentCore, a modular platform (runtime, memory, gateway, identity, observability, and browser and code-interpreter tools) that works with open-source frameworks and MCP. The original Bedrock Agents, now called Bedrock Agents Classic, is no longer open to new customers, though existing customers can keep using it. Amazon SageMaker AI covers custom training and hosting. Around these sit IAM, VPC endpoints, KMS, CloudWatch and S3. Depth: AWS Bedrock interview questions.

53. Give a high-level map of AI services on Azure.

Answer: Microsoft Foundry (previously Azure AI Foundry, and before that Azure AI Studio) is the platform for model access (including Azure OpenAI models), agents, evaluation and observability. Azure AI Search is the usual retrieval layer. For agent code, Microsoft Agent Framework is the open-source successor to Semantic Kernel and AutoGen, built by the same teams. Identity runs through Microsoft Entra ID, with private endpoints, Key Vault and Azure Monitor around it. Product names in this space change often, so check current Microsoft Learn documentation before an interview. Depth: Azure AI interview questions.

54. Give a high-level map of AI services on Google Cloud.

Answer: In 2026 Google rebranded and expanded Vertex AI as Gemini Enterprise Agent Platform. Google's own pages describe it as "formerly Vertex AI", and the Vertex AI services now come under that platform. It covers model access through Model Garden (Gemini plus third-party and open models), the Agent Development Kit (ADK), a managed agent runtime, and governance features such as an agent registry, agent identity and an agent gateway, plus evaluation and monitoring. You will still see "Vertex AI" in older docs, SDK names and job descriptions, so mention both names. Check current Google Cloud documentation for exact component names. See GCP for AI engineers.

55. How do you choose a cloud AI platform for a project?

Answer: Usually the customer's existing cloud, identity and data location decide it, not model benchmarks. If the data and Entra ID live in Azure, putting the AI there keeps networking, identity and compliance simple. Beyond that, compare the models available in the required region, data residency, private networking, quotas, managed RAG and agent features, observability, and pricing at your expected volume. To keep options open, put a thin abstraction or gateway between your code and the provider SDK.

56. What does "private" access to a cloud LLM involve?

Answer: Private access involves:

  • Calling the model endpoint through private networking (VPC endpoints or Private Link, private endpoints) so traffic doesn't cross the public internet.
  • Using workload identity (IAM roles, managed identities) instead of static API keys.
  • Encrypting data with customer-managed keys where policy demands.
  • Choosing regions to meet residency rules.
  • Confirming the provider's data-use terms, for example that prompts are not used for training.

Security teams will ask about each of these, so know how to explain them for at least one cloud.

Deployment, MLOps and LLMOps

57. What is LLMOps, and how does it differ from classic MLOps?

Answer: Classic MLOps centres on training pipelines, feature stores, model registries and drift in input data. LLMOps keeps the deployment and monitoring discipline but changes what you version and watch. You version prompts, retrieval configuration, tool definitions and model IDs, in addition to code. You evaluate with test sets and judges rather than one accuracy number. You monitor tokens, cost, latency, guardrail triggers and answer quality. And you manage upstream model changes you don't control. Depth: MLOps and LLMOps interview questions.

58. How would you design CI/CD for an LLM application?

Answer:

PR: code / prompt / config change
  -> lint + unit tests (mocked model)
  -> build container image + scan
  -> eval subset vs baseline (quality gate)
  -> deploy to staging -> full eval + load test
  -> canary in prod -> monitor quality + cost
  -> promote, or roll back

The AI-specific step is the eval gate. A prompt change that lowers faithfulness should fail the pipeline exactly the way a failing unit test does. Use GitHub Actions or similar for CI, and GitOps (Argo CD) for Kubernetes deployments.

59. How do you safely upgrade to a new model version?

Answer: Pin model versions explicitly rather than using floating aliases. Run the full eval suite on the new version and compare per category, not just the average, because a model can improve overall and still regress on your most important slice. Check format compliance, refusal behaviour, latency and cost too. Roll out behind a flag or to a canary slice of traffic, watch online metrics, and keep the old version ready for rollback. Track provider deprecation dates so you are not forced into an untested upgrade.

60. How would you deploy an AI service on Kubernetes?

Answer: Package the API and workers as containers. Set resource requests and limits, readiness and liveness probes, and horizontal autoscaling on request rate or queue depth rather than CPU alone, since LLM calls are I/O-bound. Use workload identity for cloud access and secrets from a vault. Put long-running agent tasks on a queue with workers instead of holding HTTP connections open. If you self-host models, GPU node pools and a serving engine bring their own scaling problems. Manage it all as code with Terraform and Helm or Kustomize.

61. What would you put on an LLMOps dashboard?

Answer:

  • Traffic and reliability: request volume, error rates by type, time to first token and total latency percentiles.
  • Usage and cost: tokens in and out, cost per request and per tenant or feature.
  • Quality: user feedback rate, sampled judge scores, escalation rate.
  • Safety: guardrail and refusal trigger rates.
  • Retrieval: empty-result rate, index freshness.
  • Agents: steps per task and tool error rates.

Alert on sudden changes, because a spike in refusals or a fall in retrieval hits usually means something upstream broke.

Security

62. What is prompt injection, and how do you defend against it?

Answer: Prompt injection is untrusted text (user input, or content inside retrieved documents, emails and web pages) that contains instructions the model then follows, overriding your intent. Indirect injection through retrieved content is the more dangerous form in enterprise systems. No single defence is enough, so layer them:

  • Treat all retrieved and tool content as data, and separate it clearly in the prompt.
  • Give the model least-privilege tools.
  • Require human approval for high-impact actions.
  • Validate tool arguments server-side.
  • Filter outputs, for example blocking unexpected URLs or data exfiltration patterns.
  • Red-team regularly.

Depth: AI security interview questions.

63. How do you secure the tools an agent can call?

Answer: Give the agent its own identity with scoped credentials, never a shared admin key. Where possible, act on behalf of the user so the agent can never do more than the user could. Expose narrow tools ("create_incident") rather than generic ones ("run_sql"). Validate every argument against a schema and business rules. Rate-limit tools, log every call with who, what and when for audit, and put destructive or financial actions behind explicit approval.

64. What are guardrails, and where do they sit?

Answer: Guardrails are checks around the model:

  • Input checks: jailbreak and injection detection, topic scope, PII detection.
  • Output checks: toxicity, PII leakage, groundedness, format and policy compliance.
  • Action checks: tool permissions and approvals.

They can be rules, classifiers, managed services (for example Bedrock Guardrails or Azure AI Content Safety) or LLM-based checks. Tune them on real traffic, because over-blocking frustrates users as much as under-blocking creates risk.

65. How do you prevent sensitive data leaking through an LLM application?

Answer: Control what goes in, and check what comes out. Enforce retrieval permissions, minimise and redact PII in context, and keep secrets out of prompts. Use providers and regions whose data-use terms meet your policy, restrict logs that contain prompts, and scan outputs for sensitive patterns. In multi-tenant systems, isolate data and caches per tenant so one tenant's context can never appear in another's answer.

66. What is AI red teaming, and what would you test first?

Answer: AI red teaming means deliberately attacking your AI system to find failures before attackers or users do. Test first the attacks that hurt your specific system most:

  • Indirect injection through documents the system will retrieve.
  • Attempts to make agents call tools outside their scope.
  • Extraction of other users' data.
  • Policy-violating outputs in your domain, such as financial or medical advice it shouldn't give.
  • System prompt leakage.

Turn every finding into an automated regression test.

Cost and latency

67. How do you estimate the cost of an LLM feature before building it?

Answer: Multiply expected requests by average input tokens (system prompt + history + retrieved context + tool definitions) and output tokens, at the model's per-token prices. Then add embedding, vector storage, reranking, guardrail calls and agent steps. Agents multiply the bill because each step is another call. Build a small spreadsheet with clearly labelled assumptions, for example an illustrative case of a few thousand requests a day at a few thousand tokens each. Then validate it against real measurements in the pilot.

68. What are the main levers for reducing LLM cost?

Answer: The main levers:

  • Route each task to the smallest model that passes evals.
  • Trim context: fewer, better chunks, compact tool results, summarised history.
  • Use provider prompt caching for stable prefixes.
  • Cache responses for repeated queries where answers don't depend on the user.
  • Batch offline jobs.
  • Cap output length.
  • Limit agent steps.

Measure cost per successful task, not per call. A cheaper model that needs retries or more human escalations can cost more overall.

69. Where does latency come from in a RAG or agent request, and how do you reduce it?

Answer: Latency builds up across query rewriting, retrieval, reranking, time to first token, generation length, guardrail calls and, for agents, each sequential tool round-trip. To reduce it:

  • Stream tokens to the user.
  • Parallelise independent retrievals and tools.
  • Use smaller models for intermediate steps.
  • Shorten prompts.
  • Cache.
  • Co-locate services in the same region.

Trace each step so you optimise the step that is actually slow. See LLM latency optimisation.

70. What is semantic caching, and what are its risks?

Answer: Semantic caching returns a stored answer when a new query's embedding is close enough to a previous query's, which saves a model call. The risks:

  • Serving a wrong answer when two questions look alike but differ ("cancel my card" vs "cancel my card replacement").
  • Serving one user's personalised or permissioned answer to another user.
  • Serving stale answers after the source changes.

Use it only for non-personalised content, scope the cache by tenant and permission set, set a strict similarity threshold and short TTLs, and invalidate when the source updates.

71. How does model routing work?

Answer: A router sends each request to a different model based on its difficulty, type or risk. Simple FAQs go to a small, fast model, and complex or high-stakes requests go to a stronger model. The router can be rules, a small classifier, or an escalation check (try the small model, and escalate if its confidence or validation fails). Evaluate the router itself, since mis-routing hard questions to a weak model is a quality bug. A gateway is a natural place to implement routing.

System design

72. Design an internal knowledge assistant for a 20,000-employee company.

Answer: Requirements first: sources (SharePoint, Confluence, policy PDFs), permissions, channels (web and Teams), languages, latency target, and audit needs.

Teams / Web
    |
API gateway -- Entra ID (SSO, groups)
    |
Assistant API (FastAPI)
  |-- query rewrite (small model)
  |-- hybrid search + ACL filter -- index
  |-- rerank -> top chunks
  |-- generate w/ citations (LLM via gateway)
  |-- output checks (PII, groundedness)
    |
Traces + feedback -> eval + dashboards

Ingestion: connectors -> parse -> chunk
  -> embed -> index (with ACLs), incremental

Key decisions: permission-filtered retrieval synced from the sources, citations on every answer, "not found" behaviour, an eval set per department, and a feedback button that feeds the eval set. Phase the rollout by department.

73. Design a customer-support agent that can take actions such as refunds.

Answer:

Customer -> chat channel -> auth (verify customer)
   |
Orchestrator graph
  |-- intent classify
  |-- FAQ path: RAG over help centre
  |-- action path: tools
  |     get_order, check_refund_policy,
  |     create_refund (limit + approval)
  |-- handoff path: human agent + summary
   |
Audit log + traces + quality sampling

Refunds above a threshold, or outside policy, go to a human with a prepared summary. Policy checks run in code, not in the prompt. Tools act only on the authenticated customer's orders. Measure resolution rate, wrong-action rate (which should be close to zero), handoff quality and customer satisfaction.

74. Design a document-extraction pipeline for invoices or claim forms.

Answer:

Upload / email -> queue
  -> classify doc type
  -> OCR / layout parse
  -> LLM extract to schema (Pydantic)
  -> validate (rules: totals, dates, IDs)
  -> confidence check
       ok  -> write to ERP / claims system
       low -> human review UI -> corrections
  -> corrections feed eval set

Async and queue-based, because volume spikes. Measure field-level accuracy per document type, the straight-through processing rate, and reviewer time. Keep the source page reference for each extracted field for audit.

75. Design a text-to-SQL analytics assistant.

Answer: Retrieve the relevant schema, table descriptions and example queries for the question. Generate SQL with the model, then check it before execution: parse it, allow only SELECT, permit only approved tables and columns, add row limits, and run it on a read-only replica under the user's data permissions. Return the result with the SQL shown, so analysts can verify it. Evaluate on a set of question-to-expected-result pairs, comparing results rather than SQL strings. The hard parts are business definitions ("active customer"), which a semantic layer or curated glossary solves better than prompting.

76. Design a multi-tenant AI platform for a SaaS product.

Answer:

Tenant request -> auth (tenant_id, user, plan)
  -> AI gateway
       quotas per tenant, routing, logging
  -> app services
       retrieval: index per tenant or
                  mandatory tenant filter
       prompts/config per tenant
  -> usage metering -> billing
Isolation: data, cache, memory, logs by tenant

Key topics: noisy-neighbour protection through quotas, per-tenant cost attribution, isolation of indexes and caches, tenant-specific prompt configuration without forking code, and evaluation that covers tenant-specific data.

77. Design an IT-operations assistant that triages alerts and proposes fixes.

Answer:

Alert (monitoring) -> event queue
  -> triage agent
       read: logs, metrics, recent deploys,
             runbooks (RAG), past incidents
  -> diagnosis + confidence + evidence
  -> ITSM ticket (ServiceNow / Jira)
  -> proposed fix -> human approves
       -> change via pipeline (not direct)
  -> post-incident: outcome -> eval set

Read access is broad, write access is narrow and approved. The agent proposes changes, and the existing change pipeline executes them. Measure time to triage, diagnosis accuracy against the eventual root cause, and false-escalation rate.

Behavioural and collaboration

78. Tell me about an AI feature you shipped. What went wrong, and what did you change?

Answer: Use a STAR structure, but spend most of the time on the technical decisions and the failure. A strong answer names a concrete problem, for example "retrieval missed answers in scanned annexures". It explains how you found the problem (eval results, traces, user feedback), what you changed (OCR parsing, hybrid search), and how you measured the improvement. It ends with what you would do differently. Interviewers listen for ownership, evidence and honesty about trade-offs, not a flawless story.

Interview tip: Prepare two such stories, one about quality and one about security or cost. Use projects you can discuss in detail, including your own portfolio projects if you are early in your career.

79. How do you explain LLM limitations to a non-technical stakeholder who expects perfection?

Answer: Turn the limitation into a business decision. Show measured accuracy on their own questions, explain which errors are tolerable and which are not, and propose controls: citations, human review for high-risk outputs, a scope limit, and an escalation path. Agree on success criteria before launch, such as an acceptable accuracy level on the eval set and a review process, so that "is it good enough?" has an agreed answer instead of being a matter of opinion.

80. How do you work with data, security and domain teams on an AI project?

Answer: Bring them in early, not at go-live. Ask data owners about sources, quality and permissions. Ask security for the threat model, data classification and approval of tool access. Ask domain experts to define the eval set and to judge answers. Share a short architecture and risk document in the first weeks, hold regular demos on real data, and keep a decision log. Many AI projects stall in security review because the security team saw the design too late.

Real-world scenarios

81. Your RAG assistant gave a confidently wrong answer to a senior executive. What do you do?

Answer: Contain the problem, diagnose from the trace, fix the root cause, and add a regression test. Acknowledge the error quickly, then find out whether retrieval or generation failed before changing anything.

What I would check:

  1. The trace: which chunks were retrieved, and was the correct document among them?
  2. Whether the source itself was outdated or contradictory.
  3. Whether the answer was faithful to the retrieved context, or the model added facts.
  4. Whether similar questions fail too (run a quick batch).

Production consideration: Add the question to the eval set. If the source was stale, fix the freshness pipeline. If the model invented facts, tighten the grounding rules and add a faithfulness check that falls back to "I couldn't find this in the sources."

82. Latency doubled overnight with no code change. How do you investigate?

Answer: No code change does not mean nothing changed. Use traces to find which step got slower, then look for upstream changes.

What I would check:

  1. The per-step span durations: model, retrieval, reranker, tools.
  2. Provider status pages and throttling or retry counts.
  3. Whether a model alias moved to a new version, or output length grew.
  4. Index size growth, a data change that inflated context, or a traffic spike.

Production consideration: Pin model versions, alert on latency percentiles per step, and keep a fallback model or region configured so you can shift traffic while you investigate.

83. Monthly LLM spend is far above the estimate. What do you do?

Answer: Break the cost down by feature, tenant, model and token type before cutting anything. Overruns usually come from a few causes: bloated context, agent loops, retries, or a single heavy user or integration.

What I would check:

  1. Cost per request distribution: is a long tail driving the total?
  2. Average input tokens over time (history growth, too many chunks).
  3. Agent steps per task and retry rates.
  4. Whether a batch job or test environment is hitting the production model.

Production consideration: Add per-tenant budgets and alerts at the gateway, route simple traffic to smaller models, enable prompt caching, and track cost per successful task as a standing KPI.

84. A user got the assistant to reveal content from a confidential HR document. Respond.

Answer: Treat it as a security incident: contain it, assess the scope, fix the access control and report through the incident process. This is an authorisation failure, not a prompt problem.

What I would check:

  1. Whether the document's ACL metadata was indexed correctly.
  2. Whether the retrieval filter was applied for this user and code path.
  3. Caches (semantic or response) that may have served another user's answer.
  4. Audit logs for other affected queries and users.

Production consideration: Make the permission filter mandatory in the retrieval layer (no code path can skip it), add automated permission tests to CI, and scope caches by permission set.

85. A pilot works well, but the business asks for production in four weeks. How do you plan?

Answer: Agree on a narrow production scope and a list of non-negotiables, then plan backwards. The usual gaps between a pilot and production are security review, real permissions, evaluation, monitoring and support ownership.

What I would check:

  1. Whether the eval set reflects real users, or only demo questions.
  2. Security and data-protection review lead times.
  3. Integration and SSO readiness, and load expectations.
  4. Who supports the system after launch, and the rollback plan.

Production consideration: Launch to a limited user group with monitoring and feedback, and expand as the metrics hold. A phased launch on time beats a full launch that fails publicly.

86. Your agent created duplicate tickets in ServiceNow. Diagnose and fix.

Answer: Duplicates usually come from retries without idempotency, or from the agent not recognising that its first call succeeded.

What I would check:

  1. Traces: did the tool time out and get retried, even though the first call succeeded?
  2. Whether the tool result clearly reported success and the ticket number to the model.
  3. Repeat calls with identical arguments within one run.
  4. Concurrency: two workers processing the same event.

Production consideration: Add idempotency keys (for example a hash of the conversation and the intent), check for an existing open ticket before creating one, and cap create-type tool calls per run.

87. After a provider model upgrade, JSON outputs started failing validation. What now?

Answer: Roll back to the pinned previous version if you can, then measure the regression and fix it in a controlled way.

What I would check:

  1. Whether you were using a floating alias instead of a pinned version.
  2. Which fields fail, and how (extra prose, missing fields, wrong types).
  3. Whether native structured-output mode was enabled.
  4. Your eval suite's format-compliance results on the new version.

Production consideration: Use schema-enforced output modes, validate and retry once with the error, and make format compliance a gate in model-upgrade testing.

88. Users say answers are "too generic". How do you improve them?

Answer: Generic answers usually mean weak or missing context, not a weak model. Check what the model actually saw.

What I would check:

  1. Whether the retrieved chunks are specific, or broad overview pages.
  2. Whether user context (role, location, product) is passed in.
  3. Whether the prompt asks for specifics: steps, numbers from sources, citations.
  4. A sample of "generic" complaints, sorted by question type.

Production consideration: Improve chunk metadata and reranking, add user attributes to the context where it is safe to do so, and add specificity to the eval rubric.

89. The business wants the assistant in Hindi and Telugu too. What changes?

Answer: Check model quality, retrieval across languages, token cost and evaluation separately for each language before you commit.

What I would check:

  1. Model quality in each language on domain questions, judged by native speakers.
  2. Whether the embedding model handles cross-lingual retrieval (a Telugu question over English documents).
  3. Token counts per language, which affect cost and latency.
  4. Code-mixed input such as Hinglish, and transliterated text.

Production consideration: Build an eval set for each language, decide whether to translate the query or retrieve cross-lingually, and keep the answer language consistent with the user's language.

90. A bank's security team rejects sending customer data to a public LLM API. What options do you propose?

Answer: Offer architectures that meet their controls rather than arguing.

What I would check:

  1. Exactly which data classes are restricted, and why (regulation, contract, policy).
  2. Whether a cloud model in their approved cloud, over private networking in-region, meets the requirement.
  3. Whether PII can be masked or tokenised before the model sees it.
  4. Whether a self-hosted open-weight model is needed for some data classes.

Production consideration: Document the data flow, retention, encryption and provider terms on one page for security sign-off. Often the outcome is a mix: masked data to a cloud model, and the most sensitive tasks on a self-hosted model.

91. Retrieval quality dropped after you re-indexed with a new embedding model. Why?

Answer: Common causes are a partial re-index (old and new vectors mixed), a query-time model that differs from the index-time model, changed dimensions or distance metric, or a new model that is simply worse on your domain vocabulary.

What I would check:

  1. That every vector in the index came from the same model version.
  2. That the query embedding uses the same model and any required prefixes or instructions.
  3. The distance metric and normalisation settings.
  4. Retrieval metrics on the eval set, old index vs new.

Production consideration: Re-index into a new index and switch over with a blue-green swap only after retrieval evals pass.

92. An agent keeps calling the wrong tool. How do you fix it?

Answer: Tool selection depends mostly on tool names, descriptions and how many tools the agent has to choose from.

What I would check:

  1. Overlapping or vague tool descriptions.
  2. Too many tools in context. Can you group them, or select tools dynamically?
  3. Whether the system prompt explains when to use each tool, with examples.
  4. Tool-selection accuracy per intent in the eval suite.

Production consideration: Rewrite descriptions to say when to use each tool and when not to, route by intent to a smaller tool set, and add tool-choice tests to CI.

93. A hospital wants an assistant that summarises patient records for doctors. What are your concerns?

Answer: Clinical safety, privacy and accountability come first. The assistant supports clinicians and never replaces their judgement.

What I would check:

  1. Data residency, consent and access rules for health records.
  2. Omission risk: a summary that drops an allergy is a serious failure.
  3. Whether each summary statement links back to the source record.
  4. Evaluation by clinicians on representative records, including edge cases.

Production consideration: Show sources inline, flag critical fields (allergies, medications) from structured data rather than generated text, log usage for audit, and keep the clinician as the decision-maker.

94. Product wants a fully autonomous agent that approves vendor payments. How do you respond?

Answer: Propose graduated autonomy: the agent prepares the payment and a human approves it. Autonomy grows only as measured accuracy and controls justify it.

What I would check:

  1. Existing approval controls and segregation-of-duties rules.
  2. Fraud risks such as manipulated invoices and injected instructions in documents.
  3. The cost of a wrong approval versus the time saved.
  4. Audit requirements.

Production consideration: Start with recommend-only mode. Consider auto-approval only for low-value, policy-matched, known-vendor cases, with limits enforced in code and every decision logged.

95. Your eval scores look great, but users are unhappy. What is happening?

Answer: The eval set no longer reflects real usage, or the metrics don't measure what users care about.

What I would check:

  1. Compare the eval questions with the real query distribution from logs.
  2. Read a sample of low-rated conversations and label the failure types.
  3. Check whether the judge agrees with human ratings.
  4. Check for UX issues: latency, formatting, missing follow-ups.

Production consideration: Refresh the eval set from production traffic regularly and add user-centred criteria such as actionability and completeness.

96. A document in the knowledge base contains hidden text telling the AI to email data externally. What do you do?

Answer: This is indirect prompt injection. Remove the document, check whether any action was taken, and harden the system so injected text cannot trigger exfiltration.

What I would check:

  1. Logs: did any tool call follow retrieval of that document?
  2. Whether the agent has any outbound tool (email, HTTP) it does not need.
  3. How the document entered the corpus, and who can add content.
  4. Other documents with similar patterns.

Production consideration: Remove unneeded tools, restrict outbound destinations, require approval for external sends, scan ingested content for injection patterns, and add the case to red-team tests.

97. Leadership asks whether the AI assistant is "worth it". How do you answer?

Answer: Answer with a measured comparison against the baseline process, not anecdotes.

What I would check:

  1. The baseline metrics recorded before launch (handling time, resolution rate, backlog).
  2. Adoption: active users and repeat usage.
  3. Quality and risk: error rate, escalations, incidents.
  4. Full cost: models, infrastructure, support and review time.

Production consideration: Define value metrics before building. If you didn't, measure a comparison group now.

98. Your team wants to fine-tune a model because "RAG isn't accurate enough". How do you evaluate the proposal?

Answer: First find out what kind of inaccuracy it is. Fine-tuning rarely fixes retrieval or knowledge-freshness problems.

What I would check:

  1. Whether the failures are retrieval misses, unfaithful generation, or format and style issues.
  2. Whether better chunking, hybrid search or reranking closes the gap.
  3. Whether there is enough clean, labelled training data.
  4. The ongoing cost of retraining as the base model or data changes.

Production consideration: Fine-tune only for a measured behavioural gap, compare against the improved RAG baseline on the same eval set, and plan how the fine-tuned model will be versioned and retrained.

99. A GCC team in Hyderabad must support the same assistant for business units in three countries. What changes?

Answer: Data residency, regulation, language and local policy differences turn one assistant into a configurable platform.

What I would check:

  1. Residency and AI regulations per country, for example the EU AI Act for European units.
  2. Whether each region needs its own model endpoint and index.
  3. Country-specific policies and documents, and how they are scoped.
  4. Support hours, ownership and escalation per region.

Production consideration: Deploy per region from the same code with region-specific configuration, scope retrieval by country, and keep separate eval sets per business unit.

100. On day one at a new client, you inherit an AI system with no tests, no monitoring and unhappy users. What are your first two weeks?

Answer: Make the system observable first, then measurable, then improve it. Don't rewrite anything before you understand it.

What I would check:

  1. Architecture, data sources, permissions and the riskiest capabilities (tools that write).
  2. Tracing and cost logging on every request, added in the first days.
  3. A first eval set of real failing and common questions, built with users.
  4. Quick security risks: over-privileged keys, missing permission filters.

Production consideration: Fix critical security gaps immediately. Then make one measured improvement at a time and share before-and-after results with stakeholders weekly. This is the everyday work of Forward Deployed Engineers. Our FDE interview questions cover that role in depth.

Key takeaways

  • AI engineer interviews in 2026 test production judgement: evaluation, security, cost and failure handling, not only model knowledge.
  • Treat the LLM as an unreliable dependency. Validate outputs, enforce limits in code, and keep authorisation out of prompts.
  • Most RAG and agent quality problems are context problems. Diagnose the failing stage from traces before changing prompts.
  • Evaluation is the skill that separates good candidates: offline gates in CI, plus online monitoring that feeds new test cases.
  • Know one cloud deeply and the others at the map level, using current names: Bedrock AgentCore, Microsoft Foundry, Gemini Enterprise Agent Platform.
  • For scenario questions, use the same pattern: contain, check the trace, fix the root cause, add a regression test.

Interview preparation checklist

  • Build one RAG project with permission-aware retrieval, citations and an eval set you can show.
  • Build one agent with at least two real tools (for example a ticketing API), step limits and an approval step.
  • Be able to draw three architectures from memory: knowledge assistant, action-taking agent, document extraction.
  • Prepare a cost estimate for one of your projects, with stated assumptions.
  • Practise explaining prompt injection defences and RAG access control without notes.
  • Review one cloud's AI services in its current official documentation.
  • Prepare two STAR stories: one quality failure you fixed, one security or cost issue.
  • Practise five scenario questions aloud with the contain, diagnose, fix, prevent structure.
  • Read the specialist guide for your target role: GenAI, agentic, RAG, cloud or MLOps.

FAQ

What skills are required to become an AI engineer in 2026?

Strong Python and API engineering, LLM fundamentals, RAG, tool calling and agents, evaluation, one major cloud, containers and CI/CD, and the basics of AI security. Communication with business and security teams matters as much as code.

Do I need a machine learning degree to become an AI engineer?

No. Most AI engineering builds on pre-trained models, so software engineering skills plus a working understanding of how models behave are usually enough to start. A portfolio of working projects carries more weight than a degree title.

How should I prepare for an AI engineer interview?

Build two or three end-to-end projects, practise explaining their architecture and failures, rehearse system design and scenario questions aloud, and review current cloud AI service names in official documentation.

Is AI engineering a good career for software developers?

For developers who enjoy building products, yes. Organisations need engineers who can integrate AI with real systems, data and security controls, and existing backend skills transfer directly.

What is the difference between an AI engineer and a GenAI developer?

The titles overlap heavily. GenAI developer usually stresses building with generative models. AI engineer is often broader and can include evaluation, deployment and sometimes traditional ML.

Which cloud should I learn for AI engineering?

Learn the one most used by the employers you are targeting, and learn it in depth, then understand the equivalent services on the other two at a high level. The core concepts carry across clouds.

Are coding rounds still part of AI engineer interviews?

Usually yes. Expect practical Python tasks such as calling an API, parsing structured output, writing an async pipeline or adding tests, often alongside a system design round.

How important are projects for freshers applying to AI engineer roles?

Very important. A deployed project with an eval set, a README explaining trade-offs and a short demo shows interviewers far more than certificates alone.

What do interviewers look for in AI system design rounds?

Clear requirements gathering, sensible component choices, permission and security handling, an evaluation plan, cost and latency awareness, and honest discussion of failure modes and trade-offs.

Ready to prepare with guided projects instead of on your own? Cloudsoft's APEX program for AI, ML, cloud and security builds the breadth this guide covers. If you want to work on customer-facing production delivery, the FDE PRO Forward Deployed Engineer course runs for 12 weeks with five enterprise projects, a simulated customer capstone, and placement support until you're placed. If you are newer to the field, start with our AI, GenAI and Agentic AI course. Classroom training is in Ameerpet, or live online. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us