New batches starting this week Β· Limited seats

Python Interview Questions for AI Engineers 2026 (60 Questions)

60 Python interview questions for AI engineering roles with model answers and tested snippets: async concurrency, Pydantic v2, FastAPI streaming, testing LLM code, the GIL and live-coding prompts.

Python interview questions for AI engineers 2026: 60 questions on async, Pydantic, FastAPI, generators, testing and live coding
Last updated Β· 46 min read Β· 10,099 words

Python interview questions for AI roles in 2026 are mostly backend questions with a model in the middle. Interviewers want to see that you can stream tokens with a generator, cap concurrent API calls with a semaphore, validate model output with Pydantic, expose it through FastAPI, test it without a live model and explain why the GIL does or does not matter for your workload. This guide collects 60 high-value, commonly asked questions on exactly that, with model answers, short illustrative snippets and six live-coding prompts written out in full.

If you are still deciding which Python to learn in the first place, start with the companion article on Python for AI engineers; this page assumes you have written some Python and want to be ready for the interview.

How to use this guide

  • Freshers and career switchers: interviewers test core language behaviour (mutable defaults, generators, exceptions, dataclasses) and whether you can read a short async function. Questions 1 to 19 are the base.
  • Mid-level developers: expect HTTP timeouts and retries, Pydantic v2, FastAPI streaming and dependency injection, pytest with mocks and logging. Questions 20 to 42 carry most of the weight.
  • Senior and Forward Deployed Engineer style roles: live coding plus production scenarios: rate limits across workers, memory growth, proxies buffering streams, CPU-bound work in an async service. Questions 43 to 60 matter most.
  • Every snippet is illustrative, written for Python 3.11 or later, and was run before publishing. Library APIs change between releases, so check the current documentation for FastAPI, Pydantic and httpx before you copy anything into a real project.

Contents

Core Python for AI work

1. Which built-in data structures do you reach for in LLM application code, and why?

Answer: dict for JSON-shaped payloads and constant-time lookups (document ID to metadata), list for ordered sequences such as chat messages, set for de-duplication and membership checks (have I already embedded this chunk hash?), and tuple for small immutable records or dictionary keys. From collections: deque(maxlen=n) for a sliding window of recent chat turns, Counter for token or label frequencies in an evaluation run, defaultdict(list) for grouping chunks by document, and OrderedDict when you need explicit move-to-end behaviour, as in the LRU cache in question 47.

Interview tip: say why, not just what. "A set because membership is average O(1) and I check every chunk against it" is the answer they want. Also mention that a plain dict preserves insertion order, so dict.fromkeys(items) is a quick order-preserving de-duplicate.

2. What is the mutable default argument bug, and how does it show up in chat code?

Answer: default values are evaluated once, when the function is defined, not on each call. A list default is therefore shared across every call that omits the argument.

def add_turn(msg, history=[]):      # bug: one shared list
    history.append(msg)
    return history

def add_turn_fixed(msg, history=None):
    history = [] if history is None else history
    history.append(msg)
    return history

In an LLM service the buggy version means one user's conversation history leaks into the next user's prompt, which is both a correctness bug and a data-privacy incident. The fix is None as the sentinel, or field(default_factory=list) in a dataclass and a plain [] default in Pydantic, which copies defaults per instance.

3. When do you use a list comprehension versus a generator expression?

Answer: a list comprehension builds the whole list in memory, which you want when you need to index it, iterate it twice or know its length. A generator expression produces items lazily, which is the right choice when you feed a single consumer such as sum(), any() or a file writer.

texts = ["  Hello ", "", " World  ", "Hello"]
clean = [s for t in texts if (s := t.strip())]
unique = list(dict.fromkeys(clean))    # dedupe, keep order
by_id = {i: t for i, t in enumerate(unique)}
total_chars = sum(len(t) for t in unique)   # no list built
print(unique, total_chars)   # ['Hello', 'World'] 10

The walrus operator avoids calling strip() twice. Keep comprehensions to one condition and at most two loops; beyond that, a normal loop is easier to review.

4. How do generators help with streaming model responses?

Answer: a generator lets you transform a stream piece by piece without holding all of it. Model APIs stream responses as a sequence of events; a generator can parse each line, skip keep-alives, stop on a terminal marker and yield only the text, while the caller forwards each piece to the browser.

import json
from typing import Iterable, Iterator

def text_deltas(lines: Iterable[str]) -> Iterator[str]:
    """Turn 'data: {...}' stream lines into text pieces."""
    for line in lines:
        if not line.startswith("data: "):
            continue
        payload = line[len("data: "):]
        if payload == "[DONE]":
            return
        yield json.loads(payload).get("text", "")

lines = ['data: {"text": "Hel"}', "", 'data: {"text": "lo"}',
         "data: [DONE]"]
print("".join(text_deltas(lines)))   # Hello

The event shape here is simplified; each provider and SDK has its own format. A return inside a generator ends iteration cleanly. The same pattern with async def and yield gives an async generator, which is what you hand to a streaming HTTP response (question 29).

5. How would you batch a large iterable for an embeddings API?

Answer: embedding APIs accept many inputs per request, so you send batches instead of one call per chunk. Use itertools.islice over an iterator so the source can itself be a generator reading from disk.

from itertools import islice

def batched(items, size):
    it = iter(items)
    while batch := list(islice(it, size)):
        yield batch

for b in batched(range(10), 4):
    print(b)   # [0, 1, 2, 3] / [4, 5, 6, 7] / [8, 9]

Python 3.12 added itertools.batched, which yields tuples; mention it, but be ready to write the version above because interviewers often ask for it. In practice batch size is limited by the provider's maximum inputs per request and by total tokens per request, so a production batcher counts tokens as well as items.

6. Write a timing decorator. What does functools.wraps do?

Answer: a decorator is a function that takes a function and returns a replacement. functools.wraps copies the original's name, docstring and other metadata onto the wrapper, so logs, tracebacks, pytest output and FastAPI's introspection still see the real function.

import functools
import logging
import time

log = logging.getLogger(__name__)

def timed(fn):
    @functools.wraps(fn)
    def wrapper(*args, **kwargs):
        start = time.perf_counter()
        try:
            return fn(*args, **kwargs)
        finally:
            ms = (time.perf_counter() - start) * 1000
            log.info("%s took %.1f ms", fn.__name__, ms)
    return wrapper

@timed
def rerank(docs: list[str]) -> list[str]:
    return sorted(docs, key=len)

The try/finally means the duration is logged even when the call raises. This version only works for regular functions; an async def function needs an async wrapper that awaits the call, otherwise you time how long it took to create the coroutine. A follow-up is usually "now write a retry decorator", covered in question 48.

7. What is a context manager, and where do you use one in AI services?

Answer: an object that defines setup and teardown around a with block, so cleanup runs even on exceptions. Files, database connections, HTTP clients, locks, temporary directories and tracing spans are the usual cases.

import time
from contextlib import contextmanager

@contextmanager
def span(name: str, sink: list):
    start = time.perf_counter()
    try:
        yield
    finally:          # runs even if the block raises
        sink.append((name, time.perf_counter() - start))

timings = []
with span("retrieve", timings):
    sum(range(10_000))
print(timings[0][0])   # retrieve

contextlib.contextmanager turns a generator into a context manager; asynccontextmanager does the same for async resources, and FastAPI's lifespan hook (question 31) uses exactly that.

8. How do you use type hints in an AI codebase, and does Python enforce them?

Answer: Python does not enforce hints at runtime; they are checked by tools such as mypy or Pyright and used by libraries like Pydantic and FastAPI, which read them to validate data. Useful constructs: list[str], X | None, Literal["low", "high"] for fixed labels, TypedDict for dict-shaped messages, and Protocol for structural interfaces.

from typing import Protocol

class Embedder(Protocol):
    def embed(self, texts: list[str]) -> list[list[float]]: ...

class FakeEmbedder:          # no inheritance needed
    def embed(self, texts: list[str]) -> list[list[float]]:
        return [[float(len(t))] for t in texts]

def index(docs: list[str], embedder: Embedder) -> int:
    return len(embedder.embed(docs))

print(index(["a", "bb"], FakeEmbedder()))   # 2

A Protocol is the clean way to swap a real embedding client for a fake in tests, or one provider for another, without an inheritance hierarchy.

9. Dataclass, TypedDict or Pydantic model: how do you choose?

Answer: use a dataclass for internal records you construct yourself (a chunk, a retrieval result); it is fast and does no validation. Use a TypedDict when the data must stay a plain dict, such as message lists passed straight to an SDK. Use a Pydantic model at trust boundaries: HTTP requests, configuration, and anything a model generated.

from dataclasses import dataclass, field

@dataclass(frozen=True, slots=True)
class Chunk:
    doc_id: str
    text: str
    page: int
    metadata: dict = field(default_factory=dict)

c = Chunk("policy-7", "Claims must be filed...", page=3)
print(c.page)          # 3; c.page = 4 raises FrozenInstanceError

frozen=True prevents accidental mutation and slots=True (3.10+) saves memory when you hold many objects. Note that a frozen dataclass with a dict field still cannot be hashed, because the dict is unhashable.

10. How do you design exceptions around model calls?

Answer: create a small hierarchy that encodes what the caller should do, not just what went wrong. The key split is retryable (rate limits, timeouts, server errors) versus non-retryable (bad request, authentication, content policy, schema errors).

class LLMError(Exception):
    """Base class for model-call failures."""

class RetryableLLMError(LLMError):
    """Rate limit, timeout or 5xx: safe to retry."""

class BadRequestLLMError(LLMError):
    """Bad prompt or schema: retrying will not help."""

def error_for(status: int) -> type[LLMError]:
    if status == 429 or status >= 500:
        return RetryableLLMError
    return BadRequestLLMError

try:
    raise error_for(503)("upstream unavailable")
except RetryableLLMError as exc:
    print("retry later:", exc)

Wrap SDK exceptions at the boundary with raise RetryableLLMError(...) from exc so the original traceback is kept. Never use a bare except:, which also catches KeyboardInterrupt and, in async code, can interfere with cancellation.

11. Explain shallow versus deep copy with a message list.

Answer: a shallow copy creates a new outer container but shares the inner objects; a deep copy duplicates everything recursively.

import copy

base = [{"role": "system", "content": "Be concise."}]
shallow = list(base)              # new list, same dicts
shallow[0]["content"] = "Be verbose."
print(base[0]["content"])         # Be verbose.  (shared)

deep = copy.deepcopy(base)
deep[0]["content"] = "Be concise."
print(base[0]["content"])         # Be verbose.  (unchanged)

Real-world example: a shared system-prompt template copied with list(template) and then edited per tenant silently changes the template for every later request. Either deep-copy, or better, build new message dicts per request instead of mutating shared ones.

12. Why do LLM client wrappers use *args, **kwargs and keyword-only parameters?

Answer: **kwargs lets a thin wrapper pass provider options (temperature, max tokens, stop sequences) through without listing each one. Keyword-only parameters, declared after a bare * as in def complete(prompt, *, model, timeout=30), force callers to name important settings, so a call site reads clearly and arguments cannot be swapped by position. The risk with **kwargs is that typos pass silently; validate the allowed keys, or use a typed options object, at the boundary where configuration enters.

Async and concurrency

13. What does async/await actually give you in an AI service?

Answer: concurrency for I/O-bound work on a single thread. While one coroutine waits for a model API, a vector database or another HTTP service, the event loop runs others. LLM applications spend most of their time waiting on the network, so async lets one process handle many in-flight requests or fan out many embedding calls. It does not make CPU work faster: code that parses, tokenises or computes on the event loop blocks everything else. await is the only point where a coroutine yields control, so a coroutine without awaits on slow operations behaves like synchronous code.

14. How do you limit concurrency when calling a rate-limited API?

Answer: wrap each call in an asyncio.Semaphore. Launching thousands of tasks with no limit triggers rate-limit errors, exhausts connection pools and can get your key throttled.

import asyncio

async def embed_one(text: str) -> list[float]:
    await asyncio.sleep(0.01)          # stand-in for an API call
    return [float(len(text))]

async def embed_all(texts: list[str], limit: int = 8):
    sem = asyncio.Semaphore(limit)

    async def guarded(t: str):
        async with sem:                # at most `limit` in flight
            return await embed_one(t)

    return await asyncio.gather(*(guarded(t) for t in texts))

print(len(asyncio.run(embed_all(["chunk"] * 50))))   # 50

A semaphore caps concurrency (calls in flight), not rate (calls or tokens per minute). Providers usually enforce rates, so production code often combines a semaphore with a token bucket (question 43). For very large inputs, also avoid creating every coroutine up front; feed a queue of workers instead (question 19).

15. What is the difference between asyncio.gather and asyncio.TaskGroup?

Answer: gather runs awaitables concurrently and returns results in order. With return_exceptions=True failures come back as values next to successes; without it, the first exception propagates but the other tasks keep running in the background. TaskGroup (Python 3.11+) gives structured concurrency: if any task fails, the group cancels the rest, waits for them, and raises an ExceptionGroup, which you handle with except*.

import asyncio

async def fetch(name: str, fail: bool = False) -> str:
    await asyncio.sleep(0.01)
    if fail:
        raise ValueError(name)
    return name

async def main():
    # gather: failures come back as values, others finish
    res = await asyncio.gather(
        fetch("a"), fetch("b", fail=True),
        return_exceptions=True)
    print(res)        # ['a', ValueError('b')]

    # TaskGroup (3.11+): a failure cancels the siblings
    try:
        async with asyncio.TaskGroup() as tg:
            tg.create_task(fetch("c"))
            tg.create_task(fetch("d", fail=True))
    except* ValueError as eg:
        print("failed:", eg.exceptions)

asyncio.run(main())

Interview tip: choose by failure semantics. Embedding a batch of chunks where some may fail and you will retry those individually suits gather(..., return_exceptions=True). Several calls that must all succeed for one answer (retrieve, look up the customer, check permissions) suit TaskGroup, because there is no point letting the others run once one has failed.

16. How do you put a timeout on an async operation?

Answer: use async with asyncio.timeout(seconds) (3.11+) or asyncio.wait_for. On expiry the inner task is cancelled and TimeoutError is raised.

import asyncio

async def slow_tool() -> str:
    await asyncio.sleep(5)
    return "done"

async def main():
    try:
        async with asyncio.timeout(0.1):    # Python 3.11+
            await slow_tool()
    except TimeoutError:
        print("tool timed out; return a fallback")

asyncio.run(main())

Have two layers: the HTTP client's own connect and read timeouts (question 20), and an overall deadline for the whole request, including retries, so a user never waits longer than your service-level target. An agent tool call that times out should return a structured error the model or the orchestrator can act on, not hang the loop.

17. What does "blocking the event loop" mean, and how do you fix it?

Answer: calling something slow and synchronous inside a coroutine, such as time.sleep, requests.post, a synchronous SDK client, a synchronous database driver or heavy CPU work. Every other request on that event loop stalls until it returns. Fixes: use async libraries (httpx, async database drivers, the provider's async client), or move unavoidable blocking calls to a thread.

import asyncio
import time

def legacy_sdk_call(prompt: str) -> str:
    time.sleep(0.2)                  # blocking network call
    return prompt.upper()

async def handler(prompt: str) -> str:
    # runs in a worker thread; the event loop stays free
    return await asyncio.to_thread(legacy_sdk_call, prompt)

print(asyncio.run(handler("hi")))    # HI

asyncio.to_thread suits blocking I/O. For CPU-heavy work, threads do not help much under the GIL, so use a process pool (question 60). Enabling asyncio debug mode logs callbacks that take too long, which is a quick way to find the offender.

18. How do async generators work, and what happens when a streaming client disconnects?

Answer: an async def function containing yield is an async generator; you consume it with async for. Each step can await I/O, so it is the natural shape for relaying model tokens. When a client disconnects mid-stream, the server framework stops iterating, typically by cancelling the task, so your generator gets a CancelledError or GeneratorExit at its current await or yield. Put cleanup (closing the upstream stream, recording partial usage for cost tracking) in finally, and do not swallow CancelledError: re-raise it after cleanup. Otherwise you keep paying for tokens nobody will read.

19. Design a producer-consumer ingestion pipeline with asyncio.

Answer: an asyncio.Queue with a maxsize gives backpressure: the producer waits when workers fall behind, so memory stays bounded however many documents you ingest.

import asyncio

async def worker(q: asyncio.Queue, out: list):
    while True:
        doc = await q.get()
        try:
            out.append(doc.lower())  # parse, chunk, embed here
        finally:
            q.task_done()

async def ingest(docs: list[str], n_workers: int = 4):
    q, out = asyncio.Queue(maxsize=100), []
    workers = [asyncio.create_task(worker(q, out))
               for _ in range(n_workers)]
    for d in docs:
        await q.put(d)        # waits when the queue is full
    await q.join()            # every item processed
    for w in workers:
        w.cancel()
    return out

print(len(asyncio.run(ingest(["DOC"] * 10))))   # 10

In a real pipeline each worker parses, chunks, embeds (through the rate limiter) and upserts; failures go to a dead-letter list with the document ID so a rerun processes only those. Python 3.13 added Queue.shutdown() as a cleaner way to stop workers; cancelling them after join() works on all supported versions.

HTTP clients, timeouts and retries

20. How should you configure an HTTP client for model and tool calls?

Answer: create one client per process and reuse it, set explicit timeouts and size the connection pool. Creating a new client per request throws away connection reuse, so every call pays for a new TCP and TLS handshake.

import httpx

TIMEOUT = httpx.Timeout(60.0, connect=5.0)
LIMITS = httpx.Limits(max_connections=20,
                      max_keepalive_connections=10)

def make_client(base_url: str) -> httpx.AsyncClient:
    # create once at startup, reuse for every request
    return httpx.AsyncClient(base_url=base_url,
                             timeout=TIMEOUT, limits=LIMITS)

Be explicit about which timeout is which: connect should be short, while read must allow for slow generation, and for streaming it is the gap between chunks rather than the total time. requests is fine in scripts, but it has no async API and no timeout unless you pass one, which is a classic review finding. Many provider SDKs wrap httpx internally and accept timeout and retry settings; know where yours are configured.

21. Which failures do you retry, and how?

Answer: retry transient failures with exponential backoff and jitter, cap the attempts and respect the server's hints. Do not retry client errors that will fail the same way again.

FailureRetry?Notes
429 rate limitedYesHonour Retry-After if present; slow the whole client, not just this call
500, 502, 503, 504Yes, limitedBackoff with jitter; open a circuit breaker if failures persist
Connect or read timeoutYes, if idempotentA timed-out write may have succeeded; use idempotency keys
400 bad request, 422NoFix the prompt, schema or payload; log it
401, 403NoCredential or permission problem; alert
Output fails validationOnce, differentlyRe-ask with the validation error, not a blind retry

Retries multiply load during an outage, so they belong in one place (the client wrapper or an LLM gateway), not stacked at the SDK, wrapper and caller levels at once.

22. How do you consume a streaming (server-sent events) response with httpx?

Answer: use client.stream(...) as an async context manager and iterate lines, so you process each event as it arrives and the connection is released when the block exits.

import httpx

async def stream_events(client: httpx.AsyncClient,
                        payload: dict):
    async with client.stream("POST", "/v1/generate",
                             json=payload) as resp:
        resp.raise_for_status()
        async for line in resp.aiter_lines():
            if line.startswith("data: "):
                yield line[len("data: "):]

Server-sent events are lines of data: ... separated by blank lines; lines starting with a colon are comments, often keep-alives. Production parsers also handle multi-line data fields and event: types, and many SDKs provide a ready-made stream iterator, so use it where one exists. Test the parser against a fake transport (httpx.MockTransport) rather than a live endpoint.

Pydantic v2 and validating LLM output

23. What changed between Pydantic v1 and v2 that you need to know?

Answer: v2 rewrote validation in a Rust core, which made it much faster, and renamed most of the public API. Interviewers check that you are not writing v1 code from old tutorials.

v1v2
Model.parse_obj(d)Model.model_validate(d)
Model.parse_raw(s)Model.model_validate_json(s)
m.dict(), m.json()m.model_dump(), m.model_dump_json()
@validator, @root_validator@field_validator, @model_validator
class Config:model_config = ConfigDict(...)
Model.schema()Model.model_json_schema()

Settings classes moved to the separate pydantic-settings package. TypeAdapter validates types that are not models, such as list[Chunk] or a union.

24. How do you validate a model's structured output with Pydantic?

Answer: define the expected shape with types and constraints, parse the raw text with model_validate_json, and treat ValidationError as a normal outcome with a defined fallback.

from typing import Literal
from pydantic import BaseModel, Field, ValidationError

class ClaimSummary(BaseModel):
    policy_number: str = Field(pattern=r"^POL-\d{6}$")
    claim_type: Literal["motor", "health", "property"]
    urgency: int = Field(ge=1, le=5)
    summary: str = Field(max_length=500)

raw = ('{"policy_number": "POL-123456", "claim_type": "motor",'
       ' "urgency": "3", "summary": "Rear-ended at a signal"}')
try:
    claim = ClaimSummary.model_validate_json(raw)
    print(claim.urgency + 1)     # 4: "3" coerced (lax mode)
except ValidationError as exc:
    print(exc.errors())          # field, type and message

Literal restricts labels, Field adds bounds and patterns, and exc.errors() gives machine-readable details you can log or send back to the model in a single repair attempt. Even when the provider supports schema-constrained output, still validate: constraints such as patterns and business rules are not always enforced by the provider. The provider-side mechanics are in function calling and structured outputs.

Real-world example: consider an insurer triaging claim emails. If validation fails twice, the claim goes to a human queue with the raw output attached, never into the claims system with a guessed policy number.

25. When do you use field_validator versus model_validator?

Answer: field_validator checks or normalises one field; model_validator(mode="after") checks rules that span fields, on the fully built object.

from pydantic import BaseModel, field_validator, model_validator

class Answer(BaseModel):
    answer: str
    citations: list[str] = []
    confidence: float

    @field_validator("answer")
    @classmethod
    def not_blank(cls, v: str) -> str:
        v = v.strip()
        if not v:
            raise ValueError("empty answer")
        return v

    @model_validator(mode="after")
    def cited_if_confident(self):
        if self.confidence > 0.5 and not self.citations:
            raise ValueError("confident answers need citations")
        return self

Raise ValueError inside validators; Pydantic wraps it into a ValidationError with the location. Keep validators free of I/O. Checking that cited document IDs actually exist in your retrieval results belongs in application code, where you have that context.

26. How do strict mode, extra="forbid" and JSON Schema generation help with LLMs?

Answer: by default Pydantic is lax and coerces "3" to 3; strict mode rejects it. extra="forbid" rejects unexpected keys, which catches a model inventing fields. model_json_schema() produces the JSON Schema you pass to a provider's structured-output or tool-calling feature, so the schema and the validator come from one definition.

from pydantic import BaseModel, ConfigDict

class Ticket(BaseModel):
    model_config = ConfigDict(extra="forbid", strict=True)
    title: str
    priority: int

schema = Ticket.model_json_schema()
print(schema["required"])              # ['title', 'priority']
print(schema["additionalProperties"])  # False

Lax mode is often the pragmatic choice for model output, since harmless coercions are fine, while extra="forbid" is almost always worth having. Providers support different subsets of JSON Schema, so check which keywords yours honours before relying on a constraint.

27. How would you parse several possible tool calls into the right type?

Answer: use a discriminated union: each model has a Literal tag field, and Pydantic picks the class from the tag instead of trying each one in turn.

from typing import Annotated, Literal, Union
from pydantic import BaseModel, Field, TypeAdapter

class Search(BaseModel):
    tool: Literal["search"]
    query: str

class CreateTicket(BaseModel):
    tool: Literal["create_ticket"]
    title: str
    priority: int = 3

ToolCall = Annotated[Union[Search, CreateTicket],
                     Field(discriminator="tool")]
adapter = TypeAdapter(ToolCall)

call = adapter.validate_python({"tool": "search", "query": "VPN"})
print(type(call).__name__)    # Search

Errors are clearer because Pydantic reports against the one matching model, and dispatch becomes a match on the type. This is the typed core of a hand-written agent loop; the AI agent developer interview questions guide builds the full loop around it.

FastAPI services

28. Write a FastAPI endpoint with request and response models.

Answer: declare Pydantic models for the body and the response; FastAPI validates input (returning 422 on failure), filters output to the response model and documents both in the generated OpenAPI schema.

from fastapi import FastAPI
from pydantic import BaseModel, Field

app = FastAPI()

class AskRequest(BaseModel):
    question: str = Field(min_length=3, max_length=2000)
    top_k: int = Field(default=5, ge=1, le=20)

class AskResponse(BaseModel):
    answer: str
    sources: list[str]

@app.post("/ask", response_model=AskResponse)
async def ask(req: AskRequest) -> AskResponse:
    # retrieve req.top_k chunks, call the model, validate
    return AskResponse(answer="stub", sources=["doc-1"])

response_model also acts as an output filter, so internal fields such as raw prompts or scores do not leak to clients. Point out the input bounds: a max_length on the question is a cheap guard against huge prompts that cost money.

29. How do you stream an LLM response from FastAPI?

Answer: return a StreamingResponse wrapping an async generator, with the text/event-stream media type for server-sent events.

from fastapi import FastAPI
from fastapi.responses import StreamingResponse

app = FastAPI()

async def token_events(question: str):
    for piece in ["Checking", " the", " policy", "..."]:
        yield f"data: {piece}\n\n"     # one SSE event
    yield "data: [DONE]\n\n"

@app.get("/ask/stream")
async def ask_stream(q: str):
    return StreamingResponse(token_events(q),
                             media_type="text/event-stream")

Recent FastAPI releases also include an EventSourceResponse helper for SSE; check the current docs for its exact usage. Points interviewers probe: send an explicit end event, emit an error event rather than silently closing if the upstream fails mid-stream, and remember that validation of a complete JSON object cannot happen until the stream ends, so streamed structured output needs a final validation step.

30. How does dependency injection work in FastAPI, and why does it matter for testing?

Answer: a dependency is a function declared with Depends; FastAPI calls it per request and passes the result into your endpoint. Use it for the LLM client, database sessions, the current user or tenant, and settings.

from typing import Annotated
from fastapi import Depends, FastAPI, Header
from pydantic import BaseModel

app = FastAPI()

class LLMClient:
    async def complete(self, prompt: str) -> str:
        raise NotImplementedError    # real provider call

def get_llm() -> LLMClient:
    return LLMClient()

def get_tenant(x_tenant_id: Annotated[str, Header()]) -> str:
    return x_tenant_id               # missing header -> 422

class SummariseIn(BaseModel):
    text: str

@app.post("/summarise")
async def summarise(
    body: SummariseIn,
    llm: Annotated[LLMClient, Depends(get_llm)],
    tenant: Annotated[str, Depends(get_tenant)],
):
    summary = await llm.complete(body.text)
    return {"tenant": tenant, "summary": summary}

The payoff is testability: app.dependency_overrides swaps the real client for a fake without patching imports.

from fastapi.testclient import TestClient

class FakeLLM:
    async def complete(self, prompt: str) -> str:
        return "short summary"

def test_summarise_uses_injected_llm():
    app.dependency_overrides[get_llm] = lambda: FakeLLM()
    client = TestClient(app)
    r = client.post("/summarise", json={"text": "long doc"},
                    headers={"X-Tenant-Id": "acme"})
    assert r.json()["summary"] == "short summary"
    app.dependency_overrides.clear()

Dependencies can also be generators that yield a resource and clean it up after the response, the standard pattern for database sessions.

31. async def or def for a FastAPI endpoint, and where do shared clients live?

Answer: FastAPI runs def endpoints in a thread pool and async def endpoints on the event loop. Use async def only if everything slow inside is awaited; a synchronous SDK call inside async def blocks the whole worker (scenario 50). Shared clients are created once in the lifespan handler and closed on shutdown.

from contextlib import asynccontextmanager
import httpx
from fastapi import FastAPI, Request

@asynccontextmanager
async def lifespan(app: FastAPI):
    app.state.http = httpx.AsyncClient(timeout=30.0)
    yield                              # app serves requests
    await app.state.http.aclose()      # clean shutdown

app = FastAPI(lifespan=lifespan)

@app.get("/health")
async def health(request: Request):
    return {"http_ready": not request.app.state.http.is_closed}

Use BackgroundTasks only for small post-response work such as writing an audit record. Anything long, retryable or important (re-indexing, batch summarisation) belongs on a real queue with its own workers.

NumPy, pandas, files and PDFs

32. Use NumPy to find the top-k most similar embeddings to a query.

Answer: normalise the vectors, compute all scores with one matrix-vector product, then use argpartition to get the top k without sorting everything.

import numpy as np

def top_k(query: np.ndarray, matrix: np.ndarray, k: int = 3):
    q = query / np.linalg.norm(query)
    m = matrix / np.linalg.norm(matrix, axis=1, keepdims=True)
    scores = m @ q                       # one score per row
    k = min(k, len(scores))
    idx = np.argpartition(-scores, k - 1)[:k]
    idx = idx[np.argsort(-scores[idx])]  # sort only the top k
    return idx, scores[idx]

In production, store vectors already normalised so a dot product is the cosine similarity, and let the vector database do this at scale. Knowing the brute-force version matters for small corpora, tests and evaluation scripts. For the concepts behind the numbers, see embeddings explained.

33. How do you use pandas to compare two prompt versions in an evaluation run?

Answer: load results into a DataFrame with one row per question per variant, aggregate per variant, and pivot to find individual regressions, which an average can hide.

import pandas as pd

runs = pd.DataFrame({
    "question_id": [1, 2, 3, 1, 2, 3],
    "prompt": ["v1"] * 3 + ["v2"] * 3,
    "correct": [1, 1, 0, 1, 0, 1],
    "latency_ms": [820, 950, 700, 880, 990, 760],
})
summary = runs.groupby("prompt").agg(
    accuracy=("correct", "mean"),
    p95_ms=("latency_ms", lambda s: s.quantile(0.95)),
)
wide = runs.pivot(index="question_id", columns="prompt",
                  values="correct")
regressed = wide[(wide["v1"] == 1) & (wide["v2"] == 0)]
print(regressed.index.tolist())   # [2]

The data here is made up for illustration. The interviewer is checking groupby with named aggregations, pivot and boolean masks. Always look at the regressed questions themselves before declaring a winner; the metrics side is covered in the LLM evaluation interview questions.

34. How do you read large or messy input files safely?

Answer: stream instead of loading everything, use pathlib, always pass an encoding and decide what to do with bad records instead of crashing the whole job.

import json
import logging
from pathlib import Path

log = logging.getLogger(__name__)

def iter_records(folder: Path):
    for path in sorted(folder.rglob("*.jsonl")):
        with path.open(encoding="utf-8") as f:
            for n, line in enumerate(f, start=1):
                if not line.strip():
                    continue
                try:
                    yield path.name, json.loads(line)
                except json.JSONDecodeError:
                    log.warning("bad JSON at %s:%d", path, n)

Points to mention: JSON Lines suits large datasets because each line is independent; errors="replace" is an option for text of unknown encoding, but log when you use it; and record the file and line number so a failure can be traced to its source.

35. How do you extract text from PDFs for a RAG pipeline?

Answer: first find out what kind of PDF it is. Digital PDFs have a text layer that libraries such as pypdf, pdfplumber or PyMuPDF can extract; scanned PDFs are images and need OCR; many enterprise documents mix both. Tables, multi-column layouts, headers and footers break naive extraction, so check output quality on a sample before you index anything. Keep page numbers with each extracted block so answers can cite a page. For uploads, enforce a size limit, check the file type by content rather than extension, process in a temporary directory and treat parser exceptions as per-document failures. Parser choice is covered in document parsing for RAG.

Interview tip: say "I would check whether pages have a text layer and route scanned ones to OCR". That one sentence shows you have handled real documents.

Testing, logging, packaging and performance

36. How do you use pytest fixtures and parametrize for parsing code?

Answer: parametrize runs one test over many inputs, which suits the long tail of odd model outputs; fixtures supply shared setup such as a fake client or a temporary directory (tmp_path).

import pytest
from app.parsing import extract_json   # see question 46

@pytest.mark.parametrize("raw, expected", [
    ('{"a": 1}', {"a": 1}),
    ('```json\n{"a": 1}\n```', {"a": 1}),
    ('Sure! {"a": 1} Hope this helps.', {"a": 1}),
    ('{"a": 1,}', {"a": 1}),
])
def test_extract_json(raw, expected):
    assert extract_json(raw) == expected

def test_no_json_raises():
    with pytest.raises(ValueError):
        extract_json("I could not find that policy.")

Every production parsing failure should become a new parametrized case. That is how a parser gets more robust over time instead of breaking again in the same way.

37. How do you test code that calls an LLM without calling the LLM?

Answer: inject the client (a parameter or a FastAPI dependency) and replace it with a fake or an AsyncMock in tests. Assert on your logic: parsing, fallbacks, retries and what you sent.

import asyncio
from unittest.mock import AsyncMock

LABELS = {"bug", "access", "billing"}

async def classify_ticket(llm, text: str) -> str:
    label = (await llm.complete(f"Classify: {text}")).strip()
    return label if label in LABELS else "other"

def test_unknown_label_falls_back_to_other():
    llm = AsyncMock()
    llm.complete.return_value = "banana"
    result = asyncio.run(classify_ticket(llm, "VPN down"))
    assert result == "other"
    llm.complete.assert_awaited_once()

Unit tests stay fast, free and deterministic. Separately, keep a small evaluation set that runs against the real model on a schedule or before a release, because mocks cannot tell you whether a prompt change made answers worse. Patch where a name is used, not where it is defined, if you do use unittest.mock.patch. QA-focused angles are in the AI testing interview questions.

38. How should an AI service log?

Answer: use the logging module with module-level loggers, structured (JSON) output in production, a request or trace ID on every line, and levels that mean something. Never use print in services.

import json
import logging

class JsonFormatter(logging.Formatter):
    def format(self, record: logging.LogRecord) -> str:
        return json.dumps({
            "level": record.levelname,
            "logger": record.name,
            "msg": record.getMessage(),
            "request_id": getattr(record, "request_id", None),
        })

handler = logging.StreamHandler()
handler.setFormatter(JsonFormatter())
logging.basicConfig(level=logging.INFO, handlers=[handler])

log = logging.getLogger("ask")
log.info("model call ok", extra={"request_id": "r-42"})

Log model name, latency, token counts, retry counts and validation outcomes. Do not log API keys, and do not log full prompts or outputs containing personal data unless they are redacted and your retention policy allows it. Use log.exception(...) inside except blocks to keep tracebacks. Logs, traces and metrics together are covered in AI observability.

39. How do you manage environments, dependencies, configuration and secrets?

Answer: one virtual environment per project, dependencies declared in pyproject.toml, and a lock file (from uv, Poetry or pip-tools) so every machine and image installs identical versions.

[project]
name = "claims-assistant"
version = "0.1.0"
requires-python = ">=3.11"
dependencies = ["fastapi", "httpx", "pydantic>=2"]

[project.optional-dependencies]
dev = ["pytest", "ruff"]

[project.scripts]
reindex = "claims_assistant.cli:main"

Configuration comes from environment variables, validated at startup with a settings class (for example pydantic-settings), so a missing key fails fast at boot rather than on the first request. Secrets come from a secret manager or the platform's injected environment, never from the repository or the image. Use a src/ layout so tests import the installed package, not files that happen to sit in the working directory.

40. What makes a good Docker image for a Python AI service?

Answer: a slim official Python base image; copy the dependency files and install them before copying source so the dependency layer is cached; install with --no-cache-dir; run as a non-root user; set PYTHONUNBUFFERED=1 so logs appear immediately; keep tests, notebooks and data out with .dockerignore; and never bake secrets or model API keys into a layer. Use multi-stage builds if you need compilers for native wheels. Run the app server (for example uvicorn) as the container's main process so it receives shutdown signals, and add a health endpoint. The full walkthrough is in Docker for AI applications.

41. What is the GIL, and does it matter for LLM applications?

Answer: the Global Interpreter Lock in standard CPython lets only one thread execute Python bytecode at a time. Threads still help for I/O, because the lock is released while waiting on the network or disk, and many C extensions such as NumPy release it during heavy computation. So for an LLM service that mostly waits on APIs, the GIL rarely matters. It matters for pure-Python CPU work: parsing, regex-heavy cleaning, custom tokenisation loops.

Recent change worth knowing: Python 3.13 introduced an experimental free-threaded build without the GIL (PEP 703), and in 3.14 that build became officially supported (PEP 779), but it is still a separate, optional build, not the default, and some C extensions may not support it yet. Say you would measure before relying on it.

42. Threads, processes or async: how do you choose?

Answer: pick by what the work waits on.

ApproachGood forWatch out for
asyncioMany concurrent network calls: model APIs, vector stores, toolsAny blocking call stalls everything; needs async libraries
ThreadsBlocking I/O with sync-only libraries; a few parallel calls in a scriptGIL limits CPU work; shared state needs locks
ProcessesCPU-bound Python: parsing, OCR post-processing, feature workStart-up cost; arguments and results must be pickled; more memory

Real systems combine them: an async API server, a thread for a sync SDK, a process pool for heavy parsing and separate worker containers for batch jobs. Since Python 3.14 the default multiprocessing start method on Linux is forkserver rather than fork, so worker functions must be importable and the entry point needs an if __name__ == "__main__": guard.

Want to turn these answers into working services, tests and deployments rather than memorised text? Cloudsoft's APEX AI, ML, Cloud and Cyber Security program builds this Python layer into real GenAI projects, in the Ameerpet classroom or live online.

Live-coding prompts

For each prompt, talk through the edge cases before typing, write the simple correct version first, then discuss what you would change for production. The code below is one reasonable solution, not the only one.

43. Implement a token-bucket rate limiter.

Answer: the bucket holds up to capacity tokens and refills continuously at rate tokens per second; a request spends tokens if enough are available. This allows short bursts while enforcing the average rate. Use time.monotonic(), never wall-clock time, which can jump.

import time

class TokenBucket:
    def __init__(self, rate: float, capacity: float):
        if rate <= 0 or capacity <= 0:
            raise ValueError("rate and capacity must be > 0")
        self.rate = rate              # tokens added per second
        self.capacity = capacity
        self.tokens = float(capacity)
        self.updated = time.monotonic()

    def _refill(self) -> None:
        now = time.monotonic()
        added = (now - self.updated) * self.rate
        self.tokens = min(self.capacity, self.tokens + added)
        self.updated = now

    def try_acquire(self, cost: float = 1.0) -> bool:
        if cost > self.capacity:
            raise ValueError("cost exceeds bucket capacity")
        self._refill()
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False

    def wait_time(self, cost: float = 1.0) -> float:
        self._refill()
        return max(0.0, (cost - self.tokens) / self.rate)

Follow-ups: make cost the estimated token count to enforce a tokens-per-minute limit; add a threading.Lock for threads, or an async acquire() that sleeps wait_time() and retries for asyncio; and note that this only limits one process. Across several workers or pods you need a shared limiter (scenario 57).

44. Chunk a document into overlapping windows.

Answer: slide a window of size units forward by size - overlap, validate the parameters so the loop always advances, and stop once a window reaches the end so you do not emit a tiny duplicate tail.

def chunk_words(text: str, size: int = 200,
                overlap: int = 40) -> list[str]:
    if size <= 0 or not 0 <= overlap < size:
        raise ValueError("need size > 0 and 0 <= overlap < size")
    words = text.split()
    step = size - overlap
    chunks = []
    for start in range(0, len(words), step):
        chunks.append(" ".join(words[start:start + size]))
        if start + size >= len(words):
            break                    # last window reached end
    return chunks

Edge cases to say aloud: empty text returns an empty list, text shorter than one window returns one chunk, and overlap >= size would loop forever, hence the check. Words are a stand-in for model tokens; production chunkers count tokens with the embedding model's tokenizer and respect headings and paragraphs. Strategy trade-offs are in RAG chunking strategies.

45. Compute cosine similarity without libraries.

Answer: the dot product divided by the product of the two vector lengths. Handle mismatched lengths and zero vectors explicitly.

import math

def cosine(a: list[float], b: list[float]) -> float:
    if len(a) != len(b):
        raise ValueError("vectors must have the same length")
    dot = sum(x * y for x, y in zip(a, b))
    norm_a = math.sqrt(sum(x * x for x in a))
    norm_b = math.sqrt(sum(y * y for y in b))
    if norm_a == 0 or norm_b == 0:
        return 0.0                   # define it; don't divide
    return dot / (norm_a * norm_b)

Mention that on Python 3.10+ zip(a, b, strict=True) raises on unequal lengths, that 3.12 added math.sumprod, and that for real workloads you would use NumPy (question 32) because a pure-Python loop is far slower on large arrays. If vectors are pre-normalised, the dot product alone is enough.

46. Parse JSON from a model response and repair common problems.

Answer: strip Markdown code fences, take the outermost braces, try strict parsing, and apply one conservative repair (trailing commas) before giving up.

import json
import re

FENCE = re.compile(r"```(?:json)?\s*(.*?)```", re.DOTALL)
TRAILING_COMMA = re.compile(r",\s*([}\]])")

def extract_json(raw: str) -> dict:
    text = raw.strip()
    if m := FENCE.search(text):
        text = m.group(1)
    start, end = text.find("{"), text.rfind("}")
    if start == -1 or end < start:
        raise ValueError("no JSON object found")
    candidate = text[start:end + 1]
    try:
        return json.loads(candidate)
    except json.JSONDecodeError:
        fixed = TRAILING_COMMA.sub(r"\1", candidate)
        return json.loads(fixed)     # may still raise

Discuss the limits honestly: the trailing-comma regex could alter a string value that contains a comma before a brace, and "first brace to last brace" fails if the text holds two objects. That is why repair is a fallback. The order of preference is provider structured output, then strict parse plus Pydantic validation, then one repair, then one re-ask with the validation error, then a human queue. The tests in question 36 exercise this function.

47. Build an LRU cache for embeddings.

Answer: an OrderedDict gives O(1) get, put and eviction: move a key to the end on access and pop from the front when over capacity. Key on a hash of model name plus text, because vectors from different embedding models are not comparable.

import hashlib
from collections import OrderedDict

class EmbeddingCache:
    def __init__(self, max_items: int = 10_000):
        self.max_items = max_items
        self._data: OrderedDict[str, list[float]] = OrderedDict()
        self.hits = self.misses = 0

    @staticmethod
    def _key(model: str, text: str) -> str:
        raw = f"{model}\x00{text}".encode("utf-8")
        return hashlib.sha256(raw).hexdigest()

    def get(self, model: str, text: str) -> list[float] | None:
        k = self._key(model, text)
        if k not in self._data:
            self.misses += 1
            return None
        self._data.move_to_end(k)     # mark most recently used
        self.hits += 1
        return self._data[k]

    def put(self, model: str, text: str, vec: list[float]):
        k = self._key(model, text)
        self._data[k] = vec
        self._data.move_to_end(k)
        if len(self._data) > self.max_items:
            self._data.popitem(last=False)   # evict oldest

Why not functools.lru_cache? It works for pure functions, but you cannot easily key on a normalised input, inspect hit rates or share it across processes, and decorating an async function with it caches the coroutine object, not the result. Follow-ups: a lock for thread safety, memory limits by bytes rather than item count, and Redis or a database table when several workers should share the cache.

48. Write an async retry decorator with exponential backoff and jitter.

Answer: retry only listed exception types, double the delay ceiling each attempt up to a cap, pick a random delay up to that ceiling (full jitter) so many clients do not retry in lockstep, and re-raise after the final attempt.

import asyncio
import functools
import random

class RetryableError(Exception):
    """Raise for 429, 5xx and connection resets."""

def retry(attempts: int = 4, base: float = 0.5, cap: float = 8.0,
          retry_on: tuple = (RetryableError, TimeoutError)):
    def decorator(fn):
        @functools.wraps(fn)
        async def wrapper(*args, **kwargs):
            for attempt in range(attempts):
                try:
                    return await fn(*args, **kwargs)
                except retry_on:
                    if attempt == attempts - 1:
                        raise
                    delay = min(cap, base * 2 ** attempt)
                    await asyncio.sleep(random.uniform(0, delay))
        return wrapper
    return decorator

Points to raise: retry_on must not include validation or bad-request errors (scenario 58); a server's Retry-After should override the computed delay; non-idempotent calls such as creating a ticket need an idempotency key before they are safe to retry; and the total time must fit inside the request deadline from question 16. Libraries such as tenacity provide this, but interviewers want to see that you understand it.

Production scenarios

49. A nightly job embeds a few hundred thousand chunks. It hits constant 429 errors and takes hours. What do you do?

Answer: stop sending calls faster than the account allows, batch more items per call and skip work that has already been done.

What I would check:

  1. The provider's request and token limits for this model and account, and whether the job shares a key with live traffic.
  2. Whether each call sends one chunk or a batch, and how close batches are to the per-request maximums.
  3. Concurrency control: is there a semaphore and a rate limiter, or a bare gather over everything?
  4. Whether unchanged chunks are re-embedded every night; content hashes let you skip them.
  5. Retry behaviour: jitter, Retry-After and a cap, or a hot loop hammering the API.

Production consideration: give batch jobs their own quota or key so they cannot starve user-facing traffic, make the job resumable from a checkpoint, and check whether the provider offers a discounted asynchronous batch interface for work that is not urgent.

50. A FastAPI chat endpoint is fast alone but latency spikes for everyone under modest load. Why?

Answer: the most common cause is blocking work inside async def: a synchronous SDK, requests, a sync database driver or CPU-heavy processing stalls the event loop, so requests queue behind each other.

What I would check:

  1. Every call inside the endpoint and its dependencies: is each one awaited on an async library?
  2. Asyncio debug mode or a profiler to find slow synchronous callbacks.
  3. Whether an HTTP client is created per request instead of once in the lifespan.
  4. Connection pool limits on the HTTP client and database, which queue requests silently.
  5. Worker count and CPU usage per worker.

Production consideration: fix the blocking call (async client, asyncio.to_thread, or change the endpoint to def), then load-test with realistic concurrency and streaming. More levers are in LLM latency optimisation.

51. In production, a small share of model responses fail JSON parsing and the endpoint returns 500. How do you fix it?

Answer: make invalid output a handled outcome, not an exception that escapes, and reduce how often it happens.

What I would check:

  1. Samples of the failing raw outputs: fences, prose around the JSON, truncation, wrong types or missing fields.
  2. Whether truncation is hitting the max output token setting.
  3. Whether the provider's structured-output or tool-calling mode is used, and whether the schema contains keywords the provider ignores.
  4. Where ValidationError and JSONDecodeError are caught, if anywhere.

Production consideration: wrap parsing in the repair-validate-re-ask-fallback chain from question 46, return a clear 502 or a degraded response instead of a 500, add each failing sample to the parser test suite, and alert on the validation failure rate so a model or prompt change that raises it is caught quickly.

52. A long-running worker's memory grows steadily until the container is killed. Where do you look?

Answer: something keeps references it should drop. In AI workers the usual suspects are unbounded caches, accumulated results and global conversation state.

What I would check:

  1. Module-level dicts or lists that grow per request, including a mutable default argument.
  2. functools.lru_cache(maxsize=None) or a hand-made cache without eviction.
  3. Results collected into one big list across a job instead of being written out as they finish.
  4. Tasks created with create_task and never awaited, or responses and file handles never closed.
  5. tracemalloc snapshots compared over time to see which lines allocate the growth.

Production consideration: bound every cache, stream results to storage, and set container memory limits plus a periodic worker restart as a safety net, not as the fix.

53. Streaming works locally, but in staging users see the whole answer appear at once. Why?

Answer: something between client and server is buffering: a reverse proxy, a load balancer, compression middleware, or the client code itself.

What I would check:

  1. Response headers: Content-Type: text/event-stream and Cache-Control: no-cache.
  2. Proxy settings: Nginx buffers proxied responses by default; disable buffering for the route or send an X-Accel-Buffering: no header.
  3. Gzip or other compression middleware applied to the stream.
  4. Idle timeouts at the load balancer that cut long streams; send periodic keep-alive comments.
  5. The frontend: is it reading the stream incrementally or awaiting the full body?

Production consideration: add a staging test that measures time to first byte through the real proxy path, and make sure client disconnects cancel the upstream model call (question 18).

54. CI is slow, costs money and fails randomly because tests call the real model. How do you restructure it?

Answer: separate deterministic unit tests from model evaluation.

What I would check:

  1. Which tests reach the network: block outbound calls in the unit suite so any leak fails loudly.
  2. Whether the LLM client is injectable everywhere, or imported directly deep inside functions.
  3. Which assertions compare exact model text; those can never be stable.

Production consideration: unit tests use fakes and recorded responses and run on every commit; a small, versioned evaluation set runs against the real model on a schedule or before release, with thresholds rather than exact matches; and results are tracked over time so drift is visible. The tests do not change; what they assert changes.

55. A hospital's document pipeline fails on some PDFs and a security review finds patient names in the logs. What do you change?

Answer: two separate problems: extraction that does not handle the document types the hospital actually has, and logging that treats patient data as debug output.

What I would check:

  1. Which PDFs fail: scanned pages without a text layer, encrypted files, very large files, or malformed ones.
  2. Whether one bad document crashes the whole batch or is isolated and reported.
  3. Every log statement and exception message that includes document text, prompts or model output.
  4. Log retention, who can read the logs, and whether third-party tracing tools receive the same data.

Production consideration: route scanned pages to OCR, quarantine failures with an ID-only error record, log document IDs and sizes instead of content, add a redaction filter as defence in depth, and agree log retention with the hospital's compliance team. In India, align the handling with the DPDP Act obligations the client's legal team defines.

56. The service image is very large, builds slowly and a scan finds an API key in a layer. What do you do?

Answer: rotate the key first, because removing it from the Dockerfile does not remove it from published layers. Then fix the build.

What I would check:

  1. How the key got in: copied .env, a build argument or a COPY . . without a .dockerignore.
  2. The base image and whether build tools, test data or notebooks ship in the final image.
  3. Layer order: are dependencies reinstalled whenever any source file changes?

Production consideration: inject secrets at runtime from a secret manager, use build secrets for anything needed during the build, add a .dockerignore, split dependency installation from source copy, use a multi-stage build and add secret scanning to CI so it does not happen again.

57. Your token bucket works, but with several pods you still exceed the provider's limit. Why, and what next?

Answer: each pod has its own bucket, so the combined rate is the per-pod rate multiplied by the number of pods, and autoscaling makes it worse.

What I would check:

  1. How limits are configured per pod against the account-wide quota.
  2. Whether other services or batch jobs share the same key.
  3. Whether the limiter counts requests while the provider also limits tokens.

Production consideration: move limiting to a shared place, either a central limiter in Redis (an atomic script that refills and decrements) or an LLM gateway that owns quotas, keys, retries and per-team budgets for every service. Keep a local semaphore in each pod as a second line of protection.

58. After a teammate's change, model spend rises sharply and the ITSM system shows duplicate tickets. What went wrong?

Answer: the likely cause is retry logic that is too broad: retrying on every exception, including validation errors and 400s, and retrying non-idempotent tool calls.

What I would check:

  1. The retry_on list or except Exception in the retry wrapper.
  2. Whether retries are stacked: SDK retries plus a wrapper plus the caller.
  3. Whether "create ticket" sends an idempotency key or checks for an existing ticket first.
  4. Logs grouped by request ID to count attempts per user request.

Production consideration: retry only transient errors, keep one retry layer, require idempotency keys for writes, and add a metric for attempts per request with an alert on its rise. Cost dashboards per feature catch this kind of change within a day instead of at month end.

59. Code review: what is wrong with this function?

import requests

def summarise_all(docs, results=[]):
    for d in docs:
        try:
            r = requests.post(URL, json={"text": d})
            results.append(r.json()["summary"])
        except:
            print("failed", d)
    return results

Answer: several issues, roughly in order of severity.

What I would check:

  1. Mutable default results=[]: results accumulate across calls and leak between users.
  2. No timeout on requests.post; one hung connection blocks forever.
  3. Bare except: hides every error, including bugs, and print loses them; there is no status check, so an error body triggers a KeyError that is then swallowed.
  4. Calls are sequential and synchronous; for many documents use an async client with a concurrency limit.
  5. A new connection per call; no shared session or client.
  6. URL is a global with no configuration or validation; there are no type hints and the response is not validated.

Production consideration: rewrite with an injected async client, explicit timeouts, raise_for_status, a Pydantic response model, targeted exceptions, logging with document IDs and a returned list of failures. Then add a test with a fake client that returns one success and one error.

60. Text cleaning and parsing is CPU-heavy and slows the async API. How do you fix it without rewriting everything?

Answer: move CPU-bound work off the event loop into a process pool, or out of the request path entirely.

import asyncio
from concurrent.futures import ProcessPoolExecutor

def heavy_parse(text: str) -> int:      # CPU-bound work
    return sum(len(w) for w in text.split() * 2000)

async def handle(pool: ProcessPoolExecutor, text: str) -> int:
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(pool, heavy_parse, text)

async def main():
    with ProcessPoolExecutor(max_workers=2) as pool:
        print(await handle(pool, "scanned page text"))

if __name__ == "__main__":     # needed for spawn/forkserver
    asyncio.run(main())

What I would check:

  1. A profile (cProfile or a sampling profiler) to confirm where the time goes before changing anything.
  2. Whether the work must happen per request at all, or can be done at ingestion time and stored.
  3. Argument and result sizes, because everything crossing a process boundary is pickled.
  4. Whether a faster library (a C-backed parser, compiled regex reuse) removes the problem.

Production consideration: size the pool to available CPUs, create it once at startup, and for heavy or long jobs prefer a separate worker service fed by a queue, which scales and fails independently of the API.

Key takeaways

  • AI engineering interviews test backend Python: generators, async, HTTP, validation, testing and deployment, with the model as one dependency among many.
  • Know the difference between limiting concurrency (semaphore) and limiting rate (token bucket), and between gather and TaskGroup failure behaviour.
  • Every outbound call needs a timeout, and only transient errors deserve a retry with backoff and jitter.
  • Validate model output with Pydantic v2 at the boundary and treat validation failure as a normal, handled path.
  • Design for testability: inject clients so unit tests use fakes and a separate evaluation set checks real model quality.
  • Async for waiting, processes for CPU, threads for sync-only I/O; measure before optimising.
  • Practise the six live-coding prompts until you can write them and talk through their edge cases at the same time.

Interview preparation checklist

  • Write the timing and async retry decorators from memory, with functools.wraps.
  • Write a semaphore-limited gather and a TaskGroup example, and explain what happens when one task fails.
  • Build a small FastAPI service with a Pydantic request model, a streaming endpoint and an injected LLM client.
  • Write tests for it with dependency_overrides, AsyncMock and a parametrized parser suite.
  • Translate a v1 Pydantic model into v2 syntax and generate its JSON Schema.
  • Implement the token bucket, chunker, cosine similarity, JSON repair, LRU cache and retry, then test each one's edge cases.
  • Containerise the service with a non-root user and no secrets in the image.
  • Prepare one story about a production bug you found in Python code: symptom, diagnosis, fix and the test you added.
  • Revise the neighbouring guides: AI engineer interview questions for system-level topics and Hugging Face interview questions if the role touches open models.

FAQ

Which Python version should I prepare with for AI interviews?

Use a currently supported Python 3 release, ideally 3.11 or later, because TaskGroup, asyncio.timeout and exception groups appear in interview questions. Mention newer features only when you know how they behave.

Do AI engineer interviews include LeetCode-style problems?

Some do, especially at product companies, but many AI application roles prefer practical tasks such as a rate limiter, a chunker, JSON repair or a small API. Prepare core data structures and these applied prompts.

How much NumPy and pandas do I need?

For application roles, enough to compute similarities, filter and group evaluation results and join tables. Deep numerical work matters more for ML research or model training roles.

Is async Python mandatory for AI engineering roles?

It is commonly expected, because LLM services make many concurrent network calls. You should be able to read and write async functions, limit concurrency and avoid blocking the event loop.

Should I learn Pydantic v1 or v2?

Learn v2, which current FastAPI and most new projects use. Know the main renamed methods so you can recognise v1 code in older projects and migrate it.

Do I need to know LangChain or LangGraph for a Python interview?

Not for the Python round itself. Interviewers usually want to see that you can build the core logic in plain Python; framework questions come in separate rounds.

How do I practise live coding for AI interviews?

Time yourself writing each prompt from a blank file, say the edge cases aloud, then write tests. Repeat until you can explain trade-offs while typing.

Is Python enough to get an AI engineering job?

Python is the base, but roles also expect HTTP APIs, SQL, Git, Docker, a cloud platform and an understanding of RAG, agents and evaluation. A deployed project with tests shows these together.

If you want guided practice that goes from these Python building blocks to deployed GenAI systems, look at the Cloudsoft APEX program. For language foundations first, the Python training in Hyderabad course covers them, and if your goal is to build and deploy AI systems inside customer environments, the AI Forward Deployed Engineer course (FDE PRO) uses Python and FastAPI across five enterprise projects with placement support until you're placed. Classes run in Ameerpet or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us