LLM inference interview questions in 2026 check whether you can run a model as a production service: one that meets a latency target at peak load, fits its GPU memory, stays affordable and recovers when something breaks. This guide has 55 high-value questions with model answers on prefill and decode, KV cache management, continuous batching, caching, speculative decoding, quantisation, parallelism, routing, autoscaling, capacity planning and cost. It ends with twelve production scenarios of the kind that serving and platform interviews commonly use. Every number in it is illustrative. None of them are product benchmarks.
How to use this guide
This page assumes you already know transformer basics: attention, what the KV cache is, sampling. If not, start with our LLM interview questions and come back. Here we look at the serving layer, meaning the engine, the scheduler, the router and the platform around them. What interviewers typically test at each level:
- Freshers and junior engineers: prefill vs decode, what TTFT and inter-token latency measure, why batching matters, and what an OpenAI-compatible endpoint is.
- Mid-level engineers: PagedAttention, prefix caching, chunked prefill, preemption, engine knobs, quantisation choices and multi-LoRA serving.
- Senior, platform and architect roles: parallelism strategy, disaggregated prefill and decode, KV-cache-aware routing, autoscaling and cold starts, capacity and cost models, and structured debugging of latency incidents.
To make these answers your own, serve a small open model, put load on it and watch the engine's metrics move.
- Serving fundamentals (Q1βQ10)
- Memory, batching and caching (Q11βQ19)
- Speculative decoding and quantisation (Q20βQ25)
- Parallelism and disaggregation (Q26βQ30)
- Routing, autoscaling and operations (Q31βQ38)
- Capacity planning and cost (Q39βQ43)
- Real-world scenario questions (Q44βQ55)
- Key takeaways
- Interview preparation checklist
- FAQ
Serving fundamentals
1. What does an LLM inference server do that a plain model.generate() loop does not?
Answer: A generate() loop runs one request, or one fixed batch, from start to finish. An inference server runs many independent requests on shared GPUs at the same time. That needs five extra things. A scheduler decides which requests run in each forward step. A KV cache manager allocates, shares and frees cache memory as requests come and go. Optimised kernels handle attention, quantised matrix multiplication and sampling. An API layer handles streaming, cancellation, tokenisation and chat templates. Operational surfaces provide health endpoints, metrics and adapter loading. Throughput and latency depend far more on the scheduler and memory manager than on the model code.
2. Why do serving engines treat prefill and decode as two different workloads?
Answer: Their arithmetic intensity (how many operations the GPU performs per byte it reads from memory) is very different. Prefill pushes hundreds or thousands of prompt tokens through every weight matrix in one pass. Each weight read from memory is reused many times, so prefill is compute-bound. Decode produces one new token per sequence per step. The GPU still has to read all the weights and that sequence's KV cache, but it does very little maths with them, so decode is memory-bandwidth-bound. Batching many sequences into one decode step reuses each weight read across the batch, which is why batching helps decode so much. The two phases also hurt each other: a long prefill placed in the same step as ongoing decodes delays every one of those users' next tokens. Chunked prefill, decode-priority scheduling and disaggregated serving all exist to manage this conflict.
3. Why does the KV cache, not the model weights, usually limit how many users a GPU can serve?
Answer: The weights are a fixed cost that every request shares. The KV cache grows with every token of every active sequence. Per-token cache size is roughly 2 Γ layers Γ KV heads Γ head dimension Γ bytes per value. Take a hypothetical model with 40 layers, 8 KV heads (grouped-query attention), head dimension 128 and 16-bit values: 2 Γ 40 Γ 8 Γ 128 Γ 2 bytes = 160 KiB per token. One 8,192-token conversation then needs about 1.25 GiB of cache. Once the weights are loaded, whatever memory is left divided by the cache per sequence sets your maximum concurrency. Grouped-query attention, KV cache quantisation and shorter contexts all raise that ceiling. Our GPU basics for AI engineers guide covers the hardware side of this sizing.
4. What problem does PagedAttention solve, and how?
Answer: Early servers reserved one contiguous block of cache per request, sized for the longest output it might produce. Most requests finish early, so much of that memory sat unused, and the variable sizes fragmented the free space. PagedAttention, introduced with vLLM, borrows the idea of virtual-memory paging. The cache is split into fixed-size blocks of a few tokens each. Each sequence keeps a block table that maps its logical token positions to physical blocks, which can sit anywhere in GPU memory. Blocks are allocated only as tokens are generated, and they are freed the moment a request ends. Several sequences can also point at the same physical blocks, for example parallel samples from one prompt or a shared system prompt, with copy-on-write when they diverge. The result is far less wasted memory, so more sequences fit and throughput rises.
5. Compare static, dynamic and continuous batching.
Answer:
- Static batching: collect a fixed batch, run it until the longest sequence finishes, then start the next one. Simple, but short requests wait for long ones and slots go idle.
- Dynamic (request-level) batching: group whatever requests arrive within a short time window. This works well for fixed-size models such as classifiers, but for generation the batch is still locked until its longest member finishes.
- Continuous (iteration-level) batching: the scheduler rebuilds the batch at every decode step. Finished sequences leave immediately and waiting requests join, often with their prefill mixed into the same step. TensorRT-LLM calls this in-flight batching.
Continuous batching is the default in modern LLM engines (vLLM, SGLang, TensorRT-LLM, and llama.cpp's server). It wins because generation lengths vary enormously.
6. Define TTFT, inter-token latency, TPOT, end-to-end latency, throughput and goodput precisely.
Answer:
- TTFT (time to first token): from the moment the request arrives until the first output token is sent. It includes queueing, tokenisation and prefill, so under load it is often mostly queueing.
- Inter-token latency (ITL): the gap between consecutive streamed tokens. Users feel it as stutter.
- TPOT (time per output token): the average decode time per token for a request, usually (end-to-end minus TTFT) divided by (output tokens minus one).
- End-to-end latency: roughly TTFT plus output tokens Γ TPOT.
- Throughput: requests per second, or tokens per second. Always say whether you mean input tokens, output tokens or both, because prefill tokens are much cheaper to process than decode tokens.
- Goodput: the rate of requests that meet all their SLOs (for example a TTFT target and a TPOT target). Pushing throughput up by violating latency targets does not count.
vLLM, for example, exposes histograms such as vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds and vllm:e2e_request_latency_seconds. The application-side view of these metrics is covered in our LLM latency optimisation guide.
7. Why set SLOs on p95 and p99 rather than averages, and what mistakes do people make computing them?
Answer: Latency distributions in LLM serving are long-tailed. A few long prompts, a burst of queueing or one preemption can make a small share of users wait many times longer than average, and those users file the complaints. Typical mistakes:
- Averaging the p95 values of several replicas. Percentiles cannot be averaged. Merge histograms first, then compute the percentile.
- Mixing workloads. Batch summarisation and chat in one histogram makes both look wrong. Label by route, model and tenant.
- Measuring only inside the engine. The client sees gateway, network and guardrail time on top. Measure at both ends.
8. How are concurrency, throughput and latency related?
Answer: Little's law: average concurrency = arrival rate Γ average time in the system. If requests arrive at 4 per second and each takes 10 seconds end to end, about 40 are in flight at any moment. For LLM serving this means two things. First, longer outputs or slower decoding directly raise concurrency, and with it KV cache demand, even when traffic stays flat. Second, raising the batch size raises throughput only until memory or compute saturates. Beyond that point extra concurrency just becomes queueing, and TTFT climbs steeply.
9. What does "OpenAI-compatible API" actually promise when you self-host?
Answer: It means the server exposes the same endpoint shapes, mainly /v1/chat/completions, /v1/completions and often /v1/embeddings and /v1/models, with server-sent-event streaming, so existing SDKs and gateways work by changing the base URL. vLLM, SGLang, TensorRT-LLM (through trtllm-serve) and llama.cpp's llama-server all provide this. It does not mean identical behaviour. Supported parameters differ. Tool-calling depends on the model's chat template and a server-side parser. Structured-output options and error formats vary, and extra engine-specific parameters are common. Token counts also differ from a hosted model because the tokenizer is different.
Interview tip: Mention contract tests (streaming, tool calls, JSON mode, error codes) for every engine behind the same API.
10. How do you choose between vLLM, SGLang, TensorRT-LLM and llama.cpp in 2026?
Answer: Start from your constraints, not from popularity.
| Engine | Typical fit | Watch out for |
|---|---|---|
| vLLM | General-purpose GPU serving with broad model and hardware support, multi-LoRA, structured outputs, many quantisation formats | Fast release cadence; pin versions and re-test on upgrade |
| SGLang | Workloads with heavy prefix sharing (agents, multi-turn, few-shot), thanks to RadixAttention; also supports PD disaggregation | Same need for version pinning; check model support |
| TensorRT-LLM | NVIDIA-only fleets that want highly optimised kernels and integration with Triton Inference Server | Tied to NVIDIA hardware; more tuning per model and GPU |
| llama.cpp | GGUF models on CPUs, Apple silicon, edge devices and laptops; small internal tools | Not designed for high-concurrency datacentre serving |
Hugging Face TGI is now in maintenance mode, according to its own documentation, which recommends vLLM, SGLang or local engines such as llama.cpp for new work. Mention this if you are asked about TGI. The operational picture of running any of these engines inside an enterprise is covered in self-hosting LLMs.
Memory, batching and caching
11. How does prefix caching work inside a serving engine?
Answer: When two requests share an identical token prefix, such as the same system prompt, tool definitions or document, the KV cache for that prefix only needs to be computed once. vLLM's automatic prefix caching hashes each full KV block together with the hash of all blocks before it. A new request whose first blocks match existing hashes reuses those blocks and only runs prefill on the rest. SGLang's RadixAttention keeps cached prefixes in a radix tree, which suits branching multi-turn and agent traffic. Cached blocks that no running request is using stay in memory until they are evicted (typically least recently used) to make room. The win is lower TTFT and less prefill compute. The catch is that matching is exact at the token level: a timestamp, user name or reordered tool list near the top of the prompt changes everything after it.
12. How is engine-level prefix caching different from a provider's prompt caching?
Answer: They use the same idea but you operate them differently. With a hosted API, the provider manages cache placement, lifetime and billing (often a reduced rate for cached input tokens), and sometimes you mark cache breakpoints explicitly. When you self-host, the cache lives in each replica's GPU memory. A prefix cached on replica A does nothing for a request routed to replica B. It also competes for the same memory as active sequences, so under memory pressure it gets evicted exactly when you need it. That makes routing (Q31) part of your caching strategy. Some stacks add tiered caches that spill blocks to CPU memory or storage (for example through LMCache or llm-d's tiered prefix cache) so they survive eviction and can be shared between instances.
13. What is chunked prefill, and what trade-off does its token budget control?
Answer: Without chunking, a 30,000-token prompt runs as one large prefill step, and every streaming user on that replica waits for it. Their inter-token latency spikes. Chunked prefill splits long prompts into pieces that fit a per-step token budget and mixes them with ongoing decodes. vLLM's documentation says chunked prefill is on by default where possible, that decode requests are scheduled first, and that prefill chunks fill the rest of the max_num_batched_tokens budget. A smaller budget protects inter-token latency. A larger budget finishes prefills sooner, which lowers TTFT and usually raises throughput. Interactive chat leans towards a smaller budget; long-document ingestion leans towards a larger one.
14. What is preemption, and how do you recognise it in production?
Answer: With continuous batching, the scheduler admits requests before knowing how long they will run. If the KV cache fills up mid-generation, the engine must evict a running sequence to let others continue. In vLLM's current engine the default is to recompute: drop the evicted sequence's blocks and rebuild its cache later. Swapping blocks to CPU memory is the alternative. The affected user sees a long pause, and the GPU repeats prefill work. Signs to look for: the engine logs a warning that not enough KV cache space is available, the preemption counter rises, KV cache usage sits near full, and tail TPOT gets worse while the median looks fine. vLLM's tuning guide suggests these fixes: give the cache more memory (gpu_memory_utilization), lower max_num_seqs or max_num_batched_tokens, or spread the model across more GPUs with tensor or pipeline parallelism.
15. Which engine settings would you tune first, and what does each one control?
Answer: Using vLLM's names (other engines have equivalents):
gpu_memory_utilization: the share of GPU memory the engine pre-allocates for weights, activations and KV cache. Higher means more cache, but leave room if anything else runs on the GPU.max_model_len: the longest context (prompt plus output) a single request may use. Setting it far above what you need wastes nothing at rest, but it allows a few huge requests to take over the cache.max_num_seqs: the cap on concurrent sequences per step. Too high causes preemption; too low leaves memory and throughput unused.max_num_batched_tokens: the per-step token budget that governs chunked prefill (Q13).- Parallelism settings (
tensor_parallel_size,pipeline_parallel_size) and quantisation or KV cache dtype.
16. What are KV cache quantisation and KV offloading, and when would you use each?
Answer: KV cache quantisation stores keys and values in fewer bits (for example FP8 instead of 16-bit). That roughly halves cache memory per token, so you get about twice the concurrency or context at the same memory. It can cost accuracy, especially on long-context retrieval tasks, so evaluate it separately from weight quantisation. KV offloading moves cold cache blocks to CPU memory or local storage and brings them back when a returning conversation or shared prefix needs them. It helps when many sessions return after a pause. The cost is transfer time, so it only pays off when reloading is cheaper than recomputing the prefill.
17. Why does cancellation matter in streaming LLM serving?
Answer: Users close tabs, press "stop" and retry. If the server keeps generating for a request nobody will read, it burns decode capacity and holds KV cache that waiting requests need. Make sure client disconnects travel through the whole path: browser to gateway, gateway to engine, then an abort inside the engine. Some proxies also buffer server-sent events, which wrecks perceived TTFT. Pair cancellation with sensible max_tokens limits so a looping generation cannot run until the context is full.
18. How does guided decoding (structured output) work inside the engine, and what can go wrong?
Answer: The engine compiles a constraint, such as a JSON schema, regex, choice list or context-free grammar, into a state machine. At every step it masks the logits of tokens that would break the constraint, so the model can only sample valid continuations. vLLM supports choice, regex, JSON schema and grammar constraints, with xgrammar and guidance as backends. llama.cpp supports grammars and JSON schemas. What goes wrong:
- Truncation: if
max_tokensruns out mid-object, the output is still invalid JSON. Check the finish reason. - Compilation cost: large or deeply nested schemas take time to compile the first time they are seen.
- Valid but wrong: the constraint enforces shape, not truth. The model can produce a perfectly formed field holding the wrong value.
The application side of this is in function calling and structured outputs.
19. How does multi-LoRA serving work, and when is it better than deploying merged models?
Answer: One base model stays in GPU memory, and many small LoRA adapters are loaded alongside it. Requests for different adapters can share the same batch, because special kernels apply each sequence's adapter to its own rows. In vLLM you start with --enable-lora, register adapters with --lora-modules name=path, and the client chooses one through the model field. --max-loras caps how many adapters can be active in one batch and --max-lora-rank sizes the buffers. Adapters can also be loaded at runtime through an API when explicitly enabled. This is very cost-efficient when you have many low-traffic variants, such as one adapter per business unit or language. A merged, dedicated deployment wins when one variant carries heavy traffic and needs every bit of latency, or when adapters have very different ranks. How the adapters are trained is covered in our fine-tuning LLM interview questions.
Speculative decoding and quantisation
20. Describe the main families of speculative decoding.
Answer: All of them have something cheap propose several tokens, then the target model checks them in one forward pass and keeps the longest accepted run. They differ in what does the proposing:
- Separate draft model: a smaller model from the same family. You must host it and keep its tokenizer aligned.
- Draft heads on the target's hidden states: EAGLE-style and Medusa-style methods attach light prediction layers to the target model, so drafts are cheap and well matched.
- Multi-token prediction (MTP): some models are trained with extra prediction modules that can serve as built-in drafters.
- Model-free methods: n-gram or prompt lookup, and suffix decoding, copy likely continuations from the prompt or earlier outputs. They are very cheap and work well for code editing, extraction and RAG answers that quote their sources.
vLLM's documentation lists model-based methods (EAGLE, MTP, draft models and others) as giving the larger latency gains, and simpler methods such as n-gram as lightweight with smaller gains. It describes the technique as lossless in theory, up to hardware numeric precision.
21. Why can speculative decoding help at low load but hurt at high load?
Answer: Speculation trades spare compute for fewer sequential steps. At low concurrency, decode is memory-bound and the GPU has idle compute, so checking several draft tokens costs little more than one normal step. When drafts are accepted, latency falls. At high concurrency, large batches already use that compute. Every rejected draft token is wasted work that could have served another user, and the drafter itself costs time and memory. Throughput can then drop. Acceptance rate also depends on the workload: predictable text accepts well, creative text at high temperature accepts poorly. So measure acceptance rate and enable it where latency matters.
22. Weight-only quantisation versus weight-and-activation quantisation: which phase does each speed up?
Answer: Weight-only schemes (for example INT4 with GPTQ or AWQ, which store low-bit weights and convert them back to 16-bit inside the kernel) shrink the bytes read per decode step. That helps memory-bound decode and makes larger models fit. The maths still runs in 16-bit, so compute-bound prefill gains little, and the conversion overhead can show up at large batch sizes. Weight-and-activation schemes (W8A8 in INT8 or FP8, and newer 4-bit floating-point formats on hardware that supports them) run the matrix multiplications themselves in low precision on tensor cores. That speeds up prefill and large-batch decode as well. The rule of thumb: weight-only for memory-constrained, latency-oriented serving at modest batch sizes; weight-and-activation for high-throughput serving on GPUs with native low-precision support.
23. Compare FP8, INT8, INT4 (GPTQ/AWQ) and the newer FP4 formats.
Answer:
| Format | What it is | Typical trade-off |
|---|---|---|
| FP8 | 8-bit floating point for weights and often activations; native on recent NVIDIA and AMD datacentre GPUs | Usually small quality loss; good first step when the hardware supports it |
| INT8 (e.g. SmoothQuant-style W8A8) | 8-bit integers, with activation outliers handled by rescaling | Works on older GPUs; needs calibration |
| INT4 GPTQ / AWQ | 4-bit weight-only; GPTQ minimises layer error using second-order information, AWQ protects weights that matter most to activations | Large memory savings; quality risk on reasoning, code and non-English text |
| FP4 variants (NVFP4, MXFP4) | 4-bit floating point with block scaling, on the newest accelerators | Hardware-specific; validate quality carefully |
vLLM's documentation lists FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ and GGUF among its supported formats. TensorRT-LLM lists FP8, FP4, INT4 AWQ and INT8 SmoothQuant. Which kernels are fast depends on the GPU generation, so check the support matrix for your exact hardware.
24. What is GGUF, and where does it fit in an enterprise serving strategy?
Answer: GGUF is the single-file model format used by llama.cpp and the tools built on it. The file bundles weights (often in quantised types with several bit widths), tokenizer data and metadata. It suits CPU, Apple silicon and edge deployment, offline developer laptops, and small internal tools where installing GPU drivers is not practical. llama.cpp's server provides OpenAI-compatible endpoints, parallel slots with continuous batching, grammar and JSON-schema constraints, speculative decoding and LoRA adapters. For high-concurrency datacentre serving, GPU engines are the usual choice, even though vLLM can load some GGUF models. A sensible split is GGUF on the edge, such as a branch-office or factory-floor assistant (see on-device and edge AI), and a GPU engine in the datacentre.
25. How do you prove that a quantised model is good enough?
Answer: Compare it with the full-precision model on your task evaluation set, with the same prompts and decoding settings. Reputation and perplexity are not enough. Include the slices where quantisation tends to hurt: multi-step reasoning, arithmetic, code, long-context retrieval, structured-output validity, tool-call accuracy, and Indian languages if your users write in Hindi, Telugu or Tamil. Compare output length too, because a quantised model that rambles costs more decode time and erases the speed gain. Define your acceptance criteria before the test, such as "no slice regresses beyond an agreed tolerance", and keep the full-precision deployment available for rollback. If 4-bit fails, the next options are 8-bit, a smaller model at full precision, or a distilled model (see model distillation explained).
Parallelism and disaggregation
26. What is tensor parallelism, and why is it usually kept within a node?
Answer: Tensor parallelism (TP) splits each layer's weight matrices across GPUs. Every GPU computes part of each matrix multiplication, and the partial results are combined with all-reduce operations several times per layer, for every token. That cuts per-GPU weight memory and per-token latency, but it puts collective communication on the critical path of every step. It works well over fast intra-node links such as NVLink, and it suffers badly across ordinary network links between nodes. So TP is normally set to at most the number of GPUs in one server.
27. When would you use pipeline parallelism for inference?
Answer: Pipeline parallelism (PP) assigns consecutive groups of layers to different GPUs or nodes, and activations flow from one stage to the next. It sends far less traffic between stages than TP, so it is the usual way to span nodes when a model is too large for one server, typically TP inside each node plus PP across nodes. The cost is pipeline bubbles, where stages wait for each other. PP improves capacity, not single-request latency: a lone request still passes through every stage in sequence.
28. Explain data parallelism and expert parallelism in a serving context.
Answer: Data parallelism in serving just means independent replicas, each holding a full copy of the model, with a load balancer in front. It is the simplest way to scale throughput and keeps failures isolated. Expert parallelism (EP) applies to mixture-of-experts models: instead of copying every expert onto every GPU, different GPUs host different experts, and tokens are sent to whichever GPUs hold their chosen experts using all-to-all communication. EP makes very large MoE models practical, but it adds communication and load-imbalance risk when some experts become "hot". Large MoE deployments often combine attention data parallelism with expert parallelism. vLLM's documentation lists tensor, pipeline, data, expert and context parallelism, and TensorRT-LLM describes "wide" expert parallelism for MoE.
29. What is disaggregated prefill and decode, and when is it worth the complexity?
Answer: It runs prefill and decode on separate instances (pools of GPUs). A prefill worker processes the prompt, the KV cache is transferred to a decode worker, and the decode worker streams the tokens.
client
|
router (KV/prefix aware)
|
prefill pool --KV transfer--> decode pool
(compute-bound, (bandwidth-bound,
sized for TTFT) sized for ITL)
| |
+--------- tokens streamed ----+
The benefit is that long prefills never interrupt decodes, and each pool can use its own parallelism settings, hardware and scaling policy to hit TTFT and ITL targets separately. vLLM supports it as an experimental feature through KV-transfer connectors (such as NIXL-, LMCache- and Mooncake-based ones). SGLang and TensorRT-LLM support PD disaggregation. llm-d, a Kubernetes-native distributed inference stack that runs engines such as vLLM, publishes it as a tested deployment pattern. vLLM's docs make a point worth repeating in interviews: disaggregation does not improve throughput by itself. It is a tool for controlling latency, especially tail ITL. It pays off with long prompts, strict latency SLOs, scale large enough to keep two pools busy, and fast interconnect for the KV transfer. For small fleets, chunked prefill is simpler and often enough.
30. A model fits on one GPU with little room left. Do you use TP=2 or two replicas?
Answer: It depends on the SLO and on traffic.
- Two replicas (TP=1 each): no communication overhead, isolated failures and simpler scaling. But each replica has very little room for KV cache, so concurrency and context per replica are low, and preemption is likely with long prompts.
- TP=2: weights are split, so far more memory is left for KV cache. That gives more concurrent sequences, longer contexts and lower per-token latency, at the cost of all-reduce overhead and a shared failure domain.
If the replicas are memory-starved, TP=2 often wins on both latency and throughput. If they have reasonable headroom, replicas usually give more total throughput. Prove it with a load test at your real lengths.
Routing, autoscaling and operations
31. Why is round-robin load balancing a poor fit for LLM servers, and what is KV-cache-aware routing?
Answer: Round-robin assumes requests cost about the same. LLM requests vary enormously in prompt and output length, and each replica holds a different set of cached prefixes. Round-robin therefore overloads some replicas while others sit idle, and it throws away cache hits. Better strategies, roughly in order of sophistication:
- Least outstanding requests or queue-aware routing, using engine metrics such as waiting requests and KV cache usage.
- Prefix- or KV-cache-aware routing: send a request to the replica most likely to hold its prefix already, while weighing that replica's load so hot prefixes do not create hot spots.
- Adapter-aware routing for multi-LoRA: prefer replicas that already have the adapter loaded.
On Kubernetes, the Gateway API Inference Extension (a Kubernetes SIG project) adds an InferencePool resource and an endpoint picker. The picker reads model-server metrics, including prefix-cache and LoRA information, to choose the endpoint behind gateways such as Envoy Gateway, kgateway or GKE Gateway. llm-d builds on this idea with precise prefix-cache-aware routing. The cluster side is covered in our Kubernetes for AI interview questions.
32. Which signals should drive autoscaling of an inference deployment?
Answer: Use signals that show saturation before users feel it: the number of waiting requests, KV cache usage, running requests per replica, and p95 TTFT or queue time. vLLM exposes vllm:num_requests_waiting, vllm:num_requests_running and vllm:kv_cache_usage_perc. CPU utilisation tells you almost nothing about a GPU pod. GPU utilisation also gives a false picture, because it reports whether any kernel was running, not how much of the GPU's capacity was in use. Feed these metrics to the HPA through a Prometheus adapter or KEDA. Scale up quickly, scale down slowly to avoid flapping, and set the scale-up threshold so a new replica is ready before the queue hurts. That means accounting for cold-start time (Q33).
33. What makes up an inference cold start, and how do you shorten it?
Answer: A new replica has to (1) get a GPU node, which may mean provisioning or waiting for capacity, (2) pull a container image that is often very large, (3) fetch model weights from object storage or a model hub, (4) load the weights into GPU memory, and (5) warm up, which can include kernel compilation, CUDA graph capture and memory profiling. Each step can take minutes for large models. Ways to shorten it:
- Keep a warm minimum of replicas (or a pre-provisioned warm node pool) for interactive routes. Scale-to-zero is for batch work.
- Stage weights on fast local or shared storage close to the GPUs, in safetensors format, and pre-fetch them when a node joins.
- Use compilation caches where the engine supports them, and run a short warm-up request before marking the pod ready.
- Reserve GPU capacity for production, because on-demand GPUs are not always available in your region.
34. How do you protect an inference service from overload?
Answer: Use admission control in layers. At the gateway: per-tenant rate and token limits, caps on prompt and output length, and priority classes so interactive traffic outranks batch jobs. At the engine: a bounded queue. When the queue is full or the predicted wait would break the SLO, fail fast with a 429 and a retry-after hint, instead of letting every request time out slowly. On the client: retries with exponential backoff and jitter, and a cap on retries. A shared LLM gateway is the natural place for quotas, routing, fallback to another model or a managed API, and audit logging. A clean error for a few users beats a missed SLO for everyone.
35. How should health checks be designed for model servers?
Answer: Use three probes. A startup probe with a generous budget, because loading a large model is slow and a liveness probe would otherwise kill it mid-load. A readiness probe that passes only once the model is loaded and has served a warm-up request, and that fails while draining. A liveness probe that checks the process is responsive, but not so strictly that a busy engine gets restarted under load, which would make an overload worse. For tensor-parallel or multi-node deployments, remember that one failed GPU or rank brings down the whole group. Treat the group as one unit for health, alerting and replacement, and watch GPU health signals such as Xid errors and ECC events through DCGM.
36. What would you put on a serving dashboard?
Answer: Four groups, each labelled by model, route and tenant:
- User experience: p50, p95 and p99 TTFT, ITL and end-to-end latency, plus error and 429 rates, and goodput against the SLO.
- Engine saturation: running and waiting requests, KV cache usage, preemptions, prefix cache hit rate (vLLM exposes prefix cache query and hit counters), and speculative acceptance rate if enabled.
- Traffic shape: prompt and output token distributions and requests per second. Many incidents are really a change in traffic shape.
- Hardware: GPU memory, SM activity, temperature, power and errors from DCGM.
Add request tracing (OpenTelemetry) across gateway, retrieval and engine so you can see which stage owns the latency. This connects to the wider practice in AI observability.
37. Managed model API or self-hosted serving: how do you decide?
Answer:
| Factor | Managed API | Self-hosted |
|---|---|---|
| Model quality ceiling | Frontier proprietary models available | Limited to open-weight models you can run |
| Operations | Provider runs GPUs, scaling and upgrades | You own GPUs, engines, on-call and capacity |
| Cost shape | Pay per token; reserved or provisioned throughput options | Pay for GPU time whether used or idle |
| Latency control | Subject to shared quotas and rate limits | Full control of batching, placement and SLOs |
| Data and network | Private endpoints and regional options; check terms | Can run fully inside your network or air-gapped |
| Customisation | Provider-supported fine-tuning only | Any adapter, quantisation, decoding constraint |
Many enterprises run both behind one gateway: managed APIs for general reasoning and spiky traffic, and a self-hosted open-weight model for steady, narrow or sensitive workloads. Licences matter for the self-hosted side; see open-weight LLMs for enterprise.
38. How do you roll out a new model version or engine upgrade safely?
Answer: Treat it like a database migration, not a config tweak. Pin everything: engine version, model revision, tokenizer and chat template, quantisation artefact. Run the offline evaluation suite and the API contract tests (Q9) first. Then load-test the new build at the production traffic shape, because engine upgrades change defaults and scheduling behaviour. Deploy as a canary behind the gateway with a small traffic share. Compare quality signals, latency percentiles, error rates and output length with the old version, and keep the old deployment warm for instant rollback. Shadow traffic (copy requests to the new version, compare offline) catches regressions without user impact.
Want to practise serving, scaling and evaluating models end to end rather than memorising definitions? Cloudsoft's APEX AI, ML, Cloud and Cyber Security program combines model fundamentals with cloud deployment, Kubernetes and security, in classroom sessions in Ameerpet or live online.
Capacity planning and cost
39. Walk through a capacity plan for a self-hosted chat assistant.
Answer: Every number below is a hypothetical input. Replace each one with your own measurements.
- Memory per replica. A GPU with 80 GiB, with the engine allowed to use 0.9 of it, gives 72 GiB. A model of about 14 billion parameters at 16-bit needs about 28 GiB for weights. Reserve about 4 GiB for activations and runtime. That leaves about 40 GiB for KV cache.
- Tokens that fit. At 160 KiB per token (the configuration in Q3): 40 GiB / 160 KiB β 262,000 tokens of cache.
- Concurrent sequences. If a typical conversation holds 4,000 tokens (prompt plus output), that is about 65 sequences in theory. Size against a long-tail length such as 8,000 tokens and plan for about 32.
- Demand. Peak arrival of 3 requests per second with an average duration of 12 seconds gives about 36 concurrent requests (Little's law, Q8).
- Measured ceiling. Suppose a load test shows one replica meets the TTFT and TPOT SLOs up to 24 concurrent requests. That is often below the memory limit, because compute and scheduling saturate first.
- Replicas. 36 / 24 = 1.5, so 2 replicas, plus one more for failure tolerance and rolling deploys: 3.
Interview tip: Stress step 5: memory sets the theoretical ceiling, the measured latency SLO sets the real one.
40. How do you estimate cost per million tokens for a self-hosted model?
Answer: Cost per million output tokens = (GPU cost per hour Γ GPUs) Γ· (output tokens served per hour Γ· 1,000,000). The trap is utilisation. Take a purely illustrative placeholder of βΉ250 per GPU-hour and a 2-GPU deployment, so βΉ500 per hour. Suppose it sustains 1,500 output tokens per second at the SLO, which is 5.4 million tokens per hour. Running flat out, that is about βΉ93 per million output tokens. If average daily demand only uses 0.4 of that capacity, the effective cost is about βΉ231 per million, two and a half times higher, because idle GPUs still bill. Then add engineering, on-call, storage, networking and spare capacity. Compare the result honestly with managed per-token pricing at your traffic shape. Steady, high, predictable traffic favours self-hosting; spiky or low traffic favours paying per token. For the FinOps side, see cloud cost optimisation for AI and our FinOps interview questions.
41. Why plan for goodput instead of throughput?
Answer: A deployment can show impressive tokens per second while most interactive users miss their TTFT target, because the scheduler packed huge batches. Goodput counts only requests that met every SLO. So it captures the real trade-off: as load increases, throughput keeps rising for a while, but goodput peaks and then falls once queueing and preemption push requests past their targets. Write SLOs per workload class, for example "p95 TTFT under T1 and p95 TPOT under T2 for chat; completion within a time window for overnight batch". Report capacity as "requests per second at target goodput", not as peak tokens per second.
42. How do you load-test an LLM serving stack properly?
Answer:
- Replay real distributions of prompt and output length, taken from logs with sensitive data removed. Fixed-length synthetic prompts give numbers that do not reflect production.
- Use an open-loop arrival process (requests arrive on a schedule regardless of responses), because closed-loop clients slow down when the server slows down and hide queueing.
- Ramp the load in steps and hold each step long enough for the KV cache and queues to settle. Find the knee where goodput starts to fall.
- Include realistic prefix sharing, and also run a cold-cache test.
- Record TTFT, ITL and TPOT percentiles, preemptions, KV cache usage and errors at every step, with the engine version and flags.
- Kill a replica mid-test and watch latency.
43. How do long contexts change capacity and cost?
Answer: KV cache grows linearly with context, so a replica that serves many short chats may serve only a handful of long-document sessions. Attention compute during prefill grows faster than linearly with prompt length, so TTFT rises steeply for very long prompts. Practical controls: set max_model_len to what the product actually needs, enforce prompt limits at the gateway, route long-context work to a dedicated pool (possibly with more GPUs per replica or context parallelism), use prefix caching for repeated documents, and ask whether retrieval could send fewer, better chunks instead of entire documents.
Real-world scenario questions
44. Scenario: TTFT p95 jumps several-fold every weekday from 10 to 11 a.m., while median TTFT barely moves. What do you do?
Answer: A stable median with a broken tail at a predictable time points to queueing or interference at peak load, not to a slower model. I would first separate queue time from prefill time.
What I would check:
- Waiting requests and KV cache usage per replica during the window. Is the deployment saturated, or is load uneven across replicas (a routing problem)?
- Prompt-length distribution at peak. A batch job or a new feature sending long documents at the same hour would cause this.
- Preemption counters. Cache pressure causes recompute pauses.
- Prefix cache hit rate. Did a prompt change (a timestamp moved to the top, say) or a scale-out reduce reuse?
- Autoscaler behaviour. Did it react on a lagging signal, and how long did new replicas take to become ready?
- Chunked prefill budget and per-tenant fairness. Is one tenant's traffic starving others?
Production consideration: Fixes usually combine pre-scaling on a schedule before the known peak, autoscaling on queue depth, moving batch work off the interactive pool, prefix-friendly prompt ordering and load-aware routing.
45. Scenario: after enabling 128K-token context for a contract-review tool, the service fails to start on some nodes and crashes or stalls on large uploads. Why?
Answer: Engines such as vLLM pre-allocate KV cache and check at startup that at least one maximum-length sequence fits. If max_model_len is raised beyond what the remaining memory can hold, startup fails. Nodes with smaller GPUs, or with another process already holding GPU memory, fail first. At runtime, very long prefills raise activation memory, and a few huge sequences take most of the cache, which causes preemption storms for everyone else.
What I would check:
- The startup error and the engine's reported KV cache capacity on each node type.
- Whether another process is using the GPU, and whether
gpu_memory_utilizationassumes the GPU is exclusive. - Chunked prefill settings and
max_num_batched_tokensfor the long prompts. - Preemption counts and tail latency for other users when a large upload runs.
- Whether the product truly needs 128K tokens, or whether retrieval over clauses would do.
Production consideration: Put long-context traffic on its own pool with enough memory (TP across GPUs, optional FP8 KV cache after evaluation), keep the default chat pool at a modest context length, and enforce size limits at the gateway with a clear error message for users.
46. Scenario: you switched to a 4-bit quantised model to save GPUs, but throughput under load went down. What happened?
Answer: Weight-only 4-bit reduces memory traffic, which helps decode at small batch sizes. At high concurrency the workload becomes compute-bound, and the extra work of converting weights back to 16-bit can make large batches slower than an FP8 or 16-bit path running on native tensor cores.
What I would check:
- Which kernel the engine actually selected. Is the quantisation format supported by fast kernels on this GPU generation, or did it fall back to a slower path?
- Throughput and latency at several concurrency levels, old versus new. The cross-over point usually explains it.
- Output length. A quantised model that writes longer answers or repeats itself lowers requests per second even if tokens per second held.
- Speculative decoding acceptance rate, if a draft model is still tuned to the old target.
- Whether engine flags changed during the switch (max sequences, CUDA graphs, KV cache dtype).
Production consideration: Pick quantisation by workload: weight-only for latency-oriented, memory-constrained serving; FP8 W8A8 for high-throughput serving on GPUs that support it. Re-run the task evaluation (Q25) whichever you choose.
47. Scenario: you must choose GPUs for an internal assistant with a strict p95 TTFT and smooth streaming. How do you decide?
Answer: Translate the SLO into hardware requirements, shortlist, then measure. Do not pick from spec sheets alone.
What I would check:
- Memory capacity: weights at the chosen precision plus KV cache for the target concurrency and context (Q39). This decides GPUs per replica.
- Memory bandwidth: this largely sets decode speed, and therefore TPOT and streaming smoothness.
- Compute and precision support: prefill speed (TTFT) and whether FP8 or FP4 kernels are available.
- Interconnect: if TP is needed, NVLink-class links between GPUs in a node.
- Availability and quota in the regions your data rules require, for example Indian cloud regions for a GCC serving Indian customers, plus reservation options.
- Cost at your expected utilisation (Q40), not list price per hour.
Production consideration: Load-test the shortlisted candidates with your real traffic shape and engine, compare goodput per rupee, and keep the choice reversible. Our GPU basics guide explains the specs in depth.
48. Scenario: you scaled a deployment from 2 to 8 replicas, and average TTFT got worse. How is that possible?
Answer: With random or round-robin routing, every replica now sees a quarter of the repeats of any given prefix that it saw before. Prefix cache hit rates fall, and more prompts need full prefill. New replicas also start with empty caches.
What I would check:
- Prefix cache hit rate per replica before and after the change.
- The load balancer's algorithm and whether it is aware of sessions or prefixes.
- Whether new replicas joined before warm-up finished (a readiness probe problem).
Production consideration: Use prefix- or KV-cache-aware routing (Q31) with load balancing built in, so a popular system prompt does not pin all traffic to one replica. Consider a shared or tiered prefix cache for very common prefixes.
49. Scenario: in a hospital's clinical-notes assistant, streaming freezes for a second or two whenever someone uploads a long discharge summary. What is happening?
Answer: Long prefills are interfering with decodes. Every active user's next token waits while the engine processes the big prompt in the same steps.
What I would check:
- ITL p99 correlated with prompt-length spikes.
- Whether chunked prefill is enabled, and the size of the per-step token budget.
- Whether uploads and chat share one pool.
Production consideration: Lower the chunked-prefill budget to protect ITL. If that is not enough, separate document ingestion from chat into different pools, or adopt disaggregated prefill and decode where scale justifies it (Q29). Put ITL p99 on the dashboard next to TTFT.
50. Scenario: a multi-LoRA deployment serving 40 business-unit adapters has good latency for popular units but poor latency for rarely used ones. Why?
Answer: Rare adapters are probably being swapped in and out. When more distinct adapters appear in a batch than the active-adapter limit allows, requests wait for a slot, and adapters that are not cached in CPU memory are loaded from storage on demand.
What I would check:
- The active-adapter limit and CPU adapter cache size against the number of distinct adapters seen in a typical minute.
- Where adapters are fetched from, and how long a cold load takes.
- Whether routing spreads every adapter across every replica, instead of grouping adapters onto replicas.
- Rank differences: one high-rank adapter can force larger buffers for all of them.
Production consideration: Use adapter-aware routing, pre-load adapters onto assigned replicas, size limits from real traffic, and consider a merged dedicated deployment for the few adapters with heavy traffic.
51. Scenario: an insurer's claims extraction service uses JSON-schema guided decoding, but some outputs are still invalid and the first request after each deploy is very slow. Explain both.
Answer: Invalid outputs under guided decoding almost always mean truncation: generation hit max_tokens or a stop sequence before the object closed. The slow first request is schema compilation, because the grammar for a large schema is built the first time it is seen.
What I would check:
- Finish reasons on failed outputs (length versus stop).
- Whether the schema allows unbounded arrays or long free-text fields that inflate output length.
- Which structured-output backend is in use, and whether it supports every schema feature you rely on.
Production consideration: Raise output limits based on the real output-length distribution, bound array sizes and string lengths in the schema, send a warm-up request with each schema on startup, and keep server-side validation with a retry path, because a valid shape can still hold a wrong value.
52. Scenario: a bank's branch-staff assistant scales to zero overnight. Every morning when branches open, the first users wait minutes. What do you change?
Answer: This is a cold start (Q33) landing on a predictable peak. Scale-to-zero fits batch work, not a known daily opening rush.
What I would check:
- The time spent in each stage: node provisioning, image pull, weight download, load and warm-up.
- Whether GPU capacity was even available at that hour without a reservation.
Production consideration: Scale up on a schedule before opening time, keep a minimum warm replica during business hours, stage weights near the GPUs and pre-pull images. If cost matters overnight, scale down to one small replica rather than zero, and send the rare overnight request to a managed API fallback through the gateway.
53. Scenario: a GCC's internal assistant runs on a managed API and starts hitting 429 rate-limit errors during month-end close. What are the options?
Answer: The provider's quota, not your code, is the bottleneck. Options range from cheap to structural.
What I would check:
- Which limit is hit (requests per minute or tokens per minute) and for which model and region.
- Whether retries are amplifying the problem (no backoff, retries at every layer).
- Token waste: oversized prompts, missing prompt caching, agents making redundant calls.
- Which traffic is interactive and which could be queued or batched.
Production consideration: Add backoff with jitter and per-team quotas at the gateway, cut tokens per request, request a quota increase or reserved or provisioned throughput, spread load across approved regions or deployments, and route deferrable or narrow tasks to a smaller or self-hosted model. A gateway makes all of these changes possible without touching client code.
54. Scenario: the dashboard shows GPU utilisation near maximum, yet the service misses its latency SLO and throughput is lower than the load test promised. What is going on?
Answer: The common GPU utilisation metric means "some kernel was running during the sample", not "the GPU's compute is fully used". A GPU can show near-maximum utilisation while it is badly underused: stuck in memory-bound decode with tiny batches, waiting on communication, or redoing preempted work.
What I would check:
- DCGM metrics for SM activity and memory bandwidth rather than headline utilisation.
- Batch size in practice: running sequences per step compared with the load test.
- Preemption and recompute counts, and KV cache usage.
- Traffic shape compared with the load test (longer outputs, fewer shared prefixes).
- For TP deployments, communication time and whether the GPUs are on the expected interconnect.
Production consideration: Build dashboards and autoscaling on engine metrics (queue, cache, latency percentiles), and re-run the load test whenever production traffic shape drifts away from the test profile.
55. Scenario: design a shared LLM serving platform for several product teams in a Hyderabad GCC.
Answer: I would start by asking what the teams need. Requirements first: workload classes (interactive chat, agents, batch), latency SLOs per class, data classification and residency rules, which models are needed (managed and open-weight), expected traffic, and who will be on call. Then a layered design:
apps / agents (OpenAI-compatible SDKs)
|
LLM gateway: auth, quotas, routing,
fallback, logging, cost attribution
| |
managed APIs self-hosted router
(private endpoints) (KV / LoRA aware)
| |
chat pool batch pool
(warm min) (scale to 0)
|
engine pods on GPU node pools
(pinned versions, metrics, traces)
Key decisions: a single gateway and API contract for every team; separate pools by workload class so batch never hurts interactive traffic; queue-aware autoscaling with warm minimums; KV- and adapter-aware routing; multi-LoRA for team-specific variants; per-team quotas and showback of cost; a model registry with pinned artefacts and evaluation gates before promotion; and dashboards built on goodput. The trade-off: a shared platform improves utilisation and governance, but needs clear priority rules and an owning team.
Production consideration: Ship it in phases. Start with the gateway and one managed model, add one self-hosted pool for the highest-volume narrow task, and grow from there. For the wider system-design view, see our AI system design interview questions.
Key takeaways
- Prefill is compute-bound and drives TTFT. Decode is memory-bandwidth-bound and drives ITL. Most serving techniques manage the conflict between the two.
- KV cache memory, not weights, usually caps concurrency. PagedAttention, prefix caching, KV quantisation and context limits all manage that budget.
- Plan and alert on goodput and tail percentiles (p95, p99) of TTFT and ITL, measured under realistic load, not on averages or peak tokens per second.
- Quantisation and speculative decoding are workload-dependent. Weight-only helps memory-bound decode; W8A8 and FP8 help compute-bound load; speculation helps most at low load.
- Use replicas for throughput, tensor parallelism when memory- or latency-bound, and disaggregated prefill and decode only when scale and strict latency SLOs justify it.
- Routing and autoscaling must be LLM-aware: route on queue, prefix cache and adapters; scale on queue depth and cache usage; plan around cold starts.
- Cost per million tokens is dominated by utilisation, so compare self-hosting with managed APIs at your real traffic shape.
Interview preparation checklist
- Serve a small open model with vLLM or SGLang on one GPU and call it with an OpenAI-compatible client, including streaming and a JSON-schema request.
- Scrape the engine's Prometheus metrics and build a small dashboard: TTFT, ITL, waiting requests, KV cache usage, preemptions, prefix cache hits.
- Run an open-loop load test with realistic prompt and output lengths, and find the goodput knee.
- Change one setting at a time (max sequences, batched-token budget, context length) and explain each effect.
- Compare a 16-bit and a quantised version of the same model on a small task evaluation and at two concurrency levels.
- Calculate KV cache per token for a model from its config file, and do a full capacity plan on paper (Q39).
- Build a one-page cost-per-million-tokens sheet with utilisation as an input.
- Practise the scenario questions aloud using the "What I would check" structure.
- Review neighbouring guides: MLOps and LLMOps interview questions and Kubernetes for AI interview questions.
FAQ
What skills are needed for an LLM inference or model serving engineer role?
Solid Python, a working understanding of transformer inference (prefill, decode, KV cache), hands-on experience with at least one serving engine such as vLLM or SGLang, Kubernetes and GPU operations, observability with Prometheus and tracing, and the ability to reason about latency, capacity and cost.
Do I need to write CUDA kernels to work in LLM serving?
Not for most roles. Platform and application-facing serving engineers configure, scale and debug engines. Kernel development is a specialised track. Understanding why kernels are memory-bound or compute-bound is expected, though.
How should I prepare for vLLM interview questions?
Read the official vLLM documentation on PagedAttention, prefix caching, chunked prefill, optimisation and metrics, then run it yourself, put load on it, and explain what each metric shows as concurrency rises.
Can I practise LLM serving without an expensive GPU?
Yes. Start with small models on llama.cpp on a laptop to learn the APIs, then rent a single cloud GPU for short sessions to practise vLLM or SGLang load tests. Shut it down after each session.
Is Hugging Face TGI still worth learning?
Its documentation says TGI is in maintenance mode and recommends engines such as vLLM and SGLang for new work. Know what it does, but spend your hands-on time on actively developed engines.
How is an LLM serving interview different from a general MLOps interview?
MLOps interviews cover the whole model lifecycle. LLM serving interviews go deeper into GPU memory, batching, caching, latency percentiles, parallelism and cost per token for generative models.
Are LLM inference skills useful for freshers?
Yes, if you can show them. A fresher who has deployed an open model, load-tested it and explained the results stands out, because most candidates have only used hosted APIs.
Is model serving a good career path for DevOps and cloud engineers?
It is a natural extension. Kubernetes, autoscaling, observability and capacity planning carry over directly. The new parts are GPU behaviour and how inference engines schedule work.
How long does it take to prepare for an LLM serving interview?
It depends on your background. Engineers who already know Kubernetes and Python often need a few focused weeks of hands-on serving practice. Freshers should budget longer and build one end-to-end project.
If you want guided, hands-on preparation for serving, scaling and securing AI systems on cloud platforms, explore Cloudsoft's APEX program for AI, ML, Cloud and Cyber Security. If you would rather deploy and integrate AI inside customer environments, the AI Forward Deployed Engineer FDE PRO program adds enterprise projects, a simulated customer capstone and placement support until you're placed. Classroom in Ameerpet or live online; call +91 96660 19191 for a free demo.



