New batches starting this week Β· Limited seats

LLM Interview Questions and Answers 2026 (75 Questions)

75 large language model interview questions with clear, accurate answers, from attention and tokenization to fine-tuning, inference optimisation, evaluation and production scenarios.

LLM interview questions 2026: 75 questions on transformers, tokenization, fine-tuning, KV cache, quantization and evaluation
Last updated Β· 44 min read Β· 9,670 words

LLM interview questions in 2026 go well beyond "what is a large language model": interviewers want to know whether you understand how attention, tokenization, training, decoding and serving actually work, and whether you can use that understanding to debug cost, latency and quality problems in production. This guide collects 75 high-value large language model interview questions with model answers, from transformer basics to fine-tuning, evaluation, reasoning models and real-world scenarios. The maths stays accessible but precise enough for follow-ups.

How to use this guide

This page goes deeper into model internals than our GenAI engineer interview questions, which focus on building applications. If you need a refresher on the basics first, read what an LLM is and how it works. What interviewers typically test at each level:

  • Freshers and junior engineers: clear explanations of attention, tokens, pretraining versus instruction tuning, and sampling parameters.
  • Mid-level engineers: KV cache, quantization, LoRA, evaluation design and hallucination causes and their trade-offs.
  • Senior and architect roles: serving economics, long-context behaviour, mixture of experts, preference tuning, and structured diagnosis of production incidents.

Try answering each question aloud before reading the model answer.

Transformer fundamentals

1. What is a large language model, in one precise sentence?

Answer: A large language model is a neural network, almost always a transformer, trained on a very large text corpus to predict the next token given the previous tokens, and then usually further trained to follow instructions and match human preferences. Its "knowledge" is stored in its weights as statistical patterns, not as a database of facts. It generates text one token at a time by sampling from a probability distribution over its vocabulary.

2. Explain the attention mechanism in plain language.

Answer: Attention lets each token build its new representation by looking at other tokens and deciding how much each one matters. Every token is projected into three vectors: a query (what am I looking for), a key (what do I offer) and a value (what information I pass on). The score between two tokens is the dot product of one token's query with the other's key. Scores are scaled, turned into weights with softmax so they sum to one, and used to take a weighted average of the value vectors. In matrix form: Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V.

Real-world example: In "The bank approved the loan because it met the criteria", attention helps the representation of "it" draw heavily on "loan" rather than "bank".

3. What is the difference between self-attention and cross-attention?

Answer: In self-attention, queries, keys and values all come from the same sequence, so tokens attend to each other within one input. In cross-attention, queries come from one sequence and keys and values from another. Encoder-decoder models such as the original transformer for translation use cross-attention so the decoder can attend to the encoded source sentence. Decoder-only LLMs use only (masked) self-attention; the prompt and the generated text are one continuous sequence.

4. Why is the dot product divided by the square root of the key dimension?

Answer: As the vector dimension grows, the variance of a dot product between random vectors grows with it, so raw scores get large. Large inputs push softmax towards a near one-hot output where gradients are tiny, which makes training slow and unstable. Dividing by sqrt(d_k) keeps the scores in a range where softmax stays smooth.

5. What is multi-head attention, and what are multi-query and grouped-query attention?

Answer: Multi-head attention runs several attention operations in parallel, each with its own learned projections, so different heads can capture different relationships (syntax, coreference, position patterns). In multi-query attention (MQA), all query heads share a single key and value head; in grouped-query attention (GQA), query heads are split into groups that each share one key/value head. Both shrink the KV cache, which is the main memory cost during generation, with GQA usually losing less quality than MQA. Fewer KV heads means more concurrent users or longer context on the same GPU, which is why many recent open-weight models use GQA.

6. Why do transformers need positional encoding, and what are the main approaches?

Answer: Self-attention compares content, not position. Without positional information, "dog bites man" and "man bites dog" would produce the same set of attention scores (the causal mask gives only a weak implicit signal). Main approaches:

  • Sinusoidal encodings (original transformer): fixed sine and cosine patterns added to the embeddings.
  • Learned absolute embeddings: one trained vector per position; simple but tied to the trained maximum length.
  • Rotary position embeddings (RoPE): rotate query and key vectors by an angle that depends on position, so their dot product depends on relative distance. Widely used in modern LLMs.
  • ALiBi: adds a distance-based penalty directly to attention scores.

Relative schemes like RoPE matter for long context because they can be stretched with techniques such as position interpolation, covered in Q37.

7. Compare encoder-only, encoder-decoder and decoder-only architectures.

Answer:

ArchitectureAttentionTypical trainingTypical use
Encoder-only (BERT-style)BidirectionalMasked token predictionClassification, embeddings, rerankers
Encoder-decoder (T5-style)Bidirectional encoder, causal decoder with cross-attentionSpan corruption, sequence-to-sequenceTranslation, summarisation
Decoder-only (GPT-style)Causal (each token sees only earlier tokens)Next-token predictionGeneral chat and generation LLMs

Decoder-only won for general LLMs largely because one simple objective scales well, every token in the corpus gives a training signal, and the same model handles any task phrased as text continuation. Encoder models remain very common for embeddings and reranking in RAG pipelines.

8. Walk through one forward pass of a decoder-only transformer.

Answer: Text is tokenized into IDs, each ID is looked up in an embedding table, and position information is applied (for RoPE, inside attention). The vectors pass through N identical blocks. Each block applies normalisation, masked multi-head self-attention, a residual connection, then normalisation, a feed-forward network (or MoE layer) and another residual. After the last block, a final norm and an "unembedding" projection produce one score (logit) per vocabulary token. Softmax turns the logits for the last position into a probability distribution, and the decoding strategy picks the next token.

token IDs -> embeddings (+ position info)
              |
   +----------v-----------+
   | norm -> masked       |
   | multi-head attention |   repeated N times
   | + residual           |
   | norm -> FFN or MoE   |
   | + residual           |
   +----------+-----------+
              v
   final norm -> unembedding -> logits
              v
   softmax -> decoding picks next token

9. What do the feed-forward layers, residual connections and layer normalisation contribute?

Answer: Attention mixes information between tokens; the feed-forward network (FFN) transforms each token's vector independently and holds a large share of the parameters. Modern models often use gated variants such as SwiGLU. Residual connections add each sub-layer's input to its output, which keeps gradients flowing through very deep stacks and lets layers learn small refinements. Normalisation (LayerNorm or the cheaper RMSNorm, usually applied before each sub-layer, "pre-norm") keeps activations at a stable scale so training does not diverge.

10. Why does attention cost grow quadratically with sequence length?

Answer: Every token computes a score against every other token it can see, so a sequence of n tokens produces an n by n score matrix per head per layer. FlashAttention computes exact attention in tiles without storing the full matrix in GPU memory, which cuts memory traffic and speeds things up, but the arithmetic is still quadratic. Separately, the KV cache grows linearly with length. Both effects explain why long prompts are slow and expensive.

Tokenization

11. What is a token, and why don't LLMs just use words or characters?

Answer: A token is a unit from the model's fixed vocabulary: a whole common word, a piece of a word, punctuation, a byte, or sometimes a run of spaces. Whole words would need an enormous vocabulary and still fail on new names and typos; characters would make sequences very long and expensive. Subword tokenization is the compromise: frequent strings get their own tokens, rare strings are split into pieces, and nothing is ever "unknown". Tokens are also the unit of pricing and context limits, which our guide to tokens and context windows covers in detail.

12. How does Byte Pair Encoding (BPE) work?

Answer: BPE builds a vocabulary by repeated merging. Start with characters or, in byte-level BPE, all 256 byte values. Count every adjacent pair in the training corpus, merge the most frequent pair into a new token, and repeat until the vocabulary reaches the target size. Byte-level BPE means any UTF-8 text, emoji or code can be represented, because in the worst case it falls back to individual bytes.

Interview tip: Contrast BPE with WordPiece (likelihood-based merges, used in BERT) and SentencePiece Unigram (starts large and prunes tokens).

13. What trade-offs come with vocabulary size?

Answer: A larger vocabulary produces shorter sequences, so prompts are cheaper and faster to process and more text fits in the context. The costs are a larger embedding and output layer, a more expensive softmax over the vocabulary, and many rare tokens that receive little training signal. Multilingual models tend to use larger vocabularies so non-English scripts are not split into tiny fragments.

14. Which model behaviours can be explained by tokenization?

Answer: Several well-known quirks come from the model seeing tokens rather than characters:

  • Counting letters or reversing words is hard because the model may never see individual letters.
  • Arithmetic on long numbers is unreliable when digits are chunked inconsistently.
  • An extra space at the start or end changes the tokens, which can change output for "the same" prompt.
  • Non-English text often uses more tokens per word, raising cost and filling context faster.

The practical fix is often a tool (calculator, code execution).

15. Scenario: a Telugu and Hindi customer-support bot costs far more per conversation than the English version. Why, and what can you do?

Answer: The likely cause is tokenizer efficiency. Many tokenizers were trained on English-heavy data, so Indic scripts are split into more tokens per word. The same conversation therefore uses more input and output tokens, costs more, runs slower and fills the context sooner.

What I would check:

  1. Measure tokens per message for each language using the provider's tokenizer or token-counting API, on real transcripts.
  2. Compare candidate models; tokenizer efficiency for Indic languages varies a lot between model families.
  3. Trim repeated system prompts and conversation history, and use prompt caching for static instructions.

Production consideration: Choose models on cost per resolved conversation in each language, not on the per-token price list. Evaluate quality per language too.

Pretraining, instruction tuning and preference tuning

16. What happens during pretraining?

Answer: The model learns to predict the next token over a huge, mixed corpus (web text, books, code, papers and more). The loss is cross-entropy: the negative log probability the model assigned to the actual next token, averaged over positions. Before training, data is deduplicated, filtered for quality and safety, and mixed in chosen proportions. This is where the model acquires language, world knowledge, coding and reasoning patterns, and it is by far the most expensive stage. The result is a base model that continues text well but does not reliably follow instructions.

17. What do scaling laws tell us?

Answer: Empirical scaling laws show that pretraining loss falls smoothly and predictably as you increase parameters, training tokens and compute, following roughly a power law. Compute-optimal research showed that, for a fixed training budget, you should grow data alongside model size rather than only adding parameters. In practice, many teams deliberately train smaller models on more tokens than "compute-optimal", because a smaller model is cheaper to serve for its whole life.

18. What is instruction tuning (supervised fine-tuning), and why is it needed?

Answer: Supervised fine-tuning (SFT) trains the base model on curated examples of prompts paired with good responses, often formatted as multi-turn conversations. The loss is usually computed only on response tokens. SFT teaches the model the assistant format: answer the question asked, follow instructions, use the chat template and special tokens, and stop at the right point. It needs far less data than pretraining, and quality and diversity matter more than volume.

19. What is a chat template, and why does it matter in practice?

Answer: A chat template is the exact text format, including special tokens, that marks system, user, assistant and tool turns for a given model. The model was fine-tuned on that format, so it relies on it to know who is speaking and when to stop. If you self-host and format prompts incorrectly (wrong role markers, missing end-of-turn token), quality drops sharply or the model rambles past the answer and never stops, a common debugging question. Hosted APIs apply it for you from structured messages.

20. Explain RLHF step by step.

Answer: Reinforcement learning from human feedback has three classic stages:

  1. SFT: start from an instruction-tuned model.
  2. Reward model: generate several responses per prompt, have people rank or pick the preferred one, and train a reward model to give higher scores to preferred responses.
  3. RL optimisation: fine-tune the policy (the LLM) with a reinforcement learning algorithm, classically PPO, to maximise reward, with a KL-divergence penalty that keeps it close to the SFT reference model.

The KL penalty matters: without it, the model drifts towards strange outputs that exploit weaknesses in the reward model, known as reward hacking. PPO-based RLHF is also heavy to run, with policy, reference, reward and usually value models in play at once.

21. What is DPO, and how is it different from RLHF with PPO?

Answer: Direct Preference Optimization trains on the same kind of data, a prompt with a preferred and a rejected response, but skips the separate reward model and the RL loop. Its authors showed the KL-constrained RLHF objective has a closed-form optimal policy, so reward can be written in terms of the policy. That turns preference learning into a simple classification-style loss: increase the likelihood of the preferred response relative to the rejected one, with both likelihoods measured against a frozen reference model. A parameter called beta controls how far the model may move from the reference.

DPO is simpler, more stable and cheaper to run, which is why it is popular for open-weight fine-tuning. It is offline: it learns from a fixed set of pairs, while PPO-style RL can sample new responses and get fresh rewards during training.

22. What are RLAIF and reinforcement learning from verifiable rewards?

Answer: RLAIF replaces some or all human preference labels with judgements from an AI model, often guided by a written set of principles. Anthropic's Constitutional AI is a well-known public example. It scales labelling but inherits the judge model's biases. Reinforcement learning from verifiable rewards uses an automatic check instead of a preference model: did the maths answer match, did the code pass the unit tests? Because the reward is hard to fake, it is a core ingredient in training reasoning models (Q39). Methods such as GRPO, which score several sampled answers to the same prompt against each other instead of training a value model, are common here.

23. Why does an LLM have a knowledge cutoff, and what are the implications?

Answer: The model only knows what was in its training data, which was collected up to some date. Later events and private company data are unknown to it. Implications: never rely on the model for current prices, regulations, product versions or policies; supply them through retrieval or tools; and put the current date in the system prompt when time matters. This is the core argument for RAG over fine-tuning for fast-changing knowledge, discussed in RAG vs fine-tuning.

Inference, decoding and optimisation

24. What are the prefill and decode phases of inference?

Answer: Prefill processes the whole prompt in one parallel forward pass, builds the KV cache and produces the first output token. It is compute-bound and largely determines time to first token (TTFT). Decode then generates one token per step, each step reusing the cache. It is usually memory-bandwidth-bound: for every token, the GPU has to read all model weights and the cache from memory while doing relatively little arithmetic. That is why output tokens are slower and usually priced higher than input tokens.

prompt (n tokens)
   |
   v
PREFILL: one parallel pass over the prompt
   builds KV cache, emits first token  (TTFT)
   |
   v
DECODE loop: one token per step
   reads weights + KV cache each step
   appends new K and V to the cache
   |
   v
stop token or max_tokens reached

25. What is the KV cache, and how big does it get?

Answer: During generation, the keys and values of earlier tokens do not change, so the model stores them instead of recomputing them for every new token. That is the KV cache. Its size per token is roughly 2 x layers x KV heads x head dimension x bytes per value. For an illustrative configuration of 32 layers, 8 KV heads, head dimension 128 and 16-bit values, that is 128 KiB per token, or about 4 GiB for a single 32,768-token sequence. Multiply by concurrent users and the cache, not the weights, often limits capacity.

26. How do temperature, top-k and top-p sampling work?

Answer: The model outputs logits; decoding decides which token to pick.

  • Temperature divides the logits before softmax. Below 1 sharpens the distribution (more predictable); above 1 flattens it (more varied). As temperature approaches zero, sampling approaches greedy decoding.
  • Top-k keeps only the k most likely tokens, renormalises, and samples among them.
  • Top-p (nucleus) keeps the smallest set of tokens whose cumulative probability reaches p, so the candidate set adapts: small when the model is confident, larger when it is unsure.

Use low temperature for extraction, classification and code; moderate values for drafting. Some providers fix or ignore these settings for certain models, especially reasoning models, so check current documentation.

27. Why isn't greedy decoding always the right choice, and what is beam search?

Answer: Greedy decoding picks the single most likely token at each step. It suits short, factual outputs, but in open-ended generation it often falls into repetitive loops and bland text, because a locally likely choice is not the globally most useful sequence. Beam search keeps several candidate sequences and expands the most probable ones; it helps in tasks with a narrow correct output, such as translation, but tends to produce generic text in chat and is rarely used for modern assistants.

28. How does speculative decoding speed up generation without changing the output?

Answer: A small, fast draft model proposes several tokens ahead. The large target model then checks all of them in one forward pass, which costs about the same as generating a single token because decode is memory-bound. Tokens are accepted left to right using a rejection-sampling rule; at the first rejection, a corrected token is sampled from the target model and drafting resumes. With the standard acceptance rule, the output distribution is identical to sampling from the target model alone; you only gain speed. The speed-up depends on how often the draft agrees with the target, so it works well on predictable text such as code or templated answers. Variants replace the draft model with extra prediction heads or n-gram lookup from the prompt.

29. What is quantization, and what are the main approaches?

Answer: Quantization stores weights (and sometimes activations and the KV cache) in fewer bits, for example 8-bit or 4-bit integers or 8-bit floating point, instead of 16-bit. A 7-billion-parameter model needs about 14 GB for weights at 16 bits, about 7 GB at 8 bits and about 3.5 GB at 4 bits, plus overhead. Because decode is memory-bandwidth-bound, smaller weights also mean faster generation. Approaches include post-training weight-only methods (GPTQ, AWQ), activation-aware schemes for 8-bit weight-and-activation inference, FP8 on GPUs that support it, and formats such as GGUF for CPU and edge runtimes.

Interview tip: Always add "and then I evaluate on my task"; general benchmark averages can hide quantization regressions in maths, code or non-English languages.

30. What is continuous batching, and why does it matter?

Answer: Static batching waits for a batch to fill and then runs until the longest request finishes, so short requests sit idle and GPUs waste capacity. Continuous (iteration-level) batching adds and removes requests at every decode step: as soon as one sequence finishes, a waiting one takes its slot. Combined with paged KV-cache memory (the PagedAttention idea popularised by vLLM), which allocates cache in blocks rather than one large contiguous region per request, it greatly improves throughput and reduces memory fragmentation. Serving engines such as vLLM, SGLang and TensorRT-LLM build on these ideas.

31. What is the trade-off between throughput and latency in LLM serving?

Answer: Larger batches make better use of the GPU, so total tokens per second (throughput) rises and cost per token falls. But each request shares compute with more neighbours, so per-user token speed and sometimes TTFT get worse. Overnight batch jobs can use large batches; interactive agents need smaller ones. Engineers tune batch and concurrency limits against latency targets such as p95 TTFT.

32. What metrics would you track for LLM inference performance?

Answer: At minimum:

  • Time to first token (TTFT): queueing plus prefill; drives perceived responsiveness.
  • Inter-token latency / time per output token: drives streaming speed.
  • End-to-end latency: roughly TTFT plus output tokens times inter-token latency, plus any tool calls.
  • Throughput: requests and tokens per second per GPU or per deployment.
  • Input and output token counts per request, plus cache hit rate if you use prompt caching.

Track percentiles (p50, p95, p99), not averages, and break them down by route, model and prompt version. Our guide on LLM latency optimisation goes deeper.

33. Scenario: users complain the assistant "feels slow" although average latency looks fine. What do you do?

Answer: Averages hide tails, and perceived speed depends mostly on time to first visible output. I would look at the distribution and at what the user actually sees.

What I would check:

  1. p95 and p99 TTFT and end-to-end latency, split by route and time of day (peak-hour queueing is common).
  2. Whether responses are streamed to the UI, or buffered until complete.
  3. Hidden steps before the model call: retrieval, reranking, guardrail checks, sequential tool calls.

Production consideration: Set latency objectives on percentiles and on TTFT specifically, and show progress (streaming or status messages) during multi-step agent work.

Context length and long-context issues

34. What is the context window, and what counts against it?

Answer: The context window is the maximum number of tokens the model can attend to in one request, covering everything: the system prompt, tool definitions, conversation history, retrieved documents, images converted to tokens, the user's message and the generated output (including reasoning tokens for reasoning models). If the total exceeds the limit, the request fails or something must be truncated. Large tool schemas and long histories can consume much of the window before the user says anything.

35. If a model supports a very long context, why not put every document in the prompt?

Answer: Because supported is not the same as used well. Long prompts cost more on every call, raise TTFT, and can lower accuracy: models are often weaker at using information buried in the middle of long inputs (the "lost in the middle" effect) and get distracted by irrelevant passages. Long context is great for a single long document the user explicitly provides; for large, changing, permissioned corpora, retrieval that selects the relevant pieces is usually more accurate and cheaper.

36. What is the "lost in the middle" problem, and how do you test for it?

Answer: Research has shown that many models retrieve and use facts placed at the beginning or end of a long context more reliably than facts in the middle. To test your own setup, run "needle" style tests, but more realistically: place the key clause of a real contract at different positions and depths, ask questions that require it, and measure accuracy by position. Mitigations include putting the most relevant chunks near the start or end, retrieving less but better, and summarising or splitting long documents into sections.

37. How do models get extended to longer context windows?

Answer: Training at full long length from the start is expensive, so models are usually pretrained at a shorter length and then extended. With RoPE, techniques such as position interpolation, NTK-aware scaling and YaRN rescale the rotation frequencies so longer positions map into ranges the model has seen, followed by a phase of continued training on long documents.

38. How do you manage context in a long-running chat or agent session?

Answer: Treat context as a budget. Common techniques: keep a fixed system prompt and cache it; keep recent turns verbatim and summarise older ones; store durable facts (user preferences, case details) in external memory and retrieve them when relevant; drop or compress bulky tool outputs after they are used; and load tool definitions only when needed.

Reasoning models and mixture of experts

39. What is a reasoning model, and how is it trained differently?

Answer: A reasoning model is an LLM trained to produce an extended internal chain of reasoning before its final answer, and to use that extra computation productively: break the problem down, try approaches, check intermediate results and backtrack. The key ingredient is large-scale reinforcement learning on tasks with verifiable outcomes such as maths and code, which rewards reasoning that actually reaches correct answers. The result is "test-time compute": you can buy better answers on hard problems by letting the model think longer. See reasoning models explained for a fuller treatment.

40. How is a reasoning model different from chain-of-thought prompting?

Answer: Chain-of-thought prompting asks an ordinary model to "think step by step" in its visible output. It often helps, but the model was not trained for long reasoning chains. A reasoning model has been optimised, through training, to generate and use long reasoning traces, so it typically handles multi-step problems, planning and self-correction better and needs less prompt scaffolding. Providers commonly advise giving reasoning models clear goals and constraints rather than step-by-step scripts.

41. When should you not use a reasoning model?

Answer: When the task is simple, latency-sensitive or high-volume: classification, extraction, routing, short FAQ answers and formatting. Reasoning tokens count as output tokens, so cost and latency rise, and extra thinking does not help when the bottleneck is missing information rather than hard logic. Many teams route requests: a fast model handles the bulk, a reasoning model handles hard or high-stakes cases, often with an adjustable reasoning effort or thinking budget where the provider supports it. Reasoning models can also overthink simple requests.

42. Is a model's visible chain of thought a faithful explanation of its decision?

Answer: Not reliably. Research has shown stated reasoning can omit factors that actually influenced the answer. Treat it as a debugging signal, not an audit-grade explanation. For regulated decisions, such as a loan or claim outcome, the explanation should come from the system's design: documented rules, cited source documents, logged tool results and human review, not from the model's narrative.

43. What is a mixture-of-experts (MoE) model?

Answer: In an MoE transformer, the feed-forward layer in some or all blocks is replaced by many parallel "expert" networks plus a small router. For each token, the router picks a few experts (top-k) and combines their outputs, weighted by router scores. So the model has a large total parameter count, but each token only uses a fraction of it, the active parameters. Training adds a load-balancing objective so the router does not send everything to a few experts. Despite the name, experts rarely map to human topics like "law" or "biology"; specialisation tends to be at the level of token patterns.

44. What are the serving trade-offs of MoE models?

Answer: Compute per token scales with active parameters, so an MoE model can be faster and cheaper per token than a dense model of the same total size. But all experts must be stored in memory, so memory requirements scale with total parameters; you need enough GPU memory, or multiple GPUs with expert parallelism, to hold them all. Routing also adds cross-GPU communication and load imbalance. For self-hosting, size hardware by total parameters and measure throughput by active parameters.

Fine-tuning: full, LoRA and QLoRA

45. What is the difference between full fine-tuning and parameter-efficient fine-tuning?

Answer: Full fine-tuning updates every weight. It is the most flexible but needs memory for weights, gradients and optimiser states (several times the model size), and produces a full new copy of the model per task. Parameter-efficient fine-tuning (PEFT) freezes the base model and trains a small number of new parameters, such as LoRA adapters. It needs far less memory and produces small artefacts. For most enterprise use cases, PEFT reaches quality close to full fine-tuning. Our fine-tuning LLMs guide covers the end-to-end process.

46. How does LoRA work?

Answer: Low-Rank Adaptation keeps a pretrained weight matrix W frozen and learns an update expressed as the product of two small matrices: W' = W + (alpha / r) x B A, where A and B have a small inner dimension r (the rank). The assumption is that the task-specific change is low-rank. B starts at zero, so training begins exactly at the base model. Key settings are the rank, the scaling factor alpha, and which layers get adapters (often the attention projections, and sometimes the MLP layers). After training, adapters can be merged into the weights for zero extra latency, or kept separate so one server can host many adapters on one base model.

47. What does QLoRA add?

Answer: QLoRA trains LoRA adapters on top of a base model quantized to 4 bits, which cuts memory enough to fine-tune fairly large models on a single GPU. The frozen base weights are stored in a 4-bit NormalFloat (NF4) format, with "double quantization" of the scaling constants to save more memory and paged optimisers to absorb memory spikes. Gradients flow through the dequantized weights into the adapters, which are kept in higher precision. Training steps are slower than plain LoRA, and merging adapters for serving needs care, so test the exact deployment format.

48. How do you prepare data for fine-tuning?

Answer: Start with the behaviour you want, then build examples that demonstrate it. Steps I follow:

  • Use real inputs, with target outputs reviewed by domain experts.
  • Format exactly like production: same chat template, same system prompt, same tool-call format.
  • Deduplicate, remove contradictions and fix low-quality labels; a small, clean set beats a large, noisy one.
  • Cover edge cases, including examples where the right answer is to refuse, ask a clarifying question or escalate.
  • Scrub or mask personal data, and confirm you have the right to use the data for training.
  • Split train, validation and a held-out test set before training, and keep the test set untouched.

49. What is catastrophic forgetting, and how do you mitigate it?

Answer: Catastrophic forgetting is when fine-tuning on a narrow task degrades abilities the model had before: general reasoning, instruction following, other languages or safety behaviour. Mitigations: use PEFT methods such as LoRA, which change less; use a lower learning rate and fewer epochs; mix in general instruction data or samples of the model's own outputs (replay); and always run a general regression suite, plus safety tests, alongside the task evaluation. If the model only needs to know new facts, retrieval avoids the problem entirely.

50. When is fine-tuning the right tool, and when is it not?

Answer: Fine-tuning is good at teaching behaviour: a consistent output format or style, domain-specific classification or extraction, tool-calling patterns, or getting a smaller, cheaper model to match a larger one on a narrow task (distillation). It is poor at teaching facts that change, because updating knowledge means retraining and the model may still hallucinate details. Try prompting and retrieval first, collect evaluation data, and fine-tune when you have a clear gap that examples can close.

51. Scenario: your fine-tuned model scores well on the training-style validation set but poorly on real traffic. What happened?

Answer: Most likely a distribution mismatch or overfitting: the training and validation data are cleaner, shorter or more uniform than production inputs, or the validation set leaked from the same source as training.

What I would check:

  1. Compare real inputs with training inputs: length, language mix, formatting, typos, attachments.
  2. Check the training curves: validation loss rising while training loss falls suggests overfitting.
  3. Confirm the production prompt and chat template match what the model was trained on.

Production consideration: Build the evaluation set from real, recent production samples, refreshed periodically, and gate every new adapter on it.

Evaluation and contamination

52. What is the difference between public benchmarks and task-specific evaluations?

Answer: Public benchmarks (general knowledge, maths, coding, reasoning suites) measure broad capability on standard datasets and are useful for shortlisting models. They do not tell you how a model will do on your documents, your users' phrasing, your output format or your risk tolerance. Task evaluations are built from your own data with your own success criteria: correct extraction, policy-compliant answers, valid JSON, correct tool calls, appropriate refusals. Decisions about which model to deploy, which prompt to ship or whether an upgrade is safe should be made on task evaluations. See our guide to LLM evaluation for how to build them.

53. What is benchmark contamination, and how can it be detected?

Answer: Contamination happens when benchmark questions or answers appear in a model's training data, so a high score reflects memorisation rather than ability. Web-scraped corpora make it hard to avoid. Detection approaches include checking n-gram overlap between training data and test items (only possible when you can see the data), looking for canary strings, testing whether the model can reproduce test items verbatim, and comparing scores on the original benchmark with freshly written or perturbed versions of the same problems. A sharp drop on rewritten versions is a warning sign. Keep your own test sets private.

54. How does LLM-as-a-judge work, and what are its weaknesses?

Answer: A strong model scores another model's output against a rubric, for example whether an answer is supported by the provided sources. Known weaknesses: position bias (favouring the first or second answer in pairwise comparisons), verbosity bias (favouring longer answers), self-preference (favouring outputs from its own model family), and inconsistency across runs. Mitigations: specific rubrics with examples, binary or narrow scales instead of vague 1–10 ratings, swapping answer order, using a different model family as judge, and regularly calibrating the judge against human labels on a sample.

55. What is perplexity, and why isn't it enough to evaluate a chat model?

Answer: Perplexity is the exponential of the average negative log-likelihood per token on a test text: roughly, how "surprised" the model is. Lower is better, and it is useful for comparing language-modelling quality of base models with the same tokenizer, or checking that quantization did not damage a model. It is not enough for chat models because it does not measure helpfulness, correctness, instruction following or safety, and preference tuning can make perplexity worse while making the assistant more useful. Values are also not comparable across tokenizers.

56. Scenario: a vendor's model tops a public leaderboard, and leadership wants to switch to it. How do you respond?

Answer: I would support evaluating it, but not switching on the leaderboard alone. Leaderboards may reflect contamination, prompt formats tuned for the benchmark, or tasks unrelated to ours, and they ignore our cost, latency, data-residency and security requirements.

What I would check:

  1. Run our task evaluation set on the new model and the current one, with prompts adapted fairly to each.
  2. Compare cost per successful task, p95 latency and output-format reliability.
  3. Check hosting options: region, data retention terms, availability in our cloud, enterprise agreements.
  4. Pilot behind a feature flag on a slice of traffic before full rollout.

Production consideration: A model abstraction layer plus a standing evaluation suite makes switching routine. Our guide on choosing an LLM for the enterprise lists the full criteria.

Hallucination, safety and alignment

57. Why do LLMs hallucinate?

Answer: Several causes combine. The training objective rewards plausible continuations, not truth, so the model produces fluent text whether or not it has reliable knowledge. Rare facts, such as a specific circular number or a niche drug interaction, are weakly represented in the weights. Training and evaluation have historically rewarded confident answers more than "I don't know", so models learn to guess. Errors can also snowball: once a wrong token is generated, later tokens are conditioned on it. In applications, poor retrieval and ambiguous questions add to it. LLM hallucinations explained goes through the causes in more depth.

58. How do you reduce hallucinations in a production system?

Answer: No single technique removes them; you layer controls.

  • Ground answers in retrieved, permission-checked sources and require citations.
  • Instruct and test for abstention: "If the sources do not contain the answer, say so."
  • Use tools for facts that can be looked up or calculated (balances, dates, maths).
  • Constrain outputs with schemas and validate them in code.
  • Route low-confidence or high-risk cases to a human.

Lowering temperature reduces random variation but does not fix a model that lacks the knowledge. Measure the hallucination rate on a labelled evaluation set so you know whether changes help.

59. What does "alignment" mean for LLMs, and how is it achieved?

Answer: Alignment means getting the model to behave as intended: helpful, honest about what it knows, and avoiding harmful outputs, even in situations the developers did not anticipate. Methods include SFT on good demonstrations, preference tuning (RLHF, DPO, principle-guided AI feedback), safety training data, red-teaming and system-level controls. Known problems include over-refusal, sycophancy (agreeing with a wrong user) and reward hacking.

60. What is the difference between a jailbreak and a prompt injection?

Answer: A jailbreak is a user deliberately crafting input to make the model ignore its safety training or instructions. A prompt injection is untrusted content, such as a web page, email, PDF or tool result, carrying instructions that the model follows as if they came from the developer or user. Injection is the bigger agent risk: the attacker never talks to the system, and impact grows with the agent's tools. Neither can be fully solved by prompting, which is why the system prompt should never be treated as a security boundary. Our AI guardrails guide covers layered defences.

61. Scenario: an email-summarising agent with "send email" permission is fed a message containing hidden instructions to forward the inbox. How do you design against this?

Answer: Assume the model can be manipulated and design so that a manipulated model cannot do serious damage.

What I would check:

  1. Least privilege: does a summariser need send permission at all? Split read and write capabilities into separate flows.
  2. Human approval for consequential actions such as sending, forwarding or deleting.
  3. Deterministic policy checks outside the model: allowed recipients, no external forwarding, rate limits.
  4. Clear separation and labelling of untrusted content in the prompt, plus injection classifiers as one layer.

Production consideration: Red-team this exact attack before launch and keep the cases in the regression suite.

Serving and cost

62. What drives the cost of an LLM application?

Answer: For hosted APIs: input tokens, output tokens (usually priced higher), reasoning tokens, cached-input pricing where available, and the number of calls per user task (agents can make many). Hidden multipliers: long system prompts, growing history, oversized retrieved context and retries. For self-hosting: GPU hours, utilisation, engineering and on-call time, and idle capacity outside peak hours. Track cost per successful task, not per token.

63. How does prompt caching work, and how do you design prompts to benefit from it?

Answer: Prompt caching stores the computed KV cache for a prompt prefix so later requests that start with the exact same prefix skip recomputing it. Providers typically bill cached input tokens at a reduced rate and TTFT drops. To benefit, put stable content first (system prompt, tool definitions, policy documents, few-shot examples) and variable content last (user message, retrieved chunks). Anything that changes early in the prompt, even a timestamp, breaks the cache for everything after it. Minimum lengths and cache lifetimes vary, so check current provider documentation.

64. How do you estimate GPU memory needed to self-host a model?

Answer: Add three parts. Weights: parameters times bytes per parameter (two bytes at 16-bit, one at 8-bit, about half a byte at 4-bit). KV cache: per-token cache size (Q25) times the total tokens of all concurrent sequences. Overhead: activations, CUDA context, framework buffers and fragmentation, so leave headroom. For example, a 7-billion-parameter model at 16-bit needs about 14 GB just for weights; whether a given GPU is enough then depends on the context length and concurrency you need. Beyond one GPU, tensor parallelism splits each layer across GPUs and pipeline parallelism splits layers into stages. See self-hosting LLMs for the operational side.

65. Scenario: should a hospital self-host an open-weight model or use a managed API?

Answer: It depends on data-handling requirements, scale, skills and quality needs, and the answer can be "both" for different workloads.

What I would check:

  1. Data requirements: can patient data be processed by a managed service in an approved region with acceptable retention terms, under the hospital's legal and DPDP Act obligations?
  2. Quality: does an open-weight model of a size we can host meet the bar on our evaluation set, compared with frontier hosted models?
  3. Volume and utilisation: steady high volume favours self-hosting economics; spiky or low volume favours pay-per-token.
  4. Team capacity to operate GPUs and model updates.

Production consideration: A common pattern is a managed model through a private endpoint in the cloud account for most tasks, plus a smaller self-hosted model for the most sensitive or high-volume narrow tasks.

If you want to learn these internals by building and measuring them, rather than memorising answers, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program combines model fundamentals with cloud deployment and security, in classroom sessions in Ameerpet or live online.

Real-world scenario questions

66. Scenario: an insurer wants an assistant that answers agents' questions about policy wordings and underwriting rules. Fine-tuning or RAG?

Answer: RAG first, and possibly a small amount of fine-tuning later for format or tone. Policy wordings and underwriting rules change, differ by product and version, and answers must cite the exact clause. Retrieval handles updates by re-indexing, supports document-level permissions and gives citations. Fine-tuning would bake in a snapshot, cannot cite sources and may blend clauses across products.

What I would check:

  1. How often documents change and whether versions must be selected by policy date.
  2. Who may see which documents (agent tier, region, internal-only underwriting guides).
  3. Whether the hard part is finding the right clause (a retrieval problem) or interpreting it in a house style (where fine-tuning might help).

Production consideration: Combine both only when evaluation shows a gap that retrieval and prompting cannot close, for example a fine-tuned small model for classifying query types at high volume. Our RAG interview questions go deeper into retrieval design.

67. Scenario: after upgrading to a newer model version, answer quality dropped for some users. How do you diagnose it?

Answer: Treat it as a regression: confirm it, localise it, then decide whether to roll back or adapt prompts. A newer model can be better on average yet differ in verbosity, refusals, tool-call format and sensitivity to prompts tuned for the old one.

What I would check:

  1. Confirm with data: run the standing evaluation set on both versions, and compare production feedback and escalation rates before and after.
  2. Segment the failures by route, language, task type, prompt version and input length to find where it regressed.
  3. Diff behaviour: output length, format violations, refusals, tool-call errors, changed defaults (temperature handling, reasoning settings, maximum output tokens).
  4. Check what else changed in the same release: prompts, retrieval index, SDK, chat template.
  5. Re-tune prompts for the new model against the evaluation set before shipping.

Production consideration: Pin model versions explicitly, upgrade behind a flag with canary traffic, keep the old version available for rollback, and log model version with every request so incidents can be traced.

68. Scenario: a banking chatbot has a slow time to first token and slow overall responses. How do you reduce latency?

Answer: Measure where time goes, then attack the largest component.

request -> auth -> retrieval -> rerank -> guardrail
        -> [queue] -> prefill -> decode (stream)
        -> output checks -> UI
measure each hop; fix the biggest one first

What I would check:

  1. Prefill: shrink the prompt (shorter system prompt, fewer and better chunks, summarised history) and enable prompt caching for the static prefix.
  2. Decode: shorter answers, a smaller or faster model for simple intents, streaming to the UI.
  3. Queueing from rate limits or under-provisioned capacity at peak hours; reserved or provisioned throughput if needed.
  4. Self-hosted: quantization, speculative decoding, batching settings.

Production consideration: Re-run quality evaluation after every latency change; a faster model that mis-answers card-blocking questions is not an improvement.

69. Scenario: the monthly LLM bill for an internal assistant at a Hyderabad GCC doubled with no increase in users. What do you investigate?

Answer: Cost rose per user, so look for more tokens per task or more calls per task.

What I would check:

  1. Token usage by route, prompt version and model, from gateway or provider logs.
  2. Recent changes: a longer system prompt, new tool definitions, more retrieved chunks, a switch to a reasoning model or higher reasoning effort.
  3. Agent loops: retries, repeated tool calls or runaway multi-step plans without step limits.
  4. Prompt-cache hit rate; a dynamic value at the start of the prompt can silently break caching.
  5. Non-user traffic such as evaluation jobs or load tests pointed at production.

Production consideration: Set per-route budgets and alerts, cap agent steps and output tokens, and report cost per successful task alongside quality.

70. Scenario: summaries of long vendor contracts miss important clauses buried in the middle. What do you change?

Answer: This is a classic long-context failure. Rather than one giant prompt, structure the work so each clause is examined with enough attention.

What I would check:

  1. Accuracy by clause position, to confirm a "lost in the middle" pattern.
  2. Switch to a section-wise approach: split by contract structure, extract key fields from each section against a schema (liability, termination, data protection, renewal), then synthesise.
  3. Add a checklist of mandatory clause types, so a missing one is reported as "not found" rather than silently skipped.

Production consideration: Show the clause references with each extracted item so reviewers can verify quickly; for legal work, the model assists a human reviewer rather than replacing one.

71. Scenario: after LoRA fine-tuning for ticket classification, the model now ignores instructions in other tasks and its answers are oddly short. What went wrong?

Answer: The adapter has over-specialised: training on short label outputs taught the model that short responses are always correct, and general instruction following has been degraded. This is a form of catastrophic forgetting.

What I would check:

  1. Whether the adapter is being applied to requests it was never meant for; routing may be wrong.
  2. Learning rate and epochs; aggressive values increase drift.
  3. A general instruction-following and safety regression suite, run on base and fine-tuned versions.

Production consideration: Keep task adapters separate and apply them only on their route (multi-adapter serving makes this easy), and mix general examples into training if one model must do several jobs.

72. Scenario: a claims pipeline expects JSON, but a small share of responses fail to parse and break downstream jobs. How do you fix it?

Answer: Stop relying on instructions alone and enforce structure at decode time and in code.

What I would check:

  1. Use the provider's structured output or tool-calling mode with a JSON schema; self-hosted engines support grammar-constrained decoding.
  2. Look at failed outputs: truncation at the maximum token limit is a common cause, as is extra prose around the JSON.
  3. Validate every response against the schema in code, retry once with the validation error, then route to a dead-letter queue for review.

Production consideration: Valid JSON is not correct JSON; also check field values (dates, amounts, enumerations) with business rules. Function calling and structured outputs covers the patterns.

73. Scenario: an auditor asks why the same prompt at temperature 0 produced two different answers. What do you tell them?

Answer: Temperature 0 reduces randomness but does not make hosted inference bit-for-bit deterministic. Floating-point operations on GPUs are not perfectly associative, results can vary with batch composition and hardware, MoE routing can amplify tiny numerical differences, and providers may update serving infrastructure. Some providers also adjust or ignore temperature for certain models.

What I would check:

  1. Whether the inputs were truly identical: retrieved context, timestamps, conversation history, model version.
  2. Whether a seed parameter is available and what the provider says about determinism.

Production consideration: For auditable decisions, log the full prompt, retrieved sources, model version, parameters and output for each request, and make the decision logic depend on structured outputs and rules rather than free text.

74. Scenario: you must serve an open-weight model on a limited GPU budget for a retailer's product-description generator. How do you plan capacity?

Answer: Start from the workload, not the model. Product descriptions are high-volume and batch-friendly, which suits a smaller quantized model.

What I would check:

  1. Workload shape: items per day, prompt and output lengths, deadline (overnight batch versus real time).
  2. Smallest model that passes the quality bar on a sample reviewed by merchandisers.
  3. Memory budget: weights at the chosen precision plus KV cache for the target batch size (Q64).
  4. Measured throughput on a serving engine with continuous batching, at realistic lengths, rather than estimates.

Production consideration: Add automated checks for banned claims and factual fields (price, size, material must come from the catalogue, not the model) and human review for a sample.

75. Scenario: an interviewer asks you to explain, end to end, what happens when a user sends a message to an LLM-powered assistant. How do you structure your answer?

Answer: Walk the path in order, naming one design decision per step.

  1. Application: authenticate the user, apply input guardrails, assemble context (system prompt, history, retrieved and permission-filtered documents, tool definitions).
  2. Tokenization: the chat template is applied and text becomes token IDs; token count decides cost and whether it fits the window.
  3. Prefill: one parallel pass builds the KV cache, reusing a cached prefix if possible.
  4. Decode: tokens are generated one at a time with the chosen sampling settings, possibly with speculative decoding, while batched with other users' requests.
  5. Tool calls: if the model emits a tool call, the application executes it with the user's permissions and returns the result for another model turn.
  6. Output: stream to the UI, validate structure, apply output guardrails, log traces, tokens, latency and model version.

Interview tip: Adjust depth to the role. For an ML role, spend longer on attention and decoding; for an engineering or Forward Deployed Engineer interview, spend longer on permissions, tools, observability and evaluation.

Key takeaways

  • Know the transformer mechanics well enough to explain attention, positional encoding and the decoder-only forward pass without notes.
  • Separate the training stages clearly: pretraining builds knowledge, SFT teaches the assistant format, preference tuning (RLHF, DPO, RL from verifiable rewards) shapes behaviour.
  • Most inference questions come back to prefill versus decode and the KV cache; master those and quantization, batching and speculative decoding follow.
  • Long context, reasoning models and MoE all have real trade-offs in cost, latency and memory; interviewers reward candidates who name them.
  • Fine-tune for behaviour, retrieve for knowledge, and evaluate both on your own task data, not public leaderboards.
  • Hallucination and prompt injection are managed with layered system design, not prompts alone.
  • In scenarios, show a method: measure, segment, form a hypothesis, change one thing, re-evaluate.

Interview preparation checklist

  • Write out the attention formula and explain each term, including the square-root scaling.
  • Tokenize the same paragraph in English, Hindi and Telugu with a real tokenizer and compare counts.
  • Calculate weight and KV-cache memory for one published model configuration.
  • Run a small open-weight model with a serving engine, then measure TTFT and tokens per second at different batch sizes and quantization levels.
  • Fine-tune a small model with LoRA on a narrow task, and run a general regression test before and after.
  • Build a 30–50 item task evaluation set with expected answers and a simple LLM-as-judge rubric.
  • Practise answering Q66–Q75 aloud in under four minutes each, using the "what I would check" structure.
  • Review neighbouring topics: MLOps and LLMOps interview questions and AI engineer interview questions.

FAQ

What skills are needed to clear an LLM interview?

You need a clear grasp of transformer basics, tokenization, training stages and inference, plus practical skills in Python, prompt and context design, evaluation, and deploying models through APIs or serving engines. For engineering roles, cloud, security and observability matter as much as model theory.

Do I need to know the maths behind transformers?

You should understand the attention formula, softmax, cross-entropy loss and why matrix sizes affect memory and speed. Applied roles rarely ask for derivations; research and ML engineer roles may go deeper.

How should a fresher prepare for LLM interview questions?

Learn the fundamentals in this guide, then build two or three small projects: a RAG assistant, a LoRA fine-tune on a narrow task and a simple evaluation harness. Be ready to defend every design choice.

Are LLM interviews different for ML engineers and AI application engineers?

Yes. ML engineer interviews focus more on training, fine-tuning, evaluation methodology and model internals. AI application and platform engineer interviews focus more on serving, latency, cost, RAG, tool use, security and production debugging.

Do I need a GPU to practise for LLM interviews?

Not necessarily. You can learn a lot with hosted APIs and small open-weight models on a laptop or CPU runtime. For fine-tuning and serving experiments, short sessions on a cloud GPU are usually enough.

Which topics are most commonly asked in LLM interviews in 2026?

Commonly asked topics include attention and the KV cache, sampling parameters, RAG versus fine-tuning, LoRA and QLoRA, hallucination causes, evaluation design, reasoning models, quantization and scenario questions about latency, cost and quality regressions.

Is an LLM engineering career a good choice for Indian engineers?

LLM skills are increasingly expected in AI, data, cloud and software roles across GCCs, product companies and services firms in Hyderabad, Bengaluru and other cities. Combining model understanding with production engineering keeps the skills useful across several job titles.

How long does it take to prepare for an LLM interview?

It depends on your starting point. An experienced developer can usually cover the fundamentals and one or two projects in a few focused weeks; a fresher should plan for longer, with most time spent building rather than reading.

Want to go from knowing how LLMs work to engineering them in production? Explore the Cloudsoft APEX program for AI, ML, cloud and security foundations, or, if you want to deliver AI systems inside customer environments, the AI Forward Deployed Engineer FDE PRO program, which includes enterprise projects and placement support until you're placed. Both run in classroom sessions in Ameerpet or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us