New batches starting this week Β· Limited seats

Fine-Tuning LLM Interview Questions and Answers 2026 (55 Questions)

55 fine-tuning LLM interview questions with practical answers, from the fine-tune vs RAG vs distil decision and clean data to LoRA internals, DPO and GRPO, regression evals, multi-LoRA serving and eleven production scenarios.

Fine-tuning LLM interview questions 2026: 55 questions on when to fine-tune, data, LoRA and QLoRA, SFT and DPO, forgetting and adapters
Last updated Β· 42 min read Β· 9,315 words

Fine-tuning LLM interview questions in 2026 test judgement more than recipes: whether you can tell when fine-tuning is the right tool at all, build a clean dataset that does not leak into your evaluation, pick between LoRA, QLoRA and full fine-tuning with a defensible memory estimate, and prove the tuned model is better without becoming less safe or less capable elsewhere. This guide collects 55 high-value questions with model answers, from the decision framework and data preparation through LoRA internals, preference tuning (RLHF, DPO, ORPO, KTO, GRPO), hyperparameters, evaluation, managed services and multi-LoRA serving, ending with eleven production scenarios.

How to use this guide

This page goes deeper than our fine-tuning LLMs guide, which explains the end-to-end process; read that first if the basic workflow is new to you. Interviewers for an LLM fine-tuning engineer role commonly probe three levels:

  • Freshers and junior engineers: what SFT, LoRA and QLoRA are, why data quality matters, train/validation/test splits and what a chat template does.
  • Mid-level engineers: LoRA rank, alpha and target modules, loss masking, packing, learning-rate behaviour, DPO data formats, regression suites and memory estimates.
  • Senior and architect roles: fine-tune versus prompt versus RAG versus distillation decisions, preference and reward-based RL trade-offs, managed versus self-hosted training, multi-adapter serving, licences and safety regressions, and structured debugging of failed runs.

Numbers in this guide (memory per parameter, example ranks) are illustrative rules of thumb for reasoning in an interview, not benchmarks. Library arguments and cloud features change, so check current documentation before relying on any specific setting.

When to fine-tune: fundamentals

1. What problems does fine-tuning actually solve, and what does it not solve?

Answer: Fine-tuning changes behaviour: output format, label scheme, tone, the way a model follows a domain-specific procedure, or how reliably it calls tools. It is weak at adding facts that must stay current, be cited or be restricted per user. Facts learnt in weights cannot be updated without retraining, cannot be traced to a source document, and cannot respect document-level access control. A clean way to say it in an interview: prompting and retrieval change what the model sees; fine-tuning changes how the model responds to what it sees. Fine-tuning can also reduce cost and latency by letting a smaller model do a narrow task that previously needed a larger model plus a long prompt.

Interview tip: Name the specific behaviour you want changed before you name any method. "We want the model to know our products" is a retrieval problem; "we want it to output our 40-field claim schema without a three-page prompt" is a fine-tuning candidate.

2. Walk through how you decide between prompting, RAG, fine-tuning and distillation.

Answer: I treat them as an escalation ladder, each step justified by evaluation on a fixed test set:

  1. Prompting first (instructions, few-shot examples, structured output features). It is cheap, reversible and sets the baseline.
  2. RAG when failures are about missing, changing or permissioned knowledge.
  3. Fine-tuning when the model has the knowledge in context but still fails at behaviour: format drift, poor label boundaries, wrong style, or prompts so long they hurt latency and cost.
  4. Distillation when a large model already does the task well and you need the same quality from a smaller, cheaper model at volume. That is fine-tuning a student on a teacher's outputs, covered in our model distillation explainer.

These combine: a RAG assistant whose generator is fine-tuned to cite retrieved passages in a fixed format is common. The deeper comparison is in RAG vs fine-tuning.

Fails on knowledge?   --yes--> RAG / tools
        | no
Fails on behaviour?   --yes--> better prompt
        |                       | still fails
        | no                    v
Too slow / costly?    --yes--> fine-tune or distil
        | no
Ship the prompted baseline

3. What is the difference between continued pre-training, supervised fine-tuning and preference tuning?

Answer: Continued pre-training runs the next-token objective on large amounts of unlabelled domain text (contracts, clinical notes, code) to shift the model's vocabulary and domain fluency. It needs far more data and compute, and it usually damages instruction-following, so it is followed by SFT. Supervised fine-tuning (SFT) trains on prompt-response pairs where the response is the exact target. Preference tuning trains on comparisons or scores (this answer is better than that one) and refines judgement, tone or safety after SFT has fixed the format. In enterprise work, SFT on an instruction-tuned model is by far the most common; continued pre-training is rare and justified only for unusual languages or very specialised text.

4. Should you fine-tune a base model or an instruction-tuned model?

Answer: For most enterprise tasks, start from the instruction-tuned (chat) variant. It already follows instructions, handles multi-turn formatting and carries the provider's safety tuning, so a small SFT dataset only has to teach your task. A base model gives more control but forces you to teach conversation format and refusal behaviour yourself, which needs more data and a safety evaluation you now own entirely. Base models make sense for continued pre-training, for non-chat tasks such as classification heads, or when the chat variant's style fights your target format.

5. How much data do you need, and why is "it depends" not a good enough answer?

Answer: The honest answer is a method rather than a number: run a learning curve. Train on increasing subsets (for example a quarter, half and all of your data) and plot the test metric. If quality is still rising steeply, more data will help; if it has flattened, more of the same data will not, and you need different data (harder cases, more classes) or a different approach. As a rough orientation, narrow format or classification tasks often work with hundreds to a few thousand clean examples, while teaching a new complex skill needs much more. What matters more than volume is coverage of real input variety and label consistency.

Interview tip: Mentioning learning curves signals you have run real experiments rather than copied a tutorial.

6. Can you fine-tune a closed, API-only model?

Answer: Yes, for some models, through the provider's managed fine-tuning service. You upload a dataset, the provider trains a private variant, and you call it like the base model. You do not get the weights, you choose only the methods and hyperparameters the provider exposes, and the tuned model is usually tied to that provider's hosting and pricing. Open-weight models give full control over method, data handling and serving, at the cost of running training and inference infrastructure yourself. The choice is often driven by data residency, cost at volume and whether you need to move the model later; our guide to open-weight LLMs for enterprise covers that trade-off.

7. What is the baseline in a fine-tuning project and why is it non-negotiable?

Answer: The baseline is the score of the untuned model with your strongest prompt (and retrieval, if used) on the same held-out test set and metric you will use after training. Without it you cannot claim fine-tuning helped; you can only claim the tuned model scores some number. A strong baseline often closes most of the gap, which saves the project. I also record the baseline's latency and cost per request, because a common justification for fine-tuning is matching the baseline's quality with a smaller model or shorter prompt.

8. What does "fine-tuning for tool calling" mean, and when is it worth it?

Answer: It means training on conversations where the assistant emits tool calls (function name plus JSON arguments) in the format the serving stack expects, then consumes tool results and responds. It is worth it when an open model chooses the wrong tool, invents arguments or breaks the call format on your specific tool set, and prompting with clear tool descriptions has not fixed it. The data must use the exact tool-call format of the model's chat template, include examples where the right choice is no tool call, and include error results so the model learns recovery. Pair it with constrained decoding or schema validation at serving time; see function calling and structured outputs.

Training data and chat templates

9. Compare instruction format and chat format for SFT data.

Answer: Instruction format stores each example as fields such as instruction, input and output, which the training code turns into a prompt string. Chat (conversational) format stores a list of messages with roles (system, user, assistant, and sometimes tool), which the tokenizer's chat template renders into the exact token sequence the model was trained with. For chat models, conversational format is safer: it reuses the model's own template, supports multi-turn examples, and matches how you will call the model in production. Whatever the format, store data as JSON Lines with one example per line, a version number and provenance fields.

{"messages": [
  {"role": "system", "content": "You classify..."},
  {"role": "user", "content": "Ticket: VPN drops..."},
  {"role": "assistant",
   "content": "{\"queue\": \"network\"}"}
]}

10. Why does the chat template matter so much, and what breaks when it is wrong?

Answer: The model learnt role boundaries and the end-of-turn signal from special tokens placed by its template. If training renders conversations differently from inference (missing system prompt, different role markers, no end-of-turn token after the assistant message), the model learns a format it will never see in production. Typical symptoms are outputs that never stop, answers that continue into a fake next user turn, or quality that looks fine in training-style evaluation but drops behind the real serving API. The fix is to render training data with the same tokenizer template the serving engine uses, include the production system prompt, and inspect a few decoded training sequences by eye before the first run.

11. What is loss masking, and should you train on the prompt tokens?

Answer: Loss masking sets the labels of non-target tokens (system prompt, user turns, retrieved context) so they do not contribute to the loss, meaning the model is trained only on producing assistant tokens. For most SFT this is preferable: you want the model to learn responses, not to memorise your system prompt or imitate users. Training on the full sequence can help slightly when data is scarce and inputs are themselves in-domain text, but it also lets long contexts dominate the gradient. Most trainers support completion-only or assistant-only loss; verify it by checking that masked positions really have the ignore label in a sample batch, because a template change can silently break the masking.

12. "Quality over quantity" is a clichΓ©. What does it mean concretely for fine-tuning data?

Answer: Concretely, four checks per dataset:

  • Agreement: sample examples and have two reviewers label them independently. Where they disagree, the labelling guide is ambiguous and the model will learn the ambiguity.
  • Target correctness: every target must be something you would accept in production, because SFT copies mistakes, verbosity and bad habits faithfully.
  • Coverage: real input variety (typos, code-mixed Hindi-English or Telugu-English, forwarded email chains, empty fields), rare classes and cases where the right answer is "cannot determine".
  • Style consistency: if half the targets are terse and half are chatty, outputs will be inconsistent.

Removing a few hundred inconsistent examples frequently improves a model more than adding thousands of new ones.

13. How do you deduplicate fine-tuning data, and why exact matching is not enough?

Answer: Exact deduplication (hash of normalised text) removes copy-paste duplicates, but enterprise data is full of near-duplicates: the same ticket template with a different employee name, the same email forwarded with a new signature. I add near-duplicate detection with MinHash or n-gram overlap for surface similarity, and embedding similarity for paraphrases, then keep one representative per cluster or down-weight large clusters. Duplicates matter for two reasons: they over-weight a few patterns, pushing the model towards memorisation, and if near-copies straddle the train/test split they inflate evaluation scores. Deduplicate before splitting, never after.

14. What is evaluation contamination in fine-tuning, and how do you prevent it?

Answer: Contamination means test items, or close variants of them, appear in training data, so scores measure memory instead of ability. In fine-tuning it arises from near-duplicate leakage, from splitting related items (one customer's ticket thread) across train and test, from synthetic data generated by a model that was shown test items, and from repeatedly fixing the training set in response to test failures. Prevention: deduplicate then split by group (customer, document, thread) or by time; freeze the test set and store it separately with restricted access; run an n-gram or embedding overlap check between train and test before every run; and keep a small "never seen" test set refreshed from recent production traffic. Public benchmark contamination is covered in our LLM interview questions; this answer is about your own private evaluation set.

15. How do you use synthetic data for fine-tuning without degrading the model?

Answer: Synthetic data (examples generated by a stronger model) is useful for coverage: rare classes, edge cases, multilingual variants and tool-call trajectories. The risks are low diversity, confidently wrong targets, the generator's stylistic tics, and licence or terms restrictions on using a provider's outputs for training. My controls are: seed generation with real anonymised examples and explicit diversity axes; filter automatically (schema validity, deduplication, a judge model) and then human-review a sample of every batch, rejecting the batch if the error rate is too high; label synthetic rows so their contribution can be measured; and never generate or evaluate on the test set with the same pipeline. Keep real data as the backbone and synthetic data as the supplement. Our guide to synthetic data for AI testing covers generation and quality control in depth.

16. How do you handle personal and sensitive data in a fine-tuning dataset?

Answer: Assume anything in training data can be reproduced by the model, and that an adapter cannot enforce per-user permissions. So: minimise (keep only fields the task needs), mask or pseudonymise identifiers such as names, phone numbers, account numbers, Aadhaar and PAN, and record the masking in a data card with source, consent basis, retention and owner. For Indian organisations, processing personal data for training must fit the purposes and notices covered under the DPDP Act, which our DPDP Act guide for AI applications explains. After training, run extraction tests: prompt the model with prefixes of training records and check it does not complete identifiers.

Methods: full, LoRA, QLoRA, adapters and merging

17. When is full fine-tuning justified over parameter-efficient methods?

Answer: Full fine-tuning updates every weight. It is justified when the behaviour change is large (a new language, continued pre-training, a big shift in domain), when you have abundant high-quality data, and when you have the compute and the need to own a separate model copy. It also has the highest risk of catastrophic forgetting because every weight can drift. For narrow enterprise tasks, LoRA usually comes close to full fine-tuning quality at a fraction of the memory, produces small artefacts and lets one base model serve many tasks. I would only propose full fine-tuning after a LoRA run at a generous rank has demonstrably plateaued below the target on the test set.

18. Explain the LoRA update mathematically and why it saves memory.

Answer: For a frozen weight matrix W of shape d Γ— k, LoRA learns a low-rank update Ξ”W = BΒ·A, where A is r Γ— k and B is d Γ— r, with rank r much smaller than d and k. The forward pass becomes h = WΒ·x + (Ξ± / r)Β·BΒ·AΒ·x. A is initialised randomly and B to zero, so training starts exactly at the base model. Only A and B are trained, so gradients and optimiser states exist only for a small number of parameters; the base weights need memory only for the forward and backward pass, not for optimiser state. That removes the largest memory cost of full fine-tuning. The adapter file is also small, so many task adapters can share one base model.

Interview tip: Mentioning that B starts at zero, so the untrained adapter is a no-op, is a common follow-up check.

19. What do rank and alpha control, and how do they interact?

Answer: Rank (r) is the adapter's capacity: how many independent directions the update can express. Alpha (α) scales the update by α / r. Because the scale depends on r, changing rank without changing alpha also changes the effective step size, which confuses comparisons. Common practice is to keep α proportional to r (for example α equal to r or twice r) so that rank changes mostly affect capacity. A variant called rank-stabilised LoRA scales by α / √r instead, which keeps learning behaviour more stable at higher ranks. Higher rank is not automatically better: on small datasets it adds capacity to memorise. Typical starting ranks are small (8 to 32), raised only if the learning curve shows under-fitting.

20. Which modules should LoRA target, and how does that choice affect results?

Answer: Early LoRA work targeted only the attention query and value projections. Current practice usually targets all linear layers in each transformer block: the attention projections (query, key, value, output) and the MLP projections (in decoder models, often called gate, up and down). Targeting the MLP layers helps for tasks that need new domain behaviour, because much of a model's stored knowledge and transformation capacity sits there. Targeting more modules increases trainable parameters and memory moderately, but is often a better use of the budget than a much higher rank on fewer modules. The embedding and output head are usually left frozen unless you add new tokens, in which case they must be trained or saved too.

21. What does QLoRA add beyond "LoRA on a quantised model"?

Answer: QLoRA stores frozen base weights in a 4-bit format designed for normally distributed weights (NF4), dequantises them on the fly to a higher-precision compute type for each matrix multiply, and trains LoRA adapters in that higher precision. It adds double quantisation (quantising the quantisation constants themselves to save more memory) and paged optimisers that move optimiser state to CPU memory during memory spikes. The trade-offs: steps are slower because of dequantisation, and the adapter is trained against a quantised base, so its behaviour can differ slightly when applied to the full-precision base. That last point matters when you merge (see Q23).

22. Describe DoRA and other LoRA variants at a level an interviewer expects.

Answer: You should know the ideas, not every paper:

  • DoRA (weight-decomposed low-rank adaptation) splits each weight into a magnitude vector and a direction matrix, applies LoRA to the direction and trains the magnitude separately. It often tracks full fine-tuning behaviour more closely than plain LoRA at the same rank, with some extra compute.
  • Rank-stabilised LoRA changes the scaling factor so higher ranks train stably.
  • LoRA+ uses a larger learning rate for B than for A.
  • Initialisation variants (for example starting the adapter from the principal singular vectors of W) aim to converge faster.
  • Other PEFT families: IAΒ³ learns scaling vectors on activations; prompt tuning and prefix tuning learn soft tokens prepended to the input. They are lighter but usually less capable than LoRA for instruction tasks.

The practical answer: start with LoRA on all linear layers, try DoRA if quality plateaus, and adopt a variant only if it wins on your evaluation and your serving engine supports it.

23. Should you merge an adapter into the base weights? What changes for QLoRA adapters?

Answer: Merging computes W + (Ξ± / r)Β·BΒ·A once and saves a standalone model. It removes adapter overhead at inference and simplifies deployment to engines without adapter support. Keeping adapters separate lets one base serve many tasks, makes rollback trivial and keeps artefacts small. For a QLoRA adapter, do not merge into the 4-bit weights: load the base in higher precision, merge, then re-quantise for serving if needed. Because the adapter was trained against quantised weights, evaluate the merged model again rather than assuming identical behaviour. The library mechanics are covered in our Hugging Face interview questions; the interview point here is that merging is a deployment decision you re-validate, not a free step.

24. What is model merging across separately fine-tuned models, and when is it useful?

Answer: Model merging combines the weights (or weight deltas, sometimes called task vectors) of models fine-tuned from the same base, to get one model with several skills without joint training. Methods range from simple weighted averaging to approaches that resolve sign conflicts and drop small, noisy deltas before combining (TIES and DARE are commonly cited names). It is useful for combining, say, a format adapter and a domain-style adapter, or for recovering general capability by blending a tuned model back towards its base. It only works reliably between models sharing the same base and architecture, and results are unpredictable, so every merge is a new candidate that runs the full evaluation and safety suite.

25. How do you add new tokens or a new language without breaking the model?

Answer: Adding tokens resizes the embedding and output matrices; the new rows start untrained, so the model produces garbage around them until trained. With LoRA you must make the embedding and output head trainable or saved, otherwise the new tokens never learn. For most tasks, avoid new tokens: special markers can be expressed with existing tokens. For a genuinely under-served language, the cost of poor tokenisation (many tokens per word) can justify vocabulary extension plus continued pre-training on that language, followed by SFT, which is a much larger project than a LoRA adapter and needs a strong regression suite in the original languages.

If you want guided, hands-on practice building and evaluating models like these on cloud infrastructure, the APEX AI, ML, Cloud and Cyber Security program covers the machine learning, cloud and security foundations behind this work.

Preference tuning and reward-based RL

26. Explain RLHF with PPO at a conceptual level, including why it is expensive.

Answer: RLHF has three stages: SFT to get a capable starting policy; training a reward model on human comparisons of responses; and optimising the policy with reinforcement learning (classically PPO) to maximise reward while a KL penalty keeps it close to a frozen reference copy of the SFT model. It is expensive and fragile because PPO keeps several models in play (policy, reference, reward model and a value model), generates responses online during training, and is sensitive to hyperparameters. The reward model can also be exploited: the policy finds outputs that score highly without being better, called reward hacking. That complexity is why most enterprise teams use simpler offline methods unless they are building a foundation model.

27. How does DPO work, and what does the beta parameter do?

Answer: DPO trains directly on triples of prompt, chosen response and rejected response, with no separate reward model and no online generation. Its loss increases the policy's log-probability of the chosen response relative to the rejected one, measured against a frozen reference model (usually the SFT model). Beta controls how far the policy may move from the reference: a low beta allows larger departures (stronger preference fitting, more risk of drift and verbosity), a high beta keeps the model conservative. Practical points interviewers probe: run DPO after SFT, not instead of it; chosen and rejected responses should differ in the quality you care about, not just in length; and the reference model costs memory unless you precompute its log-probabilities or use an adapter on a shared base.

Real-world example: Consider an insurer whose claims-summary model already outputs the right structure after SFT, but reviewers prefer summaries that state uncertainty explicitly. Reviewers pick the better of two outputs for a few hundred claims, and a DPO pass shifts tone without retraining the format.

28. How do ORPO and KTO differ from DPO, and when would you choose each?

Answer:

MethodData neededKey ideaFits when
DPOPaired chosen / rejectedRelative likelihood vs frozen referenceYou have good pairs and an SFT model
ORPOPaired chosen / rejectedAdds an odds-ratio penalty to the SFT loss; no reference modelYou want one combined stage and less memory
KTOUnpaired good / bad labelsLoss inspired by prospect theory, per exampleFeedback is thumbs up / down, not pairs
Reward-based RLPrompts plus a reward functionOnline sampling, optimise scored outcomesCorrectness can be checked automatically

KTO is attractive in enterprises because production feedback usually arrives as single ratings, not side-by-side comparisons. Other reference-free variants exist; the interview point is matching the method to the shape of the feedback you can realistically collect.

29. What is reinforcement learning with verifiable rewards, and how does GRPO fit in?

Answer: With verifiable rewards, the reward comes from a program rather than a learnt preference model: unit tests pass, a SQL query returns the expected rows, a JSON output validates and matches the gold fields, a maths answer equals the reference. Because the reward is checked, it is harder to fool than a reward model, which is why this approach is central to training reasoning models. GRPO (group relative policy optimisation) samples a group of responses for each prompt, scores them, and uses each response's reward relative to the group's average as its advantage, so it needs no separate value model. Managed services now expose similar reward-function-based fine-tuning for some models. The hard parts are writing rewards that cannot be gamed (a reward checking only JSON validity teaches valid but empty JSON), choosing prompts of the right difficulty, and watching for drift in behaviours the reward does not measure.

30. Where does preference data come from in an enterprise, and what biases does it carry?

Answer: Sources include reviewer comparisons in an annotation tool, edits that agents or analysts make to model drafts (the edited version is chosen, the draft rejected), thumbs ratings in the product, and judge-model comparisons checked by humans. Each carries bias: human raters and judge models both tend to prefer longer, more confident answers; product feedback over-represents unhappy users; edits capture what one team cares about. Mitigations are a written rating rubric, length-controlled comparisons, mixing rater groups, auditing judge agreement against humans, and tracking response length and refusal rate before and after preference tuning so you notice if you trained verbosity rather than quality.

Hyperparameters and training mechanics

31. How do you choose the learning rate for LoRA versus full fine-tuning?

Answer: LoRA adapters typically use a noticeably higher learning rate than full fine-tuning because only a small number of freshly initialised parameters are learning, while full fine-tuning nudges already-good weights and a high rate destroys them. Whatever the starting point (library defaults are reasonable), use a short warm-up and a decaying schedule (cosine or linear), and do a small sweep of three values spaced several-fold apart on a subset, picking by validation metric rather than training loss. Signs the rate is too high: loss spikes, gradient-norm spikes, sudden degradation on general-capability checks. Too low: validation metrics barely move after an epoch.

32. How many epochs, and how do you know when to stop?

Answer: For SFT on a small to medium dataset, one to three epochs is the usual range; large datasets often need only one. The deciding signal is the validation set, measured with the task metric, not only validation loss: evaluate each epoch (or every few hundred steps), keep checkpoints, and select the one where the task metric peaks. Validation loss can keep falling while outputs become more templated, or rise slightly while task accuracy still improves, because loss measures token likelihood, not correctness. Early stopping on the task metric plus keeping the top-scoring checkpoint is the safe pattern.

33. Explain batch size, gradient accumulation and why the effective batch matters.

Answer: Per-device batch size is limited by GPU memory. Gradient accumulation sums gradients over several micro-batches before one optimiser step, so the effective batch = per-device batch Γ— accumulation steps Γ— number of GPUs. The effective batch, not the micro-batch, interacts with the learning rate: larger effective batches give smoother gradients and usually tolerate a higher rate; very small effective batches are noisy. When you move a run from one GPU to four, recompute the effective batch, or results will change for reasons unrelated to your data. Also note that on small datasets a large effective batch means very few optimiser steps per epoch, which can look like under-training.

34. What is sequence packing, and what can go wrong with it?

Answer: Packing concatenates several short examples into one training sequence up to the maximum length, cutting the compute wasted on padding. Naive packing lets tokens of one example attend to the previous example in the same sequence, which leaks context between unrelated samples. Modern trainers avoid this with attention boundaries per example (padding-free or variable-length attention using position resets), which you should confirm is enabled. Other pitfalls: packing changes the number of examples per step, so learning-rate schedules and "epochs" mean something different; examples longer than the maximum get truncated and can lose their answer; and loss masking must still apply per example inside the pack.

35. How do you choose the maximum sequence length for training?

Answer: From the data, not from the model's advertised context window. Plot the token-length distribution of fully rendered examples (template, system prompt, context and answer) and set the maximum to cover nearly all of them. Truncation is dangerous because the cut usually removes the assistant answer at the end, so the model trains on prompts with no target. Activation memory grows with sequence length, so an unnecessarily long maximum wastes memory. If production inputs are longer than your training examples, include long examples deliberately, since behaviour at lengths never seen in training cannot be assumed from the base model's context support.

Evaluation, forgetting and safety

36. How do you evaluate a fine-tuned model end to end?

Answer: Four layers, all compared against the baseline on frozen sets:

  1. Task metrics: exact match, per-class F1, field-level accuracy for extraction, schema validity rate, pass rate on programmatic checks.
  2. LLM-as-a-judge for open-ended quality, with a rubric, calibrated against human labels on a sample and checked for position and length bias.
  3. Human review by domain experts on a stratified sample, especially for high-risk outputs.
  4. Online A/B or shadow testing on real traffic, measuring business outcomes (resolution rate, edit rate, escalations) with guardrail metrics.

Add the regression and safety suites in Q37 and Q40. Judge and metric design is the subject of our LLM evaluation interview questions.

37. What is a regression evaluation suite for fine-tuning, and what goes in it?

Answer: It is a fixed set of capabilities the tuned model must not lose, run on every candidate checkpoint next to the task test set. For an enterprise assistant I include: general instruction-following prompts in the formats users actually send; adjacent tasks the same model serves (if it also summarises, test summaries); multilingual prompts if users write in several languages; tool-calling and JSON tasks; refusal and safety prompts; and a few reasoning or arithmetic checks. Each item has an automatic check or judge rubric, with thresholds agreed before training: for example, the task metric must improve and no regression category may fall by more than an agreed tolerance relative to the base model.

38. What causes catastrophic forgetting, and what mitigations go beyond "use LoRA"?

Answer: Forgetting happens when gradients from a narrow distribution move weights that general skills depend on. Its main drivers are a high learning rate, too many epochs, a dataset with one rigid output style, and large-capacity updates. Mitigations: use a lower rate and fewer epochs; mix in a share of general instruction data (replay) alongside the task data; keep the system prompt and multi-turn structure realistic so the model does not learn "always output a label"; prefer LoRA with moderate rank; and when a single model must serve many tasks, consider separate adapters routed per task rather than one adapter for everything. LoRA reduces forgetting but does not prevent it, which is why the regression suite exists.

39. How would you run an A/B test of a fine-tuned model against the current production model?

Answer: Start with shadow mode: send copies of live requests to the candidate, store outputs, compare offline with judges and spot checks, without affecting users. Then a controlled rollout: route a small, randomly assigned share of users (not requests, to keep experiences consistent) to the candidate behind a feature flag, with predefined primary metrics (task success, edit rate), guardrail metrics (complaints, escalations, latency, refusal rate) and a stopping rule. Run long enough to cover weekly patterns, segment results by language and request type, and keep instant rollback to the previous adapter. Decide the success criteria before the test, not after seeing the numbers.

40. Why can fine-tuning weaken a model's safety behaviour, even on harmless data?

Answer: Safety behaviour in an instruction-tuned model is itself a learnt behaviour, and it can be thin. Fine-tuning on data with no refusals, especially data that rewards always complying and answering directly, can erode refusal behaviour as a side effect, even when no example is harmful. Research has also shown that small amounts of deliberately harmful data can remove safety tuning, which is why providers screen managed fine-tuning datasets. Mitigations: include examples where the correct response is to decline or escalate; keep the provider's safety system prompt during training and inference; run a safety evaluation and red-team set before and after tuning (see AI red teaming); and keep runtime guardrails independent of the model, so a regression in weights is not your only line of defence.

Compute, managed services, serving and licences

41. Estimate GPU memory for full fine-tuning, LoRA and QLoRA of a 7-billion-parameter model.

Answer: An illustrative back-of-envelope estimate, excluding activations, which depend on batch size and sequence length:

SetupRough bytes per parameter7B model, before activations
Full fine-tuning, mixed precision with AdamAbout 16 (weights, gradients, FP32 master copy, two optimiser moments)Over 100 GB, so multiple GPUs or sharding
LoRA on a 16-bit baseAbout 2 for the frozen base, plus a small adapter with its optimiser stateRoughly 14–16 GB
QLoRA on a 4-bit baseAbout 0.5 plus quantisation overhead, plus the adapterRoughly 4–6 GB

Then add activations, which can dominate at long sequence lengths; gradient checkpointing trades extra compute for much lower activation memory. Sharded approaches (such as ZeRO-style optimiser and parameter sharding) spread full fine-tuning across GPUs. Treat these as planning estimates and measure peak memory on a short run. Our GPU basics for AI engineers explains the hardware side.

42. What do managed fine-tuning services on AWS, Azure and Google Cloud offer, in general terms?

Answer: All three let you customise selected models without managing GPUs; supported models, methods and regions change, so check current documentation.

  • Amazon Bedrock model customisation documents supervised fine-tuning, reinforcement fine-tuning with reward functions (which can be defined in AWS Lambda) and distillation from a teacher to a student model. Training is charged on tokens processed across epochs plus monthly model storage. Bedrock also supports importing open-weight models you customised elsewhere. Hosting options for custom models vary by model.
  • Microsoft Foundry (formerly Azure AI Foundry), including Azure OpenAI models, documents supervised fine-tuning, DPO and reinforcement fine-tuning on select models, with training tiers that differ in data residency and cost (a regional standard tier, a global tier and a lower-cost developer tier without residency commitments).
  • Gemini Enterprise Agent Platform (previously Vertex AI) documents supervised fine-tuning, preference tuning and reinforcement-learning fine-tuning for Gemini models, and supervised and distillation tuning for open models.

The interview angle: managed services trade control for speed. Ask where training data and weights are processed (critical for regulated Indian clients), what hosting costs once the model is deployed, and whether you can export or move the result.

43. How does multi-LoRA serving work, and what are its limits?

Answer: One base model sits in GPU memory, and many small adapters are loaded alongside it. Each request names an adapter; the engine batches requests for different adapters together and applies each request's low-rank update with specialised kernels, so dozens of tasks or tenants share one GPU deployment. In vLLM, for example, you enable LoRA support, register adapters, cap concurrent adapters with a max-loras setting and set the maximum rank to the largest adapter you will load; clients choose an adapter through the model name in the request, and adapters can be loaded at runtime. Limits: all adapters must share the same base model and version; the configured maximum rank reserves memory for every slot; too many active adapters cause swapping and latency spikes; and unmerged adapters add some per-token overhead. Serving design is covered in our LLM inference and serving interview questions.

Request(model="claims-v3") --+
Request(model="hr-v1") ------+--> Router / engine
Request(model="base") -------+        |
                                      v
                 [ Base weights on GPU, shared ]
                 [ LoRA slots: claims-v3 | hr-v1 ]
                 [ Adapter cache in CPU / disk  ]

44. What licence and terms questions must you answer before fine-tuning and shipping?

Answer: Five questions, answered in writing with legal review:

  1. Base model licence: is commercial use allowed? Some open-weight licences are permissive (Apache 2.0, MIT); others are custom community licences with acceptable-use policies, attribution or naming requirements, or conditions for very large deployments; some are research or non-commercial only.
  2. Derivatives: does the licence apply to your adapter and merged model, and what must you pass on to downstream users?
  3. Training data rights: do you have the right to use the tickets, documents or customer data for training, and under which consent or contract?
  4. Synthetic and distilled data: do the terms of the model that generated your data allow using its outputs to train another model?
  5. Managed service terms: who owns the tuned model, is your data used to improve the provider's models, and where is it processed?

Record the answers in the model card for each adapter, alongside the base model's exact revision.

Real-world scenario questions

45. After fine-tuning for ticket summaries, the model forgot general skills: it summarises everything, even when asked to translate or draft an email. What do you do?

Answer: This is catastrophic forgetting combined with an over-narrow dataset: every training example taught "input in, summary out", so the adapter learnt to ignore the instruction.

What I would check:

  1. Run the regression suite on the base model and each checkpoint to see when the loss of skills started.
  2. Inspect the data: is the instruction identical in every example, and is the system prompt missing?
  3. Review learning rate, epochs and rank against the dataset size.
  4. Check whether one adapter is being used for traffic it was never meant to serve.

Fixes in order: retrain with varied instructions and a share of general instruction data mixed in, a lower learning rate and fewer epochs; select the checkpoint by the task metric and regression thresholds; or route only summary requests to the adapter and send everything else to the base model.

Production consideration: Make the regression suite a release gate in the training pipeline so this cannot ship again unnoticed.

46. A team has 300 labelled examples and its model scores almost perfectly on training data but poorly on new data. How do you handle overfitting on a small dataset?

Answer: With 300 examples the model can memorise them quickly, especially at high rank or with several epochs. The first question is whether fine-tuning is even needed: a strong few-shot prompt with retrieved similar examples may match it.

What I would check:

  1. Whether validation is group-split and deduplicated, so the gap is real and not leakage in reverse.
  2. Validation task metric per epoch: it probably peaked in the first epoch.
  3. Whether the 300 examples cover the real input variety or come from one week, one team or one template.
  4. Label consistency across annotators.

Remedies: lower rank and learning rate, one or two epochs, LoRA dropout, early stopping on the task metric; expand coverage with reviewed synthetic variants of rare cases; and use k-fold style evaluation because one small test split gives noisy numbers.

Production consideration: Report confidence intervals or fold variance with small test sets; a few examples changing outcome can swing the score substantially.

47. After SFT on extraction data, the model returns inconsistent JSON: missing fields, extra prose, occasionally invalid syntax. What went wrong and how do you fix it?

Answer: The model reproduces inconsistency in its targets, or the serving setup differs from training.

What I would check:

  1. Validate every training target against the schema: often a share of targets have prose prefixes, different key orders, null versus missing fields, or markdown code fences.
  2. Check truncation: long inputs may have cut off the end of the JSON in training.
  3. Confirm the chat template and end-of-turn token match between training and serving.
  4. Check sampling settings at inference; high temperature increases format drift.

Fixes: normalise all targets to one canonical schema and style, add examples with empty or unknown values represented consistently, retrain; at serving time, use constrained decoding or the platform's structured-output feature, validate against the schema, and retry or fall back on failure. Fine-tuning improves the base rate; validation makes it safe.

Production consideration: Track the schema-validity rate per adapter version as a monitored metric, not a one-time test.

48. You are asked to choose between LoRA rank 8, 32 and 128 for a new task. How do you decide?

Answer: Treat rank as a capacity hyperparameter chosen by experiment against cost, not by intuition.

What I would check:

  1. Train all three on the same data with alpha scaled consistently (or rank-stabilised scaling) and the same learning rate schedule.
  2. Compare the task metric on validation, the regression suite, and the training-versus-validation gap.
  3. Look at dataset size and task complexity: narrow format tasks usually saturate at low rank; new domain behaviour may need more.
  4. Check serving constraints: the multi-LoRA engine's maximum rank setting and memory per slot.

Pick the smallest rank within an agreed tolerance of the top score. If rank 128 wins only by over-fitting the training set while regression worsens, rank 8 or 32 is the right answer; if all three plateau at the same score, the bottleneck is data, not rank.

Production consideration: Standardise a small set of ranks across teams sharing a base model, so the serving engine's rank limit does not force re-training later.

49. A fine-tuned classifier scores far above the baseline on the test set, but the gain disappears in a pilot. What do you suspect?

Answer: Evaluation contamination or a test set that does not represent production.

What I would check:

  1. Near-duplicate overlap between train and test (MinHash or embedding similarity), and whether related tickets from the same thread or customer straddle the split.
  2. Whether synthetic training data was generated with test items in the prompt.
  3. Whether the team iterated on training data while looking at test failures.
  4. Distribution differences: test set from last year, pilot traffic from new products or a new language mix.

Rebuild the split with deduplication before a group or time-based split, freeze a new test set sampled from recent traffic, and re-run both baseline and tuned model.

Production consideration: Keep a rolling "fresh" evaluation set from production that the training team never sees.

50. After DPO, reviewers say answers became longer and more flattering but not more accurate. What happened?

Answer: The preference data probably rewarded length and confident tone, and DPO learnt exactly that. It is a known bias of both human raters and judge models.

What I would check:

  1. Average length of chosen versus rejected responses in the dataset.
  2. Whether the judge model that produced labels was length-biased.
  3. Beta: a low value allows large drift from the SFT model.
  4. Accuracy metrics, not just preference win rate, before and after.

Fixes: rebuild pairs where chosen and rejected differ in correctness but are similar in length, use a rubric that penalises unsupported claims, raise beta, or consider length-normalised methods; where correctness can be checked, move that part of the objective to verifiable rewards.

Production consideration: Track response length and refusal rate as guardrail metrics on every preference-tuning run.

51. A fine-tuned assistant for a hospital group now answers questions it previously declined, such as dosing advice outside its scope. How do you respond?

Answer: Treat it as a safety regression and a release blocker. Consider a hospital group in India that tuned a model on helpful patient-communication replies; none of the training data contained declines, so the model learnt to always answer.

What I would check:

  1. Run the safety and scope-boundary suite on the base and tuned models to quantify the change.
  2. Look for missing refusal and escalation examples, and whether the safety system prompt was dropped in training.
  3. Check whether runtime guardrails (topic filters, clinical-scope classifiers) were bypassed by the new deployment.
  4. Review recent production logs for out-of-scope answers that may need clinical follow-up.

Roll back to the previous adapter, retrain with explicit decline-and-escalate examples and the production system prompt, and add the failing prompts to the permanent safety suite. Our AI guardrails guide covers the runtime layer.

Production consideration: Require a documented safety sign-off per adapter version, owned by someone outside the training team.

52. A services firm in Hyderabad needs separate fine-tuned behaviour for many client accounts on one GPU budget. Design it.

Answer: One shared base model with one LoRA adapter per client, served by a multi-LoRA engine, with tenant isolation enforced around it.

What I would check:

  1. That every client's data and licence terms allow this, and that no client's data trains another client's adapter.
  2. Traffic shape per client, to size how many adapters must be resident and which can be loaded on demand.
  3. A standard rank and target-module recipe, so all adapters fit the engine's limits.
  4. Gateway routing that maps an authenticated tenant to its adapter name, so a client can never select another client's adapter.

Each adapter gets its own evaluation suite, version and model card; rollouts happen per tenant. When the base model is upgraded, all adapters are retrained and re-evaluated as a batch.

Production consideration: Log adapter name and version with every request so incidents and audits can be traced to the exact model behaviour a client received; our multi-tenant AI SaaS guide covers the isolation patterns.

53. A newer version of your base model is released. Can you reuse your existing adapters?

Answer: No. A LoRA adapter is a delta for one exact set of base weights; applying it to a different version (even with the same architecture) produces unpredictable behaviour.

What I would check:

  1. First, the new base with the current prompt on the frozen test set: it may already match the old tuned model, removing the need for an adapter.
  2. Whether the chat template or tokenizer changed, requiring data re-rendering.
  3. Licence changes in the new version.
  4. Serving engine support for the new architecture.

Retrain from the versioned dataset with the same recipe, compare on task, regression and safety suites, and roll out with shadow traffic. Versioned datasets, configs and evaluation sets make this a repeatable pipeline run.

Production consideration: Pin base model revisions in every adapter's metadata and refuse to load mismatched pairs.

54. An insurer's leadership wants to "fine-tune the model on all our policy documents so it knows our products". How do you respond?

Answer: I would reframe the request: policy wordings change, answers must cite the exact clause, and different users may see different products, so the knowledge belongs in retrieval, not in weights.

What I would check:

  1. What failures they actually see today: wrong facts (retrieval problem) or wrong format and tone (possible fine-tuning problem).
  2. How often documents change and whether citations are required by compliance.
  3. Access rules per document and per user role.

Propose RAG over the policy library with citations, evaluated on real agent questions. If answers are factually right but not in the house format, add a small SFT adapter that teaches response structure and citation style using retrieved context as input.

Production consideration: A tuned model that "remembers" an outdated policy is a mis-selling risk; retrieval with document versioning keeps answers auditable.

55. Training loss drops to nearly zero within the first few hundred steps, or suddenly spikes to NaN. How do you debug the run?

Answer: Near-zero loss early usually means the model is not learning the task but copying something easy; NaN spikes usually mean numerical instability.

What I would check:

  1. For near-zero loss: decoded batches, to see whether targets are trivially predictable, duplicated, or whether masking is broken so loss covers only padding or template tokens.
  2. Whether the label and input are identical by a data-pipeline bug.
  3. For NaN: learning rate and warm-up, gradient norms before the spike, mixed-precision settings (16-bit overflow), and gradient clipping.
  4. Specific bad examples at the spike step, such as extremely long sequences or corrupted characters.

Fix the data or masking bug first, then lower the rate, add warm-up and clipping, or switch to a more numerically stable precision. Resume from the last healthy checkpoint rather than restarting blindly.

Production consideration: Log loss, gradient norm, learning rate and a few decoded samples to your experiment tracker on every run, so failures are diagnosable afterwards.

Key takeaways

  • Fine-tuning changes behaviour, not knowledge; reach for it after prompting and RAG have been measured, and consider distillation when the goal is cost at volume.
  • Data decides outcomes: consistent targets, real coverage, deduplication before splitting, and a frozen, contamination-checked test set.
  • Render training data with the model's own chat template and mask the loss to assistant tokens.
  • LoRA on all linear layers at a modest rank is the default; rank, alpha and learning rate are chosen by experiment, and QLoRA buys memory at some speed and fidelity cost.
  • Match the preference method to the feedback you have: pairs (DPO, ORPO), single ratings (KTO), or checkable outcomes (verifiable-reward RL such as GRPO).
  • Every candidate passes task, regression and safety suites against the baseline before an A/B rollout.
  • Adapters are tied to one base version, one licence and one dataset version; version all three and serve them with clear tenant routing.

Interview preparation checklist

  • Be able to draw the prompt, RAG, fine-tune and distil decision ladder and justify each step with evaluation.
  • Write the LoRA forward equation from memory and explain rank, alpha, initialisation and target modules.
  • Do the memory estimate for full fine-tuning, LoRA and QLoRA of a 7B model on a whiteboard.
  • Run one small QLoRA SFT job end to end on a small open model: render with the chat template, mask the loss, evaluate against a prompted baseline.
  • Build a tiny DPO dataset from your own SFT outputs and observe how beta changes length and drift.
  • Prepare a regression and safety suite you can describe item by item.
  • Explain how managed fine-tuning works on at least one cloud, with data-residency questions you would ask.
  • Rehearse two scenario answers aloud: forgetting after SFT, and inconsistent JSON.
  • Review neighbouring topics in our NLP interview questions and MLOps and LLMOps interview questions.

FAQ

What skills does an LLM fine-tuning engineer need?

Solid Python and PyTorch, an understanding of transformers and tokenisation, hands-on use of PEFT and preference-tuning libraries, data cleaning and labelling discipline, evaluation design, GPU memory reasoning, and enough MLOps to version, serve and monitor adapters.

Do I need a powerful GPU to practise fine-tuning?

No. QLoRA on a small open model fits on a single modest GPU or a free or low-cost cloud notebook. The skills interviewers test, such as data preparation, evaluation and debugging, do not depend on large hardware.

Are LoRA interview questions mostly about theory or practice?

Both. Expect the update equation and the meaning of rank and alpha, then practical follow-ups on choosing rank, target modules, learning rate and how you verified the choice on a test set.

How important are DPO and RLHF for enterprise fine-tuning roles?

You should explain them clearly and know when each fits, but most enterprise projects rely on SFT first. Preference tuning and reward-based fine-tuning come up for tone, judgement and checkable tasks.

Is fine-tuning still relevant when models keep getting better at following prompts?

Yes, for narrower reasons: cheaper smaller models through tuning or distillation, strict output formats, domain behaviour and tool-calling reliability. Interviewers value candidates who can say when not to fine-tune.

How should a fresher prepare for fine-tuning interview questions?

Learn the fundamentals in this guide, complete one small end-to-end project with a baseline, a tuned adapter and an honest evaluation write-up, and be ready to explain every decision in it.

Do managed services like Bedrock or Foundry remove the need to understand fine-tuning?

No. They remove GPU management, but you still prepare data, choose methods, evaluate, check safety and decide whether the tuned model is worth its hosting cost.

Is LLM fine-tuning a good career direction for Indian engineers?

It is a valuable specialisation within AI engineering, especially in GCCs and product teams running open models. Pair it with evaluation, serving and cloud skills so you are not limited to training alone.

If you want structured, mentor-led preparation covering machine learning, cloud deployment and AI security, explore the Cloudsoft APEX program. If your goal is to take tuned and retrieval-based AI systems into real customer environments, the AI Forward Deployed Engineer FDE PRO program adds enterprise integration projects, a simulated customer capstone and placement support until you're placed. Classroom in Ameerpet or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us