New batches starting this week Β· Limited seats

Model Distillation Explained: Getting Big-Model Quality from Smaller, Cheaper Models

Model distillation trains a small student model to imitate a large teacher on a narrow task, cutting cost and latency. This guide covers the workflow, data quality pitfalls, provider terms and how distillation compares with quantisation and pruning.

Model distillation workflow: task and eval set, teacher outputs, filtering and review, fine-tuning the student, comparing to the teacher
Last updated Β· 15 min read Β· 3,192 words

Model distillation is the practice of training a smaller "student" model to reproduce the behaviour of a larger "teacher" model, so you can serve the student at lower cost and latency. In the LLM era, model distillation usually means running a large model over many real inputs from your task, filtering and reviewing its outputs, and fine-tuning a small model on those input-output pairs until it matches the teacher on your evaluation set. It works well for narrow, high-volume tasks. It does not make a small model generally as capable as a large one, and the terms of the teacher model's provider may limit what you are allowed to do with its outputs.

What is model distillation?

Every distillation setup has two models. The teacher is large, capable and expensive: a frontier model behind an API, or a large open-weight model on your own GPUs. The student is smaller and cheaper, and you want it to behave like the teacher on one specific job. You treat the teacher's behaviour as the training signal and teach the student to copy it.

The term covers two quite different techniques, and mixing them up causes confusion in design reviews.

Classic knowledge distillation: soft labels

The original form of knowledge distillation, popularised by Geoffrey Hinton and colleagues, trains the student on the teacher's full output probability distribution rather than only the correct label. If the teacher thinks an image is most likely a cat but slightly resembles a fox, those "soft labels" carry more information than the hard label "cat". A temperature setting softens the distribution so the small probabilities become visible, and the student learns from a mix of the teacher's soft targets and the true labels.

This needs the teacher's logits (its raw scores for every possible output) and usually a shared vocabulary. Model builders use it to train their own small models from their own large ones. Most enterprise teams cannot do it with a hosted frontier model, because the API returns text, not full distributions.

LLM-era distillation: teacher-generated training data

What most enterprise teams mean by LLM distillation is simpler. You prompt the teacher, often with a long, carefully engineered prompt and few-shot examples, to produce outputs for thousands of real inputs. Those outputs become a supervised fine-tuning dataset, and the student learns to produce the same output from the bare input, without the long prompt. This is sometimes called sequence-level distillation or response distillation.

Seen this way, distillation is a data strategy for fine-tuning: the teacher plays the role of an expensive, fast annotator. Everything in our guide to fine-tuning LLMs still applies; distillation only changes where the labels come from. Some teams also have the teacher write a short rationale before the answer, which can help on multi-step tasks but lengthens outputs.

When distillation makes sense

Distillation is a cost-and-latency optimisation for a task you already know a large model can do. It fits when most of these are true:

  • The task is narrow and well defined. Classification, routing, field extraction, intent detection and fixed-template summaries suit it. Open-ended reasoning does not.
  • Volume is high and steady. At a handful of calls a day, distillation adds maintenance for no real gain. On every email, ticket or transaction, per-call savings add up.
  • Latency is under pressure. A small model without a long system prompt responds faster, which matters for inline checks and steps inside agent loops. Our piece on LLM latency optimisation covers other levers to try first.
  • The model must run somewhere constrained. On laptops, phones, factory gateways or air-gapped servers, a large model is not an option; see our sibling article on on-device and edge AI.
  • Residency rules push you to self-host. A distilled open-weight student can run inside your own network.

It is usually the wrong move when the task changes every week, when the large model itself is not yet good enough, or when the real problem is missing knowledge, which is a retrieval problem (see RAG vs fine-tuning). Distillation also pairs well with routing and cascade patterns, covered in small language models for enterprise.

The distillation workflow, step by step

A distillation project has five stages. Teams that skip the first one usually regret it.

Define task + eval set
        |
        v
Teacher generates outputs --> filter + human review
        |
        v
Fine-tune student (LoRA / full)
        |
        v
Evaluate student vs teacher --> fail? fix data
        |
        v
Deploy + monitor drift --> refresh data, retrain

1. Define the task and build the eval set

Write down the inputs, output schema, label set and how ambiguous cases are handled. Then build a held-out evaluation set from real inputs, labelled by domain experts, not by the teacher. It must cover hard cases, rare classes and messy formatting, and you never train on it.

Run the teacher on this set first. If the teacher misses the bar, distilling it will not help; its score is the student's practical ceiling. Our guide to LLM evaluation covers metrics, judges and regression suites.

2. Generate teacher outputs, with filtering and human review

Collect a large pool of real, unlabelled inputs that reflects production traffic, mask personal data where policy requires, and run the teacher with your best prompt. Then clean the data before it reaches training:

  • Schema and rule checks: drop outputs that fail to parse or use labels outside the allowed set, and flag ones that break business rules.
  • Agreement filters: sample the teacher twice or use two prompts, and send disagreements to review.
  • Human review: experts check a random sample plus every flagged item and correct errors. Corrections become training data; error patterns show where to improve the prompt.
  • Deduplication and balance: remove near-duplicates and make sure rare-but-important classes are well represented.

Where real inputs are scarce, the teacher can generate synthetic inputs too, but treat them with suspicion; see synthetic data for AI testing for keeping them realistic and out of your test set.

3. Fine-tune the student

Choose a student that is small enough to meet your cost, latency or device target and has a licence that permits your use; our sibling guide to open-weight LLMs for enterprise covers licence checks. Train it with supervised fine-tuning on the cleaned pairs, using the short input it will see in production, not the teacher's long prompt. LoRA, or QLoRA on a quantised base, is the usual start because it is cheap and easy to version; full fine-tuning is an option when the student is small and adapters plateau. Record the dataset version, base model and hyperparameters for every run.

4. Evaluate the student against the teacher

Score the student, the teacher and the un-tuned student on the same held-out set, and look past the headline metric:

  • Per-class results: a student can match the teacher overall while failing on one rare, costly category.
  • Error overlap: shared failures point to the data; student-only failures may need more examples or a larger base.
  • Calibration: check that low confidence really goes with errors, because that is what lets you route hard cases to the teacher or a human.
  • Latency and cost measured on your real serving stack, plus robustness to malformed, out-of-scope and prompt-injection inputs.

Agree the acceptance bar with the business owner beforehand; "close enough to the teacher" should depend on what an error costs.

5. Deploy and monitor drift

A student is a snapshot of the teacher's behaviour on yesterday's data, while production brings new products, regulations and phrasing. Monitor input and label distributions, keep sampling live traffic for review, and run the teacher in shadow on a small slice to track agreement. When agreement falls, regenerate data for the new patterns and retrain. Budget for this cycle; it is an operating cost, not a one-off.

If you want hands-on practice with fine-tuning, evaluation and deployment on AWS, Azure and Google Cloud, Cloudsoft's AI, GenAI and Agentic AI course in Hyderabad works through these steps in labs.

Data quality pitfalls

The student can only be as good as the data it learns from. The common ways distillation goes wrong are all data problems.

  • Copying the teacher's errors. The student learns the teacher's mistakes as faithfully as its correct answers. If the teacher misreads a certain form layout, so will the student. Human review and an independently labelled eval set are the real defences.
  • Lack of diversity. Inputs from one week, region or channel produce a student that breaks on everything else. In India that often means sampling mixed English and regional languages, transliterated text and forwarded email chains.
  • Template collapse. If every teacher output starts the same way, the student learns the habit rather than the task.
  • Label leakage. If the teacher prompt included metadata the student will not see in production, the student is trained on a task it cannot perform.
  • Eval contamination and stale labels. Deduplicate across splits, and retire old examples when business rules change.

Licence and terms-of-service constraints

This is the step teams most often skip, and it can stop a project late. Many commercial model providers restrict using their outputs to develop models that compete with them, and some restrict training any model on outputs without permission. Wording differs between providers, between consumer and business terms, and between a direct API and the same model on a cloud platform, and it changes over time. Before generating training data:

  • Check the provider's terms for the exact teacher, tier and access route, including usage policies.
  • Check open-weight licences for both models; some add conditions on derivatives or on using outputs to improve other models.
  • Prefer sanctioned routes, such as a provider's own distillation feature or a cloud platform's managed service, which come with clearer terms for supported pairs.
  • Record the decision in your model register with legal sign-off and your data-protection assessment.

This is not legal advice: get the terms reviewed for your case and keep the evidence.

Managed distillation services

Some cloud platforms now package the workflow as a managed service. Amazon Bedrock Model Distillation, for example, lets you choose a supported teacher and a smaller student from the pairs Bedrock allows, supply prompts in JSONL files, and have the service generate teacher responses and fine-tune the student. The result is a custom model you then deploy within Bedrock. Other platforms offer comparable fine-tuning and distillation features. Supported model pairs, regions and pricing change, so check the current documentation rather than relying on a blog post.

Managed services remove infrastructure work, but not the need for your own eval set, data review and drift monitoring, and they tie you to that platform's models. Our Amazon Bedrock GenAI training covers Bedrock's model customisation options.

Distillation vs quantisation vs pruning

All three make models cheaper to run, and they are often used together: distil to a small student, then quantise it for serving.

AspectDistillationQuantisationPruning
What it doesTrains a smaller model to imitate a larger oneStores weights (and sometimes activations) at lower numeric precisionRemoves weights, attention heads or layers that contribute little
Changes architecture?Yes: the student is a different, smaller modelNo: same model, fewer bits per weightYes, partly: same model family with parts removed
Needs training?Yes: data generation plus fine-tuningOften none (post-training); some methods calibrate or retrainUsually some retraining afterwards to recover quality
ScopeUsually task-specificGeneral: keeps the model's broad abilities, with some lossGeneral or task-specific, depending on method
Main riskInherits teacher errors; narrow coverageQuality loss at very low precisionUneven damage to certain abilities
Best forHigh-volume narrow tasks; moving work off a large hosted modelFitting a chosen model onto smaller GPUs or devicesSpecialists shrinking a model further

For serving-side trade-offs such as GPU memory and batching, see self-hosting LLMs.

Illustrative example: email classification at an insurer

Consider a general insurer whose shared service centre in Hyderabad receives a heavy daily flow of customer and broker emails: new claims, status questions, policy changes, renewals, complaints, supporting documents and spam. A large hosted model with a detailed prompt already sorts them well and extracts policy number, claim number and urgency, but every email carries a long prompt, so cost grows with volume and responses are slow.

  1. Task and eval set. Operations leads write a label guide, including how to handle emails with several requests. Experienced handlers label a stratified sample of past emails, deliberately including rare, high-risk complaints. The teacher clears the bar, with a known weakness on multi-request emails.
  2. Teacher outputs. With personal data masked, the teacher labels a large pool of historical emails in a strict JSON schema. Every complaint label, every disagreement between two prompts and a random sample go to handlers, whose corrections go back into the data.
  3. Licence check. Legal confirms the teacher's terms for this access route permit training an internal classifier, and that the student's licence covers internal commercial use.
  4. Training and evaluation. A small open-weight model is fine-tuned with LoRA: email in, JSON out. It is close to the teacher overall but under-predicts complaints, so the team adds reviewed complaint examples, retrains, and routes low-confidence or possible-complaint emails to the teacher or a human.
  5. Operation. The student runs on the insurer's own infrastructure with the teacher in shadow on a slice. When a product launch shifts the renewal mix, falling agreement triggers a data refresh.

The large model stays involved as reviewer and fallback while the student does most of the work. Taking a system like this from a working pilot to an operated service inside a customer's environment is the kind of work Forward Deployed Engineers do, and Cloudsoft's FDE PRO program trains for it. For broader context on the domain, see generative AI in insurance.

Frequently asked questions

What is model distillation in simple terms?

Model distillation trains a small student model to copy a large teacher model's behaviour on a task. The teacher produces answers or probability scores, and the student is trained on them, so you get similar results on that task at lower cost and latency.

What is the difference between knowledge distillation and LLM distillation?

Classic knowledge distillation trains the student on the teacher's full output probabilities, known as soft labels, which requires access to the teacher's logits. LLM distillation as most enterprises practise it uses the teacher's generated text as fine-tuning data, which works even when the teacher is only available through an API.

Is distillation the same as fine-tuning?

Not exactly. Fine-tuning is the training step. Distillation is a way of getting the training data for that step, using a teacher model as the annotator, followed by fine-tuning the student on the teacher's outputs.

Can a distilled student model outperform its teacher?

Occasionally on a narrow task, especially when human reviewers correct the teacher's errors before training, but you should not plan on it. Treat the teacher's score on your evaluation set as the practical ceiling and aim to get close to it at lower cost.

It depends on the provider's terms. Many providers restrict using outputs to train competing models, and some restrict training any model on outputs without permission. Check the provider's terms for your exact model and access route, check the student's licence, and get legal sign-off.

How much data do I need for LLM distillation?

There is no fixed number. It depends on how narrow the task is, how many classes or output patterns it has and how capable the student already is. Start with a modest, well-reviewed dataset, measure on your held-out set, and add data where the errors are.

Should I use distillation, quantisation or pruning?

Use distillation to move a narrow, high-volume task from a large model to a much smaller one. Use quantisation to make a chosen model fit smaller hardware with little effort. Pruning is more specialised. Many teams distil first and then quantise the student for serving.

Do I still need the teacher model after distillation?

Usually yes, in a smaller role. Teams keep the teacher for low-confidence cases, for shadow comparisons to detect drift, and for generating fresh training data when the task or inputs change.

Ready to learn how fine-tuning, distillation, evaluation and deployment fit together in real systems? Cloudsoft's GenAI and Agentic AI training in Hyderabad runs in our Ameerpet classroom and live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us