Fine-tuning LLMs means continuing to train a pre-trained model on your own examples so it behaves differently: it follows a fixed output format, applies your label scheme, or uses your house style without a long prompt. For most enterprise teams the practical path is supervised fine-tuning with a parameter-efficient method such as LoRA or QLoRA, on a small, carefully cleaned dataset, judged by evaluation on a held-out test set before and after training. This guide assumes you have already decided fine-tuning is the right tool. If you have not, read our RAG vs fine-tuning decision guide first, because most knowledge problems are better solved with retrieval.
Types of fine-tuning, and which one you need
"Fine-tuning" covers several distinct techniques. Picking the wrong one is a common source of wasted GPU hours.
Full fine-tuning
Full fine-tuning updates every weight. It offers the most freedom, but needs memory for weights, gradients and optimizer state (several times the model's size) and produces a full model copy per task. For narrow enterprise tasks it is rarely necessary.
Parameter-efficient fine-tuning: LoRA
Parameter-efficient fine-tuning (PEFT) freezes the base model and trains a few extra parameters. The most widely used method is LoRA (Low-Rank Adaptation): LoRA fine-tuning adds pairs of small low-rank matrices alongside selected weight matrices, typically the attention projections, and trains only those. The result is an adapter, a small fraction of the base model's size, that is loaded on top of the original weights or merged into them. It cuts memory sharply and lets one base model carry many task-specific adapters. On narrow tasks it often approaches full fine-tuning quality; confirm that on your own evaluation.
QLoRA explained
QLoRA is LoRA trained on top of a quantized base model. The frozen base weights are stored in 4-bit precision to save memory, while the LoRA adapter is trained in higher precision. Gradients flow through the quantized base into the adapter. This lets you tune larger models on a single GPU, at the cost of slower steps and a possible small quality loss that your evaluation should measure. Quantization itself is covered in our guide to self-hosting LLMs.
Supervised fine-tuning on instruction pairs
Supervised fine-tuning (SFT) is the objective you will use most: each example pairs an input (instruction plus context) with the exact output you want. It is how you teach a label set, an extraction schema or a response structure. LoRA, QLoRA or full tuning describe which weights change; SFT describes what signal they learn from.
Preference tuning: DPO and similar methods
Sometimes you cannot write one correct answer, but you can say which of two is better. Preference tuning learns from that comparison. Reinforcement learning from human feedback trains a separate reward model and optimises the LLM against it. Direct Preference Optimization (DPO) and related methods skip the reward model: they train directly on triples of prompt, preferred response and rejected response, nudging the model to raise the likelihood of preferred answers relative to a frozen reference copy of itself. It usually comes after SFT, when format is right but judgement or tone is still off.
Managed fine-tuning from cloud providers
Amazon Bedrock, Azure OpenAI and Google Cloud's Vertex AI all offer managed fine-tuning for some of their models. You upload a dataset, pick a supported base model and a few settings, and the platform trains and hosts the result. You give up control over method and hyperparameters, but avoid managing GPUs, which makes it a fast first experiment. Supported models, regions and hosting charges change, so check current documentation.
| Approach | What changes | Typical fit |
|---|---|---|
| Full fine-tuning | All weights | Large behaviour shifts, teams with serious GPU capacity |
| LoRA | Small adapter matrices | Most narrow enterprise tasks on open-weight models |
| QLoRA | Adapter on a 4-bit base | Larger models on limited GPU memory |
| DPO / preference tuning | Weights or adapter, from comparisons | Tone and judgement after SFT already works |
| Managed fine-tuning | Provider-controlled | Fast experiments without running GPUs |
Preparing training data
Data quality decides the outcome far more than any hyperparameter. A few hundred to a few thousand clean, consistent examples usually beat a much larger noisy pile.
Format
Most tools accept JSON Lines, one example per line; for chat models each example is a list of system, user and assistant messages, with the assistant message as the target. Format with the base model's own chat template and include your production system prompt, because a training-versus-inference format mismatch quietly ruins results.
Quality over quantity
- Consistency: if two annotators would label the same ticket differently, the model learns the confusion. Write a labelling guide and resolve disagreements before training.
- Coverage: include edge cases, rare classes and messy real inputs (typos, mixed languages, forwarded chains).
- Balance: rebalance heavy class skew in training, but keep the test set at real-world proportions.
- Correctness: every target must be acceptable in production; the model copies mistakes faithfully.
Cleaning and PII
Deduplicate, strip boilerplate such as signatures, and validate that every structured target parses against its schema. Then handle personal data: anything in training data can in principle be reproduced, and an adapter cannot apply per-user access rules. Mask names, phone numbers, account numbers, Aadhaar and PAN numbers and similar identifiers unless the task needs them, and record what was done. Our guide to enterprise AI governance covers that process.
Train, validation and test splits without leakage
Use a training set to learn from, a validation set to choose checkpoints and settings, and a test set touched only for the final comparison. Prevent leakage by:
- Deduplicating before splitting, so near-copies do not land on both sides.
- Splitting by group where items are related: all messages from one ticket thread, one customer or one document go into the same split.
- Splitting by time when the task drifts, training on older data and testing on newer, which mirrors production.
- Never tuning against test-set failures and then reporting on the same test set.
How to fine-tune an LLM: the end-to-end process
Define task + success metric
|
Baseline: prompt + few-shot on test set
|
Collect, label, clean, de-identify data
|
Split: train / validation / test
|
Pick base model (licence, size, fit)
|
Train adapter (LoRA / QLoRA / managed)
|
Evaluate vs baseline + regression suite
|
Pass? --no--> fix data, then settings
| yes
v
Serve adapter, monitor, retrain on drift
The baseline is your control: without a prompted-model score on the same test set, you cannot show fine-tuning helped. And "fix data, then settings" is ordered on purpose.
An illustrative training configuration
The sketch below shows the shape of a QLoRA setup in the Hugging Face style, using the Transformers and PEFT libraries. It is illustrative only: argument names and defaults change between library versions, and the placeholder values in capitals are deliberately not real numbers. Check the current documentation for the versions you install.
# Illustrative only - not tied to library versions.
from transformers import (
AutoModelForCausalLM, BitsAndBytesConfig)
from peft import LoraConfig, get_peft_model
# QLoRA: load the frozen base in 4-bit
quant = BitsAndBytesConfig(load_in_4bit=True)
base = AutoModelForCausalLM.from_pretrained(
BASE_MODEL_ID, quantization_config=quant)
lora = LoraConfig(
r=RANK, # adapter capacity
lora_alpha=ALPHA, # update scaling
lora_dropout=DROPOUT,
target_modules=TARGET_LAYERS, # e.g. attention
task_type="CAUSAL_LM")
model = get_peft_model(base, lora)
# Sanity check: only a small fraction trains
model.print_trainable_parameters()
# Then: a trainer with LEARNING_RATE, EPOCHS, BATCH_SIZE,
# evaluation on the validation set each epoch, and saving
# the adapter (not the whole model) at the end.
For plain LoRA, drop the quantization config.
Hyperparameters, explained conceptually
There are no universal right values; toolkit defaults are a reasonable first run. What matters is knowing what each knob does.
- Learning rate: how big each weight update is. Too high and loss is unstable or general skills erode; too low and it barely learns.
- Epochs: how many passes over the training data. Too many on a small dataset leads to memorisation; keep the checkpoint where validation metrics peak.
- Batch size and gradient accumulation: how many examples contribute to each update. Accumulation simulates larger batches when memory is tight.
- LoRA rank: the capacity of the adapter. Low rank often suffices for format and classification; raising it does not fix bad data.
- LoRA alpha: scales how strongly the adapter's update is applied relative to the base weights. It interacts with learning rate, so change one at a time.
- Target modules and sequence length: more adapted layers add capacity and cost; too short a maximum length silently truncates examples.
The signal to watch is the gap between training and validation results. Training loss falling while validation loss rises means overfitting: use fewer epochs, more varied data or a lower learning rate.
Evaluation before and after training
Fine-tuning without evaluation is guesswork. Use the same held-out test set and the same scoring code for the baseline and every fine-tuned candidate.
- Task metrics: per-class precision and recall for classification (overall accuracy hides failures on rare, costly classes), field-level exact match and schema validity for extraction, rubric scores for generated text.
- Operational metrics: latency, tokens per request and cost per task.
- Error review: read failures by hand; patterns point at labelling problems.
Guarding against catastrophic forgetting and regressions
Training hard on a narrow task can erode a model's general abilities, a problem called catastrophic forgetting. It shows up as worse instruction-following, reasoning or safety behaviour. LoRA reduces the risk by freezing base weights, but does not remove it. Protect yourself with a regression suite alongside the task test set: general instructions, out-of-scope requests the model should decline, safety and prompt-injection probes, and inputs from neighbouring tasks the same deployment handles. Make it a gate: a candidate that wins on the task but fails the regression suite does not ship. Our guide to LLM evaluation covers building these suites, LLM-as-judge scoring and CI gates in depth.
Serving fine-tuned adapters
You have two main options once an adapter passes evaluation.
- Merge the adapter into the base weights and serve the result as an ordinary model. Inference has no extra overhead, but you now version and host a full model per task.
- Serve adapters dynamically on a shared base. Some inference servers, vLLM among them, can load several LoRA adapters on one base model and select the adapter per request. This suits many narrow tasks, such as one adapter per document type.
Treat the adapter as a release artefact, versioned with its dataset, base model identifier and evaluation report, with the previous one ready for rollback. After QLoRA, evaluate in the exact precision you will serve. Deployment, GPU sizing and gateways are covered in self-hosting LLMs; the release pipeline is covered in our LLMOps guide.
Cost and GPU needs, conceptually
Training memory depends on model size, weight precision, full versus adapter training, sequence length and batch size. Adapter training of small and mid-sized models is feasible on a single GPU, while full fine-tuning of large models needs multi-GPU setups.
Compute is often not the main cost. Budget for labelling and review, repeated experiments, evaluation runs, hosting (a self-hosted model costs money even when idle; managed providers may charge for custom-model hosting) and retraining when the base model or label scheme changes. This is why fine-tuning a small language model for one narrow, high-volume task is the case where the economics most often work.
Licensing of base models
Check the licence before training, not after. Some open-weight models use permissive open-source licences; others use custom community licences with acceptable-use policies, attribution requirements, usage-scale restrictions or limits on improving other models. Check commercial use, derivative works (an adapter may count) and redistribution. Hosted providers' terms may also restrict training on their outputs, which matters if a large model generates your training data. Record the base model, its licence version and your legal sign-off with every adapter.
When to stop and use RAG or prompting instead
Stop and rethink if any of these are true:
- The model needs facts that change: policies, prices, product catalogues. That is retrieval; see what RAG is.
- Users need citations, or different users must see different information. An adapter cannot provide either.
- A well-structured prompt with a few examples and the provider's structured-output mode already meets your threshold. For strict JSON, constrained decoding and schemas often fix format problems with no training; see function calling and structured outputs.
- You cannot produce a few hundred consistent, reviewed examples, or the business cannot agree what the correct output is.
- Repeated data fixes stop moving test-set metrics.
Want hands-on practice with prompting, RAG, structured outputs and model customisation on Bedrock, Azure OpenAI and Gemini? Cloudsoft's AI, GenAI and Agentic AI course is built around labs rather than slides.
Illustrative example: an IT ticket classifier
Consider a GCC IT team in Hyderabad supporting a bank's internal users. Every incoming ticket must be classified into a fixed category set, given a priority, and have a few fields extracted (application name, affected location, error code) into JSON that the ITSM workflow consumes. The facts are all in the ticket itself, so nothing needs retrieving.
- Baseline. The team prompts a large hosted model with the category definitions and several examples. It is inconsistent on rare, expensive categories (security incidents, payment outages), and the long prompt is costly at volume.
- Data. Historical routing proves inconsistent, so senior engineers relabel a sample against a written guide. Personal identifiers are masked; splits are by thread and time.
- Training. They train a LoRA adapter on a small open-weight model whose licence permits commercial internal use, with the short production system prompt included in every example.
- Evaluation. They compare per-class recall, JSON validity and field accuracy against the baseline, plus a regression suite including prompt-injection attempts inside ticket text.
- Serving. The adapter runs behind an internal gateway; low-confidence or schema-invalid outputs escalate to the larger model or a human queue, and corrections feed the next training round.
The same pattern applies to structured extraction from claim forms at an insurer or discharge summaries at a hospital, provided the de-identification step is taken seriously. Taking a system like this from notebook to a monitored production service inside a customer's environment is exactly the work Forward Deployed Engineers do, which is the focus of Cloudsoft's FDE PRO program.
Common fine-tuning mistakes
- Fine-tuning to inject knowledge. The model learns patterns, not reliable facts, and cannot cite them.
- No baseline. Without a prompted-model score on the same test set, improvement is unprovable.
- Leaky splits. Near-duplicates or related tickets on both sides inflate scores that collapse in production.
- Chat template mismatch between training and inference.
- Tuning hyperparameters to fix bad labels. Inconsistent data caps quality whatever the settings.
- Ignoring regressions in general ability and safety behaviour.
- Leaving PII in training data, where no access control can reach it.
- Skipping the licence check until after the model is in production.
- Treating the adapter as a one-off, with no versioning, rollback or retraining plan.
Frequently asked questions
What is the difference between LoRA and QLoRA?
LoRA freezes the base model and trains small low-rank adapter matrices on selected layers. QLoRA does the same, but stores the frozen base model in 4-bit precision to save memory, which lets you fine-tune larger models on a single GPU. QLoRA steps are slower and can cost a little quality, so compare both on your test set.
How much data do I need to fine-tune an LLM?
There is no fixed number. Narrow tasks such as classification or extraction often work with a few hundred to a few thousand clean, consistent, representative examples. Quality and coverage of edge cases matter more than volume, and your test-set results tell you whether you have enough.
When should I use DPO instead of supervised fine-tuning?
Use supervised fine-tuning when you can write the correct output for each input. Consider DPO or similar preference methods when supervised fine-tuning already gets the format right but you want to shift judgement or tone, and you can collect pairs of better and worse responses to the same prompt.
Can fine-tuning make a model forget what it knew?
Yes. Catastrophic forgetting can weaken general instruction-following, reasoning or safety behaviour after narrow training. LoRA reduces the risk because base weights stay frozen, but you should still run a regression suite of general, out-of-scope and safety tests before shipping any fine-tuned model.
Should I use managed fine-tuning or train my own adapter?
Managed fine-tuning on a cloud platform is faster to start and needs no GPU operations, but it limits model choice and control over training settings. Training your own LoRA adapter on an open-weight model gives full control and portability, at the cost of running training and serving infrastructure yourself.
Can I fine-tune a model to answer questions from our company documents?
You can, but it is usually the wrong tool. Fine-tuning does not reliably store facts, cannot cite sources, cannot enforce per-user access and goes stale when documents change. Retrieval-augmented generation is the better fit for document question answering; fine-tuning suits format, tone and narrow tasks.
Next steps
If you want to build, evaluate and ship fine-tuned and retrieval-based systems, explore GenAI and Agentic AI training in Hyderabad at Cloudsoft, in our Ameerpet classroom or live online. Call +91 96660 19191 to book a free demo.



