LLM Fine-Tuning
A base model already knows a lot about language, code, and the world. Fine-tuning does not always mean rewriting that knowledge from scratch.
Often it means steering the model — teaching it a tone, a format, and a way of answering that fits your domain.
Start with the difference between a base model and a fine-tuned model ↓
Prefer watching the video version?
Imagine you ask a model:
The Docker container is restarting continuously. How do I debug it?
A general base model may give a vague answer.
A model fine-tuned for DevOps may answer with a clear troubleshooting sequence and real commands.
1. What is a base model?
A base model is usually trained on a huge amount of internet-scale data.
Conceptually:
Scrape / collect large-scale data
↓
Train the model
↓
Base Model
After this stage, the model understands language, patterns, and a lot of general knowledge.
But it is not necessarily optimized for:
Following instructions cleanly Using a consistent DevOps tone Returning structured troubleshooting steps Matching your company's preferred answer style
That is where fine-tuning comes in.
2. What does fine-tuning do?
Fine-tuning takes an existing base model and continues training it on a smaller, more carefully chosen dataset.
Base Model
↓
High-quality fine-tuning dataset
↓
Fine-tuned Model
A useful mental model is:
Base Model → Trillions of examples Fine-Tuning → Maybe a few thousand examples to set the tone
Fine-tuning often only needs to slightly steer the model toward new behaviors, formats, or domains.
For example, an instruction-tuned model such as google/gemma-4-31B-it has already been adapted to follow instructions better than a raw base model.
3. SFT, mid-training, and RLHF
You will often hear a few related terms.
SFT — Supervised Fine-Tuning
Train the model on high-quality examples of the behavior you want.
For example:
User: Debug this CrashLoopBackOff Pod Assistant: Check logs, events, probes, and recent deploys...
Mid-training
Sometimes teams continue training on domain data before the final instruction-tuning stage.
RLHF — Reinforcement Learning from Human Feedback
Humans prefer one answer over another, and that preference signal helps shape the model's behavior.
For beginners, the most important starting idea is still SFT:
Show the model many good examples of the answers you want.
4. Prompt engineering, RAG, and fine-tuning
These are not the same thing.
Prompt engineering changes the instructions you send to the model.
Write a funny joke Draw a beautiful picture Answer like a senior SRE
RAG retrieves external documents and adds them to the prompt.
Fine-tuning changes the model weights themselves.
In many real systems, teams combine them:
Prompt Engineering
+
RAG
+
Fine-Tuning
A useful way to think about it:
Prompting → change the request RAG → change the context Fine-tuning → change the model's default behavior
5. Full fine-tuning, LoRA, and QLoRA
There are three approaches you will hear constantly:
1. Full fine-tuning 2. LoRA 3. QLoRA
Full fine-tuning updates a large portion — or all — of the model weights. It is powerful, but expensive in compute and memory.
LoRA (Low-Rank Adaptation) freezes most of the original model and trains small adapter matrices instead.
QLoRA combines quantization with LoRA, so you can fine-tune larger models with less GPU memory.
For many DevOps engineers experimenting on consumer GPUs, LoRA and QLoRA are the practical starting points.
6. Fine-tuning is often steering, not rewriting
One of the most useful research ideas for understanding fine-tuning comes from this paper:
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
One of the key insights from the paper is that fine-tuning is largely a form of steering rather than completely rewriting a language model.
During pretraining, the model already learns broad knowledge about language, reasoning, coding, and the world.
Fine-tuning typically does not need to modify all billions of parameters. Instead, it only needs to move the model in a small number of important directions within its parameter space — what the paper calls a low intrinsic dimension.
These relatively small changes are often enough to alter the model's behavior, priorities, and decision boundaries for a specific task.
That also helps explain why parameter-efficient methods like LoRA can achieve strong results by updating only a tiny fraction of the model's parameters.
In short:
Pretraining → learns broad capability Fine-tuning → steers that capability in a low-dimensional way LoRA / QLoRA → take advantage of that low-dimensional structure
Related reading while you are here:
- LoRA: Low-Rank Adaptation of Large Language Models
- yahma/alpaca-cleaned — a common high-quality instruction dataset for SFT experiments
7. Catastrophic forgetting
If you fine-tune too aggressively, the model may become better at your new task and worse at older skills.
That problem is often called catastrophic forgetting.
This is one reason teams prefer careful datasets, smaller learning rates, and parameter-efficient methods like LoRA.
You usually want the model to keep its general ability while learning a new style or specialty.
8. What does LoRA mean?
LoRA stands for:
Low-Rank Adaptation
Let's split that into two ideas.
Rank
The rank of a matrix is the number of linearly independent rows or columns.
In plain language:
How much independent information does this matrix actually contain?
Adaptation
Instead of updating the giant original weight matrix directly, LoRA learns a smaller update and adapts the model through that update.
9. A simple rank example
Look at this matrix:
Suppose:
C1 and C2 are linearly independent C3 = C1 + 2×C2 C4 = C1 + C2 C5 depends on earlier columns
Then the matrix may look large, but its true rank can still be small — for example, rank 2.
That is the intuition behind LoRA:
A big update matrix can often be approximated by a product of two much smaller matrices
10. Low-rank matrices instead of one giant matrix
Imagine a weight update matrix with size:
4 × 5 = 20 parameters
LoRA can approximate that update using two smaller matrices:
Conceptually:
4×2 = 8 parameters 2×5 = 10 parameters Total = 18 parameters
In real LLMs the savings are much more dramatic.
Example:
Original matrix: 4096 × 4096 = 16,777,216 parameters LoRA with rank 2: 4096×2 + 2×4096 = 16,384 parameters
That is why LoRA is called parameter-efficient fine-tuning.
11. Freeze the original weights, train the adapters
In LoRA:
Original weights → frozen LoRA adapters → trained
At inference time, the adapter update influences the original model.
A common control is:
lora_alpha rank (r)
The scaling factor is often:
scaling = alpha / rank
For example:
alpha = 16, rank = 16 → scaling = 1 alpha = 32, rank = 16 → scaling = 2
In plain English:
How strongly should the LoRA update influence the original model?
12. What LoRA rank should you use?
Ablation studies suggest that simple tasks often need only a small number of trainable parameters, while complex reasoning usually needs a higher rank.
A practical starting guide:
Task Type Typical LoRA Rank (r) ----------------------------------------------------------- Simple classification, sentiment, NER 2–8 General instruction tuning 8–16 Coding, math, reasoning 16–64 Very complex / large-domain adaptation 64–128 (sometimes higher)
Higher rank can capture more task-specific detail.
But if the rank becomes very high, training can become less stable. In those cases, options such as Rank-Stabilized LoRA (use_rslora) may help.
13. What is QLoRA?
QLoRA means Quantized LoRA.
The idea is:
Quantize the base model to use less memory
+
Train LoRA adapters on top
This matters a lot on smaller GPUs.
Conceptually, memory pressure can drop dramatically:
2B model class examples 8GB 4GB 2GB 1GB VRAM territory with aggressive quantization setups
Exact numbers depend on model size, sequence length, batch size, and tooling.
The important idea is:
QLoRA makes fine-tuning possible on hardware that cannot afford full fine-tuning.
14. The libraries you will see again and again
Modern fine-tuning usually combines several libraries:
trl peft accelerate bitsandbytes unsloth
TRL — Transformers Reinforcement Learning. Commonly used for SFT workflows.
PEFT — Parameter-Efficient Fine-Tuning. This is where LoRA adapters live.
Accelerate — helps run training efficiently across devices.
bitsandbytes — quantization support.
Unsloth — optimizes these building blocks so training can run faster and use less GPU memory, including on consumer-grade hardware such as a T4 GPU.
A useful mental model is:
trl / peft / bitsandbytes
=
building blocks
Unsloth
=
make those building blocks faster and lighter
Unsloth does not replace the ecosystem. It patches and accelerates parts of it — attention kernels, LoRA operations, gradient checkpointing, quantization workflows, and memory handling.
15. Batching and padding
Training examples are not all the same length.
So we often pad shorter sequences so they can sit in the same batch:
[1, 2, 3, 4] [0.1, 0.2, 0.3, 50256] ← pad token [0.5, 0.6, 50256, 50256] ← more padding
This is a practical detail, but it matters for GPU efficiency.
Even with good libraries, training can still be slow if the underlying Transformer operations are not optimized. That is one reason specialized kernels and tools like Unsloth became popular.
16. What healthy training looks like
When you fine-tune, you should watch the metrics.
Useful signals:
Training loss → is the model learning? Gradient norm → is training numerically stable? Learning rate → is the schedule behaving as expected? Eval loss → are we overfitting?
If evaluation is not configured, you are flying partially blind. Set an eval dataset and eval steps when you can.
17. Warmup and decay
A common learning-rate schedule looks like this:
Conceptually:
Warmup ↑ learning rate increases gradually Peak Decay ↓ learning rate decreases
Warmup helps training start more stably. Decay helps the model settle as training continues.
18. Full fine-tuning vs LoRA vs QLoRA
Full fine-tuning → strongest flexibility → highest cost → highest memory LoRA → freeze base model → train small adapters → much cheaper QLoRA → quantize + LoRA → best fit for limited VRAM
If you are learning on a laptop GPU, Colab T4, or a small cloud instance, start with QLoRA or LoRA.
19. What do you do with the fine-tuned model?
After training, common next steps include:
1. Keep Base Model + LoRA adapters separate 2. Merge adapters into the base model 3. Export to a portable format such as GGUF
Merging makes deployment simpler.
Keeping adapters separate makes experimentation easier — you can swap adapters without rebuilding the whole model.
GGUF is popular when you want to run models efficiently with local inference tools.
20. Why this matters for DevOps engineers
Fine-tuning is not only an ML research topic.
As a DevOps or platform engineer, you may need to think about:
GPU memory Training job reliability Dataset versioning Checkpoint storage Reproducibility (random seeds) Export formats Serving the resulting model Cost of full FT vs LoRA vs QLoRA
Even simple operational habits matter:
docker ps docker ps -a
because your training environment itself is part of the system you operate.
21. The most important fine-tuning mental model
Base Model = broad knowledge from large-scale training Fine-Tuning = steer the model toward a desired behavior LoRA / QLoRA = do that steering without updating every weight Prompting + RAG + Fine-Tuning = complementary tools, not competitors
If RAG answers "what information should I use right now?", fine-tuning answers "how should this model behave by default?"
22. If you remember only seven things
1. A base model is trained on huge general data. Fine-tuning adapts it for a specific behavior.
2. Fine-tuning often steers the model instead of rewriting everything.
3. Prompting, RAG, and fine-tuning solve different problems and are often combined.
4. Full fine-tuning is powerful but expensive. LoRA freezes most weights and trains small adapters.
5. Rank means how much independent information an update can represent. Higher rank can help harder tasks.
6. QLoRA adds quantization so fine-tuning fits into less GPU memory.
7. In practice you will use stacks like TRL, PEFT, bitsandbytes, Accelerate, and Unsloth.