LLM Fine-Tuning

A base model already knows a lot about language, code, and the world. Fine-tuning does not always mean rewriting that knowledge from scratch.

Often it means steering the model — teaching it a tone, a format, and a way of answering that fits your domain.

Start with the difference between a base model and a fine-tuned model ↓

Prefer watching the video version?

Video thumbnail for the Day 13 lesson: LLM Fine-Tuning
LLM Fine-Tuning — Video Tutorial Watch on YouTube ↗

Imagine you ask a model:

The Docker container is restarting continuously. How do I debug it?

A general base model may give a vague answer.

A model fine-tuned for DevOps may answer with a clear troubleshooting sequence and real commands.

Side-by-side comparison: a base model gives a generic Docker restart answer, while a fine-tuned model returns a step-by-step debugging flow with docker ps, docker logs, and docker inspect.
Fine-tuning can change how the model answers — not only what it knows.
Where it starts

1. What is a base model?

A base model is usually trained on a huge amount of internet-scale data.

Conceptually:

Scrape / collect large-scale data
        ↓
Train the model
        ↓
Base Model

After this stage, the model understands language, patterns, and a lot of general knowledge.

But it is not necessarily optimized for:

Following instructions cleanly
Using a consistent DevOps tone
Returning structured troubleshooting steps
Matching your company's preferred answer style

That is where fine-tuning comes in.

The idea

2. What does fine-tuning do?

Fine-tuning takes an existing base model and continues training it on a smaller, more carefully chosen dataset.

Base Model
     ↓
High-quality fine-tuning dataset
     ↓
Fine-tuned Model

A useful mental model is:

Base Model → Trillions of examples
Fine-Tuning → Maybe a few thousand examples to set the tone

Fine-tuning often only needs to slightly steer the model toward new behaviors, formats, or domains.

For example, an instruction-tuned model such as google/gemma-4-31B-it has already been adapted to follow instructions better than a raw base model.

Common stages

3. SFT, mid-training, and RLHF

You will often hear a few related terms.

SFT — Supervised Fine-Tuning

Train the model on high-quality examples of the behavior you want.

For example:

User: Debug this CrashLoopBackOff Pod
Assistant: Check logs, events, probes, and recent deploys...

Mid-training

Sometimes teams continue training on domain data before the final instruction-tuning stage.

RLHF — Reinforcement Learning from Human Feedback

Humans prefer one answer over another, and that preference signal helps shape the model's behavior.

For beginners, the most important starting idea is still SFT:

Show the model many good examples of the answers you want.
Choose the right tool

4. Prompt engineering, RAG, and fine-tuning

These are not the same thing.

Prompt engineering changes the instructions you send to the model.

Write a funny joke
Draw a beautiful picture
Answer like a senior SRE

RAG retrieves external documents and adds them to the prompt.

Fine-tuning changes the model weights themselves.

In many real systems, teams combine them:

Prompt Engineering
      +
RAG
      +
Fine-Tuning

A useful way to think about it:

Prompting → change the request
RAG → change the context
Fine-tuning → change the model's default behavior
How we fine-tune

5. Full fine-tuning, LoRA, and QLoRA

There are three approaches you will hear constantly:

1. Full fine-tuning
2. LoRA
3. QLoRA

Full fine-tuning updates a large portion — or all — of the model weights. It is powerful, but expensive in compute and memory.

LoRA (Low-Rank Adaptation) freezes most of the original model and trains small adapter matrices instead.

QLoRA combines quantization with LoRA, so you can fine-tune larger models with less GPU memory.

For many DevOps engineers experimenting on consumer GPUs, LoRA and QLoRA are the practical starting points.

Key insight

6. Fine-tuning is often steering, not rewriting

One of the most useful research ideas for understanding fine-tuning comes from this paper:

Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning

One of the key insights from the paper is that fine-tuning is largely a form of steering rather than completely rewriting a language model.

During pretraining, the model already learns broad knowledge about language, reasoning, coding, and the world.

Fine-tuning typically does not need to modify all billions of parameters. Instead, it only needs to move the model in a small number of important directions within its parameter space — what the paper calls a low intrinsic dimension.

These relatively small changes are often enough to alter the model's behavior, priorities, and decision boundaries for a specific task.

That also helps explain why parameter-efficient methods like LoRA can achieve strong results by updating only a tiny fraction of the model's parameters.

In short:

Pretraining
  → learns broad capability

Fine-tuning
  → steers that capability in a low-dimensional way

LoRA / QLoRA
  → take advantage of that low-dimensional structure

Related reading while you are here:

Risk

7. Catastrophic forgetting

If you fine-tune too aggressively, the model may become better at your new task and worse at older skills.

That problem is often called catastrophic forgetting.

This is one reason teams prefer careful datasets, smaller learning rates, and parameter-efficient methods like LoRA.

You usually want the model to keep its general ability while learning a new style or specialty.

LoRA

8. What does LoRA mean?

LoRA stands for:

Low-Rank Adaptation

Let's split that into two ideas.

Rank

The rank of a matrix is the number of linearly independent rows or columns.

In plain language:

How much independent information does this matrix actually contain?

Adaptation

Instead of updating the giant original weight matrix directly, LoRA learns a smaller update and adapts the model through that update.

Linear algebra intuition

9. A simple rank example

Look at this matrix:

A 4 by 5 matrix M with columns labeled C1 through C5.
Not every column adds new independent information.

Suppose:

C1 and C2 are linearly independent

C3 = C1 + 2×C2
C4 = C1 + C2
C5 depends on earlier columns

Then the matrix may look large, but its true rank can still be small — for example, rank 2.

That is the intuition behind LoRA:

A big update matrix
can often be approximated
by a product of two much smaller matrices
Why LoRA saves parameters

10. Low-rank matrices instead of one giant matrix

Imagine a weight update matrix with size:

4 × 5 = 20 parameters

LoRA can approximate that update using two smaller matrices:

Matrix A, a 4 by 2 low-rank factor.
Matrix A — one low-rank factor.
Matrix B, a 2 by 5 low-rank factor.
Matrix B — the other low-rank factor.

Conceptually:

4×2 = 8 parameters
2×5 = 10 parameters
Total = 18 parameters

In real LLMs the savings are much more dramatic.

Example:

Original matrix:
4096 × 4096 = 16,777,216 parameters

LoRA with rank 2:
4096×2 + 2×4096 = 16,384 parameters

That is why LoRA is called parameter-efficient fine-tuning.

How training works

11. Freeze the original weights, train the adapters

In LoRA:

Original weights → frozen
LoRA adapters → trained

At inference time, the adapter update influences the original model.

A common control is:

lora_alpha
rank (r)

The scaling factor is often:

scaling = alpha / rank

For example:

alpha = 16, rank = 16 → scaling = 1
alpha = 32, rank = 16 → scaling = 2

In plain English:

How strongly should the LoRA update influence the original model?
Choosing r

12. What LoRA rank should you use?

Ablation studies suggest that simple tasks often need only a small number of trainable parameters, while complex reasoning usually needs a higher rank.

A practical starting guide:

Task Type                              Typical LoRA Rank (r)
-----------------------------------------------------------
Simple classification, sentiment, NER  2–8
General instruction tuning             8–16
Coding, math, reasoning                16–64
Very complex / large-domain adaptation 64–128 (sometimes higher)

Higher rank can capture more task-specific detail.

But if the rank becomes very high, training can become less stable. In those cases, options such as Rank-Stabilized LoRA (use_rslora) may help.

Less VRAM

13. What is QLoRA?

QLoRA means Quantized LoRA.

The idea is:

Quantize the base model to use less memory
        +
Train LoRA adapters on top

This matters a lot on smaller GPUs.

Conceptually, memory pressure can drop dramatically:

2B model class examples

8GB
4GB
2GB
1GB VRAM territory with aggressive quantization setups

Exact numbers depend on model size, sequence length, batch size, and tooling.

The important idea is:

QLoRA makes fine-tuning possible on hardware that cannot afford full fine-tuning.
The stack

14. The libraries you will see again and again

Modern fine-tuning usually combines several libraries:

trl
peft
accelerate
bitsandbytes
unsloth

TRL — Transformers Reinforcement Learning. Commonly used for SFT workflows.

PEFT — Parameter-Efficient Fine-Tuning. This is where LoRA adapters live.

Accelerate — helps run training efficiently across devices.

bitsandbytes — quantization support.

Unsloth — optimizes these building blocks so training can run faster and use less GPU memory, including on consumer-grade hardware such as a T4 GPU.

A useful mental model is:

trl / peft / bitsandbytes
        =
building blocks

Unsloth
        =
make those building blocks faster and lighter

Unsloth does not replace the ecosystem. It patches and accelerates parts of it — attention kernels, LoRA operations, gradient checkpointing, quantization workflows, and memory handling.

Training details

15. Batching and padding

Training examples are not all the same length.

So we often pad shorter sequences so they can sit in the same batch:

[1, 2, 3, 4]
[0.1, 0.2, 0.3, 50256]      ← pad token
[0.5, 0.6, 50256, 50256]    ← more padding

This is a practical detail, but it matters for GPU efficiency.

Even with good libraries, training can still be slow if the underlying Transformer operations are not optimized. That is one reason specialized kernels and tools like Unsloth became popular.

Watch the run

16. What healthy training looks like

When you fine-tune, you should watch the metrics.

Training dashboard showing training loss decreasing, gradient norm stabilizing, learning rate warmup and decay, and eval loss not configured.
Loss should trend down. Gradient norms should stabilize. Learning rate usually warms up, then decays.

Useful signals:

Training loss → is the model learning?
Gradient norm → is training numerically stable?
Learning rate → is the schedule behaving as expected?
Eval loss → are we overfitting?

If evaluation is not configured, you are flying partially blind. Set an eval dataset and eval steps when you can.

Learning rate

17. Warmup and decay

A common learning-rate schedule looks like this:

Hand-drawn learning rate schedule with linear warmup followed by cosine decay over training steps.
Linear warmup, then cosine decay — a common pattern in fine-tuning.

Conceptually:

Warmup
  ↑ learning rate increases gradually

Peak

Decay
  ↓ learning rate decreases

Warmup helps training start more stably. Decay helps the model settle as training continues.

Choose carefully

18. Full fine-tuning vs LoRA vs QLoRA

Full fine-tuning
  → strongest flexibility
  → highest cost
  → highest memory

LoRA
  → freeze base model
  → train small adapters
  → much cheaper

QLoRA
  → quantize + LoRA
  → best fit for limited VRAM

If you are learning on a laptop GPU, Colab T4, or a small cloud instance, start with QLoRA or LoRA.

After the run

19. What do you do with the fine-tuned model?

After training, common next steps include:

1. Keep Base Model + LoRA adapters separate
2. Merge adapters into the base model
3. Export to a portable format such as GGUF

Merging makes deployment simpler.

Keeping adapters separate makes experimentation easier — you can swap adapters without rebuilding the whole model.

GGUF is popular when you want to run models efficiently with local inference tools.

DevOps mindset

20. Why this matters for DevOps engineers

Fine-tuning is not only an ML research topic.

As a DevOps or platform engineer, you may need to think about:

GPU memory
Training job reliability
Dataset versioning
Checkpoint storage
Reproducibility (random seeds)
Export formats
Serving the resulting model
Cost of full FT vs LoRA vs QLoRA

Even simple operational habits matter:

docker ps
docker ps -a

because your training environment itself is part of the system you operate.

Remember this

21. The most important fine-tuning mental model

Base Model
   = broad knowledge from large-scale training

Fine-Tuning
   = steer the model toward a desired behavior

LoRA / QLoRA
   = do that steering without updating every weight

Prompting + RAG + Fine-Tuning
   = complementary tools, not competitors

If RAG answers "what information should I use right now?", fine-tuning answers "how should this model behave by default?"

Before you go

22. If you remember only seven things

1. A base model is trained on huge general data. Fine-tuning adapts it for a specific behavior.

2. Fine-tuning often steers the model instead of rewriting everything.

3. Prompting, RAG, and fine-tuning solve different problems and are often combined.

4. Full fine-tuning is powerful but expensive. LoRA freezes most weights and trains small adapters.

5. Rank means how much independent information an update can represent. Higher rank can help harder tasks.

6. QLoRA adds quantization so fine-tuning fits into less GPU memory.

7. In practice you will use stacks like TRL, PEFT, bitsandbytes, Accelerate, and Unsloth.