1:1 mentoring with Big Tech AI engineers
LLM & Agentic

When to Fine-Tune: The Decision Framework

Should you fine-tune at all? A structured decision framework for prompt engineering vs RAG vs fine-tuning.

Last updated

Core7 min readFirst readHow LLMs Are BuiltThe Eval Stack

After this section you can

  • Place a problem on the optimisation ladder and pick the cheapest lever that fixes it
  • Answer the three questions that decide whether fine-tuning is worth considering
  • Spot the red flags that rule fine-tuning out before any budget is spent
44

When to Fine-Tune: The Decision Framework

Fine-tuning changes the model’s weights. It is the most expensive, slowest and least reversible way to change what a model does. Three questions tell you whether you have earned it.

Key idea

Match the lever to the problem. Missing facts go in the context (RAG, tools). Behaviour goes in the prompt first. Fine-tune only a narrow, high-volume task that prompting has measurably failed, and only once an eval can prove it helped.

The optimisation ladder: stop at the first rung that passes your eval
CLIMB ONLY WHEN THE RUNG BELOW HAS FAILED YOUR EVAL 1 · Prompt hours · change the input fixes behaviour and format 2 · RAG and tools days · change the context fixes missing or fresh facts 3 · Fine-tune weeks · change the weights narrow behaviour at volume 4 · Pre-train months · a new base model almost never your job Each rung up costs more, takes longer to iterate and is harder to undo. Most production problems are solved on rungs 1 and 2. cost · iteration time · lock-in

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium