1:1 mentoring with Big Tech AI engineers
LLM & Agentic

LoRA, QLoRA & PEFT Methods

How LoRA works inside transformer layers, QLoRA for memory-efficient training, and the full PEFT method comparison with code examples.

Last updated

Production12 min readFirst readFine-Tuning Data PreparationHow LLMs Are Built

After this section you can

  • Compute LoRA’s trainable parameters and explain what B·A, rank and α/r do
  • Estimate GPU memory for full fine-tuning, LoRA and QLoRA and pick the one that fits
  • Choose how to ship an adapter: merged, swapped or served alongside many others
47

LoRA, QLoRA & PEFT Methods

Full fine-tuning updates every weight and needs roughly 16 bytes of GPU memory per parameter. LoRA freezes the model and trains a tiny low-rank update; QLoRA also shrinks the frozen base to 4 bits.

Key idea

LoRA keeps every original weight frozen and learns each weight change as the product of two thin matrices, ΔW = B·A. Only A and B get gradients and optimizer state, so a 7B model that needs over 100 GB to fully fine-tune trains on a single 24 GB GPU.

LoRA: freeze the big matrix, train a thin low-rank detour beside it
FROZEN · READ ON EVERY PASS, NEVER WRITTEN x input, k = 4,096 W · frozen d × k = 4,096 × 4,096 16.8M weights · no gradient ADAPTER · THE ONLY THING THAT TRAINS A · down r × k = 16 × 4,096 random init B · up d × r = 4,096 × 16 starts at zero × α / r e.g. 32 / 16 = 2 + h same shape h = W x + (α/r) · B A x Trainable in this matrix: r × (d + k) = 131,072 of 16,777,216 (0.78%). After training, fold (α/r)·BA into W.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium