LoRA, QLoRA & PEFT Methods
How LoRA works inside transformer layers, QLoRA for memory-efficient training, and the full PEFT method comparison with code examples.
Last updated
After this section you can
- Compute LoRA’s trainable parameters and explain what B·A, rank and α/r do
- Estimate GPU memory for full fine-tuning, LoRA and QLoRA and pick the one that fits
- Choose how to ship an adapter: merged, swapped or served alongside many others
LoRA, QLoRA & PEFT Methods
Full fine-tuning updates every weight and needs roughly 16 bytes of GPU memory per parameter. LoRA freezes the model and trains a tiny low-rank update; QLoRA also shrinks the frozen base to 4 bits.
LoRA keeps every original weight frozen and learns each weight change as the product of two thin matrices, ΔW = B·A. Only A and B get gradients and optimizer state, so a 7B model that needs over 100 GB to fully fine-tune trains on a single 24 GB GPU.