Part 4 · 1 chapters · ~8 min

Fine-Tuning with LoRA

Full fine-tuning and its memory cost, low-rank adaptation (LoRA) and the parameter arithmetic, QLoRA with a 4-bit base model, preparing training data, evaluation before and after, serving many adapters on one base model, and when retrieval or prompting beats fine-tuning.

5

Adapters in code

code
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct", load_in_4bit=True)   # QLoRA-style base
cfg = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], lora_dropout=0.05)
model = get_peft_model(base, cfg)
model.print_trainable_parameters()      # a few tens of millions trainable out of ~8 billion

# data: a few thousand high-quality examples of the exact task, e.g. classifying support tickets
{"messages": [{"role": "user", "content": "Transfer debited but not received after 2 hours"},
              {"role": "assistant", "content": "{\"category\": \"pending_credit\", \"urgency\": \"high\"}"}]}

Serving adapters: the base model stays loaded once; many LoRA adapters (one per customer or task) can be swapped per request (vLLM and similar servers support multi-LoRA batching). Measure first: build the evaluation set (part 7) before fine-tuning, or you cannot tell whether it helped.

LoRA: TRAIN A SMALL UPDATE, FREEZE THE REST
a rank-r correction to each weight matrix
frozen W4,096 × 4,096 = 16.8M paramsW + B AA: r × 4,096trainableB: 4,096 × rtrainable, starts at 0r = 16: 131k params0.78% of W
swipe the figure sideways, or tap expand for full screen
1/4
the idea
Fine-tuning every weight of an 8B model needs memory for weights, gradients and optimiser state: well over 100 GB. LoRA freezes W and learns a low-rank update B·A instead.
freeze W, learn B·AHu et al. 2021