· via dev.to (home feed)
Why fine-tuning a 7B model takes 112 GB when the model itself is only 14 GB
Mixed-precision training with Adam costs about 16 bytes per parameter, so a 7B model needs 112 GB before activations. A dev.to breakdown explains where it goes and how LoRA and QLoRA cut it.

A post on dev.to walks through the arithmetic of fine-tuning memory and lands on a figure that surprises many practitioners: training a 7-billion-parameter model with mixed precision and the Adam optimizer takes roughly 112 GB of memory, of which the model itself accounts for only 14 GB.
Where the other 98 GB goes
The accounting, as the post explains, traces back to the ZeRO paper. In mixed-precision training with Adam, every parameter carries 16 bytes of state:
- 2 bytes for the fp16 weight
- 2 bytes for the fp16 gradient
- 12 bytes of optimizer state: an fp32 master copy of the weight plus Adam's two moment buffers, 4 bytes each
Multiply 16 bytes by 7 billion parameters and you reach 112 GB — and that is before a single activation has been stored. The optimizer state alone is 84 GB, six times the size of the weights. The post's central point is that training memory is driven less by the model than by the state required to update it.
LoRA removes the optimizer bill
LoRA, from Hu et al., attacks the dominant cost. Since gradients and optimizer state make up 14 of the 16 bytes per parameter, the strategy is to have far fewer parameters that need them. LoRA freezes every pretrained weight and trains small low-rank matrices injected into each layer instead. The frozen base still occupies memory at fp16 — 14 GB for a 7B model — but it accumulates no gradients and no optimizer state. Only the adapters do.
The dev.to post works through the numbers for a Llama-style 7B model with 32 layers and a hidden size of 4096, placing rank-8 adapters on the query and value projections. Each adapter holds 8 × (4096 + 4096) = 65,536 values, which comes to about 4.2 million trainable parameters across the network — roughly 0.06 percent of the model.
Hu et al. report that, compared with GPT-3 175B fine-tuned with Adam, LoRA cuts the number of trainable parameters by a factor of 10,000 and GPU memory requirements by three, while matching or exceeding full fine-tuning quality on the models they tested. Because the low-rank update can be merged into the base weight after training, it adds no inference latency — the property that distinguishes it from earlier adapter designs, which left extra layers permanently in the forward pass.
QLoRA shrinks what LoRA cannot avoid
LoRA still leaves the frozen base sitting at 14 GB in fp16. QLoRA, from Dettmers et al., targets exactly that by storing the base model in 4-bit precision, bringing it down to about 3.5 GB, with a little extra for quantization constants. The paper's double quantization technique trims that overhead from 0.5 bits to roughly 0.127 bits per parameter.
Gradients still flow through the 4-bit base into LoRA adapters kept at higher precision. The headline result: enough memory saved to fine-tune a 65B model on a single 48 GB GPU while preserving full 16-bit fine-tuning task performance. The trade-off is speed — 4-bit weights must be dequantized before any arithmetic, on every pass, so QLoRA runs slower than plain LoRA.
Three methods, one memory budget
The post frames all three approaches as variations on a single decision: which parts of the 16-byte-per-parameter bill each method eliminates. Full fine-tuning pays all 16 bytes. LoRA stops paying the 14 bytes of gradients and optimizer state on the base model while keeping the 2-byte fp16 copy. QLoRA compresses that remaining 2 bytes to roughly half a byte.
Why it matters
The 14 GB figure people quote for a 7B model is an inference number; training the same model with Adam in mixed precision costs eight times as much in weights, gradients and optimizer state alone, before activations are counted. Understanding that accounting demystifies why LoRA and QLoRA exist and what each actually optimizes — LoRA removes the per-parameter bookkeeping, QLoRA shrinks the frozen weights. For practitioners choosing between them, the decision reduces to hardware limits and how far the target task sits from the model's original training distribution, with QLoRA trading training speed for the ability to fit large models onto small GPUs.
- #fine-tuning
- #lora
- #qlora
- #large-language-models
- #gpu-memory