· via dev.to (home feed)
Guide: run 100B+ MoE models on a single 24GB RTX 4090 with llama.cpp expert offloading
A dev.to guide explains how llama.cpp expert offloading runs 125B-parameter MoE models on a 24GB RTX 4090 by keeping attention on the GPU and moving expert weights into system RAM.

How MoE gets past the VRAM wall
A practical guide published on dev.to, inspired by a Hacker News post that gathered more than 600 points, walks through running a 125-billion-parameter Mixture of Experts model on a single RTX 4090. That sounds implausible at first: the card has 24GB of VRAM, while a 125B model at 4-bit quantization occupies roughly 70-75GB. The solution, according to the author, is not better compression but smarter placement of tensors across the GPU and system RAM.
In a dense model, every token passes through every parameter, so any weights that do not fit in VRAM must be computed on the CPU — and throughput collapses because system RAM bandwidth is far lower than VRAM bandwidth. MoE models break that assumption. A router in each layer selects only a few experts per token, commonly 8 out of 128 or 256, which means a 125B model may have just 10-15B active parameters per token. The components touched on every step — attention, the KV cache, embeddings, shared experts and the router — are comparatively small, while the expert FFN weights are enormous but only partially read. The guide's core idea: keep the always-used parts on the GPU and push the experts into system RAM.
Hardware that actually matters
The test setup combined an RTX 4090 with a Ryzen 9 7950X, 128GB of DDR5-5600 arranged as two 64GB sticks in dual channel, and a Gen4 NVMe drive for fast model loading. The author suggests 96GB as a realistic minimum for a roughly 120B model at Q4.
One counterintuitive finding: with expert offloading, RAM bandwidth matters more than CPU core count. Populating four DIMMs on a consumer platform typically drags the memory bus down to 4800 or lower, so two large, fast sticks outperform four smaller ones.
On the software side, the guide builds llama.cpp with CUDA enabled and sets the compute architecture flag to 89 for Ada-generation cards, which speeds up compilation. For weights, it recommends pre-quantized GGUF files using dynamic or imatrix schemes, noting that IQ4_XS lands close to Q4_K_M in quality while being 10-15% smaller.
The flags that do the work
llama.cpp offers two offloading mechanisms. The simple one is --n-cpu-moe N, which keeps the experts of the first N layers on the CPU. The flexible one is -ot, a regex-based override that maps specific tensor patterns to specific devices for fine-grained tuning. Both are combined with -ngl 99 — offload everything — after which the expert rules pull the heavy tensors back off the GPU, leaving attention and shared weights behind.
Supporting flags matter too: -fa on enables Flash Attention and shrinks the KV cache, q8_0 KV cache quantization roughly halves cache memory with negligible quality loss, and -t should match physical rather than logical cores, since over-threading causes RAM bandwidth contention. Rule ordering for -ot is significant: earlier rules take precedence, so rules that move tensors onto the GPU belong above the catch-all rule that pushes experts to the CPU.
The recommended tuning loop starts with --n-cpu-moe set to the total layer count, watches nvidia-smi, and lowers N step by step until VRAM usage sits around 22-23GB, keeping 1-2GB of headroom to avoid out-of-memory errors when the context fills up.
Measure instead of trusting benchmarks
The author explicitly warns against copying throughput numbers from the internet, including the original Hacker News thread, because results depend heavily on RAM, context length and quantization type. Instead, the guide pairs llama-bench with a small streaming script that hits the OpenAI-compatible endpoint and records time-to-first-token and decode tokens per second.
Prefill emerges as the weak point of expert offloading, since an entire prompt must pass through CPU-resident experts. For long-prompt RAG workloads, the guide suggests raising batch sizes with -b and -ub, for example -ub 2048, so llama.cpp can shift more of that prefill expert computation to the GPU.
Pitfalls to avoid
Three failure modes are highlighted. If RAM is insufficient, the operating system swaps or mmaps experts from disk and speed falls below 1 token per second — watch free -h during inference. On Windows, WSL2 claims only half of system RAM by default, so a .wslconfig entry raising the limit is needed. And very long contexts are costly even with a quantized cache, so the advice is to start at 32K and increase only when necessary.
Why it matters
Expert offloading reframes the hardware question for large-model inference: instead of an 80GB-class GPU, a mid-range card plus abundant system RAM is enough, and DDR5 is far cheaper per gigabyte than VRAM. For developers who want to run 100B-class models locally — keeping internal code or customer data off cloud GPUs — this turns a previously out-of-reach workload into a desktop setup. The caveat is that the technique only applies to MoE architectures; dense models gain nothing from it.
- #llama-cpp
- #local-llm
- #mixture-of-experts
- #gpu
- #quantization