deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

The 125B at 100 tok/s on one RTX 4090 claim is real, but the fine print is 64GB of RAM

A repo called Strata hit Hacker News claiming a 125B model runs at 100 tok/s on one RTX 4090. Independent numbers back the speed — if you bring 64GB+ of system RAM, and ignore viral retellings that garble the units.

The 125B at 100 tok/s on one RTX 4090 claim is real, but the fine print is 64GB of RAM

A project called Strata reached 788 points on Hacker News with an improbable headline: Qwen 3.8 Flash Next, a 125-billion-parameter model, decoding at 100 tokens per second on a single RTX 4090. A fact-check post on dev.to by Ashraf Chowdury concludes the figure is genuine — but only with a workstation-class memory setup, a specific quantization, and workloads where speculative decoding happens to shine.

What actually makes it possible

A 24 GB card still cannot hold a 125B model; what changed is the model's architecture. According to the dev.to breakdown, Qwen 3.8 Flash Next is a sparse mixture-of-experts design with roughly 180B total parameters — about 125B in the main model, 51B in n-gram embeddings and 4B in a multi-token-prediction head — yet only around 6B active per token. Each token routes to 10 of 512 experts plus a shared expert, across 48 layers, with a 262k native context.

Strata's approach is conceptually simple: keep attention, shared weights and the most-used experts in VRAM, hold the remaining experts in system RAM, and fetch them over PCIe as needed. The n-gram table helps because it is a lookup rather than a matrix multiplication — one deployment reportedly kept a 47.7 GiB quantized table on an SSD with no throughput loss. The multi-token-prediction head adds speculative decoding: draft several tokens, verify them in a single pass.

The fine print: system RAM

Strata's own requirements, as summarized in the fact-check, start at 32 GB of RAM for a coder variant with half the experts removed (about 91% of full SWE-bench Verified performance), rising to 48 GB for an IQ2_XS build, 64 GB for IQ3_S — the recommended balance — and 96 GB or more for UD-IQ4_XS.

The headline number comes from an independent 4090 measurement repo, cited in the post, which ran on a Ryzen 7900 with 192 GB of DDR5 and the GPU capped at 280 W. It recorded 110.8 tok/s of text decode at IQ3_S with low reasoning, 97.65 tok/s with vision enabled, and roughly 4,750 tok/s of cold prefill at 32K context. On a 30-case executable DevOps suite the model scored 28/30, while a Q4_K_XL build reached 30/30 at low reasoning but fell to 27/30 with reasoning off. The repo's author is explicit about the caveats: all experts and the n-gram table must stay RAM-resident, quality results are single-seed, and only single-request serving is optimized.

Speed depends on what you generate

Speculative decoding acceptance varies sharply by workload: about 0.59 on prose, 0.88 on code and 0.93 on structured output, per the dev.to post. On a Strix Halo machine the same model moved from roughly 22 tok/s on prose to 82 tok/s on code. The widely shared 100 tok/s is effectively a code-and-structured-output figure; conversational prose will be noticeably slower.

A viral retelling garbles the story

The two dev.to posts disagree on basic facts. A second article's headline claims the RTX 4090 reaches 100 trillion tokens per second — a unit error contradicted by its own results table, which lists roughly 100 tokens per second. That version also asserts the quantized model fits in about 19 GB of VRAM with no CPU offload, credits a separate 0.5B drafter model with about 70% acceptance, and reports 350 W power draw. All of this conflicts with the independent measurement, which depends on streaming from 192 GB of system RAM with the card capped at 280 W. The drift is a useful reminder of how quickly benchmark claims mutate as they travel.

Quality, licensing and quirks

On the nine benchmarks it shares with Qwen3.8-27B, Flash-Next wins all of them, averaging +4.19 points, with the largest gap in agentic work: DeepSWE climbs from 42.2 to 58.7, according to the fact-check. That agentic-coding advantage, rather than raw speed, is the practical reason to run it locally.

The license is Qwen Community License 1.0, not Apache 2.0: self-hosting and commercial use are permitted, but products above 100M monthly users or $20M in revenue must display the model name, and serving the model as a service requires separate licensing. llama.cpp users need a build from Aug 27, 2026 or later; older builds reject the architecture. Day-to-day quirks include a 1–3 minute freeze on first startup while 35–55 GB loads into RAM, about a minute to first token for a 30K-token prompt on a 12 GB card, and parallel requests disabled by default. Hacker News commenters also pointed out that a 4090 launched at $1,600, which stretches the phrase consumer hardware.

Why it matters

The real lesson is architectural. Sparse expert routing, SSD-resident lookup tables and speculative decoding together shift the bottleneck for local inference from VRAM capacity to system RAM capacity — and system memory is far cheaper per gigabyte. That reframes what one desktop can do: agentic coding that clearly outperforms a 27B dense alternative, with no cloud API, provided you spec the memory. The episode is also a study in hype mechanics: one repo's carefully caveated measurement became a Hacker News headline, then a viral post that misstated both the units and the mechanism. Anyone evaluating the claim should benchmark their own workload, because prose, code and structured output produce three very different token-per-second figures from the same model.

  • #local-llm
  • #mixture-of-experts
  • #quantization
  • #rtx-4090
  • #open-source

Related posts