deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

FP8 quantization made vLLM look 47% cheaper by silently breaking Qwen2.5 output

A dev.to benchmark on an AMD MI300X found vLLM's on-the-fly FP8 flag made Qwen2.5-72B look 47% cheaper because it emitted garbage tokens at full speed; pre-quantized checkpoints cut real costs by about a third.

FP8 quantization made vLLM look 47% cheaper by silently breaking Qwen2.5 output

A saving that was too clean

A developer benchmarking inference costs on a single AMD MI300X got what looked like a decisive win: enabling FP8 quantization in vLLM cut the measured cost of serving Qwen2.5-72B-Instruct by 47%. The measurement was real, but the model behind it was broken. According to a firsthand account on dev.to, the FP8 server was emitting streams of exclamation marks instead of answers and hitting the output-token cap on every request, so the apparent saving came from generating nonsense tokens at speed.

The setup

The author had one hour on the GPU and wanted a simple answer: how much a single output token costs. The test ran vLLM's ROCm build against Qwen2.5 at 7B, 32B and 72B, with 32 concurrent requests, a 256-token output cap and the provider's $2.99-per-GPU-hour list rate; every figure is wall-clock time multiplied by that rate and divided by tokens produced. Baselines in BF16, four repeats each, came out at $0.227 per million output tokens for the 7B model, $0.77 for the 32B and $1.67 for the 72B, with the 72B repeats agreeing to within 0.4% — a stable machine.

The trap

Since FP8 stores weights in half the bytes, the next step was to flip vLLM's on-the-fly --quantization fp8 flag. Over three runs each, the 32B fell from $0.77 to $0.449 per million tokens (a 41% cut) and the 72B from $1.67 to $0.872 (47%).

The giveaway sat in the per-block detail rather than the headline. At temperature 0 with the same 32 prompts, BF16 32B wrote roughly 6,900 output tokens per block, while FP8 wrote exactly 8,192 every time — 32 requests times the 256-token cap, meaning every request ran to the limit. Asked to summarize the causes of World War I in about 80 words, BF16 returned an 89-token answer; the FP8 server printed exclamation marks for all 256 tokens. The cost number was pricing the wrong behavior.

The numbers that hold up

Repeating the comparison with checkpoints quantized ahead of time — RedHatAI's Qwen2.5 FP8-dynamic releases — and verifying answers first produced a smaller but usable result. Output length matched BF16 within about 1% on the 72B (6,361 versus 6,328 tokens per block), spot-check answers were correct, and costs fell 18% on the 7B ($0.227 to $0.185 per million tokens) and 32% on the 72B ($1.67 to $1.13), with noise bounds under 3% and under 1% respectively.

Caveats the author flags

The post is candid about its limits. It covers one session and one workload — short prompts, 256 output tokens, concurrency 32 — so long-context or long-output traffic may behave differently. The 72B FP8 server ran on the VM's second GPU while the baseline used the first; same card model, not literally the same card. The broken on-the-fly 72B was diagnosed by output length alone, without reading its answers. On-the-fly FP8 on the 7B was actually fine, with six of six answers correct at an 18% saving. And six prompts are a sanity check, not an accuracy benchmark — if FP8 matters to a product, the author argues, run real evals. It remains unclear whether the failure is specific to this ROCm image or to larger Qwen2.5 models.

Why it matters

Cost per token is the default metric for comparing serving configurations, and this is a concrete case where it could not tell a cheaper model from a broken one; the 47% figure would have sailed through a naive benchmark. The author's countermeasures are cheap. Watch average output length and maxed-out requests, because that data already sits in benchmark output. Read a handful of fixed prompts at temperature 0. Repeat the baseline so you know your own noise floor before calling a small delta a win. For teams running vLLM with quantization flags, especially on ROCm hardware, the practical takeaway is to validate output quality before trusting a cost delta, and to prefer pre-quantized checkpoints over on-the-fly conversion until the failure mode is understood. The measurements were made with Throttle, an open-source CLI whose next release, the author says, will report OUTPUT CHANGED instead of CHEAPER when output length shifts between runs.

  • #vllm
  • #fp8
  • #quantization
  • #gpu
  • #inference