· via dev.to (home feed)
QAT Gemma 4 repack serves 12B model at 675 tokens/sec on one TPU v5e
A dev.to guide shows how to repack Google's QAT Gemma 4 weights into vLLM-servable formats and serve the 12B model at 675 output tokens per second on a single TPU v5e chip, matching bf16 quality.

A guide published on dev.to lays out how to repack Google's quantization-aware-trained (QAT) Gemma 4 weights into formats vLLM can serve, then run them on a single Google Cloud TPU v5e chip. The headline result: the 12B model, stored as int8 weights with int4 embedding tables, occupies 11.31 GiB and generates 675 output tokens per second at 16 concurrent requests while scoring on par with its bf16 reference across a 3,880-record evaluation suite. Every script, log and per-record output is committed to a public repository.
Why repacking is needed
A single v5e chip, provisioned as v5litepod-1, has 15.75 GiB of HBM. According to the guide, Gemma 4 E4B in bf16 already needs 14.9 GiB and the 12B model 22.4 GiB, so anything above the smallest E2B size must be quantized to fit. Google trains 4-bit variants of every Gemma 4 size and publishes the trained values as bf16 checkpoints tagged qat-q4_0-unquantized, in which each group of 32 weights already lies on a 16-level grid. The repack converts those values into servable formats without rounding them a second time, preserving the positions QAT training produced.
Four formats and three patches
The guide builds four variants: q4w4a16 (int4 weights on the QAT grid with 16-bit activations), q4w4a16emb4 (the same with int4 vocabulary tables), w8a8 (per-channel int8 weights with per-token int8 activations) and w8a8emb4 (int8 weights plus int4 vocabulary tables). Because v5e multiplies int8 by int8 in hardware, the int8 builds are the quick options on this chip. Serving them required three additions to vLLM's tpu_inference backend: an int8 W8A8 method on the JAX path, embedding tables kept compressed in HBM and expanded only for the rows each step touches, and an int4 lm_head.
How it runs
Prerequisites are a Cloud project with v5e flex-start quota in us-west4-a, the gcloud CLI, a Cloud Storage bucket and a Hugging Face token stored in Secret Manager. Each experiment is a flex-start queued resource that boots, patches a pinned vLLM image, serves a list of builds in sequence and uploads the results, capped at four hours; the 12B build reported ready after 405 seconds. Finished checkpoints are published on Hugging Face under the xbill9 account, and the code sits in the gemma4-dev repository on GitHub.
What the numbers show
On the 3,880-record suite from Bespoke Labs, scored record-for-record against bf16, every build through 12B lands within the reported error ranges of bf16; the 26B mixture-of-experts repack reads 1.1 points lower. Throughput at 1, 4 and 16 parallel requests: the 12B int8 build does 57, 219 and 675 tokens per second, the 12B 4-bit build 35, 127 and 407, and the fastest E2B build reaches 3,086. The 12B int8 model leaves the chip's remaining memory to a KV cache of 9,728 tokens, enough for 2.38x concurrency at 4,096 tokens per request. The 26B repack serves from 13.58 GiB with a 2,176-token cache at 201 tokens per second. bf16 references for E4B, 12B and 26B ran on v6e chips because those sizes do not fit a v5e.
Compared with Google's own qat-w4a16-ct exports, which re-round every weight group, the repacks score 2.4 points higher at E2B, 1.3 at E4B and 0.6 at 12B, at identical speed: roughly 1,910 tokens per second for E2B either way at 16 requests.
Math and tool calling
The 12B int8 build scores 0.964 on GSM8K (all 1,319 problems, zero-shot chain of thought, greedy decoding) and 0.955 on BFCL v3 simple (400 records, one tool offered). At E2B, the repacks trail bf16 by 0.9 to 2.0 points on GSM8K and stay within 1.3 points on tool calling; because the QAT weights stored at bf16 themselves trail by 1.3 points on GSM8K, repacking adds at most 0.8.
Why it matters
A single v5e chip is among the cheapest ways to rent TPU capacity, and the guide demonstrates a 12B model running on one at bf16-level quality with hundreds of tokens per second of throughput, a practical template for cost-sensitive inference and development. Two broader lessons stand out. First, QAT-trained weights carry value beyond their intended 4-bit format: int8 derived from them beats int8 rounded straight from bf16 by 1.5 points on the suite and 7.2 points on BFCL, where the bf16-rounded build wrapped string arguments in surplus quotation marks in 16 of 400 records. Second, export choices such as re-rounding leave measurable accuracy on the table with no speed benefit in return.
- #gemma
- #quantization
- #tpu
- #vllm
- #inference