deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Tuning Qwen 27B on an RTX 3090 nearly doubles generation speed to 60 tokens/s

A dev.to write-up details how IQ3_S weights, a Q8 KV cache and four-token speculative decoding lifted Qwen coding generation from roughly 33 to about 60 tokens/s on a single 24GB card.

Tuning Qwen 27B on an RTX 3090 nearly doubles generation speed to 60 tokens/s

Nearly double the generation speed

A developer writing on dev.to reports that tuning a Qwen3.8-27B model on a single 24GB RTX 3090 raised generation speed during coding-agent tasks from roughly 33 tokens per second to about 60. The current preferred recipe combines IQ3_S quantized weights, an eight-bit KV cache and speculative decoding with at most four draft tokens, all inside a 128K total context window.

The speed figure covers generation only: total generated tokens divided by decode time, excluding prompt processing. Whole-task time also includes file access, test runs and tool waits, so a faster token rate does not guarantee faster delivery. The author notes one configuration that generated tokens faster yet needed about three extra minutes to hand over its code.

Quantization choices and their caveats

The original Q4_K_M weights occupied about 15.66 GiB; the selected ISTA-DASLab GSQ-RCO IQ3_S file, which bundles an MTP prediction head, is about 11.29 GiB. Smaller weights leave more GPU memory for caches and runtime buffers. Two settings are easy to confuse, the author cautions: IQ3 compresses the weights while the Q8 KV cache stores attention keys and values at eight bits, and they can be combined. The IQ3 filename also does not mean every tensor is exactly three bits.

With eight tasks repeated three times per configuration, a 128K window, a 32K output cap and xhigh reasoning, the recorded results were: original Q4_K_M with MTP off at 33.29 tokens/s and 23/24 delivered passes; GSQ-RCO IQ3_S with MTP3 and Q8 KV at 59.87 tokens/s and 21/24; and Bartowski IQ4_XS with a locally added MTP head at MTP2 and Q8 KV at 62.36 tokens/s and 22/24.

That is roughly 80% faster generation for the IQ3 bundle, but the improvement cannot be credited to IQ3 alone because weights and speculative decoding changed at the same time, and the comparison cohorts came from earlier runs rather than a contemporaneous randomized test. A later IQ4-versus-IQ3 run with a fresh random seed kept IQ3 in place: IQ4 generated 7.8% faster but delivered 6/8 tasks versus 7/8, with total task time about 2.1% longer, and one IQ4 answer contained a genuine overflow-handling defect. Other tweaks, including fewer CPU threads, altered CUDA wait behavior and added ngram speculation, produced no consistent benefit.

Picking a speculative decoding depth

MTP drafting proposes several upcoming tokens for the main model to verify, and accepted drafts reduce the model's token-by-token work. Deeper drafting costs compute because rejected positions waste effort; the depth number is a cap on proposed tokens, not guaranteed acceptances. Running each of the eight coding tasks at depths 2 through 6, with IQ3_S, Q8 KV, a 128K window and a 64K output cap fixed, put depth 4 on top at 61.31 tokens/s versus 57.98 at depth 2, about 5.7% faster, with depths 5 and 6 slower still. A short-input screen had favored depth 2, which the author cites as a reason to prioritize tests that resemble real workloads. The result reflects one pass over eight familiar tasks on one machine, not a universal optimum.

A failing task pointed at the test, not the model

A repeatedly failing Ledger task had two separate causes. One response genuinely hit the 32K output cap; raising it to 64K produced a 40,612-token answer that passed all checks, while the task grew from roughly 628 to 1,140 seconds. Separately, five Ledger answers in the depth comparison failed without exhausting any budget because the test's file and output mocks lacked normal interfaces, causing valid solutions to be rejected. After correcting those interfaces, without relaxing the requirement that read-error handling be present, the unchanged answers passed 13/13 checks, and regrading lifted the MTP4 cohort from 7/8 to 8/8 without generating anything new. The author now treats repeated failures as a prompt to inspect truncation, code defects and whether a test actually exercises what it claims to measure.

Long context: 220K works, 256K does not with MTP

Because input and output share one window, a 170K input budget plus 50K of output needs a 220K total window. According to the post, IQ3_S with Q8 KV and MTP4 processed 179,532 input tokens followed by short output under a 220K window, peaking near 22.80 GiB of GPU memory. At 256K, MTP4 failed allocation during startup; disabling MTP and reducing batch sizes allowed the larger configuration to start. Full coding validation at 220K remains incomplete. The author also discloses that the write-up was drafted with AI assistance from recorded local experiments and then reviewed.

Why it matters

Running a 27B-class model at usable speeds on one consumer GPU is the difference between a local coding agent that feels responsive and one that does not. The post offers a concrete, reproducible starting point for practitioners: a specific quantization file, KV cache precision and draft depth, with numbers attached and honest caveats about what the comparisons can and cannot prove. Just as valuable is the methodology, which separates generation speed from delivery time, treats short screens as candidate filters rather than verdicts, and demonstrates how a defective benchmark can silently misrank configurations. Anyone evaluating local models on their own tasks will recognize how easily those failure modes appear.

  • #local-llm
  • #quantization
  • #qwen
  • #speculative-decoding
  • #gpu

Related posts