deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Qdrant Turbo4 benchmark: 9x storage saving costs up to 3 points of recall versus TurboQuant

A reproducible benchmark on 500,000 Wikipedia abstracts finds Qdrant 1.19's Turbo4 cuts storage 9x versus TurboQuant but gives up between 0.5 and 3 points of recall@10, with the gap shrinking as embedding dimensions grow.

Qdrant Turbo4 benchmark: 9x storage saving costs up to 3 points of recall versus TurboQuant

What the benchmark measured

A reproducible benchmark published on dev.to puts hard numbers on the trade-off behind Turbo4, the storage-only vector datatype introduced in Qdrant 1.19. Turbo4 keeps just a 4-bit compressed version of each vector and discards the float32 originals, shrinking a collection's storage footprint by a deterministic 9x. The catch is structural: with no full-precision copy left on disk, there is nothing to rescore search candidates against, so the distance computed on the 4-bit representation becomes the final score. The safety net that traditionally corrected quantization errors is gone by design.

The baseline is the configuration most production users ran on Qdrant 1.18: TurboQuant, a 4-bit quantization layer built on a Hadamard rotation technique credited to Google Research that spreads information evenly across coordinates before compression. TurboQuant is a dual-copy architecture — a 4-bit copy in RAM for fast scanning across the HNSW graph, plus the original float32 vectors on disk for a final rescoring pass. According to the dev.to write-up, that setup reached 98.2% recall on a 1-million-document OpenAI collection, but the full-precision copies are expensive: at 1,536 dimensions the raw coordinates alone take 6.9 GB, and a 10-million-document expansion would demand over 65 GB just to keep the rescoring option alive.

How the test was run

The author evaluated five embedding models spanning 384 to 3,072 dimensions over 500,000 Wikipedia abstracts — deliberately messy, varied data rather than synthetic sets that would flatter the results. All runs shared the same hardware, HNSW settings (m=16, ef_construct=128), cosine distance and query sets; the only variable was the storage datatype. Ground truth came from an exhaustive brute-force search, and recall was averaged over 1,000 queries per configuration.

Recall: dimensionality decides

At top-10 retrieval, the setting most RAG pipelines use, Turbo4 recovered between 96.1% and 98.9% of the correct results depending on the model, leaving a gap of roughly half a percentage point to nearly three points against TurboQuant with rescoring enabled.

The embedding dimension turned out to be the deciding variable. Rounding errors in a 384-dimension vector matter proportionally more, while in a 3,072-dimension vector they largely cancel out — an effect the benchmark attributes to the law of large numbers. In practical terms, the 384-dimension gap means roughly one in 37 queries loses a relevant document; at 1,536 dimensions and above that falls to about one in 100, and a reranker behind the retrieval step tends to absorb even that residual difference. At deeper retrieval (k=100) the gap widens slightly, again mostly on shorter vectors.

Storage and throughput

The storage ratio is arithmetic rather than measurement: TurboQuant stores 36 bits per coordinate — the 32-bit original plus the 4-bit compressed copy — while Turbo4 stores only the 4-bit value, which produces the 9x reduction. A 10-million-vector collection using OpenAI's largest model needs about 130 GB under TurboQuant versus roughly 14.4 GB under Turbo4, a difference the dev.to post frames as the choice between a dedicated high-memory server and an ordinary cloud instance.

Throughput also moved in Turbo4's favour, for a mechanical reason. TurboQuant's rescoring pass has to pull original vectors from disk for its top candidates, with oversampling fetching extra candidates before the re-check (for example, 20 candidates for a limit of 10). Turbo4 never reads originals because none exist, and the benchmark reports improved search speed on that basis.

Why it matters

The benchmark turns a vendor value proposition into an engineering decision with quantified sides. For pipelines built on high-dimensional embeddings — 1,536 dimensions and up — with a reranker downstream, the recall cost of Turbo4 sits around one lost document per hundred queries, which many teams will trade willingly for a 9x storage cut and a simpler, faster single-copy search path. For short 384-dimension vectors or deep retrieval at k=100, the accuracy loss is material enough that the dual-copy TurboQuant setup and its disk overhead may remain the responsible default. More broadly, it makes explicit the architectural choice now facing vector-database users: compressed-only storage as a first-class datatype, versus quantization as an accelerator layered on top of full-precision data.

  • #qdrant
  • #vector-databases
  • #quantization
  • #benchmarks
  • #embeddings

Related posts