deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Exact brute-force search makes vector databases optional below 100k vectors

A dev.to benchmark on a 2-vCPU instance shows exact numpy search answering 100k-vector queries in 3.5 ms, with sizing rules for when an ANN index becomes worth it.

Exact brute-force search makes vector databases optional below 100k vectors

What was measured

Most retrieval-augmented generation tutorials begin by installing a vector database. A benchmark published on dev.to by the RecallRun account asks whether that dependency is actually necessary, and answers with timings rather than opinion: below roughly 100,000 chunks, a plain numpy matrix multiply is quick enough for nearly every RAG workload.

The test pits exact search — one matrix-vector product over normalized embeddings plus a partial sort — against faiss IndexFlatIP and against HNSW indexes built with hnswlib (M=16, ef_construction=200). Three common embedding dimensions were covered: 384 for MiniLM-class models, 768 for BERT-base-class encoders, and 1536, the default size of OpenAI text-embedding-3-small. The hardware was deliberately small — a 2-vCPU Intel Xeon at 2.1 GHz with 7 GB of RAM — and all latencies are medians over synthetic, normalized float32 vectors.

Exact search stays cheap for a long time

The exact-search numbers are the core of the argument. At 100,000 vectors of 384 dimensions, numpy returned the top 10 neighbours in 3.5 ms, while faiss took 8.3 ms. Scaling to 1 million vectors, exact search still completed in 62 ms at 384 dimensions and 114 ms at 768 dimensions. According to the dev.to post, numpy beat faiss by roughly 2x across the board for single queries because faiss is tuned for batches — a caveat the author flags for anyone serving batched requests.

Latency scaled linearly with the number of vectors times the dimension, which the author attributes to the operation being bound by memory bandwidth rather than compute. That makes RAM the real limit: 1 million vectors at 1,536 dimensions need about 6 GB of float32 before any metadata, which is why that configuration was skipped on the 7 GB test machine.

What HNSW buys, and what it costs

Approximate indexes are much faster per query — the measured HNSW lookups ran 10 to 100 times faster than exact search at these sizes — but the benchmark highlights two costs.

Build time is the first: constructing the 500,000-vector index took over 4.5 minutes on 2 cores, an expense repeated on every full re-index, for instance after switching embedding models.

Recall is the second, and it varies dramatically with data shape. On synthetic clustered data with 2,000 tight clusters, ef=64 returned 98.5 to 99.5 percent of the true top 10 at 100k vectors. On uniformly random directions, the same settings collapsed: 20.6 percent recall at 100k vectors and 384 dimensions, and just 5.7 percent at 500k. Real embeddings sit between these extremes, usually closer to clustered, so the author's advice is to measure recall on your own vectors before trusting default settings.

A sizing rule for practitioners

The benchmark puts the search step in context: a typical RAG request spends 1 to 5 seconds waiting for the LLM to generate, making a 3.5 ms lookup invisible. Even 60 to 110 ms for 1 million vectors matters only if you track p99 latency or push many queries per second through one machine. For scale: one PDF page is roughly 2 to 3 chunks, so 100,000 chunks corresponds to about 30,000 to 50,000 pages — more than many internal documentation-chat projects ever reach.

The resulting rule of thumb from the post: under 100k chunks, use numpy or unindexed pgvector, keeping vectors in a .npy file or Postgres; between 100k and 1M, exact search still works at low query rates, adding HNSW when p95 latency or throughput demands it; beyond 1M vectors or with many tenants, move to a proper ANN setup such as pgvector HNSW, Qdrant, Milvus or OpenSearch. The author also recommends keeping the vector store behind a single interface so the backend can be swapped later without touching the rest of the pipeline.

Why it matters

This is a concrete answer to a common premature-optimization problem. Exact search guarantees perfect recall for free, removing one variable when debugging poor RAG answers, and it defers an infrastructure dependency many teams adopt on day one. The caveats are stated plainly: the numbers come from synthetic vectors on one small machine. Exact-search latency transfers well because it does not depend on what the vectors mean, but HNSW recall absolutely does — which is precisely why the benchmark ends by telling readers to measure on their own embeddings rather than trust anyone's defaults.

  • #vector-databases
  • #rag
  • #embeddings
  • #benchmarks
  • #numpy