deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

AWS S3 Vectors adds metadata pre-filtering to stop silent result loss in multi-tenant RAG

AWS S3 Vectors now resolves metadata filters before similarity search. The old mode silently returned fewer or worse results with selective filters — a common pitfall in multi-tenant RAG pipelines.

AWS S3 Vectors adds metadata pre-filtering to stop silent result loss in multi-tenant RAG

On September 30, 2026, AWS shipped metadata pre-filtering for S3 Vectors, fixing a failure mode that was easy to miss: with a selective filter, such as scoping a query to one tenant, the service could return fewer results than requested, or results that were not the true nearest neighbors, without raising any error. According to a dev.to write-up, the update introduces two index modes, a new prefix operator and a per-query condition cap, at no additional cost.

How the old mode lost results

Vector search at scale is approximate: the engine explores a bounded pool of promising candidates rather than comparing a query against every vector. In what AWS now calls CLASSIC mode, metadata filters were evaluated while that exploration ran. When a filter matches a large share of the corpus, most candidates survive and results look correct. When a filter is highly selective, most explored candidates get discarded, the search budget runs out, and the caller receives fewer than the requested number of results — or a full set that is not the best match within the filtered subset. AWS's own illustration, cited in the post, is 8 million support tickets where a single customer owns 400: pre-filtering searches those 400 rather than candidates drawn from all 8 million.

The dev.to article points to an AWS re:Post measurement on one million synthetic vectors where recall@10 dropped from 0.86 unfiltered to 0.83 at 10% filter selectivity, 0.42 at 1% and 0.15 at 0.01%. The common workaround was splitting data into separate indexes by low-cardinality fields such as tenant, language or region, which reportedly recovered 15–31 recall points at the cost of operating many more indexes. The article stresses a related point: result count was never a quality signal. Receiving 10 results in CLASSIC mode never guaranteed they were the best 10.

ENHANCED mode and the migration path

The new ENHANCED mode resolves the filter first, then runs similarity search over the matching subset. AWS says this returns up to 5x more matching vectors when the filter is selective. Vector buckets created on or after September 30, 2026 are ENHANCED-only. Buckets created earlier stay on CLASSIC until an explicit UpdateIndexMode call, and no re-ingestion is needed; reverting to CLASSIC is possible through the CLI, SDK or API, but not the console. Teams can also trial the behavior per query with queryMode set to ENHANCED before switching an index, and the post includes a brute-force recall harness — comparing indexed results against a ground-truth top-K computed directly from the vectors — for measuring both modes on a test copy of real data.

Filter changes and limits

Filter syntax is otherwise unchanged, with one addition: a $startsWith string-prefix operator, available only on ENHANCED indexes, which suits collections organised by path. ENHANCED indexes also cap filters at 100 conditions, counted per value — an in-list with three regions consumes three conditions. The article warns teams that build dynamic filters from long permission lists, for example every project a user can access, to audit those filters before migrating, because exceeding the cap produces a validation error.

Per-vector metadata limits are unchanged: up to 40 KB of metadata per vector, of which up to 2 KB is filterable, with at most 50 keys and up to 10 non-filterable keys fixed at index creation. The recommendation is to keep filter fields short and store chunk text in non-filterable metadata, a pattern Bedrock Knowledge Bases already follows with reserved keys.

A permissions gotcha

For teams wiring this into a retrieval service, the post flags an IAM detail: the QueryVectors permission alone only works for queries without filters and without metadata returned. Adding either requires the GetVectors permission as well, or the call fails with a 403.

Why it matters

Silent quality degradation is the hardest failure to catch in RAG systems: nothing errors, the model just answers with less context than it should have had, and multi-tenant setups with small tenants are exactly where CLASSIC mode degrades most. Teams running RAG on S3 Vectors should check their index mode, measure recall on their own data before and after switching, verify that dynamically generated filters fit the 100-condition cap, and update IAM policies. Since the feature carries no extra charge and is available in every commercial Region offering S3 Vectors plus the China Regions, the main cost of adopting it is validation rather than licensing.

  • #aws
  • #s3-vectors
  • #vector-search
  • #rag
  • #cloud

Related posts