DEV Community

Cover image for S3 Vectors Now Filters Before It Searches: Metadata Pre-Filtering for Multi-Tenant RAG
Matias Martinez
Matias Martinez

Posted on Originally published at blog.cloudacademy.ar

S3 Vectors Now Filters Before It Searches: Metadata Pre-Filtering for Multi-Tenant RAG

If you run RAG on Amazon S3 Vectors with metadata filters, you may have been losing results without knowing it. Ask for the 10 nearest chunks for one tenant and you might get 3 back, or 10 that aren't actually the closest ones in that tenant's data. No error, no alarm, just a model answering with less context than it should have had.

On September 30, 2026, AWS shipped metadata pre-filtering for S3 Vectors. Indexes now come in two modes, and the new one evaluates your filter before the similarity search instead of during it. AWS says that returns up to 5x more matching vectors when the filter is selective. It also adds a $startsWith prefix operator, sets a 100-condition limit per query, and costs nothing extra.

Here's what changed, why the old behavior dropped results, how to measure the difference on your own index, and what a multi-tenant retrieval service on Amazon EKS looks like with it.

Diagram comparing CLASSIC and ENHANCED query processing in S3 Vectors
CLASSIC validates the filter while searching; ENHANCED narrows the set first. (Diagram labels in Spanish, from the original article.)

TL;DR

  • S3 Vectors indexes now have a mode: CLASSIC (filter evaluated in tandem with the approximate search) or ENHANCED (filter resolved first, then similarity search over the matching subset).
  • With selective filters, ENHANCED returns up to 5x more matching vectors, per AWS.
  • Vector buckets created on or after Sep 30, 2026 are ENHANCED-only. Older buckets stay CLASSIC until you run UpdateIndexMode; no re-ingestion needed.
  • You can test per query with queryMode=ENHANCED before switching the index.
  • New $startsWith operator; ENHANCED indexes cap filters at 100 conditions (each $in value counts).
  • No additional cost. Available in every commercial Region with S3 Vectors, including São Paulo, plus the China Regions.

Why the old mode lost results

Vector search at scale is approximate: the engine explores a neighborhood of promising candidates on a bounded budget rather than comparing your query against every vector.

In CLASSIC mode, filters are checked while that exploration runs. If your filter matches half the corpus, most candidates pass and nothing bad happens. If it matches 0.005%, almost every candidate gets thrown away, the budget runs out, and you get fewer than K results, or K results that aren't the true nearest neighbors within the filtered set. The docs say it plainly: in CLASSIC, filtered queries may return fewer than K results when few vectors match.

An AWS re:Post article measured this on one million synthetic vectors: recall@10 went from 0.86 unfiltered to 0.83 at 10% selectivity, 0.42 at 1%, and 0.15 at 0.01%. The workaround was to partition data into separate indexes by low-cardinality fields (tenant, language, Region), which recovered 15–31 recall points at the cost of managing more indexes.

ENHANCED flips the order. AWS's example: 8 million support tickets, one customer owns 400. With pre-filtering, the similarity search runs over those 400, not over candidates drawn from all 8 million.

One takeaway worth repeating: result count is not a quality signal. Getting 10 results back in CLASSIC never guaranteed they were the best 10.

Index modes at a glance

Situation Default mode Can you change it?
Buckets created on/after Sep 30, 2026 ENHANCED No way back to CLASSIC
Buckets created before Sep 30, 2026 CLASSIC Yes: to ENHANCED via UpdateIndexMode, and back to CLASSIC via CLI/SDK/API only
New indexes in an older bucket Bucket default Set with PutVectorBucketDefaultIndexMode

UpdateIndexMode affects only that index. The console offers Enable enhanced index mode on the index detail page; reverting is not available in the console. The current mode shows up as indexMode in GetIndex.

Filters: operators and the 100-condition limit

Filter syntax is unchanged: $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $exists, $and, $or, plus the new $startsWith (string prefix, ENHANCED only). Prefix matching is handy for path-organized collections: {"s3_path": {"$startsWith": "matter-4417/exhibits/"}}.

The 100-condition cap applies to ENHANCED indexes and is counted per value: {"region": {"$in": ["us-east-1", "us-west-2", "eu-west-1"]}} is 3 conditions. If you build dynamic filters with long $in lists (say, every project a user can access), audit them before migrating: over 100, an ENHANCED index returns a validation error.

Metadata limits haven't changed: up to 40 KB per vector, of which up to 2 KB is filterable; up to 50 keys; up to 10 non-filterable keys per index, fixed at index creation. Keep filter fields short and put chunk text in non-filterable metadata. Bedrock Knowledge Bases already does this with AMAZON_BEDROCK_TEXT and AMAZON_BEDROCK_METADATA.

S3 console index creation form showing non-filterable metadata keys
Non-filterable keys are declared at index creation and can't be changed later. Image: AWS.

Hands-on

You'll need a recent AWS CLI and boto3.

1. Check the mode

aws s3vectors get-index \
  --vector-bucket-name my-vector-bucket \
  --index-name tickets \
  --query 'index.indexMode' --output text
Enter fullscreen mode Exit fullscreen mode

2. Measure recall before migrating. Pick a small tenant, compute its true top-K by brute force, and compare against the index in each mode. This reads the whole index with ListVectors, so do it on a test copy or a sample.

import boto3, json, numpy as np

BUCKET, INDEX, TENANT, K = "my-vector-bucket", "tickets", "t-10428", 10
s3v = boto3.client("s3vectors")
bedrock = boto3.client("bedrock-runtime")

def embed(text):
    body = json.dumps({"inputText": text, "dimensions": 1024, "normalize": True})
    r = bedrock.invoke_model(modelId="amazon.titan-embed-text-v2:0", body=body)
    return json.loads(r["body"].read())["embedding"]

keys, mat, token = [], [], None
while True:
    kw = dict(vectorBucketName=BUCKET, indexName=INDEX, maxResults=1000,
              returnData=True, returnMetadata=True)
    if token:
        kw["nextToken"] = token
    page = s3v.list_vectors(**kw)
    for v in page["vectors"]:
        if v.get("metadata", {}).get("tenant_id") == TENANT:
            keys.append(v["key"]); mat.append(v["data"]["float32"])
    token = page.get("nextToken")
    if not token:
        break
mat = np.array(mat)

def ground_truth(q):
    q = np.array(q)
    sims = mat @ q / (np.linalg.norm(mat, axis=1) * np.linalg.norm(q))
    return {keys[i] for i in np.argsort(-sims)[:K]}

def query(q, mode):
    r = s3v.query_vectors(vectorBucketName=BUCKET, indexName=INDEX,
                          queryVector={"float32": q}, topK=K,
                          filter={"tenant_id": TENANT}, queryMode=mode)
    return {v["key"] for v in r["vectors"]}

for text in ["can't log in", "billing error", "plan change"]:
    q = embed(text); gt = ground_truth(q)
    for mode in ("CLASSIC", "ENHANCED"):
        got = query(q, mode)
        print(f"{mode:9} {text:14} returned={len(got):2} recall@{K}={len(got & gt)/len(gt):.2f}")
Enter fullscreen mode Exit fullscreen mode

This assumes cosine distance and 1024-dim Titan Text Embeddings V2; adjust to your index.

3. Migrate

aws s3vectors update-index-mode \
  --vector-bucket-name my-vector-bucket \
  --index-name tickets \
  --index-mode ENHANCED

aws s3vectors put-vector-bucket-default-index-mode \
  --vector-bucket-name my-vector-bucket \
  --default-index-mode ENHANCED
Enter fullscreen mode Exit fullscreen mode

4. Retrieval on EKS with Pod Identity. The tenant ID comes from the user's token, never from a client-controlled parameter. The API builds the filter; the pod queries S3 Vectors with a role scoped to one index.

Multi-tenant retrieval architecture on Amazon EKS querying an ENHANCED S3 Vectors index
The tenant filter is built server-side and resolved before the similarity search. (Diagram labels in Spanish.)

A permissions gotcha: s3vectors:QueryVectors alone only works without filters and without returnMetadata. With either, you also need s3vectors:GetVectors, or you get a 403.

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3vectors:QueryVectors", "s3vectors:GetVectors"],
    "Resource": "arn:aws:s3vectors:sa-east-1:111122223333:bucket/my-vector-bucket/index/tickets"
  }]
}
Enter fullscreen mode Exit fullscreen mode

Trust policy principal is pods.eks.amazonaws.com with sts:AssumeRole and sts:TagSession, and the cluster needs the eks-pod-identity-agent add-on. Then:

aws eks create-pod-identity-association \
  --cluster-name my-cluster --namespace rag \
  --service-account retrieval \
  --role-arn arn:aws:iam::111122223333:role/rag-retrieval
Enter fullscreen mode Exit fullscreen mode

And the search function:

import os, boto3
s3v = boto3.client("s3vectors")

def search(tenant_id, query_embedding, folder=None, k=8):
    conds = [{"tenant_id": tenant_id}, {"active": True}]
    if folder:
        conds.append({"s3_path": {"$startsWith": folder}})
    r = s3v.query_vectors(
        vectorBucketName=os.environ["VECTOR_BUCKET"],
        indexName=os.environ["VECTOR_INDEX"],
        queryVector={"float32": query_embedding},
        topK=k, filter={"$and": conds},
        returnMetadata=True, returnDistance=True,
    )
    return r["vectors"]
Enter fullscreen mode Exit fullscreen mode

Common traps: treating a tenant filter as isolation (a bug that drops it exposes the whole index; use per-tenant indexes with separate IAM for strict isolation); exceeding 100 conditions; using $startsWith on a CLASSIC index without queryMode=ENHANCED; storing dates as strings (range operators are numeric only); and skipping latency tests with broad filters, since AWS notes latency grows with index size, filter breadth, and condition count.

Pricing

No extra charge for pre-filtering. S3 Vectors pricing in us-east-1: storage $0.06/GB-month; PUT $0.20/GB uploaded; $2.50 per million query API calls; plus per-TB data-processed charges tiered by index size, and the first 512 KB returned per query free. AWS's published example (10M 1024-dim vectors across 40 indexes, 1M queries/month at top-100) comes to $11.38/month.

Note from that example: processing is driven by the size of the index you query, not by how many vectors pass the filter. A selective filter improves recall but doesn't make the query cheaper. Large tenants that always query their own data are still good candidates for a dedicated index.

Availability and caveats

  • Regions: all commercial Regions where S3 Vectors runs (including sa-east-1) and the China Regions.
  • Data residency: vectors stay in the bucket's Region; check where your embedding model runs, especially with cross-Region inference profiles.
  • Cross-Region traffic: keep your compute and your vector bucket in the same Region.
  • Bedrock Knowledge Bases: the announcement doesn't detail how Retrieve filters interact with index mode. If your knowledge base sits on a pre-Sep 30 bucket, check the index mode and measure.

My take

This fixes a limitation that forced you to design around the engine instead of the problem. Until now, the honest answer to "can all my tenants share one index?" was "yes, but small tenants will lose results." With pre-filtering, a shared index becomes reasonable for the long tail, while large tenants or those needing strict IAM isolation keep dedicated indexes. What doesn't change: the filter is still your code, and in multi-tenant RAG a missing filter is a data leak. I'd run the recall test above, migrate CLASSIC indexes, and use the occasion to audit where tenant filters get built.

Resources

Originally published in Spanish on CloudAcademy.ar.

Top comments (0)