Metadata Filtering in Vector Search
Real vector searches almost never mean "find the nearest vectors among everything". They mean "find the most relevant chunks from this customer's documents, in English, updated this year, that this user is allowed to see". Metadata filtering combines semantic similarity with structured conditions, and it's where many vector search implementations quietly fail: returning too few results, slowing down dramatically, or leaking data across tenants.
The core difficulty is that ANN indexes are built for similarity, not for arbitrary filters. How an engine combines the two (pre-filtering, post-filtering, or filter-aware traversal) determines correctness and performance, especially for highly selective filters that match a small fraction of vectors.
TL;DR
- Post-filtering (ANN first, filter after) is fast but can return too few results when filters are selective.
- Pre-filtering (filter first, then search the matches) is exact, but can be slow, or bypass the ANN index.
- Filter-aware ANN (in-graph filtering, as in Qdrant, Weaviate, Pinecone, and Elasticsearch/OpenSearch kNN with filters) traverses the index while respecting filters. It's the best general approach.
- Index your metadata fields (payload indexes) so filters are cheap.
- Access control and tenancy filters are security boundaries: enforce them server-side on every query.
- For very selective or per-tenant filters, consider partitioning (collections, namespaces, partial indexes) instead.
Quick Example
Filtered search in Qdrant with an indexed payload:
The same idea in PostgreSQL with pgvector, using iterative index scans to avoid under-filled results:
Core Concepts
Post-Filtering
Run ANN search for the top K candidates, then drop those failing the filter. It's simple and fast, but if the filter matches 1% of vectors, the top 100 candidates may contain only one match, or zero. Workarounds (fetching 10× or 100× more candidates) raise latency and still fail for very selective filters.
Pre-Filtering
Find all vectors matching the filter first (using metadata indexes), then compute similarity only among them. Results are exact, but:
- If the matching set is large, brute-force scoring is slow.
- Naively, it can't use the ANN graph built over all vectors.
It works well when the filter is highly selective (a few thousand matches): brute force over the small set is fast and exact. Many engines switch automatically to exact search in this case.
Filter-Aware ANN
Modern engines integrate filtering into the index traversal: HNSW search skips (but may still traverse through) nodes that don't match, and extra graph links, or payload-based subgraphs, keep connectivity high under filters. Qdrant (filterable HNSW with payload indexes), Weaviate (ACORN-style filtered HNSW), Pinecone, Milvus, and Elasticsearch and OpenSearch kNN with filters all implement variants. Query planners choose between filtered ANN and exact search based on filter selectivity.
pgvector historically post-filtered within index scans, and could return fewer rows than LIMIT. Version 0.8 added iterative index scans, which continue scanning until enough filtered results are found. Partial indexes and partitioning also help. See pgvector.
Metadata (Payload) Indexes
Filters are only fast if metadata fields are indexed: keyword and equality indexes for IDs and categories, range indexes for numbers and dates, and full-text indexes for text fields. Unindexed filters force scanning metadata for every candidate. Index the fields used in filters, especially tenant, permission, and category fields.
Multi-Tenancy and Access Control
In RAG and SaaS search, filters enforce who can see what:
- Tenant isolation: every query must include the tenant filter, applied server-side from the authenticated user, never from client input. See broken access control.
- Document permissions: store ACLs (allowed users or groups) as metadata, and filter by the user's groups. Keep them in sync when permissions change.
- Isolation strategies:
Hybrid Search With Filters
When combining vector and keyword retrieval (hybrid search), apply the same filters to both retrievers. Filtering only after fusion can leak unauthorized candidates into intermediate results, or starve the final list.
Best Practices
Decide Filters at Schema Design Time
Identify filter fields (tenant, language, product, date, visibility) up front, store them as typed metadata, and index them. Retrofitting metadata means re-ingesting data.
Test With Realistic Selectivity
Benchmark filtered queries at the selectivities you'll see in production: 50%, 1%, and 0.01% matches. Performance and recall can differ drastically from unfiltered benchmarks.
Enforce Security Filters Centrally
Build filter construction into a retrieval layer that derives tenant and permission conditions from the authenticated identity, so individual features can't forget them. Log and test cross-tenant isolation.
Consider Partitioning for Very Selective Filters
If most queries filter to one small tenant or category, per-tenant partitions or collections (or partial indexes) give both better performance and stronger isolation.
Common Mistakes
Post-Filtering With a Small K
Requesting 10 results, then filtering to one tenant, returns fewer than 10 (sometimes none) for small tenants. Use filter-aware search, or iterative scans.
Trusting Client-Supplied Filters
Letting the frontend send tenant_id or allowed_groups lets attackers query other tenants' data. Derive security filters from server-side authentication.
Stale Permission Metadata
Document permissions change, but vector metadata isn't updated, so revoked users still retrieve content. Sync ACL changes to the index promptly, or check permissions again after retrieval.
FAQ
What's the difference between pre-filtering and post-filtering?
Pre-filtering narrows the candidate set to vectors matching metadata conditions before similarity search, which is exact but potentially slow. Post-filtering runs similarity search first and removes non-matching results afterwards, which is fast but can return too few results for selective filters. Filter-aware ANN combines the two during index traversal.
Why does my filtered vector search return fewer results than requested?
The engine likely post-filters: the ANN search found K nearest vectors overall, and most were removed by your filter. Use an engine or setting with filter-aware search (pgvector iterative scans, or payload-indexed filtering in dedicated vector databases), or increase candidates, or partition the data.
How do I implement multi-tenant vector search securely?
Always apply a tenant filter derived from the authenticated user, index the tenant field, and consider stronger isolation (namespaces, partitions, separate collections, or database row-level security) for sensitive data. Test that queries can't return other tenants' vectors.
Should permissions be enforced in the vector database or the application?
Both, ideally: filter by ACL metadata in the vector query for efficiency and to avoid leaking content into prompts, and verify access in the application layer, especially when permissions change frequently or are complex.
Related Topics
- Vector Databases — Engines and their filtering capabilities
- ANN Indexes — How filters interact with graphs and clusters
- RAG — Filtering retrieval by tenant, language, and permissions
- pgvector — Iterative scans and partial indexes
- Row-Level Security — Database-enforced tenant isolation
- Hybrid Search — Consistent filters across retrievers