Qdrant multitenancy is not one filter added to an otherwise shared search request. A production design needs three explicit boundaries: authorization decides which tenant the caller may access; retrieval limits candidates to that tenant; and sparse scoring computes term rarity from the correct tenant corpus. Custom shards can add placement and failure-domain control for large or regulated tenants, but they do not replace application authorization. The safest starting point for many SaaS products is one homogeneous collection, a trusted `tenant_id` filter on every operation, a keyword payload index marked `is_tenant=true`, and isolation tests that fail the build if another tenant ever appears.
Qdrant's multitenancy, indexing and search documentation was verified on 7 October 2026. Version availability and documented engine behavior are Qdrant facts. The authorization service, rollout gates, test suite and tenant-promotion policy below are engineering recommendations, not claims about Qdrant Cloud account isolation or guarantees for an untested deployment.
Separate the four meanings of isolation
Teams often say “tenant isolation” while referring to different controls. Query isolation means a request cannot return another tenant's points. Administrative isolation means one customer's operator cannot list, delete or reconfigure another customer's data. Performance isolation limits noisy-neighbor effects. Placement isolation controls which shard, node or region stores the data. One mechanism does not provide all four.
A payload filter supports query isolation only when trusted code derives it from an authenticated principal and attaches it to every read, write, update, scroll, count and delete. An `is_tenant` payload index helps storage locality and filtered performance. A shard key limits which physical partitions an operation reaches. A separate collection can provide a stronger configuration and lifecycle boundary. Rate limits, quotas and admission control handle resource fairness.
Write these guarantees as a matrix before choosing topology. If a regulated customer requires a dedicated encryption key, independent backup schedule or region, a shared payload partition may not meet the administrative contract even if query filtering is correct. If thousands of small customers share the same schema and embedding model, one collection per customer may create avoidable operational overhead.
Requirement Primary control
Never return another tenant's point Trusted mandatory payload filter
Keep one tenant's vectors close on disk Keyword index with is_tenant=true
Keep sparse term rarity tenant-specific params.idf.corpus tenant filter
Route traffic to selected partitions Custom shard key selector
Separate schema, model or lifecycle Separate collection
Limit a noisy tenant's resource use Quota, rate limit and admission controlPut tenant context before the vector database
The client must never choose its own `tenant_id`. Resolve tenant context after authentication from a membership or entitlement store. The API gateway should pass a signed internal identity, and the retrieval service should construct the Qdrant filter from that trusted context. If a user belongs to several organizations, require one active organization per request or a separately authorized cross-organization workflow.
Use the same path for mutations. An update by point ID without a tenant condition can be just as dangerous as an unfiltered search. Give point IDs globally unique or tenant-namespaced values, but do not treat ID shape as authorization. Before deleting by document, include tenant, document and version conditions. Before hydrating result text from PostgreSQL or object storage, recheck tenant authorization at that source of truth. Defense in depth matters because a vector store result may outlive a changed membership.
function tenantScope(principal, requestedOrg) {
const tenantId = entitlements.requireMembership(principal, requestedOrg);
return {
must: [
{ key: 'tenant_id', match: { value: tenantId } },
{ key: 'state', match: { value: 'published' } }
]
};
}
const result = await qdrant.query('knowledge', {
query: queryVector,
using: 'dense_v3',
filter: tenantScope(session.user, request.org),
limit: 20,
with_payload: ['document_id', 'chunk_id', 'version']
});Keep this builder small, centralized and impossible to bypass through a generic client exposed to route handlers. Do not let prompts, tool arguments or user-supplied JSON provide the filter. The model may suggest search terms; it does not decide authorization scope.
Build the tenant payload index before ingestion
Qdrant payload indexes are not merely lookup accelerators. They give the query planner cardinality estimates and help the filterable HNSW graph support filtered traversal. Create indexes for fields used frequently in filters, especially the tenant key and publication state, before bulk ingestion when possible. Indexing every payload field wastes memory and build time; start with the fields that materially restrict the candidate set.
For the tenant field, Qdrant supports a keyword payload index with `is_tenant=true` from v1.11. The hint tells Qdrant that values identify tenant-like groups and allows tenant points to be colocated for more sequential reads. It does not authenticate callers and it does not add a filter automatically.
PUT /collections/knowledge/index
{
"field_name": "tenant_id",
"field_schema": {
"type": "keyword",
"is_tenant": true
}
}Create the collection and required indexes through versioned infrastructure code, then ingest. In staging, inspect collection state and run filtered latency and recall tests after every index change. If an index is added to an existing collection, plan optimizer work and verify the built segments before declaring success. Qdrant's documentation notes that global requests become slower in a tenant-focused HNSW configuration, so a global admin search should be a separate, explicitly authorized workload rather than an accidental query without `tenant_id`.
Keep retrieval scope and sparse IDF scope distinct
Dense similarity compares vectors and does not use term-frequency statistics. BM25 and miniCOIL sparse retrieval do: inverse document frequency raises the weight of terms that are rare in the scoring corpus. With payload-partitioned tenants, Qdrant historically computed IDF from the whole shard. That can mix vocabularies: a product code rare for tenant A may look common because tenants B through Z use it often.
Qdrant v1.19 adds an `idf` search parameter whose `corpus` accepts a payload filter. Retrieval and scoring filters remain separate. A query may retrieve only tenant A's published Arabic documents while computing IDF over all of tenant A's documents. Making the IDF corpus identical to a tiny year or category filter can produce unstable rarity statistics; making it cluster-wide lets other tenants change the score distribution.
POST /collections/knowledge/points/query
{
"query": { "text": "invoice NY-4821", "model": "qdrant/bm25" },
"using": "title-bm25",
"filter": {
"must": [
{ "key": "tenant_id", "match": { "value": "tenant-a" } },
{ "key": "state", "match": { "value": "published" } },
{ "key": "language", "match": { "value": "en" } }
]
},
"params": {
"idf": {
"corpus": {
"must": [
{ "key": "tenant_id", "match": { "value": "tenant-a" } }
]
}
}
},
"limit": 10
}For hybrid search, apply the trusted retrieval scope to both dense and sparse candidate legs and the tenant corpus to the sparse leg. An authorization bug in either prefetch can consume the candidate budget with inaccessible points even if a late-stage filter removes them. See the Qdrant hybrid-search guide for fusion and candidate-depth evaluation; the isolation contract belongs before fusion.
Choose shared, tiered or dedicated topology deliberately
A shared collection is usually the best starting point when tenants use the same vector dimensions, distance metric, payload schema and retention policy. It amortizes collection overhead and makes fleet-wide configuration manageable. Pair it with mandatory payload filters, tenant indexes and resource quotas.
Tiered multitenancy promotes large or regulated tenants to custom shard keys while keeping the long tail shared. Qdrant custom sharding lets the application choose a `shard_key_selector` for upserts and queries. A shard key can represent a region or capacity tier, while `tenant_id` still identifies the customer inside that shard. This separates placement from authorization. Maintain one authoritative directory that maps tenant to collection, shard key and policy version; never derive routing from an untrusted request field.
Use a separate collection when a tenant needs a different embedding model or vector dimension, independent schema changes, dedicated replication, retention, backup, encryption or operational ownership. Use a separate cluster or account when the contractual boundary exceeds what shared administration should carry. Avoid collection-per-user by reflex: large collection counts can multiply optimizer, indexing and monitoring work.
Tenant directory
tenant-a -> collection=knowledge, shard=shared-eu, policy=v7
tenant-b -> collection=knowledge, shard=dedicated-b, policy=v3
tenant-c -> collection=regulated-c, shard=eu-central, policy=v9Promotion must be a migration workflow: dual-read or shadow-read, backfill, consistency check, route switch and rollback window. Reuse the principles in the Qdrant embedding migration guide, but measure cross-topology completeness as well as relevance.
Treat custom sharding as placement, not permission
A shard selector reduces fan-out and can isolate a large tenant's resource usage or meet region-aware placement. It is valuable for data locality and failure domains. It is not an authorization token. If the application selects the wrong shard or omits the tenant filter inside a shared shard, the database cannot infer the intended customer boundary from business identity.
Every operation should resolve both values from the same trusted tenant directory: the retrieval filter contains tenant identity; the shard selector contains placement. Validate that the chosen tenant is allowed on the chosen shard, and reject an unknown or migrating state. Log tenant, shard, collection, policy version and request ID without logging private query text by default.
Cross-tenant analytics should not reuse the product retrieval endpoint with the filter removed. Build a separate batch pipeline, service identity and output store with its own purpose limitation. Tenant-focused HNSW settings may make unfiltered search slower anyway. More importantly, an explicit analytics path can aggregate counts or embeddings without giving interactive callers a global search primitive.
Test isolation as an invariant, not a demo
Create synthetic tenants whose documents deliberately look alike. Give tenant A and tenant B identical text but distinct canary IDs. For every API operation—query, scroll, recommend, count, update, set payload and delete—assert that an A credential cannot observe or mutate B. Run the suite through every public route, background job and agent tool, not just through one retrieval function.
Property-based tests can generate tenants, documents, roles and filter combinations. Mutation tests should remove the tenant condition and prove that the test fails. Add a production sentinel query containing only synthetic canaries and alert on any foreign tenant result. A cross-tenant hit is a security incident, even when the answer shown to the user later omits it.
Relevance evaluation also needs tenant boundaries. Track Recall@k, nDCG, zero-result rate and latency per tenant size class. For sparse search, compare global IDF with tenant-scoped IDF on terms whose frequency differs across tenants. For HNSW filters, compare approximate results against exact search on a labeled sample. Qdrant documents ACORN for combinations of strict filters that fragment ordinary HNSW traversal; enable it only after a benchmark shows an accuracy problem worth its extra search cost.
Operate migrations, deletes and failures safely
Tenant deletion is a workflow, not one delete call. Stop new writes, revoke memberships, delete or tombstone points with the tenant condition, verify the count is zero, remove secondary content and backups according to policy, and retain a minimal audit receipt. A failed step should resume idempotently. Never accept a tenant ID from a deletion request without resolving ownership and high-impact authorization.
During re-embedding or backfill, preserve tenant and shard metadata on every new point. A migration that copies vectors but drops payload indexes or routing fields can pass relevance tests while breaking isolation. Shadow queries must use the same trusted context on old and new paths. The production RAG pipeline guide explains the broader sequence from chunking through context construction.
Monitor filtered-query latency, HNSW versus exact execution, optimizer backlog, shard imbalance, per-tenant storage, rate-limit rejections, empty allowed candidate sets and authorization denials. Do not label every zero-result query a search failure; it may be the correct result after access is revoked. Keep security signals separate from relevance metrics.
When not to use Qdrant multitenant vector search
Use SQL when the answer is an exact entitlement, balance, order state or inventory quantity. Use conventional full-text search when the corpus is small, lexical matching is sufficient and vector semantics add no measured value. Keep highly sensitive workloads out of a shared collection when regulation or customer contracts require independent administration, keys, retention or backups.
Do not use embeddings as an access-control list. Do not rely on prompt instructions such as “only answer from this tenant.” Do not use post-filtering after an unscoped nearest-neighbor query. Do not expose a generic Qdrant client to an AI agent. And do not assume a shard key, `is_tenant` index or globally unique point ID proves the caller is allowed to access the point.
The executive design is compact: derive tenant context from authenticated membership; resolve placement from a trusted directory; attach a mandatory filter to every database operation; index the tenant field with `is_tenant`; scope sparse IDF to the tenant; recheck authorization when hydrating content; promote heavy tenants through a measured migration; and continuously test cross-tenant canaries. For implementation help across vector retrieval and product boundaries, see AI integrations and retrieval systems.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




