Changing an embedding model is a data migration, not a configuration edit. The new model can change vector dimensions, similarity behavior, query/document prompting, data type, latency and ranking quality. A safe Qdrant migration keeps the old search path live, dual-writes new content, backfills historical points idempotently, compares both systems on the same tenant-scoped queries, switches traffic atomically and preserves a tested rollback window.
Qdrant documents two valid mechanisms: a blue-green migration that builds a second collection and switches an alias, and—in Qdrant 1.18 or later—a named-vector migration that adds another vector space to an existing collection. The official migration guide is the factual basis for these capabilities. The release gates, observability and recovery rules below are application-level engineering recommendations.
Why an embedding change can break retrieval
Vectors from different models do not share a meaningful coordinate system merely because they have the same length. Never embed a query with the new model and search old document vectors, unless the model provider explicitly documents cross-model compatibility. Even within a compatible family, keep the exact model, output dimension, input role, truncation and normalization choices in a versioned contract.
Dimensions are part of the Qdrant vector schema. The Collections documentation says vectors within one named space share a dimension and distance metric; named vectors can carry different dimensions and metrics. A switch from 1024 to 512 dimensions therefore needs a new named vector or a new collection, not an in-place overwrite of the old space.
Provider parameters also affect meaning. Voyage's Embeddings API distinguishes `input_type: document` from `input_type: query`, supports selected output dimensions, and can return float, integer or binary data for supported models. Treat these as part of `embedding_contract_version`, not incidental request options. Silent truncation can also change what a chunk represents; record and measure it.
Choose named vectors or blue-green collections
Use named vectors when the existing collection already uses named vectors, Qdrant is version 1.18 or later, payload and point IDs should stay in one place, and temporary storage for a second vector per point is acceptable. Add a new vector name with its own size and distance, populate it, route searches with the `using` parameter, then delete the old name only after the rollback window.
Use blue-green collections when the source collection is not eligible for the named-vector path, the schema or sharding plan changes materially, you want strict physical separation, or you need to tune HNSW, quantization and storage independently before cutover. Build `documents_v2` while traffic keeps using the stable alias, then atomically move the alias from v1 to v2.
The trade-off is operational. Named vectors reduce payload duplication and make per-point updates simpler, but the collection carries two vector indexes during migration and cleanup becomes irreversible when the old vector name is deleted. Blue-green consumes a second collection and requires consistent writes to both, but rollback is an alias switch and the old collection remains intact.
Freeze the embedding contract first
Before backfill, write a migration manifest containing source and target collection or vector name, provider, exact model, output dimension, data type, distance metric, query and document input roles, chunking version, normalization, tenant filter version and text-hash algorithm. Give the manifest an immutable migration ID.
{
"migration_id": "emb-2026-10-v2",
"source": {"collection": "docs_v1", "vector": "dense_v1"},
"target": {"collection": "docs_v2", "vector": "dense_v2"},
"embedding": {
"model": "approved-model-snapshot",
"dimension": 1024,
"document_input_type": "document",
"query_input_type": "query"
},
"chunking_version": "chunk-v4"
}This is an application manifest, not a Qdrant or Voyage response. Pinning it prevents two workers from producing incompatible target vectors under the same name. Store the model revision and contract version with each migration checkpoint; do not place secrets or source text in the manifest.
Decide whether this is only an embedding migration. If chunk boundaries change too, point IDs and result granularity may change, making pairwise comparison harder. Prefer migrating one independent variable at a time, or explicitly treat the work as a full retrieval-index rebuild with its own relevance labels.
Start dual writes before historical backfill
First update the ingestion path so every create, update and delete reaches both old and new targets. Derive stable point IDs from the source document and chunk identity. Use an outbox or durable job record so an embedding-provider timeout cannot leave the new index permanently behind the source of truth.
For named vectors, an upsert can carry both old and new vectors for a point. For blue-green, execute two idempotent upserts from the same committed document revision. Do not rely on two best-effort network calls inside a request. Persist the required target states and retry them independently.
Tag every point payload with `tenant_id`, `document_id`, `content_revision`, `content_hash`, `chunking_version` and `embedding_contract_version`. Preserve the same tenant filter in old search, new search, backfill and evaluation. A better model does not compensate for a missing authorization boundary.
Deletes need the same dual path. A migration that copies historical points but misses a deletion can resurrect confidential or stale content at cutover. Record tombstones or delete revisions until both targets acknowledge them.
Backfill as a resumable, rate-limited workflow
Read source documents from the authoritative database when possible, not by treating stored vectors as recoverable text. If Qdrant payload contains the approved canonical text, the Points API supports scrolling and batch operations; request only the payload fields needed and exclude old vectors to reduce transfer.
Partition work by tenant and stable point-ID ranges. Each job should read a bounded batch, confirm the current document revision, embed with the pinned contract, upsert the target vector and persist a checkpoint. On retry, the same point ID and contract version make the write idempotent. Add exponential backoff and jitter for embedding or Qdrant rate limits, and quarantine deterministic failures such as invalid source text.
Do not declare completion from `indexed_vectors_count`; Qdrant documents that internal point and indexed-vector counters can be approximate during optimization. Use the exact Count API with a filter on the target contract version, compare against the authoritative eligible-source count, and sample point IDs for missing or stale revisions. Wait for index construction and optimization to settle before measuring tail latency.
Throttle by provider token limits, Qdrant write latency, optimizer pressure, CPU, memory and disk headroom. Backfill is a production workload. A migration that preserves availability but exhausts disk or makes p99 search latency unacceptable is not zero-downtime.
Shadow retrieval, not just row counts
Completeness only proves that vectors exist. It does not prove that they retrieve useful evidence. Build a frozen evaluation set of real query shapes with relevance judgments, tenant and filter context, language, document type and difficult negatives. Embed every query with the matching query contract for each candidate index.
Compare Recall@k, nDCG@k, MRR, zero-result rate, filter correctness, citation coverage and downstream grounded-answer acceptance. Also measure embedding latency, Qdrant p50/p95/p99, candidate count, payload hydration cost and total cost per accepted answer. Report metrics by cohort; an overall gain can hide a regression in Arabic, code, tables or one tenant's domain.
Run shadow queries asynchronously so user latency remains tied to the old path. Log IDs and scores only under the site's data policy; relevance analysis may expose document identifiers even without text. Never let shadow results perform side effects or reach the user.
Score scales can change between models, so do not carry a similarity threshold forward blindly. Recalibrate minimum scores, candidate depth and reranker inputs on the new score distribution. When a reranker follows retrieval, evaluate both pre-rerank recall and final ranking quality.
Cut over atomically and keep rollback warm
For blue-green, Qdrant collection aliases are designed to switch vector versions without stopping concurrent requests. Submit delete-old-alias and create-new-alias in one alias update so the move is atomic. For named vectors, change one versioned route configuration that contains both the embedding model and the `using` vector name; switching only one of them creates a cross-space query bug.
Canary before full cutover when the application can route a stable tenant or request cohort directly to the target. Compare user-visible outcomes and operational metrics, then move the shared route. Record route version, resolved collection or vector name, embedding contract and index build revision on each trace.
Keep dual writes and the old path for an observation window. Roll back on critical relevance regressions, authorization or filter mismatch, target lag, increased zero-result rate, p99 latency, error rate or cost per accepted answer. Rollback changes the route; it should not require another full backfill.
Qdrant's consistency documentation explains the availability cost of stronger write ordering and read consistency in replicated clusters. Choose settings for the application's conflict model rather than enabling the strongest option reflexively. Your durable ingestion ledger remains the evidence that both targets received a document revision.
Delete the old path only after evidence
Stop dual writes only after all traffic uses the new path, lag is zero, the evaluation bar remains green through a representative period, and a recovery snapshot or authoritative re-embedding path is tested. With named vectors, deleting the old vector name removes its schema and data; the migration guide calls this the point of no return without re-embedding.
With blue-green, remove the old collection after the retention and rollback policy expires. First verify no alias, worker, dashboard or replay job references it. Deletion frees resources but also removes the fastest rollback, so make it an explicit approved operation rather than an automatic cleanup on deployment success.
When not to migrate
Do not migrate because one public benchmark ranks a model higher. If current retrieval already meets the product SLO and the new model does not improve workload-shaped evaluations enough to justify embedding cost, storage, operational risk and future lock-in, keep the existing contract.
A lexical or SQL query can be better when users need exact IDs, totals, ranges or deterministic filters. Hybrid search can be better when dense embeddings miss identifiers and rare terms; this migration pattern still applies to the dense branch, while the sparse branch remains independently versioned. And if the failure is poor chunking or missing source data, changing the embedding model will not repair the evidence boundary.
Production checklist
Choose the migration shape, pin the complete embedding contract, enable durable dual writes including deletes, run a resumable tenant-scoped backfill, verify exact coverage, wait for indexing, evaluate retrieval and answer quality, canary the complete route bundle, switch atomically, observe with rollback warm, and delete the old path only after explicit approval.
The safest migration is intentionally boring: every transition has a checkpoint, every comparison uses the same authorization filters, and no cleanup happens before rollback is no longer needed. Pair this workflow with the Qdrant hybrid-search guide and the production RAG pipeline to keep indexing, retrieval and generation contracts independently testable.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




