6. Vector DBs in production¶
Intermediate · 13 min read
Loading documents into a vector database once is easy. Keeping it correct as documents change, secure across customers, affordable as it grows, and good at retrieval over months — that's the production work. Most RAG quality incidents trace back to this layer: stale chunks, duplicates, missing documents, mixed embedding models.
6.1 The ingestion pipeline¶
flowchart LR
S[Sources: Drive, Confluence,<br/>S3, website, DB] --> L[Load & parse]
L --> C[Clean & chunk]
C --> H{Changed?<br/>content hash}
H -->|new / changed| E[Embed] --> U[Upsert]
H -->|unchanged| K[Skip]
C --> D[Find removed chunks] --> X[Delete]
Three properties make it safe to re-run any time:
- Deterministic IDs — the same chunk always gets the same ID (
doc_id#chunk_index), so re-running upserts instead of creating duplicates. - Content hashes — skip re-embedding chunks whose text didn't change (saves time and money).
- Deletion sync — chunks that no longer exist in the source are removed from the index.
6.2 Syncing a changed document¶
When a document is edited, some chunks change, some stay, and some disappear (the document got shorter). Compute exactly what to do:
import hashlib
import re
def chunk(text: str) -> list[str]:
"""One chunk per sentence — tiny, to keep the example readable."""
return [sent for sent in re.split(r"(?<=\.)\s+", text.strip()) if sent]
def chunk_records(doc_id: str, text: str) -> dict[str, dict]:
return {f"{doc_id}#{i}": {"text": c, "hash": hashlib.sha256(c.encode()).hexdigest()[:12]}
for i, c in enumerate(chunk(text))}
def plan_sync(doc_id: str, new_text: str, index_state: dict[str, str]) -> dict[str, list[str]]:
"""index_state: chunk_id -> content hash currently stored in the vector DB for this document."""
new = chunk_records(doc_id, new_text)
existing = {cid: h for cid, h in index_state.items() if cid.startswith(f"{doc_id}#")}
return {
"embed_and_upsert": sorted(cid for cid, r in new.items() if existing.get(cid) != r["hash"]),
"unchanged": sorted(cid for cid, r in new.items() if existing.get(cid) == r["hash"]),
"delete": sorted(set(existing) - set(new)),
}
v1 = ("Refunds are processed within 5 working days of approval. Damaged items get a full refund "
"including shipping. Contact support for refunds older than 30 days.")
index_state = {cid: r["hash"] for cid, r in chunk_records("refunds.md", v1).items()} # what v1 left in the DB
v2 = ("Refunds are processed within 3 working days of approval. Damaged items get a full refund "
"including shipping.") # edited and shortened
for action, ids in plan_sync("refunds.md", v2, index_state).items():
print(f"{action:17} {ids}")
Without the delete step, the old "Contact support for refunds older than 30 days" chunk would stay searchable forever
— the classic stale chunk bug where the bot quotes a policy that was removed months ago.
Fixed-size chunking makes edits expensive
This example chunks by sentence, so an edit only touches its own chunk. With fixed-size chunking, changing one word near the start shifts every later chunk boundary, so all of them get new hashes and are re-embedded. Structure-aware chunking (by heading or paragraph, see NLP for GenAI) keeps unchanged sections stable.
Also store doc_id, source_url, updated_at and embedding_model in each record's metadata — you need them for
citations, freshness filters, deletions by document, and migrations.
6.3 Re-embedding migrations (zero downtime)¶
You will change the embedding model (better model, different dimensions) or the chunking strategy. Vectors from the old and new setup can't be mixed, so treat it like a database migration — blue/green:
flowchart LR
A[Live: index v1<br/>old model] --> B[Build index v2 in background<br/>re-chunk + re-embed everything]
B --> C[Evaluate v2 on the eval set<br/>recall@k, answer quality]
C -->|better| D[Switch reads to v2<br/>via config / alias]
D --> E[Keep v1 for rollback, then delete]
C -->|worse| A
- Write new documents to both indexes while v2 is being built, or replay the changes afterwards.
- Make the active index a config value (or an alias, where the database supports it), so switching and rolling back is instant.
- Budget the time: re-embedding millions of chunks hits API rate limits — estimate it up front.
6.4 Tenant isolation checklist¶
- The tenant is taken from the authenticated session on the server — never from the request body or the LLM.
- Every query path goes through one function that always applies the tenant (like
ChromaStore.queryin the hands-on page). - Document-level permissions (who can see which file) are stored as metadata and filtered, or enforced with database row-level security.
- An automated test queries as tenant A for content that only tenant B has, and fails if anything comes back.
- Deleting a customer deletes their vectors (namespace/collection drop or filter delete) — needed for contracts and data-protection law.
6.5 Sizing: memory and cost¶
HNSW indexes are fastest when the vectors and the graph fit in RAM. Estimate before you choose a plan:
def hnsw_memory_gb(n_vectors: int, dims: int, bytes_per_value: float = 4, m: int = 16, metadata_bytes: int = 500) -> float:
vectors = n_vectors * dims * bytes_per_value
graph = n_vectors * m * 2 * 4 # ~2·M neighbour ids per vector on the bottom layer, 4 bytes each
metadata = n_vectors * metadata_bytes
return (vectors + graph + metadata) / 1e9
for n, d, bpv, label in [(1_000_000, 1536, 4, "1M × 1536, float32"),
(10_000_000, 1536, 4, "10M × 1536, float32"),
(10_000_000, 1536, 1, "10M × 1536, int8"),
(10_000_000, 512, 1, "10M × 512 (Matryoshka), int8")]:
print(f"{label:32} ≈ {hnsw_memory_gb(n, d, bpv):5.1f} GB")
1M × 1536, float32 ≈ 6.8 GB
10M × 1536, float32 ≈ 67.7 GB
10M × 1536, int8 ≈ 21.6 GB
10M × 512 (Matryoshka), int8 ≈ 11.4 GB
Fewer dimensions and quantisation (see Choosing embedding models and quantisation) are the biggest cost levers; managed services price on storage, reads and writes instead, but the same levers apply.
6.6 Monitoring retrieval quality¶
The vector database can be "up" while answers silently get worse. Watch:
| Signal | What it catches |
|---|---|
| Canary queries — a fixed set of questions with known answer chunks, run daily; alert if recall@k drops | broken ingestion, a bad deploy, a mixed-model index |
| Top-1 score distribution over live queries | drift: new topics the corpus doesn't cover, embedding problems |
| No-good-result rate (top score below threshold) | content gaps — questions your documents can't answer |
| Index freshness — age of the newest record per source | a stuck ingestion job |
| Record counts per source / tenant | accidental mass deletes, duplicates |
| Query latency p95 and errors | capacity problems, slow filters |
import numpy as np
rng = np.random.default_rng(3)
last_week = rng.normal(0.62, 0.08, 1000).clip(0, 1) # top-1 similarity of each live query
this_week = np.concatenate([rng.normal(0.62, 0.08, 800), rng.normal(0.35, 0.05, 200)]).clip(0, 1)
THRESHOLD = 0.45
for name, s in [("last week", last_week), ("this week", this_week)]:
print(f"{name}: median top-1 score {np.median(s):.2f}, below threshold {np.mean(s < THRESHOLD):.1%}")
last week: median top-1 score 0.62, below threshold 2.0%
this week: median top-1 score 0.60, below threshold 20.8%
The median barely moved, but one query in five now finds nothing relevant. Reading a sample of those queries usually reveals the cause — a new product launch with no docs yet, or an ingestion job that failed.
6.7 Production checklist¶
- Deterministic chunk IDs, content hashes, deletion sync — ingestion is safe to re-run.
-
doc_id,source,updated_at,tenant,embedding_modelstored with every record. - Query and document embeddings from the same model version; model changes are blue/green migrations.
- Tenant and permission filters applied server-side on every query, with an automated leak test.
- Filter fields indexed; filtered search tested with selective filters.
- Memory/cost estimate; quantisation or smaller dimensions considered.
- Backups / snapshots, and a tested restore — or the ability to rebuild from source.
- Canary queries, score-distribution and freshness monitoring with alerts.
Interview questions¶
How do you keep a vector index in sync with changing documents?
Use deterministic chunk IDs and content hashes. On each sync, re-chunk the document, upsert chunks whose hash changed, skip unchanged ones, and delete IDs that no longer exist. Store doc_id and updated_at so you can delete or filter by document and freshness. Prefer structure-aware chunking so small edits don't shift every chunk.
How would you migrate to a new embedding model without downtime?
Build a new index alongside the old one, re-embedding all documents with the new model while dual-writing new changes. Evaluate it on a retrieval eval set, switch reads via a config flag or alias, keep the old index for rollback, then delete it. Never mix vectors from different models in one index.
How do you monitor a RAG retriever in production?
Canary queries with known answers (recall@k over time), the distribution of top similarity scores on live traffic, the rate of queries with no good match, index freshness and record counts per source, plus latency and errors. Investigate samples whenever these shift.
Practice¶
- Extend
plan_syncto handle a document that was deleted from the source entirely. - Write the tenant-leak test from 6.4 for the
ChromaStoreclass in the hands-on page.
Next: RAG — putting chunking, embeddings, vector search and LLMs together (coming soon). Back to: Embeddings & Vector Databases overview · Notes overview