Embeddings & Vector Databases¶
Intermediate · 6 topics
Every RAG system, semantic search box, recommendation feed and agent memory works the same way underneath: turn content into vectors (embeddings), store them, and quickly find the vectors closest to a query. This series covers both halves — the embeddings (what they capture, which model to use, what they cost) and the vector database (how search stays fast at millions of vectors, filtering, multi-tenancy and operations).
| # | Topic | Sub-topics |
|---|---|---|
| 1 | Embeddings in depth | What a vector captures · how embedding models are trained · similarity metrics · normalisation · query vs document · what embeddings miss |
| 2 | Choosing & using embedding models | Model options · MTEB · evaluating on your data · dimensions & Matryoshka · batching, caching, cost · storage size · re-embedding |
| 3 | Vector search & indexes | Exact k-NN · recall vs speed · IVF from scratch · HNSW · quantisation (int8, binary, PQ) · tuning |
| 4 | Vector databases | What a DB adds to an index · metadata filtering · namespaces & multi-tenancy · hybrid search · comparing Pinecone, pgvector, Qdrant, Weaviate, Milvus, Chroma |
| 5 | Hands-on: Chroma, Pinecone, pgvector | A local Chroma store · the same operations in Pinecone, pgvector and Qdrant · upsert, query, filter, delete |
| 6 | Vector DBs in production | Ingestion & IDs · updates and deletes · re-embedding migrations · tenant isolation · sizing & cost · monitoring retrieval quality |
flowchart LR
D[Documents] --> C[Chunk] --> E[Embedding model] --> V[(Vector DB<br/>vectors + metadata)]
Q[Query] --> E2[Same embedding model] --> S[Nearest-neighbour search<br/>+ metadata filter]
V --> S --> R[Top-k chunks] --> L[LLM / app]
Before you start
The basics are already in the notes: cosine similarity and top-k with NumPy, TF-IDF vs embeddings, how embedding models pool token vectors and BM25 + hybrid search. This series builds on them.
Back to: Notes overview