Skip to content

Embeddings & Vector Databases

Intermediate · 6 topics

Every RAG system, semantic search box, recommendation feed and agent memory works the same way underneath: turn content into vectors (embeddings), store them, and quickly find the vectors closest to a query. This series covers both halves — the embeddings (what they capture, which model to use, what they cost) and the vector database (how search stays fast at millions of vectors, filtering, multi-tenancy and operations).

pip install numpy scikit-learn chromadb      # everything on these pages runs locally, no API keys
# Topic Sub-topics
1 Embeddings in depth What a vector captures · how embedding models are trained · similarity metrics · normalisation · query vs document · what embeddings miss
2 Choosing & using embedding models Model options · MTEB · evaluating on your data · dimensions & Matryoshka · batching, caching, cost · storage size · re-embedding
3 Vector search & indexes Exact k-NN · recall vs speed · IVF from scratch · HNSW · quantisation (int8, binary, PQ) · tuning
4 Vector databases What a DB adds to an index · metadata filtering · namespaces & multi-tenancy · hybrid search · comparing Pinecone, pgvector, Qdrant, Weaviate, Milvus, Chroma
5 Hands-on: Chroma, Pinecone, pgvector A local Chroma store · the same operations in Pinecone, pgvector and Qdrant · upsert, query, filter, delete
6 Vector DBs in production Ingestion & IDs · updates and deletes · re-embedding migrations · tenant isolation · sizing & cost · monitoring retrieval quality
flowchart LR
    D[Documents] --> C[Chunk] --> E[Embedding model] --> V[(Vector DB<br/>vectors + metadata)]
    Q[Query] --> E2[Same embedding model] --> S[Nearest-neighbour search<br/>+ metadata filter]
    V --> S --> R[Top-k chunks] --> L[LLM / app]

Before you start

The basics are already in the notes: cosine similarity and top-k with NumPy, TF-IDF vs embeddings, how embedding models pool token vectors and BM25 + hybrid search. This series builds on them.

Back to: Notes overview