1. Embeddings in depth¶
Intermediate · 12 min read
An embedding is a list of numbers (a vector, typically 384–3,072 long) that represents the meaning of a piece of text — or an image, or a product. Texts with similar meaning get vectors that point in similar directions. That one property powers semantic search, RAG retrieval, clustering, deduplication, recommendations and classification.
1.1 A stand-in embedding model¶
Real embedding models are neural networks (an API call or a ~100 MB download). To keep every example on these pages runnable offline, we use a small stand-in: TF-IDF over character n-grams, normalised to unit length. It captures spelling similarity, not real meaning — good enough to demonstrate the mechanics, and the code has the same shape as a real model.
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
CORPUS = [
"Refunds are processed within 5 working days of approval.",
"You can return damaged items for a full refund.",
"Orders ship within 48 hours from our Pune warehouse.",
"Express delivery is available in metro cities.",
"Reset your password from Settings, then Security.",
"Two-factor authentication protects your account.",
"Invoices are emailed after every payment.",
"We accept UPI, credit cards and net banking.",
]
class StandInEmbedder:
"""Same interface as a real model: .encode(list_of_texts) -> (n, dim) array of unit vectors."""
def __init__(self, corpus):
self.tfidf = TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5), sublinear_tf=True).fit(corpus)
def encode(self, texts):
v = self.tfidf.transform(texts).toarray()
return v / np.maximum(np.linalg.norm(v, axis=1, keepdims=True), 1e-9)
model = StandInEmbedder(CORPUS)
E = model.encode(CORPUS)
print(E.shape, np.linalg.norm(E, axis=1).round(3)[:3])
The stand-in's dimension is the number of distinct character n-grams in the corpus. A real model always returns the same fixed size — e.g. 384 or 1,536 — whatever the input.
The real thing has the same shape:
# no-run — pip install sentence-transformers (downloads ~90 MB once)
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
E = model.encode(CORPUS, normalize_embeddings=True) # (8, 384)
1.2 Search = nearest neighbours¶
With unit-length vectors, cosine similarity is a dot product, and one matrix multiply scores every document:
def search(query: str, k: int = 3):
q = model.encode([query])[0]
scores = E @ q
for i in np.argsort(-scores)[:k]:
print(f"{scores[i]:.2f} {CORPUS[i]}")
search("how long do refunds take")
0.37 Refunds are processed within 5 working days of approval.
0.23 You can return damaged items for a full refund.
0.03 Orders ship within 48 hours from our Pune warehouse.
That is the whole idea behind vector search. Everything else on these pages is about doing it well (better embeddings) and fast (millions of vectors).
1.3 How embedding models learn meaning¶
Embedding models are transformers (usually encoder-only, like BERT) fine-tuned with contrastive learning:
flowchart LR
Q["query: 'refund time?'"] --> M[Same model]
P["matching passage:<br/>'Refunds take 5 days'"] --> M
N["other passages in the batch<br/>(negatives)"] --> M
M --> L["Loss: pull query and its passage together,<br/>push the negatives apart"]
Training data is millions of (query, relevant passage) pairs — search logs, question/answer sites, title/body pairs, and increasingly LLM-generated pairs. Each training step makes matching pairs more similar and everything else in the batch less similar. The geometry that results is what makes "money back for broken items" land near "refund for damaged products" with no shared words.
def contrastive_loss(q, passages, positive_index, temperature=0.05):
"""InfoNCE: softmax over similarities; the loss is low when the true passage wins clearly."""
logits = passages @ q / temperature
logits -= logits.max()
probs = np.exp(logits) / np.exp(logits).sum()
return float(-np.log(probs[positive_index]))
q = model.encode(["how long do refunds take"])[0]
print("refund passage is the positive: ", round(contrastive_loss(q, E, positive_index=0), 3))
print("shipping passage is the positive:", round(contrastive_loss(q, E, positive_index=2), 3))
Training adjusts the model's weights to make the first kind of number small for every real pair.
Hard negatives — passages that look similar but are wrong ("refunds for digital products are not available") — are what teach a model fine distinctions. They are also how you fine-tune an embedding model on your own domain.
1.4 Similarity metrics¶
| Metric | Formula | Notes |
|---|---|---|
| Cosine similarity | a·b / (‖a‖‖b‖) | direction only; the default for text |
| Dot product | a·b | equals cosine for unit vectors, and is faster; most DBs' fastest option |
| Euclidean (L2) distance | ‖a − b‖ | for unit vectors, ranks exactly like cosine: ‖a−b‖² = 2 − 2·cos |
a, b = E[0], E[1]
cos = a @ b
l2 = np.linalg.norm(a - b)
print(round(float(cos), 4), round(float(l2 ** 2), 4), round(float(2 - 2 * cos), 4))
So for normalised vectors the choice of metric doesn't change the ranking. It matters when vectors aren't normalised: then dot product also rewards vector length, which some models use on purpose. Rule: use the metric in the model's documentation, and configure the same metric in your vector database.
1.5 Query vs document embeddings¶
Questions and passages look different — "refund time?" vs a paragraph of policy text. Many models are trained to embed them asymmetrically and expect a hint:
| Model family | How you mark it |
|---|---|
| E5 | prefix "query: " or "passage: " |
| BGE | instruction prefix for queries, e.g. "Represent this sentence for searching relevant passages: " |
| Nomic, some others | "search_query: " / "search_document: " |
| Cohere, Voyage, Gemini APIs | an input_type / task_type parameter (search_query vs search_document) |
OpenAI text-embedding-3-* |
no prefix needed |
Forgetting the prefix is a silent bug: everything still works, just with noticeably worse retrieval.
1.6 What embeddings miss¶
Embeddings compress meaning into one vector, and some things get lost:
| Weakness | Example | Fix |
|---|---|---|
| Exact identifiers | error code E-4012, SKU TSH-COT-RED-M, a person's name |
hybrid search with BM25 (NLP notes) |
| Negation | "flights without a layover" retrieves layover pages | rerankers; filters; query rewriting |
| Numbers and ranges | "phones under ₹20,000" | metadata filters (price < 20000), not vectors |
| Long texts | a 20-page document becomes one blurry vector | chunk before embedding |
| Domain jargon | internal acronyms the model never saw | domain-tuned or fine-tuned embedding models; glossaries |
| Freshness / permissions | "latest policy", "documents I can see" | metadata (updated_at, tenant_id) + filters |
0.13 Express delivery is available in metro cities.
0.04 Invoices are emailed after every payment.
Nothing in the corpus mentions that error, so the "nearest" vectors are still returned — with low scores. Vector search always returns something; your app must decide when a result is too weak to use (a score threshold, or a reranker).
Interview questions¶
What is an embedding and how is it created?
A fixed-length dense vector representing the meaning of an input. Text embedding models are transformer encoders whose token outputs are pooled into one vector, fine-tuned with contrastive learning on (query, relevant passage) pairs so that related texts end up close together and unrelated ones far apart.
Cosine, dot product or Euclidean?
For normalised vectors all three give the same ranking (dot product is cheapest). For unnormalised vectors they differ; use what the model was trained with and configure the same metric in the vector database.
Why does semantic search miss exact matches like error codes?
Rare identifiers carry little learned meaning and get blended into a general vector. Keyword search (BM25) matches them exactly, which is why production retrieval uses hybrid search.
Practice¶
- Add five FAQ sentences to
CORPUS, rebuild the stand-in model, and run three searches of your own. - With a real model (
sentence-transformers), compare scores for "money back for broken items" against the refund and shipping sentences.
Next: Choosing & using embedding models — which model, how many dimensions, and what it costs.