Skip to content

1. Embeddings in depth

Intermediate · 12 min read

An embedding is a list of numbers (a vector, typically 384–3,072 long) that represents the meaning of a piece of text — or an image, or a product. Texts with similar meaning get vectors that point in similar directions. That one property powers semantic search, RAG retrieval, clustering, deduplication, recommendations and classification.

1.1 A stand-in embedding model

Real embedding models are neural networks (an API call or a ~100 MB download). To keep every example on these pages runnable offline, we use a small stand-in: TF-IDF over character n-grams, normalised to unit length. It captures spelling similarity, not real meaning — good enough to demonstrate the mechanics, and the code has the same shape as a real model.

import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer

CORPUS = [
    "Refunds are processed within 5 working days of approval.",
    "You can return damaged items for a full refund.",
    "Orders ship within 48 hours from our Pune warehouse.",
    "Express delivery is available in metro cities.",
    "Reset your password from Settings, then Security.",
    "Two-factor authentication protects your account.",
    "Invoices are emailed after every payment.",
    "We accept UPI, credit cards and net banking.",
]

class StandInEmbedder:
    """Same interface as a real model: .encode(list_of_texts) -> (n, dim) array of unit vectors."""
    def __init__(self, corpus):
        self.tfidf = TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5), sublinear_tf=True).fit(corpus)

    def encode(self, texts):
        v = self.tfidf.transform(texts).toarray()
        return v / np.maximum(np.linalg.norm(v, axis=1, keepdims=True), 1e-9)

model = StandInEmbedder(CORPUS)
E = model.encode(CORPUS)
print(E.shape, np.linalg.norm(E, axis=1).round(3)[:3])
Output
(8, 710) [1. 1. 1.]

The stand-in's dimension is the number of distinct character n-grams in the corpus. A real model always returns the same fixed size — e.g. 384 or 1,536 — whatever the input.

The real thing has the same shape:

# no-run — pip install sentence-transformers (downloads ~90 MB once)
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
E = model.encode(CORPUS, normalize_embeddings=True)       # (8, 384)

1.2 Search = nearest neighbours

With unit-length vectors, cosine similarity is a dot product, and one matrix multiply scores every document:

def search(query: str, k: int = 3):
    q = model.encode([query])[0]
    scores = E @ q
    for i in np.argsort(-scores)[:k]:
        print(f"{scores[i]:.2f}  {CORPUS[i]}")

search("how long do refunds take")
Output
0.37  Refunds are processed within 5 working days of approval.
0.23  You can return damaged items for a full refund.
0.03  Orders ship within 48 hours from our Pune warehouse.

That is the whole idea behind vector search. Everything else on these pages is about doing it well (better embeddings) and fast (millions of vectors).

1.3 How embedding models learn meaning

Embedding models are transformers (usually encoder-only, like BERT) fine-tuned with contrastive learning:

flowchart LR
    Q["query: 'refund time?'"] --> M[Same model]
    P["matching passage:<br/>'Refunds take 5 days'"] --> M
    N["other passages in the batch<br/>(negatives)"] --> M
    M --> L["Loss: pull query and its passage together,<br/>push the negatives apart"]

Training data is millions of (query, relevant passage) pairs — search logs, question/answer sites, title/body pairs, and increasingly LLM-generated pairs. Each training step makes matching pairs more similar and everything else in the batch less similar. The geometry that results is what makes "money back for broken items" land near "refund for damaged products" with no shared words.

def contrastive_loss(q, passages, positive_index, temperature=0.05):
    """InfoNCE: softmax over similarities; the loss is low when the true passage wins clearly."""
    logits = passages @ q / temperature
    logits -= logits.max()
    probs = np.exp(logits) / np.exp(logits).sum()
    return float(-np.log(probs[positive_index]))

q = model.encode(["how long do refunds take"])[0]
print("refund passage is the positive:  ", round(contrastive_loss(q, E, positive_index=0), 3))
print("shipping passage is the positive:", round(contrastive_loss(q, E, positive_index=2), 3))
Output
refund passage is the positive:   0.073
shipping passage is the positive: 6.906

Training adjusts the model's weights to make the first kind of number small for every real pair.

Hard negatives — passages that look similar but are wrong ("refunds for digital products are not available") — are what teach a model fine distinctions. They are also how you fine-tune an embedding model on your own domain.

1.4 Similarity metrics

Metric Formula Notes
Cosine similarity a·b / (‖a‖‖b‖) direction only; the default for text
Dot product a·b equals cosine for unit vectors, and is faster; most DBs' fastest option
Euclidean (L2) distance ‖a − b‖ for unit vectors, ranks exactly like cosine: ‖a−b‖² = 2 − 2·cos
a, b = E[0], E[1]
cos = a @ b
l2 = np.linalg.norm(a - b)
print(round(float(cos), 4), round(float(l2 ** 2), 4), round(float(2 - 2 * cos), 4))
Output
0.105 1.7901 1.7901

So for normalised vectors the choice of metric doesn't change the ranking. It matters when vectors aren't normalised: then dot product also rewards vector length, which some models use on purpose. Rule: use the metric in the model's documentation, and configure the same metric in your vector database.

1.5 Query vs document embeddings

Questions and passages look different — "refund time?" vs a paragraph of policy text. Many models are trained to embed them asymmetrically and expect a hint:

Model family How you mark it
E5 prefix "query: " or "passage: "
BGE instruction prefix for queries, e.g. "Represent this sentence for searching relevant passages: "
Nomic, some others "search_query: " / "search_document: "
Cohere, Voyage, Gemini APIs an input_type / task_type parameter (search_query vs search_document)
OpenAI text-embedding-3-* no prefix needed

Forgetting the prefix is a silent bug: everything still works, just with noticeably worse retrieval.

1.6 What embeddings miss

Embeddings compress meaning into one vector, and some things get lost:

Weakness Example Fix
Exact identifiers error code E-4012, SKU TSH-COT-RED-M, a person's name hybrid search with BM25 (NLP notes)
Negation "flights without a layover" retrieves layover pages rerankers; filters; query rewriting
Numbers and ranges "phones under ₹20,000" metadata filters (price < 20000), not vectors
Long texts a 20-page document becomes one blurry vector chunk before embedding
Domain jargon internal acronyms the model never saw domain-tuned or fine-tuned embedding models; glossaries
Freshness / permissions "latest policy", "documents I can see" metadata (updated_at, tenant_id) + filters
search("error E-4012 in checkout", k=2)
Output
0.13  Express delivery is available in metro cities.
0.04  Invoices are emailed after every payment.

Nothing in the corpus mentions that error, so the "nearest" vectors are still returned — with low scores. Vector search always returns something; your app must decide when a result is too weak to use (a score threshold, or a reranker).

Interview questions

What is an embedding and how is it created?

A fixed-length dense vector representing the meaning of an input. Text embedding models are transformer encoders whose token outputs are pooled into one vector, fine-tuned with contrastive learning on (query, relevant passage) pairs so that related texts end up close together and unrelated ones far apart.

Cosine, dot product or Euclidean?

For normalised vectors all three give the same ranking (dot product is cheapest). For unnormalised vectors they differ; use what the model was trained with and configure the same metric in the vector database.

Why does semantic search miss exact matches like error codes?

Rare identifiers carry little learned meaning and get blended into a general vector. Keyword search (BM25) matches them exactly, which is why production retrieval uses hybrid search.

Practice

  • Add five FAQ sentences to CORPUS, rebuild the stand-in model, and run three searches of your own.
  • With a real model (sentence-transformers), compare scores for "money back for broken items" against the refund and shipping sentences.

Next: Choosing & using embedding models — which model, how many dimensions, and what it costs.