Skip to content

2. Choosing & using embedding models

Intermediate · 12 min read

The embedding model decides what "similar" means in your system. A better model can lift retrieval quality more than any amount of prompt tuning — and the choice is sticky: change it later and you must re-embed everything.

Model names are examples

New embedding models appear every few months. Use the names below as starting points, then check the current leaderboard and your provider's docs.

2.1 The options

Type Examples Pros Cons
API models OpenAI text-embedding-3-small/large, Cohere Embed, Voyage, Google Gemini embeddings strong quality, zero infrastructure, multilingual per-token cost, data leaves your servers, rate limits
Open models BGE, E5, GTE, Nomic, Jina, all-MiniLM-L6-v2 (sentence-transformers) free to run, data stays local, can be fine-tuned you host them (CPU is fine for small models; GPU for volume)
Domain / multilingual code embeddings, legal/medical-tuned models, multilingual E5/BGE-M3 better on their domain or languages narrower

For Indian-language or mixed (Hinglish) content, test multilingual models specifically — English-only models quietly fail on it.

2.2 Benchmarks: MTEB

The MTEB (Massive Text Embedding Benchmark) leaderboard on Hugging Face scores models on retrieval, classification, clustering and more, across many languages. Use it to build a shortlist — look at the retrieval scores, the model size and the dimensions, not just the overall rank. As with LLM benchmarks, the deciding test is your data.

2.3 Evaluate on your own data

Build a small set of real queries with the chunk(s) that answer them, embed with each candidate, and compare recall@k and MRR (defined in NLP for GenAI). Here two stand-in "models" — word-level and character-level TF-IDF — play the candidates:

import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer

DOCS = [
    "Refunds are processed within 5 working days of approval.",
    "You can return damaged items for a full refund.",
    "Orders ship within 48 hours from our Pune warehouse.",
    "Express delivery is available in metro cities.",
    "Reset your password from Settings, then Security.",
    "Two-factor authentication protects your account.",
    "Invoices are emailed after every payment.",
    "We accept UPI, credit cards and net banking.",
]
EVAL = [  # (query, index of the relevant doc)
    ("how many days until I get my refund", 0),
    ("returning a broken product", 1),
    ("when will my order be shipped", 2),
    ("fast delivery to Mumbai", 3),
    ("forgot my password", 4),
    ("2FA login security", 5),
    ("where is my invoice", 6),
    ("can I pay with UPI", 7),
    ("refnd status", 0),                       # a typo, as real users type
]

def make_model(**tfidf_args):
    vec = TfidfVectorizer(sublinear_tf=True, **tfidf_args).fit(DOCS)
    def encode(texts):
        v = vec.transform(texts).toarray()
        return v / np.maximum(np.linalg.norm(v, axis=1, keepdims=True), 1e-9)
    return encode

def evaluate(encode, k=3):
    D = encode(DOCS)
    recall, rr = [], []
    for query, gold in EVAL:
        ranking = list(np.argsort(-(D @ encode([query])[0])))
        recall.append(gold in ranking[:k])
        rr.append(1 / (ranking.index(gold) + 1))
    return np.mean(recall), np.mean(rr)

for name, model in [("word-level", make_model(analyzer="word", stop_words="english")),
                    ("char-level", make_model(analyzer="char_wb", ngram_range=(3, 5)))]:
    r, mrr = evaluate(model)
    print(f"{name:10}  recall@3={r:.2f}  MRR={mrr:.2f}")
Output
word-level  recall@3=0.78  MRR=0.63
char-level  recall@3=0.78  MRR=0.77

Recall@3 is a tie, but MRR shows the difference: the character-level model puts the right document first for 6 of the 9 queries, the word-level one for only 4 — it matches word variants ("shipped" vs "ship", "invoice" vs "Invoices"). Look closer and one of the word-level model's "hits" is luck: it knows no word in "refnd status", so every score is 0 and the first document wins by default. That's why you check more than one metric, and read individual failures.

With real candidates the loop is identical — swap make_model(...) for an API or sentence-transformers call. Thirty to a hundred queries are usually enough to see which model fits your data.

2.4 Dimensions and Matryoshka embeddings

More dimensions can hold more nuance, but cost more memory and slower search:

Dimensions float32 size per vector 10M vectors
384 1.5 KB 15 GB
768 3 KB 31 GB
1,536 6 KB 61 GB
3,072 12 KB 123 GB

Many recent models are trained as Matryoshka embeddings: the most important information is packed into the first dimensions, so you can cut a 3,072-dim vector down to 256 or 512 and re-normalise, losing little quality. OpenAI exposes this as a dimensions parameter. The mechanics:

def truncate(vectors: np.ndarray, dims: int) -> np.ndarray:
    v = vectors[:, :dims]
    return v / np.linalg.norm(v, axis=1, keepdims=True)       # re-normalise after cutting

full = np.random.default_rng(0).normal(size=(4, 3072))
small = truncate(full, 256)
print(full.shape, "→", small.shape, f"({small.nbytes / full.nbytes:.0%} of the memory)")
Output
(4, 3072) → (4, 256) (8% of the memory)

Only do this with models trained for it (check the model card), and measure recall at the reduced size on your eval set. A common pattern: search with short vectors, then re-score the top 100 with the full vectors.

2.5 Batching, caching and cost

Embedding happens twice: once per chunk at ingestion (bulk), and once per query (latency-sensitive).

# no-run — batching with the OpenAI API
from openai import OpenAI
client = OpenAI()

def embed_batch(texts: list[str], model="text-embedding-3-small", batch_size=256) -> list[list[float]]:
    out = []
    for i in range(0, len(texts), batch_size):          # one request per batch, not per text
        resp = client.embeddings.create(model=model, input=texts[i:i + batch_size])
        out.extend(d.embedding for d in resp.data)
    return out

Never embed the same text twice — cache by a hash of model + text:

import hashlib

class EmbeddingCache:
    def __init__(self, encode, model_name):
        self.encode, self.model_name, self.store, self.calls = encode, model_name, {}, 0

    def key(self, text):
        return hashlib.sha256(f"{self.model_name}\x00{text}".encode()).hexdigest()

    def get(self, texts):
        missing = [t for t in dict.fromkeys(texts) if self.key(t) not in self.store]
        if missing:
            self.calls += 1
            for t, v in zip(missing, self.encode(missing)):
                self.store[self.key(t)] = v
        return np.array([self.store[self.key(t)] for t in texts])

cache = EmbeddingCache(make_model(analyzer="char_wb", ngram_range=(3, 5)), "char-tfidf-v1")
cache.get(DOCS)
cache.get(DOCS[:4] + ["A new FAQ entry about gift cards."])     # only the new text is embedded
print(len(cache.store), "vectors cached,", cache.calls, "embedding calls")
Output
9 vectors cached, 2 embedding calls

Cost is usually small next to LLM generation, but worth estimating:

def embedding_cost(n_chunks, tokens_per_chunk, usd_per_million):
    return n_chunks * tokens_per_chunk * usd_per_million / 1e6

# 200,000 chunks of ~400 tokens at an illustrative $0.02 per 1M tokens
print(f"one-time ingestion: ${embedding_cost(200_000, 400, 0.02):.2f}")
print(f"re-embedding with a 6.5x pricier model: ${embedding_cost(200_000, 400, 0.13):.2f}")
Output
one-time ingestion: $1.60
re-embedding with a 6.5x pricier model: $10.40

Embedding cost is rarely the problem. Storage, memory and re-indexing time at scale are — next section and topic 6.

2.6 Storage: compress the vectors

def storage_gb(n_vectors, dims, bytes_per_value):
    return n_vectors * dims * bytes_per_value / 1e9

n, d = 10_000_000, 1536
for name, bpv in [("float32", 4), ("float16", 2), ("int8", 1), ("binary", 1 / 8)]:
    print(f"{name:8} {storage_gb(n, d, bpv):6.1f} GB")
Output
float32    61.4 GB
float16    30.7 GB
int8       15.4 GB
binary      1.9 GB

Quantised vectors (int8, binary) trade a little accuracy for big memory savings; how that works, and how to win the accuracy back, is in Vector search & indexes.

2.7 Version your embeddings

Vectors from different models (or even different versions of one model) are not comparable — mixing them silently destroys search quality. Store the model name and version with every vector, and treat a model change as a migration:

  • Record embedding_model (e.g. text-embedding-3-small@1536) in each record's metadata or in the index name.
  • Embed queries with exactly the model that embedded the documents.
  • To switch models: build a new index in the background, re-embed everything, evaluate, then switch traffic — never re-embed in place (details in Vector DBs in production).

Interview questions

How do you choose an embedding model?

Shortlist from MTEB retrieval scores, languages, size and dimensions, and constraints (data residency → open model; no infra → API). Then evaluate candidates on your own queries with recall@k and MRR, and weigh cost, latency, memory and multilingual needs. Remember switching later means re-embedding everything.

What are Matryoshka embeddings?

Embeddings trained so that prefixes of the vector are themselves good embeddings. You can truncate (e.g. 3,072 → 256 dims) and re-normalise to cut memory and speed up search with small quality loss — useful for a fast first pass followed by full-vector re-scoring.

What happens if you change the embedding model?

Old and new vectors live in different spaces and can't be compared, so every document must be re-embedded into a new index. Do it as a migration: build in parallel, evaluate, switch traffic, keep the old index for rollback.

Practice

  • Add five queries from your own project to EVAL and compare two real embedding models.
  • Estimate storage for your corpus at 384, 1,536 and 3,072 dimensions in float32 and int8.

Next: Vector search & indexes — how search stays fast at millions of vectors.