Skip to content

5. NumPy for GenAI

Intermediate · 12 min read

Everything here is what a vector database does under the hood — small enough to understand, and enough to build semantic search for thousands of documents with no extra library.

5.1 An embeddings matrix

An embedding model turns each text into a vector. Stack them: one row per document.

import numpy as np

rng = np.random.default_rng(7)
docs = ["refund policy", "shipping times", "password reset", "pricing plans", "api limits"]
E = rng.normal(size=(len(docs), 8)).astype(np.float32)   # 5 docs x 8 dims (real models: 384–3072)
print(E.shape, E.dtype)
Output
(5, 8) float32

In a real app the rows come from an embeddings API:

# no-run — needs an API key
from openai import OpenAI
client = OpenAI()
resp = client.embeddings.create(model="text-embedding-3-small", input=docs)
E = np.array([d.embedding for d in resp.data], dtype=np.float32)

5.2 Cosine similarity

Cosine similarity compares direction, not length: 1 = same meaning, 0 = unrelated, −1 = opposite.

def cosine(a, b):
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))

print(round(cosine(np.array([1, 0]), np.array([1, 0])), 3))   # same direction
print(round(cosine(np.array([1, 0]), np.array([0, 1])), 3))   # 90°
print(round(cosine(np.array([1, 0]), np.array([1, 1])), 3))   # 45°
Output
1.0
0.0
0.707

5.3 Search: one query against every document

Normalise once; then one matrix multiply scores every document.

def normalise(x):
    return x / np.linalg.norm(x, axis=-1, keepdims=True)

E_unit = normalise(E)                          # do this once, when you index
query = E[2] + rng.normal(scale=0.3, size=8)   # a query close to "password reset"
scores = E_unit @ normalise(query)             # cosine score per document

order = np.argsort(-scores)                    # best first
for i in order[:3]:
    print(f"{scores[i]:.3f}  {docs[i]}")
Output
0.976  password reset
0.717  pricing plans
0.475  refund policy

5.4 Top-k fast at scale — argpartition

Sorting a million scores just to get the best 5 wastes time. argpartition finds the top-k without fully sorting, then you sort only those k.

big = rng.normal(size=1_000_000)
k = 5
top = np.argpartition(-big, k)[:k]          # the k best, in any order
top = top[np.argsort(-big[top])]            # sort just those k
print(big[top].round(3))
print(np.array_equal(big[top], np.sort(big)[::-1][:k]))   # same answer as a full sort
Output
[4.948 4.584 4.545 4.523 4.366]
True

5.5 Many queries at once

Q = normalise(rng.normal(size=(3, 8)))      # 3 queries
S = Q @ E_unit.T                            # (3 queries x 5 docs)
print(S.shape)
print(S.argmax(axis=1))                     # best document for each query
Output
(3, 5)
[1 1 0]

5.6 Softmax and temperature

Models turn raw scores (logits) into probabilities with softmax. Temperature makes the distribution sharper (low) or flatter (high) — the same knob you set when calling an LLM.

def softmax(logits, temperature=1.0):
    z = np.asarray(logits) / temperature
    z = z - z.max()                  # numerical stability: avoid huge exp()
    p = np.exp(z)
    return p / p.sum()

logits = np.array([2.0, 1.0, 0.1])
for t in (0.5, 1.0, 2.0):
    print(f"T={t}:", softmax(logits, t).round(3))
Output
T=0.5: [0.864 0.117 0.019]
T=1.0: [0.659 0.242 0.099]
T=2.0: [0.502 0.304 0.194]

5.7 Saving and loading vectors

np.save("embeddings.npy", E)                # fast binary file
loaded = np.load("embeddings.npy")
print(np.array_equal(E, loaded), loaded.shape)

np.savez_compressed("index.npz", vectors=E, ids=np.arange(len(docs)))
z = np.load("index.npz")
print(sorted(z.files), z["ids"])
Output
True (5, 8)
['ids', 'vectors'] [0 1 2 3 4]

Memory: use float32

1 million vectors × 1,536 dims × 4 bytes (float32) ≈ 6 GB; with float64 it's 12 GB. Beyond a few hundred thousand vectors, move to a vector database (FAISS, pgvector, Pinecone) — they use exactly these operations plus smarter indexes.

Practice

  • Add a 6th document to E, re-normalise, and search again.
  • Write top_k(query, E_unit, k) that returns (index, score) pairs using argpartition.

Next: Pandas — the table side of data work.