Skip to content

3. Text representation

Intermediate · 12 min read

Computers compare numbers, not words. Every search engine, classifier and RAG system first turns text into a vector. This page walks from the simplest representation to the one LLM apps use — and shows why the older ones are still useful.

3.1 Bag of words

Count how often each vocabulary word appears. Word order is thrown away — hence "bag".

from sklearn.feature_extraction.text import CountVectorizer

docs = [
    "refund processed in days",
    "refund policy for damaged items",
    "shipping takes days",
]
bow = CountVectorizer()
X = bow.fit_transform(docs)                       # sparse matrix: docs × vocabulary
print(bow.get_feature_names_out())
print(X.toarray())
Output
['damaged' 'days' 'for' 'in' 'items' 'policy' 'processed' 'refund'
 'shipping' 'takes']
[[0 1 0 1 0 0 1 1 0 0]
 [1 0 1 0 1 1 0 1 0 0]
 [0 1 0 0 0 0 0 0 1 1]]

Problems: common words get the same weight as rare, informative ones, and meaning is invisible — "refund" and "money back" share no columns.

3.2 TF-IDF — weight words by how informative they are

TF-IDF = term frequency × inverse document frequency. A word scores high when it is frequent in this document but rare across the collection.

tfidf(word, doc) = tf(word, doc) × idf(word)
idf(word)        = log((1 + N) / (1 + documents containing word)) + 1        # scikit-learn's version
from sklearn.feature_extraction.text import TfidfVectorizer
import numpy as np

tfidf = TfidfVectorizer(stop_words="english")       # drop "for", "the", … — fine for keyword search
T = tfidf.fit_transform(docs)
vocab = tfidf.get_feature_names_out()
for i, d in enumerate(docs):
    row = T[i].toarray().ravel()
    top = np.argsort(-row)[:2]
    print(f"{d!r:36} → {[(vocab[j], round(float(row[j]), 2)) for j in top]}")
Output
'refund processed in days'           → [('processed', 0.68), ('days', 0.52)]
'refund policy for damaged items'    → [('damaged', 0.53), ('items', 0.53)]
'shipping takes days'                → [('shipping', 0.62), ('takes', 0.62)]

"processed", "damaged" and "shipping" appear in only one document, so they define it; "refund" and "days" appear in two, so they count for less. (stop_words="english" also removed "in" and "for".)

Searching with TF-IDF

from sklearn.metrics.pairwise import cosine_similarity

query = tfidf.transform(["how many days for a refund"])
scores = cosine_similarity(query, T).ravel()
for i in np.argsort(-scores):
    print(f"{scores[i]:.2f}  {docs[i]}")
Output
0.73  refund processed in days
0.33  shipping takes days
0.28  refund policy for damaged items

n-grams — keep a little word order

bi = TfidfVectorizer(ngram_range=(1, 2))
bi.fit(["not good", "good not bad"])
print(bi.get_feature_names_out())
Output
['bad' 'good' 'good not' 'not' 'not bad' 'not good']

With bigrams, "not good" becomes its own feature — enough to fix many negation mistakes in classic classifiers.

TF-IDF is still worth knowing

It needs no GPU, no API and no training; it's explainable (you can see which words matched); and it is excellent at exact terms — product codes, error messages, names — which embeddings often miss. Its cousin BM25 is the keyword half of hybrid search (see NLP for GenAI).

3.3 The vocabulary gap

q = tfidf.transform(["money back for broken products"])
print(cosine_similarity(q, T).round(2))
Output
[[0. 0. 0.]]

"money back" means refund and "broken" means damaged, yet every score is zero — not one word matches. Counting words can't see synonyms. Fixing this needs meaning, which is what embeddings capture.

3.4 Word embeddings

Word2Vec (2013) and GloVe learned a dense vector (typically 100–300 numbers) per word by predicting neighbouring words. Words used in similar contexts end up close together — and directions in the space carry meaning. A tiny hand-made example of the famous analogy:

# dims: [royalty, maleness, person, fruit]
emb = {
    "king":  np.array([0.95, 0.90, 1.0, 0.0]),
    "queen": np.array([0.95, 0.05, 1.0, 0.0]),
    "man":   np.array([0.05, 0.90, 1.0, 0.0]),
    "woman": np.array([0.05, 0.05, 1.0, 0.0]),
    "apple": np.array([0.00, 0.00, 0.0, 1.0]),
}

def cos(a, b):
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))

target = emb["king"] - emb["man"] + emb["woman"]
print(max((w for w in emb if w not in {"king", "man", "woman"}), key=lambda w: cos(target, emb[w])))
print(round(cos(emb["king"], emb["queen"]), 2), round(cos(emb["king"], emb["apple"]), 2))
Output
queen
0.86 0.0

Real embeddings have hundreds of dimensions that are learned, not labelled — but the geometry works the same way.

Limitation: one vector per word, whatever the context. "bank" in river bank and bank account gets the same vector.

3.5 Contextual and sentence embeddings

Transformer models (BERT, then today's embedding models) produce vectors that depend on context, and can embed a whole sentence or chunk into one vector. This is what RAG uses.

# no-run — needs `pip install sentence-transformers` (downloads a ~90 MB model)
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")             # 384-dim, runs on CPU
docs = ["refund processed in days", "refund policy for damaged items", "shipping takes days"]
D = model.encode(docs, normalize_embeddings=True)
q = model.encode(["money back for broken products"], normalize_embeddings=True)
print((D @ q.T).ravel().round(2))
# the "damaged items" policy now ranks first — with zero words in common with the query
# no-run — needs an API key
from openai import OpenAI
resp = OpenAI().embeddings.create(model="text-embedding-3-small", input=docs)
D = np.array([d.embedding for d in resp.data])               # 1536 dims, already normalised

3.6 Choosing a representation

Bag of words / TF-IDF Word embeddings Sentence embeddings
Vector sparse, vocabulary-sized dense, 100–300 per word dense, 384–3072 per text
Understands synonyms ❌ partly ✅
Understands context / word order ❌ (n-grams: a little) ❌ ✅
Exact terms, codes, names ✅ excellent weak often weak
Cost ~free, CPU cheap model or API call per text
Explainable ✅ partly ❌
Typical use today keyword search (BM25), baselines, features rarely used directly semantic search, RAG, clustering, dedup

The best production retrieval uses both: keyword scores for exact matches + embeddings for meaning — that's hybrid search, covered in NLP for GenAI.

3.7 Similarity measures

a, b = np.array([1.0, 2.0, 3.0]), np.array([2.0, 4.0, 6.0])     # same direction, b is twice as long
print("cosine   ", round(cos(a, b), 3))
print("dot      ", a @ b)
print("euclidean", round(float(np.linalg.norm(a - b)), 3))
Output
cosine    1.0
dot       28.0
euclidean 3.742
  • Cosine — direction only; the default for text embeddings.
  • Dot product — equals cosine when vectors are normalised (length 1), and is faster. Most embedding APIs return normalised vectors, so vector DBs often use dot product.
  • Euclidean — straight-line distance; sensitive to length.

Use the metric the embedding model was trained for (it's in the model card) and set the same one in your vector database.

Practice

  • Fit a TfidfVectorizer(stop_words="english") on ten FAQ answers and build a tiny search function.
  • Add the query "money back" and see TF-IDF fail; note which documents an embedding model would need to find.

Next: Core NLP tasks — classification, NER and summarisation, classic vs LLM.