3. Text representation¶
Intermediate · 12 min read
Computers compare numbers, not words. Every search engine, classifier and RAG system first turns text into a vector. This page walks from the simplest representation to the one LLM apps use — and shows why the older ones are still useful.
3.1 Bag of words¶
Count how often each vocabulary word appears. Word order is thrown away — hence "bag".
from sklearn.feature_extraction.text import CountVectorizer
docs = [
"refund processed in days",
"refund policy for damaged items",
"shipping takes days",
]
bow = CountVectorizer()
X = bow.fit_transform(docs) # sparse matrix: docs × vocabulary
print(bow.get_feature_names_out())
print(X.toarray())
['damaged' 'days' 'for' 'in' 'items' 'policy' 'processed' 'refund'
'shipping' 'takes']
[[0 1 0 1 0 0 1 1 0 0]
[1 0 1 0 1 1 0 1 0 0]
[0 1 0 0 0 0 0 0 1 1]]
Problems: common words get the same weight as rare, informative ones, and meaning is invisible — "refund" and "money back" share no columns.
3.2 TF-IDF — weight words by how informative they are¶
TF-IDF = term frequency × inverse document frequency. A word scores high when it is frequent in this document but rare across the collection.
tfidf(word, doc) = tf(word, doc) × idf(word)
idf(word) = log((1 + N) / (1 + documents containing word)) + 1 # scikit-learn's version
from sklearn.feature_extraction.text import TfidfVectorizer
import numpy as np
tfidf = TfidfVectorizer(stop_words="english") # drop "for", "the", … — fine for keyword search
T = tfidf.fit_transform(docs)
vocab = tfidf.get_feature_names_out()
for i, d in enumerate(docs):
row = T[i].toarray().ravel()
top = np.argsort(-row)[:2]
print(f"{d!r:36} → {[(vocab[j], round(float(row[j]), 2)) for j in top]}")
'refund processed in days' → [('processed', 0.68), ('days', 0.52)]
'refund policy for damaged items' → [('damaged', 0.53), ('items', 0.53)]
'shipping takes days' → [('shipping', 0.62), ('takes', 0.62)]
"processed", "damaged" and "shipping" appear in only one document, so they define it; "refund" and
"days" appear in two, so they count for less. (stop_words="english" also removed "in" and "for".)
Searching with TF-IDF¶
from sklearn.metrics.pairwise import cosine_similarity
query = tfidf.transform(["how many days for a refund"])
scores = cosine_similarity(query, T).ravel()
for i in np.argsort(-scores):
print(f"{scores[i]:.2f} {docs[i]}")
n-grams — keep a little word order¶
bi = TfidfVectorizer(ngram_range=(1, 2))
bi.fit(["not good", "good not bad"])
print(bi.get_feature_names_out())
With bigrams, "not good" becomes its own feature — enough to fix many negation mistakes in classic classifiers.
TF-IDF is still worth knowing
It needs no GPU, no API and no training; it's explainable (you can see which words matched); and it is excellent at exact terms — product codes, error messages, names — which embeddings often miss. Its cousin BM25 is the keyword half of hybrid search (see NLP for GenAI).
3.3 The vocabulary gap¶
"money back" means refund and "broken" means damaged, yet every score is zero — not one word matches. Counting words can't see synonyms. Fixing this needs meaning, which is what embeddings capture.
3.4 Word embeddings¶
Word2Vec (2013) and GloVe learned a dense vector (typically 100–300 numbers) per word by predicting neighbouring words. Words used in similar contexts end up close together — and directions in the space carry meaning. A tiny hand-made example of the famous analogy:
# dims: [royalty, maleness, person, fruit]
emb = {
"king": np.array([0.95, 0.90, 1.0, 0.0]),
"queen": np.array([0.95, 0.05, 1.0, 0.0]),
"man": np.array([0.05, 0.90, 1.0, 0.0]),
"woman": np.array([0.05, 0.05, 1.0, 0.0]),
"apple": np.array([0.00, 0.00, 0.0, 1.0]),
}
def cos(a, b):
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
target = emb["king"] - emb["man"] + emb["woman"]
print(max((w for w in emb if w not in {"king", "man", "woman"}), key=lambda w: cos(target, emb[w])))
print(round(cos(emb["king"], emb["queen"]), 2), round(cos(emb["king"], emb["apple"]), 2))
Real embeddings have hundreds of dimensions that are learned, not labelled — but the geometry works the same way.
Limitation: one vector per word, whatever the context. "bank" in river bank and bank account gets the same vector.
3.5 Contextual and sentence embeddings¶
Transformer models (BERT, then today's embedding models) produce vectors that depend on context, and can embed a whole sentence or chunk into one vector. This is what RAG uses.
# no-run — needs `pip install sentence-transformers` (downloads a ~90 MB model)
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2") # 384-dim, runs on CPU
docs = ["refund processed in days", "refund policy for damaged items", "shipping takes days"]
D = model.encode(docs, normalize_embeddings=True)
q = model.encode(["money back for broken products"], normalize_embeddings=True)
print((D @ q.T).ravel().round(2))
# the "damaged items" policy now ranks first — with zero words in common with the query
# no-run — needs an API key
from openai import OpenAI
resp = OpenAI().embeddings.create(model="text-embedding-3-small", input=docs)
D = np.array([d.embedding for d in resp.data]) # 1536 dims, already normalised
3.6 Choosing a representation¶
| Bag of words / TF-IDF | Word embeddings | Sentence embeddings | |
|---|---|---|---|
| Vector | sparse, vocabulary-sized | dense, 100–300 per word | dense, 384–3072 per text |
| Understands synonyms | ❌ | partly | ✅ |
| Understands context / word order | ❌ (n-grams: a little) | ❌ | ✅ |
| Exact terms, codes, names | ✅ excellent | weak | often weak |
| Cost | ~free, CPU | cheap | model or API call per text |
| Explainable | ✅ | partly | ❌ |
| Typical use today | keyword search (BM25), baselines, features | rarely used directly | semantic search, RAG, clustering, dedup |
The best production retrieval uses both: keyword scores for exact matches + embeddings for meaning — that's hybrid search, covered in NLP for GenAI.
3.7 Similarity measures¶
a, b = np.array([1.0, 2.0, 3.0]), np.array([2.0, 4.0, 6.0]) # same direction, b is twice as long
print("cosine ", round(cos(a, b), 3))
print("dot ", a @ b)
print("euclidean", round(float(np.linalg.norm(a - b)), 3))
- Cosine — direction only; the default for text embeddings.
- Dot product — equals cosine when vectors are normalised (length 1), and is faster. Most embedding APIs return normalised vectors, so vector DBs often use dot product.
- Euclidean — straight-line distance; sensitive to length.
Use the metric the embedding model was trained for (it's in the model card) and set the same one in your vector database.
Practice¶
- Fit a
TfidfVectorizer(stop_words="english")on ten FAQ answers and build a tiny search function. - Add the query "money back" and see TF-IDF fail; note which documents an embedding model would need to find.
Next: Core NLP tasks — classification, NER and summarisation, classic vs LLM.