NLP¶
Beginner → Intermediate · 5 topics
NLP (Natural Language Processing) is the field LLMs came out of. You don't need the whole textbook to build GenAI systems, but you do need the parts that show up every day: cleaning text before indexing, counting tokens for cost and context limits, representing text as vectors, knowing when a classic model beats an LLM call, chunking documents, keyword + vector search, and measuring whether answers are any good. This series covers exactly that — with code that runs.
| # | Topic | Sub-topics |
|---|---|---|
| 1 | Text preprocessing | Unicode & whitespace · HTML · regex · PII masking · stopwords, stemming & lemmas — and when to skip them |
| 2 | Tokenization | Words vs sub-words · BPE from scratch · tiktoken · counting tokens & cost · fitting the context window |
| 3 | Text representation | Bag of words · TF-IDF · n-grams · word embeddings · sentence embeddings · cosine similarity |
| 4 | Core NLP tasks | Classification · sentiment · NER · summarisation · classic models vs LLMs |
| 5 | NLP for GenAI | Chunking strategies · BM25 · hybrid search with RRF · precision/recall/F1 · ROUGE · recall@k & MRR |
flowchart LR
T[Raw text] --> P[1. Clean]
P --> K[2. Tokenize / count]
K --> C[5. Chunk]
C --> R[3. Represent: TF-IDF / embeddings]
R --> S[5. Search: BM25 + vectors]
S --> L[LLM answer]
L --> E[5. Evaluate]
Already comfortable with NumPy?
Cosine similarity, top-k search and softmax are covered with code in NumPy for GenAI — this series builds on them.
Next: Transformers — the architecture inside every LLM.