Skip to content

NLP

Beginner → Intermediate · 5 topics

NLP (Natural Language Processing) is the field LLMs came out of. You don't need the whole textbook to build GenAI systems, but you do need the parts that show up every day: cleaning text before indexing, counting tokens for cost and context limits, representing text as vectors, knowing when a classic model beats an LLM call, chunking documents, keyword + vector search, and measuring whether answers are any good. This series covers exactly that — with code that runs.

pip install scikit-learn tiktoken rank-bm25 numpy
# Topic Sub-topics
1 Text preprocessing Unicode & whitespace · HTML · regex · PII masking · stopwords, stemming & lemmas — and when to skip them
2 Tokenization Words vs sub-words · BPE from scratch · tiktoken · counting tokens & cost · fitting the context window
3 Text representation Bag of words · TF-IDF · n-grams · word embeddings · sentence embeddings · cosine similarity
4 Core NLP tasks Classification · sentiment · NER · summarisation · classic models vs LLMs
5 NLP for GenAI Chunking strategies · BM25 · hybrid search with RRF · precision/recall/F1 · ROUGE · recall@k & MRR
flowchart LR
    T[Raw text] --> P[1. Clean]
    P --> K[2. Tokenize / count]
    K --> C[5. Chunk]
    C --> R[3. Represent: TF-IDF / embeddings]
    R --> S[5. Search: BM25 + vectors]
    S --> L[LLM answer]
    L --> E[5. Evaluate]

Already comfortable with NumPy?

Cosine similarity, top-k search and softmax are covered with code in NumPy for GenAI — this series builds on them.

Next: Transformers — the architecture inside every LLM.