Transformers¶
Intermediate · 5 topics
Every LLM — GPT, Claude, Gemini, Llama — and every embedding model is a transformer. You don't need to train one from scratch to build GenAI apps, but knowing how it works explains almost everything you meet in practice: why context windows are limited, why long prompts are slow and expensive, what temperature really does, why a KV cache matters, and what LoRA or quantisation changes. It is also one of the most asked topics in GenAI interviews.
Every piece here is built in a few lines of NumPy, so you can run it and see the numbers.
| # | Topic | Sub-topics |
|---|---|---|
| 1 | Why transformers | Sequence models before · RNN limits · the big idea · encoder, decoder and encoder–decoder |
| 2 | Self-attention from scratch | Queries, keys & values · scaled dot-product · causal mask · multi-head attention |
| 3 | The transformer block | Token & position embeddings · RoPE · layer norm · residuals · feed-forward · counting parameters |
| 4 | How LLMs generate text | Next-token prediction · logits · greedy, temperature, top-k, top-p · KV cache · cost of long context |
| 5 | Transformers for GenAI | Pre-training → SFT → RLHF/DPO · BERT vs GPT vs T5 · Hugging Face · LoRA · quantisation |
flowchart LR
T[Text] --> K[Tokens] --> E[Embeddings + positions]
E --> B1[Transformer block 1]
B1 --> B2[...]
B2 --> BN[Transformer block N]
BN --> L[Logits over vocabulary] --> S[Sample next token]
S -->|append and repeat| K
Before you start
You'll get the most from this series after Tokenization,
Text representation and NumPy linear algebra
— matrix multiply (@) and softmax are all the maths you need.
Next: LLMs — using models well in real applications.