Skip to content

Transformers

Intermediate · 5 topics

Every LLM — GPT, Claude, Gemini, Llama — and every embedding model is a transformer. You don't need to train one from scratch to build GenAI apps, but knowing how it works explains almost everything you meet in practice: why context windows are limited, why long prompts are slow and expensive, what temperature really does, why a KV cache matters, and what LoRA or quantisation changes. It is also one of the most asked topics in GenAI interviews.

Every piece here is built in a few lines of NumPy, so you can run it and see the numbers.

pip install numpy          # that's all — the optional Hugging Face examples also need transformers + torch
# Topic Sub-topics
1 Why transformers Sequence models before · RNN limits · the big idea · encoder, decoder and encoder–decoder
2 Self-attention from scratch Queries, keys & values · scaled dot-product · causal mask · multi-head attention
3 The transformer block Token & position embeddings · RoPE · layer norm · residuals · feed-forward · counting parameters
4 How LLMs generate text Next-token prediction · logits · greedy, temperature, top-k, top-p · KV cache · cost of long context
5 Transformers for GenAI Pre-training → SFT → RLHF/DPO · BERT vs GPT vs T5 · Hugging Face · LoRA · quantisation
flowchart LR
    T[Text] --> K[Tokens] --> E[Embeddings + positions]
    E --> B1[Transformer block 1]
    B1 --> B2[...]
    B2 --> BN[Transformer block N]
    BN --> L[Logits over vocabulary] --> S[Sample next token]
    S -->|append and repeat| K

Before you start

You'll get the most from this series after Tokenization, Text representation and NumPy linear algebra — matrix multiply (@) and softmax are all the maths you need.

Next: LLMs — using models well in real applications.