1. Why transformers¶
Beginner · 8 min read
The 2017 paper "Attention Is All You Need" introduced the transformer for machine translation. Within a few years it replaced almost every other architecture in NLP — and then in vision, speech and code. To see why, look at what came before.
1.1 Before transformers: reading one word at a time¶
Recurrent neural networks (RNNs, and their improved versions LSTM and GRU) read a sentence word by word, carrying a hidden state — a fixed-size vector that is supposed to remember everything so far:
import numpy as np
rng = np.random.default_rng(0)
d = 4 # hidden state size
W_h, W_x = rng.normal(scale=0.5, size=(d, d)), rng.normal(scale=0.5, size=(d, d))
def rnn(word_vectors):
h = np.zeros(d)
for x in word_vectors: # strictly one step after another
h = np.tanh(W_h @ h + W_x @ x)
return h # the whole sentence squeezed into 4 numbers
sentence = rng.normal(size=(6, d)) # 6 words, each a 4-dim vector
print(rnn(sentence).round(3))
Three problems follow from that loop:
| Problem | Why it happens | Consequence |
|---|---|---|
| No parallelism | step t needs the result of step t − 1 | training can't use a GPU's thousands of cores well → can't scale to huge data |
| Forgetting | everything must pass through one small vector, step by step | information from 50 words ago fades ("vanishing gradients") |
| Long paths | word 1 influences word 100 only through 99 updates | long-range relationships are hard to learn |
1.2 The big idea: let every word look at every other word¶
A transformer drops the loop. For each word it asks: which other words in the text matter for me, and how much? — and computes that for all words at the same time with matrix multiplication. That mechanism is self-attention.
"The animal didn't cross the street because it was too tired."
To understand "it", the model needs "animal" (not "street"). An RNN must have kept "animal" alive in its hidden state across five steps. Self-attention lets "it" look straight at "animal" — a path of length 1.
| RNN / LSTM | Transformer | |
|---|---|---|
| Processes tokens | one at a time | all at once (in training) |
| Path between distant words | grows with distance | always 1 step |
| GPU utilisation | poor | excellent → trained on trillions of tokens |
| Cost with sequence length n | grows like n | attention grows like n² (the price of "everyone looks at everyone") |
That last row is why context windows are limited and long prompts cost more — more in How LLMs generate text.
1.3 The architecture at a glance¶
flowchart TB
I[Input tokens] --> E[Token embeddings + position information]
E --> A[Self-attention: tokens exchange information]
A --> N1[Add & normalise]
N1 --> F[Feed-forward network: each token processed on its own]
F --> N2[Add & normalise]
N2 -->|repeat N times: 12 to 100+ blocks| A
N2 --> O[Output: one vector per token]
Each block = attention (tokens talk to each other) + a feed-forward network (each token "thinks" on its own). Stack dozens of blocks and you have an LLM. Topics 2 and 3 build both parts in NumPy.
1.4 Three families¶
The original transformer had two halves: an encoder that reads the input and a decoder that writes the output. Modern models usually keep only one half.
| Family | Attention | Trained to | Examples | Used for |
|---|---|---|---|---|
| Encoder-only | every token sees every token (bidirectional) | fill in masked words | BERT, RoBERTa, most embedding models | embeddings, classification, NER, rerankers |
| Decoder-only | each token sees only earlier tokens (causal) | predict the next token | GPT, Claude, Llama, Mistral, Gemini | chat, generation, reasoning, agents |
| Encoder–decoder | encoder bidirectional, decoder causal + looks at encoder | turn input text into output text | T5, BART, original translation models | translation, summarisation |
flowchart LR
subgraph Encoder-only
A1[The] --- A2[cat] --- A3[MASK] --- A4[on]
end
subgraph Decoder-only
B1[The] --> B2[cat] --> B3[sat] --> B4[?]
end
Decoder-only won for generation because next-token prediction needs no labelled data — any text on the internet is training data — and one model can then do every task by being prompted.
Why this matters in a RAG system
A typical RAG stack uses both families: an encoder model embeds chunks and queries (and an encoder cross-encoder reranks results), and a decoder LLM writes the answer.
Interview questions¶
Why did transformers replace RNNs/LSTMs?
Self-attention processes all tokens in parallel (fast training on GPUs, so models could scale to huge datasets) and connects any two positions directly, so long-range dependencies are easy to learn. RNNs process sequentially and squeeze history through a fixed-size state, which forgets and doesn't parallelise.
What's the downside of self-attention?
Every token attends to every other token, so compute and memory grow quadratically with sequence length. That limits context length and makes long prompts slower and costlier. Many optimisations (FlashAttention, sliding-window and sparse attention, KV caching) target this.
BERT vs GPT?
BERT is encoder-only: bidirectional attention, trained with masked-language modelling, used to produce representations (embeddings, classification). GPT is decoder-only: causal attention, trained to predict the next token, used to generate text.
Practice¶
- For each of these, say which family fits best: semantic search embeddings, a chatbot, English→Hindi translation, spam detection.
- Explain in two sentences why "it" in the animal sentence is easier for a transformer than for an RNN.
Next: Self-attention from scratch — the mechanism, in 15 lines of NumPy.