Skip to content

1. Why transformers

Beginner · 8 min read

The 2017 paper "Attention Is All You Need" introduced the transformer for machine translation. Within a few years it replaced almost every other architecture in NLP — and then in vision, speech and code. To see why, look at what came before.

1.1 Before transformers: reading one word at a time

Recurrent neural networks (RNNs, and their improved versions LSTM and GRU) read a sentence word by word, carrying a hidden state — a fixed-size vector that is supposed to remember everything so far:

import numpy as np

rng = np.random.default_rng(0)
d = 4                                            # hidden state size
W_h, W_x = rng.normal(scale=0.5, size=(d, d)), rng.normal(scale=0.5, size=(d, d))

def rnn(word_vectors):
    h = np.zeros(d)
    for x in word_vectors:                       # strictly one step after another
        h = np.tanh(W_h @ h + W_x @ x)
    return h                                     # the whole sentence squeezed into 4 numbers

sentence = rng.normal(size=(6, d))               # 6 words, each a 4-dim vector
print(rnn(sentence).round(3))
Output
[0.109 0.908 0.635 0.87 ]

Three problems follow from that loop:

Problem Why it happens Consequence
No parallelism step t needs the result of step t − 1 training can't use a GPU's thousands of cores well → can't scale to huge data
Forgetting everything must pass through one small vector, step by step information from 50 words ago fades ("vanishing gradients")
Long paths word 1 influences word 100 only through 99 updates long-range relationships are hard to learn

1.2 The big idea: let every word look at every other word

A transformer drops the loop. For each word it asks: which other words in the text matter for me, and how much? — and computes that for all words at the same time with matrix multiplication. That mechanism is self-attention.

"The animal didn't cross the street because it was too tired."

To understand "it", the model needs "animal" (not "street"). An RNN must have kept "animal" alive in its hidden state across five steps. Self-attention lets "it" look straight at "animal" — a path of length 1.

RNN / LSTM Transformer
Processes tokens one at a time all at once (in training)
Path between distant words grows with distance always 1 step
GPU utilisation poor excellent → trained on trillions of tokens
Cost with sequence length n grows like n attention grows like n² (the price of "everyone looks at everyone")

That last row is why context windows are limited and long prompts cost more — more in How LLMs generate text.

1.3 The architecture at a glance

flowchart TB
    I[Input tokens] --> E[Token embeddings + position information]
    E --> A[Self-attention: tokens exchange information]
    A --> N1[Add & normalise]
    N1 --> F[Feed-forward network: each token processed on its own]
    F --> N2[Add & normalise]
    N2 -->|repeat N times: 12 to 100+ blocks| A
    N2 --> O[Output: one vector per token]

Each block = attention (tokens talk to each other) + a feed-forward network (each token "thinks" on its own). Stack dozens of blocks and you have an LLM. Topics 2 and 3 build both parts in NumPy.

1.4 Three families

The original transformer had two halves: an encoder that reads the input and a decoder that writes the output. Modern models usually keep only one half.

Family Attention Trained to Examples Used for
Encoder-only every token sees every token (bidirectional) fill in masked words BERT, RoBERTa, most embedding models embeddings, classification, NER, rerankers
Decoder-only each token sees only earlier tokens (causal) predict the next token GPT, Claude, Llama, Mistral, Gemini chat, generation, reasoning, agents
Encoder–decoder encoder bidirectional, decoder causal + looks at encoder turn input text into output text T5, BART, original translation models translation, summarisation
flowchart LR
    subgraph Encoder-only
      A1[The] --- A2[cat] --- A3[MASK] --- A4[on]
    end
    subgraph Decoder-only
      B1[The] --> B2[cat] --> B3[sat] --> B4[?]
    end

Decoder-only won for generation because next-token prediction needs no labelled data — any text on the internet is training data — and one model can then do every task by being prompted.

Why this matters in a RAG system

A typical RAG stack uses both families: an encoder model embeds chunks and queries (and an encoder cross-encoder reranks results), and a decoder LLM writes the answer.

Interview questions

Why did transformers replace RNNs/LSTMs?

Self-attention processes all tokens in parallel (fast training on GPUs, so models could scale to huge datasets) and connects any two positions directly, so long-range dependencies are easy to learn. RNNs process sequentially and squeeze history through a fixed-size state, which forgets and doesn't parallelise.

What's the downside of self-attention?

Every token attends to every other token, so compute and memory grow quadratically with sequence length. That limits context length and makes long prompts slower and costlier. Many optimisations (FlashAttention, sliding-window and sparse attention, KV caching) target this.

BERT vs GPT?

BERT is encoder-only: bidirectional attention, trained with masked-language modelling, used to produce representations (embeddings, classification). GPT is decoder-only: causal attention, trained to predict the next token, used to generate text.

Practice

  • For each of these, say which family fits best: semantic search embeddings, a chatbot, English→Hindi translation, spam detection.
  • Explain in two sentences why "it" in the animal sentence is easier for a transformer than for an RNN.

Next: Self-attention from scratch — the mechanism, in 15 lines of NumPy.