3. The transformer block¶
Intermediate · 14 min read
Attention lets tokens exchange information. A full block wraps it with a few more pieces that make deep stacks trainable. By the end of this page you'll run a forward pass through a small decoder-only model.
flowchart TB
X[x: one vector per token] --> N1[LayerNorm]
N1 --> A[Masked multi-head self-attention]
A --> R1((+))
X --> R1
R1 --> N2[LayerNorm]
N2 --> F[Feed-forward: expand ×4, activation, shrink]
F --> R2((+))
R1 --> R2
R2 --> Y[to the next block]
This is the pre-norm layout used by GPT-2 onwards and Llama: normalise before each sub-layer, then add the result back to the input (the residual).
3.1 Token embeddings¶
Token IDs from the tokenizer index into a learned embedding matrix — one row per vocabulary entry:
import numpy as np
np.set_printoptions(precision=3, suppress=True)
rng = np.random.default_rng(0)
vocab_size, d_model = 50, 8
E = rng.normal(scale=0.02, size=(vocab_size, d_model)) # learned during training
token_ids = np.array([7, 23, 7, 41]) # e.g. "the cat the dog" after tokenisation
x = E[token_ids] # a lookup, no maths
print(x.shape)
print(np.array_equal(x[0], x[2])) # same token → same vector, wherever it appears
3.2 Position information¶
Attention by itself is order-blind: shuffle the tokens and each one gets exactly the same set of attention inputs. "dog bites man" and "man bites dog" would look identical. The model needs to be told where each token is.
Sinusoidal encoding (original transformer)¶
Add a fixed pattern of sines and cosines at different frequencies — every position gets a unique "fingerprint":
def sinusoidal_positions(n_positions, d_model):
pos = np.arange(n_positions)[:, None]
i = np.arange(0, d_model, 2)[None, :]
angles = pos / (10000 ** (i / d_model))
pe = np.zeros((n_positions, d_model))
pe[:, 0::2], pe[:, 1::2] = np.sin(angles), np.cos(angles)
return pe
PE = sinusoidal_positions(4, d_model)
print(PE.round(2))
x_pos = x + PE # add position to meaning
print(np.array_equal(x_pos[0], x_pos[2])) # the two "the"s are now different
[[ 0. 1. 0. 1. 0. 1. 0. 1. ]
[ 0.84 0.54 0.1 1. 0.01 1. 0. 1. ]
[ 0.91 -0.42 0.2 0.98 0.02 1. 0. 1. ]
[ 0.14 -0.99 0.3 0.96 0.03 1. 0. 1. ]]
False
GPT-2 and BERT instead learn a position-embedding table — simple, but it can't go beyond the trained length.
RoPE — rotary position embeddings (Llama, Mistral, Qwen, most modern LLMs)¶
Rather than adding a vector, RoPE rotates each query and key by an angle proportional to its position. The useful consequence: the dot product q·k depends only on the distance between the two tokens, not on where they are in the text.
def rope(v, position, base=10000):
d = v.shape[-1]
theta = position / base ** (np.arange(0, d, 2) / d)
cos, sin = np.cos(theta), np.sin(theta)
v1, v2 = v[0::2], v[1::2]
out = np.empty_like(v)
out[0::2], out[1::2] = v1 * cos - v2 * sin, v1 * sin + v2 * cos
return out
q, k = rng.normal(size=8), rng.normal(size=8)
print(round(rope(q, 5) @ rope(k, 2), 6)) # positions 5 and 2 → distance 3
print(round(rope(q, 105) @ rope(k, 102), 6)) # positions 105 and 102 → distance 3
Same distance, same score. Relative positions generalise better, and tricks that rescale RoPE's angles (position interpolation, YaRN) are how models get context-window extensions from 8K to 128K+.
3.3 Layer normalisation¶
Each token vector is rescaled to mean 0 and variance 1, then multiplied and shifted by learned parameters. This keeps numbers in a stable range through dozens of layers.
def layer_norm(x, gamma, beta, eps=1e-5):
mean = x.mean(axis=-1, keepdims=True)
var = x.var(axis=-1, keepdims=True)
return gamma * (x - mean) / np.sqrt(var + eps) + beta
h = np.array([[1.0, 200.0, -50.0, 3.0]])
out = layer_norm(h, gamma=np.ones(4), beta=np.zeros(4))
print(out, out.mean().round(6), out.std().round(3))
Llama-style models use RMSNorm, a cheaper variant that skips subtracting the mean:
x / sqrt(mean(x²) + eps) * gamma.
3.4 Residual connections¶
Each sub-layer's output is added to its input: x = x + attention(norm(x)). The block only has to
learn a change to the representation, and gradients flow straight back through the + to early layers.
Without residuals, stacks of 30–100 layers don't train. A helpful mental model: the residual stream is a
shared "notepad" that every layer reads from and writes a little into.
3.5 The feed-forward network (MLP)¶
After attention mixes information between tokens, the feed-forward network processes each token on its own: expand to ~4× the width, apply a non-linearity, project back.
def gelu(x):
return 0.5 * x * (1 + np.tanh(np.sqrt(2 / np.pi) * (x + 0.044715 * x ** 3)))
def feed_forward(x, W1, b1, W2, b2):
return gelu(x @ W1 + b1) @ W2 + b2 # (n, d) → (n, 4d) → (n, d)
About two-thirds of an LLM's parameters live in these MLPs, and research suggests much of the model's factual knowledge is stored there. Llama uses a gated variant, SwiGLU, with three matrices instead of two.
3.6 One full block — and a tiny model¶
Putting it together, reusing attention from the previous topic:
def softmax(z, axis=-1):
z = z - z.max(axis=axis, keepdims=True)
e = np.exp(z)
return e / e.sum(axis=axis, keepdims=True)
def init_block(d_model, n_heads, rng):
s = d_model ** -0.5
return {
"n_heads": n_heads,
"ln1": (np.ones(d_model), np.zeros(d_model)), "ln2": (np.ones(d_model), np.zeros(d_model)),
"W_qkv": rng.normal(scale=s, size=(d_model, 3 * d_model)), "W_o": rng.normal(scale=s, size=(d_model, d_model)),
"W1": rng.normal(scale=s, size=(d_model, 4 * d_model)), "b1": np.zeros(4 * d_model),
"W2": rng.normal(scale=s / 2, size=(4 * d_model, d_model)), "b2": np.zeros(d_model),
}
def causal_self_attention(x, p):
n, d = x.shape
h = p["n_heads"]
q, k, v = np.split(x @ p["W_qkv"], 3, axis=-1)
q, k, v = (m.reshape(n, h, d // h).transpose(1, 0, 2) for m in (q, k, v))
scores = q @ k.transpose(0, 2, 1) / np.sqrt(d // h)
scores = np.where(np.tril(np.ones((n, n), dtype=bool)), scores, -np.inf)
out = softmax(scores) @ v
return out.transpose(1, 0, 2).reshape(n, d) @ p["W_o"]
def block(x, p):
x = x + causal_self_attention(layer_norm(x, *p["ln1"]), p) # tokens talk
x = x + feed_forward(layer_norm(x, *p["ln2"]), p["W1"], p["b1"], p["W2"], p["b2"]) # tokens think
return x
def tiny_gpt(token_ids, E, blocks, ln_f):
x = E[token_ids] + sinusoidal_positions(len(token_ids), E.shape[1])
for p in blocks:
x = block(x, p)
x = layer_norm(x, *ln_f)
return x @ E.T # logits: a score for every vocabulary token, at every position
n_layers, n_heads = 4, 2
blocks = [init_block(d_model, n_heads, rng) for _ in range(n_layers)]
ln_f = (np.ones(d_model), np.zeros(d_model))
logits = tiny_gpt(token_ids, E, blocks, ln_f)
print(logits.shape) # (tokens, vocab)
probs = softmax(logits[-1]) # prediction for the token AFTER the last one
print("next-token guess:", probs.argmax(), f"(p={probs.max():.3f}, uniform would be {1 / vocab_size:.3f})")
That is the whole forward pass of a GPT. Untrained, its guess is barely better than uniform; training
adjusts every matrix so that logits[i] puts high probability on the real token at position i + 1.
Note the last line of tiny_gpt: the output projection reuses the embedding matrix E (weight tying),
a common trick that saves vocab × d_model parameters.
3.7 Counting parameters¶
Per block: attention has 4 · d² (Q, K, V, O), the MLP has 2 · 4d · d = 8d². So roughly 12 · d² per layer, plus the embedding table:
def approx_params(n_layers, d_model, vocab):
return 12 * n_layers * d_model ** 2 + vocab * d_model
for name, L, d, V in [("GPT-2 small", 12, 768, 50257), ("GPT-2 XL", 48, 1600, 50257), ("GPT-3", 96, 12288, 50257)]:
print(f"{name:12} ≈ {approx_params(L, d, V) / 1e9:6.2f} B parameters")
Close to the published 124M, 1.5B and 175B. Memory to just hold the weights = parameters × bytes per parameter: a 7B model needs ~14 GB in 16-bit, ~3.5 GB at 4-bit — which is why quantisation matters (see Transformers for GenAI).
Interview questions¶
Why do transformers need positional encodings?
Self-attention is permutation-equivariant: it treats the input as a set, so without position information word order is lost. Positions are added (sinusoidal or learned embeddings) or injected into Q and K by rotation (RoPE), which encodes relative distance and extrapolates better to longer contexts.
What are residual connections and layer norm for?
Residuals add each sub-layer's output to its input, so layers learn small updates and gradients flow directly to early layers — essential for deep stacks. Layer norm keeps activations at a stable scale. Pre-norm (normalising before each sub-layer) trains more stably than the original post-norm.
What does the feed-forward layer do if attention already mixes tokens?
Attention moves information between positions; the MLP transforms each position independently with a large non-linear function. Most parameters are in the MLPs, and they are thought to store much of the model's factual knowledge.
Practice¶
- Shuffle
token_idsand runtiny_gptwith and without the positional encoding. When do the per-token outputs (in shuffled order) stay the same? - Use
approx_paramsto estimate Llama 3 8B (32 layers, d_model 4096, vocab 128,256). Why is the real number higher? (Hint: SwiGLU's MLP is wider.)
Next: How LLMs generate text — from logits to words, one token at a time.