4. How LLMs generate text¶
Intermediate · 14 min read
An LLM does exactly one thing: given the tokens so far, output a score for every token in the
vocabulary — the logits. Everything else — chat, reasoning, tool calls, JSON output — comes from running
that step in a loop and choosing a token each time. The parameters you set on every API call (temperature,
top_p, max_tokens, stop) all control this loop.
4.1 A toy language model¶
To see the loop clearly, use a stand-in "model" that returns hand-made logits based on the last token. A real LLM computes these with the transformer blocks from all previous tokens.
import numpy as np
VOCAB = ["<end>", "the", "cat", "dog", "sat", "ran", "on", "mat", "away", "quickly"]
ID = {w: i for i, w in enumerate(VOCAB)}
NEXT = { # plausible continuations for each last word (logit values)
"the": {"cat": 3.0, "dog": 2.6, "mat": 1.5},
"cat": {"sat": 3.0, "ran": 2.0, "quickly": 0.5},
"dog": {"ran": 3.0, "sat": 1.8},
"sat": {"on": 3.5, "quickly": 0.8},
"ran": {"away": 3.0, "quickly": 2.5},
"on": {"the": 4.0},
"mat": {"<end>": 3.0},
"away": {"<end>": 3.0},
"quickly": {"<end>": 2.0, "away": 1.0},
}
def model(tokens: list[str]) -> np.ndarray:
logits = np.full(len(VOCAB), -2.0) # every token stays possible, just unlikely
for word, score in NEXT.get(tokens[-1], {}).items():
logits[ID[word]] = score
return logits
def softmax(z):
z = z - z.max()
e = np.exp(z)
return e / e.sum()
probs = softmax(model(["the"]))
for i in np.argsort(-probs)[:4]:
print(f"{VOCAB[i]:8} {probs[i]:.3f}")
4.2 Greedy decoding¶
Always take the most likely token:
def generate(prompt: list[str], pick, max_tokens: int = 10) -> str:
tokens = list(prompt)
for _ in range(max_tokens): # max_tokens: a hard stop
next_id = pick(model(tokens))
if VOCAB[next_id] == "<end>": # stop token: the model says it's done
break
tokens.append(VOCAB[next_id])
return " ".join(tokens)
greedy = lambda logits: int(np.argmax(logits))
print(generate(["the"], greedy))
Deterministic — and stuck in a loop. Greedy decoding always repeats once it re-enters a state it has seen. Real models loop less than this toy (they see the whole context, not just the last word), but greedy outputs are known to be repetitive and bland.
4.3 Temperature¶
Divide the logits by a temperature T before softmax. T < 1 sharpens the distribution (more confident, more predictable); T > 1 flattens it (more varied, more random); T → 0 is greedy.
logits = model(["the"])
for T in (0.2, 0.7, 1.0, 1.5):
p = softmax(logits / T)
print(f"T={T:<4}", " ".join(f"{VOCAB[i]}={p[i]:.2f}" for i in (ID["cat"], ID["dog"], ID["mat"])))
T=0.2 cat=0.88 dog=0.12 mat=0.00
T=0.7 cat=0.59 dog=0.33 mat=0.07
T=1.0 cat=0.52 dog=0.35 mat=0.11
T=1.5 cat=0.42 dog=0.32 mat=0.15
Sampling with temperature:
rng = np.random.default_rng(7)
def sample(T: float):
return lambda logits: int(rng.choice(len(logits), p=softmax(logits / T)))
for T in (0.3, 1.0, 2.0):
print(f"T={T}:", [generate(["the"], sample(T), max_tokens=6) for _ in range(2)])
T=0.3: ['the cat sat on the cat sat', 'the cat sat on the cat sat']
T=1.0: ['the cat sat on the away', 'the dog mat']
T=2.0: ['the cat ran the the dog ran', 'the mat cat sat on the']
At T=0.3 the samples are as repetitive as greedy decoding. At T=1.0 there is variety, but also odd turns ("the dog mat"). At T=2.0 the unlikely tokens (logit −2) come through regularly and the text falls apart ("the cat ran the the dog ran"). High temperature lets unlikely tokens through — and one bad token derails everything generated after it.
| Use case | Typical temperature |
|---|---|
| Extraction, classification, JSON, code, RAG answers, tool calls | 0 – 0.3 |
| General chat | 0.5 – 0.8 |
| Brainstorming, creative writing | 0.8 – 1.2 |
Temperature 0 isn't always perfectly deterministic
Over an API, the same prompt at temperature 0 can occasionally give different outputs: floating-point differences across GPUs and batches can flip near-ties between tokens. Don't rely on exact repeatability — validate the output instead.
4.4 Top-k and top-p (nucleus) sampling¶
Temperature reshapes the whole distribution; top-k and top-p cut off the long tail of junk tokens first.
def top_k_filter(logits, k):
cutoff = np.sort(logits)[-k]
return np.where(logits >= cutoff, logits, -np.inf)
def top_p_filter(logits, p):
probs = softmax(logits)
order = np.argsort(-probs)
cumulative = np.cumsum(probs[order])
keep = order[: np.searchsorted(cumulative, p) + 1] # smallest set whose total ≥ p
out = np.full_like(logits, -np.inf)
out[keep] = logits[keep]
return out
logits = model(["ran"])
for name, filtered in [("full", logits), ("top_k=2", top_k_filter(logits, 2)), ("top_p=0.9", top_p_filter(logits, 0.9))]:
p = softmax(filtered)
print(f"{name:9}", {VOCAB[i]: round(float(p[i]), 2) for i in np.argsort(-p) if p[i] > 0.005})
full {'away': 0.6, 'quickly': 0.37}
top_k=2 {'away': 0.62, 'quickly': 0.38}
top_p=0.9 {'away': 0.62, 'quickly': 0.38}
- top-k keeps a fixed number of candidates — too few when the model is unsure, too many when it's sure.
- top-p keeps the smallest set covering p of the probability — it adapts: 1–2 tokens when the model
is confident, dozens when it isn't. That's why
top_pis the common API parameter.
The usual order inside a sampler: logits → (repetition/frequency penalties) → temperature → top-k / top-p → sample.
4.5 How the model was trained: cross-entropy and perplexity¶
Training shows the model real text and, at every position, measures how much probability it gave the actual next token. The loss is the negative log of that probability (cross-entropy):
text = ["the", "cat", "sat", "on", "the", "mat"]
def sequence_loss(tokens):
losses = [-np.log(softmax(model(tokens[:i]))[ID[tokens[i]]]) for i in range(1, len(tokens))]
return float(np.mean(losses))
for sentence in (text, ["the", "mat", "sat", "on", "the", "dog"]):
loss = sequence_loss(sentence)
print(f"{' '.join(sentence):24} loss={loss:.2f} perplexity={np.exp(loss):.1f}")
the cat sat on the mat loss=0.67 perplexity=2.0
the mat sat on the dog loss=1.68 perplexity=5.4
The fluent sentence gets a low loss; the scrambled one is unlikely under the model, so its loss is much higher. Perplexity = e^loss ≈ "how many tokens the model was choosing between, on average". Lower is better; it's the main metric during pre-training.
Training = adjust all the weights by gradient descent to lower this loss over trillions of tokens. Thanks to the causal mask, one forward pass yields a loss term at every position at once.
4.6 The KV cache — why generation is fast enough¶
Generating token 101 needs attention over tokens 1–100. Without caching, the model would recompute the keys and values for all 100 previous tokens — and again for token 102, and so on. Since past tokens never change, their K and V are cached and only the new token's Q, K, V are computed:
def attention_work(n_new_tokens: int, prompt_len: int, cache: bool) -> int:
"""Count token-positions whose K/V must be computed during generation."""
total = prompt_len # "prefill": process the prompt once
for step in range(n_new_tokens):
total += 1 if cache else prompt_len + step + 1 # new token only vs. everything again
return total
for cache in (False, True):
print(f"cache={cache!s:5} K/V computations for 500 new tokens after a 2,000-token prompt:",
f"{attention_work(500, 2000, cache):,}")
cache=False K/V computations for 500 new tokens after a 2,000-token prompt: 1,127,250
cache=True K/V computations for 500 new tokens after a 2,000-token prompt: 2,500
The price is memory. Per token, every layer stores a K and a V vector for each KV head:
def kv_cache_bytes(tokens, layers, kv_heads, d_head, bytes_per_value=2): # 2 bytes = fp16/bf16
return 2 * layers * kv_heads * d_head * bytes_per_value * tokens # 2 = K and V
for name, layers, kv_heads in [("Llama 3 8B (GQA, 8 KV heads)", 32, 8), ("same model without GQA", 32, 32)]:
gb = kv_cache_bytes(128_000, layers, kv_heads, 128) / 1e9
print(f"{name:30} 128K-token context → {gb:5.1f} GB of KV cache per sequence")
Llama 3 8B (GQA, 8 KV heads) 128K-token context → 16.8 GB of KV cache per sequence
same model without GQA 128K-token context → 67.1 GB of KV cache per sequence
That's more than the model's own 16 GB of weights — for one user. This is why grouped-query attention, KV-cache quantisation and paged memory (vLLM's PagedAttention) exist, and why providers charge for long context.
4.7 Prefill vs decode: what latency and price mean¶
| Phase | What happens | Speed | You see it as |
|---|---|---|---|
| Prefill | the whole prompt is processed in parallel; KV cache filled | fast per token (GPU busy, parallel) | time to first token (TTFT) |
| Decode | one token per forward pass, reading the whole KV cache each time | slow per token (memory-bound) | tokens per second while streaming |
This explains everyday facts about LLM APIs:
- Output tokens cost several times more than input tokens — each one needs a full sequential forward pass.
- Long prompts raise time-to-first-token (attention over n tokens costs ~n²), and long contexts slow every decode step a little (bigger KV cache to read).
- Prompt caching discounts a repeated prompt prefix — the provider reuses its stored KV cache instead of re-running prefill. Put stable content (system prompt, documents, tool definitions) first to benefit.
- Streaming doesn't make generation faster; it shows the decode phase as it happens.
Interview questions¶
Explain temperature, top-k and top-p.
Temperature divides logits before softmax: below 1 sharpens the distribution, above 1 flattens it, near 0 is greedy. Top-k keeps only the k most likely tokens; top-p keeps the smallest set whose cumulative probability reaches p, adapting to the model's confidence. All three trade predictability against diversity.
What is the KV cache?
During autoregressive decoding, the keys and values of previous tokens don't change, so they're stored and reused; each step computes Q/K/V only for the new token. It turns quadratic recomputation into linear work per step, at the cost of memory proportional to layers × KV heads × head size × context length.
Why are output tokens more expensive than input tokens?
Input tokens are processed in one parallel prefill pass. Output tokens are generated one at a time, each needing its own forward pass that reads all the weights and the KV cache — far less efficient use of the GPU.
What is perplexity?
The exponential of the average cross-entropy loss on held-out text — roughly the effective number of choices the model hesitates between per token. Lower means the model predicts the text better.
Practice¶
- Add a repetition penalty to the greedy loop (subtract 2.0 from the logits of tokens already generated) — does the loop go away?
- Compute the KV cache for 32 concurrent users at 8K context each for Llama 3 8B. Does it fit on an 80 GB GPU next to the weights?
Next: Transformers for GenAI — training stages, Hugging Face, LoRA and quantisation.