Skip to content

4. How LLMs generate text

Intermediate · 14 min read

An LLM does exactly one thing: given the tokens so far, output a score for every token in the vocabulary — the logits. Everything else — chat, reasoning, tool calls, JSON output — comes from running that step in a loop and choosing a token each time. The parameters you set on every API call (temperature, top_p, max_tokens, stop) all control this loop.

4.1 A toy language model

To see the loop clearly, use a stand-in "model" that returns hand-made logits based on the last token. A real LLM computes these with the transformer blocks from all previous tokens.

import numpy as np

VOCAB = ["<end>", "the", "cat", "dog", "sat", "ran", "on", "mat", "away", "quickly"]
ID = {w: i for i, w in enumerate(VOCAB)}

NEXT = {                      # plausible continuations for each last word (logit values)
    "the": {"cat": 3.0, "dog": 2.6, "mat": 1.5},
    "cat": {"sat": 3.0, "ran": 2.0, "quickly": 0.5},
    "dog": {"ran": 3.0, "sat": 1.8},
    "sat": {"on": 3.5, "quickly": 0.8},
    "ran": {"away": 3.0, "quickly": 2.5},
    "on": {"the": 4.0},
    "mat": {"<end>": 3.0},
    "away": {"<end>": 3.0},
    "quickly": {"<end>": 2.0, "away": 1.0},
}

def model(tokens: list[str]) -> np.ndarray:
    logits = np.full(len(VOCAB), -2.0)                 # every token stays possible, just unlikely
    for word, score in NEXT.get(tokens[-1], {}).items():
        logits[ID[word]] = score
    return logits

def softmax(z):
    z = z - z.max()
    e = np.exp(z)
    return e / e.sum()

probs = softmax(model(["the"]))
for i in np.argsort(-probs)[:4]:
    print(f"{VOCAB[i]:8} {probs[i]:.3f}")
Output
cat      0.515
dog      0.345
mat      0.115
the      0.003

4.2 Greedy decoding

Always take the most likely token:

def generate(prompt: list[str], pick, max_tokens: int = 10) -> str:
    tokens = list(prompt)
    for _ in range(max_tokens):                       # max_tokens: a hard stop
        next_id = pick(model(tokens))
        if VOCAB[next_id] == "<end>":                 # stop token: the model says it's done
            break
        tokens.append(VOCAB[next_id])
    return " ".join(tokens)

greedy = lambda logits: int(np.argmax(logits))
print(generate(["the"], greedy))
Output
the cat sat on the cat sat on the cat sat

Deterministic — and stuck in a loop. Greedy decoding always repeats once it re-enters a state it has seen. Real models loop less than this toy (they see the whole context, not just the last word), but greedy outputs are known to be repetitive and bland.

4.3 Temperature

Divide the logits by a temperature T before softmax. T < 1 sharpens the distribution (more confident, more predictable); T > 1 flattens it (more varied, more random); T → 0 is greedy.

logits = model(["the"])
for T in (0.2, 0.7, 1.0, 1.5):
    p = softmax(logits / T)
    print(f"T={T:<4}", "  ".join(f"{VOCAB[i]}={p[i]:.2f}" for i in (ID["cat"], ID["dog"], ID["mat"])))
Output
T=0.2  cat=0.88  dog=0.12  mat=0.00
T=0.7  cat=0.59  dog=0.33  mat=0.07
T=1.0  cat=0.52  dog=0.35  mat=0.11
T=1.5  cat=0.42  dog=0.32  mat=0.15

Sampling with temperature:

rng = np.random.default_rng(7)

def sample(T: float):
    return lambda logits: int(rng.choice(len(logits), p=softmax(logits / T)))

for T in (0.3, 1.0, 2.0):
    print(f"T={T}:", [generate(["the"], sample(T), max_tokens=6) for _ in range(2)])
Output
T=0.3: ['the cat sat on the cat sat', 'the cat sat on the cat sat']
T=1.0: ['the cat sat on the away', 'the dog mat']
T=2.0: ['the cat ran the the dog ran', 'the mat cat sat on the']

At T=0.3 the samples are as repetitive as greedy decoding. At T=1.0 there is variety, but also odd turns ("the dog mat"). At T=2.0 the unlikely tokens (logit −2) come through regularly and the text falls apart ("the cat ran the the dog ran"). High temperature lets unlikely tokens through — and one bad token derails everything generated after it.

Use case Typical temperature
Extraction, classification, JSON, code, RAG answers, tool calls 0 – 0.3
General chat 0.5 – 0.8
Brainstorming, creative writing 0.8 – 1.2

Temperature 0 isn't always perfectly deterministic

Over an API, the same prompt at temperature 0 can occasionally give different outputs: floating-point differences across GPUs and batches can flip near-ties between tokens. Don't rely on exact repeatability — validate the output instead.

4.4 Top-k and top-p (nucleus) sampling

Temperature reshapes the whole distribution; top-k and top-p cut off the long tail of junk tokens first.

def top_k_filter(logits, k):
    cutoff = np.sort(logits)[-k]
    return np.where(logits >= cutoff, logits, -np.inf)

def top_p_filter(logits, p):
    probs = softmax(logits)
    order = np.argsort(-probs)
    cumulative = np.cumsum(probs[order])
    keep = order[: np.searchsorted(cumulative, p) + 1]     # smallest set whose total ≥ p
    out = np.full_like(logits, -np.inf)
    out[keep] = logits[keep]
    return out

logits = model(["ran"])
for name, filtered in [("full", logits), ("top_k=2", top_k_filter(logits, 2)), ("top_p=0.9", top_p_filter(logits, 0.9))]:
    p = softmax(filtered)
    print(f"{name:9}", {VOCAB[i]: round(float(p[i]), 2) for i in np.argsort(-p) if p[i] > 0.005})
Output
full      {'away': 0.6, 'quickly': 0.37}
top_k=2   {'away': 0.62, 'quickly': 0.38}
top_p=0.9 {'away': 0.62, 'quickly': 0.38}
  • top-k keeps a fixed number of candidates — too few when the model is unsure, too many when it's sure.
  • top-p keeps the smallest set covering p of the probability — it adapts: 1–2 tokens when the model is confident, dozens when it isn't. That's why top_p is the common API parameter.

The usual order inside a sampler: logits → (repetition/frequency penalties) → temperature → top-k / top-p → sample.

4.5 How the model was trained: cross-entropy and perplexity

Training shows the model real text and, at every position, measures how much probability it gave the actual next token. The loss is the negative log of that probability (cross-entropy):

text = ["the", "cat", "sat", "on", "the", "mat"]

def sequence_loss(tokens):
    losses = [-np.log(softmax(model(tokens[:i]))[ID[tokens[i]]]) for i in range(1, len(tokens))]
    return float(np.mean(losses))

for sentence in (text, ["the", "mat", "sat", "on", "the", "dog"]):
    loss = sequence_loss(sentence)
    print(f"{' '.join(sentence):24} loss={loss:.2f}  perplexity={np.exp(loss):.1f}")
Output
the cat sat on the mat   loss=0.67  perplexity=2.0
the mat sat on the dog   loss=1.68  perplexity=5.4

The fluent sentence gets a low loss; the scrambled one is unlikely under the model, so its loss is much higher. Perplexity = e^loss ≈ "how many tokens the model was choosing between, on average". Lower is better; it's the main metric during pre-training.

Training = adjust all the weights by gradient descent to lower this loss over trillions of tokens. Thanks to the causal mask, one forward pass yields a loss term at every position at once.

4.6 The KV cache — why generation is fast enough

Generating token 101 needs attention over tokens 1–100. Without caching, the model would recompute the keys and values for all 100 previous tokens — and again for token 102, and so on. Since past tokens never change, their K and V are cached and only the new token's Q, K, V are computed:

def attention_work(n_new_tokens: int, prompt_len: int, cache: bool) -> int:
    """Count token-positions whose K/V must be computed during generation."""
    total = prompt_len                                     # "prefill": process the prompt once
    for step in range(n_new_tokens):
        total += 1 if cache else prompt_len + step + 1     # new token only vs. everything again
    return total

for cache in (False, True):
    print(f"cache={cache!s:5}  K/V computations for 500 new tokens after a 2,000-token prompt:",
          f"{attention_work(500, 2000, cache):,}")
Output
cache=False  K/V computations for 500 new tokens after a 2,000-token prompt: 1,127,250
cache=True   K/V computations for 500 new tokens after a 2,000-token prompt: 2,500

The price is memory. Per token, every layer stores a K and a V vector for each KV head:

def kv_cache_bytes(tokens, layers, kv_heads, d_head, bytes_per_value=2):    # 2 bytes = fp16/bf16
    return 2 * layers * kv_heads * d_head * bytes_per_value * tokens       # 2 = K and V

for name, layers, kv_heads in [("Llama 3 8B (GQA, 8 KV heads)", 32, 8), ("same model without GQA", 32, 32)]:
    gb = kv_cache_bytes(128_000, layers, kv_heads, 128) / 1e9
    print(f"{name:30} 128K-token context → {gb:5.1f} GB of KV cache per sequence")
Output
Llama 3 8B (GQA, 8 KV heads)   128K-token context →  16.8 GB of KV cache per sequence
same model without GQA         128K-token context →  67.1 GB of KV cache per sequence

That's more than the model's own 16 GB of weights — for one user. This is why grouped-query attention, KV-cache quantisation and paged memory (vLLM's PagedAttention) exist, and why providers charge for long context.

4.7 Prefill vs decode: what latency and price mean

Phase What happens Speed You see it as
Prefill the whole prompt is processed in parallel; KV cache filled fast per token (GPU busy, parallel) time to first token (TTFT)
Decode one token per forward pass, reading the whole KV cache each time slow per token (memory-bound) tokens per second while streaming

This explains everyday facts about LLM APIs:

  • Output tokens cost several times more than input tokens — each one needs a full sequential forward pass.
  • Long prompts raise time-to-first-token (attention over n tokens costs ~n²), and long contexts slow every decode step a little (bigger KV cache to read).
  • Prompt caching discounts a repeated prompt prefix — the provider reuses its stored KV cache instead of re-running prefill. Put stable content (system prompt, documents, tool definitions) first to benefit.
  • Streaming doesn't make generation faster; it shows the decode phase as it happens.

Interview questions

Explain temperature, top-k and top-p.

Temperature divides logits before softmax: below 1 sharpens the distribution, above 1 flattens it, near 0 is greedy. Top-k keeps only the k most likely tokens; top-p keeps the smallest set whose cumulative probability reaches p, adapting to the model's confidence. All three trade predictability against diversity.

What is the KV cache?

During autoregressive decoding, the keys and values of previous tokens don't change, so they're stored and reused; each step computes Q/K/V only for the new token. It turns quadratic recomputation into linear work per step, at the cost of memory proportional to layers × KV heads × head size × context length.

Why are output tokens more expensive than input tokens?

Input tokens are processed in one parallel prefill pass. Output tokens are generated one at a time, each needing its own forward pass that reads all the weights and the KV cache — far less efficient use of the GPU.

What is perplexity?

The exponential of the average cross-entropy loss on held-out text — roughly the effective number of choices the model hesitates between per token. Lower means the model predicts the text better.

Practice

  • Add a repetition penalty to the greedy loop (subtract 2.0 from the logits of tokens already generated) — does the loop go away?
  • Compute the KV cache for 32 concurrent users at 8K context each for Llama 3 8B. Does it fit on an 80 GB GPU next to the weights?

Next: Transformers for GenAI — training stages, Hugging Face, LoRA and quantisation.