1. What is an LLM¶
Beginner · 11 min read
A large language model is a transformer with billions of parameters, trained on trillions of tokens of text to do one thing: predict the next token. Everything else — answering questions, writing code, calling tools, following instructions — emerges from doing that one thing extremely well, plus some fine-tuning to behave like an assistant.
Holding that one fact in mind explains most of what LLMs get right and wrong.
1.1 A text predictor that learned a lot about the world¶
To predict the next word of a chemistry textbook, a legal contract or a Python file well, a model has to pick up grammar, facts, reasoning patterns and coding conventions. Nobody programs these in — they are compressed into the weights during pre-training.
flowchart LR
D[Trillions of tokens:<br/>web, books, code] -->|pre-training:<br/>predict next token| B[Base model]
B -->|instruction tuning +<br/>preference tuning| A[Assistant model]
A -->|your prompt| O[Generated text]
How a base model becomes an assistant is covered in Transformers for GenAI; how it picks each token in How LLMs generate text.
1.2 In-context learning¶
LLMs can learn a new task from the prompt alone, with no training — show a few examples and the model continues the pattern. This is why prompting works at all:
Convert to the internal SKU format.
red cotton t-shirt, size M → TSH-COT-RED-M
blue denim jeans, size 32 → JNS-DEN-BLU-32
black wool sweater, size L →
A good model answers SWT-WOL-BLK-L — a format it has never seen before this prompt. The weights don't change;
the "learning" lives only in the context window and is gone on the next call.
1.3 What LLMs are good and bad at¶
| Strong | Weak |
|---|---|
| Understanding and writing natural language in many languages | Exact arithmetic on large numbers, counting characters or words |
| Summarising, rewriting, translating, changing tone | Facts after the knowledge cutoff, or niche facts rarely seen in training |
| Extracting structure from messy text | Knowing what it doesn't know — it answers confidently anyway |
| Writing and explaining code | Long chains of precise steps without tools or checking |
| Classification and routing with no training data | Being exactly repeatable |
| Following instructions and formats | Keeping secrets in the prompt, resisting manipulation (prompt injection) |
The pattern for the weak column: don't ask the model to do what a tool does better. Give it a calculator, a search engine, a database query or a code interpreter, and let it decide when to use them (tool calling).
import re
def calculator(expression: str) -> float:
if not re.fullmatch(r"[\d\s.+\-*/()]+", expression): # only arithmetic, nothing else
raise ValueError("unsupported expression")
return eval(expression)
# Ask an LLM "what is 48,213 × 7,919?" and it may be off by a few digits.
# Ask it to call calculator("48213 * 7919") and the answer is exact:
print(f"{calculator('48213 * 7919'):,.0f}")
1.4 Why LLMs hallucinate¶
A hallucination is fluent, confident output that is false — a made-up citation, an API parameter that doesn't exist, a wrong date. It happens because of how the model was trained:
- It is trained to produce a plausible continuation, not a verified one. "The paper by Sharma et al. (2019) in Nature showed…" is extremely plausible text, whether or not the paper exists.
- It has no built-in lookup of where a fact came from — knowledge is blended into billions of weights.
- Training and fine-tuning reward answering; saying "I don't know" has historically been rewarded less.
- Rare facts (a small company's policy, a niche library's API) were seen few times, so the model's memory of them is fuzzy — and fuzzy memory fills gaps with plausible guesses.
How to reduce it in an application (details in Evaluation & guardrails):
- Ground answers in retrieved documents (RAG) and say "answer only from the context".
- Allow "I don't know" explicitly — and skip the LLM entirely when retrieval finds nothing.
- Require citations and check they point to real sources.
- Use tools for facts and numbers instead of the model's memory.
- Lower the temperature for factual tasks.
- Measure the hallucination rate on an eval set.
1.5 Knowledge cutoff¶
A model knows nothing that happened after its training data was collected — the knowledge cutoff. It also doesn't know today's date unless you tell it.
from datetime import date
def system_prompt(today: date) -> str:
return (f"Today's date is {today:%d %B %Y}. Your training data may be older than this. "
"For anything time-sensitive (prices, news, versions, schedules), use the search tool "
"instead of your memory, and say when you're unsure how current your information is.")
print(system_prompt(date(2026, 10, 6)))
Today's date is 06 October 2026. Your training data may be older than this. For anything time-sensitive (prices, news, versions, schedules), use the search tool instead of your memory, and say when you're unsure how current your information is.
Library and API versions are a classic trap: models write code for the version they saw most in training, which may be outdated. Paste the current docs into the prompt, or retrieve them.
1.6 Non-determinism¶
The same prompt can give different answers:
- Sampling: with temperature > 0, the next token is drawn at random from the distribution — that's by design.
- Temperature 0 isn't a guarantee: floating-point differences across GPUs and batch sizes can flip near-ties.
- Model updates: a provider's model alias (e.g. "latest") can change behaviour overnight. Pin a dated model version in production and re-run your evals before upgrading.
So never test an LLM feature with one run of one example. Test many inputs, sometimes several runs each, and validate outputs in code.
1.7 Why bigger is (usually) better: scaling laws¶
Research found that a model's loss falls smoothly and predictably as you increase parameters (N), training tokens (D) and compute (C). A useful rule of thumb: training compute ≈ 6 × N × D floating-point operations, and a compute-efficient ("Chinchilla-optimal") model is trained on roughly 20 tokens per parameter.
def training_flops(params: float, tokens: float) -> float:
return 6 * params * tokens
for name, n_params in [("1B", 1e9), ("8B", 8e9), ("70B", 70e9)]:
tokens_optimal = 20 * n_params
print(f"{name:4} compute-optimal ≈ {tokens_optimal / 1e12:5.2f}T tokens, "
f"{training_flops(n_params, tokens_optimal):.1e} FLOPs")
1B compute-optimal ≈ 0.02T tokens, 1.2e+20 FLOPs
8B compute-optimal ≈ 0.16T tokens, 7.7e+21 FLOPs
70B compute-optimal ≈ 1.40T tokens, 5.9e+23 FLOPs
In practice, modern small models are trained on far more than 20 tokens per parameter (often 15T+ tokens for an 8B model): it costs more to train, but gives a smaller model that is cheaper to serve for its whole life. That's why today's 8B models beat much larger models from a few years ago.
Abilities such as multi-step reasoning or following complex instructions tend to improve sharply with scale — and more recently with reasoning training that spends extra compute at answer time (see Reasoning models).
Interview questions¶
What is an LLM, in one minute?
A large transformer trained on a huge text corpus to predict the next token. Doing that well forces it to learn language, facts and reasoning patterns. Instruction and preference tuning then turn it into an assistant that follows prompts. It generates text one token at a time by sampling from its predicted distribution.
Why do LLMs hallucinate, and how do you reduce it?
They're trained to produce plausible text, not verified facts, and have no record of their sources; rare facts are stored imprecisely. Reduce it with retrieval grounding, explicit permission to say "I don't know", citations that are checked, tools for facts and calculations, low temperature, and measuring on an eval set.
What is in-context learning?
The ability to perform a new task from instructions or examples given in the prompt, without updating weights. It's the basis of zero-shot and few-shot prompting.
What is a knowledge cutoff and how do you work around it?
The date after which the model has no training data. Work around it with retrieval or search tools for current information, by putting today's date and current documentation in the prompt, and by pinning model versions.
Practice¶
- Ask a model for three academic papers on a niche topic, then check whether each one exists.
- Ask a model which version of a library you use is the latest, and compare with the real answer.
Next: Choosing a model — closed vs open, tiers, context windows and cost.