5. Transformers for GenAI¶
Intermediate · 15 min read
You now know what's inside a transformer. This page connects it to the decisions you make as a GenAI engineer: which model to use, how a base model becomes a chat assistant, when to fine-tune, and how to run models cheaply.
5.1 From base model to assistant: the training stages¶
flowchart LR
P[Pre-training<br/>trillions of tokens<br/>next-token prediction] --> B[Base model<br/>continues text]
B --> S[Supervised fine-tuning SFT<br/>~10k–1M instruction/answer pairs]
S --> I[Instruct model<br/>follows instructions]
I --> R[Preference tuning<br/>RLHF / DPO]
R --> C[Chat model<br/>helpful, harmless, honest]
C --> RL[Optional: RL on verifiable tasks<br/>maths, code → reasoning models]
| Stage | Data | What the model learns |
|---|---|---|
| Pre-training | web pages, books, code — trillions of tokens, no labels | language, facts, reasoning patterns; costs millions of dollars of GPU time |
| SFT (supervised fine-tuning) | prompt → ideal response pairs, written or curated by people | the format of being an assistant: answer the question, follow instructions, use the chat template |
| RLHF (reinforcement learning from human feedback) | humans rank several answers; a reward model learns their preferences; the LLM is optimised against it | which answers people prefer — helpfulness, tone, refusing harmful requests |
| DPO (direct preference optimisation) | the same "chosen vs rejected" pairs | the same goal as RLHF without a separate reward model — simpler, now very common |
| RL with verifiable rewards | maths problems, code with tests | long step-by-step reasoning ("thinking" models) |
A base model asked "What is the capital of France?" may continue with "What is the capital of Germany?" — it's completing a list of quiz questions. The later stages are what make it answer.
5.2 Chat templates — what the model actually sees¶
Chat APIs take a list of messages, but the model reads one token sequence. A chat template flattens the messages with special marker tokens learned during SFT:
def apply_chat_template(messages: list[dict], add_generation_prompt: bool = True) -> str:
"""A simplified ChatML-style template (used by several open models)."""
text = "".join(f"<|im_start|>{m['role']}\n{m['content']}<|im_end|>\n" for m in messages)
if add_generation_prompt:
text += "<|im_start|>assistant\n" # the model continues from here
return text
print(apply_chat_template([
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is RAG?"},
]))
<|im_start|>system
You are a concise assistant.<|im_end|>
<|im_start|>user
What is RAG?<|im_end|>
<|im_start|>assistant
Each model family has its own template (Llama, Mistral, Qwen and Gemma all differ). Using the wrong one with an open model silently degrades quality — always use the tokenizer's built-in template:
# no-run — pip install transformers torch
from transformers import AutoTokenizer, AutoModelForCausalLM
name = "Qwen/Qwen2.5-0.5B-Instruct" # small enough for a laptop CPU
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name)
messages = [{"role": "user", "content": "Explain RAG in one sentence."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs, max_new_tokens=60, do_sample=True, temperature=0.3, top_p=0.9)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
Tool calling works the same way: tool definitions and tool calls are rendered into the sequence with special
tokens the model was fine-tuned on — the API's tools parameter is a convenience layer on top.
5.3 Embedding models: from token vectors to one vector¶
An encoder outputs one vector per token. To get one vector per sentence — for semantic search and RAG — the token vectors are pooled, usually by averaging (ignoring padding) and then normalising:
import numpy as np
rng = np.random.default_rng(0)
token_vectors = rng.normal(size=(2, 6, 8)) # 2 sentences, padded to 6 tokens, 8 dims
attention_mask = np.array([[1, 1, 1, 1, 0, 0], # sentence 1 has 4 real tokens
[1, 1, 1, 1, 1, 1]]) # sentence 2 has 6
def mean_pool(vectors, mask):
mask = mask[..., None]
pooled = (vectors * mask).sum(axis=1) / mask.sum(axis=1) # average real tokens only
return pooled / np.linalg.norm(pooled, axis=1, keepdims=True) # unit length → dot product = cosine
emb = mean_pool(token_vectors, attention_mask)
print(emb.shape, np.linalg.norm(emb, axis=1).round(3))
Embedding models (e.g. the sentence-transformers family, BGE, E5, OpenAI's text-embedding-3-*) are
transformers fine-tuned with contrastive learning — pulling matching question/passage pairs together and
pushing others apart — so that this pooled vector captures meaning.
| Need | Model type | Why |
|---|---|---|
| Embed chunks and queries for retrieval | bi-encoder (embedding model) | encode each text once, compare millions with a dot product |
| Re-order the top 20–50 results precisely | cross-encoder (reranker) | reads query + passage together through full attention — more accurate, too slow for millions |
| Classify or tag at high volume | small fine-tuned encoder (BERT-style) | fast, cheap, accurate with labelled data |
| Generate, reason, call tools | decoder LLM | — |
5.4 Prompting, RAG or fine-tuning?¶
| Your problem | Best first tool |
|---|---|
| The model doesn't know your data (docs, tickets, products) | RAG — knowledge changes; fine-tuning is a poor way to add facts |
| The model knows enough but the output format/style is off | prompting with examples, structured output |
| A narrow task at high volume needs a smaller/cheaper/faster model | fine-tuning (often distilling from a big model's outputs) |
| A consistent tone, persona or domain language prompts can't hold | fine-tuning |
| Needs the latest information | RAG / tools (web search, APIs) |
Rule of thumb: prompt → RAG → fine-tune, in that order, with an eval set at each step to prove the next one is needed.
5.5 LoRA — fine-tuning without updating every weight¶
Full fine-tuning updates billions of parameters and needs the memory for their gradients and optimiser state
(several times the model size). LoRA (low-rank adaptation) freezes the original weight matrix W and learns
a small update ΔW = B @ A, where A and B are thin matrices of rank r (typically 8–64):
d = 4096 # one projection matrix in a 7–8B model
r = 16
W = rng.normal(scale=0.02, size=(d, d)) # frozen
A = rng.normal(scale=0.01, size=(r, d)) # trainable
B = np.zeros((d, r)) # trainable, starts at 0 → the model starts unchanged
x = rng.normal(size=(1, d))
y = x @ W.T + x @ A.T @ B.T # original path + low-rank update
print(np.allclose(y, x @ W.T)) # B = 0 → identical output before training
full, lora = W.size, A.size + B.size
print(f"full matrix: {full:,} params | LoRA r={r}: {lora:,} params ({lora / full:.2%})")
- Training touches < 1% of the parameters → fits on one GPU; adapters are a few MB to share and swap.
- After training,
B @ Acan be merged intoW— zero extra inference cost. - QLoRA = LoRA on top of a 4-bit-quantised frozen base model → fine-tune a 7–8B model on a single 24 GB GPU, or even a free Colab GPU for small models.
- Libraries: Hugging Face PEFT + TRL (
SFTTrainer,DPOTrainer), Unsloth, Axolotl.
5.6 Quantisation — smaller, faster models¶
Weights are trained in 16- or 32-bit floats. Quantisation stores them in fewer bits — 8, 4, even 2 — with a scale factor to map back:
w = rng.normal(scale=0.02, size=4096).astype(np.float32) # one row of weights
def quantise_int8(x):
scale = np.abs(x).max() / 127 # map the largest weight to ±127
return np.round(x / scale).astype(np.int8), scale
def quantise_int4(x, group=64): # 4-bit, one scale per group of 64
g = x.reshape(-1, group)
scale = np.abs(g).max(axis=1, keepdims=True) / 7
return np.clip(np.round(g / scale), -8, 7).astype(np.int8), scale
q8, s8 = quantise_int8(w)
q4, s4 = quantise_int4(w)
for name, restored, bits in [("int8", q8 * s8, 8), ("int4", (q4 * s4).ravel(), 4)]:
rel_err = np.abs(restored - w).mean() / np.abs(w).mean()
print(f"{name}: {bits / 16:.0%} of fp16 size, mean relative error {rel_err:.1%}")
int8: 50% of fp16 size, mean relative error 1.0%
int4: 25% of fp16 size, mean relative error 11.6%
An ~11% error per weight sounds large, but it is random noise that largely averages out across billions of
weights — and real 4-bit methods (GPTQ, AWQ, GGUF Q4_K_M) are smarter than this naive rounding. They usually
lose only a little quality, while cutting memory by 4× and speeding up decoding (which is
limited by how fast weights are read from memory).
for params_b in (8, 70):
print(f"{params_b}B model:", " ".join(f"{name} ≈ {params_b * bits / 8:.0f} GB"
for name, bits in [("fp16", 16), ("int8", 8), ("4-bit", 4)]))
8B model: fp16 ≈ 16 GB int8 ≈ 8 GB 4-bit ≈ 4 GB
70B model: fp16 ≈ 140 GB int8 ≈ 70 GB 4-bit ≈ 35 GB
Weights only — add the KV cache from How LLMs generate text.
5.7 Serving open models¶
| Tool | Use it for |
|---|---|
| Ollama, LM Studio, llama.cpp | running quantised (GGUF) models on a laptop; local dev and private demos |
| vLLM, SGLang, TGI | production GPU serving: continuous batching, PagedAttention KV cache, OpenAI-compatible API |
Hugging Face transformers |
experiments, fine-tuning, evaluation — not optimised for high-throughput serving |
Serving tricks you'll hear about: continuous batching (new requests join a running batch every step), prefix caching (reuse the KV cache of a shared system prompt), and speculative decoding (a small model drafts several tokens, the big one verifies them in one pass — same output, 2–3× faster).
Because vLLM and Ollama expose an OpenAI-compatible endpoint, your app code barely changes:
# no-run — `ollama run llama3.1` must be running locally
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
r = client.chat.completions.create(model="llama3.1", messages=[{"role": "user", "content": "Hi!"}])
print(r.choices[0].message.content)
5.8 Fine-tuning in practice¶
Section 5.4 covered when to fine-tune; 5.5 covered LoRA. This is the workflow, step by step.
flowchart LR
B[1. Baseline:<br/>best prompt + eval set] --> D[2. Collect &<br/>clean data]
D --> S[3. Split<br/>train / validation / test]
S --> T[4. Train<br/>API or LoRA]
T --> E[5. Evaluate vs baseline<br/>on the held-out test set]
E -->|better| SH[6. Ship + monitor]
E -->|not better| D
1. Baseline first. Write the best prompt you can (with few-shot examples) and score it on an eval set. Fine-tuning has to beat this, not a lazy prompt.
2. Data. Training examples are full conversations in the model's chat format, usually as JSONL (one JSON object per line). Quality beats quantity: a few hundred clean, consistent, diverse examples often beat thousands of noisy ones. A common source is distillation — outputs from a large model, reviewed by people, used to train a small one.
import json, random
raw = [ # (ticket, label) pairs — in practice from labelled logs or reviewed LLM outputs
("App crashes when uploading a 20 MB PDF", "bug"), ("Please add dark mode", "feature_request"),
("Charged twice for March", "billing"), ("Login page shows error 500", "bug"),
("Can you support UPI autopay?", "feature_request"), ("Refund not received after 10 days", "billing"),
("Export to Excel is broken", "bug"), ("Invoice shows the wrong GST number", "billing"),
("Charged twice for March", "billing"), # a duplicate
("", "bug"), # an empty example
]
SYSTEM = "Classify the support ticket as bug, feature_request or billing. Reply with the label only."
def to_example(ticket, label):
return {"messages": [{"role": "system", "content": SYSTEM},
{"role": "user", "content": ticket},
{"role": "assistant", "content": label}]}
clean = list(dict.fromkeys((t.strip(), l) for t, l in raw if t.strip())) # drop empties and duplicates
random.Random(0).shuffle(clean)
n_val = max(1, len(clean) // 5)
train, val = clean[n_val:], clean[:n_val]
with open("train.jsonl", "w") as f:
for t, l in train:
f.write(json.dumps(to_example(t, l)) + "\n")
print(f"{len(raw)} raw → {len(clean)} clean → {len(train)} train / {len(val)} validation")
print(open("train.jsonl").readline()[:120], "…")
10 raw → 8 clean → 7 train / 1 validation
{"messages": [{"role": "system", "content": "Classify the support ticket as bug, feature_request or billing. Reply with …
Checklist for the data: every label well represented, the hard and ambiguous cases included, no personal data you aren't allowed to use, and no overlap between training data and the test set (otherwise the scores lie).
3–4. Train. Two routes:
# no-run — OpenAI fine-tuning; other providers and hosts offer similar APIs
from openai import OpenAI
client = OpenAI()
file = client.files.create(file=open("train.jsonl", "rb"), purpose="fine-tune")
job = client.fine_tuning.jobs.create(training_file=file.id, model="gpt-4o-mini-2024-07-18")
# when the job finishes: use job.fine_tuned_model as the `model` in normal chat calls
No GPUs to manage; you pay per training token and a higher price per call; weights stay with the provider.
# no-run — pip install transformers peft trl datasets ; needs a GPU
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer
trainer = SFTTrainer(
model="Qwen/Qwen2.5-1.5B-Instruct",
train_dataset=load_dataset("json", data_files="train.jsonl")["train"],
peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM"),
args=SFTConfig(output_dir="ticket-classifier", num_train_epochs=3, learning_rate=2e-4),
)
trainer.train()
Full control and the weights are yours; you run the GPUs and the serving.
5. Evaluate. Score the fine-tuned model and the prompted baseline on the same held-out test set. Also check that general abilities didn't break — fine-tuning on a narrow task can make a model worse at everything else ("catastrophic forgetting").
| Pitfall | Symptom | Fix |
|---|---|---|
| Overfitting | training loss keeps falling, validation loss rises | fewer epochs, more varied data |
| Inconsistent labels | model is "randomly" wrong on similar inputs | label guidelines, review disagreements |
| Train/test overlap | great scores, poor production results | deduplicate across splits |
| Prompt mismatch | works in testing, not in the app | use the same system prompt at training and inference |
| Trying to add knowledge | still makes up facts | use RAG for knowledge; fine-tune for behaviour |
Interview questions¶
How does a base LLM become a chat assistant?
Pre-training on huge unlabelled text teaches next-token prediction. Supervised fine-tuning on instruction → response pairs (in a chat template) teaches it to follow instructions. Preference tuning (RLHF with a reward model, or DPO directly on chosen/rejected pairs) aligns it with human preferences. Reasoning models add RL on tasks with verifiable answers.
When would you fine-tune instead of using RAG?
RAG for knowledge — facts that change or must be cited. Fine-tuning for behaviour — a consistent format, style or domain language, or to make a small model do a narrow task cheaply. Often both: a fine-tuned model that is good at answering from retrieved context. Always justify with an eval set.
What is LoRA and why is it popular?
It freezes the pretrained weights and learns a low-rank update B·A for selected matrices, training well under 1% of the parameters. That cuts GPU memory and storage dramatically, adapters can be swapped per task, and the update can be merged for zero inference overhead. QLoRA applies it on a 4-bit base model.
Bi-encoder vs cross-encoder?
A bi-encoder embeds query and document separately into vectors — fast, pre-computable, used for retrieval over millions of chunks. A cross-encoder feeds query and document together through the transformer and outputs a relevance score — more accurate but must run per pair, so it's used to rerank a short list.
What does quantisation trade off?
Lower precision weights (8/4-bit) cut memory 2–4× and speed up memory-bound decoding, at the cost of a small accuracy loss that grows at very low bit-widths. Group-wise scales and methods like GPTQ/AWQ keep the loss small at 4 bits.
Walk me through fine-tuning an LLM for a task.
Establish a prompted baseline and an eval set. Collect and clean examples in the chat format (deduplicate, balance labels, include hard cases, remove PII), split into train/validation/test without overlap. Train with a provider API or LoRA on an open model, watching validation loss for overfitting. Evaluate against the baseline on the held-out test set and check general abilities didn't regress. Ship behind a flag and monitor.
Practice¶
- Compute how many LoRA parameters you'd train for r=8 on the Q and V projections of all 32 layers of a model with d_model 4096.
- Run a small instruct model with Ollama and point the FastAPI chat endpoint at it via
base_url.
Next: LLMs — choosing, prompting, tool calling and evaluating models in practice.