Skip to content

5. Transformers for GenAI

Intermediate · 15 min read

You now know what's inside a transformer. This page connects it to the decisions you make as a GenAI engineer: which model to use, how a base model becomes a chat assistant, when to fine-tune, and how to run models cheaply.

5.1 From base model to assistant: the training stages

flowchart LR
    P[Pre-training<br/>trillions of tokens<br/>next-token prediction] --> B[Base model<br/>continues text]
    B --> S[Supervised fine-tuning SFT<br/>~10k–1M instruction/answer pairs]
    S --> I[Instruct model<br/>follows instructions]
    I --> R[Preference tuning<br/>RLHF / DPO]
    R --> C[Chat model<br/>helpful, harmless, honest]
    C --> RL[Optional: RL on verifiable tasks<br/>maths, code → reasoning models]
Stage Data What the model learns
Pre-training web pages, books, code — trillions of tokens, no labels language, facts, reasoning patterns; costs millions of dollars of GPU time
SFT (supervised fine-tuning) prompt → ideal response pairs, written or curated by people the format of being an assistant: answer the question, follow instructions, use the chat template
RLHF (reinforcement learning from human feedback) humans rank several answers; a reward model learns their preferences; the LLM is optimised against it which answers people prefer — helpfulness, tone, refusing harmful requests
DPO (direct preference optimisation) the same "chosen vs rejected" pairs the same goal as RLHF without a separate reward model — simpler, now very common
RL with verifiable rewards maths problems, code with tests long step-by-step reasoning ("thinking" models)

A base model asked "What is the capital of France?" may continue with "What is the capital of Germany?" — it's completing a list of quiz questions. The later stages are what make it answer.

5.2 Chat templates — what the model actually sees

Chat APIs take a list of messages, but the model reads one token sequence. A chat template flattens the messages with special marker tokens learned during SFT:

def apply_chat_template(messages: list[dict], add_generation_prompt: bool = True) -> str:
    """A simplified ChatML-style template (used by several open models)."""
    text = "".join(f"<|im_start|>{m['role']}\n{m['content']}<|im_end|>\n" for m in messages)
    if add_generation_prompt:
        text += "<|im_start|>assistant\n"            # the model continues from here
    return text

print(apply_chat_template([
    {"role": "system", "content": "You are a concise assistant."},
    {"role": "user", "content": "What is RAG?"},
]))
Output
<|im_start|>system
You are a concise assistant.<|im_end|>
<|im_start|>user
What is RAG?<|im_end|>
<|im_start|>assistant

Each model family has its own template (Llama, Mistral, Qwen and Gemma all differ). Using the wrong one with an open model silently degrades quality — always use the tokenizer's built-in template:

# no-run — pip install transformers torch
from transformers import AutoTokenizer, AutoModelForCausalLM

name = "Qwen/Qwen2.5-0.5B-Instruct"                  # small enough for a laptop CPU
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name)

messages = [{"role": "user", "content": "Explain RAG in one sentence."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs, max_new_tokens=60, do_sample=True, temperature=0.3, top_p=0.9)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

Tool calling works the same way: tool definitions and tool calls are rendered into the sequence with special tokens the model was fine-tuned on — the API's tools parameter is a convenience layer on top.

5.3 Embedding models: from token vectors to one vector

An encoder outputs one vector per token. To get one vector per sentence — for semantic search and RAG — the token vectors are pooled, usually by averaging (ignoring padding) and then normalising:

import numpy as np

rng = np.random.default_rng(0)
token_vectors = rng.normal(size=(2, 6, 8))           # 2 sentences, padded to 6 tokens, 8 dims
attention_mask = np.array([[1, 1, 1, 1, 0, 0],       # sentence 1 has 4 real tokens
                           [1, 1, 1, 1, 1, 1]])      # sentence 2 has 6

def mean_pool(vectors, mask):
    mask = mask[..., None]
    pooled = (vectors * mask).sum(axis=1) / mask.sum(axis=1)       # average real tokens only
    return pooled / np.linalg.norm(pooled, axis=1, keepdims=True)  # unit length → dot product = cosine

emb = mean_pool(token_vectors, attention_mask)
print(emb.shape, np.linalg.norm(emb, axis=1).round(3))
Output
(2, 8) [1. 1.]

Embedding models (e.g. the sentence-transformers family, BGE, E5, OpenAI's text-embedding-3-*) are transformers fine-tuned with contrastive learning — pulling matching question/passage pairs together and pushing others apart — so that this pooled vector captures meaning.

Need Model type Why
Embed chunks and queries for retrieval bi-encoder (embedding model) encode each text once, compare millions with a dot product
Re-order the top 20–50 results precisely cross-encoder (reranker) reads query + passage together through full attention — more accurate, too slow for millions
Classify or tag at high volume small fine-tuned encoder (BERT-style) fast, cheap, accurate with labelled data
Generate, reason, call tools decoder LLM —

5.4 Prompting, RAG or fine-tuning?

Your problem Best first tool
The model doesn't know your data (docs, tickets, products) RAG — knowledge changes; fine-tuning is a poor way to add facts
The model knows enough but the output format/style is off prompting with examples, structured output
A narrow task at high volume needs a smaller/cheaper/faster model fine-tuning (often distilling from a big model's outputs)
A consistent tone, persona or domain language prompts can't hold fine-tuning
Needs the latest information RAG / tools (web search, APIs)

Rule of thumb: prompt → RAG → fine-tune, in that order, with an eval set at each step to prove the next one is needed.

5.5 LoRA — fine-tuning without updating every weight

Full fine-tuning updates billions of parameters and needs the memory for their gradients and optimiser state (several times the model size). LoRA (low-rank adaptation) freezes the original weight matrix W and learns a small update ΔW = B @ A, where A and B are thin matrices of rank r (typically 8–64):

d = 4096                                   # one projection matrix in a 7–8B model
r = 16

W = rng.normal(scale=0.02, size=(d, d))    # frozen
A = rng.normal(scale=0.01, size=(r, d))    # trainable
B = np.zeros((d, r))                       # trainable, starts at 0 → the model starts unchanged

x = rng.normal(size=(1, d))
y = x @ W.T + x @ A.T @ B.T                # original path + low-rank update
print(np.allclose(y, x @ W.T))             # B = 0 → identical output before training

full, lora = W.size, A.size + B.size
print(f"full matrix: {full:,} params | LoRA r={r}: {lora:,} params ({lora / full:.2%})")
Output
True
full matrix: 16,777,216 params | LoRA r=16: 131,072 params (0.78%)
  • Training touches < 1% of the parameters → fits on one GPU; adapters are a few MB to share and swap.
  • After training, B @ A can be merged into W — zero extra inference cost.
  • QLoRA = LoRA on top of a 4-bit-quantised frozen base model → fine-tune a 7–8B model on a single 24 GB GPU, or even a free Colab GPU for small models.
  • Libraries: Hugging Face PEFT + TRL (SFTTrainer, DPOTrainer), Unsloth, Axolotl.

5.6 Quantisation — smaller, faster models

Weights are trained in 16- or 32-bit floats. Quantisation stores them in fewer bits — 8, 4, even 2 — with a scale factor to map back:

w = rng.normal(scale=0.02, size=4096).astype(np.float32)        # one row of weights

def quantise_int8(x):
    scale = np.abs(x).max() / 127                                # map the largest weight to ±127
    return np.round(x / scale).astype(np.int8), scale

def quantise_int4(x, group=64):                                  # 4-bit, one scale per group of 64
    g = x.reshape(-1, group)
    scale = np.abs(g).max(axis=1, keepdims=True) / 7
    return np.clip(np.round(g / scale), -8, 7).astype(np.int8), scale

q8, s8 = quantise_int8(w)
q4, s4 = quantise_int4(w)
for name, restored, bits in [("int8", q8 * s8, 8), ("int4", (q4 * s4).ravel(), 4)]:
    rel_err = np.abs(restored - w).mean() / np.abs(w).mean()
    print(f"{name}: {bits / 16:.0%} of fp16 size, mean relative error {rel_err:.1%}")
Output
int8: 50% of fp16 size, mean relative error 1.0%
int4: 25% of fp16 size, mean relative error 11.6%

An ~11% error per weight sounds large, but it is random noise that largely averages out across billions of weights — and real 4-bit methods (GPTQ, AWQ, GGUF Q4_K_M) are smarter than this naive rounding. They usually lose only a little quality, while cutting memory by 4× and speeding up decoding (which is limited by how fast weights are read from memory).

for params_b in (8, 70):
    print(f"{params_b}B model:", "  ".join(f"{name} ≈ {params_b * bits / 8:.0f} GB"
                                          for name, bits in [("fp16", 16), ("int8", 8), ("4-bit", 4)]))
Output
8B model: fp16 ≈ 16 GB  int8 ≈ 8 GB  4-bit ≈ 4 GB
70B model: fp16 ≈ 140 GB  int8 ≈ 70 GB  4-bit ≈ 35 GB

Weights only — add the KV cache from How LLMs generate text.

5.7 Serving open models

Tool Use it for
Ollama, LM Studio, llama.cpp running quantised (GGUF) models on a laptop; local dev and private demos
vLLM, SGLang, TGI production GPU serving: continuous batching, PagedAttention KV cache, OpenAI-compatible API
Hugging Face transformers experiments, fine-tuning, evaluation — not optimised for high-throughput serving

Serving tricks you'll hear about: continuous batching (new requests join a running batch every step), prefix caching (reuse the KV cache of a shared system prompt), and speculative decoding (a small model drafts several tokens, the big one verifies them in one pass — same output, 2–3× faster).

Because vLLM and Ollama expose an OpenAI-compatible endpoint, your app code barely changes:

# no-run — `ollama run llama3.1` must be running locally
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
r = client.chat.completions.create(model="llama3.1", messages=[{"role": "user", "content": "Hi!"}])
print(r.choices[0].message.content)

5.8 Fine-tuning in practice

Section 5.4 covered when to fine-tune; 5.5 covered LoRA. This is the workflow, step by step.

flowchart LR
    B[1. Baseline:<br/>best prompt + eval set] --> D[2. Collect &<br/>clean data]
    D --> S[3. Split<br/>train / validation / test]
    S --> T[4. Train<br/>API or LoRA]
    T --> E[5. Evaluate vs baseline<br/>on the held-out test set]
    E -->|better| SH[6. Ship + monitor]
    E -->|not better| D

1. Baseline first. Write the best prompt you can (with few-shot examples) and score it on an eval set. Fine-tuning has to beat this, not a lazy prompt.

2. Data. Training examples are full conversations in the model's chat format, usually as JSONL (one JSON object per line). Quality beats quantity: a few hundred clean, consistent, diverse examples often beat thousands of noisy ones. A common source is distillation — outputs from a large model, reviewed by people, used to train a small one.

import json, random

raw = [   # (ticket, label) pairs — in practice from labelled logs or reviewed LLM outputs
    ("App crashes when uploading a 20 MB PDF", "bug"), ("Please add dark mode", "feature_request"),
    ("Charged twice for March", "billing"), ("Login page shows error 500", "bug"),
    ("Can you support UPI autopay?", "feature_request"), ("Refund not received after 10 days", "billing"),
    ("Export to Excel is broken", "bug"), ("Invoice shows the wrong GST number", "billing"),
    ("Charged twice for March", "billing"),                                   # a duplicate
    ("", "bug"),                                                              # an empty example
]
SYSTEM = "Classify the support ticket as bug, feature_request or billing. Reply with the label only."

def to_example(ticket, label):
    return {"messages": [{"role": "system", "content": SYSTEM},
                         {"role": "user", "content": ticket},
                         {"role": "assistant", "content": label}]}

clean = list(dict.fromkeys((t.strip(), l) for t, l in raw if t.strip()))   # drop empties and duplicates
random.Random(0).shuffle(clean)
n_val = max(1, len(clean) // 5)
train, val = clean[n_val:], clean[:n_val]

with open("train.jsonl", "w") as f:
    for t, l in train:
        f.write(json.dumps(to_example(t, l)) + "\n")

print(f"{len(raw)} raw → {len(clean)} clean → {len(train)} train / {len(val)} validation")
print(open("train.jsonl").readline()[:120], "…")
Output
10 raw → 8 clean → 7 train / 1 validation
{"messages": [{"role": "system", "content": "Classify the support ticket as bug, feature_request or billing. Reply with  …

Checklist for the data: every label well represented, the hard and ambiguous cases included, no personal data you aren't allowed to use, and no overlap between training data and the test set (otherwise the scores lie).

3–4. Train. Two routes:

# no-run — OpenAI fine-tuning; other providers and hosts offer similar APIs
from openai import OpenAI
client = OpenAI()
file = client.files.create(file=open("train.jsonl", "rb"), purpose="fine-tune")
job = client.fine_tuning.jobs.create(training_file=file.id, model="gpt-4o-mini-2024-07-18")
# when the job finishes: use job.fine_tuned_model as the `model` in normal chat calls

No GPUs to manage; you pay per training token and a higher price per call; weights stay with the provider.

# no-run — pip install transformers peft trl datasets ; needs a GPU
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    train_dataset=load_dataset("json", data_files="train.jsonl")["train"],
    peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM"),
    args=SFTConfig(output_dir="ticket-classifier", num_train_epochs=3, learning_rate=2e-4),
)
trainer.train()

Full control and the weights are yours; you run the GPUs and the serving.

5. Evaluate. Score the fine-tuned model and the prompted baseline on the same held-out test set. Also check that general abilities didn't break — fine-tuning on a narrow task can make a model worse at everything else ("catastrophic forgetting").

Pitfall Symptom Fix
Overfitting training loss keeps falling, validation loss rises fewer epochs, more varied data
Inconsistent labels model is "randomly" wrong on similar inputs label guidelines, review disagreements
Train/test overlap great scores, poor production results deduplicate across splits
Prompt mismatch works in testing, not in the app use the same system prompt at training and inference
Trying to add knowledge still makes up facts use RAG for knowledge; fine-tune for behaviour

Interview questions

How does a base LLM become a chat assistant?

Pre-training on huge unlabelled text teaches next-token prediction. Supervised fine-tuning on instruction → response pairs (in a chat template) teaches it to follow instructions. Preference tuning (RLHF with a reward model, or DPO directly on chosen/rejected pairs) aligns it with human preferences. Reasoning models add RL on tasks with verifiable answers.

When would you fine-tune instead of using RAG?

RAG for knowledge — facts that change or must be cited. Fine-tuning for behaviour — a consistent format, style or domain language, or to make a small model do a narrow task cheaply. Often both: a fine-tuned model that is good at answering from retrieved context. Always justify with an eval set.

What is LoRA and why is it popular?

It freezes the pretrained weights and learns a low-rank update B·A for selected matrices, training well under 1% of the parameters. That cuts GPU memory and storage dramatically, adapters can be swapped per task, and the update can be merged for zero inference overhead. QLoRA applies it on a 4-bit base model.

Bi-encoder vs cross-encoder?

A bi-encoder embeds query and document separately into vectors — fast, pre-computable, used for retrieval over millions of chunks. A cross-encoder feeds query and document together through the transformer and outputs a relevance score — more accurate but must run per pair, so it's used to rerank a short list.

What does quantisation trade off?

Lower precision weights (8/4-bit) cut memory 2–4× and speed up memory-bound decoding, at the cost of a small accuracy loss that grows at very low bit-widths. Group-wise scales and methods like GPTQ/AWQ keep the loss small at 4 bits.

Walk me through fine-tuning an LLM for a task.

Establish a prompted baseline and an eval set. Collect and clean examples in the chat format (deduplicate, balance labels, include hard cases, remove PII), split into train/validation/test without overlap. Train with a provider API or LoRA on an open model, watching validation loss for overfitting. Evaluate against the baseline on the held-out test set and check general abilities didn't regress. Ship behind a flag and monitor.

Practice

  • Compute how many LoRA parameters you'd train for r=8 on the Q and V projections of all 32 layers of a model with d_model 4096.
  • Run a small instruct model with Ollama and point the FastAPI chat endpoint at it via base_url.

Next: LLMs — choosing, prompting, tool calling and evaluating models in practice.