2. Choosing a model¶
Beginner · 10 min read
There is no "best LLM" — there is the best model for a task, a budget and a latency target. Picking well is one of the highest-leverage decisions in a GenAI project: the wrong choice can make an app 20× more expensive or noticeably worse, with no code changes at all.
Model names change every few months
The names and prices below are examples to show how to reason. Always check the provider's current model list and pricing page — the reasoning stays the same.
2.1 Closed (API) models vs open-weight models¶
| Closed / API models | Open-weight models | |
|---|---|---|
| Examples | OpenAI GPT, Anthropic Claude, Google Gemini | Llama, Mistral, Qwen, Gemma, DeepSeek |
| How you use them | HTTP API, pay per token | download weights; run yourself or via a host (Groq, Together, Fireworks, Bedrock) |
| Quality at the top end | usually the strongest | close behind, improving fast |
| Setup | minutes | GPUs, serving stack (vLLM, Ollama) — or a hosted API |
| Data control | data goes to the provider (check retention terms) | can stay entirely on your servers |
| Customisation | prompting; limited fine-tuning | full fine-tuning, LoRA, quantisation |
| Cost shape | per token — cheap to start, grows with usage | fixed GPU cost — cheaper at high, steady volume |
A common pattern: prototype on a strong API model, then move high-volume, narrow tasks to a smaller or open model once you have evals proving it's good enough.
2.2 Model tiers¶
Providers offer families in several sizes. Roughly:
| Tier | Typical use | Relative speed | Relative cost |
|---|---|---|---|
| Small / fast ("mini", "flash", "haiku", 7–8B open models) | classification, routing, extraction, summaries, simple chat, high volume | fastest | 1× |
| Mid ("sonnet", "pro", 70B open models) | most production apps: RAG answers, agents, coding help | medium | ~5–20× |
| Large / frontier ("opus", top GPT/Gemini) | hard reasoning, complex agents, difficult code, evaluation (as a judge) | slower | ~20–100× |
| Reasoning / thinking modes | maths, multi-step planning, tricky logic — spends extra "thinking" tokens first | slowest | pay for thinking tokens too |
Start with the smallest model that passes your evals, not the biggest one available.
2.3 Context window¶
The context window is the maximum number of tokens — prompt plus output — the model can handle in one call. Today it ranges from ~8K to 1M+ tokens.
def fits(prompt_tokens: int, max_output: int, context_window: int) -> str:
total = prompt_tokens + max_output
status = "fits" if total <= context_window else "TOO LONG"
return f"{prompt_tokens:>7,} + {max_output:>5,} = {total:>7,} / {context_window:>9,} → {status}"
print(fits(6_000, 1_000, 8_192))
print(fits(7_500, 1_000, 8_192))
print(fits(150_000, 4_000, 200_000))
6,000 + 1,000 = 7,000 / 8,192 → fits
7,500 + 1,000 = 8,500 / 8,192 → TOO LONG
150,000 + 4,000 = 154,000 / 200,000 → fits
A big window is not a reason to fill it
Long prompts cost more (you pay for every input token), are slower (time to first token grows), and models use information in the middle of a long context less reliably than at the start or end ("lost in the middle"). Retrieving the right 2–5 chunks usually beats pasting 200 pages. See Context, memory & cost.
There is also a separate maximum output limit (often 4K–64K tokens) — long reports may need to be generated in sections.
2.4 Pricing and estimating cost¶
APIs charge per million tokens, with output tokens costing several times more than input tokens (the reason is in How LLMs generate text). Estimate before you build:
PRICES = { # USD per 1M tokens (input, output) — illustrative only
"small": (0.15, 0.60),
"mid": (3.00, 15.00),
"large": (15.00, 75.00),
}
def monthly_cost(tier: str, requests_per_day: int, in_tokens: int, out_tokens: int) -> float:
p_in, p_out = PRICES[tier]
per_request = (in_tokens * p_in + out_tokens * p_out) / 1_000_000
return per_request * requests_per_day * 30
# A RAG support bot: 2,000-token prompt (question + 4 chunks + system), 300-token answer, 5,000 requests/day
for tier in PRICES:
print(f"{tier:6} ${monthly_cost(tier, 5_000, 2_000, 300):>9,.2f} / month")
Same app, 100× cost difference between tiers. That's why model choice — and routing easy requests to a cheaper model — matters so much.
Hidden cost multipliers to remember:
- Agents make several LLM calls per user request (plan, tool call, tool call, answer) — multiply accordingly.
- Reasoning models bill their thinking tokens as output tokens.
- Retries and LLM-as-judge evals add calls.
- Prompt caching and batch APIs (often ~50% off for non-urgent jobs) reduce cost.
2.5 Benchmarks vs your own evals¶
Public benchmarks are useful for a shortlist, not for a decision:
| Benchmark type | Examples | Tells you |
|---|---|---|
| Knowledge & reasoning | MMLU, GPQA | broad knowledge, expert-level questions |
| Maths | GSM8K, AIME | multi-step reasoning |
| Code | HumanEval, SWE-bench | writing and fixing code |
| Human preference | LMArena (Chatbot Arena) | which answers people prefer in chat |
| Long context | needle-in-a-haystack, RULER | finding information in long inputs |
Problems: models may have seen benchmark questions in training, scores saturate near 100%, and none of them is your task. The deciding test is always a small eval set of your own — 30–100 real inputs with expected outputs — run against 2–4 candidate models. See Evaluation & guardrails.
2.6 A selection checklist¶
CANDIDATES = [
{"name": "small-api", "quality": 0.82, "p95_latency_s": 1.1, "cost_month": 72, "data_stays_inhouse": False},
{"name": "mid-api", "quality": 0.91, "p95_latency_s": 2.4, "cost_month": 1575, "data_stays_inhouse": False},
{"name": "open-8b-self-hosted", "quality": 0.79, "p95_latency_s": 0.9, "cost_month": 600, "data_stays_inhouse": True},
{"name": "open-70b-self-hosted", "quality": 0.88, "p95_latency_s": 2.0, "cost_month": 2400, "data_stays_inhouse": True},
]
def choose(candidates, min_quality, max_latency_s, budget, need_inhouse=False):
ok = [c for c in candidates
if c["quality"] >= min_quality and c["p95_latency_s"] <= max_latency_s
and c["cost_month"] <= budget and (c["data_stays_inhouse"] or not need_inhouse)]
return min(ok, key=lambda c: c["cost_month"])["name"] if ok else "nothing fits — relax a requirement"
print(choose(CANDIDATES, min_quality=0.80, max_latency_s=3, budget=2000))
print(choose(CANDIDATES, min_quality=0.85, max_latency_s=3, budget=2000))
print(choose(CANDIDATES, min_quality=0.85, max_latency_s=3, budget=3000, need_inhouse=True))
print(choose(CANDIDATES, min_quality=0.95, max_latency_s=3, budget=5000))
quality here is the score on your eval set. The questions to answer for every project:
- Quality — what score on my eval set is good enough?
- Latency — what p95 response time can users accept? Does it need streaming?
- Cost — requests/day × tokens/request × price; what's the budget?
- Context — how many tokens do the biggest prompts need?
- Data — can this data leave our servers? Which region? What retention terms?
- Features — tool calling, structured output, vision/audio input, batch API, prompt caching?
- Fallback — which second provider if the first is down or rate-limits us?
Interview questions¶
How do you choose an LLM for a new product feature?
Define the task and success criteria, build a small eval set from real inputs, shortlist 2–4 models across tiers (and an open model if data control matters), and measure quality, latency and cost per request on the eval set. Pick the cheapest model that meets the bar, and plan a fallback provider and a route to a bigger model for hard cases.
When would you choose an open-weight model?
When data must stay in-house, when volume is high and steady enough that fixed GPU cost beats per-token pricing, when you need deep customisation (fine-tuning, quantisation), or to avoid vendor lock-in. The cost is running and scaling the serving infrastructure yourself (or paying a host).
Why not always use the model with the biggest context window?
Cost and latency grow with prompt length, and models use information buried in long contexts less reliably. Good retrieval that sends only relevant context is usually cheaper, faster and more accurate.
Practice¶
- Estimate the monthly cost of your current project with the
monthly_costfunction and real prices from your provider's pricing page. - Write down your project's answers to the checklist in 1.6.
Next: Calling LLM APIs — messages, parameters, streaming and a client you can swap.