Skip to content

2. Choosing a model

Beginner · 10 min read

There is no "best LLM" — there is the best model for a task, a budget and a latency target. Picking well is one of the highest-leverage decisions in a GenAI project: the wrong choice can make an app 20× more expensive or noticeably worse, with no code changes at all.

Model names change every few months

The names and prices below are examples to show how to reason. Always check the provider's current model list and pricing page — the reasoning stays the same.

2.1 Closed (API) models vs open-weight models

Closed / API models Open-weight models
Examples OpenAI GPT, Anthropic Claude, Google Gemini Llama, Mistral, Qwen, Gemma, DeepSeek
How you use them HTTP API, pay per token download weights; run yourself or via a host (Groq, Together, Fireworks, Bedrock)
Quality at the top end usually the strongest close behind, improving fast
Setup minutes GPUs, serving stack (vLLM, Ollama) — or a hosted API
Data control data goes to the provider (check retention terms) can stay entirely on your servers
Customisation prompting; limited fine-tuning full fine-tuning, LoRA, quantisation
Cost shape per token — cheap to start, grows with usage fixed GPU cost — cheaper at high, steady volume

A common pattern: prototype on a strong API model, then move high-volume, narrow tasks to a smaller or open model once you have evals proving it's good enough.

2.2 Model tiers

Providers offer families in several sizes. Roughly:

Tier Typical use Relative speed Relative cost
Small / fast ("mini", "flash", "haiku", 7–8B open models) classification, routing, extraction, summaries, simple chat, high volume fastest 1×
Mid ("sonnet", "pro", 70B open models) most production apps: RAG answers, agents, coding help medium ~5–20×
Large / frontier ("opus", top GPT/Gemini) hard reasoning, complex agents, difficult code, evaluation (as a judge) slower ~20–100×
Reasoning / thinking modes maths, multi-step planning, tricky logic — spends extra "thinking" tokens first slowest pay for thinking tokens too

Start with the smallest model that passes your evals, not the biggest one available.

2.3 Context window

The context window is the maximum number of tokens — prompt plus output — the model can handle in one call. Today it ranges from ~8K to 1M+ tokens.

def fits(prompt_tokens: int, max_output: int, context_window: int) -> str:
    total = prompt_tokens + max_output
    status = "fits" if total <= context_window else "TOO LONG"
    return f"{prompt_tokens:>7,} + {max_output:>5,} = {total:>7,} / {context_window:>9,} → {status}"

print(fits(6_000, 1_000, 8_192))
print(fits(7_500, 1_000, 8_192))
print(fits(150_000, 4_000, 200_000))
Output
  6,000 + 1,000 =   7,000 /     8,192 → fits
  7,500 + 1,000 =   8,500 /     8,192 → TOO LONG
150,000 + 4,000 = 154,000 /   200,000 → fits

A big window is not a reason to fill it

Long prompts cost more (you pay for every input token), are slower (time to first token grows), and models use information in the middle of a long context less reliably than at the start or end ("lost in the middle"). Retrieving the right 2–5 chunks usually beats pasting 200 pages. See Context, memory & cost.

There is also a separate maximum output limit (often 4K–64K tokens) — long reports may need to be generated in sections.

2.4 Pricing and estimating cost

APIs charge per million tokens, with output tokens costing several times more than input tokens (the reason is in How LLMs generate text). Estimate before you build:

PRICES = {                       # USD per 1M tokens (input, output) — illustrative only
    "small": (0.15, 0.60),
    "mid": (3.00, 15.00),
    "large": (15.00, 75.00),
}

def monthly_cost(tier: str, requests_per_day: int, in_tokens: int, out_tokens: int) -> float:
    p_in, p_out = PRICES[tier]
    per_request = (in_tokens * p_in + out_tokens * p_out) / 1_000_000
    return per_request * requests_per_day * 30

# A RAG support bot: 2,000-token prompt (question + 4 chunks + system), 300-token answer, 5,000 requests/day
for tier in PRICES:
    print(f"{tier:6} ${monthly_cost(tier, 5_000, 2_000, 300):>9,.2f} / month")
Output
small  $    72.00 / month
mid    $ 1,575.00 / month
large  $ 7,875.00 / month

Same app, 100× cost difference between tiers. That's why model choice — and routing easy requests to a cheaper model — matters so much.

Hidden cost multipliers to remember:

  • Agents make several LLM calls per user request (plan, tool call, tool call, answer) — multiply accordingly.
  • Reasoning models bill their thinking tokens as output tokens.
  • Retries and LLM-as-judge evals add calls.
  • Prompt caching and batch APIs (often ~50% off for non-urgent jobs) reduce cost.

2.5 Benchmarks vs your own evals

Public benchmarks are useful for a shortlist, not for a decision:

Benchmark type Examples Tells you
Knowledge & reasoning MMLU, GPQA broad knowledge, expert-level questions
Maths GSM8K, AIME multi-step reasoning
Code HumanEval, SWE-bench writing and fixing code
Human preference LMArena (Chatbot Arena) which answers people prefer in chat
Long context needle-in-a-haystack, RULER finding information in long inputs

Problems: models may have seen benchmark questions in training, scores saturate near 100%, and none of them is your task. The deciding test is always a small eval set of your own — 30–100 real inputs with expected outputs — run against 2–4 candidate models. See Evaluation & guardrails.

2.6 A selection checklist

CANDIDATES = [
    {"name": "small-api", "quality": 0.82, "p95_latency_s": 1.1, "cost_month": 72, "data_stays_inhouse": False},
    {"name": "mid-api", "quality": 0.91, "p95_latency_s": 2.4, "cost_month": 1575, "data_stays_inhouse": False},
    {"name": "open-8b-self-hosted", "quality": 0.79, "p95_latency_s": 0.9, "cost_month": 600, "data_stays_inhouse": True},
    {"name": "open-70b-self-hosted", "quality": 0.88, "p95_latency_s": 2.0, "cost_month": 2400, "data_stays_inhouse": True},
]

def choose(candidates, min_quality, max_latency_s, budget, need_inhouse=False):
    ok = [c for c in candidates
          if c["quality"] >= min_quality and c["p95_latency_s"] <= max_latency_s
          and c["cost_month"] <= budget and (c["data_stays_inhouse"] or not need_inhouse)]
    return min(ok, key=lambda c: c["cost_month"])["name"] if ok else "nothing fits — relax a requirement"

print(choose(CANDIDATES, min_quality=0.80, max_latency_s=3, budget=2000))
print(choose(CANDIDATES, min_quality=0.85, max_latency_s=3, budget=2000))
print(choose(CANDIDATES, min_quality=0.85, max_latency_s=3, budget=3000, need_inhouse=True))
print(choose(CANDIDATES, min_quality=0.95, max_latency_s=3, budget=5000))
Output
small-api
mid-api
open-70b-self-hosted
nothing fits — relax a requirement

quality here is the score on your eval set. The questions to answer for every project:

  • Quality — what score on my eval set is good enough?
  • Latency — what p95 response time can users accept? Does it need streaming?
  • Cost — requests/day × tokens/request × price; what's the budget?
  • Context — how many tokens do the biggest prompts need?
  • Data — can this data leave our servers? Which region? What retention terms?
  • Features — tool calling, structured output, vision/audio input, batch API, prompt caching?
  • Fallback — which second provider if the first is down or rate-limits us?

Interview questions

How do you choose an LLM for a new product feature?

Define the task and success criteria, build a small eval set from real inputs, shortlist 2–4 models across tiers (and an open model if data control matters), and measure quality, latency and cost per request on the eval set. Pick the cheapest model that meets the bar, and plan a fallback provider and a route to a bigger model for hard cases.

When would you choose an open-weight model?

When data must stay in-house, when volume is high and steady enough that fixed GPU cost beats per-token pricing, when you need deep customisation (fine-tuning, quantisation), or to avoid vendor lock-in. The cost is running and scaling the serving infrastructure yourself (or paying a host).

Why not always use the model with the biggest context window?

Cost and latency grow with prompt length, and models use information buried in long contexts less reliably. Good retrieval that sends only relevant context is usually cheaper, faster and more accurate.

Practice

  • Estimate the monthly cost of your current project with the monthly_cost function and real prices from your provider's pricing page.
  • Write down your project's answers to the checklist in 1.6.

Next: Calling LLM APIs — messages, parameters, streaming and a client you can swap.