5. Reasoning models¶
Intermediate · 10 min read
A reasoning model (also called a "thinking" model) generates a long internal chain of thought before it writes the answer. It trades time and tokens for accuracy on problems that need several steps: maths, planning, tricky code, analysing a contract against a policy, deciding what an agent should do next.
5.1 How they differ from standard models¶
Prompt engineering showed that asking a standard model to "think step by step" improves multi-step answers — every generated token can build on the ones before. Reasoning models take this much further:
| Standard model + "think step by step" | Reasoning model | |
|---|---|---|
| Who decides to reason | you, through the prompt | the model, trained to do it |
| How it learned | imitation of reasoning in training data | reinforcement learning on problems with checkable answers (maths, code tests) |
| Length of reasoning | usually short | can be thousands of tokens: tries approaches, checks itself, backtracks |
| Visibility | the reasoning is in the output | often hidden or summarised; you pay for it anyway |
| Control | prompt wording | an effort level or thinking-token budget parameter |
The idea behind them: spending more compute at answer time ("test-time compute") can improve results as much as training a bigger model.
5.2 Thinking tokens cost money and time¶
Reasoning tokens are billed as output tokens and generated one at a time, so they dominate cost and latency:
PRICE_OUT = 10.00 # USD per 1M output tokens, illustrative
TOKENS_PER_SECOND = 80 # decode speed, illustrative
def request_profile(answer_tokens: int, thinking_tokens: int) -> str:
total = answer_tokens + thinking_tokens
cost = total * PRICE_OUT / 1_000_000
seconds = total / TOKENS_PER_SECOND
return f"{total:>6,} output tokens ${cost:.4f} ~{seconds:5.1f}s"
print("no thinking ", request_profile(300, 0))
print("low effort ", request_profile(300, 1_000))
print("high effort ", request_profile(300, 12_000))
no thinking 300 output tokens $0.0030 ~ 3.8s
low effort 1,300 output tokens $0.0130 ~ 16.2s
high effort 12,300 output tokens $0.1230 ~153.8s
Same visible answer, 40× the cost and a 2½-minute wait at high effort. Always check usage — providers
report reasoning tokens separately (for example completion_tokens_details.reasoning_tokens in OpenAI's API).
5.3 Controlling reasoning in the APIs¶
# no-run — needs an API key and a reasoning-capable model
from openai import OpenAI
client = OpenAI()
r = client.chat.completions.create(
model="o4-mini", # check the current reasoning models
reasoning_effort="medium", # "low" | "medium" | "high"
messages=[{"role": "user", "content": "Plan a 3-city delivery route that minimises distance: ..."}],
max_completion_tokens=8000, # must leave room for reasoning AND the answer
)
print(r.choices[0].message.content)
print(r.usage.completion_tokens_details.reasoning_tokens, "reasoning tokens")
# no-run — needs an API key; extended thinking on a supported Claude model
import anthropic
client = anthropic.Anthropic()
r = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000, # must be larger than the thinking budget
thinking={"type": "enabled", "budget_tokens": 8000},
messages=[{"role": "user", "content": "Plan a 3-city delivery route that minimises distance: ..."}],
)
for block in r.content:
print(block.type) # "thinking" blocks, then "text"
Parameter names change between model generations — check the provider docs. Two rules hold everywhere:
- Set the output limit high enough for reasoning plus answer, or you get a truncated (or empty) answer.
- With tool use, pass the thinking blocks back unchanged in the conversation, as the provider's docs require.
5.4 Prompting reasoning models¶
Prompting habits from standard models can hurt here:
| Do | Avoid |
|---|---|
| State the goal, constraints and what a good answer looks like | Scripting the steps ("first do X, then Y…") — the model plans better itself |
| Give all the relevant data and context | Adding "think step by step" — it already does |
| Ask for the final format you need (JSON, a table) | Long few-shot examples of reasoning — often unnecessary |
| Use a low effort first; raise it only if evals improve | Defaulting to maximum effort everywhere |
5.5 Self-consistency: a cheap reasoning boost for any model¶
Sample several answers at a moderate temperature and take the majority vote. Independent reasoning paths that agree are more likely to be right:
from collections import Counter
import random
def fake_model_answer(question: str, rng: random.Random) -> str:
"""Stand-in for an LLM that is right 60% of the time and wrong in different ways otherwise."""
return "42" if rng.random() < 0.6 else rng.choice(["24", "40", "44", "48"])
def self_consistency(question: str, n: int, rng: random.Random) -> str:
votes = Counter(fake_model_answer(question, rng) for _ in range(n))
return votes.most_common(1)[0][0]
rng = random.Random(0)
trials = 1000
single = sum(fake_model_answer("q", rng) == "42" for _ in range(trials)) / trials
voted = sum(self_consistency("q", 5, rng) == "42" for _ in range(trials)) / trials
print(f"one sample: {single:.0%} correct majority of 5: {voted:.0%} correct")
It works when answers can be compared (a number, a label, a choice) and costs n× the tokens. It's the same idea reasoning models apply internally.
5.6 When to use a reasoning model¶
| Use one for | Don't, for |
|---|---|
| Maths, logic, puzzles, scheduling and optimisation | Extraction, classification, summarisation, rewriting |
| Hard code: debugging, refactoring across files, algorithms | Simple chat and FAQ answers |
| Multi-step analysis: "does this contract violate our policy?" | Anything where latency must be under a couple of seconds |
| Agent planning steps and hard decisions | High-volume, low-margin requests |
| Grading / judging complex answers (LLM-as-judge) | Tasks a fast model already passes on your eval set |
A common architecture: a reasoning model plans, fast models execute the individual steps — or a fast model answers first and escalates to a reasoning model only when validation fails.
Interview questions¶
What is a reasoning model and how is it trained?
An LLM trained — largely with reinforcement learning on tasks with verifiable answers — to produce a long chain of thought before its final answer. It learns to explore approaches, check its work and backtrack. This spends more compute at inference time to get better answers on multi-step problems.
What are the trade-offs of reasoning models?
Better accuracy on complex problems, at the cost of many more output tokens (billed), much higher latency, and sometimes over-thinking simple tasks. Control it with effort/budget settings, use it only where evals show a gain, and leave enough output-token room for reasoning plus the answer.
What is self-consistency?
Sampling several independent answers to the same question and taking the majority. It improves accuracy for questions with comparable answers, at n× the cost.
Practice¶
- Run 10 maths word problems through a fast model and a reasoning model at low effort; compare accuracy, tokens and time.
- Change
fake_model_answerto be right only 40% of the time. Does majority voting still help? Why?
Next: Structured output & tool calling — data and actions, not just text.