3. Calling LLM APIs¶
Beginner → Intermediate · 12 min read
Every provider's chat API has the same shape: you send a list of messages, you get back a message plus usage numbers. Once you understand that shape, switching providers or adding streaming is a small change.
First call and API keys are covered in Call an LLM from Python; retries, timeouts and fallbacks in Handle LLM API errors. This page is about the design of the calls themselves.
3.1 Messages and roles¶
An LLM API is stateless: it remembers nothing between calls. The "conversation" is a list you send in full every time.
| Role | Who writes it | Purpose |
|---|---|---|
system |
you, the developer | instructions, persona, rules, output format — applies to the whole chat |
user |
the end user (or your code on their behalf) | the question or task; retrieved context often goes here too |
assistant |
the model (earlier replies) | conversation history; you can also write one to show an example |
tool |
your code | the result of a tool the model asked to call (topic 6) |
messages = [
{"role": "system", "content": "You are a support assistant for Acme. Answer in at most 2 sentences."},
{"role": "user", "content": "Do you ship to Pune?"},
{"role": "assistant", "content": "Yes — we ship across India, usually within 3–5 days."},
{"role": "user", "content": "And how much does it cost?"}, # "it" only makes sense with history
]
print(len(messages), "messages;", sum(len(m["content"]) for m in messages), "characters sent on this call")
Because the whole history is sent every time, long chats get more expensive per turn — managing that is covered in Context, memory & cost.
3.2 The parameters that matter¶
| Parameter | What it does | Typical setting |
|---|---|---|
model |
which model | the cheapest one that passes your evals |
max_tokens / max_completion_tokens |
cap on output length (and cost) | always set it |
temperature |
randomness (how it works) | 0–0.3 for facts, extraction, code; 0.7+ for creative |
top_p |
nucleus sampling cut-off | leave default; change temperature or top_p, not both |
stop |
strings that end generation | e.g. ["\n\n"] for one-paragraph answers |
stream |
send tokens as they're generated | True for chat UIs |
response_format / tools |
structured JSON or tool calls | topic 6 |
seed |
best-effort repeatability (some providers) | tests and evals |
timeout |
client-side time limit | 20–60 s, longer for reasoning models |
3.3 OpenAI and Anthropic side by side¶
The ideas are identical; the field names differ slightly:
# no-run — needs OPENAI_API_KEY
from openai import OpenAI
client = OpenAI(timeout=30, max_retries=2)
r = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "system", "content": "Be concise."},
{"role": "user", "content": "What is RAG?"}],
temperature=0.2,
max_tokens=200,
)
print(r.choices[0].message.content)
print(r.usage.prompt_tokens, r.usage.completion_tokens)
print(r.choices[0].finish_reason) # "stop" or "length"
# no-run — needs ANTHROPIC_API_KEY
import anthropic
client = anthropic.Anthropic(timeout=30, max_retries=2)
r = client.messages.create(
model="claude-sonnet-5-5",
system="Be concise.", # system prompt is a separate field
messages=[{"role": "user", "content": "What is RAG?"}],
temperature=0.2,
max_tokens=200, # required
)
print(r.content[0].text) # content is a list of blocks
print(r.usage.input_tokens, r.usage.output_tokens)
print(r.stop_reason) # "end_turn" or "max_tokens"
Always check why generation stopped
finish_reason == "length" (OpenAI) or stop_reason == "max_tokens" (Anthropic) means the answer was
cut off. For JSON this means invalid output; for text, a truncated answer shown to the user. Detect it and
retry with a higher limit or ask for a shorter answer.
3.4 A provider-agnostic client¶
Don't scatter SDK calls across your codebase. Wrap them behind one small interface, so you can switch providers, add logging and cost tracking in one place, and use a fake in tests:
from dataclasses import dataclass, field
from typing import Protocol
import time
@dataclass
class LLMResponse:
text: str
model: str
input_tokens: int
output_tokens: int
latency_s: float
truncated: bool = False
class LLMClient(Protocol):
def complete(self, messages: list[dict], *, max_tokens: int = 512, temperature: float = 0.2) -> LLMResponse: ...
@dataclass
class FakeLLM:
"""Deterministic stand-in for tests and examples — replies from a lookup table."""
replies: dict[str, str] = field(default_factory=dict)
calls: int = 0
def complete(self, messages, *, max_tokens=512, temperature=0.2) -> LLMResponse:
self.calls += 1
start = time.perf_counter()
question = messages[-1]["content"]
text = next((a for q, a in self.replies.items() if q in question.lower()), "I don't know.")
words = text.split()
truncated = len(words) > max_tokens
return LLMResponse(" ".join(words[:max_tokens]), "fake-1", sum(len(m["content"].split()) for m in messages),
min(len(words), max_tokens), round(time.perf_counter() - start, 1), truncated)
llm: LLMClient = FakeLLM({"rag": "RAG retrieves relevant documents and passes them to the LLM as context.",
"ship": "We ship across India in 3-5 days."})
r = llm.complete([{"role": "user", "content": "What is RAG?"}])
print(r.text)
print(r.input_tokens, r.output_tokens, r.truncated)
print(llm.complete([{"role": "user", "content": "What is RAG?"}], max_tokens=4))
RAG retrieves relevant documents and passes them to the LLM as context.
3 12 False
LLMResponse(text='RAG retrieves relevant documents', model='fake-1', input_tokens=3, output_tokens=4, latency_s=0.0, truncated=True)
The real implementation for one provider:
# no-run
from openai import OpenAI
class OpenAIClient:
def __init__(self, model: str = "gpt-4o-mini"):
self.client, self.model = OpenAI(timeout=30, max_retries=2), model
def complete(self, messages, *, max_tokens=512, temperature=0.2) -> LLMResponse:
start = time.perf_counter()
r = self.client.chat.completions.create(model=self.model, messages=messages,
max_tokens=max_tokens, temperature=temperature)
choice = r.choices[0]
return LLMResponse(choice.message.content or "", r.model, r.usage.prompt_tokens,
r.usage.completion_tokens, time.perf_counter() - start,
truncated=choice.finish_reason == "length")
Libraries like LiteLLM provide this abstraction across 100+ providers if you'd rather not maintain it.
3.5 Streaming¶
Streaming sends tokens as they are generated. Total time is the same, but the user sees the answer start in under a second instead of waiting for all of it:
def fake_stream(text: str):
for word in text.split(" "):
time.sleep(0.01)
yield word + " "
start = time.perf_counter()
first_token_at = None
answer = ""
for chunk in fake_stream("Streaming makes slow answers feel fast to the user."):
if first_token_at is None:
first_token_at = time.perf_counter() - start
answer += chunk
total = time.perf_counter() - start
print(answer.strip())
print(f"first token after {first_token_at:.2f}s, full answer after {total:.2f}s") # timings vary per run
Streaming makes slow answers feel fast to the user.
first token after 0.01s, full answer after 0.09s
With the OpenAI SDK, pass stream=True and read chunk.choices[0].delta.content; usage arrives in the last
chunk if you set stream_options={"include_usage": True}. Serving a stream to a browser is covered in
FastAPI async & streaming.
3.6 Parallel calls¶
Independent LLM calls — summarising 20 documents, classifying a batch, running 3 judges — should run concurrently, with a limit so you don't hit rate limits:
import asyncio
async def fake_call(doc: str) -> str:
await asyncio.sleep(0.2) # stands in for an async SDK call (~seconds in reality)
return f"summary of {doc}"
async def summarise_all(docs: list[str], max_concurrent: int = 5) -> list[str]:
sem = asyncio.Semaphore(max_concurrent)
async def one(doc):
async with sem: # at most 5 calls in flight
return await fake_call(doc)
return await asyncio.gather(*(one(d) for d in docs))
docs = [f"doc{i}" for i in range(20)]
start = time.perf_counter()
summaries = asyncio.run(summarise_all(docs))
print(len(summaries), summaries[0], f"in {time.perf_counter() - start:.1f}s (sequential would be 4.0s)") # timings vary slightly
Use the async clients (AsyncOpenAI, AsyncAnthropic). For thousands of non-urgent requests (nightly evals,
bulk enrichment), the providers' batch APIs are usually around half price.
Interview questions¶
LLM APIs are stateless — how does a chatbot remember the conversation?
The application stores the message history and sends it (or a trimmed/summarised version) with every request. The model only "remembers" what is in the current request's context.
What would you put in a wrapper around the LLM SDK?
One interface for all providers; timeouts and retries; logging of model, tokens, latency and cost per call; truncation detection; a fallback provider; and the ability to swap in a fake for tests.
How do you make many LLM calls fast without hitting rate limits?
Use async clients with asyncio.gather and a semaphore to cap concurrency; back off on 429s; use batch APIs
for offline jobs; and cache repeated requests.
Practice¶
- Add a
cost_usdproperty toLLMResponseusing a price table. - Write an
AnthropicClientwith the samecomplete()signature asOpenAIClient.
Next: Prompt engineering — writing prompts that work reliably.