Skip to content

3. Calling LLM APIs

Beginner → Intermediate · 12 min read

Every provider's chat API has the same shape: you send a list of messages, you get back a message plus usage numbers. Once you understand that shape, switching providers or adding streaming is a small change.

First call and API keys are covered in Call an LLM from Python; retries, timeouts and fallbacks in Handle LLM API errors. This page is about the design of the calls themselves.

3.1 Messages and roles

An LLM API is stateless: it remembers nothing between calls. The "conversation" is a list you send in full every time.

Role Who writes it Purpose
system you, the developer instructions, persona, rules, output format — applies to the whole chat
user the end user (or your code on their behalf) the question or task; retrieved context often goes here too
assistant the model (earlier replies) conversation history; you can also write one to show an example
tool your code the result of a tool the model asked to call (topic 6)
messages = [
    {"role": "system", "content": "You are a support assistant for Acme. Answer in at most 2 sentences."},
    {"role": "user", "content": "Do you ship to Pune?"},
    {"role": "assistant", "content": "Yes — we ship across India, usually within 3–5 days."},
    {"role": "user", "content": "And how much does it cost?"},          # "it" only makes sense with history
]
print(len(messages), "messages;", sum(len(m["content"]) for m in messages), "characters sent on this call")
Output
4 messages; 166 characters sent on this call

Because the whole history is sent every time, long chats get more expensive per turn — managing that is covered in Context, memory & cost.

3.2 The parameters that matter

Parameter What it does Typical setting
model which model the cheapest one that passes your evals
max_tokens / max_completion_tokens cap on output length (and cost) always set it
temperature randomness (how it works) 0–0.3 for facts, extraction, code; 0.7+ for creative
top_p nucleus sampling cut-off leave default; change temperature or top_p, not both
stop strings that end generation e.g. ["\n\n"] for one-paragraph answers
stream send tokens as they're generated True for chat UIs
response_format / tools structured JSON or tool calls topic 6
seed best-effort repeatability (some providers) tests and evals
timeout client-side time limit 20–60 s, longer for reasoning models

3.3 OpenAI and Anthropic side by side

The ideas are identical; the field names differ slightly:

# no-run — needs OPENAI_API_KEY
from openai import OpenAI
client = OpenAI(timeout=30, max_retries=2)

r = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "system", "content": "Be concise."},
              {"role": "user", "content": "What is RAG?"}],
    temperature=0.2,
    max_tokens=200,
)
print(r.choices[0].message.content)
print(r.usage.prompt_tokens, r.usage.completion_tokens)
print(r.choices[0].finish_reason)            # "stop" or "length"
# no-run — needs ANTHROPIC_API_KEY
import anthropic
client = anthropic.Anthropic(timeout=30, max_retries=2)

r = client.messages.create(
    model="claude-sonnet-5-5",
    system="Be concise.",                      # system prompt is a separate field
    messages=[{"role": "user", "content": "What is RAG?"}],
    temperature=0.2,
    max_tokens=200,                            # required
)
print(r.content[0].text)                       # content is a list of blocks
print(r.usage.input_tokens, r.usage.output_tokens)
print(r.stop_reason)                           # "end_turn" or "max_tokens"

Always check why generation stopped

finish_reason == "length" (OpenAI) or stop_reason == "max_tokens" (Anthropic) means the answer was cut off. For JSON this means invalid output; for text, a truncated answer shown to the user. Detect it and retry with a higher limit or ask for a shorter answer.

3.4 A provider-agnostic client

Don't scatter SDK calls across your codebase. Wrap them behind one small interface, so you can switch providers, add logging and cost tracking in one place, and use a fake in tests:

from dataclasses import dataclass, field
from typing import Protocol
import time

@dataclass
class LLMResponse:
    text: str
    model: str
    input_tokens: int
    output_tokens: int
    latency_s: float
    truncated: bool = False

class LLMClient(Protocol):
    def complete(self, messages: list[dict], *, max_tokens: int = 512, temperature: float = 0.2) -> LLMResponse: ...

@dataclass
class FakeLLM:
    """Deterministic stand-in for tests and examples — replies from a lookup table."""
    replies: dict[str, str] = field(default_factory=dict)
    calls: int = 0

    def complete(self, messages, *, max_tokens=512, temperature=0.2) -> LLMResponse:
        self.calls += 1
        start = time.perf_counter()
        question = messages[-1]["content"]
        text = next((a for q, a in self.replies.items() if q in question.lower()), "I don't know.")
        words = text.split()
        truncated = len(words) > max_tokens
        return LLMResponse(" ".join(words[:max_tokens]), "fake-1", sum(len(m["content"].split()) for m in messages),
                           min(len(words), max_tokens), round(time.perf_counter() - start, 1), truncated)

llm: LLMClient = FakeLLM({"rag": "RAG retrieves relevant documents and passes them to the LLM as context.",
                          "ship": "We ship across India in 3-5 days."})

r = llm.complete([{"role": "user", "content": "What is RAG?"}])
print(r.text)
print(r.input_tokens, r.output_tokens, r.truncated)
print(llm.complete([{"role": "user", "content": "What is RAG?"}], max_tokens=4))
Output
RAG retrieves relevant documents and passes them to the LLM as context.
3 12 False
LLMResponse(text='RAG retrieves relevant documents', model='fake-1', input_tokens=3, output_tokens=4, latency_s=0.0, truncated=True)

The real implementation for one provider:

# no-run
from openai import OpenAI

class OpenAIClient:
    def __init__(self, model: str = "gpt-4o-mini"):
        self.client, self.model = OpenAI(timeout=30, max_retries=2), model

    def complete(self, messages, *, max_tokens=512, temperature=0.2) -> LLMResponse:
        start = time.perf_counter()
        r = self.client.chat.completions.create(model=self.model, messages=messages,
                                                max_tokens=max_tokens, temperature=temperature)
        choice = r.choices[0]
        return LLMResponse(choice.message.content or "", r.model, r.usage.prompt_tokens,
                           r.usage.completion_tokens, time.perf_counter() - start,
                           truncated=choice.finish_reason == "length")

Libraries like LiteLLM provide this abstraction across 100+ providers if you'd rather not maintain it.

3.5 Streaming

Streaming sends tokens as they are generated. Total time is the same, but the user sees the answer start in under a second instead of waiting for all of it:

def fake_stream(text: str):
    for word in text.split(" "):
        time.sleep(0.01)
        yield word + " "

start = time.perf_counter()
first_token_at = None
answer = ""
for chunk in fake_stream("Streaming makes slow answers feel fast to the user."):
    if first_token_at is None:
        first_token_at = time.perf_counter() - start
    answer += chunk
total = time.perf_counter() - start
print(answer.strip())
print(f"first token after {first_token_at:.2f}s, full answer after {total:.2f}s")   # timings vary per run
Output
Streaming makes slow answers feel fast to the user.
first token after 0.01s, full answer after 0.09s

With the OpenAI SDK, pass stream=True and read chunk.choices[0].delta.content; usage arrives in the last chunk if you set stream_options={"include_usage": True}. Serving a stream to a browser is covered in FastAPI async & streaming.

3.6 Parallel calls

Independent LLM calls — summarising 20 documents, classifying a batch, running 3 judges — should run concurrently, with a limit so you don't hit rate limits:

import asyncio

async def fake_call(doc: str) -> str:
    await asyncio.sleep(0.2)                           # stands in for an async SDK call (~seconds in reality)
    return f"summary of {doc}"

async def summarise_all(docs: list[str], max_concurrent: int = 5) -> list[str]:
    sem = asyncio.Semaphore(max_concurrent)
    async def one(doc):
        async with sem:                                # at most 5 calls in flight
            return await fake_call(doc)
    return await asyncio.gather(*(one(d) for d in docs))

docs = [f"doc{i}" for i in range(20)]
start = time.perf_counter()
summaries = asyncio.run(summarise_all(docs))
print(len(summaries), summaries[0], f"in {time.perf_counter() - start:.1f}s (sequential would be 4.0s)")   # timings vary slightly
Output
20 summary of doc0 in 0.8s (sequential would be 4.0s)

Use the async clients (AsyncOpenAI, AsyncAnthropic). For thousands of non-urgent requests (nightly evals, bulk enrichment), the providers' batch APIs are usually around half price.

Interview questions

LLM APIs are stateless — how does a chatbot remember the conversation?

The application stores the message history and sends it (or a trimmed/summarised version) with every request. The model only "remembers" what is in the current request's context.

What would you put in a wrapper around the LLM SDK?

One interface for all providers; timeouts and retries; logging of model, tokens, latency and cost per call; truncation detection; a fallback provider; and the ability to swap in a fake for tests.

How do you make many LLM calls fast without hitting rate limits?

Use async clients with asyncio.gather and a semaphore to cap concurrency; back off on 429s; use batch APIs for offline jobs; and cache repeated requests.

Practice

  • Add a cost_usd property to LLMResponse using a price table.
  • Write an AnthropicClient with the same complete() signature as OpenAIClient.

Next: Prompt engineering — writing prompts that work reliably.