Handle LLM API errors¶
Intermediate · 15 min · tested with the openai SDK — no API key needed
Your demo works on your laptop. Then real users arrive and you meet 429 Too Many Requests, timeouts, "out of credits", prompts that are too long and replies that aren't valid JSON. This tutorial makes your app handle each one on purpose.
Works with OpenAI, Groq, and other OpenAI-compatible APIs
The examples use the official openai package (pip install openai). Groq, Together, Ollama and
many others expose the same API, so the same error classes apply.
1. Know which errors to retry¶
The golden rule: retry temporary problems, fix permanent ones. Retrying a wrong API key just fails four times instead of once.
Error class (openai.…) |
HTTP | What it usually means | Retry? |
|---|---|---|---|
RateLimitError |
429 | Too many requests or tokens per minute | ✅ after waiting |
RateLimitError with code insufficient_quota |
429 | Out of credits / billing not set up | ❌ add billing |
APITimeoutError |
— | No reply within your timeout | ✅ |
APIConnectionError |
— | Network problem, DNS, VPN, proxy | ✅ |
InternalServerError |
5xx | The provider is having problems | ✅ |
AuthenticationError |
401 | Wrong or missing API key | ❌ fix the key |
PermissionDeniedError |
403 | Your account/region can't use this | ❌ |
NotFoundError |
404 | Wrong model name (or retired model) | ❌ fix the name |
BadRequestError |
400 | Invalid request — often the prompt is too long | ❌ change the request |
All of them inherit from openai.APIError, so except openai.APIError catches any API failure.
2. Use the SDK's built-in timeout and retries¶
The openai client already retries temporary errors (429, 5xx, timeouts, connection errors) with
increasing waits. Set the numbers yourself so you know what happens:
from openai import OpenAI
client = OpenAI(
timeout=30, # give up on one request after 30 seconds
max_retries=3, # then retry up to 3 times, waiting longer each time
)
# Need different settings for one call? (e.g. a long summary)
slow_client = client.with_options(timeout=120, max_retries=1)
For most apps, this is enough. Write your own retry loop (next section) only when you need custom behaviour — logging each attempt, a hard deadline, or not retrying quota errors.
3. Your own retry loop: exponential backoff + jitter¶
Exponential backoff waits 1 s, 2 s, 4 s… between attempts, giving the API time to recover. Jitter adds a little randomness so a thousand clients don't all retry at the same moment.
import random
import time
import openai
# Temporary problems worth retrying (APITimeoutError is a kind of APIConnectionError).
RETRYABLE = (openai.RateLimitError, openai.APIConnectionError, openai.InternalServerError)
def with_retries(call, max_attempts=4, base_delay=1.0):
"""Run call(); on a temporary error, wait 1s, 2s, 4s… (+ jitter) and try again."""
for attempt in range(1, max_attempts + 1):
try:
return call()
except RETRYABLE as err:
out_of_credits = isinstance(err, openai.RateLimitError) and err.code == "insufficient_quota"
if out_of_credits or attempt == max_attempts:
raise # retrying won't help — let the caller handle it
delay = base_delay * 2 ** (attempt - 1) + random.uniform(0, base_delay / 2)
print(f"{type(err).__name__} → retry {attempt} of {max_attempts - 1}")
time.sleep(delay)
Use it by wrapping the call in a lambda, and turn off the SDK's own retries so they don't stack:
client = OpenAI(timeout=30, max_retries=0) # our loop does the retrying
reply = with_retries(lambda: client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Say hello"}],
))
Test it without an API key¶
You can create the SDK's real error objects yourself — perfect for testing retry logic for free:
import httpx
request = httpx.Request("POST", "https://api.openai.com/v1/chat/completions")
def make_rate_limit_error(code=None):
"""A real openai.RateLimitError, as the SDK would raise it."""
response = httpx.Response(429, request=request)
return openai.RateLimitError("Rate limit reached", response=response, body={"code": code})
# A fake API call: fails twice with 429, then succeeds.
outcomes = iter([make_rate_limit_error(), make_rate_limit_error(), "Hello!"])
def flaky_call():
result = next(outcomes)
if isinstance(result, Exception):
raise result
return result
print(with_retries(flaky_call, base_delay=0.01))
# → RateLimitError → retry 1 of 3
# → RateLimitError → retry 2 of 3
# → Hello!
And an out-of-credits error is not retried:
def no_credits():
raise make_rate_limit_error("insufficient_quota")
try:
with_retries(no_credits)
except openai.RateLimitError as err:
print("gave up immediately:", err.code) # → gave up immediately: insufficient_quota
4. Turn errors into messages people can act on¶
Users shouldn't see a stack trace — and you want to know exactly what to fix.
def explain(err):
"""Turn an openai SDK error into a clear, actionable message."""
# Order matters: APITimeoutError is a subclass of APIConnectionError, so check it first.
if isinstance(err, openai.APITimeoutError):
return "The AI took too long to answer — please try again."
if isinstance(err, openai.APIConnectionError):
return "Couldn't reach the AI service — check your internet connection."
if isinstance(err, openai.AuthenticationError):
return "Invalid API key — check OPENAI_API_KEY in your .env file."
if isinstance(err, openai.NotFoundError):
return "Model not found — check the model name in your settings."
if isinstance(err, openai.RateLimitError):
if err.code == "insufficient_quota":
return "Out of API credits — add billing on your provider's dashboard."
return "Too many requests right now — please wait a few seconds."
if isinstance(err, openai.BadRequestError):
if err.code == "context_length_exceeded":
return "Your question plus documents is too long — try a shorter question."
return f"The request was rejected: {err.message}"
if isinstance(err, openai.InternalServerError):
return "The AI service is having problems — please try again shortly."
return "Something went wrong with the AI service."
print(explain(make_rate_limit_error("insufficient_quota")))
# → Out of API credits — add billing on your provider's dashboard.
print(explain(openai.APITimeoutError(request=request)))
# → The AI took too long to answer — please try again.
Use it at the edge of your app — the place that talks to the user:
import logging
def safe_answer(question):
try:
return with_retries(lambda: ask_llm(question))
except openai.APIError as err:
logging.exception("LLM call failed") # full details in your logs…
return explain(err) # …a friendly message for the user
5. The model returned invalid JSON¶
Even when asked for JSON, models sometimes wrap it in ```json fences or add "Sure! Here it is:".
First, ask the API for JSON mode — it guarantees syntactically valid JSON (mention "JSON" in your prompt):
reply = client.chat.completions.create(
model="gpt-4o-mini",
response_format={"type": "json_object"},
messages=[{"role": "user", "content": 'Classify this ticket. Reply in JSON: {"priority": 1-5}'}],
)
Then parse defensively, and if it still fails, tell the model what went wrong and ask again:
import json
import re
def parse_json_reply(text):
"""Pull the JSON object out of a reply, even with ```json fences or chatter around it."""
match = re.search(r"\{.*\}", text, re.DOTALL)
if not match:
raise ValueError(f"no JSON object in reply: {text[:40]!r}")
return json.loads(match.group()) # JSONDecodeError is a kind of ValueError
def ask_json(llm, prompt, attempts=3):
"""Call llm(messages) until it returns parseable JSON; feed each error back to the model."""
messages = [{"role": "user", "content": prompt}]
for _ in range(attempts):
reply = llm(messages)
try:
return parse_json_reply(reply)
except ValueError as err:
messages += [
{"role": "assistant", "content": reply},
{"role": "user", "content": f"That was not valid JSON ({err}). Reply with only the JSON object."},
]
raise ValueError(f"no valid JSON after {attempts} attempts")
print(parse_json_reply('Sure! ```json\n{"priority": 2}\n```')) # → {'priority': 2}
# A fake model that gets it wrong once, then right:
replies = iter(["The priority is 2.", '{"priority": 2}'])
print(ask_json(lambda messages: next(replies), "Classify this ticket")) # → {'priority': 2}
To also check the values (is priority really a number from 1–5?), validate with Pydantic — see
Avoid common coding mistakes #12.
6. "Context length exceeded" — the prompt is too long¶
Every model has a maximum number of tokens it can read (prompt + answer). A long chat history or too
many retrieved chunks hits it with a BadRequestError whose code is context_length_exceeded.
Fix the request, don't retry it. Keep the system message and only the most recent turns:
def trim_history(messages, max_chars=12_000):
"""Keep the system message plus the newest messages that fit (≈ 4 characters per token)."""
system, turns = messages[0], messages[1:]
kept, total = [], len(system["content"])
for message in reversed(turns): # newest first
total += len(message["content"])
if total > max_chars:
break
kept.append(message)
return [system, *reversed(kept)]
history = [{"role": "system", "content": "Be brief."}] + [
{"role": "user", "content": f"message {i} " + "x" * 3_000} for i in range(10)
]
trimmed = trim_history(history)
print(len(history), "→", len(trimmed)) # → 11 → 4
print(trimmed[1]["content"][:9]) # → message 7
Other fixes: send fewer RAG chunks (k=3 instead of k=8), summarise old turns, or count tokens
exactly with the tiktoken package.
7. Fall back to another provider¶
If your main provider is down, a backup keeps the app working. Because Groq (and others) speak the OpenAI API, it's just a second client:
def ask_with_fallback(calls):
"""Try each (name, call) in order; return the first answer that works."""
failures = []
for name, call in calls:
try:
return call()
except openai.APIError as err:
failures.append(f"{name}: {type(err).__name__}")
raise RuntimeError("all providers failed — " + "; ".join(failures))
import os
openai_client = OpenAI(timeout=20, max_retries=2)
groq_client = OpenAI(api_key=os.getenv("GROQ_API_KEY"), base_url="https://api.groq.com/openai/v1",
timeout=20, max_retries=2)
messages = [{"role": "user", "content": "Say hello"}]
reply = ask_with_fallback([
("openai", lambda: openai_client.chat.completions.create(model="gpt-4o-mini", messages=messages)),
("groq", lambda: groq_client.chat.completions.create(model="llama-3.1-8b-instant", messages=messages)),
])
Testing it with fakes:
def down():
raise openai.InternalServerError("Service unavailable",
response=httpx.Response(503, request=request), body=None)
print(ask_with_fallback([("openai", down), ("groq", lambda: "Hi from the backup!")]))
# → Hi from the backup!
Fallbacks change behaviour
A different model may answer differently. Log which provider answered, and test your prompts on both. Model names change over time — check each provider's model list.
Production checklist¶
- Every client has an explicit
timeoutandmax_retries. - Only temporary errors (429, 5xx, timeouts, connection) are retried — with backoff and jitter.
- Out-of-credit, auth and bad-request errors are not retried, and are logged loudly.
- Users see a friendly message, logs get the full error (
logging.exception). - JSON replies are parsed defensively and validated; failures are retried with the error message.
- Long chats and big contexts are trimmed before sending.
- (Optional) A backup provider for when the main one is down.