6. Async programming (asyncio)¶
Advanced · 10 min read
LLM and API calls spend most of their time waiting on the network. async code lets one
program wait on many requests at once instead of one after another.
6.1 async and await¶
import asyncio
async def call_model(prompt: str, seconds: float) -> str: # "async def" = a coroutine
await asyncio.sleep(seconds) # `await` = "pause me here; run something else meanwhile"
return f"reply to {prompt!r}"
async def main():
reply = await call_model("hi", 0.1)
print(reply) # → reply to 'hi'
asyncio.run(main()) # start the event loop and run main()
In Jupyter, the loop is already running — use await main() directly instead of asyncio.run.
6.2 Running calls concurrently with gather¶
import asyncio
import time
async def call_model(prompt: str) -> str:
await asyncio.sleep(0.2) # pretend each call takes 200 ms
return prompt.upper()
async def main():
prompts = ["summarise", "translate", "classify"]
start = time.perf_counter()
# Start all three at once and wait for every result (returned in the same order).
results = await asyncio.gather(*(call_model(p) for p in prompts))
elapsed = time.perf_counter() - start
print(results) # → ['SUMMARISE', 'TRANSLATE', 'CLASSIFY']
print(elapsed < 0.4) # → True ~0.2 s total, not 0.6 s
asyncio.run(main())
6.3 Limiting concurrency with a semaphore¶
Firing 500 requests at once will hit rate limits. A semaphore caps how many run together.
import asyncio
async def embed(text: str, limit: asyncio.Semaphore) -> int:
async with limit: # at most N tasks inside this block at a time
await asyncio.sleep(0.05)
return len(text)
async def main():
limit = asyncio.Semaphore(5) # 5 requests in flight, max
texts = [f"doc {i}" for i in range(20)]
sizes = await asyncio.gather(*(embed(t, limit) for t in texts))
print(len(sizes), sum(sizes)) # → 20 110
asyncio.run(main())
6.4 Timeouts and errors¶
import asyncio
async def slow_call():
await asyncio.sleep(5)
return "late"
async def main():
try:
await asyncio.wait_for(slow_call(), timeout=0.1) # give up after 100 ms
except asyncio.TimeoutError:
print("timed out") # → timed out
# return_exceptions=True: one failure doesn't cancel the rest
async def ok():
return "ok"
async def boom():
raise ValueError("bad input")
results = await asyncio.gather(ok(), boom(), return_exceptions=True)
print([type(r).__name__ for r in results]) # → ['str', 'ValueError']
asyncio.run(main())
6.5 Real SDKs¶
Most LLM SDKs ship an async client — e.g. from openai import AsyncOpenAI, then
await client.chat.completions.create(...). Web frameworks like FastAPI run async def
endpoints natively, so one server can handle many slow LLM calls at once.
Don't block the loop
Inside async def, use async libraries (httpx.AsyncClient, asyncio.sleep). A normal
time.sleep() or requests.get() freezes every other task until it finishes.
Why it matters for GenAI
Embedding 10,000 chunks, evaluating a model on 200 questions, or calling three tools in
parallel goes from minutes to seconds with gather + a semaphore.
Practice¶
- Use
asyncio.gatherto run five fake calls that each sleep 0.1 s, and confirm the total time is about 0.1 s.