async / await in plain English¶
Intermediate · 15 min · every example runs
async code confuses almost everyone at first. The idea behind it is simple, though — and in GenAI
apps it can make your code 5–10× faster, because most of the time is spent waiting for an API.
The idea: don't stand still while you wait¶
Imagine making tea and toast. A synchronous cook does one thing at a time:
boil water (3 min, staring at the kettle) → then make toast (2 min) → total 5 min
An asynchronous cook starts the kettle, and while it's boiling puts the bread in the toaster:
start kettle → start toaster → wait for both → total 3 min
Same cook, no extra hands — just no standing around. That's async: one thread that switches to other work whenever it's waiting for something (a network reply, a file, a database).
Calling an LLM is exactly this situation: your code sends a request and then waits 1–10 seconds doing nothing. With async, it can send the other requests in the meantime.
Async is not \"more CPUs\"
Async helps when your program is waiting (API calls, downloads, databases). It does not speed up heavy calculations — for that you need multiple processes. See Concurrency.
The four words you need¶
| Word | Plain English |
|---|---|
async def |
"This function may pause while it waits." Calling it gives you a task to run, not the result. |
await |
"Pause here until this is done — and let other work run meanwhile." Only allowed inside async def. |
asyncio.run(main()) |
"Start the async world and run main() until it finishes." Usually called once, at the bottom of the file. |
asyncio.gather(a, b, c) |
"Run these at the same time and give me all the results, in order." |
See the speed-up¶
asyncio.sleep() stands in for an LLM call that takes 0.3 seconds.
import time
def fake_llm(question):
time.sleep(0.3) # waiting for the API…
return f"answer to {question}"
start = time.perf_counter()
answers = [fake_llm(q) for q in ["q1", "q2", "q3"]]
print(answers) # → ['answer to q1', 'answer to q2', 'answer to q3']
print(f"took {time.perf_counter() - start:.1f}s") # → took 0.9s
import asyncio
import time
async def fake_llm(question):
await asyncio.sleep(0.3) # waiting — but other tasks run meanwhile
return f"answer to {question}"
async def main():
return await asyncio.gather(fake_llm("q1"), fake_llm("q2"), fake_llm("q3"))
start = time.perf_counter()
answers = asyncio.run(main())
print(answers) # → ['answer to q1', 'answer to q2', 'answer to q3']
print(f"took {time.perf_counter() - start:.1f}s") # → took 0.3s
Three calls in the time of one. With 30 calls it would be 9 seconds vs about 0.3.
The 5 classic mistakes¶
1. Forgetting await¶
Calling an async function doesn't run it — it just creates a coroutine (a paused task).
import asyncio
async def get_answer():
return 42
async def main():
result = get_answer() # no await!
print(type(result).__name__) # → coroutine
result.close() # (tidy up so Python doesn't warn)
asyncio.run(main())
If you see <coroutine object …> printed or RuntimeWarning: coroutine '…' was never awaited,
you forgot an await.
2. await one by one — still slow¶
await waits for each call to finish before starting the next. Awaiting in a loop is
no faster than normal code.
The * "unpacks" the list, so gather(*tasks) is the same as gather(task1, task2, task3).
3. Blocking calls inside async def¶
time.sleep(), requests.get() and the normal OpenAI() client block the whole program —
nothing else can run while they wait, so gather gains nothing.
import asyncio
import time
async def fake_llm(q):
time.sleep(0.3) # blocking! the other tasks can't run
return q
async def main():
await asyncio.gather(fake_llm("q1"), fake_llm("q2"), fake_llm("q3"))
start = time.perf_counter()
asyncio.run(main())
print(f"took {time.perf_counter() - start:.1f}s") # → took 0.9s
| Instead of (blocking) | Use (async) |
|---|---|
time.sleep() |
await asyncio.sleep() |
requests.get() |
await httpx.AsyncClient().get() |
OpenAI() |
AsyncOpenAI() |
| a blocking library with no async version | await asyncio.to_thread(func, args) |
4. asyncio.run() inside Jupyter¶
Notebooks already run an event loop, so asyncio.run() fails with
RuntimeError: asyncio.run() cannot be called from a running event loop.
In a notebook, just write await main() directly in a cell.
5. Using async where it doesn't help¶
If your script makes one API call and waits for the answer, async adds complexity for no speed-up. Use it when you have many waits that can overlap.
Real use: many LLM calls, safely¶
Firing 500 requests at once will hit rate limits. A semaphore is a counter that lets only N tasks in at a time:
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI(timeout=30, max_retries=3)
limit = asyncio.Semaphore(5) # at most 5 requests in flight
async def summarise(text):
async with limit: # waits here if 5 are already running
reply = await client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": f"Summarise in one line: {text}"}],
)
return reply.choices[0].message.content
async def main(texts):
# return_exceptions=True: one failure doesn't throw away all the other results
return await asyncio.gather(*(summarise(t) for t in texts), return_exceptions=True)
# results = asyncio.run(main(["doc one…", "doc two…", "doc three…"]))
Here is the same pattern with a fake API, so you can watch the limit work:
import asyncio
import time
limit = asyncio.Semaphore(2) # only 2 at a time
async def fake_call(i):
async with limit:
await asyncio.sleep(0.2)
if i == 3:
raise ValueError("bad input") # one task fails
return i * 10
async def main():
return await asyncio.gather(*(fake_call(i) for i in range(4)), return_exceptions=True)
start = time.perf_counter()
results = asyncio.run(main())
print(results) # → [0, 10, 20, ValueError('bad input')]
print(f"took {time.perf_counter() - start:.1f}s") # → took 0.4s
4 tasks, 2 at a time → two "rounds" of 0.2 s. And the failure comes back as a value instead of
crashing everything — check with isinstance(r, Exception).
Add a time limit¶
import asyncio
async def slow_llm():
await asyncio.sleep(5)
return "finally!"
async def main():
try:
async with asyncio.timeout(0.5): # Python 3.11+ (older: asyncio.wait_for)
return await slow_llm()
except TimeoutError:
return "timed out — showing a cached answer instead"
print(asyncio.run(main())) # → timed out — showing a cached answer instead
Where you'll meet async in GenAI¶
| Place | Why it's async |
|---|---|
FastAPI endpoints (async def chat(...)) |
One server handles many users while each waits on the LLM |
Streaming replies (async for chunk in stream) |
Show tokens as they arrive |
| Batch jobs — embedding or summarising 1,000 docs | gather + a semaphore: minutes instead of hours |
| Agents calling several tools | Run independent tool calls at the same time |
Cheat sheet¶
import asyncio
async def work(x): # 1. define with async def
await asyncio.sleep(0.1) # 2. await anything that waits
return x * 2
async def main():
one = await work(1) # a single call
many = await asyncio.gather(*(work(i) for i in range(3))) # many at once
return one, many
print(asyncio.run(main())) # 3. start it once → (2, [0, 2, 4])
Deeper dive: Async programming in the Python Advanced notes.
Previous: Handle LLM API errors ←