Skip to content

async / await in plain English

Intermediate · 15 min · every example runs

async code confuses almost everyone at first. The idea behind it is simple, though — and in GenAI apps it can make your code 5–10× faster, because most of the time is spent waiting for an API.


The idea: don't stand still while you wait

Imagine making tea and toast. A synchronous cook does one thing at a time:

boil water (3 min, staring at the kettle) → then make toast (2 min) → total 5 min

An asynchronous cook starts the kettle, and while it's boiling puts the bread in the toaster:

start kettle → start toaster → wait for both → total 3 min

Same cook, no extra hands — just no standing around. That's async: one thread that switches to other work whenever it's waiting for something (a network reply, a file, a database).

Calling an LLM is exactly this situation: your code sends a request and then waits 1–10 seconds doing nothing. With async, it can send the other requests in the meantime.

Async is not \"more CPUs\"

Async helps when your program is waiting (API calls, downloads, databases). It does not speed up heavy calculations — for that you need multiple processes. See Concurrency.


The four words you need

Word Plain English
async def "This function may pause while it waits." Calling it gives you a task to run, not the result.
await "Pause here until this is done — and let other work run meanwhile." Only allowed inside async def.
asyncio.run(main()) "Start the async world and run main() until it finishes." Usually called once, at the bottom of the file.
asyncio.gather(a, b, c) "Run these at the same time and give me all the results, in order."

See the speed-up

asyncio.sleep() stands in for an LLM call that takes 0.3 seconds.

import time

def fake_llm(question):
    time.sleep(0.3)                       # waiting for the API…
    return f"answer to {question}"

start = time.perf_counter()
answers = [fake_llm(q) for q in ["q1", "q2", "q3"]]
print(answers)                            # → ['answer to q1', 'answer to q2', 'answer to q3']
print(f"took {time.perf_counter() - start:.1f}s")   # → took 0.9s
import asyncio
import time

async def fake_llm(question):
    await asyncio.sleep(0.3)              # waiting — but other tasks run meanwhile
    return f"answer to {question}"

async def main():
    return await asyncio.gather(fake_llm("q1"), fake_llm("q2"), fake_llm("q3"))

start = time.perf_counter()
answers = asyncio.run(main())
print(answers)                            # → ['answer to q1', 'answer to q2', 'answer to q3']
print(f"took {time.perf_counter() - start:.1f}s")   # → took 0.3s

Three calls in the time of one. With 30 calls it would be 9 seconds vs about 0.3.


The 5 classic mistakes

1. Forgetting await

Calling an async function doesn't run it — it just creates a coroutine (a paused task).

import asyncio

async def get_answer():
    return 42

async def main():
    result = get_answer()                 # no await!
    print(type(result).__name__)          # → coroutine
    result.close()                        # (tidy up so Python doesn't warn)

asyncio.run(main())

If you see <coroutine object …> printed or RuntimeWarning: coroutine '…' was never awaited, you forgot an await.

import asyncio

async def get_answer():
    return 42

async def main():
    result = await get_answer()
    print(result)                         # → 42

asyncio.run(main())

2. await one by one — still slow

await waits for each call to finish before starting the next. Awaiting in a loop is no faster than normal code.

import asyncio
import time

async def fake_llm(q):
    await asyncio.sleep(0.3)
    return q

async def main():
    return [await fake_llm(q) for q in ["q1", "q2", "q3"]]   # one after another

start = time.perf_counter()
asyncio.run(main())
print(f"took {time.perf_counter() - start:.1f}s")            # → took 0.9s
import asyncio
import time

async def fake_llm(q):
    await asyncio.sleep(0.3)
    return q

async def main():
    return await asyncio.gather(*(fake_llm(q) for q in ["q1", "q2", "q3"]))   # all at once

start = time.perf_counter()
asyncio.run(main())
print(f"took {time.perf_counter() - start:.1f}s")            # → took 0.3s

The * "unpacks" the list, so gather(*tasks) is the same as gather(task1, task2, task3).

3. Blocking calls inside async def

time.sleep(), requests.get() and the normal OpenAI() client block the whole program — nothing else can run while they wait, so gather gains nothing.

import asyncio
import time

async def fake_llm(q):
    time.sleep(0.3)                       # blocking! the other tasks can't run
    return q

async def main():
    await asyncio.gather(fake_llm("q1"), fake_llm("q2"), fake_llm("q3"))

start = time.perf_counter()
asyncio.run(main())
print(f"took {time.perf_counter() - start:.1f}s")   # → took 0.9s
Instead of (blocking) Use (async)
time.sleep() await asyncio.sleep()
requests.get() await httpx.AsyncClient().get()
OpenAI() AsyncOpenAI()
a blocking library with no async version await asyncio.to_thread(func, args)

4. asyncio.run() inside Jupyter

Notebooks already run an event loop, so asyncio.run() fails with RuntimeError: asyncio.run() cannot be called from a running event loop. In a notebook, just write await main() directly in a cell.

5. Using async where it doesn't help

If your script makes one API call and waits for the answer, async adds complexity for no speed-up. Use it when you have many waits that can overlap.


Real use: many LLM calls, safely

Firing 500 requests at once will hit rate limits. A semaphore is a counter that lets only N tasks in at a time:

import asyncio

from openai import AsyncOpenAI

client = AsyncOpenAI(timeout=30, max_retries=3)
limit = asyncio.Semaphore(5)                          # at most 5 requests in flight


async def summarise(text):
    async with limit:                                 # waits here if 5 are already running
        reply = await client.chat.completions.create(
            model="gpt-4o-mini",
            messages=[{"role": "user", "content": f"Summarise in one line: {text}"}],
        )
        return reply.choices[0].message.content


async def main(texts):
    # return_exceptions=True: one failure doesn't throw away all the other results
    return await asyncio.gather(*(summarise(t) for t in texts), return_exceptions=True)

# results = asyncio.run(main(["doc one…", "doc two…", "doc three…"]))

Here is the same pattern with a fake API, so you can watch the limit work:

import asyncio
import time

limit = asyncio.Semaphore(2)                          # only 2 at a time

async def fake_call(i):
    async with limit:
        await asyncio.sleep(0.2)
        if i == 3:
            raise ValueError("bad input")             # one task fails
        return i * 10

async def main():
    return await asyncio.gather(*(fake_call(i) for i in range(4)), return_exceptions=True)

start = time.perf_counter()
results = asyncio.run(main())
print(results)                                        # → [0, 10, 20, ValueError('bad input')]
print(f"took {time.perf_counter() - start:.1f}s")     # → took 0.4s

4 tasks, 2 at a time → two "rounds" of 0.2 s. And the failure comes back as a value instead of crashing everything — check with isinstance(r, Exception).

Add a time limit

import asyncio

async def slow_llm():
    await asyncio.sleep(5)
    return "finally!"

async def main():
    try:
        async with asyncio.timeout(0.5):              # Python 3.11+ (older: asyncio.wait_for)
            return await slow_llm()
    except TimeoutError:
        return "timed out — showing a cached answer instead"

print(asyncio.run(main()))                            # → timed out — showing a cached answer instead

Where you'll meet async in GenAI

Place Why it's async
FastAPI endpoints (async def chat(...)) One server handles many users while each waits on the LLM
Streaming replies (async for chunk in stream) Show tokens as they arrive
Batch jobs — embedding or summarising 1,000 docs gather + a semaphore: minutes instead of hours
Agents calling several tools Run independent tool calls at the same time

Cheat sheet

import asyncio

async def work(x):                  # 1. define with async def
    await asyncio.sleep(0.1)        # 2. await anything that waits
    return x * 2

async def main():
    one = await work(1)                                # a single call
    many = await asyncio.gather(*(work(i) for i in range(3)))   # many at once
    return one, many

print(asyncio.run(main()))          # 3. start it once  → (2, [0, 2, 4])

Deeper dive: Async programming in the Python Advanced notes.

Previous: Handle LLM API errors ←