10. Performance & caching¶
Advanced · 8 min read
The golden rule: measure first. In LLM apps the slow part is usually the network call — so caching often beats any code tweak.
10.1 Measuring¶
import timeit
items = list(range(10_000))
as_set = set(items)
# Time 1,000 membership checks in a list vs a set.
list_time = timeit.timeit(lambda: 9_999 in items, number=1_000)
set_time = timeit.timeit(lambda: 9_999 in as_set, number=1_000)
print(set_time < list_time) # → True a set lookup doesn't scan every item
To find where a whole program spends its time, run it under the profiler:
10.2 Choosing the right structure¶
| Need | Slow | Fast |
|---|---|---|
| "Is X in this collection?" | x in big_list |
x in big_set |
| Look up by key | search a list of dicts | dict[key] |
| Build a long string | s += piece in a loop |
"".join(pieces) |
| Count things | manual dict updates | collections.Counter |
10.3 In-memory caching with lru_cache¶
import functools
calls = 0
@functools.lru_cache(maxsize=1024) # keep up to 1024 recent results
def embed(text: str) -> tuple[float, ...]:
"""Pretend embedding call — the expensive part we don't want to repeat."""
global calls
calls += 1
return tuple(float(ord(c)) for c in text[:3])
for q in ["refund", "shipping", "refund", "refund"]:
embed(q)
print(calls) # → 2 two unique texts, two real calls
print(embed.cache_info().hits) # → 2
Cached arguments must be hashable (strings, numbers, tuples) — not lists or dicts. Return tuples rather than lists so callers can't modify the cached value.
10.4 Caching across runs (on disk)¶
import hashlib
import json
from pathlib import Path
CACHE = Path(".llm_cache")
CACHE.mkdir(exist_ok=True)
def cached_completion(prompt: str, model: str = "gpt-4o-mini") -> str:
"""Return a saved answer if this exact (model, prompt) was asked before."""
key = hashlib.sha256(f"{model}\n{prompt}".encode()).hexdigest() # stable file name
path = CACHE / f"{key}.json"
if path.exists():
return json.loads(path.read_text(encoding="utf-8"))["answer"]
answer = f"(model answer to: {prompt})" # ← the real API call goes here
path.write_text(json.dumps({"answer": answer}), encoding="utf-8")
return answer
print(cached_completion("What is RAG?") == cached_completion("What is RAG?")) # → True
During development this makes re-running a pipeline free and instant. In production, a shared cache (Redis, a database) does the same across servers.
10.5 Batching¶
APIs often accept many inputs per request — embedding 100 texts in one call instead of 100 calls
cuts latency and overhead dramatically. Combine with the batched() generator from
Iterators & generators.
Why it matters for GenAI
Cache embeddings and repeated answers, batch requests, run independent calls concurrently (async) — these three cut LLM-app latency and cost far more than micro-optimising Python.
Practice¶
- Profile a script that chunks a large text file. Is most of the time in reading, splitting or joining?