8. Testing with pytest¶
Intermediate · 9 min read
Tests prove your code works — and keep proving it after every change. pytest
(pip install pytest) is the standard tool.
8.1 Your first test¶
text_utils.py
def chunk_words(text: str, size: int) -> list[str]:
"""Split text into chunks of `size` words."""
if size <= 0:
raise ValueError("size must be positive")
words = text.split()
return [" ".join(words[i:i + size]) for i in range(0, len(words), size)]
test_text_utils.py
from text_utils import chunk_words
def test_splits_into_chunks(): # any function named test_* is a test
assert chunk_words("a b c d e", 2) == ["a b", "c d", "e"]
def test_empty_text_gives_no_chunks():
assert chunk_words("", 3) == []
Run pytest in the folder — it finds files named test_*.py and reports each pass or failure,
showing exactly which values differed.
8.2 Testing errors¶
test_text_utils.py
import pytest
from text_utils import chunk_words
def test_rejects_zero_size():
with pytest.raises(ValueError, match="positive"): # the block must raise this error
chunk_words("a b", 0)
8.3 Many cases with parametrize¶
test_text_utils.py
import pytest
from text_utils import chunk_words
@pytest.mark.parametrize("text, size, expected", [
("a b c", 1, 3), # each tuple becomes its own test
("a b c", 2, 2),
("a b c", 10, 1),
])
def test_chunk_count(text, size, expected):
assert len(chunk_words(text, size)) == expected
8.4 Fixtures: shared setup¶
test_retrieval.py
import pytest
@pytest.fixture
def docs(): # runs before each test that asks for it
return ["refunds within 30 days", "free shipping over 500"]
def test_finds_refund_doc(docs): # ask for a fixture by naming it as a parameter
assert any("refund" in d for d in docs)
def test_temp_files(tmp_path): # built-in fixture: a fresh temporary folder
f = tmp_path / "note.txt"
f.write_text("hello", encoding="utf-8")
assert f.read_text(encoding="utf-8") == "hello"
8.5 Faking the LLM with monkeypatch¶
Tests must not call real APIs — that's slow, costs money and gives different answers each time. Replace the call with a fake:
app.py
def call_llm(prompt: str) -> str:
raise RuntimeError("would call the real API") # the real network call lives here
def summarise(text: str) -> str:
return call_llm(f"Summarise: {text}").strip()
test_app.py
import app
def test_summarise_uses_llm_and_strips(monkeypatch):
seen = []
def fake_llm(prompt):
seen.append(prompt) # record what we were asked
return " short summary " # canned reply
monkeypatch.setattr(app, "call_llm", fake_llm) # swap it in for this test only
assert app.summarise("long text") == "short summary"
assert seen == ["Summarise: long text"]
Why it matters for GenAI
Test the deterministic parts — chunking, prompt building, parsing model output, retries — with fakes. Evaluate the model's quality separately, with a fixed question set.
Practice¶
- Write a parametrized test for a
clean(text)function that collapses whitespace, with three input/output pairs.