2. Request & response models¶
Beginner · 9 min read
A function argument typed as a Pydantic model becomes the JSON request body. FastAPI parses it, validates it, and documents it — your code only ever sees clean data.
2.1 A request body¶
from typing import Literal
from fastapi import FastAPI
from fastapi.testclient import TestClient
from pydantic import BaseModel, Field
app = FastAPI()
client = TestClient(app)
class Message(BaseModel):
role: Literal["system", "user", "assistant"]
content: str = Field(min_length=1)
class ChatRequest(BaseModel):
messages: list[Message] = Field(min_length=1)
model: str = "gpt-4o-mini"
temperature: float = Field(0.2, ge=0, le=2)
@app.post("/chat")
def chat(req: ChatRequest):
last = req.messages[-1].content
return {"model": req.model, "reply": f"You said: {last}"}
r = client.post("/chat", json={"messages": [{"role": "user", "content": "Hello"}]})
print(r.status_code, r.json())
Bad input never reaches your function:
r = client.post("/chat", json={"messages": [{"role": "bot", "content": ""}], "temperature": 5})
print(r.status_code)
for err in r.json()["detail"]:
print(err["loc"], "→", err["msg"])
422
['body', 'messages', 0, 'role'] → Input should be 'system', 'user' or 'assistant'
['body', 'messages', 0, 'content'] → String should have at least 1 character
['body', 'temperature'] → Input should be less than or equal to 2
2.2 Response models¶
response_model controls what goes out: it validates your return value and drops any field that
isn't in the model — an easy way to never leak internal data.
class Usage(BaseModel):
prompt_tokens: int
completion_tokens: int
class ChatResponse(BaseModel):
reply: str
model: str
usage: Usage
@app.post("/v2/chat", response_model=ChatResponse)
def chat_v2(req: ChatRequest):
return {
"reply": "RAG combines retrieval with generation.",
"model": req.model,
"usage": {"prompt_tokens": 42, "completion_tokens": 9},
"internal_cost_usd": 0.00031, # not in ChatResponse → removed
"system_prompt": "You are a helpful...", # not in ChatResponse → removed
}
print(client.post("/v2/chat", json={"messages": [{"role": "user", "content": "What is RAG?"}]}).json())
{'reply': 'RAG combines retrieval with generation.', 'model': 'gpt-4o-mini', 'usage': {'prompt_tokens': 42, 'completion_tokens': 9}}
You can also write the return type instead: def chat_v2(req: ChatRequest) -> ChatResponse: — FastAPI
treats it the same way.
2.3 Status codes¶
from fastapi import status
class IngestRequest(BaseModel):
doc_id: str
text: str
DOCS: dict[str, str] = {}
@app.post("/documents", status_code=status.HTTP_201_CREATED)
def ingest(req: IngestRequest):
DOCS[req.doc_id] = req.text
return {"doc_id": req.doc_id, "chars": len(req.text)}
@app.delete("/documents/{doc_id}", status_code=status.HTTP_204_NO_CONTENT)
def delete(doc_id: str):
DOCS.pop(doc_id, None)
r = client.post("/documents", json={"doc_id": "policy", "text": "Refunds take 5 days."})
print(r.status_code, r.json())
print(client.delete("/documents/policy").status_code)
| Code | Meaning | Typical GenAI case |
|---|---|---|
| 200 | OK | answer returned |
| 201 | Created | document ingested |
| 400 | Bad request | prompt too long for the model |
| 401 / 403 | Not logged in / not allowed | missing or wrong API key |
| 404 | Not found | unknown document or session |
| 422 | Validation failed | wrong JSON shape (automatic) |
| 429 | Too many requests | rate limit hit |
| 502 / 503 / 504 | Upstream failed / busy / timed out | LLM provider error |
2.4 Errors — HTTPException¶
Raise it anywhere to stop and send an error response:
from fastapi import HTTPException
MAX_PROMPT_CHARS = 20
@app.get("/documents/{doc_id}")
def read_doc(doc_id: str):
if doc_id not in DOCS:
raise HTTPException(status_code=404, detail=f"Document {doc_id!r} not found")
return {"doc_id": doc_id, "text": DOCS[doc_id]}
@app.post("/v3/chat")
def chat_v3(req: ChatRequest):
if sum(len(m.content) for m in req.messages) > MAX_PROMPT_CHARS:
raise HTTPException(status_code=400, detail="Prompt too long — trim the conversation history")
return {"reply": "ok"}
r = client.get("/documents/missing")
print(r.status_code, r.json())
r = client.post("/v3/chat", json={"messages": [{"role": "user", "content": "x" * 50}]})
print(r.status_code, r.json())
404 {'detail': "Document 'missing' not found"}
400 {'detail': 'Prompt too long — trim the conversation history'}
2.5 Turning your own exceptions into responses¶
Keep business code free of HTTP details: raise your own exception, map it to a response once. Register handlers (and middleware) when you create the app — before it serves any request.
from fastapi import Request
from fastapi.responses import JSONResponse
app = FastAPI() # a fresh app for this example
client = TestClient(app)
class LLMProviderError(Exception):
def __init__(self, provider: str):
self.provider = provider
@app.exception_handler(LLMProviderError)
async def llm_error_handler(request: Request, exc: LLMProviderError):
return JSONResponse(status_code=502, content={"error": f"{exc.provider} is unavailable, try again"})
@app.post("/v4/chat")
def chat_v4(req: ChatRequest):
raise LLMProviderError("openai") # pretend the upstream call failed
r = client.post("/v4/chat", json={"messages": [{"role": "user", "content": "hi"}]})
print(r.status_code, r.json())
2.6 Form data and file uploads¶
RAG apps need uploads. Install python-multipart (included in fastapi[standard]):
from fastapi import UploadFile
@app.post("/upload")
async def upload(file: UploadFile):
data = await file.read()
return {"filename": file.filename, "content_type": file.content_type, "bytes": len(data)}
r = client.post("/upload", files={"file": ("notes.txt", b"RAG = retrieval + generation", "text/plain")})
print(r.json())
Limit what you accept
Check file.content_type and the size before parsing. A 500 MB PDF or an unexpected file type
should get a 400 / 413, not a crashed worker.
Practice¶
- Add
POST /embedtaking{"texts": [...]}(1–100 strings) and returning fake vectors of length 4. - Make
GET /documents/{doc_id}return aDocumentOutresponse model without the raw text.
Next: Dependencies & middleware — auth, settings and shared clients.