Skip to content

2. Request & response models

Beginner · 9 min read

A function argument typed as a Pydantic model becomes the JSON request body. FastAPI parses it, validates it, and documents it — your code only ever sees clean data.

2.1 A request body

from typing import Literal
from fastapi import FastAPI
from fastapi.testclient import TestClient
from pydantic import BaseModel, Field

app = FastAPI()
client = TestClient(app)

class Message(BaseModel):
    role: Literal["system", "user", "assistant"]
    content: str = Field(min_length=1)

class ChatRequest(BaseModel):
    messages: list[Message] = Field(min_length=1)
    model: str = "gpt-4o-mini"
    temperature: float = Field(0.2, ge=0, le=2)

@app.post("/chat")
def chat(req: ChatRequest):
    last = req.messages[-1].content
    return {"model": req.model, "reply": f"You said: {last}"}

r = client.post("/chat", json={"messages": [{"role": "user", "content": "Hello"}]})
print(r.status_code, r.json())
Output
200 {'model': 'gpt-4o-mini', 'reply': 'You said: Hello'}

Bad input never reaches your function:

r = client.post("/chat", json={"messages": [{"role": "bot", "content": ""}], "temperature": 5})
print(r.status_code)
for err in r.json()["detail"]:
    print(err["loc"], "→", err["msg"])
Output
422
['body', 'messages', 0, 'role'] → Input should be 'system', 'user' or 'assistant'
['body', 'messages', 0, 'content'] → String should have at least 1 character
['body', 'temperature'] → Input should be less than or equal to 2

2.2 Response models

response_model controls what goes out: it validates your return value and drops any field that isn't in the model — an easy way to never leak internal data.

class Usage(BaseModel):
    prompt_tokens: int
    completion_tokens: int

class ChatResponse(BaseModel):
    reply: str
    model: str
    usage: Usage

@app.post("/v2/chat", response_model=ChatResponse)
def chat_v2(req: ChatRequest):
    return {
        "reply": "RAG combines retrieval with generation.",
        "model": req.model,
        "usage": {"prompt_tokens": 42, "completion_tokens": 9},
        "internal_cost_usd": 0.00031,             # not in ChatResponse → removed
        "system_prompt": "You are a helpful...",  # not in ChatResponse → removed
    }

print(client.post("/v2/chat", json={"messages": [{"role": "user", "content": "What is RAG?"}]}).json())
Output
{'reply': 'RAG combines retrieval with generation.', 'model': 'gpt-4o-mini', 'usage': {'prompt_tokens': 42, 'completion_tokens': 9}}

You can also write the return type instead: def chat_v2(req: ChatRequest) -> ChatResponse: — FastAPI treats it the same way.

2.3 Status codes

from fastapi import status

class IngestRequest(BaseModel):
    doc_id: str
    text: str

DOCS: dict[str, str] = {}

@app.post("/documents", status_code=status.HTTP_201_CREATED)
def ingest(req: IngestRequest):
    DOCS[req.doc_id] = req.text
    return {"doc_id": req.doc_id, "chars": len(req.text)}

@app.delete("/documents/{doc_id}", status_code=status.HTTP_204_NO_CONTENT)
def delete(doc_id: str):
    DOCS.pop(doc_id, None)

r = client.post("/documents", json={"doc_id": "policy", "text": "Refunds take 5 days."})
print(r.status_code, r.json())
print(client.delete("/documents/policy").status_code)
Output
201 {'doc_id': 'policy', 'chars': 20}
204
Code Meaning Typical GenAI case
200 OK answer returned
201 Created document ingested
400 Bad request prompt too long for the model
401 / 403 Not logged in / not allowed missing or wrong API key
404 Not found unknown document or session
422 Validation failed wrong JSON shape (automatic)
429 Too many requests rate limit hit
502 / 503 / 504 Upstream failed / busy / timed out LLM provider error

2.4 Errors — HTTPException

Raise it anywhere to stop and send an error response:

from fastapi import HTTPException

MAX_PROMPT_CHARS = 20

@app.get("/documents/{doc_id}")
def read_doc(doc_id: str):
    if doc_id not in DOCS:
        raise HTTPException(status_code=404, detail=f"Document {doc_id!r} not found")
    return {"doc_id": doc_id, "text": DOCS[doc_id]}

@app.post("/v3/chat")
def chat_v3(req: ChatRequest):
    if sum(len(m.content) for m in req.messages) > MAX_PROMPT_CHARS:
        raise HTTPException(status_code=400, detail="Prompt too long — trim the conversation history")
    return {"reply": "ok"}

r = client.get("/documents/missing")
print(r.status_code, r.json())
r = client.post("/v3/chat", json={"messages": [{"role": "user", "content": "x" * 50}]})
print(r.status_code, r.json())
Output
404 {'detail': "Document 'missing' not found"}
400 {'detail': 'Prompt too long — trim the conversation history'}

2.5 Turning your own exceptions into responses

Keep business code free of HTTP details: raise your own exception, map it to a response once. Register handlers (and middleware) when you create the app — before it serves any request.

from fastapi import Request
from fastapi.responses import JSONResponse

app = FastAPI()                               # a fresh app for this example
client = TestClient(app)

class LLMProviderError(Exception):
    def __init__(self, provider: str):
        self.provider = provider

@app.exception_handler(LLMProviderError)
async def llm_error_handler(request: Request, exc: LLMProviderError):
    return JSONResponse(status_code=502, content={"error": f"{exc.provider} is unavailable, try again"})

@app.post("/v4/chat")
def chat_v4(req: ChatRequest):
    raise LLMProviderError("openai")          # pretend the upstream call failed

r = client.post("/v4/chat", json={"messages": [{"role": "user", "content": "hi"}]})
print(r.status_code, r.json())
Output
502 {'error': 'openai is unavailable, try again'}

2.6 Form data and file uploads

RAG apps need uploads. Install python-multipart (included in fastapi[standard]):

from fastapi import UploadFile

@app.post("/upload")
async def upload(file: UploadFile):
    data = await file.read()
    return {"filename": file.filename, "content_type": file.content_type, "bytes": len(data)}

r = client.post("/upload", files={"file": ("notes.txt", b"RAG = retrieval + generation", "text/plain")})
print(r.json())
Output
{'filename': 'notes.txt', 'content_type': 'text/plain', 'bytes': 28}

Limit what you accept

Check file.content_type and the size before parsing. A 500 MB PDF or an unexpected file type should get a 400 / 413, not a crashed worker.

Practice

  • Add POST /embed taking {"texts": [...]} (1–100 strings) and returning fake vectors of length 4.
  • Make GET /documents/{doc_id} return a DocumentOut response model without the raw text.

Next: Dependencies & middleware — auth, settings and shared clients.