Skip to content

5. Pydantic for GenAI

Intermediate · 12 min read

An LLM returns text. Your code needs data. Pydantic is the bridge in both directions: it describes the shape you want (as JSON Schema the model can read) and checks what comes back.

flowchart LR
    M[Pydantic model] -->|model_json_schema| S[JSON Schema in the prompt / tool definition]
    S --> L[LLM]
    L -->|JSON text| V{model_validate_json}
    V -->|valid| D[Typed Python object]
    V -->|ValidationError| R[Send errors back, retry]
    R --> L

5.1 Model → JSON Schema

import json
from typing import Literal
from pydantic import BaseModel, Field

class TicketTriage(BaseModel):
    """Classify a customer support ticket."""
    category: Literal["billing", "bug", "feature_request", "other"]
    priority: int = Field(ge=1, le=5, description="1 = lowest, 5 = urgent")
    summary: str = Field(max_length=120, description="One-line summary for the agent dashboard")

print(json.dumps(TicketTriage.model_json_schema(), indent=2))
Output
{
  "description": "Classify a customer support ticket.",
  "properties": {
    "category": {
      "enum": [
        "billing",
        "bug",
        "feature_request",
        "other"
      ],
      "title": "Category",
      "type": "string"
    },
    "priority": {
      "description": "1 = lowest, 5 = urgent",
      "maximum": 5,
      "minimum": 1,
      "title": "Priority",
      "type": "integer"
    },
    "summary": {
      "description": "One-line summary for the agent dashboard",
      "maxLength": 120,
      "title": "Summary",
      "type": "string"
    }
  },
  "required": [
    "category",
    "priority",
    "summary"
  ],
  "title": "TicketTriage",
  "type": "object"
}

The docstring and every description end up in the schema — write them for the LLM, they are part of your prompt.

5.2 Structured output with an SDK

Modern SDKs take the Pydantic model directly and return a validated object:

# no-run — needs an API key
from openai import OpenAI
client = OpenAI()

completion = client.chat.completions.parse(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "My card was charged twice for one order!"}],
    response_format=TicketTriage,
)
triage = completion.choices[0].message.parsed      # a TicketTriage instance
print(triage.category, triage.priority)
# no-run — needs an API key
import anthropic
client = anthropic.Anthropic()

msg = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=512,
    tools=[{"name": "triage", "description": TicketTriage.__doc__,
            "input_schema": TicketTriage.model_json_schema()}],
    tool_choice={"type": "tool", "name": "triage"},
    messages=[{"role": "user", "content": "My card was charged twice for one order!"}],
)
tool_use = next(b for b in msg.content if b.type == "tool_use")
triage = TicketTriage.model_validate(tool_use.input)

5.3 Validate and retry (works with any LLM)

If your model or provider has no structured-output mode, ask for JSON, validate, and on failure send the errors back. Here a fake LLM stands in for the real call:

from pydantic import ValidationError

fake_replies = iter([
    '{"category": "payments", "priority": 9, "summary": "Charged twice"}',      # wrong
    '{"category": "billing", "priority": 4, "summary": "Charged twice for one order"}',
])

def call_llm(messages: list[dict]) -> str:
    return next(fake_replies)                      # swap for a real API call

def get_structured(prompt: str, schema: type[BaseModel], max_attempts: int = 3) -> BaseModel:
    messages = [
        {"role": "system", "content": "Reply with JSON only, matching this schema:\n"
                                      + json.dumps(schema.model_json_schema())},
        {"role": "user", "content": prompt},
    ]
    for attempt in range(1, max_attempts + 1):
        reply = call_llm(messages)
        try:
            return schema.model_validate_json(reply)
        except ValidationError as e:
            problems = "; ".join(f"{'.'.join(map(str, err['loc']))}: {err['msg']}" for err in e.errors())
            print(f"attempt {attempt} failed → {problems}")
            messages += [{"role": "assistant", "content": reply},
                         {"role": "user", "content": f"Invalid JSON: {problems}. Fix it and reply again."}]
    raise RuntimeError("LLM never returned valid JSON")

print(get_structured("My card was charged twice!", TicketTriage))
Output
attempt 1 failed → category: Input should be 'billing', 'bug', 'feature_request' or 'other'; priority: Input should be less than or equal to 5
category='billing' priority=4 summary='Charged twice for one order'

Libraries that do this for you

Instructor, PydanticAI and LangChain's with_structured_output() wrap exactly this loop. Knowing the loop yourself makes their behaviour — and their bugs — easy to understand.

5.4 Tool calling — tool arguments as models

An agent's tools are functions whose arguments the LLM fills in. Describe the arguments with a model, send the schema, and validate the call before running anything:

class GetWeather(BaseModel):
    """Get the current weather for a city."""
    city: str = Field(description="City name, e.g. 'Bengaluru'")
    unit: Literal["celsius", "fahrenheit"] = "celsius"

class SearchDocs(BaseModel):
    """Search the internal knowledge base."""
    query: str
    top_k: int = Field(3, ge=1, le=10)

TOOLS = {m.__name__: m for m in (GetWeather, SearchDocs)}

def tool_definitions() -> list[dict]:        # the OpenAI "tools" format
    return [{"type": "function", "function": {
                "name": name, "description": m.__doc__, "parameters": m.model_json_schema()}}
            for name, m in TOOLS.items()]

print([t["function"]["name"] for t in tool_definitions()])
print(tool_definitions()[0]["function"]["parameters"]["required"])
Output
['GetWeather', 'SearchDocs']
['city']

When the LLM answers with a tool call, validate the arguments, then run the real function:

def run_tool(name: str, arguments_json: str) -> str:
    args = TOOLS[name].model_validate_json(arguments_json)    # raises on bad arguments
    if isinstance(args, GetWeather):
        return f"28° {args.unit} and cloudy in {args.city}"
    if isinstance(args, SearchDocs):
        return f"top {args.top_k} results for {args.query!r}"

print(run_tool("GetWeather", '{"city": "Bengaluru"}'))
print(run_tool("SearchDocs", '{"query": "refund policy", "top_k": "5"}'))
try:
    run_tool("SearchDocs", '{"query": "refund policy", "top_k": 50}')
except ValidationError as e:
    print("rejected:", e.errors()[0]["msg"])
Output
28° celsius and cloudy in Bengaluru
top 5 results for 'refund policy'
rejected: Input should be less than or equal to 10

Never trust tool arguments

The LLM chose those arguments — possibly under the influence of a prompt injection in a retrieved document. Validation (limits, allowed values, patterns) is your first guard before a tool touches a database, a file system or a payment API.

5.5 Typed agent state

Agent frameworks such as LangGraph pass a state object from step to step. A model keeps it honest:

class ToolCall(BaseModel):
    name: str
    arguments: dict
    result: str | None = None

class AgentState(BaseModel):
    question: str
    messages: list[dict] = []
    tool_calls: list[ToolCall] = []
    steps: int = Field(0, le=5)           # hard cap on loop iterations
    final_answer: str | None = None

state = AgentState(question="Weather in Bengaluru?")
state.tool_calls.append(ToolCall(name="GetWeather", arguments={"city": "Bengaluru"},
                                 result=run_tool("GetWeather", '{"city": "Bengaluru"}')))
state.steps += 1
state.final_answer = state.tool_calls[-1].result
print(state.model_dump(exclude={"messages"}))
Output
{'question': 'Weather in Bengaluru?', 'tool_calls': [{'name': 'GetWeather', 'arguments': {'city': 'Bengaluru'}, 'result': '28° celsius and cloudy in Bengaluru'}], 'steps': 1, 'final_answer': '28° celsius and cloudy in Bengaluru'}

model_dump_json() turns the state into a string you can save after every step — so a crashed agent can resume, and you can replay a run when debugging.

5.6 Extraction — free text → records

The most common GenAI job in business: pull fields out of emails, invoices, resumes.

from datetime import date

class Invoice(BaseModel):
    vendor: str
    invoice_number: str
    total_inr: float = Field(gt=0)
    due_date: date
    line_items: list[str] = []

llm_reply = '{"vendor": "Acme Cloud", "invoice_number": "INV-0042", "total_inr": "18500", "due_date": "2026-10-31"}'
inv = Invoice.model_validate_json(llm_reply)
print(inv.total_inr, inv.due_date, (inv.due_date - date(2026, 10, 5)).days, "days left")
Output
18500.0 2026-10-31 26 days left

The date arrives as a string and leaves as a real date you can do maths with — that conversion is the whole point.

Interview questions

Why use Pydantic instead of json.loads for LLM output?

json.loads only checks that the text is JSON. Pydantic also checks the shape — required keys, types, allowed values, ranges — and converts types (strings to dates, numbers). It gives precise error messages you can feed back to the LLM for a retry, and the same model generates the JSON Schema that tells the LLM what to produce.

How do you make an LLM reliably return structured data?

1) Use the provider's structured-output / tool-calling mode with a schema from a Pydantic model. 2) Write clear field descriptions. 3) Validate the reply. 4) On failure, retry with the validation errors in the prompt, with a max attempt count. 5) Keep the schema small and flat; use Literal for labels.

Why validate tool-call arguments in an agent?

The LLM can produce wrong types, out-of-range values or malicious values (prompt injection). Validating before execution stops bad calls from reaching real systems and gives the agent an error it can correct.

Practice

  • Write a ResumeSummary model (name, years of experience, skills list, seniority: Literal[...]) and print its schema.
  • Extend get_structured to raise a custom exception that includes the last reply and errors.

Next: FastAPI — serve these models over HTTP.