Skip to content

3. Nested models & JSON

Intermediate · 9 min read

Real data is a tree: a chat request holds messages, a RAG answer holds sources, an agent step holds tool calls. Pydantic validates the whole tree in one go and turns it into JSON and back.

3.1 Models inside models

from pydantic import BaseModel

class Source(BaseModel):
    doc_id: str
    page: int
    score: float

class RAGAnswer(BaseModel):
    answer: str
    sources: list[Source]

raw = {
    "answer": "Refunds take 5 working days.",
    "sources": [
        {"doc_id": "policy.pdf", "page": "3", "score": 0.91},
        {"doc_id": "faq.pdf", "page": 1, "score": "0.78"},
    ],
}
ans = RAGAnswer.model_validate(raw)
print(ans.sources[0].page + 1)                 # "3" was converted to 3
print([s.doc_id for s in ans.sources])
Output
4
['policy.pdf', 'faq.pdf']

Errors point to the exact place in the tree:

from pydantic import ValidationError

raw["sources"][1]["score"] = "high"
try:
    RAGAnswer.model_validate(raw)
except ValidationError as e:
    print(e.errors()[0]["loc"])
Output
('sources', 1, 'score')

3.2 To dict and JSON — model_dump

ans = RAGAnswer(answer="Yes.", sources=[Source(doc_id="a.pdf", page=2, score=0.8)])

print(ans.model_dump())
print(ans.model_dump_json())
print(ans.model_dump(include={"answer"}))
print(ans.model_dump(exclude={"sources": {0: {"score"}}}))
Output
{'answer': 'Yes.', 'sources': [{'doc_id': 'a.pdf', 'page': 2, 'score': 0.8}]}
{"answer":"Yes.","sources":[{"doc_id":"a.pdf","page":2,"score":0.8}]}
{'answer': 'Yes.'}
{'answer': 'Yes.', 'sources': [{'doc_id': 'a.pdf', 'page': 2}]}

Useful model_dump options:

Option Effect
exclude_none=True drop fields whose value is None (smaller API payloads)
exclude_unset=True keep only fields the caller actually set — good for PATCH updates
mode="json" return JSON-safe types (datetime → string, UUID → string)
by_alias=True use the alias names (next section)
from datetime import datetime

class LogEntry(BaseModel):
    at: datetime
    model: str
    error: str | None = None

e = LogEntry(at=datetime(2026, 10, 5, 9, 30), model="gpt-4o-mini")
print(e.model_dump(exclude_none=True))
print(e.model_dump(mode="json", exclude_none=True))
Output
{'at': datetime.datetime(2026, 10, 5, 9, 30), 'model': 'gpt-4o-mini'}
{'at': '2026-10-05T09:30:00', 'model': 'gpt-4o-mini'}

3.3 From JSON text — model_validate_json

LLMs and HTTP APIs give you a string. Parse and validate in one step (faster than json.loads followed by model_validate):

reply = '{"answer": "Use pgvector.", "sources": [{"doc_id": "db.md", "page": 1, "score": 0.88}]}'
ans = RAGAnswer.model_validate_json(reply)
print(ans.answer, ans.sources[0].score)
Output
Use pgvector. 0.88

Broken JSON is a ValidationError too, so one except handles both cases:

try:
    RAGAnswer.model_validate_json('{"answer": "cut off mid-str')
except ValidationError as e:
    print(e.errors()[0]["type"])
Output
json_invalid

3.4 Lists of models — TypeAdapter

When the top level is a list (a JSONL file, a batch response), there is no model to call. Use a TypeAdapter:

from pydantic import TypeAdapter

sources = TypeAdapter(list[Source]).validate_json(
    '[{"doc_id": "a", "page": 1, "score": 0.9}, {"doc_id": "b", "page": 4, "score": 0.7}]'
)
print(len(sources), sources[1].page)
Output
2 4

3.5 Aliases — camelCase JSON, snake_case Python

APIs often send camelCase. Keep Python names clean and map them:

from pydantic import ConfigDict, Field
from pydantic.alias_generators import to_camel

class TokenUsage(BaseModel):
    model_config = ConfigDict(alias_generator=to_camel, populate_by_name=True)
    prompt_tokens: int
    completion_tokens: int
    cached_tokens: int = 0

u = TokenUsage.model_validate({"promptTokens": 900, "completionTokens": 120})
print(u.prompt_tokens)
print(u.model_dump(by_alias=True))
print(TokenUsage(prompt_tokens=1, completion_tokens=2).prompt_tokens)   # works thanks to populate_by_name
Output
900
{'promptTokens': 900, 'completionTokens': 120, 'cachedTokens': 0}
1

For a single odd key use Field(alias="..."), e.g. id_: str = Field(alias="id").

3.6 Union types — one field, several shapes

Agent messages can be text or a tool call. Give each shape a type tag and Pydantic picks the right model:

from typing import Literal, Union, Annotated

class TextPart(BaseModel):
    type: Literal["text"]
    text: str

class ToolCallPart(BaseModel):
    type: Literal["tool_call"]
    name: str
    arguments: dict

Part = Annotated[Union[TextPart, ToolCallPart], Field(discriminator="type")]

class AssistantTurn(BaseModel):
    parts: list[Part]

turn = AssistantTurn.model_validate({"parts": [
    {"type": "text", "text": "Let me check the weather."},
    {"type": "tool_call", "name": "get_weather", "arguments": {"city": "Bengaluru"}},
]})
for p in turn.parts:
    print(type(p).__name__, "→", p.text if isinstance(p, TextPart) else p.name)
Output
TextPart → Let me check the weather.
ToolCallPart → get_weather

This is a discriminated union — the same pattern the OpenAI and Anthropic SDKs use for message content blocks.

Practice

  • Write a Document model holding a list of Chunk models, then dump it to JSON with exclude_none=True.
  • Save three Source objects to a .jsonl file (one model_dump_json() per line) and load them back.

Next: Settings & secrets — config and API keys, validated at startup.