3. Nested models & JSON¶
Intermediate · 9 min read
Real data is a tree: a chat request holds messages, a RAG answer holds sources, an agent step holds tool calls. Pydantic validates the whole tree in one go and turns it into JSON and back.
3.1 Models inside models¶
from pydantic import BaseModel
class Source(BaseModel):
doc_id: str
page: int
score: float
class RAGAnswer(BaseModel):
answer: str
sources: list[Source]
raw = {
"answer": "Refunds take 5 working days.",
"sources": [
{"doc_id": "policy.pdf", "page": "3", "score": 0.91},
{"doc_id": "faq.pdf", "page": 1, "score": "0.78"},
],
}
ans = RAGAnswer.model_validate(raw)
print(ans.sources[0].page + 1) # "3" was converted to 3
print([s.doc_id for s in ans.sources])
Errors point to the exact place in the tree:
from pydantic import ValidationError
raw["sources"][1]["score"] = "high"
try:
RAGAnswer.model_validate(raw)
except ValidationError as e:
print(e.errors()[0]["loc"])
3.2 To dict and JSON — model_dump¶
ans = RAGAnswer(answer="Yes.", sources=[Source(doc_id="a.pdf", page=2, score=0.8)])
print(ans.model_dump())
print(ans.model_dump_json())
print(ans.model_dump(include={"answer"}))
print(ans.model_dump(exclude={"sources": {0: {"score"}}}))
{'answer': 'Yes.', 'sources': [{'doc_id': 'a.pdf', 'page': 2, 'score': 0.8}]}
{"answer":"Yes.","sources":[{"doc_id":"a.pdf","page":2,"score":0.8}]}
{'answer': 'Yes.'}
{'answer': 'Yes.', 'sources': [{'doc_id': 'a.pdf', 'page': 2}]}
Useful model_dump options:
| Option | Effect |
|---|---|
exclude_none=True |
drop fields whose value is None (smaller API payloads) |
exclude_unset=True |
keep only fields the caller actually set — good for PATCH updates |
mode="json" |
return JSON-safe types (datetime → string, UUID → string) |
by_alias=True |
use the alias names (next section) |
from datetime import datetime
class LogEntry(BaseModel):
at: datetime
model: str
error: str | None = None
e = LogEntry(at=datetime(2026, 10, 5, 9, 30), model="gpt-4o-mini")
print(e.model_dump(exclude_none=True))
print(e.model_dump(mode="json", exclude_none=True))
{'at': datetime.datetime(2026, 10, 5, 9, 30), 'model': 'gpt-4o-mini'}
{'at': '2026-10-05T09:30:00', 'model': 'gpt-4o-mini'}
3.3 From JSON text — model_validate_json¶
LLMs and HTTP APIs give you a string. Parse and validate in one step (faster than json.loads
followed by model_validate):
reply = '{"answer": "Use pgvector.", "sources": [{"doc_id": "db.md", "page": 1, "score": 0.88}]}'
ans = RAGAnswer.model_validate_json(reply)
print(ans.answer, ans.sources[0].score)
Broken JSON is a ValidationError too, so one except handles both cases:
try:
RAGAnswer.model_validate_json('{"answer": "cut off mid-str')
except ValidationError as e:
print(e.errors()[0]["type"])
3.4 Lists of models — TypeAdapter¶
When the top level is a list (a JSONL file, a batch response), there is no model to call. Use a
TypeAdapter:
from pydantic import TypeAdapter
sources = TypeAdapter(list[Source]).validate_json(
'[{"doc_id": "a", "page": 1, "score": 0.9}, {"doc_id": "b", "page": 4, "score": 0.7}]'
)
print(len(sources), sources[1].page)
3.5 Aliases — camelCase JSON, snake_case Python¶
APIs often send camelCase. Keep Python names clean and map them:
from pydantic import ConfigDict, Field
from pydantic.alias_generators import to_camel
class TokenUsage(BaseModel):
model_config = ConfigDict(alias_generator=to_camel, populate_by_name=True)
prompt_tokens: int
completion_tokens: int
cached_tokens: int = 0
u = TokenUsage.model_validate({"promptTokens": 900, "completionTokens": 120})
print(u.prompt_tokens)
print(u.model_dump(by_alias=True))
print(TokenUsage(prompt_tokens=1, completion_tokens=2).prompt_tokens) # works thanks to populate_by_name
For a single odd key use Field(alias="..."), e.g. id_: str = Field(alias="id").
3.6 Union types — one field, several shapes¶
Agent messages can be text or a tool call. Give each shape a type tag and Pydantic picks the right
model:
from typing import Literal, Union, Annotated
class TextPart(BaseModel):
type: Literal["text"]
text: str
class ToolCallPart(BaseModel):
type: Literal["tool_call"]
name: str
arguments: dict
Part = Annotated[Union[TextPart, ToolCallPart], Field(discriminator="type")]
class AssistantTurn(BaseModel):
parts: list[Part]
turn = AssistantTurn.model_validate({"parts": [
{"type": "text", "text": "Let me check the weather."},
{"type": "tool_call", "name": "get_weather", "arguments": {"city": "Bengaluru"}},
]})
for p in turn.parts:
print(type(p).__name__, "→", p.text if isinstance(p, TextPart) else p.name)
This is a discriminated union — the same pattern the OpenAI and Anthropic SDKs use for message content blocks.
Practice¶
- Write a
Documentmodel holding a list ofChunkmodels, then dump it to JSON withexclude_none=True. - Save three
Sourceobjects to a.jsonlfile (onemodel_dump_json()per line) and load them back.
Next: Settings & secrets — config and API keys, validated at startup.