1. Models & validation¶
Beginner · 8 min read
A model is a class that lists fields and their types. When you create one, Pydantic checks every value — and converts it when that is safe — so the rest of your code can trust the data.
1.1 Your first model¶
from pydantic import BaseModel
class ChatMessage(BaseModel):
role: str
content: str
msg = ChatMessage(role="user", content="What is RAG?")
print(msg)
print(msg.role, "|", msg.content)
It looks like a dataclass — the difference shows up when the data is wrong.
1.2 Conversion and errors¶
Pydantic converts values that clearly mean the right thing, and rejects the rest.
class LLMConfig(BaseModel):
model: str
temperature: float
max_tokens: int
cfg = LLMConfig(model="gpt-4o-mini", temperature="0.2", max_tokens="512") # strings from a form / env var
print(cfg.temperature, type(cfg.temperature).__name__)
print(cfg.max_tokens, type(cfg.max_tokens).__name__)
from pydantic import ValidationError
try:
LLMConfig(model="gpt-4o-mini", temperature="hot", max_tokens=5.5)
except ValidationError as e:
print(e.error_count(), "errors")
for err in e.errors():
print(err["loc"], "→", err["msg"])
2 errors
('temperature',) → Input should be a valid number, unable to parse string as a number
('max_tokens',) → Input should be a valid integer, got a number with a fractional part
Read errors with e.errors()
Each error has loc (which field), msg (what went wrong) and type. That list is exactly what you
send back to an LLM when its JSON is invalid — see Pydantic for GenAI.
1.3 Defaults and optional fields¶
class SearchRequest(BaseModel):
query: str # required
top_k: int = 5 # default
namespace: str | None = None # optional — may be missing or None
print(SearchRequest(query="refund policy"))
print(SearchRequest(query="refund policy", top_k=3, namespace="support"))
query='refund policy' top_k=5 namespace=None
query='refund policy' top_k=3 namespace='support'
str | None without a default is still required
namespace: str | None means "you must pass it, but it may be None". Add = None to make it
truly optional.
Mutable defaults are safe in Pydantic — each instance gets its own copy:
class Conversation(BaseModel):
messages: list[ChatMessage] = []
a, b = Conversation(), Conversation()
a.messages.append(ChatMessage(role="user", content="hi"))
print(len(a.messages), len(b.messages))
1.4 Creating models from dicts¶
Data usually arrives as a dict (parsed JSON). Use model_validate:
data = {"role": "assistant", "content": "RAG = retrieval-augmented generation."}
msg = ChatMessage.model_validate(data)
print(msg.content)
ChatMessage(**data) also works, but model_validate reads better and accepts other objects too.
1.5 Strict mode — no conversion¶
Sometimes "5" should not become 5. Turn conversion off for a model:
from pydantic import ConfigDict
class StrictScore(BaseModel):
model_config = ConfigDict(strict=True)
score: float
print(StrictScore(score=0.9).score)
try:
StrictScore(score="0.9")
except ValidationError as e:
print(e.errors()[0]["msg"])
1.6 Forbid unknown fields¶
By default extra keys are silently ignored. For LLM output that is often a bug you want to see:
class Answer(BaseModel):
model_config = ConfigDict(extra="forbid")
answer: str
try:
Answer.model_validate({"answer": "42", "confidence": 0.8})
except ValidationError as e:
print(e.errors()[0]["loc"], e.errors()[0]["msg"])
1.7 Copying and changing a model¶
Models are mutable by default, but model_copy(update=...) is the cleaner way to make a variant:
base = LLMConfig(model="gpt-4o-mini", temperature=0.2, max_tokens=512)
creative = base.model_copy(update={"temperature": 0.9})
print(base.temperature, creative.temperature)
model_copy(update=...) does not validate
The update values are copied as-is. To validate, build a new model:
LLMConfig.model_validate(base.model_dump() | {"temperature": 0.9}).
Practice¶
- Write a
Documentmodel withid: str,text: str,source: str | None = None,tokens: int = 0. - Try
Document(id=1, text="hi")— what happens toid? Then try it withstrict=True.
Next: Fields & validators — limits, allowed values and custom checks.