Skip to content

1. Models & validation

Beginner · 8 min read

A model is a class that lists fields and their types. When you create one, Pydantic checks every value — and converts it when that is safe — so the rest of your code can trust the data.

1.1 Your first model

from pydantic import BaseModel

class ChatMessage(BaseModel):
    role: str
    content: str

msg = ChatMessage(role="user", content="What is RAG?")
print(msg)
print(msg.role, "|", msg.content)
Output
role='user' content='What is RAG?'
user | What is RAG?

It looks like a dataclass — the difference shows up when the data is wrong.

1.2 Conversion and errors

Pydantic converts values that clearly mean the right thing, and rejects the rest.

class LLMConfig(BaseModel):
    model: str
    temperature: float
    max_tokens: int

cfg = LLMConfig(model="gpt-4o-mini", temperature="0.2", max_tokens="512")   # strings from a form / env var
print(cfg.temperature, type(cfg.temperature).__name__)
print(cfg.max_tokens, type(cfg.max_tokens).__name__)
Output
0.2 float
512 int
from pydantic import ValidationError

try:
    LLMConfig(model="gpt-4o-mini", temperature="hot", max_tokens=5.5)
except ValidationError as e:
    print(e.error_count(), "errors")
    for err in e.errors():
        print(err["loc"], "→", err["msg"])
Output
2 errors
('temperature',) → Input should be a valid number, unable to parse string as a number
('max_tokens',) → Input should be a valid integer, got a number with a fractional part

Read errors with e.errors()

Each error has loc (which field), msg (what went wrong) and type. That list is exactly what you send back to an LLM when its JSON is invalid — see Pydantic for GenAI.

1.3 Defaults and optional fields

class SearchRequest(BaseModel):
    query: str                      # required
    top_k: int = 5                  # default
    namespace: str | None = None    # optional — may be missing or None

print(SearchRequest(query="refund policy"))
print(SearchRequest(query="refund policy", top_k=3, namespace="support"))
Output
query='refund policy' top_k=5 namespace=None
query='refund policy' top_k=3 namespace='support'

str | None without a default is still required

namespace: str | None means "you must pass it, but it may be None". Add = None to make it truly optional.

Mutable defaults are safe in Pydantic — each instance gets its own copy:

class Conversation(BaseModel):
    messages: list[ChatMessage] = []

a, b = Conversation(), Conversation()
a.messages.append(ChatMessage(role="user", content="hi"))
print(len(a.messages), len(b.messages))
Output
1 0

1.4 Creating models from dicts

Data usually arrives as a dict (parsed JSON). Use model_validate:

data = {"role": "assistant", "content": "RAG = retrieval-augmented generation."}
msg = ChatMessage.model_validate(data)
print(msg.content)
Output
RAG = retrieval-augmented generation.

ChatMessage(**data) also works, but model_validate reads better and accepts other objects too.

1.5 Strict mode — no conversion

Sometimes "5" should not become 5. Turn conversion off for a model:

from pydantic import ConfigDict

class StrictScore(BaseModel):
    model_config = ConfigDict(strict=True)
    score: float

print(StrictScore(score=0.9).score)
try:
    StrictScore(score="0.9")
except ValidationError as e:
    print(e.errors()[0]["msg"])
Output
0.9
Input should be a valid number

1.6 Forbid unknown fields

By default extra keys are silently ignored. For LLM output that is often a bug you want to see:

class Answer(BaseModel):
    model_config = ConfigDict(extra="forbid")
    answer: str

try:
    Answer.model_validate({"answer": "42", "confidence": 0.8})
except ValidationError as e:
    print(e.errors()[0]["loc"], e.errors()[0]["msg"])
Output
('confidence',) Extra inputs are not permitted

1.7 Copying and changing a model

Models are mutable by default, but model_copy(update=...) is the cleaner way to make a variant:

base = LLMConfig(model="gpt-4o-mini", temperature=0.2, max_tokens=512)
creative = base.model_copy(update={"temperature": 0.9})
print(base.temperature, creative.temperature)
Output
0.2 0.9

model_copy(update=...) does not validate

The update values are copied as-is. To validate, build a new model: LLMConfig.model_validate(base.model_dump() | {"temperature": 0.9}).

Practice

  • Write a Document model with id: str, text: str, source: str | None = None, tokens: int = 0.
  • Try Document(id=1, text="hi") — what happens to id? Then try it with strict=True.

Next: Fields & validators — limits, allowed values and custom checks.