Skip to content

2. Fields & validators

Intermediate · 10 min read

Types say what kind of value a field holds. Fields and validators say which values are acceptable — and they double as documentation the LLM reads when you use the model as a schema.

2.1 Field — limits, defaults and descriptions

from pydantic import BaseModel, Field, ValidationError

class GenerationParams(BaseModel):
    temperature: float = Field(0.7, ge=0, le=2, description="Sampling temperature")
    max_tokens: int = Field(512, gt=0, le=8192)
    stop: list[str] = Field(default_factory=list, max_length=4)
    user_id: str = Field(min_length=3, pattern=r"^[a-z0-9_]+$")

print(GenerationParams(user_id="priya_01"))

try:
    GenerationParams(temperature=3, max_tokens=0, user_id="P!")
except ValidationError as e:
    for err in e.errors():
        print(err["loc"][0], "→", err["msg"])
Output
temperature=0.7 max_tokens=512 stop=[] user_id='priya_01'
temperature → Input should be less than or equal to 2
max_tokens → Input should be greater than 0
user_id → String should have at least 3 characters
Field argument Means
gt, ge, lt, le greater than, ≥, less than, ≤ (numbers)
min_length, max_length length of a string or list
pattern a regular expression the string must match
default_factory a function that builds the default (list, dict, uuid4…)
description text that appears in the JSON Schema — the LLM reads it

2.2 Allowed values — Literal and Enum

from typing import Literal

class Message(BaseModel):
    role: Literal["system", "user", "assistant", "tool"]
    content: str

print(Message(role="tool", content="42°C"))
try:
    Message(role="admin", content="hi")
except ValidationError as e:
    print(e.errors()[0]["msg"])
Output
role='tool' content='42°C'
Input should be 'system', 'user', 'assistant' or 'tool'

An Enum does the same and gives you a named constant to use in code:

from enum import Enum

class Intent(str, Enum):
    REFUND = "refund"
    SHIPPING = "shipping"
    OTHER = "other"

class Classification(BaseModel):
    intent: Intent
    confidence: float = Field(ge=0, le=1)

c = Classification(intent="refund", confidence=0.93)
print(c.intent, c.intent == Intent.REFUND, c.intent.value)
Output
Intent.REFUND True refund

Classification with an LLM

Literal / Enum is how you force an LLM to pick from a fixed list of labels. The allowed values appear in the schema, and anything else fails validation.

2.3 field_validator — custom checks on one field

A validator is a class method that receives the value and returns it (possibly changed) or raises ValueError.

from pydantic import field_validator

class Chunk(BaseModel):
    text: str
    source: str

    @field_validator("text")
    @classmethod
    def clean_text(cls, v: str) -> str:
        v = " ".join(v.split())               # collapse whitespace
        if not v:
            raise ValueError("chunk text is empty")
        return v

    @field_validator("source")
    @classmethod
    def must_be_pdf_or_url(cls, v: str) -> str:
        if not (v.endswith(".pdf") or v.startswith("http")):
            raise ValueError("source must be a .pdf file or a URL")
        return v

print(Chunk(text="  Refunds   are\n processed in 5 days. ", source="policy.pdf"))
try:
    Chunk(text="   ", source="notes.txt")
except ValidationError as e:
    for err in e.errors():
        print(err["loc"][0], "→", err["msg"])
Output
text='Refunds are processed in 5 days.' source='policy.pdf'
text → Value error, chunk text is empty
source → Value error, source must be a .pdf file or a URL

mode="before" runs your function before type checking — handy for messy input:

class Tags(BaseModel):
    tags: list[str]

    @field_validator("tags", mode="before")
    @classmethod
    def split_csv(cls, v):
        return [t.strip() for t in v.split(",")] if isinstance(v, str) else v

print(Tags(tags="rag, agents ,llm").tags)
print(Tags(tags=["already", "a list"]).tags)
Output
['rag', 'agents', 'llm']
['already', 'a list']

2.4 model_validator — checks across fields

When a rule involves more than one field, validate the whole model:

from pydantic import model_validator

class RetrievalConfig(BaseModel):
    top_k: int = 20
    rerank_top_n: int = 5

    @model_validator(mode="after")
    def rerank_not_bigger_than_retrieve(self):
        if self.rerank_top_n > self.top_k:
            raise ValueError("rerank_top_n cannot be larger than top_k")
        return self

print(RetrievalConfig())
try:
    RetrievalConfig(top_k=3, rerank_top_n=10)
except ValidationError as e:
    print(e.errors()[0]["msg"])
Output
top_k=20 rerank_top_n=5
Value error, rerank_top_n cannot be larger than top_k

2.5 Computed fields

A value derived from other fields, included when you export the model:

from pydantic import computed_field

class Usage(BaseModel):
    prompt_tokens: int
    completion_tokens: int

    @computed_field
    @property
    def total_tokens(self) -> int:
        return self.prompt_tokens + self.completion_tokens

u = Usage(prompt_tokens=1200, completion_tokens=300)
print(u.total_tokens)
print(u.model_dump())
Output
1500
{'prompt_tokens': 1200, 'completion_tokens': 300, 'total_tokens': 1500}

2.6 Reusable constrained types — Annotated

Write a rule once and reuse it in many models:

from typing import Annotated

Score = Annotated[float, Field(ge=0, le=1)]
NonEmpty = Annotated[str, Field(min_length=1)]

class Judgement(BaseModel):
    faithfulness: Score
    relevance: Score
    reason: NonEmpty

print(Judgement(faithfulness=0.9, relevance=1, reason="Grounded in chunk 2."))
Output
faithfulness=0.9 relevance=1.0 reason='Grounded in chunk 2.'

Practice

  • Add a field_validator to Message that strips whitespace from content and rejects empty strings.
  • Write a DateRange model whose model_validator checks start <= end.

Next: Nested models & JSON — real LLM responses are trees, not flat dicts.