6. Structured output & tool calling¶
Intermediate · 14 min read
Text is for people. Your code needs data (structured output) and the model sometimes needs to do things — look up an order, search documents, call an API (tool calling). Both use the same mechanism: a JSON schema that the model's output must follow. Tool calling in a loop is what an agent is.
6.1 Three levels of structured output¶
| Level | How | Reliability |
|---|---|---|
| Ask nicely | "Reply with JSON only: {…}" in the prompt | often works; sometimes adds prose, code fences, or wrong keys |
| JSON mode | response_format={"type": "json_object"} |
always valid JSON, but any shape |
| Schema-constrained | response_format with a JSON schema / Pydantic model, or a forced tool call |
valid JSON matching your schema — use this |
The Pydantic side — defining the schema, validating, retrying with errors — is covered in depth in Pydantic for GenAI. Even with a schema-constrained mode, validate the result: the schema can't check business rules (a date in the past, an ID that doesn't exist).
When you must parse "ask nicely" output (older or local models), strip the usual wrappers first:
import json, re
def extract_json(reply: str) -> dict:
"""Pull a JSON object out of a reply that may include prose or ```json fences."""
fenced = re.search(r"```(?:json)?\s*(\{.*?\})\s*```", reply, re.S)
candidate = fenced.group(1) if fenced else reply[reply.find("{"): reply.rfind("}") + 1]
return json.loads(candidate)
print(extract_json('Sure! Here is the data:\n```json\n{"city": "Pune", "days": 3}\n```\nAnything else?'))
print(extract_json('The answer is {"city": "Delhi", "days": 2} as requested.'))
6.2 How tool calling works¶
The model never runs your code. It asks you to run a tool by returning a structured tool call; your code runs it and sends the result back; the model continues.
sequenceDiagram
participant App
participant LLM
participant Tool as Your function
App->>LLM: messages + tool definitions
LLM-->>App: tool_call: get_order_status(order_id="A-1029")
App->>Tool: get_order_status("A-1029")
Tool-->>App: {"status": "shipped", "eta": "2026-10-08"}
App->>LLM: messages + assistant tool_call + tool result
LLM-->>App: "Your order A-1029 has shipped and should arrive on 8 October."
A tool definition is a name, a description and a JSON schema for the arguments (OpenAI format):
TOOLS = [
{"type": "function", "function": {
"name": "get_order_status",
"description": "Look up the shipping status of a customer's order. Use when the user asks where their order is.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string", "description": "Order ID like 'A-1029'"}},
"required": ["order_id"],
},
}},
{"type": "function", "function": {
"name": "search_policy",
"description": "Search the store's policy documents (refunds, shipping, warranty). Returns relevant passages.",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
}},
]
print([t["function"]["name"] for t in TOOLS])
The model decides whether to call a tool and which one mainly from the name and description — they
are part of your prompt. Generate the parameters from a Pydantic model rather than writing them by hand.
6.3 The message protocol¶
A tool call adds two kinds of messages to the history: the assistant's tool call (with an id), and a
tool message carrying the result for that id:
messages = [
{"role": "system", "content": "You are Acme's support assistant. Use tools to look things up."},
{"role": "user", "content": "Where is my order A-1029?"},
# ← returned by the model:
{"role": "assistant", "content": None, "tool_calls": [
{"id": "call_1", "type": "function",
"function": {"name": "get_order_status", "arguments": '{"order_id": "A-1029"}'}}]},
# ← added by your code after running the tool:
{"role": "tool", "tool_call_id": "call_1", "content": '{"status": "shipped", "eta": "2026-10-08"}'},
]
for m in messages:
print(f"{m['role']:9}", m.get("content") or m["tool_calls"][0]["function"])
system You are Acme's support assistant. Use tools to look things up.
user Where is my order A-1029?
assistant {'name': 'get_order_status', 'arguments': '{"order_id": "A-1029"}'}
tool {"status": "shipped", "eta": "2026-10-08"}
Note: arguments arrives as a JSON string — parse and validate it before use. Anthropic's API uses the same
idea with different names (tool_use and tool_result content blocks).
6.4 The tool loop¶
Call the model; if it asks for tools, run them, append the results, and call again — until it answers in text or you hit a step limit. This loop is an agent. Here a scripted fake model plays the LLM's part so you can see every step:
ORDERS = {"A-1029": {"status": "shipped", "eta": "2026-10-08"}}
POLICY = {"refund": "Refunds are processed within 5 working days of approval.",
"warranty": "Electronics carry a 1-year warranty."}
def get_order_status(order_id: str) -> dict:
if order_id not in ORDERS:
raise ValueError(f"No order with id {order_id}")
return ORDERS[order_id]
def search_policy(query: str) -> str:
hits = [text for key, text in POLICY.items() if key in query.lower()]
return " ".join(hits) or "No matching policy found."
REGISTRY = {"get_order_status": get_order_status, "search_policy": search_policy}
class ScriptedLLM:
"""Returns pre-written turns, like a real model deciding to call tools then answer."""
def __init__(self, turns): self.turns = iter(turns)
def __call__(self, messages, tools): return next(self.turns)
def call(id_, name, **args):
return {"id": id_, "type": "function", "function": {"name": name, "arguments": json.dumps(args)}}
def run_tool(tool_call) -> str:
fn = REGISTRY.get(tool_call["function"]["name"])
try:
if fn is None:
raise ValueError("unknown tool")
result = fn(**json.loads(tool_call["function"]["arguments"]))
return json.dumps(result) if not isinstance(result, str) else result
except Exception as e: # errors go BACK to the model, not up to the user
return f"ERROR: {e}"
def agent(user_msg: str, llm, max_steps: int = 5) -> str:
messages = [{"role": "system", "content": "You are Acme's support assistant."},
{"role": "user", "content": user_msg}]
for step in range(1, max_steps + 1):
reply = llm(messages, TOOLS)
messages.append(reply)
if not reply.get("tool_calls"):
return reply["content"]
for tc in reply["tool_calls"]: # may be several: parallel tool calls
result = run_tool(tc)
print(f" step {step}: {tc['function']['name']}({tc['function']['arguments']}) → {result}")
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
return "Sorry, I couldn't finish that request." # step limit: never loop forever
llm = ScriptedLLM([
{"role": "assistant", "content": None, "tool_calls": [
call("c1", "get_order_status", order_id="A-1092"), # typo in the id
call("c2", "search_policy", query="refund timeline")]}, # two tools in one turn
{"role": "assistant", "content": None, "tool_calls": [
call("c3", "get_order_status", order_id="A-1029")]}, # model corrects itself after the error
{"role": "assistant", "content": "Order A-1029 has shipped (ETA 8 Oct). If you return it, "
"refunds are processed within 5 working days."},
])
print(agent("Where's my order A-1029, and how long would a refund take?", llm))
step 1: get_order_status({"order_id": "A-1092"}) → ERROR: No order with id A-1092
step 1: search_policy({"query": "refund timeline"}) → Refunds are processed within 5 working days of approval.
step 2: get_order_status({"order_id": "A-1029"}) → {"status": "shipped", "eta": "2026-10-08"}
Order A-1029 has shipped (ETA 8 Oct). If you return it, refunds are processed within 5 working days.
Three production habits are visible here:
- Tool errors are returned to the model as the tool result — it can fix the argument and retry.
- Parallel tool calls — the model asked for two independent lookups in one turn. You can run them concurrently.
- A step limit — a confused model can loop; always cap it.
Replacing ScriptedLLM with a real call is one function: client.chat.completions.create(model=..., messages=messages, tools=TOOLS), then append response.choices[0].message.
6.5 Designing good tools¶
| Do | Why |
|---|---|
| Few, clearly different tools with descriptive names and descriptions | the model chooses from the description; overlapping tools confuse it |
| Say when to use the tool, not just what it does | "Use when the user asks where their order is" |
| Small, typed arguments with enums and limits | fewer invalid calls; validate them anyway |
| Return concise, relevant results (not a 5,000-row dump) | tool results consume context tokens |
| Return helpful error messages | "No order with id A-1092 — IDs look like A-1029" lets the model recover |
| Make dangerous tools ask for confirmation | refunds, emails, deletions: show the action, let a human approve |
| Give each agent only the tools it needs | least privilege limits the damage of a wrong or injected call |
Tool arguments are untrusted input
The model chose them — possibly under the influence of text it retrieved (prompt injection). Validate, authorise against the current user's permissions (not the model's say-so), and never pass arguments into SQL, shell commands or file paths unchecked.
6.6 Controlling tool use¶
| Setting (OpenAI names) | Effect | Use for |
|---|---|---|
tool_choice="auto" |
model decides | normal chat with optional tools |
tool_choice="required" |
must call some tool | agent steps that must act |
tool_choice={"type": "function", "function": {"name": "X"}} |
must call tool X | structured extraction: a "tool" whose arguments are your schema |
tool_choice="none" |
no tools | final answer step |
parallel_tool_calls=False |
one call per turn | tools that depend on each other's order |
6.7 MCP — tools as a shared standard¶
The Model Context Protocol (MCP) is an open standard for packaging tools (and resources and prompts) as a server that any compatible app or agent framework can connect to — GitHub, Slack, databases, your internal APIs. Instead of writing the same tool wrappers for every app, you run or install an MCP server once. Under the hood it's still the same thing: a name, a description, a JSON schema, and a function that returns a result.
Interview questions¶
How does function/tool calling work?
You send tool definitions (name, description, JSON schema of arguments) with the request. The model may return a structured tool call instead of text. Your code validates the arguments, runs the function, and appends the result as a tool message linked by the call id. The model then continues — possibly calling more tools — until it produces a final answer.
What can go wrong in a tool-calling agent and how do you guard against it?
Wrong tool or wrong arguments (clear descriptions, schemas, validation, helpful errors), infinite loops (step limits, budgets), dangerous actions (least privilege, human confirmation), prompt injection through tool results (treat as data, authorise by user identity), and context blow-up from large tool outputs (truncate and summarise results).
Structured output mode vs forcing a tool call — which do you use for extraction?
Both constrain the output to a schema. Use the provider's structured-output mode when available; forcing a single tool whose parameters are the schema works across providers. Validate the result either way.
Practice¶
- Add a
cancel_ordertool that returns"NEEDS_CONFIRMATION"instead of acting, and handle it inagent(). - Run the two tool calls of one turn concurrently with
concurrent.futures.ThreadPoolExecutor.
Next: Multimodal models — images, PDFs and audio.