7. Multimodal models¶
Intermediate · 12 min read
Modern LLMs aren't text-only. Most frontier models accept images and PDFs, many handle audio, and separate models generate images and speech. In business apps this unlocks a lot: reading invoices and ID cards, answering questions about screenshots and charts, voice assistants, and searching photos by description.
7.1 How images get into a transformer¶
A vision model cuts the image into small patches (e.g. 14×14 or 16×16 pixels), turns each patch into a vector with an image encoder, and feeds those vectors into the transformer alongside the text tokens. To the model, an image is just more tokens — which is why images cost tokens and fill the context window.
flowchart LR
I[Image] --> P[Split into patches]
P --> E[Vision encoder]
E --> V[Image tokens]
T[Text prompt] --> K[Text tokens]
V --> L[LLM]
K --> L
L --> O[Text answer]
7.2 Sending an image¶
Images are sent as message content parts — either a public URL or the bytes encoded in base64:
import base64, json
def image_part_openai(image_bytes: bytes, mime: str = "image/png", detail: str = "auto") -> dict:
b64 = base64.b64encode(image_bytes).decode()
return {"type": "image_url", "image_url": {"url": f"data:{mime};base64,{b64}", "detail": detail}}
def image_part_anthropic(image_bytes: bytes, mime: str = "image/png") -> dict:
b64 = base64.b64encode(image_bytes).decode()
return {"type": "image", "source": {"type": "base64", "media_type": mime, "data": b64}}
fake_png = b"\x89PNG\r\n\x1a\n" + b"\x00" * 24 # stand-in for open("invoice.png", "rb").read()
message = {"role": "user", "content": [
{"type": "text", "text": "What is the total amount on this invoice?"},
image_part_openai(fake_png),
]}
preview = json.loads(json.dumps(message)) # shorten the base64 for display
preview["content"][1]["image_url"]["url"] = preview["content"][1]["image_url"]["url"][:40] + "…"
print(json.dumps(preview, indent=1, ensure_ascii=False))
{
"role": "user",
"content": [
{
"type": "text",
"text": "What is the total amount on this invoice?"
},
{
"type": "image_url",
"image_url": {
"url": "data:image/png;base64,iVBORw0KGgoAAAAAAA…",
"detail": "auto"
}
}
]
}
Then call the API as usual: client.chat.completions.create(model=..., messages=[message]). Several images can go
in one message (e.g. "compare these two screenshots").
7.3 What an image costs¶
Each provider converts image size to tokens with its own formula. Two published examples:
import math
def anthropic_image_tokens(width: int, height: int) -> int:
"""Claude: roughly (width × height) / 750 tokens; large images are scaled down first."""
return math.ceil(width * height / 750)
def openai_tile_tokens(width: int, height: int, base: int = 85, per_tile: int = 170) -> int:
"""GPT-4o-style 'high detail': fit in 2048×2048, shortest side to 768, then count 512-px tiles."""
scale = min(1, 2048 / max(width, height)); w, h = width * scale, height * scale
scale = min(1, 768 / min(w, h)); w, h = w * scale, h * scale
return base + per_tile * math.ceil(w / 512) * math.ceil(h / 512)
for w, h in [(800, 600), (1568, 1568), (3000, 4000)]:
print(f"{w}×{h}: ~{anthropic_image_tokens(min(w, 1568), min(h, 1568)):,} (Claude-style) "
f"~{openai_tile_tokens(w, h):,} (tile-style)")
800×600: ~640 (Claude-style) ~765 (tile-style)
1568×1568: ~3,279 (Claude-style) ~765 (tile-style)
3000×4000: ~3,279 (Claude-style) ~765 (tile-style)
The formulas are illustrative (check current docs), but the lesson holds: an image is hundreds to a few thousand tokens. Resize images to what the task needs before sending — a receipt doesn't need a 12-megapixel photo — and use a low-detail mode for simple questions ("is this a cat?").
7.4 Documents: vision model, OCR, or both?¶
PDFs and scans are the most common multimodal input in business. Options:
| Approach | How | Good for | Watch out for |
|---|---|---|---|
| Text extraction (pypdf, pdfplumber) | read the PDF's text layer | digital PDFs; cheap, fast, exact | no text layer in scans; tables and layout get mangled |
| OCR (Tesseract, cloud OCR, document AI services) | image → text | scans at volume; cheap | handwriting, complex layouts; errors in numbers |
| Vision LLM on page images | send page images + a schema | messy layouts, tables, forms, charts, handwriting; understands context | cost per page; can misread digits; hallucination risk |
| Native PDF input (some APIs accept PDFs directly) | the provider does text + page images | quick prototypes | cost grows with page count |
Common production pattern: text extraction first, a vision LLM only for pages where that fails (scans, tables, charts), and validation of the extracted fields:
# no-run — vision extraction with a Pydantic schema (see the Pydantic notes)
from pydantic import BaseModel, Field
from datetime import date
class Invoice(BaseModel):
vendor: str
invoice_number: str
invoice_date: date
total_inr: float = Field(gt=0)
gstin: str | None = Field(None, pattern=r"^\d{2}[A-Z]{5}\d{4}[A-Z][1-9A-Z]Z[0-9A-Z]$")
completion = client.chat.completions.parse(
model="gpt-4o-mini", response_format=Invoice, temperature=0,
messages=[{"role": "user", "content": [
{"type": "text", "text": "Extract the invoice fields. Use null if a field is not visible."},
image_part_openai(open("invoice.png", "rb").read()),
]}],
)
invoice = completion.choices[0].message.parsed
Check the numbers
Vision models occasionally misread digits (8 vs 3, 1 vs 7) and can "fill in" a field that isn't visible. For
money and IDs: validate formats, cross-check totals (line items sum to the total; tax = rate × amount), allow
null, and send low-confidence documents to a human.
7.5 Audio and voice¶
Two architectures for voice assistants:
flowchart LR
subgraph Pipeline
M1[Mic] --> S[Speech-to-text<br/>Whisper etc.] --> L1[LLM] --> T[Text-to-speech] --> SP1[Speaker]
end
subgraph Realtime
M2[Mic] --> R[Speech-to-speech model<br/>realtime API] --> SP2[Speaker]
end
| STT → LLM → TTS pipeline | Realtime speech-to-speech model | |
|---|---|---|
| Latency | higher (three hops) — streaming each stage helps | lowest; natural interruptions |
| Control | full: log the transcript, use any LLM, RAG, tools, guardrails on text | less: fewer model choices |
| Cost | usually lower | usually higher |
| Good for | call summaries, voice notes, IVR with complex logic | live conversational agents, tutors, interview practice |
Speech-to-text alone is also hugely useful: transcribe meetings or calls, then summarise, extract action items or score them with an LLM.
# no-run — transcription then summarisation
from openai import OpenAI
client = OpenAI()
with open("sales_call.mp3", "rb") as f:
transcript = client.audio.transcriptions.create(model="whisper-1", file=f).text
summary = client.chat.completions.create(model="gpt-4o-mini", messages=[
{"role": "user", "content": f"Summarise this call and list action items with owners:\n\n{transcript}"},
]).choices[0].message.content
Indian-language support varies a lot between speech models — test with your real users' accents and code-mixed speech (Hinglish) before committing.
7.6 Multimodal embeddings: search images with text¶
Models like CLIP (and newer multimodal embedding models) map images and text into the same vector space, so a text query can find matching images — product search, photo libraries, finding a slide in a deck:
import numpy as np
# Stand-in vectors; a real CLIP model produces ~512–1024 dims for both images and text
image_index = {
"red_sneakers.jpg": np.array([0.9, 0.1, 0.0, 0.2]),
"blue_jacket.jpg": np.array([0.1, 0.9, 0.1, 0.0]),
"red_dress.jpg": np.array([0.8, 0.0, 0.6, 0.1]),
}
text_query = np.array([0.85, 0.05, 0.1, 0.3]) # embedding of "red running shoes"
def cos(a, b):
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
for name, vec in sorted(image_index.items(), key=lambda kv: -cos(text_query, kv[1])):
print(f"{cos(text_query, vec):.2f} {name}")
The same vector database and search techniques used for text RAG apply. An alternative for documents: have a vision LLM caption each image or chart, and index the captions as text.
7.7 Image generation¶
Text-to-image models (diffusion models such as Stable Diffusion, and the image models from OpenAI, Google and others) are different models from chat LLMs, though chat apps now call them as tools. Typical business uses: marketing variations, product mock-ups, thumbnails, illustrations for docs.
Practical points:
- Prompts work best as concrete descriptions: subject, setting, style, lighting, composition, aspect ratio.
- Editing (change the background, add an object, keep the product the same) is often more useful than generating from scratch.
- Check the provider's licensing and usage policy, avoid real people's likenesses and trademarks, and label AI-generated images where your platform or the law requires it.
Interview questions¶
How do vision-language models process images?
A vision encoder splits the image into patches and turns them into embedding vectors that are projected into the LLM's token space. The LLM then attends over image tokens and text tokens together. Images therefore consume tokens, and resolution drives cost.
How would you build invoice extraction from scanned PDFs?
Extract the text layer where it exists; for scans, render pages to images and use OCR or a vision LLM with a strict schema (structured output). Validate formats and arithmetic, allow nulls, route low-confidence documents to human review, and measure field-level accuracy on a labelled set.
Pipeline vs realtime model for a voice agent?
A speech-to-text → LLM → text-to-speech pipeline gives more control (any LLM, tools, RAG, logging, guardrails on text) and usually lower cost, but higher latency. A realtime speech-to-speech model gives the lowest latency and natural turn-taking, with less flexibility and typically higher cost.
Practice¶
- Photograph a printed receipt and extract vendor, date and total with a vision model; check the numbers.
- Compute the token cost of sending a 10-page scanned PDF as images at your provider's current rates.
Next: Context, memory & cost — keep long conversations and agents affordable.