Skip to content

7. Multimodal models

Intermediate · 12 min read

Modern LLMs aren't text-only. Most frontier models accept images and PDFs, many handle audio, and separate models generate images and speech. In business apps this unlocks a lot: reading invoices and ID cards, answering questions about screenshots and charts, voice assistants, and searching photos by description.

7.1 How images get into a transformer

A vision model cuts the image into small patches (e.g. 14×14 or 16×16 pixels), turns each patch into a vector with an image encoder, and feeds those vectors into the transformer alongside the text tokens. To the model, an image is just more tokens — which is why images cost tokens and fill the context window.

flowchart LR
    I[Image] --> P[Split into patches]
    P --> E[Vision encoder]
    E --> V[Image tokens]
    T[Text prompt] --> K[Text tokens]
    V --> L[LLM]
    K --> L
    L --> O[Text answer]

7.2 Sending an image

Images are sent as message content parts — either a public URL or the bytes encoded in base64:

import base64, json

def image_part_openai(image_bytes: bytes, mime: str = "image/png", detail: str = "auto") -> dict:
    b64 = base64.b64encode(image_bytes).decode()
    return {"type": "image_url", "image_url": {"url": f"data:{mime};base64,{b64}", "detail": detail}}

def image_part_anthropic(image_bytes: bytes, mime: str = "image/png") -> dict:
    b64 = base64.b64encode(image_bytes).decode()
    return {"type": "image", "source": {"type": "base64", "media_type": mime, "data": b64}}

fake_png = b"\x89PNG\r\n\x1a\n" + b"\x00" * 24          # stand-in for open("invoice.png", "rb").read()
message = {"role": "user", "content": [
    {"type": "text", "text": "What is the total amount on this invoice?"},
    image_part_openai(fake_png),
]}
preview = json.loads(json.dumps(message))                 # shorten the base64 for display
preview["content"][1]["image_url"]["url"] = preview["content"][1]["image_url"]["url"][:40] + "…"
print(json.dumps(preview, indent=1, ensure_ascii=False))
Output
{
 "role": "user",
 "content": [
  {
   "type": "text",
   "text": "What is the total amount on this invoice?"
  },
  {
   "type": "image_url",
   "image_url": {
    "url": "data:image/png;base64,iVBORw0KGgoAAAAAAA…",
    "detail": "auto"
   }
  }
 ]
}

Then call the API as usual: client.chat.completions.create(model=..., messages=[message]). Several images can go in one message (e.g. "compare these two screenshots").

7.3 What an image costs

Each provider converts image size to tokens with its own formula. Two published examples:

import math

def anthropic_image_tokens(width: int, height: int) -> int:
    """Claude: roughly (width × height) / 750 tokens; large images are scaled down first."""
    return math.ceil(width * height / 750)

def openai_tile_tokens(width: int, height: int, base: int = 85, per_tile: int = 170) -> int:
    """GPT-4o-style 'high detail': fit in 2048×2048, shortest side to 768, then count 512-px tiles."""
    scale = min(1, 2048 / max(width, height)); w, h = width * scale, height * scale
    scale = min(1, 768 / min(w, h)); w, h = w * scale, h * scale
    return base + per_tile * math.ceil(w / 512) * math.ceil(h / 512)

for w, h in [(800, 600), (1568, 1568), (3000, 4000)]:
    print(f"{w}×{h}: ~{anthropic_image_tokens(min(w, 1568), min(h, 1568)):,} (Claude-style)  "
          f"~{openai_tile_tokens(w, h):,} (tile-style)")
Output
800×600: ~640 (Claude-style)  ~765 (tile-style)
1568×1568: ~3,279 (Claude-style)  ~765 (tile-style)
3000×4000: ~3,279 (Claude-style)  ~765 (tile-style)

The formulas are illustrative (check current docs), but the lesson holds: an image is hundreds to a few thousand tokens. Resize images to what the task needs before sending — a receipt doesn't need a 12-megapixel photo — and use a low-detail mode for simple questions ("is this a cat?").

7.4 Documents: vision model, OCR, or both?

PDFs and scans are the most common multimodal input in business. Options:

Approach How Good for Watch out for
Text extraction (pypdf, pdfplumber) read the PDF's text layer digital PDFs; cheap, fast, exact no text layer in scans; tables and layout get mangled
OCR (Tesseract, cloud OCR, document AI services) image → text scans at volume; cheap handwriting, complex layouts; errors in numbers
Vision LLM on page images send page images + a schema messy layouts, tables, forms, charts, handwriting; understands context cost per page; can misread digits; hallucination risk
Native PDF input (some APIs accept PDFs directly) the provider does text + page images quick prototypes cost grows with page count

Common production pattern: text extraction first, a vision LLM only for pages where that fails (scans, tables, charts), and validation of the extracted fields:

# no-run — vision extraction with a Pydantic schema (see the Pydantic notes)
from pydantic import BaseModel, Field
from datetime import date

class Invoice(BaseModel):
    vendor: str
    invoice_number: str
    invoice_date: date
    total_inr: float = Field(gt=0)
    gstin: str | None = Field(None, pattern=r"^\d{2}[A-Z]{5}\d{4}[A-Z][1-9A-Z]Z[0-9A-Z]$")

completion = client.chat.completions.parse(
    model="gpt-4o-mini", response_format=Invoice, temperature=0,
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "Extract the invoice fields. Use null if a field is not visible."},
        image_part_openai(open("invoice.png", "rb").read()),
    ]}],
)
invoice = completion.choices[0].message.parsed

Check the numbers

Vision models occasionally misread digits (8 vs 3, 1 vs 7) and can "fill in" a field that isn't visible. For money and IDs: validate formats, cross-check totals (line items sum to the total; tax = rate × amount), allow null, and send low-confidence documents to a human.

7.5 Audio and voice

Two architectures for voice assistants:

flowchart LR
    subgraph Pipeline
      M1[Mic] --> S[Speech-to-text<br/>Whisper etc.] --> L1[LLM] --> T[Text-to-speech] --> SP1[Speaker]
    end
    subgraph Realtime
      M2[Mic] --> R[Speech-to-speech model<br/>realtime API] --> SP2[Speaker]
    end
STT → LLM → TTS pipeline Realtime speech-to-speech model
Latency higher (three hops) — streaming each stage helps lowest; natural interruptions
Control full: log the transcript, use any LLM, RAG, tools, guardrails on text less: fewer model choices
Cost usually lower usually higher
Good for call summaries, voice notes, IVR with complex logic live conversational agents, tutors, interview practice

Speech-to-text alone is also hugely useful: transcribe meetings or calls, then summarise, extract action items or score them with an LLM.

# no-run — transcription then summarisation
from openai import OpenAI
client = OpenAI()

with open("sales_call.mp3", "rb") as f:
    transcript = client.audio.transcriptions.create(model="whisper-1", file=f).text

summary = client.chat.completions.create(model="gpt-4o-mini", messages=[
    {"role": "user", "content": f"Summarise this call and list action items with owners:\n\n{transcript}"},
]).choices[0].message.content

Indian-language support varies a lot between speech models — test with your real users' accents and code-mixed speech (Hinglish) before committing.

7.6 Multimodal embeddings: search images with text

Models like CLIP (and newer multimodal embedding models) map images and text into the same vector space, so a text query can find matching images — product search, photo libraries, finding a slide in a deck:

import numpy as np

# Stand-in vectors; a real CLIP model produces ~512–1024 dims for both images and text
image_index = {
    "red_sneakers.jpg":  np.array([0.9, 0.1, 0.0, 0.2]),
    "blue_jacket.jpg":   np.array([0.1, 0.9, 0.1, 0.0]),
    "red_dress.jpg":     np.array([0.8, 0.0, 0.6, 0.1]),
}
text_query = np.array([0.85, 0.05, 0.1, 0.3])     # embedding of "red running shoes"

def cos(a, b):
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))

for name, vec in sorted(image_index.items(), key=lambda kv: -cos(text_query, kv[1])):
    print(f"{cos(text_query, vec):.2f}  {name}")
Output
0.99  red_sneakers.jpg
0.84  red_dress.jpg
0.17  blue_jacket.jpg

The same vector database and search techniques used for text RAG apply. An alternative for documents: have a vision LLM caption each image or chart, and index the captions as text.

7.7 Image generation

Text-to-image models (diffusion models such as Stable Diffusion, and the image models from OpenAI, Google and others) are different models from chat LLMs, though chat apps now call them as tools. Typical business uses: marketing variations, product mock-ups, thumbnails, illustrations for docs.

Practical points:

  • Prompts work best as concrete descriptions: subject, setting, style, lighting, composition, aspect ratio.
  • Editing (change the background, add an object, keep the product the same) is often more useful than generating from scratch.
  • Check the provider's licensing and usage policy, avoid real people's likenesses and trademarks, and label AI-generated images where your platform or the law requires it.

Interview questions

How do vision-language models process images?

A vision encoder splits the image into patches and turns them into embedding vectors that are projected into the LLM's token space. The LLM then attends over image tokens and text tokens together. Images therefore consume tokens, and resolution drives cost.

How would you build invoice extraction from scanned PDFs?

Extract the text layer where it exists; for scans, render pages to images and use OCR or a vision LLM with a strict schema (structured output). Validate formats and arithmetic, allow nulls, route low-confidence documents to human review, and measure field-level accuracy on a labelled set.

Pipeline vs realtime model for a voice agent?

A speech-to-text → LLM → text-to-speech pipeline gives more control (any LLM, tools, RAG, logging, guardrails on text) and usually lower cost, but higher latency. A realtime speech-to-speech model gives the lowest latency and natural turn-taking, with less flexibility and typically higher cost.

Practice

  • Photograph a printed receipt and extract vendor, date and total with a vision model; check the numbers.
  • Compute the token cost of sending a 10-page scanned PDF as images at your provider's current rates.

Next: Context, memory & cost — keep long conversations and agents affordable.