Skip to content

LLMs

Beginner → Intermediate · 9 topics

Transformers showed what happens inside a model. This series is about using one well: what LLMs can and can't do, picking the right model for the job, calling it reliably, writing prompts that work, getting structured data and tool calls back, keeping context and cost under control, and proving the app actually works. It's the day-to-day skill set of a GenAI engineer — and the foundation for RAG and agents.

pip install openai anthropic pydantic tiktoken     # the examples run without API keys, using a fake LLM
# Topic Sub-topics
1 What is an LLM Next-token prediction · in-context learning · strengths & limits · hallucination · knowledge cutoff · scaling laws
2 Choosing a model Closed vs open models · model tiers · context windows · pricing · benchmarks · a selection checklist
3 Calling LLM APIs Messages & roles · key parameters · streaming · usage · a provider-agnostic client · parallel calls
4 Prompt engineering Anatomy of a prompt · system prompts · few-shot · delimiters · reasoning · prompt templates
5 Reasoning models Thinking tokens · effort & budgets · cost and latency · prompting them · self-consistency
6 Structured output & tool calling JSON schemas · the tool-calling protocol · the tool loop · parallel tools · MCP
7 Multimodal models Images & PDFs · image token cost · vision vs OCR · voice pipelines · image search · image generation
8 Context, memory & cost Token budgets · conversation memory · prompt caching · response caching · model routing
9 Evaluation & guardrails Eval sets · LLM-as-judge · prompt injection · guardrails · observability · responsible AI & privacy
flowchart LR
    U[User input] --> G1[Input guardrails]
    G1 --> P[Prompt template + context + memory]
    P --> M{Model router}
    M --> L[LLM API]
    L --> T{Tool call?}
    T -->|yes| X[Run tool] --> L
    T -->|no| V[Validate output]
    V --> G2[Output guardrails] --> R[Response]
    L -.-> O[(Logs, cost, evals)]

Quick start first?

If you've never called an LLM from code, start with the tutorial Call an LLM from Python (10 minutes), then come back.

Back to: Notes overview