LLMs¶
Beginner → Intermediate · 9 topics
Transformers showed what happens inside a model. This series is about using one well: what LLMs can and can't do, picking the right model for the job, calling it reliably, writing prompts that work, getting structured data and tool calls back, keeping context and cost under control, and proving the app actually works. It's the day-to-day skill set of a GenAI engineer — and the foundation for RAG and agents.
pip install openai anthropic pydantic tiktoken # the examples run without API keys, using a fake LLM
| # | Topic | Sub-topics |
|---|---|---|
| 1 | What is an LLM | Next-token prediction · in-context learning · strengths & limits · hallucination · knowledge cutoff · scaling laws |
| 2 | Choosing a model | Closed vs open models · model tiers · context windows · pricing · benchmarks · a selection checklist |
| 3 | Calling LLM APIs | Messages & roles · key parameters · streaming · usage · a provider-agnostic client · parallel calls |
| 4 | Prompt engineering | Anatomy of a prompt · system prompts · few-shot · delimiters · reasoning · prompt templates |
| 5 | Reasoning models | Thinking tokens · effort & budgets · cost and latency · prompting them · self-consistency |
| 6 | Structured output & tool calling | JSON schemas · the tool-calling protocol · the tool loop · parallel tools · MCP |
| 7 | Multimodal models | Images & PDFs · image token cost · vision vs OCR · voice pipelines · image search · image generation |
| 8 | Context, memory & cost | Token budgets · conversation memory · prompt caching · response caching · model routing |
| 9 | Evaluation & guardrails | Eval sets · LLM-as-judge · prompt injection · guardrails · observability · responsible AI & privacy |
flowchart LR
U[User input] --> G1[Input guardrails]
G1 --> P[Prompt template + context + memory]
P --> M{Model router}
M --> L[LLM API]
L --> T{Tool call?}
T -->|yes| X[Run tool] --> L
T -->|no| V[Validate output]
V --> G2[Output guardrails] --> R[Response]
L -.-> O[(Logs, cost, evals)]
Quick start first?
If you've never called an LLM from code, start with the tutorial Call an LLM from Python (10 minutes), then come back.
Back to: Notes overview