Skip to content

Production RAG architecture

Intermediate · 20 min read · diagrams + worked examples

The RAG chatbot template is about 60 lines of Python. A RAG system serving thousands of people at a company needs more parts — and this page explains every part, why it exists, and how to choose between the options.

Before you read

You'll get the most out of this if you've built the RAG chatbot template or know what chunks, retrieval and prompts are. Every other term is explained as it appears.

What you'll learn

  • What changes between a demo and a production RAG system
  • The two paths every RAG system has — and each component on them
  • How hybrid search and re-ranking work (with a worked example and code)
  • How to size a system: chunks, storage and tokens
  • How to keep it secure, measurable and affordable

1. From demo to production — what changes?

Demo (the template) Production
Documents A few files in a folder Thousands to millions, from Drive, Confluence, websites, databases
Freshness Re-read everything at start-up Documents change daily — only re-process what changed
Users Just you Many people, each allowed to see different documents
Search Keyword search Keyword and meaning-based search, then a re-ranker
Quality "Looks right" Measured on a fixed test set before every release
Cost & speed Doesn't matter Budgeted per question; answers stream in under a few seconds
Failures Crash and restart Retries, fallbacks, alerts — and logs to explain bad answers

Everything on the rest of this page exists to solve one of the rows in this table.

2. The idea in plain English

Think of a company library with a research assistant:

  • Librarians keep the shelves up to date — new books arrive, old editions are removed. (the offline ingestion path)
  • The catalogue lets you find books by exact title and by topic. (keyword index + vector index)
  • When someone asks a question, the assistant pulls a pile of likely books, picks the best few pages, and writes an answer quoting those pages. (retrieval → re-ranking → generation with citations)
  • Some shelves are restricted — the assistant only uses books you are allowed to read. (permissions)

3. The big picture

Every RAG system has two paths that meet at the indexes.

Offline path — runs whenever documents change

flowchart LR
    S[Sources<br/>docs · web · DB] --> P[Parse & clean] --> C[Chunk + metadata]
    C --> E[Embed] --> V[(Vector index)]
    C --> K[(Keyword index<br/>BM25)]

Online path — runs for every question

flowchart LR
    U([User]) --> G[Gateway<br/>auth · rate limit] --> Q[Query rewrite] --> R[Hybrid retrieve]
    I[(Vector + BM25<br/>indexes)] -.-> R
    R --> RR[Re-rank] --> PB[Prompt builder] --> L[LLM] --> GR[Guardrails<br/>+ citations] --> A([Answer])

The next two sections walk through each path one step at a time.


4. Walkthrough: follow one document in

Imagine HR updates leave-policy.pdf in Google Drive. Here's what happens to it.

Step 1 — Detect the change. A connector checks the source on a schedule (or receives a webhook) and notices the file's modified time or content hash changed. Only this file is queued — nothing else is re-processed.

Step 2 — Parse and clean. The PDF becomes text. Headers, footers, page numbers and cookie banners are removed; headings and tables are kept, because they carry meaning.

Step 3 — Chunk. The text is split into pieces of roughly 200–500 words, following the document's headings where possible. Each chunk gets metadata:

{
  "chunk_id": "leave-policy.pdf#4",
  "text": "Employees get 24 days of paid leave per year...",
  "source": "drive://HR/leave-policy.pdf",
  "section": "Annual leave",
  "allowed_groups": ["all-employees"],
  "updated_at": "2026-09-30T10:15:00Z"
}

Step 4 — Embed. Each chunk's text is turned into an embedding — a list of numbers (for example 1,536 of them) that captures its meaning. Texts about similar topics get similar numbers.

Step 5 — Index. The embedding goes into the vector index; the text goes into the keyword (BM25) index. The old chunks of leave-policy.pdf are deleted from both — otherwise the bot would quote the outdated policy.

Run ingestion as a queue

Put one message per changed document on a queue (e.g. a database table, Redis, SQS) and let workers process them. A broken file then retries on its own instead of blocking everything else.

5. Walkthrough: follow one question out

An employee asks: "How many leave days can I carry over to next year?"

# Step What happens Typical time
1 Gateway Checks who the user is, applies rate limits, attaches their groups (e.g. all-employees, engineering) ~10 ms
2 Query rewrite (optional) If it's a follow-up ("what about interns?"), an LLM rewrites it into a full question using the chat history 200–500 ms
3 Hybrid retrieve Runs keyword and vector search in parallel, filtered to documents the user may see, then merges the two result lists 50–150 ms
4 Re-rank A re-ranking model reads the top ~30 candidates next to the question and keeps the best 4–8 100–300 ms
5 Prompt builder Puts instructions + numbered chunks + question together, within a token budget ~1 ms
6 LLM Generates the answer, streaming it word by word first words in 300–800 ms
7 Guardrails Checks citations point to real chunks, redacts personal data, refuses if the context didn't contain the answer ~10 ms

The user sees text appearing in under a second — streaming hides most of the remaining time. (Timings are typical ranges; measure your own.)


6. The components, one by one

For each part: what it does → why it's needed → your options → how to choose.

6.1 Connectors & sync

  • What: pull documents from where they live (Google Drive, SharePoint, Confluence, Notion, websites, databases).
  • Why: documents change constantly; the index must follow without a full rebuild.
  • Options: scheduled polling · webhooks from the source · nightly full sync as a safety net.
  • Choose: webhooks where available for speed, plus a periodic full comparison to catch anything missed — including deletions.

6.2 Parsing & cleaning

  • What: turn PDFs, Word files and HTML into clean text, keeping headings and tables.
  • Why: this caps your quality — if a table is scrambled here, no model can answer from it later.
  • Options: simple text extraction (pypdf, python-docx) · layout-aware parsers for complex PDFs · OCR for scans.
  • Choose: start simple; switch to layout-aware parsing for documents with tables and multi-column pages.

6.3 Chunking

  • What: split text into retrievable pieces, each with metadata (source, section, permissions, date).
  • Why: search works on chunks — too big and they mix topics; too small and they lose context.
  • Options: fixed size with overlap · split by headings (structure-aware) · split where the topic changes (semantic).
  • Choose: split by headings and prefix each chunk with its heading path (Leave policy > Carry-over), ~200–500 words with 10–20 % overlap. Then tune on real questions.

6.4 Embeddings & the vector index

  • What: store each chunk's meaning as numbers so you can find chunks by meaning, not exact words.
  • Why: "Can I take my unused holidays into January?" should find a chunk that says "carry-over of annual leave".
  • Options: hosted embedding APIs or open-source models · vector stores such as pgvector (Postgres), Pinecone, Qdrant, Weaviate, Elasticsearch/OpenSearch.
  • Choose: if you already run Postgres, pgvector is enough for up to a few million chunks. Use a managed vector database when you need more scale or less operations work.

6.5 The keyword index (BM25)

  • What: classic search on exact words — the same idea as in the template.
  • Why: embeddings are weak on exact terms: product codes (SKU-4471), names, error messages, acronyms.
  • Options: Elasticsearch/OpenSearch, Postgres full-text search, or an in-memory BM25 for small sets.

6.6 Hybrid retrieval — combining both searches

Keyword and vector search each return a ranked list, and their scores aren't comparable. The standard way to merge them is Reciprocal Rank Fusion (RRF): each document earns 1 / (60 + rank) from every list it appears in, and the totals are sorted. Being near the top of both lists wins.

Worked example

Document Keyword rank Vector rank RRF score
A 1 2 1/61 + 1/62 = 0.03252
C 3 1 1/63 + 1/61 = 0.03227
B 2 — 1/62 = 0.01613
D — 3 1/63 = 0.01587

Final order: A, C, B, D — A and C appear in both lists, so they beat documents found by only one search.

def reciprocal_rank_fusion(*ranked_lists, k=60):
    """Merge several ranked lists of IDs into one, rewarding IDs that rank well in many lists."""
    scores = {}
    for ranking in ranked_lists:
        for rank, doc_id in enumerate(ranking, start=1):          # ranks start at 1
            scores[doc_id] = scores.get(doc_id, 0.0) + 1 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)           # best total first

keyword_results = ["A", "B", "C"]
vector_results = ["C", "A", "D"]
print(reciprocal_rank_fusion(keyword_results, vector_results))   # → ['A', 'C', 'B', 'D']

6.7 Re-ranking

  • What: a second, more careful model (a cross-encoder) reads each candidate chunk together with the question and scores how well it answers it.
  • Why: the first search is fast but rough. Re-ranking the top 30–50 and keeping the best 4–8 sends the LLM fewer, better chunks — better answers and lower cost.
  • Choose: add it when the right chunk is usually somewhere in the top 30 but not reliably in the top 5. It adds 100–300 ms, so measure the gain.

6.8 Prompt builder & token budget

  • What: assembles instructions, numbered chunks, chat history and the question.
  • Why: models have a limited context window, and every token costs money and time.
  • How: set a budget and fill it in priority order. For example, with an 8,000-token budget:
Part Tokens
Instructions (system prompt) ~300
Recent chat history (summarised if long) up to 1,000
Retrieved chunks (best first, stop when full) up to 6,000
Question ~50
Room left for the answer the rest

Always wrap chunks in clear delimiters (e.g. <context>…</context>) and tell the model to treat them as data.

6.9 The LLM

  • Stream the answer so users see text immediately.
  • Use low temperature (0–0.3) for factual answers.
  • Keep a fallback model from another provider for outages, and set timeouts.
  • Use a smaller, cheaper model for helper steps such as query rewriting.

6.10 Guardrails & citations

  • Before generation: block obvious prompt-injection attempts and strip personal data you shouldn't send out.
  • After generation: check that every [n] citation points to a chunk that was actually retrieved; if the answer has no citations, either retry or say you don't know.
  • Always: treat retrieved text as untrusted — a document saying "ignore your instructions" must not change behaviour.

6.11 Caching

Cache Saves When it helps
Embeddings of chunks Re-embedding unchanged text Always — key by content hash
Embeddings of frequent questions One API call per question Many repeated questions
Full answers The whole pipeline FAQ-style traffic; expire when documents change

6.12 Observability & evaluation

  • Log every request: question, retrieved chunk IDs, prompt size, model, latency, tokens, answer, user feedback.
  • Build a golden set: 50–200 real questions with the expected source documents (and ideally reference answers).
  • Measure on every change (new chunking, new model, new prompt):
Metric Question it answers
Recall@k Was the right chunk among the top k retrieved?
MRR How high did the right chunk rank?
Faithfulness Is every claim in the answer supported by the retrieved chunks?
Answer relevance Does the answer actually address the question?
Latency & cost p50/p95 response time and tokens per question

7. Worked example: sizing a system

Scenario: a company with 50,000 documents (average 10 pages ≈ 5,000 words each) and 5,000 questions per day.

Quantity Calculation Result
Chunks 5,000 words ÷ ~400 words per chunk ≈ 13 per document × 50,000 ≈ 650,000 chunks
Vector storage 650,000 × 1,536 numbers × 4 bytes ≈ 4 GB (plus index overhead)
Tokens per question ~300 instructions + 6 chunks × ~500 + ~50 question + ~300 answer ≈ 3,650 tokens
Tokens per day 3,650 × 5,000 questions ≈ 18 million tokens/day

What this tells you: the index fits comfortably in pgvector or a managed vector store, and LLM tokens are the main running cost — so re-ranking (fewer chunks) and answer caching pay off quickly. Multiply the tokens by your model's price to get the daily cost.


8. Security & permissions

  • Filter by permission during retrieval, using the allowed_groups metadata on each chunk. Never retrieve everything and ask the LLM to "hide" what the user shouldn't see — it can't be trusted to.
  • Separate tenants (customers) with separate indexes or namespaces, plus filters.
  • Keep secrets out of prompts and redact personal data before logging.
  • Propagate deletions — when a document is removed or access is revoked, remove its chunks quickly.

9. Common mistakes

Mistake What happens Fix
Only vector search Misses exact codes, names and acronyms Add keyword search and merge with RRF
Chunks without headings Retrieved chunks lack context ("it costs ₹900" — what does?) Prefix each chunk with its heading path
Re-indexing everything nightly Slow, expensive, stale during the day Incremental sync by content hash
Old chunks not deleted Bot quotes outdated policies Delete a document's old chunks when it changes
No golden set Every change is a guess Measure recall@k and faithfulness before each release
Too many chunks in the prompt Higher cost, worse answers Re-rank and send only the best 4–8

10. Production checklist

  • Incremental ingestion with deletes, running on a queue
  • Structure-aware chunking with heading prefixes and metadata
  • Hybrid search (keyword + vector) with permission filters
  • Re-ranking measured against the golden set
  • Prompt with delimiters, a token budget and an "I don't know" rule
  • Streaming, timeouts and a fallback model
  • Citation checks and personal-data redaction
  • Logging of chunks, tokens, latency and feedback
  • Golden-set evaluation in CI before each release

Interview questions

Walk me through a production RAG architecture.

Two paths. Offline: connectors detect changes → parse and clean → structure-aware chunking with metadata → embed → write to vector and keyword indexes, deleting old chunks. Online: gateway (auth, rate limits, user groups) → optional query rewrite → hybrid retrieval with permission filters → RRF merge → re-rank → prompt with a token budget → streaming LLM → guardrails and citations. Around it: caching, logging, and golden-set evaluation.

Why hybrid search instead of only embeddings?

Embeddings match meaning but miss exact strings — product codes, names, error messages. BM25 matches those exactly. Running both and merging with reciprocal rank fusion gives the best of each, with no need to make their scores comparable.

How do you keep users from seeing documents they shouldn't?

Store permissions (groups, tenant) as metadata on each chunk and filter at retrieval time using the authenticated user's groups. Never rely on the prompt to hide data the model has already been given.

The bot gives a wrong answer. How do you debug it?

Check the logs for that request. If the right chunk wasn't retrieved, it's a retrieval problem — parsing, chunking, search or re-ranking. If it was retrieved but the answer is wrong, it's a generation problem — prompt instructions, too many noisy chunks, or the model. Then add the question to the golden set so it can't regress.

How would you reduce cost by 50 %?

Send fewer, better chunks (re-ranking), cache embeddings and repeated answers, use a smaller model for helper steps, summarise long chat histories, and check the logs for wasted tokens such as duplicate chunks.

Next: AI agent architecture →