RAG interview questions¶
Click a question to reveal the model answer. Try answering out loud first.
1. What is RAG and why would you use it instead of just prompting an LLM?
RAG retrieves relevant passages from your own data at question time and puts them in the prompt, so the model answers grounded in current, private information and can cite sources. Plain prompting only has the model's training data — which is frozen and doesn't include your documents — and it hallucinates more when it doesn't know.
Follow-up: When is RAG not enough? — When you need the model to change behaviour or format consistently (fine-tuning), or when answers need reasoning across the whole corpus rather than a few passages.
2. Walk me through a RAG pipeline end to end.
Indexing: load → clean → chunk (with overlap and metadata) → index (embeddings in a vector store and/or BM25). Query time: (optionally rewrite the query) → retrieve top-k → (optionally re-rank) → build the prompt with instructions + chunks + question → generate → return the answer with sources.
Follow-up: Where would you add caching? — Embeddings of documents (computed once), embeddings of frequent queries, and full answers for repeated questions.
3. Dense vs sparse retrieval — when do you use each?
Sparse (BM25) scores exact keyword overlap — great for names, codes, error messages, rare terms. Dense (embeddings) matches meaning — great for paraphrases and natural questions. Hybrid combines both (e.g. reciprocal rank fusion) and is a strong default for production.
4. How do you choose chunk size and overlap?
Balance context vs precision: too small loses meaning, too large mixes topics and wastes prompt space. Start around 300–500 tokens with 10–20 % overlap, split on document structure where possible, then measure recall@k on a labelled question set and tune.
5. How do you evaluate a RAG system?
Evaluate the two stages separately.
- Retrieval: recall@k / hit rate (is the right chunk in the top-k?), MRR (how high it ranks).
- Generation: faithfulness (is every claim supported by the retrieved context?), answer relevance (does it address the question?), and correctness against reference answers.
Use a fixed test set of real questions, run it on every change, and add LLM-as-judge scoring with spot-checked human review.
6. The answer is wrong. How do you debug it?
First check retrieval: was the right chunk retrieved at all? If not — fix chunking, add hybrid search, increase k, or add a re-ranker. If it was retrieved, it's a generation problem — tighten the instructions, reduce noisy chunks, put the best chunk first, or lower temperature. Logging the retrieved chunks for every answer makes this quick.
7. What is re-ranking and when is it worth it?
A second, more accurate model (usually a cross-encoder) scores each (question, chunk) pair from a larger first-stage candidate set (say top-50) and keeps the best few. It improves precision noticeably but adds latency and cost, so use it when first-stage retrieval returns the right chunk somewhere in the top-50 but not reliably in the top-5.
8. How do you stop prompt injection through retrieved documents?
Treat retrieved text as untrusted data: wrap it in clear delimiters, tell the model never to follow instructions inside it, keep tools and actions behind server-side permission checks, validate outputs (e.g. allow-listed actions only), and never put secrets in the prompt.
Practise live: AI Mock Interview · Go deeper: Production RAG architecture