5. NLP for GenAI¶
Intermediate · 15 min read
Most "the RAG bot gave a wrong answer" bugs are not LLM bugs. They are chunking bugs (the answer was split across two chunks), retrieval bugs (the right chunk ranked 9th), or unmeasured changes. This page covers the three NLP skills that fix them.
5.1 Why chunking matters¶
You embed and retrieve chunks, not documents. The chunk size is a trade-off:
| Chunks too small | Chunks too large |
|---|---|
| answer split across chunks; missing context ("it", "this policy") | one chunk mixes several topics → its embedding is a blurry average |
| more chunks to search and store | fewer fit in the prompt; more tokens = more cost |
A good starting point: 200–500 tokens with 10–20% overlap, split on natural boundaries — then measure (section 5.8) and tune.
5.2 Fixed-size token chunks with overlap¶
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
def chunk_by_tokens(text: str, size: int = 40, overlap: int = 8) -> list[str]:
ids = enc.encode(text)
step = size - overlap
return [enc.decode(ids[i:i + size]) for i in range(0, max(len(ids) - overlap, 1), step)]
policy = ("Refunds. Customers can request a refund within 30 days of delivery. Damaged items are refunded "
"in full, including shipping. Refunds are processed within 5 working days of approval. "
"Shipping. Orders ship within 48 hours. Express delivery is available in metro cities for an extra fee. "
"International shipping takes 7 to 14 days.")
chunks = chunk_by_tokens(policy)
for c in chunks:
print(f"{len(enc.encode(c)):2} tokens | {c!r}")
40 tokens | 'Refunds. Customers can request a refund within 30 days of delivery. Damaged items are refunded in full, including shipping. Refunds are processed within 5 working days of approval. Shipping.'
37 tokens | '5 working days of approval. Shipping. Orders ship within 48 hours. Express delivery is available in metro cities for an extra fee. International shipping takes 7 to 14 days.'
Simple and predictable — but the second chunk starts mid-sentence ("5 working days of approval…"), and both chunks mix the refund and shipping topics. The overlap repeats a few tokens so a sentence cut at a boundary still appears whole in one of the chunks.
5.3 Recursive chunking — split on natural boundaries¶
Try the biggest separator first (paragraphs), and only fall back to smaller ones (lines, sentences, words)
for pieces that are still too long. This is what LangChain's RecursiveCharacterTextSplitter does:
import re
def n_tokens(s: str) -> int:
return len(enc.encode(s))
def recursive_split(text: str, max_tokens: int = 60, seps=("\n\n", "\n", ". ", " ")) -> list[str]:
if n_tokens(text) <= max_tokens:
return [text.strip()] if text.strip() else []
sep, *rest = seps
parts = text.split(sep)
chunks, current = [], ""
for part in parts:
candidate = f"{current}{sep}{part}" if current else part
if n_tokens(candidate) <= max_tokens:
current = candidate
else:
if current:
chunks.append(current.strip())
current = part if n_tokens(part) <= max_tokens else ""
if not current: # a single part is still too big → go smaller
chunks.extend(recursive_split(part, max_tokens, tuple(rest) or (" ",)))
if current.strip():
chunks.append(current.strip())
return chunks
doc = """# Refunds
Customers can request a refund within 30 days of delivery. Damaged items are refunded in full, including shipping.
Refunds are processed within 5 working days of approval.
# Shipping
Orders ship within 48 hours. Express delivery is available in metro cities for an extra fee.
International shipping takes 7 to 14 days."""
for c in recursive_split(doc, max_tokens=40):
print(f"{n_tokens(c):2} tokens | {c!r}")
39 tokens | '# Refunds\nCustomers can request a refund within 30 days of delivery. Damaged items are refunded in full, including shipping.\nRefunds are processed within 5 working days of approval.'
32 tokens | '# Shipping\nOrders ship within 48 hours. Express delivery is available in metro cities for an extra fee.\nInternational shipping takes 7 to 14 days.'
One topic per chunk, no broken sentences.
5.4 Structure-aware chunks with metadata¶
Documents have structure — headings, sections, pages. Keep it as metadata and prepend the heading to the chunk text, so a chunk that says "within 5 working days" still knows it's about refunds:
def chunk_markdown(md: str, source: str) -> list[dict]:
chunks, heading = [], None
for block in re.split(r"\n(?=# )", md):
lines = block.strip().splitlines()
if lines and lines[0].startswith("# "):
heading, body = lines[0][2:], "\n".join(lines[1:])
else:
body = block
for i, piece in enumerate(recursive_split(body, max_tokens=25)):
chunks.append({
"id": f"{source}#{heading}-{i}",
"text": f"{heading}: {piece}", # heading travels with the text that gets embedded
"metadata": {"source": source, "section": heading},
})
return chunks
for c in chunk_markdown(doc, "policy.md"):
print(c["id"], "|", c["text"][:70])
policy.md#Refunds-0 | Refunds: Customers can request a refund within 30 days of delivery. Da
policy.md#Refunds-1 | Refunds: Refunds are processed within 5 working days of approval.
policy.md#Shipping-0 | Shipping: Orders ship within 48 hours. Express delivery is available i
policy.md#Shipping-1 | Shipping: International shipping takes 7 to 14 days.
Metadata also enables filters at query time (section == "Refunds", source in user_allowed_docs) — the
standard way to enforce document permissions in RAG.
| Strategy | Use when |
|---|---|
| Fixed tokens + overlap | unstructured text, quick baseline |
| Recursive | general default for prose |
| Structure-aware (headings, pages, HTML/Markdown sections) | docs, manuals, policies, wikis |
| Semantic (split where embedding similarity between sentences drops) | long text without structure, e.g. transcripts |
| Parent–child (retrieve small chunks, send their bigger parent section to the LLM) | precise retrieval and enough context |
5.5 BM25 — the keyword half of search¶
BM25 is TF-IDF's stronger successor and still the default keyword ranking in Elasticsearch,
OpenSearch and most vector databases' "sparse" mode. It adds two fixes: repeating a word has
diminishing returns (k1), and long documents are normalised so they don't win just by being long (b).
from rank_bm25 import BM25Okapi
corpus = [
"Refunds are processed within 5 working days of approval.",
"Damaged items are refunded in full, including shipping.",
"Orders ship within 48 hours.",
"Express delivery is available in metro cities.",
"Error E-4012 means the payment gateway timed out.",
"Reset your password from Settings > Security.",
]
def tokenize(s: str) -> list[str]:
return re.findall(r"[a-z0-9-]+", s.lower())
bm25 = BM25Okapi([tokenize(d) for d in corpus])
def bm25_search(query: str, k: int = 3) -> list[int]:
scores = bm25.get_scores(tokenize(query))
return sorted(range(len(corpus)), key=lambda i: -scores[i])[:k]
for q in ["what does error E-4012 mean", "money back for broken products"]:
print(q, "→", [corpus[i][:40] for i in bm25_search(q, 2)])
what does error E-4012 mean → ['Error E-4012 means the payment gateway t', 'Refunds are processed within 5 working d']
money back for broken products → ['Refunds are processed within 5 working d', 'Damaged items are refunded in full, incl']
BM25 nails the exact error code — something embeddings often fumble. But for "money back for broken products" no word matches, so its ranking is arbitrary (all scores are 0) — that query needs embeddings.
5.6 Hybrid search with reciprocal rank fusion (RRF)¶
Run keyword and vector search, then merge the two rankings. Scores from BM25 and cosine similarity
are on different scales, so instead of adding scores, RRF adds 1 / (k + rank) for each list:
def rrf(rankings: list[list[int]], k: int = 60) -> list[tuple[int, float]]:
fused: dict[int, float] = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
fused[doc_id] = fused.get(doc_id, 0.0) + 1 / (k + rank)
return sorted(fused.items(), key=lambda x: -x[1])
query = "money back for damaged order"
keyword = bm25_search(query, k=4)
dense = [1, 0, 3, 2] # ranking from an embedding model + vector DB (doc ids), for this query
for doc_id, score in rrf([keyword, dense])[:3]:
print(f"{score:.4f} {corpus[doc_id]}")
0.0328 Damaged items are refunded in full, including shipping.
0.0323 Refunds are processed within 5 working days of approval.
0.0315 Orders ship within 48 hours.
A document ranked well by both methods rises to the top. Hybrid search is the single most reliable retrieval upgrade for production RAG; add a cross-encoder reranker on the top 20–50 results for the next jump in quality.
5.7 Classification metrics — precision, recall, F1¶
For any "pick a label" or "is this relevant?" task, count four outcomes for the class you care about:
y_true = ["refund", "refund", "refund", "other", "other", "refund", "other", "other"]
y_pred = ["refund", "other", "refund", "refund", "other", "refund", "other", "other"]
tp = sum(t == p == "refund" for t, p in zip(y_true, y_pred)) # said refund, was refund
fp = sum(p == "refund" and t != "refund" for t, p in zip(y_true, y_pred)) # said refund, wasn't
fn = sum(t == "refund" and p != "refund" for t, p in zip(y_true, y_pred)) # missed a refund
precision = tp / (tp + fp)
recall = tp / (tp + fn)
f1 = 2 * precision * recall / (precision + recall)
print(f"TP={tp} FP={fp} FN={fn}")
print(f"precision={precision:.2f} recall={recall:.2f} f1={f1:.2f}")
- Precision — of everything flagged, how much was right? (cost of false alarms)
- Recall — of everything that should be flagged, how much was found? (cost of misses)
- F1 — the harmonic mean; high only if both are high.
Which matters more depends on the cost: a PII detector needs high recall (a miss leaks data); an auto-refund agent needs high precision (a false positive costs money).
5.8 Retrieval metrics — recall@k and MRR¶
To evaluate retrieval, build a small set of questions with the chunk(s) that contain the answer, then check where those chunks land in your results:
eval_set = [ # (question, ids of the relevant chunks)
("how long do refunds take", {0}),
("are damaged products refunded", {1}),
("what is error E-4012", {4}),
("how fast is shipping", {2, 3}),
("I forgot my password", {5}),
]
def recall_at_k(retrieved: list[int], relevant: set[int], k: int) -> float:
return len(set(retrieved[:k]) & relevant) / len(relevant)
def reciprocal_rank(retrieved: list[int], relevant: set[int]) -> float:
for rank, doc_id in enumerate(retrieved, start=1):
if doc_id in relevant:
return 1 / rank
return 0.0
results = {q: bm25_search(q, k=len(corpus)) for q, _ in eval_set}
for k in (1, 3):
r = sum(recall_at_k(results[q], rel, k) for q, rel in eval_set) / len(eval_set)
print(f"recall@{k} = {r:.2f}")
mrr = sum(reciprocal_rank(results[q], rel) for q, rel in eval_set) / len(eval_set)
print(f"MRR = {mrr:.2f}")
- recall@k — did the right chunk make it into the top k you send to the LLM? If it isn't retrieved, the LLM can't use it. This is the number to watch when you change chunking or embeddings.
- MRR (mean reciprocal rank) — how high the first right chunk is on average (1.0 = always first).
- nDCG — like MRR but rewards several relevant results in a good order; used in search benchmarks.
Build the eval set first
20–50 real questions with known answer chunks is enough to compare chunk sizes, BM25 vs embeddings vs hybrid, and rerankers — in minutes, without guessing. Re-run it on every change.
5.9 Generation metrics — exact match, token F1, ROUGE¶
Comparing a generated answer with a reference answer:
from collections import Counter
def normalise(s: str) -> list[str]:
s = re.sub(r"\b(a|an|the)\b", " ", s.lower())
return re.findall(r"[a-z0-9]+", s)
def exact_match(pred: str, ref: str) -> float:
return float(normalise(pred) == normalise(ref))
def token_f1(pred: str, ref: str) -> float: # the SQuAD metric
p, r = normalise(pred), normalise(ref)
common = sum((Counter(p) & Counter(r)).values())
if common == 0:
return 0.0
precision, recall = common / len(p), common / len(r)
return 2 * precision * recall / (precision + recall)
def rouge_l(pred: str, ref: str) -> float: # longest common subsequence, F-measure
p, r = normalise(pred), normalise(ref)
dp = [[0] * (len(r) + 1) for _ in range(len(p) + 1)]
for i in range(len(p)):
for j in range(len(r)):
dp[i + 1][j + 1] = dp[i][j] + 1 if p[i] == r[j] else max(dp[i][j + 1], dp[i + 1][j])
lcs = dp[-1][-1]
return 0.0 if lcs == 0 else 2 * lcs / (len(p) + len(r))
ref = "Refunds are processed within 5 working days."
for pred in ["Refunds are processed within 5 working days.",
"Your refund will be processed in 5 working days.",
"You'll get your money back in under a week.",
"Refunds are processed within 30 working days."]:
print(f"EM={exact_match(pred, ref):.0f} F1={token_f1(pred, ref):.2f} ROUGE-L={rouge_l(pred, ref):.2f} {pred}")
EM=1 F1=1.00 ROUGE-L=1.00 Refunds are processed within 5 working days.
EM=0 F1=0.50 ROUGE-L=0.50 Your refund will be processed in 5 working days.
EM=0 F1=0.00 ROUGE-L=0.00 You'll get your money back in under a week.
EM=0 F1=0.86 ROUGE-L=0.86 Refunds are processed within 30 working days.
Read the last two rows carefully: a correct paraphrase scores 0, and a wrong answer (30 days instead of 5) scores 0.86. Word-overlap metrics measure wording, not truth.
| Metric | Good for | Blind spot |
|---|---|---|
| Exact match | short factual answers (names, numbers, labels) | any rewording |
| Token F1 / ROUGE / BLEU | summaries and extractive QA at scale; regression checks | paraphrases; small factual errors |
| Embedding similarity | paraphrase-tolerant comparison | still misses "5" vs "30" |
| LLM-as-judge | correctness, faithfulness to context, helpfulness | cost; needs its own validation |
For RAG, the standard set (as in RAGAS / DeepEval) is: context recall (was the needed info retrieved?), faithfulness (is every claim supported by the retrieved context?), and answer relevance — mostly computed with an LLM judge, with cheap metrics above as fast regression checks.
Interview questions¶
How do you choose a chunk size for RAG?
Start around 200–500 tokens with 10–20% overlap, split on structure (headings, paragraphs), and attach metadata (source, section, page). Then build a small eval set and measure recall@k and answer quality for a few sizes. Smaller chunks → precise retrieval but less context; larger → more context but blurrier embeddings and higher cost. Parent–child retrieval gets both.
Why use hybrid search instead of only vector search?
Embeddings capture meaning and synonyms but are weak on exact tokens — error codes, SKUs, names, rare terms. BM25 is strong on exactly those. Fusing the rankings (RRF) gets the best of both and is robust because it needs no score calibration.
Your RAG bot gives wrong answers. How do you debug it?
Separate retrieval from generation. Check recall@k on an eval set: if the right chunk isn't retrieved, fix chunking, add hybrid search or a reranker, add metadata filters. If it is retrieved but the answer is wrong, look at the prompt, context order and length, and measure faithfulness. Change one thing at a time and re-run the eval.
Why is ROUGE a poor metric for LLM answers?
It counts overlapping words. Correct paraphrases score low and answers with a single wrong fact (a number, a negation) score high. Use it for cheap regression checks; judge correctness and faithfulness with human review or a validated LLM judge.
Practice¶
- Run
recursive_spliton one of your own documents at 200, 400 and 800 tokens and compare the number of chunks. - Add 10 questions to
eval_setand compare recall@3 for BM25 against a TF-IDF search from Text representation.
Next: Transformers — the architecture inside every LLM and embedding model.