Skip to content

5. NLP for GenAI

Intermediate · 15 min read

Most "the RAG bot gave a wrong answer" bugs are not LLM bugs. They are chunking bugs (the answer was split across two chunks), retrieval bugs (the right chunk ranked 9th), or unmeasured changes. This page covers the three NLP skills that fix them.

5.1 Why chunking matters

You embed and retrieve chunks, not documents. The chunk size is a trade-off:

Chunks too small Chunks too large
answer split across chunks; missing context ("it", "this policy") one chunk mixes several topics → its embedding is a blurry average
more chunks to search and store fewer fit in the prompt; more tokens = more cost

A good starting point: 200–500 tokens with 10–20% overlap, split on natural boundaries — then measure (section 5.8) and tune.

5.2 Fixed-size token chunks with overlap

import tiktoken

enc = tiktoken.get_encoding("o200k_base")

def chunk_by_tokens(text: str, size: int = 40, overlap: int = 8) -> list[str]:
    ids = enc.encode(text)
    step = size - overlap
    return [enc.decode(ids[i:i + size]) for i in range(0, max(len(ids) - overlap, 1), step)]

policy = ("Refunds. Customers can request a refund within 30 days of delivery. Damaged items are refunded "
          "in full, including shipping. Refunds are processed within 5 working days of approval. "
          "Shipping. Orders ship within 48 hours. Express delivery is available in metro cities for an extra fee. "
          "International shipping takes 7 to 14 days.")

chunks = chunk_by_tokens(policy)
for c in chunks:
    print(f"{len(enc.encode(c)):2} tokens | {c!r}")
Output
40 tokens | 'Refunds. Customers can request a refund within 30 days of delivery. Damaged items are refunded in full, including shipping. Refunds are processed within 5 working days of approval. Shipping.'
37 tokens | '5 working days of approval. Shipping. Orders ship within 48 hours. Express delivery is available in metro cities for an extra fee. International shipping takes 7 to 14 days.'

Simple and predictable — but the second chunk starts mid-sentence ("5 working days of approval…"), and both chunks mix the refund and shipping topics. The overlap repeats a few tokens so a sentence cut at a boundary still appears whole in one of the chunks.

5.3 Recursive chunking — split on natural boundaries

Try the biggest separator first (paragraphs), and only fall back to smaller ones (lines, sentences, words) for pieces that are still too long. This is what LangChain's RecursiveCharacterTextSplitter does:

import re

def n_tokens(s: str) -> int:
    return len(enc.encode(s))

def recursive_split(text: str, max_tokens: int = 60, seps=("\n\n", "\n", ". ", " ")) -> list[str]:
    if n_tokens(text) <= max_tokens:
        return [text.strip()] if text.strip() else []
    sep, *rest = seps
    parts = text.split(sep)
    chunks, current = [], ""
    for part in parts:
        candidate = f"{current}{sep}{part}" if current else part
        if n_tokens(candidate) <= max_tokens:
            current = candidate
        else:
            if current:
                chunks.append(current.strip())
            current = part if n_tokens(part) <= max_tokens else ""
            if not current:                              # a single part is still too big → go smaller
                chunks.extend(recursive_split(part, max_tokens, tuple(rest) or (" ",)))
    if current.strip():
        chunks.append(current.strip())
    return chunks

doc = """# Refunds
Customers can request a refund within 30 days of delivery. Damaged items are refunded in full, including shipping.
Refunds are processed within 5 working days of approval.

# Shipping
Orders ship within 48 hours. Express delivery is available in metro cities for an extra fee.
International shipping takes 7 to 14 days."""

for c in recursive_split(doc, max_tokens=40):
    print(f"{n_tokens(c):2} tokens | {c!r}")
Output
39 tokens | '# Refunds\nCustomers can request a refund within 30 days of delivery. Damaged items are refunded in full, including shipping.\nRefunds are processed within 5 working days of approval.'
32 tokens | '# Shipping\nOrders ship within 48 hours. Express delivery is available in metro cities for an extra fee.\nInternational shipping takes 7 to 14 days.'

One topic per chunk, no broken sentences.

5.4 Structure-aware chunks with metadata

Documents have structure — headings, sections, pages. Keep it as metadata and prepend the heading to the chunk text, so a chunk that says "within 5 working days" still knows it's about refunds:

def chunk_markdown(md: str, source: str) -> list[dict]:
    chunks, heading = [], None
    for block in re.split(r"\n(?=# )", md):
        lines = block.strip().splitlines()
        if lines and lines[0].startswith("# "):
            heading, body = lines[0][2:], "\n".join(lines[1:])
        else:
            body = block
        for i, piece in enumerate(recursive_split(body, max_tokens=25)):
            chunks.append({
                "id": f"{source}#{heading}-{i}",
                "text": f"{heading}: {piece}",          # heading travels with the text that gets embedded
                "metadata": {"source": source, "section": heading},
            })
    return chunks

for c in chunk_markdown(doc, "policy.md"):
    print(c["id"], "|", c["text"][:70])
Output
policy.md#Refunds-0 | Refunds: Customers can request a refund within 30 days of delivery. Da
policy.md#Refunds-1 | Refunds: Refunds are processed within 5 working days of approval.
policy.md#Shipping-0 | Shipping: Orders ship within 48 hours. Express delivery is available i
policy.md#Shipping-1 | Shipping: International shipping takes 7 to 14 days.

Metadata also enables filters at query time (section == "Refunds", source in user_allowed_docs) — the standard way to enforce document permissions in RAG.

Strategy Use when
Fixed tokens + overlap unstructured text, quick baseline
Recursive general default for prose
Structure-aware (headings, pages, HTML/Markdown sections) docs, manuals, policies, wikis
Semantic (split where embedding similarity between sentences drops) long text without structure, e.g. transcripts
Parent–child (retrieve small chunks, send their bigger parent section to the LLM) precise retrieval and enough context

BM25 is TF-IDF's stronger successor and still the default keyword ranking in Elasticsearch, OpenSearch and most vector databases' "sparse" mode. It adds two fixes: repeating a word has diminishing returns (k1), and long documents are normalised so they don't win just by being long (b).

from rank_bm25 import BM25Okapi

corpus = [
    "Refunds are processed within 5 working days of approval.",
    "Damaged items are refunded in full, including shipping.",
    "Orders ship within 48 hours.",
    "Express delivery is available in metro cities.",
    "Error E-4012 means the payment gateway timed out.",
    "Reset your password from Settings > Security.",
]

def tokenize(s: str) -> list[str]:
    return re.findall(r"[a-z0-9-]+", s.lower())

bm25 = BM25Okapi([tokenize(d) for d in corpus])

def bm25_search(query: str, k: int = 3) -> list[int]:
    scores = bm25.get_scores(tokenize(query))
    return sorted(range(len(corpus)), key=lambda i: -scores[i])[:k]

for q in ["what does error E-4012 mean", "money back for broken products"]:
    print(q, "→", [corpus[i][:40] for i in bm25_search(q, 2)])
Output
what does error E-4012 mean → ['Error E-4012 means the payment gateway t', 'Refunds are processed within 5 working d']
money back for broken products → ['Refunds are processed within 5 working d', 'Damaged items are refunded in full, incl']

BM25 nails the exact error code — something embeddings often fumble. But for "money back for broken products" no word matches, so its ranking is arbitrary (all scores are 0) — that query needs embeddings.

5.6 Hybrid search with reciprocal rank fusion (RRF)

Run keyword and vector search, then merge the two rankings. Scores from BM25 and cosine similarity are on different scales, so instead of adding scores, RRF adds 1 / (k + rank) for each list:

def rrf(rankings: list[list[int]], k: int = 60) -> list[tuple[int, float]]:
    fused: dict[int, float] = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking, start=1):
            fused[doc_id] = fused.get(doc_id, 0.0) + 1 / (k + rank)
    return sorted(fused.items(), key=lambda x: -x[1])

query = "money back for damaged order"
keyword = bm25_search(query, k=4)
dense = [1, 0, 3, 2]          # ranking from an embedding model + vector DB (doc ids), for this query

for doc_id, score in rrf([keyword, dense])[:3]:
    print(f"{score:.4f}  {corpus[doc_id]}")
Output
0.0328  Damaged items are refunded in full, including shipping.
0.0323  Refunds are processed within 5 working days of approval.
0.0315  Orders ship within 48 hours.

A document ranked well by both methods rises to the top. Hybrid search is the single most reliable retrieval upgrade for production RAG; add a cross-encoder reranker on the top 20–50 results for the next jump in quality.

5.7 Classification metrics — precision, recall, F1

For any "pick a label" or "is this relevant?" task, count four outcomes for the class you care about:

y_true = ["refund", "refund", "refund", "other", "other", "refund", "other", "other"]
y_pred = ["refund", "other",  "refund", "refund", "other", "refund", "other", "other"]

tp = sum(t == p == "refund" for t, p in zip(y_true, y_pred))              # said refund, was refund
fp = sum(p == "refund" and t != "refund" for t, p in zip(y_true, y_pred))  # said refund, wasn't
fn = sum(t == "refund" and p != "refund" for t, p in zip(y_true, y_pred))  # missed a refund

precision = tp / (tp + fp)
recall = tp / (tp + fn)
f1 = 2 * precision * recall / (precision + recall)
print(f"TP={tp} FP={fp} FN={fn}")
print(f"precision={precision:.2f} recall={recall:.2f} f1={f1:.2f}")
Output
TP=3 FP=1 FN=1
precision=0.75 recall=0.75 f1=0.75
  • Precision — of everything flagged, how much was right? (cost of false alarms)
  • Recall — of everything that should be flagged, how much was found? (cost of misses)
  • F1 — the harmonic mean; high only if both are high.

Which matters more depends on the cost: a PII detector needs high recall (a miss leaks data); an auto-refund agent needs high precision (a false positive costs money).

5.8 Retrieval metrics — recall@k and MRR

To evaluate retrieval, build a small set of questions with the chunk(s) that contain the answer, then check where those chunks land in your results:

eval_set = [   # (question, ids of the relevant chunks)
    ("how long do refunds take", {0}),
    ("are damaged products refunded", {1}),
    ("what is error E-4012", {4}),
    ("how fast is shipping", {2, 3}),
    ("I forgot my password", {5}),
]

def recall_at_k(retrieved: list[int], relevant: set[int], k: int) -> float:
    return len(set(retrieved[:k]) & relevant) / len(relevant)

def reciprocal_rank(retrieved: list[int], relevant: set[int]) -> float:
    for rank, doc_id in enumerate(retrieved, start=1):
        if doc_id in relevant:
            return 1 / rank
    return 0.0

results = {q: bm25_search(q, k=len(corpus)) for q, _ in eval_set}
for k in (1, 3):
    r = sum(recall_at_k(results[q], rel, k) for q, rel in eval_set) / len(eval_set)
    print(f"recall@{k} = {r:.2f}")
mrr = sum(reciprocal_rank(results[q], rel) for q, rel in eval_set) / len(eval_set)
print(f"MRR      = {mrr:.2f}")
Output
recall@1 = 0.90
recall@3 = 0.90
MRR      = 1.00
  • recall@k — did the right chunk make it into the top k you send to the LLM? If it isn't retrieved, the LLM can't use it. This is the number to watch when you change chunking or embeddings.
  • MRR (mean reciprocal rank) — how high the first right chunk is on average (1.0 = always first).
  • nDCG — like MRR but rewards several relevant results in a good order; used in search benchmarks.

Build the eval set first

20–50 real questions with known answer chunks is enough to compare chunk sizes, BM25 vs embeddings vs hybrid, and rerankers — in minutes, without guessing. Re-run it on every change.

5.9 Generation metrics — exact match, token F1, ROUGE

Comparing a generated answer with a reference answer:

from collections import Counter

def normalise(s: str) -> list[str]:
    s = re.sub(r"\b(a|an|the)\b", " ", s.lower())
    return re.findall(r"[a-z0-9]+", s)

def exact_match(pred: str, ref: str) -> float:
    return float(normalise(pred) == normalise(ref))

def token_f1(pred: str, ref: str) -> float:          # the SQuAD metric
    p, r = normalise(pred), normalise(ref)
    common = sum((Counter(p) & Counter(r)).values())
    if common == 0:
        return 0.0
    precision, recall = common / len(p), common / len(r)
    return 2 * precision * recall / (precision + recall)

def rouge_l(pred: str, ref: str) -> float:           # longest common subsequence, F-measure
    p, r = normalise(pred), normalise(ref)
    dp = [[0] * (len(r) + 1) for _ in range(len(p) + 1)]
    for i in range(len(p)):
        for j in range(len(r)):
            dp[i + 1][j + 1] = dp[i][j] + 1 if p[i] == r[j] else max(dp[i][j + 1], dp[i + 1][j])
    lcs = dp[-1][-1]
    return 0.0 if lcs == 0 else 2 * lcs / (len(p) + len(r))

ref = "Refunds are processed within 5 working days."
for pred in ["Refunds are processed within 5 working days.",
             "Your refund will be processed in 5 working days.",
             "You'll get your money back in under a week.",
             "Refunds are processed within 30 working days."]:
    print(f"EM={exact_match(pred, ref):.0f}  F1={token_f1(pred, ref):.2f}  ROUGE-L={rouge_l(pred, ref):.2f}  {pred}")
Output
EM=1  F1=1.00  ROUGE-L=1.00  Refunds are processed within 5 working days.
EM=0  F1=0.50  ROUGE-L=0.50  Your refund will be processed in 5 working days.
EM=0  F1=0.00  ROUGE-L=0.00  You'll get your money back in under a week.
EM=0  F1=0.86  ROUGE-L=0.86  Refunds are processed within 30 working days.

Read the last two rows carefully: a correct paraphrase scores 0, and a wrong answer (30 days instead of 5) scores 0.86. Word-overlap metrics measure wording, not truth.

Metric Good for Blind spot
Exact match short factual answers (names, numbers, labels) any rewording
Token F1 / ROUGE / BLEU summaries and extractive QA at scale; regression checks paraphrases; small factual errors
Embedding similarity paraphrase-tolerant comparison still misses "5" vs "30"
LLM-as-judge correctness, faithfulness to context, helpfulness cost; needs its own validation

For RAG, the standard set (as in RAGAS / DeepEval) is: context recall (was the needed info retrieved?), faithfulness (is every claim supported by the retrieved context?), and answer relevance — mostly computed with an LLM judge, with cheap metrics above as fast regression checks.

Interview questions

How do you choose a chunk size for RAG?

Start around 200–500 tokens with 10–20% overlap, split on structure (headings, paragraphs), and attach metadata (source, section, page). Then build a small eval set and measure recall@k and answer quality for a few sizes. Smaller chunks → precise retrieval but less context; larger → more context but blurrier embeddings and higher cost. Parent–child retrieval gets both.

Why use hybrid search instead of only vector search?

Embeddings capture meaning and synonyms but are weak on exact tokens — error codes, SKUs, names, rare terms. BM25 is strong on exactly those. Fusing the rankings (RRF) gets the best of both and is robust because it needs no score calibration.

Your RAG bot gives wrong answers. How do you debug it?

Separate retrieval from generation. Check recall@k on an eval set: if the right chunk isn't retrieved, fix chunking, add hybrid search or a reranker, add metadata filters. If it is retrieved but the answer is wrong, look at the prompt, context order and length, and measure faithfulness. Change one thing at a time and re-run the eval.

Why is ROUGE a poor metric for LLM answers?

It counts overlapping words. Correct paraphrases score low and answers with a single wrong fact (a number, a negation) score high. Use it for cheap regression checks; judge correctness and faithfulness with human review or a validated LLM judge.

Practice

  • Run recursive_split on one of your own documents at 200, 400 and 800 tokens and compare the number of chunks.
  • Add 10 questions to eval_set and compare recall@3 for BM25 against a TF-IDF search from Text representation.

Next: Transformers — the architecture inside every LLM and embedding model.