Skip to content

1. Text preprocessing

Beginner · 10 min read

Garbage in, garbage out — and in RAG, garbage in means wrong chunks retrieved. Preprocessing is the unglamorous step that decides whether search works. The rule for the LLM era:

Clean the noise, keep the meaning. Remove what no reader needs (HTML tags, repeated headers, broken whitespace). Keep what an LLM uses (casing, punctuation, stopwords, numbers).

1.1 Unicode normalisation

Text from PDFs and web pages contains look-alike characters that break exact matching and keyword search:

import unicodedata

raw = "Refund file — “30 days” only"     # full-width letters, 'fi' ligature, smart quotes, non-breaking space
nfkc = unicodedata.normalize("NFKC", raw)
print(nfkc)
print("Refund" in raw, "Refund" in nfkc)
Output
Refund file — “30 days” only
False True

NFKC turns compatibility characters (full-width, ligatures, non-breaking spaces) into their plain forms. Smart quotes and dashes are kept — they are real characters. Replace them only if your search needs it:

SMART = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'", "—": "-", "–": "-"})
print(nfkc.translate(SMART))
Output
Refund file - "30 days" only

1.2 Whitespace and broken lines

PDF extraction leaves hyphenated line breaks, hard wraps and runs of spaces:

import re

pdf_text = "The refund win-\ndow is 30\ndays.   Contact\tsupport@acme.com\n\n\nfor help."

def clean_whitespace(text: str) -> str:
    text = re.sub(r"-\n(\w)", r"\1", text)        # join words split across lines: "win-\ndow" → "window"
    text = re.sub(r"\n{2,}", "\n\n", text)        # keep paragraph breaks (useful for chunking!)
    text = re.sub(r"(?<!\n)\n(?!\n)", " ", text)   # single newlines inside a paragraph → space
    text = re.sub(r"[ \t]+", " ", text)            # collapse spaces and tabs
    return text.strip()

print(clean_whitespace(pdf_text))
Output
The refund window is 30 days. Contact support@acme.com

for help.

1.3 Stripping HTML

import html

page = "<div class='nav'>Home | Pricing</div><p>Refunds take <b>5&nbsp;days</b>.</p><script>track()</script>"

def html_to_text(s: str) -> str:
    s = re.sub(r"<(script|style)\b.*?</\1>", " ", s, flags=re.S | re.I)   # drop code, not just tags
    s = re.sub(r"<[^>]+>", " ", s)                                          # remove remaining tags
    s = html.unescape(s)                                                    # &nbsp; &amp; → characters
    return re.sub(r"\s+", " ", unicodedata.normalize("NFKC", s)).strip()

print(html_to_text(page))
Output
Home | Pricing Refunds take 5 days .

Regex is fine for a quick job. For real web pages use BeautifulSoup or trafilatura — they handle broken HTML and can drop navigation, footers and ads (the "Home | Pricing" noise above).

1.4 Removing boilerplate

Headers, footers and page numbers repeat on every PDF page. They pollute every chunk and match every query. Find lines that repeat across pages and drop them:

from collections import Counter

pages = [
    "ACME Corp — Confidential\nRefunds take 5 days.\nPage 1 of 3",
    "ACME Corp — Confidential\nShipping takes 2 days.\nPage 2 of 3",
    "ACME Corp — Confidential\nSupport is 24x7.\nPage 3 of 3",
]

def strip_boilerplate(pages: list[str], min_share: float = 0.6) -> list[str]:
    norm = lambda line: re.sub(r"\d+", "#", line.strip())          # "Page 1 of 3" ≈ "Page 2 of 3"
    counts = Counter(norm(l) for p in pages for l in set(p.splitlines()))
    repeated = {l for l, c in counts.items() if c / len(pages) >= min_share}
    return ["\n".join(l for l in p.splitlines() if norm(l) not in repeated) for p in pages]

print(strip_boilerplate(pages))
Output
['Refunds take 5 days.', 'Shipping takes 2 days.', 'Support is 24x7.']

1.5 Regex for extraction

Some fields are faster and cheaper to pull out with a pattern than with an LLM:

text = "Order #A-10293 placed 2026-09-14 for ₹1,499.00. Email priya@example.com or call +91 98765 43210."

patterns = {
    "order_id": r"#([A-Z]-\d+)",
    "date": r"\b(\d{4}-\d{2}-\d{2})\b",
    "amount": r"₹([\d,]+(?:\.\d{2})?)",
    "email": r"[\w.+-]+@[\w-]+\.[\w.]+",
    "phone": r"\+91[\s-]?\d{5}[\s-]?\d{5}",
}
for name, p in patterns.items():
    print(f"{name:9}", re.findall(p, text))
Output
order_id  ['A-10293']
date      ['2026-09-14']
amount    ['1,499.00']
email     ['priya@example.com']
phone     ['+91 98765 43210']

1.6 Masking PII before it reaches an LLM

Emails, phone numbers and IDs often shouldn't be sent to a third-party API or stored in logs. Replace them with placeholders — and keep a map if you need to restore them in the answer:

PII = {
    "EMAIL": r"[\w.+-]+@[\w-]+\.[\w.]+",
    "PHONE": r"\+?\d[\d\s-]{8,}\d",
    "PAN": r"\b[A-Z]{5}\d{4}[A-Z]\b",                  # Indian PAN card format
}

def mask_pii(text: str) -> tuple[str, dict[str, str]]:
    mapping: dict[str, str] = {}
    for label, pattern in PII.items():
        for i, value in enumerate(dict.fromkeys(re.findall(pattern, text)), start=1):
            token = f"<{label}_{i}>"
            mapping[token] = value
            text = text.replace(value, token)
    return text, mapping

masked, mapping = mask_pii("Refund to priya@example.com, PAN ABCDE1234F, phone +91 98765 43210.")
print(masked)
print(mapping)
Output
Refund to <EMAIL_1>, PAN <PAN_1>, phone <PHONE_1>.
{'<EMAIL_1>': 'priya@example.com', '<PHONE_1>': '+91 98765 43210', '<PAN_1>': 'ABCDE1234F'}

For production, use a dedicated tool such as Microsoft Presidio, which combines patterns with NER models (names and addresses can't be caught by regex).

1.7 Classic steps — and when to skip them

Older NLP pipelines (and keyword search) used heavy normalisation:

STOPWORDS = {"the", "is", "a", "an", "of", "to", "in", "for", "and", "not", "no", "do", "i"}

def classic_normalise(text: str) -> list[str]:
    words = re.findall(r"[a-z0-9]+", text.lower())     # lowercase + strip punctuation
    return [w for w in words if w not in STOPWORDS]

def crude_stem(word: str) -> str:                       # real stemmers: nltk PorterStemmer, Snowball
    for suffix in ("ing", "ed", "es", "s"):
        if word.endswith(suffix) and len(word) - len(suffix) >= 3:
            return word[: -len(suffix)]
    return word

s = "I do NOT want the refund processed in 5 days."
tokens = classic_normalise(s)
print(tokens)
print([crude_stem(t) for t in tokens])
Output
['want', 'refund', 'processed', '5', 'days']
['want', 'refund', 'process', '5', 'day']

Look at what happened: "NOT" disappeared — the meaning flipped from "I do not want" to "want refund".

Step Classic NLP / keyword search Before an LLM or embedding model
Lowercasing ✅ usually ❌ keep — casing carries meaning (US vs us, Apple vs apple)
Removing punctuation ✅ usually ❌ keep — models use it
Removing stopwords ✅ for BoW/TF-IDF ❌ never — "not", "no", "without" flip meaning
Stemming / lemmatising ✅ helps keyword match ❌ skip — models understand word forms
Unicode / whitespace / HTML cleanup ✅ ✅ always
Boilerplate removal ✅ ✅ always
PII masking depends ✅ when sending to third parties

Stemming chops suffixes by rule ("studies" → "studi"). Lemmatisation uses a dictionary to find the real base word ("studies" → "study", "better" → "good"); use spaCy for it:

# no-run — pip install spacy && python -m spacy download en_core_web_sm
import spacy
nlp = spacy.load("en_core_web_sm")
print([t.lemma_ for t in nlp("The agents were running better searches")])
# ['the', 'agent', 'be', 'run', 'well', 'search']

Practice

  • Write clean_for_rag(text) that applies NFKC, HTML stripping and whitespace cleanup, and test it on a copied web page.
  • Extend mask_pii with an Aadhaar-number pattern (12 digits, often written in groups of 4).

Next: Tokenization — how LLMs actually see your text, and what it costs.