1. Text preprocessing¶
Beginner · 10 min read
Garbage in, garbage out — and in RAG, garbage in means wrong chunks retrieved. Preprocessing is the unglamorous step that decides whether search works. The rule for the LLM era:
Clean the noise, keep the meaning. Remove what no reader needs (HTML tags, repeated headers, broken whitespace). Keep what an LLM uses (casing, punctuation, stopwords, numbers).
1.1 Unicode normalisation¶
Text from PDFs and web pages contains look-alike characters that break exact matching and keyword search:
import unicodedata
raw = "Refund file — “30 days” only" # full-width letters, 'fi' ligature, smart quotes, non-breaking space
nfkc = unicodedata.normalize("NFKC", raw)
print(nfkc)
print("Refund" in raw, "Refund" in nfkc)
NFKC turns compatibility characters (full-width, ligatures, non-breaking spaces) into their plain forms.
Smart quotes and dashes are kept — they are real characters. Replace them only if your search needs it:
SMART = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'", "—": "-", "–": "-"})
print(nfkc.translate(SMART))
1.2 Whitespace and broken lines¶
PDF extraction leaves hyphenated line breaks, hard wraps and runs of spaces:
import re
pdf_text = "The refund win-\ndow is 30\ndays. Contact\tsupport@acme.com\n\n\nfor help."
def clean_whitespace(text: str) -> str:
text = re.sub(r"-\n(\w)", r"\1", text) # join words split across lines: "win-\ndow" → "window"
text = re.sub(r"\n{2,}", "\n\n", text) # keep paragraph breaks (useful for chunking!)
text = re.sub(r"(?<!\n)\n(?!\n)", " ", text) # single newlines inside a paragraph → space
text = re.sub(r"[ \t]+", " ", text) # collapse spaces and tabs
return text.strip()
print(clean_whitespace(pdf_text))
1.3 Stripping HTML¶
import html
page = "<div class='nav'>Home | Pricing</div><p>Refunds take <b>5 days</b>.</p><script>track()</script>"
def html_to_text(s: str) -> str:
s = re.sub(r"<(script|style)\b.*?</\1>", " ", s, flags=re.S | re.I) # drop code, not just tags
s = re.sub(r"<[^>]+>", " ", s) # remove remaining tags
s = html.unescape(s) # & → characters
return re.sub(r"\s+", " ", unicodedata.normalize("NFKC", s)).strip()
print(html_to_text(page))
Regex is fine for a quick job. For real web pages use BeautifulSoup or trafilatura — they handle broken HTML and can drop navigation, footers and ads (the "Home | Pricing" noise above).
1.4 Removing boilerplate¶
Headers, footers and page numbers repeat on every PDF page. They pollute every chunk and match every query. Find lines that repeat across pages and drop them:
from collections import Counter
pages = [
"ACME Corp — Confidential\nRefunds take 5 days.\nPage 1 of 3",
"ACME Corp — Confidential\nShipping takes 2 days.\nPage 2 of 3",
"ACME Corp — Confidential\nSupport is 24x7.\nPage 3 of 3",
]
def strip_boilerplate(pages: list[str], min_share: float = 0.6) -> list[str]:
norm = lambda line: re.sub(r"\d+", "#", line.strip()) # "Page 1 of 3" ≈ "Page 2 of 3"
counts = Counter(norm(l) for p in pages for l in set(p.splitlines()))
repeated = {l for l, c in counts.items() if c / len(pages) >= min_share}
return ["\n".join(l for l in p.splitlines() if norm(l) not in repeated) for p in pages]
print(strip_boilerplate(pages))
1.5 Regex for extraction¶
Some fields are faster and cheaper to pull out with a pattern than with an LLM:
text = "Order #A-10293 placed 2026-09-14 for ₹1,499.00. Email priya@example.com or call +91 98765 43210."
patterns = {
"order_id": r"#([A-Z]-\d+)",
"date": r"\b(\d{4}-\d{2}-\d{2})\b",
"amount": r"₹([\d,]+(?:\.\d{2})?)",
"email": r"[\w.+-]+@[\w-]+\.[\w.]+",
"phone": r"\+91[\s-]?\d{5}[\s-]?\d{5}",
}
for name, p in patterns.items():
print(f"{name:9}", re.findall(p, text))
order_id ['A-10293']
date ['2026-09-14']
amount ['1,499.00']
email ['priya@example.com']
phone ['+91 98765 43210']
1.6 Masking PII before it reaches an LLM¶
Emails, phone numbers and IDs often shouldn't be sent to a third-party API or stored in logs. Replace them with placeholders — and keep a map if you need to restore them in the answer:
PII = {
"EMAIL": r"[\w.+-]+@[\w-]+\.[\w.]+",
"PHONE": r"\+?\d[\d\s-]{8,}\d",
"PAN": r"\b[A-Z]{5}\d{4}[A-Z]\b", # Indian PAN card format
}
def mask_pii(text: str) -> tuple[str, dict[str, str]]:
mapping: dict[str, str] = {}
for label, pattern in PII.items():
for i, value in enumerate(dict.fromkeys(re.findall(pattern, text)), start=1):
token = f"<{label}_{i}>"
mapping[token] = value
text = text.replace(value, token)
return text, mapping
masked, mapping = mask_pii("Refund to priya@example.com, PAN ABCDE1234F, phone +91 98765 43210.")
print(masked)
print(mapping)
Refund to <EMAIL_1>, PAN <PAN_1>, phone <PHONE_1>.
{'<EMAIL_1>': 'priya@example.com', '<PHONE_1>': '+91 98765 43210', '<PAN_1>': 'ABCDE1234F'}
For production, use a dedicated tool such as Microsoft Presidio, which combines patterns with NER models (names and addresses can't be caught by regex).
1.7 Classic steps — and when to skip them¶
Older NLP pipelines (and keyword search) used heavy normalisation:
STOPWORDS = {"the", "is", "a", "an", "of", "to", "in", "for", "and", "not", "no", "do", "i"}
def classic_normalise(text: str) -> list[str]:
words = re.findall(r"[a-z0-9]+", text.lower()) # lowercase + strip punctuation
return [w for w in words if w not in STOPWORDS]
def crude_stem(word: str) -> str: # real stemmers: nltk PorterStemmer, Snowball
for suffix in ("ing", "ed", "es", "s"):
if word.endswith(suffix) and len(word) - len(suffix) >= 3:
return word[: -len(suffix)]
return word
s = "I do NOT want the refund processed in 5 days."
tokens = classic_normalise(s)
print(tokens)
print([crude_stem(t) for t in tokens])
Look at what happened: "NOT" disappeared — the meaning flipped from "I do not want" to "want refund".
| Step | Classic NLP / keyword search | Before an LLM or embedding model |
|---|---|---|
| Lowercasing | ✅ usually | ❌ keep — casing carries meaning (US vs us, Apple vs apple) |
| Removing punctuation | ✅ usually | ❌ keep — models use it |
| Removing stopwords | ✅ for BoW/TF-IDF | ❌ never — "not", "no", "without" flip meaning |
| Stemming / lemmatising | ✅ helps keyword match | ❌ skip — models understand word forms |
| Unicode / whitespace / HTML cleanup | ✅ | ✅ always |
| Boilerplate removal | ✅ | ✅ always |
| PII masking | depends | ✅ when sending to third parties |
Stemming chops suffixes by rule ("studies" → "studi"). Lemmatisation uses a dictionary to find the real base word ("studies" → "study", "better" → "good"); use spaCy for it:
# no-run — pip install spacy && python -m spacy download en_core_web_sm
import spacy
nlp = spacy.load("en_core_web_sm")
print([t.lemma_ for t in nlp("The agents were running better searches")])
# ['the', 'agent', 'be', 'run', 'well', 'search']
Practice¶
- Write
clean_for_rag(text)that applies NFKC, HTML stripping and whitespace cleanup, and test it on a copied web page. - Extend
mask_piiwith an Aadhaar-number pattern (12 digits, often written in groups of 4).
Next: Tokenization — how LLMs actually see your text, and what it costs.