4. Core NLP tasks¶
Intermediate · 12 min read
Agents and RAG systems are full of small NLP tasks: route a message to the right tool, detect an angry customer, pull names and dates out of a ticket, summarise a long thread. An LLM can do all of them — but a small classic model is often faster, cheaper and more consistent. Knowing both is what interviewers look for.
4.1 Text classification — intent routing¶
The task: given a user message, pick one label. In an agent this decides which tool or sub-agent handles the request.
Classic: TF-IDF + logistic regression¶
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
train = [
("I want my money back", "refund"),
("refund for a damaged item", "refund"),
("how do I return this and get a refund", "refund"),
("can I cancel and be refunded", "refund"),
("where is my order", "shipping"),
("my package has not arrived", "shipping"),
("track my delivery", "shipping"),
("when will my parcel be delivered", "shipping"),
("I forgot my password", "account"),
("cannot log in to my account", "account"),
("reset my password please", "account"),
("change the email on my account", "account"),
]
texts, labels = zip(*train)
clf = make_pipeline(TfidfVectorizer(ngram_range=(1, 2)), LogisticRegression(max_iter=1000))
clf.fit(texts, labels)
for msg in ["my parcel still hasn't arrived", "please refund me", "I can't log in"]:
probs = clf.predict_proba([msg])[0]
print(f"{msg!r:34} → {clf.predict([msg])[0]:8} (confidence {probs.max():.2f})")
"my parcel still hasn't arrived" → shipping (confidence 0.44)
'please refund me' → refund (confidence 0.38)
"I can't log in" → account (confidence 0.38)
All three are right — from 12 examples, in under a millisecond each. Confidences are low because the training set is tiny; with a few hundred real examples per label this kind of model is very strong.
Measuring it¶
Always evaluate on examples the model hasn't seen:
from sklearn.metrics import classification_report
test = [("refund my order", "refund"), ("item broken, want refund", "refund"),
("where's my shipment", "shipping"), ("delivery is late", "shipping"),
("password not working", "account"), ("update my phone number on my account", "account")]
X_test, y_test = zip(*test)
print(classification_report(y_test, clf.predict(X_test), zero_division=0))
precision recall f1-score support
account 1.00 1.00 1.00 2
refund 1.00 0.50 0.67 2
shipping 0.67 1.00 0.80 2
accuracy 0.83 6
macro avg 0.89 0.83 0.82 6
weighted avg 0.89 0.83 0.82 6
"refund my order" was routed to shipping: the phrase "my order" appears in a shipping training example ("where is my order") and outweighed "refund". That is the classic model's weakness — it matches words it has seen, not meaning — and why more (and more varied) training data matters. Precision, recall and F1 are explained in NLP for GenAI.
LLM: zero-shot classification¶
No training data — describe the labels and let the model choose. Constrain the output so it can only be a valid label:
# no-run — needs an API key
from typing import Literal
from pydantic import BaseModel
from openai import OpenAI
class Route(BaseModel):
intent: Literal["refund", "shipping", "account", "other"]
confidence: float
r = OpenAI().chat.completions.parse(
model="gpt-4o-mini", temperature=0, response_format=Route,
messages=[{"role": "system", "content": "Classify the customer message. refund = money back or returns; "
"shipping = delivery status or delays; account = login or profile; other = anything else."},
{"role": "user", "content": "delivery is late"}],
)
print(r.choices[0].message.parsed) # e.g. intent='shipping' confidence=0.9
See Pydantic for GenAI for why the Literal matters.
Embeddings + nearest label (the middle option)¶
Embed a few examples per label once; classify a new message by its nearest examples. Understands synonyms like an LLM, but costs one cheap embedding call and is deterministic — a good router for agents.
4.2 Sentiment¶
A special case of classification (positive / negative / neutral, or a 1–5 score). Useful for escalating angry customers to a human, or for analysing feedback at scale.
reviews = [
("absolutely love it, works perfectly", "pos"), ("great value and fast delivery", "pos"),
("excellent support, very happy", "pos"), ("good quality, would buy again", "pos"),
("terrible, stopped working after a day", "neg"), ("waste of money, very disappointed", "neg"),
("awful support and slow refund", "neg"), ("poor quality, broke quickly", "neg"),
]
sent = make_pipeline(TfidfVectorizer(ngram_range=(1, 2)), LogisticRegression(max_iter=1000))
sent.fit(*zip(*reviews))
for r in ["very happy with the quality", "slow delivery and poor support", "not good at all"]:
print(f"{r!r:34} → {sent.predict([r])[0]}")
'very happy with the quality' → pos
'slow delivery and poor support' → neg
'not good at all' → pos
"not good at all" → pos: the model saw "good" in positive reviews and never saw "not good". Negation, sarcasm ("great, another delay 🙄") and mixed reviews are where LLMs (or fine-tuned transformers) clearly win.
# no-run — pip install transformers torch ; a pre-trained transformer handles negation
from transformers import pipeline
clf = pipeline("sentiment-analysis", model="distilbert-base-uncased-finetuned-sst-2-english")
print(clf("not good at all")) # → NEGATIVE, with high confidence
4.3 Named-entity recognition (NER)¶
NER finds spans of text and labels them: PERSON, ORG, DATE, MONEY, LOCATION, and custom types like ORDER_ID or PRODUCT. In GenAI it powers metadata extraction for RAG filters, PII detection and turning tickets into structured records.
Rules (regex) for well-formed entities¶
import re
ticket = "Hi, I'm Priya Sharma. Order A-10293 from 14 Sept 2026 (₹1,499) arrived damaged in Bengaluru."
rules = {
"ORDER_ID": r"\b[A-Z]-\d{5}\b",
"MONEY": r"₹[\d,]+(?:\.\d{2})?",
"DATE": r"\b\d{1,2} (?:Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sept?|Oct|Nov|Dec)[a-z]* \d{4}\b",
}
entities = [(m.group(), label) for label, p in rules.items() for m in re.finditer(p, ticket)]
print(entities)
Names and places ("Priya Sharma", "Bengaluru") have no fixed pattern — they need a model.
Statistical NER with spaCy¶
# no-run — pip install spacy && python -m spacy download en_core_web_sm
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp(ticket)
print([(e.text, e.label_) for e in doc.ents])
# typically finds PERSON (Priya Sharma), DATE, MONEY and GPE (Bengaluru); IDs like A-10293 are hit-and-miss
Fast (thousands of documents per second on CPU), but fixed label types and it can mislabel IDs.
LLM extraction¶
For custom entity types and messy text, describe the schema and let the LLM fill it:
# no-run — needs an API key
class TicketEntities(BaseModel):
customer_name: str | None
order_id: str | None
order_date: str | None # ISO date
amount_inr: float | None
city: str | None
issue: Literal["damaged", "late", "wrong_item", "other"]
A strong combination: regex for well-formed IDs, an LLM for everything else, Pydantic to validate the result.
4.4 Summarisation¶
| Type | How | Pros | Cons |
|---|---|---|---|
| Extractive | pick the most important existing sentences | can't invent facts; cheap; traceable | choppy; can't merge ideas |
| Abstractive | an LLM writes a new summary | fluent, concise, can merge ideas | can hallucinate; costs tokens |
A classic extractive summariser in a few lines — score each sentence by how similar it is to the document as a whole (the average TF-IDF vector), keep the top ones in their original order:
article = """Acme launched a new refund portal on Monday. Customers can now request refunds without calling support.
The portal processes most refunds within two working days. Previously refunds took up to ten days.
The office canteen also introduced a new menu this week. Acme expects support call volume to fall by a third."""
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+", article) if s.strip()]
from sklearn.metrics.pairwise import cosine_similarity
S = TfidfVectorizer(stop_words="english").fit_transform(sentences)
centroid = np.asarray(S.mean(axis=0)) # "what the document is about"
scores = cosine_similarity(S, centroid).ravel()
print(scores.round(2))
keep = sorted(np.argsort(-scores)[:3]) # top 3, in original order
print(" ".join(sentences[i] for i in keep))
[0.49 0.49 0.56 0.51 0.39 0.45]
Acme launched a new refund portal on Monday. The portal processes most refunds within two working days. Previously refunds took up to ten days.
The canteen sentence (5th score, 0.39) is the least like the rest of the story, so it is dropped.
For long documents with an LLM, use map-reduce: summarise each chunk, then summarise the summaries. Ask for citations or extracted quotes when faithfulness matters.
4.5 Classic model or LLM?¶
| Question | Prefer a classic / small model | Prefer an LLM |
|---|---|---|
| Labelled data available? | hundreds+ examples | none or a handful |
| Volume and latency | millions of calls, < 50 ms | lower volume, seconds are OK |
| Cost per call | ~0 | per token |
| Labels change often? | retraining needed | edit the prompt |
| Needs reasoning, world knowledge, negation, sarcasm? | weak | strong |
| Must be deterministic and auditable? | ✅ | harder (use temperature 0 + validation) |
| Data can't leave your servers? | ✅ runs locally | needs a self-hosted model |
A common production pattern: start with an LLM to ship fast and to label data, then train a small model on those labels for the high-volume path, and keep the LLM as a fallback for low-confidence cases.
Interview questions¶
How would you build an intent router for a customer-support agent?
Define a small closed label set (plus "other"). Start with an LLM using structured output constrained
to the labels (Literal), at temperature 0, with label descriptions. Log predictions; build a labelled
set from them; evaluate with per-class precision/recall. If volume or latency demands, train a small
classifier (TF-IDF + LR, or embeddings + LR) and route low-confidence cases to the LLM.
Extractive vs abstractive summarisation?
Extractive selects existing sentences — faithful and cheap, but choppy. Abstractive generates new text — fluent and concise, but can hallucinate. For legal/medical/financial content, prefer extractive or abstractive with citations and a faithfulness check.
Why did a bag-of-words sentiment model call 'not good' positive?
Unigram features treat "not" and "good" independently, and "good" was strongly positive in training. Fixes: bigrams ("not good" as a feature), more training data with negations, or a contextual model (fine-tuned transformer or an LLM).
Practice¶
- Add a few refund examples that mention "my order" and re-run the classification report.
- Turn the extractive summariser into a function
summarise(text, n_sentences)and try it on a news article.
Next: NLP for GenAI — chunking, hybrid search and evaluation metrics.