Skip to content

4. Core NLP tasks

Intermediate · 12 min read

Agents and RAG systems are full of small NLP tasks: route a message to the right tool, detect an angry customer, pull names and dates out of a ticket, summarise a long thread. An LLM can do all of them — but a small classic model is often faster, cheaper and more consistent. Knowing both is what interviewers look for.

4.1 Text classification — intent routing

The task: given a user message, pick one label. In an agent this decides which tool or sub-agent handles the request.

Classic: TF-IDF + logistic regression

import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

train = [
    ("I want my money back", "refund"),
    ("refund for a damaged item", "refund"),
    ("how do I return this and get a refund", "refund"),
    ("can I cancel and be refunded", "refund"),
    ("where is my order", "shipping"),
    ("my package has not arrived", "shipping"),
    ("track my delivery", "shipping"),
    ("when will my parcel be delivered", "shipping"),
    ("I forgot my password", "account"),
    ("cannot log in to my account", "account"),
    ("reset my password please", "account"),
    ("change the email on my account", "account"),
]
texts, labels = zip(*train)

clf = make_pipeline(TfidfVectorizer(ngram_range=(1, 2)), LogisticRegression(max_iter=1000))
clf.fit(texts, labels)

for msg in ["my parcel still hasn't arrived", "please refund me", "I can't log in"]:
    probs = clf.predict_proba([msg])[0]
    print(f"{msg!r:34} → {clf.predict([msg])[0]:8} (confidence {probs.max():.2f})")
Output
"my parcel still hasn't arrived"   → shipping (confidence 0.44)
'please refund me'                 → refund   (confidence 0.38)
"I can't log in"                   → account  (confidence 0.38)

All three are right — from 12 examples, in under a millisecond each. Confidences are low because the training set is tiny; with a few hundred real examples per label this kind of model is very strong.

Measuring it

Always evaluate on examples the model hasn't seen:

from sklearn.metrics import classification_report

test = [("refund my order", "refund"), ("item broken, want refund", "refund"),
        ("where's my shipment", "shipping"), ("delivery is late", "shipping"),
        ("password not working", "account"), ("update my phone number on my account", "account")]
X_test, y_test = zip(*test)
print(classification_report(y_test, clf.predict(X_test), zero_division=0))
Output
              precision    recall  f1-score   support

     account       1.00      1.00      1.00         2
      refund       1.00      0.50      0.67         2
    shipping       0.67      1.00      0.80         2

    accuracy                           0.83         6
   macro avg       0.89      0.83      0.82         6
weighted avg       0.89      0.83      0.82         6

"refund my order" was routed to shipping: the phrase "my order" appears in a shipping training example ("where is my order") and outweighed "refund". That is the classic model's weakness — it matches words it has seen, not meaning — and why more (and more varied) training data matters. Precision, recall and F1 are explained in NLP for GenAI.

LLM: zero-shot classification

No training data — describe the labels and let the model choose. Constrain the output so it can only be a valid label:

# no-run — needs an API key
from typing import Literal
from pydantic import BaseModel
from openai import OpenAI

class Route(BaseModel):
    intent: Literal["refund", "shipping", "account", "other"]
    confidence: float

r = OpenAI().chat.completions.parse(
    model="gpt-4o-mini", temperature=0, response_format=Route,
    messages=[{"role": "system", "content": "Classify the customer message. refund = money back or returns; "
                                            "shipping = delivery status or delays; account = login or profile; other = anything else."},
              {"role": "user", "content": "delivery is late"}],
)
print(r.choices[0].message.parsed)       # e.g. intent='shipping' confidence=0.9

See Pydantic for GenAI for why the Literal matters.

Embeddings + nearest label (the middle option)

Embed a few examples per label once; classify a new message by its nearest examples. Understands synonyms like an LLM, but costs one cheap embedding call and is deterministic — a good router for agents.

4.2 Sentiment

A special case of classification (positive / negative / neutral, or a 1–5 score). Useful for escalating angry customers to a human, or for analysing feedback at scale.

reviews = [
    ("absolutely love it, works perfectly", "pos"), ("great value and fast delivery", "pos"),
    ("excellent support, very happy", "pos"), ("good quality, would buy again", "pos"),
    ("terrible, stopped working after a day", "neg"), ("waste of money, very disappointed", "neg"),
    ("awful support and slow refund", "neg"), ("poor quality, broke quickly", "neg"),
]
sent = make_pipeline(TfidfVectorizer(ngram_range=(1, 2)), LogisticRegression(max_iter=1000))
sent.fit(*zip(*reviews))
for r in ["very happy with the quality", "slow delivery and poor support", "not good at all"]:
    print(f"{r!r:34} → {sent.predict([r])[0]}")
Output
'very happy with the quality'      → pos
'slow delivery and poor support'   → neg
'not good at all'                  → pos

"not good at all" → pos: the model saw "good" in positive reviews and never saw "not good". Negation, sarcasm ("great, another delay 🙄") and mixed reviews are where LLMs (or fine-tuned transformers) clearly win.

# no-run — pip install transformers torch ; a pre-trained transformer handles negation
from transformers import pipeline
clf = pipeline("sentiment-analysis", model="distilbert-base-uncased-finetuned-sst-2-english")
print(clf("not good at all"))          # → NEGATIVE, with high confidence

4.3 Named-entity recognition (NER)

NER finds spans of text and labels them: PERSON, ORG, DATE, MONEY, LOCATION, and custom types like ORDER_ID or PRODUCT. In GenAI it powers metadata extraction for RAG filters, PII detection and turning tickets into structured records.

Rules (regex) for well-formed entities

import re

ticket = "Hi, I'm Priya Sharma. Order A-10293 from 14 Sept 2026 (₹1,499) arrived damaged in Bengaluru."
rules = {
    "ORDER_ID": r"\b[A-Z]-\d{5}\b",
    "MONEY": r"₹[\d,]+(?:\.\d{2})?",
    "DATE": r"\b\d{1,2} (?:Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sept?|Oct|Nov|Dec)[a-z]* \d{4}\b",
}
entities = [(m.group(), label) for label, p in rules.items() for m in re.finditer(p, ticket)]
print(entities)
Output
[('A-10293', 'ORDER_ID'), ('₹1,499', 'MONEY'), ('14 Sept 2026', 'DATE')]

Names and places ("Priya Sharma", "Bengaluru") have no fixed pattern — they need a model.

Statistical NER with spaCy

# no-run — pip install spacy && python -m spacy download en_core_web_sm
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp(ticket)
print([(e.text, e.label_) for e in doc.ents])
# typically finds PERSON (Priya Sharma), DATE, MONEY and GPE (Bengaluru); IDs like A-10293 are hit-and-miss

Fast (thousands of documents per second on CPU), but fixed label types and it can mislabel IDs.

LLM extraction

For custom entity types and messy text, describe the schema and let the LLM fill it:

# no-run — needs an API key
class TicketEntities(BaseModel):
    customer_name: str | None
    order_id: str | None
    order_date: str | None          # ISO date
    amount_inr: float | None
    city: str | None
    issue: Literal["damaged", "late", "wrong_item", "other"]

A strong combination: regex for well-formed IDs, an LLM for everything else, Pydantic to validate the result.

4.4 Summarisation

Type How Pros Cons
Extractive pick the most important existing sentences can't invent facts; cheap; traceable choppy; can't merge ideas
Abstractive an LLM writes a new summary fluent, concise, can merge ideas can hallucinate; costs tokens

A classic extractive summariser in a few lines — score each sentence by how similar it is to the document as a whole (the average TF-IDF vector), keep the top ones in their original order:

article = """Acme launched a new refund portal on Monday. Customers can now request refunds without calling support.
The portal processes most refunds within two working days. Previously refunds took up to ten days.
The office canteen also introduced a new menu this week. Acme expects support call volume to fall by a third."""

sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+", article) if s.strip()]
from sklearn.metrics.pairwise import cosine_similarity

S = TfidfVectorizer(stop_words="english").fit_transform(sentences)
centroid = np.asarray(S.mean(axis=0))                            # "what the document is about"
scores = cosine_similarity(S, centroid).ravel()
print(scores.round(2))
keep = sorted(np.argsort(-scores)[:3])                           # top 3, in original order
print(" ".join(sentences[i] for i in keep))
Output
[0.49 0.49 0.56 0.51 0.39 0.45]
Acme launched a new refund portal on Monday. The portal processes most refunds within two working days. Previously refunds took up to ten days.

The canteen sentence (5th score, 0.39) is the least like the rest of the story, so it is dropped.

For long documents with an LLM, use map-reduce: summarise each chunk, then summarise the summaries. Ask for citations or extracted quotes when faithfulness matters.

4.5 Classic model or LLM?

Question Prefer a classic / small model Prefer an LLM
Labelled data available? hundreds+ examples none or a handful
Volume and latency millions of calls, < 50 ms lower volume, seconds are OK
Cost per call ~0 per token
Labels change often? retraining needed edit the prompt
Needs reasoning, world knowledge, negation, sarcasm? weak strong
Must be deterministic and auditable? ✅ harder (use temperature 0 + validation)
Data can't leave your servers? ✅ runs locally needs a self-hosted model

A common production pattern: start with an LLM to ship fast and to label data, then train a small model on those labels for the high-volume path, and keep the LLM as a fallback for low-confidence cases.

Interview questions

How would you build an intent router for a customer-support agent?

Define a small closed label set (plus "other"). Start with an LLM using structured output constrained to the labels (Literal), at temperature 0, with label descriptions. Log predictions; build a labelled set from them; evaluate with per-class precision/recall. If volume or latency demands, train a small classifier (TF-IDF + LR, or embeddings + LR) and route low-confidence cases to the LLM.

Extractive vs abstractive summarisation?

Extractive selects existing sentences — faithful and cheap, but choppy. Abstractive generates new text — fluent and concise, but can hallucinate. For legal/medical/financial content, prefer extractive or abstractive with citations and a faithfulness check.

Why did a bag-of-words sentiment model call 'not good' positive?

Unigram features treat "not" and "good" independently, and "good" was strongly positive in training. Fixes: bigrams ("not good" as a feature), more training data with negations, or a contextual model (fine-tuned transformer or an LLM).

Practice

  • Add a few refund examples that mention "my order" and re-run the classification report.
  • Turn the extractive summariser into a function summarise(text, n_sentences) and try it on a news article.

Next: NLP for GenAI — chunking, hybrid search and evaluation metrics.