8. Context, memory & cost¶
Intermediate · 13 min read
Everything the model knows about the current task must fit in one context window, and you pay for every token of it on every call. Context management is therefore both a quality problem (the right information, not too much noise) and a cost problem. This page covers the techniques that solve both.
8.1 Budget the context window¶
Decide in advance how many tokens each part of the prompt may use:
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
def n_tokens(text: str) -> int:
return len(enc.encode(text))
BUDGET = { # tokens, for a model with a 16K window
"system": 800,
"tools": 1_200,
"retrieved_context": 6_000,
"history": 4_000,
"user_message": 1_000,
"output": 2_000,
}
print("total planned:", f"{sum(BUDGET.values()):,}", "of 16,384")
def fit(text: str, max_tokens: int) -> str:
ids = enc.encode(text)
return text if len(ids) <= max_tokens else enc.decode(ids[:max_tokens]) + " …[truncated]"
long_user_message = "Please check this log: " + "ERROR timeout at db.query " * 400
print(n_tokens(long_user_message), "→", n_tokens(fit(long_user_message, BUDGET["user_message"])), "tokens")
The budget makes trade-offs explicit: more retrieved chunks means less history, and the output reserve must never be eaten by the prompt (or answers get cut off).
8.2 Conversation memory strategies¶
Chat history grows every turn. Four ways to keep it in budget:
| Strategy | How | Good | Bad |
|---|---|---|---|
| Full history | send everything | perfect recall | cost grows every turn; hits the limit |
| Sliding window | keep the last N turns / tokens | simple, predictable | forgets early facts ("my name is Priya") |
| Summary memory | summarise older turns into one message | keeps the gist cheaply | lossy; an extra LLM call |
| Retrieval memory | store turns/facts in a vector store; retrieve relevant ones | scales to long-term memory | more moving parts; can retrieve the wrong thing |
Production chatbots usually combine them: summary of older turns + last few turns verbatim + retrieved long-term facts.
def summarise(previous_summary: str, old_messages: list[dict]) -> str:
"""Stand-in for an LLM call: 'Update this summary with the key facts from these messages.'"""
facts = [m["content"] for m in old_messages if m["role"] == "user"]
return " | ".join(f for f in [previous_summary, *facts] if f)
class ChatMemory:
def __init__(self, keep_last: int = 4, max_history_tokens: int = 80):
self.summary, self.turns = "", []
self.keep_last, self.max_history_tokens = keep_last, max_history_tokens
def add(self, role: str, content: str):
self.turns.append({"role": role, "content": content})
if sum(n_tokens(t["content"]) for t in self.turns) > self.max_history_tokens:
old, self.turns = self.turns[:-self.keep_last], self.turns[-self.keep_last:]
self.summary = summarise(self.summary, old)
def context(self, system: str) -> list[dict]:
msgs = [{"role": "system", "content": system}]
if self.summary:
msgs.append({"role": "system", "content": "Summary of earlier conversation: " + self.summary})
return msgs + self.turns
mem = ChatMemory()
conversation = [
("user", "Hi, my name is Priya and I'm on the Pro plan."),
("assistant", "Hi Priya! How can I help with your Pro plan today?"),
("user", "I need to export my data to CSV every week, ideally automatically, with all custom fields."),
("assistant", "You can schedule weekly exports under Settings > Data > Scheduled exports; custom fields are included."),
("user", "Great. Also, my invoices should go to accounts@example.com instead of my email."),
("assistant", "Done — invoices will now go to accounts@example.com."),
("user", "What plan am I on again?"),
]
for role, text in conversation:
mem.add(role, text)
for m in mem.context("You are a helpful assistant."):
print(f"{m['role']:9} {m['content'][:95]}")
system You are a helpful assistant.
system Summary of earlier conversation: Hi, my name is Priya and I'm on the Pro plan. | I need to expo
assistant You can schedule weekly exports under Settings > Data > Scheduled exports; custom fields are in
user Great. Also, my invoices should go to accounts@example.com instead of my email.
assistant Done — invoices will now go to accounts@example.com.
user What plan am I on again?
The opening turn was dropped from the verbatim history, but "Pro plan" survives in the summary — so the model can still answer the last question.
Agents need memory management too
Every tool call and tool result is appended to an agent's context. Long-running agents trim or summarise old tool outputs, keep only the latest state, and write important findings to a notes file or store — otherwise they run out of context (or budget) half-way through a task.
8.3 Prompt caching¶
Providers can cache the processed prefix of a prompt (its KV cache — see
How LLMs generate text).
If the next request starts with exactly the same tokens, that part is billed at a large discount and processed
faster. OpenAI does this automatically for long prompts; Anthropic uses explicit cache_control breakpoints.
The rule: stable content first, variable content last.
def cost_per_request(cached_prefix: int, uncached: int, output: int,
p_in=3.00, p_cached=0.30, p_out=15.00) -> float: # USD per 1M tokens, illustrative
return (cached_prefix * p_cached + uncached * p_in + output * p_out) / 1_000_000
prefix = 8_000 # system prompt + tool definitions + a product manual: identical on every call
question = 200
output = 400
no_cache = cost_per_request(0, prefix + question, output)
with_cache = cost_per_request(prefix, question, output)
print(f"without caching ${no_cache:.4f} with caching ${with_cache:.4f} saving {1 - with_cache / no_cache:.0%}")
What breaks the cache: putting a timestamp, the user's name or retrieved chunks before the stable part, reordering tools, or editing the system prompt — any change in the prefix starts a fresh cache.
8.4 Response caching¶
Prompt caching makes repeated prefixes cheaper. Response caching skips the LLM call entirely when the same (or a very similar) question was answered before.
import hashlib, json
class ExactCache:
def __init__(self): self.store, self.hits = {}, 0
def key(self, model: str, messages: list[dict], **params) -> str:
return hashlib.sha256(json.dumps([model, messages, params], sort_keys=True).encode()).hexdigest()
def get_or_call(self, call, model, messages, **params):
k = self.key(model, messages, **params)
if k in self.store:
self.hits += 1
else:
self.store[k] = call(messages)
return self.store[k]
calls = []
fake_llm = lambda messages: calls.append(1) or f"answer to: {messages[-1]['content']}"
cache = ExactCache()
for q in ["What is your refund policy?", "What is your refund policy?", "Do you ship to Pune?"]:
cache.get_or_call(fake_llm, "small", [{"role": "user", "content": q}], temperature=0)
print(f"{len(calls)} LLM calls, {cache.hits} cache hit")
A semantic cache also matches paraphrases by comparing embeddings of the questions:
from sklearn.feature_extraction.text import TfidfVectorizer # stand-in for an embedding model
from sklearn.metrics.pairwise import cosine_similarity
cached_questions = ["what is your refund policy", "how long does shipping take", "how do I reset my password"]
cached_answers = ["Refunds within 30 days.", "2–5 days in India.", "Use Settings > Security."]
vec = TfidfVectorizer().fit(cached_questions)
Q = vec.transform(cached_questions)
def semantic_lookup(question: str, threshold: float = 0.6):
sims = cosine_similarity(vec.transform([question]), Q).ravel()
best = sims.argmax()
return (cached_answers[best], round(float(sims[best]), 2)) if sims[best] >= threshold else (None, round(float(sims[best]), 2))
for q in ["What's your refund policy?", "refund policy for damaged items?", "can I pay with UPI"]:
print(f"{q!r:36} → {semantic_lookup(q)}")
"What's your refund policy?" → ('Refunds within 30 days.', 0.89)
'refund policy for damaged items?' → ('Refunds within 30 days.', 0.63)
'can I pay with UPI' → (None, 0.0)
Look at the second row: "refund policy for damaged items" got the generic refund answer — a semantically close but different question. Semantic caches need a strict threshold, real embeddings, and should never be used for personalised or time-sensitive answers.
8.5 Model routing¶
Most traffic is easy. Send easy requests to a small model and only hard ones to a big model:
def route(message: str) -> str:
hard_signals = ["why", "compare", "plan", "debug", "explain", "design", "step by step"]
if len(message.split()) > 60 or any(s in message.lower() for s in hard_signals):
return "large"
return "small"
traffic = ["Where is my order?", "Reset my password", "Do you ship to Pune?", "Cancel my subscription",
"Compare the Pro and Business plans for a 20-person team and explain which is cheaper over a year",
"Thanks!", "Change my email", "Why was I charged twice? Debug what happened with my last 3 invoices"]
routes = [route(m) for m in traffic]
print(routes)
COST = {"small": 0.0004, "large": 0.02} # USD per request, illustrative
routed = sum(COST[r] for r in routes)
all_large = COST["large"] * len(traffic)
print(f"routed ${routed:.4f} vs all-large ${all_large:.4f} → {1 - routed / all_large:.0%} cheaper")
['small', 'small', 'small', 'small', 'large', 'small', 'small', 'large']
routed $0.0424 vs all-large $0.1600 → 74% cheaper
Real routers use a small classifier or a cheap LLM call instead of keywords, and escalate when the small model's answer fails validation or confidence checks. Measure quality per route on your eval set.
8.6 Track cost per request¶
You can't manage what you don't measure. Log tokens and cost for every call, tagged by feature and user:
from collections import defaultdict
PRICE = {"small": (0.15, 0.60), "large": (3.00, 15.00)} # USD per 1M (input, output)
ledger = defaultdict(float)
def record(feature: str, model: str, input_tokens: int, output_tokens: int):
p_in, p_out = PRICE[model]
ledger[feature] += (input_tokens * p_in + output_tokens * p_out) / 1_000_000
for _ in range(1000):
record("support_chat", "small", 2_500, 300)
for _ in range(50):
record("report_writer", "large", 12_000, 2_500)
for feature, usd in sorted(ledger.items(), key=lambda x: -x[1]):
print(f"{feature:14} ${usd:6.2f}")
50 report requests cost more than 1,000 chat requests — the kind of finding that only shows up when you log it. Tools like Langfuse, LangSmith and Helicone do this per trace automatically.
Interview questions¶
How do you give a chatbot long-term memory?
Keep recent turns verbatim, summarise older turns, and store durable facts (preferences, account details, decisions) in a database or vector store, retrieving the relevant ones each turn. Keep it within a token budget, and let users see or delete what's remembered.
How would you cut the LLM bill of an app by half?
Measure cost per feature first. Then: route easy requests to a smaller model; trim prompts (fewer, better retrieved chunks, shorter history); order prompts for prompt caching; cache repeated responses; cap output length; use batch APIs for offline jobs; and re-check quality on the eval set after each change.
What is prompt caching and how do you benefit from it?
The provider reuses the processed state of an identical prompt prefix, billing it at a discount and reducing latency. Put stable content (system prompt, tools, documents) first and variable content (user question, retrieved chunks, timestamps) last, and keep the prefix byte-for-byte identical.
Practice¶
- Change
ChatMemoryto keep the summary under 100 tokens by re-summarising when it grows. - Add a
cost_usdcolumn to your app's logs and find your most expensive feature.
Next: Evaluation & guardrails — prove it works, and keep it safe.