Skip to content

8. Context, memory & cost

Intermediate · 13 min read

Everything the model knows about the current task must fit in one context window, and you pay for every token of it on every call. Context management is therefore both a quality problem (the right information, not too much noise) and a cost problem. This page covers the techniques that solve both.

8.1 Budget the context window

Decide in advance how many tokens each part of the prompt may use:

import tiktoken
enc = tiktoken.get_encoding("o200k_base")

def n_tokens(text: str) -> int:
    return len(enc.encode(text))

BUDGET = {                     # tokens, for a model with a 16K window
    "system": 800,
    "tools": 1_200,
    "retrieved_context": 6_000,
    "history": 4_000,
    "user_message": 1_000,
    "output": 2_000,
}
print("total planned:", f"{sum(BUDGET.values()):,}", "of 16,384")

def fit(text: str, max_tokens: int) -> str:
    ids = enc.encode(text)
    return text if len(ids) <= max_tokens else enc.decode(ids[:max_tokens]) + " …[truncated]"

long_user_message = "Please check this log: " + "ERROR timeout at db.query " * 400
print(n_tokens(long_user_message), "→", n_tokens(fit(long_user_message, BUDGET["user_message"])), "tokens")
Output
total planned: 15,000 of 16,384
2006 → 1005 tokens

The budget makes trade-offs explicit: more retrieved chunks means less history, and the output reserve must never be eaten by the prompt (or answers get cut off).

8.2 Conversation memory strategies

Chat history grows every turn. Four ways to keep it in budget:

Strategy How Good Bad
Full history send everything perfect recall cost grows every turn; hits the limit
Sliding window keep the last N turns / tokens simple, predictable forgets early facts ("my name is Priya")
Summary memory summarise older turns into one message keeps the gist cheaply lossy; an extra LLM call
Retrieval memory store turns/facts in a vector store; retrieve relevant ones scales to long-term memory more moving parts; can retrieve the wrong thing

Production chatbots usually combine them: summary of older turns + last few turns verbatim + retrieved long-term facts.

def summarise(previous_summary: str, old_messages: list[dict]) -> str:
    """Stand-in for an LLM call: 'Update this summary with the key facts from these messages.'"""
    facts = [m["content"] for m in old_messages if m["role"] == "user"]
    return " | ".join(f for f in [previous_summary, *facts] if f)

class ChatMemory:
    def __init__(self, keep_last: int = 4, max_history_tokens: int = 80):
        self.summary, self.turns = "", []
        self.keep_last, self.max_history_tokens = keep_last, max_history_tokens

    def add(self, role: str, content: str):
        self.turns.append({"role": role, "content": content})
        if sum(n_tokens(t["content"]) for t in self.turns) > self.max_history_tokens:
            old, self.turns = self.turns[:-self.keep_last], self.turns[-self.keep_last:]
            self.summary = summarise(self.summary, old)

    def context(self, system: str) -> list[dict]:
        msgs = [{"role": "system", "content": system}]
        if self.summary:
            msgs.append({"role": "system", "content": "Summary of earlier conversation: " + self.summary})
        return msgs + self.turns

mem = ChatMemory()
conversation = [
    ("user", "Hi, my name is Priya and I'm on the Pro plan."),
    ("assistant", "Hi Priya! How can I help with your Pro plan today?"),
    ("user", "I need to export my data to CSV every week, ideally automatically, with all custom fields."),
    ("assistant", "You can schedule weekly exports under Settings > Data > Scheduled exports; custom fields are included."),
    ("user", "Great. Also, my invoices should go to accounts@example.com instead of my email."),
    ("assistant", "Done — invoices will now go to accounts@example.com."),
    ("user", "What plan am I on again?"),
]
for role, text in conversation:
    mem.add(role, text)
for m in mem.context("You are a helpful assistant."):
    print(f"{m['role']:9} {m['content'][:95]}")
Output
system    You are a helpful assistant.
system    Summary of earlier conversation: Hi, my name is Priya and I'm on the Pro plan. | I need to expo
assistant You can schedule weekly exports under Settings > Data > Scheduled exports; custom fields are in
user      Great. Also, my invoices should go to accounts@example.com instead of my email.
assistant Done — invoices will now go to accounts@example.com.
user      What plan am I on again?

The opening turn was dropped from the verbatim history, but "Pro plan" survives in the summary — so the model can still answer the last question.

Agents need memory management too

Every tool call and tool result is appended to an agent's context. Long-running agents trim or summarise old tool outputs, keep only the latest state, and write important findings to a notes file or store — otherwise they run out of context (or budget) half-way through a task.

8.3 Prompt caching

Providers can cache the processed prefix of a prompt (its KV cache — see How LLMs generate text). If the next request starts with exactly the same tokens, that part is billed at a large discount and processed faster. OpenAI does this automatically for long prompts; Anthropic uses explicit cache_control breakpoints.

The rule: stable content first, variable content last.

def cost_per_request(cached_prefix: int, uncached: int, output: int,
                     p_in=3.00, p_cached=0.30, p_out=15.00) -> float:      # USD per 1M tokens, illustrative
    return (cached_prefix * p_cached + uncached * p_in + output * p_out) / 1_000_000

prefix = 8_000          # system prompt + tool definitions + a product manual: identical on every call
question = 200
output = 400

no_cache = cost_per_request(0, prefix + question, output)
with_cache = cost_per_request(prefix, question, output)
print(f"without caching ${no_cache:.4f}   with caching ${with_cache:.4f}   saving {1 - with_cache / no_cache:.0%}")
Output
without caching $0.0306   with caching $0.0090   saving 71%

What breaks the cache: putting a timestamp, the user's name or retrieved chunks before the stable part, reordering tools, or editing the system prompt — any change in the prefix starts a fresh cache.

8.4 Response caching

Prompt caching makes repeated prefixes cheaper. Response caching skips the LLM call entirely when the same (or a very similar) question was answered before.

import hashlib, json

class ExactCache:
    def __init__(self): self.store, self.hits = {}, 0

    def key(self, model: str, messages: list[dict], **params) -> str:
        return hashlib.sha256(json.dumps([model, messages, params], sort_keys=True).encode()).hexdigest()

    def get_or_call(self, call, model, messages, **params):
        k = self.key(model, messages, **params)
        if k in self.store:
            self.hits += 1
        else:
            self.store[k] = call(messages)
        return self.store[k]

calls = []
fake_llm = lambda messages: calls.append(1) or f"answer to: {messages[-1]['content']}"
cache = ExactCache()
for q in ["What is your refund policy?", "What is your refund policy?", "Do you ship to Pune?"]:
    cache.get_or_call(fake_llm, "small", [{"role": "user", "content": q}], temperature=0)
print(f"{len(calls)} LLM calls, {cache.hits} cache hit")
Output
2 LLM calls, 1 cache hit

A semantic cache also matches paraphrases by comparing embeddings of the questions:

from sklearn.feature_extraction.text import TfidfVectorizer       # stand-in for an embedding model
from sklearn.metrics.pairwise import cosine_similarity

cached_questions = ["what is your refund policy", "how long does shipping take", "how do I reset my password"]
cached_answers = ["Refunds within 30 days.", "2–5 days in India.", "Use Settings > Security."]
vec = TfidfVectorizer().fit(cached_questions)
Q = vec.transform(cached_questions)

def semantic_lookup(question: str, threshold: float = 0.6):
    sims = cosine_similarity(vec.transform([question]), Q).ravel()
    best = sims.argmax()
    return (cached_answers[best], round(float(sims[best]), 2)) if sims[best] >= threshold else (None, round(float(sims[best]), 2))

for q in ["What's your refund policy?", "refund policy for damaged items?", "can I pay with UPI"]:
    print(f"{q!r:36} → {semantic_lookup(q)}")
Output
"What's your refund policy?"         → ('Refunds within 30 days.', 0.89)
'refund policy for damaged items?'   → ('Refunds within 30 days.', 0.63)
'can I pay with UPI'                 → (None, 0.0)

Look at the second row: "refund policy for damaged items" got the generic refund answer — a semantically close but different question. Semantic caches need a strict threshold, real embeddings, and should never be used for personalised or time-sensitive answers.

8.5 Model routing

Most traffic is easy. Send easy requests to a small model and only hard ones to a big model:

def route(message: str) -> str:
    hard_signals = ["why", "compare", "plan", "debug", "explain", "design", "step by step"]
    if len(message.split()) > 60 or any(s in message.lower() for s in hard_signals):
        return "large"
    return "small"

traffic = ["Where is my order?", "Reset my password", "Do you ship to Pune?", "Cancel my subscription",
           "Compare the Pro and Business plans for a 20-person team and explain which is cheaper over a year",
           "Thanks!", "Change my email", "Why was I charged twice? Debug what happened with my last 3 invoices"]
routes = [route(m) for m in traffic]
print(routes)

COST = {"small": 0.0004, "large": 0.02}                 # USD per request, illustrative
routed = sum(COST[r] for r in routes)
all_large = COST["large"] * len(traffic)
print(f"routed ${routed:.4f} vs all-large ${all_large:.4f} → {1 - routed / all_large:.0%} cheaper")
Output
['small', 'small', 'small', 'small', 'large', 'small', 'small', 'large']
routed $0.0424 vs all-large $0.1600 → 74% cheaper

Real routers use a small classifier or a cheap LLM call instead of keywords, and escalate when the small model's answer fails validation or confidence checks. Measure quality per route on your eval set.

8.6 Track cost per request

You can't manage what you don't measure. Log tokens and cost for every call, tagged by feature and user:

from collections import defaultdict

PRICE = {"small": (0.15, 0.60), "large": (3.00, 15.00)}            # USD per 1M (input, output)
ledger = defaultdict(float)

def record(feature: str, model: str, input_tokens: int, output_tokens: int):
    p_in, p_out = PRICE[model]
    ledger[feature] += (input_tokens * p_in + output_tokens * p_out) / 1_000_000

for _ in range(1000):
    record("support_chat", "small", 2_500, 300)
for _ in range(50):
    record("report_writer", "large", 12_000, 2_500)

for feature, usd in sorted(ledger.items(), key=lambda x: -x[1]):
    print(f"{feature:14} ${usd:6.2f}")
Output
report_writer  $  3.68
support_chat   $  0.56

50 report requests cost more than 1,000 chat requests — the kind of finding that only shows up when you log it. Tools like Langfuse, LangSmith and Helicone do this per trace automatically.

Interview questions

How do you give a chatbot long-term memory?

Keep recent turns verbatim, summarise older turns, and store durable facts (preferences, account details, decisions) in a database or vector store, retrieving the relevant ones each turn. Keep it within a token budget, and let users see or delete what's remembered.

How would you cut the LLM bill of an app by half?

Measure cost per feature first. Then: route easy requests to a smaller model; trim prompts (fewer, better retrieved chunks, shorter history); order prompts for prompt caching; cache repeated responses; cap output length; use batch APIs for offline jobs; and re-check quality on the eval set after each change.

What is prompt caching and how do you benefit from it?

The provider reuses the processed state of an identical prompt prefix, billing it at a discount and reducing latency. Put stable content (system prompt, tools, documents) first and variable content (user question, retrieved chunks, timestamps) last, and keep the prefix byte-for-byte identical.

Practice

  • Change ChatMemory to keep the summary under 100 tokens by re-summarising when it grows.
  • Add a cost_usd column to your app's logs and find your most expensive feature.

Next: Evaluation & guardrails — prove it works, and keep it safe.