2. Layout, state & caching¶
Intermediate · 10 min read
2.1 Sidebar — settings go here¶
Anything called on st.sidebar appears in the left panel. Convention: settings in the sidebar, the
conversation in the main area.
import streamlit as st
with st.sidebar:
st.header("Settings")
model = st.selectbox("Model", ["gpt-4o-mini", "claude-sonnet-5-5"])
temperature = st.slider("Temperature", 0.0, 1.0, 0.2)
top_k = st.number_input("Chunks to retrieve (top_k)", 1, 20, 4)
st.title("Docs assistant")
st.caption(f"{model} · temperature {temperature} · top_k {top_k}")
2.2 Columns, tabs, expanders, containers¶
import streamlit as st
st.set_page_config(page_title="RAG debugger", page_icon="🔎", layout="wide") # first st call
left, right = st.columns([2, 1]) # 2:1 width ratio
left.text_area("Question", "What is the refund window?")
right.metric("Retrieved chunks", 4)
answer_tab, sources_tab, debug_tab = st.tabs(["Answer", "Sources", "Debug"])
with answer_tab:
st.write("Refunds are accepted within **30 days** of delivery.")
with sources_tab:
for i, src in enumerate(["policy.pdf · p3", "faq.md · §2"], start=1):
with st.expander(f"[{i}] {src}"):
st.write("…the chunk text…")
with debug_tab:
st.json({"prompt_tokens": 1240, "latency_ms": 820})
box = st.container(border=True)
box.write("A bordered card. `st.empty()` gives a slot you can overwrite later — useful for live updates.")
2.3 Forms — submit once, not on every keystroke¶
Normally each widget change reruns the script. Inside a form, nothing happens until the submit button:
import streamlit as st
with st.form("eval_case"):
question = st.text_input("Question")
expected = st.text_area("Expected answer")
difficulty = st.radio("Difficulty", ["easy", "medium", "hard"], horizontal=True)
submitted = st.form_submit_button("Add test case")
if submitted:
st.success(f"Added a {difficulty} case: {question!r}")
Use forms whenever a change would trigger an expensive call — you don't want an LLM request per keystroke.
2.4 st.session_state — memory across reruns¶
st.session_state is a dict that survives reruns for one browser tab (one user session).
import streamlit as st
# 1) initialise once
st.session_state.setdefault("history", [])
st.session_state.setdefault("total_tokens", 0)
# 2) update on events
question = st.text_input("Ask", key="question")
if st.button("Send") and question:
answer = f"(fake answer to: {question})"
st.session_state.history.append((question, answer))
st.session_state.total_tokens += len(question.split()) + len(answer.split())
# 3) read anywhere
st.metric("Tokens used this session", st.session_state.total_tokens)
for q, a in reversed(st.session_state.history):
st.write(f"**Q:** {q} \n**A:** {a}")
if st.button("Clear history"):
st.session_state.history = []
st.session_state.total_tokens = 0
st.rerun() # redraw immediately with the cleared state
- Widgets with a
key=store their value inst.session_state[key]automatically. - State is per tab and lives in memory: refresh the page and it's gone; other users never see it. For history that must persist, save to a database.
Callbacks¶
on_click / on_change run before the rerun — handy for clearing an input after sending:
import streamlit as st
st.session_state.setdefault("log", [])
def send():
st.session_state.log.append(st.session_state.draft)
st.session_state.draft = "" # clearing a widget's value is only allowed in a callback
st.text_input("Message", key="draft")
st.button("Send", on_click=send)
st.write(st.session_state.log)
2.5 Caching — don't redo slow work on every rerun¶
There are two caches, and picking the right one matters:
@st.cache_data |
@st.cache_resource |
|
|---|---|---|
| Caches | data: the return value, copied for each caller | one shared object, not copied |
| Use for | API results, DataFrames, parsed PDFs, embeddings arrays | LLM / embedding clients, loaded ML models, vector-store connections, DB pools |
| Shared between users? | yes (each gets a copy) | yes (the same object) |
| Must be | serialisable (pickle) | thread-safe if users share it |
import time
import streamlit as st
@st.cache_resource # created once per server process
def get_embedder():
time.sleep(2) # stand-in for loading a sentence-transformers model
return lambda texts: [[len(t), t.count(" ")] for t in texts]
@st.cache_data(ttl=3600, show_spinner="Embedding documents…") # re-computed after 1 hour
def embed_corpus(texts: tuple[str, ...]):
return get_embedder()(list(texts))
docs = ("Refunds take 5 days.", "Shipping is free over ₹499.")
vectors = embed_corpus(docs) # slow the first time, instant after
st.write(f"{len(vectors)} vectors ready")
if st.button("Clear caches"):
st.cache_data.clear()
Cached functions are keyed on their arguments — change docs and it recomputes; same docs → instant.
Never cache per-user secrets in cache_resource
A cached client created with one user's API key is shared with every user. Cache clients built
from server-side secrets only; keep a user's own key in st.session_state.
2.6 Multipage apps¶
Put pages in a list and let Streamlit build the navigation:
import streamlit as st
def chat_page():
st.title("Chat")
def eval_page():
st.title("Evaluation results")
nav = st.navigation([st.Page(chat_page, title="Chat", icon="💬"),
st.Page(eval_page, title="Evals", icon="📊")])
nav.run()
In a bigger app each page is its own file: st.Page("pages/chat.py", title="Chat").
st.session_state is shared across pages.
Practice¶
- Move the model settings of your token-counter app into the sidebar.
- Add a
cache_datafunction that "loads" a CSV of eval results, and show it in a tab withst.dataframe.
Next: Chat apps — a ChatGPT-style UI with streaming.