4. Streamlit for GenAI¶
Intermediate · 14 min read
This page builds the app every GenAI engineer is asked for at least once: upload documents, ask questions, get answers with sources. Retrieval uses TF-IDF from scikit-learn so it runs with no API key; the answer step uses an LLM if a key is set, and falls back to showing the best passages if not.
4.1 The plan¶
flowchart LR
U[Upload PDF / TXT] --> X[Extract text]
X --> C[Split into chunks]
C --> I[Build index<br/>cached]
Q[Question in chat_input] --> R[Retrieve top-k chunks]
I --> R
R --> L[LLM: answer from context]
L --> A[Answer + sources]
| Step | Streamlit piece |
|---|---|
| Upload | st.file_uploader(accept_multiple_files=True) |
| Extract + chunk + index | @st.cache_data — runs once per set of files, not on every message |
| Settings | st.sidebar sliders for chunk size and top-k |
| Chat | st.chat_input, st.chat_message, history in st.session_state |
| Answer | st.write_stream for the LLM reply |
| Sources | st.expander with chunk text and score |
4.2 The complete app¶
import io
import re
import numpy as np
import streamlit as st
from sklearn.feature_extraction.text import TfidfVectorizer
st.set_page_config(page_title="Chat with your documents", page_icon="📄")
st.title("📄 Chat with your documents")
def secret(name: str, default=None):
"""Read from .streamlit/secrets.toml, or return `default` if there is no file or key."""
try:
return st.secrets[name]
except Exception:
return default
# ── Sidebar: settings ───────────────────────────────────────────────────────
with st.sidebar:
st.header("Settings")
files = st.file_uploader("Upload PDF or TXT files", type=["pdf", "txt"], accept_multiple_files=True)
chunk_size = st.slider("Chunk size (words)", 50, 400, 120, step=10)
overlap = st.slider("Overlap (words)", 0, 100, 20, step=5)
top_k = st.slider("Chunks per answer (top_k)", 1, 8, 3)
if st.button("Clear chat"):
st.session_state.messages = []
# ── Ingestion: extract → chunk → index (cached) ─────────────────────────────
def extract_text(name: str, data: bytes) -> list[tuple[str, str]]:
"""Return (location, text) pairs — one per PDF page, or one for a text file."""
if name.lower().endswith(".pdf"):
from pypdf import PdfReader
reader = PdfReader(io.BytesIO(data))
return [(f"{name} · p{i + 1}", page.extract_text() or "") for i, page in enumerate(reader.pages)]
return [(name, data.decode("utf-8", errors="ignore"))]
def chunk_words(text: str, size: int, overlap: int) -> list[str]:
words = re.sub(r"\s+", " ", text).strip().split(" ")
step = max(size - overlap, 1)
return [" ".join(words[i:i + size]) for i in range(0, len(words), step) if words[i:i + size] != [""]]
@st.cache_data(show_spinner="Indexing documents…")
def build_index(uploads: tuple[tuple[str, bytes], ...], size: int, overlap: int):
chunks, sources = [], []
for name, data in uploads:
for location, text in extract_text(name, data):
for piece in chunk_words(text, size, overlap):
chunks.append(piece)
sources.append(location)
vectorizer = TfidfVectorizer(stop_words="english", ngram_range=(1, 2), sublinear_tf=True)
matrix = vectorizer.fit_transform(chunks)
return chunks, sources, vectorizer, matrix
def retrieve(question: str, index, k: int) -> list[dict]:
chunks, sources, vectorizer, matrix = index
scores = (matrix @ vectorizer.transform([question]).T).toarray().ravel() # rows are L2-normalised → cosine
best = np.argsort(-scores)[:k]
return [{"text": chunks[i], "source": sources[i], "score": float(scores[i])} for i in best if scores[i] > 0]
# ── Answering ───────────────────────────────────────────────────────────────
def answer_stream(question: str, hits: list[dict]):
context = "\n\n".join(f"[{i + 1}] ({h['source']}) {h['text']}" for i, h in enumerate(hits))
api_key = secret("OPENAI_API_KEY")
if not api_key: # no key → extractive fallback
yield "**No LLM key configured — most relevant passages:**\n\n"
for i, h in enumerate(hits, start=1):
yield f"[{i}] {h['text'][:300]}…\n\n"
return
from openai import OpenAI
stream = OpenAI(api_key=api_key).chat.completions.create(
model="gpt-4o-mini", stream=True, temperature=0,
messages=[
{"role": "system", "content": "Answer ONLY from the numbered context. Cite sources like [1]. "
"If the answer is not in the context, say you don't know."},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
],
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
yield chunk.choices[0].delta.content
# ── Main area ───────────────────────────────────────────────────────────────
if not files:
st.info("⬅️ Upload one or more documents to start.")
st.stop()
index = build_index(tuple((f.name, f.getvalue()) for f in files), chunk_size, overlap)
st.caption(f"Indexed **{len(index[0])} chunks** from {len(files)} file(s).")
st.session_state.setdefault("messages", [])
for m in st.session_state.messages:
with st.chat_message(m["role"]):
st.markdown(m["content"])
for h in m.get("sources", []):
with st.expander(f"{h['source']} · score {h['score']:.2f}"):
st.write(h["text"])
if question := st.chat_input("Ask about your documents"):
st.session_state.messages.append({"role": "user", "content": question})
with st.chat_message("user"):
st.markdown(question)
hits = retrieve(question, index, top_k)
with st.chat_message("assistant"):
if not hits:
reply = "I couldn't find anything about that in your documents."
st.markdown(reply)
else:
reply = st.write_stream(answer_stream(question, hits))
for h in hits:
with st.expander(f"{h['source']} · score {h['score']:.2f}"):
st.write(h["text"])
st.session_state.messages.append({"role": "assistant", "content": reply, "sources": hits})
Run it with streamlit run pdf_chat.py, upload a PDF, and ask a question.
What to notice:
build_indexis cached on the file bytes and chunk settings. Chatting doesn't re-index; changing the chunk-size slider does.- No hits → no LLM call. The cheapest, most reliable way to avoid made-up answers.
- Sources are saved with each message, so they still show after reruns.
- To upgrade retrieval, swap TF-IDF for embeddings + a vector DB — only
build_indexandretrievechange. See NLP for GenAI for hybrid search.
4.3 Secrets¶
Never hard-code API keys. Streamlit reads .streamlit/secrets.toml:
OPENAI_API_KEY = "sk-..."
[vector_db]
url = "https://my-index.svc.pinecone.io"
api_key = "pc-..."
- Add
.streamlit/secrets.tomlto.gitignore. - On Streamlit Community Cloud, paste the same TOML into App settings → Secrets.
- Secrets are also exported as environment variables, so SDKs that read
OPENAI_API_KEYfrom the environment find it automatically.
For a public demo, let visitors bring their own key (st.text_input(type="password"), stored only in
st.session_state) so you don't pay for strangers' usage.
4.4 Streamlit as a frontend for FastAPI¶
Once the logic grows, keep Streamlit thin: UI only, every LLM/RAG call goes to your FastAPI service. Benefits: the same backend serves Streamlit, a web app and Slack; secrets stay on the server; you can scale them separately.
import uuid
import httpx
import streamlit as st
try:
API = st.secrets["API_URL"]
except Exception: # no secrets file / key → local backend
API = "http://127.0.0.1:8000"
st.session_state.setdefault("session_id", uuid.uuid4().hex)
question = st.text_input("Question")
if st.button("Ask", disabled=not question):
try:
r = httpx.post(f"{API}/rag/query", json={"question": question, "top_k": 3}, timeout=60)
r.raise_for_status()
except httpx.HTTPError as e:
st.error(f"Backend error: {e}")
st.stop()
data = r.json()
st.markdown(data["answer"])
for s in data["sources"]:
st.caption(f"{s['doc_id']} · {s['score']:.2f}")
4.5 Other GenAI apps you'll build with the same pieces¶
| App | Key elements |
|---|---|
| Prompt playground | text_area for system/user prompts, sidebar model + temperature, side-by-side columns for two models |
| Eval dashboard | cache_data loading a results CSV, st.dataframe, st.bar_chart, filters in the sidebar |
| Labelling tool | st.data_editor to mark answers correct/incorrect, download_button for the labelled CSV |
| Agent viewer | st.status showing each tool call as the agent runs, final answer in a chat bubble |
| Extraction demo | file_uploader → LLM with a Pydantic schema → st.json / st.dataframe of the fields |
4.6 Deploying¶
| Option | Good for |
|---|---|
| Streamlit Community Cloud | free; connect a public GitHub repo, pick the file, add secrets — done |
| Hugging Face Spaces | free demos with a Streamlit SDK option; good for ML portfolios |
| Docker on Railway / Render / Cloud Run | private apps, custom domains, more memory |
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8501
CMD ["streamlit", "run", "pdf_chat.py", "--server.port=8501", "--server.address=0.0.0.0"]
Know the limits
Each user session runs your script on the server; heavy models in cache_resource are shared, but
every rerun costs CPU. Streamlit has no built-in user accounts (use an auth proxy or
st.login with an identity provider), and long jobs block that user's session. For a public product
with many users, move the logic to FastAPI and keep Streamlit — or a JS frontend — as the UI.
Interview questions¶
Why does a Streamlit app re-embed my PDF on every message, and how do you fix it?
Streamlit reruns the whole script on every interaction, so un-cached ingestion runs again each time.
Wrap extraction/chunking/indexing in @st.cache_data (keyed on file bytes and settings) and the
embedding model/client in @st.cache_resource.
Where do you keep chat history in Streamlit, and what are its limits?
In st.session_state — per browser tab, in server memory. It is lost on refresh or restart and isn't
shared between users or server replicas. Persist to a database (keyed by user/session) for real history.
cache_data vs cache_resource?
cache_data stores return values and gives every caller a copy — for data (DataFrames, API results,
arrays). cache_resource stores one shared object — for clients, models and connections. Never put
per-user state or per-user API keys in cache_resource.
Practice¶
- Add a "Show retrieved chunks only" toggle that skips the LLM — handy for debugging retrieval.
- Add a sidebar metric showing how many questions were answered with no hits (a gap in your documents).
Next: NLP — tokens, embeddings and the language basics under every LLM app.