Skip to content

4. Streamlit for GenAI

Intermediate · 14 min read

This page builds the app every GenAI engineer is asked for at least once: upload documents, ask questions, get answers with sources. Retrieval uses TF-IDF from scikit-learn so it runs with no API key; the answer step uses an LLM if a key is set, and falls back to showing the best passages if not.

pip install streamlit pypdf scikit-learn openai

4.1 The plan

flowchart LR
    U[Upload PDF / TXT] --> X[Extract text]
    X --> C[Split into chunks]
    C --> I[Build index<br/>cached]
    Q[Question in chat_input] --> R[Retrieve top-k chunks]
    I --> R
    R --> L[LLM: answer from context]
    L --> A[Answer + sources]
Step Streamlit piece
Upload st.file_uploader(accept_multiple_files=True)
Extract + chunk + index @st.cache_data — runs once per set of files, not on every message
Settings st.sidebar sliders for chunk size and top-k
Chat st.chat_input, st.chat_message, history in st.session_state
Answer st.write_stream for the LLM reply
Sources st.expander with chunk text and score

4.2 The complete app

pdf_chat.py
import io
import re

import numpy as np
import streamlit as st
from sklearn.feature_extraction.text import TfidfVectorizer

st.set_page_config(page_title="Chat with your documents", page_icon="📄")
st.title("📄 Chat with your documents")


def secret(name: str, default=None):
    """Read from .streamlit/secrets.toml, or return `default` if there is no file or key."""
    try:
        return st.secrets[name]
    except Exception:
        return default

# ── Sidebar: settings ───────────────────────────────────────────────────────
with st.sidebar:
    st.header("Settings")
    files = st.file_uploader("Upload PDF or TXT files", type=["pdf", "txt"], accept_multiple_files=True)
    chunk_size = st.slider("Chunk size (words)", 50, 400, 120, step=10)
    overlap = st.slider("Overlap (words)", 0, 100, 20, step=5)
    top_k = st.slider("Chunks per answer (top_k)", 1, 8, 3)
    if st.button("Clear chat"):
        st.session_state.messages = []


# ── Ingestion: extract → chunk → index (cached) ─────────────────────────────
def extract_text(name: str, data: bytes) -> list[tuple[str, str]]:
    """Return (location, text) pairs — one per PDF page, or one for a text file."""
    if name.lower().endswith(".pdf"):
        from pypdf import PdfReader
        reader = PdfReader(io.BytesIO(data))
        return [(f"{name} · p{i + 1}", page.extract_text() or "") for i, page in enumerate(reader.pages)]
    return [(name, data.decode("utf-8", errors="ignore"))]


def chunk_words(text: str, size: int, overlap: int) -> list[str]:
    words = re.sub(r"\s+", " ", text).strip().split(" ")
    step = max(size - overlap, 1)
    return [" ".join(words[i:i + size]) for i in range(0, len(words), step) if words[i:i + size] != [""]]


@st.cache_data(show_spinner="Indexing documents…")
def build_index(uploads: tuple[tuple[str, bytes], ...], size: int, overlap: int):
    chunks, sources = [], []
    for name, data in uploads:
        for location, text in extract_text(name, data):
            for piece in chunk_words(text, size, overlap):
                chunks.append(piece)
                sources.append(location)
    vectorizer = TfidfVectorizer(stop_words="english", ngram_range=(1, 2), sublinear_tf=True)
    matrix = vectorizer.fit_transform(chunks)
    return chunks, sources, vectorizer, matrix


def retrieve(question: str, index, k: int) -> list[dict]:
    chunks, sources, vectorizer, matrix = index
    scores = (matrix @ vectorizer.transform([question]).T).toarray().ravel()   # rows are L2-normalised → cosine
    best = np.argsort(-scores)[:k]
    return [{"text": chunks[i], "source": sources[i], "score": float(scores[i])} for i in best if scores[i] > 0]


# ── Answering ───────────────────────────────────────────────────────────────
def answer_stream(question: str, hits: list[dict]):
    context = "\n\n".join(f"[{i + 1}] ({h['source']}) {h['text']}" for i, h in enumerate(hits))
    api_key = secret("OPENAI_API_KEY")
    if not api_key:                                   # no key → extractive fallback
        yield "**No LLM key configured — most relevant passages:**\n\n"
        for i, h in enumerate(hits, start=1):
            yield f"[{i}] {h['text'][:300]}…\n\n"
        return
    from openai import OpenAI
    stream = OpenAI(api_key=api_key).chat.completions.create(
        model="gpt-4o-mini", stream=True, temperature=0,
        messages=[
            {"role": "system", "content": "Answer ONLY from the numbered context. Cite sources like [1]. "
                                          "If the answer is not in the context, say you don't know."},
            {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
        ],
    )
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            yield chunk.choices[0].delta.content


# ── Main area ───────────────────────────────────────────────────────────────
if not files:
    st.info("⬅️ Upload one or more documents to start.")
    st.stop()

index = build_index(tuple((f.name, f.getvalue()) for f in files), chunk_size, overlap)
st.caption(f"Indexed **{len(index[0])} chunks** from {len(files)} file(s).")

st.session_state.setdefault("messages", [])
for m in st.session_state.messages:
    with st.chat_message(m["role"]):
        st.markdown(m["content"])
        for h in m.get("sources", []):
            with st.expander(f"{h['source']} · score {h['score']:.2f}"):
                st.write(h["text"])

if question := st.chat_input("Ask about your documents"):
    st.session_state.messages.append({"role": "user", "content": question})
    with st.chat_message("user"):
        st.markdown(question)

    hits = retrieve(question, index, top_k)
    with st.chat_message("assistant"):
        if not hits:
            reply = "I couldn't find anything about that in your documents."
            st.markdown(reply)
        else:
            reply = st.write_stream(answer_stream(question, hits))
            for h in hits:
                with st.expander(f"{h['source']} · score {h['score']:.2f}"):
                    st.write(h["text"])
    st.session_state.messages.append({"role": "assistant", "content": reply, "sources": hits})

Run it with streamlit run pdf_chat.py, upload a PDF, and ask a question.

What to notice:

  • build_index is cached on the file bytes and chunk settings. Chatting doesn't re-index; changing the chunk-size slider does.
  • No hits → no LLM call. The cheapest, most reliable way to avoid made-up answers.
  • Sources are saved with each message, so they still show after reruns.
  • To upgrade retrieval, swap TF-IDF for embeddings + a vector DB — only build_index and retrieve change. See NLP for GenAI for hybrid search.

4.3 Secrets

Never hard-code API keys. Streamlit reads .streamlit/secrets.toml:

.streamlit/secrets.toml
OPENAI_API_KEY = "sk-..."

[vector_db]
url = "https://my-index.svc.pinecone.io"
api_key = "pc-..."
# no-run
key = st.secrets["OPENAI_API_KEY"]
db_url = st.secrets["vector_db"]["url"]
  • Add .streamlit/secrets.toml to .gitignore.
  • On Streamlit Community Cloud, paste the same TOML into App settings → Secrets.
  • Secrets are also exported as environment variables, so SDKs that read OPENAI_API_KEY from the environment find it automatically.

For a public demo, let visitors bring their own key (st.text_input(type="password"), stored only in st.session_state) so you don't pay for strangers' usage.

4.4 Streamlit as a frontend for FastAPI

Once the logic grows, keep Streamlit thin: UI only, every LLM/RAG call goes to your FastAPI service. Benefits: the same backend serves Streamlit, a web app and Slack; secrets stay on the server; you can scale them separately.

frontend.py
import uuid

import httpx
import streamlit as st

try:
    API = st.secrets["API_URL"]
except Exception:                       # no secrets file / key → local backend
    API = "http://127.0.0.1:8000"
st.session_state.setdefault("session_id", uuid.uuid4().hex)

question = st.text_input("Question")
if st.button("Ask", disabled=not question):
    try:
        r = httpx.post(f"{API}/rag/query", json={"question": question, "top_k": 3}, timeout=60)
        r.raise_for_status()
    except httpx.HTTPError as e:
        st.error(f"Backend error: {e}")
        st.stop()
    data = r.json()
    st.markdown(data["answer"])
    for s in data["sources"]:
        st.caption(f"{s['doc_id']} · {s['score']:.2f}")

4.5 Other GenAI apps you'll build with the same pieces

App Key elements
Prompt playground text_area for system/user prompts, sidebar model + temperature, side-by-side columns for two models
Eval dashboard cache_data loading a results CSV, st.dataframe, st.bar_chart, filters in the sidebar
Labelling tool st.data_editor to mark answers correct/incorrect, download_button for the labelled CSV
Agent viewer st.status showing each tool call as the agent runs, final answer in a chat bubble
Extraction demo file_uploader → LLM with a Pydantic schema → st.json / st.dataframe of the fields

4.6 Deploying

Option Good for
Streamlit Community Cloud free; connect a public GitHub repo, pick the file, add secrets — done
Hugging Face Spaces free demos with a Streamlit SDK option; good for ML portfolios
Docker on Railway / Render / Cloud Run private apps, custom domains, more memory
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8501
CMD ["streamlit", "run", "pdf_chat.py", "--server.port=8501", "--server.address=0.0.0.0"]

Know the limits

Each user session runs your script on the server; heavy models in cache_resource are shared, but every rerun costs CPU. Streamlit has no built-in user accounts (use an auth proxy or st.login with an identity provider), and long jobs block that user's session. For a public product with many users, move the logic to FastAPI and keep Streamlit — or a JS frontend — as the UI.

Interview questions

Why does a Streamlit app re-embed my PDF on every message, and how do you fix it?

Streamlit reruns the whole script on every interaction, so un-cached ingestion runs again each time. Wrap extraction/chunking/indexing in @st.cache_data (keyed on file bytes and settings) and the embedding model/client in @st.cache_resource.

Where do you keep chat history in Streamlit, and what are its limits?

In st.session_state — per browser tab, in server memory. It is lost on refresh or restart and isn't shared between users or server replicas. Persist to a database (keyed by user/session) for real history.

cache_data vs cache_resource?

cache_data stores return values and gives every caller a copy — for data (DataFrames, API results, arrays). cache_resource stores one shared object — for clients, models and connections. Never put per-user state or per-user API keys in cache_resource.

Practice

  • Add a "Show retrieved chunks only" toggle that skips the LLM — handy for debugging retrieval.
  • Add a sidebar metric showing how many questions were answered with no hits (a gap in your documents).

Next: NLP — tokens, embeddings and the language basics under every LLM app.