Skip to content

5. Data structures

Beginner · 10 min read

Four built-in containers cover almost everything. LLM APIs speak in lists of dictionaries, so these are essential.

Type Example Ordered Changeable Duplicates
list [1, 2, 2] ✅ ✅ ✅
tuple (1, 2) ✅ ❌ ✅
set {1, 2} ❌ ✅ ❌
dict {"k": 1} ✅ (insertion order) ✅ keys unique

5.1 Lists

chunks = ["intro", "pricing", "faq"]

chunks.append("contact")           # add to the end
chunks.insert(0, "title")          # add at a position
print(chunks)                      # → ['title', 'intro', 'pricing', 'faq', 'contact']
print(chunks[1], chunks[-1])       # → intro contact
print(chunks[1:3])                 # → ['intro', 'pricing']   slicing works like strings

chunks.remove("faq")               # remove by value
last = chunks.pop()                # remove and return the last item
print(last, len(chunks))           # → contact 3

scores = [0.4, 0.9, 0.7]
print(sorted(scores, reverse=True))   # → [0.9, 0.7, 0.4]   new sorted list
print(max(scores), sum(scores))       # → 0.9 2.0

5.2 Tuples

point = (12.97, 77.59)             # fixed group of values — can't be changed
lat, lon = point                   # "unpacking" into variables
print(lat)                         # → 12.97

def min_max(values):
    return min(values), max(values)   # functions often return tuples

low, high = min_max([3, 9, 1])
print(low, high)                   # → 1 9

5.3 Sets

tags = {"rag", "agents", "rag"}    # duplicates are dropped automatically
print(len(tags))                   # → 2
print("rag" in tags)               # → True   very fast membership check

a, b = {"python", "sql"}, {"python", "docker"}
print(a & b)                       # → {'python'}   in both (intersection)
print(sorted(a | b))               # → ['docker', 'python', 'sql']   in either (union)

5.4 Dictionaries

message = {"role": "user", "content": "What is RAG?"}   # key → value pairs

print(message["role"])                       # → user
print(message.get("name", "anonymous"))      # → anonymous   .get() avoids errors for missing keys

message["content"] = "Explain RAG simply"    # update a value
message["name"] = "Priya"                    # add a new key

for key, value in message.items():           # loop over pairs
    print(key, "=", value)

print(list(message.keys()))                  # → ['role', 'content', 'name']

5.5 Lists of dictionaries — the LLM message format

messages = [
    {"role": "system", "content": "You are concise."},
    {"role": "user", "content": "Hi!"},
]
messages.append({"role": "assistant", "content": "Hello!"})   # keep the conversation going

user_turns = [m["content"] for m in messages if m["role"] == "user"]
print(user_turns)                            # → ['Hi!']

5.6 Comprehensions

A comprehension builds a new collection in one readable line.

words = ["RAG", "agents", "LLM", "eval"]

lengths = [len(w) for w in words]                 # list: transform each item
print(lengths)                                    # → [3, 6, 3, 4]

short = [w for w in words if len(w) <= 3]         # list: filter
print(short)                                      # → ['RAG', 'LLM']

word_len = {w: len(w) for w in words}             # dict comprehension
print(word_len["agents"])                         # → 6

unique_lengths = {len(w) for w in words}          # set comprehension
print(sorted(unique_lengths))                     # → [3, 4, 6]

Why it matters for GenAI

Chat history is a list of dicts, retrieved chunks are a list of dicts with text and source, and API responses are nested dicts. Comprehensions are how you reshape them.

Practice

  • From docs = [{"source": "a.md", "score": 0.9}, {"source": "b.md", "score": 0.3}], build a list of sources with a score above 0.5.
Answer
docs = [{"source": "a.md", "score": 0.9}, {"source": "b.md", "score": 0.3}]
print([d["source"] for d in docs if d["score"] > 0.5])   # → ['a.md']