1. Series & DataFrames¶
Beginner · 9 min read
- A Series is one column: values with an index (labels).
- A DataFrame is a table: several Series that share the same index.
1.1 A Series¶
import pandas as pd
latency = pd.Series([820, 410, 1350], index=["gpt-4o-mini", "llama-3.1-8b", "gpt-4o"], name="latency_ms")
print(latency)
print(latency["gpt-4o"]) # by label
print(latency.mean().round(1)) # maths works like NumPy
1.2 A DataFrame from a dict of columns¶
df = pd.DataFrame({
"model": ["gpt-4o-mini", "llama-3.1-8b", "gpt-4o", "gpt-4o-mini", "llama-3.1-8b", "gpt-4o"],
"question": ["What is RAG?", "What is RAG?", "What is RAG?", "Explain agents", "Explain agents", "Explain agents"],
"latency_ms": [820, 410, 1350, 960, 450, 1500],
"tokens": [210, 190, 240, 260, 230, 300],
"cost_usd": [0.00013, 0.00002, 0.0024, 0.00016, 0.00002, 0.003],
"rating": [5, 4, 5, 4, 3, 5],
})
print(df)
Output
model question latency_ms tokens cost_usd rating
0 gpt-4o-mini What is RAG? 820 210 0.00013 5
1 llama-3.1-8b What is RAG? 410 190 0.00002 4
2 gpt-4o What is RAG? 1350 240 0.00240 5
3 gpt-4o-mini Explain agents 960 260 0.00016 4
4 llama-3.1-8b Explain agents 450 230 0.00002 3
5 gpt-4o Explain agents 1500 300 0.00300 5
1.3 From a list of dicts (like JSON records)¶
API responses and log lines usually arrive as one dict per row:
records = [
{"model": "gpt-4o-mini", "tokens": 210},
{"model": "gpt-4o", "tokens": 240, "cached": True}, # extra key → new column
]
print(pd.DataFrame(records))
1.4 Reading files¶
from io import StringIO
csv_text = """model,question,latency_ms,rating
gpt-4o-mini,What is RAG?,820,5
gpt-4o,What is RAG?,1350,5
"""
small = pd.read_csv(StringIO(csv_text)) # normally: pd.read_csv("results.csv")
print(small)
Output
model question latency_ms rating
0 gpt-4o-mini What is RAG? 820 5
1 gpt-4o What is RAG? 1350 5
| Format | Read | Write |
|---|---|---|
| CSV | pd.read_csv("f.csv") |
df.to_csv("f.csv", index=False) |
| JSON lines | pd.read_json("f.jsonl", lines=True) |
df.to_json("f.jsonl", orient="records", lines=True) |
| Excel | pd.read_excel("f.xlsx") |
df.to_excel("f.xlsx", index=False) |
| Parquet (fast, typed) | pd.read_parquet("f.parquet") |
df.to_parquet("f.parquet") |
1.5 Inspecting a DataFrame — always do this first¶
print(df.shape) # (rows, columns)
print(list(df.columns))
print(df.head(2)) # first rows (tail() = last rows)
Output
(6, 6)
['model', 'question', 'latency_ms', 'tokens', 'cost_usd', 'rating']
model question latency_ms tokens cost_usd rating
0 gpt-4o-mini What is RAG? 820 210 0.00013 5
1 llama-3.1-8b What is RAG? 410 190 0.00002 4
Output
<class 'pandas.DataFrame'>
RangeIndex: 6 entries, 0 to 5
Data columns (total 6 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 model 6 non-null str
1 question 6 non-null str
2 latency_ms 6 non-null int64
3 tokens 6 non-null int64
4 cost_usd 6 non-null float64
5 rating 6 non-null int64
dtypes: float64(1), int64(3), str(2)
memory usage: 556.0 bytes
Output
latency_ms tokens cost_usd rating
count 6.0000 6.0000 6.0000 6.0000
mean 915.0000 238.3333 0.0010 4.3333
std 450.2777 38.6868 0.0014 0.8165
min 410.0000 190.0000 0.0000 3.0000
25% 542.5000 215.0000 0.0000 4.0000
50% 890.0000 235.0000 0.0001 4.5000
75% 1252.5000 255.0000 0.0018 5.0000
max 1500.0000 300.0000 0.0030 5.0000
1.6 Saving¶
df.to_csv("results.csv", index=False) # index=False: don't write row numbers
df.to_json("results.jsonl", orient="records", lines=True)
print(pd.read_csv("results.csv").shape)
print(open("results.jsonl", encoding="utf-8").readline().strip())
Output
(6, 6)
{"model":"gpt-4o-mini","question":"What is RAG?","latency_ms":820,"tokens":210,"cost_usd":0.00013,"rating":5}
Text columns in pandas 3
pandas 3 shows text columns as str; pandas 2 shows object. Everything on these pages works
the same in both.
Practice¶
- Build a DataFrame of 3 documents with columns
title,sourceandn_words. - Save it to CSV, read it back and print
info().
Next: Selecting & filtering →