Skip to content

1. Series & DataFrames

Beginner · 9 min read

  • A Series is one column: values with an index (labels).
  • A DataFrame is a table: several Series that share the same index.

1.1 A Series

import pandas as pd

latency = pd.Series([820, 410, 1350], index=["gpt-4o-mini", "llama-3.1-8b", "gpt-4o"], name="latency_ms")
print(latency)
print(latency["gpt-4o"])          # by label
print(latency.mean().round(1))    # maths works like NumPy
Output
gpt-4o-mini      820
llama-3.1-8b     410
gpt-4o          1350
Name: latency_ms, dtype: int64
1350
860.0

1.2 A DataFrame from a dict of columns

df = pd.DataFrame({
    "model":      ["gpt-4o-mini", "llama-3.1-8b", "gpt-4o", "gpt-4o-mini", "llama-3.1-8b", "gpt-4o"],
    "question":   ["What is RAG?", "What is RAG?", "What is RAG?", "Explain agents", "Explain agents", "Explain agents"],
    "latency_ms": [820, 410, 1350, 960, 450, 1500],
    "tokens":     [210, 190, 240, 260, 230, 300],
    "cost_usd":   [0.00013, 0.00002, 0.0024, 0.00016, 0.00002, 0.003],
    "rating":     [5, 4, 5, 4, 3, 5],
})
print(df)
Output
          model        question  latency_ms  tokens  cost_usd  rating
0   gpt-4o-mini    What is RAG?         820     210   0.00013       5
1  llama-3.1-8b    What is RAG?         410     190   0.00002       4
2        gpt-4o    What is RAG?        1350     240   0.00240       5
3   gpt-4o-mini  Explain agents         960     260   0.00016       4
4  llama-3.1-8b  Explain agents         450     230   0.00002       3
5        gpt-4o  Explain agents        1500     300   0.00300       5

1.3 From a list of dicts (like JSON records)

API responses and log lines usually arrive as one dict per row:

records = [
    {"model": "gpt-4o-mini", "tokens": 210},
    {"model": "gpt-4o", "tokens": 240, "cached": True},   # extra key → new column
]
print(pd.DataFrame(records))
Output
         model  tokens cached
0  gpt-4o-mini     210    NaN
1       gpt-4o     240   True

1.4 Reading files

from io import StringIO

csv_text = """model,question,latency_ms,rating
gpt-4o-mini,What is RAG?,820,5
gpt-4o,What is RAG?,1350,5
"""
small = pd.read_csv(StringIO(csv_text))      # normally: pd.read_csv("results.csv")
print(small)
Output
         model      question  latency_ms  rating
0  gpt-4o-mini  What is RAG?         820       5
1       gpt-4o  What is RAG?        1350       5
Format Read Write
CSV pd.read_csv("f.csv") df.to_csv("f.csv", index=False)
JSON lines pd.read_json("f.jsonl", lines=True) df.to_json("f.jsonl", orient="records", lines=True)
Excel pd.read_excel("f.xlsx") df.to_excel("f.xlsx", index=False)
Parquet (fast, typed) pd.read_parquet("f.parquet") df.to_parquet("f.parquet")

1.5 Inspecting a DataFrame — always do this first

print(df.shape)              # (rows, columns)
print(list(df.columns))
print(df.head(2))            # first rows (tail() = last rows)
Output
(6, 6)
['model', 'question', 'latency_ms', 'tokens', 'cost_usd', 'rating']
          model      question  latency_ms  tokens  cost_usd  rating
0   gpt-4o-mini  What is RAG?         820     210   0.00013       5
1  llama-3.1-8b  What is RAG?         410     190   0.00002       4
df.info()                    # column types and missing values
Output
<class 'pandas.DataFrame'>
RangeIndex: 6 entries, 0 to 5
Data columns (total 6 columns):
 #   Column      Non-Null Count  Dtype  
---  ------      --------------  -----  
 0   model       6 non-null      str    
 1   question    6 non-null      str    
 2   latency_ms  6 non-null      int64  
 3   tokens      6 non-null      int64  
 4   cost_usd    6 non-null      float64
 5   rating      6 non-null      int64  
dtypes: float64(1), int64(3), str(2)
memory usage: 556.0 bytes
print(df.describe().round(4))    # stats for numeric columns
Output
       latency_ms    tokens  cost_usd  rating
count      6.0000    6.0000    6.0000  6.0000
mean     915.0000  238.3333    0.0010  4.3333
std      450.2777   38.6868    0.0014  0.8165
min      410.0000  190.0000    0.0000  3.0000
25%      542.5000  215.0000    0.0000  4.0000
50%      890.0000  235.0000    0.0001  4.5000
75%     1252.5000  255.0000    0.0018  5.0000
max     1500.0000  300.0000    0.0030  5.0000

1.6 Saving

df.to_csv("results.csv", index=False)           # index=False: don't write row numbers
df.to_json("results.jsonl", orient="records", lines=True)
print(pd.read_csv("results.csv").shape)
print(open("results.jsonl", encoding="utf-8").readline().strip())
Output
(6, 6)
{"model":"gpt-4o-mini","question":"What is RAG?","latency_ms":820,"tokens":210,"cost_usd":0.00013,"rating":5}

Text columns in pandas 3

pandas 3 shows text columns as str; pandas 2 shows object. Everything on these pages works the same in both.

Practice

  • Build a DataFrame of 3 documents with columns title, source and n_words.
  • Save it to CSV, read it back and print info().

Next: Selecting & filtering →