Skip to content

1. Arrays

Beginner · 8 min read

An array (ndarray) is a grid of numbers that all have the same type. A 1-D array is a list of numbers; a 2-D array is a table (rows × columns) — like a matrix of embeddings.

1.1 Creating arrays

import numpy as np

a = np.array([1, 2, 3])                 # from a Python list
m = np.array([[1, 2, 3], [4, 5, 6]])    # 2-D: list of rows
print(a)
print(m)
print(type(a).__name__)
Output
[1 2 3]
[[1 2 3]
 [4 5 6]]
ndarray

1.2 Shape, size and dtype

print(m.shape)    # (rows, columns)
print(m.ndim)     # number of dimensions
print(m.size)     # total number of values
print(m.dtype)    # type of every value
f = np.array([1, 2.5, 3])
print(f.dtype)    # one float makes the whole array float
Output
(2, 3)
2
6
int64
float64

Why one type?

Every value having the same type is what makes arrays fast and compact. Embeddings are usually stored as float32 — half the memory of the default float64, with plenty of precision.

emb = np.array([0.12, -0.53, 0.88], dtype=np.float32)
print(emb.dtype, emb.nbytes, "bytes")
print(emb.astype(np.float64).nbytes, "bytes as float64")
Output
float32 12 bytes
24 bytes as float64

1.3 Arrays from helper functions

print(np.zeros(3))               # all zeros
print(np.ones((2, 2)))           # all ones, 2 x 2
print(np.full(3, 7))             # all sevens
print(np.arange(0, 10, 2))       # like range(): start, stop (excluded), step
print(np.linspace(0, 1, 5))      # 5 evenly spaced values, end included
print(np.eye(3))                 # identity matrix
Output
[0. 0. 0.]
[[1. 1.]
 [1. 1.]]
[7 7 7]
[0 2 4 6 8]
[0.   0.25 0.5  0.75 1.  ]
[[1. 0. 0.]
 [0. 1. 0.]
 [0. 0. 1.]]

1.4 Reshaping

reshape changes how the same values are arranged — the total size must stay the same.

x = np.arange(6)
print(x.reshape(2, 3))       # 2 rows, 3 columns
print(x.reshape(3, -1))      # -1 = "work it out" → 3 x 2
print(x.reshape(2, 3).T)     # .T = transpose (swap rows and columns)
print(x.reshape(2, 3).ravel())   # back to 1-D
Output
[[0 1 2]
 [3 4 5]]
[[0 1]
 [2 3]
 [4 5]]
[[0 3]
 [1 4]
 [2 5]]
[0 1 2 3 4 5]

1.5 Random numbers (reproducible)

rng = np.random.default_rng(42)      # seeded: same numbers every run
print(rng.integers(1, 7, size=5))    # 5 dice rolls
print(rng.random(3).round(3))        # floats in [0, 1)
print(rng.normal(0, 1, size=(2, 3)).round(2))   # bell-curve values
print(rng.choice(["rag", "agents", "llm"], size=2, replace=False))
Output
[1 5 4 3 3]
[0.697 0.094 0.976]
[[ 0.13 -0.32 -0.02]
 [-0.85  0.88  0.78]]
['agents' 'llm']

Use default_rng, not np.random.seed

np.random.default_rng(seed) is the modern API: each generator is independent, so tests and experiments stay reproducible.

Practice

  • Create a 3 × 4 array of zeros with dtype=np.float32 and print its nbytes.
  • Turn np.arange(12) into a 4 × 3 array, then transpose it.

Next: Indexing & slicing →