NUMPY:Aggregations, axes, reshape, and stacking
Mastering aggregations, axes, reshape, and stacking concepts and implementation.
Axis 0 goes down, axis 1 goes across
sum, mean, min, max, and std can reduce a whole array or a single axis. For a 2-D array, axis=0 collapses rows, so you get one number per column. axis=1 collapses columns, so you get one number per row.
import numpy as np
scores = np.array([
[70, 80, 90],
[60, 75, 85],
])
print(scores.mean())
print(scores.mean(axis=0))
print(scores.mean(axis=1))
Output:
76.66666666666667
[65. 77.5 87.5]
[80. 73.33333333]
The overall mean is one number. The column means are the average of the two students on each assignment. The row means are each student's average.
keepdims=True keeps the reduced axis as length 1, which makes broadcasting the mean back onto the original array straightforward:
import numpy as np
scores = np.array([
[70, 80, 90],
[60, 75, 85],
])
centered = scores - scores.mean(axis=1, keepdims=True)
print(centered)
Output:
[[-10. 0. 10. ]
[-13.33333333 1.66666667 11.66666667]]
Reshape does not copy the story, it rewraps it
reshape rereads the same elements in order. The new shape must hold the same number of elements.
import numpy as np
flat = np.arange(6)
print(flat.reshape(2, 3))
print(flat.reshape(3, -1))
Output:
[[0 1 2]
[3 4 5]]
[[0 1]
[2 3]
[4 5]]
-1 means "figure this axis out". reshape(3, -1) on 6 elements is 3 rows and 2 columns.
Stacking joins arrays
import numpy as np
a = np.array([[1, 2]])
b = np.array([[3, 4]])
print(np.vstack([a, b]))
print(np.hstack([a, b]))
Output:
[[1 2]
[3 4]]
[[1 2 3 4]]
vstack adds rows. hstack adds columns. The axes you are not stacking along have to match.
What to notice
- If a mean has the wrong length, the axis is backwards. Swap 0 and 1 before you change the data.
reshapefails when the element count changes. Fix the count, not the error message.- Centering a row needs
keepdims=Trueor an explicit reshape to(n, 1). Without it, broadcasting aligns the mean with columns and you subtract the wrong numbers.
Where to practice
Open the NumPy notebooks and run the array-math practice, then the Iris exercise. The lab already has NumPy, and the Iris file is on disk as iris.csv.
After NumPy, the Pandas course puts these arrays into labeled tables.