MATPLOTLIB:Subplots for a small comparison
Mastering subplots for a small comparison concepts and implementation.
One page, several panels
plt.subplots(nrows, ncols) returns a figure and an array of axes. Each axes is a panel with its own title. Share an axis when the panels must be compared on the same scale.
import matplotlib.pyplot as plt
import pandas as pd
anscombe = pd.read_csv("anscombe.csv")
fig, axes = plt.subplots(2, 2, figsize=(7, 6), sharex=True, sharey=True)
datasets = ["I", "II", "III", "IV"]
for ax, name in zip(axes.flat, datasets):
group = anscombe.loc[anscombe["dataset"] == name]
ax.scatter(group["x"], group["y"], color="#6b21a8")
ax.set_title(f"Dataset {name}")
fig.suptitle("Anscombe's quartet")
fig.tight_layout()
axes is a 2 by 2 array. axes.flat walks the panels in reading order. sharex and sharey lock the scales, which is the whole point of this chart: the four clouds occupy the same frame and still look nothing alike. Their means and correlations are almost the same. The plot is the argument.
A row of histograms
import matplotlib.pyplot as plt
import pandas as pd
iris = pd.read_csv("iris.csv")
columns = ["sepal_length", "sepal_width", "petal_length", "petal_width"]
fig, axes = plt.subplots(1, 4, figsize=(12, 3))
for ax, column in zip(axes, columns):
ax.hist(iris[column], bins=12, color="#0f766e", edgecolor="white")
ax.set_title(column.replace("_", " "))
fig.tight_layout()
When nrows and ncols are both greater than 1, axes is 2-D and you need .flat or axes[row, col]. A single row returns a 1-D array, so zip(axes, columns) works directly.
What to notice
sharey=Trueis how you stop a small group from looking as spread out as a large one just because its axis zoomed in.fig.suptitleis the title of the page.ax.set_titleis the title of one panel. Use both when the page has a single question and each panel is a case.zipstops at the shorter list. If you have more panels than series, the extra panels stay empty, which is a signal to check the loop.
Where to practice
Open the Matplotlib notebooks. The Anscombe exercise asks you to build this four-panel figure from anscombe.csv. The penguins scatter is the one to run when you want a chart that comes from a real table.
The pieces connect in that order: NumPy for the numbers, Pandas for the table, Matplotlib for the picture.