PANDAS:Series, DataFrames, and when to use them

Mastering series, dataframes, and when to use them concepts and implementation.

The hook

A NumPy array knows positions. It does not know that column 0 is a species name and column 4 is a body mass. Pandas adds those labels. A Series is one labeled column. A DataFrame is a table of columns that share an index.

Use NumPy when every value is a number and you care about shape. Use Pandas when the data has column names, mixed types, or missing entries. You can always pull a numeric column back out as an array with .to_numpy().

Install Pandas with pip install pandas. In the data analysis lab it is already available. Import it as pd.

A Series

import pandas as pd

mass = pd.Series([3750, 3800, 3250], name="body_mass_g")
print(mass)
print(mass.mean())

Output:

0    3750
1    3800
2    3250
Name: body_mass_g, dtype: int64
3600.0

The left numbers are the index, not a second column. The default index is 0, 1, 2, ....

A DataFrame

import pandas as pd

birds = pd.DataFrame({
    "species": ["Adelie", "Adelie", "Chinstrap"],
    "island": ["Torgersen", "Biscoe", "Dream"],
    "body_mass_g": [3750, 3800, 3250],
})
print(birds)
print(birds.shape)
print(birds.columns.tolist())

Output:

     species     island  body_mass_g
0     Adelie  Torgersen         3750
1     Adelie     Biscoe         3800
2  Chinstrap      Dream         3250
(3, 3)
['species', 'island', 'body_mass_g']

Each key in the dictionary becomes a column. Columns in one DataFrame can have different dtypes: text in species, integers in body_mass_g.

What to notice

  • shape is still (rows, columns), the same idea as NumPy.
  • A Series has one dtype. A DataFrame has one dtype per column.
  • Building a small DataFrame by hand is how you test an idea before you load a file.

Try this

Make a three-row DataFrame with columns name, score, and passed (passed should be a boolean). Print .dtypes and the mean of score only.

Next: reading a real CSV and the first questions you ask a table.