PANDAS:Series, DataFrames, and when to use them
Mastering series, dataframes, and when to use them concepts and implementation.
The hook
A NumPy array knows positions. It does not know that column 0 is a species name and column 4 is a body mass. Pandas adds those labels. A Series is one labeled column. A DataFrame is a table of columns that share an index.
Use NumPy when every value is a number and you care about shape. Use Pandas when the data has column names, mixed types, or missing entries. You can always pull a numeric column back out as an array with .to_numpy().
Install Pandas with pip install pandas. In the data analysis lab it is already available. Import it as pd.
A Series
import pandas as pd
mass = pd.Series([3750, 3800, 3250], name="body_mass_g")
print(mass)
print(mass.mean())
Output:
0 3750
1 3800
2 3250
Name: body_mass_g, dtype: int64
3600.0
The left numbers are the index, not a second column. The default index is 0, 1, 2, ....
A DataFrame
import pandas as pd
birds = pd.DataFrame({
"species": ["Adelie", "Adelie", "Chinstrap"],
"island": ["Torgersen", "Biscoe", "Dream"],
"body_mass_g": [3750, 3800, 3250],
})
print(birds)
print(birds.shape)
print(birds.columns.tolist())
Output:
species island body_mass_g
0 Adelie Torgersen 3750
1 Adelie Biscoe 3800
2 Chinstrap Dream 3250
(3, 3)
['species', 'island', 'body_mass_g']
Each key in the dictionary becomes a column. Columns in one DataFrame can have different dtypes: text in species, integers in body_mass_g.
What to notice
shapeis still(rows, columns), the same idea as NumPy.- A Series has one dtype. A DataFrame has one dtype per column.
- Building a small DataFrame by hand is how you test an idea before you load a file.
Try this
Make a three-row DataFrame with columns name, score, and passed (passed should be a boolean). Print .dtypes and the mean of score only.
Next: reading a real CSV and the first questions you ask a table.