PANDAS:Reading a CSV and inspecting a table

Mastering reading a csv and inspecting a table concepts and implementation.

Load, then look

read_csv is the front door. It guesses dtypes, builds an index, and turns the header into column names. Before you filter or plot, print a few facts so you know what arrived.

The lab ships Palmer penguins as penguins.csv. The same file is the running example in this chapter.

import pandas as pd

penguins = pd.read_csv("penguins.csv")
print(penguins.shape)
print(penguins.columns.tolist())
print(penguins.head(3))

Output:

(344, 7)
['species', 'island', 'bill_length_mm', 'bill_depth_mm', 'flipper_length_mm', 'body_mass_g', 'sex']
  species     island  bill_length_mm  ...  flipper_length_mm  body_mass_g   sex
0  Adelie  Torgersen            39.1  ...              181.0       3750.0  MALE
1  Adelie  Torgersen            39.5  ...              186.0       3800.0  FEMALE
2  Adelie  Torgersen            40.3  ...              195.0       3250.0  FEMALE

344 rows, 7 columns. head(3) shows the first three rows and abbreviates wide tables with ....

The inspection set

import pandas as pd

penguins = pd.read_csv("penguins.csv")
print(penguins.dtypes)
print(penguins["species"].value_counts())
print(penguins.isna().sum())

Output:

species               object
island                object
bill_length_mm       float64
bill_depth_mm        float64
flipper_length_mm    float64
body_mass_g          float64
sex                   object
dtype: object

species
Adelie       152
Gentoo       124
Chinstrap     68
Name: count, dtype: int64

species               0
island                0
bill_length_mm        2
bill_depth_mm         2
flipper_length_mm     2
body_mass_g           2
sex                  11
dtype: int64

Three species. Two rows are missing every measurement. Eleven rows are missing sex. You want to see that before a mean quietly skips those rows, or before a merge surprises you.

info() prints the same story in one block: row count, non-null counts, and dtypes.

What to notice

  • read_csv looks for the file relative to the working directory. In the lab, the CSV names are penguins.csv, iris.csv, and anscombe.csv.
  • value_counts() is for categories. describe() is for numbers. Use the one that matches the column.
  • isna().sum() counts missing values per column. A 0 there means that column is complete.

Try this

Load iris.csv, print the shape, and count how many rows belong to each species.

Next: choosing rows and columns without copying the whole table into your head.