Pandas Series and DataFrame Basics:

  • Pandas introduces two new data types to Python: Series and DataFrame.
    • The DataFrame will represent your entire spreadsheet or rectangular data, whereas the Series is a single column of the DataFrame.
      • In other words, Series are one-dimensional arrays of data, and DataFrames are two-dimensional arrays of data that are roughly similar to a spreadsheet.
      • A Pandas DataFrame can also be thought of as a dictionary or collection of Series.

The pandas DataFrame Anatomy:

  • Visually, the output display of a pandas DataFrame appears to be nothing more than an ordinary table of data consisting of rows and columns.
  • Hiding beneath the surface are the three components - the index, columns, and data that you must be aware of to maximize the DataFrame’s full potential1.
  • The below reads in the movie dataset into a pandas DataFrame and provides a labeled diagram of all its major components.
python
1import pandas as pd
2
3df = pd.read_csv('movie.csv')
4print(df)

Anatomy:

Labeled diagram of a pandas DataFrame showing index on axis 0, columns on axis 1, and data cells with NaN for missing values

  • The labels in index and column names allow for pulling out data based on the index and column name.
  • Index is axis 0 and the columns are axis 1.
  • Pandas uses NaN (not a number) to represent missing values.

Footnotes

  1. McKinney, W. (2022). Python for data analysis: Data wrangling with pandas, NumPy, and Jupyter (3rd ed.). O’Reilly Media. ↩