Refresher - Constructing DataFrames:
- Recall that a DataFrame represents a rectangular table of data and contains an ordered, named collection of columns, each of which can be a different value type.
- The pd.DataFrame() function can be used to create a DataFrame.
- For example, the data for the DataFrame below is hard coded into a list of lists. Each inner list represents a row of data. The first list contains the column headers, and the subsequent lists contain the actual data.
python
1import pandas as pd
2
3# Hardcode a dataframe
4df = pd.DataFrame(
5 [
6 [
7 'BusinessEntityID', 'EmployeeName', 'EmailAddress',
8 'PhoneNumber', 'StartDate', 'CurrentSalary'
9 ],
10 [
11 1699, 'Robinett, David', 'david22@adventure-works.com',
12 '(827) 525-0100', '06-05-2010', '$80,950'
13 ],
14 [
15 1700, 'Robinson, Rebecca', 'rebecca5 @adventure-works.com',
16 '(829) 525-0101', '05-01-2015', '$70,950'
17 ],
18 [
19 1701, 'Robinson, Dorothy', 'dorothy3@adventure-works.com',
20 '(828) 555-0102', '03-01-2017', '$50,00'
21 ],
22 ]
23)
24
25print(df.shape)
26dfOutput
(4, 6)
| 0 | 1 | 2 | 3 | 4 | 5 | |
|---|---|---|---|---|---|---|
| 0 | BusinessEntityID | EmployeeName | EmailAddress | PhoneNumber | StartDate | CurrentSalary |
| 1 | 1699 | Robinett, David | david22@adventure-works.com | (827) 525-0100 | 06-05-2010 | $80,950 |
| 2 | 1700 | Robinson, Rebecca | rebecca5 @adventure-works.com | (829) 525-0101 | 05-01-2015 | $70,950 |
| 3 | 1701 | Robinson, Dorothy | dorothy3@adventure-works.com | (828) 555-0102 | 03-01-2017 | $50,00 |
Using First Row as Header/Column Names
- If the column names we want are actually the first row of the data, like in the example above, we can:
- Grab the Column Names out of the First Row of the DataFrame using iloc.
- Set the columns property of the DataFrame equal to the column names we extracted in step 1.
- Get rid of the first row of the DataFrame using df[1:].