Dropping a Column:
- To drop a column from a DataFrame, we can use the drop() method.
- The drop() method accepts an argument that specifies the column to drop.
- The inplace parameter can be set to True to indicate that the DataFrame should be modified in place.
- In the following example, we drop the Security column from the DataFrame df.
python
1import pandas as pd
2
3# Read the CSV file into a DataFrame
4df = pd.read_csv("Data/Demo_3.8_P1_SP500_Stocks.csv")
5
6# Drop the 'Security' column
7df.drop(columns=['Security'], inplace=True)
8
9# Display the DataFrame
10print(df.shape)
11df.head()Output:
(503, 7)
| Symbol | GICS Sector | GICS Sub-Industry | Headquarters Location | Date added | CIK | Founded | |
|---|---|---|---|---|---|---|---|
| 0 | MMM | Industrials | Industrial Conglomerates | Saint Paul, Minnesota | 1957-03-04 | 66740 | 1902 |
| 1 | AOS | Industrials | Building Products | Milwaukee, Wisconsin | 2017-07-26 | 91142 | 1916 |
| 2 | ABT | Health Care | Health Care Equipment | North Chicago, Illinois | 1957-03-04 | 1800 | 1888 |
| 3 | ABBV | Health Care | Biotechnology | North Chicago, Illinois | 2012-12-31 | 1551152 | 2013 (1888) |
| 4 | ACN | Information Technology | IT Consulting & Other Services | Dublin, Ireland | 2011-07-06 | 1467373 | 1989 |
Renaming a Column:
- The rename() method can be used to rename a column in a DataFrame.
- The method accepts an argument that specifies the column to rename.
- The columns parameter can be set to a dictionary that maps the old column name to the new column name.
- The inplace parameter can be set to True to indicate that the DataFrame should be modified in place.
- In the following example, we rename the GICS Sub-Industry column to Sub-Industry in the DataFrame df.
python
12df.rename(columns={'GICS Sub-Industry': 'Sub-Industry'}, inplace=True)
13print(df.shape)
14df.head()Output:
(503, 7)
| Symbol | GICS Sector | Sub-Industry | Headquarters Location | Date added | CIK | Founded | |
|---|---|---|---|---|---|---|---|
| 0 | MMM | Industrials | Industrial Conglomerates | Saint Paul, Minnesota | 1957-03-04 | 66740 | 1902 |
| 1 | AOS | Industrials | Building Products | Milwaukee, Wisconsin | 2017-07-26 | 91142 | 1916 |
| 2 | ABT | Health Care | Health Care Equipment | North Chicago, Illinois | 1957-03-04 | 1800 | 1888 |
| 3 | ABBV | Health Care | Biotechnology | North Chicago, Illinois | 2012-12-31 | 1551152 | 2013 (1888) |
| 4 | ACN | Information Technology | IT Consulting & Other Services | Dublin, Ireland | 2011-07-06 | 1467373 | 1989 |
Removing Leading and Trailing Spaces with strip()
- Suppose some of the values in the columns have uninteded leading or trailing spaces. We can use the strip() method to clean these values.
- The strip() method removes leading and trailing spaces from a string.
python
15df.columns = df.columns.str.strip()replace()
- The replace() method can be used with a regular expression to remove errant characters.
- As shown below, we can use the replace() method to remove certain characters. First, we’ll introduce some junk data for demonstration purposes.
python
16# Introduce some junk data for demonstration
17df.loc[0, 'Sector'] = 'Industrials!'
18df.loc[1, 'Sector'] = '#Industrials'
19
20# Display the values
21print("Values in 'Sector' column:")
22print(df['Sector'].head())Output
1Values in 'Sector' column:
20 Industrials!
31 #Industrials
42 Health Care
53 Health Care
64 Information Technology- Notice that some of the values in the Sector column now have some unwanted characters. We can use the replace() method to remove these characters.
- In the following example, we remove the characters # and ! from the Sector column.
- The replace() method accepts a dictionary that maps the characters to be replaced to the replacement characters.
- The regex parameter can be set to True to indicate that the characters to be replaced are regular expressions.
- In regular expressions, the dollar sign ($) has a special meaning: it represents the end of a line. To specify that we are referring to the literal dollar sign character rather than its special meaning, we need to escape it with two backslashes.
- In the following example, we remove the characters # and ! from the Sector column.
python
23df['Sector'] = df['Sector'].replace({'#': '', '!': ''}, regex=True)
24# .replace({'\\$': '',',': ''}, regex=True) Would remove the dollar sign and comma characters from a string, for example.
25print(df.shape)
26df.head()(503, 7)
| Symbol | Sector | GICS Sub-Industry | Headquarters Location | Date added | CIK | Founded | |
|---|---|---|---|---|---|---|---|
| 0 | MMM | Industrials | Industrial Conglomerates | Saint Paul, Minnesota | 1957-03-04 | 66740 | 1902 |
| 1 | AOS | Industrials | Building Products | Milwaukee, Wisconsin | 2017-07-26 | 91142 | 1916 |
| 2 | ABT | Health Care | Health Care Equipment | North Chicago, Illinois | 1957-03-04 | 1800 | 1888 |
| 3 | ABBV | Health Care | Biotechnology | North Chicago, Illinois | 2012-12-31 | 1551152 | 2013 (1888) |
| 4 | ACN | Information Technology | IT Consulting & Other Services | Dublin, Ireland | 2011-07-06 | 1467373 | 1989 |