Dropping a Column:

  • To drop a column from a DataFrame, we can use the drop() method.
    • The drop() method accepts an argument that specifies the column to drop.
    • The inplace parameter can be set to True to indicate that the DataFrame should be modified in place.
  • In the following example, we drop the Security column from the DataFrame df.
python
1import pandas as pd
2
3# Read the CSV file into a DataFrame
4df = pd.read_csv("Data/Demo_3.8_P1_SP500_Stocks.csv")
5
6# Drop the 'Security' column
7df.drop(columns=['Security'],  inplace=True)
8
9# Display the DataFrame
10print(df.shape)
11df.head()

Output:

(503, 7)

SymbolGICS SectorGICS Sub-IndustryHeadquarters LocationDate addedCIKFounded
0MMMIndustrialsIndustrial ConglomeratesSaint Paul, Minnesota1957-03-04667401902
1AOSIndustrialsBuilding ProductsMilwaukee, Wisconsin2017-07-26911421916
2ABTHealth CareHealth Care EquipmentNorth Chicago, Illinois1957-03-0418001888
3ABBVHealth CareBiotechnologyNorth Chicago, Illinois2012-12-3115511522013 (1888)
4ACNInformation TechnologyIT Consulting & Other ServicesDublin, Ireland2011-07-0614673731989

Renaming a Column:

  • The rename() method can be used to rename a column in a DataFrame.
    • The method accepts an argument that specifies the column to rename.
    • The columns parameter can be set to a dictionary that maps the old column name to the new column name.
    • The inplace parameter can be set to True to indicate that the DataFrame should be modified in place.
  • In the following example, we rename the GICS Sub-Industry column to Sub-Industry in the DataFrame df.
python
12df.rename(columns={'GICS Sub-Industry': 'Sub-Industry'}, inplace=True)
13print(df.shape)
14df.head()

Output:

(503, 7)

SymbolGICS SectorSub-IndustryHeadquarters LocationDate addedCIKFounded
0MMMIndustrialsIndustrial ConglomeratesSaint Paul, Minnesota1957-03-04667401902
1AOSIndustrialsBuilding ProductsMilwaukee, Wisconsin2017-07-26911421916
2ABTHealth CareHealth Care EquipmentNorth Chicago, Illinois1957-03-0418001888
3ABBVHealth CareBiotechnologyNorth Chicago, Illinois2012-12-3115511522013 (1888)
4ACNInformation TechnologyIT Consulting & Other ServicesDublin, Ireland2011-07-0614673731989

Removing Leading and Trailing Spaces with strip()

  • Suppose some of the values in the columns have uninteded leading or trailing spaces. We can use the strip() method to clean these values.
    • The strip() method removes leading and trailing spaces from a string.
python
15df.columns = df.columns.str.strip()

replace()

  • The replace() method can be used with a regular expression to remove errant characters.
    • As shown below, we can use the replace() method to remove certain characters. First, we’ll introduce some junk data for demonstration purposes.
python
16# Introduce some junk data for demonstration
17df.loc[0, 'Sector'] = 'Industrials!'
18df.loc[1, 'Sector'] = '#Industrials'
19
20# Display the values
21print("Values in 'Sector' column:")
22print(df['Sector'].head())

Output

1Values in 'Sector' column:
20              Industrials!
31              #Industrials
42               Health Care
53               Health Care
64    Information Technology
  • Notice that some of the values in the Sector column now have some unwanted characters. We can use the replace() method to remove these characters.
    • In the following example, we remove the characters # and ! from the Sector column.
      • The replace() method accepts a dictionary that maps the characters to be replaced to the replacement characters.
      • The regex parameter can be set to True to indicate that the characters to be replaced are regular expressions.
        • In regular expressions, the dollar sign ($) has a special meaning: it represents the end of a line. To specify that we are referring to the literal dollar sign character rather than its special meaning, we need to escape it with two backslashes.
python
23df['Sector'] = df['Sector'].replace({'#': '', '!': ''}, regex=True) 
24# .replace({'\\$': '',',': ''}, regex=True) Would remove the dollar sign and comma characters from a string, for example.
25print(df.shape)
26df.head()

(503, 7)

SymbolSectorGICS Sub-IndustryHeadquarters LocationDate addedCIKFounded
0MMMIndustrialsIndustrial ConglomeratesSaint Paul, Minnesota1957-03-04667401902
1AOSIndustrialsBuilding ProductsMilwaukee, Wisconsin2017-07-26911421916
2ABTHealth CareHealth Care EquipmentNorth Chicago, Illinois1957-03-0418001888
3ABBVHealth CareBiotechnologyNorth Chicago, Illinois2012-12-3115511522013 (1888)
4ACNInformation TechnologyIT Consulting & Other ServicesDublin, Ireland2011-07-0614673731989