removing duplicates rows of a dataframe python

Question

I need to remove duplicate rows from a dataset. Basically, I should perform

proc sort data=mydata noduprecs dupout=mydata_dup;run;

I need to remove duplicates as well as save those duplicate rows in a separate dataframe. How can I do that?

Bubble Bubble Bubble Gut · Accepted Answer · 2017-07-06 18:02:44Z

2

Assuming your dataset is a pandas dataframe.

To remove the duplicated rows:

data = data.drop_duplicates()

To select all the duplicated rows:

dup = data.ix[data.duplicated(), :]

Hope it helps.

answered Jul 6, 2017 at 18:02

Bubble Bubble Bubble Gut

3,37617 silver badges30 bronze badges

Sign up to request clarification or add additional context in comments.

Comments

Antonio · Accepted Answer · 2021-02-26 22:37:30Z

A few examples from Pandas docs:

> df = pd.DataFrame({

    'brand': ['Yum Yum', 'Yum Yum', 'Indomie', 'Indomie', 'Indomie'],

    'style': ['cup', 'cup', 'cup', 'pack', 'pack'],

    'rating': [4, 4, 3.5, 15, 5]

})

> df
    brand style  rating
0  Yum Yum   cup     4.0
1  Yum Yum   cup     4.0
2  Indomie   cup     3.5
3  Indomie  pack    15.0
4  Indomie  pack     5.0

By default, it removes duplicate rows based on all columns.

> df.drop_duplicates()
    brand style  rating
0  Yum Yum   cup     4.0
2  Indomie   cup     3.5
3  Indomie  pack    15.0
4  Indomie  pack     5.0

To remove duplicates on specific column(s), use subset.

> df.drop_duplicates(subset=['brand'])
    brand style  rating
0  Yum Yum   cup     4.0
2  Indomie   cup     3.5

To remove duplicates and keep last occurrences, use keep.

> df.drop_duplicates(subset=['brand', 'style'], keep='last')
    brand style  rating
1  Yum Yum   cup     4.0
2  Indomie   cup     3.5
4  Indomie  pack     5.0

Collectives™ on Stack Overflow

removing duplicates rows of a dataframe python

2 Answers 2

Comments

Comments

Your Answer

Hot Network Questions

Collectives™ on Stack Overflow

2 Answers 2

Comments

Comments

Your Answer

Sign up or log in

Post as a guest

Related