Follow

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Contact

How can I delete a specific column in a tab seperated csv?

I found some code online that deletes a specific column by name using pandas:

# import pandas with shortcut 'pd'
import pandas as pd  
  
# read_csv function which is used to read the required CSV file
data = pd.read_csv('TradedInstrument_20230331_test.csv')
  
  
# drop function which is used in removing or deleting rows or columns from the CSV files
data.drop('isin', inplace=True, axis=1)
  

Any ideas if this actually works for a tab seperated csv?

If I run the above against my file, I get a UnicodeDecodeError:

MEDevel.com: Open-source for Healthcare and Education

Collecting and validating open-source software for healthcare, education, enterprise, development, medical imaging, medical records, and digital pathology.

Visit Medevel

Traceback (most recent call last):
  ....
  File "pandas/_libs/parsers.pyx", line 1917, in pandas._libs.parsers.raise_parser_error
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfc in position 38835: invalid start byte

>Solution :

You have an encoding error. It looks like 0xfc is ü character (latin small letter u with diaeresis), so try to use encoding='latin1' when your read your csv file:

data = pd.read_csv('TradedInstrument_20230331_test.csv', sep='\t', encoding='latin1')
data.drop(columns='isin').to_csv('output.csv', index=False)
Add a comment

Leave a Reply

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Discover more from Dev solutions

Subscribe now to keep reading and get access to the full archive.

Continue reading