I found some code online that deletes a specific column by name using pandas:
# import pandas with shortcut 'pd'
import pandas as pd
# read_csv function which is used to read the required CSV file
data = pd.read_csv('TradedInstrument_20230331_test.csv')
# drop function which is used in removing or deleting rows or columns from the CSV files
data.drop('isin', inplace=True, axis=1)
Any ideas if this actually works for a tab seperated csv?
If I run the above against my file, I get a UnicodeDecodeError:
Traceback (most recent call last):
....
File "pandas/_libs/parsers.pyx", line 1917, in pandas._libs.parsers.raise_parser_error
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfc in position 38835: invalid start byte
>Solution :
You have an encoding error. It looks like 0xfc is ü character (latin small letter u with diaeresis), so try to use encoding='latin1' when your read your csv file:
data = pd.read_csv('TradedInstrument_20230331_test.csv', sep='\t', encoding='latin1')
data.drop(columns='isin').to_csv('output.csv', index=False)