UnicodeDecodeError when reading
pd.read_csv expects UTF-8. If the file is saved in Windows-1252 (ANSI), pandas stops with UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 – 0xE9 is é in Windows-1252. Then state the encoding:
pd.read_csv('data.csv', encoding='cp1252')for files from Excel and Windows.encoding='latin-1'never fails, but reads€and typographic quotes wrongly –cp1252is usually the better choice.- If you don’t know the encoding, drop the file into Mojibuster: it detects it and saves the file as UTF-8.
é instead of é in the DataFrame
If the DataFrame shows é, a UTF-8 file was read as latin-1 or cp1252 – or the data was broken before. In the first case, encoding='utf-8' is enough. If the data itself is broken, value.encode('cp1252').decode('utf-8') helps for single values. With mixed data, though, it fails as soon as a cell is already correct. Easier: repair the file with Mojibuster, then read it.
Before
name,city,revenue Renée Dubois,Montréal,1250 € José GarcÃa,São Paulo,980 €
After
name,city,revenue Renée Dubois,Montréal,1250 € José García,São Paulo,980 €
to_csv: Excel shows the file wrongly
df.to_csv('export.csv') writes UTF-8 without a BOM. Excel opens such files as ANSI, and accented letters turn into é. With encoding='utf-8-sig', pandas writes a BOM and Excel recognizes UTF-8. More on Excel showing special characters wrong.
open() on Windows
open('file.txt') without encoding uses the system encoding on Windows, usually cp1252. A script that works on Linux then reads the same UTF-8 file on Windows as é. So always pass encoding='utf-8' or set the environment variable PYTHONUTF8=1. With PEP 686, UTF-8 becomes the default in future Python versions anyway.
Web data with requests
requests derives response.text from the Content-Type header. If a text response doesn’t state a character set there, it assumes ISO-8859-1, and UTF-8 pages show é. Set response.encoding = 'utf-8' first, or use response.json(), which detects UTF-8 by itself.
If text was written with errors='replace' into an encoding without the needed characters, the file contains ? – then the information is lost, see question marks instead of characters.
Frequently asked questions
Is there a Python library for this?
Yes, ftfy fixes mojibake in Python code. For single files without programming, Mojibuster is quicker and shows you every change before saving.
Which encoding values do I need most?
utf-8, utf-8-sig (UTF-8 with BOM), cp1252 (Windows, Western European), latin-1 and utf-16. Case and hyphens don’t matter.
Are delimiters and numbers preserved?
Yes. Mojibuster only changes broken characters. Commas, quotes and line breaks stay exactly as they are.