Fix character encoding in Python and pandas

pandas reads é instead of é, stops with a UnicodeDecodeError, or Excel shows your exported CSV wrongly? In Python, the encoding parameter decides between right and broken. Here are the typical cases – and you can repair broken files right in your browser above.

Works with text and these files:

  • CSV
  • TXT
  • JSON
  • XML
  • SRT
  • VTT
  • TSV
  • MD
  • HTML
  • SQL
  • LOG

For example Excel exports, database dumps, subtitles or website code.

Repair

  1. Drop a file or paste text
  2. It gets repaired automatically
  3. Copy or download
or simply drag a file here

UnicodeDecodeError when reading

pd.read_csv expects UTF-8. If the file is saved in Windows-1252 (ANSI), pandas stops with UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 – 0xE9 is é in Windows-1252. Then state the encoding:

  • pd.read_csv('data.csv', encoding='cp1252') for files from Excel and Windows.
  • encoding='latin-1' never fails, but reads € and typographic quotes wrongly – cp1252 is usually the better choice.
  • If you don’t know the encoding, drop the file into Mojibuster: it detects it and saves the file as UTF-8.

é instead of é in the DataFrame

If the DataFrame shows é, a UTF-8 file was read as latin-1 or cp1252 – or the data was broken before. In the first case, encoding='utf-8' is enough. If the data itself is broken, value.encode('cp1252').decode('utf-8') helps for single values. With mixed data, though, it fails as soon as a cell is already correct. Easier: repair the file with Mojibuster, then read it.

A CSV from pandas, opened in Excel without a BOM

Before

name,city,revenue
Renée Dubois,Montréal,1250 €
José García,São Paulo,980 €

After

name,city,revenue
Renée Dubois,Montréal,1250 €
José García,São Paulo,980 €

to_csv: Excel shows the file wrongly

df.to_csv('export.csv') writes UTF-8 without a BOM. Excel opens such files as ANSI, and accented letters turn into é. With encoding='utf-8-sig', pandas writes a BOM and Excel recognizes UTF-8. More on Excel showing special characters wrong.

open() on Windows

open('file.txt') without encoding uses the system encoding on Windows, usually cp1252. A script that works on Linux then reads the same UTF-8 file on Windows as é. So always pass encoding='utf-8' or set the environment variable PYTHONUTF8=1. With PEP 686, UTF-8 becomes the default in future Python versions anyway.

Web data with requests

requests derives response.text from the Content-Type header. If a text response doesn’t state a character set there, it assumes ISO-8859-1, and UTF-8 pages show é. Set response.encoding = 'utf-8' first, or use response.json(), which detects UTF-8 by itself.

If text was written with errors='replace' into an encoding without the needed characters, the file contains ? – then the information is lost, see question marks instead of characters.

Frequently asked questions

Is there a Python library for this?

Yes, ftfy fixes mojibake in Python code. For single files without programming, Mojibuster is quicker and shows you every change before saving.

Which encoding values do I need most?

utf-8, utf-8-sig (UTF-8 with BOM), cp1252 (Windows, Western European), latin-1 and utf-16. Case and hyphens don’t matter.

Are delimiters and numbers preserved?

Yes. Mojibuster only changes broken characters. Commas, quotes and line breaks stay exactly as they are.

More guides