Question marks instead of characters: what ? and � mean

Instead of é, ü or ’, your text shows ? or the symbol �? They look alike but have different causes. Often the file is still perfectly fine – sometimes the characters are already lost. Here’s how to tell which case you have.

Works with text and these files:

  • CSV
  • TXT
  • JSON
  • XML
  • SRT
  • VTT
  • TSV
  • MD
  • HTML
  • SQL
  • LOG

For example Excel exports, database dumps, subtitles or website code.

Repair

  1. Drop a file or paste text
  2. It gets repaired automatically
  3. Copy or download
or simply drag a file here

Case 1: � in a file that is actually fine

If an editor, player or browser shows Caf�, the file is usually saved in Windows-1252 (ANSI) and read as UTF-8. The byte of the accented letter is invalid in UTF-8, so the program inserts the replacement character � when displaying it. The file itself still contains the right byte.

Solution: Drop the original file into the box above. Mojibuster detects the encoding while reading and converts the file to UTF-8. Important: use the file itself, don’t copy the text out of the program – copying makes the � permanent.

Case 2: � or ? is stored in the text

If a program saved or passed on the text with the replacement character, the original character is lost. The same goes for ?: programs insert it when a character doesn’t exist in the target encoding – for example when saving as ASCII or writing to a database column that has no é.

Nothing can be calculated back from that, because every accented letter became the same character. Mojibuster tells you how many spots are affected instead of guessing. Only the source can help: repeat the export or use a backup.

Typical causes of question marks

  • Database: the column or connection uses latin1 or ascii, while the data arrives as UTF-8. MySQL then replaces unknown characters with ? – see MySQL and phpMyAdmin.
  • PowerShell 5.1: Export-Csv writes ASCII by default and replaces accented letters with ? – see PowerShell.
  • Python: encode('ascii', errors='replace') silently replaces characters with ? – see Python and pandas.
  • Console: the Windows console shows characters its code page doesn’t know as ?. Usually only the display is wrong, not the file.

How to tell which case you have

  1. Open the file in an editor that shows the encoding, such as Notepad++. If the status bar says “ANSI” or “Windows-1252”, it’s case 1.
  2. Drop the file into Mojibuster. If the result says “Read as Windows-1252 (ANSI)” and the characters are back, it was case 1.
  3. If Mojibuster reports lost characters, it’s case 2. Only a new export helps then.

If you see sequences like é instead of question marks, that’s a different error – and it can almost always be repaired: é instead of é.

Frequently asked questions

Can Mojibuster turn ? back into letters?

No. A ? can stand for any letter, the information is gone. Mojibuster doesn’t touch such spots but tells you how many there are.

Why do I see � in one program but not in another?

Then the file is fine, and only one of the programs reads it with the wrong encoding. Convert it to UTF-8 with Mojibuster, and both will show it correctly.

What is the � character?

The replacement character U+FFFD. Programs insert it when they can’t read a byte as a character. It points to an encoding error and isn’t part of the actual text.

More guides