Fix double-encoded characters: é instead of é

If your text shows é instead of é or ’ instead of ’, it wasn't converted wrongly once but twice. Mojibuster undoes both steps – even in text that is only partly broken.

Works with text and these files:

  • CSV
  • TXT
  • JSON
  • XML
  • SRT
  • VTT
  • TSV
  • MD
  • HTML
  • SQL
  • LOG

For example Excel exports, database dumps, subtitles or website code.

Repair

  1. Drop a file or paste text
  2. It gets repaired automatically
  3. Copy or download
or simply drag a file here

What double-encoded text looks like

A sentence converted wrongly twice

Before

The café’s naïve crème – open Monday

After

The café’s naïve crème – open Monday

How it happens

In UTF-8, é consists of two bytes. When a program reads those bytes as Windows-1252, you get two characters: é. If it saves that result as UTF-8 again, one character has become four bytes. When the same thing happens a second time, each of those characters is split up again – é turns into é.

Typical triggers for the second round:

  • A database is migrated and the WordPress or SQL export was already broken before.
  • A CSV is opened in Excel, saved, and later imported the wrong way once more.
  • A script or API converts text that already is UTF-8.

The most common patterns at a glance

Double-encoded

é
è
ü
ñ
€
’

Correct

é
è
ü
ñ
€
’

Triple encoding makes it even longer: é becomes é. Mojibuster repairs that level automatically, too.

Why find and replace isn't enough

With find and replace you'd have to know every pattern – one for each level and each special character. One is easily missed, or you accidentally replace correct text. Instead, Mojibuster calculates the original bytes and checks whether the result is valid UTF-8. Only then is the spot replaced.

  1. Paste the text above or drop the file into the box.
  2. Mojibuster detects how often the text was converted and repairs up to three levels.
  3. Under What changed? you see every spot. Then copy the result or download the file.

Text that is already correct stays unchanged – even when broken and correct lines are mixed in the same file. Single-level errors like é are covered on Fix UTF-8 characters.

Frequently asked questions

How do I recognize double encoding?

By longer sequences with Ã, †or Â. Single-level errors look shorter, like é.

How many levels can Mojibuster repair?

Up to three. More hardly ever happens in practice.

What if only part of the text is double-encoded?

No problem. Mojibuster repairs each spot on its own and leaves correct parts unchanged.

Can I check the result first?

Yes. Under What changed? Mojibuster lists every repaired spot with its line number before you copy or download.

More guides