Hindi, Tamil or Bengali text shows ठcharacters? Fix it here

Your Hindi text shows हिंदी instead of हिंदी? Then UTF-8 text was read as Windows-1252 – the letters aren’t lost. Paste the text above or drop in the CSV file, and Mojibuster turns it back into readable Hindi, Tamil, Bengali or any other Indian script.

Works with text and these files:

  • CSV
  • TXT
  • JSON
  • XML
  • SRT
  • VTT
  • TSV
  • MD
  • HTML
  • SQL
  • LOG

For example Excel exports, database dumps, subtitles or website code.

Repair

  1. Drop a file or paste text
  2. It gets repaired automatically
  3. Copy or download
or simply drag a file here

What the broken text looks like

Every Indian script has its own fingerprint. Each letter becomes three characters, and the first two are always the same pair for a script – so you can tell the language even from the broken text:

  • Hindi, Marathi, Nepali (Devanagari): ठand à¥, e.g. नाम for नाम.
  • Bengali and Assamese: ঠand à§, e.g. বাংলা for বাংলা.
  • Tamil: à® and à¯, Telugu à° and à±, Kannada ಠand à³, Malayalam à´ and àµ.
  • Gujarati: ઠand à«, Punjabi (Gurmukhi) ਠand à©.
  • Rupee sign: ₹ instead of ₹. The danda । becomes ।.
Customer list from a CSV file, opened in Excel

Broken

नाम,शहर,भाषा
सीता,कोलकाता,বাংলা
राम,भारत,हिंदी
கோவை,മലയാളം,₹1499

Repaired

नाम,शहर,भाषा
सीता,कोलकाता,বাংলা
राम,भारत,हिंदी
கோவை,മലയാളം,₹1499

Why it happens

In UTF-8, a letter of an Indian script is stored as three bytes – ह is E0 A4 B9. A program that expects Windows-1252, the old Western character table, shows one character per byte: à, ¤ and ¹. The bytes are still all there, they are only read the wrong way. That is why Mojibuster can turn them back into the right letters.

The most common source is Excel: a CSV exported from Google Sheets, a web shop, a CRM or a government portal is UTF-8 without a BOM, the invisible marker at the start of a file. Excel then guesses Windows-1252. Older websites, WordPress databases with latin1 tables and e-mails with a wrong charset cause the same error.

How to fix a CSV file for Excel

  1. Drop the CSV file into the box above. Mojibuster repairs the broken letters and shows a preview.
  2. Keep Save with BOM switched on – it’s preselected for CSV files – and click Download.
  3. Double-click the downloaded file – Excel now shows Hindi, Tamil or Bengali correctly.

If the file itself is fine and only Excel shows it wrongly, you can also import it: Data → From Text/CSV, then choose 65001: Unicode (UTF-8) under File Origin. More on this: fix special characters in Excel and fix CSV encoding.

How to save CSV files with Indian scripts correctly

  • In Excel, choose Save as → CSV UTF-8 (Comma delimited). The plain “CSV (Comma delimited)” type uses the old Windows character table and replaces Hindi letters with ?.
  • In Google Sheets, File → Download → CSV gives UTF-8 without a BOM. Run the file through Mojibuster with the BOM option before you open it in Excel.
  • In databases, use utf8mb4 for tables and connections – see MySQL and phpMyAdmin encoding.

When the text can’t be repaired

If Excel or another program saved the file in the old Western character table, every Hindi letter was replaced by ?. The information is gone and can’t be restored – Mojibuster counts these spots so you know where to look. Go back to the original source and export it again as UTF-8. More: question marks instead of characters.

Some bytes – for example the virama ् and the vowel sign ु – turn into invisible control characters when read as Windows-1252. If a program removed them, Mojibuster works out the missing letter from its neighbours. Where both readings are possible, it leaves ॠin place instead of guessing – repair the original file if you can: it still contains every byte.

Kruti Dev, Chanakya and other legacy fonts

Text typed in Kruti Dev, Chanakya, DevLys or Bamini looks like fgUnh when the font is missing – that is not an encoding error but a font that maps Latin letters to Hindi shapes. Mojibuster doesn’t convert these fonts yet. It repairs text that was saved as Unicode (UTF-8) and read the wrong way.

Background: what is mojibake? – and if the text shows à along with à¤, it was converted twice: fix double-encoded UTF-8.

Common broken characters

What garbled characters typically look like – and what they should say.
Broken Correct
हिंदीहिंदी
मराठीमराठी
भारतभारत
कोलकाताकोलकाता
বাংলাবাংলা
ঢাকাঢাকা
கோவைகோவை
తెలంగాణతెలంగాణ
ಭಾರತಭಾರತ
અમદાવાદઅમદાવાદ
മലയാളംമലയാളം
ਪੰਜਾਬੀਪੰਜਾਬੀ
₹₹

Frequently asked questions

Which Indian languages are supported?

All scripts in Unicode: Hindi, Marathi, Nepali, Bengali, Assamese, Punjabi, Gujarati, Odia, Tamil, Telugu, Kannada, Malayalam, Sinhala and Urdu. English, numbers and punctuation in the same text stay as they are.

Can I convert Kruti Dev to Unicode here?

Not yet. Kruti Dev is a font, not an encoding, and needs its own converter. Mojibuster repairs Unicode text that shows ठor ஠characters.

Will the CSV still open in Excel after the repair?

Yes. Commas, quotes and line breaks stay unchanged. Download the file with the BOM option, and Excel opens it as UTF-8.

Is my file uploaded?

No. The repair runs entirely in your browser, also on a phone. The file never leaves your device.

More guides