Three tools for when Japanese text lands on your desk

You do not have to read Japanese to end up handling it. A customer's name has to go on a form in Latin letters, a CSV from a Japanese system refuses to match your product codes, or an email arrives as ã‚ãªãŸã®åå‰ã¯ï¼Ÿ. These three pages on Omikuji Tools deal with those cases, and all of them process the text in your browser.

Romaji Converter: which spelling of a name?

Romaji is Japanese written in Latin letters, and more than one system for it is in use. The converter offers three. Hepburn follows English pronunciation (shi, chi, tsu, fu, ji) and is what you see on signs and in dictionaries. Kunrei-shiki follows the kana table systematically (si, ti, tu, hu, zi); it is the ISO 3602 system and the one taught in Japanese schools. Passport follows the rules of Japan's Ministry of Foreign Affairs. Long vowels are dropped, so さいとう becomes SAITO. An n before b, m or p becomes M, so 新聞 becomes SHIMBUN. Capital letters are the default for passport output.

With Hepburn and Kunrei-shiki you also choose how long vowels are written (ō or ô, ou, or dropped), and every style lets you pick lower case, Capitalised Words or UPPER CASE. Kana convert directly. Kanji need a reading, and that comes from a morphological dictionary (kuromoji), which your browser downloads once: about 17 MB, and only when the text contains kanji. The page is frank that names and place names can be misread. If you know the correct reading, type it in hiragana instead.

Full-width ⇔ Half-width: the characters that look the same

Japanese text has two widths of Latin letters, digits and symbols. Full-width characters (ABC, 123) take up the same square as a kanji. Half-width ones (ABC, 123) take up half of it. On screen the difference is small, but they are different code points, so 123 does not equal 123 in a search, a lookup or a form validator. Katakana comes in both widths too (アイウ vs アイウ), and the half-width kind still turns up in exports from older systems.

The converter lets you tick exactly what to change — letters, digits, symbols, katakana, spaces — and converts as you type, with a count of how many characters changed. Half-width katakana with voicing marks, such as ガ, are two characters, and they are merged into the single character ガ on the way to full-width. There are two details a non-Japanese reader would not guess: the full-width minus and hyphen both become -, and the full-width yen sign becomes a backslash. Hiragana ⇔ katakana conversion is included as well.

Mojibake Generator & Repair: garbled, but maybe not lost

Mojibake is the Japanese word for garbled text: bytes written in one encoding and read in another. The Garble mode shows how a piece of UTF-8 text comes out when it is read in each of ten encodings: Shift_JIS, EUC-JP, Windows-1252/Latin-1, Windows-1251, KOI8-R, GBK, Big5, EUC-KR, and UTF-16 LE/BE.

Repair runs the process backwards. It tries every candidate and ranks the results by readability. The ã‚ãªãŸ shape is UTF-8 read as Windows-1252. No information was lost there, so it almost always comes back. The honest part is the other case. If the text already contains U+FFFD replacement characters, the page says those bytes are gone and no encoding can recover them, and you need the original file.

👉 Romaji Converter · Full-width ⇔ Half-width · Mojibake Generator & Repair — free, no sign-up, processed in your browser.

Comments

Popular posts from this blog

Amidakuji (Ghost Leg) maker: a fair random pairing tool for up to 200 people, entirely in the browser

omikuji.dev: free browser-based tools, puzzle games and reference sites made in Japan

Hakarasenai: a Firefox extension that only stops Google Analytics from measuring you