Full-width ⇄ Half-width Converter
Convert full-width letters, digits, symbols and katakana to half-width and back. Essential for cleaning data exported from Japanese systems.
Convert between the full-width and half-width forms of Latin letters, digits, symbols and katakana, choosing which categories to touch. This is the fix for data that arrives from Japanese systems with numbers and letters in their wide forms.
// Full-width ⇄ Half-width Converter: Features
What full-width characters are
Japanese typography is built on a square grid where each character occupies one cell. Latin letters and digits, being narrower, got duplicate encodings that occupy a full cell so they line up with surrounding Japanese text. Unicode preserved both, so the digit one exists twice: U+0031 as an ordinary digit and U+FF11 as its full-width twin. They look similar, sort differently, and are not equal to each other in any string comparison.
Why this breaks data processing
A full-width digit is not a digit as far as most code is concerned. A numeric parse fails or returns zero, a regular expression matching digits finds nothing, a database numeric cast errors, and a lookup against a reference table silently misses. The values look right in a spreadsheet, which is what makes the bug so irritating to track down: the data appears correct and every operation on it quietly does the wrong thing.
Half-width katakana, the other direction of the problem
Half-width katakana came from early systems constrained to single-byte characters, and it persists in banking records, point-of-sale receipts and legacy exports. It causes the mirror-image problem: voiced sounds are written as two characters rather than one, so a character count is wrong, and text matching against normal katakana fails. Converting to full-width recombines those pairs into single characters, which is what any modern system expects.
Choose what to convert
Katakana, letters and digits, and symbols can each be included or excluded. That control matters because a blanket conversion is often wrong: in a product code you may want the digits normalised while leaving katakana alone, and in a Japanese address you may want the reverse. Converting symbols is the one to think hardest about, since it changes brackets, punctuation and currency marks that may be intentional.
Cleaning customer and address data offline
The conversion is a character mapping performed in your browser, so customer records, address data and exports from an internal system can all be cleaned here without being transmitted, stored or logged. For normalising other aspects of text, such as case or sorting, the text tools in the Japanese edition cover more ground.
// Full-width ⇄ Half-width Converter: FAQ
What is the difference between full-width and half-width?
- They are different Unicode characters that look similar. Full-width forms occupy one full cell of the Japanese character grid and have their own code points; half-width forms are the ordinary ASCII characters you type on a Western keyboard. They are never equal in a string comparison, which is the source of most of the trouble they cause.
Why does my number fail to parse?
- Almost certainly because it is written with full-width digits. They are not recognised by numeric parsing, by a digit character class in a regular expression, or by a database numeric cast. Converting the digits to half-width before processing resolves it, and is a sensible step to apply to any user-supplied numeric input.
Should I normalise on input or on output?
- On input, as close to the boundary of your system as possible. Normalising once when data arrives means everything downstream can assume a single form. Normalising on output leaves inconsistent values in storage, which means every query, comparison and export has to handle both.
Which categories should I convert?
- Depends on the data. For numeric fields and identifiers, convert letters and digits and leave katakana alone. For text destined for a modern system, converting half-width katakana to full-width is usually right. Be cautious with symbols, since brackets and punctuation may be intentional in Japanese text.
What happens to voiced katakana such as ga and pa?
- In half-width they are written as two characters, the base sound followed by a separate voicing mark. Converting to full-width recombines them into the single character a modern system expects, which also corrects the character count. Going the other way splits them again.
Does this affect Japanese text itself?
- Hiragana and kanji have no half-width forms and are left alone. Only katakana, Latin letters, digits and symbols have both widths, so those are the only categories that change.
Is this the same as Unicode NFKC normalisation?
- Similar in effect but broader in scope. NFKC folds full-width forms to their standard equivalents along with many other compatibility mappings, some of which you may not want, such as ligatures and superscripts being rewritten. This tool changes only the width categories you select, which is more predictable when you are cleaning a specific field.
Why do I see this in data from Japanese systems?
- Because Japanese input methods produce full-width characters by default in many contexts, and users typing into a form frequently leave them that way. Legacy systems compound it by storing half-width katakana. Both forms therefore turn up routinely in exports, and neither is an error on the part of the person who entered them.
Can I convert only part of a text?
- Select the categories rather than the range. If you need to convert one field of a record but not another, convert them separately, which is also the safer habit when the fields have different rules.
Is the text I clean sent to a server?
- No. The conversion is a character mapping performed in your browser, and nothing you paste is transmitted, stored or logged.
// How to Use Full-width ⇄ Half-width Converter
-
Full-width or half-width
Select full-width to half-width for cleaning numbers and letters that arrived in their wide forms, or half-width to full-width for normalising legacy half-width katakana.
-
Select the categories
Enable katakana, letters and digits, and symbols individually. Converting only what you need avoids changing punctuation or katakana that was intentional.
-
Paste and copy
Paste the text and the result appears as you type. Copy it with one click, and apply the same conversion at the point where data enters your system rather than repeatedly afterwards.
Category Text