Unicode Converter
Inspect any text character by character to see its code point, UTF-8 bytes and escape forms, or convert text to and from JavaScript, HTML and CSS escapes.
Three modes in one tool: inspect text character by character to see its code point and UTF-8 bytes, escape text into JavaScript, HTML or CSS form, or unescape text that arrived in any of those forms.
💡 Type some text to see each character’s Unicode code point, its escape forms and its UTF-8 bytes.
// Unicode Converter: Features
Inspecting, character by character
Paste text and every character is listed with its code point in U+ notation, its escape forms and its UTF-8 byte sequence. This is the mode that answers the question that brings most people here: what exactly is this character? A quotation mark that will not match, a space that behaves oddly, a letter that sorts wrongly. Seeing the code point usually identifies the culprit immediately.
The five escape formats
JavaScript escapes use backslash-u notation, which is what appears in source code and JSON strings. HTML has both hexadecimal and decimal numeric references, and which one you meet depends on the library that produced the document. CSS escapes use a backslash followed by hexadecimal digits, with their own rules about the trailing space. U+ notation is the form used in specifications and documentation rather than in code. Being able to produce any of them means you can match whatever context the value is going into.
Invisible characters are the usual suspects
A great many baffling bugs come down to a character that looks like something else or looks like nothing. A non-breaking space pasted from a web page is not a space as far as trimming and splitting are concerned. A zero-width joiner sits invisibly inside a string and breaks an equality check. A curly apostrophe from a word processor is a different character from the straight one your validation expects. The inspect mode makes all of them visible, because each gets its own row with a code point.
Code points, UTF-16 and surrogate pairs
A code point above the basic plane, which includes most emoji and a number of rare CJK characters, is stored in JavaScript as two UTF-16 units called a surrogate pair. That is why such a character reports a length of two, why slicing a string can split it in half and produce a replacement character, and why counting characters naively goes wrong with emoji. Seeing both the code point and the byte sequence makes the distinction concrete.
All three modes work offline
All three modes run locally and nothing you paste is transmitted, stored or logged. To escape the five HTML markup characters specifically rather than arbitrary code points, the HTML Escape tool handles that case directly.
// Unicode Converter: FAQ
What does the inspect mode show?
- One row per character, with its code point in U+ notation, the escape forms for that character and its UTF-8 byte sequence. It is the fastest way to identify a character that looks ordinary and is not.
Which escape formats can it produce?
- JavaScript backslash-u escapes, HTML numeric references in both hexadecimal and decimal, CSS backslash escapes, and U+ notation. Choose the one that matches where the value is going, since each context accepts a different form.
How do I find an invisible character in my text?
- Paste the text into inspect mode and look for rows whose character appears blank or unfamiliar. Non-breaking spaces, zero-width joiners and various control characters all show up with their own code points, which is usually enough to identify what has been pasted in.
Why does my emoji count as two characters?
- Because it lives above the basic multilingual plane and is stored as a surrogate pair in UTF-16, which is what JavaScript strings use. The code point is one value; the string length is two. This is also why slicing a string can cut an emoji in half.
What is the difference between a code point and a UTF-8 byte?
- A code point is the abstract number identifying a character. UTF-8 is one way of writing that number as bytes, using one byte for ASCII, two for most European scripts, three for CJK and four for characters above the basic plane. The tool shows both, which is why the byte count and character count differ.
Why does my quotation mark not match?
- Because it is probably a typographic quote rather than a straight one. Word processors and some content systems substitute curly quotes automatically, and they are entirely different characters. Inspecting the text shows the code point and confirms it in seconds.
When should I use CSS escapes?
- When a selector or a content value contains a character that CSS would otherwise interpret, or when you need to be certain a character survives a stylesheet that may be served in an uncertain encoding. The trailing space after a CSS escape is significant, which is a common source of confusion.
Can it unescape text that mixes formats?
- It recognises the supported formats, so a string containing both JavaScript and HTML escapes is handled. Anything it does not recognise is left unchanged rather than mangled, which makes it easy to see what remains.
Is it safe to paste data containing personal information?
- Yes, in the sense that nothing is transmitted, stored or logged. The inspection happens entirely in your browser. As with any tool, treat content you have shown on a shared screen as exposed.
Is the text I inspect sent to a server?
- No. All three modes run in the page, and closing the tab discards whatever you pasted.
// How to Use Unicode Converter
-
Pick a mode
Choose inspect to examine text character by character, escape to convert text into an escape format, or unescape to turn escaped text back into readable characters.
-
Paste the text to inspect
In inspect mode every character is listed immediately with its code point and bytes. In escape mode, select the target format: JavaScript, HTML hexadecimal or decimal, CSS, or U+ notation.
-
Read or copy the result
Use the character table to identify an unexpected character, or copy the escaped or unescaped output for use in your code, stylesheet or document.
Category Text