Skip to main content

Encoding 5

Base64, URL encoding, HTML escaping and image Data URIs

Text that has to survive a trip through a system that only understands a restricted set of characters. Base64, percent encoding and HTML entities all solve that problem, in different places and with different rules.

// Tools in this category

// Three encodings, three jobs

They are easy to confuse because they all turn readable text into something less readable. Base64 represents arbitrary bytes using 64 printable characters, and exists so binary data can travel through text-only channels such as email or a JSON string. Percent encoding escapes the characters a URL cannot carry literally, and is governed by which part of the URL the value sits in. HTML entities escape the characters a browser would otherwise read as markup. Using the wrong one produces output that looks plausible and fails somewhere downstream.

// Encoding is not encryption

This is worth stating plainly because the mistake is common and expensive. None of these transformations needs a key, and all of them are reversed by a single function call in any language. A password stored as Base64 is stored in plain text as far as an attacker is concerned, and a Basic auth header is readable by anyone who can see the request. Encoding solves a transport problem; confidentiality needs encryption, and stored passwords need a slow one-way hash.

// Where the bugs actually come from

Almost every encoding bug is one of three things. Something was encoded twice, so a percent sign became %25 and a single decode is no longer enough. The encoding used to turn text into bytes did not match on both ends, which is what produces mojibake. Or the right function was applied in the wrong context, most often encodeURI on a value that contains an ampersand, which silently splits it into two parameters. Seeing the intermediate form is usually faster than reasoning about which of the three it was.

// What to reach for

The Base64 Encoder converts text in both directions with full UTF-8 support and shows the size cost. The URL Encoder offers both encodeURI and encodeURIComponent so you can pick the right one for the position. The URL Parser breaks an address into its parts and lists the decoded query parameters. HTML Escape converts markup characters to entities and back. Image to Base64 turns a file into a complete data: URI, and previews one you paste in.

// Encoding: frequently asked questions

Why does my text turn into question marks or strange symbols?
That is mojibake, and it means the bytes were produced with one character encoding and interpreted with another, typically UTF-8 against a legacy encoding such as Shift_JIS or Latin-1. The tools here all use UTF-8, which is what modern systems expect. If a value only decodes correctly under an older encoding, it was produced by an older system and needs converting there.
How do I tell if a string has been encoded twice?
Look for %25 in a URL, or for a visible entity such as < in HTML. Both are the escape character having been escaped again. Decode once, confirm you have something readable, and decode again only if you do not.

// Other categories