Skip to main content

Text 10

Character counting, case conversion, sorting and replacing

Counting, converting, sorting and cleaning text. Unglamorous work that turns up constantly, and where a small mistake propagates quietly through everything downstream.

// Tools in this category

Character Counter

Count characters with and without spaces, plus words, lines, paragraphs and UTF-8 bytes, updating as you type. Nothing you paste is uploaded.

Text Statistics

Analyse text in your browser: character counts by script, the most frequent characters and words, whitespace totals and an estimated reading time.

Case Converter

Convert text between camelCase, PascalCase, snake_case, kebab-case, CONSTANT_CASE, dot.case, Title Case, uppercase and lowercase, all at once.

Text Sorter

Sort lines alphabetically, remove duplicates, trim whitespace and drop blank lines in one pass. Runs in your browser with nothing uploaded.

Bulk Text Replacer

Apply up to ten find-and-replace rules in one pass, with optional regular expressions and case-insensitive matching, plus presets for common cleanups.

Dummy Text Generator

Generate Lorem Ipsum or Japanese placeholder text by paragraph, sentence or word. Useful for layout testing where real copy is not ready yet.

Full-width ⇄ Half-width Converter

Convert full-width letters, digits, symbols and katakana to half-width and back. Essential for cleaning data exported from Japanese systems.

Unicode Converter

Inspect any text character by character to see its code point, UTF-8 bytes and escape forms, or convert text to and from JavaScript, HTML and CSS escapes.

ASCII Art Generator

Turn text into ASCII art banners using ten FIGlet fonts. Ideal for README headings, terminal startup banners and CLI splash screens.

Emoji Search

Search emoji by keyword or browse by category, and copy one with a click. Recently used emoji are remembered in your browser for next time.

// Characters are not as simple as they look

A character count sounds unambiguous until you ask what a character is. An emoji with a skin tone modifier is one thing on screen and several code points underneath. An accented letter can be one code point or two depending on how it was normalised. A full-width digit looks like a digit and is not one as far as any numeric parser is concerned. Tools that take these distinctions seriously give different answers from tools that do not, and the difference matters whenever a limit is being enforced.

// Normalise at the boundary, once

The recurring lesson in text handling is to convert as data enters your system rather than repeatedly afterwards. If half the records hold full-width digits and half hold ordinary ones, every query, comparison, sort and export has to cope with both forever. Converting once on input means everything downstream can assume a single form, which removes an entire category of intermittent bug.

// Case conventions carry meaning

camelCase, snake_case, kebab-case and CONSTANT_CASE are not interchangeable decoration: each signals something about where an identifier lives. Crossing a boundary between a JSON API, a database schema and a codebase usually means converting between them, and doing it by hand across dozens of field names is where typos come from.

// What to reach for

The Character Counter gives six counts at once, including the UTF-8 byte total that database and protocol limits actually use. The Case Converter shows nine naming conventions together. The Text Sorter sorts, deduplicates and trims in one pass. The Bulk Text Replacer applies up to ten rules with optional regular expressions. Text Statistics breaks a document down by script and frequency. The Unicode Converter identifies the invisible character causing your bug. There are also a full-width converter, a Lorem Ipsum generator, an ASCII art generator and an emoji search.

// Text: frequently asked questions

Why does my character count differ between tools?
Because they count different things. One may count UTF-16 code units, another Unicode code points, another user-perceived characters. An emoji can be one, two or more depending on which definition is used. When a limit matters, find out which unit the system enforcing it counts in.
My numbers will not parse, but they look fine. Why?
Almost certainly full-width digits, which are separate Unicode characters that merely resemble ordinary ones. They fail numeric parsing, digit character classes and database casts while displaying correctly in a spreadsheet. Convert them to half-width before processing.

// Other categories