Text Statistics
Analyse text in your browser: character counts by script, the most frequent characters and words, whitespace totals and an estimated reading time.
Break a piece of text down: how many characters of each script it contains, which characters and words appear most often, how much of it is whitespace, and how long it takes to read. Useful for checking the shape of a document rather than just its length.
// Text Statistics: Features
Breakdown by character type
Characters are counted separately as hiragana, katakana, kanji, Latin letters, digits and symbols. For mixed-script text this reveals proportions that a single count hides: a document that is nominally Japanese but forty per cent Latin letters is probably heavy with technical terms, and one with an unexpectedly high symbol count usually has markup or formatting artefacts left in it.
Frequency analysis
The twenty most common characters and the fifteen most common words are listed with their counts. Word frequency is the more useful of the two for editing: a word appearing far more often than the rest is either the subject of the document or a crutch you have leaned on, and seeing the list is usually enough to tell which. Character frequency is more diagnostic, surfacing stray characters and unexpected punctuation.
Reading time, and how it is estimated
The estimate combines two rates, because mixed text reads at neither pure rate: roughly 500 Japanese characters per minute and 230 English words per minute. Those figures are conventional averages for silent reading of ordinary prose. Dense technical material reads considerably slower and light narrative faster, so treat the number as a planning aid for an article or a talk rather than a measurement.
Whitespace is counted too
The whitespace total is more useful than it sounds. A large gap between the total character count and the count excluding whitespace points at heavy indentation, double spacing or text that has been pasted with formatting. When a document is unexpectedly close to a byte or character limit, whitespace is frequently where the excess is hiding.
Analysing an unpublished draft offline
Everything is computed as you type, and nothing is transmitted, stored or logged. Unpublished drafts, internal reports and customer correspondence can all be analysed here. For a simple length check without the breakdown, the Character Counter gives the six standard counts at a glance.
// Text Statistics: FAQ
What does the character type breakdown show?
- How many characters fall into each of six categories: hiragana, katakana, kanji, Latin letters, digits and symbols. For English text the first three will be zero and the breakdown mainly separates letters from digits and punctuation; for mixed Japanese and English it shows the proportions directly.
How is reading time calculated?
- By combining two rates, about 500 Japanese characters and 230 English words per minute, so mixed text is estimated sensibly. These are conventional averages for silent reading of ordinary prose, not measurements of your particular readers.
Is the reading time accurate?
- It is an estimate and should be treated as one. Dense technical writing, text with code samples, and material a reader needs to stop and think about all take substantially longer than the average. It is useful for sizing an article or a talk, not for a promise.
What counts as a word for the frequency list?
- Whitespace-separated runs containing letters. That works well for languages that separate words with spaces. Japanese and Chinese do not, so word frequency is not meaningful for them; the character frequency list is the useful one there.
Why would I look at character frequency?
- Mostly for diagnosis. It surfaces characters you did not expect to be present: a smart quote where a straight one belongs, a non-breaking space pasted from a web page, or a stray control character. In editing terms it is less useful than word frequency, but it catches things nothing else does.
What does a high whitespace count mean?
- Usually indentation, double spacing, or text pasted with its original formatting. It matters when you are close to a limit, since whitespace counts towards character and byte limits just as much as visible text does.
Can I use the word frequency for writing quality?
- As a prompt rather than a verdict. A word appearing far more often than its neighbours is worth looking at: it is either the genuine subject of the piece or a habit worth varying. The list tells you which words to examine; it cannot tell you whether the repetition is a problem.
How are numbers and punctuation categorised?
- Digits have their own category, and everything that is not a letter, a digit or a Japanese script character falls into symbols, which includes punctuation, mathematical operators and emoji. A surprisingly large symbol count usually means leftover markup.
Does it handle very long documents?
- Yes, though analysis re-runs on every keystroke, so a book-length document will feel sluggish. For anything of that size, analysing a representative chapter gives you the same insight far more comfortably.
Is the text I analyse sent to a server?
- No. All analysis happens in your browser, and the text is never transmitted, stored or logged.
// How to Use Text Statistics
-
Paste the text to analyse
Put the document, article or excerpt into the input box, or load the sample to see what the analysis looks like before using your own material.
-
Read the summary
The character count excluding whitespace, the whitespace total and the estimated reading time sit at the top. A large whitespace figure is worth investigating if you are near a limit.
-
Look at the breakdowns
The character type chart shows the balance of scripts, and the frequency lists show which characters and words dominate. The word list is the one to read when editing.
Category Text