Kanji Frequency & Probability Explorer

how many kanji cover 95% of the characters in real Japanese text?

Loading interactive simulation...

Lesson

The theory — Kanji Frequency & Probability Explorer

This is a coverage question, not a counting one: given that a few characters are used constantly and most are used rarely, how many of the commonest do you need to know to recognise a given share of everything written? The answer is far fewer than the total, and the reason is that character frequency is extremely uneven.

What each symbol means

rank
a character’s position when all are sorted by how often they appear. Rank 1 is the commonest.
coverage
the share of all occurrences accounted for by everything up to some rank — a running total of frequency, not a count of characters.
required
how many of the commonest characters you need for a target coverage. The table above turns a coverage you choose into that number.
known
how many you are assuming you already know, which the page uses to report what your current coverage would be.

How to read what you see

The table reads as a price list for fluency: 50% coverage costs 158 kanji, 80% costs 558, 90% costs 872, 95% costs 1,162, and 99% costs 1,651. Read the gaps rather than the rows and the shape of the language appears.

Assumes
One fixed corpus. Frequency is a property of the texts counted, not of the language in the abstract — a newspaper corpus, a manga corpus and a legal corpus would each reorder the tail and shift every number in that table.
Breaks when
The returns collapse, and they collapse in a way worth planning around. The first 158 characters buy you half of everything. Going from 90% to 95% costs another 290, and the final 1% — from 99% coverage to complete — costs more than the entire journey from 50% to 90%. There is no rank at which the tail ends, only a point where each additional character stops being worth the hour it takes.

The list decides how often you can possibly revise 🖖

Frequency rank is usually read as importance, but it fixes something more practical: how long you wait between sightings. In this tool's corpus 人 at rank 1 turns up once every 62 kanji, 休 at rank 500 once every 2,051, and 騎 at rank 1,500 once every 15,428. Read a 10,000-character story and your chance of meeting that rank-1,500 character even once is 47.7% — so learning it and then never running into it again is the ordinary outcome rather than bad luck, and it is an argument for following frequency order rather than a textbook's. Nineteen of the 2,136 jōyō kanji do not appear at all in the 709,707 characters sampled. Worth knowing what the sample is, though: 60 works of mostly early-twentieth-century literary fiction, not contemporary news or technical writing.

Coverage is not comprehension 🖖

Coverage is the share of the kanji occurrences in a text that your known set accounts for. Learn the roughly 1,162 most frequent kanji and you will recognize about 95% of the kanji characters in this corpus — yet a page full of familiar characters can still be unreadable if the grammar and vocabulary are new. Coverage measures the symbols you meet, not how much you understand.

The same kanji, two probabilities 🖖

"How likely is a kanji to have 15 or more strokes?" has two correct answers. Draw a kanji uniformly from the dictionary (types) and complex characters are a sizeable slice; draw an occurrence from running text (tokens) and they are far rarer, because the characters you actually read skew simple (人, 日, 一). This gap between counting types and counting tokens is frequency-weighted sampling — the same mechanism behind the friendship paradox.

Problem solved in full

  1. 500 kanji giving 77.39% coverage of this corpus 6 steps

    500 kanji gives 77.39% coverage of this corpus. Work out what that means for a page of text, and then what it costs to close the gap.

    1. Coverage and its complement are the same fact. The reciprocal of the complement is the mean gap between unknowns, which is the form that tells you what reading actually feels like.

    2. Scale that to a whole work. The corpus header gives the occurrence count and the number of works, so the arithmetic needs nothing from outside the page.

    3. Now the cost curve. The panel lists the kanji count needed for five coverage targets, and the interesting quantity is the difference between consecutive rows.

    4. Divide the coverage gained by the kanji spent. The rate halves between the two intervals.

    5. The remainder of the Jōyō list is the extreme case: a fifth of the characters for a hundredth of the text.

    6. Turn the goal around. Reading with one unknown per work rather than one per 4.4 kanji is a coverage figure with four nines in it, which no fixed list of 2,136 characters can deliver.

    Answer

    One unknown kanji every 4.4, and the last 1% costs eight times as much per kanji as the stretch from 90% to 95%. The first 158 kanji buy half the corpus. The 290 that take you from 90% to 95% buy 0.017 points each. The 489 that take you from 95% to 99% buy 0.008. And the 485 left over buy 0.002 — a fifth of the alphabet for a hundredth of the text. That shape is what a frequency distribution with a long tail always does, and it decides how a language is worth learning: the first thousand characters are the best investment available anywhere in education, and the last five hundred are a completionist's hobby. It also explains why 77% coverage does not feel like 77%. Missing one kanji in 4.4 means roughly 2,700 stumbles in an average work in this corpus, which is not reading.

References (2)

Example problems