Lesson
The theory — Kanji Frequency & Probability Explorer
This is a coverage question, not a counting one: given that a few characters are used constantly and most are used rarely, how many of the commonest do you need to know to recognise a given share of everything written? The answer is far fewer than the total, and the reason is that character frequency is extremely uneven.
What each symbol means
rank- a character’s position when all are sorted by how often they appear. Rank 1 is the commonest.
coverage- the share of all occurrences accounted for by everything up to some rank — a running total of frequency, not a count of characters.
required- how many of the commonest characters you need for a target coverage. The table above turns a coverage you choose into that number.
known- how many you are assuming you already know, which the page uses to report what your current coverage would be.
How to read what you see
The table reads as a price list for fluency: 50% coverage costs 158 kanji, 80% costs 558, 90% costs 872, 95% costs 1,162, and 99% costs 1,651. Read the gaps rather than the rows and the shape of the language appears.
- Assumes
- One fixed corpus. Frequency is a property of the texts counted, not of the language in the abstract — a newspaper corpus, a manga corpus and a legal corpus would each reorder the tail and shift every number in that table.
- Breaks when
- The returns collapse, and they collapse in a way worth planning around. The first
158characters buy you half of everything. Going from 90% to 95% costs another290, and the final 1% — from 99% coverage to complete — costs more than the entire journey from 50% to 90%. There is no rank at which the tail ends, only a point where each additional character stops being worth the hour it takes.
Problem solved in full
-
500 kanji giving 77.39% coverage of this corpus 6 steps
500 kanji gives 77.39% coverage of this corpus. Work out what that means for a page of text, and then what it costs to close the gap.
-
Coverage and its complement are the same fact. The reciprocal of the complement is the mean gap between unknowns, which is the form that tells you what reading actually feels like.
-
Scale that to a whole work. The corpus header gives the occurrence count and the number of works, so the arithmetic needs nothing from outside the page.
-
Now the cost curve. The panel lists the kanji count needed for five coverage targets, and the interesting quantity is the difference between consecutive rows.
-
Divide the coverage gained by the kanji spent. The rate halves between the two intervals.
-
The remainder of the Jōyō list is the extreme case: a fifth of the characters for a hundredth of the text.
-
Turn the goal around. Reading with one unknown per work rather than one per 4.4 kanji is a coverage figure with four nines in it, which no fixed list of 2,136 characters can deliver.
Answer
One unknown kanji every 4.4, and the last 1% costs eight times as much per kanji as the stretch from 90% to 95%. The first 158 kanji buy half the corpus. The 290 that take you from 90% to 95% buy 0.017 points each. The 489 that take you from 95% to 99% buy 0.008. And the 485 left over buy 0.002 — a fifth of the alphabet for a hundredth of the text. That shape is what a frequency distribution with a long tail always does, and it decides how a language is worth learning: the first thousand characters are the best investment available anywhere in education, and the last five hundred are a completionist's hobby. It also explains why 77% coverage does not feel like 77%. Missing one kanji in 4.4 means roughly 2,700 stumbles in an average work in this corpus, which is not reading.
-
References (2)
- The corpus every number here is counted from — a sample of public-domain Japanese literary works: Aozora Bunko (青空文庫), public-domain Japanese texts. Counts restricted to the Jōyō kanji set.
- Per-character metadata (readings, grade, stroke count): KANJIDIC2, Electronic Dictionary Research and Development Group (EDRDG), used under CC BY-SA 4.0.