Kanji Frequency & Probability Explorer

how many kanji cover 95% of the characters in real Japanese text?

Loading interactive simulation...

Zipfian Distributions 🖖

Symbol frequency in natural language follows a power-law distribution...

Coverage is not comprehension 🖖

Coverage is the share of the kanji occurrences in a text that your known set accounts for. Learn the roughly 1,162 most frequent kanji and you will recognize about 95% of the kanji characters in this corpus — yet a page full of familiar characters can still be unreadable if the grammar and vocabulary are new. Coverage measures the symbols you meet, not how much you understand.

The same kanji, two probabilities 🖖

"How likely is a kanji to have 15 or more strokes?" has two correct answers. Draw a kanji uniformly from the dictionary (types) and complex characters are a sizeable slice; draw an occurrence from running text (tokens) and they are far rarer, because the characters you actually read skew simple (人, 日, 一). This gap between counting types and counting tokens is frequency-weighted sampling — the same mechanism behind the friendship paradox.

Example problems