Problem solved in full
-
Twelve candidate continuations that cover 13.6% of what actually follows 5 steps
Twelve candidate continuations cover 13.6% of what actually follows this context. Work out what the other 86% looks like, and why that shape decides how sampling has to work.
-
The corpus offers 1280 distinct continuations of this context. That is the real distribution the model is trying to imitate, and it is far wider than any shortlist.
-
The top twelve — the ones the panel scores — account for 13.6% of all occurrences between them.
-
So the remaining 1268 continuations share 86.4%. Not one of them is likely on its own, and collectively they are most of what happens.
-
That is 99.1% of the options holding 86.4% of the probability: a long tail, not a peak with noise around it.
-
With top-k off all twelve are kept and the panel reads 100% probability kept. That 100% is measured inside the shortlist while the 13.6% is measured against the corpus, so every filter on this page cuts the head: the 1,268 tail continuations were gone before any slider moved.
Answer
The tool prints 12 / 12 kept, 100.0% of probability retained, 1,280 distinct continuations and 13.6% coverage. The tail is why sampling is hard. Always take the most likely token and the text collapses into repetition, because you are choosing from the 13.6% forever. Sample from everything and you eventually draw something absurd, because 86.4% of the mass sits on options that are individually terrible. Top-k and top-p exist to cut the tail without flattening the head, and the numbers above are the trade-off they are negotiating. Watch which way temperature moves it: set top-p to 90 and ten of the twelve survive, then raise the temperature to 1.50 and an eleventh comes back, because flattening the distribution makes the nucleus wider rather than narrower.
-
Learning path
How an LLM picks the next word
References (2)
- Where the temperature in a softmax comes from — it is the one in the Boltzmann distribution: D. H. Ackley, G. E. Hinton and T. J. Sejnowski, "A Learning Algorithm for Boltzmann Machines." Cognitive Science 9(1), 147–169, 1985.
- Why lowering the temperature is not the same as improving the text, and what nucleus sampling does instead: A. Holtzman, J. Buys, L. Du, M. Forbes and Y. Choi, "The Curious Case of Neural Text Degeneration." ICLR 2020.