Problems solved in full
-
One bean, one report of “left”, and four different right answers 8 steps
Throw a single bean and it lands left of the hidden marker. The panel says 0.6667. The first insight block on this page says 0.7500. The answer most people give out loud is 1, because one out of one landed left. All three are correct. Which question does each one answer?
-
The setup first, because it is the reason this particular thought experiment is famous. A marker is dropped blind onto a flat table at some unknown position p, and every later bean lands to its left with probability p. Nothing is ever measured; each bean reports one bit. Richard Price read the essay to the Royal Society in December 1763, two years after Bayes had died, and the uniform prior on p is a physical fact about how the marker was thrown rather than a modelling choice. That is rarer than it sounds, and the third insight block on this page is about how rarely you get it anywhere else.
-
The evidence. Getting k on the left out of n throws has probability pk(1−p)n−k, times a binomial coefficient that carries no p at all. Multiply by the flat prior, then normalise, and the coefficient divides out.
-
What is left is a Beta whose two parameters are those exponents plus one. One left out of one gives Beta(2, 1), which is a straight line: zero at the left edge, rising to 2 at the right.
-
Question one: where is the marker most likely to be? The density 2p is largest at the right-hand edge, so the mode is p = 1. The blurted answer k/n = 1 arrives at the same place by a different road, and both are answers to a question nobody asked.
-
Question two: is the marker past halfway? That is an area, so integrate. The cumulative function of Beta(2, 1) is p², so the mass above one half is 1 − ¼. This is the 0.7500 in the first insight block.
-
Question three: will the next bean land left? This one is not about where p is at all — it is about what happens next, and it averages p over everything you now believe about p. That average is the posterior mean.
-
With n = 1 and k = 1 the mean is two-thirds. Laplace derived this in a memoir of 1774 and it has been called the rule of succession ever since. The +1 and +2 are not a fudge for small samples; they are what integrating against a flat prior gives you.
-
A fourth number, which the tool prints nowhere. The median sits at 1/√2, between the mean and the mode, in that order. Every posterior this tool can produce is a Beta, and a Beta that leans right puts its three centres in exactly that order, and the gaps close again as the evidence concentrates the curve. Balance the throws and all three collapse onto one another.
Answer
Each of the three is right, and each answers a different question. 1 is where the marker most probably lies. 0.7500 is the chance it lies past halfway. 0.6667 is the chance the next bean lands left, and that is the one the panel puts in the headline card because it is the only one of the three that says anything about the next throw. The 95% interval runs from 0.1581 to 0.9874: one report has narrowed almost nothing, and the panel says 341 more throws before that interval is ±5 points wide.
Laplace took the same formula as far as it will go. In the Essai philosophique sur les probabilités of 1814 he set n = 1,826,213 — five thousand years of recorded sunrises — and got odds of 1 in 1,826,215 that the sun fails to rise tomorrow. He has been mocked for it ever since, mostly by people who stopped reading. In the same passage he says the number is for someone who knows nothing of what makes the days turn, and that anyone who does understand the mechanism should give a far larger one. That caveat is step 1 stated by the man who invented the calculation. -
-
What the 95% on the interval is a promise about 6 steps
Run the 400-throw preset. The panel puts the marker between 0.4512 and 0.5488 with 95% credibility, and the marker really is at 0.4913. Show that the 95% is exact rather than approximate — then work out what is left of it if the table is not flat.
-
Count the beans again, including the one nobody looks at. The marker was thrown blind onto the table and so were the 400 that followed; all 401 came off the same shoulder onto the same flat surface. There is nothing statistical that distinguishes the marker from a throw. It only has a name.
-
So take the 401 positions and sort them. The marker is equally likely to be first, or seventh, or last, because a label cannot change a rank, and the number the panel calls k is just its rank minus one. Every left-count the tool can print — all 401 of them, 0 through 400 — has probability 1/401. Which is not what a count of coin flips looks like, and the integral confirms it: the binomial coefficient cancels against the beta function and n drops out except as n + 1.
-
The same symmetry settles the interval. Feed the true marker position into its own posterior distribution function and the result is uniform on the unit interval — that is what it means for a posterior to be the honest conditional law of p. So the chance the truth falls between the 2.5th and 97.5th percentiles is 0.95, and there is no approximation anywhere in that sentence. Simulating ten throws a hundred thousand times gives 0.94967.
-
It holds at every sample size, which is the part worth pausing on. The 400-throw interval is 0.0976 wide and the marker sits comfortably inside it. The one-throw interval runs from 0.1581 to 0.9874, which excludes almost nothing — and it is exactly as valid. Being useless and being wrong are different failures.
-
Now break the one thing it rests on. Suppose the table is not flat but crowned, so the marker tends to roll toward a rail, while the panel goes on computing from a uniform prior. Integrate the coverage over that true distribution and the 95% label is worth 68.7% after one throw, 78.7% after ten, 87.1% after a hundred.
-
Note which way it failed, and how slowly evidence repaired it. A hundredfold increase in data closed a little over half the gap. And a table crowned the other way — one that gathers the marker near the middle — leaves the interval too wide instead: 99.9%, 95.8%, 95.1%. It is the edges that hurt, because that is where a flat prior is most confident and most wrong.
Answer
Exactly 95%, at every n, for the reason the third insight block gives: here the uniform prior is a fact about the throw and not a modelling choice. That is a rare thing to be able to say, and it makes this the one place where a Bayesian interval and a frequentist confidence interval are the same object rather than two things that usually agree.
The 1/401 is the tidier half of the same symmetry, and it is the line to keep. Before a single bean is thrown, every possible left-count is equally likely — the flat prior does not merely sit under the calculation, it flattens the data too. That is why the +1 and +2 in the succession rule are exact rather than a small-sample patch.
Which leaves the question the table cannot answer for anything else. Prevalence in a screening programme, a base rate for a rare fault, the fraction of a population with a trait: each plays p's structural role and none of them was thrown blind onto anything. Get that distribution wrong at the edges and the 95% on the printout is 69%, and no amount of further sampling within reach will tell you, because a hundred throws only recovered half of what a crowned table cost. -
Learning path
Priors, and where they come from
References (2)
- Bayes' original paper presenting the billiard thought experiment and uniform prior formulation: T. Bayes, "LII. An essay towards solving a problem in the doctrine of chances." Philosophical Transactions of the Royal Society of London, 370–418, 1763.
- The 1958 Biometrika reprint accessible to modern readers: T. Bayes, "An essay towards solving a problem in the doctrine of chances." Biometrika, 45(3–4):296–315, 1958.