Earth-like Planet Probability

Population-level Earth-like occurrence estimate: true count vs likely detected count.

Loading interactive simulation...

Lesson

The theory β€” Earth-like Planet Probability

This is an expected count, not a probability: start from a number of stars and multiply by a chain of fractions, each answering β€œof those, what share also…?”. The result is how many objects you would expect if every fraction were correct β€” and it is a Drake-equation-shaped calculation, with the same strength and the same weakness.

N* Γ— f_p Γ— f_HZ Γ— fβŠ• Γ— f_rocky Γ— p_det 2 500 100 156 256 78 128 1 2 3 4

Six stacked bars, each half the width of the one above it, from the surveyed stars down to the expected detections.

  1. The stars surveyed β€” the whole bar, and the only figure here that is counted rather than assumed.
  2. Each fraction takes a share of the bar above it, never of the top bar. That is what makes multiplying legitimate, and it is the step people skip.
  3. After four fractions the bar is a sixteenth of where it started: 156256 expected to exist.
  4. Detection is applied after the truth, as a separate stage β€” 78128. What exists and what a survey would see are different bars.

What each symbol means

N*
the number of stars surveyed, 150,000 at the defaults.
fractions
the chain β€” f_p with planets, f_HZ in the habitable zone, fβŠ• Earth-sized, f_rocky rocky. Each is a conditional share of the previous line, not of the whole.
E
the expected population β€” how many are out there: 3,029.
p_det
detection efficiency: the share you would actually find. It converts a truth into an observation.
E_det
the expected detections, E Γ— p_det = 242.

Where the formula comes from

  1. Multiplying is the right operation only because each fraction is conditional on the one before it. f_HZ is the share of planet-hosting stars with a habitable-zone planet, not the share of all stars.
  2. Chaining them: E = N* Β· f_p Β· f_HZ Β· fβŠ• Β· f_rocky. At the defaults that is 150,000 Γ— 0.85 Γ— 0.18 Γ— 0.22 Γ— 0.6 = 3,029.
  3. Detection is a separate stage, applied after the truth: E_det = E Β· p_det = 3,029 Γ— 0.08 = 242. Keeping it separate is the point β€” it distinguishes what exists from what a survey would see.
Assumes
That every fraction is independent of the others given the previous step, and that a single number can stand for each. Neither is safe β€” planet occurrence correlates with stellar type and metallicity, so these fractions are not really free to be set one at a time.
Breaks when
Read the precision sceptically. The output is quoted to the unit β€” 242 β€” from five fractions each set by a slider to one decimal place, and multiplying five guesses multiplies their errors rather than averaging them. Halve two of the fractions and the answer quarters. The band shown is narrow because it explores the survey size, not the fractions, and the fractions are where nearly all the real uncertainty lives.

Transit surveys miss about 214 planets out of every 215 🖖

A transiting planet only shows up if its orbit happens to lie edge-on from our direction, and the odds of that are roughly the star’s radius divided by the orbital distance. For an Earth around a Sun that is 696 340 km over 1 AU β€” about 0.465%, or one alignment in 215. Every other Earth analogue out there is real and completely invisible to this method, not because the telescope is weak but because the geometry never lines up. That is why an occurrence rate is never the raw count of detections: the count has to be divided by the alignment probability, and by the survey’s own detection efficiency, before it means anything. The number this tool reports is a corrected estimate, and the correction is much larger than the measurement.

A funnel from billions of stars 🖖

Start with every star your telescope watches, then keep only a fraction at each step: stars that host planets, planets in the habitable zone, and among those the Earth-sized, rocky ones. Multiplying these fractions is like sifting sand through nested sieves β€” the final number is far smaller than where you began. The detected count is smaller still: a low detection efficiency means many real Earth-like worlds never show up in the data at all.

Why the error bar borrows from radioactivity 🖖

The rough low-to-high band here comes from Poisson statistics, the same math that describes clicks on a Geiger counter or photons hitting a detector. When you count rare, independent events, the typical spread is about √E, so the tool's band is roughly E ± 1.96√E for 95% confidence. Counting habitable planets and counting radioactive decays obey the identical uncertainty law.

Practice

Check yourself

Predict the answer first, then use the controls above to find out. Reveal only after you have committed to a guess β€” that is what makes it practice.

  1. Set every fraction to 0.5 and note the two counts. Now raise just one of fp, fHZ, fEarth or frocky to 0.6 β€” then put it back and raise pdetect to 0.6 instead. What is different about the fifth one?

    Show answer
    The four population fractions are interchangeable. Whichever one you raise, detected goes 78,128 β†’ 93,754 and the population truth goes 156,256 β†’ 187,508, because multiplication does not care about order: a 20% improvement is worth 20% wherever you apply it. Raising pdetect gives the same 93,754 detected but leaves the truth at 156,256 β€” detection efficiency is a fact about your telescope, not about the galaxy. That is exactly what the two rows are for: four factors say how many planets are out there, and the fifth says how many you would see. It is also why arguments of this kind always turn on the *most uncertain* factor rather than the smallest one. The leverage is identical, so only the error bars distinguish them.
  2. The band reads 155,481 – 157,031 around a central 156,256 β€” give or take 775. Where does 775 come from, and what does it leave out?

    Show answer
    √156,256 = 395.3, and 1.96 Γ— 395.3 = 775: a 95% interval for Poisson counting noise. It is the scatter you would see if the five fractions were known exactly and nothing but chance decided how many planets fell into your sample. It leaves out any uncertainty in the fractions themselves β€” and those dominate completely. fHZ and fEarth are argued over by factors of two in the literature, not by half a percent. Move fHZ from 0.5 to 0.6 and the estimate jumps by 31,252, about forty times the width of the band the tool had just printed. The band is honest about the one thing it measures and wildly optimistic about the number as a whole.

Problem solved in full

  1. Earth-like planets to detect from a survey of 150,000 stars 6 steps

    Survey 150,000 stars. With the fractions this tool starts from, work out how many Earth-like planets are out there and how many you would actually detect β€” then work out how much of the printed uncertainty band you should believe.

    1. The model is a single product: start with the star count and multiply by each fraction in turn. Nothing in it is conditional on anything else, which is itself an assumption.

    2. Work through it left to right. Each factor cuts the survivors, and the last three between them discard 96% of the stars that had planets at all.

    3. Detection is a sixth factor, applied to the population rather than to the stars. Roughly one Earth-like planet in twelve is found.

    4. Now the band. Treating the count as Poisson makes its standard deviation the square root of the mean, and two of those either side is the interval the tool prints.

    5. Compare that detection efficiency with the geometry alone. A transit requires the orbit to be edge-on to within Rβ˜‰/1 au, which is one chance in 215 β€” so 0.08 is not a geometric probability but a whole survey strategy rolled into one digit, and the pure-transit answer would be 14 planets rather than 242.

    6. Finally, vary a single fraction over a range no astronomer would call unreasonable, and watch the answer move by a factor of four.

    Answer

    3,029 present, 242 detected, and the Β± 108 band is worth almost nothing. That band is Poisson counting noise: it is the spread you would see if the five fractions were exactly right and only the dice varied. They are not exactly right. Let just one of them β€” the habitable-zone fraction β€” be uncertain by a factor of two either way, and the estimate runs from 1,515 to 6,059, a range twenty times wider than the one on the page. This is the standard failure of any multiplied-fractions model, the Drake equation included: the arithmetic delivers four-digit precision from inputs that are known to one, and the precision is entirely an artefact of the multiplication.

References (1)

Example problems

  • Kepler-like - 150,000 stars work out to about 3,029 Earth-like worlds, of which the survey sees 242 at 8% efficiency. The other 2,787 are real and unseen.
  • TESS-like - It watches 200,000 stars, more than Kepler, and detects 63 against Kepler's 242 - because its efficiency is 3% rather than 8%. A bigger survey with a smaller yield.
  • Future survey - A million stars at 25% efficiency: 9,702 detections. Forty times Kepler's count, and the 95% band is four percent wide instead of twenty-five, because the spread goes as the square root.