Bayesian Biosignature Update

how base rate and test quality shape the posterior after a positive biosignature signal

Loading interactive simulation...

A better telescope cannot lower the false-positive rate 🖖

It is natural to assume a sharper instrument makes a detection more trustworthy, but the false-positive rate in this calculation is mostly not an instrument property at all. Oxygen can accumulate without life, from water vapour split by ultraviolet light. Phosphine has proposed abiotic routes, which is why the 2020 Venus claim moved straight into a dispute about geochemistry rather than about telescope quality. Methane comes out of rocks as readily as out of organisms. When the confusion lives in the chemistry, resolution and sensitivity do not touch it — they only make you more certain the molecule is there, which was never the uncertain part. The only thing that shifts the posterior is a second line of evidence that fails differently.

Why a hit isn't a confirmation 🖖

How sharp your instrument is matters far less than how rare life is to begin with. If only 1 in 1000 planets host life, even a 90% sensitive detector with a 2% false-positive rate flags mostly lifeless worlds — the posterior probability lands near 4%, because false alarms drown out the rare real find. Before trusting a positive signal, ask how likely life was before you looked.

The same error that jails the innocent 🖖

Reading a positive as "life confirmed" commits the prosecutor's fallacy — swapping P(signal | no life) for P(no life | signal). Courts have made exactly this slip: in the 1999 Sally Clark case, an expert's tiny probability of the evidence arising by chance was wrongly presented as the probability of her innocence, helping convict a mother later exonerated in 2003. Your detector's low false-positive rate is not the chance the world is lifeless.

Problem solved in full

  1. A positive biosignature test when the chance of life is 16% 5 steps

    A biosignature test comes back positive — and the chance of life is 16%. Work out why, and then find which of the test's two numbers you would actually pay to improve.

    1. Bayes' rule weighs the two ways a positive can happen: a real detection, or a false alarm. Both are probabilities of the evidence, and the prior decides how much of each there is to begin with.

    2. Substituting gives 0.0150 of true positives against 0.0784 of false ones. The false alarms outnumber the real detections five to one, and nothing about the test is broken — the prior is simply small.

    3. So a positive leaves 16.06% for life and 83.94% for a false alarm. This is the same arithmetic as a medical screen for a rare condition, and it surprises people in exactly the same way.

    4. The odds form makes the structure visible: prior odds multiplied by the likelihood ratio. The test is worth a factor of 9.375, which is real evidence — it just starts from odds of 1 in 49.

    5. Now vary the two inputs. Raising sensitivity from 0.75 towards 1 can multiply the evidence by at most 1.33. Cutting the false-positive rate from 0.08 to 0.01 multiplies it by eight, and the posterior goes to 60.5%.

    Answer

    The tool prints a posterior of 16.06%, a false-alarm share of 83.94% and a Bayes factor of 9.38. The actionable part is step 5. Sensitivity is bounded above by 1, so improving it has a ceiling; the false-positive rate has no floor, so it is where the leverage is. At a false-positive rate of 0.001 the same positive would mean 93.9%. That is why claims of biosignatures are argued over ruling out abiotic explanations rather than over detection strength — the denominator is the whole fight.

Learning path

Priors, and where they come from

References (2)

Example problems

  • Optimistic - A prior of 0.2 with an 85% sensitive test reaches a posterior of 81.0% - and even here, one positive in five is a false alarm.
  • Balanced - The worked problem's case: prior 0.02, 75% sensitivity, 8% false positives, posterior 16.1%. Eighty-four of every hundred positives are false alarms.
  • Rare life - A prior of 1 in 1000 with a 90% sensitive detector gives 4.3%. The instrument is nearly perfect and 96 of every 100 positives are still wrong.
  • Strict test - Worse than Balanced on the prior (0.01 against 0.02) and on sensitivity (70% against 75%), and it still reaches 58.6% against that preset's 16.1%. The only thing it does better is the false-positive rate: 0.5% against 8%.