Generated by Codex with GPT 5.6 Sol XHigh
Accuracy Is Not the Question
A statistic can sound decisive while answering the wrong question. A test that is 99 percent accurate appears to make a positive result almost conclusive. But that percentage describes how the test behaves when a condition is already known; a person receiving a result wants to know the reverse: given this positive result, how likely is the condition?
The gap between those questions is the base rate fallacy. People focus on vivid evidence about a particular case and neglect how common the underlying category is. The article begins with psychologist Daniel Kahneman’s example of a shy, orderly and detail-oriented person named Steve. His description fits a stereotype of a librarian, so many people classify him that way. Yet U.S. farmers outnumber librarians by more than 11 to one. Even if the personality sketch is somewhat more typical of librarians, the much larger pool of farmers may still make farmer the more likely answer.
The same error becomes consequential in medicine. Suppose a disease affects one person in 1,000. A blood test catches every real case and correctly clears 99 percent of people who do not have the disease. A positive result feels alarming because both performance figures are excellent. The rarity of the illness, however, changes what that result means.
One Signal among Ten False Alarms
Testing 1,000 randomly selected people makes the arithmetic visible. The one person with the disease produces one true positive. Among the other 999 people, a 1 percent false-positive rate produces about 10 false alarms. The expected pool of 11 positive results therefore contains only one real case. For a randomly selected person who tests positive, the probability of actually having the disease is about 9 percent, not 99 percent.
There is no contradiction. The test can be highly accurate while most of its positive results are wrong because it is searching for something extremely rare. When the population without the disease is hundreds of times larger than the population with it, even a small error rate applied to that large group can overwhelm the true signals.
This is also why context before testing matters. A person selected at random begins with the disease’s population prevalence as the relevant base rate. Someone who seeks care with a distinctive rash and high fever belongs to a narrower comparison group in which the disease may be much more common. That higher pretest likelihood makes a positive result more informative. The lesson is not to distrust medical testing, but to interpret a result alongside the reason the test was ordered.
Population-wide screening for a rare disease must therefore weigh more than the chance of catching cases. False positives can trigger anxiety, repeat tests, medical procedures and financial costs. A screening program may still be worthwhile, but its designers must compare those burdens with the benefit of finding otherwise missed disease.
Rare Events at Massive Scale
The same mathematics governs automated surveillance. At the 2017 UEFA Champions League final in Cardiff, Welsh police used facial-recognition software to scan about 170,000 fans. It flagged 2,470 people as potential matches; 2,297 of those flags were false positives. A modest error rate, multiplied across a huge crowd, created a large volume of bad leads.
Systems intended to discover terrorist plots face an even harsher version of the problem. Genuine plots are extraordinarily rare, and suspicious behavioral patterns are imperfect indicators. The article cites security expert Bruce Schneier’s back-of-the-envelope reasoning that a broad data-mining system could generate tens of millions of false alarms for every real threat it uncovered. Investigators would then have to search a vast pile of noise, spending resources and intruding on innocent people while trying to find the signal.
None of this means that every high-false-alarm system should be abandoned. Fire alarms are usually false, yet society accepts the inconvenience because a real fire has enormous consequences and an alarm is relatively cheap to check. The proper decision depends on the event’s prevalence, the detector’s error rates, and the relative costs of missed cases and false alarms.
The article’s central warning is simple: the accuracy of a detector is not the probability that the event it reports has occurred. To judge news about a medical test, an algorithm or a security system, the first question should be how often the target event happens in the population being examined. Without that base rate, an impressive percentage may reveal far less than it seems.