Cognitive Biases · CB-26
Fire the shots first, then paint the target around wherever they landed — and call it marksmanship.
A reasoning error in which a pattern is identified in data after the fact, and a hypothesis is then constructed specifically to fit that pattern, rather than being tested against data collected independently — named for the image of a shooter firing randomly at a barn wall, then painting a target around the tightest cluster of hits and claiming to be a sharpshooter.
A folk-derived term widely used in statistics and epidemiology to describe post-hoc pattern-matching, closely related to and often discussed alongside data dredging and the multiple-comparisons problem in formal statistical methodology.
The Mechanism
Draw the bullseye after the shots, and any cluster looks like skill
The 'accuracy' is manufactured entirely after the fact — the shooter had no real aim, but painting the target around wherever the shots happened to land the most densely creates the illusion of skill from what was actually pure chance.
01 · IT'S THE DATA-ANALYSIS VERSION OF FORMULATING A HYPOTHESIS AFTER SEEING THE RESULTS
Testing a hypothesis on the same data that generated it proves nothing
A hypothesis discovered by scanning a dataset for any interesting-looking pattern, and then 'confirmed' using that same dataset, is close to guaranteed to appear to work — the correct test requires either a genuinely independent, held-out dataset, or a hypothesis specified in advance of seeing the data.
02 · IT'S CLOSELY TIED TO THE MULTIPLE-COMPARISONS PROBLEM
Testing enough hypotheses guarantees some will look significant by chance alone
When many different patterns, subgroups, or correlations are tested against the same dataset, some will appear statistically significant purely by chance — without correcting for the number of comparisons made, 'discovering' an apparently meaningful pattern this way is close to inevitable and uninformative.
03 · IT'S A RECURRING PROBLEM IN MEDICAL, SOCIAL SCIENCE, AND BUSINESS ANALYTICS RESEARCH
Pre-registration and out-of-sample testing are the standard fixes
The replication crisis in several scientific fields has been linked in part to exactly this pattern — researchers exploring datasets for any significant-looking result and reporting it as a confirmed finding — which is why pre-registered hypotheses and held-out validation datasets have become increasingly standard requirements in rigorous research.
Where It Fails / Inversion
Where it fails / inversion
Exploratory data analysis is a legitimate and valuable first step in research — the fallacy specifically lies in presenting a pattern discovered through open-ended exploration as if it had been a confirmed, independently-tested hypothesis, not in exploring data at all. Properly labeled exploratory findings, flagged as needing independent confirmation, are perfectly valid science.
How To Use It
Worked example · testing a marketing insight discovered by mining customer data
If a data team notices, by exploring purchase records after the fact, that customers born in a particular month buy more of a product, that pattern should be treated as a hypothesis to test on new, independent data — not immediately acted upon as a confirmed insight, since scanning enough demographic slices of any dataset will turn up some apparently meaningful-looking pattern purely by chance.
How to use it
Before acting on a pattern discovered by exploring a dataset, ask whether the hypothesis was specified before or after you saw the data. If after, treat it as a lead to test on fresh, independent data — not as a confirmed finding, however striking the pattern looks in the data that generated it.
See Also