SEPTEMBER 2026 · THE TOOL DESKPractical technology. No theatre.

Practical guide · verified against the real thing

Data snooping: the more ideas you test, the more fake winners you find

In one line: Test a hundred random rules and a few will look great by pure luck. Here is why that happens and the discipline that keeps you from shipping a coincidence.

If you flip enough coins, one of them will land heads ten times in a row, and it would be a mistake to conclude that coin is special. Data snooping is exactly this, in research: test enough ideas against the same historical data and a handful will look excellent purely by chance. The danger is that the lucky ones look identical to the real ones.

Why the winners are partly luck

Every test you run is a chance to find a spurious pattern. Run one hypothesis and a strong result is meaningful; run two hundred and the best of them is mostly noise that happened to fit. The number of ideas you tried — including the ones you discarded — is part of the evidence, and it is the part people forget to count. This is the same intuition behind degrees of freedom: every knob you tuned is a way the result could be fitting the past rather than the future.

The discipline that helps

Decide your hypothesis before you look, not after. Keep an honest tally of how many variants you tried, and discount the result accordingly. Hold out data you have never tested against and only look at it once, at the end — the walk-forward pattern is one structured way to do this. And prefer a result that survives a simple, stable rule over one that needs five tuned parameters to shine; the simpler result has fewer places for luck to hide.

The honest question

Before believing any backtest, ask: "how many things did I try before I found this?" If the answer is large and untracked, the result is suspect no matter how good the chart looks. The validation checklist is where that question gets written down instead of waved away.

Next

Related on this desk.