First-hand · verified against the real thing
Overfitting: how to catch yourself fitting the past
In one line: The most expensive mistake in quantitative work doesn't feel like a mistake — it feels like success. The tells, the controls, and the promotion vocabulary that keeps a research journal honest.
Overfitting is not an error you commit; it is a gravity you fall toward. Every choice in research — try another variant, tune a threshold, keep the good run — feels like diligence while quietly folding the model onto the past. After ~140 logged experiments, QuantLab's most useful artifact may be the list of tells that a result is curve-fit, because they were all learned by falling. Here is the detection guide.
The tells, from the journal
The result needed the exact configuration. A parameter grid where performance spikes at one setting and dies a step away is not an optimum; it is an address where the past happened to live. The lab's frozen champion config looked like this from the outside — and the blind per-year re-test showed no static variant of the sweep survived all years after costs. Performance is strongest exactly where you looked. Universe-specificity: edges that worked on the symbols they were developed on and failed on unseen ones (logged by run number: R044, R049, R078). The drawdowns are too polite. A backtest with shallow, rare drawdowns has usually not met a hostile regime; the honest full-universe audit of the surviving champion found a drawdown several times deeper than its first report. It only wins on average. Aggregate results that hide a losing year are the classic pre-blind weakness — which is why the lab's blind campaign gates per year, not in total.
The controls that work
Seal a holdout and touch it once. The lab's version: a strategy freeze, then a separate branch whose entire purpose was testing frozen conclusions on sealed years. The freeze matters as much as the holdout — tuning stopped is tuning you cannot undo the value of. Walk-forward, not in-sample/out-sample. Train only on the past, slide forward, and judge the joined result. Leave-one-out everything. Symbols, folds, years: if removing any one element kills the result, the result was that element. Pre-register the question. One script per numbered run, one falsifiable question, verdict logged before moving on — the journal structure in how the lab was built exists to make result-shopping impossible. Statistical humility: bootstrap intervals and Monte Carlo on every promoted number, monthly stability checks, cost gates from the first prototype.
The vocabulary is the cure
The lab's verdicts — PROMOTE, WATCHLIST, REJECT, RETRACTED, OVERFIT — do the quiet work: WATCHLIST exists so a promising-but-thin result has somewhere honest to sit instead of being promoted by enthusiasm; RETRACTED exists so a withdrawn result stays in the log with its withdrawal dated. The project's biggest published conclusions are negative ones — including a family of retracted results and a failed freeze — and that record is worth more than any single win, because it is the record a stranger can audit. The moment your research journal can embarrass you, it starts protecting you.
If you take one habit from this: when a result looks best, that is the moment to schedule its most hostile audit — the lab's own champion fell to exactly that (the lookahead-bias story), and the failures it joins are sorted in why backtests fail.
Research note: QuantLab experiments are presented for educational and research purposes. Historical backtests and simulations do not guarantee future results and should not be interpreted as investment advice.
Sources
Next