Stats › Backtesting and evidence › Walk forward
A reused holdout reported 63% accuracy on pure noise
| The leak | What was measured | Apparent result | Once the leak is closed | Source |
|---|---|---|---|---|
| Reusing the holdout | A classifier fitted to data containing no signal at all, 10,000 samples and 10,000 variables | Reported accuracy above 63% | 50% is the ceiling any classifier can truly achieve on that data | Dwork and co-authors, Science |
| Temporal leakage in published models | Machine learning models for predicting civil war onset | Reported AUC of 0.94 to 0.95, and 0.91 in a second study | 0.64 and 0.75, and the complex models no longer beat plain logistic regression | Kapoor and Narayanan, Patterns |
| Survivorship in the instrument universe | Missing delisting returns in CRSP: only 11.7% of NYSE and AMEX stocks delisted for performance reasons had one, and none of the 3,750 performance delisted Nasdaq stocks did | The published smallest size decile return | Lower by 1.45 percentage points a year using collected delisting returns, and by 5.21 points if every performance delisting is set to -100% | Shumway, Journal of Finance |
| Fundamental data timing | Point in time fundamentals against two and three month lagged fundamentals, about 225,000 quarterly observations on more than 5,500 companies from 1994 | Spreads built on lagged data | Two month lags overstated point in time returns by up to 15 bp a month, three month lags understated them by up to 14 bp, with cumulative differences reaching +22.3% | S&P Capital IQ, vendor research |
| Which database the fundamentals came from | Accruals sorted on as filed SEC data against the same signal from standardised Compustat data, 19,615 firm years from 2012 to 2018 | 0.296% a month from the Compustat version | 0.673% a month from the as filed version, a difference of 0.377% with a t statistic of 2.46, and about 32% of the stocks in the portfolios differ | Du, Huddart and Jiang, Journal of Accounting and Economics |
The delisting figures are Shumway's two bounds, one using the delisting returns he collected and one assuming a total loss. The point in time figures come from a data vendor's own research note and are labelled as such. The accruals row runs in the opposite direction to the others, in that the cleaner data source gave the stronger result, which is a useful reminder that a data problem is not always a flattering one.
A test set stops being out of sample the moment it is looked at twice. On data containing no signal at all, a reused holdout reported accuracy above 63% where 50% was the ceiling. About twenty quiet iterations manufacture a result significant at 5%.
Members only
The rest of this page is for members
Below this point there are 6 sections, 1 chart, 1 table and 8 named sources, roughly 1350 words of it. Every figure carries the source it came from and the date the data is from.
You can keep browsing every statistic in the library for free. The intro and the summary are always open.