Stats Backtesting and evidence Walk forward

A reused holdout reported 63% accuracy on pure noise

The leakWhat was measuredApparent resultOnce the leak is closedSource
Reusing the holdoutA classifier fitted to data containing no signal at all, 10,000 samples and 10,000 variablesReported accuracy above 63%50% is the ceiling any classifier can truly achieve on that dataDwork and co-authors, Science
Temporal leakage in published modelsMachine learning models for predicting civil war onsetReported AUC of 0.94 to 0.95, and 0.91 in a second study0.64 and 0.75, and the complex models no longer beat plain logistic regressionKapoor and Narayanan, Patterns
Survivorship in the instrument universeMissing delisting returns in CRSP: only 11.7% of NYSE and AMEX stocks delisted for performance reasons had one, and none of the 3,750 performance delisted Nasdaq stocks didThe published smallest size decile returnLower by 1.45 percentage points a year using collected delisting returns, and by 5.21 points if every performance delisting is set to -100%Shumway, Journal of Finance
Fundamental data timingPoint in time fundamentals against two and three month lagged fundamentals, about 225,000 quarterly observations on more than 5,500 companies from 1994Spreads built on lagged dataTwo month lags overstated point in time returns by up to 15 bp a month, three month lags understated them by up to 14 bp, with cumulative differences reaching +22.3%S&P Capital IQ, vendor research
Which database the fundamentals came fromAccruals sorted on as filed SEC data against the same signal from standardised Compustat data, 19,615 firm years from 2012 to 20180.296% a month from the Compustat version0.673% a month from the as filed version, a difference of 0.377% with a t statistic of 2.46, and about 32% of the stocks in the portfolios differDu, Huddart and Jiang, Journal of Accounting and Economics

The delisting figures are Shumway's two bounds, one using the delisting returns he collected and one assuming a total loss. The point in time figures come from a data vendor's own research note and are labelled as such. The accruals row runs in the opposite direction to the others, in that the cleaner data source gave the stronger result, which is a useful reminder that a data problem is not always a flattering one.

A test set stops being out of sample the moment it is looked at twice. On data containing no signal at all, a reused holdout reported accuracy above 63% where 50% was the ceiling. About twenty quiet iterations manufacture a result significant at 5%.

above 63%Reused holdout accuracy
50%True ceiling
about 20Iterations to a false find
about 1% of sampleEmbargo length

Members only

The rest of this page is for members

Below this point there are 6 sections, 1 chart, 1 table and 8 named sources, roughly 1350 words of it. Every figure carries the source it came from and the date the data is from.

You can keep browsing every statistic in the library for free. The intro and the summary are always open.