DEV Community

Arhan Canli
Arhan Canli

Posted on Originally published at canlicapital.com AI-assisted

I simulated 160,000 backtest searches to see which overfitting corrections keep their promise

If you try twenty versions of a trading strategy and keep the best one, its Sharpe ratio is flattered by the search. Everyone who backtests knows this, and there are standard corrections for it: the deflated Sharpe ratio, the Harvey–Liu haircut, Bonferroni and Šidák adjustments, and bootstrap tests.

Each of those corrections makes a promise. At a 5% level, a strategy with no skill should pass no more than 5% of the time. But each one is derived under assumptions that real returns break: normal returns, independent observations, independent strategies. So I built a benchmark to check the promises where the truth is known.

The setup

The Null Zoo draws complete research searches: 20 strategies, each with 504 daily returns (two years). In the "null" arm, none of the strategies has any skill, so every one of them has a true Sharpe ratio of zero. In the "power" arm, one strategy gets a true annualized Sharpe ratio of 2.

The returns come from eight families, each breaking one assumption:

  • normal returns (the textbook case)
  • fat tails (Student t with 4 degrees of freedom)
  • negative and positive skew
  • volatility clustering (GARCH)
  • autocorrelation (AR(1) with coefficient 0.2)
  • two volatility regimes
  • correlated strategies (every pair correlated at 0.5)

Each validator looks at the whole search, picks the best strategy and returns a p-value. I count how often it calls the best strategy skilled: that is its size when there is no skill, and its power when there is. Each cell runs 10,000 searches, so the Monte Carlo error at 5% is about 0.2 percentage points. That's 160,000 searches in all.

What broke

Autocorrelation is the big one. With a lag-one autocorrelation of 0.2, the tests that treat returns as independent called noise skilled 18–20% of the time at a nominal 5% (the deflated Sharpe ratio is the exception only because it barely passes anything). That includes a bootstrap, because resampling single days scrambles the very dependence that inflates the Sharpe ratio. Deflating the Sharpe ratio for autocorrelation (Lo's factor) brought it back to about 5.5%. The catch: AR(1) is exactly the case that factor assumes, so other kinds of dependence need a different fix.

Skew and fat tails are the stubborn ones. Under negative skew, the standard t-based tests rejected about 8% at a nominal 5%. A non-normal standard error overcorrected (slightly conservative under negative skew) and was liberal under fat tails. A bootstrap was close to right under negative skew, slightly conservative under positive skew and liberal under fat tails. No single correction was right for all three shapes.

A detail I didn't expect: under fat tails, the strategy that wins a search tends to be the one that got a lucky outlier. Its sample skewness averaged +0.6, while the other strategies averaged about zero. Corrections that plug the winner's skewness back in then shrink its standard error, exactly when they shouldn't.

The deflated Sharpe ratio tests a harder question. With the usual rule of accepting when DSR ≥ 0.95, it almost never passed a skill-less strategy. It also passed only about 10% of strategies with a true Sharpe ratio of 2 over two years. That's not a bug in the statistic. The rule asks whether the best strategy beats the expected best of 20 lucky ones, and then asks for 95% confidence on top. That's a much higher bar than "is there any skill". The skilled strategy also widens the spread of the whole search, which raises its own bar. Size-adjusted, the statistic ranks strategies reasonably well; the threshold is what costs the power.

The number of bootstrap resamples matters more than you'd think. With 400 resamples and a correction for 20 strategies, the bootstrap can only reject when none of the 400 resamples beats the observed Sharpe ratio. On the same searches, going to 1,999 resamples raised its power by 4–5 points.

What I'd do differently now

  1. Check the selected strategy's returns for autocorrelation before any multiple-testing correction, and deflate for it if it's there.
  2. When comparing corrections, match their sides (one-sided or two-sided) and their number of resamples. Otherwise you're comparing conventions.
  3. Treat DSR ≥ 0.95 as the demanding rule it is.
  4. If your returns are strongly skewed or fat-tailed, simulate searches shaped like your own returns and check your test's false-positive rate before trusting it.

A conflict of interest, stated

One of the statistics scored here, luck-equivalent trials, is mine. It fails under autocorrelation like the others, and the paper says so. Everything is public so anyone can check or extend it.

If you have a return shape or a search structure you think would break these corrections, tell me and I'll add it to the next version.

Top comments (1)

Collapse
 
ssapable profile image
ssapable •

The "winner of a search is the one that got the lucky outlier" detail matches a much smaller thing I ran this week, at the retail end of the scale. I fitted a fixed take-profit/stop bracket to my partner's futures entries: a 7x7 grid of target and stop multiples, pick the best on the older 70% of trades, score it on the newer 30%. In-sample the winner "made" $2,447 against her actual $1,562 on the same trades, which is just the search talking. Out of sample it was +$445 vs her +$271, and with only 88 holdout trades I'm not willing to call that skill either. What your benchmark makes me want is the honest sentence for the report: with 49 candidates and 88 test trades, how often would noise produce a $170 gap? I don't have that number yet, and I suspect the answer is "often". The autocorrelation result is the one I'll steal: her trades cluster in time, and I was treating them as independent.