A backtest reports what a rule would have produced over historical data. That sounds like evidence, and it is, provided the rule was specified before anyone looked at the results.
In practice rules are rarely specified that way. Parameters get adjusted, date ranges get shifted, filters get added and removed, and the version that gets reported is the one that performed best. Each of those adjustments is a separate test, and the final result reflects the search as much as the market.
This is a statistical problem with a measurable size, and it is distinct from the question of whether a genuine edge erodes once others find it. Overfitting happens before anyone else is involved.
Why Model-Driven Approaches Amplify It
The appeal of an AI investment strategy is the ability to search a large space of possible relationships quickly.
That capability is exactly what makes overfitting more likely rather than less. The severity of the problem scales with the number of configurations tried, and automated search raises that number by orders of magnitude compared with manual work.
The trials that inflate a result include more than the obvious ones:
The last category is the one that breaks most analyses. A researcher reports the surviving strategy and rarely reports how many were tried.
What the Number of Trials Does
The relationship between trials and apparent performance has been formalised, and the implications are uncomfortable.
Work published in the mathematics literature sets out how the minimum backtest length needed to avoid overfitting varies as a function of the number of strategy configurations tested, showing that with enough trials an optimal in-sample Sharpe ratio can be produced from data containing no genuine signal at all.
The Threshold Nobody Reports
The practical form of that finding is a rule of thumb most published backtests fail. If a researcher tests a large number of configurations over a short history, a high Sharpe ratio is the expected outcome even under a null hypothesis of no predictability.
The same work demonstrates this with a seasonal trading strategy built from a modest parameter grid covering entry day, holding period, stop loss and direction. Searching that grid produced a Sharpe ratio above 1.2 on data where no such relationship existed.
That number would pass most informal screening. It is a product of the search, not of the market.
Correcting the Reported Statistic
Once the problem is stated this way, the remedy is a statistical adjustment rather than a change of method.
The relevant correction adjusts the Sharpe ratio for two leading sources of performance inflation, namely selection bias under multiple testing and non-normally distributed returns, helping separate legitimate empirical findings from statistical flukes.
The adjustment requires one input that is almost never disclosed: how many strategies were tested before the reported one was chosen. Without that number, a Sharpe ratio cannot be interpreted, because the same figure means very different things after five trials and after five thousand.
Anyone evaluating a systematic strategy, their own or someone else’s, has a single high-value question available. How many configurations were tried?
Practices That Reduce the Risk
None of this is unavoidable. The mitigations are well established and mostly procedural:
The third point fails most often in practice. Data reserved for validation stops being out-of-sample the moment it informs a second round of adjustments, and few workflows track how many times it has been consulted.
Why This Isn’t an Argument Against Systematic Methods
The findings do not say that quantitative approaches cannot work. They say that a backtest, reported without context about the search that produced it, carries almost no information about future performance.
That is a claim about evidence quality rather than about method. A systematic strategy with a pre-registered hypothesis, a disclosed trial count and a corrected performance statistic is stronger evidence than a discretionary judgment, precisely because those things are checkable.
The practical shift is in what gets asked of a strategy. Not how well did it perform historically, but how many attempts produced that result, and what does the number look like once the search is accounted for. Both questions have answers, and neither requires distrusting the approach itself.
