Article

How many trials did you run? The number your backtest hides

By Ignacio Arias, founder of TRAVIDENCE and author of the methodology · Last reviewed: August 2026


A backtest reports the result of the version you kept, not the ones you discarded. The more versions you try, the better the best one looks — even when none has a real edge. Counting honestly how many you tried is the first input for knowing whether your Sharpe is real, and it is the one almost nobody writes down.

A backtest reports the trial that survived

Nobody develops a strategy by testing a single version. You tune a parameter, add a filter, change the timeframe, try another instrument. Each change produces a new backtest, and in the end you keep the one that looks best. The other versions disappear from the screen — but not from the math.

That is the problem. The Sharpe your platform reports is not a single measurement: it is the best of everything you tried. And the best of many attempts always looks better than reality, even when no version has a real edge. Statistics calls this selection bias; in trading it is the most common signal of overfitting — the strategy fitted the noise of the past, not something that repeats.

The logic is the same as any record. If a thousand people call the direction of the market ten days in a row with no information at all, someone will get nine out of ten. Not because they know something — because there were a thousand attempts. Your best backtest may be that person.

How much the bar rises with every trial

Statistical noise produces Sharpe ratios too. Given a short sample and many versions, it produces high ones. The useful question is not "is my Sharpe good?" but "what Sharpe would noise have produced with this many trials and this much data?" That value is the noise bar: if your result does not clear it decisively, it cannot be told apart from what noise alone produces.

Two cases, computed with the Overfitting Calculator (Bailey & López de Prado, 2014 method):

Your Sharpe Trials tested Sample Noise bar Probability it is real
1.8 50 3 years of daily data (756 observations) ≈ 1.32 79.7 %
0.9 1,000 1 year of daily data (252 observations) ≈ 3.26 0.9 %

And a third that needs no calculation: with a single trial there is no search to discount. The bar is zero and the calculation reduces to the probability that the true Sharpe exceeds zero.

Two lessons come out of the table. First: a thousand trials on one year of data push the bar to 3.26 — a Sharpe of 0.9 that looked decent shows up from noise alone almost every time. Second: the sample matters as much as the count. The fewer observations behind the Sharpe, the higher the bar for the same number of trials. Noise grows slowly with each version, but it never stops growing — and it jumps when the data is thin.

The number your backtest hides

A backtest report does not include how many backtests came before it. That number lives in your process — and it is usually larger than you remember. To count it honestly, include everything that produced a different backtest:

  • every parameter combination you evaluated, not just the final one;
  • every filter or rule you added or removed;
  • every instrument and every timeframe you tried it on;
  • every version you discarded because it "did not look right".

With that count, the Bailey & López de Prado (2014) method does two things: it estimates the noise bar for your number of trials and your sample, and it turns your Sharpe into a probability of being real — the Deflated Sharpe Ratio (DSR), which also corrects for the skewness and kurtosis of returns, the fat tails typical of trading. The paper proposes 95 % as the confidence standard; below it, the result cannot be told apart from noise with enough confidence.

A nuance the industry tends to skip: there is no fixed discount. Harvey & Liu (2015) argue that the rule of thumb of "cutting the backtested Sharpe in half" is a mistake: the penalty for testing many versions is not linear. High Sharpes are trimmed a little; marginal ones, a lot. That is why the number of trials matters most exactly where it hurts — in the strategies that look "pretty good".

What to do with that number

  1. Write it down, starting today. A simple log — date, version, what changed, result — turns the number into data instead of memory. It is the discipline that separates a hypothesis from a search.
  2. Compute it. The Overfitting Calculator runs in your browser, free and without signup: Sharpe, sample size, frequency and trials tested. It returns the noise bar and the probability that your Sharpe is real.
  3. Read it as a signal, not a verdict. Clearing the bar does not certify the strategy: it means it warrants a full validation. And failing to clear it does not condemn it: more data, or fewer versions in the search, change the reading.

That number matters enough that it is one of the inputs the TRAVIDENCE audit asks for alongside your trade record.

The calculator evaluates a single overfitting signal: selection bias on the Sharpe ratio. It doesn't analyze your trade record trade by trade, or stability across subperiods, or consistency across volatility regimes, or per-component contribution. That's what the full audit does.

Audit your strategy before you risk capital →

Compute yours: Overfitting calculator →

If your result clears the bar, the next question is how much the path can hurt: Monte Carlo simulator →

How this method is applied inside a full audit: Research →


References

  • Bailey, D. H., & López de Prado, M. (2014). "The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality". The Journal of Portfolio Management, 40(5), 94–107.
  • Harvey, C. R., & Liu, Y. (2015). "Backtesting". The Journal of Portfolio Management, 42(1), 13–28.
Notice. TRAVIDENCE is an independent validation service. It is not financial advice or an investment recommendation. Past performance does not guarantee future results; no validation eliminates the risk of loss.