How we validate

Methodology


The question isn't whether a strategy worked in the past — it's whether it will work on data it has never seen. We evaluate it from your trade history, using the statistical methods that detect the signals of overfitting — the methods and the principles, not the calibration.

The problem: a good backtest is not a good strategy

A strategy can look profitable simply because it was fitted to past data. Test enough variants — parameters, rules, timeframes — and one of them will look excellent by pure chance. That's overfitting: mistaking a pattern that memorized history for an edge that will repeat. On new data, the overfit strategy fails — and that's the difference between passing in a demo and surviving live conditions.

The healthy development cycle is idea → backtest → validate → decide. The backtest proposes; validation disposes. Our job is the third stage: subjecting the backtest's results to tests designed to separate signal from noise — before you risk capital.

Validating on data the strategy has never seen

The antidote to overfitting is out-of-sample evaluation: measuring performance on data that played no part in the strategy's design or tuning. One variant is walk-forward analysis, which moves forward through time and re-tests successively. Running those tests requires the strategy's logic, and that discipline belongs to your development process: if the edge disappears out of sample, it was never an edge.

TRAVIDENCE works on what that process produces: your trade history. With no access to your logic or your parameters, we apply the statistical tests that detect the signals of overfitting in the results — the adjustment for the number of variants tested (DSR), stability across subperiods, and consistency across volatility regimes. That's why we never ask for your strategy: the evidence is in the results.

The methods

Here we describe the core concepts. These are published, peer-reviewed methods; we don't include formulas with constants, or the system's calibration.

Deflated Sharpe Ratio (DSR)

The Sharpe ratio measures risk-adjusted performance, but it's easy to inflate: test many configurations and keep the best one, and that "winning" Sharpe is biased by the search itself. The Deflated Sharpe Ratio corrects that bias — discounting the effect of multiple trials and the non-normality of returns — to estimate whether the edge is real or an artifact of selection.

Probability of Backtest Overfitting (PBO)

The probability of backtest overfitting answers an uncomfortable question: if a configuration was the best in-sample, how likely is it to be merely mediocre out of sample? A high probability is the signature of overfitting — in-sample performance that doesn't survive once the data changes.

Purged cross-validation with embargo (CPCV)

Combinatorial purged cross-validation tests the strategy across many combinations of training and testing segments instead of a single split. To keep the test honest with time series, it purges observations that overlap in time with the test segment and imposes an embargo: a margin discarded around each segment, so that information nearby in time can't leak in and make the validation look better than it is.

Regime stress and component isolation

A strategy can look solid on average and fall apart under specific conditions. That's why we evaluate its performance across different market contexts — in particular, high- and low-volatility regimes — rather than as a single aggregate number. And when a strategy combines several components, we isolate them to distinguish which ones contribute a real edge and which are noise that survived by chance. The certificate reflects these controls in its gate status section, without exposing the internal thresholds.

Tiered bands, rating-agency style

The result is expressed as a band — Robust, Conditional Robustness, Limited Robustness, or Not Robust — and a Robustness Score from 0 to 100. The band is public; the calibration that determines it is not. Publishing the scale while withholding the calibration is deliberate, and it's the same principle a rating agency follows: the market needs to understand what each level means, but the system's integrity depends on the exact thresholds not being gameable from the outside. That's why the cutoffs between bands, the weight of each dimension, and the thresholds of each control are not published.

Independence and audit-grade rigor

We bring to the retail world the same principle behind institutional operational due diligence (ODD): an independent review, performed by a third party with no stake in the outcome. We don't trade strategies, we don't sell funded accounts or challenge evaluations, and we don't charge based on the verdict. We validate from the trade record — not from your code or the parameters that give you your edge. Your strategy stays confidential — and stays yours.

Reproducibility

Every certificate includes a Reproducibility Hash: a cryptographic fingerprint of the verdict and the inputs that produced it. Given the same inputs, the system produces the same verdict and the same hash. That makes the certificate verifiable, not an opinion: anyone can confirm the document wasn't altered after it was issued.


Reference framework

Our methodology draws on published academic work — in particular Marcos López de Prado's: the Deflated Sharpe Ratio (Journal of Portfolio Management, 2014) and the overfitting-detection and purged cross-validation methods described in Advances in Financial Machine Learning (Wiley, 2018). We cite the methods; we don't reproduce proprietary calibrations or parameters.

Notice. TRAVIDENCE is an independent validation service. It is not financial advice or an investment recommendation. Past performance does not guarantee future results; no validation eliminates the risk of loss.

Look up the terms in the glossary →