Home › Documentation › Validation & robustness
Analysis 7

Overfitting detection

It opens by right clicking the Mass Test → 🧪 Analyze the sweep's overfitting… grid. Unlike the Survival Test (which judges one strategy), here the whole sweep is on trial: when you try hundreds of thousands of combinations, the "best" one almost always looks spectacular by pure chance. Two metrics expose it — PBO and DSR — and neither needs you to have set Out-of-Sample data aside beforehand.

PBO — Probability of Backtest Overfitting (0–100%)

It measures how much your own selection process is fooling you. It recombines the data into a great many in-sample/out-of-sample partitions and checks how often the best strategy inside falls below the median outside.

→ 0% — the selection holds up out of sample.
→ around 50% — picking the "best" is a coin toss.
→ >50% — an overfitted sweep.

DSR — Deflated Sharpe Ratio (0–1)

It takes the winner's Sharpe and charges it the toll of having tried many: it adjusts for the number of trials, the spread of Sharpes among them and the skew/kurtosis of the returns. It is the probability that the true Sharpe is > 0.

→ close to 1 — it survives multiple testing.
→ close to 0 — probably just luck from trying so many.
Enter the sweep's real N (the true number of combinations you tried, not the ones left in the grid) and click Recompute DSR: with billions of trials the bar rises and the DSR drops to its honest value. It does not re-run any backtest, it is instant.

CSCV — how the PBO is computed

Combinatorially Symmetric Cross-Validation. The trick is that the "out of sample" is manufactured by recombining the same data, symmetrically — which is why you do not need to set OOS aside beforehand:

1. A performance matrix is assembled of S time blocks × N candidates (each cell = the profit of the trades closed in that block).
2. Every way of splitting the S blocks into half In-Sample / half Out-of-Sample is formed (with S=16 → C(16,8) = 12,870 partitions).
3. In each partition the N candidates are ranked by their IS performance → the best one is taken → and you look at what rank it lands in OUT of sample.
4. That relative rank is turned into a logit (λ). A λ ≤ 0 means the best-IS ended up below the OOS median.
5. PBO = the fraction of partitions with λ ≤ 0. The multiple-testing effect enters on its own, through N.

IN vs OUT scatter

Each point is one partition: its in-sample performance (X axis) against its out-of-sample one (Y axis), with the y = x diagonal for reference.

■ Red — the best-IS sinks outside (λ ≤ 0).
■ Green — it holds up out of sample (λ > 0).

A cloud hugging the underside of the diagonal = degradation = overfitting.

Histogram of logits

The distribution of the λ values across all partitions, with an amber line at λ = 0.

■ Mass to the left of 0 = the PBO.
■ Mass to the right = healthy partitions.

A bell centered well to the right of 0 is what you are after.

PBO bands

< 40%robust
40 – 60%doubtful
> 60%overfitted

DSR bands

≥ 0.90solid
0.50 – 0.90weak
< 0.50noise
⚠️
Honest limitations
  • · The matrix is built from the candidates retained in the grid (top K + sample), not from the billions in the sweep. The true N enters through the DSR's real N field.
  • · The performance per block uses the trades closed in that stretch; with very few trades per block the PBO is noisier.
  • · PBO and DSR complement each other, they do not replace each other: the first judges the selection process; the second, whether the particular winner survives the toll of N.

Method based on Bailey, Borwein, López de Prado & Zhu (PBO/CSCV) and Bailey & López de Prado (Deflated Sharpe Ratio).

Try it yourself

AniQuant can be tried free for 30 days, with every module and no card.

← Previous
Walk-forward analysis
Next →
Multi-timeframe analysis
More in Validation & robustness