Overfitting detection
It opens by right clicking the Mass Test → 🧪 Analyze the sweep's overfitting… grid. Unlike the Survival Test (which judges one strategy), here the whole sweep is on trial: when you try hundreds of thousands of combinations, the "best" one almost always looks spectacular by pure chance. Two metrics expose it — PBO and DSR — and neither needs you to have set Out-of-Sample data aside beforehand.
PBO — Probability of Backtest Overfitting (0–100%)
It measures how much your own selection process is fooling you. It recombines the data into a great many in-sample/out-of-sample partitions and checks how often the best strategy inside falls below the median outside.
DSR — Deflated Sharpe Ratio (0–1)
It takes the winner's Sharpe and charges it the toll of having tried many: it adjusts for the number of trials, the spread of Sharpes among them and the skew/kurtosis of the returns. It is the probability that the true Sharpe is > 0.
CSCV — how the PBO is computed
Combinatorially Symmetric Cross-Validation. The trick is that the "out of sample" is manufactured by recombining the same data, symmetrically — which is why you do not need to set OOS aside beforehand:
IN vs OUT scatter
Each point is one partition: its in-sample performance (X axis) against its out-of-sample one (Y axis), with the y = x diagonal for reference.
A cloud hugging the underside of the diagonal = degradation = overfitting.
Histogram of logits
The distribution of the λ values across all partitions, with an amber line at λ = 0.
A bell centered well to the right of 0 is what you are after.
PBO bands
DSR bands
- · The matrix is built from the candidates retained in the grid (top K + sample), not from the billions in the sweep. The true N enters through the DSR's real N field.
- · The performance per block uses the trades closed in that stretch; with very few trades per block the PBO is noisier.
- · PBO and DSR complement each other, they do not replace each other: the first judges the selection process; the second, whether the particular winner survives the toll of N.
Method based on Bailey, Borwein, López de Prado & Zhu (PBO/CSCV) and Bailey & López de Prado (Deflated Sharpe Ratio).
AniQuant can be tried free for 30 days, with every module and no card.