Skip to content

BVI FSC Approved Investment ManagerIBR/AIM/26/2644

Research Note / Quantitative Methodology

Combinatorial Purged Cross-Validation: Why Standard Backtests Overfit

Overlapping labels, serial correlation and a single historical path make ordinary backtests flatter a strategy. Purging, embargo and combinatorial splits are the standard remedy.

Research domain
Cross-validation for financial time series
Key concepts
Purging, embargo, CPCV, Deflated Sharpe Ratio
Illustration of a researcher at a workstation reading a Fourier spectral analysis on a transparent display in a data centre

A backtest that looks excellent is often a backtest that has seen too much of the data. Ordinary cross-validation assumes observations are independent. Financial observations are not: labels overlap, features carry memory, and history offers a single path. This note explains why standard k-fold and walk-forward backtests overfit, how purging and embargo remove the leakage, and how combinatorial purged cross-validation (CPCV) turns one history into many testable paths. The method is the one set out by Marcos López de Prado. It is educational, written for professional investors, and is not an offer or solicitation.

Key takeaways
  • Standard k-fold cross-validation leaks information in financial data because labels overlap in time and features are serially correlated.
  • Purging removes training observations whose labels overlap the test set. An embargo removes a further buffer after it.
  • CPCV tests every combination of k out of N time groups, so a single history yields several complete out-of-sample paths and a distribution of results.
  • CPCV does not correct for the number of trials. It must be paired with a trial record, a Deflated Sharpe Ratio and data the research never touched.

1. Why standard backtests overfit

Three problems compound each other. The first is overlapping labels. If a label is the return over the next 20 bars, an observation at time t and one at t + 5 share most of their label. Put one in training and the other in the test set, and the model has already seen the answer to the question it is being tested on.

The second is serial correlation. Price series, volatility estimates and moving-average features change slowly, so observations just before or after a test window are close relatives of the test data even when their labels do not overlap.

The third is that history provides one path. A walk-forward test trains on the past and tests on the next block, which respects time order but produces a single sequence of results. One path is one draw. It cannot distinguish a robust edge from a fortunate ordering of events, and every parameter adjustment made after looking at that path adds to the selection bias.

2. Purging and embargo

Purging and embargo are the two repairs applied to the training set around each test block.

  • Purging: every training observation whose label window overlaps the time span of the test set is deleted. The model never trains on a label that was partly determined by test-period prices.
  • Embargo: a further buffer of observations immediately after the test set is also dropped from training. It covers leakage through serial correlation that persists beyond the label horizon, which a purge on label overlap alone does not catch.

The embargo is sized to the memory of the data rather than fixed in principle. The note on causal factor-absence placebo testing describes the protocol Qlumina applies, with a buffer of 10 trading days after every test set.

3. How combinatorial purged cross-validation works

CPCV divides the history into N contiguous groups of observations. For each split, k of those groups are held out as the test set and the other N minus k form the training set, with purging and embargo applied around every test block. Repeating this for every possible choice of k groups gives C(N, k) splits.

Because each group appears in the test set in several different splits, the out-of-sample predictions can be stitched together into complete paths, each covering the whole history exactly once but assembled from different models.

Eq. 3.1: number of complete backtest pathsCPCV
φ(N, k) = (k / N) · C(N, k)
N is the number of groups, k the number of test groups per split, C(N, k) the number of splits, and each group is tested in C(N − 1, k − 1) of them.
Groups (N)Test groups (k)Train and test splitsTests per groupComplete paths
621555
63201010
822877
1024599

The distinction between splits and paths matters. With N = 6 and k = 2 there are 15 train and test splits, and from them 5 complete backtest paths can be assembled. The model count is large. The independent paths are fewer.

4. What to do with a distribution of paths

The point of many paths is to stop reporting a single number. A strategy whose Sharpe ratio is high on every path is a different proposition from one whose average is the same but whose paths range from strongly positive to negative. Report the dispersion, the share of paths with a loss, and the worst path.

Two related measures carry the logic further. The Probability of Backtest Overfitting estimates how often the configuration that looks best in sample ranks poorly out of sample. The Deflated Sharpe Ratio adjusts an observed Sharpe ratio for the number of trials, the length of the sample and the shape of the returns. Both depend on an honest count of every strategy and parameter set tried, which is why a trial record is part of the evidence an allocator should ask for.

5. Where CPCV still goes wrong

The method is only as good as its implementation. These are the failure modes worth checking:

1. Preprocessing fitted on the full sample
Normalising, scaling or selecting features with statistics computed over the whole history leaks the test period into training before any split is made. Fit every transformation inside the training fold only.
2. Labels with a longer horizon than the purge assumes
If a label looks 20 bars ahead but the purge uses 5, overlapping observations survive in the training set. The purge window must follow the real label horizon.
3. An embargo shorter than the memory of the features
Features built from long moving averages or volatility estimates stay correlated with the test period for as long as their lookback. The embargo should reflect that memory.
4. Tuning on the CPCV result
Choosing parameters by their CPCV score and then reporting that same score re-creates the overfitting the method was meant to expose. Every trial still counts.
5. Ignoring costs and execution
A clean split structure says nothing about slippage, fees or fill quality. Validation of the signal and validation of the trade are separate questions.

6. CPCV inside a wider validation stack

No cross-validation scheme substitutes for data the research process has never seen. CPCV, however carefully built, is still run on the history the researchers can inspect. For that reason it is one layer among several: placebo controls that test whether a signal survives when the factor is removed, a Deflated Sharpe Ratio that charges for trials, and a physical quarantine of a long out-of-sample window. Qlumina describes the last of these in The 20-Year Blind Out-of-Sample Air-Gap, and the wider cost of a flattering simulation in Microstructure Feasibility.

The same applies to the models themselves: leakage is also how language-model alpha claims fail, as discussed in why LLM alpha models suffer regime collapse.

7. What allocators should ask

An allocator does not need to run the validation to test whether it was done properly. Useful questions for a manager presenting a backtest:

  • How many strategies and parameter sets were tried before this one was chosen?
  • Were training observations purged for label overlap, and what embargo was used and why?
  • Is the result a single walk-forward path or a distribution across many paths?
  • Was any part of the data held back and never examined during research?
  • How do live fills compare with what the simulation assumed?

These sit inside a wider checklist, set out in how to verify a quant track record and, for the broader audit, the 38-point forensic due diligence framework. The investor guide summarises the validation controls applied to programs in its risk and operations section.

8. Further reading

  • Marcos López de Prado, Advances in Financial Machine Learning (Wiley, 2018): the chapters on cross-validation in finance and on backtesting.
  • David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, “The Probability of Backtest Overfitting”, Journal of Computational Finance, 2017.
  • David H. Bailey and Marcos López de Prado, “The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality”, Journal of Portfolio Management, 2014.

Frequently asked questions

01

What is purged cross-validation?

Purged cross-validation is a form of cross-validation for financial time series in which any training observation whose label overlaps in time with the test set is removed before the model is fitted. It prevents information about the test period leaking into training through overlapping labels.

02

What is an embargo in cross-validation?

An embargo removes a further buffer of training observations that immediately follow the test set. It protects against leakage through serial correlation in features and labels that extends beyond the label horizon, which purging alone does not remove.

03

How many backtest paths does combinatorial purged cross-validation produce?

With N groups and k test groups per split, CPCV produces C(N, k) train and test splits and assembles k/N multiplied by C(N, k) complete out-of-sample backtest paths. For N = 6 and k = 2 that is 15 splits and 5 paths.

04

How is CPCV different from walk-forward testing?

Walk-forward testing produces a single historical path, so it yields one performance estimate that depends on that one ordering of events. CPCV produces many paths from the same data, which gives a distribution of outcomes instead of a single number.

05

Does CPCV remove the risk of backtest overfitting?

No. CPCV reduces leakage and shows the dispersion of results, but it does not account for the number of strategies or parameter sets tried. That requires a record of every trial and a correction such as the Deflated Sharpe Ratio, plus data the research process has never touched.

Executive Takeaway

One history is one draw. Ask for the distribution.

Purging and embargo remove the leakage that makes ordinary backtests flattering, and CPCV replaces a single result with a spread of results. It is a necessary discipline, not a sufficient one: a credible record also counts its trials and keeps some data untouched.

Investing involves substantial risk of loss. The value of investments and the income derived from them can fall as well as rise, and investors may not recover the amount originally invested. Past performance is no guarantee of future returns. This note is for professional investors only and is not financial, legal, tax or investment advice. See the risk disclosure.

Validation Standard

See how programs are assessed before they are listed

Managers qualify through a quantitative review, an operational review and a live pilot. The diligence record stays attached to the program.

Research notes are published for professional readers. Not an offer or solicitation. Risk disclosure