notices - See details
Notices
Enterprising Investor Theme - Data and Tech Hero
THEME: TECHNOLOGY
1 September 2026 Enterprising Investor Blog

Would the Backtest Survive a Different Specification?

What institutional investors should ask to uncover configuration risk in systematic strategies

Enterprising Investor Blogs logo thumbnail

  • A strong backtest may depend heavily on one particular model configuration.

  • Sensitivity testing reveals whether results survive reasonable alternative specifications.

  • Allocators should evaluate the range of outcomes, not just the configuration presented.

Every backtest an allocator is shown is a survivor, the one that worked. The configurations that did not work have been quietly set aside. This is not usually dishonest. It is how research proceeds. But it means the equity curve on the page is one path through a very large space of defensible choices, and the manager presenting it may not know how different the picture looks one or two reasonable decisions away.

The question is not simply whether a backtest performs well, but how sensitive that performance is to other reasonable specifications. As the number of variables, parameters, and design decisions grows, so does the risk that an apparently robust result depends heavily on one particular configuration.

This is the fourth and final dimension in a series outlining four dimensions in a broader specification risk framework (SPEC).The Question That Exposes Weak Quant Models examined the specification choices or how variables enter a model; the second, Did the Manager Change the Model or Just the Settings?, looks at the why behind performance failures; and the third, The Question That Reveals Whether a Quant Manager Can Explain a Single Trade considers how the manager explains and accounts for the model’s behavior.

This final blog looks at the configuration choices that could just as easily have been made differently and examines: 

  • Which variables, parameters, and design choices drive the result

  • How performance changes under reasonable alternative specifications

  • Whether the strategy depends too heavily on any single component

  • How wide the range of outcomes is across defensible configurations

The Pattern

Interrogating your manager’s process matters. A manager who has done the work can tell you which components are critical to performance and which are supplementary, and roughly how much performance degrades without each. 

A manager who has not will either not know or will reveal that the strategy collapses without one specific feature, which means the backtest was never a strategy. That risk is not merely theoretical. Research shows that even when analysts start with the same data and the same objective, reasonable methodological choices can lead to materially different results.

In the Nonstandard Errors study published in the Journal of Finance more than 160 research teams were given the same data and the same six hypotheses to test. The dispersion in their results was large, on the order of the standard errors themselves. Same data, same questions, defensible methods throughout, and the answers still spread widely. A single backtest could be one draw from that distribution. The allocator is usually shown the draw and never the spread.

Where Current Frameworks Stop Short

Due diligence focuses on the result and its construction: the Sharpe ratio, the drawdown, how the universe was defined, and how risk is controlled. These questions interrogate the configuration the manager chose. They rarely interrogate the configurations the manager did not choose, which is where fragility lives.

Standard robustness checks, where they exist, are usually performed and reported by the manager on the manager's own terms. The strategy is shown to survive a handful of sensible perturbations. It does not, however, tell the allocator how the strategy behaves across a range of reasonable alternatives another team would have made, or which single component the whole result rests on.

subscribe button

One Question That Changes the Conversation

Ask the manager: “What happens to your model when you remove one variable? How sensitive are the results to specification changes?”

The answer the manager gives is telling: 

  • A strong answer: The manager has run systematic robustness testing and can describe which variables are load bearing and which are supplementary. They can explain performance degradation under variable removal in terms that connect to the model's economic structure, not just its statistics. Crucially, removing any single component may weaken performance, but it does not cause the strategy to collapse. The result is distributed across the design rather than balanced on one point.

  • A standard answer: The manager acknowledges that some signals matter more than others and can offer a qualitative ranking but has not run formal sensitivity analysis. This is common in live strategies and is not disqualifying. It is a gap in structural validation, and the right follow-up is whether the absence of testing reflects a deliberate choice or simply work that was never done.

  • A concerning answer takes one of two forms. The first is evasion about whether sensitivity testing was done at all, which could mean it was not. The second is the admission, sometimes inadvertent, that performance collapses when a specific variable is removed. A model that is catastrophically dependent on a single feature is a concentrated bet on that feature's continued relevance, however sophisticated the surrounding machinery makes it look.

Why This Matters Now

The space of choices has exploded. Modern systematic strategies involve a large number of design decisions, each individually reasonable, that multiply into an enormous space of possible models. The Nonstandard Errors result shows by what degree outcomes can vary across that space even with the same data and the same intent. 

Decay compounds the problem. McLean and Pontiff found that predictor returns fall by 26 percent out of sample and 58 percent after publication. A strategy resting on one dominant feature is exposed twice: once to the chance that the feature was an artifact of data mining, and again to the chance that it was real but is now crowded and decaying. Decay can magnify configuration fragility, particularly when performance depends heavily on one feature. 

This is not an optional test for “robustness.” Sensitivity to specification should be part of understanding a model’s assumptions and limitations. CFA Institute's Standard V(A) requires members to have a reasonable and adequate basis for investment recommendations, including an understanding of the assumptions and limitations of quantitative models.

Before Your Next Meeting

This completes the four specification risk dimensions: how variables enter the model, what managers learn when it fails, whether the model can explain its decisions, and how fragile it is to changes in its configuration. Each dimension fails independently. 

Don't evaluate only the model the manager chose. Evaluate how the result behaves across other reasonable choices the manager could have made.

If you liked this post, don’t forget to subscribe to the Enterprising Investor.

All posts are the opinion of the author. As such, they should not be construed as investment advice, nor do the opinions expressed necessarily reflect the views of CFA Institute or the author’s employer.

Image credit: ©Getty Images