Does the Edge Persist? A Decade of Refits
The test we'd been putting off: stop hand-picking crisis windows — you chose them, so they flatter the model — and run the strategy straight through a decade, refitting as we go. It turned up a real edge that is fading, and most of the week went on trying to fix the fade and failing, which is how we found out what it is.
A walk-forward decade
We rebuilt the test as a continuous walk-forward backtest. Each year gets its own model, cold-fit on the data available up to that point and no further, then run forward into a year it has never seen; refit, repeat, 2015 to today. Twelve fits, no cherry-picked window, nothing trained on its own test period.
| Series | cumulative return | Sharpe | first-half Sharpe | second-half Sharpe |
|---|---|---|---|---|
| passive (buy & hold) | +334% | +0.83 | — | — |
| regime strategy | +898% | +1.88 | +2.29 | +1.41 |
Over roughly eleven years the strategy roughly triples passive's cumulative return at more than double the risk-adjusted return — a real edge. But split the decade in half and the risk-adjusted return falls by a third, +2.29 to +1.41. That second number is the one that matters.

Localising the decay
The decay is real, so the next job is to say what it is — and the way to do that is to try to fix it and see what doesn't work. Three things it isn't.
Not the crises. The big bear-year alpha — 2018, 2020, 2022 — is intact across the whole decade; as a hedge against stress the model still does its job. On decomposition, the 2025 loss is entirely a recovery-leg problem: the selloff leg still beat passive by about 18 points, the recovery leg gave back about 41. (The 2020 recovery, a clean V, worked fine; 2025's choppy recovery whipsawed every short the model attempted.) What has eroded is the calm-period edge — the years it used to quietly beat a rising market, it now quietly lags.
Not the sizing. The same decay appears whether positions are sized by a conservative fraction of Kelly or clipped to a fixed band. Sizing moves the leverage, not the slope.
Not estimation noise — and this is the one we worked at. We tried three repairs and all three failed. A more thorough state-count search (exhaustive rather than greedy) found a different model but didn't improve the recent years. Pooling thousands of extra training bars per regime — which would average away any mere fitting noise — didn't move 2024 or 2025 at all. And a trust-gated sizing scheme made more raw return but at the same Sharpe, with the decay still there. If the problem were the fit, at least one of those would have helped. None did. So it's the feature, and the model's other candidate feature, from the same family, decays in lockstep:

Does the model know when it's wrong?
A decaying average edge is survivable if the model knows when it has one. The first instinct here was the wrong one: sweep the leverage and watch the returns grow. That tells you nothing — risk-adjusted return is scale-invariant, so multiplying the bet just multiplies the headline. The real question is whether the model bets bigger exactly when it is more often right. So we sorted every bar by the size of position the model wanted and checked both how often it was directionally right and what it earned per unit of leverage.

So the confidence is mostly honest, with one dead spot. A side-note from the same diagnostic: the long bets are cleanly calibrated, while the short bets are right only about half the time and make their money through magnitude when they are — the model calls that something is wrong better than it calls which way.
The leverage we don't use
One result that governs how any of this would be deployed: the bet size we run is, in pure growth terms, far below the mathematically optimal one — on purpose.
| Leverage setting | risk-adjusted return | worst drawdown |
|---|---|---|
| cautious (what we run) | baseline | ≈ −7% |
| growth-optimal | ≈ unchanged | ≈ −86% |
Risk-adjusted return is flat across leverage; only the drawdown moves, from about −7% to about −86% as you push toward the growth-optimal point. Each doubling of leverage buys a roughly fixed increment of long-run growth and roughly doubles the worst drawdown — a straight-line trade, and we sit hard on the cautious end of it.
A smaller version of the same care showed up in refit timing: a model trained right up to the eve of the 2026 shock handled it worse than one trained to a month earlier. The days just before a shock carry its fingerprints into the fit, so when to refit is a real parameter, not a detail.
What persists
The plain reading: there is a real edge here, we can measure it without fooling ourselves, and it is fading in the calm regimes while holding in the stressed ones. The fixes we reached for were sizing and architecture, and the answer was neither — it is the signal itself, aging. A model that decays in a way you can see, localise, and explain is still worth more than one whose backtest you can't trust. The question we carry forward is the uncomfortable one: whether a framework built on stable per-regime behaviour survives if the regimes themselves keep drifting.
