72 Research

No Arbitrary Rules

Last week we stripped the inherited model back to a single-feature emission to get its sampler moving. It worked, but the choice was arbitrary — a scalar emission is no more "the model the data justifies" than a seven-feature one. So this week opened by turning that discomfort into a rule, and then spent itself letting measurements, rather than arguments, decide what to build. The result is not where we expected: the working model is the plainest thing on the table, not the most elaborate.

The rule

We wrote down the principle we'd violated and the habits it bans. The rule: build the model the data justifies, not the model that is easy to read. In practice it forbids three moves we kept reaching for while debugging:

  • Don't thin the features to make the likelihood behave. If a model misfits in

many dimensions, the fix is the distribution, not deleting the dimensions that carry the regime information.

  • Don't add a knob to pin the number of regimes. The model's own prior is the

correct statement of how many states the data supports; a penalty that forces a tidy count is a preference smuggled in as code.

  • Don't switch a component off to make a chart look clean. Disabling a piece to

debug another is a temporary diagnostic, never a fix.

Behind all three is one rule that governed the rest of the week: a healthy diagnostic is not a correct model. A sampler that mixes, a state count that settles, a parameter that lands on a sensible value — all necessary, none sufficient.

The principled fix we dropped

The honest answer to last week's dimensionality problem was not to delete features but to change the emission distribution: swap the Gaussian for a heavier-tailed Student-t, so a point in the "wilderness" between states is unlikely rather than impossible, and the sampler can move. We built it — per-state tails, learned from the data — and it worked in the right direction: on the frozen window the emission gap fell about 30%. Not enough to unfreeze the chain, and, more to the point, we caught ourselves optimising the wrong thing. The emission gap is a number inside the sampler; the strategy's job is to read regimes in time to size down before a shock and back up into a recovery. Whether the chain prefers three states or twelve is irrelevant if three detect the shock fast enough. We paused the emission work — the implementation is sound and kept — rather than tune a proxy for its own sake.

A scorecard for the actual job, and a fast way to run it

So we built the scorecard the work actually needed. Mark each day from price alone: a "shock" day is one trading well below its recent high, a "calm" day one barely off it. Then ask how often the model has left its baseline state on the shock days and how often it sits in baseline on the calm days, and average the two. Fifty per cent is a coin flip. It rewards being in the right kind of state at the right time and nothing else.

To run that over and over we built a replay harness: fit a model once on history up to the eve of a known episode, freeze it, and forward-filter through the episode one bar at a time. A run that took hours now takes seconds, and fits are cached. The first thing it caught was a bug worth recording: on an early smoke test the model switched state 187 times in 188 days. It had quietly memorised the time of day — every morning bar to one state, every afternoon bar to another — because those features sat in the emission and split the data cleanly. A model with a perfectly healthy-looking fit that was reading the clock. Moving the time features out of the emission and into the transitions fixed it: three switches instead of 187. Exactly the "healthy diagnostic, wrong model" trap the rule was written for.

Letting the data pick the architecture

With a fast loop and an honest metric, the architecture question stopped being a debate. The first cut put the elaborate inherited model — the nonparametric HDP-HMM we'd spent two weeks on — against a plain input-output HMM and the hand-coded rules. The inherited model came last: never above 61%, too sticky to move off baseline. The plain IO-HMM won outright, catching 94% of the COVID shock days while still returning most calm days to baseline. So the working model pivoted off the inherited one. Then the fuller bake-off, four architectures on three windows, one metric:

Modellate 2018COVID 20202022 grindmean
hand-coded rulebook50.049.967.355.7
IO-HMM (frozen)75.581.561.872.9
nonparametric (truncated HDP)57.057.751.655.4
split-merge IO-HMM78.881.561.874.1

Balanced accuracy, percent; 50 is chance. The two input-output HMMs win clearly. We adopted the split-merge variant — narrowly ahead of the frozen one on 2018 and otherwise tied, and it gives flexibility in the state count for free if richer emissions ever justify it. One honesty note: the nonparametric model's row is not a fair verdict on it. A default pruning threshold meant its sparsity mechanism never actually engaged, so that 55% is a statement about one hyperparameter, not about the approach. We logged it as inconclusive, not as a loss.

Why everything collapses to two regimes

The awkward finding underneath the table: every learned model picks two states. Not by accident — it's structural. Adding a third regime costs the model-selection penalty far more than the fit it buys (a few hundred nats of penalty against about a dozen of likelihood), and the input-driven transition matrix grows with the square of the state count, which makes it worse. Even on synthetic data with four regimes planted by hand, it picks two. So the lever for richer regime structure is not the architecture and not the state-count search — it is the emission, what each state is allowed to see. On returns alone, two regimes is all there is.

Standardised SPY daily returns, alone and split by the volatility term structure.
Left: daily SPY returns on their own — one fat-tailed cluster, which is why a model fed only returns separates little more than calm from volatile. Right: the same returns, split by whether the term structure closed in contango (calm) or backwardation (stress) the day before — standard deviation 0.76 versus 2.11, nearly three times wider. The regime structure is plainly there, but along the term-structure axis, not in the returns. Real data, daily, 2006 to date.

What it does, and doesn't, do

Two honest limits to close on. First, the rules still beat the learned models on the slow 2022 grind (67.3 vs 61.8) — because the rules read the term structure directly, while the HMM emits on returns only and a small persistent drift looks like calm to a return-only Gaussian. That is an emission-feature gap, not an architecture failure, and it points the same way as everything above. Second, we tested the model as an actual forecaster — predict tomorrow's direction — and it sits at a coin flip:

Windowmodel sign accuracyalways-up baseline
late 201853.5%51.2%
COVID 202052.3%56.5%
2022 grind48.2%48.8%

Around 50.6% on aggregate, inside the sampling noise. The model is a regime labeler, not a direction forecaster — which is fine for sizing, where being in the right state matters more than calling the next day's sign, but should not be dressed up as prediction.

The week leaves two results side by side: an architecture chosen by measurement rather than argument — and the plainest one, not the inherited showpiece — and a frank account of how little returns alone contain. Both point the next work at the same place, the emission: put the term structure back into what each state sees, carefully, at a dimensionality the model can survive. Which is exactly where the inherited model went wrong by overloading it.