Stripping Back a Bayesian Regime Model
This week we took delivery of the most ambitious model in the project — an inherited codebase meant to be the finished article — and spent the week establishing that it does not work, and why. It was built as four complex subsystems wired together with no test isolating any one of them, so rather than guess at the cause we wrote down ten ways it could be broken, ranked by how much damage each would do if real, and worked the list from the top.
The inherited model
It sorts the market into volatility regimes like the rule-based strategy, but nothing is hand-set. It is a sticky HDP-HMM: a hidden Markov model whose number of states is inferred rather than fixed, with a "sticky" bias toward staying put so regimes persist rather than flicker. The transitions are input-output — driven by the observed features. A small Transformer reads a recent window and feeds a context vector into those transitions. The emissions — what each state is assumed to look like — are multivariate Gaussians over a vector of volatility features. A per-regime Kelly book sits on top. Four substantial subsystems, coupled, with no isolation tests between them.
The symptom: a sampler that won't move
A model like this is fit by sampling: sweep the data over and over, reassign every bar to a state each pass, and the sequence of labelings is meant to trace the posterior — the spread of explanations the data supports. The first thing to check is whether that exploration is happening at all.
It is not. Over 300 sweeps the labels barely move — about 85 reassignments out of 150,000 chances. Two runs from different random seeds end at identical labelings, which can only mean the chain never left where it started; the stickiness and concentration parameters both collapse toward zero. The sampler is not exploring a posterior, it is sitting on its initializer. A model that cannot move cannot learn, so every richer feature of the design — the inferred state count, the sticky persistence, the input-driven transitions — is decoration on a frozen core.
The cause: emission dimensionality
Each reassignment weighs two things against each other: how well each state explains the current bar (the emission term) versus how likely each state is given where the model has just been (the transition term). For the chain to move, those have to be roughly commensurate — the data has to be able to outvote the prior sometimes. Here the emission term does not just win, it wins by a margin nothing else can touch.
The reason is dimensionality. A Gaussian density narrows by multiplication: every feature added to the emission multiplies another factor below one into the likelihood, so the gap between the best-fitting state and the next-best widens with each dimension.

At the seven-feature emission the gap is roughly ten nats — the assigned state is about twenty thousand times more likely than any rival, and no single-bar reassignment overturns 20,000:1. We ran the obvious alternative explanation to ground: maybe this is really a full-covariance estimation problem, which is genuinely hard in high dimensions. It is not — a diagonal covariance, which sidesteps that estimation entirely, freezes just the same. The pathology is the dimension count, not the covariance shape. And a one-feature control mixes freely, but then the number of states runs away to its ceiling — the opposite failure. The emission was doing two jobs at once.
| Emission | median gap | best vs runner-up | sampler |
|---|---|---|---|
| 7 features, full covariance | ~9.9 nats | ~20,000 : 1 | frozen — ~85 moves in 150,000 |
| 7 features, diagonal covariance | ~9.6 nats | ~15,000 : 1 | frozen — ~71 moves |
| 1 feature (return only) | ~0 nats | ~1 : 1 | mixes — but the state count runs away |
| 1 feature, features in the transitions | 0.2–0.3 nats | ~1.3 : 1 | mixes; settles on 3 states |
The cascade
That reframed the whole flaw list. Several of the ten suspected problems — the state count running away, the improvised hyperparameter updates, the lock-in between the transition weights and the labels — are not independent bugs. They are downstream of one decision: putting a high-dimensional feature vector into the emission. The features were doing double duty — fingerprinting the states in the emission and steering the transitions — and in the emission they were so decisive that everything else became cosmetic.
| Suspected flaw (of ten) | How we tested it | Verdict |
|---|---|---|
| the sampler is too slow to mix | count reassignments over 300 sweeps, two seeds | real — but a symptom |
| the emission is too high-dimensional to compete | likelihood gap vs number of features | the root cause (~20,000:1 at seven) |
| it's the full covariance, not the dimensions | rerun with a diagonal covariance | rejected — same freeze |
| the features are doing two jobs at once | move them out of the emission | confirmed by the fix below |
Stripping back
So we forked a stripped version with three changes. The emission carries a single feature — the return — instead of the full vector. The volatility features move to the transition kernel, where they push the model between states without dictating the labels outright. And the state-creation prior is set to its textbook value rather than an improvised one. The Transformer was switched off for this test — disabled, not removed; putting it back is its own problem once the core is healthy.
The sampler starts to move. Eleven to fifteen thousand reassignments per run instead of eighty-five. The number of states settles on three by itself, with no ceiling fighting to hold it down. The stickiness learns to a market-reasonable level rather than collapsing to nothing. For the first time the machine is exploring the space of explanations instead of restating its starting guess.
What this does and doesn't prove
A sampler that mixes is necessary, not sufficient — a wrong model can mix beautifully. Three things keep us honest about it. The training likelihood still drifts rather than settling. The stickiness may now be set too high, which would make the model sluggish exactly when a regime turns. And most of all, a single return is a thin description of a market state — the entire premise of the project is that the term structure carries regime information a bare return does not. We removed the pathology, but we may have removed signal along with it.
So the strip-back is a diagnosis, not the destination. A scalar emission is itself an arbitrary choice, and "no arbitrary choices" is meant to be the rule. The question it leaves open is the one the next week takes up: what emission distribution lets us put the features back — at a dimensionality the sampler can survive — without freezing it again.
