Sealed 2026-08-26T11:16:56Z. The visual form of the MANDATE 4.7 checks.
10 of 10 bins occupied; empty bins are absent, not interpolated. Marker size is the number of observations in the bin.
| Brier (multiclass, 0 to 2) | 0.60139 |
| base-rate reference | 0.47282 |
| improvement over reference | -0.12857 |
| persistence reference | 0.92951 |
| improvement over persistence | 0.32812 |
| Brier (per state, 0 to 1) | 0.30069 |
| base-rate reference | 0.23641 |
| improvement over reference | -0.06428 |
| persistence reference | 0.46476 |
| improvement over persistence | 0.16406 |
Scored against the outcome rule recorded in the bundle: next_bar_state. Bars scored: 1291 of 1333 emitted.
While every detector emits a degenerate distribution, the per-bar multiclass Brier is 0 if the state persists to the next bar and 2 if it flips — so brier_multiclass is exactly twice the state-change rate, and the skill over the base rate is just the state's autocorrelation restated. A detector that never changed state at all would score near-perfectly here while saying nothing whatever. That is not a leak (no future value reaches any emission; the emissions are bit-identical to the pre-E.4 ones) and it is not a reason to withhold the number, but it does mean this metric only begins to measure information content once a detector puts real mass on more than one state. The machinery is being proven now precisely so that it is already trusted when that happens.
Consecutive-bar run lengths per state, mean marked. The same occupancy can be one long stretch or a hundred single-bar twitches.