Modal detector
36/90 led · p = 0.0001
Held-out half
24/45 led · p = 0.0002
Best generic
13/90 led · p = 0.185
False alarm held
9.1% of 10% target
Transitions led at a matched 10% false alarm — 90 real instabilities
Every detector reads the same two-second pre-onset segments of the same
scenarios, calibrated on the same damped nulls. The tick on each bar is the count expected
by chance at that detector’s own alarm rate; the p-value is the label-permutation
probability of reaching the observed count.
domain-specific modal detector
generic early-warning suite
expected by chance
Why the comparison is fair
- Identical split. Same segments, same labels, same matched false alarm on
the same damped null scenarios — only the detector differs.
- Non-circular labels. A transition is a generator-trip scenario,
a null a damped bus-fault or branch-trip: the label is the disturbance type, a physical
annotation independent of the growth statistic being scored.
- Disclosed data-quality gate. Three scenarios with a non-physical time
column are dropped and counted in the sealed record, never silently.
- Pre-registered operating point. The aggregation and recency weighting
were selected on a development half and validated on the held-out half (24/45,
p = 0.0002) before this full-corpus comparison.
The sealed record
The whole comparison — corpus, operating point, every detector’s count and
p-value, the drops, the verdict — is one content-addressed artefact in the public
repository (examples/real_data/psml_modal_growth/), guarded by an integrity
test that recomputes the hash from the committed payload.
Sealed head-to-head record
bc6895879088b31b…
Claim boundary (read before quoting)
- This certifies the detector on this corpus. The certified numeric operating
point is dataset-specific: our own cross-dataset test on ISO-NE PMU captures showed the
frozen threshold does not port — deployment calibrates per system on its own
ambient data, and that step is part of the product, not a caveat hidden in a footnote.
- “Generic detectors at chance” is a statement about this task and operating
point, sealed with p-values — not a blanket dismissal of those methods elsewhere.
- Streaming deployment is stricter than per-window scoring: the sealed stream operating
point holds 11/45 held-out at a 10% stream false alarm, measured and sealed separately.