Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign

Trigger-vs-clock parity (TOST) — does event-driven recalibration reach equivalence with fixed cadence at ¼ the budget?

c-ble-trigger-equivalence · exp-csi-static

Archive snapshot, as of 18 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Sessions

state
Session State Runs Synthesis Criticism Figures Verdict
4MQH07QJ 2026-07-07T10:18 finished 84 Equivalence NOT established at n=6 (TOST fails), and the corpus cannot rule out that the trigger is genuinely worse under adverse drift — not merely underpowered. Paired (trigger[xmodal]−fixed) diff on losin = +0.034 persons (90% CI [+0.002, +0.066], TOST p=0.178 at ±0.05); CI upper bound exceeds the margin. Trigger is cheaper (1.17 recals vs clock 2.0, ≤2 ✓) but worse by an amount not inside ±0.05 — two-edged, not parity. The pilot's n=3 (seeds 0-2, σ=0.0096) sampled only the benign regime; new seeds 3-5 (σ_d=0.0386) exposed a detectability-failure regime — seed 3 fires 0/6 (xmodal 0.0062-0.0066 genuinely sub-threshold 0.0073) while error climbs to 0.67, statistic and error decouple. Secondary detectability also fails at the honest seed-cluster unit: xmodal ρ=0.68 CI [+0.10,+0.90] (lower<0.6); off-path quiet 98.6% passes. Criteria 1/3 (C3 only). recalibration-trigger-from-drift held chosen/plausible, no upgrade. IP-106 warning: the statistic can fail to fire on drift that hurts; log per-state statistic+threshold+true-error together. 168 runs gate-passed (84 reused seeds0-2 + 84 new seeds3-5).

Brief

Why this campaign exists

c-ble-drift-trigger booked its policy criterion as a "clean negative" because triggered recalibration lost to fixed cadence by +0.0079 persons — a margin the 2026-07-06 audit showed is smaller than the seed noise (SD 0.0096) and ≈60× below the ε=0.5 threshold that matters. The superiority framing was the wrong question. The right one is equivalence: does the trigger reach parity with the clock while spending ~¼ of the recalibration budget? Parity is cheaply establishable precisely because the difference is tiny: TOST power at margin ±0.05 with the observed σ_d needs 3–5 seeds; pin 6 (also clears the n=3 sign-test floor for the correlation re-confirm).

Design

Three arms — fixed-cadence / xmodal-triggered / oracle-triggered (ceiling) — paired by seed over the ordered in-path (losin) + off-path (losperp, corner) displacement sequence, reusing the c-csi-layout-drift cast extended to 6 seeds. ~4 policy arms × 6 seeds × 7 states ≈ 168 runs.

Artefact contract (hard gate)

trigger_equivalence.parquet (per seed × arm × state), TOST both-one-sided p-values + the 90% CI on the paired difference, seed-block-bootstrap ρ CIs, all pushed to the session S3 prefix.

Figure render request (mandatory)

  • paired_diff — per-seed (trigger − fixed) dots with the ±0.05 equivalence band drawn (the honest visual for "inside the margin").
  • seed_trace — error vs displacement per arm.
  • The corrected per-arm small-multiple policy figure (all 4 arms; the original dropped the signal arm via zip truncation).

Honest scope

In-silico, one synthetic corridor, self-authored displacement schedule. Tests the parity mechanism, not a field detection rate. May move recalibration-trigger-from-drift candidate→chosen (strength stays plausible); the field ROC remains the IP-106 defeater.