c-ble-trigger-equivalence — session 01KWY1C0XCSCAXYNX24MQH07QJ (2026-07-07)
Verdict: strict ±0.05-person equivalence is NOT established at n=6, and the corpus cannot rule out that the trigger is genuinely worse under adverse drift — not merely underpowered. The paired (trigger[xmodal] − fixed) occupancy-error difference on the signal arm is +0.034 persons (90% CI [+0.002, +0.066], TOST p=0.178 at ±0.05) — the CI's upper bound exceeds the +0.05 margin, so TOST fails on the upper one-sided test. The trigger is cheaper (1.17 recalibrations vs the clock's 2.0, ≤2 ✓) but worse by an amount not contained within the margin — a two-edged result, not a parity win. Criteria: 1/3 met (C3); C1 and C2 fail. The pilot's n=3 (seeds 0–2) sampled only the benign regime; the three new placement seeds exposed a detectability-failure regime the pilot never saw.
What we tested
The corpus's first equivalence design. Physics: the c-csi-layout-drift 4-arm × 7-state displacement cast (corner/los/losperp/losin), extended from 3 to 6 placement seeds — 84 runs reused (seeds 0–2, immutable, from 01KWHXSZ… + 01KWMF5D…) + 84 new runs (seeds 3–5), all gate-passed, identical seed-independent furniture schedule, only the placement seed varies. Policies (never/fixed/trigger[xmodal]/oracle-error) are computed in the reduction over each (arm, seed) ordered sequence; estimand = per-seed paired (trigger − fixed) mean occupancy-error diff, unit = seed, n=6, TOST at ±0.05 (t-based, df=5). Reduction csi_trigger_equivalence.py.
What we found
- Equivalence not established (C1 fail) — and the cause is a regime, not a power miss. Per-seed diffs: the pilot's seeds 0/1/2 are tiny (+0.004, +0.001, +0.019 — their spread reproduces the pilot's σ=0.0096 exactly), but the new placements 3/4 are +0.100 / +0.059. This is not the pilot mis-estimating variance — it faithfully characterized a non-representative subpopulation. Seed 3 is the smoking gun: its xmodal statistic sits at 0.0062–0.0066 across every drifted state, genuinely below the 0.0073 calibration threshold (no dropout, no NaN), so the trigger fires 0/6 times — yet true error climbs to 0.67 persons. On that placement the cross-modal disagreement and the true error decouple: the drift hurts accuracy without moving the statistic. So realized σ_d = 0.0386 (4× the pilot's 0.0096) reflects a structural detectability-failure regime the pilot never sampled. The corpus cannot distinguish "underpowered around a benign +0.034 mean" from "the trigger is genuinely worse under adverse in-path drift" — and the right-skewed per-seed diffs plus the seed-3 zero-fire support the latter. This negative should not be softened into a power problem.
- Cheaper, but not at parity accuracy. losin mean error: never 0.519 → fixed 0.402 (2.0 recals) → trigger 0.436 (1.17 recals) → oracle 0.420 (1.0). The trigger costs ~58% of the clock's recalibration spend and is near-silent off-path (losperp 0.17 recals, corner 0.0) — the estate-wide budget argument survives — but it buys that saving at a per-zone accuracy cost that is not contained within ±0.05. The "captures most of the benefit at ~58% of spend" point ratio is seed-fragile (its variance is dominated by 2 of the 6 seeds) and carries no interval; do not read it as an established efficiency figure.
- Detectability re-confirm fails at the honest unit (C2 fail). xmodal Spearman ρ = +0.68 but the seed-cluster bootstrap CI is [+0.10, +0.90] — lower bound 0.10 < the required 0.6. The pilot's ρ=0.93 [0.72, 1.0] was a cell-level bootstrap; at the seed-cluster unit with 6 seeds the correlation is far less certain. The off-path-quiet half passes cleanly (98.6% below threshold ≥ 80%).
- Figure integrity (C3 met). All 4 arms render (n_panels == n_arms asserted on every figure). 6 byte-identical saturation state-pairs flagged — losin state5≡6 on every seed: the displacement cast plateaus at Δ=1.25 m, so losin carries 6 distinct states, not 7. Flagged, not collapsed — harmless for the paired difference (the duplicate state cancels in trigger−fixed) but the per-arm mean traces double-weight the plateau, so "effective n honest" is delivered on the primary, disclosed on the curves.
What it means for the thesis chain
Two powered re-runs in two days (c-ble-graded-count-powered and this one) tell the same story the 2026-07-06 overhaul predicted: n=3 pilots sampled favourable subpopulations that broader replication deflates. The sharper lesson here is that the deflation is not just wider error bars — the new seeds surfaced a qualitatively new failure mode: on some placements the cross-modal disagreement statistic stays below its calibration threshold while accuracy genuinely degrades (seed 3: 0/6 fires, error → 0.67). recalibration-trigger-from-drift stays chosen / plausible — no upgrade (it is already chosen; the equivalence miss and the detectability-decoupling argue against strengthening). The deployable reading is unchanged from the condition-switch work: the trigger's value is budget under a multi-zone estate where most zones never drift, not per-zone accuracy parity. The IP-106 capture must measure the field detection rate (the real defeater) and now carries a concrete warning: the statistic can fail to fire on drift that hurts — log the per-state statistic, threshold, and true error together so the field ROC captures the decoupling regime, not just the aggregate correlation.
Honest scope
In-silico, one synthetic corridor (test-lab-synth-floor-0), self-authored deterministic displacement schedule, σ=2 dB Gaussian BLE noise, idealized refit-on-labelled-state recalibration. TOST is t-based at df=5 (scipy-free incomplete-beta, self-tested against t_{.95,5}=2.015); n=6 seeds is the replication unit. Reduction + per-(arm,seed,state) parquet + policy parquet + 3 figures (paired_diff with equivalence band, seed_trace, policy_smallmultiple) under this session's artefacts/ prefix; 84 new runs attached, 84 prior runs reused by reference.