Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KWY1C0XCSCAXYNX24MQH07QJ

finished 2026-07-07 10:18:23.532119+00:00 → 2026-07-07 11:00:39.565540+00:00 · 84 runs · supervisor: react-agent

“Equivalence NOT established at n=6 (TOST fails), and the corpus cannot rule out that the trigger is genuinely worse under adverse drift — not merely underpowered. Paired (trigger[xmodal]−fixed) diff on losin = +0.034 persons (90% CI [+0.002, +0.066], TOST p=0.178 at ±0.05); CI upper bound exceeds the margin. Trigger is cheaper (1.17 recals vs clock 2.0, ≤2 ✓) but worse by an amount not inside ±0.05 — two-edged, not parity. The pilot's n=3 (seeds 0-2, σ=0.0096) sampled only the benign regime; new seeds 3-5 (σ_d=0.0386) exposed a detectability-failure regime — seed 3 fires 0/6 (xmodal 0.0062-0.0066 genuinely sub-threshold 0.0073) while error climbs to 0.67, statistic and error decouple. Secondary detectability also fails at the honest seed-cluster unit: xmodal ρ=0.68 CI [+0.10,+0.90] (lower<0.6); off-path quiet 98.6% passes. Criteria 1/3 (C3 only). recalibration-trigger-from-drift held chosen/plausible, no upgrade. IP-106 warning: the statistic can fail to fire on drift that hurts; log per-state statistic+threshold+true-error together. 168 runs gate-passed (84 reused seeds0-2 + 84 new seeds3-5).”

Archive snapshot, as of 19 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Success criteria

CriterionResolved
Primary (equivalence): 90% CI on paired (trigger−fixed) diff within ±0.05 persons (TOST both-one-sided p<0.05), trigger ≤2 of 6 recal states. — FAIL: diff +0.034, 90% CI [+0.002, +0.066] exceeds +0.05, TOST p=0.178. Budget half met (1.17≤2) but conjunction false. Seed-3 detectability-failure regime (0/6 fires, error→0.67) the n=3 pilot never sampled. no
Secondary (correlation re-confirm): xmodal Spearman seed-block-bootstrap CI lower ≥0.6 on losin AND statistic below threshold on ≥80% off-path cells. — FAIL: ρ=0.68 but seed-cluster CI [+0.10,+0.90], lower 0.10<0.6. Off-path quiet 98.6% passes; conjunction fails. no
Figure integrity: policy figure renders all 4 arms; byte-identical losin state5≡6 saturation rows collapsed/flagged. — MET: 4 arms render (asserted); 6 losin state5≡6 saturation pairs flagged (cast plateaus at Δ=1.25m); cancels in paired diff, disclosed on per-arm traces. yes

Synthesis

c-ble-trigger-equivalence — session 01KWY1C0XCSCAXYNX24MQH07QJ (2026-07-07)

Verdict: strict ±0.05-person equivalence is NOT established at n=6, and the corpus cannot rule out that the trigger is genuinely worse under adverse drift — not merely underpowered. The paired (trigger[xmodal] − fixed) occupancy-error difference on the signal arm is +0.034 persons (90% CI [+0.002, +0.066], TOST p=0.178 at ±0.05) — the CI's upper bound exceeds the +0.05 margin, so TOST fails on the upper one-sided test. The trigger is cheaper (1.17 recalibrations vs the clock's 2.0, ≤2 ✓) but worse by an amount not contained within the margin — a two-edged result, not a parity win. Criteria: 1/3 met (C3); C1 and C2 fail. The pilot's n=3 (seeds 0–2) sampled only the benign regime; the three new placement seeds exposed a detectability-failure regime the pilot never saw.

What we tested

The corpus's first equivalence design. Physics: the c-csi-layout-drift 4-arm × 7-state displacement cast (corner/los/losperp/losin), extended from 3 to 6 placement seeds — 84 runs reused (seeds 0–2, immutable, from 01KWHXSZ… + 01KWMF5D…) + 84 new runs (seeds 3–5), all gate-passed, identical seed-independent furniture schedule, only the placement seed varies. Policies (never/fixed/trigger[xmodal]/oracle-error) are computed in the reduction over each (arm, seed) ordered sequence; estimand = per-seed paired (trigger − fixed) mean occupancy-error diff, unit = seed, n=6, TOST at ±0.05 (t-based, df=5). Reduction csi_trigger_equivalence.py.

What we found

  1. Equivalence not established (C1 fail) — and the cause is a regime, not a power miss. Per-seed diffs: the pilot's seeds 0/1/2 are tiny (+0.004, +0.001, +0.019 — their spread reproduces the pilot's σ=0.0096 exactly), but the new placements 3/4 are +0.100 / +0.059. This is not the pilot mis-estimating variance — it faithfully characterized a non-representative subpopulation. Seed 3 is the smoking gun: its xmodal statistic sits at 0.0062–0.0066 across every drifted state, genuinely below the 0.0073 calibration threshold (no dropout, no NaN), so the trigger fires 0/6 times — yet true error climbs to 0.67 persons. On that placement the cross-modal disagreement and the true error decouple: the drift hurts accuracy without moving the statistic. So realized σ_d = 0.0386 (4× the pilot's 0.0096) reflects a structural detectability-failure regime the pilot never sampled. The corpus cannot distinguish "underpowered around a benign +0.034 mean" from "the trigger is genuinely worse under adverse in-path drift" — and the right-skewed per-seed diffs plus the seed-3 zero-fire support the latter. This negative should not be softened into a power problem.
  2. Cheaper, but not at parity accuracy. losin mean error: never 0.519 → fixed 0.402 (2.0 recals) → trigger 0.436 (1.17 recals) → oracle 0.420 (1.0). The trigger costs ~58% of the clock's recalibration spend and is near-silent off-path (losperp 0.17 recals, corner 0.0) — the estate-wide budget argument survives — but it buys that saving at a per-zone accuracy cost that is not contained within ±0.05. The "captures most of the benefit at ~58% of spend" point ratio is seed-fragile (its variance is dominated by 2 of the 6 seeds) and carries no interval; do not read it as an established efficiency figure.
  3. Detectability re-confirm fails at the honest unit (C2 fail). xmodal Spearman ρ = +0.68 but the seed-cluster bootstrap CI is [+0.10, +0.90] — lower bound 0.10 < the required 0.6. The pilot's ρ=0.93 [0.72, 1.0] was a cell-level bootstrap; at the seed-cluster unit with 6 seeds the correlation is far less certain. The off-path-quiet half passes cleanly (98.6% below threshold ≥ 80%).
  4. Figure integrity (C3 met). All 4 arms render (n_panels == n_arms asserted on every figure). 6 byte-identical saturation state-pairs flagged — losin state5≡6 on every seed: the displacement cast plateaus at Δ=1.25 m, so losin carries 6 distinct states, not 7. Flagged, not collapsed — harmless for the paired difference (the duplicate state cancels in trigger−fixed) but the per-arm mean traces double-weight the plateau, so "effective n honest" is delivered on the primary, disclosed on the curves.

What it means for the thesis chain

Two powered re-runs in two days (c-ble-graded-count-powered and this one) tell the same story the 2026-07-06 overhaul predicted: n=3 pilots sampled favourable subpopulations that broader replication deflates. The sharper lesson here is that the deflation is not just wider error bars — the new seeds surfaced a qualitatively new failure mode: on some placements the cross-modal disagreement statistic stays below its calibration threshold while accuracy genuinely degrades (seed 3: 0/6 fires, error → 0.67). recalibration-trigger-from-drift stays chosen / plausible — no upgrade (it is already chosen; the equivalence miss and the detectability-decoupling argue against strengthening). The deployable reading is unchanged from the condition-switch work: the trigger's value is budget under a multi-zone estate where most zones never drift, not per-zone accuracy parity. The IP-106 capture must measure the field detection rate (the real defeater) and now carries a concrete warning: the statistic can fail to fire on drift that hurts — log the per-state statistic, threshold, and true error together so the field ROC captures the decoupling regime, not just the aggregate correlation.

Honest scope

In-silico, one synthetic corridor (test-lab-synth-floor-0), self-authored deterministic displacement schedule, σ=2 dB Gaussian BLE noise, idealized refit-on-labelled-state recalibration. TOST is t-based at df=5 (scipy-free incomplete-beta, self-tested against t_{.95,5}=2.015); n=6 seeds is the replication unit. Reduction + per-(arm,seed,state) parquet + policy parquet + 3 figures (paired_diff with equivalence band, seed_trace, policy_smallmultiple) under this session's artefacts/ prefix; 84 new runs attached, 84 prior runs reused by reference.

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

campaign-critic — c-ble-trigger-equivalence / 01KWY1C0XCSCAXYNX24MQH07QJ

Verdict: agrees with synthesis (equivalence NOT established at n=6). Severity: medium — one load-bearing framing correction, not a criterion misrepresentation. Every headline number matches the artefact; the TOST is textbook-correct (both-one-sided at ±0.05, p_TOST=max(p_lower=0.0016, p_upper=0.178)=0.178; the 90% CI is the right (1−2α) equivalence interval; upper bound +0.066 > +0.05 → not equivalent).

Claim audit

  • C1 (equivalence + ≤2 recals): correctly failed. Conjunction is FALSE (equivalence fails though the ≤2-recals half passes at 1.17). TOST correctly applied.
  • C2 (detectability): correctly failed. ρ=0.68, seed-cluster CI lower 0.10 < 0.6; off-path 98.6% ≥ 80% passes; conjunction fails.
  • C3 (figure integrity): met, lightly. 4 arms render; 6 saturation pairs flagged, not collapsed — harmless for the paired difference (cancels), disclosed on the per-arm mean traces.

Findings integrated into the synthesis

  1. The "pilot underestimated σ / underpowered" framing was mechanism-flattering (the medium-severity catch). The pilot σ=0.0096 reproduces the seeds-0–2 spread exactly in this run — a faithful estimate of a non-representative subpopulation, not a mis-estimate. The 4× inflation is a structural detectability-failure regime seeds 3–5 introduced. Supervisor verified seed 3: xmodal statistic 0.0062–0.0066 across all drifted states, genuinely sub-threshold (0.0073), 0/6 fires, no NaN — a real decoupling of statistic from error, not a reduction dropout. Synthesis rewritten: the corpus cannot distinguish "underpowered around a benign mean" from "genuinely worse under adverse drift," evidence favouring the latter.
  2. "Budget parity holds" foregrounded the favourable axis. Rewritten as "cheaper but at a per-zone accuracy cost not contained within ±0.05."
  3. "~72% benefit at ~58% spend" was a bare point ratio. Now marked seed-fragile (variance dominated by 2 of 6 seeds), no interval claimed.
  4. Coincidence spot-check. Seed-3 diff (0.10031061) and ρ-CI-lower (0.10031104) coincide to 5 sig figs — supervisor confirmed genuinely coincidental (differ 4.3e-7, independent code paths). No cross-wire.

Honest scope

In-silico, one synthetic corridor, self-authored displacement schedule, idealized refit recalibration, n=6 seeds, t-based TOST df=5. The negative is real and is not merely a power problem — the right-skewed per-seed diffs and the seed-3 zero-fire actively support a genuine trigger-detectability failure under adverse in-path drift.

Attached runs

Run Gate Purpose Replay
Y6392K75
GKZ5B12B
C3KSMGB8
WCSPNH5Y
PXGBYS2A
S8FZ6PKR
149VX599
G5KNB4XA
NPR7AFW0
KKR30279
9N9T7AVB
WG9KW2WW
672E0MBQ
5FPW59AW
C1J6RWT6
AHN27M2D
6PJ7BY0P
3B6CEDNN
MGXGRCZS
XCPC2J8V
H61WJM4S
08DTAX2R
80ZEJE3D
2BPH4Q0B
CB31YXQB
07X5CJ1X
WFQERAPE
EGDMW5YY
2YSRFSG6
4ZDYPWHW
AZGMW8DK
CFH9WA4J
6FE8A5ZY
N0PQ4QAM
F5GTANY3
44QYDWNX
P2P8543S
X9N2EDZ8
KV10N1XX
1N9A2ZYW
V2ED9Y76
BYT226TC
GD1XH9RG
YQ2FD2YB
BR3Q53JK
N3G2HE4Q
CE1RSEXP
BV0SNDET
9V01M0ZE
HEE1RQ1Z
MB3EWHKB
YQRPR97D
XEZ4G7T4
FGH6B2Y9
GA13HY3D
ZV5F2H36
38EX6E0X
1HZ0JBBZ
8DH3SD6P
AD8SVT44
HPMBRKAD
B2N42H9R
846MGBMH
WRZMS5DZ
F3NQ23Q1
NXA12T2V
0HZBJVXF
TX3BBX11
5F8TQP2W
719C1Y91
7WR6YH1R
GCWNSBYX
6QNQBZXQ
QV3MQ9HA
X5FMH2QM
9QNGEHG2
GE5X3C23
JDF1J06G
8BCRNZDA
ECEHY0MP
DBJ07TQE
37Y2J8J5
QE3NBWHF
9Z4TW6W4