Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KWMF5DW89Q4PAG63YWKKT1KT

finished 2026-07-03 17:07:03.176898+00:00 → 2026-07-03 17:14:44.275244+00:00 · 21 runs · supervisor: claude-code

“Trigger mechanism CONFIRMED in-silico, with an important policy negative. The deployable cross-modal statistic (xmodal: |CSI-map - BLE-map| disagreement) tracks true occupancy error at Spearman 0.93 [0.72,1.0], detects 100% of error>epsilon states, stays quiet on 100% of control and 94% of error-irrelevant-shift cells — while CSI self-monitoring false-fires on 17% of error-irrelevant shifts. BUT at matched budget even the perfect-knowledge trigger does not lower error below fixed cadence (0.421 vs 0.399); the payoff is parity at one-quarter of the calibration spend (2.0 vs 8.0 recals estate-wide). recalibration-trigger-from-drift: candidate -> chosen, strength stays plausible. 4/5 criteria; critic severity medium, verdict upheld.”

Archive snapshot, as of 10 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Success criteria

CriterionResolved
Corpus: >=2 furniture arms x >=6 ordered displacement states x >=3 placement seeds with ble.enabled; co-registered links + ble_links, gate-passed. yes
drift_trigger reduction emits trigger_eval.parquet with >=3 online statistics + oracle ceiling + calibration-only bootstrap thresholds. yes
Detectability: best non-oracle statistic Spearman >=0.6 in-path AND below-threshold on >=80% of off-path cells. yes
Policy: event-driven trigger mean error <= fixed cadence at matched budget, ~0 recals off-path. no
Framing discipline: in-silico mechanism test; hypothesis moves at most candidate->chosen, strength stays plausible. yes

Synthesis

c-ble-drift-trigger — session 1 (trigger evaluation over the session-2 drift corpus)

Verdict

The trigger mechanism exists in the physics — and the payoff is budget, not accuracy. On an 84-run layout-drift corpus with co-registered BLE (4 arms × 7 states × 3 seeds; 63 runs attached to c-csi-layout-drift session 01KWHXSZKAC8M8T5D9TKX112D8, the 21 signal-arm losin runs attached here), the cross-modal disagreement statistic xmodal — the gap between independently calibrated CSI→count and BLE-RSSI→count maps, computable online without ground truth — passes both detectability halves: Spearman ρ = 0.93 [CI 0.72–1.0] against true occupancy error on the in-path arm, and quiet on 100% of control cells and 94.4% of error-irrelevant-shift cells. recalibration-trigger-from-drift moves candidate → chosen; strength stays plausible (in-silico).

The discrimination result

The corpus deliberately separates environment change from model degradation: the losperp arm shifts per-link mean amplitude ~27 dB with flat count error, while losin (displacement into deep LOS blockage) grows error with spread across placement seeds — never-recalibrate sequence means 0.66 / 0.39 / 0.66 persons by seed (seed 1's placements happen to be drift-robust; the 0.29→0.93 excursion is the seed-0 case, not typical).

  • csi-ks (CSI self-monitoring, no BLE hardware): ρ = 0.61 [0.15, 0.94], but false-fires on 17% of error-irrelevant shifts — a raw distribution-shift detector confuses "the room changed" with "the model broke".
  • ble-ks (BLE channel shift): ρ = 0.62, 0% on losperp.
  • xmodal: ρ = 0.93, 5.6% on losperp — the BLE channel's value here is not as a people-counter (both maps are weak) but as a second, differently-drifting physical reference whose disagreement with CSI isolates prediction-relevant drift.
  • oracle-counter (idealized BLE device-counter, the ble-ground-truth-sufficiency premise): ρ = 0.97 [0.87, 0.99]. The deployable xmodal is statistically indistinguishable from this ceiling at n=18 cells/arm (CIs overlap heavily). Pooled-arm AUCs (0.94–0.997) are reported in the metrics but inflated by the easy no-error arms; the per-arm pair above is the honest test.

Policy at matched budget — the honest negative

Over the 6 drifted states of the signal arm (budget 2 recals), mean error / recals spent:

policy losin error losin recals estate-wide recals (4 arms)
never recalibrate 0.566 0 0
fixed evenly-spaced 0.399 2.0 8.0
trigger[xmodal] 0.407 1.33 2.0
oracle-error (perfect knowledge) 0.421 1.0 2.0

Criterion 4's strict inequality fails (0.407 > 0.399, per-seed: ties on 2/3, +0.019 on one). More telling: even the perfect-knowledge trigger does not beat fixed cadence on error at this budget — recalibrating on schedule also helps below-threshold drift the trigger deliberately ignores. The event-driven payoff is therefore parity at one-quarter of the calibration spend (2.0 vs 8.0 recals across an estate where 3 of 4 zones never drift), not error reduction. Footnote: the policy loop fires slightly more often than the open-loop detectability table (e.g. losperp 0.67 recals/sequence vs 5.6% cell fire-rate) because after each recalibration statistics and thresholds are recomputed against the new calibration state; policy-side false-fires are the deployment-relevant number.

Honest scope

In-silico mechanism test on one synthetic corridor with the analytic, deliberately weak estimator (slope −0.005..−0.03 persons/dB); the displacement arms are self-authored extremes bracketing best/worst case — the discrimination result shows the mechanism is constructible, not that it fires correctly on realistic furniture moves. BLE noise is the σ=2 dB Gaussian stand-in; the oracle σ=0.5 persons is hand-set; recalibration is an idealized refit on labelled state data. The IP-106 hardware capture should log the xmodal statistic alongside its BLE calibration campaigns; the policy claim (event-driven beats fixed) is not earned even in-silico and must not travel into the IP-106 brief.

Artefacts

trigger_eval.parquet, policy_eval.parquet, drift_trigger.metrics.json, fig_drift_trigger_{curves,detect,policy}.png under this session's artefacts/; corpus provenance in c-csi-layout-drift session 01KWHXSZKAC8M8T5D9TKX112D8.

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

Criticism — c-ble-drift-trigger / session 01KWMF5DW89Q4PAG63YWKKT1KT (campaign-critic, severity: medium, verdict upheld)

Criteria audit

  • C1 corpus — supported. 84 runs, 4 arms × 7 states × 3 seeds, ble.enabled, gate-passed. Caveat: the brief's original in-path arm los turned out flat; the effective signal arm is the design-evolution arm losin — the synthesis must say so plainly.
  • C2 reduction — supported. All four statistics + calibration-only thresholds present in the artefacts.
  • C3 detectability — supported. xmodal ρ=0.9275, CI lower bound 0.7227 > 0.6; off-path quiet 100% (corner/los) and 94.4% (losperp). The load-bearing result.
  • C4 policy — unsupported, correctly marked false. trigger[xmodal] 0.40674 > fixed 0.39886 on the signal arm.
  • C5 framing — supported. Promotion candidate→chosen is keyed to the detectability halves per the brief's Expected-interpretation #1, so it survives the C4 failure.

Findings

  • H1 [HIGH] — the draft omitted the oracle-error policy row (0.421 > fixed 0.399): even perfect drift knowledge does not beat evenly-spaced cadence at this budget. Payoff must be reframed as budget-saving at parity, not error reduction. → Applied.
  • H2 [MEDIUM] — estate-wide recal arithmetic was wrong (claimed 2.67/one-third; correct 2.0/one-quarter of fixed's 8). → Applied.
  • H3 [MEDIUM] — detectability fire-rates and policy recal counts don't reconcile (losperp 5.6% cells vs 0.67 recals/seq) because the policy recomputes statistics/thresholds against the moving calibration state. → Footnoted; policy-side false-fires flagged as the deployment-relevant number.
  • H4 [MEDIUM] — verdict led with pooled AUC 0.98, inflated by easy no-error arms. → Demoted; per-arm C3 pair leads.
  • H5 [LOW] — oracle AUC is 0.997 not 1.00; xmodal and oracle CIs overlap heavily — "nearly matching the ceiling" qualified as statistically indistinguishable at this n. → Applied.
  • H6 [LOW] — "0.29→0.93" was the seed-0 best case; seed 1 is nearly flat (→0.43). → Per-seed spread reported.

Scope caveats (carried into synthesis)

One synthetic corridor, one deliberately shallow analytic estimator, self-authored extreme displacement arms — the discrimination result proves the mechanism is constructible on hand-chosen arms, not that it fires on realistic rearrangement. Corpus provenance: 63 runs attached to the parent c-csi-layout-drift session, the 21 losin runs attached to this session (the draft's "84 attached" was corrected). The xmodal statistic is the right thing for the IP-106 hardware capture to log; the policy claim must not travel to the hardware brief.


Addendum — post-hoc statistical re-analysis (2026-07-06, statistics skill, operator session)

Independent recomputation from trigger_eval.parquet, policy_eval.parquet, and drift_trigger.metrics.json under the house statistics contract (.claude/skills/statistics/), cross-checked against a prior audit pass. Every point estimate reproduces exactly with scipy. Verdict direction upheld: the xmodal detectability result survives every stress test — and two of them leave it stronger than the shipped write-up claimed. Four precisifications and one withdrawn caveat follow.

  • [STRENGTHENS C3] The headline detector is not a displacement artefact. The 18 in-path cells are 3 monotone displacement trajectories, not 18 independent draws (per-seed Spearman 0.94–1.00), so the pooled ρ risks being a restatement of "displacement drives both axes" (thompson2012_e41f, ch. 12). Partialling delta_layout out settles which detectors survive that critique: csi-ks collapses to partial-ρ −0.06 and ble-ks to −0.02 (their pooled 0.61/0.62 is just the shared covariate), but xmodal holds at partial-ρ +0.916 and oracle at +0.944. The deployable winner genuinely tracks occupancy error after the displacement trend is removed; the two weak candidates the campaign already discarded do not. This is the honest form of the C3 claim and it is affirmative.

  • [STRENGTHENS C3] The correct-unit resample raises, not lowers, the bound. Re-running the correlation CI at the true replication unit — resampling the 3 seed-trajectories (a seed-block bootstrap; only C(5,3)=10 distinct multisets exist at n=3) rather than individual rows — moves the lower bound up for every deployable statistic: csi-ks 0.15→0.57, ble-ks 0.21→0.62, xmodal 0.72→0.82. Because the within-seed relation is near-perfectly monotone, row-level resampling breaks that structure and manufactures artificially low draws, so the shipped row-level CI is conservative, not too narrow (wasserman2004_ea08, ch. 8.3). xmodal's ρ ≥ 0.6 bar clears comfortably down to n=3 seeds. The seed-block figures are a range over 10 multisets, not a percentile CI, and should be read as such.

  • [MEDIUM] Precision language should name the real n. The synthesis's "statistically indistinguishable … at n=18 cells/arm" overstates the independent information those rows carry — the replication unit is the seed (n=3), and no bootstrap rescues n≈3 (SKILL.md §2; wasserman2004_ea08, ch. 8.3). Correct phrasing: effective replication is 3 placement seeds; the detectability conclusion is robust in direction but its precision is an n=3 statement. The losin arm is thinner still — states 5 and 6 (Δlayout 1.25/1.50 m) are byte-identical across all seeds (a documented Rician-K saturation under total blockage), so the in-path arm carries ≈15 distinct simulated outcomes, not 18. The "100% in-path detection" figure additionally rests on positives from 2 of 3 seeds (seed 1's schedule never crosses ε=0.5).

  • [LOW] The C4 "fail" is a correctly-signed tie, and the session already reads it that way. trigger[xmodal] − fixed = +0.0079 persons (per-seed +0.0040/+0.0009/+0.0188, all same sign), against a fixed-policy seed-to-seed SD of 0.063 — the margin is ⅛ of the noise floor. The exact sign test / Wilcoxon gives p = 0.25, the smallest attainable two-sided p at n=3 pairs (2·0.5³), so no n=3 outcome could reach significance by construction (gentle2020_1ba7, ch. 7.6–7.7). Oracle-error is also worse than fixed on all 3 seeds (+0.0218). The synthesis already reframes this as "parity at ¼ of the calibration spend, not error reduction" and marks C4 false — the substance is disclosed; this bullet only pins the attainable-p ceiling. Note the effect under discussion (0.008 persons) is ≈60× below ε: "parity" is the terminal read, not a hypothesis awaiting a larger n.

  • [MEDIUM — shipped-asset defect] The policy figure never renders the signal arm. _figures() zips a 2-panel subplot against the 4 sorted arms, so fig_drift_trigger_policy.png shows only corner and los — the two flat control arms — and silently drops losin (the in-path arm, the only arm where C4 fails) and losperp. Verified by static read and by rendering the shipped PNG. No claim rests on it — the losin numbers are in the synthesis prose table and metrics.json — but the delivered asset contains no evidence for the campaign's own policy finding, exactly the small-multiple-missing-the-key-arm defect the figures contract exists to catch (.claude/skills/figures/references/statistical-figures.md). Corrected render: fig-policy-corrected.png (all 4 arms, losin first and accented, per-seed min–max ranges instead of bare bars; source fig-policy-corrected.py).

  • [DESIGN] Follow-up sizing. A correlation-only re-confirmation (C3 without C4) is cheap: the seed-block bound above already holds xmodal above 0.6 at n=3, so 5–8 seeds (still caveated per SKILL.md §2) would tighten it materially. A policy re-test (C4) would need ~12–15 seeds for 80–90% power to resolve the observed 0.0079-person margin (paired analogue of diez2015_8380 ch. 5, σ_D=0.0096) — but that margin is practically negligible, so the value is in the budget comparison (2.0 vs 8.0 recals estate-wide at parity), not in chasing significance on the error difference. Any confirmation session should pre-declare the in-path cut and treat the 6 states within a seed as clustered, not as 6 independent observations.

Corrected figures (built from the same parquets, house style tokens): fig-detect-corrected.png (per-seed trajectory scatter making the n=3 clustering visible; source fig-detect-corrected.py) and fig-policy-corrected.png (the dropped-arm fix). Corpus: thompson2012_e41f ch. 12; wasserman2004_ea08 ch. 8.3; gentle2020_1ba7 ch. 7.6–7.7; diez2015_8380 ch. 5; devore2012_62c8 ch. 10.2.

Attached runs

Run Gate Purpose Replay
D1QXAFYJ
PMQ1CYMV
X5CCM542
5KT0AD8Y
MG4ZY2TE
0AKQZ60C
RYE8NDED
F4JVAAV4
9HJRSKH1
9KX2M4VN
SGT20NRH
Z5GJTJVQ
MKBT0SG3
BQF6YWTY
6ZDF6977
FG9Q5Y1T
GS3AQ5HN
6WTF6ZH7
FENWAHQZ
PARJGXWA
AV9W8J7V