Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KT98WKH2HYQYP4EP79G09CEG

finished 2026-06-04 12:14:20.194313+00:00 → 2026-06-04 12:58:50.895280+00:00 · 72 runs · supervisor: react-agent

“All three criteria met on the 72-run design of record: 12/12 per-floor monotonicity gates (ρ ≥ +0.83), and cross-geometry drift demonstrated at LOFO inflation ×3.8 (2.4 GHz) / ×7.7 (5.0 GHz) — scoped to in-silico/ResPlan/N=6/seed-0 — establishing the geometric pillar for periodic BLE recalibration; hypothesis stays 'plausible' pending multi-site field data. Critic caught a misattributed Wi-CaL ratio pre-seal; corrected to the derived ~×1.2–×2.3 range, explicitly marked as ours.”

Archive snapshot, as of 9 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Success criteria

CriterionResolved
Every floor's scene stages cleanly (≥1 Tx, ≥1 Rx, walls from PostGIS) and every (floor, n_agents, band) run returns the IP-085 scalar floor + links.parquet; no NaN/Inf. yes
Within each floor, CV rises monotonically with occupancy (Spearman ρ(n_agents, CV) ≥ +0.5) — the per-environment sanity gate. yes
The cross-floor spread of the CV-vs-occupancy slope is reported with a leave-one-floor-out MAE: if the held-out MAE inflates ≥ 1.5× the within-floor MAE, cross-geometry drift is demonstrated (supports periodic recalibration); if not, a geometry-robust proxy is reported instead. yes

Synthesis

What we tested

Criterion 1 (scene staging + artefact completeness): Six floors from the ResPlan dataset — resplan-12439, resplan-7421, resplan-12419, resplan-1374, resplan-147440, resplan-16157 — were each swept across 6 occupancy levels (n_agents ∈ {0, 2, 4, 6, 8, 12}) and 2 bands (2.4 GHz, 5.0 GHz) at n_placements = 48, seed 0, body radius 0.20 m / body loss 5.0 dB (6 × 6 × 2 = 72 runs). Criterion 2 (per-floor CV monotonicity): all 12 (floor, band) pairs were tested via Spearman ρ(n_agents, CV). Criterion 3 (cross-floor drift): a leave-one-floor-out (LOFO) linear fit on CV-vs-occupancy slope was evaluated at both bands.

What we found

Criterion 1: All 72 runs returned gate_passed = true (exit code 0); every run carries links.parquet, csi.hdf5, and the IP-085 schema file with no reported NaN/Inf in the reduction. The pre-2026-06-04-pm batch (floors 12439, 7421, 12419; ~36 runs) lacked the delay_spread_ns column; the later 36 runs carry it. This is an additive schema difference with no impact on CV or LOFO statistics. Criterion 2: Spearman ρ ≥ +0.83 on every (floor, band) cell; 10 of 12 cells reach ρ ≥ +0.94. The weakest cell (resplan-12439, 5.0 GHz, ρ = +0.83) still clears the +0.5 gate. All 12/12 cells pass. Criterion 3: At 2.4 GHz, within-floor MAE = 0.019 vs. LOFO MAE = 0.073 — an inflation of ×3.81. At 5.0 GHz, within-floor MAE = 0.017 vs. LOFO MAE = 0.134 — an inflation of ×7.74. Both bands exceed the ×1.5 threshold; cross-geometry drift is demonstrated. Three robustness limits bound this conclusion: the LOFO statistic rests on N=6 floors, a single seed (seed 0), and the ~2× spread between the bands' inflation (×3.81 at 2.4 GHz vs ×7.74 at 5.0 GHz) is unexplained — the next session (ZInD floors, additional seeds) should tighten all three.

What it means for the thesis chain

All three success criteria are met. The result feeds the thesis/csi-sensing chapter argument for ble-periodic-calibration: a model trained on one ResPlan apartment geometry does not transfer to another — the in-silico geometric component of the deployment-drift penalty is ×3.8–×7.7 (in-silico, ResPlan, N=6 floors, seed 0, pooled linear CV~N fit). For reference, Wi-CaL (choi2022_17c2) reports within-session k-fold MAE (0.13–0.18 seminar / 0.32–0.36 meeting) and leave-one-session-out MAE (0.35 / 0.41) separately; the paper publishes no inflation ratio, but ratios derived by us from those MAEs span ~×1.2 (seminar room) to ~×2.3 (meeting room). The comparison is directionally consistent, but these are not the same quantity: Wi-CaL's sessions share a single geometry so its drift is temporal and environmental; ours is purely geometric with protocol and body model held fixed. Our in-silico geometric drift and Wi-CaL's real-world drift point the same direction but measure different quantities (geometric vs temporal/environmental), so neither magnitude validates the other. The hypothesis strength remains 'plausible': the real defeater is multi-site field data. A natural follow-on is adding ZInD floors as a second dataset, and a furnished-floor seated variant once IP-092 occupiables exist.

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

Claim audit

  • C1 (staging + artefacts): supported. Critic self-verified 4 runs across 3 floors at N {0,4,12}, both bands — all gate_passed, complete artefact inventories, sensible domain metrics. 72 runs attached = 6×6×2. The delay_spread_ns schema delta between batches is real (differing image digests) and correctly flagged additive.
  • C2 (per-floor monotonicity): supported, with the caveat that the exact ρ table comes from the writer's host-side reduction (per-run variance_monotonicity corroborates; not independently recomputed from the 72 parquets).
  • C3 (cross-floor drift): borderline → resolved with scoping. The LOFO numbers clear the ×1.5 gate decisively, but rest on a 5-point pooled linear fit predicting one held-out floor — N=6 floors, single seed, and the ~2× band gap (×3.81 vs ×7.74) unexplained. These limits are now stated in the synthesis where criterion 3 concludes.

Weak claims (all integrated by the patcher before close)

  • CRITICAL — Wi-CaL misattribution: "reports LOSO MAE inflation of ×2.3–×2.7" was not in choi2022_17c2; the paper reports MAEs separately and never a ratio; honest derived range is ~×1.2 (seminar) to ~×2.3 (meeting). Replaced with the derived range explicitly marked as ours.
  • "which is expected" removed; the comparison compressed to one sentence with no mutual magnitude validation.
  • The ×3.8–×7.7 headline scoped to (in-silico, ResPlan, N=6 floors, seed 0, pooled linear CV~N fit).
  • Robustness limits (N=6, single seed, band gap) added at the criterion-3 conclusion.

Honest scope

The corpus isolates pure geometric drift: 6 ResPlan floors, protocol/body model/seed fixed, ray-traced (not measured) CSI. It cannot speak to temporal/environmental drift — the axis Wi-CaL measures and the axis periodic BLE recalibration ultimately targets — nor to furnished/seated occupancy (IP-092 occupiables absent on these floors), other datasets, or hardware fidelity. The headline is a geometric component, not a deployment-drift magnitude.

(4/4 findings applied by the campaign-patcher: 1 critical, 3 medium — see patch-log.json in the session scratch dir.)


Addendum — post-hoc statistical re-analysis (2026-07-06, statistics skill, operator session)

Independent recomputation from the 72 raw links.parquet + metadata.json artefacts (no persisted reduction table existed to check against) under the house statistics contract (.claude/skills/statistics/). The campaign's own arithmetic reproduces to the last decimal — within_mae/lofo_mae and the ×3.81 (2.4 GHz) / ×7.74 (5.0 GHz) inflation ratios match the sealed synthesis. Verdict direction upheld (cross-geometry drift is real and supports periodic recalibration); four precisifications, two of which qualify the magnitude the thesis chain will quote and its robustness at 2.4 GHz.

  • [HIGH] The ×1.5 decision statistic compares an in-sample residual to an out-of-sample error. within_mae is a degree-1 CV~N fit evaluated on the same six points it was fit to; lofo_mae is genuinely held-out (fit on 5 floors, test on the 6th). The ratio conflates the mechanical train/test gap with cross-geometry drift and is biased upward regardless of true drift (wasserman2004_ea08, ch. 13.6/22.8; SKILL.md §6 "never quote training error"). Replacing the denominator with an honest leave-one-occupancy-out CV within the same floor gives ×2.41 (2.4 GHz) and ×4.69 (5.0 GHz) — a −37%/−39% correction. Both still clear ×1.5, so the qualitative conclusion survives; the headline magnitude does not. Any thesis-chain restatement should quote the honest ratio or both.

  • [HIGH] The one number the ble-periodic-calibration chain leans on carries no uncertainty, and the honest 2.4 GHz interval straddles the gate. The replication unit is N=6 floors, not 72 runs (frames/links within a run are technical replicates of one geometry — campaign-design.md §2.4/§4). A percentile bootstrap over the 6 floors (B=1999; distinct-resample ceiling C(11,6)=462, so the interval is coarse and stays caveated per SKILL.md §2, n≤8) gives 2.4 GHz 95% CI (1.33, 11.39) — lower bound below the campaign's own ×1.5 gate — and 5.0 GHz (3.89, 18.97), clear. On the honest comparator above the 2.4 GHz interval widens further to (0.82, 7.64) with 25.7% of resamples below ×1.5; 5.0 GHz stays clear at (2.30, 11.85). The deterministic drop-one-floor jackknife agrees in direction (2.4 GHz per-floor-dropped ratios 2.17–5.13x; jackknife-normal SE crosses zero, itself evidence the CLT interval is inappropriate at n=6). The point-estimate gate passes; the CI-lower-bound-clears-gate standard does not at 2.4 GHz. "Both bands exceed the threshold" should be softened to "5.0 GHz clears the gate with margin; 2.4 GHz is consistent with clearing it but the floor-level interval reaches below it — a multi-seed, multi-dataset rerun is required." (The synthesis already flagged N=6/single-seed qualitatively; this quantifies it.)

  • [MEDIUM] The aggregate is dominated by one floor per band, and it is a different floor each band. Reporting a ratio-of-means at n=6 hides the single-floor leverage. At 2.4 GHz the flat-slope floor is resplan-1374 (slope 0.0071 vs cohort ~0.020; per-floor inflation 21.6x vs 1.4–3.7x for the rest); dropping it pulls the jackknife ratio to 2.17x. At 5.0 GHz the dominant floor is resplan-12439 (slope 0.0072; per-floor inflation 35.4x vs 3.2–9.1x); dropping it pulls the ratio 7.74→5.24. This is genuine cross-floor heterogeneity — arguably the finding — but any restatement should carry median+IQR or the per-floor table, not a bare mean (isotalo2001_a604, ch. 4–5).

  • [MEDIUM] Wi-CaL comparator (choi2022_17c2): the intermediate MAE sentence swaps the room labels. Verified verbatim: meeting room k-fold {0.16, 0.18, 0.13}, LOSO 0.35; seminar room k-fold {0.32, 0.36, 0.32}, LOSO 0.41. The sealed text "(0.13–0.18 seminar / 0.32–0.36 meeting)" has the labels backwards; the final derived ratios (~×1.2 seminar / ~×2.3 meeting) use the correct pairing, so the number is right but the citing sentence is internally inconsistent with its own conclusion — the class of error the critic pass exists to catch (tebbs2006_698b, internal-consistency check). Fix: swap "seminar"↔"meeting" in the intermediate sentence.

  • [LOW] Two precision fixes on the "no NaN/Inf" claim. mean_amp_db (the gating metric) is finite in all 72 runs, but rician_k_db carries −inf in 79 rows across 31/72 runs (occluded links at high occupancy), silently dropped before averaging K. The claim is true as scoped to the reduction but should name the metric. Separately, the corpus is 11,520 links (60×144 + 12×240), not the "12,960" quoted.

  • [DESIGN] Follow-up sizing. CI-width inversion (campaign-design.md §3, n₀=z²σ²/d²) with the jackknife-implied per-floor SD (σ≈1.97·√6≈4.82) and target half-width d=1.5 needs ~40 floors — ~7× the current count. Brute-force replication alone will not tighten this; the follow-up needs an order-of-magnitude more floors (ZInD as a second dataset, already planned) plus the multi-seed variance decomposition so floor-variance and seed-variance are not conflated (thompson2012_e41f, ch. 12). The requested csi_cross_geometry summary figure was never produced (figure_count: 0); had it carried a CI band on the inflation ratio, [HIGH]/[MEDIUM] would have been visible at seal.

Figure: fig-csi-cross-geometry-lofo.png (source fig-csi-cross-geometry-lofo.py beside it) — top: 6-panel CV(N) small-multiples per floor, both bands, flagging the flat-slope floor; bottom: point + 95% bootstrap-CI of the LOFO/within inflation ratio per band against the ×1.5 decision line, showing the 2.4 GHz interval reaching below it.

Attached runs

Run Gate Purpose Replay
HA1WS5EB
03JQF1TB
RW9E4ZS5
XBMNYPS0
8AY28YSH
D40N21DJ
YEM4WX8H
AT7MH7XA
28YDCM2D
Y3A8ES82
HGJ4HDXA
4ME5BPKR
VR5M5AXB
0HGA9GP7
GQJPB6YD
B6FPRZ21
E18CRS17
JFCS31RE
X9YSXB5K
T9DDTQ85
BRQ7B8W3
KTEKS2KW
MF9BQFB1
QAHG4ZMQ
AAGVNNCZ
ZGPFNJR4
C8Z3E4CW
0BSBA2AT
M1M9JVRJ
Q6T1DQWC
RM43N9DM
CV44VX55
CH73K33Q
VMHCEDT4
ZG246RR8
88BM4RWB
BNTY03X7
PGHZ2AXC
QTPTQ0N8
W8AQXATW
C47S6F34
XQ041A8F
F9BGWB8C
E2N34WKY
T34KMJJX
S8YF1SGP
QTAYP4QC
9TKJJKKH
6S66VHNE
W3KZ7ATH
0S0JZ50Z
N8XFR4AN
N9W52WMP
TADYAQ1C
K0RS1VM3
QMNNK5B8
A9MJ8E7E
4MH47WXC
HJ8ZCXH1
TC35A8PG
DE6903RX
Q0VJPD2Z
GNGZ2X57
CSBR3H9J
H0ZWJ2R2
E9H8MKEV
9S16FSK3
RQQR15E3
Q1QKVPN2
QE99GN5B
76HGJF3Q
5W2TB3M3