What we tested
Criterion 1 (scene staging + artefact completeness): Six floors from the ResPlan dataset — resplan-12439, resplan-7421, resplan-12419, resplan-1374, resplan-147440, resplan-16157 — were each swept across 6 occupancy levels (n_agents ∈ {0, 2, 4, 6, 8, 12}) and 2 bands (2.4 GHz, 5.0 GHz) at n_placements = 48, seed 0, body radius 0.20 m / body loss 5.0 dB (6 × 6 × 2 = 72 runs). Criterion 2 (per-floor CV monotonicity): all 12 (floor, band) pairs were tested via Spearman ρ(n_agents, CV). Criterion 3 (cross-floor drift): a leave-one-floor-out (LOFO) linear fit on CV-vs-occupancy slope was evaluated at both bands.
What we found
Criterion 1: All 72 runs returned gate_passed = true (exit code 0); every run carries links.parquet, csi.hdf5, and the IP-085 schema file with no reported NaN/Inf in the reduction. The pre-2026-06-04-pm batch (floors 12439, 7421, 12419; ~36 runs) lacked the delay_spread_ns column; the later 36 runs carry it. This is an additive schema difference with no impact on CV or LOFO statistics. Criterion 2: Spearman ρ ≥ +0.83 on every (floor, band) cell; 10 of 12 cells reach ρ ≥ +0.94. The weakest cell (resplan-12439, 5.0 GHz, ρ = +0.83) still clears the +0.5 gate. All 12/12 cells pass. Criterion 3: At 2.4 GHz, within-floor MAE = 0.019 vs. LOFO MAE = 0.073 — an inflation of ×3.81. At 5.0 GHz, within-floor MAE = 0.017 vs. LOFO MAE = 0.134 — an inflation of ×7.74. Both bands exceed the ×1.5 threshold; cross-geometry drift is demonstrated. Three robustness limits bound this conclusion: the LOFO statistic rests on N=6 floors, a single seed (seed 0), and the ~2× spread between the bands' inflation (×3.81 at 2.4 GHz vs ×7.74 at 5.0 GHz) is unexplained — the next session (ZInD floors, additional seeds) should tighten all three.
What it means for the thesis chain
All three success criteria are met. The result feeds the thesis/csi-sensing chapter argument for ble-periodic-calibration: a model trained on one ResPlan apartment geometry does not transfer to another — the in-silico geometric component of the deployment-drift penalty is ×3.8–×7.7 (in-silico, ResPlan, N=6 floors, seed 0, pooled linear CV~N fit). For reference, Wi-CaL (choi2022_17c2) reports within-session k-fold MAE (0.13–0.18 seminar / 0.32–0.36 meeting) and leave-one-session-out MAE (0.35 / 0.41) separately; the paper publishes no inflation ratio, but ratios derived by us from those MAEs span ~×1.2 (seminar room) to ~×2.3 (meeting room). The comparison is directionally consistent, but these are not the same quantity: Wi-CaL's sessions share a single geometry so its drift is temporal and environmental; ours is purely geometric with protocol and body model held fixed. Our in-silico geometric drift and Wi-CaL's real-world drift point the same direction but measure different quantities (geometric vs temporal/environmental), so neither magnitude validates the other. The hypothesis strength remains 'plausible': the real defeater is multi-site field data. A natural follow-on is adding ZInD floors as a second dataset, and a furnished-floor seated variant once IP-092 occupiables exist.