Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KW0P2QCEJ8QEBPRY2EJS49XP

finished 2026-06-26 00:43:05.998101+00:00 → 2026-06-26 01:12:51.335257+00:00 · 0 runs · supervisor: react-agent

“Residue resolved at 20 seeds/arm: the matched-headcount fingerprint is a tiny, borderline effect (KS D=0.135, p=0.0475) that only crosses significance because n grew 4x — and is partly a residual ~1-person gap. The spatial-behaviour fingerprint is negligible next to headcount.”

Archive snapshot, as of 3 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Synthesis

Residue resolution — 20 seeds/arm on the matched-headcount A/B

Why. The matched-headcount run-set (session 01KW0NNZRBA…) left the spatial-arrangement fingerprint at D=0.24, p=0.095 (5 seeds) — a same-direction residue that was underpowered. This session adds 15 more seeds per arm (scripted + reactive-matched, seeds 5–19; 30 new coupled runs, all gate-passed), pooling to 20 seeds/arm = 200 per-link variance observations each, to resolve it.

Result — the residue is real but negligible.

  • Constructed-match KS (reactive-matched vs scripted, n=200/arm): D=0.135, p=0.0475. Quadrupling the power vs the 5-seed run (D=0.24, p=0.095) shrank the effect size (0.24→0.135) while the p-value crossed 0.05 only because n grew — the textbook signature of a tiny effect made marginally "significant" by sample size. The two distributions overlap ~86%.
  • Peak occupancy: scripted 15.3 [14.7, 15.9] vs reactive-matched 16.4 [15.8, 17.1] — a residual ~1-person gap persists (the time-capped exit fires on dwell-end, so the reactive arm overruns slightly). So part of even this small D is lingering headcount, not spatial arrangement.
  • Direction is consistent across every analysis (reactive variance higher: 347 vs 237 mean per-link) — the dispersal mechanism (room_full diverts) is genuine; it just barely registers on the sensors.

Conclusion (settles the campaign). Across three controls — occupancy-stratified (D=0.26), constructed-match 5 seeds (D=0.24), constructed-match 20 seeds (D=0.135) — the matched-headcount spatial fingerprint is small, borderline, and partly confounded by a residual headcount gap. For a crowd-counting model the actionable sim-to-real risk is co-present-headcount distribution shift between behaviour regimes, not spatial arrangement. Behaviour realism changes how many people the sensor sees far more than the signature of where they stand.

Caveats. The constructed match is imperfect (16.4 vs 15.3 peak); a tighter cap (per-frame occupancy governor rather than dwell-end exit) would close the last ~1-person gap and likely push D below significance. The residual-KS p-values throughout are over-powered (links×seeds not independent) — the per-(link,seed) variance KS is the honest unit. Quasi-static RT, single floor, derived (not independently-authored) arms.</synthesis_md> ["Add seeds to resolve the borderline matched-headcount residue (D=0.24, p=0.095 at 5 seeds)", "Report the high-power effect size + significance", "Interpret honestly (effect size vs n-driven significance)"]

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

Self-critique

  • p=0.0475 is a knife-edge, not a result. It sits a hair under 0.05 with a small D=0.135 and a known residual ~1-person headcount gap. Reporting it as "significant" would be misleading; the honest framing is "tiny effect, borderline, partly confounded" — which is what the synthesis says. A pre-registered alpha + a tighter occupancy match would be needed to call it either way.
  • The match is still imperfect. 16.4 vs 15.3 peak. The dwell-end exit can't cap mid-dwell, so a per-frame occupancy governor (cap arrivals when N≥target) is the real fix — deferred; the effect is already shown to be small.
  • Variance-KS unit. 200 per-(link,seed) variances pools 10 correlated links × 20 seeds; the effective n is < 200, so even p=0.0475 is slightly optimistic — reinforcing "not a robust positive".
  • Quasi-static RT, derived arms, single evening — as throughout.

Attached runs

The attached runs are not in the current atlas snapshot — rebuild via web --build-sim-atlas.