Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KW0NNZRBA139168CQ762A3KE

finished 2026-06-26 00:36:08.587587+00:00 → 2026-06-26 00:37:14.688405+00:00 · 10 runs · supervisor: react-agent

“True matched-headcount run-set confirms the stratified result: with co-presence equalised by construction (time-capped FSM exits, peak 16.4 vs 15.4), the per-link CSI fingerprint is not distinguishable (KS D=0.24, p=0.095, down from D=0.40 p=4e-4). The §1 fingerprint was headcount.”

Archive snapshot, as of 5 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Synthesis

True matched-headcount run-set — the airtight spatial-fingerprint test

Why a new run-set. Session 2 matched occupancy by analysis-side stratification (conservative: it trims the high-N tail). This session matches it by construction: a new reactive arm where each FSM gets a per-persona injected time_after(T_persona) → exit, with T_persona set from the observed scripted active duration of that persona (median over the 5 scripted runs). Identical arrival windows + matched active durations ⇒ the occupancy time-profile tracks scripted, not just the marginal. 5 new seeds (01KW0N97BGNCH…, …NAV66K, …NCBW17, …NDYCCX, …NFB8B2), cast_reactive_matched.yaml.

The match holds. Peak occupancy reactive-matched 16.4 [14.7, 18.1] vs scripted 15.4 [13.9, 16.9] (CIs overlap; original reactive was 21.0 [19.9, 22.1], disjoint). Mean occupancy profile: reactive-matched 10.7 vs scripted 9.9 (original reactive 14.1). The time-capped exits pulled co-presence down onto the scripted profile — see the §06 profile overlay.

The fingerprint vanishes below significance. Per-(link,seed) CSI amplitude-variance, reactive-matched vs scripted, unstratified: KS D=0.24, p=0.095 — not significant, down from the original D=0.40, p=4.2×10⁻⁴. Two independent matched approaches agree:

  • §06a occupancy-stratified (same runs): D=0.26, p=0.056
  • §06b matched run-set (these new runs): D=0.24, p=0.095

Conclusion (settles the campaign question). On this floor + 10-receiver layout, the distinguishable CSI signature between the scripted and reactive crowds was predominantly co-present headcount, not spatial arrangement. Once headcount is equalised — by stratification and by construction — the per-link variance distributions are statistically indistinguishable (p≈0.06–0.10). A small same-direction residue persists (reactive disperses the crowd: 47% vs 59% in living-0 at matched N; reactive links sit ~2.7 dB lower at equal N) but is underpowered at n=50 link-observations/arm. For a counting model the actionable risk is occupancy-distribution shift between behaviour regimes, not a strong spatial fingerprint.

Caveats. The match is good but imperfect (reactive-matched peak 16.4 slightly > scripted 15.4 — the time-capped exit fires only on dwell-end, so some overrun). The exit-injection alters the late-evening dynamics of the reactive arm (forced departures), so this arm is "reactive-until-departure", not free-running reactive. Quasi-static RT; derived scripted/matched arms share persona DNA; single evening. The residue's significance is sample-size-limited — more seeds could resolve the ~2.7 dB / D≈0.24 effect either way.</synthesis_md> ["Build a reactive arm with headcount matched to scripted by construction (not analysis-side)", "Verify the occupancy time-profile actually matches", "Test the per-link CSI fingerprint at constructed-matched headcount", "Reconcile with the stratified result"]

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

Self-critique

  • The matched arm is "reactive-until-forced-departure". The injected exits change the late-evening behaviour (guests leave on a schedule rather than per their FSM), so the contrast is not "free reactive vs scripted" but "reactive-with-matched-departures vs scripted". This is the right control for the headcount confound but introduces a mild behavioural artefact in the tail.
  • p=0.095 is not evidence of no effect. With n=50 link-obs/arm the test is underpowered for a D≈0.24 effect; the consistent same-direction residue (higher reactive variance, −2.7 dB at matched N) may be real and would need ~3-5× the seeds to resolve. The honest statement is "not distinguishable at this power", not "identical".
  • Match imperfect at the peak (16.4 vs 15.4) — a small residual headcount gap remains, biasing slightly toward finding a (headcount) difference, yet the test is still null — strengthening the conclusion.
  • Quasi-static RT, derived arms, single evening — as before.

Attached runs

Run Gate Purpose Replay
Z07VX9XH replay
1XK20G9W replay
RYPEJJZ6 replay
XHRKNEG1 replay
VQRZB9Z4 replay
JAWVRNTB replay
MQ6AZ89Q replay
KJK4NESK replay
MVM00ZZK replay
VV6CSAZG replay