Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KY06KYBD90TXXG971XF7CSFB

finished 2026-07-20 16:44:16.621432+00:00 → 2026-07-20 17:03:55.873315+00:00 · 5 runs · supervisor: react-agent

“H1 (cover≠count) holds robustly: coverage-opt {3,6,7,8,11} vs counting-opt {3,4,7,11,13}, 1−Jaccard=0.571 ≥ 0.15. Trade asymmetric (counting-first −37% footfall; coverage-first −15% counting). H2 directional only (top MI at dwell/flow mounts) — CI-separation criterion unmet, no per-candidate bootstrap CIs. H3 refuted at k≤5: footfall reaches only 0.568 at 5 APs, curve not plateaued, needs denser candidates for ≥90%. The 5-seed extension corrected an optimistic pilot (rank ρ 0.62→0.474; 1−J 0.75→0.571) and showed the count-optimal AP identities are seed-sensitive ({2,3,6,9,12}→{3,4,7,11,13}, only site 3 survived). Sim-only hypothesis generator; single floor, seed as unit (n=5); no hardware anchor. Sealed with all critic findings (1 critical + 4 medium) applied.”

Archive snapshot, as of 23 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Success criteria

CriterionResolved
STAGE 0-1: crowd non-empty, n_outside_walkable==0, Weidmann speed band, all personas present, H(C)>0 yes
STAGE 2: coverage submodularity audit + greedy≥1-1/e + counting-vs-coverage Pareto + optimal k* for footfall coverage yes
STAGE 3: per-candidate measured I(C;Φ) with bootstrap CI whose lower bound separates high- from low-flow mounts no
STAGE 5: placement sets (operational/experimental/both) + 5-panel figure + crowd replay to /map yes
FRAMING: generative-evaluation, hypothesis generator not hardware detection-rate claim, sim-only until real AX210/Pi5 anchor yes

Synthesis

The shopping mall as a placement oracle — session synthesis

Campaign: c-mall-archcad · Session: 01KY06KYBD90TXXG971XF7CSFB · Runs: 5 coupled walk→ray-traced-CSI seeds (seed0–seed4) on mall-archcad-floor-0

Executive summary

We built a simulated shopping mall from real architectural drawings — a concourse, six shops, a food court and a lounge, with the actual concrete columns of a real mall lifted from CAD — and populated it with a mix of shoppers, diners, people waiting, and people just passing through. We then asked: if you had to mount five Wi-Fi access points on the ceiling, would the spots that give the best radio coverage be the same spots that let you count how many people are in the room just from how the Wi-Fi signal wobbles? The answer, consistent across five independent crowd seeds, is no — the two objectives disagree on more than half of their chosen mounting points. The trade is asymmetric: a coverage-first placement gives up only ~15% of counting quality, whereas a counting-first placement gives up ~37% of physical footfall coverage. And no matter how five access points are arranged, they reach only about 57% of where the crowd actually walks — well short of the 90% we had hoped for at this candidate density. This is a simulation study, not a hardware measurement — it tells us where to look, not yet what a real receiver would detect. Extending the earlier 3-seed pilot to 5 seeds also corrected the pilot, which had overstated how stable its rankings were.

Abstract

We ran a generative-evaluation placement study on a real-ArchCAD-grounded 84×39 m mall floor with 14 ceiling-mounted candidate access points, coupling a multi-persona pedestrian agenda simulator to ray-traced (Sionna, Metal backend) CSI over five independent crowd seeds. We estimate, per candidate, the mutual information I(C;Φ) between in-zone occupancy count and the candidate's CSI amplitude feature, and compare the resulting counting-optimal AP set to a coverage-optimal set under a shared k=5 budget. Coverage-optimal and counting-optimal placements diverge substantially (1 − Jaccard = 0.571, above the pre-registered 0.15 effect-size gate), confirming H1; counting-informative mounts sit at food-court and concourse-adjacent positions over static shop interiors, consistent with H2 at the level of point estimates; footfall coverage rises submodularly but reaches only 0.568 at k=5, refuting the pre-registered ≥90% small-k saturation target (H3) at this candidate density. Rank stability across seeds is moderate (Spearman ρ = 0.474), so the divergence is robust but the identity of the specific count-optimal mounts is seed-sensitive. Caveat: single floor, sim-only, seed is the unit of replication (n=5).

Why this ran

This session sits in the thesis's system-design chain: before recommending where operators should mount CSI-sensing access points, we need evidence that a coverage-driven placement (the industry default) is not automatically also a good counting placement. The prior 3-seed pilot on this same mall found a promising signal (1−Jaccard 0.75, rank stability ρ=0.62) but with only 3 seeds those numbers were fragile. We predicted, going in, that adding two more seeds would hold the qualitative divergence (H1) but tighten toward a more honest, likely smaller, effect — because pilot estimates from n=3 are known to run hot. That is exactly what happened, and it is itself the finding worth reporting: extending replication corrected an overstated pilot claim rather than merely confirming it.

What we measured

The core quantity is mutual information between crowd count and CSI signal: I(C;Φ) = Σ p(c,φ)·log[p(c,φ)/(p(c)p(φ))], where C is the number of people in a zone (binned into 20-frame windows) and Φ is a candidate AP's mean CSI amplitude (dB). A candidate is "count-informative" if this number is large — its signal reliably tracks how crowded the room is. We also compute Jaccard overlap J = |A ∩ B| / |A ∪ B| between two chosen AP sets A (coverage-optimal) and B (counting-optimal); 1−J is the divergence the campaign is really testing. Restating the criteria: (1) does the simulated crowd actually vary in a way that makes counting meaningful (H(C)>0)? (2) does spreading APs for coverage show the expected diminishing returns, and how many APs would it take to cover ≥90% of where people walk? (3) which specific ceiling mounts are informative about crowd size, and does that separate cleanly from the uninformative ones? (4) does the deliverable package (chosen AP sets, figure, replay) actually get produced? These four are pre-registered as staged effect-size thresholds (e.g. 1−J ≥ 0.15), not as multiple significance tests — so the brief's BH-FDR multiplicity note is largely moot here; there are no corrected p-values to report, only gate pass/fail.

Results

Crowd validity / MI precondition (Stage 0-1). In-zone occupancy count ranged [1, 31] across seeds with standard deviation 8.45 — strong, real variation, so H(C) > 0 holds cleanly and the MI estimate downstream is well-defined. Resolved: yes.

Coverage-only Pareto and AP count (Stage 2). Footfall coverage rises submodularly with AP count — k=2: 0.276, k=3: 0.393, k=4: 0.493, k=5: 0.568 — with cleanly shrinking marginal gains (0.117 → 0.100 → 0.075 per additional AP), the textbook diminishing-returns signature. Crucially the curve has not plateaued (the 5th AP still adds +0.075), yet at the full 5-AP budget it reaches only ~0.57: 14 ceiling candidates with line-of-sight-limited reach on an 84×39 m floor do not get near 90% footfall coverage within k≤5. Resolved: no — the diminishing-returns shape holds, but the pre-registered small-k ≥90% target is refuted at this candidate density; more or denser candidate mounts (not merely a re-arrangement of these 14) would be required, and that requirement is itself a useful negative result for deployment planning.

Which mounts are count-informative (Stage 3). Per-candidate MI ranges 0.079–0.259 nats (spread 0.180 nats) across the 14 candidates. Only the extremes are clearly separated: candidate 3 (0.259) is the single stand-out count-informative mount, with candidate 9 (0.215) second; the mid-pack (candidates 11≈6≈4 at 0.170/0.167/0.155) is a near-tie within ~0.02 nats, and the least-informative are candidate 5 (0.079) and candidate 8 (0.088), both quiet shop-interior mounts. The Stage-3 success criterion explicitly required a per-candidate bootstrap CI and a lower-CI-bound separation of high- from low-flow mounts; those CIs were not carried into this reduction, so the mid-pack ordering is unresolved at n=5 and the criterion is only partially met. Resolved: partial — directionally consistent with H2 (the informative mounts are dwell/flow-adjacent, the uninformative are static shop interiors) on point estimates, but the CI-separation the criterion demands is not delivered.

Deliverables and the headline divergence (Stage 5). The coverage-optimal set {3,6,7,8,11} (footfall 0.568, informativeness 0.761) and counting-optimal set {3,4,7,11,13} (footfall 0.359, informativeness 0.896) share only sites {3,7,11}: 1 − Jaccard = 0.571, above the pre-registered minimum effect size of 0.15, so H1 holds. The sacrifice is asymmetric: switching to the counting-optimal set costs 37% of footfall coverage, while switching to the coverage-optimal set costs only 15% of counting informativeness — coverage is the more expensive objective to abandon. Coverage uniquely keeps concourse mounts {6,8}; counting swaps in food-court/lounge dwell mounts {4,13}. Seed-rank stability of the candidate ordering is moderate (Spearman ρ=0.474, down from the 3-seed pilot's ρ=0.62). That moderate stability has a direct consequence: the count-optimal set itself moved between the pilot and this run ({2,3,6,9,12} at 3 seeds → {3,4,7,11,13} at 5 seeds — only site 3 survived). The H1 divergence is robust; the exact identity of the count-optimal APs is seed-sensitive and should be read as candidates, not verdicts. All deliverables (placement sets, five-panel figure, crowd replay) were produced. Resolved: yes (divergence); AP identities provisional.

What it means

For an operator deciding where to mount APs in a real mall-like space: defaulting to a coverage-first placement is not free for crowd-counting — but it is the cheaper compromise, costing ~15% of achievable counting signal. Optimising purely for counting is the expensive direction, giving up over a third of physical coverage, and it concentrates APs on food-court and lounge dwell zones rather than shop aisles. The diminishing-returns curve says five ceiling APs are not enough to blanket this floor's footfall — a fact that should inform any recommendation for denser future AX210/Pi5 hardware deployments (IP-106/IP-112). This supports the thesis-chain claim (thesis/system-design, thesis/csi-sensing) that placement should be treated as a dual-objective problem, but the ranking itself remains a simulated hypothesis until validated on a real hardware floor.

Confidence & caveats

This is a single-floor, sim-only case study — magnitudes (0.57 coverage at k=5, 0.18-nat MI spread) are illustrative, not generalizable claims about malls in general; the unit of replication is the seed (n=5), not the floor, and seeds 0–2 are the reused pilot runs while seeds 3–4 are fresh. The moderate seed-rank stability (ρ=0.474, itself uncertain at n=5 with no interval) means the specific candidate ordering should not be treated as fixed — a sixth seed could plausibly reshuffle the mid-pack — though the top/bottom extremes (candidate 3 vs candidates 5/8) are far enough apart to be more robust. No per-candidate bootstrap CIs were carried into this synthesis, so Stage-3 separation claims are directional and that criterion is only partially met. The honest correction from the 3-seed pilot (ρ 0.62→0.47, 1−J 0.75→0.57) is a methodological point in its own right: small-n pilots in this campaign family run optimistic, and every finding here should be read as the corrected, not the original, estimate. No real AX210/Pi5 measurement anchors any of these numbers yet.

Recommended figure: fig_placement_oracle_mall-archcad-floor-0 (candidate grid · footfall · coverage · per-candidate informativeness · chosen APs). Panel 3 (per-candidate informativeness) needs a reading note in-figure: bars are point-estimate nats, not CI-bounded, so only the extremes (candidate 3 high; candidates 5/8 low) are separated — the mid-pack is within noise at n=5.

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

Criticism — session 01KY06KYBD90TXXG971XF7CSFB (c-mall-archcad placement oracle)

Findings

🔴 Critical

C1 — Exec-summary headline misstated the sacrifice (direction + magnitude). The draft claimed "coverage-first placements sacrifice up to 40% of counting quality." The numbers say the opposite: the counting-first set gives up ~37% of footfall (0.359 vs 0.568), while the coverage-first set gives up only ~15% of counting (0.761 vs 0.896). FIXED: exec summary now states the asymmetric trade correctly.

⚠️ Medium

M1 — Stage-3 only partially met; H2 downgraded. The Stage-3 criterion requires per-candidate bootstrap CIs with lower-bound separation; none were carried. FIXED: Stage-3 marked partial, H2 downgraded to directional (point MI), CI-separation criterion unmet. M2 — Top-candidate ranking not defensible without CIs. cand11≈cand6≈cand4 within ~0.02 nats. FIXED: only cand3 (and cand9) reported as clearly separated; mid-pack described as unresolved at n=5. M3 — "ceiling ~0.57" over-claimed an asymptote. Marginal gains still +0.075 at k=5. FIXED: "ceiling" dropped; curve stated as not plateaued; H3 refuted as small-k-saturation. M4 — Moderate ρ not connected to AP-set fragility. FIXED: added that count-optimal set moved {2,3,6,9,12}→{3,4,7,11,13} (only site 3 survived); divergence robust, identities seed-sensitive.

ℹ️ Low

L1 — Multiplicity plan unacknowledged. FIXED: added sentence noting staged criteria are effect-size thresholds, not multiple significance tests, so BH-FDR is moot. L2 — No CI on ρ. FIXED: "ρ=0.474, itself uncertain at n=5 with no interval." L3 — "confirmed across five" overstated. FIXED: "consistent across five seeds"; seeds 0-2 reused, 3-4 fresh, noted.

Overall verdict

seal-with-edits — 1 critical + 4 medium honesty tightenings. Underlying numbers and H1/H3/framing verdicts sound. All findings applied before sealing.

Attached runs

Run Gate Purpose Replay
CJG3SPCP replay
842JP1A3 replay
NR5PC4D8 replay
RNP14W64 replay
WT5GHYE0 replay