Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.
Critic (severity: medium, did NOT agree with the original causal verdict — integrated)
Borderline criterion: C4. The 0.570 number is real but the synthesis originally asserted a single mechanism ("because hub anchors miss the through-traffic"). At N=1, resplan-12439 is confounded on ≥4 axes simultaneously — 6 anchors (not 4), larger area, the only furnished floor (16 occupiable, so furniture also occludes LOS), and placement strategy. The final synthesis now says associated with, names the confound, and reframes the divergence as an open hypothesis requiring a controlled sweep.
Other corrections integrated: (1) C1 marked true partly on gate_passed; the load-bearing 12439 floor actually departed 13/14 with wedge/jam QC warnings — now stated. (2) no-NaN scoped to the four coverage metrics (the walk metrics carry by-design NaNs out of scope). (3) Area-provenance mismatch named (coverage floor_area_m2 = envelope interior, e.g. test-lab 64 m² vs walk walkable_qc 15.7 m²).
Scope: N=7, single seed, single LOS range (8 m) and grid bin (0.25 m), near-free-flow (no congestion stress); LOS ≠ RF detection. The ranking and the LOS framing are sound; the single-cause attribution was the weak point and has been removed.
Addendum — post-hoc statistical re-analysis (2026-07-06, statistics skill, operator session)
Independent recomputation from the session's raw run artefacts (trajectory.parquet,
floor_geometry.json, walk-notebook.metrics.json per floor) driven through the
unmodified published reduction (monad_knowledge/notebooks/python/coverage_meets_crowds.py)
under the house statistics contract (.claude/skills/statistics/). The headline numbers
reproduce (blindspot_footfall_frac within 0.40 pp on every floor), but re-derivation
surfaces a measurement artefact that changes the campaign's central ranking fact. Verdict
direction on C4 ("associated with, not caused by") is upheld and reinforced; one ranking
claim is overturned; four precisifications follow.
-
[HIGH → overturns C3/C4 ranking] The 57% blind-spot at resplan-12439 is a stuck-agent
artefact, not a placement signal. One agent (id 1) contributes 50.2% of the floor's
7,937 footfall frames — it is wedged within 0.1 m of its spawn from t = 1.0 s and never
displaces more than 0.49 m across the full 200 s (239 m of in-place jitter path, tail
position std 0.030 m), sitting in a cell with LOS to no anchor. Because the footfall
estimator is a raw per-frame 2-D histogram with no dwell cap or per-agent weight
normalization, the mean's unbounded influence function (gentle2020_1ba7 ch. 8.7;
assumptions.md robustness row) lets that single frozen agent carry 76.3% of the blind-spot
value. Excluding it: blindspot_footfall_frac 0.5695 → 0.1349, which moves 12439 from
rank #1 to ~#3 in the corpus — tying resplan-1374 (0.138) and below resplan-147440 (0.218).
The claim "worst in the corpus (57%)" (C3/C4, supervisor verdict) does not survive
de-contamination. A systematic sweep of all 7 floors flags only 12439 (every other
floor's busiest agent ≤ 11.3% of frames with 3.8–6.8 m tail std — ordinary walking). The
floor's own walkable_qc had already emitted a wedge-risk warning (0.48 m doorway, core
splits into 2 parts) and the run departed 13/14 — the mechanism was in hand at C1 and never
propagated to the coverage metric. Correct phrasing: 12439's raw blind-spot fraction
is dominated by one wedged agent; the de-contaminated value (~13.5%) is mid-pack, and the
corpus ranking does not single 12439 out once the artefact is removed. Fix the estimator
(bounded per-agent contribution / unique-cell weighting) before any 12439 conclusion.
-
[MEDIUM] The "(mean 14%)" attributed to the dispersed floors is computed on the wrong
set. mean(blindspot) over the six 4-anchor dispersed floors = 7.06%; over all seven
(incl. the 6-anchor 12439) = 14.18%. The stated "(mean 14%)" is the all-seven mean —
which is itself inflated by the artefact above — not the dispersed-only subset the sentence
names (tebbs2006_698b ch. 9, internal-consistency check on a headline number). Report the
dispersed-only 7% or drop the parenthetical.
-
[MEDIUM] The area/anchor-budget trend is directionally real but not robustly significant
at n = 6. The synthesis asserts "rises broadly with area" with no statistic, contrary to
the house rule "estimate + CI + effect size, then p" (SKILL §1.5). Quantified over the six
dispersed floors: Spearman ρ = 0.943, exact permutation p = 0.0167 (720 relabelings;
the asymptotic p = 0.0048 is optimistic at this n — SKILL §2). But the significance hinges
on two soft choices: Pearson on the same six gives r = 0.805, p = 0.053 (n.s.);
and excluding test-lab-synth (a synthetic control with blindspot = 0 by construction, not
a ResPlan draw) the five real floors give ρ = 0.900, exact p = 0.083 (n.s.). The trend
is not strictly monotone either (147440, 312 m², 21.8% exceeds 1374, 389 m², 13.8%). Report
it as a suggestive, unreplicated, floor-level (ecological) association at n = 6, not an
established relationship (peck2008_2ba0 ch. 2: purposive sample licenses no generalization).
-
[MEDIUM] No headline number carries an interval; the only available resampling unit shows
overlapping "ranks." With seeds: [0] there is no seed-level replication; the sole
legitimate unit is the 14 agents per floor. Agent-bootstrap 95% CIs (B = 2000): 12419
1.1% [0.0, 2.8]; 16157 1.1% [0.2, 2.4]; 7421 4.6% [0.0, 9.8]; 1374 13.8% [5.5, 22.5];
147440 21.8% [14.4, 28.4]; 12439 57.0% [8.4, 82.0]. The 12439 interval spanning nearly
the whole [0,1] (boot-mean 0.465 ≪ point 0.570, because agent 1 is absent in ≈35% of
resamples) is a direct fingerprint of the artefact above; several adjacent floors (7421 vs
12419/16157) overlap and should not be read as cleanly separated. Every mean over runs must
carry a labelled interval (figures/references/statistical-figures.md); the corrected
summary figure is the required default.
-
[DESIGN] Follow-up sizing (fix the estimator first). Using the clean per-floor agent
bootstrap SDs (σ ≈ 3.5–4.3 pp) as the prior, a 15 pp dispersed-vs-hub effect needs only
~1–2 seeds/arm and a 5 pp effect ~10 seeds/arm (campaign-design.md §3,
n = 2σ²(z₀.₈+z₀.₀₂₅)²/ES²). If the estimator is not fixed, the contaminated variance
regime (σ ≈ 0.26 from 12439's bootstrap) inflates the same 15 pp target to ~47 seeds/arm — a
~40× cost for skipping step 1. Isolate the four confounds (6 vs 4 anchors, area, furniture,
placement) with a controlled anchor-layout factorial on fixed geometry (devore2012_62c8
ch. 11), ≥8 seeds/cell (assumptions.md small-n floor), and pre-declare the comparison.
Figures: fig-stuck-agent-contamination.png (per-agent frame share for 12439 + full-vs-excl
blind-spot across all 7 floors), fig-blindspot-summary-ci.png (corrected cross-floor ranking
with agent-bootstrap CIs). Sources beside them in the audit directory.
Corpus citations: gentle2020_1ba7 (unbounded influence, ch. 8.7), wasserman2004_ea08 (bootstrap
percentile CI, ch. 8; small-n honesty, ch. 8.3), diez2015_8380 (exact permutation, ch. 6),
tebbs2006_698b (headline-number consistency, ch. 9), peck2008_2ba0 (scope of inference, ch. 2),
devore2012_62c8 (factorial screening, ch. 11), campaign-design.md (paired power), and
figures/references/statistical-figures.md (bare-point-estimate rule).