What we tested. All nine cells of a 3-intensity × 3-seed grid (intensity ∈ {0.5, 1.0, 2.0}) were executed on resplan-12439-floor-0 using the walk-notebook simulator (three heterogeneous groups: morning-wave seat-seekers with after_dwell: return, a late-wave reusing those freed seats, and through-traffic). Phase-A gate run 01KT952AVT0H103B5TBB2MGBZ3 addressed criteria 1–2 before the grid; all eight Phase-B cells covered criteria 3–4 across the intensity range. A superseded gate take (01KT94Y4EHTBD7YBS8WW42K054, peak_walk_speed_m_s = 9.19, park-snap artefact) was excluded from all quantitative claims.
What we found. Criterion 1 — movement surfaces: all 9 runs report trajectory.html, footfall_heatmap.png, and speed_vs_time.png in domain_metrics.figures, with non-empty trajectory.parquet (180–443 kB); the crowd is visible in every cell. Criterion 2 — kinematic plausibility: n_outside_walkable == 0 on every run; mean_walk_speed_m_s ranges 1.16–1.20 m/s (within the 0.8–1.6 m/s band); peak_walk_speed_m_s ≤ 1.20 on all runs. Criterion 3 — lifecycle at intensity ≤ 1.0: across all six cells (01KT958HG44PJ72BGSC3A2Q43B, ...6RGC, ...EF, 01KT952AVT0H103B5TBB2MGBZ3, 01KT9599WAWNCYDRTY7SFZW4AM, 01KT959JTA614C7HYK78WH1KF5), morning-wave departed == spawned, late-wave seated == spawned, through-traffic departed == spawned, and turned_away == 0 on all groups. At intensity 2.0 the late-wave incurs turned_away == 2 per seed, consistent with 36 seat-seekers exceeding 14 reachable seats — the expected capacity-gating behaviour, not a criterion failure. Criterion 4 — congestion monotonicity: pooled mean_time_to_seat_s is 3.13–4.54 s at intensity 0.5, 5.74–6.77 s at 1.0, and 9.21–10.05 s at 2.0; Spearman ρ = 0.95 across all 9 cells, well above the ρ ≥ 0.60 threshold.
What it means for the thesis chain. All four success criteria are resolved. The walk-notebook simulator now supports heterogeneous, time-structured crowd scenarios on a real multi-room floor with a working seat-reuse ledger (occupancy_rate reaching 2.43 at intensity 2.0 by design), agent speed clamped at the CFSM desired velocity (peak_walk_speed_m_s a hard 1.20 m/s on every run, mean_walk_speed_m_s barely varying 1.16–1.20 m/s) so the Weidmann band is met by construction rather than emergent — no speed–density curve exists in this corpus — and full visual replays on every cell. This validates the scenario layer — groups, waves, dwell, return — above the integrator physics already verified in EXP-S1. The capacity-gated regime at intensity 2.0 (turned_away, super-linear time-to-seat) is a seat-contention queueing signature: time-to-seat rises because 14 reachable seats are contended by progressively more seat-seekers, while walking speed stays flat across intensities — this is queueing for fixed seats, not a fundamental-diagram slowdown (an FD slowdown would require speed to fall with density, which this corpus does not show). What the corpus validates for downstream use is the scenario layer (groups, waves, dwell/return, seat-reuse ledger) and the figure/provenance artefact contract that a coupled CSI campaign such as c-csi-crowd-temporal would consume — not the coupled physics itself, which no run here tests. Every capacity and seat-reuse claim above (criteria 3 and 4) was tested against n_goals = 14, not the intended 16: two of the 16 placed seats fall outside the clearance-buffered walkable polygon and are unreachable. The 14-seat capacity is therefore a stated bound on these results — the intensity-2.0 turned_away == 2 and the time-to-seat queueing are measured against 14 reachable seats, and the intended 16-seat layout was never exercised in this corpus.