EXP-C1 — IP-106 capture design
Constraint that drives everything: there are no students to draft. Every
participant is a favour asked of a friend or colleague, so participant-count is the
scarcest resource in the thesis and must be spent deliberately. This note answers
how many people, for how long, doing what — measured, not guessed, by subsampling
the real 990-window WiMANS designs down to smaller captures
(monad_knowledge/notebooks/python/csi_capture_design.py, results in
_attachments/wimans-zeroshot/spd/capture_design.json).
The answer, up front
| You can get | What it unlocks | What stays impossible |
|---|---|---|
| 2 people | drift / periodic-recalibration anchor passes (these need weeks, not bodies) | anything about graded occupancy |
| 5–6 people | 4 of the 5 claimed contributions: cross-room calibration, band comparison, hardware shift, conformal intervals | the saturation law |
| 10–12 people, one day | scattering-saturation-link — the derived-mechanism contribution | — |
Recruit 5–6 for the repeated sessions, and call in one favour for a single "crowd day" of 10–12. That crowd day is the only thing that makes the saturation law testable at all, and it is a single afternoon.
Three findings that change the plan
1. Error is a constant fraction of the occupancy range — so people buy range, not accuracy
Leave-one-room-out MAE grows almost exactly linearly with the maximum staged occupancy K, and the ratio is flat:
| max occupancy K | LORO MAE (5 GHz) | MAE / K |
|---|---|---|
| 2 | 0.39 | 0.195 |
| 3 | 0.57 | 0.190 |
| 4 | 0.75 | 0.188 |
| 5 | 0.94 | 0.188 |
The counter delivers ~19 % of full-scale error whatever the range. So adding people does not degrade the system — it extends what the system can say. This kills the temptation to keep K small "to stay under the 1-person bar": staying under it by shrinking the range is not an achievement. Report MAE and MAE/K.
2. Recording longer buys almost nothing — new arrangements do
At K=5, going from 10 to 55 windows per occupancy level moves LORO MAE from 0.94 to 0.94, and ρ not at all. The curve is flat by ~10 windows per level.
This is the pseudo-replication signature, and it is the same lesson that deflated the drift trigger (ρ 0.93 → 0.68 once the unit became the seed cluster rather than the window). Windows drawn from one arrangement of one group of people are not independent samples of "occupancy = 3" — they are one sample, measured repeatedly.
Design consequence: the sample size is the number of distinct arrangements,
not minutes. Between takes, re-randomise positions, postures and who-stands-where.
Ten short takes with different arrangements are worth far more than one long take.
Log an arrangement_id per take and analyse clustered on it.
3. The saturation law is untestable below ~8 people
A concave saturating fit was AIC-disfavoured against linear at every K ≤ 5 on both bands, and the fitted knee degenerated to ~1 person (i.e. it was fitting the empty-vs-occupied step, not graded curvature). The evidence gap narrows steadily with K — ΔAIC −10.3, −6.4, −4.2, −2.6 for K = 2..5 at 5 GHz, roughly +2.5 per added person — which extrapolates to curvature becoming detectable around K ≈ 8.
This does not contradict the existing 7-environment concavity result; it bounds where curvature is testable. Counts 0–5 sit in the law's linear regime. Any capture capped at 5 people cannot test scattering-saturation-link, which is a novelty-80 claimed contribution.
What each hypothesis actually needs
The scarce resources are different per hypothesis, and confusing them wastes favours:
| Hypothesis | Needs | Not |
|---|---|---|
| scattering-saturation-link | people (K ≳ 8, one day) | rooms, time |
| hierarchical-calibration-shrinkage | rooms (≥3, ideally ≥4) | more people |
| ble-periodic-calibration, recalibration-trigger-from-drift | calendar time (weeks, repeat visits) | more people |
| hardware-shift-dominates-layout-shift | two NICs, matched count range | more people |
| band-5ghz-occupancy-discriminability | concurrent dual-band on one node | more people |
| count-signal-in-temporal-doppler | stillness (seated block, ~51 s windows) | more people |
| spectral-count-eigenvalue-threshold | long continuous windows (γ sweep) + complex I/Q | more people |
Only one line in that table says "people". That is the whole planning insight: the crowd day is a one-off, and everything else is rooms, weeks, and hardware.
The session plan
S1 — Core sessions (5–6 people, ×3 rooms, repeatable). Levels 0..5, ≥10 distinct arrangements per level, ≥30 s continuous per arrangement, concurrent 2.4 + 5 GHz, complex I/Q persisted. ≈30 min of recording per room, ≈75 min wall-clock with resets. Serves calibration, band, conformal, hardware-shift.
S2 — Empty-room blocks (0 people). Per room × band × hardware config, sized to be split in half. The spike-count estimator calibrates its thresholds on held-out count = 0 windows — this is measurement, not warm-up.
S3 — Still block (same 5–6 people, seated). The only source in existence for stationary crowds at graded occupancy; WiMANS collapses to 12 windows at count ≥ 2. Log per-person activity so a strict "all still" slice is recoverable.
S4 — Crowd day (10–12 people, once). Levels 0..K in steps of 2 to save time — curvature needs range, not resolution. Unlocks the saturation law.
S5 — Drift passes (2 people, every 1–2 weeks). Short anchored re-measurements. Needs persistence, not people.
What to skip
- Riemannian recentring — no dedicated capture. Its LOEO test ran on 2026-08-03 and lost to a single scalar; it now rides along with S1's complex I/Q as one confirmation leg. See riemannian-covariance-recalibration.
- Receiver-density sweeps — receiver-density-drives-recoverability is deferred as a premise (novelty 20; adeel2019 covers it). Do not spend sessions varying sniffer count.
- Sim-to-real pretraining arms — synthetic-csi-sim-to-real-transfer is deferred as a reported negative. Do not add a synthetic-pretraining arm.
- State-space filter arms — state-space-fusion-optimality is absorbed; the filter can be fitted post-hoc on S1/S5 data, so it needs no capture design of its own.
- Long single takes. Finding 2: they are nearly worthless. Split the same minutes across more arrangements.
Software stack — readiness before arrival
Verified 2026-08-03 (read-only checks):
- ✅ Analysis path live —
csi_sourcesreports 5 source schemes and 13 registered views; fleet captures from monad01 / monad02 / monad05 are on S3 and readable. - ✅ Capture daemon proven —
csidsessions exist includingexp-band-24anddrift-overnight-illum(2026-07-27). - ✅ Device/serial surface —
device {profiles, ports, capture, exec}present. - ⚠️ No labelled-occupancy session exists. Every capture on S3 is smoke, console,
bench or drift. The count-labelled path — app ground truth joined to a
csiqsession — has never been exercised end to end.
The one thing to rehearse before people are in the room: a full dry run of
capture → ground-truth label → joined artefact → reduction, with one person walking,
for five minutes. Every failure mode found during that rehearsal is a favour not
wasted. Additional pre-arrival items: confirm the app emits an arrangement_id
(Finding 2 depends on it), and confirm dual-band concurrent capture actually works on
one node rather than being a sequential sweep (A2 in the IP-106 addendum).
Provenance: design study run 2026-08-03 over the WiMANS 5 GHz and 2.4 GHz substrates built for the SPD/RMT falsifier session. Subsampling over K ∈ {2..5} × m ∈ {5..55}, 40 repetitions per cell, leave-one-room-out throughout.
Toolchain status (2026-08-04) — what the instrument can and cannot do
The capture design above assumed a working instrument. It was verified on 2026-08-04 by building the app and running it against a live backend rather than by reading it; see 2026-08-04 - The day the app became an instrument - ten defects, one coupling, and an honest readiness answer. Ten defects were found and fixed, two remain open, and the split matters for scheduling.
The design study is now enforced by the software, not by discipline
Three findings from this note are no longer things an operator has to remember:
- A take is a step. Each
waitstep carriesoccupancy_count,arrangement_idandposture, and its start/end are written as markers onto the recording timeline. Slicing a session by condition no longer depends on a clipboard. - Arrangements are counted, not assumed.
arrangement_idis authored per take, so "distinct arrangements per occupancy" — the quantity this study identifies as the real sample size — is a query rather than an estimate. - The rehearsal is a quest.
EXP-C1 · Dry run (no radio)names no AP, stays witness-only, and therefore runs on a desk or an emulator. Walking it is the cheapest way to find wording and timing problems before anyone is standing in the room.
Quests now covering the protocol
| Session (this note) | Quest | Notes |
|---|---|---|
| S1 core | EXP-C1 · Core session (occupancy 0–5) |
17 steps; 12 takes, two arrangements per level |
| S2 empty-room | EXP-C1 · Empty-room baseline |
8 takes, sized to split for threshold calibration |
| S3 still block | EXP-C1 · Still-crowd block (seated) |
51 s takes per the Doppler gate |
| S5 drift passes | EXP-C1 · BLE recalibration pass (weekly) |
5 spread anchors; built for repetition, not for one run |
| — | EXP-C1 · Dual-band A/B |
matched occupancy on both bands, back to back |
| — | EXP-C1 · Dry run (no radio) |
the rehearsal |
| — | Room geometry scan (LiDAR), UWB position survey |
capability-gated; withheld from handsets that lack the sensor |
What still blocks a crowd day
Markers do not reach analysis yet. Closed 2026-08-04. markers.tsv uploads alongside the
other streams and was verified in the bucket, with occupancy_count / arrangement_id / posture
carried per take. Retry-after-failure is verified too: sessions retained through a backend outage
uploaded cleanly afterwards.
The instrument-error screen is a dead end. Closed 2026-08-04. An instrument abort is now a
dismissible banner with Retry instrument; the step list stays on screen and the quest remains
walkable without the radio.
New, and it affects the deployment: a broken env fallback in services.yaml made every session
upload return 500. Fixed locally and verified, but the deployed API still carries the bug — it
needs an image build and redeploy before any phone can upload to production.
Gates 3–5 are unproven. Socket pinning, clock discipline over the data socket, and emission have never run against a real AP; an emulator structurally cannot reach them. So has the upload path. Every radio claim in this note rests on legs that are wired but unexercised.
Revised order of operations
The sequencing this note already recommended survives, with the software gates made explicit:
Ship the marker export and the error-recovery fix.Done 2026-08-04. Remaining: redeploy the API with theservices.yamlupload fix, and seed the quests on production.- Five-minute rehearsal — one person walking,
Dry runthen a real quest, all the way through capture → label → joined artefact → reduction. - Two-person hardware bring-up against a real AP: proves gates 3–5 and the upload. Needs hardware, not people, so it can happen before any recruiting.
- S2 empty-room and S5 drift passes — both are two-person jobs and both are prerequisites for the estimators the rest of the analysis uses.
- S1 core sessions with the 5–6 core team.
- The crowd day (10–12), last, once everything above has run clean at least once.
Steps 1–4 need no participants at all. That is the useful thing about this ordering: nothing on the critical path before step 5 costs anybody a favour.