Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KTPNH0PVBPG2SQ7M651DTWXZ

finished 2026-06-09 17:05:19.579187+00:00 → 2026-06-09 17:36:41.717228+00:00 · 13 runs · supervisor: react-agent

“Impairment stage is correct + strictly additive (clean oracle preserved to 5+ s.f., effective-SNR on target, no NaN/Inf), but it reveals the synthetic occupancy signature is FRAGILE: even at SNR=30 dB mean blockage atten halves (27.5->12.3 dB/person) and the Rician-K occupancy slope sign-flips (+10.12->-2.40). The brief's survival_ratio metric is non-monotone in impairment (an AWGN-shadow artifact); no defensible SNR capture budget is quotable. synthetic-csi-sim-to-real-transfer moves speculative->testable, but the result STRENGTHENS the case for real calibration rather than discharging it.”

Archive snapshot, as of 23 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Success criteria

CriterionResolved
Each (SNR,CFO) cell returns IP-085 scalar floor + csi_impaired.hdf5 + domain_metrics.impairment; no NaN/Inf; effective_snr_db within ~1 dB of target yes
Clean control (impairment_profile null) byte-identical to pre-IP-101 output — impairment stage strictly additive yes
Occupancy blockage signature degrades monotonically as SNR falls / CFO grows; min SNR @CFO=0 above survival floor 0.5 reported no
States a capture spec for exp-csi-calibration; concludes synthetic-csi-sim-to-real-transfer speculative->testable; no measured-fidelity claim no

Synthesis

CSI hardware-impairment grid — session 01KTPNH0PV (corrected post-critic)

Corpus. 12 impaired cells (SNR {5,10,20,30} dB × CFO {0,2000,10000} Hz; seed 0; n_agents=6, n_placements=48) + 1 clean control (sim://run/01KTPNN2S8YQD175H4MEBF133P), all gate-passed, exp-csi-static, dt=2026-06-09.

Pipeline correct & additive (C1 ✓, C2 ✓ at scalar level). Every impaired cell returns finite metrics (no NaN/Inf), a domain_metrics.impairment block, and csi_impaired.hdf5; effective_snr_db matches the configured target to <0.001 dB. The impairment stage is strictly additive: the clean mean_atten_db_per_person=27.50627 and k_vs_occupancy_slope=10.1247 are reproduced identically (5+ s.f.) in all 12 impaired runs, and the control carries no impairment block. Caveat: csi.hdf5 byte-identity was not diffed and runs span ≥4 simulator digests, so "byte-identical" is supported at the scalar/schema level, not binary-verified.

Headline finding — the signature is fragile, not graceful (C3 ✗). Under even the mildest tested impairment (SNR=30 dB; 0.3 dB IQ imbalance, 1° phase noise, 10-bit), the occupancy signature is materially altered: mean blockage attenuation halves (27.5→12.3 dB/person at SNR=30; →7.9 at SNR=5) and the Rician-K occupancy slope sign-flips (+10.12 clean → −2.40 at SNR=30, −1.75 at SNR=5; negative across the grid). The survival_ratio = impaired/clean atten metric the brief proposed is non-monotone in impairment (snr30/cfo2000 0.449 > snr30/cfo0 0.447 — adding CFO raised it), an AWGN-shadow-bias artifact; it must be redefined before any budget is quoted. No cell clears the 0.5 floor, so no in-grid SNR crossover exists. CFO ≤10 kHz adds negligible independent degradation (≤0.003 within an SNR tier).

Capture spec (C4 ✗ as a budget; framing ✓). The one clean, monotone scalar is residual_occupancy_snr_db ≈ SNR − 9.14 (zero-crossing ~9.1 dB) — not the "≥15 dB" the draft asserted, which contradicts its own floor data. No defensible SNR capture budget follows from this grid. synthetic-csi-sim-to-real-transfer does move speculative→testable (the commodity-NIC model now applies cleanly), but the result strengthens the case for real calibration: the in-silico clean↔impaired gap is large, so fidelity must be fit against real capture (IP-101 Q3), not asserted.

Recommended follow-on. (1) Redefine robustness on a sign-stable feature (variance monotonicity / K-slope magnitude), not the AWGN-biased attenuation ratio. (2) Extend SNR upward (40/50 dB) only after the metric is fixed. (3) Add seeds (single-seed grid). (4) Investigate why the clean K-slope is +10.12 while the documented invariant expects K to decrease with occupancy — the sign convention itself needs pinning before the flip is interpreted.

Recommended figure: csi_impairment_K_slope_vs_snr (the K-slope sign-flip) in preference to the artifact-prone survival heatmap.

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

campaign-critic — severity HIGH, disagrees with draft verdict

The draft synthesis omitted the load-bearing finding and reported an unsound metric as a capture budget. Corrected by the supervisor before sealing.

Claim audit

  • C1 (finite + effective_snr ~1 dB): supported — effective_snr_db {30.0, 19.9999…, 5.0}, all finite, <0.001 dB error. csi_impaired.hdf5 present (65–67 kB) but tensor interior not byte-inspected.
  • C2 (clean control byte-identical): borderline — control has no impairment block and clean scalars reproduce to 5+ s.f., BUT corpus spans ≥4 digests (ea74…, caac…, 7667…/1ed7…, d23d…); never byte-diffed. Schema-additive ≠ byte-identical.
  • C3 (monotone degradation + SNR floor): the draft's headline buried the real result. k_vs_occupancy_slope inverts +10.12 (clean) → −1.75…−2.40 (every impaired cell). The signature does not degrade gracefully — it sign-flips. Reporting variance_monotonicity=−1.0 as "slope preserved" conflates two different metrics.
  • C4 (capture spec): "packet SNR≥15 dB" is unsupported — floor met at NO tested SNR (max 0.447 @30 dB); residual zero-crossing ≈9.1 dB, not 15. Internal contradiction. The speculative→testable framing itself is sound.

Weak claims

  • survival_ratio = impaired/clean atten is non-monotone in impairment (snr30_cfo2000 0.449 > snr30_cfo0 0.447 — CFO raised it); AWGN inflates the shadow estimate → measurement artifact, not a budget.
  • K-slope sign-flip omitted from the draft entirely.
  • "floor crossover above 30 dB" extrapolates a non-monotone metric — unjustified.
  • residual threshold stated "~7–8 dB"; data give ~9.1 dB (residual ≈ SNR − 9.14).
  • C2 "byte-identical" never diffed across ≥4 digests; C1 hdf5 presence asserted, not byte-verified.

Honest scope

12 single-seed cells, one static scene, one fixed impairment prior. Cannot speak to seed variance, dynamic/crowd scenes, or measured hardware fidelity (prior unvalidated, IP-101 Q3). Cannot support a capture-SNR budget. Honest conclusion: the stage runs and is schema-additive; the occupancy signature collapses (K-slope sign flip) under all tested impairment; the survival metric needs redefinition before any SNR budget is quoted.

Authoritative evidence: s3://monad-knowledge/sim/exp-csi-static/dt=2026-06-09/<run_id>/artefacts/sionna-csi-runner.metrics.json for 01KTPNN2S8YQD175H4MEBF133P (clean), 01KTPNRJ93Q9BTPE1W1BNE976B (snr30_cfo0, K-slope −2.40), 01KTPNRTFH59ZKG7V5TT3FC9YV (snr5_cfo0, K-slope −1.75).


Addendum — post-hoc statistical re-analysis (2026-07-06, statistics skill, operator session)

Independent adversarial recomputation from sionna-csi-runner.metrics.json (13 runs), links.parquet (clean control 01KTPNN2S8YQD175H4MEBF133P), config/resolved.yaml, and the runner source dockerfiles/simulators/sionna-csi-runner/entrypoint.py, under the house statistics contract (.claude/skills/statistics/). Verdict direction upheld — the corrected post-critic synthesis is sober and its numbers all reproduce. Two precisifications go deeper than the campaign-critic did; the remaining caveats it already carried are confirmed, not extended.

  • [HIGH] The clean K-slope "+10.12 dB/person invariant" is a single degenerate estimate, not a trend. The k_vs_occupancy_slope regresses rician_k_db on n_occluders over 139 finite points (111/23/4/1 at occ 0/1/2/3). Per-level finite-K means are 23.2 / 27.5 / 22.9 / 142.6 dB — flat and non-monotone across the three well-supported levels, with the entire positive slope produced by the lone occ=3 sample at 142.6 dB. That value is a degenerate output of _rician_k_db (entrypoint.py:456, the root<1 → k_lin=1e6 clamp), not physics. Dropping it moves the reported clean baseline +10.12 → +2.33 dB/person (77% swing; a single high-leverage y-outlier, worst-case for OLS — watkins2016_5a65 ch. 3 Anscombe; peck2008_2ba0 ch. 13). The sign-flip under impairment survives (+2.33 → −2.40 is still a flip), but the headline magnitude "+10.12 → −2.40" borrows drama from one clamped point. Correct phrasing: the clean K-slope is unstable to a single degenerate high-occupancy estimate; the sign flip is real, its baseline magnitude is not established. Fix the estimator before re-citing: require ≥5 samples per occupancy level or switch to a leverage-resistant slope (Theil–Sen), and guard _rician_k_db against the 1e6 clamp leaking into the fit. This predates the campaign (metric inherited from c-csi-crowd-occlusion) and is load-bearing for every campaign that cites +10.12 as ground truth.

  • [MEDIUM] residual_occupancy_snr_db ≈ SNR − 9.14 is a closed-form identity, not a grid-earned budget. OLS over all 12 cells: slope 1.000000, intercept −9.142234, R² 0.99999999, max|resid| 1.4e-3 dB. entrypoint.py:1083-1092 sets residual = 10·log10(occ_signal_power / mean_noise_var) with occ_signal_power a fixed clean-scene constant and mean_noise_var derived from the target SNR — so residual = SNR + const by construction. A single cell at any SNR reproduces the identical line; the 12 impaired runs contribute nothing to it, and it carries zero evidence of generalization to another floor, placement density, or antenna count. An R² at floating-point precision is not a validated trend (statistics skill §5: goodness-of-fit does not validate a model). The corrected synthesis's number is arithmetically right and its "no defensible SNR capture budget" verdict stands; the "9.14 dB" should be reported as a single-scene closed-form constant, not "the one clean, monotone scalar."

  • [LOW, confirms critic] Clean-control "byte-identical" is not literally true — but drift is negligible. All 13 runs carry distinct simulator digests (critic's "≥4" is a floor) and 13 distinct csi.hdf5 hashes; direct CFR tensor diff is ~3.5e-8 (complex64 build non-determinism, six orders below scalar sensitivity). Scalar additivity (5–6 s.f.) holds. Synthesis and criticism already disclose "not binary-verified"; the residual gap is only that the outcome checkbox still ticks "byte-identical" — drop the word for a numeric tolerance.

  • [LOW, confirms critic] The survival_ratio non-monotonicity is systematic, not a one-off. survival_ratio = atten_impaired/atten_clean is non-monotone in CFO at all four SNR tiers (peak at CFO=2 kHz), wiggle 0.0072→0.0017 as SNR rises, max 0.4489 (<0.5 floor, never cleared). This strengthens the critic's "redefine before quoting a budget"; the wiggle is 0.4% relative under a single shared-seed noise draw, so it is consistent-with — not proof of — a real CFO×AWGN interaction.

  • [LOW, confirms critic] Zero seed replication anywhere. seed=0 and impairment_profile.seed=0 for all 12 cells → effective RNG seed 0^0=0 everywhere; no seed-to-seed variance is estimable for any headline number. The campaign correctly omits CI/bootstrap at n=1 (statistics skill §2: n≈5–8 stays caveated even after bootstrapping — no method rescues n=1). The critic's "Honest scope" already states this; only the terse outcome-frontmatter phrasing ("reveals … FRAGILE", "under all tested impairment") reads more general than the one-scene/n=1 design licenses.

  • [DESIGN] Follow-up sizing (paired power). No σ is available from n=1, so the actual gap is a variance pilot, not more grid compute. Concrete: (1) fix the estimator first (≥5 samples/level or Theil–Sen; guard the K clamp); (2) pilot ≥5 independent seeds at SNR∈{5,30} dB, CFO=0 to estimate seed-to-seed σ of the fixed metric; (3) size the full grid with the paired-difference formula n = 2σ²(z_power+z_{α/2})²/ES² (diez2015_8380 ch. 5) — at ES=2 dB/person, 80% power, α=0.05: σ=1→4 seeds, σ=2→16, σ=3→35 per cell; (4) report power as a curve over σ until the pilot lands (watkins2016_5a65 ch. 18–19).

Figures (statistical-figures contract; per-cell point estimates only, no CI licensed at n=1, caption states so): fig-kslope-sign-flip.png (clean K-slope dashed reference vs 12 impaired cells, faceted by CFO) and fig-survival-ratio.png (survival_ratio vs SNR with the undefeated 0.5 floor + SNR=30 inset zoom on the non-monotone CFO wiggle) — both under scratchpad/overhaul/c-csi-impairment-sim-to-real/ beside their .py sources.

Corpus: clean control sim://run/01KTPNN2S8YQD175H4MEBF133P; impaired anchors 01KTPNRJ93Q9BTPE1W1BNE976B (snr30/cfo0, K-slope −2.40), 01KTPNRTFH59ZKG7V5TT3FC9YV (snr5/cfo0, K-slope −1.75). Sources: wasserman2004_ea08 (ch. 8.3 small-n, §5 GOF), watkins2016_5a65 (ch. 3 leverage/Anscombe, ch. 18–19 power), peck2008_2ba0 (ch. 13 regression diagnostics), diez2015_8380 (ch. 5 power formula), devore2012_62c8 (ch. 10.1 homoscedasticity).

Attached runs

Run Gate Purpose Replay
MEBF133P clean-control
1BNE976B snr30_cfo0
78E738FE snr20_cfo0
BPDGGMN7 snr10_cfo0
TT3FC9YV snr5_cfo0
K2ZD1YYV snr20_cfo2000
QZRP28BH snr30_cfo2000
FGPZ2DZK snr10_cfo2000
3J1YX996 snr5_cfo2000
M81AA3T2 snr10_cfo10000
DQ9AA92V snr20_cfo10000
QDQ8Q3P5 snr30_cfo10000
RTMED9JY snr5_cfo10000