Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
Campaign session

01KWT7YKMXHJC5WPZQ9ZX4WVKR

finished 2026-07-05 22:56:26.269927+00:00 → 2026-07-05 23:12:38.354265+00:00 · 6 runs · supervisor: claude-code

“Condition-switch fusion: a regime, not the crown. The disagreement-weighted switch trends below the tau=16s clock on the bursty arm at sparse cadence (all 3 seed diffs <=0; bootstrap CI [-0.26,-0.001] persons, fragile at n=3) but LOSES to the clock at dense cadence on both arms (up to +0.34 vs best-tau) because the disagreement statistic is symmetric — it cannot tell a stale anchor from a wrong CSI estimate. Burstiness manipulation verified (0.89 vs 0.53). Design consequence: the deployable architecture is clock-GATED freshness + condition-MODULATED staleness (hybrid), and IP-106 should log both the cross-modal statistic and anchor age. Criteria 3/5 (C3 borderline-fail, C4 fail); critic severity medium, verdict upheld.”

Archive snapshot, as of 22 h ago — the run corpus is rebuilt once a day, so this page is not a live reading. The fleet panel is the live one; it refreshes every 30 s.

Success criteria

CriterionResolved
Corpus: 2 dynamics arms x >=3 seeds coupled runs, gate-pass, burstiness separation >=0.15 verified, comparable occupancy range yes
fusion_condition_switch reduction emits both parquets with all five estimators per (arm, seed, cadence) yes
Headline: bursty sparse-cadence condition <= clock; smooth within 0.1 person no
Tuning honesty: condition within 0.1 of best-tau clock at every cadence on both arms no
Framing discipline: in-silico, modelled BLE counter, mechanism evidence only, no strength upgrades yes

Synthesis

c-csi-ble-condition-switch — completion session (runs + reduction over the sealed session's casts)

Verdict

The condition switch earns a regime, not the crown — and the loss is architectural. Weighting the held BLE anchor by observed cross-modal disagreement (w = exp(-max(0, d−d0)/κ), d0/κ from the calibration window, no time constant) trends below the τ=16 s clock switch exactly where c-ble-drift-trigger predicted — bursty occupancy dynamics at sparse anchor cadence (≥25 s: 1.21 vs 1.32 persons mean; per-seed diffs −0.26 / −0.00 / −0.08, all ≤ 0; seed-bootstrap CI [−0.26, −0.001]) — but loses to the clock at dense cadence on both arms (vs the best-τ clock from the sensitivity sweep: +0.17 to +0.34 persons at cadences ≤ 6 s; it beats the best-τ clock only at ≥51 s on the bursty arm, −0.13/−0.08). Honesty on the win: n=3 seeds with per-cadence SDs (0.32–0.43) larger than the 0.11 mean gap — the direction is consistent across seeds but the magnitude is not established; this is expected-interpretation #3 (payoff confined to low-BLE-availability regimes, crossover ≈ 25–50 s cadence), not a general win.

The mechanism reading (interpretation, not measured)

The disagreement statistic is symmetric: d = |csi_adapted − ble_held| rises when the anchor goes stale and when the CSI estimate is wrong. At dense cadence the anchor is nearly always fresh, yet CSI noise keeps d above d0, so the condition switch blends in the worse modality — while the clock gets freshness for free (anchor age is directly observable in deployment; the dense-cadence clock is numerically ≡ BLE-held, confirming it holds w≈1 there). The κ sensitivity is consistent, and is itself mild evidence against aggressive switching: wider tolerance monotonically wins (κ×2 → 1.236, κ×1 → 1.276, κ×0.5 → 1.338 mean MAE). We did not emit per-frame d/d0/switch-fraction traces, so this reading is inference from the aggregate tables, flagged as such.

Design consequence for the Hybrid-Fusion chapter: the deployable architecture is a hybrid — clock-gated freshness (trust a fresh anchor unconditionally) with condition-modulated staleness (let disagreement steer only once the anchor is old enough that staleness dominates the statistic). That composition is a designed follow-on, not this session's claim. The IP-106 capture recommendation is refined accordingly: log the cross-modal statistic AND the anchor-age clock — the hybrid needs both.

Criteria

  • C1 ✓ (one clause partial, disclosed). All 6 coupled runs gate-passed with links + ble_links + carried trajectory; burstiness verified: 0.893 (bursty) vs 0.526 (smooth), separation 0.367 ≥ 0.15. The occupancy-range comparability clause is partial: means match (4.2–4.6 persons both arms) but peaks differ (smooth 7–9 vs bursty 5–6 — smooth's overlapping windows stack arrivals).
  • C2 ✓. Both parquets with all five estimators per (arm, seed, cadence).
  • C3 ✗ (borderline). Bursty half passes. Smooth sparse parity misses the ≤0.1 clause by 0.0007 (Δ=0.1007) and the miss is driven entirely by the 103 s extreme (Δ=0.27, where BLE-held itself degrades to 2.9); at 26/51 s smooth is at parity (Δ=0.03/−0.00). Booked false per the letter of the criterion, read as parity-except-the-extreme.
  • C4 ✗. The condition switch is not within 0.1 of the best-τ clock at every cadence — the dense-cadence architectural loss (verified directly from fusion_condition_switch_sensitivity.parquet in this session: max Δ 0.34 at smooth 1.3 s).
  • C5 ✓. In-silico, modelled BLE device-counter (p_detect 0.9, σ 0.8), self-authored crowd; no strength changes to ble-periodic-calibration / recalibration-trigger-from-drift.

Corpus + provenance (and one disclosed deviation)

Two-session campaign: 01KWT73W0A0B0FVC40DWPRTN1C (CI, budget-sealed) authored and schema-validated the casts; this completion session ran them. Deviation from the sealed design, deliberate and load-bearing: the sealed casts put all arrival windows in the first 300 of 3600 units, which would have confined the arms' dynamics to the calibration window and left the evaluation window flat on both arms. Windows were respread across the full duration (identical 3×5-agent group structure, dwell 900, headcount 15 ≤ 16 seats on both arms; only window width differs: smooth [0,1000]/[800,1800]/[1600,2600] vs bursty [0,30]/[1200,1230]/[2400,2430]). This is a post-hoc design choice made before any fusion numbers were computed. Empirical timing: 133 RT frames over 170 s (1.29 s/frame); cadence axis 1.3–103 s. Additional scope caveats from the critic: the 6 cells were run with --where local rebuilds and carry non-identical simulator digests (a frozen-image rerun would remove that confound); both spot-checked runs carry a walkable_qc_warning (0.5 m² of sub-40 cm channels) whose jam risk loads most on bursty egress. Oracle envelope sits at 0.91 (bursty sparse) / 1.12 (smooth sparse) — both switches leave ~0.3–0.4 persons on the table.

Artefacts

fusion_condition_switch.parquet, fusion_condition_switch_sensitivity.parquet, fusion_condition_switch.metrics.json, fig_fusion_condition_switch.png under this session's artefacts/; 6 runs attached; casts recorded in the sealed session's synthesis and this session's run configs.

Criticism adversarial review

Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.

Criticism — c-csi-ble-condition-switch / session 01KWT7YKMXHJC5WPZQ9ZX4WVKR (campaign-critic, severity: medium, verdict upheld)

Criteria audit

  • C1 supported (burstiness 0.367 separation; spot-checked runs gate-passed) — but the occupancy-range-comparability clause was unverified in the draft. → Resolved by the supervisor post-critic: means 4.2–4.6 both arms, peaks 7–9 vs 5–6 (partial; disclosed).
  • C2 supported.
  • C3 borderline, not clean fail: smooth Δ=0.1007 misses ≤0.1 by 0.0007, driven entirely by the 103 s extreme; 26/51 s at parity. → Booked false per the letter, with the borderline reading stated.
  • C4 unsupported — but the draft's evidence mixed tables: the cited Δ must come from the best-τ sensitivity sweep, not the τ=16 table. → Supervisor verified the sensitivity parquet directly (max best-τ Δ 0.34 dense; −0.13/−0.08 bursty ≥51 s).
  • C5 supported.

Findings applied

  • [HIGH] Significance: the sparse bursty win (Δ=0.11) is within seed variance at n=3 (SDs 0.32–0.43; paired t p≈0.28). → Headline softened to "trends below, direction consistent, magnitude not established"; bootstrap CI [−0.26, −0.001] reported with its fragility.
  • [MEDIUM] "Casts are the sealed session's design" was misleading — the arrival windows were materially respread (justified: sealed windows sat inside the calibration window). → Disclosed as a deliberate, pre-analysis deviation.
  • [MEDIUM] Symmetric-statistic mechanism is inferred, not measured (no d/d0/switch-fraction traces; clock≡BLE-held at dense cadence is the supporting observation). → Labelled as interpretation; κ-monotonicity noted as mild evidence against aggressive switching.
  • [MEDIUM] Provenance confounds: non-identical simulator digests across the 6 locally-rebuilt cells; walkable_qc_warning (sub-40 cm channels) loading on bursty egress. → Added to scope.
  • [LOW] C1 range clause unreported. → Verified and disclosed (partial).

Verdict direction ("a regime, not the crown"; interpretation-ladder #3) sound; arithmetic checked correct; severity medium — corrections right-size confidence, none overturn the finding.


Addendum — post-hoc statistical re-analysis (2026-07-06, statistics skill, operator session)

Independent recomputation from fusion_condition_switch{,_sensitivity}.parquet under the house statistics contract (.claude/skills/statistics/). Verdict direction upheld; five precisifications, one of which strengthens the session's weakest-looking claim.

  • [HIGH → precisified] The sparse-bursty "win" sits at the mathematical floor of n=3. Paired per-seed diffs reproduce exactly (−0.256/−0.001/−0.081, mean −0.112). The exact sign-flip permutation test gives p = 0.25 — the smallest achievable two-sided p with 3 seeds (2³ = 8 sign patterns), so "all seeds ≤ 0" cannot be significant at this n by construction. The reported bootstrap CI [−0.26, −0.001] should be withdrawn as evidence: n=3 admits only 10 distinct resamples (wasserman2004_ea08, ch. 8.3; gentle2020_1ba7, ch. 3.6.2). Correct phrasing: direction consistent across seeds; significance unattainable and magnitude unestablished at n=3.
  • [STRENGTHENS C4] The dense-cadence loss is robust to the comparator's selection bias. "Best-τ" is chosen per-cell after seeing all three clocks — a selection-biased comparator. Rechecked against the fair, fixed τ=16 clock: mean +0.211 (bursty, max +0.448) and +0.280 (smooth, max +0.394) at ≤6.4 s; best-τ numbers are nearly identical (+0.221/+0.288). The architectural loss is the statistically strongest claim in the session — nearly every seed trace is above zero at dense cadence.
  • [MEDIUM] κ-monotonicity is aggregate-only. "Wider tolerance monotonically wins" holds on grand means (1.236/1.276/1.338) but in only 27/42 = 64% of (arm, seed, cadence) cells. Scope the claim to the aggregate.
  • [MEDIUM] C3's fail is a coin flip, not a finding. The pass/fail margin (0.0007 person) is ~300–600× smaller than the per-cell seed SDs (0.16–0.41). Criteria at n=3 should be phrased as effect-size-with-interval, not point thresholds; the "parity-except-the-103s-extreme" reading is the defensible one. The paired-difference figure additionally shows the smooth 103 s cell is dominated by a single seed spiking to +0.75.
  • [DESIGN] Follow-up sizing (paired power). Observed paired-diff SD ≈ 0.13, target effect 0.11 → ~11 seeds for 80% power at α = 0.05 on the paired comparison (n = σ_D²(z₀.₈+z₀.₀₂₅)²/Δ²; diez2015_8380, ch. 5). The unpaired version of the same calculation needs ~160 seeds — pairing by seed is what makes confirmation feasible. A confirmation session should also pre-declare the sparse-cadence cut, treat cadence points within a seed as clustered (they share the run), and declare BH-FDR across the 7-cadence × 2-arm grid.

Figure: artefacts/fig_condition_vs_clock_paired.png (source fig_condition_vs_clock_paired.py beside it) — paired per-seed Δ(condition − clock τ=16) vs cadence, both arms; below zero = condition better.

Attached runs

Run Gate Purpose Replay
435TVSFW replay
E7NK23ZB replay
ZKB4RWA5 replay
8EE7652E replay
6HDBH6N3 replay
WKGSQS6C replay