Written by the campaign-critic subagent against the brief's success criteria — read it as the counter-position to the synthesis above.
Self-critical notes
The headline (0.09->0.36) is genuine and positive, but bounded: (a) the MLP overfit (R2 -1.05 raw,
-0.43 tamed) — I report the LINEAR rich-feature as the result and the MLP as data-limited, NOT as "learning
works"; claiming a learned-feature win would be unsupported at 3 runs; (b) Doppler is a synthesised fading
prior matched to a nominal walking speed — it could over-state real intra-frame structure, so the lift is
an in-silico upper-ish estimate pending real CSI; (c) R2=0.36 still leaves ~64% of count variance
unexplained — CSI remains the weaker modality, consistent with the BLE-led switch; (d) single floor, 3
seeds, wide R2 error bars. The honest claim: temporal (Doppler) structure + a richer feature lift the CSI
floor materially and the bottleneck was the feature not the geometry — with real-channel validation and
more data as the load-bearing next steps.
Addendum — post-hoc statistical re-analysis (2026-07-06, statistics skill, operator session)
Independent adversarial recomputation from the raw run artefacts (links.parquet /
trajectory.parquet / ble_links.parquet and *__resolved.yaml for the three attached runs
01KW0994QJKVV5DS72XGD0515W, 01KW09BWR0JXF3D5GQ1WVCEDGT, 01KW09EFX2NVS32QACHG9GTM8J), under the
house statistics contract (.claude/skills/statistics/). Numbers reproduced by importing the
shipped monad_knowledge/notebooks/python/csi_learned_count.py functions unmodified — any gap is
evaluation design, not a different pipeline. Session verdict direction upheld and hardened: the
recomputed point estimates are honest, but this session tested a different experiment than the one
the campaign note reports as resolved, and two of its three headline R² values are split-inflated.
Eight findings confirmed, none refuted.
-
[HIGH] The vault-visible "4/4 resolved, POSITIVE" badge attaches to a question this session
never touched. The campaign brief and this session's byte-identical brief.md declare four
success_criteria: fan out exp-csi-crowd over {3 floors} × {2.4, 5.0 GHz} with
ble.enabled, run fusion_count.py (CSI-only / BLE-anchored / fused estimators), and
report the cross-floor transfer-MAE reduction as the Hybrid-Fusion chapter's first empirical.
All three attached runs' resolved.yaml instead carry floor: resplan-12439-floor-0,
carrier_freq_hz: 2.4e9, varying only scenario.seed ∈ {0,1,2} — 1 of the declared 6
(floor×band) cells, single band, and even that at a 3-seed depth the brief's sim_params pins
to seeds: [0]. The numbers come from csi_learned_count.py, a CSI-only feature probe that
never opens ble_links.parquet (grep: 2 docstring-only campaign-name mentions, zero
functional BLE references). fusion_count.py — which implements exactly the chartered design (65
ble/transfer/fused references) — exists in the repo and was never invoked here or in the
predecessor session 01KW094BB5ZKAZ13TEDRM17SVP. The synced criteria_resolved_count: 4 scores a
self-authored rubric, not the pre-registered criteria; figure_count: 0. The chartered mechanism
(does periodic BLE recalibration bound cross-environment drift?) has zero empirical content
across two consecutive sessions on the same brief (peck2008_2ba0 ch. 2 — a design substitution is
not a smaller version of the chartered study; more seeds never converge on the wrong question).
-
[HIGH] Two of three headline R² values are inflated by an i.i.d. split over an
autocorrelated series; the flagship rich-ridge number survives. Occupancy lag-1 autocorrelation
is 0.981 / 0.982 / 0.982 (n = 188/187/186), yet csi_learned_count.py (lines 147–149)
shuffles each run's ~188 macro-frames before a 70/30 split — a held-out point sits between
near-identical neighbours in time (assumptions.md: correlated adjacent measurements are a
cluster, not iid draws). Re-running the same shipped _build/_ridge/_mlp/_metrics under
leave-one-run-out (leave-one-crowd-out; the three runs share floor/layout/band, so LORO is
cross-seed generalization, wasserman2004_ea08 ch. 22.8 spirit):
| model |
reported (random split) |
LORO honest |
Δ |
| hand-CV ridge |
0.196 ± 0.115 |
0.104 ± 0.155 |
−0.092 |
| Doppler-rich ridge |
0.358 ± 0.079 |
0.389 ± 0.090 |
+0.031 |
| Doppler-rich MLP |
−0.425 ± 0.464 |
−0.012 ± 0.094 |
+0.413 |
The crowned rich-ridge R²=0.36 is robust (LORO 0.389 ≥ reported). The casualties are (a) the
hand-CV baseline, ~2× inflated (0.20 reported vs 0.10 honest) — which undercuts the "0.09→0.36
~4× lift" framing at its lower endpoint — and (b) the MLP "catastrophic overfit" narrative: LORO
−0.012 is a null, not the reported −0.43 collapse. A chronological within-run holdout collapses
further (down to R² ≈ −14.9 for 2/3 seeds), but that is partly genuine within-run distribution
shift, so LORO is the fair estimate. The synthesis's qualitative "use the linear model, the MLP
doesn't help" verdict stands — but because the evaluation over-rewards flexible in-run fitting, not
because the MLP is unusually bad.
-
[MEDIUM] Every headline is a bare point estimate; the SD was one groupby().std() away.
Reproduced spreads: hand-CV SD 0.115 (59% of the mean), rich-ridge 0.079, MLP 0.464 >
|mean| 0.425. Per-seed MLP R² = +0.053 / −0.454 / −0.874 — the sign flips across seeds, never
disclosed numerically. criticism.md flags spread only qualitatively ("wide R2 error bars",
"3 seeds, one floor", MLP "−1.05 raw → −0.43 tamed"), so uncertainty was acknowledged but not
quantified — n=3 is below the bootstrap gate (assumptions.md §2, n ≳ 10) and no CI is admissible;
the correct report is estimate + per-seed dots + explicit n=3 caveat, not a mean.
-
[MEDIUM] "Exactly as hypothesised" is unattainable at n=3. The rich-ridge > hand-CV LORO paired
ΔR² = {+0.262, +0.368, +0.226} is 3/3 directionally consistent, but the exact sign test and the
exact sign-flip permutation test (the house small-n choice) both floor at two-sided
p = 0.250 (2 of 8 patterns as extreme). A paired t reports p = 0.021, but that is exactly the
n<30 normal-theory optimism SKILL.md §2 forbids ("no method rescues" n≈5–8, a fortiori n=3).
Confirmatory language is not supported by any admissible test at this n; report as
direction consistent, significance unattainable.
-
[MEDIUM] The "fusion" benefit is stitched, not measured. No BLE feature, BLE-anchored, or fused
estimator exists anywhere in the code path that produced these numbers (ble_links.parquet present
in every run, never read). "Makes CSI a materially more useful fusion complement" and "lifts the
staleness-switch's CSI-side floor" reference components living in unrelated sessions' notebooks
(fusion_switch.py / fusion_condition_switch.py), never re-run with this feature. A facet of the
Finding-1 substitution.
-
[LOW] Success-criterion 1 ("no NaN/Inf") is violated and unchecked. rician_k_db carries
-inf in 53 / 42 / 67 rows (of 30080 / 29920 / 29760) across the three runs. Immaterial to the
R² (the feature pipeline never reads that column) — but the self-authored substitute criteria don't
check for NaN/Inf at all, so the original gate silently lapsed.
-
[LOW] The cross-session comparator shares the same leaky protocol, undisclosed. The 0.09/0.26
baseline in "0.09→0.20→0.36" comes from a different session's fusion_dense.py, whose line 80
(idx = rng.permutation(len(y)); cut = int(len(y)*0.7)) is the identical i.i.d.-over-autocorrelated
split — so the "4×" splices two numbers that likely carry the same inflation, never jointly
re-verified under one protocol and never flagged as cross-session.
-
[LOW] "~380 train samples, 30 features" overstates the per-fit data. main() loops per-run with
no pooling: each rich fit trains on one run's ~130 rows against 3·n_links = 30 features
(~4.3 samples/feature), not the ~12.7 the pooled 391-row total implies. The overfitting
conclusion holds; the stated ratio is more reassuring than the reality.
-
[DESIGN] What a properly powered follow-up needs. Separate the two conflated questions.
(a) Does the Doppler-rich feature beat hand-CV, honestly? — route through the existing
lofo_cv sim_reduction_run method (leave-one-seed-out), not a bespoke random split; the LORO
paired ΔR² pilot (mean +0.285, SD ≈ 0.075, n=3) is too thin to size on, so budget ≥8–10 seeds
for a defensible variance estimate before any confirmatory sizing (assumptions.md §2), then a
percentile bootstrap CI once n ≥ 10. (b) The chartered question — actually run fusion_count.py
on {3 floors} × {2.4, 5.0 GHz} with ble.enabled, pair by seed (same crowd → both estimators),
test the per-seed transfer-MAE difference with a label-shuffle permutation test, and apply BH-FDR
across the 6 cells before crowning any (SKILL.md §4). Effect-size floor for "bounds the drift":
recovering ≥50% of the CSI-only-transferred inflation, mirroring the real-data cross-env LOEO ~42%
figure already in the vault.
Corpus: wasserman2004_ea08 (bootstrap/CV, ch. 8/22.8), assumptions.md (autocorrelation-as-cluster,
n≳10 bootstrap gate), SKILL.md §2/§4/§5 (small-n honesty, multiplicity, reporting language),
test-chooser.md (permutation at small n, lofo_cv), campaign-design.md §1–2 (design substitution,
pairing), diez2015_8380 ch. 5 (paired power formula), peck2008_2ba0 ch. 2.
Figure: fig-csi-fusion-honest-r2.png (source fig-csi-fusion-honest-r2.py, this directory) —
per-seed dots (not aggregated bars) for reported random-split vs honest leave-one-run-out R², three
models, n=3 caveat in the caption; the chronological-holdout collapse (down to R² ≈ −14.9) is reported
in text rather than plotted so its scale doesn't swamp the panel.