c-ble-graded-count-powered — session 01KWWWEXATGECA6XSP7J8TJQP6 (2026-07-07)
Verdict: the pilot's graded-slope claim does NOT survive powering. REFUTED at n=8. The paired (topology − dispersed) occupancy-response slope difference is −0.126 dB/person (95% seed-block bootstrap CI [−0.412, +0.113], exact sign-flip p=0.461, paired d=−0.32) — the CI includes 0, so the primary criterion fails on both clauses (excludes-0 and magnitude-lower-bound > 0.30). The pilot's −0.498 dB/person (c-ble-anchor-placement, n=3) lies below the powered CI's lower bound (−0.412) — it is CI-excluded, not merely unconfirmed: the design powered to δ=0.498 rules the pilot point out.
What we tested
Exact re-run of the pilot design, only the seed count grows: 2 anchor layouts (csi-link-resplan-12439-multiroom dispersed vs …-topo-anchors hub+bedroom-1), paired by seed, N ∈ {0,2,4,6,8,12}, noise_db 2.0, 2.4 GHz, n_placements 24 — 96 runs (2×6×8), all gate-passed. Estimand = per-seed OLS slope of same-seed-baseline-normalised mean ΔRSSI on occupied N, dB/person; unit of replication = seed (not the 11,520 pooled frame×link events). Reduction csi_ble_graded_power.py (seed-clustered; 10k seed bootstrap; exact 256-flip sign test; 20-cell within-occupied ΔN grid with BH-FDR).
What we found
- Primary refuted (C1 fail). Paired slope diff −0.126 dB/person, CI [−0.412, +0.113]; median −0.084, MAD 0.235 (mean/median agree — not outlier-driven). Per-seed diffs straddle zero: −0.92, −0.39, −0.18, −0.15, −0.02, +0.08, +0.24, +0.33. p=0.461 vs the 0.0078 attainable floor.
- The pilot's mechanism story is itself wrong at power. The pilot claimed dispersed saturates flat (+0.08 dB/person) while topology alone is monotone (−0.48). At n=8 both placements carry a negative occupancy response: dispersed mean −0.227 (7/8 seeds' fitted slope negative, sign p=0.070, one strongly positive seed drives the spread), topology mean −0.353 (8/8 fitted slopes negative). Two honesty caveats on the topology figure: its sign-test p=0.0078 is the exact n=8 two-sided attainable floor (2·0.5⁸), the minimum the design can emit, not a margin; and it is an uncorrected, post-hoc, single-arm directional test outside the two pre-declared endpoints — read as descriptive, not as a third confirmed result. So topology gives a more consistently negative fitted slope, dispersed does not saturate, and the between-arm gap is within seed noise. The pilot's n=3 caught a favourable dispersed subset.
- Graded grid fails cleanly (C2 fail). 0/20 within-occupied ΔN cells survive BH-FDR(0.05); the "Graded" ΔN=2 topology cells resolve 0/3. Fine graded counting is not supported at power under either placement.
- Reporting (C3 met). Every headline carries estimate + CI + effect size at n=8; median/MAD reported alongside mean/SD; N=12 flagged (
involves_n12) in the per-cell grid — the end-difference sensitivity (−0.131, CI [−0.442, +0.127]) matches the OLS slope, confirming the null is not an artefact of the slope definition.
What it means for the thesis chain
This is the powered adjudication the 2026-07-06 statistical overhaul ordered, and it lands as a negative that corrects the record: topology-informed anchor placement does not buy a resolvably steeper graded count than dispersed placement — the pilot's headline was an n=3 seed-variance artefact. What does survive is narrower and honest: a hub+large-room anchor layout produces a consistently negative occupancy→ΔRSSI fitted slope (8/8 seeds, a descriptive response-consistency observation, not a confirmed endpoint), where dispersed placement is directionally similar but seed-noisy. For ble-periodic-calibration (held plausible, unchanged) the deliverable flips from "concentrate anchors for graded count" to "concentrate anchors for response consistency — graded headcount from a single narrowband Tx is not in evidence at power." The IP-106 hardware protocol should not budget for fine graded BLE counting off anchor placement alone. No hypothesis strength changes.
Honest scope
One floor (resplan-12439), one occlusion geometry, single narrowband Tx, uniform-drywall, self-authored Gaussian BLE noise — in-silico, so this refutes the pilot's in-silico claim, not a field result. n=8 was powered (~90%) against the pilot's δ=0.498/σ=0.379; the observed effect (~0.13) is genuinely small, not merely underpowered — but a δ≈0.13 effect would need n≈75 seeds to resolve, and is below the 0.30 dB/person threshold the brief declared as practically meaningful regardless. This is a powered negative for the pilot-sized effect, not proof the true difference is zero — the CI still admits differences out to −0.412 dB/person. Sentinel-absorption check (critic-requested, now run): all 96 cells carry exactly 120 events, no NaN cell means, and the N=12 cells span −0.99 to −9.26 dB with no constant-value signature — the null is not a degenerate-data artefact. Reduction + per-cell grid + 3 figures (seed_trace, paired_diff, perm_null) under this session's artefacts/ prefix.