Headline
In an empty room the two-chain CSI ratio held to a median 0.0056 dB (standard deviation, one-minute block medians, ten hours, twenty node-bands).
That is the drift floor. For scale: a person walking through a link perturbs CSI amplitude by whole dB. The channel's own wander is more than two orders of magnitude below the signal occupancy sensing looks for — over a night, without any calibration at all.
> > | node | received | delivery | > |---|---|---| > | monad04 | 1,020,810 | 94.52 % | > | monad06 | 1,007,968 | 93.33 % | > | monad02 | 965,106 | 89.36 % | > | monad03 | 958,142 | 88.72 % | > | monad05 | 919,081 | 85.10 % | > | **fleet** | | **90.21 %** | > > The original "24.0 Hz received vs 25 Hz sent → 96 %" is not reproducible from > any artefact: every segment is exactly 30 one-minute blocks over 1799.1–1799.5 s > wall, giving 21.3–23.6 Hz. **Use 90.2 %.** The spread across nodes (85–95 %) is > itself a result the single fleet figure hides — and it is not yet attributable, > since positions were unsurveyed.| Quantity | Value |
|---|---|
| Records captured | 4,871,107 across 5 observers |
| Duration | 12.00 h per node |
| Delivery | 90.2 % (received / 1,080,002 TX-counted frames; range 85.1–94.5 %) |
| Stalls > 1 s | 0 on every node (censoring rate 0.000000) |
| Longest gap | 0.20–0.28 s |
| Inter-arrival CV | 0.204 (< 0.5 ⇒ quantitative-grade time axis) |
| Quiet-period s.d. (h0.5–10) | median 0.0056 dB, range 0.0022–0.0122 |
| Final-hour shift | median 0.007–0.018 dB, mixed direction across nodes |
Method
Observable is the two-chain CSI ratio, never raw |H|: on the AX210/iax path
amplitude is AGC-relative, so a raw-magnitude series is largely gain-control
history. The ratio divides the common AGC term out and cancels CFO.
Every statistic runs on one-minute block medians, not raw packets. Packets 40 ms apart are serially dependent; a trend test on them reports wildly inflated significance from what are effectively a few independent looks.
Battery: X-bar/S control chart, tabular CUSUM (k = δ/2, h = 4σ), Mann-Kendall
with Theil-Sen slope, Pettitt changepoint, stall census.
Implementation: monad_knowledge/csi/drift.py (13 planted-signal tests).
Driver: notebooks/exp_lib_01_drift.py; figure notebooks/fig-exp-lib-01-drift.py.
Results
1. Capture integrity was perfect, against expectation
The capture profile's own comment predicted "~12 % of wall time lost to >1 s stalls". We lost none — zero stalls on all five observers across 12 h, with the longest gap anywhere at 0.28 s. The illuminated design is the reason: an injected 25 Hz source does not depend on ambient traffic, and the room had none (measured 0.0 Hz on ch11 before launch).
Delivery held at 90.2 % — below the ABBA pilot's 2.4 GHz figure of ~96 %, not matching it (the earlier claim of a match rested on the withdrawn 96 % number). The gap is worth a look before it is designed around: the pilot was a desk pair at short range; this is five observers at unsurveyed library distances, and the per-node spread is 85.1–94.5 %. Range is the obvious candidate and the position survey tests it directly.
2. Drift is statistically certain and physically negligible
Mann-Kendall rejects "no trend" in nearly every node-band, some at p ≈ 10⁻⁸⁹. This is power, not magnitude: with 720 blocks the test detects arbitrarily small monotone components. The Theil-Sen slopes are 0.0001–0.0033 dB/h, i.e. thousandths of a dB per hour. Report the effect size; the p-value alone would badly mislead here.
3. No shared environmental event — in amplitude
Final-hour shifts are mixed in direction in every band (4 up/1 down, 4/1, 2/3, 3/2). A building waking — HVAC, lighting, first arrivals — would move the nodes together and consistently. It did not. Whatever moved is local to each node.
4. One isolated anomaly
monad02 band 3 fell −0.138 dB in the final hour — 10.4× its own quiet s.d. Every other node-band stayed within ±4.4σ, mostly ±2σ. It is a single node, a single quarter-band, an order of magnitude beyond the rest of the fleet.
Tempting but unsupported: monad02 is the fleet's hottest node (77 °C), yet its Mimir temperature trace was flat across the whole run, so temperature does not obviously track the excursion. Cause unknown; flagged for the position survey, which may reveal something physical near that node.
5. Radios need a warm-up exclusion
The first ~30 minutes are a settling transient, clearest on monad06 band 2 (a −0.13 dB step inside the first 15 min, then stable). Including it inflates the measured drift:
| median quiet s.d. | |
|---|---|
| including first 30 min | 0.0088 dB |
| excluding first 30 min | 0.0056 dB |
Nine of twenty node-bands improve by >1.5×, monad02 band 0 by 4×. Discard a warm-up window before fitting any drift baseline — otherwise the transient is counted as channel wander.
6. Phase drifts with a fleet-wide direction; amplitude does not
Added 2026-08-13 from the mk_phase battery already in
exp_lib_01_drift_results.json, which the first pass did not report.
Amplitude and phase behave differently, and only phase looks shared:
| channel | direction census | two-sided binomial |
|---|---|---|
| amplitude | 13 up / 7 down | p = 0.26 — consistent with no shared direction |
| phase | 3 up / 17 down | p = 0.0026 |
That band-level p-value is pseudo-replicated and must not be quoted. The four tone-groups within one node share an oscillator and a receive front-end, so they are not four independent looks. At the honest unit — the node — the evidence is suggestive, not established:
| node | median phase slope (rad/h) | bands negative |
|---|---|---|
| monad02 | −0.000787 | 4/4 |
| monad03 | −0.000555 | 4/4 |
| monad04 | −0.000376 | 4/4 |
| monad05 | −0.000119 | 3/4 |
| monad06 | +0.000060 | 2/4 |
4/5 node medians negative → two-sided binomial p = 0.375. On its own that is nothing. What is striking is not the sign test but that the five node medians are perfectly rank-ordered — a monotone gradient across the fleet.
Why this matters more than its p-value: recalibration-trigger-from-drift names this exact confound as its central risk — "separating genuine environmental drift from hardware drift (CFO/STO ageing) is non-trivial — the metric could track the wrong thing." An empty room is the one arm where environmental drift is definitionally absent, so a systematic term surviving here is a candidate first measurement of the hardware-drift component. The two-chain ratio cancels CFO and the transmitter's contribution (both chains see one TX), but it does not cancel differential phase between the two receive chains — which is temperature- and ageing-sensitive.
Note the ordering does not obviously follow temperature change: monad02 is the fleet's hottest node (77 °C) and sits at the extreme, but its Mimir trace was flat all night. A thermal soak (a level effect) rather than a thermal ramp would fit that pattern — untested.
Not yet a finding — a lead with a defined test. Regress the node ordering
against position and against per-node temperature once the position survey lands
(2026-08-14). Candidate drivers: distance from monad01, front-end unit
variation, thermal soak. n = 5 will not settle it alone; a second night
(exp-lib-02) doubles the evidence for free.
What this constrains for BLE calibration
- The correction target is small: sub-0.01 dB over hours. Calibration cadence can be relaxed far below what a "drift is a big problem" prior would suggest — in a static, unoccupied room. This says nothing yet about an occupied or thermally cycling one.
- Per-band, not scalar. Bands within one node drift with different signs and magnitudes, so a single per-node offset cannot absorb it.
- Watch for isolated excursions. The monad02 event is far larger than the systematic drift, so an outlier detector matters more than a slow-drift model.
Limits
- No spatial claims — positions unsurveyed.
- Empty room only. This is the reference arm; the occupied comparison is a separate run and nothing here transfers to it.
- One band, 20 MHz. 2.4 GHz ch11 HT20, 52 tones. 5 GHz wideband behaviour is untested here.
- One night. No claim about night-to-night repeatability.
- Trend and changepoint tests both mis-describe this shape. The series are flat then move late; MK prices that as a constant hourly rate and Pettitt as a single step, and neither is what happened. The figure, not the statistic, carried the finding — see the tooling notes.
§7 Follow-up (2026-08-15): the phase gradient IS thermal
§6 closed with "Not yet a finding — a lead with a defined test", naming a thermal soak (a level effect, not a ramp) as the candidate and flagging it untested. It is now tested.
Mean SoC temperature during the run (Mimir,
avg_over_time(csid_node_temp_celsius[12h]) at the run's end) against §6's
per-node phase slopes:
| node | SoC °C during run | phase slope (rad/h) |
|---|---|---|
| monad02 | 75.67 | −0.000787 |
| monad03 | 71.02 | −0.000555 |
| monad04 | 70.10 | −0.000376 |
| monad05 | 52.88 | −0.000119 |
| monad06 | 52.84 | +0.000060 |
Spearman ρ = −1.000, exact two-sided p = 0.0167 (all 120 orderings). The "perfectly rank-ordered" gradient §6 could not explain is ordered by temperature, and the ordering is monotone with no exceptions.
fig-thermal-phase-drift.png
What makes this more than a coincidence of five numbers: temperature orders the phase channel and nothing else.
| outcome vs SoC temperature | ρ | exact p |
|---|---|---|
| phase drift slope | −1.000 | 0.0167 |
| e-process alarm count | −0.791 | 0.133 |
| amplitude noise (quiet sd) | 0.700 | 0.233 |
| delivery rate | −0.100 | 0.950 |
Delivery is range-dominated, so if "hot" were merely a stand-in for "differently placed" it should have moved delivery too. It does not (ρ = −0.10). And the split mirrors §6's own finding exactly: phase has a shared fleet-wide direction, amplitude does not.
Two further things the temperature traces settle:
- The 30-minute warm-up rule (§5) has a mechanism. The first hour shows a thermal settling transient of −1.6 to −7.2 °C per node (monad01 74.9 → 67.8, monad02 81.0 → 76.6). The empirically-derived exclusion window is the node's thermal time constant, not an arbitrary cutoff.
- Within-night variation cannot settle causation. After settling, each node holds ±1–2 °C, against a 23 °C spread between nodes. There is no usable within-node lever in this data.
Why this matters beyond the fleet. recalibration-trigger-from-drift names
its central risk as "separating genuine environmental drift from hardware drift
(CFO/STO ageing) is non-trivial — the metric could track the wrong thing." This
is a candidate separation on real hardware: the phase channel carries the
thermal/hardware component and the amplitude channel does not. If it holds
up under manipulation, part of the drift a calibration cycle currently pays for
is predictable from a sensor the node already has.
Operational consequence for the August session: log per-node temperature alongside CSI and carry it as a covariate. It is already in Mimir — the analysis just has to join it.
§8 Follow-up (2026-08-16): G4b measured for the first time — 2.4 ms against a 250 ms budget
This night's time_transfer.parquet files were never analysed. They carry the
first fleet-wide answer to the G4b clock gate, which the prereg specifies
(< 0.25 s) and the Lab Session Plan 2026-08 described as "measurable rather
than assumed" — accurately, because nobody had measured it. This section is that
measurement. It is a by-product of this run, not part of its hypothesis: the
drift headline above stands unchanged either way.
Five receivers, 12 h, one common transmitter, 4,852,104 FTM-paired rows,
every session kernel-stamped (rx_stamp_src == "kernel" throughout — no
scheduler jitter in the receive stamps, which would otherwise sit at the same
order as the quantity being measured).
| node | offset to monad06 | uncertainty (2·stderr) | MAD | n common |
|---|---|---|---|---|
| monad02 | 2.363 ms | ±0.0022 | 0.580 | 944,573 |
| monad03 | 1.780 ms | ±0.0030 | 0.770 | 939,787 |
| monad04 | 1.241 ms | ±0.0025 | 0.675 | 982,524 |
| monad05 | 1.005 ms | ±0.0054 | 1.394 | 907,310 |
| monad06 | 0 (reference) | — | — | — |
Positive = that node's clock runs behind monad06; add the offset to land on monad06's timeline. MAD is the per-packet spread of the difference series, not the precision of the median — those differ by ~200×, and conflating them is the classic error here.
Verdict: G4b PASSES. Worst pair monad02/monad06, |median| + 2·stderr =
2.365 ms against the 250 ms budget — a 106× margin. All 10 of 10 node
pairs were measured (the estimator omits an under-sampled pair rather than
reporting it as zero, so 10/10 is a real statement), all stationary under the
half-split test, 0 rows discarded.
fig-g4b-fleet-skew.png
An independent check that this is not noise. Pairwise medians must be additive — skew(A,B) + skew(B,C) = skew(A,C). Over all ten triads the residual is 305 µs median, 524 µs max, eight times smaller than the signal. That is what licenses the ordering below, not just the magnitude.
The finding that changes practice: offsets do not persist
The same measurement on the other two nights that can support it:
| night | nodes | pairs | worst bound | margin | stationary |
|---|---|---|---|---|---|
| 2026-08-11 rehearsal, office | 3 | 3 | 3.03 ms | 82× | 3/3 |
| 2026-08-12 this run | 5 | 10 | 2.37 ms | 106× | 10/10 |
| 2026-08-16 rate ladder | 5 | 10 | 2.69 ms | 93× | 9/10 |
The magnitude replicates — 2.4–3.0 ms on three separate nights. The structure does not. Against the same reference, four days apart, in the same room:
| monad02 | monad03 | monad04 | monad05 | |
|---|---|---|---|---|
| 08-12 | 2.363 | 1.780 | 1.241 | 1.005 |
| 08-16 | 2.131 | 2.187 | 2.065 | 2.691 |
This run is a clean monotone ladder with monad05 nearest monad06; four days later three nodes collapse into a 0.12 ms cluster and monad05 is farthest. Every one of those numbers is certain to ±0.007 ms, so the reordering is real, not measurement noise.
Limits
- A fixed difference between two nodes' RX pipeline floors survives the median and is indistinguishable from real skew. On six identical Pi 5 + AX210 boxes running one image it should be common-mode; on a mixed fleet it would not be, and this method could not tell.
- 305 µs is the resolution floor, from the triangle residual. Sub-millisecond structure is certified; tens-of-microseconds structure is not.
- The 2026-08-16 comparison rests on a weaker substrate. The rate ladder
restarted its injector for each of ten steps, resetting
seqto 0, so the same sequence numbers recur ten times in one receiver session. The plausibility guard caught the cross-step matches (24,554 rows discarded, ~10%) and n per pair fell from ~940k to ~27k. Its one non-stationary pair (monad03/monad05) is most likely that artefact rather than a clock step; it is not the worst pair, so the verdict is unaffected. - This says nothing about the phone. The affine-fit residual for a participant handset (EXP-P3, prereg §3.5) remains unmeasured on real hardware, and it is the part of the budget with real risk.
What it settles about monad02
The EXP-F2 CSI Drift and Periodic BLE Calibration rehearsal note worried that monad02's chrony flapping made it uncertifiable for cross-node claims. Measured on 08-11, that worry is half right: all three pairs are stationary — the flapping produced no clock step — but monad02's MAD is 1.85–2.23 ms against monad03's 0.86, and its min/max envelope is ±16–18 ms against ±8 on the clean night. Its median offset is sound; its spread is roughly double everyone else's.
Artefacts
_attachments/figures/src/fig-g4b-fleet-skew.py→_attachments/figures/fig-g4b-fleet-skew.png(§8)notebooks/python/fig_thermal_phase_drift.py→_attachments/aa-null-control/fig-thermal-phase-drift.png(§7)notebooks/exp_lib_01_drift_results.json— per node-band batterynotebooks/exp_lib_01_blocks_<node>.npz— block series (expensive to rebuild)notebooks/exp_lib_01_changepoints.jsonnotebooks/fig-exp-lib-01-drift.png
Related
- monad01 (illuminator) · monad02 · monad03 · monad04 · monad05 · monad06