Asking the fleet what it is doing…
monad-knowledge Wi-Fi sensing lab · FIIT STU
answered EXP-LIB-01

Overnight CSI drift baseline, empty library

In plain words

Left alone overnight in an empty library, how much does the radio measurement wander by itself?

the room by itself, all night one person walking through change in the reading twelve hours, empty library
The empty room's own wander over a night, drawn beside the change a walking person would cause. The design measures how far below the person the wander stays; that gap decides how often the sensor needs recalibrating. Schematic of the design, not a measurement.
Why it matters
If the empty-room wander were as large as the change a person causes, the sensor would need constant recalibration. The size of that floor decides how often to calibrate.
How it is done
One transmitter at 25 frames a second for twelve hours, receivers listening, and the two- chain CSI ratio (the one stable observable on this hardware) tracked in one-minute blocks.
Where it stands
Measured. The ratio wandered by a median 0.0056 dB over ten hours and twenty node-bands, more than two orders of magnitude below the whole-decibel change a walking person causes. Bands within one node drift with different signs, so one offset per node cannot absorb it, and the phase drift is thermal. Empty room only.

Asks, formallyOver twelve hours in an unoccupied room the channel's own wander — measured as the two-chain CSI ratio, the only stable observable on this hardware — stays far enough below the amplitude perturbation a walking person causes that uncorrected drift is not what limits overnight occupancy sensing. Falsified if the empty-room drift floor approaches the person-effect scale, which would make periodic recalibration a precondition rather than an optimisation.

FoundDrift floor measured: the two-chain CSI ratio held to a median 0.0056 dB (SD of one-minute block medians, ten hours, twenty node-bands) in an empty room, more than two orders of magnitude below the whole-dB perturbation a person causes — so a night of uncorrected drift does not threaten occupancy s…

StandingA result is in and written up

Headline

In an empty room the two-chain CSI ratio held to a median 0.0056 dB (standard deviation, one-minute block medians, ten hours, twenty node-bands).

That is the drift floor. For scale: a person walking through a link perturbs CSI amplitude by whole dB. The channel's own wander is more than two orders of magnitude below the signal occupancy sensing looks for — over a night, without any calibration at all.

> > | node | received | delivery | > |---|---|---| > | monad04 | 1,020,810 | 94.52 % | > | monad06 | 1,007,968 | 93.33 % | > | monad02 | 965,106 | 89.36 % | > | monad03 | 958,142 | 88.72 % | > | monad05 | 919,081 | 85.10 % | > | **fleet** | | **90.21 %** | > > The original "24.0 Hz received vs 25 Hz sent → 96 %" is not reproducible from > any artefact: every segment is exactly 30 one-minute blocks over 1799.1–1799.5 s > wall, giving 21.3–23.6 Hz. **Use 90.2 %.** The spread across nodes (85–95 %) is > itself a result the single fleet figure hides — and it is not yet attributable, > since positions were unsurveyed.
Quantity Value
Records captured 4,871,107 across 5 observers
Duration 12.00 h per node
Delivery 90.2 % (received / 1,080,002 TX-counted frames; range 85.1–94.5 %)
Stalls > 1 s 0 on every node (censoring rate 0.000000)
Longest gap 0.20–0.28 s
Inter-arrival CV 0.204 (< 0.5 ⇒ quantitative-grade time axis)
Quiet-period s.d. (h0.5–10) median 0.0056 dB, range 0.0022–0.0122
Final-hour shift median 0.007–0.018 dB, mixed direction across nodes

Method

Observable is the two-chain CSI ratio, never raw |H|: on the AX210/iax path amplitude is AGC-relative, so a raw-magnitude series is largely gain-control history. The ratio divides the common AGC term out and cancels CFO.

Every statistic runs on one-minute block medians, not raw packets. Packets 40 ms apart are serially dependent; a trend test on them reports wildly inflated significance from what are effectively a few independent looks.

Battery: X-bar/S control chart, tabular CUSUM (k = δ/2, h = 4σ), Mann-Kendall with Theil-Sen slope, Pettitt changepoint, stall census. Implementation: monad_knowledge/csi/drift.py (13 planted-signal tests). Driver: notebooks/exp_lib_01_drift.py; figure notebooks/fig-exp-lib-01-drift.py.

Results

1. Capture integrity was perfect, against expectation

The capture profile's own comment predicted "~12 % of wall time lost to >1 s stalls". We lost none — zero stalls on all five observers across 12 h, with the longest gap anywhere at 0.28 s. The illuminated design is the reason: an injected 25 Hz source does not depend on ambient traffic, and the room had none (measured 0.0 Hz on ch11 before launch).

Delivery held at 90.2 %below the ABBA pilot's 2.4 GHz figure of ~96 %, not matching it (the earlier claim of a match rested on the withdrawn 96 % number). The gap is worth a look before it is designed around: the pilot was a desk pair at short range; this is five observers at unsurveyed library distances, and the per-node spread is 85.1–94.5 %. Range is the obvious candidate and the position survey tests it directly.

2. Drift is statistically certain and physically negligible

Mann-Kendall rejects "no trend" in nearly every node-band, some at p ≈ 10⁻⁸⁹. This is power, not magnitude: with 720 blocks the test detects arbitrarily small monotone components. The Theil-Sen slopes are 0.0001–0.0033 dB/h, i.e. thousandths of a dB per hour. Report the effect size; the p-value alone would badly mislead here.

3. No shared environmental event — in amplitude

Final-hour shifts are mixed in direction in every band (4 up/1 down, 4/1, 2/3, 3/2). A building waking — HVAC, lighting, first arrivals — would move the nodes together and consistently. It did not. Whatever moved is local to each node.

4. One isolated anomaly

monad02 band 3 fell −0.138 dB in the final hour — 10.4× its own quiet s.d. Every other node-band stayed within ±4.4σ, mostly ±2σ. It is a single node, a single quarter-band, an order of magnitude beyond the rest of the fleet.

Tempting but unsupported: monad02 is the fleet's hottest node (77 °C), yet its Mimir temperature trace was flat across the whole run, so temperature does not obviously track the excursion. Cause unknown; flagged for the position survey, which may reveal something physical near that node.

5. Radios need a warm-up exclusion

The first ~30 minutes are a settling transient, clearest on monad06 band 2 (a −0.13 dB step inside the first 15 min, then stable). Including it inflates the measured drift:

median quiet s.d.
including first 30 min 0.0088 dB
excluding first 30 min 0.0056 dB

Nine of twenty node-bands improve by >1.5×, monad02 band 0 by 4×. Discard a warm-up window before fitting any drift baseline — otherwise the transient is counted as channel wander.

6. Phase drifts with a fleet-wide direction; amplitude does not

Added 2026-08-13 from the mk_phase battery already in exp_lib_01_drift_results.json, which the first pass did not report.

Amplitude and phase behave differently, and only phase looks shared:

channel direction census two-sided binomial
amplitude 13 up / 7 down p = 0.26 — consistent with no shared direction
phase 3 up / 17 down p = 0.0026

That band-level p-value is pseudo-replicated and must not be quoted. The four tone-groups within one node share an oscillator and a receive front-end, so they are not four independent looks. At the honest unit — the node — the evidence is suggestive, not established:

node median phase slope (rad/h) bands negative
monad02 −0.000787 4/4
monad03 −0.000555 4/4
monad04 −0.000376 4/4
monad05 −0.000119 3/4
monad06 +0.000060 2/4

4/5 node medians negative → two-sided binomial p = 0.375. On its own that is nothing. What is striking is not the sign test but that the five node medians are perfectly rank-ordered — a monotone gradient across the fleet.

Why this matters more than its p-value: recalibration-trigger-from-drift names this exact confound as its central risk — "separating genuine environmental drift from hardware drift (CFO/STO ageing) is non-trivial — the metric could track the wrong thing." An empty room is the one arm where environmental drift is definitionally absent, so a systematic term surviving here is a candidate first measurement of the hardware-drift component. The two-chain ratio cancels CFO and the transmitter's contribution (both chains see one TX), but it does not cancel differential phase between the two receive chains — which is temperature- and ageing-sensitive.

Note the ordering does not obviously follow temperature change: monad02 is the fleet's hottest node (77 °C) and sits at the extreme, but its Mimir trace was flat all night. A thermal soak (a level effect) rather than a thermal ramp would fit that pattern — untested.

Not yet a finding — a lead with a defined test. Regress the node ordering against position and against per-node temperature once the position survey lands (2026-08-14). Candidate drivers: distance from monad01, front-end unit variation, thermal soak. n = 5 will not settle it alone; a second night (exp-lib-02) doubles the evidence for free.

What this constrains for BLE calibration

  • The correction target is small: sub-0.01 dB over hours. Calibration cadence can be relaxed far below what a "drift is a big problem" prior would suggest — in a static, unoccupied room. This says nothing yet about an occupied or thermally cycling one.
  • Per-band, not scalar. Bands within one node drift with different signs and magnitudes, so a single per-node offset cannot absorb it.
  • Watch for isolated excursions. The monad02 event is far larger than the systematic drift, so an outlier detector matters more than a slow-drift model.

Limits

  1. No spatial claims — positions unsurveyed.
  2. Empty room only. This is the reference arm; the occupied comparison is a separate run and nothing here transfers to it.
  3. One band, 20 MHz. 2.4 GHz ch11 HT20, 52 tones. 5 GHz wideband behaviour is untested here.
  4. One night. No claim about night-to-night repeatability.
  5. Trend and changepoint tests both mis-describe this shape. The series are flat then move late; MK prices that as a constant hourly rate and Pettitt as a single step, and neither is what happened. The figure, not the statistic, carried the finding — see the tooling notes.

§7 Follow-up (2026-08-15): the phase gradient IS thermal

§6 closed with "Not yet a finding — a lead with a defined test", naming a thermal soak (a level effect, not a ramp) as the candidate and flagging it untested. It is now tested.

Mean SoC temperature during the run (Mimir, avg_over_time(csid_node_temp_celsius[12h]) at the run's end) against §6's per-node phase slopes:

node SoC °C during run phase slope (rad/h)
monad02 75.67 −0.000787
monad03 71.02 −0.000555
monad04 70.10 −0.000376
monad05 52.88 −0.000119
monad06 52.84 +0.000060

Spearman ρ = −1.000, exact two-sided p = 0.0167 (all 120 orderings). The "perfectly rank-ordered" gradient §6 could not explain is ordered by temperature, and the ordering is monotone with no exceptions.

fig-thermal-phase-drift.png

What makes this more than a coincidence of five numbers: temperature orders the phase channel and nothing else.

outcome vs SoC temperature ρ exact p
phase drift slope −1.000 0.0167
e-process alarm count −0.791 0.133
amplitude noise (quiet sd) 0.700 0.233
delivery rate −0.100 0.950

Delivery is range-dominated, so if "hot" were merely a stand-in for "differently placed" it should have moved delivery too. It does not (ρ = −0.10). And the split mirrors §6's own finding exactly: phase has a shared fleet-wide direction, amplitude does not.

Two further things the temperature traces settle:

  • The 30-minute warm-up rule (§5) has a mechanism. The first hour shows a thermal settling transient of −1.6 to −7.2 °C per node (monad01 74.9 → 67.8, monad02 81.0 → 76.6). The empirically-derived exclusion window is the node's thermal time constant, not an arbitrary cutoff.
  • Within-night variation cannot settle causation. After settling, each node holds ±1–2 °C, against a 23 °C spread between nodes. There is no usable within-node lever in this data.

Why this matters beyond the fleet. recalibration-trigger-from-drift names its central risk as "separating genuine environmental drift from hardware drift (CFO/STO ageing) is non-trivial — the metric could track the wrong thing." This is a candidate separation on real hardware: the phase channel carries the thermal/hardware component and the amplitude channel does not. If it holds up under manipulation, part of the drift a calibration cycle currently pays for is predictable from a sensor the node already has.

Operational consequence for the August session: log per-node temperature alongside CSI and carry it as a covariate. It is already in Mimir — the analysis just has to join it.

§8 Follow-up (2026-08-16): G4b measured for the first time — 2.4 ms against a 250 ms budget

This night's time_transfer.parquet files were never analysed. They carry the first fleet-wide answer to the G4b clock gate, which the prereg specifies (< 0.25 s) and the Lab Session Plan 2026-08 described as "measurable rather than assumed" — accurately, because nobody had measured it. This section is that measurement. It is a by-product of this run, not part of its hypothesis: the drift headline above stands unchanged either way.

Five receivers, 12 h, one common transmitter, 4,852,104 FTM-paired rows, every session kernel-stamped (rx_stamp_src == "kernel" throughout — no scheduler jitter in the receive stamps, which would otherwise sit at the same order as the quantity being measured).

node offset to monad06 uncertainty (2·stderr) MAD n common
monad02 2.363 ms ±0.0022 0.580 944,573
monad03 1.780 ms ±0.0030 0.770 939,787
monad04 1.241 ms ±0.0025 0.675 982,524
monad05 1.005 ms ±0.0054 1.394 907,310
monad06 0 (reference)

Positive = that node's clock runs behind monad06; add the offset to land on monad06's timeline. MAD is the per-packet spread of the difference series, not the precision of the median — those differ by ~200×, and conflating them is the classic error here.

Verdict: G4b PASSES. Worst pair monad02/monad06, |median| + 2·stderr = 2.365 ms against the 250 ms budget — a 106× margin. All 10 of 10 node pairs were measured (the estimator omits an under-sampled pair rather than reporting it as zero, so 10/10 is a real statement), all stationary under the half-split test, 0 rows discarded.

fig-g4b-fleet-skew.png

An independent check that this is not noise. Pairwise medians must be additive — skew(A,B) + skew(B,C) = skew(A,C). Over all ten triads the residual is 305 µs median, 524 µs max, eight times smaller than the signal. That is what licenses the ordering below, not just the magnitude.

The finding that changes practice: offsets do not persist

The same measurement on the other two nights that can support it:

night nodes pairs worst bound margin stationary
2026-08-11 rehearsal, office 3 3 3.03 ms 82× 3/3
2026-08-12 this run 5 10 2.37 ms 106× 10/10
2026-08-16 rate ladder 5 10 2.69 ms 93× 9/10

The magnitude replicates — 2.4–3.0 ms on three separate nights. The structure does not. Against the same reference, four days apart, in the same room:

monad02 monad03 monad04 monad05
08-12 2.363 1.780 1.241 1.005
08-16 2.131 2.187 2.065 2.691

This run is a clean monotone ladder with monad05 nearest monad06; four days later three nodes collapse into a 0.12 ms cluster and monad05 is farthest. Every one of those numbers is certain to ±0.007 ms, so the reordering is real, not measurement noise.

Limits

  1. A fixed difference between two nodes' RX pipeline floors survives the median and is indistinguishable from real skew. On six identical Pi 5 + AX210 boxes running one image it should be common-mode; on a mixed fleet it would not be, and this method could not tell.
  2. 305 µs is the resolution floor, from the triangle residual. Sub-millisecond structure is certified; tens-of-microseconds structure is not.
  3. The 2026-08-16 comparison rests on a weaker substrate. The rate ladder restarted its injector for each of ten steps, resetting seq to 0, so the same sequence numbers recur ten times in one receiver session. The plausibility guard caught the cross-step matches (24,554 rows discarded, ~10%) and n per pair fell from ~940k to ~27k. Its one non-stationary pair (monad03/monad05) is most likely that artefact rather than a clock step; it is not the worst pair, so the verdict is unaffected.
  4. This says nothing about the phone. The affine-fit residual for a participant handset (EXP-P3, prereg §3.5) remains unmeasured on real hardware, and it is the part of the budget with real risk.

What it settles about monad02

The EXP-F2 CSI Drift and Periodic BLE Calibration rehearsal note worried that monad02's chrony flapping made it uncertifiable for cross-node claims. Measured on 08-11, that worry is half right: all three pairs are stationary — the flapping produced no clock step — but monad02's MAD is 1.85–2.23 ms against monad03's 0.86, and its min/max envelope is ±16–18 ms against ±8 on the clean night. Its median offset is sound; its spread is roughly double everyone else's.

Artefacts

  • _attachments/figures/src/fig-g4b-fleet-skew.py_attachments/figures/fig-g4b-fleet-skew.png (§8)
  • notebooks/python/fig_thermal_phase_drift.py_attachments/aa-null-control/fig-thermal-phase-drift.png (§7)
  • notebooks/exp_lib_01_drift_results.json — per node-band battery
  • notebooks/exp_lib_01_blocks_<node>.npz — block series (expensive to rebuild)
  • notebooks/exp_lib_01_changepoints.json
  • notebooks/fig-exp-lib-01-drift.png
  • monad01 (illuminator) · monad02 · monad03 · monad04 · monad05 · monad06