sensorlint 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. sensorlint-0.1.0/.gitignore +32 -0
  2. sensorlint-0.1.0/CHANGELOG.md +79 -0
  3. sensorlint-0.1.0/LICENSE +21 -0
  4. sensorlint-0.1.0/PKG-INFO +508 -0
  5. sensorlint-0.1.0/README.md +451 -0
  6. sensorlint-0.1.0/examples/silent_empty_stage.py +173 -0
  7. sensorlint-0.1.0/examples/truncated_gzip.py +141 -0
  8. sensorlint-0.1.0/pyproject.toml +152 -0
  9. sensorlint-0.1.0/src/sensorlint/__init__.py +162 -0
  10. sensorlint-0.1.0/src/sensorlint/__main__.py +377 -0
  11. sensorlint-0.1.0/src/sensorlint/_util.py +192 -0
  12. sensorlint-0.1.0/src/sensorlint/checks/__init__.py +106 -0
  13. sensorlint-0.1.0/src/sensorlint/checks/clipping.py +278 -0
  14. sensorlint-0.1.0/src/sensorlint/checks/coverage.py +276 -0
  15. sensorlint-0.1.0/src/sensorlint/checks/decode.py +410 -0
  16. sensorlint-0.1.0/src/sensorlint/checks/length.py +164 -0
  17. sensorlint-0.1.0/src/sensorlint/checks/nonzero.py +399 -0
  18. sensorlint-0.1.0/src/sensorlint/checks/sample_rate.py +242 -0
  19. sensorlint-0.1.0/src/sensorlint/checks/scale.py +364 -0
  20. sensorlint-0.1.0/src/sensorlint/checks/staleness.py +267 -0
  21. sensorlint-0.1.0/src/sensorlint/checks/transposition.py +493 -0
  22. sensorlint-0.1.0/src/sensorlint/errors.py +108 -0
  23. sensorlint-0.1.0/src/sensorlint/py.typed +0 -0
  24. sensorlint-0.1.0/src/sensorlint/report.py +261 -0
  25. sensorlint-0.1.0/src/sensorlint/result.py +133 -0
  26. sensorlint-0.1.0/tests/conftest.py +50 -0
  27. sensorlint-0.1.0/tests/test_assert_contract.py +213 -0
  28. sensorlint-0.1.0/tests/test_cli.py +257 -0
  29. sensorlint-0.1.0/tests/test_clipping.py +227 -0
  30. sensorlint-0.1.0/tests/test_coverage.py +246 -0
  31. sensorlint-0.1.0/tests/test_decode.py +302 -0
  32. sensorlint-0.1.0/tests/test_length.py +140 -0
  33. sensorlint-0.1.0/tests/test_nonzero.py +298 -0
  34. sensorlint-0.1.0/tests/test_report.py +175 -0
  35. sensorlint-0.1.0/tests/test_result_and_errors.py +235 -0
  36. sensorlint-0.1.0/tests/test_sample_rate.py +229 -0
  37. sensorlint-0.1.0/tests/test_scale.py +181 -0
  38. sensorlint-0.1.0/tests/test_staleness.py +213 -0
  39. sensorlint-0.1.0/tests/test_transposition.py +387 -0
  40. sensorlint-0.1.0/tests/test_util.py +209 -0
@@ -0,0 +1,32 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ *$py.class
4
+ *.so
5
+
6
+ build/
7
+ dist/
8
+ sdist/
9
+ wheels/
10
+ *.egg-info/
11
+ *.egg
12
+ .eggs/
13
+
14
+ .venv/
15
+ venv/
16
+ env/
17
+ .python-version
18
+
19
+ .pytest_cache/
20
+ .mypy_cache/
21
+ .ruff_cache/
22
+ .coverage
23
+ .coverage.*
24
+ coverage.xml
25
+ htmlcov/
26
+ .tox/
27
+ .nox/
28
+
29
+ .DS_Store
30
+ .idea/
31
+ .vscode/
32
+ *.swp
@@ -0,0 +1,79 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here.
4
+
5
+ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
+ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
+
8
+ ## [Unreleased]
9
+
10
+ Nothing yet.
11
+
12
+ ## [0.1.0] - 2026-08-28
13
+
14
+ First release. Nine checks, each with a raising `assert_*` form and a
15
+ structured-result `check_*` form.
16
+
17
+ ### Added
18
+
19
+ - **`assert_decoded_fully` / `check_decoded_fully` / `safe_gunzip`** — detects
20
+ truncated or partial decompression by reading `decompressobj.eof`,
21
+ `unused_data` and the gzip CRC32/ISIZE trailer rather than inspecting the
22
+ decoded bytes. Handles gzip, zlib, raw deflate, concatenated members and
23
+ trailing NUL padding. `safe_gunzip` refuses to return partial data.
24
+ `probe_compressed` and `sniff_container` expose the underlying detail.
25
+ - **`assert_expected_length` / `check_expected_length`** — short reads and
26
+ buffer truncation, with absolute and percentage tolerances and an `axis`
27
+ option for 2-D blocks. `measure_length` handles arrays, sequences and bytes.
28
+ - **`assert_not_clipped` / `check_not_clipped`** — converter saturation, with a
29
+ separate leading-window test (`leading_samples`) for settling transients that
30
+ a whole-array fraction threshold cannot see. `autodetect_rails` infers rails
31
+ from repeated exact extreme values when `adc_max` is not supplied.
32
+ - **`assert_scale_declared` / `check_scale_declared`** — refuses raw ADC counts
33
+ with no declared scale factor and physical unit. `looks_like_raw_counts`
34
+ scores integer dtype, all-integral floats, proximity to a common converter
35
+ full scale, magnitude, and a unit quantization step.
36
+ - **`assert_not_stale` / `check_not_stale`** — frozen runs and
37
+ linearly-interpolated stretches, separated by the first difference (zero
38
+ slope is frozen, constant non-zero slope is interpolated) and detected via a
39
+ near-zero second difference. `longest_frozen_run` and
40
+ `longest_interpolated_run` are exposed directly.
41
+ - **`assert_sample_rate` / `check_sample_rate`** — claimed rate versus the
42
+ median observed interval, plus jitter, gap detection and non-monotonic
43
+ timestamp detection. Accepts float seconds, integer counts in a given unit,
44
+ `numpy.datetime64` and `datetime.datetime`. `estimate_sample_rate` exposes the
45
+ estimator.
46
+ - **`detect_channel_transposition` / `check_channel_transposition` /
47
+ `assert_channels_not_transposed`** — inter-channel energy-ratio discriminant
48
+ scored as two hypotheses (normal versus swapped) with a margin, refusing to
49
+ judge pairs whose expected separation is too small.
50
+ `estimate_expected_ratio_db` learns the baseline from reference pairs using a
51
+ median.
52
+ - **`transposition_run_lengths`** — run-length statistics over a chronological
53
+ sequence of per-file verdicts, yielding `estimated_false_positive_rate` with
54
+ no ground truth, plus `n_isolated_expected_if_random` as an i.i.d. baseline
55
+ and a `clustered` verdict.
56
+ - **`assert_coverage` / `check_coverage`** — interior density over a date range,
57
+ so endpoint-only coverage with a hole in the middle fails. Reports the longest
58
+ run of empty bins. `coverage_histogram` exposes the binning.
59
+ - **`assert_nonzero_kept` / `check_nonzero_kept`** — the meta-check: a stage that
60
+ kept zero records fails rather than reporting success. Also available as the
61
+ `@nonzero_kept` decorator and the `pipeline_stage` context manager, with an
62
+ optional `min_keep_fraction` for the nearly-empty case.
63
+ - **`CheckResult`** and **`Severity`** — the structured result every check
64
+ returns, carrying check name, pass/fail, severity, message and a details dict.
65
+ - **`SensorLintError`** hierarchy — one subclass per check family, all
66
+ subclassing `AssertionError`, each carrying the `CheckResult` that produced it.
67
+ - **`sensorlint.report`** — `summarize`, `format_table`, `format_report` and
68
+ `worst_offenders` for aggregating results across a fleet.
69
+ - **CLI** — `python -m sensorlint check <file.npy|file.csv>` with `--fs`,
70
+ `--adc-max`, `--expected-length`, `--scale`, `--unit`, `--leading-samples`,
71
+ `--json` and more. Exits non-zero on failure. Skipped checks are reported
72
+ explicitly rather than counting silently as passes.
73
+ - **`py.typed`** marker and full type annotations; mypy strict clean.
74
+ - **Examples** — `examples/truncated_gzip.py` and
75
+ `examples/silent_empty_stage.py`, both runnable, both demonstrating the real
76
+ failure before showing the assertion that stops it.
77
+
78
+ [Unreleased]: https://github.com/johnhagedorncs/sensorlint/compare/v0.1.0...HEAD
79
+ [0.1.0]: https://github.com/johnhagedorncs/sensorlint/releases/tag/v0.1.0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 John Hagedorn
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,508 @@
1
+ Metadata-Version: 2.5
2
+ Name: sensorlint
3
+ Version: 0.1.0
4
+ Summary: Assertions for sensor-data pipelines. Every check aborts instead of warning.
5
+ Project-URL: Homepage, https://github.com/johnhagedorncs/sensorlint
6
+ Project-URL: Source, https://github.com/johnhagedorncs/sensorlint
7
+ Project-URL: Issues, https://github.com/johnhagedorncs/sensorlint/issues
8
+ Project-URL: Changelog, https://github.com/johnhagedorncs/sensorlint/blob/main/CHANGELOG.md
9
+ Author: John Hagedorn
10
+ Maintainer: John Hagedorn
11
+ License: MIT License
12
+
13
+ Copyright (c) 2026 John Hagedorn
14
+
15
+ Permission is hereby granted, free of charge, to any person obtaining a copy
16
+ of this software and associated documentation files (the "Software"), to deal
17
+ in the Software without restriction, including without limitation the rights
18
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
19
+ copies of the Software, and to permit persons to whom the Software is
20
+ furnished to do so, subject to the following conditions:
21
+
22
+ The above copyright notice and this permission notice shall be included in all
23
+ copies or substantial portions of the Software.
24
+
25
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
26
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
27
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
28
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
29
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
30
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
31
+ SOFTWARE.
32
+ License-File: LICENSE
33
+ Keywords: adc,assertions,data-pipeline,data-quality,sensor,signal-processing,telemetry,time-series,validation,vibration
34
+ Classifier: Development Status :: 4 - Beta
35
+ Classifier: Intended Audience :: Developers
36
+ Classifier: Intended Audience :: Science/Research
37
+ Classifier: License :: OSI Approved :: MIT License
38
+ Classifier: Operating System :: OS Independent
39
+ Classifier: Programming Language :: Python :: 3
40
+ Classifier: Programming Language :: Python :: 3 :: Only
41
+ Classifier: Programming Language :: Python :: 3.10
42
+ Classifier: Programming Language :: Python :: 3.11
43
+ Classifier: Programming Language :: Python :: 3.12
44
+ Classifier: Topic :: Scientific/Engineering
45
+ Classifier: Topic :: Scientific/Engineering :: Information Analysis
46
+ Classifier: Topic :: Software Development :: Quality Assurance
47
+ Classifier: Topic :: Software Development :: Testing
48
+ Classifier: Typing :: Typed
49
+ Requires-Python: >=3.10
50
+ Requires-Dist: numpy>=1.22
51
+ Provides-Extra: dev
52
+ Requires-Dist: mypy>=1.8; extra == 'dev'
53
+ Requires-Dist: pytest-cov>=4.1; extra == 'dev'
54
+ Requires-Dist: pytest>=7.4; extra == 'dev'
55
+ Requires-Dist: ruff>=0.3; extra == 'dev'
56
+ Description-Content-Type: text/markdown
57
+
58
+ # sensorlint
59
+
60
+ Assertions for sensor-data pipelines. Every check aborts instead of warning.
61
+
62
+ ```python
63
+ from sensorlint import safe_gunzip, assert_expected_length, assert_not_clipped
64
+
65
+ payload = safe_gunzip(raw_bytes) # refuses to return partial data
66
+ samples = np.frombuffer(payload, dtype=np.int16)
67
+ assert_expected_length(samples, 20480)
68
+ assert_not_clipped(samples, adc_max=32767, leading_samples=2048)
69
+ ```
70
+
71
+ ---
72
+
73
+ ## The thesis
74
+
75
+ Sensor pipelines do not usually fail by crashing. They fail by producing
76
+ believable numbers.
77
+
78
+ I built this because the bugs that cost me the most time all had the same
79
+ shape: nothing raised, nothing logged, the output was the right dtype and the
80
+ right shape and full of finite values in a plausible range, and it was wrong.
81
+ There was no traceback to work backwards from, and by the time anyone noticed,
82
+ the window for reproducing what had happened had closed.
83
+
84
+ Two examples, both of which have their own check in this library.
85
+
86
+ **A partial decode that one decoder calls an error and another calls a small
87
+ file.** `gzip.decompress` raises `EOFError` on a truncated stream. On the exact
88
+ same bytes, `zlib.decompressobj().decompress` returns whatever it managed to
89
+ inflate and does not raise — which is correct behaviour for that API, because it
90
+ has no way to know you are not about to feed it the rest of the stream. If any
91
+ layer of an ingest path falls back from the strict decoder to the permissive
92
+ one, truncated objects stop being errors and start being small files. A batch
93
+ job over a corrupted prefix then processes a fraction of what it thinks it
94
+ processed, or nothing at all, and exits zero. `examples/truncated_gzip.py`
95
+ demonstrates this end to end.
96
+
97
+ **A payload in raw converter counts whose scale factors live somewhere else.**
98
+ The counts are integers in a valid range. The spectrum is correct. The trend
99
+ over time is correct. Every relative comparison you make is correct. The values
100
+ are wrong by exactly one constant factor, and there is nothing to inspect,
101
+ because the output is plausible at any scale. A threshold of 0.4 g gets compared
102
+ against a value of 8,100 counts and fires, or does not fire, for reasons
103
+ unrelated to vibration.
104
+
105
+ The common property is that the wrong answer is indistinguishable from the right
106
+ one by looking at it. So the defence has to be an assertion made at the point
107
+ where the information still exists, not an inspection made later.
108
+
109
+ ---
110
+
111
+ ## Install
112
+
113
+ ```bash
114
+ pip install git+https://github.com/<user>/sensorlint
115
+ ```
116
+
117
+ Requires Python 3.10+ and numpy. Nothing else. (Not yet on PyPI; the wheel
118
+ builds and installs cleanly from a checkout.)
119
+
120
+ From a checkout:
121
+
122
+ ```bash
123
+ make install # creates .venv and installs with dev extras
124
+ make all # ruff, mypy --strict, pytest
125
+ ```
126
+
127
+ ---
128
+
129
+ ## 30-second quickstart
130
+
131
+ ```python
132
+ import numpy as np
133
+ import sensorlint as sl
134
+
135
+ # 1. Decode without accepting partial data.
136
+ payload = sl.safe_gunzip(raw_bytes)
137
+ samples = np.frombuffer(payload, dtype=np.int16)
138
+
139
+ # 2. Assert what you were told to expect.
140
+ sl.assert_expected_length(samples, 20480)
141
+ sl.assert_not_clipped(samples, adc_max=32767, leading_samples=2048)
142
+ sl.assert_scale_declared(samples, scale=6.1e-5, unit="g")
143
+ sl.assert_not_stale(samples, max_frozen_run=64)
144
+ sl.assert_sample_rate(timestamps, claimed_fs=20_000, tolerance_pct=0.5)
145
+
146
+ # 3. At the end of every stage, the cheapest check in the library.
147
+ sl.assert_nonzero_kept(len(inputs), len(outputs), "featurize")
148
+ ```
149
+
150
+ Every `assert_*` has a `check_*` twin that returns a structured result instead
151
+ of raising, so you can run the same checks over a fleet and get a report rather
152
+ than a traceback:
153
+
154
+ ```python
155
+ results = [sl.check_not_clipped(load(f), adc_max=32767, target=f) for f in files]
156
+ print(sl.format_report(results))
157
+ ```
158
+
159
+ ```
160
+ sensorlint
161
+ ==========
162
+
163
+ FAIL 1187/1204 checks passed (98.6%)
164
+
165
+ check passed failed pass rate
166
+ ----------- ------ ------ ---------
167
+ not_clipped 1187 17 98.6%
168
+ TOTAL 1187 17 98.6%
169
+
170
+ failures by severity: error=17
171
+
172
+ worst offenders:
173
+ target failures
174
+ -------------- --------
175
+ node-14/ch0 9
176
+ node-03/ch1 5
177
+ ```
178
+
179
+ ---
180
+
181
+ ## The checks
182
+
183
+ | # | Check | Catches | Key option |
184
+ |---|-------|---------|------------|
185
+ | 1 | `assert_decoded_fully` / `safe_gunzip` | Truncated or partial decompression a permissive decoder swallowed | reads `decompressobj.eof`, not the data |
186
+ | 2 | `assert_expected_length` | Short reads, dropped samples, buffer truncation | `tolerance`, `tolerance_pct` |
187
+ | 3 | `assert_not_clipped` | Samples railed at the converter limit | `leading_samples` — the variant that matters |
188
+ | 4 | `assert_scale_declared` | Raw ADC counts with no scale factor and unit | `looks_like_raw_counts` heuristic |
189
+ | 5 | `assert_not_stale` | Frozen sensors, and linearly interpolated stretches posing as measurements | `max_interpolated_run` |
190
+ | 6 | `assert_sample_rate` | A claimed rate that the timestamps contradict; gaps; jitter | `gap_factor`, `allow_gaps` |
191
+ | 7 | `detect_channel_transposition` | Swapped channel pairs, plus a **false-positive rate with no ground truth** | `transposition_run_lengths` |
192
+ | 8 | `assert_coverage` | Data *around* a date range rather than *on* it | `bin_seconds`, `min_density` |
193
+ | 9 | `assert_nonzero_kept` | A stage that processed zero records and reported success | `@nonzero_kept`, `pipeline_stage` |
194
+
195
+ ---
196
+
197
+ ## Design principles
198
+
199
+ **1. Abort, never warn.** There is no warning mode and there will not be one. A
200
+ warning in a batch job is a log line nobody reads, and the entire premise of
201
+ this library is that these failures are already invisible. Every `assert_*`
202
+ raises. `SensorLintError` subclasses `AssertionError`, but unlike the `assert`
203
+ statement it is not removed by `python -O`, which matters because production
204
+ batch jobs are where these checks earn their keep.
205
+
206
+ **2. Every check exists twice.** The `check_*` form returns a `CheckResult`
207
+ dataclass — check name, pass/fail, severity, message, and a `details` dict of
208
+ the numbers behind the verdict. The `assert_*` form is a thin wrapper that calls
209
+ it and raises. This is not decoration: a single file failing a check is an
210
+ exception, and ten thousand files failing a check is a report, and you cannot
211
+ build a report out of tracebacks. `sensorlint.report` aggregates results by
212
+ check, by severity and by target.
213
+
214
+ **3. Zero heavy dependencies.** numpy and the standard library. Nothing else,
215
+ ever. A validation library that is annoying to install is a validation library
216
+ that gets skipped at exactly the ingest boundary where it was supposed to run.
217
+
218
+ **4. Every check has a test that constructs the real failure.** Not a mock. The
219
+ gzip tests truncate actual compressed bytes and confirm that `zlib` tolerates
220
+ what `gzip` rejects. The clipping tests actually saturate arrays. The
221
+ transposition tests actually swap two real signal arrays with different energy.
222
+ The staleness tests actually overwrite a stretch of a noisy signal with
223
+ `np.linspace` between its endpoints. If the failure cannot be constructed, I do
224
+ not trust the check to detect it.
225
+
226
+ ---
227
+
228
+ ## The checks in detail
229
+
230
+ ### 1. Decoded fully
231
+
232
+ The tell is not the data, it is the decompressor state. A complete stream leaves
233
+ `decompressobj.eof` set to `True`; a truncated one leaves it `False` with nothing
234
+ in `unused_data`. That single boolean is the whole check. `probe_compressed`
235
+ re-decodes strictly and reports `eof`, member count, unconsumed tail, and the
236
+ gzip CRC32/ISIZE trailer; `check_decoded_fully` compares that against what your
237
+ decoder handed you.
238
+
239
+ `safe_gunzip` is the function to reach for instead of `gzip.decompress` or a
240
+ hand-rolled `decompressobj` loop. It verifies end-of-stream before returning
241
+ anything, so a truncated object raises rather than becoming a short buffer that
242
+ looks like a small file. Multi-member gzip and trailing NUL padding are handled;
243
+ neither is corruption.
244
+
245
+ ### 2. Expected length
246
+
247
+ The most boring check here and one of the most common failures. A short read
248
+ gives you an array of entirely real measurements — of the first 200 ms of a
249
+ one-second record. Your RMS is a real RMS. Your spectrum has a fifth of the
250
+ frequency resolution you documented. Nothing downstream can tell. Write the
251
+ expected length down at the point where you still know it.
252
+
253
+ ### 3. Not clipped
254
+
255
+ A saturated converter returns its ceiling, repeatedly, while the input stays out
256
+ of range. The result has a *lower* peak-to-peak than the truth and a *higher*
257
+ apparent harmonic content, so it flatters every summary statistic you compute.
258
+
259
+ The reason this is more than a one-liner is the leading window. Railing very
260
+ often happens only at the start of a recording, while a coupling capacitor
261
+ charges or a gain stage settles, and then decays into perfectly normal values.
262
+ Over a 60-second record, 200 ms of solid rail is 0.3% of the samples — under any
263
+ whole-array threshold you would actually set. So the first block of every file is
264
+ garbage, every file, forever, and the per-file statistics look fine. Pass
265
+ `leading_samples=N` and the window is tested separately.
266
+
267
+ When `adc_max` is not supplied, rails are auto-detected from *repeated exact*
268
+ extreme values. Noisy data reaches its maximum once; a saturated converter
269
+ returns bit-identical extremes many times over. The signature is the repeat
270
+ count, not the magnitude.
271
+
272
+ ### 4. Scale declared
273
+
274
+ Refuses to proceed on data that looks like raw counts with no declared scale
275
+ factor and physical unit. `looks_like_raw_counts` scores several signals:
276
+ integer dtype, float values that are all whole numbers (the usual way counts get
277
+ past a type check), an extreme sitting just under a common converter full scale,
278
+ and a quantization step of exactly 1.
279
+
280
+ It is deliberately conservative in one direction. It would rather make you
281
+ declare a unit for data that was already in physical units than let undeclared
282
+ counts through, because there is no way to detect the second mistake afterwards.
283
+
284
+ ### 5. Not stale
285
+
286
+ Frozen data has zero variance over the stuck stretch, which reads to most
287
+ anomaly detectors as "very stable". Interpolated data hides better, because it
288
+ still moves.
289
+
290
+ The discriminant for interpolation is the second difference. A linearly
291
+ interpolated stretch has a constant first difference, so its second difference is
292
+ zero to floating-point precision, across many consecutive samples. Real sensor
293
+ data essentially never does this — even a very clean, very slowly varying signal
294
+ has noise in the last bits, and that is enough. A long run of `d2 == 0` is not a
295
+ property of quiet data, it is a signature of synthesis. Zero slope is reported as
296
+ frozen; constant non-zero slope is reported as interpolated.
297
+
298
+ ### 6. Sample rate
299
+
300
+ The sampling rate is nearly always metadata; the timestamps are data. A spectrum
301
+ computed with a claimed 20 kHz on data actually sampled at 19,531 Hz puts every
302
+ peak 2.4% low, so a fault frequency lands next to where you were looking rather
303
+ than on it. The plot is beautiful.
304
+
305
+ The check compares the claim against the median observed interval (median, so
306
+ one dropped block does not move the estimate), reports jitter, and refuses to
307
+ call a record contiguous when it has holes. Concatenating across a hole produces
308
+ a discontinuity that reads as broadband energy — indistinguishable from an
309
+ impact event, which is often exactly what the pipeline is looking for.
310
+
311
+ ### 7. Channel transposition — the one I would want to talk about
312
+
313
+ Two channels get swapped: a connector goes back the wrong way, a channel map is
314
+ edited, a wiring change is made and the metadata is not. Both channels still
315
+ contain real signal. Every per-channel statistic afterwards is a correct
316
+ measurement of the wrong thing, so a rising trend appears on the wrong location
317
+ and the one that is actually degrading looks stable.
318
+
319
+ The detector uses the inter-channel energy ratio, scored as two hypotheses
320
+ rather than one threshold. If A normally runs 8 dB above B, then a file where A
321
+ runs 8 dB *below* B is not "a quiet day on A". The verdict compares the distance
322
+ from the observed ratio to `+expected` against the distance to `-expected`, and
323
+ requires a margin. Ratios in the ambiguous middle return `transposed=False` with
324
+ low confidence, which is the honest answer, and channels whose expected
325
+ separation is too small to discriminate raise rather than emit coin flips.
326
+
327
+ That detector has a false-positive rate and no ground truth to measure it
328
+ against. Nobody is going to physically inspect a thousand installations to tell
329
+ you how often you were wrong.
330
+
331
+ But transposition has a property noise does not: **it is temporally contiguous.**
332
+ A cable that got swapped stays swapped until someone unplugs it. So across a
333
+ chronologically ordered sequence of files from one installation, real
334
+ transposition appears as a *run* — dozens or hundreds of consecutive positives —
335
+ while a detector firing on noise produces isolated single-file positives
336
+ scattered through long stretches of negatives.
337
+
338
+ That gives you an error rate for free:
339
+
340
+ ```python
341
+ verdicts = [detect_channel_transposition(a, b, expected_ratio_db=8.0) for a, b in files]
342
+ report = transposition_run_lengths(verdicts, min_run_length=5)
343
+
344
+ report.contiguous_runs # ((150, 40),) — the actual wiring change
345
+ report.n_isolated # 3 — lone flags, believed spurious
346
+ report.estimated_false_positive_rate # 3 / 360 = 0.0083, with no labels
347
+ report.n_isolated_expected_if_random # what pure noise would have produced
348
+ report.clustered # True: these flags carry real structure
349
+ ```
350
+
351
+ `estimated_false_positive_rate` is the count of positives sitting in short runs
352
+ divided by the number of files believed to be true negatives.
353
+ `n_isolated_expected_if_random` is the same count under an i.i.d. Bernoulli model
354
+ at the observed flag rate — if your detector's isolated count matches it, the
355
+ positives have no temporal structure and are consistent with pure chance.
356
+
357
+ The same structure tells you when to believe a positive. One flagged file in a
358
+ year of clean ones is your noise floor. Forty consecutive flagged files starting
359
+ the week after a site visit is a wiring change.
360
+
361
+ Use `estimate_expected_ratio_db(reference_pairs)` to learn the baseline from
362
+ known-good files. It takes the median, so a single already-swapped file in the
363
+ reference set cannot drag the baseline toward zero and quietly disarm the check.
364
+
365
+ ### 8. Coverage
366
+
367
+ You ask for a date range. Something returns files whose timestamps fall inside
368
+ it. The earliest is at or before the start, the latest is at or after the end,
369
+ both are true, and you conclude the range is covered.
370
+
371
+ It is not. A file on the first day and a file on the last day satisfy every
372
+ endpoint test ever written while leaving a month-long hole in between. The
373
+ statistics you compute over "the last 90 days" are computed over two days, the
374
+ trend line has two points, and the baseline is a baseline of nothing.
375
+
376
+ This check bins the range and requires the bins to be populated. Endpoint
377
+ coverage is necessary and nowhere near sufficient. It reports the longest run of
378
+ empty bins, so the failure message names the outage rather than just its
379
+ existence.
380
+
381
+ ### 9. Nonzero kept
382
+
383
+ **This is the cheapest assertion in the library and it catches the most
384
+ expensive class of bug.**
385
+
386
+ A stage iterates over its inputs, filters, transforms, writes its outputs. The
387
+ list was empty, or every record was filtered out by a predicate that stopped
388
+ matching after an upstream schema change. The loop body never runs. There is
389
+ nothing to raise. The stage writes an empty output, logs "stage complete", and
390
+ exits zero. Everything downstream then works perfectly on nothing. The job is
391
+ green. The dashboard is green.
392
+
393
+ It is one comparison against zero. It needs no thresholds, no domain knowledge,
394
+ and no numerical thinking, and it costs nothing at runtime. And a silent empty
395
+ result can survive for weeks, because it does not look like a failure — it looks
396
+ like a quiet period. By the time somebody notices the gap, the logs have rotated
397
+ and the window for working out what went wrong has closed.
398
+
399
+ Put it at the end of every stage. Not the important ones. Every one. Three forms
400
+ are provided so there is never a reason to skip it:
401
+
402
+ ```python
403
+ assert_nonzero_kept(len(inputs), len(outputs), "featurize")
404
+
405
+ @nonzero_kept("decode_batch")
406
+ def decode_batch(objects): ...
407
+
408
+ with pipeline_stage("decode", n_input=len(objects)) as stage:
409
+ for obj in objects:
410
+ stage.keep()
411
+ ```
412
+
413
+ The context manager does not mask an exception raised inside the block: a stage
414
+ that crashed already failed loudly, and replacing its traceback with an
415
+ empty-output complaint would hide the real cause.
416
+
417
+ `examples/silent_empty_stage.py` runs a three-stage pipeline before and after an
418
+ upstream field rename, with and without the assertion.
419
+
420
+ ---
421
+
422
+ ## Command line
423
+
424
+ ```bash
425
+ python -m sensorlint check recording.npy \
426
+ --expected-length 20480 \
427
+ --adc-max 32767 --leading-samples 2048 \
428
+ --scale 6.1e-5 --unit g
429
+
430
+ python -m sensorlint check series.csv \
431
+ --column 1 --timestamp-column 0 --fs 20000 --json
432
+ ```
433
+
434
+ Exits `0` when every applicable check passed, `1` when any failed, `2` on a
435
+ usage or loading error. Checks with no corresponding flag are skipped rather
436
+ than run against a guessed threshold, and skipped checks are printed explicitly
437
+ so they cannot hide inside a green run.
438
+
439
+ ---
440
+
441
+ ## Errors
442
+
443
+ One exception class per check family, so a caller who genuinely wants to
444
+ tolerate one class of problem can do so narrowly instead of writing
445
+ `except Exception` and thereby also swallowing the problem that matters.
446
+
447
+ ```
448
+ SensorLintError(AssertionError)
449
+ ├── DecodeError ├── StalenessError
450
+ ├── LengthError ├── SampleRateError
451
+ ├── ClippingError ├── TranspositionError
452
+ ├── ScaleError ├── CoverageError
453
+ └── EmptyStageError
454
+ ```
455
+
456
+ Every exception carries the `CheckResult` that produced it, so the structured
457
+ detail survives the raise:
458
+
459
+ ```python
460
+ try:
461
+ assert_not_clipped(samples, adc_max=32767, leading_samples=2048)
462
+ except ClippingError as exc:
463
+ exc.details["leading_railed_fraction"] # 0.37
464
+ exc.result.target # "node-14/ch0"
465
+ ```
466
+
467
+ ---
468
+
469
+ ## Notes
470
+
471
+ **On the Python version.** The package declares `requires-python = ">=3.10"`,
472
+ but every module carries `from __future__ import annotations` and avoids
473
+ 3.10-only syntax at runtime, so the source also imports cleanly on 3.9. The ruff
474
+ config disables the pyupgrade rules that would break that. This is deliberate:
475
+ ingest code is often the last thing on an old interpreter, and a validation
476
+ library that cannot be installed there is a validation library that does not run.
477
+
478
+ **Where this belongs.** At the ingest boundary: the first thing that touches a
479
+ payload after it is read, before anything computes on it. Wiring it into the
480
+ ingest stage of my [`bearing-watch`](https://github.com/johnhagedorncs/bearing-watch)
481
+ project is next.
482
+
483
+ **Everything in this repository is synthetic.** All example data, thresholds and
484
+ signals are generated in-repo. No proprietary code, data or results from any
485
+ employer appear anywhere in it.
486
+
487
+ ---
488
+
489
+ ## Development
490
+
491
+ ```bash
492
+ make install # .venv + dev extras
493
+ make test # pytest
494
+ make test-cov # pytest with coverage
495
+ make lint # ruff
496
+ make typecheck # mypy --strict
497
+ make examples # run both example scripts
498
+ make all # lint + typecheck + test
499
+ ```
500
+
501
+ 464 tests, 97% branch coverage, mypy strict clean. CI runs ruff, mypy strict, and
502
+ pytest on 3.10 / 3.11 / 3.12.
503
+
504
+ ---
505
+
506
+ ## License
507
+
508
+ MIT. See [LICENSE](LICENSE).