trainspotter 0.1.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. trainspotter-0.1.2/.gitignore +16 -0
  2. trainspotter-0.1.2/CHANGELOG.md +53 -0
  3. trainspotter-0.1.2/LICENSE +21 -0
  4. trainspotter-0.1.2/PKG-INFO +333 -0
  5. trainspotter-0.1.2/README.md +297 -0
  6. trainspotter-0.1.2/pyproject.toml +118 -0
  7. trainspotter-0.1.2/src/trainspotter/__init__.py +20 -0
  8. trainspotter-0.1.2/src/trainspotter/cli.py +133 -0
  9. trainspotter-0.1.2/src/trainspotter/detectors/__init__.py +106 -0
  10. trainspotter-0.1.2/src/trainspotter/detectors/base.py +147 -0
  11. trainspotter-0.1.2/src/trainspotter/detectors/divergence.py +144 -0
  12. trainspotter-0.1.2/src/trainspotter/detectors/eval_noise.py +81 -0
  13. trainspotter-0.1.2/src/trainspotter/detectors/grad_norm.py +145 -0
  14. trainspotter-0.1.2/src/trainspotter/detectors/loss_floor.py +104 -0
  15. trainspotter-0.1.2/src/trainspotter/detectors/lr_schedule.py +202 -0
  16. trainspotter-0.1.2/src/trainspotter/detectors/overfitting.py +125 -0
  17. trainspotter-0.1.2/src/trainspotter/detectors/plateau.py +176 -0
  18. trainspotter-0.1.2/src/trainspotter/detectors/spikes.py +91 -0
  19. trainspotter-0.1.2/src/trainspotter/detectors/throughput.py +136 -0
  20. trainspotter-0.1.2/src/trainspotter/model.py +68 -0
  21. trainspotter-0.1.2/src/trainspotter/py.typed +0 -0
  22. trainspotter-0.1.2/src/trainspotter/readers/__init__.py +82 -0
  23. trainspotter-0.1.2/src/trainspotter/readers/common.py +66 -0
  24. trainspotter-0.1.2/src/trainspotter/readers/hf.py +82 -0
  25. trainspotter-0.1.2/src/trainspotter/readers/lightning.py +22 -0
  26. trainspotter-0.1.2/src/trainspotter/readers/tabular.py +91 -0
  27. trainspotter-0.1.2/src/trainspotter/readers/tensorboard.py +39 -0
  28. trainspotter-0.1.2/src/trainspotter/readers/wandb.py +50 -0
  29. trainspotter-0.1.2/src/trainspotter/report/__init__.py +9 -0
  30. trainspotter-0.1.2/src/trainspotter/report/html_report.py +599 -0
  31. trainspotter-0.1.2/src/trainspotter/report/json_report.py +71 -0
  32. trainspotter-0.1.2/src/trainspotter/report/terminal.py +84 -0
@@ -0,0 +1,16 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ *.egg-info/
4
+ .eggs/
5
+ build/
6
+ dist/
7
+ .venv/
8
+ venv/
9
+ .pytest_cache/
10
+ .mypy_cache/
11
+ .ruff_cache/
12
+ .coverage
13
+ htmlcov/
14
+ *.log
15
+ .DS_Store
16
+ site/
@@ -0,0 +1,53 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented in this file.
4
+ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
5
+ and this project uses [Semantic Versioning](https://semver.org/).
6
+
7
+ ## [0.1.2] - 2026-10-01
8
+
9
+ ### Added
10
+
11
+ - Published to PyPI: `pip install trainspotter`. The README's images and links
12
+ are rewritten to absolute URLs at build time so they work on the project
13
+ page.
14
+ - `trainspotter --version`.
15
+
16
+ ## [0.1.1] - 2026-09-30
17
+
18
+ ### Fixed
19
+
20
+ - `lr_schedule` reported a short warmup as an "LR discontinuity" when the log
21
+ was coarse (the real `transformers.Trainer` example: a 20-step warmup
22
+ logged every 5 steps, so each logged interval covers 25% of the LR range).
23
+ A jump is now reported only if it stands out from the logged intervals
24
+ around it, and the message gives the step gap instead of always saying
25
+ "in a single step".
26
+ - Slope p-values print as `p<0.0001` instead of `p=0` when the normal
27
+ approximation underflows.
28
+ - An enormous rise reads as "300x its minimum" instead of "29891.0%", and a
29
+ one-step span reads "step 320" instead of "steps 320-320".
30
+ - HTML report: rounding the axis ends outward drew a gridline and tick label
31
+ above the chart frame (the divergence example's loss chart showed a "5"
32
+ over its legend); ticks now stay inside the plotted range. Off-scale value
33
+ labels at the right edge sit beside their markers instead of on the frame.
34
+
35
+ ## [0.1.0] - 2026-09-24
36
+
37
+ ### Added
38
+
39
+ - Readers for Hugging Face `trainer_state.json`, generic CSV/JSONL metric
40
+ logs, Weights & Biases history CSV exports, PyTorch Lightning
41
+ `CSVLogger` `metrics.csv`, and TensorBoard event files (optional
42
+ `tensorboard` extra).
43
+ - Nine detectors: loss spikes, divergence (NaN/Inf and sustained growth),
44
+ plateaus, overfitting onset, LR schedule anomalies, gradient-norm
45
+ explosions/clipping saturation, throughput regressions, eval-metric
46
+ noise, and a loss-floor/leakage heuristic.
47
+ - `trainspotter analyze`: terminal, JSON, and self-contained HTML reports;
48
+ `--fail-on warning|error` for CI gating.
49
+ - `trainspotter watch`: tails a growing log file and prints new findings
50
+ as they appear.
51
+ - Real example logs under `examples/`, generated by training a small
52
+ numpy MLP on scikit-learn's digits dataset with deliberately induced
53
+ pathologies (`examples/generate_examples.py`).
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Anton Soloviev
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,333 @@
1
+ Metadata-Version: 2.5
2
+ Name: trainspotter
3
+ Version: 0.1.2
4
+ Summary: Automatic diagnosis of loss curves and training logs: spot divergence, spikes, overfitting, LR and throughput problems before you burn another GPU-day.
5
+ Project-URL: Homepage, https://antonsoo.github.io/trainspotter/
6
+ Project-URL: Repository, https://github.com/antonsoo/trainspotter
7
+ Project-URL: Issues, https://github.com/antonsoo/trainspotter/issues
8
+ Project-URL: Changelog, https://github.com/antonsoo/trainspotter/blob/main/CHANGELOG.md
9
+ Author-email: Anton Soloviev <anton@praviel.com>
10
+ License: MIT
11
+ License-File: LICENSE
12
+ Keywords: cli,diagnostics,huggingface,loss-curve,machine-learning,mlops,pytorch,training
13
+ Classifier: Development Status :: 4 - Beta
14
+ Classifier: Intended Audience :: Science/Research
15
+ Classifier: License :: OSI Approved :: MIT License
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.10
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
21
+ Classifier: Typing :: Typed
22
+ Requires-Python: >=3.10
23
+ Requires-Dist: numpy>=1.24
24
+ Provides-Extra: dev
25
+ Requires-Dist: mypy>=1.10; extra == 'dev'
26
+ Requires-Dist: pytest-cov>=5.0; extra == 'dev'
27
+ Requires-Dist: pytest>=8.0; extra == 'dev'
28
+ Requires-Dist: ruff>=0.6; extra == 'dev'
29
+ Requires-Dist: scikit-learn>=1.3; extra == 'dev'
30
+ Requires-Dist: tbparse>=0.0.8; extra == 'dev'
31
+ Provides-Extra: examples
32
+ Requires-Dist: scikit-learn>=1.3; extra == 'examples'
33
+ Provides-Extra: tensorboard
34
+ Requires-Dist: tbparse>=0.0.8; extra == 'tensorboard'
35
+ Description-Content-Type: text/markdown
36
+
37
+ # trainspotter
38
+
39
+ **Spot what went wrong in a training run before you burn another GPU-day.**
40
+ Automatic diagnosis of loss curves and training logs.
41
+
42
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://github.com/antonsoo/trainspotter/blob/main/LICENSE)
43
+ ![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)
44
+ [![Live demo](https://img.shields.io/badge/live%20demo-sample%20reports-4fc3f7)](https://antonsoo.github.io/trainspotter/)
45
+ [![Hugging Face](https://img.shields.io/badge/Hugging%20Face-workbench-ffd21e)](https://huggingface.co/spaces/antonsoloviev/trainspotter)
46
+
47
+ Every ML engineer has stared at a TensorBoard or W&B chart trying to answer
48
+ "is this run okay?" -- and caught the answer late: a spike that should have
49
+ triggered a restart three hours ago, an eval curve that's been overfitting
50
+ since epoch 4, a learning-rate schedule that never actually warmed up. The
51
+ signal was in the log the whole time. trainspotter reads the log a run
52
+ already writes and turns eyeballing into a report: which pathologies were
53
+ detected, where (step ranges), the evidence behind each, and the usual
54
+ fixes. It runs post-hoc on a finished log, live via `watch` on a growing
55
+ one, or as a CI gate that fails the build.
56
+
57
+ ## Report
58
+
59
+ ![Top of the trainspotter HTML report on the divergence example: a header with the run's source and step range, an error/warning/info summary with a health-strip timeline, and the combined train/loss + eval/loss chart -- 320 steps of ordinary cosine-schedule training, shaded regions marking detected findings, then a labeled spike as the run breaks -- followed by the findings log, error findings first: two loss spikes and the divergence itself.](https://raw.githubusercontent.com/antonsoo/trainspotter/main/docs/assets/html-report-divergence.png)
60
+
61
+ *(top of the report -- [full report, all 5 charts + all 6 findings](https://github.com/antonsoo/trainspotter/blob/main/docs/assets/html-report-divergence-full.png))*
62
+
63
+ ![Terminal output of `trainspotter analyze divergence.trainer_state.json --fail-on error`, showing all 6 findings ordered by severity -- three ERROR findings (two loss spikes, the divergence) first, then one WARNING (the gradient-norm explosion), then two INFO (the downgraded overfitting onset and the LR discontinuity) -- each with step range, detector name, message, and a one-line fix suggestion.](https://raw.githubusercontent.com/antonsoo/trainspotter/main/docs/assets/terminal-divergence.png)
64
+
65
+ Browse every example's full HTML report at
66
+ **[antonsoo.github.io/trainspotter](https://antonsoo.github.io/trainspotter/)**
67
+ (built by `scripts/build_site.py`). Both images above are real output from `examples/divergence.trainer_state.json` (see
68
+ [Real demo data](#real-demo-data) for exactly how that log was produced).
69
+
70
+ ## Quickstart
71
+
72
+ ```bash
73
+ pip install trainspotter
74
+ git clone --depth 1 https://github.com/antonsoo/trainspotter && cd trainspotter
75
+ trainspotter analyze examples/divergence.trainer_state.json
76
+ ```
77
+
78
+ Or point it at your own log:
79
+
80
+ ```bash
81
+ trainspotter analyze path/to/trainer_state.json --output html --out report.html
82
+ ```
83
+
84
+ ## Features
85
+
86
+ - **Five readers**, auto-detected from the file: Hugging Face
87
+ `trainer_state.json`, generic CSV/JSONL, Weights & Biases history CSV
88
+ exports, PyTorch Lightning `CSVLogger` `metrics.csv`, and TensorBoard
89
+ event files (`pip install trainspotter[tensorboard]`).
90
+ - **Nine detectors** -- see the [table below](#detectors) -- each with a
91
+ documented algorithm, tunable thresholds, and stated false-positive modes.
92
+ - **Three report formats**: a color-coded terminal report, JSON (stable
93
+ schema, for tooling), and a self-contained HTML report (inline SVG
94
+ charts, no CDN, no JavaScript -- opens from disk, works offline forever).
95
+ - **CI gating**: `--fail-on warning|error` exits non-zero if any finding at
96
+ or above that severity was found.
97
+ - **Live tailing**: `trainspotter watch <file>` re-reads a growing log on
98
+ an interval and prints only newly-appeared findings.
99
+ - **numpy is the only required dependency.** No pandas, no plotting
100
+ library, no web framework. (`tensorboard` support is an optional extra
101
+ since parsing the event-file format properly needs `tbparse`.)
102
+
103
+ ## Usage
104
+
105
+ ```bash
106
+ # Terminal report (default), with a CI-style exit code
107
+ trainspotter analyze run/trainer_state.json --fail-on error
108
+
109
+ # Machine-readable JSON
110
+ trainspotter analyze run/trainer_state.json --output json > findings.json
111
+
112
+ # Self-contained HTML report
113
+ trainspotter analyze run/trainer_state.json --output html --out report.html
114
+
115
+ # Force a reader instead of auto-detecting from the extension
116
+ trainspotter analyze run/metrics.csv --format lightning
117
+
118
+ # Tail a log that's still being written
119
+ trainspotter watch run/trainer_state.json --interval 5 --fail-on error
120
+ ```
121
+
122
+ Real terminal output (from `examples/overfitting.trainer_state.json`, one
123
+ finding shown):
124
+
125
+ ```
126
+ [WARN ] steps 125-799 overfitting Overfitting onset
127
+ Best eval/loss was 0.526 at step 125; it's since risen to 0.5554 (5.6%, slope p<0.0001).
128
+ Meanwhile train/loss kept falling over the same range (slope < 0) -- the classic
129
+ overfitting signature.
130
+ metric: eval/loss
131
+ fix: Use the checkpoint at step 125 (best eval/loss), not the last one.
132
+ ```
133
+
134
+ ### CI usage
135
+
136
+ ```yaml
137
+ - name: Diagnose the training run
138
+ run: trainspotter analyze runs/latest/trainer_state.json --fail-on error
139
+ # exits 1 if any error-severity pathology was detected
140
+ ```
141
+
142
+ ## Detectors
143
+
144
+ Every detector's full algorithm and false-positive list lives in its
145
+ module's docstring (`src/trainspotter/detectors/*.py`); this table is the
146
+ summary.
147
+
148
+ | Detector | Detects | How | Stated false-positive modes |
149
+ |---|---|---|---|
150
+ | `spikes` | A sudden jump in a loss metric | Modified z-score (`0.6745·(x−median)/MAD`) against a trailing rolling median/MAD; flags \|z\| > 6 | Deliberate LR restarts/curriculum jumps; bumpy metrics with a too-small window |
151
+ | `divergence` | NaN/Inf, or sustained upward drift | Any non-finite value; or an OLS slope > 0 with p < 0.05 over the tail window, and the last value ≥ 1.5× the running minimum | Metrics meant to increase (only checks loss-type metrics by default); a temporary cyclic-schedule upswing |
152
+ | `plateau` | A metric that's stalled, not just converged | Sliding-window OLS slope test (p ≥ 0.2 = not significant) plus a < 2% relative-change guard; suppressed if the window has already recovered ≥ 80% of the metric's total drop, or `lr` has decayed to ≤ 30% of its peak | A real stall near a coincidentally low value, or right as an unrelated LR decay finishes, can be wrongly suppressed by the convergence check |
153
+ | `overfitting` | Train still improving while eval worsens | Finds eval's running-minimum step; tests whether the slope after it is significantly positive (p < 0.1) and ≥ 1% above the minimum | A noisy eval set can show a false uptick; a mid-run change in eval data |
154
+ | `lr_schedule` | Missing warmup, LR rising after its peak, discontinuities | Ramp check on the first value vs. peak; running-max monotonicity after the peak; a jump of ≥ 20% of the LR's total range that isn't one leg of a steady ramp | Cyclic/warm-restart schedules trip the last two by design; step-decay drops are reported as `info` |
155
+ | `grad_norm` | Gradient-norm explosions and clipping saturation | Same z-score as `spikes` (one-sided) for explosions, threshold 10 (higher than `spikes`' 6 -- a single z of 6-8 is common early-training noise); fraction of a trailing window within 0.5% of its own max, for saturation | Norms logged post-clip are saturated by construction; an isolated explosion with no coinciding loss spike is weaker evidence than one that lines up with a `spikes` finding |
156
+ | `throughput` | Step-time / throughput regressions | Median step time, early-run baseline window vs. recent window, flags ≥ 1.4× | Checkpoint/eval steps inflate single points; early-run kernel/dataloader warmup can pollute the baseline |
157
+ | `eval_noise` | An eval metric too noisy to rank checkpoints by | Median absolute step-to-step jitter as a fraction of the run's total improvement; flags ≥ 15% | An already-converged run has small total improvement by construction |
158
+ | `loss_floor` | Implausibly low train loss very early (heuristic) | Loss < 0.05 absolute **and** < 5% of its starting value within the first 10% of steps | Explicitly a heuristic: an easy task, a small-scale loss, or a resumed checkpoint all look identical to this |
159
+
160
+ ### Trust the healthy case
161
+
162
+ A tool that flags an unremarkable run gets ignored, and recall on real
163
+ incidents matters less than that first impression -- so trainspotter is
164
+ tested for silence, not just for alarms. **On the healthy baseline run,
165
+ trainspotter reports no warnings**, and `tests/test_examples_regression.py`
166
+ asserts this on every commit, alongside asserting that the other four
167
+ examples still flag the pathology they were built to demonstrate. Findings
168
+ are also ordered errors-first, then warnings, then info (chronological
169
+ within a level) -- so on a run that *did* break, the thing that broke
170
+ leads the report instead of sitting below routine early-training noise.
171
+
172
+ ## How it works
173
+
174
+ **Data model.** Every reader converges on one shape
175
+ (`trainspotter.model.Run`): a dict of named metric series, each a list of
176
+ `(step, value, wall_time)` points. Detectors and reports only ever see
177
+ this shape -- they have no idea whether the log came from `transformers`,
178
+ a CSV, or TensorBoard. Readers normalize common spellings (`loss` /
179
+ `train_loss` / `training_loss` all become `train/loss`; `learning_rate` /
180
+ `lr` become `lr`) but pass anything unrecognized straight through, so a
181
+ custom metric like `eval/bleu` still reaches the detectors under its own
182
+ name.
183
+
184
+ **Robust statistics, not raw thresholds.** Spikes and gradient-norm
185
+ explosions are judged by a *modified z-score* against a trailing rolling
186
+ median and MAD (median absolute deviation), not a mean/stddev: a mean and
187
+ stddev are themselves dragged around by the very outlier you're trying to
188
+ detect, while the median and MAD barely move. This is the same construction
189
+ Iglewicz & Hoaglin describe for outlier labeling. Plateau and overfitting
190
+ use an ordinary-least-squares slope with a normal-approximation t-test
191
+ (`math.erf`-based, no `scipy` dependency) rather than eyeballing "did it go
192
+ up or down."
193
+
194
+ **Charts don't let one spike flatten the rest of the curve.** The HTML
195
+ report's axis logic (`report/html_report.py`) picks tick spacing with
196
+ Heckbert's "nice numbers" algorithm (`Graphics Gems`, 1990 -- 1/2/5x10^n
197
+ steps, not whatever an even split of the range happens to produce) and
198
+ floors the axis at 0 for the non-negative metrics trainspotter charts
199
+ (loss, LR, grad_norm, accuracy, step time), instead of padding below zero.
200
+ If the max is more than 1.5x the 99th percentile, the axis caps at that
201
+ instead of stretching to fit one outlier -- the point is still drawn,
202
+ clamped to the top edge with a small triangle and its real value labeled
203
+ next to it, not hidden.
204
+
205
+ **Every finding is honest about its own limits.** Each detector's
206
+ docstring states its false-positive modes in plain language (see the table
207
+ above, or the source for the full version) -- this isn't boilerplate, it's
208
+ meant to be read before you act on a finding. `loss_floor` is explicitly
209
+ labeled a heuristic, not a statistical test, because there's no
210
+ distribution-free way to know a loss is "too low" without knowing the task.
211
+
212
+ ## Real demo data
213
+
214
+ Nothing in `examples/` is fabricated. All five logs were produced by
215
+ `examples/generate_examples.py`, which trains a small hand-written numpy
216
+ MLP (64 → 32 ReLU → 10, softmax cross-entropy, every gradient computed and
217
+ applied by hand -- no autodiff framework) on scikit-learn's `load_digits`
218
+ dataset (1797 real 8×8 handwritten-digit images, bundled with
219
+ scikit-learn, no download). It's deterministic (seeded) and takes about
220
+ three seconds to regenerate on this machine:
221
+
222
+ ```bash
223
+ uv sync --extra examples # or: pip install scikit-learn
224
+ uv run python examples/generate_examples.py
225
+ ```
226
+
227
+ | File | What it is | How the pathology was actually induced |
228
+ |---|---|---|
229
+ | `baseline` | A normal, unremarkable run | Full dataset, weight decay, cosine LR with warmup. **On this run, trainspotter reports no warnings** -- see [Trust the healthy case](#trust-the-healthy-case) for why that's asserted by a test, not just eyeballed |
230
+ | `divergence` | A genuine crash after real training | 320 steps (80% of the run) of ordinary warmup + cosine-decay training -- loss ~2.8 -> ~0.15, eval accuracy up to ~96% -- then a simulated incident at step 320: the LR schedule jumps to 30 **and** the loss function switches to a version with the classic missing-max-subtraction softmax bug. Plain high LR alone was tried first and *didn't* produce real divergence -- see the note in `divergence_run()`'s docstring for why (softmax cross-entropy's gradient is bounded by construction) |
231
+ | `overfitting` | Genuine overfitting | Only 4 training examples per class (40 total), no regularization, 800 steps -- the model memorizes the training set while held-out eval loss turns upward after step 125 |
232
+ | `missing_warmup` | No LR ramp-up | LR set to a constant 0.5 from step 0, no warmup phase at all |
233
+ | `throughput_drop` | A real, measured slowdown | Steps 150-259 insert an actual `time.sleep(0.02)` per step (simulating e.g. I/O contention) -- `step_time` in the log is genuinely measured wall-clock time, not a fabricated number |
234
+
235
+ Each file is committed in both HF `trainer_state.json` format
236
+ (`<name>.trainer_state.json`) and generic CSV (`<name>.csv`), written from
237
+ the same in-memory log so the two are guaranteed consistent.
238
+
239
+ ### A real `transformers.Trainer` run
240
+
241
+ `examples/transformers_tinygpt2.trainer_state.json` is a genuine
242
+ `trainer_state.json` written by `transformers.Trainer` itself (not the
243
+ numpy MLP above): a real 2-layer, 64-dim GPT-2 (`GPT2LMHeadModel` from a
244
+ from-scratch `GPT2Config`, ~330K parameters), trained for 6 epochs / 726
245
+ steps on CPU on a small hand-written toy corpus (short sentences about
246
+ training, tokenized with the real `gpt2` tokenizer), in about 90 seconds
247
+ on this machine. Regenerate it with
248
+ `examples/generate_transformers_example.py` (needs `torch` + `transformers`
249
+ in a separate environment -- see the script's docstring for the exact
250
+ install commands; they are *not* project dependencies).
251
+
252
+ Running trainspotter against it caught a real bug in the HF reader during
253
+ development: `Trainer.train()` appends one extra `log_history` entry after
254
+ training ends -- a run summary with `train_runtime`, `train_samples_per_second`,
255
+ and a `train_loss` key that is the *average* loss over the whole run, not a
256
+ per-step reading. The reader's first version merged that average straight
257
+ into the `train/loss` series (since `train_loss` is one of the recognized
258
+ spellings of `loss`), which showed up as a fake spike on the run's last
259
+ step. The fix -- detect that entry by its unique `train_runtime` key and
260
+ route it to run metadata instead of a metric point -- is
261
+ `src/trainspotter/readers/hf.py`'s `_SUMMARY_MARKER_KEY`, and
262
+ `tests/test_readers.py::test_hf_trainer_state_excludes_the_final_run_summary_entry`
263
+ pins it. This is exactly the kind of gap a hand-written fixture can miss
264
+ and a real framework's output finds immediately -- the reason this example
265
+ was worth the extra `torch` install.
266
+
267
+ It also exposed a false positive, now fixed: the first version reported an
268
+ `lr_schedule` "LR discontinuity" at steps 5-10, because `logging_steps=5`
269
+ logs the 20-step warmup so coarsely that each logged interval covers 25% of
270
+ the run's LR range. The discontinuity check now reports a jump only when it
271
+ stands out from the logged intervals around it (a neighbor moving the same
272
+ way at no less than half the per-step rate makes it one leg of a steady
273
+ ramp), and says how many steps the interval spans. On this run trainspotter
274
+ reports nothing, which
275
+ `tests/test_detector_lr_schedule.py::test_the_real_transformers_run_has_no_lr_finding`
276
+ pins.
277
+
278
+ ## Accuracy and limitations
279
+
280
+ - Every threshold in this README and in the detector table is a default,
281
+ not a law of nature -- they're tuned to be reasonable on the example
282
+ runs above, not validated against a large corpus of real training
283
+ incidents. Expect to adjust `window`, `threshold`, `alpha`, etc. for your
284
+ own runs' scale and logging frequency.
285
+ - **Readers are tested against real files**, not remembered schemas: every
286
+ reader in `tests/test_readers.py` is exercised against a file this repo's
287
+ test suite writes to disk in the documented format. The TensorBoard
288
+ reader is tested against a real event file written by TensorBoard's own
289
+ `EventFileWriter`.
290
+ - **Detectors are tested against synthetic signals with known ground
291
+ truth**: a spike planted at step *k* must be found at step *k*; a clean
292
+ curve must produce zero findings. See `tests/test_detector_*.py`. A
293
+ separate regression test (`tests/test_examples_regression.py`) runs the
294
+ whole detector set against all five real example logs and asserts the
295
+ healthy baseline stays quiet while the other four still flag what
296
+ they're supposed to.
297
+ - No detector here is a substitute for understanding your training run.
298
+ They're heuristics built to reduce how often you have to stare at a
299
+ chart by eye, not a certified diagnosis -- read the false-positive modes
300
+ in the [detector table](#detectors) before trusting a finding blindly.
301
+ - `trainspotter watch` re-reads the whole file on every poll (default
302
+ every 2s) rather than tailing new bytes; fine for the log sizes this
303
+ tool targets (thousands to tens of thousands of steps), not designed for
304
+ gigabyte-scale logs.
305
+ - The OLS significance tests (`plateau`, `overfitting`, `divergence`'s
306
+ sustained-growth check) use a normal approximation to the t-distribution,
307
+ accurate for windows of about 30+ points; on shorter windows they're
308
+ slightly anti-conservative, which is why each detector also requires a
309
+ minimum window size before trusting the test.
310
+
311
+ ## Development
312
+
313
+ ```bash
314
+ uv sync --all-extras --dev
315
+ uv run pytest # 64 tests
316
+ uv run ruff check src tests examples
317
+ uv run mypy src
318
+ ```
319
+
320
+ See [CONTRIBUTING.md](https://github.com/antonsoo/trainspotter/blob/main/CONTRIBUTING.md) for the detector-contribution
321
+ pattern.
322
+
323
+ ## Contributing
324
+
325
+ Issues and PRs are welcome -- see [CONTRIBUTING.md](https://github.com/antonsoo/trainspotter/blob/main/CONTRIBUTING.md).
326
+
327
+ ## License
328
+
329
+ [MIT](https://github.com/antonsoo/trainspotter/blob/main/LICENSE) &copy; 2026 Anton Soloviev
330
+
331
+ ---
332
+
333
+ <sub>Part of [Officina](https://antonsoo.github.io/officina/), a set of small open-source tools by [Anton Soloviev](https://github.com/antonsoo).</sub>