trainspotter 0.1.2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- trainspotter-0.1.2/.gitignore +16 -0
- trainspotter-0.1.2/CHANGELOG.md +53 -0
- trainspotter-0.1.2/LICENSE +21 -0
- trainspotter-0.1.2/PKG-INFO +333 -0
- trainspotter-0.1.2/README.md +297 -0
- trainspotter-0.1.2/pyproject.toml +118 -0
- trainspotter-0.1.2/src/trainspotter/__init__.py +20 -0
- trainspotter-0.1.2/src/trainspotter/cli.py +133 -0
- trainspotter-0.1.2/src/trainspotter/detectors/__init__.py +106 -0
- trainspotter-0.1.2/src/trainspotter/detectors/base.py +147 -0
- trainspotter-0.1.2/src/trainspotter/detectors/divergence.py +144 -0
- trainspotter-0.1.2/src/trainspotter/detectors/eval_noise.py +81 -0
- trainspotter-0.1.2/src/trainspotter/detectors/grad_norm.py +145 -0
- trainspotter-0.1.2/src/trainspotter/detectors/loss_floor.py +104 -0
- trainspotter-0.1.2/src/trainspotter/detectors/lr_schedule.py +202 -0
- trainspotter-0.1.2/src/trainspotter/detectors/overfitting.py +125 -0
- trainspotter-0.1.2/src/trainspotter/detectors/plateau.py +176 -0
- trainspotter-0.1.2/src/trainspotter/detectors/spikes.py +91 -0
- trainspotter-0.1.2/src/trainspotter/detectors/throughput.py +136 -0
- trainspotter-0.1.2/src/trainspotter/model.py +68 -0
- trainspotter-0.1.2/src/trainspotter/py.typed +0 -0
- trainspotter-0.1.2/src/trainspotter/readers/__init__.py +82 -0
- trainspotter-0.1.2/src/trainspotter/readers/common.py +66 -0
- trainspotter-0.1.2/src/trainspotter/readers/hf.py +82 -0
- trainspotter-0.1.2/src/trainspotter/readers/lightning.py +22 -0
- trainspotter-0.1.2/src/trainspotter/readers/tabular.py +91 -0
- trainspotter-0.1.2/src/trainspotter/readers/tensorboard.py +39 -0
- trainspotter-0.1.2/src/trainspotter/readers/wandb.py +50 -0
- trainspotter-0.1.2/src/trainspotter/report/__init__.py +9 -0
- trainspotter-0.1.2/src/trainspotter/report/html_report.py +599 -0
- trainspotter-0.1.2/src/trainspotter/report/json_report.py +71 -0
- trainspotter-0.1.2/src/trainspotter/report/terminal.py +84 -0
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented in this file.
|
|
4
|
+
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
5
|
+
and this project uses [Semantic Versioning](https://semver.org/).
|
|
6
|
+
|
|
7
|
+
## [0.1.2] - 2026-10-01
|
|
8
|
+
|
|
9
|
+
### Added
|
|
10
|
+
|
|
11
|
+
- Published to PyPI: `pip install trainspotter`. The README's images and links
|
|
12
|
+
are rewritten to absolute URLs at build time so they work on the project
|
|
13
|
+
page.
|
|
14
|
+
- `trainspotter --version`.
|
|
15
|
+
|
|
16
|
+
## [0.1.1] - 2026-09-30
|
|
17
|
+
|
|
18
|
+
### Fixed
|
|
19
|
+
|
|
20
|
+
- `lr_schedule` reported a short warmup as an "LR discontinuity" when the log
|
|
21
|
+
was coarse (the real `transformers.Trainer` example: a 20-step warmup
|
|
22
|
+
logged every 5 steps, so each logged interval covers 25% of the LR range).
|
|
23
|
+
A jump is now reported only if it stands out from the logged intervals
|
|
24
|
+
around it, and the message gives the step gap instead of always saying
|
|
25
|
+
"in a single step".
|
|
26
|
+
- Slope p-values print as `p<0.0001` instead of `p=0` when the normal
|
|
27
|
+
approximation underflows.
|
|
28
|
+
- An enormous rise reads as "300x its minimum" instead of "29891.0%", and a
|
|
29
|
+
one-step span reads "step 320" instead of "steps 320-320".
|
|
30
|
+
- HTML report: rounding the axis ends outward drew a gridline and tick label
|
|
31
|
+
above the chart frame (the divergence example's loss chart showed a "5"
|
|
32
|
+
over its legend); ticks now stay inside the plotted range. Off-scale value
|
|
33
|
+
labels at the right edge sit beside their markers instead of on the frame.
|
|
34
|
+
|
|
35
|
+
## [0.1.0] - 2026-09-24
|
|
36
|
+
|
|
37
|
+
### Added
|
|
38
|
+
|
|
39
|
+
- Readers for Hugging Face `trainer_state.json`, generic CSV/JSONL metric
|
|
40
|
+
logs, Weights & Biases history CSV exports, PyTorch Lightning
|
|
41
|
+
`CSVLogger` `metrics.csv`, and TensorBoard event files (optional
|
|
42
|
+
`tensorboard` extra).
|
|
43
|
+
- Nine detectors: loss spikes, divergence (NaN/Inf and sustained growth),
|
|
44
|
+
plateaus, overfitting onset, LR schedule anomalies, gradient-norm
|
|
45
|
+
explosions/clipping saturation, throughput regressions, eval-metric
|
|
46
|
+
noise, and a loss-floor/leakage heuristic.
|
|
47
|
+
- `trainspotter analyze`: terminal, JSON, and self-contained HTML reports;
|
|
48
|
+
`--fail-on warning|error` for CI gating.
|
|
49
|
+
- `trainspotter watch`: tails a growing log file and prints new findings
|
|
50
|
+
as they appear.
|
|
51
|
+
- Real example logs under `examples/`, generated by training a small
|
|
52
|
+
numpy MLP on scikit-learn's digits dataset with deliberately induced
|
|
53
|
+
pathologies (`examples/generate_examples.py`).
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Anton Soloviev
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,333 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: trainspotter
|
|
3
|
+
Version: 0.1.2
|
|
4
|
+
Summary: Automatic diagnosis of loss curves and training logs: spot divergence, spikes, overfitting, LR and throughput problems before you burn another GPU-day.
|
|
5
|
+
Project-URL: Homepage, https://antonsoo.github.io/trainspotter/
|
|
6
|
+
Project-URL: Repository, https://github.com/antonsoo/trainspotter
|
|
7
|
+
Project-URL: Issues, https://github.com/antonsoo/trainspotter/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/antonsoo/trainspotter/blob/main/CHANGELOG.md
|
|
9
|
+
Author-email: Anton Soloviev <anton@praviel.com>
|
|
10
|
+
License: MIT
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Keywords: cli,diagnostics,huggingface,loss-curve,machine-learning,mlops,pytorch,training
|
|
13
|
+
Classifier: Development Status :: 4 - Beta
|
|
14
|
+
Classifier: Intended Audience :: Science/Research
|
|
15
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
21
|
+
Classifier: Typing :: Typed
|
|
22
|
+
Requires-Python: >=3.10
|
|
23
|
+
Requires-Dist: numpy>=1.24
|
|
24
|
+
Provides-Extra: dev
|
|
25
|
+
Requires-Dist: mypy>=1.10; extra == 'dev'
|
|
26
|
+
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
|
|
27
|
+
Requires-Dist: pytest>=8.0; extra == 'dev'
|
|
28
|
+
Requires-Dist: ruff>=0.6; extra == 'dev'
|
|
29
|
+
Requires-Dist: scikit-learn>=1.3; extra == 'dev'
|
|
30
|
+
Requires-Dist: tbparse>=0.0.8; extra == 'dev'
|
|
31
|
+
Provides-Extra: examples
|
|
32
|
+
Requires-Dist: scikit-learn>=1.3; extra == 'examples'
|
|
33
|
+
Provides-Extra: tensorboard
|
|
34
|
+
Requires-Dist: tbparse>=0.0.8; extra == 'tensorboard'
|
|
35
|
+
Description-Content-Type: text/markdown
|
|
36
|
+
|
|
37
|
+
# trainspotter
|
|
38
|
+
|
|
39
|
+
**Spot what went wrong in a training run before you burn another GPU-day.**
|
|
40
|
+
Automatic diagnosis of loss curves and training logs.
|
|
41
|
+
|
|
42
|
+
[](https://github.com/antonsoo/trainspotter/blob/main/LICENSE)
|
|
43
|
+

|
|
44
|
+
[](https://antonsoo.github.io/trainspotter/)
|
|
45
|
+
[](https://huggingface.co/spaces/antonsoloviev/trainspotter)
|
|
46
|
+
|
|
47
|
+
Every ML engineer has stared at a TensorBoard or W&B chart trying to answer
|
|
48
|
+
"is this run okay?" -- and caught the answer late: a spike that should have
|
|
49
|
+
triggered a restart three hours ago, an eval curve that's been overfitting
|
|
50
|
+
since epoch 4, a learning-rate schedule that never actually warmed up. The
|
|
51
|
+
signal was in the log the whole time. trainspotter reads the log a run
|
|
52
|
+
already writes and turns eyeballing into a report: which pathologies were
|
|
53
|
+
detected, where (step ranges), the evidence behind each, and the usual
|
|
54
|
+
fixes. It runs post-hoc on a finished log, live via `watch` on a growing
|
|
55
|
+
one, or as a CI gate that fails the build.
|
|
56
|
+
|
|
57
|
+
## Report
|
|
58
|
+
|
|
59
|
+

|
|
60
|
+
|
|
61
|
+
*(top of the report -- [full report, all 5 charts + all 6 findings](https://github.com/antonsoo/trainspotter/blob/main/docs/assets/html-report-divergence-full.png))*
|
|
62
|
+
|
|
63
|
+

|
|
64
|
+
|
|
65
|
+
Browse every example's full HTML report at
|
|
66
|
+
**[antonsoo.github.io/trainspotter](https://antonsoo.github.io/trainspotter/)**
|
|
67
|
+
(built by `scripts/build_site.py`). Both images above are real output from `examples/divergence.trainer_state.json` (see
|
|
68
|
+
[Real demo data](#real-demo-data) for exactly how that log was produced).
|
|
69
|
+
|
|
70
|
+
## Quickstart
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
pip install trainspotter
|
|
74
|
+
git clone --depth 1 https://github.com/antonsoo/trainspotter && cd trainspotter
|
|
75
|
+
trainspotter analyze examples/divergence.trainer_state.json
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
Or point it at your own log:
|
|
79
|
+
|
|
80
|
+
```bash
|
|
81
|
+
trainspotter analyze path/to/trainer_state.json --output html --out report.html
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
## Features
|
|
85
|
+
|
|
86
|
+
- **Five readers**, auto-detected from the file: Hugging Face
|
|
87
|
+
`trainer_state.json`, generic CSV/JSONL, Weights & Biases history CSV
|
|
88
|
+
exports, PyTorch Lightning `CSVLogger` `metrics.csv`, and TensorBoard
|
|
89
|
+
event files (`pip install trainspotter[tensorboard]`).
|
|
90
|
+
- **Nine detectors** -- see the [table below](#detectors) -- each with a
|
|
91
|
+
documented algorithm, tunable thresholds, and stated false-positive modes.
|
|
92
|
+
- **Three report formats**: a color-coded terminal report, JSON (stable
|
|
93
|
+
schema, for tooling), and a self-contained HTML report (inline SVG
|
|
94
|
+
charts, no CDN, no JavaScript -- opens from disk, works offline forever).
|
|
95
|
+
- **CI gating**: `--fail-on warning|error` exits non-zero if any finding at
|
|
96
|
+
or above that severity was found.
|
|
97
|
+
- **Live tailing**: `trainspotter watch <file>` re-reads a growing log on
|
|
98
|
+
an interval and prints only newly-appeared findings.
|
|
99
|
+
- **numpy is the only required dependency.** No pandas, no plotting
|
|
100
|
+
library, no web framework. (`tensorboard` support is an optional extra
|
|
101
|
+
since parsing the event-file format properly needs `tbparse`.)
|
|
102
|
+
|
|
103
|
+
## Usage
|
|
104
|
+
|
|
105
|
+
```bash
|
|
106
|
+
# Terminal report (default), with a CI-style exit code
|
|
107
|
+
trainspotter analyze run/trainer_state.json --fail-on error
|
|
108
|
+
|
|
109
|
+
# Machine-readable JSON
|
|
110
|
+
trainspotter analyze run/trainer_state.json --output json > findings.json
|
|
111
|
+
|
|
112
|
+
# Self-contained HTML report
|
|
113
|
+
trainspotter analyze run/trainer_state.json --output html --out report.html
|
|
114
|
+
|
|
115
|
+
# Force a reader instead of auto-detecting from the extension
|
|
116
|
+
trainspotter analyze run/metrics.csv --format lightning
|
|
117
|
+
|
|
118
|
+
# Tail a log that's still being written
|
|
119
|
+
trainspotter watch run/trainer_state.json --interval 5 --fail-on error
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
Real terminal output (from `examples/overfitting.trainer_state.json`, one
|
|
123
|
+
finding shown):
|
|
124
|
+
|
|
125
|
+
```
|
|
126
|
+
[WARN ] steps 125-799 overfitting Overfitting onset
|
|
127
|
+
Best eval/loss was 0.526 at step 125; it's since risen to 0.5554 (5.6%, slope p<0.0001).
|
|
128
|
+
Meanwhile train/loss kept falling over the same range (slope < 0) -- the classic
|
|
129
|
+
overfitting signature.
|
|
130
|
+
metric: eval/loss
|
|
131
|
+
fix: Use the checkpoint at step 125 (best eval/loss), not the last one.
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
### CI usage
|
|
135
|
+
|
|
136
|
+
```yaml
|
|
137
|
+
- name: Diagnose the training run
|
|
138
|
+
run: trainspotter analyze runs/latest/trainer_state.json --fail-on error
|
|
139
|
+
# exits 1 if any error-severity pathology was detected
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
## Detectors
|
|
143
|
+
|
|
144
|
+
Every detector's full algorithm and false-positive list lives in its
|
|
145
|
+
module's docstring (`src/trainspotter/detectors/*.py`); this table is the
|
|
146
|
+
summary.
|
|
147
|
+
|
|
148
|
+
| Detector | Detects | How | Stated false-positive modes |
|
|
149
|
+
|---|---|---|---|
|
|
150
|
+
| `spikes` | A sudden jump in a loss metric | Modified z-score (`0.6745·(x−median)/MAD`) against a trailing rolling median/MAD; flags \|z\| > 6 | Deliberate LR restarts/curriculum jumps; bumpy metrics with a too-small window |
|
|
151
|
+
| `divergence` | NaN/Inf, or sustained upward drift | Any non-finite value; or an OLS slope > 0 with p < 0.05 over the tail window, and the last value ≥ 1.5× the running minimum | Metrics meant to increase (only checks loss-type metrics by default); a temporary cyclic-schedule upswing |
|
|
152
|
+
| `plateau` | A metric that's stalled, not just converged | Sliding-window OLS slope test (p ≥ 0.2 = not significant) plus a < 2% relative-change guard; suppressed if the window has already recovered ≥ 80% of the metric's total drop, or `lr` has decayed to ≤ 30% of its peak | A real stall near a coincidentally low value, or right as an unrelated LR decay finishes, can be wrongly suppressed by the convergence check |
|
|
153
|
+
| `overfitting` | Train still improving while eval worsens | Finds eval's running-minimum step; tests whether the slope after it is significantly positive (p < 0.1) and ≥ 1% above the minimum | A noisy eval set can show a false uptick; a mid-run change in eval data |
|
|
154
|
+
| `lr_schedule` | Missing warmup, LR rising after its peak, discontinuities | Ramp check on the first value vs. peak; running-max monotonicity after the peak; a jump of ≥ 20% of the LR's total range that isn't one leg of a steady ramp | Cyclic/warm-restart schedules trip the last two by design; step-decay drops are reported as `info` |
|
|
155
|
+
| `grad_norm` | Gradient-norm explosions and clipping saturation | Same z-score as `spikes` (one-sided) for explosions, threshold 10 (higher than `spikes`' 6 -- a single z of 6-8 is common early-training noise); fraction of a trailing window within 0.5% of its own max, for saturation | Norms logged post-clip are saturated by construction; an isolated explosion with no coinciding loss spike is weaker evidence than one that lines up with a `spikes` finding |
|
|
156
|
+
| `throughput` | Step-time / throughput regressions | Median step time, early-run baseline window vs. recent window, flags ≥ 1.4× | Checkpoint/eval steps inflate single points; early-run kernel/dataloader warmup can pollute the baseline |
|
|
157
|
+
| `eval_noise` | An eval metric too noisy to rank checkpoints by | Median absolute step-to-step jitter as a fraction of the run's total improvement; flags ≥ 15% | An already-converged run has small total improvement by construction |
|
|
158
|
+
| `loss_floor` | Implausibly low train loss very early (heuristic) | Loss < 0.05 absolute **and** < 5% of its starting value within the first 10% of steps | Explicitly a heuristic: an easy task, a small-scale loss, or a resumed checkpoint all look identical to this |
|
|
159
|
+
|
|
160
|
+
### Trust the healthy case
|
|
161
|
+
|
|
162
|
+
A tool that flags an unremarkable run gets ignored, and recall on real
|
|
163
|
+
incidents matters less than that first impression -- so trainspotter is
|
|
164
|
+
tested for silence, not just for alarms. **On the healthy baseline run,
|
|
165
|
+
trainspotter reports no warnings**, and `tests/test_examples_regression.py`
|
|
166
|
+
asserts this on every commit, alongside asserting that the other four
|
|
167
|
+
examples still flag the pathology they were built to demonstrate. Findings
|
|
168
|
+
are also ordered errors-first, then warnings, then info (chronological
|
|
169
|
+
within a level) -- so on a run that *did* break, the thing that broke
|
|
170
|
+
leads the report instead of sitting below routine early-training noise.
|
|
171
|
+
|
|
172
|
+
## How it works
|
|
173
|
+
|
|
174
|
+
**Data model.** Every reader converges on one shape
|
|
175
|
+
(`trainspotter.model.Run`): a dict of named metric series, each a list of
|
|
176
|
+
`(step, value, wall_time)` points. Detectors and reports only ever see
|
|
177
|
+
this shape -- they have no idea whether the log came from `transformers`,
|
|
178
|
+
a CSV, or TensorBoard. Readers normalize common spellings (`loss` /
|
|
179
|
+
`train_loss` / `training_loss` all become `train/loss`; `learning_rate` /
|
|
180
|
+
`lr` become `lr`) but pass anything unrecognized straight through, so a
|
|
181
|
+
custom metric like `eval/bleu` still reaches the detectors under its own
|
|
182
|
+
name.
|
|
183
|
+
|
|
184
|
+
**Robust statistics, not raw thresholds.** Spikes and gradient-norm
|
|
185
|
+
explosions are judged by a *modified z-score* against a trailing rolling
|
|
186
|
+
median and MAD (median absolute deviation), not a mean/stddev: a mean and
|
|
187
|
+
stddev are themselves dragged around by the very outlier you're trying to
|
|
188
|
+
detect, while the median and MAD barely move. This is the same construction
|
|
189
|
+
Iglewicz & Hoaglin describe for outlier labeling. Plateau and overfitting
|
|
190
|
+
use an ordinary-least-squares slope with a normal-approximation t-test
|
|
191
|
+
(`math.erf`-based, no `scipy` dependency) rather than eyeballing "did it go
|
|
192
|
+
up or down."
|
|
193
|
+
|
|
194
|
+
**Charts don't let one spike flatten the rest of the curve.** The HTML
|
|
195
|
+
report's axis logic (`report/html_report.py`) picks tick spacing with
|
|
196
|
+
Heckbert's "nice numbers" algorithm (`Graphics Gems`, 1990 -- 1/2/5x10^n
|
|
197
|
+
steps, not whatever an even split of the range happens to produce) and
|
|
198
|
+
floors the axis at 0 for the non-negative metrics trainspotter charts
|
|
199
|
+
(loss, LR, grad_norm, accuracy, step time), instead of padding below zero.
|
|
200
|
+
If the max is more than 1.5x the 99th percentile, the axis caps at that
|
|
201
|
+
instead of stretching to fit one outlier -- the point is still drawn,
|
|
202
|
+
clamped to the top edge with a small triangle and its real value labeled
|
|
203
|
+
next to it, not hidden.
|
|
204
|
+
|
|
205
|
+
**Every finding is honest about its own limits.** Each detector's
|
|
206
|
+
docstring states its false-positive modes in plain language (see the table
|
|
207
|
+
above, or the source for the full version) -- this isn't boilerplate, it's
|
|
208
|
+
meant to be read before you act on a finding. `loss_floor` is explicitly
|
|
209
|
+
labeled a heuristic, not a statistical test, because there's no
|
|
210
|
+
distribution-free way to know a loss is "too low" without knowing the task.
|
|
211
|
+
|
|
212
|
+
## Real demo data
|
|
213
|
+
|
|
214
|
+
Nothing in `examples/` is fabricated. All five logs were produced by
|
|
215
|
+
`examples/generate_examples.py`, which trains a small hand-written numpy
|
|
216
|
+
MLP (64 → 32 ReLU → 10, softmax cross-entropy, every gradient computed and
|
|
217
|
+
applied by hand -- no autodiff framework) on scikit-learn's `load_digits`
|
|
218
|
+
dataset (1797 real 8×8 handwritten-digit images, bundled with
|
|
219
|
+
scikit-learn, no download). It's deterministic (seeded) and takes about
|
|
220
|
+
three seconds to regenerate on this machine:
|
|
221
|
+
|
|
222
|
+
```bash
|
|
223
|
+
uv sync --extra examples # or: pip install scikit-learn
|
|
224
|
+
uv run python examples/generate_examples.py
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
| File | What it is | How the pathology was actually induced |
|
|
228
|
+
|---|---|---|
|
|
229
|
+
| `baseline` | A normal, unremarkable run | Full dataset, weight decay, cosine LR with warmup. **On this run, trainspotter reports no warnings** -- see [Trust the healthy case](#trust-the-healthy-case) for why that's asserted by a test, not just eyeballed |
|
|
230
|
+
| `divergence` | A genuine crash after real training | 320 steps (80% of the run) of ordinary warmup + cosine-decay training -- loss ~2.8 -> ~0.15, eval accuracy up to ~96% -- then a simulated incident at step 320: the LR schedule jumps to 30 **and** the loss function switches to a version with the classic missing-max-subtraction softmax bug. Plain high LR alone was tried first and *didn't* produce real divergence -- see the note in `divergence_run()`'s docstring for why (softmax cross-entropy's gradient is bounded by construction) |
|
|
231
|
+
| `overfitting` | Genuine overfitting | Only 4 training examples per class (40 total), no regularization, 800 steps -- the model memorizes the training set while held-out eval loss turns upward after step 125 |
|
|
232
|
+
| `missing_warmup` | No LR ramp-up | LR set to a constant 0.5 from step 0, no warmup phase at all |
|
|
233
|
+
| `throughput_drop` | A real, measured slowdown | Steps 150-259 insert an actual `time.sleep(0.02)` per step (simulating e.g. I/O contention) -- `step_time` in the log is genuinely measured wall-clock time, not a fabricated number |
|
|
234
|
+
|
|
235
|
+
Each file is committed in both HF `trainer_state.json` format
|
|
236
|
+
(`<name>.trainer_state.json`) and generic CSV (`<name>.csv`), written from
|
|
237
|
+
the same in-memory log so the two are guaranteed consistent.
|
|
238
|
+
|
|
239
|
+
### A real `transformers.Trainer` run
|
|
240
|
+
|
|
241
|
+
`examples/transformers_tinygpt2.trainer_state.json` is a genuine
|
|
242
|
+
`trainer_state.json` written by `transformers.Trainer` itself (not the
|
|
243
|
+
numpy MLP above): a real 2-layer, 64-dim GPT-2 (`GPT2LMHeadModel` from a
|
|
244
|
+
from-scratch `GPT2Config`, ~330K parameters), trained for 6 epochs / 726
|
|
245
|
+
steps on CPU on a small hand-written toy corpus (short sentences about
|
|
246
|
+
training, tokenized with the real `gpt2` tokenizer), in about 90 seconds
|
|
247
|
+
on this machine. Regenerate it with
|
|
248
|
+
`examples/generate_transformers_example.py` (needs `torch` + `transformers`
|
|
249
|
+
in a separate environment -- see the script's docstring for the exact
|
|
250
|
+
install commands; they are *not* project dependencies).
|
|
251
|
+
|
|
252
|
+
Running trainspotter against it caught a real bug in the HF reader during
|
|
253
|
+
development: `Trainer.train()` appends one extra `log_history` entry after
|
|
254
|
+
training ends -- a run summary with `train_runtime`, `train_samples_per_second`,
|
|
255
|
+
and a `train_loss` key that is the *average* loss over the whole run, not a
|
|
256
|
+
per-step reading. The reader's first version merged that average straight
|
|
257
|
+
into the `train/loss` series (since `train_loss` is one of the recognized
|
|
258
|
+
spellings of `loss`), which showed up as a fake spike on the run's last
|
|
259
|
+
step. The fix -- detect that entry by its unique `train_runtime` key and
|
|
260
|
+
route it to run metadata instead of a metric point -- is
|
|
261
|
+
`src/trainspotter/readers/hf.py`'s `_SUMMARY_MARKER_KEY`, and
|
|
262
|
+
`tests/test_readers.py::test_hf_trainer_state_excludes_the_final_run_summary_entry`
|
|
263
|
+
pins it. This is exactly the kind of gap a hand-written fixture can miss
|
|
264
|
+
and a real framework's output finds immediately -- the reason this example
|
|
265
|
+
was worth the extra `torch` install.
|
|
266
|
+
|
|
267
|
+
It also exposed a false positive, now fixed: the first version reported an
|
|
268
|
+
`lr_schedule` "LR discontinuity" at steps 5-10, because `logging_steps=5`
|
|
269
|
+
logs the 20-step warmup so coarsely that each logged interval covers 25% of
|
|
270
|
+
the run's LR range. The discontinuity check now reports a jump only when it
|
|
271
|
+
stands out from the logged intervals around it (a neighbor moving the same
|
|
272
|
+
way at no less than half the per-step rate makes it one leg of a steady
|
|
273
|
+
ramp), and says how many steps the interval spans. On this run trainspotter
|
|
274
|
+
reports nothing, which
|
|
275
|
+
`tests/test_detector_lr_schedule.py::test_the_real_transformers_run_has_no_lr_finding`
|
|
276
|
+
pins.
|
|
277
|
+
|
|
278
|
+
## Accuracy and limitations
|
|
279
|
+
|
|
280
|
+
- Every threshold in this README and in the detector table is a default,
|
|
281
|
+
not a law of nature -- they're tuned to be reasonable on the example
|
|
282
|
+
runs above, not validated against a large corpus of real training
|
|
283
|
+
incidents. Expect to adjust `window`, `threshold`, `alpha`, etc. for your
|
|
284
|
+
own runs' scale and logging frequency.
|
|
285
|
+
- **Readers are tested against real files**, not remembered schemas: every
|
|
286
|
+
reader in `tests/test_readers.py` is exercised against a file this repo's
|
|
287
|
+
test suite writes to disk in the documented format. The TensorBoard
|
|
288
|
+
reader is tested against a real event file written by TensorBoard's own
|
|
289
|
+
`EventFileWriter`.
|
|
290
|
+
- **Detectors are tested against synthetic signals with known ground
|
|
291
|
+
truth**: a spike planted at step *k* must be found at step *k*; a clean
|
|
292
|
+
curve must produce zero findings. See `tests/test_detector_*.py`. A
|
|
293
|
+
separate regression test (`tests/test_examples_regression.py`) runs the
|
|
294
|
+
whole detector set against all five real example logs and asserts the
|
|
295
|
+
healthy baseline stays quiet while the other four still flag what
|
|
296
|
+
they're supposed to.
|
|
297
|
+
- No detector here is a substitute for understanding your training run.
|
|
298
|
+
They're heuristics built to reduce how often you have to stare at a
|
|
299
|
+
chart by eye, not a certified diagnosis -- read the false-positive modes
|
|
300
|
+
in the [detector table](#detectors) before trusting a finding blindly.
|
|
301
|
+
- `trainspotter watch` re-reads the whole file on every poll (default
|
|
302
|
+
every 2s) rather than tailing new bytes; fine for the log sizes this
|
|
303
|
+
tool targets (thousands to tens of thousands of steps), not designed for
|
|
304
|
+
gigabyte-scale logs.
|
|
305
|
+
- The OLS significance tests (`plateau`, `overfitting`, `divergence`'s
|
|
306
|
+
sustained-growth check) use a normal approximation to the t-distribution,
|
|
307
|
+
accurate for windows of about 30+ points; on shorter windows they're
|
|
308
|
+
slightly anti-conservative, which is why each detector also requires a
|
|
309
|
+
minimum window size before trusting the test.
|
|
310
|
+
|
|
311
|
+
## Development
|
|
312
|
+
|
|
313
|
+
```bash
|
|
314
|
+
uv sync --all-extras --dev
|
|
315
|
+
uv run pytest # 64 tests
|
|
316
|
+
uv run ruff check src tests examples
|
|
317
|
+
uv run mypy src
|
|
318
|
+
```
|
|
319
|
+
|
|
320
|
+
See [CONTRIBUTING.md](https://github.com/antonsoo/trainspotter/blob/main/CONTRIBUTING.md) for the detector-contribution
|
|
321
|
+
pattern.
|
|
322
|
+
|
|
323
|
+
## Contributing
|
|
324
|
+
|
|
325
|
+
Issues and PRs are welcome -- see [CONTRIBUTING.md](https://github.com/antonsoo/trainspotter/blob/main/CONTRIBUTING.md).
|
|
326
|
+
|
|
327
|
+
## License
|
|
328
|
+
|
|
329
|
+
[MIT](https://github.com/antonsoo/trainspotter/blob/main/LICENSE) © 2026 Anton Soloviev
|
|
330
|
+
|
|
331
|
+
---
|
|
332
|
+
|
|
333
|
+
<sub>Part of [Officina](https://antonsoo.github.io/officina/), a set of small open-source tools by [Anton Soloviev](https://github.com/antonsoo).</sub>
|