trainmeter 0.0.2__tar.gz → 0.0.4__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (100) hide show
  1. {trainmeter-0.0.2 → trainmeter-0.0.4}/CLAUDE.md +4 -3
  2. {trainmeter-0.0.2 → trainmeter-0.0.4}/PKG-INFO +10 -3
  3. {trainmeter-0.0.2 → trainmeter-0.0.4}/README.md +9 -2
  4. {trainmeter-0.0.2 → trainmeter-0.0.4}/docs/architecture.md +44 -33
  5. {trainmeter-0.0.2 → trainmeter-0.0.4}/docs/hardware-checklist.md +21 -3
  6. trainmeter-0.0.4/docs/images/dashboard-dark.png +0 -0
  7. trainmeter-0.0.4/docs/images/dashboard-light.png +0 -0
  8. trainmeter-0.0.4/docs/roadmap.md +136 -0
  9. {trainmeter-0.0.2 → trainmeter-0.0.4}/docs/spec.md +28 -16
  10. trainmeter-0.0.4/docs/users.md +138 -0
  11. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/__init__.py +3 -3
  12. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/doctor.py +35 -17
  13. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/engine.py +40 -10
  14. trainmeter-0.0.4/src/trainmeter/hardware/__init__.py +138 -0
  15. trainmeter-0.0.4/src/trainmeter/hardware/specs/README.md +41 -0
  16. trainmeter-0.0.4/src/trainmeter/hardware/specs/a100_sxm_80gb.json +19 -0
  17. trainmeter-0.0.4/src/trainmeter/hardware/specs/b200.json +22 -0
  18. trainmeter-0.0.4/src/trainmeter/hardware/specs/gb10.json +20 -0
  19. trainmeter-0.0.4/src/trainmeter/hardware/specs/gb200.json +19 -0
  20. trainmeter-0.0.4/src/trainmeter/hardware/specs/h100_pcie.json +19 -0
  21. trainmeter-0.0.4/src/trainmeter/hardware/specs/h100_sxm.json +20 -0
  22. trainmeter-0.0.4/src/trainmeter/hardware/specs/h200.json +23 -0
  23. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/passport.py +15 -3
  24. trainmeter-0.0.4/src/trainmeter/sources/gpu/__init__.py +56 -0
  25. trainmeter-0.0.2/src/trainmeter/sources/gpu.py → trainmeter-0.0.4/src/trainmeter/sources/gpu/backend.py +7 -172
  26. trainmeter-0.0.4/src/trainmeter/sources/gpu/platform.py +79 -0
  27. trainmeter-0.0.4/src/trainmeter/sources/gpu/sampler.py +134 -0
  28. trainmeter-0.0.4/src/trainmeter/sources/gpu/selection.py +58 -0
  29. trainmeter-0.0.4/src/trainmeter/sources/gpu/server.py +40 -0
  30. trainmeter-0.0.4/src/trainmeter/sources/gpu/spark.py +79 -0
  31. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/web/server.py +4 -1
  32. trainmeter-0.0.4/src/trainmeter/web/static/app.js +726 -0
  33. trainmeter-0.0.4/src/trainmeter/web/static/favicon.svg +1 -0
  34. trainmeter-0.0.4/src/trainmeter/web/static/index.html +65 -0
  35. trainmeter-0.0.4/src/trainmeter/web/static/style.css +369 -0
  36. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_dashboard.py +56 -51
  37. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_engine.py +42 -0
  38. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_gpu.py +86 -0
  39. trainmeter-0.0.4/tests/test_hardware.py +105 -0
  40. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_passport.py +22 -0
  41. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_web.py +4 -0
  42. trainmeter-0.0.2/src/trainmeter/peaks.py +0 -92
  43. trainmeter-0.0.2/src/trainmeter/web/static/app.js +0 -571
  44. trainmeter-0.0.2/src/trainmeter/web/static/index.html +0 -43
  45. trainmeter-0.0.2/src/trainmeter/web/static/style.css +0 -157
  46. trainmeter-0.0.2/tests/test_peaks.py +0 -34
  47. {trainmeter-0.0.2 → trainmeter-0.0.4}/.github/workflows/ci.yml +0 -0
  48. {trainmeter-0.0.2 → trainmeter-0.0.4}/.github/workflows/release.yml +0 -0
  49. {trainmeter-0.0.2 → trainmeter-0.0.4}/.gitignore +0 -0
  50. {trainmeter-0.0.2 → trainmeter-0.0.4}/LICENSE +0 -0
  51. {trainmeter-0.0.2 → trainmeter-0.0.4}/examples/README.md +0 -0
  52. {trainmeter-0.0.2 → trainmeter-0.0.4}/examples/plan_run.py +0 -0
  53. {trainmeter-0.0.2 → trainmeter-0.0.4}/examples/train_tiny_gpt.py +0 -0
  54. {trainmeter-0.0.2 → trainmeter-0.0.4}/pyproject.toml +0 -0
  55. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/__main__.py +0 -0
  56. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/agent/__init__.py +0 -0
  57. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/agent/bootstrap/sitecustomize.py +0 -0
  58. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/agent/core.py +0 -0
  59. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/agent/sender.py +0 -0
  60. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/agent/taps.py +0 -0
  61. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/agent/wrap.py +0 -0
  62. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/catalog.py +0 -0
  63. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/cli.py +0 -0
  64. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/commands.py +0 -0
  65. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/config.py +0 -0
  66. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/emit.py +0 -0
  67. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/export/__init__.py +0 -0
  68. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/export/files.py +0 -0
  69. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/export/wandb.py +0 -0
  70. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/facts.py +0 -0
  71. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/flops.py +0 -0
  72. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/metrics.py +0 -0
  73. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/records.py +0 -0
  74. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/replay.py +0 -0
  75. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/report.py +0 -0
  76. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/sources/__init__.py +0 -0
  77. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/sources/host.py +0 -0
  78. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/supervisor/__init__.py +0 -0
  79. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/supervisor/ingest.py +0 -0
  80. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/supervisor/launcher.py +0 -0
  81. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/supervisor/live.py +0 -0
  82. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/supervisor/store.py +0 -0
  83. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/timeline.py +0 -0
  84. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/viewer.py +0 -0
  85. {trainmeter-0.0.2 → trainmeter-0.0.4}/src/trainmeter/web/__init__.py +0 -0
  86. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_agent.py +0 -0
  87. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_cli.py +0 -0
  88. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_config.py +0 -0
  89. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_examples.py +0 -0
  90. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_facts.py +0 -0
  91. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_flops.py +0 -0
  92. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_host.py +0 -0
  93. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_metrics.py +0 -0
  94. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_probes_real.py +0 -0
  95. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_records.py +0 -0
  96. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_release.py +0 -0
  97. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_store.py +0 -0
  98. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_taps.py +0 -0
  99. {trainmeter-0.0.2 → trainmeter-0.0.4}/tests/test_viewer.py +0 -0
  100. {trainmeter-0.0.2 → trainmeter-0.0.4}/uv.lock +0 -0
@@ -68,6 +68,7 @@ These are the kind that locally reasonable code breaks quietly.
68
68
  New entry points inside the training process get the same wrapper and a test that forces an internal error.
69
69
  2. **Peaks are dense, and every peak cites its datasheet.**
70
70
  NVIDIA headlines the 2:4 sparse figure, which is exactly twice the dense one; using it halves MFU.
71
+ A peak lives in a spec file in `hardware/specs/` next to the datasheet lines it comes from, and a figure no datasheet gives is `null` with a note, never an estimate.
71
72
  An unknown device returns `None`, and MFU is then not reported.
72
73
  Never fall back to another card's peak.
73
74
  3. **A FLOPs number always carries its convention.**
@@ -99,13 +100,13 @@ These are the kind that locally reasonable code breaks quietly.
99
100
 
100
101
  What exists today:
101
102
 
102
- - `peaks.py`, `flops.py`, `metrics.py` - peaks (dense BF16, HBM), FLOPs per token under named conventions, pure formulas (MFU, OFU, arithmetic intensity as `step_intensity`, ETA, EMA, and dollar helpers that only overlays may use).
103
+ - `hardware/` (device specs, one JSON file per GPU in `hardware/specs/` with the datasheet lines they come from), `flops.py`, `metrics.py` - dense BF16 peaks and memory bandwidth, FLOPs per token under named conventions, pure formulas (MFU, OFU, arithmetic intensity as `step_intensity`, ETA, EMA, and dollar helpers that only overlays may use).
103
104
  - `records.py` (wire schema and redaction), `facts.py` (provenance, precedence), `catalog.py` (series aliases and the metric catalog), `engine.py` (records in, snapshots out), `timeline.py` (chart history), `passport.py` (the model-card passport), `replay.py` (rebuild a run from its log), `config.py` (declared flags), `report.py` (the summary line) - the pure core.
104
- - `cli.py`, `doctor.py` (`tm doctor`), `supervisor/` (`launcher.py`, `store.py`, `ingest.py`), `sources/host.py`, `sources/gpu.py` (NVML and GPM behind a `Backend` seam, with a fake), `supervisor/live.py` (engine and timeline under one lock), `web/` (`server.py` and the static dashboard), `viewer.py` (`tm view`, `tm attach`), `commands.py` (`tm passport`, `tm export`) and `export/` (passport files, W&B) - the `tm` process.
105
+ - `cli.py`, `doctor.py` (`tm doctor`), `supervisor/` (`launcher.py`, `store.py`, `ingest.py`), `sources/host.py`, `sources/gpu/` (`backend.py`: NVML and GPM behind a `Backend` seam, with a fake; `platform.py`: the adapter interface and the row schema; `server.py` and `spark.py`: one adapter per kind of machine, named by the device's spec; `sampler.py`), `supervisor/live.py` (engine and timeline under one lock), `web/` (`server.py` and the static dashboard), `viewer.py` (`tm view`, `tm attach`), `commands.py` (`tm passport`, `tm export`) and `export/` (passport files, W&B) - the `tm` process.
105
106
  - `agent/` (`core.py`, `taps.py`, `wrap.py`, `sender.py`, `bootstrap/sitecustomize.py`) and `emit.py` - what runs inside the training process.
106
107
  - `tests/` - one test file per module; numbers for the reference 1.04B model are pinned in `test_flops.py`. Tests that need torch skip without it, and CI runs them in a second job; those that use wandb, TensorBoard or MLflow also skip without the tracker (`uv pip install torch tensorboard wandb mlflow-skinny`).
107
108
 
108
- What is planned, from `docs/architecture.md`: the remaining probes and views.
109
+ What is planned: `docs/roadmap.md`, which follows from the user analysis in `docs/users.md`.
109
110
  `docs/hardware-checklist.md` lists what only a real GPU node can verify.
110
111
 
111
112
  The dependency rule: edges import the core, and the core imports nothing but the standard library.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: trainmeter
3
- Version: 0.0.2
3
+ Version: 0.0.4
4
4
  Summary: Where your training FLOPs go and how well the GPUs are used: loss, FLOPs, MFU and hardware counters for a wrapped training run.
5
5
  Project-URL: Homepage, https://github.com/Almaz-KG/trainmeter
6
6
  Project-URL: Repository, https://github.com/Almaz-KG/trainmeter
@@ -31,13 +31,20 @@ trainmeter starts your command, watches the node from outside, and serves a live
31
31
  The dashboard shows what a training run needs to be judged: training and validation loss, FLOPs spent and required, MFU with its FLOPs convention spelled out, and how busy the GPUs really are, from `GPU-Util` down to MFU, next to arithmetic intensity.
32
32
  Your training code stays as it is.
33
33
 
34
- Status: early scaffold, version 0.0.1.
34
+ <picture>
35
+ <source media="(prefers-color-scheme: dark)" srcset="docs/images/dashboard-dark.png">
36
+ <img alt="The trainmeter dashboard: progress, throughput, MFU and OFU, goodput and the loss curve of a run on four H100s" src="docs/images/dashboard-light.png">
37
+ </picture>
38
+
39
+ <sub>The dashboard on a synthetic run.</sub>
40
+
41
+ Status: early scaffold.
35
42
  `tm` runs any command untouched, records the run, and turns what a PyTorch loop does into tokens, FLOPs spent and required, ETA, MFU (given a device) and the losses you `emit()`.
36
43
  GPU counters (NVML and GPM: the utilization ladder, OFU, observed FLOPs, arithmetic intensity) and `tm doctor` are built and tested against fakes, not yet on a real GPU.
37
44
  Series your loop already logs to wandb, TensorBoard or MLflow are picked up without a code change (`--no-tap` turns it off), and DataLoader waits, checkpoint saves and eval spans feed a "where the time goes" view.
38
45
  One real step is also counted with `FlopCounterMode` (shown beside MFU as the `counted` convention), and layers and width are read from a model config when there is one.
39
46
  Each run ends with a run passport (`passport.md` and `passport.json`) for a model card.
40
- The dashboard (Cockpit, Efficiency, Hardware, Where the time goes, Plan vs actual and Run passport) is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run.
47
+ The dashboard, one page of progress, tiles and charts with the run details and model card folded away, is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run.
41
48
  Not yet used on a real run.
42
49
  To try it, `examples/` has a tiny GPT to run under `tm` on a laptop and a run planner that needs no GPU.
43
50
 
@@ -7,13 +7,20 @@ trainmeter starts your command, watches the node from outside, and serves a live
7
7
  The dashboard shows what a training run needs to be judged: training and validation loss, FLOPs spent and required, MFU with its FLOPs convention spelled out, and how busy the GPUs really are, from `GPU-Util` down to MFU, next to arithmetic intensity.
8
8
  Your training code stays as it is.
9
9
 
10
- Status: early scaffold, version 0.0.1.
10
+ <picture>
11
+ <source media="(prefers-color-scheme: dark)" srcset="docs/images/dashboard-dark.png">
12
+ <img alt="The trainmeter dashboard: progress, throughput, MFU and OFU, goodput and the loss curve of a run on four H100s" src="docs/images/dashboard-light.png">
13
+ </picture>
14
+
15
+ <sub>The dashboard on a synthetic run.</sub>
16
+
17
+ Status: early scaffold.
11
18
  `tm` runs any command untouched, records the run, and turns what a PyTorch loop does into tokens, FLOPs spent and required, ETA, MFU (given a device) and the losses you `emit()`.
12
19
  GPU counters (NVML and GPM: the utilization ladder, OFU, observed FLOPs, arithmetic intensity) and `tm doctor` are built and tested against fakes, not yet on a real GPU.
13
20
  Series your loop already logs to wandb, TensorBoard or MLflow are picked up without a code change (`--no-tap` turns it off), and DataLoader waits, checkpoint saves and eval spans feed a "where the time goes" view.
14
21
  One real step is also counted with `FlopCounterMode` (shown beside MFU as the `counted` convention), and layers and width are read from a model config when there is one.
15
22
  Each run ends with a run passport (`passport.md` and `passport.json`) for a model card.
16
- The dashboard (Cockpit, Efficiency, Hardware, Where the time goes, Plan vs actual and Run passport) is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run.
23
+ The dashboard, one page of progress, tiles and charts with the run details and model card folded away, is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run.
17
24
  Not yet used on a real run.
18
25
  To try it, `examples/` has a tiny GPT to run under `tm` on a laptop and a run planner that needs no GPU.
19
26
 
@@ -65,7 +65,7 @@ How the invariants in `CLAUDE.md` map onto this:
65
65
  | Invariant | In this architecture |
66
66
  | --- | --- |
67
67
  | 1. Never raises into the training loop | Rule 6. Agent probes, taps and `emit()` are wrapped and self-disabling, and everything else runs in another process. Each of them gets a test that forces an internal error. |
68
- | 2. Peaks are dense and cite a datasheet | Unchanged. `peaks.py` stays in the pure core, the GPU sampler reports the device name, and an unknown device means no MFU. |
68
+ | 2. Peaks are dense and cite a datasheet | Each device has a spec file in `hardware/specs/` with its figures and the datasheet lines they come from. The GPU sampler reports the device name, and an unknown device, or one whose datasheet gives no dense BF16 figure, means no MFU. |
69
69
  | 3. A FLOPs number carries its convention | Rule 4. Value and convention are one fact and are never separated in the log, the API or the dashboard. |
70
70
  | 4. Counter metrics are not model metrics | The engine keeps modeled keys (MFU, FLOPs spent), sampled keys (OFU, observed FLOPs, DRAM active) and arithmetic intensity, which uses both, apart. The dashboard shows them side by side and labels each by origin. |
71
71
  | 5. Only rank 0 writes | Generalized by rule 5: only the supervisor writes. Rank 0 is the step clock and the other ranks feed straggler detection. |
@@ -341,42 +341,37 @@ Overlays are computed from the primitives at view time and never stored.
341
341
 
342
342
  The dashboard is one static page fed by the read-only API of section 10.
343
343
  It renders from the catalog, so a new series appears without a page change.
344
- The views, the rendering rules and the x-axis switch are specified in `docs/spec.md` (Screens).
344
+ The page, the rendering rules and the x-axis selector are specified in `docs/spec.md` (Screens).
345
345
  Plotting loss against FLOPs makes runs comparable across batch sizes and sequence lengths.
346
- Cockpit, Efficiency and Hardware arrive in slice 4, Where the time goes in slice 5, and the run passport and Plan vs actual in slice 7.
346
+ The views of slices 4, 5 and 7 were later folded into this one page without tabs.
347
347
 
348
- Layout sketches, with placeholders instead of numbers:
348
+ Layout sketch, with placeholders instead of numbers:
349
349
 
350
350
  ```text
351
351
  +----------------------------------------------------------------------+
352
- | <command> | <n> x <GPU> | FLOPs per token <value> (<convention>) |
353
- +----------------------------------+-----------------------------------+
354
- | PROGRESS | THROUGHPUT |
355
- | FLOPs spent of FLOPs required | tokens/s and step time |
356
- | tokens seen, ETA | data wait share of the step |
357
- +----------------------------------+-----------------------------------+
358
- | TRAIN LOSS and VAL LOSS | MFU and OFU |
359
- | x axis: step, tokens, FLOPs | the gap between them |
360
- | or wall time | latest utilization ladder |
361
- +----------------------------------+-----------------------------------+
362
- | LR, GRAD NORM and every other series the loop logs |
352
+ | trainmeter <state> <n> x <GPU> (theme)|
353
+ +----------------------------------------------------------------------+
354
+ | PROGRESS <percent> tokens, FLOPs of required, ETA, elapsed |
355
+ +-----------+-----------+-----------+-----------+-----------+----------+
356
+ | tokens/s | step time | MFU | OFU | goodput | FLOPs/tok|
357
+ | sparkline | sparkline | sparkline | sparkline | sparkline | |
358
+ +-----------+-----------+-----------+-----------+-----------+----------+
359
+ | TRAINING [x axis: step|tokens|...] |
360
+ | loss (full width), tokens/s, MFU and OFU, lr, grad norm, |
361
+ | every other logged series, plan against actual |
362
+ +----------------------------------------------------------------------+
363
+ | GPU EFFICIENCY: ladder now and over time, roofline, arithmetic |
364
+ | intensity, FLOPs spent (modeled next to observed) |
365
+ +----------------------------------------------------------------------+
366
+ | NODE: time split, host CPU and memory |
367
+ | GPUS: one small chart per counter, one line per GPU; ranks |
368
+ +----------------------------------------------------------------------+
369
+ | > Run details (collapsed) > Model card (collapsed) |
363
370
  +----------------------------------------------------------------------+
364
371
  ```
365
372
 
366
- ```text
367
- +----------------------------------+-----------------------------------+
368
- | UTILIZATION LADDER over time | ROOFLINE |
369
- | GPU-Util, SM active, | arithmetic intensity against |
370
- | tensor active, OFU, MFU | achieved FLOP/s, with the ridge |
371
- | | and the peak of the device |
372
- +----------------------------------+-----------------------------------+
373
- | FLOPs SPENT | ARITHMETIC INTENSITY over time |
374
- | modeled next to observed | aggregate over the window |
375
- +----------------------------------+-----------------------------------+
376
- ```
377
-
378
- The optional price field is the one place where money exists.
379
- It computes dollars in the browser from `elapsed`, `gpu_hours` and `tokens_seen`, keeps the price in that browser, and never sends it to the supervisor or writes it to the run log.
373
+ The dashboard shows no price and no dollars.
374
+ Money stays an overlay over `elapsed`, `gpu_hours` and `tokens_seen`, and a price never reaches the supervisor or the run log.
380
375
 
381
376
  ## 8. Run lifecycle
382
377
 
@@ -536,7 +531,7 @@ flowchart TB
536
531
  subgraph sup["Inside the tm process"]
537
532
  cli["cli.py, config.py"]
538
533
  supervisor["supervisor/<br/>launcher, ingest, store"]
539
- sources["sources/<br/>gpu.py, host.py"]
534
+ sources["sources/<br/>gpu/, host.py"]
540
535
  web["web/<br/>server, dashboard"]
541
536
  end
542
537
 
@@ -545,7 +540,7 @@ flowchart TB
545
540
  catalog["catalog.py<br/>metric names,<br/>units, aliases"]
546
541
  facts["facts.py<br/>provenance,<br/>precedence"]
547
542
  records["records.py<br/>wire schema,<br/>codec"]
548
- formulas["peaks.py, flops.py,<br/>metrics.py<br/>exist today"]
543
+ formulas["hardware/, flops.py,<br/>metrics.py<br/>exist today"]
549
544
  end
550
545
 
551
546
  inproc -->|"import records.py only"| core
@@ -562,13 +557,13 @@ The dependency rule is one sentence: edges import the core, and the core imports
562
557
 
563
558
  | Module | Layer | Role | State |
564
559
  | --- | --- | --- | --- |
565
- | `peaks.py`, `flops.py`, `metrics.py` | pure core | peaks, FLOPs conventions, formulas | exists |
560
+ | `hardware/`, `flops.py`, `metrics.py` | pure core | device specs and peaks, FLOPs conventions, formulas | exists |
566
561
  | `records.py` | pure core | wire schema, codec, schema version | new |
567
562
  | `facts.py` | pure core | provenance, precedence, conflict warnings | new |
568
563
  | `catalog.py` | pure core | canonical metric names, units, labels, aliases, served to the dashboard | new |
569
564
  | `engine.py` | pure core | records in, snapshots out; takes over the arithmetic inside `Meter._step` | new |
570
565
  | `supervisor/` | tm process | launcher, ingest, store | new |
571
- | `sources/` | tm process | `gpu.py` (NVML behind a seam), `host.py` (`/proc`) | new |
566
+ | `sources/` | tm process | `gpu/` (NVML behind a seam, one adapter per kind of machine), `host.py` (`/proc`) | new |
572
567
  | `web/` | tm process | standard-library server, static dashboard | new |
573
568
  | `agent/` | training process | bootstrap, probes and taps | new |
574
569
  | `emit.py` | training process | the public `emit()` | new |
@@ -579,6 +574,22 @@ The engine is a pure function of records and an injected clock.
579
574
  Live it is fed by the ingest queue, and in `tm view` it is fed by `events.jsonl`, so both paths produce the same numbers.
580
575
  A recorded run is also the test fixture: golden logs go in, snapshots come out.
581
576
 
577
+ ### GPU platforms
578
+
579
+ Machines differ in what NVML can tell about them, so the GPU sampler has three layers, and only the middle one knows which machine it is on.
580
+
581
+ | Layer | Module | Job |
582
+ | --- | --- | --- |
583
+ | Device access | `sources/gpu/backend.py` | NVML and GPM calls; a field the device does not offer is left out |
584
+ | Platform adapter | `sources/gpu/platform.py`, `server.py`, `spark.py` | Turns what the device offers into one row schema, and names each field the machine never gives with the reason |
585
+ | Sampler | `sources/gpu/sampler.py` | Picks the GPUs, asks each one's adapter, emits `sample.gpu` |
586
+
587
+ The row schema in `platform.py` is the contract: the engine, the dashboard and the passport read only those fields, never the machine.
588
+ The device's spec in `hardware/specs/` names its adapter, and a device without a spec is read as a server GPU.
589
+ `server` passes NVML through, which is right for discrete cards with their own memory.
590
+ `spark` covers the DGX Spark, whose GB10 shares the node's memory with the CPU: it reads memory from `/proc/meminfo` and leaves out the memory clock, PCIe and NVLink, which do not describe that chip.
591
+ A new kind of machine is a new adapter and a spec file; nothing downstream changes.
592
+
582
593
  ## 12. Decisions and alternatives
583
594
 
584
595
  | # | Decision | Why | Rejected |
@@ -10,16 +10,23 @@ Every item names the code it verifies and what a wrong answer looks like.
10
10
  - [ ] No root and no DCGM are needed.
11
11
  If something asks for them, that is a bug.
12
12
 
13
+ ## Device specs (`hardware/specs/`)
14
+
15
+ The spec files were transcribed without access to the linked pages from the build environment.
16
+
17
+ - [ ] For each device you can reach, open every `sources` link in its spec file, compare each figure, and set `checked` to the date.
18
+ A figure that disagrees with the page is a bug, and a sparse figure taken as dense halves MFU.
19
+
13
20
  ## `tm doctor`
14
21
 
15
22
  - [ ] `tm doctor` prints `ok` for NVML with the driver version.
16
- - [ ] Each GPU line shows the right model and its dense BF16 peak from the table.
23
+ - [ ] Each GPU line shows the right model and its dense BF16 peak from its spec, and the `platform` line names the right adapter.
17
24
  - [ ] Each `GPU n GPM` line is `ok` and prints a live read.
18
25
  - [ ] On a provider that blocks GPM, the line is `warn` with the reason, and `tm` still runs.
19
26
  - [ ] `CUDA_VISIBLE_DEVICES=1 tm doctor` selects GPU 1, and so does `tm doctor --gpus 1`.
20
27
  With `CUDA_DEVICE_ORDER` unset on a mixed-GPU node, the ordering warning appears.
21
28
 
22
- ## Scale of the counters (`sources/gpu.py`, `NvmlBackend.gpm`)
29
+ ## Scale of the counters (`sources/gpu/backend.py`, `NvmlBackend.gpm`)
23
30
 
24
31
  The code divides SM, tensor and DRAM activity by 100 because the NVML headers document percent.
25
32
  Nothing has confirmed it on a device.
@@ -54,12 +61,23 @@ Nothing has confirmed it on a device.
54
61
  - [ ] The counted step shows no extra device synchronization (compare step time of step 6 with steps 5 and 7) and no extra GPU memory.
55
62
  - [ ] With `torch.compile` the probe reports that nothing was counted and the counted card is absent, instead of showing a small number.
56
63
 
57
- ## `tm attach` (`sources/gpu.py`, `processes`)
64
+ ## `tm attach` (`sources/gpu/backend.py`, `processes`)
58
65
 
59
66
  - [ ] `tm attach <pid>` on a program running on GPU 1 samples GPU 1 only and `n_gpus` says `detected`.
60
67
  - [ ] Inside a container the pids do not match, and the fallback to every GPU shows its warning; `--gpus` fixes it.
61
68
  - [ ] Attaching does not change the program's step time, and stopping `tm` with Ctrl-C leaves it running.
62
69
 
70
+ ## DGX Spark (`sources/gpu/spark.py`)
71
+
72
+ The GB10 adapter is built from public reports of what NVML returns on that chip, not from a run on one.
73
+
74
+ - [ ] `tm doctor` names the `spark` adapter, says the memory is unified, and lists the memory clock, PCIe and NVLink as not on this machine.
75
+ - [ ] The GPU line says GB10 has no dense BF16 figure, and MFU is absent with the same reason on the dashboard.
76
+ - [ ] The `GPU 0 GPM` line says whether GPM works on GB10; record the answer here, since NVIDIA does not publish it.
77
+ If it works, run the dense bf16 matmul loop from the counters section and check the scale the same way.
78
+ - [ ] During training, the "Memory used, shared with the CPU" chart follows `free -b` (total minus available) within a few percent.
79
+ - [ ] Utilization, power, temperature and SM clock move with the load, and nothing in the GPU section is an empty chart.
80
+
63
81
  ## Agent device probe (`agent/core.py`)
64
82
 
65
83
  - [ ] `n_gpus` in `summary.json` says `detected` under `torchrun --nproc_per_node=N`, and `assumed` with a warning under a non-Python trainer.
@@ -0,0 +1,136 @@
1
+ # Roadmap
2
+
3
+ This plan follows from the user analysis in [`users.md`](users.md).
4
+ Every item is tied to a pain or a point of friction from there.
5
+ Sizes are an estimate, not a commitment: S is up to a couple of days, M up to a week, L more than a week.
6
+ There are no dates, because they depend on access to a real GPU node.
7
+
8
+ The slide version is the deck [trainmeter: development plan](https://claude.ai/artifact/QDxC41C2o4AxuVX9onhdN4) (Russian, private until shared).
9
+
10
+ ## Three themes, ordered by the cost of a mistake
11
+
12
+ 1. **Trust in the numbers.** A wrong MFU costs more than a missing one, because budgets are planned on it.
13
+ If a number is wrong, everything else only delivers the mistake faster.
14
+ 2. **A short path to the first MFU.** Until the user reaches a number, the value of the product is invisible.
15
+ 3. **A quiet night.** A hang or a throughput drop costs hours of rental, and the run should not vanish with the node.
16
+ It is the most expensive pain, but it matters to those who already use the product.
17
+
18
+ ## Theme 1: trust in the numbers
19
+
20
+ | What to do | What the user gets | Size |
21
+ | --- | --- | --- |
22
+ | Go through `hardware-checklist.md` on an H100 | MFU, OFU and GPM checked against a real node | M |
23
+ | Peaks for popular rented GPUs | MFU on L40S, RTX 4090 and others, each with its datasheet | S |
24
+ | "How it is computed" on every tile | The formula, its inputs and where they came from, checkable by hand | M |
25
+ | Counted FLOPs under `torch.compile` | Counting before compilation instead of an empty card | L |
26
+ | Node calibration | How much of the datasheet peak this particular node delivers | M |
27
+
28
+ The checklist comes first: until it passes, no number is confirmed.
29
+ New peaks are dense BF16 with a datasheet only, or MFU is not shown (invariant 2).
30
+ Counting FLOPs under `torch.compile` is marked in the architecture as needing a spike, hence L.
31
+ Node calibration runs CUDA and conflicts with invariant 9; see the decisions below.
32
+
33
+ ## Theme 2: a short path
34
+
35
+ | What to do | The friction it removes | Size |
36
+ | --- | --- | --- |
37
+ | An observe-only banner with the reason | The agent is silent: say which Python lacks trainmeter | S |
38
+ | `tm doctor --python` | Check the training environment, not the one `tm` runs in | S |
39
+ | A URL that does not get lost, and `tm open` | The link and a ready `ssh -L` command in the run directory | S |
40
+ | Ready flags when there is no MFU | The reason and a command to copy | S |
41
+ | Choosing the convention in the dashboard | The MFU range becomes a number in one click | M |
42
+
43
+ Choosing the convention works at view time and writes nothing to the run log.
44
+ It is possible because FLOPs per token under all four conventions follow from the model shape.
45
+
46
+ The first run, step by step:
47
+
48
+ | Step | Today | Target |
49
+ | --- | --- | --- |
50
+ | 1. Install | `pip` refuses, `uv tool` goes to the wrong environment | One `uv add` into the training environment |
51
+ | 2. Run | The agent may fail to attach without saying so | The banner says at once what is seen and what is not |
52
+ | 3. Open | The URL is in the log, the tunnel is set up by hand | `tm open` and a ready `ssh -L` command |
53
+ | 4. See MFU | A range, or nothing without flags | A number with its convention, or an exact hint |
54
+
55
+ The measure is the spec's: the first MFU on screen within 10 minutes of installation, timed on a clean rented node.
56
+
57
+ ## Theme 3: a quiet night
58
+
59
+ | What to do | What the user gets | Size |
60
+ | --- | --- | --- |
61
+ | Alerts: hung, throughput dropped | Slice 6 finished; the rules are already on the paused branch | S |
62
+ | A heartbeat to a URL | An outside service notices that the node died | M |
63
+ | A notification through a webhook | A message in Telegram, Slack or ntfy | S |
64
+ | `tm report RUN` | The run's dashboard as one HTML file | M |
65
+ | A copy of the run off the node | The run directory survives the end of the rental | M |
66
+
67
+ A dead node cannot report its own death, so it needs an outside observer.
68
+ A heartbeat every minute to a URL such as healthchecks.io raises the alarm when the signals stop.
69
+ The heartbeat is off by default: data leaves the node only on an explicit flag.
70
+ `tm report` replaces what the W&B export was expected to give: our own dashboard with every label, as one file that opens anywhere.
71
+
72
+ ## Profilers alongside `tm`
73
+
74
+ `tm` shows that the GPUs are underused across the whole run; a profiler explains why, inside a window of a few steps.
75
+ We do not build a profiler of our own, but we stay correct next to one and read what it produces.
76
+
77
+ | What to do | What the user gets | Size |
78
+ | --- | --- | --- |
79
+ | Mark steps taken under `torch.profiler` | Profiled steps leave the EMA, ETA and MFU averages and show as a band on the charts | S |
80
+ | Detect `ncu` and `nsys` and pause GPM | Counter metrics read "profiler attached" instead of wrong values | S |
81
+ | `ncu` as the reference in `hardware-checklist.md` | Observed FLOPs, arithmetic intensity and peaks checked kernel by kernel on a real node | S |
82
+ | Import a `torch.profiler` trace | The step time split into GEMM, attention, exposed NCCL, copies and GPU idle, top kernels, and a per-kernel roofline on our peak table | M |
83
+ | Import an `ncu` report | Speed-of-light and roofline for the kernels of one step | M |
84
+ | Trace capture on demand | A "trace the next steps" button with no change to the training code | L |
85
+
86
+ `torch.profiler` matters more to us than `ncu`.
87
+ `ncu` says how well one kernel runs, while a trace says where the step time goes, which is where low MFU usually comes from.
88
+ Under `ncu` kernels are replayed and serialized and clocks are locked, so its timings never mix with the run's own.
89
+ Every imported result is labeled as a window under a profiler, never shown in place of the run's metrics (invariants 3 and 4).
90
+ Trace analysis reads JSON with the standard library and runs in `tm`, so nothing new enters the training process; traces reach hundreds of megabytes and are parsed as a stream.
91
+ Capture on demand would have the agent start a profiler inside the run: the profiler may synchronize the device and writes its own file, against invariants 5 and 8, so it waits for a decision below.
92
+ Kineto's on-demand tracing through dynolog is the path to study first, because it puts no code of ours in the process.
93
+
94
+ ## Releases
95
+
96
+ | Release | Goal | Contents |
97
+ | --- | --- | --- |
98
+ | 0.1 | Correct numbers, quick start | The checklist on an H100, peaks for rented GPUs, the observe-only banner, `tm doctor --python`, `tm open` and the URL, profiled steps marked and GPM paused under a profiler |
99
+ | 0.2 | A quiet night | Alerts from slice 6, heartbeat and webhook, `tm report`, a copy of the run off the node |
100
+ | 0.3 | Deeper into the numbers | Choosing the convention, "how it is computed", counted FLOPs under compile, run comparison, `torch.profiler` trace import |
101
+
102
+ Run comparison sits in 0.3 conditionally: it needs the hypothesis from `users.md` confirmed first.
103
+
104
+ ## How we will know it worked
105
+
106
+ The four thresholds are already in the success criteria of `docs/spec.md`:
107
+
108
+ | Threshold | Measure | Closed by |
109
+ | --- | --- | --- |
110
+ | 10 minutes | From installation to the first MFU on screen | Theme 2 |
111
+ | 1 percentage point | MFU against a hand calculation under the same convention | Theme 1 |
112
+ | 5 minutes | From the node's death to an alert, with the heartbeat on | Theme 3 |
113
+ | Under 0.5% | Overhead per training step | Theme 1, on a real node |
114
+
115
+ And the standing condition: the loss curve with and without `tm` is identical.
116
+
117
+ ## Deliberately not in this plan
118
+
119
+ - Multi-node runs, Kubernetes, Slurm: the primary segment works on one node.
120
+ - AMD, TPU, Apple Silicon: NVIDIA first, and checked on it.
121
+ - A profiler of our own: kernels are for `torch.profiler` and Nsight, and `tm` only reads what they produce.
122
+ - Pausing or stopping a run: `tm` only observes and changes nothing.
123
+ - Billing and prices: money stays a view-time overlay.
124
+ - Panels in W&B: `tm report` instead.
125
+
126
+ The first five match the out-of-scope list in `docs/spec.md`.
127
+
128
+ ## Decisions needed from the owner before 0.1
129
+
130
+ 1. Where and when to get an H100 node for the checklist: [provider, hours].
131
+ 2. Node calibration runs CUDA: is it acceptable as a separate command run before training, outside invariant 9?
132
+ 3. Which GPUs to add to the peak table first: [L40S, RTX 4090, ...].
133
+ 4. The notification channel: a generic webhook, or Telegram and Slack directly?
134
+ 5. The W&B export: keep it, write into the run the loop already logs to, or remove it?
135
+ 6. Trace capture on demand: is an opt-in exception to invariants 5 and 8 acceptable, or does capture stay with the user?
136
+ 7. Trace analysis: our own breakdown on the standard library, or HTA or TraceLens as an optional extra?
@@ -73,7 +73,7 @@ Rules:
73
73
 
74
74
  1. Unknown stays unknown.
75
75
  A metric that lacks an input is absent and names the input that would enable it, and a zero is never shown in its place.
76
- 2. An assumed fact carries a visible badge.
76
+ 2. An assumed fact is marked `assumed` in the run details, and the dashboard raises a notice for it.
77
77
  A declared value that disagrees with a detected one is used, and a warning shows both.
78
78
  3. `n_gpus` is the number of GPUs the run uses, taken from the ranks.
79
79
  Falling back to the visible devices is an assumed fact with a warning.
@@ -107,7 +107,7 @@ The engine maps the raw key the loop used to a canonical series:
107
107
  | `lr` | `lr`, `learning_rate`, `train/lr`, `train/learning_rate` |
108
108
  | `grad_norm` | `grad_norm`, `train/grad_norm`, `gradient_norm` |
109
109
 
110
- A mapped series carries a `mapped by name` badge and its original key.
110
+ A mapped series shows its original key, with `mapped by name` in its tooltip.
111
111
  `tm.toml` can pin or override a mapping.
112
112
  A number that matches no alias still appears as a custom series under its own key.
113
113
 
@@ -149,7 +149,7 @@ The last column is the label the dashboard shows beside the number.
149
149
  | MFU | FLOPs per token x tokens/s / (GPUs x dense BF16 peak) | model, peak table | convention and peak source |
150
150
  | OFU | tensor active x SM clock / max SM clock | GPM | `counters` |
151
151
  | MFU-OFU gap | OFU minus MFU | both | both labels |
152
- | Arithmetic intensity | ridge x MFU / DRAM active, in FLOP per byte of HBM traffic | MFU, GPM, peak table | `aggregate over the window, not per kernel` |
152
+ | Arithmetic intensity | ridge x MFU / DRAM active, in FLOP per byte of memory traffic | MFU, GPM, peak table | `aggregate over the window, not per kernel` |
153
153
  | Roofline point | arithmetic intensity against achieved FLOP/s (MFU x peak), on the device's ridge and dense peak | as above | as above |
154
154
 
155
155
  OFU follows the May 2026 preprint "Instant GPU Efficiency Visibility at Fleet Scale" (arXiv 2605.20799).
@@ -186,7 +186,7 @@ Across ranks: the spread of step time, for straggler detection.
186
186
  Overlays are computed from the primitives at view time and never stored in the run log.
187
187
  Dollars are GPU-hours times a price per GPU-hour, or elapsed hours times a price per node-hour.
188
188
  GPU energy is the integral of the sampled power.
189
- A price is typed into the dashboard, stays in that browser, and is never sent to the supervisor.
189
+ The dashboard shows no dollars, and a price is never sent to the supervisor.
190
190
 
191
191
  ## Run log
192
192
 
@@ -257,23 +257,35 @@ The run ends when the program does, with no exit status, because `tm` did not st
257
257
 
258
258
  ## Screens
259
259
 
260
- | Screen | Shows | Answers |
260
+ The dashboard is one page with no tabs.
261
+ Its header holds only the icon, the run state, the rig (GPU count and model) and the theme toggle.
262
+ The x-axis selector sits in the heading of the Training section, next to the charts it changes.
263
+ The look follows beszel: one warm light theme and one cool dark theme, cards with a title and one muted line under it, flat translucent area charts and meter bars.
264
+ Below it, top to bottom:
265
+
266
+ | Part | Shows | Answers |
261
267
  | --- | --- | --- |
262
- | Cockpit | progress (FLOPs spent of required, tokens, ETA), train and validation loss, tokens/s and step time, MFU and OFU with their gap, lr, grad norm and every other logged series | Is the run on plan? |
263
- | Efficiency | the utilization ladder over time, MFU against OFU, arithmetic intensity over time, the roofline point, FLOPs spent (modeled next to observed) | Are the GPUs used well, and where does the gap open? |
264
- | Hardware | per GPU: tensor, DRAM and SM active, power, clocks, memory, temperature, NVLink, stragglers; and the process tree's CPU (cores busy) and resident memory | Which card is slow? |
265
- | Where the time goes | node time split: boot, startup and compile, data, compute, NCCL, checkpoints, eval, idle | What did the node spend time on without training? |
266
- | Run passport | architecture, FLOPs per token, hardware, topology, versions, final losses | What goes into the model card? |
267
- | Plan vs actual | the declared budget and sizing assumptions (MFU, time) next to measured values | How wrong was my plan? |
268
+ | Progress | percent of the token budget, tokens and FLOPs spent of required, ETA, elapsed and GPU-hours | Is the run on plan? |
269
+ | Tiles | tokens/s, step time, MFU, OFU, goodput and FLOPs per token, each a number with a sparkline | How is it going right now? |
270
+ | Training | train and validation loss, tokens/s, MFU against OFU, lr, grad norm, every other logged series, and plan against actual | Is the model learning, and how wrong was my plan? |
271
+ | GPU efficiency | the utilization ladder now and over time, the roofline point, arithmetic intensity over time, FLOPs spent (modeled next to observed) | Are the GPUs used well, and where does the gap open? |
272
+ | Node | node time split (startup and compile, data, checkpoints, eval, compute), and the process tree's CPU (cores busy) and resident memory | What did the node spend time on without training? |
273
+ | GPUs | per GPU: tensor, DRAM and SM active, power, clocks, memory, temperature, NVLink, and step time per rank | Which card is slow? |
274
+ | Run details (collapsed) | the command, host, directory, declared flags, model, FLOPs per token under every convention, hardware and versions | What exactly ran? |
275
+ | Model card (collapsed) | the run passport as Markdown, with a copy button | What goes into the model card? |
276
+
277
+ A chart with no data is left out, and a part with nothing to draw is one line that names what would fill it.
268
278
 
269
279
  Rendering rules:
270
280
 
271
- 1. Every number shows its convention and provenance, and `assumed` and `mapped by name` carry a visible badge.
272
- 2. A metric that lacks an input names the input that would enable it, such as "budget not declared: pass --budget-tokens", and never shows a zero.
281
+ 1. Every MFU and FLOPs number shows its convention next to it in plain text, and a series mapped by name shows its raw key.
282
+ Provenance (`declared`, `assumed`, `measured`) is listed next to each fact in the run details, and a guessed input also raises a notice at the top of the page.
283
+ A tile is a label, a number, one short line and a sparkline, with no badges; any longer explanation is in its tooltip.
284
+ 2. A metric that lacks an input shows "n/a", names the input that would enable it in its tooltip, such as "budget not declared: pass --budget-tokens", and never shows a zero.
273
285
  3. Arithmetic intensity is labeled as an aggregate over the sampling window.
274
- 4. Every time series has an x-axis switch: step, tokens, FLOPs spent (modeled, with its convention shown) or wall time.
286
+ 4. Every time series follows one x-axis selector: step, tokens, FLOPs spent (modeled, with its convention shown) or wall time.
275
287
  Until a step has been seen, and for a program watched from outside, the charts use wall time, because every point would otherwise sit at zero.
276
- 5. The header has an optional price field that adds dollars spent and dollars per billion tokens, as described under Overlays.
288
+ 5. The dashboard shows no price and no dollars.
277
289
  6. The dashboard is static files with hand-written SVG charts, no build step and no CDN.
278
290
 
279
291
  ## MVP
@@ -281,7 +293,7 @@ Rendering rules:
281
293
  The first build is slices 1 to 4 of the build order in `docs/architecture.md`.
282
294
  It answers two questions: how is the model training, and how well are the GPUs used.
283
295
 
284
- - [x] Peak table (dense BF16 and HBM bandwidth) with datasheet links
296
+ - [x] Device specs (dense BF16 and memory bandwidth), one file per GPU with the datasheet lines they come from
285
297
  - [x] FLOPs per token under four named conventions
286
298
  - [x] Pure formulas in `metrics.py`: MFU, OFU, arithmetic intensity, ETA, EMA
287
299
  - [x] Launcher: `tm -- cmd` with the signal and exit-status contract, run directory, run log, host sampler and exit summary (slice 1)