fm-bench 0.7.1 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,7 +1,8 @@
1
1
  # fm-bench
2
2
 
3
- [![CI](https://github.com/devinoldenburg/fm-bench/actions/workflows/ci.yml/badge.svg)](https://github.com/devinoldenburg/fm-bench/actions/workflows/ci.yml)
3
+ [![CI](https://github.com/dvnold/fm-bench/actions/workflows/ci.yml/badge.svg)](https://github.com/dvnold/fm-bench/actions/workflows/ci.yml)
4
4
  [![npm](https://img.shields.io/npm/v/fm-bench.svg)](https://www.npmjs.com/package/fm-bench)
5
+ [![node](https://img.shields.io/node/v/fm-bench.svg)](https://nodejs.org/)
5
6
 
6
7
  Benchmark Apple's `fm` command on macOS 27+.
7
8
 
@@ -30,28 +31,27 @@ npm install -g fm-bench
30
31
  fm-bench
31
32
  ```
32
33
 
33
- One command discovers your models, runs the standard prompt suite, and prints a full benchmark report (real output from a macOS 27.0 Apple M5 Pro):
34
+ One command discovers your models, warms each one up, runs the standard prompt suite, and prints a full benchmark report. Real output of `fm-bench --runs 3` on macOS 27.2 (26B5091g), Apple M5 Pro:
34
35
 
35
36
  ```text
36
- fm-bench 0.7.0 | darwin/arm64 | fm
37
- prompts 3 | runs 3 | concurrency 1 | stream on | measured 9 | failed 0 | skipped 0 | elapsed 6.37s | SLO
38
- TTFT<=1.00s,E2E<=3.00s
39
- unavailable: quota unavailable — this fm build exposes no quota command
37
+ fm-bench 0.8.0 | darwin/arm64 | fm
38
+ prompts 3 | runs 3 | concurrency 1 | stream on | measured 9 | failed 0 | skipped 0 | elapsed 11.07s
39
+ models: system = AFM 3 Core Advanced
40
40
 
41
- ┌───┬────────┬────────┬─────┬──────┬──────────┬───────┬───────┬─────────┬────────┬───────┬─────┬──────┐
42
- │ C │ MODEL │ STATUS │ OK │ GOOD │ GOOD RPS │ TTFT │ E2E │ E2E P95 │ USER/S │ SYS/S │ CV │ NOTE │
43
- ├───┼────────┼────────┼─────┼──────┼──────────┼───────┼───────┼─────────┼────────┼───────┼─────┼──────┤
44
- │ 1 │ system │ ok │ 9/9 │ 100% │ 1.7 │ 377ms │ 487ms │ 749ms │ 34.0 │ 32.8 │ 25% │ │
45
- └───┴────────┴────────┴─────┴──────┴──────────┴───────┴───────┴─────────┴────────┴───────┴─────┴──────┘
41
+ ┌───┬────────┬────────┬─────┬───────┬───────┬─────────┬────────┬───────┬────┬──────┐
42
+ │ C │ MODEL │ STATUS │ OK │ TTFT │ E2E │ E2E P95 │ USER/S │ SYS/S │ CV │ NOTE │
43
+ ├───┼────────┼────────┼─────┼───────┼───────┼─────────┼────────┼───────┼────┼──────┤
44
+ │ 1 │ system │ ok │ 9/9 │ 620ms │ 687ms │ 1.43s │ 40.6 │ 37.7 │ 3% │ │
45
+ └───┴────────┴────────┴─────┴───────┴───────┴─────────┴────────┴───────┴────┴──────┘
46
46
 
47
47
  ┌───┬────────┬────────┬─────────┬───────────┬──────────┬───────────┬─────────┬────────┐
48
48
  │ C │ MODEL │ IN AVG │ OUT AVG │ PREFILL/S │ DECODE/S │ CHUNK P95 │ E2E P99 │ REPEAT │
49
49
  ├───┼────────┼────────┼─────────┼───────────┼──────────┼───────────┼─────────┼────────┤
50
- │ 1 │ system │ 13 │ 20 │ 35.9 │ 113 │ 197ms │ 749ms │ 100% │
50
+ │ 1 │ system │ 27 │ 38 │ 40.7 │ 78.1 │ 197ms │ 1.45s │ 100% │
51
51
  └───┴────────┴────────┴─────────┴───────────┴──────────┴───────────┴─────────┴────────┘
52
52
  ```
53
53
 
54
- Default profile is `standard` (3 prompts). Use `--runs 5`, `--sweep-concurrency 1,2`, or `--profile client` for heavier suites.
54
+ Default profile is `standard` (3 prompts) with one run each, so run-to-run CV shows `-` until you pass `--runs 2` or more. Use `--runs 5`, `--sweep-concurrency 1,2`, or `--profile client` for heavier suites. When a model is unavailable (for example `pcc` outside the Terminal app), it is listed as skipped with fm's reason; `--available-only` hides it.
55
55
 
56
56
  Wide terminals add TTFT P95, TPOT, and RPS columns; medium terminals tighten the table; narrow terminals switch to compact model cards automatically. `--width <n>` previews any layout.
57
57
 
@@ -64,7 +64,7 @@ npm install -g fm-bench
64
64
  Install directly from GitHub (always latest):
65
65
 
66
66
  ```sh
67
- npm install -g --install-links git+https://github.com/devinoldenburg/fm-bench.git
67
+ npm install -g --install-links git+https://github.com/dvnold/fm-bench.git
68
68
  ```
69
69
 
70
70
  Local development:
@@ -74,24 +74,26 @@ npm install && npm link
74
74
  fm-bench doctor # verify your setup
75
75
  ```
76
76
 
77
- **Requirements:** macOS 27+, Node.js 20+, Apple Intelligence enabled.
77
+ **Requirements:** macOS 27+, Node.js 22+, Apple Intelligence enabled.
78
78
 
79
79
  ## Commands
80
80
 
81
81
  | Command | What it does |
82
82
  |---------|-------------|
83
83
  | `fm-bench` | Run the full benchmark (default) |
84
- | `fm-bench models` | Show the detected `fm` capabilities, discovered models, availability, and quota when supported |
84
+ | `fm-bench models` | Show the detected `fm` capabilities, discovered models, their identity (for example `AFM 3 Core Advanced`), availability, and quota |
85
85
  | `fm-bench compare <a.json> <b.json>` | Regression diff with suite/hardware/macOS warnings; `--strict` exits 2 when suites differ |
86
86
  | `fm-bench history [dir]` | Trend table from saved reports (sorted by time, tags visible) |
87
87
  | `fm-bench validate <report.json>` | Verify report JSON (schema v1) before sharing; `--json` for CI |
88
88
  | `fm-bench export <report.json>` | Standalone HTML report with embedded JSON |
89
89
  | `fm-bench legend` | Definitions, provenance (measured/proxy/derived), and color rules for every column |
90
- | `fm-bench doctor` | Environment and `fm` capability check; `--json` for scripts |
90
+ | `fm-bench doctor` | Environment, `fm` license, and capability check; `--json` for scripts |
91
91
  | `fm-bench metrics` | Alias for `legend` |
92
92
 
93
93
  Exit codes: `0` success, `1` operational failure (failed `--ci` gate, invalid reports), `2` usage or environment error (bad flags, unsupported macOS, unusable `fm`, no runnable model). Interrupting a run with Ctrl+C terminates in-flight `fm` processes.
94
94
 
95
+ Mistyped commands and flags are caught with a suggestion (`Unknown command "modles". Did you mean "models"?`) instead of being benchmarked as a prompt. To benchmark such a word anyway, put it after `--`.
96
+
95
97
  ## Common Recipes
96
98
 
97
99
  ```sh
@@ -130,7 +132,7 @@ fm-bench --format csv --out bench.csv
130
132
  |------|---------|-------------|
131
133
  | `-m, --models <list>` | discovered | Comma-separated or repeated model names |
132
134
  | `-r, --runs <n>` | 1 | Measured runs per prompt/model |
133
- | `--warmup <n>` | 0 | Unmeasured warmup runs per model before measurement |
135
+ | `--warmup <n>` | 1 | Unmeasured warmup runs per model before measurement; `0` measures cold start |
134
136
  | `-c, --concurrency <n>` | 1 | Parallel `fm` processes |
135
137
  | `--sweep-concurrency <list>` | — | Separate operating points, e.g. `1,2,4` |
136
138
  | `--request-rate <rps>` | — | Pace request starts at a target rate |
@@ -200,27 +202,32 @@ Nine built-in suites, choose the one that matches your use case:
200
202
 
201
203
  ## Metrics
202
204
 
203
- **Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT, 95% confidence interval, coefficient of variation (CV).
205
+ **Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT, 95% confidence interval, and run-to-run CV (per prompt across repeated runs, so mixing short and long prompts does not read as instability).
204
206
 
205
207
  **Throughput** — prefill tokens/s, decode tokens/s, output tokens/s per request, aggregate system tokens/s, requests per second.
206
208
 
207
- **Streaming quality** — second-chunk delay, chunk-gap p95, captured from stdout chunk arrival during streaming runs.
209
+ **Streaming quality** — first-chunk tokens, second-chunk delay, chunk-gap p95, captured from stdout chunk arrival during streaming runs.
208
210
 
209
211
  **Reliability** — success rate, goodput rate and RPS against SLO budgets, repeatability (most common output hash frequency across repeated runs).
210
212
 
213
+ **Model identity** — the identity `fm models` reports (for example `system = AFM 3 Core Advanced`) is printed in the report header and saved in JSON, and `compare` warns when the model behind a name changed between two runs.
214
+
211
215
  Every metric is labelled by provenance in JSON and in `fm-bench legend`:
212
216
 
213
217
  - **measured** — process wall clock, exit codes, chunk arrival, `fm` token counts
214
218
  - **proxy** — TTFT and chunk gaps (chunk granularity, not token timestamps); prefill tokens/s
215
219
  - **derived** — TPOT, decode tokens/s, throughput, CV, confidence intervals, goodput
216
220
 
217
- Token counts come from the `fm` build's own token-counting command (`count-tokens`, or `token-count` on older builds). If the build cannot count tokens, token metrics render as `-`, JSON carries `null`, and `metrics.promptTokens.available` is `false`. Spread statistics (`CV`, 95% CI) need at least two successful samples; with one sample they are unavailable rather than `0`. See [docs/methodology.md](docs/methodology.md).
221
+ Token counts come from the `fm` build's own token-counting command (`count-tokens`, or `token-count` on older builds). `count-tokens` adds one framing token to every count; fm-bench calibrates that overhead once per run and removes it from output counts (`tokenCounter` in JSON). If the build cannot count tokens, token metrics render as `-`, JSON carries `null`, and `metrics.promptTokens.available` is `false`. Spread statistics (`CV`, 95% CI) need at least two successful samples; with one sample they are unavailable rather than `0`.
222
+
223
+ `fm` streams coarse deltas — on macOS 27.2 the first stdout chunk carries about 20 tokens — so TTFT is time to the first chunk, and TPOT and decode tokens/s are computed only over the tokens that arrive after it. See [docs/methodology.md](docs/methodology.md).
218
224
 
219
225
  ## fm compatibility
220
226
 
221
227
  `fm-bench` probes `fm --help` and `fm respond --help` once per run and adapts:
222
228
 
223
- - Subcommand names (`count-tokens` vs legacy `token-count`) are detected, not hardcoded.
229
+ - Subcommand names (`count-tokens` vs legacy `token-count`, `models` vs the deprecated `available`) are detected, not hardcoded.
230
+ - Model availability and identity come from one `fm models` call; reasons such as `Private Cloud Compute is not available in this context. Please use the Terminal app.` are shown in full.
224
231
  - Flags the build does not document (`--stream`, `--use-case`, `--guardrails`, `--model`) are not passed through.
225
232
  - Unsupported models are rejected before a benchmark starts, with the supported list in the error.
226
233
  - Raw `fm` argument-error text never reaches reports, tables, or JSON.
@@ -293,8 +300,9 @@ fm-bench legend --json # machine-readable
293
300
  ## Requirements
294
301
 
295
302
  - macOS 27.0 or newer (Apple's `fm` CLI is preinstalled there).
296
- - Node.js 20 or newer.
297
- - Apple Intelligence enabled on the device.
303
+ - Node.js 22 or newer (CI covers Node 22, 24, and 26).
304
+ - Apple Intelligence enabled on the device, and the `fm` terms accepted once with `fm license` (`fm-bench doctor` checks this).
305
+ - To benchmark the Private Cloud Compute model (`pcc`), run fm-bench from the Terminal app: `fm` reports `pcc` as unavailable in other contexts such as editor terminals.
298
306
 
299
307
  Benchmark commands refuse to start on older macOS versions and report the detected version plus the latest supported macOS — see [docs/supported-platforms.md](docs/supported-platforms.md).
300
308
 
package/bin/fm-bench.js CHANGED
@@ -4,18 +4,29 @@ import { killActiveChildren } from '../src/process.js';
4
4
 
5
5
  // Terminate any in-flight `fm` child process before exiting, then use the
6
6
  // conventional 128 + signal exit code. A second signal exits immediately.
7
- let interrupting = false;
7
+ let signalsReceived = 0;
8
+ let interrupted = false;
8
9
  for (const [signal, exitCode] of [['SIGINT', 130], ['SIGTERM', 143]]) {
9
10
  process.on(signal, () => {
11
+ signalsReceived += 1;
12
+ interrupted = true;
13
+ // Carry the status on the process too, so the code survives even if the
14
+ // exit below is preempted.
15
+ process.exitCode = exitCode;
10
16
  const killed = killActiveChildren(signal);
11
- if (interrupting) process.exit(exitCode);
12
- interrupting = true;
13
- const graceMs = killed > 0 ? 500 : 0;
14
- setTimeout(() => process.exit(exitCode), graceMs).unref();
17
+ if (signalsReceived > 1 || killed === 0) process.exit(exitCode);
18
+ // Deliberately not unref'd: this timer performs the exit after giving the
19
+ // killed children a moment to be reaped.
20
+ setTimeout(() => process.exit(exitCode), 500);
15
21
  });
16
22
  }
17
23
 
18
24
  runCli(process.argv.slice(2)).catch((error) => {
25
+ if (interrupted) {
26
+ // The user asked to stop; a follow-up benchmark error is noise.
27
+ process.exitCode = typeof error?.exitCode === 'number' ? error.exitCode : 1;
28
+ return;
29
+ }
19
30
  const message = error?.message || String(error);
20
31
  console.error(`fm-bench: ${message}`);
21
32
  process.exitCode = typeof error?.exitCode === 'number' ? error.exitCode : 1;
@@ -1,15 +1,18 @@
1
1
  # `fm` compatibility
2
2
 
3
- `fm-bench` is a client of whatever `fm` binary is installed. Apple has already changed that surface between macOS 27 builds: current builds expose `count-tokens`, older ones exposed `token-count`, and quota reporting is not available everywhere. fm-bench therefore detects capabilities at runtime instead of assuming a fixed subcommand list.
3
+ `fm-bench` is a client of whatever `fm` binary is installed. Apple has already changed that surface between macOS 27 builds: `count-tokens` replaced `token-count`, macOS 27.2 renamed `available` to `models` (the old name still works there but prints a deprecation warning) and added `quota-usage`, `config`, and the `pcc` model. fm-bench therefore detects capabilities at runtime instead of assuming a fixed subcommand list.
4
4
 
5
5
  ## How detection works
6
6
 
7
- Once per run, `fm-bench` spawns two cheap calls:
7
+ Once per run, `fm-bench` spawns three cheap calls:
8
8
 
9
- 1. `fm --help` — the command list, the `MODELS` section, and (when present) a `--model` option list.
9
+ 1. `fm --help` — the command list, the `MODELS` section, and (when present) a `--model` option list. The `<model>` custom-provider placeholder is not treated as a model.
10
10
  2. `fm respond --help` — which flags `respond` actually accepts (`--model`, `--[no-]stream`, `--instructions`, `--greedy`, `--use-case`, `--guardrails`, `--image`, `--tool`, `--schema`).
11
+ 3. `fm models` — one call that reports every model's availability, its identity (for example `✓ system (AFM 3 Core Advanced)`), and the reason for anything unavailable. Builds without `models` fall back to `fm available --model <name>` per model.
11
12
 
12
- The result is recorded in the report (`capabilities`) and shown by `fm-bench models`, `fm-bench doctor`, and `fm-bench doctor --json`.
13
+ `warning:` lines that `fm` prints (such as the `available` rename notice) are discarded before parsing, so they can never be read as a model status or an error cause.
14
+
15
+ The result is recorded in the report (`capabilities`, `models[].identity`) and shown by `fm-bench models`, `fm-bench doctor`, and `fm-bench doctor --json`. `doctor` also runs `fm license --status`, because `fm respond` cannot run until the terms are accepted.
13
16
 
14
17
  Detection is tolerant of formatting changes: section headers are matched case-insensitively on uppercase lines, model lines by their column layout, and boolean flags in either `--flag` or Apple's `--[no-]flag` spelling. ANSI escapes are stripped before parsing.
15
18
 
@@ -18,6 +21,7 @@ Detection is tolerant of formatting changes: section headers are matched case-in
18
21
  | Capability | When missing |
19
22
  |------------|--------------|
20
23
  | Token counting (`count-tokens`, fallback `token-count`) | The run still benchmarks latency and streaming. Every token metric is `null` with `available: false` in `metrics`, the table prints an `unavailable:` note, and the report carries a warning. `fm count-tokens` is never called. |
24
+ | Model list (`models`, fallback `available`) | With neither command, models from `fm --help` are assumed runnable and any failure is recorded per run. Model identity is only available from `models`. |
21
25
  | Streaming | TTFT, generation time, TPOT, prefill, and chunk-gap metrics are unavailable, and `--stream`/`--no-stream` is not passed through. E2E latency, success rate, and RPS still work. |
22
26
  | Quota command | The `models` table omits the quota column and reports quota as unavailable. |
23
27
  | Model selection (`--model`) | The flag is not passed; `fm` uses its default model. |
@@ -41,6 +45,7 @@ The macOS 27 requirement is enforced only when `fm-bench` resolves the default `
41
45
 
42
46
  | Platform | `fm` surface | Notes |
43
47
  |----------|--------------|-------|
44
- | macOS 27.0 (26A5425a), Apple M5 Pro (Mac17,9) | `available`, `chat`, `count-tokens`, `license`, `respond`, `schema`, `serve`; model `system` only; no quota command; no `--version` | Real-machine smoke tests in 0.7.0. `test/fixtures/fm-help-macos27.txt` captures this help output as a regression baseline. |
48
+ | macOS 27.2 (26B5091g), Apple M5 Pro (Mac17,9) | `chat`, `config`, `count-tokens`, `license`, `models`, `quota-usage`, `respond`, `schema`, `serve` (plus deprecated `available`); models `system` (`AFM 3 Core Advanced`) and `pcc`; `--model-provider` for custom Chat Completions providers; no `--version` | Real-machine runs in 0.8.0. `pcc` reports "not available in this context. Please use the Terminal app." outside Terminal. `count-tokens` adds one framing token per count. The first streamed chunk carries about 20 tokens. Fixtures: `test/fixtures/*-macos27.2.txt`. |
49
+ | macOS 27.0 (26A5425a), Apple M5 Pro (Mac17,9) | `available`, `chat`, `count-tokens`, `license`, `respond`, `schema`, `serve`; model `system` only; no quota command; no `--version` | Real-machine smoke tests in 0.7.0. `test/fixtures/fm-help-macos27.txt` captures this help output as a regression baseline; the fake `fm` emulates it with `FAKE_FM_SCENARIO=legacy-available`. |
45
50
 
46
51
  If your `fm` build differs, `fm-bench doctor` shows exactly what was detected, and the report's `capabilities` block records it. Please open an issue with the `doctor --json` output and the report digest when something is missing.
@@ -31,24 +31,25 @@ The same classification is machine-readable in every JSON report under `metrics`
31
31
 
32
32
  For a column-by-column terminal reference, run `fm-bench legend`.
33
33
 
34
- - `TTFT` — *proxy*. Time from starting `fm respond` to the first streamed stdout chunk. `fm` exposes no per-token timestamps, so this is a terminal-side approximation of time to first token; the first chunk may already contain several tokens.
34
+ - `TTFT` — *proxy*. Time from starting `fm respond` to the first streamed stdout chunk. `fm` exposes no per-token timestamps, so this is a terminal-side approximation of time to first token. `fm` streams coarse deltas: on macOS 27.2 the first chunk consistently carries about 20 tokens (the same through a pseudo-terminal and through `fm serve`), so TTFT is closer to "time to the first ~20 tokens". `first_chunk_tokens` records the exact count per run.
35
35
  - `E2E latency` — *measured*. Time from starting `fm respond` until the process exits and the full response is captured.
36
- - `generation_ms` — *derived*. `E2E - TTFT`, reported only when the answer arrives in more than one stdout chunk. When a single chunk carries the whole answer, prefill and decode are not separable and the value is unavailable instead of near zero.
37
- - `TPOT` — *derived*. `(E2E - TTFT) / (output_tokens - 1)`, reported when the answer has at least three output tokens so the interval is an average over two or more decode tokens rather than the inverse of a single chunk gap. Requires streaming and a token-counting `fm`.
38
- - `second_chunk_ms` — *proxy*. Time between the first and second streamed stdout chunks: a terminal-side signal for startup smoothness.
39
- - `chunk_gap` — *proxy*. Distribution of time between consecutive streamed stdout chunks. Useful for spotting streaming jitter; chunk-based, not token-based.
36
+ - **Deliveries.** `fm` often writes the tail of a short answer as several `write()` calls well under a millisecond apart. Stdout chunks that arrive less than 5 ms after the previous one are therefore grouped into one *delivery*. Real streaming deltas on macOS 27.2 arrive about 50–250 ms apart, so this only removes write bursts. Generation time, TPOT, decode rate, `second_chunk_ms`, and `chunk_gap` are computed over deliveries; `stdout_chunks` still counts raw chunks.
37
+ - `generation_ms` — *derived*. Last streamed delivery minus first streamed delivery, reported only when the answer arrives in more than one delivery. When a single chunk carries the whole answer, prefill and decode are not separable and the value is unavailable instead of near zero. Process teardown after the last chunk is excluded.
38
+ - `TPOT` — *derived*. `generation_ms / (output_tokens - first_chunk_tokens)`: the tokens that arrived during the generation window divided into it. Reported when at least two tokens arrived after the first chunk, so the interval is an average rather than the inverse of a single chunk gap. Requires streaming and a token-counting `fm`. Before 0.8.0 the denominator was `output_tokens - 1`, which credited the ~20 first-chunk tokens to the decode window and overstated decode speed by up to several times on short answers.
39
+ - `second_chunk_ms` — *proxy*. Time between the first and second streamed deliveries: a terminal-side signal for startup smoothness.
40
+ - `chunk_gap` — *proxy*. Distribution of time between consecutive streamed deliveries. Useful for spotting streaming jitter; delivery-based, not token-based.
40
41
  - `prefill_tokens_per_second` — *proxy*. Input prompt tokens divided by TTFT seconds. Because TTFT includes process startup and first-token latency, this systematically understates true prefill speed and should be read as an upper-bound-constrained estimate, not a kernel measurement.
41
42
  - `tokens_per_second` — *derived*. Output tokens divided by E2E seconds for one request.
42
- - `decode_tokens_per_second` — *derived*. Output tokens after the first token divided by generation seconds.
43
+ - `decode_tokens_per_second` — *derived*. Per run: output tokens that arrived after the first delivery divided by generation seconds. Short answers finish in two or three deliveries, so their per-run rate rests on few tokens. The table's DECODE/S column is therefore the token-weighted `decodeThroughput` (decode tokens summed over runs divided by the summed generation time), so long generations dominate. The `throughput` profile gives the most trustworthy decode figure (about 51–53 tokens/s on an M5 Pro with macOS 27.2).
43
44
  - `output token throughput` — *derived*. All successful output tokens for a model divided by that model's measured wall-clock window.
44
45
  - `total token throughput` — *derived*. Successful prompt and output tokens divided by the model's measured wall-clock window.
45
46
  - `RPS` — *measured*. Successful requests divided by the model's measured wall-clock window.
46
47
  - `goodput` — *derived*. Successful requests that also satisfy every provided SLO threshold. A run whose SLO metric is unavailable counts as not good, so an unverifiable SLO can never inflate goodput.
47
48
  - `goodput RPS` — *derived*. SLO-passing requests divided by the model's measured wall-clock window. Zero is reported when SLOs are set and nothing passes.
48
49
  - `repeatability` — *derived*. For repeated runs of the same prompt, the average share of runs that produced the most common normalized output hash.
49
- - `CV` — *derived*. Sample standard deviation divided by the mean. Reported only with two or more successful samples.
50
+ - `CV` — *derived*. Run-to-run stability: the E2E coefficient of variation (sample standard deviation divided by the mean) of each prompt across its repeated runs, averaged over prompts (`stabilityCv` in JSON). Reported once at least one prompt has two successful runs. The pooled `latency.cv` over every sample is still in the JSON, but it mostly reflects that prompts have different lengths, so the table no longer shows it. Before 0.8.0 the CV column used the pooled value.
50
51
  - `95% CI` — *derived*. t-distribution confidence interval around the sample mean, using exact t critical values up to 30 degrees of freedom and the standard 2.0 / 1.96 approximations beyond that. Reported only with two or more successful samples. Treat it as context, not proof, at small sample sizes.
51
- - `quota` — *measured*, when the `fm` build exposes a quota command. Most current builds do not, in which case `fm-bench` reports quota as unavailable rather than empty.
52
+ - `quota` — *measured*, when the `fm` build exposes a quota command. macOS 27.2 has `quota-usage` (the on-device `system` model reports "Not applicable", since quota only applies to `pcc`); macOS 27.0 has none, in which case `fm-bench` reports quota as unavailable rather than empty.
52
53
 
53
54
  ### Small samples
54
55
 
@@ -66,6 +67,8 @@ Use `--request-rate <rps>` to pace request starts independently of concurrency.
66
67
 
67
68
  `--warmup <n>` runs `n` unmeasured calls per model at the start of each operating point. The first `fm respond` after a cold start includes model load time, which can dominate a short prompt (observed at several hundred milliseconds on Apple silicon). Warmups are never mixed into the measured results.
68
69
 
70
+ Since 0.8.0 the default is one warmup per model, so a default run reports steady-state latency. Pass `--warmup 0` to measure cold start deliberately. `warmup` is part of the suite key, so `compare` flags a 0.7.x report (warmup 0) against a 0.8.0 default run as a different suite.
71
+
69
72
  ## Failures and retries
70
73
 
71
74
  A failed call is recorded as a failed measurement with its error text, and is excluded from latency and throughput statistics. `--retry <n>` retries failed calls with exponential backoff before recording the failure; the report keeps the run count and records the number of attempts separately, so a retried success is still one sample and never silently duplicates work.
@@ -76,7 +79,9 @@ The `client` profile is a pragmatic local-machine mix inspired by MLPerf Client'
76
79
 
77
80
  ## Caveats
78
81
 
79
- Token counts come from the `fm` build's own token-counting command (`count-tokens` on current builds, `token-count` on older ones). If a build has no such command, token-derived metrics are reported as unavailable rather than estimated. `fm-bench` cannot judge semantic quality unless you provide your own prompt suite and inspect captured outputs with `--capture-output`.
82
+ Token counts come from the `fm` build's own token-counting command (`count-tokens` on current builds, `token-count` on older ones). If a build has no such command, token-derived metrics are reported as unavailable rather than estimated.
83
+
84
+ `count-tokens` counts its input as a prompt and adds a constant framing token: on macOS 27.2 it reports 2 for `a`, 3 for `a a`, and 4 for `a a a`. Once per run fm-bench counts those three strings, derives the overhead from the consistent per-word step (`overhead = count(a) − step`), and subtracts it from output and first-chunk token counts. The result is saved as `tokenCounter: { command, overhead, calibrated }`. If the three counts are not consistent, no correction is applied and `calibrated` is `false`. Prompt token counts are left exactly as `fm` reports them, because the framing token is part of what the model processes. `count-tokens` always uses the on-device system tokenizer, so token counts for `pcc` are system-tokenizer counts. `fm-bench` cannot judge semantic quality unless you provide your own prompt suite and inspect captured outputs with `--capture-output`.
80
85
 
81
86
  Client-side measurements include process startup, local queueing, model prefill, streaming, detokenization, and terminal pipe overhead. That is intentional for a command-line benchmark, but it is not the same as an internal model-kernel benchmark.
82
87
 
package/docs/releasing.md CHANGED
@@ -1,61 +1,78 @@
1
1
  # Releasing
2
2
 
3
- `fm-bench` uses semver tags (`v*.*.*`). Pushing a tag runs the **Release** workflow: lint/test/package check, npm publish (with provenance), and a GitHub release whose body is taken from the matching section in **`CHANGELOG.md`** (plus a compare link).
3
+ `fm-bench` uses semver tags (`v*.*.*`). A release tag runs the **Release** workflow: lint/test/package check, npm publish with provenance, a GitHub release whose body is the matching section of **`CHANGELOG.md`** (plus a compare link), and an install check of the published package.
4
4
 
5
- ## Prerequisites
5
+ ## npm authentication
6
6
 
7
- - Repository secret **`NPM_TOKEN`**: npm automation token with publish access to `fm-bench`.
8
- - **`main`** is green on CI.
7
+ The Release workflow runs on Node 24 (npm 11) with `id-token: write`, so it publishes through **npm trusted publishing** (OIDC) when the package trusts this workflow. No long-lived token is needed, and provenance is attached automatically. Set it up once (requires an npm login with publish rights and 2FA):
9
8
 
10
- If `NPM_TOKEN` is missing or expired, the Release workflow warns, skips the npm publish step, and still creates the GitHub release. The release notes then state that npm publish was skipped. Re-run the Release workflow with the tag once the token is fixed — an already-published version is skipped automatically.
9
+ ```sh
10
+ npm login
11
+ npm trust github fm-bench --repo dvnold/fm-bench --file release.yml
12
+ npm trust list fm-bench
13
+ ```
14
+
15
+ or on npmjs.com: **fm-bench → Settings → Trusted publishing → GitHub Actions**, organization/user `dvnold`, repository `fm-bench`, workflow `release.yml`.
16
+
17
+ Without trusted publishing, npm falls back to the **`NPM_TOKEN`** repository secret (a granular access token with publish rights). npm limits those tokens to 90 days, so an old token fails with `E401`/`E403`/`E404`.
18
+
19
+ When publishing fails, the workflow still creates or updates the GitHub release (its notes say npm publish failed) and then fails the run, so the problem is visible. Fix the authentication, then re-run the workflow for the tag. An already-published version is skipped.
20
+
21
+ Provenance also requires `repository.url` in `package.json` to match the repository the workflow runs in (`git+https://github.com/dvnold/fm-bench.git`). Update it if the repository is renamed or transferred.
11
22
 
12
23
  ## Before tagging
13
24
 
14
25
  ```sh
15
26
  npm run check # lint + tests + npm pack integrity
16
27
  npm run publish:dry-run
28
+ actionlint # if workflows changed
17
29
  ```
18
30
 
19
- Then bump `CHANGELOG.md` with a `## X.Y.Z` section; the release notes come from that section.
31
+ Then add a `## X.Y.Z` section to `CHANGELOG.md`; the release notes come from that section.
20
32
 
21
33
  ## Option A — GitHub Actions (recommended)
22
34
 
23
35
  1. Open **Actions → Version → Run workflow**.
24
36
  2. Choose `patch`, `minor`, `major`, or an exact semver.
25
- 3. The job runs `npm version`, pushes the commit and tag to `main`.
26
- 4. The tag push triggers **Release** automatically.
37
+ 3. The job runs `npm version`, pushes the commit and tag to `main`, and dispatches **Release** for the new tag. (A tag pushed with the workflow's `GITHUB_TOKEN` does not trigger other workflows by itself, so Version starts Release explicitly.)
27
38
 
28
39
  ## Option B — Local
29
40
 
30
41
  ```sh
31
42
  npm ci && npm run check
32
43
  npm version minor # or patch / major
33
- git push origin main --follow-tags
44
+ git push origin main --follow-tags # the tag push triggers Release
34
45
  ```
35
46
 
36
47
  ## Re-run Release without republishing
37
48
 
38
- If npm already has the version but the GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.6.3`). The workflow skips npm publish when that version is already on the registry. If the GitHub release already exists, it **updates the release notes** from `CHANGELOG.md`.
49
+ If npm already has the version but the GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.8.0`). The workflow skips npm publish when that version is already on the registry. If the GitHub release already exists, it **updates the release notes** from `CHANGELOG.md`.
50
+
51
+ ```sh
52
+ gh workflow run release.yml --repo dvnold/fm-bench -f tag=v0.8.0
53
+ ```
39
54
 
40
55
  Refresh notes locally without re-publishing:
41
56
 
42
57
  ```sh
43
- node scripts/changelog-release-notes.mjs 0.6.3 > notes.md
44
- gh release edit v0.6.3 --notes-file notes.md --repo devinoldenburg/fm-bench
58
+ node scripts/changelog-release-notes.mjs 0.8.0 > notes.md
59
+ gh release edit v0.8.0 --notes-file notes.md --repo dvnold/fm-bench
45
60
  ```
46
61
 
47
62
  ## Verifying a release
48
63
 
64
+ The workflow installs the published version into a clean directory and runs it against the fake `fm`. To check by hand:
65
+
49
66
  ```sh
50
- git tag --list 'v0.7.0'
51
- gh release view v0.7.0 --repo devinoldenburg/fm-bench
67
+ gh release view v0.8.0 --repo dvnold/fm-bench
52
68
  npm view fm-bench version # registry version
69
+ npm view fm-bench@0.8.0 dist.attestations --json # provenance
53
70
 
54
71
  mkdir -p fm-bench-check && cd fm-bench-check
55
- npm install fm-bench@0.7.0
72
+ npm install fm-bench@0.8.0
56
73
  ./node_modules/.bin/fm-bench --version
57
- ./node_modules/.bin/fm-bench --help > /dev/null
58
74
  ./node_modules/.bin/fm-bench doctor
75
+ ./node_modules/.bin/fm-bench --profile quick --runs 3
59
76
  ```
60
77
 
61
78
  ## Dry run
@@ -1,6 +1,6 @@
1
1
  # Report format (schema v1)
2
2
 
3
- Every measured run can be saved as JSON. Reports from fm-bench **0.6.0+** include a versioned schema so you can validate, share, and compare results across machines. The schema version stayed at `1` in 0.7.0: the new `capabilities`, `metrics`, and per-result `attempts` fields are additive, so older readers and older reports both keep working.
3
+ Every measured run can be saved as JSON. Reports from fm-bench **0.6.0+** include a versioned schema so you can validate, share, and compare results across machines. The schema version is still `1`: the `capabilities`, `metrics`, and per-result `attempts` fields (0.7.0) and the `tokenCounter`, `models[].identity`, `firstChunkTokens`, and `stabilityCv` fields (0.8.0) are additive, so older readers and older reports both keep working.
4
4
 
5
5
  ## Top-level fields
6
6
 
@@ -15,10 +15,11 @@ Every measured run can be saved as JSON. Reports from fm-bench **0.6.0+** includ
15
15
  | `environment` | Host fingerprint: platform, arch, Node, hardware model, CPU, memory, macOS version/build, `fm` help digest, thermal/power snapshot |
16
16
  | `capabilities` | What the installed `fm` build exposes (see below) |
17
17
  | `metrics` | Per-metric availability and provenance for this run (see below) |
18
- | `suite` | Derived suite key + fingerprint for apples-to-apples comparison |
18
+ | `tokenCounter` | `{ command, overhead, calibrated }`: the token-counting command and the framing overhead removed from output counts (see [methodology](./methodology.md#caveats)) |
19
+ | `suite` | Derived suite key + fingerprint for apples-to-apples comparison. `suite.fingerprint.modelIdentities` maps model names to identities, for example `{ "system": "AFM 3 Core Advanced" }` |
19
20
  | `prompts` | Prompt ids, text, and token counts |
20
- | `models` | Discovered models and availability |
21
- | `summary` | Per-model / per-concurrency roll-up statistics |
21
+ | `models` | Discovered models, availability, `reason` when unavailable, and `identity` as reported by `fm models` |
22
+ | `summary` | Per-model / per-concurrency roll-up statistics, including `stabilityCv` (run-to-run CV), `decodeThroughput` (token-weighted decode tokens/s), and `firstChunkTokens` |
22
23
  | `results` | Per-run measurements (optional `output` when `--capture-output`) |
23
24
 
24
25
  ## `capabilities`
@@ -29,13 +30,18 @@ Detected once per run from `fm --help` and `fm respond --help`.
29
30
  {
30
31
  "capabilities": {
31
32
  "bin": "fm",
32
- "digest": "921a7839714e3707",
33
- "commands": ["available", "chat", "count-tokens", "license", "respond", "schema", "serve"],
34
- "models": [{ "name": "system", "description": "On-device Apple Foundation Model" }],
33
+ "digest": "a45f2838f399c5da",
34
+ "commands": ["chat", "config", "count-tokens", "license", "models", "quota-usage", "respond", "schema", "serve"],
35
+ "models": [
36
+ { "name": "system", "description": "On-device Apple Foundation Model" },
37
+ { "name": "pcc", "description": "Apple Foundation Model on Private Cloud Compute" }
38
+ ],
35
39
  "features": {
36
40
  "tokenCounting": true,
37
41
  "tokenCountCommand": "count-tokens",
38
- "quota": false,
42
+ "modelListCommand": "models",
43
+ "license": true,
44
+ "quota": true,
39
45
  "streaming": true,
40
46
  "modelSelection": true,
41
47
  "instructions": true,
@@ -47,11 +53,13 @@ Detected once per run from `fm --help` and `fm respond --help`.
47
53
  "structuredOutput": true,
48
54
  "server": true
49
55
  },
50
- "warnings": ["this fm build exposes no quota command, so quota is not reported"]
56
+ "warnings": []
51
57
  }
52
58
  }
53
59
  ```
54
60
 
61
+ On a macOS 27.0 build the same block has `"modelListCommand": "available"`, `"quota": false`, and the warning `this fm build exposes no quota command, so quota is not reported`.
62
+
55
63
  `digest` is a short hash of the normalized `fm --help` output, so a report records which CLI surface produced it. Reports from two different `fm` builds are not directly comparable even on the same machine.
56
64
 
57
65
  ## `metrics`
@@ -72,8 +80,8 @@ Each entry states how a metric was obtained for this run and whether it was avai
72
80
  "label": "quota",
73
81
  "kind": "measured",
74
82
  "source": "fm quota-usage",
75
- "available": false,
76
- "unavailableReason": "this fm build exposes no quota command"
83
+ "available": true,
84
+ "unavailableReason": ""
77
85
  }
78
86
  }
79
87
  }
@@ -87,13 +95,14 @@ Per-run rows in `results` carry:
87
95
  |-------|-------|
88
96
  | `attempts` | Total `fm` invocations for this measured run, including retries. `1` when no retry was needed. |
89
97
  | `ok` | Whether the call produced a response and exited cleanly. |
90
- | `firstTokenMs`, `generationMs`, `tpotMs` | `null` when the run cannot supply them (no streaming, single-chunk answer, failed run). |
91
- | `promptTokens`, `outputTokens`, `tokensPerSecond`, `decodeTokensPerSecond`, `prefillTokensPerSecond` | `null` when the `fm` build cannot count tokens. |
98
+ | `firstTokenMs`, `generationMs`, `tpotMs` | `null` when the run cannot supply them (no streaming, single-chunk answer, failed run). `generationMs` runs from the first to the last streamed chunk. |
99
+ | `promptTokens`, `outputTokens`, `tokensPerSecond`, `decodeTokensPerSecond`, `prefillTokensPerSecond` | `null` when the `fm` build cannot count tokens. `promptTokens` is fm's own count; `outputTokens` has `tokenCounter.overhead` removed. |
100
+ | `firstChunkTokens`, `decodeTokens` | Output tokens carried by the first streamed delivery (overhead removed), and the tokens that arrived after it. `tpotMs` and `decodeTokensPerSecond` use only `decodeTokens`. `null` for single-delivery or non-streamed answers. |
92
101
  | `error` | Actionable failure text (`timed out after 30000ms`, `fm exited with code 3`, the first actionable line of `fm` stderr). |
93
102
 
94
103
  ## CSV
95
104
 
96
- `--format csv` and `--out runs.csv` write per-run rows. Columns are stable; `attempts` was added in 0.7.0. Text cells that begin with `=`, `+`, `@`, or a non-numeric `-` are prefixed with a single quote so spreadsheets do not execute prompt or model output as a formula.
105
+ `--format csv` and `--out runs.csv` write per-run rows. Columns are stable; `attempts` was added in 0.7.0, and `first_chunk_tokens` was appended as the last column in 0.8.0 so positional readers keep working. Text cells that begin with `=`, `+`, `@`, or a non-numeric `-` are prefixed with a single quote so spreadsheets do not execute prompt or model output as a formula.
97
106
 
98
107
  ## Sharing results
99
108
 
@@ -5,7 +5,7 @@
5
5
  ## Requirements
6
6
 
7
7
  - **macOS 27.0 or newer** — Apple's `fm` CLI is preinstalled starting with macOS 27. Older macOS releases do not ship `fm`, so `fm-bench` cannot run there.
8
- - **Node.js 20 or newer**.
8
+ - **Node.js 22 or newer**. Node 20 reached end of life in April 2026; CI tests Node 22, 24, and 26.
9
9
  - **Apple Intelligence enabled** on the device.
10
10
 
11
11
  ## Version enforcement
@@ -31,13 +31,18 @@ Support is capability-based rather than version-pinned. The CLI probes the insta
31
31
 
32
32
  ## Models
33
33
 
34
- `fm-bench` benchmarks exactly the models the installed `fm` reports — it does not assume that any particular cloud or adapter model exists. On the verified macOS 27.0 build that is the on-device `system` model only. If a build adds models (for example a Private Cloud Compute model), they are discovered automatically and reported with their own availability.
34
+ `fm-bench` benchmarks exactly the models the installed `fm` reports — it does not assume that any particular cloud or adapter model exists. macOS 27.0 ships the on-device `system` model only; macOS 27.2 adds `pcc` (Apple Foundation Model on Private Cloud Compute). New models are discovered automatically and reported with their own availability and identity.
35
35
 
36
- Requesting only models the build cannot run exits with code `2` before any benchmark starts, and `fm-bench models` shows the same reasons without failing:
36
+ `fm` only serves `pcc` to the Terminal app. From editor terminals and other contexts it reports `Private Cloud Compute is not available in this context. Please use the Terminal app.`, and fm-bench shows that reason unchanged. Run fm-bench from Terminal to benchmark `pcc`.
37
+
38
+ Requesting only models that cannot run right now exits with code `2` before any benchmark starts, and `fm-bench models` shows the same reasons without failing:
37
39
 
38
40
  ```text
39
41
  fm-bench: No benchmark was run: none of the requested models are usable right now.
40
42
  requested: pcc
41
- pcc: not supported by this fm build (supported: system)
43
+ models reported by this fm build: system, pcc
44
+ pcc: Private Cloud Compute is not available in this context. Please use the Terminal app.
42
45
  run "fm-bench models" to see availability and reasons
43
46
  ```
47
+
48
+ On a build that does not have the model at all, the reason reads `not supported by this fm build (supported: system)` instead.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "fm-bench",
3
- "version": "0.7.1",
3
+ "version": "0.8.0",
4
4
  "description": "Dynamic benchmark CLI for Apple's fm command on macOS 27+.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -26,25 +26,31 @@
26
26
  },
27
27
  "keywords": [
28
28
  "apple",
29
+ "apple-intelligence",
30
+ "apple-silicon",
29
31
  "foundation-models",
30
32
  "fm",
33
+ "llm",
31
34
  "benchmark",
35
+ "benchmarking",
36
+ "latency",
37
+ "throughput",
32
38
  "macos",
33
39
  "cli"
34
40
  ],
35
41
  "author": "Devin Oldenburg",
36
42
  "license": "MIT",
37
43
  "engines": {
38
- "node": ">=20"
44
+ "node": ">=22"
39
45
  },
40
46
  "repository": {
41
47
  "type": "git",
42
- "url": "git+https://github.com/devinoldenburg/fm-bench.git"
48
+ "url": "git+https://github.com/dvnold/fm-bench.git"
43
49
  },
44
50
  "bugs": {
45
- "url": "https://github.com/devinoldenburg/fm-bench/issues"
51
+ "url": "https://github.com/dvnold/fm-bench/issues"
46
52
  },
47
- "homepage": "https://github.com/devinoldenburg/fm-bench#readme",
53
+ "homepage": "https://github.com/dvnold/fm-bench#readme",
48
54
  "publishConfig": {
49
55
  "access": "public",
50
56
  "registry": "https://registry.npmjs.org/"