fm-bench 0.6.3 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,7 @@
1
1
  # fm-bench
2
2
 
3
3
  [![CI](https://github.com/devinoldenburg/fm-bench/actions/workflows/ci.yml/badge.svg)](https://github.com/devinoldenburg/fm-bench/actions/workflows/ci.yml)
4
+ [![npm](https://img.shields.io/npm/v/fm-bench.svg)](https://www.npmjs.com/package/fm-bench)
4
5
 
5
6
  Benchmark Apple's `fm` command on macOS 27+.
6
7
 
@@ -8,7 +9,7 @@ Measure latency, throughput, streaming smoothness, stability, and goodput across
8
9
 
9
10
  ## Why fm-bench exists
10
11
 
11
- Apple's Foundation Models can run on-device, through Private Cloud Compute, and through model adapters. That makes raw model quality only half the story.
12
+ Apple's Foundation Models run on-device, and the `fm` CLI is the terminal and script interface to them. That makes raw model quality only half the story.
12
13
 
13
14
  For real apps, the important questions are:
14
15
 
@@ -16,10 +17,12 @@ For real apps, the important questions are:
16
17
  - Does streaming stay smooth?
17
18
  - How stable is latency over repeated runs?
18
19
  - What happens under concurrency?
19
- - Which model/hardware pair meets an interactive SLO?
20
+ - Does this Mac meet an interactive SLO?
20
21
 
21
22
  `fm-bench` answers those questions with repeatable local benchmarks. Think of it as GeekBench for Apple Foundation Models — run it, get numbers, compare across hardware, models, and macOS updates.
22
23
 
24
+ It is deliberately honest about what it can measure: the installed `fm` is probed for its actual capabilities, and any metric that build cannot supply is reported as unavailable rather than estimated. See [docs/compatibility.md](docs/compatibility.md).
25
+
23
26
  ## Quick Start
24
27
 
25
28
  ```sh
@@ -27,22 +30,30 @@ npm install -g fm-bench
27
30
  fm-bench
28
31
  ```
29
32
 
30
- One command discovers your models, runs the standard prompt suite, and prints a full benchmark report:
33
+ One command discovers your models, runs the standard prompt suite, and prints a full benchmark report (real output from a macOS 27.0 Apple M5 Pro):
31
34
 
32
35
  ```text
33
- fm-bench 0.6.3 | darwin/arm64 | fm
34
- prompts 3 | runs 5 | concurrency 1 | stream on | measured 15 | failed 0 | skipped 0 | elapsed 38.20s
35
-
36
- ┌───┬────────┬────────┬─────────┬──────┬──────┬──────────┬──────┬──────────┬─────┬─────┐
37
- │ C │ MODEL │ STATUS │ OK/RUNS │ SUCC │ GOOD │ GOOD RPS │ TTFT │ E2E P95 │ SYS │ CV │
38
- ├───┼────────┼────────┼─────────┼──────┼──────┼──────────┼──────┼──────────┼─────┼─────┤
39
- │ 1 │ system │ ok │ 15/15 │ 100% │ 93% │ 0.4 │ 318ms│ 3.20s │ 42 │ 12% │
40
- └───┴────────┴────────┴─────────┴──────┴──────┴──────────┴──────┴──────────┴─────┴─────┘
36
+ fm-bench 0.7.0 | darwin/arm64 | fm
37
+ prompts 3 | runs 3 | concurrency 1 | stream on | measured 9 | failed 0 | skipped 0 | elapsed 6.37s | SLO
38
+ TTFT<=1.00s,E2E<=3.00s
39
+ unavailable: quota unavailable — this fm build exposes no quota command
40
+
41
+ ┌───┬────────┬────────┬─────┬──────┬──────────┬───────┬───────┬─────────┬────────┬───────┬─────┬──────┐
42
+ │ C │ MODEL │ STATUS │ OK │ GOOD │ GOOD RPS │ TTFT │ E2E │ E2E P95 │ USER/S │ SYS/S │ CV │ NOTE │
43
+ ├───┼────────┼────────┼─────┼──────┼──────────┼───────┼───────┼─────────┼────────┼───────┼─────┼──────┤
44
+ │ 1 │ system │ ok │ 9/9 │ 100% │ 1.7 │ 377ms │ 487ms │ 749ms │ 34.0 │ 32.8 │ 25% │ │
45
+ └───┴────────┴────────┴─────┴──────┴──────────┴───────┴───────┴─────────┴────────┴───────┴─────┴──────┘
46
+
47
+ ┌───┬────────┬────────┬─────────┬───────────┬──────────┬───────────┬─────────┬────────┐
48
+ │ C │ MODEL │ IN AVG │ OUT AVG │ PREFILL/S │ DECODE/S │ CHUNK P95 │ E2E P99 │ REPEAT │
49
+ ├───┼────────┼────────┼─────────┼───────────┼──────────┼───────────┼─────────┼────────┤
50
+ │ 1 │ system │ 13 │ 20 │ 35.9 │ 113 │ 197ms │ 749ms │ 100% │
51
+ └───┴────────┴────────┴─────────┴───────────┴──────────┴───────────┴─────────┴────────┘
41
52
  ```
42
53
 
43
54
  Default profile is `standard` (3 prompts). Use `--runs 5`, `--sweep-concurrency 1,2`, or `--profile client` for heavier suites.
44
55
 
45
- Wide terminals add TTFT P95, TPOT, decode/prefill throughput, chunk-gap smoothness, and 95% CI columns. Narrow terminals switch to compact model cards automatically.
56
+ Wide terminals add TTFT P95, TPOT, and RPS columns; medium terminals tighten the table; narrow terminals switch to compact model cards automatically. `--width <n>` previews any layout.
46
57
 
47
58
  ## Install
48
59
 
@@ -70,13 +81,16 @@ fm-bench doctor # verify your setup
70
81
  | Command | What it does |
71
82
  |---------|-------------|
72
83
  | `fm-bench` | Run the full benchmark (default) |
73
- | `fm-bench models` | List discovered models, availability, and quota |
74
- | `fm-bench compare <a.json> <b.json>` | Regression diff with suite/hardware warnings; `--strict` for CI |
84
+ | `fm-bench models` | Show the detected `fm` capabilities, discovered models, availability, and quota when supported |
85
+ | `fm-bench compare <a.json> <b.json>` | Regression diff with suite/hardware/macOS warnings; `--strict` exits 2 when suites differ |
75
86
  | `fm-bench history [dir]` | Trend table from saved reports (sorted by time, tags visible) |
76
- | `fm-bench validate <report.json>` | Verify report JSON (schema v1) before sharing |
87
+ | `fm-bench validate <report.json>` | Verify report JSON (schema v1) before sharing; `--json` for CI |
77
88
  | `fm-bench export <report.json>` | Standalone HTML report with embedded JSON |
78
- | `fm-bench legend` | Definitions for every table column and color rule |
79
- | `fm-bench doctor` | Environment check: Node, macOS, `fm`, CPU, memory, thermals, battery |
89
+ | `fm-bench legend` | Definitions, provenance (measured/proxy/derived), and color rules for every column |
90
+ | `fm-bench doctor` | Environment and `fm` capability check; `--json` for scripts |
91
+ | `fm-bench metrics` | Alias for `legend` |
92
+
93
+ Exit codes: `0` success, `1` operational failure (failed `--ci` gate, invalid reports), `2` usage or environment error (bad flags, unsupported macOS, unusable `fm`, no runnable model). Interrupting a run with Ctrl+C terminates in-flight `fm` processes.
80
94
 
81
95
  ## Common Recipes
82
96
 
@@ -84,15 +98,12 @@ fm-bench doctor # verify your setup
84
98
  # Quick smoke test
85
99
  fm-bench --profile quick
86
100
 
87
- # Standard 5-run benchmark with SLO budgets
88
- fm-bench --runs 5 --slo-ttft-ms 750 --slo-e2e-ms 4000
101
+ # Standard 5-run benchmark with SLO budgets and warmups
102
+ fm-bench --runs 5 --warmup 1 --slo-ttft-ms 750 --slo-e2e-ms 4000
89
103
 
90
104
  # Sweep concurrency to find your throughput ceiling
91
105
  fm-bench --sweep-concurrency 1,2,4 --runs 3
92
106
 
93
- # Stress both on-device and PCC models
94
- fm-bench --models system,pcc --runs 3 --profile stress
95
-
96
107
  # Reasoning and coding workloads
97
108
  fm-bench --profile reasoning --runs 5
98
109
  fm-bench --profile coding --runs 3 --histogram
@@ -103,7 +114,7 @@ fm-bench --output-dir reports/ --tag after-update --export-html
103
114
  fm-bench validate reports/*.json
104
115
  fm-bench compare reports/fm-bench_*before*.json reports/fm-bench_*after*.json --strict
105
116
 
106
- # Fail CI when SLOs regress
117
+ # Fail CI when SLOs regress or any run fails
107
118
  fm-bench --ci --slo-ttft-ms 750 --slo-e2e-ms 4000 --runs 5
108
119
 
109
120
  # Save JSON for automation
@@ -117,9 +128,9 @@ fm-bench --format csv --out bench.csv
117
128
 
118
129
  | Flag | Default | Description |
119
130
  |------|---------|-------------|
120
- | `-m, --models <list>` | all | Comma-separated or repeated model names |
131
+ | `-m, --models <list>` | discovered | Comma-separated or repeated model names |
121
132
  | `-r, --runs <n>` | 1 | Measured runs per prompt/model |
122
- | `--warmup <n>` | 0 | Warmup runs per model before measurement |
133
+ | `--warmup <n>` | 0 | Unmeasured warmup runs per model before measurement |
123
134
  | `-c, --concurrency <n>` | 1 | Parallel `fm` processes |
124
135
  | `--sweep-concurrency <list>` | — | Separate operating points, e.g. `1,2,4` |
125
136
  | `--request-rate <rps>` | — | Pace request starts at a target rate |
@@ -129,7 +140,9 @@ fm-bench --format csv --out bench.csv
129
140
  | `--profile <name>` | standard | Built-in prompt suite (see Profiles) |
130
141
  | `-p, --prompt <text>` | — | Custom prompt, repeatable |
131
142
  | `--prompt-file <file>` | — | JSON, JSONL, or blank-line separated prompts |
132
- | `-i, --instructions <text>` | — | Passed to `fm respond` |
143
+ | `-i, --instructions <text>` | — | Passed to `fm respond` when the build supports it |
144
+ | `--use-case <case>` | — | System model use case, when supported |
145
+ | `--guardrails <level>` | — | System model guardrail level, when supported |
133
146
 
134
147
  **Quality Gates**
135
148
 
@@ -159,11 +172,15 @@ fm-bench --format csv --out bench.csv
159
172
 
160
173
  | Flag | Description |
161
174
  |------|-------------|
175
+ | `--no-stream` | Disable streaming; TTFT and decode metrics are reported as unavailable |
176
+ | `--greedy` / `--no-greedy` | Request or omit greedy sampling (default: greedy) |
177
+ | `--available-only` | Hide unavailable discovered models |
162
178
  | `--color` / `--no-color` | Force or disable ANSI colors (auto on TTYs) |
163
179
  | `--ascii` | Plain ASCII table borders instead of Unicode |
164
180
  | `--compact` | Force narrow terminal layout |
165
181
  | `--width <n>` | Render as if the terminal is `n` columns wide |
166
182
  | `--progress` / `--no-progress` | Force or disable the live progress line |
183
+ | `--fm-bin <path>` | `fm` binary to execute (default: `FM_BIN` or `fm`) |
167
184
 
168
185
  ## Prompt Profiles
169
186
 
@@ -181,9 +198,38 @@ Nine built-in suites, choose the one that matches your use case:
181
198
  | `coding` | 5 | Code review, refactoring, algorithms, system design |
182
199
  | `creative` | 5 | Product copy, analogies, commit messages, docs |
183
200
 
201
+ ## Metrics
202
+
203
+ **Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT, 95% confidence interval, coefficient of variation (CV).
204
+
205
+ **Throughput** — prefill tokens/s, decode tokens/s, output tokens/s per request, aggregate system tokens/s, requests per second.
206
+
207
+ **Streaming quality** — second-chunk delay, chunk-gap p95, captured from stdout chunk arrival during streaming runs.
208
+
209
+ **Reliability** — success rate, goodput rate and RPS against SLO budgets, repeatability (most common output hash frequency across repeated runs).
210
+
211
+ Every metric is labelled by provenance in JSON and in `fm-bench legend`:
212
+
213
+ - **measured** — process wall clock, exit codes, chunk arrival, `fm` token counts
214
+ - **proxy** — TTFT and chunk gaps (chunk granularity, not token timestamps); prefill tokens/s
215
+ - **derived** — TPOT, decode tokens/s, throughput, CV, confidence intervals, goodput
216
+
217
+ Token counts come from the `fm` build's own token-counting command (`count-tokens`, or `token-count` on older builds). If the build cannot count tokens, token metrics render as `-`, JSON carries `null`, and `metrics.promptTokens.available` is `false`. Spread statistics (`CV`, 95% CI) need at least two successful samples; with one sample they are unavailable rather than `0`. See [docs/methodology.md](docs/methodology.md).
218
+
219
+ ## fm compatibility
220
+
221
+ `fm-bench` probes `fm --help` and `fm respond --help` once per run and adapts:
222
+
223
+ - Subcommand names (`count-tokens` vs legacy `token-count`) are detected, not hardcoded.
224
+ - Flags the build does not document (`--stream`, `--use-case`, `--guardrails`, `--model`) are not passed through.
225
+ - Unsupported models are rejected before a benchmark starts, with the supported list in the error.
226
+ - Raw `fm` argument-error text never reaches reports, tables, or JSON.
227
+
228
+ `fm-bench doctor` shows exactly what was detected. Policy and the verified-build table: [docs/compatibility.md](docs/compatibility.md).
229
+
184
230
  ## Sharing and comparing results
185
231
 
186
- Reports from 0.6.0+ include **schema v1**: `reportId`, hardware fingerprint, and a **suite key** so you can tell if two JSON files used the same prompts and run settings. See [docs/report-format.md](docs/report-format.md).
232
+ Reports from 0.6.0+ include **schema v1**: `reportId`, hardware fingerprint, a **suite key** so you can tell if two JSON files used the same prompts and run settings, plus the detected `fm` capabilities and per-metric availability. See [docs/report-format.md](docs/report-format.md).
187
233
 
188
234
  - Share **HTML** with teammates who do not use the CLI: `fm-bench export bench.json -o bench.html`
189
235
  - Gate uploads in CI: `fm-bench validate artifact.json`
@@ -194,33 +240,24 @@ Reports from 0.6.0+ include **schema v1**: `reportId`, hardware fingerprint, and
194
240
  Track performance across macOS updates, model changes, or hardware swaps:
195
241
 
196
242
  ```sh
197
- # Before
198
243
  fm-bench --profile coding --runs 5 --output-dir reports/ --tag before
199
-
200
- # After the change
201
244
  fm-bench --profile coding --runs 5 --output-dir reports/ --tag after
202
-
203
- # See what changed
204
245
  fm-bench compare reports/fm-bench_*before*.json reports/fm-bench_*after*.json
205
- ```
206
-
207
- The compare output shows each model/concurrency row with the before value, a color-coded percent delta (green = improvement, red = regression), and the after value — for TTFT, E2E, TPOT, tokens/s, RPS, success rate, and CV.
208
-
209
- ```sh
210
- # View the full trend over time
211
246
  fm-bench history reports/
212
247
  ```
213
248
 
249
+ The compare output shows each model/concurrency row with the before value, a color-coded percent delta (green = improvement, red = regression), and the after value — for TTFT, E2E, TPOT, tokens/s, RPS, success rate, and CV. It warns when hardware, macOS version/build, or the prompt suite differ.
250
+
214
251
  ## CI Integration
215
252
 
216
253
  Gate deployments or model updates on benchmark quality:
217
254
 
218
255
  ```sh
219
- # Fails with exit code 1 if TTFT > 750ms or E2E > 4s on any run
256
+ # Fails with exit code 1 if TTFT > 750ms, E2E > 4s, or any run fails
220
257
  fm-bench --ci --slo-ttft-ms 750 --slo-e2e-ms 4000 --runs 5
221
258
  ```
222
259
 
223
- Prints `fm-bench ci: PASS` or `fm-bench ci: FAIL — <reason>` to stderr. Designed for GitHub Actions, Buildkite, or any shell-based pipeline.
260
+ Prints `fm-bench ci: PASS` or `fm-bench ci: FAIL — <reason>` to stderr. Designed for GitHub Actions, Buildkite, or any shell-based pipeline. `--json` and `--csv` write only data to stdout; progress and diagnostics go to stderr.
224
261
 
225
262
  ## Prompt Files
226
263
 
@@ -240,36 +277,18 @@ JSONL:
240
277
  {"id":"latency","prompt":"Explain p95 latency in one sentence."}
241
278
  ```
242
279
 
243
- Plain text files are split on blank lines.
244
-
245
- ## Metrics
246
-
247
- **Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT (p50/p95), 95% confidence interval, coefficient of variation (CV).
248
-
249
- **Throughput** — prefill tokens/s, decode tokens/s, output tokens/s per request, aggregate system tokens/s, requests per second.
250
-
251
- **Streaming quality** — second-chunk delay, chunk-gap p95. Captured from `stdout` chunks during streaming runs.
252
-
253
- **Reliability** — success rate, goodput rate and RPS against SLO budgets, repeatability (most common output hash frequency across repeated runs).
254
-
255
- **Stability** — CV (stddev/mean for E2E latency); green ≤10%, yellow ≤25%, red >25%.
256
-
257
- Token counts come from `fm token-count --quiet`. If `fm` cannot count tokens, those fields are blank while character throughput is still reported.
258
-
259
- Terminal layout is responsive: wide → full scoreboard + detail tables, medium → tighter single table, narrow → compact model cards. Use `--width` to preview any layout and `--ascii` for log-friendly output.
280
+ Plain text files are split on blank lines. Parse errors name the file and, for JSONL, the line number.
260
281
 
261
282
  ## Colors and Legend
262
283
 
263
284
  Table output is color-coded on interactive terminals — **green** is better/passing, **yellow** is marginal/partial, **red** is failing/unstable. Fixed thresholds apply to success rate, goodput, CV, and repeatability. Latency uses SLO thresholds when set, otherwise lower-is-better relative ranking. Throughput uses higher-is-better relative ranking.
264
285
 
265
286
  ```sh
266
- fm-bench legend # full column definitions and color rules
287
+ fm-bench legend # definitions, provenance, and color rules
267
288
  fm-bench legend --json # machine-readable
268
289
  ```
269
290
 
270
- `NO_COLOR=1` disables color; `FORCE_COLOR=1` or `--color` enables it. `--ascii` switches to plain ASCII borders for log systems.
271
-
272
- A live single-line progress indicator runs on stderr during interactive sessions. The final report always goes to stdout — `--json`, `--csv`, and `--out` stay automation-friendly.
291
+ `NO_COLOR=1` disables color; `FORCE_COLOR=1` or `--color` enables it. `--ascii` switches to plain ASCII borders for log systems. A live single-line progress indicator runs on stderr during interactive sessions; the final report always goes to stdout.
273
292
 
274
293
  ## Requirements
275
294
 
@@ -279,19 +298,18 @@ A live single-line progress indicator runs on stderr during interactive sessions
279
298
 
280
299
  Benchmark commands refuse to start on older macOS versions and report the detected version plus the latest supported macOS — see [docs/supported-platforms.md](docs/supported-platforms.md).
281
300
 
282
- `pcc` (Private Cloud Compute) availability depends on Apple's current eligibility. `fm-bench` shows it as skipped if `fm available --model pcc` reports unavailable.
283
-
284
- See [docs/methodology.md](docs/methodology.md) for benchmark methodology and metric references.
285
-
286
301
  ## Development
287
302
 
288
303
  ```sh
289
304
  npm install
290
- npm test # node --test
291
- npm run lint
305
+ npm test # node --test (unit + integration against a fake fm)
306
+ npm run lint # node --check on every source file
307
+ npm run check # lint + tests + npm pack integrity — run this before pushing
292
308
  ```
293
309
 
294
- No runtime npm dependencies.
310
+ No runtime npm dependencies. Integration tests drive the real CLI against `test/fixtures/fake-fm.mjs`, which emulates normal, streaming, slow, malformed, failing, timed-out, partial, and interrupted `fm` behaviour, so the suite runs on machines without `fm`.
311
+
312
+ Documentation: [methodology](docs/methodology.md) · [report format](docs/report-format.md) · [compatibility](docs/compatibility.md) · [supported platforms](docs/supported-platforms.md) · [releasing](docs/releasing.md).
295
313
 
296
314
  ## License
297
315
 
package/bin/fm-bench.js CHANGED
@@ -1,5 +1,19 @@
1
1
  #!/usr/bin/env node
2
2
  import { runCli } from '../src/cli.js';
3
+ import { killActiveChildren } from '../src/process.js';
4
+
5
+ // Terminate any in-flight `fm` child process before exiting, then use the
6
+ // conventional 128 + signal exit code. A second signal exits immediately.
7
+ let interrupting = false;
8
+ for (const [signal, exitCode] of [['SIGINT', 130], ['SIGTERM', 143]]) {
9
+ process.on(signal, () => {
10
+ const killed = killActiveChildren(signal);
11
+ if (interrupting) process.exit(exitCode);
12
+ interrupting = true;
13
+ const graceMs = killed > 0 ? 500 : 0;
14
+ setTimeout(() => process.exit(exitCode), graceMs).unref();
15
+ });
16
+ }
3
17
 
4
18
  runCli(process.argv.slice(2)).catch((error) => {
5
19
  const message = error?.message || String(error);
@@ -0,0 +1,46 @@
1
+ # `fm` compatibility
2
+
3
+ `fm-bench` is a client of whatever `fm` binary is installed. Apple has already changed that surface between macOS 27 builds: current builds expose `count-tokens`, older ones exposed `token-count`, and quota reporting is not available everywhere. fm-bench therefore detects capabilities at runtime instead of assuming a fixed subcommand list.
4
+
5
+ ## How detection works
6
+
7
+ Once per run, `fm-bench` spawns two cheap calls:
8
+
9
+ 1. `fm --help` — the command list, the `MODELS` section, and (when present) a `--model` option list.
10
+ 2. `fm respond --help` — which flags `respond` actually accepts (`--model`, `--[no-]stream`, `--instructions`, `--greedy`, `--use-case`, `--guardrails`, `--image`, `--tool`, `--schema`).
11
+
12
+ The result is recorded in the report (`capabilities`) and shown by `fm-bench models`, `fm-bench doctor`, and `fm-bench doctor --json`.
13
+
14
+ Detection is tolerant of formatting changes: section headers are matched case-insensitively on uppercase lines, model lines by their column layout, and boolean flags in either `--flag` or Apple's `--[no-]flag` spelling. ANSI escapes are stripped before parsing.
15
+
16
+ ## Capability policy
17
+
18
+ | Capability | When missing |
19
+ |------------|--------------|
20
+ | Token counting (`count-tokens`, fallback `token-count`) | The run still benchmarks latency and streaming. Every token metric is `null` with `available: false` in `metrics`, the table prints an `unavailable:` note, and the report carries a warning. `fm count-tokens` is never called. |
21
+ | Streaming | TTFT, generation time, TPOT, prefill, and chunk-gap metrics are unavailable, and `--stream`/`--no-stream` is not passed through. E2E latency, success rate, and RPS still work. |
22
+ | Quota command | The `models` table omits the quota column and reports quota as unavailable. |
23
+ | Model selection (`--model`) | The flag is not passed; `fm` uses its default model. |
24
+ | `--use-case`, `--guardrails`, `--instructions`, `--greedy` | Silently not passed, instead of failing the run with an argument error. |
25
+ | No usable `fm` at all | The command exits `2` with the spawn error and a hint to install `fm` or point `--fm-bin` / `FM_BIN` at a compatible binary. |
26
+
27
+ Models the build does not list are refused before any benchmark starts, with the message `not supported by this fm build (supported: ...)`. Raw `fm` argument-error text (usage blocks, `Unknown command`) never reaches reports, tables, or JSON output.
28
+
29
+ ## Exit codes
30
+
31
+ | Code | Meaning |
32
+ |------|---------|
33
+ | `0` | Success. A failed measured run is data, not a CLI failure — use `--ci` to make failures fatal. |
34
+ | `1` | Operational failure: `--ci` gate failed, reports failed validation, or an `fm` run could not produce data. |
35
+ | `2` | Usage or environment error: unknown flag, missing argument, unsupported macOS, unusable `fm`, malformed report input, or `compare --strict` suite mismatch. |
36
+ | `130` / `143` | Interrupted by SIGINT / SIGTERM. In-flight `fm` child processes are terminated first. |
37
+
38
+ The macOS 27 requirement is enforced only when `fm-bench` resolves the default `fm` from `PATH`. Supplying `--fm-bin <path>` or `FM_BIN` lets the CLI run on any host and instead fails on the binary's actual capabilities (exit `2` when it exposes no usable commands).
39
+
40
+ ## Verified builds
41
+
42
+ | Platform | `fm` surface | Notes |
43
+ |----------|--------------|-------|
44
+ | macOS 27.0 (26A5425a), Apple M5 Pro (Mac17,9) | `available`, `chat`, `count-tokens`, `license`, `respond`, `schema`, `serve`; model `system` only; no quota command; no `--version` | Real-machine smoke tests in 0.7.0. `test/fixtures/fm-help-macos27.txt` captures this help output as a regression baseline. |
45
+
46
+ If your `fm` build differs, `fm-bench doctor` shows exactly what was detected, and the report's `capabilities` block records it. Please open an issue with the `doctor --json` output and the report digest when something is missing.
@@ -14,27 +14,45 @@ The metric set follows common LLM inference benchmark practice:
14
14
  - MLCommons describes varying concurrency and reporting verified operating points for TTFT, throughput, interactivity, and response latency rather than interpolated performance: <https://mlcommons.org/2026/03/mlperf-endpoints-gen-ai-benchmarking/>
15
15
  - MLPerf Client emphasizes local client workloads with multiple task types and varying prompt/response lengths: <https://mlcommons.org/benchmarks/client/>
16
16
 
17
+ ## What `fm` can actually tell us
18
+
19
+ `fm-bench` never invents precision the CLI cannot provide. Every metric is classified by how it is obtained:
20
+
21
+ | Kind | Meaning |
22
+ |------|---------|
23
+ | **measured** | Observed directly: process wall clock, exit codes, stdout chunk arrival, `fm`-reported token counts. |
24
+ | **proxy** | Observed at a coarser granularity than the ideal metric (for example chunk arrival instead of token timestamps). |
25
+ | **derived** | Computed from measured values (for example output tokens divided by elapsed seconds). |
26
+ | **controlled** | An input setting rather than a measurement, such as the concurrency operating point. |
27
+
28
+ The same classification is machine-readable in every JSON report under `metrics` and in the CLI via `fm-bench legend` (SOURCE column). A metric the installed `fm` build cannot support is `available: false` with a reason, and renders as `-` rather than a plausible-looking number.
29
+
17
30
  ## Metrics
18
31
 
19
32
  For a column-by-column terminal reference, run `fm-bench legend`.
20
33
 
21
- - `TTFT`: time from starting `fm respond` to the first streamed stdout chunk. This is a practical terminal-side proxy for time to first token.
22
- - `E2E latency`: time from starting `fm respond` until the process exits and the full response is captured.
23
- - `generation_ms`: `E2E - TTFT`.
24
- - `TPOT`: `(E2E - TTFT) / (output_tokens - 1)`. The first output token is excluded so TPOT focuses on decode cadence.
25
- - `second_chunk_ms`: time between the first and second streamed stdout chunks. This is a terminal-side proxy for time-to-second-token style startup smoothness.
26
- - `chunk_gap`: the distribution of time between consecutive streamed stdout chunks. It is useful for spotting streaming jitter, but it is chunk-based rather than token-based because the `fm` CLI writes stdout chunks, not token timestamp events.
27
- - `prefill_tokens_per_second`: input prompt tokens divided by TTFT seconds. This estimates prompt-processing speed for streaming runs.
28
- - `tokens_per_second`: output tokens divided by E2E seconds for one request.
29
- - `decode_tokens_per_second`: output tokens after the first token divided by generation seconds.
30
- - `total output token throughput`: all successful output tokens for a model divided by that model's measured wall-clock window.
31
- - `total token throughput`: successful prompt and output tokens divided by that model's measured wall-clock window.
32
- - `RPS`: successful requests divided by that model's measured wall-clock window.
33
- - `goodput`: successful requests that also satisfy all provided SLO thresholds.
34
- - `goodput RPS`: SLO-passing requests divided by that model's measured wall-clock window. If SLOs are set and no requests pass, this is reported as zero.
35
- - `repeatability`: for repeated runs of the same prompt, the average share of runs that produced the most common normalized output hash.
36
- - `CV`: coefficient of variation, or sample standard deviation divided by the mean. Lower values indicate steadier latency for that metric.
37
- - `95% CI`: a t-distribution confidence interval around the sample mean. Treat it as useful context, not proof, especially with very small sample sizes.
34
+ - `TTFT` — *proxy*. Time from starting `fm respond` to the first streamed stdout chunk. `fm` exposes no per-token timestamps, so this is a terminal-side approximation of time to first token; the first chunk may already contain several tokens.
35
+ - `E2E latency` — *measured*. Time from starting `fm respond` until the process exits and the full response is captured.
36
+ - `generation_ms` — *derived*. `E2E - TTFT`, reported only when the answer arrives in more than one stdout chunk. When a single chunk carries the whole answer, prefill and decode are not separable and the value is unavailable instead of near zero.
37
+ - `TPOT` — *derived*. `(E2E - TTFT) / (output_tokens - 1)`, reported when the answer has at least three output tokens so the interval is an average over two or more decode tokens rather than the inverse of a single chunk gap. Requires streaming and a token-counting `fm`.
38
+ - `second_chunk_ms` — *proxy*. Time between the first and second streamed stdout chunks: a terminal-side signal for startup smoothness.
39
+ - `chunk_gap` — *proxy*. Distribution of time between consecutive streamed stdout chunks. Useful for spotting streaming jitter; chunk-based, not token-based.
40
+ - `prefill_tokens_per_second` — *proxy*. Input prompt tokens divided by TTFT seconds. Because TTFT includes process startup and first-token latency, this systematically understates true prefill speed and should be read as an upper-bound-constrained estimate, not a kernel measurement.
41
+ - `tokens_per_second` — *derived*. Output tokens divided by E2E seconds for one request.
42
+ - `decode_tokens_per_second` — *derived*. Output tokens after the first token divided by generation seconds.
43
+ - `output token throughput` — *derived*. All successful output tokens for a model divided by that model's measured wall-clock window.
44
+ - `total token throughput` — *derived*. Successful prompt and output tokens divided by the model's measured wall-clock window.
45
+ - `RPS` — *measured*. Successful requests divided by the model's measured wall-clock window.
46
+ - `goodput` — *derived*. Successful requests that also satisfy every provided SLO threshold. A run whose SLO metric is unavailable counts as not good, so an unverifiable SLO can never inflate goodput.
47
+ - `goodput RPS` — *derived*. SLO-passing requests divided by the model's measured wall-clock window. Zero is reported when SLOs are set and nothing passes.
48
+ - `repeatability` — *derived*. For repeated runs of the same prompt, the average share of runs that produced the most common normalized output hash.
49
+ - `CV` — *derived*. Sample standard deviation divided by the mean. Reported only with two or more successful samples.
50
+ - `95% CI` — *derived*. t-distribution confidence interval around the sample mean, using exact t critical values up to 30 degrees of freedom and the standard 2.0 / 1.96 approximations beyond that. Reported only with two or more successful samples. Treat it as context, not proof, at small sample sizes.
51
+ - `quota` — *measured*, when the `fm` build exposes a quota command. Most current builds do not, in which case `fm-bench` reports quota as unavailable rather than empty.
52
+
53
+ ### Small samples
54
+
55
+ Statistics are computed over the successful samples of one model/concurrency row. With zero samples the row reports `null`; with one sample, percentiles and the mean are reported while spread metrics (`stddev`, `cv`, `ci95Low`, `ci95High`) are `null`. A single run cannot demonstrate stability, so `fm-bench` does not print `0%` variation for it.
38
56
 
39
57
  ## Operating Points
40
58
 
@@ -44,20 +62,28 @@ Use `--request-rate <rps>` to pace request starts independently of concurrency.
44
62
 
45
63
  `fm-bench` does not interpolate between operating points. It reports only what was actually measured.
46
64
 
65
+ ## Warmups
66
+
67
+ `--warmup <n>` runs `n` unmeasured calls per model at the start of each operating point. The first `fm respond` after a cold start includes model load time, which can dominate a short prompt (observed at several hundred milliseconds on Apple silicon). Warmups are never mixed into the measured results.
68
+
69
+ ## Failures and retries
70
+
71
+ A failed call is recorded as a failed measurement with its error text, and is excluded from latency and throughput statistics. `--retry <n>` retries failed calls with exponential backoff before recording the failure; the report keeps the run count and records the number of attempts separately, so a retried success is still one sample and never silently duplicates work.
72
+
47
73
  ## Prompt Profiles
48
74
 
49
75
  The `client` profile is a pragmatic local-machine mix inspired by MLPerf Client's emphasis on multiple task categories and prompt/response lengths. It includes short chat, content generation, structured extraction, light summarization, and code analysis prompts. It is not a formal MLPerf submission suite; it is a convenient built-in workload for comparing your own Mac, OS build, and `fm` models over time.
50
76
 
51
77
  ## Caveats
52
78
 
53
- `fm-bench` uses `fm token-count --quiet` as the source of token counts, so token values follow Apple's local tokenizer behavior. It does not judge semantic quality unless you provide your own prompt suite and inspect captured outputs with `--capture-output`.
79
+ Token counts come from the `fm` build's own token-counting command (`count-tokens` on current builds, `token-count` on older ones). If a build has no such command, token-derived metrics are reported as unavailable rather than estimated. `fm-bench` cannot judge semantic quality unless you provide your own prompt suite and inspect captured outputs with `--capture-output`.
54
80
 
55
81
  Client-side measurements include process startup, local queueing, model prefill, streaming, detokenization, and terminal pipe overhead. That is intentional for a command-line benchmark, but it is not the same as an internal model-kernel benchmark.
56
82
 
57
- Stream smoothness metrics use stdout chunk arrival times. A chunk can contain more than one token, and terminal or pipe buffering can affect chunk boundaries. Treat `second_chunk_ms` and `chunk_gap` as user-visible streaming diagnostics, not raw decoder telemetry.
83
+ Stream smoothness metrics use stdout chunk arrival times. A chunk can contain more than one token, and pipe buffering can affect chunk boundaries. Treat `second_chunk_ms` and `chunk_gap` as user-visible streaming diagnostics, not raw decoder telemetry.
58
84
 
59
85
  For serious comparisons, prefer at least three runs per prompt, include warmups, benchmark both interactive and throughput or client profiles, compare models at the same concurrency operating points, set SLOs that match your real UX budget, and save JSON reports for later analysis.
60
86
 
61
87
  ## Report artifacts
62
88
 
63
- Saved JSON includes client-side environment metadata (hardware model, macOS build, `fm` help digest, power/thermal snapshot) so shared results remain interpretable on other machines. Use `fm-bench validate` before publishing and `fm-bench compare --strict` when you require identical prompt suites. Format details: [report-format.md](./report-format.md).
89
+ Saved JSON includes client-side environment metadata (hardware model, macOS build, `fm` help digest, power/thermal snapshot) plus the detected `fm` capabilities and per-metric availability, so shared results remain interpretable on other machines. Use `fm-bench validate` before publishing and `fm-bench compare --strict` when you require identical prompt suites. Format details: [report-format.md](./report-format.md).
package/docs/releasing.md CHANGED
@@ -1,12 +1,23 @@
1
1
  # Releasing
2
2
 
3
- `fm-bench` uses semver tags (`v*.*.*`). Pushing a tag runs the **Release** workflow: test, lint, npm publish (with provenance), and a GitHub release whose body is taken from the matching section in **`CHANGELOG.md`** (plus a compare link).
3
+ `fm-bench` uses semver tags (`v*.*.*`). Pushing a tag runs the **Release** workflow: lint/test/package check, npm publish (with provenance), and a GitHub release whose body is taken from the matching section in **`CHANGELOG.md`** (plus a compare link).
4
4
 
5
5
  ## Prerequisites
6
6
 
7
7
  - Repository secret **`NPM_TOKEN`**: npm automation token with publish access to `fm-bench`.
8
8
  - **`main`** is green on CI.
9
9
 
10
+ If `NPM_TOKEN` is missing or expired, the Release workflow warns, skips the npm publish step, and still creates the GitHub release. The release notes then state that npm publish was skipped. Re-run the Release workflow with the tag once the token is fixed — an already-published version is skipped automatically.
11
+
12
+ ## Before tagging
13
+
14
+ ```sh
15
+ npm run check # lint + tests + npm pack integrity
16
+ npm run publish:dry-run
17
+ ```
18
+
19
+ Then bump `CHANGELOG.md` with a `## X.Y.Z` section; the release notes come from that section.
20
+
10
21
  ## Option A — GitHub Actions (recommended)
11
22
 
12
23
  1. Open **Actions → Version → Run workflow**.
@@ -17,20 +28,34 @@
17
28
  ## Option B — Local
18
29
 
19
30
  ```sh
20
- npm ci && npm test && npm run lint && npm run check:pack
21
- npm version patch # or minor / major
22
- git push --follow-tags
31
+ npm ci && npm run check
32
+ npm version minor # or patch / major
33
+ git push origin main --follow-tags
23
34
  ```
24
35
 
25
36
  ## Re-run Release without republishing
26
37
 
27
- If npm already has the version but GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.5.3`). The workflow skips npm publish when that version is already on the registry. If the GitHub release already exists, it **updates the release notes** from `CHANGELOG.md`.
38
+ If npm already has the version but the GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.6.3`). The workflow skips npm publish when that version is already on the registry. If the GitHub release already exists, it **updates the release notes** from `CHANGELOG.md`.
28
39
 
29
40
  Refresh notes locally without re-publishing:
30
41
 
31
42
  ```sh
32
- node scripts/changelog-release-notes.mjs 0.6.0 > /tmp/notes.md
33
- gh release edit v0.6.0 --notes-file /tmp/notes.md --repo devinoldenburg/fm-bench
43
+ node scripts/changelog-release-notes.mjs 0.6.3 > notes.md
44
+ gh release edit v0.6.3 --notes-file notes.md --repo devinoldenburg/fm-bench
45
+ ```
46
+
47
+ ## Verifying a release
48
+
49
+ ```sh
50
+ git tag --list 'v0.7.0'
51
+ gh release view v0.7.0 --repo devinoldenburg/fm-bench
52
+ npm view fm-bench version # registry version
53
+
54
+ mkdir -p fm-bench-check && cd fm-bench-check
55
+ npm install fm-bench@0.7.0
56
+ ./node_modules/.bin/fm-bench --version
57
+ ./node_modules/.bin/fm-bench --help > /dev/null
58
+ ./node_modules/.bin/fm-bench doctor
34
59
  ```
35
60
 
36
61
  ## Dry run
@@ -40,4 +65,4 @@ npm run publish:dry-run
40
65
  npm run check:pack
41
66
  ```
42
67
 
43
- CI runs the same dry-run on every push to `main` and on pull requests.
68
+ CI runs the dry run on every push to `main` and on pull requests.