fm-bench 0.6.3 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +86 -68
- package/bin/fm-bench.js +14 -0
- package/docs/compatibility.md +46 -0
- package/docs/methodology.md +46 -20
- package/docs/releasing.md +33 -8
- package/docs/report-format.md +79 -2
- package/docs/supported-platforms.md +19 -2
- package/package.json +5 -4
- package/src/bench.js +129 -37
- package/src/capabilities.js +196 -0
- package/src/cli.js +153 -68
- package/src/compare.js +11 -3
- package/src/fm-help.js +131 -0
- package/src/fm.js +110 -106
- package/src/history.js +2 -1
- package/src/macos.js +2 -1
- package/src/metrics.js +212 -0
- package/src/process.js +39 -9
- package/src/prompts.js +22 -4
- package/src/report.js +13 -2
- package/src/schema.js +17 -1
- package/src/stats.js +43 -14
- package/src/table.js +124 -50
package/README.md
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
# fm-bench
|
|
2
2
|
|
|
3
3
|
[](https://github.com/devinoldenburg/fm-bench/actions/workflows/ci.yml)
|
|
4
|
+
[](https://www.npmjs.com/package/fm-bench)
|
|
4
5
|
|
|
5
6
|
Benchmark Apple's `fm` command on macOS 27+.
|
|
6
7
|
|
|
@@ -8,7 +9,7 @@ Measure latency, throughput, streaming smoothness, stability, and goodput across
|
|
|
8
9
|
|
|
9
10
|
## Why fm-bench exists
|
|
10
11
|
|
|
11
|
-
Apple's Foundation Models
|
|
12
|
+
Apple's Foundation Models run on-device, and the `fm` CLI is the terminal and script interface to them. That makes raw model quality only half the story.
|
|
12
13
|
|
|
13
14
|
For real apps, the important questions are:
|
|
14
15
|
|
|
@@ -16,10 +17,12 @@ For real apps, the important questions are:
|
|
|
16
17
|
- Does streaming stay smooth?
|
|
17
18
|
- How stable is latency over repeated runs?
|
|
18
19
|
- What happens under concurrency?
|
|
19
|
-
-
|
|
20
|
+
- Does this Mac meet an interactive SLO?
|
|
20
21
|
|
|
21
22
|
`fm-bench` answers those questions with repeatable local benchmarks. Think of it as GeekBench for Apple Foundation Models — run it, get numbers, compare across hardware, models, and macOS updates.
|
|
22
23
|
|
|
24
|
+
It is deliberately honest about what it can measure: the installed `fm` is probed for its actual capabilities, and any metric that build cannot supply is reported as unavailable rather than estimated. See [docs/compatibility.md](docs/compatibility.md).
|
|
25
|
+
|
|
23
26
|
## Quick Start
|
|
24
27
|
|
|
25
28
|
```sh
|
|
@@ -27,22 +30,30 @@ npm install -g fm-bench
|
|
|
27
30
|
fm-bench
|
|
28
31
|
```
|
|
29
32
|
|
|
30
|
-
One command discovers your models, runs the standard prompt suite, and prints a full benchmark report:
|
|
33
|
+
One command discovers your models, runs the standard prompt suite, and prints a full benchmark report (real output from a macOS 27.0 Apple M5 Pro):
|
|
31
34
|
|
|
32
35
|
```text
|
|
33
|
-
fm-bench 0.
|
|
34
|
-
prompts 3 | runs
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
│
|
|
40
|
-
|
|
36
|
+
fm-bench 0.7.0 | darwin/arm64 | fm
|
|
37
|
+
prompts 3 | runs 3 | concurrency 1 | stream on | measured 9 | failed 0 | skipped 0 | elapsed 6.37s | SLO
|
|
38
|
+
TTFT<=1.00s,E2E<=3.00s
|
|
39
|
+
unavailable: quota unavailable — this fm build exposes no quota command
|
|
40
|
+
|
|
41
|
+
┌───┬────────┬────────┬─────┬──────┬──────────┬───────┬───────┬─────────┬────────┬───────┬─────┬──────┐
|
|
42
|
+
│ C │ MODEL │ STATUS │ OK │ GOOD │ GOOD RPS │ TTFT │ E2E │ E2E P95 │ USER/S │ SYS/S │ CV │ NOTE │
|
|
43
|
+
├───┼────────┼────────┼─────┼──────┼──────────┼───────┼───────┼─────────┼────────┼───────┼─────┼──────┤
|
|
44
|
+
│ 1 │ system │ ok │ 9/9 │ 100% │ 1.7 │ 377ms │ 487ms │ 749ms │ 34.0 │ 32.8 │ 25% │ │
|
|
45
|
+
└───┴────────┴────────┴─────┴──────┴──────────┴───────┴───────┴─────────┴────────┴───────┴─────┴──────┘
|
|
46
|
+
|
|
47
|
+
┌───┬────────┬────────┬─────────┬───────────┬──────────┬───────────┬─────────┬────────┐
|
|
48
|
+
│ C │ MODEL │ IN AVG │ OUT AVG │ PREFILL/S │ DECODE/S │ CHUNK P95 │ E2E P99 │ REPEAT │
|
|
49
|
+
├───┼────────┼────────┼─────────┼───────────┼──────────┼───────────┼─────────┼────────┤
|
|
50
|
+
│ 1 │ system │ 13 │ 20 │ 35.9 │ 113 │ 197ms │ 749ms │ 100% │
|
|
51
|
+
└───┴────────┴────────┴─────────┴───────────┴──────────┴───────────┴─────────┴────────┘
|
|
41
52
|
```
|
|
42
53
|
|
|
43
54
|
Default profile is `standard` (3 prompts). Use `--runs 5`, `--sweep-concurrency 1,2`, or `--profile client` for heavier suites.
|
|
44
55
|
|
|
45
|
-
Wide terminals add TTFT P95, TPOT,
|
|
56
|
+
Wide terminals add TTFT P95, TPOT, and RPS columns; medium terminals tighten the table; narrow terminals switch to compact model cards automatically. `--width <n>` previews any layout.
|
|
46
57
|
|
|
47
58
|
## Install
|
|
48
59
|
|
|
@@ -70,13 +81,16 @@ fm-bench doctor # verify your setup
|
|
|
70
81
|
| Command | What it does |
|
|
71
82
|
|---------|-------------|
|
|
72
83
|
| `fm-bench` | Run the full benchmark (default) |
|
|
73
|
-
| `fm-bench models` |
|
|
74
|
-
| `fm-bench compare <a.json> <b.json>` | Regression diff with suite/hardware warnings; `--strict`
|
|
84
|
+
| `fm-bench models` | Show the detected `fm` capabilities, discovered models, availability, and quota when supported |
|
|
85
|
+
| `fm-bench compare <a.json> <b.json>` | Regression diff with suite/hardware/macOS warnings; `--strict` exits 2 when suites differ |
|
|
75
86
|
| `fm-bench history [dir]` | Trend table from saved reports (sorted by time, tags visible) |
|
|
76
|
-
| `fm-bench validate <report.json>` | Verify report JSON (schema v1) before sharing |
|
|
87
|
+
| `fm-bench validate <report.json>` | Verify report JSON (schema v1) before sharing; `--json` for CI |
|
|
77
88
|
| `fm-bench export <report.json>` | Standalone HTML report with embedded JSON |
|
|
78
|
-
| `fm-bench legend` | Definitions
|
|
79
|
-
| `fm-bench doctor` | Environment
|
|
89
|
+
| `fm-bench legend` | Definitions, provenance (measured/proxy/derived), and color rules for every column |
|
|
90
|
+
| `fm-bench doctor` | Environment and `fm` capability check; `--json` for scripts |
|
|
91
|
+
| `fm-bench metrics` | Alias for `legend` |
|
|
92
|
+
|
|
93
|
+
Exit codes: `0` success, `1` operational failure (failed `--ci` gate, invalid reports), `2` usage or environment error (bad flags, unsupported macOS, unusable `fm`, no runnable model). Interrupting a run with Ctrl+C terminates in-flight `fm` processes.
|
|
80
94
|
|
|
81
95
|
## Common Recipes
|
|
82
96
|
|
|
@@ -84,15 +98,12 @@ fm-bench doctor # verify your setup
|
|
|
84
98
|
# Quick smoke test
|
|
85
99
|
fm-bench --profile quick
|
|
86
100
|
|
|
87
|
-
# Standard 5-run benchmark with SLO budgets
|
|
88
|
-
fm-bench --runs 5 --slo-ttft-ms 750 --slo-e2e-ms 4000
|
|
101
|
+
# Standard 5-run benchmark with SLO budgets and warmups
|
|
102
|
+
fm-bench --runs 5 --warmup 1 --slo-ttft-ms 750 --slo-e2e-ms 4000
|
|
89
103
|
|
|
90
104
|
# Sweep concurrency to find your throughput ceiling
|
|
91
105
|
fm-bench --sweep-concurrency 1,2,4 --runs 3
|
|
92
106
|
|
|
93
|
-
# Stress both on-device and PCC models
|
|
94
|
-
fm-bench --models system,pcc --runs 3 --profile stress
|
|
95
|
-
|
|
96
107
|
# Reasoning and coding workloads
|
|
97
108
|
fm-bench --profile reasoning --runs 5
|
|
98
109
|
fm-bench --profile coding --runs 3 --histogram
|
|
@@ -103,7 +114,7 @@ fm-bench --output-dir reports/ --tag after-update --export-html
|
|
|
103
114
|
fm-bench validate reports/*.json
|
|
104
115
|
fm-bench compare reports/fm-bench_*before*.json reports/fm-bench_*after*.json --strict
|
|
105
116
|
|
|
106
|
-
# Fail CI when SLOs regress
|
|
117
|
+
# Fail CI when SLOs regress or any run fails
|
|
107
118
|
fm-bench --ci --slo-ttft-ms 750 --slo-e2e-ms 4000 --runs 5
|
|
108
119
|
|
|
109
120
|
# Save JSON for automation
|
|
@@ -117,9 +128,9 @@ fm-bench --format csv --out bench.csv
|
|
|
117
128
|
|
|
118
129
|
| Flag | Default | Description |
|
|
119
130
|
|------|---------|-------------|
|
|
120
|
-
| `-m, --models <list>` |
|
|
131
|
+
| `-m, --models <list>` | discovered | Comma-separated or repeated model names |
|
|
121
132
|
| `-r, --runs <n>` | 1 | Measured runs per prompt/model |
|
|
122
|
-
| `--warmup <n>` | 0 |
|
|
133
|
+
| `--warmup <n>` | 0 | Unmeasured warmup runs per model before measurement |
|
|
123
134
|
| `-c, --concurrency <n>` | 1 | Parallel `fm` processes |
|
|
124
135
|
| `--sweep-concurrency <list>` | — | Separate operating points, e.g. `1,2,4` |
|
|
125
136
|
| `--request-rate <rps>` | — | Pace request starts at a target rate |
|
|
@@ -129,7 +140,9 @@ fm-bench --format csv --out bench.csv
|
|
|
129
140
|
| `--profile <name>` | standard | Built-in prompt suite (see Profiles) |
|
|
130
141
|
| `-p, --prompt <text>` | — | Custom prompt, repeatable |
|
|
131
142
|
| `--prompt-file <file>` | — | JSON, JSONL, or blank-line separated prompts |
|
|
132
|
-
| `-i, --instructions <text>` | — | Passed to `fm respond` |
|
|
143
|
+
| `-i, --instructions <text>` | — | Passed to `fm respond` when the build supports it |
|
|
144
|
+
| `--use-case <case>` | — | System model use case, when supported |
|
|
145
|
+
| `--guardrails <level>` | — | System model guardrail level, when supported |
|
|
133
146
|
|
|
134
147
|
**Quality Gates**
|
|
135
148
|
|
|
@@ -159,11 +172,15 @@ fm-bench --format csv --out bench.csv
|
|
|
159
172
|
|
|
160
173
|
| Flag | Description |
|
|
161
174
|
|------|-------------|
|
|
175
|
+
| `--no-stream` | Disable streaming; TTFT and decode metrics are reported as unavailable |
|
|
176
|
+
| `--greedy` / `--no-greedy` | Request or omit greedy sampling (default: greedy) |
|
|
177
|
+
| `--available-only` | Hide unavailable discovered models |
|
|
162
178
|
| `--color` / `--no-color` | Force or disable ANSI colors (auto on TTYs) |
|
|
163
179
|
| `--ascii` | Plain ASCII table borders instead of Unicode |
|
|
164
180
|
| `--compact` | Force narrow terminal layout |
|
|
165
181
|
| `--width <n>` | Render as if the terminal is `n` columns wide |
|
|
166
182
|
| `--progress` / `--no-progress` | Force or disable the live progress line |
|
|
183
|
+
| `--fm-bin <path>` | `fm` binary to execute (default: `FM_BIN` or `fm`) |
|
|
167
184
|
|
|
168
185
|
## Prompt Profiles
|
|
169
186
|
|
|
@@ -181,9 +198,38 @@ Nine built-in suites, choose the one that matches your use case:
|
|
|
181
198
|
| `coding` | 5 | Code review, refactoring, algorithms, system design |
|
|
182
199
|
| `creative` | 5 | Product copy, analogies, commit messages, docs |
|
|
183
200
|
|
|
201
|
+
## Metrics
|
|
202
|
+
|
|
203
|
+
**Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT, 95% confidence interval, coefficient of variation (CV).
|
|
204
|
+
|
|
205
|
+
**Throughput** — prefill tokens/s, decode tokens/s, output tokens/s per request, aggregate system tokens/s, requests per second.
|
|
206
|
+
|
|
207
|
+
**Streaming quality** — second-chunk delay, chunk-gap p95, captured from stdout chunk arrival during streaming runs.
|
|
208
|
+
|
|
209
|
+
**Reliability** — success rate, goodput rate and RPS against SLO budgets, repeatability (most common output hash frequency across repeated runs).
|
|
210
|
+
|
|
211
|
+
Every metric is labelled by provenance in JSON and in `fm-bench legend`:
|
|
212
|
+
|
|
213
|
+
- **measured** — process wall clock, exit codes, chunk arrival, `fm` token counts
|
|
214
|
+
- **proxy** — TTFT and chunk gaps (chunk granularity, not token timestamps); prefill tokens/s
|
|
215
|
+
- **derived** — TPOT, decode tokens/s, throughput, CV, confidence intervals, goodput
|
|
216
|
+
|
|
217
|
+
Token counts come from the `fm` build's own token-counting command (`count-tokens`, or `token-count` on older builds). If the build cannot count tokens, token metrics render as `-`, JSON carries `null`, and `metrics.promptTokens.available` is `false`. Spread statistics (`CV`, 95% CI) need at least two successful samples; with one sample they are unavailable rather than `0`. See [docs/methodology.md](docs/methodology.md).
|
|
218
|
+
|
|
219
|
+
## fm compatibility
|
|
220
|
+
|
|
221
|
+
`fm-bench` probes `fm --help` and `fm respond --help` once per run and adapts:
|
|
222
|
+
|
|
223
|
+
- Subcommand names (`count-tokens` vs legacy `token-count`) are detected, not hardcoded.
|
|
224
|
+
- Flags the build does not document (`--stream`, `--use-case`, `--guardrails`, `--model`) are not passed through.
|
|
225
|
+
- Unsupported models are rejected before a benchmark starts, with the supported list in the error.
|
|
226
|
+
- Raw `fm` argument-error text never reaches reports, tables, or JSON.
|
|
227
|
+
|
|
228
|
+
`fm-bench doctor` shows exactly what was detected. Policy and the verified-build table: [docs/compatibility.md](docs/compatibility.md).
|
|
229
|
+
|
|
184
230
|
## Sharing and comparing results
|
|
185
231
|
|
|
186
|
-
Reports from 0.6.0+ include **schema v1**: `reportId`, hardware fingerprint,
|
|
232
|
+
Reports from 0.6.0+ include **schema v1**: `reportId`, hardware fingerprint, a **suite key** so you can tell if two JSON files used the same prompts and run settings, plus the detected `fm` capabilities and per-metric availability. See [docs/report-format.md](docs/report-format.md).
|
|
187
233
|
|
|
188
234
|
- Share **HTML** with teammates who do not use the CLI: `fm-bench export bench.json -o bench.html`
|
|
189
235
|
- Gate uploads in CI: `fm-bench validate artifact.json`
|
|
@@ -194,33 +240,24 @@ Reports from 0.6.0+ include **schema v1**: `reportId`, hardware fingerprint, and
|
|
|
194
240
|
Track performance across macOS updates, model changes, or hardware swaps:
|
|
195
241
|
|
|
196
242
|
```sh
|
|
197
|
-
# Before
|
|
198
243
|
fm-bench --profile coding --runs 5 --output-dir reports/ --tag before
|
|
199
|
-
|
|
200
|
-
# After the change
|
|
201
244
|
fm-bench --profile coding --runs 5 --output-dir reports/ --tag after
|
|
202
|
-
|
|
203
|
-
# See what changed
|
|
204
245
|
fm-bench compare reports/fm-bench_*before*.json reports/fm-bench_*after*.json
|
|
205
|
-
```
|
|
206
|
-
|
|
207
|
-
The compare output shows each model/concurrency row with the before value, a color-coded percent delta (green = improvement, red = regression), and the after value — for TTFT, E2E, TPOT, tokens/s, RPS, success rate, and CV.
|
|
208
|
-
|
|
209
|
-
```sh
|
|
210
|
-
# View the full trend over time
|
|
211
246
|
fm-bench history reports/
|
|
212
247
|
```
|
|
213
248
|
|
|
249
|
+
The compare output shows each model/concurrency row with the before value, a color-coded percent delta (green = improvement, red = regression), and the after value — for TTFT, E2E, TPOT, tokens/s, RPS, success rate, and CV. It warns when hardware, macOS version/build, or the prompt suite differ.
|
|
250
|
+
|
|
214
251
|
## CI Integration
|
|
215
252
|
|
|
216
253
|
Gate deployments or model updates on benchmark quality:
|
|
217
254
|
|
|
218
255
|
```sh
|
|
219
|
-
# Fails with exit code 1 if TTFT > 750ms
|
|
256
|
+
# Fails with exit code 1 if TTFT > 750ms, E2E > 4s, or any run fails
|
|
220
257
|
fm-bench --ci --slo-ttft-ms 750 --slo-e2e-ms 4000 --runs 5
|
|
221
258
|
```
|
|
222
259
|
|
|
223
|
-
Prints `fm-bench ci: PASS` or `fm-bench ci: FAIL — <reason>` to stderr. Designed for GitHub Actions, Buildkite, or any shell-based pipeline.
|
|
260
|
+
Prints `fm-bench ci: PASS` or `fm-bench ci: FAIL — <reason>` to stderr. Designed for GitHub Actions, Buildkite, or any shell-based pipeline. `--json` and `--csv` write only data to stdout; progress and diagnostics go to stderr.
|
|
224
261
|
|
|
225
262
|
## Prompt Files
|
|
226
263
|
|
|
@@ -240,36 +277,18 @@ JSONL:
|
|
|
240
277
|
{"id":"latency","prompt":"Explain p95 latency in one sentence."}
|
|
241
278
|
```
|
|
242
279
|
|
|
243
|
-
Plain text files are split on blank lines.
|
|
244
|
-
|
|
245
|
-
## Metrics
|
|
246
|
-
|
|
247
|
-
**Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT (p50/p95), 95% confidence interval, coefficient of variation (CV).
|
|
248
|
-
|
|
249
|
-
**Throughput** — prefill tokens/s, decode tokens/s, output tokens/s per request, aggregate system tokens/s, requests per second.
|
|
250
|
-
|
|
251
|
-
**Streaming quality** — second-chunk delay, chunk-gap p95. Captured from `stdout` chunks during streaming runs.
|
|
252
|
-
|
|
253
|
-
**Reliability** — success rate, goodput rate and RPS against SLO budgets, repeatability (most common output hash frequency across repeated runs).
|
|
254
|
-
|
|
255
|
-
**Stability** — CV (stddev/mean for E2E latency); green ≤10%, yellow ≤25%, red >25%.
|
|
256
|
-
|
|
257
|
-
Token counts come from `fm token-count --quiet`. If `fm` cannot count tokens, those fields are blank while character throughput is still reported.
|
|
258
|
-
|
|
259
|
-
Terminal layout is responsive: wide → full scoreboard + detail tables, medium → tighter single table, narrow → compact model cards. Use `--width` to preview any layout and `--ascii` for log-friendly output.
|
|
280
|
+
Plain text files are split on blank lines. Parse errors name the file and, for JSONL, the line number.
|
|
260
281
|
|
|
261
282
|
## Colors and Legend
|
|
262
283
|
|
|
263
284
|
Table output is color-coded on interactive terminals — **green** is better/passing, **yellow** is marginal/partial, **red** is failing/unstable. Fixed thresholds apply to success rate, goodput, CV, and repeatability. Latency uses SLO thresholds when set, otherwise lower-is-better relative ranking. Throughput uses higher-is-better relative ranking.
|
|
264
285
|
|
|
265
286
|
```sh
|
|
266
|
-
fm-bench legend #
|
|
287
|
+
fm-bench legend # definitions, provenance, and color rules
|
|
267
288
|
fm-bench legend --json # machine-readable
|
|
268
289
|
```
|
|
269
290
|
|
|
270
|
-
`NO_COLOR=1` disables color; `FORCE_COLOR=1` or `--color` enables it. `--ascii` switches to plain ASCII borders for log systems.
|
|
271
|
-
|
|
272
|
-
A live single-line progress indicator runs on stderr during interactive sessions. The final report always goes to stdout — `--json`, `--csv`, and `--out` stay automation-friendly.
|
|
291
|
+
`NO_COLOR=1` disables color; `FORCE_COLOR=1` or `--color` enables it. `--ascii` switches to plain ASCII borders for log systems. A live single-line progress indicator runs on stderr during interactive sessions; the final report always goes to stdout.
|
|
273
292
|
|
|
274
293
|
## Requirements
|
|
275
294
|
|
|
@@ -279,19 +298,18 @@ A live single-line progress indicator runs on stderr during interactive sessions
|
|
|
279
298
|
|
|
280
299
|
Benchmark commands refuse to start on older macOS versions and report the detected version plus the latest supported macOS — see [docs/supported-platforms.md](docs/supported-platforms.md).
|
|
281
300
|
|
|
282
|
-
`pcc` (Private Cloud Compute) availability depends on Apple's current eligibility. `fm-bench` shows it as skipped if `fm available --model pcc` reports unavailable.
|
|
283
|
-
|
|
284
|
-
See [docs/methodology.md](docs/methodology.md) for benchmark methodology and metric references.
|
|
285
|
-
|
|
286
301
|
## Development
|
|
287
302
|
|
|
288
303
|
```sh
|
|
289
304
|
npm install
|
|
290
|
-
npm test
|
|
291
|
-
npm run lint
|
|
305
|
+
npm test # node --test (unit + integration against a fake fm)
|
|
306
|
+
npm run lint # node --check on every source file
|
|
307
|
+
npm run check # lint + tests + npm pack integrity — run this before pushing
|
|
292
308
|
```
|
|
293
309
|
|
|
294
|
-
No runtime npm dependencies.
|
|
310
|
+
No runtime npm dependencies. Integration tests drive the real CLI against `test/fixtures/fake-fm.mjs`, which emulates normal, streaming, slow, malformed, failing, timed-out, partial, and interrupted `fm` behaviour, so the suite runs on machines without `fm`.
|
|
311
|
+
|
|
312
|
+
Documentation: [methodology](docs/methodology.md) · [report format](docs/report-format.md) · [compatibility](docs/compatibility.md) · [supported platforms](docs/supported-platforms.md) · [releasing](docs/releasing.md).
|
|
295
313
|
|
|
296
314
|
## License
|
|
297
315
|
|
package/bin/fm-bench.js
CHANGED
|
@@ -1,5 +1,19 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
2
|
import { runCli } from '../src/cli.js';
|
|
3
|
+
import { killActiveChildren } from '../src/process.js';
|
|
4
|
+
|
|
5
|
+
// Terminate any in-flight `fm` child process before exiting, then use the
|
|
6
|
+
// conventional 128 + signal exit code. A second signal exits immediately.
|
|
7
|
+
let interrupting = false;
|
|
8
|
+
for (const [signal, exitCode] of [['SIGINT', 130], ['SIGTERM', 143]]) {
|
|
9
|
+
process.on(signal, () => {
|
|
10
|
+
const killed = killActiveChildren(signal);
|
|
11
|
+
if (interrupting) process.exit(exitCode);
|
|
12
|
+
interrupting = true;
|
|
13
|
+
const graceMs = killed > 0 ? 500 : 0;
|
|
14
|
+
setTimeout(() => process.exit(exitCode), graceMs).unref();
|
|
15
|
+
});
|
|
16
|
+
}
|
|
3
17
|
|
|
4
18
|
runCli(process.argv.slice(2)).catch((error) => {
|
|
5
19
|
const message = error?.message || String(error);
|
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
# `fm` compatibility
|
|
2
|
+
|
|
3
|
+
`fm-bench` is a client of whatever `fm` binary is installed. Apple has already changed that surface between macOS 27 builds: current builds expose `count-tokens`, older ones exposed `token-count`, and quota reporting is not available everywhere. fm-bench therefore detects capabilities at runtime instead of assuming a fixed subcommand list.
|
|
4
|
+
|
|
5
|
+
## How detection works
|
|
6
|
+
|
|
7
|
+
Once per run, `fm-bench` spawns two cheap calls:
|
|
8
|
+
|
|
9
|
+
1. `fm --help` — the command list, the `MODELS` section, and (when present) a `--model` option list.
|
|
10
|
+
2. `fm respond --help` — which flags `respond` actually accepts (`--model`, `--[no-]stream`, `--instructions`, `--greedy`, `--use-case`, `--guardrails`, `--image`, `--tool`, `--schema`).
|
|
11
|
+
|
|
12
|
+
The result is recorded in the report (`capabilities`) and shown by `fm-bench models`, `fm-bench doctor`, and `fm-bench doctor --json`.
|
|
13
|
+
|
|
14
|
+
Detection is tolerant of formatting changes: section headers are matched case-insensitively on uppercase lines, model lines by their column layout, and boolean flags in either `--flag` or Apple's `--[no-]flag` spelling. ANSI escapes are stripped before parsing.
|
|
15
|
+
|
|
16
|
+
## Capability policy
|
|
17
|
+
|
|
18
|
+
| Capability | When missing |
|
|
19
|
+
|------------|--------------|
|
|
20
|
+
| Token counting (`count-tokens`, fallback `token-count`) | The run still benchmarks latency and streaming. Every token metric is `null` with `available: false` in `metrics`, the table prints an `unavailable:` note, and the report carries a warning. `fm count-tokens` is never called. |
|
|
21
|
+
| Streaming | TTFT, generation time, TPOT, prefill, and chunk-gap metrics are unavailable, and `--stream`/`--no-stream` is not passed through. E2E latency, success rate, and RPS still work. |
|
|
22
|
+
| Quota command | The `models` table omits the quota column and reports quota as unavailable. |
|
|
23
|
+
| Model selection (`--model`) | The flag is not passed; `fm` uses its default model. |
|
|
24
|
+
| `--use-case`, `--guardrails`, `--instructions`, `--greedy` | Silently not passed, instead of failing the run with an argument error. |
|
|
25
|
+
| No usable `fm` at all | The command exits `2` with the spawn error and a hint to install `fm` or point `--fm-bin` / `FM_BIN` at a compatible binary. |
|
|
26
|
+
|
|
27
|
+
Models the build does not list are refused before any benchmark starts, with the message `not supported by this fm build (supported: ...)`. Raw `fm` argument-error text (usage blocks, `Unknown command`) never reaches reports, tables, or JSON output.
|
|
28
|
+
|
|
29
|
+
## Exit codes
|
|
30
|
+
|
|
31
|
+
| Code | Meaning |
|
|
32
|
+
|------|---------|
|
|
33
|
+
| `0` | Success. A failed measured run is data, not a CLI failure — use `--ci` to make failures fatal. |
|
|
34
|
+
| `1` | Operational failure: `--ci` gate failed, reports failed validation, or an `fm` run could not produce data. |
|
|
35
|
+
| `2` | Usage or environment error: unknown flag, missing argument, unsupported macOS, unusable `fm`, malformed report input, or `compare --strict` suite mismatch. |
|
|
36
|
+
| `130` / `143` | Interrupted by SIGINT / SIGTERM. In-flight `fm` child processes are terminated first. |
|
|
37
|
+
|
|
38
|
+
The macOS 27 requirement is enforced only when `fm-bench` resolves the default `fm` from `PATH`. Supplying `--fm-bin <path>` or `FM_BIN` lets the CLI run on any host and instead fails on the binary's actual capabilities (exit `2` when it exposes no usable commands).
|
|
39
|
+
|
|
40
|
+
## Verified builds
|
|
41
|
+
|
|
42
|
+
| Platform | `fm` surface | Notes |
|
|
43
|
+
|----------|--------------|-------|
|
|
44
|
+
| macOS 27.0 (26A5425a), Apple M5 Pro (Mac17,9) | `available`, `chat`, `count-tokens`, `license`, `respond`, `schema`, `serve`; model `system` only; no quota command; no `--version` | Real-machine smoke tests in 0.7.0. `test/fixtures/fm-help-macos27.txt` captures this help output as a regression baseline. |
|
|
45
|
+
|
|
46
|
+
If your `fm` build differs, `fm-bench doctor` shows exactly what was detected, and the report's `capabilities` block records it. Please open an issue with the `doctor --json` output and the report digest when something is missing.
|
package/docs/methodology.md
CHANGED
|
@@ -14,27 +14,45 @@ The metric set follows common LLM inference benchmark practice:
|
|
|
14
14
|
- MLCommons describes varying concurrency and reporting verified operating points for TTFT, throughput, interactivity, and response latency rather than interpolated performance: <https://mlcommons.org/2026/03/mlperf-endpoints-gen-ai-benchmarking/>
|
|
15
15
|
- MLPerf Client emphasizes local client workloads with multiple task types and varying prompt/response lengths: <https://mlcommons.org/benchmarks/client/>
|
|
16
16
|
|
|
17
|
+
## What `fm` can actually tell us
|
|
18
|
+
|
|
19
|
+
`fm-bench` never invents precision the CLI cannot provide. Every metric is classified by how it is obtained:
|
|
20
|
+
|
|
21
|
+
| Kind | Meaning |
|
|
22
|
+
|------|---------|
|
|
23
|
+
| **measured** | Observed directly: process wall clock, exit codes, stdout chunk arrival, `fm`-reported token counts. |
|
|
24
|
+
| **proxy** | Observed at a coarser granularity than the ideal metric (for example chunk arrival instead of token timestamps). |
|
|
25
|
+
| **derived** | Computed from measured values (for example output tokens divided by elapsed seconds). |
|
|
26
|
+
| **controlled** | An input setting rather than a measurement, such as the concurrency operating point. |
|
|
27
|
+
|
|
28
|
+
The same classification is machine-readable in every JSON report under `metrics` and in the CLI via `fm-bench legend` (SOURCE column). A metric the installed `fm` build cannot support is `available: false` with a reason, and renders as `-` rather than a plausible-looking number.
|
|
29
|
+
|
|
17
30
|
## Metrics
|
|
18
31
|
|
|
19
32
|
For a column-by-column terminal reference, run `fm-bench legend`.
|
|
20
33
|
|
|
21
|
-
- `TTFT
|
|
22
|
-
- `E2E latency
|
|
23
|
-
- `generation_ms
|
|
24
|
-
- `TPOT
|
|
25
|
-
- `second_chunk_ms
|
|
26
|
-
- `chunk_gap
|
|
27
|
-
- `prefill_tokens_per_second
|
|
28
|
-
- `tokens_per_second
|
|
29
|
-
- `decode_tokens_per_second
|
|
30
|
-
- `
|
|
31
|
-
- `total token throughput
|
|
32
|
-
- `RPS
|
|
33
|
-
- `goodput
|
|
34
|
-
- `goodput RPS
|
|
35
|
-
- `repeatability
|
|
36
|
-
- `CV
|
|
37
|
-
- `95% CI
|
|
34
|
+
- `TTFT` — *proxy*. Time from starting `fm respond` to the first streamed stdout chunk. `fm` exposes no per-token timestamps, so this is a terminal-side approximation of time to first token; the first chunk may already contain several tokens.
|
|
35
|
+
- `E2E latency` — *measured*. Time from starting `fm respond` until the process exits and the full response is captured.
|
|
36
|
+
- `generation_ms` — *derived*. `E2E - TTFT`, reported only when the answer arrives in more than one stdout chunk. When a single chunk carries the whole answer, prefill and decode are not separable and the value is unavailable instead of near zero.
|
|
37
|
+
- `TPOT` — *derived*. `(E2E - TTFT) / (output_tokens - 1)`, reported when the answer has at least three output tokens so the interval is an average over two or more decode tokens rather than the inverse of a single chunk gap. Requires streaming and a token-counting `fm`.
|
|
38
|
+
- `second_chunk_ms` — *proxy*. Time between the first and second streamed stdout chunks: a terminal-side signal for startup smoothness.
|
|
39
|
+
- `chunk_gap` — *proxy*. Distribution of time between consecutive streamed stdout chunks. Useful for spotting streaming jitter; chunk-based, not token-based.
|
|
40
|
+
- `prefill_tokens_per_second` — *proxy*. Input prompt tokens divided by TTFT seconds. Because TTFT includes process startup and first-token latency, this systematically understates true prefill speed and should be read as an upper-bound-constrained estimate, not a kernel measurement.
|
|
41
|
+
- `tokens_per_second` — *derived*. Output tokens divided by E2E seconds for one request.
|
|
42
|
+
- `decode_tokens_per_second` — *derived*. Output tokens after the first token divided by generation seconds.
|
|
43
|
+
- `output token throughput` — *derived*. All successful output tokens for a model divided by that model's measured wall-clock window.
|
|
44
|
+
- `total token throughput` — *derived*. Successful prompt and output tokens divided by the model's measured wall-clock window.
|
|
45
|
+
- `RPS` — *measured*. Successful requests divided by the model's measured wall-clock window.
|
|
46
|
+
- `goodput` — *derived*. Successful requests that also satisfy every provided SLO threshold. A run whose SLO metric is unavailable counts as not good, so an unverifiable SLO can never inflate goodput.
|
|
47
|
+
- `goodput RPS` — *derived*. SLO-passing requests divided by the model's measured wall-clock window. Zero is reported when SLOs are set and nothing passes.
|
|
48
|
+
- `repeatability` — *derived*. For repeated runs of the same prompt, the average share of runs that produced the most common normalized output hash.
|
|
49
|
+
- `CV` — *derived*. Sample standard deviation divided by the mean. Reported only with two or more successful samples.
|
|
50
|
+
- `95% CI` — *derived*. t-distribution confidence interval around the sample mean, using exact t critical values up to 30 degrees of freedom and the standard 2.0 / 1.96 approximations beyond that. Reported only with two or more successful samples. Treat it as context, not proof, at small sample sizes.
|
|
51
|
+
- `quota` — *measured*, when the `fm` build exposes a quota command. Most current builds do not, in which case `fm-bench` reports quota as unavailable rather than empty.
|
|
52
|
+
|
|
53
|
+
### Small samples
|
|
54
|
+
|
|
55
|
+
Statistics are computed over the successful samples of one model/concurrency row. With zero samples the row reports `null`; with one sample, percentiles and the mean are reported while spread metrics (`stddev`, `cv`, `ci95Low`, `ci95High`) are `null`. A single run cannot demonstrate stability, so `fm-bench` does not print `0%` variation for it.
|
|
38
56
|
|
|
39
57
|
## Operating Points
|
|
40
58
|
|
|
@@ -44,20 +62,28 @@ Use `--request-rate <rps>` to pace request starts independently of concurrency.
|
|
|
44
62
|
|
|
45
63
|
`fm-bench` does not interpolate between operating points. It reports only what was actually measured.
|
|
46
64
|
|
|
65
|
+
## Warmups
|
|
66
|
+
|
|
67
|
+
`--warmup <n>` runs `n` unmeasured calls per model at the start of each operating point. The first `fm respond` after a cold start includes model load time, which can dominate a short prompt (observed at several hundred milliseconds on Apple silicon). Warmups are never mixed into the measured results.
|
|
68
|
+
|
|
69
|
+
## Failures and retries
|
|
70
|
+
|
|
71
|
+
A failed call is recorded as a failed measurement with its error text, and is excluded from latency and throughput statistics. `--retry <n>` retries failed calls with exponential backoff before recording the failure; the report keeps the run count and records the number of attempts separately, so a retried success is still one sample and never silently duplicates work.
|
|
72
|
+
|
|
47
73
|
## Prompt Profiles
|
|
48
74
|
|
|
49
75
|
The `client` profile is a pragmatic local-machine mix inspired by MLPerf Client's emphasis on multiple task categories and prompt/response lengths. It includes short chat, content generation, structured extraction, light summarization, and code analysis prompts. It is not a formal MLPerf submission suite; it is a convenient built-in workload for comparing your own Mac, OS build, and `fm` models over time.
|
|
50
76
|
|
|
51
77
|
## Caveats
|
|
52
78
|
|
|
53
|
-
`fm
|
|
79
|
+
Token counts come from the `fm` build's own token-counting command (`count-tokens` on current builds, `token-count` on older ones). If a build has no such command, token-derived metrics are reported as unavailable rather than estimated. `fm-bench` cannot judge semantic quality unless you provide your own prompt suite and inspect captured outputs with `--capture-output`.
|
|
54
80
|
|
|
55
81
|
Client-side measurements include process startup, local queueing, model prefill, streaming, detokenization, and terminal pipe overhead. That is intentional for a command-line benchmark, but it is not the same as an internal model-kernel benchmark.
|
|
56
82
|
|
|
57
|
-
Stream smoothness metrics use stdout chunk arrival times. A chunk can contain more than one token, and
|
|
83
|
+
Stream smoothness metrics use stdout chunk arrival times. A chunk can contain more than one token, and pipe buffering can affect chunk boundaries. Treat `second_chunk_ms` and `chunk_gap` as user-visible streaming diagnostics, not raw decoder telemetry.
|
|
58
84
|
|
|
59
85
|
For serious comparisons, prefer at least three runs per prompt, include warmups, benchmark both interactive and throughput or client profiles, compare models at the same concurrency operating points, set SLOs that match your real UX budget, and save JSON reports for later analysis.
|
|
60
86
|
|
|
61
87
|
## Report artifacts
|
|
62
88
|
|
|
63
|
-
Saved JSON includes client-side environment metadata (hardware model, macOS build, `fm` help digest, power/thermal snapshot) so shared results remain interpretable on other machines. Use `fm-bench validate` before publishing and `fm-bench compare --strict` when you require identical prompt suites. Format details: [report-format.md](./report-format.md).
|
|
89
|
+
Saved JSON includes client-side environment metadata (hardware model, macOS build, `fm` help digest, power/thermal snapshot) plus the detected `fm` capabilities and per-metric availability, so shared results remain interpretable on other machines. Use `fm-bench validate` before publishing and `fm-bench compare --strict` when you require identical prompt suites. Format details: [report-format.md](./report-format.md).
|
package/docs/releasing.md
CHANGED
|
@@ -1,12 +1,23 @@
|
|
|
1
1
|
# Releasing
|
|
2
2
|
|
|
3
|
-
`fm-bench` uses semver tags (`v*.*.*`). Pushing a tag runs the **Release** workflow: test
|
|
3
|
+
`fm-bench` uses semver tags (`v*.*.*`). Pushing a tag runs the **Release** workflow: lint/test/package check, npm publish (with provenance), and a GitHub release whose body is taken from the matching section in **`CHANGELOG.md`** (plus a compare link).
|
|
4
4
|
|
|
5
5
|
## Prerequisites
|
|
6
6
|
|
|
7
7
|
- Repository secret **`NPM_TOKEN`**: npm automation token with publish access to `fm-bench`.
|
|
8
8
|
- **`main`** is green on CI.
|
|
9
9
|
|
|
10
|
+
If `NPM_TOKEN` is missing or expired, the Release workflow warns, skips the npm publish step, and still creates the GitHub release. The release notes then state that npm publish was skipped. Re-run the Release workflow with the tag once the token is fixed — an already-published version is skipped automatically.
|
|
11
|
+
|
|
12
|
+
## Before tagging
|
|
13
|
+
|
|
14
|
+
```sh
|
|
15
|
+
npm run check # lint + tests + npm pack integrity
|
|
16
|
+
npm run publish:dry-run
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Then bump `CHANGELOG.md` with a `## X.Y.Z` section; the release notes come from that section.
|
|
20
|
+
|
|
10
21
|
## Option A — GitHub Actions (recommended)
|
|
11
22
|
|
|
12
23
|
1. Open **Actions → Version → Run workflow**.
|
|
@@ -17,20 +28,34 @@
|
|
|
17
28
|
## Option B — Local
|
|
18
29
|
|
|
19
30
|
```sh
|
|
20
|
-
npm ci && npm
|
|
21
|
-
npm version
|
|
22
|
-
git push --follow-tags
|
|
31
|
+
npm ci && npm run check
|
|
32
|
+
npm version minor # or patch / major
|
|
33
|
+
git push origin main --follow-tags
|
|
23
34
|
```
|
|
24
35
|
|
|
25
36
|
## Re-run Release without republishing
|
|
26
37
|
|
|
27
|
-
If npm already has the version but GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.
|
|
38
|
+
If npm already has the version but the GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.6.3`). The workflow skips npm publish when that version is already on the registry. If the GitHub release already exists, it **updates the release notes** from `CHANGELOG.md`.
|
|
28
39
|
|
|
29
40
|
Refresh notes locally without re-publishing:
|
|
30
41
|
|
|
31
42
|
```sh
|
|
32
|
-
node scripts/changelog-release-notes.mjs 0.6.
|
|
33
|
-
gh release edit v0.6.
|
|
43
|
+
node scripts/changelog-release-notes.mjs 0.6.3 > notes.md
|
|
44
|
+
gh release edit v0.6.3 --notes-file notes.md --repo devinoldenburg/fm-bench
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Verifying a release
|
|
48
|
+
|
|
49
|
+
```sh
|
|
50
|
+
git tag --list 'v0.7.0'
|
|
51
|
+
gh release view v0.7.0 --repo devinoldenburg/fm-bench
|
|
52
|
+
npm view fm-bench version # registry version
|
|
53
|
+
|
|
54
|
+
mkdir -p fm-bench-check && cd fm-bench-check
|
|
55
|
+
npm install fm-bench@0.7.0
|
|
56
|
+
./node_modules/.bin/fm-bench --version
|
|
57
|
+
./node_modules/.bin/fm-bench --help > /dev/null
|
|
58
|
+
./node_modules/.bin/fm-bench doctor
|
|
34
59
|
```
|
|
35
60
|
|
|
36
61
|
## Dry run
|
|
@@ -40,4 +65,4 @@ npm run publish:dry-run
|
|
|
40
65
|
npm run check:pack
|
|
41
66
|
```
|
|
42
67
|
|
|
43
|
-
CI runs the
|
|
68
|
+
CI runs the dry run on every push to `main` and on pull requests.
|