fm-bench 0.7.1 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +32 -24
- package/bin/fm-bench.js +16 -5
- package/docs/compatibility.md +10 -5
- package/docs/methodology.md +14 -9
- package/docs/releasing.md +33 -16
- package/docs/report-format.md +23 -14
- package/docs/supported-platforms.md +9 -4
- package/package.json +11 -5
- package/src/bench.js +84 -25
- package/src/capabilities.js +26 -9
- package/src/cli.js +80 -8
- package/src/compare.js +10 -1
- package/src/export.js +6 -4
- package/src/fm-help.js +62 -7
- package/src/fm.js +123 -2
- package/src/metrics.js +15 -7
- package/src/process.js +7 -1
- package/src/report.js +3 -2
- package/src/schema.js +25 -0
- package/src/stats.js +36 -0
- package/src/table.js +60 -26
package/README.md
CHANGED
|
@@ -1,7 +1,8 @@
|
|
|
1
1
|
# fm-bench
|
|
2
2
|
|
|
3
|
-
[](https://github.com/dvnold/fm-bench/actions/workflows/ci.yml)
|
|
4
4
|
[](https://www.npmjs.com/package/fm-bench)
|
|
5
|
+
[](https://nodejs.org/)
|
|
5
6
|
|
|
6
7
|
Benchmark Apple's `fm` command on macOS 27+.
|
|
7
8
|
|
|
@@ -30,28 +31,27 @@ npm install -g fm-bench
|
|
|
30
31
|
fm-bench
|
|
31
32
|
```
|
|
32
33
|
|
|
33
|
-
One command discovers your models, runs the standard prompt suite, and prints a full benchmark report
|
|
34
|
+
One command discovers your models, warms each one up, runs the standard prompt suite, and prints a full benchmark report. Real output of `fm-bench --runs 3` on macOS 27.2 (26B5091g), Apple M5 Pro:
|
|
34
35
|
|
|
35
36
|
```text
|
|
36
|
-
fm-bench 0.
|
|
37
|
-
prompts 3 | runs 3 | concurrency 1 | stream on | measured 9 | failed 0 | skipped 0 | elapsed
|
|
38
|
-
|
|
39
|
-
unavailable: quota unavailable — this fm build exposes no quota command
|
|
37
|
+
fm-bench 0.8.0 | darwin/arm64 | fm
|
|
38
|
+
prompts 3 | runs 3 | concurrency 1 | stream on | measured 9 | failed 0 | skipped 0 | elapsed 11.07s
|
|
39
|
+
models: system = AFM 3 Core Advanced
|
|
40
40
|
|
|
41
|
-
|
|
42
|
-
│ C │ MODEL │ STATUS │ OK │
|
|
43
|
-
|
|
44
|
-
│ 1 │ system │ ok │ 9/9 │
|
|
45
|
-
|
|
41
|
+
┌───┬────────┬────────┬─────┬───────┬───────┬─────────┬────────┬───────┬────┬──────┐
|
|
42
|
+
│ C │ MODEL │ STATUS │ OK │ TTFT │ E2E │ E2E P95 │ USER/S │ SYS/S │ CV │ NOTE │
|
|
43
|
+
├───┼────────┼────────┼─────┼───────┼───────┼─────────┼────────┼───────┼────┼──────┤
|
|
44
|
+
│ 1 │ system │ ok │ 9/9 │ 620ms │ 687ms │ 1.43s │ 40.6 │ 37.7 │ 3% │ │
|
|
45
|
+
└───┴────────┴────────┴─────┴───────┴───────┴─────────┴────────┴───────┴────┴──────┘
|
|
46
46
|
|
|
47
47
|
┌───┬────────┬────────┬─────────┬───────────┬──────────┬───────────┬─────────┬────────┐
|
|
48
48
|
│ C │ MODEL │ IN AVG │ OUT AVG │ PREFILL/S │ DECODE/S │ CHUNK P95 │ E2E P99 │ REPEAT │
|
|
49
49
|
├───┼────────┼────────┼─────────┼───────────┼──────────┼───────────┼─────────┼────────┤
|
|
50
|
-
│ 1 │ system │
|
|
50
|
+
│ 1 │ system │ 27 │ 38 │ 40.7 │ 78.1 │ 197ms │ 1.45s │ 100% │
|
|
51
51
|
└───┴────────┴────────┴─────────┴───────────┴──────────┴───────────┴─────────┴────────┘
|
|
52
52
|
```
|
|
53
53
|
|
|
54
|
-
Default profile is `standard` (3 prompts). Use `--runs 5`, `--sweep-concurrency 1,2`, or `--profile client` for heavier suites.
|
|
54
|
+
Default profile is `standard` (3 prompts) with one run each, so run-to-run CV shows `-` until you pass `--runs 2` or more. Use `--runs 5`, `--sweep-concurrency 1,2`, or `--profile client` for heavier suites. When a model is unavailable (for example `pcc` outside the Terminal app), it is listed as skipped with fm's reason; `--available-only` hides it.
|
|
55
55
|
|
|
56
56
|
Wide terminals add TTFT P95, TPOT, and RPS columns; medium terminals tighten the table; narrow terminals switch to compact model cards automatically. `--width <n>` previews any layout.
|
|
57
57
|
|
|
@@ -64,7 +64,7 @@ npm install -g fm-bench
|
|
|
64
64
|
Install directly from GitHub (always latest):
|
|
65
65
|
|
|
66
66
|
```sh
|
|
67
|
-
npm install -g --install-links git+https://github.com/
|
|
67
|
+
npm install -g --install-links git+https://github.com/dvnold/fm-bench.git
|
|
68
68
|
```
|
|
69
69
|
|
|
70
70
|
Local development:
|
|
@@ -74,24 +74,26 @@ npm install && npm link
|
|
|
74
74
|
fm-bench doctor # verify your setup
|
|
75
75
|
```
|
|
76
76
|
|
|
77
|
-
**Requirements:** macOS 27+, Node.js
|
|
77
|
+
**Requirements:** macOS 27+, Node.js 22+, Apple Intelligence enabled.
|
|
78
78
|
|
|
79
79
|
## Commands
|
|
80
80
|
|
|
81
81
|
| Command | What it does |
|
|
82
82
|
|---------|-------------|
|
|
83
83
|
| `fm-bench` | Run the full benchmark (default) |
|
|
84
|
-
| `fm-bench models` | Show the detected `fm` capabilities, discovered models, availability, and quota
|
|
84
|
+
| `fm-bench models` | Show the detected `fm` capabilities, discovered models, their identity (for example `AFM 3 Core Advanced`), availability, and quota |
|
|
85
85
|
| `fm-bench compare <a.json> <b.json>` | Regression diff with suite/hardware/macOS warnings; `--strict` exits 2 when suites differ |
|
|
86
86
|
| `fm-bench history [dir]` | Trend table from saved reports (sorted by time, tags visible) |
|
|
87
87
|
| `fm-bench validate <report.json>` | Verify report JSON (schema v1) before sharing; `--json` for CI |
|
|
88
88
|
| `fm-bench export <report.json>` | Standalone HTML report with embedded JSON |
|
|
89
89
|
| `fm-bench legend` | Definitions, provenance (measured/proxy/derived), and color rules for every column |
|
|
90
|
-
| `fm-bench doctor` | Environment
|
|
90
|
+
| `fm-bench doctor` | Environment, `fm` license, and capability check; `--json` for scripts |
|
|
91
91
|
| `fm-bench metrics` | Alias for `legend` |
|
|
92
92
|
|
|
93
93
|
Exit codes: `0` success, `1` operational failure (failed `--ci` gate, invalid reports), `2` usage or environment error (bad flags, unsupported macOS, unusable `fm`, no runnable model). Interrupting a run with Ctrl+C terminates in-flight `fm` processes.
|
|
94
94
|
|
|
95
|
+
Mistyped commands and flags are caught with a suggestion (`Unknown command "modles". Did you mean "models"?`) instead of being benchmarked as a prompt. To benchmark such a word anyway, put it after `--`.
|
|
96
|
+
|
|
95
97
|
## Common Recipes
|
|
96
98
|
|
|
97
99
|
```sh
|
|
@@ -130,7 +132,7 @@ fm-bench --format csv --out bench.csv
|
|
|
130
132
|
|------|---------|-------------|
|
|
131
133
|
| `-m, --models <list>` | discovered | Comma-separated or repeated model names |
|
|
132
134
|
| `-r, --runs <n>` | 1 | Measured runs per prompt/model |
|
|
133
|
-
| `--warmup <n>` |
|
|
135
|
+
| `--warmup <n>` | 1 | Unmeasured warmup runs per model before measurement; `0` measures cold start |
|
|
134
136
|
| `-c, --concurrency <n>` | 1 | Parallel `fm` processes |
|
|
135
137
|
| `--sweep-concurrency <list>` | — | Separate operating points, e.g. `1,2,4` |
|
|
136
138
|
| `--request-rate <rps>` | — | Pace request starts at a target rate |
|
|
@@ -200,27 +202,32 @@ Nine built-in suites, choose the one that matches your use case:
|
|
|
200
202
|
|
|
201
203
|
## Metrics
|
|
202
204
|
|
|
203
|
-
**Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT, 95% confidence interval,
|
|
205
|
+
**Latency** — TTFT (p50/p95), E2E (p50/p95/p99), TPOT, 95% confidence interval, and run-to-run CV (per prompt across repeated runs, so mixing short and long prompts does not read as instability).
|
|
204
206
|
|
|
205
207
|
**Throughput** — prefill tokens/s, decode tokens/s, output tokens/s per request, aggregate system tokens/s, requests per second.
|
|
206
208
|
|
|
207
|
-
**Streaming quality** — second-chunk delay, chunk-gap p95, captured from stdout chunk arrival during streaming runs.
|
|
209
|
+
**Streaming quality** — first-chunk tokens, second-chunk delay, chunk-gap p95, captured from stdout chunk arrival during streaming runs.
|
|
208
210
|
|
|
209
211
|
**Reliability** — success rate, goodput rate and RPS against SLO budgets, repeatability (most common output hash frequency across repeated runs).
|
|
210
212
|
|
|
213
|
+
**Model identity** — the identity `fm models` reports (for example `system = AFM 3 Core Advanced`) is printed in the report header and saved in JSON, and `compare` warns when the model behind a name changed between two runs.
|
|
214
|
+
|
|
211
215
|
Every metric is labelled by provenance in JSON and in `fm-bench legend`:
|
|
212
216
|
|
|
213
217
|
- **measured** — process wall clock, exit codes, chunk arrival, `fm` token counts
|
|
214
218
|
- **proxy** — TTFT and chunk gaps (chunk granularity, not token timestamps); prefill tokens/s
|
|
215
219
|
- **derived** — TPOT, decode tokens/s, throughput, CV, confidence intervals, goodput
|
|
216
220
|
|
|
217
|
-
Token counts come from the `fm` build's own token-counting command (`count-tokens`, or `token-count` on older builds). If the build cannot count tokens, token metrics render as `-`, JSON carries `null`, and `metrics.promptTokens.available` is `false`. Spread statistics (`CV`, 95% CI) need at least two successful samples; with one sample they are unavailable rather than `0`.
|
|
221
|
+
Token counts come from the `fm` build's own token-counting command (`count-tokens`, or `token-count` on older builds). `count-tokens` adds one framing token to every count; fm-bench calibrates that overhead once per run and removes it from output counts (`tokenCounter` in JSON). If the build cannot count tokens, token metrics render as `-`, JSON carries `null`, and `metrics.promptTokens.available` is `false`. Spread statistics (`CV`, 95% CI) need at least two successful samples; with one sample they are unavailable rather than `0`.
|
|
222
|
+
|
|
223
|
+
`fm` streams coarse deltas — on macOS 27.2 the first stdout chunk carries about 20 tokens — so TTFT is time to the first chunk, and TPOT and decode tokens/s are computed only over the tokens that arrive after it. See [docs/methodology.md](docs/methodology.md).
|
|
218
224
|
|
|
219
225
|
## fm compatibility
|
|
220
226
|
|
|
221
227
|
`fm-bench` probes `fm --help` and `fm respond --help` once per run and adapts:
|
|
222
228
|
|
|
223
|
-
- Subcommand names (`count-tokens` vs legacy `token-count`) are detected, not hardcoded.
|
|
229
|
+
- Subcommand names (`count-tokens` vs legacy `token-count`, `models` vs the deprecated `available`) are detected, not hardcoded.
|
|
230
|
+
- Model availability and identity come from one `fm models` call; reasons such as `Private Cloud Compute is not available in this context. Please use the Terminal app.` are shown in full.
|
|
224
231
|
- Flags the build does not document (`--stream`, `--use-case`, `--guardrails`, `--model`) are not passed through.
|
|
225
232
|
- Unsupported models are rejected before a benchmark starts, with the supported list in the error.
|
|
226
233
|
- Raw `fm` argument-error text never reaches reports, tables, or JSON.
|
|
@@ -293,8 +300,9 @@ fm-bench legend --json # machine-readable
|
|
|
293
300
|
## Requirements
|
|
294
301
|
|
|
295
302
|
- macOS 27.0 or newer (Apple's `fm` CLI is preinstalled there).
|
|
296
|
-
- Node.js
|
|
297
|
-
- Apple Intelligence enabled on the device.
|
|
303
|
+
- Node.js 22 or newer (CI covers Node 22, 24, and 26).
|
|
304
|
+
- Apple Intelligence enabled on the device, and the `fm` terms accepted once with `fm license` (`fm-bench doctor` checks this).
|
|
305
|
+
- To benchmark the Private Cloud Compute model (`pcc`), run fm-bench from the Terminal app: `fm` reports `pcc` as unavailable in other contexts such as editor terminals.
|
|
298
306
|
|
|
299
307
|
Benchmark commands refuse to start on older macOS versions and report the detected version plus the latest supported macOS — see [docs/supported-platforms.md](docs/supported-platforms.md).
|
|
300
308
|
|
package/bin/fm-bench.js
CHANGED
|
@@ -4,18 +4,29 @@ import { killActiveChildren } from '../src/process.js';
|
|
|
4
4
|
|
|
5
5
|
// Terminate any in-flight `fm` child process before exiting, then use the
|
|
6
6
|
// conventional 128 + signal exit code. A second signal exits immediately.
|
|
7
|
-
let
|
|
7
|
+
let signalsReceived = 0;
|
|
8
|
+
let interrupted = false;
|
|
8
9
|
for (const [signal, exitCode] of [['SIGINT', 130], ['SIGTERM', 143]]) {
|
|
9
10
|
process.on(signal, () => {
|
|
11
|
+
signalsReceived += 1;
|
|
12
|
+
interrupted = true;
|
|
13
|
+
// Carry the status on the process too, so the code survives even if the
|
|
14
|
+
// exit below is preempted.
|
|
15
|
+
process.exitCode = exitCode;
|
|
10
16
|
const killed = killActiveChildren(signal);
|
|
11
|
-
if (
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
setTimeout(() => process.exit(exitCode),
|
|
17
|
+
if (signalsReceived > 1 || killed === 0) process.exit(exitCode);
|
|
18
|
+
// Deliberately not unref'd: this timer performs the exit after giving the
|
|
19
|
+
// killed children a moment to be reaped.
|
|
20
|
+
setTimeout(() => process.exit(exitCode), 500);
|
|
15
21
|
});
|
|
16
22
|
}
|
|
17
23
|
|
|
18
24
|
runCli(process.argv.slice(2)).catch((error) => {
|
|
25
|
+
if (interrupted) {
|
|
26
|
+
// The user asked to stop; a follow-up benchmark error is noise.
|
|
27
|
+
process.exitCode = typeof error?.exitCode === 'number' ? error.exitCode : 1;
|
|
28
|
+
return;
|
|
29
|
+
}
|
|
19
30
|
const message = error?.message || String(error);
|
|
20
31
|
console.error(`fm-bench: ${message}`);
|
|
21
32
|
process.exitCode = typeof error?.exitCode === 'number' ? error.exitCode : 1;
|
package/docs/compatibility.md
CHANGED
|
@@ -1,15 +1,18 @@
|
|
|
1
1
|
# `fm` compatibility
|
|
2
2
|
|
|
3
|
-
`fm-bench` is a client of whatever `fm` binary is installed. Apple has already changed that surface between macOS 27 builds:
|
|
3
|
+
`fm-bench` is a client of whatever `fm` binary is installed. Apple has already changed that surface between macOS 27 builds: `count-tokens` replaced `token-count`, macOS 27.2 renamed `available` to `models` (the old name still works there but prints a deprecation warning) and added `quota-usage`, `config`, and the `pcc` model. fm-bench therefore detects capabilities at runtime instead of assuming a fixed subcommand list.
|
|
4
4
|
|
|
5
5
|
## How detection works
|
|
6
6
|
|
|
7
|
-
Once per run, `fm-bench` spawns
|
|
7
|
+
Once per run, `fm-bench` spawns three cheap calls:
|
|
8
8
|
|
|
9
|
-
1. `fm --help` — the command list, the `MODELS` section, and (when present) a `--model` option list.
|
|
9
|
+
1. `fm --help` — the command list, the `MODELS` section, and (when present) a `--model` option list. The `<model>` custom-provider placeholder is not treated as a model.
|
|
10
10
|
2. `fm respond --help` — which flags `respond` actually accepts (`--model`, `--[no-]stream`, `--instructions`, `--greedy`, `--use-case`, `--guardrails`, `--image`, `--tool`, `--schema`).
|
|
11
|
+
3. `fm models` — one call that reports every model's availability, its identity (for example `✓ system (AFM 3 Core Advanced)`), and the reason for anything unavailable. Builds without `models` fall back to `fm available --model <name>` per model.
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
`warning:` lines that `fm` prints (such as the `available` rename notice) are discarded before parsing, so they can never be read as a model status or an error cause.
|
|
14
|
+
|
|
15
|
+
The result is recorded in the report (`capabilities`, `models[].identity`) and shown by `fm-bench models`, `fm-bench doctor`, and `fm-bench doctor --json`. `doctor` also runs `fm license --status`, because `fm respond` cannot run until the terms are accepted.
|
|
13
16
|
|
|
14
17
|
Detection is tolerant of formatting changes: section headers are matched case-insensitively on uppercase lines, model lines by their column layout, and boolean flags in either `--flag` or Apple's `--[no-]flag` spelling. ANSI escapes are stripped before parsing.
|
|
15
18
|
|
|
@@ -18,6 +21,7 @@ Detection is tolerant of formatting changes: section headers are matched case-in
|
|
|
18
21
|
| Capability | When missing |
|
|
19
22
|
|------------|--------------|
|
|
20
23
|
| Token counting (`count-tokens`, fallback `token-count`) | The run still benchmarks latency and streaming. Every token metric is `null` with `available: false` in `metrics`, the table prints an `unavailable:` note, and the report carries a warning. `fm count-tokens` is never called. |
|
|
24
|
+
| Model list (`models`, fallback `available`) | With neither command, models from `fm --help` are assumed runnable and any failure is recorded per run. Model identity is only available from `models`. |
|
|
21
25
|
| Streaming | TTFT, generation time, TPOT, prefill, and chunk-gap metrics are unavailable, and `--stream`/`--no-stream` is not passed through. E2E latency, success rate, and RPS still work. |
|
|
22
26
|
| Quota command | The `models` table omits the quota column and reports quota as unavailable. |
|
|
23
27
|
| Model selection (`--model`) | The flag is not passed; `fm` uses its default model. |
|
|
@@ -41,6 +45,7 @@ The macOS 27 requirement is enforced only when `fm-bench` resolves the default `
|
|
|
41
45
|
|
|
42
46
|
| Platform | `fm` surface | Notes |
|
|
43
47
|
|----------|--------------|-------|
|
|
44
|
-
| macOS 27.
|
|
48
|
+
| macOS 27.2 (26B5091g), Apple M5 Pro (Mac17,9) | `chat`, `config`, `count-tokens`, `license`, `models`, `quota-usage`, `respond`, `schema`, `serve` (plus deprecated `available`); models `system` (`AFM 3 Core Advanced`) and `pcc`; `--model-provider` for custom Chat Completions providers; no `--version` | Real-machine runs in 0.8.0. `pcc` reports "not available in this context. Please use the Terminal app." outside Terminal. `count-tokens` adds one framing token per count. The first streamed chunk carries about 20 tokens. Fixtures: `test/fixtures/*-macos27.2.txt`. |
|
|
49
|
+
| macOS 27.0 (26A5425a), Apple M5 Pro (Mac17,9) | `available`, `chat`, `count-tokens`, `license`, `respond`, `schema`, `serve`; model `system` only; no quota command; no `--version` | Real-machine smoke tests in 0.7.0. `test/fixtures/fm-help-macos27.txt` captures this help output as a regression baseline; the fake `fm` emulates it with `FAKE_FM_SCENARIO=legacy-available`. |
|
|
45
50
|
|
|
46
51
|
If your `fm` build differs, `fm-bench doctor` shows exactly what was detected, and the report's `capabilities` block records it. Please open an issue with the `doctor --json` output and the report digest when something is missing.
|
package/docs/methodology.md
CHANGED
|
@@ -31,24 +31,25 @@ The same classification is machine-readable in every JSON report under `metrics`
|
|
|
31
31
|
|
|
32
32
|
For a column-by-column terminal reference, run `fm-bench legend`.
|
|
33
33
|
|
|
34
|
-
- `TTFT` — *proxy*. Time from starting `fm respond` to the first streamed stdout chunk. `fm` exposes no per-token timestamps, so this is a terminal-side approximation of time to first token
|
|
34
|
+
- `TTFT` — *proxy*. Time from starting `fm respond` to the first streamed stdout chunk. `fm` exposes no per-token timestamps, so this is a terminal-side approximation of time to first token. `fm` streams coarse deltas: on macOS 27.2 the first chunk consistently carries about 20 tokens (the same through a pseudo-terminal and through `fm serve`), so TTFT is closer to "time to the first ~20 tokens". `first_chunk_tokens` records the exact count per run.
|
|
35
35
|
- `E2E latency` — *measured*. Time from starting `fm respond` until the process exits and the full response is captured.
|
|
36
|
-
- `
|
|
37
|
-
- `
|
|
38
|
-
- `
|
|
39
|
-
- `
|
|
36
|
+
- **Deliveries.** `fm` often writes the tail of a short answer as several `write()` calls well under a millisecond apart. Stdout chunks that arrive less than 5 ms after the previous one are therefore grouped into one *delivery*. Real streaming deltas on macOS 27.2 arrive about 50–250 ms apart, so this only removes write bursts. Generation time, TPOT, decode rate, `second_chunk_ms`, and `chunk_gap` are computed over deliveries; `stdout_chunks` still counts raw chunks.
|
|
37
|
+
- `generation_ms` — *derived*. Last streamed delivery minus first streamed delivery, reported only when the answer arrives in more than one delivery. When a single chunk carries the whole answer, prefill and decode are not separable and the value is unavailable instead of near zero. Process teardown after the last chunk is excluded.
|
|
38
|
+
- `TPOT` — *derived*. `generation_ms / (output_tokens - first_chunk_tokens)`: the tokens that arrived during the generation window divided into it. Reported when at least two tokens arrived after the first chunk, so the interval is an average rather than the inverse of a single chunk gap. Requires streaming and a token-counting `fm`. Before 0.8.0 the denominator was `output_tokens - 1`, which credited the ~20 first-chunk tokens to the decode window and overstated decode speed by up to several times on short answers.
|
|
39
|
+
- `second_chunk_ms` — *proxy*. Time between the first and second streamed deliveries: a terminal-side signal for startup smoothness.
|
|
40
|
+
- `chunk_gap` — *proxy*. Distribution of time between consecutive streamed deliveries. Useful for spotting streaming jitter; delivery-based, not token-based.
|
|
40
41
|
- `prefill_tokens_per_second` — *proxy*. Input prompt tokens divided by TTFT seconds. Because TTFT includes process startup and first-token latency, this systematically understates true prefill speed and should be read as an upper-bound-constrained estimate, not a kernel measurement.
|
|
41
42
|
- `tokens_per_second` — *derived*. Output tokens divided by E2E seconds for one request.
|
|
42
|
-
- `decode_tokens_per_second` — *derived*.
|
|
43
|
+
- `decode_tokens_per_second` — *derived*. Per run: output tokens that arrived after the first delivery divided by generation seconds. Short answers finish in two or three deliveries, so their per-run rate rests on few tokens. The table's DECODE/S column is therefore the token-weighted `decodeThroughput` (decode tokens summed over runs divided by the summed generation time), so long generations dominate. The `throughput` profile gives the most trustworthy decode figure (about 51–53 tokens/s on an M5 Pro with macOS 27.2).
|
|
43
44
|
- `output token throughput` — *derived*. All successful output tokens for a model divided by that model's measured wall-clock window.
|
|
44
45
|
- `total token throughput` — *derived*. Successful prompt and output tokens divided by the model's measured wall-clock window.
|
|
45
46
|
- `RPS` — *measured*. Successful requests divided by the model's measured wall-clock window.
|
|
46
47
|
- `goodput` — *derived*. Successful requests that also satisfy every provided SLO threshold. A run whose SLO metric is unavailable counts as not good, so an unverifiable SLO can never inflate goodput.
|
|
47
48
|
- `goodput RPS` — *derived*. SLO-passing requests divided by the model's measured wall-clock window. Zero is reported when SLOs are set and nothing passes.
|
|
48
49
|
- `repeatability` — *derived*. For repeated runs of the same prompt, the average share of runs that produced the most common normalized output hash.
|
|
49
|
-
- `CV` — *derived*.
|
|
50
|
+
- `CV` — *derived*. Run-to-run stability: the E2E coefficient of variation (sample standard deviation divided by the mean) of each prompt across its repeated runs, averaged over prompts (`stabilityCv` in JSON). Reported once at least one prompt has two successful runs. The pooled `latency.cv` over every sample is still in the JSON, but it mostly reflects that prompts have different lengths, so the table no longer shows it. Before 0.8.0 the CV column used the pooled value.
|
|
50
51
|
- `95% CI` — *derived*. t-distribution confidence interval around the sample mean, using exact t critical values up to 30 degrees of freedom and the standard 2.0 / 1.96 approximations beyond that. Reported only with two or more successful samples. Treat it as context, not proof, at small sample sizes.
|
|
51
|
-
- `quota` — *measured*, when the `fm` build exposes a quota command.
|
|
52
|
+
- `quota` — *measured*, when the `fm` build exposes a quota command. macOS 27.2 has `quota-usage` (the on-device `system` model reports "Not applicable", since quota only applies to `pcc`); macOS 27.0 has none, in which case `fm-bench` reports quota as unavailable rather than empty.
|
|
52
53
|
|
|
53
54
|
### Small samples
|
|
54
55
|
|
|
@@ -66,6 +67,8 @@ Use `--request-rate <rps>` to pace request starts independently of concurrency.
|
|
|
66
67
|
|
|
67
68
|
`--warmup <n>` runs `n` unmeasured calls per model at the start of each operating point. The first `fm respond` after a cold start includes model load time, which can dominate a short prompt (observed at several hundred milliseconds on Apple silicon). Warmups are never mixed into the measured results.
|
|
68
69
|
|
|
70
|
+
Since 0.8.0 the default is one warmup per model, so a default run reports steady-state latency. Pass `--warmup 0` to measure cold start deliberately. `warmup` is part of the suite key, so `compare` flags a 0.7.x report (warmup 0) against a 0.8.0 default run as a different suite.
|
|
71
|
+
|
|
69
72
|
## Failures and retries
|
|
70
73
|
|
|
71
74
|
A failed call is recorded as a failed measurement with its error text, and is excluded from latency and throughput statistics. `--retry <n>` retries failed calls with exponential backoff before recording the failure; the report keeps the run count and records the number of attempts separately, so a retried success is still one sample and never silently duplicates work.
|
|
@@ -76,7 +79,9 @@ The `client` profile is a pragmatic local-machine mix inspired by MLPerf Client'
|
|
|
76
79
|
|
|
77
80
|
## Caveats
|
|
78
81
|
|
|
79
|
-
Token counts come from the `fm` build's own token-counting command (`count-tokens` on current builds, `token-count` on older ones). If a build has no such command, token-derived metrics are reported as unavailable rather than estimated.
|
|
82
|
+
Token counts come from the `fm` build's own token-counting command (`count-tokens` on current builds, `token-count` on older ones). If a build has no such command, token-derived metrics are reported as unavailable rather than estimated.
|
|
83
|
+
|
|
84
|
+
`count-tokens` counts its input as a prompt and adds a constant framing token: on macOS 27.2 it reports 2 for `a`, 3 for `a a`, and 4 for `a a a`. Once per run fm-bench counts those three strings, derives the overhead from the consistent per-word step (`overhead = count(a) − step`), and subtracts it from output and first-chunk token counts. The result is saved as `tokenCounter: { command, overhead, calibrated }`. If the three counts are not consistent, no correction is applied and `calibrated` is `false`. Prompt token counts are left exactly as `fm` reports them, because the framing token is part of what the model processes. `count-tokens` always uses the on-device system tokenizer, so token counts for `pcc` are system-tokenizer counts. `fm-bench` cannot judge semantic quality unless you provide your own prompt suite and inspect captured outputs with `--capture-output`.
|
|
80
85
|
|
|
81
86
|
Client-side measurements include process startup, local queueing, model prefill, streaming, detokenization, and terminal pipe overhead. That is intentional for a command-line benchmark, but it is not the same as an internal model-kernel benchmark.
|
|
82
87
|
|
package/docs/releasing.md
CHANGED
|
@@ -1,61 +1,78 @@
|
|
|
1
1
|
# Releasing
|
|
2
2
|
|
|
3
|
-
`fm-bench` uses semver tags (`v*.*.*`).
|
|
3
|
+
`fm-bench` uses semver tags (`v*.*.*`). A release tag runs the **Release** workflow: lint/test/package check, npm publish with provenance, a GitHub release whose body is the matching section of **`CHANGELOG.md`** (plus a compare link), and an install check of the published package.
|
|
4
4
|
|
|
5
|
-
##
|
|
5
|
+
## npm authentication
|
|
6
6
|
|
|
7
|
-
-
|
|
8
|
-
- **`main`** is green on CI.
|
|
7
|
+
The Release workflow runs on Node 24 (npm 11) with `id-token: write`, so it publishes through **npm trusted publishing** (OIDC) when the package trusts this workflow. No long-lived token is needed, and provenance is attached automatically. Set it up once (requires an npm login with publish rights and 2FA):
|
|
9
8
|
|
|
10
|
-
|
|
9
|
+
```sh
|
|
10
|
+
npm login
|
|
11
|
+
npm trust github fm-bench --repo dvnold/fm-bench --file release.yml
|
|
12
|
+
npm trust list fm-bench
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
or on npmjs.com: **fm-bench → Settings → Trusted publishing → GitHub Actions**, organization/user `dvnold`, repository `fm-bench`, workflow `release.yml`.
|
|
16
|
+
|
|
17
|
+
Without trusted publishing, npm falls back to the **`NPM_TOKEN`** repository secret (a granular access token with publish rights). npm limits those tokens to 90 days, so an old token fails with `E401`/`E403`/`E404`.
|
|
18
|
+
|
|
19
|
+
When publishing fails, the workflow still creates or updates the GitHub release (its notes say npm publish failed) and then fails the run, so the problem is visible. Fix the authentication, then re-run the workflow for the tag. An already-published version is skipped.
|
|
20
|
+
|
|
21
|
+
Provenance also requires `repository.url` in `package.json` to match the repository the workflow runs in (`git+https://github.com/dvnold/fm-bench.git`). Update it if the repository is renamed or transferred.
|
|
11
22
|
|
|
12
23
|
## Before tagging
|
|
13
24
|
|
|
14
25
|
```sh
|
|
15
26
|
npm run check # lint + tests + npm pack integrity
|
|
16
27
|
npm run publish:dry-run
|
|
28
|
+
actionlint # if workflows changed
|
|
17
29
|
```
|
|
18
30
|
|
|
19
|
-
Then
|
|
31
|
+
Then add a `## X.Y.Z` section to `CHANGELOG.md`; the release notes come from that section.
|
|
20
32
|
|
|
21
33
|
## Option A — GitHub Actions (recommended)
|
|
22
34
|
|
|
23
35
|
1. Open **Actions → Version → Run workflow**.
|
|
24
36
|
2. Choose `patch`, `minor`, `major`, or an exact semver.
|
|
25
|
-
3. The job runs `npm version`, pushes the commit and tag to `main
|
|
26
|
-
4. The tag push triggers **Release** automatically.
|
|
37
|
+
3. The job runs `npm version`, pushes the commit and tag to `main`, and dispatches **Release** for the new tag. (A tag pushed with the workflow's `GITHUB_TOKEN` does not trigger other workflows by itself, so Version starts Release explicitly.)
|
|
27
38
|
|
|
28
39
|
## Option B — Local
|
|
29
40
|
|
|
30
41
|
```sh
|
|
31
42
|
npm ci && npm run check
|
|
32
43
|
npm version minor # or patch / major
|
|
33
|
-
git push origin main --follow-tags
|
|
44
|
+
git push origin main --follow-tags # the tag push triggers Release
|
|
34
45
|
```
|
|
35
46
|
|
|
36
47
|
## Re-run Release without republishing
|
|
37
48
|
|
|
38
|
-
If npm already has the version but the GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.
|
|
49
|
+
If npm already has the version but the GitHub release failed (or vice versa), use **Actions → Release → Run workflow** and enter the existing tag (for example `v0.8.0`). The workflow skips npm publish when that version is already on the registry. If the GitHub release already exists, it **updates the release notes** from `CHANGELOG.md`.
|
|
50
|
+
|
|
51
|
+
```sh
|
|
52
|
+
gh workflow run release.yml --repo dvnold/fm-bench -f tag=v0.8.0
|
|
53
|
+
```
|
|
39
54
|
|
|
40
55
|
Refresh notes locally without re-publishing:
|
|
41
56
|
|
|
42
57
|
```sh
|
|
43
|
-
node scripts/changelog-release-notes.mjs 0.
|
|
44
|
-
gh release edit v0.
|
|
58
|
+
node scripts/changelog-release-notes.mjs 0.8.0 > notes.md
|
|
59
|
+
gh release edit v0.8.0 --notes-file notes.md --repo dvnold/fm-bench
|
|
45
60
|
```
|
|
46
61
|
|
|
47
62
|
## Verifying a release
|
|
48
63
|
|
|
64
|
+
The workflow installs the published version into a clean directory and runs it against the fake `fm`. To check by hand:
|
|
65
|
+
|
|
49
66
|
```sh
|
|
50
|
-
|
|
51
|
-
gh release view v0.7.0 --repo devinoldenburg/fm-bench
|
|
67
|
+
gh release view v0.8.0 --repo dvnold/fm-bench
|
|
52
68
|
npm view fm-bench version # registry version
|
|
69
|
+
npm view fm-bench@0.8.0 dist.attestations --json # provenance
|
|
53
70
|
|
|
54
71
|
mkdir -p fm-bench-check && cd fm-bench-check
|
|
55
|
-
npm install fm-bench@0.
|
|
72
|
+
npm install fm-bench@0.8.0
|
|
56
73
|
./node_modules/.bin/fm-bench --version
|
|
57
|
-
./node_modules/.bin/fm-bench --help > /dev/null
|
|
58
74
|
./node_modules/.bin/fm-bench doctor
|
|
75
|
+
./node_modules/.bin/fm-bench --profile quick --runs 3
|
|
59
76
|
```
|
|
60
77
|
|
|
61
78
|
## Dry run
|
package/docs/report-format.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Report format (schema v1)
|
|
2
2
|
|
|
3
|
-
Every measured run can be saved as JSON. Reports from fm-bench **0.6.0+** include a versioned schema so you can validate, share, and compare results across machines. The schema version
|
|
3
|
+
Every measured run can be saved as JSON. Reports from fm-bench **0.6.0+** include a versioned schema so you can validate, share, and compare results across machines. The schema version is still `1`: the `capabilities`, `metrics`, and per-result `attempts` fields (0.7.0) and the `tokenCounter`, `models[].identity`, `firstChunkTokens`, and `stabilityCv` fields (0.8.0) are additive, so older readers and older reports both keep working.
|
|
4
4
|
|
|
5
5
|
## Top-level fields
|
|
6
6
|
|
|
@@ -15,10 +15,11 @@ Every measured run can be saved as JSON. Reports from fm-bench **0.6.0+** includ
|
|
|
15
15
|
| `environment` | Host fingerprint: platform, arch, Node, hardware model, CPU, memory, macOS version/build, `fm` help digest, thermal/power snapshot |
|
|
16
16
|
| `capabilities` | What the installed `fm` build exposes (see below) |
|
|
17
17
|
| `metrics` | Per-metric availability and provenance for this run (see below) |
|
|
18
|
-
| `
|
|
18
|
+
| `tokenCounter` | `{ command, overhead, calibrated }`: the token-counting command and the framing overhead removed from output counts (see [methodology](./methodology.md#caveats)) |
|
|
19
|
+
| `suite` | Derived suite key + fingerprint for apples-to-apples comparison. `suite.fingerprint.modelIdentities` maps model names to identities, for example `{ "system": "AFM 3 Core Advanced" }` |
|
|
19
20
|
| `prompts` | Prompt ids, text, and token counts |
|
|
20
|
-
| `models` | Discovered models and
|
|
21
|
-
| `summary` | Per-model / per-concurrency roll-up statistics |
|
|
21
|
+
| `models` | Discovered models, availability, `reason` when unavailable, and `identity` as reported by `fm models` |
|
|
22
|
+
| `summary` | Per-model / per-concurrency roll-up statistics, including `stabilityCv` (run-to-run CV), `decodeThroughput` (token-weighted decode tokens/s), and `firstChunkTokens` |
|
|
22
23
|
| `results` | Per-run measurements (optional `output` when `--capture-output`) |
|
|
23
24
|
|
|
24
25
|
## `capabilities`
|
|
@@ -29,13 +30,18 @@ Detected once per run from `fm --help` and `fm respond --help`.
|
|
|
29
30
|
{
|
|
30
31
|
"capabilities": {
|
|
31
32
|
"bin": "fm",
|
|
32
|
-
"digest": "
|
|
33
|
-
"commands": ["
|
|
34
|
-
"models": [
|
|
33
|
+
"digest": "a45f2838f399c5da",
|
|
34
|
+
"commands": ["chat", "config", "count-tokens", "license", "models", "quota-usage", "respond", "schema", "serve"],
|
|
35
|
+
"models": [
|
|
36
|
+
{ "name": "system", "description": "On-device Apple Foundation Model" },
|
|
37
|
+
{ "name": "pcc", "description": "Apple Foundation Model on Private Cloud Compute" }
|
|
38
|
+
],
|
|
35
39
|
"features": {
|
|
36
40
|
"tokenCounting": true,
|
|
37
41
|
"tokenCountCommand": "count-tokens",
|
|
38
|
-
"
|
|
42
|
+
"modelListCommand": "models",
|
|
43
|
+
"license": true,
|
|
44
|
+
"quota": true,
|
|
39
45
|
"streaming": true,
|
|
40
46
|
"modelSelection": true,
|
|
41
47
|
"instructions": true,
|
|
@@ -47,11 +53,13 @@ Detected once per run from `fm --help` and `fm respond --help`.
|
|
|
47
53
|
"structuredOutput": true,
|
|
48
54
|
"server": true
|
|
49
55
|
},
|
|
50
|
-
"warnings": [
|
|
56
|
+
"warnings": []
|
|
51
57
|
}
|
|
52
58
|
}
|
|
53
59
|
```
|
|
54
60
|
|
|
61
|
+
On a macOS 27.0 build the same block has `"modelListCommand": "available"`, `"quota": false`, and the warning `this fm build exposes no quota command, so quota is not reported`.
|
|
62
|
+
|
|
55
63
|
`digest` is a short hash of the normalized `fm --help` output, so a report records which CLI surface produced it. Reports from two different `fm` builds are not directly comparable even on the same machine.
|
|
56
64
|
|
|
57
65
|
## `metrics`
|
|
@@ -72,8 +80,8 @@ Each entry states how a metric was obtained for this run and whether it was avai
|
|
|
72
80
|
"label": "quota",
|
|
73
81
|
"kind": "measured",
|
|
74
82
|
"source": "fm quota-usage",
|
|
75
|
-
"available":
|
|
76
|
-
"unavailableReason": "
|
|
83
|
+
"available": true,
|
|
84
|
+
"unavailableReason": ""
|
|
77
85
|
}
|
|
78
86
|
}
|
|
79
87
|
}
|
|
@@ -87,13 +95,14 @@ Per-run rows in `results` carry:
|
|
|
87
95
|
|-------|-------|
|
|
88
96
|
| `attempts` | Total `fm` invocations for this measured run, including retries. `1` when no retry was needed. |
|
|
89
97
|
| `ok` | Whether the call produced a response and exited cleanly. |
|
|
90
|
-
| `firstTokenMs`, `generationMs`, `tpotMs` | `null` when the run cannot supply them (no streaming, single-chunk answer, failed run). |
|
|
91
|
-
| `promptTokens`, `outputTokens`, `tokensPerSecond`, `decodeTokensPerSecond`, `prefillTokensPerSecond` | `null` when the `fm` build cannot count tokens. |
|
|
98
|
+
| `firstTokenMs`, `generationMs`, `tpotMs` | `null` when the run cannot supply them (no streaming, single-chunk answer, failed run). `generationMs` runs from the first to the last streamed chunk. |
|
|
99
|
+
| `promptTokens`, `outputTokens`, `tokensPerSecond`, `decodeTokensPerSecond`, `prefillTokensPerSecond` | `null` when the `fm` build cannot count tokens. `promptTokens` is fm's own count; `outputTokens` has `tokenCounter.overhead` removed. |
|
|
100
|
+
| `firstChunkTokens`, `decodeTokens` | Output tokens carried by the first streamed delivery (overhead removed), and the tokens that arrived after it. `tpotMs` and `decodeTokensPerSecond` use only `decodeTokens`. `null` for single-delivery or non-streamed answers. |
|
|
92
101
|
| `error` | Actionable failure text (`timed out after 30000ms`, `fm exited with code 3`, the first actionable line of `fm` stderr). |
|
|
93
102
|
|
|
94
103
|
## CSV
|
|
95
104
|
|
|
96
|
-
`--format csv` and `--out runs.csv` write per-run rows. Columns are stable; `attempts` was added in 0.7.0. Text cells that begin with `=`, `+`, `@`, or a non-numeric `-` are prefixed with a single quote so spreadsheets do not execute prompt or model output as a formula.
|
|
105
|
+
`--format csv` and `--out runs.csv` write per-run rows. Columns are stable; `attempts` was added in 0.7.0, and `first_chunk_tokens` was appended as the last column in 0.8.0 so positional readers keep working. Text cells that begin with `=`, `+`, `@`, or a non-numeric `-` are prefixed with a single quote so spreadsheets do not execute prompt or model output as a formula.
|
|
97
106
|
|
|
98
107
|
## Sharing results
|
|
99
108
|
|
|
@@ -5,7 +5,7 @@
|
|
|
5
5
|
## Requirements
|
|
6
6
|
|
|
7
7
|
- **macOS 27.0 or newer** — Apple's `fm` CLI is preinstalled starting with macOS 27. Older macOS releases do not ship `fm`, so `fm-bench` cannot run there.
|
|
8
|
-
- **Node.js
|
|
8
|
+
- **Node.js 22 or newer**. Node 20 reached end of life in April 2026; CI tests Node 22, 24, and 26.
|
|
9
9
|
- **Apple Intelligence enabled** on the device.
|
|
10
10
|
|
|
11
11
|
## Version enforcement
|
|
@@ -31,13 +31,18 @@ Support is capability-based rather than version-pinned. The CLI probes the insta
|
|
|
31
31
|
|
|
32
32
|
## Models
|
|
33
33
|
|
|
34
|
-
`fm-bench` benchmarks exactly the models the installed `fm` reports — it does not assume that any particular cloud or adapter model exists.
|
|
34
|
+
`fm-bench` benchmarks exactly the models the installed `fm` reports — it does not assume that any particular cloud or adapter model exists. macOS 27.0 ships the on-device `system` model only; macOS 27.2 adds `pcc` (Apple Foundation Model on Private Cloud Compute). New models are discovered automatically and reported with their own availability and identity.
|
|
35
35
|
|
|
36
|
-
|
|
36
|
+
`fm` only serves `pcc` to the Terminal app. From editor terminals and other contexts it reports `Private Cloud Compute is not available in this context. Please use the Terminal app.`, and fm-bench shows that reason unchanged. Run fm-bench from Terminal to benchmark `pcc`.
|
|
37
|
+
|
|
38
|
+
Requesting only models that cannot run right now exits with code `2` before any benchmark starts, and `fm-bench models` shows the same reasons without failing:
|
|
37
39
|
|
|
38
40
|
```text
|
|
39
41
|
fm-bench: No benchmark was run: none of the requested models are usable right now.
|
|
40
42
|
requested: pcc
|
|
41
|
-
|
|
43
|
+
models reported by this fm build: system, pcc
|
|
44
|
+
pcc: Private Cloud Compute is not available in this context. Please use the Terminal app.
|
|
42
45
|
run "fm-bench models" to see availability and reasons
|
|
43
46
|
```
|
|
47
|
+
|
|
48
|
+
On a build that does not have the model at all, the reason reads `not supported by this fm build (supported: system)` instead.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "fm-bench",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.8.0",
|
|
4
4
|
"description": "Dynamic benchmark CLI for Apple's fm command on macOS 27+.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -26,25 +26,31 @@
|
|
|
26
26
|
},
|
|
27
27
|
"keywords": [
|
|
28
28
|
"apple",
|
|
29
|
+
"apple-intelligence",
|
|
30
|
+
"apple-silicon",
|
|
29
31
|
"foundation-models",
|
|
30
32
|
"fm",
|
|
33
|
+
"llm",
|
|
31
34
|
"benchmark",
|
|
35
|
+
"benchmarking",
|
|
36
|
+
"latency",
|
|
37
|
+
"throughput",
|
|
32
38
|
"macos",
|
|
33
39
|
"cli"
|
|
34
40
|
],
|
|
35
41
|
"author": "Devin Oldenburg",
|
|
36
42
|
"license": "MIT",
|
|
37
43
|
"engines": {
|
|
38
|
-
"node": ">=
|
|
44
|
+
"node": ">=22"
|
|
39
45
|
},
|
|
40
46
|
"repository": {
|
|
41
47
|
"type": "git",
|
|
42
|
-
"url": "git+https://github.com/
|
|
48
|
+
"url": "git+https://github.com/dvnold/fm-bench.git"
|
|
43
49
|
},
|
|
44
50
|
"bugs": {
|
|
45
|
-
"url": "https://github.com/
|
|
51
|
+
"url": "https://github.com/dvnold/fm-bench/issues"
|
|
46
52
|
},
|
|
47
|
-
"homepage": "https://github.com/
|
|
53
|
+
"homepage": "https://github.com/dvnold/fm-bench#readme",
|
|
48
54
|
"publishConfig": {
|
|
49
55
|
"access": "public",
|
|
50
56
|
"registry": "https://registry.npmjs.org/"
|