claude-autorouter 0.3.6 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/.env.example +12 -7
  2. package/CONTRIBUTING.md +37 -0
  3. package/README.md +43 -70
  4. package/bin/autorouter.mjs +40 -56
  5. package/docs/development.md +50 -2
  6. package/docs/hardware-benchmark.md +29 -0
  7. package/docs/hardware-comparison.md +55 -0
  8. package/docs/hardware-results-16gb.json +4002 -0
  9. package/docs/hardware-results-16gb.md +26 -0
  10. package/docs/hardware-results-64gb.json +4020 -0
  11. package/docs/reference.md +83 -40
  12. package/docs/releasing.md +76 -34
  13. package/docs/router-performance.json +1697 -0
  14. package/docs/router-performance.md +50 -0
  15. package/docs/status-performance.json +363 -0
  16. package/docs/status-performance.md +44 -0
  17. package/package.json +57 -9
  18. package/src/auto-routing.mjs +214 -0
  19. package/src/bounded-json.mjs +57 -0
  20. package/src/cli-help.mjs +87 -0
  21. package/src/config-command.mjs +141 -0
  22. package/src/config.mjs +52 -27
  23. package/src/contracts.mjs +123 -0
  24. package/src/evaluation-report.mjs +114 -0
  25. package/src/local-diagnostic.mjs +191 -0
  26. package/src/model-catalog.mjs +96 -0
  27. package/src/model-request.mjs +10 -6
  28. package/src/ollama-evaluator.mjs +9 -27
  29. package/src/onboarding.mjs +82 -23
  30. package/src/request-validation.mjs +54 -0
  31. package/src/response-observer.mjs +126 -18
  32. package/src/router.mjs +174 -70
  33. package/src/savings.mjs +74 -16
  34. package/src/server.mjs +79 -12
  35. package/src/session-history.mjs +261 -0
  36. package/src/session-log.mjs +9 -58
  37. package/src/status-state.mjs +110 -62
  38. package/src/statusline.mjs +57 -27
  39. package/src/telemetry-event.mjs +196 -0
  40. package/src/token-counter.mjs +3 -1
  41. package/src/turn-state.mjs +132 -0
  42. package/src/user-config.mjs +18 -8
@@ -0,0 +1,50 @@
1
+ # Synthetic router performance
2
+
3
+ On October 5, 2026, eight simultaneous identical requests made one evaluator call after coalescing, compared with eight before it. This holds across different agent scopes. Distinct requests still evaluate independently, and pinned tool continuations still evaluate on every request. [The recorded baseline and candidate](router-performance.json) contain three rounds of measurements and their comparison gates.
4
+
5
+ The baseline was recorded before changing routing or evaluator transport. The final candidate was measured at 10:25 UTC after the compatibility guards and response handling were frozen; the baseline and comparison thresholds were retained unchanged. Both versions used the same synthetic harness: 60 iterations per scenario, concurrency eight, a mocked evaluator with a one-millisecond timer, and a 286,807-byte request containing 300 synthetic tools. No network services, user prompts, provider credentials, model inference, or downloads were involved.
6
+
7
+ Measurements came from one 16 GiB Apple M4 Mac with ten logical CPUs, Darwin 25.6.0, and Node v26.10.0. The workstation remained in normal use. Its one-minute load average was 2.32 after the baseline and 2.22 after the final candidate; free memory and the remaining load averages are in the JSON. The final run preceded the full regression suite to avoid interference from that workload. No user process names or arguments were collected.
8
+
9
+ ## Results
10
+
11
+ The table reports the largest p50 and p95 across the three rounds, in milliseconds. Counts exclude setup calls that warm the cache or establish the initial pin.
12
+
13
+ | Scenario | Before p50 / p95 | After p50 / p95 | Evaluator calls per round, before → after |
14
+ | --- | ---: | ---: | ---: |
15
+ | Small request | 1.331 / 1.646 | 1.261 / 1.396 | 60 → 60 |
16
+ | Large tool catalog | 3.421 / 4.415 | 2.810 / 3.037 | 60 → 60 |
17
+ | Cached request | 0.025 / 0.035 | 0.010 / 0.014 | 0 → 0 |
18
+ | Pinned tool continuation | 1.373 / 1.778 | 1.261 / 1.346 | 60 → 60 |
19
+ | Eight identical concurrent requests | 1.936 / 2.815 | 1.915 / 2.397 | 480 → 60 |
20
+ | Identical requests from eight agents | 1.715 / 2.219 | 1.566 / 1.815 | 480 → 60 |
21
+ | Distinct requests from eight agents | 1.686 / 2.201 | 1.523 / 1.907 | 480 → 480 |
22
+ | Immediate caller cancellation | 0.097 / 0.124 | 0.028 / 0.056 | 60 → 0 |
23
+
24
+ The evaluator-call reduction is deterministic. Timing differences also reflect timer scheduling, garbage collection, JIT warmup, and background load; these measurements do not establish faster Claude responses or a model inference speedup. Route time includes the synthetic evaluator delay and JSON processing. `policy_overhead_ms` separately reports route timing minus classification timing; it is an approximate measure using the router's rounded telemetry. There is no network or real model latency in either measure.
25
+
26
+ The JSON includes retained heap/RSS changes, sampled peak heap/RSS changes, and event-loop delay. Sampling runs every millisecond, with explicit garbage collection before and after each workload. It can miss allocations inside one synchronous interval. The event-loop histogram has ten-millisecond resolution, so values near ten milliseconds are the sampling floor. RSS is descriptive because allocator page reuse can dominate differences between short runs.
27
+
28
+ ## Local regression gates
29
+
30
+ For each scenario, the p95 route-time limit is twice the largest p95 across the three pre-change rounds. Peak sampled heap and p95 event-loop delay use the same baseline-relative margin. For example, the small-request route limit is 3.292 ms, the large-catalog limit is 8.831 ms, and the identical-agent limit is 4.439 ms. These are local regression guards, not portable latency promises or CI timing thresholds. The comparison also requires matching workload parameters, Node version, and hardware, and exactly one evaluator call per group of eight identical requests. All recorded gates pass.
31
+
32
+ Repeat this process with a new output file before and after a proposed change:
33
+
34
+ ```sh
35
+ node --expose-gc scripts/benchmark-router.mjs baseline /tmp/router-comparison.json
36
+ # Apply the proposed change.
37
+ node --expose-gc scripts/benchmark-router.mjs candidate /tmp/router-comparison.json
38
+ ```
39
+
40
+ The script refuses to replace an existing baseline and exits nonzero when a comparison gate fails. Its results should be reviewed alongside deterministic concurrency/cancellation tests. Actual local-model measurements on 16 GiB and 64 GiB Macs are recorded separately in [the hardware comparison](hardware-comparison.md), including classifier failures and comparability limits. Status-file storage performance is measured separately.
41
+
42
+ ## Bounds and preserved behavior
43
+
44
+ The router retains at most 256 pending evaluations and 1,024 subscribers. Capacity exhaustion uses the conservative fallback and reports `classifier_error: capacity_exhausted`. Each subscriber can cancel independently; the final cancellation aborts the evaluator, removes its pending entry, and prevents a late answer from entering the cache. Failed work is never cached. Coalescing shares only classification; each caller still applies its own scoped turn, context, and compatibility policy.
45
+
46
+ Keys retain the full request body, including the requested model/floor, and include the evaluator endpoint, model, credentials, deadline, excerpt budget, confidence policy or keep-alive setting, and a frozen rubric hash. Only their SHA-256 digests are stored as keys; credentials and request contents are not benchmark output. Using a smaller excerpt-only key would need a separate compatibility review.
47
+
48
+ Both evaluators use the same bounded JSON reader: 64 KiB for decisions and 1 MiB for local-model metadata. Streamed bytes enforce the limit even with missing or inaccurate length headers. A single growable buffer bounds chunk metadata, malformed JSON produces a stable error category, and cancellation never waits on an uncooperative body-cancellation promise. Jev's deadline now includes body reading; Ollama's configured zero deadline remains disabled while caller cancellation and byte limits remain active.
49
+
50
+ Prior-pin classification reuse was considered and intentionally deferred. The benchmark shows approximately 1.3 ms median total routing time for synthetic pinned turns, while eliminating the evaluation would alter the existing evaluate-per-request behavior. Coalescing removes duplicate simultaneous evaluations without making that semantic change. New human tasks and sequential tool continuations retain their current evaluation and safety checks.
@@ -0,0 +1,363 @@
1
+ {
2
+ "schema_version": 1,
3
+ "runs": [
4
+ {
5
+ "schema_version": 1,
6
+ "label": "synchronous-baseline",
7
+ "recorded_at": "2026-10-05T08:22:55.868Z",
8
+ "environment": {
9
+ "node": "v26.10.0",
10
+ "platform": "darwin",
11
+ "release": "25.6.0",
12
+ "arch": "arm64",
13
+ "cpu": "Apple M4",
14
+ "logical_cpus": 10,
15
+ "memory_bytes": 17179869184,
16
+ "load_average": [
17
+ 3.75537109375,
18
+ 3.2236328125,
19
+ 3.1318359375
20
+ ],
21
+ "background_load": "Other workspace activity was not suspended; no process identities or personal data captured."
22
+ },
23
+ "method": {
24
+ "rounds_per_run": 60,
25
+ "repetitions": 3,
26
+ "events_per_round": 6,
27
+ "concurrent_sessions": 20,
28
+ "producer_interval_ms": 5,
29
+ "mock_stream_interval_ms": 5,
30
+ "injection": "One 20ms delay per snapshot write: blocking wait for synchronous writer, timer wait for asynchronous writer. Metadata operations are not delayed.",
31
+ "note": "Stream timer lateness measures event-loop interference, not a network benchmark. Flush completion may include coalesced subsequent snapshots."
32
+ },
33
+ "scenarios": [
34
+ {
35
+ "storage": "normal_temporary_directory",
36
+ "injected_write_delay_ms": 0,
37
+ "timings": {
38
+ "update_ms": {
39
+ "samples": 180,
40
+ "p50": 0.09,
41
+ "p95": 0.189,
42
+ "max": 1.263
43
+ },
44
+ "flush_call_ms": {
45
+ "samples": 180,
46
+ "p50": 1.07,
47
+ "p95": 1.955,
48
+ "max": 2.467
49
+ },
50
+ "flush_completion_ms": {
51
+ "samples": 180,
52
+ "p50": 1.082,
53
+ "p95": 2.007,
54
+ "max": 2.542
55
+ },
56
+ "stream_timer_lateness_ms": {
57
+ "samples": 186,
58
+ "p50": 0.996,
59
+ "p95": 2.36,
60
+ "max": 3.6
61
+ },
62
+ "startup_ms": {
63
+ "samples": 3,
64
+ "p50": 1.069,
65
+ "p95": 1.069,
66
+ "max": 1.25
67
+ },
68
+ "close_ms": {
69
+ "samples": 3,
70
+ "p50": 0.958,
71
+ "p95": 0.958,
72
+ "max": 1.062
73
+ }
74
+ },
75
+ "event_loop_delay_ms_by_run": [
76
+ {
77
+ "p50": 1.151,
78
+ "p95": 2.357,
79
+ "max": 4.073
80
+ },
81
+ {
82
+ "p50": 1.153,
83
+ "p95": 2.591,
84
+ "max": 3.912
85
+ },
86
+ {
87
+ "p50": 1.154,
88
+ "p95": 2.726,
89
+ "max": 3.903
90
+ }
91
+ ],
92
+ "snapshot_writes_by_run": [
93
+ 61,
94
+ 61,
95
+ 61
96
+ ],
97
+ "rss_peak_bytes": 22137664
98
+ },
99
+ {
100
+ "storage": "injected_slow",
101
+ "injected_write_delay_ms": 20,
102
+ "timings": {
103
+ "update_ms": {
104
+ "samples": 180,
105
+ "p50": 0.219,
106
+ "p95": 0.304,
107
+ "max": 0.442
108
+ },
109
+ "flush_call_ms": {
110
+ "samples": 180,
111
+ "p50": 26.068,
112
+ "p95": 27.621,
113
+ "max": 28.401
114
+ },
115
+ "flush_completion_ms": {
116
+ "samples": 180,
117
+ "p50": 26.108,
118
+ "p95": 27.719,
119
+ "max": 28.528
120
+ },
121
+ "stream_timer_lateness_ms": {
122
+ "samples": 186,
123
+ "p50": 26.445,
124
+ "p95": 28.508,
125
+ "max": 29.835
126
+ },
127
+ "startup_ms": {
128
+ "samples": 3,
129
+ "p50": 25.292,
130
+ "p95": 25.292,
131
+ "max": 27.452
132
+ },
133
+ "close_ms": {
134
+ "samples": 3,
135
+ "p50": 1.203,
136
+ "p95": 1.203,
137
+ "max": 1.303
138
+ }
139
+ },
140
+ "event_loop_delay_ms_by_run": [
141
+ {
142
+ "p50": 1.151,
143
+ "p95": 27.853,
144
+ "max": 29.508
145
+ },
146
+ {
147
+ "p50": 1.156,
148
+ "p95": 28.639,
149
+ "max": 30.015
150
+ },
151
+ {
152
+ "p50": 1.158,
153
+ "p95": 28.836,
154
+ "max": 30.31
155
+ }
156
+ ],
157
+ "snapshot_writes_by_run": [
158
+ 61,
159
+ 61,
160
+ 61
161
+ ],
162
+ "rss_peak_bytes": 22678336
163
+ }
164
+ ]
165
+ },
166
+ {
167
+ "schema_version": 1,
168
+ "label": "asynchronous-single-flight",
169
+ "recorded_at": "2026-10-05T08:28:48.754Z",
170
+ "environment": {
171
+ "node": "v26.10.0",
172
+ "platform": "darwin",
173
+ "release": "25.6.0",
174
+ "arch": "arm64",
175
+ "cpu": "Apple M4",
176
+ "logical_cpus": 10,
177
+ "memory_bytes": 17179869184,
178
+ "load_average": [
179
+ 4.52685546875,
180
+ 3.9130859375,
181
+ 3.46826171875
182
+ ],
183
+ "background_load": "Other workspace activity was not suspended; no process identities or personal data captured."
184
+ },
185
+ "method": {
186
+ "rounds_per_run": 60,
187
+ "repetitions": 3,
188
+ "events_per_round": 6,
189
+ "concurrent_sessions": 20,
190
+ "producer_interval_ms": 5,
191
+ "mock_stream_interval_ms": 5,
192
+ "injection": "One 20ms delay per snapshot write: blocking wait for synchronous writer, timer wait for asynchronous writer. Metadata operations are not delayed.",
193
+ "note": "Stream timer lateness measures event-loop interference, not a network benchmark. Flush completion may include coalesced subsequent snapshots."
194
+ },
195
+ "scenarios": [
196
+ {
197
+ "storage": "normal_temporary_directory",
198
+ "injected_write_delay_ms": 0,
199
+ "timings": {
200
+ "update_ms": {
201
+ "samples": 180,
202
+ "p50": 0.121,
203
+ "p95": 0.185,
204
+ "max": 5.922
205
+ },
206
+ "flush_call_ms": {
207
+ "samples": 180,
208
+ "p50": 0.004,
209
+ "p95": 0.006,
210
+ "max": 0.016
211
+ },
212
+ "flush_completion_ms": {
213
+ "samples": 180,
214
+ "p50": 1.387,
215
+ "p95": 2.341,
216
+ "max": 4.168
217
+ },
218
+ "stream_timer_lateness_ms": {
219
+ "samples": 189,
220
+ "p50": 0.215,
221
+ "p95": 0.955,
222
+ "max": 1.381
223
+ },
224
+ "startup_ms": {
225
+ "samples": 3,
226
+ "p50": 2.103,
227
+ "p95": 2.103,
228
+ "max": 3.156
229
+ },
230
+ "close_ms": {
231
+ "samples": 3,
232
+ "p50": 2.203,
233
+ "p95": 2.203,
234
+ "max": 2.358
235
+ }
236
+ },
237
+ "event_loop_delay_ms_by_run": [
238
+ {
239
+ "p50": 1.143,
240
+ "p95": 1.747,
241
+ "max": 7.68
242
+ },
243
+ {
244
+ "p50": 1.151,
245
+ "p95": 1.838,
246
+ "max": 4.184
247
+ },
248
+ {
249
+ "p50": 1.152,
250
+ "p95": 1.766,
251
+ "max": 2.052
252
+ }
253
+ ],
254
+ "snapshot_writes_by_run": [
255
+ 61,
256
+ 61,
257
+ 61
258
+ ],
259
+ "rss_peak_bytes": 23284496
260
+ },
261
+ {
262
+ "storage": "injected_slow",
263
+ "injected_write_delay_ms": 20,
264
+ "timings": {
265
+ "update_ms": {
266
+ "samples": 180,
267
+ "p50": 0.125,
268
+ "p95": 0.193,
269
+ "max": 0.28
270
+ },
271
+ "flush_call_ms": {
272
+ "samples": 180,
273
+ "p50": 0.004,
274
+ "p95": 0.006,
275
+ "max": 0.025
276
+ },
277
+ "flush_completion_ms": {
278
+ "samples": 180,
279
+ "p50": 199.18,
280
+ "p95": 342.723,
281
+ "max": 361.781
282
+ },
283
+ "stream_timer_lateness_ms": {
284
+ "samples": 212,
285
+ "p50": 0.065,
286
+ "p95": 1.042,
287
+ "max": 1.299
288
+ },
289
+ "startup_ms": {
290
+ "samples": 3,
291
+ "p50": 22.754,
292
+ "p95": 22.754,
293
+ "max": 23.073
294
+ },
295
+ "close_ms": {
296
+ "samples": 3,
297
+ "p50": 1.595,
298
+ "p95": 1.595,
299
+ "max": 2.428
300
+ }
301
+ },
302
+ "event_loop_delay_ms_by_run": [
303
+ {
304
+ "p50": 1.158,
305
+ "p95": 1.509,
306
+ "max": 2.224
307
+ },
308
+ {
309
+ "p50": 1.154,
310
+ "p95": 1.451,
311
+ "max": 2.001
312
+ },
313
+ {
314
+ "p50": 1.154,
315
+ "p95": 1.501,
316
+ "max": 2.061
317
+ }
318
+ ],
319
+ "snapshot_writes_by_run": [
320
+ 17,
321
+ 17,
322
+ 17
323
+ ],
324
+ "rss_peak_bytes": 23874368
325
+ }
326
+ ],
327
+ "baseline_comparison": {
328
+ "baseline_label": "synchronous-baseline",
329
+ "scope": "Local synthetic regression thresholds derived from this recorded baseline; not portable latency promises or ordinary CI gates.",
330
+ "gates": [
331
+ {
332
+ "name": "slow_storage_stream_p95_ms",
333
+ "actual": 1.042,
334
+ "maximum": 7.127,
335
+ "passed": true,
336
+ "rationale": "Require at least 75% less interference than the measured blocking-storage baseline."
337
+ },
338
+ {
339
+ "name": "slow_storage_flush_call_p95_ms",
340
+ "actual": 0.006,
341
+ "maximum": 2.762,
342
+ "passed": true,
343
+ "rationale": "Require at least 90% less synchronous caller time when storage is slow."
344
+ },
345
+ {
346
+ "name": "normal_storage_stream_p95_ms",
347
+ "actual": 0.955,
348
+ "maximum": 2.95,
349
+ "passed": true,
350
+ "rationale": "Allow 25% baseline headroom for scheduler noise while checking normal-storage regression."
351
+ },
352
+ {
353
+ "name": "normal_storage_update_p95_ms",
354
+ "actual": 0.185,
355
+ "maximum": 0.378,
356
+ "passed": true,
357
+ "rationale": "Allow twice the small pre-existing normalization/update cost; this is not a disk-latency bound."
358
+ }
359
+ ]
360
+ }
361
+ }
362
+ ]
363
+ }
@@ -0,0 +1,44 @@
1
+ # Status persistence benchmark
2
+
3
+ Status snapshots now use asynchronous atomic replacement with one writer and one dirty flag. Slow storage can delay the status display without blocking request handling or response streaming. Events received during a write are folded into the latest bounded session state; the writer does not queue a snapshot for every event.
4
+
5
+ The [recorded results](status-performance.json) compare the original synchronous writer, captured before implementation, with the asynchronous writer on an Apple M4, 16 GiB RAM, macOS Darwin 25.6.0, Node 26.10.0. Other workspace activity continued during both runs; load averages and peak process RSS are recorded. These are local synthetic observations, not provider latency or a cross-machine performance claim.
6
+
7
+ Each storage scenario has three runs of 60 bursts, six telemetry events per burst, 20 rotating sessions, and a 5 ms producer interval. A separate 5 ms timer feeds a local mock stream. The slow-storage case injects 20 ms into each snapshot write: a blocking wait for the original writer and an asynchronous wait for its replacement. It does not delay filesystem metadata operations. The stream measurement is lateness beyond its expected timer interval, not network throughput.
8
+
9
+ | Measurement, milliseconds unless stated | Normal synchronous | Normal asynchronous | Slow synchronous | Slow asynchronous |
10
+ | --- | ---: | ---: | ---: | ---: |
11
+ | Update burst p50 / p95 | 0.090 / 0.189 | 0.121 / 0.185 | 0.219 / 0.304 | 0.125 / 0.193 |
12
+ | `flush()` caller p50 / p95 | 1.070 / 1.955 | 0.004 / 0.006 | 26.068 / 27.621 | 0.004 / 0.006 |
13
+ | Mock stream lateness p50 / p95 | 0.996 / 2.360 | 0.215 / 0.955 | 26.445 / 28.508 | 0.065 / 1.042 |
14
+ | Snapshot writes per run, including initial file | 61 | 61 | 61 | 17 |
15
+ | Peak process RSS, MiB | 21.11 | 22.21 | 21.63 | 22.77 |
16
+
17
+ Under injected slow storage, stream lateness p95 fell by 96.3%. The asynchronous `flush()` promise still waits for persistence: its slow-storage completion p95 was 342.723 ms because it drains updates arriving during that write sequence. Request handling calls `update()` and does not await this drain. The JSON includes completion latency, initialization, shutdown, event-loop histograms, and maximum values, including scheduler outliers omitted from the compact table.
18
+
19
+ ## Local regression gates
20
+
21
+ The harness derives these gates from the recorded synchronous baseline. All four passed on the measured machine:
22
+
23
+ | Gate | Maximum | Measured |
24
+ | --- | ---: | ---: |
25
+ | Slow-storage stream lateness p95: at least 75% lower | 7.127 ms | 1.042 ms |
26
+ | Slow-storage synchronous flush cost p95: at least 90% lower | 2.762 ms | 0.006 ms |
27
+ | Normal-storage stream lateness p95: baseline plus 25% | 2.950 ms | 0.955 ms |
28
+ | Normal-storage update p95: twice baseline | 0.378 ms | 0.185 ms |
29
+
30
+ These numerical checks are opt-in and sensitive to host load; they do not run as ordinary CI timing tests. Repeat locally with:
31
+
32
+ ```sh
33
+ node scripts/benchmark-status.mjs --label local-check --check
34
+ ```
35
+
36
+ This appends or replaces that label in `docs/status-performance.json`. `--output PATH` selects a separate report; `--module PATH` can benchmark another checkout's `src/status-state.mjs`. `--check` requires an existing `synchronous-baseline` entry in the selected report. No sockets, inference, credentials, or user session data are used.
37
+
38
+ ## Persistence contract
39
+
40
+ `createStatusState()` returns immediately. Callers await `ready` before reading `path`; the path becomes available only after the first complete snapshot exists. Initial storage failure or a one-second readiness deadline disables the optional status display. Accepted snapshots use a private directory with mode `0700`, exclusive temporary files with mode `0600`, and atomic rename. Filesystem errors do not reject `ready`, `flush()`, or `close()`.
41
+
42
+ `close()` stops new updates, waits for the active write and latest accepted state, then removes the private directory. Native filesystem operations cannot be cancelled, so shutdown retains ownership of an outstanding operation even if startup readiness has already timed out. It never removes the directory while a detached writer could recreate a file.
43
+
44
+ Deterministic tests hold a write unresolved while an entire mock inference stream finishes; 2,000 update bursts produce only one subsequent snapshot. Other tests cover concurrent close calls, late updates, readiness timeout, temporary-file cleanup, recovery after a failed rename, permissions, and foreground/session isolation.
package/package.json CHANGED
@@ -1,16 +1,58 @@
1
1
  {
2
2
  "name": "claude-autorouter",
3
- "version": "0.3.6",
3
+ "version": "0.4.0",
4
4
  "license": "Apache-2.0",
5
5
  "type": "module",
6
6
  "description": "A Claude Code model router with Jev and local Ollama System One evaluators",
7
- "repository": { "type": "git", "url": "git+https://github.com/frapposelli/claude-autorouter.git" },
8
- "bin": { "claude-autorouter": "bin/autorouter.mjs" },
9
- "engines": { "node": ">=22" },
10
- "os": ["darwin", "linux"],
11
- "keywords": ["claude", "claude-code", "model-routing", "jev", "ollama", "nimble", "tev1", "systemone", "cli"],
12
- "files": ["bin/*.mjs", "src/*.mjs", "docs/reference.md", "docs/development.md", "docs/releasing.md", "docs/ollama-evaluation.md", ".env.example", "LICENSE"],
13
- "publishConfig": { "access": "public", "registry": "https://registry.npmjs.org/" },
7
+ "repository": {
8
+ "type": "git",
9
+ "url": "git+https://github.com/frapposelli/claude-autorouter.git"
10
+ },
11
+ "bin": {
12
+ "claude-autorouter": "bin/autorouter.mjs"
13
+ },
14
+ "engines": {
15
+ "node": ">=22"
16
+ },
17
+ "os": [
18
+ "darwin",
19
+ "linux"
20
+ ],
21
+ "keywords": [
22
+ "claude",
23
+ "claude-code",
24
+ "model-routing",
25
+ "jev",
26
+ "ollama",
27
+ "nimble",
28
+ "tev1",
29
+ "systemone",
30
+ "cli"
31
+ ],
32
+ "files": [
33
+ "bin/*.mjs",
34
+ "src/*.mjs",
35
+ "docs/reference.md",
36
+ "docs/development.md",
37
+ "docs/releasing.md",
38
+ "docs/ollama-evaluation.md",
39
+ ".env.example",
40
+ "LICENSE",
41
+ "CONTRIBUTING.md",
42
+ "docs/router-performance.md",
43
+ "docs/router-performance.json",
44
+ "docs/status-performance.md",
45
+ "docs/status-performance.json",
46
+ "docs/hardware-benchmark.md",
47
+ "docs/hardware-results-16gb.md",
48
+ "docs/hardware-results-16gb.json",
49
+ "docs/hardware-comparison.md",
50
+ "docs/hardware-results-64gb.json"
51
+ ],
52
+ "publishConfig": {
53
+ "access": "public",
54
+ "registry": "https://registry.npmjs.org/"
55
+ },
14
56
  "scripts": {
15
57
  "start": "node bin/autorouter.mjs serve",
16
58
  "claude": "node bin/autorouter.mjs claude",
@@ -21,6 +63,12 @@
21
63
  "eval": "node --env-file=.env scripts/evaluate.mjs",
22
64
  "eval:ollama": "node scripts/evaluate-ollama.mjs",
23
65
  "test:live": "node --env-file=.env scripts/live-validation.mjs",
24
- "check": "node scripts/check.mjs"
66
+ "check": "node scripts/check.mjs",
67
+ "benchmark:router": "node --expose-gc scripts/benchmark-router.mjs",
68
+ "benchmark:status": "node scripts/benchmark-status.mjs"
69
+ },
70
+ "devDependencies": {
71
+ "@types/node": "22.20.5",
72
+ "typescript": "6.0.3"
25
73
  }
26
74
  }