loadout-ai 0.7.0 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/CHANGELOG.md +72 -0
  2. package/README.md +33 -33
  3. package/catalog/discovered.json +26880 -24184
  4. package/dist/src/cli.js +5 -0
  5. package/dist/src/commands/catalog.js +103 -116
  6. package/dist/src/core/agents/agent-inspection.js +26 -4
  7. package/dist/src/core/catalog/registry.js +58 -10
  8. package/dist/src/core/catalog/safety.js +36 -7
  9. package/dist/src/core/install/source.js +21 -7
  10. package/dist/src/core/reporting/cli-guide.js +3 -3
  11. package/dist/src/core/reporting/completion.js +42 -95
  12. package/dist/src/core/reporting/doctor.js +3 -5
  13. package/dist/src/core/routing/handoff.js +94 -58
  14. package/dist/src/core/routing/policy.js +147 -0
  15. package/dist/src/core/routing/route.js +25 -153
  16. package/docs/CANDIDATE_INTELLIGENCE.md +9 -2
  17. package/docs/CATALOG.md +1 -1
  18. package/docs/CREDENTIAL_AND_UPDATE_POLICY.md +1 -1
  19. package/docs/DISCOVERED.md +252 -251
  20. package/docs/FEATURE_TEST_MATRIX.md +7 -260
  21. package/docs/GITHUB_AUTHORIZATION.md +5 -0
  22. package/docs/PROVENANCE_AND_COMPARISON.md +1 -1
  23. package/docs/RELEASE_REVIEW.md +0 -1
  24. package/package.json +6 -4
  25. package/skills/loadout-router/SKILL.md +43 -78
  26. package/MASTER_PLAN.md +0 -2207
  27. package/docs/ACTIVE_SET.md +0 -53
  28. package/docs/COMPATIBILITY_POLICY.md +0 -22
  29. package/docs/CONVERSION_AND_SANDBOX.md +0 -27
  30. package/docs/EVALUATION_PROTOCOL_V1.md +0 -300
  31. package/docs/HEAD_TO_HEAD_EVALUATION.md +0 -79
  32. package/docs/PROVIDER_CONFIGURATION.md +0 -45
  33. package/docs/README_RESEARCH.md +0 -36
  34. package/docs/REPOSITORY_STABILIZATION.md +0 -190
  35. package/docs/SAFE_UPDATE_DEMO.md +0 -25
  36. package/docs/SCHEMA_DECISIONS.md +0 -25
  37. package/docs/SUBMISSION_COPY.md +0 -90
  38. package/docs/TEAM_POLICY.md +0 -18
  39. package/docs/superpowers/plans/2026-07-19-relatable-readme-hero.md +0 -283
  40. package/docs/superpowers/plans/2026-07-20-loadout-readme-explainer.md +0 -116
  41. package/docs/superpowers/plans/2026-07-20-project-activation-safety.md +0 -469
  42. package/docs/superpowers/specs/2026-07-19-relatable-readme-hero-design.md +0 -80
  43. package/docs/superpowers/specs/2026-07-20-loadout-readme-explainer-design.md +0 -55
  44. package/docs/superpowers/specs/2026-07-20-project-activation-safety-design.md +0 -228
@@ -98,7 +98,6 @@ Confirm isolation before any applied command:
98
98
 
99
99
  ```bash
100
100
  loadout status --json
101
- loadout capabilities
102
101
  ```
103
102
 
104
103
  Expected: paths, if shown, are below `TEST_ROOT`; Codex and Claude Code are detected;
@@ -158,8 +157,6 @@ loadout library --json
158
157
  loadout health --explain --json
159
158
  loadout report --json
160
159
  loadout card --json
161
- loadout outcomes --json
162
- loadout capabilities --inspect --json
163
160
  ```
164
161
 
165
162
  Expected:
@@ -184,7 +181,6 @@ Network variants must be run deliberately:
184
181
  loadout catalog --refresh --json
185
182
  loadout health --updates --json
186
183
  loadout scan --agents codex,claude-code --refresh-provenance --json
187
- loadout compare brainstorming --offline --json
188
184
  loadout discover --source mcp-registry --limit 10 --json
189
185
  loadout discover --source skills-sh --limit 10 --json
190
186
  ```
@@ -201,79 +197,7 @@ skills.sh path needs its request-scoped `VERCEL_OIDC_TOKEN`; without one it must
201
197
  previous complete cache or return an attributed `unavailable` result without making
202
198
  an unauthenticated request. Neither source installs or promotes a lead.
203
199
 
204
- ## 3. Package, manifest, lock, portability, and registry track (S/A)
205
-
206
- Create a package _inside_ the test project so its manifest can be exported portably:
207
-
208
- ```bash
209
- mkdir -p "$TEST_PROJECT/packages"
210
- loadout create "$TEST_PROJECT/packages/matrix-demo" \
211
- --name matrix-demo --description "Disposable matrix package"
212
- loadout pack "$TEST_PROJECT/packages/matrix-demo" --json
213
- loadout publish "$TEST_PROJECT/packages/matrix-demo" --local
214
- loadout search matrix-demo --json
215
-
216
- loadout init --path "$TEST_PROJECT/loadout.json" --name matrix \
217
- --agents codex,claude-code --scope project
218
- (
219
- cd "$TEST_PROJECT"
220
- loadout add local-demo --manifest loadout.json --local \
221
- --path packages/matrix-demo --agents codex,claude-code
222
- loadout sync --manifest loadout.json --lock loadout.lock
223
- loadout sync --manifest loadout.json --lock loadout.lock --yes
224
- loadout audit --manifest loadout.json --lock loadout.lock --json
225
- )
226
- ```
227
-
228
- Expected: `pack` returns a deterministic digest; local publication is immutable;
229
- the first `sync` is a dry run; applied sync prints one snapshot; `audit` returns
230
- `"valid": true`; skills appear only below the disposable profile.
231
-
232
- Exercise desired-state editing and portability:
233
-
234
- ```bash
235
- (
236
- cd "$TEST_PROJECT"
237
- loadout lock --manifest loadout.json --output loadout.lock
238
- loadout export portable.json --manifest loadout.json --lock loadout.lock
239
- loadout import portable.json --manifest imported.json --lock imported.lock
240
- loadout import portable.json --manifest imported.json --lock imported.lock --yes
241
- loadout unadd local-demo --manifest imported.json
242
- )
243
- ```
244
-
245
- Expected: import previews before writing; applied import snapshots destinations;
246
- `unadd` changes desired state only and does not delete installed files.
247
-
248
- The full authenticated remote-registry protocol, including wrong-token rejection,
249
- immutable-version conflict rejection, exact digest download, risk approval, and the
250
- HTTPS/non-loopback boundary, is reproducibly covered by:
251
-
252
- ```bash
253
- npx vitest run tests/registry-api.test.ts tests/package.test.ts
254
- ```
255
-
256
- Manual `registry-serve`/remote `publish` is an X test. After completing the throwaway
257
- credential setup in track 9, start this in one terminal:
258
-
259
- ```bash
260
- loadout registry-serve --port 7331 \
261
- --credential-keychain loadout-registry-test --credential-account tester
262
- ```
263
-
264
- Then publish from another terminal:
265
-
266
- ```bash
267
- loadout publish "$TEST_PROJECT/packages/matrix-demo" \
268
- --registry-url http://127.0.0.1:7331 \
269
- --credential-keychain loadout-registry-test --credential-account tester
270
- ```
271
-
272
- Never place the token on the command line. An identical version/content publish may
273
- be accepted idempotently; changed content at the same version must be rejected. Stop
274
- the server with Ctrl-C.
275
-
276
- ## 4. Install, active-set, outcome, and rollback track (A; Maximum is N)
200
+ ## 3. Install, active-set, and rollback track (A; Maximum is N)
277
201
 
278
202
  Exercise the new-user golden path before its constituent commands:
279
203
 
@@ -321,9 +245,6 @@ loadout disable direct-demo --agents codex
321
245
  loadout disable direct-demo --agents codex --yes --json
322
246
  loadout enable direct-demo --agents codex
323
247
  loadout enable direct-demo --agents codex --yes --json
324
- loadout outcome direct-demo/direct-demo --agent codex --task testing \
325
- --result success
326
- loadout outcomes --json
327
248
  loadout share "$TEST_PROJECT/share.json"
328
249
  loadout remove direct-demo
329
250
  loadout remove direct-demo --yes
@@ -355,7 +276,7 @@ and preserves matching active Stable units at the same reviewed commit; MCP-only
355
276
  packages remain explicit setup items. `--approve-risk` acknowledges displayed static
356
277
  findings but does not execute third-party repository scripts.
357
278
 
358
- ## 5. Existing-skill provenance, adoption, comparison, and freshness (R/S/A)
279
+ ## 4. Existing-skill provenance, adoption, and freshness (R/S/A)
359
280
 
360
281
  Copy one harmless skill into the disposable unmanaged profile, then inspect it:
361
282
 
@@ -366,7 +287,6 @@ cp "$TEST_PROJECT/packages/matrix-demo/skills/matrix-demo/SKILL.md" \
366
287
  loadout scan --agents codex --json
367
288
  loadout adopt unmanaged-demo --agent codex --json
368
289
  loadout adopt unmanaged-demo --agent codex --yes --json
369
- loadout compare unmanaged-demo --agent codex --offline --json
370
290
  ```
371
291
 
372
292
  Expected: scan labels the copy unmanaged; adoption preview does not change its bytes;
@@ -401,13 +321,12 @@ Only run the second command when the first reports a real reviewed update. Expec
401
321
  exact diff/safety plan, a snapshot on success, and refusal when new risky findings are
402
322
  not acknowledged with `--approve-risk`.
403
323
 
404
- ## 6. Discovery and human review queue (N/S)
324
+ ## 5. Discovery and human review queue (N/S)
405
325
 
406
326
  ```bash
407
327
  loadout discover --source hacker-news --limit 20 --min-score 20 --json
408
328
  loadout discover --source github --limit 20 --queue --json
409
329
  loadout discover --source all --limit 20 --queue --json
410
- loadout review-queue --decision pending --json
411
330
  ```
412
331
 
413
332
  Expected: results contain source evidence and public repository identifiers; queueing
@@ -415,12 +334,6 @@ deduplicates leads; nothing is promoted, cloned into an agent, or installed.
415
334
 
416
335
  For a repository printed by the queue:
417
336
 
418
- ```bash
419
- loadout review owner/repository --decision shortlisted
420
- loadout review-queue --decision shortlisted --json
421
- loadout review owner/repository --decision ignored
422
- ```
423
-
424
337
  Private GitHub discovery is opt-in and reads `GITHUB_TOKEN` only when `--private` is
425
338
  present. Prefer a native credential reference:
426
339
 
@@ -431,7 +344,7 @@ loadout discover --source github --private \
431
344
 
432
345
  Use a low-scope test token. The output and state must never contain its value.
433
346
 
434
- ## 7. Static inspection, MCP, conversion, canary, and sandbox (R/S/X)
347
+ ## 6. Static inspection, MCP, and conversion (R/S)
435
348
 
436
349
  Static package analysis never executes package content:
437
350
 
@@ -439,8 +352,6 @@ Static package analysis never executes package content:
439
352
  loadout inspect --source "$TEST_PROJECT/packages/matrix-demo" --json
440
353
  loadout evaluate --source "$TEST_PROJECT/packages/matrix-demo" --json
441
354
  loadout mcp --source "$TEST_PROJECT/packages/matrix-demo" --json
442
- loadout canary --source "$TEST_PROJECT/packages/matrix-demo" \
443
- --package matrix-demo --json
444
355
  ```
445
356
 
446
357
  Repeat `inspect`, `evaluate`, or `mcp` with `--repository owner/repository` for the
@@ -469,10 +380,6 @@ loadout mcp-config --config "$TEST_PROJECT/mcp.json" --name local-example \
469
380
  --command node --arg server.js
470
381
  loadout mcp-config --config "$TEST_PROJECT/mcp.json" --name local-example \
471
382
  --command node --arg server.js --yes
472
- loadout codex-mcp-config --config "$TEST_PROJECT/config.toml" \
473
- --name remote-example --url https://example.com/mcp
474
- loadout codex-mcp-config --config "$TEST_PROJECT/config.toml" \
475
- --name remote-example --url https://example.com/mcp --yes
476
383
  loadout mcp-recipe --json
477
384
  ```
478
385
 
@@ -486,148 +393,11 @@ requirements you have reviewed.
486
393
 
487
394
  Docker sandbox execution is intentionally separate:
488
395
 
489
- ```bash
490
- loadout sandbox-run --source "$TEST_PROJECT/packages/matrix-demo" \
491
- --image '<reviewed-image>@sha256:<digest>' \
492
- --command node --command --version --json
493
- loadout sandbox-run --source "$TEST_PROJECT/packages/matrix-demo" \
494
- --image '<reviewed-image>@sha256:<digest>' \
495
- --command node --command --version \
496
- --approve-risk --timeout 30000 --json
497
- ```
498
-
499
396
  Expected: the first invocation refuses/only plans without approval; the approved
500
397
  container has a read-only source mount, no inherited secrets, no Docker socket, no
501
398
  network, and a time bound. The image may need to be pulled beforehand.
502
399
 
503
- ## 8. Signing and head-to-head evidence (S)
504
-
505
- Validate the model-free benchmark campaign and card/compare surfaces with their
506
- deterministic automated contracts:
507
-
508
- ```bash
509
- npx vitest run tests/benchmark-campaign.test.ts tests/benchmark-cli.test.ts \
510
- tests/loadout-card.test.ts tests/share-report.test.ts
511
- ```
512
-
513
- Expected: campaign hashes and paired order are deterministic, every retry is included
514
- in the worst-case budget, over-budget plans are blocked, resumable metadata contains
515
- no prompt/output/credential bytes, and aggregate comparison never invents a quality
516
- delta. See `docs/EVALUATION_PROTOCOL_V1.md` for the campaign JSON contract. These
517
- tests do not call a model provider and consume no provider credit.
518
-
519
- ```bash
520
- loadout keygen --private-key "$TEST_ROOT/private.pem" \
521
- --public-key "$TEST_ROOT/public.pem"
522
- loadout catalog-sign --catalog "$LOADOUT_ROOT/catalog/packages.json" \
523
- --private-key "$TEST_ROOT/private.pem" --output "$TEST_ROOT/catalog.signed.json"
524
- loadout catalog-verify --snapshot "$TEST_ROOT/catalog.signed.json" \
525
- --public-key "$TEST_ROOT/public.pem"
526
- ```
527
-
528
- Expected: the private key is owner-only and outside the repository; verification
529
- succeeds; changing any byte in the signed payload makes verification fail.
530
-
531
- Preview and apply the same signed catalog inside the disposable profile:
532
-
533
- ```bash
534
- loadout catalog-update --source "$TEST_ROOT/catalog.signed.json" \
535
- --public-key "$TEST_ROOT/public.pem"
536
- loadout catalog-update --source "$TEST_ROOT/catalog.signed.json" \
537
- --public-key "$TEST_ROOT/public.pem" --yes
538
- loadout catalog --coverage --json
539
- ```
540
-
541
- Expected: preview prints an exact signed diff without mutation; apply creates a
542
- snapshot and trusted state; the effective catalog re-verifies the stored envelope.
543
- Repeating `--yes` refuses a replay. Test removal only in the disposable profile and
544
- only with the separate `--allow-removals` acknowledgement.
545
-
546
- The repository's generated feed can be triaged without network access:
547
-
548
- ```bash
549
- loadout candidate list --limit 5 --json
550
- loadout candidate list --query "codex skills"
551
- loadout capabilities --gaps --json
552
- loadout recommend --project "$TEST_PROJECT" --agent codex --json
553
- ```
554
-
555
- `candidate inspect owner/repository --output ./candidate-dossier.json` is a networked
556
- test: it performs a real public Git clone and writes a static immutable dossier to
557
- disposable Loadout state. Review that output before exercising `candidate propose`;
558
- proposal preview and approved proposal output never mutate the catalog.
559
-
560
- Graphify is an explicit executable recipe rather than a broad-setup component. With
561
- `uv` installed, exercise it only inside the disposable profile:
562
-
563
- ```bash
564
- loadout tool
565
- loadout tool graphify --agents codex
566
- loadout tool graphify --agents codex --yes --approve-risk
567
- "$LOADOUT_HOME/runtime/graphify/bin/graphify" --version
568
- test -f "$LOADOUT_USER_HOME/.codex/skills/graphify/SKILL.md"
569
- loadout tool graphify --remove
570
- loadout tool graphify --remove --yes --approve-risk
571
- test ! -e "$LOADOUT_USER_HOME/.codex/skills/graphify"
572
- test ! -e "$LOADOUT_HOME/runtime/graphify"
573
- ```
574
-
575
- Expected: preview identifies the exact wheel hash and all commands; apply reports
576
- Graphify 0.9.17, writes only the disposable target and isolated runtime, and removal
577
- restores the original target. The installer subprocess must not inherit API keys.
578
-
579
- Create a deterministic workflow fixture and five declared trials per candidate. This
580
- is harness input, not model-generated evidence, and it executes no candidate content:
581
-
582
- ```bash
583
- node --input-type=module <<'NODE'
584
- import { writeFile } from "node:fs/promises";
585
- const root = process.env.TEST_PROJECT;
586
- const fixture = {
587
- id: "matrix-workflow",
588
- version: "1",
589
- category: "workflow-adherence",
590
- requiredActions: ["inspect", "edit", "verify"],
591
- forbiddenActions: ["delete-unrelated"]
592
- };
593
- const trials = Array.from({ length: 5 }, () => [
594
- {
595
- candidateId: "baseline",
596
- fixtureId: fixture.id,
597
- observations: ["inspect", "edit", "verify"],
598
- durationMs: 10
599
- },
600
- {
601
- candidateId: "improved",
602
- fixtureId: fixture.id,
603
- observations: ["inspect", "edit", "verify", "report-uncertainty"],
604
- durationMs: 10
605
- }
606
- ]).flat();
607
- await writeFile(`${root}/fixture.json`, JSON.stringify(fixture, null, 2));
608
- await writeFile(`${root}/trials.json`, JSON.stringify(trials, null, 2));
609
- NODE
610
- ```
611
-
612
- Sign and inspect the resulting evidence:
613
-
614
- ```bash
615
- loadout head-to-head --fixture "$TEST_PROJECT/fixture.json" \
616
- --trials "$TEST_PROJECT/trials.json" --private-key "$TEST_ROOT/private.pem" \
617
- --output "$TEST_ROOT/evidence.json" --json
618
- loadout alerts --evidence "$TEST_ROOT/evidence.json" \
619
- --public-key "$TEST_ROOT/public.pem" --json
620
- ```
621
-
622
- The harness scores declared observations only. It never executes candidate content.
623
- The authoritative schema, safety-failure, minimum-trial, tamper, and practical-delta
624
- tests are also directly runnable:
625
-
626
- ```bash
627
- npx vitest run tests/head-to-head.test.ts tests/signing.test.ts
628
- ```
629
-
630
- ## 9. Credentials and model-provider verification (H/$)
400
+ ## 7. Credentials and model-provider verification (H/$)
631
401
 
632
402
  This track touches the real operating-system credential store even when
633
403
  `LOADOUT_HOME` is disposable. Use a unique throwaway service name and delete it.
@@ -668,7 +438,7 @@ unset OPENROUTER_API_KEY
668
438
  `models verify` makes one minimal request and may consume provider credit ($). Inspect
669
439
  provider billing before and after; do not run it in a loop.
670
440
 
671
- ## 10. Watchers, native scheduling, completions, and loopback UI/API (X/H)
441
+ ## 8. Watchers, native scheduling, completions, and loopback UI/API (X/H)
672
442
 
673
443
  One-shot update watching is safe and networked:
674
444
 
@@ -707,33 +477,10 @@ and `unschedule` commands remain available for job-specific control.
707
477
 
708
478
  The optional read-only loopback API can be checked separately:
709
479
 
710
- ```bash
711
- loadout serve --port 0
712
- ```
713
-
714
480
  Confirm that it binds only to `127.0.0.1`, inspect the API response, and stop it with
715
481
  Ctrl-C. It must not bind a public interface.
716
482
 
717
- ## 11. Improvement-cycle records (S)
718
-
719
- ```bash
720
- loadout improve --json
721
- loadout improve --write --output "$TEST_PROJECT/improvements" --json
722
- ```
723
-
724
- Copy the exact cycle id printed by the second command:
725
-
726
- ```bash
727
- loadout improve-feedback --id <cycle-id> --outcome partial \
728
- --note "Disposable matrix verification" \
729
- --directory "$TEST_PROJECT/improvements"
730
- ```
731
-
732
- Expected: the first command is read-only; `--write` persists a local prompt/cycle
733
- record only; feedback requires a human-selected outcome and stores no project source
734
- or prompt transcript.
735
-
736
- ## 12. Cleanup and pass criteria
483
+ ## 9. Cleanup and pass criteria
737
484
 
738
485
  First verify snapshot availability and roll back any remaining disposable mutation:
739
486
 
@@ -19,6 +19,11 @@ paste a token into a manifest, catalog, log, or command argument.
19
19
 
20
20
  ## Local flow and failure modes
21
21
 
22
+ > **Partly unimplemented.** `loadout connect` and `loadout disconnect` do not
23
+ > exist in the CLI today. `loadout discover --private` is real and reads
24
+ > credentials the user has already configured. Steps 1, 2, 3, and 5 describe the
25
+ > intended flow, not current behaviour.
26
+
22
27
  1. `loadout connect github` opens a browser using PKCE and a loopback callback.
23
28
  2. The local process verifies `state`, PKCE verifier, expiration, and callback host.
24
29
  3. It stores only an OS-keychain reference to the refresh/session material.
@@ -26,7 +26,7 @@ their `SKILL.md` files, and writes a local index below `LOADOUT_HOME/provenance`
26
26
 
27
27
  ## Relationship classification
28
28
 
29
- `loadout compare` uses deterministic relationships:
29
+ `loadout scan` uses deterministic relationships:
30
30
 
31
31
  - `exact-copy`: identical instruction fingerprint;
32
32
  - `divergent-same-name`: same normalized name, different instructions;
@@ -19,7 +19,6 @@ For current behavior and evidence, use:
19
19
  catalog/support facts;
20
20
  - the [changelog](../CHANGELOG.md) for released behavior;
21
21
  - the [feature test matrix](./FEATURE_TEST_MATRIX.md) for adapter evidence;
22
- - the [repository stabilization record](./REPOSITORY_STABILIZATION.md) and
23
22
  [sanitized July 19 live checks](./evidence/live-checks-2026-07-19.json) only when
24
23
  investigating that historical release line.
25
24
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "loadout-ai",
3
- "version": "0.7.0",
3
+ "version": "0.8.0",
4
4
  "private": false,
5
5
  "license": "MIT",
6
6
  "description": "Universal upgrade manager for AI coding agents",
@@ -15,7 +15,6 @@
15
15
  "docs",
16
16
  "README.md",
17
17
  "CHANGELOG.md",
18
- "MASTER_PLAN.md",
19
18
  "SECURITY.md",
20
19
  "LICENSE"
21
20
  ],
@@ -57,6 +56,7 @@
57
56
  "test:e2e:readme": "node scripts/readme-product-flow.mjs",
58
57
  "test:package": "node scripts/package-smoke.mjs",
59
58
  "pretest:performance": "npm run build",
59
+ "test:coverage": "vitest run --coverage",
60
60
  "test:performance": "node scripts/scan-benchmark.mjs",
61
61
  "typecheck": "tsc -p tsconfig.json --noEmit",
62
62
  "readme:update": "node scripts/update-readme-facts.mjs",
@@ -65,8 +65,9 @@
65
65
  "check:live": "node scripts/check-live-evidence.mjs",
66
66
  "check:allowlists": "npm run build && node scripts/check-mode-allowlists.mjs",
67
67
  "check:catalog-freshness": "npm run build && node scripts/check-catalog-freshness.mjs",
68
- "check:evidence": "node scripts/check-catalog-attribution.mjs && node scripts/check-discovery-artifacts.mjs && npm run check:readme-claims",
69
- "verify": "npm run format:check && npm run lint && npm run typecheck && npm run check:evidence && npm test -- --run && npm run test:e2e:cli && npm run test:e2e:readme && npm run test:package && npm run test:performance",
68
+ "check:audit": "npm audit --audit-level=high",
69
+ "check:evidence": "node scripts/check-catalog-attribution.mjs && node scripts/check-discovery-artifacts.mjs && node scripts/check-documented-commands.mjs && npm run check:readme-claims",
70
+ "verify": "npm run format:check && npm run lint && npm run typecheck && npm run check:audit && npm run check:evidence && npm test -- --run && npm run test:e2e:cli && npm run test:e2e:readme && npm run test:package && npm run test:performance",
70
71
  "verify:full": "npm run verify"
71
72
  },
72
73
  "dependencies": {
@@ -76,6 +77,7 @@
76
77
  "devDependencies": {
77
78
  "@eslint/js": "^9.39.5",
78
79
  "@types/node": "^22.10.0",
80
+ "@vitest/coverage-v8": "^4.1.10",
79
81
  "eslint": "^9.39.5",
80
82
  "eslint-config-prettier": "^10.1.8",
81
83
  "prettier": "^3.9.5",
@@ -1,120 +1,85 @@
1
1
  ---
2
2
  name: loadout-router
3
- description: Choose the right model tier and agent for a coding task, and hand work off between Claude Code and Codex. Use when the user asks which model to use, mentions running low on usage or quota, wants to save tokens or cost, asks whether to switch to Opus/Sonnet/Haiku or a GPT tier, or wants to delegate a task to another agent.
3
+ description: Decide which model to use for a coding task and hand work to another agent. Use when the user asks which model to use, mentions running low on usage or cost, asks whether to switch to Opus or Sonnet or a GPT tier, or wants to delegate a task to Codex or Claude Code.
4
4
  ---
5
5
 
6
6
  # Loadout Router
7
7
 
8
- Route each coding task to the cheapest model that still does it well, and hand
9
- work between agents when a different one is better suited.
8
+ Pick the model that fits the work, using the user's own routing policy rather
9
+ than your guess or mine.
10
10
 
11
- This skill wraps the `loadout` CLI, so the model catalog and pricing stay
12
- current with the installed version rather than going stale in this file.
13
-
14
- ## Prerequisite
15
-
16
- Check once per session:
17
-
18
- ```bash
19
- loadout --version
20
- ```
21
-
22
- If that fails, tell the user to install it (`npm install --global loadout-ai`)
23
- and answer from general knowledge instead of guessing at specifics.
24
-
25
- ## Choosing a model
26
-
27
- Run the router with the task described in plain words:
11
+ ## The policy is the user's, not yours
28
12
 
29
13
  ```bash
30
- loadout route <task description>
14
+ loadout route
31
15
  ```
32
16
 
33
- It classifies the task into one of six phases — plan, implement, review, test,
34
- debug, document — and prints the recommended tier, the models in that tier with
35
- current prices, which of the user's installed agents can run them, and a cheaper
36
- fallback with its tradeoff.
17
+ That prints three buckets and the model the user has chosen for each:
37
18
 
38
- Report the recommendation and the reason. Name the actual model, not just the
39
- tier. If the output lists a `Conserve:` alternative, mention it only when the
40
- user cares about cost or quota, otherwise it is noise.
19
+ - **hard** — architecture, security, migrations, tricky debugging, risky review
20
+ - **normal** — most implementation, ordinary debugging, refactors
21
+ - **cheap** — tests, docs, boilerplate, renames, mechanical edits
41
22
 
42
- When the user is explicit about the phase, skip classification:
23
+ Read the policy before advising. If the user disagrees with a recommendation,
24
+ the fix is to change the policy, not to argue:
43
25
 
44
26
  ```bash
45
- loadout route --phase review
27
+ loadout route --set cheap=claude-sonnet-5
46
28
  ```
47
29
 
48
- ## When the user is low on usage
30
+ ## Your job is the bucket, not the model
49
31
 
50
- If the user mentions running out, being rate limited, conserving quota, or
51
- stretching a plan, add `--conserve`:
32
+ The CLI can guess a bucket from wording, and it says so when it does. **You
33
+ should do better**, because you have the conversation, the code, and the stakes.
34
+ Decide the bucket yourself and state it:
52
35
 
53
36
  ```bash
54
- loadout route --conserve <task description>
37
+ loadout route --bucket hard
55
38
  ```
56
39
 
57
- This drops each phase one tier and prints the tradeoff you are accepting. Say
58
- what is being given up — "shallower architectural reasoning, so review the plan
59
- more carefully" — rather than presenting it as a free win.
40
+ Judge by consequence, not vocabulary:
60
41
 
61
- Neither Claude Code nor Codex exposes remaining quota programmatically, so never
62
- claim to know how much the user has left. `--conserve` is a user-driven choice,
63
- not a measurement.
64
-
65
- ## Comparing models and cost
42
+ - Anything touching auth, payments, migrations, or data deletion is **hard**,
43
+ however small the diff.
44
+ - Unfamiliar code is harder than familiar code doing the same thing.
45
+ - A one-line change in a hot path is not cheap.
46
+ - Genuinely mechanical work — a rename, a docstring, a test for code you just
47
+ wrote — is **cheap**, and paying frontier prices for it is waste.
66
48
 
67
- ```bash
68
- loadout route --models
69
- loadout route --models --provider anthropic
70
- loadout route --models --tier fast
71
- loadout route --cost
72
- ```
49
+ Report the model and why that bucket. One or two sentences.
73
50
 
74
- Use these when the user asks what is available or what something costs. Prices
75
- are per-million-token list rates; actual spend depends on prompt size, so give
76
- ratios ("roughly 4x cheaper") rather than predicting a dollar total.
51
+ ## When cost matters
77
52
 
78
- ## Handing work to another agent
53
+ If the user mentions running low, being rate limited, or wanting to spend less,
54
+ say what a cheaper bucket would cost them in quality rather than presenting it
55
+ as free. `loadout route` shows real per-million prices for the comparison.
79
56
 
80
- When a different agent suits the task better — or the user asks to delegate —
81
- use the handoff log. Check it is set up:
57
+ Neither Claude Code nor Codex exposes remaining quota programmatically, so never
58
+ claim to know how much the user has left.
82
59
 
83
- ```bash
84
- loadout handoff status
85
- ```
60
+ ## Handing work to the other agent
86
61
 
87
- If uninitialized, run `loadout handoff init` first. Then send the task:
62
+ One command sends a task; it sets up the shared log on first use:
88
63
 
89
64
  ```bash
90
- loadout handoff send codex "write unit tests for the auth module" --context "see src/auth.ts"
65
+ loadout handoff codex "write vitest coverage for src/auth.ts" --context "zod schemas already exist"
91
66
  ```
92
67
 
93
- Put anything the receiving agent needs into `--context`: file paths, the
94
- decision you already made, what you deliberately left out. The other agent has
95
- none of this conversation.
68
+ Put everything the receiver needs into `--context` — file paths, decisions you
69
+ already made, what you deliberately left out. It has none of this conversation.
96
70
 
97
- Sending a task is a real side effect on a shared file. Confirm with the user
98
- before sending unless they asked for the handoff themselves.
71
+ Sending writes to a shared file in the user's repository, so confirm first
72
+ unless they asked for the handoff themselves.
99
73
 
100
74
  ## Reading your own inbox
101
75
 
102
- At the start of a session, and after finishing a task, check whether another
103
- agent left you work:
76
+ At session start, and after finishing a task:
104
77
 
105
78
  ```bash
106
- loadout handoff inbox claude-code
79
+ loadout handoff claude-code
107
80
  ```
108
81
 
109
- If it lists tasks, work them in order and run the `loadout handoff done <id>`
110
- command it prints for each one. If it reports none, continue as normal and do
111
- not mention it.
112
-
113
- ## Judgment this skill does not replace
82
+ Work anything listed in order, then run the `loadout handoff --done <id>`
83
+ command it prints. If nothing is pending, say nothing and carry on.
114
84
 
115
- - A "simple" task in an unfamiliar or high-risk area still deserves a stronger
116
- model. The classifier reads keywords, not stakes.
117
- - Security-sensitive, migration, and data-loss paths are worth frontier tier
118
- regardless of what phase they classify as.
119
- - If the user has already chosen a model, do not argue unless the choice is
120
- clearly wrong for the work.
85
+ `loadout handoff` with no arguments shows every pending task, both directions.