loadout-ai 0.7.0 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +72 -0
- package/README.md +33 -33
- package/catalog/discovered.json +26880 -24184
- package/dist/src/cli.js +5 -0
- package/dist/src/commands/catalog.js +103 -116
- package/dist/src/core/agents/agent-inspection.js +26 -4
- package/dist/src/core/catalog/registry.js +58 -10
- package/dist/src/core/catalog/safety.js +36 -7
- package/dist/src/core/install/source.js +21 -7
- package/dist/src/core/reporting/cli-guide.js +3 -3
- package/dist/src/core/reporting/completion.js +42 -95
- package/dist/src/core/reporting/doctor.js +3 -5
- package/dist/src/core/routing/handoff.js +94 -58
- package/dist/src/core/routing/policy.js +147 -0
- package/dist/src/core/routing/route.js +25 -153
- package/docs/CANDIDATE_INTELLIGENCE.md +9 -2
- package/docs/CATALOG.md +1 -1
- package/docs/CREDENTIAL_AND_UPDATE_POLICY.md +1 -1
- package/docs/DISCOVERED.md +252 -251
- package/docs/FEATURE_TEST_MATRIX.md +7 -260
- package/docs/GITHUB_AUTHORIZATION.md +5 -0
- package/docs/PROVENANCE_AND_COMPARISON.md +1 -1
- package/docs/RELEASE_REVIEW.md +0 -1
- package/package.json +6 -4
- package/skills/loadout-router/SKILL.md +43 -78
- package/MASTER_PLAN.md +0 -2207
- package/docs/ACTIVE_SET.md +0 -53
- package/docs/COMPATIBILITY_POLICY.md +0 -22
- package/docs/CONVERSION_AND_SANDBOX.md +0 -27
- package/docs/EVALUATION_PROTOCOL_V1.md +0 -300
- package/docs/HEAD_TO_HEAD_EVALUATION.md +0 -79
- package/docs/PROVIDER_CONFIGURATION.md +0 -45
- package/docs/README_RESEARCH.md +0 -36
- package/docs/REPOSITORY_STABILIZATION.md +0 -190
- package/docs/SAFE_UPDATE_DEMO.md +0 -25
- package/docs/SCHEMA_DECISIONS.md +0 -25
- package/docs/SUBMISSION_COPY.md +0 -90
- package/docs/TEAM_POLICY.md +0 -18
- package/docs/superpowers/plans/2026-07-19-relatable-readme-hero.md +0 -283
- package/docs/superpowers/plans/2026-07-20-loadout-readme-explainer.md +0 -116
- package/docs/superpowers/plans/2026-07-20-project-activation-safety.md +0 -469
- package/docs/superpowers/specs/2026-07-19-relatable-readme-hero-design.md +0 -80
- package/docs/superpowers/specs/2026-07-20-loadout-readme-explainer-design.md +0 -55
- package/docs/superpowers/specs/2026-07-20-project-activation-safety-design.md +0 -228
|
@@ -98,7 +98,6 @@ Confirm isolation before any applied command:
|
|
|
98
98
|
|
|
99
99
|
```bash
|
|
100
100
|
loadout status --json
|
|
101
|
-
loadout capabilities
|
|
102
101
|
```
|
|
103
102
|
|
|
104
103
|
Expected: paths, if shown, are below `TEST_ROOT`; Codex and Claude Code are detected;
|
|
@@ -158,8 +157,6 @@ loadout library --json
|
|
|
158
157
|
loadout health --explain --json
|
|
159
158
|
loadout report --json
|
|
160
159
|
loadout card --json
|
|
161
|
-
loadout outcomes --json
|
|
162
|
-
loadout capabilities --inspect --json
|
|
163
160
|
```
|
|
164
161
|
|
|
165
162
|
Expected:
|
|
@@ -184,7 +181,6 @@ Network variants must be run deliberately:
|
|
|
184
181
|
loadout catalog --refresh --json
|
|
185
182
|
loadout health --updates --json
|
|
186
183
|
loadout scan --agents codex,claude-code --refresh-provenance --json
|
|
187
|
-
loadout compare brainstorming --offline --json
|
|
188
184
|
loadout discover --source mcp-registry --limit 10 --json
|
|
189
185
|
loadout discover --source skills-sh --limit 10 --json
|
|
190
186
|
```
|
|
@@ -201,79 +197,7 @@ skills.sh path needs its request-scoped `VERCEL_OIDC_TOKEN`; without one it must
|
|
|
201
197
|
previous complete cache or return an attributed `unavailable` result without making
|
|
202
198
|
an unauthenticated request. Neither source installs or promotes a lead.
|
|
203
199
|
|
|
204
|
-
## 3.
|
|
205
|
-
|
|
206
|
-
Create a package _inside_ the test project so its manifest can be exported portably:
|
|
207
|
-
|
|
208
|
-
```bash
|
|
209
|
-
mkdir -p "$TEST_PROJECT/packages"
|
|
210
|
-
loadout create "$TEST_PROJECT/packages/matrix-demo" \
|
|
211
|
-
--name matrix-demo --description "Disposable matrix package"
|
|
212
|
-
loadout pack "$TEST_PROJECT/packages/matrix-demo" --json
|
|
213
|
-
loadout publish "$TEST_PROJECT/packages/matrix-demo" --local
|
|
214
|
-
loadout search matrix-demo --json
|
|
215
|
-
|
|
216
|
-
loadout init --path "$TEST_PROJECT/loadout.json" --name matrix \
|
|
217
|
-
--agents codex,claude-code --scope project
|
|
218
|
-
(
|
|
219
|
-
cd "$TEST_PROJECT"
|
|
220
|
-
loadout add local-demo --manifest loadout.json --local \
|
|
221
|
-
--path packages/matrix-demo --agents codex,claude-code
|
|
222
|
-
loadout sync --manifest loadout.json --lock loadout.lock
|
|
223
|
-
loadout sync --manifest loadout.json --lock loadout.lock --yes
|
|
224
|
-
loadout audit --manifest loadout.json --lock loadout.lock --json
|
|
225
|
-
)
|
|
226
|
-
```
|
|
227
|
-
|
|
228
|
-
Expected: `pack` returns a deterministic digest; local publication is immutable;
|
|
229
|
-
the first `sync` is a dry run; applied sync prints one snapshot; `audit` returns
|
|
230
|
-
`"valid": true`; skills appear only below the disposable profile.
|
|
231
|
-
|
|
232
|
-
Exercise desired-state editing and portability:
|
|
233
|
-
|
|
234
|
-
```bash
|
|
235
|
-
(
|
|
236
|
-
cd "$TEST_PROJECT"
|
|
237
|
-
loadout lock --manifest loadout.json --output loadout.lock
|
|
238
|
-
loadout export portable.json --manifest loadout.json --lock loadout.lock
|
|
239
|
-
loadout import portable.json --manifest imported.json --lock imported.lock
|
|
240
|
-
loadout import portable.json --manifest imported.json --lock imported.lock --yes
|
|
241
|
-
loadout unadd local-demo --manifest imported.json
|
|
242
|
-
)
|
|
243
|
-
```
|
|
244
|
-
|
|
245
|
-
Expected: import previews before writing; applied import snapshots destinations;
|
|
246
|
-
`unadd` changes desired state only and does not delete installed files.
|
|
247
|
-
|
|
248
|
-
The full authenticated remote-registry protocol, including wrong-token rejection,
|
|
249
|
-
immutable-version conflict rejection, exact digest download, risk approval, and the
|
|
250
|
-
HTTPS/non-loopback boundary, is reproducibly covered by:
|
|
251
|
-
|
|
252
|
-
```bash
|
|
253
|
-
npx vitest run tests/registry-api.test.ts tests/package.test.ts
|
|
254
|
-
```
|
|
255
|
-
|
|
256
|
-
Manual `registry-serve`/remote `publish` is an X test. After completing the throwaway
|
|
257
|
-
credential setup in track 9, start this in one terminal:
|
|
258
|
-
|
|
259
|
-
```bash
|
|
260
|
-
loadout registry-serve --port 7331 \
|
|
261
|
-
--credential-keychain loadout-registry-test --credential-account tester
|
|
262
|
-
```
|
|
263
|
-
|
|
264
|
-
Then publish from another terminal:
|
|
265
|
-
|
|
266
|
-
```bash
|
|
267
|
-
loadout publish "$TEST_PROJECT/packages/matrix-demo" \
|
|
268
|
-
--registry-url http://127.0.0.1:7331 \
|
|
269
|
-
--credential-keychain loadout-registry-test --credential-account tester
|
|
270
|
-
```
|
|
271
|
-
|
|
272
|
-
Never place the token on the command line. An identical version/content publish may
|
|
273
|
-
be accepted idempotently; changed content at the same version must be rejected. Stop
|
|
274
|
-
the server with Ctrl-C.
|
|
275
|
-
|
|
276
|
-
## 4. Install, active-set, outcome, and rollback track (A; Maximum is N)
|
|
200
|
+
## 3. Install, active-set, and rollback track (A; Maximum is N)
|
|
277
201
|
|
|
278
202
|
Exercise the new-user golden path before its constituent commands:
|
|
279
203
|
|
|
@@ -321,9 +245,6 @@ loadout disable direct-demo --agents codex
|
|
|
321
245
|
loadout disable direct-demo --agents codex --yes --json
|
|
322
246
|
loadout enable direct-demo --agents codex
|
|
323
247
|
loadout enable direct-demo --agents codex --yes --json
|
|
324
|
-
loadout outcome direct-demo/direct-demo --agent codex --task testing \
|
|
325
|
-
--result success
|
|
326
|
-
loadout outcomes --json
|
|
327
248
|
loadout share "$TEST_PROJECT/share.json"
|
|
328
249
|
loadout remove direct-demo
|
|
329
250
|
loadout remove direct-demo --yes
|
|
@@ -355,7 +276,7 @@ and preserves matching active Stable units at the same reviewed commit; MCP-only
|
|
|
355
276
|
packages remain explicit setup items. `--approve-risk` acknowledges displayed static
|
|
356
277
|
findings but does not execute third-party repository scripts.
|
|
357
278
|
|
|
358
|
-
##
|
|
279
|
+
## 4. Existing-skill provenance, adoption, and freshness (R/S/A)
|
|
359
280
|
|
|
360
281
|
Copy one harmless skill into the disposable unmanaged profile, then inspect it:
|
|
361
282
|
|
|
@@ -366,7 +287,6 @@ cp "$TEST_PROJECT/packages/matrix-demo/skills/matrix-demo/SKILL.md" \
|
|
|
366
287
|
loadout scan --agents codex --json
|
|
367
288
|
loadout adopt unmanaged-demo --agent codex --json
|
|
368
289
|
loadout adopt unmanaged-demo --agent codex --yes --json
|
|
369
|
-
loadout compare unmanaged-demo --agent codex --offline --json
|
|
370
290
|
```
|
|
371
291
|
|
|
372
292
|
Expected: scan labels the copy unmanaged; adoption preview does not change its bytes;
|
|
@@ -401,13 +321,12 @@ Only run the second command when the first reports a real reviewed update. Expec
|
|
|
401
321
|
exact diff/safety plan, a snapshot on success, and refusal when new risky findings are
|
|
402
322
|
not acknowledged with `--approve-risk`.
|
|
403
323
|
|
|
404
|
-
##
|
|
324
|
+
## 5. Discovery and human review queue (N/S)
|
|
405
325
|
|
|
406
326
|
```bash
|
|
407
327
|
loadout discover --source hacker-news --limit 20 --min-score 20 --json
|
|
408
328
|
loadout discover --source github --limit 20 --queue --json
|
|
409
329
|
loadout discover --source all --limit 20 --queue --json
|
|
410
|
-
loadout review-queue --decision pending --json
|
|
411
330
|
```
|
|
412
331
|
|
|
413
332
|
Expected: results contain source evidence and public repository identifiers; queueing
|
|
@@ -415,12 +334,6 @@ deduplicates leads; nothing is promoted, cloned into an agent, or installed.
|
|
|
415
334
|
|
|
416
335
|
For a repository printed by the queue:
|
|
417
336
|
|
|
418
|
-
```bash
|
|
419
|
-
loadout review owner/repository --decision shortlisted
|
|
420
|
-
loadout review-queue --decision shortlisted --json
|
|
421
|
-
loadout review owner/repository --decision ignored
|
|
422
|
-
```
|
|
423
|
-
|
|
424
337
|
Private GitHub discovery is opt-in and reads `GITHUB_TOKEN` only when `--private` is
|
|
425
338
|
present. Prefer a native credential reference:
|
|
426
339
|
|
|
@@ -431,7 +344,7 @@ loadout discover --source github --private \
|
|
|
431
344
|
|
|
432
345
|
Use a low-scope test token. The output and state must never contain its value.
|
|
433
346
|
|
|
434
|
-
##
|
|
347
|
+
## 6. Static inspection, MCP, and conversion (R/S)
|
|
435
348
|
|
|
436
349
|
Static package analysis never executes package content:
|
|
437
350
|
|
|
@@ -439,8 +352,6 @@ Static package analysis never executes package content:
|
|
|
439
352
|
loadout inspect --source "$TEST_PROJECT/packages/matrix-demo" --json
|
|
440
353
|
loadout evaluate --source "$TEST_PROJECT/packages/matrix-demo" --json
|
|
441
354
|
loadout mcp --source "$TEST_PROJECT/packages/matrix-demo" --json
|
|
442
|
-
loadout canary --source "$TEST_PROJECT/packages/matrix-demo" \
|
|
443
|
-
--package matrix-demo --json
|
|
444
355
|
```
|
|
445
356
|
|
|
446
357
|
Repeat `inspect`, `evaluate`, or `mcp` with `--repository owner/repository` for the
|
|
@@ -469,10 +380,6 @@ loadout mcp-config --config "$TEST_PROJECT/mcp.json" --name local-example \
|
|
|
469
380
|
--command node --arg server.js
|
|
470
381
|
loadout mcp-config --config "$TEST_PROJECT/mcp.json" --name local-example \
|
|
471
382
|
--command node --arg server.js --yes
|
|
472
|
-
loadout codex-mcp-config --config "$TEST_PROJECT/config.toml" \
|
|
473
|
-
--name remote-example --url https://example.com/mcp
|
|
474
|
-
loadout codex-mcp-config --config "$TEST_PROJECT/config.toml" \
|
|
475
|
-
--name remote-example --url https://example.com/mcp --yes
|
|
476
383
|
loadout mcp-recipe --json
|
|
477
384
|
```
|
|
478
385
|
|
|
@@ -486,148 +393,11 @@ requirements you have reviewed.
|
|
|
486
393
|
|
|
487
394
|
Docker sandbox execution is intentionally separate:
|
|
488
395
|
|
|
489
|
-
```bash
|
|
490
|
-
loadout sandbox-run --source "$TEST_PROJECT/packages/matrix-demo" \
|
|
491
|
-
--image '<reviewed-image>@sha256:<digest>' \
|
|
492
|
-
--command node --command --version --json
|
|
493
|
-
loadout sandbox-run --source "$TEST_PROJECT/packages/matrix-demo" \
|
|
494
|
-
--image '<reviewed-image>@sha256:<digest>' \
|
|
495
|
-
--command node --command --version \
|
|
496
|
-
--approve-risk --timeout 30000 --json
|
|
497
|
-
```
|
|
498
|
-
|
|
499
396
|
Expected: the first invocation refuses/only plans without approval; the approved
|
|
500
397
|
container has a read-only source mount, no inherited secrets, no Docker socket, no
|
|
501
398
|
network, and a time bound. The image may need to be pulled beforehand.
|
|
502
399
|
|
|
503
|
-
##
|
|
504
|
-
|
|
505
|
-
Validate the model-free benchmark campaign and card/compare surfaces with their
|
|
506
|
-
deterministic automated contracts:
|
|
507
|
-
|
|
508
|
-
```bash
|
|
509
|
-
npx vitest run tests/benchmark-campaign.test.ts tests/benchmark-cli.test.ts \
|
|
510
|
-
tests/loadout-card.test.ts tests/share-report.test.ts
|
|
511
|
-
```
|
|
512
|
-
|
|
513
|
-
Expected: campaign hashes and paired order are deterministic, every retry is included
|
|
514
|
-
in the worst-case budget, over-budget plans are blocked, resumable metadata contains
|
|
515
|
-
no prompt/output/credential bytes, and aggregate comparison never invents a quality
|
|
516
|
-
delta. See `docs/EVALUATION_PROTOCOL_V1.md` for the campaign JSON contract. These
|
|
517
|
-
tests do not call a model provider and consume no provider credit.
|
|
518
|
-
|
|
519
|
-
```bash
|
|
520
|
-
loadout keygen --private-key "$TEST_ROOT/private.pem" \
|
|
521
|
-
--public-key "$TEST_ROOT/public.pem"
|
|
522
|
-
loadout catalog-sign --catalog "$LOADOUT_ROOT/catalog/packages.json" \
|
|
523
|
-
--private-key "$TEST_ROOT/private.pem" --output "$TEST_ROOT/catalog.signed.json"
|
|
524
|
-
loadout catalog-verify --snapshot "$TEST_ROOT/catalog.signed.json" \
|
|
525
|
-
--public-key "$TEST_ROOT/public.pem"
|
|
526
|
-
```
|
|
527
|
-
|
|
528
|
-
Expected: the private key is owner-only and outside the repository; verification
|
|
529
|
-
succeeds; changing any byte in the signed payload makes verification fail.
|
|
530
|
-
|
|
531
|
-
Preview and apply the same signed catalog inside the disposable profile:
|
|
532
|
-
|
|
533
|
-
```bash
|
|
534
|
-
loadout catalog-update --source "$TEST_ROOT/catalog.signed.json" \
|
|
535
|
-
--public-key "$TEST_ROOT/public.pem"
|
|
536
|
-
loadout catalog-update --source "$TEST_ROOT/catalog.signed.json" \
|
|
537
|
-
--public-key "$TEST_ROOT/public.pem" --yes
|
|
538
|
-
loadout catalog --coverage --json
|
|
539
|
-
```
|
|
540
|
-
|
|
541
|
-
Expected: preview prints an exact signed diff without mutation; apply creates a
|
|
542
|
-
snapshot and trusted state; the effective catalog re-verifies the stored envelope.
|
|
543
|
-
Repeating `--yes` refuses a replay. Test removal only in the disposable profile and
|
|
544
|
-
only with the separate `--allow-removals` acknowledgement.
|
|
545
|
-
|
|
546
|
-
The repository's generated feed can be triaged without network access:
|
|
547
|
-
|
|
548
|
-
```bash
|
|
549
|
-
loadout candidate list --limit 5 --json
|
|
550
|
-
loadout candidate list --query "codex skills"
|
|
551
|
-
loadout capabilities --gaps --json
|
|
552
|
-
loadout recommend --project "$TEST_PROJECT" --agent codex --json
|
|
553
|
-
```
|
|
554
|
-
|
|
555
|
-
`candidate inspect owner/repository --output ./candidate-dossier.json` is a networked
|
|
556
|
-
test: it performs a real public Git clone and writes a static immutable dossier to
|
|
557
|
-
disposable Loadout state. Review that output before exercising `candidate propose`;
|
|
558
|
-
proposal preview and approved proposal output never mutate the catalog.
|
|
559
|
-
|
|
560
|
-
Graphify is an explicit executable recipe rather than a broad-setup component. With
|
|
561
|
-
`uv` installed, exercise it only inside the disposable profile:
|
|
562
|
-
|
|
563
|
-
```bash
|
|
564
|
-
loadout tool
|
|
565
|
-
loadout tool graphify --agents codex
|
|
566
|
-
loadout tool graphify --agents codex --yes --approve-risk
|
|
567
|
-
"$LOADOUT_HOME/runtime/graphify/bin/graphify" --version
|
|
568
|
-
test -f "$LOADOUT_USER_HOME/.codex/skills/graphify/SKILL.md"
|
|
569
|
-
loadout tool graphify --remove
|
|
570
|
-
loadout tool graphify --remove --yes --approve-risk
|
|
571
|
-
test ! -e "$LOADOUT_USER_HOME/.codex/skills/graphify"
|
|
572
|
-
test ! -e "$LOADOUT_HOME/runtime/graphify"
|
|
573
|
-
```
|
|
574
|
-
|
|
575
|
-
Expected: preview identifies the exact wheel hash and all commands; apply reports
|
|
576
|
-
Graphify 0.9.17, writes only the disposable target and isolated runtime, and removal
|
|
577
|
-
restores the original target. The installer subprocess must not inherit API keys.
|
|
578
|
-
|
|
579
|
-
Create a deterministic workflow fixture and five declared trials per candidate. This
|
|
580
|
-
is harness input, not model-generated evidence, and it executes no candidate content:
|
|
581
|
-
|
|
582
|
-
```bash
|
|
583
|
-
node --input-type=module <<'NODE'
|
|
584
|
-
import { writeFile } from "node:fs/promises";
|
|
585
|
-
const root = process.env.TEST_PROJECT;
|
|
586
|
-
const fixture = {
|
|
587
|
-
id: "matrix-workflow",
|
|
588
|
-
version: "1",
|
|
589
|
-
category: "workflow-adherence",
|
|
590
|
-
requiredActions: ["inspect", "edit", "verify"],
|
|
591
|
-
forbiddenActions: ["delete-unrelated"]
|
|
592
|
-
};
|
|
593
|
-
const trials = Array.from({ length: 5 }, () => [
|
|
594
|
-
{
|
|
595
|
-
candidateId: "baseline",
|
|
596
|
-
fixtureId: fixture.id,
|
|
597
|
-
observations: ["inspect", "edit", "verify"],
|
|
598
|
-
durationMs: 10
|
|
599
|
-
},
|
|
600
|
-
{
|
|
601
|
-
candidateId: "improved",
|
|
602
|
-
fixtureId: fixture.id,
|
|
603
|
-
observations: ["inspect", "edit", "verify", "report-uncertainty"],
|
|
604
|
-
durationMs: 10
|
|
605
|
-
}
|
|
606
|
-
]).flat();
|
|
607
|
-
await writeFile(`${root}/fixture.json`, JSON.stringify(fixture, null, 2));
|
|
608
|
-
await writeFile(`${root}/trials.json`, JSON.stringify(trials, null, 2));
|
|
609
|
-
NODE
|
|
610
|
-
```
|
|
611
|
-
|
|
612
|
-
Sign and inspect the resulting evidence:
|
|
613
|
-
|
|
614
|
-
```bash
|
|
615
|
-
loadout head-to-head --fixture "$TEST_PROJECT/fixture.json" \
|
|
616
|
-
--trials "$TEST_PROJECT/trials.json" --private-key "$TEST_ROOT/private.pem" \
|
|
617
|
-
--output "$TEST_ROOT/evidence.json" --json
|
|
618
|
-
loadout alerts --evidence "$TEST_ROOT/evidence.json" \
|
|
619
|
-
--public-key "$TEST_ROOT/public.pem" --json
|
|
620
|
-
```
|
|
621
|
-
|
|
622
|
-
The harness scores declared observations only. It never executes candidate content.
|
|
623
|
-
The authoritative schema, safety-failure, minimum-trial, tamper, and practical-delta
|
|
624
|
-
tests are also directly runnable:
|
|
625
|
-
|
|
626
|
-
```bash
|
|
627
|
-
npx vitest run tests/head-to-head.test.ts tests/signing.test.ts
|
|
628
|
-
```
|
|
629
|
-
|
|
630
|
-
## 9. Credentials and model-provider verification (H/$)
|
|
400
|
+
## 7. Credentials and model-provider verification (H/$)
|
|
631
401
|
|
|
632
402
|
This track touches the real operating-system credential store even when
|
|
633
403
|
`LOADOUT_HOME` is disposable. Use a unique throwaway service name and delete it.
|
|
@@ -668,7 +438,7 @@ unset OPENROUTER_API_KEY
|
|
|
668
438
|
`models verify` makes one minimal request and may consume provider credit ($). Inspect
|
|
669
439
|
provider billing before and after; do not run it in a loop.
|
|
670
440
|
|
|
671
|
-
##
|
|
441
|
+
## 8. Watchers, native scheduling, completions, and loopback UI/API (X/H)
|
|
672
442
|
|
|
673
443
|
One-shot update watching is safe and networked:
|
|
674
444
|
|
|
@@ -707,33 +477,10 @@ and `unschedule` commands remain available for job-specific control.
|
|
|
707
477
|
|
|
708
478
|
The optional read-only loopback API can be checked separately:
|
|
709
479
|
|
|
710
|
-
```bash
|
|
711
|
-
loadout serve --port 0
|
|
712
|
-
```
|
|
713
|
-
|
|
714
480
|
Confirm that it binds only to `127.0.0.1`, inspect the API response, and stop it with
|
|
715
481
|
Ctrl-C. It must not bind a public interface.
|
|
716
482
|
|
|
717
|
-
##
|
|
718
|
-
|
|
719
|
-
```bash
|
|
720
|
-
loadout improve --json
|
|
721
|
-
loadout improve --write --output "$TEST_PROJECT/improvements" --json
|
|
722
|
-
```
|
|
723
|
-
|
|
724
|
-
Copy the exact cycle id printed by the second command:
|
|
725
|
-
|
|
726
|
-
```bash
|
|
727
|
-
loadout improve-feedback --id <cycle-id> --outcome partial \
|
|
728
|
-
--note "Disposable matrix verification" \
|
|
729
|
-
--directory "$TEST_PROJECT/improvements"
|
|
730
|
-
```
|
|
731
|
-
|
|
732
|
-
Expected: the first command is read-only; `--write` persists a local prompt/cycle
|
|
733
|
-
record only; feedback requires a human-selected outcome and stores no project source
|
|
734
|
-
or prompt transcript.
|
|
735
|
-
|
|
736
|
-
## 12. Cleanup and pass criteria
|
|
483
|
+
## 9. Cleanup and pass criteria
|
|
737
484
|
|
|
738
485
|
First verify snapshot availability and roll back any remaining disposable mutation:
|
|
739
486
|
|
|
@@ -19,6 +19,11 @@ paste a token into a manifest, catalog, log, or command argument.
|
|
|
19
19
|
|
|
20
20
|
## Local flow and failure modes
|
|
21
21
|
|
|
22
|
+
> **Partly unimplemented.** `loadout connect` and `loadout disconnect` do not
|
|
23
|
+
> exist in the CLI today. `loadout discover --private` is real and reads
|
|
24
|
+
> credentials the user has already configured. Steps 1, 2, 3, and 5 describe the
|
|
25
|
+
> intended flow, not current behaviour.
|
|
26
|
+
|
|
22
27
|
1. `loadout connect github` opens a browser using PKCE and a loopback callback.
|
|
23
28
|
2. The local process verifies `state`, PKCE verifier, expiration, and callback host.
|
|
24
29
|
3. It stores only an OS-keychain reference to the refresh/session material.
|
|
@@ -26,7 +26,7 @@ their `SKILL.md` files, and writes a local index below `LOADOUT_HOME/provenance`
|
|
|
26
26
|
|
|
27
27
|
## Relationship classification
|
|
28
28
|
|
|
29
|
-
`loadout
|
|
29
|
+
`loadout scan` uses deterministic relationships:
|
|
30
30
|
|
|
31
31
|
- `exact-copy`: identical instruction fingerprint;
|
|
32
32
|
- `divergent-same-name`: same normalized name, different instructions;
|
package/docs/RELEASE_REVIEW.md
CHANGED
|
@@ -19,7 +19,6 @@ For current behavior and evidence, use:
|
|
|
19
19
|
catalog/support facts;
|
|
20
20
|
- the [changelog](../CHANGELOG.md) for released behavior;
|
|
21
21
|
- the [feature test matrix](./FEATURE_TEST_MATRIX.md) for adapter evidence;
|
|
22
|
-
- the [repository stabilization record](./REPOSITORY_STABILIZATION.md) and
|
|
23
22
|
[sanitized July 19 live checks](./evidence/live-checks-2026-07-19.json) only when
|
|
24
23
|
investigating that historical release line.
|
|
25
24
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "loadout-ai",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.8.0",
|
|
4
4
|
"private": false,
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"description": "Universal upgrade manager for AI coding agents",
|
|
@@ -15,7 +15,6 @@
|
|
|
15
15
|
"docs",
|
|
16
16
|
"README.md",
|
|
17
17
|
"CHANGELOG.md",
|
|
18
|
-
"MASTER_PLAN.md",
|
|
19
18
|
"SECURITY.md",
|
|
20
19
|
"LICENSE"
|
|
21
20
|
],
|
|
@@ -57,6 +56,7 @@
|
|
|
57
56
|
"test:e2e:readme": "node scripts/readme-product-flow.mjs",
|
|
58
57
|
"test:package": "node scripts/package-smoke.mjs",
|
|
59
58
|
"pretest:performance": "npm run build",
|
|
59
|
+
"test:coverage": "vitest run --coverage",
|
|
60
60
|
"test:performance": "node scripts/scan-benchmark.mjs",
|
|
61
61
|
"typecheck": "tsc -p tsconfig.json --noEmit",
|
|
62
62
|
"readme:update": "node scripts/update-readme-facts.mjs",
|
|
@@ -65,8 +65,9 @@
|
|
|
65
65
|
"check:live": "node scripts/check-live-evidence.mjs",
|
|
66
66
|
"check:allowlists": "npm run build && node scripts/check-mode-allowlists.mjs",
|
|
67
67
|
"check:catalog-freshness": "npm run build && node scripts/check-catalog-freshness.mjs",
|
|
68
|
-
"check:
|
|
69
|
-
"
|
|
68
|
+
"check:audit": "npm audit --audit-level=high",
|
|
69
|
+
"check:evidence": "node scripts/check-catalog-attribution.mjs && node scripts/check-discovery-artifacts.mjs && node scripts/check-documented-commands.mjs && npm run check:readme-claims",
|
|
70
|
+
"verify": "npm run format:check && npm run lint && npm run typecheck && npm run check:audit && npm run check:evidence && npm test -- --run && npm run test:e2e:cli && npm run test:e2e:readme && npm run test:package && npm run test:performance",
|
|
70
71
|
"verify:full": "npm run verify"
|
|
71
72
|
},
|
|
72
73
|
"dependencies": {
|
|
@@ -76,6 +77,7 @@
|
|
|
76
77
|
"devDependencies": {
|
|
77
78
|
"@eslint/js": "^9.39.5",
|
|
78
79
|
"@types/node": "^22.10.0",
|
|
80
|
+
"@vitest/coverage-v8": "^4.1.10",
|
|
79
81
|
"eslint": "^9.39.5",
|
|
80
82
|
"eslint-config-prettier": "^10.1.8",
|
|
81
83
|
"prettier": "^3.9.5",
|
|
@@ -1,120 +1,85 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: loadout-router
|
|
3
|
-
description:
|
|
3
|
+
description: Decide which model to use for a coding task and hand work to another agent. Use when the user asks which model to use, mentions running low on usage or cost, asks whether to switch to Opus or Sonnet or a GPT tier, or wants to delegate a task to Codex or Claude Code.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Loadout Router
|
|
7
7
|
|
|
8
|
-
|
|
9
|
-
|
|
8
|
+
Pick the model that fits the work, using the user's own routing policy rather
|
|
9
|
+
than your guess or mine.
|
|
10
10
|
|
|
11
|
-
|
|
12
|
-
current with the installed version rather than going stale in this file.
|
|
13
|
-
|
|
14
|
-
## Prerequisite
|
|
15
|
-
|
|
16
|
-
Check once per session:
|
|
17
|
-
|
|
18
|
-
```bash
|
|
19
|
-
loadout --version
|
|
20
|
-
```
|
|
21
|
-
|
|
22
|
-
If that fails, tell the user to install it (`npm install --global loadout-ai`)
|
|
23
|
-
and answer from general knowledge instead of guessing at specifics.
|
|
24
|
-
|
|
25
|
-
## Choosing a model
|
|
26
|
-
|
|
27
|
-
Run the router with the task described in plain words:
|
|
11
|
+
## The policy is the user's, not yours
|
|
28
12
|
|
|
29
13
|
```bash
|
|
30
|
-
loadout route
|
|
14
|
+
loadout route
|
|
31
15
|
```
|
|
32
16
|
|
|
33
|
-
|
|
34
|
-
debug, document — and prints the recommended tier, the models in that tier with
|
|
35
|
-
current prices, which of the user's installed agents can run them, and a cheaper
|
|
36
|
-
fallback with its tradeoff.
|
|
17
|
+
That prints three buckets and the model the user has chosen for each:
|
|
37
18
|
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
19
|
+
- **hard** — architecture, security, migrations, tricky debugging, risky review
|
|
20
|
+
- **normal** — most implementation, ordinary debugging, refactors
|
|
21
|
+
- **cheap** — tests, docs, boilerplate, renames, mechanical edits
|
|
41
22
|
|
|
42
|
-
|
|
23
|
+
Read the policy before advising. If the user disagrees with a recommendation,
|
|
24
|
+
the fix is to change the policy, not to argue:
|
|
43
25
|
|
|
44
26
|
```bash
|
|
45
|
-
loadout route --
|
|
27
|
+
loadout route --set cheap=claude-sonnet-5
|
|
46
28
|
```
|
|
47
29
|
|
|
48
|
-
##
|
|
30
|
+
## Your job is the bucket, not the model
|
|
49
31
|
|
|
50
|
-
|
|
51
|
-
|
|
32
|
+
The CLI can guess a bucket from wording, and it says so when it does. **You
|
|
33
|
+
should do better**, because you have the conversation, the code, and the stakes.
|
|
34
|
+
Decide the bucket yourself and state it:
|
|
52
35
|
|
|
53
36
|
```bash
|
|
54
|
-
loadout route --
|
|
37
|
+
loadout route --bucket hard
|
|
55
38
|
```
|
|
56
39
|
|
|
57
|
-
|
|
58
|
-
what is being given up — "shallower architectural reasoning, so review the plan
|
|
59
|
-
more carefully" — rather than presenting it as a free win.
|
|
40
|
+
Judge by consequence, not vocabulary:
|
|
60
41
|
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
42
|
+
- Anything touching auth, payments, migrations, or data deletion is **hard**,
|
|
43
|
+
however small the diff.
|
|
44
|
+
- Unfamiliar code is harder than familiar code doing the same thing.
|
|
45
|
+
- A one-line change in a hot path is not cheap.
|
|
46
|
+
- Genuinely mechanical work — a rename, a docstring, a test for code you just
|
|
47
|
+
wrote — is **cheap**, and paying frontier prices for it is waste.
|
|
66
48
|
|
|
67
|
-
|
|
68
|
-
loadout route --models
|
|
69
|
-
loadout route --models --provider anthropic
|
|
70
|
-
loadout route --models --tier fast
|
|
71
|
-
loadout route --cost
|
|
72
|
-
```
|
|
49
|
+
Report the model and why that bucket. One or two sentences.
|
|
73
50
|
|
|
74
|
-
|
|
75
|
-
are per-million-token list rates; actual spend depends on prompt size, so give
|
|
76
|
-
ratios ("roughly 4x cheaper") rather than predicting a dollar total.
|
|
51
|
+
## When cost matters
|
|
77
52
|
|
|
78
|
-
|
|
53
|
+
If the user mentions running low, being rate limited, or wanting to spend less,
|
|
54
|
+
say what a cheaper bucket would cost them in quality rather than presenting it
|
|
55
|
+
as free. `loadout route` shows real per-million prices for the comparison.
|
|
79
56
|
|
|
80
|
-
|
|
81
|
-
|
|
57
|
+
Neither Claude Code nor Codex exposes remaining quota programmatically, so never
|
|
58
|
+
claim to know how much the user has left.
|
|
82
59
|
|
|
83
|
-
|
|
84
|
-
loadout handoff status
|
|
85
|
-
```
|
|
60
|
+
## Handing work to the other agent
|
|
86
61
|
|
|
87
|
-
|
|
62
|
+
One command sends a task; it sets up the shared log on first use:
|
|
88
63
|
|
|
89
64
|
```bash
|
|
90
|
-
loadout handoff
|
|
65
|
+
loadout handoff codex "write vitest coverage for src/auth.ts" --context "zod schemas already exist"
|
|
91
66
|
```
|
|
92
67
|
|
|
93
|
-
Put
|
|
94
|
-
|
|
95
|
-
none of this conversation.
|
|
68
|
+
Put everything the receiver needs into `--context` — file paths, decisions you
|
|
69
|
+
already made, what you deliberately left out. It has none of this conversation.
|
|
96
70
|
|
|
97
|
-
Sending
|
|
98
|
-
|
|
71
|
+
Sending writes to a shared file in the user's repository, so confirm first
|
|
72
|
+
unless they asked for the handoff themselves.
|
|
99
73
|
|
|
100
74
|
## Reading your own inbox
|
|
101
75
|
|
|
102
|
-
At
|
|
103
|
-
agent left you work:
|
|
76
|
+
At session start, and after finishing a task:
|
|
104
77
|
|
|
105
78
|
```bash
|
|
106
|
-
loadout handoff
|
|
79
|
+
loadout handoff claude-code
|
|
107
80
|
```
|
|
108
81
|
|
|
109
|
-
|
|
110
|
-
command it prints
|
|
111
|
-
not mention it.
|
|
112
|
-
|
|
113
|
-
## Judgment this skill does not replace
|
|
82
|
+
Work anything listed in order, then run the `loadout handoff --done <id>`
|
|
83
|
+
command it prints. If nothing is pending, say nothing and carry on.
|
|
114
84
|
|
|
115
|
-
|
|
116
|
-
model. The classifier reads keywords, not stakes.
|
|
117
|
-
- Security-sensitive, migration, and data-loss paths are worth frontier tier
|
|
118
|
-
regardless of what phase they classify as.
|
|
119
|
-
- If the user has already chosen a model, do not argue unless the choice is
|
|
120
|
-
clearly wrong for the work.
|
|
85
|
+
`loadout handoff` with no arguments shows every pending task, both directions.
|