loadout-ai 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/CHANGELOG.md +65 -0
  2. package/MASTER_PLAN.md +173 -9
  3. package/README.md +57 -19
  4. package/catalog/discovered.json +8422 -6957
  5. package/catalog/packages.json +26 -0
  6. package/dist/src/cli.js +140 -68
  7. package/dist/src/core/active-policy.js +216 -42
  8. package/dist/src/core/active-set.js +14 -5
  9. package/dist/src/core/adopt.js +1 -1
  10. package/dist/src/core/catalog-coverage.js +2 -1
  11. package/dist/src/core/catalog-install.js +23 -8
  12. package/dist/src/core/cli-guide.js +5 -2
  13. package/dist/src/core/codex-mcp.js +49 -5
  14. package/dist/src/core/completion.js +0 -3
  15. package/dist/src/core/health.js +19 -11
  16. package/dist/src/core/install.js +6 -39
  17. package/dist/src/core/mcp-recipes.js +13 -4
  18. package/dist/src/core/mcp.js +6 -5
  19. package/dist/src/core/model-config.js +1 -1
  20. package/dist/src/core/portable.js +1 -1
  21. package/dist/src/core/recommend.js +129 -27
  22. package/dist/src/core/remove.js +28 -2
  23. package/dist/src/core/runtime-tool-recipe.js +2 -2
  24. package/dist/src/core/runtime-tools.js +15 -4
  25. package/dist/src/core/snapshot.js +20 -0
  26. package/dist/src/core/state.js +9 -0
  27. package/dist/src/core/sync.js +1 -1
  28. package/dist/src/core/target-occupancy.js +50 -0
  29. package/dist/src/core/transaction.js +2 -1
  30. package/dist/src/shared/schemas.js +1 -0
  31. package/docs/CATALOG.md +3 -2
  32. package/docs/DISCOVERED.md +253 -255
  33. package/docs/FEATURE_TEST_MATRIX.md +19 -32
  34. package/docs/GITHUB_AUTHORIZATION.md +2 -2
  35. package/docs/TESTING.md +3 -16
  36. package/docs/USER_TEST_GUIDE.md +39 -13
  37. package/docs/assets/loadout-workflow.png +0 -0
  38. package/docs/evidence/readme-claims.json +0 -8
  39. package/docs/superpowers/plans/2026-07-20-loadout-readme-explainer.md +116 -0
  40. package/docs/superpowers/plans/2026-07-20-project-activation-safety.md +469 -0
  41. package/docs/superpowers/specs/2026-07-20-loadout-readme-explainer-design.md +55 -0
  42. package/docs/superpowers/specs/2026-07-20-project-activation-safety-design.md +228 -0
  43. package/package.json +4 -6
  44. package/dashboard/app.js +0 -607
  45. package/dashboard/index.html +0 -249
  46. package/dashboard/styles.css +0 -384
  47. package/dist/src/core/demo.js +0 -136
  48. package/dist/src/dashboard.js +0 -421
@@ -112,26 +112,20 @@ Run the entire required gate with one command:
112
112
  npm run verify
113
113
  ```
114
114
 
115
- Use `npm run verify:full` when Playwright Chromium is installed and you also want the
116
- optional dashboard browser check. The individual stages are listed below for focused
117
- reruns and diagnosis.
118
-
119
- | Command | Coverage | Expected result |
120
- | ---------------------------- | ----------------------------------------------------------------------- | -------------------------------------------------- |
121
- | `npm run format:check` | Repository formatting | Exit 0; no files changed. |
122
- | `npm run lint` | TypeScript lint rules | Exit 0. |
123
- | `npm run typecheck` | TypeScript contract | Exit 0. |
124
- | `npm run check:evidence` | Catalog/discovery attribution, README claims, and release boundaries | Exit 0; no claim is silently promoted. |
125
- | `npm test` | Unit, integration, native filesystem, safety, and regression suites | All tests pass. |
126
- | `npm run test:e2e:cli` | Disposable scan → compare → optimize → apply → rollback journey | Prints a successful CLI product flow. |
127
- | `npm run test:e2e:readme` | Isolated library/activation/manifest/card/rollback journey | Prints README product flow success. |
128
- | `npm run test:package` | `npm pack`, install outside the checkout, packaged CLI install/rollback | Prints package smoke success. |
129
- | `npm run test:performance` | Seven scans of 1,000 real on-disk skill directories | p95 remains below the enforced five-second budget. |
130
- | `npm run test:e2e:dashboard` | Loopback dashboard first-run browser test | Playwright passes; no real profile is used. |
131
-
132
- The dashboard test needs a Playwright browser. If the browser executable is absent,
133
- run `npx playwright install chromium` once; that download is not a Loadout product
134
- side effect.
115
+ `npm run verify:full` is an alias for the same CLI release gate. The individual stages
116
+ are listed below for focused reruns and diagnosis.
117
+
118
+ | Command | Coverage | Expected result |
119
+ | -------------------------- | ----------------------------------------------------------------------- | -------------------------------------------------- |
120
+ | `npm run format:check` | Repository formatting | Exit 0; no files changed. |
121
+ | `npm run lint` | TypeScript lint rules | Exit 0. |
122
+ | `npm run typecheck` | TypeScript contract | Exit 0. |
123
+ | `npm run check:evidence` | Catalog/discovery attribution, README claims, and release boundaries | Exit 0; no claim is silently promoted. |
124
+ | `npm test` | Unit, integration, native filesystem, safety, and regression suites | All tests pass. |
125
+ | `npm run test:e2e:cli` | Disposable scan → compare → optimize → apply → rollback journey | Prints a successful CLI product flow. |
126
+ | `npm run test:e2e:readme` | Isolated library/activation/manifest/card/rollback journey | Prints README product flow success. |
127
+ | `npm run test:package` | `npm pack`, install outside the checkout, packaged CLI install/rollback | Prints package smoke success. |
128
+ | `npm run test:performance` | Seven scans of 1,000 real on-disk skill directories | p95 remains below the enforced five-second budget. |
135
129
 
136
130
  The focused regression contract for the v0.3.x profile lifecycle is:
137
131
 
@@ -361,11 +355,6 @@ and preserves matching active Stable units at the same reviewed commit; MCP-only
361
355
  packages remain explicit setup items. `--approve-risk` acknowledges displayed static
362
356
  findings but does not execute third-party repository scripts.
363
357
 
364
- `loadout demo --json` is a shorter N smoke test. It creates its own isolated profile,
365
- performs a real public-repository install and rollback, and must report that local
366
- agent configuration was untouched. Add `--keep` only when you intend to inspect and
367
- manually remove the printed demo directory.
368
-
369
358
  ## 5. Existing-skill provenance, adoption, comparison, and freshness (R/S/A)
370
359
 
371
360
  Copy one harmless skill into the disposable unmanaged profile, then inspect it:
@@ -716,16 +705,14 @@ systemd/cron facility selected on Linux. `LOADOUT_HOME` does not make the native
716
705
  scheduler disposable, so never skip the remove command. The lower-level `schedule`
717
706
  and `unschedule` commands remain available for job-specific control.
718
707
 
719
- Loopback services are X but do not require the dashboard for normal product use:
708
+ The optional read-only loopback API can be checked separately:
720
709
 
721
710
  ```bash
722
711
  loadout serve --port 0
723
- loadout dashboard --port 0
724
712
  ```
725
713
 
726
- Run them separately, open the printed `127.0.0.1` URL, confirm status/health/catalog
727
- render from disposable state, and stop each with Ctrl-C. They must not bind a public
728
- interface.
714
+ Confirm that it binds only to `127.0.0.1`, inspect the API response, and stop it with
715
+ Ctrl-C. It must not bind a public interface.
729
716
 
730
717
  ## 11. Improvement-cycle records (S)
731
718
 
@@ -815,5 +802,5 @@ sections. Parenthesized numbers identify the track above.
815
802
  `catalog-verify`, `catalog-update` (8, 9).
816
803
  - MCP/conversion/sandbox: `mcp`, `inspect`, `evaluate`, `mcp-recipe`, `mcp-config`,
817
804
  `codex-mcp-config`, `convert`, `sandbox-run` (7).
818
- - Host/secondary surfaces: `completion`, `autopilot`, `schedule`, `unschedule`,
819
- `serve`, `dashboard` (10).
805
+ - Host/automation surfaces: `completion`, `autopilot`, `schedule`, `unschedule`,
806
+ `serve` (10).
@@ -1,7 +1,7 @@
1
1
  # Optional GitHub authorization
2
2
 
3
3
  Public discovery remains anonymous. Private discovery is opt-in and begins only
4
- when the user explicitly connects GitHub from the local CLI or loopback dashboard.
4
+ when the user explicitly connects GitHub from the local CLI.
5
5
 
6
6
  ## Authorization model
7
7
 
@@ -15,7 +15,7 @@ issues, pull-request write, or repository-write permission is requested.
15
15
  The app client ID, redirect URI, and application slug may be committed as public
16
16
  configuration. The private key and client secret cannot: they belong to a hosted
17
17
  broker or a user-controlled local credential manager. Loadout never asks the user to
18
- paste a token into a manifest, catalog, log, or dashboard URL.
18
+ paste a token into a manifest, catalog, log, or command argument.
19
19
 
20
20
  ## Local flow and failure modes
21
21
 
package/docs/TESTING.md CHANGED
@@ -27,7 +27,7 @@ snapshot back:
27
27
  npm run test:e2e:cli
28
28
  ```
29
29
 
30
- This test does not use the dashboard, network, mock command output, or any real agent
30
+ This test does not use the network, mock command output, or any real agent
31
31
  profile. It is a required CI gate on Ubuntu; the manual cross-platform workflow runs
32
32
  the broader native filesystem suite.
33
33
 
@@ -67,9 +67,8 @@ is universally safe. The opt-in `LOADOUT_TEST_LIVE_CATALOG=1` extension separate
67
67
  checks the current pinned Stable sources and remains network-dependent.
68
68
 
69
69
  Run `npm run verify` for formatting, lint, types, deterministic evidence checks, all
70
- Vitest suites, both CLI journeys, package smoke, and the performance gate. Run
71
- `npm run verify:full` only when Playwright Chromium is installed and the optional
72
- dashboard browser test is also wanted.
70
+ Vitest suites, both CLI journeys, package smoke, and the performance gate.
71
+ `npm run verify:full` is retained as an alias for the same complete CLI gate.
73
72
 
74
73
  Current npm publication, the current pinned Stable repositories, and GitHub repository
75
74
  settings are external state. Check them separately with:
@@ -281,18 +280,6 @@ output must never contain the token value; the config should contain only
281
280
  explicit `--connect --approve-risk` verification path because an arbitrary MCP host
282
281
  cannot resolve Loadout's keychain reference by itself.
283
282
 
284
- ## Optional dashboard
285
-
286
- The dashboard is a secondary inspection surface, not the onboarding requirement:
287
-
288
- ```bash
289
- npx . dashboard
290
- ```
291
-
292
- Open the printed loopback URL. CLI setup, updates, removal, discovery, and rollback all
293
- work without it. Browser automation is also optional and runs only when manually
294
- dispatched in CI; locally, use `npm run test:e2e:dashboard`.
295
-
296
283
  ## Final cleanup test
297
284
 
298
285
  ```bash
@@ -26,27 +26,29 @@ Loadout-managed skills from your own pre-existing skills. `health` checks local
26
26
  loadout catalog --json
27
27
  loadout candidate list --limit 10
28
28
  loadout recommend --project .
29
- loadout optimize --project .
29
+ loadout optimize --project . --agents codex,claude-code --limit 30
30
30
  loadout tool
31
31
  loadout tool graphify
32
32
  ```
33
33
 
34
34
  Run the project commands from the project you care about, or replace `.` with its
35
- absolute path. `optimize` is still a preview until `--yes` is supplied. `tool
36
- graphify` is also a preview; Graphify is a reviewed runtime tool and does not
37
- need an OpenAI or Anthropic API key for its code-only install.
35
+ absolute path. `recommend` labels ordinary skill libraries separately from MCP or
36
+ runtime integrations that require explicit setup. `optimize` is still a preview
37
+ until `--yes` is supplied. Its limit applies separately to every agent and includes
38
+ both Loadout-managed and pre-existing unmanaged skills, so Claude and Codex can
39
+ receive different numbers of additions. `tool graphify` is also a preview; Graphify
40
+ is a reviewed runtime tool and does not need an OpenAI or Anthropic API key for its
41
+ code-only install.
38
42
 
39
- ## 3. Open the optional dashboard
43
+ ## 3. Understand the three scopes
40
44
 
41
45
  ```bash
42
- loadout dashboard
46
+ loadout profiles
43
47
  ```
44
48
 
45
- Open the `http://127.0.0.1:PORT` address it prints. It never listens on the
46
- network. The dashboard shows status, health, installed packages, updates, local
47
- project recommendations, profiles, and the catalog. Its Apply and Undo buttons
48
- require an in-page preview, acknowledgement, and a private local session token.
49
- Stop the server with `Control-C`.
49
+ Stable keeps the active set at 30. Power deliberately activates a larger toolkit.
50
+ Maximum downloads the broadest screened skill library but keeps new entries disabled
51
+ until project optimization or an explicit enable action selects them.
50
52
 
51
53
  ## 4. Preview and install a profile
52
54
 
@@ -57,13 +59,18 @@ confirmation before changing files. A mutation creates a snapshot first.
57
59
  # Recommended everyday skills
58
60
  loadout setup --mode stable --agents codex,claude-code
59
61
 
60
- # Broader daily-use selection (50 curated skill directories)
62
+ # Broader daily-use selection (roughly 50 curated skills per agent)
61
63
  loadout setup --mode power --agents codex,claude-code
62
64
 
63
65
  # Download the broad reviewed library while keeping the active set controlled
64
66
  loadout setup --mode maximum --agents codex,claude-code
67
+ loadout setup --mode maximum --agents codex,claude-code --details
65
68
  ```
66
69
 
70
+ Maximum stores reviewed copies in Loadout's disabled library; it does not expose the
71
+ whole catalog to each agent. Follow it with `loadout optimize --project .
72
+ --agents codex,claude-code --limit 30` to preview a compact project-aware working set.
73
+
67
74
  At the API-access question, choose `None` unless you separately pay for a
68
75
  provider API. A ChatGPT Plus or Claude Pro subscription is not an API key. Core
69
76
  skill profiles do not require one; credentialed MCP and runtime operations stay
@@ -134,6 +141,21 @@ them.
134
141
 
135
142
  ```bash
136
143
  loadout mcp-recipe --no-key
144
+ loadout mcp-recipe --credential-free
145
+
146
+ # Preview, configure, verify, then remove Playwright for Codex
147
+ loadout mcp-recipe playwright --agent codex
148
+ loadout mcp-recipe playwright --agent codex --yes
149
+ loadout mcp-recipe playwright --agent codex --verify
150
+ loadout remove mcp-recipe:playwright:codex
151
+ loadout remove mcp-recipe:playwright:codex --yes
152
+
153
+ # Repeat independently for Claude Code
154
+ loadout mcp-recipe playwright --agent claude-code
155
+ loadout mcp-recipe playwright --agent claude-code --yes
156
+ loadout mcp-recipe playwright --agent claude-code --verify
157
+ loadout remove mcp-recipe:playwright:claude-code
158
+ loadout remove mcp-recipe:playwright:claude-code --yes
137
159
  ```
138
160
 
139
161
  Expect Playwright MCP, Chrome DevTools MCP, and GitHub read-only. None requires a
@@ -161,7 +183,7 @@ cleanup deliberately deletes Loadout's snapshots, so it is the last lifecycle te
161
183
  ## Troubleshooting and recovery
162
184
 
163
185
  - **`loadout` is not found after installation:** confirm `npm install --global
164
- loadout-ai@0.4.0` completed, run `hash -r`, and confirm npm's global binary
186
+ loadout-ai@0.5.0` completed, run `hash -r`, and confirm npm's global binary
165
187
  directory is on `PATH`. For a source checkout, run `npm run build` and `npm link`.
166
188
  - **A preview asks for `--approve-risk`:** read the reported scripts, domains,
167
189
  credentials, binaries, or instruction findings. If you accept that specific plan,
@@ -172,6 +194,10 @@ loadout-ai@0.4.0` completed, run `hash -r`, and confirm npm's global binary
172
194
  legacy snapshot without post-mutation evidence. Run `loadout health --explain` and
173
195
  inspect the affected path before deciding whether an explicit force option is
174
196
  appropriate; do not delete the path merely to make the command pass.
197
+ - **Activation reports fewer additions for one agent:** this is expected when that
198
+ agent already has unmanaged or managed skills. `--limit` is a total per-agent
199
+ ceiling, not a request to add that many new skills. Recursively empty rollback
200
+ directories do not consume capacity and are safe for Loadout to reuse.
175
201
  - **A fetch, discovery, or update check fails:** retry only after checking network,
176
202
  proxy, DNS, and source-host access. Local inventory, library, health, rollback, and
177
203
  offline fixture tests remain separate; an unavailable live check is not a pass.
Binary file
@@ -214,14 +214,6 @@
214
214
  "tests/benchmark-campaign.test.ts"
215
215
  ]
216
216
  },
217
- {
218
- "id": "dashboard.local",
219
- "section": "Current beta limits",
220
- "summary": "The dashboard presents local Loadout data and does not establish hosted-service availability.",
221
- "evidenceClass": "integration-verified",
222
- "status": "bounded",
223
- "evidence": ["src/dashboard.ts", "tests/dashboard.test.ts"]
224
- },
225
217
  {
226
218
  "id": "uninstall.complete",
227
219
  "section": "Troubleshooting and uninstall",
@@ -0,0 +1,116 @@
1
+ # Loadout README Explainer Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Generate and integrate an original Loadout workflow infographic that accurately explains the preview-first, recoverable extension lifecycle.
6
+
7
+ **Architecture:** A single project-owned PNG becomes the README hero while the existing SVG remains available. The README product-flow test locks the asset path and accessible description, and the normal verification suite guards product claims and repository integrity.
8
+
9
+ **Tech Stack:** Built-in image generation, PNG asset, Markdown/HTML README, Vitest, npm verification scripts, GitHub Actions.
10
+
11
+ ## Global Constraints
12
+
13
+ - Use the exact five-stage order: Choose, Inspect, Preview, Apply, Undo.
14
+ - Name only checked-in supported agents: Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot, OpenCode, Windsurf, plus “+ 5 more adapters”.
15
+ - Use the exact footer: “Preview first · Managed changes · Snapshot-backed undo”.
16
+ - Do not add numerical safety, performance, popularity, or compatibility claims.
17
+ - Preserve `docs/assets/loadout-hero.svg`; create `docs/assets/loadout-workflow.png`.
18
+ - Use the exact alt text from the approved design spec.
19
+
20
+ ---
21
+
22
+ ### Task 1: Lock the README integration contract
23
+
24
+ **Files:**
25
+
26
+ - Modify: `tests/readme-product-flow.test.ts`
27
+ - Test: `tests/readme-product-flow.test.ts`
28
+
29
+ **Interfaces:**
30
+
31
+ - Consumes: the current README hero assertion.
32
+ - Produces: a test contract for `./docs/assets/loadout-workflow.png` and the approved alt text.
33
+
34
+ - [ ] **Step 1: Update the hero assertion**
35
+
36
+ ```ts
37
+ expect(readme).toContain(
38
+ '<img src="./docs/assets/loadout-workflow.png" alt="Loadout workflow: choose extensions, inspect sources, preview changes, apply through a managed snapshot, and undo safely across supported AI coding agents." width="960">',
39
+ );
40
+ ```
41
+
42
+ - [ ] **Step 2: Run the focused test and verify red**
43
+
44
+ Run: `npx vitest run tests/readme-product-flow.test.ts`
45
+
46
+ Expected: FAIL because `README.md` still references `loadout-hero.svg`.
47
+
48
+ ### Task 2: Generate and validate the infographic
49
+
50
+ **Files:**
51
+
52
+ - Create: `docs/assets/loadout-workflow.png`
53
+ - Preserve: `docs/assets/loadout-hero.svg`
54
+
55
+ **Interfaces:**
56
+
57
+ - Consumes: `docs/superpowers/specs/2026-07-20-loadout-readme-explainer-design.md` and the three supplied visual references.
58
+ - Produces: a wide, readable, original PNG suitable for a 960-pixel README presentation.
59
+
60
+ - [ ] **Step 1: Generate the project artwork**
61
+
62
+ Use built-in image generation with the supplied images as style references and this content contract: title “Your Agent Extensions, Under Control”; Choose → Inspect → Preview → Apply → Undo; a Loadout inventory card; the seven named agents; “+ 5 more adapters”; footer “Preview first · Managed changes · Snapshot-backed undo”; warm white hand-drawn infographic; no unsupported claims, logos, watermarks, gradients, shadows, or garbled text.
63
+
64
+ - [ ] **Step 2: Save the selected output**
65
+
66
+ Copy the generated artifact to `docs/assets/loadout-workflow.png` without modifying `docs/assets/loadout-hero.svg`.
67
+
68
+ - [ ] **Step 3: Inspect the saved PNG**
69
+
70
+ Open the workspace asset at original detail and verify every required label, composition, padding, and legibility. If any required text is wrong, perform one targeted image edit and inspect again.
71
+
72
+ ### Task 3: Integrate and verify the README
73
+
74
+ **Files:**
75
+
76
+ - Modify: `README.md`
77
+ - Test: `tests/readme-product-flow.test.ts`
78
+
79
+ **Interfaces:**
80
+
81
+ - Consumes: `docs/assets/loadout-workflow.png` and the test contract from Task 1.
82
+ - Produces: the rendered README hero reference and accessible description.
83
+
84
+ - [ ] **Step 1: Replace the README hero reference**
85
+
86
+ ```html
87
+ <img
88
+ src="./docs/assets/loadout-workflow.png"
89
+ alt="Loadout workflow: choose extensions, inspect sources, preview changes, apply through a managed snapshot, and undo safely across supported AI coding agents."
90
+ width="960"
91
+ />
92
+ ```
93
+
94
+ - [ ] **Step 2: Run the focused test and verify green**
95
+
96
+ Run: `npx vitest run tests/readme-product-flow.test.ts`
97
+
98
+ Expected: 9 passed and 1 skipped, with zero failures.
99
+
100
+ - [ ] **Step 3: Run the complete local gate**
101
+
102
+ Run: `npm run verify`
103
+
104
+ Expected: formatting, lint, type checking, evidence gates, 114 test files, CLI flow, README flow, package smoke, and performance checks all pass.
105
+
106
+ - [ ] **Step 4: Commit and push**
107
+
108
+ ```bash
109
+ git add README.md docs/assets/loadout-workflow.png tests/readme-product-flow.test.ts docs/superpowers/plans/2026-07-20-loadout-readme-explainer.md
110
+ git commit -m "docs: add Loadout workflow explainer"
111
+ git push origin main
112
+ ```
113
+
114
+ - [ ] **Step 5: Verify remote CI**
115
+
116
+ Run the normal push CI, then dispatch the full CI workflow on the final commit. Confirm the fast gate, Playwright diagnostics, and all six native OS/Node jobs complete successfully.