loadout-ai 0.3.1 → 0.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +67 -0
- package/MASTER_PLAN.md +177 -13
- package/README.md +157 -284
- package/dashboard/app.js +4 -4
- package/dashboard/index.html +4 -4
- package/dist/src/cli.js +27 -21
- package/dist/src/core/active-policy.js +172 -38
- package/dist/src/core/active-set.js +13 -4
- package/dist/src/core/adapters.js +10 -0
- package/dist/src/core/adopt.js +165 -32
- package/dist/src/core/agent-health-score.js +2 -2
- package/dist/src/core/catalog-coverage.js +2 -1
- package/dist/src/core/catalog-install.js +8 -1
- package/dist/src/core/catalog-release.js +2 -1
- package/dist/src/core/conformance.js +74 -0
- package/dist/src/core/install.js +11 -11
- package/dist/src/core/profiles.js +9 -4
- package/dist/src/core/ranking.js +1 -1
- package/dist/src/core/readme-claims.js +10 -0
- package/dist/src/core/readme-facts.js +40 -0
- package/dist/src/core/recommend.js +104 -12
- package/dist/src/core/runtime-tools.js +5 -2
- package/dist/src/core/scheduler.js +2 -1
- package/dist/src/core/snapshot.js +58 -13
- package/dist/src/core/state.js +8 -1
- package/dist/src/core/target-occupancy.js +50 -0
- package/dist/src/core/transaction.js +2 -1
- package/dist/src/core/uninstall.js +26 -2
- package/dist/src/dashboard.js +5 -2
- package/dist/src/shared/schemas.js +57 -0
- package/docs/FEATURE_TEST_MATRIX.md +16 -0
- package/docs/README_RESEARCH.md +36 -0
- package/docs/RELEASE_REVIEW.md +31 -5
- package/docs/REPOSITORY_STABILIZATION.md +190 -0
- package/docs/TESTING.md +50 -0
- package/docs/USER_TEST_GUIDE.md +42 -4
- package/docs/assets/loadout-hero.svg +259 -0
- package/docs/assets/loadout-mark.svg +54 -0
- package/docs/evidence/live-checks-2026-07-19.json +22 -0
- package/docs/evidence/live-checks.schema.json +28 -0
- package/docs/evidence/readme-claims.json +286 -0
- package/docs/superpowers/plans/2026-07-19-relatable-readme-hero.md +283 -0
- package/docs/superpowers/plans/2026-07-20-project-activation-safety.md +469 -0
- package/docs/superpowers/specs/2026-07-19-relatable-readme-hero-design.md +80 -0
- package/docs/superpowers/specs/2026-07-20-project-activation-safety-design.md +228 -0
- package/package.json +8 -4
- package/SIMPLE_PLAN.md +0 -44
- package/docs/plans/2026-07-18-release-0.3.md +0 -42
- package/docs/superpowers/plans/2026-07-18-cli-ux-polish.md +0 -86
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
# Project Activation Safety and Relevance Design
|
|
2
|
+
|
|
3
|
+
**Status:** Approved direction; awaiting written-spec review
|
|
4
|
+
**Date:** 2026-07-20
|
|
5
|
+
**Release target:** Loadout 0.4.1
|
|
6
|
+
|
|
7
|
+
## Problem
|
|
8
|
+
|
|
9
|
+
The founder acceptance run of `loadout activate --project . --agents
|
|
10
|
+
codex,claude-code --limit 30` exposed three connected defects:
|
|
11
|
+
|
|
12
|
+
1. The preview reported `0/30 active` even though Claude Code already had 12
|
|
13
|
+
unmanaged skills. The planner counted only active Loadout records, so applying the
|
|
14
|
+
proposal could exceed the disclosed per-agent limit.
|
|
15
|
+
2. An earlier rollback correctly restored recursively empty skill directories. The
|
|
16
|
+
activation planner treated the existence of those empty directories as occupied
|
|
17
|
+
content and blocked safe activation for both agents.
|
|
18
|
+
3. Project detection recognized only JavaScript/TypeScript and Playwright. It missed
|
|
19
|
+
strong local evidence that this repository is a Node CLI, an npm package, a Vitest
|
|
20
|
+
project, an MCP-related tool, and a release/supply-chain project. Ranking then used
|
|
21
|
+
the entire 30-slot budget as a quota and proposed redundant or mismatched skills,
|
|
22
|
+
including several overlapping Playwright workflows and a Jest skill for a Vitest
|
|
23
|
+
repository.
|
|
24
|
+
|
|
25
|
+
No activation was applied. The Maximum library remains disabled, and the founder's
|
|
26
|
+
12 unmanaged Claude skills remain unchanged.
|
|
27
|
+
|
|
28
|
+
## Goals
|
|
29
|
+
|
|
30
|
+
- Make `--limit` a truthful ceiling over all active skills visible to each selected
|
|
31
|
+
agent, whether managed by Loadout or not.
|
|
32
|
+
- Plan and explain capacity separately for every selected agent.
|
|
33
|
+
- Treat recursively empty rollback residue as unoccupied while continuing to refuse
|
|
34
|
+
every non-empty, symlinked, unreadable, or unsupported target.
|
|
35
|
+
- Improve local project recognition enough to distinguish this Node CLI/npm/Vitest
|
|
36
|
+
repository from a generic Playwright web application.
|
|
37
|
+
- Select a compact, diverse working set instead of filling every available slot with
|
|
38
|
+
weak or redundant matches.
|
|
39
|
+
- Preserve a preview-first, transactional, rollback-safe mutation path.
|
|
40
|
+
|
|
41
|
+
## Non-goals
|
|
42
|
+
|
|
43
|
+
- No model or external API call is required for recommendation or activation.
|
|
44
|
+
- No project source, filename, dependency, or outcome data leaves the machine.
|
|
45
|
+
- This work does not claim that a deterministic recommendation proves global package
|
|
46
|
+
quality.
|
|
47
|
+
- This work does not install deferred MCP servers or collect credentials.
|
|
48
|
+
- Dashboard removal remains a separate 0.4.1 task.
|
|
49
|
+
|
|
50
|
+
## Design
|
|
51
|
+
|
|
52
|
+
### 1. One shared definition of an occupied skill target
|
|
53
|
+
|
|
54
|
+
Loadout will use one filesystem predicate for both initial installation and later
|
|
55
|
+
library activation:
|
|
56
|
+
|
|
57
|
+
- A missing path is unoccupied.
|
|
58
|
+
- A real directory containing no entries at any depth is unoccupied.
|
|
59
|
+
- A regular file, symlink, special entry, unreadable directory, or directory with any
|
|
60
|
+
non-directory descendant is occupied.
|
|
61
|
+
- Recursive inspection is bounded to 10,000 entries. Reaching the bound is treated as
|
|
62
|
+
occupied, never as empty.
|
|
63
|
+
|
|
64
|
+
Activation will re-check this predicate inside the transaction immediately before
|
|
65
|
+
copying. This closes the preview/apply race without deleting or adopting user content.
|
|
66
|
+
Empty target directories may be removed immediately before the copy because they
|
|
67
|
+
contain no bytes to preserve; rollback still restores the pre-transaction topology.
|
|
68
|
+
|
|
69
|
+
### 2. Per-agent total active capacity
|
|
70
|
+
|
|
71
|
+
For each requested agent, the planner will scan the agent's real skill root and count
|
|
72
|
+
every directory containing a valid `SKILL.md`. This inventory already distinguishes
|
|
73
|
+
managed and unmanaged skills and will be the capacity source of truth.
|
|
74
|
+
|
|
75
|
+
For agent `a`:
|
|
76
|
+
|
|
77
|
+
```text
|
|
78
|
+
available(a) = max(0, limit - inventory(a).total)
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
Candidate ranking is deterministic and shared, but selection is sliced independently
|
|
82
|
+
to each agent's available capacity. In the founder fixture, Claude Code has 12 active
|
|
83
|
+
unmanaged skills and may receive at most 18 additions; Codex has zero active skills and
|
|
84
|
+
may receive at most 30. An agent at its limit receives no changes and a clear warning;
|
|
85
|
+
it does not prevent another requested agent from using its own available capacity.
|
|
86
|
+
|
|
87
|
+
The plan must never count an empty directory as an active skill. Existing managed
|
|
88
|
+
skills already in the desired state count toward capacity but are not proposed again.
|
|
89
|
+
|
|
90
|
+
### 3. Project signals
|
|
91
|
+
|
|
92
|
+
Project scanning remains deterministic, bounded, and local-only. It will read known
|
|
93
|
+
root manifests/configuration and a bounded set of repository filenames. For Node
|
|
94
|
+
projects it will derive explicit signals from `package.json`:
|
|
95
|
+
|
|
96
|
+
- `bin` -> Node CLI
|
|
97
|
+
- package name plus `publishConfig` or non-private package -> npm package/release
|
|
98
|
+
- `vitest` -> Vitest
|
|
99
|
+
- `jest` -> Jest
|
|
100
|
+
- `@playwright/test` or Playwright config -> Playwright
|
|
101
|
+
- `commander` -> command-line application
|
|
102
|
+
- `zod` -> schema validation
|
|
103
|
+
- scripts containing `prepack`, package-smoke, or release checks -> package/release
|
|
104
|
+
|
|
105
|
+
Repository filenames add delivery, security, MCP, and documentation signals only when
|
|
106
|
+
matching explicit reviewed patterns such as `.github/workflows`, `SECURITY.md`, MCP-
|
|
107
|
+
named modules, and package/release scripts. The scanner does not recursively read
|
|
108
|
+
arbitrary project source or send any data elsewhere.
|
|
109
|
+
|
|
110
|
+
Human output will name the important detected roles, for example:
|
|
111
|
+
|
|
112
|
+
```text
|
|
113
|
+
Detected: TypeScript, Node CLI, npm package, Vitest, Playwright, MCP tooling
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
### 4. Relevant and diverse selection
|
|
117
|
+
|
|
118
|
+
The existing source-priority tie-break remains, but a capacity ceiling is not a quota.
|
|
119
|
+
Selection stops when no candidate reaches the evidence threshold.
|
|
120
|
+
|
|
121
|
+
Ranking will:
|
|
122
|
+
|
|
123
|
+
- reward exact tool and role matches such as Vitest, CLI design, npm packaging,
|
|
124
|
+
TypeScript, MCP security, release verification, and documentation;
|
|
125
|
+
- reject framework mismatches such as Jest-only guidance when Vitest is detected and
|
|
126
|
+
Jest is not;
|
|
127
|
+
- group close alternatives into capability families such as browser testing, code
|
|
128
|
+
review, documentation, planning, security, and architecture;
|
|
129
|
+
- select the highest-evidence representative before taking another member of the same
|
|
130
|
+
family;
|
|
131
|
+
- cap weak generic foundation choices so they cannot crowd out project-specific
|
|
132
|
+
evidence;
|
|
133
|
+
- continue to honor explicit `--pin` selectors, while disclosing when a pin consumes a
|
|
134
|
+
slot or conflicts with the limit.
|
|
135
|
+
|
|
136
|
+
The same skill name from multiple repositories remains an alternative, not two active
|
|
137
|
+
copies. Multiple distinct skills in one family are allowed only when each has separate
|
|
138
|
+
strong project evidence.
|
|
139
|
+
|
|
140
|
+
### 5. Recommendation type clarity
|
|
141
|
+
|
|
142
|
+
Package-level `loadout recommend` output will label every suggestion as one of:
|
|
143
|
+
|
|
144
|
+
- `skill library` — eligible for disabled-library activation;
|
|
145
|
+
- `MCP/runtime setup` — requires a separate explicit preview and may require a
|
|
146
|
+
non-model credential;
|
|
147
|
+
- `unavailable` — not prepared locally, with the reason.
|
|
148
|
+
|
|
149
|
+
Playwright MCP and GitHub MCP must not look like ordinary automatically activatable
|
|
150
|
+
skills. The command will show the next preview command for explicit integrations and
|
|
151
|
+
will continue to state that recommendations are rule-selected, not proof of quality.
|
|
152
|
+
|
|
153
|
+
### 6. Preview and apply output
|
|
154
|
+
|
|
155
|
+
The default human preview will begin with one compact block per agent:
|
|
156
|
+
|
|
157
|
+
```text
|
|
158
|
+
Claude Code: 12 active (0 managed, 12 unmanaged), 18/30 slots available
|
|
159
|
+
Codex: 0 active, 30/30 slots available
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
It will then show additions grouped by agent and capability, followed by blockers. A
|
|
163
|
+
recursively empty target is not printed as a blocker. A real conflict prints the exact
|
|
164
|
+
target and a non-destructive next action. JSON output retains full scores, reasons,
|
|
165
|
+
alternatives, per-agent budgets, targets, and blocker details.
|
|
166
|
+
|
|
167
|
+
`--yes` is rejected when any proposed target is blocked. A successful apply creates
|
|
168
|
+
one transaction snapshot covering every selected agent and prints the exact rollback
|
|
169
|
+
command.
|
|
170
|
+
|
|
171
|
+
## Data flow
|
|
172
|
+
|
|
173
|
+
1. Scan project metadata into local project signals.
|
|
174
|
+
2. Scan each requested agent's active skill inventory.
|
|
175
|
+
3. Read reviewed disabled-library records and local-only outcomes.
|
|
176
|
+
4. Score and diversify eligible candidates.
|
|
177
|
+
5. Slice the ordered candidates independently to each agent's available capacity.
|
|
178
|
+
6. Build exact activation changes and inspect targets with the shared occupancy rule.
|
|
179
|
+
7. Print the preview; stop unless `--yes` is present.
|
|
180
|
+
8. Inside the transaction, re-read state, re-check capacity and target occupancy, copy
|
|
181
|
+
reviewed library bytes, fingerprint the result, persist activation state, and emit
|
|
182
|
+
the rollback snapshot.
|
|
183
|
+
|
|
184
|
+
## Failure handling
|
|
185
|
+
|
|
186
|
+
- If inventory scanning fails for an agent, activation for that agent is blocked; its
|
|
187
|
+
capacity is never guessed.
|
|
188
|
+
- If the project manifest is malformed, show the exact manifest error and continue
|
|
189
|
+
only with signals that remain trustworthy.
|
|
190
|
+
- If state or filesystem contents change between preview and apply, abort the entire
|
|
191
|
+
multi-agent transaction without partial activation.
|
|
192
|
+
- If fewer candidates meet the evidence threshold than available slots, activate only
|
|
193
|
+
those candidates and explain that unused capacity is intentional.
|
|
194
|
+
|
|
195
|
+
## Tests and acceptance criteria
|
|
196
|
+
|
|
197
|
+
Automated tests must prove:
|
|
198
|
+
|
|
199
|
+
1. A recursively empty target does not block preview or apply.
|
|
200
|
+
2. A target containing a file, symlink, unreadable entry, or more than the inspection
|
|
201
|
+
bound remains blocked and unchanged.
|
|
202
|
+
3. Empty-directory handling is identical in initial setup and project activation.
|
|
203
|
+
4. Twelve unmanaged Claude skills plus an 18-skill proposal never exceeds a limit of 30.
|
|
204
|
+
5. The same plan may select 18 additions for Claude and 30 for an empty Codex profile.
|
|
205
|
+
6. An agent already at capacity receives no additions while another agent can proceed.
|
|
206
|
+
7. Apply re-checks inventory and occupancy and aborts atomically when either changes.
|
|
207
|
+
8. The Loadout repository fixture detects Node CLI, npm package, Vitest, Playwright,
|
|
208
|
+
Commander, Zod, release, security, and MCP signals.
|
|
209
|
+
9. A Vitest-only fixture never recommends a Jest-only skill.
|
|
210
|
+
10. Redundant Playwright/browser candidates do not consume most of the active set.
|
|
211
|
+
11. Recommendation output distinguishes skill libraries from explicit MCP/runtime
|
|
212
|
+
setup.
|
|
213
|
+
12. Human output shows per-agent managed/unmanaged counts and unused capacity; JSON
|
|
214
|
+
exposes the same facts structurally.
|
|
215
|
+
13. Existing Stable, Power, Maximum, rollback, and non-overwrite tests remain green.
|
|
216
|
+
|
|
217
|
+
Founder acceptance resumes only after the published release candidate previews a
|
|
218
|
+
non-blocked plan, reports Claude's existing 12 skills, respects both per-agent limits,
|
|
219
|
+
applies transactionally, and restores its explicit snapshot without changing unmanaged
|
|
220
|
+
content.
|
|
221
|
+
|
|
222
|
+
## Compatibility
|
|
223
|
+
|
|
224
|
+
Existing `loadout activate` and `loadout optimize` flags remain valid. The meaning of
|
|
225
|
+
`--limit` is corrected from an implicit managed-only count to the documented total
|
|
226
|
+
active skill ceiling. JSON consumers must migrate from one global `activeBefore` and
|
|
227
|
+
`capacity` pair to per-agent budget records; the 0.4.1 changelog will call out this
|
|
228
|
+
schema correction.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "loadout-ai",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.4.1",
|
|
4
4
|
"private": false,
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"description": "Universal upgrade manager for AI coding agents",
|
|
@@ -16,7 +16,6 @@
|
|
|
16
16
|
"README.md",
|
|
17
17
|
"CHANGELOG.md",
|
|
18
18
|
"MASTER_PLAN.md",
|
|
19
|
-
"SIMPLE_PLAN.md",
|
|
20
19
|
"SECURITY.md",
|
|
21
20
|
"LICENSE"
|
|
22
21
|
],
|
|
@@ -55,13 +54,18 @@
|
|
|
55
54
|
"test:e2e:dashboard": "playwright test",
|
|
56
55
|
"pretest:e2e:cli": "npm run build",
|
|
57
56
|
"test:e2e:cli": "node scripts/cli-product-flow.mjs",
|
|
57
|
+
"test:e2e:readme": "node scripts/readme-product-flow.mjs",
|
|
58
58
|
"test:package": "node scripts/package-smoke.mjs",
|
|
59
59
|
"pretest:performance": "npm run build",
|
|
60
60
|
"test:performance": "node scripts/scan-benchmark.mjs",
|
|
61
61
|
"typecheck": "tsc -p tsconfig.json --noEmit",
|
|
62
|
-
"
|
|
62
|
+
"readme:update": "node scripts/update-readme-facts.mjs",
|
|
63
|
+
"readme:check": "node scripts/update-readme-facts.mjs --check",
|
|
64
|
+
"check:readme-claims": "npm run build && node scripts/check-readme-claims.mjs",
|
|
65
|
+
"check:live": "node scripts/check-live-evidence.mjs",
|
|
66
|
+
"check:evidence": "node scripts/check-catalog-attribution.mjs && node scripts/check-discovery-artifacts.mjs && npm run check:readme-claims && node --import tsx scripts/check-release-claims.ts",
|
|
63
67
|
"feed:build": "node --import tsx scripts/build-intelligence-feed.ts",
|
|
64
|
-
"verify": "npm run format:check && npm run lint && npm run typecheck && npm run check:evidence && npm test -- --run && npm run test:e2e:cli && npm run test:package && npm run test:performance",
|
|
68
|
+
"verify": "npm run format:check && npm run lint && npm run typecheck && npm run check:evidence && npm test -- --run && npm run test:e2e:cli && npm run test:e2e:readme && npm run test:package && npm run test:performance",
|
|
65
69
|
"verify:full": "npm run verify && npm run test:e2e:dashboard"
|
|
66
70
|
},
|
|
67
71
|
"dependencies": {
|
package/SIMPLE_PLAN.md
DELETED
|
@@ -1,44 +0,0 @@
|
|
|
1
|
-
# Loadout plan — simple version
|
|
2
|
-
|
|
3
|
-
Loadout's primary experience is CLI-first:
|
|
4
|
-
|
|
5
|
-
```bash
|
|
6
|
-
npx loadout-ai
|
|
7
|
-
```
|
|
8
|
-
|
|
9
|
-
It detects installed agents, scans the actual skills already present, and recommends a
|
|
10
|
-
small Stable foundation. Maximum Library and Custom are explicit alternatives. It
|
|
11
|
-
previews reviewed pinned sources, exact overlaps, capacity, and safety findings before
|
|
12
|
-
one rollback-safe transaction. The dashboard is optional diagnostics.
|
|
13
|
-
|
|
14
|
-
Loadout also provides the package-manager operations:
|
|
15
|
-
|
|
16
|
-
- Find, install, update, remove, create, share, and synchronize AI-agent add-ons.
|
|
17
|
-
- Work with skills, commands, rules, agents, plugins, and MCP tools.
|
|
18
|
-
- Support Codex, Claude Code, Cursor, and more from one setup file.
|
|
19
|
-
|
|
20
|
-
Loadout's four major advantages are:
|
|
21
|
-
|
|
22
|
-
1. **Safety:** scan first, explain every change, block dangerous behavior, and never
|
|
23
|
-
touch unrelated files.
|
|
24
|
-
2. **Recovery:** back up before changes and provide one-command undo.
|
|
25
|
-
3. **Guidance:** check setup health and recommend tested add-on collections for the
|
|
26
|
-
user's project.
|
|
27
|
-
4. **Optimization:** keep a broad reviewed library but expose only the best supported,
|
|
28
|
-
non-overlapping active set for the current agent and project.
|
|
29
|
-
|
|
30
|
-
The original build order was:
|
|
31
|
-
|
|
32
|
-
1. Finish the reliable package-manager foundation.
|
|
33
|
-
2. Match OpenPackage's package types, sources, synchronization, and publishing.
|
|
34
|
-
3. Add health checks, security scanning, safe updates, recommendations, and profiles.
|
|
35
|
-
4. Keep the dashboard optional and prove every supported platform with tests.
|
|
36
|
-
|
|
37
|
-
That foundation is now integrated on `main`. [MASTER_PLAN.md](./MASTER_PLAN.md) is the
|
|
38
|
-
only canonical checklist; contributor branches and notes are historical inputs, not
|
|
39
|
-
separate sources of project status. Phase 12 now tracks provenance for unmanaged
|
|
40
|
-
skills, evidence-backed comparison, library-versus-active-set state, safe adoption and
|
|
41
|
-
enable/disable, project activation, guided optimization, category evaluations, daily
|
|
42
|
-
review queues, provider/MCP workflows, CLI polish, npm publication, and public-beta
|
|
43
|
-
testing. Catalog expansion, legal review, keychains, additional adapters, and submission
|
|
44
|
-
work remain explicit rather than hidden behind earlier checked implementation proofs.
|
|
@@ -1,42 +0,0 @@
|
|
|
1
|
-
# Loadout 0.3 Product-Hardening Plan
|
|
2
|
-
|
|
3
|
-
## Goal
|
|
4
|
-
|
|
5
|
-
Ship a user-testable CLI release with one-command cleanup, honest whole-profile update
|
|
6
|
-
checks, explicit no-key MCP discovery, and a verified local dashboard.
|
|
7
|
-
|
|
8
|
-
## Public contracts
|
|
9
|
-
|
|
10
|
-
1. `loadout uninstall` is a preview. `loadout uninstall --yes` removes scheduled
|
|
11
|
-
Loadout jobs, runtime tools, managed agent files, disabled library copies, cache,
|
|
12
|
-
snapshots, and state while preserving unmanaged or modified files. `--remove-cli`
|
|
13
|
-
additionally runs the global npm uninstall.
|
|
14
|
-
2. `loadout update` checks every tracked package and compares the last installed
|
|
15
|
-
Stable/Power/Maximum profile with the current trusted catalog. It never mutates by
|
|
16
|
-
default.
|
|
17
|
-
3. `loadout update --yes` reapplies trusted profile changes and applies all
|
|
18
|
-
statically-safe active package updates; blocked, disabled, or failed updates are
|
|
19
|
-
reported and skipped. `--package <id> --yes` remains the precise path.
|
|
20
|
-
4. Stable, Power, and Maximum are re-evaluated against the trusted catalog whenever
|
|
21
|
-
`update` runs. Discovery can nominate new candidates daily, but candidates cannot
|
|
22
|
-
enter a profile until immutable source, license, compatibility, and safety review
|
|
23
|
-
are recorded.
|
|
24
|
-
5. `loadout mcp-recipe --no-key` lists reviewed recipes requiring no separately
|
|
25
|
-
billed AI/model API key. It includes GitHub read-only while separately disclosing
|
|
26
|
-
the GitHub token. `--credential-free` is the stricter zero-credential filter.
|
|
27
|
-
6. `loadout dashboard` remains loopback-only and is verified by the browser E2E test.
|
|
28
|
-
|
|
29
|
-
## Implementation sequence
|
|
30
|
-
|
|
31
|
-
- [x] Add failing tests for uninstall preview/application and CLI help.
|
|
32
|
-
- [x] Implement uninstall orchestration with dependency injection and path guards.
|
|
33
|
-
- [x] Add failing tests for persisted profile state and profile drift reporting.
|
|
34
|
-
- [x] Record setup mode in snapshot-restorable install state and include it in update.
|
|
35
|
-
- [x] Add failing tests for bulk safe update selection and clear skipped results.
|
|
36
|
-
- [x] Implement `update --yes`, retaining `--apply` as a compatible alias.
|
|
37
|
-
- [x] Add and test Chrome DevTools as a pinned no-key MCP recipe.
|
|
38
|
-
- [x] Add `mcp-recipe --no-key` and beginner-readable recipe output.
|
|
39
|
-
- [x] Update README, test guide, master plan, changelog, completion, and version.
|
|
40
|
-
- [ ] Run unit, lint, type, evidence, package, CLI E2E, performance, dashboard E2E,
|
|
41
|
-
and npm dry-run gates.
|
|
42
|
-
- [ ] Review, commit, merge to main, push, tag, and publish to npm.
|
|
@@ -1,86 +0,0 @@
|
|
|
1
|
-
# CLI UX Polish Implementation Plan
|
|
2
|
-
|
|
3
|
-
> **For agentic workers:** REQUIRED SUB-SKILL: Use `executing-plans` to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
|
4
|
-
|
|
5
|
-
**Goal:** Make Loadout's everyday CLI path understandable for a first-time user while preserving its existing safe, advanced capabilities.
|
|
6
|
-
|
|
7
|
-
**Architecture:** Keep the CLI as the primary interface. Add one concise, read-only `guide` command and a focused help footer; retain advanced commands but remove them from the first-screen help. Fix JSON output and package-scoped update previews as explicit public CLI contracts. Document a safe test journey and keep the dashboard as an optional local companion.
|
|
8
|
-
|
|
9
|
-
**Tech Stack:** Node.js 20+, TypeScript, Commander, Vitest, framework-free local dashboard.
|
|
10
|
-
|
|
11
|
-
## Global Constraints
|
|
12
|
-
|
|
13
|
-
- Never mutate a user's agent configuration from a default or preview command.
|
|
14
|
-
- Existing command names remain valid even if hidden from first-screen help.
|
|
15
|
-
- Machine-readable commands must emit valid JSON when `--json` is accepted.
|
|
16
|
-
- Do not require an API key or GitHub account for the core test journey.
|
|
17
|
-
- Every behavior change gets a test before its production code.
|
|
18
|
-
|
|
19
|
-
---
|
|
20
|
-
|
|
21
|
-
### Task 1: Define the beginner CLI contract
|
|
22
|
-
|
|
23
|
-
**Files:**
|
|
24
|
-
|
|
25
|
-
- Modify: `tests/cli-help.test.ts`
|
|
26
|
-
- Modify: `src/cli.ts`
|
|
27
|
-
|
|
28
|
-
**Interfaces:**
|
|
29
|
-
|
|
30
|
-
- Produces: `loadout guide`, a read-only command that explains setup, discovery, project recommendations, recovery, dashboard, and help.
|
|
31
|
-
|
|
32
|
-
- [x] **Step 1: Write failing CLI contract tests** for `loadout guide`, a focused top-level help footer, and retained access to an advanced command.
|
|
33
|
-
- [x] **Step 2: Run** `npm test -- tests/cli-help.test.ts` **and confirm the new assertions fail.**
|
|
34
|
-
- [x] **Step 3: Implement** `guide`, hide maintainer-only commands from first-screen help, and add a short help footer that points to the guide.
|
|
35
|
-
- [x] **Step 4: Run** `npm test -- tests/cli-help.test.ts` **and confirm it passes.**
|
|
36
|
-
- [ ] **Step 5: Commit** the focused CLI discoverability change.
|
|
37
|
-
|
|
38
|
-
### Task 2: Repair machine-readable and scoped preview contracts
|
|
39
|
-
|
|
40
|
-
**Files:**
|
|
41
|
-
|
|
42
|
-
- Modify: `tests/cli-help.test.ts`
|
|
43
|
-
- Modify: `tests/update.test.ts`
|
|
44
|
-
- Modify: `src/cli.ts`
|
|
45
|
-
- Modify: `src/core/update.ts`
|
|
46
|
-
|
|
47
|
-
**Interfaces:**
|
|
48
|
-
|
|
49
|
-
- `loadout catalog --json` returns a JSON array.
|
|
50
|
-
- `loadout update --package <id>` plans only that managed package and rejects an unknown installed package clearly.
|
|
51
|
-
|
|
52
|
-
- [x] **Step 1: Write failing tests** for catalog JSON and package-scoped update planning.
|
|
53
|
-
- [x] **Step 2: Run the focused test files** and confirm the assertions fail for the existing implementation.
|
|
54
|
-
- [x] **Step 3: Implement the smallest compatible fixes.**
|
|
55
|
-
- [x] **Step 4: Run the focused tests** and confirm they pass.
|
|
56
|
-
- [ ] **Step 5: Commit** the contract fixes separately.
|
|
57
|
-
|
|
58
|
-
### Task 3: Record user testing and current product scope
|
|
59
|
-
|
|
60
|
-
**Files:**
|
|
61
|
-
|
|
62
|
-
- Create: `docs/USER_TEST_GUIDE.md`
|
|
63
|
-
- Modify: `MASTER_PLAN.md`
|
|
64
|
-
|
|
65
|
-
**Interfaces:**
|
|
66
|
-
|
|
67
|
-
- The guide gives a real-profile-safe order of commands, says which commands change state, and gives rollback instructions.
|
|
68
|
-
- `MASTER_PLAN.md` begins with the one authoritative unfinished-work section and moves stale immediate tasks out of the active path.
|
|
69
|
-
|
|
70
|
-
- [x] **Step 1: Add a concise user testing guide** for daily use, discovery, optional dashboard, safe mutation, rollback, and advanced validation.
|
|
71
|
-
- [x] **Step 2: Add the current remaining work section** and mark historic, non-essential ideas as deferred rather than pretending they are active product requirements.
|
|
72
|
-
- [x] **Step 3: Run markdown and CLI smoke checks** to make sure command names match the real surface.
|
|
73
|
-
- [ ] **Step 4: Commit** documentation separately.
|
|
74
|
-
|
|
75
|
-
### Task 4: Validate the actual user journey
|
|
76
|
-
|
|
77
|
-
**Files:**
|
|
78
|
-
|
|
79
|
-
- Test: `tests/cli-help.test.ts`
|
|
80
|
-
- Test: `tests/update.test.ts`
|
|
81
|
-
- Test: `tests/dashboard.test.ts`
|
|
82
|
-
|
|
83
|
-
- [x] **Step 1: Build and run** the CLI guide, catalog JSON, health, library, project recommendation preview, and dashboard endpoint checks without changing agent files.
|
|
84
|
-
- [ ] **Step 2: Run** `npm run verify:full` **and inspect every failure if any.**
|
|
85
|
-
- [ ] **Step 3: Review the diff for accidental profile/cache/secrets changes.**
|
|
86
|
-
- [ ] **Step 4: Commit, then present the exact install/test/recovery commands to the user.**
|