loadout-ai 0.7.0 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/CHANGELOG.md +72 -0
  2. package/README.md +33 -33
  3. package/catalog/discovered.json +26880 -24184
  4. package/dist/src/cli.js +5 -0
  5. package/dist/src/commands/catalog.js +103 -116
  6. package/dist/src/core/agents/agent-inspection.js +26 -4
  7. package/dist/src/core/catalog/registry.js +58 -10
  8. package/dist/src/core/catalog/safety.js +36 -7
  9. package/dist/src/core/install/source.js +21 -7
  10. package/dist/src/core/reporting/cli-guide.js +3 -3
  11. package/dist/src/core/reporting/completion.js +42 -95
  12. package/dist/src/core/reporting/doctor.js +3 -5
  13. package/dist/src/core/routing/handoff.js +94 -58
  14. package/dist/src/core/routing/policy.js +147 -0
  15. package/dist/src/core/routing/route.js +25 -153
  16. package/docs/CANDIDATE_INTELLIGENCE.md +9 -2
  17. package/docs/CATALOG.md +1 -1
  18. package/docs/CREDENTIAL_AND_UPDATE_POLICY.md +1 -1
  19. package/docs/DISCOVERED.md +252 -251
  20. package/docs/FEATURE_TEST_MATRIX.md +7 -260
  21. package/docs/GITHUB_AUTHORIZATION.md +5 -0
  22. package/docs/PROVENANCE_AND_COMPARISON.md +1 -1
  23. package/docs/RELEASE_REVIEW.md +0 -1
  24. package/package.json +6 -4
  25. package/skills/loadout-router/SKILL.md +43 -78
  26. package/MASTER_PLAN.md +0 -2207
  27. package/docs/ACTIVE_SET.md +0 -53
  28. package/docs/COMPATIBILITY_POLICY.md +0 -22
  29. package/docs/CONVERSION_AND_SANDBOX.md +0 -27
  30. package/docs/EVALUATION_PROTOCOL_V1.md +0 -300
  31. package/docs/HEAD_TO_HEAD_EVALUATION.md +0 -79
  32. package/docs/PROVIDER_CONFIGURATION.md +0 -45
  33. package/docs/README_RESEARCH.md +0 -36
  34. package/docs/REPOSITORY_STABILIZATION.md +0 -190
  35. package/docs/SAFE_UPDATE_DEMO.md +0 -25
  36. package/docs/SCHEMA_DECISIONS.md +0 -25
  37. package/docs/SUBMISSION_COPY.md +0 -90
  38. package/docs/TEAM_POLICY.md +0 -18
  39. package/docs/superpowers/plans/2026-07-19-relatable-readme-hero.md +0 -283
  40. package/docs/superpowers/plans/2026-07-20-loadout-readme-explainer.md +0 -116
  41. package/docs/superpowers/plans/2026-07-20-project-activation-safety.md +0 -469
  42. package/docs/superpowers/specs/2026-07-19-relatable-readme-hero-design.md +0 -80
  43. package/docs/superpowers/specs/2026-07-20-loadout-readme-explainer-design.md +0 -55
  44. package/docs/superpowers/specs/2026-07-20-project-activation-safety-design.md +0 -228
package/MASTER_PLAN.md DELETED
@@ -1,2207 +0,0 @@
1
- # Loadout Master Plan
2
-
3
- Status: Approved baseline for implementation
4
- Hackathon: OpenAI Build Week 2026
5
- Category: Developer Tools
6
- Team size: 3
7
- Target submission: July 21, 2026 at 5:00 PM Pacific / July 22 at 4:00 AM Dubai
8
-
9
- ## Launch finish line (authoritative, July 21, 2026)
10
-
11
- This is the only active checklist. Everything under **Archived implementation
12
- history** is evidence of how the product was built, not unfinished launch scope.
13
-
14
- ### Product complete
15
-
16
- - [x] Publish the CLI as `loadout-ai`; keep the product usable without cloning
17
- this repository or providing an OpenAI/Anthropic API key.
18
- - [x] Ship Stable, Power, Maximum, and Custom profiles with preview-first apply,
19
- snapshots, drift protection, explicit rollback, removal, and complete uninstall.
20
- - [x] Detect and manage skills across supported agent adapters while preserving
21
- unrelated and unmanaged files.
22
- - [x] Ship a 53-source pinned catalog, daily read-only discovery/update checks,
23
- existing-skill reconciliation, and project-aware recommendation/optimization.
24
- - [x] Credit and support the MIT Humanizer and Obsidian Skills sources. Humanizer is
25
- available through Custom mode; Obsidian is recommended for detected vaults.
26
- - [x] Ship three explicit MCP recipes for Playwright, Chrome DevTools, and read-only
27
- GitHub, plus the separately reviewed Graphify runtime-tool recipe.
28
- - [x] Remove the conflicting dashboard and make the CLI the only product surface.
29
- - [x] Publish clear npm setup, testing, upstream attribution, trust boundaries, and a
30
- detailed account of how Codex and GPT-5.6 were used in the README.
31
- - [x] Pass the complete local release gate and the latest hosted CI run on `main`.
32
- - [x] Preserve pre-existing adopted skills during removal and complete uninstall;
33
- distinguish additive Custom installation from whole-profile Custom setup, and
34
- show truthful progress during large approved uninstalls.
35
-
36
- ### Final founder acceptance
37
-
38
- - [x] Install the exact public npm release, verify `loadout --version`, and scan the
39
- founder's restored real profiles. Stable, Power, Maximum, project optimization,
40
- Humanizer, Graphify, both-host Playwright MCP lifecycles, rollback, complete
41
- uninstall, and clean reinstall were exercised during founder acceptance.
42
- - [x] Keep the GitHub repository name `loadout`. The product remains **Loadout** and
43
- the npm package/CLI identities remain `loadout-ai` / `loadout`.
44
- - [x] Record explicit decisions for all six `NOASSERTION` sources in
45
- `docs/UPSTREAM_LICENSE_DECISIONS.md` without inventing a license claim.
46
-
47
- ### Submission work
48
-
49
- - [x] Record the real CLI demo using `docs/DEMO_SCRIPT.md`; keep it under three
50
- minutes, include a voiceover, upload it publicly to YouTube, and verify the URL.
51
- - [x] Replace the README demo placeholder with the final YouTube link.
52
- - [x] In the voiceover, explain what Loadout does and how Codex and GPT-5.6 were used.
53
- - [ ] Run `/feedback` in Codex, copy the resulting session ID, and enter it in Devpost.
54
- - [x] Confirm the public repository URL is accessible to Devpost and OpenAI.
55
- - [x] Confirm all team invitations are accepted. Both teammates are collaborators and
56
- there are no pending repository invitations.
57
- - [ ] Select **Developer Tools**, complete the edited Devpost description, and submit
58
- rather than leaving a draft.
59
-
60
- ### Deliberately not part of this submission
61
-
62
- These are not launch blockers and should not be rebuilt before submission:
63
-
64
- - A dashboard or second frontend surface.
65
- - Hosted accounts, GitHub OAuth, cloud sync, analytics, or enterprise administration.
66
- - Automatic installation of newly discovered projects or arbitrary third-party
67
- installer execution.
68
- - A universal quality score, fabricated benchmark results, or automatic claims that
69
- stars prove safety or usefulness.
70
- - A production-hosted signed intelligence service, a second executable runtime tool,
71
- speculative adapter features without user demand, or ten external user studies.
72
- - Restoring GitHub-hosted Actions capacity before submission; the full gate can run
73
- locally, and the latest completed `main` CI is already passing.
74
-
75
- ## Archived implementation diary (historical, not active scope)
76
-
77
- Everything below this heading is a frozen chronological implementation record. Its
78
- dashboard references, old release numbers, unchecked boxes, and superseded acceptance
79
- paths describe earlier product states; they are neither current capabilities nor
80
- active launch tasks. Use only the launch finish line above for current status.
81
-
82
- ### Implemented product work
83
-
84
- - [x] `P18-01 [TERRA]` Make the first CLI screen beginner-readable: a read-only
85
- `loadout guide`, concise default library summary, focused first-screen help,
86
- and retained access to every advanced command through `loadout advanced` or
87
- `<command> --help`.
88
- - [x] `P18-02 [TERRA]` Make core machine output honest: `catalog --json` returns
89
- JSON, and `update --package <id>` checks only that package rather than every
90
- tracked installation.
91
- - [x] `P18-04 [TERRA]` Bound and explain live network checks in `health --updates`
92
- and `watch --once`; they must show progress, a clear timeout, and an actionable
93
- result even when a large Maximum library is present.
94
- - [x] `P18-05 [TERRA]` Make Agent Health distinguish active skills from disabled
95
- Maximum-library copies, so a broad download is not presented as a broken active
96
- configuration.
97
- - [x] `P18-06 [TERRA+LUNA]` Audit the local dashboard with real founder state at
98
- desktop and mobile widths. Keep it optional, reduce jargon, and only add UI
99
- actions that retain preview, explicit acknowledgement, snapshot, and rollback.
100
- - [x] `P18-07 [TERRA]` Add a concise beginner section to README that links to the
101
- testing guide and explains Stable, Power, Maximum, project recommendations,
102
- daily discovery, and rollback in plain language.
103
- - [x] `P18-15 [TERRA]` Bind material README claims to checked repository evidence,
104
- add generated facts and adapter lifecycle coverage, test the documented product
105
- flow, and separate deterministic, live, and human evidence.
106
- - [x] `P18-16 [TERRA]` Make risky setup previews include every required approval
107
- flag, reject unknown top-level commands, and make user-requested rollback fail
108
- closed when files changed after a mutation or when a legacy snapshot lacks
109
- post-mutation evidence. Internal failed-transaction recovery remains
110
- authoritative.
111
- - [x] `P18-19 [SOL+TERRA]` Re-audit the final unmerged development branch before deletion and preserve its
112
- one valuable adoption-safety intent in the current architecture. Adoption now
113
- binds the complete safe tree, rejects drift and special entries, records only
114
- final verified bytes, conservatively attributes catalog review, deep-freezes
115
- previews, and accepts only the exact same-process planner-issued plan
116
- (`186daa0`, `09e0e0c`, `3f2cafe`, `10d6109`). No valuable work remains on the
117
- retired branch; its local and remote copies were deleted after synchronization while
118
- the user's ignored `.superpowers/` directory was preserved.
119
- - [x] `P18-20 [SOL+TERRA]` Fix the fresh-clone live Stable rollback regression
120
- without weakening drift protection. `5f8e38e` records installed-profile state
121
- inside the catalog transaction before post-mutation evidence and makes the live
122
- evidence journey roll Stable back before unrelated fixture transactions. An
123
- unsafe proposal to ignore `state.json` drift was rejected because it could
124
- orphan later installations.
125
-
126
- All intended work through P18-20 is on remote `main`, and there are no open PRs. The
127
- merged `codex/relatable-readme-hero` remote branch remains safe to delete after the
128
- 0.4.0 release checkpoint.
129
-
130
- ### Deterministic repository verification
131
-
132
- - [x] `P18-08 [TERRA]` Run the complete clean-state verification gate after plan and
133
- documentation consolidation, including formatting, lint, type checking, build,
134
- evidence checks, unit and integration tests, CLI and README product flows,
135
- package smoke, performance, dashboard Playwright, package contents, and
136
- `git diff --check`. The July 20 integrated-main audit passed 113 test files with
137
- 597 tests passing and one intentionally skipped, both product flows, package
138
- smoke, the 1,000-skill performance gate at 2.40 seconds p95, and both dashboard
139
- viewports.
140
- The additional live Stable gate installed four pinned packages and completed
141
- state and filesystem rollback assertions at `5f8e38e`.
142
-
143
- ### Branch cleanup status
144
-
145
- - [x] Review all unique history on the final unmerged development branch before
146
- deletion; the adoption gap above was the only remaining valuable intent.
147
- - [x] Reimplement and regression-test that intent in the current architecture.
148
- - [x] Complete the second remote synchronization, confirm the four adoption commits
149
- are reachable from remote `main`, and delete the superseded development
150
- branches. The temporary worktree was removed and the unrelated ignored
151
- `.superpowers/` directory was preserved. The later merged README hero branch is
152
- retained only until the 0.4.0 release checkpoint.
153
-
154
- ### Human and external work (not provable by local tests)
155
-
156
- - [ ] `P18-03 [HUMAN+TERRA]` Run the founder acceptance path in
157
- `docs/USER_TEST_GUIDE.md` on real Codex and Claude profiles. Record each
158
- observed failure and turn reproducible ones into regression tests.
159
- - [ ] `P18-13 [HUMAN+TERRA]` Publish `loadout-ai@0.4.0` to npm, then test that exact
160
- registry tarball in a fresh terminal through Stable -> rollback -> Power ->
161
- rollback -> Maximum -> project optimization -> complete uninstall.
162
- Publication completed on July 20. A clean temporary install resolved the exact
163
- registry tarball, reported version `0.4.0`, and completed `loadout demo` with
164
- rollback verification. The founder's real Codex/Claude CLI lifecycle test
165
- remains required before this item can be checked; dashboard testing was removed
166
- after founder review rejected it as a conflicting product surface.
167
- - [ ] `P18-17 [HUMAN]` Resolve the GitHub account billing or spending-limit condition
168
- that prevents Actions jobs from starting, rerun CI and daily discovery on the
169
- integrated commit, and record the result. CI runs
170
- [`29691581581`](https://github.com/VirajMishra1/loadout/actions/runs/29691581581)
171
- (`4fdc473`) and [`29692535521`](https://github.com/VirajMishra1/loadout/actions/runs/29692535521)
172
- (`e74ba16`) failed before any step ran; their annotations cite failed recent
173
- account payments or a spending limit that must be increased. They are external
174
- runner-account failures, not product-test failures.
175
- - [ ] `P18-18 [HUMAN]` Decide and configure appropriate `main` branch protection.
176
- GitHub currently returns 404 for the protection endpoint, so protection is
177
- absent or not observable; do not describe it as enabled.
178
-
179
- ### Founder acceptance findings and 0.4.1 corrections
180
-
181
- - [x] `P18-21A [HUMAN+TERRA]` Verify the published 0.4.0 Stable lifecycle on the
182
- founder's real Claude Code and Codex paths. Stable installed four pinned sources
183
- as 60 managed activations, preserved all 12 unmanaged Claude skills, reported no
184
- managed drift, and restored the explicit pre-install snapshot successfully.
185
- - [x] `P18-21B [TERRA]` Make rollback history understandable. The first
186
- founder test exposed that bare `loadout rollback` selected a newer no-op
187
- dashboard/sync snapshot whose pre-state already contained Stable. No data was
188
- lost, but `rollback --list` showed opaque IDs only and `Restored snapshot` did
189
- not disclose that zero effective files changed. Rollback history now displays
190
- timestamp, mutation label, affected roots, effective entry count, latest marker,
191
- and explicit no-op guidance. New mutations carry user-facing labels. A CLI
192
- regression journey proves adjacent no-op and install snapshots plus explicit
193
- older-snapshot restoration.
194
- - [x] `P18-22 [TERRA]` Remove the dashboard before the public release. Founder review
195
- confirmed that it presents recommendation presets (`stable`, `web`,
196
- `collaboration`, `maximum`) as policy profiles while the real CLI contract is
197
- `stable`, `power`, `maximum`, and `custom`; its manifest-sync mutation model is
198
- not the CLI setup workflow. Remove the command, loopback server, browser assets,
199
- dashboard-only dependencies/tests/docs/evidence, and npm package contents.
200
- The CLI retains all useful inspection, recommendation, configuration, health,
201
- rollback, and uninstall capabilities; there is no competing dashboard workflow.
202
- - [ ] `P18-23 [HUMAN+TERRA]` Complete the remaining CLI-only founder path on the exact
203
- npm package. Power and its explicit rollback are complete: snapshot
204
- `1784546929191-9cfb0705dbed` restored zero Loadout-managed activations, retained
205
- all 12 unmanaged Claude skills, returned Codex to zero skills, and left no
206
- duplicate groups. Maximum setup is also complete: snapshot
207
- `1784547473319-7581eb19ee9c` stored 2,316 screened skill copies from 29 packages
208
- in the disabled library, activated none of them, and again preserved the 12
209
- unmanaged Claude skills. Remaining path: project activation -> recommendation/
210
- optimization -> Graphify install/remove -> credential-free MCP inventory ->
211
- read-only update/discovery -> complete uninstall -> reinstall. Record every
212
- mutation snapshot ID and use explicit rollback IDs during acceptance.
213
- - [x] `P18-24 [SOL+TERRA]` Make Power match the intended product hierarchy based on
214
- the real founder run.
215
- The published transaction itself passed: eight immutable sources installed for
216
- Claude Code and Codex, 100 managed activations were created, 12 unmanaged Claude
217
- skills were preserved, duplicate targets were resolved, six rejected skill
218
- units were excluded, and 1,170 managed files reported zero drift. The profile
219
- policy failed acceptance: it activates 50 skills per agent despite Loadout's
220
- recommended limit of 30; every selected package retained blocking static-risk
221
- findings (79 total); only five of eight sources have an asserted SPDX license;
222
- and one coarse prompt approves all package findings. The founder clarified that
223
- Power is intentionally the larger active mode, not another 30-skill Stable.
224
- Stable remains bounded at 30; Power explicitly warns about its larger context
225
- footprint; Maximum stores the broadest screened library disabled and activates
226
- only project-relevant subsets. Invalid units remain quarantined and detailed
227
- findings stay available behind `--details`. Existing transaction and founder
228
- evidence proves unmanaged-skill preservation and exact rollback.
229
- - [x] `P18-25 [LUNA+TERRA]` Make health output honest and focused for empty or unmanaged
230
- profiles. After the successful Power rollback, `loadout health --explain` led
231
- with `Loadout health: healthy` and then reported 28/100 critical evidence for
232
- Claude Code, 0/100 unknown for Codex, and similarly verbose sections for every
233
- detected agent. Empty managed state now reports `not configured`, the default
234
- remains concise, and `--explain --agents <ids>` keeps deeper evidence scoped.
235
- Health also understands managed Codex TOML MCP entries instead of falsely
236
- reporting JSON drift.
237
- - [x] `P18-26 [LUNA+TERRA]` Make Maximum's preview understandable without weakening
238
- its safe defaults. The founder run proved the disabled-library contract, but
239
- printed 50 unit-level quarantine blocks, 19 expected MCP/runtime deferrals, two
240
- `No SKILL.md` preparation failures, and one coarse 28-package risk approval in
241
- the default path. The subsequent scan reported 11 repository failures without
242
- naming them. Default to a concise grouped summary with exact counts, severity,
243
- source, and next commands; retain complete findings behind an explicit details
244
- or JSON view. MCP/runtime-only records are explicit deferred setup rather than
245
- failed skill preparation; actual repository preparation failures remain named
246
- with their reason. The default preview groups quarantine and deferral counts,
247
- while `--details` shows every unit. Regression tests cover both views.
248
- - [x] `P18-27 [SOL+TERRA, LUNA copy review]` Correct project-aware activation safety
249
- and relevance using
250
- `docs/superpowers/specs/2026-07-20-project-activation-safety-design.md`.
251
- Founder preview proved that activation currently ignores 12 unmanaged Claude
252
- skills when calculating `--limit 30`, treats rollback-restored empty directories
253
- as occupied for both agents, and proposes a generic, redundant 30-skill set from
254
- only JavaScript/TypeScript and Playwright signals. Implement the approved shared
255
- empty-target predicate, per-agent total active capacity, bounded Node CLI/npm/
256
- Vitest/MCP project signals, diverse evidence-threshold selection, integration-
257
- type labels, atomic apply revalidation, and the specified regression suite.
258
- Implementation is complete on `codex/project-activation-safety`: the full local
259
- release gate passes, including 114 unit-test files with 605 passing tests and
260
- one intentional skip; the disposable two-agent CLI journey proves 12 unmanaged
261
- Claude skills, distinct Claude/Codex budgets, empty rollback residue, one atomic
262
- apply, and explicit rollback; the packed artifact is 468.0 kB with 145 files;
263
- and the fresh scan benchmark passed at 1,583.5 ms p95 across seven real CLI
264
- runs. Do
265
- not resume real-profile activation until this branch is reviewed, integrated,
266
- versioned, and published as the corrected release.
267
- - [x] `P18-28 [HUMAN+TERRA]` Merge the project-activation corrections through PR #4
268
- and publish the exact verified release as `loadout-ai@0.4.1`. Registry
269
- verification returned version `0.4.1` and integrity
270
- `sha512-k8WTNh6kTIaaBFTPGsl/QD/7/LQ1Gg9uMReM87gAkBPdcG0IxanM9NznyjemsYa124yYPczAY5MVYan4i91MtA==`;
271
- a clean temporary global install reported `0.4.1` and detected TypeScript,
272
- Playwright, Node CLI, npm package, release, MCP, security, Commander, Zod, and
273
- Vitest signals from this repository. Tag `v0.4.1` points to release commit
274
- `0d25b8e`. This is the corrected founder-testing release, not the final public
275
- launch candidate: dashboard removal, Power policy, health output, and Maximum
276
- preview work remain open.
277
- - [x] `P18-29 [SOL+TERRA]` Block ecosystem-mismatched project activation candidates
278
- before scoring. The exact published `0.4.1` founder preview correctly detected
279
- this TypeScript Node CLI, respected both agents' real capacity, and produced no
280
- false occupied-target blockers, but still proposed `mcp-csharp-publish`,
281
- `mcp-csharp-test`, `uv-package-manager`, `social-publishing`, and
282
- `vercel-cli-with-tokens`, then exposed `msstore-cli`, `phoenix-cli`, and
283
- `publish-to-pages` once those higher-ranked mismatches were removed. A second
284
- preview exposed the general cause: any domain-specific name ending in `-cli`,
285
- including `datadog-cli`, inherited the full Node CLI and Commander score; a
286
- third exposed substring matching of `npm` inside `pnpm`, plus backend and web
287
- design guidance without corresponding project roles; a fourth caught generic
288
- `schema` admitting database design for Zod and universal accessibility guidance
289
- in a CLI project. Add deterministic language/provider/specialization
290
- compatibility gates, bounded generic CLI and schema evidence, token-aware
291
- package-manager matching, and role-gated backend/frontend guidance. Generic
292
- `mcp`, `cli`, `package`, and `publish` words must not override compatibility;
293
- preserve explicit pins as a deliberate escape hatch; prove the exact regression
294
- with tests and repeat the read-only founder preview before any real activation.
295
- Complete on `codex/activation-compatibility-gates`: the exact live Maximum
296
- library preview now proposes 24 relevant Codex skills and 18 capacity-bounded
297
- Claude Code skills with zero occupied-target blockers and none of the observed
298
- ecosystem mismatches. The full release gate passes with 114 test files, 607
299
- passing tests, one intentional skip, both CLI product journeys, packaged CLI
300
- smoke, evidence checks, and a 1,518.4 ms p95 scan benchmark across seven real
301
- runs. The preview remained read-only; publish the merged correction before the
302
- founder applies it to real agent profiles.
303
-
304
- ### Product-first release candidate work
305
-
306
- - [x] `P18-30 [TERRA]` Add `kepano/obsidian-skills` at an immutable MIT-licensed
307
- revision as the 51st credited catalog source. Detect `.obsidian` vaults and
308
- recommend/activate the five Obsidian-oriented skills only for relevant projects;
309
- do not burden universal Stable with a niche tool.
310
- - [x] `P18-31 [TERRA]` Make reviewed MCP recipes usable from the normal CLI for both
311
- Codex and Claude Code. `mcp-recipe --agent <host>` now chooses the real host
312
- config path and format, previews before writing, stores managed fingerprints,
313
- verifies presence, reports drift in health, and removes only Loadout's entry.
314
- Config plus ownership state commit in one rollback transaction, so rollback
315
- cannot leave a stale managed MCP record. Disposable end-to-end tests cover
316
- preview, apply, verify, health, rollback, reapply, and removal.
317
- - [x] `P18-32 [TERRA]` Present evidence maturity without the misleading exclusive
318
- `0 discovered` headline. The README now reports 51 sourced/inspected records,
319
- four Stable sources, and the dated discovery feed separately. It clearly says
320
- independent human-review and comparative benchmark publications are future
321
- promotion stages rather than implying the catalog is unusable.
322
- - [ ] `P18-33 [HUMAN+TERRA]` Publish the next verified CLI-only release and run the
323
- founder path from that exact npm tarball: Stable, Power, Maximum/project
324
- activation, Obsidian recommendation, Graphify, both-host credential-free MCP,
325
- read-only update/discovery, rollback history, uninstall, and clean reinstall.
326
- - [ ] `P18-34 [HUMAN]` Resolve or bypass exhausted GitHub-hosted Actions minutes by
327
- running the documented complete gate locally now; later restore hosted CI with
328
- billing/minutes, a self-hosted runner, or a teammate-owned fork. Never label an
329
- unstarted hosted job as a code failure.
330
- - [x] `P18-35 [TERRA]` Run the complete CLI-only release gate locally after all
331
- product-first changes against the exact `0.5.0` package. On July 21 it passed formatting, lint, type checking,
332
- catalog/discovery/README/release evidence, 113 test files with 603 passing and
333
- one intentional skip, both CLI product journeys, packed npm smoke, and seven
334
- real 1,000-skill scans at 1,468.2 ms p95.
335
- - [x] `P18-36 [TERRA]` Close the stale-build packaging gap caught by the first npm
336
- publish attempt. Every build now removes `dist` first on every platform, and the
337
- package smoke test rejects removed dashboard/demo JavaScript if it ever leaks
338
- back into the tarball. The failed OTP-gated attempt published nothing.
339
- - [x] `P18-37 [TERRA]` Fix the founder-discovered disabled-library uninstall bug.
340
- Removal now resolves each managed file to its disabled Maximum-library copy,
341
- never an unmanaged skill that later occupies the original active path. A
342
- regression test preserves the replacement bytes, the real founder state now
343
- previews an unblocked removal of 29 packages/2,316 disabled records, and health
344
- reports `library ready (nothing active)` with explicit counts. Release as
345
- `0.5.1` before continuing the complete-uninstall acceptance step.
346
- - [x] `P18-38 [TERRA]` Fix the founder-discovered asymmetric project-activation
347
- scope bug. The `0.5.1` preview correctly budgeted 22 Codex and 18 Claude
348
- additions, but apply-time revalidation broadened four Codex-only selectors
349
- onto Claude. Revalidation now preserves the exact package, skill, and agent
350
- tuple from the preview; a two-agent regression proves exact 30/18 applied
351
- counts. CLI-only TypeScript projects also reject browser-testing and
352
- Playwright false positives, and the default preview removes unexplained scores
353
- and duplicated delta output. Release as `0.5.2` before resuming activation.
354
- - [x] `P18-39 [TERRA]` Integrate reviewed runtime tools into the same inventory and
355
- health truth as skill repositories. Founder testing proved Graphify installed
356
- correctly for Claude Code and Codex, but the scanner counted Claude's target as
357
- unmanaged and omitted Codex's official `~/.codex/skills` target because its
358
- standard collection root is `~/.agents/skills`. Scan every registered runtime
359
- target, attribute it to `runtime-tool:<id>`, count both host targets in health,
360
- surface runtime/MCP totals and missing targets, label runtime snapshots, and
361
- raise the default cold MCP handshake window to the tested 30 seconds. Release
362
- as `0.5.3` before continuing Graphify removal and discovery/update acceptance.
363
- - [x] `P18-40 [TERRA]` Scan both Codex's canonical and current compatibility skill
364
- roots, exclude host-bundled `.system` entries, and count runtime-tool targets
365
- once. The real founder preview now reports 47 visible skills: 13 Claude, 12
366
- Codex, 11 Cursor, and 11 Windsurf.
367
- - [x] `P18-41 [TERRA]` Credit and pin the official Apache-2.0 Cloudflare Skills
368
- collection and MIT Humanizer repository as catalog sources. Index safe siblings
369
- independently so one quarantined unit cannot hide every valid skill in a large
370
- collection.
371
- - [x] `P18-42 [TERRA]` Add grouped existing-skill reconciliation. Exact complete-tree
372
- matches can be adopted without rewriting bytes; Git origins disambiguate
373
- same-name candidates; outdated replacement remains explicit, risk-gated,
374
- byte-verified, and one rollback-safe transaction.
375
- - [x] `P18-43 [TERRA]` Preserve an adopted skill's exact unit and existing host path
376
- during future updates, including Codex compatibility-root installations. The
377
- real read-only preview found eight exact groups, two unambiguous old groups,
378
- and left two quarantined/unknown groups untouched. The complete local release
379
- gate passes 114 test files with 616 passing tests and one intentional skip,
380
- both CLI product journeys, package smoke, and a 1,000-skill scan at 1,422.7 ms
381
- p95 across seven real CLI runs.
382
- - [ ] `P18-44 [HUMAN+TERRA]` Publish and install `loadout-ai@0.5.4`, run
383
- `loadout reconcile --refresh`, adopt the eight exact groups, inspect the
384
- Humanizer and Turnstile update findings before any replacement, then finish
385
- read-only update/discovery, Graphify removal, complete uninstall, and clean
386
- reinstall acceptance.
387
- - [x] `P18-45 [TERRA]` Correct the founder-discovered update-reporting defects after
388
- the real `0.5.4` Maximum/reconciliation run. Update planning now deduplicates
389
- repeated agent names and shared-repository fetches, separates active updates
390
- from newer commits held for disabled library copies, compacts safety output,
391
- recognizes adopted sources in profile state, and labels partial quarantine
392
- evidence instead of calling it a repository failure. The real profile confirms
393
- 39 active managed skills, 2,326 disabled copies, one Graphify runtime, zero
394
- managed drift, and no active adopted-skill update. Release this correction
395
- before the remaining Graphify removal, discovery, uninstall, and reinstall
396
- acceptance steps.
397
- - [x] `P18-46 [TERRA]` Eliminate the repeated large-repository timeout exposed by the
398
- published `0.5.5` founder check. Resolve remote HEAD metadata before fetching
399
- content, download only changed revisions that require static safety analysis,
400
- reuse immutable cached snapshots, and allow those bounded reviews 120 seconds.
401
- The real 40-record Maximum state now completes with 19 disabled-library
402
- changes, 21 current packages, zero active updates, and zero unavailable checks;
403
- both previously failing large repositories complete without changing or
404
- activating any agent-visible skill.
405
-
406
- ### Release 0.3 lifecycle hardening
407
-
408
- - [x] `P18-09 [SOL+TERRA]` Add preview-first complete uninstall with modified-file
409
- protection, native-job cleanup, runtime restoration, guarded state deletion,
410
- and optional global npm removal.
411
- - [x] `P18-10 [TERRA]` Persist the installed profile and make `loadout update`
412
- evaluate both profile drift and every managed repository.
413
- - [x] `P18-11 [SOL+TERRA]` Add explicit bulk safe updates while holding disabled,
414
- risky, and failed packages for review; scheduled checks remain read-only.
415
- - [x] `P18-12 [TERRA]` Add a pinned Chrome DevTools MCP recipe and distinguish
416
- separately billed AI/model API keys from unrelated service credentials. The
417
- no-model-key inventory includes GitHub read-only with its token disclosed.
418
- - [x] `P18-14 [TERRA]` Treat recursively empty directories from legacy cleanup as
419
- unoccupied without weakening unmanaged-file protection, and make complete
420
- uninstall remove empty nested directory shells.
421
-
422
- ### Explicitly deferred (do not expand during this usability pass)
423
-
424
- - Hosted accounts, GitHub OAuth, cloud sync, analytics, and enterprise policy.
425
- - Required API keys, automatic provider spending, or treating chat subscriptions as
426
- API access.
427
- - Automatically installing newly discovered repositories or executing arbitrary
428
- third-party installers.
429
- - A universal quality score, social-network scraping, or support claims for an agent
430
- that has not passed a real adapter test.
431
-
432
- ## Archived implementation history
433
-
434
- Sections 1–20 below preserve the original product design, allocation, completed work,
435
- and deferred exploration. They are historical context. Only the current-status section
436
- above defines active work.
437
-
438
- ## 1. Executive summary
439
-
440
- Loadout is a universal extension manager for AI coding agents. It detects the agents
441
- installed on a user's computer, discovers trusted skills and MCP tools from official
442
- catalogs and high-signal GitHub repositories, installs them in the correct format,
443
- keeps configurations synchronized, checks for updates, and restores a known-good
444
- snapshot if an update fails.
445
-
446
- The product is intentionally consumer-first and CLI-first. A user should not need to
447
- understand `SKILL.md`, MCP configuration, plugin manifests, platform-specific
448
- directories, or GitHub repository layouts. Interactive setup starts with one command:
449
-
450
- ```bash
451
- npx loadout-ai
452
- ```
453
-
454
- ```text
455
- Choose a loadout: [1] Stable Boost (recommended), [2] Maximum Library, [3] Custom
456
- ```
457
-
458
- The executable installed by the package is still named `loadout`. The `loadout` npm
459
- package name belongs to an unrelated project, so the publishable package is
460
- `loadout-ai`.
461
-
462
- The hackathon MVP proves the safe package lifecycle with a curated catalog. The
463
- product-defining loop is broader: scan what the user already has -> discover and
464
- review candidates -> compare evidence -> recommend a small active set -> install or
465
- activate -> verify -> update -> block an unsafe or incompatible update -> rollback.
466
-
467
- ## 2. Product thesis
468
-
469
- Developers increasingly use multiple AI coding agents, but the ecosystem of skills,
470
- plugins, agents, rules, and MCP tools is fragmented across GitHub, official
471
- marketplaces, social media, and independent registries. Existing package managers
472
- focus on files and packages. Loadout focuses on the outcome a user wants: make every
473
- installed agent more capable without requiring manual discovery or configuration.
474
-
475
- Loadout wins through:
476
-
477
- 1. A one-command consumer experience.
478
- 2. Broad agent and operating-system support.
479
- 3. A maintained Stable and Trending catalog.
480
- 4. A large reviewed library plus a small conflict-aware active set instead of blindly
481
- exposing every downloaded package to every agent.
482
- 5. Human-readable update and permission diffs.
483
- 6. Snapshots and rollback.
484
- 7. Optional project-aware recommendations without requiring GitHub access.
485
-
486
- ## 3. Product and delivery scope
487
-
488
- ### 3.1 Submission-critical vertical slice
489
-
490
- - TypeScript CLI packaged for `npx loadout-ai` after an owner publishes it to npm.
491
- - CLI-first interactive and non-interactive Stable/Maximum/Custom setup.
492
- - CLI-only primary experience. The existing framework-free loopback dashboard is a
493
- secondary diagnostic surface and is not required for onboarding or daily use.
494
- - Windows 11, macOS, and Linux support.
495
- - Agent detection for:
496
- - Claude Code
497
- - Codex
498
- - Cursor
499
- - Gemini CLI
500
- - OpenCode
501
- - Hermes
502
- - Windsurf
503
- - Cline
504
- - GitHub Copilot
505
- - Roo Code
506
- - Kiro CLI
507
- - Junie
508
- - Skill installation for all twelve agents where their documented layout is known.
509
- - MCP configuration for Claude Code, Codex, and Cursor.
510
- - Curated catalog containing 50 pinned real repositories.
511
- - Stable, Trending, Official, and Community tier support; the current bundled review
512
- set contains Official and Stable records.
513
- - Stable Boost, Maximum Boost, and Custom modes.
514
- - Immutable package records pinned by Git commit SHA.
515
- - Snapshot before the first mutation.
516
- - Update detection and human-readable diff.
517
- - Block at least one incompatible or risky update in the demo.
518
- - One-command rollback.
519
- - No GitHub login required for the core experience.
520
- - Optional local-folder scan for project-aware recommendations.
521
- - Clear display of `native`, `adapted`, and `unsupported` components.
522
-
523
- ### 3.2 Deferred product exploration (not current launch scope)
524
-
525
- These are ideas worth revisiting after the core CLI has passed real founder and
526
- external user testing. They are not commitments for this hackathon release and must
527
- not be presented as production-ready simply because a prototype or command exists.
528
-
529
- - GitHub OAuth for private repositories and personalized discovery, using minimal
530
- read-only scopes by default.
531
- - Community Loadout publishing, sharing, importing, versioning, and reporting.
532
- - Historical star, fork, contributor, release, and download velocity charts.
533
- - Model/provider configuration and comparison, including OpenRouter.
534
- - Automated category-specific evaluations with repeatable fixtures and confidence
535
- information.
536
- - Background catalog and update notifications.
537
- - Signed catalog snapshots and signature verification in the client.
538
- - Best-effort compilation of hooks, commands, agents, and subagents between platforms,
539
- with explicit loss reports instead of false compatibility claims.
540
- - Sandboxed execution for third-party installers that genuinely require execution,
541
- with no host credentials and no automatic promotion from the sandbox.
542
- - A user-controlled encrypted credential vault backed by the operating-system keychain;
543
- the service must not store plaintext user secrets.
544
- - Policy-gated autonomous updates for MCP servers, hooks, and executables after
545
- sandbox tests, permission comparison, and rollback preparation.
546
- - An adapter SDK and community adapter registry for broad agent support.
547
- - Category-specific capability scoring and comparison; never one misleading universal
548
- number claiming scientific certainty across unrelated tasks.
549
- - Discovery connectors for major social and community sources where their APIs and
550
- terms permit access.
551
- - Team and enterprise policy administration, including allowlists, denylists, required
552
- versions, audit history, and shared Loadouts.
553
-
554
- ### 3.3 Non-negotiable safety boundaries
555
-
556
- The ambitious scope does not authorize unsafe shortcuts:
557
-
558
- - Never execute untrusted installation scripts directly on the host during discovery.
559
- - Never store plaintext secrets in the repository, catalog, logs, analytics, or hosted
560
- database.
561
- - Never claim perfect conversion when platform semantics differ; show a loss report.
562
- - Never silently grant new filesystem, network, account, hook, or executable powers.
563
- - Never market a category score as universal scientific truth.
564
- - Never scrape a source in violation of its API rules, robots policy, or terms.
565
- - Never claim support for an agent until its adapter passes the published conformance
566
- suite.
567
-
568
- ## 4. Primary users
569
-
570
- ### 4.1 Multifunctional power user
571
-
572
- Uses Claude, Codex, Cursor, or other agents for many kinds of work and wants the
573
- largest useful capability set without manually visiting repositories.
574
-
575
- Default path: audit the existing setup, retain a broad reviewed library, and activate
576
- only the best evidence-backed global and project-specific subset. Maximum Library is
577
- explicit stress/power-user mode, not the default active set.
578
-
579
- ### 4.2 New agent user
580
-
581
- Has installed one or more agents but does not understand extension formats or MCP.
582
-
583
- Default path: Stable Boost.
584
-
585
- ### 4.3 Project-focused developer
586
-
587
- Wants recommendations for a particular local repository.
588
-
589
- Default path: Stable or Maximum Boost plus optional local-folder analysis.
590
-
591
- ### 4.4 Team
592
-
593
- Wants a reproducible configuration shared through source control.
594
-
595
- Post-MVP path: commit `loadout.lock` and restore it on another machine.
596
-
597
- ## 5. User experience
598
-
599
- ### 5.1 First run
600
-
601
- 1. User runs `npx loadout-ai` after publication, or `npx .` from a clone.
602
- 2. Loadout detects supported installed agents and read-only scans their existing
603
- skills, separating Loadout-managed content from unmanaged content without assuming
604
- unmanaged means unsafe.
605
- 3. It recommends Stable, Maximum Library, Custom, or an evidence-backed optimization
606
- of the existing setup. It filters out components that require explicit
607
- credentials/configuration, then
608
- concurrently fetches only reviewed skill repositories at their pinned commits.
609
- 4. It resolves overlapping skill targets deterministically, keeps the higher-ranked
610
- reviewed source, and reports every deferred duplicate.
611
- 5. It shows repository counts, actual skill-directory counts, safety findings,
612
- deferred MCP/executable packages, and detected targets before mutation.
613
- 6. The user confirms the loadout and separately approves script/domain/instruction
614
- findings when present.
615
- 7. Loadout snapshots all targets and installs the entire loadout as one durable,
616
- rollback-safe transaction; caught or interrupted failures restore prior state.
617
- 8. Daily use continues through `scan`, `status`, fast local `health`, `update`,
618
- `discover`, `recommend`, `compare`, `optimize`, `remove`, and `rollback`. Commands
619
- not yet implemented remain Phase 12 backlog items. The dashboard remains optional.
620
-
621
- Account capability rule: interactive setup asks whether the user has separately billed
622
- OpenAI API, Anthropic API, OpenRouter, other provider access, or none. ChatGPT and
623
- Claude subscriptions do not count. The answer contains provider names only, is not a
624
- credential, is not persisted by setup, and cannot weaken safety policy. Static skills
625
- remain available without model API access; credentialed MCP/runtime/model operations
626
- remain explicit and fail closed until a named environment or OS-keychain reference is
627
- appropriate for that execution boundary.
628
-
629
- ### 5.2 Normal use
630
-
631
- - `loadout status`: agents, packages, conflicts, and update health.
632
- - `loadout scan`: read-only inventory of existing skills, ownership, fingerprints,
633
- duplicates, and capacity warnings.
634
- - `loadout setup --mode stable`: preview the small reviewed daily-use foundation.
635
- - `loadout setup --mode maximum`: preview the broad reviewed loadout.
636
- - `loadout setup --mode maximum --yes --approve-risk`: install it non-interactively
637
- after review.
638
- - `loadout add <package>`: plan and add a package.
639
- - `loadout remove <package>`: remove only files managed by Loadout.
640
- - `loadout update`: fetch package update information and display a read-only plan by
641
- default.
642
- - `loadout rollback`: restore the previous snapshot.
643
- - `loadout doctor`: validate configurations and dependencies.
644
- - `loadout dashboard`: optional secondary diagnostic surface; never required by the
645
- CLI-first journey.
646
-
647
- ### 5.3 No-account guarantee
648
-
649
- The following must work without signup or GitHub OAuth:
650
-
651
- - Agent detection
652
- - Catalog browsing
653
- - Stable and Maximum Boost
654
- - Public package installation
655
- - Updates
656
- - Local snapshots
657
- - Rollback
658
-
659
- GitHub access is optional and used only for private repositories, GitHub operations,
660
- or personalized project discovery.
661
-
662
- ## 6. Catalog policy
663
-
664
- ### 6.1 Admission tiers
665
-
666
- #### Official
667
-
668
- Accepted without a star minimum when publisher identity is verifiable and the source
669
- is an official vendor or standards organization.
670
-
671
- #### Stable
672
-
673
- Default discovery has no star floor. Stable normally requires:
674
-
675
- - At least 1,000 GitHub stars, or a documented exception based on verified publisher,
676
- package adoption, maintainer reputation, or independent evaluation.
677
- - Clear installable component.
678
- - Non-archived repository.
679
- - Recent meaningful maintenance.
680
- - License metadata present or explicitly reviewed.
681
- - Supported source can be pinned to an immutable commit.
682
- - Basic security and compatibility checks pass.
683
-
684
- Packages above 5,000 stars receive a `Popular` signal, not automatic trust or an
685
- exclusive right to enter the catalog.
686
-
687
- #### Trending
688
-
689
- - Normally at least 100 stars, or an explicitly approved exception for an official or
690
- independently verified release.
691
- - Strong recent star velocity or adoption signal.
692
- - Active maintenance.
693
- - Basic safety and compatibility checks pass.
694
- - Never enabled silently in Stable mode.
695
-
696
- #### Community
697
-
698
- - Any star count, including zero-star newly published packages.
699
- - Searchable or manually installable.
700
- - Requires explicit user selection.
701
-
702
- ### 6.2 Discovery sources
703
-
704
- MVP:
705
-
706
- - Curated seed list in the repository.
707
- - GitHub Search API for known filenames and topics.
708
- - OpenAI skills catalog.
709
- - Anthropic official plugin marketplace.
710
- - Official MCP Registry.
711
- - skills.sh metadata where permitted.
712
-
713
- Full-product ingestion:
714
-
715
- - GitHub star snapshots and acceleration.
716
- - GitHub release feeds.
717
- - npm and PyPI download/release signals.
718
- - Hacker News API.
719
- - Reddit and other community sources where API terms permit.
720
-
721
- ### 6.3 Search signatures
722
-
723
- - `SKILL.md`
724
- - `.claude-plugin/plugin.json`
725
- - `.codex-plugin/plugin.json`
726
- - `.mcp.json`
727
- - `mcp.json`
728
- - `topic:agent-skills`
729
- - `topic:claude-code`
730
- - `topic:codex`
731
- - `topic:mcp-server`
732
-
733
- ### 6.4 Ranking
734
-
735
- Do not compare packages from unrelated categories. Rank within a capability category.
736
-
737
- Initial score:
738
-
739
- - 30% community adoption: logarithmic stars, forks, contributors.
740
- - 20% momentum: recent growth and releases.
741
- - 20% maintenance: meaningful recency, responsiveness, multiple maintainers.
742
- - 15% compatibility: agents, operating systems, clean install result.
743
- - 15% trust: official identity, license, pinned dependencies, absence of risky patterns.
744
-
745
- The score is a recommendation aid, not a claim of objective superiority.
746
-
747
- ## 7. Default catalog categories
748
-
749
- - Engineering workflow
750
- - Documentation retrieval
751
- - Codebase intelligence
752
- - Browser automation and verification
753
- - Frontend design
754
- - Context and token optimization
755
- - Memory
756
- - Source-control integrations
757
- - Security
758
- - Research
759
- - Data and documents
760
- - Product and marketing
761
-
762
- The initial catalog should include representative packages discussed during product
763
- research, such as Superpowers, ECC, Karpathy-inspired guidance, Context7, Graphify,
764
- Serena, Playwright MCP, Chrome DevTools MCP, UI UX Pro Max, Taste Skill, RTK, GitHub
765
- MCP, Planning with Files, and official OpenAI and Anthropic catalogs. Exact inclusion
766
- requires license and install-shape verification.
767
-
768
- ## 8. Conflict policy
769
-
770
- Loadout may download or register many packages but should not activate overlapping
771
- packages blindly.
772
-
773
- Initial conflict families:
774
-
775
- - Major workflow harnesses: Superpowers, ECC, GSD, Compound Engineering.
776
- - Codebase intelligence: Graphify, Serena, Understand Anything, Codebase Memory.
777
- - Browser control: Playwright MCP, Chrome DevTools MCP, Browser MCP.
778
- - Frontend guidance: UI UX Pro Max, Taste Skill, Hallmark.
779
- - Persistent memory: Claude Mem, Beads, Agent Memory, Engram.
780
- - Output compression: RTK, Headroom, Context Mode, Caveman.
781
-
782
- Rules:
783
-
784
- 1. Stable Boost selects at most one primary package in each conflicting family.
785
- 2. Maximum Boost may download all approved candidates but activates one default.
786
- 3. Custom mode may override a soft conflict after a warning.
787
- 4. Hard conflicts block confirmation until one candidate is removed.
788
- 5. Conflict explanations must use plain language.
789
-
790
- ## 9. Technical architecture
791
-
792
- ### 9.1 Implemented repository layout
793
-
794
- ```text
795
- loadout/
796
- ├── src/
797
- │ ├── cli.ts # Commands and packaged executable
798
- │ ├── dashboard.ts # Loopback HTTP server and authenticated API
799
- │ ├── core/ # Catalog, adapters, transactions, policy, registry
800
- │ └── shared/ # Types and runtime schemas
801
- ├── dashboard/ # Dependency-free HTML, CSS, and JavaScript UI
802
- ├── catalog/
803
- │ └── packages.json # Reviewed catalog with immutable source evidence
804
- ├── docs/ # Security, compatibility, and operating policies
805
- ├── tests/ # Unit, integration, fixtures, and Playwright E2E
806
- ├── README.md
807
- ├── MASTER_PLAN.md
808
- └── package.json
809
- ```
810
-
811
- The flat package is deliberate for the hackathon: it avoids workspace build and
812
- publishing complexity while keeping modules separated by responsibility.
813
-
814
- ### 9.2 Implemented stack
815
-
816
- - Node.js 20+
817
- - TypeScript
818
- - npm with a committed lockfile
819
- - Commander for CLI
820
- - Browser-native HTML, CSS, and JavaScript for the dashboard
821
- - Zod for runtime schemas
822
- - Vitest for unit/integration tests
823
- - Playwright for dashboard end-to-end tests
824
- - Conservative append-only Codex TOML support and unrelated-key-preserving JSON writes
825
- - GitHub Actions on Windows, macOS, and Linux
826
-
827
- ### 9.3 Local state
828
-
829
- ```text
830
- ~/.loadout/
831
- ├── cache/<package>/<commit>/
832
- ├── snapshots/<timestamp>/
833
- ├── staging/<transaction-id>/
834
- ├── catalog.json
835
- ├── state.json
836
- └── logs/
837
- ```
838
-
839
- Never store secret values in state, logs, snapshots, or telemetry.
840
-
841
- ## 10. Core data models
842
-
843
- ### 10.1 Catalog package
844
-
845
- ```ts
846
- type CatalogPackage = {
847
- id: string;
848
- displayName: string;
849
- source: { type: "github"; repo: string; ref: string };
850
- tier: "official" | "stable" | "trending" | "community";
851
- category: string;
852
- description: string;
853
- license?: string;
854
- stars?: number;
855
- components: Component[];
856
- platforms: Record<PlatformId, "native" | "adapted" | "unsupported">;
857
- operatingSystems: Array<"windows" | "macos" | "linux">;
858
- permissions: PermissionSummary;
859
- conflicts: string[];
860
- };
861
- ```
862
-
863
- ### 10.2 Installed package
864
-
865
- ```ts
866
- type InstalledPackage = {
867
- id: string;
868
- source: string;
869
- commit: string;
870
- files: Array<{ path: string; sha256: string }>;
871
- installedAt: string;
872
- platforms: PlatformId[];
873
- snapshotId: string;
874
- };
875
- ```
876
-
877
- ### 10.3 Mutation plan
878
-
879
- ```ts
880
- type MutationPlan = {
881
- id: string;
882
- creates: PlannedFile[];
883
- updates: PlannedFile[];
884
- deletes: PlannedFile[];
885
- configChanges: ConfigChange[];
886
- warnings: PlanWarning[];
887
- requiresRestart: PlatformId[];
888
- };
889
- ```
890
-
891
- ## 11. Adapter contract
892
-
893
- Each adapter must implement:
894
-
895
- ```ts
896
- interface AgentAdapter {
897
- id: PlatformId;
898
- detect(): Promise<DetectionResult>;
899
- inspect(): Promise<InstalledComponent[]>;
900
- planInstall(pkg: NormalizedPackage): Promise<AdapterPlan>;
901
- planRemove(pkg: InstalledPackage): Promise<AdapterPlan>;
902
- validate(plan: AdapterPlan): Promise<ValidationResult>;
903
- smokeTest(): Promise<SmokeTestResult>;
904
- }
905
- ```
906
-
907
- Adapters must never directly write during `planInstall`. The core transaction engine
908
- owns all mutations.
909
-
910
- ## 12. Installation transaction
911
-
912
- 1. Resolve package to an immutable commit.
913
- 2. Download without executing repository scripts.
914
- 3. Reject paths escaping the package root.
915
- 4. Parse and normalize supported components.
916
- 5. Ask adapters for mutation plans.
917
- 6. Merge plans and detect collisions.
918
- 7. Display preview.
919
- 8. Snapshot every target file that exists.
920
- 9. Write new files to staging.
921
- 10. Validate staged files and configuration.
922
- 11. Commit changes.
923
- 12. Run smoke tests.
924
- 13. On failure, automatically restore snapshot.
925
- 14. On success, update lockfile and state.
926
-
927
- ## 13. Update and rollback
928
-
929
- ### 13.1 Update
930
-
931
- 1. Fetch catalog update.
932
- 2. Resolve installed package source.
933
- 3. Compare pinned commit with candidate commit.
934
- 4. Download candidate to cache.
935
- 5. Produce file, instruction, command, domain, and permission diff.
936
- 6. Run static checks.
937
- 7. Plan install as a replacement transaction.
938
- 8. Require approval for scripts, hooks, MCP changes, executables, new domains, or new
939
- environment-variable requirements.
940
- 9. Apply transaction.
941
- 10. Run smoke tests.
942
- 11. Restore automatically if verification fails.
943
-
944
- ### 13.2 Rollback acceptance criterion
945
-
946
- After `loadout rollback`, every file touched by the last transaction must equal its
947
- pre-transaction bytes. Files not managed by Loadout must remain untouched.
948
-
949
- ## 14. Security baseline
950
-
951
- Reject or flag:
952
-
953
- - Absolute paths or `../` traversal escaping package root.
954
- - Symlinks escaping package root.
955
- - Embedded secrets.
956
- - Obfuscated executable payloads.
957
- - `curl | bash` and equivalent remote bootstrap execution.
958
- - Package-manager lifecycle scripts during discovery/install.
959
- - Newly introduced hooks, binaries, domains, environment-variable reads, or broad
960
- filesystem permissions.
961
- - Unpinned remote dependencies where pinning is expected.
962
-
963
- Rules:
964
-
965
- - Never execute third-party code during catalog ingestion.
966
- - Never log secret values.
967
- - Never auto-approve new permissions.
968
- - Pin installed packages to commits and store per-file hashes.
969
- - Keep at least the previous known-good snapshot.
970
- - Treat star count as popularity, not proof of safety.
971
-
972
- ## 15. Dashboard requirements
973
-
974
- ### 15.1 Home
975
-
976
- - Detected agents.
977
- - Operating system.
978
- - Stable/Maximum/Custom call to action.
979
- - Installed package count.
980
- - Updates and conflicts.
981
- - Last known-good snapshot.
982
-
983
- ### 15.2 Discover
984
-
985
- - Outcome-first categories.
986
- - Stable, Trending, Official badges.
987
- - Stars, publisher, platforms, permissions.
988
- - Add/Remove action.
989
- - Technical details drawer.
990
-
991
- ### 15.3 Installed
992
-
993
- - Package list.
994
- - Active platforms.
995
- - Native/adapted/unsupported labels.
996
- - Exact pinned commit.
997
- - Remove and inspect actions.
998
-
999
- ### 15.4 Updates
1000
-
1001
- - Old and candidate versions.
1002
- - Plain-language summary.
1003
- - File/config/permission diff.
1004
- - Update, ignore, or rollback.
1005
-
1006
- ### 15.5 Design constraints
1007
-
1008
- - Local-only application for MVP.
1009
- - No signup wall.
1010
- - First meaningful screen in under five seconds after server launch.
1011
- - Keyboard accessible.
1012
- - Responsive down to tablet width.
1013
- - Never expose secret values in UI.
1014
-
1015
- ## 16. Work allocation
1016
-
1017
- ### Track A: Catalog and discovery — Member 1
1018
-
1019
- Owns catalog schema, initial package records, GitHub metadata retrieval, tiers,
1020
- scoring, conflict families, and static source checks.
1021
-
1022
- ### Track B: Core and adapters — Member 2
1023
-
1024
- Owns CLI, detection, transaction engine, snapshots, lockfile, updates, rollback, and
1025
- agent adapters.
1026
-
1027
- ### Track C: Dashboard and submission — Member 3
1028
-
1029
- Owns dashboard, onboarding, API contract integration, visual polish, demo fixtures,
1030
- README/setup instructions, video, and Devpost content.
1031
-
1032
- Shared decisions require a short decision record in the PR or `docs/decisions/`.
1033
-
1034
- ## 17. Model delegation guide
1035
-
1036
- Task labels:
1037
-
1038
- - `[LUNA]`: bounded, mechanical, clear expected output, low architectural judgment.
1039
- - `[TERRA]`: normal feature implementation with defined interfaces and tests.
1040
- - `[SOL]`: architecture, security, ambiguous integration, conflict resolution, or
1041
- cross-cutting review.
1042
- - `[HUMAN]`: product choice, external permission, legal/licensing judgment, final
1043
- acceptance, or credential handling.
1044
-
1045
- Luna tasks must include exact files, inputs, expected output, and acceptance checks.
1046
- Terra tasks must include an interface or behavior contract and test expectations.
1047
- Sol tasks should produce a decision or implementation plus tradeoffs and failure
1048
- modes.
1049
-
1050
- ## 18. Detailed backlog
1051
-
1052
- ### Phase 0: Repository and team setup
1053
-
1054
- - [x] `P0-01 [LUNA]` Add the npm/TypeScript project skeleton matching section 9.1.
1055
- - Acceptance: `npm ci`, build, lint, typecheck, and tests succeed from the root.
1056
- - [x] `P0-02 [LUNA]` Add `.gitignore`, `.editorconfig`, Prettier, and ESLint defaults.
1057
- - Acceptance: formatting and lint commands run at repository root.
1058
- - [x] `P0-03 [TERRA]` Add GitHub Actions matrix for Node on Windows, macOS, Linux.
1059
- - Acceptance: install, lint, typecheck, and tests run on all three.
1060
- - [x] `P0-04 [HUMAN]` Add all three teammates to the private repository.
1061
- - Verified collaborators: `VirajMishra1`, `cars3`, and `reddynitish`.
1062
- - [ ] `P0-05 [HUMAN]` Protect `main` after the first working CI run.
1063
- - Blocked by GitHub's branch-protection restriction for this private repository on
1064
- the current plan (API returned HTTP 403). Revisit after making the repository
1065
- public or enabling a plan that supports protection.
1066
-
1067
- ### Phase 1: Shared types and catalog
1068
-
1069
- - [x] `P1-01 [SOL]` Finalize catalog, installed-state, plan, and lockfile schemas.
1070
- - Acceptance: schema decision documented; no secret-value fields exist.
1071
- - [x] `P1-02 [TERRA]` Implement Zod schemas and inferred TypeScript types.
1072
- - Acceptance: valid fixtures parse; invalid fixtures fail with actionable errors.
1073
- - [x] `P1-03 [LUNA]` Create valid/invalid catalog fixtures.
1074
- - Acceptance: at least five valid and ten invalid cases.
1075
- - [x] `P1-04 [TERRA]` Implement seed catalog loader.
1076
- - Acceptance: loads bundled catalog offline and returns categories/packages.
1077
- - [x] `P1-05 [LUNA]` Add first ten verified catalog records.
1078
- - Acceptance: source, category, tier, license, commit/ref, components, platforms.
1079
- - [x] `P1-06 [LUNA]` Add next ten verified catalog records.
1080
- - [x] `P1-07 [TERRA]` Implement GitHub metadata fetch with cache and rate-limit errors.
1081
- - [x] `P1-08 [TERRA]` Implement tier and ranking functions.
1082
- - [x] `P1-09 [SOL]` Review scoring for obvious gaming and bias failure modes.
1083
- - [x] `P1-10 [TERRA]` Implement conflict-family resolver.
1084
- - Acceptance: Stable picks one default; hard conflicts block; Custom can override
1085
- soft conflicts.
1086
- - [x] `P1-11 [TERRA]` Expand from 20 to at least 50 fully reviewed catalog records.
1087
- - Every record needs immutable commit, license review, component evidence, platform
1088
- evidence, and an install/config path; popularity alone is insufficient.
1089
- - Fifty evidence-complete records now pin immutable commits. All new skill-bearing
1090
- repositories passed real Loadout discovery/frontmatter inspection; unsafe symlinked
1091
- collections were rejected rather than weakened into the catalog.
1092
-
1093
- ### Phase 2: Agent detection
1094
-
1095
- - [x] `P2-01 [SOL]` Finalize adapter contract and platform capability matrix.
1096
- - [x] `P2-02 [TERRA]` Implement shared filesystem/path utilities.
1097
- - Acceptance: tests cover Windows paths, POSIX paths, WSL distinction, home dirs.
1098
- - [x] `P2-03 [LUNA]` Add fake home-directory fixtures for all platforms.
1099
- - [x] `P2-04 [TERRA]` Implement Claude Code detection and inspection.
1100
- - [x] `P2-05 [TERRA]` Implement Codex detection and inspection.
1101
- - [x] `P2-06 [TERRA]` Implement Cursor detection and inspection.
1102
- - [x] `P2-07 [LUNA]` Implement Gemini CLI detection from approved path table.
1103
- - [x] `P2-08 [LUNA]` Implement OpenCode detection from approved path table.
1104
- - [x] `P2-09 [LUNA]` Implement Hermes detection from approved path table.
1105
- - [x] `P2-10 [TERRA]` Build `loadout doctor` detection report.
1106
-
1107
- ### Phase 3: Package parsing and normalization
1108
-
1109
- - [x] `P3-01 [SOL]` Define normalized package/component representation.
1110
- - [x] `P3-02 [TERRA]` Implement `SKILL.md` parser and validation.
1111
- - [x] `P3-03 [TERRA]` Implement Claude plugin manifest parser.
1112
- - [x] `P3-04 [TERRA]` Implement Codex plugin manifest parser.
1113
- - [x] `P3-05 [TERRA]` Implement MCP JSON parser.
1114
- - [x] `P3-06 [LUNA]` Add parser fixtures from sanitized real layouts.
1115
- - [x] `P3-07 [TERRA]` Map parsed skills to universal component records.
1116
- - [x] `P3-08 [SOL]` Define native/adapted/unsupported rules for MVP platforms.
1117
- - [x] `P3-09 [TERRA]` Generate compatibility summary from normalized package.
1118
-
1119
- ### Phase 4: Transaction engine
1120
-
1121
- - [x] `P4-01 [SOL]` Threat-model the mutation transaction.
1122
- - [x] `P4-02 [TERRA]` Implement immutable package cache by commit.
1123
- - [x] `P4-03 [TERRA]` Implement per-file SHA-256 calculation.
1124
- - [x] `P4-04 [TERRA]` Implement snapshot creator and manifest.
1125
- - [x] `P4-05 [TERRA]` Implement staging directory and planned writes.
1126
- - [x] `P4-06 [TERRA]` Implement path traversal and escaping-symlink rejection.
1127
- - [x] `P4-07 [TERRA]` Implement plan collision detection.
1128
- - [x] `P4-08 [SOL]` Review atomic commit behavior across all three operating systems.
1129
- - Accepted for the supported local-filesystem scope after CI run `29401149042` exercised the atomic and transaction suites on Node 20/22 for Windows, macOS, and Linux. See `docs/RELEASE_REVIEW.md` for the explicit power-loss boundary.
1130
- - [x] `P4-09 [TERRA]` Implement commit with automatic restore on failure.
1131
- - [x] `P4-10 [TERRA]` Implement `loadout rollback`.
1132
- - [x] `P4-11 [LUNA]` Add interrupted-write and corrupted-stage fixtures.
1133
- - [x] `P4-12 [TERRA]` Verify rollback restores byte-identical files.
1134
-
1135
- ### Phase 5: Agent adapters
1136
-
1137
- - [x] `P5-01 [TERRA]` Claude skill install/remove planner.
1138
- - [x] `P5-02 [TERRA]` Codex skill install/remove planner.
1139
- - [x] `P5-03 [TERRA]` Cursor skill install/remove planner.
1140
- - [x] `P5-04 [LUNA]` Gemini skill planner using approved layout.
1141
- - [x] `P5-05 [LUNA]` OpenCode skill planner using approved layout.
1142
- - [x] `P5-06 [LUNA]` Hermes skill planner using approved layout.
1143
- - [x] `P5-07 [SOL]` Review adapters for lossy or false compatibility claims.
1144
- - [x] `P5-08 [TERRA]` Claude MCP config planner preserving unrelated entries.
1145
- - [x] `P5-09 [TERRA]` Codex MCP config planner preserving unrelated entries/comments.
1146
- - Implementation appends only new official TOML tables; replacement of an existing table remains intentionally unsupported until a comment-preserving TOML editor is added.
1147
- - [x] `P5-10 [TERRA]` Cursor MCP config planner preserving unrelated entries.
1148
- - [x] `P5-11 [TERRA]` Smoke-test interface and results.
1149
- - The native-skill adapter smoke suite plans, installs, and removes a real `SKILL.md` fixture for each declared filesystem layout. It does not claim plugin, hook, MCP, or executable runtime support beyond the capability matrix.
1150
-
1151
- ### Phase 6: CLI
1152
-
1153
- - [x] `P6-01 [TERRA]` CLI bootstrap, version, help, structured error handling.
1154
- - [x] `P6-02 [TERRA]` `loadout status`.
1155
- - [x] `P6-03 [TERRA]` `loadout doctor`.
1156
- - [x] `P6-04 [TERRA]` `loadout plan --mode stable|maximum|custom`.
1157
- - [x] `P6-05 [TERRA]` Confirmed `loadout install --yes` and `loadout sync --yes`
1158
- mutation paths.
1159
- - [x] `P6-06 [TERRA]` `loadout add` and `loadout remove`.
1160
- - [x] `P6-07 [TERRA]` Read-only `loadout update` planning by default.
1161
- - [x] `P6-08 [TERRA]` `loadout rollback`.
1162
- - [x] `P6-09 [LUNA]` CLI snapshot tests for help and error messages.
1163
- - [x] `P6-10 [SOL]` Review destructive command confirmation and recovery behavior.
1164
- - [x] `P6-11 [TERRA]` Make interactive CLI setup the primary product path.
1165
- - Maximum/Stable/Custom detect targets, concurrently prepare pinned reviewed
1166
- commits, defer explicit MCP setup, resolve lower-ranked duplicate skill targets,
1167
- show safety findings, and install as one transaction.
1168
-
1169
- ### Phase 7: Local API and dashboard
1170
-
1171
- - [x] `P7-01 [SOL]` Define local API contract and threat boundary.
1172
- - [x] `P7-02 [TERRA]` Start local server on random loopback port with session token.
1173
- - [x] `P7-03 [TERRA]` Agents/status endpoint.
1174
- - [x] `P7-04 [TERRA]` Catalog/list/detail endpoints.
1175
- - [x] `P7-05 [TERRA]` Plan/apply/progress endpoints.
1176
- - [x] `P7-06 [TERRA]` Updates/diff/rollback endpoints.
1177
- - [x] `P7-07 [LUNA]` Dashboard shell, routing, typography, color tokens.
1178
- - [x] `P7-08 [TERRA]` Home screen.
1179
- - [x] `P7-09 [TERRA]` Discover screen.
1180
- - [x] `P7-10 [TERRA]` Installed screen.
1181
- - [x] `P7-11 [TERRA]` Updates and diff screen.
1182
- - [x] `P7-12 [LUNA]` Empty, loading, and error states.
1183
- - [x] `P7-13 [LUNA]` Keyboard and accessible-label pass.
1184
- - [x] `P7-14 [TERRA]` Playwright first-run happy-path test.
1185
- - Runs Chromium against the real loopback dashboard with an empty disposable Loadout home; it previews and applies a safe first-run manifest without touching user configuration.
1186
- - [x] `P7-15 [SOL]` Product and security review of complete flow.
1187
- - Reviewed 2026-07-15; see `docs/RELEASE_REVIEW.md` for boundaries, fixes, and release conditions.
1188
-
1189
- ### Phase 8: Updates and safety demo
1190
-
1191
- - [x] `P8-01 [TERRA]` Detect candidate commit for installed package.
1192
- - [x] `P8-02 [TERRA]` Generate changed-file diff.
1193
- - [x] `P8-03 [TERRA]` Generate instruction/script/domain/env summary.
1194
- - [x] `P8-04 [TERRA]` Implement approval policy for sensitive changes.
1195
- - [x] `P8-05 [LUNA]` Create benign Ponytail-style update fixture.
1196
- - [x] `P8-06 [LUNA]` Create risky update fixture adding a hook and domain.
1197
- - [x] `P8-07 [TERRA]` Demonstrate safe update acceptance.
1198
- - [x] `P8-08 [TERRA]` Demonstrate risky update quarantine.
1199
- - [x] `P8-09 [TERRA]` Demonstrate rollback after simulated smoke-test failure.
1200
-
1201
- ### Phase 9: Cross-platform verification
1202
-
1203
- - [x] `P9-01 [TERRA]` Windows native install test.
1204
- - CI run `29401149042` passed on `windows-latest` with Node 20 and 22. The native-filesystem smoke test used disposable `LOADOUT_USER_HOME` and `LOADOUT_HOME` directories to plan, install, byte-verify, and remove a real skill through every declared agent-owned skills layout.
1205
- - [x] `P9-02 [TERRA]` WSL behavior test or documented compatibility boundary.
1206
- - [x] `P9-03 [TERRA]` macOS install test.
1207
- - CI run `29401149042` passed on `macos-latest` with Node 20 and 22 using the host path implementation, not a simulated layout.
1208
- - [x] `P9-04 [TERRA]` Linux install test.
1209
- - CI run `29401149042` passed on `ubuntu-latest` with Node 20 and 22 using the host path implementation, not a simulated layout.
1210
- - [x] `P9-05 [LUNA]` CRLF/LF fixture coverage.
1211
- - [x] `P9-06 [LUNA]` `.cmd` executable-resolution fixture coverage.
1212
- - [x] `P9-07 [SOL]` Cross-platform go/no-go review.
1213
- - Reviewed 2026-07-15 after successful CI run `29401149042`: go for the bounded claim that Loadout can plan, install, verify, and remove native `SKILL.md` directories on Windows, macOS, and Linux. No-go remains for a universal runtime claim covering plugins, hooks, executables, or arbitrary MCP servers.
1214
-
1215
- ### Phase 10: Submission
1216
-
1217
- - [x] `P10-01 [LUNA]` Expand README with install and supported-platform table.
1218
- - [x] `P10-02 [LUNA]` Add sample catalog data and judge test instructions.
1219
- - [x] `P10-03 [TERRA]` Add one-command demo mode using isolated fake home dirs.
1220
- - [x] `P10-04 [SOL]` Final architecture and threat-model review.
1221
- - [ ] `P10-05 [HUMAN]` Verify licenses and attribution for included sources.
1222
- - [ ] `P10-06 [HUMAN]` Record under-three-minute demo.
1223
- - [ ] `P10-07 [HUMAN]` Explain where Codex and GPT-5.6 were used.
1224
- - [ ] `P10-08 [HUMAN]` Capture required `/feedback` Codex session ID.
1225
- - [ ] `P10-09 [HUMAN]` Complete Devpost description, category, repository, and video.
1226
- - [ ] `P10-10 [HUMAN]` Submit before deadline with buffer.
1227
-
1228
- ### Phase 11: Advanced committed capabilities
1229
-
1230
- - [x] `P11-01 [SOL]` Design GitHub OAuth and minimal-scope authorization model.
1231
- - [x] `P11-02 [TERRA]` Implement optional private-repository discovery.
1232
- - `loadout discover --private` uses an explicit caller-provided `GITHUB_TOKEN`, returns metadata only, and never persists or logs the token; OAuth/App brokering remains deployment-configured per `docs/GITHUB_AUTHORIZATION.md`.
1233
- - [x] `P11-03 [TERRA]` Implement Community Loadout export/import with versioning.
1234
- - [x] `P11-04 [TERRA]` Implement star/release/download snapshot storage and charts.
1235
- - [x] `P11-05 [SOL]` Define provider-neutral model configuration schema.
1236
- - [x] `P11-06 [TERRA]` Implement OpenRouter provider adapter without storing keys in
1237
- application state.
1238
- - The adapter resolves a credential reference at request time and never serializes or logs the raw token.
1239
- - [x] `P11-07 [SOL]` Define category-specific evaluation protocol and uncertainty.
1240
- - [x] `P11-08 [TERRA]` Implement first two automated evaluation categories.
1241
- - Static skill hygiene and MCP manifest evaluations are deterministic and never execute package code; see `docs/EVALUATION_PROTOCOL.md`.
1242
- - [x] `P11-09 [TERRA]` Implement a read-only update watcher and notifications.
1243
- - `loadout watch` performs read-only interval checks and emits human or JSON notifications; it never applies updates automatically.
1244
- - [x] `P11-10 [SOL]` Define catalog signing, rotation, and compromise recovery.
1245
- - [x] `P11-11 [TERRA]` Implement catalog signing and client-side verification tools.
1246
- - `keygen`, `catalog-sign`, and `catalog-verify` are covered by tests. A real release
1247
- key and signed-release publishing step remain owner-controlled release work; CI
1248
- does not contain or manufacture the production signing identity.
1249
- - [x] `P11-12 [SOL]` Design cross-platform hook/subagent compiler with loss reports.
1250
- - [x] `P11-13 [TERRA]` Implement first two hook/subagent conversion targets.
1251
- - `loadout convert` creates a loss-reported static skill from a subagent or a non-executable review artifact from a hook; it never synthesizes executable hook behavior and requires manual approval.
1252
- - [x] `P11-14 [SOL]` Design sandbox threat model for third-party installers.
1253
- - [x] `P11-15 [TERRA]` Implement disposable sandbox runner with no host secrets.
1254
- - `loadout sandbox-run` uses explicit approval, a reviewed image, read-only source mount, no network, dropped capabilities, resource limits, and a scrubbed environment; Docker remains an explicit local prerequisite.
1255
- - [x] `P11-16 [SOL]` Design OS-keychain-backed credential interface.
1256
- - [x] `P11-17 [TERRA]` Implement macOS, Windows, and Linux credential backends.
1257
- - `credentials` uses macOS Keychain, Linux Secret Service, or Windows Credential
1258
- Manager through bounded no-shell processes. Writes use stdin, errors are redacted,
1259
- and secret values never enter plans, arguments, snapshots, or JSON output.
1260
- - [x] `P11-18 [SOL]` Define autonomous-update permission policies and recovery rules.
1261
- - [x] `P11-19 [TERRA]` Implement a policy-gated canary planning pipeline.
1262
- - `loadout canary` performs a non-mutating static gate; promotion requires explicit approval plus injected verification and transaction callbacks, so it cannot silently update agent files.
1263
- - [x] `P11-20 [SOL]` Define the internal adapter contract and conformance tests.
1264
- - `src/core/adapters.ts`, the shared capability matrix, compatibility policy, and
1265
- conformance tests are the implemented contract. A separately versioned public SDK
1266
- package and community registry are not yet published.
1267
- - [x] `P11-21 [TERRA]` Add the next six agent adapters through the SDK.
1268
- - Windsurf, Cline, GitHub Copilot, Roo Code, Kiro CLI, and Junie use documented
1269
- vendor-specific Agent Skills roots. Only native skill support is claimed; all
1270
- unverified component types remain explicitly unsupported.
1271
- - [x] `P11-22 [TERRA]` Add compliant Hacker News and community-source connectors.
1272
- - Hacker News Firebase and GitHub REST repository search are read-only connectors; neither mutates the catalog or installs a lead.
1273
- - [x] `P11-23 [SOL]` Design team/enterprise policy and audit schemas.
1274
- - [x] `P11-24 [TERRA]` Implement shared Loadouts, allowlists, denylists, and audit view.
1275
- - Manifest policy now enforces package/repository allowlists and denylists before synchronization; existing audit output remains the read-only decision view.
1276
-
1277
- ### Phase 12: Best-available optimization and public beta
1278
-
1279
- This phase turns the safe installer into the product thesis: Loadout continuously
1280
- understands what a user already has, maintains a broad reviewed library, and exposes
1281
- only a small evidence-backed active set for the current agent and project. “Best”
1282
- always means best supported choice under disclosed evidence and uncertainty, never a
1283
- universal or permanent truth.
1284
-
1285
- - [x] `P12-01 [SOL]` Correct the default product posture from Maximum-first to
1286
- Stable-first.
1287
- - Stable is the small `superpowers + context7` foundation when those records exist.
1288
- Maximum remains an explicit broad-library/stress mode.
1289
- - [x] `P12-02 [TERRA]` Add read-only `loadout scan` for existing skill directories.
1290
- - Acceptance: report actual `SKILL.md` count, normalized names, content
1291
- fingerprints, Loadout ownership, unmanaged content, within-agent duplicates,
1292
- cross-agent mirrors, per-agent totals, and capacity warnings without executing or
1293
- changing instructions.
1294
- - [x] `P12-03 [TERRA]` Close integration defects found by the packaged-CLI audit.
1295
- - New packages include an installable skill skeleton; skipped enabled packages fail
1296
- synchronization rather than producing a misleading lock; absent evaluation
1297
- categories are `not-applicable`; empty canaries still block; signing creates parent
1298
- directories; portable absolute-path errors provide a remedy.
1299
- - [x] `P12-04 [TERRA]` Warn when a prepared loadout exceeds 30 active skill
1300
- directories per agent.
1301
- - The warning is a capacity heuristic, not a claim that the agent cannot load more.
1302
- - [x] `P12-05 [SOL]` Define the provenance confidence model for existing unmanaged
1303
- content.
1304
- - Levels: exact Loadout record, exact catalog hash, embedded repository/commit,
1305
- heuristic source match, and unknown. Never invent provenance from a folder name.
1306
- - [x] `P12-06 [TERRA]` Implement catalog-hash and embedded-metadata provenance
1307
- matching in `loadout scan`.
1308
- - Acceptance: every match includes evidence and confidence; network access is
1309
- optional; unknown remains a first-class result.
1310
- - [x] `P12-07 [SOL]` Define semantic duplicate and capability-family rules.
1311
- - Separate exact duplicate, same-name divergent content, cross-agent mirror,
1312
- overlapping workflow, complementary capability, and verified hard conflict.
1313
- - [x] `P12-08 [TERRA]` Implement `loadout compare <skill-or-package>`.
1314
- - Show installed candidate, reviewed alternatives, provenance, maintenance,
1315
- adoption velocity, permissions, compatibility, evaluation evidence, uncertainty,
1316
- and a plain-language recommendation. No mutation.
1317
- - [x] `P12-09 [SOL]` Define the reviewed-library versus active-set state model and
1318
- migration boundary.
1319
- - Downloaded/cached, reviewed, installed, active, disabled, quarantined, and removed
1320
- are distinct states. Existing user files are never silently adopted or deleted.
1321
- - [x] `P12-10 [TERRA]` Implement transactional `loadout enable` and `loadout disable`.
1322
- - Acceptance: only Loadout-managed links/files change; one snapshot covers a batch;
1323
- disabling preserves the library copy; rollback restores byte-identical state.
1324
- - [x] `P12-11 [TERRA]` Implement `loadout adopt` for explicitly selected unmanaged
1325
- skills.
1326
- - Preview provenance and hashes, snapshot first, preserve original content, and
1327
- require confirmation. Bulk adoption without review is forbidden.
1328
- - Adoption is one-skill-only, dry-run by default, rechecks the fingerprint before
1329
- the state transaction, and marks only exact catalog fingerprints as reviewed.
1330
- - [x] `P12-12 [SOL]` Define active-set selection policy.
1331
- - Inputs include user-pinned capabilities, project signals, agent compatibility,
1332
- conflicts, task families, capacity budget, evaluation confidence, and prior human
1333
- outcomes. Popularity cannot override safety or user pins.
1334
- - The complete ordering and neutral boundaries for not-yet-available evaluation and
1335
- outcome evidence are documented in `docs/ACTIVE_SET_POLICY.md`.
1336
- - [x] `P12-13 [TERRA]` Implement project-aware `loadout activate --project <path>`.
1337
- - Preview the delta between global and project active sets; do not require GitHub;
1338
- do not expose irrelevant library content to the agent.
1339
- - Selection is per skill (not per mega-repository), local-only, capacity-bounded,
1340
- pin-aware, agent-scoped, and dry-run by default.
1341
- - [x] `P12-14 [TERRA]` Implement `loadout optimize` as the primary guided workflow.
1342
- - Flow: scan -> explain findings -> compare alternatives -> propose active set ->
1343
- preview exact changes -> confirm -> verify -> provide one-command rollback.
1344
- - The guided CLI prints project signals, scores and reasons, equivalent-source
1345
- alternatives, the exact enable delta, verified snapshot id, and rollback command.
1346
- - [x] `P12-15 [SOL]` Design representative, category-specific head-to-head
1347
- evaluations.
1348
- - Start with workflow adherence, code-review coverage, documentation retrieval, and
1349
- browser-test planning. Record fixtures, rubrics, model/version, variance, cost, and
1350
- uncertainty; never execute untrusted host code.
1351
- - `docs/HEAD_TO_HEAD_EVALUATION.md` defines fixtures, weighted rubrics, trial
1352
- controls, variance/effect thresholds, cost evidence, uncertainty, non-execution
1353
- boundaries, and signed snapshot requirements for all four categories.
1354
- - [x] `P12-16 [TERRA]` Implement the first two head-to-head evaluation harnesses and
1355
- persist signed evidence snapshots.
1356
- - `loadout head-to-head` scores synthetic workflow-adherence and code-review-coverage
1357
- trial observations against declared fixtures, persists an Ed25519-signed evidence
1358
- envelope, and never executes candidate content. Results do not silently replace a
1359
- user's active capability.
1360
- - [x] `P12-17 [TERRA]` Add daily candidate ingestion and review queues.
1361
- - Combine official sources, GitHub search, release/activity observations, star
1362
- velocity, compliant community connectors, deduplication, rate-limit handling, and
1363
- a human promotion gate. Discovery never installs automatically.
1364
- - `discover --source all --queue` aggregates the documented GitHub REST and
1365
- Hacker News Firebase sources, preserves partial-source failures, deduplicates
1366
- leads, and keeps human shortlist/ignore decisions. Repeated GitHub observations
1367
- calculate disclosed per-day star velocity; `schedule --job discovery` runs only
1368
- this read-only candidate queue refresh. Discovery never installs or promotes a
1369
- candidate.
1370
- - [x] `P12-18 [TERRA]` Add freshness and replacement alerts.
1371
- - Explain when an installed source is archived, materially stale, permission-expanded,
1372
- superseded, or outperformed by reviewed evidence. Offer compare/ignore/pin actions.
1373
- - `alerts` reports archived, one-year-stale, reviewed-commit-change,
1374
- permission-expansion, and verified signed-evidence outperformance findings with
1375
- compare/update/disable actions and local ignore. `alert-pin`, `alert-unpin`, and
1376
- `alert-pins` persist explicit local replacement preferences without changing the
1377
- active set.
1378
- - [x] `P12-19 [SOL]` Define privacy-preserving local outcome signals.
1379
- - Default local-only: explicit accept/reject, rollback, disable, repeated activation,
1380
- and task-category success. No source code, prompts, filenames, or secrets leave the
1381
- machine without separate informed consent.
1382
- - The bounded local store accepts only exact package/skill selectors, agent ids,
1383
- task families, outcome enums, and timestamps; paths and arbitrary notes are rejected.
1384
- - [x] `P12-20 [TERRA]` Connect improvement feedback to ranking evidence without
1385
- creating a popularity feedback loop.
1386
- - Human outcomes are scoped by task and agent; one user's preference cannot globally
1387
- crown a package.
1388
- - The active-set policy applies capped adjustments only to the same selector,
1389
- agent, and task family. Strong rejection/rollback evidence suppresses automatic
1390
- selection, while an explicit pin remains the user's override.
1391
- - [x] `P12-21 [TERRA]` Expose provider-neutral model/OpenRouter configuration through
1392
- validated CLI commands.
1393
- - The current adapter is library-level only. Acceptance requires plan/apply/status,
1394
- redacted output, credential references, and disposable mocked-network tests before
1395
- any real-key test.
1396
- - `loadout models set/status/verify` stores only validated metadata and credential
1397
- references; apply is snapshotted, output is redacted, and provider requests resolve
1398
- an environment or native OS credential only at the explicit verification boundary.
1399
- - [x] `P12-22 [TERRA]` Add reviewed MCP setup recipes and connection verification.
1400
- - `mcp-recipe` provides immutable, source-linked Playwright and GitHub read-only recipes,
1401
- previews commands, permissions, environment names, and target config without
1402
- printing values, and separates authorization from configuration. It preserves
1403
- unrelated JSON keys and verifies configured references without launching. An
1404
- explicitly approved `--connect` path starts only the exact reviewed npm version or
1405
- OCI digest, resolves credentials just-in-time, performs a bounded MCP initialize
1406
- handshake, redacts failures, and cleans up the process; the existing Codex MCP
1407
- planner remains the TOML-preserving path.
1408
- - [x] `P12-23 [TERRA]` Complete P11-17 keychain backends and connect them to provider,
1409
- private-discovery, registry, and MCP workflows.
1410
- - Provider verification, private GitHub discovery, remote registry publishing/
1411
- serving, and explicit MCP connection checks share environment/native-keychain
1412
- references. MCP secrets enter only the short-lived verified subprocess environment
1413
- and are never written into agent configuration or Loadout state.
1414
- - [x] `P12-24 [SOL]` Design a cross-platform daily scheduler that invokes read-only
1415
- discovery/update checks.
1416
- - macOS LaunchAgent, Windows Task Scheduler, and Linux systemd/cron implementations
1417
- must be opt-in, inspectable, removable, rate-limited, and unable to apply updates.
1418
- - The native plans schedule only `loadout watch --once --json`; generated files and
1419
- native actions are shown in the dry run, and no apply-capable command is present.
1420
- - [x] `P12-25 [TERRA]` Implement `loadout schedule` and `loadout unschedule` with native
1421
- disposable/configuration tests.
1422
- - macOS LaunchAgent, Linux systemd user timer, and Windows Task Scheduler XML are
1423
- generated natively, snapshotted, installed only with `--yes`, and removable.
1424
- - [x] `P12-26 [TERRA]` Polish the CLI as the sole required product surface.
1425
- - Consistent progress, compact tables, accessible color/no-color output, actionable
1426
- errors, interruption handling, shell completion, noninteractive JSON, and terminal
1427
- widths from 80 to 200 columns.
1428
- - Bash, Zsh, Fish, and PowerShell completion cover top-level and nested credential/
1429
- model commands. Help flushes fully through pipes; automation can request structured
1430
- JSON errors; capability tables are deterministic, ANSI-free, and bounded at 80,
1431
- 120, and 200 columns. Long-running services clean up signals, and mutations retain
1432
- transactional interruption recovery.
1433
- - [x] `P12-27 [LUNA]` Reframe README, testing, and demo around scan/compare/optimize;
1434
- move dashboard instructions to an optional diagnostics section.
1435
- - README and disposable testing now lead with the scan -> compare -> optimize ->
1436
- rollback workflow; dashboard instructions are explicitly optional diagnostics.
1437
- - [x] `P12-28 [TERRA]` Add a privacy-safe `loadout report`/`loadout share` artifact.
1438
- - Default output contains package ids, versions, evidence, and compatibility only;
1439
- exclude usernames, absolute paths, private repositories, project names, and secrets.
1440
- - The artifact contains package ids, commits, agent compatibility, aggregate
1441
- activation/review counts, and MCP package ids; repository names, server names,
1442
- paths, filenames, projects, prompts, code, and credential data are excluded.
1443
- - [x] `P12-29 [TERRA]` Complete P1-11 with at least 50 reviewed records and capability
1444
- coverage metrics.
1445
- - Measure unique capabilities, overlap, licenses, immutable commits, install shape,
1446
- platforms, activity, and evaluation readiness rather than raw repository count.
1447
- - `catalog --coverage [--json]` reports 50 immutable records across 37 categories,
1448
- including component/install shapes, overlaps, licenses, source-inspection platforms,
1449
- activity observations, and evaluation readiness.
1450
- - [ ] `P12-30 [HUMAN]` Complete legal/license attribution review, including all current
1451
- `NOASSERTION` records, before distribution.
1452
- - [ ] `P12-31 [TERRA]` Publish `loadout-ai` to npm and run clean-machine package tests
1453
- from outside the repository on macOS, Windows, Linux, and Node 20/22.
1454
- - Partial: `loadout-ai@0.2.0` was published publicly on 2026-07-17 while the GitHub
1455
- repository remained private. Registry version/integrity, a clean external npm
1456
- install, version/help, 50-record catalog coverage, and the real disposable
1457
- Superpowers install/rollback demo pass on macOS. Hosted matrix evidence already
1458
- covers Windows/macOS/Linux and Node 20/22; clean independent-user installs on
1459
- Windows and Linux remain required.
1460
- - [ ] `P12-32 [HUMAN]` Run moderated founder testing on the real Claude and Codex
1461
- profiles with snapshots and explicit rollback checkpoints.
1462
- - Partial 2026-07-16: the founder ran the published npm CLI against the real macOS
1463
- profile. Read-only doctor/scan/status/health and snapshot listing completed; the
1464
- run exposed and regression-tested the Codex Desktop detection fix. A real Stable
1465
- mutation and rollback remain pending.
1466
- - [ ] `P12-33 [HUMAN]` Run at least ten external user tests spanning new users, power
1467
- users, Windows, macOS, Linux, one-agent, and multi-agent setups.
1468
- - [ ] `P12-34 [SOL]` Public-beta go/no-go review.
1469
- - Required: zero known destructive data-loss defects; every mutation previewed and
1470
- recoverable; no false “best” or compatibility claims; install/optimize/rollback
1471
- success on all supported platforms; p95 local scan under five seconds for 1,000
1472
- skills; actionable failure messages; npm provenance and attribution complete.
1473
- - Partial evidence: seven real CLI runs over 1,000 on-disk skills measured a local
1474
- p95 of 1.28 seconds on 2026-07-16. A disposable real Maximum flow also prepared all
1475
- 31 skill-bearing repositories, exposed 1,219 skill directories, resolved 48
1476
- overlaps, exercised optimization/apply/rollback, and left no test profile behind.
1477
- Transaction, rollback, package-tarball, and CLI product-flow gates pass locally;
1478
- hosted OS/Node release evidence and attribution approval remain required before
1479
- go-live.
1480
-
1481
- ### Phase 13: Pre-testing hardening and continuous discovery
1482
-
1483
- - [x] `P13-01 [TERRA]` Repair default GitHub discovery against the live API.
1484
- - Replace the ineffective parenthesized topic query with three valid rolling topic
1485
- searches, merge case-insensitive duplicates, rank deterministically, preserve
1486
- exact custom-query semantics, and test rate-limit and malformed-response failures.
1487
- - [x] `P13-02 [TERRA]` Publish a bounded daily discovery evidence feed.
1488
- - One five-minute GitHub Actions job performs eight real GitHub searches, records
1489
- current and retained candidates in `catalog/discovered.json`, and regenerates
1490
- `docs/DISCOVERED.md` without installing or promoting a repository.
1491
- - Day-one ordering uses an explicitly labelled lifetime star average; after one
1492
- complete day, real observed star velocity takes precedence. Empty or malformed
1493
- runs cannot replace the last healthy artifact, and CI validates every generated
1494
- record.
1495
- - [x] `P13-03 [LUNA]` Rebuild the README and credit every reviewed upstream source.
1496
- - The README is CLI-first, accurately distinguishes preview/network/mutation
1497
- boundaries, links the daily candidate feed, and records npm publication accurately.
1498
- - `docs/CATALOG.md` links all 50 repositories and all 50 immutable reviewed commits;
1499
- CI prevents silent attribution drift and identifies six `NOASSERTION` records for
1500
- human legal review.
1501
- - [x] `P13-04 [TERRA]` Harden beginner test-drive and rollback recovery.
1502
- - Catalog-backed demos fetch the exact reviewed commit, `test-drive` aliases the
1503
- isolated install/rollback exercise, and all `--agents` arguments use one validated
1504
- parser.
1505
- - Rollback lists snapshots read-only and rejects traversal, malformed bytes,
1506
- filesystem/home/state roots, overlapping roots, escaping paths, duplicates, and
1507
- inconsistent persisted records before mutation.
1508
- - [x] `P13-05 [SOL]` Document an executable full-feature founder test matrix.
1509
- - `docs/FEATURE_TEST_MATRIX.md` covers every top-level command, authority boundary,
1510
- network/process/credit side effect, expected result, platform integration, and
1511
- cleanup path using disposable profiles wherever possible.
1512
- - [x] `P13-06 [TERRA]` Pass the integrated pre-testing gate and repeat the real live
1513
- discovery plus immutable isolated test-drive after all Phase 13 changes.
1514
- - `npm run verify` passes 73 test files/270 tests, CLI E2E, installed-tarball
1515
- smoke, generated-evidence validation, and a 1,000-skill p95 of 0.97 seconds.
1516
- Playwright passes independently; live discovery returns real leads; the reviewed
1517
- two-agent test-drive installs at the pinned Superpowers commit and rolls back its
1518
- temporary profile.
1519
-
1520
- ### Phase 14: Candidate intelligence and trusted catalog delivery
1521
-
1522
- - [x] `P14-01 [TERRA]` Turn the generated discovery feed into an explainable candidate
1523
- triage CLI.
1524
- - `candidate list` validates the feed, supports bounded search/limits, distinguishes
1525
- measured velocity from lifetime averages, and discloses that adoption evidence is
1526
- not quality or safety evidence.
1527
- - [x] `P14-02 [TERRA]` Build immutable static candidate dossiers.
1528
- - `candidate inspect` resolves a real public repository to a full commit and records
1529
- portable component paths, static evaluation, license, growth, catalog overlap,
1530
- uncertainty, and human-review blockers without executing candidate content.
1531
- - Git fetch is time-bounded; system/global Git config, templates, hooks, credential
1532
- helpers, inherited `GIT_*` overrides, and LFS smudging are isolated. A bounded
1533
- GitHub tree-API preflight must prove every blob size and a total below 100 MiB and
1534
- 20,000 files before checkout.
1535
- - [x] `P14-03 [TERRA]` Add a human-gated catalog proposal boundary.
1536
- - Blocked dossiers cannot produce proposals; platform/category/id claims are
1537
- explicit; the pinned source inspection and evaluation are recomputed before
1538
- admission; preview is default; approved output remains an isolated record and
1539
- never mutates the catalog or installs the source.
1540
- - [x] `P14-04 [TERRA]` Make signed remote catalog releases operational.
1541
- - `catalog-update` accepts local files or bounded HTTPS, verifies Ed25519 signatures
1542
- and complete immutable evidence, previews exact additions/updates/removals, blocks
1543
- replay and implicit removals, then snapshots and atomically applies trusted state.
1544
- - Apply revalidates and recomputes under an exclusive lock, pins the first signing
1545
- key, preserves replay high-water across catalog rollback, and never lets unsigned
1546
- cached metadata clear a signed archive decision.
1547
- - Effective catalog loads re-verify the persisted envelope and merge only mutable
1548
- GitHub refresh metadata over the trusted signed base.
1549
- - [x] `P14-05 [TERRA]` Connect local outcomes to normal recommendations and expose
1550
- adapter expansion gaps honestly.
1551
- - `recommend --agent` applies capped agent/task-scoped local evidence without adding
1552
- unreviewed candidates. `capabilities --gaps` lists unsupported combinations and
1553
- the documentation, preservation, transaction, and smoke evidence required before
1554
- support can be claimed.
1555
- - [x] `P14-06 [TERRA]` Pass the full integrated gate and a live candidate dossier flow.
1556
- - Required: full verify, package smoke, CLI product flow, real current-feed list, one
1557
- live immutable candidate inspection, no source execution, and clean git state.
1558
- - `npm run verify:full` passes 77 test files/310 tests, evidence validation, CLI
1559
- product flow, package smoke, a 1,000-skill p95 of 1.54 seconds, and Playwright;
1560
- the catalog carries 50 credited immutable records and the current discovery feed
1561
- contains 242 validated leads. A disposable live inspection pinned
1562
- `Leonxlnx/taste-skill` at commit
1563
- `b17742737e796305d829b3ad39eda3add0d79060`, found 13 skills and one plugin,
1564
- surfaced script/network/environment review findings and five catalog overlaps,
1565
- and executed or installed none of the source.
1566
-
1567
- ### Phase 15: Stable release candidate and daily autopilot
1568
-
1569
- - [x] `P15-01 [TERRA]` Make Stable the strongest low-risk default, not a minimal demo.
1570
- - Stable now selects 30 high-value engineering, documentation, frontend,
1571
- observability, performance, planning, testing, review, Git, architecture, and
1572
- current-documentation skills from four immutable SPDX-identified sources.
1573
- - A live Codex-targeted preparation found zero static-risk approvals, collisions,
1574
- or skipped components. Power and Maximum remain broader opt-in libraries.
1575
- - [x] `P15-02 [TERRA]` Separate technical screening from recommendation trust.
1576
- - The catalog and coverage API expose `discovered`, `inspected`, `human-reviewed`,
1577
- `benchmarked`, and `recommended` stages. The bundled release reports 50
1578
- technically screened records and four recommended Stable sources without
1579
- manufacturing human-review or benchmark evidence.
1580
- - [x] `P15-03 [TERRA]` Correct candidate installability and runtime-tool detection.
1581
- - Candidate dossiers classify `portable-components`, `explicit-runtime-setup`, or
1582
- `unsupported-source-shape`. Conventional component directories accept real
1583
- manifest files only, preventing nested reference folders from being mislabeled
1584
- as agents; the live Graphify inspection now reports an unsupported runtime-tool
1585
- shape instead of a portable agent bundle.
1586
- - [x] `P15-04 [TERRA]` Add one-command daily autopilot on macOS, Linux, and Windows.
1587
- - `autopilot` previews, transactionally applies, or removes both daily update and
1588
- discovery schedules. Scheduled commands use the pinned npm package version and
1589
- remain read-only: they never install, promote, or update content automatically.
1590
- - [x] `P15-05 [TERRA]` Harden npm release metadata and trusted-publishing workflow.
1591
- - The package has normalized repository metadata, Node 24/npm 11 trusted-publish
1592
- tooling, full `npm run verify`, tag/version consistency, public provenance, and a
1593
- first-publish token fallback.
1594
- - [x] `P15-05A [TERRA]` Pass the complete local release-candidate gate.
1595
- - `npm run verify` passes 79 test files/317 tests, catalog/discovery attribution,
1596
- the real CLI product flow, installed npm-tarball smoke, and seven real scans of
1597
- 1,000 on-disk skills at a 1.29-second p95 on macOS/Node 23.
1598
- - [ ] `P15-06 [HUMAN]` Approve public repository visibility, complete the six
1599
- `NOASSERTION` license decisions, authenticate npm, and publish `loadout-ai`.
1600
- - Partial 2026-07-17: npm authentication and the public `loadout-ai@0.2.0`
1601
- publication succeeded. Repository visibility remains private at the owner's
1602
- request, and the six root-level `NOASSERTION` decisions remain a human release
1603
- gate; therefore this combined item is not ticked.
1604
- - [x] `P15-07A [TERRA]` Run the hosted macOS/Windows/Linux Node matrix.
1605
- - GitHub Actions run `29502324100` passes fast verification, dashboard browser
1606
- diagnostics, and native install/package flows on Windows, macOS, and Ubuntu with
1607
- Node 20 and 22. The first run exposed a Windows-only root-path test assumption;
1608
- the portable regression fix passes on both Windows versions.
1609
- - [ ] `P15-07B [HUMAN]` After npm publication, verify a clean external
1610
- `npx loadout-ai` installation on Windows, macOS, and Linux.
1611
- - [x] `P15-08 [SOL]` Implement the first bounded explicit runtime-tool recipe.
1612
- - `loadout tool graphify` previews and installs Graphify 0.9.17 from an exact
1613
- SHA-256-pinned PyPI wheel linked to a reviewed Git commit. It uses an isolated
1614
- `uv` tool directory, fixed commands, a five-minute timeout, a credential-stripped
1615
- subprocess environment, exact agent targets, version verification, generated
1616
- runtime pinning, snapshots, rollback on failure, and reversible removal.
1617
- - A real disposable Codex-profile exercise installed the pinned wheel, generated
1618
- the skill, verified version 0.9.17 and the pinned lookup, removed it, restored the
1619
- prior profile, and left no runtime or active-profile residue.
1620
- - [ ] `P15-09 [SOL]` Generalize the reviewed runtime-recipe schema only as new tools
1621
- earn admission. Add OS sandbox backends where available; never execute an
1622
- arbitrary candidate repository installer or inherit provider credentials.
1623
-
1624
- ### Phase 16: Evidence-driven product moat and simple upgrade journey
1625
-
1626
- Phase 16 changes the product category from “another agent package manager” into the
1627
- trust, discovery, evaluation, optimization, and rollback layer above skills.sh,
1628
- OpenPackage, Microsoft APM, the official MCP Registry, and raw GitHub repositories.
1629
- Loadout may ingest those ecosystems as read-only evidence sources or use their
1630
- declarative formats as inputs, but it must not duplicate mature dependency-manager
1631
- behavior merely to increase its command or repository count.
1632
-
1633
- #### Phase 16 model-routing and credit policy
1634
-
1635
- - **Sol** owns ambiguous, security-sensitive, research-heavy, or product-defining
1636
- work: evaluation methodology, threat models, ranking/promotion policy,
1637
- compatibility semantics, runtime execution boundaries, and final integration
1638
- reviews. Sol work must produce decisions, invariants, adversarial cases, or review
1639
- evidence—not bulk mechanical edits.
1640
- - **Terra** owns the main implementation path: CLI orchestration, connectors,
1641
- schemas, state machines, transactions, provider adapters, scoring engines,
1642
- cross-platform integration, and regression fixes. Terra is the default for
1643
- production code with clear acceptance criteria.
1644
- - **Luna** owns bounded repeatable work: fixtures, schema examples, deterministic
1645
- transforms, documentation matrices, command completion, generated reports, and
1646
- repetitive adapter tests. Luna output still requires Terra or Sol integration
1647
- review when it affects trust, mutation, or public claims.
1648
- - Prefer Standard speed. Fast mode is reserved for a submission-critical wall-clock
1649
- emergency because GPT-5.6 Fast consumes credits at a higher multiplier.
1650
- - Reserve at least 15% of remaining Codex credit for integration failures, founder
1651
- feedback, cross-platform regressions, release review, and submission polish.
1652
- - ChatGPT/Codex credit must not be represented as OpenAI API credit. A model-backed
1653
- evaluation runner may be built without credentials, but paid benchmark execution
1654
- requires a separately verified provider credential and explicit per-run budget.
1655
- - No task may consume credit merely to exhaust the grant. Every model-backed run must
1656
- have a hypothesis, maximum trials/tokens/cost, deterministic acceptance test, and
1657
- persisted evidence artifact.
1658
-
1659
- #### Product contract
1660
-
1661
- The default journey must answer, in order:
1662
-
1663
- 1. What agents and extensions are already installed?
1664
- 2. Which files are owned, duplicated, stale, unsupported, or risky?
1665
- 3. Which reviewed additions fit this project and agent?
1666
- 4. Why is each addition preferred to a real alternative?
1667
- 5. What exact permissions and filesystem changes will occur?
1668
- 6. Can the complete change be restored byte-for-byte?
1669
- 7. Did the resulting loadout improve a measurable task outcome?
1670
-
1671
- The intended first-run surface is one command, with advanced commands retained:
1672
-
1673
- ```text
1674
- loadout upgrade
1675
- -> scan
1676
- -> diagnose and score
1677
- -> recommend and compare
1678
- -> preview permissions and exact targets
1679
- -> snapshot and transactionally apply
1680
- -> optimize the bounded active set
1681
- -> verify health
1682
- -> show an evidence-linked before/after report
1683
- ```
1684
-
1685
- - [x] `P16-01 [SOL]` Specify the Loadout Evaluation Protocol v1.
1686
- - Define paired baseline-versus-skill trials, pinned repositories, task families,
1687
- hidden or immutable acceptance tests, randomized run ordering, minimum repeats,
1688
- model/provider/version capture, temperature/reasoning capture where available,
1689
- latency, tokens, reported cost, pass rate, regressions, and uncertainty.
1690
- - Separate deterministic task verification from model-based judging. A model judge
1691
- can annotate qualitative dimensions but cannot override a failed executable test.
1692
- - Define contamination, prompt-injection, evaluator-tampering, flaky-test, timeout,
1693
- and partial-result policies before executing a community skill.
1694
- - Acceptance: the protocol is versioned, fixture-hash bound, reproducible, privacy
1695
- bounded, and explicit about what it cannot prove.
1696
- - Completed 2026-07-16 in `docs/EVALUATION_PROTOCOL_V1.md` with strict campaign and
1697
- resumable-run schemas, deterministic paired scheduling, blinded order, content
1698
- hashes, retry-inclusive budgets, interruption/tamper rules, privacy boundaries,
1699
- and nine adversarial protocol tests.
1700
-
1701
- - [ ] `P16-02 [TERRA]` Implement a provider-neutral benchmark run schema and budget
1702
- gate.
1703
- - Add exact model/provider/endpoint references without persisting secrets; store
1704
- input/output token counts, latency, reported cost, exit state, fixture hash,
1705
- candidate hash, agent version, and trial seed.
1706
- - Require `--max-cost`, `--max-trials`, and `--approve-model-spend` before a paid
1707
- runner starts. Preview must calculate the maximum possible spend and run count.
1708
- - Resume an interrupted campaign without duplicating completed trials; atomically
1709
- persist append-only evidence and reject edited or mismatched artifacts.
1710
- - Acceptance: no API key, prompt content, project source, or credential value enters
1711
- logs, lockfiles, reports, snapshots, or signed public evidence.
1712
- - Engineering complete 2026-07-16: campaign/run validation, canonical hashes,
1713
- deterministic recovery, worst-case budget preview, mode-0600 metadata, and
1714
- `loadout benchmark plan` are implemented. The runner adds a strict JSONL hash
1715
- chain, fsync, a cross-process lock, hard aggregate ceilings, retry accounting,
1716
- output hashes only, tamper/torn-log rejection, and explicit interrupted-provider
1717
- reconciliation. It has no default provider or secret-bearing endpoint.
1718
- - Acceptance still open: a real paid-provider adapter and live reconciliation run
1719
- must prove provider usage agrees with the local ceilings.
1720
-
1721
- - [ ] `P16-03 [SOL+TERRA]` Implement the isolated paired evaluation runner.
1722
- - Use disposable worktrees or copied fixtures and the strongest available local OS
1723
- sandbox. Network is denied by default; any required domain is declared and
1724
- separately approved. Candidate install scripts are never run implicitly.
1725
- - Run baseline and candidate with identical repository bytes, acceptance tests,
1726
- model settings, budgets, and tool policy; randomize ordering and record failures.
1727
- - Candidate output cannot edit evaluator code, hidden tests, prior evidence, or the
1728
- opposing trial. Restore or destroy the sandbox after every trial.
1729
- - Acceptance: adversarial fixtures prove evaluator isolation, timeout, budget,
1730
- interruption recovery, and tamper rejection on macOS, Linux, and Windows-capable
1731
- fallback paths.
1732
- - Engineering complete 2026-07-16: the provider-neutral runner requires injected
1733
- provider/isolation executors plus explicit spend approval, randomizes deterministic
1734
- pairs, pauses on unknown paid-provider state, hashes outputs, tears down every
1735
- request, and selects Docker then Podman with no host fallback. Acceptance remains
1736
- open for a concrete container/worktree executor and live cross-platform matrix.
1737
-
1738
- - [ ] `P16-04 [LUNA, SOL review]` Create the first real benchmark fixture suite.
1739
- - Start with planning/workflow adherence, code review, frontend accessibility,
1740
- debugging, documentation freshness, API design, and safe migration tasks.
1741
- - Every fixture has a pinned permissively licensed repository or synthetic source,
1742
- deterministic setup, explicit acceptance criteria, expected runtime, and license.
1743
- - Include negative-control skills, deliberately outdated guidance, overlapping
1744
- skills, and no-skill baselines so the harness can detect zero or negative value.
1745
- - Acceptance: at least five trials per compared candidate in the release evidence;
1746
- fixtures themselves contain no mock performance claims or fabricated outcomes.
1747
- - Engineering complete 2026-07-16: seven synthetic MIT-licensed task families plus
1748
- no-skill, negative, outdated, and overlapping controls have exact file/fixture/
1749
- rubric/control/suite hashes, deterministic materialization and grading, bounded
1750
- cross-platform metadata, and tamper/symlink/path/inventory tests. No outcome was
1751
- invented; acceptance remains open for five real paired trials per candidate.
1752
-
1753
- - [ ] `P16-05 [SOL]` Connect benchmark evidence to trust and Stable promotion.
1754
- - `benchmarked` requires signed protocol-conformant evidence; `recommended` requires
1755
- human license/trust review plus no blocking security finding and meaningful gain
1756
- in at least one declared task family without an unacceptable regression.
1757
- - Stars, install telemetry, and popularity can prioritize evaluation but cannot
1758
- establish quality. Missing evidence contributes zero rather than a neutral score.
1759
- - Acceptance: catalog coverage explains every promotion/demotion and retains the
1760
- prior signed evidence so a recommendation cannot silently change.
1761
- - Engineering complete 2026-07-16: signed evidence validation recomputes paired
1762
- task-family deltas from hash-bound completions; recommendation requires meaningful
1763
- gain, no unacceptable regression, no blocking security finding, and a commit-bound
1764
- signed human attestation. Hash-chained decisions retain prior evidence and demote
1765
- stale revisions. Real trials and genuine human review remain external gates.
1766
-
1767
- - [x] `P16-06 [SOL design, TERRA implementation]` Add the unified `loadout upgrade`
1768
- golden path.
1769
- - One preview combines scan, health, capability gaps, project signals, local
1770
- outcomes, Stable/Power choices, exact alternatives, risk findings, file targets,
1771
- deferred MCP/runtime steps, and the rollback point.
1772
- - Apply remains explicit. It uses one durable transaction and never turns daily
1773
- discovery into automatic installation. Existing unmanaged files are preserved.
1774
- - Non-interactive `--json`, `--yes`, risk approval, selected-agent, and project-root
1775
- behavior must be deterministic. Advanced constituent commands remain supported.
1776
- - Acceptance: a new user can preview in under one minute, understand the five most
1777
- important decisions, apply to a disposable profile, verify, and roll back.
1778
- - Completed 2026-07-16: `loadout upgrade` combines local health, explainable
1779
- scores, project signals, recommendations, immutable preparation, exact targets,
1780
- collision/risk evidence, one transaction, post-apply health, JSON, selected-agent,
1781
- custom-mode, and approval behavior. It includes local-outcome personalization,
1782
- capability gaps, deterministic alternatives, deferred MCP/runtime actions, and a
1783
- bounded-active versus disabled-library policy. A disposable network exercise
1784
- installed 30 Stable skills and rolled every managed byte back; a later fresh
1785
- preview again prepared all 30 without touching the real profile.
1786
-
1787
- - [x] `P16-07 [SOL policy, TERRA implementation]` Add an explainable Agent Health
1788
- Score and `loadout health --explain`.
1789
- - Score only evidenced dimensions: immutable provenance, license state, safety
1790
- findings, drift, duplicates, staleness, active-set capacity, native compatibility,
1791
- project relevance, benchmark evidence, local outcomes, and recoverability.
1792
- - Show the exact contribution, cap, evidence date, uncertainty, and remediation for
1793
- every point. Never estimate a post-upgrade score from unexecuted performance.
1794
- - Acceptance: deterministic fixtures cover perfect, empty, overloaded, drifted,
1795
- unlicensed, incompatible, and mixed managed/unmanaged profiles.
1796
- - Completed 2026-07-16: ten independently capped dimensions total 100; every
1797
- dimension exposes contribution, evidence, uncertainty, remediation, and evidence
1798
- coverage. Local collection uses pinned catalog/state/hash/inventory/outcome/
1799
- snapshot evidence; unavailable static-risk or benchmark evidence remains unknown
1800
- and earns zero. `health --explain` and upgrade before/after output share the policy.
1801
-
1802
- - [x] `P16-08 [TERRA]` Add a read-only skills.sh discovery connector.
1803
- - Ingest permitted public metadata and immutable GitHub source references; preserve
1804
- source attribution, observation time, ranking meaning, and telemetry uncertainty.
1805
- - Deduplicate against GitHub/Hacker News observations and the reviewed catalog.
1806
- skills.sh popularity is an install signal, not safety or performance evidence.
1807
- - Connector failure is partial and cannot block offline catalog use or installation.
1808
- - Completed 2026-07-16 against the documented API with bounded pagination,
1809
- response/time limits, attribution, rate-limit evidence, deduplication, strict
1810
- schema validation, complete-cache fallback, and `discover --source skills-sh`.
1811
- The current upstream contract requires a request-scoped Vercel OIDC token and
1812
- exposes mutable repository identity rather than a commit, so Loadout says so and
1813
- requires later immutable dossier review instead of inventing pin evidence.
1814
-
1815
- - [x] `P16-09 [TERRA]` Add an official MCP Registry discovery connector.
1816
- - Validate registry responses against a bounded schema; preserve namespace,
1817
- publication version, distribution type, repository, and verification evidence.
1818
- - Registry membership establishes identity/distribution evidence only. Loadout still
1819
- previews credentials, permissions, transports, commands, domains, and target
1820
- configurations before any MCP recipe can be admitted or applied.
1821
- - Acceptance: pagination, duplicate versions, malformed records, rate limits,
1822
- replay, and offline cache behavior have deterministic tests.
1823
- - Completed 2026-07-16 with official v0.1 cursor pagination, bounded schemas,
1824
- namespace/version/distribution/lifecycle evidence, duplicate resolution, partial
1825
- results, cursor-replay defense, complete-cache fallback, and
1826
- `discover --source mcp-registry`. Registry presence is never labeled popularity,
1827
- safety approval, or Loadout recommendation.
1828
-
1829
- - [x] `P16-10 [SOL design, TERRA implementation]` Treat Microsoft APM and OpenPackage
1830
- as interoperable inputs rather than enemies to reimplement.
1831
- - Inspect/import supported declarative manifests and lock evidence without invoking
1832
- either external CLI. Map primitives loss-reportingly into Loadout capability and
1833
- trust records; retain the original source and unsupported fields.
1834
- - Consider an optional backend adapter only after the preview, ownership,
1835
- transaction, and rollback contracts can remain true. Never claim Loadout created
1836
- or independently reviewed third-party registry evidence.
1837
- - Completed 2026-07-16: bounded read-only planners map Microsoft APM manifests/locks
1838
- and OpenPackage manifests/workspace indexes while preserving exact bytes/SHA-256,
1839
- unsupported fields, source uncertainty, and declared-but-unverified hashes.
1840
- `loadout interop apm|openpackage` never invokes an external CLI, resolves a
1841
- registry, writes, installs, or claims third-party review.
1842
-
1843
- - [x] `P16-11 [SOL design, TERRA implementation]` Add agent and model compatibility
1844
- intelligence.
1845
- - Detect installed agent CLI/application versions using bounded read-only commands
1846
- or version files with timeouts. Never start an agent session or inherit secrets.
1847
- - Maintain a signed compatibility feed for path/config/format changes, deprecated
1848
- surfaces, supported model/provider changes, and known recipe breakage.
1849
- - Add `versions` and `compatibility` output with current version, evidence source,
1850
- freshness, affected managed content, migration preview, and uncertainty.
1851
- - Acceptance: fixtures cover missing binaries, prereleases, malformed output,
1852
- timeouts, Windows executable resolution, offline state, and a breaking path change.
1853
- - Completed 2026-07-16: `loadout versions` detects installed agent CLI versions using
1854
- fixed read-only commands, a sanitized environment, five-second timeout, semantic
1855
- version parsing, explicit missing/malformed/timeout evidence, prerelease
1856
- uncertainty, and Windows executable resolution. Strict signed notices cover
1857
- freshness/offline/stale/invalid states, version ranges, affected managed install/
1858
- activation/MCP content, and approval-only migration previews. `loadout
1859
- compatibility` consumes verified intelligence without mutating agent state.
1860
-
1861
- - [ ] `P16-12 [SOL trust design, TERRA implementation]` Publish a bounded signed daily
1862
- intelligence feed.
1863
- - Generate discovery observations, compatibility notices, candidate inspection
1864
- summaries, benchmark changes, and signed catalog-release pointers centrally.
1865
- - The public feed contains no user telemetry or private repository information.
1866
- Clients verify signatures, size, schema, freshness, sequence, and replay state.
1867
- - Feed consumption never installs, promotes, updates, or executes a candidate.
1868
- Trusted catalog membership changes only in a separately reviewed signed release.
1869
- - Acceptance: local file and HTTPS preview/apply, key pinning/rotation policy,
1870
- downgrade/replay rejection, stale fallback, and compromise recovery are tested.
1871
- - Engineering complete 2026-07-16: strict public-only schemas, Ed25519 signing,
1872
- local/HTTPS preview, bounded reads, expiry, sequence high-water marks, explicit
1873
- next-key authorization, verified stale fallback, compromise reset, cache-only
1874
- apply, central discovery projection, and `loadout intelligence` are tested. Apply
1875
- cannot install, promote, update, or execute. Acceptance remains open until a human
1876
- provisions the production signing key and public host for the daily workflow.
1877
-
1878
- - [x] `P16-13 [SOL security design, TERRA implementation]` Upgrade skill security and
1879
- specification validation.
1880
- - Validate Agent Skills frontmatter, naming, size, progressive-disclosure structure,
1881
- symlinks, executable files, dependencies, remote instruction loads, domains,
1882
- environment references, Unicode controls, prompt-injection/exfiltration language,
1883
- and capability/permission declarations.
1884
- - Generate an SBOM-like inventory for executable recipes and report disagreements
1885
- between deterministic and optional model-assisted scanners instead of collapsing
1886
- them into an unjustified safe/unsafe label.
1887
- - Acceptance: malicious and benign adversarial fixtures measure false positives and
1888
- false negatives; critical findings fail closed unless a narrowly scoped explicit
1889
- override is supported and recorded.
1890
- - Completed 2026-07-16: `loadout skill-audit` validates Agent Skills metadata and
1891
- disclosure bounds, symlinks, executables, dependencies, remote loads, domains,
1892
- environment names, Unicode controls, injection/exfiltration patterns, and declared
1893
- capabilities. It emits a content-hashed SBOM-like inventory and reports assisted
1894
- scanner disagreement separately. Selected critical content fails closed; allowlist
1895
- selection happens before validation so unselected collection bytes cannot enter or
1896
- block a plan. Benign/malicious regression fixtures record expected error counts.
1897
-
1898
- - [ ] `P16-14 [TERRA, LUNA fixtures]` Create privacy-safe viral CLI artifacts.
1899
- - Add a deterministic Markdown/JSON `loadout card` with agents, active skills,
1900
- provenance coverage, health dimensions, update date, and zero project paths,
1901
- prompts, code, repository names from private sources, or secrets.
1902
- - Add `compare-loadouts` for two explicit privacy-safe reports and a static badge
1903
- endpoint specification that does not require telemetry.
1904
- - Acceptance: snapshot/redaction tests and a beginner comprehension test prove that
1905
- the artifact is useful without implying a universal quality score.
1906
- - Engineering complete 2026-07-16: deterministic Markdown/JSON `loadout card` and aggregate-only
1907
- `compare-loadouts` are implemented with redaction tests, evidence coverage, claim
1908
- boundaries, and zero project/repository/path/prompt/code/credential detail. A
1909
- telemetry-free Shields endpoint artifact covers evidence, active-skill,
1910
- managed-package, and MCP aggregates. Only the real beginner comprehension study
1911
- remains open.
1912
-
1913
- - [ ] `P16-15 [SOL design, TERRA implementation]` Generalize reviewed runtime recipes.
1914
- - Define a versioned declarative schema for exact artifacts, hashes/signatures,
1915
- dependency cutoffs, permissions, sanitized environment, fixed commands, health
1916
- checks, agent targets, timeouts, snapshot roots, removal, and supported OSes.
1917
- - Migrate Graphify without changing its current reviewed behavior. Admit another
1918
- tool only after independent usefulness, license, security, and rollback review.
1919
- - Acceptance: schema validation, Graphify parity, malicious recipe rejection,
1920
- Windows path behavior, failure rollback, and removal restoration pass.
1921
- - Engineering complete 2026-07-16: a strict v1 schema covers exact artifacts and
1922
- hashes, source/license/trust, dependency cutoffs, permissions, sanitized env,
1923
- direct commands, health checks, agent targets, timeouts, snapshots/removal, and OS
1924
- binaries. Graphify generates the prior exact plan/SKILL bytes; malicious recipe,
1925
- plan-integrity, Windows path, rollback, and removal tests pass. An independently
1926
- reviewed second tool and a live Windows install remain open.
1927
-
1928
- - [ ] `P16-16 [TERRA implementation, LUNA fixtures]` Deepen adapters only where
1929
- official documentation and user demand justify it.
1930
- - Prioritize commands/agents/rules/plugins/MCP gaps for the agents actually observed
1931
- in founder tests. Every claim needs source documentation, preservation fixtures,
1932
- transaction coverage, and a real disposable smoke test.
1933
- - Do not broaden a compatibility badge merely because a directory can be copied.
1934
- - Waiting on P16-18 evidence by design: the generic capability-gap report and
1935
- preservation/transaction fixtures exist, but no new adapter surface is claimed
1936
- until founder demand and official documentation justify it.
1937
-
1938
- - [x] `P16-17 [SOL]` Add a supply-chain and product claim review gate.
1939
- - Verify every “best,” “safe,” “compatible,” “daily,” “official,” and “supported”
1940
- statement against current stored evidence. Reject release artifacts containing
1941
- stale counts, fabricated benchmark data, unreviewed licenses, or silent execution.
1942
- - Produce a machine-readable release evidence index linking claims to tests,
1943
- immutable sources, benchmark artifacts, and human decisions.
1944
- - Completed 2026-07-16: `loadout claims` and `npm run check:evidence` emit a
1945
- machine-readable six-claim index, verify catalog counts/evidence files and
1946
- universal-best/no-benchmark boundaries, and reject unsupported release claims.
1947
-
1948
- - [ ] `P16-18 [HUMAN+SOL]` Complete founder and external product validation.
1949
- - Run the complete matrix first in disposable profiles, then with snapshots on the
1950
- founder's real Codex and Claude installations. Record confusion and time-to-value,
1951
- not only command success.
1952
- - Run at least ten external sessions across beginners, power users, one/many agents,
1953
- and Windows/macOS/Linux. Convert reproducible failures into regression tests.
1954
- - Public beta requires zero known destructive-loss defects, successful rollback,
1955
- provenance publication, license decisions, and no unsupported universal-best claim.
1956
-
1957
- #### Phase 16 execution waves
1958
-
1959
- Integration checkpoint 2026-07-16: the complete required gate passes 101 test files/
1960
- 448 tests, catalog and discovery attribution, the real CLI product flow, installed
1961
- npm-tarball smoke, and seven 1,000-skill scans at a 1.27-second p95. The optional
1962
- Chromium first-run test and `npm publish --dry-run` also pass. Live bounded smoke tests
1963
- returned current official MCP Registry records, failed skills.sh closed without its
1964
- required token, installed and rolled back all 30 Stable skills in a disposable Codex
1965
- profile, and changed no real agent profile. A fresh post-security-upgrade preview
1966
- prepared all 30 selected Stable directories from four immutable pins; its first run
1967
- exposed and then regression-tested selection-before-validation for collection repos.
1968
-
1969
- 1. **Wave A — proof foundation:** P16-01 through P16-05.
1970
- 2. **Wave B — killer first run:** P16-06 and P16-07.
1971
- 3. **Wave C — ecosystem intelligence:** P16-08 through P16-12.
1972
- 4. **Wave D — trust depth and sharing:** P16-13 through P16-17.
1973
- 5. **Wave E — validation and release:** P16-18, P12-30 through P12-34, and P15-06/07B.
1974
-
1975
- Work in later waves may scaffold interfaces in parallel, but public claims and
1976
- recommendation promotion cannot bypass earlier evidence/trust gates. More catalog
1977
- records, a new frontend, automatic candidate installation, and broad arbitrary
1978
- runtime execution are explicitly lower priority than these waves.
1979
-
1980
- ### Phase 17: Credential-aware onboarding and install/update correctness
1981
-
1982
- This phase was added after founder review identified a common misconception: paid
1983
- ChatGPT or Claude chat subscriptions are not provider API billing. Loadout must give a
1984
- useful no-key experience while making credentialed integrations impossible to apply by
1985
- accident.
1986
-
1987
- - [x] `P17-01 [SOL policy, TERRA implementation]` Define a non-secret setup access
1988
- profile for `openai`, `anthropic`, `openrouter`, `other`, or `none`.
1989
- - Interactive onboarding asks in plain language and explicitly excludes chat
1990
- subscriptions. `--api-access` accepts provider names, rejects unknown/key-like
1991
- values, and never persists a credential.
1992
- - The answer is eligibility/explanation context only; it cannot bypass trust,
1993
- compatibility, license, safety, approval, or transaction gates.
1994
- - [x] `P17-02 [TERRA]` Keep broad setup credential-free by construction.
1995
- - Stable, Power, Maximum, and Custom automatically copy only statically inspected
1996
- skill directories at immutable commits. MCP-only and executable records remain
1997
- explicit regardless of declared API access.
1998
- - Product copy explains that a skill mentioning OpenAI or Anthropic does not itself
1999
- require a model key; the specific runtime operation decides that requirement.
2000
- - [x] `P17-03 [SOL security, TERRA implementation]` Fail closed on credentialed MCP
2001
- configuration.
2002
- - A credentialed recipe cannot be applied until every required environment reference
2003
- resolves. Config output stores `${VARIABLE}` only and never a value.
2004
- - OS-keychain references are accepted only at Loadout-controlled execution
2005
- boundaries such as the explicit bounded connection verifier; arbitrary host
2006
- configs are not falsely claimed to resolve Loadout keychain entries.
2007
- - [x] `P17-04 [TERRA]` Quarantine invalid Maximum units rather than entire collections.
2008
- - Deterministic critical validation still fails closed for the unit. Safe siblings at
2009
- the same immutable source revision remain available, and every rejected unit and
2010
- reason appears in the preview.
2011
- - [x] `P17-05 [SOL invariant, TERRA implementation]` Support Maximum after Stable.
2012
- - Matching active units stay active only when the incoming and installed reviewed
2013
- commits match. Additional units enter the disabled library. Revision mismatch or a
2014
- missing active unit blocks the transaction instead of relabeling or overwriting it.
2015
- - [x] `P17-06 [SOL invariant, TERRA implementation]` Scope collection updates to the
2016
- exact managed unit set.
2017
- - Update diff, risk analysis, planning, copying, verification, and state replacement
2018
- ignore unrelated repository siblings and root files. Missing or added managed-unit
2019
- targets block instead of silently changing the active set.
2020
- - [x] `P17-07 [TERRA]` Persist static assessment evidence in install state and expose it
2021
- to Agent Health.
2022
- - Store status, finding count, assessment time, and policy version without secret or
2023
- source-content values. Recompute it on every applied revision.
2024
- - [ ] `P17-08 [HUMAN+SOL]` Complete founder verification on the published package.
2025
- - Test no-key Stable, Stable-to-Maximum overlay, quarantine output, project
2026
- optimization, scoped update, credential-gated MCP configuration, health evidence,
2027
- rollback, and clean uninstall on the founder's snapshotted real Codex profile.
2028
- - Repeat the safe no-key journey with Claude, then collect one Windows and one Linux
2029
- external session. Convert every reproducible failure into a regression test before
2030
- public-release claims.
2031
- - Partial 2026-07-17: the registry and clean external macOS install are verified for
2032
- `0.2.0`; the published CLI loads the complete command surface, reports all 50
2033
- catalog records, and completes its disposable install/rollback demo. Real founder
2034
- Stable/Power/Maximum and Claude-profile exercises remain deliberately pending.
2035
- - [x] `P17-09 [TERRA implementation, LUNA copy review]` Make Power resilient and
2036
- explain the product in beginner language before the `0.2.0` release.
2037
- - Power now quarantines a rejected selected skill while retaining safe selected
2038
- siblings from the same immutable collection, matching Maximum's unit-level
2039
- isolation without weakening Stable's exact bounded default.
2040
- - A live Codex preview prepares all eight Power collections and 50 skill directories,
2041
- quarantines six individual units, and reports every overlap and approval boundary.
2042
- - README now leads with install/mode/project/daily workflows, directly credits all
2043
- 50 upstream repositories, and distinguishes a new lead, an available update, and
2044
- a comparison-backed replacement instead of treating popularity as proof.
2045
-
2046
- Engineering checkpoint 2026-07-17: the complete local gate passes 103 test files/460
2047
- tests, packaged CLI smoke, the real CLI product journey, release-claim checks, and seven
2048
- 1,000-skill scans at a 1.88-second p95. A real read-only Maximum preview against the 50
2049
- pinned catalog records prepared 29 repositories/1,158 Codex skill directories,
2050
- quarantined 50 invalid units while retaining safe siblings, deferred 19 MCP-only
2051
- records, and resolved 44 lower-ranked overlaps. Comparing that prepared plan with the
2052
- founder's 30 active Stable units found zero missing units and zero commit mismatches;
2053
- no real agent file was mutated.
2054
-
2055
- ## 19. Seven-day schedule
2056
-
2057
- ### Day 1: Foundation
2058
-
2059
- - Repository, CI, schemas, fixtures, adapter contract, dashboard wireframe.
2060
- - Freeze MVP decisions by end of day.
2061
-
2062
- ### Day 2: Detection and catalog
2063
-
2064
- - Six agent detectors.
2065
- - Seed catalog and conflict families.
2066
- - Dashboard shell and agent status.
2067
-
2068
- ### Day 3: Installation
2069
-
2070
- - Package parsing, cache, transaction, snapshots.
2071
- - Claude and Codex skill adapters.
2072
-
2073
- ### Day 4: Breadth
2074
-
2075
- - Remaining skill adapters.
2076
- - Claude/Codex/Cursor MCP planning.
2077
- - Stable and Maximum flows.
2078
-
2079
- ### Day 5: Updates and UI
2080
-
2081
- - Update diff, sensitive-change policy, rollback.
2082
- - Finish four dashboard screens.
2083
-
2084
- ### Day 6: Verification
2085
-
2086
- - Windows/macOS/Linux tests.
2087
- - Demo mode, risky update fixture, polish.
2088
- - Put incomplete advanced capabilities behind explicit experimental flags; preserve
2089
- their backlog and code without exposing broken paths in the judge experience.
2090
-
2091
- ### Day 7: Submission
2092
-
2093
- - Fix only critical defects.
2094
- - README, video, Devpost, feedback session ID.
2095
- - Submit with several hours of buffer.
2096
-
2097
- ## 20. Definition of done
2098
-
2099
- The MVP is done only when a judge can:
2100
-
2101
- 1. Clone the repository.
2102
- 2. Run documented setup successfully.
2103
- 3. Launch Loadout without providing an account.
2104
- 4. See detected agents or use isolated demo mode.
2105
- 5. Scan existing skills and select Stable, Maximum Library, or Custom.
2106
- 6. Preview exact planned changes.
2107
- 7. Apply a real skill to at least Claude and Codex fixtures or installations.
2108
- 8. Verify unrelated configuration survives.
2109
- 9. View an update diff.
2110
- 10. See a risky update blocked.
2111
- 11. Roll back to byte-identical prior configuration.
2112
- 12. Understand supported platforms and limitations from the UI and README.
2113
-
2114
- ## 21. Required tests
2115
-
2116
- - Catalog schema validation.
2117
- - Ranking determinism.
2118
- - Conflict selection.
2119
- - Platform path resolution.
2120
- - Agent detection in fake home directories.
2121
- - Existing-skill inventory, ownership, fingerprint, duplicate, and capacity reporting.
2122
- - Skill parsing.
2123
- - MCP parsing.
2124
- - Path traversal rejection.
2125
- - Escaping-symlink rejection.
2126
- - Plan collision detection.
2127
- - Snapshot integrity.
2128
- - Interrupted transaction recovery.
2129
- - Unrelated config preservation.
2130
- - Idempotent second install.
2131
- - Tampered cache/hash rejection.
2132
- - Sensitive update classification.
2133
- - Rollback byte equality.
2134
- - Dashboard first-run flow.
2135
- - Windows, macOS, Linux CI.
2136
-
2137
- ## 22. Recorded product walkthrough
2138
-
2139
- 1. Show Claude, Codex, and Cursor with inconsistent/manual setup.
2140
- 2. Run `npx loadout-ai` after npm publication, or `npx .` from the cloned repository.
2141
- 3. Run the read-only scan and show actual skills, unmanaged content, duplicates, and
2142
- overloaded profiles without changing anything.
2143
- 4. Select Stable for daily use; show Maximum Library as an explicit stress/power-user
2144
- option rather than the default.
2145
- 5. Review repository count, actual skill count, overlaps, deferred MCP setup, and
2146
- safety findings.
2147
- 6. Approve; show the single-transaction success and restore point.
2148
- 7. Show a newly discovered Trending repository.
2149
- 8. Show a benign update and approve it.
2150
- 9. Show a second update adding a hook and external domain; Loadout blocks it.
2151
- 10. Trigger rollback and show restored healthy state.
2152
- 11. Close with supported agents, operating systems, and future catalog vision.
2153
-
2154
- ## 23. Risks and mitigations
2155
-
2156
- ### Too much platform breadth
2157
-
2158
- Mitigation: full skill support for six; MCP only for three; mark unsupported honestly.
2159
-
2160
- ### Corrupting user configuration
2161
-
2162
- Mitigation: plan-only adapters, snapshots, staging, validation, automatic restore,
2163
- fixture-based tests, and disposable test homes isolated from the real profile.
2164
-
2165
- ### Supply-chain risk
2166
-
2167
- Mitigation: curated catalog, immutable commits, hashes, no lifecycle scripts, static
2168
- checks, approval for new powers, no claims that stars imply safety.
2169
-
2170
- ### Weak differentiation from OpenPackage/skills installers
2171
-
2172
- Mitigation: lead with one-command diagnosis of what the user already has, honest
2173
- provenance, evidence-backed comparison, reviewed-library versus active-set separation,
2174
- project optimization, automatic discovery tiers, update explanation, and rollback.
2175
-
2176
- ### Competing frontend and CLI behavior
2177
-
2178
- Mitigation: keep one authoritative CLI product surface. The superseded dashboard and
2179
- its separate profile/mutation model were removed before the public release candidate.
2180
-
2181
- ### GitHub API rate limits
2182
-
2183
- Mitigation: bundled offline catalog, caching, authenticated CI bot later, graceful
2184
- stale-data indicator.
2185
-
2186
- ### Team integration failure
2187
-
2188
- Mitigation: shared schemas on day one, daily integration, ownership boundaries, CI,
2189
- small PRs, no long-lived branches.
2190
-
2191
- ## 24. Git workflow
2192
-
2193
- - Default branch: `main`.
2194
- - Branches: `codex/<short-task>` or `<member>/<short-task>`.
2195
- - One backlog ID per PR where practical.
2196
- - PR description includes task ID, test evidence, screenshots for UI, and risks.
2197
- - Rebase or update before merge.
2198
- - Do not commit secrets, tokens, generated caches, or real user configuration.
2199
- - Require one teammate review for core transaction/security changes.
2200
- - Tag demo-ready checkpoints.
2201
-
2202
- ## 25. Immediate next tasks
2203
-
2204
- Use **Launch finish line** at the top of this document. The next actions are the short
2205
- founder acceptance remainder, the recorded demo, `/feedback`, and Devpost submission.
2206
- Do not add a hosted service, provider dependency, dashboard, or speculative feature
2207
- before the deadline.