loadout-ai 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/CHANGELOG.md +148 -1
  2. package/README.md +160 -315
  3. package/SECURITY.md +21 -1
  4. package/catalog/discovered.json +31269 -26003
  5. package/catalog/packages.json +4 -4
  6. package/dist/src/cli.js +13 -0
  7. package/dist/src/commands/agents.js +3 -1
  8. package/dist/src/commands/catalog-candidate.js +150 -0
  9. package/dist/src/commands/catalog-workflows.js +147 -0
  10. package/dist/src/commands/catalog.js +9 -367
  11. package/dist/src/commands/coordinate.js +586 -0
  12. package/dist/src/commands/coordination-discussions.js +194 -0
  13. package/dist/src/commands/coordination-sessions.js +197 -0
  14. package/dist/src/commands/inventory.js +4 -2
  15. package/dist/src/core/agents/agent-inspection.js +26 -4
  16. package/dist/src/core/{routing → agents}/model-config.js +7 -2
  17. package/dist/src/core/catalog/catalog.js +1 -0
  18. package/dist/src/core/catalog/registry.js +44 -10
  19. package/dist/src/core/catalog/safety.js +36 -7
  20. package/dist/src/core/coordination/adapters/claude-code.js +154 -0
  21. package/dist/src/core/coordination/adapters/codex.js +137 -0
  22. package/dist/src/core/coordination/adapters/types.js +64 -0
  23. package/dist/src/core/coordination/auth.js +59 -0
  24. package/dist/src/core/coordination/bridge-lease.js +79 -0
  25. package/dist/src/core/coordination/conflict-preview.js +167 -0
  26. package/dist/src/core/coordination/contract-diff.js +106 -0
  27. package/dist/src/core/coordination/coordinator.js +552 -0
  28. package/dist/src/core/coordination/crash-recovery.js +169 -0
  29. package/dist/src/core/coordination/daemon.js +667 -0
  30. package/dist/src/core/coordination/discussion.js +340 -0
  31. package/dist/src/core/coordination/events.js +161 -0
  32. package/dist/src/core/coordination/http-api.js +118 -0
  33. package/dist/src/core/coordination/interrupt-policy.js +89 -0
  34. package/dist/src/core/coordination/lock.js +110 -0
  35. package/dist/src/core/coordination/mcp-server.js +424 -0
  36. package/dist/src/core/coordination/redaction.js +85 -0
  37. package/dist/src/core/coordination/replay.js +188 -0
  38. package/dist/src/core/coordination/retention.js +128 -0
  39. package/dist/src/core/coordination/runtime.js +21 -0
  40. package/dist/src/core/coordination/session-manager.js +366 -0
  41. package/dist/src/core/coordination/watcher.js +128 -0
  42. package/dist/src/core/{routing → delegation}/first-party-skills.js +25 -13
  43. package/dist/src/core/{routing → delegation}/handoff.js +130 -63
  44. package/dist/src/core/discovery/candidate-intelligence-evidence.js +108 -0
  45. package/dist/src/core/discovery/candidate-intelligence-types.js +1 -0
  46. package/dist/src/core/discovery/candidate-intelligence-validation.js +108 -0
  47. package/dist/src/core/discovery/candidate-intelligence.js +3 -214
  48. package/dist/src/core/discovery/community.js +12 -4
  49. package/dist/src/core/discovery/github-discovery.js +6 -2
  50. package/dist/src/core/discovery/private-discovery.js +6 -2
  51. package/dist/src/core/install/reconcile.js +1 -1
  52. package/dist/src/core/install/source.js +21 -7
  53. package/dist/src/core/install/uninstall.js +39 -3
  54. package/dist/src/core/reporting/cli-guide.js +10 -6
  55. package/dist/src/core/reporting/completion.js +61 -109
  56. package/dist/src/core/reporting/doctor.js +4 -6
  57. package/dist/src/core/runtime/bounded-json.js +65 -0
  58. package/dist/src/core/runtime/github.js +30 -16
  59. package/dist/src/core/runtime/mcp-recipes.js +1 -1
  60. package/dist/src/core/workspace/active-policy.js +1 -1
  61. package/docs/CANDIDATE_INTELLIGENCE.md +9 -2
  62. package/docs/CATALOG.md +1 -1
  63. package/docs/CREDENTIAL_AND_UPDATE_POLICY.md +1 -1
  64. package/docs/DEMO_SCRIPT.md +19 -23
  65. package/docs/DISCOVERED.md +251 -250
  66. package/docs/FEATURE_TEST_MATRIX.md +10 -263
  67. package/docs/GITHUB_AUTHORIZATION.md +5 -0
  68. package/docs/LIVE_COLLABORATION.md +234 -0
  69. package/docs/PROVENANCE_AND_COMPARISON.md +1 -1
  70. package/docs/REFERENCE.md +163 -0
  71. package/docs/RELEASE_REVIEW.md +1 -2
  72. package/docs/TESTING.md +17 -1
  73. package/docs/USER_TEST_GUIDE.md +138 -3
  74. package/docs/assets/loadout-discover-activate.webp +0 -0
  75. package/docs/assets/loadout-handoff-coordinate.webp +0 -0
  76. package/docs/assets/loadout-social-preview.png +0 -0
  77. package/docs/decisions/001-coordination-jsonl-locking.md +35 -0
  78. package/docs/decisions/002-local-daemon-authentication.md +32 -0
  79. package/docs/decisions/003-bounded-agent-discussions.md +91 -0
  80. package/docs/specs/BOUNDED_AGENT_DISCUSSIONS.md +169 -0
  81. package/docs/superpowers/plans/2026-09-03-coordination-hardening.md +231 -0
  82. package/docs/superpowers/plans/2026-09-03-release-readiness.md +232 -0
  83. package/docs/superpowers/plans/2026-09-04-bounded-agent-discussions.md +121 -0
  84. package/package.json +18 -7
  85. package/skills/loadout-handoff/SKILL.md +238 -0
  86. package/MASTER_PLAN.md +0 -2207
  87. package/dist/src/core/routing/route.js +0 -539
  88. package/docs/ACTIVE_SET.md +0 -53
  89. package/docs/COMPATIBILITY_POLICY.md +0 -22
  90. package/docs/CONVERSION_AND_SANDBOX.md +0 -27
  91. package/docs/EVALUATION_PROTOCOL_V1.md +0 -300
  92. package/docs/HEAD_TO_HEAD_EVALUATION.md +0 -79
  93. package/docs/PROVIDER_CONFIGURATION.md +0 -45
  94. package/docs/README_RESEARCH.md +0 -36
  95. package/docs/REPOSITORY_STABILIZATION.md +0 -190
  96. package/docs/SAFE_UPDATE_DEMO.md +0 -25
  97. package/docs/SCHEMA_DECISIONS.md +0 -25
  98. package/docs/SUBMISSION_COPY.md +0 -90
  99. package/docs/TEAM_POLICY.md +0 -18
  100. package/docs/assets/loadout-workflow.png +0 -0
  101. package/docs/superpowers/plans/2026-07-19-relatable-readme-hero.md +0 -283
  102. package/docs/superpowers/plans/2026-07-20-loadout-readme-explainer.md +0 -116
  103. package/docs/superpowers/plans/2026-07-20-project-activation-safety.md +0 -469
  104. package/docs/superpowers/specs/2026-07-19-relatable-readme-hero-design.md +0 -80
  105. package/docs/superpowers/specs/2026-07-20-loadout-readme-explainer-design.md +0 -55
  106. package/docs/superpowers/specs/2026-07-20-project-activation-safety-design.md +0 -228
  107. package/skills/loadout-router/SKILL.md +0 -120
  108. /package/dist/src/core/{routing → agents}/credentials.js +0 -0
@@ -1,228 +0,0 @@
1
- # Project Activation Safety and Relevance Design
2
-
3
- **Status:** Approved direction; awaiting written-spec review
4
- **Date:** 2026-07-20
5
- **Release target:** Loadout 0.4.1
6
-
7
- ## Problem
8
-
9
- The founder acceptance run of `loadout activate --project . --agents
10
- codex,claude-code --limit 30` exposed three connected defects:
11
-
12
- 1. The preview reported `0/30 active` even though Claude Code already had 12
13
- unmanaged skills. The planner counted only active Loadout records, so applying the
14
- proposal could exceed the disclosed per-agent limit.
15
- 2. An earlier rollback correctly restored recursively empty skill directories. The
16
- activation planner treated the existence of those empty directories as occupied
17
- content and blocked safe activation for both agents.
18
- 3. Project detection recognized only JavaScript/TypeScript and Playwright. It missed
19
- strong local evidence that this repository is a Node CLI, an npm package, a Vitest
20
- project, an MCP-related tool, and a release/supply-chain project. Ranking then used
21
- the entire 30-slot budget as a quota and proposed redundant or mismatched skills,
22
- including several overlapping Playwright workflows and a Jest skill for a Vitest
23
- repository.
24
-
25
- No activation was applied. The Maximum library remains disabled, and the founder's
26
- 12 unmanaged Claude skills remain unchanged.
27
-
28
- ## Goals
29
-
30
- - Make `--limit` a truthful ceiling over all active skills visible to each selected
31
- agent, whether managed by Loadout or not.
32
- - Plan and explain capacity separately for every selected agent.
33
- - Treat recursively empty rollback residue as unoccupied while continuing to refuse
34
- every non-empty, symlinked, unreadable, or unsupported target.
35
- - Improve local project recognition enough to distinguish this Node CLI/npm/Vitest
36
- repository from a generic Playwright web application.
37
- - Select a compact, diverse working set instead of filling every available slot with
38
- weak or redundant matches.
39
- - Preserve a preview-first, transactional, rollback-safe mutation path.
40
-
41
- ## Non-goals
42
-
43
- - No model or external API call is required for recommendation or activation.
44
- - No project source, filename, dependency, or outcome data leaves the machine.
45
- - This work does not claim that a deterministic recommendation proves global package
46
- quality.
47
- - This work does not install deferred MCP servers or collect credentials.
48
- - Dashboard removal remains a separate 0.4.1 task.
49
-
50
- ## Design
51
-
52
- ### 1. One shared definition of an occupied skill target
53
-
54
- Loadout will use one filesystem predicate for both initial installation and later
55
- library activation:
56
-
57
- - A missing path is unoccupied.
58
- - A real directory containing no entries at any depth is unoccupied.
59
- - A regular file, symlink, special entry, unreadable directory, or directory with any
60
- non-directory descendant is occupied.
61
- - Recursive inspection is bounded to 10,000 entries. Reaching the bound is treated as
62
- occupied, never as empty.
63
-
64
- Activation will re-check this predicate inside the transaction immediately before
65
- copying. This closes the preview/apply race without deleting or adopting user content.
66
- Empty target directories may be removed immediately before the copy because they
67
- contain no bytes to preserve; rollback still restores the pre-transaction topology.
68
-
69
- ### 2. Per-agent total active capacity
70
-
71
- For each requested agent, the planner will scan the agent's real skill root and count
72
- every directory containing a valid `SKILL.md`. This inventory already distinguishes
73
- managed and unmanaged skills and will be the capacity source of truth.
74
-
75
- For agent `a`:
76
-
77
- ```text
78
- available(a) = max(0, limit - inventory(a).total)
79
- ```
80
-
81
- Candidate ranking is deterministic and shared, but selection is sliced independently
82
- to each agent's available capacity. In the founder fixture, Claude Code has 12 active
83
- unmanaged skills and may receive at most 18 additions; Codex has zero active skills and
84
- may receive at most 30. An agent at its limit receives no changes and a clear warning;
85
- it does not prevent another requested agent from using its own available capacity.
86
-
87
- The plan must never count an empty directory as an active skill. Existing managed
88
- skills already in the desired state count toward capacity but are not proposed again.
89
-
90
- ### 3. Project signals
91
-
92
- Project scanning remains deterministic, bounded, and local-only. It will read known
93
- root manifests/configuration and a bounded set of repository filenames. For Node
94
- projects it will derive explicit signals from `package.json`:
95
-
96
- - `bin` -> Node CLI
97
- - package name plus `publishConfig` or non-private package -> npm package/release
98
- - `vitest` -> Vitest
99
- - `jest` -> Jest
100
- - `@playwright/test` or Playwright config -> Playwright
101
- - `commander` -> command-line application
102
- - `zod` -> schema validation
103
- - scripts containing `prepack`, package-smoke, or release checks -> package/release
104
-
105
- Repository filenames add delivery, security, MCP, and documentation signals only when
106
- matching explicit reviewed patterns such as `.github/workflows`, `SECURITY.md`, MCP-
107
- named modules, and package/release scripts. The scanner does not recursively read
108
- arbitrary project source or send any data elsewhere.
109
-
110
- Human output will name the important detected roles, for example:
111
-
112
- ```text
113
- Detected: TypeScript, Node CLI, npm package, Vitest, Playwright, MCP tooling
114
- ```
115
-
116
- ### 4. Relevant and diverse selection
117
-
118
- The existing source-priority tie-break remains, but a capacity ceiling is not a quota.
119
- Selection stops when no candidate reaches the evidence threshold.
120
-
121
- Ranking will:
122
-
123
- - reward exact tool and role matches such as Vitest, CLI design, npm packaging,
124
- TypeScript, MCP security, release verification, and documentation;
125
- - reject framework mismatches such as Jest-only guidance when Vitest is detected and
126
- Jest is not;
127
- - group close alternatives into capability families such as browser testing, code
128
- review, documentation, planning, security, and architecture;
129
- - select the highest-evidence representative before taking another member of the same
130
- family;
131
- - cap weak generic foundation choices so they cannot crowd out project-specific
132
- evidence;
133
- - continue to honor explicit `--pin` selectors, while disclosing when a pin consumes a
134
- slot or conflicts with the limit.
135
-
136
- The same skill name from multiple repositories remains an alternative, not two active
137
- copies. Multiple distinct skills in one family are allowed only when each has separate
138
- strong project evidence.
139
-
140
- ### 5. Recommendation type clarity
141
-
142
- Package-level `loadout recommend` output will label every suggestion as one of:
143
-
144
- - `skill library` — eligible for disabled-library activation;
145
- - `MCP/runtime setup` — requires a separate explicit preview and may require a
146
- non-model credential;
147
- - `unavailable` — not prepared locally, with the reason.
148
-
149
- Playwright MCP and GitHub MCP must not look like ordinary automatically activatable
150
- skills. The command will show the next preview command for explicit integrations and
151
- will continue to state that recommendations are rule-selected, not proof of quality.
152
-
153
- ### 6. Preview and apply output
154
-
155
- The default human preview will begin with one compact block per agent:
156
-
157
- ```text
158
- Claude Code: 12 active (0 managed, 12 unmanaged), 18/30 slots available
159
- Codex: 0 active, 30/30 slots available
160
- ```
161
-
162
- It will then show additions grouped by agent and capability, followed by blockers. A
163
- recursively empty target is not printed as a blocker. A real conflict prints the exact
164
- target and a non-destructive next action. JSON output retains full scores, reasons,
165
- alternatives, per-agent budgets, targets, and blocker details.
166
-
167
- `--yes` is rejected when any proposed target is blocked. A successful apply creates
168
- one transaction snapshot covering every selected agent and prints the exact rollback
169
- command.
170
-
171
- ## Data flow
172
-
173
- 1. Scan project metadata into local project signals.
174
- 2. Scan each requested agent's active skill inventory.
175
- 3. Read reviewed disabled-library records and local-only outcomes.
176
- 4. Score and diversify eligible candidates.
177
- 5. Slice the ordered candidates independently to each agent's available capacity.
178
- 6. Build exact activation changes and inspect targets with the shared occupancy rule.
179
- 7. Print the preview; stop unless `--yes` is present.
180
- 8. Inside the transaction, re-read state, re-check capacity and target occupancy, copy
181
- reviewed library bytes, fingerprint the result, persist activation state, and emit
182
- the rollback snapshot.
183
-
184
- ## Failure handling
185
-
186
- - If inventory scanning fails for an agent, activation for that agent is blocked; its
187
- capacity is never guessed.
188
- - If the project manifest is malformed, show the exact manifest error and continue
189
- only with signals that remain trustworthy.
190
- - If state or filesystem contents change between preview and apply, abort the entire
191
- multi-agent transaction without partial activation.
192
- - If fewer candidates meet the evidence threshold than available slots, activate only
193
- those candidates and explain that unused capacity is intentional.
194
-
195
- ## Tests and acceptance criteria
196
-
197
- Automated tests must prove:
198
-
199
- 1. A recursively empty target does not block preview or apply.
200
- 2. A target containing a file, symlink, unreadable entry, or more than the inspection
201
- bound remains blocked and unchanged.
202
- 3. Empty-directory handling is identical in initial setup and project activation.
203
- 4. Twelve unmanaged Claude skills plus an 18-skill proposal never exceeds a limit of 30.
204
- 5. The same plan may select 18 additions for Claude and 30 for an empty Codex profile.
205
- 6. An agent already at capacity receives no additions while another agent can proceed.
206
- 7. Apply re-checks inventory and occupancy and aborts atomically when either changes.
207
- 8. The Loadout repository fixture detects Node CLI, npm package, Vitest, Playwright,
208
- Commander, Zod, release, security, and MCP signals.
209
- 9. A Vitest-only fixture never recommends a Jest-only skill.
210
- 10. Redundant Playwright/browser candidates do not consume most of the active set.
211
- 11. Recommendation output distinguishes skill libraries from explicit MCP/runtime
212
- setup.
213
- 12. Human output shows per-agent managed/unmanaged counts and unused capacity; JSON
214
- exposes the same facts structurally.
215
- 13. Existing Stable, Power, Maximum, rollback, and non-overwrite tests remain green.
216
-
217
- Founder acceptance resumes only after the published release candidate previews a
218
- non-blocked plan, reports Claude's existing 12 skills, respects both per-agent limits,
219
- applies transactionally, and restores its explicit snapshot without changing unmanaged
220
- content.
221
-
222
- ## Compatibility
223
-
224
- Existing `loadout activate` and `loadout optimize` flags remain valid. The meaning of
225
- `--limit` is corrected from an implicit managed-only count to the documented total
226
- active skill ceiling. JSON consumers must migrate from one global `activeBefore` and
227
- `capacity` pair to per-agent budget records; the 0.4.1 changelog will call out this
228
- schema correction.
@@ -1,120 +0,0 @@
1
- ---
2
- name: loadout-router
3
- description: Choose the right model tier and agent for a coding task, and hand work off between Claude Code and Codex. Use when the user asks which model to use, mentions running low on usage or quota, wants to save tokens or cost, asks whether to switch to Opus/Sonnet/Haiku or a GPT tier, or wants to delegate a task to another agent.
4
- ---
5
-
6
- # Loadout Router
7
-
8
- Route each coding task to the cheapest model that still does it well, and hand
9
- work between agents when a different one is better suited.
10
-
11
- This skill wraps the `loadout` CLI, so the model catalog and pricing stay
12
- current with the installed version rather than going stale in this file.
13
-
14
- ## Prerequisite
15
-
16
- Check once per session:
17
-
18
- ```bash
19
- loadout --version
20
- ```
21
-
22
- If that fails, tell the user to install it (`npm install --global loadout-ai`)
23
- and answer from general knowledge instead of guessing at specifics.
24
-
25
- ## Choosing a model
26
-
27
- Run the router with the task described in plain words:
28
-
29
- ```bash
30
- loadout route <task description>
31
- ```
32
-
33
- It classifies the task into one of six phases — plan, implement, review, test,
34
- debug, document — and prints the recommended tier, the models in that tier with
35
- current prices, which of the user's installed agents can run them, and a cheaper
36
- fallback with its tradeoff.
37
-
38
- Report the recommendation and the reason. Name the actual model, not just the
39
- tier. If the output lists a `Conserve:` alternative, mention it only when the
40
- user cares about cost or quota, otherwise it is noise.
41
-
42
- When the user is explicit about the phase, skip classification:
43
-
44
- ```bash
45
- loadout route --phase review
46
- ```
47
-
48
- ## When the user is low on usage
49
-
50
- If the user mentions running out, being rate limited, conserving quota, or
51
- stretching a plan, add `--conserve`:
52
-
53
- ```bash
54
- loadout route --conserve <task description>
55
- ```
56
-
57
- This drops each phase one tier and prints the tradeoff you are accepting. Say
58
- what is being given up — "shallower architectural reasoning, so review the plan
59
- more carefully" — rather than presenting it as a free win.
60
-
61
- Neither Claude Code nor Codex exposes remaining quota programmatically, so never
62
- claim to know how much the user has left. `--conserve` is a user-driven choice,
63
- not a measurement.
64
-
65
- ## Comparing models and cost
66
-
67
- ```bash
68
- loadout route --models
69
- loadout route --models --provider anthropic
70
- loadout route --models --tier fast
71
- loadout route --cost
72
- ```
73
-
74
- Use these when the user asks what is available or what something costs. Prices
75
- are per-million-token list rates; actual spend depends on prompt size, so give
76
- ratios ("roughly 4x cheaper") rather than predicting a dollar total.
77
-
78
- ## Handing work to another agent
79
-
80
- When a different agent suits the task better — or the user asks to delegate —
81
- use the handoff log. Check it is set up:
82
-
83
- ```bash
84
- loadout handoff status
85
- ```
86
-
87
- If uninitialized, run `loadout handoff init` first. Then send the task:
88
-
89
- ```bash
90
- loadout handoff send codex "write unit tests for the auth module" --context "see src/auth.ts"
91
- ```
92
-
93
- Put anything the receiving agent needs into `--context`: file paths, the
94
- decision you already made, what you deliberately left out. The other agent has
95
- none of this conversation.
96
-
97
- Sending a task is a real side effect on a shared file. Confirm with the user
98
- before sending unless they asked for the handoff themselves.
99
-
100
- ## Reading your own inbox
101
-
102
- At the start of a session, and after finishing a task, check whether another
103
- agent left you work:
104
-
105
- ```bash
106
- loadout handoff inbox claude-code
107
- ```
108
-
109
- If it lists tasks, work them in order and run the `loadout handoff done <id>`
110
- command it prints for each one. If it reports none, continue as normal and do
111
- not mention it.
112
-
113
- ## Judgment this skill does not replace
114
-
115
- - A "simple" task in an unfamiliar or high-risk area still deserves a stronger
116
- model. The classifier reads keywords, not stakes.
117
- - Security-sensitive, migration, and data-loss paths are worth frontier tier
118
- regardless of what phase they classify as.
119
- - If the user has already chosen a model, do not argue unless the choice is
120
- clearly wrong for the work.