loadout-ai 0.4.0 → 0.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +25 -0
- package/MASTER_PLAN.md +94 -3
- package/README.md +2 -2
- package/dist/src/cli.js +1 -1
- package/dist/src/core/active-policy.js +172 -38
- package/dist/src/core/active-set.js +13 -4
- package/dist/src/core/install.js +3 -36
- package/dist/src/core/recommend.js +95 -9
- package/dist/src/core/target-occupancy.js +50 -0
- package/docs/USER_TEST_GUIDE.md +17 -5
- package/docs/superpowers/plans/2026-07-20-project-activation-safety.md +469 -0
- package/docs/superpowers/specs/2026-07-20-project-activation-safety-design.md +228 -0
- package/package.json +1 -1
|
@@ -0,0 +1,469 @@
|
|
|
1
|
+
# Project Activation Safety Implementation Plan
|
|
2
|
+
|
|
3
|
+
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
|
4
|
+
|
|
5
|
+
**Goal:** Make project activation respect every agent's real active-skill capacity, tolerate recursively empty rollback residue, and recommend a compact project-relevant set.
|
|
6
|
+
|
|
7
|
+
**Architecture:** Extract the existing bounded target-occupancy rule into one shared filesystem module used by setup and activation. Build per-agent budgets from the read-only installed-skill inventory, then score and slice reviewed library candidates per agent. Extend deterministic local project signals and recommendation metadata without introducing network or model calls.
|
|
8
|
+
|
|
9
|
+
**Tech Stack:** TypeScript 5.7, Node.js 20+, Commander 12, Vitest 4, existing Loadout transaction/state/inventory modules.
|
|
10
|
+
|
|
11
|
+
## Global Constraints
|
|
12
|
+
|
|
13
|
+
- No model or external API call is required for recommendation or activation.
|
|
14
|
+
- No project source, filename, dependency, or outcome data leaves the machine.
|
|
15
|
+
- Never execute package code while scanning, recommending, or activating.
|
|
16
|
+
- Missing and recursively empty targets are unoccupied; files, symlinks, special entries, unreadable directories, and scans beyond 10,000 entries are occupied.
|
|
17
|
+
- `--limit` is a per-agent ceiling over managed plus unmanaged skills containing `SKILL.md`.
|
|
18
|
+
- Preview remains read-only; apply revalidates inside one rollback-safe transaction.
|
|
19
|
+
- Existing CLI flags remain valid; the project-plan JSON schema may add per-agent budgets in 0.4.1.
|
|
20
|
+
|
|
21
|
+
---
|
|
22
|
+
|
|
23
|
+
### Task 1: Share the bounded target-occupancy rule
|
|
24
|
+
|
|
25
|
+
**Files:**
|
|
26
|
+
|
|
27
|
+
- Create: `src/core/target-occupancy.ts`
|
|
28
|
+
- Create: `tests/target-occupancy.test.ts`
|
|
29
|
+
- Modify: `src/core/install.ts`
|
|
30
|
+
- Modify: `src/core/active-set.ts`
|
|
31
|
+
- Modify: `tests/active-set.test.ts`
|
|
32
|
+
|
|
33
|
+
**Interfaces:**
|
|
34
|
+
|
|
35
|
+
- Produces: `inspectTargetOccupancy(path: string, maximumEntries?: number): Promise<TargetOccupancy>`.
|
|
36
|
+
- Consumed by: setup collision checks and activation preview/apply checks.
|
|
37
|
+
|
|
38
|
+
- [x] **Step 1: Write failing occupancy tests**
|
|
39
|
+
|
|
40
|
+
```ts
|
|
41
|
+
import { mkdir, mkdtemp, symlink, writeFile } from "node:fs/promises";
|
|
42
|
+
import { tmpdir } from "node:os";
|
|
43
|
+
import { join } from "node:path";
|
|
44
|
+
import { expect, it } from "vitest";
|
|
45
|
+
import { inspectTargetOccupancy } from "../src/core/target-occupancy.js";
|
|
46
|
+
|
|
47
|
+
it("treats missing and recursively empty targets as unoccupied", async () => {
|
|
48
|
+
const root = await mkdtemp(join(tmpdir(), "loadout-target-"));
|
|
49
|
+
const empty = join(root, "skill");
|
|
50
|
+
await mkdir(join(empty, "nested"), { recursive: true });
|
|
51
|
+
await expect(
|
|
52
|
+
inspectTargetOccupancy(join(root, "missing")),
|
|
53
|
+
).resolves.toMatchObject({ occupied: false });
|
|
54
|
+
await expect(inspectTargetOccupancy(empty)).resolves.toMatchObject({
|
|
55
|
+
occupied: false,
|
|
56
|
+
});
|
|
57
|
+
});
|
|
58
|
+
|
|
59
|
+
it("treats content, symlinks, and the inspection bound as occupied", async () => {
|
|
60
|
+
const root = await mkdtemp(join(tmpdir(), "loadout-target-"));
|
|
61
|
+
const content = join(root, "content");
|
|
62
|
+
const linked = join(root, "linked");
|
|
63
|
+
await mkdir(content);
|
|
64
|
+
await writeFile(join(content, "SKILL.md"), "content");
|
|
65
|
+
await symlink(content, linked);
|
|
66
|
+
await expect(inspectTargetOccupancy(content)).resolves.toMatchObject({
|
|
67
|
+
occupied: true,
|
|
68
|
+
reason: "content",
|
|
69
|
+
});
|
|
70
|
+
await expect(inspectTargetOccupancy(linked)).resolves.toMatchObject({
|
|
71
|
+
occupied: true,
|
|
72
|
+
reason: "symlink",
|
|
73
|
+
});
|
|
74
|
+
await expect(inspectTargetOccupancy(join(root), 1)).resolves.toMatchObject({
|
|
75
|
+
occupied: true,
|
|
76
|
+
reason: "inspection-limit",
|
|
77
|
+
});
|
|
78
|
+
});
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
- [x] **Step 2: Run the occupancy tests and verify RED**
|
|
82
|
+
|
|
83
|
+
Run: `npx vitest run tests/target-occupancy.test.ts`
|
|
84
|
+
Expected: FAIL because `src/core/target-occupancy.ts` does not exist.
|
|
85
|
+
|
|
86
|
+
- [x] **Step 3: Implement the shared predicate**
|
|
87
|
+
|
|
88
|
+
```ts
|
|
89
|
+
import { lstat, readdir } from "node:fs/promises";
|
|
90
|
+
import { join } from "node:path";
|
|
91
|
+
|
|
92
|
+
export interface TargetOccupancy {
|
|
93
|
+
occupied: boolean;
|
|
94
|
+
reason?:
|
|
95
|
+
"content" | "symlink" | "unsupported" | "unreadable" | "inspection-limit";
|
|
96
|
+
}
|
|
97
|
+
|
|
98
|
+
export async function inspectTargetOccupancy(
|
|
99
|
+
path: string,
|
|
100
|
+
maximumEntries = 10_000,
|
|
101
|
+
): Promise<TargetOccupancy> {
|
|
102
|
+
let root;
|
|
103
|
+
try {
|
|
104
|
+
root = await lstat(path);
|
|
105
|
+
} catch (error) {
|
|
106
|
+
if (
|
|
107
|
+
error &&
|
|
108
|
+
typeof error === "object" &&
|
|
109
|
+
"code" in error &&
|
|
110
|
+
(error as { code?: string }).code === "ENOENT"
|
|
111
|
+
)
|
|
112
|
+
return { occupied: false };
|
|
113
|
+
return { occupied: true, reason: "unreadable" };
|
|
114
|
+
}
|
|
115
|
+
if (root.isSymbolicLink()) return { occupied: true, reason: "symlink" };
|
|
116
|
+
if (!root.isDirectory()) return { occupied: true, reason: "unsupported" };
|
|
117
|
+
const queue = [path];
|
|
118
|
+
let inspected = 0;
|
|
119
|
+
while (queue.length) {
|
|
120
|
+
const directory = queue.pop()!;
|
|
121
|
+
let entries;
|
|
122
|
+
try {
|
|
123
|
+
entries = await readdir(directory, { withFileTypes: true });
|
|
124
|
+
} catch {
|
|
125
|
+
return { occupied: true, reason: "unreadable" };
|
|
126
|
+
}
|
|
127
|
+
for (const entry of entries) {
|
|
128
|
+
inspected += 1;
|
|
129
|
+
if (inspected > maximumEntries)
|
|
130
|
+
return { occupied: true, reason: "inspection-limit" };
|
|
131
|
+
if (entry.isDirectory() && !entry.isSymbolicLink())
|
|
132
|
+
queue.push(join(directory, entry.name));
|
|
133
|
+
else
|
|
134
|
+
return {
|
|
135
|
+
occupied: true,
|
|
136
|
+
reason: entry.isSymbolicLink() ? "symlink" : "content",
|
|
137
|
+
};
|
|
138
|
+
}
|
|
139
|
+
}
|
|
140
|
+
return { occupied: false };
|
|
141
|
+
}
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
- [x] **Step 4: Replace duplicate setup logic and activation `pathExists` checks**
|
|
145
|
+
|
|
146
|
+
In `src/core/install.ts`, import `inspectTargetOccupancy` and replace the recursive block in `assertActiveTargetsUnoccupied` with:
|
|
147
|
+
|
|
148
|
+
```ts
|
|
149
|
+
for (const target of targets) {
|
|
150
|
+
if ((await inspectTargetOccupancy(target)).occupied) occupied.push(target);
|
|
151
|
+
}
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
In `src/core/active-set.ts`, use the same predicate when enabling and include the reason in the blocker. Re-run the check immediately before the transaction copies any target; remove only targets proven recursively empty.
|
|
155
|
+
|
|
156
|
+
- [x] **Step 5: Add the activation regression**
|
|
157
|
+
|
|
158
|
+
Extend `tests/active-set.test.ts` so a disabled library entry with `empty/nested` under its active target previews without blockers and applies, while a target containing `notes.txt` remains blocked and unchanged.
|
|
159
|
+
|
|
160
|
+
- [x] **Step 6: Verify GREEN and commit**
|
|
161
|
+
|
|
162
|
+
Run: `npx vitest run tests/target-occupancy.test.ts tests/install.test.ts tests/active-set.test.ts`
|
|
163
|
+
Expected: all tests pass.
|
|
164
|
+
|
|
165
|
+
```bash
|
|
166
|
+
git add src/core/target-occupancy.ts src/core/install.ts src/core/active-set.ts tests/target-occupancy.test.ts tests/active-set.test.ts
|
|
167
|
+
git commit -m "fix: share safe target occupancy checks"
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
### Task 2: Enforce per-agent total active capacity
|
|
171
|
+
|
|
172
|
+
**Files:**
|
|
173
|
+
|
|
174
|
+
- Modify: `src/core/active-policy.ts`
|
|
175
|
+
- Modify: `src/core/active-set.ts`
|
|
176
|
+
- Modify: `tests/active-policy.test.ts`
|
|
177
|
+
|
|
178
|
+
**Interfaces:**
|
|
179
|
+
|
|
180
|
+
- Consumes: `detectAgents()`, `scanInstalledSkills()`, and reviewed activation records.
|
|
181
|
+
- Produces: `AgentActiveSetPlan` records with total, managed, unmanaged, capacity, and selected candidates.
|
|
182
|
+
|
|
183
|
+
- [x] **Step 1: Write failing mixed-agent capacity tests**
|
|
184
|
+
|
|
185
|
+
Add a fixture with 12 unmanaged `SKILL.md` directories under Claude's skill root, zero under Codex, and 40 reviewed disabled candidates per agent. Assert:
|
|
186
|
+
|
|
187
|
+
```ts
|
|
188
|
+
const plan = await planProjectActivation(project, {
|
|
189
|
+
agents: ["claude-code", "codex"],
|
|
190
|
+
limit: 30,
|
|
191
|
+
});
|
|
192
|
+
expect(
|
|
193
|
+
plan.agentPlans.find((item) => item.agent === "claude-code"),
|
|
194
|
+
).toMatchObject({
|
|
195
|
+
activeBefore: 12,
|
|
196
|
+
unmanagedBefore: 12,
|
|
197
|
+
capacity: 18,
|
|
198
|
+
});
|
|
199
|
+
expect(
|
|
200
|
+
plan.agentPlans.find((item) => item.agent === "claude-code")!.selected,
|
|
201
|
+
).toHaveLength(18);
|
|
202
|
+
expect(plan.agentPlans.find((item) => item.agent === "codex")).toMatchObject({
|
|
203
|
+
activeBefore: 0,
|
|
204
|
+
capacity: 30,
|
|
205
|
+
});
|
|
206
|
+
expect(
|
|
207
|
+
plan.agentPlans.find((item) => item.agent === "codex")!.selected,
|
|
208
|
+
).toHaveLength(30);
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
Add a second test where Claude already has 30 unmanaged skills and Codex has none. Claude must receive zero additions and Codex must still receive candidates.
|
|
212
|
+
|
|
213
|
+
- [x] **Step 2: Run the capacity tests and verify RED**
|
|
214
|
+
|
|
215
|
+
Run: `npx vitest run tests/active-policy.test.ts`
|
|
216
|
+
Expected: FAIL because the plan has one managed-only global budget and no `agentPlans`.
|
|
217
|
+
|
|
218
|
+
- [x] **Step 3: Add per-agent plan types and inventory-backed budgets**
|
|
219
|
+
|
|
220
|
+
```ts
|
|
221
|
+
export interface AgentActiveSetPlan {
|
|
222
|
+
agent: AgentId;
|
|
223
|
+
activeBefore: number;
|
|
224
|
+
managedBefore: number;
|
|
225
|
+
unmanagedBefore: number;
|
|
226
|
+
capacity: number;
|
|
227
|
+
selected: ActiveSetCandidate[];
|
|
228
|
+
}
|
|
229
|
+
|
|
230
|
+
export interface ProjectActiveSetPlan {
|
|
231
|
+
project: ProjectSignals;
|
|
232
|
+
limit: number;
|
|
233
|
+
agents?: AgentId[];
|
|
234
|
+
agentPlans: AgentActiveSetPlan[];
|
|
235
|
+
activation?: ActivationPlan;
|
|
236
|
+
warnings: string[];
|
|
237
|
+
}
|
|
238
|
+
```
|
|
239
|
+
|
|
240
|
+
Resolve requested agents through `detectAgents`, call `scanInstalledSkills`, and create one budget from each inventory summary. Score only that agent's reviewed disabled records, diversify them, and slice to that agent's capacity.
|
|
241
|
+
|
|
242
|
+
- [x] **Step 4: Merge exact per-agent activation plans**
|
|
243
|
+
|
|
244
|
+
Call `planActivationChange("enable", selectors, { agents: [agent] })` once per non-empty agent selection and combine `changes`, `skipped`, `warnings`, `blocked`, and unique package IDs into one transaction plan. `applyProjectActivation` continues to call `applyActivationChange` exactly once.
|
|
245
|
+
|
|
246
|
+
Before copying, re-scan inventory and abort if any agent's current total would make the planned additions exceed `limit`.
|
|
247
|
+
|
|
248
|
+
- [x] **Step 5: Format truthful per-agent budgets**
|
|
249
|
+
|
|
250
|
+
Replace the global budget line with:
|
|
251
|
+
|
|
252
|
+
```text
|
|
253
|
+
Claude Code: 12 active (0 managed, 12 unmanaged); 18/30 slots available
|
|
254
|
+
Codex: 0 active (0 managed, 0 unmanaged); 30/30 slots available
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
Group additions beneath their target agent and do not duplicate one global list that implies identical capacities.
|
|
258
|
+
|
|
259
|
+
- [x] **Step 6: Verify GREEN and commit**
|
|
260
|
+
|
|
261
|
+
Run: `npx vitest run tests/active-policy.test.ts tests/active-set.test.ts tests/skill-inventory.test.ts`
|
|
262
|
+
Expected: all tests pass, including 18 Claude additions and 30 Codex additions.
|
|
263
|
+
|
|
264
|
+
```bash
|
|
265
|
+
git add src/core/active-policy.ts src/core/active-set.ts tests/active-policy.test.ts
|
|
266
|
+
git commit -m "fix: enforce per-agent active skill limits"
|
|
267
|
+
```
|
|
268
|
+
|
|
269
|
+
### Task 3: Detect bounded Node CLI and package signals
|
|
270
|
+
|
|
271
|
+
**Files:**
|
|
272
|
+
|
|
273
|
+
- Modify: `src/shared/types.ts`
|
|
274
|
+
- Modify: `src/core/recommend.ts`
|
|
275
|
+
- Modify: `tests/recommend.test.ts`
|
|
276
|
+
|
|
277
|
+
**Interfaces:**
|
|
278
|
+
|
|
279
|
+
- Extends `ProjectSignals` with `roles: string[]` and `tools: string[]`.
|
|
280
|
+
- Consumed by package recommendations and skill ranking.
|
|
281
|
+
|
|
282
|
+
- [x] **Step 1: Write failing signal tests**
|
|
283
|
+
|
|
284
|
+
Create a package fixture containing `bin`, `publishConfig`, `commander`, `zod`, `vitest`, `@playwright/test`, `prepack`, and an `mcp` keyword, plus `SECURITY.md`. Assert:
|
|
285
|
+
|
|
286
|
+
```ts
|
|
287
|
+
expect(signals.roles).toEqual(
|
|
288
|
+
expect.arrayContaining([
|
|
289
|
+
"node-cli",
|
|
290
|
+
"npm-package",
|
|
291
|
+
"release",
|
|
292
|
+
"mcp",
|
|
293
|
+
"security",
|
|
294
|
+
]),
|
|
295
|
+
);
|
|
296
|
+
expect(signals.tools).toEqual(
|
|
297
|
+
expect.arrayContaining(["commander", "zod", "vitest", "playwright"]),
|
|
298
|
+
);
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
- [x] **Step 2: Run the recommendation tests and verify RED**
|
|
302
|
+
|
|
303
|
+
Run: `npx vitest run tests/recommend.test.ts`
|
|
304
|
+
Expected: FAIL because `roles` and `tools` are absent.
|
|
305
|
+
|
|
306
|
+
- [x] **Step 3: Extend project signals and parse only known metadata**
|
|
307
|
+
|
|
308
|
+
```ts
|
|
309
|
+
export interface ProjectSignals {
|
|
310
|
+
root: string;
|
|
311
|
+
languages: string[];
|
|
312
|
+
frameworks: string[];
|
|
313
|
+
roles: string[];
|
|
314
|
+
tools: string[];
|
|
315
|
+
files: string[];
|
|
316
|
+
}
|
|
317
|
+
```
|
|
318
|
+
|
|
319
|
+
In `scanProject`, derive roles and tools only from root entry names, known manifest fields, dependencies/devDependencies, scripts, publish metadata, and keywords. Do not recursively read arbitrary source content.
|
|
320
|
+
|
|
321
|
+
- [x] **Step 4: Format readable detected roles**
|
|
322
|
+
|
|
323
|
+
Add display labels so human output says `TypeScript, Node CLI, npm package, Vitest, Playwright, MCP tooling` while JSON retains stable lowercase identifiers.
|
|
324
|
+
|
|
325
|
+
- [x] **Step 5: Verify GREEN and commit**
|
|
326
|
+
|
|
327
|
+
Run: `npx vitest run tests/recommend.test.ts tests/outcomes.test.ts`
|
|
328
|
+
Expected: all tests pass.
|
|
329
|
+
|
|
330
|
+
```bash
|
|
331
|
+
git add src/shared/types.ts src/core/recommend.ts tests/recommend.test.ts
|
|
332
|
+
git commit -m "feat: detect local cli and package signals"
|
|
333
|
+
```
|
|
334
|
+
|
|
335
|
+
### Task 4: Diversify project skill selection
|
|
336
|
+
|
|
337
|
+
**Files:**
|
|
338
|
+
|
|
339
|
+
- Modify: `src/core/active-policy.ts`
|
|
340
|
+
- Modify: `tests/active-policy.test.ts`
|
|
341
|
+
|
|
342
|
+
**Interfaces:**
|
|
343
|
+
|
|
344
|
+
- Consumes: extended `ProjectSignals`.
|
|
345
|
+
- Produces: evidence-threshold candidates grouped by capability family.
|
|
346
|
+
|
|
347
|
+
- [x] **Step 1: Write failing relevance tests**
|
|
348
|
+
|
|
349
|
+
Add reviewed candidates named `javascript-typescript-jest`, `vitest-testing`, five Playwright variants, `cli-design`, `npm-package`, and `mcp-security`. For a Vitest Node CLI fixture, assert Jest is absent, CLI/npm/MCP candidates are present, and no more than three browser-testing candidates are selected.
|
|
350
|
+
|
|
351
|
+
- [x] **Step 2: Run the active-policy tests and verify RED**
|
|
352
|
+
|
|
353
|
+
Run: `npx vitest run tests/active-policy.test.ts`
|
|
354
|
+
Expected: FAIL because current ranking selects Jest and redundant Playwright variants.
|
|
355
|
+
|
|
356
|
+
- [x] **Step 3: Add exact signal rules and mismatch rejection**
|
|
357
|
+
|
|
358
|
+
Add role/tool rules for Node CLI, npm package, Vitest, Commander, Zod, MCP, release, and security. Reject a candidate matching `jest` when `vitest` is present and `jest` is absent.
|
|
359
|
+
|
|
360
|
+
- [x] **Step 4: Add deterministic family caps**
|
|
361
|
+
|
|
362
|
+
Implement `candidateFamily(unitId)` and deterministic caps: browser testing 3; documentation 2; code review 2; architecture 2; planning 3; security 3; language/tooling 5; uncategorized 3. Explicit full-selector pins bypass family caps but still consume capacity. Stop after eligible candidates are exhausted; never fill unused slots with candidates below the existing evidence threshold.
|
|
363
|
+
|
|
364
|
+
- [x] **Step 5: Verify GREEN and commit**
|
|
365
|
+
|
|
366
|
+
Run: `npx vitest run tests/active-policy.test.ts`
|
|
367
|
+
Expected: all relevance, diversity, pin, outcome, and per-agent capacity tests pass.
|
|
368
|
+
|
|
369
|
+
```bash
|
|
370
|
+
git add src/core/active-policy.ts tests/active-policy.test.ts
|
|
371
|
+
git commit -m "feat: diversify project skill selection"
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
### Task 5: Label recommendation component types
|
|
375
|
+
|
|
376
|
+
**Files:**
|
|
377
|
+
|
|
378
|
+
- Modify: `src/shared/types.ts`
|
|
379
|
+
- Modify: `src/core/recommend.ts`
|
|
380
|
+
- Modify: `tests/recommend.test.ts`
|
|
381
|
+
|
|
382
|
+
**Interfaces:**
|
|
383
|
+
|
|
384
|
+
- Extends `PackageRecommendation` with `kind: "skill-library" | "mcp-runtime" | "unavailable"`.
|
|
385
|
+
- Uses catalog `components` to classify suggestions.
|
|
386
|
+
|
|
387
|
+
- [x] **Step 1: Write failing type-label tests**
|
|
388
|
+
|
|
389
|
+
Assert Superpowers formats as `skill library`, while Playwright MCP and GitHub MCP format as `MCP/runtime setup` and include a separate preview/setup hint rather than activation wording.
|
|
390
|
+
|
|
391
|
+
- [x] **Step 2: Run the recommendation tests and verify RED**
|
|
392
|
+
|
|
393
|
+
Run: `npx vitest run tests/recommend.test.ts`
|
|
394
|
+
Expected: FAIL because recommendations have no `kind`.
|
|
395
|
+
|
|
396
|
+
- [x] **Step 3: Classify catalog-backed recommendations**
|
|
397
|
+
|
|
398
|
+
```ts
|
|
399
|
+
function recommendationKind(
|
|
400
|
+
pkg: CatalogPackage,
|
|
401
|
+
): PackageRecommendation["kind"] {
|
|
402
|
+
if (pkg.components?.includes("skill")) return "skill-library";
|
|
403
|
+
if (pkg.components?.some((item) => item === "mcp" || item === "plugin"))
|
|
404
|
+
return "mcp-runtime";
|
|
405
|
+
return "unavailable";
|
|
406
|
+
}
|
|
407
|
+
```
|
|
408
|
+
|
|
409
|
+
Attach the kind when adding each catalog recommendation and preserve it through local-outcome personalization.
|
|
410
|
+
|
|
411
|
+
- [x] **Step 4: Format kinds and next actions**
|
|
412
|
+
|
|
413
|
+
Human output must label every line and end MCP/runtime suggestions with the read-only recipe or explicit setup command supported by the package. It must not claim those integrations are automatically activatable skills.
|
|
414
|
+
|
|
415
|
+
- [x] **Step 5: Verify GREEN and commit**
|
|
416
|
+
|
|
417
|
+
Run: `npx vitest run tests/recommend.test.ts tests/outcomes.test.ts`
|
|
418
|
+
Expected: all tests pass.
|
|
419
|
+
|
|
420
|
+
```bash
|
|
421
|
+
git add src/shared/types.ts src/core/recommend.ts tests/recommend.test.ts
|
|
422
|
+
git commit -m "feat: label recommendation component types"
|
|
423
|
+
```
|
|
424
|
+
|
|
425
|
+
### Task 6: Verify the product path and prepare the release candidate
|
|
426
|
+
|
|
427
|
+
**Files:**
|
|
428
|
+
|
|
429
|
+
- Modify: `CHANGELOG.md`
|
|
430
|
+
- Modify: `master_plan.md`
|
|
431
|
+
- Modify: `docs/USER_TEST_GUIDE.md`
|
|
432
|
+
- Modify: `scripts/cli-product-flow.mjs`
|
|
433
|
+
|
|
434
|
+
**Interfaces:**
|
|
435
|
+
|
|
436
|
+
- Consumes: all corrected preview/apply behavior.
|
|
437
|
+
- Produces: repeatable release and founder acceptance evidence.
|
|
438
|
+
|
|
439
|
+
- [x] **Step 1: Extend the CLI product-flow regression**
|
|
440
|
+
|
|
441
|
+
Add a journey that creates one agent with unmanaged skills and one empty agent, restores recursively empty snapshot residue, installs a disabled library, previews project activation, applies it, and rolls back the explicit activation snapshot. Assert unmanaged bytes are unchanged at every point.
|
|
442
|
+
|
|
443
|
+
- [x] **Step 2: Run the focused product flow after the unit-level RED/GREEN cycles**
|
|
444
|
+
|
|
445
|
+
Run: `npm run test:e2e:cli`
|
|
446
|
+
Expected: PASS for per-agent capacity, empty-target activation, atomic apply, and explicit rollback.
|
|
447
|
+
|
|
448
|
+
- [x] **Step 3: Update user-facing documentation**
|
|
449
|
+
|
|
450
|
+
Document that `--limit` includes unmanaged skills, Maximum remains disabled by default, recommendations distinguish skills from MCP/runtime setup, and project activation may choose different set sizes per agent.
|
|
451
|
+
|
|
452
|
+
- [x] **Step 4: Run the full local release gate**
|
|
453
|
+
|
|
454
|
+
Run: `npm run verify`
|
|
455
|
+
Expected: formatting, lint, typecheck, evidence checks, unit tests, CLI/readme/package flows, and performance checks all pass.
|
|
456
|
+
|
|
457
|
+
- [x] **Step 5: Build and inspect the packed npm artifact without publishing**
|
|
458
|
+
|
|
459
|
+
Run: `npm pack --dry-run`
|
|
460
|
+
Expected: the package contains the corrected compiled CLI and no untracked secret or local-state files.
|
|
461
|
+
|
|
462
|
+
- [x] **Step 6: Record completion and commit**
|
|
463
|
+
|
|
464
|
+
Mark only the implemented portions of `P18-27` complete, record exact test counts and remaining founder/npm steps, and keep real-profile activation blocked until the corrected package is published.
|
|
465
|
+
|
|
466
|
+
```bash
|
|
467
|
+
git add CHANGELOG.md MASTER_PLAN.md docs/USER_TEST_GUIDE.md scripts/cli-product-flow.mjs
|
|
468
|
+
git commit -m "test: cover safe project activation journey"
|
|
469
|
+
```
|
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
# Project Activation Safety and Relevance Design
|
|
2
|
+
|
|
3
|
+
**Status:** Approved direction; awaiting written-spec review
|
|
4
|
+
**Date:** 2026-07-20
|
|
5
|
+
**Release target:** Loadout 0.4.1
|
|
6
|
+
|
|
7
|
+
## Problem
|
|
8
|
+
|
|
9
|
+
The founder acceptance run of `loadout activate --project . --agents
|
|
10
|
+
codex,claude-code --limit 30` exposed three connected defects:
|
|
11
|
+
|
|
12
|
+
1. The preview reported `0/30 active` even though Claude Code already had 12
|
|
13
|
+
unmanaged skills. The planner counted only active Loadout records, so applying the
|
|
14
|
+
proposal could exceed the disclosed per-agent limit.
|
|
15
|
+
2. An earlier rollback correctly restored recursively empty skill directories. The
|
|
16
|
+
activation planner treated the existence of those empty directories as occupied
|
|
17
|
+
content and blocked safe activation for both agents.
|
|
18
|
+
3. Project detection recognized only JavaScript/TypeScript and Playwright. It missed
|
|
19
|
+
strong local evidence that this repository is a Node CLI, an npm package, a Vitest
|
|
20
|
+
project, an MCP-related tool, and a release/supply-chain project. Ranking then used
|
|
21
|
+
the entire 30-slot budget as a quota and proposed redundant or mismatched skills,
|
|
22
|
+
including several overlapping Playwright workflows and a Jest skill for a Vitest
|
|
23
|
+
repository.
|
|
24
|
+
|
|
25
|
+
No activation was applied. The Maximum library remains disabled, and the founder's
|
|
26
|
+
12 unmanaged Claude skills remain unchanged.
|
|
27
|
+
|
|
28
|
+
## Goals
|
|
29
|
+
|
|
30
|
+
- Make `--limit` a truthful ceiling over all active skills visible to each selected
|
|
31
|
+
agent, whether managed by Loadout or not.
|
|
32
|
+
- Plan and explain capacity separately for every selected agent.
|
|
33
|
+
- Treat recursively empty rollback residue as unoccupied while continuing to refuse
|
|
34
|
+
every non-empty, symlinked, unreadable, or unsupported target.
|
|
35
|
+
- Improve local project recognition enough to distinguish this Node CLI/npm/Vitest
|
|
36
|
+
repository from a generic Playwright web application.
|
|
37
|
+
- Select a compact, diverse working set instead of filling every available slot with
|
|
38
|
+
weak or redundant matches.
|
|
39
|
+
- Preserve a preview-first, transactional, rollback-safe mutation path.
|
|
40
|
+
|
|
41
|
+
## Non-goals
|
|
42
|
+
|
|
43
|
+
- No model or external API call is required for recommendation or activation.
|
|
44
|
+
- No project source, filename, dependency, or outcome data leaves the machine.
|
|
45
|
+
- This work does not claim that a deterministic recommendation proves global package
|
|
46
|
+
quality.
|
|
47
|
+
- This work does not install deferred MCP servers or collect credentials.
|
|
48
|
+
- Dashboard removal remains a separate 0.4.1 task.
|
|
49
|
+
|
|
50
|
+
## Design
|
|
51
|
+
|
|
52
|
+
### 1. One shared definition of an occupied skill target
|
|
53
|
+
|
|
54
|
+
Loadout will use one filesystem predicate for both initial installation and later
|
|
55
|
+
library activation:
|
|
56
|
+
|
|
57
|
+
- A missing path is unoccupied.
|
|
58
|
+
- A real directory containing no entries at any depth is unoccupied.
|
|
59
|
+
- A regular file, symlink, special entry, unreadable directory, or directory with any
|
|
60
|
+
non-directory descendant is occupied.
|
|
61
|
+
- Recursive inspection is bounded to 10,000 entries. Reaching the bound is treated as
|
|
62
|
+
occupied, never as empty.
|
|
63
|
+
|
|
64
|
+
Activation will re-check this predicate inside the transaction immediately before
|
|
65
|
+
copying. This closes the preview/apply race without deleting or adopting user content.
|
|
66
|
+
Empty target directories may be removed immediately before the copy because they
|
|
67
|
+
contain no bytes to preserve; rollback still restores the pre-transaction topology.
|
|
68
|
+
|
|
69
|
+
### 2. Per-agent total active capacity
|
|
70
|
+
|
|
71
|
+
For each requested agent, the planner will scan the agent's real skill root and count
|
|
72
|
+
every directory containing a valid `SKILL.md`. This inventory already distinguishes
|
|
73
|
+
managed and unmanaged skills and will be the capacity source of truth.
|
|
74
|
+
|
|
75
|
+
For agent `a`:
|
|
76
|
+
|
|
77
|
+
```text
|
|
78
|
+
available(a) = max(0, limit - inventory(a).total)
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
Candidate ranking is deterministic and shared, but selection is sliced independently
|
|
82
|
+
to each agent's available capacity. In the founder fixture, Claude Code has 12 active
|
|
83
|
+
unmanaged skills and may receive at most 18 additions; Codex has zero active skills and
|
|
84
|
+
may receive at most 30. An agent at its limit receives no changes and a clear warning;
|
|
85
|
+
it does not prevent another requested agent from using its own available capacity.
|
|
86
|
+
|
|
87
|
+
The plan must never count an empty directory as an active skill. Existing managed
|
|
88
|
+
skills already in the desired state count toward capacity but are not proposed again.
|
|
89
|
+
|
|
90
|
+
### 3. Project signals
|
|
91
|
+
|
|
92
|
+
Project scanning remains deterministic, bounded, and local-only. It will read known
|
|
93
|
+
root manifests/configuration and a bounded set of repository filenames. For Node
|
|
94
|
+
projects it will derive explicit signals from `package.json`:
|
|
95
|
+
|
|
96
|
+
- `bin` -> Node CLI
|
|
97
|
+
- package name plus `publishConfig` or non-private package -> npm package/release
|
|
98
|
+
- `vitest` -> Vitest
|
|
99
|
+
- `jest` -> Jest
|
|
100
|
+
- `@playwright/test` or Playwright config -> Playwright
|
|
101
|
+
- `commander` -> command-line application
|
|
102
|
+
- `zod` -> schema validation
|
|
103
|
+
- scripts containing `prepack`, package-smoke, or release checks -> package/release
|
|
104
|
+
|
|
105
|
+
Repository filenames add delivery, security, MCP, and documentation signals only when
|
|
106
|
+
matching explicit reviewed patterns such as `.github/workflows`, `SECURITY.md`, MCP-
|
|
107
|
+
named modules, and package/release scripts. The scanner does not recursively read
|
|
108
|
+
arbitrary project source or send any data elsewhere.
|
|
109
|
+
|
|
110
|
+
Human output will name the important detected roles, for example:
|
|
111
|
+
|
|
112
|
+
```text
|
|
113
|
+
Detected: TypeScript, Node CLI, npm package, Vitest, Playwright, MCP tooling
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
### 4. Relevant and diverse selection
|
|
117
|
+
|
|
118
|
+
The existing source-priority tie-break remains, but a capacity ceiling is not a quota.
|
|
119
|
+
Selection stops when no candidate reaches the evidence threshold.
|
|
120
|
+
|
|
121
|
+
Ranking will:
|
|
122
|
+
|
|
123
|
+
- reward exact tool and role matches such as Vitest, CLI design, npm packaging,
|
|
124
|
+
TypeScript, MCP security, release verification, and documentation;
|
|
125
|
+
- reject framework mismatches such as Jest-only guidance when Vitest is detected and
|
|
126
|
+
Jest is not;
|
|
127
|
+
- group close alternatives into capability families such as browser testing, code
|
|
128
|
+
review, documentation, planning, security, and architecture;
|
|
129
|
+
- select the highest-evidence representative before taking another member of the same
|
|
130
|
+
family;
|
|
131
|
+
- cap weak generic foundation choices so they cannot crowd out project-specific
|
|
132
|
+
evidence;
|
|
133
|
+
- continue to honor explicit `--pin` selectors, while disclosing when a pin consumes a
|
|
134
|
+
slot or conflicts with the limit.
|
|
135
|
+
|
|
136
|
+
The same skill name from multiple repositories remains an alternative, not two active
|
|
137
|
+
copies. Multiple distinct skills in one family are allowed only when each has separate
|
|
138
|
+
strong project evidence.
|
|
139
|
+
|
|
140
|
+
### 5. Recommendation type clarity
|
|
141
|
+
|
|
142
|
+
Package-level `loadout recommend` output will label every suggestion as one of:
|
|
143
|
+
|
|
144
|
+
- `skill library` — eligible for disabled-library activation;
|
|
145
|
+
- `MCP/runtime setup` — requires a separate explicit preview and may require a
|
|
146
|
+
non-model credential;
|
|
147
|
+
- `unavailable` — not prepared locally, with the reason.
|
|
148
|
+
|
|
149
|
+
Playwright MCP and GitHub MCP must not look like ordinary automatically activatable
|
|
150
|
+
skills. The command will show the next preview command for explicit integrations and
|
|
151
|
+
will continue to state that recommendations are rule-selected, not proof of quality.
|
|
152
|
+
|
|
153
|
+
### 6. Preview and apply output
|
|
154
|
+
|
|
155
|
+
The default human preview will begin with one compact block per agent:
|
|
156
|
+
|
|
157
|
+
```text
|
|
158
|
+
Claude Code: 12 active (0 managed, 12 unmanaged), 18/30 slots available
|
|
159
|
+
Codex: 0 active, 30/30 slots available
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
It will then show additions grouped by agent and capability, followed by blockers. A
|
|
163
|
+
recursively empty target is not printed as a blocker. A real conflict prints the exact
|
|
164
|
+
target and a non-destructive next action. JSON output retains full scores, reasons,
|
|
165
|
+
alternatives, per-agent budgets, targets, and blocker details.
|
|
166
|
+
|
|
167
|
+
`--yes` is rejected when any proposed target is blocked. A successful apply creates
|
|
168
|
+
one transaction snapshot covering every selected agent and prints the exact rollback
|
|
169
|
+
command.
|
|
170
|
+
|
|
171
|
+
## Data flow
|
|
172
|
+
|
|
173
|
+
1. Scan project metadata into local project signals.
|
|
174
|
+
2. Scan each requested agent's active skill inventory.
|
|
175
|
+
3. Read reviewed disabled-library records and local-only outcomes.
|
|
176
|
+
4. Score and diversify eligible candidates.
|
|
177
|
+
5. Slice the ordered candidates independently to each agent's available capacity.
|
|
178
|
+
6. Build exact activation changes and inspect targets with the shared occupancy rule.
|
|
179
|
+
7. Print the preview; stop unless `--yes` is present.
|
|
180
|
+
8. Inside the transaction, re-read state, re-check capacity and target occupancy, copy
|
|
181
|
+
reviewed library bytes, fingerprint the result, persist activation state, and emit
|
|
182
|
+
the rollback snapshot.
|
|
183
|
+
|
|
184
|
+
## Failure handling
|
|
185
|
+
|
|
186
|
+
- If inventory scanning fails for an agent, activation for that agent is blocked; its
|
|
187
|
+
capacity is never guessed.
|
|
188
|
+
- If the project manifest is malformed, show the exact manifest error and continue
|
|
189
|
+
only with signals that remain trustworthy.
|
|
190
|
+
- If state or filesystem contents change between preview and apply, abort the entire
|
|
191
|
+
multi-agent transaction without partial activation.
|
|
192
|
+
- If fewer candidates meet the evidence threshold than available slots, activate only
|
|
193
|
+
those candidates and explain that unused capacity is intentional.
|
|
194
|
+
|
|
195
|
+
## Tests and acceptance criteria
|
|
196
|
+
|
|
197
|
+
Automated tests must prove:
|
|
198
|
+
|
|
199
|
+
1. A recursively empty target does not block preview or apply.
|
|
200
|
+
2. A target containing a file, symlink, unreadable entry, or more than the inspection
|
|
201
|
+
bound remains blocked and unchanged.
|
|
202
|
+
3. Empty-directory handling is identical in initial setup and project activation.
|
|
203
|
+
4. Twelve unmanaged Claude skills plus an 18-skill proposal never exceeds a limit of 30.
|
|
204
|
+
5. The same plan may select 18 additions for Claude and 30 for an empty Codex profile.
|
|
205
|
+
6. An agent already at capacity receives no additions while another agent can proceed.
|
|
206
|
+
7. Apply re-checks inventory and occupancy and aborts atomically when either changes.
|
|
207
|
+
8. The Loadout repository fixture detects Node CLI, npm package, Vitest, Playwright,
|
|
208
|
+
Commander, Zod, release, security, and MCP signals.
|
|
209
|
+
9. A Vitest-only fixture never recommends a Jest-only skill.
|
|
210
|
+
10. Redundant Playwright/browser candidates do not consume most of the active set.
|
|
211
|
+
11. Recommendation output distinguishes skill libraries from explicit MCP/runtime
|
|
212
|
+
setup.
|
|
213
|
+
12. Human output shows per-agent managed/unmanaged counts and unused capacity; JSON
|
|
214
|
+
exposes the same facts structurally.
|
|
215
|
+
13. Existing Stable, Power, Maximum, rollback, and non-overwrite tests remain green.
|
|
216
|
+
|
|
217
|
+
Founder acceptance resumes only after the published release candidate previews a
|
|
218
|
+
non-blocked plan, reports Claude's existing 12 skills, respects both per-agent limits,
|
|
219
|
+
applies transactionally, and restores its explicit snapshot without changing unmanaged
|
|
220
|
+
content.
|
|
221
|
+
|
|
222
|
+
## Compatibility
|
|
223
|
+
|
|
224
|
+
Existing `loadout activate` and `loadout optimize` flags remain valid. The meaning of
|
|
225
|
+
`--limit` is corrected from an implicit managed-only count to the documented total
|
|
226
|
+
active skill ceiling. JSON consumers must migrate from one global `activeBefore` and
|
|
227
|
+
`capacity` pair to per-agent budget records; the 0.4.1 changelog will call out this
|
|
228
|
+
schema correction.
|