@mmerterden/multi-agent-pipeline 17.5.0 → 17.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/CHANGELOG.md +177 -0
  2. package/README.md +24 -0
  3. package/README.tr.md +24 -0
  4. package/docs/features.md +44 -3
  5. package/docs/token-budget-history.md +1 -1
  6. package/install/templates/claude-hooks.json +13 -1
  7. package/package.json +1 -1
  8. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  9. package/pipeline/commands/multi-agent/feedback/SKILL.md +7 -1
  10. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  11. package/pipeline/commands/multi-agent/issue/SKILL.md +13 -1
  12. package/pipeline/commands/multi-agent/jira/SKILL.md +13 -1
  13. package/pipeline/commands/multi-agent/resume/SKILL.md +16 -1
  14. package/pipeline/commands/multi-agent/setup/SKILL.md +14 -16
  15. package/pipeline/commands/multi-agent/update/SKILL.md +13 -56
  16. package/pipeline/multi-agent-refs/features/code-graph.md +20 -0
  17. package/pipeline/multi-agent-refs/features/doctor.md +23 -0
  18. package/pipeline/multi-agent-refs/features/maturity-followup.md +166 -0
  19. package/pipeline/multi-agent-refs/features/package-manager.md +80 -0
  20. package/pipeline/multi-agent-refs/features/usage-reporting.md +79 -0
  21. package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
  22. package/pipeline/multi-agent-refs/phases/phase-0-init.md +5 -2
  23. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +8 -2
  24. package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -1
  25. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  26. package/pipeline/preferences-template.json +1 -1
  27. package/pipeline/schemas/agent-state.schema.json +122 -11
  28. package/pipeline/schemas/prefs.schema.json +35 -0
  29. package/pipeline/schemas/token-budget.json +2 -2
  30. package/pipeline/scripts/doctor.mjs +65 -0
  31. package/pipeline/scripts/feedback-send.mjs +1 -1
  32. package/pipeline/scripts/graph-report.mjs +155 -1
  33. package/pipeline/scripts/maturity-followup.mjs +294 -0
  34. package/pipeline/scripts/package-manager.mjs +310 -0
  35. package/pipeline/scripts/usage-register.mjs +271 -0
  36. package/pipeline/scripts/usage-report.mjs +2 -2
  37. package/pipeline/skills/.skill-manifest.json +5 -5
  38. package/pipeline/skills/shared/core/multi-agent-issue/SKILL.md +14 -0
  39. package/pipeline/skills/shared/core/multi-agent-jira/SKILL.md +14 -0
  40. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +13 -0
  41. package/pipeline/skills/shared/core/multi-agent-update/SKILL.md +6 -0
@@ -90,12 +90,24 @@ After issue selection, inspect `maturity` from `~/.claude/lib/issue-fetcher.sh`:
90
90
 
91
91
  | Outcome | Behavior |
92
92
  |---|---|
93
- | `blockers` non-empty | **Halt** (autopilot too): show summary, advise "fix the issue first" |
93
+ | `blockers` non-empty | **Ask, then halt** - see "Blockers" below. Default is still the halt |
94
94
  | `warnings` contains `description_empty_parent_available` | Tailored parent-description question above (not the generic one) |
95
95
  | `warnings` contains `description_empty_sibling_available` | Same question, sibling key substituted for the parent's |
96
96
  | other `warnings` only | AskUserQuestion: show summary + "Continue?". Autopilot auto-continues; warnings logged to `agent-log.md` |
97
97
  | both empty (score ≥ 90) | Continue silently |
98
98
 
99
+ **Blockers: ask, do not just stop.** `$HOME/.claude/multi-agent-refs/features/maturity-followup.md`
100
+ owns this. In short: an interactive run asks at this step (open the item and fix it /
101
+ continue without it, recording which gap was waved through / abort) instead of ending
102
+ with a summary; an autopilot run with `prefs.global.maturityFollowup.autopilotCommentsOnIssue`
103
+ on posts ONE comment on the item asking for what is missing, then halts on the circuit
104
+ breaker with `state.waitingFor = "maturity"` so `resume` re-enters THIS step with the item
105
+ re-fetched. Both defaults keep today's behaviour: the comment is off, the halt is the halt.
106
+ An edit is a reason to re-check, never proof the gap closed - the check re-runs on the new
107
+ content, and only a DIFFERENT gap set ever earns a second comment. Read the item's existing
108
+ comments FIRST and derive the prior ask from the newest one of ours (`priorFromComments`):
109
+ a scan is a new run with a fresh state file, so state alone would make every scan a first ask.
110
+
99
111
  > **Resolution check**: If Jira `resolution` is set (`Fixed`, `Done`, `Won't Do`, `Resolved`...), `already_resolved` is added - the issue was closed already. To reopen, clear `resolution` in Jira first.
100
112
 
101
113
  > **fixVersions**: When `descriptor.fixVersions` is non-empty (e.g. `["v1.48.0"]`) and a matching `release/v1.48.0` branch exists in `prefs.projects[*].branches`, suggest that branch. Otherwise fall back to `develop`/cwd current.
@@ -38,6 +38,21 @@ Resume a paused or failed task from the last successful phase.
38
38
  - Claude Code: `TaskCreate` every phase from the state file in phase order, then `TaskUpdate` each to its stored status, and replace every `tasklist_id` meta with the new IDs.
39
39
  - Other CLIs: a single `bash $HOME/.claude/scripts/phase-tracker.sh render`.
40
40
 
41
- 5. **Continue the pipeline** - start from the next phase (same pipeline as the main multi-agent command).
41
+ 5. **Continue the pipeline.** Read `state.waitingFor` FIRST: when it names a step, the
42
+ run re-enters THAT step rather than the next phase. `currentPhase + 1` is the fallback,
43
+ not the rule - a run that stopped mid-phase to ask a human has `currentPhase` pointing at
44
+ the phase it is still inside, so resuming past it skips the question permanently. That
45
+ was already true of Phase 7's channels pause, which documented itself as resumable
46
+ through this field while this file never mentioned it.
47
+
48
+ | `waitingFor` | Re-entry |
49
+ |---|---|
50
+ | `maturity` | Phase 0, the maturity step, with the item **re-fetched** and re-scored - an edit is a reason to look again, never proof the gap closed (`$HOME/.claude/multi-agent-refs/features/maturity-followup.md`) |
51
+ | `user-channels-choice` | Phase 7, the channels multi-select, with the stored `channelsInput` |
52
+ | absent | `currentPhase + 1`, as before (same pipeline as the main multi-agent command) |
53
+
54
+ Clear `waitingFor` in the same write that records the answer, the moment the step is
55
+ re-entered. A field that outlives the question it asked sends every later resume back
56
+ to the step the user already answered.
42
57
 
43
58
  6. **Log**: `🔄 Resumed {JIRA-KEY}-{id} from Phase {N}`
@@ -209,9 +209,7 @@ Full key list and shape: `$HOME/.claude/multi-agent-refs/keychain.md`.
209
209
 
210
210
  `null` = not mapped (missing or skipped). Pipeline phases read this mapping to retrieve tokens dynamically - never hardcoded key names.
211
211
 
212
- ### Step 2.7 - Operational reporting token (optional, opt-in)
213
-
214
- Only relevant when the admin has issued this user a token. Since v15.8.0, `/multi-agent:update` self-registers a per-machine token automatically when none is onboarded (opt-out: `usageLog.optOut=true`), so Skip here is never a dead end; an admin-issued token pasted now simply takes precedence.
212
+ ### Step 2.7 - Operational reporting
215
213
 
216
214
  Ask (in `outputLanguage`), and proceed only on an explicit yes:
217
215
 
@@ -220,27 +218,27 @@ Do you have an operational-reporting token from your admin?
220
218
  [ Paste token / Skip ]
221
219
  ```
222
220
 
223
- On paste, store the secret in the credential store ONLY - never in a file, prefs value, git, or any synced/published tree. Use the standard per-user key name so it is revocable independently and consistent across the user's machines:
221
+ On paste, store it in the credential store ONLY - never in a file, prefs value,
222
+ git or any synced tree - under the standard per-user name, so it is revocable on
223
+ its own:
224
224
 
225
225
  ```bash
226
226
  ~/.claude/lib/credential-store.sh set "${USER}_Usage_Ingest_Token" "<pasted-token>"
227
227
  ```
228
228
 
229
- Then map it and enable logging (the token itself stays in the credential store; only the logical mapping + the on-switch land in prefs):
229
+ Then, **whether they pasted or skipped**, run:
230
230
 
231
231
  ```bash
232
- node -e '
233
- const fs=require("fs"),os=require("os"),p=os.homedir()+"/.claude/multi-agent-preferences.json";
234
- const j=JSON.parse(fs.readFileSync(p,"utf8"));
235
- j.global=j.global||{}; j.global.keychainMapping=j.global.keychainMapping||{};
236
- j.global.keychainMapping.usage_ingest=process.argv[1];
237
- j.global.usageLog=Object.assign({enabled:true},j.global.usageLog||{},{enabled:true});
238
- fs.writeFileSync(p,JSON.stringify(j,null,2)+"\n");
239
- ' "${USER}_Usage_Ingest_Token"
240
- echo " -> operational reporting configured (token in credential store)"
232
+ node "$HOME/.claude/scripts/usage-register.mjs"
241
233
  ```
242
234
 
243
- Security notes to surface to the user: the token is **write-only** (append-only to the endpoint - no read access, no other scope), stored **only in the OS credential store**, and **per-user** so the admin can revoke this one token without affecting anyone else. `usage-report.mjs` reads it from the credential store at runtime via the `usage_ingest` mapping; it is never written to a file or transmitted except over TLS to the ingest endpoint.
235
+ It maps whatever token exists and switches reporting on, and requests a
236
+ per-machine write-only one when none resolves. This call is why the step exists:
237
+ registration used to happen only inside `/multi-agent:update`, so a user who
238
+ installed, ran setup and never ran update never registered and never reported -
239
+ which reads in the panel exactly like nobody using the pipeline. What is sent,
240
+ what never is, and `usageLog.optOut`:
241
+ `$HOME/.claude/multi-agent-refs/features/usage-reporting.md`.
244
242
 
245
243
  ### Auto-learned fields (no setup step needed)
246
244
 
@@ -802,7 +800,7 @@ All tokens are optional in the sense that every service can be answered with Ski
802
800
 
803
801
  ### Step 8 - Enforcement hooks (optional, Claude Code)
804
802
 
805
- Offer to merge `$HOME/.claude/templates/claude-hooks.json`: three `PreToolUse` gates that block on a non-zero exit (secret scan, agent-guard, read-size) plus two capture hooks that block nothing (`SessionEnd`, `SessionStart`). What each does: `$HOME/.claude/multi-agent-refs/picker-contract.md`.
803
+ Offer to merge `$HOME/.claude/templates/claude-hooks.json`: three `PreToolUse` gates that block on a non-zero exit (secret scan, agent-guard, read-size) plus three capture hooks that block nothing (`PreCompact`, `SessionEnd`, `SessionStart`). What each does: `$HOME/.claude/multi-agent-refs/picker-contract.md`.
806
804
 
807
805
  - Ask (picker): "Install the pipeline's hooks into `~/.claude/settings.json`?" Default Yes.
808
806
  - On Yes, deep-merge EVERY event in the template's `hooks` object, not `PreToolUse` alone - merging one event silently drops the capture hooks, and a run killed before Phase 7 then loses its findings exactly as it did before they existed. Preserve existing hooks; never duplicate a matcher already calling the same script.
@@ -111,65 +111,22 @@ A git clone of the pipeline repo is a maintainer workspace, kept in sync by `/mu
111
111
  fi
112
112
  ```
113
113
 
114
- 5b. **Auto-configure operational reporting.** Off means silent: the emitter still
115
- no-ops unless `enabled` is true AND a token resolves. This step never SHIPS a
116
- secret; when no token is onboarded it REQUESTS a per-machine write-only token
117
- from the reporting endpoint's `/register` route (self-registration, v15.8.0+),
118
- stores it only in the credential store, and flips `usageLog.enabled` on. It
119
- reports coarse run metadata only - never prompts, code, diffs, or paths.
120
- Resolution order: env `MULTI_AGENT_USAGE_TOKEN`, then `usageLog.token`, then
121
- the Keychain item named by `keychainMapping.usage_ingest`, then
122
- self-registration. Registration failing (offline, endpoint down, admin turned
123
- ingest off) leaves reporting off with one status line - never an error.
124
- **Opt-out is `usageLog.optOut: true`**: it blocks both the auto-enable and the
125
- self-registration permanently; print the opt-out hint on first auto-enable.
114
+ 5b. **Auto-configure operational reporting.** One call, and it is the same call
115
+ `/multi-agent:setup` and a first run make - the registration used to live here
116
+ as forty lines of shell, so a user who installed, ran setup and never ran
117
+ update was never registered and never reported, which reads in the panel
118
+ exactly like nobody using the pipeline.
119
+
126
120
  ```bash
127
- PREFS="$HOME/.claude/multi-agent-preferences.json"
128
- # Reads go through node, not jq. node is a declared engine (>=20.11) so it
129
- # is always there; jq is not, and gating this block on it meant a machine
130
- # without jq silently never registered and never reported - which reads in
131
- # the panel exactly like nobody using the pipeline.
132
- if [ -f "$PREFS" ]; then
133
- pref() { node -e 'const fs=require("fs");let v;try{v=process.argv[2].split(".").reduce((a,k)=>a?.[k],JSON.parse(fs.readFileSync(process.argv[1],"utf8")))}catch{};process.stdout.write(v==null?"":String(v))' "$PREFS" "$1" 2>/dev/null; }
134
- ENABLED=$(pref global.usageLog.enabled)
135
- OPTOUT=$(pref global.usageLog.optOut)
136
- if [ "$ENABLED" != "true" ] && [ "$OPTOUT" != "true" ]; then
137
- TOK="${MULTI_AGENT_USAGE_TOKEN:-}"
138
- [ -z "$TOK" ] && TOK=$(pref global.usageLog.token)
139
- if [ -z "$TOK" ]; then
140
- KNAME=$(pref global.keychainMapping.usage_ingest)
141
- [ -n "$KNAME" ] && TOK=$(bash "$HOME/.claude/lib/credential-store.sh" get "$KNAME" 2>/dev/null)
142
- fi
143
- if [ -z "$TOK" ]; then
144
- EP=$(pref global.usageLog.endpoint); [ -z "$EP" ] && EP="https://mmerterden.vercel.app/api/usage/ingest"
145
- REG_EP="${EP%/ingest}/register"
146
- # Telemetry identity is the GitHub account name, never the git
147
- # identity.name (which can carry a corporate title). username -> live
148
- # gh login -> OS user.
149
- RUSER=$(pref global.identities.0.username)
150
- [ -z "$RUSER" ] && RUSER=$(gh api user --jq .login 2>/dev/null || echo "")
151
- [ -z "$RUSER" ] && RUSER="$USER"
152
- RESP=$(curl -sSL -m 10 -X POST -H "Content-Type: application/json" \
153
- --data "{\"u\":\"$RUSER\",\"c\":\"$(hostname -s 2>/dev/null || echo unknown)\"}" \
154
- "$REG_EP" 2>/dev/null)
155
- TOK=$(printf '%s' "$RESP" | node -e 'let b="";process.stdin.on("data",d=>b+=d).on("end",()=>{try{process.stdout.write(String(JSON.parse(b).token||""))}catch{}})' 2>/dev/null)
156
- if [ -n "$TOK" ]; then
157
- printf '%s' "$TOK" | bash "$HOME/.claude/lib/credential-store.sh" set "${USER}_Usage_Ingest_Token" "$(cat)"
158
- node -e 'const fs=require("fs"),p=process.argv[1];const j=JSON.parse(fs.readFileSync(p,"utf8"));j.global=j.global||{};j.global.keychainMapping=j.global.keychainMapping||{};j.global.keychainMapping.usage_ingest=process.argv[2];fs.writeFileSync(p,JSON.stringify(j,null,2)+"\n");' "$PREFS" "${USER}_Usage_Ingest_Token"
159
- echo " -> operational reporting: registered this machine (write-only token in credential store)"
160
- echo " opt out any time: set global.usageLog.optOut=true in multi-agent-preferences.json"
161
- fi
162
- fi
163
- if [ -n "$TOK" ]; then
164
- node -e 'const fs=require("fs"),p=process.argv[1];const j=JSON.parse(fs.readFileSync(p,"utf8"));j.global=j.global||{};j.global.usageLog=j.global.usageLog||{};j.global.usageLog.enabled=true;fs.writeFileSync(p,JSON.stringify(j,null,2)+"\n");' "$PREFS"
165
- echo " -> operational reporting configured"
166
- else
167
- echo " -> operational reporting left off (no token; registration unreachable or disabled)"
168
- fi
169
- fi
170
- fi
121
+ node "$HOME/.claude/scripts/usage-register.mjs"
171
122
  ```
172
123
 
124
+ It requests a per-machine **write-only** token when none resolves, stores it
125
+ only in the credential store, and flips `usageLog.enabled` on. `usageLog.optOut:
126
+ true` blocks it permanently. Offline, endpoint down or ingest disabled leaves
127
+ reporting off with one status line - never an error. Contract and the data it
128
+ sends: `$HOME/.claude/multi-agent-refs/features/usage-reporting.md`.
129
+
173
130
  6. **Show the new version and its changes** (from the packaged CHANGELOG - there is no git history on this channel):
174
131
  ```bash
175
132
  NEW=$(tr -d '[:space:]' < "$HOME/.claude/.pipeline-version" 2>/dev/null)
@@ -3,6 +3,7 @@
3
3
  <!-- toc -->
4
4
  - [Phase 1 Step 2.6 - query before dispatching Explore](#phase-1-step-26---query-before-dispatching-explore)
5
5
  - [Phase 7 Step 3 - refresh after the branch changed code](#phase-7-step-3---refresh-after-the-branch-changed-code)
6
+ - [What the report answers that a file-level view cannot](#what-the-report-answers-that-a-file-level-view-cannot)
6
7
  - [The graph is drawable, and one place already asks for it](#the-graph-is-drawable-and-one-place-already-asks-for-it)
7
8
  <!-- /toc -->
8
9
 
@@ -74,6 +75,25 @@ heuristic. A non-zero validator exit keeps the previous graph and logs
74
75
  `knowledge.graph_invalid`; it never fails the run - a stale graph is a degraded
75
76
  Phase 1, not a broken deliverable.
76
77
 
78
+ ### What the report answers that a file-level view cannot
79
+
80
+ `GRAPH_REPORT.md` ends with **Symbols nothing else references**: symbols no other
81
+ file in the repo names. The older "unconnected files" section only ever found
82
+ files with NO edge at all, so a file imported for one symbol while three of its
83
+ other exports were dead looked healthy.
84
+
85
+ It is candidates, never verdicts, and it gates nothing. The extractor is regex
86
+ over comment-stripped source, not a parser (ADR-0010), so dynamic dispatch,
87
+ reflection, string-keyed lookup and a public API consumed outside this repo are
88
+ indistinguishable from dead code here. Four classes are therefore excluded and
89
+ COUNTED rather than listed, because they could not carry a reference edge however
90
+ heavily they are used: a kind outside the stack's `referenceKinds`, a name
91
+ declared in more than one place (the builder drops ambiguous tokens), a nested
92
+ declaration, and anything declared in a test file. Symbols referenced ONLY from
93
+ tests are listed separately - that is not dead code, it is code whose only
94
+ consumer is its own test, which is worth knowing before a plan calls it
95
+ load-bearing.
96
+
77
97
  ### The graph is drawable, and one place already asks for it
78
98
 
79
99
  The PR body's Impact Analysis, part 3, asks which symbols and files a change
@@ -212,6 +212,29 @@ package in the npx cache left the server dying on `Cannot find module` at
212
212
  startup, which the client surfaces only as `CONNECTION_CLOSED`, and a check that
213
213
  read the registration and stopped would have called that healthy.
214
214
 
215
+ ### mcp-surface
216
+
217
+ How many MCP servers are charged in the project the caller is standing in,
218
+ counted across three places that all cost the same: the global blocks in
219
+ `~/.claude.json` and `~/.claude/settings.json`, the per-project block inside
220
+ `~/.claude.json`, and a `.mcp.json` committed to the repo. INFO once the count
221
+ exceeds `prefs.global.mcpSurface.infoAbove` (default 8); `--explain` lists the
222
+ names with the scope each came from.
223
+
224
+ Servers registered by a marketplace plugin are NOT counted: nothing in the
225
+ config names them, and inventing a number is worse than reporting the one that
226
+ is countable.
227
+
228
+ The reason it exists: every registered server sends its tool list on every turn,
229
+ the user adds them one at a time, and nobody ever sees the running total - our
230
+ own toolkit contributes 99 tools by itself. This is the same argument that made
231
+ the pre-run context budget a measured number rather than an intention.
232
+
233
+ It only reports. It never disables a server, never blocks and never warns: how
234
+ many servers are worth their context is the user's call, not a health failure.
235
+ The threshold is judgement, which is why it is a pref instead of a constant in
236
+ the script - a number nobody can see is a number nobody can argue with.
237
+
215
238
  ### disk-space
216
239
 
217
240
  Free space on the volume holding `$HOME`. WARN under 2 GB: a worktree plus a
@@ -0,0 +1,166 @@
1
+ # Feature: Maturity Follow-Up
2
+
3
+ <!-- toc -->
4
+ - [1. The rule everything else follows](#1-the-rule-everything-else-follows)
5
+ - [2. Interactive: ask at the step, do not halt at it](#2-interactive-ask-at-the-step-do-not-halt-at-it)
6
+ - [3. Autopilot: ask on the item, then stop](#3-autopilot-ask-on-the-item-then-stop)
7
+ - [4. Resuming into the step, not past it](#4-resuming-into-the-step-not-past-it)
8
+ - [5. State](#5-state)
9
+ <!-- /toc -->
10
+
11
+ **Pattern**: the maturity check has always produced a machine-readable gap list -
12
+ stable codes in `blockers[]` and `warnings[]` - and then thrown most of it away.
13
+ A blocker halted the run, an autopilot queue moved to the next item, and the
14
+ issue stayed exactly as immature as it was found. Nobody was told, so nothing
15
+ changed, so the next scan halted on the same issue for the same reason. The
16
+ check was doing its job and producing no effect.
17
+
18
+ Three behaviours, one decision function
19
+ (`$HOME/.claude/scripts/maturity-followup.mjs`, pure - no network, no issue API,
20
+ no clock unless handed one). Asserted by `smoke-maturity-followup.sh` and
21
+ `test/maturity-followup.test.mjs`.
22
+
23
+ ## 1. The rule everything else follows
24
+
25
+ **An edit is a reason to look again. It is never proof that the gap closed.**
26
+
27
+ A reply reading "will do later" moves the artifact's timestamp and fixes
28
+ nothing. So a changed artifact re-runs the maturity check against the new
29
+ content and the CHECK decides. Nothing in this feature infers maturity from the
30
+ fact that something moved, and the decision function is handed a freshly scored
31
+ `maturity` on every pass for exactly that reason.
32
+
33
+ The corollary is the second comment. A run that re-comments on every scan turns
34
+ an issue into a wall of identical bot text, so:
35
+
36
+ | Situation | What happens |
37
+ |---|---|
38
+ | First pass, gaps present | comment once |
39
+ | Same gaps, artifact untouched | say nothing |
40
+ | Same gaps, artifact edited | say nothing - the re-check already ran and they survived |
41
+ | **Different** gaps | comment - a different question is new information |
42
+ | No gaps | proceed; development starts |
43
+
44
+ **Where "have we already asked" comes from.** The item, not our state file. An
45
+ autopilot scan is a NEW run with a fresh `agent-state.json`, so deriving it from
46
+ state alone would make every scan a first ask - the wall of identical bot
47
+ comments this table exists to prevent. So the comment carries its own gap set on
48
+ a last line, `multi-agent gaps: code,code`, and the next pass reads the item's
49
+ comments and takes the newest one of ours (`priorFromComments`). `state.maturityFollowup`
50
+ is a cache of the same answer for the run that wrote it, never the source.
51
+
52
+ A comment of ours carrying no gap line - written before v17.6.0, or edited by
53
+ hand - reads as "asked, about something we can no longer name": an empty gap set,
54
+ which never equals a live one, so the next scan asks again WITH the codes instead
55
+ of staying silent forever on an unreadable record.
56
+
57
+ "Cannot tell whether it moved" (a tracker whose API omits the timestamp, an
58
+ unparseable value) resolves to *re-check*, never to *wait*. Folding unknown into
59
+ "nothing changed" parks a run forever on a host that never told us anything.
60
+
61
+ ## 2. Interactive: ask at the step, do not halt at it
62
+
63
+ A blocker used to end the run with a summary. It now asks, at the maturity step,
64
+ with the gap as the question. The options are real choices and meet the
65
+ two-option floor on their own (`picker-contract.md`, "Two options or it is not a
66
+ question"):
67
+
68
+ | Option | What it does |
69
+ |---|---|
70
+ | Open the item and fix it | halts, prints the item URL, resumes into this same step |
71
+ | Continue without it | proceeds, and records WHICH gap was accepted in `state.maturity.accepted[]` |
72
+ | Abort | no worktree, no branch, no state file |
73
+
74
+ `prefs.global.maturityFollowup.askInteractively` (default `true`) turns this back into
75
+ the old halt.
76
+
77
+ **What an answer here does not do.** An answer typed into a picker improves this
78
+ run and leaves the item as immature as it was for the next person. That is a
79
+ real cost, not an oversight, and the step says so: after an answer that supplies
80
+ missing content, it offers to write that content back to the item - as a
81
+ separate, individually approved write, per the standing rule that every Jira
82
+ write is approved on its own.
83
+
84
+ ## 3. Autopilot: ask on the item, then stop
85
+
86
+ `autopilotCommentsOnIssue` (**default `false`**) lets an autopilot run post one
87
+ comment on the item asking for what is missing. It is an outward-facing write,
88
+ so it carries the same fence as every other one in this pipeline:
89
+
90
+ - **Off by default.** Nothing posts unless the user turned it on.
91
+ - **A question, never a state change.** No transition, no resolution, no
92
+ assignee, no label, no close - ever. The standing rule that this pipeline
93
+ never auto-closes an issue is not relaxed by a feature that writes comments.
94
+ - **One comment.** The marker line makes the next scan able to recognise its own
95
+ prior comment; matching on the marker rather than on authorship is what keeps
96
+ that working when the token belongs to a shared service account.
97
+ - **No square brackets in the marker or the gap line.** `[text]` is a LINK in
98
+ Jira wiki markup, and this comment is most likely to be posted exactly there,
99
+ so a bracketed marker renders as a broken link to a page nobody created.
100
+ - **`Ref:`, never `Closes:`/`Fixes:`/`Resolves:`**, so no platform-side
101
+ automation reads a question as an instruction.
102
+ - **Human-facing copy follows `outputLanguage`**, and the gap wording is the
103
+ fetcher's own `maturity.summary` verbatim. Re-deriving those labels here would
104
+ give the project two copies of one table and only one would be maintained.
105
+ - **Then it stops.** The run halts on the circuit breaker (`features/autopilot-circuit-breaker.md`),
106
+ which is the sanctioned autopilot pause: state recorded, one actionable line
107
+ printed, waiting for `resume`. Posting a question and continuing on a guess is
108
+ worse than not asking - the guess lands in a branch while the question sits
109
+ unanswered.
110
+
111
+ **Not a second readiness reviewer.** `/multi-agent:review-jira` and
112
+ `/multi-agent:review-issue` also post a gap list, and they are a different thing: a
113
+ human invokes them ON PURPOSE to review an item, with the full readiness rubric
114
+ (`readiness-review.md`) behind the verdict. This comment is a side effect of a
115
+ development run that could not start, carries only the fetcher's own blocker codes,
116
+ and posts at most once. Both obey the same tone contract (`channels/issue-comment.md`):
117
+ no AI attribution, `Ref:` never a closing keyword, copy in `outputLanguage`.
118
+
119
+ **Warnings still auto-continue.** Converting every warning into a halt would
120
+ stall queues overnight on items that ran fine yesterday, so blockers are
121
+ actionable by default and `prefs.global.maturityFollowup.commentOnWarnings` raises
122
+ warnings to the same treatment. Either way the gaps are recorded, so the next pass can compare.
123
+
124
+ ## 4. Resuming into the step, not past it
125
+
126
+ `/multi-agent:resume` starts from `currentPhase + 1`. A run that halted at the
127
+ maturity step has `currentPhase: 0`, so resuming would start at Phase 1 and skip
128
+ the check - the halt would be permanent in the one direction that matters.
129
+
130
+ So resume reads `state.waitingFor` first: when it names a step, the run re-enters
131
+ THAT step rather than the next phase. `waitingFor` already existed and Phase 7's
132
+ channels pause already documented itself as resumable through it
133
+ (`phases/phase-7-report.md`), while `resume/SKILL.md` never mentioned the field -
134
+ so that pause had the same gap and this fixes both.
135
+
136
+ | `waitingFor` | Re-entry |
137
+ |---|---|
138
+ | `maturity` | Phase 0, the maturity step, with the item re-fetched |
139
+ | `user-channels-choice` | Phase 7, the channels menu |
140
+ | absent | `currentPhase + 1`, as before |
141
+
142
+ `waitingFor` is cleared by the write that records the answer. A field that
143
+ outlives its question sends every later resume back to the step the user already
144
+ answered.
145
+
146
+ ## 5. State
147
+
148
+ ```jsonc
149
+ "maturity": {
150
+ "score": 60, // null for free-text: nothing to score
151
+ "blockers": ["description_empty"],
152
+ "warnings": [],
153
+ "summary": "...", // localized by the fetcher, used verbatim
154
+ "accepted": ["short_description"] // gaps a human waved through, interactive only
155
+ },
156
+ "maturityFollowup": {
157
+ "gaps": ["description_empty"], // sorted + deduplicated, so comparison is stable
158
+ "askedAt": "2026-09-15T11:00:00Z",
159
+ "target": { "kind": "jira", "key": "PROJ-1234", "url": "..." },
160
+ "commentUrl": "..."
161
+ }
162
+ ```
163
+
164
+ `maturityFollowup` exists only after a comment was posted, and it is a cache: the
165
+ authoritative record of what was asked is the comment on the item itself, because
166
+ that is the only store the next run can see.
@@ -0,0 +1,80 @@
1
+ # Feature: Package Manager Resolution
2
+
3
+ **Pattern**: the node-shaped arms of Phase 3 and the verify-by-test loop typed
4
+ `npm` into the command line. A repo on pnpm, yarn or bun then gets one of two
5
+ outcomes, both bad: the command fails outright, or npm resolves against a lock
6
+ file it does not own and the run continues on a tree the repo's own tooling
7
+ would never have produced. Either way it happens in Phase 3, with a worktree and
8
+ a branch already created - the failure shape `docs/adr/0012-macos-only.md`
9
+ rejected for platforms.
10
+
11
+ `$HOME/.claude/scripts/package-manager.mjs` resolves it from the repo. Node core
12
+ only (ADR-0004): no corepack call, no spawn, no network - a resolver that shelled
13
+ out would need a working install of the very tool it is identifying. Asserted by
14
+ `smoke-package-manager.sh` and `test/package-manager.test.mjs`.
15
+
16
+ ## 1. Resolution order
17
+
18
+ | # | Evidence | Reported `source` |
19
+ |---|---|---|
20
+ | 1 | `$MA_PACKAGE_MANAGER` | `env` |
21
+ | 2 | `package.json` `"packageManager"` (corepack's own field) | `packageManager-field` |
22
+ | 3 | a lock file (`pnpm-lock.yaml`, `yarn.lock`, `bun.lockb`/`bun.lock`, `package-lock.json`, `npm-shrinkwrap.json`) | `lockfile` |
23
+ | 4 | npm | `default` |
24
+
25
+ What the repo **said** outranks what the repo **left behind**: a stale lock file
26
+ outlives a migration and a declaration does not. The default is reported AS a
27
+ default, never as evidence - "npm because nothing said otherwise" and "npm
28
+ because the repo committed a package-lock" are different answers to the same
29
+ question, and only one of them is safe to act on twice.
30
+
31
+ The walk goes upward from the given directory and stops after the directory
32
+ holding `.git`. A monorepo keeps its lock file at the root while the task edits a
33
+ package three levels down, so stopping at the starting directory would resolve to
34
+ the default for most real repos; going past the repo root would let a stray
35
+ `yarn.lock` in a home directory decide how somebody's project builds.
36
+
37
+ **Two lock files** means a migration left one behind. The newest wins and BOTH
38
+ are reported (`source: lockfile-newest`, `ambiguous: [...]`): silently picking one
39
+ of two committed lock files is how a repo ends up building with the manager it
40
+ migrated away from.
41
+
42
+ ## 2. The command lines
43
+
44
+ - **`run` for every manager**, always: `pnpm build` and `yarn build` work only
45
+ until a script shares a name with a builtin (`test`, `add`, `install`), and
46
+ then the builtin wins and the repo's own script never runs.
47
+ - **Only npm needs `--`** before pass-through arguments. Adding it for the others
48
+ hands the test runner a literal `--` to ignore.
49
+ - **`bun run test`, never `bun test`**: the latter is bun's own runner and would
50
+ ignore the script the repo declared.
51
+ - **No `--frozen-lockfile` / `--immutable`**: that is a CI decision, not ours.
52
+ - **Never `eval "$(pm ...)"` on its own.** Exit 3 empties the command
53
+ substitution, and `eval ""` SUCCEEDS - so a repo with no build script would
54
+ report a build that never ran, which is the failure this feature exists to
55
+ stop, wearing different clothes. Capture first, then eval on success:
56
+ `CMD=$(... ) && eval "$CMD" || echo "no build script"`.
57
+ - **Exit 3 means the repo declares no such script.** That is the `--if-present`
58
+ case, answered by an exit code rather than by a flag whose support differs per
59
+ manager. The caller skips the step and says so; it never substitutes a
60
+ different command.
61
+
62
+ ## 3. The resolved name goes through `eval`
63
+
64
+ The phase runs the printed line through `eval`, so the name is held to the shape
65
+ a manager's binary actually has (`^[a-z][a-z0-9-]*$`). A `packageManager` field
66
+ or an `MA_PACKAGE_MANAGER` value that does not match is dropped with a warning
67
+ and the resolution continues from the repo's own evidence. `smoke-package-manager.sh`
68
+ proves this the only way that counts: it evals the produced line with every real
69
+ manager stubbed out and asserts the crafted payload did not run.
70
+
71
+ An unknown but well-formed name (`deno`, say) resolves and is reported with
72
+ `known: false`, so the caller can say WHICH unrecognised manager it saw instead
73
+ of quietly falling back to npm.
74
+
75
+ ## 4. What is out of scope
76
+
77
+ iOS and Android are untouched: `xcodebuild` and `./gradlew` are not package
78
+ managers and nothing about this changes them. Installing dependencies is not
79
+ automated either - `installCommand()` exists for a caller that has decided to
80
+ install, and no phase calls it today.
@@ -0,0 +1,79 @@
1
+ # Feature: Operational Reporting
2
+
3
+ **Pattern**: reporting needs two things - `usageLog.enabled` true AND a token
4
+ that resolves - and both were arranged automatically in exactly one place:
5
+ `/multi-agent:update`, as forty lines of shell embedded in the skill. A user who
6
+ installed the package, ran `/multi-agent:setup` and worked for weeks never ran
7
+ update, so they never registered, never reported, and the panel could not tell
8
+ them apart from nobody using the pipeline at all.
9
+
10
+ Registration is now one call (`$HOME/.claude/scripts/usage-register.mjs`) made
11
+ from the three places a machine can first become real: **setup**, **update**, and
12
+ **the Phase 0 exit gate** of a run on a machine that reached neither. Asserted by
13
+ `smoke-usage-register.sh`.
14
+
15
+ ## 1. What is sent, and what never is
16
+
17
+ `usage-report.mjs` emits coarse run metadata: task id, phase, status, durations,
18
+ token counts, the credential-health summary. Never prompts, never code, never
19
+ diffs, never absolute paths. The registration call sends two fields: the
20
+ reporting user and the short hostname.
21
+
22
+ The reporting user is the **GitHub login** - `identities[0].username`, then
23
+ `gh api user`, then the OS user. Never `identity.name`, which carries a person's
24
+ real name and sometimes a corporate title.
25
+
26
+ ## 2. The token
27
+
28
+ Requested, never shipped. `/register` mints a per-machine **write-only** token:
29
+ append-only to the ingest endpoint, no read access, no other scope. Only its
30
+ sha256 hash is stored server-side, so a database leak exposes no usable
31
+ credential, and the owner can revoke one row without touching anyone else.
32
+
33
+ It lands in the OS credential store under `<user>_Usage_Ingest_Token`. Prefs hold
34
+ the NAME of that entry (`keychainMapping.usage_ingest`) and the on-switch, never
35
+ the secret. Resolution order at emit time: `$MULTI_AGENT_USAGE_TOKEN`, then
36
+ `usageLog.token`, then the credential-store entry.
37
+
38
+ ## 3. Opting out, and the two silences
39
+
40
+ `usageLog.optOut: true` blocks registration permanently and is checked before
41
+ anything else - before the network call, before the credential store.
42
+
43
+ The other silence is not a choice: offline, endpoint down, ingest disabled by the
44
+ admin, or a credential store that refuses the write. That leaves reporting off
45
+ with one status line and exit 0. **A caller is never failed over bookkeeping**,
46
+ which is the same rule the capture hooks follow.
47
+
48
+ Both are reported distinguishably (`--json` gives `status`: `skipped` with the
49
+ reason, `unavailable` with the cause, `registered`, `enabled`, `dry-run`) because
50
+ "you turned it off" and "we could not reach the endpoint" are different facts
51
+ about the same empty panel.
52
+
53
+ `prefs.global.usageLog.endpoint` overrides where both calls go - the register URL
54
+ is derived from it, so a self-hosted ingest gets its own registration rather than
55
+ this one's. Absent means the shipped default, and every shipped default names the
56
+ same host on purpose: a machine that registers against one host and reports to
57
+ another shows up as a token that never sends anything.
58
+
59
+ ## 3b. Feedback is not telemetry
60
+
61
+ `/multi-agent:feedback` rides the same token, and `optOut` does not silence it:
62
+ passive collection is a choice, a message somebody typed to be read is not. So a
63
+ feedback run may register (`--feedback`) on a machine that opted out - and when it
64
+ does, it writes the credential-store entry and **leaves `usageLog.enabled` alone**.
65
+ The opt-out still holds for everything it was about; the person just gets their
66
+ message delivered.
67
+
68
+ ## 4. The half-configured case
69
+
70
+ A token in the credential store with `enabled: false` produces exactly the same
71
+ silence as no token at all, and it happens whenever a run is interrupted between
72
+ the two writes. The call repairs it: when a token already resolves but the switch
73
+ is off, it turns the switch on and says so rather than reporting "unchanged".
74
+
75
+ ## 5. Where it is NOT called
76
+
77
+ The installer. `install.js` lays down files and nothing else; seeding state is the
78
+ one thing the install contract forbids, and a fresh machine has no preferences
79
+ file for the registration to write into. Setup creates it; registration follows.
@@ -5,7 +5,7 @@
5
5
  **Gated by `prefs.global.verifyByTest.enabled`** (default: `false`). When enabled, after triage 3.6 and before Step 4, IF the validated triage output contains at least one `accepted` blocking finding:
6
6
 
7
7
  1. Dispatch ONE verifier sub-agent for the iteration (model: `verifyByTest.model`, default `sonnet`) - never one dispatch per finding. Input: up to `verifyByTest.maxFindings` (default 3) accepted blocking findings, the diff hunks for their files, and the Phase 1 test conventions.
8
- 2. Per finding, the verifier writes ONE minimal repro test asserting the correct behavior the finding claims is broken, then runs ONLY that test via the Phase 3 single-test invocation (`xcodebuild test -only-testing:`, `pytest {file}::{name}`, `npm test -- --testPathPattern=`, `./gradlew test --tests`) under `acquire_build_lock`/`release_build_lock`. One green run is not trusted: the same single-test invocation runs `verifyByTest.repeatCount` times (default 3; value 1 disables the repeat) as a shell loop, each run's log tee'd to `$WORKTREE/.pipeline/verify-<i>-<k>.test.log` for `k` in `1..repeatCount`. ALL runs must pass, and each log must clear `evidence-gate.mjs --claim test --status passed`, before the outcome counts as "passes". A first run that fails as predicted ends the loop early (`confirmed` needs no repeat).
8
+ 2. Per finding, the verifier writes ONE minimal repro test asserting the correct behavior the finding claims is broken, then runs ONLY that test via the Phase 3 single-test invocation (`xcodebuild test -only-testing:`, `pytest {file}::{name}`, the resolved node command from `scripts/package-manager.mjs test` (npm/pnpm/yarn/bun, never assumed), `./gradlew test --tests`) under `acquire_build_lock`/`release_build_lock`. One green run is not trusted: the same single-test invocation runs `verifyByTest.repeatCount` times (default 3; value 1 disables the repeat) as a shell loop, each run's log tee'd to `$WORKTREE/.pipeline/verify-<i>-<k>.test.log` for `k` in `1..repeatCount`. ALL runs must pass, and each log must clear `evidence-gate.mjs --claim test --status passed`, before the outcome counts as "passes". A first run that fails as predicted ends the loop early (`confirmed` needs no repeat).
9
9
  3. Stamp each processed finding with a `verification` object (triage-output schema v3.2.0) and re-run `validate-triage.mjs` on the mutated triage file under the standard 3.2.1 gate protocol.
10
10
  4. Findings beyond `maxFindings` keep their judgment-only verdict (log `verify_by_test=cap-exceeded`).
11
11
  5. The whole step is bounded by `verifyByTest.stepTimeoutSec` (default 600); on breach or verifier crash, remaining findings keep judgment-only verdicts and the pipeline proceeds. Never blocks.
@@ -706,11 +706,14 @@ Phase 0 owns `agent-state.json`. Do not call
706
706
 
707
707
  ```bash
708
708
  node "$HOME/.claude/scripts/phase0-exit-gate.mjs" "$TASK_ID" --input "$ORIGINAL_INPUT"
709
+ node "$HOME/.claude/scripts/usage-register.mjs" --quiet >/dev/null 2>&1 || true
709
710
  node "$HOME/.claude/scripts/usage-report.mjs" --task-id "$TASK_ID" >/dev/null 2>&1 || true
710
711
  ```
711
712
 
712
- The second line reports the run as started: reporting only from Phase 7 reported
713
- only runs that finish, and few do. Phase 7 upserts the same key over it.
713
+ The third line reports the run as started: reporting only from Phase 7 reported
714
+ only runs that finish, and few do. Phase 7 upserts the same key over it. The
715
+ second is the backstop for a machine that reached neither setup nor update - it
716
+ is a no-op once a token resolves, and permanently so under `usageLog.optOut`.
714
717
 
715
718
  It asserts five things, each of which has failed silently in a real run:
716
719