@mmerterden/multi-agent-pipeline 17.5.1 → 18.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (134) hide show
  1. package/CHANGELOG.md +276 -0
  2. package/README.md +59 -1
  3. package/README.tr.md +57 -0
  4. package/docs/adr/0011-dormant-ci.md +25 -1
  5. package/docs/features.md +24 -0
  6. package/docs/server-readiness.md +188 -0
  7. package/docs/token-budget-history.md +1 -1
  8. package/index.js +16 -1
  9. package/install/_common.mjs +42 -17
  10. package/install/_dev-only-files.mjs +8 -0
  11. package/install/_unattended-profile.mjs +113 -0
  12. package/install/index.mjs +48 -0
  13. package/install/templates/claude-hooks.json +13 -1
  14. package/manifest.json +1049 -0
  15. package/package.json +5 -2
  16. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  17. package/pipeline/commands/multi-agent/feedback/SKILL.md +7 -1
  18. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  19. package/pipeline/commands/multi-agent/issue/SKILL.md +13 -1
  20. package/pipeline/commands/multi-agent/jira/SKILL.md +13 -1
  21. package/pipeline/commands/multi-agent/resume/SKILL.md +16 -1
  22. package/pipeline/commands/multi-agent/setup/SKILL.md +14 -16
  23. package/pipeline/commands/multi-agent/status/SKILL.md +52 -21
  24. package/pipeline/commands/multi-agent/update/SKILL.md +13 -56
  25. package/pipeline/lib/_jira-auth.sh +8 -0
  26. package/pipeline/lib/analysis-jira-write.sh +32 -0
  27. package/pipeline/lib/ask-choice.sh +13 -2
  28. package/pipeline/lib/autopilot-state.sh +8 -0
  29. package/pipeline/lib/fatal.mjs +129 -0
  30. package/pipeline/lib/figma-mcp-refresh.sh +18 -0
  31. package/pipeline/lib/figma-screenshot.sh +18 -0
  32. package/pipeline/lib/invoked-directly.mjs +43 -0
  33. package/pipeline/lib/jira-publish.sh +42 -0
  34. package/pipeline/lib/md2confluence-v3.py +47 -0
  35. package/pipeline/lib/outbound-gate.mjs +175 -0
  36. package/pipeline/lib/plan-todos.sh +27 -6
  37. package/pipeline/lib/post-pr-review.sh +77 -8
  38. package/pipeline/lib/repo-hygiene.sh +8 -3
  39. package/pipeline/lib/require-jq.sh +40 -0
  40. package/pipeline/lib/run-paths.sh +335 -0
  41. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +70 -0
  42. package/pipeline/multi-agent-refs/features/code-graph.md +20 -0
  43. package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
  44. package/pipeline/multi-agent-refs/features/doctor.md +68 -0
  45. package/pipeline/multi-agent-refs/features/maturity-followup.md +166 -0
  46. package/pipeline/multi-agent-refs/features/package-manager.md +80 -0
  47. package/pipeline/multi-agent-refs/features/usage-reporting.md +79 -0
  48. package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
  49. package/pipeline/multi-agent-refs/features/verify.md +83 -0
  50. package/pipeline/multi-agent-refs/phases/operations.md +13 -2
  51. package/pipeline/multi-agent-refs/phases/phase-0-init.md +6 -3
  52. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +8 -2
  53. package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -1
  54. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  55. package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
  56. package/pipeline/preferences-template.json +1 -1
  57. package/pipeline/schemas/agent-state.schema.json +122 -11
  58. package/pipeline/schemas/prefs.schema.json +35 -0
  59. package/pipeline/schemas/token-budget.json +2 -2
  60. package/pipeline/scripts/_run-paths.mjs +372 -0
  61. package/pipeline/scripts/aggregate-metrics.mjs +64 -64
  62. package/pipeline/scripts/autopilot-arming.mjs +2 -1
  63. package/pipeline/scripts/autopilot-intake.mjs +2 -1
  64. package/pipeline/scripts/autopilot-runner.mjs +206 -2
  65. package/pipeline/scripts/build-references.mjs +2 -1
  66. package/pipeline/scripts/build-stack-plugins.mjs +10 -2
  67. package/pipeline/scripts/capture-evidence.sh +7 -2
  68. package/pipeline/scripts/classify-plan-safety.mjs +2 -1
  69. package/pipeline/scripts/cost-analyze.mjs +600 -0
  70. package/pipeline/scripts/cost-budget-check.mjs +4 -12
  71. package/pipeline/scripts/council-view.mjs +2 -1
  72. package/pipeline/scripts/crush-json.mjs +2 -1
  73. package/pipeline/scripts/diff-explain.mjs +6 -9
  74. package/pipeline/scripts/diff-risk-score.mjs +2 -1
  75. package/pipeline/scripts/doctor.mjs +203 -4
  76. package/pipeline/scripts/evidence-gate.mjs +9 -3
  77. package/pipeline/scripts/feedback-send.mjs +13 -3
  78. package/pipeline/scripts/gc-abandoned.sh +29 -13
  79. package/pipeline/scripts/gc-worktrees.sh +11 -4
  80. package/pipeline/scripts/github-ssh-setup.sh +64 -7
  81. package/pipeline/scripts/graph-mermaid.mjs +4 -2
  82. package/pipeline/scripts/graph-report.mjs +155 -1
  83. package/pipeline/scripts/keychain-save.sh +101 -30
  84. package/pipeline/scripts/learn-from-transcripts.mjs +2 -1
  85. package/pipeline/scripts/learning-curve.mjs +34 -29
  86. package/pipeline/scripts/make-manifest.mjs +199 -0
  87. package/pipeline/scripts/maturity-followup.mjs +294 -0
  88. package/pipeline/scripts/migrate-prefs.mjs +2 -1
  89. package/pipeline/scripts/migrate-state.mjs +94 -4
  90. package/pipeline/scripts/package-manager.mjs +310 -0
  91. package/pipeline/scripts/phase-banner.sh +6 -2
  92. package/pipeline/scripts/phase-tracker.sh +41 -3
  93. package/pipeline/scripts/plan-coverage-gate.mjs +6 -2
  94. package/pipeline/scripts/pre-commit-check.sh +7 -0
  95. package/pipeline/scripts/pre-push-check.sh +7 -0
  96. package/pipeline/scripts/purge.sh +23 -6
  97. package/pipeline/scripts/render-agent-log-cost.sh +9 -2
  98. package/pipeline/scripts/render-cost-summary.sh +9 -2
  99. package/pipeline/scripts/render-work-summary.sh +11 -4
  100. package/pipeline/scripts/review-file-filter.mjs +4 -2
  101. package/pipeline/scripts/review-scope.mjs +2 -1
  102. package/pipeline/scripts/routine-registry.mjs +2 -1
  103. package/pipeline/scripts/run-aggregator.mjs +13 -14
  104. package/pipeline/scripts/run-metrics.mjs +3 -1
  105. package/pipeline/scripts/runs-index.mjs +343 -0
  106. package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
  107. package/pipeline/scripts/search-logs.sh +18 -0
  108. package/pipeline/scripts/test-gap-scan.mjs +2 -1
  109. package/pipeline/scripts/test-integrity-gate.mjs +2 -1
  110. package/pipeline/scripts/update-issue-progress.sh +56 -7
  111. package/pipeline/scripts/usage-register.mjs +271 -0
  112. package/pipeline/scripts/usage-report.mjs +14 -3
  113. package/pipeline/scripts/validate-analysis-doc.mjs +2 -1
  114. package/pipeline/scripts/validate-code-graph.mjs +6 -3
  115. package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
  116. package/pipeline/scripts/validate-diff-risk.mjs +6 -3
  117. package/pipeline/scripts/validate-test-gap.mjs +6 -3
  118. package/pipeline/scripts/validate-triage.mjs +3 -1
  119. package/pipeline/scripts/verify-citations.mjs +4 -2
  120. package/pipeline/scripts/verify.mjs +327 -0
  121. package/pipeline/scripts/worktree-finalize.sh +13 -4
  122. package/pipeline/scripts/write-state.mjs +154 -15
  123. package/pipeline/skills/.skill-manifest.json +6 -6
  124. package/pipeline/skills/.skills-index.json +56 -1
  125. package/pipeline/skills/shared/README.md +8 -3
  126. package/pipeline/skills/shared/core/multi-agent-issue/SKILL.md +14 -0
  127. package/pipeline/skills/shared/core/multi-agent-jira/SKILL.md +14 -0
  128. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +13 -0
  129. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +33 -9
  130. package/pipeline/skills/shared/core/multi-agent-update/SKILL.md +6 -0
  131. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
  132. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
  133. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
  134. package/pipeline/skills/skills-index.md +6 -1
@@ -0,0 +1,166 @@
1
+ # Feature: Maturity Follow-Up
2
+
3
+ <!-- toc -->
4
+ - [1. The rule everything else follows](#1-the-rule-everything-else-follows)
5
+ - [2. Interactive: ask at the step, do not halt at it](#2-interactive-ask-at-the-step-do-not-halt-at-it)
6
+ - [3. Autopilot: ask on the item, then stop](#3-autopilot-ask-on-the-item-then-stop)
7
+ - [4. Resuming into the step, not past it](#4-resuming-into-the-step-not-past-it)
8
+ - [5. State](#5-state)
9
+ <!-- /toc -->
10
+
11
+ **Pattern**: the maturity check has always produced a machine-readable gap list -
12
+ stable codes in `blockers[]` and `warnings[]` - and then thrown most of it away.
13
+ A blocker halted the run, an autopilot queue moved to the next item, and the
14
+ issue stayed exactly as immature as it was found. Nobody was told, so nothing
15
+ changed, so the next scan halted on the same issue for the same reason. The
16
+ check was doing its job and producing no effect.
17
+
18
+ Three behaviours, one decision function
19
+ (`$HOME/.claude/scripts/maturity-followup.mjs`, pure - no network, no issue API,
20
+ no clock unless handed one). Asserted by `smoke-maturity-followup.sh` and
21
+ `test/maturity-followup.test.mjs`.
22
+
23
+ ## 1. The rule everything else follows
24
+
25
+ **An edit is a reason to look again. It is never proof that the gap closed.**
26
+
27
+ A reply reading "will do later" moves the artifact's timestamp and fixes
28
+ nothing. So a changed artifact re-runs the maturity check against the new
29
+ content and the CHECK decides. Nothing in this feature infers maturity from the
30
+ fact that something moved, and the decision function is handed a freshly scored
31
+ `maturity` on every pass for exactly that reason.
32
+
33
+ The corollary is the second comment. A run that re-comments on every scan turns
34
+ an issue into a wall of identical bot text, so:
35
+
36
+ | Situation | What happens |
37
+ |---|---|
38
+ | First pass, gaps present | comment once |
39
+ | Same gaps, artifact untouched | say nothing |
40
+ | Same gaps, artifact edited | say nothing - the re-check already ran and they survived |
41
+ | **Different** gaps | comment - a different question is new information |
42
+ | No gaps | proceed; development starts |
43
+
44
+ **Where "have we already asked" comes from.** The item, not our state file. An
45
+ autopilot scan is a NEW run with a fresh `agent-state.json`, so deriving it from
46
+ state alone would make every scan a first ask - the wall of identical bot
47
+ comments this table exists to prevent. So the comment carries its own gap set on
48
+ a last line, `multi-agent gaps: code,code`, and the next pass reads the item's
49
+ comments and takes the newest one of ours (`priorFromComments`). `state.maturityFollowup`
50
+ is a cache of the same answer for the run that wrote it, never the source.
51
+
52
+ A comment of ours carrying no gap line - written before v17.6.0, or edited by
53
+ hand - reads as "asked, about something we can no longer name": an empty gap set,
54
+ which never equals a live one, so the next scan asks again WITH the codes instead
55
+ of staying silent forever on an unreadable record.
56
+
57
+ "Cannot tell whether it moved" (a tracker whose API omits the timestamp, an
58
+ unparseable value) resolves to *re-check*, never to *wait*. Folding unknown into
59
+ "nothing changed" parks a run forever on a host that never told us anything.
60
+
61
+ ## 2. Interactive: ask at the step, do not halt at it
62
+
63
+ A blocker used to end the run with a summary. It now asks, at the maturity step,
64
+ with the gap as the question. The options are real choices and meet the
65
+ two-option floor on their own (`picker-contract.md`, "Two options or it is not a
66
+ question"):
67
+
68
+ | Option | What it does |
69
+ |---|---|
70
+ | Open the item and fix it | halts, prints the item URL, resumes into this same step |
71
+ | Continue without it | proceeds, and records WHICH gap was accepted in `state.maturity.accepted[]` |
72
+ | Abort | no worktree, no branch, no state file |
73
+
74
+ `prefs.global.maturityFollowup.askInteractively` (default `true`) turns this back into
75
+ the old halt.
76
+
77
+ **What an answer here does not do.** An answer typed into a picker improves this
78
+ run and leaves the item as immature as it was for the next person. That is a
79
+ real cost, not an oversight, and the step says so: after an answer that supplies
80
+ missing content, it offers to write that content back to the item - as a
81
+ separate, individually approved write, per the standing rule that every Jira
82
+ write is approved on its own.
83
+
84
+ ## 3. Autopilot: ask on the item, then stop
85
+
86
+ `autopilotCommentsOnIssue` (**default `false`**) lets an autopilot run post one
87
+ comment on the item asking for what is missing. It is an outward-facing write,
88
+ so it carries the same fence as every other one in this pipeline:
89
+
90
+ - **Off by default.** Nothing posts unless the user turned it on.
91
+ - **A question, never a state change.** No transition, no resolution, no
92
+ assignee, no label, no close - ever. The standing rule that this pipeline
93
+ never auto-closes an issue is not relaxed by a feature that writes comments.
94
+ - **One comment.** The marker line makes the next scan able to recognise its own
95
+ prior comment; matching on the marker rather than on authorship is what keeps
96
+ that working when the token belongs to a shared service account.
97
+ - **No square brackets in the marker or the gap line.** `[text]` is a LINK in
98
+ Jira wiki markup, and this comment is most likely to be posted exactly there,
99
+ so a bracketed marker renders as a broken link to a page nobody created.
100
+ - **`Ref:`, never `Closes:`/`Fixes:`/`Resolves:`**, so no platform-side
101
+ automation reads a question as an instruction.
102
+ - **Human-facing copy follows `outputLanguage`**, and the gap wording is the
103
+ fetcher's own `maturity.summary` verbatim. Re-deriving those labels here would
104
+ give the project two copies of one table and only one would be maintained.
105
+ - **Then it stops.** The run halts on the circuit breaker (`features/autopilot-circuit-breaker.md`),
106
+ which is the sanctioned autopilot pause: state recorded, one actionable line
107
+ printed, waiting for `resume`. Posting a question and continuing on a guess is
108
+ worse than not asking - the guess lands in a branch while the question sits
109
+ unanswered.
110
+
111
+ **Not a second readiness reviewer.** `/multi-agent:review-jira` and
112
+ `/multi-agent:review-issue` also post a gap list, and they are a different thing: a
113
+ human invokes them ON PURPOSE to review an item, with the full readiness rubric
114
+ (`readiness-review.md`) behind the verdict. This comment is a side effect of a
115
+ development run that could not start, carries only the fetcher's own blocker codes,
116
+ and posts at most once. Both obey the same tone contract (`channels/issue-comment.md`):
117
+ no AI attribution, `Ref:` never a closing keyword, copy in `outputLanguage`.
118
+
119
+ **Warnings still auto-continue.** Converting every warning into a halt would
120
+ stall queues overnight on items that ran fine yesterday, so blockers are
121
+ actionable by default and `prefs.global.maturityFollowup.commentOnWarnings` raises
122
+ warnings to the same treatment. Either way the gaps are recorded, so the next pass can compare.
123
+
124
+ ## 4. Resuming into the step, not past it
125
+
126
+ `/multi-agent:resume` starts from `currentPhase + 1`. A run that halted at the
127
+ maturity step has `currentPhase: 0`, so resuming would start at Phase 1 and skip
128
+ the check - the halt would be permanent in the one direction that matters.
129
+
130
+ So resume reads `state.waitingFor` first: when it names a step, the run re-enters
131
+ THAT step rather than the next phase. `waitingFor` already existed and Phase 7's
132
+ channels pause already documented itself as resumable through it
133
+ (`phases/phase-7-report.md`), while `resume/SKILL.md` never mentioned the field -
134
+ so that pause had the same gap and this fixes both.
135
+
136
+ | `waitingFor` | Re-entry |
137
+ |---|---|
138
+ | `maturity` | Phase 0, the maturity step, with the item re-fetched |
139
+ | `user-channels-choice` | Phase 7, the channels menu |
140
+ | absent | `currentPhase + 1`, as before |
141
+
142
+ `waitingFor` is cleared by the write that records the answer. A field that
143
+ outlives its question sends every later resume back to the step the user already
144
+ answered.
145
+
146
+ ## 5. State
147
+
148
+ ```jsonc
149
+ "maturity": {
150
+ "score": 60, // null for free-text: nothing to score
151
+ "blockers": ["description_empty"],
152
+ "warnings": [],
153
+ "summary": "...", // localized by the fetcher, used verbatim
154
+ "accepted": ["short_description"] // gaps a human waved through, interactive only
155
+ },
156
+ "maturityFollowup": {
157
+ "gaps": ["description_empty"], // sorted + deduplicated, so comparison is stable
158
+ "askedAt": "2026-09-15T11:00:00Z",
159
+ "target": { "kind": "jira", "key": "PROJ-1234", "url": "..." },
160
+ "commentUrl": "..."
161
+ }
162
+ ```
163
+
164
+ `maturityFollowup` exists only after a comment was posted, and it is a cache: the
165
+ authoritative record of what was asked is the comment on the item itself, because
166
+ that is the only store the next run can see.
@@ -0,0 +1,80 @@
1
+ # Feature: Package Manager Resolution
2
+
3
+ **Pattern**: the node-shaped arms of Phase 3 and the verify-by-test loop typed
4
+ `npm` into the command line. A repo on pnpm, yarn or bun then gets one of two
5
+ outcomes, both bad: the command fails outright, or npm resolves against a lock
6
+ file it does not own and the run continues on a tree the repo's own tooling
7
+ would never have produced. Either way it happens in Phase 3, with a worktree and
8
+ a branch already created - the failure shape `docs/adr/0012-macos-only.md`
9
+ rejected for platforms.
10
+
11
+ `$HOME/.claude/scripts/package-manager.mjs` resolves it from the repo. Node core
12
+ only (ADR-0004): no corepack call, no spawn, no network - a resolver that shelled
13
+ out would need a working install of the very tool it is identifying. Asserted by
14
+ `smoke-package-manager.sh` and `test/package-manager.test.mjs`.
15
+
16
+ ## 1. Resolution order
17
+
18
+ | # | Evidence | Reported `source` |
19
+ |---|---|---|
20
+ | 1 | `$MA_PACKAGE_MANAGER` | `env` |
21
+ | 2 | `package.json` `"packageManager"` (corepack's own field) | `packageManager-field` |
22
+ | 3 | a lock file (`pnpm-lock.yaml`, `yarn.lock`, `bun.lockb`/`bun.lock`, `package-lock.json`, `npm-shrinkwrap.json`) | `lockfile` |
23
+ | 4 | npm | `default` |
24
+
25
+ What the repo **said** outranks what the repo **left behind**: a stale lock file
26
+ outlives a migration and a declaration does not. The default is reported AS a
27
+ default, never as evidence - "npm because nothing said otherwise" and "npm
28
+ because the repo committed a package-lock" are different answers to the same
29
+ question, and only one of them is safe to act on twice.
30
+
31
+ The walk goes upward from the given directory and stops after the directory
32
+ holding `.git`. A monorepo keeps its lock file at the root while the task edits a
33
+ package three levels down, so stopping at the starting directory would resolve to
34
+ the default for most real repos; going past the repo root would let a stray
35
+ `yarn.lock` in a home directory decide how somebody's project builds.
36
+
37
+ **Two lock files** means a migration left one behind. The newest wins and BOTH
38
+ are reported (`source: lockfile-newest`, `ambiguous: [...]`): silently picking one
39
+ of two committed lock files is how a repo ends up building with the manager it
40
+ migrated away from.
41
+
42
+ ## 2. The command lines
43
+
44
+ - **`run` for every manager**, always: `pnpm build` and `yarn build` work only
45
+ until a script shares a name with a builtin (`test`, `add`, `install`), and
46
+ then the builtin wins and the repo's own script never runs.
47
+ - **Only npm needs `--`** before pass-through arguments. Adding it for the others
48
+ hands the test runner a literal `--` to ignore.
49
+ - **`bun run test`, never `bun test`**: the latter is bun's own runner and would
50
+ ignore the script the repo declared.
51
+ - **No `--frozen-lockfile` / `--immutable`**: that is a CI decision, not ours.
52
+ - **Never `eval "$(pm ...)"` on its own.** Exit 3 empties the command
53
+ substitution, and `eval ""` SUCCEEDS - so a repo with no build script would
54
+ report a build that never ran, which is the failure this feature exists to
55
+ stop, wearing different clothes. Capture first, then eval on success:
56
+ `CMD=$(... ) && eval "$CMD" || echo "no build script"`.
57
+ - **Exit 3 means the repo declares no such script.** That is the `--if-present`
58
+ case, answered by an exit code rather than by a flag whose support differs per
59
+ manager. The caller skips the step and says so; it never substitutes a
60
+ different command.
61
+
62
+ ## 3. The resolved name goes through `eval`
63
+
64
+ The phase runs the printed line through `eval`, so the name is held to the shape
65
+ a manager's binary actually has (`^[a-z][a-z0-9-]*$`). A `packageManager` field
66
+ or an `MA_PACKAGE_MANAGER` value that does not match is dropped with a warning
67
+ and the resolution continues from the repo's own evidence. `smoke-package-manager.sh`
68
+ proves this the only way that counts: it evals the produced line with every real
69
+ manager stubbed out and asserts the crafted payload did not run.
70
+
71
+ An unknown but well-formed name (`deno`, say) resolves and is reported with
72
+ `known: false`, so the caller can say WHICH unrecognised manager it saw instead
73
+ of quietly falling back to npm.
74
+
75
+ ## 4. What is out of scope
76
+
77
+ iOS and Android are untouched: `xcodebuild` and `./gradlew` are not package
78
+ managers and nothing about this changes them. Installing dependencies is not
79
+ automated either - `installCommand()` exists for a caller that has decided to
80
+ install, and no phase calls it today.
@@ -0,0 +1,79 @@
1
+ # Feature: Operational Reporting
2
+
3
+ **Pattern**: reporting needs two things - `usageLog.enabled` true AND a token
4
+ that resolves - and both were arranged automatically in exactly one place:
5
+ `/multi-agent:update`, as forty lines of shell embedded in the skill. A user who
6
+ installed the package, ran `/multi-agent:setup` and worked for weeks never ran
7
+ update, so they never registered, never reported, and the panel could not tell
8
+ them apart from nobody using the pipeline at all.
9
+
10
+ Registration is now one call (`$HOME/.claude/scripts/usage-register.mjs`) made
11
+ from the three places a machine can first become real: **setup**, **update**, and
12
+ **the Phase 0 exit gate** of a run on a machine that reached neither. Asserted by
13
+ `smoke-usage-register.sh`.
14
+
15
+ ## 1. What is sent, and what never is
16
+
17
+ `usage-report.mjs` emits coarse run metadata: task id, phase, status, durations,
18
+ token counts, the credential-health summary. Never prompts, never code, never
19
+ diffs, never absolute paths. The registration call sends two fields: the
20
+ reporting user and the short hostname.
21
+
22
+ The reporting user is the **GitHub login** - `identities[0].username`, then
23
+ `gh api user`, then the OS user. Never `identity.name`, which carries a person's
24
+ real name and sometimes a corporate title.
25
+
26
+ ## 2. The token
27
+
28
+ Requested, never shipped. `/register` mints a per-machine **write-only** token:
29
+ append-only to the ingest endpoint, no read access, no other scope. Only its
30
+ sha256 hash is stored server-side, so a database leak exposes no usable
31
+ credential, and the owner can revoke one row without touching anyone else.
32
+
33
+ It lands in the OS credential store under `<user>_Usage_Ingest_Token`. Prefs hold
34
+ the NAME of that entry (`keychainMapping.usage_ingest`) and the on-switch, never
35
+ the secret. Resolution order at emit time: `$MULTI_AGENT_USAGE_TOKEN`, then
36
+ `usageLog.token`, then the credential-store entry.
37
+
38
+ ## 3. Opting out, and the two silences
39
+
40
+ `usageLog.optOut: true` blocks registration permanently and is checked before
41
+ anything else - before the network call, before the credential store.
42
+
43
+ The other silence is not a choice: offline, endpoint down, ingest disabled by the
44
+ admin, or a credential store that refuses the write. That leaves reporting off
45
+ with one status line and exit 0. **A caller is never failed over bookkeeping**,
46
+ which is the same rule the capture hooks follow.
47
+
48
+ Both are reported distinguishably (`--json` gives `status`: `skipped` with the
49
+ reason, `unavailable` with the cause, `registered`, `enabled`, `dry-run`) because
50
+ "you turned it off" and "we could not reach the endpoint" are different facts
51
+ about the same empty panel.
52
+
53
+ `prefs.global.usageLog.endpoint` overrides where both calls go - the register URL
54
+ is derived from it, so a self-hosted ingest gets its own registration rather than
55
+ this one's. Absent means the shipped default, and every shipped default names the
56
+ same host on purpose: a machine that registers against one host and reports to
57
+ another shows up as a token that never sends anything.
58
+
59
+ ## 3b. Feedback is not telemetry
60
+
61
+ `/multi-agent:feedback` rides the same token, and `optOut` does not silence it:
62
+ passive collection is a choice, a message somebody typed to be read is not. So a
63
+ feedback run may register (`--feedback`) on a machine that opted out - and when it
64
+ does, it writes the credential-store entry and **leaves `usageLog.enabled` alone**.
65
+ The opt-out still holds for everything it was about; the person just gets their
66
+ message delivered.
67
+
68
+ ## 4. The half-configured case
69
+
70
+ A token in the credential store with `enabled: false` produces exactly the same
71
+ silence as no token at all, and it happens whenever a run is interrupted between
72
+ the two writes. The call repairs it: when a token already resolves but the switch
73
+ is off, it turns the switch on and says so rather than reporting "unchanged".
74
+
75
+ ## 5. Where it is NOT called
76
+
77
+ The installer. `install.js` lays down files and nothing else; seeding state is the
78
+ one thing the install contract forbids, and a fresh machine has no preferences
79
+ file for the registration to write into. Setup creates it; registration follows.
@@ -5,7 +5,7 @@
5
5
  **Gated by `prefs.global.verifyByTest.enabled`** (default: `false`). When enabled, after triage 3.6 and before Step 4, IF the validated triage output contains at least one `accepted` blocking finding:
6
6
 
7
7
  1. Dispatch ONE verifier sub-agent for the iteration (model: `verifyByTest.model`, default `sonnet`) - never one dispatch per finding. Input: up to `verifyByTest.maxFindings` (default 3) accepted blocking findings, the diff hunks for their files, and the Phase 1 test conventions.
8
- 2. Per finding, the verifier writes ONE minimal repro test asserting the correct behavior the finding claims is broken, then runs ONLY that test via the Phase 3 single-test invocation (`xcodebuild test -only-testing:`, `pytest {file}::{name}`, `npm test -- --testPathPattern=`, `./gradlew test --tests`) under `acquire_build_lock`/`release_build_lock`. One green run is not trusted: the same single-test invocation runs `verifyByTest.repeatCount` times (default 3; value 1 disables the repeat) as a shell loop, each run's log tee'd to `$WORKTREE/.pipeline/verify-<i>-<k>.test.log` for `k` in `1..repeatCount`. ALL runs must pass, and each log must clear `evidence-gate.mjs --claim test --status passed`, before the outcome counts as "passes". A first run that fails as predicted ends the loop early (`confirmed` needs no repeat).
8
+ 2. Per finding, the verifier writes ONE minimal repro test asserting the correct behavior the finding claims is broken, then runs ONLY that test via the Phase 3 single-test invocation (`xcodebuild test -only-testing:`, `pytest {file}::{name}`, the resolved node command from `scripts/package-manager.mjs test` (npm/pnpm/yarn/bun, never assumed), `./gradlew test --tests`) under `acquire_build_lock`/`release_build_lock`. One green run is not trusted: the same single-test invocation runs `verifyByTest.repeatCount` times (default 3; value 1 disables the repeat) as a shell loop, each run's log tee'd to `$WORKTREE/.pipeline/verify-<i>-<k>.test.log` for `k` in `1..repeatCount`. ALL runs must pass, and each log must clear `evidence-gate.mjs --claim test --status passed`, before the outcome counts as "passes". A first run that fails as predicted ends the loop early (`confirmed` needs no repeat).
9
9
  3. Stamp each processed finding with a `verification` object (triage-output schema v3.2.0) and re-run `validate-triage.mjs` on the mutated triage file under the standard 3.2.1 gate protocol.
10
10
  4. Findings beyond `maxFindings` keep their judgment-only verdict (log `verify_by_test=cap-exceeded`).
11
11
  5. The whole step is bounded by `verifyByTest.stepTimeoutSec` (default 600); on breach or verifier crash, remaining findings keep judgment-only verdicts and the pipeline proceeds. Never blocks.
@@ -0,0 +1,83 @@
1
+ # verify - is this install the thing that was published
2
+
3
+ The install is a COPY. `install.js` writes the pipeline tree into `~/.claude`,
4
+ `~/.copilot` and `~/.codex`, and from that moment the two halves drift
5
+ independently. Both directions produce bugs that are hard to name:
6
+
7
+ - an edit made in the installed copy is a behaviour with no source, and the next
8
+ update silently reverts it;
9
+ - a file the installer failed to write is a script the docs describe and nobody
10
+ has, which reads as a documentation error.
11
+
12
+ `multi-agent-pipeline verify` answers both mechanically.
13
+
14
+ ```bash
15
+ npx @mmerterden/multi-agent-pipeline verify # package + install
16
+ npx @mmerterden/multi-agent-pipeline verify --package # package integrity only
17
+ npx @mmerterden/multi-agent-pipeline verify --install # install drift only
18
+ npx @mmerterden/multi-agent-pipeline verify --json
19
+ ```
20
+
21
+ | Code | Meaning |
22
+ |---|---|
23
+ | 0 | everything matches |
24
+ | 1 | a difference was found, named file by file |
25
+ | 2 | nothing to verify - a source checkout, or a version published before manifests existed |
26
+
27
+ Exit 2 is not a pass and not a failure. A dev checkout has no manifest by
28
+ design, and reporting that as either would be a lie in one direction or the
29
+ other.
30
+
31
+ ## The manifest
32
+
33
+ `manifest.json` is written at pack time by `prepack`, never committed. A
34
+ manifest in git is stale one commit after it is written, and a stale manifest
35
+ reports honest edits as tampering - which is worse than having none, because
36
+ people learn to ignore it.
37
+
38
+ The file list is not guessed. It comes from `npm pack --dry-run --json`, so by
39
+ construction it is the same set npm publishes, `files` globs and all. The gate
40
+ asserts the two counts agree, which is what catches a `files` entry and a
41
+ manifest that have stopped describing the same package.
42
+
43
+ Two things it cannot cover, said here rather than discovered later: it cannot
44
+ hash itself, and a signature over it does not authenticate the tarball.
45
+
46
+ ## What a green result proves, and what it does not
47
+
48
+ It proves the bytes match what the publisher recorded. It is not proof of WHO
49
+ published them. The manifest, the signature and the verifier all travel inside
50
+ the same tarball, so anyone able to rewrite one can rewrite the others.
51
+ Provenance belongs to npm's own integrity field.
52
+
53
+ What this does catch is the set of failures that actually happen: a damaged or
54
+ partial install, a file edited after install, and an update that did not land.
55
+
56
+ Signing is optional. `make-manifest.mjs --sign` reads an ed25519 private key
57
+ from the credential store (or `MULTI_AGENT_SIGNING_KEY` on a build host with no
58
+ store) and writes `manifest.sig`; `verify` checks it against
59
+ `MULTI_AGENT_SIGNING_PUBKEY` when one is pinned. Without a key it says "signed,
60
+ no public key to check it against" rather than claiming valid - a signature
61
+ nobody can check is not a signature that passed.
62
+
63
+ ## How each tree is compared
64
+
65
+ | Tree | Mode | Why |
66
+ |---|---|---|
67
+ | `scripts` | bytes | verbatim copy, minus the dev-only set |
68
+ | `lib` | bytes | verbatim copy |
69
+ | `multi-agent-refs` | bytes | verbatim copy |
70
+ | `agents` | bytes | verbatim copy |
71
+ | `commands/multi-agent` | presence | `install.js` rewrites each SKILL.md `description` into the user's `outputLanguage` |
72
+
73
+ Byte-comparing `commands/` reports every command as drift on a perfectly
74
+ healthy machine. Measured here: all 57 command files differ, and 56 of them
75
+ differ by nothing except the translated description. A report that is wrong by
76
+ default is a report nobody reads.
77
+
78
+ The dev-only filter matters just as much: smokes, linters and fixtures ship in
79
+ the package and are deliberately NOT installed. Without excluding them, `verify`
80
+ would report 252 files as "the installer skipped this".
81
+
82
+ `~/.copilot` and `~/.codex` are reported as present, not compared: the installer
83
+ rewrites paths for both on purpose, so a byte difference there is the design.
@@ -78,8 +78,9 @@ Update `agent-state.json` at EVERY phase transition.
78
78
 
79
79
  ### Writing `agent-state.json` (required mechanism)
80
80
 
81
- Every state update goes through `write-state.mjs`. Never write the file with a
82
- plain read-modify-write (`jq ... > tmp && mv`, an editor tool, `cat >`):
81
+ Every state write goes through `write-state.mjs`, **including the first one in
82
+ Phase 0**. Never write the file with a plain read-modify-write (`jq ... > tmp &&
83
+ mv`, an editor tool, `cat >`):
83
84
 
84
85
  ```bash
85
86
  # Merge a patch into the current state (the normal case).
@@ -97,6 +98,16 @@ the first landed. `write-state.mjs` does tmpfile + rename (atomic on POSIX) unde
97
98
  an advisory `.lock`, reclaims a lock whose holder PID is dead, and releases the
98
99
  lock on every error path.
99
100
 
101
+ Why the CREATE matters as much as the updates: the writer stamps `rev` on every
102
+ write and `schemaVersion` on the first one. A document written by hand in Phase 0
103
+ starts with neither, so every later writer compares against an absent revision
104
+ and `migrate-state.mjs` can never place the file on a migration path. Measured on
105
+ a real install before this was fixed: 39 of 43 `agent-state.json` files carried no
106
+ `rev` and 43 of 43 carried no `schemaVersion`, which is the whole of
107
+ `$HOME/.claude/schemas/migrations/` sitting unreachable. `migrate-state.mjs --all`
108
+ reports the legacy ones; it does not stamp them, because a stamp would assert a
109
+ conformance nothing checked.
110
+
100
111
  Exit codes the caller must handle: `0` written, `1` invalid JSON on stdin, `2`
101
112
  lock timeout (another writer held it past the acquire window - retry once, then
102
113
  halt per the halt-visibility rule), `3` I/O error.
@@ -514,7 +514,7 @@ done
514
514
 
515
515
  State file in multi-repo mode:
516
516
  - Single shared `agent-state.json` lives at `$HOME/.claude/logs/multi-agent/{first-project}/{task-id}/agent-state.json` (anchored on the first repo for back-compat with `multi-agent log`/`status` commands)
517
- - Every update to it, here and in every later phase, goes through `node $HOME/.claude/scripts/write-state.mjs` - the required mechanism in `operations.md` "Writing `agent-state.json`", and the race a per-repo read-modify-write loses `projects[]` entries to.
517
+ - Every write to it, creation included, goes through `node $HOME/.claude/scripts/write-state.mjs` - the required mechanism in `operations.md` "Writing `agent-state.json`", and the race a per-repo read-modify-write loses `projects[]` entries to.
518
518
  - `state.projects[]` holds per-repo `{name, root, worktreePath, branch, baseBranch, identity, platform, baseFetchStatus, commit, pr, pushAttempts, buildStatus}` - see `agent-state.schema.json`
519
519
  - Scalar fields (`project`, `projectRoot`, `worktreePath`, `branch`, `baseBranch`, `identity`) mirror `projects[0]` so legacy phases that read scalars keep working
520
520
  - Atomicity: if any repo's worktree creation fails (collision aborted, fetch aborted, disk full), roll back already-created worktrees: `git -C $proj worktree remove --force $WT_PATH; git -C $proj branch -D $BRANCH`. Never leave a partial multi-repo state.
@@ -706,11 +706,14 @@ Phase 0 owns `agent-state.json`. Do not call
706
706
 
707
707
  ```bash
708
708
  node "$HOME/.claude/scripts/phase0-exit-gate.mjs" "$TASK_ID" --input "$ORIGINAL_INPUT"
709
+ node "$HOME/.claude/scripts/usage-register.mjs" --quiet >/dev/null 2>&1 || true
709
710
  node "$HOME/.claude/scripts/usage-report.mjs" --task-id "$TASK_ID" >/dev/null 2>&1 || true
710
711
  ```
711
712
 
712
- The second line reports the run as started: reporting only from Phase 7 reported
713
- only runs that finish, and few do. Phase 7 upserts the same key over it.
713
+ The third line reports the run as started: reporting only from Phase 7 reported
714
+ only runs that finish, and few do. Phase 7 upserts the same key over it. The
715
+ second is the backstop for a machine that reached neither setup nor update - it
716
+ is a no-op once a token resolves, and permanently so under `usageLog.optOut`.
714
717
 
715
718
  It asserts five things, each of which has failed silently in a real run:
716
719
 
@@ -133,9 +133,15 @@ For each task (respecting dependency order):
133
133
  release_build_lock ;;
134
134
  android) ./gradlew test --tests "{testClass}.{testMethod}" 2>&1 | tail -5 ;;
135
135
  backend) pytest "{test_file}::{test_name}" 2>&1 | tail -5 ;;
136
- web) npm test -- --testPathPattern="{file}" 2>&1 | tail -5 ;;
136
+ web) CMD=$(node $HOME/.claude/scripts/package-manager.mjs test \
137
+ --dir "{worktreePath}" --pattern "--testPathPattern={file}") \
138
+ && eval "$CMD" 2>&1 | tail -5 || echo "no test script declared" ;;
137
139
  esac
138
140
  ```
141
+ - The node arm resolves the manager instead of typing `npm`; exit 3 means the repo
142
+ declares no such script - say so, never substitute one, and never let the empty
143
+ command substitution pass for a pass (`features/package-manager.md`).
144
+
139
145
  - Must fail for the RIGHT reason (expected assertion, not compilation error)
140
146
 
141
147
  **GREEN - Minimal code to pass:**
@@ -172,7 +178,7 @@ For each task (respecting dependency order):
172
178
  - **ios, preferred (MCP, multi-agent-toolkit >= 3.0.0)**: `acquire_build_lock` → `mcp__multi-agent-toolkit__ios_xcodebuild({project|workspace, scheme, configuration: "Release", destination: "generic/platform=iOS", derived_data_path: "{worktreePath}/.DerivedData"})` → `release_build_lock`. Returns one line `Build: SUCCESS|FAILURE (E errors, W warnings) [xcresult-<id>]`; on failure drill in via `mcp__multi-agent-toolkit__ios_xcresult({id, mode: "errors"})`, never dump the full log.
173
179
  - **ios, fallback (raw)**: same lock pair around `xcodebuild build -scheme "{scheme}" -destination "generic/platform=iOS" -derivedDataPath "{worktreePath}/.DerivedData" 2>&1 | tail -5`.
174
180
  - **android**: lock pair around `./gradlew assembleDebug 2>&1 | tail -5` (the Gradle daemon and `build/` outputs contend across parallel worktrees exactly as DerivedData does - the lock applies).
175
- - **backend / web**: `python -m compileall .` / `npm run build --if-present 2>&1 | tail -5`; no lock.
181
+ - **backend / web**: `python -m compileall .` / `CMD=$(node $HOME/.claude/scripts/package-manager.mjs run --dir "{worktreePath}" --script build) && eval "$CMD" 2>&1 | tail -5 || echo "no build script"`; no lock.
176
182
  5. If build fails → fix → rebuild (max 3 attempts, track `retryCount` in state).
177
183
  6. **Intermediate commit** (after each completed task in the plan):
178
184
  ```bash
@@ -19,7 +19,7 @@ If any gate fails → fix first, don't waste AI tokens reviewing broken code.
19
19
  # Gate 1: Build (xcodebuild/gradle assemble/tsc/py compile - stack-dependent; Xcode uses the build queue lock, see Phase 3) - tee output to a log
20
20
  <build-command> 2>&1 | tee "$WORKTREE/.build.log"
21
21
  # Gate 2: Lint (swiftlint/ktlint/ruff/eslint - stack-dependent)
22
- # Gate 3: Tests pass (xcodebuild test/gradle test/pytest/npm test) - tee output to a log
22
+ # Gate 3: Tests pass (xcodebuild/gradle/pytest/the resolved node command) - tee output to a log
23
23
  <test-command> 2>&1 | tee "$WORKTREE/.test.log"
24
24
  # Gate 4: Secrets - run the scanner against the staged diff
25
25
  bash $HOME/.claude/scripts/pre-commit-check.sh
@@ -160,6 +160,6 @@ asked anything, so the only thing that keeps it accountable is being readable af
160
160
 
161
161
  ## Deterministic gates note
162
162
 
163
- Claude Code's `PreToolUse` exit-2 hooks are the HARD blocking gates. Three ship, none needing run-specific arguments so they are naturally hookable: (1) `pre-commit-check.sh` scans the staged diff on every `git commit` and blocks on a detected secret; (2) `agent-guard.sh` runs on `git commit` + `git push` and blocks AI/assistant attribution in a commit message and force-push to a protected branch (main/master/develop); (3) `check-read-size.sh` runs on `Read` and on the shell commands that read a file whole, and routes an oversized read to a cheap worker (`bulk-read.sh`) instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `bulkRead.mode` is set to `observe` or `enforce`, so merging the block changes nothing until the user opts in. Its `observe` mode blocks nothing and only logs, which is how the baseline is measured before anything is routed. All three are self-contained, fail-open on internal error, and never execute the inspected command. Two capture hooks ship in the same block and block nothing: `SessionEnd` runs `capture-flush.sh --if-stale` (writing a killed run's findings into the per-repo stores, since every durable write used to live in Phase 7 - the phase a run is least likely to reach) plus `note-session.sh` (the mechanical shape of a non-pipeline session: tools used, commands that failed, calls the user refused - never an argument, never any output), and `SessionStart` runs `capture-resume.sh`, at most two lines about an unfinished run and a stale observation queue. Neither calls a model; both exit 0 on every path. The recommended hook block ships at `install/templates/claude-hooks.json`; `multi-agent:setup` offers to merge it into `~/.claude/settings.json`. The other deterministic gates (evidence, consensus, intent, learnings) are invoked by the pipeline phases with per-run arguments (a build-log path, the triage JSON, the free-text input), so they are phase-enforced by contract, not OS-hookable.
163
+ Claude Code's `PreToolUse` exit-2 hooks are the HARD blocking gates. Three ship, none needing run-specific arguments so they are naturally hookable: (1) `pre-commit-check.sh` scans the staged diff on every `git commit` and blocks on a detected secret; (2) `agent-guard.sh` runs on `git commit` + `git push` and blocks AI/assistant attribution in a commit message and force-push to a protected branch (main/master/develop); (3) `check-read-size.sh` runs on `Read` and on the shell commands that read a file whole, and routes an oversized read to a cheap worker (`bulk-read.sh`) instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `bulkRead.mode` is set to `observe` or `enforce`, so merging the block changes nothing until the user opts in. Its `observe` mode blocks nothing and only logs, which is how the baseline is measured before anything is routed. All three are self-contained, fail-open on internal error, and never execute the inspected command. Three capture hooks ship in the same block and block nothing: `SessionEnd` runs `capture-flush.sh --if-stale` (writing a killed run's findings into the per-repo stores, since every durable write used to live in Phase 7 - the phase a run is least likely to reach) plus `note-session.sh` (the mechanical shape of a non-pipeline session: tools used, commands that failed, calls the user refused - never an argument, never any output), `PreCompact` runs `capture-flush.sh` without `--if-stale` (a compaction summarizes a long phase mid-flight, so it is the moment unflushed findings are at risk; the staleness test exists only so a session exit does not re-flush a finished run, and both store writes are idempotent), and `SessionStart` runs `capture-resume.sh`, at most two lines about an unfinished run and a stale observation queue. None of them calls a model; all exit 0 on every path. The recommended hook block ships at `install/templates/claude-hooks.json`; `multi-agent:setup` offers to merge it into `~/.claude/settings.json`. The other deterministic gates (evidence, consensus, intent, learnings) are invoked by the pipeline phases with per-run arguments (a build-log path, the triage JSON, the free-text input), so they are phase-enforced by contract, not OS-hookable.
164
164
 
165
165
  Copilot CLI has no `PreToolUse` equivalent, so the secret scan there is workflow-enforced (run as a phase step, not OS-blocked) plus a CI smoke-gate step.