@mmerterden/multi-agent-pipeline 17.5.1 → 18.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +276 -0
- package/README.md +59 -1
- package/README.tr.md +57 -0
- package/docs/adr/0011-dormant-ci.md +25 -1
- package/docs/features.md +24 -0
- package/docs/server-readiness.md +188 -0
- package/docs/token-budget-history.md +1 -1
- package/index.js +16 -1
- package/install/_common.mjs +42 -17
- package/install/_dev-only-files.mjs +8 -0
- package/install/_unattended-profile.mjs +113 -0
- package/install/index.mjs +48 -0
- package/install/templates/claude-hooks.json +13 -1
- package/manifest.json +1049 -0
- package/package.json +5 -2
- package/pipeline/commands/multi-agent/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/feedback/SKILL.md +7 -1
- package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/issue/SKILL.md +13 -1
- package/pipeline/commands/multi-agent/jira/SKILL.md +13 -1
- package/pipeline/commands/multi-agent/resume/SKILL.md +16 -1
- package/pipeline/commands/multi-agent/setup/SKILL.md +14 -16
- package/pipeline/commands/multi-agent/status/SKILL.md +52 -21
- package/pipeline/commands/multi-agent/update/SKILL.md +13 -56
- package/pipeline/lib/_jira-auth.sh +8 -0
- package/pipeline/lib/analysis-jira-write.sh +32 -0
- package/pipeline/lib/ask-choice.sh +13 -2
- package/pipeline/lib/autopilot-state.sh +8 -0
- package/pipeline/lib/fatal.mjs +129 -0
- package/pipeline/lib/figma-mcp-refresh.sh +18 -0
- package/pipeline/lib/figma-screenshot.sh +18 -0
- package/pipeline/lib/invoked-directly.mjs +43 -0
- package/pipeline/lib/jira-publish.sh +42 -0
- package/pipeline/lib/md2confluence-v3.py +47 -0
- package/pipeline/lib/outbound-gate.mjs +175 -0
- package/pipeline/lib/plan-todos.sh +27 -6
- package/pipeline/lib/post-pr-review.sh +77 -8
- package/pipeline/lib/repo-hygiene.sh +8 -3
- package/pipeline/lib/require-jq.sh +40 -0
- package/pipeline/lib/run-paths.sh +335 -0
- package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +70 -0
- package/pipeline/multi-agent-refs/features/code-graph.md +20 -0
- package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
- package/pipeline/multi-agent-refs/features/doctor.md +68 -0
- package/pipeline/multi-agent-refs/features/maturity-followup.md +166 -0
- package/pipeline/multi-agent-refs/features/package-manager.md +80 -0
- package/pipeline/multi-agent-refs/features/usage-reporting.md +79 -0
- package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
- package/pipeline/multi-agent-refs/features/verify.md +83 -0
- package/pipeline/multi-agent-refs/phases/operations.md +13 -2
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +6 -3
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +8 -2
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -1
- package/pipeline/multi-agent-refs/picker-contract.md +1 -1
- package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
- package/pipeline/preferences-template.json +1 -1
- package/pipeline/schemas/agent-state.schema.json +122 -11
- package/pipeline/schemas/prefs.schema.json +35 -0
- package/pipeline/schemas/token-budget.json +2 -2
- package/pipeline/scripts/_run-paths.mjs +372 -0
- package/pipeline/scripts/aggregate-metrics.mjs +64 -64
- package/pipeline/scripts/autopilot-arming.mjs +2 -1
- package/pipeline/scripts/autopilot-intake.mjs +2 -1
- package/pipeline/scripts/autopilot-runner.mjs +206 -2
- package/pipeline/scripts/build-references.mjs +2 -1
- package/pipeline/scripts/build-stack-plugins.mjs +10 -2
- package/pipeline/scripts/capture-evidence.sh +7 -2
- package/pipeline/scripts/classify-plan-safety.mjs +2 -1
- package/pipeline/scripts/cost-analyze.mjs +600 -0
- package/pipeline/scripts/cost-budget-check.mjs +4 -12
- package/pipeline/scripts/council-view.mjs +2 -1
- package/pipeline/scripts/crush-json.mjs +2 -1
- package/pipeline/scripts/diff-explain.mjs +6 -9
- package/pipeline/scripts/diff-risk-score.mjs +2 -1
- package/pipeline/scripts/doctor.mjs +203 -4
- package/pipeline/scripts/evidence-gate.mjs +9 -3
- package/pipeline/scripts/feedback-send.mjs +13 -3
- package/pipeline/scripts/gc-abandoned.sh +29 -13
- package/pipeline/scripts/gc-worktrees.sh +11 -4
- package/pipeline/scripts/github-ssh-setup.sh +64 -7
- package/pipeline/scripts/graph-mermaid.mjs +4 -2
- package/pipeline/scripts/graph-report.mjs +155 -1
- package/pipeline/scripts/keychain-save.sh +101 -30
- package/pipeline/scripts/learn-from-transcripts.mjs +2 -1
- package/pipeline/scripts/learning-curve.mjs +34 -29
- package/pipeline/scripts/make-manifest.mjs +199 -0
- package/pipeline/scripts/maturity-followup.mjs +294 -0
- package/pipeline/scripts/migrate-prefs.mjs +2 -1
- package/pipeline/scripts/migrate-state.mjs +94 -4
- package/pipeline/scripts/package-manager.mjs +310 -0
- package/pipeline/scripts/phase-banner.sh +6 -2
- package/pipeline/scripts/phase-tracker.sh +41 -3
- package/pipeline/scripts/plan-coverage-gate.mjs +6 -2
- package/pipeline/scripts/pre-commit-check.sh +7 -0
- package/pipeline/scripts/pre-push-check.sh +7 -0
- package/pipeline/scripts/purge.sh +23 -6
- package/pipeline/scripts/render-agent-log-cost.sh +9 -2
- package/pipeline/scripts/render-cost-summary.sh +9 -2
- package/pipeline/scripts/render-work-summary.sh +11 -4
- package/pipeline/scripts/review-file-filter.mjs +4 -2
- package/pipeline/scripts/review-scope.mjs +2 -1
- package/pipeline/scripts/routine-registry.mjs +2 -1
- package/pipeline/scripts/run-aggregator.mjs +13 -14
- package/pipeline/scripts/run-metrics.mjs +3 -1
- package/pipeline/scripts/runs-index.mjs +343 -0
- package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
- package/pipeline/scripts/search-logs.sh +18 -0
- package/pipeline/scripts/test-gap-scan.mjs +2 -1
- package/pipeline/scripts/test-integrity-gate.mjs +2 -1
- package/pipeline/scripts/update-issue-progress.sh +56 -7
- package/pipeline/scripts/usage-register.mjs +271 -0
- package/pipeline/scripts/usage-report.mjs +14 -3
- package/pipeline/scripts/validate-analysis-doc.mjs +2 -1
- package/pipeline/scripts/validate-code-graph.mjs +6 -3
- package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
- package/pipeline/scripts/validate-diff-risk.mjs +6 -3
- package/pipeline/scripts/validate-test-gap.mjs +6 -3
- package/pipeline/scripts/validate-triage.mjs +3 -1
- package/pipeline/scripts/verify-citations.mjs +4 -2
- package/pipeline/scripts/verify.mjs +327 -0
- package/pipeline/scripts/worktree-finalize.sh +13 -4
- package/pipeline/scripts/write-state.mjs +154 -15
- package/pipeline/skills/.skill-manifest.json +6 -6
- package/pipeline/skills/.skills-index.json +56 -1
- package/pipeline/skills/shared/README.md +8 -3
- package/pipeline/skills/shared/core/multi-agent-issue/SKILL.md +14 -0
- package/pipeline/skills/shared/core/multi-agent-jira/SKILL.md +14 -0
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +13 -0
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +33 -9
- package/pipeline/skills/shared/core/multi-agent-update/SKILL.md +6 -0
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
- package/pipeline/skills/skills-index.md +6 -1
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,282 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [18.0.0] - 2026-09-17
|
|
18
|
+
|
|
19
|
+
Major, for two behaviour changes rather than a renamed command: a run's state now
|
|
20
|
+
has ONE canonical directory, and four publish paths refuse to send when the leak
|
|
21
|
+
gate cannot be loaded.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- **`multi-agent-pipeline verify`** answers "is this install the thing that was
|
|
26
|
+
published". The install is a copy, and from the moment `install.js` writes it
|
|
27
|
+
the two halves drift independently: an edit in the installed tree is behaviour
|
|
28
|
+
with no source, and a file the installer skipped is a script the docs describe
|
|
29
|
+
and nobody has. `manifest.json` (SHA-256 per shipped file, version, source
|
|
30
|
+
commit) is built at pack time by `prepack` and never committed - a manifest in
|
|
31
|
+
git is stale one commit after it is written, and a stale manifest reports
|
|
32
|
+
honest edits as tampering.
|
|
33
|
+
|
|
34
|
+
`commands/` is compared by presence, not bytes, because `install.js` rewrites
|
|
35
|
+
each description into the user's `outputLanguage`: measured here, all 57
|
|
36
|
+
command files differ and 56 of them differ by nothing else. Dev-only files are
|
|
37
|
+
excluded, or 252 smokes and linters read as "the installer skipped this".
|
|
38
|
+
|
|
39
|
+
What a green result proves is stated in the output's own reference: the bytes
|
|
40
|
+
match what the publisher recorded. Not who published them - the manifest, the
|
|
41
|
+
signature and the verifier travel in the same tarball, so provenance belongs
|
|
42
|
+
to npm's integrity field. Signing is optional, and an unverifiable signature
|
|
43
|
+
says so rather than claiming valid.
|
|
44
|
+
|
|
45
|
+
- **`cost-analyze.mjs`** - projection, anomaly, burn and diff. `cost-budget-check`
|
|
46
|
+
watches one run against one ceiling, which is blind to both ways a budget
|
|
47
|
+
actually empties: a drift no single run trips, and one session that burns a
|
|
48
|
+
week in an hour while every run stays under its cap.
|
|
49
|
+
|
|
50
|
+
The series is not where the schema says it is. `tracker-state.json` carries
|
|
51
|
+
per-phase token fields that are written only when a phase reports them; of 98
|
|
52
|
+
trackers on a working machine, zero carry any. The dense series is the host's
|
|
53
|
+
own transcripts, so that is what this reads - and the consequences are printed
|
|
54
|
+
rather than buried: it covers everything Claude Code did on the machine, the
|
|
55
|
+
figures are LIST-price estimates rather than a bill, and a host with no
|
|
56
|
+
transcripts reports UNMEASURED instead of zero.
|
|
57
|
+
|
|
58
|
+
Anomalies use the median and the MAD, because the expensive session the check
|
|
59
|
+
exists to find is the observation that inflates a mean and a standard
|
|
60
|
+
deviation - it hides inside the statistic measured against it.
|
|
61
|
+
|
|
62
|
+
- **`install --unattended`** writes a documented permission profile, and prints
|
|
63
|
+
it with a reason per line before writing. autopilot passes
|
|
64
|
+
`--permission-prompts none`, which stops Claude Code asking and grants
|
|
65
|
+
nothing; on a fresh machine the run stops at the first tool call with no
|
|
66
|
+
prompt for anyone to answer, which is indistinguishable from an empty queue. A
|
|
67
|
+
default install still writes no permissions at all.
|
|
68
|
+
|
|
69
|
+
- **`doctor --profile=server`** adds four checks that only matter when nobody is
|
|
70
|
+
at the keyboard: the unattended contract, the permission posture, the
|
|
71
|
+
scheduler, and the keychain. The default run is unchanged - same 17 checks,
|
|
72
|
+
same verdict - because a laptop told it fails a server check learns to ignore
|
|
73
|
+
doctor.
|
|
74
|
+
|
|
75
|
+
- **`docs/server-readiness.md`** sets up nothing and says so first. It covers the
|
|
76
|
+
permission posture, why the scheduler is a LaunchAgent rather than a
|
|
77
|
+
LaunchDaemon (the keychain is locked until login, and that failure arrives
|
|
78
|
+
much later wearing a 401), and which credential each phase needs.
|
|
79
|
+
|
|
80
|
+
- **`scorecard-snapshot.mjs --diff`** keeps what the scorecard said and reports
|
|
81
|
+
what moved. Deliberately no 0-100 score: the scorecard reports twelve measured
|
|
82
|
+
metrics AND four it refuses to measure, and one figure would hide both halves.
|
|
83
|
+
|
|
84
|
+
### Changed
|
|
85
|
+
|
|
86
|
+
- **A run's state has one canonical directory.** `{project}/{taskId}/` is
|
|
87
|
+
canonical and the flat `{taskId}/` is read for compatibility; every reader
|
|
88
|
+
resolves through `lib/run-paths.sh` / `scripts/_run-paths.mjs`, so a run that
|
|
89
|
+
exists in both layouts is counted once.
|
|
90
|
+
|
|
91
|
+
- **Five publish paths refuse rather than send when the leak gate is missing.**
|
|
92
|
+
Jira comments and descriptions, PR review bodies and inline comments, the
|
|
93
|
+
GitHub issue progress comment, the issue-CREATION path, and Confluence page
|
|
94
|
+
create/update all run their outbound text through `lib/outbound-gate.mjs`
|
|
95
|
+
first. A missing gate file refuses; opening the gate because the gate is not
|
|
96
|
+
there would be the one failure mode that matters. The count is part of the
|
|
97
|
+
gate: a sixth publisher cannot ship without the smoke's list naming it.
|
|
98
|
+
|
|
99
|
+
- **`set -e` on the four destructive scripts, and deliberately not on the two
|
|
100
|
+
collectors.** `gc-abandoned`, `gc-worktrees`, `purge` and `worktree-finalize`
|
|
101
|
+
delete things, so the command after an unnoticed failure is the dangerous one.
|
|
102
|
+
`pre-push-check` runs the gates and counts failures, and `pre-commit-check` is
|
|
103
|
+
a hook built out of greps that are supposed to find nothing - under `-e` the
|
|
104
|
+
first clean detector would end the scan and report "no secrets" for a file it
|
|
105
|
+
never finished reading.
|
|
106
|
+
|
|
107
|
+
### Fixed
|
|
108
|
+
|
|
109
|
+
- **A script that dies now prints one line instead of a stack dump, and releases
|
|
110
|
+
what it held.** `lib/fatal.mjs` catches the sync throw, the rejection nobody
|
|
111
|
+
awaited and the throw from inside a callback; the last two are invisible to a
|
|
112
|
+
try/catch around `main()`. `write-state` releases its advisory lock on the way
|
|
113
|
+
out, so the next writer never has to judge a lock on age alone. EPIPE is
|
|
114
|
+
deliberately not fatal: `runs-index.mjs --json | head` is the ordinary way to
|
|
115
|
+
read a large output.
|
|
116
|
+
|
|
117
|
+
- **Build junk no longer reaches an install.** A CI runner shipped 147 files
|
|
118
|
+
where every developer tree shipped 146, for four rounds, and the extra was
|
|
119
|
+
`__pycache__/*.pyc` - untracked, so no diff of the source could show it, and
|
|
120
|
+
copied verbatim into all three install trees. `copyDir` now filters
|
|
121
|
+
`__pycache__`, `*.pyc` and `.DS_Store` on every path.
|
|
122
|
+
|
|
123
|
+
- **`jq` is required rather than optional on the nine paths that publish or
|
|
124
|
+
decide.** A missing `jq` renders as empty DATA and the work carries on with
|
|
125
|
+
it; those nine now exit 3.
|
|
126
|
+
|
|
127
|
+
- **The autopilot runner survives a month unwatched.** `runner.log` is truncated
|
|
128
|
+
in place past 5MB (renaming it leaves launchd's `O_APPEND` descriptor writing
|
|
129
|
+
into the renamed file while the new one stays empty - the rotation that looks
|
|
130
|
+
right and silently stops logging), three consecutive empty attempts stop new
|
|
131
|
+
work being taken, and one JSON line per tick goes to `ticks.jsonl`.
|
|
132
|
+
|
|
133
|
+
- **`github-ssh-setup.sh` no longer waits forever on a headless machine**, and
|
|
134
|
+
`write-state.mjs` gained `--if-rev=<n>` compare-and-swap so a second writer
|
|
135
|
+
cannot silently overwrite the first.
|
|
136
|
+
|
|
137
|
+
- **A direct-run guard that compared `import.meta.url` to `file://${argv[1]}`**
|
|
138
|
+
is false whenever argv[1] is not already resolved - a `/var` path that
|
|
139
|
+
resolves to `/private/var`, or the symlink `install --link` writes. The script
|
|
140
|
+
then does nothing at all, silently.
|
|
141
|
+
|
|
142
|
+
---
|
|
143
|
+
|
|
144
|
+
## [17.6.0] - 2026-09-15
|
|
145
|
+
|
|
146
|
+
### Added
|
|
147
|
+
|
|
148
|
+
- **The maturity check now has an effect (`prefs.global.maturityFollowup`).** It
|
|
149
|
+
has always produced a machine-readable gap list - stable codes in `blockers[]`
|
|
150
|
+
and `warnings[]` - and then thrown most of it away. A blocker halted the run,
|
|
151
|
+
an autopilot queue moved to the next item, and the item stayed exactly as
|
|
152
|
+
immature as it was found. Nobody was told, so nothing changed, so the next
|
|
153
|
+
scan halted on the same item for the same reason. The check was doing its job
|
|
154
|
+
and producing no effect.
|
|
155
|
+
|
|
156
|
+
**Interactive runs ask at the step instead of ending at it**: open the item and
|
|
157
|
+
fix it, continue without it (recording in `state.maturity.accepted[]` *which*
|
|
158
|
+
gap was waved through, which is what separates an informed continue from a
|
|
159
|
+
skipped check), or abort. `askInteractively: false` restores the old halt.
|
|
160
|
+
|
|
161
|
+
**Autopilot can ask on the item itself**, behind
|
|
162
|
+
`autopilotCommentsOnIssue` (**off by default**, because it is an outward-facing
|
|
163
|
+
write). One comment naming what is missing, then a halt on the circuit breaker
|
|
164
|
+
with `state.waitingFor = "maturity"`. A question, never a state change: no
|
|
165
|
+
transition, no resolution, no assignee, no label, no close. `Ref:` never
|
|
166
|
+
`Closes:`. Copy in `outputLanguage`, and the gap wording is the fetcher's own
|
|
167
|
+
`maturity.summary` verbatim rather than a second copy of that table.
|
|
168
|
+
|
|
169
|
+
**"Have we already asked" is read off the ITEM, not off our state file.** An
|
|
170
|
+
autopilot scan is a new run with a fresh `agent-state.json`, so a state-only
|
|
171
|
+
record would make every scan a first ask - the wall of identical bot comments
|
|
172
|
+
this feature exists to prevent. The comment therefore carries its own gap set
|
|
173
|
+
on a last line (`multi-agent gaps: code,code`) and the next pass takes the
|
|
174
|
+
newest comment of ours. Neither the marker nor that line uses square brackets:
|
|
175
|
+
`[text]` is a link in Jira wiki markup, and Jira is where this comment is most
|
|
176
|
+
likely to land. The contract is carried on all three hosts - the Claude command
|
|
177
|
+
tree, and the `shared/core` skills that Copilot dispatches and Codex reads.
|
|
178
|
+
|
|
179
|
+
**The rule that shapes the rest: an edit is a reason to look again, never proof
|
|
180
|
+
the gap closed.** A reply reading "will do later" moves the timestamp and fixes
|
|
181
|
+
nothing, so a changed item is re-fetched and re-scored and the *check* decides.
|
|
182
|
+
Only a genuinely DIFFERENT gap set earns a second comment - otherwise an
|
|
183
|
+
unattended queue turns an item into a wall of identical bot text. "Cannot tell
|
|
184
|
+
whether it moved" resolves to re-check, never to wait, because folding unknown
|
|
185
|
+
into "nothing changed" parks a run forever on a tracker that omits the field.
|
|
186
|
+
|
|
187
|
+
Warnings still auto-continue under autopilot. Converting them to halts in a
|
|
188
|
+
release would stall queues overnight on items that ran fine yesterday;
|
|
189
|
+
`commentOnWarnings` raises them opt-in, and the gaps are recorded either way.
|
|
190
|
+
|
|
191
|
+
> Four of the additions below started as ideas in a public Claude Code
|
|
192
|
+
> configuration repo (`worldflowai/everything-claude-code`). No code was taken:
|
|
193
|
+
> that copy ships no LICENSE file and the GitHub API reports none, and the same
|
|
194
|
+
> call was already made twice here (ADR-0010 on `graphify`, and swiftlens). Each
|
|
195
|
+
> idea was re-derived against this pipeline's own constraints - zero runtime
|
|
196
|
+
> dependencies, macOS only, stdout is the MCP server's JSON-RPC channel - and
|
|
197
|
+
> most of what that repo does was already covered here, in a gated form.
|
|
198
|
+
|
|
199
|
+
- **The node stacks no longer type `npm` (`pipeline/scripts/package-manager.mjs`).**
|
|
200
|
+
Phase 3's web test arm and its build step were hardcoded, so a repo on pnpm,
|
|
201
|
+
yarn or bun failed in Phase 3 - with a worktree and a branch already created -
|
|
202
|
+
or, worse, npm resolved against a lock file it does not own and the run
|
|
203
|
+
continued on a tree the repo's own tooling would never have produced. The
|
|
204
|
+
manager is now resolved from the repo: `$MA_PACKAGE_MANAGER`, then
|
|
205
|
+
`package.json#packageManager`, then a lock file, then npm reported AS a
|
|
206
|
+
default rather than as evidence. Node core only, per ADR-0004. Two lock files
|
|
207
|
+
means a migration left one behind: the newest wins and both are named. Exit 3
|
|
208
|
+
means the repo declares no such script, which is the `--if-present` case
|
|
209
|
+
answered by an exit code instead of a flag whose support differs per manager.
|
|
210
|
+
The resolved name goes through the phase's `eval`, so it is held to the shape
|
|
211
|
+
a binary actually has and the gate proves it by eval'ing the produced line
|
|
212
|
+
with every manager stubbed out. iOS and Android are untouched.
|
|
213
|
+
`refs/features/package-manager.md`.
|
|
214
|
+
|
|
215
|
+
- **A `PreCompact` hook, so a compaction does not eat what a phase learned.**
|
|
216
|
+
`SessionEnd` capture exists because every durable write used to live in Phase
|
|
217
|
+
7, the phase a run is least likely to reach. An auto-compaction is that same
|
|
218
|
+
failure one level down: it summarizes the conversation while a long Phase 3 or
|
|
219
|
+
Phase 4 is still running, and anything not yet flushed is gone before
|
|
220
|
+
`SessionEnd` ever fires. The template now wires `capture-flush.sh` there,
|
|
221
|
+
deliberately WITHOUT `--if-stale` - that test exists so a session exit does not
|
|
222
|
+
re-flush a finished run, and at a compaction what matters is whether anything
|
|
223
|
+
is unflushed, not whether the run finished. Both store writes are idempotent.
|
|
224
|
+
`doctor`'s `hook-coverage` check reports it as missing until the block is
|
|
225
|
+
merged, at no extra cost, because that check diffs the template generically.
|
|
226
|
+
|
|
227
|
+
- **`doctor` now counts the MCP servers this host has registered (`mcp-surface`).**
|
|
228
|
+
`mcp-registration` answers "is ours registered". Nobody was answering the
|
|
229
|
+
question the user never gets asked: every registered server sends its tool
|
|
230
|
+
list on every turn, they are added one at a time, and our own toolkit is 99
|
|
231
|
+
tools by itself. The check only ever REPORTS - INFO above
|
|
232
|
+
`prefs.global.mcpSurface.infoAbove` (default 8), `--explain` lists the names,
|
|
233
|
+
and it never blocks, never warns and never disables anything, because how many
|
|
234
|
+
servers are worth their context is the user's call and not a health failure.
|
|
235
|
+
The threshold is judgement, which is why it is a pref instead of a constant
|
|
236
|
+
nobody can see.
|
|
237
|
+
|
|
238
|
+
- **`GRAPH_REPORT.md` gained "Symbols nothing else references".** The report
|
|
239
|
+
found unconnected FILES; a file imported for one symbol while three of its
|
|
240
|
+
other exports were dead has edges, so those exports were invisible. The join
|
|
241
|
+
is a query over data the graph already held. Candidates, never verdicts, and
|
|
242
|
+
it gates nothing: ADR-0010 records that the extractor is regex over
|
|
243
|
+
comment-stripped source, not a parser, so dynamic dispatch, reflection,
|
|
244
|
+
string-keyed lookup and a public API consumed outside the repo all look like
|
|
245
|
+
dead code from here. Four classes are excluded and COUNTED rather than listed,
|
|
246
|
+
because they could not carry a reference edge however heavily used they are: a
|
|
247
|
+
kind outside the stack's `referenceKinds`, a name declared in more than one
|
|
248
|
+
place, a nested declaration, and anything in a test file. Symbols referenced
|
|
249
|
+
only from tests are listed separately - not dead code, but code whose only
|
|
250
|
+
consumer is its own test.
|
|
251
|
+
|
|
252
|
+
### Fixed
|
|
253
|
+
|
|
254
|
+
- **`/multi-agent:resume` started from `currentPhase + 1` unconditionally**, so a
|
|
255
|
+
run that stopped mid-phase to ask a human resumed past the question. It reads
|
|
256
|
+
`state.waitingFor` first now. This was already broken for Phase 7's channels
|
|
257
|
+
pause, which documented itself as resumable through that field while
|
|
258
|
+
`resume/SKILL.md` never mentioned it - one fix, two callers.
|
|
259
|
+
|
|
260
|
+
- **Only one machine in the world was reporting usage, and nothing was red.**
|
|
261
|
+
Operational reporting needs `usageLog.enabled` AND a token that resolves, and
|
|
262
|
+
both were arranged in exactly one place: `/multi-agent:update`, as forty lines
|
|
263
|
+
of shell embedded in the skill. A user who installed the package, ran
|
|
264
|
+
`/multi-agent:setup` (which said registration happens in update) and worked for
|
|
265
|
+
weeks never registered, never reported, and an empty panel reads exactly like
|
|
266
|
+
nobody using the pipeline. Copilot and Codex were worse off still: their own
|
|
267
|
+
`setup` and `update` skills never mentioned registration at all.
|
|
268
|
+
|
|
269
|
+
It is now one script (`pipeline/scripts/usage-register.mjs`) called from the
|
|
270
|
+
five surfaces where a machine can first become real - setup and update on
|
|
271
|
+
Claude Code, the same two on the cross-CLI surface, and the Phase 0 exit gate
|
|
272
|
+
as the backstop for a machine that reached neither. Same fence as before, now
|
|
273
|
+
enforced in one place: the token is REQUESTED and write-only, it lands in the
|
|
274
|
+
credential store and never in a file, `usageLog.optOut: true` blocks everything
|
|
275
|
+
permanently and is checked before the network call, and an unreachable endpoint
|
|
276
|
+
leaves reporting off with one line and exit 0 - a caller is never failed over
|
|
277
|
+
bookkeeping. `--json` distinguishes "you opted out" from "we could not reach
|
|
278
|
+
it", because those are different facts about the same empty panel. A machine
|
|
279
|
+
that has a token but `enabled: false` (a run interrupted between the two
|
|
280
|
+
writes) is repaired rather than left silent. `refs/features/usage-reporting.md`,
|
|
281
|
+
`smoke-usage-register.sh`.
|
|
282
|
+
|
|
283
|
+
- **`state.maturity` was written by every issue-shaped run and forbidden by the
|
|
284
|
+
schema.** `agent-state.schema.json` has `additionalProperties: false`, and
|
|
285
|
+
`maturity` was not declared - it survived only on the grandfather list of
|
|
286
|
+
`smoke-state-keys-declared.sh`. Declared now, along with `maturityFollowup`
|
|
287
|
+
and `waitingFor`, whose two values (`maturity`, `user-channels-choice`) are the
|
|
288
|
+
two places a run pauses INSIDE a phase rather than between two. The grandfather
|
|
289
|
+
list is two entries shorter, and it only ever shrinks.
|
|
290
|
+
|
|
291
|
+
---
|
|
292
|
+
|
|
17
293
|
## [17.5.1] - 2026-09-15
|
|
18
294
|
|
|
19
295
|
### Documentation
|
package/README.md
CHANGED
|
@@ -17,7 +17,7 @@ Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtim
|
|
|
17
17
|
### Prerequisites
|
|
18
18
|
|
|
19
19
|
- **Node.js >= 20.11** - required; the pipeline's own tooling runs on it.
|
|
20
|
-
- **`jq`** - optional
|
|
20
|
+
- **`jq`** - required for nine paths, optional for the rest. 78 shell files call it. The nine that publish or decide - the autopilot queue, Jira comments, PR reviews, issue updates, the plan file, both Figma fetchers, log search and Jira auth - now refuse with exit 3 rather than run, because a missing `jq` renders as empty DATA and the work carries on with it. Everywhere else it still degrades. The install prints a note when it is missing.
|
|
21
21
|
- **`gh`** - for GitHub issue and PR work. Its built-in `--jq` is independent of the `jq` binary.
|
|
22
22
|
|
|
23
23
|
## Quick Start
|
|
@@ -33,6 +33,41 @@ npx @mmerterden/multi-agent-pipeline install --all # all three
|
|
|
33
33
|
/multi-agent:setup # keychain token scan + git identity + default stack
|
|
34
34
|
```
|
|
35
35
|
|
|
36
|
+
### No `npx` on that machine?
|
|
37
|
+
|
|
38
|
+
`npx` ships with npm, and npm ships with Node - so "npx: command not found" almost
|
|
39
|
+
always means Node is missing from that shell, not that anything is wrong with the
|
|
40
|
+
package. Check first, then pick the row that matches:
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
node -v; npm -v; command -v node npm npx
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
| What you see | What to do |
|
|
47
|
+
|---|---|
|
|
48
|
+
| nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
|
|
49
|
+
| `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
|
|
50
|
+
| nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
|
|
51
|
+
| npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
|
|
52
|
+
|
|
53
|
+
And the path that needs neither `npx` nor a global install - clone and run the
|
|
54
|
+
installer directly:
|
|
55
|
+
|
|
56
|
+
```bash
|
|
57
|
+
git clone https://github.com/mmerterden/multi-agent-pipeline.git
|
|
58
|
+
cd multi-agent-pipeline
|
|
59
|
+
node index.js install --claude # add --dry-run first to see what it would write
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
`npm exec @mmerterden/multi-agent-pipeline install --claude` also works on any npm
|
|
63
|
+
7+ without `npx` on `PATH`.
|
|
64
|
+
|
|
65
|
+
**One error that is not an npx problem.** This package declares `os: ["darwin"]`,
|
|
66
|
+
so npm refuses to install it anywhere else and says `npm ERR! notsup Unsupported
|
|
67
|
+
platform`. That is deliberate ([ADR-0012](./docs/adr/0012-macos-only.md)), not a
|
|
68
|
+
missing tool: every credential read shells `security`, every iOS build
|
|
69
|
+
`xcodebuild`, every piece of visual evidence `simctl`.
|
|
70
|
+
|
|
36
71
|
Tool flags combine (`--claude --codex`). With no tool flag at all, the installer targets Claude Code only. Other flags: `--dry-run` (show what would be written, write nothing), `--platform=ios|android|all` (skip the stack skills you do not need), `--link` (symlink instead of copy, for local development).
|
|
37
72
|
|
|
38
73
|
Run a task - the input type is auto-detected:
|
|
@@ -72,6 +107,29 @@ checkout):
|
|
|
72
107
|
|
|
73
108
|
`/multi-agent:analysis` runs its own shorter chain and, since v16.12.0, reviews what it wrote before publishing it: the draft goes through the same three-reviewer set and triage as a code diff, a blocking finding returns it to synthesis with dispatch closed, and the gaps that survive are either searched, asked about, or recorded with an owner. It used to publish behind a structural validator alone.
|
|
74
109
|
|
|
110
|
+
### 18.0.0: one state directory, and a way to ask whether your install is real
|
|
111
|
+
|
|
112
|
+
- **`multi-agent-pipeline verify`.** The install is a copy, and from the moment it is written the two halves drift independently: an edit in the installed tree is behaviour with no source, and a file the installer skipped is a script the docs describe and nobody has. `verify` compares both against a manifest built at pack time. What a green result proves is stated plainly - the bytes match what the publisher recorded, not who published them.
|
|
113
|
+
- **Cost, past the single run.** The per-task ceiling cannot see the two ways a budget actually empties: a drift that trips nothing, and one session that burns a week in an hour while every run stays under its cap. `cost-analyze` projects, finds days out of family by median absolute deviation, and reports acceleration - as LIST-price estimates, which it says on every run rather than in a footnote.
|
|
114
|
+
- **A run's state lives in one place.** `{project}/{taskId}/` is canonical, the flat layout is still read, and a run that exists in both is counted once.
|
|
115
|
+
- **Server readiness, entirely opt-in.** `doctor --profile=server` adds four checks that only matter when nobody is at the keyboard, `install --unattended` writes a permission profile after printing it, and the autopilot runner now survives a month unwatched. The default install writes no permissions and the default doctor run is unchanged, because a laptop told it fails a server check learns to ignore doctor.
|
|
116
|
+
|
|
117
|
+
### Your package manager, your hooks, your MCP surface
|
|
118
|
+
|
|
119
|
+
Three smaller things in 17.6.0, each closing a gap where the pipeline assumed instead of looking:
|
|
120
|
+
|
|
121
|
+
- **Phase 3 stopped typing `npm`.** A repo on pnpm, yarn or bun used to fail in the development phase, with a worktree and a branch already created. The manager is resolved from the repo now - an env override, then `package.json#packageManager`, then the lock file, then npm reported as a default rather than as evidence. iOS and Android are untouched.
|
|
122
|
+
- **A compaction no longer eats what a phase learned.** The capture hook ran at session end; an auto-compaction summarizes a long review or development phase while it is still running, and anything not yet written was gone before session end ever fired. The hooks template now flushes at `PreCompact` too.
|
|
123
|
+
- **`doctor` counts your MCP servers.** Every registered server sends its tool list on every turn and they are added one at a time, so nobody ever sees the total. It reports the count and nothing else: no warning, no blocking, no disabling.
|
|
124
|
+
|
|
125
|
+
### Maturity: the check now has an effect
|
|
126
|
+
|
|
127
|
+
Phase 0 scores how ready the item is before anything is built, and until 17.6.0 it scored and stopped: a blocker halted the run, an autopilot queue moved on, and the item stayed exactly as immature as it was found. Nobody was told, so nothing changed, so the next scan halted on the same item for the same reason.
|
|
128
|
+
|
|
129
|
+
An interactive run now **asks at that step** - open the item and fix it, continue without it, or abort - and records which gap you waved through rather than just that you continued. An autopilot run can be told to **ask on the item itself** (`prefs.global.maturityFollowup.autopilotCommentsOnIssue`, off by default): one comment naming what is missing, then it stops and waits. Never a status change, never an assignee, never a close, `Ref:` and never `Closes:`.
|
|
130
|
+
|
|
131
|
+
Edit the item and the next scan re-checks it; if the gap closed, development starts on its own. An edit is only a reason to look again - a reply reading "will do later" moves the timestamp and fixes nothing, so the check re-runs against the new content and decides. The same question is never asked twice.
|
|
132
|
+
|
|
75
133
|
### Base branch: evidence, then a question
|
|
76
134
|
|
|
77
135
|
Phase 0 Step 3 collects the candidates **with the evidence behind each one** before it asks. An issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version - so that becomes a ranked row whose description says why, next to rows that say "the repository's default branch" or "you used this last time". You still choose; the evidence only reorders.
|
package/README.tr.md
CHANGED
|
@@ -33,6 +33,40 @@ npx @mmerterden/multi-agent-pipeline install --all # üçü birden
|
|
|
33
33
|
/multi-agent:setup # keychain token taraması + git kimliği + varsayılan stack
|
|
34
34
|
```
|
|
35
35
|
|
|
36
|
+
### O makinede `npx` yoksa
|
|
37
|
+
|
|
38
|
+
`npx` npm ile, npm de Node ile gelir - yani "npx: command not found" neredeyse her
|
|
39
|
+
zaman o kabukta Node olmadığı anlamına gelir, pakette bir sorun olduğu değil. Önce
|
|
40
|
+
bak, sonra sana uyan satırı uygula:
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
node -v; npm -v; command -v node npm npx
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
| Gördüğün | Yapılacak |
|
|
47
|
+
|---|---|
|
|
48
|
+
| hiçbiri yok | Node >= 20.11 kur: `brew install node` ya da nodejs.org'dan LTS installer |
|
|
49
|
+
| `node` çalışıyor, `npx` çalışmıyor | `npm i -g @mmerterden/multi-agent-pipeline` sonra `multi-agent-pipeline install --claude` |
|
|
50
|
+
| nvm kurulu ama kabuk görmüyor | `source ~/.nvm/nvm.sh && nvm use --lts`, ya da yeni bir terminal aç |
|
|
51
|
+
| npm çok eski (< 5.2, npx'ten önceki sürümler) | `npm i -g npm@latest`, ya da üstteki global kurulum satırı |
|
|
52
|
+
|
|
53
|
+
Ne `npx` ne de global kurulum isteyen yol - klonla ve installer'ı doğrudan çalıştır:
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
git clone https://github.com/mmerterden/multi-agent-pipeline.git
|
|
57
|
+
cd multi-agent-pipeline
|
|
58
|
+
node index.js install --claude # önce --dry-run ile ne yazacağını görebilirsin
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
`npm exec @mmerterden/multi-agent-pipeline install --claude` de, `PATH`'te `npx`
|
|
62
|
+
olmayan her npm 7+ üzerinde çalışır.
|
|
63
|
+
|
|
64
|
+
**npx sorunu olmayan bir hata.** Bu paket `os: ["darwin"]` beyan eder; npm başka
|
|
65
|
+
hiçbir yerde kurmaz ve `npm ERR! notsup Unsupported platform` der. Bu kasıtlıdır
|
|
66
|
+
([ADR-0012](./docs/adr/0012-macos-only.md)), eksik bir araç değil: her credential
|
|
67
|
+
okuması `security`, her iOS build'i `xcodebuild`, her görsel kanıt `simctl`
|
|
68
|
+
çağırıyor.
|
|
69
|
+
|
|
36
70
|
Tool flag'leri birleştirilebilir (`--claude --codex`). Hiç tool flag'i verilmezse installer sadece Claude Code'u hedefler. Diğer flag'ler: `--dry-run` (ne yazılacağını gösterir, hiçbir şey yazmaz), `--platform=ios|android|all` (ihtiyacın olmayan stack skill'lerini atlar), `--link` (kopyalamak yerine symlink, lokal geliştirme için).
|
|
37
71
|
|
|
38
72
|
Bir görev çalıştır - girdi tipi otomatik algılanır:
|
|
@@ -72,6 +106,29 @@ mu):
|
|
|
72
106
|
|
|
73
107
|
`/multi-agent:analysis` kendi kısa zincirini koşar ve v16.12.0'dan beri yazdığını yayınlamadan önce review ediyor: taslak, bir kod diff'iyle aynı üç-reviewer setinden ve triyajdan geçiyor, bloklayıcı bulgu dokümanı sentez fazına geri gönderip dispatch'i kapatıyor, hayatta kalan boşluklar ya aranıyor ya sana soruluyor ya da sahibiyle birlikte kayda giriyor. Önceden yalnızca yapısal bir validator'ın arkasından yayınlıyordu.
|
|
74
108
|
|
|
109
|
+
### 18.0.0: tek bir durum dizini, ve kurulumun gerçekten o kurulum olup olmadığını sorma yolu
|
|
110
|
+
|
|
111
|
+
- **`multi-agent-pipeline verify`.** Kurulum bir kopyadır ve yazıldığı andan itibaren iki taraf birbirinden bağımsız kayar: kurulu ağaçtaki bir düzenleme kaynağı olmayan bir davranıştır, kurulumun atladığı dosya ise dokümanın anlattığı ama kimsede olmayan bir script. `verify` ikisini de paketleme anında üretilen bir manifest'e karşı karşılaştırır. Yeşil sonucun ne kanıtladığı açıkça yazılıdır: baytlar yayıncının kaydettiğiyle aynıdır, kimin yayınladığı değil.
|
|
112
|
+
- **Maliyet, tek koşunun ötesinde.** Görev başına tavan, bütçeyi gerçekten bitiren iki şeyi göremez: hiçbir koşuyu aşmayan kayma, ve her koşu tavanın altında kalırken bir saatte bir haftayı yakan tek oturum. `cost-analyze` projeksiyon yapar, medyan mutlak sapmayla aileden ayrılan günü bulur ve ivmeyi raporlar - liste fiyatından tahmin olarak, ve bunu dipnotta değil her koşuda söyler.
|
|
113
|
+
- **Bir koşunun durumu tek yerde.** `{project}/{taskId}/` kanonik, düz yerleşim hâlâ okunuyor, ve iki yerde birden duran koşu bir kez sayılıyor.
|
|
114
|
+
- **Sunucu hazırlığı, tamamen opt-in.** `doctor --profile=server` yalnızca klavyede kimse yokken anlamı olan dört kontrol ekler, `install --unattended` izin profilini yazmadan önce gösterir, autopilot runner artık bir ay gözetimsiz ayakta kalır. Varsayılan kurulum hiçbir izin yazmaz ve varsayılan doctor koşusu değişmez: sunucu kontrolünden kaldığı söylenen bir dizüstü, doctor'ı görmezden gelmeyi öğrenir.
|
|
115
|
+
|
|
116
|
+
### Paket yöneticisi, hook'lar ve MCP yüzeyi
|
|
117
|
+
|
|
118
|
+
17.6.0'da üç küçük iş; üçü de pipeline'ın bakmak yerine varsaydığı bir yeri kapatıyor:
|
|
119
|
+
|
|
120
|
+
- **Faz 3 artık `npm` yazmıyor.** pnpm, yarn ya da bun kullanan bir repo geliştirme fazında, worktree ve branch zaten açılmışken patlıyordu. Yönetici artık reponun kendisinden çözülüyor: önce ortam değişkeni, sonra `package.json#packageManager`, sonra lock dosyası, sonra kanıt değil varsayılan olarak bildirilen npm. iOS ve Android'e dokunulmadı.
|
|
121
|
+
- **Compaction bir fazın öğrendiğini artık yemiyor.** Yakalama hook'u oturum sonunda koşuyordu; otomatik compaction ise uzun bir review ya da geliştirme fazını tam koşarken özetliyor ve henüz yazılmamış olan her şey oturum sonu hiç gelmeden kayboluyordu. Hook şablonu artık `PreCompact`'te de yazıyor.
|
|
122
|
+
- **`doctor` MCP sunucularını sayıyor.** Kayıtlı her sunucu her turda araç listesini gönderiyor ve sunucular teker teker ekleniyor, yani toplamı kimse görmüyor. Sadece sayıyı bildiriyor: uyarı yok, engelleme yok, kapatma yok.
|
|
123
|
+
|
|
124
|
+
### Olgunluk: kontrolün artık bir sonucu var
|
|
125
|
+
|
|
126
|
+
Faz 0 hiçbir şey inşa edilmeden önce maddenin ne kadar hazır olduğunu puanlıyor, ve 17.6.0'a kadar puanlayıp duruyordu: bir blocker koşuyu durduruyor, autopilot sırası bir sonraki maddeye geçiyor, madde bulunduğu kadar olgunlaşmamış kalıyordu. Kimseye söylenmediği için hiçbir şey değişmiyor, bir sonraki tarama aynı maddede aynı sebeple duruyordu.
|
|
127
|
+
|
|
128
|
+
İnteraktif koşu artık **o adımda soruyor** - maddeyi açıp düzelt, onsuz devam et, ya da iptal - ve yalnızca devam ettiğini değil, hangi eksiği görmezden geldiğini kaydediyor. Autopilot koşusuna ise **maddenin kendisine sorması** söylenebiliyor (`prefs.global.maturityFollowup.autopilotCommentsOnIssue`, varsayılan kapalı): eksiği adıyla söyleyen tek bir yorum, sonra durup bekliyor. Durum değişmiyor, atanan kişi değişmiyor, kapatma yok; `Ref:` var, `Closes:` yok.
|
|
129
|
+
|
|
130
|
+
Maddeyi düzenlediğinde bir sonraki tarama yeniden kontrol ediyor; eksik kapandıysa geliştirme kendiliğinden başlıyor. Düzenleme yalnızca yeniden bakma sebebi - "sonra yaparım" diyen bir yorum da damgayı ilerletir ve hiçbir şeyi düzeltmez, o yüzden kontrol yeni içerik üzerinde yeniden koşup karar veriyor. Aynı soru iki kez sorulmuyor.
|
|
131
|
+
|
|
75
132
|
### Temel branch: önce kanıt, sonra soru
|
|
76
133
|
|
|
77
134
|
Faz 0 Adım 3 sormadan önce adayları **her birinin arkasındaki kanıtla** topluyor. Hedef sürüm alanı taşıyan ya da release'i temsil eden ayrı bir maddeye bağlı bir issue, release dallarını sürümle adlandıran bir repoda branch'i zaten söylüyor - bu, nedenini açıklayan sıralı bir satır oluyor; yanında "reponun varsayılan dalı" ya da "geçen sefer bunu kullandın" diyen satırlar duruyor. Seçim yine senin; kanıt yalnızca sıralamayı değiştiriyor.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# 11. CI stays in the repository but dormant; the pre-push gate is primary
|
|
2
2
|
|
|
3
|
-
**Status:** Accepted · 2026-08-31 · amended 2026-09-13
|
|
3
|
+
**Status:** Accepted · 2026-08-31 · amended 2026-09-13 · amended 2026-09-17
|
|
4
4
|
|
|
5
5
|
> **Amendment, 2026-09-13 (ADR-0012, macOS only).** Linux and Windows support
|
|
6
6
|
> was removed. `package.json` now declares `os: ["darwin"]`, which npm enforces
|
|
@@ -91,6 +91,30 @@ gate already covers everything except the three items listed above.
|
|
|
91
91
|
only macOS, disappears when the laptop is closed, and a self-hosted runner on a
|
|
92
92
|
public repository executes code from any fork's pull request.
|
|
93
93
|
|
|
94
|
+
> **Amendment, 2026-09-17.** Two of those three reasons have moved.
|
|
95
|
+
>
|
|
96
|
+
> "Covers only macOS" stopped being a cost when ADR-0012 removed the other
|
|
97
|
+
> platforms: macOS is now the only target, so a macOS-only runner covers all of
|
|
98
|
+
> it.
|
|
99
|
+
>
|
|
100
|
+
> The fork argument is weaker than it was but is NOT gone, and the difference is
|
|
101
|
+
> worth stating precisely because the plan that revisited this got it wrong.
|
|
102
|
+
> The repository is private, which is the important half. It is not
|
|
103
|
+
> fork-free: measured on 2026-09-17 it has **one fork**, owned by another
|
|
104
|
+
> account. GitHub requires approval before a fork's pull request runs a
|
|
105
|
+
> workflow, so the exposure is gated - but the gate is a setting and an
|
|
106
|
+
> approval given by reflex puts that code on the runner, with the runner's
|
|
107
|
+
> keychain.
|
|
108
|
+
>
|
|
109
|
+
> "Disappears when the laptop is closed" is the reason this is still not the
|
|
110
|
+
> default. It becomes a real option only on a machine that stays awake, which
|
|
111
|
+
> is what `docs/server-readiness.md` describes and why that page is a
|
|
112
|
+
> prerequisite rather than a companion.
|
|
113
|
+
>
|
|
114
|
+
> The decision is unchanged: CI stays dormant and the pre-push gate stays
|
|
115
|
+
> primary. What changed is that a self-hosted runner is now a supported option
|
|
116
|
+
> for someone who has read that page, rather than a rejected one.
|
|
117
|
+
|
|
94
118
|
**Delete the workflows.** Honest but lossy. The files encode which steps matter
|
|
95
119
|
and in what order, and the possible issue-triggered direction would have to
|
|
96
120
|
rebuild them from nothing.
|
package/docs/features.md
CHANGED
|
@@ -48,6 +48,8 @@ A deterministic, LLM-free map of what a repo declares and what refers to what, e
|
|
|
48
48
|
|
|
49
49
|
Phase 1 queries it to hand Explore a ranked starting file set instead of a full scan, and Phase 7 rebuilds it after the branch changed code - a rebuild is seconds, so staleness is a `baseCommit` comparison rather than a date heuristic. Off by default behind `prefs.global.codeGraph.enabled`; with it off the pipeline behaves exactly as before.
|
|
50
50
|
|
|
51
|
+
`GRAPH_REPORT.md` also ends with **Symbols nothing else references**: symbols no other file in the repo names, split from the ones referenced only by their own tests. Candidates, never verdicts - the extractor is regex, not a parser, so the four classes that could not carry a reference edge either way (a kind outside the stack's `referenceKinds`, a name declared twice, a nested declaration, a test file) are counted and excluded rather than listed, and nothing gates on the result.
|
|
52
|
+
|
|
51
53
|
Measured on a 4,300-file Swift app against a grep-and-read baseline at the same 30,000-token retrieval budget: 80.4% key-fact coverage at 18,465 tokens against 66.0% at 24,555. The gain is entirely in searches phrased in domain words (63.3% against 32.0%, at under half the cost). When the task already names an exact type, `grep -lw` is still slightly better and slightly cheaper, and the command says so rather than overselling. Reasoning, trade and limits: `docs/adr/0010-own-code-graph.md`.
|
|
52
54
|
|
|
53
55
|
### Stack Auto-Detection
|
|
@@ -76,6 +78,26 @@ Stack skill sets ship as versioned plugins in the `multi-agent-plugins` marketpl
|
|
|
76
78
|
/multi-agent:stack all # every stack plugin
|
|
77
79
|
```
|
|
78
80
|
|
|
81
|
+
### Package Manager Resolution (Phase 3, node-shaped stacks)
|
|
82
|
+
|
|
83
|
+
Phase 3's web test arm and its build step used to type `npm`. A repo on pnpm, yarn or bun then failed in Phase 3 - with a worktree and a branch already created - or, worse, npm resolved against a lock file it does not own and the run continued on a tree the repo's own tooling would never have produced.
|
|
84
|
+
|
|
85
|
+
`scripts/package-manager.mjs` resolves it from the repo instead: `$MA_PACKAGE_MANAGER`, then `package.json#packageManager`, then a lock file, then npm - reported AS a default, never as evidence, because "npm because nothing said otherwise" and "npm because the repo committed a package-lock" are different answers. Node core only (ADR-0004): a resolver that shelled out would need a working install of the tool it is identifying. The walk goes up to the directory holding `.git` and stops there, so a monorepo's root lock file is found and a stray one in a home directory is not. Two lock files means a migration left one behind: the newest wins and both are named.
|
|
86
|
+
|
|
87
|
+
Every manager gets the explicit `run` form (a script named `test` or `add` would otherwise lose to the builtin), only npm gets the `--` separator, and `bun run test` never `bun test`. Exit 3 means the repo declares no such script - the `--if-present` case, answered by an exit code rather than a flag whose support differs per manager. The resolved name is pasted into the phase's `eval`, so it is held to the shape a binary actually has, and the gate proves that by eval'ing the produced line with every manager stubbed out. iOS and Android are untouched. `refs/features/package-manager.md`.
|
|
88
|
+
|
|
89
|
+
### Maturity Follow-Up (`prefs.global.maturityFollowup`)
|
|
90
|
+
|
|
91
|
+
The maturity check has always produced a machine-readable gap list - stable codes in `blockers[]` and `warnings[]` - and then thrown most of it away. A blocker halted the run, an autopilot queue moved to the next item, and the item stayed exactly as immature as it was found. Nobody was told, so nothing changed, so the next scan halted on the same item for the same reason.
|
|
92
|
+
|
|
93
|
+
**Interactive runs ask at the step rather than ending at it**: open the item and fix it, continue without it, or abort. Continuing records which gap was waved through in `state.maturity.accepted[]` - that is what separates an informed continue from a skipped check. An answer typed into a picker improves this run and leaves the item as immature for the next person, so the step offers to write the supplied content back, as a separately approved write.
|
|
94
|
+
|
|
95
|
+
**Autopilot can ask on the item itself**, behind `autopilotCommentsOnIssue` (off by default, because it is an outward-facing write). One comment naming what is missing, then a halt on the circuit breaker with `state.waitingFor = "maturity"` so `resume` re-enters that step. A question, never a state change: no transition, no resolution, no assignee, no label, no close; `Ref:` never `Closes:`; copy in `outputLanguage`, with the gap wording taken verbatim from the fetcher's own summary rather than re-derived.
|
|
96
|
+
|
|
97
|
+
**An edit is a reason to look again, never proof the gap closed.** A reply reading "will do later" moves the timestamp and fixes nothing, so a changed item is re-fetched and re-scored and the check decides. Only a different gap set earns a second comment; "cannot tell whether it moved" re-checks rather than waiting, because folding unknown into "nothing changed" parks a run forever on a tracker that omits the field.
|
|
98
|
+
|
|
99
|
+
Warnings still auto-continue under autopilot - converting them to halts would stall queues on items that ran fine yesterday. `commentOnWarnings` raises them opt-in.
|
|
100
|
+
|
|
79
101
|
### Base-Branch Evidence (Phase 0 Step 3, `prefs.global.baseBranchEvidence.enabled`)
|
|
80
102
|
|
|
81
103
|
Step 3 used to ask one question with a list it could not vouch for. `git fetch origin` ran, its exit code was discarded, and `git branch -r` printed the remote-tracking cache either way - so on a restricted network a weeks-old local list was presented as the remote's answer, with nothing saying so. And the answer was usually derivable: an issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version.
|
|
@@ -240,6 +262,8 @@ Phase 3 treats the issue-tracker status update as a required step with a post-mu
|
|
|
240
262
|
|
|
241
263
|
- **Pre-Commit Secret Detection** (12 patterns): `PreToolUse` hook scans staged files for API keys/tokens, AWS access keys, private keys, `.env` files, service account JSON. Commit **blocked** if found.
|
|
242
264
|
- **Read-Size Gate** (opt-in, `prefs.global.bulkRead.mode`): a `PreToolUse` hook inspects `Read` and the shell commands that read a file whole. In `observe` it only logs what it would have caught - the baseline you measure before routing anything. In `enforce` a file over `minLines` (default 350) is blocked and delegated to a haiku-rung worker (`bulk-read.sh`), which returns a line-numbered summary so the follow-up is a bounded `Read(offset:limit:)` instead of the whole file; the full text is parked under `.multi-agent/refs/`. The development phase and any file the run has already touched are exempt, because Claude Code's `Edit` requires its own `Read` first.
|
|
265
|
+
- **Capture Hooks** (`SessionEnd`, `PreCompact`, `SessionStart`): every durable write used to live in Phase 7, the phase a run is least likely to reach. `SessionEnd` flushes a run that never got there; `PreCompact` flushes before an auto-compaction summarizes a long phase mid-flight, which is the same loss one level down; `SessionStart` prints at most two lines about an unfinished run. None calls a model, none reads a payload, and all exit 0 on every path - a hook that fails a session over bookkeeping is worse than the bookkeeping.
|
|
266
|
+
- **Operational Reporting** (`prefs.global.usageLog`): coarse run metadata - task id, phase, status, durations, token counts - and never prompts, code, diffs or absolute paths. The per-machine token is REQUESTED from the endpoint by `usage-register.mjs` (setup, update, and the Phase 0 exit gate as a backstop), is write-only, and lives in the OS credential store; prefs hold only the entry name and the switch. `usageLog.optOut: true` blocks registration permanently and is checked before the network call. An unreachable endpoint leaves reporting off with one line and exit 0 - a run is never failed over bookkeeping.
|
|
243
267
|
- **Build Queue**: All `xcodebuild` calls acquire a lock. Each worktree uses own `-derivedDataPath`. Stale locks auto-clean after 15 min. Non-Xcode builds don't need the lock.
|
|
244
268
|
- **Context Management**: `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=65` - compaction at 65% usage (prevents degradation in 8-phase sessions).
|
|
245
269
|
- **3-Iteration Hard Kill**: Any retry loop stops after 3 attempts, then pauses for user. No infinite loops.
|