@mmerterden/multi-agent-pipeline 17.5.1 → 18.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (134) hide show
  1. package/CHANGELOG.md +276 -0
  2. package/README.md +59 -1
  3. package/README.tr.md +57 -0
  4. package/docs/adr/0011-dormant-ci.md +25 -1
  5. package/docs/features.md +24 -0
  6. package/docs/server-readiness.md +188 -0
  7. package/docs/token-budget-history.md +1 -1
  8. package/index.js +16 -1
  9. package/install/_common.mjs +42 -17
  10. package/install/_dev-only-files.mjs +8 -0
  11. package/install/_unattended-profile.mjs +113 -0
  12. package/install/index.mjs +48 -0
  13. package/install/templates/claude-hooks.json +13 -1
  14. package/manifest.json +1049 -0
  15. package/package.json +5 -2
  16. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  17. package/pipeline/commands/multi-agent/feedback/SKILL.md +7 -1
  18. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  19. package/pipeline/commands/multi-agent/issue/SKILL.md +13 -1
  20. package/pipeline/commands/multi-agent/jira/SKILL.md +13 -1
  21. package/pipeline/commands/multi-agent/resume/SKILL.md +16 -1
  22. package/pipeline/commands/multi-agent/setup/SKILL.md +14 -16
  23. package/pipeline/commands/multi-agent/status/SKILL.md +52 -21
  24. package/pipeline/commands/multi-agent/update/SKILL.md +13 -56
  25. package/pipeline/lib/_jira-auth.sh +8 -0
  26. package/pipeline/lib/analysis-jira-write.sh +32 -0
  27. package/pipeline/lib/ask-choice.sh +13 -2
  28. package/pipeline/lib/autopilot-state.sh +8 -0
  29. package/pipeline/lib/fatal.mjs +129 -0
  30. package/pipeline/lib/figma-mcp-refresh.sh +18 -0
  31. package/pipeline/lib/figma-screenshot.sh +18 -0
  32. package/pipeline/lib/invoked-directly.mjs +43 -0
  33. package/pipeline/lib/jira-publish.sh +42 -0
  34. package/pipeline/lib/md2confluence-v3.py +47 -0
  35. package/pipeline/lib/outbound-gate.mjs +175 -0
  36. package/pipeline/lib/plan-todos.sh +27 -6
  37. package/pipeline/lib/post-pr-review.sh +77 -8
  38. package/pipeline/lib/repo-hygiene.sh +8 -3
  39. package/pipeline/lib/require-jq.sh +40 -0
  40. package/pipeline/lib/run-paths.sh +335 -0
  41. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +70 -0
  42. package/pipeline/multi-agent-refs/features/code-graph.md +20 -0
  43. package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
  44. package/pipeline/multi-agent-refs/features/doctor.md +68 -0
  45. package/pipeline/multi-agent-refs/features/maturity-followup.md +166 -0
  46. package/pipeline/multi-agent-refs/features/package-manager.md +80 -0
  47. package/pipeline/multi-agent-refs/features/usage-reporting.md +79 -0
  48. package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
  49. package/pipeline/multi-agent-refs/features/verify.md +83 -0
  50. package/pipeline/multi-agent-refs/phases/operations.md +13 -2
  51. package/pipeline/multi-agent-refs/phases/phase-0-init.md +6 -3
  52. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +8 -2
  53. package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -1
  54. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  55. package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
  56. package/pipeline/preferences-template.json +1 -1
  57. package/pipeline/schemas/agent-state.schema.json +122 -11
  58. package/pipeline/schemas/prefs.schema.json +35 -0
  59. package/pipeline/schemas/token-budget.json +2 -2
  60. package/pipeline/scripts/_run-paths.mjs +372 -0
  61. package/pipeline/scripts/aggregate-metrics.mjs +64 -64
  62. package/pipeline/scripts/autopilot-arming.mjs +2 -1
  63. package/pipeline/scripts/autopilot-intake.mjs +2 -1
  64. package/pipeline/scripts/autopilot-runner.mjs +206 -2
  65. package/pipeline/scripts/build-references.mjs +2 -1
  66. package/pipeline/scripts/build-stack-plugins.mjs +10 -2
  67. package/pipeline/scripts/capture-evidence.sh +7 -2
  68. package/pipeline/scripts/classify-plan-safety.mjs +2 -1
  69. package/pipeline/scripts/cost-analyze.mjs +600 -0
  70. package/pipeline/scripts/cost-budget-check.mjs +4 -12
  71. package/pipeline/scripts/council-view.mjs +2 -1
  72. package/pipeline/scripts/crush-json.mjs +2 -1
  73. package/pipeline/scripts/diff-explain.mjs +6 -9
  74. package/pipeline/scripts/diff-risk-score.mjs +2 -1
  75. package/pipeline/scripts/doctor.mjs +203 -4
  76. package/pipeline/scripts/evidence-gate.mjs +9 -3
  77. package/pipeline/scripts/feedback-send.mjs +13 -3
  78. package/pipeline/scripts/gc-abandoned.sh +29 -13
  79. package/pipeline/scripts/gc-worktrees.sh +11 -4
  80. package/pipeline/scripts/github-ssh-setup.sh +64 -7
  81. package/pipeline/scripts/graph-mermaid.mjs +4 -2
  82. package/pipeline/scripts/graph-report.mjs +155 -1
  83. package/pipeline/scripts/keychain-save.sh +101 -30
  84. package/pipeline/scripts/learn-from-transcripts.mjs +2 -1
  85. package/pipeline/scripts/learning-curve.mjs +34 -29
  86. package/pipeline/scripts/make-manifest.mjs +199 -0
  87. package/pipeline/scripts/maturity-followup.mjs +294 -0
  88. package/pipeline/scripts/migrate-prefs.mjs +2 -1
  89. package/pipeline/scripts/migrate-state.mjs +94 -4
  90. package/pipeline/scripts/package-manager.mjs +310 -0
  91. package/pipeline/scripts/phase-banner.sh +6 -2
  92. package/pipeline/scripts/phase-tracker.sh +41 -3
  93. package/pipeline/scripts/plan-coverage-gate.mjs +6 -2
  94. package/pipeline/scripts/pre-commit-check.sh +7 -0
  95. package/pipeline/scripts/pre-push-check.sh +7 -0
  96. package/pipeline/scripts/purge.sh +23 -6
  97. package/pipeline/scripts/render-agent-log-cost.sh +9 -2
  98. package/pipeline/scripts/render-cost-summary.sh +9 -2
  99. package/pipeline/scripts/render-work-summary.sh +11 -4
  100. package/pipeline/scripts/review-file-filter.mjs +4 -2
  101. package/pipeline/scripts/review-scope.mjs +2 -1
  102. package/pipeline/scripts/routine-registry.mjs +2 -1
  103. package/pipeline/scripts/run-aggregator.mjs +13 -14
  104. package/pipeline/scripts/run-metrics.mjs +3 -1
  105. package/pipeline/scripts/runs-index.mjs +343 -0
  106. package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
  107. package/pipeline/scripts/search-logs.sh +18 -0
  108. package/pipeline/scripts/test-gap-scan.mjs +2 -1
  109. package/pipeline/scripts/test-integrity-gate.mjs +2 -1
  110. package/pipeline/scripts/update-issue-progress.sh +56 -7
  111. package/pipeline/scripts/usage-register.mjs +271 -0
  112. package/pipeline/scripts/usage-report.mjs +14 -3
  113. package/pipeline/scripts/validate-analysis-doc.mjs +2 -1
  114. package/pipeline/scripts/validate-code-graph.mjs +6 -3
  115. package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
  116. package/pipeline/scripts/validate-diff-risk.mjs +6 -3
  117. package/pipeline/scripts/validate-test-gap.mjs +6 -3
  118. package/pipeline/scripts/validate-triage.mjs +3 -1
  119. package/pipeline/scripts/verify-citations.mjs +4 -2
  120. package/pipeline/scripts/verify.mjs +327 -0
  121. package/pipeline/scripts/worktree-finalize.sh +13 -4
  122. package/pipeline/scripts/write-state.mjs +154 -15
  123. package/pipeline/skills/.skill-manifest.json +6 -6
  124. package/pipeline/skills/.skills-index.json +56 -1
  125. package/pipeline/skills/shared/README.md +8 -3
  126. package/pipeline/skills/shared/core/multi-agent-issue/SKILL.md +14 -0
  127. package/pipeline/skills/shared/core/multi-agent-jira/SKILL.md +14 -0
  128. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +13 -0
  129. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +33 -9
  130. package/pipeline/skills/shared/core/multi-agent-update/SKILL.md +6 -0
  131. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
  132. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
  133. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
  134. package/pipeline/skills/skills-index.md +6 -1
@@ -0,0 +1,335 @@
1
+ #!/bin/bash
2
+ # run-paths.sh - the single resolver for pipeline run-state paths (shell side).
3
+ #
4
+ # Shell twin of pipeline/scripts/_run-paths.mjs. The two must agree; the
5
+ # contract is asserted by pipeline/scripts/smoke-run-path-canonical.sh, which
6
+ # runs both over the same fixture tree and diffs their answers.
7
+ #
8
+ # A run's files live under the log root in ONE of two layouts:
9
+ #
10
+ # nested <root>/<project>/<task_id>/ documented canonical
11
+ # flat <root>/<task_id>/ what phase-tracker.sh writes
12
+ #
13
+ # Both are populated on real machines, so every reader accepts both. This file
14
+ # does NOT change where anything is written: measured on a real install, 90 of
15
+ # 103 runs are flat and every tracker-state.json is. A task id present in both
16
+ # layouts is ONE run; the most recently written directory wins. Relocation is
17
+ # opt-in and explicit: `migrate-state.mjs --relocate`.
18
+ #
19
+ # No `set -e` on purpose: this is a sourced library and changing the caller's
20
+ # shell options is not this file's business.
21
+ #
22
+ # Usage:
23
+ # . "$(dirname "$0")/../lib/run-paths.sh"
24
+ # root=$(ma_logs_root)
25
+ # dir=$(ma_resolve_run_dir "{JIRA_KEY}-123") || echo "unknown run"
26
+ # f=$(ma_resolve_run_file "{JIRA_KEY}-123" tracker-state.json)
27
+ # ma_list_runs # task_id<TAB>project<TAB>dir<TAB>layout, one per run
28
+ #
29
+ # Bash 3.2 compatible: no associative arrays, no `mapfile`, no `grep -P`.
30
+
31
+ # Marker files and reserved names are spelled out at each test rather than held
32
+ # in a space-separated variable iterated with `for x in $VAR`. bash word-splits
33
+ # an unquoted variable; zsh does not, so there the loop would run ONCE with the
34
+ # whole string as a single word, every test would fail, and the library would
35
+ # answer "no runs" instead of erroring. This file is sourced, so the caller's
36
+ # shell decides - and answering zero is the failure mode that looks like data.
37
+
38
+ ma_logs_root() {
39
+ printf '%s\n' "${LOGS_ROOT:-$HOME/.claude/logs/multi-agent}"
40
+ }
41
+
42
+ # ma_is_run_dir <dir> -> 0 when the directory holds at least one run marker
43
+ # Phase 6 removes the worktree and salvages the run's files into `artifacts/`
44
+ # inside the same run directory, so a shipped run keeps its state one level
45
+ # deeper. A reader that only looks at the top level reports it as stateless.
46
+ MA_ARTIFACTS_SUBDIR=artifacts
47
+
48
+ ma_is_run_dir() {
49
+ local d="$1"
50
+ [ -f "$d/agent-state.json" ] && return 0
51
+ [ -f "$d/tracker-state.json" ] && return 0
52
+ [ -f "$d/agent-log.md" ] && return 0
53
+ [ -f "$d/$MA_ARTIFACTS_SUBDIR/agent-state.json" ] && return 0
54
+ [ -f "$d/$MA_ARTIFACTS_SUBDIR/tracker-state.json" ] && return 0
55
+ [ -f "$d/$MA_ARTIFACTS_SUBDIR/agent-log.md" ] && return 0
56
+ return 1
57
+ }
58
+
59
+ ma_is_reserved() {
60
+ case "$1" in
61
+ review-watch | jira-backups | shadow-git | _analysis-jira | prompts) return 0 ;;
62
+ esac
63
+ return 1
64
+ }
65
+
66
+ # Newest mtime across a run directory's markers, as epoch seconds. GNU-first
67
+ # `stat -c` ordering is required: `stat -f` is a valid GNU flag
68
+ # (--file-system) that SUCCEEDS, so BSD-first would silently win on Linux.
69
+ ma_run_mtime() {
70
+ local d="$1" m t newest=0 b
71
+ for b in "$d" "$d/$MA_ARTIFACTS_SUBDIR"; do
72
+ for m in agent-state.json tracker-state.json agent-log.md; do
73
+ [ -f "$b/$m" ] || continue
74
+ # -L follows symlinks. Some nested run directories are symlink bridges to
75
+ # the flat copy of the same run (3 such pairs on the install this was
76
+ # measured on); without -L the shell reads the LINK's mtime while the
77
+ # .mjs twin's fs.statSync reads the TARGET's, and the two picked
78
+ # different winners for one record. Both must read the same clock.
79
+ t=$(stat -L -c %Y "$b/$m" 2>/dev/null || stat -L -f %m "$b/$m" 2>/dev/null || echo 0)
80
+ [ -n "$t" ] || t=0
81
+ [ "$t" -gt "$newest" ] && newest="$t"
82
+ done
83
+ done
84
+ printf '%s\n' "$newest"
85
+ }
86
+
87
+ # ma_canonical_run_dir <task_id> [project] -> where a NEW run should be written
88
+ ma_canonical_run_dir() {
89
+ local task_id="$1" project="${2:-}" root
90
+ root=$(ma_logs_root)
91
+ if [ -n "$project" ]; then
92
+ printf '%s\n' "$root/$project/$task_id"
93
+ else
94
+ printf '%s\n' "$root/$task_id"
95
+ fi
96
+ }
97
+
98
+ # ma_run_dir_candidates <task_id> [project] -> candidate dirs, most specific
99
+ # first. Existence is not checked here.
100
+ ma_run_dir_candidates() {
101
+ local task_id="$1" project="${2:-}" root d n
102
+ root=$(ma_logs_root)
103
+ [ -n "$project" ] && printf '%s\n' "$root/$project/$task_id"
104
+ printf '%s\n' "$root/$task_id"
105
+ [ -n "$project" ] && return 0
106
+ # No project given: the run may still be nested under one.
107
+ #
108
+ # `find` rather than a `*/` glob on purpose: an unmatched glob is a literal
109
+ # string in bash and a hard error in zsh, and this file is sourced, so the
110
+ # caller's shell decides. find behaves identically in both.
111
+ while IFS= read -r d; do
112
+ [ -n "$d" ] || continue
113
+ n=$(basename "$d")
114
+ ma_is_reserved "$n" && continue
115
+ [ "$n" = "$task_id" ] && continue
116
+ ma_is_run_dir "$root/$n/$task_id" && printf '%s\n' "$root/$n/$task_id"
117
+ done <<EOF
118
+ $(find "$root" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
119
+ EOF
120
+ return 0
121
+ }
122
+
123
+ # ma_realpath <path> -> the path with every symlink resolved, or the input when
124
+ # it cannot be resolved. BSD realpath first, python3 as the fallback (both are
125
+ # already hard requirements of this tree).
126
+ ma_realpath() {
127
+ realpath "$1" 2>/dev/null ||
128
+ python3 -c 'import os,sys; print(os.path.realpath(sys.argv[1]))' "$1" 2>/dev/null ||
129
+ printf '%s\n' "$1"
130
+ }
131
+
132
+ # ma_is_bridge_dir <dir> -> 0 when any marker in it is a symlink elsewhere.
133
+ # Some nested run directories are symlink bridges into the flat copy of the
134
+ # same run. They are one record seen twice, not two copies.
135
+ ma_is_bridge_dir() {
136
+ local d="$1" m
137
+ for m in agent-state.json tracker-state.json agent-log.md; do
138
+ [ -L "$d/$m" ] && return 0
139
+ done
140
+ return 1
141
+ }
142
+
143
+ # ma_same_record <dir_a> <dir_b> -> 0 when both are views of ONE record.
144
+ ma_same_record() {
145
+ local a="$1" b="$2" m ra rb
146
+ for m in agent-state.json tracker-state.json agent-log.md; do
147
+ [ -e "$a/$m" ] && [ -e "$b/$m" ] || continue
148
+ ra=$(ma_realpath "$a/$m")
149
+ rb=$(ma_realpath "$b/$m")
150
+ [ "$ra" = "$rb" ] && return 0
151
+ done
152
+ return 1
153
+ }
154
+
155
+ # ma_rank_fields <dir> -> "<mtime>\t<has_state>\t<depth>", the three ranking
156
+ # columns used to pick between candidate directories for one task id.
157
+ #
158
+ # ma_sort_key packs the same three into one space-separated key for callers
159
+ # that sort a single column; both must match compareCandidates in
160
+ # _run-paths.mjs exactly.
161
+ #
162
+ # mtime alone is not an order: two directories written in the same second tie,
163
+ # and each implementation then fell back to its own traversal order - which is
164
+ # how the two came to disagree about a run present in both layouts. The tail of
165
+ # the key defines the answer: richer record first (agent-state.json is what
166
+ # every reader wants), then the documented nested layout, then the path.
167
+ ma_rank_fields() {
168
+ local d="$1" t has_state depth real
169
+ t=$(ma_run_mtime "$d")
170
+ has_state=0
171
+ [ -f "$d/agent-state.json" ] && has_state=1
172
+ depth=$(printf '%s' "$d" | awk -F/ '{print NF}')
173
+ real=1
174
+ ma_is_bridge_dir "$d" && real=0
175
+ printf '%s\t%s\t%s\t%s\n' "$real" "$t" "$has_state" "$depth"
176
+ }
177
+
178
+ ma_sort_key() {
179
+ local d="$1" t has_state depth real
180
+ t=$(ma_run_mtime "$d")
181
+ has_state=0
182
+ [ -f "$d/agent-state.json" ] && has_state=1
183
+ depth=$(printf '%s' "$d" | awk -F/ '{print NF}')
184
+ real=1
185
+ ma_is_bridge_dir "$d" && real=0
186
+ printf '%d %012d %d %04d %s\n' "$real" "$t" "$has_state" "$depth" "$d"
187
+ }
188
+
189
+ # ma_resolve_run_dir <task_id> [project] -> the directory, or rc 1 when unknown.
190
+ ma_resolve_run_dir() {
191
+ local task_id="$1" project="${2:-}" c best
192
+ best=$(
193
+ while IFS= read -r c; do
194
+ [ -n "$c" ] || continue
195
+ ma_is_run_dir "$c" || continue
196
+ ma_sort_key "$c"
197
+ done <<EOF
198
+ $(ma_run_dir_candidates "$task_id" "$project")
199
+ EOF
200
+ )
201
+ [ -n "$best" ] || return 1
202
+ # Descending on mtime, then has-state, then depth; ascending on path.
203
+ printf '%s\n' "$best" | sort -k1,1nr -k2,2nr -k3,3nr -k4,4nr -k5,5 | head -1 | cut -d' ' -f5-
204
+ }
205
+
206
+ # ma_task_id_variants <id> -> the spellings one task id has been written under,
207
+ # in preference order. `#316` arrives from a GitHub issue reference, `316` from
208
+ # the bare-number input class, `task-316` from an older directory convention.
209
+ # Callers used to inline this list; dropping a spelling silently stops old runs
210
+ # from resolving.
211
+ ma_task_id_variants() {
212
+ local raw="$1" bare
213
+ bare="${raw#\#}"
214
+ printf '%s\n' "$raw"
215
+ [ "$bare" != "$raw" ] && printf '%s\n' "$bare"
216
+ printf '%s\n' "task-$bare"
217
+ }
218
+
219
+ # ma_resolve_run_dir_any <task_id> [project] -> dir for any spelling, or rc 1
220
+ ma_resolve_run_dir_any() {
221
+ local task_id="$1" project="${2:-}" v dir
222
+ while IFS= read -r v; do
223
+ [ -n "$v" ] || continue
224
+ dir=$(ma_resolve_run_dir "$v" "$project") && {
225
+ printf '%s\n' "$dir"
226
+ return 0
227
+ }
228
+ done <<EOF
229
+ $(ma_task_id_variants "$task_id")
230
+ EOF
231
+ return 1
232
+ }
233
+
234
+ # ma_resolve_run_file <task_id> <filename> [project] -> path, or rc 1.
235
+ # Every spelling of the id is tried.
236
+ ma_resolve_run_file() {
237
+ local task_id="$1" filename="$2" project="${3:-}" v dir
238
+ while IFS= read -r v; do
239
+ [ -n "$v" ] || continue
240
+ dir=$(ma_resolve_run_dir "$v" "$project") || continue
241
+ if [ -f "$dir/$filename" ]; then
242
+ printf '%s\n' "$dir/$filename"
243
+ return 0
244
+ fi
245
+ # The salvaged copy Phase 6 leaves behind.
246
+ if [ -f "$dir/$MA_ARTIFACTS_SUBDIR/$filename" ]; then
247
+ printf '%s\n' "$dir/$MA_ARTIFACTS_SUBDIR/$filename"
248
+ return 0
249
+ fi
250
+ done <<EOF
251
+ $(ma_task_id_variants "$task_id")
252
+ EOF
253
+ return 1
254
+ }
255
+
256
+ # ma_list_runs -> one TAB-separated row per run, deduplicated by task id:
257
+ # task_id<TAB>project<TAB>dir<TAB>layout
258
+ # `project` is "-" when the layout does not name one.
259
+ ma_list_runs() {
260
+ local root d n kd kn rows row id
261
+ root=$(ma_logs_root)
262
+ [ -d "$root" ] || return 0
263
+ rows=""
264
+ while IFS= read -r d; do
265
+ [ -n "$d" ] || continue
266
+ n=$(basename "$d")
267
+ ma_is_reserved "$n" && continue
268
+ if ma_is_run_dir "$root/$n"; then
269
+ rows="$rows$n - $root/$n flat $(ma_rank_fields "$root/$n")
270
+ "
271
+ continue
272
+ fi
273
+ while IFS= read -r kd; do
274
+ [ -n "$kd" ] || continue
275
+ kn=$(basename "$kd")
276
+ ma_is_run_dir "$root/$n/$kn" || continue
277
+ rows="$rows$kn $n $root/$n/$kn nested $(ma_rank_fields "$root/$n/$kn")
278
+ "
279
+ done <<EOF
280
+ $(find "$root/$n" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
281
+ EOF
282
+ done <<EOF
283
+ $(find "$root" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
284
+ EOF
285
+ # Dedup by task id using the same total order as resolveRunDir and as
286
+ # compareCandidates in the .mjs twin: mtime desc, has-state desc, depth desc,
287
+ # path ASC. The sub-keys are separate columns on purpose - a single reverse
288
+ # sort over a composite key also reverses the path, which is the one
289
+ # component that has to ascend.
290
+ printf '%s' "$rows" | grep -v '^$' |
291
+ sort -t' ' -k1,1 -k5,5nr -k6,6nr -k7,7nr -k8,8nr -k3,3 | awk -F' ' '
292
+ !seen[$1]++ { print $1 "\t" $2 "\t" $3 "\t" $4 }
293
+ ' | sort -t' ' -k1,1
294
+ }
295
+
296
+ # ma_duplicate_run_ids -> task ids that exist as two SEPARATE records in the two
297
+ # layouts, one per line. A symlink bridge is excluded: it is one record seen
298
+ # twice, and reporting it overstates the drift.
299
+ ma_duplicate_run_ids() {
300
+ local root d n kd kn pairs id dirs a b
301
+ root=$(ma_logs_root)
302
+ [ -d "$root" ] || return 0
303
+ pairs=$(
304
+ while IFS= read -r d; do
305
+ [ -n "$d" ] || continue
306
+ n=$(basename "$d")
307
+ ma_is_reserved "$n" && continue
308
+ if ma_is_run_dir "$root/$n"; then
309
+ printf '%s\t%s\n' "$n" "$root/$n"
310
+ continue
311
+ fi
312
+ while IFS= read -r kd; do
313
+ [ -n "$kd" ] || continue
314
+ kn=$(basename "$kd")
315
+ ma_is_run_dir "$root/$n/$kn" && printf '%s\t%s\n' "$kn" "$root/$n/$kn"
316
+ done <<EOF
317
+ $(find "$root/$n" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
318
+ EOF
319
+ done <<EOF
320
+ $(find "$root" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
321
+ EOF
322
+ )
323
+ # Ids seen more than once, minus the ones whose two directories resolve to the
324
+ # same underlying files (a symlink bridge is one record, not drift).
325
+ while IFS= read -r id; do
326
+ [ -n "$id" ] || continue
327
+ dirs=$(printf '%s\n' "$pairs" | awk -F'\t' -v k="$id" '$1==k {print $2}')
328
+ a=$(printf '%s\n' "$dirs" | sed -n 1p)
329
+ b=$(printf '%s\n' "$dirs" | sed -n 2p)
330
+ [ -n "$b" ] || continue
331
+ ma_same_record "$a" "$b" || printf '%s\n' "$id"
332
+ done <<EOF
333
+ $(printf '%s\n' "$pairs" | awk -F'\t' '{print $1}' | sort | uniq -d)
334
+ EOF
335
+ }
@@ -1,5 +1,14 @@
1
1
  # Feature: Autopilot Circuit-Breaker
2
2
 
3
+ <!-- toc -->
4
+ - [Trip conditions](#trip-conditions)
5
+ - [Wiring status](#wiring-status)
6
+ - [Action on trip](#action-on-trip)
7
+ - [Why this is the right autopilot exception](#why-this-is-the-right-autopilot-exception)
8
+ - [The runner-level breaker (a different failure, a different layer)](#the-runner-level-breaker-a-different-failure-a-different-layer)
9
+ - [Two more things a long-running runner needs](#two-more-things-a-long-running-runner-needs)
10
+ <!-- /toc -->
11
+
3
12
  **Pattern**: autopilot runs with zero interaction, which is exactly when a silent failure loop is most expensive - an agent can burn a budget re-attempting the same broken fix, or thrash between two phases, with nobody watching. A circuit-breaker converts "keep going no matter what" into "keep going until a defined unsafe condition, then halt and hand back to the user." This is the sanctioned autopilot pause (same class as the Phase 7 channels pause): the run stops, records why, and waits for an explicit `resume`.
4
13
 
5
14
  **Gated by `prefs.global.autopilotCircuitBreaker`** (`enabled` default true, `identicalFindingCycles` default 2, `maxReworkCycles` default 3; `schemas/prefs.schema.json`). Halting is always safe, so the breaker itself defaults on. Disable per-run only with an explicit override. Complements, does not replace, the existing autopilot safety rules (build-fail max 3 retries, Phase 4 blocking-finding rework, destructive-op confirmations).
@@ -43,3 +52,64 @@ State shape: `state.circuitBreaker = {tripped, trigger, detail, checkpoint: {pha
43
52
  Autopilot's contract is "no interaction on the happy path." The circuit-breaker fires only off the happy path, where continuing unattended is the *less* safe choice: repeating a proven-broken action, or spending past a ceiling the user set, is not autonomy, it is a runaway. Halting with a precise reason is cheaper than the tokens (and trust) a silent loop burns.
44
53
 
45
54
  Inspired by autonomous-loop operators that gate on explicit stop conditions (no-progress, identical-stack-trace repetition, cost drift, conflict blocking), adapted to this pipeline's phase/rework/state model.
55
+
56
+ ## The runner-level breaker (a different failure, a different layer)
57
+
58
+ Everything above is the IN-RUN breaker: one run, one state file, triggers that
59
+ mean "this task is not converging". It cannot see the failure a server actually
60
+ hits, which is not about any one task.
61
+
62
+ An expired token, a `claude` binary that no longer launches, a full disk: every
63
+ item fails identically, a few minutes apart, until the queue is empty. Nothing
64
+ above notices, because each individual run failed for its own apparently local
65
+ reason. What is left afterwards is a ledger of failures with no indication of
66
+ which came first and no queue to resume.
67
+
68
+ `autopilot-runner.mjs` therefore refuses to take NEW work after
69
+ `MA_AP_BREAKER_LIMIT` consecutive attempts produced nothing (default 3, zero
70
+ disables). The count is CONSECUTIVE and read from `attempted.jsonl` rather than
71
+ kept as a counter, because a counter is a second piece of state that can
72
+ disagree with the ledger - and the ledger is what a person reads when they ask
73
+ why the runner stopped.
74
+
75
+ Two outcomes are deliberately not failures:
76
+
77
+ - `needs-input` is a run that did its work and is waiting for a person. That is
78
+ the system behaving correctly, and counting it would stop a machine whose only
79
+ problem is that somebody has not answered yet.
80
+ - `blocked-*` items never ran at all - arming refused them - so they say nothing
81
+ about whether a run would have worked.
82
+
83
+ When it opens, the reason goes into `queue.json.blockedReason` (so
84
+ `autopilot-status` shows it rather than reporting an idle queue), a
85
+ `breaker-open` line goes into `ticks.jsonl`, and the log names the outcome that
86
+ opened it plus the override.
87
+
88
+ It is HALF-OPEN rather than latched, and that distinction is load-bearing. The
89
+ breaker returns before the only writer of `attempted.jsonl`, so a version that
90
+ simply stopped could never record another attempt, the consecutive count could
91
+ never fall, and the runner would be off for good - while its own log said
92
+ "until an attempt succeeds". After `MA_AP_BREAKER_COOLDOWN_SEC` (30 min by
93
+ default) one probe is let through: a recovered machine succeeds and the count
94
+ resets on its own, and a still-broken one fails, becomes the newest attempt, and
95
+ closes the breaker for another cooldown. One wasted run per cooldown is the
96
+ price of not needing a person to notice.
97
+
98
+ ## Two more things a long-running runner needs
99
+
100
+ **The log.** launchd appends the runner's stdout to `runner.log` forever. On a
101
+ laptop that file is never read; on a machine ticking every few minutes for
102
+ months it is what fills the disk, and a full disk stops the runs the log existed
103
+ to record. Each tick truncates it past `MA_AP_LOG_MAX_BYTES` (default 5MB),
104
+ keeping the last 256KB in `runner.log.1`.
105
+
106
+ Truncating the same inode is the mechanism, not renaming: launchd holds the file
107
+ open with `O_APPEND`, so a rename leaves that descriptor writing into the renamed
108
+ file while the new one stays empty forever. That is the rotation that looks
109
+ correct and silently stops logging.
110
+
111
+ **Telemetry.** `ticks.jsonl` gets one JSON line per tick - what the tick did,
112
+ what it took, what came out, how many consecutive failures preceded it. The
113
+ human log answers "what happened just now"; this answers "how has this been
114
+ behaving for a week", which is the question a server raises and prose cannot be
115
+ asked.
@@ -3,6 +3,7 @@
3
3
  <!-- toc -->
4
4
  - [Phase 1 Step 2.6 - query before dispatching Explore](#phase-1-step-26---query-before-dispatching-explore)
5
5
  - [Phase 7 Step 3 - refresh after the branch changed code](#phase-7-step-3---refresh-after-the-branch-changed-code)
6
+ - [What the report answers that a file-level view cannot](#what-the-report-answers-that-a-file-level-view-cannot)
6
7
  - [The graph is drawable, and one place already asks for it](#the-graph-is-drawable-and-one-place-already-asks-for-it)
7
8
  <!-- /toc -->
8
9
 
@@ -74,6 +75,25 @@ heuristic. A non-zero validator exit keeps the previous graph and logs
74
75
  `knowledge.graph_invalid`; it never fails the run - a stale graph is a degraded
75
76
  Phase 1, not a broken deliverable.
76
77
 
78
+ ### What the report answers that a file-level view cannot
79
+
80
+ `GRAPH_REPORT.md` ends with **Symbols nothing else references**: symbols no other
81
+ file in the repo names. The older "unconnected files" section only ever found
82
+ files with NO edge at all, so a file imported for one symbol while three of its
83
+ other exports were dead looked healthy.
84
+
85
+ It is candidates, never verdicts, and it gates nothing. The extractor is regex
86
+ over comment-stripped source, not a parser (ADR-0010), so dynamic dispatch,
87
+ reflection, string-keyed lookup and a public API consumed outside this repo are
88
+ indistinguishable from dead code here. Four classes are therefore excluded and
89
+ COUNTED rather than listed, because they could not carry a reference edge however
90
+ heavily they are used: a kind outside the stack's `referenceKinds`, a name
91
+ declared in more than one place (the builder drops ambiguous tokens), a nested
92
+ declaration, and anything declared in a test file. Symbols referenced ONLY from
93
+ tests are listed separately - that is not dead code, it is code whose only
94
+ consumer is its own test, which is worth knowing before a plan calls it
95
+ load-bearing.
96
+
77
97
  ### The graph is drawable, and one place already asks for it
78
98
 
79
99
  The PR body's Impact Analysis, part 3, asks which symbols and files a change
@@ -0,0 +1,93 @@
1
+ # cost analysis - projection, anomaly, burn, diff
2
+
3
+ `cost-budget-check.mjs` answers one question, live, for one run: is THIS task
4
+ about to cross its ceiling. Two ways a budget actually empties are invisible to
5
+ it. A slow drift never trips any single run's ceiling. And one pathological
6
+ session can burn a week's worth in an hour while every individual run stays
7
+ comfortably under its cap.
8
+
9
+ `cost-analyze.mjs` answers the other four questions off one series.
10
+
11
+ | Sub-command | Question |
12
+ |---|---|
13
+ | `projection` | at this rate, what does a week / a month / a quarter cost, and when does a stated budget run out |
14
+ | `anomaly` | which day or session is out of family |
15
+ | `burn` | is the last day accelerating against the preceding week |
16
+ | `diff` | two saved snapshots, side by side |
17
+
18
+ ```bash
19
+ node "$HOME/.claude/scripts/cost-analyze.mjs" projection --days 14 [--monthly-usd 200]
20
+ node "$HOME/.claude/scripts/cost-analyze.mjs" anomaly --days 30 [--by day|session] [--threshold 3.5]
21
+ node "$HOME/.claude/scripts/cost-analyze.mjs" burn [--factor 3]
22
+ node "$HOME/.claude/scripts/cost-analyze.mjs" --save [--days 30]
23
+ node "$HOME/.claude/scripts/cost-analyze.mjs" diff [a.json b.json]
24
+ ```
25
+
26
+ Exit 0 means nothing to report, 10 means a finding, 2 is a usage error. Every
27
+ sub-command takes `--json`, and the human output and the JSON are rendered from
28
+ the same numbers.
29
+
30
+ ## Where the numbers come from, and what that makes them worth
31
+
32
+ The per-run token accumulators in `tracker-state.json` are the pipeline's own
33
+ ledger, and they are the obvious source. They are also empty: the schema has
34
+ `tokens_in` / `tokens_out` / `tokens_cached` per phase, they are written only
35
+ when a phase reports them, and on a real machine no tracker carries them at all.
36
+
37
+ The series that does exist is the host's own transcripts, one usage record per
38
+ assistant turn, under `~/.claude/projects/<slug>/<session>.jsonl`. So that is
39
+ what this reads, and three consequences follow that are stated here rather than
40
+ discovered later:
41
+
42
+ - It covers everything Claude Code did on this machine, not only pipeline runs.
43
+ - The figures are ESTIMATES priced from `cost-table.json`, at LIST price. On a
44
+ subscription they are the right number for comparing one day to another and
45
+ the wrong number to call a bill. The human output says so on every run.
46
+ - A host with no such transcripts (Copilot, Codex) reports `UNMEASURED`, never
47
+ zero. Zero would read as "you spent nothing" when it means "I could not see".
48
+
49
+ Records whose model is not in `cost-table.json` are COUNTED as unpriced and left
50
+ out of the total. Folding them in at zero would make an unknown model look free,
51
+ which is the direction that hurts.
52
+
53
+ Cache writes are priced at 1.25x input, derived rather than read from the table,
54
+ because `cost-table.json` carries a cache READ rate only. Dropping the component
55
+ would make every long-context session look cheap.
56
+
57
+ ## Why the median and the MAD, not the mean
58
+
59
+ The single expensive session `anomaly` exists to find is also the observation
60
+ that inflates a mean and a standard deviation - it hides inside the statistic
61
+ measured against it. With six ordinary days and one at fifty times the rest, a
62
+ mean-based z-score puts the outlier at about 2.3 sigma, under every usual
63
+ threshold. The median does not move for one observation, and neither does the
64
+ median absolute deviation.
65
+
66
+ The score is the modified z-score of Iglewicz and Hoaglin, `0.6745 * (x - median)
67
+ / MAD`, flagged at 3.5.
68
+
69
+ When more than half the values are identical the MAD is zero. That is not an
70
+ error and not a licence to divide by zero: the published fallback is the mean
71
+ absolute deviation scaled by 1.253314. When THAT is zero too there is no
72
+ dispersion at all, and the honest answer is that no anomaly can be called -
73
+ which is what it says.
74
+
75
+ Fewer than five points is refused outright, with the number it needs.
76
+
77
+ ## Why projection divides by the calendar window
78
+
79
+ Total over the window, divided by the WINDOW, not by the days that happen to
80
+ have data. Dividing by active days answers a different question - what a working
81
+ day costs - and then projecting a month as thirty working days overstates it by
82
+ roughly a third.
83
+
84
+ `--monthly-usd` is optional and there is no default. Without it the projection
85
+ is reported and the exhaustion date is not, with the reason given. Inventing a
86
+ budget to compare against would produce a number that looks measured.
87
+
88
+ ## Snapshots
89
+
90
+ `--save` writes `~/.claude/state/cost/<iso>.json`: the window, the total, the
91
+ per-day rate and the per-day breakdown. `diff` compares two, by path or the last
92
+ two by default, and says when the windows differ - the per-day rate is
93
+ comparable across different windows, the total is not.
@@ -8,6 +8,7 @@
8
8
  - [The line shape](#the-line-shape)
9
9
  - [It recommends, it never fixes](#it-recommends-it-never-fixes)
10
10
  - [Checks](#checks)
11
+ - [The server profile (`--profile=server`)](#the-server-profile---profileserver)
11
12
  <!-- /toc -->
12
13
 
13
14
  Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
@@ -212,6 +213,29 @@ package in the npx cache left the server dying on `Cannot find module` at
212
213
  startup, which the client surfaces only as `CONNECTION_CLOSED`, and a check that
213
214
  read the registration and stopped would have called that healthy.
214
215
 
216
+ ### mcp-surface
217
+
218
+ How many MCP servers are charged in the project the caller is standing in,
219
+ counted across three places that all cost the same: the global blocks in
220
+ `~/.claude.json` and `~/.claude/settings.json`, the per-project block inside
221
+ `~/.claude.json`, and a `.mcp.json` committed to the repo. INFO once the count
222
+ exceeds `prefs.global.mcpSurface.infoAbove` (default 8); `--explain` lists the
223
+ names with the scope each came from.
224
+
225
+ Servers registered by a marketplace plugin are NOT counted: nothing in the
226
+ config names them, and inventing a number is worse than reporting the one that
227
+ is countable.
228
+
229
+ The reason it exists: every registered server sends its tool list on every turn,
230
+ the user adds them one at a time, and nobody ever sees the running total - our
231
+ own toolkit contributes 99 tools by itself. This is the same argument that made
232
+ the pre-run context budget a measured number rather than an intention.
233
+
234
+ It only reports. It never disables a server, never blocks and never warns: how
235
+ many servers are worth their context is the user's call, not a health failure.
236
+ The threshold is judgement, which is why it is a pref instead of a constant in
237
+ the script - a number nobody can see is a number nobody can argue with.
238
+
215
239
  ### disk-space
216
240
 
217
241
  Free space on the volume holding `$HOME`. WARN under 2 GB: a worktree plus a
@@ -229,3 +253,47 @@ a full second checkout, so the total is measured in gigabytes rather than
229
253
  megabytes. SKIP outside a git repository - a project name in prefs is a name,
230
254
  not a path, and guessing checkout locations to produce a number is how a
231
255
  diagnostic starts lying.
256
+
257
+ ## The server profile (`--profile=server`)
258
+
259
+ Four checks that only run under `--profile=server`. They are ADDITIVE: a
260
+ default run is byte-for-byte what it was, and `--list-checks` reports whichever
261
+ set the invocation would actually perform, because a list that disagrees with
262
+ the run is worse than no list.
263
+
264
+ None of them may BLOCK. The blocking set is closed at six and these are
265
+ readiness, not correctness: a machine that fails all four still runs the
266
+ pipeline perfectly well with someone watching it. What they catch is the
267
+ failure mode of an unwatched run - stopping without saying so.
268
+
269
+ ### unattended-contract
270
+
271
+ Is `multi-agent-refs/unattended-contract.md` installed. It is the only place
272
+ that says which entry points honour `MULTI_AGENT_UNATTENDED=1` and what each
273
+ one resolves to. Without it an operator setting up a server has to read the
274
+ scripts to find out, which is how the wrong assumption gets made.
275
+
276
+ ### unattended-permissions
277
+
278
+ Does `settings.json` carry `permissions.allow` entries covering Bash, Edit and
279
+ Write. autopilot spawns its child with `--permission-prompts none`, but the
280
+ tools that child then calls still have to be allowed, and on a fresh machine
281
+ they are not - the run stops at the first prompt with nobody there to answer
282
+ it, printing nothing. This is the single most likely reason a server sits idle
283
+ and looks healthy.
284
+
285
+ ### scheduler
286
+
287
+ Is a multi-agent launchd agent loaded. Something has to start the work; a
288
+ server with no agent and an empty queue looks exactly like a server that has
289
+ finished everything.
290
+
291
+ ### keychain-unlock
292
+
293
+ Is a default keychain reachable from this session. Every credential the
294
+ pipeline reads lives there, and a LaunchDaemon started before login sees a
295
+ LOCKED keychain: each fetch fails, and the failures surface much later as 401s
296
+ that blame the token rather than the lock. The check is a proxy - reading a
297
+ real credential would be a network call and a side effect - so it asks whether
298
+ a login keychain is reachable at all, which is the condition a boot-time
299
+ daemon fails.