@mmerterden/multi-agent-pipeline 17.6.0 → 18.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +127 -0
- package/README.md +43 -1
- package/README.tr.md +41 -0
- package/docs/adr/0011-dormant-ci.md +25 -1
- package/docs/server-readiness.md +188 -0
- package/index.js +16 -1
- package/install/_common.mjs +42 -17
- package/install/_dev-only-files.mjs +8 -0
- package/install/_unattended-profile.mjs +113 -0
- package/install/index.mjs +48 -0
- package/manifest.json +1049 -0
- package/package.json +5 -2
- package/pipeline/commands/multi-agent/status/SKILL.md +52 -21
- package/pipeline/lib/_jira-auth.sh +8 -0
- package/pipeline/lib/analysis-jira-write.sh +32 -0
- package/pipeline/lib/ask-choice.sh +13 -2
- package/pipeline/lib/autopilot-state.sh +8 -0
- package/pipeline/lib/fatal.mjs +129 -0
- package/pipeline/lib/figma-mcp-refresh.sh +18 -0
- package/pipeline/lib/figma-screenshot.sh +18 -0
- package/pipeline/lib/invoked-directly.mjs +43 -0
- package/pipeline/lib/jira-publish.sh +42 -0
- package/pipeline/lib/md2confluence-v3.py +47 -0
- package/pipeline/lib/outbound-gate.mjs +175 -0
- package/pipeline/lib/plan-todos.sh +27 -6
- package/pipeline/lib/post-pr-review.sh +77 -8
- package/pipeline/lib/repo-hygiene.sh +8 -3
- package/pipeline/lib/require-jq.sh +40 -0
- package/pipeline/lib/run-paths.sh +335 -0
- package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +70 -0
- package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
- package/pipeline/multi-agent-refs/features/doctor.md +45 -0
- package/pipeline/multi-agent-refs/features/verify.md +83 -0
- package/pipeline/multi-agent-refs/phases/operations.md +13 -2
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
- package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
- package/pipeline/scripts/_run-paths.mjs +372 -0
- package/pipeline/scripts/aggregate-metrics.mjs +64 -64
- package/pipeline/scripts/autopilot-arming.mjs +2 -1
- package/pipeline/scripts/autopilot-intake.mjs +2 -1
- package/pipeline/scripts/autopilot-runner.mjs +206 -2
- package/pipeline/scripts/build-references.mjs +2 -1
- package/pipeline/scripts/build-stack-plugins.mjs +10 -2
- package/pipeline/scripts/capture-evidence.sh +7 -2
- package/pipeline/scripts/classify-plan-safety.mjs +2 -1
- package/pipeline/scripts/cost-analyze.mjs +600 -0
- package/pipeline/scripts/cost-budget-check.mjs +4 -12
- package/pipeline/scripts/council-view.mjs +2 -1
- package/pipeline/scripts/crush-json.mjs +2 -1
- package/pipeline/scripts/diff-explain.mjs +6 -9
- package/pipeline/scripts/diff-risk-score.mjs +2 -1
- package/pipeline/scripts/doctor.mjs +138 -4
- package/pipeline/scripts/evidence-gate.mjs +9 -3
- package/pipeline/scripts/feedback-send.mjs +12 -2
- package/pipeline/scripts/gc-abandoned.sh +29 -13
- package/pipeline/scripts/gc-worktrees.sh +11 -4
- package/pipeline/scripts/github-ssh-setup.sh +64 -7
- package/pipeline/scripts/graph-mermaid.mjs +4 -2
- package/pipeline/scripts/keychain-save.sh +101 -30
- package/pipeline/scripts/learn-from-transcripts.mjs +2 -1
- package/pipeline/scripts/learning-curve.mjs +34 -29
- package/pipeline/scripts/make-manifest.mjs +199 -0
- package/pipeline/scripts/migrate-prefs.mjs +2 -1
- package/pipeline/scripts/migrate-state.mjs +94 -4
- package/pipeline/scripts/phase-banner.sh +6 -2
- package/pipeline/scripts/phase-tracker.sh +41 -3
- package/pipeline/scripts/plan-coverage-gate.mjs +6 -2
- package/pipeline/scripts/pre-commit-check.sh +7 -0
- package/pipeline/scripts/pre-push-check.sh +7 -0
- package/pipeline/scripts/purge.sh +23 -6
- package/pipeline/scripts/render-agent-log-cost.sh +9 -2
- package/pipeline/scripts/render-cost-summary.sh +9 -2
- package/pipeline/scripts/render-work-summary.sh +11 -4
- package/pipeline/scripts/review-file-filter.mjs +4 -2
- package/pipeline/scripts/review-scope.mjs +2 -1
- package/pipeline/scripts/routine-registry.mjs +2 -1
- package/pipeline/scripts/run-aggregator.mjs +13 -14
- package/pipeline/scripts/run-metrics.mjs +3 -1
- package/pipeline/scripts/runs-index.mjs +343 -0
- package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
- package/pipeline/scripts/search-logs.sh +18 -0
- package/pipeline/scripts/test-gap-scan.mjs +2 -1
- package/pipeline/scripts/test-integrity-gate.mjs +2 -1
- package/pipeline/scripts/update-issue-progress.sh +56 -7
- package/pipeline/scripts/usage-report.mjs +12 -1
- package/pipeline/scripts/validate-analysis-doc.mjs +2 -1
- package/pipeline/scripts/validate-code-graph.mjs +6 -3
- package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
- package/pipeline/scripts/validate-diff-risk.mjs +6 -3
- package/pipeline/scripts/validate-test-gap.mjs +6 -3
- package/pipeline/scripts/validate-triage.mjs +3 -1
- package/pipeline/scripts/verify-citations.mjs +4 -2
- package/pipeline/scripts/verify.mjs +327 -0
- package/pipeline/scripts/worktree-finalize.sh +13 -4
- package/pipeline/scripts/write-state.mjs +154 -15
- package/pipeline/skills/.skill-manifest.json +2 -2
- package/pipeline/skills/.skills-index.json +56 -1
- package/pipeline/skills/shared/README.md +8 -3
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +33 -9
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
- package/pipeline/skills/skills-index.md +6 -1
|
@@ -0,0 +1,335 @@
|
|
|
1
|
+
#!/bin/bash
|
|
2
|
+
# run-paths.sh - the single resolver for pipeline run-state paths (shell side).
|
|
3
|
+
#
|
|
4
|
+
# Shell twin of pipeline/scripts/_run-paths.mjs. The two must agree; the
|
|
5
|
+
# contract is asserted by pipeline/scripts/smoke-run-path-canonical.sh, which
|
|
6
|
+
# runs both over the same fixture tree and diffs their answers.
|
|
7
|
+
#
|
|
8
|
+
# A run's files live under the log root in ONE of two layouts:
|
|
9
|
+
#
|
|
10
|
+
# nested <root>/<project>/<task_id>/ documented canonical
|
|
11
|
+
# flat <root>/<task_id>/ what phase-tracker.sh writes
|
|
12
|
+
#
|
|
13
|
+
# Both are populated on real machines, so every reader accepts both. This file
|
|
14
|
+
# does NOT change where anything is written: measured on a real install, 90 of
|
|
15
|
+
# 103 runs are flat and every tracker-state.json is. A task id present in both
|
|
16
|
+
# layouts is ONE run; the most recently written directory wins. Relocation is
|
|
17
|
+
# opt-in and explicit: `migrate-state.mjs --relocate`.
|
|
18
|
+
#
|
|
19
|
+
# No `set -e` on purpose: this is a sourced library and changing the caller's
|
|
20
|
+
# shell options is not this file's business.
|
|
21
|
+
#
|
|
22
|
+
# Usage:
|
|
23
|
+
# . "$(dirname "$0")/../lib/run-paths.sh"
|
|
24
|
+
# root=$(ma_logs_root)
|
|
25
|
+
# dir=$(ma_resolve_run_dir "{JIRA_KEY}-123") || echo "unknown run"
|
|
26
|
+
# f=$(ma_resolve_run_file "{JIRA_KEY}-123" tracker-state.json)
|
|
27
|
+
# ma_list_runs # task_id<TAB>project<TAB>dir<TAB>layout, one per run
|
|
28
|
+
#
|
|
29
|
+
# Bash 3.2 compatible: no associative arrays, no `mapfile`, no `grep -P`.
|
|
30
|
+
|
|
31
|
+
# Marker files and reserved names are spelled out at each test rather than held
|
|
32
|
+
# in a space-separated variable iterated with `for x in $VAR`. bash word-splits
|
|
33
|
+
# an unquoted variable; zsh does not, so there the loop would run ONCE with the
|
|
34
|
+
# whole string as a single word, every test would fail, and the library would
|
|
35
|
+
# answer "no runs" instead of erroring. This file is sourced, so the caller's
|
|
36
|
+
# shell decides - and answering zero is the failure mode that looks like data.
|
|
37
|
+
|
|
38
|
+
ma_logs_root() {
|
|
39
|
+
printf '%s\n' "${LOGS_ROOT:-$HOME/.claude/logs/multi-agent}"
|
|
40
|
+
}
|
|
41
|
+
|
|
42
|
+
# ma_is_run_dir <dir> -> 0 when the directory holds at least one run marker
|
|
43
|
+
# Phase 6 removes the worktree and salvages the run's files into `artifacts/`
|
|
44
|
+
# inside the same run directory, so a shipped run keeps its state one level
|
|
45
|
+
# deeper. A reader that only looks at the top level reports it as stateless.
|
|
46
|
+
MA_ARTIFACTS_SUBDIR=artifacts
|
|
47
|
+
|
|
48
|
+
ma_is_run_dir() {
|
|
49
|
+
local d="$1"
|
|
50
|
+
[ -f "$d/agent-state.json" ] && return 0
|
|
51
|
+
[ -f "$d/tracker-state.json" ] && return 0
|
|
52
|
+
[ -f "$d/agent-log.md" ] && return 0
|
|
53
|
+
[ -f "$d/$MA_ARTIFACTS_SUBDIR/agent-state.json" ] && return 0
|
|
54
|
+
[ -f "$d/$MA_ARTIFACTS_SUBDIR/tracker-state.json" ] && return 0
|
|
55
|
+
[ -f "$d/$MA_ARTIFACTS_SUBDIR/agent-log.md" ] && return 0
|
|
56
|
+
return 1
|
|
57
|
+
}
|
|
58
|
+
|
|
59
|
+
ma_is_reserved() {
|
|
60
|
+
case "$1" in
|
|
61
|
+
review-watch | jira-backups | shadow-git | _analysis-jira | prompts) return 0 ;;
|
|
62
|
+
esac
|
|
63
|
+
return 1
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
# Newest mtime across a run directory's markers, as epoch seconds. GNU-first
|
|
67
|
+
# `stat -c` ordering is required: `stat -f` is a valid GNU flag
|
|
68
|
+
# (--file-system) that SUCCEEDS, so BSD-first would silently win on Linux.
|
|
69
|
+
ma_run_mtime() {
|
|
70
|
+
local d="$1" m t newest=0 b
|
|
71
|
+
for b in "$d" "$d/$MA_ARTIFACTS_SUBDIR"; do
|
|
72
|
+
for m in agent-state.json tracker-state.json agent-log.md; do
|
|
73
|
+
[ -f "$b/$m" ] || continue
|
|
74
|
+
# -L follows symlinks. Some nested run directories are symlink bridges to
|
|
75
|
+
# the flat copy of the same run (3 such pairs on the install this was
|
|
76
|
+
# measured on); without -L the shell reads the LINK's mtime while the
|
|
77
|
+
# .mjs twin's fs.statSync reads the TARGET's, and the two picked
|
|
78
|
+
# different winners for one record. Both must read the same clock.
|
|
79
|
+
t=$(stat -L -c %Y "$b/$m" 2>/dev/null || stat -L -f %m "$b/$m" 2>/dev/null || echo 0)
|
|
80
|
+
[ -n "$t" ] || t=0
|
|
81
|
+
[ "$t" -gt "$newest" ] && newest="$t"
|
|
82
|
+
done
|
|
83
|
+
done
|
|
84
|
+
printf '%s\n' "$newest"
|
|
85
|
+
}
|
|
86
|
+
|
|
87
|
+
# ma_canonical_run_dir <task_id> [project] -> where a NEW run should be written
|
|
88
|
+
ma_canonical_run_dir() {
|
|
89
|
+
local task_id="$1" project="${2:-}" root
|
|
90
|
+
root=$(ma_logs_root)
|
|
91
|
+
if [ -n "$project" ]; then
|
|
92
|
+
printf '%s\n' "$root/$project/$task_id"
|
|
93
|
+
else
|
|
94
|
+
printf '%s\n' "$root/$task_id"
|
|
95
|
+
fi
|
|
96
|
+
}
|
|
97
|
+
|
|
98
|
+
# ma_run_dir_candidates <task_id> [project] -> candidate dirs, most specific
|
|
99
|
+
# first. Existence is not checked here.
|
|
100
|
+
ma_run_dir_candidates() {
|
|
101
|
+
local task_id="$1" project="${2:-}" root d n
|
|
102
|
+
root=$(ma_logs_root)
|
|
103
|
+
[ -n "$project" ] && printf '%s\n' "$root/$project/$task_id"
|
|
104
|
+
printf '%s\n' "$root/$task_id"
|
|
105
|
+
[ -n "$project" ] && return 0
|
|
106
|
+
# No project given: the run may still be nested under one.
|
|
107
|
+
#
|
|
108
|
+
# `find` rather than a `*/` glob on purpose: an unmatched glob is a literal
|
|
109
|
+
# string in bash and a hard error in zsh, and this file is sourced, so the
|
|
110
|
+
# caller's shell decides. find behaves identically in both.
|
|
111
|
+
while IFS= read -r d; do
|
|
112
|
+
[ -n "$d" ] || continue
|
|
113
|
+
n=$(basename "$d")
|
|
114
|
+
ma_is_reserved "$n" && continue
|
|
115
|
+
[ "$n" = "$task_id" ] && continue
|
|
116
|
+
ma_is_run_dir "$root/$n/$task_id" && printf '%s\n' "$root/$n/$task_id"
|
|
117
|
+
done <<EOF
|
|
118
|
+
$(find "$root" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
|
|
119
|
+
EOF
|
|
120
|
+
return 0
|
|
121
|
+
}
|
|
122
|
+
|
|
123
|
+
# ma_realpath <path> -> the path with every symlink resolved, or the input when
|
|
124
|
+
# it cannot be resolved. BSD realpath first, python3 as the fallback (both are
|
|
125
|
+
# already hard requirements of this tree).
|
|
126
|
+
ma_realpath() {
|
|
127
|
+
realpath "$1" 2>/dev/null ||
|
|
128
|
+
python3 -c 'import os,sys; print(os.path.realpath(sys.argv[1]))' "$1" 2>/dev/null ||
|
|
129
|
+
printf '%s\n' "$1"
|
|
130
|
+
}
|
|
131
|
+
|
|
132
|
+
# ma_is_bridge_dir <dir> -> 0 when any marker in it is a symlink elsewhere.
|
|
133
|
+
# Some nested run directories are symlink bridges into the flat copy of the
|
|
134
|
+
# same run. They are one record seen twice, not two copies.
|
|
135
|
+
ma_is_bridge_dir() {
|
|
136
|
+
local d="$1" m
|
|
137
|
+
for m in agent-state.json tracker-state.json agent-log.md; do
|
|
138
|
+
[ -L "$d/$m" ] && return 0
|
|
139
|
+
done
|
|
140
|
+
return 1
|
|
141
|
+
}
|
|
142
|
+
|
|
143
|
+
# ma_same_record <dir_a> <dir_b> -> 0 when both are views of ONE record.
|
|
144
|
+
ma_same_record() {
|
|
145
|
+
local a="$1" b="$2" m ra rb
|
|
146
|
+
for m in agent-state.json tracker-state.json agent-log.md; do
|
|
147
|
+
[ -e "$a/$m" ] && [ -e "$b/$m" ] || continue
|
|
148
|
+
ra=$(ma_realpath "$a/$m")
|
|
149
|
+
rb=$(ma_realpath "$b/$m")
|
|
150
|
+
[ "$ra" = "$rb" ] && return 0
|
|
151
|
+
done
|
|
152
|
+
return 1
|
|
153
|
+
}
|
|
154
|
+
|
|
155
|
+
# ma_rank_fields <dir> -> "<mtime>\t<has_state>\t<depth>", the three ranking
|
|
156
|
+
# columns used to pick between candidate directories for one task id.
|
|
157
|
+
#
|
|
158
|
+
# ma_sort_key packs the same three into one space-separated key for callers
|
|
159
|
+
# that sort a single column; both must match compareCandidates in
|
|
160
|
+
# _run-paths.mjs exactly.
|
|
161
|
+
#
|
|
162
|
+
# mtime alone is not an order: two directories written in the same second tie,
|
|
163
|
+
# and each implementation then fell back to its own traversal order - which is
|
|
164
|
+
# how the two came to disagree about a run present in both layouts. The tail of
|
|
165
|
+
# the key defines the answer: richer record first (agent-state.json is what
|
|
166
|
+
# every reader wants), then the documented nested layout, then the path.
|
|
167
|
+
ma_rank_fields() {
|
|
168
|
+
local d="$1" t has_state depth real
|
|
169
|
+
t=$(ma_run_mtime "$d")
|
|
170
|
+
has_state=0
|
|
171
|
+
[ -f "$d/agent-state.json" ] && has_state=1
|
|
172
|
+
depth=$(printf '%s' "$d" | awk -F/ '{print NF}')
|
|
173
|
+
real=1
|
|
174
|
+
ma_is_bridge_dir "$d" && real=0
|
|
175
|
+
printf '%s\t%s\t%s\t%s\n' "$real" "$t" "$has_state" "$depth"
|
|
176
|
+
}
|
|
177
|
+
|
|
178
|
+
ma_sort_key() {
|
|
179
|
+
local d="$1" t has_state depth real
|
|
180
|
+
t=$(ma_run_mtime "$d")
|
|
181
|
+
has_state=0
|
|
182
|
+
[ -f "$d/agent-state.json" ] && has_state=1
|
|
183
|
+
depth=$(printf '%s' "$d" | awk -F/ '{print NF}')
|
|
184
|
+
real=1
|
|
185
|
+
ma_is_bridge_dir "$d" && real=0
|
|
186
|
+
printf '%d %012d %d %04d %s\n' "$real" "$t" "$has_state" "$depth" "$d"
|
|
187
|
+
}
|
|
188
|
+
|
|
189
|
+
# ma_resolve_run_dir <task_id> [project] -> the directory, or rc 1 when unknown.
|
|
190
|
+
ma_resolve_run_dir() {
|
|
191
|
+
local task_id="$1" project="${2:-}" c best
|
|
192
|
+
best=$(
|
|
193
|
+
while IFS= read -r c; do
|
|
194
|
+
[ -n "$c" ] || continue
|
|
195
|
+
ma_is_run_dir "$c" || continue
|
|
196
|
+
ma_sort_key "$c"
|
|
197
|
+
done <<EOF
|
|
198
|
+
$(ma_run_dir_candidates "$task_id" "$project")
|
|
199
|
+
EOF
|
|
200
|
+
)
|
|
201
|
+
[ -n "$best" ] || return 1
|
|
202
|
+
# Descending on mtime, then has-state, then depth; ascending on path.
|
|
203
|
+
printf '%s\n' "$best" | sort -k1,1nr -k2,2nr -k3,3nr -k4,4nr -k5,5 | head -1 | cut -d' ' -f5-
|
|
204
|
+
}
|
|
205
|
+
|
|
206
|
+
# ma_task_id_variants <id> -> the spellings one task id has been written under,
|
|
207
|
+
# in preference order. `#316` arrives from a GitHub issue reference, `316` from
|
|
208
|
+
# the bare-number input class, `task-316` from an older directory convention.
|
|
209
|
+
# Callers used to inline this list; dropping a spelling silently stops old runs
|
|
210
|
+
# from resolving.
|
|
211
|
+
ma_task_id_variants() {
|
|
212
|
+
local raw="$1" bare
|
|
213
|
+
bare="${raw#\#}"
|
|
214
|
+
printf '%s\n' "$raw"
|
|
215
|
+
[ "$bare" != "$raw" ] && printf '%s\n' "$bare"
|
|
216
|
+
printf '%s\n' "task-$bare"
|
|
217
|
+
}
|
|
218
|
+
|
|
219
|
+
# ma_resolve_run_dir_any <task_id> [project] -> dir for any spelling, or rc 1
|
|
220
|
+
ma_resolve_run_dir_any() {
|
|
221
|
+
local task_id="$1" project="${2:-}" v dir
|
|
222
|
+
while IFS= read -r v; do
|
|
223
|
+
[ -n "$v" ] || continue
|
|
224
|
+
dir=$(ma_resolve_run_dir "$v" "$project") && {
|
|
225
|
+
printf '%s\n' "$dir"
|
|
226
|
+
return 0
|
|
227
|
+
}
|
|
228
|
+
done <<EOF
|
|
229
|
+
$(ma_task_id_variants "$task_id")
|
|
230
|
+
EOF
|
|
231
|
+
return 1
|
|
232
|
+
}
|
|
233
|
+
|
|
234
|
+
# ma_resolve_run_file <task_id> <filename> [project] -> path, or rc 1.
|
|
235
|
+
# Every spelling of the id is tried.
|
|
236
|
+
ma_resolve_run_file() {
|
|
237
|
+
local task_id="$1" filename="$2" project="${3:-}" v dir
|
|
238
|
+
while IFS= read -r v; do
|
|
239
|
+
[ -n "$v" ] || continue
|
|
240
|
+
dir=$(ma_resolve_run_dir "$v" "$project") || continue
|
|
241
|
+
if [ -f "$dir/$filename" ]; then
|
|
242
|
+
printf '%s\n' "$dir/$filename"
|
|
243
|
+
return 0
|
|
244
|
+
fi
|
|
245
|
+
# The salvaged copy Phase 6 leaves behind.
|
|
246
|
+
if [ -f "$dir/$MA_ARTIFACTS_SUBDIR/$filename" ]; then
|
|
247
|
+
printf '%s\n' "$dir/$MA_ARTIFACTS_SUBDIR/$filename"
|
|
248
|
+
return 0
|
|
249
|
+
fi
|
|
250
|
+
done <<EOF
|
|
251
|
+
$(ma_task_id_variants "$task_id")
|
|
252
|
+
EOF
|
|
253
|
+
return 1
|
|
254
|
+
}
|
|
255
|
+
|
|
256
|
+
# ma_list_runs -> one TAB-separated row per run, deduplicated by task id:
|
|
257
|
+
# task_id<TAB>project<TAB>dir<TAB>layout
|
|
258
|
+
# `project` is "-" when the layout does not name one.
|
|
259
|
+
ma_list_runs() {
|
|
260
|
+
local root d n kd kn rows row id
|
|
261
|
+
root=$(ma_logs_root)
|
|
262
|
+
[ -d "$root" ] || return 0
|
|
263
|
+
rows=""
|
|
264
|
+
while IFS= read -r d; do
|
|
265
|
+
[ -n "$d" ] || continue
|
|
266
|
+
n=$(basename "$d")
|
|
267
|
+
ma_is_reserved "$n" && continue
|
|
268
|
+
if ma_is_run_dir "$root/$n"; then
|
|
269
|
+
rows="$rows$n - $root/$n flat $(ma_rank_fields "$root/$n")
|
|
270
|
+
"
|
|
271
|
+
continue
|
|
272
|
+
fi
|
|
273
|
+
while IFS= read -r kd; do
|
|
274
|
+
[ -n "$kd" ] || continue
|
|
275
|
+
kn=$(basename "$kd")
|
|
276
|
+
ma_is_run_dir "$root/$n/$kn" || continue
|
|
277
|
+
rows="$rows$kn $n $root/$n/$kn nested $(ma_rank_fields "$root/$n/$kn")
|
|
278
|
+
"
|
|
279
|
+
done <<EOF
|
|
280
|
+
$(find "$root/$n" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
|
|
281
|
+
EOF
|
|
282
|
+
done <<EOF
|
|
283
|
+
$(find "$root" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
|
|
284
|
+
EOF
|
|
285
|
+
# Dedup by task id using the same total order as resolveRunDir and as
|
|
286
|
+
# compareCandidates in the .mjs twin: mtime desc, has-state desc, depth desc,
|
|
287
|
+
# path ASC. The sub-keys are separate columns on purpose - a single reverse
|
|
288
|
+
# sort over a composite key also reverses the path, which is the one
|
|
289
|
+
# component that has to ascend.
|
|
290
|
+
printf '%s' "$rows" | grep -v '^$' |
|
|
291
|
+
sort -t' ' -k1,1 -k5,5nr -k6,6nr -k7,7nr -k8,8nr -k3,3 | awk -F' ' '
|
|
292
|
+
!seen[$1]++ { print $1 "\t" $2 "\t" $3 "\t" $4 }
|
|
293
|
+
' | sort -t' ' -k1,1
|
|
294
|
+
}
|
|
295
|
+
|
|
296
|
+
# ma_duplicate_run_ids -> task ids that exist as two SEPARATE records in the two
|
|
297
|
+
# layouts, one per line. A symlink bridge is excluded: it is one record seen
|
|
298
|
+
# twice, and reporting it overstates the drift.
|
|
299
|
+
ma_duplicate_run_ids() {
|
|
300
|
+
local root d n kd kn pairs id dirs a b
|
|
301
|
+
root=$(ma_logs_root)
|
|
302
|
+
[ -d "$root" ] || return 0
|
|
303
|
+
pairs=$(
|
|
304
|
+
while IFS= read -r d; do
|
|
305
|
+
[ -n "$d" ] || continue
|
|
306
|
+
n=$(basename "$d")
|
|
307
|
+
ma_is_reserved "$n" && continue
|
|
308
|
+
if ma_is_run_dir "$root/$n"; then
|
|
309
|
+
printf '%s\t%s\n' "$n" "$root/$n"
|
|
310
|
+
continue
|
|
311
|
+
fi
|
|
312
|
+
while IFS= read -r kd; do
|
|
313
|
+
[ -n "$kd" ] || continue
|
|
314
|
+
kn=$(basename "$kd")
|
|
315
|
+
ma_is_run_dir "$root/$n/$kn" && printf '%s\t%s\n' "$kn" "$root/$n/$kn"
|
|
316
|
+
done <<EOF
|
|
317
|
+
$(find "$root/$n" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
|
|
318
|
+
EOF
|
|
319
|
+
done <<EOF
|
|
320
|
+
$(find "$root" -mindepth 1 -maxdepth 1 -type d 2>/dev/null)
|
|
321
|
+
EOF
|
|
322
|
+
)
|
|
323
|
+
# Ids seen more than once, minus the ones whose two directories resolve to the
|
|
324
|
+
# same underlying files (a symlink bridge is one record, not drift).
|
|
325
|
+
while IFS= read -r id; do
|
|
326
|
+
[ -n "$id" ] || continue
|
|
327
|
+
dirs=$(printf '%s\n' "$pairs" | awk -F'\t' -v k="$id" '$1==k {print $2}')
|
|
328
|
+
a=$(printf '%s\n' "$dirs" | sed -n 1p)
|
|
329
|
+
b=$(printf '%s\n' "$dirs" | sed -n 2p)
|
|
330
|
+
[ -n "$b" ] || continue
|
|
331
|
+
ma_same_record "$a" "$b" || printf '%s\n' "$id"
|
|
332
|
+
done <<EOF
|
|
333
|
+
$(printf '%s\n' "$pairs" | awk -F'\t' '{print $1}' | sort | uniq -d)
|
|
334
|
+
EOF
|
|
335
|
+
}
|
|
@@ -1,5 +1,14 @@
|
|
|
1
1
|
# Feature: Autopilot Circuit-Breaker
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [Trip conditions](#trip-conditions)
|
|
5
|
+
- [Wiring status](#wiring-status)
|
|
6
|
+
- [Action on trip](#action-on-trip)
|
|
7
|
+
- [Why this is the right autopilot exception](#why-this-is-the-right-autopilot-exception)
|
|
8
|
+
- [The runner-level breaker (a different failure, a different layer)](#the-runner-level-breaker-a-different-failure-a-different-layer)
|
|
9
|
+
- [Two more things a long-running runner needs](#two-more-things-a-long-running-runner-needs)
|
|
10
|
+
<!-- /toc -->
|
|
11
|
+
|
|
3
12
|
**Pattern**: autopilot runs with zero interaction, which is exactly when a silent failure loop is most expensive - an agent can burn a budget re-attempting the same broken fix, or thrash between two phases, with nobody watching. A circuit-breaker converts "keep going no matter what" into "keep going until a defined unsafe condition, then halt and hand back to the user." This is the sanctioned autopilot pause (same class as the Phase 7 channels pause): the run stops, records why, and waits for an explicit `resume`.
|
|
4
13
|
|
|
5
14
|
**Gated by `prefs.global.autopilotCircuitBreaker`** (`enabled` default true, `identicalFindingCycles` default 2, `maxReworkCycles` default 3; `schemas/prefs.schema.json`). Halting is always safe, so the breaker itself defaults on. Disable per-run only with an explicit override. Complements, does not replace, the existing autopilot safety rules (build-fail max 3 retries, Phase 4 blocking-finding rework, destructive-op confirmations).
|
|
@@ -43,3 +52,64 @@ State shape: `state.circuitBreaker = {tripped, trigger, detail, checkpoint: {pha
|
|
|
43
52
|
Autopilot's contract is "no interaction on the happy path." The circuit-breaker fires only off the happy path, where continuing unattended is the *less* safe choice: repeating a proven-broken action, or spending past a ceiling the user set, is not autonomy, it is a runaway. Halting with a precise reason is cheaper than the tokens (and trust) a silent loop burns.
|
|
44
53
|
|
|
45
54
|
Inspired by autonomous-loop operators that gate on explicit stop conditions (no-progress, identical-stack-trace repetition, cost drift, conflict blocking), adapted to this pipeline's phase/rework/state model.
|
|
55
|
+
|
|
56
|
+
## The runner-level breaker (a different failure, a different layer)
|
|
57
|
+
|
|
58
|
+
Everything above is the IN-RUN breaker: one run, one state file, triggers that
|
|
59
|
+
mean "this task is not converging". It cannot see the failure a server actually
|
|
60
|
+
hits, which is not about any one task.
|
|
61
|
+
|
|
62
|
+
An expired token, a `claude` binary that no longer launches, a full disk: every
|
|
63
|
+
item fails identically, a few minutes apart, until the queue is empty. Nothing
|
|
64
|
+
above notices, because each individual run failed for its own apparently local
|
|
65
|
+
reason. What is left afterwards is a ledger of failures with no indication of
|
|
66
|
+
which came first and no queue to resume.
|
|
67
|
+
|
|
68
|
+
`autopilot-runner.mjs` therefore refuses to take NEW work after
|
|
69
|
+
`MA_AP_BREAKER_LIMIT` consecutive attempts produced nothing (default 3, zero
|
|
70
|
+
disables). The count is CONSECUTIVE and read from `attempted.jsonl` rather than
|
|
71
|
+
kept as a counter, because a counter is a second piece of state that can
|
|
72
|
+
disagree with the ledger - and the ledger is what a person reads when they ask
|
|
73
|
+
why the runner stopped.
|
|
74
|
+
|
|
75
|
+
Two outcomes are deliberately not failures:
|
|
76
|
+
|
|
77
|
+
- `needs-input` is a run that did its work and is waiting for a person. That is
|
|
78
|
+
the system behaving correctly, and counting it would stop a machine whose only
|
|
79
|
+
problem is that somebody has not answered yet.
|
|
80
|
+
- `blocked-*` items never ran at all - arming refused them - so they say nothing
|
|
81
|
+
about whether a run would have worked.
|
|
82
|
+
|
|
83
|
+
When it opens, the reason goes into `queue.json.blockedReason` (so
|
|
84
|
+
`autopilot-status` shows it rather than reporting an idle queue), a
|
|
85
|
+
`breaker-open` line goes into `ticks.jsonl`, and the log names the outcome that
|
|
86
|
+
opened it plus the override.
|
|
87
|
+
|
|
88
|
+
It is HALF-OPEN rather than latched, and that distinction is load-bearing. The
|
|
89
|
+
breaker returns before the only writer of `attempted.jsonl`, so a version that
|
|
90
|
+
simply stopped could never record another attempt, the consecutive count could
|
|
91
|
+
never fall, and the runner would be off for good - while its own log said
|
|
92
|
+
"until an attempt succeeds". After `MA_AP_BREAKER_COOLDOWN_SEC` (30 min by
|
|
93
|
+
default) one probe is let through: a recovered machine succeeds and the count
|
|
94
|
+
resets on its own, and a still-broken one fails, becomes the newest attempt, and
|
|
95
|
+
closes the breaker for another cooldown. One wasted run per cooldown is the
|
|
96
|
+
price of not needing a person to notice.
|
|
97
|
+
|
|
98
|
+
## Two more things a long-running runner needs
|
|
99
|
+
|
|
100
|
+
**The log.** launchd appends the runner's stdout to `runner.log` forever. On a
|
|
101
|
+
laptop that file is never read; on a machine ticking every few minutes for
|
|
102
|
+
months it is what fills the disk, and a full disk stops the runs the log existed
|
|
103
|
+
to record. Each tick truncates it past `MA_AP_LOG_MAX_BYTES` (default 5MB),
|
|
104
|
+
keeping the last 256KB in `runner.log.1`.
|
|
105
|
+
|
|
106
|
+
Truncating the same inode is the mechanism, not renaming: launchd holds the file
|
|
107
|
+
open with `O_APPEND`, so a rename leaves that descriptor writing into the renamed
|
|
108
|
+
file while the new one stays empty forever. That is the rotation that looks
|
|
109
|
+
correct and silently stops logging.
|
|
110
|
+
|
|
111
|
+
**Telemetry.** `ticks.jsonl` gets one JSON line per tick - what the tick did,
|
|
112
|
+
what it took, what came out, how many consecutive failures preceded it. The
|
|
113
|
+
human log answers "what happened just now"; this answers "how has this been
|
|
114
|
+
behaving for a week", which is the question a server raises and prose cannot be
|
|
115
|
+
asked.
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
# cost analysis - projection, anomaly, burn, diff
|
|
2
|
+
|
|
3
|
+
`cost-budget-check.mjs` answers one question, live, for one run: is THIS task
|
|
4
|
+
about to cross its ceiling. Two ways a budget actually empties are invisible to
|
|
5
|
+
it. A slow drift never trips any single run's ceiling. And one pathological
|
|
6
|
+
session can burn a week's worth in an hour while every individual run stays
|
|
7
|
+
comfortably under its cap.
|
|
8
|
+
|
|
9
|
+
`cost-analyze.mjs` answers the other four questions off one series.
|
|
10
|
+
|
|
11
|
+
| Sub-command | Question |
|
|
12
|
+
|---|---|
|
|
13
|
+
| `projection` | at this rate, what does a week / a month / a quarter cost, and when does a stated budget run out |
|
|
14
|
+
| `anomaly` | which day or session is out of family |
|
|
15
|
+
| `burn` | is the last day accelerating against the preceding week |
|
|
16
|
+
| `diff` | two saved snapshots, side by side |
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
node "$HOME/.claude/scripts/cost-analyze.mjs" projection --days 14 [--monthly-usd 200]
|
|
20
|
+
node "$HOME/.claude/scripts/cost-analyze.mjs" anomaly --days 30 [--by day|session] [--threshold 3.5]
|
|
21
|
+
node "$HOME/.claude/scripts/cost-analyze.mjs" burn [--factor 3]
|
|
22
|
+
node "$HOME/.claude/scripts/cost-analyze.mjs" --save [--days 30]
|
|
23
|
+
node "$HOME/.claude/scripts/cost-analyze.mjs" diff [a.json b.json]
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
Exit 0 means nothing to report, 10 means a finding, 2 is a usage error. Every
|
|
27
|
+
sub-command takes `--json`, and the human output and the JSON are rendered from
|
|
28
|
+
the same numbers.
|
|
29
|
+
|
|
30
|
+
## Where the numbers come from, and what that makes them worth
|
|
31
|
+
|
|
32
|
+
The per-run token accumulators in `tracker-state.json` are the pipeline's own
|
|
33
|
+
ledger, and they are the obvious source. They are also empty: the schema has
|
|
34
|
+
`tokens_in` / `tokens_out` / `tokens_cached` per phase, they are written only
|
|
35
|
+
when a phase reports them, and on a real machine no tracker carries them at all.
|
|
36
|
+
|
|
37
|
+
The series that does exist is the host's own transcripts, one usage record per
|
|
38
|
+
assistant turn, under `~/.claude/projects/<slug>/<session>.jsonl`. So that is
|
|
39
|
+
what this reads, and three consequences follow that are stated here rather than
|
|
40
|
+
discovered later:
|
|
41
|
+
|
|
42
|
+
- It covers everything Claude Code did on this machine, not only pipeline runs.
|
|
43
|
+
- The figures are ESTIMATES priced from `cost-table.json`, at LIST price. On a
|
|
44
|
+
subscription they are the right number for comparing one day to another and
|
|
45
|
+
the wrong number to call a bill. The human output says so on every run.
|
|
46
|
+
- A host with no such transcripts (Copilot, Codex) reports `UNMEASURED`, never
|
|
47
|
+
zero. Zero would read as "you spent nothing" when it means "I could not see".
|
|
48
|
+
|
|
49
|
+
Records whose model is not in `cost-table.json` are COUNTED as unpriced and left
|
|
50
|
+
out of the total. Folding them in at zero would make an unknown model look free,
|
|
51
|
+
which is the direction that hurts.
|
|
52
|
+
|
|
53
|
+
Cache writes are priced at 1.25x input, derived rather than read from the table,
|
|
54
|
+
because `cost-table.json` carries a cache READ rate only. Dropping the component
|
|
55
|
+
would make every long-context session look cheap.
|
|
56
|
+
|
|
57
|
+
## Why the median and the MAD, not the mean
|
|
58
|
+
|
|
59
|
+
The single expensive session `anomaly` exists to find is also the observation
|
|
60
|
+
that inflates a mean and a standard deviation - it hides inside the statistic
|
|
61
|
+
measured against it. With six ordinary days and one at fifty times the rest, a
|
|
62
|
+
mean-based z-score puts the outlier at about 2.3 sigma, under every usual
|
|
63
|
+
threshold. The median does not move for one observation, and neither does the
|
|
64
|
+
median absolute deviation.
|
|
65
|
+
|
|
66
|
+
The score is the modified z-score of Iglewicz and Hoaglin, `0.6745 * (x - median)
|
|
67
|
+
/ MAD`, flagged at 3.5.
|
|
68
|
+
|
|
69
|
+
When more than half the values are identical the MAD is zero. That is not an
|
|
70
|
+
error and not a licence to divide by zero: the published fallback is the mean
|
|
71
|
+
absolute deviation scaled by 1.253314. When THAT is zero too there is no
|
|
72
|
+
dispersion at all, and the honest answer is that no anomaly can be called -
|
|
73
|
+
which is what it says.
|
|
74
|
+
|
|
75
|
+
Fewer than five points is refused outright, with the number it needs.
|
|
76
|
+
|
|
77
|
+
## Why projection divides by the calendar window
|
|
78
|
+
|
|
79
|
+
Total over the window, divided by the WINDOW, not by the days that happen to
|
|
80
|
+
have data. Dividing by active days answers a different question - what a working
|
|
81
|
+
day costs - and then projecting a month as thirty working days overstates it by
|
|
82
|
+
roughly a third.
|
|
83
|
+
|
|
84
|
+
`--monthly-usd` is optional and there is no default. Without it the projection
|
|
85
|
+
is reported and the exhaustion date is not, with the reason given. Inventing a
|
|
86
|
+
budget to compare against would produce a number that looks measured.
|
|
87
|
+
|
|
88
|
+
## Snapshots
|
|
89
|
+
|
|
90
|
+
`--save` writes `~/.claude/state/cost/<iso>.json`: the window, the total, the
|
|
91
|
+
per-day rate and the per-day breakdown. `diff` compares two, by path or the last
|
|
92
|
+
two by default, and says when the windows differ - the per-day rate is
|
|
93
|
+
comparable across different windows, the total is not.
|
|
@@ -8,6 +8,7 @@
|
|
|
8
8
|
- [The line shape](#the-line-shape)
|
|
9
9
|
- [It recommends, it never fixes](#it-recommends-it-never-fixes)
|
|
10
10
|
- [Checks](#checks)
|
|
11
|
+
- [The server profile (`--profile=server`)](#the-server-profile---profileserver)
|
|
11
12
|
<!-- /toc -->
|
|
12
13
|
|
|
13
14
|
Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
|
|
@@ -252,3 +253,47 @@ a full second checkout, so the total is measured in gigabytes rather than
|
|
|
252
253
|
megabytes. SKIP outside a git repository - a project name in prefs is a name,
|
|
253
254
|
not a path, and guessing checkout locations to produce a number is how a
|
|
254
255
|
diagnostic starts lying.
|
|
256
|
+
|
|
257
|
+
## The server profile (`--profile=server`)
|
|
258
|
+
|
|
259
|
+
Four checks that only run under `--profile=server`. They are ADDITIVE: a
|
|
260
|
+
default run is byte-for-byte what it was, and `--list-checks` reports whichever
|
|
261
|
+
set the invocation would actually perform, because a list that disagrees with
|
|
262
|
+
the run is worse than no list.
|
|
263
|
+
|
|
264
|
+
None of them may BLOCK. The blocking set is closed at six and these are
|
|
265
|
+
readiness, not correctness: a machine that fails all four still runs the
|
|
266
|
+
pipeline perfectly well with someone watching it. What they catch is the
|
|
267
|
+
failure mode of an unwatched run - stopping without saying so.
|
|
268
|
+
|
|
269
|
+
### unattended-contract
|
|
270
|
+
|
|
271
|
+
Is `multi-agent-refs/unattended-contract.md` installed. It is the only place
|
|
272
|
+
that says which entry points honour `MULTI_AGENT_UNATTENDED=1` and what each
|
|
273
|
+
one resolves to. Without it an operator setting up a server has to read the
|
|
274
|
+
scripts to find out, which is how the wrong assumption gets made.
|
|
275
|
+
|
|
276
|
+
### unattended-permissions
|
|
277
|
+
|
|
278
|
+
Does `settings.json` carry `permissions.allow` entries covering Bash, Edit and
|
|
279
|
+
Write. autopilot spawns its child with `--permission-prompts none`, but the
|
|
280
|
+
tools that child then calls still have to be allowed, and on a fresh machine
|
|
281
|
+
they are not - the run stops at the first prompt with nobody there to answer
|
|
282
|
+
it, printing nothing. This is the single most likely reason a server sits idle
|
|
283
|
+
and looks healthy.
|
|
284
|
+
|
|
285
|
+
### scheduler
|
|
286
|
+
|
|
287
|
+
Is a multi-agent launchd agent loaded. Something has to start the work; a
|
|
288
|
+
server with no agent and an empty queue looks exactly like a server that has
|
|
289
|
+
finished everything.
|
|
290
|
+
|
|
291
|
+
### keychain-unlock
|
|
292
|
+
|
|
293
|
+
Is a default keychain reachable from this session. Every credential the
|
|
294
|
+
pipeline reads lives there, and a LaunchDaemon started before login sees a
|
|
295
|
+
LOCKED keychain: each fetch fails, and the failures surface much later as 401s
|
|
296
|
+
that blame the token rather than the lock. The check is a proxy - reading a
|
|
297
|
+
real credential would be a network call and a side effect - so it asks whether
|
|
298
|
+
a login keychain is reachable at all, which is the condition a boot-time
|
|
299
|
+
daemon fails.
|
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
# verify - is this install the thing that was published
|
|
2
|
+
|
|
3
|
+
The install is a COPY. `install.js` writes the pipeline tree into `~/.claude`,
|
|
4
|
+
`~/.copilot` and `~/.codex`, and from that moment the two halves drift
|
|
5
|
+
independently. Both directions produce bugs that are hard to name:
|
|
6
|
+
|
|
7
|
+
- an edit made in the installed copy is a behaviour with no source, and the next
|
|
8
|
+
update silently reverts it;
|
|
9
|
+
- a file the installer failed to write is a script the docs describe and nobody
|
|
10
|
+
has, which reads as a documentation error.
|
|
11
|
+
|
|
12
|
+
`multi-agent-pipeline verify` answers both mechanically.
|
|
13
|
+
|
|
14
|
+
```bash
|
|
15
|
+
npx @mmerterden/multi-agent-pipeline verify # package + install
|
|
16
|
+
npx @mmerterden/multi-agent-pipeline verify --package # package integrity only
|
|
17
|
+
npx @mmerterden/multi-agent-pipeline verify --install # install drift only
|
|
18
|
+
npx @mmerterden/multi-agent-pipeline verify --json
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
| Code | Meaning |
|
|
22
|
+
|---|---|
|
|
23
|
+
| 0 | everything matches |
|
|
24
|
+
| 1 | a difference was found, named file by file |
|
|
25
|
+
| 2 | nothing to verify - a source checkout, or a version published before manifests existed |
|
|
26
|
+
|
|
27
|
+
Exit 2 is not a pass and not a failure. A dev checkout has no manifest by
|
|
28
|
+
design, and reporting that as either would be a lie in one direction or the
|
|
29
|
+
other.
|
|
30
|
+
|
|
31
|
+
## The manifest
|
|
32
|
+
|
|
33
|
+
`manifest.json` is written at pack time by `prepack`, never committed. A
|
|
34
|
+
manifest in git is stale one commit after it is written, and a stale manifest
|
|
35
|
+
reports honest edits as tampering - which is worse than having none, because
|
|
36
|
+
people learn to ignore it.
|
|
37
|
+
|
|
38
|
+
The file list is not guessed. It comes from `npm pack --dry-run --json`, so by
|
|
39
|
+
construction it is the same set npm publishes, `files` globs and all. The gate
|
|
40
|
+
asserts the two counts agree, which is what catches a `files` entry and a
|
|
41
|
+
manifest that have stopped describing the same package.
|
|
42
|
+
|
|
43
|
+
Two things it cannot cover, said here rather than discovered later: it cannot
|
|
44
|
+
hash itself, and a signature over it does not authenticate the tarball.
|
|
45
|
+
|
|
46
|
+
## What a green result proves, and what it does not
|
|
47
|
+
|
|
48
|
+
It proves the bytes match what the publisher recorded. It is not proof of WHO
|
|
49
|
+
published them. The manifest, the signature and the verifier all travel inside
|
|
50
|
+
the same tarball, so anyone able to rewrite one can rewrite the others.
|
|
51
|
+
Provenance belongs to npm's own integrity field.
|
|
52
|
+
|
|
53
|
+
What this does catch is the set of failures that actually happen: a damaged or
|
|
54
|
+
partial install, a file edited after install, and an update that did not land.
|
|
55
|
+
|
|
56
|
+
Signing is optional. `make-manifest.mjs --sign` reads an ed25519 private key
|
|
57
|
+
from the credential store (or `MULTI_AGENT_SIGNING_KEY` on a build host with no
|
|
58
|
+
store) and writes `manifest.sig`; `verify` checks it against
|
|
59
|
+
`MULTI_AGENT_SIGNING_PUBKEY` when one is pinned. Without a key it says "signed,
|
|
60
|
+
no public key to check it against" rather than claiming valid - a signature
|
|
61
|
+
nobody can check is not a signature that passed.
|
|
62
|
+
|
|
63
|
+
## How each tree is compared
|
|
64
|
+
|
|
65
|
+
| Tree | Mode | Why |
|
|
66
|
+
|---|---|---|
|
|
67
|
+
| `scripts` | bytes | verbatim copy, minus the dev-only set |
|
|
68
|
+
| `lib` | bytes | verbatim copy |
|
|
69
|
+
| `multi-agent-refs` | bytes | verbatim copy |
|
|
70
|
+
| `agents` | bytes | verbatim copy |
|
|
71
|
+
| `commands/multi-agent` | presence | `install.js` rewrites each SKILL.md `description` into the user's `outputLanguage` |
|
|
72
|
+
|
|
73
|
+
Byte-comparing `commands/` reports every command as drift on a perfectly
|
|
74
|
+
healthy machine. Measured here: all 57 command files differ, and 56 of them
|
|
75
|
+
differ by nothing except the translated description. A report that is wrong by
|
|
76
|
+
default is a report nobody reads.
|
|
77
|
+
|
|
78
|
+
The dev-only filter matters just as much: smokes, linters and fixtures ship in
|
|
79
|
+
the package and are deliberately NOT installed. Without excluding them, `verify`
|
|
80
|
+
would report 252 files as "the installer skipped this".
|
|
81
|
+
|
|
82
|
+
`~/.copilot` and `~/.codex` are reported as present, not compared: the installer
|
|
83
|
+
rewrites paths for both on purpose, so a byte difference there is the design.
|
|
@@ -78,8 +78,9 @@ Update `agent-state.json` at EVERY phase transition.
|
|
|
78
78
|
|
|
79
79
|
### Writing `agent-state.json` (required mechanism)
|
|
80
80
|
|
|
81
|
-
Every state
|
|
82
|
-
plain read-modify-write (`jq ... > tmp &&
|
|
81
|
+
Every state write goes through `write-state.mjs`, **including the first one in
|
|
82
|
+
Phase 0**. Never write the file with a plain read-modify-write (`jq ... > tmp &&
|
|
83
|
+
mv`, an editor tool, `cat >`):
|
|
83
84
|
|
|
84
85
|
```bash
|
|
85
86
|
# Merge a patch into the current state (the normal case).
|
|
@@ -97,6 +98,16 @@ the first landed. `write-state.mjs` does tmpfile + rename (atomic on POSIX) unde
|
|
|
97
98
|
an advisory `.lock`, reclaims a lock whose holder PID is dead, and releases the
|
|
98
99
|
lock on every error path.
|
|
99
100
|
|
|
101
|
+
Why the CREATE matters as much as the updates: the writer stamps `rev` on every
|
|
102
|
+
write and `schemaVersion` on the first one. A document written by hand in Phase 0
|
|
103
|
+
starts with neither, so every later writer compares against an absent revision
|
|
104
|
+
and `migrate-state.mjs` can never place the file on a migration path. Measured on
|
|
105
|
+
a real install before this was fixed: 39 of 43 `agent-state.json` files carried no
|
|
106
|
+
`rev` and 43 of 43 carried no `schemaVersion`, which is the whole of
|
|
107
|
+
`$HOME/.claude/schemas/migrations/` sitting unreachable. `migrate-state.mjs --all`
|
|
108
|
+
reports the legacy ones; it does not stamp them, because a stamp would assert a
|
|
109
|
+
conformance nothing checked.
|
|
110
|
+
|
|
100
111
|
Exit codes the caller must handle: `0` written, `1` invalid JSON on stdin, `2`
|
|
101
112
|
lock timeout (another writer held it past the acquire window - retry once, then
|
|
102
113
|
halt per the halt-visibility rule), `3` I/O error.
|