@mmerterden/multi-agent-pipeline 17.0.0 → 17.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +32 -0
- package/README.md +49 -4
- package/README.tr.md +50 -4
- package/docs/architecture.md +3 -3
- package/docs/ecosystem.md +5 -5
- package/install/templates/multi-agent-autopilot.plist.template +79 -0
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +64 -0
- package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +173 -0
- package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/channels/SKILL.md +41 -12
- package/pipeline/commands/multi-agent/help/SKILL.md +41 -35
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/sync/SKILL.md +10 -9
- package/pipeline/commands/multi-agent/update/SKILL.md +1 -1
- package/pipeline/lib/autopilot-activation.sh +117 -0
- package/pipeline/lib/autopilot-state.sh +150 -0
- package/pipeline/lib/issue-fetcher.sh +18 -1
- package/pipeline/lib/plan-todos.sh +18 -0
- package/pipeline/multi-agent-refs/channels/jira.md +80 -20
- package/pipeline/multi-agent-refs/channels/pr.md +65 -19
- package/pipeline/multi-agent-refs/cross-cli-contract.md +6 -5
- package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -7
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +17 -15
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +1 -1
- package/pipeline/multi-agent-refs/readiness-review.md +7 -1
- package/pipeline/multi-agent-refs/rules.md +3 -11
- package/pipeline/multi-agent-refs/tracker-contract.md +32 -0
- package/pipeline/schemas/autopilot-config.schema.json +149 -0
- package/pipeline/schemas/token-budget.json +2 -2
- package/pipeline/scripts/autopilot-arming.mjs +147 -0
- package/pipeline/scripts/autopilot-intake.mjs +383 -0
- package/pipeline/scripts/autopilot-menubar.swift +361 -0
- package/pipeline/scripts/autopilot-runner.mjs +349 -0
- package/pipeline/scripts/autopilot-status.sh +212 -0
- package/pipeline/scripts/jira-search.sh +70 -0
- package/pipeline/scripts/phase-tracker.sh +134 -12
- package/pipeline/scripts/probe-evidence-capability.sh +27 -3
- package/pipeline/scripts/run-ui-tests.sh +113 -4
- package/pipeline/skills/.skill-manifest.json +16 -4
- package/pipeline/skills/shared/core/multi-agent-autopilot-off/SKILL.md +67 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-on/SKILL.md +146 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +64 -0
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +62 -11
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +9 -8
|
@@ -0,0 +1,117 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# autopilot-activation.sh - which repos may this machine pick work up from.
|
|
3
|
+
#
|
|
4
|
+
# The answer is NEVER "all of them" and never "whatever has the label". Measured
|
|
5
|
+
# on one machine: 66 repos grant push, and a good half belong to other people -
|
|
6
|
+
# a colleague's backend, someone's side project. A label is a filter, not a gate:
|
|
7
|
+
# anyone able to open an issue in a repo you happen to have push on could put
|
|
8
|
+
# `agent-queue` on it. The gate is this picker, and the picker is per machine.
|
|
9
|
+
#
|
|
10
|
+
# Two lists come out, and the difference matters:
|
|
11
|
+
#
|
|
12
|
+
# eligible push rights AND a checkout on this machine
|
|
13
|
+
# unavailable push rights, no checkout - LISTED, with that as the reason
|
|
14
|
+
#
|
|
15
|
+
# The second list is printed rather than dropped. Silently hiding 40 repos from a
|
|
16
|
+
# 66-repo answer looks exactly like a permissions problem, and the user then goes
|
|
17
|
+
# looking for a token fault that does not exist.
|
|
18
|
+
#
|
|
19
|
+
# Usage:
|
|
20
|
+
# bash autopilot-activation.sh list # TSV: state, nameWithOwner, localPath
|
|
21
|
+
# bash autopilot-activation.sh list --json
|
|
22
|
+
# . autopilot-activation.sh; ma_ap_local_path <owner/repo>
|
|
23
|
+
#
|
|
24
|
+
# Exit: 0 on success (including an empty list), 2 on usage, 3 when gh cannot answer.
|
|
25
|
+
|
|
26
|
+
set -uo pipefail
|
|
27
|
+
|
|
28
|
+
# Repo roots to search for checkouts. $HOME's immediate children covers the
|
|
29
|
+
# normal layout; MA_AUTOPILOT_REPO_ROOTS overrides it for tests and for anyone
|
|
30
|
+
# who keeps checkouts somewhere else.
|
|
31
|
+
MA_AP_ROOTS="${MA_AUTOPILOT_REPO_ROOTS:-$HOME}"
|
|
32
|
+
|
|
33
|
+
# owner/repo for a checkout, read from its origin remote. Both URL forms, and a
|
|
34
|
+
# trailing .git is stripped - `git@github.com:o/r.git` and
|
|
35
|
+
# `https://github.com/o/r` are the same repo and a picker that treats them as
|
|
36
|
+
# two shows the user a duplicate they cannot explain.
|
|
37
|
+
ma_ap_remote_slug() {
|
|
38
|
+
local dir="$1" url
|
|
39
|
+
url=$(git -C "$dir" remote get-url origin 2>/dev/null) || return 1
|
|
40
|
+
url="${url%.git}"
|
|
41
|
+
case "$url" in
|
|
42
|
+
*github.com[:/]*) printf '%s\n' "${url##*github.com[:/]}" ;;
|
|
43
|
+
*) return 1 ;;
|
|
44
|
+
esac
|
|
45
|
+
}
|
|
46
|
+
|
|
47
|
+
# Every checkout under the roots, as `slug<TAB>path`. One pass, so the lookup
|
|
48
|
+
# below is a grep rather than a walk per repo.
|
|
49
|
+
ma_ap_checkout_index() {
|
|
50
|
+
local root d slug
|
|
51
|
+
for root in $MA_AP_ROOTS; do
|
|
52
|
+
[ -d "$root" ] || continue
|
|
53
|
+
for d in "$root"/*/; do
|
|
54
|
+
[ -d "$d/.git" ] || continue
|
|
55
|
+
slug=$(ma_ap_remote_slug "${d%/}") || continue
|
|
56
|
+
printf '%s\t%s\n' "$slug" "${d%/}"
|
|
57
|
+
done
|
|
58
|
+
done
|
|
59
|
+
}
|
|
60
|
+
|
|
61
|
+
ma_ap_local_path() { # $1 = owner/repo
|
|
62
|
+
ma_ap_checkout_index | awk -F'\t' -v s="$1" 'tolower($1)==tolower(s){print $2; exit}'
|
|
63
|
+
}
|
|
64
|
+
|
|
65
|
+
# Repos the authenticated account may push to.
|
|
66
|
+
#
|
|
67
|
+
# The query string form is deliberate: `gh api user/repos --field affiliation=...`
|
|
68
|
+
# sends the value as a BODY field on a GET, which this endpoint ignores - it
|
|
69
|
+
# answers with the default affiliation and the caller never learns the filter did
|
|
70
|
+
# not apply. Measured: the --field form returned 0 rows here, the query string
|
|
71
|
+
# form returned 66.
|
|
72
|
+
ma_ap_writable_repos() {
|
|
73
|
+
gh api "user/repos?affiliation=owner,collaborator,organization_member&per_page=100" \
|
|
74
|
+
--paginate -q '.[] | select(.permissions.push == true) | .full_name' 2>/dev/null
|
|
75
|
+
}
|
|
76
|
+
|
|
77
|
+
ma_ap_list() {
|
|
78
|
+
local json="${1:-}" slug path index
|
|
79
|
+
index=$(ma_ap_checkout_index)
|
|
80
|
+
local repos
|
|
81
|
+
repos=$(ma_ap_writable_repos)
|
|
82
|
+
if [ -z "$repos" ]; then
|
|
83
|
+
echo "autopilot-activation: gh returned no writable repos - check 'gh auth status'" >&2
|
|
84
|
+
return 3
|
|
85
|
+
fi
|
|
86
|
+
|
|
87
|
+
local first=1
|
|
88
|
+
[ "$json" = "--json" ] && printf '['
|
|
89
|
+
while IFS= read -r slug; do
|
|
90
|
+
[ -n "$slug" ] || continue
|
|
91
|
+
path=$(printf '%s\n' "$index" | awk -F'\t' -v s="$slug" 'tolower($1)==tolower(s){print $2; exit}')
|
|
92
|
+
local state="eligible"
|
|
93
|
+
[ -z "$path" ] && state="unavailable"
|
|
94
|
+
if [ "$json" = "--json" ]; then
|
|
95
|
+
[ "$first" -eq 0 ] && printf ','
|
|
96
|
+
first=0
|
|
97
|
+
printf '{"state":"%s","nameWithOwner":"%s","localPath":"%s"}' "$state" "$slug" "$path"
|
|
98
|
+
else
|
|
99
|
+
printf '%s\t%s\t%s\n' "$state" "$slug" "$path"
|
|
100
|
+
fi
|
|
101
|
+
done <<EOF
|
|
102
|
+
$repos
|
|
103
|
+
EOF
|
|
104
|
+
[ "$json" = "--json" ] && printf ']\n'
|
|
105
|
+
return 0
|
|
106
|
+
}
|
|
107
|
+
|
|
108
|
+
if [ "${BASH_SOURCE[0]:-$0}" = "$0" ]; then
|
|
109
|
+
case "${1:-}" in
|
|
110
|
+
list) ma_ap_list "${2:-}" ;;
|
|
111
|
+
"" | -h | --help) grep -E '^#( |$)' "$0" | sed -E 's/^# ?//' ;;
|
|
112
|
+
*)
|
|
113
|
+
echo "autopilot-activation: unknown command ${1}" >&2
|
|
114
|
+
exit 2
|
|
115
|
+
;;
|
|
116
|
+
esac
|
|
117
|
+
fi
|
|
@@ -0,0 +1,150 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# autopilot-state.sh - where continuous mode keeps its state, and who may read it.
|
|
3
|
+
#
|
|
4
|
+
# Everything lives under ~/.claude/autopilot/ and NOT in
|
|
5
|
+
# multi-agent-preferences.json. That file has `additionalProperties: false` and a
|
|
6
|
+
# migration chain, so a block there would push keys into every local user's
|
|
7
|
+
# preferences forever - including everyone who never turns this on. Separation
|
|
8
|
+
# also gives the off state a definition that cannot be got wrong: the directory
|
|
9
|
+
# is absent.
|
|
10
|
+
#
|
|
11
|
+
# config.json the repo selection. Survives autopilot-off on purpose, so
|
|
12
|
+
# turning the mode back on does not re-ask which repos.
|
|
13
|
+
# queue.json the current ordered list plus whatever is in flight
|
|
14
|
+
# status.json what the menu bar renders. Written by the runner, read-only
|
|
15
|
+
# to everything else
|
|
16
|
+
# attempted.jsonl append-only: {source,id,taskId,outcome,prUrl,at}
|
|
17
|
+
# runner.pid pid + the in-flight sessionId + kern.boottime
|
|
18
|
+
# runner.log
|
|
19
|
+
# bin/menubar built on demand from scripts/autopilot-menubar.swift
|
|
20
|
+
#
|
|
21
|
+
# 0700 throughout. The queue names real tickets and real repos, and on a shared
|
|
22
|
+
# machine that is somebody's roadmap.
|
|
23
|
+
#
|
|
24
|
+
# Usage:
|
|
25
|
+
# . autopilot-state.sh
|
|
26
|
+
# ma_ap_root # prints the dir, creating it 0700 only when asked
|
|
27
|
+
# ma_ap_is_on # exit 0 when configured, 1 otherwise. NO side effects
|
|
28
|
+
# ma_ap_write <file> <- # atomic write from stdin, 0600
|
|
29
|
+
|
|
30
|
+
set -uo pipefail
|
|
31
|
+
|
|
32
|
+
MA_AP_ROOT="${MA_AUTOPILOT_ROOT:-$HOME/.claude/autopilot}"
|
|
33
|
+
|
|
34
|
+
ma_ap_root() { printf '%s\n' "$MA_AP_ROOT"; }
|
|
35
|
+
|
|
36
|
+
MA_AP_LABEL="${MA_AUTOPILOT_LABEL:-com.multi-agent.autopilot}"
|
|
37
|
+
MA_AP_PLIST="${MA_AUTOPILOT_PLIST:-$HOME/Library/LaunchAgents/$MA_AP_LABEL.plist}"
|
|
38
|
+
|
|
39
|
+
# Two different questions, and conflating them was a real bug in the first draft
|
|
40
|
+
# of this file: `autopilot-off` deliberately KEEPS config.json so turning the
|
|
41
|
+
# mode back on does not re-ask which repos, so a predicate reading the config
|
|
42
|
+
# would still answer "on" after you turned it off.
|
|
43
|
+
#
|
|
44
|
+
# configured = a repo selection exists
|
|
45
|
+
# on = launchd holds the job
|
|
46
|
+
#
|
|
47
|
+
# On is the launchd job because that is the thing that actually makes work
|
|
48
|
+
# happen; anything else is a claim about a file.
|
|
49
|
+
ma_ap_is_configured() { [ -f "$MA_AP_ROOT/config.json" ]; }
|
|
50
|
+
|
|
51
|
+
ma_ap_is_on() { [ -f "$MA_AP_PLIST" ] && launchctl list 2>/dev/null | grep -q "$MA_AP_LABEL"; }
|
|
52
|
+
|
|
53
|
+
# The failure mode a `resume` command would have papered over: configured and
|
|
54
|
+
# meant to be running, but launchd does not hold the job - an OS update dropped
|
|
55
|
+
# the plist, or it was booted out by hand. doctor reports this; there is no
|
|
56
|
+
# command to remember.
|
|
57
|
+
ma_ap_is_orphaned() { ma_ap_is_configured && ! ma_ap_is_on && [ -f "$MA_AP_PLIST" ]; }
|
|
58
|
+
|
|
59
|
+
# Deliberately no side effects in any of the three. A predicate that creates its
|
|
60
|
+
# own directory turns "is autopilot on?" into "autopilot is now half on", and
|
|
61
|
+
# every status command, hook and doctor check calls these.
|
|
62
|
+
|
|
63
|
+
ma_ap_ensure_root() {
|
|
64
|
+
[ -d "$MA_AP_ROOT" ] || mkdir -p "$MA_AP_ROOT" || return 1
|
|
65
|
+
chmod 700 "$MA_AP_ROOT" 2>/dev/null
|
|
66
|
+
[ -d "$MA_AP_ROOT/bin" ] || mkdir -p "$MA_AP_ROOT/bin" 2>/dev/null
|
|
67
|
+
chmod 700 "$MA_AP_ROOT/bin" 2>/dev/null
|
|
68
|
+
return 0
|
|
69
|
+
}
|
|
70
|
+
|
|
71
|
+
# Write via a temp file in the SAME directory then rename. The menu bar polls
|
|
72
|
+
# status.json every few seconds, and a reader that catches a half-written file
|
|
73
|
+
# shows an empty menu; rename is atomic on the same filesystem, so it never sees
|
|
74
|
+
# a partial one.
|
|
75
|
+
ma_ap_write() { # $1 = filename under the root; content on stdin
|
|
76
|
+
ma_ap_ensure_root || return 1
|
|
77
|
+
local dst="$MA_AP_ROOT/$1" tmp="$MA_AP_ROOT/.$1.$$"
|
|
78
|
+
cat > "$tmp" || {
|
|
79
|
+
rm -f "$tmp"
|
|
80
|
+
return 1
|
|
81
|
+
}
|
|
82
|
+
chmod 600 "$tmp" 2>/dev/null
|
|
83
|
+
mv -f "$tmp" "$dst"
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
ma_ap_append() { # $1 = filename, content on stdin - for the jsonl
|
|
87
|
+
ma_ap_ensure_root || return 1
|
|
88
|
+
local dst="$MA_AP_ROOT/$1"
|
|
89
|
+
cat >> "$dst" || return 1
|
|
90
|
+
chmod 600 "$dst" 2>/dev/null
|
|
91
|
+
}
|
|
92
|
+
|
|
93
|
+
ma_ap_read() { # $1 = filename; empty and exit 1 when absent
|
|
94
|
+
local src="$MA_AP_ROOT/$1"
|
|
95
|
+
[ -f "$src" ] || return 1
|
|
96
|
+
cat "$src"
|
|
97
|
+
}
|
|
98
|
+
|
|
99
|
+
# jq with a default, so a caller never has to distinguish "key absent" from
|
|
100
|
+
# "file absent" from "file unparseable" - all three mean "use the default".
|
|
101
|
+
ma_ap_cfg() { # $1 = jq path, $2 = default
|
|
102
|
+
local v
|
|
103
|
+
v=$(ma_ap_read config.json 2>/dev/null | jq -r "$1 // empty" 2>/dev/null)
|
|
104
|
+
[ -n "$v" ] && printf '%s\n' "$v" || printf '%s\n' "$2"
|
|
105
|
+
}
|
|
106
|
+
|
|
107
|
+
# AC or battery. The mode holds a sleep assertion only on AC: a queue that keeps
|
|
108
|
+
# a laptop awake on battery is a bug, and `StartInterval` does not wake a
|
|
109
|
+
# sleeping Mac anyway, so on battery the work resumes when you plug in.
|
|
110
|
+
ma_ap_power() {
|
|
111
|
+
case "$(pmset -g ps 2>/dev/null | head -1)" in
|
|
112
|
+
*"AC Power"*) printf 'ac\n' ;;
|
|
113
|
+
*) printf 'battery\n' ;;
|
|
114
|
+
esac
|
|
115
|
+
}
|
|
116
|
+
|
|
117
|
+
# Boot time, so a stale runner.pid cannot be mistaken for a live runner. After a
|
|
118
|
+
# restart pids start low and the recorded 4711 may belong to something unrelated;
|
|
119
|
+
# a pid recorded BEFORE the current boot is stale by definition, with no probing.
|
|
120
|
+
# The pattern is ANCHORED on purpose. `kern.boottime` prints
|
|
121
|
+
# `{ sec = 1788181644, usec = 618574 } Mon Aug 31 ...`, and `.*sec = ` is greedy:
|
|
122
|
+
# it walks past `sec` to `usec` and captures 618574, the microseconds. The
|
|
123
|
+
# original form here did exactly that, so this returned a six-digit number that
|
|
124
|
+
# looked plausible and never matched the same fact read anywhere else - which is
|
|
125
|
+
# how the runner's staleness check compared two different numbers and called a
|
|
126
|
+
# live runner dead.
|
|
127
|
+
ma_ap_boottime() {
|
|
128
|
+
sysctl -n kern.boottime 2>/dev/null | sed -n 's/^{ *sec = \([0-9]*\).*/\1/p'
|
|
129
|
+
}
|
|
130
|
+
|
|
131
|
+
# The machine's honest ceiling, so raising `slots` is a measured decision. RAM
|
|
132
|
+
# and cores bound the concurrent Claude sessions; disk bounds the worktrees
|
|
133
|
+
# (0.75 GB each, measured at 11 GB across 15). The hard cap is 4 because beyond
|
|
134
|
+
# that the bottleneck stops being this machine.
|
|
135
|
+
ma_ap_slot_ceiling() {
|
|
136
|
+
local ram_gb cores free_gb a b c
|
|
137
|
+
ram_gb=$(( $(sysctl -n hw.memsize 2>/dev/null || echo 0) / 1073741824 ))
|
|
138
|
+
cores=$(sysctl -n hw.ncpu 2>/dev/null || echo 2)
|
|
139
|
+
free_gb=$(df -g "$HOME" 2>/dev/null | awk 'NR==2{print $4}')
|
|
140
|
+
[ -n "$free_gb" ] || free_gb=0
|
|
141
|
+
a=$(((ram_gb - 4) * 10 / 25))
|
|
142
|
+
b=$((cores / 2))
|
|
143
|
+
c=$(((free_gb - 20) * 100 / 75))
|
|
144
|
+
local m=$a
|
|
145
|
+
[ "$b" -lt "$m" ] && m=$b
|
|
146
|
+
[ "$c" -lt "$m" ] && m=$c
|
|
147
|
+
[ "$m" -gt 4 ] && m=4
|
|
148
|
+
[ "$m" -lt 1 ] && m=1
|
|
149
|
+
printf '%s\n' "$m"
|
|
150
|
+
}
|
|
@@ -49,8 +49,23 @@
|
|
|
49
49
|
|
|
50
50
|
set -euo pipefail
|
|
51
51
|
|
|
52
|
+
# Sourcing this file hands a caller the provider helpers - `jira_search`,
|
|
53
|
+
# `fetch_jira`, `fetch_github` and the `jira_curl` underneath them - without
|
|
54
|
+
# resolving anything. Executing it resolves one input, exactly as before.
|
|
55
|
+
#
|
|
56
|
+
# The guard exists because the alternative is worse than it looks: a second
|
|
57
|
+
# consumer of `jira_search` with no way to source would have to copy
|
|
58
|
+
# `jira_curl`, and that function is not a URL - it is the credential resolution,
|
|
59
|
+
# the host lookup and the rule that keeps a token off argv. Three copies of that
|
|
60
|
+
# is three places for a token to leak.
|
|
61
|
+
MA_IF_EXECUTED=0
|
|
62
|
+
[ "${BASH_SOURCE[0]}" = "$0" ] && MA_IF_EXECUTED=1
|
|
63
|
+
|
|
52
64
|
INPUT="${1:-}"
|
|
53
|
-
[
|
|
65
|
+
if [ "$MA_IF_EXECUTED" = 1 ] && [ -z "$INPUT" ]; then
|
|
66
|
+
echo '{"error":"input required"}'
|
|
67
|
+
exit 1
|
|
68
|
+
fi
|
|
54
69
|
|
|
55
70
|
JIRA_TOKEN_KEY="${ACCOUNT_JIRA_TOKEN_KEY:-}"
|
|
56
71
|
JIRA_HOST="${ACCOUNT_JIRA_HOST:-}"
|
|
@@ -361,6 +376,7 @@ PY
|
|
|
361
376
|
}
|
|
362
377
|
|
|
363
378
|
# --- Dispatch -----------------------------------------------------------------
|
|
379
|
+
if [ "$MA_IF_EXECUTED" = 1 ]; then
|
|
364
380
|
case "$KIND" in
|
|
365
381
|
jira-id)
|
|
366
382
|
KEY="$INPUT"
|
|
@@ -592,3 +608,4 @@ sys.stdout.write("\x1f".join(parts))
|
|
|
592
608
|
"description=" "branchHint=$branch"
|
|
593
609
|
;;
|
|
594
610
|
esac
|
|
611
|
+
fi
|
|
@@ -94,6 +94,24 @@ do_set() {
|
|
|
94
94
|
if [ -z "$plan_blob" ] || [ "$plan_blob" = "-" ]; then
|
|
95
95
|
plan_blob=$(cat)
|
|
96
96
|
fi
|
|
97
|
+
# A planning-output document is accepted directly and converted here.
|
|
98
|
+
#
|
|
99
|
+
# The conversion used to live as a jq blob inside phase-2-planning.md, which
|
|
100
|
+
# made the mapping from `tasks[]` to `todos[]` a thing two files defined - and
|
|
101
|
+
# the phase doc was the copy nothing tested. Accepting both shapes costs four
|
|
102
|
+
# lines and removes the second definition.
|
|
103
|
+
if jq -e '.tasks and (.todos | not)' <<<"$plan_blob" >/dev/null 2>&1; then
|
|
104
|
+
plan_blob=$(jq '{
|
|
105
|
+
title: (.summary // .title // "plan"),
|
|
106
|
+
todos: [ .tasks[] | {
|
|
107
|
+
id: .id,
|
|
108
|
+
task: (.title // .subject // ""),
|
|
109
|
+
status: "pending",
|
|
110
|
+
deps: (.dependsOn // .blockedBy // [])
|
|
111
|
+
} ]
|
|
112
|
+
}' <<<"$plan_blob")
|
|
113
|
+
fi
|
|
114
|
+
|
|
97
115
|
# Validate against schema (best-effort - jq syntax check, then required-field probe).
|
|
98
116
|
if ! jq -e '.title and (.todos | type == "array")' <<<"$plan_blob" >/dev/null 2>&1; then
|
|
99
117
|
echo "plan-todos: input must be an object with .title (string) and .todos (array)" >&2
|
|
@@ -10,35 +10,89 @@ Every Jira comment posted by this adapter follows the same section order. Sectio
|
|
|
10
10
|
|
|
11
11
|
| # | Section key | Heading (`tr`) | Heading (`en`) | Required? |
|
|
12
12
|
|---|---|---|---|---|
|
|
13
|
-
| 1 | `summary` | `##
|
|
13
|
+
| 1 | `summary` | `## Geliştirme Özeti` | `## Development Summary` | always |
|
|
14
14
|
| 2 | `test_scenarios` | `## Test Senaryoları` | `## Test Scenarios` | always (use " - " placeholder line if truly N/A) |
|
|
15
|
-
| 3 | `
|
|
15
|
+
| 3 | `impact` | `## Etki Analizi` | `## Impact Analysis` | always |
|
|
16
|
+
| 4 | `context_refs` | `## Bağlantılar` | `## References` | when any link exists |
|
|
16
17
|
|
|
17
18
|
> Heading levels in the source markdown are `##`. The wiki-markup converter (next section) rewrites them to `h2.` for Jira rendering.
|
|
18
19
|
|
|
20
|
+
**Jira is not the PR, and this is the contract that keeps getting that wrong.**
|
|
21
|
+
The PR is read by a reviewer holding the diff. The Jira comment is read by the
|
|
22
|
+
person who filed the ticket and by the tester who has to verify it, and neither
|
|
23
|
+
of them has the diff open. So a Jira body carries no identifiers, no file paths,
|
|
24
|
+
no stack frames, no diff hunks and no framework names. Counts in prose are fine
|
|
25
|
+
and are often the most useful sentence in the comment ("two files, eight lines
|
|
26
|
+
removed, no additions"); a class name is not. A sentence that needs a symbol to
|
|
27
|
+
make its point is describing the change at the wrong level for this reader -
|
|
28
|
+
restate the behaviour, not the mechanism. The technical account has a home, and
|
|
29
|
+
it is the PR body (`channels/pr.md`).
|
|
30
|
+
|
|
19
31
|
### Section content rules
|
|
20
32
|
|
|
21
33
|
**`summary`** - 2-5 sentences in `outputLanguage`. What changed, why, and the user-visible impact. No "we", no marketing tone. Past tense (the work is done at the time the comment goes up).
|
|
22
34
|
|
|
23
|
-
**Visual evidence inside these
|
|
35
|
+
**Visual evidence inside these sections.** When `state.visualEvidence` carries artefacts, they render INSIDE `summary` and `test_scenarios` - never as a section of their own, which the fixed section order forbids. Phase 6 Step 2.9 has already uploaded them and written the returned name to `visualEvidence.*[].jiraFilename`: reference that name, and call `jira-attach.sh <issue> <file>...` only for an artefact whose `jiraFilename` is absent. Uploading unconditionally here attaches every file a second time whenever the comment is re-rendered - which is exactly what a post-hoc `/multi-agent:channels` run does.
|
|
24
36
|
|
|
25
37
|
- `summary`, after its sentences: one line naming the pair in `outputLanguage` (`Düzeltme öncesi / Düzeltme sonrası`), then the thumbnails on the next line - `!<file>-before.png|thumbnail! !<file>-after.png|thumbnail!`.
|
|
26
38
|
- `test_scenarios`, under the scenario the recording demonstrates: `!<file>-flow.mp4!` plus one line stating the tier used.
|
|
27
39
|
|
|
28
40
|
A `gaps[]` entry prints its reason on the line where the artefact would have been (`Düzeltme öncesi: ticket'ta görsel yok`). Never an empty thumbnail, never a silent omission. Contract: `$HOME/.claude/multi-agent-refs/features/visual-evidence.md`.
|
|
29
41
|
|
|
30
|
-
**`test_scenarios`** -
|
|
42
|
+
**`test_scenarios`** - one titled scenario per acceptance criterion, each a
|
|
43
|
+
numbered list of steps ending in the expected result. Given/When/Then was the
|
|
44
|
+
shape here for several releases and it reads as translated English to the person
|
|
45
|
+
who actually runs these; a tester wants a list they can follow with the app open.
|
|
46
|
+
The heading and the text render in `outputLanguage`; this file shows the English
|
|
47
|
+
skeleton:
|
|
31
48
|
|
|
32
49
|
```markdown
|
|
33
50
|
## Test Scenarios
|
|
34
51
|
|
|
35
|
-
1.
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
52
|
+
**1. <what this scenario exercises>**
|
|
53
|
+
1. <step the tester performs, in the app's own words>
|
|
54
|
+
2. <step>
|
|
55
|
+
3. **Expected:** <what they should see>
|
|
56
|
+
|
|
57
|
+
**2. <regression scenario>**
|
|
58
|
+
1. <step>
|
|
59
|
+
2. **Expected:** <what should still behave as before>
|
|
39
60
|
```
|
|
40
61
|
|
|
41
|
-
|
|
62
|
+
Order is load-bearing. The scenarios that reproduce the reported behaviour come
|
|
63
|
+
first, then the regression scenarios for whatever the change could have disturbed -
|
|
64
|
+
anything sharing the component that was touched. A comment that lists only the fix
|
|
65
|
+
leaves the tester to guess the blast radius, and that guess is the one thing they
|
|
66
|
+
cannot make from the ticket.
|
|
67
|
+
|
|
68
|
+
Screens, fields and menu items are named the way the app shows them, in
|
|
69
|
+
`outputLanguage`. Function names and file paths do not appear at all, per the rule
|
|
70
|
+
above. When there are no executable scenarios (pure doc change, dependency bump),
|
|
71
|
+
write a single line: `- No manual test required - <short reason>`.
|
|
72
|
+
|
|
73
|
+
**`impact`** - four fixed numbered parts, each answered, never left as a
|
|
74
|
+
placeholder. This is what a test lead reads before deciding how wide to test and
|
|
75
|
+
what a release manager reads before deciding whether it ships this week, so it is
|
|
76
|
+
written in the same plain register as the rest of the comment:
|
|
77
|
+
|
|
78
|
+
```markdown
|
|
79
|
+
## Impact Analysis
|
|
80
|
+
|
|
81
|
+
**1 - The problem**
|
|
82
|
+
<what was wrong and since when, in the reader's words>
|
|
83
|
+
|
|
84
|
+
**2 - What was changed**
|
|
85
|
+
<what the change does, and why the behaviour around it is unchanged>
|
|
86
|
+
|
|
87
|
+
**3 - Affected areas**
|
|
88
|
+
<the screens and flows that must be tested, including anything sharing the changed component>
|
|
89
|
+
|
|
90
|
+
**4 - Effect on other systems**
|
|
91
|
+
<none, or which service, contract or channel is affected>
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
Part 4 is answered "none" far more often than it is answered at all, and "none" is
|
|
95
|
+
a real answer worth writing: it is what lets the reader stop looking.
|
|
42
96
|
|
|
43
97
|
**`context_refs`** - flat bullet list of external links the work depends on or produced. PR URL is **not** repeated here; it lives on the first line (see *Cross-link injection*). Common entries (only when present in the task):
|
|
44
98
|
|
|
@@ -58,7 +112,7 @@ The links are pulled from `agent-state.json.contextLinks[]` (see Phase 0 link ex
|
|
|
58
112
|
|
|
59
113
|
```
|
|
60
114
|
1. Read agent-state.json (taskId, contextLinks, prUrls, language).
|
|
61
|
-
2. Build section bodies in markdown
|
|
115
|
+
2. Build section bodies in markdown, in table order: summary, test_scenarios, impact, context_refs.
|
|
62
116
|
3. Run the assembled body through the `humanizer` skill (see Hard rules below).
|
|
63
117
|
4. Apply Cross-link injection (PR URL on line 1).
|
|
64
118
|
5. Run Wiki markup conversion.
|
|
@@ -116,7 +170,7 @@ Do not hand-apply this table. It was hand-applied for several releases and a smi
|
|
|
116
170
|
|
|
117
171
|
Lines outside these patterns pass through verbatim. Multi-paragraph blocks are joined with one blank line.
|
|
118
172
|
|
|
119
|
-
Every row above is load-bearing, including the ones that look cosmetic. The table must cover each construct the section templates in this file actually emit - `##` for the
|
|
173
|
+
Every row above is load-bearing, including the ones that look cosmetic. The table must cover each construct the section templates in this file actually emit - `##` for the four required headings, `**1. title**` on a scenario and `**Expected:**` inside it, the `**1 - ...**` part labels in the impact section, and `1.`-numbered steps. A missing row does not degrade gracefully: the "pass through verbatim" fallback POSTs `## Test Senaryoları` and `**Expected:**` as literal text, so the comment renders with visible `##` and stray asterisks. Single `*bold*` in Markdown means *italic* in Jira wiki - never map `**bold**` to `*bold*` by dropping one asterisk mechanically without checking the source was bold, not italic.
|
|
120
174
|
|
|
121
175
|
There is no markdown→Jira-wiki converter program in the pipeline (`lib/` ships `md2confluence-v3.py` for Confluence only, and `scripts/jira-wiki-escape.mjs` covers the emoticon step alone). This conversion table is applied by the model, by hand, which is exactly why it has to be complete.
|
|
122
176
|
|
|
@@ -149,19 +203,25 @@ Authorization: Bearer $JIRA_TOKEN
|
|
|
149
203
|
Content-Type: application/json
|
|
150
204
|
```
|
|
151
205
|
|
|
152
|
-
|
|
206
|
+
**Post through `jira-publish.sh`. Never hand-roll the `curl`.**
|
|
153
207
|
|
|
154
208
|
```bash
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
jq -n --rawfile body /tmp/channels-$TASK_ID-jira-escaped.txt '{body: $body}' \
|
|
158
|
-
> /tmp/channels-$TASK_ID-jira-payload.json
|
|
159
|
-
curl -s -X POST -H "Authorization: Bearer $JIRA_TOKEN" \
|
|
160
|
-
-H "Content-Type: application/json" \
|
|
161
|
-
--data-binary @/tmp/channels-$TASK_ID-jira-payload.json \
|
|
162
|
-
"$JIRA_BASE/rest/api/2/issue/$JIRA_ID/comment"
|
|
209
|
+
bash "$HOME/.claude/lib/jira-publish.sh" --issue "$JIRA_ID" \
|
|
210
|
+
--body-file /tmp/channels-$TASK_ID-jira.txt --target comment
|
|
163
211
|
```
|
|
164
212
|
|
|
213
|
+
That script escapes every body unconditionally (`lib/jira-publish.sh:109`) and
|
|
214
|
+
passes the token through a `-K` config rather than argv, so neither the escape
|
|
215
|
+
nor the token handling depends on an agent remembering a step.
|
|
216
|
+
|
|
217
|
+
This block used to be a hand-written `jq` + `curl` pair with the escape as a
|
|
218
|
+
separate `--check` line above it, and a shipped comment rendered the Swift
|
|
219
|
+
selector `hash(into:)` as `hash(into` plus a smiley - twice. The escaper was
|
|
220
|
+
present, correct, and simply not run on that path: `:)` reached Jira intact and
|
|
221
|
+
Jira's wiki renderer turned it into an emoticon, which is what section
|
|
222
|
+
*Emoticon escaping* below is about. A defence that has to be invoked by hand is
|
|
223
|
+
a defence that is eventually not invoked.
|
|
224
|
+
|
|
165
225
|
The dispatch summary line includes the comment URL with `?focusedCommentId=...` so the user can paste it into Slack/Teams.
|
|
166
226
|
|
|
167
227
|
### Writing the issue description (not a comment)
|
|
@@ -10,20 +10,31 @@ The PR description targets code reviewers - it stays technical. Every adapter
|
|
|
10
10
|
|
|
11
11
|
| # | Section key | Heading (`tr`) | Heading (`en`) | Required? |
|
|
12
12
|
|---|---|---|---|---|
|
|
13
|
-
| 1 | `summary` | `##
|
|
14
|
-
| 2 | `
|
|
13
|
+
| 1 | `summary` | `## Geliştirme Özeti` | `## Development Summary` | always |
|
|
14
|
+
| 2 | `technical` | `## Teknik Açıklama` | `## Technical Explanation` | always |
|
|
15
15
|
| 3 | `architecture` | `## Mimari Kararlar` | `## Architecture Decisions` | when a non-trivial design choice was made |
|
|
16
|
-
| 4 | `
|
|
17
|
-
| 5 | `
|
|
18
|
-
| 6 | `
|
|
19
|
-
| 7 | `
|
|
20
|
-
| 8 | `
|
|
16
|
+
| 4 | `impact` | `## Etki Analizi` | `## Impact Analysis` | always |
|
|
17
|
+
| 5 | `test_scenarios` | `## Test Senaryoları` | `## Test Scenarios` | always |
|
|
18
|
+
| 6 | `visuals` | `## Görsel Kanıt` | `## Visual Evidence` | when `state.visualEvidence.required` |
|
|
19
|
+
| 7 | `risk` | `## Risk ve Güvenlik` | `## Risk and Security` | when `state.diffRisk.signals` carries a high-stakes signal (`security_path`, `migration`, `public_api`, `no_test_change`, `test_lines_removed`) |
|
|
20
|
+
| 8 | `dependencies` | `## Bağımlılıklar` | `## Dependencies` | when deps added/removed/bumped |
|
|
21
|
+
| 9 | `build` | `## Build` | `## Build` | always |
|
|
22
|
+
| 10 | `related` | `## İlgili` | `## Related` | always (Jira/issue ref; never `Closes/Fixes`) |
|
|
23
|
+
|
|
24
|
+
`summary`, `impact` and `test_scenarios` carry the same three section keys the
|
|
25
|
+
Jira adapter uses, and that pairing is deliberate: the same three questions get
|
|
26
|
+
answered on both surfaces, at the register each reader needs. The PR versions name
|
|
27
|
+
symbols, files, line counts and build shas. The Jira versions name screens and
|
|
28
|
+
behaviour and nothing else (`channels/jira.md`). `technical` has no Jira twin at
|
|
29
|
+
all - it is the section whose absence over there is the point.
|
|
21
30
|
|
|
22
31
|
### Section content rules
|
|
23
32
|
|
|
24
|
-
**`summary`** -
|
|
33
|
+
**`summary`** - 2-4 sentences in `outputLanguage`. What broke or what was added, where the user meets it, and the size of the effect when it is measurable (a crash count, a share of a known total, a version range). Past tense, no marketing voice. Code identifiers stay verbatim. This is the one section a reviewer reads before deciding whether to read the rest, so it names the user-visible behaviour before the mechanism.
|
|
25
34
|
|
|
26
|
-
**`
|
|
35
|
+
**`technical`** - the account for someone holding the diff: a short paragraph on the mechanism (what the code was actually doing wrong, or what the new code does), then the per-file bullets, then one line of diff stat (`2 files, 8 deletions, 0 insertions`). Where a change is safe for a reason that is not obvious from the diff - an equality relation preserved, an invariant kept, a call site left alone deliberately - that reason belongs here in a sentence, because it is the question the reviewer would otherwise ask in a comment. Tables are welcome when several symbols share a property worth listing side by side.
|
|
36
|
+
|
|
37
|
+
The bullet list is one item per logically distinct change. Each bullet starts with the touched component and ends with a one-line "what". The source is `$WORKTREE/.pipeline/scope-check.json` `files[].reason` (Phase 3 Step 3.7): a file the dev could not justify there is a file this list cannot describe either, so the bullet quotes the gate output instead of inventing a reason. Use the stack's native file extensions / module paths - the example below shows the **shape**, not a stack lock-in:
|
|
27
38
|
|
|
28
39
|
```markdown
|
|
29
40
|
## Changes
|
|
@@ -37,20 +48,55 @@ Across stacks the same shape produces, for example: `LoginView.swift - ...` (i
|
|
|
37
48
|
|
|
38
49
|
**`architecture`** - only when the change involves a non-trivial decision (new abstraction, pattern change, data flow shift, dependency direction). Format: short paragraph stating the decision and the alternative considered. Skip the section entirely for mechanical refactors / dependency bumps / formatting passes.
|
|
39
50
|
|
|
40
|
-
**`
|
|
51
|
+
**`impact`** - four fixed numbered parts, each answered, never a placeholder. Same four questions as the Jira `impact` section, answered here with the identifiers and numbers that section is not allowed to carry:
|
|
52
|
+
|
|
53
|
+
```markdown
|
|
54
|
+
## Impact Analysis
|
|
55
|
+
|
|
56
|
+
**1 - The problem**
|
|
57
|
+
<what was wrong, since which version, with the crash/report identifiers and counts>
|
|
58
|
+
|
|
59
|
+
**2 - What was changed**
|
|
60
|
+
<the change, and why behaviour around it is unchanged>
|
|
61
|
+
|
|
62
|
+
**3 - Affected functions**
|
|
63
|
+
<the screens, flows and symbols that must be tested, including shared components the change reaches>
|
|
64
|
+
|
|
65
|
+
**4 - Effect on other systems**
|
|
66
|
+
<none, or which service, contract or channel>
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
When the change deliberately fixes part of a wider problem, a closing **Risk and remaining scope** paragraph names what is still open and why it was left - a reviewer who can see the rest of the pattern in the repo will ask otherwise, and the honest answer is cheaper written down than defended in a thread.
|
|
41
70
|
|
|
42
|
-
|
|
71
|
+
**`test_scenarios`** - the same titled-scenario shape the Jira adapter uses, so the tester reads one list on both surfaces, with symbols allowed here:
|
|
43
72
|
|
|
44
73
|
```markdown
|
|
45
|
-
##
|
|
74
|
+
## Test Scenarios
|
|
46
75
|
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
76
|
+
**1. <what this scenario exercises>**
|
|
77
|
+
1. <step>
|
|
78
|
+
2. **Expected:** <observable outcome>
|
|
79
|
+
|
|
80
|
+
**2. <regression scenario>**
|
|
81
|
+
1. <step>
|
|
82
|
+
2. **Expected:** <what should still behave as before>
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
Reproduction scenarios first, then regressions for whatever the change could have disturbed. When a UI test target covers a scenario, say so on the scenario line and let `## Build` carry the run result - a scenario a machine already ran is not the same request as one a human has to perform, and conflating them wastes the reviewer's time.
|
|
86
|
+
|
|
87
|
+
**`build`** - what was built, on what, and what came out. Not a promise that it builds; the recorded result of the run that happened:
|
|
88
|
+
|
|
89
|
+
```markdown
|
|
90
|
+
## Build
|
|
91
|
+
|
|
92
|
+
- <build command or scheme> on <base branch>@<sha>: BUILD SUCCEEDED, 0 errors
|
|
93
|
+
- Tests: <test command>: <N> passed, <M> failed
|
|
94
|
+
- UI tests: <target>: <status> | not run - <reason from state.uiTest.notRunReason>
|
|
51
95
|
```
|
|
52
96
|
|
|
53
|
-
|
|
97
|
+
The base sha matters because "it builds" is a claim about a merge base, and the reviewer's local tree is usually not that one. `state.uiTest` supplies the third line verbatim, including its `notRunReason` - "no UI test target in this repo" is a result, not a gap, and writing it stops the same question being asked on every PR.
|
|
98
|
+
|
|
99
|
+
Pick commands for the project's stack - the pipeline supports iOS (Swift/Xcode), Android (Gradle), web (npm/pnpm/yarn) and backend (pytest/jest/go test/etc). Multi-repo PRs (one PR per repo) emit the commands for that repo's stack only - never mix iOS + Android commands into a single PR body.
|
|
54
100
|
|
|
55
101
|
**`risk`** - only when `state.diffRisk.signals` (Phase 4 Step 1.75) contains a high-stakes signal. Four fixed lines, each answered, never left as a placeholder; the source is Phase 1 `touchedAreas` plus the signals themselves, and when a signal is present the absence of this section is a Phase 6 Step 3 blocker:
|
|
56
102
|
|
|
@@ -136,7 +182,7 @@ Never use `Closes #N`, `Fixes #N`, `Resolves PROJ-X`. Issues require 4-approval
|
|
|
136
182
|
|
|
137
183
|
```
|
|
138
184
|
1. Read agent-state.json (taskId, contextLinks, identity, language).
|
|
139
|
-
2. Build section bodies in markdown - summary first, then in the table order, skipping conditional sections that don't apply. Section order is fixed: `summary` → `
|
|
185
|
+
2. Build section bodies in markdown - summary first, then in the table order, skipping conditional sections that don't apply. Section order is fixed: `summary` → `technical` → `architecture` (cond.) → `impact` → `test_scenarios` → `visuals` (cond.) → `risk` (cond.) → `dependencies` (cond.) → `build` → `related`.
|
|
140
186
|
3. Run the assembled body through the `humanizer` skill.
|
|
141
187
|
4. Apply Multi-repo cross-links (## Related PRs prepend when projects.length > 1).
|
|
142
188
|
5. Dispatch per the Behaviour-by-remote table.
|
|
@@ -220,7 +266,7 @@ The Bitbucket REST API returns `409 Conflict` if `version` is stale. Adapter beh
|
|
|
220
266
|
## Hard rules (must not regress)
|
|
221
267
|
|
|
222
268
|
- Real newlines, no HTML entities - heredoc + `jq --rawfile` + `curl --data-binary @file`. Never embed `\n` literally; Bitbucket stores the literal `\n` characters.
|
|
223
|
-
- Section order is fixed: `summary` → `
|
|
269
|
+
- Section order is fixed: `summary` → `technical` → `architecture` (cond.) → `impact` → `test_scenarios` → `visuals` (cond.) → `risk` (cond.) → `dependencies` (cond.) → `build` → `related`. Conditional sections may be omitted but never reordered or inserted between fixed ones.
|
|
224
270
|
- Humanizer pass runs **after** body assembly and **before** dispatch - every body line carries technical, non-AI tone. References at least one symbol/file/line drawn from the diff or pipeline log.
|
|
225
271
|
- Body content language follows `prefs.global.outputLanguage`. Code identifiers, file paths, branch names, PR titles, commands, type names, and `Closes/Fixes`-style keywords stay verbatim English. The template file (this doc) is English because `promptLanguage="en"` is locked; the body is rendered in the user's language at write-time.
|
|
226
272
|
- No body markers (`<!-- channels:start -->`) - channels does a full replace each run.
|
|
@@ -6,13 +6,14 @@
|
|
|
6
6
|
|
|
7
7
|
---
|
|
8
8
|
|
|
9
|
-
## 1. Command Inventory (
|
|
9
|
+
## 1. Command Inventory (56 commands)
|
|
10
10
|
|
|
11
11
|
```
|
|
12
|
-
analysis, analysis-jira, analysis-resolve, autopilot,
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
12
|
+
analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
|
|
13
|
+
autopilot-on, autopilot-status, build-optimize, channels, complaint-analysis,
|
|
14
|
+
create-jira, design-check, diff-explain, doctor, feedback, forget,
|
|
15
|
+
garbage-collect, graph, help, ios-coding-standard, issue, jira, kill,
|
|
16
|
+
language, local, local-autopilot, log, manual-test, prune-logs,
|
|
16
17
|
prune-prompts, purge, refactor, resume, resume-local, review,
|
|
17
18
|
review-analysis, review-issue, review-jira, routines, save, scan, search,
|
|
18
19
|
setup, stack, status, steer, store-ready, sync, test, test-accessibility,
|