@mmerterden/multi-agent-pipeline 17.0.0 → 17.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/CHANGELOG.md +32 -0
  2. package/README.md +49 -4
  3. package/README.tr.md +50 -4
  4. package/docs/architecture.md +3 -3
  5. package/docs/ecosystem.md +5 -5
  6. package/install/templates/multi-agent-autopilot.plist.template +79 -0
  7. package/package.json +1 -1
  8. package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +64 -0
  9. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +173 -0
  10. package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +74 -0
  11. package/pipeline/commands/multi-agent/channels/SKILL.md +41 -12
  12. package/pipeline/commands/multi-agent/help/SKILL.md +41 -35
  13. package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
  14. package/pipeline/commands/multi-agent/sync/SKILL.md +10 -9
  15. package/pipeline/commands/multi-agent/update/SKILL.md +1 -1
  16. package/pipeline/lib/autopilot-activation.sh +117 -0
  17. package/pipeline/lib/autopilot-state.sh +150 -0
  18. package/pipeline/lib/issue-fetcher.sh +18 -1
  19. package/pipeline/lib/plan-todos.sh +18 -0
  20. package/pipeline/multi-agent-refs/channels/jira.md +80 -20
  21. package/pipeline/multi-agent-refs/channels/pr.md +65 -19
  22. package/pipeline/multi-agent-refs/cross-cli-contract.md +6 -5
  23. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -7
  24. package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
  25. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +17 -15
  26. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +1 -1
  27. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +1 -1
  28. package/pipeline/multi-agent-refs/readiness-review.md +7 -1
  29. package/pipeline/multi-agent-refs/rules.md +3 -11
  30. package/pipeline/multi-agent-refs/tracker-contract.md +32 -0
  31. package/pipeline/schemas/autopilot-config.schema.json +149 -0
  32. package/pipeline/schemas/token-budget.json +2 -2
  33. package/pipeline/scripts/autopilot-arming.mjs +147 -0
  34. package/pipeline/scripts/autopilot-intake.mjs +383 -0
  35. package/pipeline/scripts/autopilot-menubar.swift +361 -0
  36. package/pipeline/scripts/autopilot-runner.mjs +349 -0
  37. package/pipeline/scripts/autopilot-status.sh +212 -0
  38. package/pipeline/scripts/jira-search.sh +70 -0
  39. package/pipeline/scripts/phase-tracker.sh +134 -12
  40. package/pipeline/scripts/probe-evidence-capability.sh +27 -3
  41. package/pipeline/scripts/run-ui-tests.sh +113 -4
  42. package/pipeline/skills/.skill-manifest.json +16 -4
  43. package/pipeline/skills/shared/core/multi-agent-autopilot-off/SKILL.md +67 -0
  44. package/pipeline/skills/shared/core/multi-agent-autopilot-on/SKILL.md +146 -0
  45. package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +64 -0
  46. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +62 -11
  47. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +9 -8
@@ -0,0 +1,117 @@
1
+ #!/usr/bin/env bash
2
+ # autopilot-activation.sh - which repos may this machine pick work up from.
3
+ #
4
+ # The answer is NEVER "all of them" and never "whatever has the label". Measured
5
+ # on one machine: 66 repos grant push, and a good half belong to other people -
6
+ # a colleague's backend, someone's side project. A label is a filter, not a gate:
7
+ # anyone able to open an issue in a repo you happen to have push on could put
8
+ # `agent-queue` on it. The gate is this picker, and the picker is per machine.
9
+ #
10
+ # Two lists come out, and the difference matters:
11
+ #
12
+ # eligible push rights AND a checkout on this machine
13
+ # unavailable push rights, no checkout - LISTED, with that as the reason
14
+ #
15
+ # The second list is printed rather than dropped. Silently hiding 40 repos from a
16
+ # 66-repo answer looks exactly like a permissions problem, and the user then goes
17
+ # looking for a token fault that does not exist.
18
+ #
19
+ # Usage:
20
+ # bash autopilot-activation.sh list # TSV: state, nameWithOwner, localPath
21
+ # bash autopilot-activation.sh list --json
22
+ # . autopilot-activation.sh; ma_ap_local_path <owner/repo>
23
+ #
24
+ # Exit: 0 on success (including an empty list), 2 on usage, 3 when gh cannot answer.
25
+
26
+ set -uo pipefail
27
+
28
+ # Repo roots to search for checkouts. $HOME's immediate children covers the
29
+ # normal layout; MA_AUTOPILOT_REPO_ROOTS overrides it for tests and for anyone
30
+ # who keeps checkouts somewhere else.
31
+ MA_AP_ROOTS="${MA_AUTOPILOT_REPO_ROOTS:-$HOME}"
32
+
33
+ # owner/repo for a checkout, read from its origin remote. Both URL forms, and a
34
+ # trailing .git is stripped - `git@github.com:o/r.git` and
35
+ # `https://github.com/o/r` are the same repo and a picker that treats them as
36
+ # two shows the user a duplicate they cannot explain.
37
+ ma_ap_remote_slug() {
38
+ local dir="$1" url
39
+ url=$(git -C "$dir" remote get-url origin 2>/dev/null) || return 1
40
+ url="${url%.git}"
41
+ case "$url" in
42
+ *github.com[:/]*) printf '%s\n' "${url##*github.com[:/]}" ;;
43
+ *) return 1 ;;
44
+ esac
45
+ }
46
+
47
+ # Every checkout under the roots, as `slug<TAB>path`. One pass, so the lookup
48
+ # below is a grep rather than a walk per repo.
49
+ ma_ap_checkout_index() {
50
+ local root d slug
51
+ for root in $MA_AP_ROOTS; do
52
+ [ -d "$root" ] || continue
53
+ for d in "$root"/*/; do
54
+ [ -d "$d/.git" ] || continue
55
+ slug=$(ma_ap_remote_slug "${d%/}") || continue
56
+ printf '%s\t%s\n' "$slug" "${d%/}"
57
+ done
58
+ done
59
+ }
60
+
61
+ ma_ap_local_path() { # $1 = owner/repo
62
+ ma_ap_checkout_index | awk -F'\t' -v s="$1" 'tolower($1)==tolower(s){print $2; exit}'
63
+ }
64
+
65
+ # Repos the authenticated account may push to.
66
+ #
67
+ # The query string form is deliberate: `gh api user/repos --field affiliation=...`
68
+ # sends the value as a BODY field on a GET, which this endpoint ignores - it
69
+ # answers with the default affiliation and the caller never learns the filter did
70
+ # not apply. Measured: the --field form returned 0 rows here, the query string
71
+ # form returned 66.
72
+ ma_ap_writable_repos() {
73
+ gh api "user/repos?affiliation=owner,collaborator,organization_member&per_page=100" \
74
+ --paginate -q '.[] | select(.permissions.push == true) | .full_name' 2>/dev/null
75
+ }
76
+
77
+ ma_ap_list() {
78
+ local json="${1:-}" slug path index
79
+ index=$(ma_ap_checkout_index)
80
+ local repos
81
+ repos=$(ma_ap_writable_repos)
82
+ if [ -z "$repos" ]; then
83
+ echo "autopilot-activation: gh returned no writable repos - check 'gh auth status'" >&2
84
+ return 3
85
+ fi
86
+
87
+ local first=1
88
+ [ "$json" = "--json" ] && printf '['
89
+ while IFS= read -r slug; do
90
+ [ -n "$slug" ] || continue
91
+ path=$(printf '%s\n' "$index" | awk -F'\t' -v s="$slug" 'tolower($1)==tolower(s){print $2; exit}')
92
+ local state="eligible"
93
+ [ -z "$path" ] && state="unavailable"
94
+ if [ "$json" = "--json" ]; then
95
+ [ "$first" -eq 0 ] && printf ','
96
+ first=0
97
+ printf '{"state":"%s","nameWithOwner":"%s","localPath":"%s"}' "$state" "$slug" "$path"
98
+ else
99
+ printf '%s\t%s\t%s\n' "$state" "$slug" "$path"
100
+ fi
101
+ done <<EOF
102
+ $repos
103
+ EOF
104
+ [ "$json" = "--json" ] && printf ']\n'
105
+ return 0
106
+ }
107
+
108
+ if [ "${BASH_SOURCE[0]:-$0}" = "$0" ]; then
109
+ case "${1:-}" in
110
+ list) ma_ap_list "${2:-}" ;;
111
+ "" | -h | --help) grep -E '^#( |$)' "$0" | sed -E 's/^# ?//' ;;
112
+ *)
113
+ echo "autopilot-activation: unknown command ${1}" >&2
114
+ exit 2
115
+ ;;
116
+ esac
117
+ fi
@@ -0,0 +1,150 @@
1
+ #!/usr/bin/env bash
2
+ # autopilot-state.sh - where continuous mode keeps its state, and who may read it.
3
+ #
4
+ # Everything lives under ~/.claude/autopilot/ and NOT in
5
+ # multi-agent-preferences.json. That file has `additionalProperties: false` and a
6
+ # migration chain, so a block there would push keys into every local user's
7
+ # preferences forever - including everyone who never turns this on. Separation
8
+ # also gives the off state a definition that cannot be got wrong: the directory
9
+ # is absent.
10
+ #
11
+ # config.json the repo selection. Survives autopilot-off on purpose, so
12
+ # turning the mode back on does not re-ask which repos.
13
+ # queue.json the current ordered list plus whatever is in flight
14
+ # status.json what the menu bar renders. Written by the runner, read-only
15
+ # to everything else
16
+ # attempted.jsonl append-only: {source,id,taskId,outcome,prUrl,at}
17
+ # runner.pid pid + the in-flight sessionId + kern.boottime
18
+ # runner.log
19
+ # bin/menubar built on demand from scripts/autopilot-menubar.swift
20
+ #
21
+ # 0700 throughout. The queue names real tickets and real repos, and on a shared
22
+ # machine that is somebody's roadmap.
23
+ #
24
+ # Usage:
25
+ # . autopilot-state.sh
26
+ # ma_ap_root # prints the dir, creating it 0700 only when asked
27
+ # ma_ap_is_on # exit 0 when configured, 1 otherwise. NO side effects
28
+ # ma_ap_write <file> <- # atomic write from stdin, 0600
29
+
30
+ set -uo pipefail
31
+
32
+ MA_AP_ROOT="${MA_AUTOPILOT_ROOT:-$HOME/.claude/autopilot}"
33
+
34
+ ma_ap_root() { printf '%s\n' "$MA_AP_ROOT"; }
35
+
36
+ MA_AP_LABEL="${MA_AUTOPILOT_LABEL:-com.multi-agent.autopilot}"
37
+ MA_AP_PLIST="${MA_AUTOPILOT_PLIST:-$HOME/Library/LaunchAgents/$MA_AP_LABEL.plist}"
38
+
39
+ # Two different questions, and conflating them was a real bug in the first draft
40
+ # of this file: `autopilot-off` deliberately KEEPS config.json so turning the
41
+ # mode back on does not re-ask which repos, so a predicate reading the config
42
+ # would still answer "on" after you turned it off.
43
+ #
44
+ # configured = a repo selection exists
45
+ # on = launchd holds the job
46
+ #
47
+ # On is the launchd job because that is the thing that actually makes work
48
+ # happen; anything else is a claim about a file.
49
+ ma_ap_is_configured() { [ -f "$MA_AP_ROOT/config.json" ]; }
50
+
51
+ ma_ap_is_on() { [ -f "$MA_AP_PLIST" ] && launchctl list 2>/dev/null | grep -q "$MA_AP_LABEL"; }
52
+
53
+ # The failure mode a `resume` command would have papered over: configured and
54
+ # meant to be running, but launchd does not hold the job - an OS update dropped
55
+ # the plist, or it was booted out by hand. doctor reports this; there is no
56
+ # command to remember.
57
+ ma_ap_is_orphaned() { ma_ap_is_configured && ! ma_ap_is_on && [ -f "$MA_AP_PLIST" ]; }
58
+
59
+ # Deliberately no side effects in any of the three. A predicate that creates its
60
+ # own directory turns "is autopilot on?" into "autopilot is now half on", and
61
+ # every status command, hook and doctor check calls these.
62
+
63
+ ma_ap_ensure_root() {
64
+ [ -d "$MA_AP_ROOT" ] || mkdir -p "$MA_AP_ROOT" || return 1
65
+ chmod 700 "$MA_AP_ROOT" 2>/dev/null
66
+ [ -d "$MA_AP_ROOT/bin" ] || mkdir -p "$MA_AP_ROOT/bin" 2>/dev/null
67
+ chmod 700 "$MA_AP_ROOT/bin" 2>/dev/null
68
+ return 0
69
+ }
70
+
71
+ # Write via a temp file in the SAME directory then rename. The menu bar polls
72
+ # status.json every few seconds, and a reader that catches a half-written file
73
+ # shows an empty menu; rename is atomic on the same filesystem, so it never sees
74
+ # a partial one.
75
+ ma_ap_write() { # $1 = filename under the root; content on stdin
76
+ ma_ap_ensure_root || return 1
77
+ local dst="$MA_AP_ROOT/$1" tmp="$MA_AP_ROOT/.$1.$$"
78
+ cat > "$tmp" || {
79
+ rm -f "$tmp"
80
+ return 1
81
+ }
82
+ chmod 600 "$tmp" 2>/dev/null
83
+ mv -f "$tmp" "$dst"
84
+ }
85
+
86
+ ma_ap_append() { # $1 = filename, content on stdin - for the jsonl
87
+ ma_ap_ensure_root || return 1
88
+ local dst="$MA_AP_ROOT/$1"
89
+ cat >> "$dst" || return 1
90
+ chmod 600 "$dst" 2>/dev/null
91
+ }
92
+
93
+ ma_ap_read() { # $1 = filename; empty and exit 1 when absent
94
+ local src="$MA_AP_ROOT/$1"
95
+ [ -f "$src" ] || return 1
96
+ cat "$src"
97
+ }
98
+
99
+ # jq with a default, so a caller never has to distinguish "key absent" from
100
+ # "file absent" from "file unparseable" - all three mean "use the default".
101
+ ma_ap_cfg() { # $1 = jq path, $2 = default
102
+ local v
103
+ v=$(ma_ap_read config.json 2>/dev/null | jq -r "$1 // empty" 2>/dev/null)
104
+ [ -n "$v" ] && printf '%s\n' "$v" || printf '%s\n' "$2"
105
+ }
106
+
107
+ # AC or battery. The mode holds a sleep assertion only on AC: a queue that keeps
108
+ # a laptop awake on battery is a bug, and `StartInterval` does not wake a
109
+ # sleeping Mac anyway, so on battery the work resumes when you plug in.
110
+ ma_ap_power() {
111
+ case "$(pmset -g ps 2>/dev/null | head -1)" in
112
+ *"AC Power"*) printf 'ac\n' ;;
113
+ *) printf 'battery\n' ;;
114
+ esac
115
+ }
116
+
117
+ # Boot time, so a stale runner.pid cannot be mistaken for a live runner. After a
118
+ # restart pids start low and the recorded 4711 may belong to something unrelated;
119
+ # a pid recorded BEFORE the current boot is stale by definition, with no probing.
120
+ # The pattern is ANCHORED on purpose. `kern.boottime` prints
121
+ # `{ sec = 1788181644, usec = 618574 } Mon Aug 31 ...`, and `.*sec = ` is greedy:
122
+ # it walks past `sec` to `usec` and captures 618574, the microseconds. The
123
+ # original form here did exactly that, so this returned a six-digit number that
124
+ # looked plausible and never matched the same fact read anywhere else - which is
125
+ # how the runner's staleness check compared two different numbers and called a
126
+ # live runner dead.
127
+ ma_ap_boottime() {
128
+ sysctl -n kern.boottime 2>/dev/null | sed -n 's/^{ *sec = \([0-9]*\).*/\1/p'
129
+ }
130
+
131
+ # The machine's honest ceiling, so raising `slots` is a measured decision. RAM
132
+ # and cores bound the concurrent Claude sessions; disk bounds the worktrees
133
+ # (0.75 GB each, measured at 11 GB across 15). The hard cap is 4 because beyond
134
+ # that the bottleneck stops being this machine.
135
+ ma_ap_slot_ceiling() {
136
+ local ram_gb cores free_gb a b c
137
+ ram_gb=$(( $(sysctl -n hw.memsize 2>/dev/null || echo 0) / 1073741824 ))
138
+ cores=$(sysctl -n hw.ncpu 2>/dev/null || echo 2)
139
+ free_gb=$(df -g "$HOME" 2>/dev/null | awk 'NR==2{print $4}')
140
+ [ -n "$free_gb" ] || free_gb=0
141
+ a=$(((ram_gb - 4) * 10 / 25))
142
+ b=$((cores / 2))
143
+ c=$(((free_gb - 20) * 100 / 75))
144
+ local m=$a
145
+ [ "$b" -lt "$m" ] && m=$b
146
+ [ "$c" -lt "$m" ] && m=$c
147
+ [ "$m" -gt 4 ] && m=4
148
+ [ "$m" -lt 1 ] && m=1
149
+ printf '%s\n' "$m"
150
+ }
@@ -49,8 +49,23 @@
49
49
 
50
50
  set -euo pipefail
51
51
 
52
+ # Sourcing this file hands a caller the provider helpers - `jira_search`,
53
+ # `fetch_jira`, `fetch_github` and the `jira_curl` underneath them - without
54
+ # resolving anything. Executing it resolves one input, exactly as before.
55
+ #
56
+ # The guard exists because the alternative is worse than it looks: a second
57
+ # consumer of `jira_search` with no way to source would have to copy
58
+ # `jira_curl`, and that function is not a URL - it is the credential resolution,
59
+ # the host lookup and the rule that keeps a token off argv. Three copies of that
60
+ # is three places for a token to leak.
61
+ MA_IF_EXECUTED=0
62
+ [ "${BASH_SOURCE[0]}" = "$0" ] && MA_IF_EXECUTED=1
63
+
52
64
  INPUT="${1:-}"
53
- [ -z "$INPUT" ] && { echo '{"error":"input required"}'; exit 1; }
65
+ if [ "$MA_IF_EXECUTED" = 1 ] && [ -z "$INPUT" ]; then
66
+ echo '{"error":"input required"}'
67
+ exit 1
68
+ fi
54
69
 
55
70
  JIRA_TOKEN_KEY="${ACCOUNT_JIRA_TOKEN_KEY:-}"
56
71
  JIRA_HOST="${ACCOUNT_JIRA_HOST:-}"
@@ -361,6 +376,7 @@ PY
361
376
  }
362
377
 
363
378
  # --- Dispatch -----------------------------------------------------------------
379
+ if [ "$MA_IF_EXECUTED" = 1 ]; then
364
380
  case "$KIND" in
365
381
  jira-id)
366
382
  KEY="$INPUT"
@@ -592,3 +608,4 @@ sys.stdout.write("\x1f".join(parts))
592
608
  "description=" "branchHint=$branch"
593
609
  ;;
594
610
  esac
611
+ fi
@@ -94,6 +94,24 @@ do_set() {
94
94
  if [ -z "$plan_blob" ] || [ "$plan_blob" = "-" ]; then
95
95
  plan_blob=$(cat)
96
96
  fi
97
+ # A planning-output document is accepted directly and converted here.
98
+ #
99
+ # The conversion used to live as a jq blob inside phase-2-planning.md, which
100
+ # made the mapping from `tasks[]` to `todos[]` a thing two files defined - and
101
+ # the phase doc was the copy nothing tested. Accepting both shapes costs four
102
+ # lines and removes the second definition.
103
+ if jq -e '.tasks and (.todos | not)' <<<"$plan_blob" >/dev/null 2>&1; then
104
+ plan_blob=$(jq '{
105
+ title: (.summary // .title // "plan"),
106
+ todos: [ .tasks[] | {
107
+ id: .id,
108
+ task: (.title // .subject // ""),
109
+ status: "pending",
110
+ deps: (.dependsOn // .blockedBy // [])
111
+ } ]
112
+ }' <<<"$plan_blob")
113
+ fi
114
+
97
115
  # Validate against schema (best-effort - jq syntax check, then required-field probe).
98
116
  if ! jq -e '.title and (.todos | type == "array")' <<<"$plan_blob" >/dev/null 2>&1; then
99
117
  echo "plan-todos: input must be an object with .title (string) and .todos (array)" >&2
@@ -10,35 +10,89 @@ Every Jira comment posted by this adapter follows the same section order. Sectio
10
10
 
11
11
  | # | Section key | Heading (`tr`) | Heading (`en`) | Required? |
12
12
  |---|---|---|---|---|
13
- | 1 | `summary` | `## Yapılan Çalışma Özeti` | `## Work Summary` | always |
13
+ | 1 | `summary` | `## Geliştirme Özeti` | `## Development Summary` | always |
14
14
  | 2 | `test_scenarios` | `## Test Senaryoları` | `## Test Scenarios` | always (use " - " placeholder line if truly N/A) |
15
- | 3 | `context_refs` | `## Bağlantılar` | `## References` | when any link exists |
15
+ | 3 | `impact` | `## Etki Analizi` | `## Impact Analysis` | always |
16
+ | 4 | `context_refs` | `## Bağlantılar` | `## References` | when any link exists |
16
17
 
17
18
  > Heading levels in the source markdown are `##`. The wiki-markup converter (next section) rewrites them to `h2.` for Jira rendering.
18
19
 
20
+ **Jira is not the PR, and this is the contract that keeps getting that wrong.**
21
+ The PR is read by a reviewer holding the diff. The Jira comment is read by the
22
+ person who filed the ticket and by the tester who has to verify it, and neither
23
+ of them has the diff open. So a Jira body carries no identifiers, no file paths,
24
+ no stack frames, no diff hunks and no framework names. Counts in prose are fine
25
+ and are often the most useful sentence in the comment ("two files, eight lines
26
+ removed, no additions"); a class name is not. A sentence that needs a symbol to
27
+ make its point is describing the change at the wrong level for this reader -
28
+ restate the behaviour, not the mechanism. The technical account has a home, and
29
+ it is the PR body (`channels/pr.md`).
30
+
19
31
  ### Section content rules
20
32
 
21
33
  **`summary`** - 2-5 sentences in `outputLanguage`. What changed, why, and the user-visible impact. No "we", no marketing tone. Past tense (the work is done at the time the comment goes up).
22
34
 
23
- **Visual evidence inside these two sections.** When `state.visualEvidence` carries artefacts, they render INSIDE `summary` and `test_scenarios` - never as a fourth section, which the fixed section order forbids. Upload first (`jira-attach.sh <issue> <file>...`), then reference by the returned filename:
35
+ **Visual evidence inside these sections.** When `state.visualEvidence` carries artefacts, they render INSIDE `summary` and `test_scenarios` - never as a section of their own, which the fixed section order forbids. Phase 6 Step 2.9 has already uploaded them and written the returned name to `visualEvidence.*[].jiraFilename`: reference that name, and call `jira-attach.sh <issue> <file>...` only for an artefact whose `jiraFilename` is absent. Uploading unconditionally here attaches every file a second time whenever the comment is re-rendered - which is exactly what a post-hoc `/multi-agent:channels` run does.
24
36
 
25
37
  - `summary`, after its sentences: one line naming the pair in `outputLanguage` (`Düzeltme öncesi / Düzeltme sonrası`), then the thumbnails on the next line - `!<file>-before.png|thumbnail! !<file>-after.png|thumbnail!`.
26
38
  - `test_scenarios`, under the scenario the recording demonstrates: `!<file>-flow.mp4!` plus one line stating the tier used.
27
39
 
28
40
  A `gaps[]` entry prints its reason on the line where the artefact would have been (`Düzeltme öncesi: ticket'ta görsel yok`). Never an empty thumbnail, never a silent omission. Contract: `$HOME/.claude/multi-agent-refs/features/visual-evidence.md`.
29
41
 
30
- **`test_scenarios`** - Given/When/Then numbered list. One scenario per acceptance criterion. The heading and scenario text are rendered in `outputLanguage` at write-time; the template itself (this file) shows the English skeleton:
42
+ **`test_scenarios`** - one titled scenario per acceptance criterion, each a
43
+ numbered list of steps ending in the expected result. Given/When/Then was the
44
+ shape here for several releases and it reads as translated English to the person
45
+ who actually runs these; a tester wants a list they can follow with the app open.
46
+ The heading and the text render in `outputLanguage`; this file shows the English
47
+ skeleton:
31
48
 
32
49
  ```markdown
33
50
  ## Test Scenarios
34
51
 
35
- 1. **Given** a signed-in user
36
- **When** they navigate to the Profile screen
37
- **Then** their saved preferences appear in the configured order
38
- 2. **Given** ...
52
+ **1. <what this scenario exercises>**
53
+ 1. <step the tester performs, in the app's own words>
54
+ 2. <step>
55
+ 3. **Expected:** <what they should see>
56
+
57
+ **2. <regression scenario>**
58
+ 1. <step>
59
+ 2. **Expected:** <what should still behave as before>
39
60
  ```
40
61
 
41
- Code identifiers (function names, file paths) stay verbatim - they are not translated. When there are no executable scenarios (pure doc change, dependency bump), write a single line: `- No manual test required - <short reason>`.
62
+ Order is load-bearing. The scenarios that reproduce the reported behaviour come
63
+ first, then the regression scenarios for whatever the change could have disturbed -
64
+ anything sharing the component that was touched. A comment that lists only the fix
65
+ leaves the tester to guess the blast radius, and that guess is the one thing they
66
+ cannot make from the ticket.
67
+
68
+ Screens, fields and menu items are named the way the app shows them, in
69
+ `outputLanguage`. Function names and file paths do not appear at all, per the rule
70
+ above. When there are no executable scenarios (pure doc change, dependency bump),
71
+ write a single line: `- No manual test required - <short reason>`.
72
+
73
+ **`impact`** - four fixed numbered parts, each answered, never left as a
74
+ placeholder. This is what a test lead reads before deciding how wide to test and
75
+ what a release manager reads before deciding whether it ships this week, so it is
76
+ written in the same plain register as the rest of the comment:
77
+
78
+ ```markdown
79
+ ## Impact Analysis
80
+
81
+ **1 - The problem**
82
+ <what was wrong and since when, in the reader's words>
83
+
84
+ **2 - What was changed**
85
+ <what the change does, and why the behaviour around it is unchanged>
86
+
87
+ **3 - Affected areas**
88
+ <the screens and flows that must be tested, including anything sharing the changed component>
89
+
90
+ **4 - Effect on other systems**
91
+ <none, or which service, contract or channel is affected>
92
+ ```
93
+
94
+ Part 4 is answered "none" far more often than it is answered at all, and "none" is
95
+ a real answer worth writing: it is what lets the reader stop looking.
42
96
 
43
97
  **`context_refs`** - flat bullet list of external links the work depends on or produced. PR URL is **not** repeated here; it lives on the first line (see *Cross-link injection*). Common entries (only when present in the task):
44
98
 
@@ -58,7 +112,7 @@ The links are pulled from `agent-state.json.contextLinks[]` (see Phase 0 link ex
58
112
 
59
113
  ```
60
114
  1. Read agent-state.json (taskId, contextLinks, prUrls, language).
61
- 2. Build section bodies in markdown - summary first, test_scenarios next, context_refs last.
115
+ 2. Build section bodies in markdown, in table order: summary, test_scenarios, impact, context_refs.
62
116
  3. Run the assembled body through the `humanizer` skill (see Hard rules below).
63
117
  4. Apply Cross-link injection (PR URL on line 1).
64
118
  5. Run Wiki markup conversion.
@@ -116,7 +170,7 @@ Do not hand-apply this table. It was hand-applied for several releases and a smi
116
170
 
117
171
  Lines outside these patterns pass through verbatim. Multi-paragraph blocks are joined with one blank line.
118
172
 
119
- Every row above is load-bearing, including the ones that look cosmetic. The table must cover each construct the section templates in this file actually emit - `##` for the three required headings, `**Given**` / `**When**` / `**Then**` in the test-scenario skeleton, and `1.`-numbered scenarios. A missing row does not degrade gracefully: the "pass through verbatim" fallback POSTs `## Test Senaryoları` and `**Given**` as literal text, so the comment renders with visible `##` and stray asterisks. Single `*bold*` in Markdown means *italic* in Jira wiki - never map `**bold**` to `*bold*` by dropping one asterisk mechanically without checking the source was bold, not italic.
173
+ Every row above is load-bearing, including the ones that look cosmetic. The table must cover each construct the section templates in this file actually emit - `##` for the four required headings, `**1. title**` on a scenario and `**Expected:**` inside it, the `**1 - ...**` part labels in the impact section, and `1.`-numbered steps. A missing row does not degrade gracefully: the "pass through verbatim" fallback POSTs `## Test Senaryoları` and `**Expected:**` as literal text, so the comment renders with visible `##` and stray asterisks. Single `*bold*` in Markdown means *italic* in Jira wiki - never map `**bold**` to `*bold*` by dropping one asterisk mechanically without checking the source was bold, not italic.
120
174
 
121
175
  There is no markdown→Jira-wiki converter program in the pipeline (`lib/` ships `md2confluence-v3.py` for Confluence only, and `scripts/jira-wiki-escape.mjs` covers the emoticon step alone). This conversion table is applied by the model, by hand, which is exactly why it has to be complete.
122
176
 
@@ -149,19 +203,25 @@ Authorization: Bearer $JIRA_TOKEN
149
203
  Content-Type: application/json
150
204
  ```
151
205
 
152
- Body assembled with `jq --rawfile` + `curl --data-binary @file`, from the escaped file that *Emoticon escaping* produced:
206
+ **Post through `jira-publish.sh`. Never hand-roll the `curl`.**
153
207
 
154
208
  ```bash
155
- node "$HOME/.claude/scripts/jira-wiki-escape.mjs" --check /tmp/channels-$TASK_ID-jira-escaped.txt \
156
- || { echo "unescaped Jira emoticon in the body - do not POST" >&2; exit 1; }
157
- jq -n --rawfile body /tmp/channels-$TASK_ID-jira-escaped.txt '{body: $body}' \
158
- > /tmp/channels-$TASK_ID-jira-payload.json
159
- curl -s -X POST -H "Authorization: Bearer $JIRA_TOKEN" \
160
- -H "Content-Type: application/json" \
161
- --data-binary @/tmp/channels-$TASK_ID-jira-payload.json \
162
- "$JIRA_BASE/rest/api/2/issue/$JIRA_ID/comment"
209
+ bash "$HOME/.claude/lib/jira-publish.sh" --issue "$JIRA_ID" \
210
+ --body-file /tmp/channels-$TASK_ID-jira.txt --target comment
163
211
  ```
164
212
 
213
+ That script escapes every body unconditionally (`lib/jira-publish.sh:109`) and
214
+ passes the token through a `-K` config rather than argv, so neither the escape
215
+ nor the token handling depends on an agent remembering a step.
216
+
217
+ This block used to be a hand-written `jq` + `curl` pair with the escape as a
218
+ separate `--check` line above it, and a shipped comment rendered the Swift
219
+ selector `hash(into:)` as `hash(into` plus a smiley - twice. The escaper was
220
+ present, correct, and simply not run on that path: `:)` reached Jira intact and
221
+ Jira's wiki renderer turned it into an emoticon, which is what section
222
+ *Emoticon escaping* below is about. A defence that has to be invoked by hand is
223
+ a defence that is eventually not invoked.
224
+
165
225
  The dispatch summary line includes the comment URL with `?focusedCommentId=...` so the user can paste it into Slack/Teams.
166
226
 
167
227
  ### Writing the issue description (not a comment)
@@ -10,20 +10,31 @@ The PR description targets code reviewers - it stays technical. Every adapter
10
10
 
11
11
  | # | Section key | Heading (`tr`) | Heading (`en`) | Required? |
12
12
  |---|---|---|---|---|
13
- | 1 | `summary` | `## Özet` | `## Summary` | always |
14
- | 2 | `changes` | `## Değişiklikler` | `## Changes` | always |
13
+ | 1 | `summary` | `## Geliştirme Özeti` | `## Development Summary` | always |
14
+ | 2 | `technical` | `## Teknik Açıklama` | `## Technical Explanation` | always |
15
15
  | 3 | `architecture` | `## Mimari Kararlar` | `## Architecture Decisions` | when a non-trivial design choice was made |
16
- | 4 | `verification` | `## Doğrulama` | `## Verification` | always |
17
- | 5 | `visuals` | `## Görsel Kanıt` | `## Visual Evidence` | when `state.visualEvidence.required` |
18
- | 6 | `risk` | `## Risk ve Güvenlik` | `## Risk and Security` | when `state.diffRisk.signals` carries a high-stakes signal (`security_path`, `migration`, `public_api`, `no_test_change`, `test_lines_removed`) |
19
- | 7 | `dependencies` | `## Bağımlılıklar` | `## Dependencies` | when deps added/removed/bumped |
20
- | 8 | `related` | `## İlgili` | `## Related` | always (Jira/issue ref; never `Closes/Fixes`) |
16
+ | 4 | `impact` | `## Etki Analizi` | `## Impact Analysis` | always |
17
+ | 5 | `test_scenarios` | `## Test Senaryoları` | `## Test Scenarios` | always |
18
+ | 6 | `visuals` | `## Görsel Kanıt` | `## Visual Evidence` | when `state.visualEvidence.required` |
19
+ | 7 | `risk` | `## Risk ve Güvenlik` | `## Risk and Security` | when `state.diffRisk.signals` carries a high-stakes signal (`security_path`, `migration`, `public_api`, `no_test_change`, `test_lines_removed`) |
20
+ | 8 | `dependencies` | `## Bağımlılıklar` | `## Dependencies` | when deps added/removed/bumped |
21
+ | 9 | `build` | `## Build` | `## Build` | always |
22
+ | 10 | `related` | `## İlgili` | `## Related` | always (Jira/issue ref; never `Closes/Fixes`) |
23
+
24
+ `summary`, `impact` and `test_scenarios` carry the same three section keys the
25
+ Jira adapter uses, and that pairing is deliberate: the same three questions get
26
+ answered on both surfaces, at the register each reader needs. The PR versions name
27
+ symbols, files, line counts and build shas. The Jira versions name screens and
28
+ behaviour and nothing else (`channels/jira.md`). `technical` has no Jira twin at
29
+ all - it is the section whose absence over there is the point.
21
30
 
22
31
  ### Section content rules
23
32
 
24
- **`summary`** - 1-3 sentences in `outputLanguage`. The "why" of the change. Past tense, no marketing voice. Code identifiers stay verbatim.
33
+ **`summary`** - 2-4 sentences in `outputLanguage`. What broke or what was added, where the user meets it, and the size of the effect when it is measurable (a crash count, a share of a known total, a version range). Past tense, no marketing voice. Code identifiers stay verbatim. This is the one section a reviewer reads before deciding whether to read the rest, so it names the user-visible behaviour before the mechanism.
25
34
 
26
- **`changes`** - bullet list, one item per logically distinct change. Each bullet starts with the touched component and ends with a one-line "what". The source is `$WORKTREE/.pipeline/scope-check.json` `files[].reason` (Phase 3 Step 3.7): a file the dev could not justify there is a file this list cannot describe either, so the bullet quotes the gate output instead of inventing a reason. Use the stack's native file extensions / module paths - the example below shows the **shape**, not a stack lock-in:
35
+ **`technical`** - the account for someone holding the diff: a short paragraph on the mechanism (what the code was actually doing wrong, or what the new code does), then the per-file bullets, then one line of diff stat (`2 files, 8 deletions, 0 insertions`). Where a change is safe for a reason that is not obvious from the diff - an equality relation preserved, an invariant kept, a call site left alone deliberately - that reason belongs here in a sentence, because it is the question the reviewer would otherwise ask in a comment. Tables are welcome when several symbols share a property worth listing side by side.
36
+
37
+ The bullet list is one item per logically distinct change. Each bullet starts with the touched component and ends with a one-line "what". The source is `$WORKTREE/.pipeline/scope-check.json` `files[].reason` (Phase 3 Step 3.7): a file the dev could not justify there is a file this list cannot describe either, so the bullet quotes the gate output instead of inventing a reason. Use the stack's native file extensions / module paths - the example below shows the **shape**, not a stack lock-in:
27
38
 
28
39
  ```markdown
29
40
  ## Changes
@@ -37,20 +48,55 @@ Across stacks the same shape produces, for example: `LoginView.swift - ...` (i
37
48
 
38
49
  **`architecture`** - only when the change involves a non-trivial decision (new abstraction, pattern change, data flow shift, dependency direction). Format: short paragraph stating the decision and the alternative considered. Skip the section entirely for mechanical refactors / dependency bumps / formatting passes.
39
50
 
40
- **`verification`** - what the reviewer should run to confirm the change works. Commands first, manual steps next. Pick the commands for the project's stack - the pipeline supports iOS (Swift/Xcode), Android (Gradle), web (npm/pnpm/yarn), and backend (varies: pytest/jest/go test/etc). Do not hardcode one stack in the body; emit only the commands relevant to the repos touched by this PR.
51
+ **`impact`** - four fixed numbered parts, each answered, never a placeholder. Same four questions as the Jira `impact` section, answered here with the identifiers and numbers that section is not allowed to carry:
52
+
53
+ ```markdown
54
+ ## Impact Analysis
55
+
56
+ **1 - The problem**
57
+ <what was wrong, since which version, with the crash/report identifiers and counts>
58
+
59
+ **2 - What was changed**
60
+ <the change, and why behaviour around it is unchanged>
61
+
62
+ **3 - Affected functions**
63
+ <the screens, flows and symbols that must be tested, including shared components the change reaches>
64
+
65
+ **4 - Effect on other systems**
66
+ <none, or which service, contract or channel>
67
+ ```
68
+
69
+ When the change deliberately fixes part of a wider problem, a closing **Risk and remaining scope** paragraph names what is still open and why it was left - a reviewer who can see the rest of the pattern in the repo will ask otherwise, and the honest answer is cheaper written down than defended in a thread.
41
70
 
42
- Skeleton (the adapter fills the body with the actual stack-appropriate lines at write-time; the template lists the shape only):
71
+ **`test_scenarios`** - the same titled-scenario shape the Jira adapter uses, so the tester reads one list on both surfaces, with symbols allowed here:
43
72
 
44
73
  ```markdown
45
- ## Verification
74
+ ## Test Scenarios
46
75
 
47
- - Build: <stack-appropriate build command>
48
- - Tests: <stack-appropriate test command, scoped to the touched targets/packages>
49
- - Lint / type-check: <if the project has one>
50
- - Manual: <one or more user-facing steps the reviewer can follow without setup>
76
+ **1. <what this scenario exercises>**
77
+ 1. <step>
78
+ 2. **Expected:** <observable outcome>
79
+
80
+ **2. <regression scenario>**
81
+ 1. <step>
82
+ 2. **Expected:** <what should still behave as before>
83
+ ```
84
+
85
+ Reproduction scenarios first, then regressions for whatever the change could have disturbed. When a UI test target covers a scenario, say so on the scenario line and let `## Build` carry the run result - a scenario a machine already ran is not the same request as one a human has to perform, and conflating them wastes the reviewer's time.
86
+
87
+ **`build`** - what was built, on what, and what came out. Not a promise that it builds; the recorded result of the run that happened:
88
+
89
+ ```markdown
90
+ ## Build
91
+
92
+ - <build command or scheme> on <base branch>@<sha>: BUILD SUCCEEDED, 0 errors
93
+ - Tests: <test command>: <N> passed, <M> failed
94
+ - UI tests: <target>: <status> | not run - <reason from state.uiTest.notRunReason>
51
95
  ```
52
96
 
53
- Multi-repo PRs (one PR per repo) emit verification commands for that repo's stack only - never mix iOS + Android commands into a single PR body.
97
+ The base sha matters because "it builds" is a claim about a merge base, and the reviewer's local tree is usually not that one. `state.uiTest` supplies the third line verbatim, including its `notRunReason` - "no UI test target in this repo" is a result, not a gap, and writing it stops the same question being asked on every PR.
98
+
99
+ Pick commands for the project's stack - the pipeline supports iOS (Swift/Xcode), Android (Gradle), web (npm/pnpm/yarn) and backend (pytest/jest/go test/etc). Multi-repo PRs (one PR per repo) emit the commands for that repo's stack only - never mix iOS + Android commands into a single PR body.
54
100
 
55
101
  **`risk`** - only when `state.diffRisk.signals` (Phase 4 Step 1.75) contains a high-stakes signal. Four fixed lines, each answered, never left as a placeholder; the source is Phase 1 `touchedAreas` plus the signals themselves, and when a signal is present the absence of this section is a Phase 6 Step 3 blocker:
56
102
 
@@ -136,7 +182,7 @@ Never use `Closes #N`, `Fixes #N`, `Resolves PROJ-X`. Issues require 4-approval
136
182
 
137
183
  ```
138
184
  1. Read agent-state.json (taskId, contextLinks, identity, language).
139
- 2. Build section bodies in markdown - summary first, then in the table order, skipping conditional sections that don't apply. Section order is fixed: `summary` → `changes` → `architecture` (cond.) → `verification` → `risk` (cond.) → `dependencies` (cond.) → `related`.
185
+ 2. Build section bodies in markdown - summary first, then in the table order, skipping conditional sections that don't apply. Section order is fixed: `summary` → `technical` → `architecture` (cond.) → `impact` → `test_scenarios` → `visuals` (cond.) → `risk` (cond.) → `dependencies` (cond.) → `build` → `related`.
140
186
  3. Run the assembled body through the `humanizer` skill.
141
187
  4. Apply Multi-repo cross-links (## Related PRs prepend when projects.length > 1).
142
188
  5. Dispatch per the Behaviour-by-remote table.
@@ -220,7 +266,7 @@ The Bitbucket REST API returns `409 Conflict` if `version` is stale. Adapter beh
220
266
  ## Hard rules (must not regress)
221
267
 
222
268
  - Real newlines, no HTML entities - heredoc + `jq --rawfile` + `curl --data-binary @file`. Never embed `\n` literally; Bitbucket stores the literal `\n` characters.
223
- - Section order is fixed: `summary` → `changes` → `architecture` (cond.) → `verification` → `dependencies` (cond.) → `related`. Conditional sections may be omitted but never reordered or inserted between fixed ones.
269
+ - Section order is fixed: `summary` → `technical` → `architecture` (cond.) → `impact` → `test_scenarios` → `visuals` (cond.) → `risk` (cond.) → `dependencies` (cond.) → `build` → `related`. Conditional sections may be omitted but never reordered or inserted between fixed ones.
224
270
  - Humanizer pass runs **after** body assembly and **before** dispatch - every body line carries technical, non-AI tone. References at least one symbol/file/line drawn from the diff or pipeline log.
225
271
  - Body content language follows `prefs.global.outputLanguage`. Code identifiers, file paths, branch names, PR titles, commands, type names, and `Closes/Fixes`-style keywords stay verbatim English. The template file (this doc) is English because `promptLanguage="en"` is locked; the body is rendered in the user's language at write-time.
226
272
  - No body markers (`<!-- channels:start -->`) - channels does a full replace each run.
@@ -6,13 +6,14 @@
6
6
 
7
7
  ---
8
8
 
9
- ## 1. Command Inventory (53 commands)
9
+ ## 1. Command Inventory (56 commands)
10
10
 
11
11
  ```
12
- analysis, analysis-jira, analysis-resolve, autopilot, build-optimize, channels,
13
- complaint-analysis, create-jira, design-check, doctor, diff-explain, feedback,
14
- forget, garbage-collect, graph, help, ios-coding-standard, issue, jira,
15
- kill, language, local, local-autopilot, log, manual-test, prune-logs,
12
+ analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
13
+ autopilot-on, autopilot-status, build-optimize, channels, complaint-analysis,
14
+ create-jira, design-check, diff-explain, doctor, feedback, forget,
15
+ garbage-collect, graph, help, ios-coding-standard, issue, jira, kill,
16
+ language, local, local-autopilot, log, manual-test, prune-logs,
16
17
  prune-prompts, purge, refactor, resume, resume-local, review,
17
18
  review-analysis, review-issue, review-jira, routines, save, scan, search,
18
19
  setup, stack, status, steer, store-ready, sync, test, test-accessibility,