@mmerterden/multi-agent-pipeline 17.3.0 → 17.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +203 -0
- package/README.md +23 -5
- package/README.tr.md +23 -5
- package/docs/adr/0013-lsp-code-intelligence.md +102 -0
- package/docs/adr/README.md +1 -0
- package/docs/token-budget-history.md +1 -1
- package/install/templates/copilot-instructions.md +9 -3
- package/package.json +1 -1
- package/pipeline/agents/code-reviewer.md +35 -1
- package/pipeline/commands/multi-agent/analysis/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +5 -3
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/local/SKILL.md +17 -6
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +3 -3
- package/pipeline/lib/multi-repo-pipeline.sh +26 -0
- package/pipeline/multi-agent-refs/analysis/locked.md +4 -4
- package/pipeline/multi-agent-refs/analysis/render.md +2 -1
- package/pipeline/multi-agent-refs/channels/pr.md +26 -0
- package/pipeline/multi-agent-refs/cross-cli-contract.md +22 -0
- package/pipeline/multi-agent-refs/features/base-branch-evidence.md +222 -0
- package/pipeline/multi-agent-refs/features/code-graph.md +40 -0
- package/pipeline/multi-agent-refs/features/code-intelligence.md +80 -0
- package/pipeline/multi-agent-refs/features/design-conformance.md +14 -0
- package/pipeline/multi-agent-refs/features/review-file-set.md +132 -0
- package/pipeline/multi-agent-refs/phases/modes.md +23 -3
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +96 -71
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +31 -23
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +1 -1
- package/pipeline/multi-agent-refs/phases.md +7 -2
- package/pipeline/multi-agent-refs/picker-contract.md +37 -5
- package/pipeline/multi-agent-refs/tracker-contract.md +25 -14
- package/pipeline/schemas/agent-state.schema.json +88 -4
- package/pipeline/schemas/prefs.schema.json +22 -0
- package/pipeline/schemas/review-file-exclusions.json +137 -0
- package/pipeline/schemas/reviewer-output.schema.json +27 -1
- package/pipeline/schemas/token-budget.json +2 -2
- package/pipeline/scripts/autopilot-runner.mjs +292 -45
- package/pipeline/scripts/base-branch-candidates.mjs +599 -0
- package/pipeline/scripts/diff-risk-score.mjs +1 -36
- package/pipeline/scripts/gc-abandoned.sh +5 -3
- package/pipeline/scripts/gen-mode-dispatch.mjs +39 -16
- package/pipeline/scripts/git-path.mjs +63 -0
- package/pipeline/scripts/glob-match.mjs +62 -0
- package/pipeline/scripts/graph-mermaid.mjs +251 -0
- package/pipeline/scripts/phase-tracker.sh +39 -2
- package/pipeline/scripts/phase0-exit-gate.mjs +128 -0
- package/pipeline/scripts/review-file-filter.mjs +180 -0
- package/pipeline/scripts/skill-conformance.mjs +1 -31
- package/pipeline/scripts/validate-analysis-doc.mjs +53 -0
- package/pipeline/scripts/validate-reviewer.mjs +90 -1
- package/pipeline/scripts/verify-citations.mjs +428 -0
- package/pipeline/skills/.skill-manifest.json +2 -2
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +1 -1
|
@@ -25,13 +25,15 @@ Return ONLY a JSON object:
|
|
|
25
25
|
{
|
|
26
26
|
"findings": [{"severity": "blocking|important|suggestion", "file": "...", "line": N, "issue": "...", "fix": "...", "ruleId": "SEC-01", "criteriaSource": "ios-coding-standard"}],
|
|
27
27
|
"conformance": [{"ruleId": "SEC-01", "verdict": "conformant|violated|not-applicable", "file": "Sources/X.swift", "line": 42, "reason": "..."}],
|
|
28
|
+
"fileCoverage": [{"path": "Sources/X.swift", "verdict": "reviewed|skipped", "reason": "..."}],
|
|
28
29
|
"approved": true|false
|
|
29
30
|
}
|
|
30
31
|
```
|
|
31
32
|
|
|
32
33
|
`ruleId` + `criteriaSource` are required on a finding that comes from a cited
|
|
33
34
|
rule and omitted otherwise. `conformance` is required whenever a `${CRITERIA}`
|
|
34
|
-
block was supplied -
|
|
35
|
+
block was supplied, and `fileCoverage` whenever a `${REVIEW_FILES}` block was -
|
|
36
|
+
see below.
|
|
35
37
|
|
|
36
38
|
## Severity Classification
|
|
37
39
|
|
|
@@ -90,6 +92,38 @@ obvious breach; `judgement` rules are yours, and each needs the measurement its
|
|
|
90
92
|
When no `${CRITERIA}` block is supplied, omit `conformance` entirely and review
|
|
91
93
|
on the focus areas above.
|
|
92
94
|
|
|
95
|
+
## Review file set (required when supplied)
|
|
96
|
+
|
|
97
|
+
The orchestrator may pass a `${REVIEW_FILES}` block: the changed files you are
|
|
98
|
+
accountable for, computed before you were dispatched. Generated output,
|
|
99
|
+
lockfiles, recorded snapshots, vendored source and binaries are already removed,
|
|
100
|
+
so what is left is human-authored code and the list is not negotiable.
|
|
101
|
+
|
|
102
|
+
```
|
|
103
|
+
${REVIEW_FILES}
|
|
104
|
+
reviewed:
|
|
105
|
+
<repo-relative path>
|
|
106
|
+
...
|
|
107
|
+
excluded (not yours, listed so you can see the decision):
|
|
108
|
+
<repo-relative path> - <reason>
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Return one `fileCoverage` row **per path under `reviewed`**, and no rows for
|
|
112
|
+
paths outside it. Rules:
|
|
113
|
+
|
|
114
|
+
- `reviewed` means you read the file's diff in full. There is no `partial`: what
|
|
115
|
+
you read but did not understand belongs in a `findings` entry, not in a
|
|
116
|
+
hedged verdict.
|
|
117
|
+
- `skipped` requires a `reason` naming why you could not read it - the diff was
|
|
118
|
+
truncated, the content was unreadable. "Not relevant" is not a skip reason; a
|
|
119
|
+
file you read and found nothing in is `reviewed`.
|
|
120
|
+
- A path that is neither read nor explicitly skipped fails the stage, for the
|
|
121
|
+
same reason an unanswered rule ID does: a reviewer that opened one file of ten
|
|
122
|
+
and one that read all ten return identical empty `findings` arrays, and this
|
|
123
|
+
list is what tells them apart.
|
|
124
|
+
|
|
125
|
+
When no `${REVIEW_FILES}` block is supplied, omit `fileCoverage` entirely.
|
|
126
|
+
|
|
93
127
|
## Priority Files (advisory)
|
|
94
128
|
|
|
95
129
|
When the orchestrator passes a `${PRIORITY_FILES}` block, treat it as a heuristic
|
|
@@ -169,10 +169,10 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
|
169
169
|
|
|
170
170
|
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
171
171
|
|
|
172
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate
|
|
172
|
+
**TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
173
173
|
|
|
174
174
|
```text
|
|
175
|
-
#
|
|
175
|
+
# Register one tile per phase, capture the taskId, persist it:
|
|
176
176
|
for each phase in 0:Init, 1:Analysis, 2:Planning, 4:Review, 6:Commit, 7:Report:
|
|
177
177
|
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
178
178
|
-> returns taskId
|
|
@@ -194,7 +194,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
|
194
194
|
|
|
195
195
|
#### TaskCreate ordering (strict)
|
|
196
196
|
|
|
197
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `analysis` that means: Phase 0 → Phase 1 → Phase 2 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
197
|
+
**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `analysis` that means: Phase 0 → Phase 1 → Phase 2 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
198
198
|
|
|
199
199
|
### Visual channel - Copilot CLI / plain shell
|
|
200
200
|
|
|
@@ -83,10 +83,10 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
|
83
83
|
|
|
84
84
|
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
85
85
|
|
|
86
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate
|
|
86
|
+
**TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
87
87
|
|
|
88
88
|
```text
|
|
89
|
-
#
|
|
89
|
+
# Register one tile per phase, capture the taskId, persist it:
|
|
90
90
|
for each phase in 0:Init, 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report:
|
|
91
91
|
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
92
92
|
-> returns taskId
|
|
@@ -108,7 +108,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
|
108
108
|
|
|
109
109
|
#### TaskCreate ordering (strict)
|
|
110
110
|
|
|
111
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
111
|
+
**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
112
112
|
|
|
113
113
|
### Visual channel - Copilot CLI / plain shell
|
|
114
114
|
|
|
@@ -45,9 +45,11 @@ pkill -f "$HOME/.claude/autopilot/bin/menubar" 2>/dev/null || true
|
|
|
45
45
|
|
|
46
46
|
Without `--now` nothing here runs. With it, the child session is stopped and the
|
|
47
47
|
run is marked `abandoned`, its worktree is removed **unless it holds uncommitted
|
|
48
|
-
work**, in which case the work
|
|
49
|
-
|
|
50
|
-
is
|
|
48
|
+
work**, in which case the work goes into a stash entry labelled
|
|
49
|
+
`autopilot/abandoned/<task-id>` (find it with `git stash list`) and the worktree
|
|
50
|
+
is kept. Same contract as `gc-abandoned.sh` and `autopilot-runner.mjs`; losing a
|
|
51
|
+
day of edits is worse than 750 MB. No branch is created - one made after
|
|
52
|
+
`stash push` would point at HEAD and contain none of the work.
|
|
51
53
|
|
|
52
54
|
### 5. Keep the selection
|
|
53
55
|
|
|
@@ -115,7 +115,7 @@ deletes nothing until you confirm.
|
|
|
115
115
|
| Line | Meaning |
|
|
116
116
|
|---|---|
|
|
117
117
|
| `would remove` / `would remove worktree` | reapable: stopped past its age limit, or finished with its worktree left behind |
|
|
118
|
-
| `would stash` | uncommitted work found - it
|
|
118
|
+
| `would stash` | uncommitted work found - it goes into a stash entry labelled `autopilot/abandoned/<task-id>` (`git stash list`) and the worktree is **kept** |
|
|
119
119
|
| `would mark abandoned` | the state is stale but there is no worktree left to remove |
|
|
120
120
|
| `unattributed` | a worktree with no run state. **Never removed, whatever the flags.** `.worktrees/` is not exclusively ours, and nothing distinguishes a hand-made worktree from a pipeline one whose log was pruned. Report the list and its size; the user decides |
|
|
121
121
|
|
|
@@ -67,10 +67,21 @@ Two channels run in parallel at every phase boundary:
|
|
|
67
67
|
```bash
|
|
68
68
|
# Phase 0, very first shell call (every CLI):
|
|
69
69
|
bash $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
|
|
70
|
-
|
|
70
|
+
bash $HOME/.claude/scripts/phase-tracker.sh add 0 "Init"
|
|
71
|
+
bash $HOME/.claude/scripts/phase-tracker.sh tiles
|
|
72
|
+
bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
|
|
73
|
+
|
|
74
|
+
# Phase 0 Step 7.5, immediately after the depth answer - the first moment this
|
|
75
|
+
# mode knows its phase set. Full:
|
|
76
|
+
for p in "1:Analysis" "2:Planning" "3:Dev" "4:Review" "6:Commit" "7:Report"; do
|
|
71
77
|
bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
72
78
|
done
|
|
73
|
-
|
|
79
|
+
# Short (Analysis and Planning are not run, so they get no tile at all):
|
|
80
|
+
for p in "3:Dev" "4:Review" "6:Commit" "7:Report"; do
|
|
81
|
+
bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
82
|
+
done
|
|
83
|
+
# Then the widget, narrowed to the phases that do not have a tile yet:
|
|
84
|
+
bash $HOME/.claude/scripts/phase-tracker.sh tiles --new
|
|
74
85
|
|
|
75
86
|
# Every phase boundary (every CLI):
|
|
76
87
|
bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress|completed|failed|skipped
|
|
@@ -83,11 +94,11 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
|
83
94
|
|
|
84
95
|
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
85
96
|
|
|
86
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is
|
|
97
|
+
**TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. This mode registers in two batches (Step -1, then Step 7.5), so `tiles --new` narrows the second one and the Phase 0 tile is never created twice. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
87
98
|
|
|
88
99
|
```text
|
|
89
|
-
#
|
|
90
|
-
for each phase in 0:Init, 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report:
|
|
100
|
+
# Register one tile per phase, capture the taskId, persist it:
|
|
101
|
+
for each phase in 0:Init at Step -1, then 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report (Full) or 3:Dev, 4:Review, 6:Commit, 7:Report (Short) at Step 7.5:
|
|
91
102
|
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
92
103
|
-> returns taskId
|
|
93
104
|
bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"
|
|
@@ -108,7 +119,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
|
108
119
|
|
|
109
120
|
#### TaskCreate ordering (strict)
|
|
110
121
|
|
|
111
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
122
|
+
**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local` that means: Phase 0 at Step -1, then the rest in ascending order at Step 7.5 (Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7 minus whatever the depth answer drops). The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
112
123
|
|
|
113
124
|
### Visual channel - Copilot CLI / plain shell
|
|
114
125
|
|
|
@@ -104,10 +104,10 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
|
104
104
|
|
|
105
105
|
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
106
106
|
|
|
107
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate
|
|
107
|
+
**TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
108
108
|
|
|
109
109
|
```text
|
|
110
|
-
#
|
|
110
|
+
# Register one tile per phase, capture the taskId, persist it:
|
|
111
111
|
for each phase in 0:Init, 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report:
|
|
112
112
|
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
113
113
|
-> returns taskId
|
|
@@ -129,7 +129,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
|
129
129
|
|
|
130
130
|
#### TaskCreate ordering (strict)
|
|
131
131
|
|
|
132
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
132
|
+
**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
133
133
|
|
|
134
134
|
### Visual channel - Copilot CLI / plain shell
|
|
135
135
|
|
|
@@ -292,6 +292,12 @@ agent_state = {
|
|
|
292
292
|
"worktreePath": os.path.join(worktree_root, primary["name"]),
|
|
293
293
|
"branch": state["branch"],
|
|
294
294
|
"baseBranch": state.get("baseBranch", "develop"),
|
|
295
|
+
"baseBranchSource": state.get("baseBranchSource", "asked"),
|
|
296
|
+
# The multi-repo bridge always builds worktrees - that is what it is for -
|
|
297
|
+
# so the workspace was never in question here. It is recorded anyway,
|
|
298
|
+
# because phase0-exit-gate.mjs refuses to close Phase 0 on a state file
|
|
299
|
+
# that cannot say who decided.
|
|
300
|
+
"workspaceSource": state.get("workspaceSource", "command"),
|
|
295
301
|
"remoteType": primary.get("provider", "github"),
|
|
296
302
|
"currentPhase": 0,
|
|
297
303
|
"status": "in_progress",
|
|
@@ -308,6 +314,26 @@ related = issue.get("relatedIssues") or []
|
|
|
308
314
|
if isinstance(related, list) and related:
|
|
309
315
|
agent_state["relatedIssues"] = related
|
|
310
316
|
|
|
317
|
+
# Always written, `[]` included: the empty array is the record that the
|
|
318
|
+
# dev-context step ran, and phase0-exit-gate.mjs refuses to close Phase 0
|
|
319
|
+
# without it. Only repos this run will NOT modify belong here - the extras that
|
|
320
|
+
# get a worktree are already in projects[].
|
|
321
|
+
STACKS = {"ios", "android", "node", "python", "go"}
|
|
322
|
+
worktreed = {r.get("name") for r in all_repos}
|
|
323
|
+
siblings = []
|
|
324
|
+
for r in (state.get("readonlySiblings") or []) + extras:
|
|
325
|
+
name = r.get("name")
|
|
326
|
+
if not name or name in worktreed:
|
|
327
|
+
continue
|
|
328
|
+
stack = r.get("stack", "unknown")
|
|
329
|
+
siblings.append({
|
|
330
|
+
"name": name,
|
|
331
|
+
"root": r.get("localPath") or None,
|
|
332
|
+
"stack": stack if stack in STACKS else "unknown",
|
|
333
|
+
"canPush": bool(r.get("canPush", False)),
|
|
334
|
+
})
|
|
335
|
+
agent_state["siblings"] = siblings
|
|
336
|
+
|
|
311
337
|
logs_dir = os.path.expanduser(os.path.join("~/.claude/logs/multi-agent", primary["name"], state["taskId"]))
|
|
312
338
|
os.makedirs(logs_dir, exist_ok=True)
|
|
313
339
|
out_path = os.path.join(logs_dir, "agent-state.json")
|
|
@@ -1,16 +1,16 @@
|
|
|
1
|
-
# Locked decisions (
|
|
1
|
+
# Locked decisions (37)
|
|
2
2
|
|
|
3
|
-
> The
|
|
3
|
+
> The 37 Locked decisions of the analysis flow. Loaded by `/multi-agent:analysis`, by `/multi-agent:analysis-resolve` (which inherits them) and by pipeline Phase 1 when it runs the analysis engine. Numbering is canonical: cite as `Locked <n> (<short label>)`.
|
|
4
4
|
|
|
5
5
|
### Index by category (v9.1.0+)
|
|
6
6
|
|
|
7
|
-
Browse-friendly grouping of the
|
|
7
|
+
Browse-friendly grouping of the 37 Locked decisions. Numbering stays canonical (matches the list below); the index is read-only navigation.
|
|
8
8
|
|
|
9
9
|
| Category | Decisions | Concern |
|
|
10
10
|
|---|---|---|
|
|
11
11
|
| **A. Governance** | 1, 5, 6, 7, 10, 26, 27, 32, 36 | Run-level process rules: one feature per run, default output, auto-commit ban, punctuation policy, output picker timing, Pass B preview, evidence digest cache, analysis profile, document reviewed before publish |
|
|
12
12
|
| **B. Citation and Evidence** | 3, 4, 8, 11, 24, 30, 34 | Every fact in the doc traces back to a source: citation discipline, forward-looking spec, standards binding, repo-evidence reuse-first, Pass B footnote mandatory, analysis self-contained (pipeline-wide), references built from the evidence record |
|
|
13
|
-
| **C. Output Format and Structure** | 2, 9, 13, 14, 16, 17, 20, 21, 25, 33, 35 | How the document is laid out: section omission rule, per-platform output split, Gherkin user stories, Goals + Non-Goals paired, Files-to-Add tag, API response variants exhaustive, localization mode (ownership-aware), References at the bottom, Lite mode, corporate backbone always renders, stack-optional render |
|
|
13
|
+
| **C. Output Format and Structure** | 2, 9, 13, 14, 16, 17, 20, 21, 25, 33, 35, 37 | How the document is laid out: section omission rule, per-platform output split, Gherkin user stories, Goals + Non-Goals paired, Files-to-Add tag, API response variants exhaustive, localization mode (ownership-aware), References at the bottom, Lite mode, corporate backbone always renders, stack-optional render, redesign records v1 before planning v2 |
|
|
14
14
|
| **D. Design Source and Pipeline Architecture** | 12, 22, 23 | Where design comes from and how the pipeline renders: Figma 3-tier access (BLOCKING), platform-agnostic template + Pass B render, convention extraction (Phase 1c) |
|
|
15
15
|
| **E. UI, Variant, and Test Coverage** | 15, 18, 19, 28, 29, 31 | UI artefact rules: SVG default for new assets, screenshots embedded, all Figma variants drilled, SwiftUI Preview block (iOS), variant usage explicit, business-rule to acceptance-criterion to test traceability |
|
|
16
16
|
|
|
@@ -124,10 +124,11 @@ Result: `state.analysisSpec.outputs.requested[]`.
|
|
|
124
124
|
for f in /tmp/analysis-<feature-slug>-<ts>/*.md; do
|
|
125
125
|
node "$HOME/.claude/scripts/validate-analysis-doc.mjs" "$f" || GATE_FAILED=1
|
|
126
126
|
node "$HOME/.claude/scripts/build-references.mjs" <state.json> --check "$f" || GATE_FAILED=1
|
|
127
|
+
node "$HOME/.claude/scripts/verify-citations.mjs" "$f" --repo "$REPO_ROOT" || GATE_FAILED=1
|
|
127
128
|
done
|
|
128
129
|
```
|
|
129
130
|
|
|
130
|
-
`validate-analysis-doc.mjs` enforces the mechanically-checkable Locked decisions on the emitted markdown itself (front-matter completeness, never-omitted sections per Locked 2, humanizer punctuation per Locked 7, Full-mode business-rule traceability per Locked 31, and in the corporate profile the backbone presence and `EKLENECEK`-to-Section-20 pairing per Locked 33). `build-references.mjs --check` runs the References coverage gate (Locked 34): a source the run consumed but did not list, or a listed row with no evidence behind it, blocks dispatch. Any ERROR blocks dispatch: fix the draft and re-validate. Warnings are advisory (run with `--strict` to treat them as blocking). This turns the "fails the dispatch gate" prose into a real, model-independent check.
|
|
131
|
+
`validate-analysis-doc.mjs` enforces the mechanically-checkable Locked decisions on the emitted markdown itself (front-matter completeness, never-omitted sections per Locked 2, humanizer punctuation per Locked 7, Full-mode business-rule traceability per Locked 31, and in the corporate profile the backbone presence and `EKLENECEK`-to-Section-20 pairing per Locked 33). `build-references.mjs --check` runs the References coverage gate (Locked 34): a source the run consumed but did not list, or a listed row with no evidence behind it, blocks dispatch. `verify-citations.mjs` resolves each claimed `file:line` at HEAD with `git cat-file`, so `Foo.swift:9999` in a repo with no Foo.swift blocks dispatch (Locked 3, Locked 37). Any ERROR blocks dispatch: fix the draft and re-validate. Warnings are advisory (run with `--strict` to treat them as blocking). This turns the "fails the dispatch gate" prose into a real, model-independent check.
|
|
131
132
|
|
|
132
133
|
Iterate `state.analysisSpec.outputs.requested`. For each target:
|
|
133
134
|
|
|
@@ -77,6 +77,32 @@ Across stacks the same shape produces, for example: `LoginView.swift - ...` (i
|
|
|
77
77
|
<none, or which service, contract or channel>
|
|
78
78
|
```
|
|
79
79
|
|
|
80
|
+
Part 3 is the one part of this body that is MEASURED rather than recalled. The
|
|
81
|
+
code graph already answers it, so draw the answer instead of re-typing it:
|
|
82
|
+
|
|
83
|
+
```bash
|
|
84
|
+
node "$HOME/.claude/scripts/graph-mermaid.mjs" "<changed symbol[,symbol]>"
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
Append the fenced block it prints under part 3, above the prose. GitHub renders
|
|
88
|
+
mermaid natively in pull requests, so this costs no renderer and no plugin. The
|
|
89
|
+
prose stays: the diagram says which symbols the change reaches, the sentence says
|
|
90
|
+
which screens and flows a tester must open, and neither answers the other.
|
|
91
|
+
|
|
92
|
+
Exit 1 means the repo has no graph yet (`/multi-agent:graph` builds it) or the
|
|
93
|
+
symbol is not in it. That is a gap with a reason, not a failure: write the prose
|
|
94
|
+
alone and say the graph was unavailable. Never hand-draw the diagram - a drawn
|
|
95
|
+
blast radius nobody measured is worse than none, because a diagram is read as
|
|
96
|
+
fact.
|
|
97
|
+
|
|
98
|
+
The commit line the script prints stays with it. A graph built before the change
|
|
99
|
+
draws the radius of an older tree, and the reader has no other way to notice.
|
|
100
|
+
|
|
101
|
+
This is a GitHub-only section. `channels/jira.md` has no mermaid handling at all:
|
|
102
|
+
a fence there converts to a literal `{code:mermaid}` block, so the Jira impact
|
|
103
|
+
section keeps its prose. Confluence renders it through the `ac:name="mermaid"`
|
|
104
|
+
macro (`md2confluence-v3.py`) when the space has the plugin.
|
|
105
|
+
|
|
80
106
|
When the change deliberately fixes part of a wider problem, a closing **Risk and remaining scope** paragraph names what is still open and why it was left - a reviewer who can see the rest of the pattern in the repo will ask otherwise, and the honest answer is cheaper written down than defended in a thread.
|
|
81
107
|
|
|
82
108
|
**`test_scenarios`** - the same titled-scenario shape the Jira adapter uses, so the tester reads one list on both surfaces, with symbols allowed here:
|
|
@@ -214,6 +214,28 @@ on skill directories would demand exactly the layout that breaks it.
|
|
|
214
214
|
|
|
215
215
|
Future changes that break an item in the "stay identical" list must update **both** files in the same commit. `smoke-cross-cli-behavior.sh` enforces the identity-preserving axis (input parsing, routing, output shape); structural differences are left to manual review because enforcing them would require forcing the files to the same shape, which we intentionally don't want.
|
|
216
216
|
|
|
217
|
+
### Panel diversity per host
|
|
218
|
+
|
|
219
|
+
Phase 4 runs three reviewers everywhere, but the diversity those three buy is not the
|
|
220
|
+
same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
|
|
221
|
+
beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
|
|
222
|
+
Anthropic models on one, three OpenAI models on the other - so the same three-way
|
|
223
|
+
agreement is weaker evidence there, and Phase 4 says so in the triage note on a
|
|
224
|
+
borderline finding.
|
|
225
|
+
|
|
226
|
+
Where the budget goes instead, when vendor diversity is unavailable:
|
|
227
|
+
|
|
228
|
+
| Host | Reviewer 1 | Reviewer 2 | Reviewer 3 |
|
|
229
|
+
|---|---|---|---|
|
|
230
|
+
| Copilot CLI | Fable/Opus, security + architecture | GPT-5.4, edge cases (cross-vendor) | Sonnet, quality |
|
|
231
|
+
| Claude Code | Fable, security + architecture | Opus, edge cases | Sonnet, quality |
|
|
232
|
+
| Codex | `xhigh`, security + architecture | a different family member, edge cases | `medium`, quality |
|
|
233
|
+
|
|
234
|
+
On Codex the axis is reasoning effort as much as model identity, because the family
|
|
235
|
+
members available there are closer to each other than Fable and Sonnet are. That is a
|
|
236
|
+
weaker axis, not an equivalent one, and treating it as equivalent is the error this
|
|
237
|
+
section exists to prevent.
|
|
238
|
+
|
|
217
239
|
## 3. Frontmatter Transform Rules (Claude ↔ Copilot)
|
|
218
240
|
|
|
219
241
|
Each file has a different frontmatter schema. The sync flow transforms between them:
|
|
@@ -0,0 +1,222 @@
|
|
|
1
|
+
# Feature: Base-Branch Evidence
|
|
2
|
+
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [1. The fetch is a fact, not a formality](#1-the-fetch-is-a-fact-not-a-formality)
|
|
5
|
+
- [2. Evidence sources](#2-evidence-sources)
|
|
6
|
+
- [2b. A filter is not a ranking](#2b-a-filter-is-not-a-ranking)
|
|
7
|
+
- [3. The convention is learned from the remote](#3-the-convention-is-learned-from-the-remote)
|
|
8
|
+
- [4. The picker](#4-the-picker)
|
|
9
|
+
- [5. Autopilot](#5-autopilot)
|
|
10
|
+
- [6. State](#6-state)
|
|
11
|
+
<!-- /toc -->
|
|
12
|
+
|
|
13
|
+
**Pattern**: Phase 0 Step 3 used to ask one question with a list it could not
|
|
14
|
+
vouch for. `git fetch origin` ran, its exit code was ignored, and `git branch -r`
|
|
15
|
+
printed the remote-tracking cache either way - so on a restricted network a
|
|
16
|
+
weeks-old local list was presented as the remote's answer, with nothing in the
|
|
17
|
+
output saying so. And the answer was usually derivable: an issue that carries a
|
|
18
|
+
target version, or that links a separate issue representing the release, already
|
|
19
|
+
names the branch on a repo whose release branches encode the version. Nothing
|
|
20
|
+
derived it.
|
|
21
|
+
|
|
22
|
+
The shape here is **evidence collection, then a picker** - deliberately not a
|
|
23
|
+
rule engine. Every candidate carries why it is a candidate; the ranking orders
|
|
24
|
+
them; a human chooses. A rule engine would have to be right, and the inputs
|
|
25
|
+
(field names, branch spellings, board conventions) differ per repo and change
|
|
26
|
+
under it. An evidence list only has to be honest.
|
|
27
|
+
|
|
28
|
+
Implemented by `$HOME/.claude/scripts/base-branch-candidates.mjs` (pure, no git,
|
|
29
|
+
no network - it is handed ref lists and issue evidence). Gated by `prefs.global.baseBranchEvidence.enabled`
|
|
30
|
+
(default `true`; off falls back to the rule-6 sort order, still through the picker). Asserted by `smoke-base-branch-evidence.sh`
|
|
31
|
+
and `test/base-branch-candidates.test.mjs`.
|
|
32
|
+
|
|
33
|
+
## 1. The fetch is a fact, not a formality
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
git -C "$PROJECT_ROOT" fetch origin; FETCH_RC=$?
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
`FETCH_RC` decides `refProvenance`, which is a required input to the collector
|
|
40
|
+
and rides on every candidate as its own evidence row:
|
|
41
|
+
|
|
42
|
+
| `FETCH_RC` | `refProvenance` | Ref list | What the user is told |
|
|
43
|
+
|---|---|---|---|
|
|
44
|
+
| 0 | `remote` | `git branch -r` | nothing extra; the list is current |
|
|
45
|
+
| non-zero | `local` | `git branch -r` (cache) **and** `git branch` | the picker question itself says the fetch failed and these refs may be stale |
|
|
46
|
+
|
|
47
|
+
A degraded list is still a list - falling back is correct, hiding it is not. The
|
|
48
|
+
degraded picker gains a **Retry the fetch** row, and the fetch-fail picker
|
|
49
|
+
already in Step 3 (Connect VPN / cached / local branch / abort) still owns the
|
|
50
|
+
`baseFetchStatus` value. The two fit together: that picker records *what the run
|
|
51
|
+
is working from*, this one records *what the candidate list is worth*.
|
|
52
|
+
|
|
53
|
+
Local-only refs never silently become remote ones. `state.baseBranchEvidence.refProvenance`
|
|
54
|
+
must be `local` whenever `baseFetchStatus` is `cached-stale` or `local-branch`,
|
|
55
|
+
and `phase0-exit-gate.mjs` refuses to close Phase 0 otherwise. That assertion is
|
|
56
|
+
the whole of problem 1: the run may degrade, it may not misreport.
|
|
57
|
+
|
|
58
|
+
## 2. Evidence sources
|
|
59
|
+
|
|
60
|
+
| Kind | Where it comes from | Weight |
|
|
61
|
+
|---|---|---|
|
|
62
|
+
| `issue-version` | a version-typed field on the issue whose value matches a branch on the ref list | 100 exact, 60 same major.minor, x0.6 for an affects-version field |
|
|
63
|
+
| `linked-release` | a linked issue or parent whose own fix-version or summary names a version that matches a branch | 90 exact (110 when `baseBranchEvidence.preferLinkedRelease`) |
|
|
64
|
+
| `version-convention` | the release-branch template inferred from the ref list | annotation only, no points |
|
|
65
|
+
| `recent` | `prefs.global.recentBranches[{projectKey}]`, inside `settings.branchTtlDays` | 40 + up to 10 for recency |
|
|
66
|
+
| `repo-default` | `origin/HEAD` | 20 |
|
|
67
|
+
| `sort-order` | the develop / release / main families Step 3 always sorted by | 15 / 10 / 5 |
|
|
68
|
+
| `ref-provenance` | the fetch outcome above | annotation only, on every candidate |
|
|
69
|
+
|
|
70
|
+
### Field discovery, not a field table
|
|
71
|
+
|
|
72
|
+
No field id is hardcoded, and none can be: a board's "target version" is a
|
|
73
|
+
custom field whose id differs per Jira instance. What is stable is the *schema*.
|
|
74
|
+
`GET /rest/api/2/field` returns `[{id, name, custom, schema: {type, items, custom}}]`,
|
|
75
|
+
and any field whose schema resolves to `version` - directly, or as `array` with
|
|
76
|
+
`items: "version"` - is read, whatever it is called. The two system fields
|
|
77
|
+
(`fixVersions`, `versions`) are read unconditionally; an affects-version field
|
|
78
|
+
is weighted lower than a fix-version one because it describes where the bug was
|
|
79
|
+
seen, not where the fix lands.
|
|
80
|
+
|
|
81
|
+
`GET /rest/api/2/issue/<key>?expand=names` returns the same display names inline
|
|
82
|
+
and saves the second call when only labels are needed.
|
|
83
|
+
|
|
84
|
+
### The linked release issue
|
|
85
|
+
|
|
86
|
+
`fields.issuelinks[]` carries `{type: {inward, outward}, inwardIssue, outwardIssue}`
|
|
87
|
+
and `fields.parent` carries the parent of a sub-task. On a board where opening a
|
|
88
|
+
development sub-task requires selecting the related release issue, the link is
|
|
89
|
+
the authoritative answer and the version field is the corroboration - so
|
|
90
|
+
`prefs.global.baseBranchEvidence.preferLinkedRelease` (default `false`) raises the
|
|
91
|
+
linked-release weight above the version-field one rather than adding a second rule. A linked issue counts as
|
|
92
|
+
a release issue when a version can be read out of its own fix-version field or
|
|
93
|
+
its summary; the issue *type name* is not consulted, because type names are
|
|
94
|
+
per-board copy.
|
|
95
|
+
|
|
96
|
+
GitHub's analogue is the milestone title, read the same way.
|
|
97
|
+
|
|
98
|
+
## 2b. A filter is not a ranking
|
|
99
|
+
|
|
100
|
+
Step 3's rule-5 list used to be narrowed with `grep -E '(develop|release|main|master)'`,
|
|
101
|
+
which is the prefix table this whole feature exists not to have - and worse than a
|
|
102
|
+
table, because it DISCARDS. A repo whose release branches read `stabilise-2.7`
|
|
103
|
+
matched none of the four words, so none of its branches reached the picker, the
|
|
104
|
+
convention-learner had nothing to learn from, and the version on the issue could
|
|
105
|
+
never match anything. The feature would have been inert on exactly the repos it
|
|
106
|
+
was written for.
|
|
107
|
+
|
|
108
|
+
The rule is the distinction: **a word list that ranks is fine, a word list that
|
|
109
|
+
filters is not.** Ranking only reorders rows the user can still see past; filtering
|
|
110
|
+
removes answers with nothing saying so. So the four family words stay where they
|
|
111
|
+
belong - rule 6's sort and this collector's 15/10/5 `sort-order` weights - and the
|
|
112
|
+
filter gained a version alternative that admits any branch carrying a version
|
|
113
|
+
token, whatever it is called.
|
|
114
|
+
|
|
115
|
+
The excluded prefixes (`feature/`, `bugfix/`, `fix/`, `hotfix/`, `chore/`) are a
|
|
116
|
+
different thing again: those are the task branches **this pipeline creates itself**,
|
|
117
|
+
in Step 4. Excluding your own output is not a convention assumption.
|
|
118
|
+
|
|
119
|
+
## 3. The convention is learned from the remote
|
|
120
|
+
|
|
121
|
+
A table of branch prefixes would be wrong for most repos the day it was written.
|
|
122
|
+
So the template is inferred from the refs that exist: every branch carrying a
|
|
123
|
+
version token is reduced to its shape by replacing the token with `<version>`,
|
|
124
|
+
and the most common shape is this repo's convention. A repo whose release
|
|
125
|
+
branches read `<prefix>/develop_<version>` produces that template; a repo that
|
|
126
|
+
spells them `release-<version>` produces that one; a repo with no versioned
|
|
127
|
+
branches produces `null` and the whole source goes quiet.
|
|
128
|
+
|
|
129
|
+
The inference is reported with its member count (`learned from 4 branch(es)`),
|
|
130
|
+
because an inference from one branch and an inference from a dozen are different
|
|
131
|
+
claims and the picker row should not flatten them.
|
|
132
|
+
|
|
133
|
+
**A predicted branch is never offered.** When the learned template predicts a
|
|
134
|
+
name that is not on the ref list, that is a note, not an option:
|
|
135
|
+
|
|
136
|
+
```
|
|
137
|
+
version 1.51.0 (field "Target Version") has no branch on the remote;
|
|
138
|
+
the convention <prefix>/develop_<version> learned from 4 branch(es) would spell it
|
|
139
|
+
<prefix>/develop_1.51.0
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
That sentence is more useful than a candidate would be - it tells the user the
|
|
143
|
+
release branch has not been cut yet, which is a real answer to "which base?" -
|
|
144
|
+
and an option the user picks has to be checkoutable. Offering a name that does
|
|
145
|
+
not exist just moves the failure to Step 8, where it reads as a git error.
|
|
146
|
+
|
|
147
|
+
## 4. The picker
|
|
148
|
+
|
|
149
|
+
`toPickerOptions()` builds the rows, and the **evidence is the description**. A
|
|
150
|
+
row reading `matches version 1.51.0 from the issue field "Target Version"` and a
|
|
151
|
+
row reading `the repository's default branch` are different answers to the same
|
|
152
|
+
question; a picker that hides which one it is cannot be chosen on its merits.
|
|
153
|
+
|
|
154
|
+
The two-option floor is enforced inside that function rather than left to the
|
|
155
|
+
caller: it never returns fewer than two rows, because `AskUserQuestion` refuses a
|
|
156
|
+
question with fewer than two declared options *and discards every question
|
|
157
|
+
batched with it*, and the host's injected Other row does not count toward the
|
|
158
|
+
schema minimum. The escape row (`Show all branches`) is a genuine second choice,
|
|
159
|
+
not an `OK` button - see picker-contract.md, "Two options or it is not a
|
|
160
|
+
question".
|
|
161
|
+
|
|
162
|
+
**Interactive runs always ask.** A derived candidate is a better-ordered list,
|
|
163
|
+
never a skipped question. `baseBranchSource` stays `asked`, because a human
|
|
164
|
+
answered; the derivation is recorded separately in `state.baseBranchEvidence` so
|
|
165
|
+
it stays readable afterwards.
|
|
166
|
+
|
|
167
|
+
## 5. Autopilot
|
|
168
|
+
|
|
169
|
+
Autopilot cannot be asked anything, so it resolves in this order and records
|
|
170
|
+
which rule fired in `baseBranchSource`:
|
|
171
|
+
|
|
172
|
+
1. `remembered` - `recentBranches` inside the TTL and still on the ref list (memory outranks the default, per the picker contract).
|
|
173
|
+
2. `derived` - the top candidate carries `issue-version` or `linked-release` evidence **and** no other candidate ties its score. Ambiguity is not resolved by coin-flip.
|
|
174
|
+
3. `default` - the sort order.
|
|
175
|
+
|
|
176
|
+
`derived` is an autopilot-only value, exactly like `remembered` and `default`;
|
|
177
|
+
an interactive run recording it has skipped its picker, and the exit gate fails
|
|
178
|
+
it. The gate additionally refuses `derived` unless `state.baseBranchEvidence`
|
|
179
|
+
actually contains issue-derived evidence for the chosen branch - a `derived` that
|
|
180
|
+
derived from nothing is `default` wearing a hat.
|
|
181
|
+
|
|
182
|
+
### Asking on the issue (`prefs.global.baseBranchEvidence.autopilotAsksOnIssue`, default OFF)
|
|
183
|
+
|
|
184
|
+
When the derivation is **ambiguous** and nothing is remembered, autopilot can
|
|
185
|
+
post one comment on the Jira issue or GitHub issue asking which branch to
|
|
186
|
+
develop from, then halt.
|
|
187
|
+
|
|
188
|
+
This is an outward-facing write, so it is fenced:
|
|
189
|
+
|
|
190
|
+
- **Off by default.** Nothing posts unless the user turned it on for this project.
|
|
191
|
+
- **A question, never a state change.** No transition, no resolution, no assignee change, no label change, no close - ever. The standing rule that the pipeline never auto-closes an issue is not relaxed by this feature, and a comment is the only write it is allowed to make.
|
|
192
|
+
- **Human-facing copy follows `outputLanguage`**, like every other comment the pipeline writes.
|
|
193
|
+
- **`Ref:`, never `Closes:`/`Fixes:`/`Resolves:`** in the body, so no platform-side automation reads it as an instruction.
|
|
194
|
+
- **It renders the same evidence the picker would have shown**, candidate by candidate, so the person answering sees what the run saw.
|
|
195
|
+
- **Then the run halts.** Posting a question and continuing on a guess is worse than not asking: the guess lands in a branch while the question sits unanswered. So the comment trips the circuit breaker (trigger 6), which is the sanctioned autopilot pause - state recorded, one actionable line printed, waiting for an explicit `resume`. Phase 0 does not close and no worktree is created.
|
|
196
|
+
|
|
197
|
+
The nearest-safe reading of "post a comment asking which branch" is therefore:
|
|
198
|
+
ask once, change nothing, stop. An autopilot that posts a question and then
|
|
199
|
+
answers it itself has not asked anything.
|
|
200
|
+
|
|
201
|
+
## 6. State
|
|
202
|
+
|
|
203
|
+
```jsonc
|
|
204
|
+
"baseBranchSource": "asked" | "input" | "remembered" | "default" | "derived",
|
|
205
|
+
"baseBranchEvidence": {
|
|
206
|
+
"refProvenance": "remote" | "local", // required whenever this object exists
|
|
207
|
+
"chosen": "<branch>",
|
|
208
|
+
"ambiguous": false,
|
|
209
|
+
"convention": { "template": "<prefix>/develop_<version>", "members": 4 },
|
|
210
|
+
"candidates": [
|
|
211
|
+
{ "branch": "<branch>", "score": 115,
|
|
212
|
+
"evidence": [{ "kind": "issue-version", "detail": "..." }] }
|
|
213
|
+
],
|
|
214
|
+
"notes": ["..."],
|
|
215
|
+
"askedOnIssue": { "target": "PROJ-1234", "url": "...", "at": "..." }
|
|
216
|
+
}
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
`baseBranchEvidence` is required when `baseBranchSource` is `derived`, and when
|
|
220
|
+
`baseFetchStatus` is `cached-stale` or `local-branch`. Everywhere else it is
|
|
221
|
+
optional: a run that took the base from the task reference has no evidence to
|
|
222
|
+
record and should not be made to invent some.
|
|
@@ -1,5 +1,11 @@
|
|
|
1
1
|
## Code Graph (Phase 1 Step 2.6 + Phase 7 Step 3)
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [Phase 1 Step 2.6 - query before dispatching Explore](#phase-1-step-26---query-before-dispatching-explore)
|
|
5
|
+
- [Phase 7 Step 3 - refresh after the branch changed code](#phase-7-step-3---refresh-after-the-branch-changed-code)
|
|
6
|
+
- [The graph is drawable, and one place already asks for it](#the-graph-is-drawable-and-one-place-already-asks-for-it)
|
|
7
|
+
<!-- /toc -->
|
|
8
|
+
|
|
3
9
|
A deterministic, LLM-free map of what a repo declares and what refers to what,
|
|
4
10
|
written to `~/.claude/knowledge/<project>/code-graph.json`. Gated by
|
|
5
11
|
`prefs.global.codeGraph.enabled` (default `false`); with it off, Phase 1 and
|
|
@@ -67,3 +73,37 @@ The rebuild costs no API tokens, so it runs every task rather than on a stalenes
|
|
|
67
73
|
heuristic. A non-zero validator exit keeps the previous graph and logs
|
|
68
74
|
`knowledge.graph_invalid`; it never fails the run - a stale graph is a degraded
|
|
69
75
|
Phase 1, not a broken deliverable.
|
|
76
|
+
|
|
77
|
+
### The graph is drawable, and one place already asks for it
|
|
78
|
+
|
|
79
|
+
The PR body's Impact Analysis, part 3, asks which symbols and files a change
|
|
80
|
+
reaches. That is `graph-affected.mjs`'s question, and until now the answer was
|
|
81
|
+
re-typed as prose by a model while the measurement sat on disk unread.
|
|
82
|
+
|
|
83
|
+
```bash
|
|
84
|
+
node $HOME/.claude/scripts/graph-mermaid.mjs "<symbol[,symbol]>" [--depth N] [--max-nodes N]
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
It emits a fenced `flowchart` and nothing else - no renderer, no plugin, no
|
|
88
|
+
dependency, because mermaid is text and GitHub renders it natively in pull
|
|
89
|
+
requests, issues and markdown files. Traversal is not reimplemented: `findByName`
|
|
90
|
+
and `affected` are imported from `graph-affected.mjs`, so the diagram and the
|
|
91
|
+
text report cannot disagree about what is affected.
|
|
92
|
+
|
|
93
|
+
Three properties that are enforced rather than promised
|
|
94
|
+
(`smoke-graph-mermaid.sh`):
|
|
95
|
+
|
|
96
|
+
- Every drawn node and edge resolves back into `code-graph.json`, with the edge
|
|
97
|
+
kind it claims. A diagram is read as fact and checked less than prose, so an
|
|
98
|
+
invented edge is the expensive failure.
|
|
99
|
+
- Over `--max-nodes` the leftover count is printed inside the diagram, not
|
|
100
|
+
dropped. A small picture of a large blast radius reads as reassurance.
|
|
101
|
+
- The graph's `baseCommit` is printed beside it. A graph built before the change
|
|
102
|
+
draws an older tree, and nothing else in the PR would reveal that.
|
|
103
|
+
|
|
104
|
+
Exit 1 with a reason on stderr means no graph or no such symbol. The caller
|
|
105
|
+
records the gap and writes the prose alone; it never hand-draws a replacement.
|
|
106
|
+
|
|
107
|
+
Jira is not a target: its renderer turns the fence into a literal
|
|
108
|
+
`{code:mermaid}` block. Confluence renders it through the `ac:name="mermaid"`
|
|
109
|
+
macro when the space carries the plugin (`channels/confluence.md`).
|