@mmerterden/multi-agent-pipeline 16.27.0 → 16.28.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. package/CHANGELOG.md +71 -1
  2. package/package.json +4 -4
  3. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  4. package/pipeline/commands/multi-agent/issue/SKILL.md +1 -0
  5. package/pipeline/commands/multi-agent/jira/SKILL.md +3 -0
  6. package/pipeline/commands/multi-agent/log/SKILL.md +7 -1
  7. package/pipeline/commands/multi-agent/review-issue/SKILL.md +1 -0
  8. package/pipeline/lib/issue-fetcher.sh +134 -5
  9. package/pipeline/lib/multi-repo-pipeline.sh +8 -0
  10. package/pipeline/multi-agent-refs/analysis/evidence.md +1 -1
  11. package/pipeline/multi-agent-refs/analysis/intake.md +11 -2
  12. package/pipeline/multi-agent-refs/analysis/locked.md +1 -0
  13. package/pipeline/multi-agent-refs/analysis/redesign.md +112 -0
  14. package/pipeline/multi-agent-refs/analysis/render.md +5 -0
  15. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -0
  16. package/pipeline/multi-agent-refs/analysis/review.md +15 -0
  17. package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
  18. package/pipeline/multi-agent-refs/analysis-template-corporate.md +3 -3
  19. package/pipeline/multi-agent-refs/analysis-template.md +36 -0
  20. package/pipeline/multi-agent-refs/cross-cli-contract.md +3 -0
  21. package/pipeline/multi-agent-refs/features/jira-context.md +101 -0
  22. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +1 -1
  23. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +1 -1
  24. package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -1
  25. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  26. package/pipeline/schemas/agent-state.schema.json +19 -0
  27. package/pipeline/schemas/prefs.schema.json +26 -0
  28. package/pipeline/scripts/anonymize-findings.mjs +24 -0
  29. package/pipeline/scripts/build-references.mjs +10 -6
  30. package/pipeline/scripts/council-view.mjs +144 -0
  31. package/pipeline/scripts/phase-tracker.sh +18 -1
  32. package/pipeline/scripts/skill-siblings.mjs +41 -8
  33. package/pipeline/scripts/validate-analysis-doc.mjs +371 -7
  34. package/pipeline/scripts/validate-analysis.mjs +7 -5
  35. package/pipeline/skills/shared/core/multi-agent-issue/SKILL.md +1 -0
  36. package/pipeline/skills/shared/core/multi-agent-review-issue/SKILL.md +1 -0
@@ -86,6 +86,21 @@ Fable triages the pooled findings exactly as in `/multi-agent:review`: drop dupl
86
86
 
87
87
  The verdict names the profile it judged against, the counts per severity, and what was NOT checked (the references gate without a state file, anything the fetch could not reach). Follow the pipeline rule on claiming: state which findings are mechanically proven and which are judgement.
88
88
 
89
+ **Every rubric class is printed, including the empty ones.** A verdict that lists only the classes that fired says nothing about the rest: the reader cannot tell "checked, clean" from "never looked". Print A through F in order with a count each, and write `none` where the count is zero. Same rule for the deterministic gates: name each one and its verdict, including the ones that skipped and why.
90
+
91
+ **Colophon.** Close the verdict with the reviewed document's `evidence_digest` and `base_commit`, taken from its front-matter, plus the reviewer models used. Two reviews of the same feature on the same day are otherwise indistinguishable, and the first question anyone asks of an older verdict is which version of the document it judged. `validate-analysis-doc.mjs --report` prints the per-check verdicts to paste under it.
92
+
93
+ ```
94
+ Verdict: 2 Blocker, 3 Important, 1 Suggestion (profile: global, mode: full)
95
+
96
+ A evidence 2 D altitude none
97
+ B backbone none E admitted gaps 1
98
+ C ... F contradiction 3
99
+
100
+ Not checked: references coverage (no state file passed)
101
+ Provenance: evidence_digest sha256:1f3a9c2 base_commit 4c1b2de
102
+ ```
103
+
89
104
  ## Phase 4 - Output
90
105
 
91
106
  Default is the chat report. Nothing is written anywhere without an explicit choice.
@@ -88,7 +88,7 @@ For each `platform` in `state.analysisSpec.platforms[]`:
88
88
  - `web` → `evidence.standards[]` entries matching `react`, `vue`, `next`, `sveltekit` → `~/.claude/rules/code-style.md`
89
89
  2. **Apply per-platform omission rules.** Backend-only file drops Sections 5, 6, 7, 8, 16. Web with no UI inventory still keeps 5 (UI exists in code). Sections 1, 2, 4, 9, 13, 14, 20, 21 always present per Locked decision 2 + 13.
90
90
  3. **Resolve mode.** If user passed `--lite` → Lite. If user passed `--full` → Full. Otherwise use `state.analysisSpec.liteModeAuto`. Lite mode renders only Sections 1, 2, 4, 9, 13, 14, 21 plus optional 23.
91
- 4. **Produce YAML front-matter header** (see `$HOME/.claude/multi-agent-refs/analysis-template.md`). Include `profile: <state.analysisSpec.profile | global>` and `platform: <platform | none>` so the validator applies the right contract per profile (Locked 32) and recognises the stack-optional render (Locked 35), `mode: full | lite`, plus `ui_tests: <state.analysisSpec.options.uiTests | false>` and `a11y_depth: <state.analysisSpec.options.a11yDepth | basic>` so the pre-dispatch validator can enforce the opt-in coverage (15.6 present when ui_tests, 16.2 walkthrough present when a11y_depth is full).
91
+ 4. **Produce YAML front-matter header** (see `$HOME/.claude/multi-agent-refs/analysis-template.md`). Include `profile: <state.analysisSpec.profile | global>` and `platform: <platform | none>` so the validator applies the right contract per profile (Locked 32) and recognises the stack-optional render (Locked 35), `mode: full | lite`, plus `ui_tests: <state.analysisSpec.options.uiTests | false>`, `a11y_depth: <state.analysisSpec.options.a11yDepth | basic>` and `redesign: <state.analysisSpec.options.redesign | false>` so the pre-dispatch validator can enforce the opt-in coverage (15.6 present when ui_tests, 16.2 walkthrough present when a11y_depth is full, 4.5 / 4.6 / 9.5 present when redesign). `status` is written only by `/multi-agent:analysis-resolve`; a rendered document is a draft.
92
92
  5. **Read conventions for this platform's repo.** For each cell Pass B fills in Section 13 and in any per-platform projection (Sections 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17), read `state.analysisSpec.evidence.conventions[<repo>].<field>` and emit the value with a footnote (Locked 24). If `conventionOverrides` has an entry for that field, use the override and footnote with `^[user-override: <reason>]` instead of evidence path.
93
93
  6. **Concatenate non-null sections in canonical order.** Numbering stays sequential `1..N` over the rendered set (omitted sections do not create gaps).
94
94
  7. **Schema validation** on the per-platform spec object:
@@ -433,12 +433,12 @@ Never omitted.
433
433
  ## 20. Riskler ve Açık Sorular <!-- TR -->
434
434
  ## 20. Risks and Open Questions <!-- EN -->
435
435
 
436
- | # | Konu | Neden açık | Kime sorulacak | Etkilediği bölüm |
436
+ | Id | Konu | Neden açık | Kime sorulacak | Etkilediği bölüm |
437
437
  |---|---|---|---|---|
438
- | 1 | <question> | <what evidence is missing> | <role or team> | 2.3 |
438
+ | AS-01 | <question> | <what evidence is missing> | <role or team> | 2.3 |
439
439
  ```
440
440
 
441
- The source corporate documents keep open questions out of the page and raise them in conversation instead. This profile keeps them in the document deliberately: an `EKLENECEK` with no matching row here is an unanswered question nobody owns. Every `EKLENECEK` and every unverified assumption in 2.4 emits a row (Locked 33).
441
+ The source corporate documents keep open questions out of the page and raise them in conversation instead. This profile keeps them in the document deliberately: an `EKLENECEK` with no matching `AS-NN` row here is an unanswered question nobody owns. Every `EKLENECEK` and every unverified assumption in 2.4 emits a row (Locked 33).
442
442
 
443
443
  ## 21. Referanslar / References
444
444
 
@@ -46,6 +46,8 @@ platform: ios | android | web | backend | mobile | none
46
46
  profile: global | corporate
47
47
  language: tr | en
48
48
  mode: full | lite
49
+ status: draft | final
50
+ redesign: true | false
49
51
  ui_tests: true | false
50
52
  a11y_depth: basic | full
51
53
  generated: <ISO 8601 UTC timestamp>
@@ -64,6 +66,8 @@ template_version: v3
64
66
 
65
67
  `ui_tests` and `a11y_depth` record the Phase 0 Step 5a opt-ins (defaults `false` / `basic`) so the coverage choice is auditable and the pre-dispatch validator can enforce it: `ui_tests: true` requires Section 15.6, `a11y_depth: full` requires the Section 16.2 walkthrough.
66
68
 
69
+ `status` is the publication claim; absent means `draft`. Under `final` an unresolved `AS-NN` or open Section 20 row is an ERROR, under `draft` it is counted. `/multi-agent:analysis-resolve` flips it, after a confirmation.
70
+
67
71
  `profile` names the template the document was rendered against (Locked 32) and `platform: none` marks the stack-optional render (Locked 35). Both are read by `validate-analysis-doc.mjs`, which applies a different contract per profile: without the key a corporate document would be judged against the global rules and its backbone would read as a pile of Locked 2 violations.
68
72
 
69
73
  Phase 1 compares `evidence_digest` against an existing document to decide whether to reuse it (Locked 27). Phase 3 reads `platform` to verify file match, `mode` to know which section set to expect, and both `evidence_digest` and `base_commit` to judge freshness: the digest says the evidence changed, `base_commit` says the repo moved. Phase 2 parses the block but gates only on `template_version`.
@@ -284,6 +288,27 @@ EARS states the rule; Gherkin states how you check it. Each rule still maps to a
284
288
 
285
289
  Each scenario in 4.1-4.3 cites the matching `BR-<slug>-NN` id (and, when a server code applies, the Section 9.4 code). Plain-prose user stories without Given/When/Then are rejected by the renderer.
286
290
 
291
+ ### 4.5 Mevcut Davranış / Current Behaviour (OPTIONAL)
292
+
293
+ `redesign: true` only. Locked 37; contract and checks in `analysis/redesign.md`. Evidence and certainty are derived from `repoEvidence`, never graded by the writer.
294
+
295
+ ```markdown
296
+ | Id | Davranış / Behaviour | Kanıt / Evidence | Kesinlik / Certainty |
297
+ |---|---|---|---|
298
+ | CB-<slug>-01 | <what v1 does, present tense, no v2 language> | <repo>/<path>:<line> | confirmed |
299
+ | CB-<slug>-02 | <what v1 does> | <module> (cross-cutting) | uncertain |
300
+ ```
301
+
302
+ ### 4.6 Endpoint Eşlemesi / Endpoint Mapping (OPTIONAL)
303
+
304
+ `redesign: true` only. An unchanged path still gets its row: "unchanged" is an answer, its absence is not.
305
+
306
+ ```markdown
307
+ | v1 | v2 | Değişiklik / Change |
308
+ |---|---|---|
309
+ | GET /v1/<path> | GET /v2/<path> | <shape change, or unchanged> |
310
+ ```
311
+
287
312
  ## 5. Tasarım Referansı / Design Reference
288
313
 
289
314
  Per-platform projection. Omitted if no Figma URL AND no Code Connect mapping. Locked 17 + 18.
@@ -484,6 +509,17 @@ List every HTTP status code one by one.
484
509
  | ERR-211 | Empty result set | <screen> |
485
510
  ```
486
511
 
512
+ ### 9.5 Fark Listesi / Difference List (OPTIONAL)
513
+
514
+ `redesign: true` only. One row per `CB-` id, both directions checked. Closed status vocabulary - `Moved`, `Partial`, `Missing`, `New`, `Out of scope` - in a bilingual cell like Section 20's `Açık / Open`. Every `Missing` and `Partial` owes a Section 20 row by `AS-NN`.
515
+
516
+ ```markdown
517
+ | Id | Durum / Status | Nereye / Where | AS-NN |
518
+ |---|---|---|---|
519
+ | CB-<slug>-01 | Taşındı / Moved | 4.2 | - |
520
+ | CB-<slug>-02 | Eksik / Missing | - | AS-03 |
521
+ ```
522
+
487
523
  ## 10. Lokalizasyon Anahtarları / Localization Keys
488
524
 
489
525
  Per-platform projection. Locked 20 - shape depends on the project `figma-config` `localization.ownership`. The locale set comes from `localization.locales` (default `tr, en, ar, de, es, fr, it, ru`); never hardcode a locale list.
@@ -230,6 +230,9 @@ argument-hint: "<input hint>"
230
230
  |---|---|
231
231
  | Claude → Copilot | Add `name`, `user-invocable: true`, `argument-hint`; remove `allowed-tools`; `description` stays English (localization happens only in the installed Claude tree via `description-tr` + `localize-commands.mjs`, never in repo or Copilot files) |
232
232
  | Copilot → Claude | Remove `name`, `user-invocable`, `argument-hint`; add `allowed-tools` (inferred from command's actual tool calls) |
233
+ | Both directions | `not-for: <sibling>` carries across unchanged |
234
+
235
+ **`not-for`** is optional and lists sibling commands this one must not be chosen for, as bare names (`not-for: jira, review-jira`). It is written for the model reading the routing surface, so it belongs beside the description rather than in prose further down. `lint-skills.mjs` also consumes it: a trigger-vocabulary collision either side has named drops off the warning list and is reported as settled instead, and a name that resolves to no sibling on the same surface is an error - a typo would otherwise leave the pair unanswered while the author believes it is handled. Both trees spell the value the same way; the `multi-agent-` prefix on the Copilot side is stripped before matching.
233
236
 
234
237
  **Invariant**: the English `description` must have the same meaning on both sides, and a `description-tr` line (when present) must be a faithful translation of it. Semantic drift is not allowed. `description-en` sidecars are an installed-tree artifact and must never appear in repo or Copilot files.
235
238
 
@@ -0,0 +1,101 @@
1
+ # Related-issue context at intake
2
+
3
+ A development sub-task is often filed with no description of its own. The
4
+ requirement sits on the parent, and the rest of the picture - the analysis, the
5
+ test scope - sits on the sibling sub-tasks beside it. The fetcher already read
6
+ the parent; it did not read the siblings, so a board that keeps its analysis in a
7
+ separate sub-task handed the pipeline an empty task.
8
+
9
+ `issue-fetcher.sh` now reads those siblings and exposes them as
10
+ `descriptor.relatedIssues[]`.
11
+
12
+ ## What it costs
13
+
14
+ One search request, and only when the issue has a parent:
15
+
16
+ ```
17
+ GET /rest/api/2/search?jql=parent="<PARENT>"&fields=summary,issuetype,status,description&maxResults=20
18
+ ```
19
+
20
+ An issue with no parent has no siblings, so the lookup is skipped entirely and
21
+ that run pays nothing. The count does not grow with the number of siblings: the
22
+ issue's own `subtasks` field lists ITS children rather than its siblings, and the
23
+ parent's lists siblings without their descriptions, so either of those shapes
24
+ would cost one GET per sibling. `smoke-jira-context.sh` counts the requests, so
25
+ this stays a measurement rather than a claim.
26
+
27
+ The request lands before the maturity verdict reaches the user, which is one
28
+ extra round trip on the intake path. `jiraContext.enabled: false` is the way out
29
+ when latency matters more than the context.
30
+
31
+ ## Shape
32
+
33
+ ```json
34
+ "relatedIssues": [
35
+ {
36
+ "key": "PROJ-1002",
37
+ "relation": "sibling",
38
+ "type": "Sub-task",
39
+ "status": "Done",
40
+ "summary": "...",
41
+ "description": "...",
42
+ "truncated": false
43
+ }
44
+ ]
45
+ ```
46
+
47
+ `type` is carried through verbatim from Jira and is never matched against in
48
+ code. Boards name their sub-task types differently and the query does not need
49
+ to know: `parent = <key>` returns all of them.
50
+
51
+ Siblings that carry a description are ordered first, so `maxItems` drops the
52
+ empty ones rather than the useful ones. An empty sibling is still listed - its
53
+ key, type and status are information - with an empty `description`.
54
+
55
+ ## Settings
56
+
57
+ `prefs.global.jiraContext`:
58
+
59
+ | Key | Default | Effect |
60
+ |---|---|---|
61
+ | `enabled` | `true` | `false` skips the search entirely |
62
+ | `maxItems` | `6` | Hard cap; `0` behaves like `enabled: false` |
63
+ | `maxCharsPerItem` | `1200` | Per-sibling truncation, marked `truncated: true` |
64
+
65
+ Worst case payload is `maxItems x maxCharsPerItem`, next to the 4000-character
66
+ cap on the issue's own description.
67
+
68
+ ## Maturity
69
+
70
+ One new code, `description_empty_sibling_available`, and it fires narrowly: the
71
+ issue's own description is empty, the parent's is empty or unreachable, AND a
72
+ sibling carries content. Today that combination produces the hard
73
+ `description_empty` blocker; it becomes a warning instead, the same downgrade
74
+ `description_empty_parent_available` already makes one level up.
75
+
76
+ The narrowness is deliberate. `score` is `100 - 40*blockers - 10*warnings`, and
77
+ the picker continues silently only at `score >= 90`, so a warning that fired
78
+ whenever an issue merely HAD siblings would cost ten points on every parented
79
+ issue and turn a silent intake into a question across the board.
80
+
81
+ ## Where it goes
82
+
83
+ The descriptor is carried verbatim into the picker state as `issueRef`, and the
84
+ bridge copies `relatedIssues` onto `agent-state.json` when it is non-empty.
85
+
86
+ Persisting it is the point: `/multi-agent:resume` rebuilds context from durable
87
+ artefacts and never from the conversation, so an intake-only enrichment would be
88
+ gone by the first resume. The task's own `description` and `maturity` are still
89
+ not persisted, which is a known asymmetry rather than an oversight - this
90
+ change did not widen the bridge beyond the one field it needed.
91
+
92
+ Sibling text is added as its own labelled block. It never replaces the working
93
+ description the way an accepted `parentDescription` does: substituting it would
94
+ erase the provenance that Phase 2's parent-story scope-drift check reads.
95
+
96
+ ## Not included
97
+
98
+ Following the parent's `issuelinks` in a second search. A linked bug can be the
99
+ whole requirement while a linked test-execution record is noise, and the fetcher
100
+ cannot tell them apart, so the scope stops at sub-tasks. There is no setting for
101
+ it: a switch that does nothing is worse than an absent one.
@@ -109,7 +109,7 @@ back. "Copy X and rename it" is the reuse answer, not a hint - name X's files.
109
109
 
110
110
  #### Step 1.5 - External Context Injection (`state.contextLinks[]`)
111
111
 
112
- Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher (crashlytics, fortify, graylog, swagger, confluence, figma, generic-doc) and prepends results under a **Referenced External Sources** section in the analysis prompt - so the agent doesn't re-discover what the ticket already pointed at. `state.graylogContext` is injected there too, as diagnostic context (advisory only). Failures never fatal (a non-zero fetcher exit is marked skipped and the analysis still runs, exactly as for crashlytics); pending refs are advisories. Full dispatch table, exit-code handling, prompt injection shape, log line shape: `$HOME/.claude/multi-agent-refs/features/external-context-injection.md`.
112
+ Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher (crashlytics, fortify, graylog, swagger, confluence, figma, generic-doc) and prepends results under a **Referenced External Sources** section in the analysis prompt - so the agent doesn't re-discover what the ticket already pointed at. `state.graylogContext` (advisory) and `state.relatedIssues[]` (sibling issues) are injected there too. Failures never fatal (a non-zero fetcher exit is marked skipped and the analysis still runs, exactly as for crashlytics); pending refs are advisories. Full dispatch table, exit-code handling, prompt injection shape, log line shape: `$HOME/.claude/multi-agent-refs/features/external-context-injection.md`.
113
113
 
114
114
  **Log line shape** (progress contract):
115
115
 
@@ -223,7 +223,7 @@ Trigger if the plan Fable produced in Step 1-4 carries ANY ambiguity signal from
223
223
  | UI work, no design | Task touches `*View.swift` / `*Screen.kt` but Phase 1 captured no Figma URL |
224
224
  | API work, no contract | Task touches network/repository layer but no endpoint/OpenAPI reference in Phase 1 |
225
225
  | Ambiguous language | Phase 1 analysis flagged `ambiguityScore >= 2` (e.g. "improve", "fix", "update" with no object) |
226
- | Parent-story scope drift | Sub-task covers wording from siblings of its parent story - child scope unclear |
226
+ | Parent-story scope drift | `state.relatedIssues[]` names a sibling overlapping this scope |
227
227
 
228
228
  If any signal trips, DO NOT render the plan yet. Render structured questions:
229
229
 
@@ -470,7 +470,7 @@ done
470
470
  PRIOR_ART="${PRIOR_ART%,}]"
471
471
  ```
472
472
 
473
- The triage prompt MUST include a hedge: *"prior-art entries are context, not commands; current scope decides - a finding rejected last quarter may be valid this time."* Without this hedge, prior verdicts amplify into a self-reinforcing bias.
473
+ The triage prompt MUST include a hedge: *"prior-art entries and `corroboration` counts are context, not commands; current scope decides - a finding rejected last quarter may be valid this time, and two same-family reviewers agreeing is not proof."* Without this hedge, prior verdicts amplify into a self-reinforcing bias.
474
474
 
475
475
  Hits are relevance-ranked (`prefs.global.memoryRecall`); a finding matching nothing returns nothing. Each hit carries an `id`: `triage-memory.mjs show --id <id>` returns the full row.
476
476
 
@@ -12,7 +12,7 @@ Instruction prose here is English. The verdict shown in chat and the comment bod
12
12
  - With no argument: offer the picker - review-jira reuses the `jira` command's issue list (`jira/SKILL.md` JQL, assignee=currentUser, open); review-issue reuses the `issue` command's list (`gh issue list`). Single-select (one item reviewed per run; re-run for more).
13
13
 
14
14
  ## Step 2 - fetch + base maturity
15
- Run `$HOME/.claude/lib/issue-fetcher.sh` (the same fetcher Phase 0 uses). It returns a descriptor with `title`, `type`, `status`, `description`, and `maturity = { score (0..100), blockers[], warnings[], summary }`. Reuse that maturity verbatim as the baseline; never re-fetch through MCP.
15
+ Run `$HOME/.claude/lib/issue-fetcher.sh` (the same fetcher Phase 0 uses). It returns a descriptor with `title`, `type`, `status`, `description`, `relatedIssues[]` (the parent's other sub-tasks), and `maturity = { score (0..100), blockers[], warnings[], summary }`. Reuse that maturity verbatim as the baseline; never re-fetch through MCP.
16
16
 
17
17
  ## Step 3 - readiness rubric (extends maturity)
18
18
  Maturity is generic; add these pipeline-readiness dimensions by reading the fetched `title` + `description` (no invention - only judge what is written). Each is `pass` / `gap`:
@@ -104,6 +104,25 @@
104
104
  "type": ["string", "null"],
105
105
  "description": "Set when a phase halts on a hard error (validator failed twice, no subagent returned, dispatch error past fallback, lock irrecoverable). Format '<phase>:<cause>'. Surfaced to the user and cleared on successful resume. See operations.md 'Halt visibility'."
106
106
  },
107
+ "relatedIssues": {
108
+ "type": "array",
109
+ "maxItems": 20,
110
+ "description": "Jira issues fetched alongside the task at intake: the other sub-tasks under this issue's parent, where a board keeps the analysis and the test scope. NOT the same field as siblings[] below, which is read-only sibling REPOS. Written by the picker bridge from descriptor.relatedIssues; read by Phase 1 as ground truth and by Phase 2's parent-story scope-drift check. Durable on purpose: /multi-agent:resume rebuilds context from artefacts, never from the conversation, so intake-only enrichment would vanish on the first resume. The task's own description and maturity are still NOT persisted here, which is a known asymmetry, not an oversight.",
111
+ "items": {
112
+ "type": "object",
113
+ "additionalProperties": false,
114
+ "required": ["key", "relation"],
115
+ "properties": {
116
+ "key": { "type": "string" },
117
+ "relation": { "type": "string", "enum": ["sibling"] },
118
+ "type": { "type": "string" },
119
+ "status": { "type": "string" },
120
+ "summary": { "type": "string" },
121
+ "description": { "type": "string" },
122
+ "truncated": { "type": "boolean", "default": false }
123
+ }
124
+ }
125
+ },
107
126
  "siblings": {
108
127
  "type": "array",
109
128
  "maxItems": 10,
@@ -1358,6 +1358,32 @@
1358
1358
  "description": "v12.8+ - Phase 4 Step 1.77 reviewer-scope gate. Decides the reviewer count from the deterministic diff-risk report via `pipeline/scripts/review-scope.mjs`: a diff under 20 lines of churn with max_score < 3.0 and no security_path / migration / public_api / no_test_change / test_lines_removed signal runs ONE reviewer instead of the full CLI-aware set (2 on Claude Code, 3 on Copilot CLI). No LLM. Fails safe in one direction only - any error, empty report or validator rejection resolves to the full set, because skipping a reviewer trades coverage for cost. Set false to force the full set on every diff.",
1359
1359
  "$comment": "Shipped inert for a release: the script existed, was unit- and smoke-tested, and no phase doc referenced it, so every diff paid for the full reviewer set. Wired in Step 1.77; smoke-gate-wiring.sh keeps it reachable."
1360
1360
  },
1361
+ "jiraContext": {
1362
+ "type": "object",
1363
+ "additionalProperties": false,
1364
+ "description": "v16.28+ - when a Jira issue has a parent, the intake fetcher also reads the parent's other sub-tasks in ONE extra search request and exposes them as descriptor.relatedIssues[]. A development sub-task filed with an empty description whose analysis sibling carries the real requirement is the case this exists for. Costs nothing on an issue with no parent, because an issue with no parent has no siblings. The extra round trip lands before the maturity verdict is shown, so turn it off when intake latency matters more than the context.",
1365
+ "properties": {
1366
+ "enabled": {
1367
+ "type": "boolean",
1368
+ "default": true,
1369
+ "description": "Master switch. When false the fetcher never issues the search and relatedIssues[] is always empty."
1370
+ },
1371
+ "maxItems": {
1372
+ "type": "integer",
1373
+ "default": 6,
1374
+ "minimum": 0,
1375
+ "maximum": 20,
1376
+ "description": "Hard cap on relatedIssues[]. Siblings that carry a description are kept first, so the cap drops the empty ones. 0 disables the lookup exactly like enabled: false."
1377
+ },
1378
+ "maxCharsPerItem": {
1379
+ "type": "integer",
1380
+ "default": 1200,
1381
+ "minimum": 0,
1382
+ "maximum": 4000,
1383
+ "description": "Per-sibling description truncation, marked with truncated: true. Worst case payload is maxItems x this, next to the 4000-char cap on the issue's own description."
1384
+ }
1385
+ }
1386
+ },
1361
1387
  "priorArtEnrichment": {
1362
1388
  "type": "object",
1363
1389
  "additionalProperties": false,
@@ -122,6 +122,30 @@ export function anonymize(input, { seed } = {}) {
122
122
  }
123
123
  });
124
124
 
125
+ // How many DIFFERENT reviewers reported this finding, counted from the
126
+ // fingerprints the validator already stamped. It costs nothing: every
127
+ // reviewer's findings are in this one payload and nowhere else are they
128
+ // together. `count` names no reviewer, so it survives the identity strip
129
+ // without reopening what the strip closed.
130
+ //
131
+ // It is context for triage, never a rule. ADR-0001 rejected majority voting
132
+ // precisely because the same hallucination shows up in two of three
133
+ // same-family reviewers, so a high count is not evidence of truth. The
134
+ // reading that pays is the other end: count 1 is the finding only one
135
+ // reviewer saw, which is where a judge should look hardest.
136
+ const seen = new Map();
137
+ for (const f of tagged) {
138
+ const k = typeof f?.fingerprint === "string" && f.fingerprint ? f.fingerprint : null;
139
+ if (!k) continue;
140
+ if (!seen.has(k)) seen.set(k, new Set());
141
+ seen.get(k).add(f.foundBy);
142
+ }
143
+ const of = reviewers.length;
144
+ for (const f of tagged) {
145
+ const k = typeof f?.fingerprint === "string" && f.fingerprint ? f.fingerprint : null;
146
+ f.corroboration = { count: k ? seen.get(k).size : 1, of };
147
+ }
148
+
125
149
  tagged.sort((a, b) => {
126
150
  const ka = stableKey(a);
127
151
  const kb = stableKey(b);
@@ -338,7 +338,8 @@ function main() {
338
338
  process.stderr.write(
339
339
  "usage: build-references.mjs <state.json|-> [--lang tr|en] [--check <doc.md>]\n",
340
340
  );
341
- process.exit(2);
341
+ process.exitCode = 2;
342
+ return;
342
343
  }
343
344
 
344
345
  let spec;
@@ -346,13 +347,15 @@ function main() {
346
347
  spec = readState(statePath);
347
348
  } catch (err) {
348
349
  process.stderr.write(`ERROR: cannot read state: ${err.message}\n`);
349
- process.exit(2);
350
+ process.exitCode = 2;
351
+ return;
350
352
  }
351
353
 
352
354
  const lang = argOf(argv, "--lang") ?? spec.language ?? "tr";
353
355
  if (!HEADERS[lang]) {
354
356
  process.stderr.write(`ERROR: unknown language "${lang}"; expected tr or en\n`);
355
- process.exit(2);
357
+ process.exitCode = 2;
358
+ return;
356
359
  }
357
360
 
358
361
  const docPath = argOf(argv, "--check");
@@ -361,10 +364,11 @@ function main() {
361
364
  if (problems.length) {
362
365
  for (const p of problems) process.stderr.write(`${p}\n`);
363
366
  process.stderr.write(`\n${problems.length} references coverage failure(s)\n`);
364
- process.exit(1);
367
+ process.exitCode = 1;
368
+ return;
365
369
  }
366
370
  process.stdout.write("references coverage ok\n");
367
- process.exit(0);
371
+ return;
368
372
  }
369
373
 
370
374
  const rows = rowsFrom(spec, lang);
@@ -376,7 +380,7 @@ function main() {
376
380
  ? "Bu koşuda getirilen kaynak yok.\n"
377
381
  : "No sources were fetched for this run.\n",
378
382
  );
379
- process.exit(0);
383
+ return;
380
384
  }
381
385
  process.stdout.write(`${renderTable(rows, lang)}\n`);
382
386
  }
@@ -0,0 +1,144 @@
1
+ #!/usr/bin/env node
2
+ // council-view.mjs - what each reviewer found, and what triage did with it.
3
+ //
4
+ // Phase 4 dispatches three reviewers, anonymizes their findings so the judge
5
+ // cannot mark its own homework, and reports the verdict. Everything needed to
6
+ // answer "which model found this, and was it kept" is already on the state -
7
+ // reviewers[].findings, anonymizationMap.labelToModel, triage.accepted - and
8
+ // nothing ever renders it. run-metrics.mjs reduces it to a ratio; the rows
9
+ // behind the ratio are not visible anywhere.
10
+ //
11
+ // This is the panel view llm-council shows as its stage-1 tabs, built from
12
+ // state that already exists. No model call, no new field, no new phase.
13
+ //
14
+ // De-anonymization is deliberate and safe here: the map is read AFTER triage
15
+ // has ruled, so nothing this prints can bias the judgement it describes. The
16
+ // map never goes into a prompt; this is a report.
17
+ //
18
+ // Usage:
19
+ // council-view.mjs <agent-state.json> [--json] [--iteration N]
20
+ //
21
+ // Exit codes: 0 rendered, 2 nothing to render (no review iterations), 64 usage.
22
+
23
+ import { readFileSync } from "node:fs";
24
+
25
+ const argv = process.argv.slice(2);
26
+ const JSON_OUT = argv.includes("--json");
27
+ const iterFlag = argv.indexOf("--iteration");
28
+ const ONLY = iterFlag !== -1 ? Number(argv[iterFlag + 1]) : null;
29
+ const file = argv.find((a) => !a.startsWith("--") && a !== String(ONLY));
30
+ // `--iteration` with nothing after it made ONLY NaN, and `it.iteration === NaN`
31
+ // is never true, so a state full of iterations exited 2 - which /multi-agent:log
32
+ // is told to read as "the run never reached Phase 4" and quietly omit. A usage
33
+ // error has to look like one.
34
+ const BAD_ITERATION = iterFlag !== -1 && !Number.isInteger(ONLY);
35
+
36
+ function die(msg, code) {
37
+ process.stderr.write(msg + "\n");
38
+ process.exitCode = code;
39
+ }
40
+
41
+ function verdictOf(triage, fingerprint, issue, fileName) {
42
+ for (const bucket of ["accepted", "deferred", "rejected"]) {
43
+ for (const entry of Array.isArray(triage?.[bucket]) ? triage[bucket] : []) {
44
+ // Only `accepted[]` is a flat finding. triage-output.schema.json defines a
45
+ // deferred or rejected item as `{ finding, reason }` - the reason is the
46
+ // point of those buckets - so reading the id off the wrapper made every
47
+ // ruled-out finding render as `unruled`, which is exactly the column this
48
+ // view exists to show.
49
+ const f = entry?.finding ?? entry;
50
+ if (fingerprint && f?.fingerprint === fingerprint) return bucket;
51
+ if (!fingerprint && f?.file === fileName && f?.issue === issue) return bucket;
52
+ }
53
+ }
54
+ return "unruled";
55
+ }
56
+
57
+ function main() {
58
+ if (!file) return die("usage: council-view.mjs <agent-state.json> [--json] [--iteration N]", 64);
59
+ if (BAD_ITERATION) return die("council-view: --iteration needs an integer", 64);
60
+ let state;
61
+ try {
62
+ state = JSON.parse(readFileSync(file, "utf8"));
63
+ } catch (e) {
64
+ return die(`council-view: cannot read ${file}: ${e.message}`, 64);
65
+ }
66
+
67
+ const iterations = Array.isArray(state?.reviewIterations) ? state.reviewIterations : [];
68
+ const wanted = ONLY == null ? iterations : iterations.filter((it) => it?.iteration === ONLY);
69
+ if (wanted.length === 0) return die("council-view: no review iterations on this state", 2);
70
+
71
+ const out = [];
72
+ for (const it of wanted) {
73
+ // Absent map is the normal state for a run written before anonymization,
74
+ // and for one where every reviewer timed out. Say "unknown" rather than
75
+ // guessing a model, the same way run-metrics.mjs degrades.
76
+ const map = it?.anonymizationMap?.labelToModel || {};
77
+ const triage = it?.triage || {};
78
+ const rows = [];
79
+ for (const r of Array.isArray(it?.reviewers) ? it.reviewers : []) {
80
+ const model = typeof r?.model === "string" && r.model ? r.model : "unknown";
81
+ for (const f of Array.isArray(r?.findings) ? r.findings : []) {
82
+ rows.push({
83
+ model,
84
+ severity: f?.severity || "",
85
+ file: f?.file || "",
86
+ line: f?.line ?? "",
87
+ issue: f?.issue || "",
88
+ fingerprint: f?.fingerprint || null,
89
+ corroboration: f?.corroboration?.count ?? null,
90
+ of: f?.corroboration?.of ?? null,
91
+ verdict: verdictOf(triage, f?.fingerprint, f?.issue, f?.file),
92
+ });
93
+ }
94
+ }
95
+ out.push({
96
+ iteration: it?.iteration ?? null,
97
+ reviewers: Object.keys(map).length ? map : null,
98
+ findings: rows,
99
+ });
100
+ }
101
+
102
+ if (JSON_OUT) {
103
+ process.stdout.write(JSON.stringify(out, null, 2) + "\n");
104
+ return;
105
+ }
106
+
107
+ const lines = [];
108
+ for (const block of out) {
109
+ lines.push(`## Council view - iteration ${block.iteration ?? "?"}`);
110
+ lines.push("");
111
+ if (block.findings.length === 0) {
112
+ lines.push("No reviewer returned a finding in this iteration.");
113
+ lines.push("");
114
+ continue;
115
+ }
116
+ // Per-model totals first: the ratio is the summary, the rows are the evidence.
117
+ const byModel = new Map();
118
+ for (const r of block.findings) {
119
+ const e = byModel.get(r.model) || { raw: 0, accepted: 0 };
120
+ e.raw += 1;
121
+ if (r.verdict === "accepted") e.accepted += 1;
122
+ byModel.set(r.model, e);
123
+ }
124
+ lines.push("| Reviewer | Raised | Accepted |");
125
+ lines.push("| -------- | ------ | -------- |");
126
+ for (const [m, e] of byModel) lines.push(`| ${m} | ${e.raw} | ${e.accepted} |`);
127
+ lines.push("");
128
+ lines.push("| Reviewer | Severity | Where | Finding | Seen by | Triage |");
129
+ lines.push("| -------- | -------- | ----- | ------- | ------- | ------ |");
130
+ for (const r of block.findings) {
131
+ const seen = r.corroboration == null ? "-" : `${r.corroboration}/${r.of}`;
132
+ const where = r.file ? `${r.file}:${r.line}` : "-";
133
+ lines.push(
134
+ `| ${r.model} | ${r.severity} | ${where} | ${r.issue.replace(/\|/g, "\\|")} | ${seen} | ${r.verdict} |`,
135
+ );
136
+ }
137
+ lines.push("");
138
+ }
139
+ // `process.exitCode`, never `process.exit()`: stdout to a pipe is async and
140
+ // exiting on the next line truncates a long table at the pipe buffer.
141
+ process.stdout.write(lines.join("\n"));
142
+ }
143
+
144
+ main();
@@ -570,7 +570,10 @@ tracker_next_hint() {
570
570
  [ "${TRACKER_QUIET:-0}" = "1" ] && return 0
571
571
  name=$(load_state | jq -r --arg id "$pid" '(.phases[] | select(.id == $id) | .name) // ""')
572
572
  case "$(host_kind)" in
573
- claude) mirror="TaskUpdate(\"Phase $pid $name\", status=\"$status\")" ;;
573
+ # Not every Claude Code build ships the tile API, and the shell cannot see
574
+ # the model's tool list. Naming the fallback on the same line is what keeps a
575
+ # build without TaskUpdate from advancing eight phases in silence.
576
+ claude) mirror="TaskUpdate(\"Phase $pid $name\", status=\"$status\") - or, if TaskUpdate is not one of your tools, reprint the card above inside your reply text" ;;
574
577
  codex) mirror="update_plan: set step \"Phase $pid $name\" to $status (send the FULL step list, it is not a delta)" ;;
575
578
  *) mirror="no task widget on this host - reprint the card above inside your reply text" ;;
576
579
  esac
@@ -1166,10 +1169,24 @@ GATE
1166
1169
  }
1167
1170
  case "$(host_kind)" in
1168
1171
  claude)
1172
+ # The native tile API is not in every Claude Code build. This script
1173
+ # cannot probe the model's tool list, and a run that fires nothing and
1174
+ # says nothing is the failure that reached a user: the tracker state was
1175
+ # written correctly, every phase advanced, and the screen stayed empty
1176
+ # for the whole run. So the fallback is named here rather than assumed,
1177
+ # and the branch is taken where the information actually lives - the
1178
+ # model knows which tools it has, the shell does not.
1169
1179
  echo "REQUIRED - create one native tile per phase, in this exact order,"
1170
1180
  echo "BEFORE any TaskUpdate. The widget renders by creation order, not by"
1171
1181
  echo "phase number, so an out-of-order call scrambles the stack."
1172
1182
  echo "$tiles_state" | jq -r '.phases[] | " TaskCreate(subject: \"Phase \(.id) \(.name)\")"'
1183
+ echo
1184
+ echo "IF TaskCreate IS NOT ONE OF YOUR TOOLS this build has no native tile"
1185
+ echo "API. Do not skip the signal: the card below IS the widget then, and"
1186
+ echo "it must be reprinted inside your reply at every phase boundary - "
1187
+ echo "tool output is collapsed, so a card left in stdout never arrives."
1188
+ echo
1189
+ render
1173
1190
  ;;
1174
1191
  codex)
1175
1192
  echo "REQUIRED - register the plan in ONE update_plan call with this step"