@mmerterden/multi-agent-pipeline 16.27.0 → 16.28.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +71 -1
- package/package.json +4 -4
- package/pipeline/commands/multi-agent/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/issue/SKILL.md +1 -0
- package/pipeline/commands/multi-agent/jira/SKILL.md +3 -0
- package/pipeline/commands/multi-agent/log/SKILL.md +7 -1
- package/pipeline/commands/multi-agent/review-issue/SKILL.md +1 -0
- package/pipeline/lib/issue-fetcher.sh +134 -5
- package/pipeline/lib/multi-repo-pipeline.sh +8 -0
- package/pipeline/multi-agent-refs/analysis/evidence.md +1 -1
- package/pipeline/multi-agent-refs/analysis/intake.md +11 -2
- package/pipeline/multi-agent-refs/analysis/locked.md +1 -0
- package/pipeline/multi-agent-refs/analysis/redesign.md +112 -0
- package/pipeline/multi-agent-refs/analysis/render.md +5 -0
- package/pipeline/multi-agent-refs/analysis/resolve.md +1 -0
- package/pipeline/multi-agent-refs/analysis/review.md +15 -0
- package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
- package/pipeline/multi-agent-refs/analysis-template-corporate.md +3 -3
- package/pipeline/multi-agent-refs/analysis-template.md +36 -0
- package/pipeline/multi-agent-refs/cross-cli-contract.md +3 -0
- package/pipeline/multi-agent-refs/features/jira-context.md +101 -0
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -1
- package/pipeline/multi-agent-refs/readiness-review.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +19 -0
- package/pipeline/schemas/prefs.schema.json +26 -0
- package/pipeline/scripts/anonymize-findings.mjs +24 -0
- package/pipeline/scripts/build-references.mjs +10 -6
- package/pipeline/scripts/council-view.mjs +144 -0
- package/pipeline/scripts/phase-tracker.sh +18 -1
- package/pipeline/scripts/skill-siblings.mjs +41 -8
- package/pipeline/scripts/validate-analysis-doc.mjs +371 -7
- package/pipeline/scripts/validate-analysis.mjs +7 -5
- package/pipeline/skills/shared/core/multi-agent-issue/SKILL.md +1 -0
- package/pipeline/skills/shared/core/multi-agent-review-issue/SKILL.md +1 -0
|
@@ -86,6 +86,21 @@ Fable triages the pooled findings exactly as in `/multi-agent:review`: drop dupl
|
|
|
86
86
|
|
|
87
87
|
The verdict names the profile it judged against, the counts per severity, and what was NOT checked (the references gate without a state file, anything the fetch could not reach). Follow the pipeline rule on claiming: state which findings are mechanically proven and which are judgement.
|
|
88
88
|
|
|
89
|
+
**Every rubric class is printed, including the empty ones.** A verdict that lists only the classes that fired says nothing about the rest: the reader cannot tell "checked, clean" from "never looked". Print A through F in order with a count each, and write `none` where the count is zero. Same rule for the deterministic gates: name each one and its verdict, including the ones that skipped and why.
|
|
90
|
+
|
|
91
|
+
**Colophon.** Close the verdict with the reviewed document's `evidence_digest` and `base_commit`, taken from its front-matter, plus the reviewer models used. Two reviews of the same feature on the same day are otherwise indistinguishable, and the first question anyone asks of an older verdict is which version of the document it judged. `validate-analysis-doc.mjs --report` prints the per-check verdicts to paste under it.
|
|
92
|
+
|
|
93
|
+
```
|
|
94
|
+
Verdict: 2 Blocker, 3 Important, 1 Suggestion (profile: global, mode: full)
|
|
95
|
+
|
|
96
|
+
A evidence 2 D altitude none
|
|
97
|
+
B backbone none E admitted gaps 1
|
|
98
|
+
C ... F contradiction 3
|
|
99
|
+
|
|
100
|
+
Not checked: references coverage (no state file passed)
|
|
101
|
+
Provenance: evidence_digest sha256:1f3a9c2 base_commit 4c1b2de
|
|
102
|
+
```
|
|
103
|
+
|
|
89
104
|
## Phase 4 - Output
|
|
90
105
|
|
|
91
106
|
Default is the chat report. Nothing is written anywhere without an explicit choice.
|
|
@@ -88,7 +88,7 @@ For each `platform` in `state.analysisSpec.platforms[]`:
|
|
|
88
88
|
- `web` → `evidence.standards[]` entries matching `react`, `vue`, `next`, `sveltekit` → `~/.claude/rules/code-style.md`
|
|
89
89
|
2. **Apply per-platform omission rules.** Backend-only file drops Sections 5, 6, 7, 8, 16. Web with no UI inventory still keeps 5 (UI exists in code). Sections 1, 2, 4, 9, 13, 14, 20, 21 always present per Locked decision 2 + 13.
|
|
90
90
|
3. **Resolve mode.** If user passed `--lite` → Lite. If user passed `--full` → Full. Otherwise use `state.analysisSpec.liteModeAuto`. Lite mode renders only Sections 1, 2, 4, 9, 13, 14, 21 plus optional 23.
|
|
91
|
-
4. **Produce YAML front-matter header** (see `$HOME/.claude/multi-agent-refs/analysis-template.md`). Include `profile: <state.analysisSpec.profile | global>` and `platform: <platform | none>` so the validator applies the right contract per profile (Locked 32) and recognises the stack-optional render (Locked 35), `mode: full | lite`, plus `ui_tests: <state.analysisSpec.options.uiTests | false
|
|
91
|
+
4. **Produce YAML front-matter header** (see `$HOME/.claude/multi-agent-refs/analysis-template.md`). Include `profile: <state.analysisSpec.profile | global>` and `platform: <platform | none>` so the validator applies the right contract per profile (Locked 32) and recognises the stack-optional render (Locked 35), `mode: full | lite`, plus `ui_tests: <state.analysisSpec.options.uiTests | false>`, `a11y_depth: <state.analysisSpec.options.a11yDepth | basic>` and `redesign: <state.analysisSpec.options.redesign | false>` so the pre-dispatch validator can enforce the opt-in coverage (15.6 present when ui_tests, 16.2 walkthrough present when a11y_depth is full, 4.5 / 4.6 / 9.5 present when redesign). `status` is written only by `/multi-agent:analysis-resolve`; a rendered document is a draft.
|
|
92
92
|
5. **Read conventions for this platform's repo.** For each cell Pass B fills in Section 13 and in any per-platform projection (Sections 5, 6, 7, 8, 10, 11, 13, 14, 15, 16, 17), read `state.analysisSpec.evidence.conventions[<repo>].<field>` and emit the value with a footnote (Locked 24). If `conventionOverrides` has an entry for that field, use the override and footnote with `^[user-override: <reason>]` instead of evidence path.
|
|
93
93
|
6. **Concatenate non-null sections in canonical order.** Numbering stays sequential `1..N` over the rendered set (omitted sections do not create gaps).
|
|
94
94
|
7. **Schema validation** on the per-platform spec object:
|
|
@@ -433,12 +433,12 @@ Never omitted.
|
|
|
433
433
|
## 20. Riskler ve Açık Sorular <!-- TR -->
|
|
434
434
|
## 20. Risks and Open Questions <!-- EN -->
|
|
435
435
|
|
|
436
|
-
|
|
|
436
|
+
| Id | Konu | Neden açık | Kime sorulacak | Etkilediği bölüm |
|
|
437
437
|
|---|---|---|---|---|
|
|
438
|
-
|
|
|
438
|
+
| AS-01 | <question> | <what evidence is missing> | <role or team> | 2.3 |
|
|
439
439
|
```
|
|
440
440
|
|
|
441
|
-
The source corporate documents keep open questions out of the page and raise them in conversation instead. This profile keeps them in the document deliberately: an `EKLENECEK` with no matching row here is an unanswered question nobody owns. Every `EKLENECEK` and every unverified assumption in 2.4 emits a row (Locked 33).
|
|
441
|
+
The source corporate documents keep open questions out of the page and raise them in conversation instead. This profile keeps them in the document deliberately: an `EKLENECEK` with no matching `AS-NN` row here is an unanswered question nobody owns. Every `EKLENECEK` and every unverified assumption in 2.4 emits a row (Locked 33).
|
|
442
442
|
|
|
443
443
|
## 21. Referanslar / References
|
|
444
444
|
|
|
@@ -46,6 +46,8 @@ platform: ios | android | web | backend | mobile | none
|
|
|
46
46
|
profile: global | corporate
|
|
47
47
|
language: tr | en
|
|
48
48
|
mode: full | lite
|
|
49
|
+
status: draft | final
|
|
50
|
+
redesign: true | false
|
|
49
51
|
ui_tests: true | false
|
|
50
52
|
a11y_depth: basic | full
|
|
51
53
|
generated: <ISO 8601 UTC timestamp>
|
|
@@ -64,6 +66,8 @@ template_version: v3
|
|
|
64
66
|
|
|
65
67
|
`ui_tests` and `a11y_depth` record the Phase 0 Step 5a opt-ins (defaults `false` / `basic`) so the coverage choice is auditable and the pre-dispatch validator can enforce it: `ui_tests: true` requires Section 15.6, `a11y_depth: full` requires the Section 16.2 walkthrough.
|
|
66
68
|
|
|
69
|
+
`status` is the publication claim; absent means `draft`. Under `final` an unresolved `AS-NN` or open Section 20 row is an ERROR, under `draft` it is counted. `/multi-agent:analysis-resolve` flips it, after a confirmation.
|
|
70
|
+
|
|
67
71
|
`profile` names the template the document was rendered against (Locked 32) and `platform: none` marks the stack-optional render (Locked 35). Both are read by `validate-analysis-doc.mjs`, which applies a different contract per profile: without the key a corporate document would be judged against the global rules and its backbone would read as a pile of Locked 2 violations.
|
|
68
72
|
|
|
69
73
|
Phase 1 compares `evidence_digest` against an existing document to decide whether to reuse it (Locked 27). Phase 3 reads `platform` to verify file match, `mode` to know which section set to expect, and both `evidence_digest` and `base_commit` to judge freshness: the digest says the evidence changed, `base_commit` says the repo moved. Phase 2 parses the block but gates only on `template_version`.
|
|
@@ -284,6 +288,27 @@ EARS states the rule; Gherkin states how you check it. Each rule still maps to a
|
|
|
284
288
|
|
|
285
289
|
Each scenario in 4.1-4.3 cites the matching `BR-<slug>-NN` id (and, when a server code applies, the Section 9.4 code). Plain-prose user stories without Given/When/Then are rejected by the renderer.
|
|
286
290
|
|
|
291
|
+
### 4.5 Mevcut Davranış / Current Behaviour (OPTIONAL)
|
|
292
|
+
|
|
293
|
+
`redesign: true` only. Locked 37; contract and checks in `analysis/redesign.md`. Evidence and certainty are derived from `repoEvidence`, never graded by the writer.
|
|
294
|
+
|
|
295
|
+
```markdown
|
|
296
|
+
| Id | Davranış / Behaviour | Kanıt / Evidence | Kesinlik / Certainty |
|
|
297
|
+
|---|---|---|---|
|
|
298
|
+
| CB-<slug>-01 | <what v1 does, present tense, no v2 language> | <repo>/<path>:<line> | confirmed |
|
|
299
|
+
| CB-<slug>-02 | <what v1 does> | <module> (cross-cutting) | uncertain |
|
|
300
|
+
```
|
|
301
|
+
|
|
302
|
+
### 4.6 Endpoint Eşlemesi / Endpoint Mapping (OPTIONAL)
|
|
303
|
+
|
|
304
|
+
`redesign: true` only. An unchanged path still gets its row: "unchanged" is an answer, its absence is not.
|
|
305
|
+
|
|
306
|
+
```markdown
|
|
307
|
+
| v1 | v2 | Değişiklik / Change |
|
|
308
|
+
|---|---|---|
|
|
309
|
+
| GET /v1/<path> | GET /v2/<path> | <shape change, or unchanged> |
|
|
310
|
+
```
|
|
311
|
+
|
|
287
312
|
## 5. Tasarım Referansı / Design Reference
|
|
288
313
|
|
|
289
314
|
Per-platform projection. Omitted if no Figma URL AND no Code Connect mapping. Locked 17 + 18.
|
|
@@ -484,6 +509,17 @@ List every HTTP status code one by one.
|
|
|
484
509
|
| ERR-211 | Empty result set | <screen> |
|
|
485
510
|
```
|
|
486
511
|
|
|
512
|
+
### 9.5 Fark Listesi / Difference List (OPTIONAL)
|
|
513
|
+
|
|
514
|
+
`redesign: true` only. One row per `CB-` id, both directions checked. Closed status vocabulary - `Moved`, `Partial`, `Missing`, `New`, `Out of scope` - in a bilingual cell like Section 20's `Açık / Open`. Every `Missing` and `Partial` owes a Section 20 row by `AS-NN`.
|
|
515
|
+
|
|
516
|
+
```markdown
|
|
517
|
+
| Id | Durum / Status | Nereye / Where | AS-NN |
|
|
518
|
+
|---|---|---|---|
|
|
519
|
+
| CB-<slug>-01 | Taşındı / Moved | 4.2 | - |
|
|
520
|
+
| CB-<slug>-02 | Eksik / Missing | - | AS-03 |
|
|
521
|
+
```
|
|
522
|
+
|
|
487
523
|
## 10. Lokalizasyon Anahtarları / Localization Keys
|
|
488
524
|
|
|
489
525
|
Per-platform projection. Locked 20 - shape depends on the project `figma-config` `localization.ownership`. The locale set comes from `localization.locales` (default `tr, en, ar, de, es, fr, it, ru`); never hardcode a locale list.
|
|
@@ -230,6 +230,9 @@ argument-hint: "<input hint>"
|
|
|
230
230
|
|---|---|
|
|
231
231
|
| Claude → Copilot | Add `name`, `user-invocable: true`, `argument-hint`; remove `allowed-tools`; `description` stays English (localization happens only in the installed Claude tree via `description-tr` + `localize-commands.mjs`, never in repo or Copilot files) |
|
|
232
232
|
| Copilot → Claude | Remove `name`, `user-invocable`, `argument-hint`; add `allowed-tools` (inferred from command's actual tool calls) |
|
|
233
|
+
| Both directions | `not-for: <sibling>` carries across unchanged |
|
|
234
|
+
|
|
235
|
+
**`not-for`** is optional and lists sibling commands this one must not be chosen for, as bare names (`not-for: jira, review-jira`). It is written for the model reading the routing surface, so it belongs beside the description rather than in prose further down. `lint-skills.mjs` also consumes it: a trigger-vocabulary collision either side has named drops off the warning list and is reported as settled instead, and a name that resolves to no sibling on the same surface is an error - a typo would otherwise leave the pair unanswered while the author believes it is handled. Both trees spell the value the same way; the `multi-agent-` prefix on the Copilot side is stripped before matching.
|
|
233
236
|
|
|
234
237
|
**Invariant**: the English `description` must have the same meaning on both sides, and a `description-tr` line (when present) must be a faithful translation of it. Semantic drift is not allowed. `description-en` sidecars are an installed-tree artifact and must never appear in repo or Copilot files.
|
|
235
238
|
|
|
@@ -0,0 +1,101 @@
|
|
|
1
|
+
# Related-issue context at intake
|
|
2
|
+
|
|
3
|
+
A development sub-task is often filed with no description of its own. The
|
|
4
|
+
requirement sits on the parent, and the rest of the picture - the analysis, the
|
|
5
|
+
test scope - sits on the sibling sub-tasks beside it. The fetcher already read
|
|
6
|
+
the parent; it did not read the siblings, so a board that keeps its analysis in a
|
|
7
|
+
separate sub-task handed the pipeline an empty task.
|
|
8
|
+
|
|
9
|
+
`issue-fetcher.sh` now reads those siblings and exposes them as
|
|
10
|
+
`descriptor.relatedIssues[]`.
|
|
11
|
+
|
|
12
|
+
## What it costs
|
|
13
|
+
|
|
14
|
+
One search request, and only when the issue has a parent:
|
|
15
|
+
|
|
16
|
+
```
|
|
17
|
+
GET /rest/api/2/search?jql=parent="<PARENT>"&fields=summary,issuetype,status,description&maxResults=20
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
An issue with no parent has no siblings, so the lookup is skipped entirely and
|
|
21
|
+
that run pays nothing. The count does not grow with the number of siblings: the
|
|
22
|
+
issue's own `subtasks` field lists ITS children rather than its siblings, and the
|
|
23
|
+
parent's lists siblings without their descriptions, so either of those shapes
|
|
24
|
+
would cost one GET per sibling. `smoke-jira-context.sh` counts the requests, so
|
|
25
|
+
this stays a measurement rather than a claim.
|
|
26
|
+
|
|
27
|
+
The request lands before the maturity verdict reaches the user, which is one
|
|
28
|
+
extra round trip on the intake path. `jiraContext.enabled: false` is the way out
|
|
29
|
+
when latency matters more than the context.
|
|
30
|
+
|
|
31
|
+
## Shape
|
|
32
|
+
|
|
33
|
+
```json
|
|
34
|
+
"relatedIssues": [
|
|
35
|
+
{
|
|
36
|
+
"key": "PROJ-1002",
|
|
37
|
+
"relation": "sibling",
|
|
38
|
+
"type": "Sub-task",
|
|
39
|
+
"status": "Done",
|
|
40
|
+
"summary": "...",
|
|
41
|
+
"description": "...",
|
|
42
|
+
"truncated": false
|
|
43
|
+
}
|
|
44
|
+
]
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
`type` is carried through verbatim from Jira and is never matched against in
|
|
48
|
+
code. Boards name their sub-task types differently and the query does not need
|
|
49
|
+
to know: `parent = <key>` returns all of them.
|
|
50
|
+
|
|
51
|
+
Siblings that carry a description are ordered first, so `maxItems` drops the
|
|
52
|
+
empty ones rather than the useful ones. An empty sibling is still listed - its
|
|
53
|
+
key, type and status are information - with an empty `description`.
|
|
54
|
+
|
|
55
|
+
## Settings
|
|
56
|
+
|
|
57
|
+
`prefs.global.jiraContext`:
|
|
58
|
+
|
|
59
|
+
| Key | Default | Effect |
|
|
60
|
+
|---|---|---|
|
|
61
|
+
| `enabled` | `true` | `false` skips the search entirely |
|
|
62
|
+
| `maxItems` | `6` | Hard cap; `0` behaves like `enabled: false` |
|
|
63
|
+
| `maxCharsPerItem` | `1200` | Per-sibling truncation, marked `truncated: true` |
|
|
64
|
+
|
|
65
|
+
Worst case payload is `maxItems x maxCharsPerItem`, next to the 4000-character
|
|
66
|
+
cap on the issue's own description.
|
|
67
|
+
|
|
68
|
+
## Maturity
|
|
69
|
+
|
|
70
|
+
One new code, `description_empty_sibling_available`, and it fires narrowly: the
|
|
71
|
+
issue's own description is empty, the parent's is empty or unreachable, AND a
|
|
72
|
+
sibling carries content. Today that combination produces the hard
|
|
73
|
+
`description_empty` blocker; it becomes a warning instead, the same downgrade
|
|
74
|
+
`description_empty_parent_available` already makes one level up.
|
|
75
|
+
|
|
76
|
+
The narrowness is deliberate. `score` is `100 - 40*blockers - 10*warnings`, and
|
|
77
|
+
the picker continues silently only at `score >= 90`, so a warning that fired
|
|
78
|
+
whenever an issue merely HAD siblings would cost ten points on every parented
|
|
79
|
+
issue and turn a silent intake into a question across the board.
|
|
80
|
+
|
|
81
|
+
## Where it goes
|
|
82
|
+
|
|
83
|
+
The descriptor is carried verbatim into the picker state as `issueRef`, and the
|
|
84
|
+
bridge copies `relatedIssues` onto `agent-state.json` when it is non-empty.
|
|
85
|
+
|
|
86
|
+
Persisting it is the point: `/multi-agent:resume` rebuilds context from durable
|
|
87
|
+
artefacts and never from the conversation, so an intake-only enrichment would be
|
|
88
|
+
gone by the first resume. The task's own `description` and `maturity` are still
|
|
89
|
+
not persisted, which is a known asymmetry rather than an oversight - this
|
|
90
|
+
change did not widen the bridge beyond the one field it needed.
|
|
91
|
+
|
|
92
|
+
Sibling text is added as its own labelled block. It never replaces the working
|
|
93
|
+
description the way an accepted `parentDescription` does: substituting it would
|
|
94
|
+
erase the provenance that Phase 2's parent-story scope-drift check reads.
|
|
95
|
+
|
|
96
|
+
## Not included
|
|
97
|
+
|
|
98
|
+
Following the parent's `issuelinks` in a second search. A linked bug can be the
|
|
99
|
+
whole requirement while a linked test-execution record is noise, and the fetcher
|
|
100
|
+
cannot tell them apart, so the scope stops at sub-tasks. There is no setting for
|
|
101
|
+
it: a switch that does nothing is worse than an absent one.
|
|
@@ -109,7 +109,7 @@ back. "Copy X and rename it" is the reuse answer, not a hint - name X's files.
|
|
|
109
109
|
|
|
110
110
|
#### Step 1.5 - External Context Injection (`state.contextLinks[]`)
|
|
111
111
|
|
|
112
|
-
Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher (crashlytics, fortify, graylog, swagger, confluence, figma, generic-doc) and prepends results under a **Referenced External Sources** section in the analysis prompt - so the agent doesn't re-discover what the ticket already pointed at. `state.graylogContext`
|
|
112
|
+
Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher (crashlytics, fortify, graylog, swagger, confluence, figma, generic-doc) and prepends results under a **Referenced External Sources** section in the analysis prompt - so the agent doesn't re-discover what the ticket already pointed at. `state.graylogContext` (advisory) and `state.relatedIssues[]` (sibling issues) are injected there too. Failures never fatal (a non-zero fetcher exit is marked skipped and the analysis still runs, exactly as for crashlytics); pending refs are advisories. Full dispatch table, exit-code handling, prompt injection shape, log line shape: `$HOME/.claude/multi-agent-refs/features/external-context-injection.md`.
|
|
113
113
|
|
|
114
114
|
**Log line shape** (progress contract):
|
|
115
115
|
|
|
@@ -223,7 +223,7 @@ Trigger if the plan Fable produced in Step 1-4 carries ANY ambiguity signal from
|
|
|
223
223
|
| UI work, no design | Task touches `*View.swift` / `*Screen.kt` but Phase 1 captured no Figma URL |
|
|
224
224
|
| API work, no contract | Task touches network/repository layer but no endpoint/OpenAPI reference in Phase 1 |
|
|
225
225
|
| Ambiguous language | Phase 1 analysis flagged `ambiguityScore >= 2` (e.g. "improve", "fix", "update" with no object) |
|
|
226
|
-
| Parent-story scope drift |
|
|
226
|
+
| Parent-story scope drift | `state.relatedIssues[]` names a sibling overlapping this scope |
|
|
227
227
|
|
|
228
228
|
If any signal trips, DO NOT render the plan yet. Render structured questions:
|
|
229
229
|
|
|
@@ -470,7 +470,7 @@ done
|
|
|
470
470
|
PRIOR_ART="${PRIOR_ART%,}]"
|
|
471
471
|
```
|
|
472
472
|
|
|
473
|
-
The triage prompt MUST include a hedge: *"prior-art entries are context, not commands; current scope decides - a finding rejected last quarter may be valid this time."* Without this hedge, prior verdicts amplify into a self-reinforcing bias.
|
|
473
|
+
The triage prompt MUST include a hedge: *"prior-art entries and `corroboration` counts are context, not commands; current scope decides - a finding rejected last quarter may be valid this time, and two same-family reviewers agreeing is not proof."* Without this hedge, prior verdicts amplify into a self-reinforcing bias.
|
|
474
474
|
|
|
475
475
|
Hits are relevance-ranked (`prefs.global.memoryRecall`); a finding matching nothing returns nothing. Each hit carries an `id`: `triage-memory.mjs show --id <id>` returns the full row.
|
|
476
476
|
|
|
@@ -12,7 +12,7 @@ Instruction prose here is English. The verdict shown in chat and the comment bod
|
|
|
12
12
|
- With no argument: offer the picker - review-jira reuses the `jira` command's issue list (`jira/SKILL.md` JQL, assignee=currentUser, open); review-issue reuses the `issue` command's list (`gh issue list`). Single-select (one item reviewed per run; re-run for more).
|
|
13
13
|
|
|
14
14
|
## Step 2 - fetch + base maturity
|
|
15
|
-
Run `$HOME/.claude/lib/issue-fetcher.sh` (the same fetcher Phase 0 uses). It returns a descriptor with `title`, `type`, `status`, `description`, and `maturity = { score (0..100), blockers[], warnings[], summary }`. Reuse that maturity verbatim as the baseline; never re-fetch through MCP.
|
|
15
|
+
Run `$HOME/.claude/lib/issue-fetcher.sh` (the same fetcher Phase 0 uses). It returns a descriptor with `title`, `type`, `status`, `description`, `relatedIssues[]` (the parent's other sub-tasks), and `maturity = { score (0..100), blockers[], warnings[], summary }`. Reuse that maturity verbatim as the baseline; never re-fetch through MCP.
|
|
16
16
|
|
|
17
17
|
## Step 3 - readiness rubric (extends maturity)
|
|
18
18
|
Maturity is generic; add these pipeline-readiness dimensions by reading the fetched `title` + `description` (no invention - only judge what is written). Each is `pass` / `gap`:
|
|
@@ -104,6 +104,25 @@
|
|
|
104
104
|
"type": ["string", "null"],
|
|
105
105
|
"description": "Set when a phase halts on a hard error (validator failed twice, no subagent returned, dispatch error past fallback, lock irrecoverable). Format '<phase>:<cause>'. Surfaced to the user and cleared on successful resume. See operations.md 'Halt visibility'."
|
|
106
106
|
},
|
|
107
|
+
"relatedIssues": {
|
|
108
|
+
"type": "array",
|
|
109
|
+
"maxItems": 20,
|
|
110
|
+
"description": "Jira issues fetched alongside the task at intake: the other sub-tasks under this issue's parent, where a board keeps the analysis and the test scope. NOT the same field as siblings[] below, which is read-only sibling REPOS. Written by the picker bridge from descriptor.relatedIssues; read by Phase 1 as ground truth and by Phase 2's parent-story scope-drift check. Durable on purpose: /multi-agent:resume rebuilds context from artefacts, never from the conversation, so intake-only enrichment would vanish on the first resume. The task's own description and maturity are still NOT persisted here, which is a known asymmetry, not an oversight.",
|
|
111
|
+
"items": {
|
|
112
|
+
"type": "object",
|
|
113
|
+
"additionalProperties": false,
|
|
114
|
+
"required": ["key", "relation"],
|
|
115
|
+
"properties": {
|
|
116
|
+
"key": { "type": "string" },
|
|
117
|
+
"relation": { "type": "string", "enum": ["sibling"] },
|
|
118
|
+
"type": { "type": "string" },
|
|
119
|
+
"status": { "type": "string" },
|
|
120
|
+
"summary": { "type": "string" },
|
|
121
|
+
"description": { "type": "string" },
|
|
122
|
+
"truncated": { "type": "boolean", "default": false }
|
|
123
|
+
}
|
|
124
|
+
}
|
|
125
|
+
},
|
|
107
126
|
"siblings": {
|
|
108
127
|
"type": "array",
|
|
109
128
|
"maxItems": 10,
|
|
@@ -1358,6 +1358,32 @@
|
|
|
1358
1358
|
"description": "v12.8+ - Phase 4 Step 1.77 reviewer-scope gate. Decides the reviewer count from the deterministic diff-risk report via `pipeline/scripts/review-scope.mjs`: a diff under 20 lines of churn with max_score < 3.0 and no security_path / migration / public_api / no_test_change / test_lines_removed signal runs ONE reviewer instead of the full CLI-aware set (2 on Claude Code, 3 on Copilot CLI). No LLM. Fails safe in one direction only - any error, empty report or validator rejection resolves to the full set, because skipping a reviewer trades coverage for cost. Set false to force the full set on every diff.",
|
|
1359
1359
|
"$comment": "Shipped inert for a release: the script existed, was unit- and smoke-tested, and no phase doc referenced it, so every diff paid for the full reviewer set. Wired in Step 1.77; smoke-gate-wiring.sh keeps it reachable."
|
|
1360
1360
|
},
|
|
1361
|
+
"jiraContext": {
|
|
1362
|
+
"type": "object",
|
|
1363
|
+
"additionalProperties": false,
|
|
1364
|
+
"description": "v16.28+ - when a Jira issue has a parent, the intake fetcher also reads the parent's other sub-tasks in ONE extra search request and exposes them as descriptor.relatedIssues[]. A development sub-task filed with an empty description whose analysis sibling carries the real requirement is the case this exists for. Costs nothing on an issue with no parent, because an issue with no parent has no siblings. The extra round trip lands before the maturity verdict is shown, so turn it off when intake latency matters more than the context.",
|
|
1365
|
+
"properties": {
|
|
1366
|
+
"enabled": {
|
|
1367
|
+
"type": "boolean",
|
|
1368
|
+
"default": true,
|
|
1369
|
+
"description": "Master switch. When false the fetcher never issues the search and relatedIssues[] is always empty."
|
|
1370
|
+
},
|
|
1371
|
+
"maxItems": {
|
|
1372
|
+
"type": "integer",
|
|
1373
|
+
"default": 6,
|
|
1374
|
+
"minimum": 0,
|
|
1375
|
+
"maximum": 20,
|
|
1376
|
+
"description": "Hard cap on relatedIssues[]. Siblings that carry a description are kept first, so the cap drops the empty ones. 0 disables the lookup exactly like enabled: false."
|
|
1377
|
+
},
|
|
1378
|
+
"maxCharsPerItem": {
|
|
1379
|
+
"type": "integer",
|
|
1380
|
+
"default": 1200,
|
|
1381
|
+
"minimum": 0,
|
|
1382
|
+
"maximum": 4000,
|
|
1383
|
+
"description": "Per-sibling description truncation, marked with truncated: true. Worst case payload is maxItems x this, next to the 4000-char cap on the issue's own description."
|
|
1384
|
+
}
|
|
1385
|
+
}
|
|
1386
|
+
},
|
|
1361
1387
|
"priorArtEnrichment": {
|
|
1362
1388
|
"type": "object",
|
|
1363
1389
|
"additionalProperties": false,
|
|
@@ -122,6 +122,30 @@ export function anonymize(input, { seed } = {}) {
|
|
|
122
122
|
}
|
|
123
123
|
});
|
|
124
124
|
|
|
125
|
+
// How many DIFFERENT reviewers reported this finding, counted from the
|
|
126
|
+
// fingerprints the validator already stamped. It costs nothing: every
|
|
127
|
+
// reviewer's findings are in this one payload and nowhere else are they
|
|
128
|
+
// together. `count` names no reviewer, so it survives the identity strip
|
|
129
|
+
// without reopening what the strip closed.
|
|
130
|
+
//
|
|
131
|
+
// It is context for triage, never a rule. ADR-0001 rejected majority voting
|
|
132
|
+
// precisely because the same hallucination shows up in two of three
|
|
133
|
+
// same-family reviewers, so a high count is not evidence of truth. The
|
|
134
|
+
// reading that pays is the other end: count 1 is the finding only one
|
|
135
|
+
// reviewer saw, which is where a judge should look hardest.
|
|
136
|
+
const seen = new Map();
|
|
137
|
+
for (const f of tagged) {
|
|
138
|
+
const k = typeof f?.fingerprint === "string" && f.fingerprint ? f.fingerprint : null;
|
|
139
|
+
if (!k) continue;
|
|
140
|
+
if (!seen.has(k)) seen.set(k, new Set());
|
|
141
|
+
seen.get(k).add(f.foundBy);
|
|
142
|
+
}
|
|
143
|
+
const of = reviewers.length;
|
|
144
|
+
for (const f of tagged) {
|
|
145
|
+
const k = typeof f?.fingerprint === "string" && f.fingerprint ? f.fingerprint : null;
|
|
146
|
+
f.corroboration = { count: k ? seen.get(k).size : 1, of };
|
|
147
|
+
}
|
|
148
|
+
|
|
125
149
|
tagged.sort((a, b) => {
|
|
126
150
|
const ka = stableKey(a);
|
|
127
151
|
const kb = stableKey(b);
|
|
@@ -338,7 +338,8 @@ function main() {
|
|
|
338
338
|
process.stderr.write(
|
|
339
339
|
"usage: build-references.mjs <state.json|-> [--lang tr|en] [--check <doc.md>]\n",
|
|
340
340
|
);
|
|
341
|
-
process.
|
|
341
|
+
process.exitCode = 2;
|
|
342
|
+
return;
|
|
342
343
|
}
|
|
343
344
|
|
|
344
345
|
let spec;
|
|
@@ -346,13 +347,15 @@ function main() {
|
|
|
346
347
|
spec = readState(statePath);
|
|
347
348
|
} catch (err) {
|
|
348
349
|
process.stderr.write(`ERROR: cannot read state: ${err.message}\n`);
|
|
349
|
-
process.
|
|
350
|
+
process.exitCode = 2;
|
|
351
|
+
return;
|
|
350
352
|
}
|
|
351
353
|
|
|
352
354
|
const lang = argOf(argv, "--lang") ?? spec.language ?? "tr";
|
|
353
355
|
if (!HEADERS[lang]) {
|
|
354
356
|
process.stderr.write(`ERROR: unknown language "${lang}"; expected tr or en\n`);
|
|
355
|
-
process.
|
|
357
|
+
process.exitCode = 2;
|
|
358
|
+
return;
|
|
356
359
|
}
|
|
357
360
|
|
|
358
361
|
const docPath = argOf(argv, "--check");
|
|
@@ -361,10 +364,11 @@ function main() {
|
|
|
361
364
|
if (problems.length) {
|
|
362
365
|
for (const p of problems) process.stderr.write(`${p}\n`);
|
|
363
366
|
process.stderr.write(`\n${problems.length} references coverage failure(s)\n`);
|
|
364
|
-
process.
|
|
367
|
+
process.exitCode = 1;
|
|
368
|
+
return;
|
|
365
369
|
}
|
|
366
370
|
process.stdout.write("references coverage ok\n");
|
|
367
|
-
|
|
371
|
+
return;
|
|
368
372
|
}
|
|
369
373
|
|
|
370
374
|
const rows = rowsFrom(spec, lang);
|
|
@@ -376,7 +380,7 @@ function main() {
|
|
|
376
380
|
? "Bu koşuda getirilen kaynak yok.\n"
|
|
377
381
|
: "No sources were fetched for this run.\n",
|
|
378
382
|
);
|
|
379
|
-
|
|
383
|
+
return;
|
|
380
384
|
}
|
|
381
385
|
process.stdout.write(`${renderTable(rows, lang)}\n`);
|
|
382
386
|
}
|
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// council-view.mjs - what each reviewer found, and what triage did with it.
|
|
3
|
+
//
|
|
4
|
+
// Phase 4 dispatches three reviewers, anonymizes their findings so the judge
|
|
5
|
+
// cannot mark its own homework, and reports the verdict. Everything needed to
|
|
6
|
+
// answer "which model found this, and was it kept" is already on the state -
|
|
7
|
+
// reviewers[].findings, anonymizationMap.labelToModel, triage.accepted - and
|
|
8
|
+
// nothing ever renders it. run-metrics.mjs reduces it to a ratio; the rows
|
|
9
|
+
// behind the ratio are not visible anywhere.
|
|
10
|
+
//
|
|
11
|
+
// This is the panel view llm-council shows as its stage-1 tabs, built from
|
|
12
|
+
// state that already exists. No model call, no new field, no new phase.
|
|
13
|
+
//
|
|
14
|
+
// De-anonymization is deliberate and safe here: the map is read AFTER triage
|
|
15
|
+
// has ruled, so nothing this prints can bias the judgement it describes. The
|
|
16
|
+
// map never goes into a prompt; this is a report.
|
|
17
|
+
//
|
|
18
|
+
// Usage:
|
|
19
|
+
// council-view.mjs <agent-state.json> [--json] [--iteration N]
|
|
20
|
+
//
|
|
21
|
+
// Exit codes: 0 rendered, 2 nothing to render (no review iterations), 64 usage.
|
|
22
|
+
|
|
23
|
+
import { readFileSync } from "node:fs";
|
|
24
|
+
|
|
25
|
+
const argv = process.argv.slice(2);
|
|
26
|
+
const JSON_OUT = argv.includes("--json");
|
|
27
|
+
const iterFlag = argv.indexOf("--iteration");
|
|
28
|
+
const ONLY = iterFlag !== -1 ? Number(argv[iterFlag + 1]) : null;
|
|
29
|
+
const file = argv.find((a) => !a.startsWith("--") && a !== String(ONLY));
|
|
30
|
+
// `--iteration` with nothing after it made ONLY NaN, and `it.iteration === NaN`
|
|
31
|
+
// is never true, so a state full of iterations exited 2 - which /multi-agent:log
|
|
32
|
+
// is told to read as "the run never reached Phase 4" and quietly omit. A usage
|
|
33
|
+
// error has to look like one.
|
|
34
|
+
const BAD_ITERATION = iterFlag !== -1 && !Number.isInteger(ONLY);
|
|
35
|
+
|
|
36
|
+
function die(msg, code) {
|
|
37
|
+
process.stderr.write(msg + "\n");
|
|
38
|
+
process.exitCode = code;
|
|
39
|
+
}
|
|
40
|
+
|
|
41
|
+
function verdictOf(triage, fingerprint, issue, fileName) {
|
|
42
|
+
for (const bucket of ["accepted", "deferred", "rejected"]) {
|
|
43
|
+
for (const entry of Array.isArray(triage?.[bucket]) ? triage[bucket] : []) {
|
|
44
|
+
// Only `accepted[]` is a flat finding. triage-output.schema.json defines a
|
|
45
|
+
// deferred or rejected item as `{ finding, reason }` - the reason is the
|
|
46
|
+
// point of those buckets - so reading the id off the wrapper made every
|
|
47
|
+
// ruled-out finding render as `unruled`, which is exactly the column this
|
|
48
|
+
// view exists to show.
|
|
49
|
+
const f = entry?.finding ?? entry;
|
|
50
|
+
if (fingerprint && f?.fingerprint === fingerprint) return bucket;
|
|
51
|
+
if (!fingerprint && f?.file === fileName && f?.issue === issue) return bucket;
|
|
52
|
+
}
|
|
53
|
+
}
|
|
54
|
+
return "unruled";
|
|
55
|
+
}
|
|
56
|
+
|
|
57
|
+
function main() {
|
|
58
|
+
if (!file) return die("usage: council-view.mjs <agent-state.json> [--json] [--iteration N]", 64);
|
|
59
|
+
if (BAD_ITERATION) return die("council-view: --iteration needs an integer", 64);
|
|
60
|
+
let state;
|
|
61
|
+
try {
|
|
62
|
+
state = JSON.parse(readFileSync(file, "utf8"));
|
|
63
|
+
} catch (e) {
|
|
64
|
+
return die(`council-view: cannot read ${file}: ${e.message}`, 64);
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
const iterations = Array.isArray(state?.reviewIterations) ? state.reviewIterations : [];
|
|
68
|
+
const wanted = ONLY == null ? iterations : iterations.filter((it) => it?.iteration === ONLY);
|
|
69
|
+
if (wanted.length === 0) return die("council-view: no review iterations on this state", 2);
|
|
70
|
+
|
|
71
|
+
const out = [];
|
|
72
|
+
for (const it of wanted) {
|
|
73
|
+
// Absent map is the normal state for a run written before anonymization,
|
|
74
|
+
// and for one where every reviewer timed out. Say "unknown" rather than
|
|
75
|
+
// guessing a model, the same way run-metrics.mjs degrades.
|
|
76
|
+
const map = it?.anonymizationMap?.labelToModel || {};
|
|
77
|
+
const triage = it?.triage || {};
|
|
78
|
+
const rows = [];
|
|
79
|
+
for (const r of Array.isArray(it?.reviewers) ? it.reviewers : []) {
|
|
80
|
+
const model = typeof r?.model === "string" && r.model ? r.model : "unknown";
|
|
81
|
+
for (const f of Array.isArray(r?.findings) ? r.findings : []) {
|
|
82
|
+
rows.push({
|
|
83
|
+
model,
|
|
84
|
+
severity: f?.severity || "",
|
|
85
|
+
file: f?.file || "",
|
|
86
|
+
line: f?.line ?? "",
|
|
87
|
+
issue: f?.issue || "",
|
|
88
|
+
fingerprint: f?.fingerprint || null,
|
|
89
|
+
corroboration: f?.corroboration?.count ?? null,
|
|
90
|
+
of: f?.corroboration?.of ?? null,
|
|
91
|
+
verdict: verdictOf(triage, f?.fingerprint, f?.issue, f?.file),
|
|
92
|
+
});
|
|
93
|
+
}
|
|
94
|
+
}
|
|
95
|
+
out.push({
|
|
96
|
+
iteration: it?.iteration ?? null,
|
|
97
|
+
reviewers: Object.keys(map).length ? map : null,
|
|
98
|
+
findings: rows,
|
|
99
|
+
});
|
|
100
|
+
}
|
|
101
|
+
|
|
102
|
+
if (JSON_OUT) {
|
|
103
|
+
process.stdout.write(JSON.stringify(out, null, 2) + "\n");
|
|
104
|
+
return;
|
|
105
|
+
}
|
|
106
|
+
|
|
107
|
+
const lines = [];
|
|
108
|
+
for (const block of out) {
|
|
109
|
+
lines.push(`## Council view - iteration ${block.iteration ?? "?"}`);
|
|
110
|
+
lines.push("");
|
|
111
|
+
if (block.findings.length === 0) {
|
|
112
|
+
lines.push("No reviewer returned a finding in this iteration.");
|
|
113
|
+
lines.push("");
|
|
114
|
+
continue;
|
|
115
|
+
}
|
|
116
|
+
// Per-model totals first: the ratio is the summary, the rows are the evidence.
|
|
117
|
+
const byModel = new Map();
|
|
118
|
+
for (const r of block.findings) {
|
|
119
|
+
const e = byModel.get(r.model) || { raw: 0, accepted: 0 };
|
|
120
|
+
e.raw += 1;
|
|
121
|
+
if (r.verdict === "accepted") e.accepted += 1;
|
|
122
|
+
byModel.set(r.model, e);
|
|
123
|
+
}
|
|
124
|
+
lines.push("| Reviewer | Raised | Accepted |");
|
|
125
|
+
lines.push("| -------- | ------ | -------- |");
|
|
126
|
+
for (const [m, e] of byModel) lines.push(`| ${m} | ${e.raw} | ${e.accepted} |`);
|
|
127
|
+
lines.push("");
|
|
128
|
+
lines.push("| Reviewer | Severity | Where | Finding | Seen by | Triage |");
|
|
129
|
+
lines.push("| -------- | -------- | ----- | ------- | ------- | ------ |");
|
|
130
|
+
for (const r of block.findings) {
|
|
131
|
+
const seen = r.corroboration == null ? "-" : `${r.corroboration}/${r.of}`;
|
|
132
|
+
const where = r.file ? `${r.file}:${r.line}` : "-";
|
|
133
|
+
lines.push(
|
|
134
|
+
`| ${r.model} | ${r.severity} | ${where} | ${r.issue.replace(/\|/g, "\\|")} | ${seen} | ${r.verdict} |`,
|
|
135
|
+
);
|
|
136
|
+
}
|
|
137
|
+
lines.push("");
|
|
138
|
+
}
|
|
139
|
+
// `process.exitCode`, never `process.exit()`: stdout to a pipe is async and
|
|
140
|
+
// exiting on the next line truncates a long table at the pipe buffer.
|
|
141
|
+
process.stdout.write(lines.join("\n"));
|
|
142
|
+
}
|
|
143
|
+
|
|
144
|
+
main();
|
|
@@ -570,7 +570,10 @@ tracker_next_hint() {
|
|
|
570
570
|
[ "${TRACKER_QUIET:-0}" = "1" ] && return 0
|
|
571
571
|
name=$(load_state | jq -r --arg id "$pid" '(.phases[] | select(.id == $id) | .name) // ""')
|
|
572
572
|
case "$(host_kind)" in
|
|
573
|
-
|
|
573
|
+
# Not every Claude Code build ships the tile API, and the shell cannot see
|
|
574
|
+
# the model's tool list. Naming the fallback on the same line is what keeps a
|
|
575
|
+
# build without TaskUpdate from advancing eight phases in silence.
|
|
576
|
+
claude) mirror="TaskUpdate(\"Phase $pid $name\", status=\"$status\") - or, if TaskUpdate is not one of your tools, reprint the card above inside your reply text" ;;
|
|
574
577
|
codex) mirror="update_plan: set step \"Phase $pid $name\" to $status (send the FULL step list, it is not a delta)" ;;
|
|
575
578
|
*) mirror="no task widget on this host - reprint the card above inside your reply text" ;;
|
|
576
579
|
esac
|
|
@@ -1166,10 +1169,24 @@ GATE
|
|
|
1166
1169
|
}
|
|
1167
1170
|
case "$(host_kind)" in
|
|
1168
1171
|
claude)
|
|
1172
|
+
# The native tile API is not in every Claude Code build. This script
|
|
1173
|
+
# cannot probe the model's tool list, and a run that fires nothing and
|
|
1174
|
+
# says nothing is the failure that reached a user: the tracker state was
|
|
1175
|
+
# written correctly, every phase advanced, and the screen stayed empty
|
|
1176
|
+
# for the whole run. So the fallback is named here rather than assumed,
|
|
1177
|
+
# and the branch is taken where the information actually lives - the
|
|
1178
|
+
# model knows which tools it has, the shell does not.
|
|
1169
1179
|
echo "REQUIRED - create one native tile per phase, in this exact order,"
|
|
1170
1180
|
echo "BEFORE any TaskUpdate. The widget renders by creation order, not by"
|
|
1171
1181
|
echo "phase number, so an out-of-order call scrambles the stack."
|
|
1172
1182
|
echo "$tiles_state" | jq -r '.phases[] | " TaskCreate(subject: \"Phase \(.id) \(.name)\")"'
|
|
1183
|
+
echo
|
|
1184
|
+
echo "IF TaskCreate IS NOT ONE OF YOUR TOOLS this build has no native tile"
|
|
1185
|
+
echo "API. Do not skip the signal: the card below IS the widget then, and"
|
|
1186
|
+
echo "it must be reprinted inside your reply at every phase boundary - "
|
|
1187
|
+
echo "tool output is collapsed, so a card left in stdout never arrives."
|
|
1188
|
+
echo
|
|
1189
|
+
render
|
|
1173
1190
|
;;
|
|
1174
1191
|
codex)
|
|
1175
1192
|
echo "REQUIRED - register the plan in ONE update_plan call with this step"
|