orchestrator-workflow 0.33.0 → 0.34.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,16 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.34.0] - 2026-09-13
|
|
11
|
+
|
|
12
|
+
- Reviewer findings now carry `introduced_by_delta: yes | no | unknown`.
|
|
13
|
+
A `no` attribution requires a named base build and replay in `reproduction`;
|
|
14
|
+
it remains in the ordinary finding gate and Findings table, while only `yes`
|
|
15
|
+
and `unknown` participate in bounded halt and escalation rules. The legacy
|
|
16
|
+
five-cell table and placeholder row remain byte-compatible with the
|
|
17
|
+
completeness reader; concrete rows record attribution parenthetically in
|
|
18
|
+
their Description field.
|
|
19
|
+
|
|
10
20
|
## [0.33.0] - 2026-09-12
|
|
11
21
|
|
|
12
22
|
### Changed
|
|
@@ -73,7 +73,7 @@ Check, at minimum:
|
|
|
73
73
|
review round, classify each finding as `new` or `repeated` against the
|
|
74
74
|
earlier rounds you were told about; on a first round every finding is
|
|
75
75
|
`new` by definition. The orchestrator uses this to detect the
|
|
76
|
-
review-round escalation budget's trigger.
|
|
76
|
+
review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
|
|
77
77
|
- GitHub Actions shell replay: for any diff that adds or changes a GitHub
|
|
78
78
|
Actions `run:` step, replay it yourself under the shell the step actually
|
|
79
79
|
runs: `bash --noprofile --norc -eo pipefail` when `shell: bash` is set on
|
|
@@ -145,7 +145,6 @@ Rules:
|
|
|
145
145
|
read its summary before opening full logs.
|
|
146
146
|
|
|
147
147
|
Return exactly this structure as your final output, nothing else:
|
|
148
|
-
|
|
149
148
|
```yaml
|
|
150
149
|
status: reviewed
|
|
151
150
|
role: reviewer
|
|
@@ -158,6 +157,7 @@ findings:
|
|
|
158
157
|
description: ""
|
|
159
158
|
suggested_fix: ""
|
|
160
159
|
recurrence: new | repeated
|
|
160
|
+
introduced_by_delta: yes | no | unknown
|
|
161
161
|
acceptance_recommendation: accept | accept_with_notes | fix_required | reject
|
|
162
162
|
missing_tests:
|
|
163
163
|
- ""
|
|
@@ -128,9 +128,9 @@ trivial change.
|
|
|
128
128
|
the exhausted tier path falls straight to the merge-hold), or an
|
|
129
129
|
operator merge-hold, and adds a row (task, choice, reason) to
|
|
130
130
|
`03-decisions.md`'s Review-round escalation table, then sets the
|
|
131
|
-
`review-round-escalation` marker to the most recent choice. A
|
|
132
|
-
|
|
133
|
-
|
|
131
|
+
`review-round-escalation` marker to the most recent choice. A negative round
|
|
132
|
+
has an `acceptance_recommendation` of `fix_required` or `reject`; a misfired
|
|
133
|
+
review is not a round. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the three is
|
|
134
134
|
picked is judgment; that one is picked and recorded is not. Escalating
|
|
135
135
|
never substitutes for a review round and comes in addition to the halt
|
|
136
136
|
rule's split-or-redesign response, not instead of it.
|
package/assets/skill/SKILL.md
CHANGED
|
@@ -312,7 +312,7 @@ directory and the subagents.
|
|
|
312
312
|
marks each finding's `recurrence` as `new` or `repeated` against the earlier
|
|
313
313
|
rounds it was told about, which is what lets the orchestrator detect the
|
|
314
314
|
Review-round escalation budget's trigger (see below) without re-deriving it
|
|
315
|
-
by hand. When the implementer's report replays a prior round's mutation
|
|
315
|
+
by hand. The reviewer classifies every finding with the `introduced_by_delta` field (`yes`, `no`, or `unknown`); it sets `no` only after naming a base build and replaying the same reproduction in `reproduction`, and transfers it through the ordinary gate (not bounded-round guidance, which considers only `yes`/`unknown`). When findings are transferred, record the classification parenthetically in the `Description` field as `(introduced_by_delta: yes|no|unknown)`. When the implementer's report replays a prior round's mutation
|
|
316
316
|
probe, the orchestrator's reviewer briefing names the replayed probes the
|
|
317
317
|
implementer reports as killed together with their mutant definition
|
|
318
318
|
(`file`, `anchor`, `before`, `after`) and `verified_applied_via` value,
|
|
@@ -556,7 +556,6 @@ background monitor is no substitute for those returns.
|
|
|
556
556
|
The output shape remains the same for either selected contract. Compare the
|
|
557
557
|
delegated versioned records and producer evidence under Contract selection
|
|
558
558
|
above; a recommendation does not replace orchestrator acceptance.
|
|
559
|
-
|
|
560
559
|
```yaml
|
|
561
560
|
status: reviewed
|
|
562
561
|
role: reviewer
|
|
@@ -569,6 +568,7 @@ findings:
|
|
|
569
568
|
description: ""
|
|
570
569
|
suggested_fix: ""
|
|
571
570
|
recurrence: new | repeated
|
|
571
|
+
introduced_by_delta: yes | no | unknown
|
|
572
572
|
acceptance_recommendation: accept | accept_with_notes | fix_required | reject
|
|
573
573
|
missing_tests:
|
|
574
574
|
- ""
|
|
@@ -584,7 +584,6 @@ withdrawn:
|
|
|
584
584
|
- description: ""
|
|
585
585
|
reason: ""
|
|
586
586
|
```
|
|
587
|
-
|
|
588
587
|
`acceptance_recommendation` is mandatory: every reviewer return must set it.
|
|
589
588
|
When it is missing, the orchestrator asks the reviewer to resupply it
|
|
590
589
|
instead of inferring one from the findings list.
|
|
@@ -594,6 +593,7 @@ task: `new` for a defect class not previously found here, `repeated` for
|
|
|
594
593
|
one that already appeared in an earlier round. On a task's first review
|
|
595
594
|
round every finding is `new` by definition. This is what feeds the
|
|
596
595
|
Review-round escalation budget's trigger.
|
|
596
|
+
`introduced_by_delta` records whether a finding is attributable to the reviewed delta: `no` requires a named base build and replay in `reproduction`, is transferred parenthetically in the `Description` field of `05-review-findings.md` without renaming `Severity`/`Decision`, and follows the ordinary gate; only `yes`/`unknown` participate in bounded-round rules.
|
|
597
597
|
|
|
598
598
|
`method_applied` echoes the `review_method` named in the briefing (see step
|
|
599
599
|
7); `withdrawn` lists each finding the reviewer proposed and then retracted
|
|
@@ -765,7 +765,7 @@ review and never satisfies the review gate, since review is never skipped.
|
|
|
765
765
|
The signal: a review round finds a new defect of the same class a previous
|
|
766
766
|
round's fix already addressed, so the class has recurred once after being
|
|
767
767
|
fixed, and the next fix would again be case-by-case enumeration (boundary
|
|
768
|
-
tokens, spellings, and similar one-off patches). Stop the first time this
|
|
768
|
+
tokens, spellings, and similar one-off patches). Apply this signal only to `introduced_by_delta: yes`/`unknown`; `no` continues through the ordinary finding gate. Stop the first time this
|
|
769
769
|
signal fires: the recurrence is already the class's second occurrence, so
|
|
770
770
|
do not wait for a third one before stopping. Name the structural cause in
|
|
771
771
|
one sentence, and decide to split or redesign rather than keep accreting
|
|
@@ -783,11 +783,9 @@ and across repeated review rounds, so effort does not keep accumulating
|
|
|
783
783
|
unaided: by the second round-2 halt signal on the same task, or by the
|
|
784
784
|
third `fix_required` review round on the same task, whichever comes
|
|
785
785
|
first, choose one of three escalations instead of running another round
|
|
786
|
-
the same way. A
|
|
787
|
-
`
|
|
788
|
-
|
|
789
|
-
chosen once the third such round has returned, before the next attempt
|
|
790
|
-
starts. The escalation is chosen in addition to the halt rule's
|
|
786
|
+
the same way. A negative round has an `acceptance_recommendation` of
|
|
787
|
+
`fix_required` or `reject`; a misfired review is not a round (see Subagent
|
|
788
|
+
misfire rule). A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. The escalation is chosen in addition to the halt rule's
|
|
791
789
|
split-or-redesign response, not instead of it.
|
|
792
790
|
|
|
793
791
|
- **Tier or model escalation**: raise the implementer to at least
|
|
@@ -13,7 +13,7 @@ critical waivers remain distinct. -->
|
|
|
13
13
|
|
|
14
14
|
## Review-round escalation
|
|
15
15
|
|
|
16
|
-
<!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third
|
|
16
|
+
<!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third negative round on that task. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. A run carries multiple tasks, so this table can carry multiple rows. Leave the single placeholder row as n/a when no task in this run has triggered the budget. -->
|
|
17
17
|
|
|
18
18
|
| Task | Choice | Reason |
|
|
19
19
|
|---|---|---|
|
|
@@ -16,7 +16,8 @@ by the grounding-mcp completeness reader yet).
|
|
|
16
16
|
| Severity | Category | Description | Suggested Fix | Decision |
|
|
17
17
|
|---|---|---|---|---|
|
|
18
18
|
| low/medium/high/critical | correctness/architecture/security/tests/maintainability/performance/docs | <!-- finding --> | <!-- fix --> | accepted/defer |
|
|
19
|
-
<!-- This row
|
|
19
|
+
<!-- This legacy five-cell table and placeholder row are the shipped template, not a finding: the orchestrator-workflow completeness reader matches the row byte-for-byte and fails the completeness gate closed when it survives untouched with no concrete finding row, the same way a `TODO` marker does. During findings transfer (step 7), replace this row with each reviewer finding and record its attribution parenthetically in the Description field as `(introduced_by_delta: yes|no|unknown)`. For a genuine zero-findings review, delete this row instead — a header row with no data rows is a valid, complete table; leaving this row next to real finding rows is also fine. This mirrors grounding-mcp's placeholder-row detection; keep the two in sync. -->
|
|
20
|
+
<!-- A `no` classification requires the named base build and replay recorded in the reviewer's `reproduction`; it follows the ordinary finding gate, while only `yes`/`unknown` feed bounded-round halt and escalation guidance. The load-bearing Severity and Decision headers remain unchanged. -->
|
|
20
21
|
|
|
21
22
|
## Missing Tests
|
|
22
23
|
|
|
@@ -35,4 +36,3 @@ accept | accept_with_notes | fix_required | reject
|
|
|
35
36
|
<!-- Reproduction note: when a finding rests on empirical or probabilistic evidence (flake rates, benchmarks, "n runs green", performance/timing numbers), record the reviewer's independent reproduction (method, sample size, result vs. the implementer's claim) in the reviewer output contract's `reproduction` field (SKILL.md step 7). Deterministic checks (a single test run, tsc, lint) do not require it. -->
|
|
36
37
|
|
|
37
38
|
<!-- Recurrence note: each finding in the reviewer output contract also carries a `recurrence` field (new or repeated), letting the orchestrator read the Review-round escalation budget's trigger (SKILL.md, Review-round escalation budget) off the reviewer's own return instead of reconstructing it by hand. A repeated finding here is what feeds that budget's round count. -->
|
|
38
|
-
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "orchestrator-workflow",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.34.0",
|
|
4
4
|
"description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
|
|
5
5
|
"main": "dist/index.js",
|
|
6
6
|
"type": "module",
|