orchestrator-workflow 0.33.0 → 0.34.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,16 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.34.0] - 2026-09-13
11
+
12
+ - Reviewer findings now carry `introduced_by_delta: yes | no | unknown`.
13
+ A `no` attribution requires a named base build and replay in `reproduction`;
14
+ it remains in the ordinary finding gate and Findings table, while only `yes`
15
+ and `unknown` participate in bounded halt and escalation rules. The legacy
16
+ five-cell table and placeholder row remain byte-compatible with the
17
+ completeness reader; concrete rows record attribution parenthetically in
18
+ their Description field.
19
+
10
20
  ## [0.33.0] - 2026-09-12
11
21
 
12
22
  ### Changed
@@ -73,7 +73,7 @@ Check, at minimum:
73
73
  review round, classify each finding as `new` or `repeated` against the
74
74
  earlier rounds you were told about; on a first round every finding is
75
75
  `new` by definition. The orchestrator uses this to detect the
76
- review-round escalation budget's trigger.
76
+ review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
77
77
  - GitHub Actions shell replay: for any diff that adds or changes a GitHub
78
78
  Actions `run:` step, replay it yourself under the shell the step actually
79
79
  runs: `bash --noprofile --norc -eo pipefail` when `shell: bash` is set on
@@ -145,7 +145,6 @@ Rules:
145
145
  read its summary before opening full logs.
146
146
 
147
147
  Return exactly this structure as your final output, nothing else:
148
-
149
148
  ```yaml
150
149
  status: reviewed
151
150
  role: reviewer
@@ -158,6 +157,7 @@ findings:
158
157
  description: ""
159
158
  suggested_fix: ""
160
159
  recurrence: new | repeated
160
+ introduced_by_delta: yes | no | unknown
161
161
  acceptance_recommendation: accept | accept_with_notes | fix_required | reject
162
162
  missing_tests:
163
163
  - ""
@@ -128,9 +128,9 @@ trivial change.
128
128
  the exhausted tier path falls straight to the merge-hold), or an
129
129
  operator merge-hold, and adds a row (task, choice, reason) to
130
130
  `03-decisions.md`'s Review-round escalation table, then sets the
131
- `review-round-escalation` marker to the most recent choice. A counted
132
- round is a completed reviewer return recommending `fix_required` or
133
- `reject`; a misfired review is not a round. Which of the three is
131
+ `review-round-escalation` marker to the most recent choice. A negative round
132
+ has an `acceptance_recommendation` of `fix_required` or `reject`; a misfired
133
+ review is not a round. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the three is
134
134
  picked is judgment; that one is picked and recorded is not. Escalating
135
135
  never substitutes for a review round and comes in addition to the halt
136
136
  rule's split-or-redesign response, not instead of it.
@@ -312,7 +312,7 @@ directory and the subagents.
312
312
  marks each finding's `recurrence` as `new` or `repeated` against the earlier
313
313
  rounds it was told about, which is what lets the orchestrator detect the
314
314
  Review-round escalation budget's trigger (see below) without re-deriving it
315
- by hand. When the implementer's report replays a prior round's mutation
315
+ by hand. The reviewer classifies every finding with the `introduced_by_delta` field (`yes`, `no`, or `unknown`); it sets `no` only after naming a base build and replaying the same reproduction in `reproduction`, and transfers it through the ordinary gate (not bounded-round guidance, which considers only `yes`/`unknown`). When findings are transferred, record the classification parenthetically in the `Description` field as `(introduced_by_delta: yes|no|unknown)`. When the implementer's report replays a prior round's mutation
316
316
  probe, the orchestrator's reviewer briefing names the replayed probes the
317
317
  implementer reports as killed together with their mutant definition
318
318
  (`file`, `anchor`, `before`, `after`) and `verified_applied_via` value,
@@ -556,7 +556,6 @@ background monitor is no substitute for those returns.
556
556
  The output shape remains the same for either selected contract. Compare the
557
557
  delegated versioned records and producer evidence under Contract selection
558
558
  above; a recommendation does not replace orchestrator acceptance.
559
-
560
559
  ```yaml
561
560
  status: reviewed
562
561
  role: reviewer
@@ -569,6 +568,7 @@ findings:
569
568
  description: ""
570
569
  suggested_fix: ""
571
570
  recurrence: new | repeated
571
+ introduced_by_delta: yes | no | unknown
572
572
  acceptance_recommendation: accept | accept_with_notes | fix_required | reject
573
573
  missing_tests:
574
574
  - ""
@@ -584,7 +584,6 @@ withdrawn:
584
584
  - description: ""
585
585
  reason: ""
586
586
  ```
587
-
588
587
  `acceptance_recommendation` is mandatory: every reviewer return must set it.
589
588
  When it is missing, the orchestrator asks the reviewer to resupply it
590
589
  instead of inferring one from the findings list.
@@ -594,6 +593,7 @@ task: `new` for a defect class not previously found here, `repeated` for
594
593
  one that already appeared in an earlier round. On a task's first review
595
594
  round every finding is `new` by definition. This is what feeds the
596
595
  Review-round escalation budget's trigger.
596
+ `introduced_by_delta` records whether a finding is attributable to the reviewed delta: `no` requires a named base build and replay in `reproduction`, is transferred parenthetically in the `Description` field of `05-review-findings.md` without renaming `Severity`/`Decision`, and follows the ordinary gate; only `yes`/`unknown` participate in bounded-round rules.
597
597
 
598
598
  `method_applied` echoes the `review_method` named in the briefing (see step
599
599
  7); `withdrawn` lists each finding the reviewer proposed and then retracted
@@ -765,7 +765,7 @@ review and never satisfies the review gate, since review is never skipped.
765
765
  The signal: a review round finds a new defect of the same class a previous
766
766
  round's fix already addressed, so the class has recurred once after being
767
767
  fixed, and the next fix would again be case-by-case enumeration (boundary
768
- tokens, spellings, and similar one-off patches). Stop the first time this
768
+ tokens, spellings, and similar one-off patches). Apply this signal only to `introduced_by_delta: yes`/`unknown`; `no` continues through the ordinary finding gate. Stop the first time this
769
769
  signal fires: the recurrence is already the class's second occurrence, so
770
770
  do not wait for a third one before stopping. Name the structural cause in
771
771
  one sentence, and decide to split or redesign rather than keep accreting
@@ -783,11 +783,9 @@ and across repeated review rounds, so effort does not keep accumulating
783
783
  unaided: by the second round-2 halt signal on the same task, or by the
784
784
  third `fix_required` review round on the same task, whichever comes
785
785
  first, choose one of three escalations instead of running another round
786
- the same way. A counted round is a completed reviewer return whose
787
- `acceptance_recommendation` is `fix_required` or `reject`; a misfired
788
- review is not a round (see Subagent misfire rule); the escalation is
789
- chosen once the third such round has returned, before the next attempt
790
- starts. The escalation is chosen in addition to the halt rule's
786
+ the same way. A negative round has an `acceptance_recommendation` of
787
+ `fix_required` or `reject`; a misfired review is not a round (see Subagent
788
+ misfire rule). A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. The escalation is chosen in addition to the halt rule's
791
789
  split-or-redesign response, not instead of it.
792
790
 
793
791
  - **Tier or model escalation**: raise the implementer to at least
@@ -13,7 +13,7 @@ critical waivers remain distinct. -->
13
13
 
14
14
  ## Review-round escalation
15
15
 
16
- <!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third fix_required review round on that task. A run carries multiple tasks, so this table can carry multiple rows. Leave the single placeholder row as n/a when no task in this run has triggered the budget. -->
16
+ <!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third negative round on that task. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. A run carries multiple tasks, so this table can carry multiple rows. Leave the single placeholder row as n/a when no task in this run has triggered the budget. -->
17
17
 
18
18
  | Task | Choice | Reason |
19
19
  |---|---|---|
@@ -16,7 +16,8 @@ by the grounding-mcp completeness reader yet).
16
16
  | Severity | Category | Description | Suggested Fix | Decision |
17
17
  |---|---|---|---|---|
18
18
  | low/medium/high/critical | correctness/architecture/security/tests/maintainability/performance/docs | <!-- finding --> | <!-- fix --> | accepted/defer |
19
- <!-- This row is the shipped template placeholder, not a finding: the orchestrator-workflow completeness reader fails the completeness gate closed when this exact row survives untouched and no concrete finding row has been added, the same way a `TODO` marker does. During findings transfer (step 7), replace this row with each reviewer finding. For a genuine zero-findings review, delete this row instead — a header row with no data rows is a valid, complete table; leaving this row next to real finding rows is also fine. This mirrors grounding-mcp's placeholder-row detection; keep the two in sync. -->
19
+ <!-- This legacy five-cell table and placeholder row are the shipped template, not a finding: the orchestrator-workflow completeness reader matches the row byte-for-byte and fails the completeness gate closed when it survives untouched with no concrete finding row, the same way a `TODO` marker does. During findings transfer (step 7), replace this row with each reviewer finding and record its attribution parenthetically in the Description field as `(introduced_by_delta: yes|no|unknown)`. For a genuine zero-findings review, delete this row instead — a header row with no data rows is a valid, complete table; leaving this row next to real finding rows is also fine. This mirrors grounding-mcp's placeholder-row detection; keep the two in sync. -->
20
+ <!-- A `no` classification requires the named base build and replay recorded in the reviewer's `reproduction`; it follows the ordinary finding gate, while only `yes`/`unknown` feed bounded-round halt and escalation guidance. The load-bearing Severity and Decision headers remain unchanged. -->
20
21
 
21
22
  ## Missing Tests
22
23
 
@@ -35,4 +36,3 @@ accept | accept_with_notes | fix_required | reject
35
36
  <!-- Reproduction note: when a finding rests on empirical or probabilistic evidence (flake rates, benchmarks, "n runs green", performance/timing numbers), record the reviewer's independent reproduction (method, sample size, result vs. the implementer's claim) in the reviewer output contract's `reproduction` field (SKILL.md step 7). Deterministic checks (a single test run, tsc, lint) do not require it. -->
36
37
 
37
38
  <!-- Recurrence note: each finding in the reviewer output contract also carries a `recurrence` field (new or repeated), letting the orchestrator read the Review-round escalation budget's trigger (SKILL.md, Review-round escalation budget) off the reviewer's own return instead of reconstructing it by hand. A repeated finding here is what feeds that budget's round count. -->
38
-
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.33.0",
3
+ "version": "0.34.0",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",