orchestrator-workflow 0.32.0 → 0.34.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,64 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.34.0] - 2026-09-13
11
+
12
+ - Reviewer findings now carry `introduced_by_delta: yes | no | unknown`.
13
+ A `no` attribution requires a named base build and replay in `reproduction`;
14
+ it remains in the ordinary finding gate and Findings table, while only `yes`
15
+ and `unknown` participate in bounded halt and escalation rules. The legacy
16
+ five-cell table and placeholder row remain byte-compatible with the
17
+ completeness reader; concrete rows record attribution parenthetically in
18
+ their Description field.
19
+
20
+ ## [0.33.0] - 2026-09-12
21
+
22
+ ### Changed
23
+
24
+ - Implementer reports now paste non-empty commit lists from `git log --reverse
25
+ --format=%H <base>..HEAD`, and foreground verification/probe plans and
26
+ repeat tallies return in the same turn as the last check; a background
27
+ monitor does not substitute. Anchored by pandora batch48 evidence.
28
+
29
+ - The implementer output contract now enumerates mutation-probe `result` as
30
+ `killed | survived | not_applicable` in both the installed prompt and the
31
+ SKILL.md reference. The docs-consistency guard separately pins each copy's
32
+ complete field block, including that enum, and now recognizes versioned
33
+ parenthesized headings in `see <name> below` forward pointers.
34
+
35
+ - The implementer and reviewer prompts (`assets/agents/implementer.md`,
36
+ `assets/agents/reviewer.md`, mirrored in SKILL.md) now say: cite a
37
+ coverage gate's threshold and pass/fail counts, not a run-specific
38
+ coverage percentage; cite a percentage only together with the exact
39
+ commit and the run count, since branch coverage varies between runs of
40
+ the same commit. Anchored by pandora run
41
+ `.ai/runs/2026-09-11-memory-sync-wipe`.
42
+ - The implementer `mutation_probes` output field (`assets/agents/
43
+ implementer.md`, mirrored in SKILL.md, and the `04-implementation-
44
+ summary.md` template's Mutation Probes table) now carries the mutant's
45
+ definition, not only its label: `file`, `anchor` (a line number or a
46
+ unique string), `before`, and `after` alongside the existing
47
+ `verified_applied_via`, `result`, `restored_verified`, and `replayed`
48
+ fields, so a later round can mechanically reapply the same edit instead
49
+ of only reading a prose description. Added `expectation: met | violated
50
+ | not_applicable` beside `result`, reporting whether the probe's
51
+ `result` matched its `--expect`, scoped to a measured `killed` or
52
+ `survived` `result` (a routine negative-control probe now reports
53
+ `result: survived, expectation: met`, which is not a regression) and
54
+ `not_applicable` otherwise (for example when the mutant could not be
55
+ applied and no `result` was measured). Added an eleventh sub-field,
56
+ `reason`: free text, required when `result` is `not_applicable`, empty
57
+ otherwise, carrying one of two canonical strings that distinguish a
58
+ non-regression from a regression: `no definition recorded` (a
59
+ prior-round probe recorded with only an id, no definition to reapply)
60
+ and `target text no longer present` (a replayed probe whose mutant can
61
+ no longer be applied). The fix-round replay rule now names a
62
+ probe to replay by its definition, not merely by its id, and treats a
63
+ replayed probe whose `expectation` is now `violated` (or which can no
64
+ longer be applied) as the regression signal, not `result` alone.
65
+ Anchored by pandora run `.ai/runs/2026-09-11-memory-sync-wipe` and the
66
+ agent-primitives `probe` result/expectation split (0.3.0).
67
+
10
68
  ## [0.32.0] - 2026-09-11
11
69
 
12
70
  ### Added
@@ -31,28 +31,56 @@ Rules:
31
31
  - Touch only the files relevant to the assigned task. Respect the
32
32
  allowed_changes and forbidden_changes lists in your task contract.
33
33
  - Add or update tests where appropriate. Run the tests you touched and report
34
- the result honestly; if you could not run them, say why.
34
+ the result honestly; if you could not run them, say why. Cite a coverage
35
+ gate's threshold and pass/fail counts, not a run-specific coverage
36
+ percentage; cite a percentage only together with the exact commit and the
37
+ run count, since branch coverage can vary between runs of the same commit.
35
38
  - When the task assignment names mutation probes to run, run each one and
36
- report it in the `mutation_probes` field of your output (mutant,
37
- verified_applied_via, result, restored_verified); an output missing that
38
- field when probes were named is treated as a misfire, not evidence. When
39
- the assignment names no mutation probes, return `mutation_probes: []`
40
- rather than omitting the field. Each item also carries `replayed`:
41
- `false` for a probe newly introduced this round.
39
+ report it in the `mutation_probes` field of your output (mutant, file,
40
+ anchor, before, after, verified_applied_via, result, expectation,
41
+ reason, restored_verified); an output missing that field when probes
42
+ were named is treated as a misfire, not evidence. `file` and `anchor`
43
+ (a line number or a unique surrounding string) locate the mutant;
44
+ `before` and `after` are the exact text swapped there, so a later round
45
+ can reapply the same edit without guessing instead of only a prose
46
+ description. `expectation` records whether `result` matched what the
47
+ probe was expected to do (`met`) or not (`violated`), independent of
48
+ `result` itself, only alongside a measured `killed` or `survived`
49
+ `result`; it is `not_applicable` otherwise (for example when the mutant
50
+ could not be applied and no `result` was measured). `reason` is free
51
+ text, required when `result` is `not_applicable`, empty otherwise,
52
+ carrying one of two canonical strings that distinguish a non-regression
53
+ from a regression: `no definition recorded` (a prior-round probe
54
+ recorded with only an id, no definition to reapply) and `target text no
55
+ longer present` (a replayed probe whose mutant can no longer be
56
+ applied). When the assignment names no mutation probes, return
57
+ `mutation_probes: []` rather than omitting the field.
58
+ Each item also carries `replayed`: `false` for a probe newly
59
+ introduced this round.
42
60
  - On any round after the task's first, the assignment also names every
43
61
  mutation probe named in an earlier round of this task (on the task's
44
62
  first round there are none), drawn from the run's
45
- `04-implementation-summary.md`. Replay each one, not only this round's
46
- new probes, before returning your report, and report each replayed
47
- probe in `mutation_probes` with the four evidence fields plus
48
- `replayed: true`. A replayed probe whose mutant now survives or can no
49
- longer be applied is a regression signal: report it as such (`result`
50
- `survived` or `not_applicable` with the reason) and resolve it before
51
- the next reviewer spawn.
63
+ `04-implementation-summary.md`, naming each by its mutant definition
64
+ (file, anchor, before, after), not merely by its id; a probe recorded
65
+ with only an id and no definition to reapply cannot be replayed and is
66
+ `not_applicable` (reason: `no definition recorded`), not a regression.
67
+ Replay each one, not only this round's new probes, before returning your
68
+ report, and report each replayed probe in `mutation_probes` with the
69
+ evidence fields plus `replayed: true`. A replayed probe whose
70
+ `expectation` is now `violated`, or which can no longer be applied
71
+ (reason: `target text no longer present`), is the regression signal;
72
+ `result` alone is not: report it as such (`result` `survived` or
73
+ `not_applicable` with the reason) and resolve it before the next
74
+ reviewer spawn.
52
75
  - When a verify runner is available, run it for the checks the acceptance
53
76
  criteria name and report its summary under `tests.executed`; when a
54
77
  mutation-probe runner is available, run the named probes through it and
55
- copy its fields into `mutation_probes`.
78
+ copy its fields into `mutation_probes`; when the runner reports a
79
+ probe's mutant record (`file`, `anchor`, `before`, `after`) separately
80
+ from its result fields (`verified_applied_via`, `result`, `expectation`,
81
+ `reason`, `restored_verified`), take the definition fields from that
82
+ mutant record so the copied report still carries all eleven
83
+ `mutation_probes` sub-fields.
56
84
  - Run every long test, build, or mutation-probe command in the foreground
57
85
  and wait for it to finish before returning. When one foreground call
58
86
  cannot hold it to completion, poll the backgrounded run to completion
@@ -80,6 +108,11 @@ Rules:
80
108
  when the task assignment asked for a commit is treated as a misfire, not
81
109
  evidence. When the task produced no commit, return `commits: []` rather
82
110
  than omitting the field.
111
+ - Populate a non-empty `commits` field by pasting `git log --reverse
112
+ --format=%H <base>..HEAD`; never type or hand-complete commit shas.
113
+ - Verification plans, probe plans, and repeat tallies run in the foreground,
114
+ and the implementer reports their returns in the same turn as the last
115
+ check. A background monitor is no substitute for those returns.
83
116
  - Only write a verification claim (for example "Verified by ...") in a code
84
117
  comment, commit message, or your report for a check you actually ran and
85
118
  measured yourself; never claim a run you did not execute.
@@ -131,8 +164,14 @@ tests:
131
164
  not_executed_reason: ""
132
165
  mutation_probes:
133
166
  - mutant: ""
167
+ file: ""
168
+ anchor: ""
169
+ before: ""
170
+ after: ""
134
171
  verified_applied_via: ""
135
- result: ""
172
+ result: killed | survived | not_applicable
173
+ expectation: met | violated | not_applicable
174
+ reason: ""
136
175
  restored_verified: ""
137
176
  replayed: false | true
138
177
  risks:
@@ -73,7 +73,7 @@ Check, at minimum:
73
73
  review round, classify each finding as `new` or `repeated` against the
74
74
  earlier rounds you were told about; on a first round every finding is
75
75
  `new` by definition. The orchestrator uses this to detect the
76
- review-round escalation budget's trigger.
76
+ review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
77
77
  - GitHub Actions shell replay: for any diff that adds or changes a GitHub
78
78
  Actions `run:` step, replay it yourself under the shell the step actually
79
79
  runs: `bash --noprofile --norc -eo pipefail` when `shell: bash` is set on
@@ -135,13 +135,16 @@ Rules:
135
135
  shell replay above is a second, explicitly non-probabilistic trigger for
136
136
  the same field: report it in `reproduction` too, with `sample_size:
137
137
  not_applicable` when the replay itself has no meaningful sample size.
138
+ - When citing a coverage gate, cite the threshold and pass/fail counts, not
139
+ a run-specific coverage percentage; cite a percentage only together with
140
+ the exact commit and the run count, since branch coverage can vary
141
+ between runs of the same commit.
138
142
  - When a mutation-probe runner is available in the session, run probes
139
143
  through it instead of editing files by hand, and carry its result fields
140
144
  into your findings and `reproduction`; when a verify runner is available,
141
145
  read its summary before opening full logs.
142
146
 
143
147
  Return exactly this structure as your final output, nothing else:
144
-
145
148
  ```yaml
146
149
  status: reviewed
147
150
  role: reviewer
@@ -154,6 +157,7 @@ findings:
154
157
  description: ""
155
158
  suggested_fix: ""
156
159
  recurrence: new | repeated
160
+ introduced_by_delta: yes | no | unknown
157
161
  acceptance_recommendation: accept | accept_with_notes | fix_required | reject
158
162
  missing_tests:
159
163
  - ""
@@ -128,9 +128,9 @@ trivial change.
128
128
  the exhausted tier path falls straight to the merge-hold), or an
129
129
  operator merge-hold, and adds a row (task, choice, reason) to
130
130
  `03-decisions.md`'s Review-round escalation table, then sets the
131
- `review-round-escalation` marker to the most recent choice. A counted
132
- round is a completed reviewer return recommending `fix_required` or
133
- `reject`; a misfired review is not a round. Which of the three is
131
+ `review-round-escalation` marker to the most recent choice. A negative round
132
+ has an `acceptance_recommendation` of `fix_required` or `reject`; a misfired
133
+ review is not a round. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the three is
134
134
  picked is judgment; that one is picked and recorded is not. Escalating
135
135
  never substitutes for a review round and comes in addition to the halt
136
136
  rule's split-or-redesign response, not instead of it.
@@ -211,20 +211,34 @@ directory and the subagents.
211
211
  for real, observe the named test fail, restore, re-verify). Hold the
212
212
  implementer's report to the claim-only-what-was-measured rule too: treat any
213
213
  verification claim there that is not backed by a check it actually ran as
214
- unverified. On any round after the task's first, the briefing also names
214
+ unverified. The installed `implementer.md` prompt has the implementer cite
215
+ a coverage gate's threshold and pass/fail counts, not a run-specific
216
+ coverage percentage, citing a percentage only together with the exact
217
+ commit and the run count, since branch coverage can vary between runs of
218
+ the same commit. On any round after the task's first, the briefing also names
215
219
  every mutation probe named in an earlier round of this task (on the
216
220
  task's first round there are none), drawn from the run's
217
- `04-implementation-summary.md`; the implementer replays each one, not
218
- only the round's new probes, before the next reviewer spawn, and
219
- reports each in `mutation_probes` with the four evidence fields plus
220
- `replayed: true`. A replayed probe whose mutant now survives or can no
221
- longer be applied is a regression signal, reported as such (`result`
222
- `survived` or `not_applicable` with the reason) and resolved before the
223
- next reviewer spawn. Record meaningful decisions in
221
+ `04-implementation-summary.md`, naming each by its mutant definition
222
+ (file, anchor, before, after), not merely by its id; a probe recorded
223
+ with only an id and no definition to reapply cannot be replayed and is
224
+ `not_applicable` (reason: `no definition recorded`), not a regression.
225
+ The implementer replays each one, not only the round's new probes,
226
+ before the next reviewer spawn, and reports each in `mutation_probes`
227
+ with the evidence fields plus `replayed: true`. A replayed probe whose
228
+ `expectation` is now `violated`, or which can no longer be applied
229
+ (reason: `target text no longer present`), is the regression signal;
230
+ `result` alone is not: reported as such (`result` `survived` or
231
+ `not_applicable` with the reason) and resolved before the next reviewer
232
+ spawn. Record meaningful decisions in
224
233
  `03-decisions.md` and consolidate evidence in
225
234
  `04-implementation-summary.md`, recording each probe the implementer
226
235
  reports as a row in `04-implementation-summary.md`'s Mutation Probes
227
- subsection, with the round it was named in. For any diff that adds or
236
+ subsection, with the round it was named in. Each row's Before/After
237
+ cells hold a single-line excerpt; when the mutant's actual before/after
238
+ text is multi-line or contains an unescaped `|`, or the mutant is a
239
+ patch/diff rather than a text swap, the full text or diff goes in the
240
+ implementer report or a fenced block placed directly under the table,
241
+ with the row noting where it lives. For any diff that adds or
228
242
  changes a GitHub Actions `run:` step, the installed `implementer.md`
229
243
  prompt requires replaying it locally under the shell the step actually
230
244
  runs, with the expected-success and the expected-failure inputs, before
@@ -258,54 +272,60 @@ directory and the subagents.
258
272
  operator flags high-risk; `normal` only for docs, renames, or batch
259
273
  cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
260
274
  never substitutes for it: do not pair `adversarial` with the `-medium`
261
- reviewer tier, a budget mismatch that names probes without the effort to
262
- run them; tiers themselves are unchanged by this axis. When the
263
- reviewer's environment cannot use version
264
- control to see the diff (for example a policy-gated repository), supply the
265
- diff as a pre-generated file in the briefing instead of expecting the
266
- reviewer to derive it, and have the reviewer report explicitly if it could
267
- only reconstruct the delta some other way, rather than silently reviewing
268
- less than the full change. The reviewer checks spec compliance, architecture
269
- consistency, edge cases, security, test adequacy (including whether new
270
- tests would fail if the change were reverted), and maintainability. Findings
271
- go to `05-review-findings.md`; transfer each finding from the reviewer
272
- output contract into the table's columns as-is, keeping the Severity and
273
- Decision headers unchanged, since those two are what the
274
- orchestrator-workflow completeness reader verifies. Replace the shipped
275
- placeholder/legend row with the transferred findings; for a genuine
276
- zero-findings review, delete that row instead of leaving it in place, since
277
- the completeness reader treats an untouched placeholder row with no finding
278
- rows as the template never having been filled in. When acceptance rests on
279
- empirical or probabilistic evidence (flake rates, benchmarks, "n runs
280
- green", performance/timing numbers), the reviewer must independently
281
- reproduce it — its own runs or measurements, not a re-read of the
282
- implementer's log — and record the method, sample size, and result against
283
- the implementer's claim in the reviewer output contract's `reproduction`
284
- field. This does not apply to deterministic checks (a single test run,
285
- `tsc`, lint): only claims that could vary run to run trigger it. The GitHub
286
- Actions shell replay named in step 6 is a second, explicitly
275
+ reviewer tier, a budget mismatch that names probes without the effort to run
276
+ them; tiers themselves are unchanged by this axis. When the reviewer's
277
+ environment cannot use version control to see the diff (for example a
278
+ policy-gated repository), supply the diff as a pre-generated file in the
279
+ briefing instead of expecting the reviewer to derive it, and have the
280
+ reviewer report explicitly if it could only reconstruct the delta some other
281
+ way, rather than silently reviewing less than the full change. The reviewer
282
+ checks spec compliance, architecture consistency, edge cases, security, test
283
+ adequacy (including whether new tests would fail if the change were
284
+ reverted), and maintainability. Findings go to `05-review-findings.md`;
285
+ transfer each finding from the reviewer output contract into the table's
286
+ columns as-is, keeping the Severity and Decision headers unchanged, since
287
+ those two are what the orchestrator-workflow completeness reader verifies.
288
+ Replace the shipped placeholder/legend row with the transferred findings;
289
+ for a genuine zero-findings review, delete that row instead of leaving it in
290
+ place, since the completeness reader treats an untouched placeholder row
291
+ with no finding rows as the template never having been filled in. When
292
+ acceptance rests on empirical or probabilistic evidence (flake rates,
293
+ benchmarks, "n runs green", performance/timing numbers), the reviewer must
294
+ independently reproduce it — its own runs or measurements, not a re-read of
295
+ the implementer's log — and record the method, sample size, and result
296
+ against the implementer's claim in the reviewer output contract's
297
+ `reproduction` field. This does not apply to deterministic checks (a single
298
+ test run, `tsc`, lint): only claims that could vary run to run trigger it.
299
+ The GitHub Actions shell replay named in step 6 is a second, explicitly
287
300
  non-probabilistic trigger for the same field, with `sample_size:
288
301
  not_applicable` allowed when the replay itself has no meaningful sample
289
- size. A change that deletes or renames an exported identifier, type, config
290
- key, or file is also checked for identifier drift (docs or comments still
291
- describing the old name as current), by the reviewer or by the orchestrator
292
- itself when it reviews a trivial rename per Scaling delegation, using a
293
- connected drift check when one exists. When this is not the task's first
294
- review round, name the round number in the briefing; the reviewer marks each
295
- finding's `recurrence` as `new` or `repeated` against the earlier rounds it
296
- was told about, which is what lets the orchestrator detect the Review-round
297
- escalation budget's trigger (see below) without re-deriving it by hand. When
298
- the implementer's report replays a prior round's mutation probe, the
299
- orchestrator's reviewer briefing names the replayed probes the implementer
300
- reports as killed together with their `mutant` and `verified_applied_via`
301
- values; the reviewer may then skip re-running those. The reviewer output
302
- contract itself is unchanged. Never run mutation probes in place against a
303
- worktree a reviewer subagent is concurrently reviewing; isolate the probe in
304
- a separate worktree or wait until the reviewer has returned before probing
305
- that tree again. For an explicitly adopted v1 run, ask the reviewer to
306
- compare the frozen delegated criteria with the referenced evidence and judge
307
- semantic adequacy, including whether a manual check is actually concrete and
308
- reasoned.
302
+ size. When citing a coverage gate, the installed `reviewer.md` prompt has
303
+ the reviewer cite the threshold and pass/fail counts, not a run-specific
304
+ coverage percentage, citing a percentage only together with the exact commit
305
+ and the run count, since branch coverage can vary between runs of the same
306
+ commit. A change that deletes or renames an exported identifier, type,
307
+ config key, or file is also checked for identifier drift (docs or comments
308
+ still describing the old name as current), by the reviewer or by the
309
+ orchestrator itself when it reviews a trivial rename per Scaling delegation,
310
+ using a connected drift check when one exists. When this is not the task's
311
+ first review round, name the round number in the briefing; the reviewer
312
+ marks each finding's `recurrence` as `new` or `repeated` against the earlier
313
+ rounds it was told about, which is what lets the orchestrator detect the
314
+ Review-round escalation budget's trigger (see below) without re-deriving it
315
+ by hand. The reviewer classifies every finding with the `introduced_by_delta` field (`yes`, `no`, or `unknown`); it sets `no` only after naming a base build and replaying the same reproduction in `reproduction`, and transfers it through the ordinary gate (not bounded-round guidance, which considers only `yes`/`unknown`). When findings are transferred, record the classification parenthetically in the `Description` field as `(introduced_by_delta: yes|no|unknown)`. When the implementer's report replays a prior round's mutation
316
+ probe, the orchestrator's reviewer briefing names the replayed probes the
317
+ implementer reports as killed together with their mutant definition
318
+ (`file`, `anchor`, `before`, `after`) and `verified_applied_via` value,
319
+ not merely their id; a probe recorded with only an id and no definition
320
+ cannot be skipped this way and is `not_applicable`. The reviewer may
321
+ then skip re-running the ones named by definition.
322
+ The reviewer output contract itself is unchanged. Never run mutation probes
323
+ in place against a worktree a reviewer subagent is concurrently reviewing;
324
+ isolate the probe in a separate worktree or wait until the reviewer has
325
+ returned before probing that tree again. For an explicitly adopted v1 run,
326
+ ask the reviewer to compare the frozen delegated criteria with the
327
+ referenced evidence and judge semantic adequacy, including whether a manual
328
+ check is actually concrete and reasoned.
309
329
  8. **Decide acceptance.** Accept, request fixes, defer, or escalate to the
310
330
  operator. High or critical findings block acceptance until fixed or
311
331
  explicitly waived: critical findings require operator sign-off; high
@@ -313,8 +333,8 @@ directory and the subagents.
313
333
  or critical finding counts as a waiver and follows the same rules. Record
314
334
  all decisions and waivers in `03-decisions.md` and summarize waivers in
315
335
  the Accepted Waivers section of `06-handoff.md`. A reviewer recommendation is not orchestrator acceptance and cannot authorize a critical waiver; only the operator may authorize a critical waiver. For newly created decision records, identify a stable ID, trigger/evidence, decision, accountable authority/source with concrete approval evidence, consequences, and a superseded decision ID when revising a prior decision. Link baseline revisions and waivers to those decision IDs. Established runs retain their recorded decision format; absent fields never create a retroactive blocker. Routine decisions within the delegated contract remain the orchestrator's responsibility; an out-of-scope change requires an operator decision. Markdown records evidence of real authority and never grant it by themselves. Do not accept while a
316
- required baseline criterion in an explicitly adopted v1 run has an open residual; a residual retains its ID
317
- and cannot be converted away. After independent review,
336
+ required baseline criterion in an explicitly adopted v1 run has an open
337
+ residual; a residual retains its ID and cannot be converted away. After independent review,
318
338
  the orchestrator may close a docs-only delta without another reviewer round only
319
339
  when the entire unreviewed delta contains only explanatory
320
340
  documentation, comments, or citations; contains no source- or test-file
@@ -454,8 +474,14 @@ tests:
454
474
  not_executed_reason: ""
455
475
  mutation_probes:
456
476
  - mutant: ""
477
+ file: ""
478
+ anchor: ""
479
+ before: ""
480
+ after: ""
457
481
  verified_applied_via: ""
458
- result: ""
482
+ result: killed | survived | not_applicable
483
+ expectation: met | violated | not_applicable
484
+ reason: ""
459
485
  restored_verified: ""
460
486
  replayed: false | true
461
487
  risks:
@@ -481,19 +507,37 @@ identifies the reviewed artifact and revision, reviewer, method, pass/fail
481
507
  standard, reasoned result, and baseline/criterion identities; it stays manual.
482
508
 
483
509
  When the task assignment names mutation probes to run, the implementer
484
- reports each one in the `mutation_probes` field (mutant,
485
- verified_applied_via, result, restored_verified); when the assignment
486
- names none, it returns `mutation_probes: []` rather than omitting the
487
- field, so 'none asked for' is distinguishable from 'asked for and not
488
- reported'. Each item also carries `replayed`: `false` for a probe newly
489
- introduced this round, `true` for a prior round's probe replayed this
490
- round under the replay rule in step 6. On any round after the task's
491
- first, the implementer replays every probe named in an earlier round of
492
- this task (on the task's first round there are none), not only this
493
- round's new probes, before the next reviewer spawn, reporting each one in
494
- `mutation_probes` alongside the round's new probes. A replayed probe
495
- whose mutant now survives or can no longer be applied is a regression
496
- signal, reported as such and resolved before the next reviewer spawn.
510
+ reports each one in the `mutation_probes` field (mutant, file, anchor,
511
+ before, after, verified_applied_via, result, expectation, reason,
512
+ restored_verified); `file` and `anchor` (a line number or a unique
513
+ surrounding string) locate the mutant, `before` and `after` are the
514
+ exact text swapped there, and `expectation` records whether `result`
515
+ matched what the probe was expected to do (`met`) or not (`violated`),
516
+ independent of `result` itself, only alongside a measured `killed` or
517
+ `survived` `result`; it is `not_applicable` otherwise (for example when
518
+ the mutant could not be applied and no `result` was measured). `reason`
519
+ is free text, required when `result` is `not_applicable`, empty
520
+ otherwise, carrying one of two canonical strings that distinguish a
521
+ non-regression from a regression: `no definition recorded` (a
522
+ prior-round probe recorded with only an id, no definition to reapply)
523
+ and `target text no longer present` (a replayed probe whose mutant can
524
+ no longer be applied). When the assignment names none, it returns
525
+ `mutation_probes: []` rather than
526
+ omitting the field, so 'none asked for' is distinguishable from 'asked
527
+ for and not reported'. Each item also carries `replayed`: `false` for a
528
+ probe newly introduced this round, `true` for a prior round's probe
529
+ replayed this round under the replay rule in step 6. On any round after
530
+ the task's first, the implementer replays every probe named in an
531
+ earlier round of this task (on the task's first round there are none),
532
+ naming each by its mutant definition, not merely by its id, not only
533
+ this round's new probes, before the next reviewer spawn, reporting each
534
+ one in `mutation_probes` alongside the round's new probes. A replayed
535
+ probe whose `expectation` is now `violated`, or which can no longer be
536
+ applied (reason: `target text no longer present`), is the regression
537
+ signal, reported as such and resolved before the next reviewer spawn;
538
+ `result` alone is not a regression signal, and a probe recorded with
539
+ only an id and no definition to reapply is `not_applicable` (reason:
540
+ `no definition recorded`).
497
541
 
498
542
  The `commits` field lists the full sha of every commit the implementer
499
543
  produced on the task branch, in the order produced; when the task
@@ -501,12 +545,17 @@ produced no commit, the implementer returns `commits: []` rather than
501
545
  omitting the field, so 'did not commit' is distinguishable from
502
546
  'forgot to report'.
503
547
 
548
+ For a non-empty `commits` field, the implementer pastes `git log
549
+ --reverse --format=%H <base>..HEAD`; it never types or hand-completes commit
550
+ shas. Verification plans, probe plans, and repeat tallies run in the foreground, and the
551
+ implementer reports their returns in the same turn as the last check. A
552
+ background monitor is no substitute for those returns.
553
+
504
554
  ## Reviewer output contract
505
555
 
506
556
  The output shape remains the same for either selected contract. Compare the
507
557
  delegated versioned records and producer evidence under Contract selection
508
558
  above; a recommendation does not replace orchestrator acceptance.
509
-
510
559
  ```yaml
511
560
  status: reviewed
512
561
  role: reviewer
@@ -519,6 +568,7 @@ findings:
519
568
  description: ""
520
569
  suggested_fix: ""
521
570
  recurrence: new | repeated
571
+ introduced_by_delta: yes | no | unknown
522
572
  acceptance_recommendation: accept | accept_with_notes | fix_required | reject
523
573
  missing_tests:
524
574
  - ""
@@ -534,7 +584,6 @@ withdrawn:
534
584
  - description: ""
535
585
  reason: ""
536
586
  ```
537
-
538
587
  `acceptance_recommendation` is mandatory: every reviewer return must set it.
539
588
  When it is missing, the orchestrator asks the reviewer to resupply it
540
589
  instead of inferring one from the findings list.
@@ -544,6 +593,7 @@ task: `new` for a defect class not previously found here, `repeated` for
544
593
  one that already appeared in an earlier round. On a task's first review
545
594
  round every finding is `new` by definition. This is what feeds the
546
595
  Review-round escalation budget's trigger.
596
+ `introduced_by_delta` records whether a finding is attributable to the reviewed delta: `no` requires a named base build and replay in `reproduction`, is transferred parenthetically in the `Description` field of `05-review-findings.md` without renaming `Severity`/`Decision`, and follows the ordinary gate; only `yes`/`unknown` participate in bounded-round rules.
547
597
 
548
598
  `method_applied` echoes the `review_method` named in the briefing (see step
549
599
  7); `withdrawn` lists each finding the reviewer proposed and then retracted
@@ -715,7 +765,7 @@ review and never satisfies the review gate, since review is never skipped.
715
765
  The signal: a review round finds a new defect of the same class a previous
716
766
  round's fix already addressed, so the class has recurred once after being
717
767
  fixed, and the next fix would again be case-by-case enumeration (boundary
718
- tokens, spellings, and similar one-off patches). Stop the first time this
768
+ tokens, spellings, and similar one-off patches). Apply this signal only to `introduced_by_delta: yes`/`unknown`; `no` continues through the ordinary finding gate. Stop the first time this
719
769
  signal fires: the recurrence is already the class's second occurrence, so
720
770
  do not wait for a third one before stopping. Name the structural cause in
721
771
  one sentence, and decide to split or redesign rather than keep accreting
@@ -733,11 +783,9 @@ and across repeated review rounds, so effort does not keep accumulating
733
783
  unaided: by the second round-2 halt signal on the same task, or by the
734
784
  third `fix_required` review round on the same task, whichever comes
735
785
  first, choose one of three escalations instead of running another round
736
- the same way. A counted round is a completed reviewer return whose
737
- `acceptance_recommendation` is `fix_required` or `reject`; a misfired
738
- review is not a round (see Subagent misfire rule); the escalation is
739
- chosen once the third such round has returned, before the next attempt
740
- starts. The escalation is chosen in addition to the halt rule's
786
+ the same way. A negative round has an `acceptance_recommendation` of
787
+ `fix_required` or `reject`; a misfired review is not a round (see Subagent
788
+ misfire rule). A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. The escalation is chosen in addition to the halt rule's
741
789
  split-or-redesign response, not instead of it.
742
790
 
743
791
  - **Tier or model escalation**: raise the implementer to at least
@@ -13,7 +13,7 @@ critical waivers remain distinct. -->
13
13
 
14
14
  ## Review-round escalation
15
15
 
16
- <!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third fix_required review round on that task. A run carries multiple tasks, so this table can carry multiple rows. Leave the single placeholder row as n/a when no task in this run has triggered the budget. -->
16
+ <!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third negative round on that task. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. A run carries multiple tasks, so this table can carry multiple rows. Leave the single placeholder row as n/a when no task in this run has triggered the budget. -->
17
17
 
18
18
  | Task | Choice | Reason |
19
19
  |---|---|---|
@@ -68,9 +68,15 @@ an optional row cannot stand in for a required criterion.
68
68
 
69
69
  ### Mutation Probes
70
70
 
71
- | Round | Mutant | Verified Applied Via | Result | Restored Verified | Replayed |
72
- |---|---|---|---|---|---|
73
- | <!-- round --> | <!-- mutant --> | <!-- verified_applied_via --> | <!-- result --> | <!-- restored_verified --> | <!-- replayed --> |
71
+ Before/After cells hold a single-line excerpt. When the mutant's actual
72
+ before/after text is multi-line or contains an unescaped `|`, or the mutant
73
+ is a patch/diff rather than a text swap, put the full text or diff in the
74
+ implementer report or a fenced block directly under the table, and note
75
+ where it lives in the row's own cell.
76
+
77
+ | Round | Mutant | File | Anchor | Before | After | Verified Applied Via | Result | Expectation | Reason | Restored Verified | Replayed |
78
+ |---|---|---|---|---|---|---|---|---|---|---|---|
79
+ | <!-- round --> | <!-- mutant --> | <!-- file --> | <!-- anchor --> | <!-- before --> | <!-- after --> | <!-- verified_applied_via --> | <!-- result --> | <!-- expectation --> | <!-- reason --> | <!-- restored_verified --> | <!-- replayed --> |
74
80
 
75
81
  ## Risks / Notes
76
82
 
@@ -16,7 +16,8 @@ by the grounding-mcp completeness reader yet).
16
16
  | Severity | Category | Description | Suggested Fix | Decision |
17
17
  |---|---|---|---|---|
18
18
  | low/medium/high/critical | correctness/architecture/security/tests/maintainability/performance/docs | <!-- finding --> | <!-- fix --> | accepted/defer |
19
- <!-- This row is the shipped template placeholder, not a finding: the orchestrator-workflow completeness reader fails the completeness gate closed when this exact row survives untouched and no concrete finding row has been added, the same way a `TODO` marker does. During findings transfer (step 7), replace this row with each reviewer finding. For a genuine zero-findings review, delete this row instead — a header row with no data rows is a valid, complete table; leaving this row next to real finding rows is also fine. This mirrors grounding-mcp's placeholder-row detection; keep the two in sync. -->
19
+ <!-- This legacy five-cell table and placeholder row are the shipped template, not a finding: the orchestrator-workflow completeness reader matches the row byte-for-byte and fails the completeness gate closed when it survives untouched with no concrete finding row, the same way a `TODO` marker does. During findings transfer (step 7), replace this row with each reviewer finding and record its attribution parenthetically in the Description field as `(introduced_by_delta: yes|no|unknown)`. For a genuine zero-findings review, delete this row instead — a header row with no data rows is a valid, complete table; leaving this row next to real finding rows is also fine. This mirrors grounding-mcp's placeholder-row detection; keep the two in sync. -->
20
+ <!-- A `no` classification requires the named base build and replay recorded in the reviewer's `reproduction`; it follows the ordinary finding gate, while only `yes`/`unknown` feed bounded-round halt and escalation guidance. The load-bearing Severity and Decision headers remain unchanged. -->
20
21
 
21
22
  ## Missing Tests
22
23
 
@@ -35,4 +36,3 @@ accept | accept_with_notes | fix_required | reject
35
36
  <!-- Reproduction note: when a finding rests on empirical or probabilistic evidence (flake rates, benchmarks, "n runs green", performance/timing numbers), record the reviewer's independent reproduction (method, sample size, result vs. the implementer's claim) in the reviewer output contract's `reproduction` field (SKILL.md step 7). Deterministic checks (a single test run, tsc, lint) do not require it. -->
36
37
 
37
38
  <!-- Recurrence note: each finding in the reviewer output contract also carries a `recurrence` field (new or repeated), letting the orchestrator read the Review-round escalation budget's trigger (SKILL.md, Review-round escalation budget) off the reviewer's own return instead of reconstructing it by hand. A repeated finding here is what feeds that budget's round count. -->
38
-
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.32.0",
3
+ "version": "0.34.0",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",