orchestrator-workflow 0.38.1 → 0.39.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,107 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.39.0] - 2026-09-23
11
+
12
+ - The docs-only review default now states its condition directly instead
13
+ of borrowing step 8's closure term: a review round whose entire delta
14
+ contains only explanatory documentation, comments, or citations, with no
15
+ source- or test-file edits and no semantic change to executable commands,
16
+ configuration, policy, instructions, or behavior, defaults to the
17
+ `-medium` reviewer tier with `review_method: normal`. The pinned-prose
18
+ cap counts a test-adequacy review round as one whose returned findings are
19
+ all `low` or `medium` `tests` findings about pin gaps; every normative
20
+ sentence a change adds or alters at its site is pinned, and one left
21
+ unpinned is named with the reason it is not load-bearing; a reviewer
22
+ respects a briefing that bounds the prose mutant space to a claim list
23
+ (#332).
24
+
25
+ - Clarified the review-round tier/model escalation path for `single` and
26
+ rewrapped the installed policy fence with its cited documentation.
27
+
28
+ - Mutation-probe verdict reporting now distinguishes runner-supplied,
29
+ result-only, and no-verdict cases without changing the output contract.
30
+ `result: killed` means the probe's test command reacted to the mutant
31
+ under the runner's pass predicate, or the test pass predicate declared
32
+ in the task assignment or probe plan when no runner supplies a verdict;
33
+ `survived` means it did not. `expectation: met` means the measured
34
+ result matches the expected result declared in the task assignment or
35
+ probe plan, and `violated` means it does not; both fields are
36
+ `not_applicable` when no result was measured. When a mutation-probe
37
+ runner is available, run the named probes through it and copy every
38
+ supplied `result` and `expectation` verbatim into `mutation_probes`,
39
+ never substituting your interpretation of its test output. Quote each
40
+ supplied verdict in `tests.executed`; when it supplies only `result`,
41
+ derive `expectation` from the expected result declared in the task
42
+ assignment or probe plan, and identify that declaration and derivation
43
+ there. When no machine-readable verdict is available, state that
44
+ explicitly in `tests.executed`, identify the declared test pass
45
+ predicate and expected result, and quote the observed baseline and
46
+ mutant outcomes. Derive `result` from those observations only when the
47
+ baseline passed, mutant application was verified, and the mutant test
48
+ completed under the same command and predicate; derive `expectation` by
49
+ comparing that result with the declared expected result, and label both
50
+ derivations as manual. Before transferring a probe row, compare each
51
+ copied field with the quoted verdict and each derived field with its
52
+ stated declaration and evidence. An explicit absence of a
53
+ machine-readable verdict requires the manual comparison, not resupply of
54
+ a nonexistent verdict. On a mismatch or missing required evidence,
55
+ obtain corrected evidence from the implementer or rerun the probe in
56
+ isolation, record the action in `03-decisions.md`, and keep the row
57
+ blocked from transfer until the comparison succeeds; if the evidence
58
+ cannot be obtained, record the unresolved proof rather than repeatedly
59
+ requesting an unavailable verdict. Never invent a verdict, override a
60
+ supplied field, or fill an unsupported derivation. Apply the same
61
+ evidence reporting and comparison to probes you run yourself before
62
+ recording their rows in `04-implementation-summary.md`. For probes you
63
+ run, apply the implementer's verdict-copy and manual-derivation rules to
64
+ your own measurements, reporting the quoted verdict or explicit verdict
65
+ absence and derivation evidence in `reproduction` and carrying the same
66
+ reported values into any associated finding. A quoted probe verdict is
67
+ not a named result of the verification set, so the set's
68
+ missing-or-extra rule does not apply to it. It reports per probe, in
69
+ `reproduction`, the probe, the replayed verdict or explicit verdict
70
+ absence with manual derivation evidence, and whether the measured
71
+ `result` and `expectation` match the recorded fields; a mismatch is a
72
+ finding of at least `high` and sets `matches_implementer_claim:
73
+ mismatched`. The implementer legend renders into every implementer tier
74
+ and Codex developer instructions; the revised test suite exercises all
75
+ three cases and the corresponding transfer decisions. Evidence for task
76
+ cb4ff78f-b0cb-402e-8387-f996f676f964 is recorded in the astra-old-eight
77
+ run, T-005 implementation report.
78
+
79
+ - A fix round now closes the defect class instead of the reported
80
+ instance. `assets/agents/implementer.md` (and every tier variant
81
+ rendered from it) states, for any round after a task's first, the
82
+ class-enumeration obligation (a search command plus its hit list in the
83
+ report, or a source-level closure with the reason) and a
84
+ one-mutation-probe-per-fixed-finding obligation. The implementer output
85
+ contract (`assets/skill/references/contracts.md`, mirrored in
86
+ `assets/agents/implementer.md`) gains a `class_closure` field (`kind:
87
+ enumerated | source | not_applicable` plus `command`, `sites` and
88
+ `closed: true | false`, with an unclosed site named in `risks` with the
89
+ reason); the subagent misfire rule now names the omission of
90
+ `class_closure` on any round after the task's first as a misfire, the
91
+ way it already names `mutation_probes` and `commits`.
92
+ `assets/agents/reviewer.md` gains the matching independent
93
+ class-enumeration obligation for round N+1 and now requires
94
+ `recurrence: repeated` whenever a finding's class matches an earlier
95
+ round's finding, even at a new site. `SKILL.md` step 8
96
+ (`references/evidence-and-probes.md`) now says the orchestrator halts
97
+ at the first `recurrence: repeated` finding whose `introduced_by_delta`
98
+ is `yes` or `unknown` and names split or redesign in `03-decisions.md`
99
+ before any further implementer spawn; the Round-2 halt rule paragraph in
100
+ `references/review-and-recovery.md` cross-references that halt at the
101
+ same scope, both sites pinned through one shared test constant.
102
+ `assets/templates/04-implementation-summary.md`
103
+ gains a Class Closure row per fix round (class, enumeration command,
104
+ sites, closure kind). Anchored by two observed batches: batch 48 saw 8
105
+ repeated findings of 32 review rounds, and batch 51 saw 13
106
+ repeated-finding mentions across 22 rounds, 7 of those rounds
107
+ attributable to case-level fixes or inert fixes rather than a closed
108
+ class. Anchored by an observed run: this fix-round contract was itself
109
+ applied in the fix rounds of the run that shipped it, before merge.
110
+
10
111
  ## [0.38.1] - 2026-09-20
11
112
 
12
113
  - Repository lint, not shipped in the package: rule 1 of
@@ -108,27 +209,27 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
108
209
  the check itself unchanged is a `03-decisions.md` entry, not a revision";
109
210
  the orchestrator records that entry, states in it why no evidence is
110
211
  invalidated, and communicates the corrected wording in the next
111
- delegation. Step 7 now says: "For a review round whose entire delta is a
112
- docs-only delta in the sense of step 8's docs-only closure, default to the
113
- `-medium` reviewer tier with `review_method: normal` where tier variants
114
- are installed". That refines the general tier default for this one class
115
- only, a round that touches an instruction, policy, template or prompt file
116
- keeps the general default, and the minimum review methods are untouched;
212
+ delegation. Step 7 now says: "For a review round whose entire delta contains only explanatory documentation, comments, or citations and contains no source- or test-file edits and no semantic change to executable commands, configuration, policy, instructions, or behavior, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed". That refines the general tier default for this one class only. "A review round that touches an instruction, policy, template or prompt file (for example a SKILL.md instruction) keeps the general default, whatever the file type, and the minimums named above are unaffected";
117
213
  the AGENTS.md section does not yet point to this refinement.
118
214
  `references/review-and-recovery.md` gains a "Pinned-prose changes" section
119
215
  for a change whose acceptance rests on tests that pin documentation
120
216
  wording: "A prose mutant survives exactly when its bytes sit in no
121
217
  assertion", so review rounds that hunt for the next unpinned sentence do
122
- not converge. The section asks for one normative site per rule, a claim
123
- list in the acceptance criterion as the pin obligation (every normative
124
- sentence the change adds or alters at that site is a claim, an omission is
125
- named with its reason), a reviewer briefing that bounds the prose mutant
126
- space to that list, copies bound to the normative site by one shared test
127
- constant, and: "Cap test-adequacy review rounds on the change at two." It
128
- defines the capped round, exempts semantic findings, and changes neither
129
- the Round-2 halt rule, the escalation budget, the Fix-regression decision
130
- point nor the review gate. Step 7 points to the section without restating
131
- it. Evidence (issue #300 and the change that added the Fix-regression
218
+ not converge. The section asks for one normative site per rule, a claim list in the acceptance criterion as the pin obligation. "Every normative sentence the change adds or alters at that site is pinned; one left unpinned is named in the criterion with the reason it is not load-bearing." A reviewer briefing
219
+ bounds the prose mutant space to that list, and "When the briefing
220
+ bounds the
221
+ prose mutant space to a claim list, respect that bound and put scope notes in
222
+ `residual_risks`, unless an unlisted sentence is shown to be load-bearing."
223
+ Copies are bound to the normative site by one shared test constant, and: "Cap
224
+ test-adequacy review rounds on the change at two." "A test-adequacy review
225
+ round is one whose returned findings are all `tests` findings of severity
226
+ `low` or `medium` about pin gaps on the pinned prose; a round returning any
227
+ other finding is an ordinary round outside the cap." "The cap changes neither
228
+ the Round-2 halt rule, the Review-round escalation budget nor the
229
+ Fix-regression decision point: a test-adequacy review round still counts as a
230
+ negative round where it is one." It exempts semantic findings and leaves the
231
+ review gate as it is. Step 7 points to the section without
232
+ restating it. Evidence (issue #300 and the change that added the Fix-regression
132
233
  decision point; one repository each, not a benchmark): the issue reports
133
234
  baseline revisions r1 to r3 for two wording precisions of a verification
134
235
  method, and a run in which the top reviewer tier was about half the day's
@@ -87,6 +87,23 @@ Rules:
87
87
  `result` alone is not: report it as such (`result` `survived` or
88
88
  `not_applicable` with the reason) and resolve it before the next
89
89
  reviewer spawn.
90
+ - On any round after the task's first, when the round fixes a review
91
+ finding, enumerate the defect's class before returning: run a search
92
+ command for the pattern the finding's fix addresses and list every hit
93
+ in the report, or state a source-level closure (the fix removes the
94
+ pattern at its one source, closing the whole class without a search)
95
+ and say why in `summary`. Report the result in the output contract's
96
+ `class_closure` field: `kind: enumerated | source | not_applicable`
97
+ (`not_applicable` only on the task's first round, when there is no
98
+ review finding yet to fix), the search `command` that produced the hit
99
+ list (empty when `kind` is not `enumerated`), and the `sites` list of
100
+ every hit found (empty when `kind` is not `enumerated`), and `closed:
101
+ true` when every site the round found (by search or by source-level
102
+ closure) is fixed this round, `false` when a found site is not; name an unclosed site in `risks` with
103
+ the reason. On the task's first round `kind` is `not_applicable`, `command` and `sites` are empty,
104
+ and `closed` is `true`, since no site was found to leave open. Run one mutation probe per review
105
+ finding fixed in the round, in addition to any probe the assignment names, and report each one in
106
+ `mutation_probes`. A fix-round return without `class_closure` is a misfire per the misfire rule.
90
107
  - A persisted probe-plan reference may stand in for a repeated inline mutant
91
108
  definition when it resolves to a path plus immutable revision or hash and the
92
109
  mutant locator/index. Resolve it before running; a missing, stale, or
@@ -96,14 +113,33 @@ Rules:
96
113
  Never rewrite a prior plan for new code; record intentional supersession and
97
114
  rationale in run state before using a replacement.
98
115
  - When a verify runner is available, run it for the checks the acceptance
99
- criteria name and report its summary under `tests.executed`; when a
100
- mutation-probe runner is available, run the named probes through it and
101
- copy its fields into `mutation_probes`; when the runner reports a
102
- probe's mutant record (`file`, `anchor`, `before`, `after`) separately
116
+ criteria name and report its summary under `tests.executed`. When a
117
+ mutation-probe runner is available, run the named probes through it and copy
118
+ every supplied `result` and `expectation` verbatim into `mutation_probes`,
119
+ never substituting your interpretation of its test output. When the
120
+ runner reports a probe's mutant record (`file`, `anchor`, `before`,
121
+ `after`) separately
103
122
  from its result fields (`verified_applied_via`, `result`, `expectation`,
104
123
  `reason`, `restored_verified`), take the definition fields from that
105
124
  mutant record so the copied report still carries all eleven
106
- `mutation_probes` sub-fields. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
125
+ `mutation_probes` sub-fields. `result: killed` means the probe's test
126
+ command reacted to the mutant under the runner's pass predicate, or the
127
+ test pass predicate declared in the task assignment or probe plan when
128
+ no runner supplies a verdict; `survived` means it did not. `expectation:
129
+ met` means the measured result matches the expected result declared in
130
+ the task assignment or probe plan, and `violated` means it does not;
131
+ both fields are `not_applicable` when no result was measured. Quote each
132
+ supplied verdict in `tests.executed`; when it supplies only `result`,
133
+ derive `expectation` from the expected result declared in the task
134
+ assignment or probe plan, and identify that declaration and derivation
135
+ there. When no machine-readable verdict is available, state that
136
+ explicitly in `tests.executed`, identify the declared test pass
137
+ predicate and expected result, and quote the observed baseline and
138
+ mutant outcomes. Derive `result` from those observations only when the
139
+ baseline passed, mutant application was verified, and the mutant test
140
+ completed under the same command and predicate; derive `expectation` by
141
+ comparing that result with the declared expected result, and label both
142
+ derivations as manual.
107
143
  - Run every long test, build, or mutation-probe command in the foreground
108
144
  and wait for it to finish before returning. When one foreground call
109
145
  cannot hold it to completion, poll the backgrounded run to completion
@@ -210,6 +246,12 @@ mutation_probes:
210
246
  reason: ""
211
247
  restored_verified: ""
212
248
  replayed: false | true
249
+ class_closure:
250
+ kind: enumerated | source | not_applicable
251
+ command: ""
252
+ sites:
253
+ - ""
254
+ closed: true | false
213
255
  risks:
214
256
  - severity: low | medium | high
215
257
  description: ""
@@ -77,7 +77,7 @@ Check, at minimum:
77
77
  calibrated to sit inside the output's own run-to-run noise (timing digits,
78
78
  temporary-directory names) is not a regression test; the fix is to pin
79
79
  the argument under test in-process, or assert the actual contract (a
80
- bound, or the presence of a warning), never a byte ceiling.
80
+ bound, or the presence of a warning), never a byte ceiling. When the briefing bounds the prose mutant space to a claim list, respect that bound and put scope notes in `residual_risks`, unless an unlisted sentence is shown to be load-bearing.
81
81
  - Maintainability: naming, dead code, needless abstraction, doc drift.
82
82
  - Placement: does the change add org-, machine-, or point-in-time-bound
83
83
  evidence (dates, sample sizes, task ids, home paths, incident tallies) to a
@@ -87,8 +87,21 @@ Check, at minimum:
87
87
  - Recurrence: when the briefing tells you this is not the task's first
88
88
  review round, classify each finding as `new` or `repeated` against the
89
89
  earlier rounds you were told about; on a first round every finding is
90
- `new` by definition. The orchestrator uses this to detect the
91
- review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
90
+ `new` by definition. Class match, not site match, decides this: set
91
+ `repeated` whenever a finding's defect class matches an earlier round's
92
+ finding, even when this instance sits at a site the earlier round never
93
+ touched. On any round after the task's first, run your own
94
+ class-enumeration search independent of the implementer's
95
+ `class_closure` report: search for the pattern the earlier finding's fix
96
+ addressed, and compare what your own search returns against the
97
+ implementer's `class_closure.sites` list (or its `source` closure
98
+ reason); a site your search finds that the implementer's report omits
99
+ is itself a finding, classified under the ordinary severity gate by
100
+ the underlying defect's own severity, not by a fixed floor. In run
101
+ mode `single` there is no separate implementer `class_closure` report
102
+ to compare against; compare your search's hits against the Class Closure row of
103
+ `04-implementation-summary.md` plus its Risks / Notes section instead. The orchestrator
104
+ uses this to detect the review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
92
105
  - GitHub Actions shell replay: for any diff that adds or changes a GitHub
93
106
  Actions `run:` step, replay it yourself under the shell the step actually
94
107
  runs: `bash --noprofile --norc -eo pipefail` when `shell: bash` is set on
@@ -155,15 +168,35 @@ Rules:
155
168
  the exact commit and the run count, since branch coverage can vary
156
169
  between runs of the same commit.
157
170
  - When a mutation-probe runner is available in the session, run probes
158
- through it instead of editing files by hand, and carry its result fields
159
- into your findings and `reproduction`; when a verify runner is available,
160
- read its summary before opening full logs.
171
+ through it instead of editing files by hand. For probes you run, apply the
172
+ implementer's verdict-copy and manual-derivation rules to your own
173
+ measurements, reporting the quoted verdict or explicit verdict absence and
174
+ derivation evidence in `reproduction` and carrying the same reported values
175
+ into any associated finding. When a verify runner is available, read its
176
+ summary before opening full logs.
161
177
  - A reviewer briefing may identify a replayed probe through a resolved
162
178
  immutable probe-plan reference (path plus revision/hash and mutant
163
179
  locator/index) rather than repeat its inline definition. Verify the plan and
164
180
  result bind the checked state, cwd, attempt, expectation, application, and
165
181
  restoration; a plan alone, stale reference, or unresolved reference is not
166
- evidence. Legacy inline probe reports remain valid. When the briefing names run mode `single`, the orchestrator implemented the change itself and nobody has cross-checked its probe evidence: replay every named orchestrator probe, where named means the briefing gives its full definition or a resolved immutable plan-and-result reference (in a scratch copy or an isolating probe runner, never in the reviewed tree), and state in `reproduction`, per probe, the replayed verdict and whether it matches the recorded `result` and `expectation`. Do not skip a named probe in that mode, under any `review_method`; any mismatch also sets `matches_implementer_claim: mismatched`. A mismatch is a finding of at least `high`; a probe given only by id is `not_applicable` and is missing evidence, not a pass, and so is a briefing in that mode that names no probe. Without that mode line in the briefing this obligation does not exist.
182
+ evidence. Legacy inline probe reports remain valid.
183
+ When the briefing names run mode `single`, the orchestrator implemented
184
+ the change itself
185
+ and nobody has cross-checked its probe evidence: replay every named
186
+ orchestrator probe, where named means the briefing gives its full
187
+ definition or a resolved immutable plan-and-result reference (in a
188
+ scratch copy or an isolating probe runner, never in the reviewed tree).
189
+ It reports per probe, in `reproduction`, the probe, the replayed verdict
190
+ or explicit verdict absence with manual derivation evidence, and whether
191
+ the measured `result` and `expectation` match the recorded fields; a
192
+ mismatch is a finding of at least `high` and sets
193
+ `matches_implementer_claim: mismatched`. Do not skip a named probe in
194
+ that mode, under any `review_method`; any mismatch also sets
195
+ `matches_implementer_claim: mismatched`. A mismatch is a finding of at
196
+ least `high`; a probe given only by id is `not_applicable` and is
197
+ missing evidence, not a pass, and so is a briefing in that mode that
198
+ names no probe. Without that mode line in the briefing this obligation
199
+ does not exist.
167
200
 
168
201
  Return exactly this structure as your final output, nothing else:
169
202
  ```yaml
@@ -1,12 +1,14 @@
1
1
  <!-- orchestrator-workflow:begin -->
2
2
  ## Agentic Coding Workflow
3
3
 
4
- This repository uses an orchestrator-led agent workflow, installed and updated by
4
+ This repository uses an orchestrator-led agent workflow, installed and
5
+ updated by
5
6
  [orchestrator-workflow](https://github.com/LanNguyenSi/agent-dx/tree/master/packages/orchestrator-workflow).
6
7
 
7
8
  The primary agent acts as the orchestrator. It owns the goal, planning, task
8
- validation, delegation, final acceptance, and the operator handoff. Non-trivial
9
- review is delegated to a narrow subagent; which agent implements non-trivial work follows from the run mode (Core rules). The full procedure
9
+ validation, delegation, final acceptance, and the operator handoff.
10
+ Non-trivial review is delegated to a narrow subagent; which agent implements
11
+ non-trivial work follows from the run mode (Core rules). The full procedure
10
12
  and the subagent I/O contracts live in the `orchestrator-workflow` skill.
11
13
 
12
14
  ### Core rules
@@ -21,10 +23,14 @@ and the subagent I/O contracts live in the `orchestrator-workflow` skill.
21
23
  inline with the same read-only discipline instead.
22
24
  - The orchestrator plans features itself. It may delegate task slicing, but it
23
25
  validates the sliced tasks before implementation starts.
24
- - Non-trivial implementation follows the run mode recorded in `00-goal.md`. `delegated`, the default, sends it to narrow implementer subagents, one task
25
- per subagent; in `single` the orchestrator implements one coherent workstream itself; `batch` runs implementers in parallel worktrees. The skill's Run mode section defines the modes and how to choose one.
26
+ - Non-trivial implementation follows the run mode recorded in `00-goal.md`.
27
+ `delegated`, the default, sends it to narrow implementer subagents, one
28
+ task per subagent; in `single` the orchestrator implements one coherent
29
+ workstream itself; `batch` runs implementers in parallel worktrees. The
30
+ skill's Run mode section defines the modes and how to choose one.
26
31
  - Non-trivial review goes to a separate reviewer subagent (see Scaling
27
- delegation). Review itself is never skipped, in any run mode, not even for docs or bulk
32
+ delegation). Review itself is never skipped, in any run mode, not even for
33
+ docs or bulk
28
34
  changes.
29
35
  - Final acceptance and the final answer to the operator stay with the
30
36
  orchestrator.
@@ -41,7 +47,8 @@ default, not a ritual.
41
47
  solution; skip it when the change is well understood. Under a `minimal`
42
48
  profile there is no explorer subagent to spawn; run this step inline
43
49
  instead.
44
- - Slicing and, in run modes `delegated` and `batch`, implementer subagents are for non-trivial work: multiple files,
50
+ - Slicing and, in run modes `delegated` and `batch`, implementer subagents are
51
+ for non-trivial work: multiple files,
45
52
  real logic, or anything that benefits from decomposition or a fresh context.
46
53
  Under a `minimal` profile there is no task-slicer subagent; the orchestrator
47
54
  slices inline with the same contract.
@@ -55,7 +62,9 @@ default, not a ritual.
55
62
  scripts, hand-edited lockfiles, cross-major overrides, or anything the
56
63
  operator flags high-risk; `normal` fits only docs, renames, or batch
57
64
  cosmetics; `rigorous` is the default otherwise. Never pair `adversarial`
58
- with the `-medium` reviewer tier; tiers themselves are unchanged. A docs-only delta has its own review default; the skill's Delegate review step states it.
65
+ with the `-medium` reviewer tier; tiers themselves are unchanged. A
66
+ docs-only delta has its own review default; the skill's Delegate review step
67
+ states it.
59
68
  - When tier variants are installed (manifest `tiers: true`), the orchestrator
60
69
  picks the effort tier per task by complexity and risk, at its own judgment.
61
70
  The unsuffixed default subagent is the normal case; `-high`/`-xhigh` fit
@@ -111,7 +120,8 @@ trivial change.
111
120
  - Medium and low findings are addressed or consciously accepted at the
112
121
  orchestrator's judgment.
113
122
  - After independent review, the orchestrator may close a docs-only delta
114
- without another reviewer round only when its entire unreviewed delta is explanatory
123
+ without another reviewer round only when its entire unreviewed delta is
124
+ explanatory
115
125
  documentation, comments, or citations; has no source- or test-file edits or
116
126
  semantic changes to executable commands, configuration, policy,
117
127
  instructions, or behavior; and closes only low/medium documentation or
@@ -128,9 +138,13 @@ trivial change.
128
138
  the exhausted tier path falls straight to the merge-hold), or an
129
139
  operator merge-hold, and adds a row (task, choice, reason) to
130
140
  `03-decisions.md`'s Review-round escalation table, then sets the
131
- `review-round-escalation` marker to the most recent choice. A negative round
141
+ `review-round-escalation` marker to the most recent choice.
142
+ In `single`, tier/model escalation requires a recorded switch to `delegated`.
143
+ A negative round
132
144
  has an `acceptance_recommendation` of `fix_required` or `reject`; a misfired
133
- review is not a round. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the three is
145
+ review is not a round. A negative round counts only with at least one
146
+ introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the
147
+ three is
134
148
  picked is judgment; that one is picked and recorded is not. Escalating
135
149
  never substitutes for a review round and comes in addition to the halt
136
150
  rule's split-or-redesign response, not instead of it.
@@ -175,7 +189,8 @@ Workflow state lives under `.ai/`:
175
189
  routing selections.
176
190
  - Every worktree a run touches carries a `.ai/run` pointer (absolute path of
177
191
  the run directory, gitignored) and `00-goal.md` carries one
178
- `run-base[<repo-basename>]` marker per repository for multi-repo runs, next to the run's `mode` marker.
192
+ `run-base[<repo-basename>]` marker per repository for multi-repo runs, next
193
+ to the run's `mode` marker.
179
194
 
180
195
  ### Models
181
196
 
@@ -113,6 +113,12 @@ mutation_probes:
113
113
  reason: ""
114
114
  restored_verified: ""
115
115
  replayed: false | true
116
+ class_closure:
117
+ kind: enumerated | source | not_applicable
118
+ command: ""
119
+ sites:
120
+ - ""
121
+ closed: true | false
116
122
  risks:
117
123
  - severity: low | medium | high
118
124
  description: ""
@@ -125,8 +131,27 @@ commits:
125
131
 
126
132
  Follow [evidence-and-probes.md workflow step 6](evidence-and-probes.md#workflow)
127
133
  for implementation evidence, verification, mutation probes, and replay. For
128
- output-field semantics and commit reporting, follow the installed
129
- implementer role prompt. Return the selected contract's YAML envelope. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
134
+ commit reporting, follow the installed implementer role prompt. Return the
135
+ selected contract's YAML envelope. `result: killed` means the probe's test
136
+ command reacted to the mutant under the runner's pass predicate, or the
137
+ test pass predicate declared in the task assignment or probe plan when no
138
+ runner supplies a verdict; `survived` means it did not. `expectation: met`
139
+ means the measured result matches the expected result declared in the task
140
+ assignment or probe plan, and `violated` means it does not; both fields
141
+ are `not_applicable` when no result was measured. When a mutation-probe
142
+ runner is available, run the named probes through it and copy every
143
+ supplied `result` and `expectation` verbatim into `mutation_probes`, never
144
+ substituting your interpretation of its test output. Quote each supplied
145
+ verdict in `tests.executed`; when it supplies only `result`, derive
146
+ `expectation` from the expected result declared in the task assignment or
147
+ probe plan, and identify that declaration and derivation there. When no
148
+ machine-readable verdict is available, state that explicitly in
149
+ `tests.executed`, identify the declared test pass predicate and expected
150
+ result, and quote the observed baseline and mutant outcomes. Derive
151
+ `result` from those observations only when the baseline passed, mutant
152
+ application was verified, and the mutant test completed under the same
153
+ command and predicate; derive `expectation` by comparing that result with
154
+ the declared expected result, and label both derivations as manual.
130
155
 
131
156
  ## Reviewer output contract
132
157
 
@@ -112,7 +112,21 @@ directory and the subagents.
112
112
  `03-decisions.md` and consolidate evidence in
113
113
  `04-implementation-summary.md`, recording each probe the implementer
114
114
  reports as a row in `04-implementation-summary.md`'s Mutation Probes
115
- subsection, with the round it was named in. Before transferring a probe row, compare its `result` and `expectation` with the runner verdict quoted in `tests.executed`; on a mismatch, or when a verdict the runner states is not quoted, resupply it (ask the same implementer for the verdict, respawn one when it is gone, or rerun the probe yourself in isolation), record the resupply in `03-decisions.md`, and treat it as a transfer blocker rather than a misfire, since the return itself parses; never infer either field. A quoted probe verdict is not a named result of the verification set, so the set's missing-or-extra rule does not apply to it. Each row's Before/After
115
+ subsection, with the round it was named in. Before transferring a probe
116
+ row, compare each copied field with the quoted verdict and each derived
117
+ field with its stated declaration and evidence. An explicit absence of
118
+ a machine-readable verdict requires the manual comparison, not resupply
119
+ of a nonexistent verdict. On a mismatch or missing required evidence,
120
+ obtain corrected evidence from the implementer or rerun the probe in
121
+ isolation, record the action in `03-decisions.md`, and keep the row
122
+ blocked from transfer until the comparison succeeds; if the evidence
123
+ cannot be obtained, record the unresolved proof rather than repeatedly
124
+ requesting an unavailable verdict. Never invent a verdict, override a
125
+ supplied field, or fill an unsupported derivation. Apply the same
126
+ evidence reporting and comparison to probes you run yourself before
127
+ recording their rows in `04-implementation-summary.md`. A quoted probe
128
+ verdict is not a named result of the verification set, so the set's
129
+ missing-or-extra rule does not apply to it. Each row's Before/After
116
130
  cells hold a single-line excerpt; when the mutant's actual before/after
117
131
  text is multi-line or contains an unescaped `|`, or the mutant is a
118
132
  patch/diff rather than a text swap, the full text or diff goes in the
@@ -152,7 +166,7 @@ directory and the subagents.
152
166
  cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
153
167
  never substitutes for it: do not pair `adversarial` with the `-medium`
154
168
  reviewer tier, a budget mismatch that names probes without the effort to run
155
- them; tiers themselves are unchanged by this axis. For a review round whose entire delta is a docs-only delta in the sense of step 8's docs-only closure, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed. This refines the general tier default above for that one class only: there `-medium` is the default and a higher tier is the non-default choice recorded with a one-line reason. A review round that touches an instruction, policy, template or prompt file keeps the general default, whatever the file type, and the minimums named above are unaffected. For a change whose acceptance rests on tests that pin documentation wording, write the briefing as the Pinned-prose changes section of [review and recovery](review-and-recovery.md) requires. When the reviewer's
169
+ them; tiers themselves are unchanged by this axis. For a review round whose entire delta contains only explanatory documentation, comments, or citations and contains no source- or test-file edits and no semantic change to executable commands, configuration, policy, instructions, or behavior, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed. This refines the general tier default above for that one class only: there `-medium` is the default and a higher tier is the non-default choice recorded with a one-line reason. A review round that touches an instruction, policy, template or prompt file (for example a SKILL.md instruction) keeps the general default, whatever the file type, and the minimums named above are unaffected. For a change whose acceptance rests on tests that pin documentation wording, write the briefing as the Pinned-prose changes section of [review and recovery](review-and-recovery.md) requires. When the reviewer's
156
170
  environment cannot use version control to see the diff (for example a
157
171
  policy-gated repository), supply the diff as a pre-generated file in the
158
172
  briefing instead of expecting the reviewer to derive it, and have the
@@ -198,8 +212,25 @@ directory and the subagents.
198
212
  not merely their id; a probe recorded with only an id and no definition
199
213
  cannot be skipped this way and is `not_applicable`. The reviewer may
200
214
  then skip re-running the ones named by definition.
201
- The reviewer output contract itself is unchanged. In run mode `single` that skip permission does not apply: nobody but the orchestrator has seen its probe evidence. A probe counts as named when the briefing gives its full definition or a resolved immutable plan-and-result reference; an id alone does not name a probe. The orchestrator records every probe it ran in one of those two forms in `04-implementation-summary.md` before requesting review, the reviewer briefing states the run mode and names each of those probes, and the reviewer must replay every named orchestrator probe, through the probe runner when one is available and never in the reviewed tree. It reports per probe, in `reproduction`, the probe, the replayed runner verdict, and whether that verdict matches the recorded `result` and `expectation`; a mismatch is a finding of at least `high` and sets `matches_implementer_claim: mismatched`. A probe given only by id is `not_applicable` and counts as missing evidence, not as a pass, and so does a `single` briefing that names no probe at all. Never run mutation probes
202
- in place against a worktree a reviewer subagent is concurrently reviewing;
215
+ The reviewer output contract itself is unchanged. In run mode `single`
216
+ that skip permission does not apply: nobody but the orchestrator has
217
+ seen its probe evidence. A probe counts as named when the briefing
218
+ gives its full definition or a resolved immutable plan-and-result
219
+ reference; an id alone does not name a probe. The orchestrator records
220
+ every probe it ran in one of those two forms in
221
+ `04-implementation-summary.md` before requesting review, the reviewer
222
+ briefing states the run mode and names each of those probes, and the
223
+ reviewer must replay every named orchestrator probe, through the probe
224
+ runner when one is available and never in the reviewed tree. It reports
225
+ per probe, in `reproduction`, the probe, the replayed verdict or
226
+ explicit verdict absence with manual derivation evidence, and whether
227
+ the measured `result` and `expectation` match the recorded fields; a
228
+ mismatch is a finding of at least `high` and sets
229
+ `matches_implementer_claim: mismatched`. A probe given only by id is
230
+ `not_applicable` and counts as missing evidence, not as a pass, and so
231
+ does a `single` briefing that names no probe at all. Never run mutation
232
+ probes in place against a worktree a reviewer subagent is concurrently
233
+ reviewing;
203
234
  isolate the probe in a separate worktree or wait until the reviewer has
204
235
  returned before probing that tree again. For an explicitly adopted v1 run,
205
236
  ask the reviewer to compare the frozen delegated criteria with the
@@ -222,9 +253,11 @@ directory and the subagents.
222
253
  documentation or maintainability findings. This option never closes a
223
254
  high/critical or other ineligible finding. Record the concrete verification
224
255
  in a `05-review-findings.md` row, keeping its Severity and Decision headers
225
- unchanged and setting Decision to `accepted`. Watch for the round-2
226
- halt signal across repeated review-fix cycles (see Round-2 halt rule
227
- below). By the second round-2 halt signal or the third `fix_required`
256
+ unchanged and setting Decision to `accepted`. Halt at the first
257
+ `recurrence: repeated` finding whose class a previous round's fix already addressed and
258
+ whose `introduced_by_delta` is `yes` or `unknown` (see Round-2 halt rule below): before any
259
+ further implementer spawn on that task, name split or redesign in `03-decisions.md`. By the
260
+ second round-2 halt signal or the third `fix_required`
228
261
  review round on the same task, apply the Review-round escalation budget
229
262
  (see below) instead of running another round unaided. When a fix round's review meets the trigger of the Fix-regression decision point (defined only in [review and recovery](review-and-recovery.md), not restated here), record the Fix-regression decision point before another fix round starts. At an advisor
230
263
  trigger (architectural uncertainty, conflicting
@@ -305,7 +338,8 @@ is an honest failure, not a misfire. `skip`, `acknowledged`, `limitation`, and
305
338
  inconclusive results remain non-passes and cannot be silently accepted. When a
306
339
  repository has `docs/okf/`, include its bundle check in every set regardless of
307
340
  which files changed. This is a documented convention, not an OW execution
308
- engine or runtime schema validator.
341
+ engine or runtime schema validator. A quoted probe verdict is not a named result
342
+ of the verification set, so the set's missing-or-extra rule does not apply to it.
309
343
 
310
344
  # Persisted probe plans
311
345
 
@@ -7,7 +7,8 @@ A subagent return is a misfire, not evidence, when its output does not parse
7
7
  against its role's output contract, including an implementer return that
8
8
  omits the `mutation_probes` field even though the task assignment named
9
9
  mutation probes to run, or that omits the `commits` field even though the
10
- task assignment asked for a commit. When a subagent returns near-instantly
10
+ task assignment asked for a commit, or that omits the `class_closure`
11
+ field on any round after the task's first. When a subagent returns near-instantly
11
12
  with no tool activity, treat that as a misfire signal rather than proof:
12
13
  check the output against the contract with extra suspicion, and accept it
13
14
  only if it is contract-valid and the assignment was answerable from the
@@ -44,7 +45,11 @@ cases. Ship the healthy half on its own verification, and refile the
44
45
  removed half as its own task carrying the measurement history that led to
45
46
  the split. Acceptance criteria that cannot be satisfied this way go to the
46
47
  operator as a merge-hold (hold the change unmerged and hand the decision to
47
- the operator).
48
+ the operator). Step 8 of the detailed workflow states the operational
49
+ halt: stop at the first `recurrence: repeated` finding whose class a
50
+ previous round's fix already addressed and whose `introduced_by_delta`
51
+ is `yes` or `unknown`, before any further implementer spawn on the
52
+ task, and record the split-or-redesign decision in `03-decisions.md`.
48
53
 
49
54
  ## Review-round escalation budget
50
55
 
@@ -65,6 +70,9 @@ split-or-redesign response, not instead of it.
65
70
  option is exhausted; under a `full` profile the choice falls to the
66
71
  advisor spawn or the merge-hold, under a `minimal` profile (no advisor
67
72
  subagent to spawn) it falls straight to the merge-hold.
73
+ In `single`, tier/model escalation requires a recorded switch to `delegated`.
74
+ Apply the Run mode switch procedure before assigning the next attempt to an
75
+ implementer raised according to this tier/model escalation option.
68
76
  - **Advisor spawn** (where the advisor is installed, `full` profile):
69
77
  send the advisor subagent the question "redesign, split, or hold?" and
70
78
  weigh its recommendation before deciding.
@@ -124,8 +132,8 @@ the next unpinned sentence do not converge. For such a change:
124
132
  - List the load-bearing claims of the normative site in the acceptance
125
133
  criterion, and pin each one as the whole sentence or clause that carries
126
134
  it. That list is the pin obligation. Every normative sentence the change
127
- adds or alters at that site is a claim; one left off the list is named
128
- in the criterion with the reason it is not load-bearing.
135
+ adds or alters at that site is pinned; one left unpinned is named in the
136
+ criterion with the reason it is not load-bearing.
129
137
  - Bound the reviewer's prose mutant space to that list in the briefing. A
130
138
  survivor outside the list is a scope note in the reviewer's
131
139
  `residual_risks`, not a finding, unless the reviewer shows that the
@@ -133,16 +141,16 @@ the next unpinned sentence do not converge. For such a change:
133
141
  - Bind each copy to the normative site through one shared test constant,
134
142
  and let a pointer point without restating the rule.
135
143
  - Cap test-adequacy review rounds on the change at two. A test-adequacy
136
- review round is one whose only unresolved findings are `tests` findings
137
- about pin gaps on the pinned prose; a round with any other unresolved
138
- finding is an ordinary round outside the cap. Pin gaps that remain
144
+ review round is one whose returned findings are all `tests` findings of
145
+ severity `low` or `medium` about pin gaps on the pinned prose; a round
146
+ returning any other finding is an ordinary round outside the cap. Pin gaps that remain
139
147
  become accepted notes or a follow-up.
140
148
 
141
149
  Semantic findings are exempt from the bound and from the cap: two sites
142
150
  stating different rules, a contradiction with another rule, and a false
143
151
  claim are defects at whatever severity they deserve. The cap changes
144
152
  neither the Round-2 halt rule, the Review-round escalation budget nor the
145
- Fix-regression decision point: a capped round still counts as a negative
153
+ Fix-regression decision point: a test-adequacy review round still counts as a negative
146
154
  round where it is one. The review gate is unchanged: a high or critical
147
155
  finding of any category still blocks and is never capped away, and
148
156
  accepting one follows the waiver rules. Anchored by an observed run; see
@@ -92,6 +92,21 @@ where it lives in the row's own cell.
92
92
  |---|---|---|---|---|---|---|---|---|---|---|---|
93
93
  | <!-- round --> | <!-- mutant --> | <!-- file --> | <!-- anchor --> | <!-- before --> | <!-- after --> | <!-- verified_applied_via --> | <!-- result --> | <!-- expectation --> | <!-- reason --> | <!-- restored_verified --> | <!-- replayed --> |
94
94
 
95
+ ### Class Closure
96
+
97
+ One row per fix round (any round after the task's first) that fixed a review
98
+ finding: the defect class it closed, the search command run to enumerate other
99
+ sites of that class (blank when `Closure Kind` is not `enumerated`), the sites
100
+ the search found (blank when `Closure Kind` is not `enumerated`), and the
101
+ closure kind (`enumerated | source`) the implementer reported in
102
+ `class_closure`; `not_applicable` never appears in this table, since a row
103
+ exists only for a round that fixed a finding, never for the task's first round.
104
+ An unclosed site of the row's class is named in Risks / Notes with the reason.
105
+
106
+ | Round | Class | Enumeration Command | Sites | Closure Kind |
107
+ |---|---|---|---|---|
108
+ | <!-- round --> | <!-- class --> | <!-- enumeration_command --> | <!-- sites --> | <!-- closure_kind --> |
109
+
95
110
  ### Optional Probe Plan and Result Index
96
111
 
97
112
  An optional runner-supported probe plan may be referenced here by relative
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.38.1",
3
+ "version": "0.39.0",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",