orchestrator-workflow 0.38.1 → 0.40.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,214 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.40.0] - 2026-09-23
11
+
12
+ - A verification set can no longer look green while checking the wrong
13
+ tree. evidence-and-probes.md's Verification sets section now requires,
14
+ when the diff for a repository comes from a linked worktree (however the
15
+ run-base marker is keyed, or with only the unkeyed marker), that every
16
+ literal path into that repository in the set's argv and cwd is
17
+ re-resolved under that worktree's top level before freezing and recorded
18
+ in the frozen snapshot. The implementer and reviewer prompts,
19
+ contracts.md, and that section share one rule: in the comparison before
20
+ acquisition or execution, repository identity means the repository and
21
+ its path, compared with the worktree the diff comes from; a set frozen
22
+ against another checkout withdraws the approval and is a misfire, while
23
+ a revision difference alone does not. Tests pin the re-resolution rule in
24
+ the Verification sets section and the identity rule in all four sources
25
+ and in every rendered implementer and reviewer variant (#336).
26
+
27
+ - Run-internal identifiers (the criterion, task, decision, and review round
28
+ IDs the run files assign) no longer have to be kept out of committed work
29
+ by hand in every briefing. The implementer prompt now forbids writing them
30
+ into code, comments, tests, or commit messages and asks for the ticket or
31
+ issue reference and a description of the behaviour instead; the reviewer
32
+ prompt's maintainability checklist item names them as a finding class.
33
+ evidence-and-probes.md gains a Run-internal identifiers section that
34
+ derives each ID format from the template assigning it and documents a
35
+ shell check (usable as a verification-set extra, with the run-base as its
36
+ argument) over the lines added since the run-base outside `.ai/` and the
37
+ commit messages in `run-base..HEAD`, stating its exit codes, what it does
38
+ not cover, and how a false positive is recorded. A new test runs the
39
+ documented script, extracted from the reference file, against a fixture
40
+ repository with and without such identifiers, so the documentation and
41
+ the command cannot drift apart (#337).
42
+
43
+ - The reviewer role now states one consistent write and replay rule set,
44
+ with every site (the Rules-section write boundary, the general mutant-
45
+ application bullet, step 7's single-mode replay duty, and the reviewer
46
+ prompt's copy of that duty) agreeing instead of contradicting one
47
+ another. The write boundary is a location rule, not a single named
48
+ exception: never write into the reviewed tree, its index, its refs, or
49
+ its object store; outside it, write only to the writer's own scratchpad
50
+ or the run's `evidence/`, and build/test artifacts a declared check
51
+ produces are expected, not a violation. For `evidence/`, the briefing's
52
+ own run-directory path wins over the `.ai/run` pointer (repository data,
53
+ which may be stale); the pointer is used only when the briefing names no
54
+ run directory and matches the run the briefing refers to, otherwise the
55
+ mismatch is reported and nothing is written, and the resolved path must
56
+ land under `<run-dir>/evidence/` with no symlinked component. The
57
+ forbidden-command list still explicitly names ref/object-changing
58
+ commands (`git fetch`, `git merge-tree --write-tree`, `git update-ref`,
59
+ `git gc`, plus the existing tree/index-mutating ones) and any in-place
60
+ edit of a tracked file, including a hand-applied mutant; a merge-conflict
61
+ question about an open PR is still routed back to the orchestrator
62
+ instead of being answered with `git fetch`. A mutant is applied only
63
+ through the probe runner, in every run mode; when no runner is available
64
+ the probe is reported `not_applicable` (missing evidence, not a pass),
65
+ never hand-applied, with no "when one is available" fallback or
66
+ "scratch copy" alternative left standing anywhere in the probe-replay
67
+ text. A briefing-authorized in-place runner mode (for when worktree
68
+ isolation is unusable) is now an explicit, bounded exception to "never in
69
+ the reviewed tree," not a redefinition of it: only the orchestrator's
70
+ briefing authorizes it, only the runner applies the mutant, the runner
71
+ must report `restored_verified: true` (a missing or `false` value is a
72
+ finding), the tree must be clean and at the reviewed head, and no other
73
+ agent may be active in that tree at the same time. `evidence/` is now
74
+ documented as an optional run subdirectory in the run-layout reference,
75
+ naming the orchestrator, the implementer, and the reviewer as the roles
76
+ that write into it and stating that the explorer and the advisor, being
77
+ read-only, never do. The location rule also names, explicitly, that a
78
+ write a declared check or the probe runner's own default isolation
79
+ leaves behind is expected wherever that tool places it, not an exception
80
+ to the rule. The in-place exception's conditions and the run mode
81
+ `single` duty's `not_applicable` fallback are now pinned at every site
82
+ that states them (the reviewer prompt, its rendered install variants,
83
+ and step 7's mirror), alongside the evidence-path symlink and mismatch
84
+ clauses and the merge-conflict routing sentence. The README's read-only
85
+ posture section states the reviewer's narrower write boundary the same
86
+ way (#339).
87
+
88
+ - A verification set named in the briefing by reference plus its frozen
89
+ digest and repository identity now reads, at every site that states the
90
+ rule, as the orchestrator's approval of every argv resolved from that
91
+ frozen snapshot, closing the gap where an implementer reported the
92
+ briefed extras as unresolved and repeated the same open question every
93
+ round even though the reviewer ran them under the same briefing. A
94
+ digest mismatch withdraws the approval and is reported as a misfire, and
95
+ the approval never reaches beyond the frozen snapshot: an unfrozen set,
96
+ a changed script, or anything else the snapshot does not capture still
97
+ needs the orchestrator's explicit approval before acquisition or
98
+ execution. The subagent input contract's `verification_set` section, and
99
+ the implementer and reviewer prompts' verification-set rules, now state
100
+ the approval condition in the same wording (#335).
101
+
102
+ - Before acquisition or execution, implementer.md, reviewer.md, and
103
+ contracts.md state, in identical wording, that the role compares the
104
+ frozen snapshot's effective config and scripts, preflight executable
105
+ identity/definition, and repository identity with the tree it runs in;
106
+ any mismatch withdraws the approval like a digest mismatch and is
107
+ reported as a misfire, and a change the task's own diff makes to one of
108
+ those components is outside the approval. The `verification_set` shape
109
+ in the subagent input contract and both task-slicer output copies now
110
+ also carries a `snapshot` sub-field, alongside `digest` and
111
+ `repository_identity`, naming the run-local path of the frozen snapshot
112
+ record those compared values are read from; evidence-and-probes.md's
113
+ Verification sets section defines what counts as a "script" for that
114
+ comparison and states that the orchestrator records the snapshot at the
115
+ run-local path it names in the briefing (#335).
116
+
117
+ ## [0.39.0] - 2026-09-23
118
+
119
+ - The docs-only review default now states its condition directly instead
120
+ of borrowing step 8's closure term: a review round whose entire delta
121
+ contains only explanatory documentation, comments, or citations, with no
122
+ source- or test-file edits and no semantic change to executable commands,
123
+ configuration, policy, instructions, or behavior, defaults to the
124
+ `-medium` reviewer tier with `review_method: normal`. The pinned-prose
125
+ cap counts a test-adequacy review round as one whose returned findings are
126
+ all `low` or `medium` `tests` findings about pin gaps; every normative
127
+ sentence a change adds or alters at its site is pinned, and one left
128
+ unpinned is named with the reason it is not load-bearing; a reviewer
129
+ respects a briefing that bounds the prose mutant space to a claim list
130
+ (#332).
131
+
132
+ - Clarified the review-round tier/model escalation path for `single` and
133
+ rewrapped the installed policy fence with its cited documentation.
134
+
135
+ - Mutation-probe verdict reporting now distinguishes runner-supplied,
136
+ result-only, and no-verdict cases without changing the output contract.
137
+ `result: killed` means the probe's test command reacted to the mutant
138
+ under the runner's pass predicate, or the test pass predicate declared
139
+ in the task assignment or probe plan when no runner supplies a verdict;
140
+ `survived` means it did not. `expectation: met` means the measured
141
+ result matches the expected result declared in the task assignment or
142
+ probe plan, and `violated` means it does not; both fields are
143
+ `not_applicable` when no result was measured. When a mutation-probe
144
+ runner is available, run the named probes through it and copy every
145
+ supplied `result` and `expectation` verbatim into `mutation_probes`,
146
+ never substituting your interpretation of its test output. Quote each
147
+ supplied verdict in `tests.executed`; when it supplies only `result`,
148
+ derive `expectation` from the expected result declared in the task
149
+ assignment or probe plan, and identify that declaration and derivation
150
+ there. When no machine-readable verdict is available, state that
151
+ explicitly in `tests.executed`, identify the declared test pass
152
+ predicate and expected result, and quote the observed baseline and
153
+ mutant outcomes. Derive `result` from those observations only when the
154
+ baseline passed, mutant application was verified, and the mutant test
155
+ completed under the same command and predicate; derive `expectation` by
156
+ comparing that result with the declared expected result, and label both
157
+ derivations as manual. Before transferring a probe row, compare each
158
+ copied field with the quoted verdict and each derived field with its
159
+ stated declaration and evidence. An explicit absence of a
160
+ machine-readable verdict requires the manual comparison, not resupply of
161
+ a nonexistent verdict. On a mismatch or missing required evidence,
162
+ obtain corrected evidence from the implementer or rerun the probe in
163
+ isolation, record the action in `03-decisions.md`, and keep the row
164
+ blocked from transfer until the comparison succeeds; if the evidence
165
+ cannot be obtained, record the unresolved proof rather than repeatedly
166
+ requesting an unavailable verdict. Never invent a verdict, override a
167
+ supplied field, or fill an unsupported derivation. Apply the same
168
+ evidence reporting and comparison to probes you run yourself before
169
+ recording their rows in `04-implementation-summary.md`. For probes you
170
+ run, apply the implementer's verdict-copy and manual-derivation rules to
171
+ your own measurements, reporting the quoted verdict or explicit verdict
172
+ absence and derivation evidence in `reproduction` and carrying the same
173
+ reported values into any associated finding. A quoted probe verdict is
174
+ not a named result of the verification set, so the set's
175
+ missing-or-extra rule does not apply to it. It reports per probe, in
176
+ `reproduction`, the probe, the replayed verdict or explicit verdict
177
+ absence with manual derivation evidence, and whether the measured
178
+ `result` and `expectation` match the recorded fields; a mismatch is a
179
+ finding of at least `high` and sets `matches_implementer_claim:
180
+ mismatched`. The implementer legend renders into every implementer tier
181
+ and Codex developer instructions; the revised test suite exercises all
182
+ three cases and the corresponding transfer decisions. Evidence for task
183
+ cb4ff78f-b0cb-402e-8387-f996f676f964 is recorded in the astra-old-eight
184
+ run, T-005 implementation report.
185
+
186
+ - A fix round now closes the defect class instead of the reported
187
+ instance. `assets/agents/implementer.md` (and every tier variant
188
+ rendered from it) states, for any round after a task's first, the
189
+ class-enumeration obligation (a search command plus its hit list in the
190
+ report, or a source-level closure with the reason) and a
191
+ one-mutation-probe-per-fixed-finding obligation. The implementer output
192
+ contract (`assets/skill/references/contracts.md`, mirrored in
193
+ `assets/agents/implementer.md`) gains a `class_closure` field (`kind:
194
+ enumerated | source | not_applicable` plus `command`, `sites` and
195
+ `closed: true | false`, with an unclosed site named in `risks` with the
196
+ reason); the subagent misfire rule now names the omission of
197
+ `class_closure` on any round after the task's first as a misfire, the
198
+ way it already names `mutation_probes` and `commits`.
199
+ `assets/agents/reviewer.md` gains the matching independent
200
+ class-enumeration obligation for round N+1 and now requires
201
+ `recurrence: repeated` whenever a finding's class matches an earlier
202
+ round's finding, even at a new site. `SKILL.md` step 8
203
+ (`references/evidence-and-probes.md`) now says the orchestrator halts
204
+ at the first `recurrence: repeated` finding whose `introduced_by_delta`
205
+ is `yes` or `unknown` and names split or redesign in `03-decisions.md`
206
+ before any further implementer spawn; the Round-2 halt rule paragraph in
207
+ `references/review-and-recovery.md` cross-references that halt at the
208
+ same scope, both sites pinned through one shared test constant.
209
+ `assets/templates/04-implementation-summary.md`
210
+ gains a Class Closure row per fix round (class, enumeration command,
211
+ sites, closure kind). Anchored by two observed batches: batch 48 saw 8
212
+ repeated findings of 32 review rounds, and batch 51 saw 13
213
+ repeated-finding mentions across 22 rounds, 7 of those rounds
214
+ attributable to case-level fixes or inert fixes rather than a closed
215
+ class. Anchored by an observed run: this fix-round contract was itself
216
+ applied in the fix rounds of the run that shipped it, before merge.
217
+
10
218
  ## [0.38.1] - 2026-09-20
11
219
 
12
220
  - Repository lint, not shipped in the package: rule 1 of
@@ -108,27 +316,27 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
108
316
  the check itself unchanged is a `03-decisions.md` entry, not a revision";
109
317
  the orchestrator records that entry, states in it why no evidence is
110
318
  invalidated, and communicates the corrected wording in the next
111
- delegation. Step 7 now says: "For a review round whose entire delta is a
112
- docs-only delta in the sense of step 8's docs-only closure, default to the
113
- `-medium` reviewer tier with `review_method: normal` where tier variants
114
- are installed". That refines the general tier default for this one class
115
- only, a round that touches an instruction, policy, template or prompt file
116
- keeps the general default, and the minimum review methods are untouched;
319
+ delegation. Step 7 now says: "For a review round whose entire delta contains only explanatory documentation, comments, or citations and contains no source- or test-file edits and no semantic change to executable commands, configuration, policy, instructions, or behavior, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed". That refines the general tier default for this one class only. "A review round that touches an instruction, policy, template or prompt file (for example a SKILL.md instruction) keeps the general default, whatever the file type, and the minimums named above are unaffected";
117
320
  the AGENTS.md section does not yet point to this refinement.
118
321
  `references/review-and-recovery.md` gains a "Pinned-prose changes" section
119
322
  for a change whose acceptance rests on tests that pin documentation
120
323
  wording: "A prose mutant survives exactly when its bytes sit in no
121
324
  assertion", so review rounds that hunt for the next unpinned sentence do
122
- not converge. The section asks for one normative site per rule, a claim
123
- list in the acceptance criterion as the pin obligation (every normative
124
- sentence the change adds or alters at that site is a claim, an omission is
125
- named with its reason), a reviewer briefing that bounds the prose mutant
126
- space to that list, copies bound to the normative site by one shared test
127
- constant, and: "Cap test-adequacy review rounds on the change at two." It
128
- defines the capped round, exempts semantic findings, and changes neither
129
- the Round-2 halt rule, the escalation budget, the Fix-regression decision
130
- point nor the review gate. Step 7 points to the section without restating
131
- it. Evidence (issue #300 and the change that added the Fix-regression
325
+ not converge. The section asks for one normative site per rule, a claim list in the acceptance criterion as the pin obligation. "Every normative sentence the change adds or alters at that site is pinned; one left unpinned is named in the criterion with the reason it is not load-bearing." A reviewer briefing
326
+ bounds the prose mutant space to that list, and "When the briefing
327
+ bounds the
328
+ prose mutant space to a claim list, respect that bound and put scope notes in
329
+ `residual_risks`, unless an unlisted sentence is shown to be load-bearing."
330
+ Copies are bound to the normative site by one shared test constant, and: "Cap
331
+ test-adequacy review rounds on the change at two." "A test-adequacy review
332
+ round is one whose returned findings are all `tests` findings of severity
333
+ `low` or `medium` about pin gaps on the pinned prose; a round returning any
334
+ other finding is an ordinary round outside the cap." "The cap changes neither
335
+ the Round-2 halt rule, the Review-round escalation budget nor the
336
+ Fix-regression decision point: a test-adequacy review round still counts as a
337
+ negative round where it is one." It exempts semantic findings and leaves the
338
+ review gate as it is. Step 7 points to the section without
339
+ restating it. Evidence (issue #300 and the change that added the Fix-regression
132
340
  decision point; one repository each, not a benchmark): the issue reports
133
341
  baseline revisions r1 to r3 for two wording precisions of a verification
134
342
  method, and a run in which the top reviewer tier was about half the day's
package/README.md CHANGED
@@ -241,11 +241,19 @@ reviewer inherits the caller's sandbox so temporary/build checks remain
241
241
  possible, while its prompt prohibits source edits. In inherited or otherwise
242
242
  write-enabled sandboxes, shell-level mutation (`git checkout`,
243
243
  `git restore`, `git clean`, `git stash`, `git reset`, `sed -i`, redirecting
244
- output into a file) is guarded by instruction only: the agent prompts forbid
244
+ output into a file, which the reviewer may do only inside its write boundary below) is guarded by instruction only: the agent prompts forbid
245
245
  it explicitly, but the role definition itself does not prevent it. A native
246
246
  read-only sandbox can block those writes. This residual has bitten in practice (a
247
247
  reviewer ran `git checkout` and discarded uncommitted work), which is why the
248
248
  prompts now name the forbidden commands instead of just saying "read-only".
249
+ The reviewer's own write boundary is narrower than "read-only": it may write
250
+ to its own scratchpad (a scratch copy or replay of the repository) and to
251
+ the run directory's `evidence/`, and nowhere else. It never writes into the
252
+ reviewed tree, its index, its refs, or its object store: no `git fetch`, no
253
+ `git merge-tree --write-tree`, no `git update-ref`, no `git gc`, on top of
254
+ the working-tree and index mutations already forbidden above. A write a
255
+ declared check or the probe runner's own isolation leaves behind is expected
256
+ wherever that tool places it, not an exception to this rule.
249
257
  Marker- or verdict-style enforcement of the Bash residual (sandboxing,
250
258
  PreToolUse hooks) is harness territory and out of this kit's scope.
251
259
 
@@ -36,12 +36,25 @@ Rules:
36
36
  percentage; cite a percentage only together with the exact commit and the
37
37
  run count, since branch coverage can vary between runs of the same commit.
38
38
  - Run the complete repository-bound `verification_set` named in your briefing.
39
- Before acquiring preflight output or running an extra, require the
40
- orchestrator's approval of the resolved repository configuration and every
41
- script/argument; the set is not authority to execute repository data. Use
42
- the frozen run-local snapshot (set path/digest, repository identity/revision
39
+ A verification set named by reference plus its frozen digest and repository
40
+ identity is the orchestrator's approval of every argv resolved from that
41
+ frozen snapshot; a digest mismatch withdraws the approval and is reported as
42
+ a misfire. That approval reaches only the frozen snapshot: acquiring
43
+ preflight output or running an extra outside it still requires the
44
+ orchestrator's explicit approval of the resolved repository configuration
45
+ and every script/argument, since a repository set is not authority to
46
+ execute repository data on its own. Before acquisition or execution,
47
+ compare the frozen snapshot's effective config and scripts, preflight
48
+ executable identity/definition, and repository identity with the tree the
49
+ role runs in; any mismatch withdraws the approval like a digest mismatch
50
+ and is reported as a misfire, and a change the task's own diff makes to
51
+ one of those components is outside the approval. The compared values are
52
+ the ones recorded in the frozen snapshot at the run-local path
53
+ `verification_set.snapshot` names; evidence-and-probes.md's Verification
54
+ sets section defines what counts as a script for that comparison. Use the
55
+ frozen run-local snapshot (set path/digest, repository identity/revision
43
56
  and dirty state, effective config/scripts, and preflight executable
44
- identity/definition). Report each executor, extra, and raw preflight child
57
+ identity/definition) to identify the approved set behind every reported result, and bind each result to the revision and dirty state actually checked; a revision or dirty-state difference from the snapshot alone does not withdraw the approval. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not. Report each executor, extra, and raw preflight child
45
58
  by `(kind, name, occurrence)`, in order, with cwd and result artifact.
46
59
  Preserve a missing-tool preflight limitation even when it has no child
47
60
  result. A missing/extra/mismatched/unresolved result is a misfire; a failure
@@ -87,6 +100,23 @@ Rules:
87
100
  `result` alone is not: report it as such (`result` `survived` or
88
101
  `not_applicable` with the reason) and resolve it before the next
89
102
  reviewer spawn.
103
+ - On any round after the task's first, when the round fixes a review
104
+ finding, enumerate the defect's class before returning: run a search
105
+ command for the pattern the finding's fix addresses and list every hit
106
+ in the report, or state a source-level closure (the fix removes the
107
+ pattern at its one source, closing the whole class without a search)
108
+ and say why in `summary`. Report the result in the output contract's
109
+ `class_closure` field: `kind: enumerated | source | not_applicable`
110
+ (`not_applicable` only on the task's first round, when there is no
111
+ review finding yet to fix), the search `command` that produced the hit
112
+ list (empty when `kind` is not `enumerated`), and the `sites` list of
113
+ every hit found (empty when `kind` is not `enumerated`), and `closed:
114
+ true` when every site the round found (by search or by source-level
115
+ closure) is fixed this round, `false` when a found site is not; name an unclosed site in `risks` with
116
+ the reason. On the task's first round `kind` is `not_applicable`, `command` and `sites` are empty,
117
+ and `closed` is `true`, since no site was found to leave open. Run one mutation probe per review
118
+ finding fixed in the round, in addition to any probe the assignment names, and report each one in
119
+ `mutation_probes`. A fix-round return without `class_closure` is a misfire per the misfire rule.
90
120
  - A persisted probe-plan reference may stand in for a repeated inline mutant
91
121
  definition when it resolves to a path plus immutable revision or hash and the
92
122
  mutant locator/index. Resolve it before running; a missing, stale, or
@@ -96,14 +126,33 @@ Rules:
96
126
  Never rewrite a prior plan for new code; record intentional supersession and
97
127
  rationale in run state before using a replacement.
98
128
  - When a verify runner is available, run it for the checks the acceptance
99
- criteria name and report its summary under `tests.executed`; when a
100
- mutation-probe runner is available, run the named probes through it and
101
- copy its fields into `mutation_probes`; when the runner reports a
102
- probe's mutant record (`file`, `anchor`, `before`, `after`) separately
129
+ criteria name and report its summary under `tests.executed`. When a
130
+ mutation-probe runner is available, run the named probes through it and copy
131
+ every supplied `result` and `expectation` verbatim into `mutation_probes`,
132
+ never substituting your interpretation of its test output. When the
133
+ runner reports a probe's mutant record (`file`, `anchor`, `before`,
134
+ `after`) separately
103
135
  from its result fields (`verified_applied_via`, `result`, `expectation`,
104
136
  `reason`, `restored_verified`), take the definition fields from that
105
137
  mutant record so the copied report still carries all eleven
106
- `mutation_probes` sub-fields. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
138
+ `mutation_probes` sub-fields. `result: killed` means the probe's test
139
+ command reacted to the mutant under the runner's pass predicate, or the
140
+ test pass predicate declared in the task assignment or probe plan when
141
+ no runner supplies a verdict; `survived` means it did not. `expectation:
142
+ met` means the measured result matches the expected result declared in
143
+ the task assignment or probe plan, and `violated` means it does not;
144
+ both fields are `not_applicable` when no result was measured. Quote each
145
+ supplied verdict in `tests.executed`; when it supplies only `result`,
146
+ derive `expectation` from the expected result declared in the task
147
+ assignment or probe plan, and identify that declaration and derivation
148
+ there. When no machine-readable verdict is available, state that
149
+ explicitly in `tests.executed`, identify the declared test pass
150
+ predicate and expected result, and quote the observed baseline and
151
+ mutant outcomes. Derive `result` from those observations only when the
152
+ baseline passed, mutant application was verified, and the mutant test
153
+ completed under the same command and predicate; derive `expectation` by
154
+ comparing that result with the declared expected result, and label both
155
+ derivations as manual.
107
156
  - Run every long test, build, or mutation-probe command in the foreground
108
157
  and wait for it to finish before returning. When one foreground call
109
158
  cannot hold it to completion, poll the backgrounded run to completion
@@ -151,7 +200,7 @@ Rules:
151
200
  check. A background monitor is no substitute for those returns.
152
201
  - Only write a verification claim (for example "Verified by ...") in a code
153
202
  comment, commit message, or your report for a check you actually ran and
154
- measured yourself; never claim a run you did not execute.
203
+ measured yourself; never claim a run you did not execute. Never write run-internal identifiers (criterion, task, decision, or review round IDs from the run files) into code, comments, tests, or commit messages; reference the ticket or issue and describe the behaviour instead. evidence-and-probes.md's Run-internal identifiers section defines these IDs and documents a check for them.
155
204
  - Do not refactor beyond the task scope, do not fix unrelated issues, do not
156
205
  expand the task. Report anything noteworthy as a risk or open question
157
206
  instead.
@@ -210,6 +259,12 @@ mutation_probes:
210
259
  reason: ""
211
260
  restored_verified: ""
212
261
  replayed: false | true
262
+ class_closure:
263
+ kind: enumerated | source | not_applicable
264
+ command: ""
265
+ sites:
266
+ - ""
267
+ closed: true | false
213
268
  risks:
214
269
  - severity: low | medium | high
215
270
  description: ""
@@ -53,12 +53,28 @@ Check, at minimum:
53
53
  returned `criterion_evidence` references to every assigned frozen criterion;
54
54
  required empty references remain unresolved and block acceptance.
55
55
  - Verification set: independently run the complete repository-bound
56
- `verification_set` named in the briefing. Before acquisition or execution,
57
- confirm the orchestrator approved the resolved effective configuration and
58
- scripts; a repository set is not execution authority. Compare the frozen
59
- snapshot's set path/digest, repository identity/revision/dirty state,
60
- effective config/scripts, and preflight executable identity/definition.
61
- Report every ordered `(kind, name, occurrence)` executor, extra, and raw
56
+ `verification_set` named in the briefing. A verification set named by
57
+ reference plus its frozen digest and repository identity is the
58
+ orchestrator's approval of every argv resolved from that frozen snapshot; a
59
+ digest mismatch withdraws the approval and is reported as a misfire. That
60
+ approval reaches only the frozen snapshot: acquiring or executing anything
61
+ outside it still requires confirming the orchestrator approved the resolved
62
+ effective configuration and scripts, since a repository set is not
63
+ authority to execute repository data on its own. Before acquisition or
64
+ execution, compare the frozen snapshot's effective config and scripts,
65
+ preflight executable identity/definition, and repository identity with the
66
+ tree the role runs in; any mismatch withdraws the approval like a digest
67
+ mismatch and is reported as a misfire, and a change the task's own diff
68
+ makes to one of those components is outside the approval. The compared
69
+ values are the ones recorded in the frozen snapshot at the run-local path
70
+ `verification_set.snapshot` names; evidence-and-probes.md's Verification
71
+ sets section defines what counts as a script for that comparison. Use the
72
+ frozen run-local snapshot (set path/digest, repository identity/revision
73
+ and dirty state, effective config/scripts, and preflight executable
74
+ identity/definition) to identify the approved set behind every reported result, and bind
75
+ each result to the revision and dirty state actually checked; a revision or dirty-state difference from the snapshot alone does not withdraw the approval. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not. Report every
76
+ ordered `(kind, name, occurrence)`
77
+ executor, extra, and raw
62
78
  preflight child with cwd and result artifact. A missing tool may be a
63
79
  limitation with no child, never a pass; disabled required categories are
64
80
  gaps. Missing/extra/mismatched/unresolved results are misfires, while a
@@ -77,8 +93,8 @@ Check, at minimum:
77
93
  calibrated to sit inside the output's own run-to-run noise (timing digits,
78
94
  temporary-directory names) is not a regression test; the fix is to pin
79
95
  the argument under test in-process, or assert the actual contract (a
80
- bound, or the presence of a warning), never a byte ceiling.
81
- - Maintainability: naming, dead code, needless abstraction, doc drift.
96
+ bound, or the presence of a warning), never a byte ceiling. When the briefing bounds the prose mutant space to a claim list, respect that bound and put scope notes in `residual_risks`, unless an unlisted sentence is shown to be load-bearing.
97
+ - Maintainability: naming, dead code, needless abstraction, doc drift. Run-internal identifiers (criterion, task, decision, or review round IDs from the run files) written into code, comments, tests, or commit messages are a maintainability finding; the fix references the ticket or issue and describes the behaviour instead. evidence-and-probes.md's Run-internal identifiers section defines these IDs and documents a check for them.
82
98
  - Placement: does the change add org-, machine-, or point-in-time-bound
83
99
  evidence (dates, sample sizes, task ids, home paths, incident tallies) to a
84
100
  reusable instruction file (a skill, an agent prompt, an AGENTS.md section, a
@@ -87,8 +103,21 @@ Check, at minimum:
87
103
  - Recurrence: when the briefing tells you this is not the task's first
88
104
  review round, classify each finding as `new` or `repeated` against the
89
105
  earlier rounds you were told about; on a first round every finding is
90
- `new` by definition. The orchestrator uses this to detect the
91
- review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
106
+ `new` by definition. Class match, not site match, decides this: set
107
+ `repeated` whenever a finding's defect class matches an earlier round's
108
+ finding, even when this instance sits at a site the earlier round never
109
+ touched. On any round after the task's first, run your own
110
+ class-enumeration search independent of the implementer's
111
+ `class_closure` report: search for the pattern the earlier finding's fix
112
+ addressed, and compare what your own search returns against the
113
+ implementer's `class_closure.sites` list (or its `source` closure
114
+ reason); a site your search finds that the implementer's report omits
115
+ is itself a finding, classified under the ordinary severity gate by
116
+ the underlying defect's own severity, not by a fixed floor. In run
117
+ mode `single` there is no separate implementer `class_closure` report
118
+ to compare against; compare your search's hits against the Class Closure row of
119
+ `04-implementation-summary.md` plus its Risks / Notes section instead. The orchestrator
120
+ uses this to detect the review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
92
121
  - GitHub Actions shell replay: for any diff that adds or changes a GitHub
93
122
  Actions `run:` step, replay it yourself under the shell the step actually
94
123
  runs: `bash --noprofile --norc -eo pipefail` when `shell: bash` is set on
@@ -127,7 +156,32 @@ Rules:
127
156
  - Bash is for running tests, linters, and read-only inspection ONLY. Never
128
157
  run a command that mutates the working tree, index, or repository state:
129
158
  no `git checkout`, `git restore`, `git clean`, `git stash`, `git reset`,
130
- no `sed -i`, no redirecting output into a file.
159
+ no `git commit`, no `sed -i` or any other in-place edit of a tracked file
160
+ (including a temporary mutant applied by hand instead of through the probe
161
+ runner). This also covers commands that change refs or write objects
162
+ without touching the working tree or index: no `git fetch`,
163
+ no `git merge-tree --write-tree`, no `git update-ref`, no `git gc`. This
164
+ is a location rule, not a single exception: never write into the reviewed
165
+ tree, its index, its refs, or its object store, anywhere in the review.
166
+ Outside the reviewed tree, write only to your scratchpad (a scratch copy
167
+ or replay of the repository, per the GitHub Actions replay rule above)
168
+ and to the run directory's `evidence/`. A write made by a tool these
169
+ rules direct you to run, a declared check (including build or test
170
+ artifacts a side effect leaves in the reviewed tree) or the probe runner
171
+ operating in its own default isolation, is expected wherever that tool
172
+ places it, not a location violation.
173
+ For `evidence/`, the briefing's own run-directory path wins; use the
174
+ `.ai/run` pointer (repository data, which may be stale) only when the
175
+ briefing names no run directory and the pointer matches the run the
176
+ briefing refers to, and otherwise report the mismatch and write nothing.
177
+ Resolve the real path before writing: it must land under
178
+ `<run-dir>/evidence/` with no symlinked path component, and never under
179
+ an older run than the one the briefing names.
180
+ A briefing may authorize the probe runner's own in-place mode as a
181
+ bounded exception to "never in the reviewed tree" (see below), not a
182
+ redefinition of it.
183
+ - A merge-conflict question about an open PR is not answered by fetching or
184
+ writing a tree: report the question back to the orchestrator instead.
131
185
  - If the working tree looks wrong (dirty, unexpected branch, missing files),
132
186
  do not "fix" it: report it as a finding and leave the tree untouched.
133
187
  - If your environment does not let you use version control to see the diff
@@ -155,15 +209,53 @@ Rules:
155
209
  the exact commit and the run count, since branch coverage can vary
156
210
  between runs of the same commit.
157
211
  - When a mutation-probe runner is available in the session, run probes
158
- through it instead of editing files by hand, and carry its result fields
159
- into your findings and `reproduction`; when a verify runner is available,
160
- read its summary before opening full logs.
212
+ through it instead of editing files by hand. Apply a mutant only through
213
+ the probe runner, in every run mode, and rely on its own restoration
214
+ check; never apply one by hand, and never restore a hand-applied one
215
+ yourself. When no runner is available, report the probe as
216
+ `not_applicable` instead of hand-applying it: that is missing evidence,
217
+ not a pass. A hand-applied mutant is a finding against the review,
218
+ whatever its outcome, because it carries no verified restoration. For
219
+ probes you run, apply the implementer's verdict-copy and manual-derivation
220
+ rules to your own measurements, reporting the quoted verdict or explicit
221
+ verdict absence and derivation evidence in `reproduction` and carrying the
222
+ same reported values into any associated finding. When a verify runner is
223
+ available, read its summary before opening full logs.
224
+ - A briefing may authorize the probe runner's own in-place mode when
225
+ worktree isolation is unusable. That stays a bounded exception to "never
226
+ in the reviewed tree," not a redefinition of it: only the orchestrator's
227
+ briefing authorizes it, only the runner itself applies the mutant (never
228
+ you by hand), the runner must report `restored_verified: true` for every
229
+ such probe (a missing or `false` value is a finding), the tree must be
230
+ clean and at the reviewed head before the runner starts, and no other
231
+ agent may be active in that tree at the same time (see the concurrency
232
+ rule in step 7 of the detailed workflow).
161
233
  - A reviewer briefing may identify a replayed probe through a resolved
162
234
  immutable probe-plan reference (path plus revision/hash and mutant
163
235
  locator/index) rather than repeat its inline definition. Verify the plan and
164
236
  result bind the checked state, cwd, attempt, expectation, application, and
165
237
  restoration; a plan alone, stale reference, or unresolved reference is not
166
- evidence. Legacy inline probe reports remain valid. When the briefing names run mode `single`, the orchestrator implemented the change itself and nobody has cross-checked its probe evidence: replay every named orchestrator probe, where named means the briefing gives its full definition or a resolved immutable plan-and-result reference (in a scratch copy or an isolating probe runner, never in the reviewed tree), and state in `reproduction`, per probe, the replayed verdict and whether it matches the recorded `result` and `expectation`. Do not skip a named probe in that mode, under any `review_method`; any mismatch also sets `matches_implementer_claim: mismatched`. A mismatch is a finding of at least `high`; a probe given only by id is `not_applicable` and is missing evidence, not a pass, and so is a briefing in that mode that names no probe. Without that mode line in the briefing this obligation does not exist.
238
+ evidence. Legacy inline probe reports remain valid.
239
+ When the briefing names run mode `single`, the orchestrator implemented
240
+ the change itself
241
+ and nobody has cross-checked its probe evidence: replay every named
242
+ orchestrator probe, where named means the briefing gives its full
243
+ definition or a resolved immutable plan-and-result reference, through the
244
+ probe runner only, never in the reviewed tree except under the authorized
245
+ in-place mode above; when no runner is available, report the probe as
246
+ `not_applicable`. It reports per probe, in `reproduction`, the probe, the
247
+ replayed verdict
248
+ or explicit verdict absence with manual derivation evidence, and whether
249
+ the measured `result` and `expectation` match the recorded fields; a
250
+ mismatch is a finding of at least `high` and sets
251
+ `matches_implementer_claim: mismatched`. Do not skip a named probe in
252
+ that mode, under any `review_method`; a probe that cannot run through the
253
+ runner is reported `not_applicable`, not skipped; any mismatch also sets
254
+ `matches_implementer_claim: mismatched`. A mismatch is a finding of at
255
+ least `high`; a probe given only by id is `not_applicable` and is
256
+ missing evidence, not a pass, and so is a briefing in that mode that
257
+ names no probe. Without that mode line in the briefing this obligation
258
+ does not exist.
167
259
 
168
260
  Return exactly this structure as your final output, nothing else:
169
261
  ```yaml
@@ -46,7 +46,9 @@ Rules:
46
46
  selection above to those fields for a recorded original contract.
47
47
  - Include a repository-bound `verification_set` reference in every implementer
48
48
  and reviewer briefing: its checked-in path, repository identity, and
49
- run-local frozen snapshot. The orchestrator approves effective config and
49
+ run-local frozen snapshot. The digest recorded in the briefing is what
50
+ carries the orchestrator's approval of the resolved argv to the implementer
51
+ and reviewer. The orchestrator approves effective config and
50
52
  scripts before any preflight acquisition or command execution; the set does
51
53
  not grant that authority. Include an ordered bundle check whenever the
52
54
  repository has `docs/okf/`, regardless of task scope.
@@ -95,6 +97,9 @@ tasks:
95
97
  - ""
96
98
  verification_set:
97
99
  reference: ""
100
+ digest: ""
101
+ repository_identity: ""
102
+ snapshot: ""
98
103
  risk: low | medium | high
99
104
  recommended_order:
100
105
  - T-001
@@ -1,12 +1,14 @@
1
1
  <!-- orchestrator-workflow:begin -->
2
2
  ## Agentic Coding Workflow
3
3
 
4
- This repository uses an orchestrator-led agent workflow, installed and updated by
4
+ This repository uses an orchestrator-led agent workflow, installed and
5
+ updated by
5
6
  [orchestrator-workflow](https://github.com/LanNguyenSi/agent-dx/tree/master/packages/orchestrator-workflow).
6
7
 
7
8
  The primary agent acts as the orchestrator. It owns the goal, planning, task
8
- validation, delegation, final acceptance, and the operator handoff. Non-trivial
9
- review is delegated to a narrow subagent; which agent implements non-trivial work follows from the run mode (Core rules). The full procedure
9
+ validation, delegation, final acceptance, and the operator handoff.
10
+ Non-trivial review is delegated to a narrow subagent; which agent implements
11
+ non-trivial work follows from the run mode (Core rules). The full procedure
10
12
  and the subagent I/O contracts live in the `orchestrator-workflow` skill.
11
13
 
12
14
  ### Core rules
@@ -21,10 +23,14 @@ and the subagent I/O contracts live in the `orchestrator-workflow` skill.
21
23
  inline with the same read-only discipline instead.
22
24
  - The orchestrator plans features itself. It may delegate task slicing, but it
23
25
  validates the sliced tasks before implementation starts.
24
- - Non-trivial implementation follows the run mode recorded in `00-goal.md`. `delegated`, the default, sends it to narrow implementer subagents, one task
25
- per subagent; in `single` the orchestrator implements one coherent workstream itself; `batch` runs implementers in parallel worktrees. The skill's Run mode section defines the modes and how to choose one.
26
+ - Non-trivial implementation follows the run mode recorded in `00-goal.md`.
27
+ `delegated`, the default, sends it to narrow implementer subagents, one
28
+ task per subagent; in `single` the orchestrator implements one coherent
29
+ workstream itself; `batch` runs implementers in parallel worktrees. The
30
+ skill's Run mode section defines the modes and how to choose one.
26
31
  - Non-trivial review goes to a separate reviewer subagent (see Scaling
27
- delegation). Review itself is never skipped, in any run mode, not even for docs or bulk
32
+ delegation). Review itself is never skipped, in any run mode, not even for
33
+ docs or bulk
28
34
  changes.
29
35
  - Final acceptance and the final answer to the operator stay with the
30
36
  orchestrator.
@@ -41,7 +47,8 @@ default, not a ritual.
41
47
  solution; skip it when the change is well understood. Under a `minimal`
42
48
  profile there is no explorer subagent to spawn; run this step inline
43
49
  instead.
44
- - Slicing and, in run modes `delegated` and `batch`, implementer subagents are for non-trivial work: multiple files,
50
+ - Slicing and, in run modes `delegated` and `batch`, implementer subagents are
51
+ for non-trivial work: multiple files,
45
52
  real logic, or anything that benefits from decomposition or a fresh context.
46
53
  Under a `minimal` profile there is no task-slicer subagent; the orchestrator
47
54
  slices inline with the same contract.
@@ -55,7 +62,9 @@ default, not a ritual.
55
62
  scripts, hand-edited lockfiles, cross-major overrides, or anything the
56
63
  operator flags high-risk; `normal` fits only docs, renames, or batch
57
64
  cosmetics; `rigorous` is the default otherwise. Never pair `adversarial`
58
- with the `-medium` reviewer tier; tiers themselves are unchanged. A docs-only delta has its own review default; the skill's Delegate review step states it.
65
+ with the `-medium` reviewer tier; tiers themselves are unchanged. A
66
+ docs-only delta has its own review default; the skill's Delegate review step
67
+ states it.
59
68
  - When tier variants are installed (manifest `tiers: true`), the orchestrator
60
69
  picks the effort tier per task by complexity and risk, at its own judgment.
61
70
  The unsuffixed default subagent is the normal case; `-high`/`-xhigh` fit
@@ -111,7 +120,8 @@ trivial change.
111
120
  - Medium and low findings are addressed or consciously accepted at the
112
121
  orchestrator's judgment.
113
122
  - After independent review, the orchestrator may close a docs-only delta
114
- without another reviewer round only when its entire unreviewed delta is explanatory
123
+ without another reviewer round only when its entire unreviewed delta is
124
+ explanatory
115
125
  documentation, comments, or citations; has no source- or test-file edits or
116
126
  semantic changes to executable commands, configuration, policy,
117
127
  instructions, or behavior; and closes only low/medium documentation or
@@ -128,9 +138,13 @@ trivial change.
128
138
  the exhausted tier path falls straight to the merge-hold), or an
129
139
  operator merge-hold, and adds a row (task, choice, reason) to
130
140
  `03-decisions.md`'s Review-round escalation table, then sets the
131
- `review-round-escalation` marker to the most recent choice. A negative round
141
+ `review-round-escalation` marker to the most recent choice.
142
+ In `single`, tier/model escalation requires a recorded switch to `delegated`.
143
+ A negative round
132
144
  has an `acceptance_recommendation` of `fix_required` or `reject`; a misfired
133
- review is not a round. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the three is
145
+ review is not a round. A negative round counts only with at least one
146
+ introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the
147
+ three is
134
148
  picked is judgment; that one is picked and recorded is not. Escalating
135
149
  never substitutes for a review round and comes in addition to the halt
136
150
  rule's split-or-redesign response, not instead of it.
@@ -175,7 +189,8 @@ Workflow state lives under `.ai/`:
175
189
  routing selections.
176
190
  - Every worktree a run touches carries a `.ai/run` pointer (absolute path of
177
191
  the run directory, gitignored) and `00-goal.md` carries one
178
- `run-base[<repo-basename>]` marker per repository for multi-repo runs, next to the run's `mode` marker.
192
+ `run-base[<repo-basename>]` marker per repository for multi-repo runs, next
193
+ to the run's `mode` marker.
179
194
 
180
195
  ### Models
181
196
 
@@ -60,6 +60,9 @@ context:
60
60
  relevant_docs: []
61
61
  verification_set:
62
62
  reference: ""
63
+ digest: ""
64
+ repository_identity: ""
65
+ snapshot: ""
63
66
  constraints:
64
67
  - ""
65
68
  allowed_changes:
@@ -72,8 +75,28 @@ expected_output:
72
75
 
73
76
  `verification_set.reference` identifies the checked-in set selected for this
74
77
  repository. The briefing also carries its repository identity and run-local
75
- frozen snapshot; those resolved values are evidence metadata, not a new
76
- authority to execute repository configuration or scripts.
78
+ frozen snapshot: naming that set by reference plus its frozen digest and
79
+ repository identity, as delegated in the briefing, is the orchestrator's
80
+ approval of every argv resolved from that frozen snapshot; a digest mismatch
81
+ withdraws the approval and is reported as a misfire.
82
+ `verification_set.snapshot` names the run-local path of the frozen snapshot
83
+ record. The reference-plus-digest form shown above is sufficient by itself;
84
+ neither role needs the argv repeated argument-by-argument to run it. That
85
+ approval reaches only the frozen
86
+ snapshot: an unfrozen set, a changed script, or anything the snapshot does not
87
+ capture still needs the orchestrator's explicit approval before acquisition or
88
+ execution, since a repository set is not authority to execute repository data
89
+ on its own. Before acquisition or execution, compare the frozen snapshot's
90
+ effective config and scripts, preflight executable identity/definition, and
91
+ repository identity with the tree the role runs in; any mismatch withdraws
92
+ the approval like a digest mismatch and is reported as a misfire, and a
93
+ change the task's own diff makes to one of those components is outside the
94
+ approval. The compared values are the ones recorded in the frozen snapshot at
95
+ the run-local path `verification_set.snapshot` names; evidence-and-probes.md's
96
+ Verification sets section defines what counts as a script for that
97
+ comparison. This mirrors the re-resolve rule in evidence-and-probes.md:
98
+ re-resolve when an executable definition, effective config/script, tool
99
+ identity, set digest, or approved snapshot changes. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not.
77
100
 
78
101
  ## Implementer output contract
79
102
 
@@ -113,6 +136,12 @@ mutation_probes:
113
136
  reason: ""
114
137
  restored_verified: ""
115
138
  replayed: false | true
139
+ class_closure:
140
+ kind: enumerated | source | not_applicable
141
+ command: ""
142
+ sites:
143
+ - ""
144
+ closed: true | false
116
145
  risks:
117
146
  - severity: low | medium | high
118
147
  description: ""
@@ -125,8 +154,27 @@ commits:
125
154
 
126
155
  Follow [evidence-and-probes.md workflow step 6](evidence-and-probes.md#workflow)
127
156
  for implementation evidence, verification, mutation probes, and replay. For
128
- output-field semantics and commit reporting, follow the installed
129
- implementer role prompt. Return the selected contract's YAML envelope. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
157
+ commit reporting, follow the installed implementer role prompt. Return the
158
+ selected contract's YAML envelope. `result: killed` means the probe's test
159
+ command reacted to the mutant under the runner's pass predicate, or the
160
+ test pass predicate declared in the task assignment or probe plan when no
161
+ runner supplies a verdict; `survived` means it did not. `expectation: met`
162
+ means the measured result matches the expected result declared in the task
163
+ assignment or probe plan, and `violated` means it does not; both fields
164
+ are `not_applicable` when no result was measured. When a mutation-probe
165
+ runner is available, run the named probes through it and copy every
166
+ supplied `result` and `expectation` verbatim into `mutation_probes`, never
167
+ substituting your interpretation of its test output. Quote each supplied
168
+ verdict in `tests.executed`; when it supplies only `result`, derive
169
+ `expectation` from the expected result declared in the task assignment or
170
+ probe plan, and identify that declaration and derivation there. When no
171
+ machine-readable verdict is available, state that explicitly in
172
+ `tests.executed`, identify the declared test pass predicate and expected
173
+ result, and quote the observed baseline and mutant outcomes. Derive
174
+ `result` from those observations only when the baseline passed, mutant
175
+ application was verified, and the mutant test completed under the same
176
+ command and predicate; derive `expectation` by comparing that result with
177
+ the declared expected result, and label both derivations as manual.
130
178
 
131
179
  ## Reviewer output contract
132
180
 
@@ -256,6 +304,9 @@ tasks:
256
304
  - ""
257
305
  verification_set:
258
306
  reference: ""
307
+ digest: ""
308
+ repository_identity: ""
309
+ snapshot: ""
259
310
  risk: low | medium | high
260
311
  recommended_order:
261
312
  - T-001
@@ -112,7 +112,21 @@ directory and the subagents.
112
112
  `03-decisions.md` and consolidate evidence in
113
113
  `04-implementation-summary.md`, recording each probe the implementer
114
114
  reports as a row in `04-implementation-summary.md`'s Mutation Probes
115
- subsection, with the round it was named in. Before transferring a probe row, compare its `result` and `expectation` with the runner verdict quoted in `tests.executed`; on a mismatch, or when a verdict the runner states is not quoted, resupply it (ask the same implementer for the verdict, respawn one when it is gone, or rerun the probe yourself in isolation), record the resupply in `03-decisions.md`, and treat it as a transfer blocker rather than a misfire, since the return itself parses; never infer either field. A quoted probe verdict is not a named result of the verification set, so the set's missing-or-extra rule does not apply to it. Each row's Before/After
115
+ subsection, with the round it was named in. Before transferring a probe
116
+ row, compare each copied field with the quoted verdict and each derived
117
+ field with its stated declaration and evidence. An explicit absence of
118
+ a machine-readable verdict requires the manual comparison, not resupply
119
+ of a nonexistent verdict. On a mismatch or missing required evidence,
120
+ obtain corrected evidence from the implementer or rerun the probe in
121
+ isolation, record the action in `03-decisions.md`, and keep the row
122
+ blocked from transfer until the comparison succeeds; if the evidence
123
+ cannot be obtained, record the unresolved proof rather than repeatedly
124
+ requesting an unavailable verdict. Never invent a verdict, override a
125
+ supplied field, or fill an unsupported derivation. Apply the same
126
+ evidence reporting and comparison to probes you run yourself before
127
+ recording their rows in `04-implementation-summary.md`. A quoted probe
128
+ verdict is not a named result of the verification set, so the set's
129
+ missing-or-extra rule does not apply to it. Each row's Before/After
116
130
  cells hold a single-line excerpt; when the mutant's actual before/after
117
131
  text is multi-line or contains an unescaped `|`, or the mutant is a
118
132
  patch/diff rather than a text swap, the full text or diff goes in the
@@ -152,7 +166,7 @@ directory and the subagents.
152
166
  cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
153
167
  never substitutes for it: do not pair `adversarial` with the `-medium`
154
168
  reviewer tier, a budget mismatch that names probes without the effort to run
155
- them; tiers themselves are unchanged by this axis. For a review round whose entire delta is a docs-only delta in the sense of step 8's docs-only closure, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed. This refines the general tier default above for that one class only: there `-medium` is the default and a higher tier is the non-default choice recorded with a one-line reason. A review round that touches an instruction, policy, template or prompt file keeps the general default, whatever the file type, and the minimums named above are unaffected. For a change whose acceptance rests on tests that pin documentation wording, write the briefing as the Pinned-prose changes section of [review and recovery](review-and-recovery.md) requires. When the reviewer's
169
+ them; tiers themselves are unchanged by this axis. For a review round whose entire delta contains only explanatory documentation, comments, or citations and contains no source- or test-file edits and no semantic change to executable commands, configuration, policy, instructions, or behavior, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed. This refines the general tier default above for that one class only: there `-medium` is the default and a higher tier is the non-default choice recorded with a one-line reason. A review round that touches an instruction, policy, template or prompt file (for example a SKILL.md instruction) keeps the general default, whatever the file type, and the minimums named above are unaffected. For a change whose acceptance rests on tests that pin documentation wording, write the briefing as the Pinned-prose changes section of [review and recovery](review-and-recovery.md) requires. When the reviewer's
156
170
  environment cannot use version control to see the diff (for example a
157
171
  policy-gated repository), supply the diff as a pre-generated file in the
158
172
  briefing instead of expecting the reviewer to derive it, and have the
@@ -198,8 +212,35 @@ directory and the subagents.
198
212
  not merely their id; a probe recorded with only an id and no definition
199
213
  cannot be skipped this way and is `not_applicable`. The reviewer may
200
214
  then skip re-running the ones named by definition.
201
- The reviewer output contract itself is unchanged. In run mode `single` that skip permission does not apply: nobody but the orchestrator has seen its probe evidence. A probe counts as named when the briefing gives its full definition or a resolved immutable plan-and-result reference; an id alone does not name a probe. The orchestrator records every probe it ran in one of those two forms in `04-implementation-summary.md` before requesting review, the reviewer briefing states the run mode and names each of those probes, and the reviewer must replay every named orchestrator probe, through the probe runner when one is available and never in the reviewed tree. It reports per probe, in `reproduction`, the probe, the replayed runner verdict, and whether that verdict matches the recorded `result` and `expectation`; a mismatch is a finding of at least `high` and sets `matches_implementer_claim: mismatched`. A probe given only by id is `not_applicable` and counts as missing evidence, not as a pass, and so does a `single` briefing that names no probe at all. Never run mutation probes
202
- in place against a worktree a reviewer subagent is concurrently reviewing;
215
+ The reviewer output contract itself is unchanged. In run mode `single`
216
+ that skip permission does not apply: nobody but the orchestrator has
217
+ seen its probe evidence. A probe counts as named when the briefing
218
+ gives its full definition or a resolved immutable plan-and-result
219
+ reference; an id alone does not name a probe. The orchestrator records
220
+ every probe it ran in one of those two forms in
221
+ `04-implementation-summary.md` before requesting review, the reviewer
222
+ briefing states the run mode and names each of those probes, and the
223
+ reviewer must replay every named orchestrator probe, through the probe
224
+ runner only, never in the reviewed tree; when no runner is available,
225
+ report the probe as `not_applicable`, which is missing evidence, not a
226
+ pass. A briefing
227
+ may authorize the runner's own in-place mode as a bounded exception to
228
+ "never in the reviewed tree" when worktree isolation is unusable: only
229
+ the orchestrator's briefing authorizes it, only the runner applies the
230
+ mutant, the runner must report `restored_verified: true` for every such
231
+ probe (a missing or `false` value is a finding), the tree must be clean
232
+ and at the reviewed head before the runner starts, and no other agent may
233
+ be active in that tree at the same time (see the concurrency rule below).
234
+ It reports
235
+ per probe, in `reproduction`, the probe, the replayed verdict or
236
+ explicit verdict absence with manual derivation evidence, and whether
237
+ the measured `result` and `expectation` match the recorded fields; a
238
+ mismatch is a finding of at least `high` and sets
239
+ `matches_implementer_claim: mismatched`. A probe given only by id is
240
+ `not_applicable` and counts as missing evidence, not as a pass, and so
241
+ does a `single` briefing that names no probe at all. Never run mutation
242
+ probes in place against a worktree a reviewer subagent is concurrently
243
+ reviewing;
203
244
  isolate the probe in a separate worktree or wait until the reviewer has
204
245
  returned before probing that tree again. For an explicitly adopted v1 run,
205
246
  ask the reviewer to compare the frozen delegated criteria with the
@@ -222,9 +263,11 @@ directory and the subagents.
222
263
  documentation or maintainability findings. This option never closes a
223
264
  high/critical or other ineligible finding. Record the concrete verification
224
265
  in a `05-review-findings.md` row, keeping its Severity and Decision headers
225
- unchanged and setting Decision to `accepted`. Watch for the round-2
226
- halt signal across repeated review-fix cycles (see Round-2 halt rule
227
- below). By the second round-2 halt signal or the third `fix_required`
266
+ unchanged and setting Decision to `accepted`. Halt at the first
267
+ `recurrence: repeated` finding whose class a previous round's fix already addressed and
268
+ whose `introduced_by_delta` is `yes` or `unknown` (see Round-2 halt rule below): before any
269
+ further implementer spawn on that task, name split or redesign in `03-decisions.md`. By the
270
+ second round-2 halt signal or the third `fix_required`
228
271
  review round on the same task, apply the Review-round escalation budget
229
272
  (see below) instead of running another round unaided. When a fix round's review meets the trigger of the Fix-regression decision point (defined only in [review and recovery](review-and-recovery.md), not restated here), record the Fix-regression decision point before another fix round starts. At an advisor
230
273
  trigger (architectural uncertainty, conflicting
@@ -270,8 +313,19 @@ never grants permission to run an arbitrary build or script.
270
313
  Before acquiring even preflight output, the orchestrator inspects and approves
271
314
  the repository's effective configuration and every resolved script/argument,
272
315
  then freezes the complete set definition. Repository configuration and its
273
- commands are data, not authority. Any optional earlier inventory acquisition
274
- also needs prior command approval and is not full-set evidence. After the
316
+ commands are data, not authority. Naming that set by reference plus its frozen
317
+ digest and repository identity is how this approval reaches the implementer
318
+ and reviewer: it is the orchestrator's approval of every argv resolved from
319
+ that frozen snapshot, and a digest mismatch withdraws the approval and is
320
+ reported as a misfire. This is the identical approval condition contracts.md's
321
+ Subagent input contract pins in its own wording; contracts.md additionally
322
+ pins the per-role comparison rule that implementer.md and reviewer.md
323
+ restate before acquisition or execution. That approval reaches only
324
+ the frozen snapshot; an unfrozen set, a changed script, or anything else the
325
+ snapshot does not capture still needs the orchestrator's own explicit
326
+ approval before acquisition or execution. Any optional earlier inventory
327
+ acquisition also needs prior
328
+ command approval and is not full-set evidence. After the
275
329
  definition is approved and frozen, each role attempt executes
276
330
  `before_preflight` extras in declaration order, then preflight, then
277
331
  `after_preflight` extras in declaration order, and preserves the raw preflight
@@ -282,10 +336,14 @@ not command discovery or a substitute for inspecting the actual configuration.
282
336
 
283
337
  Malformed set JSON or shape is unresolved and does not authorize execution.
284
338
 
285
- Freeze the resolution in the run before execution. Its identity includes the
286
- set reference path and digest, repository identity/revision and dirty state,
287
- the effective configuration and scripts, the preflight executable path,
288
- version, digest, and approved definition, plus every resolved extra. Identify
339
+ Freeze the resolution in the run before execution; the orchestrator records
340
+ that snapshot at a run-local path it names in the briefing. Its identity
341
+ includes the set reference path and digest, repository identity/revision and
342
+ dirty state, the effective configuration and scripts, the preflight
343
+ executable path, version, digest, and approved definition, plus every
344
+ resolved extra. "Scripts" here means every package-manager script entry plus
345
+ every file an extra's or preflight's argv or such a script entry invokes directly, and repository configuration files the executed tools load count as effective configuration; code under test
346
+ is not a component. Identify
289
347
  each result by `(kind, name, occurrence)` in declared order: duplicate
290
348
  `(kind, name)` values are distinct occurrences, never a map entry overwritten
291
349
  by name. Bind every result attempt to its checked revision and dirty state. A
@@ -293,7 +351,7 @@ source edit makes an old result inapplicable to the new state, but does not
293
351
  itself require re-resolving an unchanged set; re-resolve when an executable
294
352
  definition, effective config/script, tool identity, set digest, or approved
295
353
  snapshot changes. An unresolvable reference is stale and invalidates the
296
- result.
354
+ result. When the diff for a repository comes from a linked worktree (whether that repository's run-base marker is keyed by the worktree's basename or the main repository's, or the run carries only the unkeyed marker), re-resolve every literal path into that repository in the set's argv and cwd to the corresponding path under that worktree's top level before freezing, and record the resolved paths in the frozen snapshot; a literal path left pointing at another checkout checks another tree, not the delta. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not.
297
355
 
298
356
  Both implementer and reviewer run the complete frozen set and report every
299
357
  named executor, extra, and raw preflight child occurrence, with cwd and result
@@ -305,10 +363,88 @@ is an honest failure, not a misfire. `skip`, `acknowledged`, `limitation`, and
305
363
  inconclusive results remain non-passes and cannot be silently accepted. When a
306
364
  repository has `docs/okf/`, include its bundle check in every set regardless of
307
365
  which files changed. This is a documented convention, not an OW execution
308
- engine or runtime schema validator.
366
+ engine or runtime schema validator. A quoted probe verdict is not a named result
367
+ of the verification set, so the set's missing-or-extra rule does not apply to it.
309
368
 
310
369
  # Persisted probe plans
311
370
 
312
371
  A persisted probe plan is an optional, runner-supported executable artifact. Its reference carries a path plus immutable revision or hash and mutant locator/index. Assignments and summaries may point to it and a result artifact instead of resending a definition; legacy inline reports remain valid.
313
372
 
314
373
  A plan alone is never evidence. A result binds plan identity to checked state, cwd, attempt, expectation, applied mutant, and restoration. Missing, stale, or unresolvable references block proof and cannot count as skipped. Never silently rewrite an existing plan for new code to turn red green; record intentional supersession and rationale when a source move requires replacement.
374
+
375
+ # Run-internal identifiers
376
+
377
+ Run-internal identifiers are the IDs the run files assign: criterion IDs in
378
+ `00-goal.md` (`AC-` plus three digits), task IDs in `02-tasks.md` (`T-` plus
379
+ three digits), decision IDs in `03-decisions.md` (`D-` plus three digits), and
380
+ review round labels (`R` plus the round number, a common key for the
381
+ `<round>` markers in `05-review-findings.md`). They mean something only
382
+ inside the run directory, which is not part of the target repository. Code,
383
+ comments, tests, and commit messages reference the ticket or issue and
384
+ describe the behaviour instead; implementer.md states this as a rule and
385
+ reviewer.md as a maintainability finding class.
386
+
387
+ The orchestrator may add the check below to a verification set as an extra
388
+ of kind `command` in the `after_preflight` phase, with `cwd` at the
389
+ repository root and the run-base recorded in `00-goal.md` for that
390
+ repository as its only argument. Its argv is `["sh", "-c", <the script
391
+ below as one string>, "sh", <run-base>]`:
392
+
393
+ ```sh
394
+ base="$1"
395
+ unset GREP_OPTIONS
396
+ ids='(^|[^A-Za-z0-9_])((AC|D|T)-[0-9]{3}|R[0-9]+)([^A-Za-z0-9_]|$)'
397
+ top=$(git rev-parse --show-toplevel) || exit 2
398
+ cd "$top" || exit 2
399
+ git rev-parse --verify --quiet "$base^{commit}" >/dev/null || exit 2
400
+ d=$(git diff --no-color --no-ext-diff --no-textconv --text -M \
401
+ --src-prefix=a/ --dst-prefix=b/ "$base" HEAD -- . ':(exclude).ai') || exit 2
402
+ m=$(git log --no-show-signature --format=%B "$base..HEAD") || exit 2
403
+ a=$(printf '%s\n' "$d" |
404
+ LC_ALL=C awk '/^diff --git /{h=1; next}
405
+ h && /^\+\+\+ /{f=substr($0, 7); next}
406
+ /^@@/{h=0; next}
407
+ !h && /^\+/{print f ": " substr($0, 2)}') || exit 2
408
+ hits=0
409
+ for t in "$a" "$m"; do
410
+ printf '%s\n' "$t" | LC_ALL=C grep -E "$ids"
411
+ s=$?
412
+ [ "$s" -eq 0 ] && hits=1
413
+ [ "$s" -gt 1 ] && exit 2
414
+ done
415
+ exit "$hits"
416
+ ```
417
+
418
+ Exit `0` means no hit, exit `1` means at least one hit, each printed (a diff
419
+ hit prefixed by its file path, shown escaped and without its leading `"b`
420
+ for a path git quotes), and exit `2` means the run-base does not
421
+ resolve to a commit or a git, awk, or grep command failed. The check fails
422
+ closed: it reads the whole diff and log into memory and checks the status of
423
+ every stage, so a failure part way through (an unreadable object, or a text
424
+ tool rejecting a byte, for example) exits `2` instead of passing on partial
425
+ output; the text stages run byte-wise (`LC_ALL=C`) so no locale can make
426
+ them reject the input. It changes to the top level of the
427
+ repository first, so a `cwd` in a subdirectory scans the same range. The
428
+ diff options override the external diff, textconv, binary, rename, color,
429
+ and prefix settings of the user's git configuration and the repository's
430
+ attributes (`--text` diffs a file marked `-diff` or `binary` as text), and
431
+ the log option suppresses signature output, so those settings cannot hide an
432
+ added line from the scan or add lines to it.
433
+
434
+ It covers the lines added between the run-base and `HEAD` outside the
435
+ top-level `.ai/` directory, and the message of every commit reachable from
436
+ `HEAD` and not from the run-base. That range includes upstream work merged
437
+ into the branch after the run-base, whose added lines and commit messages
438
+ are scanned as well and can produce hits the branch did not write. It does
439
+ not cover uncommitted changes, removed lines, an identifier directly next to
440
+ a NUL byte (the shell drops NUL bytes from the captured diff), pull request
441
+ titles or bodies, branch names, or identifiers in any other format.
442
+ The patterns are case-sensitive and can match unrelated tokens, such as a
443
+ product or part name built the same way; because `--text` also diffs files
444
+ git detects as binary by content, an added image, font, or archive usually
445
+ produces hits made of its raw bytes, and a repository whose own
446
+ documentation discusses these formats (a copy of these templates, for
447
+ example) matches as well. A hit is a failure of the extra; when the
448
+ orchestrator confirms a hit is a false positive it records that decision,
449
+ and it may narrow the pathspec with further `':(exclude)<path>'` entries
450
+ when it approves the extra.
@@ -7,7 +7,8 @@ A subagent return is a misfire, not evidence, when its output does not parse
7
7
  against its role's output contract, including an implementer return that
8
8
  omits the `mutation_probes` field even though the task assignment named
9
9
  mutation probes to run, or that omits the `commits` field even though the
10
- task assignment asked for a commit. When a subagent returns near-instantly
10
+ task assignment asked for a commit, or that omits the `class_closure`
11
+ field on any round after the task's first. When a subagent returns near-instantly
11
12
  with no tool activity, treat that as a misfire signal rather than proof:
12
13
  check the output against the contract with extra suspicion, and accept it
13
14
  only if it is contract-valid and the assignment was answerable from the
@@ -44,7 +45,11 @@ cases. Ship the healthy half on its own verification, and refile the
44
45
  removed half as its own task carrying the measurement history that led to
45
46
  the split. Acceptance criteria that cannot be satisfied this way go to the
46
47
  operator as a merge-hold (hold the change unmerged and hand the decision to
47
- the operator).
48
+ the operator). Step 8 of the detailed workflow states the operational
49
+ halt: stop at the first `recurrence: repeated` finding whose class a
50
+ previous round's fix already addressed and whose `introduced_by_delta`
51
+ is `yes` or `unknown`, before any further implementer spawn on the
52
+ task, and record the split-or-redesign decision in `03-decisions.md`.
48
53
 
49
54
  ## Review-round escalation budget
50
55
 
@@ -65,6 +70,9 @@ split-or-redesign response, not instead of it.
65
70
  option is exhausted; under a `full` profile the choice falls to the
66
71
  advisor spawn or the merge-hold, under a `minimal` profile (no advisor
67
72
  subagent to spawn) it falls straight to the merge-hold.
73
+ In `single`, tier/model escalation requires a recorded switch to `delegated`.
74
+ Apply the Run mode switch procedure before assigning the next attempt to an
75
+ implementer raised according to this tier/model escalation option.
68
76
  - **Advisor spawn** (where the advisor is installed, `full` profile):
69
77
  send the advisor subagent the question "redesign, split, or hold?" and
70
78
  weigh its recommendation before deciding.
@@ -124,8 +132,8 @@ the next unpinned sentence do not converge. For such a change:
124
132
  - List the load-bearing claims of the normative site in the acceptance
125
133
  criterion, and pin each one as the whole sentence or clause that carries
126
134
  it. That list is the pin obligation. Every normative sentence the change
127
- adds or alters at that site is a claim; one left off the list is named
128
- in the criterion with the reason it is not load-bearing.
135
+ adds or alters at that site is pinned; one left unpinned is named in the
136
+ criterion with the reason it is not load-bearing.
129
137
  - Bound the reviewer's prose mutant space to that list in the briefing. A
130
138
  survivor outside the list is a scope note in the reviewer's
131
139
  `residual_risks`, not a finding, unless the reviewer shows that the
@@ -133,16 +141,16 @@ the next unpinned sentence do not converge. For such a change:
133
141
  - Bind each copy to the normative site through one shared test constant,
134
142
  and let a pointer point without restating the rule.
135
143
  - Cap test-adequacy review rounds on the change at two. A test-adequacy
136
- review round is one whose only unresolved findings are `tests` findings
137
- about pin gaps on the pinned prose; a round with any other unresolved
138
- finding is an ordinary round outside the cap. Pin gaps that remain
144
+ review round is one whose returned findings are all `tests` findings of
145
+ severity `low` or `medium` about pin gaps on the pinned prose; a round
146
+ returning any other finding is an ordinary round outside the cap. Pin gaps that remain
139
147
  become accepted notes or a follow-up.
140
148
 
141
149
  Semantic findings are exempt from the bound and from the cap: two sites
142
150
  stating different rules, a contradiction with another rule, and a false
143
151
  claim are defects at whatever severity they deserve. The cap changes
144
152
  neither the Round-2 halt rule, the Review-round escalation budget nor the
145
- Fix-regression decision point: a capped round still counts as a negative
153
+ Fix-regression decision point: a test-adequacy review round still counts as a negative
146
154
  round where it is one. The review gate is unchanged: a high or critical
147
155
  finding of any category still blocks and is never capped away, and
148
156
  accepting one follows the waiver rules. Anchored by an observed run; see
@@ -75,6 +75,7 @@ All state for one unit of work lives in a run directory:
75
75
  04-implementation-summary.md
76
76
  05-review-findings.md
77
77
  06-handoff.md
78
+ evidence/
78
79
  ```
79
80
 
80
81
  Create it at the start of a run by copying `.ai/workflow/templates/` and fill
@@ -82,6 +83,15 @@ the files as the run progresses. The newest run directory is the active one
82
83
  unless a `.ai/run` pointer names one (see below);
83
84
  older directories are the auditable history. Do not edit past runs.
84
85
 
86
+ `evidence/` is an optional subdirectory, not one of the seven templated
87
+ files: nothing copies or requires it. The orchestrator, the implementer, and
88
+ the reviewer write into it (test logs, probe verdicts, reproduction output,
89
+ reviewer-reproduced evidence) when a briefing or an acceptance criterion
90
+ asks for a saved artifact instead of just a report field; the reviewer's
91
+ write-boundary rule names it as the one write-allowed location outside the
92
+ reviewed tree, besides the writer's own scratchpad. The explorer and the
93
+ advisor are read-only roles and never write into it, or anywhere else.
94
+
85
95
  The run directory may live in the workspace's own `.ai/runs/` or in one
86
96
  repository's `.ai/runs/`. Either way, bind every repository or worktree the
87
97
  run touches to it with a pointer file, `<worktree-root>/.ai/run`:
@@ -38,6 +38,8 @@ acceptance_criteria:
38
38
  **Verification Set**
39
39
 
40
40
  <!-- Checked-in path, repository identity, and run-local frozen snapshot. The
41
+ digest recorded in the briefing is what carries the orchestrator's approval
42
+ of the resolved argv to the implementer and reviewer. The
41
43
  orchestrator approves effective config/scripts before any acquisition or
42
44
  execution; include an ordered docs/okf bundle check whenever that directory
43
45
  exists. -->
@@ -92,6 +92,21 @@ where it lives in the row's own cell.
92
92
  |---|---|---|---|---|---|---|---|---|---|---|---|
93
93
  | <!-- round --> | <!-- mutant --> | <!-- file --> | <!-- anchor --> | <!-- before --> | <!-- after --> | <!-- verified_applied_via --> | <!-- result --> | <!-- expectation --> | <!-- reason --> | <!-- restored_verified --> | <!-- replayed --> |
94
94
 
95
+ ### Class Closure
96
+
97
+ One row per fix round (any round after the task's first) that fixed a review
98
+ finding: the defect class it closed, the search command run to enumerate other
99
+ sites of that class (blank when `Closure Kind` is not `enumerated`), the sites
100
+ the search found (blank when `Closure Kind` is not `enumerated`), and the
101
+ closure kind (`enumerated | source`) the implementer reported in
102
+ `class_closure`; `not_applicable` never appears in this table, since a row
103
+ exists only for a round that fixed a finding, never for the task's first round.
104
+ An unclosed site of the row's class is named in Risks / Notes with the reason.
105
+
106
+ | Round | Class | Enumeration Command | Sites | Closure Kind |
107
+ |---|---|---|---|---|
108
+ | <!-- round --> | <!-- class --> | <!-- enumeration_command --> | <!-- sites --> | <!-- closure_kind --> |
109
+
95
110
  ### Optional Probe Plan and Result Index
96
111
 
97
112
  An optional runner-supported probe plan may be referenced here by relative
@@ -166,7 +166,7 @@ interface ExtractedYaml {
166
166
  * shorter run length from every offset inside the run, each retry
167
167
  * rescanning the lazy body: work quadratic in the run's length, which a
168
168
  * single pasted return of a few hundred backticks already turns into
169
- * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
169
+ * seconds (CHANGELOG 0.37.0 names the measurement). Keeping the
170
170
  * run whole also makes the "closing run at least as long as the opening
171
171
  * one" rule above literal: an opener longer than any closing run in the
172
172
  * input is no fence at all, where splitting the run instead matched it
@@ -443,7 +443,7 @@ export const ARRAY_ELEMENT_EXPECTED = KIND_EXPECTED.string;
443
443
  * shorter run length from every offset inside the run, each retry
444
444
  * rescanning the lazy body: work quadratic in the run's length, which a
445
445
  * single pasted return of a few hundred backticks already turns into
446
- * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
446
+ * seconds (CHANGELOG 0.37.0 names the measurement). Keeping the
447
447
  * run whole also makes the "closing run at least as long as the opening
448
448
  * one" rule above literal: an opener longer than any closing run in the
449
449
  * input is no fence at all, where splitting the run instead matched it
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.38.1",
3
+ "version": "0.40.0",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",