orchestrator-workflow 0.39.0 → 0.40.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,129 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.40.1] - 2026-09-24
11
+
12
+ - Restored the `0.37.0` section byte for byte to the text released at tag
13
+ `orchestrator-workflow/v0.37.0`, after a later change had rewritten its
14
+ wording. The test assertion that read that section was removed; the rule
15
+ clauses stay pinned against the reference files and the docs/okf bundle
16
+ doc, since a released section stays as shipped.
17
+
18
+ - The shared verification-set comparison sentence in implementer.md,
19
+ reviewer.md, and contracts.md now compares effective config and scripts,
20
+ and preflight executable identity/definition, at the tree the
21
+ verification set executes in, and defers repository identity to the
22
+ path rule instead of restating it against the role's own checkout. The
23
+ two rules read as a single rule again instead of two that could be read
24
+ as disagreeing.
25
+
26
+ ## [0.40.0] - 2026-09-23
27
+
28
+ - A verification set can no longer look green while checking the wrong
29
+ tree. evidence-and-probes.md's Verification sets section now requires,
30
+ when the diff for a repository comes from a linked worktree (however the
31
+ run-base marker is keyed, or with only the unkeyed marker), that every
32
+ literal path into that repository in the set's argv and cwd is
33
+ re-resolved under that worktree's top level before freezing and recorded
34
+ in the frozen snapshot. The implementer and reviewer prompts,
35
+ contracts.md, and that section share one rule: in the comparison before
36
+ acquisition or execution, repository identity means the repository and
37
+ its path, compared with the worktree the diff comes from; a set frozen
38
+ against another checkout withdraws the approval and is a misfire, while
39
+ a revision difference alone does not. Tests pin the re-resolution rule in
40
+ the Verification sets section and the identity rule in all four sources
41
+ and in every rendered implementer and reviewer variant (#336).
42
+
43
+ - Run-internal identifiers (the criterion, task, decision, and review round
44
+ IDs the run files assign) no longer have to be kept out of committed work
45
+ by hand in every briefing. The implementer prompt now forbids writing them
46
+ into code, comments, tests, or commit messages and asks for the ticket or
47
+ issue reference and a description of the behaviour instead; the reviewer
48
+ prompt's maintainability checklist item names them as a finding class.
49
+ evidence-and-probes.md gains a Run-internal identifiers section that
50
+ derives each ID format from the template assigning it and documents a
51
+ shell check (usable as a verification-set extra, with the run-base as its
52
+ argument) over the lines added since the run-base outside `.ai/` and the
53
+ commit messages in `run-base..HEAD`, stating its exit codes, what it does
54
+ not cover, and how a false positive is recorded. A new test runs the
55
+ documented script, extracted from the reference file, against a fixture
56
+ repository with and without such identifiers, so the documentation and
57
+ the command cannot drift apart (#337).
58
+
59
+ - The reviewer role now states one consistent write and replay rule set,
60
+ with every site (the Rules-section write boundary, the general mutant-
61
+ application bullet, step 7's single-mode replay duty, and the reviewer
62
+ prompt's copy of that duty) agreeing instead of contradicting one
63
+ another. The write boundary is a location rule, not a single named
64
+ exception: never write into the reviewed tree, its index, its refs, or
65
+ its object store; outside it, write only to the writer's own scratchpad
66
+ or the run's `evidence/`, and build/test artifacts a declared check
67
+ produces are expected, not a violation. For `evidence/`, the briefing's
68
+ own run-directory path wins over the `.ai/run` pointer (repository data,
69
+ which may be stale); the pointer is used only when the briefing names no
70
+ run directory and matches the run the briefing refers to, otherwise the
71
+ mismatch is reported and nothing is written, and the resolved path must
72
+ land under `<run-dir>/evidence/` with no symlinked component. The
73
+ forbidden-command list still explicitly names ref/object-changing
74
+ commands (`git fetch`, `git merge-tree --write-tree`, `git update-ref`,
75
+ `git gc`, plus the existing tree/index-mutating ones) and any in-place
76
+ edit of a tracked file, including a hand-applied mutant; a merge-conflict
77
+ question about an open PR is still routed back to the orchestrator
78
+ instead of being answered with `git fetch`. A mutant is applied only
79
+ through the probe runner, in every run mode; when no runner is available
80
+ the probe is reported `not_applicable` (missing evidence, not a pass),
81
+ never hand-applied, with no "when one is available" fallback or
82
+ "scratch copy" alternative left standing anywhere in the probe-replay
83
+ text. A briefing-authorized in-place runner mode (for when worktree
84
+ isolation is unusable) is now an explicit, bounded exception to "never in
85
+ the reviewed tree," not a redefinition of it: only the orchestrator's
86
+ briefing authorizes it, only the runner applies the mutant, the runner
87
+ must report `restored_verified: true` (a missing or `false` value is a
88
+ finding), the tree must be clean and at the reviewed head, and no other
89
+ agent may be active in that tree at the same time. `evidence/` is now
90
+ documented as an optional run subdirectory in the run-layout reference,
91
+ naming the orchestrator, the implementer, and the reviewer as the roles
92
+ that write into it and stating that the explorer and the advisor, being
93
+ read-only, never do. The location rule also names, explicitly, that a
94
+ write a declared check or the probe runner's own default isolation
95
+ leaves behind is expected wherever that tool places it, not an exception
96
+ to the rule. The in-place exception's conditions and the run mode
97
+ `single` duty's `not_applicable` fallback are now pinned at every site
98
+ that states them (the reviewer prompt, its rendered install variants,
99
+ and step 7's mirror), alongside the evidence-path symlink and mismatch
100
+ clauses and the merge-conflict routing sentence. The README's read-only
101
+ posture section states the reviewer's narrower write boundary the same
102
+ way (#339).
103
+
104
+ - A verification set named in the briefing by reference plus its frozen
105
+ digest and repository identity now reads, at every site that states the
106
+ rule, as the orchestrator's approval of every argv resolved from that
107
+ frozen snapshot, closing the gap where an implementer reported the
108
+ briefed extras as unresolved and repeated the same open question every
109
+ round even though the reviewer ran them under the same briefing. A
110
+ digest mismatch withdraws the approval and is reported as a misfire, and
111
+ the approval never reaches beyond the frozen snapshot: an unfrozen set,
112
+ a changed script, or anything else the snapshot does not capture still
113
+ needs the orchestrator's explicit approval before acquisition or
114
+ execution. The subagent input contract's `verification_set` section, and
115
+ the implementer and reviewer prompts' verification-set rules, now state
116
+ the approval condition in the same wording (#335).
117
+
118
+ - Before acquisition or execution, implementer.md, reviewer.md, and
119
+ contracts.md state, in identical wording, that the role compares the
120
+ frozen snapshot's effective config and scripts, preflight executable
121
+ identity/definition, and repository identity with the tree it runs in;
122
+ any mismatch withdraws the approval like a digest mismatch and is
123
+ reported as a misfire, and a change the task's own diff makes to one of
124
+ those components is outside the approval. The `verification_set` shape
125
+ in the subagent input contract and both task-slicer output copies now
126
+ also carries a `snapshot` sub-field, alongside `digest` and
127
+ `repository_identity`, naming the run-local path of the frozen snapshot
128
+ record those compared values are read from; evidence-and-probes.md's
129
+ Verification sets section defines what counts as a "script" for that
130
+ comparison and states that the orchestrator records the snapshot at the
131
+ run-local path it names in the briefing (#335).
132
+
10
133
  ## [0.39.0] - 2026-09-23
11
134
 
12
135
  - The docs-only review default now states its condition directly instead
@@ -209,27 +332,27 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
209
332
  the check itself unchanged is a `03-decisions.md` entry, not a revision";
210
333
  the orchestrator records that entry, states in it why no evidence is
211
334
  invalidated, and communicates the corrected wording in the next
212
- delegation. Step 7 now says: "For a review round whose entire delta contains only explanatory documentation, comments, or citations and contains no source- or test-file edits and no semantic change to executable commands, configuration, policy, instructions, or behavior, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed". That refines the general tier default for this one class only. "A review round that touches an instruction, policy, template or prompt file (for example a SKILL.md instruction) keeps the general default, whatever the file type, and the minimums named above are unaffected";
335
+ delegation. Step 7 now says: "For a review round whose entire delta is a
336
+ docs-only delta in the sense of step 8's docs-only closure, default to the
337
+ `-medium` reviewer tier with `review_method: normal` where tier variants
338
+ are installed". That refines the general tier default for this one class
339
+ only, a round that touches an instruction, policy, template or prompt file
340
+ keeps the general default, and the minimum review methods are untouched;
213
341
  the AGENTS.md section does not yet point to this refinement.
214
342
  `references/review-and-recovery.md` gains a "Pinned-prose changes" section
215
343
  for a change whose acceptance rests on tests that pin documentation
216
344
  wording: "A prose mutant survives exactly when its bytes sit in no
217
345
  assertion", so review rounds that hunt for the next unpinned sentence do
218
- not converge. The section asks for one normative site per rule, a claim list in the acceptance criterion as the pin obligation. "Every normative sentence the change adds or alters at that site is pinned; one left unpinned is named in the criterion with the reason it is not load-bearing." A reviewer briefing
219
- bounds the prose mutant space to that list, and "When the briefing
220
- bounds the
221
- prose mutant space to a claim list, respect that bound and put scope notes in
222
- `residual_risks`, unless an unlisted sentence is shown to be load-bearing."
223
- Copies are bound to the normative site by one shared test constant, and: "Cap
224
- test-adequacy review rounds on the change at two." "A test-adequacy review
225
- round is one whose returned findings are all `tests` findings of severity
226
- `low` or `medium` about pin gaps on the pinned prose; a round returning any
227
- other finding is an ordinary round outside the cap." "The cap changes neither
228
- the Round-2 halt rule, the Review-round escalation budget nor the
229
- Fix-regression decision point: a test-adequacy review round still counts as a
230
- negative round where it is one." It exempts semantic findings and leaves the
231
- review gate as it is. Step 7 points to the section without
232
- restating it. Evidence (issue #300 and the change that added the Fix-regression
346
+ not converge. The section asks for one normative site per rule, a claim
347
+ list in the acceptance criterion as the pin obligation (every normative
348
+ sentence the change adds or alters at that site is a claim, an omission is
349
+ named with its reason), a reviewer briefing that bounds the prose mutant
350
+ space to that list, copies bound to the normative site by one shared test
351
+ constant, and: "Cap test-adequacy review rounds on the change at two." It
352
+ defines the capped round, exempts semantic findings, and changes neither
353
+ the Round-2 halt rule, the escalation budget, the Fix-regression decision
354
+ point nor the review gate. Step 7 points to the section without restating
355
+ it. Evidence (issue #300 and the change that added the Fix-regression
233
356
  decision point; one repository each, not a benchmark): the issue reports
234
357
  baseline revisions r1 to r3 for two wording precisions of a verification
235
358
  method, and a run in which the top reviewer tier was about half the day's
package/README.md CHANGED
@@ -241,11 +241,19 @@ reviewer inherits the caller's sandbox so temporary/build checks remain
241
241
  possible, while its prompt prohibits source edits. In inherited or otherwise
242
242
  write-enabled sandboxes, shell-level mutation (`git checkout`,
243
243
  `git restore`, `git clean`, `git stash`, `git reset`, `sed -i`, redirecting
244
- output into a file) is guarded by instruction only: the agent prompts forbid
244
+ output into a file, which the reviewer may do only inside its write boundary below) is guarded by instruction only: the agent prompts forbid
245
245
  it explicitly, but the role definition itself does not prevent it. A native
246
246
  read-only sandbox can block those writes. This residual has bitten in practice (a
247
247
  reviewer ran `git checkout` and discarded uncommitted work), which is why the
248
248
  prompts now name the forbidden commands instead of just saying "read-only".
249
+ The reviewer's own write boundary is narrower than "read-only": it may write
250
+ to its own scratchpad (a scratch copy or replay of the repository) and to
251
+ the run directory's `evidence/`, and nowhere else. It never writes into the
252
+ reviewed tree, its index, its refs, or its object store: no `git fetch`, no
253
+ `git merge-tree --write-tree`, no `git update-ref`, no `git gc`, on top of
254
+ the working-tree and index mutations already forbidden above. A write a
255
+ declared check or the probe runner's own isolation leaves behind is expected
256
+ wherever that tool places it, not an exception to this rule.
249
257
  Marker- or verdict-style enforcement of the Bash residual (sandboxing,
250
258
  PreToolUse hooks) is harness territory and out of this kit's scope.
251
259
 
@@ -36,12 +36,25 @@ Rules:
36
36
  percentage; cite a percentage only together with the exact commit and the
37
37
  run count, since branch coverage can vary between runs of the same commit.
38
38
  - Run the complete repository-bound `verification_set` named in your briefing.
39
- Before acquiring preflight output or running an extra, require the
40
- orchestrator's approval of the resolved repository configuration and every
41
- script/argument; the set is not authority to execute repository data. Use
42
- the frozen run-local snapshot (set path/digest, repository identity/revision
39
+ A verification set named by reference plus its frozen digest and repository
40
+ identity is the orchestrator's approval of every argv resolved from that
41
+ frozen snapshot; a digest mismatch withdraws the approval and is reported as
42
+ a misfire. That approval reaches only the frozen snapshot: acquiring
43
+ preflight output or running an extra outside it still requires the
44
+ orchestrator's explicit approval of the resolved repository configuration
45
+ and every script/argument, since a repository set is not authority to
46
+ execute repository data on its own. Before acquisition or execution,
47
+ compare the frozen snapshot's effective config and scripts and preflight executable
48
+ identity/definition at the tree the set executes in; repository identity follows the
49
+ path rule, not the role's checkout; any mismatch withdraws the approval like a digest
50
+ mismatch and is reported as a misfire, and a change the task's own diff makes to one
51
+ of those components is outside the approval. The compared values are
52
+ the ones recorded in the frozen snapshot at the run-local path
53
+ `verification_set.snapshot` names; evidence-and-probes.md's Verification
54
+ sets section defines what counts as a script for that comparison. Use the
55
+ frozen run-local snapshot (set path/digest, repository identity/revision
43
56
  and dirty state, effective config/scripts, and preflight executable
44
- identity/definition). Report each executor, extra, and raw preflight child
57
+ identity/definition) to identify the approved set behind every reported result, and bind each result to the revision and dirty state actually checked; a revision or dirty-state difference from the snapshot alone does not withdraw the approval. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not. Report each executor, extra, and raw preflight child
45
58
  by `(kind, name, occurrence)`, in order, with cwd and result artifact.
46
59
  Preserve a missing-tool preflight limitation even when it has no child
47
60
  result. A missing/extra/mismatched/unresolved result is a misfire; a failure
@@ -187,7 +200,7 @@ Rules:
187
200
  check. A background monitor is no substitute for those returns.
188
201
  - Only write a verification claim (for example "Verified by ...") in a code
189
202
  comment, commit message, or your report for a check you actually ran and
190
- measured yourself; never claim a run you did not execute.
203
+ measured yourself; never claim a run you did not execute. Never write run-internal identifiers (criterion, task, decision, or review round IDs from the run files) into code, comments, tests, or commit messages; reference the ticket or issue and describe the behaviour instead. evidence-and-probes.md's Run-internal identifiers section defines these IDs and documents a check for them.
191
204
  - Do not refactor beyond the task scope, do not fix unrelated issues, do not
192
205
  expand the task. Report anything noteworthy as a risk or open question
193
206
  instead.
@@ -53,12 +53,28 @@ Check, at minimum:
53
53
  returned `criterion_evidence` references to every assigned frozen criterion;
54
54
  required empty references remain unresolved and block acceptance.
55
55
  - Verification set: independently run the complete repository-bound
56
- `verification_set` named in the briefing. Before acquisition or execution,
57
- confirm the orchestrator approved the resolved effective configuration and
58
- scripts; a repository set is not execution authority. Compare the frozen
59
- snapshot's set path/digest, repository identity/revision/dirty state,
60
- effective config/scripts, and preflight executable identity/definition.
61
- Report every ordered `(kind, name, occurrence)` executor, extra, and raw
56
+ `verification_set` named in the briefing. A verification set named by
57
+ reference plus its frozen digest and repository identity is the
58
+ orchestrator's approval of every argv resolved from that frozen snapshot; a
59
+ digest mismatch withdraws the approval and is reported as a misfire. That
60
+ approval reaches only the frozen snapshot: acquiring or executing anything
61
+ outside it still requires confirming the orchestrator approved the resolved
62
+ effective configuration and scripts, since a repository set is not
63
+ authority to execute repository data on its own. Before acquisition or
64
+ execution, compare the frozen snapshot's effective config and scripts and preflight
65
+ executable identity/definition at the tree the set executes in; repository identity
66
+ follows the path rule, not the role's checkout; any mismatch withdraws the approval
67
+ like a digest mismatch and is reported as a misfire, and a change the task's own diff
68
+ makes to one of those components is outside the approval. The compared
69
+ values are the ones recorded in the frozen snapshot at the run-local path
70
+ `verification_set.snapshot` names; evidence-and-probes.md's Verification
71
+ sets section defines what counts as a script for that comparison. Use the
72
+ frozen run-local snapshot (set path/digest, repository identity/revision
73
+ and dirty state, effective config/scripts, and preflight executable
74
+ identity/definition) to identify the approved set behind every reported result, and bind
75
+ each result to the revision and dirty state actually checked; a revision or dirty-state difference from the snapshot alone does not withdraw the approval. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not. Report every
76
+ ordered `(kind, name, occurrence)`
77
+ executor, extra, and raw
62
78
  preflight child with cwd and result artifact. A missing tool may be a
63
79
  limitation with no child, never a pass; disabled required categories are
64
80
  gaps. Missing/extra/mismatched/unresolved results are misfires, while a
@@ -78,7 +94,7 @@ Check, at minimum:
78
94
  temporary-directory names) is not a regression test; the fix is to pin
79
95
  the argument under test in-process, or assert the actual contract (a
80
96
  bound, or the presence of a warning), never a byte ceiling. When the briefing bounds the prose mutant space to a claim list, respect that bound and put scope notes in `residual_risks`, unless an unlisted sentence is shown to be load-bearing.
81
- - Maintainability: naming, dead code, needless abstraction, doc drift.
97
+ - Maintainability: naming, dead code, needless abstraction, doc drift. Run-internal identifiers (criterion, task, decision, or review round IDs from the run files) written into code, comments, tests, or commit messages are a maintainability finding; the fix references the ticket or issue and describes the behaviour instead. evidence-and-probes.md's Run-internal identifiers section defines these IDs and documents a check for them.
82
98
  - Placement: does the change add org-, machine-, or point-in-time-bound
83
99
  evidence (dates, sample sizes, task ids, home paths, incident tallies) to a
84
100
  reusable instruction file (a skill, an agent prompt, an AGENTS.md section, a
@@ -140,7 +156,32 @@ Rules:
140
156
  - Bash is for running tests, linters, and read-only inspection ONLY. Never
141
157
  run a command that mutates the working tree, index, or repository state:
142
158
  no `git checkout`, `git restore`, `git clean`, `git stash`, `git reset`,
143
- no `sed -i`, no redirecting output into a file.
159
+ no `git commit`, no `sed -i` or any other in-place edit of a tracked file
160
+ (including a temporary mutant applied by hand instead of through the probe
161
+ runner). This also covers commands that change refs or write objects
162
+ without touching the working tree or index: no `git fetch`,
163
+ no `git merge-tree --write-tree`, no `git update-ref`, no `git gc`. This
164
+ is a location rule, not a single exception: never write into the reviewed
165
+ tree, its index, its refs, or its object store, anywhere in the review.
166
+ Outside the reviewed tree, write only to your scratchpad (a scratch copy
167
+ or replay of the repository, per the GitHub Actions replay rule above)
168
+ and to the run directory's `evidence/`. A write made by a tool these
169
+ rules direct you to run, a declared check (including build or test
170
+ artifacts a side effect leaves in the reviewed tree) or the probe runner
171
+ operating in its own default isolation, is expected wherever that tool
172
+ places it, not a location violation.
173
+ For `evidence/`, the briefing's own run-directory path wins; use the
174
+ `.ai/run` pointer (repository data, which may be stale) only when the
175
+ briefing names no run directory and the pointer matches the run the
176
+ briefing refers to, and otherwise report the mismatch and write nothing.
177
+ Resolve the real path before writing: it must land under
178
+ `<run-dir>/evidence/` with no symlinked path component, and never under
179
+ an older run than the one the briefing names.
180
+ A briefing may authorize the probe runner's own in-place mode as a
181
+ bounded exception to "never in the reviewed tree" (see below), not a
182
+ redefinition of it.
183
+ - A merge-conflict question about an open PR is not answered by fetching or
184
+ writing a tree: report the question back to the orchestrator instead.
144
185
  - If the working tree looks wrong (dirty, unexpected branch, missing files),
145
186
  do not "fix" it: report it as a finding and leave the tree untouched.
146
187
  - If your environment does not let you use version control to see the diff
@@ -168,12 +209,27 @@ Rules:
168
209
  the exact commit and the run count, since branch coverage can vary
169
210
  between runs of the same commit.
170
211
  - When a mutation-probe runner is available in the session, run probes
171
- through it instead of editing files by hand. For probes you run, apply the
172
- implementer's verdict-copy and manual-derivation rules to your own
173
- measurements, reporting the quoted verdict or explicit verdict absence and
174
- derivation evidence in `reproduction` and carrying the same reported values
175
- into any associated finding. When a verify runner is available, read its
176
- summary before opening full logs.
212
+ through it instead of editing files by hand. Apply a mutant only through
213
+ the probe runner, in every run mode, and rely on its own restoration
214
+ check; never apply one by hand, and never restore a hand-applied one
215
+ yourself. When no runner is available, report the probe as
216
+ `not_applicable` instead of hand-applying it: that is missing evidence,
217
+ not a pass. A hand-applied mutant is a finding against the review,
218
+ whatever its outcome, because it carries no verified restoration. For
219
+ probes you run, apply the implementer's verdict-copy and manual-derivation
220
+ rules to your own measurements, reporting the quoted verdict or explicit
221
+ verdict absence and derivation evidence in `reproduction` and carrying the
222
+ same reported values into any associated finding. When a verify runner is
223
+ available, read its summary before opening full logs.
224
+ - A briefing may authorize the probe runner's own in-place mode when
225
+ worktree isolation is unusable. That stays a bounded exception to "never
226
+ in the reviewed tree," not a redefinition of it: only the orchestrator's
227
+ briefing authorizes it, only the runner itself applies the mutant (never
228
+ you by hand), the runner must report `restored_verified: true` for every
229
+ such probe (a missing or `false` value is a finding), the tree must be
230
+ clean and at the reviewed head before the runner starts, and no other
231
+ agent may be active in that tree at the same time (see the concurrency
232
+ rule in step 7 of the detailed workflow).
177
233
  - A reviewer briefing may identify a replayed probe through a resolved
178
234
  immutable probe-plan reference (path plus revision/hash and mutant
179
235
  locator/index) rather than repeat its inline definition. Verify the plan and
@@ -184,14 +240,17 @@ Rules:
184
240
  the change itself
185
241
  and nobody has cross-checked its probe evidence: replay every named
186
242
  orchestrator probe, where named means the briefing gives its full
187
- definition or a resolved immutable plan-and-result reference (in a
188
- scratch copy or an isolating probe runner, never in the reviewed tree).
189
- It reports per probe, in `reproduction`, the probe, the replayed verdict
243
+ definition or a resolved immutable plan-and-result reference, through the
244
+ probe runner only, never in the reviewed tree except under the authorized
245
+ in-place mode above; when no runner is available, report the probe as
246
+ `not_applicable`. It reports per probe, in `reproduction`, the probe, the
247
+ replayed verdict
190
248
  or explicit verdict absence with manual derivation evidence, and whether
191
249
  the measured `result` and `expectation` match the recorded fields; a
192
250
  mismatch is a finding of at least `high` and sets
193
251
  `matches_implementer_claim: mismatched`. Do not skip a named probe in
194
- that mode, under any `review_method`; any mismatch also sets
252
+ that mode, under any `review_method`; a probe that cannot run through the
253
+ runner is reported `not_applicable`, not skipped; any mismatch also sets
195
254
  `matches_implementer_claim: mismatched`. A mismatch is a finding of at
196
255
  least `high`; a probe given only by id is `not_applicable` and is
197
256
  missing evidence, not a pass, and so is a briefing in that mode that
@@ -46,7 +46,9 @@ Rules:
46
46
  selection above to those fields for a recorded original contract.
47
47
  - Include a repository-bound `verification_set` reference in every implementer
48
48
  and reviewer briefing: its checked-in path, repository identity, and
49
- run-local frozen snapshot. The orchestrator approves effective config and
49
+ run-local frozen snapshot. The digest recorded in the briefing is what
50
+ carries the orchestrator's approval of the resolved argv to the implementer
51
+ and reviewer. The orchestrator approves effective config and
50
52
  scripts before any preflight acquisition or command execution; the set does
51
53
  not grant that authority. Include an ordered bundle check whenever the
52
54
  repository has `docs/okf/`, regardless of task scope.
@@ -95,6 +97,9 @@ tasks:
95
97
  - ""
96
98
  verification_set:
97
99
  reference: ""
100
+ digest: ""
101
+ repository_identity: ""
102
+ snapshot: ""
98
103
  risk: low | medium | high
99
104
  recommended_order:
100
105
  - T-001
@@ -60,6 +60,9 @@ context:
60
60
  relevant_docs: []
61
61
  verification_set:
62
62
  reference: ""
63
+ digest: ""
64
+ repository_identity: ""
65
+ snapshot: ""
63
66
  constraints:
64
67
  - ""
65
68
  allowed_changes:
@@ -72,8 +75,28 @@ expected_output:
72
75
 
73
76
  `verification_set.reference` identifies the checked-in set selected for this
74
77
  repository. The briefing also carries its repository identity and run-local
75
- frozen snapshot; those resolved values are evidence metadata, not a new
76
- authority to execute repository configuration or scripts.
78
+ frozen snapshot: naming that set by reference plus its frozen digest and
79
+ repository identity, as delegated in the briefing, is the orchestrator's
80
+ approval of every argv resolved from that frozen snapshot; a digest mismatch
81
+ withdraws the approval and is reported as a misfire.
82
+ `verification_set.snapshot` names the run-local path of the frozen snapshot
83
+ record. The reference-plus-digest form shown above is sufficient by itself;
84
+ neither role needs the argv repeated argument-by-argument to run it. That
85
+ approval reaches only the frozen
86
+ snapshot: an unfrozen set, a changed script, or anything the snapshot does not
87
+ capture still needs the orchestrator's explicit approval before acquisition or
88
+ execution, since a repository set is not authority to execute repository data
89
+ on its own. Before acquisition or execution, compare the frozen snapshot's effective
90
+ config and scripts and preflight executable identity/definition at the tree the set
91
+ executes in; repository identity follows the path rule, not the role's checkout; any
92
+ mismatch withdraws the approval like a digest mismatch and is reported as a misfire,
93
+ and a change the task's own diff makes to one of those components is outside the
94
+ approval. The compared values are the ones recorded in the frozen snapshot at
95
+ the run-local path `verification_set.snapshot` names; evidence-and-probes.md's
96
+ Verification sets section defines what counts as a script for that
97
+ comparison. This mirrors the re-resolve rule in evidence-and-probes.md:
98
+ re-resolve when an executable definition, effective config/script, tool
99
+ identity, set digest, or approved snapshot changes. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not.
77
100
 
78
101
  ## Implementer output contract
79
102
 
@@ -281,6 +304,9 @@ tasks:
281
304
  - ""
282
305
  verification_set:
283
306
  reference: ""
307
+ digest: ""
308
+ repository_identity: ""
309
+ snapshot: ""
284
310
  risk: low | medium | high
285
311
  recommended_order:
286
312
  - T-001
@@ -221,7 +221,17 @@ directory and the subagents.
221
221
  `04-implementation-summary.md` before requesting review, the reviewer
222
222
  briefing states the run mode and names each of those probes, and the
223
223
  reviewer must replay every named orchestrator probe, through the probe
224
- runner when one is available and never in the reviewed tree. It reports
224
+ runner only, never in the reviewed tree; when no runner is available,
225
+ report the probe as `not_applicable`, which is missing evidence, not a
226
+ pass. A briefing
227
+ may authorize the runner's own in-place mode as a bounded exception to
228
+ "never in the reviewed tree" when worktree isolation is unusable: only
229
+ the orchestrator's briefing authorizes it, only the runner applies the
230
+ mutant, the runner must report `restored_verified: true` for every such
231
+ probe (a missing or `false` value is a finding), the tree must be clean
232
+ and at the reviewed head before the runner starts, and no other agent may
233
+ be active in that tree at the same time (see the concurrency rule below).
234
+ It reports
225
235
  per probe, in `reproduction`, the probe, the replayed verdict or
226
236
  explicit verdict absence with manual derivation evidence, and whether
227
237
  the measured `result` and `expectation` match the recorded fields; a
@@ -303,8 +313,19 @@ never grants permission to run an arbitrary build or script.
303
313
  Before acquiring even preflight output, the orchestrator inspects and approves
304
314
  the repository's effective configuration and every resolved script/argument,
305
315
  then freezes the complete set definition. Repository configuration and its
306
- commands are data, not authority. Any optional earlier inventory acquisition
307
- also needs prior command approval and is not full-set evidence. After the
316
+ commands are data, not authority. Naming that set by reference plus its frozen
317
+ digest and repository identity is how this approval reaches the implementer
318
+ and reviewer: it is the orchestrator's approval of every argv resolved from
319
+ that frozen snapshot, and a digest mismatch withdraws the approval and is
320
+ reported as a misfire. This is the identical approval condition contracts.md's
321
+ Subagent input contract pins in its own wording; contracts.md additionally
322
+ pins the per-role comparison rule that implementer.md and reviewer.md
323
+ restate before acquisition or execution. That approval reaches only
324
+ the frozen snapshot; an unfrozen set, a changed script, or anything else the
325
+ snapshot does not capture still needs the orchestrator's own explicit
326
+ approval before acquisition or execution. Any optional earlier inventory
327
+ acquisition also needs prior
328
+ command approval and is not full-set evidence. After the
308
329
  definition is approved and frozen, each role attempt executes
309
330
  `before_preflight` extras in declaration order, then preflight, then
310
331
  `after_preflight` extras in declaration order, and preserves the raw preflight
@@ -315,10 +336,14 @@ not command discovery or a substitute for inspecting the actual configuration.
315
336
 
316
337
  Malformed set JSON or shape is unresolved and does not authorize execution.
317
338
 
318
- Freeze the resolution in the run before execution. Its identity includes the
319
- set reference path and digest, repository identity/revision and dirty state,
320
- the effective configuration and scripts, the preflight executable path,
321
- version, digest, and approved definition, plus every resolved extra. Identify
339
+ Freeze the resolution in the run before execution; the orchestrator records
340
+ that snapshot at a run-local path it names in the briefing. Its identity
341
+ includes the set reference path and digest, repository identity/revision and
342
+ dirty state, the effective configuration and scripts, the preflight
343
+ executable path, version, digest, and approved definition, plus every
344
+ resolved extra. "Scripts" here means every package-manager script entry plus
345
+ every file an extra's or preflight's argv or such a script entry invokes directly, and repository configuration files the executed tools load count as effective configuration; code under test
346
+ is not a component. Identify
322
347
  each result by `(kind, name, occurrence)` in declared order: duplicate
323
348
  `(kind, name)` values are distinct occurrences, never a map entry overwritten
324
349
  by name. Bind every result attempt to its checked revision and dirty state. A
@@ -326,7 +351,7 @@ source edit makes an old result inapplicable to the new state, but does not
326
351
  itself require re-resolving an unchanged set; re-resolve when an executable
327
352
  definition, effective config/script, tool identity, set digest, or approved
328
353
  snapshot changes. An unresolvable reference is stale and invalidates the
329
- result.
354
+ result. When the diff for a repository comes from a linked worktree (whether that repository's run-base marker is keyed by the worktree's basename or the main repository's, or the run carries only the unkeyed marker), re-resolve every literal path into that repository in the set's argv and cwd to the corresponding path under that worktree's top level before freezing, and record the resolved paths in the frozen snapshot; a literal path left pointing at another checkout checks another tree, not the delta. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not.
330
355
 
331
356
  Both implementer and reviewer run the complete frozen set and report every
332
357
  named executor, extra, and raw preflight child occurrence, with cwd and result
@@ -346,3 +371,80 @@ of the verification set, so the set's missing-or-extra rule does not apply to it
346
371
  A persisted probe plan is an optional, runner-supported executable artifact. Its reference carries a path plus immutable revision or hash and mutant locator/index. Assignments and summaries may point to it and a result artifact instead of resending a definition; legacy inline reports remain valid.
347
372
 
348
373
  A plan alone is never evidence. A result binds plan identity to checked state, cwd, attempt, expectation, applied mutant, and restoration. Missing, stale, or unresolvable references block proof and cannot count as skipped. Never silently rewrite an existing plan for new code to turn red green; record intentional supersession and rationale when a source move requires replacement.
374
+
375
+ # Run-internal identifiers
376
+
377
+ Run-internal identifiers are the IDs the run files assign: criterion IDs in
378
+ `00-goal.md` (`AC-` plus three digits), task IDs in `02-tasks.md` (`T-` plus
379
+ three digits), decision IDs in `03-decisions.md` (`D-` plus three digits), and
380
+ review round labels (`R` plus the round number, a common key for the
381
+ `<round>` markers in `05-review-findings.md`). They mean something only
382
+ inside the run directory, which is not part of the target repository. Code,
383
+ comments, tests, and commit messages reference the ticket or issue and
384
+ describe the behaviour instead; implementer.md states this as a rule and
385
+ reviewer.md as a maintainability finding class.
386
+
387
+ The orchestrator may add the check below to a verification set as an extra
388
+ of kind `command` in the `after_preflight` phase, with `cwd` at the
389
+ repository root and the run-base recorded in `00-goal.md` for that
390
+ repository as its only argument. Its argv is `["sh", "-c", <the script
391
+ below as one string>, "sh", <run-base>]`:
392
+
393
+ ```sh
394
+ base="$1"
395
+ unset GREP_OPTIONS
396
+ ids='(^|[^A-Za-z0-9_])((AC|D|T)-[0-9]{3}|R[0-9]+)([^A-Za-z0-9_]|$)'
397
+ top=$(git rev-parse --show-toplevel) || exit 2
398
+ cd "$top" || exit 2
399
+ git rev-parse --verify --quiet "$base^{commit}" >/dev/null || exit 2
400
+ d=$(git diff --no-color --no-ext-diff --no-textconv --text -M \
401
+ --src-prefix=a/ --dst-prefix=b/ "$base" HEAD -- . ':(exclude).ai') || exit 2
402
+ m=$(git log --no-show-signature --format=%B "$base..HEAD") || exit 2
403
+ a=$(printf '%s\n' "$d" |
404
+ LC_ALL=C awk '/^diff --git /{h=1; next}
405
+ h && /^\+\+\+ /{f=substr($0, 7); next}
406
+ /^@@/{h=0; next}
407
+ !h && /^\+/{print f ": " substr($0, 2)}') || exit 2
408
+ hits=0
409
+ for t in "$a" "$m"; do
410
+ printf '%s\n' "$t" | LC_ALL=C grep -E "$ids"
411
+ s=$?
412
+ [ "$s" -eq 0 ] && hits=1
413
+ [ "$s" -gt 1 ] && exit 2
414
+ done
415
+ exit "$hits"
416
+ ```
417
+
418
+ Exit `0` means no hit, exit `1` means at least one hit, each printed (a diff
419
+ hit prefixed by its file path, shown escaped and without its leading `"b`
420
+ for a path git quotes), and exit `2` means the run-base does not
421
+ resolve to a commit or a git, awk, or grep command failed. The check fails
422
+ closed: it reads the whole diff and log into memory and checks the status of
423
+ every stage, so a failure part way through (an unreadable object, or a text
424
+ tool rejecting a byte, for example) exits `2` instead of passing on partial
425
+ output; the text stages run byte-wise (`LC_ALL=C`) so no locale can make
426
+ them reject the input. It changes to the top level of the
427
+ repository first, so a `cwd` in a subdirectory scans the same range. The
428
+ diff options override the external diff, textconv, binary, rename, color,
429
+ and prefix settings of the user's git configuration and the repository's
430
+ attributes (`--text` diffs a file marked `-diff` or `binary` as text), and
431
+ the log option suppresses signature output, so those settings cannot hide an
432
+ added line from the scan or add lines to it.
433
+
434
+ It covers the lines added between the run-base and `HEAD` outside the
435
+ top-level `.ai/` directory, and the message of every commit reachable from
436
+ `HEAD` and not from the run-base. That range includes upstream work merged
437
+ into the branch after the run-base, whose added lines and commit messages
438
+ are scanned as well and can produce hits the branch did not write. It does
439
+ not cover uncommitted changes, removed lines, an identifier directly next to
440
+ a NUL byte (the shell drops NUL bytes from the captured diff), pull request
441
+ titles or bodies, branch names, or identifiers in any other format.
442
+ The patterns are case-sensitive and can match unrelated tokens, such as a
443
+ product or part name built the same way; because `--text` also diffs files
444
+ git detects as binary by content, an added image, font, or archive usually
445
+ produces hits made of its raw bytes, and a repository whose own
446
+ documentation discusses these formats (a copy of these templates, for
447
+ example) matches as well. A hit is a failure of the extra; when the
448
+ orchestrator confirms a hit is a false positive it records that decision,
449
+ and it may narrow the pathspec with further `':(exclude)<path>'` entries
450
+ when it approves the extra.
@@ -75,6 +75,7 @@ All state for one unit of work lives in a run directory:
75
75
  04-implementation-summary.md
76
76
  05-review-findings.md
77
77
  06-handoff.md
78
+ evidence/
78
79
  ```
79
80
 
80
81
  Create it at the start of a run by copying `.ai/workflow/templates/` and fill
@@ -82,6 +83,15 @@ the files as the run progresses. The newest run directory is the active one
82
83
  unless a `.ai/run` pointer names one (see below);
83
84
  older directories are the auditable history. Do not edit past runs.
84
85
 
86
+ `evidence/` is an optional subdirectory, not one of the seven templated
87
+ files: nothing copies or requires it. The orchestrator, the implementer, and
88
+ the reviewer write into it (test logs, probe verdicts, reproduction output,
89
+ reviewer-reproduced evidence) when a briefing or an acceptance criterion
90
+ asks for a saved artifact instead of just a report field; the reviewer's
91
+ write-boundary rule names it as the one write-allowed location outside the
92
+ reviewed tree, besides the writer's own scratchpad. The explorer and the
93
+ advisor are read-only roles and never write into it, or anywhere else.
94
+
85
95
  The run directory may live in the workspace's own `.ai/runs/` or in one
86
96
  repository's `.ai/runs/`. Either way, bind every repository or worktree the
87
97
  run touches to it with a pointer file, `<worktree-root>/.ai/run`:
@@ -38,6 +38,8 @@ acceptance_criteria:
38
38
  **Verification Set**
39
39
 
40
40
  <!-- Checked-in path, repository identity, and run-local frozen snapshot. The
41
+ digest recorded in the briefing is what carries the orchestrator's approval
42
+ of the resolved argv to the implementer and reviewer. The
41
43
  orchestrator approves effective config/scripts before any acquisition or
42
44
  execution; include an ordered docs/okf bundle check whenever that directory
43
45
  exists. -->
@@ -166,7 +166,7 @@ interface ExtractedYaml {
166
166
  * shorter run length from every offset inside the run, each retry
167
167
  * rescanning the lazy body: work quadratic in the run's length, which a
168
168
  * single pasted return of a few hundred backticks already turns into
169
- * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
169
+ * seconds (CHANGELOG 0.37.0 names the measurement). Keeping the
170
170
  * run whole also makes the "closing run at least as long as the opening
171
171
  * one" rule above literal: an opener longer than any closing run in the
172
172
  * input is no fence at all, where splitting the run instead matched it
@@ -443,7 +443,7 @@ export const ARRAY_ELEMENT_EXPECTED = KIND_EXPECTED.string;
443
443
  * shorter run length from every offset inside the run, each retry
444
444
  * rescanning the lazy body: work quadratic in the run's length, which a
445
445
  * single pasted return of a few hundred backticks already turns into
446
- * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
446
+ * seconds (CHANGELOG 0.37.0 names the measurement). Keeping the
447
447
  * run whole also makes the "closing run at least as long as the opening
448
448
  * one" rule above literal: an opener longer than any closing run in the
449
449
  * input is no fence at all, where splitting the run instead matched it
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.39.0",
3
+ "version": "0.40.1",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",