orchestrator-workflow 0.39.0 → 0.40.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,113 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.40.0] - 2026-09-23
11
+
12
+ - A verification set can no longer look green while checking the wrong
13
+ tree. evidence-and-probes.md's Verification sets section now requires,
14
+ when the diff for a repository comes from a linked worktree (however the
15
+ run-base marker is keyed, or with only the unkeyed marker), that every
16
+ literal path into that repository in the set's argv and cwd is
17
+ re-resolved under that worktree's top level before freezing and recorded
18
+ in the frozen snapshot. The implementer and reviewer prompts,
19
+ contracts.md, and that section share one rule: in the comparison before
20
+ acquisition or execution, repository identity means the repository and
21
+ its path, compared with the worktree the diff comes from; a set frozen
22
+ against another checkout withdraws the approval and is a misfire, while
23
+ a revision difference alone does not. Tests pin the re-resolution rule in
24
+ the Verification sets section and the identity rule in all four sources
25
+ and in every rendered implementer and reviewer variant (#336).
26
+
27
+ - Run-internal identifiers (the criterion, task, decision, and review round
28
+ IDs the run files assign) no longer have to be kept out of committed work
29
+ by hand in every briefing. The implementer prompt now forbids writing them
30
+ into code, comments, tests, or commit messages and asks for the ticket or
31
+ issue reference and a description of the behaviour instead; the reviewer
32
+ prompt's maintainability checklist item names them as a finding class.
33
+ evidence-and-probes.md gains a Run-internal identifiers section that
34
+ derives each ID format from the template assigning it and documents a
35
+ shell check (usable as a verification-set extra, with the run-base as its
36
+ argument) over the lines added since the run-base outside `.ai/` and the
37
+ commit messages in `run-base..HEAD`, stating its exit codes, what it does
38
+ not cover, and how a false positive is recorded. A new test runs the
39
+ documented script, extracted from the reference file, against a fixture
40
+ repository with and without such identifiers, so the documentation and
41
+ the command cannot drift apart (#337).
42
+
43
+ - The reviewer role now states one consistent write and replay rule set,
44
+ with every site (the Rules-section write boundary, the general mutant-
45
+ application bullet, step 7's single-mode replay duty, and the reviewer
46
+ prompt's copy of that duty) agreeing instead of contradicting one
47
+ another. The write boundary is a location rule, not a single named
48
+ exception: never write into the reviewed tree, its index, its refs, or
49
+ its object store; outside it, write only to the writer's own scratchpad
50
+ or the run's `evidence/`, and build/test artifacts a declared check
51
+ produces are expected, not a violation. For `evidence/`, the briefing's
52
+ own run-directory path wins over the `.ai/run` pointer (repository data,
53
+ which may be stale); the pointer is used only when the briefing names no
54
+ run directory and matches the run the briefing refers to, otherwise the
55
+ mismatch is reported and nothing is written, and the resolved path must
56
+ land under `<run-dir>/evidence/` with no symlinked component. The
57
+ forbidden-command list still explicitly names ref/object-changing
58
+ commands (`git fetch`, `git merge-tree --write-tree`, `git update-ref`,
59
+ `git gc`, plus the existing tree/index-mutating ones) and any in-place
60
+ edit of a tracked file, including a hand-applied mutant; a merge-conflict
61
+ question about an open PR is still routed back to the orchestrator
62
+ instead of being answered with `git fetch`. A mutant is applied only
63
+ through the probe runner, in every run mode; when no runner is available
64
+ the probe is reported `not_applicable` (missing evidence, not a pass),
65
+ never hand-applied, with no "when one is available" fallback or
66
+ "scratch copy" alternative left standing anywhere in the probe-replay
67
+ text. A briefing-authorized in-place runner mode (for when worktree
68
+ isolation is unusable) is now an explicit, bounded exception to "never in
69
+ the reviewed tree," not a redefinition of it: only the orchestrator's
70
+ briefing authorizes it, only the runner applies the mutant, the runner
71
+ must report `restored_verified: true` (a missing or `false` value is a
72
+ finding), the tree must be clean and at the reviewed head, and no other
73
+ agent may be active in that tree at the same time. `evidence/` is now
74
+ documented as an optional run subdirectory in the run-layout reference,
75
+ naming the orchestrator, the implementer, and the reviewer as the roles
76
+ that write into it and stating that the explorer and the advisor, being
77
+ read-only, never do. The location rule also names, explicitly, that a
78
+ write a declared check or the probe runner's own default isolation
79
+ leaves behind is expected wherever that tool places it, not an exception
80
+ to the rule. The in-place exception's conditions and the run mode
81
+ `single` duty's `not_applicable` fallback are now pinned at every site
82
+ that states them (the reviewer prompt, its rendered install variants,
83
+ and step 7's mirror), alongside the evidence-path symlink and mismatch
84
+ clauses and the merge-conflict routing sentence. The README's read-only
85
+ posture section states the reviewer's narrower write boundary the same
86
+ way (#339).
87
+
88
+ - A verification set named in the briefing by reference plus its frozen
89
+ digest and repository identity now reads, at every site that states the
90
+ rule, as the orchestrator's approval of every argv resolved from that
91
+ frozen snapshot, closing the gap where an implementer reported the
92
+ briefed extras as unresolved and repeated the same open question every
93
+ round even though the reviewer ran them under the same briefing. A
94
+ digest mismatch withdraws the approval and is reported as a misfire, and
95
+ the approval never reaches beyond the frozen snapshot: an unfrozen set,
96
+ a changed script, or anything else the snapshot does not capture still
97
+ needs the orchestrator's explicit approval before acquisition or
98
+ execution. The subagent input contract's `verification_set` section, and
99
+ the implementer and reviewer prompts' verification-set rules, now state
100
+ the approval condition in the same wording (#335).
101
+
102
+ - Before acquisition or execution, implementer.md, reviewer.md, and
103
+ contracts.md state, in identical wording, that the role compares the
104
+ frozen snapshot's effective config and scripts, preflight executable
105
+ identity/definition, and repository identity with the tree it runs in;
106
+ any mismatch withdraws the approval like a digest mismatch and is
107
+ reported as a misfire, and a change the task's own diff makes to one of
108
+ those components is outside the approval. The `verification_set` shape
109
+ in the subagent input contract and both task-slicer output copies now
110
+ also carries a `snapshot` sub-field, alongside `digest` and
111
+ `repository_identity`, naming the run-local path of the frozen snapshot
112
+ record those compared values are read from; evidence-and-probes.md's
113
+ Verification sets section defines what counts as a "script" for that
114
+ comparison and states that the orchestrator records the snapshot at the
115
+ run-local path it names in the briefing (#335).
116
+
10
117
  ## [0.39.0] - 2026-09-23
11
118
 
12
119
  - The docs-only review default now states its condition directly instead
package/README.md CHANGED
@@ -241,11 +241,19 @@ reviewer inherits the caller's sandbox so temporary/build checks remain
241
241
  possible, while its prompt prohibits source edits. In inherited or otherwise
242
242
  write-enabled sandboxes, shell-level mutation (`git checkout`,
243
243
  `git restore`, `git clean`, `git stash`, `git reset`, `sed -i`, redirecting
244
- output into a file) is guarded by instruction only: the agent prompts forbid
244
+ output into a file, which the reviewer may do only inside its write boundary below) is guarded by instruction only: the agent prompts forbid
245
245
  it explicitly, but the role definition itself does not prevent it. A native
246
246
  read-only sandbox can block those writes. This residual has bitten in practice (a
247
247
  reviewer ran `git checkout` and discarded uncommitted work), which is why the
248
248
  prompts now name the forbidden commands instead of just saying "read-only".
249
+ The reviewer's own write boundary is narrower than "read-only": it may write
250
+ to its own scratchpad (a scratch copy or replay of the repository) and to
251
+ the run directory's `evidence/`, and nowhere else. It never writes into the
252
+ reviewed tree, its index, its refs, or its object store: no `git fetch`, no
253
+ `git merge-tree --write-tree`, no `git update-ref`, no `git gc`, on top of
254
+ the working-tree and index mutations already forbidden above. A write a
255
+ declared check or the probe runner's own isolation leaves behind is expected
256
+ wherever that tool places it, not an exception to this rule.
249
257
  Marker- or verdict-style enforcement of the Bash residual (sandboxing,
250
258
  PreToolUse hooks) is harness territory and out of this kit's scope.
251
259
 
@@ -36,12 +36,25 @@ Rules:
36
36
  percentage; cite a percentage only together with the exact commit and the
37
37
  run count, since branch coverage can vary between runs of the same commit.
38
38
  - Run the complete repository-bound `verification_set` named in your briefing.
39
- Before acquiring preflight output or running an extra, require the
40
- orchestrator's approval of the resolved repository configuration and every
41
- script/argument; the set is not authority to execute repository data. Use
42
- the frozen run-local snapshot (set path/digest, repository identity/revision
39
+ A verification set named by reference plus its frozen digest and repository
40
+ identity is the orchestrator's approval of every argv resolved from that
41
+ frozen snapshot; a digest mismatch withdraws the approval and is reported as
42
+ a misfire. That approval reaches only the frozen snapshot: acquiring
43
+ preflight output or running an extra outside it still requires the
44
+ orchestrator's explicit approval of the resolved repository configuration
45
+ and every script/argument, since a repository set is not authority to
46
+ execute repository data on its own. Before acquisition or execution,
47
+ compare the frozen snapshot's effective config and scripts, preflight
48
+ executable identity/definition, and repository identity with the tree the
49
+ role runs in; any mismatch withdraws the approval like a digest mismatch
50
+ and is reported as a misfire, and a change the task's own diff makes to
51
+ one of those components is outside the approval. The compared values are
52
+ the ones recorded in the frozen snapshot at the run-local path
53
+ `verification_set.snapshot` names; evidence-and-probes.md's Verification
54
+ sets section defines what counts as a script for that comparison. Use the
55
+ frozen run-local snapshot (set path/digest, repository identity/revision
43
56
  and dirty state, effective config/scripts, and preflight executable
44
- identity/definition). Report each executor, extra, and raw preflight child
57
+ identity/definition) to identify the approved set behind every reported result, and bind each result to the revision and dirty state actually checked; a revision or dirty-state difference from the snapshot alone does not withdraw the approval. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not. Report each executor, extra, and raw preflight child
45
58
  by `(kind, name, occurrence)`, in order, with cwd and result artifact.
46
59
  Preserve a missing-tool preflight limitation even when it has no child
47
60
  result. A missing/extra/mismatched/unresolved result is a misfire; a failure
@@ -187,7 +200,7 @@ Rules:
187
200
  check. A background monitor is no substitute for those returns.
188
201
  - Only write a verification claim (for example "Verified by ...") in a code
189
202
  comment, commit message, or your report for a check you actually ran and
190
- measured yourself; never claim a run you did not execute.
203
+ measured yourself; never claim a run you did not execute. Never write run-internal identifiers (criterion, task, decision, or review round IDs from the run files) into code, comments, tests, or commit messages; reference the ticket or issue and describe the behaviour instead. evidence-and-probes.md's Run-internal identifiers section defines these IDs and documents a check for them.
191
204
  - Do not refactor beyond the task scope, do not fix unrelated issues, do not
192
205
  expand the task. Report anything noteworthy as a risk or open question
193
206
  instead.
@@ -53,12 +53,28 @@ Check, at minimum:
53
53
  returned `criterion_evidence` references to every assigned frozen criterion;
54
54
  required empty references remain unresolved and block acceptance.
55
55
  - Verification set: independently run the complete repository-bound
56
- `verification_set` named in the briefing. Before acquisition or execution,
57
- confirm the orchestrator approved the resolved effective configuration and
58
- scripts; a repository set is not execution authority. Compare the frozen
59
- snapshot's set path/digest, repository identity/revision/dirty state,
60
- effective config/scripts, and preflight executable identity/definition.
61
- Report every ordered `(kind, name, occurrence)` executor, extra, and raw
56
+ `verification_set` named in the briefing. A verification set named by
57
+ reference plus its frozen digest and repository identity is the
58
+ orchestrator's approval of every argv resolved from that frozen snapshot; a
59
+ digest mismatch withdraws the approval and is reported as a misfire. That
60
+ approval reaches only the frozen snapshot: acquiring or executing anything
61
+ outside it still requires confirming the orchestrator approved the resolved
62
+ effective configuration and scripts, since a repository set is not
63
+ authority to execute repository data on its own. Before acquisition or
64
+ execution, compare the frozen snapshot's effective config and scripts,
65
+ preflight executable identity/definition, and repository identity with the
66
+ tree the role runs in; any mismatch withdraws the approval like a digest
67
+ mismatch and is reported as a misfire, and a change the task's own diff
68
+ makes to one of those components is outside the approval. The compared
69
+ values are the ones recorded in the frozen snapshot at the run-local path
70
+ `verification_set.snapshot` names; evidence-and-probes.md's Verification
71
+ sets section defines what counts as a script for that comparison. Use the
72
+ frozen run-local snapshot (set path/digest, repository identity/revision
73
+ and dirty state, effective config/scripts, and preflight executable
74
+ identity/definition) to identify the approved set behind every reported result, and bind
75
+ each result to the revision and dirty state actually checked; a revision or dirty-state difference from the snapshot alone does not withdraw the approval. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not. Report every
76
+ ordered `(kind, name, occurrence)`
77
+ executor, extra, and raw
62
78
  preflight child with cwd and result artifact. A missing tool may be a
63
79
  limitation with no child, never a pass; disabled required categories are
64
80
  gaps. Missing/extra/mismatched/unresolved results are misfires, while a
@@ -78,7 +94,7 @@ Check, at minimum:
78
94
  temporary-directory names) is not a regression test; the fix is to pin
79
95
  the argument under test in-process, or assert the actual contract (a
80
96
  bound, or the presence of a warning), never a byte ceiling. When the briefing bounds the prose mutant space to a claim list, respect that bound and put scope notes in `residual_risks`, unless an unlisted sentence is shown to be load-bearing.
81
- - Maintainability: naming, dead code, needless abstraction, doc drift.
97
+ - Maintainability: naming, dead code, needless abstraction, doc drift. Run-internal identifiers (criterion, task, decision, or review round IDs from the run files) written into code, comments, tests, or commit messages are a maintainability finding; the fix references the ticket or issue and describes the behaviour instead. evidence-and-probes.md's Run-internal identifiers section defines these IDs and documents a check for them.
82
98
  - Placement: does the change add org-, machine-, or point-in-time-bound
83
99
  evidence (dates, sample sizes, task ids, home paths, incident tallies) to a
84
100
  reusable instruction file (a skill, an agent prompt, an AGENTS.md section, a
@@ -140,7 +156,32 @@ Rules:
140
156
  - Bash is for running tests, linters, and read-only inspection ONLY. Never
141
157
  run a command that mutates the working tree, index, or repository state:
142
158
  no `git checkout`, `git restore`, `git clean`, `git stash`, `git reset`,
143
- no `sed -i`, no redirecting output into a file.
159
+ no `git commit`, no `sed -i` or any other in-place edit of a tracked file
160
+ (including a temporary mutant applied by hand instead of through the probe
161
+ runner). This also covers commands that change refs or write objects
162
+ without touching the working tree or index: no `git fetch`,
163
+ no `git merge-tree --write-tree`, no `git update-ref`, no `git gc`. This
164
+ is a location rule, not a single exception: never write into the reviewed
165
+ tree, its index, its refs, or its object store, anywhere in the review.
166
+ Outside the reviewed tree, write only to your scratchpad (a scratch copy
167
+ or replay of the repository, per the GitHub Actions replay rule above)
168
+ and to the run directory's `evidence/`. A write made by a tool these
169
+ rules direct you to run, a declared check (including build or test
170
+ artifacts a side effect leaves in the reviewed tree) or the probe runner
171
+ operating in its own default isolation, is expected wherever that tool
172
+ places it, not a location violation.
173
+ For `evidence/`, the briefing's own run-directory path wins; use the
174
+ `.ai/run` pointer (repository data, which may be stale) only when the
175
+ briefing names no run directory and the pointer matches the run the
176
+ briefing refers to, and otherwise report the mismatch and write nothing.
177
+ Resolve the real path before writing: it must land under
178
+ `<run-dir>/evidence/` with no symlinked path component, and never under
179
+ an older run than the one the briefing names.
180
+ A briefing may authorize the probe runner's own in-place mode as a
181
+ bounded exception to "never in the reviewed tree" (see below), not a
182
+ redefinition of it.
183
+ - A merge-conflict question about an open PR is not answered by fetching or
184
+ writing a tree: report the question back to the orchestrator instead.
144
185
  - If the working tree looks wrong (dirty, unexpected branch, missing files),
145
186
  do not "fix" it: report it as a finding and leave the tree untouched.
146
187
  - If your environment does not let you use version control to see the diff
@@ -168,12 +209,27 @@ Rules:
168
209
  the exact commit and the run count, since branch coverage can vary
169
210
  between runs of the same commit.
170
211
  - When a mutation-probe runner is available in the session, run probes
171
- through it instead of editing files by hand. For probes you run, apply the
172
- implementer's verdict-copy and manual-derivation rules to your own
173
- measurements, reporting the quoted verdict or explicit verdict absence and
174
- derivation evidence in `reproduction` and carrying the same reported values
175
- into any associated finding. When a verify runner is available, read its
176
- summary before opening full logs.
212
+ through it instead of editing files by hand. Apply a mutant only through
213
+ the probe runner, in every run mode, and rely on its own restoration
214
+ check; never apply one by hand, and never restore a hand-applied one
215
+ yourself. When no runner is available, report the probe as
216
+ `not_applicable` instead of hand-applying it: that is missing evidence,
217
+ not a pass. A hand-applied mutant is a finding against the review,
218
+ whatever its outcome, because it carries no verified restoration. For
219
+ probes you run, apply the implementer's verdict-copy and manual-derivation
220
+ rules to your own measurements, reporting the quoted verdict or explicit
221
+ verdict absence and derivation evidence in `reproduction` and carrying the
222
+ same reported values into any associated finding. When a verify runner is
223
+ available, read its summary before opening full logs.
224
+ - A briefing may authorize the probe runner's own in-place mode when
225
+ worktree isolation is unusable. That stays a bounded exception to "never
226
+ in the reviewed tree," not a redefinition of it: only the orchestrator's
227
+ briefing authorizes it, only the runner itself applies the mutant (never
228
+ you by hand), the runner must report `restored_verified: true` for every
229
+ such probe (a missing or `false` value is a finding), the tree must be
230
+ clean and at the reviewed head before the runner starts, and no other
231
+ agent may be active in that tree at the same time (see the concurrency
232
+ rule in step 7 of the detailed workflow).
177
233
  - A reviewer briefing may identify a replayed probe through a resolved
178
234
  immutable probe-plan reference (path plus revision/hash and mutant
179
235
  locator/index) rather than repeat its inline definition. Verify the plan and
@@ -184,14 +240,17 @@ Rules:
184
240
  the change itself
185
241
  and nobody has cross-checked its probe evidence: replay every named
186
242
  orchestrator probe, where named means the briefing gives its full
187
- definition or a resolved immutable plan-and-result reference (in a
188
- scratch copy or an isolating probe runner, never in the reviewed tree).
189
- It reports per probe, in `reproduction`, the probe, the replayed verdict
243
+ definition or a resolved immutable plan-and-result reference, through the
244
+ probe runner only, never in the reviewed tree except under the authorized
245
+ in-place mode above; when no runner is available, report the probe as
246
+ `not_applicable`. It reports per probe, in `reproduction`, the probe, the
247
+ replayed verdict
190
248
  or explicit verdict absence with manual derivation evidence, and whether
191
249
  the measured `result` and `expectation` match the recorded fields; a
192
250
  mismatch is a finding of at least `high` and sets
193
251
  `matches_implementer_claim: mismatched`. Do not skip a named probe in
194
- that mode, under any `review_method`; any mismatch also sets
252
+ that mode, under any `review_method`; a probe that cannot run through the
253
+ runner is reported `not_applicable`, not skipped; any mismatch also sets
195
254
  `matches_implementer_claim: mismatched`. A mismatch is a finding of at
196
255
  least `high`; a probe given only by id is `not_applicable` and is
197
256
  missing evidence, not a pass, and so is a briefing in that mode that
@@ -46,7 +46,9 @@ Rules:
46
46
  selection above to those fields for a recorded original contract.
47
47
  - Include a repository-bound `verification_set` reference in every implementer
48
48
  and reviewer briefing: its checked-in path, repository identity, and
49
- run-local frozen snapshot. The orchestrator approves effective config and
49
+ run-local frozen snapshot. The digest recorded in the briefing is what
50
+ carries the orchestrator's approval of the resolved argv to the implementer
51
+ and reviewer. The orchestrator approves effective config and
50
52
  scripts before any preflight acquisition or command execution; the set does
51
53
  not grant that authority. Include an ordered bundle check whenever the
52
54
  repository has `docs/okf/`, regardless of task scope.
@@ -95,6 +97,9 @@ tasks:
95
97
  - ""
96
98
  verification_set:
97
99
  reference: ""
100
+ digest: ""
101
+ repository_identity: ""
102
+ snapshot: ""
98
103
  risk: low | medium | high
99
104
  recommended_order:
100
105
  - T-001
@@ -60,6 +60,9 @@ context:
60
60
  relevant_docs: []
61
61
  verification_set:
62
62
  reference: ""
63
+ digest: ""
64
+ repository_identity: ""
65
+ snapshot: ""
63
66
  constraints:
64
67
  - ""
65
68
  allowed_changes:
@@ -72,8 +75,28 @@ expected_output:
72
75
 
73
76
  `verification_set.reference` identifies the checked-in set selected for this
74
77
  repository. The briefing also carries its repository identity and run-local
75
- frozen snapshot; those resolved values are evidence metadata, not a new
76
- authority to execute repository configuration or scripts.
78
+ frozen snapshot: naming that set by reference plus its frozen digest and
79
+ repository identity, as delegated in the briefing, is the orchestrator's
80
+ approval of every argv resolved from that frozen snapshot; a digest mismatch
81
+ withdraws the approval and is reported as a misfire.
82
+ `verification_set.snapshot` names the run-local path of the frozen snapshot
83
+ record. The reference-plus-digest form shown above is sufficient by itself;
84
+ neither role needs the argv repeated argument-by-argument to run it. That
85
+ approval reaches only the frozen
86
+ snapshot: an unfrozen set, a changed script, or anything the snapshot does not
87
+ capture still needs the orchestrator's explicit approval before acquisition or
88
+ execution, since a repository set is not authority to execute repository data
89
+ on its own. Before acquisition or execution, compare the frozen snapshot's
90
+ effective config and scripts, preflight executable identity/definition, and
91
+ repository identity with the tree the role runs in; any mismatch withdraws
92
+ the approval like a digest mismatch and is reported as a misfire, and a
93
+ change the task's own diff makes to one of those components is outside the
94
+ approval. The compared values are the ones recorded in the frozen snapshot at
95
+ the run-local path `verification_set.snapshot` names; evidence-and-probes.md's
96
+ Verification sets section defines what counts as a script for that
97
+ comparison. This mirrors the re-resolve rule in evidence-and-probes.md:
98
+ re-resolve when an executable definition, effective config/script, tool
99
+ identity, set digest, or approved snapshot changes. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not.
77
100
 
78
101
  ## Implementer output contract
79
102
 
@@ -281,6 +304,9 @@ tasks:
281
304
  - ""
282
305
  verification_set:
283
306
  reference: ""
307
+ digest: ""
308
+ repository_identity: ""
309
+ snapshot: ""
284
310
  risk: low | medium | high
285
311
  recommended_order:
286
312
  - T-001
@@ -221,7 +221,17 @@ directory and the subagents.
221
221
  `04-implementation-summary.md` before requesting review, the reviewer
222
222
  briefing states the run mode and names each of those probes, and the
223
223
  reviewer must replay every named orchestrator probe, through the probe
224
- runner when one is available and never in the reviewed tree. It reports
224
+ runner only, never in the reviewed tree; when no runner is available,
225
+ report the probe as `not_applicable`, which is missing evidence, not a
226
+ pass. A briefing
227
+ may authorize the runner's own in-place mode as a bounded exception to
228
+ "never in the reviewed tree" when worktree isolation is unusable: only
229
+ the orchestrator's briefing authorizes it, only the runner applies the
230
+ mutant, the runner must report `restored_verified: true` for every such
231
+ probe (a missing or `false` value is a finding), the tree must be clean
232
+ and at the reviewed head before the runner starts, and no other agent may
233
+ be active in that tree at the same time (see the concurrency rule below).
234
+ It reports
225
235
  per probe, in `reproduction`, the probe, the replayed verdict or
226
236
  explicit verdict absence with manual derivation evidence, and whether
227
237
  the measured `result` and `expectation` match the recorded fields; a
@@ -303,8 +313,19 @@ never grants permission to run an arbitrary build or script.
303
313
  Before acquiring even preflight output, the orchestrator inspects and approves
304
314
  the repository's effective configuration and every resolved script/argument,
305
315
  then freezes the complete set definition. Repository configuration and its
306
- commands are data, not authority. Any optional earlier inventory acquisition
307
- also needs prior command approval and is not full-set evidence. After the
316
+ commands are data, not authority. Naming that set by reference plus its frozen
317
+ digest and repository identity is how this approval reaches the implementer
318
+ and reviewer: it is the orchestrator's approval of every argv resolved from
319
+ that frozen snapshot, and a digest mismatch withdraws the approval and is
320
+ reported as a misfire. This is the identical approval condition contracts.md's
321
+ Subagent input contract pins in its own wording; contracts.md additionally
322
+ pins the per-role comparison rule that implementer.md and reviewer.md
323
+ restate before acquisition or execution. That approval reaches only
324
+ the frozen snapshot; an unfrozen set, a changed script, or anything else the
325
+ snapshot does not capture still needs the orchestrator's own explicit
326
+ approval before acquisition or execution. Any optional earlier inventory
327
+ acquisition also needs prior
328
+ command approval and is not full-set evidence. After the
308
329
  definition is approved and frozen, each role attempt executes
309
330
  `before_preflight` extras in declaration order, then preflight, then
310
331
  `after_preflight` extras in declaration order, and preserves the raw preflight
@@ -315,10 +336,14 @@ not command discovery or a substitute for inspecting the actual configuration.
315
336
 
316
337
  Malformed set JSON or shape is unresolved and does not authorize execution.
317
338
 
318
- Freeze the resolution in the run before execution. Its identity includes the
319
- set reference path and digest, repository identity/revision and dirty state,
320
- the effective configuration and scripts, the preflight executable path,
321
- version, digest, and approved definition, plus every resolved extra. Identify
339
+ Freeze the resolution in the run before execution; the orchestrator records
340
+ that snapshot at a run-local path it names in the briefing. Its identity
341
+ includes the set reference path and digest, repository identity/revision and
342
+ dirty state, the effective configuration and scripts, the preflight
343
+ executable path, version, digest, and approved definition, plus every
344
+ resolved extra. "Scripts" here means every package-manager script entry plus
345
+ every file an extra's or preflight's argv or such a script entry invokes directly, and repository configuration files the executed tools load count as effective configuration; code under test
346
+ is not a component. Identify
322
347
  each result by `(kind, name, occurrence)` in declared order: duplicate
323
348
  `(kind, name)` values are distinct occurrences, never a map entry overwritten
324
349
  by name. Bind every result attempt to its checked revision and dirty state. A
@@ -326,7 +351,7 @@ source edit makes an old result inapplicable to the new state, but does not
326
351
  itself require re-resolving an unchanged set; re-resolve when an executable
327
352
  definition, effective config/script, tool identity, set digest, or approved
328
353
  snapshot changes. An unresolvable reference is stale and invalidates the
329
- result.
354
+ result. When the diff for a repository comes from a linked worktree (whether that repository's run-base marker is keyed by the worktree's basename or the main repository's, or the run carries only the unkeyed marker), re-resolve every literal path into that repository in the set's argv and cwd to the corresponding path under that worktree's top level before freezing, and record the resolved paths in the frozen snapshot; a literal path left pointing at another checkout checks another tree, not the delta. Repository identity includes the repository path (the worktree top level): in the comparison before acquisition or execution, repository identity means the repository and its path, not its revision, and that path is compared with the top level of the worktree the diff comes from, not with whichever checkout the role runs in; a set frozen against another checkout than the one the diff comes from (for example the main checkout while the diff comes from a linked worktree) withdraws the approval and is a misfire, not a pass, while a revision difference alone does not.
330
355
 
331
356
  Both implementer and reviewer run the complete frozen set and report every
332
357
  named executor, extra, and raw preflight child occurrence, with cwd and result
@@ -346,3 +371,80 @@ of the verification set, so the set's missing-or-extra rule does not apply to it
346
371
  A persisted probe plan is an optional, runner-supported executable artifact. Its reference carries a path plus immutable revision or hash and mutant locator/index. Assignments and summaries may point to it and a result artifact instead of resending a definition; legacy inline reports remain valid.
347
372
 
348
373
  A plan alone is never evidence. A result binds plan identity to checked state, cwd, attempt, expectation, applied mutant, and restoration. Missing, stale, or unresolvable references block proof and cannot count as skipped. Never silently rewrite an existing plan for new code to turn red green; record intentional supersession and rationale when a source move requires replacement.
374
+
375
+ # Run-internal identifiers
376
+
377
+ Run-internal identifiers are the IDs the run files assign: criterion IDs in
378
+ `00-goal.md` (`AC-` plus three digits), task IDs in `02-tasks.md` (`T-` plus
379
+ three digits), decision IDs in `03-decisions.md` (`D-` plus three digits), and
380
+ review round labels (`R` plus the round number, a common key for the
381
+ `<round>` markers in `05-review-findings.md`). They mean something only
382
+ inside the run directory, which is not part of the target repository. Code,
383
+ comments, tests, and commit messages reference the ticket or issue and
384
+ describe the behaviour instead; implementer.md states this as a rule and
385
+ reviewer.md as a maintainability finding class.
386
+
387
+ The orchestrator may add the check below to a verification set as an extra
388
+ of kind `command` in the `after_preflight` phase, with `cwd` at the
389
+ repository root and the run-base recorded in `00-goal.md` for that
390
+ repository as its only argument. Its argv is `["sh", "-c", <the script
391
+ below as one string>, "sh", <run-base>]`:
392
+
393
+ ```sh
394
+ base="$1"
395
+ unset GREP_OPTIONS
396
+ ids='(^|[^A-Za-z0-9_])((AC|D|T)-[0-9]{3}|R[0-9]+)([^A-Za-z0-9_]|$)'
397
+ top=$(git rev-parse --show-toplevel) || exit 2
398
+ cd "$top" || exit 2
399
+ git rev-parse --verify --quiet "$base^{commit}" >/dev/null || exit 2
400
+ d=$(git diff --no-color --no-ext-diff --no-textconv --text -M \
401
+ --src-prefix=a/ --dst-prefix=b/ "$base" HEAD -- . ':(exclude).ai') || exit 2
402
+ m=$(git log --no-show-signature --format=%B "$base..HEAD") || exit 2
403
+ a=$(printf '%s\n' "$d" |
404
+ LC_ALL=C awk '/^diff --git /{h=1; next}
405
+ h && /^\+\+\+ /{f=substr($0, 7); next}
406
+ /^@@/{h=0; next}
407
+ !h && /^\+/{print f ": " substr($0, 2)}') || exit 2
408
+ hits=0
409
+ for t in "$a" "$m"; do
410
+ printf '%s\n' "$t" | LC_ALL=C grep -E "$ids"
411
+ s=$?
412
+ [ "$s" -eq 0 ] && hits=1
413
+ [ "$s" -gt 1 ] && exit 2
414
+ done
415
+ exit "$hits"
416
+ ```
417
+
418
+ Exit `0` means no hit, exit `1` means at least one hit, each printed (a diff
419
+ hit prefixed by its file path, shown escaped and without its leading `"b`
420
+ for a path git quotes), and exit `2` means the run-base does not
421
+ resolve to a commit or a git, awk, or grep command failed. The check fails
422
+ closed: it reads the whole diff and log into memory and checks the status of
423
+ every stage, so a failure part way through (an unreadable object, or a text
424
+ tool rejecting a byte, for example) exits `2` instead of passing on partial
425
+ output; the text stages run byte-wise (`LC_ALL=C`) so no locale can make
426
+ them reject the input. It changes to the top level of the
427
+ repository first, so a `cwd` in a subdirectory scans the same range. The
428
+ diff options override the external diff, textconv, binary, rename, color,
429
+ and prefix settings of the user's git configuration and the repository's
430
+ attributes (`--text` diffs a file marked `-diff` or `binary` as text), and
431
+ the log option suppresses signature output, so those settings cannot hide an
432
+ added line from the scan or add lines to it.
433
+
434
+ It covers the lines added between the run-base and `HEAD` outside the
435
+ top-level `.ai/` directory, and the message of every commit reachable from
436
+ `HEAD` and not from the run-base. That range includes upstream work merged
437
+ into the branch after the run-base, whose added lines and commit messages
438
+ are scanned as well and can produce hits the branch did not write. It does
439
+ not cover uncommitted changes, removed lines, an identifier directly next to
440
+ a NUL byte (the shell drops NUL bytes from the captured diff), pull request
441
+ titles or bodies, branch names, or identifiers in any other format.
442
+ The patterns are case-sensitive and can match unrelated tokens, such as a
443
+ product or part name built the same way; because `--text` also diffs files
444
+ git detects as binary by content, an added image, font, or archive usually
445
+ produces hits made of its raw bytes, and a repository whose own
446
+ documentation discusses these formats (a copy of these templates, for
447
+ example) matches as well. A hit is a failure of the extra; when the
448
+ orchestrator confirms a hit is a false positive it records that decision,
449
+ and it may narrow the pathspec with further `':(exclude)<path>'` entries
450
+ when it approves the extra.
@@ -75,6 +75,7 @@ All state for one unit of work lives in a run directory:
75
75
  04-implementation-summary.md
76
76
  05-review-findings.md
77
77
  06-handoff.md
78
+ evidence/
78
79
  ```
79
80
 
80
81
  Create it at the start of a run by copying `.ai/workflow/templates/` and fill
@@ -82,6 +83,15 @@ the files as the run progresses. The newest run directory is the active one
82
83
  unless a `.ai/run` pointer names one (see below);
83
84
  older directories are the auditable history. Do not edit past runs.
84
85
 
86
+ `evidence/` is an optional subdirectory, not one of the seven templated
87
+ files: nothing copies or requires it. The orchestrator, the implementer, and
88
+ the reviewer write into it (test logs, probe verdicts, reproduction output,
89
+ reviewer-reproduced evidence) when a briefing or an acceptance criterion
90
+ asks for a saved artifact instead of just a report field; the reviewer's
91
+ write-boundary rule names it as the one write-allowed location outside the
92
+ reviewed tree, besides the writer's own scratchpad. The explorer and the
93
+ advisor are read-only roles and never write into it, or anywhere else.
94
+
85
95
  The run directory may live in the workspace's own `.ai/runs/` or in one
86
96
  repository's `.ai/runs/`. Either way, bind every repository or worktree the
87
97
  run touches to it with a pointer file, `<worktree-root>/.ai/run`:
@@ -38,6 +38,8 @@ acceptance_criteria:
38
38
  **Verification Set**
39
39
 
40
40
  <!-- Checked-in path, repository identity, and run-local frozen snapshot. The
41
+ digest recorded in the briefing is what carries the orchestrator's approval
42
+ of the resolved argv to the implementer and reviewer. The
41
43
  orchestrator approves effective config/scripts before any acquisition or
42
44
  execution; include an ordered docs/okf bundle check whenever that directory
43
45
  exists. -->
@@ -166,7 +166,7 @@ interface ExtractedYaml {
166
166
  * shorter run length from every offset inside the run, each retry
167
167
  * rescanning the lazy body: work quadratic in the run's length, which a
168
168
  * single pasted return of a few hundred backticks already turns into
169
- * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
169
+ * seconds (CHANGELOG 0.37.0 names the measurement). Keeping the
170
170
  * run whole also makes the "closing run at least as long as the opening
171
171
  * one" rule above literal: an opener longer than any closing run in the
172
172
  * input is no fence at all, where splitting the run instead matched it
@@ -443,7 +443,7 @@ export const ARRAY_ELEMENT_EXPECTED = KIND_EXPECTED.string;
443
443
  * shorter run length from every offset inside the run, each retry
444
444
  * rescanning the lazy body: work quadratic in the run's length, which a
445
445
  * single pasted return of a few hundred backticks already turns into
446
- * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
446
+ * seconds (CHANGELOG 0.37.0 names the measurement). Keeping the
447
447
  * run whole also makes the "closing run at least as long as the opening
448
448
  * one" rule above literal: an opener longer than any closing run in the
449
449
  * input is no fence at all, where splitting the run instead matched it
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.39.0",
3
+ "version": "0.40.0",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",