orchestrator-workflow 0.31.0 → 0.33.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,441 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.33.0] - 2026-09-12
11
+
12
+ ### Changed
13
+
14
+ - Implementer reports now paste non-empty commit lists from `git log --reverse
15
+ --format=%H <base>..HEAD`, and foreground verification/probe plans and
16
+ repeat tallies return in the same turn as the last check; a background
17
+ monitor does not substitute. Anchored by pandora batch48 evidence.
18
+
19
+ - The implementer output contract now enumerates mutation-probe `result` as
20
+ `killed | survived | not_applicable` in both the installed prompt and the
21
+ SKILL.md reference. The docs-consistency guard separately pins each copy's
22
+ complete field block, including that enum, and now recognizes versioned
23
+ parenthesized headings in `see <name> below` forward pointers.
24
+
25
+ - The implementer and reviewer prompts (`assets/agents/implementer.md`,
26
+ `assets/agents/reviewer.md`, mirrored in SKILL.md) now say: cite a
27
+ coverage gate's threshold and pass/fail counts, not a run-specific
28
+ coverage percentage; cite a percentage only together with the exact
29
+ commit and the run count, since branch coverage varies between runs of
30
+ the same commit. Anchored by pandora run
31
+ `.ai/runs/2026-09-11-memory-sync-wipe`.
32
+ - The implementer `mutation_probes` output field (`assets/agents/
33
+ implementer.md`, mirrored in SKILL.md, and the `04-implementation-
34
+ summary.md` template's Mutation Probes table) now carries the mutant's
35
+ definition, not only its label: `file`, `anchor` (a line number or a
36
+ unique string), `before`, and `after` alongside the existing
37
+ `verified_applied_via`, `result`, `restored_verified`, and `replayed`
38
+ fields, so a later round can mechanically reapply the same edit instead
39
+ of only reading a prose description. Added `expectation: met | violated
40
+ | not_applicable` beside `result`, reporting whether the probe's
41
+ `result` matched its `--expect`, scoped to a measured `killed` or
42
+ `survived` `result` (a routine negative-control probe now reports
43
+ `result: survived, expectation: met`, which is not a regression) and
44
+ `not_applicable` otherwise (for example when the mutant could not be
45
+ applied and no `result` was measured). Added an eleventh sub-field,
46
+ `reason`: free text, required when `result` is `not_applicable`, empty
47
+ otherwise, carrying one of two canonical strings that distinguish a
48
+ non-regression from a regression: `no definition recorded` (a
49
+ prior-round probe recorded with only an id, no definition to reapply)
50
+ and `target text no longer present` (a replayed probe whose mutant can
51
+ no longer be applied). The fix-round replay rule now names a
52
+ probe to replay by its definition, not merely by its id, and treats a
53
+ replayed probe whose `expectation` is now `violated` (or which can no
54
+ longer be applied) as the regression signal, not `result` alone.
55
+ Anchored by pandora run `.ai/runs/2026-09-11-memory-sync-wipe` and the
56
+ agent-primitives `probe` result/expectation split (0.3.0).
57
+
58
+ ## [0.32.0] - 2026-09-11
59
+
60
+ ### Added
61
+
62
+ - A review-method axis, orthogonal to the effort tier: every reviewer
63
+ briefing now names `review_method: normal | rigorous | adversarial`
64
+ (`assets/agents/reviewer.md`, SKILL.md step 7, the kit-fence
65
+ Scaling-delegation text). The three methods are obligation sets, not
66
+ personas: `normal` reads the diff and spec and runs the declared tests
67
+ once, for docs/renames/batch cosmetics; `rigorous` (the default when a
68
+ briefing names none) adds an independent extract, a base-attribution
69
+ control, and mandatory reproduction of every empirical claim (the
70
+ pre-existing `reproduction`/`matches_implementer_claim` requirement);
71
+ `adversarial` adds one discriminating probe or negative control per
72
+ acceptance criterion, an active search of the neighbouring scenario
73
+ space, an attempt to break the claimed invariant, and a list of break
74
+ attempts that failed. `adversarial` and `rigorous` both carry a
75
+ withdrawal rule: a finding that does not reproduce on a second attempt
76
+ with a corrected harness is withdrawn in the same round and reported
77
+ under a `withdrawn` list with the reason, so the method cannot buy false
78
+ positives; emit `withdrawn: []` when nothing was withdrawn. The reviewer
79
+ output contract (both copies, `reviewer.md` and SKILL.md) gains
80
+ `method_applied` and `withdrawn`; until the grounding-mcp reader parses
81
+ the marker (agent-grounding follow-up 5df7b809), the orchestrator checks
82
+ by hand that the return's `method_applied` matches the briefing's
83
+ `review_method`, resupplying a mismatch or omission rather than
84
+ accepting it.
85
+ `assets/templates/05-review-findings.md` gains a `Method` line per review
86
+ round, outside the pinned Findings table
87
+ (`test/template-markers.test.ts` pins it); the grounding-mcp
88
+ completeness reader does not parse it yet, tracked as a follow-up in the
89
+ agent-grounding repo. SKILL.md's selection rule: `adversarial` at
90
+ minimum for security judgment, install/deploy scripts, hand-edited
91
+ lockfiles, cross-major overrides, or anything the operator flags
92
+ high-risk; `normal` only for docs, renames, or batch cosmetics;
93
+ `rigorous` otherwise; never `adversarial` on the `-medium` reviewer tier
94
+ (budget mismatch). Tiers themselves are unchanged. Anchored by the
95
+ pandora run `2026-09-11-cve-sweep`: reviews R1, R5, R11, R13 and R14
96
+ found the Critical and High findings by probing; R4, R6 and R9 on the
97
+ default method returned no findings and R8 only two informational lows;
98
+ R5 withdrew a harness artefact. The selection rule stays advisory, not
99
+ an AGENTS.md rule, until an A/B (same tasks, run once under `rigorous`
100
+ and once under `adversarial`, counting real Critical/High findings and
101
+ findings withdrawn) is recorded.
102
+
103
+ - A citation-sibling-drift guard (`test/docs-consistency.test.ts`, next to
104
+ the existing anchor-load-bearing checks) catches a citation that resolves
105
+ and anchors correctly on its own but names the wrong sibling among a run
106
+ of near-identical citations, a class neither okf-kit's `citations-resolve`
107
+ rule nor the local anchor guards can see, because both check an anchor
108
+ only inside its own cited range. Two rules, applied per paragraph: (a) the
109
+ same `file:range#anchor` cited twice in one paragraph, unless allowlisted
110
+ with a reason; (b) a string anchor's text also occurring, uncited, at
111
+ another line of the same target within a 20-line window (widened from 10
112
+ in review round 2, see below) while the paragraph cites a sibling range
113
+ of that file, unless allowlisted. Fixtures
114
+ reproduce three review findings that shared this shape and were the
115
+ motivation for the guard: three sibling `it`-block citations collapsing
116
+ onto one range twice, the third never cited; three per-harness bullet
117
+ citations doing the same; two logically distinct assertions collapsing
118
+ onto one shared range and anchor text, the second's own line never cited.
119
+ Run over the current bundle, every hit either rule reports is read
120
+ against its own target file and the citing paragraph, then fixed or
121
+ allowlisted; the measured per-rule hit counts live in `docs/okf/log.md`
122
+ with the classification that produced them, not here, so the two sites
123
+ cannot drift apart. The recurring coincidence shapes are a doc-wide
124
+ "topic sentence, then repeat as the closing list item" convention for
125
+ rule (a), and a short or common token -- a keyword, a mirrored field on
126
+ twin interfaces, a comment restating a literal, a test-assertion idiom on
127
+ an adjacent line, a reused local variable name -- recurring near a real
128
+ citation for rule (b). Lives in this file rather than as an okf-kit rule
129
+ because this
130
+ suite already runs on every PR while an okf-kit rule needs a release and
131
+ a fleet pin bump first; an opt-in okf-kit rule for the same class is a
132
+ named follow-up candidate once this guard has proven itself.
133
+ - Review round 2 of the citation-sibling-drift guard above (task agent-dx
134
+ 9f72ae6d): the round-1 review found 7 of the round's 18 raw hits were not
135
+ coincidental at all -- real mis-pointed citations that got allowlisted
136
+ instead of fixed, because the round-1 pass classified every hit by range
137
+ and reason without re-deriving each cited claim's real evidence line by
138
+ line. All seven re-pointed to their real evidence, citation-only, no
139
+ content changes: a duplicate `docs-consistency.test.ts` self-citation
140
+ whose second claim's real test sat 26 lines below the first
141
+ (`subagent-contracts-superset.md`); an `init.test.ts` range that stopped
142
+ 6 lines short of the `model: opus` assertion it named
143
+ (`model-preselection.md`); a `SKILL.md` duplicate whose first claim's
144
+ real text sat just above the cited line (`run-state-lifecycle-and-
145
+ markers.md`); an `init.ts` comment cited in place of the real
146
+ `installKitFile` call it restates (`install-fence-mechanics.md`); an
147
+ `init.ts` range crossing from one branch of `installKitFile` into
148
+ another branch's own record line (`install-fence-mechanics.md`); an
149
+ `init.ts` range naming the wrong call for an `opencodeEffortLine(...)`
150
+ claim (`model-preselection.md`, cited from two spellings of the same
151
+ path); and a milder range that stopped 1 line short of the parameter its
152
+ own sentence's second half named (`install-fence-mechanics.md`). Widened
153
+ `SIBLING_GUARD_WINDOW` from 10 to 20 once a real case fell just outside
154
+ it (two genuinely distinct `init.ts` notes sharing one message, 18 lines
155
+ apart, both legitimate and now allowlisted per doc); re-triaged every
156
+ additional hit the wider window surfaced against the bundle, fixing or
157
+ allowlisting each with a reason stating what the cited line actually
158
+ says (see `docs/okf/log.md` for the full re-triage and the re-measured
159
+ counts). Added `anchorKey` (first 8 hex chars of a sha256 over the
160
+ finding's own anchor text, computed at test time, never stored literally
161
+ in the array) to every allowlist entry and to the match, plus a test
162
+ asserting every entry matched at least one finding on the current
163
+ bundle, closing a gap where a range-only match would silently exempt any
164
+ future, differently-anchored finding on the same range. Added a fixture
165
+ at the real batch-39 S3 geometry (a literal duplicate citation, its real
166
+ sibling 15 lines away, not the original fixture's 10-line near-miss
167
+ range) asserting both rules' behaviour, and a negative fixture pinning
168
+ that the duplicate-citation rule fires regardless of window, since it is
169
+ a pairing comparison, not a windowed one. Two coverage gaps noted in the
170
+ guard's own comment and here rather than closed this round: a path-less
171
+ continuation citation (`:N-M#"..."`, whose path is implied by the
172
+ preceding citation) never matches the citation regex, so this guard
173
+ cannot see one -- extending the regex to resolve a continuation's
174
+ implied path is a named follow-up; and a citation-shaped string inside a
175
+ fenced ` ``` ` code block is now skipped rather than matched (closing the
176
+ reverse risk of misreading a code sample as a citation), a cheap
177
+ addition alongside the rest of this round's work.
178
+ - Review round 3 of the citation-sibling-drift guard above (task agent-dx
179
+ 9f72ae6d): a second consecutive review round found allowlist entries that
180
+ certified real wrong-sibling drift, so this round changes the mechanism
181
+ rather than only the entries. An allowlist entry now records the
182
+ GEOMETRY it was cleared against -- the target-file line(s) carrying the
183
+ uncited identical anchor text for a rule-(b) entry, the doc line the
184
+ repeat sits on for a rule-(a) one -- and that geometry is part of the
185
+ match, so an entry exempts only the hit it was actually reviewed for: a
186
+ new uncited occurrence next to an already-cleared one, or a repeat that
187
+ moved to another doc line, fails instead of inheriting the old verdict. A
188
+ test re-reads those recorded lines against the current files
189
+ independently of the guard's own output (the anchor text is re-derived
190
+ from the doc's own citation, since the array deliberately stores a hash
191
+ rather than the literal text), so an entry whose situation no longer
192
+ exists goes red instead of silently exempting a different one. The
193
+ free-form `reason` field is replaced by `claim`: one sentence naming what
194
+ the citing sentence describes and why the cited line, rather than the
195
+ uncited sibling, is its evidence, written so a reviewer can falsify it by
196
+ reading exactly the two lines the entry names. Process, recorded in
197
+ `docs/okf/log.md` with each round's classification: an allowlist entry is
198
+ accepted only on an INDEPENDENT review classification of the hit, never
199
+ on the reading of whoever implemented or re-pointed the citation, which
200
+ is how both earlier rounds' wrong verdicts reached a green suite.
201
+ Re-pointed the pair this round's review found (a sentence about the
202
+ dropped-role tier-variant SUB-loop citing the enclosing loop's own note
203
+ range, in two docs) to the sub-loop's own note, dropped their entries,
204
+ and re-triaged every remaining hit at the current window against its
205
+ target file. Also citation-only: a fence-contract citation that stopped
206
+ short of the assertion its sentence names, a run-state citation one line
207
+ short of the sentence it supports, and a slicer-superset citation whose
208
+ sentence's second half is now cited from the test that actually pins it.
209
+ Two more fixtures: rule (b) firing at the guard's window and staying
210
+ silent at the round-1 value of 10 for an uncited occurrence 15 lines
211
+ outside the cited range (the window was previously pinned only
212
+ indirectly, through the no-dead-exemption test), and a doc that ends
213
+ inside an unclosed fence now failing loudly instead of silently dropping
214
+ every citation after the stray delimiter, paired with an assertion that
215
+ every bundle doc yields at least one citation.
216
+ - Review round 4 of the citation-sibling-drift guard above (task agent-dx
217
+ 9f72ae6d): a third consecutive review classified every allowlist entry
218
+ by re-reading the two lines each `claim` names, rather than trusting the
219
+ prior round's verdicts; none certified real drift, but two claims were
220
+ inaccurate and one more citation was mis-paired in a shape the guard
221
+ itself cannot see. Rule (a)'s match compared only the finding's second
222
+ citation line, which a `return true;` mutant of that comparison
223
+ survives, and which also could not tell a two-citation entry's cleared
224
+ repeat from a THIRD, unreviewed repeat sharing the same second line;
225
+ fixed with a dedicated fixture and a repeat-count check. The allowlist
226
+ entry's match and its independent geometry re-check both gained
227
+ `paragraphLine`, the doc line of the finding's own first citation: an
228
+ entry was previously keyed by (doc, kind, real target, range, anchorKey,
229
+ recorded geometry) alone, so the same coincidence recurring in a SECOND,
230
+ unreviewed paragraph of a doc would silently inherit the first
231
+ paragraph's verdict -- exactly the shape one entry was carrying (the
232
+ same `init.ts` range cited, and separately drifting, from two paragraphs
233
+ of `model-preselection.md`); the second paragraph's citation is now
234
+ re-pointed to its own, different evidence instead, so the entry covers
235
+ one paragraph only. The geometry re-check's duplicate-citation branch
236
+ dropped a near-tautological "some citation exists at the recorded line"
237
+ check (true of any citation the extractor produces, by construction) for
238
+ one that reads the group's own citations, sorts them into document
239
+ order, and checks the recorded lines by POSITION -- closing a mutant
240
+ (`if (false)` on the old guard) no existing fixture caught. The bare
241
+ `claim.length > 40` sanity check now also rejects a claim that never
242
+ names one of its own entry's recorded lines, closing the gap that let
243
+ three `subagent-contracts-superset.md` entries carry a long claim that
244
+ never actually pointed at its own geometry. One inaccurate claim
245
+ (`model-preselection.md`) said an uncited line named a "codex-only
246
+ effort field"; it is opencode's own field for a non-Claude-family,
247
+ non-Ollama provider, not a codex field at all, and no anchor exists that
248
+ can widen the citation to cover it under this file's own occurrence-cap
249
+ rule, so the claim was corrected instead. One real mis-pairing:
250
+ `install-fence-mechanics.md` cited the OUTER per-dropped-role loop's
251
+ gate/note for a sentence about the tier-variant SUB-loop, and the
252
+ sub-loop's own gate/note for a sentence about the base-file note --
253
+ swapped, citation-only, no content change. Known limit, unclosed this
254
+ round: neither rule catches a citation that resolves and anchors cleanly
255
+ but simply names the WRONG target -- no duplication, no anchor text
256
+ recurring nearby -- which is exactly the shape this round's real
257
+ mis-pairing was; both citations passed every existing check (including
258
+ this guard) because nothing about either one, read alone or against its
259
+ paragraph's siblings, looks wrong.
260
+ - Citation-sibling-drift guard, continuation-citation coverage (task
261
+ agent-dx b50fd903): closed the guard's own documented coverage gap
262
+ (round 1) that a path-less continuation citation (`:N-M#"..."`, whose
263
+ path is implied by the preceding FULL citation earlier in the same
264
+ paragraph) never matched `ANCHOR_CITATION_RE`, so the guard could not
265
+ see one at all -- the form `model-preselection.md` alone carries four
266
+ of. `extractSiblingGuardCitations` now also matches the anchored,
267
+ path-less tail on its own (`ANCHOR_CONTINUATION_CITATION_RE`, anchor
268
+ group required so a bare `:N-M` digit pair in ordinary prose is never
269
+ mistaken for one) and resolves it against `governingPathByParagraph`,
270
+ the nearest preceding full citation's own `citedPath` in the same
271
+ paragraph -- the same "nearest preceding, same paragraph" BINDING RULE
272
+ okf-kit's own short-form/continuation citations use in
273
+ `citations-resolve.ts`. That mirrors the binding rule only, not the
274
+ grammar (review round 3 correction: this bundle's anchored, path-less
275
+ `:N-M#"..."` form IS backtick-wrapped; an earlier version of this
276
+ bullet said it was not). okf-kit's own `CONT_COLON_RE` still misses it
277
+ because its own closing backtick has to follow the digit range
278
+ immediately, and this form's closing backtick follows the `#"anchor"`
279
+ tail instead; `SHORT_FORM_COLON_RE` misses it too, both because a match
280
+ right after a backtick is skipped and because no serial connective
281
+ ("and", "also", ...) precedes it either -- okf-kit sees these citations
282
+ as nothing at all, not merely as unresolved ones. Two fixtures pin the
283
+ closed gap: a path-less
284
+ continuation duplicating its governing citation's own range and anchor
285
+ is flagged by the duplicate rule (drifted/corrected), and a
286
+ discriminating fixture with two different full citations in one
287
+ paragraph before the continuation, which only passes when the
288
+ continuation binds to the NEARER of the two, not the paragraph's first.
289
+ Run against the current bundle, the newly-visible continuation
290
+ citations in `model-preselection.md` produced no new SIBLING-GUARD
291
+ finding (duplicate-citation/wrong-sibling-anchor), allowlisted or
292
+ otherwise -- see `docs/okf/log.md`. That measurement did not cover
293
+ whether each continuation's own anchor still resolved against its
294
+ target at head; it did not, by the time this round's own CHANGELOG line
295
+ shift landed a few commits later -- see the round-2 follow-up bullet
296
+ below.
297
+ - Citation-sibling-drift guard, okf-kit-porting decision (task agent-dx
298
+ b50fd903, run `.ai/runs/2026-09-08-open-pool-batch44` D-006): the guard
299
+ stays kit-local (this package's own `test/docs-consistency.test.ts`),
300
+ not ported to okf-kit as an opt-in `citations-sibling` check. Who pays:
301
+ kit-local means only this package's own OKF bundle is guarded by it,
302
+ and every other fleet bundle with sibling citations (a paragraph
303
+ repeating, or near-duplicating, one of its own citations) stays
304
+ unguarded until each such bundle's own docs-consistency-style suite
305
+ grows the same check by hand; porting would instead put the maintenance
306
+ on okf-kit's maintainers, who would then carry the rule itself, the
307
+ allowlist shape (recorded geometry, a falsifiable one-sentence claim,
308
+ and the independent-review-classification process this file's own
309
+ allowlist process block already documents) as a public, cross-repo
310
+ contract, and a fleet-wide pin bump on every fix to it. Trigger to
311
+ revisit: a second fleet bundle observed carrying real sibling-citation
312
+ drift in a review pass (not merely plausible in the abstract) reopens
313
+ the port decision.
314
+ - Citation-sibling-drift guard, continuation-citation coverage, review
315
+ round 2 (task agent-dx b50fd903): the round-1 bullet above added
316
+ continuation-citation EXTRACTION but not RESOLUTION at the three
317
+ `matchAll(ANCHOR_CITATION_RE)` sites that check "anchor on last content
318
+ line", "anchor <=3 times file-wide/exactly once in-range", "unanchored
319
+ citation", and "citation stays inside one describe/it/test block" --
320
+ and, separately (review round 3 correction: an earlier version of this
321
+ bullet blamed a same-round CHANGELOG.md line shift for staling
322
+ `src/init.ts`'s own line numbers; false, since editing CHANGELOG.md
323
+ cannot move a different file's lines, and this branch had not touched
324
+ `src/` yet at that point), all four of `model-preselection.md`'s
325
+ continuation citations were already stale at this task's own merge
326
+ base: `src/init.ts` last moved on 2026-09-05 (`60cb546`, task agent-dx
327
+ #184, native Codex routing), and nothing detected the drift, since this
328
+ round-1 bullet's own fix added continuation extraction only, not
329
+ resolution, at the three sites above -- so the round-1 bundle re-run's
330
+ "zero unallowlisted findings" never actually re-checked those four
331
+ anchors' text against `src/init.ts` at head. All four re-pointed
332
+ (citation-only) against `src/init.ts` at head; the fourth (into
333
+ `composeClaudeAgentVariant`'s own call site) needed a freshly-derived
334
+ anchor since its old text no longer occurs there at all (the call now
335
+ takes three arguments, not two). `extractSiblingGuardCitations` is now
336
+ also what the three resolution sites above call, instead of each
337
+ running its own bespoke `matchAll(ANCHOR_CITATION_RE)` loop, so a
338
+ resolved continuation is checked by those properties exactly like a
339
+ full citation is; this is what would have caught the stale
340
+ `model-preselection.md` anchors, had it existed in round 1. Four
341
+ further gaps closed in the same extractor: `governingPathByParagraph`
342
+ now resets (not merely leaves stale) on an unresolved/ambiguous full
343
+ citation, matching okf-kit's own reset behaviour; a continuation-match
344
+ overlap filter now also drops a match whose immediately preceding text
345
+ is path-shaped regardless of file extension, so a citation into an
346
+ extension `ANCHOR_CITATION_RE` does not recognise (a `.toml`, a `.tsx`)
347
+ cannot have its own tail misread as a phantom continuation; the
348
+ left-to-right, nearest-preceding ordering of full and continuation
349
+ matches on one line, and the paragraph-scoping property (a continuation
350
+ never resolves across a paragraph boundary), are now both pinned by
351
+ fixtures rather than only described in a comment; and
352
+ `ANCHOR_CONTINUATION_CITATION_RE`'s anchor alternation (previously a
353
+ hand copy of `ANCHOR_CITATION_RE`'s own group 4 pattern) is now asserted
354
+ to be a substring of it, throwing at module load on drift. Residual,
355
+ unclosed this round, named next to the pre-existing fenced-code-block
356
+ gap in the extractor's own comment: this guard's continuation form
357
+ mirrors okf-kit's short-form BINDING RULE only, not its grammar (see
358
+ the round-1 bullet above), so okf-kit's own `citations-resolve` rule
359
+ still cannot see one of these citations at all; closing that gap means
360
+ either changing okf-kit's own grammar (out of this task's scope) or
361
+ accepting the guard-only coverage as the design. Review round 3 (LOW
362
+ 5): the three resolution sites above inherit the extractor's two other
363
+ latent costs too, since they now call it -- a citation inside a fenced
364
+ code block goes unchecked at those sites as well, and a document ending
365
+ inside an unclosed fence makes them throw, same as this guard -- both
366
+ accepted as the same currently-unused-shape cost, not a new one.
367
+ - Citation-sibling-drift guard, docs/okf/log.md's own citations (task
368
+ agent-dx b50fd903, review round 3, D-037): review round 2 found that
369
+ `log.md` -- excluded from `ANCHOR_OKF_DOCS` and therefore from every
370
+ guard above, and from okf-kit's own citation grammar too -- is read by
371
+ nothing, so citation-shaped historical text written into a log entry
372
+ goes unchecked; round 1 of this task had already removed such text from
373
+ one entry (`0f054d2`) and round 2 wrote the same shape into its own
374
+ entry again. Rather than another round of rephrasing that recurs on the
375
+ next entry, `log.md` gets its own guard in
376
+ `test/docs-consistency.test.ts`: every full, anchored citation it writes
377
+ must still resolve at head (the target exists, the anchor text sits
378
+ somewhere inside the cited range), and it may never carry the bundle's
379
+ path-less continuation form at all, since that form has no
380
+ governing-citation semantics in `log.md` -- nothing resolves a
381
+ continuation written there against anything, so it can only be stale
382
+ prose dressed as a citation. Fixtures pin both rules both ways (a stale
383
+ full citation fails, an unresolvable path fails, a continuation form
384
+ fails even when it would resolve, a clean entry passes); run against
385
+ the current bundle, both checks are clean, and every citation-shaped
386
+ historical value the run reported was rephrased as plain prose, in the
387
+ round-2 entry and in an older entry from task 9f72ae6d. Review round 4
388
+ (D-050) removed this bullet's original hit counts rather than
389
+ correcting them: they were typed by hand and did not match what the
390
+ guard produces (see the round-4 bullet below for the rule and for where
391
+ the live figures live instead).
392
+ The round-2 entry's own false same-round-CHANGELOG-line-shift narration
393
+ is corrected in place (see the review round 3 correction two bullets
394
+ above for the identical fix here); `docs/okf/index.md`'s Maintenance
395
+ section now names the new guard.
396
+ - Citation scanning is paragraph-joined, and the log guard is no longer
397
+ silenceable (task agent-dx b50fd903, review round 4, D-050): both
398
+ citation scanners in `test/docs-consistency.test.ts` matched per
399
+ physical line while every doc in this bundle hard-wraps its prose, so a
400
+ citation whose own text straddled a wrap matched neither regex and was
401
+ invisible to every check built on them -- and re-running the regexes
402
+ over raw document text cannot close that, since the string-anchor
403
+ alternation forbids a newline inside the anchor by construction. Both
404
+ scanners now consume one shared `citationScanParagraphs` helper that
405
+ joins each paragraph's lines the way a hard wrap split them and maps
406
+ every joined offset back to its physical line, so findings, allowlist
407
+ geometry and failure messages still name real doc lines; a wrapped full
408
+ citation with a stale anchor and a wrapped continuation form are each
409
+ pinned by their own fixture. The same helper carries the one fence
410
+ pass, so the `log.md` guard inherits the unbalanced-fence throw the
411
+ round-3 hand copy had left behind (a single stray ``` excused every
412
+ citation after it), and its non-vacuity floor now carries the live
413
+ count in its own computed test name. The unanchored-citation brake and
414
+ the block-straddle collector take their doc set and resolver as
415
+ parameters, like the string-anchor collector already did, so the
416
+ "a resolved continuation survives this collector" property is pinned by
417
+ a synthetic doc set at all three sites instead of by source text alone;
418
+ the brake's examined count is pinned as an exact delta (one added full
419
+ citation raises it by one, one added continuation by one more) rather
420
+ than by a floor a dropped-continuation mutant could sink under. The
421
+ `log.md` resolver rejects a cited path carrying a `..` segment and
422
+ asserts repository containment on its on-disk fallback, and a bare
423
+ basename that also exists at the repository root is reported ambiguous
424
+ with both candidates named instead of silently binding to this
425
+ package's own file; the deeper repo-wide basename ambiguity okf-kit
426
+ reports is still bound unconditionally by `anchorScopeResolve()`'s own
427
+ documented design, named as the residual. Convention this round
428
+ installs (D-050): a log entry or CHANGELOG bullet writes no hand-typed
429
+ count of what a guard found; the live figures are the guards' own
430
+ computed test names, read off a passing run.
431
+
432
+ ### Fixed
433
+
434
+ - Three unpinned properties from the review round 4 lows above
435
+ (task agent-dx 4ece8e1e) now have a discriminating fixture each: the
436
+ `log.md` resolver's on-disk fallback is checked against an existing,
437
+ outside-the-repository absolute path so its own containment conjunct is
438
+ no longer redundant with the `..`-segment rejection; `extractSiblingGuard
439
+ Citations` gets its own wrapped-citation fixture (full citation and
440
+ continuation each straddling a hard line break), independent of the
441
+ `log.md` guard's; and the previously duplicated `PATH_SHAPED_BEFORE_RE`
442
+ path-shaped regex is now one module-scope const both call sites read,
443
+ so the two copies can no longer drift apart.
444
+
10
445
  ## [0.31.0] - 2026-09-07
11
446
 
12
447
  ### Changed
@@ -31,28 +31,56 @@ Rules:
31
31
  - Touch only the files relevant to the assigned task. Respect the
32
32
  allowed_changes and forbidden_changes lists in your task contract.
33
33
  - Add or update tests where appropriate. Run the tests you touched and report
34
- the result honestly; if you could not run them, say why.
34
+ the result honestly; if you could not run them, say why. Cite a coverage
35
+ gate's threshold and pass/fail counts, not a run-specific coverage
36
+ percentage; cite a percentage only together with the exact commit and the
37
+ run count, since branch coverage can vary between runs of the same commit.
35
38
  - When the task assignment names mutation probes to run, run each one and
36
- report it in the `mutation_probes` field of your output (mutant,
37
- verified_applied_via, result, restored_verified); an output missing that
38
- field when probes were named is treated as a misfire, not evidence. When
39
- the assignment names no mutation probes, return `mutation_probes: []`
40
- rather than omitting the field. Each item also carries `replayed`:
41
- `false` for a probe newly introduced this round.
39
+ report it in the `mutation_probes` field of your output (mutant, file,
40
+ anchor, before, after, verified_applied_via, result, expectation,
41
+ reason, restored_verified); an output missing that field when probes
42
+ were named is treated as a misfire, not evidence. `file` and `anchor`
43
+ (a line number or a unique surrounding string) locate the mutant;
44
+ `before` and `after` are the exact text swapped there, so a later round
45
+ can reapply the same edit without guessing instead of only a prose
46
+ description. `expectation` records whether `result` matched what the
47
+ probe was expected to do (`met`) or not (`violated`), independent of
48
+ `result` itself, only alongside a measured `killed` or `survived`
49
+ `result`; it is `not_applicable` otherwise (for example when the mutant
50
+ could not be applied and no `result` was measured). `reason` is free
51
+ text, required when `result` is `not_applicable`, empty otherwise,
52
+ carrying one of two canonical strings that distinguish a non-regression
53
+ from a regression: `no definition recorded` (a prior-round probe
54
+ recorded with only an id, no definition to reapply) and `target text no
55
+ longer present` (a replayed probe whose mutant can no longer be
56
+ applied). When the assignment names no mutation probes, return
57
+ `mutation_probes: []` rather than omitting the field.
58
+ Each item also carries `replayed`: `false` for a probe newly
59
+ introduced this round.
42
60
  - On any round after the task's first, the assignment also names every
43
61
  mutation probe named in an earlier round of this task (on the task's
44
62
  first round there are none), drawn from the run's
45
- `04-implementation-summary.md`. Replay each one, not only this round's
46
- new probes, before returning your report, and report each replayed
47
- probe in `mutation_probes` with the four evidence fields plus
48
- `replayed: true`. A replayed probe whose mutant now survives or can no
49
- longer be applied is a regression signal: report it as such (`result`
50
- `survived` or `not_applicable` with the reason) and resolve it before
51
- the next reviewer spawn.
63
+ `04-implementation-summary.md`, naming each by its mutant definition
64
+ (file, anchor, before, after), not merely by its id; a probe recorded
65
+ with only an id and no definition to reapply cannot be replayed and is
66
+ `not_applicable` (reason: `no definition recorded`), not a regression.
67
+ Replay each one, not only this round's new probes, before returning your
68
+ report, and report each replayed probe in `mutation_probes` with the
69
+ evidence fields plus `replayed: true`. A replayed probe whose
70
+ `expectation` is now `violated`, or which can no longer be applied
71
+ (reason: `target text no longer present`), is the regression signal;
72
+ `result` alone is not: report it as such (`result` `survived` or
73
+ `not_applicable` with the reason) and resolve it before the next
74
+ reviewer spawn.
52
75
  - When a verify runner is available, run it for the checks the acceptance
53
76
  criteria name and report its summary under `tests.executed`; when a
54
77
  mutation-probe runner is available, run the named probes through it and
55
- copy its fields into `mutation_probes`.
78
+ copy its fields into `mutation_probes`; when the runner reports a
79
+ probe's mutant record (`file`, `anchor`, `before`, `after`) separately
80
+ from its result fields (`verified_applied_via`, `result`, `expectation`,
81
+ `reason`, `restored_verified`), take the definition fields from that
82
+ mutant record so the copied report still carries all eleven
83
+ `mutation_probes` sub-fields.
56
84
  - Run every long test, build, or mutation-probe command in the foreground
57
85
  and wait for it to finish before returning. When one foreground call
58
86
  cannot hold it to completion, poll the backgrounded run to completion
@@ -80,6 +108,11 @@ Rules:
80
108
  when the task assignment asked for a commit is treated as a misfire, not
81
109
  evidence. When the task produced no commit, return `commits: []` rather
82
110
  than omitting the field.
111
+ - Populate a non-empty `commits` field by pasting `git log --reverse
112
+ --format=%H <base>..HEAD`; never type or hand-complete commit shas.
113
+ - Verification plans, probe plans, and repeat tallies run in the foreground,
114
+ and the implementer reports their returns in the same turn as the last
115
+ check. A background monitor is no substitute for those returns.
83
116
  - Only write a verification claim (for example "Verified by ...") in a code
84
117
  comment, commit message, or your report for a check you actually ran and
85
118
  measured yourself; never claim a run you did not execute.
@@ -131,8 +164,14 @@ tests:
131
164
  not_executed_reason: ""
132
165
  mutation_probes:
133
166
  - mutant: ""
167
+ file: ""
168
+ anchor: ""
169
+ before: ""
170
+ after: ""
134
171
  verified_applied_via: ""
135
- result: ""
172
+ result: killed | survived | not_applicable
173
+ expectation: met | violated | not_applicable
174
+ reason: ""
136
175
  restored_verified: ""
137
176
  replayed: false | true
138
177
  risks:
@@ -22,6 +22,25 @@ a version. For a recorded original string-list contract, retain the original
22
22
  and `criterion_evidence` fields; keep all existing role output fields. This
23
23
  selection governs the rules and every YAML block below.
24
24
 
25
+ Review method: the orchestrator names `review_method: normal | rigorous |
26
+ adversarial` in every briefing; treat an unnamed method as `rigorous`. The
27
+ three methods are obligation sets, not personas: they define what you must
28
+ read, reproduce, and probe, and how a non-reproducing finding is withdrawn,
29
+ not how skeptical to sound.
30
+
31
+ | Method | Obligations |
32
+ |---|---|
33
+ | `normal` | Read the diff and the spec; run the declared tests once; findings come only from what you read. `normal` adds nothing beyond the obligations already stated in the Check list and the Rules below, and suspends none of them: the empirical-reproduction rule and the GitHub Actions shell replay rule apply under every method. `normal` only means no further independent reproduction beyond what those already require. Fits docs, renames, and batch cosmetics. |
34
+ | `rigorous` (default) | Everything `normal` requires, plus: your own extract of the change, a base-attribution control, classifying every change, and reproducing every empirical claim yourself. `reproduction` and `matches_implementer_claim` are mandatory, as already required below. |
35
+ | `adversarial` | Everything `rigorous` requires, plus: one discriminating probe or negative control per acceptance criterion; an active search of the neighbouring scenario space (environment, install modes, platform, ordering, concurrency); an attempt to break the claimed invariant; and an explicit list of break attempts that failed. |
36
+
37
+ Withdrawal rule (`rigorous` and `adversarial`): a finding that does not
38
+ reproduce on a second attempt with a corrected harness is withdrawn in the
39
+ same round, not carried into the next one, and reported under `withdrawn`
40
+ with the reason; this keeps the method from buying false positives. Emit
41
+ `withdrawn: []` when nothing was withdrawn. Report the method you actually
42
+ applied in `method_applied`.
43
+
25
44
  Check, at minimum:
26
45
 
27
46
  - Acceptance baseline: for a run explicitly adopted as `acceptance-baseline/v1`,
@@ -116,6 +135,10 @@ Rules:
116
135
  shell replay above is a second, explicitly non-probabilistic trigger for
117
136
  the same field: report it in `reproduction` too, with `sample_size:
118
137
  not_applicable` when the replay itself has no meaningful sample size.
138
+ - When citing a coverage gate, cite the threshold and pass/fail counts, not
139
+ a run-specific coverage percentage; cite a percentage only together with
140
+ the exact commit and the run count, since branch coverage can vary
141
+ between runs of the same commit.
119
142
  - When a mutation-probe runner is available in the session, run probes
120
143
  through it instead of editing files by hand, and carry its result fields
121
144
  into your findings and `reproduction`; when a verify runner is available,
@@ -145,4 +168,8 @@ reproduction:
145
168
  sample_size: ""
146
169
  result: ""
147
170
  matches_implementer_claim: matched | mismatched | not_applicable
171
+ method_applied: normal | rigorous | adversarial
172
+ withdrawn:
173
+ - description: ""
174
+ reason: ""
148
175
  ```
@@ -49,6 +49,13 @@ default, not a ritual.
49
49
  orchestrator may review it itself; reserve the reviewer subagent for
50
50
  changes whose risk or size warrants an independent skeptical pass. Either
51
51
  way, review is never skipped.
52
+ - Every reviewer briefing also names a `review_method`: `normal | rigorous |
53
+ adversarial`, an obligation set orthogonal to the effort tier below.
54
+ `adversarial` is the minimum for security judgment, install/deploy
55
+ scripts, hand-edited lockfiles, cross-major overrides, or anything the
56
+ operator flags high-risk; `normal` fits only docs, renames, or batch
57
+ cosmetics; `rigorous` is the default otherwise. Never pair `adversarial`
58
+ with the `-medium` reviewer tier; tiers themselves are unchanged.
52
59
  - When tier variants are installed (manifest `tiers: true`), the orchestrator
53
60
  picks the effort tier per task by complexity and risk, at its own judgment.
54
61
  The unsuffixed default subagent is the normal case; `-high`/`-xhigh` fit
@@ -211,20 +211,34 @@ directory and the subagents.
211
211
  for real, observe the named test fail, restore, re-verify). Hold the
212
212
  implementer's report to the claim-only-what-was-measured rule too: treat any
213
213
  verification claim there that is not backed by a check it actually ran as
214
- unverified. On any round after the task's first, the briefing also names
214
+ unverified. The installed `implementer.md` prompt has the implementer cite
215
+ a coverage gate's threshold and pass/fail counts, not a run-specific
216
+ coverage percentage, citing a percentage only together with the exact
217
+ commit and the run count, since branch coverage can vary between runs of
218
+ the same commit. On any round after the task's first, the briefing also names
215
219
  every mutation probe named in an earlier round of this task (on the
216
220
  task's first round there are none), drawn from the run's
217
- `04-implementation-summary.md`; the implementer replays each one, not
218
- only the round's new probes, before the next reviewer spawn, and
219
- reports each in `mutation_probes` with the four evidence fields plus
220
- `replayed: true`. A replayed probe whose mutant now survives or can no
221
- longer be applied is a regression signal, reported as such (`result`
222
- `survived` or `not_applicable` with the reason) and resolved before the
223
- next reviewer spawn. Record meaningful decisions in
221
+ `04-implementation-summary.md`, naming each by its mutant definition
222
+ (file, anchor, before, after), not merely by its id; a probe recorded
223
+ with only an id and no definition to reapply cannot be replayed and is
224
+ `not_applicable` (reason: `no definition recorded`), not a regression.
225
+ The implementer replays each one, not only the round's new probes,
226
+ before the next reviewer spawn, and reports each in `mutation_probes`
227
+ with the evidence fields plus `replayed: true`. A replayed probe whose
228
+ `expectation` is now `violated`, or which can no longer be applied
229
+ (reason: `target text no longer present`), is the regression signal;
230
+ `result` alone is not: reported as such (`result` `survived` or
231
+ `not_applicable` with the reason) and resolved before the next reviewer
232
+ spawn. Record meaningful decisions in
224
233
  `03-decisions.md` and consolidate evidence in
225
234
  `04-implementation-summary.md`, recording each probe the implementer
226
235
  reports as a row in `04-implementation-summary.md`'s Mutation Probes
227
- subsection, with the round it was named in. For any diff that adds or
236
+ subsection, with the round it was named in. Each row's Before/After
237
+ cells hold a single-line excerpt; when the mutant's actual before/after
238
+ text is multi-line or contains an unescaped `|`, or the mutant is a
239
+ patch/diff rather than a text swap, the full text or diff goes in the
240
+ implementer report or a fenced block placed directly under the table,
241
+ with the row noting where it lives. For any diff that adds or
228
242
  changes a GitHub Actions `run:` step, the installed `implementer.md`
229
243
  prompt requires replaying it locally under the shell the step actually
230
244
  runs, with the expected-success and the expected-failure inputs, before
@@ -251,52 +265,67 @@ directory and the subagents.
251
265
  `reviewer-<tier>` subagents, if any) by the task's complexity and risk, at
252
266
  your own judgment, defaulting to the unsuffixed subagent when unsure; record
253
267
  a non-default tier choice with a one-line reason in `03-decisions.md` when
254
- the task is non-trivial. When the reviewer's environment cannot use version
255
- control to see the diff (for example a policy-gated repository), supply the
256
- diff as a pre-generated file in the briefing instead of expecting the
257
- reviewer to derive it, and have the reviewer report explicitly if it could
258
- only reconstruct the delta some other way, rather than silently reviewing
259
- less than the full change. The reviewer checks spec compliance, architecture
260
- consistency, edge cases, security, test adequacy (including whether new
261
- tests would fail if the change were reverted), and maintainability. Findings
262
- go to `05-review-findings.md`; transfer each finding from the reviewer
263
- output contract into the table's columns as-is, keeping the Severity and
264
- Decision headers unchanged, since those two are what the
265
- orchestrator-workflow completeness reader verifies. Replace the shipped
266
- placeholder/legend row with the transferred findings; for a genuine
267
- zero-findings review, delete that row instead of leaving it in place, since
268
- the completeness reader treats an untouched placeholder row with no finding
269
- rows as the template never having been filled in. When acceptance rests on
270
- empirical or probabilistic evidence (flake rates, benchmarks, "n runs
271
- green", performance/timing numbers), the reviewer must independently
272
- reproduce it — its own runs or measurements, not a re-read of the
273
- implementer's log — and record the method, sample size, and result against
274
- the implementer's claim in the reviewer output contract's `reproduction`
275
- field. This does not apply to deterministic checks (a single test run,
276
- `tsc`, lint): only claims that could vary run to run trigger it. The GitHub
277
- Actions shell replay named in step 6 is a second, explicitly
268
+ the task is non-trivial. Also name `review_method: normal | rigorous |
269
+ adversarial` in the briefing; every briefing names one. Pick it by risk
270
+ class: `adversarial` at minimum for security judgment, install/deploy
271
+ scripts, hand-edited lockfiles, cross-major overrides, or anything the
272
+ operator flags high-risk; `normal` only for docs, renames, or batch
273
+ cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
274
+ never substitutes for it: do not pair `adversarial` with the `-medium`
275
+ reviewer tier, a budget mismatch that names probes without the effort to run
276
+ them; tiers themselves are unchanged by this axis. When the reviewer's
277
+ environment cannot use version control to see the diff (for example a
278
+ policy-gated repository), supply the diff as a pre-generated file in the
279
+ briefing instead of expecting the reviewer to derive it, and have the
280
+ reviewer report explicitly if it could only reconstruct the delta some other
281
+ way, rather than silently reviewing less than the full change. The reviewer
282
+ checks spec compliance, architecture consistency, edge cases, security, test
283
+ adequacy (including whether new tests would fail if the change were
284
+ reverted), and maintainability. Findings go to `05-review-findings.md`;
285
+ transfer each finding from the reviewer output contract into the table's
286
+ columns as-is, keeping the Severity and Decision headers unchanged, since
287
+ those two are what the orchestrator-workflow completeness reader verifies.
288
+ Replace the shipped placeholder/legend row with the transferred findings;
289
+ for a genuine zero-findings review, delete that row instead of leaving it in
290
+ place, since the completeness reader treats an untouched placeholder row
291
+ with no finding rows as the template never having been filled in. When
292
+ acceptance rests on empirical or probabilistic evidence (flake rates,
293
+ benchmarks, "n runs green", performance/timing numbers), the reviewer must
294
+ independently reproduce it — its own runs or measurements, not a re-read of
295
+ the implementer's log — and record the method, sample size, and result
296
+ against the implementer's claim in the reviewer output contract's
297
+ `reproduction` field. This does not apply to deterministic checks (a single
298
+ test run, `tsc`, lint): only claims that could vary run to run trigger it.
299
+ The GitHub Actions shell replay named in step 6 is a second, explicitly
278
300
  non-probabilistic trigger for the same field, with `sample_size:
279
301
  not_applicable` allowed when the replay itself has no meaningful sample
280
- size. A change that deletes or renames an exported identifier, type, config
281
- key, or file is also checked for identifier drift (docs or comments still
282
- describing the old name as current), by the reviewer or by the orchestrator
283
- itself when it reviews a trivial rename per Scaling delegation, using a
284
- connected drift check when one exists. When this is not the task's first
285
- review round, name the round number in the briefing; the reviewer marks each
286
- finding's `recurrence` as `new` or `repeated` against the earlier rounds it
287
- was told about, which is what lets the orchestrator detect the Review-round
288
- escalation budget's trigger (see below) without re-deriving it by hand. When
289
- the implementer's report replays a prior round's mutation probe, the
290
- orchestrator's reviewer briefing names the replayed probes the implementer
291
- reports as killed together with their `mutant` and `verified_applied_via`
292
- values; the reviewer may then skip re-running those. The reviewer output
293
- contract itself is unchanged. Never run mutation probes in place against a
294
- worktree a reviewer subagent is concurrently reviewing; isolate the probe in
295
- a separate worktree or wait until the reviewer has returned before probing
296
- that tree again. For an explicitly adopted v1 run, ask the reviewer to
297
- compare the frozen delegated criteria with the referenced evidence and judge
298
- semantic adequacy, including whether a manual check is actually concrete and
299
- reasoned.
302
+ size. When citing a coverage gate, the installed `reviewer.md` prompt has
303
+ the reviewer cite the threshold and pass/fail counts, not a run-specific
304
+ coverage percentage, citing a percentage only together with the exact commit
305
+ and the run count, since branch coverage can vary between runs of the same
306
+ commit. A change that deletes or renames an exported identifier, type,
307
+ config key, or file is also checked for identifier drift (docs or comments
308
+ still describing the old name as current), by the reviewer or by the
309
+ orchestrator itself when it reviews a trivial rename per Scaling delegation,
310
+ using a connected drift check when one exists. When this is not the task's
311
+ first review round, name the round number in the briefing; the reviewer
312
+ marks each finding's `recurrence` as `new` or `repeated` against the earlier
313
+ rounds it was told about, which is what lets the orchestrator detect the
314
+ Review-round escalation budget's trigger (see below) without re-deriving it
315
+ by hand. When the implementer's report replays a prior round's mutation
316
+ probe, the orchestrator's reviewer briefing names the replayed probes the
317
+ implementer reports as killed together with their mutant definition
318
+ (`file`, `anchor`, `before`, `after`) and `verified_applied_via` value,
319
+ not merely their id; a probe recorded with only an id and no definition
320
+ cannot be skipped this way and is `not_applicable`. The reviewer may
321
+ then skip re-running the ones named by definition.
322
+ The reviewer output contract itself is unchanged. Never run mutation probes
323
+ in place against a worktree a reviewer subagent is concurrently reviewing;
324
+ isolate the probe in a separate worktree or wait until the reviewer has
325
+ returned before probing that tree again. For an explicitly adopted v1 run,
326
+ ask the reviewer to compare the frozen delegated criteria with the
327
+ referenced evidence and judge semantic adequacy, including whether a manual
328
+ check is actually concrete and reasoned.
300
329
  8. **Decide acceptance.** Accept, request fixes, defer, or escalate to the
301
330
  operator. High or critical findings block acceptance until fixed or
302
331
  explicitly waived: critical findings require operator sign-off; high
@@ -304,8 +333,8 @@ directory and the subagents.
304
333
  or critical finding counts as a waiver and follows the same rules. Record
305
334
  all decisions and waivers in `03-decisions.md` and summarize waivers in
306
335
  the Accepted Waivers section of `06-handoff.md`. A reviewer recommendation is not orchestrator acceptance and cannot authorize a critical waiver; only the operator may authorize a critical waiver. For newly created decision records, identify a stable ID, trigger/evidence, decision, accountable authority/source with concrete approval evidence, consequences, and a superseded decision ID when revising a prior decision. Link baseline revisions and waivers to those decision IDs. Established runs retain their recorded decision format; absent fields never create a retroactive blocker. Routine decisions within the delegated contract remain the orchestrator's responsibility; an out-of-scope change requires an operator decision. Markdown records evidence of real authority and never grant it by themselves. Do not accept while a
307
- required baseline criterion in an explicitly adopted v1 run has an open residual; a residual retains its ID
308
- and cannot be converted away. After independent review,
336
+ required baseline criterion in an explicitly adopted v1 run has an open
337
+ residual; a residual retains its ID and cannot be converted away. After independent review,
309
338
  the orchestrator may close a docs-only delta without another reviewer round only
310
339
  when the entire unreviewed delta contains only explanatory
311
340
  documentation, comments, or citations; contains no source- or test-file
@@ -445,8 +474,14 @@ tests:
445
474
  not_executed_reason: ""
446
475
  mutation_probes:
447
476
  - mutant: ""
477
+ file: ""
478
+ anchor: ""
479
+ before: ""
480
+ after: ""
448
481
  verified_applied_via: ""
449
- result: ""
482
+ result: killed | survived | not_applicable
483
+ expectation: met | violated | not_applicable
484
+ reason: ""
450
485
  restored_verified: ""
451
486
  replayed: false | true
452
487
  risks:
@@ -472,19 +507,37 @@ identifies the reviewed artifact and revision, reviewer, method, pass/fail
472
507
  standard, reasoned result, and baseline/criterion identities; it stays manual.
473
508
 
474
509
  When the task assignment names mutation probes to run, the implementer
475
- reports each one in the `mutation_probes` field (mutant,
476
- verified_applied_via, result, restored_verified); when the assignment
477
- names none, it returns `mutation_probes: []` rather than omitting the
478
- field, so 'none asked for' is distinguishable from 'asked for and not
479
- reported'. Each item also carries `replayed`: `false` for a probe newly
480
- introduced this round, `true` for a prior round's probe replayed this
481
- round under the replay rule in step 6. On any round after the task's
482
- first, the implementer replays every probe named in an earlier round of
483
- this task (on the task's first round there are none), not only this
484
- round's new probes, before the next reviewer spawn, reporting each one in
485
- `mutation_probes` alongside the round's new probes. A replayed probe
486
- whose mutant now survives or can no longer be applied is a regression
487
- signal, reported as such and resolved before the next reviewer spawn.
510
+ reports each one in the `mutation_probes` field (mutant, file, anchor,
511
+ before, after, verified_applied_via, result, expectation, reason,
512
+ restored_verified); `file` and `anchor` (a line number or a unique
513
+ surrounding string) locate the mutant, `before` and `after` are the
514
+ exact text swapped there, and `expectation` records whether `result`
515
+ matched what the probe was expected to do (`met`) or not (`violated`),
516
+ independent of `result` itself, only alongside a measured `killed` or
517
+ `survived` `result`; it is `not_applicable` otherwise (for example when
518
+ the mutant could not be applied and no `result` was measured). `reason`
519
+ is free text, required when `result` is `not_applicable`, empty
520
+ otherwise, carrying one of two canonical strings that distinguish a
521
+ non-regression from a regression: `no definition recorded` (a
522
+ prior-round probe recorded with only an id, no definition to reapply)
523
+ and `target text no longer present` (a replayed probe whose mutant can
524
+ no longer be applied). When the assignment names none, it returns
525
+ `mutation_probes: []` rather than
526
+ omitting the field, so 'none asked for' is distinguishable from 'asked
527
+ for and not reported'. Each item also carries `replayed`: `false` for a
528
+ probe newly introduced this round, `true` for a prior round's probe
529
+ replayed this round under the replay rule in step 6. On any round after
530
+ the task's first, the implementer replays every probe named in an
531
+ earlier round of this task (on the task's first round there are none),
532
+ naming each by its mutant definition, not merely by its id, not only
533
+ this round's new probes, before the next reviewer spawn, reporting each
534
+ one in `mutation_probes` alongside the round's new probes. A replayed
535
+ probe whose `expectation` is now `violated`, or which can no longer be
536
+ applied (reason: `target text no longer present`), is the regression
537
+ signal, reported as such and resolved before the next reviewer spawn;
538
+ `result` alone is not a regression signal, and a probe recorded with
539
+ only an id and no definition to reapply is `not_applicable` (reason:
540
+ `no definition recorded`).
488
541
 
489
542
  The `commits` field lists the full sha of every commit the implementer
490
543
  produced on the task branch, in the order produced; when the task
@@ -492,6 +545,12 @@ produced no commit, the implementer returns `commits: []` rather than
492
545
  omitting the field, so 'did not commit' is distinguishable from
493
546
  'forgot to report'.
494
547
 
548
+ For a non-empty `commits` field, the implementer pastes `git log
549
+ --reverse --format=%H <base>..HEAD`; it never types or hand-completes commit
550
+ shas. Verification plans, probe plans, and repeat tallies run in the foreground, and the
551
+ implementer reports their returns in the same turn as the last check. A
552
+ background monitor is no substitute for those returns.
553
+
495
554
  ## Reviewer output contract
496
555
 
497
556
  The output shape remains the same for either selected contract. Compare the
@@ -520,6 +579,10 @@ reproduction:
520
579
  sample_size: ""
521
580
  result: ""
522
581
  matches_implementer_claim: matched | mismatched | not_applicable
582
+ method_applied: normal | rigorous | adversarial
583
+ withdrawn:
584
+ - description: ""
585
+ reason: ""
523
586
  ```
524
587
 
525
588
  `acceptance_recommendation` is mandatory: every reviewer return must set it.
@@ -532,6 +595,15 @@ one that already appeared in an earlier round. On a task's first review
532
595
  round every finding is `new` by definition. This is what feeds the
533
596
  Review-round escalation budget's trigger.
534
597
 
598
+ `method_applied` echoes the `review_method` named in the briefing (see step
599
+ 7); `withdrawn` lists each finding the reviewer proposed and then retracted
600
+ under the withdrawal rule (`rigorous` and `adversarial` only), with its
601
+ reason; emit `withdrawn: []` when nothing was withdrawn. Until a
602
+ grounding-mcp reader parses the marker (tracked as a cross-repo
603
+ follow-up), the orchestrator checks by hand that the return's
604
+ `method_applied` matches the briefing's `review_method`; a mismatch or
605
+ omission is resupplied, not accepted.
606
+
535
607
  ## Task slicer output contract
536
608
 
537
609
  Use this v1 block subject to Contract selection above for every task.
@@ -68,9 +68,15 @@ an optional row cannot stand in for a required criterion.
68
68
 
69
69
  ### Mutation Probes
70
70
 
71
- | Round | Mutant | Verified Applied Via | Result | Restored Verified | Replayed |
72
- |---|---|---|---|---|---|
73
- | <!-- round --> | <!-- mutant --> | <!-- verified_applied_via --> | <!-- result --> | <!-- restored_verified --> | <!-- replayed --> |
71
+ Before/After cells hold a single-line excerpt. When the mutant's actual
72
+ before/after text is multi-line or contains an unescaped `|`, or the mutant
73
+ is a patch/diff rather than a text swap, put the full text or diff in the
74
+ implementer report or a fenced block directly under the table, and note
75
+ where it lives in the row's own cell.
76
+
77
+ | Round | Mutant | File | Anchor | Before | After | Verified Applied Via | Result | Expectation | Reason | Restored Verified | Replayed |
78
+ |---|---|---|---|---|---|---|---|---|---|---|---|
79
+ | <!-- round --> | <!-- mutant --> | <!-- file --> | <!-- anchor --> | <!-- before --> | <!-- after --> | <!-- verified_applied_via --> | <!-- result --> | <!-- expectation --> | <!-- reason --> | <!-- restored_verified --> | <!-- replayed --> |
74
80
 
75
81
  ## Risks / Notes
76
82
 
@@ -4,6 +4,11 @@
4
4
 
5
5
  <!-- Short summary. -->
6
6
 
7
+ <!-- review-method[<round>] = normal|rigorous|adversarial -->
8
+ Method: normal | rigorous | adversarial (the `review_method` named in this
9
+ round's briefing and the `method_applied` the reviewer returned; not parsed
10
+ by the grounding-mcp completeness reader yet).
11
+
7
12
  ## Findings
8
13
 
9
14
  <!-- The Severity and Decision column headers below are load-bearing: the orchestrator-workflow completeness reader locates this table by its header row and verifies unresolved findings from those two columns. Do not rename or drop them. -->
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.31.0",
3
+ "version": "0.33.0",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",