orchestrator-workflow 0.31.0 → 0.33.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +435 -0
- package/assets/agents/implementer.md +55 -16
- package/assets/agents/reviewer.md +27 -0
- package/assets/agents-md-section.md +7 -0
- package/assets/skill/SKILL.md +141 -69
- package/assets/templates/04-implementation-summary.md +9 -3
- package/assets/templates/05-review-findings.md +5 -0
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,441 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.33.0] - 2026-09-12
|
|
11
|
+
|
|
12
|
+
### Changed
|
|
13
|
+
|
|
14
|
+
- Implementer reports now paste non-empty commit lists from `git log --reverse
|
|
15
|
+
--format=%H <base>..HEAD`, and foreground verification/probe plans and
|
|
16
|
+
repeat tallies return in the same turn as the last check; a background
|
|
17
|
+
monitor does not substitute. Anchored by pandora batch48 evidence.
|
|
18
|
+
|
|
19
|
+
- The implementer output contract now enumerates mutation-probe `result` as
|
|
20
|
+
`killed | survived | not_applicable` in both the installed prompt and the
|
|
21
|
+
SKILL.md reference. The docs-consistency guard separately pins each copy's
|
|
22
|
+
complete field block, including that enum, and now recognizes versioned
|
|
23
|
+
parenthesized headings in `see <name> below` forward pointers.
|
|
24
|
+
|
|
25
|
+
- The implementer and reviewer prompts (`assets/agents/implementer.md`,
|
|
26
|
+
`assets/agents/reviewer.md`, mirrored in SKILL.md) now say: cite a
|
|
27
|
+
coverage gate's threshold and pass/fail counts, not a run-specific
|
|
28
|
+
coverage percentage; cite a percentage only together with the exact
|
|
29
|
+
commit and the run count, since branch coverage varies between runs of
|
|
30
|
+
the same commit. Anchored by pandora run
|
|
31
|
+
`.ai/runs/2026-09-11-memory-sync-wipe`.
|
|
32
|
+
- The implementer `mutation_probes` output field (`assets/agents/
|
|
33
|
+
implementer.md`, mirrored in SKILL.md, and the `04-implementation-
|
|
34
|
+
summary.md` template's Mutation Probes table) now carries the mutant's
|
|
35
|
+
definition, not only its label: `file`, `anchor` (a line number or a
|
|
36
|
+
unique string), `before`, and `after` alongside the existing
|
|
37
|
+
`verified_applied_via`, `result`, `restored_verified`, and `replayed`
|
|
38
|
+
fields, so a later round can mechanically reapply the same edit instead
|
|
39
|
+
of only reading a prose description. Added `expectation: met | violated
|
|
40
|
+
| not_applicable` beside `result`, reporting whether the probe's
|
|
41
|
+
`result` matched its `--expect`, scoped to a measured `killed` or
|
|
42
|
+
`survived` `result` (a routine negative-control probe now reports
|
|
43
|
+
`result: survived, expectation: met`, which is not a regression) and
|
|
44
|
+
`not_applicable` otherwise (for example when the mutant could not be
|
|
45
|
+
applied and no `result` was measured). Added an eleventh sub-field,
|
|
46
|
+
`reason`: free text, required when `result` is `not_applicable`, empty
|
|
47
|
+
otherwise, carrying one of two canonical strings that distinguish a
|
|
48
|
+
non-regression from a regression: `no definition recorded` (a
|
|
49
|
+
prior-round probe recorded with only an id, no definition to reapply)
|
|
50
|
+
and `target text no longer present` (a replayed probe whose mutant can
|
|
51
|
+
no longer be applied). The fix-round replay rule now names a
|
|
52
|
+
probe to replay by its definition, not merely by its id, and treats a
|
|
53
|
+
replayed probe whose `expectation` is now `violated` (or which can no
|
|
54
|
+
longer be applied) as the regression signal, not `result` alone.
|
|
55
|
+
Anchored by pandora run `.ai/runs/2026-09-11-memory-sync-wipe` and the
|
|
56
|
+
agent-primitives `probe` result/expectation split (0.3.0).
|
|
57
|
+
|
|
58
|
+
## [0.32.0] - 2026-09-11
|
|
59
|
+
|
|
60
|
+
### Added
|
|
61
|
+
|
|
62
|
+
- A review-method axis, orthogonal to the effort tier: every reviewer
|
|
63
|
+
briefing now names `review_method: normal | rigorous | adversarial`
|
|
64
|
+
(`assets/agents/reviewer.md`, SKILL.md step 7, the kit-fence
|
|
65
|
+
Scaling-delegation text). The three methods are obligation sets, not
|
|
66
|
+
personas: `normal` reads the diff and spec and runs the declared tests
|
|
67
|
+
once, for docs/renames/batch cosmetics; `rigorous` (the default when a
|
|
68
|
+
briefing names none) adds an independent extract, a base-attribution
|
|
69
|
+
control, and mandatory reproduction of every empirical claim (the
|
|
70
|
+
pre-existing `reproduction`/`matches_implementer_claim` requirement);
|
|
71
|
+
`adversarial` adds one discriminating probe or negative control per
|
|
72
|
+
acceptance criterion, an active search of the neighbouring scenario
|
|
73
|
+
space, an attempt to break the claimed invariant, and a list of break
|
|
74
|
+
attempts that failed. `adversarial` and `rigorous` both carry a
|
|
75
|
+
withdrawal rule: a finding that does not reproduce on a second attempt
|
|
76
|
+
with a corrected harness is withdrawn in the same round and reported
|
|
77
|
+
under a `withdrawn` list with the reason, so the method cannot buy false
|
|
78
|
+
positives; emit `withdrawn: []` when nothing was withdrawn. The reviewer
|
|
79
|
+
output contract (both copies, `reviewer.md` and SKILL.md) gains
|
|
80
|
+
`method_applied` and `withdrawn`; until the grounding-mcp reader parses
|
|
81
|
+
the marker (agent-grounding follow-up 5df7b809), the orchestrator checks
|
|
82
|
+
by hand that the return's `method_applied` matches the briefing's
|
|
83
|
+
`review_method`, resupplying a mismatch or omission rather than
|
|
84
|
+
accepting it.
|
|
85
|
+
`assets/templates/05-review-findings.md` gains a `Method` line per review
|
|
86
|
+
round, outside the pinned Findings table
|
|
87
|
+
(`test/template-markers.test.ts` pins it); the grounding-mcp
|
|
88
|
+
completeness reader does not parse it yet, tracked as a follow-up in the
|
|
89
|
+
agent-grounding repo. SKILL.md's selection rule: `adversarial` at
|
|
90
|
+
minimum for security judgment, install/deploy scripts, hand-edited
|
|
91
|
+
lockfiles, cross-major overrides, or anything the operator flags
|
|
92
|
+
high-risk; `normal` only for docs, renames, or batch cosmetics;
|
|
93
|
+
`rigorous` otherwise; never `adversarial` on the `-medium` reviewer tier
|
|
94
|
+
(budget mismatch). Tiers themselves are unchanged. Anchored by the
|
|
95
|
+
pandora run `2026-09-11-cve-sweep`: reviews R1, R5, R11, R13 and R14
|
|
96
|
+
found the Critical and High findings by probing; R4, R6 and R9 on the
|
|
97
|
+
default method returned no findings and R8 only two informational lows;
|
|
98
|
+
R5 withdrew a harness artefact. The selection rule stays advisory, not
|
|
99
|
+
an AGENTS.md rule, until an A/B (same tasks, run once under `rigorous`
|
|
100
|
+
and once under `adversarial`, counting real Critical/High findings and
|
|
101
|
+
findings withdrawn) is recorded.
|
|
102
|
+
|
|
103
|
+
- A citation-sibling-drift guard (`test/docs-consistency.test.ts`, next to
|
|
104
|
+
the existing anchor-load-bearing checks) catches a citation that resolves
|
|
105
|
+
and anchors correctly on its own but names the wrong sibling among a run
|
|
106
|
+
of near-identical citations, a class neither okf-kit's `citations-resolve`
|
|
107
|
+
rule nor the local anchor guards can see, because both check an anchor
|
|
108
|
+
only inside its own cited range. Two rules, applied per paragraph: (a) the
|
|
109
|
+
same `file:range#anchor` cited twice in one paragraph, unless allowlisted
|
|
110
|
+
with a reason; (b) a string anchor's text also occurring, uncited, at
|
|
111
|
+
another line of the same target within a 20-line window (widened from 10
|
|
112
|
+
in review round 2, see below) while the paragraph cites a sibling range
|
|
113
|
+
of that file, unless allowlisted. Fixtures
|
|
114
|
+
reproduce three review findings that shared this shape and were the
|
|
115
|
+
motivation for the guard: three sibling `it`-block citations collapsing
|
|
116
|
+
onto one range twice, the third never cited; three per-harness bullet
|
|
117
|
+
citations doing the same; two logically distinct assertions collapsing
|
|
118
|
+
onto one shared range and anchor text, the second's own line never cited.
|
|
119
|
+
Run over the current bundle, every hit either rule reports is read
|
|
120
|
+
against its own target file and the citing paragraph, then fixed or
|
|
121
|
+
allowlisted; the measured per-rule hit counts live in `docs/okf/log.md`
|
|
122
|
+
with the classification that produced them, not here, so the two sites
|
|
123
|
+
cannot drift apart. The recurring coincidence shapes are a doc-wide
|
|
124
|
+
"topic sentence, then repeat as the closing list item" convention for
|
|
125
|
+
rule (a), and a short or common token -- a keyword, a mirrored field on
|
|
126
|
+
twin interfaces, a comment restating a literal, a test-assertion idiom on
|
|
127
|
+
an adjacent line, a reused local variable name -- recurring near a real
|
|
128
|
+
citation for rule (b). Lives in this file rather than as an okf-kit rule
|
|
129
|
+
because this
|
|
130
|
+
suite already runs on every PR while an okf-kit rule needs a release and
|
|
131
|
+
a fleet pin bump first; an opt-in okf-kit rule for the same class is a
|
|
132
|
+
named follow-up candidate once this guard has proven itself.
|
|
133
|
+
- Review round 2 of the citation-sibling-drift guard above (task agent-dx
|
|
134
|
+
9f72ae6d): the round-1 review found 7 of the round's 18 raw hits were not
|
|
135
|
+
coincidental at all -- real mis-pointed citations that got allowlisted
|
|
136
|
+
instead of fixed, because the round-1 pass classified every hit by range
|
|
137
|
+
and reason without re-deriving each cited claim's real evidence line by
|
|
138
|
+
line. All seven re-pointed to their real evidence, citation-only, no
|
|
139
|
+
content changes: a duplicate `docs-consistency.test.ts` self-citation
|
|
140
|
+
whose second claim's real test sat 26 lines below the first
|
|
141
|
+
(`subagent-contracts-superset.md`); an `init.test.ts` range that stopped
|
|
142
|
+
6 lines short of the `model: opus` assertion it named
|
|
143
|
+
(`model-preselection.md`); a `SKILL.md` duplicate whose first claim's
|
|
144
|
+
real text sat just above the cited line (`run-state-lifecycle-and-
|
|
145
|
+
markers.md`); an `init.ts` comment cited in place of the real
|
|
146
|
+
`installKitFile` call it restates (`install-fence-mechanics.md`); an
|
|
147
|
+
`init.ts` range crossing from one branch of `installKitFile` into
|
|
148
|
+
another branch's own record line (`install-fence-mechanics.md`); an
|
|
149
|
+
`init.ts` range naming the wrong call for an `opencodeEffortLine(...)`
|
|
150
|
+
claim (`model-preselection.md`, cited from two spellings of the same
|
|
151
|
+
path); and a milder range that stopped 1 line short of the parameter its
|
|
152
|
+
own sentence's second half named (`install-fence-mechanics.md`). Widened
|
|
153
|
+
`SIBLING_GUARD_WINDOW` from 10 to 20 once a real case fell just outside
|
|
154
|
+
it (two genuinely distinct `init.ts` notes sharing one message, 18 lines
|
|
155
|
+
apart, both legitimate and now allowlisted per doc); re-triaged every
|
|
156
|
+
additional hit the wider window surfaced against the bundle, fixing or
|
|
157
|
+
allowlisting each with a reason stating what the cited line actually
|
|
158
|
+
says (see `docs/okf/log.md` for the full re-triage and the re-measured
|
|
159
|
+
counts). Added `anchorKey` (first 8 hex chars of a sha256 over the
|
|
160
|
+
finding's own anchor text, computed at test time, never stored literally
|
|
161
|
+
in the array) to every allowlist entry and to the match, plus a test
|
|
162
|
+
asserting every entry matched at least one finding on the current
|
|
163
|
+
bundle, closing a gap where a range-only match would silently exempt any
|
|
164
|
+
future, differently-anchored finding on the same range. Added a fixture
|
|
165
|
+
at the real batch-39 S3 geometry (a literal duplicate citation, its real
|
|
166
|
+
sibling 15 lines away, not the original fixture's 10-line near-miss
|
|
167
|
+
range) asserting both rules' behaviour, and a negative fixture pinning
|
|
168
|
+
that the duplicate-citation rule fires regardless of window, since it is
|
|
169
|
+
a pairing comparison, not a windowed one. Two coverage gaps noted in the
|
|
170
|
+
guard's own comment and here rather than closed this round: a path-less
|
|
171
|
+
continuation citation (`:N-M#"..."`, whose path is implied by the
|
|
172
|
+
preceding citation) never matches the citation regex, so this guard
|
|
173
|
+
cannot see one -- extending the regex to resolve a continuation's
|
|
174
|
+
implied path is a named follow-up; and a citation-shaped string inside a
|
|
175
|
+
fenced ` ``` ` code block is now skipped rather than matched (closing the
|
|
176
|
+
reverse risk of misreading a code sample as a citation), a cheap
|
|
177
|
+
addition alongside the rest of this round's work.
|
|
178
|
+
- Review round 3 of the citation-sibling-drift guard above (task agent-dx
|
|
179
|
+
9f72ae6d): a second consecutive review round found allowlist entries that
|
|
180
|
+
certified real wrong-sibling drift, so this round changes the mechanism
|
|
181
|
+
rather than only the entries. An allowlist entry now records the
|
|
182
|
+
GEOMETRY it was cleared against -- the target-file line(s) carrying the
|
|
183
|
+
uncited identical anchor text for a rule-(b) entry, the doc line the
|
|
184
|
+
repeat sits on for a rule-(a) one -- and that geometry is part of the
|
|
185
|
+
match, so an entry exempts only the hit it was actually reviewed for: a
|
|
186
|
+
new uncited occurrence next to an already-cleared one, or a repeat that
|
|
187
|
+
moved to another doc line, fails instead of inheriting the old verdict. A
|
|
188
|
+
test re-reads those recorded lines against the current files
|
|
189
|
+
independently of the guard's own output (the anchor text is re-derived
|
|
190
|
+
from the doc's own citation, since the array deliberately stores a hash
|
|
191
|
+
rather than the literal text), so an entry whose situation no longer
|
|
192
|
+
exists goes red instead of silently exempting a different one. The
|
|
193
|
+
free-form `reason` field is replaced by `claim`: one sentence naming what
|
|
194
|
+
the citing sentence describes and why the cited line, rather than the
|
|
195
|
+
uncited sibling, is its evidence, written so a reviewer can falsify it by
|
|
196
|
+
reading exactly the two lines the entry names. Process, recorded in
|
|
197
|
+
`docs/okf/log.md` with each round's classification: an allowlist entry is
|
|
198
|
+
accepted only on an INDEPENDENT review classification of the hit, never
|
|
199
|
+
on the reading of whoever implemented or re-pointed the citation, which
|
|
200
|
+
is how both earlier rounds' wrong verdicts reached a green suite.
|
|
201
|
+
Re-pointed the pair this round's review found (a sentence about the
|
|
202
|
+
dropped-role tier-variant SUB-loop citing the enclosing loop's own note
|
|
203
|
+
range, in two docs) to the sub-loop's own note, dropped their entries,
|
|
204
|
+
and re-triaged every remaining hit at the current window against its
|
|
205
|
+
target file. Also citation-only: a fence-contract citation that stopped
|
|
206
|
+
short of the assertion its sentence names, a run-state citation one line
|
|
207
|
+
short of the sentence it supports, and a slicer-superset citation whose
|
|
208
|
+
sentence's second half is now cited from the test that actually pins it.
|
|
209
|
+
Two more fixtures: rule (b) firing at the guard's window and staying
|
|
210
|
+
silent at the round-1 value of 10 for an uncited occurrence 15 lines
|
|
211
|
+
outside the cited range (the window was previously pinned only
|
|
212
|
+
indirectly, through the no-dead-exemption test), and a doc that ends
|
|
213
|
+
inside an unclosed fence now failing loudly instead of silently dropping
|
|
214
|
+
every citation after the stray delimiter, paired with an assertion that
|
|
215
|
+
every bundle doc yields at least one citation.
|
|
216
|
+
- Review round 4 of the citation-sibling-drift guard above (task agent-dx
|
|
217
|
+
9f72ae6d): a third consecutive review classified every allowlist entry
|
|
218
|
+
by re-reading the two lines each `claim` names, rather than trusting the
|
|
219
|
+
prior round's verdicts; none certified real drift, but two claims were
|
|
220
|
+
inaccurate and one more citation was mis-paired in a shape the guard
|
|
221
|
+
itself cannot see. Rule (a)'s match compared only the finding's second
|
|
222
|
+
citation line, which a `return true;` mutant of that comparison
|
|
223
|
+
survives, and which also could not tell a two-citation entry's cleared
|
|
224
|
+
repeat from a THIRD, unreviewed repeat sharing the same second line;
|
|
225
|
+
fixed with a dedicated fixture and a repeat-count check. The allowlist
|
|
226
|
+
entry's match and its independent geometry re-check both gained
|
|
227
|
+
`paragraphLine`, the doc line of the finding's own first citation: an
|
|
228
|
+
entry was previously keyed by (doc, kind, real target, range, anchorKey,
|
|
229
|
+
recorded geometry) alone, so the same coincidence recurring in a SECOND,
|
|
230
|
+
unreviewed paragraph of a doc would silently inherit the first
|
|
231
|
+
paragraph's verdict -- exactly the shape one entry was carrying (the
|
|
232
|
+
same `init.ts` range cited, and separately drifting, from two paragraphs
|
|
233
|
+
of `model-preselection.md`); the second paragraph's citation is now
|
|
234
|
+
re-pointed to its own, different evidence instead, so the entry covers
|
|
235
|
+
one paragraph only. The geometry re-check's duplicate-citation branch
|
|
236
|
+
dropped a near-tautological "some citation exists at the recorded line"
|
|
237
|
+
check (true of any citation the extractor produces, by construction) for
|
|
238
|
+
one that reads the group's own citations, sorts them into document
|
|
239
|
+
order, and checks the recorded lines by POSITION -- closing a mutant
|
|
240
|
+
(`if (false)` on the old guard) no existing fixture caught. The bare
|
|
241
|
+
`claim.length > 40` sanity check now also rejects a claim that never
|
|
242
|
+
names one of its own entry's recorded lines, closing the gap that let
|
|
243
|
+
three `subagent-contracts-superset.md` entries carry a long claim that
|
|
244
|
+
never actually pointed at its own geometry. One inaccurate claim
|
|
245
|
+
(`model-preselection.md`) said an uncited line named a "codex-only
|
|
246
|
+
effort field"; it is opencode's own field for a non-Claude-family,
|
|
247
|
+
non-Ollama provider, not a codex field at all, and no anchor exists that
|
|
248
|
+
can widen the citation to cover it under this file's own occurrence-cap
|
|
249
|
+
rule, so the claim was corrected instead. One real mis-pairing:
|
|
250
|
+
`install-fence-mechanics.md` cited the OUTER per-dropped-role loop's
|
|
251
|
+
gate/note for a sentence about the tier-variant SUB-loop, and the
|
|
252
|
+
sub-loop's own gate/note for a sentence about the base-file note --
|
|
253
|
+
swapped, citation-only, no content change. Known limit, unclosed this
|
|
254
|
+
round: neither rule catches a citation that resolves and anchors cleanly
|
|
255
|
+
but simply names the WRONG target -- no duplication, no anchor text
|
|
256
|
+
recurring nearby -- which is exactly the shape this round's real
|
|
257
|
+
mis-pairing was; both citations passed every existing check (including
|
|
258
|
+
this guard) because nothing about either one, read alone or against its
|
|
259
|
+
paragraph's siblings, looks wrong.
|
|
260
|
+
- Citation-sibling-drift guard, continuation-citation coverage (task
|
|
261
|
+
agent-dx b50fd903): closed the guard's own documented coverage gap
|
|
262
|
+
(round 1) that a path-less continuation citation (`:N-M#"..."`, whose
|
|
263
|
+
path is implied by the preceding FULL citation earlier in the same
|
|
264
|
+
paragraph) never matched `ANCHOR_CITATION_RE`, so the guard could not
|
|
265
|
+
see one at all -- the form `model-preselection.md` alone carries four
|
|
266
|
+
of. `extractSiblingGuardCitations` now also matches the anchored,
|
|
267
|
+
path-less tail on its own (`ANCHOR_CONTINUATION_CITATION_RE`, anchor
|
|
268
|
+
group required so a bare `:N-M` digit pair in ordinary prose is never
|
|
269
|
+
mistaken for one) and resolves it against `governingPathByParagraph`,
|
|
270
|
+
the nearest preceding full citation's own `citedPath` in the same
|
|
271
|
+
paragraph -- the same "nearest preceding, same paragraph" BINDING RULE
|
|
272
|
+
okf-kit's own short-form/continuation citations use in
|
|
273
|
+
`citations-resolve.ts`. That mirrors the binding rule only, not the
|
|
274
|
+
grammar (review round 3 correction: this bundle's anchored, path-less
|
|
275
|
+
`:N-M#"..."` form IS backtick-wrapped; an earlier version of this
|
|
276
|
+
bullet said it was not). okf-kit's own `CONT_COLON_RE` still misses it
|
|
277
|
+
because its own closing backtick has to follow the digit range
|
|
278
|
+
immediately, and this form's closing backtick follows the `#"anchor"`
|
|
279
|
+
tail instead; `SHORT_FORM_COLON_RE` misses it too, both because a match
|
|
280
|
+
right after a backtick is skipped and because no serial connective
|
|
281
|
+
("and", "also", ...) precedes it either -- okf-kit sees these citations
|
|
282
|
+
as nothing at all, not merely as unresolved ones. Two fixtures pin the
|
|
283
|
+
closed gap: a path-less
|
|
284
|
+
continuation duplicating its governing citation's own range and anchor
|
|
285
|
+
is flagged by the duplicate rule (drifted/corrected), and a
|
|
286
|
+
discriminating fixture with two different full citations in one
|
|
287
|
+
paragraph before the continuation, which only passes when the
|
|
288
|
+
continuation binds to the NEARER of the two, not the paragraph's first.
|
|
289
|
+
Run against the current bundle, the newly-visible continuation
|
|
290
|
+
citations in `model-preselection.md` produced no new SIBLING-GUARD
|
|
291
|
+
finding (duplicate-citation/wrong-sibling-anchor), allowlisted or
|
|
292
|
+
otherwise -- see `docs/okf/log.md`. That measurement did not cover
|
|
293
|
+
whether each continuation's own anchor still resolved against its
|
|
294
|
+
target at head; it did not, by the time this round's own CHANGELOG line
|
|
295
|
+
shift landed a few commits later -- see the round-2 follow-up bullet
|
|
296
|
+
below.
|
|
297
|
+
- Citation-sibling-drift guard, okf-kit-porting decision (task agent-dx
|
|
298
|
+
b50fd903, run `.ai/runs/2026-09-08-open-pool-batch44` D-006): the guard
|
|
299
|
+
stays kit-local (this package's own `test/docs-consistency.test.ts`),
|
|
300
|
+
not ported to okf-kit as an opt-in `citations-sibling` check. Who pays:
|
|
301
|
+
kit-local means only this package's own OKF bundle is guarded by it,
|
|
302
|
+
and every other fleet bundle with sibling citations (a paragraph
|
|
303
|
+
repeating, or near-duplicating, one of its own citations) stays
|
|
304
|
+
unguarded until each such bundle's own docs-consistency-style suite
|
|
305
|
+
grows the same check by hand; porting would instead put the maintenance
|
|
306
|
+
on okf-kit's maintainers, who would then carry the rule itself, the
|
|
307
|
+
allowlist shape (recorded geometry, a falsifiable one-sentence claim,
|
|
308
|
+
and the independent-review-classification process this file's own
|
|
309
|
+
allowlist process block already documents) as a public, cross-repo
|
|
310
|
+
contract, and a fleet-wide pin bump on every fix to it. Trigger to
|
|
311
|
+
revisit: a second fleet bundle observed carrying real sibling-citation
|
|
312
|
+
drift in a review pass (not merely plausible in the abstract) reopens
|
|
313
|
+
the port decision.
|
|
314
|
+
- Citation-sibling-drift guard, continuation-citation coverage, review
|
|
315
|
+
round 2 (task agent-dx b50fd903): the round-1 bullet above added
|
|
316
|
+
continuation-citation EXTRACTION but not RESOLUTION at the three
|
|
317
|
+
`matchAll(ANCHOR_CITATION_RE)` sites that check "anchor on last content
|
|
318
|
+
line", "anchor <=3 times file-wide/exactly once in-range", "unanchored
|
|
319
|
+
citation", and "citation stays inside one describe/it/test block" --
|
|
320
|
+
and, separately (review round 3 correction: an earlier version of this
|
|
321
|
+
bullet blamed a same-round CHANGELOG.md line shift for staling
|
|
322
|
+
`src/init.ts`'s own line numbers; false, since editing CHANGELOG.md
|
|
323
|
+
cannot move a different file's lines, and this branch had not touched
|
|
324
|
+
`src/` yet at that point), all four of `model-preselection.md`'s
|
|
325
|
+
continuation citations were already stale at this task's own merge
|
|
326
|
+
base: `src/init.ts` last moved on 2026-09-05 (`60cb546`, task agent-dx
|
|
327
|
+
#184, native Codex routing), and nothing detected the drift, since this
|
|
328
|
+
round-1 bullet's own fix added continuation extraction only, not
|
|
329
|
+
resolution, at the three sites above -- so the round-1 bundle re-run's
|
|
330
|
+
"zero unallowlisted findings" never actually re-checked those four
|
|
331
|
+
anchors' text against `src/init.ts` at head. All four re-pointed
|
|
332
|
+
(citation-only) against `src/init.ts` at head; the fourth (into
|
|
333
|
+
`composeClaudeAgentVariant`'s own call site) needed a freshly-derived
|
|
334
|
+
anchor since its old text no longer occurs there at all (the call now
|
|
335
|
+
takes three arguments, not two). `extractSiblingGuardCitations` is now
|
|
336
|
+
also what the three resolution sites above call, instead of each
|
|
337
|
+
running its own bespoke `matchAll(ANCHOR_CITATION_RE)` loop, so a
|
|
338
|
+
resolved continuation is checked by those properties exactly like a
|
|
339
|
+
full citation is; this is what would have caught the stale
|
|
340
|
+
`model-preselection.md` anchors, had it existed in round 1. Four
|
|
341
|
+
further gaps closed in the same extractor: `governingPathByParagraph`
|
|
342
|
+
now resets (not merely leaves stale) on an unresolved/ambiguous full
|
|
343
|
+
citation, matching okf-kit's own reset behaviour; a continuation-match
|
|
344
|
+
overlap filter now also drops a match whose immediately preceding text
|
|
345
|
+
is path-shaped regardless of file extension, so a citation into an
|
|
346
|
+
extension `ANCHOR_CITATION_RE` does not recognise (a `.toml`, a `.tsx`)
|
|
347
|
+
cannot have its own tail misread as a phantom continuation; the
|
|
348
|
+
left-to-right, nearest-preceding ordering of full and continuation
|
|
349
|
+
matches on one line, and the paragraph-scoping property (a continuation
|
|
350
|
+
never resolves across a paragraph boundary), are now both pinned by
|
|
351
|
+
fixtures rather than only described in a comment; and
|
|
352
|
+
`ANCHOR_CONTINUATION_CITATION_RE`'s anchor alternation (previously a
|
|
353
|
+
hand copy of `ANCHOR_CITATION_RE`'s own group 4 pattern) is now asserted
|
|
354
|
+
to be a substring of it, throwing at module load on drift. Residual,
|
|
355
|
+
unclosed this round, named next to the pre-existing fenced-code-block
|
|
356
|
+
gap in the extractor's own comment: this guard's continuation form
|
|
357
|
+
mirrors okf-kit's short-form BINDING RULE only, not its grammar (see
|
|
358
|
+
the round-1 bullet above), so okf-kit's own `citations-resolve` rule
|
|
359
|
+
still cannot see one of these citations at all; closing that gap means
|
|
360
|
+
either changing okf-kit's own grammar (out of this task's scope) or
|
|
361
|
+
accepting the guard-only coverage as the design. Review round 3 (LOW
|
|
362
|
+
5): the three resolution sites above inherit the extractor's two other
|
|
363
|
+
latent costs too, since they now call it -- a citation inside a fenced
|
|
364
|
+
code block goes unchecked at those sites as well, and a document ending
|
|
365
|
+
inside an unclosed fence makes them throw, same as this guard -- both
|
|
366
|
+
accepted as the same currently-unused-shape cost, not a new one.
|
|
367
|
+
- Citation-sibling-drift guard, docs/okf/log.md's own citations (task
|
|
368
|
+
agent-dx b50fd903, review round 3, D-037): review round 2 found that
|
|
369
|
+
`log.md` -- excluded from `ANCHOR_OKF_DOCS` and therefore from every
|
|
370
|
+
guard above, and from okf-kit's own citation grammar too -- is read by
|
|
371
|
+
nothing, so citation-shaped historical text written into a log entry
|
|
372
|
+
goes unchecked; round 1 of this task had already removed such text from
|
|
373
|
+
one entry (`0f054d2`) and round 2 wrote the same shape into its own
|
|
374
|
+
entry again. Rather than another round of rephrasing that recurs on the
|
|
375
|
+
next entry, `log.md` gets its own guard in
|
|
376
|
+
`test/docs-consistency.test.ts`: every full, anchored citation it writes
|
|
377
|
+
must still resolve at head (the target exists, the anchor text sits
|
|
378
|
+
somewhere inside the cited range), and it may never carry the bundle's
|
|
379
|
+
path-less continuation form at all, since that form has no
|
|
380
|
+
governing-citation semantics in `log.md` -- nothing resolves a
|
|
381
|
+
continuation written there against anything, so it can only be stale
|
|
382
|
+
prose dressed as a citation. Fixtures pin both rules both ways (a stale
|
|
383
|
+
full citation fails, an unresolvable path fails, a continuation form
|
|
384
|
+
fails even when it would resolve, a clean entry passes); run against
|
|
385
|
+
the current bundle, both checks are clean, and every citation-shaped
|
|
386
|
+
historical value the run reported was rephrased as plain prose, in the
|
|
387
|
+
round-2 entry and in an older entry from task 9f72ae6d. Review round 4
|
|
388
|
+
(D-050) removed this bullet's original hit counts rather than
|
|
389
|
+
correcting them: they were typed by hand and did not match what the
|
|
390
|
+
guard produces (see the round-4 bullet below for the rule and for where
|
|
391
|
+
the live figures live instead).
|
|
392
|
+
The round-2 entry's own false same-round-CHANGELOG-line-shift narration
|
|
393
|
+
is corrected in place (see the review round 3 correction two bullets
|
|
394
|
+
above for the identical fix here); `docs/okf/index.md`'s Maintenance
|
|
395
|
+
section now names the new guard.
|
|
396
|
+
- Citation scanning is paragraph-joined, and the log guard is no longer
|
|
397
|
+
silenceable (task agent-dx b50fd903, review round 4, D-050): both
|
|
398
|
+
citation scanners in `test/docs-consistency.test.ts` matched per
|
|
399
|
+
physical line while every doc in this bundle hard-wraps its prose, so a
|
|
400
|
+
citation whose own text straddled a wrap matched neither regex and was
|
|
401
|
+
invisible to every check built on them -- and re-running the regexes
|
|
402
|
+
over raw document text cannot close that, since the string-anchor
|
|
403
|
+
alternation forbids a newline inside the anchor by construction. Both
|
|
404
|
+
scanners now consume one shared `citationScanParagraphs` helper that
|
|
405
|
+
joins each paragraph's lines the way a hard wrap split them and maps
|
|
406
|
+
every joined offset back to its physical line, so findings, allowlist
|
|
407
|
+
geometry and failure messages still name real doc lines; a wrapped full
|
|
408
|
+
citation with a stale anchor and a wrapped continuation form are each
|
|
409
|
+
pinned by their own fixture. The same helper carries the one fence
|
|
410
|
+
pass, so the `log.md` guard inherits the unbalanced-fence throw the
|
|
411
|
+
round-3 hand copy had left behind (a single stray ``` excused every
|
|
412
|
+
citation after it), and its non-vacuity floor now carries the live
|
|
413
|
+
count in its own computed test name. The unanchored-citation brake and
|
|
414
|
+
the block-straddle collector take their doc set and resolver as
|
|
415
|
+
parameters, like the string-anchor collector already did, so the
|
|
416
|
+
"a resolved continuation survives this collector" property is pinned by
|
|
417
|
+
a synthetic doc set at all three sites instead of by source text alone;
|
|
418
|
+
the brake's examined count is pinned as an exact delta (one added full
|
|
419
|
+
citation raises it by one, one added continuation by one more) rather
|
|
420
|
+
than by a floor a dropped-continuation mutant could sink under. The
|
|
421
|
+
`log.md` resolver rejects a cited path carrying a `..` segment and
|
|
422
|
+
asserts repository containment on its on-disk fallback, and a bare
|
|
423
|
+
basename that also exists at the repository root is reported ambiguous
|
|
424
|
+
with both candidates named instead of silently binding to this
|
|
425
|
+
package's own file; the deeper repo-wide basename ambiguity okf-kit
|
|
426
|
+
reports is still bound unconditionally by `anchorScopeResolve()`'s own
|
|
427
|
+
documented design, named as the residual. Convention this round
|
|
428
|
+
installs (D-050): a log entry or CHANGELOG bullet writes no hand-typed
|
|
429
|
+
count of what a guard found; the live figures are the guards' own
|
|
430
|
+
computed test names, read off a passing run.
|
|
431
|
+
|
|
432
|
+
### Fixed
|
|
433
|
+
|
|
434
|
+
- Three unpinned properties from the review round 4 lows above
|
|
435
|
+
(task agent-dx 4ece8e1e) now have a discriminating fixture each: the
|
|
436
|
+
`log.md` resolver's on-disk fallback is checked against an existing,
|
|
437
|
+
outside-the-repository absolute path so its own containment conjunct is
|
|
438
|
+
no longer redundant with the `..`-segment rejection; `extractSiblingGuard
|
|
439
|
+
Citations` gets its own wrapped-citation fixture (full citation and
|
|
440
|
+
continuation each straddling a hard line break), independent of the
|
|
441
|
+
`log.md` guard's; and the previously duplicated `PATH_SHAPED_BEFORE_RE`
|
|
442
|
+
path-shaped regex is now one module-scope const both call sites read,
|
|
443
|
+
so the two copies can no longer drift apart.
|
|
444
|
+
|
|
10
445
|
## [0.31.0] - 2026-09-07
|
|
11
446
|
|
|
12
447
|
### Changed
|
|
@@ -31,28 +31,56 @@ Rules:
|
|
|
31
31
|
- Touch only the files relevant to the assigned task. Respect the
|
|
32
32
|
allowed_changes and forbidden_changes lists in your task contract.
|
|
33
33
|
- Add or update tests where appropriate. Run the tests you touched and report
|
|
34
|
-
the result honestly; if you could not run them, say why.
|
|
34
|
+
the result honestly; if you could not run them, say why. Cite a coverage
|
|
35
|
+
gate's threshold and pass/fail counts, not a run-specific coverage
|
|
36
|
+
percentage; cite a percentage only together with the exact commit and the
|
|
37
|
+
run count, since branch coverage can vary between runs of the same commit.
|
|
35
38
|
- When the task assignment names mutation probes to run, run each one and
|
|
36
|
-
report it in the `mutation_probes` field of your output (mutant,
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
`
|
|
39
|
+
report it in the `mutation_probes` field of your output (mutant, file,
|
|
40
|
+
anchor, before, after, verified_applied_via, result, expectation,
|
|
41
|
+
reason, restored_verified); an output missing that field when probes
|
|
42
|
+
were named is treated as a misfire, not evidence. `file` and `anchor`
|
|
43
|
+
(a line number or a unique surrounding string) locate the mutant;
|
|
44
|
+
`before` and `after` are the exact text swapped there, so a later round
|
|
45
|
+
can reapply the same edit without guessing instead of only a prose
|
|
46
|
+
description. `expectation` records whether `result` matched what the
|
|
47
|
+
probe was expected to do (`met`) or not (`violated`), independent of
|
|
48
|
+
`result` itself, only alongside a measured `killed` or `survived`
|
|
49
|
+
`result`; it is `not_applicable` otherwise (for example when the mutant
|
|
50
|
+
could not be applied and no `result` was measured). `reason` is free
|
|
51
|
+
text, required when `result` is `not_applicable`, empty otherwise,
|
|
52
|
+
carrying one of two canonical strings that distinguish a non-regression
|
|
53
|
+
from a regression: `no definition recorded` (a prior-round probe
|
|
54
|
+
recorded with only an id, no definition to reapply) and `target text no
|
|
55
|
+
longer present` (a replayed probe whose mutant can no longer be
|
|
56
|
+
applied). When the assignment names no mutation probes, return
|
|
57
|
+
`mutation_probes: []` rather than omitting the field.
|
|
58
|
+
Each item also carries `replayed`: `false` for a probe newly
|
|
59
|
+
introduced this round.
|
|
42
60
|
- On any round after the task's first, the assignment also names every
|
|
43
61
|
mutation probe named in an earlier round of this task (on the task's
|
|
44
62
|
first round there are none), drawn from the run's
|
|
45
|
-
`04-implementation-summary.md
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
`
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
63
|
+
`04-implementation-summary.md`, naming each by its mutant definition
|
|
64
|
+
(file, anchor, before, after), not merely by its id; a probe recorded
|
|
65
|
+
with only an id and no definition to reapply cannot be replayed and is
|
|
66
|
+
`not_applicable` (reason: `no definition recorded`), not a regression.
|
|
67
|
+
Replay each one, not only this round's new probes, before returning your
|
|
68
|
+
report, and report each replayed probe in `mutation_probes` with the
|
|
69
|
+
evidence fields plus `replayed: true`. A replayed probe whose
|
|
70
|
+
`expectation` is now `violated`, or which can no longer be applied
|
|
71
|
+
(reason: `target text no longer present`), is the regression signal;
|
|
72
|
+
`result` alone is not: report it as such (`result` `survived` or
|
|
73
|
+
`not_applicable` with the reason) and resolve it before the next
|
|
74
|
+
reviewer spawn.
|
|
52
75
|
- When a verify runner is available, run it for the checks the acceptance
|
|
53
76
|
criteria name and report its summary under `tests.executed`; when a
|
|
54
77
|
mutation-probe runner is available, run the named probes through it and
|
|
55
|
-
copy its fields into `mutation_probes
|
|
78
|
+
copy its fields into `mutation_probes`; when the runner reports a
|
|
79
|
+
probe's mutant record (`file`, `anchor`, `before`, `after`) separately
|
|
80
|
+
from its result fields (`verified_applied_via`, `result`, `expectation`,
|
|
81
|
+
`reason`, `restored_verified`), take the definition fields from that
|
|
82
|
+
mutant record so the copied report still carries all eleven
|
|
83
|
+
`mutation_probes` sub-fields.
|
|
56
84
|
- Run every long test, build, or mutation-probe command in the foreground
|
|
57
85
|
and wait for it to finish before returning. When one foreground call
|
|
58
86
|
cannot hold it to completion, poll the backgrounded run to completion
|
|
@@ -80,6 +108,11 @@ Rules:
|
|
|
80
108
|
when the task assignment asked for a commit is treated as a misfire, not
|
|
81
109
|
evidence. When the task produced no commit, return `commits: []` rather
|
|
82
110
|
than omitting the field.
|
|
111
|
+
- Populate a non-empty `commits` field by pasting `git log --reverse
|
|
112
|
+
--format=%H <base>..HEAD`; never type or hand-complete commit shas.
|
|
113
|
+
- Verification plans, probe plans, and repeat tallies run in the foreground,
|
|
114
|
+
and the implementer reports their returns in the same turn as the last
|
|
115
|
+
check. A background monitor is no substitute for those returns.
|
|
83
116
|
- Only write a verification claim (for example "Verified by ...") in a code
|
|
84
117
|
comment, commit message, or your report for a check you actually ran and
|
|
85
118
|
measured yourself; never claim a run you did not execute.
|
|
@@ -131,8 +164,14 @@ tests:
|
|
|
131
164
|
not_executed_reason: ""
|
|
132
165
|
mutation_probes:
|
|
133
166
|
- mutant: ""
|
|
167
|
+
file: ""
|
|
168
|
+
anchor: ""
|
|
169
|
+
before: ""
|
|
170
|
+
after: ""
|
|
134
171
|
verified_applied_via: ""
|
|
135
|
-
result:
|
|
172
|
+
result: killed | survived | not_applicable
|
|
173
|
+
expectation: met | violated | not_applicable
|
|
174
|
+
reason: ""
|
|
136
175
|
restored_verified: ""
|
|
137
176
|
replayed: false | true
|
|
138
177
|
risks:
|
|
@@ -22,6 +22,25 @@ a version. For a recorded original string-list contract, retain the original
|
|
|
22
22
|
and `criterion_evidence` fields; keep all existing role output fields. This
|
|
23
23
|
selection governs the rules and every YAML block below.
|
|
24
24
|
|
|
25
|
+
Review method: the orchestrator names `review_method: normal | rigorous |
|
|
26
|
+
adversarial` in every briefing; treat an unnamed method as `rigorous`. The
|
|
27
|
+
three methods are obligation sets, not personas: they define what you must
|
|
28
|
+
read, reproduce, and probe, and how a non-reproducing finding is withdrawn,
|
|
29
|
+
not how skeptical to sound.
|
|
30
|
+
|
|
31
|
+
| Method | Obligations |
|
|
32
|
+
|---|---|
|
|
33
|
+
| `normal` | Read the diff and the spec; run the declared tests once; findings come only from what you read. `normal` adds nothing beyond the obligations already stated in the Check list and the Rules below, and suspends none of them: the empirical-reproduction rule and the GitHub Actions shell replay rule apply under every method. `normal` only means no further independent reproduction beyond what those already require. Fits docs, renames, and batch cosmetics. |
|
|
34
|
+
| `rigorous` (default) | Everything `normal` requires, plus: your own extract of the change, a base-attribution control, classifying every change, and reproducing every empirical claim yourself. `reproduction` and `matches_implementer_claim` are mandatory, as already required below. |
|
|
35
|
+
| `adversarial` | Everything `rigorous` requires, plus: one discriminating probe or negative control per acceptance criterion; an active search of the neighbouring scenario space (environment, install modes, platform, ordering, concurrency); an attempt to break the claimed invariant; and an explicit list of break attempts that failed. |
|
|
36
|
+
|
|
37
|
+
Withdrawal rule (`rigorous` and `adversarial`): a finding that does not
|
|
38
|
+
reproduce on a second attempt with a corrected harness is withdrawn in the
|
|
39
|
+
same round, not carried into the next one, and reported under `withdrawn`
|
|
40
|
+
with the reason; this keeps the method from buying false positives. Emit
|
|
41
|
+
`withdrawn: []` when nothing was withdrawn. Report the method you actually
|
|
42
|
+
applied in `method_applied`.
|
|
43
|
+
|
|
25
44
|
Check, at minimum:
|
|
26
45
|
|
|
27
46
|
- Acceptance baseline: for a run explicitly adopted as `acceptance-baseline/v1`,
|
|
@@ -116,6 +135,10 @@ Rules:
|
|
|
116
135
|
shell replay above is a second, explicitly non-probabilistic trigger for
|
|
117
136
|
the same field: report it in `reproduction` too, with `sample_size:
|
|
118
137
|
not_applicable` when the replay itself has no meaningful sample size.
|
|
138
|
+
- When citing a coverage gate, cite the threshold and pass/fail counts, not
|
|
139
|
+
a run-specific coverage percentage; cite a percentage only together with
|
|
140
|
+
the exact commit and the run count, since branch coverage can vary
|
|
141
|
+
between runs of the same commit.
|
|
119
142
|
- When a mutation-probe runner is available in the session, run probes
|
|
120
143
|
through it instead of editing files by hand, and carry its result fields
|
|
121
144
|
into your findings and `reproduction`; when a verify runner is available,
|
|
@@ -145,4 +168,8 @@ reproduction:
|
|
|
145
168
|
sample_size: ""
|
|
146
169
|
result: ""
|
|
147
170
|
matches_implementer_claim: matched | mismatched | not_applicable
|
|
171
|
+
method_applied: normal | rigorous | adversarial
|
|
172
|
+
withdrawn:
|
|
173
|
+
- description: ""
|
|
174
|
+
reason: ""
|
|
148
175
|
```
|
|
@@ -49,6 +49,13 @@ default, not a ritual.
|
|
|
49
49
|
orchestrator may review it itself; reserve the reviewer subagent for
|
|
50
50
|
changes whose risk or size warrants an independent skeptical pass. Either
|
|
51
51
|
way, review is never skipped.
|
|
52
|
+
- Every reviewer briefing also names a `review_method`: `normal | rigorous |
|
|
53
|
+
adversarial`, an obligation set orthogonal to the effort tier below.
|
|
54
|
+
`adversarial` is the minimum for security judgment, install/deploy
|
|
55
|
+
scripts, hand-edited lockfiles, cross-major overrides, or anything the
|
|
56
|
+
operator flags high-risk; `normal` fits only docs, renames, or batch
|
|
57
|
+
cosmetics; `rigorous` is the default otherwise. Never pair `adversarial`
|
|
58
|
+
with the `-medium` reviewer tier; tiers themselves are unchanged.
|
|
52
59
|
- When tier variants are installed (manifest `tiers: true`), the orchestrator
|
|
53
60
|
picks the effort tier per task by complexity and risk, at its own judgment.
|
|
54
61
|
The unsuffixed default subagent is the normal case; `-high`/`-xhigh` fit
|
package/assets/skill/SKILL.md
CHANGED
|
@@ -211,20 +211,34 @@ directory and the subagents.
|
|
|
211
211
|
for real, observe the named test fail, restore, re-verify). Hold the
|
|
212
212
|
implementer's report to the claim-only-what-was-measured rule too: treat any
|
|
213
213
|
verification claim there that is not backed by a check it actually ran as
|
|
214
|
-
unverified.
|
|
214
|
+
unverified. The installed `implementer.md` prompt has the implementer cite
|
|
215
|
+
a coverage gate's threshold and pass/fail counts, not a run-specific
|
|
216
|
+
coverage percentage, citing a percentage only together with the exact
|
|
217
|
+
commit and the run count, since branch coverage can vary between runs of
|
|
218
|
+
the same commit. On any round after the task's first, the briefing also names
|
|
215
219
|
every mutation probe named in an earlier round of this task (on the
|
|
216
220
|
task's first round there are none), drawn from the run's
|
|
217
|
-
`04-implementation-summary.md
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
`
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
221
|
+
`04-implementation-summary.md`, naming each by its mutant definition
|
|
222
|
+
(file, anchor, before, after), not merely by its id; a probe recorded
|
|
223
|
+
with only an id and no definition to reapply cannot be replayed and is
|
|
224
|
+
`not_applicable` (reason: `no definition recorded`), not a regression.
|
|
225
|
+
The implementer replays each one, not only the round's new probes,
|
|
226
|
+
before the next reviewer spawn, and reports each in `mutation_probes`
|
|
227
|
+
with the evidence fields plus `replayed: true`. A replayed probe whose
|
|
228
|
+
`expectation` is now `violated`, or which can no longer be applied
|
|
229
|
+
(reason: `target text no longer present`), is the regression signal;
|
|
230
|
+
`result` alone is not: reported as such (`result` `survived` or
|
|
231
|
+
`not_applicable` with the reason) and resolved before the next reviewer
|
|
232
|
+
spawn. Record meaningful decisions in
|
|
224
233
|
`03-decisions.md` and consolidate evidence in
|
|
225
234
|
`04-implementation-summary.md`, recording each probe the implementer
|
|
226
235
|
reports as a row in `04-implementation-summary.md`'s Mutation Probes
|
|
227
|
-
subsection, with the round it was named in.
|
|
236
|
+
subsection, with the round it was named in. Each row's Before/After
|
|
237
|
+
cells hold a single-line excerpt; when the mutant's actual before/after
|
|
238
|
+
text is multi-line or contains an unescaped `|`, or the mutant is a
|
|
239
|
+
patch/diff rather than a text swap, the full text or diff goes in the
|
|
240
|
+
implementer report or a fenced block placed directly under the table,
|
|
241
|
+
with the row noting where it lives. For any diff that adds or
|
|
228
242
|
changes a GitHub Actions `run:` step, the installed `implementer.md`
|
|
229
243
|
prompt requires replaying it locally under the shell the step actually
|
|
230
244
|
runs, with the expected-success and the expected-failure inputs, before
|
|
@@ -251,52 +265,67 @@ directory and the subagents.
|
|
|
251
265
|
`reviewer-<tier>` subagents, if any) by the task's complexity and risk, at
|
|
252
266
|
your own judgment, defaulting to the unsuffixed subagent when unsure; record
|
|
253
267
|
a non-default tier choice with a one-line reason in `03-decisions.md` when
|
|
254
|
-
the task is non-trivial.
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
the
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
268
|
+
the task is non-trivial. Also name `review_method: normal | rigorous |
|
|
269
|
+
adversarial` in the briefing; every briefing names one. Pick it by risk
|
|
270
|
+
class: `adversarial` at minimum for security judgment, install/deploy
|
|
271
|
+
scripts, hand-edited lockfiles, cross-major overrides, or anything the
|
|
272
|
+
operator flags high-risk; `normal` only for docs, renames, or batch
|
|
273
|
+
cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
|
|
274
|
+
never substitutes for it: do not pair `adversarial` with the `-medium`
|
|
275
|
+
reviewer tier, a budget mismatch that names probes without the effort to run
|
|
276
|
+
them; tiers themselves are unchanged by this axis. When the reviewer's
|
|
277
|
+
environment cannot use version control to see the diff (for example a
|
|
278
|
+
policy-gated repository), supply the diff as a pre-generated file in the
|
|
279
|
+
briefing instead of expecting the reviewer to derive it, and have the
|
|
280
|
+
reviewer report explicitly if it could only reconstruct the delta some other
|
|
281
|
+
way, rather than silently reviewing less than the full change. The reviewer
|
|
282
|
+
checks spec compliance, architecture consistency, edge cases, security, test
|
|
283
|
+
adequacy (including whether new tests would fail if the change were
|
|
284
|
+
reverted), and maintainability. Findings go to `05-review-findings.md`;
|
|
285
|
+
transfer each finding from the reviewer output contract into the table's
|
|
286
|
+
columns as-is, keeping the Severity and Decision headers unchanged, since
|
|
287
|
+
those two are what the orchestrator-workflow completeness reader verifies.
|
|
288
|
+
Replace the shipped placeholder/legend row with the transferred findings;
|
|
289
|
+
for a genuine zero-findings review, delete that row instead of leaving it in
|
|
290
|
+
place, since the completeness reader treats an untouched placeholder row
|
|
291
|
+
with no finding rows as the template never having been filled in. When
|
|
292
|
+
acceptance rests on empirical or probabilistic evidence (flake rates,
|
|
293
|
+
benchmarks, "n runs green", performance/timing numbers), the reviewer must
|
|
294
|
+
independently reproduce it — its own runs or measurements, not a re-read of
|
|
295
|
+
the implementer's log — and record the method, sample size, and result
|
|
296
|
+
against the implementer's claim in the reviewer output contract's
|
|
297
|
+
`reproduction` field. This does not apply to deterministic checks (a single
|
|
298
|
+
test run, `tsc`, lint): only claims that could vary run to run trigger it.
|
|
299
|
+
The GitHub Actions shell replay named in step 6 is a second, explicitly
|
|
278
300
|
non-probabilistic trigger for the same field, with `sample_size:
|
|
279
301
|
not_applicable` allowed when the replay itself has no meaningful sample
|
|
280
|
-
size.
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
|
|
302
|
+
size. When citing a coverage gate, the installed `reviewer.md` prompt has
|
|
303
|
+
the reviewer cite the threshold and pass/fail counts, not a run-specific
|
|
304
|
+
coverage percentage, citing a percentage only together with the exact commit
|
|
305
|
+
and the run count, since branch coverage can vary between runs of the same
|
|
306
|
+
commit. A change that deletes or renames an exported identifier, type,
|
|
307
|
+
config key, or file is also checked for identifier drift (docs or comments
|
|
308
|
+
still describing the old name as current), by the reviewer or by the
|
|
309
|
+
orchestrator itself when it reviews a trivial rename per Scaling delegation,
|
|
310
|
+
using a connected drift check when one exists. When this is not the task's
|
|
311
|
+
first review round, name the round number in the briefing; the reviewer
|
|
312
|
+
marks each finding's `recurrence` as `new` or `repeated` against the earlier
|
|
313
|
+
rounds it was told about, which is what lets the orchestrator detect the
|
|
314
|
+
Review-round escalation budget's trigger (see below) without re-deriving it
|
|
315
|
+
by hand. When the implementer's report replays a prior round's mutation
|
|
316
|
+
probe, the orchestrator's reviewer briefing names the replayed probes the
|
|
317
|
+
implementer reports as killed together with their mutant definition
|
|
318
|
+
(`file`, `anchor`, `before`, `after`) and `verified_applied_via` value,
|
|
319
|
+
not merely their id; a probe recorded with only an id and no definition
|
|
320
|
+
cannot be skipped this way and is `not_applicable`. The reviewer may
|
|
321
|
+
then skip re-running the ones named by definition.
|
|
322
|
+
The reviewer output contract itself is unchanged. Never run mutation probes
|
|
323
|
+
in place against a worktree a reviewer subagent is concurrently reviewing;
|
|
324
|
+
isolate the probe in a separate worktree or wait until the reviewer has
|
|
325
|
+
returned before probing that tree again. For an explicitly adopted v1 run,
|
|
326
|
+
ask the reviewer to compare the frozen delegated criteria with the
|
|
327
|
+
referenced evidence and judge semantic adequacy, including whether a manual
|
|
328
|
+
check is actually concrete and reasoned.
|
|
300
329
|
8. **Decide acceptance.** Accept, request fixes, defer, or escalate to the
|
|
301
330
|
operator. High or critical findings block acceptance until fixed or
|
|
302
331
|
explicitly waived: critical findings require operator sign-off; high
|
|
@@ -304,8 +333,8 @@ directory and the subagents.
|
|
|
304
333
|
or critical finding counts as a waiver and follows the same rules. Record
|
|
305
334
|
all decisions and waivers in `03-decisions.md` and summarize waivers in
|
|
306
335
|
the Accepted Waivers section of `06-handoff.md`. A reviewer recommendation is not orchestrator acceptance and cannot authorize a critical waiver; only the operator may authorize a critical waiver. For newly created decision records, identify a stable ID, trigger/evidence, decision, accountable authority/source with concrete approval evidence, consequences, and a superseded decision ID when revising a prior decision. Link baseline revisions and waivers to those decision IDs. Established runs retain their recorded decision format; absent fields never create a retroactive blocker. Routine decisions within the delegated contract remain the orchestrator's responsibility; an out-of-scope change requires an operator decision. Markdown records evidence of real authority and never grant it by themselves. Do not accept while a
|
|
307
|
-
required baseline criterion in an explicitly adopted v1 run has an open
|
|
308
|
-
and cannot be converted away. After independent review,
|
|
336
|
+
required baseline criterion in an explicitly adopted v1 run has an open
|
|
337
|
+
residual; a residual retains its ID and cannot be converted away. After independent review,
|
|
309
338
|
the orchestrator may close a docs-only delta without another reviewer round only
|
|
310
339
|
when the entire unreviewed delta contains only explanatory
|
|
311
340
|
documentation, comments, or citations; contains no source- or test-file
|
|
@@ -445,8 +474,14 @@ tests:
|
|
|
445
474
|
not_executed_reason: ""
|
|
446
475
|
mutation_probes:
|
|
447
476
|
- mutant: ""
|
|
477
|
+
file: ""
|
|
478
|
+
anchor: ""
|
|
479
|
+
before: ""
|
|
480
|
+
after: ""
|
|
448
481
|
verified_applied_via: ""
|
|
449
|
-
result:
|
|
482
|
+
result: killed | survived | not_applicable
|
|
483
|
+
expectation: met | violated | not_applicable
|
|
484
|
+
reason: ""
|
|
450
485
|
restored_verified: ""
|
|
451
486
|
replayed: false | true
|
|
452
487
|
risks:
|
|
@@ -472,19 +507,37 @@ identifies the reviewed artifact and revision, reviewer, method, pass/fail
|
|
|
472
507
|
standard, reasoned result, and baseline/criterion identities; it stays manual.
|
|
473
508
|
|
|
474
509
|
When the task assignment names mutation probes to run, the implementer
|
|
475
|
-
reports each one in the `mutation_probes` field (mutant,
|
|
476
|
-
verified_applied_via, result,
|
|
477
|
-
|
|
478
|
-
|
|
479
|
-
|
|
480
|
-
|
|
481
|
-
|
|
482
|
-
|
|
483
|
-
|
|
484
|
-
|
|
485
|
-
|
|
486
|
-
|
|
487
|
-
|
|
510
|
+
reports each one in the `mutation_probes` field (mutant, file, anchor,
|
|
511
|
+
before, after, verified_applied_via, result, expectation, reason,
|
|
512
|
+
restored_verified); `file` and `anchor` (a line number or a unique
|
|
513
|
+
surrounding string) locate the mutant, `before` and `after` are the
|
|
514
|
+
exact text swapped there, and `expectation` records whether `result`
|
|
515
|
+
matched what the probe was expected to do (`met`) or not (`violated`),
|
|
516
|
+
independent of `result` itself, only alongside a measured `killed` or
|
|
517
|
+
`survived` `result`; it is `not_applicable` otherwise (for example when
|
|
518
|
+
the mutant could not be applied and no `result` was measured). `reason`
|
|
519
|
+
is free text, required when `result` is `not_applicable`, empty
|
|
520
|
+
otherwise, carrying one of two canonical strings that distinguish a
|
|
521
|
+
non-regression from a regression: `no definition recorded` (a
|
|
522
|
+
prior-round probe recorded with only an id, no definition to reapply)
|
|
523
|
+
and `target text no longer present` (a replayed probe whose mutant can
|
|
524
|
+
no longer be applied). When the assignment names none, it returns
|
|
525
|
+
`mutation_probes: []` rather than
|
|
526
|
+
omitting the field, so 'none asked for' is distinguishable from 'asked
|
|
527
|
+
for and not reported'. Each item also carries `replayed`: `false` for a
|
|
528
|
+
probe newly introduced this round, `true` for a prior round's probe
|
|
529
|
+
replayed this round under the replay rule in step 6. On any round after
|
|
530
|
+
the task's first, the implementer replays every probe named in an
|
|
531
|
+
earlier round of this task (on the task's first round there are none),
|
|
532
|
+
naming each by its mutant definition, not merely by its id, not only
|
|
533
|
+
this round's new probes, before the next reviewer spawn, reporting each
|
|
534
|
+
one in `mutation_probes` alongside the round's new probes. A replayed
|
|
535
|
+
probe whose `expectation` is now `violated`, or which can no longer be
|
|
536
|
+
applied (reason: `target text no longer present`), is the regression
|
|
537
|
+
signal, reported as such and resolved before the next reviewer spawn;
|
|
538
|
+
`result` alone is not a regression signal, and a probe recorded with
|
|
539
|
+
only an id and no definition to reapply is `not_applicable` (reason:
|
|
540
|
+
`no definition recorded`).
|
|
488
541
|
|
|
489
542
|
The `commits` field lists the full sha of every commit the implementer
|
|
490
543
|
produced on the task branch, in the order produced; when the task
|
|
@@ -492,6 +545,12 @@ produced no commit, the implementer returns `commits: []` rather than
|
|
|
492
545
|
omitting the field, so 'did not commit' is distinguishable from
|
|
493
546
|
'forgot to report'.
|
|
494
547
|
|
|
548
|
+
For a non-empty `commits` field, the implementer pastes `git log
|
|
549
|
+
--reverse --format=%H <base>..HEAD`; it never types or hand-completes commit
|
|
550
|
+
shas. Verification plans, probe plans, and repeat tallies run in the foreground, and the
|
|
551
|
+
implementer reports their returns in the same turn as the last check. A
|
|
552
|
+
background monitor is no substitute for those returns.
|
|
553
|
+
|
|
495
554
|
## Reviewer output contract
|
|
496
555
|
|
|
497
556
|
The output shape remains the same for either selected contract. Compare the
|
|
@@ -520,6 +579,10 @@ reproduction:
|
|
|
520
579
|
sample_size: ""
|
|
521
580
|
result: ""
|
|
522
581
|
matches_implementer_claim: matched | mismatched | not_applicable
|
|
582
|
+
method_applied: normal | rigorous | adversarial
|
|
583
|
+
withdrawn:
|
|
584
|
+
- description: ""
|
|
585
|
+
reason: ""
|
|
523
586
|
```
|
|
524
587
|
|
|
525
588
|
`acceptance_recommendation` is mandatory: every reviewer return must set it.
|
|
@@ -532,6 +595,15 @@ one that already appeared in an earlier round. On a task's first review
|
|
|
532
595
|
round every finding is `new` by definition. This is what feeds the
|
|
533
596
|
Review-round escalation budget's trigger.
|
|
534
597
|
|
|
598
|
+
`method_applied` echoes the `review_method` named in the briefing (see step
|
|
599
|
+
7); `withdrawn` lists each finding the reviewer proposed and then retracted
|
|
600
|
+
under the withdrawal rule (`rigorous` and `adversarial` only), with its
|
|
601
|
+
reason; emit `withdrawn: []` when nothing was withdrawn. Until a
|
|
602
|
+
grounding-mcp reader parses the marker (tracked as a cross-repo
|
|
603
|
+
follow-up), the orchestrator checks by hand that the return's
|
|
604
|
+
`method_applied` matches the briefing's `review_method`; a mismatch or
|
|
605
|
+
omission is resupplied, not accepted.
|
|
606
|
+
|
|
535
607
|
## Task slicer output contract
|
|
536
608
|
|
|
537
609
|
Use this v1 block subject to Contract selection above for every task.
|
|
@@ -68,9 +68,15 @@ an optional row cannot stand in for a required criterion.
|
|
|
68
68
|
|
|
69
69
|
### Mutation Probes
|
|
70
70
|
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
71
|
+
Before/After cells hold a single-line excerpt. When the mutant's actual
|
|
72
|
+
before/after text is multi-line or contains an unescaped `|`, or the mutant
|
|
73
|
+
is a patch/diff rather than a text swap, put the full text or diff in the
|
|
74
|
+
implementer report or a fenced block directly under the table, and note
|
|
75
|
+
where it lives in the row's own cell.
|
|
76
|
+
|
|
77
|
+
| Round | Mutant | File | Anchor | Before | After | Verified Applied Via | Result | Expectation | Reason | Restored Verified | Replayed |
|
|
78
|
+
|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
79
|
+
| <!-- round --> | <!-- mutant --> | <!-- file --> | <!-- anchor --> | <!-- before --> | <!-- after --> | <!-- verified_applied_via --> | <!-- result --> | <!-- expectation --> | <!-- reason --> | <!-- restored_verified --> | <!-- replayed --> |
|
|
74
80
|
|
|
75
81
|
## Risks / Notes
|
|
76
82
|
|
|
@@ -4,6 +4,11 @@
|
|
|
4
4
|
|
|
5
5
|
<!-- Short summary. -->
|
|
6
6
|
|
|
7
|
+
<!-- review-method[<round>] = normal|rigorous|adversarial -->
|
|
8
|
+
Method: normal | rigorous | adversarial (the `review_method` named in this
|
|
9
|
+
round's briefing and the `method_applied` the reviewer returned; not parsed
|
|
10
|
+
by the grounding-mcp completeness reader yet).
|
|
11
|
+
|
|
7
12
|
## Findings
|
|
8
13
|
|
|
9
14
|
<!-- The Severity and Decision column headers below are load-bearing: the orchestrator-workflow completeness reader locates this table by its header row and verifies unresolved findings from those two columns. Do not rename or drop them. -->
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "orchestrator-workflow",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.33.0",
|
|
4
4
|
"description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
|
|
5
5
|
"main": "dist/index.js",
|
|
6
6
|
"type": "module",
|