orchestrator-workflow 0.32.0 → 0.34.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +58 -0
- package/assets/agents/implementer.md +55 -16
- package/assets/agents/reviewer.md +6 -2
- package/assets/agents-md-section.md +3 -3
- package/assets/skill/SKILL.md +127 -79
- package/assets/templates/03-decisions.md +1 -1
- package/assets/templates/04-implementation-summary.md +9 -3
- package/assets/templates/05-review-findings.md +2 -2
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,64 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.34.0] - 2026-09-13
|
|
11
|
+
|
|
12
|
+
- Reviewer findings now carry `introduced_by_delta: yes | no | unknown`.
|
|
13
|
+
A `no` attribution requires a named base build and replay in `reproduction`;
|
|
14
|
+
it remains in the ordinary finding gate and Findings table, while only `yes`
|
|
15
|
+
and `unknown` participate in bounded halt and escalation rules. The legacy
|
|
16
|
+
five-cell table and placeholder row remain byte-compatible with the
|
|
17
|
+
completeness reader; concrete rows record attribution parenthetically in
|
|
18
|
+
their Description field.
|
|
19
|
+
|
|
20
|
+
## [0.33.0] - 2026-09-12
|
|
21
|
+
|
|
22
|
+
### Changed
|
|
23
|
+
|
|
24
|
+
- Implementer reports now paste non-empty commit lists from `git log --reverse
|
|
25
|
+
--format=%H <base>..HEAD`, and foreground verification/probe plans and
|
|
26
|
+
repeat tallies return in the same turn as the last check; a background
|
|
27
|
+
monitor does not substitute. Anchored by pandora batch48 evidence.
|
|
28
|
+
|
|
29
|
+
- The implementer output contract now enumerates mutation-probe `result` as
|
|
30
|
+
`killed | survived | not_applicable` in both the installed prompt and the
|
|
31
|
+
SKILL.md reference. The docs-consistency guard separately pins each copy's
|
|
32
|
+
complete field block, including that enum, and now recognizes versioned
|
|
33
|
+
parenthesized headings in `see <name> below` forward pointers.
|
|
34
|
+
|
|
35
|
+
- The implementer and reviewer prompts (`assets/agents/implementer.md`,
|
|
36
|
+
`assets/agents/reviewer.md`, mirrored in SKILL.md) now say: cite a
|
|
37
|
+
coverage gate's threshold and pass/fail counts, not a run-specific
|
|
38
|
+
coverage percentage; cite a percentage only together with the exact
|
|
39
|
+
commit and the run count, since branch coverage varies between runs of
|
|
40
|
+
the same commit. Anchored by pandora run
|
|
41
|
+
`.ai/runs/2026-09-11-memory-sync-wipe`.
|
|
42
|
+
- The implementer `mutation_probes` output field (`assets/agents/
|
|
43
|
+
implementer.md`, mirrored in SKILL.md, and the `04-implementation-
|
|
44
|
+
summary.md` template's Mutation Probes table) now carries the mutant's
|
|
45
|
+
definition, not only its label: `file`, `anchor` (a line number or a
|
|
46
|
+
unique string), `before`, and `after` alongside the existing
|
|
47
|
+
`verified_applied_via`, `result`, `restored_verified`, and `replayed`
|
|
48
|
+
fields, so a later round can mechanically reapply the same edit instead
|
|
49
|
+
of only reading a prose description. Added `expectation: met | violated
|
|
50
|
+
| not_applicable` beside `result`, reporting whether the probe's
|
|
51
|
+
`result` matched its `--expect`, scoped to a measured `killed` or
|
|
52
|
+
`survived` `result` (a routine negative-control probe now reports
|
|
53
|
+
`result: survived, expectation: met`, which is not a regression) and
|
|
54
|
+
`not_applicable` otherwise (for example when the mutant could not be
|
|
55
|
+
applied and no `result` was measured). Added an eleventh sub-field,
|
|
56
|
+
`reason`: free text, required when `result` is `not_applicable`, empty
|
|
57
|
+
otherwise, carrying one of two canonical strings that distinguish a
|
|
58
|
+
non-regression from a regression: `no definition recorded` (a
|
|
59
|
+
prior-round probe recorded with only an id, no definition to reapply)
|
|
60
|
+
and `target text no longer present` (a replayed probe whose mutant can
|
|
61
|
+
no longer be applied). The fix-round replay rule now names a
|
|
62
|
+
probe to replay by its definition, not merely by its id, and treats a
|
|
63
|
+
replayed probe whose `expectation` is now `violated` (or which can no
|
|
64
|
+
longer be applied) as the regression signal, not `result` alone.
|
|
65
|
+
Anchored by pandora run `.ai/runs/2026-09-11-memory-sync-wipe` and the
|
|
66
|
+
agent-primitives `probe` result/expectation split (0.3.0).
|
|
67
|
+
|
|
10
68
|
## [0.32.0] - 2026-09-11
|
|
11
69
|
|
|
12
70
|
### Added
|
|
@@ -31,28 +31,56 @@ Rules:
|
|
|
31
31
|
- Touch only the files relevant to the assigned task. Respect the
|
|
32
32
|
allowed_changes and forbidden_changes lists in your task contract.
|
|
33
33
|
- Add or update tests where appropriate. Run the tests you touched and report
|
|
34
|
-
the result honestly; if you could not run them, say why.
|
|
34
|
+
the result honestly; if you could not run them, say why. Cite a coverage
|
|
35
|
+
gate's threshold and pass/fail counts, not a run-specific coverage
|
|
36
|
+
percentage; cite a percentage only together with the exact commit and the
|
|
37
|
+
run count, since branch coverage can vary between runs of the same commit.
|
|
35
38
|
- When the task assignment names mutation probes to run, run each one and
|
|
36
|
-
report it in the `mutation_probes` field of your output (mutant,
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
`
|
|
39
|
+
report it in the `mutation_probes` field of your output (mutant, file,
|
|
40
|
+
anchor, before, after, verified_applied_via, result, expectation,
|
|
41
|
+
reason, restored_verified); an output missing that field when probes
|
|
42
|
+
were named is treated as a misfire, not evidence. `file` and `anchor`
|
|
43
|
+
(a line number or a unique surrounding string) locate the mutant;
|
|
44
|
+
`before` and `after` are the exact text swapped there, so a later round
|
|
45
|
+
can reapply the same edit without guessing instead of only a prose
|
|
46
|
+
description. `expectation` records whether `result` matched what the
|
|
47
|
+
probe was expected to do (`met`) or not (`violated`), independent of
|
|
48
|
+
`result` itself, only alongside a measured `killed` or `survived`
|
|
49
|
+
`result`; it is `not_applicable` otherwise (for example when the mutant
|
|
50
|
+
could not be applied and no `result` was measured). `reason` is free
|
|
51
|
+
text, required when `result` is `not_applicable`, empty otherwise,
|
|
52
|
+
carrying one of two canonical strings that distinguish a non-regression
|
|
53
|
+
from a regression: `no definition recorded` (a prior-round probe
|
|
54
|
+
recorded with only an id, no definition to reapply) and `target text no
|
|
55
|
+
longer present` (a replayed probe whose mutant can no longer be
|
|
56
|
+
applied). When the assignment names no mutation probes, return
|
|
57
|
+
`mutation_probes: []` rather than omitting the field.
|
|
58
|
+
Each item also carries `replayed`: `false` for a probe newly
|
|
59
|
+
introduced this round.
|
|
42
60
|
- On any round after the task's first, the assignment also names every
|
|
43
61
|
mutation probe named in an earlier round of this task (on the task's
|
|
44
62
|
first round there are none), drawn from the run's
|
|
45
|
-
`04-implementation-summary.md
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
`
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
63
|
+
`04-implementation-summary.md`, naming each by its mutant definition
|
|
64
|
+
(file, anchor, before, after), not merely by its id; a probe recorded
|
|
65
|
+
with only an id and no definition to reapply cannot be replayed and is
|
|
66
|
+
`not_applicable` (reason: `no definition recorded`), not a regression.
|
|
67
|
+
Replay each one, not only this round's new probes, before returning your
|
|
68
|
+
report, and report each replayed probe in `mutation_probes` with the
|
|
69
|
+
evidence fields plus `replayed: true`. A replayed probe whose
|
|
70
|
+
`expectation` is now `violated`, or which can no longer be applied
|
|
71
|
+
(reason: `target text no longer present`), is the regression signal;
|
|
72
|
+
`result` alone is not: report it as such (`result` `survived` or
|
|
73
|
+
`not_applicable` with the reason) and resolve it before the next
|
|
74
|
+
reviewer spawn.
|
|
52
75
|
- When a verify runner is available, run it for the checks the acceptance
|
|
53
76
|
criteria name and report its summary under `tests.executed`; when a
|
|
54
77
|
mutation-probe runner is available, run the named probes through it and
|
|
55
|
-
copy its fields into `mutation_probes
|
|
78
|
+
copy its fields into `mutation_probes`; when the runner reports a
|
|
79
|
+
probe's mutant record (`file`, `anchor`, `before`, `after`) separately
|
|
80
|
+
from its result fields (`verified_applied_via`, `result`, `expectation`,
|
|
81
|
+
`reason`, `restored_verified`), take the definition fields from that
|
|
82
|
+
mutant record so the copied report still carries all eleven
|
|
83
|
+
`mutation_probes` sub-fields.
|
|
56
84
|
- Run every long test, build, or mutation-probe command in the foreground
|
|
57
85
|
and wait for it to finish before returning. When one foreground call
|
|
58
86
|
cannot hold it to completion, poll the backgrounded run to completion
|
|
@@ -80,6 +108,11 @@ Rules:
|
|
|
80
108
|
when the task assignment asked for a commit is treated as a misfire, not
|
|
81
109
|
evidence. When the task produced no commit, return `commits: []` rather
|
|
82
110
|
than omitting the field.
|
|
111
|
+
- Populate a non-empty `commits` field by pasting `git log --reverse
|
|
112
|
+
--format=%H <base>..HEAD`; never type or hand-complete commit shas.
|
|
113
|
+
- Verification plans, probe plans, and repeat tallies run in the foreground,
|
|
114
|
+
and the implementer reports their returns in the same turn as the last
|
|
115
|
+
check. A background monitor is no substitute for those returns.
|
|
83
116
|
- Only write a verification claim (for example "Verified by ...") in a code
|
|
84
117
|
comment, commit message, or your report for a check you actually ran and
|
|
85
118
|
measured yourself; never claim a run you did not execute.
|
|
@@ -131,8 +164,14 @@ tests:
|
|
|
131
164
|
not_executed_reason: ""
|
|
132
165
|
mutation_probes:
|
|
133
166
|
- mutant: ""
|
|
167
|
+
file: ""
|
|
168
|
+
anchor: ""
|
|
169
|
+
before: ""
|
|
170
|
+
after: ""
|
|
134
171
|
verified_applied_via: ""
|
|
135
|
-
result:
|
|
172
|
+
result: killed | survived | not_applicable
|
|
173
|
+
expectation: met | violated | not_applicable
|
|
174
|
+
reason: ""
|
|
136
175
|
restored_verified: ""
|
|
137
176
|
replayed: false | true
|
|
138
177
|
risks:
|
|
@@ -73,7 +73,7 @@ Check, at minimum:
|
|
|
73
73
|
review round, classify each finding as `new` or `repeated` against the
|
|
74
74
|
earlier rounds you were told about; on a first round every finding is
|
|
75
75
|
`new` by definition. The orchestrator uses this to detect the
|
|
76
|
-
review-round escalation budget's trigger.
|
|
76
|
+
review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
|
|
77
77
|
- GitHub Actions shell replay: for any diff that adds or changes a GitHub
|
|
78
78
|
Actions `run:` step, replay it yourself under the shell the step actually
|
|
79
79
|
runs: `bash --noprofile --norc -eo pipefail` when `shell: bash` is set on
|
|
@@ -135,13 +135,16 @@ Rules:
|
|
|
135
135
|
shell replay above is a second, explicitly non-probabilistic trigger for
|
|
136
136
|
the same field: report it in `reproduction` too, with `sample_size:
|
|
137
137
|
not_applicable` when the replay itself has no meaningful sample size.
|
|
138
|
+
- When citing a coverage gate, cite the threshold and pass/fail counts, not
|
|
139
|
+
a run-specific coverage percentage; cite a percentage only together with
|
|
140
|
+
the exact commit and the run count, since branch coverage can vary
|
|
141
|
+
between runs of the same commit.
|
|
138
142
|
- When a mutation-probe runner is available in the session, run probes
|
|
139
143
|
through it instead of editing files by hand, and carry its result fields
|
|
140
144
|
into your findings and `reproduction`; when a verify runner is available,
|
|
141
145
|
read its summary before opening full logs.
|
|
142
146
|
|
|
143
147
|
Return exactly this structure as your final output, nothing else:
|
|
144
|
-
|
|
145
148
|
```yaml
|
|
146
149
|
status: reviewed
|
|
147
150
|
role: reviewer
|
|
@@ -154,6 +157,7 @@ findings:
|
|
|
154
157
|
description: ""
|
|
155
158
|
suggested_fix: ""
|
|
156
159
|
recurrence: new | repeated
|
|
160
|
+
introduced_by_delta: yes | no | unknown
|
|
157
161
|
acceptance_recommendation: accept | accept_with_notes | fix_required | reject
|
|
158
162
|
missing_tests:
|
|
159
163
|
- ""
|
|
@@ -128,9 +128,9 @@ trivial change.
|
|
|
128
128
|
the exhausted tier path falls straight to the merge-hold), or an
|
|
129
129
|
operator merge-hold, and adds a row (task, choice, reason) to
|
|
130
130
|
`03-decisions.md`'s Review-round escalation table, then sets the
|
|
131
|
-
`review-round-escalation` marker to the most recent choice. A
|
|
132
|
-
|
|
133
|
-
|
|
131
|
+
`review-round-escalation` marker to the most recent choice. A negative round
|
|
132
|
+
has an `acceptance_recommendation` of `fix_required` or `reject`; a misfired
|
|
133
|
+
review is not a round. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the three is
|
|
134
134
|
picked is judgment; that one is picked and recorded is not. Escalating
|
|
135
135
|
never substitutes for a review round and comes in addition to the halt
|
|
136
136
|
rule's split-or-redesign response, not instead of it.
|
package/assets/skill/SKILL.md
CHANGED
|
@@ -211,20 +211,34 @@ directory and the subagents.
|
|
|
211
211
|
for real, observe the named test fail, restore, re-verify). Hold the
|
|
212
212
|
implementer's report to the claim-only-what-was-measured rule too: treat any
|
|
213
213
|
verification claim there that is not backed by a check it actually ran as
|
|
214
|
-
unverified.
|
|
214
|
+
unverified. The installed `implementer.md` prompt has the implementer cite
|
|
215
|
+
a coverage gate's threshold and pass/fail counts, not a run-specific
|
|
216
|
+
coverage percentage, citing a percentage only together with the exact
|
|
217
|
+
commit and the run count, since branch coverage can vary between runs of
|
|
218
|
+
the same commit. On any round after the task's first, the briefing also names
|
|
215
219
|
every mutation probe named in an earlier round of this task (on the
|
|
216
220
|
task's first round there are none), drawn from the run's
|
|
217
|
-
`04-implementation-summary.md
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
`
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
221
|
+
`04-implementation-summary.md`, naming each by its mutant definition
|
|
222
|
+
(file, anchor, before, after), not merely by its id; a probe recorded
|
|
223
|
+
with only an id and no definition to reapply cannot be replayed and is
|
|
224
|
+
`not_applicable` (reason: `no definition recorded`), not a regression.
|
|
225
|
+
The implementer replays each one, not only the round's new probes,
|
|
226
|
+
before the next reviewer spawn, and reports each in `mutation_probes`
|
|
227
|
+
with the evidence fields plus `replayed: true`. A replayed probe whose
|
|
228
|
+
`expectation` is now `violated`, or which can no longer be applied
|
|
229
|
+
(reason: `target text no longer present`), is the regression signal;
|
|
230
|
+
`result` alone is not: reported as such (`result` `survived` or
|
|
231
|
+
`not_applicable` with the reason) and resolved before the next reviewer
|
|
232
|
+
spawn. Record meaningful decisions in
|
|
224
233
|
`03-decisions.md` and consolidate evidence in
|
|
225
234
|
`04-implementation-summary.md`, recording each probe the implementer
|
|
226
235
|
reports as a row in `04-implementation-summary.md`'s Mutation Probes
|
|
227
|
-
subsection, with the round it was named in.
|
|
236
|
+
subsection, with the round it was named in. Each row's Before/After
|
|
237
|
+
cells hold a single-line excerpt; when the mutant's actual before/after
|
|
238
|
+
text is multi-line or contains an unescaped `|`, or the mutant is a
|
|
239
|
+
patch/diff rather than a text swap, the full text or diff goes in the
|
|
240
|
+
implementer report or a fenced block placed directly under the table,
|
|
241
|
+
with the row noting where it lives. For any diff that adds or
|
|
228
242
|
changes a GitHub Actions `run:` step, the installed `implementer.md`
|
|
229
243
|
prompt requires replaying it locally under the shell the step actually
|
|
230
244
|
runs, with the expected-success and the expected-failure inputs, before
|
|
@@ -258,54 +272,60 @@ directory and the subagents.
|
|
|
258
272
|
operator flags high-risk; `normal` only for docs, renames, or batch
|
|
259
273
|
cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
|
|
260
274
|
never substitutes for it: do not pair `adversarial` with the `-medium`
|
|
261
|
-
reviewer tier, a budget mismatch that names probes without the effort to
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
reviewer
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
implementer's
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
Actions shell replay named in step 6 is a second, explicitly
|
|
275
|
+
reviewer tier, a budget mismatch that names probes without the effort to run
|
|
276
|
+
them; tiers themselves are unchanged by this axis. When the reviewer's
|
|
277
|
+
environment cannot use version control to see the diff (for example a
|
|
278
|
+
policy-gated repository), supply the diff as a pre-generated file in the
|
|
279
|
+
briefing instead of expecting the reviewer to derive it, and have the
|
|
280
|
+
reviewer report explicitly if it could only reconstruct the delta some other
|
|
281
|
+
way, rather than silently reviewing less than the full change. The reviewer
|
|
282
|
+
checks spec compliance, architecture consistency, edge cases, security, test
|
|
283
|
+
adequacy (including whether new tests would fail if the change were
|
|
284
|
+
reverted), and maintainability. Findings go to `05-review-findings.md`;
|
|
285
|
+
transfer each finding from the reviewer output contract into the table's
|
|
286
|
+
columns as-is, keeping the Severity and Decision headers unchanged, since
|
|
287
|
+
those two are what the orchestrator-workflow completeness reader verifies.
|
|
288
|
+
Replace the shipped placeholder/legend row with the transferred findings;
|
|
289
|
+
for a genuine zero-findings review, delete that row instead of leaving it in
|
|
290
|
+
place, since the completeness reader treats an untouched placeholder row
|
|
291
|
+
with no finding rows as the template never having been filled in. When
|
|
292
|
+
acceptance rests on empirical or probabilistic evidence (flake rates,
|
|
293
|
+
benchmarks, "n runs green", performance/timing numbers), the reviewer must
|
|
294
|
+
independently reproduce it — its own runs or measurements, not a re-read of
|
|
295
|
+
the implementer's log — and record the method, sample size, and result
|
|
296
|
+
against the implementer's claim in the reviewer output contract's
|
|
297
|
+
`reproduction` field. This does not apply to deterministic checks (a single
|
|
298
|
+
test run, `tsc`, lint): only claims that could vary run to run trigger it.
|
|
299
|
+
The GitHub Actions shell replay named in step 6 is a second, explicitly
|
|
287
300
|
non-probabilistic trigger for the same field, with `sample_size:
|
|
288
301
|
not_applicable` allowed when the replay itself has no meaningful sample
|
|
289
|
-
size.
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
|
|
307
|
-
|
|
308
|
-
|
|
302
|
+
size. When citing a coverage gate, the installed `reviewer.md` prompt has
|
|
303
|
+
the reviewer cite the threshold and pass/fail counts, not a run-specific
|
|
304
|
+
coverage percentage, citing a percentage only together with the exact commit
|
|
305
|
+
and the run count, since branch coverage can vary between runs of the same
|
|
306
|
+
commit. A change that deletes or renames an exported identifier, type,
|
|
307
|
+
config key, or file is also checked for identifier drift (docs or comments
|
|
308
|
+
still describing the old name as current), by the reviewer or by the
|
|
309
|
+
orchestrator itself when it reviews a trivial rename per Scaling delegation,
|
|
310
|
+
using a connected drift check when one exists. When this is not the task's
|
|
311
|
+
first review round, name the round number in the briefing; the reviewer
|
|
312
|
+
marks each finding's `recurrence` as `new` or `repeated` against the earlier
|
|
313
|
+
rounds it was told about, which is what lets the orchestrator detect the
|
|
314
|
+
Review-round escalation budget's trigger (see below) without re-deriving it
|
|
315
|
+
by hand. The reviewer classifies every finding with the `introduced_by_delta` field (`yes`, `no`, or `unknown`); it sets `no` only after naming a base build and replaying the same reproduction in `reproduction`, and transfers it through the ordinary gate (not bounded-round guidance, which considers only `yes`/`unknown`). When findings are transferred, record the classification parenthetically in the `Description` field as `(introduced_by_delta: yes|no|unknown)`. When the implementer's report replays a prior round's mutation
|
|
316
|
+
probe, the orchestrator's reviewer briefing names the replayed probes the
|
|
317
|
+
implementer reports as killed together with their mutant definition
|
|
318
|
+
(`file`, `anchor`, `before`, `after`) and `verified_applied_via` value,
|
|
319
|
+
not merely their id; a probe recorded with only an id and no definition
|
|
320
|
+
cannot be skipped this way and is `not_applicable`. The reviewer may
|
|
321
|
+
then skip re-running the ones named by definition.
|
|
322
|
+
The reviewer output contract itself is unchanged. Never run mutation probes
|
|
323
|
+
in place against a worktree a reviewer subagent is concurrently reviewing;
|
|
324
|
+
isolate the probe in a separate worktree or wait until the reviewer has
|
|
325
|
+
returned before probing that tree again. For an explicitly adopted v1 run,
|
|
326
|
+
ask the reviewer to compare the frozen delegated criteria with the
|
|
327
|
+
referenced evidence and judge semantic adequacy, including whether a manual
|
|
328
|
+
check is actually concrete and reasoned.
|
|
309
329
|
8. **Decide acceptance.** Accept, request fixes, defer, or escalate to the
|
|
310
330
|
operator. High or critical findings block acceptance until fixed or
|
|
311
331
|
explicitly waived: critical findings require operator sign-off; high
|
|
@@ -313,8 +333,8 @@ directory and the subagents.
|
|
|
313
333
|
or critical finding counts as a waiver and follows the same rules. Record
|
|
314
334
|
all decisions and waivers in `03-decisions.md` and summarize waivers in
|
|
315
335
|
the Accepted Waivers section of `06-handoff.md`. A reviewer recommendation is not orchestrator acceptance and cannot authorize a critical waiver; only the operator may authorize a critical waiver. For newly created decision records, identify a stable ID, trigger/evidence, decision, accountable authority/source with concrete approval evidence, consequences, and a superseded decision ID when revising a prior decision. Link baseline revisions and waivers to those decision IDs. Established runs retain their recorded decision format; absent fields never create a retroactive blocker. Routine decisions within the delegated contract remain the orchestrator's responsibility; an out-of-scope change requires an operator decision. Markdown records evidence of real authority and never grant it by themselves. Do not accept while a
|
|
316
|
-
required baseline criterion in an explicitly adopted v1 run has an open
|
|
317
|
-
and cannot be converted away. After independent review,
|
|
336
|
+
required baseline criterion in an explicitly adopted v1 run has an open
|
|
337
|
+
residual; a residual retains its ID and cannot be converted away. After independent review,
|
|
318
338
|
the orchestrator may close a docs-only delta without another reviewer round only
|
|
319
339
|
when the entire unreviewed delta contains only explanatory
|
|
320
340
|
documentation, comments, or citations; contains no source- or test-file
|
|
@@ -454,8 +474,14 @@ tests:
|
|
|
454
474
|
not_executed_reason: ""
|
|
455
475
|
mutation_probes:
|
|
456
476
|
- mutant: ""
|
|
477
|
+
file: ""
|
|
478
|
+
anchor: ""
|
|
479
|
+
before: ""
|
|
480
|
+
after: ""
|
|
457
481
|
verified_applied_via: ""
|
|
458
|
-
result:
|
|
482
|
+
result: killed | survived | not_applicable
|
|
483
|
+
expectation: met | violated | not_applicable
|
|
484
|
+
reason: ""
|
|
459
485
|
restored_verified: ""
|
|
460
486
|
replayed: false | true
|
|
461
487
|
risks:
|
|
@@ -481,19 +507,37 @@ identifies the reviewed artifact and revision, reviewer, method, pass/fail
|
|
|
481
507
|
standard, reasoned result, and baseline/criterion identities; it stays manual.
|
|
482
508
|
|
|
483
509
|
When the task assignment names mutation probes to run, the implementer
|
|
484
|
-
reports each one in the `mutation_probes` field (mutant,
|
|
485
|
-
verified_applied_via, result,
|
|
486
|
-
|
|
487
|
-
|
|
488
|
-
|
|
489
|
-
|
|
490
|
-
|
|
491
|
-
|
|
492
|
-
|
|
493
|
-
|
|
494
|
-
|
|
495
|
-
|
|
496
|
-
|
|
510
|
+
reports each one in the `mutation_probes` field (mutant, file, anchor,
|
|
511
|
+
before, after, verified_applied_via, result, expectation, reason,
|
|
512
|
+
restored_verified); `file` and `anchor` (a line number or a unique
|
|
513
|
+
surrounding string) locate the mutant, `before` and `after` are the
|
|
514
|
+
exact text swapped there, and `expectation` records whether `result`
|
|
515
|
+
matched what the probe was expected to do (`met`) or not (`violated`),
|
|
516
|
+
independent of `result` itself, only alongside a measured `killed` or
|
|
517
|
+
`survived` `result`; it is `not_applicable` otherwise (for example when
|
|
518
|
+
the mutant could not be applied and no `result` was measured). `reason`
|
|
519
|
+
is free text, required when `result` is `not_applicable`, empty
|
|
520
|
+
otherwise, carrying one of two canonical strings that distinguish a
|
|
521
|
+
non-regression from a regression: `no definition recorded` (a
|
|
522
|
+
prior-round probe recorded with only an id, no definition to reapply)
|
|
523
|
+
and `target text no longer present` (a replayed probe whose mutant can
|
|
524
|
+
no longer be applied). When the assignment names none, it returns
|
|
525
|
+
`mutation_probes: []` rather than
|
|
526
|
+
omitting the field, so 'none asked for' is distinguishable from 'asked
|
|
527
|
+
for and not reported'. Each item also carries `replayed`: `false` for a
|
|
528
|
+
probe newly introduced this round, `true` for a prior round's probe
|
|
529
|
+
replayed this round under the replay rule in step 6. On any round after
|
|
530
|
+
the task's first, the implementer replays every probe named in an
|
|
531
|
+
earlier round of this task (on the task's first round there are none),
|
|
532
|
+
naming each by its mutant definition, not merely by its id, not only
|
|
533
|
+
this round's new probes, before the next reviewer spawn, reporting each
|
|
534
|
+
one in `mutation_probes` alongside the round's new probes. A replayed
|
|
535
|
+
probe whose `expectation` is now `violated`, or which can no longer be
|
|
536
|
+
applied (reason: `target text no longer present`), is the regression
|
|
537
|
+
signal, reported as such and resolved before the next reviewer spawn;
|
|
538
|
+
`result` alone is not a regression signal, and a probe recorded with
|
|
539
|
+
only an id and no definition to reapply is `not_applicable` (reason:
|
|
540
|
+
`no definition recorded`).
|
|
497
541
|
|
|
498
542
|
The `commits` field lists the full sha of every commit the implementer
|
|
499
543
|
produced on the task branch, in the order produced; when the task
|
|
@@ -501,12 +545,17 @@ produced no commit, the implementer returns `commits: []` rather than
|
|
|
501
545
|
omitting the field, so 'did not commit' is distinguishable from
|
|
502
546
|
'forgot to report'.
|
|
503
547
|
|
|
548
|
+
For a non-empty `commits` field, the implementer pastes `git log
|
|
549
|
+
--reverse --format=%H <base>..HEAD`; it never types or hand-completes commit
|
|
550
|
+
shas. Verification plans, probe plans, and repeat tallies run in the foreground, and the
|
|
551
|
+
implementer reports their returns in the same turn as the last check. A
|
|
552
|
+
background monitor is no substitute for those returns.
|
|
553
|
+
|
|
504
554
|
## Reviewer output contract
|
|
505
555
|
|
|
506
556
|
The output shape remains the same for either selected contract. Compare the
|
|
507
557
|
delegated versioned records and producer evidence under Contract selection
|
|
508
558
|
above; a recommendation does not replace orchestrator acceptance.
|
|
509
|
-
|
|
510
559
|
```yaml
|
|
511
560
|
status: reviewed
|
|
512
561
|
role: reviewer
|
|
@@ -519,6 +568,7 @@ findings:
|
|
|
519
568
|
description: ""
|
|
520
569
|
suggested_fix: ""
|
|
521
570
|
recurrence: new | repeated
|
|
571
|
+
introduced_by_delta: yes | no | unknown
|
|
522
572
|
acceptance_recommendation: accept | accept_with_notes | fix_required | reject
|
|
523
573
|
missing_tests:
|
|
524
574
|
- ""
|
|
@@ -534,7 +584,6 @@ withdrawn:
|
|
|
534
584
|
- description: ""
|
|
535
585
|
reason: ""
|
|
536
586
|
```
|
|
537
|
-
|
|
538
587
|
`acceptance_recommendation` is mandatory: every reviewer return must set it.
|
|
539
588
|
When it is missing, the orchestrator asks the reviewer to resupply it
|
|
540
589
|
instead of inferring one from the findings list.
|
|
@@ -544,6 +593,7 @@ task: `new` for a defect class not previously found here, `repeated` for
|
|
|
544
593
|
one that already appeared in an earlier round. On a task's first review
|
|
545
594
|
round every finding is `new` by definition. This is what feeds the
|
|
546
595
|
Review-round escalation budget's trigger.
|
|
596
|
+
`introduced_by_delta` records whether a finding is attributable to the reviewed delta: `no` requires a named base build and replay in `reproduction`, is transferred parenthetically in the `Description` field of `05-review-findings.md` without renaming `Severity`/`Decision`, and follows the ordinary gate; only `yes`/`unknown` participate in bounded-round rules.
|
|
547
597
|
|
|
548
598
|
`method_applied` echoes the `review_method` named in the briefing (see step
|
|
549
599
|
7); `withdrawn` lists each finding the reviewer proposed and then retracted
|
|
@@ -715,7 +765,7 @@ review and never satisfies the review gate, since review is never skipped.
|
|
|
715
765
|
The signal: a review round finds a new defect of the same class a previous
|
|
716
766
|
round's fix already addressed, so the class has recurred once after being
|
|
717
767
|
fixed, and the next fix would again be case-by-case enumeration (boundary
|
|
718
|
-
tokens, spellings, and similar one-off patches). Stop the first time this
|
|
768
|
+
tokens, spellings, and similar one-off patches). Apply this signal only to `introduced_by_delta: yes`/`unknown`; `no` continues through the ordinary finding gate. Stop the first time this
|
|
719
769
|
signal fires: the recurrence is already the class's second occurrence, so
|
|
720
770
|
do not wait for a third one before stopping. Name the structural cause in
|
|
721
771
|
one sentence, and decide to split or redesign rather than keep accreting
|
|
@@ -733,11 +783,9 @@ and across repeated review rounds, so effort does not keep accumulating
|
|
|
733
783
|
unaided: by the second round-2 halt signal on the same task, or by the
|
|
734
784
|
third `fix_required` review round on the same task, whichever comes
|
|
735
785
|
first, choose one of three escalations instead of running another round
|
|
736
|
-
the same way. A
|
|
737
|
-
`
|
|
738
|
-
|
|
739
|
-
chosen once the third such round has returned, before the next attempt
|
|
740
|
-
starts. The escalation is chosen in addition to the halt rule's
|
|
786
|
+
the same way. A negative round has an `acceptance_recommendation` of
|
|
787
|
+
`fix_required` or `reject`; a misfired review is not a round (see Subagent
|
|
788
|
+
misfire rule). A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. The escalation is chosen in addition to the halt rule's
|
|
741
789
|
split-or-redesign response, not instead of it.
|
|
742
790
|
|
|
743
791
|
- **Tier or model escalation**: raise the implementer to at least
|
|
@@ -13,7 +13,7 @@ critical waivers remain distinct. -->
|
|
|
13
13
|
|
|
14
14
|
## Review-round escalation
|
|
15
15
|
|
|
16
|
-
<!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third
|
|
16
|
+
<!-- One row per task that triggers the Review-round escalation budget in SKILL.md: the second round-2 halt signal or the third negative round on that task. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. A run carries multiple tasks, so this table can carry multiple rows. Leave the single placeholder row as n/a when no task in this run has triggered the budget. -->
|
|
17
17
|
|
|
18
18
|
| Task | Choice | Reason |
|
|
19
19
|
|---|---|---|
|
|
@@ -68,9 +68,15 @@ an optional row cannot stand in for a required criterion.
|
|
|
68
68
|
|
|
69
69
|
### Mutation Probes
|
|
70
70
|
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
71
|
+
Before/After cells hold a single-line excerpt. When the mutant's actual
|
|
72
|
+
before/after text is multi-line or contains an unescaped `|`, or the mutant
|
|
73
|
+
is a patch/diff rather than a text swap, put the full text or diff in the
|
|
74
|
+
implementer report or a fenced block directly under the table, and note
|
|
75
|
+
where it lives in the row's own cell.
|
|
76
|
+
|
|
77
|
+
| Round | Mutant | File | Anchor | Before | After | Verified Applied Via | Result | Expectation | Reason | Restored Verified | Replayed |
|
|
78
|
+
|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
79
|
+
| <!-- round --> | <!-- mutant --> | <!-- file --> | <!-- anchor --> | <!-- before --> | <!-- after --> | <!-- verified_applied_via --> | <!-- result --> | <!-- expectation --> | <!-- reason --> | <!-- restored_verified --> | <!-- replayed --> |
|
|
74
80
|
|
|
75
81
|
## Risks / Notes
|
|
76
82
|
|
|
@@ -16,7 +16,8 @@ by the grounding-mcp completeness reader yet).
|
|
|
16
16
|
| Severity | Category | Description | Suggested Fix | Decision |
|
|
17
17
|
|---|---|---|---|---|
|
|
18
18
|
| low/medium/high/critical | correctness/architecture/security/tests/maintainability/performance/docs | <!-- finding --> | <!-- fix --> | accepted/defer |
|
|
19
|
-
<!-- This row
|
|
19
|
+
<!-- This legacy five-cell table and placeholder row are the shipped template, not a finding: the orchestrator-workflow completeness reader matches the row byte-for-byte and fails the completeness gate closed when it survives untouched with no concrete finding row, the same way a `TODO` marker does. During findings transfer (step 7), replace this row with each reviewer finding and record its attribution parenthetically in the Description field as `(introduced_by_delta: yes|no|unknown)`. For a genuine zero-findings review, delete this row instead — a header row with no data rows is a valid, complete table; leaving this row next to real finding rows is also fine. This mirrors grounding-mcp's placeholder-row detection; keep the two in sync. -->
|
|
20
|
+
<!-- A `no` classification requires the named base build and replay recorded in the reviewer's `reproduction`; it follows the ordinary finding gate, while only `yes`/`unknown` feed bounded-round halt and escalation guidance. The load-bearing Severity and Decision headers remain unchanged. -->
|
|
20
21
|
|
|
21
22
|
## Missing Tests
|
|
22
23
|
|
|
@@ -35,4 +36,3 @@ accept | accept_with_notes | fix_required | reject
|
|
|
35
36
|
<!-- Reproduction note: when a finding rests on empirical or probabilistic evidence (flake rates, benchmarks, "n runs green", performance/timing numbers), record the reviewer's independent reproduction (method, sample size, result vs. the implementer's claim) in the reviewer output contract's `reproduction` field (SKILL.md step 7). Deterministic checks (a single test run, tsc, lint) do not require it. -->
|
|
36
37
|
|
|
37
38
|
<!-- Recurrence note: each finding in the reviewer output contract also carries a `recurrence` field (new or repeated), letting the orchestrator read the Review-round escalation budget's trigger (SKILL.md, Review-round escalation budget) off the reviewer's own return instead of reconstructing it by hand. A repeated finding here is what feeds that budget's round count. -->
|
|
38
|
-
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "orchestrator-workflow",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.34.0",
|
|
4
4
|
"description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
|
|
5
5
|
"main": "dist/index.js",
|
|
6
6
|
"type": "module",
|