orchestrator-workflow 0.35.0 → 0.37.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +209 -0
- package/README.md +56 -0
- package/assets/agents/implementer.md +14 -1
- package/assets/skill/references/contracts.md +38 -1
- package/assets/skill/references/evidence-and-probes.md +4 -4
- package/assets/skill/references/review-and-recovery.md +66 -0
- package/dist/cli.js +95 -0
- package/dist/review-report.d.ts +214 -0
- package/dist/review-report.js +554 -0
- package/package.json +3 -2
package/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,215 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.37.0] - 2026-09-18
|
|
11
|
+
|
|
12
|
+
- The two verdict fields of a mutation probe now carry a legend, a source
|
|
13
|
+
and a check. `assets/agents/implementer.md`, the normative site for the
|
|
14
|
+
output-field semantics, says: "`result: killed` means the probe's test
|
|
15
|
+
command reacted to the mutant under the runner's own pass predicate and
|
|
16
|
+
`survived` means it did not; `expectation: met` means that outcome is what
|
|
17
|
+
the probe was expected to show and `violated` means it is not; both are
|
|
18
|
+
`not_applicable` when no `result` was measured." and: "When the probe
|
|
19
|
+
runner states a machine-readable verdict, copy whichever of `result` and
|
|
20
|
+
`expectation` it states from it verbatim, never from your own reading of
|
|
21
|
+
the test output; when it states only `result`, set `expectation` by
|
|
22
|
+
comparing that verdict with the probe's declared expectation. Quote the
|
|
23
|
+
runner's verdict for each probe in `tests.executed`, and say there when
|
|
24
|
+
`expectation` was set this way, so both fields can be checked against it."
|
|
25
|
+
`references/contracts.md` carries both sentences verbatim for the
|
|
26
|
+
orchestrator. Step 6 of `references/evidence-and-probes.md` adds the
|
|
27
|
+
orchestrator's side: "Before transferring a probe row, compare its
|
|
28
|
+
`result` and `expectation` with the runner verdict quoted in
|
|
29
|
+
`tests.executed`; on a mismatch, or when a verdict the runner states is
|
|
30
|
+
not quoted, resupply it (ask the same implementer for the verdict, respawn
|
|
31
|
+
one when it is gone, or rerun the probe yourself in isolation), record the
|
|
32
|
+
resupply in `03-decisions.md`, and treat it as a transfer blocker rather
|
|
33
|
+
than a misfire, since the return itself parses; never infer either field.
|
|
34
|
+
A quoted probe verdict is not a named result of the verification set, so
|
|
35
|
+
the set's missing-or-extra rule does not apply to it." The contract shape,
|
|
36
|
+
the enums and the eleven `mutation_probes` sub-fields are unchanged, no
|
|
37
|
+
tool is named, and every sentence was appended to an existing line so no
|
|
38
|
+
cited line moved (the prompt is cited by line throughout the knowledge
|
|
39
|
+
bundle); the prompt change reaches an existing install with the next kit
|
|
40
|
+
re-install, which also re-renders the tier variants. Evidence (issue #300,
|
|
41
|
+
one run, not a benchmark): an implementer returned all six probes as
|
|
42
|
+
`survived` with `expectation: violated` while its own evidence showed the
|
|
43
|
+
suite failing on every mutant, and the mislabel was caught only because
|
|
44
|
+
the orchestrator compared the fields with that evidence by hand. Pinned in
|
|
45
|
+
`test/probe-plans-recovery.test.ts`.
|
|
46
|
+
- Three ceremony rules in the skill references, none of which changes the
|
|
47
|
+
review gate, the waiver rules or the AGENTS.md section (byte-identical, so
|
|
48
|
+
no AGENTS.md re-install is needed; the rules reach an existing install
|
|
49
|
+
with the next kit re-install). Step 6 of
|
|
50
|
+
`references/evidence-and-probes.md` now says: "Record a baseline revision
|
|
51
|
+
only when scope or the normative text of a criterion changes, including a
|
|
52
|
+
change to what its verification checks; a wording precision that leaves
|
|
53
|
+
the check itself unchanged is a `03-decisions.md` entry, not a revision";
|
|
54
|
+
the orchestrator records that entry, states in it why no evidence is
|
|
55
|
+
invalidated, and communicates the corrected wording in the next
|
|
56
|
+
delegation. Step 7 now says: "For a review round whose entire delta is a
|
|
57
|
+
docs-only delta in the sense of step 8's docs-only closure, default to the
|
|
58
|
+
`-medium` reviewer tier with `review_method: normal` where tier variants
|
|
59
|
+
are installed". That refines the general tier default for this one class
|
|
60
|
+
only, a round that touches an instruction, policy, template or prompt file
|
|
61
|
+
keeps the general default, and the minimum review methods are untouched;
|
|
62
|
+
the AGENTS.md section does not yet point to this refinement.
|
|
63
|
+
`references/review-and-recovery.md` gains a "Pinned-prose changes" section
|
|
64
|
+
for a change whose acceptance rests on tests that pin documentation
|
|
65
|
+
wording: "A prose mutant survives exactly when its bytes sit in no
|
|
66
|
+
assertion", so review rounds that hunt for the next unpinned sentence do
|
|
67
|
+
not converge. The section asks for one normative site per rule, a claim
|
|
68
|
+
list in the acceptance criterion as the pin obligation (every normative
|
|
69
|
+
sentence the change adds or alters at that site is a claim, an omission is
|
|
70
|
+
named with its reason), a reviewer briefing that bounds the prose mutant
|
|
71
|
+
space to that list, copies bound to the normative site by one shared test
|
|
72
|
+
constant, and: "Cap test-adequacy review rounds on the change at two." It
|
|
73
|
+
defines the capped round, exempts semantic findings, and changes neither
|
|
74
|
+
the Round-2 halt rule, the escalation budget, the Fix-regression decision
|
|
75
|
+
point nor the review gate. Step 7 points to the section without restating
|
|
76
|
+
it. Evidence (issue #300 and the change that added the Fix-regression
|
|
77
|
+
decision point; one repository each, not a benchmark): the issue reports
|
|
78
|
+
baseline revisions r1 to r3 for two wording precisions of a verification
|
|
79
|
+
method, and a run in which the top reviewer tier was about half the day's
|
|
80
|
+
cost across nine reviewer rounds and an advisor, where the documentation
|
|
81
|
+
rounds did not need that tier; the decision point itself, about 27 lines
|
|
82
|
+
of rule text, took four review rounds with unpinned prose reported in
|
|
83
|
+
every one, until the trigger was held in one test constant and the mutant
|
|
84
|
+
space was bounded. Pinned in `test/probe-plans-recovery.test.ts`.
|
|
85
|
+
- `references/review-and-recovery.md` gains a "Fix-regression decision
|
|
86
|
+
point" between the Review-round escalation budget and the Final
|
|
87
|
+
acceptance rule: when the review of a fix round reports at least one
|
|
88
|
+
`high` or `critical` finding that the previous round's review did not
|
|
89
|
+
report, with `introduced_by_delta: yes`, the orchestrator names in one
|
|
90
|
+
sentence why the fix could introduce it and records one of four outcomes
|
|
91
|
+
in `03-decisions.md` (continue with the stated reason, redesign, split,
|
|
92
|
+
hold) before another fix round starts. It is a decision point, not a
|
|
93
|
+
halt, and changes neither the Round-2 halt rule nor the budget. The
|
|
94
|
+
rule is defined only in that section; step 8 of
|
|
95
|
+
`references/evidence-and-probes.md` points to it without restating the
|
|
96
|
+
trigger. The AGENTS.md section is unchanged, so no AGENTS.md re-install
|
|
97
|
+
is needed; the rule reaches an existing install with the next kit
|
|
98
|
+
re-install, like any other skill-reference change. Evidence (issue #300,
|
|
99
|
+
one repository, one model mix, not a benchmark): in a five-round run the
|
|
100
|
+
third review already showed that the fix had introduced new high
|
|
101
|
+
findings, but the defect class only recurred in the fourth review, so
|
|
102
|
+
the halt rule fired one round, about an hour, after the structural
|
|
103
|
+
problem was visible. Pinned in `test/probe-plans-recovery.test.ts`.
|
|
104
|
+
- `validate-review-report` now checks that every element of a string-array
|
|
105
|
+
field (`summary`, `missing_tests`, `residual_risks`) is a string, one
|
|
106
|
+
diagnostic per offending element at `<field>[<index>]`; and
|
|
107
|
+
`extractYamlSource` now prefers a fenced block tagged `yaml`/`yml` when
|
|
108
|
+
several fences are present, falling back to the first fence only when
|
|
109
|
+
none carries that tag, with a warning naming the earlier fence it
|
|
110
|
+
skipped. The fence tag is now matched against the whole info string's
|
|
111
|
+
first whitespace-delimited word rather than a leading run of letters,
|
|
112
|
+
so a tag followed by attributes (a fence opened `yaml title=x`) is
|
|
113
|
+
recognized as `yaml` instead of matching no fence at all. The "prose
|
|
114
|
+
found before" warning no longer also fires when the text preceding the
|
|
115
|
+
preferred fence is exactly the skipped fence(s) plus whitespace, so a
|
|
116
|
+
skipped fence is no longer double-reported as both skipped and prose.
|
|
117
|
+
The opening fence's backtick run is now consumed whole and the closing
|
|
118
|
+
fence must repeat at least as many backticks at column 0 with nothing
|
|
119
|
+
but whitespace after it, so a return opened with four backticks is
|
|
120
|
+
recognized as `yaml` (the run's leftover backticks previously landed in
|
|
121
|
+
the info string, making the tag itself start with a backtick) and ends
|
|
122
|
+
at its own four-backtick run rather than at a three-backtick line
|
|
123
|
+
inside it. Widening the info-string capture also widens what counts as
|
|
124
|
+
a fence at all: any attribute-bearing fence, and any fence opened with
|
|
125
|
+
more than three backticks, is now a fence, so input whose only fence
|
|
126
|
+
was opened `js title=x` is treated as fenced rather than as literal
|
|
127
|
+
YAML. Only whitespace-separated attributes count toward the tag:
|
|
128
|
+
`yaml title=x` counts, `yaml,title=x` does not, its first word being
|
|
129
|
+
the whole string. That opening run is now matched whole and never
|
|
130
|
+
re-entered at a shorter length, which bounds the scan: a run with no
|
|
131
|
+
valid closer was previously retried at every shorter run length from
|
|
132
|
+
every offset inside the run, each retry rescanning the block body, so
|
|
133
|
+
a 2000-backtick run in a 22 KB return took roughly 50 seconds through
|
|
134
|
+
`extractYamlSource` where it now takes under a millisecond. Keeping
|
|
135
|
+
the run whole also makes the closing-length rule literal: a return
|
|
136
|
+
whose opening run is longer than every closing run in it is no fence
|
|
137
|
+
at all and reaches the parser whole, where splitting the run
|
|
138
|
+
previously matched it and read its leftover backticks as the start of
|
|
139
|
+
the tag.
|
|
140
|
+
|
|
141
|
+
- `check-release-changelogs`'s `VERSION_HEADING_RE` (rule 1, version-heading)
|
|
142
|
+
now accepts the same prerelease identifier class as `parseSemver`
|
|
143
|
+
(`[0-9A-Za-z.-]`, including the hyphen), sourced from one shared
|
|
144
|
+
`PRERELEASE_IDENTIFIER_CHARS` constant so the two cannot drift apart
|
|
145
|
+
again: a package.json version with a hyphenated prerelease tag such as
|
|
146
|
+
`1.0.0-alpha-1` parsed fine through `parseSemver` but previously never
|
|
147
|
+
matched the heading regex, which only accepted `[\w.]`. The previous
|
|
148
|
+
0.35.0 hardening bullet named the new rules without naming three
|
|
149
|
+
details of their own contract: the `--expect <csv>` option overrides
|
|
150
|
+
rule 5's (checked-package-scope) expectation list, and an empty csv
|
|
151
|
+
(`--expect ""`) opts out of rule 5 entirely rather than passing it
|
|
152
|
+
vacuously; rule 5 itself only runs when that expectation list is
|
|
153
|
+
non-empty, so a fixture whose packages do not share this repo's names
|
|
154
|
+
can still run the other four rules without a spurious finding; and
|
|
155
|
+
rule 2's (fresh-unreleased) direction check compares versions by
|
|
156
|
+
semver precedence, not string inequality, with build metadata stripped
|
|
157
|
+
before the comparison since two versions differing only in build
|
|
158
|
+
metadata carry equal precedence per the semver spec. The shared class
|
|
159
|
+
is narrower than `\w`: an underscore prerelease heading such as
|
|
160
|
+
`1.0.0-alpha_1` is no longer accepted either, since semver's own
|
|
161
|
+
prerelease grammar forbids underscores and `parseSemver` already
|
|
162
|
+
rejected such a version before this change.
|
|
163
|
+
|
|
164
|
+
## [0.36.0] - 2026-09-16
|
|
165
|
+
|
|
166
|
+
- A `validate-review-report <file>` CLI subcommand (`-` reads stdin) checks a
|
|
167
|
+
reviewer return's YAML against the reviewer output contract's required
|
|
168
|
+
fields and enums, printing one diagnostic per missing or invalid field and
|
|
169
|
+
exiting `0`/`1`/`2` for valid/structurally-invalid/usage-error; `--format
|
|
170
|
+
json` prints the same diagnostics as one JSON object. The check is
|
|
171
|
+
structural only, never semantic, and never constitutes orchestrator
|
|
172
|
+
acceptance. Its schema lives in `src/review-report.ts` and is pinned in
|
|
173
|
+
`test/docs-consistency.test.ts` against the contract block in
|
|
174
|
+
`assets/agents/reviewer.md` itself, so a contract edit without a matching
|
|
175
|
+
schema edit fails the suite instead of drifting silently. An excess
|
|
176
|
+
positional argument (e.g. two file paths) now also exits `2` as a usage
|
|
177
|
+
error instead of silently validating only the first path. A fenced
|
|
178
|
+
return ends at the first closing fence that starts at column 0, so a
|
|
179
|
+
triple-backtick sequence inside a value (a reviewer quoting a fenced
|
|
180
|
+
snippet in a `description`) no longer closes the block early and hands
|
|
181
|
+
the parser a truncated document.
|
|
182
|
+
- Removed incidental blank-line padding (a run of seven consecutive blank
|
|
183
|
+
lines) and a mid-sentence paragraph split from
|
|
184
|
+
`docs/okf/subagent-contracts-superset.md`'s mutation-probe field
|
|
185
|
+
discussion, and corrected its Commits field section, which still said
|
|
186
|
+
the not-applicable `commits: []` clause is pinned "in both copies"
|
|
187
|
+
after the 0.35.0 contract-reduction refactor left it in the installed
|
|
188
|
+
implementer prompt alone. Re-pointed every citation the removed lines
|
|
189
|
+
shifted (two `docs/okf/log.md` self-citations into this doc, and the
|
|
190
|
+
`SIBLING_GUARD_BUNDLE_ALLOWLIST` explanation comment in
|
|
191
|
+
`test/docs-consistency.test.ts`, whose recorded `:1217`/`:1232`
|
|
192
|
+
coordinates no longer matched the paragraph's current citations).
|
|
193
|
+
Docs-only: no YAML contract, role prompt, guard matcher, or exemption
|
|
194
|
+
geometry changed. Anchored by agent-dx tracker task 8a55e082.
|
|
195
|
+
|
|
196
|
+
- `assets/agents/implementer.md` gains a pre-return rule bullet next to
|
|
197
|
+
the commit-reporting ones: before committing, when slop-detector is
|
|
198
|
+
available run `slop-detector check <changed file> [<changed file> ...]
|
|
199
|
+
--pack review-slop` over every changed file and `git log -1
|
|
200
|
+
--format=%B | slop-detector check --stdin-path COMMIT_MSG --pack
|
|
201
|
+
review-slop` over the commit message, with the repository-vendored
|
|
202
|
+
`node packages/slop-detector/dist/cli.js check ...` path named as the
|
|
203
|
+
alternative where the CLI is not installed on PATH; fix every
|
|
204
|
+
block-level finding before returning, or add a legitimate match to
|
|
205
|
+
`review.allow`. Only exit `0` or `1` is a result: exit `2` is a usage
|
|
206
|
+
error, not a clean check. A returned report that skipped the check on a
|
|
207
|
+
diff with block-level findings is a misfire, not evidence. Pinned by
|
|
208
|
+
`test/docs-consistency.test.ts`;
|
|
209
|
+
`docs/okf/subagent-contracts-superset.md` gained a matching sentence
|
|
210
|
+
describing the rule.
|
|
211
|
+
- Anchored by pandora batch 51 (`.ai/runs/2026-09-13-quickwins-batch51`):
|
|
212
|
+
four review rounds (or post-merge cleanup commits) across five repos
|
|
213
|
+
were spent catching run-local review tokens (finding ids, round
|
|
214
|
+
references, workspace-handoff phrases) that nothing mechanical
|
|
215
|
+
flagged before `slop-detector`'s new `review-slop` pack existed; see
|
|
216
|
+
that package's own CHANGELOG.md for the pack itself and the
|
|
217
|
+
pre-cleanup PR the fixtures were drawn from.
|
|
218
|
+
|
|
10
219
|
## [0.35.0] - 2026-09-15
|
|
11
220
|
|
|
12
221
|
- The npm tarball now ships a `LICENSE` file matching the repo root LICENSE
|
package/README.md
CHANGED
|
@@ -691,3 +691,59 @@ references.
|
|
|
691
691
|
package's version, so a release of `okf-kit` must bump those pins in the
|
|
692
692
|
same commit as the version cut; see `CONTRIBUTING.md`'s "Releasing okf-kit"
|
|
693
693
|
section (repo root) for the order.
|
|
694
|
+
|
|
695
|
+
## Reviewer-report validation
|
|
696
|
+
|
|
697
|
+
```bash
|
|
698
|
+
orchestrator-workflow validate-review-report path/to/return.yaml
|
|
699
|
+
orchestrator-workflow validate-review-report - < path/to/return.yaml
|
|
700
|
+
orchestrator-workflow validate-review-report path/to/return.yaml --format json
|
|
701
|
+
```
|
|
702
|
+
|
|
703
|
+
Checks a reviewer return's YAML against the reviewer output contract's
|
|
704
|
+
required fields and enums (see the "Reviewer output contract" section of
|
|
705
|
+
`assets/skill/references/contracts.md`, byte-identical to the contract in
|
|
706
|
+
`assets/agents/reviewer.md`), whether the return is fenced in a code
|
|
707
|
+
block (any language tag, or none) or given unfenced, and prints one
|
|
708
|
+
diagnostic per missing or invalid field. Every element of a string-array
|
|
709
|
+
field (`summary`, `missing_tests`, `residual_risks`) must itself be a
|
|
710
|
+
string; a non-string element (a number, a mapping, a boolean, or `null`
|
|
711
|
+
-- written as a bare or `~` bullet) is its own diagnostic at
|
|
712
|
+
`<field>[<index>]`. A fenced return ends at the first closing fence that
|
|
713
|
+
starts at column 0, repeats at least as many backticks as the opening
|
|
714
|
+
fence, and carries nothing but whitespace after that run, so neither a
|
|
715
|
+
reviewer quoting a fenced snippet inside a value (a `description` block
|
|
716
|
+
scalar, which YAML indents) nor one wrapping a return in four backticks
|
|
717
|
+
around a snippet fenced at column 0 truncates the return. A return
|
|
718
|
+
with no closing fence satisfying all three is not fenced at all, so its
|
|
719
|
+
whole text reaches the parser; that includes one whose opener is longer
|
|
720
|
+
than every closing run present. When the return carries more than one
|
|
721
|
+
fenced block, the first one whose fence tag's first word is `yaml` or
|
|
722
|
+
`yml` is validated, case-insensitively and counting whitespace-separated
|
|
723
|
+
attributes (`yaml title=x` counts; `yaml,title=x` does not, its first
|
|
724
|
+
word being the whole string), falling back to the first fence only when
|
|
725
|
+
none carries that word; a warning names any earlier fence skipped this
|
|
726
|
+
way. This preference can validate a later worked example instead of an
|
|
727
|
+
earlier, real but unfenced return: a reviewer who leaves their own return
|
|
728
|
+
unfenced and then quotes a `yaml`-tagged example afterward has that
|
|
729
|
+
example validated instead, which the emitted warning also names.
|
|
730
|
+
`--format json` prints the same diagnostics as a single JSON object
|
|
731
|
+
instead of human-readable text. It exits `0` when the return is
|
|
732
|
+
structurally valid, `1` when it is structurally invalid (a required field
|
|
733
|
+
is missing or its value falls outside its enum, or the input is
|
|
734
|
+
unparsable, empty, or not a mapping), and `2` for a usage error (an
|
|
735
|
+
unreadable file, an unrecognized `--format` value, a missing `<file>`
|
|
736
|
+
argument, an unknown option, or an excess positional argument).
|
|
737
|
+
`--format json` governs the validation verdict only: a commander parsing
|
|
738
|
+
error (missing argument, unknown option, excess arguments) or an
|
|
739
|
+
unrecognized `--format` value itself still prints plain text to stderr
|
|
740
|
+
with nothing on stdout, regardless of `--format`; the one exception is an
|
|
741
|
+
unreadable file, which does emit the JSON envelope on stdout. This check
|
|
742
|
+
is structural only: it never judges semantic adequacy, cannot waive a
|
|
743
|
+
finding, and passing it is never orchestrator acceptance. The
|
|
744
|
+
required-field set it checks is hand-maintained in `src/review-report.ts`
|
|
745
|
+
and pinned against the contract block itself by
|
|
746
|
+
`test/docs-consistency.test.ts`, so a contract edit without a matching
|
|
747
|
+
schema edit fails the suite instead of drifting silently; every field
|
|
748
|
+
listed there is dispatched to its own checker, so an entry added to the
|
|
749
|
+
list without a checker fails to typecheck rather than passing unchecked.
|
|
@@ -103,7 +103,7 @@ Rules:
|
|
|
103
103
|
from its result fields (`verified_applied_via`, `result`, `expectation`,
|
|
104
104
|
`reason`, `restored_verified`), take the definition fields from that
|
|
105
105
|
mutant record so the copied report still carries all eleven
|
|
106
|
-
`mutation_probes` sub-fields.
|
|
106
|
+
`mutation_probes` sub-fields. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
|
|
107
107
|
- Run every long test, build, or mutation-probe command in the foreground
|
|
108
108
|
and wait for it to finish before returning. When one foreground call
|
|
109
109
|
cannot hold it to completion, poll the backgrounded run to completion
|
|
@@ -133,6 +133,19 @@ Rules:
|
|
|
133
133
|
than omitting the field.
|
|
134
134
|
- Populate a non-empty `commits` field by pasting `git log --reverse
|
|
135
135
|
--format=%H <base>..HEAD`; never type or hand-complete commit shas.
|
|
136
|
+
- Before committing, when slop-detector is available run `slop-detector
|
|
137
|
+
check <changed file> [<changed file> ...] --pack review-slop` over every
|
|
138
|
+
changed file, and `git log -1 --format=%B | slop-detector check
|
|
139
|
+
--stdin-path COMMIT_MSG --pack review-slop` over the commit message;
|
|
140
|
+
where it is vendored in the repository rather than installed on PATH,
|
|
141
|
+
the same two invocations run as `node
|
|
142
|
+
packages/slop-detector/dist/cli.js check ...`. Fix every block-level
|
|
143
|
+
finding before returning, or add a legitimate match to `review.allow`
|
|
144
|
+
in the repository's slop.config.yml rather than deleting correct text.
|
|
145
|
+
Only exit `0` or `1` is a result; exit `2` is a usage error (a mistyped
|
|
146
|
+
invocation, or `--stdin-path` with nothing piped in), so it is not a
|
|
147
|
+
clean check. A returned report that skipped this check on a diff with
|
|
148
|
+
block-level findings is a misfire, not evidence.
|
|
136
149
|
- Verification plans, probe plans, and repeat tallies run in the foreground,
|
|
137
150
|
and the implementer reports their returns in the same turn as the last
|
|
138
151
|
check. A background monitor is no substitute for those returns.
|
|
@@ -126,7 +126,7 @@ commits:
|
|
|
126
126
|
Follow [evidence-and-probes.md workflow step 6](evidence-and-probes.md#workflow)
|
|
127
127
|
for implementation evidence, verification, mutation probes, and replay. For
|
|
128
128
|
output-field semantics and commit reporting, follow the installed
|
|
129
|
-
implementer role prompt. Return the selected contract's YAML envelope.
|
|
129
|
+
implementer role prompt. Return the selected contract's YAML envelope. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
|
|
130
130
|
|
|
131
131
|
## Reviewer output contract
|
|
132
132
|
|
|
@@ -165,6 +165,43 @@ withdrawn:
|
|
|
165
165
|
When it is missing, the orchestrator asks the reviewer to resupply it
|
|
166
166
|
instead of inferring one from the findings list.
|
|
167
167
|
|
|
168
|
+
A structural check for this exact contract ships as a CLI subcommand:
|
|
169
|
+
`orchestrator-workflow validate-review-report <file>` (pass `-` to read
|
|
170
|
+
the return from stdin instead) parses the reviewer return's YAML, fenced
|
|
171
|
+
in a code block with any language tag or none, or unfenced, and checks
|
|
172
|
+
the required fields and enums above one by one, printing a diagnostic
|
|
173
|
+
(path, expected, got) for each missing or invalid field; every element of
|
|
174
|
+
a string-array field (`summary`, `missing_tests`, `residual_risks`) must
|
|
175
|
+
itself be a string, with a diagnostic at `<field>[<index>]` for each one
|
|
176
|
+
that is not (a number, a mapping, a boolean, and `null` -- a bare or `~`
|
|
177
|
+
bullet -- are all rejected the same way). A fenced return ends at the
|
|
178
|
+
first closing fence that starts at column 0, repeats at least as many
|
|
179
|
+
backticks as the opening one, and carries nothing but whitespace after
|
|
180
|
+
that run, so a return wrapped in four backticks may quote a snippet
|
|
181
|
+
fenced in three without truncating itself; a return with no closing
|
|
182
|
+
fence satisfying all three is not fenced at all and reaches the parser
|
|
183
|
+
whole, including one whose opener is longer than every closing run
|
|
184
|
+
present. When more than one fenced block
|
|
185
|
+
is present, the first one whose fence tag's first word is `yaml` or `yml`
|
|
186
|
+
is validated, case-insensitively and counting whitespace-separated
|
|
187
|
+
attributes (`yaml title=x` counts; `yaml,title=x` does not, its first
|
|
188
|
+
word being the whole string), falling back to the first fence only when
|
|
189
|
+
none carries that word, with a warning naming any earlier fence skipped
|
|
190
|
+
this way. The same preference can validate a later worked example instead
|
|
191
|
+
of an earlier, real but unfenced return, which the emitted warning also
|
|
192
|
+
names. Add `--format json` for the same diagnostics as a single JSON
|
|
193
|
+
object. It exits `0` when the return is structurally valid, `1` when it
|
|
194
|
+
is structurally invalid (a missing or out-of-enum required field, or
|
|
195
|
+
unparsable, empty, non-mapping input), and `2` for a usage error (an
|
|
196
|
+
unreadable file, an unrecognized `--format` value, or an argument-parsing
|
|
197
|
+
error: a missing `<file>` argument, an unknown option, an excess
|
|
198
|
+
positional argument). `--format json` governs the validation verdict
|
|
199
|
+
only: an argument-parsing error or an unrecognized `--format` value still
|
|
200
|
+
prints plain text to stderr, except an unreadable file, which still emits
|
|
201
|
+
the JSON envelope on stdout. The check is structural only: it never
|
|
202
|
+
judges semantic adequacy, cannot waive a finding, and passing it is never
|
|
203
|
+
orchestrator acceptance.
|
|
204
|
+
|
|
168
205
|
`recurrence` classifies each finding against earlier rounds on the same
|
|
169
206
|
task: `new` for a defect class not previously found here, `repeated` for
|
|
170
207
|
one that already appeared in an earlier round. On a task's first review
|
|
@@ -112,7 +112,7 @@ directory and the subagents.
|
|
|
112
112
|
`03-decisions.md` and consolidate evidence in
|
|
113
113
|
`04-implementation-summary.md`, recording each probe the implementer
|
|
114
114
|
reports as a row in `04-implementation-summary.md`'s Mutation Probes
|
|
115
|
-
subsection, with the round it was named in. Each row's Before/After
|
|
115
|
+
subsection, with the round it was named in. Before transferring a probe row, compare its `result` and `expectation` with the runner verdict quoted in `tests.executed`; on a mismatch, or when a verdict the runner states is not quoted, resupply it (ask the same implementer for the verdict, respawn one when it is gone, or rerun the probe yourself in isolation), record the resupply in `03-decisions.md`, and treat it as a transfer blocker rather than a misfire, since the return itself parses; never infer either field. A quoted probe verdict is not a named result of the verification set, so the set's missing-or-extra rule does not apply to it. Each row's Before/After
|
|
116
116
|
cells hold a single-line excerpt; when the mutant's actual before/after
|
|
117
117
|
text is multi-line or contains an unescaped `|`, or the mutant is a
|
|
118
118
|
patch/diff rather than a text swap, the full text or diff goes in the
|
|
@@ -137,7 +137,7 @@ directory and the subagents.
|
|
|
137
137
|
acceptance; the coverage index is not a results database or acceptance
|
|
138
138
|
engine. Only the orchestrator can explicitly revise a baseline, recording
|
|
139
139
|
old/new revisions, affected IDs, authority and reason, invalidated evidence,
|
|
140
|
-
and verified rationale for carrying unchanged evidence forward.
|
|
140
|
+
and verified rationale for carrying unchanged evidence forward. Record a baseline revision only when scope or the normative text of a criterion changes, including a change to what its verification checks; a wording precision that leaves the check itself unchanged is a `03-decisions.md` entry, not a revision: the orchestrator records it, states in that entry why no evidence is invalidated, and communicates the corrected wording in the next delegation.
|
|
141
141
|
7. **Delegate review.** Send the diff to the reviewer subagent, naming in the
|
|
142
142
|
briefing the base and head revision the diff was generated from. When tier
|
|
143
143
|
variants are installed, pick the reviewer tier (the installed
|
|
@@ -152,7 +152,7 @@ directory and the subagents.
|
|
|
152
152
|
cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
|
|
153
153
|
never substitutes for it: do not pair `adversarial` with the `-medium`
|
|
154
154
|
reviewer tier, a budget mismatch that names probes without the effort to run
|
|
155
|
-
them; tiers themselves are unchanged by this axis. When the reviewer's
|
|
155
|
+
them; tiers themselves are unchanged by this axis. For a review round whose entire delta is a docs-only delta in the sense of step 8's docs-only closure, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed. This refines the general tier default above for that one class only: there `-medium` is the default and a higher tier is the non-default choice recorded with a one-line reason. A review round that touches an instruction, policy, template or prompt file keeps the general default, whatever the file type, and the minimums named above are unaffected. For a change whose acceptance rests on tests that pin documentation wording, write the briefing as the Pinned-prose changes section of [review and recovery](review-and-recovery.md) requires. When the reviewer's
|
|
156
156
|
environment cannot use version control to see the diff (for example a
|
|
157
157
|
policy-gated repository), supply the diff as a pre-generated file in the
|
|
158
158
|
briefing instead of expecting the reviewer to derive it, and have the
|
|
@@ -226,7 +226,7 @@ directory and the subagents.
|
|
|
226
226
|
halt signal across repeated review-fix cycles (see Round-2 halt rule
|
|
227
227
|
below). By the second round-2 halt signal or the third `fix_required`
|
|
228
228
|
review round on the same task, apply the Review-round escalation budget
|
|
229
|
-
(see below) instead of running another round unaided. At an advisor
|
|
229
|
+
(see below) instead of running another round unaided. When a fix round's review meets the trigger of the Fix-regression decision point (defined only in [review and recovery](review-and-recovery.md), not restated here), record the Fix-regression decision point before another fix round starts. At an advisor
|
|
230
230
|
trigger (architectural uncertainty, conflicting
|
|
231
231
|
requirements, a high-commitment fork among valid options, repeated
|
|
232
232
|
implementation failures, a review deadlock, a high-risk decision), the
|
|
@@ -82,6 +82,72 @@ through the reviewer subagent in full; this budget forces a change in
|
|
|
82
82
|
approach, not a shortcut past the review gate. Anchored by a measurement;
|
|
83
83
|
see the entry for this rule in the orchestrator-workflow CHANGELOG.
|
|
84
84
|
|
|
85
|
+
## Fix-regression decision point
|
|
86
|
+
|
|
87
|
+
The signal: the review of a fix round (any implementation round after the
|
|
88
|
+
task's first) reports at least one `high` or `critical` finding that the
|
|
89
|
+
previous round's review did not report, with `introduced_by_delta: yes`, so
|
|
90
|
+
the fix itself broke something. Read the qualifier off the findings of the
|
|
91
|
+
two reviews, not off `recurrence`: a `recurrence: repeated` finding that
|
|
92
|
+
the previous round's review did not report still triggers it. `unknown` and
|
|
93
|
+
`no` do not trigger this decision point: `no` continues through the
|
|
94
|
+
ordinary finding gate, and `unknown` keeps its existing treatment under the
|
|
95
|
+
Round-2 halt rule and the escalation budget. The signal needs no
|
|
96
|
+
recurrence: it fires even when the new finding's defect class has not
|
|
97
|
+
appeared on this task before, which is what separates it from the Round-2
|
|
98
|
+
halt rule above.
|
|
99
|
+
|
|
100
|
+
Before another fix round starts, name in one sentence why the fix could
|
|
101
|
+
introduce the defect (the structural cause, or the statement that there is
|
|
102
|
+
none), and record one of four outcomes as a decision in `03-decisions.md`:
|
|
103
|
+
continue with the stated reason, redesign, split, or hold (a merge-hold to
|
|
104
|
+
the operator). Spawning the advisor for this decision is optional.
|
|
105
|
+
|
|
106
|
+
This is a decision point, not a halt: continuing is a valid outcome, it is
|
|
107
|
+
not a round-2 halt signal, and it does not count toward the Review-round
|
|
108
|
+
escalation budget (the negative round itself still counts there as
|
|
109
|
+
before). It never replaces a review round. When the same review also fires
|
|
110
|
+
the Round-2 halt signal, the halt rule governs and this record is folded
|
|
111
|
+
into its split-or-redesign decision. Anchored by an observed run; see the
|
|
112
|
+
entry for this rule in the orchestrator-workflow CHANGELOG.
|
|
113
|
+
|
|
114
|
+
## Pinned-prose changes
|
|
115
|
+
|
|
116
|
+
This applies to a change whose acceptance rests on tests that pin
|
|
117
|
+
documentation wording (a rule text asserted by string match). A prose
|
|
118
|
+
mutant survives exactly when its bytes sit in no assertion, so a surviving
|
|
119
|
+
mutant alone says nothing about quality, and review rounds that hunt for
|
|
120
|
+
the next unpinned sentence do not converge. For such a change:
|
|
121
|
+
|
|
122
|
+
- Name one normative site per rule when slicing; every other site that
|
|
123
|
+
states the rule is a copy.
|
|
124
|
+
- List the load-bearing claims of the normative site in the acceptance
|
|
125
|
+
criterion, and pin each one as the whole sentence or clause that carries
|
|
126
|
+
it. That list is the pin obligation. Every normative sentence the change
|
|
127
|
+
adds or alters at that site is a claim; one left off the list is named
|
|
128
|
+
in the criterion with the reason it is not load-bearing.
|
|
129
|
+
- Bound the reviewer's prose mutant space to that list in the briefing. A
|
|
130
|
+
survivor outside the list is a scope note in the reviewer's
|
|
131
|
+
`residual_risks`, not a finding, unless the reviewer shows that the
|
|
132
|
+
unlisted sentence is load-bearing.
|
|
133
|
+
- Bind each copy to the normative site through one shared test constant,
|
|
134
|
+
and let a pointer point without restating the rule.
|
|
135
|
+
- Cap test-adequacy review rounds on the change at two. A test-adequacy
|
|
136
|
+
review round is one whose only unresolved findings are `tests` findings
|
|
137
|
+
about pin gaps on the pinned prose; a round with any other unresolved
|
|
138
|
+
finding is an ordinary round outside the cap. Pin gaps that remain
|
|
139
|
+
become accepted notes or a follow-up.
|
|
140
|
+
|
|
141
|
+
Semantic findings are exempt from the bound and from the cap: two sites
|
|
142
|
+
stating different rules, a contradiction with another rule, and a false
|
|
143
|
+
claim are defects at whatever severity they deserve. The cap changes
|
|
144
|
+
neither the Round-2 halt rule, the Review-round escalation budget nor the
|
|
145
|
+
Fix-regression decision point: a capped round still counts as a negative
|
|
146
|
+
round where it is one. The review gate is unchanged: a high or critical
|
|
147
|
+
finding of any category still blocks and is never capped away, and
|
|
148
|
+
accepting one follows the waiver rules. Anchored by an observed run; see
|
|
149
|
+
the entry for this rule in the orchestrator-workflow CHANGELOG.
|
|
150
|
+
|
|
85
151
|
## Final acceptance rule
|
|
86
152
|
|
|
87
153
|
Subagents provide evidence. The orchestrator decides. The operator receives
|
package/dist/cli.js
CHANGED
|
@@ -1172,6 +1172,101 @@ program
|
|
|
1172
1172
|
printTargetDetail(targetReport, PACKAGE_VERSION);
|
|
1173
1173
|
process.exitCode = exitCode;
|
|
1174
1174
|
});
|
|
1175
|
+
const validateReviewReportCommand = program
|
|
1176
|
+
.command("validate-review-report")
|
|
1177
|
+
.description("Check a reviewer return's YAML against the reviewer output contract's required fields and enums; reports structural validity ONLY, never semantic adequacy, and never waives a finding or constitutes orchestrator acceptance")
|
|
1178
|
+
.argument("<file>", "path to a file holding the reviewer return's YAML, or - to read stdin")
|
|
1179
|
+
.option("--format <format>", "output format: text (default) or json", "text")
|
|
1180
|
+
// commander 12's default for a subcommand is `allowExcessArguments:
|
|
1181
|
+
// true`, so `validate-review-report a.yaml extra.yaml` silently
|
|
1182
|
+
// validated only `a.yaml` and exited 0 (fix-round, review finding M1).
|
|
1183
|
+
// Disabling it turns a trailing extra argument into commander's own
|
|
1184
|
+
// "too many arguments" parsing error, which the scoped `exitOverride`
|
|
1185
|
+
// below then maps to exit 2 like every other usage error.
|
|
1186
|
+
.allowExcessArguments(false)
|
|
1187
|
+
.action(async (file, opts) => {
|
|
1188
|
+
// Imported dynamically, here rather than as a top-level static import,
|
|
1189
|
+
// so this command's addition cannot shift the line numbers of any
|
|
1190
|
+
// statement above it in this file: several docs/okf/ citations anchor
|
|
1191
|
+
// to exact lines in src/cli.ts, and a top-level import would have
|
|
1192
|
+
// re-pointed all of them for a reason unrelated to their own content.
|
|
1193
|
+
const { STRUCTURAL_ONLY_NOTE, validateReviewReport } = await import("./review-report.js");
|
|
1194
|
+
const format = opts.format ?? "text";
|
|
1195
|
+
if (format !== "text" && format !== "json") {
|
|
1196
|
+
console.error(`Unknown --format value: ${format} (expected "text" or "json")`);
|
|
1197
|
+
console.error(STRUCTURAL_ONLY_NOTE);
|
|
1198
|
+
process.exitCode = 2;
|
|
1199
|
+
return;
|
|
1200
|
+
}
|
|
1201
|
+
let raw;
|
|
1202
|
+
try {
|
|
1203
|
+
raw = readFileSync(file === "-" ? 0 : file, "utf8");
|
|
1204
|
+
}
|
|
1205
|
+
catch (error) {
|
|
1206
|
+
const message = error instanceof Error ? error.message : String(error);
|
|
1207
|
+
const label = file === "-" ? "stdin" : file;
|
|
1208
|
+
if (format === "json") {
|
|
1209
|
+
console.log(JSON.stringify({
|
|
1210
|
+
valid: false,
|
|
1211
|
+
diagnostics: [
|
|
1212
|
+
{ path: "<file>", expected: "a readable file", got: message },
|
|
1213
|
+
],
|
|
1214
|
+
warnings: [],
|
|
1215
|
+
note: STRUCTURAL_ONLY_NOTE,
|
|
1216
|
+
}));
|
|
1217
|
+
}
|
|
1218
|
+
else {
|
|
1219
|
+
console.error(`Could not read ${label}: ${message}`);
|
|
1220
|
+
console.error(STRUCTURAL_ONLY_NOTE);
|
|
1221
|
+
}
|
|
1222
|
+
process.exitCode = 2;
|
|
1223
|
+
return;
|
|
1224
|
+
}
|
|
1225
|
+
const result = validateReviewReport(raw);
|
|
1226
|
+
if (format === "json") {
|
|
1227
|
+
console.log(JSON.stringify({
|
|
1228
|
+
valid: result.valid,
|
|
1229
|
+
diagnostics: result.diagnostics,
|
|
1230
|
+
warnings: result.warnings,
|
|
1231
|
+
note: STRUCTURAL_ONLY_NOTE,
|
|
1232
|
+
}));
|
|
1233
|
+
process.exitCode = result.valid ? 0 : 1;
|
|
1234
|
+
return;
|
|
1235
|
+
}
|
|
1236
|
+
for (const warning of result.warnings) {
|
|
1237
|
+
console.error(`warning: ${warning}`);
|
|
1238
|
+
}
|
|
1239
|
+
// Both the exit-0 (valid) and exit-1 (invalid) cases are "verdict
|
|
1240
|
+
// paths": a validation verdict was actually computed, so its message
|
|
1241
|
+
// and the structural-only note that qualifies it belong on the same
|
|
1242
|
+
// stream. Previously the note printed on stderr unconditionally while
|
|
1243
|
+
// the valid-case verdict printed on stdout, splitting one reading
|
|
1244
|
+
// across two streams (fix-round, review finding L5); the two exit-2
|
|
1245
|
+
// "usage error" branches above keep stderr for both, since no verdict
|
|
1246
|
+
// was computed there.
|
|
1247
|
+
if (result.valid) {
|
|
1248
|
+
console.log("Structurally valid reviewer return.");
|
|
1249
|
+
}
|
|
1250
|
+
else {
|
|
1251
|
+
console.log(`Structurally invalid reviewer return (${result.diagnostics.length} issue${result.diagnostics.length === 1 ? "" : "s"}):`);
|
|
1252
|
+
for (const diagnostic of result.diagnostics) {
|
|
1253
|
+
console.log(` ${diagnostic.path}: expected ${diagnostic.expected}, got ${diagnostic.got}`);
|
|
1254
|
+
}
|
|
1255
|
+
}
|
|
1256
|
+
console.log(STRUCTURAL_ONLY_NOTE);
|
|
1257
|
+
process.exitCode = result.valid ? 0 : 1;
|
|
1258
|
+
});
|
|
1259
|
+
// Commander's default for a parsing failure (a missing `<file>` argument,
|
|
1260
|
+
// an unknown option, excess arguments) is exit code 1, the same code this
|
|
1261
|
+
// command otherwise reserves for "structurally invalid" -- collapsing
|
|
1262
|
+
// "you didn't invoke this right" into "the return you gave me is invalid"
|
|
1263
|
+
// (fix-round, review finding L1). Scoped to this one subcommand so every
|
|
1264
|
+
// other command's existing commander-parsing exit behavior is untouched:
|
|
1265
|
+
// commander already prints the error message itself before calling this
|
|
1266
|
+
// callback, so remapping the exit code is all that is needed here.
|
|
1267
|
+
validateReviewReportCommand.exitOverride((err) => {
|
|
1268
|
+
process.exit(err.exitCode === 0 ? 0 : 2);
|
|
1269
|
+
});
|
|
1175
1270
|
program.parseAsync(process.argv).catch((error) => {
|
|
1176
1271
|
console.error(error instanceof Error ? error.message : error);
|
|
1177
1272
|
process.exitCode = 1;
|