orchestrator-workflow 0.36.0 → 0.37.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,160 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.37.0] - 2026-09-18
11
+
12
+ - The two verdict fields of a mutation probe now carry a legend, a source
13
+ and a check. `assets/agents/implementer.md`, the normative site for the
14
+ output-field semantics, says: "`result: killed` means the probe's test
15
+ command reacted to the mutant under the runner's own pass predicate and
16
+ `survived` means it did not; `expectation: met` means that outcome is what
17
+ the probe was expected to show and `violated` means it is not; both are
18
+ `not_applicable` when no `result` was measured." and: "When the probe
19
+ runner states a machine-readable verdict, copy whichever of `result` and
20
+ `expectation` it states from it verbatim, never from your own reading of
21
+ the test output; when it states only `result`, set `expectation` by
22
+ comparing that verdict with the probe's declared expectation. Quote the
23
+ runner's verdict for each probe in `tests.executed`, and say there when
24
+ `expectation` was set this way, so both fields can be checked against it."
25
+ `references/contracts.md` carries both sentences verbatim for the
26
+ orchestrator. Step 6 of `references/evidence-and-probes.md` adds the
27
+ orchestrator's side: "Before transferring a probe row, compare its
28
+ `result` and `expectation` with the runner verdict quoted in
29
+ `tests.executed`; on a mismatch, or when a verdict the runner states is
30
+ not quoted, resupply it (ask the same implementer for the verdict, respawn
31
+ one when it is gone, or rerun the probe yourself in isolation), record the
32
+ resupply in `03-decisions.md`, and treat it as a transfer blocker rather
33
+ than a misfire, since the return itself parses; never infer either field.
34
+ A quoted probe verdict is not a named result of the verification set, so
35
+ the set's missing-or-extra rule does not apply to it." The contract shape,
36
+ the enums and the eleven `mutation_probes` sub-fields are unchanged, no
37
+ tool is named, and every sentence was appended to an existing line so no
38
+ cited line moved (the prompt is cited by line throughout the knowledge
39
+ bundle); the prompt change reaches an existing install with the next kit
40
+ re-install, which also re-renders the tier variants. Evidence (issue #300,
41
+ one run, not a benchmark): an implementer returned all six probes as
42
+ `survived` with `expectation: violated` while its own evidence showed the
43
+ suite failing on every mutant, and the mislabel was caught only because
44
+ the orchestrator compared the fields with that evidence by hand. Pinned in
45
+ `test/probe-plans-recovery.test.ts`.
46
+ - Three ceremony rules in the skill references, none of which changes the
47
+ review gate, the waiver rules or the AGENTS.md section (byte-identical, so
48
+ no AGENTS.md re-install is needed; the rules reach an existing install
49
+ with the next kit re-install). Step 6 of
50
+ `references/evidence-and-probes.md` now says: "Record a baseline revision
51
+ only when scope or the normative text of a criterion changes, including a
52
+ change to what its verification checks; a wording precision that leaves
53
+ the check itself unchanged is a `03-decisions.md` entry, not a revision";
54
+ the orchestrator records that entry, states in it why no evidence is
55
+ invalidated, and communicates the corrected wording in the next
56
+ delegation. Step 7 now says: "For a review round whose entire delta is a
57
+ docs-only delta in the sense of step 8's docs-only closure, default to the
58
+ `-medium` reviewer tier with `review_method: normal` where tier variants
59
+ are installed". That refines the general tier default for this one class
60
+ only, a round that touches an instruction, policy, template or prompt file
61
+ keeps the general default, and the minimum review methods are untouched;
62
+ the AGENTS.md section does not yet point to this refinement.
63
+ `references/review-and-recovery.md` gains a "Pinned-prose changes" section
64
+ for a change whose acceptance rests on tests that pin documentation
65
+ wording: "A prose mutant survives exactly when its bytes sit in no
66
+ assertion", so review rounds that hunt for the next unpinned sentence do
67
+ not converge. The section asks for one normative site per rule, a claim
68
+ list in the acceptance criterion as the pin obligation (every normative
69
+ sentence the change adds or alters at that site is a claim, an omission is
70
+ named with its reason), a reviewer briefing that bounds the prose mutant
71
+ space to that list, copies bound to the normative site by one shared test
72
+ constant, and: "Cap test-adequacy review rounds on the change at two." It
73
+ defines the capped round, exempts semantic findings, and changes neither
74
+ the Round-2 halt rule, the escalation budget, the Fix-regression decision
75
+ point nor the review gate. Step 7 points to the section without restating
76
+ it. Evidence (issue #300 and the change that added the Fix-regression
77
+ decision point; one repository each, not a benchmark): the issue reports
78
+ baseline revisions r1 to r3 for two wording precisions of a verification
79
+ method, and a run in which the top reviewer tier was about half the day's
80
+ cost across nine reviewer rounds and an advisor, where the documentation
81
+ rounds did not need that tier; the decision point itself, about 27 lines
82
+ of rule text, took four review rounds with unpinned prose reported in
83
+ every one, until the trigger was held in one test constant and the mutant
84
+ space was bounded. Pinned in `test/probe-plans-recovery.test.ts`.
85
+ - `references/review-and-recovery.md` gains a "Fix-regression decision
86
+ point" between the Review-round escalation budget and the Final
87
+ acceptance rule: when the review of a fix round reports at least one
88
+ `high` or `critical` finding that the previous round's review did not
89
+ report, with `introduced_by_delta: yes`, the orchestrator names in one
90
+ sentence why the fix could introduce it and records one of four outcomes
91
+ in `03-decisions.md` (continue with the stated reason, redesign, split,
92
+ hold) before another fix round starts. It is a decision point, not a
93
+ halt, and changes neither the Round-2 halt rule nor the budget. The
94
+ rule is defined only in that section; step 8 of
95
+ `references/evidence-and-probes.md` points to it without restating the
96
+ trigger. The AGENTS.md section is unchanged, so no AGENTS.md re-install
97
+ is needed; the rule reaches an existing install with the next kit
98
+ re-install, like any other skill-reference change. Evidence (issue #300,
99
+ one repository, one model mix, not a benchmark): in a five-round run the
100
+ third review already showed that the fix had introduced new high
101
+ findings, but the defect class only recurred in the fourth review, so
102
+ the halt rule fired one round, about an hour, after the structural
103
+ problem was visible. Pinned in `test/probe-plans-recovery.test.ts`.
104
+ - `validate-review-report` now checks that every element of a string-array
105
+ field (`summary`, `missing_tests`, `residual_risks`) is a string, one
106
+ diagnostic per offending element at `<field>[<index>]`; and
107
+ `extractYamlSource` now prefers a fenced block tagged `yaml`/`yml` when
108
+ several fences are present, falling back to the first fence only when
109
+ none carries that tag, with a warning naming the earlier fence it
110
+ skipped. The fence tag is now matched against the whole info string's
111
+ first whitespace-delimited word rather than a leading run of letters,
112
+ so a tag followed by attributes (a fence opened `yaml title=x`) is
113
+ recognized as `yaml` instead of matching no fence at all. The "prose
114
+ found before" warning no longer also fires when the text preceding the
115
+ preferred fence is exactly the skipped fence(s) plus whitespace, so a
116
+ skipped fence is no longer double-reported as both skipped and prose.
117
+ The opening fence's backtick run is now consumed whole and the closing
118
+ fence must repeat at least as many backticks at column 0 with nothing
119
+ but whitespace after it, so a return opened with four backticks is
120
+ recognized as `yaml` (the run's leftover backticks previously landed in
121
+ the info string, making the tag itself start with a backtick) and ends
122
+ at its own four-backtick run rather than at a three-backtick line
123
+ inside it. Widening the info-string capture also widens what counts as
124
+ a fence at all: any attribute-bearing fence, and any fence opened with
125
+ more than three backticks, is now a fence, so input whose only fence
126
+ was opened `js title=x` is treated as fenced rather than as literal
127
+ YAML. Only whitespace-separated attributes count toward the tag:
128
+ `yaml title=x` counts, `yaml,title=x` does not, its first word being
129
+ the whole string. That opening run is now matched whole and never
130
+ re-entered at a shorter length, which bounds the scan: a run with no
131
+ valid closer was previously retried at every shorter run length from
132
+ every offset inside the run, each retry rescanning the block body, so
133
+ a 2000-backtick run in a 22 KB return took roughly 50 seconds through
134
+ `extractYamlSource` where it now takes under a millisecond. Keeping
135
+ the run whole also makes the closing-length rule literal: a return
136
+ whose opening run is longer than every closing run in it is no fence
137
+ at all and reaches the parser whole, where splitting the run
138
+ previously matched it and read its leftover backticks as the start of
139
+ the tag.
140
+
141
+ - `check-release-changelogs`'s `VERSION_HEADING_RE` (rule 1, version-heading)
142
+ now accepts the same prerelease identifier class as `parseSemver`
143
+ (`[0-9A-Za-z.-]`, including the hyphen), sourced from one shared
144
+ `PRERELEASE_IDENTIFIER_CHARS` constant so the two cannot drift apart
145
+ again: a package.json version with a hyphenated prerelease tag such as
146
+ `1.0.0-alpha-1` parsed fine through `parseSemver` but previously never
147
+ matched the heading regex, which only accepted `[\w.]`. The previous
148
+ 0.35.0 hardening bullet named the new rules without naming three
149
+ details of their own contract: the `--expect <csv>` option overrides
150
+ rule 5's (checked-package-scope) expectation list, and an empty csv
151
+ (`--expect ""`) opts out of rule 5 entirely rather than passing it
152
+ vacuously; rule 5 itself only runs when that expectation list is
153
+ non-empty, so a fixture whose packages do not share this repo's names
154
+ can still run the other four rules without a spurious finding; and
155
+ rule 2's (fresh-unreleased) direction check compares versions by
156
+ semver precedence, not string inequality, with build metadata stripped
157
+ before the comparison since two versions differing only in build
158
+ metadata carry equal precedence per the semver spec. The shared class
159
+ is narrower than `\w`: an underscore prerelease heading such as
160
+ `1.0.0-alpha_1` is no longer accepted either, since semver's own
161
+ prerelease grammar forbids underscores and `parseSemver` already
162
+ rejected such a version before this change.
163
+
10
164
  ## [0.36.0] - 2026-09-16
11
165
 
12
166
  - A `validate-review-report <file>` CLI subcommand (`-` reads stdin) checks a
package/README.md CHANGED
@@ -705,17 +705,35 @@ required fields and enums (see the "Reviewer output contract" section of
705
705
  `assets/skill/references/contracts.md`, byte-identical to the contract in
706
706
  `assets/agents/reviewer.md`), whether the return is fenced in a code
707
707
  block (any language tag, or none) or given unfenced, and prints one
708
- diagnostic per missing or invalid field. A fenced return ends at the
709
- first closing fence that starts at column 0, so a reviewer quoting a
710
- fenced snippet inside a value (a `description` block scalar, which YAML
711
- indents) does not truncate the return. `--format json` prints the same
712
- diagnostics as a single JSON object instead of human-readable text. It
713
- exits `0` when the return is structurally valid, `1` when it is
714
- structurally invalid (a required field is missing or its value falls
715
- outside its enum, or the input is unparsable, empty, or not a mapping),
716
- and `2` for a usage error (an unreadable file, an unrecognized `--format`
717
- value, a missing `<file>` argument, an unknown option, or an excess
718
- positional argument).
708
+ diagnostic per missing or invalid field. Every element of a string-array
709
+ field (`summary`, `missing_tests`, `residual_risks`) must itself be a
710
+ string; a non-string element (a number, a mapping, a boolean, or `null`
711
+ -- written as a bare or `~` bullet) is its own diagnostic at
712
+ `<field>[<index>]`. A fenced return ends at the first closing fence that
713
+ starts at column 0, repeats at least as many backticks as the opening
714
+ fence, and carries nothing but whitespace after that run, so neither a
715
+ reviewer quoting a fenced snippet inside a value (a `description` block
716
+ scalar, which YAML indents) nor one wrapping a return in four backticks
717
+ around a snippet fenced at column 0 truncates the return. A return
718
+ with no closing fence satisfying all three is not fenced at all, so its
719
+ whole text reaches the parser; that includes one whose opener is longer
720
+ than every closing run present. When the return carries more than one
721
+ fenced block, the first one whose fence tag's first word is `yaml` or
722
+ `yml` is validated, case-insensitively and counting whitespace-separated
723
+ attributes (`yaml title=x` counts; `yaml,title=x` does not, its first
724
+ word being the whole string), falling back to the first fence only when
725
+ none carries that word; a warning names any earlier fence skipped this
726
+ way. This preference can validate a later worked example instead of an
727
+ earlier, real but unfenced return: a reviewer who leaves their own return
728
+ unfenced and then quotes a `yaml`-tagged example afterward has that
729
+ example validated instead, which the emitted warning also names.
730
+ `--format json` prints the same diagnostics as a single JSON object
731
+ instead of human-readable text. It exits `0` when the return is
732
+ structurally valid, `1` when it is structurally invalid (a required field
733
+ is missing or its value falls outside its enum, or the input is
734
+ unparsable, empty, or not a mapping), and `2` for a usage error (an
735
+ unreadable file, an unrecognized `--format` value, a missing `<file>`
736
+ argument, an unknown option, or an excess positional argument).
719
737
  `--format json` governs the validation verdict only: a commander parsing
720
738
  error (missing argument, unknown option, excess arguments) or an
721
739
  unrecognized `--format` value itself still prints plain text to stderr
@@ -103,7 +103,7 @@ Rules:
103
103
  from its result fields (`verified_applied_via`, `result`, `expectation`,
104
104
  `reason`, `restored_verified`), take the definition fields from that
105
105
  mutant record so the copied report still carries all eleven
106
- `mutation_probes` sub-fields.
106
+ `mutation_probes` sub-fields. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
107
107
  - Run every long test, build, or mutation-probe command in the foreground
108
108
  and wait for it to finish before returning. When one foreground call
109
109
  cannot hold it to completion, poll the backgrounded run to completion
@@ -126,7 +126,7 @@ commits:
126
126
  Follow [evidence-and-probes.md workflow step 6](evidence-and-probes.md#workflow)
127
127
  for implementation evidence, verification, mutation probes, and replay. For
128
128
  output-field semantics and commit reporting, follow the installed
129
- implementer role prompt. Return the selected contract's YAML envelope.
129
+ implementer role prompt. Return the selected contract's YAML envelope. `result: killed` means the probe's test command reacted to the mutant under the runner's own pass predicate and `survived` means it did not; `expectation: met` means that outcome is what the probe was expected to show and `violated` means it is not; both are `not_applicable` when no `result` was measured. When the probe runner states a machine-readable verdict, copy whichever of `result` and `expectation` it states from it verbatim, never from your own reading of the test output; when it states only `result`, set `expectation` by comparing that verdict with the probe's declared expectation. Quote the runner's verdict for each probe in `tests.executed`, and say there when `expectation` was set this way, so both fields can be checked against it.
130
130
 
131
131
  ## Reviewer output contract
132
132
 
@@ -166,22 +166,41 @@ When it is missing, the orchestrator asks the reviewer to resupply it
166
166
  instead of inferring one from the findings list.
167
167
 
168
168
  A structural check for this exact contract ships as a CLI subcommand:
169
- `orchestrator-workflow validate-review-report <file>` (pass `-` to read the
170
- return from stdin instead) parses the reviewer return's YAML, fenced in a
171
- code block with any language tag or none, or unfenced, and checks the
172
- required fields and enums above one by one, printing a diagnostic (path,
173
- expected, got) for each missing or invalid field; add `--format json` for
174
- the same diagnostics as a single JSON object. It exits `0` when the return
175
- is structurally valid, `1` when it is structurally invalid (a missing or
176
- out-of-enum required field, or unparsable, empty, non-mapping input), and
177
- `2` for a usage error (an unreadable file, an unrecognized `--format`
178
- value, or an argument-parsing error: a missing `<file>` argument, an
179
- unknown option, an excess positional argument). `--format json` governs
180
- the validation verdict only: an argument-parsing error or an unrecognized
181
- `--format` value still prints plain text to stderr, except an unreadable
182
- file, which still emits the JSON envelope on stdout. The check is
183
- structural only: it never judges semantic adequacy, cannot waive a
184
- finding, and passing it is never orchestrator acceptance.
169
+ `orchestrator-workflow validate-review-report <file>` (pass `-` to read
170
+ the return from stdin instead) parses the reviewer return's YAML, fenced
171
+ in a code block with any language tag or none, or unfenced, and checks
172
+ the required fields and enums above one by one, printing a diagnostic
173
+ (path, expected, got) for each missing or invalid field; every element of
174
+ a string-array field (`summary`, `missing_tests`, `residual_risks`) must
175
+ itself be a string, with a diagnostic at `<field>[<index>]` for each one
176
+ that is not (a number, a mapping, a boolean, and `null` -- a bare or `~`
177
+ bullet -- are all rejected the same way). A fenced return ends at the
178
+ first closing fence that starts at column 0, repeats at least as many
179
+ backticks as the opening one, and carries nothing but whitespace after
180
+ that run, so a return wrapped in four backticks may quote a snippet
181
+ fenced in three without truncating itself; a return with no closing
182
+ fence satisfying all three is not fenced at all and reaches the parser
183
+ whole, including one whose opener is longer than every closing run
184
+ present. When more than one fenced block
185
+ is present, the first one whose fence tag's first word is `yaml` or `yml`
186
+ is validated, case-insensitively and counting whitespace-separated
187
+ attributes (`yaml title=x` counts; `yaml,title=x` does not, its first
188
+ word being the whole string), falling back to the first fence only when
189
+ none carries that word, with a warning naming any earlier fence skipped
190
+ this way. The same preference can validate a later worked example instead
191
+ of an earlier, real but unfenced return, which the emitted warning also
192
+ names. Add `--format json` for the same diagnostics as a single JSON
193
+ object. It exits `0` when the return is structurally valid, `1` when it
194
+ is structurally invalid (a missing or out-of-enum required field, or
195
+ unparsable, empty, non-mapping input), and `2` for a usage error (an
196
+ unreadable file, an unrecognized `--format` value, or an argument-parsing
197
+ error: a missing `<file>` argument, an unknown option, an excess
198
+ positional argument). `--format json` governs the validation verdict
199
+ only: an argument-parsing error or an unrecognized `--format` value still
200
+ prints plain text to stderr, except an unreadable file, which still emits
201
+ the JSON envelope on stdout. The check is structural only: it never
202
+ judges semantic adequacy, cannot waive a finding, and passing it is never
203
+ orchestrator acceptance.
185
204
 
186
205
  `recurrence` classifies each finding against earlier rounds on the same
187
206
  task: `new` for a defect class not previously found here, `repeated` for
@@ -112,7 +112,7 @@ directory and the subagents.
112
112
  `03-decisions.md` and consolidate evidence in
113
113
  `04-implementation-summary.md`, recording each probe the implementer
114
114
  reports as a row in `04-implementation-summary.md`'s Mutation Probes
115
- subsection, with the round it was named in. Each row's Before/After
115
+ subsection, with the round it was named in. Before transferring a probe row, compare its `result` and `expectation` with the runner verdict quoted in `tests.executed`; on a mismatch, or when a verdict the runner states is not quoted, resupply it (ask the same implementer for the verdict, respawn one when it is gone, or rerun the probe yourself in isolation), record the resupply in `03-decisions.md`, and treat it as a transfer blocker rather than a misfire, since the return itself parses; never infer either field. A quoted probe verdict is not a named result of the verification set, so the set's missing-or-extra rule does not apply to it. Each row's Before/After
116
116
  cells hold a single-line excerpt; when the mutant's actual before/after
117
117
  text is multi-line or contains an unescaped `|`, or the mutant is a
118
118
  patch/diff rather than a text swap, the full text or diff goes in the
@@ -137,7 +137,7 @@ directory and the subagents.
137
137
  acceptance; the coverage index is not a results database or acceptance
138
138
  engine. Only the orchestrator can explicitly revise a baseline, recording
139
139
  old/new revisions, affected IDs, authority and reason, invalidated evidence,
140
- and verified rationale for carrying unchanged evidence forward.
140
+ and verified rationale for carrying unchanged evidence forward. Record a baseline revision only when scope or the normative text of a criterion changes, including a change to what its verification checks; a wording precision that leaves the check itself unchanged is a `03-decisions.md` entry, not a revision: the orchestrator records it, states in that entry why no evidence is invalidated, and communicates the corrected wording in the next delegation.
141
141
  7. **Delegate review.** Send the diff to the reviewer subagent, naming in the
142
142
  briefing the base and head revision the diff was generated from. When tier
143
143
  variants are installed, pick the reviewer tier (the installed
@@ -152,7 +152,7 @@ directory and the subagents.
152
152
  cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
153
153
  never substitutes for it: do not pair `adversarial` with the `-medium`
154
154
  reviewer tier, a budget mismatch that names probes without the effort to run
155
- them; tiers themselves are unchanged by this axis. When the reviewer's
155
+ them; tiers themselves are unchanged by this axis. For a review round whose entire delta is a docs-only delta in the sense of step 8's docs-only closure, default to the `-medium` reviewer tier with `review_method: normal` where tier variants are installed. This refines the general tier default above for that one class only: there `-medium` is the default and a higher tier is the non-default choice recorded with a one-line reason. A review round that touches an instruction, policy, template or prompt file keeps the general default, whatever the file type, and the minimums named above are unaffected. For a change whose acceptance rests on tests that pin documentation wording, write the briefing as the Pinned-prose changes section of [review and recovery](review-and-recovery.md) requires. When the reviewer's
156
156
  environment cannot use version control to see the diff (for example a
157
157
  policy-gated repository), supply the diff as a pre-generated file in the
158
158
  briefing instead of expecting the reviewer to derive it, and have the
@@ -226,7 +226,7 @@ directory and the subagents.
226
226
  halt signal across repeated review-fix cycles (see Round-2 halt rule
227
227
  below). By the second round-2 halt signal or the third `fix_required`
228
228
  review round on the same task, apply the Review-round escalation budget
229
- (see below) instead of running another round unaided. At an advisor
229
+ (see below) instead of running another round unaided. When a fix round's review meets the trigger of the Fix-regression decision point (defined only in [review and recovery](review-and-recovery.md), not restated here), record the Fix-regression decision point before another fix round starts. At an advisor
230
230
  trigger (architectural uncertainty, conflicting
231
231
  requirements, a high-commitment fork among valid options, repeated
232
232
  implementation failures, a review deadlock, a high-risk decision), the
@@ -82,6 +82,72 @@ through the reviewer subagent in full; this budget forces a change in
82
82
  approach, not a shortcut past the review gate. Anchored by a measurement;
83
83
  see the entry for this rule in the orchestrator-workflow CHANGELOG.
84
84
 
85
+ ## Fix-regression decision point
86
+
87
+ The signal: the review of a fix round (any implementation round after the
88
+ task's first) reports at least one `high` or `critical` finding that the
89
+ previous round's review did not report, with `introduced_by_delta: yes`, so
90
+ the fix itself broke something. Read the qualifier off the findings of the
91
+ two reviews, not off `recurrence`: a `recurrence: repeated` finding that
92
+ the previous round's review did not report still triggers it. `unknown` and
93
+ `no` do not trigger this decision point: `no` continues through the
94
+ ordinary finding gate, and `unknown` keeps its existing treatment under the
95
+ Round-2 halt rule and the escalation budget. The signal needs no
96
+ recurrence: it fires even when the new finding's defect class has not
97
+ appeared on this task before, which is what separates it from the Round-2
98
+ halt rule above.
99
+
100
+ Before another fix round starts, name in one sentence why the fix could
101
+ introduce the defect (the structural cause, or the statement that there is
102
+ none), and record one of four outcomes as a decision in `03-decisions.md`:
103
+ continue with the stated reason, redesign, split, or hold (a merge-hold to
104
+ the operator). Spawning the advisor for this decision is optional.
105
+
106
+ This is a decision point, not a halt: continuing is a valid outcome, it is
107
+ not a round-2 halt signal, and it does not count toward the Review-round
108
+ escalation budget (the negative round itself still counts there as
109
+ before). It never replaces a review round. When the same review also fires
110
+ the Round-2 halt signal, the halt rule governs and this record is folded
111
+ into its split-or-redesign decision. Anchored by an observed run; see the
112
+ entry for this rule in the orchestrator-workflow CHANGELOG.
113
+
114
+ ## Pinned-prose changes
115
+
116
+ This applies to a change whose acceptance rests on tests that pin
117
+ documentation wording (a rule text asserted by string match). A prose
118
+ mutant survives exactly when its bytes sit in no assertion, so a surviving
119
+ mutant alone says nothing about quality, and review rounds that hunt for
120
+ the next unpinned sentence do not converge. For such a change:
121
+
122
+ - Name one normative site per rule when slicing; every other site that
123
+ states the rule is a copy.
124
+ - List the load-bearing claims of the normative site in the acceptance
125
+ criterion, and pin each one as the whole sentence or clause that carries
126
+ it. That list is the pin obligation. Every normative sentence the change
127
+ adds or alters at that site is a claim; one left off the list is named
128
+ in the criterion with the reason it is not load-bearing.
129
+ - Bound the reviewer's prose mutant space to that list in the briefing. A
130
+ survivor outside the list is a scope note in the reviewer's
131
+ `residual_risks`, not a finding, unless the reviewer shows that the
132
+ unlisted sentence is load-bearing.
133
+ - Bind each copy to the normative site through one shared test constant,
134
+ and let a pointer point without restating the rule.
135
+ - Cap test-adequacy review rounds on the change at two. A test-adequacy
136
+ review round is one whose only unresolved findings are `tests` findings
137
+ about pin gaps on the pinned prose; a round with any other unresolved
138
+ finding is an ordinary round outside the cap. Pin gaps that remain
139
+ become accepted notes or a follow-up.
140
+
141
+ Semantic findings are exempt from the bound and from the cap: two sites
142
+ stating different rules, a contradiction with another rule, and a false
143
+ claim are defects at whatever severity they deserve. The cap changes
144
+ neither the Round-2 halt rule, the Review-round escalation budget nor the
145
+ Fix-regression decision point: a capped round still counts as a negative
146
+ round where it is one. The review gate is unchanged: a high or critical
147
+ finding of any category still blocks and is never capped away, and
148
+ accepting one follows the waiver rules. Anchored by an observed run; see
149
+ the entry for this rule in the orchestrator-workflow CHANGELOG.
150
+
85
151
  ## Final acceptance rule
86
152
 
87
153
  Subagents provide evidence. The orchestrator decides. The operator receives
@@ -59,7 +59,9 @@ export type SchemaFieldName = (typeof TOP_LEVEL_FIELDS)[number] | (typeof FINDIN
59
59
  * - `string`: present and a string; the empty string is accepted.
60
60
  * - `non-empty-string`: present, a string, and not blank.
61
61
  * - `scalar`: present and either a string or a number.
62
- * - `array`: present and an array; an empty array is accepted.
62
+ * - `array`: present and an array; an empty array is accepted, and every
63
+ * element must be a string (each non-string element is its own
64
+ * diagnostic at `<field>[<index>]`; see {@link checkArrayField}).
63
65
  * - `mapping-list`: `array`, and every element a mapping.
64
66
  * - `mapping`: present and a mapping.
65
67
  *
@@ -118,6 +120,8 @@ export declare const FIELD_KINDS: {
118
120
  export declare function expectedTextFor(field: SchemaFieldName): string;
119
121
  /** The `expected` text a diagnostic about one element of a `mapping-list` carries. */
120
122
  export declare const MAPPING_LIST_ELEMENT_EXPECTED: string;
123
+ /** The `expected` text a diagnostic about one non-string element of a plain `array`-kind field (`summary`, `missing_tests`, `residual_risks`) carries. */
124
+ export declare const ARRAY_ELEMENT_EXPECTED: string;
121
125
  interface ExtractedYaml {
122
126
  yamlText: string;
123
127
  warnings: string[];
@@ -138,13 +142,66 @@ interface ExtractedYaml {
138
142
  * before the opening fence or after the closing fence is tolerated, but
139
143
  * each is named as its own warning rather than silently dropped.
140
144
  *
141
- * The closing fence must start at column 0: the pattern anchors it with
142
- * `^` under the `m` flag, so a triple-backtick sequence inside a value
143
- * (a reviewer quoting a fenced snippet in a `description` block scalar,
144
- * which YAML necessarily indents) can no longer close the block early
145
- * and hand the parser a truncated document, which surfaced as
146
- * diagnostics about fields the return actually carried (fix-round,
147
- * review finding L3).
145
+ * The opening fence's whole backtick run is captured, and the closing
146
+ * fence must be a run at least as long, starting at column 0, with
147
+ * nothing but whitespace after it (CommonMark's own rule): the pattern
148
+ * backreferences the captured run and anchors it with `^` under the `m`
149
+ * flag. So a triple-backtick sequence inside a value (a reviewer quoting
150
+ * a fenced snippet in a `description` block scalar, which YAML
151
+ * necessarily indents) can no longer close the block early and hand the
152
+ * parser a truncated document, which surfaced as diagnostics about
153
+ * fields the return actually carried (fix-round, review finding L3);
154
+ * and a return a reviewer wrapped in four backticks precisely because
155
+ * it contains a fence of its own is closed by its own four-backtick run
156
+ * rather than by that inner one. Matching a fixed three backticks
157
+ * instead of the run left a longer opener's remaining backticks in the
158
+ * info string, which read as the tag `` `yaml `` and matched no
159
+ * yaml/yml fence at all. The OPENING fence keeps its own position
160
+ * discipline unchanged: it is located anywhere in the input rather than
161
+ * anchored to a line start.
162
+ *
163
+ * Lookarounds on both sides of the run keep it whole, so the opening
164
+ * run is never re-entered at a shorter length. Without them, an input
165
+ * whose long backtick run has no valid closer is retried at every
166
+ * shorter run length from every offset inside the run, each retry
167
+ * rescanning the lazy body: work quadratic in the run's length, which a
168
+ * single pasted return of a few hundred backticks already turns into
169
+ * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
170
+ * run whole also makes the "closing run at least as long as the opening
171
+ * one" rule above literal: an opener longer than any closing run in the
172
+ * input is no fence at all, where splitting the run instead matched it
173
+ * and pushed the leftover backticks into the info string. One cost
174
+ * stays: every backtick run is still tried as an opener candidate, and a
175
+ * candidate with no qualifying closer scans to the end of the input, so
176
+ * an input of many runs with no valid closer costs work quadratic in the
177
+ * number of runs (well under a second at several thousand runs; a
178
+ * well-formed return is unaffected). Anchoring the opener to a line
179
+ * start would remove it, at the price of the position-free opener the
180
+ * paragraph above keeps.
181
+ *
182
+ * When the input carries more than one fenced block (a reviewer pasting
183
+ * a worked example ahead of the real return, say), the FIRST fence whose
184
+ * info string's first whitespace-delimited word is `yaml` or `yml`
185
+ * (case-insensitive) is preferred over every earlier fence, tagged or
186
+ * not; only when none of the fences carries that word does today's
187
+ * original first-fence behaviour apply. The whole info string is
188
+ * captured, not only a leading run of letters, so a tag followed by
189
+ * attributes (` ```yaml title=x `) is still recognised as `yaml` --
190
+ * previously the capture stopped at the first non-letter and required a
191
+ * newline right after it, so an attribute-bearing info string matched no
192
+ * fence at all, tagged or not; this also widens the untagged first-fence
193
+ * fallback, so an attribute-bearing fence with no yaml/yml word is at
194
+ * least recognised as a fence. This preference rule is not free of
195
+ * surprises of its own: a reviewer whose own return is left unfenced and
196
+ * who then quotes a ```yaml example afterward has that later example
197
+ * validated instead of their real return, which the emitted warning
198
+ * names. Preferring a later, differently-positioned fence over
199
+ * `fences[0]` is itself named as a warning, distinct from the existing
200
+ * before/after prose warnings; the before-warning is suppressed when the
201
+ * text preceding the chosen fence consists only of the skipped fence(s)
202
+ * and whitespace, since calling a legitimate (if unpreferred) fenced
203
+ * block "prose" alongside the skip warning that already names it is
204
+ * redundant; real prose ahead of a skipped fence still warns as before.
148
205
  */
149
206
  export declare function extractYamlSource(raw: string): ExtractedYaml;
150
207
  /**
@@ -145,6 +145,29 @@ function checkScalarField(record, key, path, diagnostics) {
145
145
  });
146
146
  }
147
147
  }
148
+ /**
149
+ * Checks a top-level `array`-kind field (`summary`, `missing_tests`,
150
+ * `residual_risks`): the contract writes each of these as a plain list of
151
+ * strings (`- ""`), so every element that is not a string is its own
152
+ * diagnostic at `<key>[<index>]`, `expected: "string"`, alongside the
153
+ * container-level checks. All offending elements are reported, not only
154
+ * the first, the same way {@link checkFindings} reports every non-mapping
155
+ * `findings[]` entry rather than stopping at one.
156
+ *
157
+ * Emptiness is deliberately not judged here: an element that is an
158
+ * explicitly quoted empty or blank string (`- ""`, `- " "`) is still a
159
+ * string and passes, the same tolerance {@link checkStringField}'s plain
160
+ * `"string"` kind gives a top-level field (unlike
161
+ * {@link checkNonEmptyStringField}'s `"non-empty-string"` kind, which
162
+ * `task_id` uses). A bare `- ` or `- ~` bullet is not a string at all --
163
+ * YAML parses either as `null`, which this checker rejects the same as
164
+ * any other non-string element; only a quoted placeholder passes. These
165
+ * three fields are declared `"array"` in {@link FIELD_KINDS}, not
166
+ * `"non-empty-string"`, so a quoted element gets the same tolerance its
167
+ * own kind implies; a reviewer emitting a quoted placeholder blank bullet
168
+ * is a content question the orchestrator judges, not a structural one
169
+ * this validator judges.
170
+ */
148
171
  function checkArrayField(doc, key, diagnostics) {
149
172
  const value = doc[key];
150
173
  if (value === undefined) {
@@ -157,7 +180,17 @@ function checkArrayField(doc, key, diagnostics) {
157
180
  expected: "array",
158
181
  got: describeValue(value),
159
182
  });
183
+ return;
160
184
  }
185
+ value.forEach((element, index) => {
186
+ if (typeof element !== "string") {
187
+ diagnostics.push({
188
+ path: `${key}[${index}]`,
189
+ expected: "string",
190
+ got: describeValue(element),
191
+ });
192
+ }
193
+ });
161
194
  }
162
195
  /**
163
196
  * Dispatch table keyed by every name in {@link FINDING_FIELDS}. The
@@ -368,6 +401,8 @@ export function expectedTextFor(field) {
368
401
  }
369
402
  /** The `expected` text a diagnostic about one element of a `mapping-list` carries. */
370
403
  export const MAPPING_LIST_ELEMENT_EXPECTED = KIND_EXPECTED.mapping;
404
+ /** The `expected` text a diagnostic about one non-string element of a plain `array`-kind field (`summary`, `missing_tests`, `residual_risks`) carries. */
405
+ export const ARRAY_ELEMENT_EXPECTED = KIND_EXPECTED.string;
371
406
  /**
372
407
  * A reviewer return is commonly wrapped in a single fenced code block.
373
408
  * Strips one leading/trailing fence when present, whatever language tag
@@ -384,28 +419,98 @@ export const MAPPING_LIST_ELEMENT_EXPECTED = KIND_EXPECTED.mapping;
384
419
  * before the opening fence or after the closing fence is tolerated, but
385
420
  * each is named as its own warning rather than silently dropped.
386
421
  *
387
- * The closing fence must start at column 0: the pattern anchors it with
388
- * `^` under the `m` flag, so a triple-backtick sequence inside a value
389
- * (a reviewer quoting a fenced snippet in a `description` block scalar,
390
- * which YAML necessarily indents) can no longer close the block early
391
- * and hand the parser a truncated document, which surfaced as
392
- * diagnostics about fields the return actually carried (fix-round,
393
- * review finding L3).
422
+ * The opening fence's whole backtick run is captured, and the closing
423
+ * fence must be a run at least as long, starting at column 0, with
424
+ * nothing but whitespace after it (CommonMark's own rule): the pattern
425
+ * backreferences the captured run and anchors it with `^` under the `m`
426
+ * flag. So a triple-backtick sequence inside a value (a reviewer quoting
427
+ * a fenced snippet in a `description` block scalar, which YAML
428
+ * necessarily indents) can no longer close the block early and hand the
429
+ * parser a truncated document, which surfaced as diagnostics about
430
+ * fields the return actually carried (fix-round, review finding L3);
431
+ * and a return a reviewer wrapped in four backticks precisely because
432
+ * it contains a fence of its own is closed by its own four-backtick run
433
+ * rather than by that inner one. Matching a fixed three backticks
434
+ * instead of the run left a longer opener's remaining backticks in the
435
+ * info string, which read as the tag `` `yaml `` and matched no
436
+ * yaml/yml fence at all. The OPENING fence keeps its own position
437
+ * discipline unchanged: it is located anywhere in the input rather than
438
+ * anchored to a line start.
439
+ *
440
+ * Lookarounds on both sides of the run keep it whole, so the opening
441
+ * run is never re-entered at a shorter length. Without them, an input
442
+ * whose long backtick run has no valid closer is retried at every
443
+ * shorter run length from every offset inside the run, each retry
444
+ * rescanning the lazy body: work quadratic in the run's length, which a
445
+ * single pasted return of a few hundred backticks already turns into
446
+ * seconds (CHANGELOG [Unreleased] names the measurement). Keeping the
447
+ * run whole also makes the "closing run at least as long as the opening
448
+ * one" rule above literal: an opener longer than any closing run in the
449
+ * input is no fence at all, where splitting the run instead matched it
450
+ * and pushed the leftover backticks into the info string. One cost
451
+ * stays: every backtick run is still tried as an opener candidate, and a
452
+ * candidate with no qualifying closer scans to the end of the input, so
453
+ * an input of many runs with no valid closer costs work quadratic in the
454
+ * number of runs (well under a second at several thousand runs; a
455
+ * well-formed return is unaffected). Anchoring the opener to a line
456
+ * start would remove it, at the price of the position-free opener the
457
+ * paragraph above keeps.
458
+ *
459
+ * When the input carries more than one fenced block (a reviewer pasting
460
+ * a worked example ahead of the real return, say), the FIRST fence whose
461
+ * info string's first whitespace-delimited word is `yaml` or `yml`
462
+ * (case-insensitive) is preferred over every earlier fence, tagged or
463
+ * not; only when none of the fences carries that word does today's
464
+ * original first-fence behaviour apply. The whole info string is
465
+ * captured, not only a leading run of letters, so a tag followed by
466
+ * attributes (` ```yaml title=x `) is still recognised as `yaml` --
467
+ * previously the capture stopped at the first non-letter and required a
468
+ * newline right after it, so an attribute-bearing info string matched no
469
+ * fence at all, tagged or not; this also widens the untagged first-fence
470
+ * fallback, so an attribute-bearing fence with no yaml/yml word is at
471
+ * least recognised as a fence. This preference rule is not free of
472
+ * surprises of its own: a reviewer whose own return is left unfenced and
473
+ * who then quotes a ```yaml example afterward has that later example
474
+ * validated instead of their real return, which the emitted warning
475
+ * names. Preferring a later, differently-positioned fence over
476
+ * `fences[0]` is itself named as a warning, distinct from the existing
477
+ * before/after prose warnings; the before-warning is suppressed when the
478
+ * text preceding the chosen fence consists only of the skipped fence(s)
479
+ * and whitespace, since calling a legitimate (if unpreferred) fenced
480
+ * block "prose" alongside the skip warning that already names it is
481
+ * redundant; real prose ahead of a skipped fence still warns as before.
394
482
  */
395
483
  export function extractYamlSource(raw) {
396
484
  const warnings = [];
397
485
  // No BOM handling: the yaml parser accepts a leading U+FEFF and the
398
486
  // fenced path trims it away with the surrounding prose.
399
487
  const withoutBom = raw;
400
- const fenceMatch = withoutBom.match(/```[A-Za-z]*\r?\n([\s\S]*?)\r?\n?^```/m);
401
- if (fenceMatch) {
402
- const start = fenceMatch.index ?? 0;
488
+ const fences = [
489
+ ...withoutBom.matchAll(/(?<!`)(`{3,})(?!`)([^\r\n]*)\r?\n([\s\S]*?)\r?\n?^\1`*[ \t]*$/gm),
490
+ ];
491
+ if (fences.length > 0) {
492
+ const fenceTag = (info) => info.trim().split(/\s+/, 1)[0] ?? "";
493
+ const yamlTaggedIndex = fences.findIndex((match) => /^(?:yaml|yml)$/i.test(fenceTag(match[2])));
494
+ const chosenIndex = yamlTaggedIndex >= 0 ? yamlTaggedIndex : 0;
495
+ const chosen = fences[chosenIndex];
496
+ if (chosenIndex > 0) {
497
+ const skippedCount = chosenIndex;
498
+ const noun = skippedCount === 1 ? "block" : "blocks";
499
+ const verb = skippedCount === 1 ? "was" : "were";
500
+ warnings.push(`${skippedCount} earlier fenced ${noun} without a yaml/yml tag ${verb} skipped in favor of the later \`${fenceTag(chosen[2])}\` fenced block; only that later block was validated`);
501
+ }
502
+ const start = chosen.index ?? 0;
403
503
  const before = withoutBom.slice(0, start);
404
- if (before.trim().length > 0) {
504
+ const beforeIsOnlySkippedFences = chosenIndex > 0 &&
505
+ fences
506
+ .slice(0, chosenIndex)
507
+ .reduce((text, skipped) => text.replace(skipped[0], ""), before)
508
+ .trim().length === 0;
509
+ if (before.trim().length > 0 && !beforeIsOnlySkippedFences) {
405
510
  warnings.push("prose found before the opening ```yaml fence; only the fenced block was validated");
406
511
  }
407
- const inner = fenceMatch[1];
408
- const after = withoutBom.slice(start + fenceMatch[0].length);
512
+ const inner = chosen[3];
513
+ const after = withoutBom.slice(start + chosen[0].length);
409
514
  if (after.trim().length > 0) {
410
515
  warnings.push("prose found after the closing ```yaml fence; only the fenced block was validated");
411
516
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.36.0",
3
+ "version": "0.37.0",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",