orchestrator-workflow 0.33.0 → 0.35.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +68 -0
- package/INSTALL-AGENT.md +17 -4
- package/LICENSE +21 -0
- package/README.md +64 -3
- package/assets/agents/implementer.md +23 -0
- package/assets/agents/reviewer.md +23 -2
- package/assets/agents/task-slicer.md +8 -0
- package/assets/agents-md-section.md +3 -3
- package/assets/skill/SKILL.md +73 -805
- package/assets/skill/references/contracts.md +266 -0
- package/assets/skill/references/evidence-and-probes.md +314 -0
- package/assets/skill/references/review-and-recovery.md +97 -0
- package/assets/skill/references/run-state-and-harness.md +171 -0
- package/assets/templates/02-tasks.md +7 -0
- package/assets/templates/03-decisions.md +1 -1
- package/assets/templates/04-implementation-summary.md +29 -0
- package/assets/templates/05-review-findings.md +6 -6
- package/dist/assets.d.ts +9 -0
- package/dist/assets.js +19 -0
- package/dist/init.js +60 -4
- package/dist/uninstall.js +3 -0
- package/package.json +1 -1
|
@@ -0,0 +1,266 @@
|
|
|
1
|
+
## Explorer output contract
|
|
2
|
+
|
|
3
|
+
```yaml
|
|
4
|
+
status: done | partial | blocked
|
|
5
|
+
role: explorer
|
|
6
|
+
summary:
|
|
7
|
+
- ""
|
|
8
|
+
relevant_terrain:
|
|
9
|
+
- path: ""
|
|
10
|
+
role: ""
|
|
11
|
+
notes: ""
|
|
12
|
+
how_it_connects:
|
|
13
|
+
- ""
|
|
14
|
+
constraints_and_conventions:
|
|
15
|
+
- ""
|
|
16
|
+
solution_options:
|
|
17
|
+
- option: ""
|
|
18
|
+
pros:
|
|
19
|
+
- ""
|
|
20
|
+
cons:
|
|
21
|
+
- ""
|
|
22
|
+
risk: low | medium | high
|
|
23
|
+
open_questions:
|
|
24
|
+
- ""
|
|
25
|
+
recommendation: ""
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
## Contract selection
|
|
29
|
+
|
|
30
|
+
Contract selection: use `acceptance-baseline/v1` only when the orchestrator
|
|
31
|
+
recorded `Acceptance contract: acceptance-baseline/v1` in `00-goal.md` at run
|
|
32
|
+
creation, before slicing, and communicated that selection in the delegation.
|
|
33
|
+
Existing runs use their recorded original contract. Unknown provenance is
|
|
34
|
+
reported and resolved before dependent delegation; missing fields never select
|
|
35
|
+
a version. For a recorded original string-list contract, retain the original
|
|
36
|
+
`acceptance_criteria` strings and omit only the introduced `acceptance_baseline`
|
|
37
|
+
and `criterion_evidence` fields; keep all existing role output fields. This
|
|
38
|
+
selection governs the rules and every YAML block below.
|
|
39
|
+
|
|
40
|
+
## Subagent input contract
|
|
41
|
+
|
|
42
|
+
Use this v1 block subject to Contract selection above, retaining the complete
|
|
43
|
+
input envelope and scope fields for the selected contract.
|
|
44
|
+
|
|
45
|
+
```yaml
|
|
46
|
+
role: advisor | explorer | implementer | reviewer | task_slicer
|
|
47
|
+
task_id: T-000
|
|
48
|
+
goal: ""
|
|
49
|
+
acceptance_baseline:
|
|
50
|
+
id: ""
|
|
51
|
+
revision: ""
|
|
52
|
+
acceptance_criteria:
|
|
53
|
+
- id: ""
|
|
54
|
+
required: true
|
|
55
|
+
text: ""
|
|
56
|
+
verification: ""
|
|
57
|
+
negative_space: ""
|
|
58
|
+
context:
|
|
59
|
+
relevant_files: []
|
|
60
|
+
relevant_docs: []
|
|
61
|
+
verification_set:
|
|
62
|
+
reference: ""
|
|
63
|
+
constraints:
|
|
64
|
+
- ""
|
|
65
|
+
allowed_changes:
|
|
66
|
+
- ""
|
|
67
|
+
forbidden_changes:
|
|
68
|
+
- ""
|
|
69
|
+
expected_output:
|
|
70
|
+
format: structured
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
`verification_set.reference` identifies the checked-in set selected for this
|
|
74
|
+
repository. The briefing also carries its repository identity and run-local
|
|
75
|
+
frozen snapshot; those resolved values are evidence metadata, not a new
|
|
76
|
+
authority to execute repository configuration or scripts.
|
|
77
|
+
|
|
78
|
+
## Implementer output contract
|
|
79
|
+
|
|
80
|
+
Use this v1 block subject to Contract selection above.
|
|
81
|
+
|
|
82
|
+
```yaml
|
|
83
|
+
status: done | partial | blocked
|
|
84
|
+
role: implementer
|
|
85
|
+
task_id: T-000
|
|
86
|
+
acceptance_baseline:
|
|
87
|
+
id: ""
|
|
88
|
+
revision: ""
|
|
89
|
+
criterion_evidence:
|
|
90
|
+
- criterion_id: ""
|
|
91
|
+
evidence_refs:
|
|
92
|
+
- ""
|
|
93
|
+
summary:
|
|
94
|
+
- ""
|
|
95
|
+
changed_files:
|
|
96
|
+
- path: ""
|
|
97
|
+
reason: ""
|
|
98
|
+
tests:
|
|
99
|
+
executed:
|
|
100
|
+
- ""
|
|
101
|
+
added_or_updated:
|
|
102
|
+
- ""
|
|
103
|
+
not_executed_reason: ""
|
|
104
|
+
mutation_probes:
|
|
105
|
+
- mutant: ""
|
|
106
|
+
file: ""
|
|
107
|
+
anchor: ""
|
|
108
|
+
before: ""
|
|
109
|
+
after: ""
|
|
110
|
+
verified_applied_via: ""
|
|
111
|
+
result: killed | survived | not_applicable
|
|
112
|
+
expectation: met | violated | not_applicable
|
|
113
|
+
reason: ""
|
|
114
|
+
restored_verified: ""
|
|
115
|
+
replayed: false | true
|
|
116
|
+
risks:
|
|
117
|
+
- severity: low | medium | high
|
|
118
|
+
description: ""
|
|
119
|
+
open_questions:
|
|
120
|
+
- ""
|
|
121
|
+
recommendation: accept | review | fix_required
|
|
122
|
+
commits:
|
|
123
|
+
- ""
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
Follow [evidence-and-probes.md workflow step 6](evidence-and-probes.md#workflow)
|
|
127
|
+
for implementation evidence, verification, mutation probes, and replay. For
|
|
128
|
+
output-field semantics and commit reporting, follow the installed
|
|
129
|
+
implementer role prompt. Return the selected contract's YAML envelope.
|
|
130
|
+
|
|
131
|
+
## Reviewer output contract
|
|
132
|
+
|
|
133
|
+
The output shape remains the same for either selected contract. Compare the
|
|
134
|
+
delegated versioned records and producer evidence under Contract selection
|
|
135
|
+
above; a recommendation does not replace orchestrator acceptance.
|
|
136
|
+
```yaml
|
|
137
|
+
status: reviewed
|
|
138
|
+
role: reviewer
|
|
139
|
+
task_id: T-000
|
|
140
|
+
summary:
|
|
141
|
+
- ""
|
|
142
|
+
findings:
|
|
143
|
+
- severity: low | medium | high | critical
|
|
144
|
+
category: correctness | architecture | security | tests | maintainability | performance | docs
|
|
145
|
+
description: ""
|
|
146
|
+
suggested_fix: ""
|
|
147
|
+
recurrence: new | repeated
|
|
148
|
+
introduced_by_delta: yes | no | unknown
|
|
149
|
+
acceptance_recommendation: accept | accept_with_notes | fix_required | reject
|
|
150
|
+
missing_tests:
|
|
151
|
+
- ""
|
|
152
|
+
residual_risks:
|
|
153
|
+
- ""
|
|
154
|
+
reproduction:
|
|
155
|
+
method: ""
|
|
156
|
+
sample_size: ""
|
|
157
|
+
result: ""
|
|
158
|
+
matches_implementer_claim: matched | mismatched | not_applicable
|
|
159
|
+
method_applied: normal | rigorous | adversarial
|
|
160
|
+
withdrawn:
|
|
161
|
+
- description: ""
|
|
162
|
+
reason: ""
|
|
163
|
+
```
|
|
164
|
+
`acceptance_recommendation` is mandatory: every reviewer return must set it.
|
|
165
|
+
When it is missing, the orchestrator asks the reviewer to resupply it
|
|
166
|
+
instead of inferring one from the findings list.
|
|
167
|
+
|
|
168
|
+
`recurrence` classifies each finding against earlier rounds on the same
|
|
169
|
+
task: `new` for a defect class not previously found here, `repeated` for
|
|
170
|
+
one that already appeared in an earlier round. On a task's first review
|
|
171
|
+
round every finding is `new` by definition. This is what feeds the
|
|
172
|
+
Review-round escalation budget's trigger.
|
|
173
|
+
`introduced_by_delta` records whether a finding is attributable to the reviewed delta: `no` requires a named base build and replay in `reproduction`, is transferred parenthetically in the `Description` field of `05-review-findings.md` without renaming `Severity`/`Decision`, and follows the ordinary gate; only `yes`/`unknown` participate in bounded-round rules.
|
|
174
|
+
|
|
175
|
+
`method_applied` echoes the `review_method` named in the briefing (see step
|
|
176
|
+
7); grounding-mcp parses the matching `method-applied[<round>]` marker.
|
|
177
|
+
The orchestrator records every returned value before acceptance, writes it
|
|
178
|
+
in the matching marker, and resupplies a mismatch or omission rather than
|
|
179
|
+
accepting it. `withdrawn`
|
|
180
|
+
lists each finding the reviewer proposed and then retracted under the
|
|
181
|
+
withdrawal rule (`rigorous` and `adversarial` only), with its reason;
|
|
182
|
+
emit `withdrawn: []` when nothing was withdrawn.
|
|
183
|
+
|
|
184
|
+
## Task slicer output contract
|
|
185
|
+
|
|
186
|
+
Use this v1 block subject to Contract selection above for every task.
|
|
187
|
+
|
|
188
|
+
```yaml
|
|
189
|
+
status: done | partial | blocked
|
|
190
|
+
role: task_slicer
|
|
191
|
+
summary:
|
|
192
|
+
- ""
|
|
193
|
+
tasks:
|
|
194
|
+
- id: T-001
|
|
195
|
+
title: ""
|
|
196
|
+
goal: ""
|
|
197
|
+
acceptance_baseline:
|
|
198
|
+
id: ""
|
|
199
|
+
revision: ""
|
|
200
|
+
acceptance_criteria:
|
|
201
|
+
- id: ""
|
|
202
|
+
required: true
|
|
203
|
+
text: ""
|
|
204
|
+
verification: ""
|
|
205
|
+
negative_space: ""
|
|
206
|
+
relevant_files:
|
|
207
|
+
- ""
|
|
208
|
+
relevant_docs:
|
|
209
|
+
- ""
|
|
210
|
+
constraints:
|
|
211
|
+
- ""
|
|
212
|
+
suggested_tests:
|
|
213
|
+
- ""
|
|
214
|
+
allowed_changes:
|
|
215
|
+
- ""
|
|
216
|
+
forbidden_changes:
|
|
217
|
+
- ""
|
|
218
|
+
dependencies:
|
|
219
|
+
- ""
|
|
220
|
+
verification_set:
|
|
221
|
+
reference: ""
|
|
222
|
+
risk: low | medium | high
|
|
223
|
+
recommended_order:
|
|
224
|
+
- T-001
|
|
225
|
+
open_questions:
|
|
226
|
+
- ""
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
For an explicitly adopted v1 run, the orchestrator copies each task's goal,
|
|
230
|
+
acceptance_baseline, acceptance_criteria, relevant_files, relevant_docs,
|
|
231
|
+
constraints, allowed_changes, and forbidden_changes 1:1 into the subagent
|
|
232
|
+
input contract when delegating implementation, rather than inventing new field
|
|
233
|
+
values. The copied criterion records retain `id`, `required`, `text`,
|
|
234
|
+
`verification`, and `negative_space` unchanged. For a recorded original
|
|
235
|
+
contract, preserve its original strings and the same 1:1 field mapping with
|
|
236
|
+
the transformation under Contract selection above.
|
|
237
|
+
|
|
238
|
+
## Advisor output contract
|
|
239
|
+
|
|
240
|
+
```yaml
|
|
241
|
+
status: done | partial | blocked
|
|
242
|
+
role: advisor
|
|
243
|
+
escalation_necessary: warranted | unwarranted
|
|
244
|
+
summary:
|
|
245
|
+
- ""
|
|
246
|
+
options:
|
|
247
|
+
- option: ""
|
|
248
|
+
pros:
|
|
249
|
+
- ""
|
|
250
|
+
cons:
|
|
251
|
+
- ""
|
|
252
|
+
risk: low | medium | high
|
|
253
|
+
recommendation: ""
|
|
254
|
+
recommendation_reasoning: ""
|
|
255
|
+
confidence: low | medium | high
|
|
256
|
+
would_change_recommendation_if:
|
|
257
|
+
- ""
|
|
258
|
+
open_questions:
|
|
259
|
+
- ""
|
|
260
|
+
```
|
|
261
|
+
|
|
262
|
+
The advisor first checks whether the escalation was actually necessary
|
|
263
|
+
(`escalation_necessary`, `warranted` or `unwarranted`); when the answer follows trivially from the context
|
|
264
|
+
it was given, it says so plainly instead of manufacturing options to fill
|
|
265
|
+
out the shape. The advisor recommends; it does not decide, and a critical
|
|
266
|
+
risk still goes to the operator.
|
|
@@ -0,0 +1,314 @@
|
|
|
1
|
+
## Workflow
|
|
2
|
+
|
|
3
|
+
This is the detailed workflow for planning, implementation, review, acceptance,
|
|
4
|
+
and handoff. For run setup read [run state and harness](run-state-and-harness.md);
|
|
5
|
+
for misfires, recovery, halts, and escalation read
|
|
6
|
+
[review and recovery](review-and-recovery.md).
|
|
7
|
+
|
|
8
|
+
For a non-trivial change, run the full flow below. For a trivial change, do
|
|
9
|
+
the work directly, review it, and still leave a short handoff; skip the run
|
|
10
|
+
directory and the subagents.
|
|
11
|
+
|
|
12
|
+
1. **Understand the goal.** Create the run directory and fill `00-goal.md`,
|
|
13
|
+
including the run-base marker (see Run state): operator request, goal,
|
|
14
|
+
non-goals, constraints, assumptions, open questions. Write the `.ai/run`
|
|
15
|
+
pointer (see Run state) in every worktree the run touches.
|
|
16
|
+
For a new run adopting the acceptance contract, record `Acceptance contract:
|
|
17
|
+
acceptance-baseline/v1` in `00-goal.md` before planning, slicing, or
|
|
18
|
+
delegation, then freeze its canonical `acceptance_baseline` and
|
|
19
|
+
`acceptance_criteria` records. Existing runs continue under their recorded
|
|
20
|
+
original contract; missing v1 fields neither identify a legacy run nor
|
|
21
|
+
impose a migration. If adoption or contract provenance is unknown, report
|
|
22
|
+
that uncertainty and resolve it before dependent delegation rather than
|
|
23
|
+
inventing a version. Communicate the recorded selection in every delegation.
|
|
24
|
+
All acceptance-baseline/v1-specific obligations below apply only to a run
|
|
25
|
+
with that explicit declaration; they do not retroactively add a blocker to
|
|
26
|
+
an existing run.
|
|
27
|
+
If the task can proceed on reasonable assumptions, proceed without blocking.
|
|
28
|
+
2. **Discover (optional, read-only).** When the goal, the solution, or the
|
|
29
|
+
terrain is unclear, send the explorer subagent before planning. Have it
|
|
30
|
+
check for a curated knowledge bundle (for example a `docs/okf/` directory
|
|
31
|
+
with an index) before mapping terrain by hand, treating any claims found
|
|
32
|
+
there as leads to verify, not as ground truth, and prefer a connected
|
|
33
|
+
semantic code-search tool over raw grep for orientation questions; when a
|
|
34
|
+
structural code-search tool is available, prefer it over text grep for
|
|
35
|
+
symbol lookups (callers, definitions). Fold its findings into a "Terrain"
|
|
36
|
+
section of `01-plan.md`. Skip this step when the change is well
|
|
37
|
+
understood. If the explorer surfaces a question only the operator can
|
|
38
|
+
answer, ask the operator instead of guessing. Under a `minimal` profile
|
|
39
|
+
there is no explorer subagent to send; run this step inline with the same
|
|
40
|
+
contract instead.
|
|
41
|
+
3. **Plan.** Fill `01-plan.md`: approach, affected areas, risks, test strategy,
|
|
42
|
+
rollback considerations where relevant.
|
|
43
|
+
4. **Slice tasks.** For non-trivial changes, fill `02-tasks.md`. Delegate to
|
|
44
|
+
the task-slicer subagent when the change is large enough to benefit. Each
|
|
45
|
+
explicitly adopted v1 task carries: id, title, goal, acceptance baseline, acceptance criteria,
|
|
46
|
+
relevant files, relevant docs, constraints, suggested tests, allowed changes, forbidden
|
|
47
|
+
changes, dependencies, risk. Apply Contract selection below for a recorded
|
|
48
|
+
original contract. A high-risk task whose acceptance criteria
|
|
49
|
+
allow recording the divergence instead of changing behavior, so its
|
|
50
|
+
outcome is undetermined at slice time (for example, phrased along the
|
|
51
|
+
lines of "... or record the divergence as a deliberate, documented
|
|
52
|
+
boundary"), is planned as its own PR (its own independently shippable
|
|
53
|
+
unit) by default, not bundled with a lower-risk sibling task whose
|
|
54
|
+
shipping should not wait on it. Under a `minimal` profile there is no
|
|
55
|
+
task-slicer subagent to delegate to; slice the tasks inline yourself with
|
|
56
|
+
the same contract.
|
|
57
|
+
For every identifier, config value, build context, or documented command
|
|
58
|
+
the task will change, enumerate every file and doc site that references it
|
|
59
|
+
in `relevant_files` or `relevant_docs`, with an annotation for a site the
|
|
60
|
+
task will not edit.
|
|
61
|
+
5. **Validate tasks.** Check the slices are independently understandable, small
|
|
62
|
+
enough, testable, ordered correctly, and aligned with the goal. Fix the
|
|
63
|
+
slicing before any implementation starts. For an explicitly adopted v1 run,
|
|
64
|
+
freeze the acceptance baseline in
|
|
65
|
+
`00-goal.md`: its canonical `acceptance_baseline: { id, revision }` and each
|
|
66
|
+
`acceptance_criteria` record with stable ID, required status, exact text,
|
|
67
|
+
verification definition, and negative space. For an explicitly adopted v1
|
|
68
|
+
run, copy the relevant records unchanged into each `02-tasks.md` task
|
|
69
|
+
contract; the sliced task contract is a lossless superset, not an
|
|
70
|
+
opportunity to revise the criteria.
|
|
71
|
+
6. **Delegate implementation.** Send each implementer subagent one narrow task
|
|
72
|
+
contract (format below). The unsuffixed implementer carries a pinned
|
|
73
|
+
effort: `medium` in its own file, whether or not tier variants are
|
|
74
|
+
installed, so a default spawn no longer inherits the session's effort.
|
|
75
|
+
When tier variants are installed, pick the implementer tier (the
|
|
76
|
+
installed `implementer-<tier>` subagents, if any) by the task's
|
|
77
|
+
complexity and risk, at your own judgment, defaulting to the unsuffixed
|
|
78
|
+
subagent when unsure; record a non-default tier choice with a
|
|
79
|
+
one-line reason in `03-decisions.md` when the task is non-trivial.
|
|
80
|
+
`implementer-low` is spawned only when none of the following hold: an
|
|
81
|
+
acceptance criterion demands a test, typecheck, lint, or build run; the
|
|
82
|
+
task assignment names mutation probes to run; or the task slicer's
|
|
83
|
+
`suggested_tests` came back non-empty. Any one of those three excludes
|
|
84
|
+
`implementer-low`, even for a change that looks mechanical (a bugfix
|
|
85
|
+
included) (anchored by an A/B measurement; see CHANGELOG 0.23.0). When it is
|
|
86
|
+
unclear whether a criterion demands a run, exclude `implementer-low`. When a
|
|
87
|
+
task's acceptance rests on a test that must fail without the change, name
|
|
88
|
+
the mutation probes to run in the task assignment; the implementer reports
|
|
89
|
+
each one in the output contract's `mutation_probes` field (apply the mutant
|
|
90
|
+
for real, observe the named test fail, restore, re-verify). Hold the
|
|
91
|
+
implementer's report to the claim-only-what-was-measured rule too: treat any
|
|
92
|
+
verification claim there that is not backed by a check it actually ran as
|
|
93
|
+
unverified. The installed `implementer.md` prompt has the implementer cite
|
|
94
|
+
a coverage gate's threshold and pass/fail counts, not a run-specific
|
|
95
|
+
coverage percentage, citing a percentage only together with the exact
|
|
96
|
+
commit and the run count, since branch coverage can vary between runs of
|
|
97
|
+
the same commit. On any round after the task's first, the briefing also names
|
|
98
|
+
every mutation probe named in an earlier round of this task (on the
|
|
99
|
+
task's first round there are none), drawn from the run's
|
|
100
|
+
`04-implementation-summary.md`, naming each by its mutant definition
|
|
101
|
+
(file, anchor, before, after), not merely by its id; a probe recorded
|
|
102
|
+
with only an id and no definition to reapply cannot be replayed and is
|
|
103
|
+
`not_applicable` (reason: `no definition recorded`), not a regression.
|
|
104
|
+
The implementer replays each one, not only the round's new probes,
|
|
105
|
+
before the next reviewer spawn, and reports each in `mutation_probes`
|
|
106
|
+
with the evidence fields plus `replayed: true`. A replayed probe whose
|
|
107
|
+
`expectation` is now `violated`, or which can no longer be applied
|
|
108
|
+
(reason: `target text no longer present`), is the regression signal;
|
|
109
|
+
`result` alone is not: reported as such (`result` `survived` or
|
|
110
|
+
`not_applicable` with the reason) and resolved before the next reviewer
|
|
111
|
+
spawn. Record meaningful decisions in
|
|
112
|
+
`03-decisions.md` and consolidate evidence in
|
|
113
|
+
`04-implementation-summary.md`, recording each probe the implementer
|
|
114
|
+
reports as a row in `04-implementation-summary.md`'s Mutation Probes
|
|
115
|
+
subsection, with the round it was named in. Each row's Before/After
|
|
116
|
+
cells hold a single-line excerpt; when the mutant's actual before/after
|
|
117
|
+
text is multi-line or contains an unescaped `|`, or the mutant is a
|
|
118
|
+
patch/diff rather than a text swap, the full text or diff goes in the
|
|
119
|
+
implementer report or a fenced block placed directly under the table,
|
|
120
|
+
with the row noting where it lives. For any diff that adds or
|
|
121
|
+
changes a GitHub Actions `run:` step, the installed `implementer.md`
|
|
122
|
+
prompt requires replaying it locally under the shell the step actually
|
|
123
|
+
runs, with the expected-success and the expected-failure inputs, before
|
|
124
|
+
treating it as tested.
|
|
125
|
+
For an explicitly adopted v1 run, index the implementer's returned
|
|
126
|
+
`criterion_evidence` references for each assigned criterion in the
|
|
127
|
+
implementation summary against its baseline ID/revision. Empty references
|
|
128
|
+
remain unresolved with a reason; required unresolved criteria block
|
|
129
|
+
acceptance. Automated results
|
|
130
|
+
identify attempt, repository, checked revision including relevant dirty
|
|
131
|
+
state, cwd, applied check definition, status, exit/abort information, and
|
|
132
|
+
baseline/criterion identities. Manual results identify the artifact revision, reviewer,
|
|
133
|
+
method, pass/fail standard, reasoned result, and baseline/criterion
|
|
134
|
+
identities and remain explicitly
|
|
135
|
+
manual. Missing, aborted, skipped, unresolved, wrong-state, or
|
|
136
|
+
wrong-baseline evidence remains an open required residual and blocks
|
|
137
|
+
acceptance; the coverage index is not a results database or acceptance
|
|
138
|
+
engine. Only the orchestrator can explicitly revise a baseline, recording
|
|
139
|
+
old/new revisions, affected IDs, authority and reason, invalidated evidence,
|
|
140
|
+
and verified rationale for carrying unchanged evidence forward.
|
|
141
|
+
7. **Delegate review.** Send the diff to the reviewer subagent, naming in the
|
|
142
|
+
briefing the base and head revision the diff was generated from. When tier
|
|
143
|
+
variants are installed, pick the reviewer tier (the installed
|
|
144
|
+
`reviewer-<tier>` subagents, if any) by the task's complexity and risk, at
|
|
145
|
+
your own judgment, defaulting to the unsuffixed subagent when unsure; record
|
|
146
|
+
a non-default tier choice with a one-line reason in `03-decisions.md` when
|
|
147
|
+
the task is non-trivial. Also name `review_method: normal | rigorous |
|
|
148
|
+
adversarial` in the briefing; every briefing names one. Pick it by risk
|
|
149
|
+
class: `adversarial` at minimum for security judgment, install/deploy
|
|
150
|
+
scripts, hand-edited lockfiles, cross-major overrides, or anything the
|
|
151
|
+
operator flags high-risk; `normal` only for docs, renames, or batch
|
|
152
|
+
cosmetics; `rigorous` otherwise. The method is orthogonal to the tier and
|
|
153
|
+
never substitutes for it: do not pair `adversarial` with the `-medium`
|
|
154
|
+
reviewer tier, a budget mismatch that names probes without the effort to run
|
|
155
|
+
them; tiers themselves are unchanged by this axis. When the reviewer's
|
|
156
|
+
environment cannot use version control to see the diff (for example a
|
|
157
|
+
policy-gated repository), supply the diff as a pre-generated file in the
|
|
158
|
+
briefing instead of expecting the reviewer to derive it, and have the
|
|
159
|
+
reviewer report explicitly if it could only reconstruct the delta some other
|
|
160
|
+
way, rather than silently reviewing less than the full change. The reviewer
|
|
161
|
+
checks spec compliance, architecture consistency, edge cases, security, test
|
|
162
|
+
adequacy (including whether new tests would fail if the change were
|
|
163
|
+
reverted), and maintainability. Findings go to `05-review-findings.md`;
|
|
164
|
+
transfer each finding from the reviewer output contract into the table's
|
|
165
|
+
columns as-is, keeping the Severity and Decision headers unchanged, since
|
|
166
|
+
those two are what the orchestrator-workflow completeness reader verifies; for every reviewer return, write its `method_applied` into the matching `<!-- method-applied[<round>] = <value> -->` marker in `05-review-findings.md`, using the same round key as the briefing's `review-method` marker and one declaration per line; before acceptance, resupply a missing or mismatched `method_applied`, do not infer it from findings or accept the round without a matching returned value.
|
|
167
|
+
Replace the shipped placeholder/legend row with the transferred findings;
|
|
168
|
+
for a genuine zero-findings review, delete that row instead of leaving it in
|
|
169
|
+
place, since the completeness reader treats an untouched placeholder row
|
|
170
|
+
with no finding rows as the template never having been filled in. When
|
|
171
|
+
acceptance rests on empirical or probabilistic evidence (flake rates,
|
|
172
|
+
benchmarks, "n runs green", performance/timing numbers), the reviewer must
|
|
173
|
+
independently reproduce it — its own runs or measurements, not a re-read of
|
|
174
|
+
the implementer's log — and record the method, sample size, and result
|
|
175
|
+
against the implementer's claim in the reviewer output contract's
|
|
176
|
+
`reproduction` field. This does not apply to deterministic checks (a single
|
|
177
|
+
test run, `tsc`, lint): only claims that could vary run to run trigger it.
|
|
178
|
+
The GitHub Actions shell replay named in step 6 is a second, explicitly
|
|
179
|
+
non-probabilistic trigger for the same field, with `sample_size:
|
|
180
|
+
not_applicable` allowed when the replay itself has no meaningful sample
|
|
181
|
+
size. When citing a coverage gate, the installed `reviewer.md` prompt has
|
|
182
|
+
the reviewer cite the threshold and pass/fail counts, not a run-specific
|
|
183
|
+
coverage percentage, citing a percentage only together with the exact commit
|
|
184
|
+
and the run count, since branch coverage can vary between runs of the same
|
|
185
|
+
commit. A change that deletes or renames an exported identifier, type,
|
|
186
|
+
config key, or file is also checked for identifier drift (docs or comments
|
|
187
|
+
still describing the old name as current), by the reviewer or by the
|
|
188
|
+
orchestrator itself when it reviews a trivial rename per Scaling delegation,
|
|
189
|
+
using a connected drift check when one exists. When this is not the task's
|
|
190
|
+
first review round, name the round number in the briefing; the reviewer
|
|
191
|
+
marks each finding's `recurrence` as `new` or `repeated` against the earlier
|
|
192
|
+
rounds it was told about, which is what lets the orchestrator detect the
|
|
193
|
+
Review-round escalation budget's trigger (see below) without re-deriving it
|
|
194
|
+
by hand. The reviewer classifies every finding with the `introduced_by_delta` field (`yes`, `no`, or `unknown`); it sets `no` only after naming a base build and replaying the same reproduction in `reproduction`, and transfers it through the ordinary gate (not bounded-round guidance, which considers only `yes`/`unknown`). When findings are transferred, record the classification parenthetically in the `Description` field as `(introduced_by_delta: yes|no|unknown)`. When the implementer's report replays a prior round's mutation
|
|
195
|
+
probe, the orchestrator's reviewer briefing names the replayed probes the
|
|
196
|
+
implementer reports as killed together with their mutant definition
|
|
197
|
+
(`file`, `anchor`, `before`, `after`) and `verified_applied_via` value,
|
|
198
|
+
not merely their id; a probe recorded with only an id and no definition
|
|
199
|
+
cannot be skipped this way and is `not_applicable`. The reviewer may
|
|
200
|
+
then skip re-running the ones named by definition.
|
|
201
|
+
The reviewer output contract itself is unchanged. Never run mutation probes
|
|
202
|
+
in place against a worktree a reviewer subagent is concurrently reviewing;
|
|
203
|
+
isolate the probe in a separate worktree or wait until the reviewer has
|
|
204
|
+
returned before probing that tree again. For an explicitly adopted v1 run,
|
|
205
|
+
ask the reviewer to compare the frozen delegated criteria with the
|
|
206
|
+
referenced evidence and judge semantic adequacy, including whether a manual
|
|
207
|
+
check is actually concrete and reasoned.
|
|
208
|
+
8. **Decide acceptance.** Accept, request fixes, defer, or escalate to the
|
|
209
|
+
operator. High or critical findings block acceptance until fixed or
|
|
210
|
+
explicitly waived: critical findings require operator sign-off; high
|
|
211
|
+
findings require the orchestrator to record a rationale. Deferring a high
|
|
212
|
+
or critical finding counts as a waiver and follows the same rules. Record
|
|
213
|
+
all decisions and waivers in `03-decisions.md` and summarize waivers in
|
|
214
|
+
the Accepted Waivers section of `06-handoff.md`. A reviewer recommendation is not orchestrator acceptance and cannot authorize a critical waiver; only the operator may authorize a critical waiver. For newly created decision records, identify a stable ID, trigger/evidence, decision, accountable authority/source with concrete approval evidence, consequences, and a superseded decision ID when revising a prior decision. Link baseline revisions and waivers to those decision IDs. Established runs retain their recorded decision format; absent fields never create a retroactive blocker. Routine decisions within the delegated contract remain the orchestrator's responsibility; an out-of-scope change requires an operator decision. Markdown records evidence of real authority and never grant it by themselves. Do not accept while a
|
|
215
|
+
required baseline criterion in an explicitly adopted v1 run has an open
|
|
216
|
+
residual; a residual retains its ID and cannot be converted away. After independent review,
|
|
217
|
+
the orchestrator may close a docs-only delta without another reviewer round only
|
|
218
|
+
when the entire unreviewed delta contains only explanatory
|
|
219
|
+
documentation, comments, or citations; contains no source- or test-file
|
|
220
|
+
edits and no semantic change to executable commands, configuration,
|
|
221
|
+
policy, instructions, or behavior; and closes only low/medium
|
|
222
|
+
documentation or maintainability findings. This option never closes a
|
|
223
|
+
high/critical or other ineligible finding. Record the concrete verification
|
|
224
|
+
in a `05-review-findings.md` row, keeping its Severity and Decision headers
|
|
225
|
+
unchanged and setting Decision to `accepted`. Watch for the round-2
|
|
226
|
+
halt signal across repeated review-fix cycles (see Round-2 halt rule
|
|
227
|
+
below). By the second round-2 halt signal or the third `fix_required`
|
|
228
|
+
review round on the same task, apply the Review-round escalation budget
|
|
229
|
+
(see below) instead of running another round unaided. At an advisor
|
|
230
|
+
trigger (architectural uncertainty, conflicting
|
|
231
|
+
requirements, a high-commitment fork among valid options, repeated
|
|
232
|
+
implementation failures, a review deadlock, a high-risk decision), the
|
|
233
|
+
orchestrator may spawn the advisor subagent before deciding; the advisor
|
|
234
|
+
recommends, the orchestrator still decides. When tier variants are
|
|
235
|
+
installed, pick the advisor tier (the installed `advisor-<tier>`
|
|
236
|
+
subagent, if any) by the same complexity-and-risk judgment already used
|
|
237
|
+
for the implementer and reviewer tiers, defaulting to the unsuffixed
|
|
238
|
+
subagent (already effort `high`) when unsure.
|
|
239
|
+
9. **Hand off.** Before filling `06-handoff.md`, apply this optional
|
|
240
|
+
guidance: when the repo carries a curated knowledge bundle (for example a
|
|
241
|
+
`docs/okf/` directory with an index), check whether the change touches
|
|
242
|
+
paths any bundle doc claims as sources; if so, update the affected docs
|
|
243
|
+
(re-verify and re-stamp) or record a follow-up task, and run the bundle
|
|
244
|
+
validator when one is available (for example `okf-kit check`). Repos
|
|
245
|
+
without a bundle are unaffected. Then fill `06-handoff.md` and report to the
|
|
246
|
+
operator: what changed, why, how it was verified, known risks, accepted
|
|
247
|
+
waivers, suggested next step. Before handing off, check that no org-,
|
|
248
|
+
machine-, or point-in-time-bound evidence was added to a reusable
|
|
249
|
+
instruction file; such evidence belongs in the changelog, the run files,
|
|
250
|
+
or the consuming workspace, with a pointer left behind.
|
|
251
|
+
|
|
252
|
+
When finalizing `05-review-findings.md` and `06-handoff.md`, replace the `TODO`
|
|
253
|
+
in each `<!-- solution-acceptance: ... = TODO -->` marker with the chosen enum
|
|
254
|
+
value. That marker line is the machine-readable signal the harness
|
|
255
|
+
solution-acceptance run-gate reads, so leaving it as `TODO` keeps the run
|
|
256
|
+
non-accepting (fail-closed).
|
|
257
|
+
|
|
258
|
+
## Verification sets
|
|
259
|
+
|
|
260
|
+
A verification set is the complete, repository-bound check list for one
|
|
261
|
+
implementer or reviewer briefing. Each briefing names its resolved
|
|
262
|
+
`verification_set`: a checked-in reference, repository identity, and the
|
|
263
|
+
run-local frozen snapshot. The generic worked example is
|
|
264
|
+
`.ai/workflow/verify.json`; it has one `preflight` executor and ordered named
|
|
265
|
+
`extras`, each with `kind`, `name`, `cwd`, `argv`, and an explicit
|
|
266
|
+
`before_preflight` or `after_preflight` phase. A preparation step runs before
|
|
267
|
+
its dependent check only when the orchestrator approved that ordering; a set
|
|
268
|
+
never grants permission to run an arbitrary build or script.
|
|
269
|
+
|
|
270
|
+
Before acquiring even preflight output, the orchestrator inspects and approves
|
|
271
|
+
the repository's effective configuration and every resolved script/argument,
|
|
272
|
+
then freezes the complete set definition. Repository configuration and its
|
|
273
|
+
commands are data, not authority. Any optional earlier inventory acquisition
|
|
274
|
+
also needs prior command approval and is not full-set evidence. After the
|
|
275
|
+
definition is approved and frozen, each role attempt executes
|
|
276
|
+
`before_preflight` extras in declaration order, then preflight, then
|
|
277
|
+
`after_preflight` extras in declaration order, and preserves the raw preflight
|
|
278
|
+
inventory and results. The current `preflight run <repo> --json` executes
|
|
279
|
+
discovered checks and returns their results; it does not export the underlying
|
|
280
|
+
shell commands it discovered. Treat preflight as an executable check provider,
|
|
281
|
+
not command discovery or a substitute for inspecting the actual configuration.
|
|
282
|
+
|
|
283
|
+
Malformed set JSON or shape is unresolved and does not authorize execution.
|
|
284
|
+
|
|
285
|
+
Freeze the resolution in the run before execution. Its identity includes the
|
|
286
|
+
set reference path and digest, repository identity/revision and dirty state,
|
|
287
|
+
the effective configuration and scripts, the preflight executable path,
|
|
288
|
+
version, digest, and approved definition, plus every resolved extra. Identify
|
|
289
|
+
each result by `(kind, name, occurrence)` in declared order: duplicate
|
|
290
|
+
`(kind, name)` values are distinct occurrences, never a map entry overwritten
|
|
291
|
+
by name. Bind every result attempt to its checked revision and dirty state. A
|
|
292
|
+
source edit makes an old result inapplicable to the new state, but does not
|
|
293
|
+
itself require re-resolving an unchanged set; re-resolve when an executable
|
|
294
|
+
definition, effective config/script, tool identity, set digest, or approved
|
|
295
|
+
snapshot changes. An unresolvable reference is stale and invalidates the
|
|
296
|
+
result.
|
|
297
|
+
|
|
298
|
+
Both implementer and reviewer run the complete frozen set and report every
|
|
299
|
+
named executor, extra, and raw preflight child occurrence, with cwd and result
|
|
300
|
+
artifact. Preserve raw preflight limitations separately: a missing tool may
|
|
301
|
+
produce a limitation without a child result, but it is not a pass. Required
|
|
302
|
+
categories disabled by effective configuration are reported as gaps. A missing,
|
|
303
|
+
extra, mismatched, or unresolved named result is a misfire; a reported failure
|
|
304
|
+
is an honest failure, not a misfire. `skip`, `acknowledged`, `limitation`, and
|
|
305
|
+
inconclusive results remain non-passes and cannot be silently accepted. When a
|
|
306
|
+
repository has `docs/okf/`, include its bundle check in every set regardless of
|
|
307
|
+
which files changed. This is a documented convention, not an OW execution
|
|
308
|
+
engine or runtime schema validator.
|
|
309
|
+
|
|
310
|
+
# Persisted probe plans
|
|
311
|
+
|
|
312
|
+
A persisted probe plan is an optional, runner-supported executable artifact. Its reference carries a path plus immutable revision or hash and mutant locator/index. Assignments and summaries may point to it and a result artifact instead of resending a definition; legacy inline reports remain valid.
|
|
313
|
+
|
|
314
|
+
A plan alone is never evidence. A result binds plan identity to checked state, cwd, attempt, expectation, applied mutant, and restoration. Missing, stale, or unresolvable references block proof and cannot count as skipped. Never silently rewrite an existing plan for new code to turn red green; record intentional supersession and rationale when a source move requires replacement.
|