orchestrator-workflow 0.41.0 → 0.43.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,252 @@
1
+ # Model routing reference
2
+
3
+ The default per-role model table, the full `--routing` JSON shape, the Codex
4
+ default routing table, opencode model resolution, and the effort-tiers
5
+ mechanics behind the [package README](../README.md)'s "Model preselection"
6
+ section.
7
+
8
+ ## Default model routing
9
+
10
+ `--models` is the backward-compatible, per-role input for Claude Code and
11
+ opencode. It does not configure Codex. Existing manifests that contain only
12
+ `models` continue to produce the same Claude Code and opencode defaults:
13
+
14
+ | Role | Default | Why |
15
+ |---|---|---|
16
+ | explorer | `sonnet` | read-only terrain mapping is broad reading, not deep reasoning |
17
+ | task-slicer | `sonnet` | structured decomposition, no deep reasoning needed |
18
+ | implementer | `sonnet` | fast, cheap, good enough for narrow pre-sliced tasks |
19
+ | reviewer | `opus` | skeptical review benefits from the strongest model |
20
+ | advisor | `opus` | escalations happen precisely when the situation is hard, so it shares the reviewer's strongest-model default |
21
+
22
+ The orchestrator itself runs on the session's main model. For Codex, start the
23
+ orchestrator on `gpt-6-astra` at `high` effort; use `xhigh` for demanding work.
24
+ The installer does not mutate global or fleet Codex configuration to enforce
25
+ that recommendation.
26
+
27
+ ## Routing JSON shape
28
+
29
+ Routing is a harness-specific map from role and tier to a complete
30
+ `{model, effort}` selection. Pass a JSON file with `--routing`; the CLI deep
31
+ merges only the leaves you provide and records the resulting effective map in
32
+ `.ai/workflow/manifest.json`. The role's default-tier key configures its
33
+ unsuffixed file; another allowed key configures the corresponding
34
+ `<role>-<tier>` variant when `--tiers` is enabled. For example:
35
+
36
+ ```json
37
+ {
38
+ "codex": {
39
+ "implementer": {
40
+ "medium": { "model": "gpt-6.1-sol", "effort": "medium" },
41
+ "xhigh": { "model": "gpt-6-astra", "effort": "xhigh" }
42
+ }
43
+ }
44
+ }
45
+ ```
46
+
47
+ An omitted `--routing` preserves the exact persisted map on a re-install.
48
+ Changing one leaf leaves the others intact, which makes a previous manifest a
49
+ usable rollback record. Model updates are deliberate per role and tier: the
50
+ installer never interprets a newer model as automatically better and never
51
+ rewrites a preserved choice merely because another model exists.
52
+
53
+ ## Codex defaults
54
+
55
+ Codex uses native `.codex/agents/*.toml` custom agents. The file shape
56
+ follows the
57
+ [official Codex subagent configuration](https://learn.chatgpt.com/docs/agent-configuration/subagents).
58
+ The shipped routing is:
59
+
60
+ | Role | Tier | Model | Effort |
61
+ |---|---|---|---|
62
+ | explorer | low | `gpt-6-luna` | low |
63
+ | explorer | medium (default) | `gpt-6.1-sol` | medium |
64
+ | explorer | high | `gpt-6.1-sol` | high |
65
+ | task-slicer | low | `gpt-6.1-sol` | low |
66
+ | task-slicer | medium (default) | `gpt-6.1-sol` | medium |
67
+ | task-slicer | high | `gpt-6.1-sol` | high |
68
+ | implementer | low | `gpt-6-luna` | low |
69
+ | implementer | medium (default) | `gpt-6.1-sol` | medium |
70
+ | implementer | high | `gpt-6.1-sol` | high |
71
+ | implementer | xhigh | `gpt-6-astra` | xhigh |
72
+ | reviewer | medium | `gpt-6.1-sol` | medium |
73
+ | reviewer | high (default) | `gpt-6-astra` | high |
74
+ | reviewer | xhigh | `gpt-6-astra` | xhigh |
75
+ | advisor | high (default) | `gpt-6-astra` | high |
76
+ | advisor | xhigh | `gpt-6-astra` | xhigh |
77
+
78
+ When you have a deterministic Codex model catalog, pass it with
79
+ `--codex-catalog <json-file>`. The CLI validates the selected Codex model and
80
+ effort pairs before writing. Without a supplied catalog it performs no online
81
+ entitlement check; offline or account-specific availability remains unknown.
82
+ Use the harness's native capability commands, such as `codex debug models`, to
83
+ refresh a catalog before installation when appropriate. A bundled-capability
84
+ view describes what the binary knows and does not prove account entitlement.
85
+
86
+ The bundled `assets/codex-models.json` maps the internal aliases `small`,
87
+ `balanced`, and `strong` to the shipped concrete IDs. These aliases are
88
+ installer inputs only: native agent files and manifests always record the
89
+ resolved concrete ID. To migrate only selected classes, pass a sparse JSON
90
+ file through `--codex-models <json-file>`, for example
91
+ `{"balanced":"gpt-example"}`. It creates complete routing leaves with each
92
+ matching tier's standard effort; it does not alter omitted aliases or a plain
93
+ reinstall. An explicit `--routing` leaf is applied afterward and has highest
94
+ precedence, including when a custom effort must be retained.
95
+
96
+ ## opencode model resolution
97
+
98
+ opencode requires fully-qualified `provider/model-id`
99
+ strings (e.g. `github-copilot/claude-sonnet-4.6`). At install time the CLI
100
+ runs `opencode models` to fetch the live catalog and auto-detects which
101
+ provider to use (the one that offers Claude models). When exactly one such
102
+ provider exists the aliases are resolved to the highest-version matching id in
103
+ the catalog. When multiple providers offer Claude models the CLI warns and asks
104
+ you to pass `--opencode-provider <id>` to disambiguate, or to supply
105
+ fully-qualified ids per role via `--models`. If no resolution is possible
106
+ (catalog empty, `opencode` binary absent, ambiguous provider) the `model:`
107
+ frontmatter line is omitted entirely and the subagent inherits the
108
+ session/default model, a safe, portable fallback. Fully-qualified ids in
109
+ `--models` always pass through unchanged regardless of the catalog.
110
+ Nested-path providers like `openrouter` (whose ids look like
111
+ `openrouter/anthropic/claude-...`) are not auto-resolved from aliases and must
112
+ be supplied as a fully-qualified `--models` entry, e.g.
113
+ `reviewer=openrouter/anthropic/claude-opus-4.8`.
114
+
115
+ ## Effort tiers
116
+
117
+ `--tiers` renders an additional per-role subagent definition for each
118
+ non-default effort tier, alongside the one default (unsuffixed) agent file
119
+ `--profile` already installs. Each tier variant is a standalone subagent
120
+ definition, not a modification of the default file. Claude Code and opencode
121
+ use `<role>.md` / `<role>-<tier>.md`; Codex uses `<role>.toml` /
122
+ `<role>-<tier>.toml`.
123
+
124
+ **Every default file carries its own pinned effort, independent of
125
+ `--tiers`.** The harness composers add the role's own default routing
126
+ selection to the unsuffixed file. In the legacy Claude/opencode path this is
127
+ `TIER_DEFS[DEFAULT_TIER[role]].effort`: `effort: medium` for explorer,
128
+ task-slicer, and implementer; `effort: high` for reviewer and advisor
129
+ (opencode: a `variant: high` line when the resolved model is Claude-family,
130
+ following the same dispatch rule tier variants use, `reasoningEffort:
131
+ medium`/`reasoningEffort: high` for a non-Claude-family provider-qualified
132
+ model, nothing for Ollama, a provider-less id, or an unresolved model). This
133
+ pin does not depend on `tiers`, so a plain install (no `--tiers`) already
134
+ carries it; the flag only controls whether the additional `<role>-<tier>.md`
135
+ variant files are also rendered. The motivation: a default spawn used to
136
+ silently inherit the orchestrator session's own effort, so a `high`-effort
137
+ orchestrator session made every default subagent spawn at `high` too,
138
+ regardless of the role's own intended weight; the pin makes each role's
139
+ effort deterministic and independent of the caller's session. A `--tiers`-off
140
+ install (the default) has no variant files and therefore no in-install
141
+ escalation path off a default's pinned effort; run `init --tiers` afterward
142
+ if a task ever needs one.
143
+
144
+ Default off, like every optional pack in this kit: a fresh install renders
145
+ no variant files unless asked. `--tiers` turns the feature on for that run,
146
+ `--no-tiers` turns it off; a plain re-run with neither flag keeps whatever
147
+ the previous install had, the same override-vs-persist rule already used
148
+ for `--profile` and `--models`. There is no interactive prompt for it:
149
+ `tiers` is opt-in/off via the flags only. Neither Codex nor the other harnesses
150
+ get `max` or `ultra` variants from this kit.
151
+
152
+ ```bash
153
+ npx orchestrator-workflow init --tiers --yes
154
+ ```
155
+
156
+ Turning tiers back off with `--no-tiers` after having them on follows the
157
+ same pattern as a `full` to `minimal` profile downgrade: `init` prints a note
158
+ naming the now-untracked `<role>-<tier>.md` variant files and how to remove
159
+ them, rather than deleting them or leaving the leftover unexplained.
160
+
161
+ **Which tiers each role gets.** A role never gets a variant file for its own
162
+ default tier: that would collide with, and duplicate, the default file.
163
+
164
+ | Role | Tiers available | Default tier (no variant file) |
165
+ |---|---|---|
166
+ | explorer | low, medium, high | medium |
167
+ | task-slicer | low, medium, high | medium |
168
+ | implementer | low, medium, high, xhigh | medium |
169
+ | reviewer | medium, high, xhigh | high |
170
+ | advisor | high, xhigh | high |
171
+
172
+ With `--profile full` and every tier rendered, that is 5 default files plus
173
+ 10 variant files: 15 files total per harness.
174
+
175
+ **Tier to model class to effort.** Each tier resolves to a model class and an
176
+ effort value:
177
+
178
+ | Tier | Model class | Model alias | Effort requested |
179
+ |---|---|---|---|
180
+ | low | small | `haiku` | `low` |
181
+ | medium | medium | `sonnet` | `medium` |
182
+ | high | medium | `sonnet` | `high` |
183
+ | xhigh | large | `opus` | `xhigh` |
184
+
185
+ Claude Code variants carry both a `model:` line (the class's alias) and an
186
+ `effort: <tier>` line in frontmatter. Read-only roles (explorer, reviewer,
187
+ advisor) keep `disallowedTools: Edit, Write, NotebookEdit` on their variants
188
+ too.
189
+
190
+ **opencode variants key off the resolved model's family, not its provider
191
+ prefix**, since opencode's effort surface is not uniform across model
192
+ families:
193
+
194
+ - **Claude-family models** (any resolved id whose provider is
195
+ `anthropic/`, or whose segment after the provider prefix contains
196
+ `claude-`, which covers `anthropic/claude-...` as well as a Claude model
197
+ fronted by a different provider, e.g. `github-copilot/claude-sonnet-4.6`
198
+ or the nested `openrouter/anthropic/claude-opus-4.8`): only `high` and
199
+ `xhigh` get an effort field, as `variant: high` and `variant: max`
200
+ respectively; `low` and `medium` collapse to no effort field at all,
201
+ since opencode's `variant:` option does not distinguish an effort below
202
+ `high`. This collapse is deliberate and documented, not a bug: a
203
+ `low`/`medium` variant on a Claude-family model still gets its class's
204
+ `model:` line, just no `variant:` line.
205
+ - **Ollama, or an id with no provider prefix**: no effort field at all.
206
+ There is no known effort passthrough for Ollama, and an id with no `/`
207
+ resolves to no provider to key the decision on.
208
+ - **Every other non-Claude-family model**: a plain `reasoningEffort: <tier>`
209
+ line, `xhigh` included (opencode's built-in OpenAI-style variants
210
+ document an `xhigh` reasoning effort).
211
+
212
+ The variant's `model:` line is resolved the same way the base per-role model
213
+ is (an `opencode models` catalog lookup against the auto-detected or
214
+ `--opencode-provider`-specified provider), just keyed by the tier's model
215
+ class instead of by role. When that lookup cannot resolve a model for a
216
+ class, the CLI warns once on stderr and **no variant file is rendered for
217
+ that class at all**, not a file with a missing `model:` line: a variant
218
+ with no resolved model would carry neither a `model:` nor an effort line, an
219
+ indistinguishable no-op duplicate of the base file with no ledger entry to
220
+ compare it against, so `init` skips writing it entirely. This guard and its
221
+ warning are opencode-scoped only; Claude Code variants resolve `model:` from
222
+ a plain alias (`haiku`/`sonnet`/`opus`) and need no live catalog lookup, so
223
+ they are unaffected.
224
+
225
+ Codex variants carry `model` and `model_reasoning_effort` from their exact
226
+ routing leaf. The canonical role prompt becomes `developer_instructions`.
227
+ Runtime dispatch follows the client's actual capabilities: select the named
228
+ installed agent when supported; otherwise, if spawning supports explicit model
229
+ and effort, read the installed TOML and pass its selection, developer
230
+ instructions, and narrow task contract into a fresh task-local spawn. A
231
+ full-history spawn may not permit a model override. If that explicit spawn
232
+ cannot accept a sandbox override, explorer and advisor inherit the caller's
233
+ sandbox and their prompt is the edit guard. When native spawning is
234
+ unavailable, run the same contract inline and sequentially. The orchestrator
235
+ alone spawns agents. In particular, it must not choose `implementer-low` when
236
+ the task requires a test, typecheck, lint, build, or named mutation probe.
237
+
238
+ **Warning: `CLAUDE_CODE_EFFORT_LEVEL` overrides every agent's frontmatter
239
+ `effort:`, tier variants included.** Claude Code's `effort:` frontmatter
240
+ field does work: it reaches the model request as `output_config.effort`.
241
+ But when the harness environment sets `CLAUDE_CODE_EFFORT_LEVEL`, that
242
+ environment variable wins over the frontmatter `effort:` on every installed
243
+ agent, tier variants and default files alike, not just the one this feature
244
+ adds. Check for it before relying on a specific tier variant's requested
245
+ effort actually taking effect.
246
+
247
+ The pin is also emitted unconditionally regardless of which model the role
248
+ resolves to via `--models`, including a model with no effort support at all
249
+ (e.g. `--models reviewer=haiku` still renders `model: haiku` followed by
250
+ `effort: high`). On Haiku 4.5, which does not support the `effort`
251
+ parameter, the harness ignores the pinned value rather than rejecting it
252
+ (anchored by a measurement, see CHANGELOG 0.23.0).
@@ -0,0 +1,113 @@
1
+ # Operator-level install
2
+
3
+ See the [package README](../README.md) for the single-repository `init` /
4
+ `uninstall` commands this layer sits alongside.
5
+
6
+ Alongside `init`, which installs the kit into one repository from that
7
+ repository's own working directory, an operator who maintains many
8
+ repositories can set defaults once and project them onto each target
9
+ instead of re-answering the same prompts per repo. This layer adds no new
10
+ binary: `setup`, `apply`, `doctor`, and `adopt` below are subcommands of the
11
+ same `orchestrator-workflow` CLI `init` and `uninstall` already ship as, and
12
+ `init`/`uninstall` remain fully supported and unchanged for a
13
+ single-repository install.
14
+
15
+ ```bash
16
+ orchestrator-workflow setup --yes
17
+ orchestrator-workflow apply --target /path/to/repo
18
+ ```
19
+
20
+ **`setup`** writes or updates this operator's default install options
21
+ (harnesses, profile, legacy models, routing, tiers) as the baseline for future installs; it
22
+ touches no repository. A flag always wins; a flag-less re-run keeps the
23
+ previously stored values; a first-ever `setup` falls back to `claude` /
24
+ `full` / the kit's default routing / tiers off. `setup` takes the same
25
+ option flags as `init` (`--harness`, `--profile`, `--models`, `--codex-models`, `--routing`,
26
+ `--codex-catalog`, `--tiers` / `--no-tiers`, `--opencode-provider`, `--yes`).
27
+ The defaults live in
28
+ `<operator home>/manifest.json`, where the operator home is
29
+ `~/.orchestrator-workflow/` unless the `ORCHESTRATOR_WORKFLOW_HOME`
30
+ environment variable names a different directory.
31
+
32
+ **`apply --target <repo>`** projects the operator's install onto a target
33
+ repository and registers that target, by its real resolved path, in the
34
+ operator manifest. It requires a prior `orchestrator-workflow setup`;
35
+ without one it exits `1` with "No operator setup found". Option resolution
36
+ follows one precedence order: an
37
+ explicit flag wins, then the target's own previously recorded settings,
38
+ then the operator's defaults (harnesses fall back one step further, to
39
+ what `init` would have auto-detected) -- except a target whose own
40
+ manifest recorded a real `harnesses: []` (a deliberate templates-only
41
+ install, see [Install reference](install-reference.md)'s "Templates-only
42
+ mode"), which stays templates-only on
43
+ a flagless run regardless of the operator's defaults or what is on disk;
44
+ an **interactive** re-run on such a target still prompts, with the same
45
+ nothing-pre-checked behaviour described there
46
+ (it applies identically to `apply`).
47
+ Pass `--sync` to invert that for
48
+ profile, tiers, legacy models, and routing: the operator's defaults then win over whatever
49
+ the target already had recorded. A target pinned to a kit version other
50
+ than the one being applied is skipped rather than touched (see the pin
51
+ rule below). `apply` also takes the same install options as `init` (`--harness`,
52
+ `--profile`, `--models`, `--codex-models`, `--routing`, `--codex-catalog`, `--tiers` /
53
+ `--no-tiers`, `--opencode-provider`, `--force`, `--yes`), which feed the
54
+ precedence rule above. An explicit routing file is the highest-precedence
55
+ deep patch; leaves it omits retain their resolved baseline values.
56
+
57
+ **`doctor [--json] [--prune]`** reports every operator-registered target's
58
+ status: `clean`, `divergent` (from the operator defaults, including routing), `version-lag`,
59
+ `drift` (installed files edited, deleted, or unreadable since install),
60
+ `missing`, `no-manifest`, or `unverifiable`. It exits `2` when the operator
61
+ manifest is missing or unreadable, or, with `--prune`, when the operator
62
+ manifest lock cannot be acquired or the rewrite fails; `1` when any target
63
+ is `drift`, `missing`, `no-manifest`, or `unverifiable`; and `0` otherwise.
64
+ `--json` prints one JSON object instead
65
+ of human output, with one entry per target plus the operator home and
66
+ version. `--prune` removes `missing` and `no-manifest` targets from the
67
+ registry before reporting (never an `unverifiable` one, since that status
68
+ means the check itself was inconclusive, not that the target is confirmed
69
+ gone) and rewrites the manifest file in its normalized form.
70
+ For a legacy opencode leaf without a recorded provider-qualified model id,
71
+ doctor reports `Routing comparison incomplete` and includes
72
+ `routingComparisonGaps` in JSON instead of declaring a false routing
73
+ divergence. The gap alone does not change the target status.
74
+
75
+ **`adopt [dir] [--json]`** brings a repository that already has the kit installed,
76
+ by hand or by an earlier `init`, under the operator's management without
77
+ changing anything in that repository: it registers the repository
78
+ verbatim, using the repository's own recorded settings to bootstrap the
79
+ operator manifest when none exists yet, records the repository's own
80
+ installed version as its baseline, and prints that one target's `doctor`
81
+ report. It exits `1` only when the freshly adopted target itself reports
82
+ drift, and `2` for a precondition failure (no repo manifest, an unreadable
83
+ or foreign manifest, or a lock or write failure).
84
+
85
+ **The kit-version pin.** A repository's own manifest can additionally
86
+ carry an optional `pin`: a kit version that `apply` must match before it
87
+ will touch that repository again. `apply` skips (exit `0`) a target pinned
88
+ to a different version than the one being applied. `--pin <version>` sets
89
+ or replaces the pin and applies regardless of any existing one; `--unpin`
90
+ clears it and applies; `--force-pin` advances an existing pin that
91
+ differs, but has no effect on a target with no pin recorded (it stays
92
+ unpinned). `doctor` reports `version-lag` when the installed version
93
+ differs from the running kit version; on a pinned target the pin is
94
+ compared against the installed version instead, so a pin equal to the
95
+ installed version is `clean` and a pin that no longer matches it is
96
+ `version-lag`.
97
+
98
+ **The registry is implicit**, not a separate command: `apply` and `adopt`
99
+ register a target as a side effect of a real run, and `doctor --prune` is
100
+ how a registry entry is removed again; there is no bare register or
101
+ unregister command. The workspace root of a multi-repo checkout is treated
102
+ as an ordinary target, nothing special.
103
+
104
+ All writes to the operator manifest, by `setup`, `apply`, `doctor --prune`,
105
+ and `adopt` alike, go through one advisory lock in the operator home, so
106
+ concurrent orchestrator-workflow commands on the same machine cannot
107
+ corrupt each other's state.
108
+
109
+ `apply` shares the file-ownership rules `init` uses; see the [package
110
+ README](../README.md)'s "Ownership and re-runs" section.
111
+
112
+ A sparse `--codex-models` file supplies only the internal Codex aliases it
113
+ intends to migrate and remains below an explicit routing leaf.
@@ -0,0 +1,48 @@
1
+ # Role profile reference
2
+
3
+ Detail behind the `--profile` table in the [package README](../README.md):
4
+ what the advisor role is for, and exactly what happens on a profile or tier
5
+ change across re-runs.
6
+
7
+ ## Advisor (escalation)
8
+
9
+ The fifth `full`-profile role, `advisor`, is read-only and consulted only
10
+ when the orchestrator hits one of a defined set of triggers: architectural
11
+ uncertainty, requirements that contradict each other, multiple valid solution
12
+ paths where committing to one is expensive to reverse, repeated
13
+ implementation failures on the same task, a review deadlock, or a high-risk
14
+ decision. It is not a standard pipeline step; like tier choice, spawning it
15
+ is the orchestrator's own judgment call. The advisor lays out the options
16
+ with their pros, cons, and risk, and gives a recommendation; it recommends,
17
+ never decides, and never writes code; the orchestrator still decides, and a
18
+ critical risk still goes to the operator. `minimal` never installs it, the
19
+ same as explorer and task-slicer.
20
+
21
+ ## Re-runs and profile changes
22
+
23
+ A plain re-run (no `--profile` flag) keeps the profile recorded in
24
+ `.ai/workflow/manifest.json` from the previous install, the same
25
+ override-vs-persist rule already used for `--harness` and `--models`.
26
+ Passing `--profile` explicitly always overrides the recorded value,
27
+ immediately switching which per-role files the next run installs and
28
+ updating the manifest to match. Switching profiles follows the same
29
+ precedent already in place for dropping a harness from `--harness` on a
30
+ re-run: files for roles no longer in the profile are simply no longer
31
+ installed or tracked in the manifest; they are not automatically deleted
32
+ from disk. `init` detects a `full` -> `minimal` downgrade and prints a note
33
+ naming the now-untracked `task-slicer.md` / `explorer.md` / `advisor.md`
34
+ agent files and how to remove them. For a fully clean switch, run
35
+ `orchestrator-workflow uninstall` first, or remove those files by hand.
36
+ Uninstalling a `minimal` install that has never been downgraded from `full`
37
+ is always clean on its own: it only ever removes what it actually installed,
38
+ so there is nothing to report as missing for the roles that were never
39
+ written. A `minimal` install reached via a `full` -> `minimal` downgrade is
40
+ not clean in that sense: the downgrade's now-untracked
41
+ `task-slicer.md` / `explorer.md` / `advisor.md` files are not in the
42
+ manifest's file ledger, so uninstall leaves them on disk without reporting
43
+ them at all.
44
+
45
+ The same override-vs-persist and downgrade-note rules apply to `--tiers`
46
+ (rendering the additional `<role>-<tier>.md` variant files); see [Model
47
+ routing reference: Effort tiers](model-routing-reference.md#effort-tiers)
48
+ for the full role/tier table.
@@ -0,0 +1,36 @@
1
+ # Run contracts: acceptance baseline and decision authority
2
+
3
+ Reference for the `00-goal.md` acceptance-baseline contract and the
4
+ `03-decisions.md` decision-authority record that the installed run templates
5
+ carry. See the [package README](../README.md) for the rest of the run-state
6
+ layout.
7
+
8
+ ## Acceptance-baseline adoption
9
+
10
+ New runs that need a frozen acceptance contract explicitly record
11
+ `Acceptance contract: acceptance-baseline/v1` in `00-goal.md` before planning,
12
+ slicing, or delegation. The same file then carries the canonical
13
+ `acceptance_baseline` identity and full `acceptance_criteria` records; each
14
+ delegated task receives its relevant records unchanged. Existing runs remain
15
+ under their recorded contract: missing v1 fields neither trigger migration nor
16
+ license a guess about a run's provenance. Communicate the recorded selection
17
+ in delegation and resolve unknown provenance before dependent work. A recorded
18
+ original string-list contract keeps its original criterion strings and omits
19
+ only the added baseline and criterion-evidence fields.
20
+
21
+ For v1, implementers return `acceptance_baseline: { id, revision }` and one
22
+ `criterion_evidence` entry per assigned criterion, with `criterion_id` and
23
+ `evidence_refs`. References resolve from the owning run directory and point
24
+ to producer artifacts with the checked state and result metadata.
25
+ `04-implementation-summary.md` indexes those references; empty references
26
+ remain unresolved, and required unresolved criteria block acceptance. Manual
27
+ evidence stays explicitly manual. Review findings and orchestrator acceptance
28
+ remain separate from this coverage index.
29
+
30
+ ## Decision authority
31
+
32
+ `03-decisions.md` records decisions with an ID, trigger/evidence, decision,
33
+ accountable authority/source, consequences, and an optional superseded
34
+ decision. It documents real approval evidence; it does not grant authority.
35
+ A reviewer recommendation does not equal orchestrator acceptance, and only
36
+ the operator may authorize a critical waiver.
@@ -0,0 +1,57 @@
1
+ # `validate-review-report` CLI reference
2
+
3
+ See the [package README](../README.md) for the rest of the CLI surface.
4
+
5
+ ```bash
6
+ orchestrator-workflow validate-review-report path/to/return.yaml
7
+ orchestrator-workflow validate-review-report - < path/to/return.yaml
8
+ orchestrator-workflow validate-review-report path/to/return.yaml --format json
9
+ ```
10
+
11
+ Checks a reviewer return's YAML against the reviewer output contract's
12
+ required fields and enums (see the "Reviewer output contract" section of
13
+ `assets/skill/references/contracts.md`, byte-identical to the contract in
14
+ `assets/agents/reviewer.md`), whether the return is fenced in a code
15
+ block (any language tag, or none) or given unfenced, and prints one
16
+ diagnostic per missing or invalid field. Every element of a string-array
17
+ field (`summary`, `missing_tests`, `residual_risks`) must itself be a
18
+ string; a non-string element (a number, a mapping, a boolean, or `null`
19
+ (written as a bare or `~` bullet)) is its own diagnostic at
20
+ `<field>[<index>]`. A fenced return ends at the first closing fence that
21
+ starts at column 0, repeats at least as many backticks as the opening
22
+ fence, and carries nothing but whitespace after that run, so neither a
23
+ reviewer quoting a fenced snippet inside a value (a `description` block
24
+ scalar, which YAML indents) nor one wrapping a return in four backticks
25
+ around a snippet fenced at column 0 truncates the return. A return
26
+ with no closing fence satisfying all three is not fenced at all, so its
27
+ whole text reaches the parser; that includes one whose opener is longer
28
+ than every closing run present. When the return carries more than one
29
+ fenced block, the first one whose fence tag's first word is `yaml` or
30
+ `yml` is validated, case-insensitively and counting whitespace-separated
31
+ attributes (`yaml title=x` counts; `yaml,title=x` does not, its first
32
+ word being the whole string), falling back to the first fence only when
33
+ none carries that word; a warning names any earlier fence skipped this
34
+ way. This preference can validate a later worked example instead of an
35
+ earlier, real but unfenced return: a reviewer who leaves their own return
36
+ unfenced and then quotes a `yaml`-tagged example afterward has that
37
+ example validated instead, which the emitted warning also names.
38
+ `--format json` prints the same diagnostics as a single JSON object
39
+ instead of human-readable text. It exits `0` when the return is
40
+ structurally valid, `1` when it is structurally invalid (a required field
41
+ is missing or its value falls outside its enum, or the input is
42
+ unparsable, empty, or not a mapping), and `2` for a usage error (an
43
+ unreadable file, an unrecognized `--format` value, a missing `<file>`
44
+ argument, an unknown option, or an excess positional argument).
45
+ `--format json` governs the validation verdict only: a commander parsing
46
+ error (missing argument, unknown option, excess arguments) or an
47
+ unrecognized `--format` value itself still prints plain text to stderr
48
+ with nothing on stdout, regardless of `--format`; the one exception is an
49
+ unreadable file, which does emit the JSON envelope on stdout. This check
50
+ is structural only: it never judges semantic adequacy, cannot waive a
51
+ finding, and passing it is never orchestrator acceptance. The
52
+ required-field set it checks is hand-maintained in `src/review-report.ts`
53
+ and pinned against the contract block itself by
54
+ `test/docs-consistency.test.ts`, so a contract edit without a matching
55
+ schema edit fails the suite instead of drifting silently; every field
56
+ listed there is dispatched to its own checker, so an entry added to the
57
+ list without a checker fails to typecheck rather than passing unchecked.
@@ -0,0 +1,53 @@
1
+ # Verification sets
2
+
3
+ See the [package README](../README.md) for how this file fits into an
4
+ implementer or reviewer briefing.
5
+
6
+ A repository may check in `.ai/workflow/verify.json` to name the complete
7
+ verification set for an implementer or reviewer briefing. This generic worked
8
+ example uses a `preflight run <repo> --json` executor and ordered extras with
9
+ `cwd`, `argv`, and an explicit before/after-preflight phase, so an approved
10
+ build can precede a dependent test:
11
+
12
+ ```json
13
+ {
14
+ "format": "orchestrator-workflow-verification-set/v1",
15
+ "preflight": {
16
+ "kind": "preflight",
17
+ "name": "preflight",
18
+ "cwd": ".",
19
+ "argv": ["preflight", "run", ".", "--json"]
20
+ },
21
+ "extras": [
22
+ {
23
+ "kind": "command",
24
+ "name": "build",
25
+ "phase": "before_preflight",
26
+ "cwd": "packages/example",
27
+ "argv": ["npm", "run", "build"]
28
+ },
29
+ {
30
+ "kind": "command",
31
+ "name": "package-tests",
32
+ "phase": "after_preflight",
33
+ "cwd": "packages/example",
34
+ "argv": ["npm", "test"]
35
+ },
36
+ {
37
+ "kind": "bundlecheck",
38
+ "name": "knowledge-bundle",
39
+ "phase": "after_preflight",
40
+ "cwd": "packages/example",
41
+ "argv": ["npx", "okf-kit", "check", "docs/okf"]
42
+ }
43
+ ]
44
+ }
45
+ ```
46
+
47
+ The workflow does not execute or validate this file: the orchestrator first
48
+ approves the resolved effective config and scripts, then records a run-local
49
+ snapshot with the set digest, repository identity, executable identity, and
50
+ every result. Preflight JSON reports check results, not the underlying shell
51
+ commands it discovered. A repository with a configured knowledge bundle
52
+ (`knowledge` in `.ai/workflow/manifest.json`; default `docs/okf/`) includes
53
+ its bundle check in every set, even when the task did not edit documentation.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "orchestrator-workflow",
3
- "version": "0.41.0",
3
+ "version": "0.43.1",
4
4
  "description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
5
5
  "main": "dist/index.js",
6
6
  "type": "module",
@@ -12,7 +12,8 @@
12
12
  "assets",
13
13
  "README.md",
14
14
  "INSTALL-AGENT.md",
15
- "CHANGELOG.md"
15
+ "CHANGELOG.md",
16
+ "docs/*.md"
16
17
  ],
17
18
  "scripts": {
18
19
  "build": "tsc",