orchestrator-workflow 0.41.0 → 0.43.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +135 -0
- package/INSTALL-AGENT.md +6 -1
- package/README.md +132 -645
- package/assets/agents/task-slicer.md +11 -1
- package/assets/codex-models.json +5 -0
- package/assets/skill/SKILL.md +3 -2
- package/assets/skill/references/bundle-gate-in-ci.md +60 -11
- package/assets/skill/references/contracts.md +15 -1
- package/assets/skill/references/evidence-and-probes.md +52 -8
- package/assets/skill/references/review-and-recovery.md +5 -0
- package/dist/cli.js +12 -4
- package/dist/index.d.ts +2 -2
- package/dist/index.js +1 -1
- package/dist/init.d.ts +11 -3
- package/dist/init.js +56 -13
- package/dist/routing.d.ts +28 -62
- package/dist/routing.js +86 -17
- package/docs/architecture.md +44 -0
- package/docs/harnesses.md +38 -0
- package/docs/install-reference.md +95 -0
- package/docs/model-routing-reference.md +252 -0
- package/docs/operator-install.md +113 -0
- package/docs/role-profile-reference.md +48 -0
- package/docs/run-contracts.md +36 -0
- package/docs/validate-review-report.md +57 -0
- package/docs/verification-sets.md +53 -0
- package/package.json +3 -2
|
@@ -0,0 +1,252 @@
|
|
|
1
|
+
# Model routing reference
|
|
2
|
+
|
|
3
|
+
The default per-role model table, the full `--routing` JSON shape, the Codex
|
|
4
|
+
default routing table, opencode model resolution, and the effort-tiers
|
|
5
|
+
mechanics behind the [package README](../README.md)'s "Model preselection"
|
|
6
|
+
section.
|
|
7
|
+
|
|
8
|
+
## Default model routing
|
|
9
|
+
|
|
10
|
+
`--models` is the backward-compatible, per-role input for Claude Code and
|
|
11
|
+
opencode. It does not configure Codex. Existing manifests that contain only
|
|
12
|
+
`models` continue to produce the same Claude Code and opencode defaults:
|
|
13
|
+
|
|
14
|
+
| Role | Default | Why |
|
|
15
|
+
|---|---|---|
|
|
16
|
+
| explorer | `sonnet` | read-only terrain mapping is broad reading, not deep reasoning |
|
|
17
|
+
| task-slicer | `sonnet` | structured decomposition, no deep reasoning needed |
|
|
18
|
+
| implementer | `sonnet` | fast, cheap, good enough for narrow pre-sliced tasks |
|
|
19
|
+
| reviewer | `opus` | skeptical review benefits from the strongest model |
|
|
20
|
+
| advisor | `opus` | escalations happen precisely when the situation is hard, so it shares the reviewer's strongest-model default |
|
|
21
|
+
|
|
22
|
+
The orchestrator itself runs on the session's main model. For Codex, start the
|
|
23
|
+
orchestrator on `gpt-6-astra` at `high` effort; use `xhigh` for demanding work.
|
|
24
|
+
The installer does not mutate global or fleet Codex configuration to enforce
|
|
25
|
+
that recommendation.
|
|
26
|
+
|
|
27
|
+
## Routing JSON shape
|
|
28
|
+
|
|
29
|
+
Routing is a harness-specific map from role and tier to a complete
|
|
30
|
+
`{model, effort}` selection. Pass a JSON file with `--routing`; the CLI deep
|
|
31
|
+
merges only the leaves you provide and records the resulting effective map in
|
|
32
|
+
`.ai/workflow/manifest.json`. The role's default-tier key configures its
|
|
33
|
+
unsuffixed file; another allowed key configures the corresponding
|
|
34
|
+
`<role>-<tier>` variant when `--tiers` is enabled. For example:
|
|
35
|
+
|
|
36
|
+
```json
|
|
37
|
+
{
|
|
38
|
+
"codex": {
|
|
39
|
+
"implementer": {
|
|
40
|
+
"medium": { "model": "gpt-6.1-sol", "effort": "medium" },
|
|
41
|
+
"xhigh": { "model": "gpt-6-astra", "effort": "xhigh" }
|
|
42
|
+
}
|
|
43
|
+
}
|
|
44
|
+
}
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
An omitted `--routing` preserves the exact persisted map on a re-install.
|
|
48
|
+
Changing one leaf leaves the others intact, which makes a previous manifest a
|
|
49
|
+
usable rollback record. Model updates are deliberate per role and tier: the
|
|
50
|
+
installer never interprets a newer model as automatically better and never
|
|
51
|
+
rewrites a preserved choice merely because another model exists.
|
|
52
|
+
|
|
53
|
+
## Codex defaults
|
|
54
|
+
|
|
55
|
+
Codex uses native `.codex/agents/*.toml` custom agents. The file shape
|
|
56
|
+
follows the
|
|
57
|
+
[official Codex subagent configuration](https://learn.chatgpt.com/docs/agent-configuration/subagents).
|
|
58
|
+
The shipped routing is:
|
|
59
|
+
|
|
60
|
+
| Role | Tier | Model | Effort |
|
|
61
|
+
|---|---|---|---|
|
|
62
|
+
| explorer | low | `gpt-6-luna` | low |
|
|
63
|
+
| explorer | medium (default) | `gpt-6.1-sol` | medium |
|
|
64
|
+
| explorer | high | `gpt-6.1-sol` | high |
|
|
65
|
+
| task-slicer | low | `gpt-6.1-sol` | low |
|
|
66
|
+
| task-slicer | medium (default) | `gpt-6.1-sol` | medium |
|
|
67
|
+
| task-slicer | high | `gpt-6.1-sol` | high |
|
|
68
|
+
| implementer | low | `gpt-6-luna` | low |
|
|
69
|
+
| implementer | medium (default) | `gpt-6.1-sol` | medium |
|
|
70
|
+
| implementer | high | `gpt-6.1-sol` | high |
|
|
71
|
+
| implementer | xhigh | `gpt-6-astra` | xhigh |
|
|
72
|
+
| reviewer | medium | `gpt-6.1-sol` | medium |
|
|
73
|
+
| reviewer | high (default) | `gpt-6-astra` | high |
|
|
74
|
+
| reviewer | xhigh | `gpt-6-astra` | xhigh |
|
|
75
|
+
| advisor | high (default) | `gpt-6-astra` | high |
|
|
76
|
+
| advisor | xhigh | `gpt-6-astra` | xhigh |
|
|
77
|
+
|
|
78
|
+
When you have a deterministic Codex model catalog, pass it with
|
|
79
|
+
`--codex-catalog <json-file>`. The CLI validates the selected Codex model and
|
|
80
|
+
effort pairs before writing. Without a supplied catalog it performs no online
|
|
81
|
+
entitlement check; offline or account-specific availability remains unknown.
|
|
82
|
+
Use the harness's native capability commands, such as `codex debug models`, to
|
|
83
|
+
refresh a catalog before installation when appropriate. A bundled-capability
|
|
84
|
+
view describes what the binary knows and does not prove account entitlement.
|
|
85
|
+
|
|
86
|
+
The bundled `assets/codex-models.json` maps the internal aliases `small`,
|
|
87
|
+
`balanced`, and `strong` to the shipped concrete IDs. These aliases are
|
|
88
|
+
installer inputs only: native agent files and manifests always record the
|
|
89
|
+
resolved concrete ID. To migrate only selected classes, pass a sparse JSON
|
|
90
|
+
file through `--codex-models <json-file>`, for example
|
|
91
|
+
`{"balanced":"gpt-example"}`. It creates complete routing leaves with each
|
|
92
|
+
matching tier's standard effort; it does not alter omitted aliases or a plain
|
|
93
|
+
reinstall. An explicit `--routing` leaf is applied afterward and has highest
|
|
94
|
+
precedence, including when a custom effort must be retained.
|
|
95
|
+
|
|
96
|
+
## opencode model resolution
|
|
97
|
+
|
|
98
|
+
opencode requires fully-qualified `provider/model-id`
|
|
99
|
+
strings (e.g. `github-copilot/claude-sonnet-4.6`). At install time the CLI
|
|
100
|
+
runs `opencode models` to fetch the live catalog and auto-detects which
|
|
101
|
+
provider to use (the one that offers Claude models). When exactly one such
|
|
102
|
+
provider exists the aliases are resolved to the highest-version matching id in
|
|
103
|
+
the catalog. When multiple providers offer Claude models the CLI warns and asks
|
|
104
|
+
you to pass `--opencode-provider <id>` to disambiguate, or to supply
|
|
105
|
+
fully-qualified ids per role via `--models`. If no resolution is possible
|
|
106
|
+
(catalog empty, `opencode` binary absent, ambiguous provider) the `model:`
|
|
107
|
+
frontmatter line is omitted entirely and the subagent inherits the
|
|
108
|
+
session/default model, a safe, portable fallback. Fully-qualified ids in
|
|
109
|
+
`--models` always pass through unchanged regardless of the catalog.
|
|
110
|
+
Nested-path providers like `openrouter` (whose ids look like
|
|
111
|
+
`openrouter/anthropic/claude-...`) are not auto-resolved from aliases and must
|
|
112
|
+
be supplied as a fully-qualified `--models` entry, e.g.
|
|
113
|
+
`reviewer=openrouter/anthropic/claude-opus-4.8`.
|
|
114
|
+
|
|
115
|
+
## Effort tiers
|
|
116
|
+
|
|
117
|
+
`--tiers` renders an additional per-role subagent definition for each
|
|
118
|
+
non-default effort tier, alongside the one default (unsuffixed) agent file
|
|
119
|
+
`--profile` already installs. Each tier variant is a standalone subagent
|
|
120
|
+
definition, not a modification of the default file. Claude Code and opencode
|
|
121
|
+
use `<role>.md` / `<role>-<tier>.md`; Codex uses `<role>.toml` /
|
|
122
|
+
`<role>-<tier>.toml`.
|
|
123
|
+
|
|
124
|
+
**Every default file carries its own pinned effort, independent of
|
|
125
|
+
`--tiers`.** The harness composers add the role's own default routing
|
|
126
|
+
selection to the unsuffixed file. In the legacy Claude/opencode path this is
|
|
127
|
+
`TIER_DEFS[DEFAULT_TIER[role]].effort`: `effort: medium` for explorer,
|
|
128
|
+
task-slicer, and implementer; `effort: high` for reviewer and advisor
|
|
129
|
+
(opencode: a `variant: high` line when the resolved model is Claude-family,
|
|
130
|
+
following the same dispatch rule tier variants use, `reasoningEffort:
|
|
131
|
+
medium`/`reasoningEffort: high` for a non-Claude-family provider-qualified
|
|
132
|
+
model, nothing for Ollama, a provider-less id, or an unresolved model). This
|
|
133
|
+
pin does not depend on `tiers`, so a plain install (no `--tiers`) already
|
|
134
|
+
carries it; the flag only controls whether the additional `<role>-<tier>.md`
|
|
135
|
+
variant files are also rendered. The motivation: a default spawn used to
|
|
136
|
+
silently inherit the orchestrator session's own effort, so a `high`-effort
|
|
137
|
+
orchestrator session made every default subagent spawn at `high` too,
|
|
138
|
+
regardless of the role's own intended weight; the pin makes each role's
|
|
139
|
+
effort deterministic and independent of the caller's session. A `--tiers`-off
|
|
140
|
+
install (the default) has no variant files and therefore no in-install
|
|
141
|
+
escalation path off a default's pinned effort; run `init --tiers` afterward
|
|
142
|
+
if a task ever needs one.
|
|
143
|
+
|
|
144
|
+
Default off, like every optional pack in this kit: a fresh install renders
|
|
145
|
+
no variant files unless asked. `--tiers` turns the feature on for that run,
|
|
146
|
+
`--no-tiers` turns it off; a plain re-run with neither flag keeps whatever
|
|
147
|
+
the previous install had, the same override-vs-persist rule already used
|
|
148
|
+
for `--profile` and `--models`. There is no interactive prompt for it:
|
|
149
|
+
`tiers` is opt-in/off via the flags only. Neither Codex nor the other harnesses
|
|
150
|
+
get `max` or `ultra` variants from this kit.
|
|
151
|
+
|
|
152
|
+
```bash
|
|
153
|
+
npx orchestrator-workflow init --tiers --yes
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
Turning tiers back off with `--no-tiers` after having them on follows the
|
|
157
|
+
same pattern as a `full` to `minimal` profile downgrade: `init` prints a note
|
|
158
|
+
naming the now-untracked `<role>-<tier>.md` variant files and how to remove
|
|
159
|
+
them, rather than deleting them or leaving the leftover unexplained.
|
|
160
|
+
|
|
161
|
+
**Which tiers each role gets.** A role never gets a variant file for its own
|
|
162
|
+
default tier: that would collide with, and duplicate, the default file.
|
|
163
|
+
|
|
164
|
+
| Role | Tiers available | Default tier (no variant file) |
|
|
165
|
+
|---|---|---|
|
|
166
|
+
| explorer | low, medium, high | medium |
|
|
167
|
+
| task-slicer | low, medium, high | medium |
|
|
168
|
+
| implementer | low, medium, high, xhigh | medium |
|
|
169
|
+
| reviewer | medium, high, xhigh | high |
|
|
170
|
+
| advisor | high, xhigh | high |
|
|
171
|
+
|
|
172
|
+
With `--profile full` and every tier rendered, that is 5 default files plus
|
|
173
|
+
10 variant files: 15 files total per harness.
|
|
174
|
+
|
|
175
|
+
**Tier to model class to effort.** Each tier resolves to a model class and an
|
|
176
|
+
effort value:
|
|
177
|
+
|
|
178
|
+
| Tier | Model class | Model alias | Effort requested |
|
|
179
|
+
|---|---|---|---|
|
|
180
|
+
| low | small | `haiku` | `low` |
|
|
181
|
+
| medium | medium | `sonnet` | `medium` |
|
|
182
|
+
| high | medium | `sonnet` | `high` |
|
|
183
|
+
| xhigh | large | `opus` | `xhigh` |
|
|
184
|
+
|
|
185
|
+
Claude Code variants carry both a `model:` line (the class's alias) and an
|
|
186
|
+
`effort: <tier>` line in frontmatter. Read-only roles (explorer, reviewer,
|
|
187
|
+
advisor) keep `disallowedTools: Edit, Write, NotebookEdit` on their variants
|
|
188
|
+
too.
|
|
189
|
+
|
|
190
|
+
**opencode variants key off the resolved model's family, not its provider
|
|
191
|
+
prefix**, since opencode's effort surface is not uniform across model
|
|
192
|
+
families:
|
|
193
|
+
|
|
194
|
+
- **Claude-family models** (any resolved id whose provider is
|
|
195
|
+
`anthropic/`, or whose segment after the provider prefix contains
|
|
196
|
+
`claude-`, which covers `anthropic/claude-...` as well as a Claude model
|
|
197
|
+
fronted by a different provider, e.g. `github-copilot/claude-sonnet-4.6`
|
|
198
|
+
or the nested `openrouter/anthropic/claude-opus-4.8`): only `high` and
|
|
199
|
+
`xhigh` get an effort field, as `variant: high` and `variant: max`
|
|
200
|
+
respectively; `low` and `medium` collapse to no effort field at all,
|
|
201
|
+
since opencode's `variant:` option does not distinguish an effort below
|
|
202
|
+
`high`. This collapse is deliberate and documented, not a bug: a
|
|
203
|
+
`low`/`medium` variant on a Claude-family model still gets its class's
|
|
204
|
+
`model:` line, just no `variant:` line.
|
|
205
|
+
- **Ollama, or an id with no provider prefix**: no effort field at all.
|
|
206
|
+
There is no known effort passthrough for Ollama, and an id with no `/`
|
|
207
|
+
resolves to no provider to key the decision on.
|
|
208
|
+
- **Every other non-Claude-family model**: a plain `reasoningEffort: <tier>`
|
|
209
|
+
line, `xhigh` included (opencode's built-in OpenAI-style variants
|
|
210
|
+
document an `xhigh` reasoning effort).
|
|
211
|
+
|
|
212
|
+
The variant's `model:` line is resolved the same way the base per-role model
|
|
213
|
+
is (an `opencode models` catalog lookup against the auto-detected or
|
|
214
|
+
`--opencode-provider`-specified provider), just keyed by the tier's model
|
|
215
|
+
class instead of by role. When that lookup cannot resolve a model for a
|
|
216
|
+
class, the CLI warns once on stderr and **no variant file is rendered for
|
|
217
|
+
that class at all**, not a file with a missing `model:` line: a variant
|
|
218
|
+
with no resolved model would carry neither a `model:` nor an effort line, an
|
|
219
|
+
indistinguishable no-op duplicate of the base file with no ledger entry to
|
|
220
|
+
compare it against, so `init` skips writing it entirely. This guard and its
|
|
221
|
+
warning are opencode-scoped only; Claude Code variants resolve `model:` from
|
|
222
|
+
a plain alias (`haiku`/`sonnet`/`opus`) and need no live catalog lookup, so
|
|
223
|
+
they are unaffected.
|
|
224
|
+
|
|
225
|
+
Codex variants carry `model` and `model_reasoning_effort` from their exact
|
|
226
|
+
routing leaf. The canonical role prompt becomes `developer_instructions`.
|
|
227
|
+
Runtime dispatch follows the client's actual capabilities: select the named
|
|
228
|
+
installed agent when supported; otherwise, if spawning supports explicit model
|
|
229
|
+
and effort, read the installed TOML and pass its selection, developer
|
|
230
|
+
instructions, and narrow task contract into a fresh task-local spawn. A
|
|
231
|
+
full-history spawn may not permit a model override. If that explicit spawn
|
|
232
|
+
cannot accept a sandbox override, explorer and advisor inherit the caller's
|
|
233
|
+
sandbox and their prompt is the edit guard. When native spawning is
|
|
234
|
+
unavailable, run the same contract inline and sequentially. The orchestrator
|
|
235
|
+
alone spawns agents. In particular, it must not choose `implementer-low` when
|
|
236
|
+
the task requires a test, typecheck, lint, build, or named mutation probe.
|
|
237
|
+
|
|
238
|
+
**Warning: `CLAUDE_CODE_EFFORT_LEVEL` overrides every agent's frontmatter
|
|
239
|
+
`effort:`, tier variants included.** Claude Code's `effort:` frontmatter
|
|
240
|
+
field does work: it reaches the model request as `output_config.effort`.
|
|
241
|
+
But when the harness environment sets `CLAUDE_CODE_EFFORT_LEVEL`, that
|
|
242
|
+
environment variable wins over the frontmatter `effort:` on every installed
|
|
243
|
+
agent, tier variants and default files alike, not just the one this feature
|
|
244
|
+
adds. Check for it before relying on a specific tier variant's requested
|
|
245
|
+
effort actually taking effect.
|
|
246
|
+
|
|
247
|
+
The pin is also emitted unconditionally regardless of which model the role
|
|
248
|
+
resolves to via `--models`, including a model with no effort support at all
|
|
249
|
+
(e.g. `--models reviewer=haiku` still renders `model: haiku` followed by
|
|
250
|
+
`effort: high`). On Haiku 4.5, which does not support the `effort`
|
|
251
|
+
parameter, the harness ignores the pinned value rather than rejecting it
|
|
252
|
+
(anchored by a measurement, see CHANGELOG 0.23.0).
|
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
# Operator-level install
|
|
2
|
+
|
|
3
|
+
See the [package README](../README.md) for the single-repository `init` /
|
|
4
|
+
`uninstall` commands this layer sits alongside.
|
|
5
|
+
|
|
6
|
+
Alongside `init`, which installs the kit into one repository from that
|
|
7
|
+
repository's own working directory, an operator who maintains many
|
|
8
|
+
repositories can set defaults once and project them onto each target
|
|
9
|
+
instead of re-answering the same prompts per repo. This layer adds no new
|
|
10
|
+
binary: `setup`, `apply`, `doctor`, and `adopt` below are subcommands of the
|
|
11
|
+
same `orchestrator-workflow` CLI `init` and `uninstall` already ship as, and
|
|
12
|
+
`init`/`uninstall` remain fully supported and unchanged for a
|
|
13
|
+
single-repository install.
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
orchestrator-workflow setup --yes
|
|
17
|
+
orchestrator-workflow apply --target /path/to/repo
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
**`setup`** writes or updates this operator's default install options
|
|
21
|
+
(harnesses, profile, legacy models, routing, tiers) as the baseline for future installs; it
|
|
22
|
+
touches no repository. A flag always wins; a flag-less re-run keeps the
|
|
23
|
+
previously stored values; a first-ever `setup` falls back to `claude` /
|
|
24
|
+
`full` / the kit's default routing / tiers off. `setup` takes the same
|
|
25
|
+
option flags as `init` (`--harness`, `--profile`, `--models`, `--codex-models`, `--routing`,
|
|
26
|
+
`--codex-catalog`, `--tiers` / `--no-tiers`, `--opencode-provider`, `--yes`).
|
|
27
|
+
The defaults live in
|
|
28
|
+
`<operator home>/manifest.json`, where the operator home is
|
|
29
|
+
`~/.orchestrator-workflow/` unless the `ORCHESTRATOR_WORKFLOW_HOME`
|
|
30
|
+
environment variable names a different directory.
|
|
31
|
+
|
|
32
|
+
**`apply --target <repo>`** projects the operator's install onto a target
|
|
33
|
+
repository and registers that target, by its real resolved path, in the
|
|
34
|
+
operator manifest. It requires a prior `orchestrator-workflow setup`;
|
|
35
|
+
without one it exits `1` with "No operator setup found". Option resolution
|
|
36
|
+
follows one precedence order: an
|
|
37
|
+
explicit flag wins, then the target's own previously recorded settings,
|
|
38
|
+
then the operator's defaults (harnesses fall back one step further, to
|
|
39
|
+
what `init` would have auto-detected) -- except a target whose own
|
|
40
|
+
manifest recorded a real `harnesses: []` (a deliberate templates-only
|
|
41
|
+
install, see [Install reference](install-reference.md)'s "Templates-only
|
|
42
|
+
mode"), which stays templates-only on
|
|
43
|
+
a flagless run regardless of the operator's defaults or what is on disk;
|
|
44
|
+
an **interactive** re-run on such a target still prompts, with the same
|
|
45
|
+
nothing-pre-checked behaviour described there
|
|
46
|
+
(it applies identically to `apply`).
|
|
47
|
+
Pass `--sync` to invert that for
|
|
48
|
+
profile, tiers, legacy models, and routing: the operator's defaults then win over whatever
|
|
49
|
+
the target already had recorded. A target pinned to a kit version other
|
|
50
|
+
than the one being applied is skipped rather than touched (see the pin
|
|
51
|
+
rule below). `apply` also takes the same install options as `init` (`--harness`,
|
|
52
|
+
`--profile`, `--models`, `--codex-models`, `--routing`, `--codex-catalog`, `--tiers` /
|
|
53
|
+
`--no-tiers`, `--opencode-provider`, `--force`, `--yes`), which feed the
|
|
54
|
+
precedence rule above. An explicit routing file is the highest-precedence
|
|
55
|
+
deep patch; leaves it omits retain their resolved baseline values.
|
|
56
|
+
|
|
57
|
+
**`doctor [--json] [--prune]`** reports every operator-registered target's
|
|
58
|
+
status: `clean`, `divergent` (from the operator defaults, including routing), `version-lag`,
|
|
59
|
+
`drift` (installed files edited, deleted, or unreadable since install),
|
|
60
|
+
`missing`, `no-manifest`, or `unverifiable`. It exits `2` when the operator
|
|
61
|
+
manifest is missing or unreadable, or, with `--prune`, when the operator
|
|
62
|
+
manifest lock cannot be acquired or the rewrite fails; `1` when any target
|
|
63
|
+
is `drift`, `missing`, `no-manifest`, or `unverifiable`; and `0` otherwise.
|
|
64
|
+
`--json` prints one JSON object instead
|
|
65
|
+
of human output, with one entry per target plus the operator home and
|
|
66
|
+
version. `--prune` removes `missing` and `no-manifest` targets from the
|
|
67
|
+
registry before reporting (never an `unverifiable` one, since that status
|
|
68
|
+
means the check itself was inconclusive, not that the target is confirmed
|
|
69
|
+
gone) and rewrites the manifest file in its normalized form.
|
|
70
|
+
For a legacy opencode leaf without a recorded provider-qualified model id,
|
|
71
|
+
doctor reports `Routing comparison incomplete` and includes
|
|
72
|
+
`routingComparisonGaps` in JSON instead of declaring a false routing
|
|
73
|
+
divergence. The gap alone does not change the target status.
|
|
74
|
+
|
|
75
|
+
**`adopt [dir] [--json]`** brings a repository that already has the kit installed,
|
|
76
|
+
by hand or by an earlier `init`, under the operator's management without
|
|
77
|
+
changing anything in that repository: it registers the repository
|
|
78
|
+
verbatim, using the repository's own recorded settings to bootstrap the
|
|
79
|
+
operator manifest when none exists yet, records the repository's own
|
|
80
|
+
installed version as its baseline, and prints that one target's `doctor`
|
|
81
|
+
report. It exits `1` only when the freshly adopted target itself reports
|
|
82
|
+
drift, and `2` for a precondition failure (no repo manifest, an unreadable
|
|
83
|
+
or foreign manifest, or a lock or write failure).
|
|
84
|
+
|
|
85
|
+
**The kit-version pin.** A repository's own manifest can additionally
|
|
86
|
+
carry an optional `pin`: a kit version that `apply` must match before it
|
|
87
|
+
will touch that repository again. `apply` skips (exit `0`) a target pinned
|
|
88
|
+
to a different version than the one being applied. `--pin <version>` sets
|
|
89
|
+
or replaces the pin and applies regardless of any existing one; `--unpin`
|
|
90
|
+
clears it and applies; `--force-pin` advances an existing pin that
|
|
91
|
+
differs, but has no effect on a target with no pin recorded (it stays
|
|
92
|
+
unpinned). `doctor` reports `version-lag` when the installed version
|
|
93
|
+
differs from the running kit version; on a pinned target the pin is
|
|
94
|
+
compared against the installed version instead, so a pin equal to the
|
|
95
|
+
installed version is `clean` and a pin that no longer matches it is
|
|
96
|
+
`version-lag`.
|
|
97
|
+
|
|
98
|
+
**The registry is implicit**, not a separate command: `apply` and `adopt`
|
|
99
|
+
register a target as a side effect of a real run, and `doctor --prune` is
|
|
100
|
+
how a registry entry is removed again; there is no bare register or
|
|
101
|
+
unregister command. The workspace root of a multi-repo checkout is treated
|
|
102
|
+
as an ordinary target, nothing special.
|
|
103
|
+
|
|
104
|
+
All writes to the operator manifest, by `setup`, `apply`, `doctor --prune`,
|
|
105
|
+
and `adopt` alike, go through one advisory lock in the operator home, so
|
|
106
|
+
concurrent orchestrator-workflow commands on the same machine cannot
|
|
107
|
+
corrupt each other's state.
|
|
108
|
+
|
|
109
|
+
`apply` shares the file-ownership rules `init` uses; see the [package
|
|
110
|
+
README](../README.md)'s "Ownership and re-runs" section.
|
|
111
|
+
|
|
112
|
+
A sparse `--codex-models` file supplies only the internal Codex aliases it
|
|
113
|
+
intends to migrate and remains below an explicit routing leaf.
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# Role profile reference
|
|
2
|
+
|
|
3
|
+
Detail behind the `--profile` table in the [package README](../README.md):
|
|
4
|
+
what the advisor role is for, and exactly what happens on a profile or tier
|
|
5
|
+
change across re-runs.
|
|
6
|
+
|
|
7
|
+
## Advisor (escalation)
|
|
8
|
+
|
|
9
|
+
The fifth `full`-profile role, `advisor`, is read-only and consulted only
|
|
10
|
+
when the orchestrator hits one of a defined set of triggers: architectural
|
|
11
|
+
uncertainty, requirements that contradict each other, multiple valid solution
|
|
12
|
+
paths where committing to one is expensive to reverse, repeated
|
|
13
|
+
implementation failures on the same task, a review deadlock, or a high-risk
|
|
14
|
+
decision. It is not a standard pipeline step; like tier choice, spawning it
|
|
15
|
+
is the orchestrator's own judgment call. The advisor lays out the options
|
|
16
|
+
with their pros, cons, and risk, and gives a recommendation; it recommends,
|
|
17
|
+
never decides, and never writes code; the orchestrator still decides, and a
|
|
18
|
+
critical risk still goes to the operator. `minimal` never installs it, the
|
|
19
|
+
same as explorer and task-slicer.
|
|
20
|
+
|
|
21
|
+
## Re-runs and profile changes
|
|
22
|
+
|
|
23
|
+
A plain re-run (no `--profile` flag) keeps the profile recorded in
|
|
24
|
+
`.ai/workflow/manifest.json` from the previous install, the same
|
|
25
|
+
override-vs-persist rule already used for `--harness` and `--models`.
|
|
26
|
+
Passing `--profile` explicitly always overrides the recorded value,
|
|
27
|
+
immediately switching which per-role files the next run installs and
|
|
28
|
+
updating the manifest to match. Switching profiles follows the same
|
|
29
|
+
precedent already in place for dropping a harness from `--harness` on a
|
|
30
|
+
re-run: files for roles no longer in the profile are simply no longer
|
|
31
|
+
installed or tracked in the manifest; they are not automatically deleted
|
|
32
|
+
from disk. `init` detects a `full` -> `minimal` downgrade and prints a note
|
|
33
|
+
naming the now-untracked `task-slicer.md` / `explorer.md` / `advisor.md`
|
|
34
|
+
agent files and how to remove them. For a fully clean switch, run
|
|
35
|
+
`orchestrator-workflow uninstall` first, or remove those files by hand.
|
|
36
|
+
Uninstalling a `minimal` install that has never been downgraded from `full`
|
|
37
|
+
is always clean on its own: it only ever removes what it actually installed,
|
|
38
|
+
so there is nothing to report as missing for the roles that were never
|
|
39
|
+
written. A `minimal` install reached via a `full` -> `minimal` downgrade is
|
|
40
|
+
not clean in that sense: the downgrade's now-untracked
|
|
41
|
+
`task-slicer.md` / `explorer.md` / `advisor.md` files are not in the
|
|
42
|
+
manifest's file ledger, so uninstall leaves them on disk without reporting
|
|
43
|
+
them at all.
|
|
44
|
+
|
|
45
|
+
The same override-vs-persist and downgrade-note rules apply to `--tiers`
|
|
46
|
+
(rendering the additional `<role>-<tier>.md` variant files); see [Model
|
|
47
|
+
routing reference: Effort tiers](model-routing-reference.md#effort-tiers)
|
|
48
|
+
for the full role/tier table.
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
# Run contracts: acceptance baseline and decision authority
|
|
2
|
+
|
|
3
|
+
Reference for the `00-goal.md` acceptance-baseline contract and the
|
|
4
|
+
`03-decisions.md` decision-authority record that the installed run templates
|
|
5
|
+
carry. See the [package README](../README.md) for the rest of the run-state
|
|
6
|
+
layout.
|
|
7
|
+
|
|
8
|
+
## Acceptance-baseline adoption
|
|
9
|
+
|
|
10
|
+
New runs that need a frozen acceptance contract explicitly record
|
|
11
|
+
`Acceptance contract: acceptance-baseline/v1` in `00-goal.md` before planning,
|
|
12
|
+
slicing, or delegation. The same file then carries the canonical
|
|
13
|
+
`acceptance_baseline` identity and full `acceptance_criteria` records; each
|
|
14
|
+
delegated task receives its relevant records unchanged. Existing runs remain
|
|
15
|
+
under their recorded contract: missing v1 fields neither trigger migration nor
|
|
16
|
+
license a guess about a run's provenance. Communicate the recorded selection
|
|
17
|
+
in delegation and resolve unknown provenance before dependent work. A recorded
|
|
18
|
+
original string-list contract keeps its original criterion strings and omits
|
|
19
|
+
only the added baseline and criterion-evidence fields.
|
|
20
|
+
|
|
21
|
+
For v1, implementers return `acceptance_baseline: { id, revision }` and one
|
|
22
|
+
`criterion_evidence` entry per assigned criterion, with `criterion_id` and
|
|
23
|
+
`evidence_refs`. References resolve from the owning run directory and point
|
|
24
|
+
to producer artifacts with the checked state and result metadata.
|
|
25
|
+
`04-implementation-summary.md` indexes those references; empty references
|
|
26
|
+
remain unresolved, and required unresolved criteria block acceptance. Manual
|
|
27
|
+
evidence stays explicitly manual. Review findings and orchestrator acceptance
|
|
28
|
+
remain separate from this coverage index.
|
|
29
|
+
|
|
30
|
+
## Decision authority
|
|
31
|
+
|
|
32
|
+
`03-decisions.md` records decisions with an ID, trigger/evidence, decision,
|
|
33
|
+
accountable authority/source, consequences, and an optional superseded
|
|
34
|
+
decision. It documents real approval evidence; it does not grant authority.
|
|
35
|
+
A reviewer recommendation does not equal orchestrator acceptance, and only
|
|
36
|
+
the operator may authorize a critical waiver.
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# `validate-review-report` CLI reference
|
|
2
|
+
|
|
3
|
+
See the [package README](../README.md) for the rest of the CLI surface.
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
orchestrator-workflow validate-review-report path/to/return.yaml
|
|
7
|
+
orchestrator-workflow validate-review-report - < path/to/return.yaml
|
|
8
|
+
orchestrator-workflow validate-review-report path/to/return.yaml --format json
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
Checks a reviewer return's YAML against the reviewer output contract's
|
|
12
|
+
required fields and enums (see the "Reviewer output contract" section of
|
|
13
|
+
`assets/skill/references/contracts.md`, byte-identical to the contract in
|
|
14
|
+
`assets/agents/reviewer.md`), whether the return is fenced in a code
|
|
15
|
+
block (any language tag, or none) or given unfenced, and prints one
|
|
16
|
+
diagnostic per missing or invalid field. Every element of a string-array
|
|
17
|
+
field (`summary`, `missing_tests`, `residual_risks`) must itself be a
|
|
18
|
+
string; a non-string element (a number, a mapping, a boolean, or `null`
|
|
19
|
+
(written as a bare or `~` bullet)) is its own diagnostic at
|
|
20
|
+
`<field>[<index>]`. A fenced return ends at the first closing fence that
|
|
21
|
+
starts at column 0, repeats at least as many backticks as the opening
|
|
22
|
+
fence, and carries nothing but whitespace after that run, so neither a
|
|
23
|
+
reviewer quoting a fenced snippet inside a value (a `description` block
|
|
24
|
+
scalar, which YAML indents) nor one wrapping a return in four backticks
|
|
25
|
+
around a snippet fenced at column 0 truncates the return. A return
|
|
26
|
+
with no closing fence satisfying all three is not fenced at all, so its
|
|
27
|
+
whole text reaches the parser; that includes one whose opener is longer
|
|
28
|
+
than every closing run present. When the return carries more than one
|
|
29
|
+
fenced block, the first one whose fence tag's first word is `yaml` or
|
|
30
|
+
`yml` is validated, case-insensitively and counting whitespace-separated
|
|
31
|
+
attributes (`yaml title=x` counts; `yaml,title=x` does not, its first
|
|
32
|
+
word being the whole string), falling back to the first fence only when
|
|
33
|
+
none carries that word; a warning names any earlier fence skipped this
|
|
34
|
+
way. This preference can validate a later worked example instead of an
|
|
35
|
+
earlier, real but unfenced return: a reviewer who leaves their own return
|
|
36
|
+
unfenced and then quotes a `yaml`-tagged example afterward has that
|
|
37
|
+
example validated instead, which the emitted warning also names.
|
|
38
|
+
`--format json` prints the same diagnostics as a single JSON object
|
|
39
|
+
instead of human-readable text. It exits `0` when the return is
|
|
40
|
+
structurally valid, `1` when it is structurally invalid (a required field
|
|
41
|
+
is missing or its value falls outside its enum, or the input is
|
|
42
|
+
unparsable, empty, or not a mapping), and `2` for a usage error (an
|
|
43
|
+
unreadable file, an unrecognized `--format` value, a missing `<file>`
|
|
44
|
+
argument, an unknown option, or an excess positional argument).
|
|
45
|
+
`--format json` governs the validation verdict only: a commander parsing
|
|
46
|
+
error (missing argument, unknown option, excess arguments) or an
|
|
47
|
+
unrecognized `--format` value itself still prints plain text to stderr
|
|
48
|
+
with nothing on stdout, regardless of `--format`; the one exception is an
|
|
49
|
+
unreadable file, which does emit the JSON envelope on stdout. This check
|
|
50
|
+
is structural only: it never judges semantic adequacy, cannot waive a
|
|
51
|
+
finding, and passing it is never orchestrator acceptance. The
|
|
52
|
+
required-field set it checks is hand-maintained in `src/review-report.ts`
|
|
53
|
+
and pinned against the contract block itself by
|
|
54
|
+
`test/docs-consistency.test.ts`, so a contract edit without a matching
|
|
55
|
+
schema edit fails the suite instead of drifting silently; every field
|
|
56
|
+
listed there is dispatched to its own checker, so an entry added to the
|
|
57
|
+
list without a checker fails to typecheck rather than passing unchecked.
|
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
# Verification sets
|
|
2
|
+
|
|
3
|
+
See the [package README](../README.md) for how this file fits into an
|
|
4
|
+
implementer or reviewer briefing.
|
|
5
|
+
|
|
6
|
+
A repository may check in `.ai/workflow/verify.json` to name the complete
|
|
7
|
+
verification set for an implementer or reviewer briefing. This generic worked
|
|
8
|
+
example uses a `preflight run <repo> --json` executor and ordered extras with
|
|
9
|
+
`cwd`, `argv`, and an explicit before/after-preflight phase, so an approved
|
|
10
|
+
build can precede a dependent test:
|
|
11
|
+
|
|
12
|
+
```json
|
|
13
|
+
{
|
|
14
|
+
"format": "orchestrator-workflow-verification-set/v1",
|
|
15
|
+
"preflight": {
|
|
16
|
+
"kind": "preflight",
|
|
17
|
+
"name": "preflight",
|
|
18
|
+
"cwd": ".",
|
|
19
|
+
"argv": ["preflight", "run", ".", "--json"]
|
|
20
|
+
},
|
|
21
|
+
"extras": [
|
|
22
|
+
{
|
|
23
|
+
"kind": "command",
|
|
24
|
+
"name": "build",
|
|
25
|
+
"phase": "before_preflight",
|
|
26
|
+
"cwd": "packages/example",
|
|
27
|
+
"argv": ["npm", "run", "build"]
|
|
28
|
+
},
|
|
29
|
+
{
|
|
30
|
+
"kind": "command",
|
|
31
|
+
"name": "package-tests",
|
|
32
|
+
"phase": "after_preflight",
|
|
33
|
+
"cwd": "packages/example",
|
|
34
|
+
"argv": ["npm", "test"]
|
|
35
|
+
},
|
|
36
|
+
{
|
|
37
|
+
"kind": "bundlecheck",
|
|
38
|
+
"name": "knowledge-bundle",
|
|
39
|
+
"phase": "after_preflight",
|
|
40
|
+
"cwd": "packages/example",
|
|
41
|
+
"argv": ["npx", "okf-kit", "check", "docs/okf"]
|
|
42
|
+
}
|
|
43
|
+
]
|
|
44
|
+
}
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
The workflow does not execute or validate this file: the orchestrator first
|
|
48
|
+
approves the resolved effective config and scripts, then records a run-local
|
|
49
|
+
snapshot with the set digest, repository identity, executable identity, and
|
|
50
|
+
every result. Preflight JSON reports check results, not the underlying shell
|
|
51
|
+
commands it discovered. A repository with a configured knowledge bundle
|
|
52
|
+
(`knowledge` in `.ai/workflow/manifest.json`; default `docs/okf/`) includes
|
|
53
|
+
its bundle check in every set, even when the task did not edit documentation.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "orchestrator-workflow",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.43.1",
|
|
4
4
|
"description": "Installer for an orchestrator-led agent workflow: .ai/ run state, an AGENTS.md policy section, and per-harness subagent definitions for Claude Code, OpenAI Codex, and opencode",
|
|
5
5
|
"main": "dist/index.js",
|
|
6
6
|
"type": "module",
|
|
@@ -12,7 +12,8 @@
|
|
|
12
12
|
"assets",
|
|
13
13
|
"README.md",
|
|
14
14
|
"INSTALL-AGENT.md",
|
|
15
|
-
"CHANGELOG.md"
|
|
15
|
+
"CHANGELOG.md",
|
|
16
|
+
"docs/*.md"
|
|
16
17
|
],
|
|
17
18
|
"scripts": {
|
|
18
19
|
"build": "tsc",
|