orchestrator-workflow 0.41.0 → 0.43.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +135 -0
- package/INSTALL-AGENT.md +6 -1
- package/README.md +132 -645
- package/assets/agents/task-slicer.md +11 -1
- package/assets/codex-models.json +5 -0
- package/assets/skill/SKILL.md +3 -2
- package/assets/skill/references/bundle-gate-in-ci.md +60 -11
- package/assets/skill/references/contracts.md +15 -1
- package/assets/skill/references/evidence-and-probes.md +52 -8
- package/assets/skill/references/review-and-recovery.md +5 -0
- package/dist/cli.js +12 -4
- package/dist/index.d.ts +2 -2
- package/dist/index.js +1 -1
- package/dist/init.d.ts +11 -3
- package/dist/init.js +56 -13
- package/dist/routing.d.ts +28 -62
- package/dist/routing.js +86 -17
- package/docs/architecture.md +44 -0
- package/docs/harnesses.md +38 -0
- package/docs/install-reference.md +95 -0
- package/docs/model-routing-reference.md +252 -0
- package/docs/operator-install.md +113 -0
- package/docs/role-profile-reference.md +48 -0
- package/docs/run-contracts.md +36 -0
- package/docs/validate-review-report.md +57 -0
- package/docs/verification-sets.md +53 -0
- package/package.json +3 -2
package/README.md
CHANGED
|
@@ -1,88 +1,33 @@
|
|
|
1
1
|
# orchestrator-workflow
|
|
2
2
|
|
|
3
3
|
Installs an orchestrator-led agent workflow into any repository: one `.ai/`
|
|
4
|
-
directory for run state, one marker-fenced policy section in `AGENTS.md`,
|
|
5
|
-
subagent definitions with preselected models for the harnesses
|
|
6
|
-
use (Claude Code, OpenAI Codex, opencode).
|
|
7
|
-
|
|
8
|
-
The
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
For v1, implementers return `acceptance_baseline: { id, revision }` and one
|
|
27
|
-
`criterion_evidence` entry per assigned criterion, with `criterion_id` and
|
|
28
|
-
`evidence_refs`. References resolve from the owning run directory and point
|
|
29
|
-
to producer artifacts with the checked state and result metadata.
|
|
30
|
-
`04-implementation-summary.md` indexes those references; empty references
|
|
31
|
-
remain unresolved, and required unresolved criteria block acceptance. Manual
|
|
32
|
-
evidence stays explicitly manual. Review findings and orchestrator acceptance
|
|
33
|
-
remain separate from this coverage index.
|
|
34
|
-
|
|
35
|
-
### Decision authority
|
|
36
|
-
|
|
37
|
-
`03-decisions.md` records decisions with an ID, trigger/evidence, decision,
|
|
38
|
-
accountable authority/source, consequences, and an optional superseded
|
|
39
|
-
decision. It documents real approval evidence; it does not grant authority.
|
|
40
|
-
A reviewer recommendation does not equal orchestrator acceptance, and only
|
|
41
|
-
the operator may authorize a critical waiver.
|
|
42
|
-
|
|
43
|
-
## Why this shape
|
|
44
|
-
|
|
45
|
-
```text
|
|
46
|
-
Operator
|
|
47
|
-
goal | ^ handoff: what changed, how verified,
|
|
48
|
-
v | what remains open
|
|
49
|
-
explorer --> Orchestrator . . . . . .ai/runs/<date>-<slug>/
|
|
50
|
-
optional, session model 00-goal 04-implementation-summary
|
|
51
|
-
read-only plans, validates slices, 01-plan 05-review-findings
|
|
52
|
-
terrain map decides acceptance 02-tasks 06-handoff
|
|
53
|
-
| 03-decisions
|
|
54
|
-
narrow | ^ structured (state lives in files,
|
|
55
|
-
contracts v | YAML evidence not in chat history)
|
|
56
|
-
+-------------+-------------+
|
|
57
|
-
| | |
|
|
58
|
-
task-slicer implementer reviewer
|
|
59
|
-
sonnet sonnet opus
|
|
60
|
-
small, one narrow skeptical, severity-rated
|
|
61
|
-
testable task, plus findings, no rewrites
|
|
62
|
-
slices tests
|
|
63
|
-
```
|
|
64
|
-
|
|
65
|
-
Two effects fall out of this shape:
|
|
66
|
-
|
|
67
|
-
- **Token efficiency.** The orchestrator's context stays small: subagents
|
|
68
|
-
receive narrow task contracts instead of the whole conversation, return
|
|
69
|
-
structured YAML evidence instead of transcripts, and durable state lives
|
|
70
|
-
in run files that survive context compaction. The cheap models do the
|
|
71
|
-
volume work; the strongest model is spent only on orchestration decisions
|
|
72
|
-
and the skeptical review. The ceremony scales to the task: a trivial change
|
|
73
|
-
is done directly, the full flow is for non-trivial work, and a read-only
|
|
74
|
-
explorer maps the terrain first only when the solution is unclear. When
|
|
75
|
-
available, the explorer prefers each of a repo's configured knowledge
|
|
76
|
-
bundles (`knowledge` in `.ai/workflow/manifest.json`; default `docs/okf/`)
|
|
77
|
-
or a connected semantic code-search tool over hand-mapping terrain with
|
|
78
|
-
grep.
|
|
79
|
-
- **Quality through structure.** Writing and reviewing are separated by
|
|
80
|
-
role and model, task slices are validated before any implementation
|
|
81
|
-
starts, acceptance is decided on evidence (tests executed, findings
|
|
82
|
-
addressed), and every run leaves an auditable trail in `.ai/runs/`.
|
|
4
|
+
directory for run state, one marker-fenced policy section in `AGENTS.md`,
|
|
5
|
+
and per-role subagent definitions with preselected models for the harnesses
|
|
6
|
+
you actually use (Claude Code, OpenAI Codex, opencode).
|
|
7
|
+
|
|
8
|
+
The primary agent acts as the orchestrator: it owns goal, plan, task
|
|
9
|
+
validation, acceptance, and the operator handoff. Review is always
|
|
10
|
+
delegated to narrow subagents, and by default so is implementation (see
|
|
11
|
+
[Run modes](#run-modes)); subagents return structured YAML evidence, not
|
|
12
|
+
transcripts, and every unit of work leaves an auditable run directory
|
|
13
|
+
behind. See [Architecture: why this shape](docs/architecture.md) for the
|
|
14
|
+
loop diagram and the reasoning, and [Run contracts](docs/run-contracts.md)
|
|
15
|
+
for the optional frozen acceptance-baseline contract and the
|
|
16
|
+
decision-authority record every run keeps.
|
|
17
|
+
|
|
18
|
+
## Key features
|
|
19
|
+
|
|
20
|
+
- Orchestrator-led workflow: one agent plans and decides; narrow subagents implement and review.
|
|
21
|
+
- Per-harness subagent definitions with preselected, pinned per-role models and effort (Claude Code, Codex, opencode).
|
|
22
|
+
- An auditable `.ai/runs/` directory per unit of work, with an optional frozen acceptance-baseline contract.
|
|
23
|
+
- Agent-led or manual CLI install, both idempotent and conflict-safe on re-run.
|
|
24
|
+
- Operator-level install for projecting routing and profile defaults onto many repositories.
|
|
25
|
+
- A `validate-review-report` CLI to structurally check reviewer YAML returns.
|
|
83
26
|
|
|
84
27
|
## Install
|
|
85
28
|
|
|
29
|
+
Requires Node.js >= 20.
|
|
30
|
+
|
|
86
31
|
### Recommended: agent-led installation
|
|
87
32
|
|
|
88
33
|
Give a coding agent this line:
|
|
@@ -101,10 +46,8 @@ and fallback behavior auditable. The link tracks `master`; pin it to a commit
|
|
|
101
46
|
SHA for a stable audit.
|
|
102
47
|
|
|
103
48
|
The compact skill entrypoint and its routed references form one installed
|
|
104
|
-
bundle
|
|
105
|
-
|
|
106
|
-
current coherent bundle intact unless an explicitly authorized `--force` run
|
|
107
|
-
replaces the affected files.
|
|
49
|
+
bundle; see [Install reference](docs/install-reference.md) for what the
|
|
50
|
+
reinstall conflict check does.
|
|
108
51
|
|
|
109
52
|
### Manual and advanced CLI installation
|
|
110
53
|
|
|
@@ -135,78 +78,26 @@ npx orchestrator-workflow init --profile minimal --yes
|
|
|
135
78
|
**Templates-only mode.** `--harness none` (the literal word `none`, on its
|
|
136
79
|
own) installs only `.ai/workflow/**` and `.ai/runs/.gitkeep`: no
|
|
137
80
|
`AGENTS.md`, no `CLAUDE.md`, no harness-specific directory, and a manifest
|
|
138
|
-
recording `harnesses: []`.
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
`none` combined with a real harness name (`--harness none,claude`) is
|
|
142
|
-
rejected as ambiguous rather than silently picking one. A plain
|
|
143
|
-
**non-interactive** re-run (no `--harness` flag) after a templates-only
|
|
144
|
-
install stays templates-only, for `init` and `apply` alike, even when
|
|
145
|
-
`apply`'s own operator-defaults name a harness or the target has harness
|
|
146
|
-
files on disk from something else; add a harness back with an explicit
|
|
147
|
-
`--harness <list>` on a later run, the same explicit-flag-wins rule
|
|
148
|
-
`--profile`/`--models`/`--tiers` use, applied to the no-harness case. An
|
|
149
|
-
**interactive** re-run is different: it still prompts, with nothing forced
|
|
150
|
-
pre-selected, instead of silently skipping straight back to templates-only
|
|
151
|
-
without asking; deselect every checkbox to stay templates-only. `init` and
|
|
152
|
-
`apply` both pre-check nothing at all on this prompt, and both still
|
|
153
|
-
annotate what is detected on disk with a " (detected)" label; select a
|
|
154
|
-
harness to install it.
|
|
81
|
+
recording `harnesses: []`. `none` combined with a real harness name
|
|
82
|
+
(`--harness none,claude`) is rejected as ambiguous rather than silently
|
|
83
|
+
picking one.
|
|
155
84
|
|
|
156
85
|
```bash
|
|
157
86
|
npx orchestrator-workflow init --harness none --yes
|
|
158
87
|
```
|
|
159
88
|
|
|
89
|
+
See [Install reference](docs/install-reference.md) for the exact re-run
|
|
90
|
+
rules (a plain non-interactive re-run stays templates-only; an interactive
|
|
91
|
+
one still prompts).
|
|
92
|
+
|
|
160
93
|
## Verification sets
|
|
161
94
|
|
|
162
95
|
A repository may check in `.ai/workflow/verify.json` to name the complete
|
|
163
|
-
verification set
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
```json
|
|
169
|
-
{
|
|
170
|
-
"format": "orchestrator-workflow-verification-set/v1",
|
|
171
|
-
"preflight": {
|
|
172
|
-
"kind": "preflight",
|
|
173
|
-
"name": "preflight",
|
|
174
|
-
"cwd": ".",
|
|
175
|
-
"argv": ["preflight", "run", ".", "--json"]
|
|
176
|
-
},
|
|
177
|
-
"extras": [
|
|
178
|
-
{
|
|
179
|
-
"kind": "command",
|
|
180
|
-
"name": "build",
|
|
181
|
-
"phase": "before_preflight",
|
|
182
|
-
"cwd": "packages/example",
|
|
183
|
-
"argv": ["npm", "run", "build"]
|
|
184
|
-
},
|
|
185
|
-
{
|
|
186
|
-
"kind": "command",
|
|
187
|
-
"name": "package-tests",
|
|
188
|
-
"phase": "after_preflight",
|
|
189
|
-
"cwd": "packages/example",
|
|
190
|
-
"argv": ["npm", "test"]
|
|
191
|
-
},
|
|
192
|
-
{
|
|
193
|
-
"kind": "bundlecheck",
|
|
194
|
-
"name": "knowledge-bundle",
|
|
195
|
-
"phase": "after_preflight",
|
|
196
|
-
"cwd": "packages/example",
|
|
197
|
-
"argv": ["npx", "okf-kit", "check", "docs/okf"]
|
|
198
|
-
}
|
|
199
|
-
]
|
|
200
|
-
}
|
|
201
|
-
```
|
|
202
|
-
|
|
203
|
-
The workflow does not execute or validate this file: the orchestrator first
|
|
204
|
-
approves the resolved effective config and scripts, then records a run-local
|
|
205
|
-
snapshot with the set digest, repository identity, executable identity, and
|
|
206
|
-
every result. Preflight JSON reports check results, not the underlying shell
|
|
207
|
-
commands it discovered. A repository with a configured knowledge bundle
|
|
208
|
-
(`knowledge` in `.ai/workflow/manifest.json`; default `docs/okf/`) includes
|
|
209
|
-
its bundle check in every set, even when the task did not edit documentation.
|
|
96
|
+
verification set (a preflight executor plus ordered extras) for an
|
|
97
|
+
implementer or reviewer briefing; the workflow itself never executes or
|
|
98
|
+
validates this file, only the orchestrator does, recording every result.
|
|
99
|
+
See [Verification sets](docs/verification-sets.md) for the worked JSON
|
|
100
|
+
example.
|
|
210
101
|
|
|
211
102
|
## What gets installed
|
|
212
103
|
|
|
@@ -224,64 +115,22 @@ touches (a machine-local absolute path, not written by the installer); add
|
|
|
224
115
|
it to the repository's `.gitignore`.
|
|
225
116
|
|
|
226
117
|
`manifest.json` may also carry a `knowledge` list of `{ path, repoRoot }`
|
|
227
|
-
entries
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
236
|
-
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
`.ai/workflow/manifest.json` is absent or an empty list, the default
|
|
244
|
-
`docs/okf/` applies, today's behaviour. `doctor` prints a `knowledge:`
|
|
245
|
-
detail line (the `knowledgeWarnings` key in `--json`) for a configured
|
|
246
|
-
`path` or `repoRoot` that is not a directory, for each ignored invalid
|
|
247
|
-
entry, and when a non-empty list omits an existing default bundle directory.
|
|
248
|
-
These warnings never change the status or the exit code.
|
|
249
|
-
|
|
250
|
-
Per selected harness:
|
|
251
|
-
|
|
252
|
-
Each installed skill includes the compact `SKILL.md` entrypoint and every
|
|
253
|
-
regular Markdown file from its adjacent `references/` directory. The entrypoint
|
|
254
|
-
routes run-state/harness, contracts, evidence/probes, and review/recovery work
|
|
255
|
-
to those files; references are part of the installed skill, not optional docs.
|
|
256
|
-
|
|
257
|
-
| Harness | Files | Notes |
|
|
258
|
-
|---|---|---|
|
|
259
|
-
| Claude Code | `.claude/skills/orchestrator-workflow/{SKILL.md,references/*.md}`, `.claude/agents/{explorer,task-slicer,implementer,reviewer,advisor}.md`, `CLAUDE.md` | Claude Code reads `CLAUDE.md`, not `AGENTS.md`; the installer adds an additive `@AGENTS.md` import. Subagent models go into the `model:` frontmatter; the read-only explorer, reviewer, and advisor also get `disallowedTools: Edit, Write, NotebookEdit`. |
|
|
260
|
-
| OpenAI Codex | `.agents/skills/orchestrator-workflow/{SKILL.md,references/*.md}`, `.codex/agents/{explorer,task-slicer,implementer,reviewer,advisor}.toml` | Codex reads `AGENTS.md` natively. Native custom-agent files carry the canonical role instructions plus `model` and `model_reasoning_effort`. Explorer and advisor request a read-only sandbox; reviewer inherits the caller's sandbox so it can run temporary/build checks, while its prompt prohibits source edits. |
|
|
261
|
-
| opencode | `.opencode/skills/orchestrator-workflow/{SKILL.md,references/*.md}`, `.opencode/agents/{explorer,task-slicer,implementer,reviewer,advisor}.md` | opencode reads `AGENTS.md` natively. Subagents get `mode: subagent`; the read-only explorer, reviewer, and advisor also get `permission: edit: deny`. Model resolution is described below. |
|
|
262
|
-
|
|
263
|
-
**Read-only posture, honestly stated.** Claude Code disables file-mutation
|
|
264
|
-
tools for explorer, reviewer, and advisor; opencode denies edits for those
|
|
265
|
-
roles. Codex requests a read-only sandbox for explorer and advisor. Its
|
|
266
|
-
reviewer inherits the caller's sandbox so temporary/build checks remain
|
|
267
|
-
possible, while its prompt prohibits source edits. In inherited or otherwise
|
|
268
|
-
write-enabled sandboxes, shell-level mutation (`git checkout`,
|
|
269
|
-
`git restore`, `git clean`, `git stash`, `git reset`, `sed -i`, redirecting
|
|
270
|
-
output into a file, which the reviewer may do only inside its write boundary below) is guarded by instruction only: the agent prompts forbid
|
|
271
|
-
it explicitly, but the role definition itself does not prevent it. A native
|
|
272
|
-
read-only sandbox can block those writes. This residual has bitten in practice (a
|
|
273
|
-
reviewer ran `git checkout` and discarded uncommitted work), which is why the
|
|
274
|
-
prompts now name the forbidden commands instead of just saying "read-only".
|
|
275
|
-
The reviewer's own write boundary is narrower than "read-only": it may write
|
|
276
|
-
to its own scratchpad (a scratch copy or replay of the repository) and to
|
|
277
|
-
the run directory's `evidence/`, and nowhere else. It never writes into the
|
|
278
|
-
reviewed tree, its index, its refs, or its object store: no `git fetch`, no
|
|
279
|
-
`git merge-tree --write-tree`, no `git update-ref`, no `git gc`, on top of
|
|
280
|
-
the working-tree and index mutations already forbidden above. A write a
|
|
281
|
-
declared check or the probe runner's own isolation leaves behind is expected
|
|
282
|
-
wherever that tool places it, not an exception to this rule.
|
|
283
|
-
Marker- or verdict-style enforcement of the Bash residual (sandboxing,
|
|
284
|
-
PreToolUse hooks) is harness territory and out of this kit's scope.
|
|
118
|
+
entries for a repo whose bundle is not at the default location; when
|
|
119
|
+
`knowledge` in `.ai/workflow/manifest.json` is absent or an empty list, the
|
|
120
|
+
default `docs/okf/` applies. See
|
|
121
|
+
[Install reference](docs/install-reference.md) for the field's exact path
|
|
122
|
+
rules and validation behavior.
|
|
123
|
+
|
|
124
|
+
Per selected harness, the installer writes the compact `SKILL.md` entrypoint,
|
|
125
|
+
its `references/` files, and one subagent definition per role
|
|
126
|
+
(`.claude/agents/`, `.codex/agents/`, or `.opencode/agents/`). Each harness's
|
|
127
|
+
read-only posture (explorer, reviewer, and advisor) is tool-level where the
|
|
128
|
+
harness supports it. Codex's reviewer inherits the caller's sandbox instead
|
|
129
|
+
(its prompt prohibits source edits, so temporary/build checks stay
|
|
130
|
+
possible), and wherever a sandbox is writable, shell-level mutation is
|
|
131
|
+
guarded by instruction only, not enforced. See [Harnesses](docs/harnesses.md)
|
|
132
|
+
for the exact per-harness file list and the honest read-only-posture
|
|
133
|
+
writeup, including the reviewer's own narrower write boundary.
|
|
285
134
|
|
|
286
135
|
## Role profile
|
|
287
136
|
|
|
@@ -299,16 +148,11 @@ not "just implementer". There is no per-role checklist; the two profiles are
|
|
|
299
148
|
the only supported shapes.
|
|
300
149
|
|
|
301
150
|
**Advisor (escalation).** The fifth `full`-profile role, `advisor`, is
|
|
302
|
-
read-only and consulted only
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
|
|
307
|
-
tier choice, spawning it is the orchestrator's own judgment call. The advisor
|
|
308
|
-
lays out the options with their pros, cons, and risk, and gives a
|
|
309
|
-
recommendation — it recommends, never decides, and never writes code; the
|
|
310
|
-
orchestrator still decides, and a critical risk still goes to the operator.
|
|
311
|
-
`minimal` never installs it, the same as explorer and task-slicer.
|
|
151
|
+
read-only and consulted only at defined escalation triggers (architectural
|
|
152
|
+
uncertainty, a review deadlock, a high-risk decision, and similar); it
|
|
153
|
+
recommends, never decides. `minimal` never installs it, the same as explorer
|
|
154
|
+
and task-slicer. See [Role profile reference](docs/role-profile-reference.md)
|
|
155
|
+
for the full trigger list.
|
|
312
156
|
|
|
313
157
|
```bash
|
|
314
158
|
npx orchestrator-workflow init --profile minimal --yes
|
|
@@ -319,256 +163,33 @@ which profile to install — defaulting to `full`. `--profile` rejects any
|
|
|
319
163
|
value other than `minimal` or `full` with a clear error instead of silently
|
|
320
164
|
falling back to a default.
|
|
321
165
|
|
|
322
|
-
**Re-runs and profile changes.** A plain re-run
|
|
323
|
-
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
updating the manifest to match. Switching profiles follows the same
|
|
328
|
-
precedent already in place for dropping a harness from `--harness` on a
|
|
329
|
-
re-run: files for roles no longer in the profile are simply no longer
|
|
330
|
-
installed or tracked in the manifest; they are not automatically deleted
|
|
331
|
-
from disk. `init` detects a `full` → `minimal` downgrade and prints a note
|
|
332
|
-
naming the now-untracked `task-slicer.md` / `explorer.md` / `advisor.md`
|
|
333
|
-
agent files and how to remove them. For a fully clean switch, run `orchestrator-workflow
|
|
334
|
-
uninstall` first, or remove those files by hand. Uninstalling a `minimal`
|
|
335
|
-
install that has never been downgraded from `full` is always clean on its
|
|
336
|
-
own: it only ever removes what it actually installed, so there is nothing to
|
|
337
|
-
report as missing for the roles that were never written. A `minimal` install
|
|
338
|
-
reached via a `full` → `minimal` downgrade is not clean in that sense: the
|
|
339
|
-
downgrade's now-untracked `task-slicer.md` / `explorer.md` / `advisor.md`
|
|
340
|
-
files are not in the manifest's file ledger, so uninstall leaves them on disk
|
|
341
|
-
without reporting them at all.
|
|
166
|
+
**Re-runs and profile changes.** A plain re-run keeps the recorded profile;
|
|
167
|
+
`--profile` explicitly overrides it, and a `full` to `minimal` downgrade
|
|
168
|
+
leaves the now-untracked role files on disk with a printed note. See
|
|
169
|
+
[Role profile reference](docs/role-profile-reference.md) for the exact
|
|
170
|
+
override and uninstall-cleanliness rules.
|
|
342
171
|
|
|
343
172
|
## Model preselection
|
|
344
173
|
|
|
345
174
|
Routing is a harness-specific map from role and tier to a complete
|
|
346
|
-
`{model, effort}` selection
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
|
|
350
|
-
`<role>-<tier
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
|
|
357
|
-
|
|
358
|
-
|
|
359
|
-
|
|
360
|
-
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
Changing one leaf leaves the others intact, which makes a previous manifest a
|
|
365
|
-
usable rollback record. Model updates are deliberate per role and tier: the
|
|
366
|
-
installer never interprets a newer model as automatically better and never
|
|
367
|
-
rewrites a preserved choice merely because another model exists.
|
|
368
|
-
|
|
369
|
-
`--models` remains as the backward-compatible, per-role input for Claude Code
|
|
370
|
-
and opencode. It does not configure Codex. Existing manifests that contain
|
|
371
|
-
only `models` continue to produce the same Claude Code and opencode defaults:
|
|
372
|
-
|
|
373
|
-
| Role | Default | Why |
|
|
374
|
-
|---|---|---|
|
|
375
|
-
| explorer | `sonnet` | read-only terrain mapping is broad reading, not deep reasoning |
|
|
376
|
-
| task-slicer | `sonnet` | structured decomposition, no deep reasoning needed |
|
|
377
|
-
| implementer | `sonnet` | fast, cheap, good enough for narrow pre-sliced tasks |
|
|
378
|
-
| reviewer | `opus` | skeptical review benefits from the strongest model |
|
|
379
|
-
| advisor | `opus` | escalations happen precisely when the situation is hard, so it shares the reviewer's strongest-model default |
|
|
380
|
-
|
|
381
|
-
The orchestrator itself runs on the session's main model. For Codex, start the
|
|
382
|
-
orchestrator on `gpt-6-astra` at `high` effort; use `xhigh` for demanding work.
|
|
383
|
-
The installer does not mutate global or fleet Codex configuration to enforce
|
|
384
|
-
that recommendation.
|
|
385
|
-
|
|
386
|
-
**Codex defaults.** Codex uses native `.codex/agents/*.toml` custom agents.
|
|
387
|
-
The file shape follows the
|
|
388
|
-
[official Codex subagent configuration](https://learn.chatgpt.com/docs/agent-configuration/subagents).
|
|
389
|
-
The shipped routing is:
|
|
390
|
-
|
|
391
|
-
| Role | Tier | Model | Effort |
|
|
392
|
-
|---|---|---|---|
|
|
393
|
-
| explorer | low | `gpt-5.6-luna` | low |
|
|
394
|
-
| explorer | medium (default) | `gpt-5.6-sol` | medium |
|
|
395
|
-
| explorer | high | `gpt-5.6-sol` | high |
|
|
396
|
-
| task-slicer | low | `gpt-5.6-luna` | low |
|
|
397
|
-
| task-slicer | medium (default) | `gpt-5.6-sol` | medium |
|
|
398
|
-
| task-slicer | high | `gpt-5.6-sol` | high |
|
|
399
|
-
| implementer | low | `gpt-5.6-luna` | low |
|
|
400
|
-
| implementer | medium (default) | `gpt-5.6-terra` | medium |
|
|
401
|
-
| implementer | high | `gpt-5.6-terra` | high |
|
|
402
|
-
| implementer | xhigh | `gpt-6-astra` | xhigh |
|
|
403
|
-
| reviewer | medium | `gpt-5.6-terra` | medium |
|
|
404
|
-
| reviewer | high (default) | `gpt-6-astra` | high |
|
|
405
|
-
| reviewer | xhigh | `gpt-6-astra` | xhigh |
|
|
406
|
-
| advisor | high (default) | `gpt-6-astra` | high |
|
|
407
|
-
| advisor | xhigh | `gpt-6-astra` | xhigh |
|
|
408
|
-
|
|
409
|
-
When you have a deterministic Codex model catalog, pass it with
|
|
410
|
-
`--codex-catalog <json-file>`. The CLI validates the selected Codex model and
|
|
411
|
-
effort pairs before writing. Without a supplied catalog it performs no online
|
|
412
|
-
entitlement check; offline or account-specific availability remains unknown.
|
|
413
|
-
Use the harness's native capability commands, such as `codex debug models`, to
|
|
414
|
-
refresh a catalog before installation when appropriate. A bundled-capability
|
|
415
|
-
view describes what the binary knows and does not prove account entitlement.
|
|
416
|
-
|
|
417
|
-
**opencode model resolution.** opencode requires fully-qualified `provider/model-id`
|
|
418
|
-
strings (e.g. `github-copilot/claude-sonnet-4.6`). At install time the CLI
|
|
419
|
-
runs `opencode models` to fetch the live catalog and auto-detects which
|
|
420
|
-
provider to use (the one that offers Claude models). When exactly one such
|
|
421
|
-
provider exists the aliases are resolved to the highest-version matching id in
|
|
422
|
-
the catalog. When multiple providers offer Claude models the CLI warns and asks
|
|
423
|
-
you to pass `--opencode-provider <id>` to disambiguate, or to supply
|
|
424
|
-
fully-qualified ids per role via `--models`. If no resolution is possible
|
|
425
|
-
(catalog empty, `opencode` binary absent, ambiguous provider) the `model:`
|
|
426
|
-
frontmatter line is omitted entirely and the subagent inherits the
|
|
427
|
-
session/default model — a safe, portable fallback. Fully-qualified ids in
|
|
428
|
-
`--models` always pass through unchanged regardless of the catalog.
|
|
429
|
-
Nested-path providers like `openrouter` (whose ids look like
|
|
430
|
-
`openrouter/anthropic/claude-...`) are not auto-resolved from aliases and must
|
|
431
|
-
be supplied as a fully-qualified `--models` entry, e.g.
|
|
432
|
-
`reviewer=openrouter/anthropic/claude-opus-4.8`.
|
|
433
|
-
|
|
434
|
-
## Effort tiers
|
|
435
|
-
|
|
436
|
-
`--tiers` renders an additional per-role subagent definition for each
|
|
437
|
-
non-default effort tier, alongside the one default (unsuffixed) agent file
|
|
438
|
-
`--profile` already installs. Each tier variant is a standalone subagent
|
|
439
|
-
definition, not a modification of the default file. Claude Code and opencode
|
|
440
|
-
use `<role>.md` / `<role>-<tier>.md`; Codex uses `<role>.toml` /
|
|
441
|
-
`<role>-<tier>.toml`.
|
|
442
|
-
|
|
443
|
-
**Every default file carries its own pinned effort, independent of
|
|
444
|
-
`--tiers`.** The harness composers add the role's own default routing
|
|
445
|
-
selection to the unsuffixed file. In the legacy Claude/opencode path this is
|
|
446
|
-
`TIER_DEFS[DEFAULT_TIER[role]].effort`: `effort: medium` for explorer,
|
|
447
|
-
task-slicer, and implementer; `effort: high` for reviewer and advisor
|
|
448
|
-
(opencode: a `variant: high` line when the resolved model is Claude-family,
|
|
449
|
-
following the same dispatch rule tier variants use, `reasoningEffort:
|
|
450
|
-
medium`/`reasoningEffort: high` for a non-Claude-family provider-qualified
|
|
451
|
-
model, nothing for Ollama, a provider-less id, or an unresolved model). This
|
|
452
|
-
pin does not depend on `tiers`, so a plain install (no `--tiers`) already
|
|
453
|
-
carries it; the flag only controls whether the additional `<role>-<tier>.md`
|
|
454
|
-
variant files are also rendered. The motivation: a default spawn used to
|
|
455
|
-
silently inherit the orchestrator session's own effort, so a `high`-effort
|
|
456
|
-
orchestrator session made every default subagent spawn at `high` too,
|
|
457
|
-
regardless of the role's own intended weight; the pin makes each role's
|
|
458
|
-
effort deterministic and independent of the caller's session. A `--tiers`-off
|
|
459
|
-
install (the default) has no variant files and therefore no in-install
|
|
460
|
-
escalation path off a default's pinned effort; run `init --tiers` afterward
|
|
461
|
-
if a task ever needs one.
|
|
462
|
-
|
|
463
|
-
Default off, like every optional pack in this kit: a fresh install renders
|
|
464
|
-
no variant files unless asked. `--tiers` turns the feature on for that run,
|
|
465
|
-
`--no-tiers` turns it off; a plain re-run with neither flag keeps whatever
|
|
466
|
-
the previous install had, the same override-vs-persist rule already used
|
|
467
|
-
for `--profile` and `--models`. There is no interactive prompt for it:
|
|
468
|
-
`tiers` is opt-in/off via the flags only. Neither Codex nor the other harnesses
|
|
469
|
-
get `max` or `ultra` variants from this kit.
|
|
470
|
-
|
|
471
|
-
```bash
|
|
472
|
-
npx orchestrator-workflow init --tiers --yes
|
|
473
|
-
```
|
|
474
|
-
|
|
475
|
-
Turning tiers back off with `--no-tiers` after having them on follows the
|
|
476
|
-
same pattern as a `full` → `minimal` profile downgrade: `init` prints a note
|
|
477
|
-
naming the now-untracked `<role>-<tier>.md` variant files and how to remove
|
|
478
|
-
them, rather than deleting them or leaving the leftover unexplained.
|
|
479
|
-
|
|
480
|
-
**Which tiers each role gets.** A role never gets a variant file for its own
|
|
481
|
-
default tier: that would collide with, and duplicate, the default file.
|
|
482
|
-
|
|
483
|
-
| Role | Tiers available | Default tier (no variant file) |
|
|
484
|
-
|---|---|---|
|
|
485
|
-
| explorer | low, medium, high | medium |
|
|
486
|
-
| task-slicer | low, medium, high | medium |
|
|
487
|
-
| implementer | low, medium, high, xhigh | medium |
|
|
488
|
-
| reviewer | medium, high, xhigh | high |
|
|
489
|
-
| advisor | high, xhigh | high |
|
|
490
|
-
|
|
491
|
-
With `--profile full` and every tier rendered, that is 5 default files plus
|
|
492
|
-
10 variant files: 15 files total per harness.
|
|
493
|
-
|
|
494
|
-
**Tier → model class → effort.** Each tier resolves to a model class and an
|
|
495
|
-
effort value:
|
|
496
|
-
|
|
497
|
-
| Tier | Model class | Model alias | Effort requested |
|
|
498
|
-
|---|---|---|---|
|
|
499
|
-
| low | small | `haiku` | `low` |
|
|
500
|
-
| medium | medium | `sonnet` | `medium` |
|
|
501
|
-
| high | medium | `sonnet` | `high` |
|
|
502
|
-
| xhigh | large | `opus` | `xhigh` |
|
|
503
|
-
|
|
504
|
-
Claude Code variants carry both a `model:` line (the class's alias) and an
|
|
505
|
-
`effort: <tier>` line in frontmatter. Read-only roles (explorer, reviewer,
|
|
506
|
-
advisor) keep `disallowedTools: Edit, Write, NotebookEdit` on their variants
|
|
507
|
-
too.
|
|
508
|
-
|
|
509
|
-
**opencode variants key off the resolved model's family, not its provider
|
|
510
|
-
prefix**, since opencode's effort surface is not uniform across model
|
|
511
|
-
families:
|
|
512
|
-
|
|
513
|
-
- **Claude-family models** (any resolved id whose provider is
|
|
514
|
-
`anthropic/`, or whose segment after the provider prefix contains
|
|
515
|
-
`claude-`, which covers `anthropic/claude-...` as well as a Claude model
|
|
516
|
-
fronted by a different provider, e.g. `github-copilot/claude-sonnet-4.6`
|
|
517
|
-
or the nested `openrouter/anthropic/claude-opus-4.8`): only `high` and
|
|
518
|
-
`xhigh` get an effort field, as `variant: high` and `variant: max`
|
|
519
|
-
respectively; `low` and `medium` collapse to no effort field at all,
|
|
520
|
-
since opencode's `variant:` option does not distinguish an effort below
|
|
521
|
-
`high`. This collapse is deliberate and documented, not a bug: a
|
|
522
|
-
`low`/`medium` variant on a Claude-family model still gets its class's
|
|
523
|
-
`model:` line, just no `variant:` line.
|
|
524
|
-
- **Ollama, or an id with no provider prefix**: no effort field at all.
|
|
525
|
-
There is no known effort passthrough for Ollama, and an id with no `/`
|
|
526
|
-
resolves to no provider to key the decision on.
|
|
527
|
-
- **Every other non-Claude-family model**: a plain `reasoningEffort: <tier>`
|
|
528
|
-
line, `xhigh` included (opencode's built-in OpenAI-style variants
|
|
529
|
-
document an `xhigh` reasoning effort).
|
|
530
|
-
|
|
531
|
-
The variant's `model:` line is resolved the same way the base per-role model
|
|
532
|
-
is (an `opencode models` catalog lookup against the auto-detected or
|
|
533
|
-
`--opencode-provider`-specified provider), just keyed by the tier's model
|
|
534
|
-
class instead of by role. When that lookup cannot resolve a model for a
|
|
535
|
-
class, the CLI warns once on stderr and **no variant file is rendered for
|
|
536
|
-
that class at all**, not a file with a missing `model:` line: a variant
|
|
537
|
-
with no resolved model would carry neither a `model:` nor an effort line, an
|
|
538
|
-
indistinguishable no-op duplicate of the base file with no ledger entry to
|
|
539
|
-
compare it against, so `init` skips writing it entirely. This guard and its
|
|
540
|
-
warning are opencode-scoped only; Claude Code variants resolve `model:` from
|
|
541
|
-
a plain alias (`haiku`/`sonnet`/`opus`) and need no live catalog lookup, so
|
|
542
|
-
they are unaffected.
|
|
543
|
-
|
|
544
|
-
Codex variants carry `model` and `model_reasoning_effort` from their exact
|
|
545
|
-
routing leaf. The canonical role prompt becomes `developer_instructions`.
|
|
546
|
-
Runtime dispatch follows the client's actual capabilities: select the named
|
|
547
|
-
installed agent when supported; otherwise, if spawning supports explicit model
|
|
548
|
-
and effort, read the installed TOML and pass its selection, developer
|
|
549
|
-
instructions, and narrow task contract into a fresh task-local spawn. A
|
|
550
|
-
full-history spawn may not permit a model override. If that explicit spawn
|
|
551
|
-
cannot accept a sandbox override, explorer and advisor inherit the caller's
|
|
552
|
-
sandbox and their prompt is the edit guard. When native spawning is
|
|
553
|
-
unavailable, run the same contract inline and sequentially. The orchestrator
|
|
554
|
-
alone spawns agents. In particular, it must not choose `implementer-low` when
|
|
555
|
-
the task requires a test, typecheck, lint, build, or named mutation probe.
|
|
556
|
-
|
|
557
|
-
**Warning: `CLAUDE_CODE_EFFORT_LEVEL` overrides every agent's frontmatter
|
|
558
|
-
`effort:`, tier variants included.** Claude Code's `effort:` frontmatter
|
|
559
|
-
field does work: it reaches the model request as `output_config.effort`.
|
|
560
|
-
But when the harness environment sets `CLAUDE_CODE_EFFORT_LEVEL`, that
|
|
561
|
-
environment variable wins over the frontmatter `effort:` on every installed
|
|
562
|
-
agent, tier variants and default files alike, not just the one this feature
|
|
563
|
-
adds. Check for it before relying on a specific tier variant's requested
|
|
564
|
-
effort actually taking effect.
|
|
565
|
-
|
|
566
|
-
The pin is also emitted unconditionally regardless of which model the role
|
|
567
|
-
resolves to via `--models`, including a model with no effort support at all
|
|
568
|
-
(e.g. `--models reviewer=haiku` still renders `model: haiku` followed by
|
|
569
|
-
`effort: high`). On Haiku 4.5, which does not support the `effort`
|
|
570
|
-
parameter, the harness ignores the pinned value rather than rejecting it
|
|
571
|
-
(anchored by a measurement, see CHANGELOG 0.23.0).
|
|
175
|
+
`{model, effort}` selection, set with `--routing <json-file>` and deep-merged
|
|
176
|
+
into `.ai/workflow/manifest.json`; `--models` is the backward-compatible,
|
|
177
|
+
per-role input for Claude Code and opencode only (never Codex); `--codex-models <json-file>` is a sparse Codex-only alias map whose supplied aliases update their role/tier leaves below any explicit `--routing` leaf. Every
|
|
178
|
+
installed agent file carries its own pinned effort regardless of `--tiers`;
|
|
179
|
+
`--tiers` additionally renders one `<role>-<tier>.md`/`.toml` variant file
|
|
180
|
+
per non-default tier. See [Model routing reference](docs/model-routing-reference.md)
|
|
181
|
+
for the default-model table, the `--routing` JSON shape, the Codex default
|
|
182
|
+
routing table, opencode model resolution, and the full effort-tiers
|
|
183
|
+
mechanics (including the `CLAUDE_CODE_EFFORT_LEVEL` environment override).
|
|
184
|
+
|
|
185
|
+
### Effort tiers
|
|
186
|
+
|
|
187
|
+
The per-role default effort (`medium` for explorer, task-slicer, and
|
|
188
|
+
implementer; `high` for reviewer and advisor) is pinned in each agent file
|
|
189
|
+
on Claude Code and Codex; on opencode it depends on the resolved model.
|
|
190
|
+
See [Model routing reference: Effort tiers](docs/model-routing-reference.md#effort-tiers)
|
|
191
|
+
for the full role/tier table, the per-harness frontmatter shape, and the
|
|
192
|
+
tier variants `--tiers` renders.
|
|
572
193
|
|
|
573
194
|
## Run modes
|
|
574
195
|
|
|
@@ -583,128 +204,26 @@ reference
|
|
|
583
204
|
|
|
584
205
|
## Operator-level install
|
|
585
206
|
|
|
586
|
-
|
|
587
|
-
|
|
588
|
-
|
|
589
|
-
|
|
590
|
-
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
|
|
594
|
-
|
|
595
|
-
```bash
|
|
596
|
-
orchestrator-workflow setup --yes
|
|
597
|
-
orchestrator-workflow apply --target /path/to/repo
|
|
598
|
-
```
|
|
599
|
-
|
|
600
|
-
**`setup`** writes or updates this operator's default install options
|
|
601
|
-
(harnesses, profile, legacy models, routing, tiers) as the baseline for future installs; it
|
|
602
|
-
touches no repository. A flag always wins; a flag-less re-run keeps the
|
|
603
|
-
previously stored values; a first-ever `setup` falls back to `claude` /
|
|
604
|
-
`full` / the kit's default routing / tiers off. `setup` takes the same
|
|
605
|
-
option flags as `init` (`--harness`, `--profile`, `--models`, `--routing`,
|
|
606
|
-
`--codex-catalog`, `--tiers` / `--no-tiers`, `--opencode-provider`, `--yes`).
|
|
607
|
-
The defaults live in
|
|
608
|
-
`<operator home>/manifest.json`, where the operator home is
|
|
609
|
-
`~/.orchestrator-workflow/` unless the `ORCHESTRATOR_WORKFLOW_HOME`
|
|
610
|
-
environment variable names a different directory.
|
|
611
|
-
|
|
612
|
-
**`apply --target <repo>`** projects the operator's install onto a target
|
|
613
|
-
repository and registers that target, by its real resolved path, in the
|
|
614
|
-
operator manifest. It requires a prior `orchestrator-workflow setup`;
|
|
615
|
-
without one it exits `1` with "No operator setup found". Option resolution
|
|
616
|
-
follows one precedence order: an
|
|
617
|
-
explicit flag wins, then the target's own previously recorded settings,
|
|
618
|
-
then the operator's defaults (harnesses fall back one step further, to
|
|
619
|
-
what `init` would have auto-detected) -- except a target whose own
|
|
620
|
-
manifest recorded a real `harnesses: []` (a deliberate templates-only
|
|
621
|
-
install, see "Templates-only mode" above), which stays templates-only on
|
|
622
|
-
a flagless run regardless of the operator's defaults or what is on disk;
|
|
623
|
-
an **interactive** re-run on such a target still prompts, with the same
|
|
624
|
-
nothing-pre-checked behaviour described in "Templates-only mode" above
|
|
625
|
-
(it applies identically to `apply`).
|
|
626
|
-
Pass `--sync` to invert that for
|
|
627
|
-
profile, tiers, legacy models, and routing: the operator's defaults then win over whatever
|
|
628
|
-
the target already had recorded. A target pinned to a kit version other
|
|
629
|
-
than the one being applied is skipped rather than touched (see the pin
|
|
630
|
-
rule below). `apply` also takes the same install options as `init` (`--harness`,
|
|
631
|
-
`--profile`, `--models`, `--routing`, `--codex-catalog`, `--tiers` /
|
|
632
|
-
`--no-tiers`, `--opencode-provider`, `--force`, `--yes`), which feed the
|
|
633
|
-
precedence rule above. An explicit routing file is the highest-precedence
|
|
634
|
-
deep patch; leaves it omits retain their resolved baseline values.
|
|
635
|
-
|
|
636
|
-
**`doctor [--json] [--prune]`** reports every operator-registered target's
|
|
637
|
-
status: `clean`, `divergent` (from the operator defaults, including routing), `version-lag`,
|
|
638
|
-
`drift` (installed files edited, deleted, or unreadable since install),
|
|
639
|
-
`missing`, `no-manifest`, or `unverifiable`. It exits `2` when the operator
|
|
640
|
-
manifest is missing or unreadable, or, with `--prune`, when the operator
|
|
641
|
-
manifest lock cannot be acquired or the rewrite fails; `1` when any target
|
|
642
|
-
is `drift`, `missing`, `no-manifest`, or `unverifiable`; and `0` otherwise.
|
|
643
|
-
`--json` prints one JSON object instead
|
|
644
|
-
of human output, with one entry per target plus the operator home and
|
|
645
|
-
version. `--prune` removes `missing` and `no-manifest` targets from the
|
|
646
|
-
registry before reporting (never an `unverifiable` one, since that status
|
|
647
|
-
means the check itself was inconclusive, not that the target is confirmed
|
|
648
|
-
gone) and rewrites the manifest file in its normalized form.
|
|
649
|
-
For a legacy opencode leaf without a recorded provider-qualified model id,
|
|
650
|
-
doctor reports `Routing comparison incomplete` and includes
|
|
651
|
-
`routingComparisonGaps` in JSON instead of declaring a false routing
|
|
652
|
-
divergence. The gap alone does not change the target status.
|
|
653
|
-
|
|
654
|
-
**`adopt [dir] [--json]`** brings a repository that already has the kit installed,
|
|
655
|
-
by hand or by an earlier `init`, under the operator's management without
|
|
656
|
-
changing anything in that repository: it registers the repository
|
|
657
|
-
verbatim, using the repository's own recorded settings to bootstrap the
|
|
658
|
-
operator manifest when none exists yet, records the repository's own
|
|
659
|
-
installed version as its baseline, and prints that one target's `doctor`
|
|
660
|
-
report. It exits `1` only when the freshly adopted target itself reports
|
|
661
|
-
drift, and `2` for a precondition failure (no repo manifest, an unreadable
|
|
662
|
-
or foreign manifest, or a lock or write failure).
|
|
663
|
-
|
|
664
|
-
**The kit-version pin.** A repository's own manifest can additionally
|
|
665
|
-
carry an optional `pin`: a kit version that `apply` must match before it
|
|
666
|
-
will touch that repository again. `apply` skips (exit `0`) a target pinned
|
|
667
|
-
to a different version than the one being applied. `--pin <version>` sets
|
|
668
|
-
or replaces the pin and applies regardless of any existing one; `--unpin`
|
|
669
|
-
clears it and applies; `--force-pin` advances an existing pin that
|
|
670
|
-
differs, but has no effect on a target with no pin recorded (it stays
|
|
671
|
-
unpinned). `doctor` reports `version-lag` when the installed version
|
|
672
|
-
differs from the running kit version; on a pinned target the pin is
|
|
673
|
-
compared against the installed version instead, so a pin equal to the
|
|
674
|
-
installed version is `clean` and a pin that no longer matches it is
|
|
675
|
-
`version-lag`.
|
|
676
|
-
|
|
677
|
-
**The registry is implicit**, not a separate command: `apply` and `adopt`
|
|
678
|
-
register a target as a side effect of a real run, and `doctor --prune` is
|
|
679
|
-
how a registry entry is removed again; there is no bare register or
|
|
680
|
-
unregister command. The workspace root of a multi-repo checkout is treated
|
|
681
|
-
as an ordinary target, nothing special.
|
|
682
|
-
|
|
683
|
-
All writes to the operator manifest, by `setup`, `apply`, `doctor --prune`,
|
|
684
|
-
and `adopt` alike, go through one advisory lock in the operator home, so
|
|
685
|
-
concurrent orchestrator-workflow commands on the same machine cannot
|
|
686
|
-
corrupt each other's state.
|
|
207
|
+
An operator who maintains many repositories can set defaults once with
|
|
208
|
+
`setup` and project them onto each target with `apply --target <repo>`,
|
|
209
|
+
instead of re-answering the same `init` prompts per repo; `doctor` reports
|
|
210
|
+
each registered target's status and `adopt` brings an already-installed
|
|
211
|
+
repository under management without changing it. `init`/`uninstall` remain
|
|
212
|
+
fully supported and unchanged for a single-repository install. See
|
|
213
|
+
[Operator-level install](docs/operator-install.md) for the full command
|
|
214
|
+
reference (every flag, the `--sync` precedence inversion, the kit-version
|
|
215
|
+
pin, and the registry/locking model).
|
|
687
216
|
|
|
688
217
|
## Ownership and re-runs
|
|
689
218
|
|
|
690
|
-
`init` is idempotent: a second run changes nothing. `
|
|
691
|
-
|
|
692
|
-
|
|
693
|
-
|
|
694
|
-
|
|
695
|
-
|
|
696
|
-
|
|
697
|
-
|
|
698
|
-
between its own markers. A broken or duplicated marker fence is reported as
|
|
699
|
-
a conflict and left alone.
|
|
700
|
-
- Templates, skills, and subagent definitions are kit-owned. The manifest
|
|
701
|
-
records a hash of each file as installed, so a re-run after a kit upgrade
|
|
702
|
-
updates files you never touched and reports files you edited as conflicts
|
|
703
|
-
instead of overwriting them; `--force` overwrites those too.
|
|
704
|
-
- `.ai/workflow/manifest.json` is the kit's state file. It records the applied
|
|
705
|
-
version, harnesses, role profile, models, the `--tiers` flag, the optional
|
|
706
|
-
kit-version pin, and file hashes, and is rewritten whenever that state
|
|
707
|
-
changes; do not edit it by hand.
|
|
219
|
+
`init` is idempotent: a second run changes nothing. `AGENTS.md`/`CLAUDE.md`
|
|
220
|
+
belong to you (the installer only touches its own fenced section or import
|
|
221
|
+
line); templates, skills, and subagent definitions are kit-owned and
|
|
222
|
+
conflict-checked by file hash; `.ai/workflow/manifest.json` is the kit's
|
|
223
|
+
state file. `apply` installs through the same path and is subject to the
|
|
224
|
+
same rules, refreshing this target's entry in the operator manifest on
|
|
225
|
+
every run. See [Install reference](docs/install-reference.md) for the
|
|
226
|
+
exact per-file ownership rules.
|
|
708
227
|
|
|
709
228
|
## Uninstall
|
|
710
229
|
|
|
@@ -720,75 +239,43 @@ init's own boilerplate remains. Kit directories are pruned only when empty,
|
|
|
720
239
|
and run history under `.ai/runs/` is always kept. Interactive runs ask for
|
|
721
240
|
confirmation; non-interactive runs require `--yes`.
|
|
722
241
|
|
|
723
|
-
## Relation to agentic-coding-playbook
|
|
724
|
-
|
|
725
|
-
This kit ships the orchestration layer: who coordinates whom, where state
|
|
726
|
-
lives, and the I/O contracts between roles. The extended role prompts and the
|
|
727
|
-
organizational guidance (when to use agents at all, review depth, risk tiers)
|
|
728
|
-
live in the sibling package
|
|
729
|
-
[agentic-coding-playbook](../agentic-coding-playbook), which the skill
|
|
730
|
-
references.
|
|
731
|
-
|
|
732
|
-
## okf-kit version pin
|
|
733
|
-
|
|
734
|
-
`test/docs-consistency.test.ts` pins the `okf-kit@<version>` this repo's own
|
|
735
|
-
`.github/workflows/` install against the sibling `packages/okf-kit`
|
|
736
|
-
package's version, so a release of `okf-kit` must bump those pins in the
|
|
737
|
-
same commit as the version cut; see `CONTRIBUTING.md`'s "Releasing okf-kit"
|
|
738
|
-
section (repo root) for the order.
|
|
739
|
-
|
|
740
242
|
## Reviewer-report validation
|
|
741
243
|
|
|
742
244
|
```bash
|
|
743
245
|
orchestrator-workflow validate-review-report path/to/return.yaml
|
|
744
|
-
orchestrator-workflow validate-review-report - < path/to/return.yaml
|
|
745
|
-
orchestrator-workflow validate-review-report path/to/return.yaml --format json
|
|
746
246
|
```
|
|
747
247
|
|
|
748
248
|
Checks a reviewer return's YAML against the reviewer output contract's
|
|
749
|
-
required fields and enums
|
|
750
|
-
|
|
751
|
-
`
|
|
752
|
-
|
|
753
|
-
|
|
754
|
-
|
|
755
|
-
|
|
756
|
-
|
|
757
|
-
|
|
758
|
-
|
|
759
|
-
|
|
760
|
-
|
|
761
|
-
|
|
762
|
-
|
|
763
|
-
|
|
764
|
-
|
|
765
|
-
|
|
766
|
-
|
|
767
|
-
|
|
768
|
-
|
|
769
|
-
|
|
770
|
-
|
|
771
|
-
|
|
772
|
-
|
|
773
|
-
|
|
774
|
-
|
|
775
|
-
|
|
776
|
-
|
|
777
|
-
|
|
778
|
-
|
|
779
|
-
|
|
780
|
-
|
|
781
|
-
|
|
782
|
-
`--format json` governs the validation verdict only: a commander parsing
|
|
783
|
-
error (missing argument, unknown option, excess arguments) or an
|
|
784
|
-
unrecognized `--format` value itself still prints plain text to stderr
|
|
785
|
-
with nothing on stdout, regardless of `--format`; the one exception is an
|
|
786
|
-
unreadable file, which does emit the JSON envelope on stdout. This check
|
|
787
|
-
is structural only: it never judges semantic adequacy, cannot waive a
|
|
788
|
-
finding, and passing it is never orchestrator acceptance. The
|
|
789
|
-
required-field set it checks is hand-maintained in `src/review-report.ts`
|
|
790
|
-
and pinned against the contract block itself by
|
|
791
|
-
`test/docs-consistency.test.ts`, so a contract edit without a matching
|
|
792
|
-
schema edit fails the suite instead of drifting silently; every field
|
|
793
|
-
listed there is dispatched to its own checker, so an entry added to the
|
|
794
|
-
list without a checker fails to typecheck rather than passing unchecked.
|
|
249
|
+
required fields and enums, structurally only (it never judges semantic
|
|
250
|
+
adequacy or waives a finding). See
|
|
251
|
+
[`validate-review-report` CLI reference](docs/validate-review-report.md) for
|
|
252
|
+
every flag, exit code, and fence-detection edge case.
|
|
253
|
+
|
|
254
|
+
## Documentation
|
|
255
|
+
|
|
256
|
+
- [Architecture: why this shape](docs/architecture.md): the orchestrator/subagent loop diagram and rationale.
|
|
257
|
+
- [Run contracts](docs/run-contracts.md): the acceptance-baseline contract and `03-decisions.md`'s decision-authority rules.
|
|
258
|
+
- [Install reference](docs/install-reference.md): the agent-led install's conflict check, templates-only re-run rules, and the `knowledge` manifest field.
|
|
259
|
+
- [Harnesses](docs/harnesses.md): the per-harness installed-file list and the honest read-only-posture writeup.
|
|
260
|
+
- [Verification sets](docs/verification-sets.md): the worked `.ai/workflow/verify.json` JSON example.
|
|
261
|
+
- [Role profile reference](docs/role-profile-reference.md): the advisor's escalation triggers and profile/tier re-run behavior.
|
|
262
|
+
- [Model routing reference](docs/model-routing-reference.md): the default-model table, the `--routing` JSON shape, the Codex default routing table, opencode model resolution, and the effort-tiers mechanics.
|
|
263
|
+
- [Operator-level install](docs/operator-install.md): the full `setup`/`apply`/`doctor`/`adopt` command reference.
|
|
264
|
+
- [`validate-review-report` CLI reference](docs/validate-review-report.md): every flag, exit code, and fence-detection edge case.
|
|
265
|
+
- [Curated knowledge bundle](docs/okf/index.md): the OKF-format reference docs for this package's own contracts and mechanics, at `docs/okf/`, the default location a repository falls back to when `knowledge` in `.ai/workflow/manifest.json` is unset.
|
|
266
|
+
- [agentic-coding-playbook](../agentic-coding-playbook): the extended role prompts and organizational guidance this kit's skill references.
|
|
267
|
+
|
|
268
|
+
## Development
|
|
269
|
+
|
|
270
|
+
This package lives in the [agent-dx](https://github.com/LanNguyenSi/agent-dx)
|
|
271
|
+
monorepo, alongside the sibling agentic-coding-playbook package (see
|
|
272
|
+
Documentation above). `npm test` (vitest) and `npm run typecheck` run from
|
|
273
|
+
`packages/orchestrator-workflow`; see the repository root's
|
|
274
|
+
`CONTRIBUTING.md` for the full contributor workflow, including the
|
|
275
|
+
"Releasing okf-kit" order (`test/docs-consistency.test.ts` pins the
|
|
276
|
+
`okf-kit@<version>` this repo's CI installs against the sibling
|
|
277
|
+
`packages/okf-kit` package's version).
|
|
278
|
+
|
|
279
|
+
## License
|
|
280
|
+
|
|
281
|
+
MIT.
|