@orkestrel/scaffold 0.0.77 → 0.0.79

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (158) hide show
  1. package/dist/agents/skills/orkestrel-dispatch/scripts/bench.js +204 -0
  2. package/dist/agents/skills/orkestrel-dispatch/scripts/brief.js +102 -0
  3. package/dist/agents/skills/orkestrel-dispatch/scripts/cite.js +95 -0
  4. package/dist/agents/skills/orkestrel-dispatch/scripts/helpers.js +207 -0
  5. package/dist/agents/skills/orkestrel-dispatch/scripts/launch.js +108 -0
  6. package/dist/agents/skills/orkestrel-dispatch/scripts/login.js +114 -0
  7. package/dist/agents/skills/orkestrel-dispatch/scripts/result.js +108 -0
  8. package/dist/agents/skills/orkestrel-dispatch/scripts/sweep.js +156 -0
  9. package/dist/agents/skills/orkestrel-harden/scripts/discovery.js +196 -0
  10. package/dist/agents/skills/orkestrel-publish/scripts/compare.js +206 -0
  11. package/dist/agents/skills/orkestrel-publish/scripts/pins.js +93 -0
  12. package/dist/agents/skills/orkestrel-publish/scripts/wave.js +458 -0
  13. package/dist/agents/skills/orkestrel-publish/scripts/window.js +188 -0
  14. package/dist/agents/skills/orkestrel-scout/scripts/map.js +300 -0
  15. package/dist/agents/templates/brief.md +55 -0
  16. package/dist/bin/main.js +58 -6
  17. package/dist/bin/main.js.map +1 -1
  18. package/dist/host/AGENTS.md +77 -135
  19. package/dist/host/agents/orchestration.md +147 -998
  20. package/dist/host/agents/skills/enterprise-bootstrap/SKILL.md +2 -2
  21. package/dist/host/agents/skills/enterprise-bootstrap/references/inspection.md +1 -1
  22. package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/SKILL.md +6 -13
  23. package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/agents/openai.yaml +1 -1
  24. package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/references/fleet.md +5 -7
  25. package/dist/host/agents/skills/{orkestrel-build-application → orkestrel-build}/SKILL.md +11 -22
  26. package/dist/host/agents/skills/{orkestrel-build-application → orkestrel-build}/agents/openai.yaml +1 -1
  27. package/dist/host/agents/skills/orkestrel-debrief/SKILL.md +8 -16
  28. package/dist/host/agents/skills/orkestrel-debrief/references/instruction-audit.md +3 -3
  29. package/dist/host/agents/skills/orkestrel-debrief/references/retention.md +13 -13
  30. package/dist/host/agents/skills/orkestrel-dispatch/SKILL.md +61 -0
  31. package/dist/host/agents/skills/orkestrel-dispatch/agents/openai.yaml +4 -0
  32. package/dist/host/agents/skills/orkestrel-dispatch/references/bench.md +25 -0
  33. package/dist/host/agents/skills/orkestrel-dispatch/references/launch.md +32 -0
  34. package/dist/host/agents/skills/orkestrel-dispatch/scripts/bench.ts +259 -0
  35. package/dist/host/agents/skills/orkestrel-dispatch/scripts/brief.ts +110 -0
  36. package/dist/host/agents/skills/orkestrel-dispatch/scripts/cite.ts +115 -0
  37. package/dist/host/agents/skills/orkestrel-dispatch/scripts/helpers.ts +239 -0
  38. package/dist/host/agents/skills/orkestrel-dispatch/scripts/launch.ts +124 -0
  39. package/dist/host/agents/skills/orkestrel-dispatch/scripts/login.ts +123 -0
  40. package/dist/host/agents/skills/orkestrel-dispatch/scripts/result.ts +129 -0
  41. package/dist/host/agents/skills/orkestrel-dispatch/scripts/sweep.ts +157 -0
  42. package/dist/host/agents/skills/orkestrel-falsify/SKILL.md +42 -193
  43. package/dist/host/agents/skills/orkestrel-falsify/references/brief.md +38 -108
  44. package/dist/host/agents/skills/orkestrel-falsify/references/reconcile.md +35 -134
  45. package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/SKILL.md +10 -14
  46. package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/agents/openai.yaml +1 -1
  47. package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/hardening.md +3 -4
  48. package/dist/host/agents/skills/orkestrel-harden/scripts/discovery.ts +228 -0
  49. package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/SKILL.md +15 -23
  50. package/dist/host/agents/skills/orkestrel-journey/agents/openai.yaml +4 -0
  51. package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/captures.md +1 -1
  52. package/dist/host/agents/skills/{orkestrel-polish-surface → orkestrel-polish}/SKILL.md +25 -33
  53. package/dist/host/agents/skills/{orkestrel-polish-surface → orkestrel-polish}/agents/openai.yaml +1 -1
  54. package/dist/host/agents/skills/{orkestrel-polish-surface → orkestrel-polish}/references/capture-harness.md +3 -3
  55. package/dist/host/agents/skills/orkestrel-publish/SKILL.md +33 -20
  56. package/dist/host/agents/skills/orkestrel-publish/references/release.md +39 -0
  57. package/dist/host/agents/skills/orkestrel-publish/references/wave.md +22 -21
  58. package/dist/host/agents/skills/orkestrel-publish/references/window.md +27 -14
  59. package/dist/host/agents/skills/orkestrel-publish/scripts/compare.ts +220 -0
  60. package/dist/host/agents/skills/orkestrel-publish/scripts/pins.ts +114 -0
  61. package/dist/host/agents/skills/orkestrel-publish/scripts/wave.ts +629 -0
  62. package/dist/host/agents/skills/orkestrel-publish/scripts/window.ts +242 -0
  63. package/dist/host/agents/skills/orkestrel-scout/SKILL.md +28 -0
  64. package/dist/host/agents/skills/orkestrel-scout/agents/openai.yaml +4 -0
  65. package/dist/host/agents/skills/orkestrel-scout/scripts/map.ts +352 -0
  66. package/dist/host/agents/templates/brief.md +21 -142
  67. package/dist/host/agents/transports/claude-cli.md +21 -0
  68. package/dist/host/agents/transports/codex.md +38 -159
  69. package/dist/host/agents/transports/cursor.md +16 -65
  70. package/dist/host/claude/AGENTS.md +38 -0
  71. package/dist/host/claude/agents/analyst.md +14 -53
  72. package/dist/host/claude/agents/astra.md +26 -0
  73. package/dist/host/claude/agents/builder.md +14 -30
  74. package/dist/host/claude/agents/checker.md +13 -57
  75. package/dist/host/claude/agents/distiller.md +11 -26
  76. package/dist/host/claude/agents/grok.md +12 -35
  77. package/dist/host/claude/agents/opus.md +14 -30
  78. package/dist/host/claude/agents/orkestrel.md +4 -4
  79. package/dist/host/claude/agents/planner.md +10 -44
  80. package/dist/host/claude/agents/researcher.md +11 -30
  81. package/dist/host/claude/agents/reviewer.md +11 -95
  82. package/dist/host/claude/agents/scout.md +9 -23
  83. package/dist/host/claude/agents/verifier.md +15 -33
  84. package/dist/host/claude/rules/documentation.md +8 -2
  85. package/dist/host/claude/rules/portability.md +7 -1
  86. package/dist/host/claude/rules/quality.md +36 -96
  87. package/dist/host/claude/rules/styles.md +3 -0
  88. package/dist/host/claude/rules/tests.md +6 -3
  89. package/dist/host/claude/rules/workspace.md +21 -15
  90. package/dist/host/claude/rules/writing.md +57 -108
  91. package/dist/host/claude/settings.json +5 -3
  92. package/dist/host/claude/skills/enterprise-bootstrap/SKILL.md +1 -1
  93. package/dist/host/claude/skills/{orkestrel-align-packages → orkestrel-align}/SKILL.md +2 -2
  94. package/dist/host/claude/skills/{orkestrel-build-application → orkestrel-build}/SKILL.md +2 -2
  95. package/dist/host/claude/skills/orkestrel-dispatch/SKILL.md +11 -0
  96. package/dist/host/claude/skills/orkestrel-falsify/SKILL.md +2 -1
  97. package/dist/host/claude/skills/{orkestrel-harden-package → orkestrel-harden}/SKILL.md +2 -2
  98. package/dist/host/claude/skills/{orkestrel-prove-journey → orkestrel-journey}/SKILL.md +2 -2
  99. package/dist/host/claude/skills/orkestrel-polish/SKILL.md +12 -0
  100. package/dist/host/claude/skills/orkestrel-scout/SKILL.md +11 -0
  101. package/dist/host/codex/agents/analyst.toml +14 -31
  102. package/dist/host/codex/agents/astra.toml +25 -0
  103. package/dist/host/codex/agents/builder.toml +13 -20
  104. package/dist/host/codex/agents/checker.toml +13 -27
  105. package/dist/host/codex/agents/distiller.toml +9 -22
  106. package/dist/host/codex/agents/grok.toml +11 -30
  107. package/dist/host/codex/agents/opus.toml +14 -22
  108. package/dist/host/codex/agents/orkestrel.toml +1 -1
  109. package/dist/host/codex/agents/planner.toml +11 -28
  110. package/dist/host/codex/agents/researcher.toml +10 -22
  111. package/dist/host/codex/agents/reviewer.toml +11 -27
  112. package/dist/host/codex/agents/scout.toml +11 -17
  113. package/dist/host/codex/agents/verifier.toml +16 -12
  114. package/dist/host/codex/config.toml +18 -21
  115. package/dist/host/cursor/mcp.json +0 -4
  116. package/dist/host/cursor/rules/orchestration.mdc +12 -20
  117. package/dist/host/dotfiles/mcp.json +0 -4
  118. package/dist/host/dotfiles/oxlintrc.json +7 -0
  119. package/dist/host/guides/probe.md +18 -14
  120. package/dist/host/guides/scaffold.md +147 -83
  121. package/dist/host/guides/test.md +442 -148
  122. package/dist/host/manifest.json +322 -185
  123. package/dist/host/scripts/codex.sh +0 -0
  124. package/dist/host/scripts/cursor.sh +0 -0
  125. package/dist/host/scripts/deps.sh +0 -0
  126. package/dist/host/scripts/ollama.sh +0 -0
  127. package/dist/host/tests/config.test.ts +86 -55
  128. package/dist/host/tests/policy.test.ts +1 -5
  129. package/dist/host/tests/setupPolicy.ts +179 -4
  130. package/dist/src/core/index.cjs +264 -89
  131. package/dist/src/core/index.cjs.map +1 -1
  132. package/dist/src/core/index.d.cts +95 -30
  133. package/dist/src/core/index.d.ts +95 -30
  134. package/dist/src/core/index.js +262 -90
  135. package/dist/src/core/index.js.map +1 -1
  136. package/dist/src/server/index.cjs +55 -9
  137. package/dist/src/server/index.cjs.map +1 -1
  138. package/dist/src/server/index.d.cts +29 -4
  139. package/dist/src/server/index.d.ts +29 -4
  140. package/dist/src/server/index.js +56 -11
  141. package/dist/src/server/index.js.map +1 -1
  142. package/package.json +16 -12
  143. package/dist/host/CLAUDE.md +0 -61
  144. package/dist/host/agents/skills/orkestrel-prove-journey/agents/openai.yaml +0 -4
  145. package/dist/host/agents/transports/claude.md +0 -49
  146. package/dist/host/claude/agents/application.md +0 -36
  147. package/dist/host/claude/agents/sol.md +0 -61
  148. package/dist/host/claude/skills/orkestrel-polish-surface/SKILL.md +0 -12
  149. package/dist/host/codex/agents/application.toml +0 -25
  150. package/dist/host/codex/agents/sol.toml +0 -19
  151. /package/dist/host/agents/skills/{orkestrel-align-packages → orkestrel-align}/references/integration.md +0 -0
  152. /package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/centralization.md +0 -0
  153. /package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/contract.md +0 -0
  154. /package/dist/host/agents/skills/{orkestrel-harden-package → orkestrel-harden}/references/research.md +0 -0
  155. /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/decide.md +0 -0
  156. /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/layer.md +0 -0
  157. /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/statechart.md +0 -0
  158. /package/dist/host/agents/skills/{orkestrel-prove-journey → orkestrel-journey}/references/styles.md +0 -0
@@ -1,1019 +1,168 @@
1
1
  # Orchestration
2
2
 
3
- How agents are dispatched, how long-running work is supervised, and how both are accepted. Every
4
- harness follows this file.
3
+ How agents are dispatched, parallelized, reviewed, and cleaned up. Every harness follows this file. `AGENTS.md` governs code and outranks it.
5
4
 
6
5
  ## Authority
7
6
 
8
- Read in this order before acting:
7
+ Read in order: the user's instruction; `AGENTS.md` and the rules it scopes to your paths; this file; the named skill; the guide. A harness bridge (`.claude/AGENTS.md`, `.codex/config.toml`, `.cursor/rules/orchestration.mdc`) adds mechanics and restates nothing here.
9
8
 
10
- 1. The user's current instruction. It wins over every later item.
11
- 2. `AGENTS.md` and the applicable `.claude/rules/*.md` files. They govern code substance.
12
- 3. This file. It governs agent operation only and cannot weaken the coding contract.
13
- 4. The dispatch-named skill and the references it requires.
14
- 5. The governing guide or spec.
15
-
16
- `CLAUDE.md`, `.codex/config.toml`, and `.cursor/rules/` are bridges. Each points here and adds
17
- only what its harness needs. None of them restates this file.
9
+ ## Engines
18
10
 
19
- A line stays in this file when an executor who is not doing that thing is worse off without it. A
20
- line becomes a skill when it fires on a named trigger and its reader is one agent at one moment.
11
+ | Engine | Job | Posture |
12
+ | --------------------------------- | ----------------------------------------------------- | -------------------------------------------------------- |
13
+ | Cursor Grok | absorb, distill, scout, bounded research | read-only; returns evidence, never decisions |
14
+ | Opus 5.5 | subjective design, design-fit review, implementation | proposes, audits, implements; never accepts its own work |
15
+ | GPT-6 Astra | objective analysis, correctness audit, implementation | proposes, audits, implements; never accepts its own work |
16
+ | Sonnet 5.5, GPT-6 Sol, GPT-6 Luna | mechanical units, gates, conformance, locating | executes a fully specified brief |
21
17
 
22
- Every dispatch tells its executor to read every item after the user's current instruction before
23
- acting.
18
+ The harness's own engine orchestrates: Opus 5.5 in Claude Code, Astra in Codex, Grok in Cursor. The Orchestrator owns the plan, every decision, integration, and acceptance. When the size gate names a review, the review runs on an engine that did not write the work.
24
19
 
25
- ## The engines
20
+ ## Roles
26
21
 
27
- One workflow runs across all providers. Each engine has one job and never takes another's.
22
+ | Role | Job | Claude | Codex |
23
+ | ------------ | ------------------------------------------------------------------ | ------------- | ---------- |
24
+ | `grok` | absorption and distillation on Cursor Grok | sonnet driver | sol driver |
25
+ | `analyst` | objective design argument or correctness audit on Astra | sonnet driver | astra |
26
+ | `astra` | objective, constraint-heavy implementation on Astra | sonnet driver | astra |
27
+ | `planner` | subjective design on Opus | opus | sol driver |
28
+ | `reviewer` | design-fit or correctness review on Opus | opus | sol driver |
29
+ | `opus` | subjective implementation (API shape, naming, guide voice) on Opus | opus | sol driver |
30
+ | `builder` | one fully specified taste-free unit, app layer included | sonnet | sol |
31
+ | `checker` | mechanical conformance against criteria and rules | sonnet | luna |
32
+ | `verifier` | gates and evidence commands, exit-code truth | sonnet | sol |
33
+ | `scout` | locate files, symbols, seams | sonnet | luna |
34
+ | `distiller` | bulk reading when the Cursor bench is dark | sonnet | luna |
35
+ | `researcher` | primary-source research when the Cursor bench is dark | sonnet | luna |
36
+ | `orkestrel` | ecosystem reconciliation over supplied evidence | sonnet | luna |
37
+
38
+ - Name the role and its engine in every dispatch. Reach a role by its own name.
39
+ - A driver launches another provider's CLI and returns the journal path and session id with the result. It never judges or implements. Refuse a bench result that carries no journal path and session id; it ran on the driver.
40
+ - `.agents/transports/<provider>.md` owns each bench's invocation, sandbox, journal, and recovery. `.claude/agents/` and `.codex/agents/` hold one file per role.
41
+
42
+ ## Routing
43
+
44
+ - Route by judgment load. Objective, constraint-heavy, mechanical-precision work goes to `astra`. API shape, naming, and documentation voice go to `opus`. A unit whose correct implementations cannot differ goes to `builder`.
45
+ - Delegate bulk supporting reads to `grok`: terrain, prior art, diff sweeps, scattered sources. Read decision-bearing source yourself. Fall back to `distiller` or `researcher` only when the Cursor bench is dark, and record the fallback.
46
+ - Route gates to `verifier`. A writer's self-reported gate is not evidence.
47
+ - Work directly on a one-file change, a lookup, or a one-line fix. Dispatch when isolation, parallelism, a second engine, or a large read pays for the brief.
48
+
49
+ ## Size gate
50
+
51
+ `AGENTS.md` § Work loop sizes a change and fixes the precedence (large, then medium, then small). This file adds who runs what.
52
+
53
+ | Size | Design | Implementation | Review | Gates |
54
+ | ------ | ----------------------------------------------------------------------------------------------- | ---------------------------------------------- | -------------------------------------------------------------------------------------------------------------- | ---------------------------------- |
55
+ | small | none | direct or `builder` | none; the touched test is the review | touched file |
56
+ | medium | one opinion (`planner` or `analyst`) only when the shape is open | `opus`, `astra`, or `builder` by judgment load | one pass by a reviewer engine that did not write it, on numbered claims about the contract and the risky seams | `verifier` on the touched projects |
57
+ | large | one adversarial round: `planner` and `analyst`, same brief, clean contexts, blind to each other | disjoint units in parallel | one `orkestrel-falsify` round on the integrated result | `verifier` tree-wide, once |
58
+
59
+ - Never run a review on a small change. Never run a second design round on the same brief.
60
+ - A fix round re-audits the repaired claim only. A fix that adopted the reviewer's prescription verbatim closes with a mutation probe (disable the load-bearing line, watch the test fail, restore it). A fix that departs from it gets one pass by the other engine.
61
+ - Three review rounds at one seam end the depth search. Name what the audit is for, dispatch one blind lens per station of the stream the defect moves along, locate the source, and plan from there.
62
+ - An all-confirmed round ends the audit. A round needs an added or repaired claim to attack; reviewer appetite is not a subject.
63
+
64
+ ## Campaign
65
+
66
+ A large change runs these phases once each, in order, and ends when the exit criterion written at the start is met.
67
+
68
+ 1. **Absorb.** Map the terrain with the `orkestrel-scout` skill's `map.ts`, then `grok` distills it and the prior art. In an Orkestrel repository, `verifier` or another tool-capable role collects manifests, lockfiles, installed declarations, and registry readings, and `orkestrel` reconciles them.
69
+ 2. **Design.** One adversarial round: `planner` (subjective) and `analyst` (objective) on one brief, clean contexts, blind to each other. Reconcile into one plan: units, owned files, dependencies, parallel and serial order, acceptance criteria, exit criterion, routing ledger.
70
+ 3. **Implement.** Dispatch units per § Routing and § Parallelism.
71
+ 4. **Integrate.** Apply returned patches to shared files serially. A type, mechanism, or criterion found here is a successor unit, never an integration edit.
72
+ 5. **Audit.** One `orkestrel-falsify` round on the integrated result, one engine that did not write it per lane; add `checker` when the criteria are mechanical.
73
+ 6. **Verify.** `verifier` runs the tree-wide gates once.
74
+ 7. **Re-baseline.** Strike satisfied units, restate transformed ones, add units the exit criterion already required. A re-baseline never moves the exit criterion; that needs the user.
75
+ 8. **Accept.** Report outcome, decisions, evidence, remaining risk. Run `orkestrel-debrief` when the campaign changed process, then prune per § Cleanup.
76
+
77
+ ## Parallelism
78
+
79
+ - Fan out when the work is many small, fully specified, disjoint units: one cheap agent per unit with minimal context. Give the agents the same model, effort, agent type, tools, schema, and working directory when isolation and routing allow it, so their prompts share one cache prefix.
80
+ - Judgment stays narrow: a design or review pass is at most the two lanes the size gate names, each in a clean context, blind to the other, and it runs once. Never fan judgment out further, and never run it as a loop.
81
+ - One writer per checkout. Give parallel writers disjoint owned files; give a writer its own worktree under `tmp/worktrees/` only when two units must write the same files at once. Commit a checkpoint before a writing dispatch.
82
+ - Shared files are report-only; a unit returns an exact patch.
83
+ - Concurrent units validate read-only and scoped to their owned files. Never run tree-wide `format`, lint `--fix`, or `build` beside a live unit.
84
+ - One Grok lane at a time; queue the rest. An empty lane beside a live probe is starvation, not darkness.
85
+ - The Orchestrator's own sweep is a writing dispatch: run it before or after the units, never beside them.
28
86
 
29
- | Engine | Job | Posture |
30
- | --------------- | ----------------------------------------------------- | -------------------------------------------- |
31
- | **Cursor Grok** | Absorption, distillation, scouting, bounded research | Read-only; returns evidence, never decisions |
32
- | **Opus 5.5** | Subjective design, design-fit review, implementation | Proposes, audits, implements; never accepts |
33
- | **GPT-6 Astra** | Objective analysis, correctness audit, implementation | Proposes, audits, implements; never accepts |
87
+ ## Permission floor
34
88
 
35
- - Route each nontrivial implementation unit to Opus or Astra. Objective, constraint-heavy,
36
- mechanical-precision work goes to Astra. API-shape, naming, and documentation-voice work goes to
37
- Opus. Cursor Composer is not an implementation route, and no `composer` role exists.
38
- - Design runs the adversarial pass. § Execution loop's audit step fixes which lanes an audit runs.
39
-
40
- ## Orchestration by harness
89
+ - Read-only roles carry no edit or write tool. The Orchestrator supplies the diff and status to every review dispatch.
90
+ - No role commits, pushes, tags, publishes, installs, or runs a destructive command.
91
+ - No role runs `git checkout`, `git restore`, `git stash`, `git reset`, or `git clean`. Undo an edit by rewriting the text.
92
+ - No role reads, prints, copies, uploads, or packages a secret: `.env*`, `.npmrc`, `auth.json`, keys, tokens, `CURSOR_API_KEY`.
93
+ - A unit writes only under its owned files and its checkout's `tmp/`. The Orchestrator's own briefs, probes, drafts, and records live in the checkout's `tmp/` too, never in a directory outside the repository; the user follows the work through the tree.
94
+ - When a sandbox rejects a write, the unit stops and reports the rejection. It never tries another write mechanism.
95
+ - A unit deletes only the fixture directory it created, by its recorded path. It never removes a `tmp/<lane>` directory, another unit's files, or anything through `sweep.ts --tmp` or `--unit … --delete`; those are the Orchestrator's, after the last live unit exits.
41
96
 
42
- The harness's own engine orchestrates. Run it on the latest version of that model at high
43
- reasoning effort.
97
+ ## tmp layout
44
98
 
45
- | Harness | Orchestrator engine |
46
- | ----------- | ------------------- |
47
- | Claude Code | Opus 5.5 |
48
- | Codex | GPT-6 Astra |
49
- | Cursor | Cursor Grok |
50
-
51
- - The Orchestrator reconciles. No engine reconciles itself or accepts its own work.
52
- - The Orchestrator shares its engine with one lane: Opus in Claude Code, Astra in Codex. That lane is
53
- still dispatched as a separate subagent with a clean context, never run inline.
54
- - In a fix round the auditor is an engine that did not write it. When the writer's engine is the
55
- Orchestrator's engine, the auditor is the other lane.
56
- - Each harness reaches the engines it does not host through its own bridge file. The bridge owns
57
- the invocation mechanics and their cost constraints; this file owns the routing.
58
-
59
- ## The adversarial pass
60
-
61
- The subjective lane and the objective lane run on every design round; an audit round runs the lanes
62
- the execution loop's audit step names, on the same clean-context terms.
63
-
64
- | Lane | Argues |
65
- | -------------- | --------------------------------------------------------------------------- |
66
- | **Subjective** | Shape, taste, naming, ergonomics, design fit, the feel the API must present |
67
- | **Objective** | Correctness, constraints, and what the code and contracts actually permit |
68
-
69
- **A required lane always runs.** Never collapse required lanes into one. Never let an engine's
70
- absence stand in for a required lane.
71
-
72
- Call a lane the round did not dispatch **not run**. `dark` names a bench that cannot round-trip and
73
- names nothing else, so never write it of a lane. A verdict file's recorded reason uses those words.
74
-
75
- ### Clean contexts
76
-
77
- - Dispatch each lane as a fresh subagent. Never run a lane inside the Orchestrator's own context.
78
- - Give each subagent the brief and its evidence slice, and nothing else. Do not carry the
79
- Orchestrator's conversation, its working hypothesis, or the other lane's answer.
80
- - A lane run in the Orchestrator's context is the Orchestrator assessing itself, whatever model
81
- name it carries. The clean context is what makes the lane independent and unbiased, and it is
82
- what keeps the main context at decision level.
83
- - Run the lanes in parallel, blind to each other. Reconcile them yourself.
84
-
85
- ### Engine assignment
86
-
87
- By default Opus 5.5 holds the subjective lane and Astra holds the objective lane.
88
-
89
- Swap the lanes whenever the round needs an engine that is not the one running that lane. Bench
90
- darkness is one trigger and the writer's engine is another: § Execution loop's audit step requires
91
- an auditor that did not write the work, so where Astra wrote the work under audit, give the objective
92
- lane to Opus 5.5 and the subjective lane to Astra, and reverse that where Opus 5.5 wrote it. Both engines
93
- still run, so this is a lane swap rather than a substitution. Record which engine held which lane in
94
- the routing ledger.
95
-
96
- When one engine is unavailable, the remaining engine runs **every** lane — still separate
97
- subagents, still clean contexts, still blind to each other, each told which perspective it holds.
98
- Record the substitution.
99
-
100
- | Harness | Engine unavailable | Runs every lane |
101
- | ----------- | ------------------------------------- | --------------- |
102
- | Claude Code | Astra (Codex bench dark) | Opus 5.5 |
103
- | Codex | Opus 5.5 (Claude CLI dark) | GPT-6 Astra |
104
- | Cursor | Opus 5.5 and Astra (MCP servers dark) | Cursor Grok |
105
-
106
- - Never assign Grok to either lane in Claude Code or Codex. If the remaining native engine is also
107
- unavailable there, the pass cannot run: stop and report rather than substituting Grok.
108
- - Grok takes every lane only in Cursor, and only when Opus 5.5 and Astra are both unavailable.
109
- - Treat a lane that returns no verdicts as a lane that did not run. Re-probe the bench before ruling
110
- on why, per Bench laws rule "One lane at a time per bench". A bench lane reporting that its driver
111
- executed and its engine was never reached is a dark bench, not a result. Record the bench dark
112
- from that report, re-run the lane on the substitute engine from the preceding table, and name in
113
- the routing ledger which lane ran on which engine. Never accept a round with one lane empty.
114
- - Re-read bench liveness at dispatch, not at session start. A bench that probed live can be dark when
115
- the lane launches.
116
-
117
- ## Tedious work goes to Grok
118
-
119
- Absorption, distillation, repository scouting, and bounded primary-source research go to Grok
120
- first, always. Grok returns distilled evidence with `file:line` pointers, never raw dumps.
121
-
122
- Send it any task whose cost is reading: mapping terrain, surveying prior art, sweeping a large
123
- diff, reconciling scattered sources.
124
-
125
- Fall back in this order and record the substitution:
126
-
127
- 1. **Cursor Grok.** The default.
128
- 2. **Luna** (`gpt-5.6-luna`), when the Cursor bench is dark and Codex is available.
129
- 3. **Sonnet**, when both benches are dark.
130
-
131
- - Dispatch `distiller` for absorption and distillation after the ladder steps past Grok, and
132
- `researcher`, `scout`, or `checker` for the job each of those names. `scout` excludes deep
133
- reading and `researcher` excludes repository-scale absorption, so neither takes that step.
134
- - Never route absorption to the Orchestrator itself, even when the Orchestrator is Grok. Keep the
135
- main context at decision level; in Cursor that means a Grok executor session, not this one.
136
- - Never spend Opus 5.5 or Astra on it.
137
- - Grok is read-only, so a writing unit never routes there. Fully specified mechanical writing goes
138
- to `builder` or `application` on the harness's cheap native tier.
139
- - `verifier` runs commands and reports exit codes, so it stays on the native tier too.
140
-
141
- ## Orchestrator and executor
142
-
143
- - The top-level agent is the **Orchestrator**. It owns the goal, plan, decisions, cross-unit
144
- state, integration, and final acceptance.
145
- - A dispatched subagent is an **Executor**. It performs its bounded assignment directly, spawns
146
- nothing, and returns the required distillate.
147
- - Work directly on a typo, a one-line fix, or a single lookup. Orchestrate when isolation,
148
- parallelism, independent review, or substantial context justifies it.
149
- - Dispatch staging, packing, gate-chain invocation, and instrument authorship as units — `builder`
150
- for a fully specified script, `verifier` for its evidence — each with a brief and an audit like
151
- any other unit. The commit and the push stay with the Orchestrator.
152
- - A release instrument installs, commits, or runs a mutating tree-wide command, which the
153
- permission floor bars every role from, so its run is the Orchestrator's tracked command, retained
154
- with its log as the authoring unit's acceptance evidence. An upload loop that must run inside a
155
- one-time code's life with the user at the keyboard stays with the Orchestrator whole.
99
+ Every working file lives under the checkout's gitignored `tmp/`, in the directory its kind names.
156
100
 
157
- ## Roles
101
+ | Directory | Holds |
102
+ | ------------------------------------------ | -------------------------------------------------------------------------------------------------- |
103
+ | `tmp/claude/`, `tmp/codex/`, `tmp/cursor/` | bench briefs, launch scripts, journals, `.err` files, last-message files, login logs, per provider |
104
+ | `tmp/units/` | native unit briefs, reports, claims files, and campaign records |
105
+ | `tmp/probes/` | runtime probes the `probe` Vitest project collects and the `probe` MCP server arms |
106
+ | `tmp/type/` | the type stage's workspace mirror, owned by `@orkestrel/probe` |
107
+ | `tmp/captures/` | screenshots, resolved-style snapshots, and other capture portfolios |
108
+ | `tmp/worktrees/` | worktrees the Orchestrator creates for a parallel writer |
158
109
 
159
- One role set, mirrored per provider. Name the role and state its engine in every dispatch, even
160
- when the role file already pins it.
161
-
162
- | Job | Claude role (`.claude/agents/`) | Codex role (`.codex/agents/`) | Engine |
163
- | ---------------------------------------- | ------------------------------- | ----------------------------- | ----------------------------- |
164
- | Absorption, distillation, scouting | `grok` | `grok` | Cursor Grok (bridge) |
165
- | Creative design and alternatives | `planner` | `planner` | Opus 5.5 (native / bridge) |
166
- | Design-fit review and audit | `reviewer` | `reviewer` | Opus 5.5 (native / bridge) |
167
- | Objective analysis and correctness audit | `analyst` | `analyst` | GPT-6 Astra (bridge / native) |
168
- | Nontrivial implementation (objective) | `sol` | `sol` | GPT-6 Astra (bridge / native) |
169
- | Nontrivial implementation (subjective) | `opus` | `opus` | Opus 5.5 (native / bridge) |
170
- | Bulk reading and evidence distillation | `distiller` | `distiller` | Grok → Luna → Sonnet |
171
- | Bounded primary-source research | `researcher` | `researcher` | Grok → Luna → Sonnet |
172
- | Repository reconnaissance | `scout` | `scout` | Grok → Luna → Sonnet |
173
- | Mechanical conformance evidence | `checker` | `checker` | Grok → Luna → Sonnet |
174
- | Ecosystem evidence | `orkestrel` | `orkestrel` | Sonnet / Terra |
175
- | Fully specified mechanical unit | `builder` | `builder` | Sonnet / Terra |
176
- | Fully specified app-layer unit | `application` | `application` | Sonnet / Terra |
177
- | Gate evidence | `verifier` | `verifier` | Sonnet / Terra |
178
-
179
- - A **bridge** role is a cheap driver whose only work is invoking another provider's CLI. It never
180
- implements, judges, or endorses the result.
181
- - `sol` and `opus` each name one engine on both provider surfaces. The harness decides whether the
182
- role is native or a bridge; the name never does. State the engine in the dispatch anyway.
183
- - A role name and a model alias occupy different fields — a dispatch and a role file's `name` carry
184
- the role, a `model:` pin carries the alias — so the Codex `opus` bridge is a role named `opus`
185
- pinned to the `gpt-5.6-terra` model.
186
- - Give every role a file in the scaffold checkout, under `.claude/agents/` and under
187
- `.codex/agents/`. The role file is where engine, effort, tools, permissions, and charter are
188
- pinned, and the tool allowlist is what makes the read-only floor real. A role with no file has
189
- nowhere to pin either. The requirement is the canon repository's alone: a fleet target holds the
190
- catalog agent and no other role, and a session that dispatches roles starts on scaffold and
191
- attaches the target.
192
- - Reach every role by its own name. Do not rely on a remembered route.
193
- - `distiller`, `researcher`, `scout`, and `checker` are native lanes for jobs that belong to Grok
194
- first. Dispatch `grok` with their brief before using them, and use the native role only after the
195
- ladder has stepped past Grok. Record which step you are on.
196
- - `orkestrel` stays native because it carries the package catalog in its own role file. Sending its
197
- job to a bench means shipping that catalog across, which costs more than the bench saves.
198
- - A transport contract lives in `.agents/transports/`, not in an agents directory. A harness lists
199
- its dispatchable agents from that directory, so a contract that is never dispatched sits outside
200
- it. `.agents/transports/codex.md` is the shared Astra transport contract and
201
- `.agents/transports/claude.md` the shared Opus transport contract. Neither is a route: `analyst`
202
- and `sol` are the named Astra bridges, `planner`, `reviewer`, and `opus` the named Opus bridges, and
203
- each binds its own contract by reference and pins only its route and sandbox.
204
- `.agents/transports/cursor.md` is the shared Cursor transport contract, and both harnesses' `grok`
205
- bridges bind it, because Cursor is native to neither. A contract's home is the provider it carries,
206
- never the harness that reaches it.
207
- - Mirroring is by work class, not filename. A transport contract is provider-specific: the Codex
208
- contract carries the Astra transport the Claude-side bridges follow, the Claude contract carries the
209
- Opus transport the Codex-side bridges follow, and each bridge binds the contract of the provider it
210
- reaches.
211
- - Opus and Astra roles use high effort. Native cheap-tier roles use low or medium. Bridge drivers use
212
- the cheapest tier that can run a CLI.
213
- - Never route orchestration or acceptance across a bridge.
110
+ ## .orkestrel layout
214
111
 
215
- ## Permission floor
112
+ `.orkestrel/` is the campaign folder: tracked, deleted in the acceptance commit, never a home for a rule. A nested directory belongs to one package; a file at the top level belongs to the ecosystem.
113
+
114
+ | Path | Holds |
115
+ | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
116
+ | `.orkestrel/<package>/` | one package's live campaign: its `plan.md`, its routing ledger, the verdicts of its open seams, and its exposure records |
117
+ | `.orkestrel/plan.md` | the live fleet campaign's plan: goal, exit criterion, packages, layer order, and the units each package owes |
118
+ | `.orkestrel/ledger.md` | the fleet campaign's routing ledger: every dispatch, engine, and bench substitution across packages |
119
+ | `.orkestrel/release.md` | the live release wave: its layers, each package's bump ruling with its evidence, the round each package landed in, and standing readings |
216
120
 
217
- Every role honours this floor. No dispatch may widen it.
218
-
219
- - Agents run to completion. Constrain only genuine security or destruction risk. Do not gate
220
- routine work behind approval prompts or turn budgets.
221
- - Read-only roles carry no `Edit` and no `Write`. The tool allowlist is the guarantee. Because
222
- those roles cannot inspect the tree by writing to it, the Orchestrator supplies the actual diff
223
- and status evidence in every review dispatch.
224
- - `verifier` has no edit or write tools and never fixes a failure.
225
- - Run one writing role per checkout, on disjoint checkouts, each dispatched from a clean committed
226
- baseline and each owning disjoint files. In a single checkout that is one writer at a time.
227
- - Treat every shared file as report-only.
228
- - No role commits, pushes, tags, publishes, installs dependencies, or runs a destructive command.
229
- - No role runs `git checkout`, `git restore`, `git stash`, `git reset`, or `git clean`. Each discards
230
- a working-tree change silently. A role that must undo its own edit undoes exactly that edit.
231
- - A dispatch that has a unit plant a line to prove an instrument can fail names a file the unit under
232
- verification did not touch, and names how the plant is removed. Check the tree's status before
233
- choosing the file.
234
- - No role reads, prints, copies, uploads, or packages a secret — `CURSOR_API_KEY`, Codex auth
235
- files, `.env*`, `.npmrc`, `auth.json`, keys, or tokens.
236
- - Concurrent executors never run tree-wide `format`, lint `--fix`, or `build`. They validate
237
- read-only and scoped to their own files.
238
- - Keep hooks light. A Stop hook may run cheap changed-file verification such as `git diff --check`.
239
- It never duplicates the gate suite; gates belong to `verifier`.
240
-
241
- ## Context and decomposition
242
-
243
- - Keep the main context at decision level. Send large reads, repository scans, raw logs, diffs,
244
- and exploratory sweeps to `grok` and consume the distillate.
245
- - Settle a behavioural question by running it, including your own. This binds before a brief is
246
- written, not only after a report arrives, and it needs no disagreement to trigger it. A round of
247
- deliberation that a probe would have ended is the expensive failure, and it is invisible because
248
- it feels like rigour.
249
- - Check an assumption before it enters the plan. An unverified belief the Orchestrator states
250
- becomes a fact for every unit downstream, and every one of them inherits the error.
251
- - Break a repeating frame deliberately. When several rounds against one subject keep finding the
252
- same class of defect through a new door, the search is following the frame rather than the
253
- defect. Bound the scope, then fan out independent lenses over disjoint slices in one pass.
254
- Parallelism is worth more here for the framing it breaks than for the wall-clock it saves.
255
- - When the recurring class has a direction — each fix relocates it along one stream: a dependency
256
- chain, a data path, a call chain — the source sits elsewhere on that stream, and deepening the
257
- current station cannot reach it. Switch from depth to breadth: fan probes over the stream's
258
- stations in parallel, blind and clean-contexted, as far up and down as the stream runs, to locate
259
- the source. Then plan downstream from the source with the sweep's map of how far the defect
260
- reaches, so the remaining work carries a measured bound instead of an open count of rounds. The
261
- seam budget in `.claude/rules/quality.md` § Rounds and verdicts states when this fires.
262
- - The subjective and objective lanes are the adversarial pass's FLOOR, not its shape. Where a
263
- subject has more seams than that pass can attack, fan out one lens per seam over disjoint slices,
264
- keep every lens blind and
265
- clean-contexted, and number every slice's claims in one shared sequence. Change the lenses in a
266
- successor round rather than repeating them.
267
- - Decompose by required context and independently verifiable acceptance criteria, not by task type.
268
- - Send instructions down fully specified. Return findings smaller than the context consumed.
269
- - Parallelize independent work. Serialize dependencies and shared-file contention.
270
- - The Orchestrator owns the plan and every final decision. Design engines propose, writers execute,
271
- auditors advise.
272
-
273
- ## Writing concurrency
274
-
275
- Concurrent executors share a filesystem unless isolated. Follow these rules to prevent clobbered
276
- edits, formatter and build races, cache phantoms, and validation cross-talk.
277
-
278
- 1. Serialize writers as § Permission floor states: one per checkout, checkouts disjoint. Commit a
279
- checkpoint before each writing dispatch so git is the rollback mechanism.
280
- 2. Assign disjoint owned files plus explicit shared and off-limits files.
281
- 3. Keep shared files report-only. Executors return exact patches for serial integration.
282
- 4. Restrict concurrent executors to read-only, scoped validation. A tree-wide result may contain a
283
- sibling's in-flight failure, so an executor reports only its owned scope.
284
- 5. Give concurrent audit lanes worktree isolation whenever the campaign is uncommitted. Lanes
285
- sharing one working tree contaminate each other's readings in both directions.
286
- 6. After integration, clear shared caches if needed, then have one independent `verifier` run the
287
- authoritative tree-wide sweep. A writer's self-report never establishes green.
288
- 7. The Orchestrator's own sweep is a writing dispatch and queues behind the units that own those
289
- files. A script that fixes one thing across every target is the most direct way to break the
290
- serialization rule, because it does not feel like a dispatch — nobody was named, no brief was
291
- written, and it finishes in seconds. It still writes into trees a live unit owns, and a unit
292
- whose brief it invalidates will repair the same drift the other way and report a state that is
293
- already false. Run it before the units, or after them, or send the decision to every unit in
294
- flight per the mid-campaign rule under **Dispatch anatomy**. Never beside them.
295
- 8. A fleet pass that records a per-target status commits only the targets it recorded green. Reading
296
- "is the tree dirty" instead of "did this target pass" pushes a red target the moment one exists,
297
- and a flake makes that look like it worked. Refuse the failed row, name it, and re-run it alone
298
- before deciding what it was.
299
- 9. Run a fleet pass in slices that report as they finish, never as one block. A block hides its first
300
- failure behind every target that follows, so the failure surfaces after the work it exists to
301
- stop. A slice hands control back while most of the fleet is still unstarted.
302
- 10. Re-run a timing or resource failure alone before believing it. Concurrent slices, builds, and
303
- suites make a container miss deadlines it meets when idle, so a red result under load is a
304
- question rather than an answer. A unit re-running the file alone is not alone: its own exec,
305
- code-mode host, and sandbox stay resident throughout, and on a small container that residue
306
- alone misses a deadline the same file meets on an idle one. So the deciding re-run belongs to
307
- the Orchestrator after the unit exits, never to the unit, and a timing failure a unit cannot
308
- clear is carried to that reading rather than diagnosed by the unit. This makes a writer's own
309
- gate evidence systematically pessimistic on timing, which is another reason the independent
310
- `verifier` runs the authoritative gates.
311
- 11. While any unit is live, the Orchestrator's own instruments go in its scratchpad, never in the
312
- subject repository's `tmp/`. A probe that both writes and deletes inside the subject tree can
313
- remove a file it did not create, and a cleanup keyed to a caller-supplied path list is how. The
314
- same directory is where dispatched units build their instruments, so removing it destroys a live
315
- lane's work.
316
-
317
- ## Execution loop
318
-
319
- At session start, before planning, probe bench liveness and plan routing against the result. Resolve
320
- each CLI first (`codex --version`; `agent --version`, falling back to `agent.cmd --version`), then
321
- run the bench's authentication-state check where it exposes one. Neither answer is liveness. A
322
- version string proves the binary is installed, and an authentication-state check reads stored
323
- credentials, so both pass while the account is out of quota, while the routed model is unavailable to
324
- it, while the server has already revoked the credential the check read, and inside a sandbox
325
- with the network denied. Record a bench live only on a bounded round-tripped model call that came
326
- back, and record what came back beside the routing decision. Probes are read-only, and the role file
327
- owns each bench's exact probe.
328
-
329
- The local steps still run, because they route the recovery rather than decide the verdict: an
330
- unresolved CLI is an install problem, a failed authentication-state check starts the login ladder
331
- in Recovering a dark bench, and a bench that passes both and still cannot round-trip is dark for a
332
- reason no local check can see. Record every dark bench with its fallback and the lane substitution
333
- it forces, and never absorb one silently. A readiness script reports readiness and performs no
334
- model call, so the round trip belongs to the Orchestrator's own probe or to the bridge carrying
335
- the unit, never to the hook. Liveness also expires: a dispatch that fails on quota, model access,
336
- or the network is a fresh liveness result rather than a unit-level fault, so record the bench dark
337
- from there and re-plan the lane instead of re-dispatching against a session-start answer that no
338
- longer holds.
339
-
340
- 1. **Absorb.** Dispatch `grok` for terrain, prior art, and the reading the decision needs. In an
341
- Orkestrel repo dispatch `orkestrel` alongside it for live package state. Skip only when the
342
- ground is already known.
343
- 2. **Design adversarially.** Run the adversarial pass on one design brief: `planner` for
344
- the subjective lane and `analyst` for the objective lane. Reconcile them yourself into one plan:
345
- units, dependencies, ownership, parallel and serial order, acceptance criteria, risks.
346
- - Surface the plan before dispatch, including a routing ledger naming each unit's role **and**
347
- engine. Routing a unit to a Claude-native agent when its work class belongs to a bench —
348
- reading-heavy to Grok, objective audit or objective implementation to Astra — without a recorded
349
- bench-dark deviation is a dispatch deviation.
350
- - State the goal's exit criterion beside the units: the enumerated capabilities whose closure
351
- ends the campaign, each to end implemented, repaired, retained, or intentionally excluded on
352
- evidence. A plan that names work but not its end can only be abandoned, never finished.
353
- 3. **Implement.** Route each nontrivial objective unit to `sol` and each nontrivial subjective unit
354
- to `opus`, in the checkout the unit writes, one writer per checkout. Route a fully specified
355
- taste-free unit to `builder`. Never route implementation to an engine the unit's judgment load
356
- exceeds.
357
- 4. **Integrate.** Evaluate each distillate against its acceptance criteria, apply shared-file
358
- patches serially, and route cross-cutting findings. Integration applies exact returned patches
359
- and mechanical conflict resolution only. A new type, mechanism, behavior, or acceptance
360
- criterion discovered at integration is a successor brief routed to a writer, never an
361
- integration edit.
362
- 5. **Audit adversarially.** Audit every nontrivial implementation with the objective lane and the
363
- subjective lane — `analyst` and `reviewer`, the way the design step names its lanes — at least
364
- one of them on an engine that did not write the work. Dispatch `checker` in addition when the
365
- acceptance criteria are mechanical — counts, paths, parity rows, scope honesty — never in place
366
- of a lane. A round that runs fewer lanes than its brief names, or omits the checker its criteria
367
- call for, records the deviation in its verdict file with that round's own reason, never a
368
- template sentence.
369
- - State the audit's subject as numbered falsifiable claims and require per-claim verdicts with
370
- evidence, per the Falsification law in `.claude/rules/quality.md` and the `orkestrel-falsify`
371
- value set, unless the dispatch names a different skill that fixes another.
372
- - Require every lane, before confirming a claim **about a proof**, to name the mutation that would
373
- make that proof fail and to say whether its assertions distinguish that mutation from the
374
- passing case. Put the answer in the evidence. Describing what a proof does is not ruling on it:
375
- a proof that cannot fail reads exactly like one that can, and structural reading confirms both.
376
- This is the discipline that separates a lane which runs the disputed thing from one which reads
377
- it, and the reading lane is the one that confirms a hole.
378
- - In a fix round, give the unit to an auditor engine that did not write it.
379
- - Run the `orkestrel-falsify` skill for multi-round audits. It owns the brief anatomy, the
380
- successor-brief rule, the verdict shape and its single terminal line, and the reconciliation
381
- discipline.
382
- - Reconcile the lanes that ran. Drop, on the record, any finding no lane can substantiate.
383
- - Where two lanes contradict each other, check first that each one's citations resolve in the
384
- file it names, before weighing the arguments. A line number past a file's end settles the
385
- disagreement in one command, and a lane that confirmed a claim on evidence that cannot exist
386
- has its verdict on that claim discarded rather than balanced against the other lane's. Discard
387
- the claim, not the lane: verify a sample of that lane's other citations, and keep the readings
388
- that hold.
389
- 6. **Verify.** Have one independent `verifier` run the authoritative gates.
390
- 7. **Re-baseline.** Reconcile the remaining plan against what the phase revealed, before dispatching
391
- the next one.
392
- - Rule on every remaining unit: **satisfied**, the phase closed it, strike it; **transformed**,
393
- the intent stands and the work changed, restate it; **added**, the phase revealed a fork the
394
- plan did not consider; **unchanged**.
395
- - Redraw the dependency order. Strike a unit whose subject a later unit deletes. State the new
396
- prerequisite of a unit that acquired one.
397
- - Walk the remaining units once and ask of each whether what landed still supports it. A
398
- decision taken inside a unit can remove a later unit's foundation, and the unit that took it
399
- cannot see that.
400
- - Re-baseline when a probe overturns a decision the plan rests on, not only at a phase boundary.
401
- A measurement that falsifies your own reconciliation changes which units run.
402
- - Record what changed and why. An unrecorded re-baseline cannot be audited, and the next one
403
- re-derives it.
404
- 8. **Accept.** Decide, then report outcomes, decisions, evidence, and remaining risk concisely.
405
- When the design step's exit criterion is met and the gates are green, accept. The next goal is the
406
- deliverable.
407
-
408
- ### Re-baselining is not rescoping
409
-
410
- `.claude/rules/quality.md` fixes the enumerated scope when work begins and forbids reopening it. A
411
- re-baseline changes which units run. It never changes the goal's exit criterion.
412
-
413
- Strike a unit because the phase satisfied it, never because it became inconvenient. Add a unit
414
- because implementation revealed work the exit criterion already required, never because an engine
415
- thought of something else worth doing. A re-baseline that moves the exit criterion is a rescope, and
416
- that needs the user.
121
+ - Write a file about one package under that package's directory and never at the top level; write a file about two or more packages at the top level and never under a package.
122
+ - Delete a package directory in that campaign's acceptance commit and a top-level file when the wave or fleet campaign it records closes. `sweep.ts --report` lists both.
123
+
124
+ ## Dispatch
125
+
126
+ `orkestrel-dispatch` owns the brief template, launch mechanics for native and bench units, long-running commands, liveness, and recovery. Load it before a bench lane, a campaign unit, or any command that outlives the turn. The floor that follows binds every dispatch.
127
+
128
+ - A small native unit's brief is the dispatch message: scope, law references, expected result. Write a brief file with the dispatch skill's `scripts/brief.ts` at `tmp/units/<unit>-brief.md` (native) or `tmp/<bench>/<unit>-brief.md` (bench) when the unit is a bench lane, belongs to a campaign, or may need re-running. A re-run gets `<unit>-brief-<n>.md`; never edit a launched brief.
129
+ - State in every brief: role and engine, one objective, the evidence slice, owned and off-limits files, the acceptance criteria cheap-first, the return shape, the deviation contract, and that the executor performs the work itself and spawns nothing.
130
+ - Capture a campaign unit's report to `tmp/units/<unit>-report.md`, or read the bench's last-message file.
131
+ - Launch every multi-minute command as a harness-tracked background command through `node .agents/skills/orkestrel-dispatch/scripts/launch.ts` under a cap sized from prior runs, and read a bench answer with that skill's `scripts/result.ts`. Every launcher, probe, and instrument is a TypeScript file run by `node`, never a shell or Python script. Read liveness from the artifact the work produces. Kill only a recorded process id, never a pattern, and confirm the process and its descendants are gone before another writer takes its files.
132
+ - Probe a bench with the dispatch skill's `scripts/bench.ts` before its first lane, and re-read liveness at every dispatch after a failure.
133
+ - Send a mid-campaign decision to every live unit whose brief it invalidates.
417
134
 
418
135
  ## Deviation protocol
419
136
 
420
- Stop when a conflict prevents the primary objective or requires an unowned change. Resolve an
421
- ancillary choice within the owned scope, record the choice, and continue.
422
-
423
- Every charter references this section rather than restating it. A charter keeps only a stop
424
- condition its own route owns, such as a misrouted unit it must refuse.
425
-
426
- When a writer stops:
427
-
428
- 1. The writer reports: expected, found, exact evidence, done or not done, and at most one
429
- short hypothesis. It does not investigate, improvise, or alter the plan.
430
- 2. The Orchestrator triages:
431
- - obvious correction → tighten and re-dispatch;
432
- - missing mechanical evidence → dispatch `verifier`;
433
- - unknown terrain → dispatch `grok` with the report and the plan slice;
434
- - unknown design or root cause → dispatch `planner` and `analyst` on the question.
435
- 3. The Orchestrator decides, updates the plan, and re-dispatches.
436
-
437
- Workflow failures use the same ladder. Do not absorb their raw logs into the main context.
438
-
439
- ## Dispatch anatomy
440
-
441
- ### Native first
442
-
443
- Launch a model through the running harness's own mechanism whenever that harness hosts it: Claude
444
- subagents and workflows in Claude Code, Codex-native agents in a Codex session, Cursor-native
445
- sessions in Cursor. MCP and CLI transports exist only to reach a model the harness does not host.
446
- Never route a native model through its own CLI or an MCP loopback.
447
-
448
- - Use a single-agent dispatch when later control flow depends on the previous result.
449
- - Use a workflow for a known deterministic fan-out, staged pipeline, or loop. Serialize writing
450
- nodes; never run concurrent writers in the tree.
451
- - Name a role and its engine in every node.
452
-
453
- The harness bridge names the concrete mechanism for each of these.
454
-
455
- ### Every dispatch is a file before it is a launch
456
-
457
- - Write the brief to a file under `tmp/`, named for its unit, before launching the unit, whatever
458
- engine executes it. A brief composed only inside a launch argument cannot be corrected, resumed,
459
- or re-run after that call ends. A native unit's pair is `tmp/units/<unit>-brief.md` and
460
- `tmp/units/<unit>-report.md`. A bench unit's pair sits under the bench directory its role file
461
- names, beside the journal.
462
- - Write the unit's returned report in the SAME action that commits its code, never afterwards. A
463
- commit message states what changed; the report states what the unit measured, what it decided, what
464
- it could not close, and which of its own claims it flagged. An auditor's subject is the report,
465
- so a report living only in the Orchestrator's context stops the next lane on arrival.
466
- - Capture the unit's returned report to a file beside its brief under the same unit name, so a
467
- unit's instruction and its outcome are one pair on disk.
468
- - Name, inside a report that rests on a bench lane, that lane's journal path and session id, so
469
- the provenance survives the journal's sweep.
470
- - Re-run a unit with a successor brief, never with an edit to the brief it already ran. Name the
471
- successor `<unit>-brief-<n>.md`, state in it what changed and why, and leave the original in place
472
- unedited. A fix round's brief names the findings it carries and where each came from.
473
- - Name a corrected unit's effective brief and report and the pair they supersede before that unit
474
- integrates. A unit whose correction landed as a serial patch or a direct reconciliation, with no
475
- successor pair on disk, cannot be re-run from what the campaign kept.
476
- - When you copy a previous round's or unit's launch artifact to derive this one's, rewrite every
477
- field that names the subject, not only the ones the edit was about. That covers each brief's lane
478
- focus, its evidence and report paths, its subject files, and the path it gives for the law; the
479
- claims file's subject line; each driver and watch script's header, journal, and output paths; and
480
- the workflow's `meta.description` and every node's label. Read the derived copy start to finish
481
- against the round it launches before launching it. A field carried over from the source names a
482
- different subject and still reads as deliberate, so the lane works around it or rules on the wrong
483
- file, and the retained record attributes that to the round.
484
- - Write an audit round's numbered claims to `tmp/audit/<unit>-audit-claims.md` and point every lane
485
- of the round at that one file, so "both lanes ran the same brief" stays checkable after the round.
486
- Retain it as `.orkestrel/<package>/<unit>-audit-claims.md`, beside the round's verdict.
487
- `.agents/skills/orkestrel-falsify/references/brief.md` owns what the claims say.
488
- - Write the round's verdict to `.orkestrel/<package>/<unit>-audit-verdict.md`. That file is where
489
- the audit step records a lane or a checker that did not run.
490
- - Read the copy the executor will open, not the one you wrote. A brief written in the orchestrator's
491
- repository and staged into the subject's checkout so a `-C` invocation can reach it is a second
492
- file, and staging can rewrite a path or drop a clause. The executor rules on what it opens, so a
493
- staged copy whose facts are false stops a unit that was correctly briefed. Verify the staged path
494
- and its load-bearing facts before launching, and stage into a scratch directory the subject tree
495
- ignores rather than into the checkout root.
496
- - Before dispatching a successor or accepting a round, open every file the effective brief names
497
- and confirm it resolves from the executor's root. Refuse the transition when one does not.
498
- - Check the launch argument's own file list the same way. A driver prompt names records beside the
499
- brief it points at, and those names are part of the effective instruction even though they sit
500
- outside the brief file. Staging the brief and its terrain is not enough: list every path the prompt
501
- mentions and confirm each one resolves, because the unit stops on the one that does not and the
502
- stop is charged to the round rather than to the launch.
503
- - Send a decision taken mid-campaign to every unit already in flight whose brief it invalidates. An
504
- executor cannot see a change made after it was dispatched, so it writes the state its brief
505
- described and the defect surfaces as its own.
506
- - Retention is uniform for every unit, whatever engine ran it, including an Orchestrator-owned
507
- integration, fix, probe, or capture unit: copy the brief, the returned report or distillate, the
508
- audit verdict, the exact executed script or instrument, and the acceptance evidence into
509
- `.orkestrel/<package>/` as the unit is dispatched and as it returns, then sweep only the `tmp/`
510
- launch copies. Rewrite every `tmp/` path inside a copied artifact to the retained path it now
511
- names, in the same action that copies it. A retained file naming a launch copy resolves to
512
- nothing after the sweep, and the successor, the claim list, and the staged authority always do.
513
- Retention covers every lane of an audit round, including the lane whose brief produced a
514
- deviation: a round with a retained report and no retained brief cannot be re-run. A capture
515
- claim's instrument is acceptance evidence; the frames may be swept once the record transcribes
516
- them, because the committed instrument re-produces the film. The **Bench laws** rule "Ephemeral
517
- streams, durable records" owns journals and points here for everything durable.
518
- - Name a retained log with the `<unit>.log.txt` pattern, never with a bare `.log` suffix, which the
519
- root `.gitignore` file ignores.
520
- - Promote anything that must outlive the campaign into a durable artifact before the sweep — a
521
- commit message, a guide, a rule, a retrospective. What is only in a swept file did not survive,
522
- and a debrief that must quote the record verbatim has nothing to quote.
523
- - Land a process rule stated as binding mid-campaign in the owning rule or contract file in the
524
- same commit that states it. A campaign artifact is evidence, never a rule's home.
525
-
526
- ### Where campaign artifacts live
527
-
528
- - Put every campaign artifact in the **orchestrator's** repository under `.orkestrel/<package>/`,
529
- named for the package the campaign is about.
530
- - Give a campaign spanning several packages one shared `.orkestrel/campaign/` folder instead, so the
531
- wave's plan, routing ledger, and verdicts sit together rather than split across the packages they
532
- rule on.
533
- - Never put them in the package they are about. A published package's tree is its product.
534
- - Where the orchestrator's repository is itself a subject package, keep `.orkestrel/` as the
535
- artifact home and stage every landing chain by path, never with `git add -A`. Run that
536
- checkout's own CLI through the built `node dist/bin/main.js` entry, never through the `npx`
537
- launcher, and record its audit reading beside the landing instead of gating the landing on it.
538
- Commit the records before the landing's commit step, never during it.
539
- - Claim nothing outside `.orkestrel/` unless Orkestrel scaffold mandates it. Everything Orkestrel
540
- owns in a consumer's tree lives beneath that folder, so a convention can be settled there without
541
- colliding with a convention that is not Orkestrel's.
542
- - Keep the campaign narrative and every ruling in the durable artifact that owns it — the guide for
543
- product truth, a rule or role file for process truth, the commit message for the decision itself.
544
- Use `ROADMAP.md` only where the repository already keeps one.
545
- - Prefer a mechanism that recomputes a fact over a document that records it. A document recording
546
- live state is stale from the moment it is written, and the next campaign reads it as current.
547
- Where the fact can be derived, derive it: the fleet's publish order lives in the catalog table
548
- `scaffold catalog` regenerates, not in a written order anyone has to remember to update.
549
- - Prune the campaign folder in a commit at acceptance. Git history is the archive; the working tree
550
- is the workspace. `.agents/skills/orkestrel-debrief/references/retention.md` owns the procedure —
551
- the checks that close the prune, the artifact locations it sweeps, and the promotion record
552
- the commit message carries. Run it before deleting anything.
553
-
554
- ### Required sections
555
-
556
- - **Role and engine.** The named role and its explicit engine.
557
- - **Objective.** One concrete outcome.
558
- - **Context.** The evidence slice, paths, decisions, `AGENTS.md`, applicable rules, the skill name
559
- and its required references (or an explicit none), and the guide or spec. Include the host
560
- environment facts the unit will hit — the shell, path, and network constraints its commands run
561
- under — because an executor that rediscovers them pays for the discovery in round trips. State
562
- every standing condition the same way: a file expected to be dirty, a command known to fail, a
563
- shim the shell blocks. A condition the brief leaves unnamed comes back as a deviation report
564
- about something you already knew.
565
- - **Unknowns.** What the Orchestrator does not yet know that the unit needs, named as unknown, with
566
- how the unit reports back on it. A brief that cannot be fully specified says so instead of
567
- shipping a guess the executor would have to invent an answer around.
568
- - **Scope.** Owned files, shared and off-limits files, allowed tools, permission limits.
569
- - **Execution.** State that the executor performs the assignment directly and spawns nothing. Put
570
- it in every brief; an executor deep in a task does not re-read this contract. Write the sentence
571
- for the reader the transport actually delivers it to, because "directly" names a different action
572
- on each side of a bridge. To a native subagent or a bench engine reading the brief inside its own
573
- CLI, it means do the work yourself. To a bridge driver, whose entire assignment is the launch, it
574
- means carry the brief across unaltered and return the journal — the engine behind the CLI is not a
575
- subagent the driver is spawning, and a driver told to work directly answers from its own engine
576
- instead. That answer reads normal and its only tell is the missing journal, so pair this sentence
577
- with **Bench laws** rule "Journal first" and refuse a bench result whose journal path and session
578
- id are absent.
579
- - **Output.** The exact distilled return shape. No process diary.
580
- - **Deviation contract.** Point the writer at § Deviation protocol and scope it: name the ancillary
581
- choices this unit settles itself — where a paragraph sits, which heading a section takes. An
582
- unscoped contract stops a unit over a detail it was equipped to settle.
583
- - **Acceptance criteria.** Independently checkable completion conditions.
584
- - **Review evidence.** What the subject type requires, per the table in `orkestrel-falsify`. For a
585
- code change that is the actual diff and the actual status output; omitting either is a dispatch
586
- deviation. For any claim about a rendered or externally driven surface, the capture portfolio is
587
- the review input and source is corroboration. A subject may occupy more than one row — a ruling
588
- whose fixes already landed as edits is both a proposal and a code change — and it gets the
589
- evidence of every row it occupies.
590
-
591
- ### Check the brief before you send it
592
-
593
- Fill `.agents/templates/brief.md`, keeping its section and row headings verbatim, so a row you
594
- cannot close is visible as a heading with a named unknown under it rather than as an absence nobody
595
- can see. Add a section the unit needs; never drop one. Then run this checklist against what you
596
- filled.
597
-
598
- - Name the executor that will open the brief, and write the transport for that reader. A bridge
599
- driver and the bench engine inside that driver's CLI need opposite instructions.
600
- - Paste the command and its output behind every factual claim — paths, counts, registrations, file
601
- existence. Name the scope any search covered, and check a fact against the code rather than
602
- against another artifact that states it.
603
- - Search the campaign's own record for every fact a brief asserts, and reconcile any disagreement
604
- before dispatch. A brief that contradicts a measurement the campaign already took is the most
605
- expensive kind of wrong fact, because the scope read cannot save you: it checks the brief against
606
- the tree with the brief's framing in hand, so it reproduces the error as often as it catches it.
607
- The unit reading the record cold is the only reader positioned to refuse, and by then the round is
608
- spent. Grep the plan and the retained reports for the subject, and not the tree alone.
609
- - Better, give each measurement one home and keep the brief out of it. Put the measurements in a
610
- terrain record the brief names and stages beside itself, and write the brief as rulings and
611
- obligations that point at it. A measurement restated in a brief is a second copy that can drift
612
- from the first, and the drift is invisible because both artifacts look authoritative. Tell the unit
613
- which artifact wins when they disagree, and to stop rather than resolve it.
614
- - Cite a site by its symbol in any artifact meant to outlive a landing, and name a line only as
615
- approximate ("around line 40"). A record written before one unit lands and read after another does is the
616
- normal case, not the exception: a line citation that was exact when measured goes stale the moment
617
- something is inserted above it, and the reader cannot tell a stale number from a wrong one. Name
618
- the function, the mixin, the case title, or the surrounding construct, and tell the reader to
619
- locate it by that.
620
- - Take every measurement under the conditions the unit runs in, or have the unit take it before
621
- doing anything else.
622
- - Ask what the change does to every fact you measured, and fix each criterion to the state the unit
623
- finishes in.
624
- - Scope the change by the files its result makes **false**, not by the files that declare the thing
625
- changing: the test asserting the reversed behaviour, the fixture carrying the raised value, the
626
- golden digest over generated output, the consumer script naming the removed union member.
627
- - Derive that set by running the suite. Where you cannot run it, name the search's bound in the
628
- brief so the unit re-derives the set.
629
- - Find the enumerating assertions by searching for the population's EXISTING members, never by
630
- reasoning about which files look relevant. A literal set or list naming every member of a growing
631
- population goes false the moment a unit adds one, and it can sit in a file whose name suggests it
632
- holds only machinery. Grep the tree for a member the population already has; every file that comes
633
- back is a file the unit must own. Reasoning finds the obvious site and misses the rest, and the
634
- miss surfaces as the unit's stop.
635
- - Answer any question of the form "can this tree do X" the same way: search for X already being done,
636
- never by inspecting the interface that would do it. Grep for the behaviour, the export, the staged
637
- preference, the driven surface. A sibling already doing it is the answer, and its call site is the
638
- pattern the unit copies. An interface's surface is evidence about that interface and never about the
639
- tree, so a missing control there proves nothing — the mechanism is as often a function an installed
640
- package exports as a method on the object in hand. Record the answer as unreachable only after a
641
- search for the thing itself came back empty, and name the pattern that search used. A brief that
642
- asserts a capability is absent tells the unit to stop looking, so this error costs the round even
643
- when the tree is correct.
644
- - Diff the previous unit's actual status against this brief's owned set whenever a series of units
645
- adds members to one population. A grant the previous sibling needed and this brief omits is the
646
- likeliest gap, because each brief is written from the plan rather than from what the last unit
647
- touched.
648
- - Grant a behaviour with the tests that pin it, a constant with every fixture and expectation
649
- derived from it, and a template change with the materialized copy the package generates from it.
650
- - Grant the file a brief tells the unit to copy a pattern from, wherever a standing gate forbids
651
- duplicating that pattern. Naming a pattern to follow and withholding the file it lives in
652
- instructs the unit into a gate failure it cannot fix inside its scope: the honest fix touches both
653
- copies, and extracting one side alone breaks the same rule from the other direction. The unit then
654
- stops, correctly, and the round is spent on a contradiction the brief carried.
655
- - Read each criterion against the off-limits list, line by line. Grant the file a criterion needs or
656
- strike that criterion. A file the change will break that appears in neither list is unscoped.
657
- - Where a scope line names the `tests/**` or `src/**` glob, name the paths the `scaffold repair`
658
- command restores as off-limits in the same sentence. Bound a criterion wider than its Sites by
659
- the brief's scope.
660
- - Scope a unit that changes a mechanism to own the prose describing it. Where a brief scopes that
661
- prose out, name the carrier and dispatch it before the change ships.
662
- - Give a small unrelated obligation its own unit.
663
- - Name the property the unit must change, and stop. Record an expected consequence as an
664
- observation, never as a second criterion.
665
- - Order the criteria cheap-first, so an unreachable one cannot hide a typecheck or a lint criterion
666
- behind it. Where the change edits a file the repository vendors or digests, the regeneration step
667
- precedes every gate that reads the generated artifact.
668
- - Never make a timing-sensitive or whole-suite gate result a criterion for a unit running inside its
669
- own exec. Name it as an observation the unit reports with its own reading, and take the
670
- authoritative run yourself after the unit exits. A scoped run over the unit's owned files stays a
671
- legitimate criterion.
672
- - Check the brief's output mechanism and its verification method against the executor's tool
673
- allowlist. A read-only lane writes no report file and runs no probe, so hand it the rendered
674
- evidence instead.
675
- - Keep the brief's control identifiers inside the brief, and say in the brief that a test is named
676
- for what it proves rather than for the control that specified it.
677
-
678
- ### Carry every finding
679
-
680
- After reconciling findings into briefs, walk the retained finding list once. Every finding names
681
- the brief item that carries it. A finding with no carrier is a dropped finding.
682
-
683
- Name a carrier by the unit, never by a condition. "Whichever unit next touches this file" is a
684
- description that stops being true the moment the next unit is chosen, and nobody re-reads the verdict
685
- at that moment — so the finding reads as assigned and is not. Where the carrying unit is not yet
686
- known, say that plainly and re-check the assignment at the next dispatch.
687
-
688
- Every finding names exactly one carrier. A further brief claiming the same finding is not redundancy
689
- that costs a little duplicated work — it is a conflict the executor discovers mid-unit, between
690
- documents you told it to obey, with no way to tell which you meant. It will either implement the row
691
- twice or stop. Where a reconciliation table and a brief disagree about who owns a row, the brief the
692
- executor opens wins, and you fix the other one before the second unit launches.
693
-
694
- ## Long-running commands
695
-
696
- A bench exec, a Workflow, an install, a build, and a publish chain are one class of thing: a
697
- command that outlives the turn that started it. Every law here binds all of them.
698
-
699
- ### Launching
700
-
701
- - The Orchestrator launches every long command as a harness-tracked background command under a hard
702
- time cap. Never detach one from inside a dispatched agent. The harness owns the lifecycle,
703
- completion re-invokes the session, and the cap kills a wedged command loudly instead of trusting
704
- the agent to report its own failure. A wedged bridge is silent, and silence must never read as
705
- progress.
706
- - Never replicate git commits through a hosting provider's REST API when a file's content must ride
707
- inside the tool-call JSON. The per-call read cap truncates it, and a lockfile is the file that
708
- proves the point: it transits incomplete and the tree it lands in installs something else. Push
709
- over git, or move the file another way.
710
- - Write a multi-step chain to a script file and run the file. A chain composed inside one shell
711
- argument cannot be read back, corrected, or re-run, and the record of what actually ran is the
712
- argument text in a transcript rather than a file on disk.
713
- - Never edit a script file while a shell is executing it. `bash` reads a script incrementally from a
714
- byte offset rather than loading it, so an edit that shifts line numbers moves the text under that
715
- offset and the shell resumes mid-construct. The run dies on a syntax error in a line the script
716
- does not contain, which reads as a defect in the work rather than as the edit that caused it. Copy
717
- the file, edit the copy, and launch the copy for the next run. A successor's header names its own
718
- file, its own log, and what changed from the file it supersedes.
719
- - On a Windows host this binds every program-carrying command, not only long ones. Heredocs,
720
- `node -e`, `node -p`, `&&` chaining, and any argument carrying `${...}` trip the Git Bash
721
- approval classifier and turn an unattended run into a manual approval prompt. Write the program
722
- to a file, invoke the file, and keep each shell call one plain command.
723
- - Detach anything that must survive its launching shell with `setsid`. A backgrounded flow the
724
- harness reaps mid-step leaves the work half done and the exit status missing, and the reap looks
725
- identical to the step failing.
726
- - Size the cap yourself, from the observed high mark of comparable commands, plus an
727
- independently budgeted gate allowance, plus explicit slack. Never size it from the estimate
728
- alone. Never delegate it: a bridge starts with a clean context, holds no record of prior runs,
729
- and can only guess. A cap-killed exec is indistinguishable from a real failure.
730
- - Run the first use of any CLI flag, subcommand, quoting form, or stdin combination in a throwaway
731
- probe. Never inside a dispatched unit or a publish chain.
732
- - A launch is not a launch until its record grows past its header. Confirm the log advanced beyond
733
- the head before recording that the command started, and treat an instantly-dead log as a failed
734
- launch whose tail is the evidence.
735
- - Keep network-dependent work out of sandboxed bench execs. Bench sandboxes deny network, so
736
- lockfile generation, real installs, and live fetches belong to the Orchestrator's own tracked
737
- commands or to whichever of `sol`, `opus`, and `builder` runs natively in the harness, as an
738
- ordinary dispatched writing unit. A bench exec hanging on `npm` until its cap fires is the
739
- signature of this misroute, not of a slow bench.
740
- - A Workflow journals identically and dies identically, so give it the same watch — with one
741
- correction. A workflow journal writes only at agent start and result, so its mtime goes quiet for
742
- minutes during healthy work, and the liveness signal is the newest subagent transcript instead. A
743
- watch that reports only new events cannot report a death, because silence and progress look the
744
- same; the filter must fire on absence. Recover with `resumeFromRunId`, which returns every
745
- completed agent from cache and re-runs only what never finished.
746
-
747
- ### Reading liveness
748
-
749
- Read liveness from the artifact the work produces, never from its wrapper. A subagent's transcript
750
- file can report zero bytes while the agent is working normally, so an empty or stale wrapper proves
751
- nothing.
752
-
753
- - Judge a unit by what it has changed in the tree: modification times on the files it owns, the
754
- counts its suite reports, the report it was told to write.
755
- - Check that before killing anything. A healthy unit killed on a false signal loses everything it
756
- had not yet written down, and the loss is charged to the orchestrator, not the unit.
757
- - If a unit must be stopped, say plainly that it was stopped and why, then assess the tree it left
758
- rather than assuming its partial bytes are either good or worthless.
759
- - Follow the deviation ladder for a stalled journal or a cap-killed exec, using the session id from
760
- the journal head as the recovery handle.
761
-
762
- ### Confirm dead before relaunching
763
-
764
- - Prove the previous run is gone before starting another. List the processes and read the list. A
765
- second run started beside a live first one produces failures that read as the subject's — a
766
- publish chain relaunched over a live one reports `EOTP` and `E403` that are its own processes
767
- colliding, and both readings point at the registry.
768
- - Never ask `pgrep -f` or `ps | grep` whether a command is running from a shell whose own command line
769
- contains that command's text. The shell matches itself, so the answer is yes whatever the truth is.
770
- This bites in separate places and each one fails differently:
771
- - **A liveness watcher** loops forever, reporting "still running" and never delivering its completion
772
- notification, because it always finds itself.
773
- - **An elapsed-time reading** returns the watcher's age rather than the exec's, so the number falls
774
- instead of rising and reads as a relaunch that never happened.
775
- - **A pre-launch "is anything already running" check** reports a phantom concurrent writer, which is
776
- the worst of them: the honest response to it is to kill something, and there is nothing there.
777
- Read liveness from the recorded process id with `kill -0 <pid>`, or enumerate by executable name and
778
- parent with `ps -eo pid,ppid,comm` and read the rows. Both are immune; a pattern over the full command
779
- line is not.
780
- - Kill by process id, never by pattern. `pkill -f` matches the relaunch that is already starting, so
781
- the pattern that cleans up the old run kills the new one and the cleanup reads as a launch
782
- failure.
783
- - A killed `codex exec` is dead only when its process tree is dead: walk the children with
784
- `ps --ppid` and confirm the `codex-code-mode-host` child is gone. Before dispatching a substitute
785
- writer, check the owned files' modification times against the baseline — a live orphan is still
786
- writing the tree the substitute is about to own.
787
- - Read a failure against what was running when it happened, not against what you believe was
788
- running. The check costs one command and is the only thing that separates a real failure from
789
- self-inflicted contention.
790
-
791
- ## Bench laws
792
-
793
- External engines widen capacity. They never inherit authority. Treat every bench output as a
794
- proposal or hypothesis until it is verified against source and accepted by the Orchestrator.
795
-
796
- A bench is cross-provider reach only. Never send a model across a bridge when the running harness
797
- hosts it natively.
798
-
799
- A bench exec is a long-running command, so every law under **Long-running commands** binds it too.
800
- This section adds what is true of a bench and nothing else.
801
-
802
- Every bridge verifies before running that its CLI resolves and its bench is authenticated, and stops
803
- with a deviation report naming the fallback when either fails. The role file owns the exact
804
- invocation, flags, paths, probe, and recovery ladder; these laws bind every bench regardless of
805
- transport.
806
-
807
- 1. **Transport by work class.** Use an MCP transport only for a short interactive exchange — one
808
- bounded question or a follow-up on a live thread, finishing in roughly two minutes. Use the
809
- journaled CLI for audits, implementation units, and anything else multi-minute. An interrupted
810
- MCP call loses its session invisibly; a journal survives any client-side failure.
811
- 2. **Journal first.** Every bench invocation leaves a tailable on-disk record beside its brief
812
- under `tmp/<bench>/`: the event stream or output log, and the final answer. Arm exactly one
813
- Monitor per long exec, filtered to milestones — commands run, files changed, agent messages,
814
- terminal states — never the raw event stream. Exit the filter on the exec's terminal event so no
815
- watcher outlives its subject. The journal's mtime is the liveness signal; the session id in its
816
- head is the recovery handle. The journal is also the proof the bench ran: a bench unit returns
817
- its journal path and session id with its result, and the Orchestrator confirms both before
818
- using that result. A report does not carry the engine that produced it, so a bench unit with no
819
- journal ran on its driver's engine, however normal its answer reads.
820
- 3. **Tracked, never loose.** Register every bench unit in the session task registry at launch with
821
- its subject, journal path, and session id, and complete it there at acceptance. "What is
822
- running" always has a first-class answer instead of a recollection of a command.
823
- 4. **One lane at a time per bench.** Launch one `grok` lane at a time and queue the rest. Lanes past
824
- that starve the bench: it accepts each launch and the lanes come back empty, which reads as the
825
- work failing. Re-probe before ruling on an empty lane. A probe that round-trips while the lanes
826
- return empty names starvation, not darkness — cut the concurrency and re-run the lane, rather
827
- than recording the bench dark and substituting an engine.
828
- 5. **A bench sandbox spawns a child and denies that child's child.** Under `workspace-write` a bench
829
- exec runs a test suite and spawns children normally, and every operation one level deeper fails:
830
- a grandchild process is denied `EPERM`, and a nested `npm install` is denied the same way. So a
831
- proof needing a process tree, a tree-kill, a detached group, or an installed package cannot be
832
- produced inside a bench unit at all — however carefully that unit isolates. Name the limit in the
833
- brief before dispatch, tell the unit to record such a proof as an observation naming the exact
834
- settling command, and take that proof yourself on the host. Never let a unit substitute the
835
- reachable half: linking a packed tarball is not installing it, and a gate written to catch an
836
- install failure that only ever links cannot see the defect it exists for.
837
- Not every symptom names the sandbox. A nested process and a nested install fail `EPERM`, which
838
- reads as a denial; a nested `git` invocation instead reports **"not a git repository"** while the
839
- unit's own `git status` succeeds a moment earlier. That one reads as a broken checkout, and a unit
840
- acting on it will go looking for damage that is not there. Tell a unit which of its tools shell out
841
- one level down — a scaffolding CLI that probes git, a formatter that spawns a worker — so it
842
- recognises the shape instead of diagnosing the tree.
843
- The child a bench does create has unreliable stdio, and that failure wears a worse disguise than a
844
- denial. A Node process spawned by a bench unit's own Node process has been measured both buffering
845
- its pipe until EOF and publishing nothing at all, so no workaround built on either reading is
846
- dependable. Any subject whose behaviour lives in a child's pipes is therefore unmeasurable inside a
847
- bench: a stage driving a language server, a protocol fixture, a built entry driven as a spawned
848
- child. It fails as a **false green**. The stage never arms, the boot inspection times out, and that
849
- timeout produces the same rejection a genuine stage timeout produces, so a test asserting on the
850
- message passes inside the bench while the host's gate reports the honest red — and neither run
851
- reports why they disagree. Route such a subject to the harness's native writing lane, or keep it on
852
- the bench and supply every executed measurement yourself. Never dispatch it to a bench and expect
853
- it to prove its own work. The shape to recognise is a child that exits 0 almost immediately, a
854
- request to it that never resolves, and a stack landing in the spawning code's exit handler.
855
- **When a bench sandbox denies a loopback listener, `listen` fails `EPERM` on every address.** A
856
- subject needing a real local server is unmeasurable inside the bench. Name the limit in the brief
857
- before dispatch. Have the unit report the reading as an observation naming the exact command.
858
- Take the proof on the host.
859
- **When a brief assigns a bench unit a path outside the obvious source tree, name the write limit
860
- in the brief.** If the sandbox rejects the patch, the unit stops and reports the rejection. Never
861
- find another write mechanism.
862
- **A bench unit writes only under its own `--cd` root and the system temporary directory.** A
863
- report path in the orchestrator's repository is outside both, so the write is rejected after the
864
- work is already done and the unit improvises a fallback the Orchestrator has to go find. Name the
865
- report path inside the unit's own tree, or tell the unit to return the report as its final
866
- message and take it from the `--output-last-message` file. Retention still lands the report in
867
- `.orkestrel/<package>/` — the Orchestrator copies it there, and the unit never writes it there.
868
- 6. **Ephemeral streams, durable records.** A journal proves a bench is alive and recovers an
869
- interrupted session. Keep journals under `tmp/`, never commit them, and sweep them at acceptance
870
- after the final gate evidence is recorded. Durable retention — brief, distillate, verdict,
871
- instrument, acceptance evidence — is owned by **Dispatch anatomy**; this rule owns only the
872
- journal stream.
873
-
874
- ### Recovering a dark bench
875
-
876
- - A probe that finds no bench binary records the bench dark and, in the same turn, names to the user
877
- the install command and the bench it unblocks. Re-probe when the user answers.
878
- - A probe that finds a bench binary present but authentication unavailable starts recovery in the
879
- same turn. Do not record the bench dark and wait.
880
- - Background the login command with its output captured under `tmp/<bench>/`, surface the
881
- verification URL and one-time code to the user the moment they appear there, arm a watcher on
882
- completion, and re-probe when it fires. The bench comes live mid-session with no restart.
883
- - A session that sits dark until the user asks for the login has failed the probe, not the bench.
884
- - Never substitute an API key, access token, copied auth file, or another login flow. If recovery
885
- cannot complete, record the bench dark, name the fallback in the plan, and say so.
886
- - The role file owns each bench's exact login command and probe.
887
-
888
- ## Publishing the fleet
889
-
890
- Publishing is the user's decision and the user's credential. The Orchestrator prepares, surfaces
891
- the approval, and runs the publishes the user asked for. It never substitutes an API key, an access
892
- token, a copied auth file, or another login flow, and it never asks the user to paste a token into
893
- the conversation.
894
-
895
- A publish chain is a long-running command, so every law under **Long-running commands** binds it:
896
- write the chain to a file, detach it with `setsid`, and confirm the previous one is dead before
897
- starting another. Publish serially, because concurrent publishes collide on the authentication
898
- handshake and fail each other.
899
-
900
- The release itself runs from the `orkestrel-publish` skill: the wave's per-repo visit and its bump
901
- triggers in `references/wave.md`, and the preparation order, the login approval, and the
902
- five-minute upload window in `references/window.md`. Load the skill when the user asks for a
903
- release, and follow it there rather than reconstructing the procedure here. What remains in this
904
- section binds an executor who is not publishing.
905
-
906
- A wave over unpublished tips derives its order per run from the graph and records only the round
907
- each package landed in, never the order itself.
908
-
909
- ### Fixing a dependency before it publishes
910
-
911
- A defect a consumer meets sometimes lives in a package the consumer only has from the registry.
912
- Waiting for that package to publish before the consumer can prove its own fix serializes releases
913
- that could have been one. Do not wait, and do not work around it in the consumer.
914
-
915
- Build the dependency from source, pack it, and **install the tarball** into the consumer.
916
-
917
- - **Install it, never link it.** A link resolves through a directory and skips the packing, the
918
- `files` list, and the exports map — which is most of what a distribution proof exists to check. A
919
- gate written to catch an install failure that only ever linked cannot see the defect it exists for.
920
- - **Write the swap to a script and run the file**, so the build, the pack, and the install are one
921
- artifact the next run reuses rather than a command nobody can read back.
922
- - **Record the range you replaced** in the same step that replaces it. A consumer sitting on an
923
- unpublished tarball with no record of what it had is a consumer nobody can restore.
924
- - **Rebuild and repack whenever the source moves.** A stale tarball is the same defect as a stale
925
- `dist/`, and it is worse for being invisible: the consumer's gates go green against a fix that no
926
- longer exists in the dependency's tree.
927
- - **Run one unit per checkout, at that checkout's catalog layer.** Give a checkout with no rows an
928
- adopt unit only when its typecheck against the staged closure reddens.
929
- - **Fetch and merge the dependency's default branch before packing it**, wherever another session
930
- can move that branch. A pack from a stale tip ships the consumer a dependency the dependency's own
931
- repository no longer has.
932
- - **Restore the registry copy before any gate that must prove the published artifact, and before
933
- publishing anything.** A distribution proof run against a local tarball proves the local tarball.
934
- The release still follows layer order: the dependency publishes first, then the consumer re-pins to
935
- the version the registry now serves and re-runs its gates against that.
936
- - **Keep the tarballs out of the tree.** They belong under `tmp/`, they are swept at acceptance, and
937
- they are never committed.
938
-
939
- The tarball is a head start, not a shortcut. It lets the consumer's work proceed and its proofs run
940
- against the real packed artifact while the dependency's own release is still being prepared.
941
-
942
- ### What a bump obliges
943
-
944
- A runtime dependency and a development dependency have different blast radius, and confusing them
945
- either publishes packages nobody needed to publish or leaves a consumer pinned to an older release.
946
-
947
- - A **runtime** `dependencies` bump reaches every consumer of the published package. Every package
948
- downstream of it re-pins, re-runs its gates, bumps, and republishes, in layer order.
949
- - A **development** `devDependencies` bump reaches nobody. Re-pin it, prove the gates still green,
950
- and commit to `main`. Do not bump the version and do not publish.
951
- - A development bump that moves the published artifact is no longer a development bump. Prove the
952
- direction with the build, not the diff of sources: rebuild after the re-pin and compare `dist/`
953
- against the published tarball. Compare material content only — exclude sourcemaps and ignore
954
- whitespace-only differences; a superfluous diff (formatting, blank lines, map noise) moves
955
- nothing and obliges nothing. A material diff — tokens, declarations, logic — means the published
956
- surface moved — a forced `src` or `app` edit and a toolchain-changed emit both surface here — so
957
- that package bumps and publishes on its own account, and its own dependents follow the preceding
958
- runtime rule.
959
-
960
- Every package is `0.0.x`, where a caret pins one exact release. A dependent therefore sees a new
961
- version only after it re-pins and republishes, so the fleet publishes in topological layer order
962
- derived from the runtime `dependencies` and `peerDependencies` edges, never from a development
963
- edge. Layers exist for a reason a flat pass cannot fix: ranges that disagree install duplicate
964
- copies of the same package, and the compiler reads them as distinct types.
965
-
966
- Read the order from the catalog table in `.claude/agents/orkestrel.md`, which `scaffold catalog`
967
- regenerates from the registry. Its `Layer` column is the publish round. Regenerate it before
968
- sequencing a cascade rather than trusting the copy in the tree, and never write a second order down
969
- somewhere else.
970
-
971
- The tooling packages sit outside that order because nothing depends on them at runtime. `scaffold`
972
- is a development dependency of every package, including packages it depends on itself, so a runtime
973
- layering would report a cycle that does not exist. Each package builds against the already-published
974
- `scaffold`, never against an unpublished one, and a `scaffold` release therefore publishes on its own
975
- and propagates as files rather than as a cascade. A package the fleet consumes as a development
976
- dependency takes the same shape when its consumers' gates read its unpublished tip;
977
- `.agents/skills/orkestrel-publish/references/wave.md` § Rule on the bump states that trigger.
978
-
979
- `scaffold` also carries a second published surface beside `dist/src`: `package.json` ships
980
- `dist/host`, the vendored file set every target receives through `repair`.
981
-
982
- - Bump and publish `scaffold` when any vendored byte changes, or when the set of vendored paths
983
- changes. That surface moved on its own account, and `dist/src` need not move with it.
984
- - Re-pin `@orkestrel/scaffold` in each target after a vendored-only release, run `repair` there, and
985
- prove that target's gates still green. `repair` restores `tests/setupPolicy.ts` and
986
- `tests/policy.test.ts`, so a vendored-only release can turn a green target red. A target bumps
987
- only when its own published surface moved.
988
- - Keep a target's own Claude permissions in `.claude/settings.local.json`, never in the vendored
989
- `.claude/settings.json`. `repair` restores the vendored copy, so a `defaultMode` or an `allow`
990
- entry added there is reverted without warning and the operator loses grants they set
991
- deliberately. Change the vendored file only here, in the host inventory.
992
- - Never edit a vendored file inside a target. `repair` restores it, so the edit is reverted and
993
- reports as drift in `scaffold audit`. In this repository those same files are the published
994
- `dist/host` surface, so editing one forces a bump, a publish, and a re-propagation across every
995
- target. Scope a fleet-wide refactor to the files each target owns, and record the vendored
996
- exclusion in the brief rather than letting each unit rediscover it.
997
-
998
- ## Acceptance laws
999
-
1000
- - No writer's and no external engine's self-assessment is authoritative.
1001
- - Never spend Opus 5.5 or Astra on absorption, distillation, scouting, or mechanical edits. Never route
1002
- judgment-bearing implementation away from Opus 5.5 or Astra.
1003
- - Substitute an engine only when the same session records the bench dark — CLI missing, auth
1004
- expired, model unavailable. Name the fallback in the plan; never improvise it silently. The
1005
- tedious-work ladder is the only pre-approved substitution, and each step down it is still recorded.
1006
- - Never run the lanes on different briefs, and never show either one the other's answer before
1007
- both have returned.
1008
- - Never run a lane inline in the Orchestrator's context, and never drop a lane because its default
1009
- engine is unavailable. Substitute the engine, keep the lane.
1010
- - Never accept unreviewed implementation, unverified hypotheses, shared-tree writing races,
1011
- implicit engines, fixed Claude model IDs, or verbose completed-work residue.
1012
- - Evidence a claim about a rendered or externally driven surface with its capture or a real foreign
1013
- client driving it, never with source alone. Where no such surface exists this law is inert.
1014
- - When the Orchestrator writes any part of a unit, that part is briefed, owned, and audited like any
1015
- other part, and its auditor is an engine the Orchestrator does not share.
1016
- - Final acceptance belongs only to the Orchestrator, after independent audit and gate evidence.
1017
- - Accept when the plan's exit criterion is met and the gates are green, not when the last engine
1018
- runs out of appetite. Reopening an accepted criterion is the user's instruction, not an auditor's
1019
- finding.
137
+ A unit stops when a conflict blocks its objective or requires an unowned change. It reports expected, found, exact evidence, done or not done, and at most one hypothesis. It settles an ancillary choice inside its scope and records it.
138
+
139
+ The Orchestrator triages: obvious correction → tighten and re-dispatch; missing evidence → `verifier`; unknown terrain → `grok`; unknown design or cause → one `planner` or `analyst` question. Then decide, update the plan, re-dispatch.
140
+
141
+ ## Benches
142
+
143
+ - Probe a bench with the dispatch skill's `scripts/bench.ts` before its first lane in a session. A version string or a login status is not liveness.
144
+ - A dispatch that fails on auth, quota, model access, or network records the bench dark and re-plans the lane on the substitute engine: Astra dark → Opus holds the objective lane too; Opus dark in Codex → Astra holds the subjective lane too; Grok dark → Luna, then Sonnet. Record every substitution in the routing ledger.
145
+ - Never assign Grok a design or review lane in Claude Code or Codex.
146
+ - Journal every bench run under `tmp/<bench>/` with its session id. Never commit a journal.
147
+
148
+ ## Cleanup
149
+
150
+ - When a unit is accepted, carry its measurements, unresolved findings, and any instrument worth keeping into the commit message, the campaign ledger, or a test, then delete its `tmp/units/<unit>-*` files and its bench journals.
151
+ - `.orkestrel/` holds only what § .orkestrel layout names. Delete a verdict when its seam closes and the package directory in the acceptance commit; git history is the archive.
152
+ - Delete a probe from the source tree before the unit returns; promote a probe that settled a claim into a test.
153
+ - Run `node .agents/skills/orkestrel-dispatch/scripts/sweep.ts --report` at acceptance; it lists leftovers under `tmp/` and `.orkestrel/` by age and deletes nothing. Run the same script with `--tmp` after the last live unit exits; it deletes files older than its threshold and refuses while a journal is changing.
154
+ - A campaign artifact never lives in the package it is about.
155
+
156
+ ## Cache
157
+
158
+ - Fix the model at session start. Effort switches and MCP changes invalidate the cache on most models and configurations; check the rule for the active model and tool-loading mode before changing either mid-session.
159
+ - Keep the always-on instruction set small; rules load by path. An edit to a loaded instruction file takes effect after `/clear`, `/compact`, or a restart.
160
+ - Give reading and driver roles no instruction files (`omitClaudeMd` in Claude Code) when their charter carries the floor they need.
161
+ - Send bulk reads to a subagent and keep only the distillate.
162
+
163
+ ## Acceptance
164
+
165
+ - No writer's and no bench's self-assessment is authoritative. The Orchestrator accepts after the review and the gates the size gate names.
166
+ - Accept when the exit criterion is met and the gates are green. Reopening an accepted criterion is the user's instruction, not a reviewer's finding.
167
+ - Substitute an engine only when the same session recorded the bench dark; name the fallback in the ledger.
168
+ - Publishing is the user's decision and credential; `orkestrel-publish` owns the release procedure, the bump rules, and the dependency-tarball rules.