@hybridlabor-api/aos 4.1.0 → 4.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/.agents/agents.md +77 -0
  2. package/.agents/graph.md +45 -0
  3. package/.agents/nodes.json +4 -2
  4. package/.agents/state.schema.json +6 -0
  5. package/.claude/workflows/startcycle-dispatch.mjs +139 -8
  6. package/.claude/workflows/teamwork-dispatch.mjs +287 -0
  7. package/CLAUDE.md +47 -0
  8. package/GEMINI.md +9 -1
  9. package/README.md +12 -6
  10. package/THIRD_PARTY_NOTICES.md +50 -0
  11. package/docs/skills_table.md +1 -0
  12. package/installer.js +15 -0
  13. package/package.json +1 -1
  14. package/skills/basic/bdbmediastorm/SKILL.md +7 -5
  15. package/skills/basic/startcycle/SKILL.md +21 -0
  16. package/skills/basic/startcycle-graph/SKILL.md +43 -7
  17. package/skills/basic/startcycle-graph-user/SKILL.md +67 -11
  18. package/skills/basic/teamwork-preview/SKILL.md +209 -0
  19. package/skills/bdbrainstorm/SKILL.md +4 -3
  20. package/skills/global_config/ask-tim/SKILL.md +73 -6
  21. package/skills/global_config/bdbresilience/SKILL.md +216 -0
  22. package/skills/global_config/bdbresilience/contracts/nodes-integration.md +225 -0
  23. package/skills/global_config/bdbresilience/references/cicd-triage.md +179 -0
  24. package/skills/global_config/bdbresilience/references/distributed-locking.md +235 -0
  25. package/skills/global_config/bdbresilience/references/error-recovery.md +210 -0
  26. package/skills/global_config/bdbresilience/references/two-phase-go-gate.md +151 -0
  27. package/skills/global_config/domain-modeling/ADR-FORMAT.md +47 -0
  28. package/skills/global_config/domain-modeling/CONTEXT-FORMAT.md +60 -0
  29. package/skills/global_config/domain-modeling/SKILL.md +77 -0
  30. package/skills/global_config/grill-me/SKILL.md +14 -0
  31. package/skills/global_config/grill-with-docs/SKILL.md +24 -0
  32. package/skills/global_config/grilling/SKILL.md +42 -0
  33. package/skills/global_config/openwiki-skill/scripts/install_daemon.sh +55 -12
package/CLAUDE.md CHANGED
@@ -15,6 +15,53 @@ Ask one question first: **do the workers need to see each other?**
15
15
 
16
16
  "Runs in parallel" is not a reason to reach for a team — subagents already run in parallel. Peer communication and dynamic task claiming are the only things a team adds.
17
17
 
18
+ ## Delegating to an external CLI
19
+ Some work is cheaper on another provider's compute (bulk scaffolding, exhaustive
20
+ test generation, long-context reads that distil to a digest). None of that tooling
21
+ ships with AOS — it depends on CLIs and Claude Code plugins the user installed
22
+ separately, so check what is actually present instead of assuming.
23
+
24
+ **Prefer a plugin's delegation subagent over shelling out to its CLI.** Where one
25
+ is installed it already handles the wrapper flags, cost discipline, and digest
26
+ contract: `antigravity:antigravity-delegate` (agy), `opencode:opencode-rescue`,
27
+ `codex:codex-rescue`. These are Claude Code plugins — on another harness, or a
28
+ machine without them, calling the CLI directly is the only path.
29
+
30
+ **Delegate only above the break-even.** A small, self-contained, or
31
+ judgement-heavy task costs more to hand off and verify than to just do. Keep the
32
+ digest, not the raw output.
33
+
34
+ **Give it a real timeout.** Measured 2026-09: a trivial headless `agy` prompt
35
+ took **605s**. `agy-delegate` defaults to `--print-timeout 5m`, so it aborts at
36
+ 300s and reports an empty body while the answer is still coming — pass
37
+ `--timeout 15m` for anything non-trivial. A short timeout does not read as
38
+ "slow", it reads as "broken".
39
+
40
+ **Match the model to the task, not to the default.** `agy-delegate`'s tiers map
41
+ to models that can go stale (its built-in `flash` still points at Gemini 3.7
42
+ while 3.8 ships). Either pass `--model "<exact name from \`agy models\`>"` per
43
+ call, or remap the tiers once via the plugin's own options — as env vars those
44
+ belong in `~/.zshenv`, not `~/.zshrc`, since `.zshrc` is only sourced for
45
+ interactive shells and tool-invoked ones would never see them:
46
+
47
+ | Work | Model |
48
+ |---|---|
49
+ | media, fast/mechanical coding, boilerplate | `Gemini 3.8 Flash (Medium)` → `CLAUDE_PLUGIN_OPTION_TIER_FLASH` |
50
+ | trivial one-liners | `Gemini 3.8 Flash (Low)` → `CLAUDE_PLUGIN_OPTION_TIER_FLASH_LO` |
51
+ | review, architecture, hard reasoning | `Claude Sonnet 4.6 (Thinking)` → `CLAUDE_PLUGIN_OPTION_TIER_PRO` |
52
+
53
+ Adversarial review is the case that most repays a stronger model: a Flash tier
54
+ tends to agree with what it is shown, which is the one thing a reviewer must
55
+ not do. Re-check the names against `agy models` after an agy upgrade — the id
56
+ carries both the version and the effort suffix.
57
+
58
+ **Verify the result, never the status field.** A timed-out delegation returns
59
+ `{"status": "SUCCESS", "usage": {"total": 0}}` with an empty body — success by
60
+ every field except the one that matters, and the zero token counts are *not*
61
+ proof the prompt never arrived (headless usage reporting is simply unpopulated).
62
+ Check the returned content, treat an empty body as failure regardless of status,
63
+ and never report a delegated step as done on the strength of its own self-report.
64
+
18
65
  ## Safety Gate — mechanically enforced, not advisory
19
66
  `git push`, `npm publish`, `npm version`, and recursive `rm` are blocked by `.claude/hooks/go-gate.mjs` (registered in `.claude/settings.json`) unless your immediately preceding message is the literal word **GO**. This is a hook, not a rule I read and try to follow — it cannot be argued around, and it doesn't depend on this file being loaded.
20
67
  - A subagent does not inherit its orchestrator's GO.
package/GEMINI.md CHANGED
@@ -47,4 +47,12 @@ Ask one question first: **do the workers need to see each other?**
47
47
  - **Yes — they must react to each other, or claim work dynamically from a shared list** → an orchestrated agent team (via `send_message`). Currently only `/bdbrainstorm` qualifies, where the spec demands a real debate rather than parallel monologues. True Agent Teams were evaluated and deferred for `/startcycle-graph` (needs an interactive session; the graph runs headless) — see `.agents/graph.md` and F-17's addendum in `docs/sessions/audit-agents.md`.
48
48
  - **Small task** → do it yourself. A two-file edit needs no agents.
49
49
 
50
- "Runs in parallel" is not a reason to reach for a team — subagents already run in parallel. Peer communication and dynamic task claiming are the only things a team adds.
50
+ "Runs in parallel" is not a reason to reach for a team — subagents already run in parallel. Peer communication and dynamic task claiming are the only things a team adds.
51
+
52
+ ## 9. Delegating to an external CLI
53
+ None of this tooling ships with AOS — it depends on CLIs and plugins the user installed separately, so check what is present rather than assuming.
54
+ - **Prefer a plugin's delegation subagent over shelling out to its CLI.** Where installed, it already handles the wrapper flags, cost discipline, and digest contract: `antigravity:antigravity-delegate` (agy), `opencode:opencode-rescue`, `codex:codex-rescue`. These are Claude Code plugins — on another harness, calling the CLI directly is the only path.
55
+ - **Delegate only above the break-even.** A small, self-contained, or judgement-heavy task costs more to hand off and verify than to just do. Keep the digest, not the raw output.
56
+ - **Give it a real timeout.** Measured 2026-09: a trivial headless `agy` prompt took **605s**. `agy-delegate` defaults to `--print-timeout 5m`, so it aborts at 300s and reports an empty body while the answer is still coming — pass `--timeout 15m` for anything non-trivial.
57
+ - **Match the model to the task, not to the default.** The wrapper's tiers map to models that go stale (built-in `flash` still points at Gemini 3.7 while 3.8 ships). Media and fast/mechanical coding → `Gemini 3.8 Flash (Medium)`; trivial → `Gemini 3.8 Flash (Low)`; review, architecture and hard reasoning → `Claude Sonnet 4.6 (Thinking)`. Adversarial review most repays the stronger model: a Flash tier tends to agree with what it is shown, which is exactly what a reviewer must not do. Pass `--model` per call, or remap the tiers once via `CLAUDE_PLUGIN_OPTION_TIER_{FLASH,FLASH_LO,PRO}` — in `~/.zshenv`, not `~/.zshrc`, which non-interactive tool shells never source. Re-check names against `agy models` after an upgrade.
58
+ - **Verify the result, never the status field.** A timed-out delegation returns `{"status": "SUCCESS", "usage": {"total": 0}}` with an empty body — success by every field except the one that matters, and the zero token counts are *not* proof the prompt never arrived (headless usage reporting is simply unpopulated). Treat an empty body as failure regardless of status, and never report a delegated step as done on the strength of its own self-report.
package/README.md CHANGED
@@ -5,13 +5,19 @@
5
5
 
6
6
  [![CI](https://github.com/hybridlabor-api/aos/actions/workflows/ci.yml/badge.svg)](https://github.com/hybridlabor-api/aos/actions)
7
7
  [![NPM Version](https://img.shields.io/npm/v/@hybridlabor-api/aos.svg)](https://www.npmjs.com/package/@hybridlabor-api/aos)
8
- [![runtime](https://img.shields.io/badge/node-20+-blue.svg)](https://github.com/hybridlabor-api/aos)
8
+ [![NPM Downloads](https://img.shields.io/npm/dw/@hybridlabor-api/aos.svg)](https://www.npmjs.com/package/@hybridlabor-api/aos)
9
9
  [![license](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](LICENSE)
10
- [![skills](https://img.shields.io/badge/skills-154%20curated-brightgreen.svg)](https://github.com/hybridlabor-api/aos)
10
+ [![GitHub stars](https://img.shields.io/github/stars/hybridlabor-api/aos?style=flat&color=gold)](https://github.com/hybridlabor-api/aos/stargazers)
11
+ [![last commit](https://img.shields.io/github/last-commit/hybridlabor-api/aos.svg)](https://github.com/hybridlabor-api/aos/commits/main)
12
+
13
+ [![skills](https://img.shields.io/badge/skills-169%20curated-brightgreen.svg)](docs/skills_table.md)
14
+ [![MCPs](https://img.shields.io/badge/local%20MCPs-21-brightgreen.svg)](mcps/)
15
+ [![harnesses](https://img.shields.io/badge/harnesses-9%20supported-blueviolet.svg)](#-installation)
16
+ [![runtime](https://img.shields.io/badge/node-20+-blue.svg)](https://github.com/hybridlabor-api/aos)
11
17
 
12
- > **Supercharging AI coding agents with 154 hyper-curated skills, 21 local MCP wrappers, and a runnable multi-agent dispatcher graph.**
18
+ > **Supercharging AI coding agents with 169 hyper-curated skills, 21 local MCP wrappers, and a runnable multi-agent dispatcher graph.**
13
19
 
14
- Welcome to **BDB Agent OS — AOS v4.0.0**: 154 curated skills, 21 local MCP wrappers, and a dispatcher graph that turns them into a real multi-agent build pipeline, not just a prompt library. Point it at a goal and it plans, builds, reviews, and ships through seven coordinated agent nodes — with a mechanically enforced gate before anything actually goes live.
20
+ Welcome to **BDB Agent OS — AOS v4.0.0**: 169 curated skills, 21 local MCP wrappers, and a dispatcher graph that turns them into a real multi-agent build pipeline, not just a prompt library. Point it at a goal and it plans, builds, reviews, and ships through seven coordinated agent nodes — with a mechanically enforced gate before anything actually goes live.
15
21
 
16
22
  It is harness-neutral by design, not "optimized for one tool with others as an afterthought": the dispatcher graph runs on Claude Code's Dynamic Workflows, the same skills and MCP configuration install natively into **Google Antigravity, ChatGPT Codex / Codex CLI, Claude Desktop, Cursor, Aider, Roo Code, Cline, and Windsurf**, and the lightweight `/startcycle-graph-user` variant falls back to Claude Code's own subagents on any machine that has none of the above installed.
17
23
 
@@ -95,9 +101,9 @@ below for how the pipeline itself works.
95
101
 
96
102
  ---
97
103
 
98
- ## 🌟 154 Optimized Skills
104
+ ## 🌟 169 Optimized Skills
99
105
 
100
- We started with a massive pool of over 1,400 raw AI skills. After rigorous testing, filtering, and refinement, we've distilled them down to a hyper-curated set of **154 Optimized Skills** (featuring a native OpenWiki documentation engine, the **memB local semantic memory brain**, and **Universal Agent Harness synchronization**).
106
+ We started with a massive pool of over 1,400 raw AI skills. After rigorous testing, filtering, and refinement, we've distilled them down to a hyper-curated set of **169 Optimized Skills** (featuring a native OpenWiki documentation engine, the **memB local semantic memory brain**, and **Universal Agent Harness synchronization**).
101
107
 
102
108
  These skills are precision-engineered to ensure agents waste no time on redundant tasks and instead operate with maximum agency, strict architectural constraints, and robust context awareness.
103
109
 
@@ -127,6 +127,56 @@ SOFTWARE.
127
127
 
128
128
  ---
129
129
 
130
+ ## mattpocock/skills
131
+
132
+ <https://github.com/mattpocock/skills> — MIT.
133
+
134
+ The grilling family and the domain-modeling discipline it composes with:
135
+
136
+ | In AOS | Upstream |
137
+ |---|---|
138
+ | `skills/global_config/grilling/` | `skills/productivity/grilling/` |
139
+ | `skills/global_config/grill-me/` | `skills/productivity/grill-me/` |
140
+ | `skills/global_config/grill-with-docs/` | `skills/engineering/grill-with-docs/` |
141
+ | `skills/global_config/domain-modeling/` | `skills/engineering/domain-modeling/` (incl. `ADR-FORMAT.md`, `CONTEXT-FORMAT.md`) |
142
+
143
+ `grilling`'s interview protocol and `domain-modeling` are carried over
144
+ essentially verbatim; `grill-me` and `grill-with-docs` are rewritten to name AOS's
145
+ own pipelines in their hand-off sections, but keep upstream's composition — both
146
+ are thin wrappers that invoke the primitive rather than restating it. Upstream's
147
+ `agents/openai.yaml` under `domain-modeling` is harness-specific to that project
148
+ and was not carried over.
149
+
150
+ `skills/global_config/ask-tim/` is derived from upstream's `ask-matt` router. It
151
+ is not a copy: the flow it maps is AOS's own, since most of the skills on
152
+ upstream's main flow have no AOS equivalent.
153
+
154
+ ```
155
+ MIT License
156
+
157
+ Copyright (c) 2026 Matt Pocock
158
+
159
+ Permission is hereby granted, free of charge, to any person obtaining a copy
160
+ of this software and associated documentation files (the "Software"), to deal
161
+ in the Software without restriction, including without limitation the rights
162
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
163
+ copies of the Software, and to permit persons to whom the Software is
164
+ furnished to do so, subject to the following conditions:
165
+
166
+ The above copyright notice and this permission notice shall be included in all
167
+ copies or substantial portions of the Software.
168
+
169
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
170
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
171
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
172
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
173
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
174
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
175
+ SOFTWARE.
176
+ ```
177
+
178
+ ---
179
+
130
180
  ## Note on `mcps/`
131
181
 
132
182
  Sub-repositories vendored under `mcps/` carry their own `LICENSE` files in
@@ -48,6 +48,7 @@
48
48
  | `subagent-driven-development` | Use when executing implementation plans with independent tasks in the current session |
49
49
  | `tdd-workflow` | Test-Driven Development workflow principles. RED-GREEN-REFACTOR cycle. |
50
50
  | `tmux` | Expert tmux session, window, and pane management for terminal multiplexing, persistent remote workflows, and shell scripting automation. |
51
+ | `teamwork-preview` | Interactive 9-step prompt crafting and delegation protocol for autonomous multi-agent teams across harnesses. |
51
52
 
52
53
  #### 🎨 Frontend & UI/UX
53
54
  | Skill Name | Description |
package/installer.js CHANGED
@@ -2554,6 +2554,21 @@ function injectHarnessRules() {
2554
2554
  log.step(`Installed GEMINI.md to ${path.join(geminiDir, 'GEMINI.md')}`);
2555
2555
  }, 'The harness injection below still runs.');
2556
2556
 
2557
+ // Dispatcher scripts must land in ~/.claude/workflows/, because that is
2558
+ // where the skills that route to them look: startcycle-graph's SKILL.md
2559
+ // and teamwork-preview's both tell the model to call `Workflow` with
2560
+ // scriptPath `$HOME/.claude/workflows/<name>.mjs`. installProjectHarness()
2561
+ // copies them into a *project*, which only helps a project that opted into
2562
+ // the local harness -- on a plain global install those paths did not exist
2563
+ // at all, so the skill pointed at a file that was never delivered.
2564
+ installStep(`install dispatcher workflows to ${path.join(homeDir, '.claude', 'workflows')}`, () => {
2565
+ const workflowsSrc = path.join(srcDir, '.claude', 'workflows');
2566
+ if (fs.existsSync(workflowsSrc)) {
2567
+ copyDirRecursiveSync(workflowsSrc, path.join(homeDir, '.claude', 'workflows'));
2568
+ log.step(`Installed dispatcher workflows to ${path.join(homeDir, '.claude', 'workflows')}`);
2569
+ }
2570
+ }, '/startcycle-graph and /teamwork-preview fall back to their prose protocols.');
2571
+
2557
2572
  const startcycleWorkflowSrc = path.join(srcDir, '.agents', 'workflows', 'startcycle.md');
2558
2573
  const sources = installStep('read the global rule sources', () => ({
2559
2574
  globalRules: fs.readFileSync(geminiMdSrc, 'utf8'),
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@hybridlabor-api/aos",
3
- "version": "4.1.0",
3
+ "version": "4.2.0",
4
4
  "description": "AOS — A Curated AI AGENT OS. Optimized agent skills and add-ons like memB, OpenWiki, Heimdall Token Saver, and Godmode architectures.",
5
5
  "main": "installer.js",
6
6
  "bin": {
@@ -26,9 +26,11 @@ Ideation must never be performed in isolation. Spawn specialized subagents to an
26
26
 
27
27
  ---
28
28
 
29
- ## 2. Interactive `/grill-me` Technical Interview
29
+ ## 2. Interactive Technical Interview
30
30
 
31
- Before drafting signal flow diagrams or system configs, execute a mandatory `/grill-me` interactive interview. Deeply challenge the user's technical assumptions and hardware readiness by asking targeted questions:
31
+ Before drafting signal flow diagrams or system configs, **invoke the `grill-with-docs` skill** (or `grill-me` when there is no working directory) and run it to completion. Those skills hold the interview protocol — design tree, frontier rounds, numbered questions each with a recommended answer — and it is not restated here.
32
+
33
+ What this domain adds to that protocol: deeply challenge the user's technical assumptions and hardware readiness. The frontier questions for a show-control build are:
32
34
 
33
35
  * **Signal & Network Protocols:**
34
36
  - What protocols govern data movement? (OSC, Art-Net, sACN, MIDI, SMPTE Timecode, NDI)?
@@ -48,7 +50,7 @@ Before drafting signal flow diagrams or system configs, execute a mandatory `/gr
48
50
 
49
51
  ## 3. Target Directory & Scaffolding
50
52
 
51
- After aligning on system architecture through the `/grill-me` process:
53
+ After aligning on system architecture through the grilling interview:
52
54
  1. **Confirm Output Directory:** Ask the user: *"In which project directory should the output show-control plan and architecture files be stored?"*
53
55
  2. **Scaffold Foundational Files:** Once confirmed, write the core show specification files (`agent.md`, `signal-flow.md`, `network-patch.json`, `failover-matrix.md`).
54
56
 
@@ -94,7 +96,7 @@ BDB MediaStorm is the master ideation and brainstorming engine for live show-con
94
96
  - **Exclude:** Do not use for generating video timelines or 3D meshes.
95
97
 
96
98
  ## Core Process
97
- 1. Run an interactive `/grill-me` session to challenge assumptions about protocols, hardware, and bandwidth.
99
+ 1. Run `grill-with-docs` (or `grill-me`) to challenge assumptions about protocols, hardware, and bandwidth.
98
100
  2. Scaffold foundational files (`agent.md`, `signal-flow.md`, `network-patch.json`).
99
101
  3. Generate a strict Mermaid.js signal flow diagram mapping all protocols.
100
102
  4. Document a main/backup redundancy and failover matrix.
@@ -115,7 +117,7 @@ BDB MediaStorm is the master ideation and brainstorming engine for live show-con
115
117
 
116
118
  ## Verification
117
119
 
118
- - [ ] `/grill-me` session was completed with answers regarding protocols and bandwidth.
120
+ - [ ] The grilling interview was completed with answers regarding protocols and bandwidth.
119
121
  - [ ] Output includes a Mermaid.js signal flow diagram.
120
122
  - [ ] A dedicated failover/blackout mechanism is documented.
121
123
 
@@ -80,6 +80,27 @@ Run only the streams the goal actually needs. A plain backend feature does not n
80
80
 
81
81
  3c is for TouchDesigner, show-control, DMX/grandMA3, 3D, or other media-pipeline goals — most goals are not this. Skip it unless the plan actually calls for it.
82
82
 
83
+ > **Injecting a specific skill.** `/startcycle --skill=<name> <goal>`
84
+ > (repeatable, quote a name with spaces) forces that skill into this run as
85
+ > a hard requirement — for a private skill of your own that isn't part of
86
+ > any build agent's own `skills:` frontmatter. Since this variant has no
87
+ > dispatcher script or `state.json` to carry it automatically, the invoker
88
+ > does the work `startcycle-graph`'s script does for you: extract the
89
+ > `--skill=` flag(s) from the invocation text before anything else runs,
90
+ > confirm each name resolves to a real `SKILL.md` (under
91
+ > any harness's global skills directory — `~/.claude/skills/<name>/`,
92
+ > `~/.agents/skills/`, `~/.codex/skills/`, `~/.cursor/skills/` or `~/.roo/skills/`
93
+ > — or this project's own `skills/` tree) — stop
94
+ > and tell the user if one doesn't, don't silently proceed without it —
95
+ > note the validated list in `00_execution_plan.md`, and include it as a
96
+ > **hard requirement, not a suggestion** in each Build stream's dispatch
97
+ > prompt at step 3. Reviewer (step 4) checks the resulting artifacts for
98
+ > evidence the skill was actually applied, not just available, and treats
99
+ > an ignored mandate as a contract-misread finding. See
100
+ > [`.agents/graph.md`](../../../.agents/graph.md)'s "Mandatory Skill
101
+ > Injection" section for the full rationale — this is the same mechanic,
102
+ > just invoker-driven instead of script-driven.
103
+
83
104
  ### 4. Reviewer
84
105
  - **Agent**: `reviewer`
85
106
  - **Reads**: the artifacts each build stream produced (01/02/03) and the plan's stated contract (`00_execution_plan.md`) — nothing else. Never the goal directly, never a build agent's own claim that it's done; passing that claim through biases the review toward agreement.
@@ -21,12 +21,21 @@ several fields the schema doesn't define at all). Fix it first:
21
21
  mkdir -p .agents
22
22
  [ -f .agents/graph.md ] || cp "$HOME/.agents/graph.md" .agents/graph.md
23
23
  [ -f .agents/state.schema.json ] || cp "$HOME/.agents/state.schema.json" .agents/state.schema.json
24
+ [ -f .agents/nodes.json ] || cp "$HOME/.agents/nodes.json" .agents/nodes.json
24
25
  ```
25
26
 
26
- If `$HOME/.agents/graph.md` or `$HOME/.agents/state.schema.json` doesn't
27
- exist either, stop and tell the user: this machine has no canonical copy of
28
- the graph contract to bootstrap from, and `/startcycle-graph` will produce a
29
- non-conforming `state.json` until one is installed. Don't silently proceed.
27
+ `nodes.json` is not optional and is the one that fails loudest: it is the
28
+ node registry the dispatcher loads as its very first step, and without it
29
+ the run escalates immediately with *"`.agents/nodes.json` failed to load, or
30
+ is missing required node id(s)"* — before Architect or any other agent has
31
+ run. (Found exactly that way: a first run in a project that had followed
32
+ this bootstrap step as it was previously written, which copied only the
33
+ other two files.)
34
+
35
+ If any of the three doesn't exist under `$HOME/.agents/` either, stop and
36
+ tell the user: this machine has no canonical copy of the graph contract to
37
+ bootstrap from, and `/startcycle-graph` cannot run correctly until one is
38
+ installed. Don't silently proceed.
30
39
  (`.claude/agents/*.md`, the seven agent persona files, do NOT need this
31
40
  treatment — Claude Code resolves subagents from the user-level
32
41
  `~/.claude/agents/` fine without a project-local copy.)
@@ -37,9 +46,36 @@ yourself (e.g. `echo $HOME` or your own environment info) rather than
37
46
  hardcoding a username, giving
38
47
  `$HOME/.claude/workflows/startcycle-dispatch.mjs` — and `args` set to the
39
48
  goal text that follows `ARGUMENTS:` below this file's content. Pass the goal
40
- through verbatim. If there is no `ARGUMENTS:` text, pass no `args` (or
41
- `args: undefined`) — the workflow itself asks for a goal in that case rather
42
- than guessing one.
49
+ through verbatim, including any `--skill=<name>` flag(s) it contains — the
50
+ dispatcher script parses those itself (see below), do not strip or
51
+ interpret them yourself. If there is no `ARGUMENTS:` text, pass no `args`
52
+ (or `args: undefined`) — the workflow itself asks for a goal in that case
53
+ rather than guessing one.
54
+
55
+ **Injecting a specific skill.** `/startcycle-graph --skill=<name> <goal>`
56
+ (repeatable, quote a name with spaces) forces that skill into this run as a
57
+ hard requirement for the build nodes, validated to exist before anything
58
+ else runs — this is how you make the pipeline use your own private skill
59
+ that isn't part of `.agents/nodes.json`'s registry.
60
+
61
+ `<name>` is the **exact skill directory name**, not a description — `ui-component`,
62
+ not "the UI one". Two ways to find it without leaving the terminal:
63
+
64
+ - `/ask-tim` — the routing skill; start there when you know the *job* but not the name
65
+ - list the installed skills directly, if you half-remember the spelling. The installer
66
+ syncs the same set to every harness it detects, so use whichever path is yours:
67
+ `~/.claude/skills`, `~/.agents/skills`, `~/.codex/skills`, `~/.cursor/skills`, or
68
+ `~/.roo/skills`. `ls ~/.agents/skills` is the safest guess on an unknown machine —
69
+ that one is written on every install regardless of harness.
70
+
71
+ A name that does not resolve halts the run before any agent works, and the
72
+ error now lists installed near-misses rather than only saying "not found".
73
+ That is deliberate: silently running without a skill you explicitly demanded
74
+ is worse than stopping.
75
+
76
+ See [`.agents/graph.md`](../../../.agents/graph.md)'s "Mandatory Skill
77
+ Injection" section for the full mechanics; nothing about it needs handling
78
+ in this router file, since `args` is passed through as raw text either way.
43
79
 
44
80
  Use `scriptPath`, not `name: "startcycle-dispatch"` — by-name lookup for a
45
81
  custom (non-built-in) workflow script has been observed to fail with
@@ -29,6 +29,26 @@ Do not invent more structure than the task has. A two-file edit doesn't need
29
29
  a Plan node — just do it. This skill is for the cases actually shaped like a
30
30
  small graph, not an excuse to always draw one.
31
31
 
32
+ ## 1b. Injecting a specific skill (optional)
33
+
34
+ `/startcycle-graph-user --skill=<name> <task>` (repeatable, quote a name
35
+ with spaces) forces that skill into this run as a hard requirement — for a
36
+ private skill of the user's own this throwaway graph would otherwise never
37
+ know to reach for. Extract any `--skill=` flag(s) from the invocation text
38
+ before step 1, confirm each name resolves to a real `SKILL.md` (under
39
+ any harness's global skills directory — `~/.claude/skills/<name>/`,
40
+ `~/.agents/skills/`, `~/.codex/skills/`, `~/.cursor/skills/`, `~/.roo/skills/` —
41
+ or this project's own `skills/` tree if it has
42
+ one) — stop and tell the user if one doesn't, never silently proceed
43
+ without it — and include it as a **hard requirement, not a suggestion** in
44
+ the Plan node's and every Worker node's prompt. The Review node checks the
45
+ combined output for evidence the skill was actually applied, not just
46
+ mentioned, and calls that out explicitly if it wasn't. Nothing about this
47
+ gets persisted, same as everything else in this skill — it's a per-run
48
+ instruction, not a contract. See
49
+ [`.agents/graph.md`](../../../.agents/graph.md)'s "Mandatory Skill
50
+ Injection" section for the same mechanic in the durable graph variant.
51
+
32
52
  ## 2. Detect what's available — before deciding how workers run
33
53
 
34
54
  ```bash
@@ -39,21 +59,53 @@ command -v codex >/dev/null 2>&1 && echo codex
39
59
 
40
60
  This machine may have none of these — the skill (and whoever installed this
41
61
  package) cannot assume Antigravity, OpenCode, or a Codex plugin connector is
42
- present. Pick the worker path in this priority order, first one found wins:
43
-
44
- 1. **agy present** → workers run via `agy-job start --tier flash [--yolo] "<task>"`
45
- (or `pro` for harder reasoning) — see the `antigravity` skill for the exact
46
- invocation pattern and cost discipline. Separate compute pool, zero
47
- Anthropic tokens for the work itself.
48
- 2. **opencode present** → route worker tasks through it the same way (its own
49
- subagent/session primitive), if the project already uses it.
50
- 3. **codex present** → same idea, via the Codex CLI's own task-delegation
51
- surface if this project has that plugin wired up.
52
- 4. **none present** → fall back to Claude Code's own `Agent` tool for each
62
+ present.
63
+
64
+ **Prefer a plugin's own delegation subagent over shelling out to its CLI.**
65
+ If a delegation plugin is installed, it exposes a subagent that already
66
+ handles the wrapper flags, the cost discipline, and the digest contract for
67
+ you — reach for that first, and only drop to a raw CLI call when no such
68
+ subagent exists:
69
+
70
+ | CLI | Plugin subagent (preferred) | Raw fallback |
71
+ |---|---|---|
72
+ | agy | `antigravity:antigravity-delegate` | `agy-job start --tier flash [--yolo] "<task>"` |
73
+ | opencode | `opencode:opencode-rescue` | the CLI's own session primitive |
74
+ | codex | `codex:codex-rescue` | the Codex CLI's task-delegation surface |
75
+
76
+ These subagents are **Claude Code plugins**, so they exist only when that
77
+ harness is running *and* the plugin is installed. Check what is actually
78
+ available rather than assuming — on any other harness, or a machine without
79
+ the plugins, the raw CLI column is the only path. Neither column ships with
80
+ AOS: both depend on tooling the user installed separately.
81
+
82
+ Pick the worker path in this priority order, first one found wins:
83
+
84
+ 1. **A delegation plugin subagent is available** → use it (table above).
85
+ Separate compute pool, zero Anthropic tokens for the work itself, and the
86
+ wrapper reports failures in a shape the plugin already knows how to read.
87
+ 2. **The CLI is present but its plugin subagent is not** → call the CLI
88
+ directly per the raw-fallback column, following the `antigravity` skill's
89
+ invocation pattern and cost discipline.
90
+ 3. **None present** → fall back to Claude Code's own `Agent` tool for each
53
91
  worker, with an explicit `model: "haiku"` override. This is the only path
54
92
  that costs Anthropic tokens for the worker step, and the only one
55
93
  guaranteed to exist everywhere — it is the floor, not the default.
56
94
 
95
+ **Give the delegation a real timeout.** Measured 2026-09: a trivial headless
96
+ `agy` prompt took **605s**. `agy-delegate` defaults to `--print-timeout 5m`,
97
+ so it aborts at 300s and reports an empty body while the answer is still on
98
+ its way — pass `--timeout 15m` for anything non-trivial. Budget worker
99
+ wall-clock accordingly; this is the single most likely reason a fan-out
100
+ "fails" on a machine where the CLI is perfectly healthy.
101
+
102
+ **Verify the delegation actually produced content — a status string is not a
103
+ result.** A timed-out delegation returns `{"status": "SUCCESS", "usage":
104
+ {"total": 0}}` with an *empty* body: success by every field except the one
105
+ that matters. The zero token counts are not proof the prompt never arrived —
106
+ headless usage reporting is simply unpopulated. Check the returned text
107
+ itself, and treat an empty body as a failure no matter what the status says.
108
+
57
109
  **This decision happens here, in your own turn, via Bash — never inside a
58
110
  `Workflow` script.** A `Workflow` script's body has no shell or filesystem
59
111
  access (ambient `agent()`/`pipeline()` globals only), so it cannot itself
@@ -104,6 +156,8 @@ graph (`.agents/graph.md`) already does properly.
104
156
  | "The session is already on Opus, so the worker call inherits it fine." | That's exactly the cost this skill exists to avoid — force the tier explicitly every time. |
105
157
  | "This task has one obvious step, but a 3-node graph looks more thorough." | More nodes than the task needs is overhead, not rigor. Size the graph to the work. |
106
158
  | "I'll just call the Workflow tool and let the script figure out which backend to use." | The script can't — it has no shell access. That decision is yours, before the Workflow call, or not via Workflow at all. |
159
+ | "The CLI is on PATH, so I'll shell out to it directly." | If its plugin subagent is installed, that's the supported path — it already handles the wrapper flags and cost discipline. Shell out only when no subagent exists. |
160
+ | "The wrapper returned SUCCESS, so the work is done." | A failing delegation has returned `SUCCESS` with zero tokens and an empty body. Check the actual content, not the status field. |
107
161
 
108
162
  ## 7. Red Flags
109
163
 
@@ -115,5 +169,7 @@ graph (`.agents/graph.md`) already does properly.
115
169
  ## 8. Verification
116
170
 
117
171
  - [ ] Detection step actually ran (`command -v` checks), not assumed.
172
+ - [ ] A plugin delegation subagent was preferred where one was available, rather than shelling out to the CLI anyway.
118
173
  - [ ] Each node's model was explicitly forced, not inherited.
174
+ - [ ] Each worker's returned **content** was checked, not just its status field — an empty body is a failure regardless of a `SUCCESS` status.
119
175
  - [ ] Nothing persistent was left behind after the task completed.