@hybridlabor-api/aos 4.1.0 → 4.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/agents.md +77 -0
- package/.agents/graph.md +45 -0
- package/.agents/nodes.json +4 -2
- package/.agents/state.schema.json +6 -0
- package/.claude/workflows/startcycle-dispatch.mjs +139 -8
- package/.claude/workflows/teamwork-dispatch.mjs +287 -0
- package/CLAUDE.md +47 -0
- package/GEMINI.md +9 -1
- package/README.md +12 -6
- package/THIRD_PARTY_NOTICES.md +50 -0
- package/docs/skills_table.md +1 -0
- package/installer.js +15 -0
- package/package.json +1 -1
- package/skills/basic/bdbmediastorm/SKILL.md +7 -5
- package/skills/basic/startcycle/SKILL.md +21 -0
- package/skills/basic/startcycle-graph/SKILL.md +43 -7
- package/skills/basic/startcycle-graph-user/SKILL.md +67 -11
- package/skills/basic/teamwork-preview/SKILL.md +209 -0
- package/skills/bdbrainstorm/SKILL.md +4 -3
- package/skills/global_config/ask-tim/SKILL.md +73 -6
- package/skills/global_config/bdbresilience/SKILL.md +216 -0
- package/skills/global_config/bdbresilience/contracts/nodes-integration.md +225 -0
- package/skills/global_config/bdbresilience/references/cicd-triage.md +179 -0
- package/skills/global_config/bdbresilience/references/distributed-locking.md +235 -0
- package/skills/global_config/bdbresilience/references/error-recovery.md +210 -0
- package/skills/global_config/bdbresilience/references/two-phase-go-gate.md +151 -0
- package/skills/global_config/domain-modeling/ADR-FORMAT.md +47 -0
- package/skills/global_config/domain-modeling/CONTEXT-FORMAT.md +60 -0
- package/skills/global_config/domain-modeling/SKILL.md +77 -0
- package/skills/global_config/grill-me/SKILL.md +14 -0
- package/skills/global_config/grill-with-docs/SKILL.md +24 -0
- package/skills/global_config/grilling/SKILL.md +42 -0
- package/skills/global_config/openwiki-skill/scripts/install_daemon.sh +55 -12
package/CLAUDE.md
CHANGED
|
@@ -15,6 +15,53 @@ Ask one question first: **do the workers need to see each other?**
|
|
|
15
15
|
|
|
16
16
|
"Runs in parallel" is not a reason to reach for a team — subagents already run in parallel. Peer communication and dynamic task claiming are the only things a team adds.
|
|
17
17
|
|
|
18
|
+
## Delegating to an external CLI
|
|
19
|
+
Some work is cheaper on another provider's compute (bulk scaffolding, exhaustive
|
|
20
|
+
test generation, long-context reads that distil to a digest). None of that tooling
|
|
21
|
+
ships with AOS — it depends on CLIs and Claude Code plugins the user installed
|
|
22
|
+
separately, so check what is actually present instead of assuming.
|
|
23
|
+
|
|
24
|
+
**Prefer a plugin's delegation subagent over shelling out to its CLI.** Where one
|
|
25
|
+
is installed it already handles the wrapper flags, cost discipline, and digest
|
|
26
|
+
contract: `antigravity:antigravity-delegate` (agy), `opencode:opencode-rescue`,
|
|
27
|
+
`codex:codex-rescue`. These are Claude Code plugins — on another harness, or a
|
|
28
|
+
machine without them, calling the CLI directly is the only path.
|
|
29
|
+
|
|
30
|
+
**Delegate only above the break-even.** A small, self-contained, or
|
|
31
|
+
judgement-heavy task costs more to hand off and verify than to just do. Keep the
|
|
32
|
+
digest, not the raw output.
|
|
33
|
+
|
|
34
|
+
**Give it a real timeout.** Measured 2026-09: a trivial headless `agy` prompt
|
|
35
|
+
took **605s**. `agy-delegate` defaults to `--print-timeout 5m`, so it aborts at
|
|
36
|
+
300s and reports an empty body while the answer is still coming — pass
|
|
37
|
+
`--timeout 15m` for anything non-trivial. A short timeout does not read as
|
|
38
|
+
"slow", it reads as "broken".
|
|
39
|
+
|
|
40
|
+
**Match the model to the task, not to the default.** `agy-delegate`'s tiers map
|
|
41
|
+
to models that can go stale (its built-in `flash` still points at Gemini 3.7
|
|
42
|
+
while 3.8 ships). Either pass `--model "<exact name from \`agy models\`>"` per
|
|
43
|
+
call, or remap the tiers once via the plugin's own options — as env vars those
|
|
44
|
+
belong in `~/.zshenv`, not `~/.zshrc`, since `.zshrc` is only sourced for
|
|
45
|
+
interactive shells and tool-invoked ones would never see them:
|
|
46
|
+
|
|
47
|
+
| Work | Model |
|
|
48
|
+
|---|---|
|
|
49
|
+
| media, fast/mechanical coding, boilerplate | `Gemini 3.8 Flash (Medium)` → `CLAUDE_PLUGIN_OPTION_TIER_FLASH` |
|
|
50
|
+
| trivial one-liners | `Gemini 3.8 Flash (Low)` → `CLAUDE_PLUGIN_OPTION_TIER_FLASH_LO` |
|
|
51
|
+
| review, architecture, hard reasoning | `Claude Sonnet 4.6 (Thinking)` → `CLAUDE_PLUGIN_OPTION_TIER_PRO` |
|
|
52
|
+
|
|
53
|
+
Adversarial review is the case that most repays a stronger model: a Flash tier
|
|
54
|
+
tends to agree with what it is shown, which is the one thing a reviewer must
|
|
55
|
+
not do. Re-check the names against `agy models` after an agy upgrade — the id
|
|
56
|
+
carries both the version and the effort suffix.
|
|
57
|
+
|
|
58
|
+
**Verify the result, never the status field.** A timed-out delegation returns
|
|
59
|
+
`{"status": "SUCCESS", "usage": {"total": 0}}` with an empty body — success by
|
|
60
|
+
every field except the one that matters, and the zero token counts are *not*
|
|
61
|
+
proof the prompt never arrived (headless usage reporting is simply unpopulated).
|
|
62
|
+
Check the returned content, treat an empty body as failure regardless of status,
|
|
63
|
+
and never report a delegated step as done on the strength of its own self-report.
|
|
64
|
+
|
|
18
65
|
## Safety Gate — mechanically enforced, not advisory
|
|
19
66
|
`git push`, `npm publish`, `npm version`, and recursive `rm` are blocked by `.claude/hooks/go-gate.mjs` (registered in `.claude/settings.json`) unless your immediately preceding message is the literal word **GO**. This is a hook, not a rule I read and try to follow — it cannot be argued around, and it doesn't depend on this file being loaded.
|
|
20
67
|
- A subagent does not inherit its orchestrator's GO.
|
package/GEMINI.md
CHANGED
|
@@ -47,4 +47,12 @@ Ask one question first: **do the workers need to see each other?**
|
|
|
47
47
|
- **Yes — they must react to each other, or claim work dynamically from a shared list** → an orchestrated agent team (via `send_message`). Currently only `/bdbrainstorm` qualifies, where the spec demands a real debate rather than parallel monologues. True Agent Teams were evaluated and deferred for `/startcycle-graph` (needs an interactive session; the graph runs headless) — see `.agents/graph.md` and F-17's addendum in `docs/sessions/audit-agents.md`.
|
|
48
48
|
- **Small task** → do it yourself. A two-file edit needs no agents.
|
|
49
49
|
|
|
50
|
-
"Runs in parallel" is not a reason to reach for a team — subagents already run in parallel. Peer communication and dynamic task claiming are the only things a team adds.
|
|
50
|
+
"Runs in parallel" is not a reason to reach for a team — subagents already run in parallel. Peer communication and dynamic task claiming are the only things a team adds.
|
|
51
|
+
|
|
52
|
+
## 9. Delegating to an external CLI
|
|
53
|
+
None of this tooling ships with AOS — it depends on CLIs and plugins the user installed separately, so check what is present rather than assuming.
|
|
54
|
+
- **Prefer a plugin's delegation subagent over shelling out to its CLI.** Where installed, it already handles the wrapper flags, cost discipline, and digest contract: `antigravity:antigravity-delegate` (agy), `opencode:opencode-rescue`, `codex:codex-rescue`. These are Claude Code plugins — on another harness, calling the CLI directly is the only path.
|
|
55
|
+
- **Delegate only above the break-even.** A small, self-contained, or judgement-heavy task costs more to hand off and verify than to just do. Keep the digest, not the raw output.
|
|
56
|
+
- **Give it a real timeout.** Measured 2026-09: a trivial headless `agy` prompt took **605s**. `agy-delegate` defaults to `--print-timeout 5m`, so it aborts at 300s and reports an empty body while the answer is still coming — pass `--timeout 15m` for anything non-trivial.
|
|
57
|
+
- **Match the model to the task, not to the default.** The wrapper's tiers map to models that go stale (built-in `flash` still points at Gemini 3.7 while 3.8 ships). Media and fast/mechanical coding → `Gemini 3.8 Flash (Medium)`; trivial → `Gemini 3.8 Flash (Low)`; review, architecture and hard reasoning → `Claude Sonnet 4.6 (Thinking)`. Adversarial review most repays the stronger model: a Flash tier tends to agree with what it is shown, which is exactly what a reviewer must not do. Pass `--model` per call, or remap the tiers once via `CLAUDE_PLUGIN_OPTION_TIER_{FLASH,FLASH_LO,PRO}` — in `~/.zshenv`, not `~/.zshrc`, which non-interactive tool shells never source. Re-check names against `agy models` after an upgrade.
|
|
58
|
+
- **Verify the result, never the status field.** A timed-out delegation returns `{"status": "SUCCESS", "usage": {"total": 0}}` with an empty body — success by every field except the one that matters, and the zero token counts are *not* proof the prompt never arrived (headless usage reporting is simply unpopulated). Treat an empty body as failure regardless of status, and never report a delegated step as done on the strength of its own self-report.
|
package/README.md
CHANGED
|
@@ -5,13 +5,19 @@
|
|
|
5
5
|
|
|
6
6
|
[](https://github.com/hybridlabor-api/aos/actions)
|
|
7
7
|
[](https://www.npmjs.com/package/@hybridlabor-api/aos)
|
|
8
|
-
[](https://www.npmjs.com/package/@hybridlabor-api/aos)
|
|
9
9
|
[](LICENSE)
|
|
10
|
-
[](https://github.com/hybridlabor-api/aos/stargazers)
|
|
11
|
+
[](https://github.com/hybridlabor-api/aos/commits/main)
|
|
12
|
+
|
|
13
|
+
[](docs/skills_table.md)
|
|
14
|
+
[](mcps/)
|
|
15
|
+
[](#-installation)
|
|
16
|
+
[](https://github.com/hybridlabor-api/aos)
|
|
11
17
|
|
|
12
|
-
> **Supercharging AI coding agents with
|
|
18
|
+
> **Supercharging AI coding agents with 169 hyper-curated skills, 21 local MCP wrappers, and a runnable multi-agent dispatcher graph.**
|
|
13
19
|
|
|
14
|
-
Welcome to **BDB Agent OS — AOS v4.0.0**:
|
|
20
|
+
Welcome to **BDB Agent OS — AOS v4.0.0**: 169 curated skills, 21 local MCP wrappers, and a dispatcher graph that turns them into a real multi-agent build pipeline, not just a prompt library. Point it at a goal and it plans, builds, reviews, and ships through seven coordinated agent nodes — with a mechanically enforced gate before anything actually goes live.
|
|
15
21
|
|
|
16
22
|
It is harness-neutral by design, not "optimized for one tool with others as an afterthought": the dispatcher graph runs on Claude Code's Dynamic Workflows, the same skills and MCP configuration install natively into **Google Antigravity, ChatGPT Codex / Codex CLI, Claude Desktop, Cursor, Aider, Roo Code, Cline, and Windsurf**, and the lightweight `/startcycle-graph-user` variant falls back to Claude Code's own subagents on any machine that has none of the above installed.
|
|
17
23
|
|
|
@@ -95,9 +101,9 @@ below for how the pipeline itself works.
|
|
|
95
101
|
|
|
96
102
|
---
|
|
97
103
|
|
|
98
|
-
## 🌟
|
|
104
|
+
## 🌟 169 Optimized Skills
|
|
99
105
|
|
|
100
|
-
We started with a massive pool of over 1,400 raw AI skills. After rigorous testing, filtering, and refinement, we've distilled them down to a hyper-curated set of **
|
|
106
|
+
We started with a massive pool of over 1,400 raw AI skills. After rigorous testing, filtering, and refinement, we've distilled them down to a hyper-curated set of **169 Optimized Skills** (featuring a native OpenWiki documentation engine, the **memB local semantic memory brain**, and **Universal Agent Harness synchronization**).
|
|
101
107
|
|
|
102
108
|
These skills are precision-engineered to ensure agents waste no time on redundant tasks and instead operate with maximum agency, strict architectural constraints, and robust context awareness.
|
|
103
109
|
|
package/THIRD_PARTY_NOTICES.md
CHANGED
|
@@ -127,6 +127,56 @@ SOFTWARE.
|
|
|
127
127
|
|
|
128
128
|
---
|
|
129
129
|
|
|
130
|
+
## mattpocock/skills
|
|
131
|
+
|
|
132
|
+
<https://github.com/mattpocock/skills> — MIT.
|
|
133
|
+
|
|
134
|
+
The grilling family and the domain-modeling discipline it composes with:
|
|
135
|
+
|
|
136
|
+
| In AOS | Upstream |
|
|
137
|
+
|---|---|
|
|
138
|
+
| `skills/global_config/grilling/` | `skills/productivity/grilling/` |
|
|
139
|
+
| `skills/global_config/grill-me/` | `skills/productivity/grill-me/` |
|
|
140
|
+
| `skills/global_config/grill-with-docs/` | `skills/engineering/grill-with-docs/` |
|
|
141
|
+
| `skills/global_config/domain-modeling/` | `skills/engineering/domain-modeling/` (incl. `ADR-FORMAT.md`, `CONTEXT-FORMAT.md`) |
|
|
142
|
+
|
|
143
|
+
`grilling`'s interview protocol and `domain-modeling` are carried over
|
|
144
|
+
essentially verbatim; `grill-me` and `grill-with-docs` are rewritten to name AOS's
|
|
145
|
+
own pipelines in their hand-off sections, but keep upstream's composition — both
|
|
146
|
+
are thin wrappers that invoke the primitive rather than restating it. Upstream's
|
|
147
|
+
`agents/openai.yaml` under `domain-modeling` is harness-specific to that project
|
|
148
|
+
and was not carried over.
|
|
149
|
+
|
|
150
|
+
`skills/global_config/ask-tim/` is derived from upstream's `ask-matt` router. It
|
|
151
|
+
is not a copy: the flow it maps is AOS's own, since most of the skills on
|
|
152
|
+
upstream's main flow have no AOS equivalent.
|
|
153
|
+
|
|
154
|
+
```
|
|
155
|
+
MIT License
|
|
156
|
+
|
|
157
|
+
Copyright (c) 2026 Matt Pocock
|
|
158
|
+
|
|
159
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
160
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
161
|
+
in the Software without restriction, including without limitation the rights
|
|
162
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
163
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
164
|
+
furnished to do so, subject to the following conditions:
|
|
165
|
+
|
|
166
|
+
The above copyright notice and this permission notice shall be included in all
|
|
167
|
+
copies or substantial portions of the Software.
|
|
168
|
+
|
|
169
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
170
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
171
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
172
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
173
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
174
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
175
|
+
SOFTWARE.
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
---
|
|
179
|
+
|
|
130
180
|
## Note on `mcps/`
|
|
131
181
|
|
|
132
182
|
Sub-repositories vendored under `mcps/` carry their own `LICENSE` files in
|
package/docs/skills_table.md
CHANGED
|
@@ -48,6 +48,7 @@
|
|
|
48
48
|
| `subagent-driven-development` | Use when executing implementation plans with independent tasks in the current session |
|
|
49
49
|
| `tdd-workflow` | Test-Driven Development workflow principles. RED-GREEN-REFACTOR cycle. |
|
|
50
50
|
| `tmux` | Expert tmux session, window, and pane management for terminal multiplexing, persistent remote workflows, and shell scripting automation. |
|
|
51
|
+
| `teamwork-preview` | Interactive 9-step prompt crafting and delegation protocol for autonomous multi-agent teams across harnesses. |
|
|
51
52
|
|
|
52
53
|
#### 🎨 Frontend & UI/UX
|
|
53
54
|
| Skill Name | Description |
|
package/installer.js
CHANGED
|
@@ -2554,6 +2554,21 @@ function injectHarnessRules() {
|
|
|
2554
2554
|
log.step(`Installed GEMINI.md to ${path.join(geminiDir, 'GEMINI.md')}`);
|
|
2555
2555
|
}, 'The harness injection below still runs.');
|
|
2556
2556
|
|
|
2557
|
+
// Dispatcher scripts must land in ~/.claude/workflows/, because that is
|
|
2558
|
+
// where the skills that route to them look: startcycle-graph's SKILL.md
|
|
2559
|
+
// and teamwork-preview's both tell the model to call `Workflow` with
|
|
2560
|
+
// scriptPath `$HOME/.claude/workflows/<name>.mjs`. installProjectHarness()
|
|
2561
|
+
// copies them into a *project*, which only helps a project that opted into
|
|
2562
|
+
// the local harness -- on a plain global install those paths did not exist
|
|
2563
|
+
// at all, so the skill pointed at a file that was never delivered.
|
|
2564
|
+
installStep(`install dispatcher workflows to ${path.join(homeDir, '.claude', 'workflows')}`, () => {
|
|
2565
|
+
const workflowsSrc = path.join(srcDir, '.claude', 'workflows');
|
|
2566
|
+
if (fs.existsSync(workflowsSrc)) {
|
|
2567
|
+
copyDirRecursiveSync(workflowsSrc, path.join(homeDir, '.claude', 'workflows'));
|
|
2568
|
+
log.step(`Installed dispatcher workflows to ${path.join(homeDir, '.claude', 'workflows')}`);
|
|
2569
|
+
}
|
|
2570
|
+
}, '/startcycle-graph and /teamwork-preview fall back to their prose protocols.');
|
|
2571
|
+
|
|
2557
2572
|
const startcycleWorkflowSrc = path.join(srcDir, '.agents', 'workflows', 'startcycle.md');
|
|
2558
2573
|
const sources = installStep('read the global rule sources', () => ({
|
|
2559
2574
|
globalRules: fs.readFileSync(geminiMdSrc, 'utf8'),
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@hybridlabor-api/aos",
|
|
3
|
-
"version": "4.
|
|
3
|
+
"version": "4.2.0",
|
|
4
4
|
"description": "AOS — A Curated AI AGENT OS. Optimized agent skills and add-ons like memB, OpenWiki, Heimdall Token Saver, and Godmode architectures.",
|
|
5
5
|
"main": "installer.js",
|
|
6
6
|
"bin": {
|
|
@@ -26,9 +26,11 @@ Ideation must never be performed in isolation. Spawn specialized subagents to an
|
|
|
26
26
|
|
|
27
27
|
---
|
|
28
28
|
|
|
29
|
-
## 2. Interactive
|
|
29
|
+
## 2. Interactive Technical Interview
|
|
30
30
|
|
|
31
|
-
Before drafting signal flow diagrams or system configs,
|
|
31
|
+
Before drafting signal flow diagrams or system configs, **invoke the `grill-with-docs` skill** (or `grill-me` when there is no working directory) and run it to completion. Those skills hold the interview protocol — design tree, frontier rounds, numbered questions each with a recommended answer — and it is not restated here.
|
|
32
|
+
|
|
33
|
+
What this domain adds to that protocol: deeply challenge the user's technical assumptions and hardware readiness. The frontier questions for a show-control build are:
|
|
32
34
|
|
|
33
35
|
* **Signal & Network Protocols:**
|
|
34
36
|
- What protocols govern data movement? (OSC, Art-Net, sACN, MIDI, SMPTE Timecode, NDI)?
|
|
@@ -48,7 +50,7 @@ Before drafting signal flow diagrams or system configs, execute a mandatory `/gr
|
|
|
48
50
|
|
|
49
51
|
## 3. Target Directory & Scaffolding
|
|
50
52
|
|
|
51
|
-
After aligning on system architecture through the
|
|
53
|
+
After aligning on system architecture through the grilling interview:
|
|
52
54
|
1. **Confirm Output Directory:** Ask the user: *"In which project directory should the output show-control plan and architecture files be stored?"*
|
|
53
55
|
2. **Scaffold Foundational Files:** Once confirmed, write the core show specification files (`agent.md`, `signal-flow.md`, `network-patch.json`, `failover-matrix.md`).
|
|
54
56
|
|
|
@@ -94,7 +96,7 @@ BDB MediaStorm is the master ideation and brainstorming engine for live show-con
|
|
|
94
96
|
- **Exclude:** Do not use for generating video timelines or 3D meshes.
|
|
95
97
|
|
|
96
98
|
## Core Process
|
|
97
|
-
1. Run
|
|
99
|
+
1. Run `grill-with-docs` (or `grill-me`) to challenge assumptions about protocols, hardware, and bandwidth.
|
|
98
100
|
2. Scaffold foundational files (`agent.md`, `signal-flow.md`, `network-patch.json`).
|
|
99
101
|
3. Generate a strict Mermaid.js signal flow diagram mapping all protocols.
|
|
100
102
|
4. Document a main/backup redundancy and failover matrix.
|
|
@@ -115,7 +117,7 @@ BDB MediaStorm is the master ideation and brainstorming engine for live show-con
|
|
|
115
117
|
|
|
116
118
|
## Verification
|
|
117
119
|
|
|
118
|
-
- [ ]
|
|
120
|
+
- [ ] The grilling interview was completed with answers regarding protocols and bandwidth.
|
|
119
121
|
- [ ] Output includes a Mermaid.js signal flow diagram.
|
|
120
122
|
- [ ] A dedicated failover/blackout mechanism is documented.
|
|
121
123
|
|
|
@@ -80,6 +80,27 @@ Run only the streams the goal actually needs. A plain backend feature does not n
|
|
|
80
80
|
|
|
81
81
|
3c is for TouchDesigner, show-control, DMX/grandMA3, 3D, or other media-pipeline goals — most goals are not this. Skip it unless the plan actually calls for it.
|
|
82
82
|
|
|
83
|
+
> **Injecting a specific skill.** `/startcycle --skill=<name> <goal>`
|
|
84
|
+
> (repeatable, quote a name with spaces) forces that skill into this run as
|
|
85
|
+
> a hard requirement — for a private skill of your own that isn't part of
|
|
86
|
+
> any build agent's own `skills:` frontmatter. Since this variant has no
|
|
87
|
+
> dispatcher script or `state.json` to carry it automatically, the invoker
|
|
88
|
+
> does the work `startcycle-graph`'s script does for you: extract the
|
|
89
|
+
> `--skill=` flag(s) from the invocation text before anything else runs,
|
|
90
|
+
> confirm each name resolves to a real `SKILL.md` (under
|
|
91
|
+
> any harness's global skills directory — `~/.claude/skills/<name>/`,
|
|
92
|
+
> `~/.agents/skills/`, `~/.codex/skills/`, `~/.cursor/skills/` or `~/.roo/skills/`
|
|
93
|
+
> — or this project's own `skills/` tree) — stop
|
|
94
|
+
> and tell the user if one doesn't, don't silently proceed without it —
|
|
95
|
+
> note the validated list in `00_execution_plan.md`, and include it as a
|
|
96
|
+
> **hard requirement, not a suggestion** in each Build stream's dispatch
|
|
97
|
+
> prompt at step 3. Reviewer (step 4) checks the resulting artifacts for
|
|
98
|
+
> evidence the skill was actually applied, not just available, and treats
|
|
99
|
+
> an ignored mandate as a contract-misread finding. See
|
|
100
|
+
> [`.agents/graph.md`](../../../.agents/graph.md)'s "Mandatory Skill
|
|
101
|
+
> Injection" section for the full rationale — this is the same mechanic,
|
|
102
|
+
> just invoker-driven instead of script-driven.
|
|
103
|
+
|
|
83
104
|
### 4. Reviewer
|
|
84
105
|
- **Agent**: `reviewer`
|
|
85
106
|
- **Reads**: the artifacts each build stream produced (01/02/03) and the plan's stated contract (`00_execution_plan.md`) — nothing else. Never the goal directly, never a build agent's own claim that it's done; passing that claim through biases the review toward agreement.
|
|
@@ -21,12 +21,21 @@ several fields the schema doesn't define at all). Fix it first:
|
|
|
21
21
|
mkdir -p .agents
|
|
22
22
|
[ -f .agents/graph.md ] || cp "$HOME/.agents/graph.md" .agents/graph.md
|
|
23
23
|
[ -f .agents/state.schema.json ] || cp "$HOME/.agents/state.schema.json" .agents/state.schema.json
|
|
24
|
+
[ -f .agents/nodes.json ] || cp "$HOME/.agents/nodes.json" .agents/nodes.json
|
|
24
25
|
```
|
|
25
26
|
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
the
|
|
29
|
-
|
|
27
|
+
`nodes.json` is not optional and is the one that fails loudest: it is the
|
|
28
|
+
node registry the dispatcher loads as its very first step, and without it
|
|
29
|
+
the run escalates immediately with *"`.agents/nodes.json` failed to load, or
|
|
30
|
+
is missing required node id(s)"* — before Architect or any other agent has
|
|
31
|
+
run. (Found exactly that way: a first run in a project that had followed
|
|
32
|
+
this bootstrap step as it was previously written, which copied only the
|
|
33
|
+
other two files.)
|
|
34
|
+
|
|
35
|
+
If any of the three doesn't exist under `$HOME/.agents/` either, stop and
|
|
36
|
+
tell the user: this machine has no canonical copy of the graph contract to
|
|
37
|
+
bootstrap from, and `/startcycle-graph` cannot run correctly until one is
|
|
38
|
+
installed. Don't silently proceed.
|
|
30
39
|
(`.claude/agents/*.md`, the seven agent persona files, do NOT need this
|
|
31
40
|
treatment — Claude Code resolves subagents from the user-level
|
|
32
41
|
`~/.claude/agents/` fine without a project-local copy.)
|
|
@@ -37,9 +46,36 @@ yourself (e.g. `echo $HOME` or your own environment info) rather than
|
|
|
37
46
|
hardcoding a username, giving
|
|
38
47
|
`$HOME/.claude/workflows/startcycle-dispatch.mjs` — and `args` set to the
|
|
39
48
|
goal text that follows `ARGUMENTS:` below this file's content. Pass the goal
|
|
40
|
-
through verbatim
|
|
41
|
-
|
|
42
|
-
|
|
49
|
+
through verbatim, including any `--skill=<name>` flag(s) it contains — the
|
|
50
|
+
dispatcher script parses those itself (see below), do not strip or
|
|
51
|
+
interpret them yourself. If there is no `ARGUMENTS:` text, pass no `args`
|
|
52
|
+
(or `args: undefined`) — the workflow itself asks for a goal in that case
|
|
53
|
+
rather than guessing one.
|
|
54
|
+
|
|
55
|
+
**Injecting a specific skill.** `/startcycle-graph --skill=<name> <goal>`
|
|
56
|
+
(repeatable, quote a name with spaces) forces that skill into this run as a
|
|
57
|
+
hard requirement for the build nodes, validated to exist before anything
|
|
58
|
+
else runs — this is how you make the pipeline use your own private skill
|
|
59
|
+
that isn't part of `.agents/nodes.json`'s registry.
|
|
60
|
+
|
|
61
|
+
`<name>` is the **exact skill directory name**, not a description — `ui-component`,
|
|
62
|
+
not "the UI one". Two ways to find it without leaving the terminal:
|
|
63
|
+
|
|
64
|
+
- `/ask-tim` — the routing skill; start there when you know the *job* but not the name
|
|
65
|
+
- list the installed skills directly, if you half-remember the spelling. The installer
|
|
66
|
+
syncs the same set to every harness it detects, so use whichever path is yours:
|
|
67
|
+
`~/.claude/skills`, `~/.agents/skills`, `~/.codex/skills`, `~/.cursor/skills`, or
|
|
68
|
+
`~/.roo/skills`. `ls ~/.agents/skills` is the safest guess on an unknown machine —
|
|
69
|
+
that one is written on every install regardless of harness.
|
|
70
|
+
|
|
71
|
+
A name that does not resolve halts the run before any agent works, and the
|
|
72
|
+
error now lists installed near-misses rather than only saying "not found".
|
|
73
|
+
That is deliberate: silently running without a skill you explicitly demanded
|
|
74
|
+
is worse than stopping.
|
|
75
|
+
|
|
76
|
+
See [`.agents/graph.md`](../../../.agents/graph.md)'s "Mandatory Skill
|
|
77
|
+
Injection" section for the full mechanics; nothing about it needs handling
|
|
78
|
+
in this router file, since `args` is passed through as raw text either way.
|
|
43
79
|
|
|
44
80
|
Use `scriptPath`, not `name: "startcycle-dispatch"` — by-name lookup for a
|
|
45
81
|
custom (non-built-in) workflow script has been observed to fail with
|
|
@@ -29,6 +29,26 @@ Do not invent more structure than the task has. A two-file edit doesn't need
|
|
|
29
29
|
a Plan node — just do it. This skill is for the cases actually shaped like a
|
|
30
30
|
small graph, not an excuse to always draw one.
|
|
31
31
|
|
|
32
|
+
## 1b. Injecting a specific skill (optional)
|
|
33
|
+
|
|
34
|
+
`/startcycle-graph-user --skill=<name> <task>` (repeatable, quote a name
|
|
35
|
+
with spaces) forces that skill into this run as a hard requirement — for a
|
|
36
|
+
private skill of the user's own this throwaway graph would otherwise never
|
|
37
|
+
know to reach for. Extract any `--skill=` flag(s) from the invocation text
|
|
38
|
+
before step 1, confirm each name resolves to a real `SKILL.md` (under
|
|
39
|
+
any harness's global skills directory — `~/.claude/skills/<name>/`,
|
|
40
|
+
`~/.agents/skills/`, `~/.codex/skills/`, `~/.cursor/skills/`, `~/.roo/skills/` —
|
|
41
|
+
or this project's own `skills/` tree if it has
|
|
42
|
+
one) — stop and tell the user if one doesn't, never silently proceed
|
|
43
|
+
without it — and include it as a **hard requirement, not a suggestion** in
|
|
44
|
+
the Plan node's and every Worker node's prompt. The Review node checks the
|
|
45
|
+
combined output for evidence the skill was actually applied, not just
|
|
46
|
+
mentioned, and calls that out explicitly if it wasn't. Nothing about this
|
|
47
|
+
gets persisted, same as everything else in this skill — it's a per-run
|
|
48
|
+
instruction, not a contract. See
|
|
49
|
+
[`.agents/graph.md`](../../../.agents/graph.md)'s "Mandatory Skill
|
|
50
|
+
Injection" section for the same mechanic in the durable graph variant.
|
|
51
|
+
|
|
32
52
|
## 2. Detect what's available — before deciding how workers run
|
|
33
53
|
|
|
34
54
|
```bash
|
|
@@ -39,21 +59,53 @@ command -v codex >/dev/null 2>&1 && echo codex
|
|
|
39
59
|
|
|
40
60
|
This machine may have none of these — the skill (and whoever installed this
|
|
41
61
|
package) cannot assume Antigravity, OpenCode, or a Codex plugin connector is
|
|
42
|
-
present.
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
62
|
+
present.
|
|
63
|
+
|
|
64
|
+
**Prefer a plugin's own delegation subagent over shelling out to its CLI.**
|
|
65
|
+
If a delegation plugin is installed, it exposes a subagent that already
|
|
66
|
+
handles the wrapper flags, the cost discipline, and the digest contract for
|
|
67
|
+
you — reach for that first, and only drop to a raw CLI call when no such
|
|
68
|
+
subagent exists:
|
|
69
|
+
|
|
70
|
+
| CLI | Plugin subagent (preferred) | Raw fallback |
|
|
71
|
+
|---|---|---|
|
|
72
|
+
| agy | `antigravity:antigravity-delegate` | `agy-job start --tier flash [--yolo] "<task>"` |
|
|
73
|
+
| opencode | `opencode:opencode-rescue` | the CLI's own session primitive |
|
|
74
|
+
| codex | `codex:codex-rescue` | the Codex CLI's task-delegation surface |
|
|
75
|
+
|
|
76
|
+
These subagents are **Claude Code plugins**, so they exist only when that
|
|
77
|
+
harness is running *and* the plugin is installed. Check what is actually
|
|
78
|
+
available rather than assuming — on any other harness, or a machine without
|
|
79
|
+
the plugins, the raw CLI column is the only path. Neither column ships with
|
|
80
|
+
AOS: both depend on tooling the user installed separately.
|
|
81
|
+
|
|
82
|
+
Pick the worker path in this priority order, first one found wins:
|
|
83
|
+
|
|
84
|
+
1. **A delegation plugin subagent is available** → use it (table above).
|
|
85
|
+
Separate compute pool, zero Anthropic tokens for the work itself, and the
|
|
86
|
+
wrapper reports failures in a shape the plugin already knows how to read.
|
|
87
|
+
2. **The CLI is present but its plugin subagent is not** → call the CLI
|
|
88
|
+
directly per the raw-fallback column, following the `antigravity` skill's
|
|
89
|
+
invocation pattern and cost discipline.
|
|
90
|
+
3. **None present** → fall back to Claude Code's own `Agent` tool for each
|
|
53
91
|
worker, with an explicit `model: "haiku"` override. This is the only path
|
|
54
92
|
that costs Anthropic tokens for the worker step, and the only one
|
|
55
93
|
guaranteed to exist everywhere — it is the floor, not the default.
|
|
56
94
|
|
|
95
|
+
**Give the delegation a real timeout.** Measured 2026-09: a trivial headless
|
|
96
|
+
`agy` prompt took **605s**. `agy-delegate` defaults to `--print-timeout 5m`,
|
|
97
|
+
so it aborts at 300s and reports an empty body while the answer is still on
|
|
98
|
+
its way — pass `--timeout 15m` for anything non-trivial. Budget worker
|
|
99
|
+
wall-clock accordingly; this is the single most likely reason a fan-out
|
|
100
|
+
"fails" on a machine where the CLI is perfectly healthy.
|
|
101
|
+
|
|
102
|
+
**Verify the delegation actually produced content — a status string is not a
|
|
103
|
+
result.** A timed-out delegation returns `{"status": "SUCCESS", "usage":
|
|
104
|
+
{"total": 0}}` with an *empty* body: success by every field except the one
|
|
105
|
+
that matters. The zero token counts are not proof the prompt never arrived —
|
|
106
|
+
headless usage reporting is simply unpopulated. Check the returned text
|
|
107
|
+
itself, and treat an empty body as a failure no matter what the status says.
|
|
108
|
+
|
|
57
109
|
**This decision happens here, in your own turn, via Bash — never inside a
|
|
58
110
|
`Workflow` script.** A `Workflow` script's body has no shell or filesystem
|
|
59
111
|
access (ambient `agent()`/`pipeline()` globals only), so it cannot itself
|
|
@@ -104,6 +156,8 @@ graph (`.agents/graph.md`) already does properly.
|
|
|
104
156
|
| "The session is already on Opus, so the worker call inherits it fine." | That's exactly the cost this skill exists to avoid — force the tier explicitly every time. |
|
|
105
157
|
| "This task has one obvious step, but a 3-node graph looks more thorough." | More nodes than the task needs is overhead, not rigor. Size the graph to the work. |
|
|
106
158
|
| "I'll just call the Workflow tool and let the script figure out which backend to use." | The script can't — it has no shell access. That decision is yours, before the Workflow call, or not via Workflow at all. |
|
|
159
|
+
| "The CLI is on PATH, so I'll shell out to it directly." | If its plugin subagent is installed, that's the supported path — it already handles the wrapper flags and cost discipline. Shell out only when no subagent exists. |
|
|
160
|
+
| "The wrapper returned SUCCESS, so the work is done." | A failing delegation has returned `SUCCESS` with zero tokens and an empty body. Check the actual content, not the status field. |
|
|
107
161
|
|
|
108
162
|
## 7. Red Flags
|
|
109
163
|
|
|
@@ -115,5 +169,7 @@ graph (`.agents/graph.md`) already does properly.
|
|
|
115
169
|
## 8. Verification
|
|
116
170
|
|
|
117
171
|
- [ ] Detection step actually ran (`command -v` checks), not assumed.
|
|
172
|
+
- [ ] A plugin delegation subagent was preferred where one was available, rather than shelling out to the CLI anyway.
|
|
118
173
|
- [ ] Each node's model was explicitly forced, not inherited.
|
|
174
|
+
- [ ] Each worker's returned **content** was checked, not just its status field — an empty body is a failure regardless of a `SUCCESS` status.
|
|
119
175
|
- [ ] Nothing persistent was left behind after the task completed.
|