@hybridlabor-api/aos 4.1.0 → 4.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/.agents/agents.md +77 -0
  2. package/.agents/graph.md +45 -0
  3. package/.agents/nodes.json +4 -2
  4. package/.agents/state.schema.json +6 -0
  5. package/.claude/workflows/startcycle-dispatch.mjs +139 -8
  6. package/.claude/workflows/teamwork-dispatch.mjs +287 -0
  7. package/CLAUDE.md +47 -0
  8. package/GEMINI.md +9 -1
  9. package/README.md +12 -6
  10. package/THIRD_PARTY_NOTICES.md +50 -0
  11. package/docs/skills_table.md +1 -0
  12. package/installer.js +15 -0
  13. package/package.json +1 -1
  14. package/skills/basic/bdbmediastorm/SKILL.md +7 -5
  15. package/skills/basic/startcycle/SKILL.md +21 -0
  16. package/skills/basic/startcycle-graph/SKILL.md +43 -7
  17. package/skills/basic/startcycle-graph-user/SKILL.md +67 -11
  18. package/skills/basic/teamwork-preview/SKILL.md +209 -0
  19. package/skills/bdbrainstorm/SKILL.md +4 -3
  20. package/skills/global_config/ask-tim/SKILL.md +73 -6
  21. package/skills/global_config/bdbresilience/SKILL.md +216 -0
  22. package/skills/global_config/bdbresilience/contracts/nodes-integration.md +225 -0
  23. package/skills/global_config/bdbresilience/references/cicd-triage.md +179 -0
  24. package/skills/global_config/bdbresilience/references/distributed-locking.md +235 -0
  25. package/skills/global_config/bdbresilience/references/error-recovery.md +210 -0
  26. package/skills/global_config/bdbresilience/references/two-phase-go-gate.md +151 -0
  27. package/skills/global_config/domain-modeling/ADR-FORMAT.md +47 -0
  28. package/skills/global_config/domain-modeling/CONTEXT-FORMAT.md +60 -0
  29. package/skills/global_config/domain-modeling/SKILL.md +77 -0
  30. package/skills/global_config/grill-me/SKILL.md +14 -0
  31. package/skills/global_config/grill-with-docs/SKILL.md +24 -0
  32. package/skills/global_config/grilling/SKILL.md +42 -0
  33. package/skills/global_config/openwiki-skill/scripts/install_daemon.sh +55 -12
package/.agents/agents.md CHANGED
@@ -137,6 +137,83 @@ next. This file defines *what each agent is*, not *what calls what*.
137
137
  - **Output Artifacts**: `production_artifacts/04_release_report.md`
138
138
  - **Reads**: `state.artifacts.*`, `state.findings`, `state.approvals` · **Writes**: `state.gate`, `state.artifacts.report`, `state.phase: ship|done`
139
139
 
140
+ ---
141
+
142
+ # Auxiliary agents
143
+
144
+ The six below are **not** pipeline nodes — they are never in `.agents/nodes.json`,
145
+ never invoked by the dispatcher, and never part of the seven-agent routing above.
146
+ They are standalone specialists you reach for directly. They live here rather than
147
+ only in `.claude/agents/` so the installer compiles them for every harness
148
+ (Antigravity, OpenCode, Codex, Cursor, Roo) instead of leaving them Claude-Code-only.
149
+
150
+ Ported from [affaan-m/ECC](https://github.com/affaan-m/ECC) (MIT) — see
151
+ `THIRD_PARTY_NOTICES.md`.
152
+
153
+ ---
154
+
155
+ ## 🕳️ silent-failure-hunter
156
+ - **Role**: Reviews code for silent failures, swallowed errors, bad fallbacks, and missing error propagation. Finds the bugs that never raise.
157
+ - **Model**: sonnet
158
+ - **Primary Skills**:
159
+ - `systematic-debugging`
160
+ - `debugger`
161
+ - `clean-code`
162
+ - **Output Artifact**: findings returned inline (writes no file)
163
+
164
+ ---
165
+
166
+ ## 🛡️ security-reviewer
167
+ - **Role**: Security vulnerability detection and remediation. Use after writing code that handles user input, authentication, API endpoints, or sensitive data. Flags secrets, SSRF, injection, unsafe crypto, and OWASP Top 10.
168
+ - **Model**: sonnet
169
+ - **Primary Skills**:
170
+ - `systematic-debugging`
171
+ - `clean-code`
172
+ - `api-design-principles`
173
+ - **Output Artifact**: findings returned inline (writes no file)
174
+
175
+ ---
176
+
177
+ ## 🔧 go-build-resolver
178
+ - **Role**: Resolves Go build, vet, and compilation errors with minimal changes. Use when Go builds fail — relevant to `bdb-synapse`, which ships a Go binary.
179
+ - **Model**: sonnet
180
+ - **Primary Skills**:
181
+ - `golang-pro`
182
+ - `go-concurrency-patterns`
183
+ - `systematic-debugging`
184
+ - **Output Artifact**: edits the failing sources directly
185
+
186
+ ---
187
+
188
+ ## 🗄️ database-reviewer
189
+ - **Role**: PostgreSQL specialist for query optimization, schema design, security, and performance. Use when writing SQL, creating migrations, or troubleshooting database performance.
190
+ - **Model**: sonnet
191
+ - **Primary Skills**:
192
+ - `postgres-best-practices`
193
+ - `database-design`
194
+ - `drizzle-orm-expert`
195
+ - **Output Artifact**: findings returned inline (writes no file)
196
+
197
+ ---
198
+
199
+ ## 📦 opensource-forker
200
+ - **Role**: Forks a project for open-sourcing — copies files, strips secrets and credentials, replaces internal references with placeholders, generates `.env.example`, cleans git history. Run before `opensource-sanitizer`.
201
+ - **Model**: haiku
202
+ - **Primary Skills**:
203
+ - `github-repo`
204
+ - `bash-linux`
205
+ - **Output Artifact**: `FORK_REPORT.md`
206
+
207
+ ---
208
+
209
+ ## 🧼 opensource-sanitizer
210
+ - **Role**: Verifies an open-source fork is fully sanitized before release. Scans for leaked secrets, PII, internal references, and dangerous files; emits PASS/FAIL/PASS-WITH-WARNINGS. Run after `opensource-forker`, before any public release.
211
+ - **Model**: sonnet
212
+ - **Primary Skills**:
213
+ - `github-repo`
214
+ - `bash-linux`
215
+ - **Output Artifact**: `SANITIZATION_REPORT.md`
216
+
140
217
  ---
141
218
  ## 🔄 Context Boot Sequence
142
219
  Before executing any tasks, every agent MUST perform the following checks silently:
package/.agents/graph.md CHANGED
@@ -33,6 +33,51 @@ before returning — this is what replaces "hand-off," and it's why a node
33
33
  never needs another node's reasoning: `goal` and prior artifacts are always
34
34
  read from the same typed record, not re-derived from a sibling's prose.
35
35
 
36
+ ## Mandatory Skill Injection
37
+
38
+ `/startcycle-graph --skill=<name> <goal>` (repeatable: `--skill=a --skill=b
39
+ <goal>`, quote a name containing spaces) forces a specific skill into this
40
+ run — for the case where you have your own private skill (never part of
41
+ `.agents/nodes.json`'s registry, and never touched by AOS's installer per
42
+ its foreign-file conflict policy) that you need applied regardless of what
43
+ the registry's own per-node allowlist would have reached for.
44
+
45
+ - The dispatcher script (`startcycle-dispatch.mjs`) extracts every
46
+ `--skill=` flag from the invocation text before anything else runs, then
47
+ validates each name resolves to a real installed skill (a `SKILL.md`
48
+ under any harness's global skills directory — `~/.claude/skills/<name>/`,
49
+ `~/.agents/skills/`, `~/.codex/skills/`, `~/.cursor/skills/`, `~/.roo/skills/`,
50
+ all of which the installer writes — or this project's own `skills/` tree) via
51
+ a read-only lookup agent. **A name that doesn't resolve escalates
52
+ immediately** — same "never silently fall back or guess" posture as a
53
+ missing registry node id. This is a fail-fast check specifically so a
54
+ typo doesn't silently ship a run that never used the skill you asked for.
55
+ A flag written with an empty value (`--skill=` with nothing after it)
56
+ escalates for the same reason: it would otherwise inject nothing *and*
57
+ leave the literal `--skill=` glued to the goal text Architect reads.
58
+ - The validated list is persisted to `state.mandatory_skills` (set by
59
+ Architect on the first write) and passed to every build node's prompt —
60
+ and Architect's own — as a **hard requirement, not a suggestion**,
61
+ layered on top of (never replacing) the registry's own per-node skill
62
+ allowlist.
63
+ - **TechLead rejects a plan that ignores the mandate**, at the plan-approval
64
+ gate — one extra planning round instead of a wasted build cycle. Without
65
+ this the mandate is only caught downstream by Reviewer, i.e. after the
66
+ build nodes have already run against a plan that never accounted for it.
67
+ - **Reviewer checks it was actually used, not just available.** An artifact
68
+ that shows no sign of applying a mandated skill's guidance is a
69
+ `contract_misread` finding (blocking), owned by whichever build node
70
+ should have applied it — the same precedence class as misreading the
71
+ plan itself, since an ignored `--skill` flag is exactly that.
72
+ - Nodes that do **not** receive the mandate, deliberately: `shipping` (runs
73
+ mechanical gates — lint/typecheck/tests — and produces no artifact a skill
74
+ would shape).
75
+ - `/startcycle` (the linear variant, no `state.json`) and
76
+ `/startcycle-graph-user` (throwaway, nothing persistent) support the same
77
+ `--skill=<name>` syntax — see each skill's own `SKILL.md` for how the
78
+ orchestrator threads it through without a durable state file to carry it
79
+ in.
80
+
36
81
  ## Nodes
37
82
 
38
83
  Seven, up from the original five — `Planner_Orchestrator` is split into
@@ -72,7 +72,8 @@
72
72
  "drizzle-orm-expert",
73
73
  "postgres-best-practices",
74
74
  "typescript-pro",
75
- "python-pro"
75
+ "python-pro",
76
+ "bdbresilience"
76
77
  ],
77
78
  "instructions": "Implement the backend per the plan: DDD models, type-safe schemas, API routes, Clean Architecture. Write production_artifacts/02_backend_schema.md and the code."
78
79
  },
@@ -128,7 +129,8 @@
128
129
  "seo-audit",
129
130
  "wcag-audit-patterns",
130
131
  "github-repo",
131
- "clean-code"
132
+ "clean-code",
133
+ "bdbresilience"
132
134
  ],
133
135
  "instructions": null
134
136
  }
@@ -89,6 +89,12 @@
89
89
  "additionalProperties": false
90
90
  }
91
91
  },
92
+ "mandatory_skills": {
93
+ "type": "array",
94
+ "items": { "type": "string" },
95
+ "default": [],
96
+ "description": "Skill names the user required via /startcycle-graph's --skill=<name> flag (repeatable), validated to exist before the run proceeds. Empty when the user didn't ask for one. Build nodes (and Architect) are told to actually apply these, not just have them available; Reviewer checks the resulting artifacts for evidence they were used and flags a contract-misread finding if not. See .agents/graph.md's 'Mandatory Skill Injection' section."
97
+ },
92
98
  "needs_human": {
93
99
  "type": "boolean",
94
100
  "default": false,
@@ -81,6 +81,14 @@ const MAX_ITERATIONS = 3;
81
81
  // correct on the first draft.
82
82
  let iteration = 0;
83
83
 
84
+ // Populated from a --skill=<name> flag (repeatable) in the invocation text --
85
+ // see extractMandatorySkills() and .agents/graph.md's "Mandatory Skill
86
+ // Injection" section. Empty when the user didn't ask for one. Module-level
87
+ // like `iteration` above, for the same reason: skillsNote() and the Reviewer
88
+ // prompt both need it and neither is in a position to thread it through as a
89
+ // parameter without touching every call site.
90
+ let mandatorySkills = [];
91
+
84
92
  // Every sequential (non-build-role) agent gets the same note: read the full
85
93
  // state, and explicitly set state.iteration to the dispatcher's current
86
94
  // count so the persisted file (which .claude/hooks/graph-gate.mjs reads)
@@ -149,11 +157,27 @@ function reviewerStateNote() {
149
157
  // Applies to every node, build and sequential alike.
150
158
  function skillsNote(node) {
151
159
  const skills = Array.isArray(node?.skills) ? node.skills : [];
152
- if (skills.length === 0) return '';
153
- return (
154
- ` Use these skills for this work: ${skills.join(', ')}. ` +
155
- 'Do not reach for skills outside this list unless the task genuinely requires it.'
156
- );
160
+ const parts = [];
161
+ if (skills.length > 0) {
162
+ parts.push(
163
+ ` Use these skills for this work: ${skills.join(', ')}. ` +
164
+ 'Do not reach for skills outside this list unless the task genuinely requires it.'
165
+ );
166
+ }
167
+ // A --skill flag is a hard requirement from the user, not the registry's
168
+ // own suggested allowlist above -- it applies on top of, never instead of,
169
+ // that list. Only nodes that actually produce work get told to use it:
170
+ // build-role nodes, plus Architect (who should fold the skill's guidance
171
+ // into the plan itself, not just leave it for Build to discover cold).
172
+ // Reviewer gets a separate mention in its own prompt below, framed as a
173
+ // check rather than a use.
174
+ if (mandatorySkills.length > 0 && (node?.role === 'build' || node?.id === 'architect')) {
175
+ parts.push(
176
+ ` The user explicitly required this run to use the following skill(s), via /startcycle-graph's --skill flag: ${mandatorySkills.join(', ')}. ` +
177
+ "This is a hard requirement, not a suggestion -- actually apply the skill's guidance in your work, and name in your returned summary how each one was applied."
178
+ );
179
+ }
180
+ return parts.join('');
157
181
  }
158
182
 
159
183
  // "a", "a or b", "a, b, or c" -- used for NODE_NAMES, itself derived from the
@@ -297,13 +321,110 @@ const techleadNode = { id: 'techlead', ...registryNodes.techlead };
297
321
  const reviewerNode = { id: 'reviewer', ...registryNodes.reviewer };
298
322
  const shippingNode = { id: 'shipping', ...registryNodes.shipping };
299
323
 
300
- const goal = typeof args === 'string' ? args : args?.goal;
301
- if (!goal) {
324
+ const rawGoal = typeof args === 'string' ? args : args?.goal;
325
+ if (!rawGoal) {
302
326
  return escalate(
303
327
  'startcycle-graph needs a goal, e.g. "Run /startcycle-graph on: add OAuth login with Google" -- nothing was invoked.'
304
328
  );
305
329
  }
306
330
 
331
+ // --skill=<name>, repeatable, extracted out of the raw goal text before
332
+ // anything else sees it -- e.g. "--skill=my-custom-skill Add OAuth login"
333
+ // becomes goal "Add OAuth login" plus one mandated skill name. This is the
334
+ // mechanism for injecting a skill this script has never heard of (a user's
335
+ // own private skill, never part of .agents/nodes.json's registry) -- see
336
+ // .agents/graph.md's "Mandatory Skill Injection" section.
337
+ function extractMandatorySkills(text) {
338
+ // The `|--skill=(?=\s|$)` alternative deliberately matches a flag with an
339
+ // EMPTY value ("--skill= add OAuth"). Without it, `\S+` simply fails to
340
+ // match, the flag falls through as ordinary prose, and the run proceeds
341
+ // with no skill injected AND the literal "--skill=" still glued to the
342
+ // goal text handed to Architect -- a silent no-op on a typo, which is the
343
+ // exact failure mode the validation below exists to prevent. Capturing it
344
+ // as an empty name instead routes it into `malformed` and escalates.
345
+ const flagPattern = /--skill=("[^"]+"|'[^']+'|\S+)|--skill=(?=\s|$)/g;
346
+ const skills = [];
347
+ let malformed = 0;
348
+ const goal = text
349
+ .replace(flagPattern, (_, val) => {
350
+ if (val === undefined) { malformed++; return ''; }
351
+ const unquoted =
352
+ (val.startsWith('"') && val.endsWith('"')) || (val.startsWith("'") && val.endsWith("'"))
353
+ ? val.slice(1, -1)
354
+ : val;
355
+ if (unquoted.trim() === '') { malformed++; return ''; }
356
+ skills.push(unquoted);
357
+ return '';
358
+ })
359
+ .replace(/\s{2,}/g, ' ')
360
+ .trim();
361
+ return { skills, goal, malformed };
362
+ }
363
+
364
+ const { skills: skillsFromFlags, goal, malformed } = extractMandatorySkills(rawGoal);
365
+ if (malformed > 0) {
366
+ return await escalate(
367
+ `--skill was given with an empty value (${malformed} time(s)). Write --skill=<name>, e.g. --skill=my-custom-skill. ` +
368
+ 'Refusing to proceed rather than silently running without the skill you asked for.'
369
+ );
370
+ }
371
+ if (!goal) {
372
+ return await escalate(
373
+ 'startcycle-graph needs actual goal text, not just --skill flag(s) -- e.g. "--skill=my-custom-skill add OAuth login with Google", not "--skill=my-custom-skill" alone.'
374
+ );
375
+ }
376
+ // Object-form args may also carry a structured list directly, for a future
377
+ // caller that never goes through the string-flag convention at all.
378
+ const skillsFromArgs = Array.isArray(args?.mandatorySkills) ? args.mandatorySkills : [];
379
+ const mandatorySkillNames = [...new Set([...skillsFromFlags, ...skillsFromArgs])];
380
+
381
+ // Validate before anything else runs -- same "never silently fall back or
382
+ // guess" posture as the registry load above. This script has no filesystem
383
+ // access of its own (comment block item #1), so validation is itself an
384
+ // agent() call, not a local fs check.
385
+ if (mandatorySkillNames.length > 0) {
386
+ const skillCheckResult = await agent(
387
+ `Check whether each of these skill names resolves to an installed skill with a real SKILL.md: ${JSON.stringify(mandatorySkillNames)}. ` +
388
+ 'The installer syncs the same skill set to every harness it detects, so check all of these global locations, not just the first: ' +
389
+ '~/.claude/skills/<name>/SKILL.md, ~/.agents/skills/<name>/SKILL.md, ~/.codex/skills/<name>/SKILL.md, ' +
390
+ '~/.cursor/skills/<name>/SKILL.md, ~/.roo/skills/<name>/SKILL.md. A skill present in any one of them counts as installed — ' +
391
+ 'this workflow may be driven from a harness whose directory is not ~/.claude. ' +
392
+ 'If this project has its own skills/ directory, also accept skills/<name>/SKILL.md or skills/<container>/<name>/SKILL.md. ' +
393
+ 'This is a read-only lookup, not a reasoning task -- do not invent a path that does not exist, and never report a close match as `found`.\n\n' +
394
+ 'For any name that does NOT resolve, list up to five installed skills whose directory names are plausible near-misses ' +
395
+ '(substring, obvious typo, or the same words in another order) in `suggestions`. Read the real directory listing to do this -- ' +
396
+ 'suggest only names that actually exist on disk. `--skill=` requires an exact directory name, and a user who mistyped one ' +
397
+ 'has no way to discover the right spelling from an error that only says "not found".\n\n' +
398
+ 'Return only: { "found": string[], "missing": string[], "suggestions": string[] }.',
399
+ {
400
+ label: 'validate-mandatory-skills',
401
+ model: 'haiku',
402
+ schema: {
403
+ type: 'object',
404
+ required: ['found', 'missing'],
405
+ properties: {
406
+ found: { type: 'array', items: { type: 'string' } },
407
+ missing: { type: 'array', items: { type: 'string' } },
408
+ suggestions: { type: 'array', items: { type: 'string' } },
409
+ },
410
+ },
411
+ }
412
+ );
413
+ const missing = skillCheckResult?.missing ?? [];
414
+ if (missing.length > 0) {
415
+ const near = skillCheckResult?.suggestions ?? [];
416
+ return await escalate(
417
+ `--skill named skill(s) that could not be found on this machine: ${missing.join(', ')}. ` +
418
+ (near.length
419
+ ? `Did you mean: ${near.join(', ')}? `
420
+ : 'No installed skill has a similar name. ') +
421
+ '--skill= takes the exact skill directory name; run /ask-tim to find the one you want. ' +
422
+ 'Refusing to silently proceed without a mandated skill.'
423
+ );
424
+ }
425
+ mandatorySkills = skillCheckResult?.found ?? mandatorySkillNames;
426
+ }
427
+
307
428
  // ---------------------------------------------------------------------
308
429
  // Architect <-> TechLead: plan, then capability-map approval. Sequential --
309
430
  // not part of the CHANGE 1 race, writes state.json directly as before.
@@ -323,7 +444,7 @@ while (!approved) {
323
444
  : '') +
324
445
  `Turn this goal into a system plan with an explicit capability map (module boundaries, ` +
325
446
  `dependency direction, build order). Write it to production_artifacts/00_execution_plan.md. ` +
326
- `Set state.goal, state.phase = "plan", state.artifacts.plan to that path. ` +
447
+ `Set state.goal, state.phase = "plan", state.artifacts.plan to that path, and state.mandatory_skills to ${JSON.stringify(mandatorySkills)}. ` +
327
448
  `Decide whether the goal needs the Media_EventTech build node (TouchDesigner/show-control/3D/media work) -- most goals don't.\n\n` +
328
449
  `Return only: { "planPath": string, "needsMedia": boolean }.`,
329
450
  {
@@ -349,6 +470,11 @@ while (!approved) {
349
470
  `You are acting as the ${techleadNode.label} agent (${techleadNode.personaFile}). ${dispatchNote(techleadNode)}${skillsNote(techleadNode)}\n\n` +
350
471
  `Read the plan at ${planPath}. Approve it only if it has an explicit capability map: ` +
351
472
  `module boundaries, dependency direction, and build order are all stated, not implicit. ` +
473
+ (mandatorySkills.length > 0
474
+ ? `The user also required this run to use the following skill(s) via /startcycle-graph's --skill flag: ${mandatorySkills.join(', ')}. ` +
475
+ `Reject the plan if it does not actually account for them — catching that here costs one planning round, ` +
476
+ `whereas letting it through wastes a full build cycle before Reviewer flags it.\n`
477
+ : '') +
352
478
  `Record your decision in state.json (plan approval, state.phase = "build" if approved).\n\n` +
353
479
  `Return only: { "approved": boolean, "reason": string }.`,
354
480
  {
@@ -459,6 +585,11 @@ while (!reviewedClean) {
459
585
  }, and the actual code). ` +
460
586
  `Do an adversarial review against the contract: find what is wrong, do not validate, do not summarize. ` +
461
587
  `Do not assume the implementation is correct just because it exists. ` +
588
+ (mandatorySkills.length > 0
589
+ ? `The user explicitly required these skill(s) to be used this run, via /startcycle-graph's --skill flag: ${mandatorySkills.join(', ')}. ` +
590
+ `If an artifact shows no sign of applying a mandated skill's guidance, that is a contract misread finding (blocking), owned by whichever build node should have applied it. ` +
591
+ `A mandated skill being merely available is not enough -- check for it actually being used.\n`
592
+ : '') +
462
593
  `Classify every finding by precedence: contract misread > valid & actionable (blocking) > valid trade-off (advisory) > noise (discard). ` +
463
594
  `Each finding must name which node owns fixing it: ${NODE_NAMES} -- no other value is valid. ` +
464
595
  `If you are re-reviewing after a repair round and an issue you flagged before is still present and still unfixed, ` +
@@ -0,0 +1,287 @@
1
+ // Dispatcher for /teamwork-preview. Turns the 9-step prompt-crafting protocol
2
+ // from skills/basic/teamwork-preview/SKILL.md into an actual runnable sequence.
3
+ //
4
+ // Why a script and not prose: this repo has already paid for the alternative
5
+ // once, recorded verbatim in skills/basic/startcycle-graph/SKILL.md --
6
+ //
7
+ // "an earlier version of this file embedded the full pipeline description in
8
+ // prose, and the model followed it 'in spirit' inline instead of invoking the
9
+ // script -- silently skipping the whole graph, with no state.json, no
10
+ // subagents, no Reviewer, and no quality gate ever running."
11
+ //
12
+ // A 9-step protocol with acceptance criteria and integrity modes is exactly the
13
+ // kind of thing that gets followed approximately. A script either runs or it
14
+ // does not.
15
+ //
16
+ // Runtime constraints inherited from startcycle-dispatch.mjs, all four of which
17
+ // this script obeys:
18
+ // 1. No filesystem access from the script itself. Every read and write happens
19
+ // inside an agent() call; the script branches on schema-validated returns.
20
+ // 2. No module loading. `agent`, `args` are ambient globals injected by the
21
+ // runtime, not imports.
22
+ // 3. Concurrent agents writing one file race. This script is deliberately
23
+ // sequential -- each step's answers reshape the next question, so there is
24
+ // nothing to parallelise and no fragment/merge dance is needed.
25
+ // 4. Prompts are the interface. An agent that returns prose instead of the
26
+ // declared schema breaks the branch, so every step declares one.
27
+ //
28
+ // This is NOT Antigravity's /teamwork-preview. That command is compiled into the
29
+ // agy binary (its own conductor/orchestrator/auditor agent types, maintained by
30
+ // Google). This is an independent implementation of the same idea, on a
31
+ // different runtime, and it will behave differently.
32
+
33
+ export const meta = {
34
+ name: 'teamwork-dispatch',
35
+ description:
36
+ 'Interactive 9-step prompt crafting for multi-agent delegation: elicit, disambiguate, set integrity mode, draft requirements, design verification, set acceptance criteria, then assemble and validate a spec. Produces prompt_draft.md; does not build anything.',
37
+ };
38
+
39
+ // The user's answers accumulate here. Each step gets the answers so far, so a
40
+ // later question can be shaped by an earlier one -- which is the whole point of
41
+ // an interview and the reason these run sequentially rather than in parallel.
42
+ const spec = {
43
+ idea: null,
44
+ scale: null,
45
+ integrityMode: null,
46
+ requirements: [],
47
+ verification: null,
48
+ acceptanceCriteria: [],
49
+ infrastructure: null,
50
+ workingDirectory: null,
51
+ };
52
+
53
+ // One shared preamble. Every step is an interview turn, not a build turn: the
54
+ // agent asks, the human answers, nothing gets implemented. Stated once here
55
+ // rather than restated nine times, where the ninth copy would drift.
56
+ const INTERVIEW_RULE =
57
+ 'You are conducting one step of an interactive interview. Ask the user, wait for their answer, and record it. ' +
58
+ 'Do NOT implement anything, do NOT write project code, and do NOT proceed past your own step. ' +
59
+ 'Finding facts is your job, not the user\'s: if a question can be answered by reading the filesystem or running a command, ' +
60
+ 'do that yourself instead of asking. The decisions are the user\'s: put each to them and wait. ' +
61
+ 'If the user has already answered something in an earlier step, do not ask it again -- the answers so far are given below.';
62
+
63
+ function answersSoFar() {
64
+ const known = Object.entries(spec).filter(([, v]) =>
65
+ Array.isArray(v) ? v.length > 0 : v !== null
66
+ );
67
+ if (known.length === 0) return 'Nothing settled yet -- this is the first step.';
68
+ return `Answers settled so far:\n${JSON.stringify(Object.fromEntries(known), null, 2)}`;
69
+ }
70
+
71
+ async function step(n, title, instruction, schema, label) {
72
+ return await agent(
73
+ `${INTERVIEW_RULE}\n\n` +
74
+ `## Step ${n} of 9: ${title}\n\n${instruction}\n\n` +
75
+ `${answersSoFar()}\n\n` +
76
+ 'Return only the declared JSON.',
77
+ { label: label || `step-${n}`, schema }
78
+ );
79
+ }
80
+
81
+ const goal = typeof args === 'string' ? args : args?.goal;
82
+
83
+ // ---------------------------------------------------------------------
84
+ // Steps 1-3: what is this, how big, and what is it allowed to use
85
+ // ---------------------------------------------------------------------
86
+
87
+ const s1 = await step(
88
+ 1,
89
+ 'Elicit the idea',
90
+ 'Ask what the user wants to build, what its purpose is (production, demo, eval, or prototype), and who the audience is. ' +
91
+ 'Condense their answer into a description of one or two sentences -- not a paragraph, and not a restatement of the question.' +
92
+ (goal ? `\n\nThe user already said: ${JSON.stringify(goal)}. Start from that; ask only what it leaves open.` : ''),
93
+ {
94
+ type: 'object',
95
+ required: ['description', 'purpose'],
96
+ properties: {
97
+ description: { type: 'string' },
98
+ purpose: { type: 'string', enum: ['production', 'demo', 'eval', 'prototype'] },
99
+ audience: { type: 'string' },
100
+ },
101
+ }
102
+ );
103
+ spec.idea = s1;
104
+
105
+ const s2 = await step(
106
+ 2,
107
+ 'Identify ambiguity and scale',
108
+ 'Probe every point that has more than one reasonable interpretation -- data sources, third-party services, where the scope stops. ' +
109
+ 'Then establish the shape of the effort:\n' +
110
+ '- a single self-contained fix or feature (one implementer plus repeated adversarial review)\n' +
111
+ '- math, formal proofs, or a massive search space (may warrant a large agent team)\n' +
112
+ '- a standard multi-agent build\n\n' +
113
+ 'Ambiguity you leave unresolved here becomes a wrong assumption baked into the spec, so be thorough now rather than agreeable.',
114
+ {
115
+ type: 'object',
116
+ required: ['scale'],
117
+ properties: {
118
+ scale: { type: 'string', enum: ['single-focused', 'large-scale', 'standard'] },
119
+ ambiguitiesResolved: { type: 'array', items: { type: 'string' } },
120
+ },
121
+ }
122
+ );
123
+ spec.scale = s2;
124
+
125
+ const s3 = await step(
126
+ 3,
127
+ 'Determine integrity mode',
128
+ 'Clarify the operational boundaries: may code be copied from existing open-source projects? Are pre-built libraries allowed for the core logic ' +
129
+ '(as opposed to the scaffolding)? May the implementer inspect the tests before writing the code?\n\n' +
130
+ 'Map the answers: unrestricted → `development`; some shortcuts acceptable because it is a showcase → `demo`; ' +
131
+ 'strict isolation, zero external leakage → `benchmark`.\n\n' +
132
+ 'The last question matters more than it looks: an implementer who can read the tests first can satisfy them without solving the problem.',
133
+ {
134
+ type: 'object',
135
+ required: ['integrityMode'],
136
+ properties: {
137
+ integrityMode: { type: 'string', enum: ['development', 'demo', 'benchmark'] },
138
+ rationale: { type: 'string' },
139
+ },
140
+ }
141
+ );
142
+ spec.integrityMode = s3.integrityMode;
143
+
144
+ // ---------------------------------------------------------------------
145
+ // Steps 4-6: what must be true, and how anyone would know
146
+ // ---------------------------------------------------------------------
147
+
148
+ const s4 = await step(
149
+ 4,
150
+ 'Draft requirements',
151
+ 'Write two to five requirement blocks (R1, R2, ...). Each states **what** is required, never **how** to implement it.\n\n' +
152
+ 'Apply the litmus test to every one: would a senior engineer feel over-constrained by this? If yes, prune it. ' +
153
+ 'A requirement that dictates implementation removes the judgement you are hiring the implementer for.',
154
+ {
155
+ type: 'object',
156
+ required: ['requirements'],
157
+ properties: {
158
+ requirements: {
159
+ type: 'array',
160
+ minItems: 2,
161
+ maxItems: 5,
162
+ items: {
163
+ type: 'object',
164
+ required: ['id', 'text'],
165
+ properties: { id: { type: 'string' }, text: { type: 'string' } },
166
+ },
167
+ },
168
+ },
169
+ }
170
+ );
171
+ spec.requirements = s4.requirements;
172
+
173
+ const s5 = await step(
174
+ 5,
175
+ 'Design the verification mechanism',
176
+ 'This is the forcing function, and it is the step that decides whether the whole exercise works.\n\n' +
177
+ 'Its job is to create an objective target that forces a real build → test → debug loop and makes premature self-certification impossible. ' +
178
+ 'An agent that can declare its own work done, will.\n\n' +
179
+ 'Prefer something programmatic: a unit test suite, a test runner invocation, a CLI script that asserts. ' +
180
+ 'Only if that is genuinely infeasible, draft an explicit agent-as-judge rubric -- and say why programmatic was not possible. ' +
181
+ 'Ask whether the user has existing test suites, schemas, or a reference implementation to hand the implementer.',
182
+ {
183
+ type: 'object',
184
+ required: ['mechanism', 'isProgrammatic'],
185
+ properties: {
186
+ mechanism: { type: 'string' },
187
+ isProgrammatic: { type: 'boolean' },
188
+ resources: { type: 'array', items: { type: 'string' } },
189
+ },
190
+ }
191
+ );
192
+ spec.verification = s5;
193
+
194
+ const s6 = await step(
195
+ 6,
196
+ 'Set acceptance criteria',
197
+ 'Convert the verification mechanism into checkable criteria -- each one a thing that is either true or false, never a judgement call.\n\n' +
198
+ `Calibrate to the stated purpose (${spec.idea.purpose}): a demo must be achievable in a rapid time budget; ` +
199
+ 'production needs real coverage, error handling and readiness; an eval needs reproducible metrics far more than polish.\n\n' +
200
+ 'A criterion nobody can mechanically check is a wish, not a criterion.',
201
+ {
202
+ type: 'object',
203
+ required: ['criteria'],
204
+ properties: { criteria: { type: 'array', minItems: 1, items: { type: 'string' } } },
205
+ }
206
+ );
207
+ spec.acceptanceCriteria = s6.criteria;
208
+
209
+ // ---------------------------------------------------------------------
210
+ // Steps 7-8: where it runs
211
+ // ---------------------------------------------------------------------
212
+
213
+ const s7 = await step(
214
+ 7,
215
+ 'Infrastructure constraints',
216
+ 'Only if the work reaches outside the local workspace: define the sandboxing or controlled APIs for remote file operations, ' +
217
+ 'job launching, and outbound network calls.\n\n' +
218
+ 'If the project stays entirely within local workspace files, say so and skip -- do not invent constraints to fill this step.',
219
+ {
220
+ type: 'object',
221
+ required: ['applicable'],
222
+ properties: { applicable: { type: 'boolean' }, constraints: { type: 'string' } },
223
+ }
224
+ );
225
+ spec.infrastructure = s7.applicable ? s7.constraints : null;
226
+
227
+ const s8 = await step(
228
+ 8,
229
+ 'Choose the working directory',
230
+ 'Confirm where this runs. Default to a path inside the current repository if the work belongs to it, ' +
231
+ `otherwise \`~/teamwork_projects/<project_name>\`. Check that the path exists or can be created, and say which.`,
232
+ {
233
+ type: 'object',
234
+ required: ['workingDirectory'],
235
+ properties: { workingDirectory: { type: 'string' }, exists: { type: 'boolean' } },
236
+ }
237
+ );
238
+ spec.workingDirectory = s8.workingDirectory;
239
+
240
+ // ---------------------------------------------------------------------
241
+ // Step 9: assemble, validate, and stop for approval
242
+ // ---------------------------------------------------------------------
243
+
244
+ const s9 = await agent(
245
+ 'You are assembling the final specification from a completed 9-step interview. Do NOT implement any of it.\n\n' +
246
+ `Write \`prompt_draft.md\` into ${JSON.stringify(spec.workingDirectory)} with this structure:\n\n` +
247
+ '- the one-to-two sentence project description\n' +
248
+ '- `Working directory: <path>`\n' +
249
+ '- `Integrity mode: <mode>`\n' +
250
+ '- a team-scaling directive, if the scale calls for one\n' +
251
+ '- `## Requirements` — the R1..Rn blocks\n' +
252
+ '- `## Verification` — the mechanism, and the resources it may use\n' +
253
+ '- `## Acceptance Criteria` — as markdown checkboxes (`- [ ]`)\n' +
254
+ (spec.infrastructure ? '- `## Infrastructure Constraints`\n' : '') +
255
+ '\nThen validate it against three checks and report each honestly:\n' +
256
+ '1. Does every acceptance criterion trace back to a stated requirement? An orphan criterion means the interview missed a requirement.\n' +
257
+ '2. Is every criterion mechanically checkable — true or false, no judgement call?\n' +
258
+ '3. Does any requirement dictate *how* rather than *what*?\n\n' +
259
+ 'Report failures rather than quietly fixing them: a validation step that edits its own input proves nothing.\n\n' +
260
+ `The interview produced:\n${JSON.stringify(spec, null, 2)}`,
261
+ {
262
+ label: 'step-9-assemble',
263
+ schema: {
264
+ type: 'object',
265
+ required: ['draftPath', 'validationsPassed'],
266
+ properties: {
267
+ draftPath: { type: 'string' },
268
+ validationsPassed: { type: 'boolean' },
269
+ issues: { type: 'array', items: { type: 'string' } },
270
+ },
271
+ },
272
+ }
273
+ );
274
+
275
+ // Deliberately stops here. Delegation is a separate, human-approved act: the
276
+ // spec is the deliverable of this workflow, not a launch command. Handing an
277
+ // unapproved spec straight to a swarm is precisely the premature
278
+ // self-certification step 5 exists to prevent.
279
+ return {
280
+ phase: s9.validationsPassed ? 'ready_for_approval' : 'validation_failed',
281
+ draftPath: s9.draftPath,
282
+ issues: s9.issues ?? [],
283
+ integrityMode: spec.integrityMode,
284
+ reason: s9.validationsPassed
285
+ ? `Spec assembled at ${s9.draftPath}. Review it, then delegate to the execution harness of your choice — this workflow deliberately does not launch anything.`
286
+ : `Spec assembled at ${s9.draftPath} but validation found issues; resolve them before delegating.`,
287
+ };