superpowers-mcp 6.2.4 → 6.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ja.md +16 -3
- package/README.ko.md +16 -3
- package/README.md +16 -3
- package/README.zh-TW.md +16 -3
- package/out/server.js +1 -1
- package/package.json +1 -1
- package/skills/brainstorming/SKILL.md +108 -9
- package/skills/finishing-a-development-branch/SKILL.md +30 -0
- package/skills/requesting-code-review/code-reviewer.md +9 -0
- package/skills/subagent-driven-development/SKILL.md +101 -29
- package/skills/subagent-driven-development/implementer-prompt.md +12 -0
- package/skills/subagent-driven-development/re-review-prompt.md +9 -0
- package/skills/subagent-driven-development/scripts/sdd-workspace.ps1 +1 -1
- package/skills/subagent-driven-development/task-reviewer-prompt.md +25 -5
- package/skills/using-superpowers/SKILL.md +1 -0
- package/skills/using-superpowers/references/codex-tools.md +70 -1
- package/skills/using-superpowers/references/hermes-tools.md +56 -0
- package/skills/writing-plans/SKILL.md +3 -0
- package/skills/writing-skills/anthropic-best-practices.md +1 -1
- package/skills/writing-skills/render-graphs.js +3 -2
package/README.ja.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[English](README.md) | [繁體中文](README.zh-TW.md) | [日本語](README.ja.md) | [한국어](README.ko.md)
|
|
4
4
|
|
|
5
|
-
[](https://github.com/Poseidoncode/superpowers-mcp)
|
|
6
6
|
[](LICENSE)
|
|
7
7
|
|
|
8
8
|
このドキュメントは、オリジナルの Superpowers スキルライブラリを独立した MCP Toolpack にパッケージ化するための情報と使用手順をまとめたものです。
|
|
@@ -112,7 +112,7 @@
|
|
|
112
112
|
これらのスキルは、サポートされている IDE(Antigravity や Cursor など)内で複雑なメタ実行パターンをオーケストレーションするために設計されています。
|
|
113
113
|
|
|
114
114
|
- **`subagent-driven-development`**: サブエージェントを駆動してタスクを実行
|
|
115
|
-
- **使用方法**: 定義済みの計画をタスクごとに実行します。システムはタスクごとに新しい「実装」サブエージェントを生成し、その後、統合された **タスクレビューアー**(仕様準拠 + コード品質)サブエージェントと、最後に **全ブランチ最終レビュー** を実行します。**Pre-Flight Plan Review**
|
|
115
|
+
- **使用方法**: 定義済みの計画をタスクごとに実行します。システムはタスクごとに新しい「実装」サブエージェントを生成し、その後、統合された **タスクレビューアー**(仕様準拠 + コード品質)サブエージェントと、最後に **全ブランチ最終レビュー** を実行します。**Pre-Flight Plan Review** は、実行開始前のタスク競合をスキャンします。計画は plan-scoped ワークスペース(`.superpowers/sdd/<plan>/`)で実行され、コントローラーは停止せずに衝突を裁定してレジャーに記録(rulings)し、同じ形の小タスクは 1 回のディスパッチにまとめられます。
|
|
116
116
|
- **モデル選択**: タスクの複雑さに基づいてサブエージェントモデルを選択 — 機械的な作業には低コストモデル、アーキテクチャや微妙な並行性変更には高性能モデル。
|
|
117
117
|
- **例**: 「subagent-driven-development スキルを読み込んで、docs/plans/feature-plan.md にリストされているタスクを 1 つずつ実行して」
|
|
118
118
|
- **`dispatching-parallel-agents`**: タスクを並列エージェントに派遣
|
|
@@ -128,7 +128,20 @@
|
|
|
128
128
|
|
|
129
129
|
## 🆕 最近の更新
|
|
130
130
|
|
|
131
|
-
### v6.
|
|
131
|
+
### v6.3.0(最新)
|
|
132
|
+
|
|
133
|
+
- **上流 obra/superpowers v6.3.0 との同期** — 適用可能な改善をすべて採用し、フォーク固有のセキュリティ強化と PowerShell サポートは維持。
|
|
134
|
+
- **brainstorming — 3 パスルーター**: すべてのリクエストを事前に `spike` / `bounded` / `architectural` に分類し、プロセス量をタスクに合わせて調整。ただし承認ゲートは常に全パスに適用されます。実行中に隠れた複雑さが判明した場合はパスをアップグレード — ダウングレードはありません。
|
|
135
|
+
- **subagent-driven-development — 裁定、停止しない(rulings, not stalls)**: 衝突・曖昧さ・計画の欠陥はコントローラーが直接裁定しレジャーに記録(`Ruling: ...`)。停止するのは明示された 4 条件のみ。Pre-flight 競合スキャンはレジャー表を出力し、同形状の小タスクは単一ディスパッチにバッチ化され、子エージェント待機は境界付きストレッチを使用。3 つのプロンプトすべてに no-subagents 契約を追加。
|
|
136
|
+
- **Hermes Agent サポート**: 新しい `hermes-tools.md` リファレンスがスキルアクションを Hermes ツール(`delegate_task`、`skill_view` など)にマッピング。
|
|
137
|
+
- **Codex**: V1/V2 マルチエージェントの違い、`followup_task` による修正ラウンド再開、イベント購読型 `wait_agent` のガイダンス。
|
|
138
|
+
- **writing-plans**: プランテンプレートに `Spec:` フィールドを追加。
|
|
139
|
+
- **finishing-a-development-branch**: worktree 削除拒否時の手順 — 自分の判断で `--force` しない。
|
|
140
|
+
- **デュアルエージェント code review による修正**: merged パスで「Commit them to \<branch\>」を選んでもファイルがベースブランチの外に取り残されない(finishing-a-development-branch)。`sdd-workspace.ps1` のスラッグ導出は全プラットフォームで `basename` と一致(`PLAN.MD` は `PLAN.MD` のまま)。
|
|
141
|
+
- **意図的に未採用**: 上流 v6.3.0 のサーバー簡素化(loopback-only バインド、`O_NOFOLLOW` 読み取り、nonce CSP、ローカルブランド SVG の削除)— 本パッケージは強化版サーバーを維持。上流の `.ps1` 削除と plugin-only 再構成もこの MCP サーバー構成には適用されません。
|
|
142
|
+
- **テスト**: MCP フロー、render-graphs(8 アサーション)、PowerShell 完全スイート(64 アサーション)すべて合格。
|
|
143
|
+
|
|
144
|
+
### v6.2.4
|
|
132
145
|
|
|
133
146
|
- **上流に合わせた brainstorm セッションの永続化**:`--project-dir` 指定時、companion はセッションキーを `.superpowers/brainstorm/.last-token`(オーナーのみ読み取り可、.gitignore 済み)に保存し、`.last-port` と並んで再起動後も再利用します——開いたままのブラウザタブは再起動後も接続を維持し、URL の再共有は不要です。一時 `/tmp` セッションでは従来どおり起動ごとにキーをローテーションします。明示的な `BRAINSTORM_TOKEN` 環境変数は常に優先され、ファイルには書き込まれません。強制的にローテーションするには、サーバー停止後に `.last-token` を削除してください。
|
|
134
147
|
- **トークンファイル読み取り経路の強化**(`readPrivateFile`):シンボリックリンクや複数リンクの `.last-token` は拒否され、セッションキーとして採用されなくなります。読み取りは `O_NOFOLLOW` 付き fd 経由で行い、identity を再検証し 0600 に締め付けます——既に強化済みの書き込み経路との非対称性を解消しました(独立したセキュリティレビューで発見)。
|
package/README.ko.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[English](README.md) | [繁體中文](README.zh-TW.md) | [日本語](README.ja.md) | [한국어](README.ko.md)
|
|
4
4
|
|
|
5
|
-
[](https://github.com/Poseidoncode/superpowers-mcp)
|
|
6
6
|
[](LICENSE)
|
|
7
7
|
|
|
8
8
|
이 문서는 원본 Superpowers 스킬 라이브러리를 독립적인 MCP Toolpack으로 패키징하기 위한 정보와 사용 지침을 요약한 것입니다.
|
|
@@ -112,7 +112,7 @@
|
|
|
112
112
|
이러한 스킬은 지원되는 IDE(Antigravity 또는 Cursor 등) 내에서 복잡한 메타 실행 패턴을 오케스트레이션하기 위해 설계되었습니다.
|
|
113
113
|
|
|
114
114
|
- **`subagent-driven-development`**: 서브에이전트를 구동하여 작업 실행
|
|
115
|
-
- **사용법**: 미리 정의된 계획을 작업별로 실행합니다. 시스템은 각 작업마다 새로운 "구현" 서브에이전트를 생성한 후, 통합된 **작업 리뷰어**(명세 준수 + 코드 품질) 서브에이전트와 마지막에 **전체 브랜치 최종 리뷰**를 실행합니다. **Pre-Flight Plan Review**는 실행 시작 전 작업 충돌을 스캔합니다.
|
|
115
|
+
- **사용법**: 미리 정의된 계획을 작업별로 실행합니다. 시스템은 각 작업마다 새로운 "구현" 서브에이전트를 생성한 후, 통합된 **작업 리뷰어**(명세 준수 + 코드 품질) 서브에이전트와 마지막에 **전체 브랜치 최종 리뷰**를 실행합니다. **Pre-Flight Plan Review**는 실행 시작 전 작업 충돌을 스캔합니다. 계획은 plan-scoped 워크스페이스(`.superpowers/sdd/<plan>/`)에서 실행되며, 컨트롤러는 멈추지 않고 충돌을 재결정하여 레저(ledger)에 기록(rulings)하고, 같은 형태의 작은 작업은 단일 디스패치로 배치 처리됩니다.
|
|
116
116
|
- **모델 선택**: 작업 복잡성에 따라 서브에이전트 모델 선택 — 기계적인 작업에는 저비용 모델, 아키텍처 및 미묘한 동시성 변경에는 고성능 모델
|
|
117
117
|
- **예시**: "subagent-driven-development 스킬을 읽고 docs/plans/feature-plan.md에 나열된 작업을 하나씩 실행해줘"
|
|
118
118
|
- **`dispatching-parallel-agents`**: 작업을 병렬 에이전트에 할당
|
|
@@ -128,7 +128,20 @@
|
|
|
128
128
|
|
|
129
129
|
## 🆕 최근 업데이트
|
|
130
130
|
|
|
131
|
-
### v6.
|
|
131
|
+
### v6.3.0 (최신)
|
|
132
|
+
|
|
133
|
+
- **상류 obra/superpowers v6.3.0 동기화** — 적용 가능한 개선 사항을 모두 채택하고, 포크 고유의 보안 강화와 PowerShell 지원은 유지.
|
|
134
|
+
- **brainstorming — 3경로 라우터**: 모든 요청을 사전에 `spike` / `bounded` / `architectural`로 분류하며, 절차의 양은 작업 규모에 맞춰 조정됩니다. 단 승인 게이트는 모든 경로에 동일하게 적용됩니다. 실행 중 숨은 복잡성이 발견되면 경로를 업그레이드 — 다운그레이드는 없습니다.
|
|
135
|
+
- **subagent-driven-development — 멈추지 않고 재결정(rulings, not stalls)**: 충돌·모호성·계획 결함은 컨트롤러가 직접 재결정하고 레저에 기록(`Ruling: ...`). 명시된 4가지 조건에서만 실행을 중지합니다. Pre-flight 충돌 스캔은 레저 테이블을 산출하고, 같은 형태의 소규모 작업은 단일 디스패치로 배치되며, 서브에이전트 대기는 경계 있는 구간(bounded stretches)을 사용합니다. 세 프롬프트 모두 no-subagents 계약 추가.
|
|
136
|
+
- **Hermes Agent 지원**: 새 `hermes-tools.md` 참조 파일이 스킬 액션을 Hermes 도구(`delegate_task`, `skill_view` 등)에 매핑.
|
|
137
|
+
- **Codex**: V1/V2 멀티에이전트 차이, `followup_task` 수정 라운드 재개, 이벤트 구독형 `wait_agent` 가이드.
|
|
138
|
+
- **writing-plans**: 계획 템플릿에 `Spec:` 필드 추가.
|
|
139
|
+
- **finishing-a-development-branch**: worktree 제거 거부 시 절차 — 임의로 `--force`하지 않음.
|
|
140
|
+
- **이중 에이전트 code review 수정**: 병합 경로에서 "Commit them to \<branch\>"를 선택해도 파일이 베이스 브랜치 밖에 남지 않음(finishing-a-development-branch). `sdd-workspace.ps1`의 슬러그 도출은 모든 플랫폼에서 `basename`과 일치(`PLAN.MD`는 `PLAN.MD` 유지).
|
|
141
|
+
- **의도적으로 미채택**: 상류 v6.3.0의 서버 단순화(loopback-only 바인딩, `O_NOFOLLOW` 읽기, nonce CSP, 로컬 브랜드 SVG 제거) — 본 패키지는 강화된 서버를 유지합니다. 상류의 `.ps1` 삭제와 plugin-only 재구성도 이 MCP 서버 구성에는 적용되지 않습니다.
|
|
142
|
+
- **테스트**: MCP 플로우, render-graphs(8개 어서션), PowerShell 전체 스위트(64개 어서션) 모두 통과.
|
|
143
|
+
|
|
144
|
+
### v6.2.4
|
|
132
145
|
|
|
133
146
|
- **업스트림 정렬 — brainstorm 세션 영속화**: `--project-dir` 사용 시 companion이 세션 키를 `.superpowers/brainstorm/.last-token`(소유자 전용, .gitignore 적용)에 저장하고 `.last-port`와 함께 재시작 후에도 재사용합니다 — 이미 열린 브라우저 탭은 재시작 후에도 연결이 유지되며 URL을 다시 공유할 필요가 없습니다. 임시 `/tmp` 세션은 기존처럼 호출마다 키를 교체하며, 명시적 `BRAINSTORM_TOKEN` 환경 변수는 항상 우선하고 파일에 기록되지 않습니다. 강제 교체를 원하면 서버 중지 후 `.last-token`을 삭제하세요.
|
|
134
147
|
- **토큰 파일 읽기 경로 강화** (`readPrivateFile`): 심볼릭 링크 또는 다중 링크 `.last-token`은 거부되어 세션 키로 채택되지 않습니다. 읽기는 `O_NOFOLLOW` fd를 통해 수행되고 identity를 재검증하며 0600으로 강화됩니다 — 이미 강화된 쓰기 경로와의 비대칭을 해소했습니다(독립 보안 리뷰에서 발견).
|
package/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[English](README.md) | [繁體中文](README.zh-TW.md) | [日本語](README.ja.md) | [한국어](README.ko.md)
|
|
4
4
|
|
|
5
|
-
[](https://github.com/Poseidoncode/superpowers-mcp)
|
|
6
6
|
[](LICENSE)
|
|
7
7
|
|
|
8
8
|
This document summarizes the information and usage instructions for packaging the original Superpowers skills library into an independent MCP Toolpack.
|
|
@@ -112,7 +112,7 @@ To help you choose the right skill, we've categorized them into 6 logical phases
|
|
|
112
112
|
These skills are designed for orchestrating complex meta-execution patterns within supported IDEs (like Antigravity or Cursor).
|
|
113
113
|
|
|
114
114
|
- **`subagent-driven-development`**: Driving sub-agents to execute tasks
|
|
115
|
-
- **Usage**: Used to execute a predefined plan task-by-task. The system spawns a fresh "implementer" sub-agent per task, followed by a consolidated **task reviewer** (spec compliance + code quality) sub-agent, plus a **whole-branch final review** at the end. A **Pre-Flight Plan Review** scans for task conflicts before execution begins.
|
|
115
|
+
- **Usage**: Used to execute a predefined plan task-by-task. The system spawns a fresh "implementer" sub-agent per task, followed by a consolidated **task reviewer** (spec compliance + code quality) sub-agent, plus a **whole-branch final review** at the end. A **Pre-Flight Plan Review** scans for task conflicts before execution begins. Plans run in plan-scoped workspaces (`.superpowers/sdd/<plan>/`), the controller rules on conflicts and records them in the ledger instead of stopping, and small same-shape tasks are batched into a single dispatch.
|
|
116
116
|
- **Model Selection**: Choose sub-agent models based on task complexity — cheaper models for mechanical work, capable models for architecture and subtle concurrency changes.
|
|
117
117
|
- **Example**: "Read the subagent-driven-development skill, then execute the tasks listed in docs/plans/feature-plan.md one by one."
|
|
118
118
|
- **`dispatching-parallel-agents`**: Dispatching tasks to parallel agents
|
|
@@ -128,7 +128,20 @@ These skills are designed for orchestrating complex meta-execution patterns with
|
|
|
128
128
|
|
|
129
129
|
## 🆕 Recent Updates
|
|
130
130
|
|
|
131
|
-
### v6.
|
|
131
|
+
### v6.3.0 (Latest)
|
|
132
|
+
|
|
133
|
+
- **Upstream sync with obra/superpowers v6.3.0** — all applicable improvements adopted, fork-specific security hardening and PowerShell support preserved.
|
|
134
|
+
- **brainstorming — three-path router**: every request is classified `spike` / `bounded` / `architectural` up front; the ceremony scales with the task but the approval gate never does. Hidden complexity upgrades the path mid-task — never downgrades.
|
|
135
|
+
- **subagent-driven-development — rulings, not stalls**: conflicts, ambiguities, and plan defects are ruled on and ledgered (`Ruling: ...`) instead of parking the session on a human; only four named conditions stop execution. Pre-flight conflict scans produce a ledgered table, small same-shape tasks batch into one dispatch, subagent waits use bounded stretches, and all three prompts carry the no-subagents contract.
|
|
136
|
+
- **Hermes Agent support**: new `hermes-tools.md` reference maps skill actions to Hermes tools (`delegate_task`, `skill_view`, …).
|
|
137
|
+
- **Codex**: V1/V2 multi-agent differences, `followup_task` fix-round resume, and event-subscription `wait_agent` guidance.
|
|
138
|
+
- **writing-plans**: plan template now carries a `Spec:` field.
|
|
139
|
+
- **finishing-a-development-branch**: worktree removal-refused procedure — never `--force` on your own initiative.
|
|
140
|
+
- **Fixes from dual-agent code review**: merged-path "commit them to \<branch\>" no longer strands files outside the base branch (finishing-a-development-branch); `sdd-workspace.ps1` slug derivation matches `basename` on all platforms (`PLAN.MD` stays `PLAN.MD`).
|
|
141
|
+
- **Not adopted (deliberate)**: upstream's v6.3.0 server simplification (it removed loopback-only enforcement, `O_NOFOLLOW` reads, nonce CSP, and the local brand SVG) — this package keeps its hardened server; upstream's `.ps1` deletions and plugin-only restructuring also don't apply to this MCP server layout.
|
|
142
|
+
- **Tests**: MCP flow, render-graphs (8 assertions), and the full PowerShell suite (64 assertions) all pass.
|
|
143
|
+
|
|
144
|
+
### v6.2.4
|
|
132
145
|
|
|
133
146
|
- **Upstream alignment — persistent brainstorm sessions**: with `--project-dir`, the companion now persists its session key to `.superpowers/brainstorm/.last-token` (owner-only, gitignored) alongside `.last-port` and reuses it across restarts — an already-open browser tab stays connected after a restart, no URL re-sharing needed. Ephemeral `/tmp` sessions keep rotating the key per invocation, and an explicit `BRAINSTORM_TOKEN` env var still wins and is never persisted. Delete `.last-token` (server stopped) to force a fresh key.
|
|
134
147
|
- **Token-file read path hardened** (`readPrivateFile`): symlinked or multi-link `.last-token` files are rejected instead of being adopted as the session key, with the read performed through an `O_NOFOLLOW` fd whose identity is re-checked and tightened to 0600 — closing the asymmetry with the already-hardened write path (found by independent security review).
|
package/README.zh-TW.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[English](README.md) | [繁體中文](README.zh-TW.md) | [日本語](README.ja.md) | [한국어](README.ko.md)
|
|
4
4
|
|
|
5
|
-
[](https://github.com/Poseidoncode/superpowers-mcp)
|
|
6
6
|
[](LICENSE)
|
|
7
7
|
|
|
8
8
|
本文檔總結了將原始 Superpowers 技能庫打包成獨立 MCP Toolpack 的相關資訊與使用說明。
|
|
@@ -112,7 +112,7 @@
|
|
|
112
112
|
這是特別針對在支援多重代理 (Multi-Agent) 協作或具備強大推論能力的 IDE(如 Antigravity 或 Cursor)中所設計的高階操作技巧。
|
|
113
113
|
|
|
114
114
|
- **`subagent-driven-development`**: 驅動子代理執行複雜任務
|
|
115
|
-
- **具體用法**:適用於「執行已規劃好的詳細計畫」。針對每一個小任務,AI 會派發全新的「實作代理」去寫 Code,完成後啟動合併的 **task reviewer**(規格 + 品質一次審查),並在最後執行 **whole-branch final review**(全分支最終審查)。在執行前還會進行 **Pre-Flight Plan Review
|
|
115
|
+
- **具體用法**:適用於「執行已規劃好的詳細計畫」。針對每一個小任務,AI 會派發全新的「實作代理」去寫 Code,完成後啟動合併的 **task reviewer**(規格 + 品質一次審查),並在最後執行 **whole-branch final review**(全分支最終審查)。在執行前還會進行 **Pre-Flight Plan Review**,掃描任務之間的潛在衝突。計畫在 plan-scoped 工作區(`.superpowers/sdd/<plan>/`)執行;控制器對衝突直接裁決並記錄於 ledger(rulings),不再停擺等待;相同形狀的小任務會批次合併為單次派發。
|
|
116
116
|
- **模型選擇**:根據任務複雜度選擇子代理模型 — 機械性工作用低成本模型,架構設計與細微併發變更需要強大模型。
|
|
117
117
|
- **指令範例**:「用 read_skill 讀取 subagent-driven-development 技能,然後依照 docs/plans/feature-plan.md 的內容逐一派發子代理實作。」
|
|
118
118
|
- **`dispatching-parallel-agents`**: 派發平行代理同步執行任務
|
|
@@ -128,7 +128,20 @@
|
|
|
128
128
|
|
|
129
129
|
## 🆕 最近更新
|
|
130
130
|
|
|
131
|
-
### v6.
|
|
131
|
+
### v6.3.0 (最新版)
|
|
132
|
+
|
|
133
|
+
- **對齊上游 obra/superpowers v6.3.0** — 採用所有適用改進,保留本 fork 的安全強化與 PowerShell 支援。
|
|
134
|
+
- **brainstorming — 三路分類流程(three-path router)**:每個請求先分類為 `spike` / `bounded` / `architectural`;流程深度隨任務規模調整,但審批門檻永遠不變。隱藏複雜度會在執行途中升級路徑——絕不降級。
|
|
135
|
+
- **subagent-driven-development — 裁決而非停擺(rulings, not stalls)**:衝突、歧義與計畫缺陷由控制器直接裁決並記錄於 ledger(`Ruling: ...`),不再等待人類;僅四種明列情況會停止執行。Pre-flight 衝突掃描產出 ledger 表格,相同形狀的小任務批次合併為單次派發,子代理等待改用有界區間,三個 prompt 都帶有 no-subagents 契約。
|
|
136
|
+
- **Hermes Agent 支援**:新增 `hermes-tools.md` 參考檔,將技能動作對應到 Hermes 工具(`delegate_task`、`skill_view` 等)。
|
|
137
|
+
- **Codex**:V1/V2 多代理差異、`followup_task` 修復輪恢復、事件訂閱式 `wait_agent` 指引。
|
|
138
|
+
- **writing-plans**:計畫範本新增 `Spec:` 欄位。
|
|
139
|
+
- **finishing-a-development-branch**:worktree 移除被拒時的處理程序——絕不擅自 `--force`。
|
|
140
|
+
- **雙代理 code review 修復**:merged path 下「Commit them to \<branch\>」不再讓檔案遺留在 base branch 之外(finishing-a-development-branch);`sdd-workspace.ps1` 的 slug 推導在所有平台與 `basename` 一致(`PLAN.MD` 維持 `PLAN.MD`)。
|
|
141
|
+
- **刻意未採用**:上游 v6.3.0 的 server 簡化(移除 loopback-only 綁定、`O_NOFOLLOW` 讀取、nonce CSP、本地品牌 SVG)——本套件保留加固版 server;上游刪除 `.ps1` 及轉為 plugin-only 架構亦不適用於本 MCP server 佈局。
|
|
142
|
+
- **測試**:MCP 流程、render-graphs(8 斷言)、完整 PowerShell 套件(64 斷言)全部通過。
|
|
143
|
+
|
|
144
|
+
### v6.2.4
|
|
132
145
|
|
|
133
146
|
- **對齊上游 — brainstorm 持久化 session**:搭配 `--project-dir` 時,companion 現在會把 session 金鑰持久化到 `.superpowers/brainstorm/.last-token`(僅擁有者可讀、已列入 .gitignore),與 `.last-port` 並存並在重啟後重用——已開啟的瀏覽器分頁在重啟後依然保持連線,無需重新分享 URL。暫存 `/tmp` session 仍維持每次啟動輪換金鑰;明確設定的 `BRAINSTORM_TOKEN` env 依然優先且永不寫入檔案。若要強制輪換,請在停止伺服器後刪除 `.last-token`。
|
|
134
147
|
- **Token 檔讀取路徑加固**(`readPrivateFile`):symlink 或多重連結的 `.last-token` 現在會被拒絕,不再被採納為 session 金鑰;讀取透過 `O_NOFOLLOW` fd 進行,身份複查並收緊為 0600——補上與已加固寫入路徑之間的不對稱(由獨立資安審查發現)。
|
package/out/server.js
CHANGED
|
@@ -49,7 +49,7 @@ Set the \`cycles\` parameter to \`"ref"\` to resolve cyclical schemas with defs.
|
|
|
49
49
|
`}var vs=class{constructor(t=hf.default.stdin,r=hf.default.stdout){this._stdin=t,this._stdout=r,this._readBuffer=new gs,this._started=!1,this._ondata=n=>{this._readBuffer.append(n),this.processReadBuffer()},this._onerror=n=>{this.onerror?.(n)}}async start(){if(this._started)throw new Error("StdioServerTransport already started! If using Server class, note that connect() calls start() automatically.");this._started=!0,this._stdin.on("data",this._ondata),this._stdin.on("error",this._onerror)}processReadBuffer(){for(;;)try{let t=this._readBuffer.readMessage();if(t===null)break;this.onmessage?.(t)}catch(t){this.onerror?.(t)}}async close(){this._stdin.off("data",this._ondata),this._stdin.off("error",this._onerror),this._stdin.listenerCount("data")===0&&this._stdin.pause(),this._readBuffer.clear(),this.onclose?.()}send(t){return new Promise(r=>{let n=j_(t);this._stdout.write(n)?r():this._stdout.once("drain",r)})}};var ze=sr(require("fs/promises")),ye=sr(require("path")),$n=10*1024*1024,ys=class{skillsPath;cachedSkills=null;loadingPromise=null;skillMap=new Map;contentCache=new Map;constructor(t){this.skillsPath=t}stripQuotes(t){return t.replace(/^"(.*)"$|^'(.*)'$/,"$1$2").trim()}parseFrontmatter(t){let r=t;if(r.charCodeAt(0)===65279&&(r=r.slice(1)),!r.startsWith("---"))return{name:"",description:""};let n=r.split(/\r?\n/),o=[],i=!1;for(let u=1;u<n.length;u++){if(n[u].trim()==="---"){i=!0;break}o.push(n[u])}if(!i)return{name:"",description:""};let a="",s="",c=!1;for(let u of o){let l=u.match(/^name:\s*(.*?)\s*$/);if(l){a=this.stripQuotes(l[1]),c=!1;continue}let d=u.match(/^description:\s*(.*?)\s*$/);if(d){s=this.stripQuotes(d[1]),c=!0;continue}c&&/^\s+/.test(u)?s+=" "+u.trim():c=!1}return{name:a,description:s}}async exists(t){try{return await ze.access(t),!0}catch{return!1}}async listSkills(t=!1){if(this.cachedSkills&&!t)return this.cachedSkills;if(this.loadingPromise&&!t)return this.loadingPromise;t&&this.contentCache.clear();let r=this.internalListSkills();this.loadingPromise=r;try{return await r}finally{this.loadingPromise===r&&(this.loadingPromise=null)}}async internalListSkills(){if(!await this.exists(this.skillsPath))return this.cachedSkills??[];let t=[],r=new Map,n=!1;try{let o=await ze.readdir(this.skillsPath,{withFileTypes:!0});for(let i of o){if(!i.isDirectory()&&!i.isSymbolicLink())continue;let a=ye.join(this.skillsPath,i.name),s=ye.join(a,"SKILL.md");if(await this.exists(s))try{let c=await this.readFileNoFollow(s,this.skillsPath),{name:u,description:l}=this.parseFrontmatter(c),m={name:u||i.name,description:l,skillPath:s};t.push(m),r.set(m.name.toLowerCase(),m);let p=ye.basename(ye.dirname(m.skillPath)).toLowerCase();r.set(p,m)}catch{process.stderr.write(`Warning: Failed to read skill file in directory "${i.name}"
|
|
50
50
|
`)}}n=!0}catch(o){process.stderr.write(`Error reading skills directory: ${String(o)}
|
|
51
51
|
`)}return n&&(this.skillMap=r,this.cachedSkills=t.sort((o,i)=>o.name.localeCompare(i.name))),this.cachedSkills??[]}async findSkill(t){let r=typeof t=="string"?t.trim():"";if(!(!r||r==="."||r===".."||r.includes("/")||r.includes("\\")||r.includes("\0")))return this.cachedSkills||await this.listSkills(),this.skillMap.get(r.toLowerCase())}async readFileNoFollow(t,r){let n=ye.resolve(t),o=ye.resolve(r),i=async()=>{let l=await ze.realpath(n),d=await ze.realpath(o);if(!(await ze.stat(d)).isDirectory())throw new Error("Skills directory must be a directory");let p=ye.relative(d,l);if(p===".."||p.startsWith(`..${ye.sep}`)||ye.isAbsolute(p))throw new Error("File is outside skills directory");let h=await ze.stat(l);if(!h.isFile()||h.nlink!==1)throw new Error("Skill path is not a regular file");if(h.size>$n)throw new Error("Skill file exceeds size limit");return{realFilePath:l,stat:h}},a=await i(),s=await i();if(a.realFilePath!==s.realFilePath||a.stat.dev!==s.stat.dev||a.stat.ino!==s.stat.ino)throw new Error("File changed while validating its path");let c=process.platform==="win32"?0:ze.constants.O_NOFOLLOW,u=await ze.open(s.realFilePath,ze.constants.O_RDONLY|c);try{let l=await u.stat();if(!l.isFile()||l.nlink!==1||l.dev!==s.stat.dev||l.ino!==s.stat.ino||l.size>$n)throw new Error("File changed while opening");let d=[],m=64*1024,p=0;for(;p<=$n;){let v=Buffer.allocUnsafe(Math.min(m,$n+1-p)),{bytesRead:$}=await u.read(v,0,v.length,null);if($===0)break;if(p+=$,d.push(v.subarray(0,$)),p>$n)throw new Error("Skill file exceeds size limit")}let h=await u.stat();if(h.size>$n||h.dev!==s.stat.dev||h.ino!==s.stat.ino)throw new Error("File changed while reading");return Buffer.concat(d,p).toString("utf-8")}finally{await u.close()}}async readSkillContent(t,r=!1){let n=ye.resolve(t),o=ye.resolve(this.skillsPath),i=n,a=o;try{i=await ze.realpath(n),a=await ze.realpath(o)}catch{}let s=ye.relative(a,i);if(s.startsWith("..")||ye.isAbsolute(s))throw new Error(`Access denied: path "${t}" is outside skills directory`);if(this.contentCache.has(i)&&!r)return this.contentCache.get(i);try{let c=await this.readFileNoFollow(t,this.skillsPath);c.charCodeAt(0)===65279&&(c=c.slice(1));let u=c.replace(/^---\s*\r?\n[\s\S]*?\r?\n---\s*\r?\n?/,"").trim();return this.contentCache.set(i,u),u}catch(c){throw new Error(`Failed to read skill content: ${c instanceof Error?c.message:String(c)}`)}}clearCache(){this.cachedSkills=null,this.loadingPromise=null,this.skillMap.clear(),this.contentCache.clear()}};function jT(){let e=process.env.SKILLS_PATH;if(e){let r=ct.resolve(e),n=ct.normalize(r).toLowerCase(),o=ct.parse(r).root.toLowerCase();if(n===o||["/etc","/var","/bin","/sbin","/usr","/root","/sys","/proc","/dev","c:\\windows","c:\\program files","c:\\program files (x86)"].some(s=>n===s||n.startsWith(s+ct.sep)))process.stderr.write(`Warning: Potentially unsafe SKILLS_PATH: "${e}". Fallback to default.
|
|
52
|
-
`);else return r}let t=ct.join(__dirname,"..","skills");return E_.existsSync(t)?t:ct.join(__dirname,"skills")}var O_=jT(),It=new ys(O_),Pr=new hs({name:"superpowers-mcp",version:"6.
|
|
52
|
+
`);else return r}let t=ct.join(__dirname,"..","skills");return E_.existsSync(t)?t:ct.join(__dirname,"skills")}var O_=jT(),It=new ys(O_),Pr=new hs({name:"superpowers-mcp",version:"6.3.0"},{capabilities:{resources:{subscribe:!1},prompts:{},tools:{}}});Pr.setRequestHandler(od,async()=>({resources:(await It.listSkills()).map(t=>({uri:`skill://superpowers/${encodeURIComponent(t.name)}`,name:t.name,description:t.description,mimeType:"text/markdown"}))}));Pr.setRequestHandler(ad,async e=>{let t=e.params.uri,r=t.match(/^skill:\/\/superpowers\/(.+)$/);if(!r)throw new D(M.InvalidRequest,`Invalid skill URI: ${t}`);let n;try{n=decodeURIComponent(r[1])}catch{throw new D(M.InvalidRequest,`Invalid skill URI: ${t}`)}let o=await It.findSkill(n);if(!o)throw new D(M.InvalidRequest,`Skill not found: ${n}`);try{let i=await It.readSkillContent(o.skillPath);return{contents:[{uri:t,mimeType:"text/markdown",text:i}]}}catch{throw new D(M.InternalError,"Failed to read skill content safely.")}});Pr.setRequestHandler(sd,async()=>({prompts:[{name:"session-start",description:"Inject the Superpowers context into an AI agent session. Tells the agent it has superpowers and how to use the skill system."}]}));Pr.setRequestHandler(cd,async e=>{if(e.params.name!=="session-start")throw new D(M.InvalidRequest,`Unknown prompt: ${e.params.name}`);let t=await It.findSkill("using-superpowers"),r="";if(t)try{r=await It.readSkillContent(t.skillPath)}catch{r=`# Superpowers
|
|
53
53
|
|
|
54
54
|
You have superpowers. Use the read_skill and list_skills tools to discover and load skills.`}else{let o=ct.join(O_,"using-superpowers","SKILL.md");try{r=await It.readSkillContent(o)}catch{r=`# Superpowers
|
|
55
55
|
|
package/package.json
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"name": "superpowers-mcp",
|
|
3
3
|
"displayName": "Superpowers MCP",
|
|
4
4
|
"description": "Superpowers skills library (TDD, debugging, collaboration workflows) as an MCP server for VSCode and Antigravity",
|
|
5
|
-
"version": "6.
|
|
5
|
+
"version": "6.3.0",
|
|
6
6
|
"publisher": "superpowers",
|
|
7
7
|
"license": "MIT",
|
|
8
8
|
"repository": {
|
|
@@ -7,20 +7,91 @@ description: "You MUST use this before any creative work - creating features, bu
|
|
|
7
7
|
|
|
8
8
|
Help turn ideas into fully formed designs and specs through natural collaborative dialogue.
|
|
9
9
|
|
|
10
|
-
Start by
|
|
10
|
+
Start by classifying how much process the request needs, then work
|
|
11
|
+
through your path: understand the context, refine the idea, present a
|
|
12
|
+
design, and get your human partner's approval.
|
|
11
13
|
|
|
12
14
|
<HARD-GATE>
|
|
13
|
-
Do NOT invoke any implementation skill, write any code, scaffold any
|
|
15
|
+
Do NOT invoke any implementation skill, write any code, scaffold any
|
|
16
|
+
project, or take any implementation action until you have told your
|
|
17
|
+
human partner what you intend and they have approved it. This applies
|
|
18
|
+
to EVERY task on EVERY path below — the ceremony scales with the task;
|
|
19
|
+
the approval gate never does.
|
|
14
20
|
</HARD-GATE>
|
|
15
21
|
|
|
16
|
-
##
|
|
17
|
-
|
|
18
|
-
|
|
22
|
+
## Three Paths
|
|
23
|
+
|
|
24
|
+
Before your first question, classify the request and say the
|
|
25
|
+
classification out loud — "this looks bounded, so I'll present a short
|
|
26
|
+
design here rather than write a spec" — so your human partner can
|
|
27
|
+
override it:
|
|
28
|
+
|
|
29
|
+
- **Spike** — a feasibility question ("can we...", "is it possible...",
|
|
30
|
+
"quick and dirty is fine") whose output is an answer, not code you
|
|
31
|
+
keep. Present the question and what you'll try in 2-3 sentences, get
|
|
32
|
+
a nod, then find out as cheaply as correctness allows. No design
|
|
33
|
+
doc, no spec file. Report findings as a recommendation; anything you
|
|
34
|
+
built stays labeled throwaway.
|
|
35
|
+
- **Bounded** — a well-scoped change to code that already exists in
|
|
36
|
+
this repo: a new flag, a small endpoint, a one-file fix.
|
|
37
|
+
Understanding the kind of app is not enough — bounded means the flow
|
|
38
|
+
you are changing is already here to read. If there is no existing
|
|
39
|
+
flow to change, the task is not bounded. Ask the clarifying
|
|
40
|
+
questions that matter, present a short design IN CHAT (a few
|
|
41
|
+
sentences to a few short paragraphs), and STOP. Implementation
|
|
42
|
+
starts only after your human partner says yes to that design — a
|
|
43
|
+
bounded task's approval is as hard a gate as an architectural
|
|
44
|
+
one. No spec file, no implementation plan document.
|
|
45
|
+
- **Architectural** — new projects, new subsystems, changes that
|
|
46
|
+
restructure how components fit together or alter interfaces others
|
|
47
|
+
depend on. Follow the full process: questions, approaches, sectioned
|
|
48
|
+
design, written spec, then the writing-plans skill.
|
|
49
|
+
|
|
50
|
+
When in doubt between two paths, take the heavier one. The ratchet is
|
|
51
|
+
one-way: hidden complexity discovered mid-task upgrades the path —
|
|
52
|
+
stop, say so, and step up. Nothing downgrades mid-task.
|
|
53
|
+
|
|
54
|
+
## Anti-Pattern: "Too Simple To Need Approval"
|
|
55
|
+
|
|
56
|
+
Every path ends with your human partner approving your intent before
|
|
57
|
+
implementation. A todo list, a single-function utility, a config
|
|
58
|
+
change — the design may be two sentences in chat, but you MUST present
|
|
59
|
+
it and get approval. "Simple" tasks are where unexamined assumptions
|
|
60
|
+
cause the most wasted work. What scales with simplicity is the
|
|
61
|
+
artifact, never the approval.
|
|
62
|
+
|
|
63
|
+
## Red Flags
|
|
64
|
+
|
|
65
|
+
| Thought | Reality |
|
|
66
|
+
|---------|---------|
|
|
67
|
+
| "This is too simple to need a design" | Simple means a short design, not no design. Two sentences in chat, then approval. |
|
|
68
|
+
| "I'll call it bounded and skip the spec" | Reaching for a label to skip work IS the doubt — take the heavier path. |
|
|
69
|
+
| "It's bounded and the design is obvious — I'll start while they read it" | The gate is the approval, not the design's length. Present, then stop until you hear yes. |
|
|
70
|
+
| "I understand this kind of app, so it's bounded" | Bounded measures the repo, not your familiarity. A new project has no existing flow — it is architectural. |
|
|
71
|
+
| "The spike works, so I'll keep the code" | A spike's output is an answer. Keeping the code is a new request — classify it. |
|
|
72
|
+
| "It grew, but I'm almost done — no need to re-classify" | Hidden complexity upgrades the path mid-task. Stop and say so. |
|
|
73
|
+
| "They approved the spike, so the follow-up change is approved too" | Each task gets its own classification and its own approval. |
|
|
19
74
|
|
|
20
75
|
## Checklist
|
|
21
76
|
|
|
22
|
-
|
|
77
|
+
Classify first, announce the path, then create a task for each item on
|
|
78
|
+
your path and complete them in order.
|
|
79
|
+
|
|
80
|
+
**Spike:**
|
|
81
|
+
1. **Explore project context** — enough to frame the probe
|
|
82
|
+
2. **Present question + probe plan** — 2-3 sentences
|
|
83
|
+
3. **Get approval** — a nod is enough
|
|
84
|
+
4. **Investigate** — as cheaply as correctness allows
|
|
85
|
+
5. **Report findings** — a recommendation; label anything built as throwaway
|
|
86
|
+
|
|
87
|
+
**Bounded:**
|
|
88
|
+
1. **Explore project context** — check files, docs, recent commits
|
|
89
|
+
2. **Ask clarifying questions** — one at a time, the ones that matter
|
|
90
|
+
3. **Present short design in chat** — approach, files touched, testing
|
|
91
|
+
4. **Get approval** — STOP and wait for an explicit yes; presenting the design and starting in the same breath is skipping the gate
|
|
92
|
+
5. **Implement** — proceed with the normal development workflow (TDD applies); no plan document
|
|
23
93
|
|
|
94
|
+
**Architectural:**
|
|
24
95
|
1. **Explore project context** — check files, docs, recent commits
|
|
25
96
|
2. **Offer the visual companion just-in-time** — NOT upfront. The first time a question would genuinely be clearer shown than described, offer it then (its own message); on approval its browser tab opens for you. If no visual question ever arises, never offer it. See the Visual Companion section below.
|
|
26
97
|
3. **Ask clarifying questions** — one at a time, understand purpose/constraints/success criteria
|
|
@@ -35,6 +106,13 @@ You MUST create a task for each of these items and complete them in order:
|
|
|
35
106
|
|
|
36
107
|
```dot
|
|
37
108
|
digraph brainstorming {
|
|
109
|
+
"Classify: spike / bounded / architectural" [shape=diamond];
|
|
110
|
+
"Present question + probe (2-3 sentences)" [shape=box];
|
|
111
|
+
"Ask clarifying questions (bounded)" [shape=box];
|
|
112
|
+
"Present short design in chat" [shape=box];
|
|
113
|
+
"Human approves?" [shape=diamond];
|
|
114
|
+
"Investigate; report recommendation" [shape=doublecircle];
|
|
115
|
+
"Implement via normal workflow (no plan doc)" [shape=doublecircle];
|
|
38
116
|
"Explore project context" [shape=box];
|
|
39
117
|
"Ask clarifying questions" [shape=box];
|
|
40
118
|
"Propose 2-3 approaches" [shape=box];
|
|
@@ -44,7 +122,17 @@ digraph brainstorming {
|
|
|
44
122
|
"Spec self-review\n(fix inline)" [shape=box];
|
|
45
123
|
"User reviews spec?" [shape=diamond];
|
|
46
124
|
"Invoke writing-plans skill" [shape=doublecircle];
|
|
47
|
-
|
|
125
|
+
"Hidden complexity? Upgrade path" [shape=box];
|
|
126
|
+
|
|
127
|
+
"Classify: spike / bounded / architectural" -> "Present question + probe (2-3 sentences)" [label="spike"];
|
|
128
|
+
"Classify: spike / bounded / architectural" -> "Ask clarifying questions (bounded)" [label="bounded"];
|
|
129
|
+
"Classify: spike / bounded / architectural" -> "Explore project context" [label="architectural"];
|
|
130
|
+
"Present question + probe (2-3 sentences)" -> "Human approves?";
|
|
131
|
+
"Ask clarifying questions (bounded)" -> "Present short design in chat";
|
|
132
|
+
"Present short design in chat" -> "Human approves?";
|
|
133
|
+
"Human approves?" -> "Investigate; report recommendation" [label="spike: yes"];
|
|
134
|
+
"Human approves?" -> "Implement via normal workflow (no plan doc)" [label="bounded: yes"];
|
|
135
|
+
"Hidden complexity? Upgrade path" -> "Classify: spike / bounded / architectural";
|
|
48
136
|
"Explore project context" -> "Ask clarifying questions";
|
|
49
137
|
"Ask clarifying questions" -> "Propose 2-3 approaches";
|
|
50
138
|
"Propose 2-3 approaches" -> "Present design sections";
|
|
@@ -58,10 +146,21 @@ digraph brainstorming {
|
|
|
58
146
|
}
|
|
59
147
|
```
|
|
60
148
|
|
|
61
|
-
**
|
|
149
|
+
**Terminal states are path-bound.** Architectural: the ONLY skill you
|
|
150
|
+
invoke after brainstorming is writing-plans — never frontend-design,
|
|
151
|
+
mcp-builder, or any other implementation skill. Bounded: after
|
|
152
|
+
approval, implementation proceeds directly through the normal
|
|
153
|
+
development workflow; no plan document. Spike: the terminal state is a
|
|
154
|
+
reported recommendation.
|
|
62
155
|
|
|
63
156
|
## The Process
|
|
64
157
|
|
|
158
|
+
The subsections below serve the bounded and architectural paths (a
|
|
159
|
+
spike stops at "present the probe, get a nod"). Sections from
|
|
160
|
+
**Exploring approaches** onward are architectural-path depth — for
|
|
161
|
+
bounded work, context plus a few questions plus a short in-chat design
|
|
162
|
+
is the whole process.
|
|
163
|
+
|
|
65
164
|
**Understanding the idea:**
|
|
66
165
|
|
|
67
166
|
- Check out the current project state first (files, docs, recent commits)
|
|
@@ -100,7 +199,7 @@ digraph brainstorming {
|
|
|
100
199
|
- Where existing code has problems that affect the work (e.g., a file that's grown too large, unclear boundaries, tangled responsibilities), include targeted improvements as part of the design - the way a good developer improves code they're working in.
|
|
101
200
|
- Don't propose unrelated refactoring. Stay focused on what serves the current goal.
|
|
102
201
|
|
|
103
|
-
## After the Design
|
|
202
|
+
## After the Design (architectural path)
|
|
104
203
|
|
|
105
204
|
**Documentation:**
|
|
106
205
|
|
|
@@ -174,6 +174,35 @@ git worktree remove "$WORKTREE_PATH"
|
|
|
174
174
|
git worktree prune # Self-healing: clean up any stale registrations
|
|
175
175
|
```
|
|
176
176
|
|
|
177
|
+
**If removal is refused** (`contains modified or untracked files`): the
|
|
178
|
+
worktree holds files that exist nowhere else — uncommitted plans, notes,
|
|
179
|
+
or scratch work. Never `--force` on your own initiative. Show your human
|
|
180
|
+
partner what is at stake and ask:
|
|
181
|
+
|
|
182
|
+
```bash
|
|
183
|
+
git -C "$WORKTREE_PATH" status --porcelain -uall
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
```
|
|
187
|
+
Worktree removal refused — these files were never committed:
|
|
188
|
+
|
|
189
|
+
<file list>
|
|
190
|
+
|
|
191
|
+
1. Commit them to <branch> before cleanup
|
|
192
|
+
2. Move them into <main repo root>
|
|
193
|
+
3. Delete them (unrecoverable)
|
|
194
|
+
|
|
195
|
+
Which?
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
Carry out the choice, then remove the worktree.
|
|
199
|
+
|
|
200
|
+
**If you chose "Commit them to <branch>" after a merge (Option 1):** the
|
|
201
|
+
new commit sits on top of the merge, so `git branch -d <feature-branch>`
|
|
202
|
+
will refuse ("not fully merged"). Merge the branch into the base branch
|
|
203
|
+
again — or cherry-pick the new commit there — before cleanup, so the
|
|
204
|
+
saved files survive the branch deletion.
|
|
205
|
+
|
|
177
206
|
**Otherwise:** The host environment owns this workspace — leave it in
|
|
178
207
|
place. If your platform provides a workspace-exit tool, use it.
|
|
179
208
|
|
|
@@ -196,6 +225,7 @@ place. If your platform provides a workspace-exit tool, use it.
|
|
|
196
225
|
| "'Yeah, get rid of it' counts as confirmation" | Only the typed word `discard` authorizes deletion. |
|
|
197
226
|
| "The PR is up, so the worktree is clutter now" | PR feedback gets fixed in that worktree. It stays until the work lands. |
|
|
198
227
|
| "This other worktree looks stale — I'll clean it too" | Clean up only worktrees under `.worktrees/` or `worktrees/`. Everything else belongs to the host. |
|
|
228
|
+
| "Removal refused — `--force` is just finishing the cleanup" | The refusal means files exist only in that worktree. `--force` destroys them permanently. Show your human partner and ask. |
|
|
199
229
|
| "The merged-result failure is probably flaky" | A failing merged result stops everything. Branch and worktree stay put while you investigate. |
|
|
200
230
|
| "The base branch is obviously main" | Confirm the fork point or ask. Merging into the wrong base is expensive to undo. |
|
|
201
231
|
| "The push was rejected — force-push will fix it" | A rejected push means the remote moved. Investigate; force-push only on your human partner's explicit request. |
|
|
@@ -34,6 +34,15 @@ Subagent (general-purpose):
|
|
|
34
34
|
|
|
35
35
|
Your review is read-only on this checkout. Do not mutate the working tree, the index, HEAD, or branch state in any way. Use tools like `git show`, `git diff`, and `git log` to inspect history. If you need a working copy of a different revision, check it out into a separate temporary directory (e.g. `git worktree add /tmp/review-[SHA] [SHA]`) — never move HEAD on this checkout.
|
|
36
36
|
|
|
37
|
+
## You Do Not Dispatch Subagents
|
|
38
|
+
|
|
39
|
+
Do all of this review yourself. Never spawn a subagent to review part
|
|
40
|
+
of the diff, and never spawn another reviewer for a second opinion.
|
|
41
|
+
This process already provides every review seat the work gets; a
|
|
42
|
+
reviewer you spawn duplicates one of them at full cost, and its
|
|
43
|
+
verdict counts for nothing. If the diff feels too large for one
|
|
44
|
+
pass, review it in passes yourself and say so in your report.
|
|
45
|
+
|
|
37
46
|
## What to Check
|
|
38
47
|
|
|
39
48
|
**Plan alignment:**
|
|
@@ -14,7 +14,21 @@ Execute plan by dispatching a fresh implementer subagent per task, a task review
|
|
|
14
14
|
**Narration:** between tool calls, narrate at most one short line — the
|
|
15
15
|
ledger and the tool results carry the record.
|
|
16
16
|
|
|
17
|
-
**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are
|
|
17
|
+
**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
|
|
18
|
+
|
|
19
|
+
**Rulings, not stalls.** A running plan does not wait on a human. Conflicts,
|
|
20
|
+
ambiguities, plan defects, a cap you would have asked to exceed — decide
|
|
21
|
+
them. The spec is the binding authority, the plan is its argument, and your
|
|
22
|
+
judgment settles what neither answers. Record every decision in the ledger as
|
|
23
|
+
`Ruling: <what you decided> — <why> — <what it costs if wrong>`, and keep
|
|
24
|
+
going. A wrong ruling costs rework your human partner can see and undo; a
|
|
25
|
+
session parked on a question costs their whole day and buys nothing.
|
|
26
|
+
|
|
27
|
+
Four things stop you, and only these: an irreversible or destructive
|
|
28
|
+
operation; a security-sensitive action; a side effect outside this worktree
|
|
29
|
+
that norms say you ask about first (a merge, a push to a shared branch, a
|
|
30
|
+
publish); and a plan so broken that every path forward is a guess. For those,
|
|
31
|
+
stop and ask.
|
|
18
32
|
|
|
19
33
|
## When to Use
|
|
20
34
|
|
|
@@ -57,14 +71,14 @@ digraph process {
|
|
|
57
71
|
"Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" [shape=box];
|
|
58
72
|
"Spec ✅ and quality approved?" [shape=diamond];
|
|
59
73
|
"Finding conflicts with plan text?" [shape=diamond];
|
|
60
|
-
"
|
|
74
|
+
"Rule on the conflict, ledger the ruling" [shape=box];
|
|
61
75
|
"Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box];
|
|
62
76
|
"Dispatch scoped re-review (./re-review-prompt.md)" [shape=box];
|
|
63
77
|
"All findings addressed?" [shape=diamond];
|
|
64
78
|
"R = 5?" [shape=diamond];
|
|
65
79
|
"Adjudicate each open finding" [shape=box];
|
|
66
80
|
"Any load-bearing finding?" [shape=diamond];
|
|
67
|
-
"
|
|
81
|
+
"Rule and continue; stop only if every path forward is a guess" [shape=box];
|
|
68
82
|
"Park findings in ledger with rulings" [shape=box];
|
|
69
83
|
"Append completion to ledger, mark todo complete" [shape=box];
|
|
70
84
|
}
|
|
@@ -85,8 +99,8 @@ digraph process {
|
|
|
85
99
|
"Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" -> "Spec ✅ and quality approved?";
|
|
86
100
|
"Spec ✅ and quality approved?" -> "Append completion to ledger, mark todo complete" [label="yes"];
|
|
87
101
|
"Spec ✅ and quality approved?" -> "Finding conflicts with plan text?" [label="no"];
|
|
88
|
-
"Finding conflicts with plan text?" -> "
|
|
89
|
-
"
|
|
102
|
+
"Finding conflicts with plan text?" -> "Rule on the conflict, ledger the ruling" [label="yes"];
|
|
103
|
+
"Rule on the conflict, ledger the ruling" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
|
|
90
104
|
"Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"];
|
|
91
105
|
"Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)";
|
|
92
106
|
"Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?";
|
|
@@ -95,7 +109,7 @@ digraph process {
|
|
|
95
109
|
"R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"];
|
|
96
110
|
"R = 5?" -> "Adjudicate each open finding" [label="yes - breaker trips"];
|
|
97
111
|
"Adjudicate each open finding" -> "Any load-bearing finding?";
|
|
98
|
-
"Any load-bearing finding?" -> "
|
|
112
|
+
"Any load-bearing finding?" -> "Rule and continue; stop only if every path forward is a guess" [label="yes"];
|
|
99
113
|
"Any load-bearing finding?" -> "Park findings in ledger with rulings" [label="no"];
|
|
100
114
|
"Park findings in ledger with rulings" -> "Append completion to ledger, mark todo complete";
|
|
101
115
|
"Append completion to ledger, mark todo complete" -> "More tasks remain?";
|
|
@@ -140,19 +154,32 @@ a ledger file, not only in todos.
|
|
|
140
154
|
that happens, recover from `git log`.
|
|
141
155
|
|
|
142
156
|
Read the plan once, note its context and Global Constraints, and create a
|
|
143
|
-
todo per task.
|
|
157
|
+
todo per task. If the plan names a Spec, read that too: the spec is the
|
|
158
|
+
authority the plan argues from, and conflicts inside the plan resolve
|
|
159
|
+
against it. A plan with no reachable spec gets a ledger note saying so —
|
|
160
|
+
rulings made without one are provisional.
|
|
144
161
|
|
|
145
|
-
Before dispatching Task 1, scan the plan once for conflicts
|
|
162
|
+
Before dispatching Task 1, scan the plan once for conflicts, writing down
|
|
163
|
+
what you checked as you check it:
|
|
146
164
|
|
|
147
165
|
- tasks that contradict each other or the plan's Global Constraints
|
|
148
166
|
- anything the plan explicitly mandates that the review rubric treats as a
|
|
149
167
|
defect (a test that asserts nothing, verbatim duplication of a logic block)
|
|
150
168
|
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
169
|
+
The scan's output is a table, not a verdict. One row for every pair of tasks
|
|
170
|
+
that share a file or an interface: the two tasks, what one produces against
|
|
171
|
+
what the other consumes, and what you found. One row for every task: whether
|
|
172
|
+
its own text agrees with itself — the tests it specifies against the code it
|
|
173
|
+
specifies, the files it creates against the files it later touches. "The scan
|
|
174
|
+
is clean" without those rows is not a scan you ran.
|
|
175
|
+
|
|
176
|
+
Write the table to the ledger. Rule on everything you find before execution
|
|
177
|
+
begins — each finding against the plan text that mandates it — and record
|
|
178
|
+
each ruling in the ledger. If the scan is clean, proceed without comment.
|
|
179
|
+
Rule on each conflict it surfaces — the spec is the binding authority, the
|
|
180
|
+
plan is its argument — record the ruling beside its row, and dispatch
|
|
181
|
+
Task 1. The review loop remains the net for conflicts that only emerge from
|
|
182
|
+
implementation.
|
|
156
183
|
|
|
157
184
|
## Model Selection
|
|
158
185
|
|
|
@@ -193,17 +220,37 @@ that implementer. Single-file mechanical fixes also take the cheapest tier.
|
|
|
193
220
|
|
|
194
221
|
## The Task Loop
|
|
195
222
|
|
|
223
|
+
**Batch small same-shape work.** When the plan lists several tasks that are
|
|
224
|
+
each a small, independent edit of the same kind — the same one-line fix,
|
|
225
|
+
constant change, or field addition repeated across files — do not dispatch
|
|
226
|
+
one subagent per task. Compose ONE dispatch brief listing every file and
|
|
227
|
+
its change, send the whole batch to a single subagent, and review its diff
|
|
228
|
+
as one unit. Reserve one-dispatch-per-task for work that needs its own
|
|
229
|
+
judgment, its own tests, or its own review surface.
|
|
230
|
+
|
|
196
231
|
Everything you paste into a dispatch prompt — and everything a subagent
|
|
197
232
|
prints back — stays resident in your context for the rest of the session
|
|
198
233
|
and is re-read on every later turn. Hand artifacts over as files.
|
|
199
234
|
|
|
235
|
+
**Waiting on dispatched subagents:** never poll a wait interface with
|
|
236
|
+
short timeouts, and never sit in one silent, open-ended wait either.
|
|
237
|
+
While you have local work — ledger updates, packaging the next review,
|
|
238
|
+
reading reports — keep working; child results arrive on their own.
|
|
239
|
+
When you are genuinely idle, wait in bounded stretches (five to ten
|
|
240
|
+
minutes, where your platform allows), and between stretches post one
|
|
241
|
+
line of status and reconcile your live children: list them, and chase
|
|
242
|
+
any that finished without reporting. A bounded stretch keeps nearly
|
|
243
|
+
all of a long wait's efficiency while guaranteeing a stuck or lost
|
|
244
|
+
child is noticed within minutes, not at the end of the session.
|
|
245
|
+
|
|
200
246
|
### 1. Dispatch the implementer
|
|
201
247
|
|
|
202
248
|
Record BASE (`git rev-parse HEAD`) before dispatching — the review package
|
|
203
249
|
and fix-round diffs need it.
|
|
204
250
|
|
|
205
251
|
- **Task brief:** before dispatching an implementer, run this skill's
|
|
206
|
-
`scripts/task-brief PLAN_FILE N` (or `scripts/task-brief.ps1 PLAN_FILE N` on
|
|
252
|
+
`scripts/task-brief PLAN_FILE N` (or `scripts/task-brief.ps1 PLAN_FILE N` on
|
|
253
|
+
Windows PowerShell) — it extracts the task's full text to a
|
|
207
254
|
uniquely named file and prints the path. Compose the dispatch so the
|
|
208
255
|
brief stays the single source of
|
|
209
256
|
requirements. Your dispatch should contain: (1) one line on where this
|
|
@@ -223,6 +270,12 @@ and fix-round diffs need it.
|
|
|
223
270
|
later dispatches — a real session's dispatch hit 42k chars of which 99%
|
|
224
271
|
was pasted history. A fresh subagent needs its task, the interfaces it
|
|
225
272
|
touches, and the global constraints. Nothing else.
|
|
273
|
+
- The dispatch carries the no-subagents contract (it is in the
|
|
274
|
+
implementer template): the implementer never dispatches subagents —
|
|
275
|
+
not helpers, and never a reviewer. Review arrives from you, after the
|
|
276
|
+
report. In real sessions, every reviewer a worker spawned duplicated
|
|
277
|
+
the task review the controller dispatched anyway — a full extra
|
|
278
|
+
review seat per task.
|
|
226
279
|
- If an earlier task parked a finding in the area this task touches, carry
|
|
227
280
|
a pointer to that ledger entry in the dispatch.
|
|
228
281
|
- Record the implementer's agent identity from the dispatch result —
|
|
@@ -245,7 +298,7 @@ Implementer subagents report one of four statuses. Handle each appropriately:
|
|
|
245
298
|
1. If it's a context problem, provide more context and re-dispatch with the same model
|
|
246
299
|
2. If the task requires more reasoning, re-dispatch with a more capable model
|
|
247
300
|
3. If the task is too large, break it into smaller pieces
|
|
248
|
-
4. If the plan itself is wrong,
|
|
301
|
+
4. If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch
|
|
249
302
|
|
|
250
303
|
**Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.
|
|
251
304
|
|
|
@@ -262,7 +315,9 @@ required. Implementer self-review never replaces the task review; both are
|
|
|
262
315
|
needed.
|
|
263
316
|
|
|
264
317
|
- Hand the reviewer its diff as a file: run this skill's
|
|
265
|
-
`scripts/review-package PLAN_FILE BASE HEAD` (or
|
|
318
|
+
`scripts/review-package PLAN_FILE BASE HEAD` (or
|
|
319
|
+
`scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell) and
|
|
320
|
+
pass the reviewer the file path
|
|
266
321
|
it prints (or, without bash: `git log --oneline`, `git diff --stat`,
|
|
267
322
|
and `git diff -U10` for the range, redirected to one uniquely named
|
|
268
323
|
file). The output never enters your own context, and the reviewer sees
|
|
@@ -312,10 +367,11 @@ Before the loop starts, two routes leave it immediately:
|
|
|
312
367
|
before merge. A roll-up nobody reads is a silent discard. Minor findings
|
|
313
368
|
never enter the loop.
|
|
314
369
|
- A finding labeled plan-mandated — or any finding that conflicts with
|
|
315
|
-
what the plan's text requires — is
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
dispatch a fix that contradicts the plan
|
|
370
|
+
what the plan's text requires — is yours to rule on: weigh the finding
|
|
371
|
+
against the plan text, decide with the spec as the binding authority, and
|
|
372
|
+
ledger the ruling before you act on it. Do not dismiss the finding because
|
|
373
|
+
the plan mandates it, and do not dispatch a fix that contradicts the plan
|
|
374
|
+
without a recorded ruling.
|
|
319
375
|
Everything else enters the loop. A fix round is one fix dispatch plus one
|
|
320
376
|
scoped re-review. Five rounds maximum per task:
|
|
321
377
|
|
|
@@ -361,15 +417,16 @@ dispatching. Adjudicate each open finding yourself — you hold the plan and
|
|
|
361
417
|
the cross-task context the reviewer lacks:
|
|
362
418
|
|
|
363
419
|
- **The reviewer is wrong, or the point is contestable:** park it —
|
|
364
|
-
`Task <N>: parked — <finding> —
|
|
420
|
+
`Task <N>: parked — <finding> — Ruling: <why the code stands>`. The final
|
|
365
421
|
review sees both sides.
|
|
366
422
|
- **Real, but nothing downstream builds on it:** park it the same way, with
|
|
367
423
|
a ruling that says it's real and deferred.
|
|
368
424
|
- **Real and load-bearing** — a later task builds on it, or it reveals a
|
|
369
|
-
plan defect:
|
|
370
|
-
|
|
371
|
-
the
|
|
372
|
-
|
|
425
|
+
plan defect: rule on the smallest change that unblocks the dependent work,
|
|
426
|
+
ledger it as `Task <N>: Ruling: <finding> — <what you decided and why>`,
|
|
427
|
+
and carry it into the next task's dispatch. Parking a structural failure
|
|
428
|
+
silently lets every dependent task build on it. Stop only when the defect
|
|
429
|
+
leaves every path forward a guess.
|
|
373
430
|
|
|
374
431
|
Adjudicate only at the cap. Adjudicating earlier to end a loop is
|
|
375
432
|
pre-judging with a different name. Every adjudication is a ledger entry —
|
|
@@ -392,8 +449,10 @@ parked-with-ruling at the cap.
|
|
|
392
449
|
## Final Review
|
|
393
450
|
|
|
394
451
|
The final whole-branch review gets a package too: run
|
|
395
|
-
`scripts/review-package PLAN_FILE MERGE_BASE HEAD` (or
|
|
396
|
-
|
|
452
|
+
`scripts/review-package PLAN_FILE MERGE_BASE HEAD` (or
|
|
453
|
+
`scripts/review-package.ps1 PLAN_FILE MERGE_BASE HEAD` on Windows PowerShell;
|
|
454
|
+
MERGE_BASE = the commit the branch started from, e.g. `git merge-base main HEAD`)
|
|
455
|
+
and include the
|
|
397
456
|
printed path in the final review dispatch, so the final reviewer reads
|
|
398
457
|
one file instead of re-deriving the branch diff with git commands. Dispatch
|
|
399
458
|
on the most capable available model (see Model Selection), using
|
|
@@ -407,15 +466,27 @@ with the complete findings list — not one fixer per finding.
|
|
|
407
466
|
Per-finding fixers each rebuild context and re-run suites; a real
|
|
408
467
|
session's final-review fix wave cost more than all its tasks combined.
|
|
409
468
|
Then run exactly one scoped re-review of the fix wave
|
|
410
|
-
(`scripts/review-package PLAN_FILE FIX_BASE HEAD
|
|
469
|
+
(`scripts/review-package PLAN_FILE FIX_BASE HEAD`, or
|
|
470
|
+
`scripts/review-package.ps1 PLAN_FILE FIX_BASE HEAD` on Windows PowerShell,
|
|
471
|
+
over the fix range,
|
|
411
472
|
[re-review-prompt.md](re-review-prompt.md)).
|
|
412
473
|
Adjudicate any residual findings as in the task loop's breaker: park with
|
|
413
|
-
rulings, or
|
|
474
|
+
rulings, or rule on the load-bearing ones and ledger what you decided. Only
|
|
475
|
+
the four classes above stop you here. There is no second fix wave —
|
|
414
476
|
residual load-bearing findings surface to your human partner when
|
|
415
477
|
finishing-a-development-branch presents the options.
|
|
416
478
|
|
|
417
479
|
## Finish
|
|
418
480
|
|
|
481
|
+
Before you delete anything, collect every ledger line containing `Ruling:` —
|
|
482
|
+
preflight rulings, parked findings, breaker adjudications, all of them — into
|
|
483
|
+
your final message under "Rulings I made", in the order you made them, each
|
|
484
|
+
with what it costs if wrong. The list is exhaustive: if the ledger holds a
|
|
485
|
+
ruling, the list holds it. That list is the only place the decisions you
|
|
486
|
+
took on your human partner's behalf reach them — they read it and rework
|
|
487
|
+
whatever you got wrong. A ruling that dies with the workspace was a decision
|
|
488
|
+
made in secret.
|
|
489
|
+
|
|
419
490
|
When the final whole-branch review is clean and its fixes are merged,
|
|
420
491
|
delete this plan's workspace (`rm -rf <workspace>`) — the git history is
|
|
421
492
|
the record now. Sibling directories belong to other plans; leave them
|
|
@@ -435,6 +506,7 @@ Use superpowers:finishing-a-development-branch.
|
|
|
435
506
|
| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
|
|
436
507
|
| "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
|
|
437
508
|
| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
|
|
509
|
+
| "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
|
|
438
510
|
|
|
439
511
|
## Example Workflow
|
|
440
512
|
|
|
@@ -47,6 +47,18 @@ Subagent (general-purpose):
|
|
|
47
47
|
While iterating, run the focused test for what you're changing; run the
|
|
48
48
|
full suite once before committing, not after every edit.
|
|
49
49
|
|
|
50
|
+
## You Do Not Dispatch Subagents
|
|
51
|
+
|
|
52
|
+
Do all of this task's work yourself. Never spawn a subagent to
|
|
53
|
+
implement part of the task, and above all never spawn a reviewer to
|
|
54
|
+
check your work. Self-review (below) means reading your own diff.
|
|
55
|
+
Review is the controller's job: after you report, it dispatches a
|
|
56
|
+
fresh reviewer against your diff. A reviewer you spawn duplicates
|
|
57
|
+
that review at full cost, and its approval counts for nothing in
|
|
58
|
+
the process. If you catch yourself thinking "an independent review
|
|
59
|
+
would strengthen my report" — that review is already scheduled.
|
|
60
|
+
Report instead.
|
|
61
|
+
|
|
50
62
|
## Code Organization
|
|
51
63
|
|
|
52
64
|
You reason best about code you can hold in context at once, and your edits are more
|
|
@@ -43,6 +43,15 @@ Subagent (general-purpose):
|
|
|
43
43
|
Your review is read-only on this checkout. Do not mutate the working
|
|
44
44
|
tree, the index, HEAD, or branch state in any way.
|
|
45
45
|
|
|
46
|
+
## You Do Not Dispatch Subagents
|
|
47
|
+
|
|
48
|
+
Do all of this review yourself. Never spawn a subagent to review part
|
|
49
|
+
of the diff, and never spawn another reviewer for a second opinion.
|
|
50
|
+
This process already provides every review seat the work gets; a
|
|
51
|
+
reviewer you spawn duplicates one of them at full cost, and its
|
|
52
|
+
verdict counts for nothing. If the diff feels too large for one
|
|
53
|
+
pass, review it in passes yourself and say so in your report.
|
|
54
|
+
|
|
46
55
|
## Scope
|
|
47
56
|
|
|
48
57
|
Your scope is the findings list and the fix diff. Verdict every finding.
|
|
@@ -22,7 +22,7 @@ if (-not (Test-Path -LiteralPath $plan -PathType Leaf)) {
|
|
|
22
22
|
exit 2
|
|
23
23
|
}
|
|
24
24
|
|
|
25
|
-
$slug = [System.IO.Path]::GetFileName($plan) -
|
|
25
|
+
$slug = [System.IO.Path]::GetFileName($plan) -creplace '\.md$', ''
|
|
26
26
|
if ([string]::IsNullOrEmpty($slug) -or $slug -eq "." -or $slug -eq "..") {
|
|
27
27
|
[Console]::Error.WriteLine("cannot derive a workspace name from: $plan")
|
|
28
28
|
exit 2
|
|
@@ -52,6 +52,15 @@ Subagent (general-purpose):
|
|
|
52
52
|
Your review is read-only on this checkout. Do not mutate the working
|
|
53
53
|
tree, the index, HEAD, or branch state in any way.
|
|
54
54
|
|
|
55
|
+
## You Do Not Dispatch Subagents
|
|
56
|
+
|
|
57
|
+
Do all of this review yourself. Never spawn a subagent to review part
|
|
58
|
+
of the diff, and never spawn another reviewer for a second opinion.
|
|
59
|
+
This process already provides every review seat the work gets; a
|
|
60
|
+
reviewer you spawn duplicates one of them at full cost, and its
|
|
61
|
+
verdict counts for nothing. If the diff feels too large for one
|
|
62
|
+
pass, review it in passes yourself and say so in your report.
|
|
63
|
+
|
|
55
64
|
## Do Not Trust the Report
|
|
56
65
|
|
|
57
66
|
Treat the implementer's report as unverified claims about the code. It
|
|
@@ -75,6 +84,13 @@ Subagent (general-purpose):
|
|
|
75
84
|
Warnings or other noise in the implementer's reported test output are
|
|
76
85
|
findings — test output should be pristine.
|
|
77
86
|
|
|
87
|
+
Evidence you cannot see is not evidence that doesn't exist. If the
|
|
88
|
+
report or its test evidence looks truncated, or you cannot locate the
|
|
89
|
+
results it claims, re-read the file at its stated path — and if it is
|
|
90
|
+
genuinely missing or garbled, report that as a gap for the controller.
|
|
91
|
+
Re-running the suite to regenerate what you failed to read is not
|
|
92
|
+
verification; illegibility of the evidence is not invalidation of it.
|
|
93
|
+
|
|
78
94
|
## Part 1: Spec Compliance
|
|
79
95
|
|
|
80
96
|
Compare the diff against What Was Requested:
|
|
@@ -86,6 +102,12 @@ Subagent (general-purpose):
|
|
|
86
102
|
- **Misunderstood:** right feature built the wrong way, wrong problem
|
|
87
103
|
solved
|
|
88
104
|
|
|
105
|
+
If the brief lists several files each with its own change (a batched
|
|
106
|
+
dispatch), check the diff against that list file by file: every listed
|
|
107
|
+
file must have its corresponding hunk. A listed file the diff never
|
|
108
|
+
touches is a Missing finding, no matter how clean the rest of the
|
|
109
|
+
batch looks.
|
|
110
|
+
|
|
89
111
|
If a requirement cannot be verified from this diff alone (it lives in
|
|
90
112
|
unchanged code or spans tasks), report it as a ⚠️ item instead of
|
|
91
113
|
broadening your search.
|
|
@@ -167,8 +189,7 @@ Subagent (general-purpose):
|
|
|
167
189
|
|
|
168
190
|
**Placeholders:**
|
|
169
191
|
- `[MODEL]` — REQUIRED: reviewer model per SKILL.md Model Selection
|
|
170
|
-
- `[BRIEF_FILE]` — REQUIRED: the task brief file (`scripts/task-brief PLAN N`,
|
|
171
|
-
or `scripts/task-brief.ps1 PLAN N` on Windows PowerShell,
|
|
192
|
+
- `[BRIEF_FILE]` — REQUIRED: the task brief file (`scripts/task-brief PLAN N`, or `scripts/task-brief.ps1 PLAN N` on Windows PowerShell,
|
|
172
193
|
prints the path; same file the implementer worked from)
|
|
173
194
|
- `[GLOBAL_CONSTRAINTS]` — the binding requirements copied verbatim from
|
|
174
195
|
the plan's Global Constraints section or the spec: exact values, formats,
|
|
@@ -179,9 +200,8 @@ Subagent (general-purpose):
|
|
|
179
200
|
- `[BASE_SHA]` — commit before this task
|
|
180
201
|
- `[HEAD_SHA]` — current commit
|
|
181
202
|
- `[DIFF_FILE]` — REQUIRED: the path the controller wrote the review
|
|
182
|
-
package to (`scripts/review-package PLAN_FILE BASE HEAD`, or
|
|
183
|
-
|
|
184
|
-
prints the unique path it wrote; the package never enters the controller's context)
|
|
203
|
+
package to (`scripts/review-package PLAN_FILE BASE HEAD`, or `scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell, prints the unique
|
|
204
|
+
path it wrote; the package never enters the controller's context)
|
|
185
205
|
|
|
186
206
|
**Reviewer returns:** Spec Compliance verdict (✅/❌/⚠️), Strengths, Issues
|
|
187
207
|
(Critical/Important/Minor), Task quality verdict
|
|
@@ -56,6 +56,7 @@ If your harness appears here, read its reference file for special instructions:
|
|
|
56
56
|
- Codex: `references/codex-tools.md`
|
|
57
57
|
- Pi: `references/pi-tools.md`
|
|
58
58
|
- Antigravity: `references/antigravity-tools.md`
|
|
59
|
+
- Hermes Agent: `references/hermes-tools.md`
|
|
59
60
|
|
|
60
61
|
## User Instructions
|
|
61
62
|
|
|
@@ -7,7 +7,76 @@ Add to your Codex config (`~/.codex/config.toml`):
|
|
|
7
7
|
multi_agent = true
|
|
8
8
|
```
|
|
9
9
|
|
|
10
|
-
This enables
|
|
10
|
+
This enables the multi-agent tools that skills like
|
|
11
|
+
`dispatching-parallel-agents` and `subagent-driven-development` use.
|
|
12
|
+
Which tools you get depends on the multi-agent version your model
|
|
13
|
+
preset selects (current presets run V2; older ones run V1). Trust your
|
|
14
|
+
actual tool list over any table — including this one — when they
|
|
15
|
+
disagree.
|
|
16
|
+
|
|
17
|
+
- **Spawning:** give children a clean context with
|
|
18
|
+
`spawn_agent {fork_turns: "none"}`; the default `"all"` copies your
|
|
19
|
+
entire transcript into the child. On Codex 0.145+, role files under
|
|
20
|
+
`~/.codex/agents/` attach to isolated forks via `agent_type`.
|
|
21
|
+
Full-history forks accept `model` and `reasoning_effort` overrides
|
|
22
|
+
(only `agent_type` is refused there) — isolated forks are the SDD
|
|
23
|
+
default for context hygiene, not because overrides require them.
|
|
24
|
+
- **Fix rounds:** resume the implementer with `followup_task` — it
|
|
25
|
+
delivers your message, triggers a turn, and transparently reloads a
|
|
26
|
+
child the harness evicted. Never dispatch a fresh implementer on the
|
|
27
|
+
theory that a spawned agent cannot be messaged again; on V2 it
|
|
28
|
+
always can.
|
|
29
|
+
- **Lifecycle:** V2 has no `close_agent`. Finished children are
|
|
30
|
+
evicted automatically when slots are needed; leaving them unclosed
|
|
31
|
+
costs nothing. Only V1 sessions have `close_agent` — there, close
|
|
32
|
+
reviewers when their review returns, and close each implementer
|
|
33
|
+
after its task's review passes.
|
|
34
|
+
- **Model names:** never copy a model name from a skill, table, or old
|
|
35
|
+
session into `spawn_agent` without checking it against your current
|
|
36
|
+
spawn allowlist — V2 accepts only V2-capable presets and hard-errors
|
|
37
|
+
on the rest.
|
|
38
|
+
|
|
39
|
+
## Waiting on children
|
|
40
|
+
|
|
41
|
+
`wait_agent` is an event subscription, not a poll: a long wait wakes
|
|
42
|
+
the moment a child produces mailbox activity, with the same latency as
|
|
43
|
+
a short one. Short-timeout polling buys nothing and costs a tool call —
|
|
44
|
+
and a context rebill — per poll. In measured sessions, roughly
|
|
45
|
+
two-thirds of all wait calls were short polls that timed out.
|
|
46
|
+
|
|
47
|
+
- While you still have local work, do not wait at all. A completed
|
|
48
|
+
child's final answer is pushed into your mailbox and arrives with
|
|
49
|
+
your next turn.
|
|
50
|
+
- When you are genuinely idle with children outstanding, wait in
|
|
51
|
+
bounded stretches: `wait_agent` with `timeout_ms` 300000-600000
|
|
52
|
+
(5-10 minutes). After each stretch — wake or timeout — post one
|
|
53
|
+
status line, run `list_agents`, and chase any child that finished
|
|
54
|
+
without reporting. Never stack polls shorter than five minutes; the
|
|
55
|
+
event subscription wakes a bounded stretch just as fast as a short
|
|
56
|
+
one.
|
|
57
|
+
- Completion mail cannot wake an idle controller (it is delivered
|
|
58
|
+
without triggering a turn); covering that idle window is
|
|
59
|
+
`wait_agent`'s only job. A stretch that times out with no activity
|
|
60
|
+
is your cue to reconcile, not to shorten the next stretch.
|
|
61
|
+
|
|
62
|
+
## Model routing on spawns
|
|
63
|
+
|
|
64
|
+
Every `spawn_agent` you issue — including when you are yourself a
|
|
65
|
+
spawned child running a fan-out — sets `model` AND `reasoning_effort`
|
|
66
|
+
explicitly, per the Model Selection rules of the skill you are
|
|
67
|
+
executing. Setting `model` alone is a trap: the child's effort
|
|
68
|
+
silently resets to that model's default, not to yours.
|
|
69
|
+
|
|
70
|
+
Ask your human partner to add a machine-level backstop to
|
|
71
|
+
`~/.codex/config.toml` so any spawn that slips through still routes to
|
|
72
|
+
a deliberate tier instead of silently inheriting the session's most
|
|
73
|
+
expensive model:
|
|
74
|
+
|
|
75
|
+
```toml
|
|
76
|
+
[agents]
|
|
77
|
+
default_subagent_model = "<a mid-tier model from your spawn allowlist>"
|
|
78
|
+
default_subagent_reasoning_effort = "medium"
|
|
79
|
+
```
|
|
11
80
|
|
|
12
81
|
## Environment Detection
|
|
13
82
|
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# Hermes Agent Tool Mapping
|
|
2
|
+
|
|
3
|
+
Skills speak in actions ("dispatch a subagent", "create a todo", "read a file"). On Hermes Agent these resolve to the tools below.
|
|
4
|
+
|
|
5
|
+
## Tools
|
|
6
|
+
|
|
7
|
+
| Action skills request | Hermes tool |
|
|
8
|
+
|---|---|
|
|
9
|
+
| Read a file | `read_file` |
|
|
10
|
+
| Create a new file | `write_file` |
|
|
11
|
+
| Edit a file (targeted patch) | `patch` |
|
|
12
|
+
| Run a shell command | `terminal` |
|
|
13
|
+
| Search file contents | `search_files` |
|
|
14
|
+
| Find files by name | `terminal` with `find` |
|
|
15
|
+
| Fetch a URL / read a webpage | `web_extract(urls=[...])` |
|
|
16
|
+
| Search the web | `web_search(query=...)` |
|
|
17
|
+
| Dispatch a subagent | `delegate_task(goal=..., context=..., toolsets=[...], role="leaf")` |
|
|
18
|
+
| Task tracking | `todo` tool |
|
|
19
|
+
| Invoke a skill | `skill_view("skill-name")` |
|
|
20
|
+
|
|
21
|
+
## Instructions file
|
|
22
|
+
|
|
23
|
+
When a skill mentions "your instructions file," on Hermes Agent this is **`AGENTS.md`** in the project directory, or **`SOUL.md`** globally at `~/.hermes/SOUL.md`.
|
|
24
|
+
|
|
25
|
+
## Invoking a skill
|
|
26
|
+
|
|
27
|
+
Hermes Agent has a `skills` toolset with `skill_view` and `skills_list` tools.
|
|
28
|
+
To invoke a superpowers skill, use:
|
|
29
|
+
|
|
30
|
+
```
|
|
31
|
+
skill_view("brainstorming")
|
|
32
|
+
skill_view("test-driven-development")
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
If `skill_view` cannot find a superpowers skill (it may not appear in the catalog
|
|
36
|
+
until the plugin fully registers it), fall back to reading the SKILL.md directly:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
read_file(path="~/.hermes/plugins/superpowers/skills/<skill-name>/SKILL.md")
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
This fallback is the same mechanism used by other harnesses without native skill loading.
|
|
43
|
+
|
|
44
|
+
## Subagent dispatch
|
|
45
|
+
|
|
46
|
+
Use `delegate_task` to spawn isolated subagents for parallel or sequential workstreams:
|
|
47
|
+
|
|
48
|
+
```
|
|
49
|
+
delegate_task(goal="...", context="...", toolsets=[...], role="leaf")
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
If `delegate_task` is unavailable, do the work inline rather than inventing tool calls.
|
|
53
|
+
|
|
54
|
+
## Task tracking
|
|
55
|
+
|
|
56
|
+
Use the `todo` tool for task tracking within a session. For multi-agent task boards, use `hermes kanban` CLI if available. Treat older `TodoWrite` references as the task-tracking action.
|
|
@@ -66,6 +66,9 @@ independently testable deliverable.
|
|
|
66
66
|
|
|
67
67
|
**Tech Stack:** [Key technologies/libraries]
|
|
68
68
|
|
|
69
|
+
**Spec:** [path to the spec/design doc this plan implements — the plan
|
|
70
|
+
argues from the spec, so the spec travels with it; executors read both]
|
|
71
|
+
|
|
69
72
|
## Global Constraints
|
|
70
73
|
|
|
71
74
|
[The spec's project-wide requirements — version floors, dependency limits,
|
|
@@ -246,7 +246,7 @@ SKILL.md serves as an overview that points agents to detailed materials as neede
|
|
|
246
246
|
|
|
247
247
|
A basic Skill starts with just a SKILL.md file containing metadata and instructions:
|
|
248
248
|
|
|
249
|
-
<img src="https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=87782ff239b297d9a9e8e1b72ed72db9" alt="Simple SKILL.md file showing YAML frontmatter and markdown body" data-og-width="2048" width="2048" data-og-height="1153" height="1153" data-path="images/agent-skills-simple-file.png" data-optimize="true" data-opv="3" srcset="https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=280&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=c61cc33b6f5855809907f7fda94cd80e 280w, https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=560&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=
|
|
249
|
+
<img src="https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=87782ff239b297d9a9e8e1b72ed72db9" alt="Simple SKILL.md file showing YAML frontmatter and markdown body" data-og-width="2048" width="2048" data-og-height="1153" height="1153" data-path="images/agent-skills-simple-file.png" data-optimize="true" data-opv="3" srcset="https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=280&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=c61cc33b6f5855809907f7fda94cd80e 280w, https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=560&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=90d2c0c1c76b36e8d485f49e0810dbfd 560w, https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=840&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=ad17d231ac7b0bea7e5b4d58fb4aeabb 840w, https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=1100&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=f5d0a7a3c668435bb0aee9a3a8f8c329 1100w, https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=1650&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=0e927c1af9de5799cfe557d12249f6e6 1650w, https://mintcdn.com/anthropic-claude-docs/4Bny2bjzuGBK7o00/images/agent-skills-simple-file.png?w=2500&fit=max&auto=format&n=4Bny2bjzuGBK7o00&q=85&s=46bbb1a51dd4c8202a470ac8c80a893d 2500w" />
|
|
250
250
|
|
|
251
251
|
As your Skill grows, you can bundle additional content that agents load only when needed:
|
|
252
252
|
|
|
@@ -107,9 +107,10 @@ function main() {
|
|
|
107
107
|
process.exit(1);
|
|
108
108
|
}
|
|
109
109
|
|
|
110
|
-
// Check if dot is available
|
|
110
|
+
// Check if dot is available. Run the binary directly rather than probing
|
|
111
|
+
// with `which`, which is not a command on Windows.
|
|
111
112
|
try {
|
|
112
|
-
execSync('
|
|
113
|
+
execSync('dot -V', { stdio: 'ignore', encoding: 'utf-8' });
|
|
113
114
|
} catch {
|
|
114
115
|
console.error('Error: graphviz (dot) not found. Install with:');
|
|
115
116
|
console.error(' brew install graphviz # macOS');
|