agentflowctl 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -3,7 +3,7 @@
3
3
  讓 Claude Code、Codex、Gemini CLI(或任何 agent CLI)在同一條流程裡輪流寫規格、寫計畫、寫測試、寫實作、互相審查、互相修正,一路做到開 PR。
4
4
 
5
5
  ```
6
- 需求 → spec → plan ⇄ plan_review ⇄ plan_fix →(僵持時)仲裁 → implement(測試 A → 實作 B)→ verify ⇄ fix → review ⇄ fix → pr
6
+ 需求 → spec → plan ⇄ plan_review ⇄ plan_fix →(僵持時)仲裁 → implement(每個任務:測試 A → 實作 B → 任務審查 ⇄ 修正 → 任務驗證 ⇄ 修正)→ verify ⇄ fix → review ⇄ fix → pr
7
7
  ```
8
8
 
9
9
  從需求到 PR 全程由程式推進。每個步驟都是一次獨立的 agent 執行,用 `.flow/` 裡的檔案交接。是否通過一律由程式檢查:跑測試、比對 git diff、驗證 JSON。兩家模型就能完整運作;有第三家時,仲裁會交給沒參與討論的那一家。
@@ -17,7 +17,7 @@ claude # 完成登入
17
17
  codex # 完成登入
18
18
  ```
19
19
 
20
- 沒有內建的 agent,只會使用 `flow.config.json` 的 `agents` 裡設定的。用 `agent setup` 互動設定:它會偵測本機的 `claude`、`codex`、`gemini`,逐一詢問要不要加入、名稱與 model,再設定參與的 agent;確認後才一次寫入,最後自動跑一次 `doctor`:
20
+ 沒有內建的 agent,只會使用 `flow.config.json` 的 `agents` 裡設定的。用 `agent setup` 互動設定:它會先列出已設定的 agent 與 CLI 是否可以執行,再偵測本機裝了哪些支援的 CLI(目前是 `claude`、`codex`、`gemini`),只針對已安裝的逐一詢問要不要加入、名稱與 model,再設定參與的 agent;沒偵測到的只列出、不詢問。確認後才一次寫入,最後自動跑一次 `doctor`:
21
21
 
22
22
  ```bash
23
23
  npx agentflowctl agent setup
@@ -120,7 +120,7 @@ agentflowctl clean --all # 清掉所有已結束的 run 與中斷留
120
120
  | `--manual-plan` | 計畫通過審查後進入 `awaiting_approval`,等 `approve` 才開始實作 |
121
121
  | `-v` / `--verbose` | 執行時印出 agent 的文字、工具呼叫與專案指令;`run`、`resume`、`approve` 都適用,也可設 `AGENTFLOWCTL_VERBOSE=1` |
122
122
 
123
- `status` 會列出任務。進行中的任務會標出正在寫測試還是正在寫實作。
123
+ `status` 會列出任務。進行中的任務會標出目前的步驟:🧪 寫測試、🛠️ 寫實作、👀 任務審查、🔍 任務驗證、🩹 任務修正。
124
124
 
125
125
  ### 清除 worktree
126
126
 
@@ -268,7 +268,7 @@ run 因 Ctrl-C、失敗、額度暫停或等待核准而停下時,終端機會
268
268
  | `defaultModels` | `{}` | 依 `claude`、`codex`、`gemini` adapter 指定全域預設 model;agent 的 `model` 優先,兩者都沒設時使用各 CLI 的預設。`command` adapter 不套用 |
269
269
  | `fixStrategy` | `ring` | `ring`:審查意見隨機交給審查者以外的一家;`author`:交回最後作者 |
270
270
  | `tddSplit` | `true` | 測試與實作是否分開 |
271
- | `reviewQuorum` | `1` | 程式碼需要幾位不同審查者都 `approve` |
271
+ | `reviewQuorum` | `1` | 程式碼需要幾位不同審查者都 `approve`(任務審查與最後的程式碼審查都適用) |
272
272
  | `planReviewQuorum` | `1` | 計畫需要幾位不同審查者都 `approve` |
273
273
  | `planArbiter` | `true` | 計畫審查僵持時交付仲裁。關掉之後,僵持會直接讓 run 失敗 |
274
274
  | `tieBreak` | `proceed` | 兩家仲裁意見分歧時:`proceed` 繼續並記錄爭議;`stop` 停下 |
@@ -300,7 +300,7 @@ verify 失敗(型別、lint、建置)一律交回最後作者。審查意見
300
300
  `agents` 與 `cycle` 也可以用 `agent` 指令修改,不必手動編輯 JSON。每次寫入前都會先驗證整份設定:
301
301
 
302
302
  ```bash
303
- agentflowctl agent setup # 互動設定 claude、codex、gemini 與參與的 agent
303
+ agentflowctl agent setup # 偵測已安裝的 agent CLI,互動設定與參與的 agent
304
304
  agentflowctl agent list # 設定的 agent、是否已安裝、是否參與
305
305
  agentflowctl agent add claude-strong --adapter claude --model opus
306
306
  agentflowctl agent add aider --adapter command -- aider --yes-always --message {prompt}
@@ -315,7 +315,7 @@ agentflowctl agent cycle claude-strong,codex,gemini # 不帶參數時顯
315
315
  - `set --adapter` 換 adapter 時,會清掉舊 adapter 的 `model`、`extraArgs`、`command`,這次有重新指定的除外。
316
316
  - `remove` 會一併從 `cycle` 移除。`cycle` 變空就刪除這個欄位,改回從 `agents` 自動偵測。
317
317
  - `--extra-arg` 可以重複指定,會整個取代原本的 `extraArgs`。參數以 `-` 開頭時,寫成 `--extra-arg=--sandbox`。
318
- - `setup` 遇到已存在的名稱會先問要不要覆寫;不覆寫時保留原設定,但仍會參與。在非互動式環境(CI、管線)裡請改用 `agent add`。`command` adapter 要自己寫指令,不在 `setup` 裡。
318
+ - `setup` 只詢問偵測到已安裝的 CLI,一個都沒有就不變更設定。遇到已存在的名稱會先問要不要覆寫;不覆寫時保留原設定,但仍會參與。在非互動式環境(CI、管線)裡請改用 `agent add`。`command` adapter 要自己寫指令,不在 `setup` 裡。
319
319
 
320
320
  已建立的 run 會沿用建立時參與的 agent,不受這些修改影響。
321
321
 
@@ -327,6 +327,8 @@ agentflowctl agent cycle claude-strong,codex,gemini # 不帶參數時顯
327
327
 
328
328
  **跨模型審查。** 審查看需求覆蓋、驗收條件能不能測且一條只寫一個行為、任務是否只做一件事(最多兩件)與技術方向。審查者只能寫意見。若改了規格或計畫,檔案會被還原。修改者要在 `plan.md` 的「審查回應」逐條回覆;不同意要寫理由。
329
329
 
330
+ 計畫審查先核對需求與四份計畫交接檔,再查閱任務說明中要修改的既有檔案,有疑慮時才擴大範圍;審查紀錄只列會影響實作的問題,不逐條列已通過項目。計畫修訂先依 `feedback.md` 定位需要改的段落,修改驗收條件或任務時再檢查受影響的對應關係。檔案格式、任務對驗收條件的覆蓋與任務相依仍由程式驗證;原始需求的語意覆蓋由審查者判斷,以減少反覆讀取文件的 token 用量。
331
+
330
332
  **僵持時仲裁。** 兩種情況會觸發:這輪審查意見和上一輪一樣,或已達重試上限。仲裁者只判斷一件事:照這份計畫實作,能不能滿足需求。
331
333
 
332
334
  | 有幾家 | 誰來仲裁 | 結果 |
@@ -335,7 +337,7 @@ agentflowctl agent cycle claude-strong,codex,gemini # 不帶參數時顯
335
337
  | 兩家 | 兩家各自在全新 context 裡判斷 | 都核准就繼續;都不核准就依裁決意見修訂並重新審查;分歧依 `tieBreak` |
336
338
  | 一家 | 同一家 | 由它自己仲裁 |
337
339
 
338
- 兩家時的仲裁是雙盲的。仲裁者只看計畫,以及一份不含審查者名稱的爭議清單(`.flow/dispute.md`)。爭議清單用 `<issue>` 包住每則意見;給修訂者的 `.flow/feedback.md` 則用 `<opinion author="…">` 包住每位審查者的意見,避免意見內文與外層結構混淆。帶有名稱的審查紀錄移到 worktree 以外。`tieBreak` 預設 `proceed`,因為後面還有測試紅燈、綠燈、verify 與程式碼審查。
340
+ 兩家時的仲裁是雙盲的。仲裁者只看計畫,以及一份不含審查者名稱的爭議清單(`.flow/dispute.md`)。爭議清單用 `<issue>` 包住每則意見;給修訂者的 `.flow/feedback.md` 則用 `<opinion author="…">` 包住每位審查者的意見,避免意見內文與外層結構混淆。帶有名稱的審查紀錄移到 worktree 以外。`tieBreak` 預設 `proceed`,因為後面還有測試紅燈、綠燈、任務審查、verify 與程式碼審查。
339
341
 
340
342
  計畫定案或仲裁最終停止時,裁決與每位仲裁者的理由附在 `plan.md` 最後的「仲裁紀錄」。需再修訂時,裁決理由寫進 `.flow/feedback.md`,供修訂者處理;重新審查會從第一輪計數。原始審查與每輪仲裁紀錄在 `.agentflowctl/runs/<id>/reviews/`。兩家都要求修改時不因仲裁輪數而直接失敗;整個 run 仍受 `maxAgentRuns` 限制。
341
343
 
@@ -351,17 +353,25 @@ agentflowctl agent cycle claude-strong,codex,gemini # 不帶參數時顯
351
353
  | plan_fix | 依 `fixStrategy` | 修改後仍通過 plan 的格式與 DAG 檢查 | 還原並重試 |
352
354
  | 仲裁 | 見上一節 | 一致核准;分歧依 `tieBreak` | 兩家都不核准時進入 plan_fix 再審查;第三方不核准或 `tieBreak: stop` 時失敗 |
353
355
  | 人工確認 | 你(只有 `--manual-plan`) | `agentflowctl approve` | — |
354
- | implement 紅燈 | 洗牌輪流,每家各一次 | 有測試變更,而且測試執行後失敗 | 還原並重試 |
355
- | implement 綠燈 | 測試作者以外隨機一位 | 測試檔沒有任何修改,而且測試通過 | 還原,或帶著輸出重試 |
356
+ | implement 紅燈 | 洗牌輪流,每家各一次 | 有測試變更,而且測試執行後失敗;規格與計畫檔沒有被修改 | 還原並重試 |
357
+ | implement 綠燈 | 測試作者以外隨機一位 | 測試檔與規格、計畫檔都沒有被修改,而且測試通過 | 還原,或帶著輸出重試 |
358
+ | implement 任務審查 | 作者以外隨機挑(可多位,不重複) | 所有審查者都 `approve` 這個任務的變更 | 進入任務修正 |
359
+ | implement 任務驗證 | — | `install` 與所有 `checks` 通過,才換下一個任務 | 進入任務修正 |
360
+ | implement 任務修正 | 與 fix 相同 | 與 fix 相同;修完重新任務審查、任務驗證 | 還原並重試 |
356
361
  | verify | — | `install` 與所有 `checks` 通過 | 交回作者修正 |
357
- | review | 作者以外隨機挑(可多位,不重複) | 所有審查者都 `approve` | 依 `fixStrategy` 交給他人修正 |
362
+ | fix | verify 失敗交回作者;審查意見依 `fixStrategy` | 沒有刪除測試檔,也沒有修改規格與計畫檔 | 還原並重試 |
363
+ | review | 作者以外隨機挑(可多位,不重複) | 所有審查者都 `approve`(審查者對程式碼與規格、計畫檔的修改一律還原) | 依 `fixStrategy` 交給他人修正 |
358
364
  | pr | — | push 成功;有 `gh` 就開 PR | — |
359
365
 
360
366
  驗收條件寫在 `.flow/acceptance.json`(`AC-1`…),任務寫在 `.flow/tasks.json`(`T-1`…)。
367
+ 每個任務的寫測試與寫實作 prompt 只帶入該任務對應的驗收條件;agent 優先讀任務與相關程式碼,遇到資訊不足或矛盾才查規格、計畫的相關段落。agent 可先跑相關測試,紅燈與完整測試仍由外部流程執行與判定,減少重複讀取文件和全套測試輸出所用的 token。
368
+ 計畫定案後(實作、修正、程式碼審查)不可修改 `.flow/` 裡的規格與計畫檔(`spec.md`、`acceptance.json`、`plan.md`、`tasks.json`、`tasks.ordered.json`);實作與修正時被改就還原並重試,因為寫出的程式碼可能依賴被改過的規格,必須重寫;審查者只交出審查結果,修改直接還原即可,不必重跑審查。對規格有疑慮要寫進交接事項。
369
+ 每個任務綠燈後先做任務審查:審查者只看這個任務寫測試前到目前的 diff(`.flow/diff.patch`)與這個任務的驗收條件,審查紀錄存成 `.flow/review-<任務>-<審查者>.json`。審查者避開最後實作者;有第三家可選時,也優先避開測試作者,只有兩家或審查人數不足時才由測試作者審查。任務審查不要求結清所有交接事項(可能屬於後面的任務),最後的程式碼審查才會擋。審查通過後執行任務驗證(與 verify 相同的 `install` 與 `checks`)。審查要求修改或驗證失敗都進入任務修正,修正者的挑法與 fix 相同,修完回到任務審查再驗證。審查執行或任務修正若先失敗、後成功,會清除各自的連續失敗次數。所有任務完成後,仍照原本流程對整份變更跑一次 verify 與程式碼審查。
370
+ 程式碼審查仍逐條核對所有驗收條件,但只在 `review.json` 列出未通過或其他重要問題;先看 diff 與相關檔案,驗收條件不清楚時才查規格。修正階段先依 `feedback.md` 定位問題並執行相關檢查,完整檢查仍由後續 verify 執行,以減少反覆讀取完整文件與測試輸出。
361
371
 
362
372
  ## Prompt 結構
363
373
 
364
- 每個階段的 prompt 都有專屬角色:需求分析師、軟體架構師、計畫審查者、計畫修訂者、中立仲裁者、測試工程師、實作工程師、除錯工程師、程式碼審查者。內容用 XML 標籤分段:`<role>`、`<context>`、`<inputs>`、`<steps>`、`<constraints>`、`<output_format>`、`<reply_format>`。
374
+ 每個階段的 prompt 都有專屬角色:需求分析師、軟體架構師、計畫審查者、計畫修訂者、中立仲裁者、測試工程師、實作工程師、任務審查者、除錯工程師、程式碼審查者。內容用 XML 標籤分段:`<role>`、`<context>`、`<inputs>`、`<steps>`、`<constraints>`、`<output_format>`、`<reply_format>`。
365
375
 
366
376
  Agent 的最後回覆要附上 XML 中繼資料:
367
377
 
@@ -382,7 +392,7 @@ Agent 的最後回覆要附上 XML 中繼資料:
382
392
 
383
393
  程式只在原有關卡通過後接收交接回覆,並將正式紀錄原子儲存於 `.agentflowctl/runs/<id>/handoff.json`。`action` 是需要後續處理的事項;`info` 只供參考。作者只能提出已修正並附證據,審查者才能確認結案或附理由接受。XML `<concerns>` 可以供人閱讀,但重要疑慮必須寫進交接 JSON,才能交給下一位 agent。額度代打與重試不會接收失敗呼叫的交接內容;中斷後可用 `resume` 接續。
384
394
 
385
- 計畫審查或程式碼審查若核准,但該階段仍有未結的 `action`,程式會視為互相矛盾的審查結果並重試。計畫定案和開 PR 前也會再檢查一次;`info` 會提供給目標階段閱讀,但不阻擋通關。
395
+ 計畫審查或程式碼審查若核准,但該階段仍有未結的 `action`,程式會視為互相矛盾的審查結果並重試。計畫定案和開 PR 前也會再檢查一次;`info` 會提供給目標階段閱讀,但不阻擋通關。審查結果本身也要一致:核准時 `items` 不可有未通過的項目,要求修改時至少要列一筆,否則視為格式錯誤並重新審查。
386
396
 
387
397
  ## Adapter
388
398
 
package/dist/cli.js CHANGED
@@ -88,6 +88,8 @@ async function resolveCycle(flag) {
88
88
  throw new Error(`設定的 agent(${defined.join("、")})都沒有偵測到已安裝的 CLI,可用 agentflowctl doctor 檢查`);
89
89
  return found;
90
90
  }
91
+ /** status 任務清單中,進行中任務的標記 */
92
+ const TASK_PHASE_MARK = { tests: "🧪", code: "🛠️ ", review: "👀", verify: "🔍", fix: "🩹" };
91
93
  const program = new Command()
92
94
  .name("agentflowctl")
93
95
  .description("在專案資料夾內執行的 Agent 開發流程:規格 → 計畫 → TDD 實作 → 驗證 → 審查 → PR")
@@ -228,7 +230,7 @@ program
228
230
  console.log("\n任務");
229
231
  tasks.data.forEach((t, i) => {
230
232
  const active = i === run.taskIndex && run.stage === "implement";
231
- const mark = i < run.taskIndex ? "✅" : active ? (run.taskPhase === "tests" ? "🧪" : "🛠️ ") : "⬜";
233
+ const mark = i < run.taskIndex ? "✅" : active ? TASK_PHASE_MARK[run.taskPhase] : "⬜";
232
234
  console.log(` ${mark} ${t.id} ${t.title}`);
233
235
  });
234
236
  });
@@ -312,17 +314,27 @@ agent
312
314
  });
313
315
  agent
314
316
  .command("setup")
315
- .description("互動式設定:偵測已安裝的 claude、codex、gemini,逐一選擇要不要加入並設定參與的 agent")
317
+ .description("互動式設定:偵測本機已安裝的 agent CLI,逐一選擇要不要加入並設定參與的 agent")
316
318
  .action(async () => {
317
319
  if (!stdin.isTTY)
318
320
  throw new Error("agent setup 需要互動式終端機,請改用 agent add");
319
321
  const detected = {};
320
322
  for (const a of SETUP_ADAPTERS)
321
323
  detected[a] = await probeAgent({ adapter: a, extraArgs: [] });
324
+ let configured;
325
+ try {
326
+ const cfg = loadRepoConfig();
327
+ configured = {};
328
+ for (const name of Object.keys(cfg.agents))
329
+ configured[name] = await probeAgent(resolveAgent(cfg, name));
330
+ }
331
+ catch {
332
+ // 設定檔不合法時照樣進精靈,只是不列出已設定的 agent
333
+ }
322
334
  const rl = createInterface({ input: stdin, output: stdout });
323
335
  let edit;
324
336
  try {
325
- edit = await runSetup(readRawConfig(configPath()), { ask: (q) => rl.question(q), detected, log: (l) => console.log(l) });
337
+ edit = await runSetup(readRawConfig(configPath()), { ask: (q) => rl.question(q), detected, configured, log: (l) => console.log(l) });
326
338
  }
327
339
  finally {
328
340
  rl.close();
package/dist/engine.js CHANGED
@@ -12,9 +12,9 @@ import { CMD_AGENT, nextLogFile } from "./logs.js";
12
12
  import { exec } from "./proc.js";
13
13
  import { arbiterPanel, availableAgent, fixAgent, planAgent, planFixAgent, reviewers, specAgent, taskAgents } from "./roles.js";
14
14
  import { resolveAgent, runAgent, runCommand } from "./runner.js";
15
- import { AcceptanceList, ArbiterResult, RepoConfig, ReviewResult, TaskList, } from "./schemas.js";
15
+ import { AcceptanceList, ArbiterResult, ConsistentReviewResult, RepoConfig, TaskList, } from "./schemas.js";
16
16
  import { addSubstitution, addUsage, agentRuns, saveRun } from "./store.js";
17
- import { orderTasks } from "./tasks.js";
17
+ import { orderTasks, taskAcceptance } from "./tasks.js";
18
18
  import { readJsonFile, renderPrompt, tail } from "./util.js";
19
19
  // ───────────────────────── 共用工具 ─────────────────────────
20
20
  const info = (run, msg) => console.log(`[${run.id}] ${msg}`);
@@ -168,15 +168,17 @@ async function specStage(run) {
168
168
  return retry(run, "spec", handoffError, "spec");
169
169
  return succeed(run, "spec", "plan");
170
170
  }
171
- // ── 規格與計畫檔案:審查者與仲裁者只能讀,不能改 ──
171
+ // ── 規格與計畫檔案:計畫審查、仲裁與計畫定案後的所有階段只能讀,不能改 ──
172
172
  const PLAN_FILES = ["spec.md", "acceptance.json", "plan.md", "tasks.json"];
173
- function snapshotPlan(run) {
174
- return Object.fromEntries(PLAN_FILES.map((f) => [f, existsSync(flowFile(run, f)) ? readFileSync(flowFile(run, f), "utf8") : undefined]));
173
+ /** 計畫定案後(實作、修正、程式碼審查)另外依賴排好的任務順序,同樣不能被改 */
174
+ const LOCKED_FILES = [...PLAN_FILES, "tasks.ordered.json"];
175
+ function snapshotPlan(run, files = PLAN_FILES) {
176
+ return Object.fromEntries(files.map((f) => [f, existsSync(flowFile(run, f)) ? readFileSync(flowFile(run, f), "utf8") : undefined]));
175
177
  }
176
- /** 把被審查者動過的計畫檔案還原成審查前的內容 */
178
+ /** 把被動過的計畫檔案還原成快照的內容,回傳被改的檔名 */
177
179
  function restorePlan(run, snap) {
178
180
  const changed = [];
179
- for (const f of PLAN_FILES) {
181
+ for (const f of Object.keys(snap)) {
180
182
  const now = existsSync(flowFile(run, f)) ? readFileSync(flowFile(run, f), "utf8") : undefined;
181
183
  if (now === snap[f])
182
184
  continue;
@@ -271,7 +273,7 @@ async function planReviewStage(run) {
271
273
  info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
272
274
  if (!r.ok)
273
275
  return retry(run, "plan-review-run", `Agent 執行失敗:${r.summary}`, "plan_review");
274
- const review = readJsonFile(flowFile(run, "plan-review.json"), ReviewResult);
276
+ const review = readJsonFile(flowFile(run, "plan-review.json"), ConsistentReviewResult);
275
277
  if (!review.ok)
276
278
  return retry(run, "plan-review-run", review.error, "plan_review");
277
279
  const handoffError = finishHandoff(run, outcome, "reviewer", { target: "plan", verdict: review.data.verdict });
@@ -411,6 +413,9 @@ async function arbitratePlan(run) {
411
413
  }
412
414
  return { ...run, stage: "failed", failedStage: "plan_review", failureReason: `${summary},需要人工決定(見 plan.md 的仲裁紀錄)` };
413
415
  }
416
+ function planTamperedMessage(files) {
417
+ return `計畫定案後不可修改規格與計畫檔,已還原你的變更:${files.map((f) => `.flow/${f}`).join(", ")}。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。`;
418
+ }
414
419
  async function implementStage(run) {
415
420
  const tasks = loadOrderedTasks(run);
416
421
  const task = tasks[run.taskIndex];
@@ -424,18 +429,34 @@ async function implementStage(run) {
424
429
  const testCmd = `${cfg.install} && ${cfg.test}`;
425
430
  const progress = `${run.taskIndex + 1}/${tasks.length} ${task.id} ${task.title}`;
426
431
  const taskJson = JSON.stringify(task, null, 2);
432
+ const acceptance = readJsonFile(flowFile(run, "acceptance.json"), AcceptanceList);
433
+ if (!acceptance.ok)
434
+ throw new Error(acceptance.error);
435
+ const acceptanceJson = JSON.stringify(taskAcceptance(task, acceptance.data), null, 2);
436
+ if (run.taskPhase === "review")
437
+ return taskReviewStep(run, task, progress, taskJson, acceptanceJson);
438
+ if (run.taskPhase === "verify")
439
+ return taskVerifyStep(run, task, progress);
440
+ if (run.taskPhase === "fix")
441
+ return taskFixStep(run, task, progress);
427
442
  const agents = taskAgents(run.cycle, run.taskIndex, cfg.tddSplit, run.id);
428
443
  // ── 紅燈:只寫測試,而且測試必須失敗 ──
429
444
  if (run.taskPhase === "tests") {
430
445
  const key = `${task.id}:tests`;
431
446
  info(run, `🧪 [${progress}] 撰寫測試(${agents.tests})`);
432
447
  const before = await headCommit(repo);
433
- const outcome = await agentStep(run, agents.tests, `${task.id}-tests`, renderPrompt("implement-tests", { task: taskJson, testPattern: cfg.testPattern, testCmd }), { kind: "write", reset: () => resetTo(repo, before) });
448
+ const snap = snapshotPlan(run, LOCKED_FILES);
449
+ const outcome = await agentStep(run, agents.tests, `${task.id}-tests`, renderPrompt("implement-tests", { task: taskJson, acceptance: acceptanceJson, testPattern: cfg.testPattern, testCmd }), { kind: "write", reset: async () => { await resetTo(repo, before); restorePlan(run, snap); } });
434
450
  const { r, agent: testsAuthor } = outcome;
451
+ const tampered = restorePlan(run, snap);
435
452
  if (!r.ok) {
436
453
  await resetTo(repo, before);
437
454
  return retry(run, key, `Agent 執行失敗:${r.summary}`, "implement");
438
455
  }
456
+ if (tampered.length) {
457
+ await resetTo(repo, before);
458
+ return retry(run, key, planTamperedMessage(tampered), "implement");
459
+ }
439
460
  const commit = await commitAll(repo, `test(${task.id}): ${task.title} [${testsAuthor}]`);
440
461
  if (!commit)
441
462
  return retry(run, key, "沒有任何檔案變更,這個階段必須撰寫測試。", "implement");
@@ -456,7 +477,7 @@ async function implementStage(run) {
456
477
  }
457
478
  writeFileSync(flowFile(run, "red-output.txt"), red.output);
458
479
  info(run, `🔴 [${progress}] 測試如預期失敗`);
459
- return { ...succeed(run, key, "implement"), taskPhase: "code", testsCommit: commit, lastTestsAuthor: testsAuthor };
480
+ return { ...succeed(run, key, "implement"), taskPhase: "code", taskBase: before, testsCommit: commit, lastTestsAuthor: testsAuthor };
460
481
  }
461
482
  // ── 綠燈:實作到測試通過,而且不可動測試 ──
462
483
  const key = `${task.id}:code`;
@@ -465,10 +486,16 @@ async function implementStage(run) {
465
486
  throw new Error("缺少 testsCommit,狀態不一致");
466
487
  info(run, `🛠️ [${progress}] 實作(${agents.code},測試由 ${run.lastTestsAuthor ?? agents.tests} 撰寫)`);
467
488
  const redOutput = existsSync(flowFile(run, "red-output.txt")) ? readFileSync(flowFile(run, "red-output.txt"), "utf8") : "";
468
- const outcome = await agentStep(run, agents.code, `${task.id}-code`, renderPrompt("implement-code", { task: taskJson, testCmd, redOutput: tail(redOutput, 3000) }), { kind: "write", reset: () => resetTo(repo, testsCommit) });
489
+ const snap = snapshotPlan(run, LOCKED_FILES);
490
+ const outcome = await agentStep(run, agents.code, `${task.id}-code`, renderPrompt("implement-code", { task: taskJson, acceptance: acceptanceJson, testCmd, redOutput: tail(redOutput, 3000) }), { kind: "write", reset: async () => { await resetTo(repo, testsCommit); restorePlan(run, snap); } });
469
491
  const { r, agent: codeAuthor } = outcome;
492
+ const tampered = restorePlan(run, snap);
470
493
  if (!r.ok)
471
494
  return retry(run, key, `Agent 執行失敗:${r.summary}`, "implement");
495
+ if (tampered.length) {
496
+ await resetTo(repo, testsCommit);
497
+ return retry(run, key, planTamperedMessage(tampered), "implement");
498
+ }
472
499
  await commitAll(repo, `feat(${task.id}): ${task.title} [${codeAuthor}]`);
473
500
  const touched = (await changedFiles(repo, testsCommit, await headCommit(repo))).filter((f) => testRe.test(f));
474
501
  if (touched.length) {
@@ -485,28 +512,83 @@ async function implementStage(run) {
485
512
  await resetTo(repo, testsCommit);
486
513
  return retry(run, key, handoffError, "implement");
487
514
  }
488
- info(run, `🟢 [${progress}] 完成`);
515
+ info(run, `🟢 [${progress}] 測試通過`);
516
+ return { ...succeed(run, key, "implement"), taskPhase: "review", lastWriter: codeAuthor };
517
+ }
518
+ // ── 任務審查:只看這個任務的變更與驗收條件 ──
519
+ async function taskReviewStep(run, task, progress, taskJson, acceptanceJson) {
520
+ const key = `${task.id}:review`;
521
+ const base = run.taskBase ?? (run.testsCommit && `${run.testsCommit}~1`);
522
+ if (!base)
523
+ throw new Error("缺少 taskBase,狀態不一致");
524
+ const result = await codeReview(run, {
525
+ base,
526
+ seed: `${run.id}:review:${task.id}:${run.attempts[key] ?? 0}`,
527
+ label: `[${progress}] 任務審查`,
528
+ step: `${task.id}-review`,
529
+ prompt: (reviewer, authors) => renderPrompt("task-review", { reviewer, authors, task: taskJson, acceptance: acceptanceJson }),
530
+ saveAs: (reviewer) => `review-${task.id}-${reviewer}.json`,
531
+ testAuthor: run.lastTestsAuthor,
532
+ // 未結交接事項可能屬於後面的任務,由最後的整體審查把關
533
+ gate: false,
534
+ runKey: `${task.id}:review-run`,
535
+ backTo: "implement",
536
+ });
537
+ if ("run" in result)
538
+ return result.run;
539
+ const reviewed = succeed(run, `${task.id}:review-run`, "implement");
540
+ if (!result.objector)
541
+ return { ...succeed(reviewed, key, "implement"), taskPhase: "verify" };
542
+ return {
543
+ ...retry(reviewed, key, `任務審查要求修改:\n\n${result.issues.join("\n\n")}`, "implement"),
544
+ taskPhase: "fix",
545
+ fixSource: "review",
546
+ lastReviewer: result.objector,
547
+ };
548
+ }
549
+ // ── 任務驗證:通過才進入下一個任務 ──
550
+ async function taskVerifyStep(run, task, progress) {
551
+ const key = `${task.id}:verify`;
552
+ info(run, `🔍 [${progress}] 執行驗證`);
553
+ const report = await runChecks(run, `${task.id}-`);
554
+ if (report)
555
+ return { ...retry(run, key, report, "implement"), taskPhase: "fix", fixSource: "verify" };
556
+ info(run, `✅ [${progress}] 完成`);
489
557
  return {
490
558
  ...succeed(run, key, "implement"),
491
559
  taskIndex: run.taskIndex + 1,
492
560
  taskPhase: "tests",
561
+ taskBase: undefined,
493
562
  testsCommit: undefined,
494
563
  lastTestsAuthor: undefined,
495
- lastWriter: codeAuthor,
496
564
  };
497
565
  }
498
- async function verifyStage(run) {
499
- info(run, "🔍 執行驗證");
566
+ // ── 任務修正:修完重新審查、驗證 ──
567
+ async function taskFixStep(run, task, progress) {
568
+ const result = await applyFix(run, {
569
+ seed: `${run.id}:fix:${task.id}:${run.attempts[`${task.id}:review`] ?? 0}`,
570
+ label: `[${progress}] `,
571
+ step: `${task.id}-fix`,
572
+ commitScope: `fix(${task.id})`,
573
+ key: `${task.id}:fix`,
574
+ backTo: "implement",
575
+ });
576
+ if ("run" in result)
577
+ return result.run;
578
+ return { ...succeed(run, `${task.id}:fix`, "implement"), taskPhase: "review", lastWriter: result.agent };
579
+ }
580
+ /** 執行 install 與所有 checks,結果寫入 verify.json;有失敗時回傳給修正者的報告 */
581
+ async function runChecks(run, stepPrefix = "") {
500
582
  const cfg = loadRepoConfig();
501
583
  const results = [];
502
- const install = await runCommand(target(run, "install", CMD_AGENT), cfg.install);
584
+ const install = await runCommand(target(run, `${stepPrefix}install`, CMD_AGENT), cfg.install);
503
585
  if (!install.ok) {
504
586
  info(run, ` ✗ install${logHint(run, install.seq)}`);
505
587
  results.push({ name: "install", ok: false, output: tail(install.output) });
506
588
  }
507
589
  else {
508
590
  for (const check of cfg.checks) {
509
- const r = await runCommand(target(run, check.name, CMD_AGENT), check.cmd);
591
+ const r = await runCommand(target(run, `${stepPrefix}${check.name}`, CMD_AGENT), check.cmd);
510
592
  info(run, ` ${r.ok ? "✓" : "✗"} ${check.name}${r.ok ? "" : logHint(run, r.seq)}`);
511
593
  results.push({ name: check.name, ok: r.ok, output: tail(r.output, 3000) });
512
594
  }
@@ -514,11 +596,18 @@ async function verifyStage(run) {
514
596
  writeFileSync(flowFile(run, "verify.json"), JSON.stringify(results, null, 2));
515
597
  const failed = results.filter((r) => !r.ok);
516
598
  if (failed.length === 0)
599
+ return undefined;
600
+ return failed.map((f) => `## ${f.name} 失敗\n\n\`\`\`\n${f.output}\n\`\`\``).join("\n\n");
601
+ }
602
+ async function verifyStage(run) {
603
+ info(run, "🔍 執行驗證");
604
+ const report = await runChecks(run);
605
+ if (!report)
517
606
  return succeed(run, "verify", "review");
518
- const report = failed.map((f) => `## ${f.name} 失敗\n\n\`\`\`\n${f.output}\n\`\`\``).join("\n\n");
519
607
  return { ...retry(run, "verify", report, "fix"), fixSource: "verify" };
520
608
  }
521
- async function fixStage(run) {
609
+ /** 修正驗證錯誤或審查意見;成功時回傳實際修正者,未通過時回傳重試後的 run */
610
+ async function applyFix(run, opts) {
522
611
  const cfg = loadRepoConfig();
523
612
  const source = run.fixSource ?? "verify";
524
613
  const agent = fixAgent(run.cycle, {
@@ -526,77 +615,123 @@ async function fixStage(run) {
526
615
  strategy: cfg.fixStrategy,
527
616
  lastWriter: run.lastWriter,
528
617
  lastReviewer: run.lastReviewer,
529
- seed: `${run.id}:fix:${run.attempts.review ?? 0}`,
618
+ seed: opts.seed,
530
619
  });
531
620
  const why = source === "review" ? `依 ${run.lastReviewer ?? "reviewer"} 的審查意見` : "修正驗證錯誤";
532
- info(run, `🩹 ${why}(${agent})`);
621
+ info(run, `🩹 ${opts.label}${why}(${agent})`);
533
622
  const repo = worktreeDir(run.id);
534
623
  const feedback = readFeedback(run);
535
624
  const before = await headCommit(repo);
536
- const outcome = await agentStep(run, agent, "fix", renderPrompt("fix", { testPattern: cfg.testPattern }), {
625
+ const snap = snapshotPlan(run, LOCKED_FILES);
626
+ const outcome = await agentStep(run, agent, opts.step, renderPrompt("fix", { testPattern: cfg.testPattern }), {
537
627
  kind: "write",
538
- reset: () => resetTo(repo, before),
628
+ reset: async () => { await resetTo(repo, before); restorePlan(run, snap); },
539
629
  });
540
630
  const { r, agent: actual } = outcome;
631
+ const tampered = restorePlan(run, snap);
632
+ const again = (reason) => ({ run: retry(run, opts.key, reason, opts.backTo) });
541
633
  if (!r.ok)
542
- return retry(run, "fix", `${feedback}\n\n(上次修正時 Agent 執行失敗:${r.summary})`, "fix");
543
- await commitAll(repo, `fix: ${why} [${actual}]`);
634
+ return again(`${feedback}\n\n(上次修正時 Agent 執行失敗:${r.summary})`);
635
+ if (tampered.length) {
636
+ await resetTo(repo, before);
637
+ return again(`${feedback}\n\n另外:${planTamperedMessage(tampered)}`);
638
+ }
639
+ await commitAll(repo, `${opts.commitScope}: ${why} [${actual}]`);
544
640
  const testRe = new RegExp(cfg.testPattern);
545
641
  const deleted = (await changedFiles(repo, before, await headCommit(repo), "D")).filter((f) => testRe.test(f));
546
642
  if (deleted.length) {
547
643
  await resetTo(repo, before);
548
- return retry(run, "fix", `${feedback}\n\n另外:不可刪除測試檔來讓檢查通過,已還原:${deleted.join(", ")}`, "fix");
644
+ return again(`${feedback}\n\n另外:不可刪除測試檔來讓檢查通過,已還原:${deleted.join(", ")}`);
549
645
  }
550
646
  const handoffError = finishHandoff(run, outcome, "writer");
551
647
  if (handoffError) {
552
648
  await resetTo(repo, before);
553
- return retry(run, "fix", handoffError, "fix");
649
+ return again(handoffError);
554
650
  }
651
+ return { agent: actual };
652
+ }
653
+ async function fixStage(run) {
654
+ const result = await applyFix(run, {
655
+ seed: `${run.id}:fix:${run.attempts.review ?? 0}`,
656
+ label: "",
657
+ step: "fix",
658
+ commitScope: "fix",
659
+ key: "fix",
660
+ backTo: "fix",
661
+ });
662
+ if ("run" in result)
663
+ return result.run;
555
664
  // 修正者成為新的作者,下一輪審查會換成別人
556
- return { ...to(run, "verify"), lastWriter: actual };
665
+ return { ...to(run, "verify"), lastWriter: result.agent };
557
666
  }
558
- async function reviewStage(run) {
667
+ /**
668
+ * 由作者以外的審查小組審查 base 之後的變更;回傳第一位要求修改的審查者與所有意見,
669
+ * 審查本身未完成(執行失敗、格式錯誤、交接不合格)時回傳重試後的 run。
670
+ * gate 為 true 時,核准前必須結清所有程式碼類的未結交接事項。
671
+ */
672
+ async function codeReview(run, opts) {
559
673
  const cfg = loadRepoConfig();
560
674
  const repo = worktreeDir(run.id);
561
- const panel = reviewers(run.cycle, run.lastWriter, cfg.reviewQuorum, `${run.id}:review:${run.attempts.review ?? 0}`);
562
- writeFileSync(flowFile(run, "diff.patch"), await git(repo, "diff", `${run.baseBranch}...HEAD`));
563
- const authors = [...new Set((await git(repo, "log", "--format=%s", `${run.baseBranch}..HEAD`)).match(/\[[^\]]+\]$/gm) ?? [])]
675
+ const panel = reviewers(run.cycle, run.lastWriter, cfg.reviewQuorum, opts.seed, opts.testAuthor);
676
+ writeFileSync(flowFile(run, "diff.patch"), await git(repo, "diff", `${opts.base}...HEAD`));
677
+ const authors = [...new Set((await git(repo, "log", "--format=%s", `${opts.base}..HEAD`)).match(/\[[^\]]+\]$/gm) ?? [])]
564
678
  .map((s) => s.slice(1, -1));
565
679
  const issues = [];
566
- let firstObjector;
680
+ let objector;
567
681
  for (const [slot, reviewer] of panel.entries()) {
568
- info(run, `👀 程式碼審查(${reviewer})`);
682
+ info(run, `👀 ${opts.label}(${reviewer})`);
569
683
  rmSync(flowFile(run, "review.json"), { force: true });
570
- const outcome = await agentStep(run, reviewer, "review", renderPrompt("review", { reviewer, authors: authors.join("、") || "未知" }), {
571
- kind: "review", slot,
684
+ const snap = snapshotPlan(run, LOCKED_FILES);
685
+ const outcome = await agentStep(run, reviewer, opts.step, opts.prompt(reviewer, authors.join("、") || "未知"), {
686
+ kind: "review", slot, reset: async () => { await discardChanges(repo); restorePlan(run, snap); },
572
687
  });
573
688
  const { r } = outcome;
574
689
  await discardChanges(repo); // 審查者不可改程式碼
690
+ const tampered = restorePlan(run, snap); // .flow/ 不受 git 管理,要另外還原
691
+ if (tampered.length)
692
+ info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
575
693
  if (!r.ok)
576
- return retry(run, "review-run", `Agent 執行失敗:${r.summary}`, "review");
577
- const review = readJsonFile(flowFile(run, "review.json"), ReviewResult);
694
+ return { run: retry(run, opts.runKey, `Agent 執行失敗:${r.summary}`, opts.backTo) };
695
+ const review = readJsonFile(flowFile(run, "review.json"), ConsistentReviewResult);
578
696
  if (!review.ok)
579
- return retry(run, "review-run", review.error, "review");
580
- const handoffError = finishHandoff(run, outcome, "reviewer", { target: "code", verdict: review.data.verdict });
697
+ return { run: retry(run, opts.runKey, review.error, opts.backTo) };
698
+ const gate = opts.gate ? { target: "code", verdict: review.data.verdict } : undefined;
699
+ const handoffError = finishHandoff(run, outcome, "reviewer", gate);
581
700
  if (handoffError)
582
- return retry(run, "review-run", handoffError, "review");
583
- renameSync(flowFile(run, "review.json"), flowFile(run, `review-${reviewer}.json`));
701
+ return { run: retry(run, opts.runKey, handoffError, opts.backTo) };
702
+ renameSync(flowFile(run, "review.json"), flowFile(run, opts.saveAs(reviewer)));
584
703
  if (review.data.verdict === "approve") {
585
704
  info(run, ` ✓ ${reviewer} 核准`);
586
705
  continue;
587
706
  }
588
707
  info(run, ` ✗ ${reviewer} 要求修改`);
589
- firstObjector ??= reviewer;
708
+ objector ??= reviewer;
590
709
  issues.push(opinion(reviewer, review.data.items
591
710
  .filter((i) => i.status !== "met")
592
711
  .map((i) => reviewIssue(i.criterion, i.status, i.note))));
593
712
  }
594
- if (!firstObjector)
713
+ return { objector, issues };
714
+ }
715
+ async function reviewStage(run) {
716
+ const result = await codeReview(run, {
717
+ base: run.baseBranch,
718
+ seed: `${run.id}:review:${run.attempts.review ?? 0}`,
719
+ label: "程式碼審查",
720
+ step: "review",
721
+ prompt: (reviewer, authors) => renderPrompt("review", { reviewer, authors }),
722
+ saveAs: (reviewer) => `review-${reviewer}.json`,
723
+ gate: true,
724
+ runKey: "review-run",
725
+ backTo: "review",
726
+ });
727
+ if ("run" in result)
728
+ return result.run;
729
+ if (!result.objector)
595
730
  return succeed(run, "review", "pr");
596
731
  return {
597
- ...retry(run, "review", `程式碼審查要求修改:\n\n${issues.join("\n\n")}`, "fix"),
732
+ ...retry(run, "review", `程式碼審查要求修改:\n\n${result.issues.join("\n\n")}`, "fix"),
598
733
  fixSource: "review",
599
- lastReviewer: firstObjector,
734
+ lastReviewer: result.objector,
600
735
  };
601
736
  }
602
737
  async function prStage(run) {
package/dist/roles.js CHANGED
@@ -64,9 +64,12 @@ export function taskAgents(cycle, taskIndex, tddSplit, seed) {
64
64
  const tests = taskBag(cycle, seed, Math.floor(taskIndex / n))[taskIndex % n];
65
65
  return { tests, code: tddSplit ? pick(cycle, `${seed}:code:${taskIndex}`, [tests]) : tests };
66
66
  }
67
- /** 隨機挑出 quorum 位彼此不重複、而且不是作者的 reviewer */
68
- export function reviewers(cycle, lastWriter, quorum, seed) {
69
- const picked = shuffled(cycle.filter((c) => c !== lastWriter), seed).slice(0, quorum);
67
+ /** 隨機挑出 quorum 位彼此不重複、而且不是最後作者的 reviewer;任務審查優先避開測試作者 */
68
+ export function reviewers(cycle, lastWriter, quorum, seed, testAuthor) {
69
+ const candidates = shuffled(cycle.filter((c) => c !== lastWriter), seed);
70
+ const preferred = testAuthor ? candidates.filter((c) => c !== testAuthor) : candidates;
71
+ const fallback = testAuthor ? candidates.filter((c) => c === testAuthor) : [];
72
+ const picked = [...preferred, ...fallback].slice(0, quorum);
70
73
  return picked.length ? picked : [lastWriter ?? cycle[0]];
71
74
  }
72
75
  /**
package/dist/schemas.js CHANGED
@@ -83,6 +83,14 @@ export const ReviewResult = z.object({
83
83
  note: z.string().default(""),
84
84
  })),
85
85
  });
86
+ /** 計畫與程式碼審查的結果:verdict 必須和 items 一致,否則視為格式錯誤重試 */
87
+ export const ConsistentReviewResult = ReviewResult.superRefine((review, ctx) => {
88
+ const unmet = review.items.filter((i) => i.status !== "met");
89
+ if (review.verdict === "approve" && unmet.length)
90
+ ctx.addIssue({ code: "custom", message: `verdict 為 approve,但 items 仍有未通過的項目:${unmet.map((i) => i.criterion).join("、")}` });
91
+ if (review.verdict === "changes_requested" && !unmet.length)
92
+ ctx.addIssue({ code: "custom", message: "verdict 為 changes_requested,但 items 沒有列出任何未通過的項目" });
93
+ });
86
94
  /** 仲裁者可能用 reject 表示否決;讀取時正規化,保留相同的理由欄位。 */
87
95
  export const ArbiterResult = ReviewResult.extend({
88
96
  verdict: z.enum(["approve", "changes_requested", "reject"])
@@ -164,7 +172,10 @@ export const FlowRun = z.object({
164
172
  /** 各關卡的連續失敗次數 */
165
173
  attempts: z.record(z.string(), z.number()),
166
174
  taskIndex: z.number().int().nonnegative(),
167
- taskPhase: z.enum(["tests", "code"]),
175
+ /** 目前任務進行到哪一步:寫測試 → 實作 → 審查 → 驗證,審查或驗證未通過時進入修正 */
176
+ taskPhase: z.enum(["tests", "code", "review", "verify", "fix"]),
177
+ /** 目前任務寫測試前的 commit,任務審查只看這之後的變更 */
178
+ taskBase: z.string().optional(),
168
179
  testsCommit: z.string().optional(),
169
180
  /** 目前任務的測試實際由誰撰寫(可能是代打) */
170
181
  lastTestsAuthor: z.string().optional(),
package/dist/setup.js CHANGED
@@ -1,9 +1,11 @@
1
1
  import { addAgent, setAgent, setCycle } from "./agentConfig.js";
2
+ import { ADAPTERS } from "./agents/index.js";
2
3
  /**
3
4
  * `agentflowctl agent setup` 的互動精靈。
4
5
  * 問答與偵測結果由外部注入,這裡只負責流程;所有修改都在記憶體裡完成,確認後才交給呼叫端寫入。
5
6
  */
6
- export const SETUP_ADAPTERS = ["claude", "codex", "gemini"];
7
+ /** 不需自訂指令就能加入的 adapter;新增 adapter 時會自動出現在精靈裡 */
8
+ export const SETUP_ADAPTERS = Object.keys(ADAPTERS).filter((a) => a !== "command");
7
9
  /** 回傳要寫入的設定;使用者沒選任何 agent 或最後沒確認時回傳 null */
8
10
  export async function runSetup(initial, deps) {
9
11
  const { detected, log } = deps;
@@ -16,9 +18,26 @@ export async function runSetup(initial, deps) {
16
18
  const changes = [];
17
19
  const chosen = [];
18
20
  const existing = () => Object.keys((cfg.agents ?? {}));
19
- for (const adapter of SETUP_ADAPTERS) {
20
- log(`${detected[adapter] ? "✅" : "❌"} ${adapter}${detected[adapter] ? "" : "(沒有偵測到 CLI)"}`);
21
- if (!(await confirm(` 加入 ${adapter}?`, detected[adapter])))
21
+ const configured = Object.entries(deps.configured ?? {});
22
+ if (configured.length) {
23
+ const defs = (cfg.agents ?? {});
24
+ log("已設定的 agent:");
25
+ for (const [name, ok] of configured) {
26
+ log(` ${ok ? "✅" : "⚠️ "} ${name}(${defs[name]?.adapter ?? "?"}${ok ? "" : ",找不到可執行的 CLI"})`);
27
+ }
28
+ }
29
+ const installed = Object.keys(detected).filter((a) => detected[a]);
30
+ const missing = Object.keys(detected).filter((a) => !detected[a]);
31
+ if (missing.length)
32
+ log(`沒有偵測到:${missing.join("、")}(安裝後可重跑 agent setup)`);
33
+ if (!installed.length) {
34
+ log("沒有偵測到可加入的 agent CLI,設定沒有變更");
35
+ log("需要自訂指令的 CLI 請改用 agent add <name> --adapter command -- <指令>");
36
+ return null;
37
+ }
38
+ for (const adapter of installed) {
39
+ log(`✅ ${adapter}`);
40
+ if (!(await confirm(` 加入 ${adapter}?`, true)))
22
41
  continue;
23
42
  let name;
24
43
  for (;;) {
package/dist/tasks.js CHANGED
@@ -1,5 +1,15 @@
1
1
  /** 一個任務最多做兩件事:對應的驗收條件超過這個數量就要再拆 */
2
2
  export const MAX_TASK_ACCEPTANCE = 2;
3
+ /** 只取目前任務負責的驗收條件,避免每次實作都重讀整份清單。 */
4
+ export function taskAcceptance(task, acceptance) {
5
+ const byId = new Map(acceptance.map((item) => [item.id, item]));
6
+ return task.acceptance.map((id) => {
7
+ const item = byId.get(id);
8
+ if (!item)
9
+ throw new Error(`${task.id} 對應的驗收條件 ${id} 不存在`);
10
+ return item;
11
+ });
12
+ }
3
13
  /**
4
14
  * 檢查任務清單並依相依關係排序(Kahn 演算法)。
5
15
  * 回傳排序後的任務,或回傳一段可以直接回饋給 Agent 的錯誤說明。
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "agentflowctl",
3
3
  "license": "MIT",
4
- "version": "0.8.0",
4
+ "version": "0.10.0",
5
5
  "description": "跨廠商 AI 開發 harness:Claude Code、Codex、Gemini 輪流實作、審查、修正",
6
6
  "keywords": [
7
7
  "ai",
package/prompts/fix.md CHANGED
@@ -18,19 +18,21 @@
18
18
  </handoff>
19
19
 
20
20
  <inputs>
21
- - .flow/feedback.md:失敗的檢查(型別、lint、測試、建置)或審查意見
22
- - .flow/spec.md 與 .flow/acceptance.json:規格
21
+ - .flow/feedback.md:先讀取,確認失敗的檢查或審查意見與其證據
22
+ - 相關程式碼與測試:依 feedback 指出的檔案和問題範圍查閱
23
+ - .flow/acceptance.json 與 .flow/spec.md:預期行為不清楚或互相矛盾時,才查相關條件或段落
23
24
  </inputs>
24
25
 
25
26
  <steps>
26
- 1. 找出每個問題的根本原因再修正,不要只針對症狀打補丁。
27
- 2. 在沙箱內自行執行相關檢查,確認問題已解決且沒有造成新的錯誤。
27
+ 1. 優先處理 feedback 與交接事項指出的問題;先定位相關程式碼和測試,找出根本原因再修正。
28
+ 2. 在沙箱內執行與修改相關的檢查,確認問題已解決。完整檢查由後續 verify 階段執行;需要判斷跨模組影響時再自行擴大檢查範圍。
28
29
  </steps>
29
30
 
30
31
  <constraints>
31
32
  - 不可刪除測試檔(檔名符合 `{{testPattern}}`),也不可用 skip、放寬斷言、`@ts-ignore`、`eslint-disable` 等方式讓檢查通過。
32
33
  - 如果測試本身確實有誤,可以修正測試,但必須在回覆的 `<concerns>` 說明理由。
33
34
  - 不要執行 git commit(權限設定已禁止)。
35
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
34
36
  </constraints>
35
37
 
36
38
  <reply_format>
@@ -23,8 +23,15 @@
23
23
  ```
24
24
  </task>
25
25
 
26
+ <acceptance>
27
+ 這個任務負責的驗收條件:
28
+ ```json
29
+ {{acceptance}}
30
+ ```
31
+ </acceptance>
32
+
26
33
  <inputs>
27
- 完整規格與計畫請參考 .flow/spec.md、.flow/plan.md。
34
+ 先依上面的任務、驗收條件與現有測試工作。只有資訊不足或互相矛盾時,再閱讀 .flow/spec.md、.flow/plan.md 的相關段落,並在交接中指出問題;不必通讀整份文件。
28
35
  </inputs>
29
36
 
30
37
  <red_output>
@@ -36,13 +43,14 @@
36
43
  <steps>
37
44
  1. 若 .flow/feedback.md 存在,先閱讀,並依內容修正。
38
45
  2. 撰寫讓測試通過的最小實作,符合專案既有的程式風格與架構。
39
- 3. 執行 `{{testCmd}}` 確認全部測試通過(包含既有測試)。
46
+ 3. 先執行本任務相關的測試,確認實作方向。外部流程會再執行 `{{testCmd}}` 驗證全部測試(包含既有測試),不需要自行重跑全套測試。
40
47
  4. 測試通過後,在不改變行為的前提下整理程式碼。
41
48
  </steps>
42
49
 
43
50
  <constraints>
44
51
  - **不可修改任何測試檔**,修改會被自動還原並視為失敗。若認為測試本身有誤,請寫在回覆的 `<concerns>`。
45
52
  - 不要執行 git commit(權限設定已禁止)。
53
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
46
54
  </constraints>
47
55
 
48
56
  <reply_format>
@@ -23,20 +23,28 @@
23
23
  ```
24
24
  </task>
25
25
 
26
+ <acceptance>
27
+ 這個任務負責的驗收條件:
28
+ ```json
29
+ {{acceptance}}
30
+ ```
31
+ </acceptance>
32
+
26
33
  <inputs>
27
- 完整規格與計畫請參考 .flow/spec.md、.flow/plan.md。
34
+ 先依上面的任務與驗收條件工作。只有資訊不足或互相矛盾時,再閱讀 .flow/spec.md、.flow/plan.md 的相關段落,並在交接中指出問題;不必通讀整份文件。
28
35
  </inputs>
29
36
 
30
37
  <steps>
31
38
  1. 若 .flow/feedback.md 存在,先閱讀,並依內容調整做法。
32
39
  2. 依任務描述撰寫測試,檔名必須符合正規表示式 `{{testPattern}}`。
33
40
  3. 測試必須驗證這個任務要新增的行為,並且因為功能尚未實作而**失敗**。
34
- 4. 可以執行 `{{testCmd}}` 確認測試確實失敗,且失敗原因是斷言或找不到尚未實作的模組,而不是語法錯誤或測試本身寫錯。
41
+ 4. 可以先執行本任務相關的測試,確認失敗原因是斷言或找不到尚未實作的模組,而不是語法錯誤或測試本身寫錯。外部流程會再執行 `{{testCmd}}` 驗證紅燈,不需要自行重跑全套測試。
35
42
  </steps>
36
43
 
37
44
  <constraints>
38
45
  - 不可實作功能本身。可以建立讓測試能編譯所需的最小型別或空殼匯出,但不可以有真正的邏輯。
39
46
  - 不要執行 git commit(權限設定已禁止),外部流程會提交並驗證測試是否失敗。
47
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
40
48
  </constraints>
41
49
 
42
50
  <reply_format>
@@ -22,10 +22,9 @@
22
22
  </requirement>
23
23
 
24
24
  <steps>
25
- 1. 閱讀 .flow/feedback.md,裡面是其他模型的審查意見。
26
- 2. 閱讀目前的 .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json 與相關程式碼。
27
- 3. 逐條處理審查意見,直接修改上述四個檔案。
28
- 4. 在 .flow/plan.md 最後的「## 審查回應」一節,逐條說明每個意見怎麼處理;不同意的意見,請寫出具體理由,而不是忽略它。回應時只談內容,不要提到審查者或你自己是哪個模型、哪家公司,之後可能由第三方匿名仲裁。
25
+ 1. 先閱讀 .flow/feedback.md,依每則意見定位 .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json 中相關的段落或項目;技術細節需要確認時才讀相關程式碼。
26
+ 2. 逐條處理審查意見,只修改需要修訂的檔案。新增或調整驗收條件、任務時,檢查受影響的條件與任務對應。
27
+ 3. 在 .flow/plan.md 最後的「## 審查回應」一節,逐條說明每個意見怎麼處理;不同意的意見,請寫出具體理由,而不是忽略它。回應時只談內容,不要提到審查者或你自己是哪個模型、哪家公司,之後可能由第三方匿名仲裁。
29
28
  </steps>
30
29
 
31
30
  <output_format>
@@ -22,12 +22,8 @@
22
22
  </requirement>
23
23
 
24
24
  <inputs>
25
- - .flow/spec.md:規格
26
- - .flow/acceptance.json:驗收條件
27
- - .flow/plan.md:實作方式
28
- - .flow/tasks.json:任務拆解
29
-
30
- 請同時閱讀相關的既有程式碼,確認計畫符合專案的實際架構。
25
+ - 先核對原始需求、.flow/spec.md、.flow/acceptance.json、.flow/plan.md 與 .flow/tasks.json,確認需求、驗收條件與任務的對應。
26
+ - 依 .flow/tasks.json 各任務 `description` 列出要動的檔案,查閱其中既有的檔案,確認計畫符合專案的架構與慣例;有疑慮時再擴大查閱範圍。
31
27
  </inputs>
32
28
 
33
29
  <review_focus>
@@ -37,6 +33,7 @@
37
33
  4. **技術方向**:是否符合專案既有的架構與慣例?有沒有更簡單的做法,或明顯的風險?
38
34
 
39
35
  措辭、格式這類不影響實作結果的小問題,不需要要求修改。
36
+ 只在證據不足時擴大查閱範圍;檔案格式、任務對驗收條件的覆蓋與任務相依已有程式關卡檢查,不必自行重跑完整檢查。仍須自行判斷原始需求是否被規格與驗收條件涵蓋。
40
37
  </review_focus>
41
38
 
42
39
  <output_format>
@@ -44,15 +41,16 @@
44
41
 
45
42
  ```json
46
43
  {
47
- "verdict": "approve",
44
+ "verdict": "changes_requested",
48
45
  "items": [
49
- { "criterion": "需求覆蓋", "status": "met", "note": "" },
50
46
  { "criterion": "AC-2", "status": "not_met", "note": "沒有涵蓋 API 逾時的情況,建議新增一條驗收條件並由 T-3 負責" }
51
47
  ]
52
48
  }
53
49
  ```
54
50
 
55
51
  - `verdict`:沒有會影響實作結果的問題時為 `approve`,否則為 `changes_requested`。
52
+ - `items` 只列會影響實作結果的問題,每筆 `status` 為 `not_met` 或 `partial`;已通過的項目不必逐條記錄。
53
+ - `approve` 時 `items` 為空陣列;`changes_requested` 時至少要有一筆。兩者不一致會被視為格式錯誤並重新審查。
56
54
  - 每個問題的 `note` 請寫出具體要改哪個檔案的哪個部分,以及建議怎麼改。
57
55
  </output_format>
58
56
 
package/prompts/review.md CHANGED
@@ -18,9 +18,9 @@
18
18
  </handoff>
19
19
 
20
20
  <inputs>
21
- - 規格:.flow/spec.md
22
- - 驗收條件:.flow/acceptance.json
23
- - 本次變更:.flow/diff.patch
21
+ - 驗收條件:.flow/acceptance.json(逐條核對)
22
+ - 本次變更:.flow/diff.patch(先看變更,再按需讀相關程式碼與測試)
23
+ - 規格:.flow/spec.md(驗收條件不清楚或互相矛盾時,才查相關段落)
24
24
  - 自動化檢查結果:.flow/verify.json(已全部通過)
25
25
  </inputs>
26
26
 
@@ -29,6 +29,8 @@
29
29
  2. 是否有明顯的錯誤、邊界情況遺漏、安全問題或效能問題。
30
30
  3. 是否符合專案既有的架構與慣例。
31
31
 
32
+ 先根據 diff 與驗收條件定位需要查閱的檔案;只在證據不足時讀取其他檔案。自動化檢查已由外部流程執行,不必為了審查重跑全套檢查。
33
+
32
34
  風格偏好與無關緊要的小問題不需要要求修改。
33
35
  </review_focus>
34
36
 
@@ -37,16 +39,16 @@
37
39
 
38
40
  ```json
39
41
  {
40
- "verdict": "approve",
42
+ "verdict": "changes_requested",
41
43
  "items": [
42
- { "criterion": "AC-1", "status": "met", "note": "" },
43
- { "criterion": "錯誤處理", "status": "not_met", "note": "API 失敗時沒有顯示錯誤訊息,見 src/form.tsx" }
44
+ { "criterion": "AC-1", "status": "not_met", "note": "src/form.tsx 缺少 API 失敗時的錯誤訊息,請在 catch 中顯示錯誤並補測試" }
44
45
  ]
45
46
  }
46
47
  ```
47
48
 
48
- - `verdict`:全部驗收條件都 `met` 且沒有嚴重問題時為 `approve`,否則為 `changes_requested`。
49
- - 每個驗收條件都要有一筆;額外發現的問題也各自列一筆,`note` 請寫出具體位置與修正方向。
49
+ - `verdict`:逐條核對所有驗收條件後,全部通過且沒有嚴重問題時為 `approve`,否則為 `changes_requested`。
50
+ - `items` 只列未通過的驗收條件與額外發現的重要問題,每筆的 `status` 為 `not_met` 或 `partial`,`note` 請寫出具體位置與修正方向。
51
+ - `approve` 時 `items` 為空陣列;`changes_requested` 時至少要有一筆。兩者不一致會被視為格式錯誤並重新審查。
50
52
  </output_format>
51
53
 
52
54
  <constraints>
@@ -0,0 +1,77 @@
1
+ <role>
2
+ 你是任務審查者({{reviewer}}),負責在單一任務完成後、進入下一個任務前,獨立審查這個任務的變更。本任務的作者:{{authors}}。若你曾撰寫本任務的測試,仍須重新檢查測試是否真正驗證驗收條件;不要預設測試或實作正確。
3
+ </role>
4
+
5
+ <context>
6
+ 目前的工作目錄就是專案(agentflowctl 為這次任務建立的專用 git worktree)。這個任務的測試已由外部流程確認通過;完整的自動化檢查會在你審查之後才執行,所有任務完成後還有一次整體審查。
7
+ </context>
8
+
9
+ <handoff>
10
+ 先閱讀 .flow/handoff-context.md,處理與本任務有關的待辦事項;與本任務無關的事項留給後續任務或整體審查。完成時寫入 .flow/handoff-response.json;即使沒有事項也必須寫出空陣列:
11
+
12
+ ```json
13
+ { "newIssues": [], "dispositions": [] }
14
+ ```
15
+
16
+ 新增事項格式:{ "kind": "action 或 info", "summary": "具體問題", "evidence": "檔案位置或檢查證據", "targetStage": "plan 或 code" }。
17
+ 處置格式:{ "id": "既有事項 ID", "status": "proposed_resolved、resolved 或 accepted", "reason": "具體處理理由", "evidence": "檔案、commit 或檢查結果" }。撰寫者只能用 proposed_resolved 提出修正;審查者可以用 resolved 或 accepted 結案。重要疑慮必須放在這份檔案,不能只寫在回覆的 <concerns>。
18
+ </handoff>
19
+
20
+ <inputs>
21
+ <task>
22
+ {{task}}
23
+ </task>
24
+
25
+ <acceptance>
26
+ {{acceptance}}
27
+ </acceptance>
28
+
29
+ - 本任務的變更:.flow/diff.patch(先看變更,再按需讀相關程式碼與測試)
30
+ - 規格與計畫:.flow/spec.md、.flow/plan.md(任務或驗收條件不清楚、互相矛盾時,才查相關段落)
31
+ </inputs>
32
+
33
+ <review_focus>
34
+ 1. 逐條確認上面每個驗收條件是否真的被實作,而且有對應的測試真正驗證它(不是空洞的測試)。
35
+ 2. 是否有明顯的錯誤、邊界情況遺漏、安全問題或效能問題。
36
+ 3. 是否符合專案既有的架構與慣例,以及任務說明的範圍(沒有做到一半,也沒有做了其他任務的事)。
37
+
38
+ 只審查本任務的變更;其他任務的驗收條件不在這次審查範圍內。先根據 diff 與驗收條件定位需要查閱的檔案,只在證據不足時讀取其他檔案。不必為了審查重跑全套檢查。
39
+
40
+ 風格偏好與無關緊要的小問題不需要要求修改。
41
+ </review_focus>
42
+
43
+ <output_format>
44
+ 寫入 .flow/review.json:
45
+
46
+ ```json
47
+ {
48
+ "verdict": "changes_requested",
49
+ "items": [
50
+ { "criterion": "AC-1", "status": "not_met", "note": "src/form.tsx 缺少 API 失敗時的錯誤訊息,請在 catch 中顯示錯誤並補測試" }
51
+ ]
52
+ }
53
+ ```
54
+
55
+ - `verdict`:逐條核對本任務的驗收條件後,全部通過且沒有嚴重問題時為 `approve`,否則為 `changes_requested`。
56
+ - `items` 只列未通過的驗收條件與額外發現的重要問題,每筆的 `status` 為 `not_met` 或 `partial`,`note` 請寫出具體位置與修正方向。
57
+ - `approve` 時 `items` 為空陣列;`changes_requested` 時至少要有一筆。兩者不一致會被視為格式錯誤並重新審查。
58
+ </output_format>
59
+
60
+ <constraints>
61
+ - 只能寫入 .flow/review.json 與 .flow/handoff-response.json,不可修改任何程式碼,其他變更都會被捨棄。
62
+ </constraints>
63
+
64
+ <reply_format>
65
+ 完成後,回覆的最後必須附上以下 XML 中繼資料(只附一次,標籤名稱不可更改):
66
+
67
+ ```xml
68
+ <result>
69
+ <status>done 或 blocked</status>
70
+ <summary>一兩句說明這次做了什麼;blocked 時說明卡在哪裡</summary>
71
+ <files_changed>
72
+ <file>每個新增或修改的檔案路徑各一行</file>
73
+ </files_changed>
74
+ <concerns>對需求、規格、計畫或測試的疑慮;沒有就留空</concerns>
75
+ </result>
76
+ ```
77
+ </reply_format>