agentflowctl 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -17,7 +17,7 @@ claude # 完成登入
17
17
  codex # 完成登入
18
18
  ```
19
19
 
20
- 沒有內建的 agent,只會使用 `flow.config.json` 的 `agents` 裡設定的。用 `agent setup` 互動設定:它會偵測本機的 `claude`、`codex`、`gemini`,逐一詢問要不要加入、名稱與 model,再設定參與的 agent;確認後才一次寫入,最後自動跑一次 `doctor`:
20
+ 沒有內建的 agent,只會使用 `flow.config.json` 的 `agents` 裡設定的。用 `agent setup` 互動設定:它會先列出已設定的 agent 與 CLI 是否可以執行,再偵測本機裝了哪些支援的 CLI(目前是 `claude`、`codex`、`gemini`),只針對已安裝的逐一詢問要不要加入、名稱與 model,再設定參與的 agent;沒偵測到的只列出、不詢問。確認後才一次寫入,最後自動跑一次 `doctor`:
21
21
 
22
22
  ```bash
23
23
  npx agentflowctl agent setup
@@ -103,7 +103,7 @@ agentflowctl status f-xxxx # 階段、上一步結果、未結交接事
103
103
  agentflowctl list
104
104
  agentflowctl logs f-xxxx # 列出每一份 log 的編號、結果、階段、步驟、agent
105
105
  agentflowctl logs f-xxxx 7 # 解析第 7 份 log,最後附上錯誤整理(--latest 看最新一份)
106
- agentflowctl logs f-xxxx 7 --full # 完整顯示多行指令與絕對路徑
106
+ agentflowctl logs f-xxxx 7 --full # 逐條顯示 shell 指令,完整顯示多行內容與絕對路徑
107
107
  agentflowctl logs f-xxxx 7 --raw # 原始內容(agent 的 JSON 行)
108
108
  agentflowctl resume f-xxxx # 從暫停、Ctrl-C 或失敗處接續
109
109
  agentflowctl cancel f-xxxx
@@ -183,14 +183,14 @@ agentflowctl clean --all # 清掉所有已結束的 run 與中斷留
183
183
  | 標記 | 內容 |
184
184
  |---|---|
185
185
  | 💬 | agent 的完整文字,不截斷 |
186
- | 🔧 | 工具呼叫。只顯示第一行,多行時註明共幾行;去掉 `/bin/zsh -lc '…'` 這類 shell 包裝,worktree 內的絕對路徑改成相對路徑 |
186
+ | 🔧 | 工具呼叫。連續的 shell 指令(多半是讀檔、搜尋)收成一行 `🔧 shell 指令 ×N`;其他工具只顯示第一行,多行時註明共幾行,worktree 內的絕對路徑改成相對路徑 |
187
187
  | 📊 | token 用量 |
188
188
  | 🏁 | 最後結果。與最後一則 💬 相同時不再重印 |
189
189
  | ⚠️ | 工具回報的錯誤。agent 通常會自己換方法繼續,所以不列進錯誤整理 |
190
190
  | ❌ | adapter 不認得的錯誤事件 |
191
191
  | 📄 | 不是 JSON 的輸出行 |
192
192
 
193
- 要看完整的工具內容與重複的最後結果,加 `--full`。adapter 不認得、也看不出錯誤跡象的 JSON 行不會顯示,只列出行數,要看全部請加 `--raw`。專案指令的 log 本來就是純文字,會原樣顯示。
193
+ 要逐條看 shell 指令、完整的工具內容與重複的最後結果,加 `--full`。adapter 不認得、也看不出錯誤跡象的 JSON 行不會顯示,只列出行數,要看全部請加 `--raw`。專案指令的 log 本來就是純文字,會原樣顯示。
194
194
 
195
195
  最後一段「錯誤」整理出結束碼、agent 回報的失敗、錯誤事件與 stderr:
196
196
 
@@ -200,7 +200,7 @@ agentflowctl clean --all # 清掉所有已結束的 run 與中斷留
200
200
  檔案 /repo/.agentflowctl/runs/f-xxxx/logs/003-plan-plan-codex.log
201
201
 
202
202
  💬 先讀 spec.md 與 acceptance.json
203
- 🔧 shell: cat .flow/spec.md
203
+ 🔧 shell 指令 ×1(--full 查看)
204
204
  🏁 失敗:stream disconnected before completion
205
205
 
206
206
  ── 錯誤 ──
@@ -300,7 +300,7 @@ verify 失敗(型別、lint、建置)一律交回最後作者。審查意見
300
300
  `agents` 與 `cycle` 也可以用 `agent` 指令修改,不必手動編輯 JSON。每次寫入前都會先驗證整份設定:
301
301
 
302
302
  ```bash
303
- agentflowctl agent setup # 互動設定 claude、codex、gemini 與參與的 agent
303
+ agentflowctl agent setup # 偵測已安裝的 agent CLI,互動設定與參與的 agent
304
304
  agentflowctl agent list # 設定的 agent、是否已安裝、是否參與
305
305
  agentflowctl agent add claude-strong --adapter claude --model opus
306
306
  agentflowctl agent add aider --adapter command -- aider --yes-always --message {prompt}
@@ -315,7 +315,7 @@ agentflowctl agent cycle claude-strong,codex,gemini # 不帶參數時顯
315
315
  - `set --adapter` 換 adapter 時,會清掉舊 adapter 的 `model`、`extraArgs`、`command`,這次有重新指定的除外。
316
316
  - `remove` 會一併從 `cycle` 移除。`cycle` 變空就刪除這個欄位,改回從 `agents` 自動偵測。
317
317
  - `--extra-arg` 可以重複指定,會整個取代原本的 `extraArgs`。參數以 `-` 開頭時,寫成 `--extra-arg=--sandbox`。
318
- - `setup` 遇到已存在的名稱會先問要不要覆寫;不覆寫時保留原設定,但仍會參與。在非互動式環境(CI、管線)裡請改用 `agent add`。`command` adapter 要自己寫指令,不在 `setup` 裡。
318
+ - `setup` 只詢問偵測到已安裝的 CLI,一個都沒有就不變更設定。遇到已存在的名稱會先問要不要覆寫;不覆寫時保留原設定,但仍會參與。在非互動式環境(CI、管線)裡請改用 `agent add`。`command` adapter 要自己寫指令,不在 `setup` 裡。
319
319
 
320
320
  已建立的 run 會沿用建立時參與的 agent,不受這些修改影響。
321
321
 
@@ -323,9 +323,11 @@ agentflowctl agent cycle claude-strong,codex,gemini # 不帶參數時顯
323
323
 
324
324
  人工確認計畫是為了擋住方向錯了還一路做下去。預設用三層機制取代它;加上 `--manual-plan` 時,三層都過了仍會停下來等你。
325
325
 
326
- **格式與覆蓋率。** 每次撰寫或修改計畫之後,都要重新通過 zod、任務相依、無循環、每條驗收條件都有任務負責。沒過就還原。
326
+ **格式與覆蓋率。** 每次撰寫或修改計畫之後,都要重新通過 zod、任務相依、無循環、每條驗收條件都有任務負責,而且每個任務最多對應兩條驗收條件(一次只做一件事,最多兩件)。沒過就還原。
327
327
 
328
- **跨模型審查。** 審查看需求覆蓋、驗收條件能不能測、任務大小與技術方向。審查者只能寫意見。若改了規格或計畫,檔案會被還原。修改者要在 `plan.md` 的「審查回應」逐條回覆;不同意要寫理由。
328
+ **跨模型審查。** 審查看需求覆蓋、驗收條件能不能測且一條只寫一個行為、任務是否只做一件事(最多兩件)與技術方向。審查者只能寫意見。若改了規格或計畫,檔案會被還原。修改者要在 `plan.md` 的「審查回應」逐條回覆;不同意要寫理由。
329
+
330
+ 計畫審查先核對需求與四份計畫交接檔,再查閱任務說明中要修改的既有檔案,有疑慮時才擴大範圍;審查紀錄只列會影響實作的問題,不逐條列已通過項目。計畫修訂先依 `feedback.md` 定位需要改的段落,修改驗收條件或任務時再檢查受影響的對應關係。檔案格式、任務對驗收條件的覆蓋與任務相依仍由程式驗證;原始需求的語意覆蓋由審查者判斷,以減少反覆讀取文件的 token 用量。
329
331
 
330
332
  **僵持時仲裁。** 兩種情況會觸發:這輪審查意見和上一輪一樣,或已達重試上限。仲裁者只判斷一件事:照這份計畫實作,能不能滿足需求。
331
333
 
@@ -346,18 +348,22 @@ agentflowctl agent cycle claude-strong,codex,gemini # 不帶參數時顯
346
348
  | 階段 | 負責的 agent | 程式認定通過的條件 | 失敗時 |
347
349
  |---|---|---|---|
348
350
  | spec | 隨機一位 | 檔案存在、zod 驗證、id 不重複 | 重試 |
349
- | plan | 與 spec 同一位 | zod、相依存在、無循環、每條驗收條件都有任務 | 重試 |
351
+ | plan | 與 spec 同一位 | zod、相依存在、無循環、每條驗收條件都有任務、每個任務最多兩條驗收條件 | 重試 |
350
352
  | plan_review | 計畫作者以外隨機挑(可多位,不重複) | 所有審查者都 `approve` | 進入 plan_fix |
351
353
  | plan_fix | 依 `fixStrategy` | 修改後仍通過 plan 的格式與 DAG 檢查 | 還原並重試 |
352
354
  | 仲裁 | 見上一節 | 一致核准;分歧依 `tieBreak` | 兩家都不核准時進入 plan_fix 再審查;第三方不核准或 `tieBreak: stop` 時失敗 |
353
355
  | 人工確認 | 你(只有 `--manual-plan`) | `agentflowctl approve` | — |
354
- | implement 紅燈 | 洗牌輪流,每家各一次 | 有測試變更,而且測試執行後失敗 | 還原並重試 |
355
- | implement 綠燈 | 測試作者以外隨機一位 | 測試檔沒有任何修改,而且測試通過 | 還原,或帶著輸出重試 |
356
+ | implement 紅燈 | 洗牌輪流,每家各一次 | 有測試變更,而且測試執行後失敗;規格與計畫檔沒有被修改 | 還原並重試 |
357
+ | implement 綠燈 | 測試作者以外隨機一位 | 測試檔與規格、計畫檔都沒有被修改,而且測試通過 | 還原,或帶著輸出重試 |
356
358
  | verify | — | `install` 與所有 `checks` 通過 | 交回作者修正 |
357
- | review | 作者以外隨機挑(可多位,不重複) | 所有審查者都 `approve` | 依 `fixStrategy` 交給他人修正 |
359
+ | fix | verify 失敗交回作者;審查意見依 `fixStrategy` | 沒有刪除測試檔,也沒有修改規格與計畫檔 | 還原並重試 |
360
+ | review | 作者以外隨機挑(可多位,不重複) | 所有審查者都 `approve`(審查者對程式碼與規格、計畫檔的修改一律還原) | 依 `fixStrategy` 交給他人修正 |
358
361
  | pr | — | push 成功;有 `gh` 就開 PR | — |
359
362
 
360
363
  驗收條件寫在 `.flow/acceptance.json`(`AC-1`…),任務寫在 `.flow/tasks.json`(`T-1`…)。
364
+ 每個任務的寫測試與寫實作 prompt 只帶入該任務對應的驗收條件;agent 優先讀任務與相關程式碼,遇到資訊不足或矛盾才查規格、計畫的相關段落。agent 可先跑相關測試,紅燈與完整測試仍由外部流程執行與判定,減少重複讀取文件和全套測試輸出所用的 token。
365
+ 計畫定案後(實作、修正、程式碼審查)不可修改 `.flow/` 裡的規格與計畫檔(`spec.md`、`acceptance.json`、`plan.md`、`tasks.json`、`tasks.ordered.json`);實作與修正時被改就還原並重試,因為寫出的程式碼可能依賴被改過的規格,必須重寫;審查者只交出審查結果,修改直接還原即可,不必重跑審查。對規格有疑慮要寫進交接事項。
366
+ 程式碼審查仍逐條核對所有驗收條件,但只在 `review.json` 列出未通過或其他重要問題;先看 diff 與相關檔案,驗收條件不清楚時才查規格。修正階段先依 `feedback.md` 定位問題並執行相關檢查,完整檢查仍由後續 verify 執行,以減少反覆讀取完整文件與測試輸出。
361
367
 
362
368
  ## Prompt 結構
363
369
 
@@ -382,7 +388,7 @@ Agent 的最後回覆要附上 XML 中繼資料:
382
388
 
383
389
  程式只在原有關卡通過後接收交接回覆,並將正式紀錄原子儲存於 `.agentflowctl/runs/<id>/handoff.json`。`action` 是需要後續處理的事項;`info` 只供參考。作者只能提出已修正並附證據,審查者才能確認結案或附理由接受。XML `<concerns>` 可以供人閱讀,但重要疑慮必須寫進交接 JSON,才能交給下一位 agent。額度代打與重試不會接收失敗呼叫的交接內容;中斷後可用 `resume` 接續。
384
390
 
385
- 計畫審查或程式碼審查若核准,但該階段仍有未結的 `action`,程式會視為互相矛盾的審查結果並重試。計畫定案和開 PR 前也會再檢查一次;`info` 會提供給目標階段閱讀,但不阻擋通關。
391
+ 計畫審查或程式碼審查若核准,但該階段仍有未結的 `action`,程式會視為互相矛盾的審查結果並重試。計畫定案和開 PR 前也會再檢查一次;`info` 會提供給目標階段閱讀,但不阻擋通關。審查結果本身也要一致:核准時 `items` 不可有未通過的項目,要求修改時至少要列一筆,否則視為格式錯誤並重新審查。
386
392
 
387
393
  ## Adapter
388
394
 
package/dist/cli.js CHANGED
@@ -312,17 +312,27 @@ agent
312
312
  });
313
313
  agent
314
314
  .command("setup")
315
- .description("互動式設定:偵測已安裝的 claude、codex、gemini,逐一選擇要不要加入並設定參與的 agent")
315
+ .description("互動式設定:偵測本機已安裝的 agent CLI,逐一選擇要不要加入並設定參與的 agent")
316
316
  .action(async () => {
317
317
  if (!stdin.isTTY)
318
318
  throw new Error("agent setup 需要互動式終端機,請改用 agent add");
319
319
  const detected = {};
320
320
  for (const a of SETUP_ADAPTERS)
321
321
  detected[a] = await probeAgent({ adapter: a, extraArgs: [] });
322
+ let configured;
323
+ try {
324
+ const cfg = loadRepoConfig();
325
+ configured = {};
326
+ for (const name of Object.keys(cfg.agents))
327
+ configured[name] = await probeAgent(resolveAgent(cfg, name));
328
+ }
329
+ catch {
330
+ // 設定檔不合法時照樣進精靈,只是不列出已設定的 agent
331
+ }
322
332
  const rl = createInterface({ input: stdin, output: stdout });
323
333
  let edit;
324
334
  try {
325
- edit = await runSetup(readRawConfig(configPath()), { ask: (q) => rl.question(q), detected, log: (l) => console.log(l) });
335
+ edit = await runSetup(readRawConfig(configPath()), { ask: (q) => rl.question(q), detected, configured, log: (l) => console.log(l) });
326
336
  }
327
337
  finally {
328
338
  rl.close();
@@ -369,7 +379,7 @@ program
369
379
  .command("logs <id> [seq]")
370
380
  .description("列出 log;指定編號(或 --latest)時顯示解析後的內容,最後附上錯誤整理")
371
381
  .option("--latest", "顯示最新一份 log", false)
372
- .option("--full", "完整顯示工具內容(多行指令、絕對路徑)與重複的最後回覆", false)
382
+ .option("--full", "逐條顯示 shell 指令,完整顯示工具內容(多行指令、絕對路徑)與重複的最後回覆", false)
373
383
  .option("--raw", "顯示原始內容(agent 的 JSON 行)", false)
374
384
  .action((id, seq, opts) => {
375
385
  mustGetRun(id);
package/dist/engine.js CHANGED
@@ -12,9 +12,9 @@ import { CMD_AGENT, nextLogFile } from "./logs.js";
12
12
  import { exec } from "./proc.js";
13
13
  import { arbiterPanel, availableAgent, fixAgent, planAgent, planFixAgent, reviewers, specAgent, taskAgents } from "./roles.js";
14
14
  import { resolveAgent, runAgent, runCommand } from "./runner.js";
15
- import { AcceptanceList, ArbiterResult, RepoConfig, ReviewResult, TaskList, } from "./schemas.js";
15
+ import { AcceptanceList, ArbiterResult, ConsistentReviewResult, RepoConfig, TaskList, } from "./schemas.js";
16
16
  import { addSubstitution, addUsage, agentRuns, saveRun } from "./store.js";
17
- import { orderTasks } from "./tasks.js";
17
+ import { orderTasks, taskAcceptance } from "./tasks.js";
18
18
  import { readJsonFile, renderPrompt, tail } from "./util.js";
19
19
  // ───────────────────────── 共用工具 ─────────────────────────
20
20
  const info = (run, msg) => console.log(`[${run.id}] ${msg}`);
@@ -35,10 +35,13 @@ const exhausted = new Set();
35
35
  function handoffTarget(run) {
36
36
  return ["spec", "plan", "plan_review", "plan_fix"].includes(run.stage) ? "plan" : "code";
37
37
  }
38
- /** 同一輪重跑使用相同 key;重試次數或 panel 位置改變時使用新 key。 */
38
+ /**
39
+ * 每次執行 agent 都用新的 key。重試次數會在後面的輪次重複出現,只靠它組 key 會撞到先前已套用的呼叫,
40
+ * 讓這次的處置被當成重播而略過,所以加上這個 run 已執行 agent 的次數。
41
+ */
39
42
  function handoffKey(run, step, slot, agent) {
40
43
  const attempts = Object.entries(run.attempts).sort(([a], [b]) => a.localeCompare(b));
41
- return JSON.stringify([run.id, run.stage, step, run.taskIndex, run.taskPhase, attempts, slot, agent]);
44
+ return JSON.stringify([run.id, run.stage, step, run.taskIndex, run.taskPhase, attempts, slot, agent, agentRuns(run.id)]);
42
45
  }
43
46
  async function agentStep(run, planned, step, prompt, mode) {
44
47
  const cfg = loadRepoConfig();
@@ -165,15 +168,17 @@ async function specStage(run) {
165
168
  return retry(run, "spec", handoffError, "spec");
166
169
  return succeed(run, "spec", "plan");
167
170
  }
168
- // ── 規格與計畫檔案:審查者與仲裁者只能讀,不能改 ──
171
+ // ── 規格與計畫檔案:計畫審查、仲裁與計畫定案後的所有階段只能讀,不能改 ──
169
172
  const PLAN_FILES = ["spec.md", "acceptance.json", "plan.md", "tasks.json"];
170
- function snapshotPlan(run) {
171
- return Object.fromEntries(PLAN_FILES.map((f) => [f, existsSync(flowFile(run, f)) ? readFileSync(flowFile(run, f), "utf8") : undefined]));
173
+ /** 計畫定案後(實作、修正、程式碼審查)另外依賴排好的任務順序,同樣不能被改 */
174
+ const LOCKED_FILES = [...PLAN_FILES, "tasks.ordered.json"];
175
+ function snapshotPlan(run, files = PLAN_FILES) {
176
+ return Object.fromEntries(files.map((f) => [f, existsSync(flowFile(run, f)) ? readFileSync(flowFile(run, f), "utf8") : undefined]));
172
177
  }
173
- /** 把被審查者動過的計畫檔案還原成審查前的內容 */
178
+ /** 把被動過的計畫檔案還原成快照的內容,回傳被改的檔名 */
174
179
  function restorePlan(run, snap) {
175
180
  const changed = [];
176
- for (const f of PLAN_FILES) {
181
+ for (const f of Object.keys(snap)) {
177
182
  const now = existsSync(flowFile(run, f)) ? readFileSync(flowFile(run, f), "utf8") : undefined;
178
183
  if (now === snap[f])
179
184
  continue;
@@ -268,7 +273,7 @@ async function planReviewStage(run) {
268
273
  info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
269
274
  if (!r.ok)
270
275
  return retry(run, "plan-review-run", `Agent 執行失敗:${r.summary}`, "plan_review");
271
- const review = readJsonFile(flowFile(run, "plan-review.json"), ReviewResult);
276
+ const review = readJsonFile(flowFile(run, "plan-review.json"), ConsistentReviewResult);
272
277
  if (!review.ok)
273
278
  return retry(run, "plan-review-run", review.error, "plan_review");
274
279
  const handoffError = finishHandoff(run, outcome, "reviewer", { target: "plan", verdict: review.data.verdict });
@@ -408,6 +413,9 @@ async function arbitratePlan(run) {
408
413
  }
409
414
  return { ...run, stage: "failed", failedStage: "plan_review", failureReason: `${summary},需要人工決定(見 plan.md 的仲裁紀錄)` };
410
415
  }
416
+ function planTamperedMessage(files) {
417
+ return `計畫定案後不可修改規格與計畫檔,已還原你的變更:${files.map((f) => `.flow/${f}`).join(", ")}。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。`;
418
+ }
411
419
  async function implementStage(run) {
412
420
  const tasks = loadOrderedTasks(run);
413
421
  const task = tasks[run.taskIndex];
@@ -421,18 +429,28 @@ async function implementStage(run) {
421
429
  const testCmd = `${cfg.install} && ${cfg.test}`;
422
430
  const progress = `${run.taskIndex + 1}/${tasks.length} ${task.id} ${task.title}`;
423
431
  const taskJson = JSON.stringify(task, null, 2);
432
+ const acceptance = readJsonFile(flowFile(run, "acceptance.json"), AcceptanceList);
433
+ if (!acceptance.ok)
434
+ throw new Error(acceptance.error);
435
+ const acceptanceJson = JSON.stringify(taskAcceptance(task, acceptance.data), null, 2);
424
436
  const agents = taskAgents(run.cycle, run.taskIndex, cfg.tddSplit, run.id);
425
437
  // ── 紅燈:只寫測試,而且測試必須失敗 ──
426
438
  if (run.taskPhase === "tests") {
427
439
  const key = `${task.id}:tests`;
428
440
  info(run, `🧪 [${progress}] 撰寫測試(${agents.tests})`);
429
441
  const before = await headCommit(repo);
430
- const outcome = await agentStep(run, agents.tests, `${task.id}-tests`, renderPrompt("implement-tests", { task: taskJson, testPattern: cfg.testPattern, testCmd }), { kind: "write", reset: () => resetTo(repo, before) });
442
+ const snap = snapshotPlan(run, LOCKED_FILES);
443
+ const outcome = await agentStep(run, agents.tests, `${task.id}-tests`, renderPrompt("implement-tests", { task: taskJson, acceptance: acceptanceJson, testPattern: cfg.testPattern, testCmd }), { kind: "write", reset: async () => { await resetTo(repo, before); restorePlan(run, snap); } });
431
444
  const { r, agent: testsAuthor } = outcome;
445
+ const tampered = restorePlan(run, snap);
432
446
  if (!r.ok) {
433
447
  await resetTo(repo, before);
434
448
  return retry(run, key, `Agent 執行失敗:${r.summary}`, "implement");
435
449
  }
450
+ if (tampered.length) {
451
+ await resetTo(repo, before);
452
+ return retry(run, key, planTamperedMessage(tampered), "implement");
453
+ }
436
454
  const commit = await commitAll(repo, `test(${task.id}): ${task.title} [${testsAuthor}]`);
437
455
  if (!commit)
438
456
  return retry(run, key, "沒有任何檔案變更,這個階段必須撰寫測試。", "implement");
@@ -462,10 +480,16 @@ async function implementStage(run) {
462
480
  throw new Error("缺少 testsCommit,狀態不一致");
463
481
  info(run, `🛠️ [${progress}] 實作(${agents.code},測試由 ${run.lastTestsAuthor ?? agents.tests} 撰寫)`);
464
482
  const redOutput = existsSync(flowFile(run, "red-output.txt")) ? readFileSync(flowFile(run, "red-output.txt"), "utf8") : "";
465
- const outcome = await agentStep(run, agents.code, `${task.id}-code`, renderPrompt("implement-code", { task: taskJson, testCmd, redOutput: tail(redOutput, 3000) }), { kind: "write", reset: () => resetTo(repo, testsCommit) });
483
+ const snap = snapshotPlan(run, LOCKED_FILES);
484
+ const outcome = await agentStep(run, agents.code, `${task.id}-code`, renderPrompt("implement-code", { task: taskJson, acceptance: acceptanceJson, testCmd, redOutput: tail(redOutput, 3000) }), { kind: "write", reset: async () => { await resetTo(repo, testsCommit); restorePlan(run, snap); } });
466
485
  const { r, agent: codeAuthor } = outcome;
486
+ const tampered = restorePlan(run, snap);
467
487
  if (!r.ok)
468
488
  return retry(run, key, `Agent 執行失敗:${r.summary}`, "implement");
489
+ if (tampered.length) {
490
+ await resetTo(repo, testsCommit);
491
+ return retry(run, key, planTamperedMessage(tampered), "implement");
492
+ }
469
493
  await commitAll(repo, `feat(${task.id}): ${task.title} [${codeAuthor}]`);
470
494
  const touched = (await changedFiles(repo, testsCommit, await headCommit(repo))).filter((f) => testRe.test(f));
471
495
  if (touched.length) {
@@ -530,13 +554,19 @@ async function fixStage(run) {
530
554
  const repo = worktreeDir(run.id);
531
555
  const feedback = readFeedback(run);
532
556
  const before = await headCommit(repo);
557
+ const snap = snapshotPlan(run, LOCKED_FILES);
533
558
  const outcome = await agentStep(run, agent, "fix", renderPrompt("fix", { testPattern: cfg.testPattern }), {
534
559
  kind: "write",
535
- reset: () => resetTo(repo, before),
560
+ reset: async () => { await resetTo(repo, before); restorePlan(run, snap); },
536
561
  });
537
562
  const { r, agent: actual } = outcome;
563
+ const tampered = restorePlan(run, snap);
538
564
  if (!r.ok)
539
565
  return retry(run, "fix", `${feedback}\n\n(上次修正時 Agent 執行失敗:${r.summary})`, "fix");
566
+ if (tampered.length) {
567
+ await resetTo(repo, before);
568
+ return retry(run, "fix", `${feedback}\n\n另外:${planTamperedMessage(tampered)}`, "fix");
569
+ }
540
570
  await commitAll(repo, `fix: ${why} [${actual}]`);
541
571
  const testRe = new RegExp(cfg.testPattern);
542
572
  const deleted = (await changedFiles(repo, before, await headCommit(repo), "D")).filter((f) => testRe.test(f));
@@ -564,14 +594,18 @@ async function reviewStage(run) {
564
594
  for (const [slot, reviewer] of panel.entries()) {
565
595
  info(run, `👀 程式碼審查(${reviewer})`);
566
596
  rmSync(flowFile(run, "review.json"), { force: true });
597
+ const snap = snapshotPlan(run, LOCKED_FILES);
567
598
  const outcome = await agentStep(run, reviewer, "review", renderPrompt("review", { reviewer, authors: authors.join("、") || "未知" }), {
568
- kind: "review", slot,
599
+ kind: "review", slot, reset: async () => { await discardChanges(repo); restorePlan(run, snap); },
569
600
  });
570
601
  const { r } = outcome;
571
602
  await discardChanges(repo); // 審查者不可改程式碼
603
+ const tampered = restorePlan(run, snap); // .flow/ 不受 git 管理,要另外還原
604
+ if (tampered.length)
605
+ info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
572
606
  if (!r.ok)
573
607
  return retry(run, "review-run", `Agent 執行失敗:${r.summary}`, "review");
574
- const review = readJsonFile(flowFile(run, "review.json"), ReviewResult);
608
+ const review = readJsonFile(flowFile(run, "review.json"), ConsistentReviewResult);
575
609
  if (!review.ok)
576
610
  return retry(run, "review-run", review.error, "review");
577
611
  const handoffError = finishHandoff(run, outcome, "reviewer", { target: "code", verdict: review.data.verdict });
package/dist/logs.js CHANGED
@@ -99,25 +99,15 @@ function findError(ev) {
99
99
  }
100
100
  /** worktree 絕對路徑改成相對路徑:`<worktree>/x` → `x`,單獨的 `<worktree>` → `.` */
101
101
  const relativeToWorktree = (s) => s.replace(/[^\s'"=]*\/\.agentflowctl\/worktrees\/[^/\s'"]+(\/)?/g, (_m, slash) => (slash ? "" : "."));
102
- /** 去掉 codex 這類 CLI 包在指令外面的 `/bin/zsh -lc '...'` */
103
- function unwrapShell(cmd) {
104
- const m = cmd.match(/^\/bin\/(?:ba|z)?sh\s+-l?c\s+(['"])([\s\S]*)\1$/);
105
- if (!m)
106
- return cmd;
107
- const [, quote, inner] = m;
108
- // 單引號包裝裡的 '"'"' 或 '\'' 是被跳脫的單引號,內容太複雜時保留原樣
109
- if (quote === "'" && inner.includes("'"))
110
- return cmd;
111
- if (quote === '"' && /(^|[^\\])"/.test(inner))
112
- return cmd;
113
- return inner;
114
- }
115
102
  /** 精簡顯示的工具內容:只留第一行,並註明原本有幾行 */
116
103
  function compactDetail(detail) {
117
- const lines = relativeToWorktree(unwrapShell(detail.trim())).split("\n");
104
+ const lines = relativeToWorktree(detail.trim()).split("\n");
118
105
  const first = clip(lines[0], 200);
119
106
  return lines.length > 1 ? `${first} …(共 ${lines.length} 行)` : first;
120
107
  }
108
+ /** 各家 CLI 執行 shell 指令的工具名稱:codex 的 shell、claude 的 Bash、gemini 的 run_shell_command */
109
+ const SHELL_TOOLS = new Set(["shell", "bash", "run_shell_command"]);
110
+ const isShellTool = (ev) => ev.kind === "tool" && SHELL_TOOLS.has(ev.name.toLowerCase());
121
111
  function renderEvent(ev, full, lastText) {
122
112
  switch (ev.kind) {
123
113
  case "text":
@@ -151,12 +141,25 @@ function analyzeLog(text, full) {
151
141
  }
152
142
  else {
153
143
  let skipped = 0;
144
+ // 精簡模式下連續的 shell 指令(多半是讀檔、搜尋)收成一行,只留數量
145
+ let shells = 0;
146
+ const flushShells = () => {
147
+ if (shells)
148
+ lines.push(`🔧 shell 指令 ×${shells}(--full 查看)`);
149
+ shells = 0;
150
+ };
154
151
  for (const line of body) {
155
152
  if (!line.trim())
156
153
  continue;
157
154
  const events = adapter.parse(line);
158
155
  if (events.length) {
159
156
  for (const ev of events) {
157
+ if (!full && isShellTool(ev)) {
158
+ shells++;
159
+ continue;
160
+ }
161
+ if (ev.kind !== "usage")
162
+ flushShells();
160
163
  const s = renderEvent(ev, full, lastText);
161
164
  if (s)
162
165
  lines.push(s);
@@ -171,6 +174,7 @@ function analyzeLog(text, full) {
171
174
  }
172
175
  const json = tryJson(line);
173
176
  if (!json) {
177
+ flushShells();
174
178
  lines.push(`📄 ${line}`);
175
179
  continue;
176
180
  }
@@ -179,10 +183,12 @@ function analyzeLog(text, full) {
179
183
  skipped++;
180
184
  continue;
181
185
  }
186
+ flushShells();
182
187
  lines.push(`${err.fatal ? "❌" : "⚠️ "} ${indent(err.message)}`);
183
188
  if (err.fatal)
184
189
  errors.push(`錯誤事件:${indent(err.message)}`);
185
190
  }
191
+ flushShells();
186
192
  if (skipped)
187
193
  lines.push(`(另有 ${skipped} 行其他事件未顯示,可用 --raw 查看)`);
188
194
  }
package/dist/schemas.js CHANGED
@@ -83,6 +83,14 @@ export const ReviewResult = z.object({
83
83
  note: z.string().default(""),
84
84
  })),
85
85
  });
86
+ /** 計畫與程式碼審查的結果:verdict 必須和 items 一致,否則視為格式錯誤重試 */
87
+ export const ConsistentReviewResult = ReviewResult.superRefine((review, ctx) => {
88
+ const unmet = review.items.filter((i) => i.status !== "met");
89
+ if (review.verdict === "approve" && unmet.length)
90
+ ctx.addIssue({ code: "custom", message: `verdict 為 approve,但 items 仍有未通過的項目:${unmet.map((i) => i.criterion).join("、")}` });
91
+ if (review.verdict === "changes_requested" && !unmet.length)
92
+ ctx.addIssue({ code: "custom", message: "verdict 為 changes_requested,但 items 沒有列出任何未通過的項目" });
93
+ });
86
94
  /** 仲裁者可能用 reject 表示否決;讀取時正規化,保留相同的理由欄位。 */
87
95
  export const ArbiterResult = ReviewResult.extend({
88
96
  verdict: z.enum(["approve", "changes_requested", "reject"])
package/dist/setup.js CHANGED
@@ -1,9 +1,11 @@
1
1
  import { addAgent, setAgent, setCycle } from "./agentConfig.js";
2
+ import { ADAPTERS } from "./agents/index.js";
2
3
  /**
3
4
  * `agentflowctl agent setup` 的互動精靈。
4
5
  * 問答與偵測結果由外部注入,這裡只負責流程;所有修改都在記憶體裡完成,確認後才交給呼叫端寫入。
5
6
  */
6
- export const SETUP_ADAPTERS = ["claude", "codex", "gemini"];
7
+ /** 不需自訂指令就能加入的 adapter;新增 adapter 時會自動出現在精靈裡 */
8
+ export const SETUP_ADAPTERS = Object.keys(ADAPTERS).filter((a) => a !== "command");
7
9
  /** 回傳要寫入的設定;使用者沒選任何 agent 或最後沒確認時回傳 null */
8
10
  export async function runSetup(initial, deps) {
9
11
  const { detected, log } = deps;
@@ -16,9 +18,26 @@ export async function runSetup(initial, deps) {
16
18
  const changes = [];
17
19
  const chosen = [];
18
20
  const existing = () => Object.keys((cfg.agents ?? {}));
19
- for (const adapter of SETUP_ADAPTERS) {
20
- log(`${detected[adapter] ? "✅" : "❌"} ${adapter}${detected[adapter] ? "" : "(沒有偵測到 CLI)"}`);
21
- if (!(await confirm(` 加入 ${adapter}?`, detected[adapter])))
21
+ const configured = Object.entries(deps.configured ?? {});
22
+ if (configured.length) {
23
+ const defs = (cfg.agents ?? {});
24
+ log("已設定的 agent:");
25
+ for (const [name, ok] of configured) {
26
+ log(` ${ok ? "✅" : "⚠️ "} ${name}(${defs[name]?.adapter ?? "?"}${ok ? "" : ",找不到可執行的 CLI"})`);
27
+ }
28
+ }
29
+ const installed = Object.keys(detected).filter((a) => detected[a]);
30
+ const missing = Object.keys(detected).filter((a) => !detected[a]);
31
+ if (missing.length)
32
+ log(`沒有偵測到:${missing.join("、")}(安裝後可重跑 agent setup)`);
33
+ if (!installed.length) {
34
+ log("沒有偵測到可加入的 agent CLI,設定沒有變更");
35
+ log("需要自訂指令的 CLI 請改用 agent add <name> --adapter command -- <指令>");
36
+ return null;
37
+ }
38
+ for (const adapter of installed) {
39
+ log(`✅ ${adapter}`);
40
+ if (!(await confirm(` 加入 ${adapter}?`, true)))
22
41
  continue;
23
42
  let name;
24
43
  for (;;) {
package/dist/tasks.js CHANGED
@@ -1,3 +1,15 @@
1
+ /** 一個任務最多做兩件事:對應的驗收條件超過這個數量就要再拆 */
2
+ export const MAX_TASK_ACCEPTANCE = 2;
3
+ /** 只取目前任務負責的驗收條件,避免每次實作都重讀整份清單。 */
4
+ export function taskAcceptance(task, acceptance) {
5
+ const byId = new Map(acceptance.map((item) => [item.id, item]));
6
+ return task.acceptance.map((id) => {
7
+ const item = byId.get(id);
8
+ if (!item)
9
+ throw new Error(`${task.id} 對應的驗收條件 ${id} 不存在`);
10
+ return item;
11
+ });
12
+ }
1
13
  /**
2
14
  * 檢查任務清單並依相依關係排序(Kahn 演算法)。
3
15
  * 回傳排序後的任務,或回傳一段可以直接回饋給 Agent 的錯誤說明。
@@ -12,6 +24,9 @@ export function orderTasks(tasks, acceptanceIds) {
12
24
  }
13
25
  const covered = new Set();
14
26
  for (const t of tasks) {
27
+ if (t.acceptance.length > MAX_TASK_ACCEPTANCE) {
28
+ errors.push(`${t.id} 對應 ${t.acceptance.length} 條驗收條件,一個任務最多 ${MAX_TASK_ACCEPTANCE} 條,請拆成更小的任務`);
29
+ }
15
30
  for (const d of t.dependsOn)
16
31
  if (!ids.has(d))
17
32
  errors.push(`${t.id} 相依的 ${d} 不存在`);
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "agentflowctl",
3
3
  "license": "MIT",
4
- "version": "0.7.0",
4
+ "version": "0.9.0",
5
5
  "description": "跨廠商 AI 開發 harness:Claude Code、Codex、Gemini 輪流實作、審查、修正",
6
6
  "keywords": [
7
7
  "ai",
package/prompts/fix.md CHANGED
@@ -18,19 +18,21 @@
18
18
  </handoff>
19
19
 
20
20
  <inputs>
21
- - .flow/feedback.md:失敗的檢查(型別、lint、測試、建置)或審查意見
22
- - .flow/spec.md 與 .flow/acceptance.json:規格
21
+ - .flow/feedback.md:先讀取,確認失敗的檢查或審查意見與其證據
22
+ - 相關程式碼與測試:依 feedback 指出的檔案和問題範圍查閱
23
+ - .flow/acceptance.json 與 .flow/spec.md:預期行為不清楚或互相矛盾時,才查相關條件或段落
23
24
  </inputs>
24
25
 
25
26
  <steps>
26
- 1. 找出每個問題的根本原因再修正,不要只針對症狀打補丁。
27
- 2. 在沙箱內自行執行相關檢查,確認問題已解決且沒有造成新的錯誤。
27
+ 1. 優先處理 feedback 與交接事項指出的問題;先定位相關程式碼和測試,找出根本原因再修正。
28
+ 2. 在沙箱內執行與修改相關的檢查,確認問題已解決。完整檢查由後續 verify 階段執行;需要判斷跨模組影響時再自行擴大檢查範圍。
28
29
  </steps>
29
30
 
30
31
  <constraints>
31
32
  - 不可刪除測試檔(檔名符合 `{{testPattern}}`),也不可用 skip、放寬斷言、`@ts-ignore`、`eslint-disable` 等方式讓檢查通過。
32
33
  - 如果測試本身確實有誤,可以修正測試,但必須在回覆的 `<concerns>` 說明理由。
33
34
  - 不要執行 git commit(權限設定已禁止)。
35
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
34
36
  </constraints>
35
37
 
36
38
  <reply_format>
@@ -23,8 +23,15 @@
23
23
  ```
24
24
  </task>
25
25
 
26
+ <acceptance>
27
+ 這個任務負責的驗收條件:
28
+ ```json
29
+ {{acceptance}}
30
+ ```
31
+ </acceptance>
32
+
26
33
  <inputs>
27
- 完整規格與計畫請參考 .flow/spec.md、.flow/plan.md。
34
+ 先依上面的任務、驗收條件與現有測試工作。只有資訊不足或互相矛盾時,再閱讀 .flow/spec.md、.flow/plan.md 的相關段落,並在交接中指出問題;不必通讀整份文件。
28
35
  </inputs>
29
36
 
30
37
  <red_output>
@@ -36,13 +43,14 @@
36
43
  <steps>
37
44
  1. 若 .flow/feedback.md 存在,先閱讀,並依內容修正。
38
45
  2. 撰寫讓測試通過的最小實作,符合專案既有的程式風格與架構。
39
- 3. 執行 `{{testCmd}}` 確認全部測試通過(包含既有測試)。
46
+ 3. 先執行本任務相關的測試,確認實作方向。外部流程會再執行 `{{testCmd}}` 驗證全部測試(包含既有測試),不需要自行重跑全套測試。
40
47
  4. 測試通過後,在不改變行為的前提下整理程式碼。
41
48
  </steps>
42
49
 
43
50
  <constraints>
44
51
  - **不可修改任何測試檔**,修改會被自動還原並視為失敗。若認為測試本身有誤,請寫在回覆的 `<concerns>`。
45
52
  - 不要執行 git commit(權限設定已禁止)。
53
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
46
54
  </constraints>
47
55
 
48
56
  <reply_format>
@@ -23,20 +23,28 @@
23
23
  ```
24
24
  </task>
25
25
 
26
+ <acceptance>
27
+ 這個任務負責的驗收條件:
28
+ ```json
29
+ {{acceptance}}
30
+ ```
31
+ </acceptance>
32
+
26
33
  <inputs>
27
- 完整規格與計畫請參考 .flow/spec.md、.flow/plan.md。
34
+ 先依上面的任務與驗收條件工作。只有資訊不足或互相矛盾時,再閱讀 .flow/spec.md、.flow/plan.md 的相關段落,並在交接中指出問題;不必通讀整份文件。
28
35
  </inputs>
29
36
 
30
37
  <steps>
31
38
  1. 若 .flow/feedback.md 存在,先閱讀,並依內容調整做法。
32
39
  2. 依任務描述撰寫測試,檔名必須符合正規表示式 `{{testPattern}}`。
33
40
  3. 測試必須驗證這個任務要新增的行為,並且因為功能尚未實作而**失敗**。
34
- 4. 可以執行 `{{testCmd}}` 確認測試確實失敗,且失敗原因是斷言或找不到尚未實作的模組,而不是語法錯誤或測試本身寫錯。
41
+ 4. 可以先執行本任務相關的測試,確認失敗原因是斷言或找不到尚未實作的模組,而不是語法錯誤或測試本身寫錯。外部流程會再執行 `{{testCmd}}` 驗證紅燈,不需要自行重跑全套測試。
35
42
  </steps>
36
43
 
37
44
  <constraints>
38
45
  - 不可實作功能本身。可以建立讓測試能編譯所需的最小型別或空殼匯出,但不可以有真正的邏輯。
39
46
  - 不要執行 git commit(權限設定已禁止),外部流程會提交並驗證測試是否失敗。
47
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
40
48
  </constraints>
41
49
 
42
50
  <reply_format>
@@ -22,15 +22,15 @@
22
22
  </requirement>
23
23
 
24
24
  <steps>
25
- 1. 閱讀 .flow/feedback.md,裡面是其他模型的審查意見。
26
- 2. 閱讀目前的 .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json 與相關程式碼。
27
- 3. 逐條處理審查意見,直接修改上述四個檔案。
28
- 4. 在 .flow/plan.md 最後的「## 審查回應」一節,逐條說明每個意見怎麼處理;不同意的意見,請寫出具體理由,而不是忽略它。回應時只談內容,不要提到審查者或你自己是哪個模型、哪家公司,之後可能由第三方匿名仲裁。
25
+ 1. 先閱讀 .flow/feedback.md,依每則意見定位 .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json 中相關的段落或項目;技術細節需要確認時才讀相關程式碼。
26
+ 2. 逐條處理審查意見,只修改需要修訂的檔案。新增或調整驗收條件、任務時,檢查受影響的條件與任務對應。
27
+ 3. 在 .flow/plan.md 最後的「## 審查回應」一節,逐條說明每個意見怎麼處理;不同意的意見,請寫出具體理由,而不是忽略它。回應時只談內容,不要提到審查者或你自己是哪個模型、哪家公司,之後可能由第三方匿名仲裁。
29
28
  </steps>
30
29
 
31
30
  <output_format>
32
31
  - acceptance.json 與 tasks.json 的格式必須維持不變(見檔案內現有內容)。
33
32
  - 每一條驗收條件都至少要有一個任務負責;`dependsOn` 不可有循環。
33
+ - 一個任務只做一件事,最多兩件:`acceptance` 最多列兩條驗收條件;驗收條件一條只描述一個行為。修改時若任務變大,請拆開,不要合併。
34
34
  - 測試檔名必須符合正規表示式 `{{testPattern}}`。
35
35
  </output_format>
36
36
 
@@ -22,21 +22,18 @@
22
22
  </requirement>
23
23
 
24
24
  <inputs>
25
- - .flow/spec.md:規格
26
- - .flow/acceptance.json:驗收條件
27
- - .flow/plan.md:實作方式
28
- - .flow/tasks.json:任務拆解
29
-
30
- 請同時閱讀相關的既有程式碼,確認計畫符合專案的實際架構。
25
+ - 先核對原始需求、.flow/spec.md、.flow/acceptance.json、.flow/plan.md 與 .flow/tasks.json,確認需求、驗收條件與任務的對應。
26
+ - 依 .flow/tasks.json 各任務 `description` 列出要動的檔案,查閱其中既有的檔案,確認計畫符合專案的架構與慣例;有疑慮時再擴大查閱範圍。
31
27
  </inputs>
32
28
 
33
29
  <review_focus>
34
30
  1. **需求覆蓋**:規格是否完整涵蓋原始需求?有沒有遺漏、誤解,或加入需求沒要求的範圍?
35
- 2. **驗收條件**:每一條是否具體、可以用自動化測試驗證?有沒有重要的邊界情況或錯誤處理沒被列入?
36
- 3. **任務拆解**:每個任務是否小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試?相依順序是否合理?
31
+ 2. **驗收條件**:每一條是否具體、可以用自動化測試驗證,而且只描述一個行為?把多個行為寫在同一條的,要求拆開。有沒有重要的邊界情況或錯誤處理沒被列入?
32
+ 3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?
37
33
  4. **技術方向**:是否符合專案既有的架構與慣例?有沒有更簡單的做法,或明顯的風險?
38
34
 
39
35
  措辭、格式這類不影響實作結果的小問題,不需要要求修改。
36
+ 只在證據不足時擴大查閱範圍;檔案格式、任務對驗收條件的覆蓋與任務相依已有程式關卡檢查,不必自行重跑完整檢查。仍須自行判斷原始需求是否被規格與驗收條件涵蓋。
40
37
  </review_focus>
41
38
 
42
39
  <output_format>
@@ -44,15 +41,16 @@
44
41
 
45
42
  ```json
46
43
  {
47
- "verdict": "approve",
44
+ "verdict": "changes_requested",
48
45
  "items": [
49
- { "criterion": "需求覆蓋", "status": "met", "note": "" },
50
46
  { "criterion": "AC-2", "status": "not_met", "note": "沒有涵蓋 API 逾時的情況,建議新增一條驗收條件並由 T-3 負責" }
51
47
  ]
52
48
  }
53
49
  ```
54
50
 
55
51
  - `verdict`:沒有會影響實作結果的問題時為 `approve`,否則為 `changes_requested`。
52
+ - `items` 只列會影響實作結果的問題,每筆 `status` 為 `not_met` 或 `partial`;已通過的項目不必逐條記錄。
53
+ - `approve` 時 `items` 為空陣列;`changes_requested` 時至少要有一筆。兩者不一致會被視為格式錯誤並重新審查。
56
54
  - 每個問題的 `note` 請寫出具體要改哪個檔案的哪個部分,以及建議怎麼改。
57
55
  </output_format>
58
56
 
package/prompts/plan.md CHANGED
@@ -46,7 +46,10 @@
46
46
  </output_format>
47
47
 
48
48
  <guidelines>
49
- - 每個任務是一個可獨立測試的垂直切片,小到一次 TDD 循環就能完成。
49
+ - 一個任務只做一件事,最多兩件:`acceptance` 最多列兩條驗收條件,超過就拆成多個任務(程式會檢查,超過會被退回)。
50
+ - 每個任務是一個可獨立測試的垂直切片,小到一次 TDD 循環就能完成;只動少數幾個檔案,測試只驗證一兩個行為。
51
+ - `title` 用一句話說出這件事;需要用「並且」「以及」串起來的,就是兩個任務。
52
+ - `description` 寫清楚要動哪些檔案、測試要驗證哪個行為,以及這個任務不做什麼。
50
53
  - 每個任務都必須能寫出「在實作前會失敗」的測試;純設定或重構類工作請併入相關任務。
51
54
  - 測試檔名必須符合正規表示式 `{{testPattern}}`。
52
55
  - 每一條驗收條件都至少要有一個任務負責;`dependsOn` 不可有循環。
package/prompts/review.md CHANGED
@@ -18,9 +18,9 @@
18
18
  </handoff>
19
19
 
20
20
  <inputs>
21
- - 規格:.flow/spec.md
22
- - 驗收條件:.flow/acceptance.json
23
- - 本次變更:.flow/diff.patch
21
+ - 驗收條件:.flow/acceptance.json(逐條核對)
22
+ - 本次變更:.flow/diff.patch(先看變更,再按需讀相關程式碼與測試)
23
+ - 規格:.flow/spec.md(驗收條件不清楚或互相矛盾時,才查相關段落)
24
24
  - 自動化檢查結果:.flow/verify.json(已全部通過)
25
25
  </inputs>
26
26
 
@@ -29,6 +29,8 @@
29
29
  2. 是否有明顯的錯誤、邊界情況遺漏、安全問題或效能問題。
30
30
  3. 是否符合專案既有的架構與慣例。
31
31
 
32
+ 先根據 diff 與驗收條件定位需要查閱的檔案;只在證據不足時讀取其他檔案。自動化檢查已由外部流程執行,不必為了審查重跑全套檢查。
33
+
32
34
  風格偏好與無關緊要的小問題不需要要求修改。
33
35
  </review_focus>
34
36
 
@@ -37,16 +39,16 @@
37
39
 
38
40
  ```json
39
41
  {
40
- "verdict": "approve",
42
+ "verdict": "changes_requested",
41
43
  "items": [
42
- { "criterion": "AC-1", "status": "met", "note": "" },
43
- { "criterion": "錯誤處理", "status": "not_met", "note": "API 失敗時沒有顯示錯誤訊息,見 src/form.tsx" }
44
+ { "criterion": "AC-1", "status": "not_met", "note": "src/form.tsx 缺少 API 失敗時的錯誤訊息,請在 catch 中顯示錯誤並補測試" }
44
45
  ]
45
46
  }
46
47
  ```
47
48
 
48
- - `verdict`:全部驗收條件都 `met` 且沒有嚴重問題時為 `approve`,否則為 `changes_requested`。
49
- - 每個驗收條件都要有一筆;額外發現的問題也各自列一筆,`note` 請寫出具體位置與修正方向。
49
+ - `verdict`:逐條核對所有驗收條件後,全部通過且沒有嚴重問題時為 `approve`,否則為 `changes_requested`。
50
+ - `items` 只列未通過的驗收條件與額外發現的重要問題,每筆的 `status` 為 `not_met` 或 `partial`,`note` 請寫出具體位置與修正方向。
51
+ - `approve` 時 `items` 為空陣列;`changes_requested` 時至少要有一筆。兩者不一致會被視為格式錯誤並重新審查。
50
52
  </output_format>
51
53
 
52
54
  <constraints>
package/prompts/spec.md CHANGED
@@ -28,6 +28,13 @@
28
28
  4. 撰寫 .flow/acceptance.json,每一條驗收條件都必須能用自動化測試驗證。
29
29
  </steps>
30
30
 
31
+ <guidelines>
32
+ - 驗收條件要寫細:一條只描述一個可觀察的行為(一個輸入或情境,對應一個預期結果)。
33
+ - 描述裡出現「並且」「同時」「以及」,或同時涵蓋成功與失敗路徑時,拆成多條。
34
+ - 邊界情況與錯誤處理各自獨立成一條,不要附在正常路徑那一條裡。
35
+ - 寧可多幾條小的,也不要少數幾條大的;後面每個任務最多只能對應兩條驗收條件。
36
+ </guidelines>
37
+
31
38
  <output_format>
32
39
  .flow/acceptance.json 的格式:
33
40