agentflowctl 0.16.0 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -33,7 +33,7 @@ npx agentflowctl run --req-file ./requirement.md
33
33
  ## 執行時會發生什麼
34
34
 
35
35
  1. agent 整理需求與驗收條件,接著寫計畫,交給其他 agent 審查。
36
- 2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent 實作到測試通過;每個任務都會經過審查與驗證。計畫 agent 會依改動內容在任務標記 `tdd`:建置流程、設定、文件、型別、純重構這類不適合先寫失敗測試的任務,會略過紅綠燈直接實作,改由任務審查與驗證把關。專案沒有測試框架時,所有任務都略過紅綠燈,也不跑 `test` 檢查。
36
+ 2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent 實作到測試通過;每個任務都會經過審查與驗證。計畫 agent 會依改動內容在任務標記 `tdd`:建置流程、設定、文件、型別、純重構,以及實作前就會通過的特徵化測試,會略過紅綠燈直接實作,改由任務審查與驗證把關。描述寫明不要求紅燈卻沒標 `tdd: false` 的計畫不會通過。已定案的計畫若仍帶著這個衝突,執行到該任務時仍會寫測試,但測試一開始就通過也算完成,不會再要求紅燈、也不會因此讓 run 失敗。專案沒有測試框架時,所有任務都略過紅綠燈,也不跑 `test` 檢查。
37
37
  3. 全部任務完成後,再執行專案檢查與整體程式碼審查。未通過的項目會交回修正。
38
38
  4. 有 `origin` 時會推送分支;若 `gh` 可用,會嘗試建立 PR。沒有 `origin` 時,完成的分支留在本機。
39
39
 
@@ -198,6 +198,7 @@ Codex 另有幾點差異:
198
198
  | `tddSplit` | `true` | 有多位 agent 時,`true` 會把同一任務的測試與實作分給不同 agent |
199
199
  | `reviewQuorum` | `1` | 任務與最終程式碼審查需要幾位不同審查者核准 |
200
200
  | `planReviewQuorum` | `1` | 計畫需要幾位不同審查者核准 |
201
+ | `reviewConcurrency` | 不限 | 同一輪審查最多幾位審查者同時執行;`1` 為一次一位;說明見表格下方,`doctor` 會顯示目前的設定 |
201
202
  | `planArbiter` | `true` | 計畫審查僵持,或修訂一次後仍被要求修改時是否啟用仲裁 |
202
203
  | `planReviewLayers` | `{ "enabled": true, "minTasks": 7, "maxGroups": 5, "tasksPerGroup": 3 }` | 任務夠多時把計畫審查拆成索引與任務群;說明見表格下方 |
203
204
  | `tieBreak` | `"proceed"` | 兩位仲裁者意見分歧時,`"proceed"` 繼續、`"stop"` 停止 |
@@ -210,6 +211,34 @@ Codex 另有幾點差異:
210
211
 
211
212
  計畫審查會依任務規模選做法。同時符合下列條件時,每輪先做一次索引審查,再只審查有變動的任務群:任務達到 `planReviewLayers.minTasks` 個;依 description 寫的檔案路徑能分成至少兩群,而且最大一群不超過三分之二;`plan.md` 每個任務都有 `## T-<數字>` 標題。索引審查讀規格、全部任務描述、驗收條件與整體做法,人數是 `planReviewQuorum`。群數最多 `maxGroups`,也不超過任務數除以 `tasksPerGroup`;每群一位審查者,含 `high` 任務的群改由 `planReviewQuorum` 位審查。改了 `plan.md` 的整體做法時所有群都重審;某一次審查失敗時只重跑還沒完成的部分。已達門檻卻不符其他條件時,終端機會印出原因並改由審查者讀完整份規格與計畫。`"planReviewLayers": { "enabled": false }` 可以關閉,`doctor` 會顯示目前的設定。
212
213
 
214
+ #### 平行審查
215
+
216
+ 同一輪的審查者(整份計畫審查、分層計畫審查的索引與各群、程式碼審查)預設同時執行,可用 `reviewConcurrency` 限制同時數量(`1` 為一次一位)。
217
+
218
+ **審查者的環境**
219
+
220
+ - 每位審查者在自己的臨時 git worktree 裡工作,固定放在 `.agentflowctl/runs/<id>/tmp-review/slot-<N>/`;只有該路徑被占用(例如上次中斷的殘骸)時才改用唯一的子目錄。
221
+ - 看到的是 run 目前的 HEAD、`.flow/` 的複本,以及指向 run worktree 頂層 `node_modules` 的 symlink(有才建);不含其他未 commit 或被 gitignore 的檔案(例如 `dist/`)。所以審查者改了什麼都不會影響 run 的 worktree 或其他審查者。
222
+ - 唯一的例外是 `node_modules`:它是共用的 symlink,審查者若在臨時 worktree 裡跑安裝,會寫到 run 真正的 `node_modules`。
223
+ - 程式碼審查開始前,會先把 run 的 worktree 還原成 HEAD(清掉驗證階段留下的未 commit 修改與未追蹤產物)。
224
+
225
+ **結果如何套用**
226
+
227
+ - 審查結果一完成就存進 `.agentflowctl/runs/<id>/parallel-review/`,全部審查者跑完後才依固定順序逐一套用結果與交接事項。
228
+ - 因此即使 `reviewConcurrency` 為 `1`,同一輪的審查者也看不到彼此本輪新增或結掉的交接事項,核准與否也依它開始時看到的帳本判斷。例如一位審查者結掉了某個未結事項,另一位核准卻沒有結掉它,後者會被判交接不合格而多重跑一次(重跑時就看得到該事項已結)。
229
+ - 同時執行時,引擎印出的訊息會加上 `[審查者]` 前綴;agent 自己的輸出(`-v` 時印出的內容、回報的疑慮)不加。
230
+
231
+ **中斷與 resume**
232
+
233
+ - 任何時候 Ctrl-C 或被中斷,之後 `resume` 只會補跑沒完成的審查者,已完成的不重跑。
234
+ - 已存檔的審查只在同一輪、計畫或程式碼沒變,而且交接帳本沒有被這一輪以外的呼叫(例如修正者)動過時沿用,否則重審。
235
+ - 中斷留下的臨時 worktree(`.agentflowctl/runs/<id>/tmp-review/`)在該 run 下次 `resume`(或 `clean`)時自動清掉。
236
+
237
+ **額度與預算**
238
+
239
+ - 若有審查者的額度用完,其他審查者仍會跑完並存檔,全部結束後才暫停;但同一家 agent 還沒啟動的審查者也會略過(這家在這次執行中已確認額度用完),`resume` 時再和額度用完的那位一起補跑。
240
+ - 審查只在 `maxAgentRuns` 剩餘的次數內啟動:預算不夠整輪時只啟動預算內的審查者(結果照樣存檔),run 以 `agent_budget` 失敗,用 `resume <id> --max-agent-runs <次數>` 調高後只補跑沒跑的。
241
+
213
242
  審查意見的處理寫在 `.flow/plan-replies.md`,每輪覆寫,不寫進 `plan.md` 文末。下一輪索引會看到整份回應;任務群只看到自己的 `## T-<數字>` 節。
214
243
 
215
244
  `install`、`test`、`checks` 未設定時,會依 `packageManager`、lockfile 和 `package.json` scripts 偵測。完整範例見 [examples/flow.config.json](examples/flow.config.json)。專案設定每一步都會重新讀取,但已建立 run 的參與 agent 與執行次數上限會沿用建立時的值;要調高後者請用 `resume --max-agent-runs`。
@@ -226,7 +255,7 @@ pnpm release minor # 確認後建立 v0.x+1.0 的 GitHub Release
226
255
  pnpm release 1.0.0 # 指定版本,必須大於目前最新的 tag
227
256
  ```
228
257
 
229
- 腳本會先檢查目前在 `main`、工作區乾淨、與 `origin/main` 同步、`gh` 已登入、新 tag 不存在,再跑 typecheck、test、build(通過只顯示 ✓,失敗才印出完整輸出),列出自上個 tag 以來的 commit 並等你輸入 `y` 確認,然後用 `gh release create --generate-notes` 建立 Release。npm 由 Release 觸發的 `npm-publish.yml` 發布;版本號取自 tag,腳本不會改 `package.json`。
258
+ 腳本會先檢查目前在 `main`、工作區乾淨、與 `origin/main` 同步、`gh` 已登入、新 tag 不存在,再跑 typecheck、test、build(通過只顯示 ✓,失敗才印出完整輸出),列出自上個 tag 以來的 commit 並等你輸入 `y` 確認,接著把 `package.json` 的 `version` 更新為新版號、commit 並 push 到 `main`,再用 `gh release create --generate-notes` 建立 Release。npm 由 Release 觸發的 `npm-publish.yml` 發布。
230
259
 
231
260
  ## 更多文件
232
261
 
@@ -1,4 +1,4 @@
1
- import { mkdirSync, writeFileSync } from "node:fs";
1
+ import { existsSync, mkdirSync, readFileSync, renameSync, writeFileSync } from "node:fs";
2
2
  import { join } from "node:path";
3
3
  import { num, str, toolDetail, tryJson } from "./types.js";
4
4
  /**
@@ -19,7 +19,13 @@ function settingsFile(o) {
19
19
  };
20
20
  mkdirSync(o.runDir, { recursive: true });
21
21
  const path = join(o.runDir, "claude-settings.json");
22
- writeFileSync(path, JSON.stringify(settings, null, 2));
22
+ const content = JSON.stringify(settings, null, 2);
23
+ // 平行審查時多個 claude 同時啟動:內容沒變就不寫;要寫時先寫暫存檔再 rename,別的行程不會讀到被截斷的檔案而在沒有 deny list 下執行
24
+ if (existsSync(path) && readFileSync(path, "utf8") === content)
25
+ return path;
26
+ const tmp = `${path}.${process.pid}.${Date.now()}.tmp`;
27
+ writeFileSync(tmp, content);
28
+ renameSync(tmp, path);
23
29
  return path;
24
30
  }
25
31
  export const claude = {
package/dist/cleanup.js CHANGED
@@ -2,6 +2,7 @@ import { existsSync, readdirSync, rmSync } from "node:fs";
2
2
  import { git, removeWorktree } from "./git.js";
3
3
  import { projectRoot, runDir, runsDir, worktreeDir, worktreesDir } from "./paths.js";
4
4
  import { getRun } from "./store.js";
5
+ import { cleanupTempWorktrees } from "./tempWorktree.js";
5
6
  /**
6
7
  * 移除一個 run 的 worktree 與紀錄,分支保留。
7
8
  * 不需要 state.json:中斷在建立 worktree 之後、寫入紀錄之前留下的孤兒也能清。
@@ -14,6 +15,8 @@ export async function cleanRun(id) {
14
15
  const root = projectRoot();
15
16
  const wt = worktreeDir(id);
16
17
  const found = existsSync(wt) || existsSync(runDir(id));
18
+ // 平行審查中斷留下的臨時 worktree:locked 的登記 prune 不會清,要先處理;tmp-review/ 已不在時也要查
19
+ await cleanupTempWorktrees(id, { force: true });
17
20
  if (existsSync(wt)) {
18
21
  // git 不認得這個資料夾時(登記已被 prune、或 worktree add 做到一半)改成直接刪
19
22
  await removeWorktree(root, wt).catch(() => rmSync(wt, { recursive: true, force: true }));
package/dist/cli.js CHANGED
@@ -544,6 +544,7 @@ async function doctor() {
544
544
  console.log(`\n單一 run 的 agent 執行上限:${cfg.maxAgentRuns} 次`);
545
545
  console.log(`修正策略:${cfg.fixStrategy} 測試與實作分開:${cfg.tddSplit ? "是" : "否"}`);
546
546
  console.log(`程式碼審查人數:${cfg.reviewQuorum} 計畫審查人數:${cfg.planReviewQuorum} 計畫仲裁:${cfg.planArbiter ? "開啟" : "關閉"}`);
547
+ console.log(`同一輪審查者同時執行上限:${cfg.reviewConcurrency ?? "不限"}`);
547
548
  const layers = cfg.planReviewLayers;
548
549
  console.log(`計畫分層審查:${layers.enabled ? `任務達 ${layers.minTasks} 個時開啟,最多 ${layers.maxGroups} 群,每群平均至少 ${layers.tasksPerGroup} 個任務` : "關閉"}`);
549
550
  }
package/dist/engine.js CHANGED
@@ -5,25 +5,27 @@ import { arbitrationDecision } from "./arbitration.js";
5
5
  import { detectProjectDefaults, usesTestFramework, withProjectDefaults } from "./detect.js";
6
6
  import { escapeXml, opinion, reviewIssue } from "./feedback.js";
7
7
  import { changedFiles, commitAll, discardChanges, git, headCommit, resetTo } from "./git.js";
8
- import { acceptHandoff, openActions, prepareHandoff, previewHandoff, readHandoff, recoverHandoff, reviewHandoffGate, validateHandoffResponse } from "./handoff.js";
8
+ import { acceptHandoff, openActions, prepareHandoff, previewHandoff, readHandoff, recoverHandoff, responsePath, reviewHandoffGate, validateHandoffResponse } from "./handoff.js";
9
9
  import { flowDir, logDir, planArbitrationPath, planReviewStatePath, projectRoot, runDir, worktreeDir } from "./paths.js";
10
10
  import { CMD_AGENT, nextLogFile } from "./logs.js";
11
11
  import { exec } from "./proc.js";
12
12
  import { arbiterPanel, availableAgent, fixAgent, planAgent, planFixAgent, reviewers, specAgent, taskAgents } from "./roles.js";
13
+ import { dropCall, loadCalls, openRound, runPool, saveCall, storedCallValid } from "./parallelReview.js";
14
+ import { cleanupTempWorktrees, withTempWorktree } from "./tempWorktree.js";
13
15
  import { resolveAgent, runAgent, runCommand } from "./runner.js";
14
16
  import { AcceptanceList, ArbiterResult, ConsistentReviewResult, RepoConfig, TaskList, } from "./schemas.js";
15
17
  import { addRetry, addSubstitution, addUsage, agentRuns, saveRun } from "./store.js";
16
18
  import { clearModelReviewFailure, clearModelReviewStage, recordModelReviewFailure, selectModel } from "./modelSelection.js";
17
19
  import { applyReviewVerdicts, dirtyGroups, dirtyTaskIds, extractPlanEvidence, groupReviewerCount, layeredReview, neighborTasks, planContentKey, planOverview, planReviewIndex, readPendingArbitration, readPlanReviewState, repliesForTasks, reviewFingerprint, roundProgress, } from "./planReview.js";
18
- import { orderTasks, taskAcceptance, validateTaskComplexity } from "./tasks.js";
20
+ import { descriptionWaivesRed, orderTasks, taskAcceptance, validateTaskComplexity, validateTddFlag } from "./tasks.js";
19
21
  import { readJsonFile, renderPrompt, tail } from "./util.js";
20
22
  // ───────────────────────── 共用工具 ─────────────────────────
21
23
  const info = (run, msg) => console.log(`[${run.id}] ${msg}`);
22
24
  const logHint = (run, seq) => `(agentflowctl logs ${run.id} ${seq})`;
23
25
  const flowFile = (run, name) => join(flowDir(run.id), name);
24
26
  const to = (run, stage) => ({ ...run, stage });
25
- function target(run, step, agent) {
26
- return { runId: run.id, cwd: worktreeDir(run.id), logFile: nextLogFile(logDir(run.id), run.stage, step, agent), stage: run.stage, step };
27
+ function target(run, step, agent, cwd = worktreeDir(run.id)) {
28
+ return { runId: run.id, cwd, logFile: nextLogFile(logDir(run.id), run.stage, step, agent), stage: run.stage, step };
27
29
  }
28
30
  /** 額度用完而必須停下:審查類步驟,或所有 agent 的額度都用完 */
29
31
  export class QuotaPause extends Error {
@@ -50,7 +52,8 @@ function handoffKey(run, step, slot, agent) {
50
52
  }
51
53
  async function agentStep(run, planned, step, prompt, mode) {
52
54
  const cfg = loadRepoConfig();
53
- const reset = mode.reset ?? (() => discardChanges(worktreeDir(run.id)));
55
+ const say = (msg) => info(run, mode.tag ? `[${mode.tag}] ${msg}` : msg);
56
+ const reset = mode.reset ?? (mode.workspace ? () => { } : () => discardChanges(worktreeDir(run.id)));
54
57
  let agent = planned;
55
58
  for (;;) {
56
59
  if (exhausted.has(agent)) {
@@ -62,16 +65,16 @@ async function agentStep(run, planned, step, prompt, mode) {
62
65
  throw new QuotaPause(`所有 agent 的額度都已用完(${[...exhausted].join("、")})`);
63
66
  const note = step.endsWith("-code") && sub === run.lastTestsAuthor ? "測試與實作由同一家負責" : undefined;
64
67
  addSubstitution(run.id, { step, planned, actual: sub, note });
65
- info(run, `🔁 ${agent} 額度已用完,${step} 由 ${sub} 代打${note ? `(注意:${note})` : ""}`);
68
+ say(`🔁 ${agent} 額度已用完,${step} 由 ${sub} 代打${note ? `(注意:${note})` : ""}`);
66
69
  agent = sub;
67
70
  }
68
71
  const selected = selectModel(run, cfg, agent, step, /^T-\d+-/.test(step) ? loadOrderedTasks(run)[run.taskIndex]?.complexity : undefined, mode.kind === "review" ? agent : undefined, mode.modelScope);
69
- info(run, `🤖 ${step}:${agent} 使用 ${selected.name ?? "CLI 預設(名稱未知)"}${selected.insufficient ? `(低於目標 ${selected.targetStrength})` : ""}`);
72
+ say(`🤖 ${step}:${agent} 使用 ${selected.name ?? "CLI 預設(名稱未知)"}${selected.insufficient ? `(低於目標 ${selected.targetStrength})` : ""}`);
70
73
  const callKey = handoffKey(run, step, mode.slot ?? 0, agent);
71
- prepareHandoff(run.id, callKey, handoffTarget(run), mode.blind ?? false);
72
- const r = await runAgent(agent, { ...resolveAgent(cfg, agent), model: selected.name }, { ...target(run, step, agent), strength: selected.strength, targetStrength: selected.targetStrength }, prompt);
74
+ prepareHandoff(run.id, callKey, handoffTarget(run), mode.blind ?? false, mode.workspace?.flow);
75
+ const r = await runAgent(agent, { ...resolveAgent(cfg, agent), model: selected.name }, { ...target(run, step, agent, mode.workspace?.dir), strength: selected.strength, targetStrength: selected.targetStrength }, prompt);
73
76
  if (r.resolvedModel && r.resolvedModel !== selected.name)
74
- info(run, ` ↳ CLI 回報實際模型:${r.resolvedModel}`);
77
+ say(` ↳ CLI 回報實際模型:${r.resolvedModel}`);
75
78
  addUsage(run.id, { stage: step, agent, model: selected.name, resolvedModel: r.resolvedModel,
76
79
  strength: selected.strength, targetStrength: selected.targetStrength, usageReported: r.usageReported,
77
80
  inputTokens: r.inputTokens, outputTokens: r.outputTokens, cacheReadTokens: r.cacheReadTokens, cacheWriteTokens: r.cacheWriteTokens });
@@ -79,14 +82,16 @@ async function agentStep(run, planned, step, prompt, mode) {
79
82
  reportMeta(run, agent, r);
80
83
  return { r, agent, step, callKey };
81
84
  }
82
- info(run, `⛽ ${agent} 的額度已用完`);
85
+ say(`⛽ ${agent} 的額度已用完`);
83
86
  exhausted.add(agent);
84
87
  await reset();
85
88
  // 迴圈回到開頭:review 會停下,write 會找代打
86
89
  }
87
90
  }
88
91
  /** 原有關卡已通過後才接受交接;失敗回覆不進入正式紀錄。 */
89
- function finishHandoff(run, outcome, role, gate) {
92
+ function finishHandoff(run, outcome, role, gate,
93
+ /** 平行審查:關卡以這個呼叫啟動時看到的帳本判斷(存檔裡的 base),套用仍落在目前的帳本 */
94
+ base) {
90
95
  const parsed = validateHandoffResponse(run.id);
91
96
  if (!parsed.ok)
92
97
  return parsed.error;
@@ -94,7 +99,7 @@ function finishHandoff(run, outcome, role, gate) {
94
99
  const source = {
95
100
  stage: run.stage, step: outcome.step, agent: outcome.agent, callKey: outcome.callKey,
96
101
  };
97
- const preview = previewHandoff(readHandoff(run.id), outcome.callKey, source, parsed.data, role);
102
+ const preview = previewHandoff(base ?? readHandoff(run.id), outcome.callKey, source, parsed.data, role);
98
103
  if (gate) {
99
104
  const error = reviewHandoffGate(preview, gate.target, gate.verdict);
100
105
  if (error)
@@ -199,6 +204,27 @@ function hasTestFramework() {
199
204
  function taskUsesTdd(task, framework) {
200
205
  return framework && task.tdd !== false;
201
206
  }
207
+ /** 已定案的任務寫明不要求紅燈:仍寫測試,但通過也算完成紅燈階段。 */
208
+ function testsRedGuidance(testCmd, waiveRed) {
209
+ if (!waiveRed) {
210
+ return {
211
+ roleGoal: "你的測試要精準描述任務要新增的行為,並且在功能實作前確實失敗;之後會由另一位工程師實作到通過,而且對方不能修改你的測試。",
212
+ redGuidance: [
213
+ "3. 測試必須驗證這個任務要新增的行為,並且因為功能尚未實作而**失敗**。",
214
+ `4. 可以先執行本任務相關的測試,確認失敗原因是斷言或找不到尚未實作的模組,而不是語法錯誤或測試本身寫錯。外部流程會再執行 \`${testCmd}\` 驗證紅燈,不需要自行重跑全套測試。`,
215
+ ].join("\n"),
216
+ verifyNote: "驗證測試是否失敗",
217
+ };
218
+ }
219
+ return {
220
+ roleGoal: "這個任務不要求紅燈。請直接寫出鎖定既有行為的測試;測試一開始就通過是預期結果,不要停下來,也不要為了製造失敗而改產品程式。",
221
+ redGuidance: [
222
+ "3. 這個任務的描述已寫明不要求紅燈。請寫出鎖定既有行為的測試;測試一開始就通過是預期結果,不要為了製造失敗而改產品程式,也不要停下來不寫。",
223
+ `4. 可以先執行本任務相關的測試,確認它們能跑完。外部流程會再執行 \`${testCmd}\`,通過即可,不需要自行重跑全套測試。`,
224
+ ].join("\n"),
225
+ verifyNote: "確認測試能跑完。這個任務不要求測試失敗",
226
+ };
227
+ }
202
228
  function loadOrderedTasks(run) {
203
229
  const r = readJsonFile(flowFile(run, "tasks.ordered.json"), TaskList);
204
230
  if (!r.ok)
@@ -269,6 +295,9 @@ function validatePlan(run) {
269
295
  const complexityError = validateTaskComplexity(tasks.data, run.modelMode ?? "balanced");
270
296
  if (complexityError)
271
297
  return complexityError;
298
+ const tddError = validateTddFlag(tasks.data);
299
+ if (tddError)
300
+ return tddError;
272
301
  return orderTasks(tasks.data, new Set(ids));
273
302
  }
274
303
  function acceptPlan(run, ordered) {
@@ -330,38 +359,115 @@ async function planReviewStage(run) {
330
359
  const layered = loadLayeredPlan(run, cfg);
331
360
  return layered ? planReviewLayered(run, layered, cfg) : planReviewFull(run, cfg);
332
361
  }
362
+ const readIfExists = (path) => (existsSync(path) ? readFileSync(path, "utf8") : null);
333
363
  /**
334
- * 整份審查與分層審查共用的單次呼叫:還原審查者改過的計畫檔、驗證裁決與交接。
335
- * 失敗時回傳已呼叫 retry 的 run、沒有 collected,呼叫端應直接回傳這個 run。
364
+ * 平行階段:每位審查者在自己的臨時 worktree 執行,成功的結果立刻存檔。
365
+ * 這個階段不改 run、不寫交接帳本、不呼叫 retry。
366
+ * - 已有同一輪存檔、且仍有效的呼叫直接沿用(resume);帳本在它啟動後被本輪以外的呼叫(例如修正者)動過就作廢重跑
367
+ * - 預算(maxAgentRuns)不夠時只啟動預算內的,回傳 undefined,呼叫端應原樣回傳 run,由 advance 判定 agent_budget
368
+ * - 有呼叫額度用完時,其他呼叫照常跑完並存檔,最後才丟 QuotaPause
336
369
  */
337
- async function collectPlanReview(run, spec) {
370
+ async function executeReviewCalls(run, cfg, scope, fingerprint, calls) {
371
+ if (calls.length === 0)
372
+ return { dir: "", finished: [] };
373
+ // openRound、loadCalls 與有效性檢查都要在 runPool 啟動任何呼叫之前(loadCalls 會刪 .tmp)
374
+ const dir = openRound(run.id, scope, fingerprint);
375
+ const saved = loadCalls(dir);
376
+ const live = readHandoff(run.id);
377
+ const round = [...saved.values()];
378
+ for (const [key, stored] of [...saved]) {
379
+ if (storedCallValid(stored, live, round))
380
+ continue;
381
+ dropCall(dir, key);
382
+ saved.delete(key);
383
+ }
384
+ let budget = run.maxAgentRuns - agentRuns(run.id);
385
+ // 已有存檔的不花預算;其餘依序佔預算,超出的這次不啟動
386
+ const launch = calls.filter((call) => saved.has(call.key) || budget-- > 0);
387
+ const limit = cfg.reviewConcurrency ?? Infinity;
388
+ const tagged = limit > 1 && launch.filter((call) => !saved.has(call.key)).length > 1;
389
+ // 預算內能跑的照跑並存檔,之後 resume 加大預算就不必重付
390
+ const executed = await runPool(launch, limit, (call) => executeOne(run, dir, saved, call, tagged));
391
+ const quota = executed.find((item) => "quota" in item);
392
+ if (quota)
393
+ throw new QuotaPause(quota.quota);
394
+ if (launch.length < calls.length)
395
+ return undefined;
396
+ return { dir, finished: executed };
397
+ }
398
+ async function executeOne(run, dir, saved, call, tagged) {
399
+ const reused = saved.get(call.key);
400
+ if (reused) {
401
+ info(run, ` ↪ 沿用本輪已完成的審查(${call.reviewer})`);
402
+ return { stored: reused, ok: true };
403
+ }
404
+ try {
405
+ return await withTempWorktree(run.id, `slot-${call.slot}`, async (ws) => {
406
+ rmSync(join(ws.flow, call.output), { force: true });
407
+ // 平行階段不動帳本,同一批呼叫看到的都是這一份;序列收尾以它評估這個呼叫的核准門檻
408
+ const base = readHandoff(run.id);
409
+ const outcome = await agentStep(run, call.reviewer, call.step, call.prompt, {
410
+ kind: "review", slot: call.slot, modelScope: call.scope, workspace: ws, tag: tagged ? call.reviewer : undefined,
411
+ });
412
+ const stored = {
413
+ key: call.key, reviewer: call.reviewer, agent: outcome.agent, step: outcome.step, callKey: outcome.callKey,
414
+ summary: outcome.r.summary,
415
+ output: readIfExists(join(ws.flow, call.output)),
416
+ handoffResponse: readIfExists(responsePath(run.id, ws.flow)),
417
+ base,
418
+ };
419
+ if (outcome.r.ok)
420
+ saveCall(dir, stored); // 先存檔,臨時 worktree 才會被移除
421
+ return { stored, ok: outcome.r.ok };
422
+ });
423
+ }
424
+ catch (err) {
425
+ if (err instanceof QuotaPause)
426
+ return { quota: err.message };
427
+ throw err;
428
+ }
429
+ }
430
+ /** 序列收尾第一步:把存檔的內容放回共用 .flow/,之後就能沿用原有的驗證與交接程式 */
431
+ function replay(run, stored, output) {
432
+ mkdirSync(flowDir(run.id), { recursive: true });
433
+ const put = (path, text) => (text === null ? rmSync(path, { force: true }) : writeFileSync(path, text));
434
+ put(flowFile(run, output), stored.output);
435
+ put(responsePath(run.id), stored.handoffResponse);
436
+ }
437
+ /** 存檔內容還原成 finishHandoff 需要的 StepOutcome */
438
+ function storedOutcome(stored, ok) {
439
+ return {
440
+ r: { ok, quotaExhausted: false, summary: stored.summary, usageReported: false },
441
+ agent: stored.agent, step: stored.step, callKey: stored.callKey,
442
+ };
443
+ }
444
+ /**
445
+ * 整份審查與分層審查共用的序列收尾:驗證裁決與交接。
446
+ * 失敗時回傳已呼叫 retry 的 run、沒有 collected,呼叫端應直接回傳這個 run;
447
+ * 不合格的存檔結果一併刪掉,否則同一輪重跑會反覆讀到同一份。
448
+ */
449
+ function applyPlanReview(run, spec, done, dir) {
338
450
  const { reviewer, step } = spec;
339
- const snap = snapshotPlan(run, PLAN_REPLY_FILES);
340
- rmSync(flowFile(run, spec.output), { force: true });
341
- const outcome = await agentStep(run, reviewer, step, spec.prompt, {
342
- kind: "review", slot: spec.slot, modelScope: spec.scope,
343
- reset: async () => { await discardChanges(worktreeDir(run.id)); restorePlan(run, snap); },
344
- });
345
- await discardChanges(worktreeDir(run.id));
346
- const tampered = restorePlan(run, snap);
347
- if (tampered.length)
348
- info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
349
- const stop = (category, reason) => ({
350
- run: retry(recordModelReviewFailure(run, step, reviewer, spec.scope), "plan-review-run", reason, "plan_review", category),
351
- });
352
- if (!outcome.r.ok)
353
- return stop("agent_error", `Agent 執行失敗:${outcome.r.summary}`);
451
+ const stop = (category, reason) => {
452
+ dropCall(dir, spec.key);
453
+ return { run: retry(recordModelReviewFailure(run, step, reviewer, spec.scope), "plan-review-run", reason, "plan_review", category) };
454
+ };
455
+ if (!done.ok)
456
+ return stop("agent_error", `Agent 執行失敗:${done.stored.summary}`);
457
+ replay(run, done.stored, spec.output);
354
458
  const review = readJsonFile(flowFile(run, spec.output), ConsistentReviewResult);
355
459
  if (!review.ok)
356
460
  return stop("format_invalid", review.error);
357
- // 群審查不帶關卡:索引要求修改並新增事項後,群的核准不算矛盾;最後由 planSettled 檢查未結事項
358
- const handoffError = finishHandoff(run, outcome, "reviewer", spec.gated ? { target: "plan", verdict: review.data.verdict } : undefined);
461
+ // 群審查不帶關卡:索引要求修改並新增事項後,群的核准不算矛盾;最後由 planSettled 檢查未結事項。
462
+ // 門檻以這個呼叫自己看到的帳本評估:同輪其他審查者剛新增(或崩潰前已套用)的事項它沒看過,不能拿來判它矛盾
463
+ const handoffError = finishHandoff(run, storedOutcome(done.stored, true), "reviewer", spec.gated ? { target: "plan", verdict: review.data.verdict } : undefined, done.stored.base);
359
464
  if (handoffError)
360
465
  return stop("handoff_invalid", handoffError);
361
466
  run = clearModelReviewFailure(run, step, reviewer, spec.scope);
362
467
  // 審查紀錄移到 worktree 外面:之後的仲裁者看不到是哪一家提的意見
363
468
  mkdirSync(join(runDir(run.id), "reviews"), { recursive: true });
364
- renameSync(flowFile(run, spec.output), join(runDir(run.id), "reviews", spec.archive));
469
+ writeFileSync(join(runDir(run.id), "reviews", spec.archive), readFileSync(flowFile(run, spec.output)));
470
+ rmSync(flowFile(run, spec.output), { force: true });
365
471
  if (review.data.verdict === "approve") {
366
472
  info(run, ` ✓ ${reviewer} 核准${spec.subject}`);
367
473
  return { run, collected: { reviewer, verdict: "approve", issueLines: [] } };
@@ -414,14 +520,19 @@ async function planReviewFull(run, cfg) {
414
520
  const author = run.planWriter ?? planAgent(run.cycle, run.id);
415
521
  const round = (run.attempts["plan-review"] ?? 0) + 1;
416
522
  const panel = reviewers(run.cycle, author, cfg.planReviewQuorum, `${run.id}:plan-review:${round}`);
523
+ const specs = panel.map((reviewer, slot) => ({
524
+ key: `full:${reviewer}`, reviewer, step: "plan-review", slot, gated: true, subject: "計畫",
525
+ prompt: renderPrompt("plan-review", { reviewer, author, requirement: run.requirement }),
526
+ output: "plan-review.json", archive: `plan-review-${round}-${reviewer}.json`,
527
+ }));
528
+ for (const spec of specs)
529
+ info(run, `🧐 計畫審查第 ${round} 輪(${spec.reviewer},作者 ${author})`);
530
+ const ran = await executeReviewCalls(run, cfg, "plan-review", `${round}:${currentPlanKey(run)}`, specs);
531
+ if (!ran)
532
+ return run;
417
533
  const calls = [];
418
- for (const [slot, reviewer] of panel.entries()) {
419
- info(run, `🧐 計畫審查第 ${round} 輪(${reviewer},作者 ${author})`);
420
- const passed = await collectPlanReview(run, {
421
- reviewer, step: "plan-review", slot, gated: true, subject: "計畫",
422
- prompt: renderPrompt("plan-review", { reviewer, author, requirement: run.requirement }),
423
- output: "plan-review.json", archive: `plan-review-${round}-${reviewer}.json`,
424
- });
534
+ for (const [i, spec] of specs.entries()) {
535
+ const passed = applyPlanReview(run, spec, ran.finished[i], ran.dir);
425
536
  run = passed.run;
426
537
  if (!passed.collected)
427
538
  return run;
@@ -490,35 +601,17 @@ async function planReviewLayered(run, layered, cfg) {
490
601
  const calls = [];
491
602
  // 沿用的呼叫也佔一格,重跑時每個呼叫的 slot 才不會變
492
603
  let nextSlot = 0;
493
- /** 執行或沿用一次呼叫;失敗時回傳 false,run 已是 retry 後的狀態 */
494
- const runCall = async (key, taskIds, label, spec) => {
495
- const slot = nextSlot++;
604
+ const items = [];
605
+ const add = (key, taskIds, label, spec) => {
496
606
  const reused = done.get(key);
497
- if (reused) {
607
+ if (reused)
498
608
  info(run, ` ↪ 沿用本輪已完成的${label}(${reused.reviewer})`);
499
- calls.push(reused);
500
- return true;
501
- }
502
- info(run, `🧐 ${label}第 ${round} 輪(${spec.reviewer},作者 ${author})`);
503
- // run 是外層參數,刻意在閉包裡更新:後續呼叫與最後的彙總都要看到 retry、clearModelReviewFailure 之後的 run
504
- const passed = await collectPlanReview(run, { ...spec, slot });
505
- run = passed.run;
506
- if (!passed.collected)
507
- return false;
508
- // 真的執行並成功就是有進展:同一輪不同呼叫輪流失敗時,不會累計到重試上限而讓 run 失敗。
509
- // 進度寫在 round 裡、不會重跑,所以一輪最多失敗「呼叫數 × maxAttempts」次。
510
- // 整份審查每次重跑整輪,不能這樣歸零,否則同一位審查者反覆失敗會無限重試。
511
- const attempts = { ...run.attempts };
512
- delete attempts["plan-review-run"];
513
- run = { ...run, attempts };
514
- const call = { key, ...passed.collected, ...(taskIds ? { taskIds } : {}) };
515
- progress.calls.push(call);
516
- calls.push(call);
517
- saveState({ version: 1, ...(reviewed ? { reviewed } : {}), round: progress });
518
- return true;
609
+ else
610
+ info(run, `🧐 ${label}第 ${round} 輪(${spec.reviewer},作者 ${author})`);
611
+ items.push({ key, taskIds, reused, spec: { ...spec, slot: nextSlot++, key } });
519
612
  };
520
613
  for (const reviewer of panel) {
521
- const ok = await runCall(`index:${reviewer}`, undefined, "計畫索引審查", {
614
+ add(`index:${reviewer}`, undefined, "計畫索引審查", {
522
615
  reviewer, step: "plan-review", gated: true, subject: "計畫索引",
523
616
  prompt: renderPrompt("plan-review-index", {
524
617
  reviewer, author, requirement: run.requirement,
@@ -529,8 +622,6 @@ async function planReviewLayered(run, layered, cfg) {
529
622
  }),
530
623
  output: "plan-review.json", archive: `plan-review-${round}-${reviewer}.json`,
531
624
  });
532
- if (!ok)
533
- return run;
534
625
  }
535
626
  for (const group of layered.dirty) {
536
627
  const groupTasks = layered.tasks.filter((task) => group.taskIds.includes(task.id));
@@ -538,7 +629,7 @@ async function planReviewLayered(run, layered, cfg) {
538
629
  const neighbors = neighborTasks(layered.tasks, group.taskIds).map(({ id, title, description, dependsOn }) => ({ id, title, description, dependsOn }));
539
630
  const groupPanel = reviewers(run.cycle, author, groupReviewerCount(groupTasks, cfg.planReviewQuorum), `${run.id}:plan-group:${group.id}:${round}`);
540
631
  for (const reviewer of groupPanel) {
541
- const ok = await runCall(`group:${group.id}:${group.taskIds.join(",")}:${reviewer}`, group.taskIds, `計畫群 ${group.id} 審查`, {
632
+ add(`group:${group.id}:${group.taskIds.join(",")}:${reviewer}`, group.taskIds, `計畫群 ${group.id} 審查`, {
542
633
  reviewer, step: "plan-review-group", gated: false, subject: `任務群 ${group.id}`, scope: group.id,
543
634
  prompt: renderPrompt("plan-review-group", {
544
635
  reviewer, author, groupId: group.id,
@@ -551,10 +642,32 @@ async function planReviewLayered(run, layered, cfg) {
551
642
  }),
552
643
  output: "plan-review-group.json", archive: `plan-review-${round}-${group.id}-${reviewer}.json`,
553
644
  });
554
- if (!ok)
555
- return run;
556
645
  }
557
646
  }
647
+ // 平行執行還沒套用的呼叫;沿用的(已套用、記在進度檔裡)不再跑
648
+ const pending = items.filter((item) => !item.reused);
649
+ const ran = await executeReviewCalls(run, cfg, "plan-review", `${round}:${planKey}`, pending.map((item) => item.spec));
650
+ if (!ran)
651
+ return run;
652
+ // 序列收尾,依原本的順序逐一套用;每套用一個就寫進度檔,中途被中斷也不會重複套用
653
+ const applied = new Map();
654
+ for (const [i, item] of pending.entries()) {
655
+ const passed = applyPlanReview(run, item.spec, ran.finished[i], ran.dir);
656
+ run = passed.run;
657
+ if (!passed.collected)
658
+ return run;
659
+ // 真的執行並成功就是有進展:同一輪不同呼叫輪流失敗時,不會累計到重試上限而讓 run 失敗。
660
+ // 進度寫在 round 裡、不會重跑,所以一輪最多失敗「呼叫數 × maxAttempts」次。
661
+ // 整份審查每次重跑整輪,不能這樣歸零,否則同一位審查者反覆失敗會無限重試。
662
+ const attempts = { ...run.attempts };
663
+ delete attempts["plan-review-run"];
664
+ run = { ...run, attempts };
665
+ const call = { key: item.key, ...passed.collected, ...(item.taskIds ? { taskIds: item.taskIds } : {}) };
666
+ progress.calls.push(call);
667
+ applied.set(item.key, call);
668
+ saveState({ version: 1, ...(reviewed ? { reviewed } : {}), round: progress });
669
+ }
670
+ calls.push(...items.map((item) => item.reused ?? applied.get(item.key)));
558
671
  // 只放這一輪真的審過的群(含沿用的),沒審到的任務才留得住前次 verdict
559
672
  const groupVerdicts = calls.flatMap((call) => call.taskIds ? [{ taskIds: call.taskIds, verdict: call.verdict }] : []);
560
673
  saveState({ version: 1, reviewed: applyReviewVerdicts(reviewed, layered.tasks, layered.acceptance, layered.planMd, groupVerdicts) });
@@ -701,6 +814,7 @@ async function implementStage(run) {
701
814
  if (!acceptance.ok)
702
815
  throw new Error(acceptance.error);
703
816
  const acceptanceJson = JSON.stringify(taskAcceptance(task, acceptance.data), null, 2);
817
+ const waiveRed = task.tdd !== false && descriptionWaivesRed(task.description);
704
818
  if (run.taskPhase === "review")
705
819
  return taskReviewStep(run, task, progress, taskJson, acceptanceJson);
706
820
  if (run.taskPhase === "verify")
@@ -719,9 +833,12 @@ async function implementStage(run) {
719
833
  if (run.taskPhase === "tests") {
720
834
  const key = `${task.id}:tests`;
721
835
  info(run, `🧪 [${progress}] 撰寫測試(${agents.tests})`);
836
+ if (waiveRed && existsSync(flowFile(run, "feedback.md")) && readFileSync(flowFile(run, "feedback.md"), "utf8").includes("請撰寫會因功能尚未實作而失敗的測試")) {
837
+ rmSync(flowFile(run, "feedback.md"));
838
+ }
722
839
  const before = await headCommit(repo);
723
840
  const snap = snapshotPlan(run, LOCKED_FILES);
724
- const outcome = await agentStep(run, agents.tests, `${task.id}-tests`, renderPrompt("implement-tests", { task: taskJson, acceptance: acceptanceJson, testPattern: cfg.testPattern, testCmd }), { kind: "write", reset: async () => { await resetTo(repo, before); restorePlan(run, snap); } });
841
+ const outcome = await agentStep(run, agents.tests, `${task.id}-tests`, renderPrompt("implement-tests", { task: taskJson, acceptance: acceptanceJson, testPattern: cfg.testPattern, testCmd, ...testsRedGuidance(testCmd, waiveRed) }), { kind: "write", reset: async () => { await resetTo(repo, before); restorePlan(run, snap); } });
725
842
  const { r, agent: testsAuthor } = outcome;
726
843
  const tampered = restorePlan(run, snap);
727
844
  if (!r.ok) {
@@ -741,7 +858,7 @@ async function implementStage(run) {
741
858
  return retry(run, key, `沒有新增或修改任何符合 /${cfg.testPattern}/ 的測試檔。`, "implement", "tests_not_written");
742
859
  }
743
860
  const red = await runCommand(target(run, `${task.id}-red`, CMD_AGENT), testCmd);
744
- if (red.ok) {
861
+ if (red.ok && !waiveRed) {
745
862
  await resetTo(repo, before);
746
863
  return retry(run, key, "測試在功能尚未實作前就全部通過,代表測試沒有驗證到新行為。請撰寫會因功能尚未實作而失敗的測試。", "implement", "tests_not_red");
747
864
  }
@@ -750,8 +867,10 @@ async function implementStage(run) {
750
867
  await resetTo(repo, before);
751
868
  return retry(run, key, handoffError, "implement", "handoff_invalid");
752
869
  }
753
- writeFileSync(flowFile(run, "red-output.txt"), red.output);
754
- info(run, `🔴 [${progress}] 測試如預期失敗`);
870
+ writeFileSync(flowFile(run, "red-output.txt"), waiveRed && red.ok
871
+ ? "此任務不要求紅燈,測試在既有實作下已經通過。不要為了製造失敗而修改產品程式;若沒有其他必須的實作,保持現況即可。"
872
+ : red.output);
873
+ info(run, waiveRed && red.ok ? `✅ [${progress}] 測試已寫好(此任務不要求紅燈)` : `🔴 [${progress}] 測試如預期失敗`);
755
874
  return { ...succeed(run, key, "implement"), taskPhase: "code", taskBase: before, testsCommit: commit, lastTestsAuthor: testsAuthor };
756
875
  }
757
876
  // ── 綠燈:實作到測試通過,而且不可動測試 ──
@@ -956,32 +1075,42 @@ async function codeReview(run, opts) {
956
1075
  const cfg = loadRepoConfig();
957
1076
  const repo = worktreeDir(run.id);
958
1077
  const panel = reviewers(run.cycle, run.lastWriter, cfg.reviewQuorum, opts.seed, opts.testAuthor, opts.prefer);
1078
+ // 審查的是 HEAD:先清掉驗證階段留下的未 commit 修改與未追蹤產物,之後修正時的 commitAll 才不會把它們帶進去(.flow/ 在 exclude 內不受影響)
1079
+ await discardChanges(repo);
959
1080
  writeFileSync(flowFile(run, "diff.patch"), await git(repo, "diff", `${opts.base}...HEAD`));
960
1081
  const authors = [...new Set((await git(repo, "log", "--format=%s", `${opts.base}..HEAD`)).match(/\[[^\]]+\]$/gm) ?? [])]
961
1082
  .map((s) => s.slice(1, -1));
1083
+ const specs = panel.map((reviewer, slot) => ({
1084
+ key: `${slot}:${reviewer}`, slot, reviewer, step: opts.step,
1085
+ prompt: opts.prompt(reviewer, authors.join("、") || "未知"),
1086
+ output: "review.json",
1087
+ }));
1088
+ for (const spec of specs)
1089
+ info(run, `👀 ${opts.label}(${spec.reviewer})`);
1090
+ // 指紋含 HEAD:程式碼被修過就不沿用舊的審查結果
1091
+ const ran = await executeReviewCalls(run, cfg, opts.step, `${opts.seed}|${await headCommit(repo)}|${opts.base}`, specs);
1092
+ if (!ran)
1093
+ return { run };
962
1094
  const issues = [];
963
1095
  let objector;
964
- for (const [slot, reviewer] of panel.entries()) {
965
- info(run, `👀 ${opts.label}(${reviewer})`);
966
- rmSync(flowFile(run, "review.json"), { force: true });
967
- const snap = snapshotPlan(run, LOCKED_FILES);
968
- const outcome = await agentStep(run, reviewer, opts.step, opts.prompt(reviewer, authors.join("、") || "未知"), {
969
- kind: "review", slot, reset: async () => { await discardChanges(repo); restorePlan(run, snap); },
970
- });
971
- const { r } = outcome;
972
- await discardChanges(repo); // 審查者不可改程式碼
973
- const tampered = restorePlan(run, snap); // .flow/ 不受 git 管理,要另外還原
974
- if (tampered.length)
975
- info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
976
- if (!r.ok)
977
- return { run: retry(opts.step === "review" ? recordModelReviewFailure(run, "review", reviewer) : run, opts.runKey, `Agent 執行失敗:${r.summary}`, opts.backTo, "agent_error") };
1096
+ for (const [i, spec] of specs.entries()) {
1097
+ const reviewer = spec.reviewer;
1098
+ const done = ran.finished[i];
1099
+ const fail = (reason, category) => {
1100
+ dropCall(ran.dir, spec.key); // 不合格的存檔不能留著,否則同一輪重跑會讀到同一份
1101
+ return { run: retry(opts.step === "review" ? recordModelReviewFailure(run, "review", reviewer) : run, opts.runKey, reason, opts.backTo, category) };
1102
+ };
1103
+ if (!done.ok)
1104
+ return fail(`Agent 執行失敗:${done.stored.summary}`, "agent_error");
1105
+ replay(run, done.stored, "review.json");
978
1106
  const review = readJsonFile(flowFile(run, "review.json"), ConsistentReviewResult);
979
1107
  if (!review.ok)
980
- return { run: retry(opts.step === "review" ? recordModelReviewFailure(run, "review", reviewer) : run, opts.runKey, review.error, opts.backTo, "format_invalid") };
1108
+ return fail(review.error, "format_invalid");
981
1109
  const gate = opts.gate ? { target: "code", verdict: review.data.verdict } : undefined;
982
- const handoffError = finishHandoff(run, outcome, "reviewer", gate);
1110
+ // 門檻以這個呼叫自己看到的帳本評估:同輪其他審查者剛新增(或崩潰前已套用)的事項它沒看過,不能拿來判它矛盾
1111
+ const handoffError = finishHandoff(run, storedOutcome(done.stored, true), "reviewer", gate, done.stored.base);
983
1112
  if (handoffError)
984
- return { run: retry(opts.step === "review" ? recordModelReviewFailure(run, "review", reviewer) : run, opts.runKey, handoffError, opts.backTo, "handoff_invalid") };
1113
+ return fail(handoffError, "handoff_invalid");
985
1114
  if (opts.step === "review")
986
1115
  run = clearModelReviewFailure(run, "review", reviewer);
987
1116
  renameSync(flowFile(run, "review.json"), flowFile(run, opts.saveAs(reviewer)));
@@ -1061,6 +1190,7 @@ const STAGES = {
1061
1190
  export async function advance(initial) {
1062
1191
  let run = initial;
1063
1192
  recoverHandoff(run.id);
1193
+ await cleanupTempWorktrees(run.id); // 上次被中斷(Ctrl-C、SIGTERM、當機)時留下的平行審查臨時 worktree
1064
1194
  for (;;) {
1065
1195
  if (["done", "failed", "awaiting_approval", "paused"].includes(run.stage))
1066
1196
  return run;
package/dist/handoff.js CHANGED
@@ -5,7 +5,8 @@ import { z } from "zod";
5
5
  import { flowDir, handoffPath, runDir } from "./paths.js";
6
6
  import { HandoffLedger, HandoffResponse, HandoffSource } from "./schemas.js";
7
7
  import { readJsonFile } from "./util.js";
8
- const responsePath = (id) => join(flowDir(id), "handoff-response.json");
8
+ /** agent 寫交接回覆的位置;平行審查者用自己臨時 worktree 的 .flow/ */
9
+ export const responsePath = (id, flow = flowDir(id)) => join(flow, "handoff-response.json");
9
10
  const receiptsDir = (id) => join(runDir(id), "handoff-receipts");
10
11
  const keyHash = (key) => createHash("sha256").update(key).digest("hex").slice(0, 16);
11
12
  const receiptPath = (id, key) => join(receiptsDir(id), `${keyHash(key)}.json`);
@@ -71,7 +72,7 @@ export function mergeHandoff(id, callKey, source, response, role) {
71
72
  return next;
72
73
  }
73
74
  /** 只把目前步驟需要處理的事項投影給 agent。 */
74
- export function prepareHandoff(id, _callKey, target, blind) {
75
+ export function prepareHandoff(id, _callKey, target, blind, flow = flowDir(id)) {
75
76
  const items = readHandoff(id).issues.filter((item) => item.targetStage === target && (item.kind === "info" || item.status === "open" || item.status === "proposed_resolved"));
76
77
  const render = (item) => {
77
78
  const source = blind ? "" : `\n來源:${item.source.stage}/${item.source.agent}`;
@@ -85,12 +86,12 @@ export function prepareHandoff(id, _callKey, target, blind) {
85
86
  actions.length ? `## 待處理事項(action,可在 dispositions 處置)\n\n${actions.join("\n\n")}` : "",
86
87
  infos.length ? `## 參考資訊(info,只供參考,不要放進 dispositions)\n\n${infos.join("\n\n")}` : "",
87
88
  ].filter(Boolean);
88
- mkdirSync(flowDir(id), { recursive: true });
89
- writeFileSync(join(flowDir(id), "handoff-context.md"), `# 待處理交接事項\n\n${sections.length ? sections.join("\n\n") : "目前沒有待處理事項。"}\n`);
90
- rmSync(responsePath(id), { force: true });
89
+ mkdirSync(flow, { recursive: true });
90
+ writeFileSync(join(flow, "handoff-context.md"), `# 待處理交接事項\n\n${sections.length ? sections.join("\n\n") : "目前沒有待處理事項。"}\n`);
91
+ rmSync(responsePath(id, flow), { force: true });
91
92
  }
92
- export function validateHandoffResponse(id) {
93
- return readJsonFile(responsePath(id), HandoffResponse);
93
+ export function validateHandoffResponse(id, flow = flowDir(id)) {
94
+ return readJsonFile(responsePath(id, flow), HandoffResponse);
94
95
  }
95
96
  /** 已通過原有關卡的回覆先記收據,再合併;中斷後可重播。 */
96
97
  export function acceptHandoff(id, callKey, source, response, role) {
@@ -0,0 +1,112 @@
1
+ import { createHash } from "node:crypto";
2
+ import { existsSync, mkdirSync, readFileSync, readdirSync, renameSync, rmSync, writeFileSync } from "node:fs";
3
+ import { join } from "node:path";
4
+ import { z } from "zod";
5
+ import { parallelReviewDir } from "./paths.js";
6
+ import { HandoffLedger } from "./schemas.js";
7
+ /**
8
+ * 依序啟動、同時最多 limit 個;結果依輸入順序回傳。
9
+ * fn 丟例外時,其餘已啟動與待啟動的項目仍會跑完(讓它們的收尾,例如移除臨時 worktree,都能完成),最後才把第一個例外丟出。
10
+ */
11
+ export async function runPool(items, limit, fn) {
12
+ const results = new Array(items.length);
13
+ let next = 0;
14
+ let failure;
15
+ const worker = async () => {
16
+ for (;;) {
17
+ const index = next++;
18
+ if (index >= items.length)
19
+ return;
20
+ try {
21
+ results[index] = await fn(items[index], index);
22
+ }
23
+ catch (error) {
24
+ failure ??= { error };
25
+ }
26
+ }
27
+ };
28
+ await Promise.all(Array.from({ length: Math.max(1, Math.min(limit, items.length)) }, worker));
29
+ if (failure)
30
+ throw failure.error;
31
+ return results;
32
+ }
33
+ /**
34
+ * 一次已執行成功的審查呼叫。放在 run 目錄(worktree 外),以輪次+輸入指紋為鍵記憶化:
35
+ * 中斷後 resume、或同一輪因別的呼叫失敗而重跑時直接沿用,不重複付費。
36
+ * 只存原始檔案內容:套用時再放回共用的 .flow/,走原有的驗證與交接程式。
37
+ */
38
+ export const StoredCall = z.object({
39
+ /** 同一輪內唯一的呼叫識別 */
40
+ key: z.string(),
41
+ reviewer: z.string(),
42
+ agent: z.string(),
43
+ step: z.string(),
44
+ callKey: z.string(),
45
+ summary: z.string(),
46
+ /** 裁決檔原文;agent 沒寫時為 null */
47
+ output: z.string().nullable(),
48
+ /** handoff-response.json 原文;agent 沒寫時為 null */
49
+ handoffResponse: z.string().nullable(),
50
+ /** 這個呼叫啟動時看到的交接帳本:核准門檻以它評估,也用來判斷存檔是否仍可沿用 */
51
+ base: HandoffLedger,
52
+ });
53
+ const hash = (text) => createHash("sha256").update(text).digest("hex").slice(0, 16);
54
+ const scopeDir = (runId, scope) => join(parallelReviewDir(runId), scope);
55
+ /** 這一輪的存檔目錄;同 scope 下其他指紋(計畫或程式碼已變、輪次已換)的結果一律作廢 */
56
+ export function openRound(runId, scope, fingerprint) {
57
+ const base = scopeDir(runId, scope);
58
+ const mine = hash(fingerprint);
59
+ if (existsSync(base)) {
60
+ for (const name of readdirSync(base)) {
61
+ if (name !== mine)
62
+ rmSync(join(base, name), { recursive: true, force: true });
63
+ }
64
+ }
65
+ const dir = join(base, mine);
66
+ mkdirSync(dir, { recursive: true });
67
+ return dir;
68
+ }
69
+ const callFile = (dir, key) => join(dir, `${hash(key)}.json`);
70
+ export function saveCall(dir, call) {
71
+ const path = callFile(dir, call.key);
72
+ const tmp = `${path}.tmp`;
73
+ writeFileSync(tmp, JSON.stringify(call));
74
+ renameSync(tmp, path);
75
+ }
76
+ /**
77
+ * 讀出這一輪已存檔的結果,並刪掉寫到一半的 .tmp 與損毀的檔案。
78
+ * 呼叫順序:必須在 runPool 啟動任何呼叫之前;之後才呼叫會刪到別的呼叫正在寫的 .tmp。
79
+ */
80
+ export function loadCalls(dir) {
81
+ const calls = new Map();
82
+ if (!existsSync(dir))
83
+ return calls;
84
+ for (const name of readdirSync(dir)) {
85
+ const path = join(dir, name);
86
+ if (!name.endsWith(".json")) {
87
+ rmSync(path, { force: true }); // 寫到一半留下的 .tmp
88
+ continue;
89
+ }
90
+ try {
91
+ const call = StoredCall.parse(JSON.parse(readFileSync(path, "utf8")));
92
+ calls.set(call.key, call);
93
+ }
94
+ catch {
95
+ rmSync(path, { force: true }); // 損毀或格式不符:當作沒有,重跑
96
+ }
97
+ }
98
+ return calls;
99
+ }
100
+ export function dropCall(dir, key) {
101
+ rmSync(callFile(dir, key), { force: true });
102
+ }
103
+ /**
104
+ * 存檔是否仍可沿用:帳本在這個呼叫啟動後多出的已套用呼叫,必須全是這一輪已存檔的審查者(崩潰或重跑時已套用的 slot)。
105
+ * 多出別的呼叫(例如修正者只回覆交接、沒改計畫或程式碼,輪次與指紋都沒變)代表審查者沒看過現在的帳本,要重跑。
106
+ */
107
+ export function storedCallValid(call, live, round) {
108
+ const seen = new Set(call.base.appliedCalls ?? []);
109
+ const sameRound = new Set([...round].map((item) => item.callKey));
110
+ return (live.appliedCalls ?? []).every((key) => seen.has(key) || sameRound.has(key));
111
+ }
112
+ //# sourceMappingURL=parallelReview.js.map
package/dist/paths.js CHANGED
@@ -22,6 +22,10 @@ export const handoffPath = (id) => join(runDir(id), "handoff.json");
22
22
  export const planReviewStatePath = (id) => join(runDir(id), "plan-review-state.json");
23
23
  /** 已交付仲裁、尚未得出裁決:暫停後 resume 直接回到仲裁。放在 worktree 外,agent 無法偽造 */
24
24
  export const planArbitrationPath = (id) => join(runDir(id), "plan-arbitration.json");
25
+ /** 平行審查的臨時 worktree(每個呼叫一個);孤兒在 advance() 開頭統一清掉 */
26
+ export const tempWorktreesDir = (id) => join(runDir(id), "tmp-review");
27
+ /** 平行審查已執行成功的呼叫存檔(依輪次與輸入指紋分目錄),resume 時沿用 */
28
+ export const parallelReviewDir = (id) => join(runDir(id), "parallel-review");
25
29
  /** 每個 run 一個 git worktree,Agent 只在這裡工作,不碰你正在編輯的檔案 */
26
30
  export const worktreesDir = () => join(agentflowctlDir(), "worktrees");
27
31
  export const worktreeDir = (id) => join(worktreesDir(), id);
@@ -192,6 +192,7 @@ export function taskFingerprint(task, acceptance, planMd) {
192
192
  complexity: task.complexity ?? null,
193
193
  dependsOn: task.dependsOn,
194
194
  acceptance: task.acceptance.map((id) => ({ id, description: byId.get(id) ?? "" })),
195
+ tdd: task.tdd ?? null,
195
196
  evidence: extractPlanEvidence(planMd, [task.id]),
196
197
  });
197
198
  }
package/dist/schemas.js CHANGED
@@ -74,7 +74,7 @@ export const TaskItem = z.object({
74
74
  dependsOn: z.array(z.string()).default([]),
75
75
  acceptance: z.array(z.string()).min(1, "每個任務至少要對應一條驗收條件"),
76
76
  complexity: z.enum(["low", "medium", "high"]).optional(),
77
- /** false=這個任務不適合先寫會失敗的測試(建置流程、設定、文件、純重構等),略過紅燈直接實作;沒寫視為 true */
77
+ /** false=這個任務不適合先寫會失敗的測試(建置流程、設定、文件、純重構、實作前就會通過的特徵化測試等),略過紅燈直接實作;沒寫視為 true。描述寫明不要求紅燈時必須為 false */
78
78
  tdd: z.boolean().optional(),
79
79
  });
80
80
  /** Agent 在 plan 階段產出的 .flow/tasks.json */
@@ -143,6 +143,8 @@ export const RepoConfig = z.object({
143
143
  reviewQuorum: z.number().int().min(1).default(1),
144
144
  /** 計畫需要幾位不同的 reviewer 都核准 */
145
145
  planReviewQuorum: z.number().int().min(1).default(1),
146
+ /** 同一輪審查最多幾位審查者同時執行;沒寫=不限,1=一次一位 */
147
+ reviewConcurrency: z.number().int().min(1).optional(),
146
148
  /** 計畫審查僵持不下(達到重試上限或意見不再變化)時,交給第三方 agent 仲裁,而不是停下來等人 */
147
149
  planArbiter: z.boolean().default(true),
148
150
  /** 任務夠多、能依檔案分群時,計畫審查改成每輪一次索引加上只審有變動的任務群 */
package/dist/tasks.js CHANGED
@@ -1,5 +1,19 @@
1
1
  /** 一個任務最多做兩件事:對應的驗收條件超過這個數量就要再拆 */
2
2
  export const MAX_TASK_ACCEPTANCE = 2;
3
+ /** 描述已明確放棄紅燈。新計畫必須把 tdd 標成 false;已定案的任務則仍寫測試,但不要求先失敗。 */
4
+ const RED_WAIVED = /不要求紅燈|不必紅燈|不需紅燈|無需紅燈|不用紅燈|略過紅燈|略過紅綠燈/;
5
+ export function descriptionWaivesRed(description) {
6
+ return RED_WAIVED.test(description);
7
+ }
8
+ /** 描述與 tdd 矛盾時退回計畫:寫明不要求紅燈就必須標 false。 */
9
+ export function validateTddFlag(tasks) {
10
+ const conflicts = tasks.filter((task) => task.tdd !== false && descriptionWaivesRed(task.description));
11
+ if (!conflicts.length)
12
+ return undefined;
13
+ return conflicts
14
+ .map((task) => `${task.id} 的描述寫明不要求紅燈,但 tdd 不是 false。這種任務在實作前就會通過,請改成 "tdd": false。`)
15
+ .join("\n");
16
+ }
3
17
  /** adaptive 的新計畫必須明確標註難度;舊 run 仍可讀取缺少欄位的 task。 */
4
18
  export function validateTaskComplexity(tasks, mode) {
5
19
  if (mode === "balanced")
@@ -0,0 +1,129 @@
1
+ import { cpSync, existsSync, mkdirSync, realpathSync, rmSync, symlinkSync } from "node:fs";
2
+ import { join, sep } from "node:path";
3
+ import { git, removeWorktree } from "./git.js";
4
+ import { flowDir, projectRoot, tempWorktreesDir, worktreeDir } from "./paths.js";
5
+ export { tempWorktreesDir };
6
+ let created = 0;
7
+ const warn = (what, err) => console.warn(`⚠️ ${what}失敗(下次執行或 clean 會再清):${err.message}`);
8
+ /**
9
+ * 建立臨時 worktree:detached、內容是 run worktree 的 HEAD,並複製目前的 .flow/(.flow/ 不受 git 管理,要另外帶),
10
+ * run worktree 頂層有 node_modules 時再建一個指向它的 symlink,讓審查者讀得到依賴。
11
+ * 路徑固定為 tmp-review/<name>:agent CLI(Claude Code、Gemini CLI)會依工作目錄記錄專案,路徑每次不同會一直累積紀錄。
12
+ * 固定路徑被占用(目錄已存在,或殘留的 locked 登記讓 git worktree add 失敗)時才改用唯一的父目錄,basename 仍是 name。
13
+ */
14
+ export async function createTempWorktree(runId, name) {
15
+ const stable = join(tempWorktreesDir(runId), name);
16
+ mkdirSync(tempWorktreesDir(runId), { recursive: true });
17
+ let ws;
18
+ if (!existsSync(stable)) {
19
+ try {
20
+ await git(worktreeDir(runId), "worktree", "add", "--detach", stable, "HEAD");
21
+ ws = { dir: stable, flow: join(stable, ".flow") };
22
+ }
23
+ catch {
24
+ // 固定路徑被占用:清掉這次可能留下的半成品目錄(原本不存在,是這次建的),改用唯一路徑
25
+ try {
26
+ rmSync(stable, { recursive: true, force: true });
27
+ }
28
+ catch (err) {
29
+ warn("清掉固定路徑的半成品 ", err);
30
+ }
31
+ }
32
+ }
33
+ if (!ws) {
34
+ const parent = join(tempWorktreesDir(runId), `${process.pid}-${Date.now()}-${created++}`);
35
+ const dir = join(parent, name);
36
+ mkdirSync(parent, { recursive: true });
37
+ await git(worktreeDir(runId), "worktree", "add", "--detach", dir, "HEAD");
38
+ ws = { dir, flow: join(dir, ".flow"), uniqueParent: parent };
39
+ }
40
+ try {
41
+ if (existsSync(flowDir(runId)))
42
+ cpSync(flowDir(runId), ws.flow, { recursive: true });
43
+ else
44
+ mkdirSync(ws.flow, { recursive: true });
45
+ const deps = join(worktreeDir(runId), "node_modules");
46
+ if (existsSync(deps))
47
+ symlinkSync(deps, join(ws.dir, "node_modules"), process.platform === "win32" ? "junction" : "dir");
48
+ }
49
+ catch (err) {
50
+ await removeTempWorktree(runId, ws); // 不會丟例外,不會蓋掉原本的錯誤
51
+ throw err;
52
+ }
53
+ return ws;
54
+ }
55
+ /**
56
+ * 移除這次建立的臨時 worktree:固定路徑只刪 dir(絕不刪 tmp-review/ 本身,裡面可能還有別的呼叫在用),唯一路徑連父目錄一起刪。
57
+ * git worktree remove 與 rmSync 都不會跟進 node_modules symlink 刪到原本的依賴。
58
+ * 永遠不丟例外:這在 withTempWorktree 的 finally 裡執行,丟出去會蓋掉審查本身的結果(已存檔,甚至可能是額度用完);
59
+ * 清不掉的只印警告,交給下次 advance 或 clean 的清理。
60
+ */
61
+ export async function removeTempWorktree(runId, ws) {
62
+ let removed = true;
63
+ try {
64
+ await removeWorktree(worktreeDir(runId), ws.dir);
65
+ }
66
+ catch {
67
+ removed = false; // 目錄可能已經不在:刪掉目錄後再讓 git 忘掉這一個登記
68
+ }
69
+ try {
70
+ rmSync(ws.uniqueParent ?? ws.dir, { recursive: true, force: true, maxRetries: 3 });
71
+ }
72
+ catch (err) {
73
+ warn("移除平行審查的臨時 worktree ", err);
74
+ return;
75
+ }
76
+ // 只處理自己這一個登記;repo 層級的 prune 可能清掉使用者其他資料夾暫時不在的 worktree
77
+ if (!removed)
78
+ await git(worktreeDir(runId), "worktree", "remove", "-f", "-f", ws.dir).catch(() => { });
79
+ }
80
+ export async function withTempWorktree(runId, name, fn) {
81
+ const ws = await createTempWorktree(runId, name);
82
+ try {
83
+ return await fn(ws);
84
+ }
85
+ finally {
86
+ await removeTempWorktree(runId, ws);
87
+ }
88
+ }
89
+ /** git worktree list 裡登記在這個 run 臨時目錄下的 worktree(含 locked 的) */
90
+ async function registeredTempWorktrees(repo, runId) {
91
+ const base = tempWorktreesDir(runId);
92
+ const prefixes = [base + sep];
93
+ try {
94
+ prefixes.push(realpathSync(base) + sep);
95
+ }
96
+ catch {
97
+ // 目錄已不在:只比對原路徑
98
+ }
99
+ return (await git(repo, "worktree", "list", "--porcelain"))
100
+ .split("\n")
101
+ .filter((line) => line.startsWith("worktree "))
102
+ .map((line) => line.slice("worktree ".length))
103
+ .filter((path) => prefixes.some((prefix) => path.startsWith(prefix)));
104
+ }
105
+ /**
106
+ * 清掉上次中斷(Ctrl-C 的 process.exit 不跑 finally、SIGTERM、kill -9、當機)留下的臨時 worktree 與 git 登記。
107
+ * git 在 worktree add 途中被強制中止會留下 locked 登記,prune 會略過它,所以先逐個 remove -f -f。
108
+ * 這個 run 沒有 tmp-review/ 就什麼都不做(不呼叫 git);prune 是 repo 層級的,只在確實找到這個 run 的登記時才執行。
109
+ * force(`clean` 用):tmp-review/ 已不在也照樣列出登記並清掉,否則只剩 locked 登記時會永遠留在 .git/worktrees/。
110
+ * 不丟例外:清不乾淨只印警告,下一次 advance 會再試;固定路徑被殘骸占用時建立會改用唯一路徑,不會被擋住。
111
+ */
112
+ export async function cleanupTempWorktrees(runId, opts = {}) {
113
+ if (!opts.force && !existsSync(tempWorktreesDir(runId)))
114
+ return;
115
+ try {
116
+ const repo = existsSync(worktreeDir(runId)) ? worktreeDir(runId) : projectRoot();
117
+ const registered = await registeredTempWorktrees(repo, runId);
118
+ for (const path of registered) {
119
+ await git(repo, "worktree", "remove", "-f", "-f", path).catch(() => { }); // 失敗交給下面的 rmSync 與 prune
120
+ }
121
+ rmSync(tempWorktreesDir(runId), { recursive: true, force: true, maxRetries: 3 });
122
+ if (registered.length)
123
+ await git(repo, "worktree", "prune");
124
+ }
125
+ catch (err) {
126
+ warn("清理平行審查的臨時 worktree ", err);
127
+ }
128
+ }
129
+ //# sourceMappingURL=tempWorktree.js.map
@@ -4,6 +4,7 @@
4
4
  "tddSplit": true,
5
5
  "reviewQuorum": 1,
6
6
  "planReviewQuorum": 1,
7
+ "reviewConcurrency": 2,
7
8
  "planArbiter": true,
8
9
  "planReviewLayers": { "enabled": true, "minTasks": 7, "maxGroups": 5, "tasksPerGroup": 3 },
9
10
  "tieBreak": "proceed",
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "agentflowctl",
3
3
  "license": "MIT",
4
- "version": "0.16.0",
4
+ "version": "0.17.1",
5
5
  "description": "跨廠商 AI 開發 harness:Claude Code、Codex、Gemini 輪流實作、審查、修正",
6
6
  "keywords": [
7
7
  "ai",
@@ -1,5 +1,5 @@
1
1
  <role>
2
- 你是任務實作者,負責完成一個**不走 TDD** 的任務。這個任務不適合先寫會失敗的測試(例如建置流程、設定、文件、型別或純重構),或專案沒有測試框架,所以沒有紅燈測試可依循。你要依任務描述與驗收條件,用符合專案風格的最小改動完成它。
2
+ 你是任務實作者,負責完成一個**不走 TDD** 的任務。這個任務不適合先寫會失敗的測試(例如建置流程、設定、文件、型別、純重構,或鎖定既有行為的特徵化測試),或專案沒有測試框架,所以沒有紅燈測試可依循。你要依任務描述與驗收條件,用符合專案風格的最小改動完成它。若任務是補一個現有實作下就會通過的測試,寫出該測試即可,不要為了製造失敗而去改產品程式。
3
3
  </role>
4
4
 
5
5
  <context>
@@ -1,5 +1,5 @@
1
1
  <role>
2
- 你是測試工程師,在 TDD 的紅燈階段**只寫測試,不寫實作**。你的測試要精準描述任務要新增的行為,並且在功能實作前確實失敗;之後會由另一位工程師實作到通過,而且對方不能修改你的測試。
2
+ 你是測試工程師,在 TDD 的紅燈階段**只寫測試,不寫實作**。{{roleGoal}}
3
3
  </role>
4
4
 
5
5
  <context>
@@ -37,13 +37,12 @@
37
37
  <steps>
38
38
  1. 若 .flow/feedback.md 存在,先閱讀,並依內容調整做法。
39
39
  2. 依任務描述撰寫測試,檔名必須符合正規表示式 `{{testPattern}}`。
40
- 3. 測試必須驗證這個任務要新增的行為,並且因為功能尚未實作而**失敗**。
41
- 4. 可以先執行本任務相關的測試,確認失敗原因是斷言或找不到尚未實作的模組,而不是語法錯誤或測試本身寫錯。外部流程會再執行 `{{testCmd}}` 驗證紅燈,不需要自行重跑全套測試。
40
+ {{redGuidance}}
42
41
  </steps>
43
42
 
44
43
  <constraints>
45
44
  - 不可實作功能本身。可以建立讓測試能編譯所需的最小型別或空殼匯出,但不可以有真正的邏輯。
46
- - 不要執行 git commit(權限設定已禁止),外部流程會提交並驗證測試是否失敗。
45
+ - 不要執行 git commit(權限設定已禁止),外部流程會提交並{{verifyNote}}。
47
46
  - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
48
47
  </constraints>
49
48
 
@@ -46,7 +46,7 @@
46
46
 
47
47
  <review_focus>
48
48
  只看這一群:
49
- 1. 每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出實作前會失敗的測試。
49
+ 1. 每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出實作前會失敗的測試。無法先失敗的(特徵化、實作前就會通過、描述寫明不要求紅燈)必須是 `tdd: false`。只在描述補註、卻留下 `tdd: true` 或沒寫 `tdd`,仍是 `changes_requested`,`note` 要要求改成 `tdd: false`。
50
50
  2. 任務描述是否對得上上面的驗收條文。
51
51
  3. 對照描述點名、已存在的檔案與難度摘錄,獨立核對 `complexity`。`low` 是沿用既有做法、侷限單一行為且失敗可由局部測試發現;`medium` 包括多模組或介面協調、非典型邊界、相容性或狀態遷移;`high` 包括跨系統契約、架構或資料模型變更、未知的關鍵路徑,或資料遺失、權限、難以回復的風險。摘錄是空的,而且高低估會影響選模時,要求在 .flow/plan.md 該 task 的「## T-<數字>」標題下補上證據。
52
52
  4. 做法是否符合那些檔案中已存在者的慣例。
@@ -30,7 +30,7 @@
30
30
  <review_focus>
31
31
  1. **需求覆蓋**:規格是否完整涵蓋原始需求?有沒有遺漏、誤解,或加入需求沒要求的範圍?
32
32
  2. **驗收條件**:每一條是否具體、可以用自動化測試驗證,而且只描述一個行為?把多個行為寫在同一條的,要求拆開。有沒有重要的邊界情況或錯誤處理沒被列入?
33
- 3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試(標 `tdd: false` 的任務除外:核對它的改動內容確實不適合先寫失敗測試,例如建置流程、設定、文件、型別、純重構,並有寫明驗收方式;會改變程式行為卻標成 `false` 的,要求改回 `true`)?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?
33
+ 3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試(標 `tdd: false` 的任務除外:核對它的改動內容確實不適合先寫失敗測試,例如建置流程、設定、文件、型別、純重構,或實作前就會通過的特徵化測試,並有寫明驗收方式;會改變程式行為卻標成 `false` 的,要求改回 `true`。描述寫明不要求紅燈,`tdd` 卻不是 `false` 的,要求改成 `false`)?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?
34
34
  同時依 .flow/plan.md 的逐項理由及實際程式碼,獨立核對每個 task 的 `complexity`:分別看影響範圍、技術不確定性與失敗後果,取最高等級。`low` 須是沿用既有做法、侷限單一行為或模組且失敗可由局部測試發現;`medium` 包括多模組或介面協調、非典型邊界、相容性或狀態遷移風險;`high` 包括跨系統契約、架構或資料模型變更、未知的關鍵技術路徑,或資料遺失、權限、難以回復的風險。不要只憑檔案數、程式碼行數或驗收條件數判定。
35
35
  理由缺漏、與程式碼不符,或高低估會影響選模時,要求修正;在 `note` 指出 task ID、具體證據、建議等級及須修改的 .flow/plan.md/.flow/tasks.json 部分。不要為缺少高價值證據的細微措辭差異要求修改。
36
36
  4. **技術方向**:是否符合專案既有的架構與慣例?有沒有更簡單的做法,或明顯的風險?
package/prompts/plan.md CHANGED
@@ -52,7 +52,7 @@
52
52
  - 每個任務是一個可獨立測試的垂直切片,小到一次 TDD 循環就能完成;只動少數幾個檔案,測試只驗證一兩個行為。
53
53
  - `title` 用一句話說出這件事;需要用「並且」「以及」串起來的,就是兩個任務。
54
54
  - `description` 寫清楚要動哪些檔案(寫含目錄的路徑,例如 `src/form.ts`,不要只寫檔名)、測試要驗證哪個行為,以及這個任務不做什麼。計畫審查會依這些路徑把任務分群。
55
- - 依改動內容標記 `tdd`:會改變程式行為、能寫出「在實作前會失敗」的測試的任務標 `true`(預設);改動內容不適合先寫失敗測試的任務標 `false`,這類任務會略過紅燈直接實作,改由任務審查與驗證指令把關。適合標 `false` 的例子:建置流程與打包設定(build、CI、bundler、tsconfig)、依賴與版本設定、文件與 prompt 文字、樣式與靜態資源、型別宣告、不改變行為的重構與搬移檔案。能併入相關行為任務的設定或重構,仍請併入,不要獨立成任務。標 `false` 時要在 .flow/plan.md 該 task 的節裡寫明理由,以及這個任務要怎麼驗收(例如「`npm run build` 通過」)。
55
+ - 依改動內容標記 `tdd`:會改變程式行為、能寫出「在實作前會失敗」的測試的任務標 `true`(預設);改動內容不適合先寫失敗測試的任務標 `false`,這類任務會略過紅燈直接實作,改由任務審查與驗證指令把關。適合標 `false` 的例子:建置流程與打包設定(build、CI、bundler、tsconfig)、依賴與版本設定、文件與 prompt 文字、樣式與靜態資源、型別宣告、不改變行為的重構與搬移檔案,以及鎖定既有行為的特徵化測試(實作前就會通過、不要求紅燈、通常不需改產品程式)。能併入相關行為任務的設定或重構,仍請併入,不要獨立成任務。描述寫了「不要求紅燈」卻沒把 `tdd` 設成 `false`,計畫不會通過。標 `false` 時要在 .flow/plan.md 該 task 的節裡寫明理由,以及這個任務要怎麼驗收(例如「`npm run build` 通過」或「新增的測試在現有實作下通過」)。
56
56
  - 專案沒有測試框架時,所有任務都會略過 TDD(程式會強制),此時仍請照實標記 `tdd`,並在 `description` 寫清楚驗收方式。
57
57
  - 每個任務先檢查預計修改的程式碼,再依「影響範圍、技術不確定性、失敗後果」三個面向判定 `complexity`,取其中最高的等級;不要只憑檔案數、程式碼行數或驗收條件數判定。
58
58
  - `low`:沿用現有做法,變更侷限在單一行為或模組,失敗容易由局部測試發現且不影響既有資料或對外契約。
@@ -35,6 +35,8 @@
35
35
  2. 是否有明顯的錯誤、邊界情況遺漏、安全問題或效能問題。
36
36
  3. 是否符合專案既有的架構與慣例,以及任務說明的範圍(沒有做到一半,也沒有做了其他任務的事)。
37
37
 
38
+ 4. 任務描述寫明不要求紅燈時,測試在既有實作下一開始就通過是允許的。不要因為沒有紅燈、或 `tdd` 不是 `false`,就 `changes_requested`。只審查測試是否鎖定任務描述的行為。
39
+
38
40
  只審查本任務的變更;其他任務的驗收條件不在這次審查範圍內。先根據 diff 與驗收條件定位需要查閱的檔案,只在證據不足時讀取其他檔案。不必為了審查重跑全套檢查。
39
41
 
40
42
  風格偏好與無關緊要的小問題不需要要求修改。