agentflowctl 0.17.1 → 0.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -33,7 +33,7 @@ npx agentflowctl run --req-file ./requirement.md
33
33
  ## 執行時會發生什麼
34
34
 
35
35
  1. agent 整理需求與驗收條件,接著寫計畫,交給其他 agent 審查。
36
- 2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent 實作到測試通過;每個任務都會經過審查與驗證。計畫 agent 會依改動內容在任務標記 `tdd`:建置流程、設定、文件、型別、純重構,以及實作前就會通過的特徵化測試,會略過紅綠燈直接實作,改由任務審查與驗證把關。描述寫明不要求紅燈卻沒標 `tdd: false` 的計畫不會通過。已定案的計畫若仍帶著這個衝突,執行到該任務時仍會寫測試,但測試一開始就通過也算完成,不會再要求紅燈、也不會因此讓 run 失敗。專案沒有測試框架時,所有任務都略過紅綠燈,也不跑 `test` 檢查。
36
+ 2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent 實作到測試通過;每個任務都會經過審查與驗證。計畫 agent 會依改動內容在任務標記 `tdd`:建置流程、設定、文件、型別、純重構,以及實作前就會通過的特徵化測試,會略過紅綠燈直接實作,改由任務審查與驗證把關。描述寫明不要求紅燈卻沒標 `tdd: false` 的計畫不會通過。已定案的計畫若仍帶著這個衝突,執行到該任務時仍會寫測試,但測試一開始就通過也算完成,不會再要求紅燈、也不會因此讓 run 失敗。專案沒有測試框架時,所有任務都略過紅綠燈,也不跑 `test` 檢查。只跑檢查、不改檔案的工作不要拆成實作任務;這種任務若沒有檔案變更會直接略過,不再要求 commit。需要人眼確認的任務標成 `kind: confirm`,計畫定案後寫進 `.flow/confirmations.json`,不進入實作,也不會把 run 停下來。`status` 在任務清單之外另列「待你確認(不進實作)」;`confirmations <id>` 只印這個區塊。計畫還沒通過首次驗證前查詢,清單會從尚未經檢查的草稿蒐集,標題會多帶「(計畫尚未定案,以下為草稿)」,項目最終可能不會定案。
37
37
  3. 全部任務完成後,再執行專案檢查與整體程式碼審查。未通過的項目會交回修正。
38
38
  4. 有 `origin` 時會推送分支;若 `gh` 可用,會嘗試建立 PR。沒有 `origin` 時,完成的分支留在本機。
39
39
 
@@ -49,7 +49,8 @@ agentflowctl 會依專案的 `packageManager`、lockfile 與 `package.json` scri
49
49
 
50
50
  ```bash
51
51
  agentflowctl list # 列出 run
52
- agentflowctl status f-xxxx # 看進度、結果與下一步
52
+ agentflowctl status f-xxxx # 看進度、結果與下一步(任務與待人確認分區)
53
+ agentflowctl confirmations f-xxxx # 只列出需要人眼確認的任務
53
54
  agentflowctl logs f-xxxx # 列出各步驟的 log
54
55
  agentflowctl logs f-xxxx --latest # 看最新一份 log
55
56
  agentflowctl stats f-xxxx # 各步驟耗時、執行與失敗次數
@@ -57,11 +58,11 @@ agentflowctl insights # 這個專案所有 run 的結果、失敗
57
58
  agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
58
59
  ```
59
60
 
60
- `status` 會列出目前階段、未結的交接事項與下一步指令;失敗或暫停時也會顯示原因。要看某一步的詳細輸出,可用 `logs <id> <編號>`;加 `--full` 看完整工具內容,或加 `--raw` 看原始輸出。
61
+ `status` 會列出目前階段、未結的交接事項與下一步指令;失敗或暫停時也會顯示原因。任務進度與「待你確認(不進實作)」分成兩個區塊:實作任務在「任務」,`kind: confirm` 的項目在「待你確認(不進實作)」,兩邊都會列出。計畫還沒通過首次驗證時,這個區塊的標題會多帶「(計畫尚未定案,以下為草稿)」,代表清單是從尚未經檢查的草稿蒐集,項目最終可能不會定案。只想看要人眼確認的項目時,用 `confirmations <id>`。要看某一步的詳細輸出,可用 `logs <id> <編號>`;加 `--full` 看完整工具內容,或加 `--raw` 看原始輸出。
61
62
 
62
- `stats` 依 log 的開始與結束時間統計每個步驟的執行次數、失敗次數、總耗時與最長一次,並分開列出 agent 與專案指令(install、測試、checks)各占多少時間,最耗時的步驟排在最前面。沒有結束紀錄的 log 列為未完成,不計入耗時;總經過時間包含暫停與等待核准。
63
+ `stats` 依 log 的開始與結束時間統計每個步驟的執行次數、失敗次數、總耗時與最長一次,並分開列出 agent 與專案指令(install、測試、checks)各占多少時間,最耗時的步驟排在最前面。沒有結束紀錄的 log 列為未完成,不計入耗時;總經過時間包含暫停與等待核准。紅燈階段的測試指令(`T-<n>-red`)失敗是預期結果,只有測試意外通過才計為失敗。
63
64
 
64
- `insights` 把所有 run 分成獨立區塊彙總:最終狀態、失敗原因(重試達上限、仲裁停止、agent 次數用完等)、用量(輸入/輸出/cache、強度占比、階段/agent,各 run 用量最高的任務)、各模型與步驟、步驟執行與失敗(與 `stats` 相同取 log 的檔頭檔尾,跨 run 不計總經過時間)、關卡重試原因。任務關卡與任務步驟不分 task id 合併計算。最後列出最多五則建議,規則由程式套門檻,不是再請 agent 分析。合計 token 不是主指標。覆蓋不足時會先警告占比可能失真。各分組是同一批呼叫的不同切片,不要跨組相加。舊 run 沒有重試或失敗原因紀錄,不會回填。單一 run 的全量明細仍用 `status <id>` 與 `stats <id>`。
65
+ `insights` 把所有 run 分成獨立區塊彙總:最終狀態、失敗原因(重試達上限、仲裁停止、agent 次數用完等)、用量(輸入/輸出/cache、強度占比、階段/agent,各 run 用量最高的任務)、各模型與步驟、步驟執行與失敗(與 `stats` 相同取 log 的檔頭檔尾,跨 run 不計總經過時間)、關卡重試原因。任務關卡與任務步驟不分 task id 合併計算。最後列出最多五則建議,規則由程式套門檻,不是再請 agent 分析。「輸入遠大於輸出」不計入 cache 讀取(Claude 的 cache 寫入仍計入),門檻是非 cache 的輸入與輸出合計至少 5000 tokens 且輸入佔 85% 以上;cache 讀取量另外列在說明中。合計 token 不是主指標。覆蓋不足時會先警告占比可能失真。各分組是同一批呼叫的不同切片,不要跨組相加。舊 run 沒有重試或失敗原因紀錄,不會回填。單一 run 的全量明細仍用 `status <id>` 與 `stats <id>`。
65
66
 
66
67
  執行紀錄在 `.agentflowctl/runs/<id>/`,工作分支在 `.agentflowctl/worktrees/<id>/`。不再需要某次 run 時,可用 `agentflowctl clean <id>` 清除 worktree 與紀錄;`agentflowctl clean --all` 一次清除所有 done、failed 的 run,以及沒有紀錄的 worktree(進行中、暫停、等待核准的不動)。`flow/<id>` 分支會保留。
67
68
 
@@ -79,6 +80,8 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
79
80
  | `failed`:測試、檢查、審查或 agent 執行失敗 | 依 `status` 提示查看失敗的 log,處理原因後執行 `agentflowctl resume <id>`;失敗階段會重試 |
80
81
  | `failed`:已達 agent 執行次數上限 | 用 `agentflowctl resume <id> --max-agent-runs 100` 調高上限後接續,數字須大於已執行次數 |
81
82
 
83
+ 寫作類步驟(規格、計畫、計畫修正、紅燈測試、綠燈實作、fix)完成且通過關卡後,若 `.flow/handoff-response.json` 不合格,會請同一家 agent 再呼叫一次,只補寫交接(不重做工作,也不換人代打);補寫期間對其他檔案的變更與 commit 一律丟棄(規格與計畫類步驟沒有 commit,補寫者會被告知範圍為空,不會把事項標成已處理,補寫期間對規格與計畫檔的修改同樣丟棄)。補寫仍不合格或額度用完,才照舊還原這一步並重試。補寫呼叫記在原步驟名稱底下,`stats` 與 `logs` 會多一筆,也會計入 `--max-agent-runs` 的次數。
84
+
82
85
  例如失敗時,可照終端機列出的 log 編號查看原因:
83
86
 
84
87
  ```bash
package/dist/cli.js CHANGED
@@ -1,6 +1,6 @@
1
1
  #!/usr/bin/env node
2
2
  import { Command } from "commander";
3
- import { readFileSync } from "node:fs";
3
+ import { existsSync, readFileSync } from "node:fs";
4
4
  import { join } from "node:path";
5
5
  import { stdin, stdout } from "node:process";
6
6
  import { createInterface } from "node:readline/promises";
@@ -12,11 +12,12 @@ import { cleanableRuns, cleanRun } from "./cleanup.js";
12
12
  import { describeDetected, detectProjectDefaults } from "./detect.js";
13
13
  import { CMD_AGENT, listLogs, localTime, logMark, nextLogFile, renderLog } from "./logs.js";
14
14
  import { flowDir, logDir, projectRoot, worktreeDir } from "./paths.js";
15
- import { ModelStage, ModelStrength, StopAfterStage, TaskList } from "./schemas.js";
15
+ import { ModelStage, ModelStrength, OrderedTaskList, StopAfterStage } from "./schemas.js";
16
16
  import { computeInsights, failureLabel, retryLabel } from "./insights.js";
17
17
  import { computeUsageInsights } from "./usageInsights.js";
18
18
  import { computeStats, formatDuration } from "./stats.js";
19
19
  import { agentRuns, getRun, listRetries, listRuns, listSubstitutions, listUsage, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
20
+ import { confirmationLines, confirmationTasks } from "./tasks.js";
20
21
  import { padDisplay, readJsonFile } from "./util.js";
21
22
  import { openActions, readHandoff } from "./handoff.js";
22
23
  import { stopReport } from "./stopReport.js";
@@ -122,6 +123,19 @@ async function resolveCycle(flag) {
122
123
  }
123
124
  /** status 任務清單中,進行中任務的標記 */
124
125
  const TASK_PHASE_MARK = { tests: "🧪", code: "🛠️ ", review: "👀", verify: "🔍", fix: "🩹" };
126
+ /**
127
+ * 這個 run 需要人眼確認的任務。確認清單已寫出(計畫通過首次驗證後)就用它,draft 為 false;
128
+ * 還沒寫出時從計畫與實作清單蒐集 kind: confirm 的任務,這些項目尚未經 validatePlan 檢查,draft 為 true。
129
+ * ordered 可傳入呼叫端已讀過的 tasks.ordered.json,避免重複讀檔。
130
+ */
131
+ function loadConfirmationTasks(id, ordered = readJsonFile(join(flowDir(id), "tasks.ordered.json"), OrderedTaskList)) {
132
+ const savedPath = join(flowDir(id), "confirmations.json");
133
+ const savedExists = existsSync(savedPath);
134
+ const saved = savedExists ? readJsonFile(savedPath, OrderedTaskList) : undefined;
135
+ const planned = readJsonFile(join(flowDir(id), "tasks.json"), OrderedTaskList);
136
+ const tasks = confirmationTasks(saved ? (saved.ok ? saved.data : []) : undefined, ordered.ok ? ordered.data : [], planned.ok ? planned.data : []);
137
+ return { tasks, draft: !savedExists };
138
+ }
125
139
  const program = new Command()
126
140
  .name("agentflowctl")
127
141
  .description("在專案資料夾內執行的 Agent 開發流程:規格 → 計畫 → TDD 實作 → 驗證 → 審查 → PR")
@@ -308,15 +322,28 @@ program
308
322
  console.log(` ${r.key.padEnd(19)} ${retryLabel(r.category)}${r.final ? "(達上限)" : `(第 ${r.attempt} 次)`}`);
309
323
  }
310
324
  }
311
- const tasks = readJsonFile(join(flowDir(id), "tasks.ordered.json"), TaskList);
312
- if (!tasks.ok)
313
- return;
314
- console.log("\n任務");
315
- tasks.data.forEach((t, i) => {
316
- const active = i === run.taskIndex && run.stage === "implement";
317
- const mark = i < run.taskIndex ? "✅" : active ? TASK_PHASE_MARK[run.taskPhase] : "⬜";
318
- console.log(` ${mark} ${t.id} ${t.title}`);
319
- });
325
+ const tasks = readJsonFile(join(flowDir(id), "tasks.ordered.json"), OrderedTaskList);
326
+ if (tasks.ok && tasks.data.some((t) => t.kind !== "confirm")) {
327
+ console.log("\n任務");
328
+ tasks.data.forEach((t, i) => {
329
+ if (t.kind === "confirm")
330
+ return;
331
+ const active = i === run.taskIndex && run.stage === "implement";
332
+ const mark = i < run.taskIndex ? "✅" : active ? TASK_PHASE_MARK[run.taskPhase] : "⬜";
333
+ console.log(` ${mark} ${t.id} ${t.title}`);
334
+ });
335
+ }
336
+ const confirm = loadConfirmationTasks(id, tasks);
337
+ if (confirm.tasks.length)
338
+ console.log(`\n${confirmationLines(confirm.tasks, confirm.draft).join("\n")}`);
339
+ });
340
+ program
341
+ .command("confirmations <id>")
342
+ .description("只列出這個 run 需要人眼確認的任務")
343
+ .action((id) => {
344
+ mustGetRun(id);
345
+ const confirm = loadConfirmationTasks(id);
346
+ console.log(confirmationLines(confirm.tasks, confirm.draft).join("\n"));
320
347
  });
321
348
  // ───────────── agent 管理:讀寫 flow.config.json 的 agents 與 cycle ─────────────
322
349
  const configPath = () => join(projectRoot(), "flow.config.json");
package/dist/engine.js CHANGED
@@ -7,13 +7,13 @@ import { escapeXml, opinion, reviewIssue } from "./feedback.js";
7
7
  import { changedFiles, commitAll, discardChanges, git, headCommit, resetTo } from "./git.js";
8
8
  import { acceptHandoff, openActions, prepareHandoff, previewHandoff, readHandoff, recoverHandoff, responsePath, reviewHandoffGate, validateHandoffResponse } from "./handoff.js";
9
9
  import { flowDir, logDir, planArbitrationPath, planReviewStatePath, projectRoot, runDir, worktreeDir } from "./paths.js";
10
- import { CMD_AGENT, nextLogFile } from "./logs.js";
10
+ import { CMD_AGENT, nextLogFile, redStepName } from "./logs.js";
11
11
  import { exec } from "./proc.js";
12
12
  import { arbiterPanel, availableAgent, fixAgent, planAgent, planFixAgent, reviewers, specAgent, taskAgents } from "./roles.js";
13
13
  import { dropCall, loadCalls, openRound, runPool, saveCall, storedCallValid } from "./parallelReview.js";
14
14
  import { cleanupTempWorktrees, withTempWorktree } from "./tempWorktree.js";
15
15
  import { resolveAgent, runAgent, runCommand } from "./runner.js";
16
- import { AcceptanceList, ArbiterResult, ConsistentReviewResult, RepoConfig, TaskList, } from "./schemas.js";
16
+ import { AcceptanceList, ArbiterResult, ConsistentReviewResult, RepoConfig, OrderedTaskList, TaskList, } from "./schemas.js";
17
17
  import { addRetry, addSubstitution, addUsage, agentRuns, saveRun } from "./store.js";
18
18
  import { clearModelReviewFailure, clearModelReviewStage, recordModelReviewFailure, selectModel } from "./modelSelection.js";
19
19
  import { applyReviewVerdicts, dirtyGroups, dirtyTaskIds, extractPlanEvidence, groupReviewerCount, layeredReview, neighborTasks, planContentKey, planOverview, planReviewIndex, readPendingArbitration, readPlanReviewState, repliesForTasks, reviewFingerprint, roundProgress, } from "./planReview.js";
@@ -57,8 +57,8 @@ async function agentStep(run, planned, step, prompt, mode) {
57
57
  let agent = planned;
58
58
  for (;;) {
59
59
  if (exhausted.has(agent)) {
60
- if (mode.kind === "review") {
61
- throw new QuotaPause(`${agent} 的額度已用完;${step} 是審查步驟,不由另一家代打`);
60
+ if (mode.kind === "review" || mode.pinned) {
61
+ throw new QuotaPause(`${agent} 的額度已用完;${step} ${mode.pinned ? "不由另一家代打" : "是審查步驟,不由另一家代打"}`);
62
62
  }
63
63
  const sub = availableAgent(run.cycle, agent, [...exhausted], `${run.id}:${step}:sub`);
64
64
  if (!sub)
@@ -112,6 +112,38 @@ base) {
112
112
  return error.message;
113
113
  }
114
114
  }
115
+ /**
116
+ * 寫作步驟的交接:關卡已通過,交接回覆不合格時不丟掉工作,請同一家 agent 只補寫一次。
117
+ * 補寫前記下 HEAD 與 files 的內容,補寫後一律 reset 回去並還原這些檔案,連補寫自行建立的 commit 也丟棄,
118
+ * 所以補寫無法改變已通過的關卡。補寫仍失敗(或額度用完)就回傳錯誤,由呼叫端照舊還原並重試。
119
+ * base 是這一步開始前的 commit;沒有 commit 的步驟(規格、計畫)省略,補寫者會被告知範圍為空。
120
+ */
121
+ async function settleHandoff(run, outcome, files, base) {
122
+ const first = finishHandoff(run, outcome, "writer");
123
+ if (!first)
124
+ return undefined;
125
+ const repo = worktreeDir(run.id);
126
+ const settled = await headCommit(repo);
127
+ const snap = snapshotPlan(run, files);
128
+ const cleanup = async () => { await resetTo(repo, settled); restorePlan(run, snap); };
129
+ info(run, `📎 交接回覆不合格,請 ${outcome.agent} 只補寫交接(不重做工作):${first}`);
130
+ let repair;
131
+ try {
132
+ // pinned 不找代打,額度用完直接暫停,所以不需要 reset;善後一律交給 finally
133
+ repair = await agentStep(run, outcome.agent, outcome.step, renderPrompt("handoff-repair", { step: outcome.step, error: first, range: `${base ?? settled}..${settled}` }), { kind: "write", pinned: true, reset: () => { } });
134
+ }
135
+ catch (error) {
136
+ if (error instanceof QuotaPause)
137
+ return first;
138
+ throw error;
139
+ }
140
+ finally {
141
+ await cleanup();
142
+ }
143
+ if (!repair.r.ok)
144
+ return `${first}(補寫交接時 Agent 執行失敗:${repair.r.summary})`;
145
+ return finishHandoff(run, repair, "writer");
146
+ }
115
147
  /** 印出回覆裡的 XML 中繼資料;只供人檢視,關卡仍由程式檢查決定 */
116
148
  function reportMeta(run, agent, r) {
117
149
  if (!r.meta)
@@ -226,11 +258,30 @@ function testsRedGuidance(testCmd, waiveRed) {
226
258
  };
227
259
  }
228
260
  function loadOrderedTasks(run) {
229
- const r = readJsonFile(flowFile(run, "tasks.ordered.json"), TaskList);
261
+ const r = readJsonFile(flowFile(run, "tasks.ordered.json"), OrderedTaskList);
230
262
  if (!r.ok)
231
263
  throw new Error(r.error);
232
264
  return r.data;
233
265
  }
266
+ function loadConfirmations(run) {
267
+ if (!existsSync(flowFile(run, "confirmations.json")))
268
+ return [];
269
+ const r = readJsonFile(flowFile(run, "confirmations.json"), OrderedTaskList);
270
+ return r.ok ? r.data : [];
271
+ }
272
+ function rememberConfirmation(run, task) {
273
+ const current = loadConfirmations(run);
274
+ if (current.some((item) => item.id === task.id))
275
+ return;
276
+ writeFileSync(flowFile(run, "confirmations.json"), JSON.stringify([...current, task], null, 2));
277
+ }
278
+ function announceConfirmations(run) {
279
+ const confirm = loadConfirmations(run);
280
+ if (!confirm.length)
281
+ return;
282
+ info(run, `👀 另有 ${confirm.length} 項需要你確認(不進實作,不會停下):${confirm.map((t) => `${t.id} ${t.title}`).join("、")}`);
283
+ info(run, ` 清單:${flowFile(run, "confirmations.json")}`);
284
+ }
234
285
  // ───────────────────────── 各階段 ─────────────────────────
235
286
  async function specStage(run) {
236
287
  const agent = specAgent(run.cycle, run.id);
@@ -248,7 +299,7 @@ async function specStage(run) {
248
299
  const ids = ac.data.map((a) => a.id);
249
300
  if (new Set(ids).size !== ids.length)
250
301
  return retry(run, "spec", "驗收條件 id 有重複", "spec", "format_invalid");
251
- const handoffError = finishHandoff(run, outcome, "writer");
302
+ const handoffError = await settleHandoff(run, outcome, PLAN_FILES);
252
303
  if (handoffError)
253
304
  return retry(run, "spec", handoffError, "spec", "handoff_invalid");
254
305
  return succeed(run, "spec", "plan");
@@ -256,7 +307,7 @@ async function specStage(run) {
256
307
  // ── 規格與計畫檔案:計畫審查、仲裁與計畫定案後的所有階段只能讀,不能改 ──
257
308
  const PLAN_FILES = ["spec.md", "acceptance.json", "plan.md", "tasks.json"];
258
309
  /** 計畫定案後(實作、修正、程式碼審查)另外依賴排好的任務順序,同樣不能被改 */
259
- const LOCKED_FILES = [...PLAN_FILES, "tasks.ordered.json"];
310
+ const LOCKED_FILES = [...PLAN_FILES, "tasks.ordered.json", "confirmations.json"];
260
311
  /** 審查、修訂與仲裁的快照另外包含審查回應;不要併進 PLAN_FILES,定案後的階段不依賴它 */
261
312
  const PLAN_REPLY_FILES = [...PLAN_FILES, "plan-replies.md"];
262
313
  function snapshotPlan(run, files = PLAN_FILES) {
@@ -301,7 +352,10 @@ function validatePlan(run) {
301
352
  return orderTasks(tasks.data, new Set(ids));
302
353
  }
303
354
  function acceptPlan(run, ordered) {
304
- writeFileSync(flowFile(run, "tasks.ordered.json"), JSON.stringify(ordered, null, 2));
355
+ const confirm = ordered.filter((t) => t.kind === "confirm");
356
+ const implement = ordered.filter((t) => t.kind !== "confirm");
357
+ writeFileSync(flowFile(run, "tasks.ordered.json"), JSON.stringify(implement, null, 2));
358
+ writeFileSync(flowFile(run, "confirmations.json"), JSON.stringify(confirm, null, 2));
305
359
  }
306
360
  function announceTasks(run, ordered) {
307
361
  const cfg = loadRepoConfig();
@@ -322,6 +376,7 @@ function planSettled(run, key) {
322
376
  const next = run.autopilot ? "implement" : "awaiting_approval";
323
377
  const ordered = loadOrderedTasks(run);
324
378
  announceTasks(run, ordered);
379
+ announceConfirmations(run);
325
380
  if (!run.autopilot) {
326
381
  info(run, `✋ 計畫已通過審查,請檢視 ${flowFile(run, "plan.md")},確認後執行 agentflowctl approve ${run.id}`);
327
382
  }
@@ -342,7 +397,7 @@ async function planStage(run) {
342
397
  const ordered = validatePlan(run);
343
398
  if (typeof ordered === "string")
344
399
  return retry(run, "plan", ordered, "plan", "format_invalid");
345
- const handoffError = finishHandoff(run, outcome, "writer");
400
+ const handoffError = await settleHandoff(run, outcome, PLAN_FILES);
346
401
  if (handoffError)
347
402
  return retry(run, "plan", handoffError, "plan", "handoff_invalid");
348
403
  acceptPlan(run, ordered);
@@ -707,7 +762,7 @@ async function planFixStage(run) {
707
762
  restorePlan(run, snap);
708
763
  return retry(run, "plan-fix", `${feedback}\n\n另外,修改後的計畫沒有通過格式檢查,已還原:\n${ordered}`, "plan_fix", "format_invalid");
709
764
  }
710
- const handoffError = finishHandoff(run, outcome, "writer");
765
+ const handoffError = await settleHandoff(run, outcome, PLAN_REPLY_FILES);
711
766
  if (handoffError) {
712
767
  restorePlan(run, snap);
713
768
  // 計畫已還原,要保留原本的審查意見,否則下一次修正不知道要改什麼
@@ -802,6 +857,7 @@ async function implementStage(run) {
802
857
  const task = tasks[run.taskIndex];
803
858
  if (!task) {
804
859
  info(run, "✅ 所有任務完成");
860
+ announceConfirmations(run);
805
861
  return to(run, "verify");
806
862
  }
807
863
  const cfg = loadRepoConfig();
@@ -809,6 +865,13 @@ async function implementStage(run) {
809
865
  const testRe = new RegExp(cfg.testPattern);
810
866
  const testCmd = `${cfg.install} && ${cfg.test}`;
811
867
  const progress = `${run.taskIndex + 1}/${tasks.length} ${task.id} ${task.title}`;
868
+ if (task.kind === "confirm") {
869
+ rememberConfirmation(run, task);
870
+ const rest = tasks.filter((item) => item.id !== task.id);
871
+ writeFileSync(flowFile(run, "tasks.ordered.json"), JSON.stringify(rest, null, 2));
872
+ info(run, `👀 [${progress}] 改放到待你確認的清單,實作繼續`);
873
+ return run;
874
+ }
812
875
  const taskJson = JSON.stringify(task, null, 2);
813
876
  const acceptance = readJsonFile(flowFile(run, "acceptance.json"), AcceptanceList);
814
877
  if (!acceptance.ok)
@@ -857,12 +920,12 @@ async function implementStage(run) {
857
920
  await resetTo(repo, before);
858
921
  return retry(run, key, `沒有新增或修改任何符合 /${cfg.testPattern}/ 的測試檔。`, "implement", "tests_not_written");
859
922
  }
860
- const red = await runCommand(target(run, `${task.id}-red`, CMD_AGENT), testCmd);
923
+ const red = await runCommand(target(run, redStepName(task.id), CMD_AGENT), testCmd);
861
924
  if (red.ok && !waiveRed) {
862
925
  await resetTo(repo, before);
863
926
  return retry(run, key, "測試在功能尚未實作前就全部通過,代表測試沒有驗證到新行為。請撰寫會因功能尚未實作而失敗的測試。", "implement", "tests_not_red");
864
927
  }
865
- const handoffError = finishHandoff(run, outcome, "writer");
928
+ const handoffError = await settleHandoff(run, outcome, LOCKED_FILES, before);
866
929
  if (handoffError) {
867
930
  await resetTo(repo, before);
868
931
  return retry(run, key, handoffError, "implement", "handoff_invalid");
@@ -894,8 +957,10 @@ async function implementStage(run) {
894
957
  return retry(run, key, planTamperedMessage(tampered), "implement", "plan_tampered");
895
958
  }
896
959
  const codeCommit = await commitAll(repo, `feat(${task.id}): ${task.title} [${codeAuthor}]`);
897
- if (!tdd && !codeCommit)
898
- return retry(run, key, "沒有任何檔案變更,這個任務必須完成實作。", "implement", "code_not_written");
960
+ if (!tdd && !codeCommit) {
961
+ info(run, `⏭️ [${progress}] 沒有檔案變更,略過這個任務`);
962
+ return finishTask(succeed(run, key, "implement"));
963
+ }
899
964
  const touched = tdd ? (await changedFiles(repo, testsCommit, await headCommit(repo))).filter((f) => testRe.test(f)) : [];
900
965
  if (touched.length) {
901
966
  await resetTo(repo, testsCommit);
@@ -907,7 +972,7 @@ async function implementStage(run) {
907
972
  info(run, ` ✗ 測試仍未通過${logHint(run, green.seq)}`);
908
973
  return retry(run, key, `測試仍未通過:\n\n\`\`\`\n${tail(green.output)}\n\`\`\``, "implement", "tests_not_green");
909
974
  }
910
- const handoffError = finishHandoff(run, outcome, "writer");
975
+ const handoffError = await settleHandoff(run, outcome, LOCKED_FILES, testsCommit);
911
976
  if (handoffError) {
912
977
  await resetTo(repo, testsCommit);
913
978
  return retry(run, key, handoffError, "implement", "handoff_invalid");
@@ -948,6 +1013,17 @@ async function taskReviewStep(run, task, progress, taskJson, acceptanceJson) {
948
1013
  lastReviewer: result.objector,
949
1014
  };
950
1015
  }
1016
+ /** 這個任務結束,下一個從寫測試開始 */
1017
+ function finishTask(run) {
1018
+ return {
1019
+ ...run,
1020
+ taskIndex: run.taskIndex + 1,
1021
+ taskPhase: "tests",
1022
+ taskBase: undefined,
1023
+ testsCommit: undefined,
1024
+ lastTestsAuthor: undefined,
1025
+ };
1026
+ }
951
1027
  // ── 任務驗證:通過才進入下一個任務 ──
952
1028
  async function taskVerifyStep(run, task, progress) {
953
1029
  const key = `${task.id}:verify`;
@@ -956,14 +1032,7 @@ async function taskVerifyStep(run, task, progress) {
956
1032
  if (report)
957
1033
  return { ...retry(run, key, report, "implement", "checks_failed"), taskPhase: "fix", fixSource: "verify" };
958
1034
  info(run, `✅ [${progress}] 完成`);
959
- return {
960
- ...succeed(run, key, "implement"),
961
- taskIndex: run.taskIndex + 1,
962
- taskPhase: "tests",
963
- taskBase: undefined,
964
- testsCommit: undefined,
965
- lastTestsAuthor: undefined,
966
- };
1035
+ return finishTask(succeed(run, key, "implement"));
967
1036
  }
968
1037
  // ── 任務修正:修完重新審查、驗證 ──
969
1038
  async function taskFixStep(run, task, progress) {
@@ -1045,7 +1114,7 @@ async function applyFix(run, opts) {
1045
1114
  await resetTo(repo, before);
1046
1115
  return again(`${feedback}\n\n另外:不可刪除測試檔來讓檢查通過,已還原:${deleted.join(", ")}`, "tests_deleted");
1047
1116
  }
1048
- const handoffError = finishHandoff(run, outcome, "writer");
1117
+ const handoffError = await settleHandoff(run, outcome, LOCKED_FILES, before);
1049
1118
  if (handoffError) {
1050
1119
  await resetTo(repo, before);
1051
1120
  return again(`${feedback}\n\n另外,交接回覆不合格,本次修正已還原:${handoffError}`, "handoff_invalid");
package/dist/logs.js CHANGED
@@ -15,6 +15,9 @@ export const FOOTER_PREFIX = "# exit ";
15
15
  export const STDERR_MARK = "[stderr]";
16
16
  /** 專案指令(install、測試、checks)的 agent 欄位 */
17
17
  export const CMD_AGENT = "cmd";
18
+ /** 紅燈測試指令的步驟名稱;engine 寫 log 與 stats 判斷預期失敗都用這裡,命名不會各改各的 */
19
+ export const redStepName = (taskId) => `${taskId}-red`;
20
+ export const isRedStep = (step) => /^T-\d+-red$/.test(step);
18
21
  const segment = (s) => s.replace(/[\s/\\:*?"<>|]+/g, "_");
19
22
  /** 下一份 log 的完整路徑;序號接在目錄裡最大的序號之後 */
20
23
  export function nextLogFile(dir, stage, step, agent) {
@@ -193,6 +193,7 @@ export function taskFingerprint(task, acceptance, planMd) {
193
193
  dependsOn: task.dependsOn,
194
194
  acceptance: task.acceptance.map((id) => ({ id, description: byId.get(id) ?? "" })),
195
195
  tdd: task.tdd ?? null,
196
+ kind: task.kind ?? null,
196
197
  evidence: extractPlanEvidence(planMd, [task.id]),
197
198
  });
198
199
  }
package/dist/schemas.js CHANGED
@@ -76,9 +76,13 @@ export const TaskItem = z.object({
76
76
  complexity: z.enum(["low", "medium", "high"]).optional(),
77
77
  /** false=這個任務不適合先寫會失敗的測試(建置流程、設定、文件、純重構、實作前就會通過的特徵化測試等),略過紅燈直接實作;沒寫視為 true。描述寫明不要求紅燈時必須為 false */
78
78
  tdd: z.boolean().optional(),
79
+ /** confirm=不進入實作佇列,另存給使用者確認;沒寫視為要實作。舊 tasks.json 沒有此欄位 */
80
+ kind: z.enum(["implement", "confirm"]).optional(),
79
81
  });
82
+ /** 排好的實作清單與另存的確認清單都可以是空的 */
83
+ export const OrderedTaskList = z.array(TaskItem);
80
84
  /** Agent 在 plan 階段產出的 .flow/tasks.json */
81
- export const TaskList = z.array(TaskItem).min(1);
85
+ export const TaskList = OrderedTaskList.min(1);
82
86
  /** Agent 在 review 階段產出的 .flow/review.json */
83
87
  export const ReviewResult = z.object({
84
88
  verdict: z.enum(["approve", "changes_requested"]),
package/dist/stats.js CHANGED
@@ -1,5 +1,7 @@
1
- import { CMD_AGENT } from "./logs.js";
1
+ import { CMD_AGENT, isRedStep } from "./logs.js";
2
2
  const time = (iso) => (iso ? new Date(iso).getTime() : NaN);
3
+ /** 紅燈階段的測試指令:失敗才是預期結果,意外通過才代表這一步沒過關 */
4
+ const isRedCommand = (kind, step) => kind === "cmd" && isRedStep(step);
3
5
  /** 從 log 的檔頭與檔尾算出每個步驟的次數、失敗與耗時,不讀 log 內容 */
4
6
  export function computeStats(entries) {
5
7
  const byKey = new Map();
@@ -23,7 +25,7 @@ export function computeStats(entries) {
23
25
  unfinished += 1;
24
26
  continue;
25
27
  }
26
- if (!footer.ok)
28
+ if (isRedCommand(kind, header.step) ? footer.ok : !footer.ok)
27
29
  s.failed += 1;
28
30
  const ms = Math.max(0, end - start);
29
31
  s.totalMs += ms;
package/dist/tasks.js CHANGED
@@ -84,4 +84,36 @@ export function orderTasks(tasks, acceptanceIds) {
84
84
  }
85
85
  return ordered;
86
86
  }
87
+ /**
88
+ * 需要人確認的任務。confirmations.json 已寫出時以它為準(空陣列代表沒有);
89
+ * 還沒寫出時,從實作清單與計畫裡蒐集 kind 為 confirm 的任務。
90
+ */
91
+ export function confirmationTasks(saved, ordered, planned) {
92
+ if (saved)
93
+ return saved;
94
+ const seen = new Set();
95
+ const out = [];
96
+ for (const task of [...ordered, ...planned]) {
97
+ if (task.kind !== "confirm" || seen.has(task.id))
98
+ continue;
99
+ seen.add(task.id);
100
+ out.push(task);
101
+ }
102
+ return out;
103
+ }
104
+ /**
105
+ * 只列出待人確認的任務;沒有時回一句說明。
106
+ * draft 代表計畫還沒通過首次驗證,清單是從未經 validatePlan 檢查的草稿蒐集來的,可能有項目最終不會定案。
107
+ */
108
+ export function confirmationLines(tasks, draft = false) {
109
+ if (!tasks.length)
110
+ return ["沒有需要人確認的任務"];
111
+ const lines = [`待你確認(不進實作)${draft ? "(計畫尚未定案,以下為草稿)" : ""}`];
112
+ for (const task of tasks) {
113
+ lines.push(` ${task.id} ${task.title}`);
114
+ lines.push(` ${task.description}`);
115
+ lines.push(` 驗收:${task.acceptance.join("、")}`);
116
+ }
117
+ return lines;
118
+ }
87
119
  //# sourceMappingURL=tasks.js.map
@@ -63,10 +63,12 @@ export function usageFindings(input) {
63
63
  detail: `重試 ${retries.length} 次(不含審查要求修改與仲裁要求修訂),最多「${top ? retryLabel(top[0]) : "未知"}」(${top?.[1] ?? 0} 次)。以平均每次 ${avg} tokens 估算,重試約 ${waste} tokens。先改該關卡,再縮 prompt。`,
64
64
  });
65
65
  }
66
- if (total.tokens >= 5000 && total.inputTokens / total.tokens >= 0.85) {
66
+ const freshInput = Math.max(0, total.inputTokens - total.cacheReadTokens);
67
+ const freshTotal = freshInput + total.outputTokens;
68
+ if (freshTotal >= 5000 && freshInput / freshTotal >= 0.85) {
67
69
  out.push({
68
- code: "input_heavy", impactTokens: total.inputTokens, title: FINDING_LABEL.input_heavy,
69
- detail: `輸入 ${total.inputTokens}、輸出 ${total.outputTokens}(輸入佔 ${pct(total.inputTokens, total.tokens)})。上下文可能太大:規格、計畫、測試輸出或一次讀太多檔。`,
70
+ code: "input_heavy", impactTokens: freshInput, title: FINDING_LABEL.input_heavy,
71
+ detail: `不含 cache 讀取的輸入 ${freshInput}、輸出 ${total.outputTokens}(輸入佔 ${pct(freshInput, freshTotal)});另有 cache 讀取 ${total.cacheReadTokens} 未計入。上下文可能太大:規格、計畫、測試輸出或一次讀太多檔。`,
70
72
  });
71
73
  }
72
74
  const stagesWithTokens = Object.values(byStage).filter((s) => s.tokens > 0).length;
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "agentflowctl",
3
3
  "license": "MIT",
4
- "version": "0.17.1",
4
+ "version": "0.19.0",
5
5
  "description": "跨廠商 AI 開發 harness:Claude Code、Codex、Gemini 輪流實作、審查、修正",
6
6
  "keywords": [
7
7
  "ai",
package/prompts/fix.md CHANGED
@@ -32,7 +32,7 @@
32
32
  - 不可刪除測試檔(檔名符合 `{{testPattern}}`),也不可用 skip、放寬斷言、`@ts-ignore`、`eslint-disable` 等方式讓檢查通過。
33
33
  - 如果測試本身確實有誤,可以修正測試,但必須在回覆的 `<concerns>` 說明理由。
34
34
  - 不要執行 git commit(權限設定已禁止)。
35
- - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
35
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json、.flow/confirmations.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
36
36
  </constraints>
37
37
 
38
38
  <reply_format>
@@ -0,0 +1,51 @@
1
+ <role>
2
+ 你是交接紀錄員,只負責補寫上一個步驟沒有寫好的 .flow/handoff-response.json。工作本身已經完成並通過檢查,不要重做。
3
+ </role>
4
+
5
+ <context>
6
+ 目前的工作目錄就是專案(agentflowctl 為這次任務建立的專用 git worktree)。
7
+ 步驟 {{step}} 的工作已完成,變更範圍是 `{{range}}`,但交接回覆沒有通過檢查,原因:
8
+
9
+ {{error}}
10
+ </context>
11
+
12
+ <handoff>
13
+ 先閱讀 .flow/handoff-context.md,確認有哪些待處理事項。再寫入 .flow/handoff-response.json;即使沒有事項也必須寫出空陣列:
14
+
15
+ ```json
16
+ { "newIssues": [], "dispositions": [] }
17
+ ```
18
+
19
+ 新增事項格式:{ "kind": "action 或 info", "summary": "具體問題", "evidence": "檔案位置或檢查證據", "targetStage": "plan 或 code" }。
20
+ 處置格式:{ "id": "既有事項 ID", "status": "proposed_resolved", "reason": "具體處理理由", "evidence": "檔案、commit 或檢查結果" }。只有「待處理事項」(action)可以處置,而且這個步驟的作者只能用 proposed_resolved;「參考資訊」(info)不要放進 dispositions。
21
+ </handoff>
22
+
23
+ <inputs>
24
+ - .flow/handoff-context.md:待處理事項與參考資訊
25
+ - `git log {{range}}` 與 `git diff {{range}}`:這一步實際做了什麼,處置的理由與證據要以它為準。範圍兩端相同代表這一步沒有產生 commit,這時不要把任何事項標成已處理
26
+ </inputs>
27
+
28
+ <steps>
29
+ 1. 讀 .flow/handoff-context.md 與上述 git 範圍,判斷每個待處理事項這一步有沒有真的處理;沒有把握的事項不要處置。
30
+ 2. 只寫入 .flow/handoff-response.json,格式如上。
31
+ </steps>
32
+
33
+ <constraints>
34
+ - 只能建立或修改 .flow/handoff-response.json。其他檔案的變更與任何 git commit 都會在你結束後被自動丟棄,不會生效。
35
+ - 不要重做或改善這一步的工作,也不要重新跑完整測試。
36
+ </constraints>
37
+
38
+ <reply_format>
39
+ 完成後,回覆的最後必須附上以下 XML 中繼資料(只附一次,標籤名稱不可更改):
40
+
41
+ ```xml
42
+ <result>
43
+ <status>done 或 blocked</status>
44
+ <summary>一兩句說明補寫了什麼;blocked 時說明卡在哪裡</summary>
45
+ <files_changed>
46
+ <file>.flow/handoff-response.json</file>
47
+ </files_changed>
48
+ <concerns>沒有就留空</concerns>
49
+ </result>
50
+ ```
51
+ </reply_format>
@@ -50,7 +50,7 @@
50
50
  <constraints>
51
51
  - **不可修改任何測試檔**,修改會被自動還原並視為失敗。若認為測試本身有誤,請寫在回覆的 `<concerns>`。
52
52
  - 不要執行 git commit(權限設定已禁止)。
53
- - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
53
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json、.flow/confirmations.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
54
54
  </constraints>
55
55
 
56
56
  <reply_format>
@@ -46,9 +46,9 @@
46
46
  </steps>
47
47
 
48
48
  <constraints>
49
- - 必須實際修改檔案;沒有任何變更會被視為失敗。
49
+ - 需要改的檔案就改。這個任務若沒有要寫進 git 的變更,不要為了產生 commit 而硬改檔案;沒有變更會略過這個任務,不會重試。
50
50
  - 不要執行 git commit(權限設定已禁止)。
51
- - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
51
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json、.flow/confirmations.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
52
52
  </constraints>
53
53
 
54
54
  <reply_format>
@@ -43,7 +43,7 @@
43
43
  <constraints>
44
44
  - 不可實作功能本身。可以建立讓測試能編譯所需的最小型別或空殼匯出,但不可以有真正的邏輯。
45
45
  - 不要執行 git commit(權限設定已禁止),外部流程會提交並{{verifyNote}}。
46
- - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
46
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json、.flow/confirmations.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
47
47
  </constraints>
48
48
 
49
49
  <reply_format>
@@ -33,6 +33,7 @@
33
33
  - 一個任務只做一件事,最多兩件:`acceptance` 最多列兩條驗收條件;驗收條件一條只描述一個行為。修改時若任務變大,請拆開,不要合併。
34
34
  - 測試檔名必須符合正規表示式 `{{testPattern}}`。
35
35
  - 修改 task 時保留或補上 `complexity`(`low`、`medium`、`high`)。依影響範圍、技術不確定性與失敗後果重新判定,取最高等級;同步更新 .flow/plan.md 中該 task 的逐項證據與最終等級。若不同意審查者建議的等級,在 .flow/plan-replies.md 對應的 `## T-<數字>` 節引用具體程式碼或測試依據。同步更新 .flow/plan.md 中該 task 的證據時,放在該 task 的 `## T-<數字>` 標題下。
36
+ - 只跑檢查、不改檔案的任務要併回會改檔的任務,不要留成獨立任務。需要人眼確認的任務設 `"kind": "confirm"`,它們會另存給使用者,不要留在實作佇列。
36
37
  </output_format>
37
38
 
38
39
  <constraints>
@@ -46,7 +46,7 @@
46
46
 
47
47
  <review_focus>
48
48
  只看這一群:
49
- 1. 每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出實作前會失敗的測試。無法先失敗的(特徵化、實作前就會通過、描述寫明不要求紅燈)必須是 `tdd: false`。只在描述補註、卻留下 `tdd: true` 或沒寫 `tdd`,仍是 `changes_requested`,`note` 要要求改成 `tdd: false`。
49
+ 1. 每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出實作前會失敗的測試。無法先失敗的(特徵化、實作前就會通過、描述寫明不要求紅燈)必須是 `tdd: false`。只在描述補註、卻留下 `tdd: true` 或沒寫 `tdd`,仍是 `changes_requested`,`note` 要要求改成 `tdd: false`。只跑既有檢查、不改檔案的工作不要獨立成任務,要求併回會改檔的任務。需要人眼確認的,要求改成 `"kind": "confirm"`,離開實作佇列另存給使用者。
50
50
  2. 任務描述是否對得上上面的驗收條文。
51
51
  3. 對照描述點名、已存在的檔案與難度摘錄,獨立核對 `complexity`。`low` 是沿用既有做法、侷限單一行為且失敗可由局部測試發現;`medium` 包括多模組或介面協調、非典型邊界、相容性或狀態遷移;`high` 包括跨系統契約、架構或資料模型變更、未知的關鍵路徑,或資料遺失、權限、難以回復的風險。摘錄是空的,而且高低估會影響選模時,要求在 .flow/plan.md 該 task 的「## T-<數字>」標題下補上證據。
52
52
  4. 做法是否符合那些檔案中已存在者的慣例。
@@ -30,7 +30,7 @@
30
30
  <review_focus>
31
31
  1. **需求覆蓋**:規格是否完整涵蓋原始需求?有沒有遺漏、誤解,或加入需求沒要求的範圍?
32
32
  2. **驗收條件**:每一條是否具體、可以用自動化測試驗證,而且只描述一個行為?把多個行為寫在同一條的,要求拆開。有沒有重要的邊界情況或錯誤處理沒被列入?
33
- 3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試(標 `tdd: false` 的任務除外:核對它的改動內容確實不適合先寫失敗測試,例如建置流程、設定、文件、型別、純重構,或實作前就會通過的特徵化測試,並有寫明驗收方式;會改變程式行為卻標成 `false` 的,要求改回 `true`。描述寫明不要求紅燈,`tdd` 卻不是 `false` 的,要求改成 `false`)?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?
33
+ 3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試(標 `tdd: false` 的任務除外:核對它的改動內容確實不適合先寫失敗測試,例如建置流程、設定、文件、型別、純重構,或實作前就會通過的特徵化測試,並有寫明驗收方式;會改變程式行為卻標成 `false` 的,要求改回 `true`。描述寫明不要求紅燈,`tdd` 卻不是 `false` 的,要求改成 `false`)?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?只跑既有檢查、不改檔案的工作不要獨立成任務,要求併回會改檔的任務。需要人眼確認、程式無法判定的,要求改成 `"kind": "confirm"`,讓它離開實作佇列、另存給使用者,不要留給 agent,也不要為此把流程停住。
34
34
  同時依 .flow/plan.md 的逐項理由及實際程式碼,獨立核對每個 task 的 `complexity`:分別看影響範圍、技術不確定性與失敗後果,取最高等級。`low` 須是沿用既有做法、侷限單一行為或模組且失敗可由局部測試發現;`medium` 包括多模組或介面協調、非典型邊界、相容性或狀態遷移風險;`high` 包括跨系統契約、架構或資料模型變更、未知的關鍵技術路徑,或資料遺失、權限、難以回復的風險。不要只憑檔案數、程式碼行數或驗收條件數判定。
35
35
  理由缺漏、與程式碼不符,或高低估會影響選模時,要求修正;在 `note` 指出 task ID、具體證據、建議等級及須修改的 .flow/plan.md/.flow/tasks.json 部分。不要為缺少高價值證據的細微措辭差異要求修改。
36
36
  4. **技術方向**:是否符合專案既有的架構與慣例?有沒有更簡單的做法,或明顯的風險?
package/prompts/plan.md CHANGED
@@ -53,6 +53,8 @@
53
53
  - `title` 用一句話說出這件事;需要用「並且」「以及」串起來的,就是兩個任務。
54
54
  - `description` 寫清楚要動哪些檔案(寫含目錄的路徑,例如 `src/form.ts`,不要只寫檔名)、測試要驗證哪個行為,以及這個任務不做什麼。計畫審查會依這些路徑把任務分群。
55
55
  - 依改動內容標記 `tdd`:會改變程式行為、能寫出「在實作前會失敗」的測試的任務標 `true`(預設);改動內容不適合先寫失敗測試的任務標 `false`,這類任務會略過紅燈直接實作,改由任務審查與驗證指令把關。適合標 `false` 的例子:建置流程與打包設定(build、CI、bundler、tsconfig)、依賴與版本設定、文件與 prompt 文字、樣式與靜態資源、型別宣告、不改變行為的重構與搬移檔案,以及鎖定既有行為的特徵化測試(實作前就會通過、不要求紅燈、通常不需改產品程式)。能併入相關行為任務的設定或重構,仍請併入,不要獨立成任務。描述寫了「不要求紅燈」卻沒把 `tdd` 設成 `false`,計畫不會通過。標 `false` 時要在 .flow/plan.md 該 task 的節裡寫明理由,以及這個任務要怎麼驗收(例如「`npm run build` 通過」或「新增的測試在現有實作下通過」)。
56
+ - 不要把「只跑檢查、不改任何會進 git 的檔案」拆成任務。建置、打包預算、lint、型別檢查已由專案的驗證指令負責;若某條驗收要靠這些檢查證明,掛在真正改檔案的那個任務上。這種工作沒有 commit 時會被略過,不會重試。
57
+ - 必須由人眼確認、程式無法用檔案或指令判定的事項,不要放進實作佇列。在該任務加上 `"kind": "confirm"`(沒寫視為要實作)。計畫定案後程式把它們寫進 .flow/confirmations.json 給使用者看,實作不會做到它們,也不會因此停下。
56
58
  - 專案沒有測試框架時,所有任務都會略過 TDD(程式會強制),此時仍請照實標記 `tdd`,並在 `description` 寫清楚驗收方式。
57
59
  - 每個任務先檢查預計修改的程式碼,再依「影響範圍、技術不確定性、失敗後果」三個面向判定 `complexity`,取其中最高的等級;不要只憑檔案數、程式碼行數或驗收條件數判定。
58
60
  - `low`:沿用現有做法,變更侷限在單一行為或模組,失敗容易由局部測試發現且不影響既有資料或對外契約。