agentflowctl 0.14.0 → 0.15.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,5 +1,8 @@
1
1
  # agentflowctl
2
2
 
3
+ [![npm version](https://img.shields.io/npm/v/agentflowctl.svg)](https://www.npmjs.com/package/agentflowctl)
4
+ [![GitHub release](https://img.shields.io/github/v/release/gogogohuang/agentflowctl)](https://github.com/gogogohuang/agentflowctl/releases)
5
+
3
6
  讓 Claude Code、Codex、Gemini CLI 等 agent 在同一個專案裡分工:整理需求、規劃、寫測試與程式、交叉審查,最後建立 PR。agentflowctl 負責推進流程,並用檔案、測試和檢查結果決定能否進到下一步。
4
7
 
5
8
  目前內建支援三種 LLM CLI:**Claude Code、Codex、Gemini CLI**。未來可擴充自定義 LLM adapter,讓其他模型供應商加入流程。目前若要串接其他 CLI,可使用 `command` adapter,自行提供執行命令;需要模型驗證時,也須提供探測命令。
@@ -31,13 +34,13 @@ npx agentflowctl run --req-file ./requirement.md
31
34
  ## 執行時會發生什麼
32
35
 
33
36
  1. agent 整理需求與驗收條件,接著寫計畫,交給其他 agent 審查。
34
- 2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent 實作到測試通過;每個任務都會經過審查與驗證。
37
+ 2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent 實作到測試通過;每個任務都會經過審查與驗證。計畫 agent 會依改動內容在任務標記 `tdd`:建置流程、設定、文件、型別、純重構這類不適合先寫失敗測試的任務,會略過紅綠燈直接實作,改由任務審查與驗證把關。專案沒有測試框架時,所有任務都略過紅綠燈,也不跑 `test` 檢查。
35
38
  3. 全部任務完成後,再執行專案檢查與整體程式碼審查。未通過的項目會交回修正。
36
39
  4. 有 `origin` 時會推送分支;若 `gh` 可用,會嘗試建立 PR。沒有 `origin` 時,完成的分支留在本機。
37
40
 
38
41
  流程預設會自動往下走。想在計畫通過審查後親自確認,可加 `--manual-plan`;確認後執行 `agentflowctl approve <id>`。
39
42
 
40
- agentflowctl 會依專案的 `packageManager`、lockfile 與 `package.json` scripts 選擇安裝、測試及檢查指令。第一次執行時,請留意終端機印出的偵測結果;需要調整可在 `flow.config.json` 指定 `install`、`test` 或 `checks`。
43
+ agentflowctl 會依專案的 `packageManager`、lockfile 與 `package.json` scripts 選擇安裝、測試及檢查指令。第一次執行時,請留意終端機印出的偵測結果;需要調整可在 `flow.config.json` 指定 `install`、`test` 或 `checks`。`package.json` 的依賴或 `test` script 看不出測試框架(且沒有手動設定 `test`)時,終端機會提示「未偵測到測試框架」,並略過紅綠燈;要改回來,在 `flow.config.json` 設定 `test`。
41
44
 
42
45
  ## 查看進度
43
46
 
@@ -59,7 +62,7 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
59
62
 
60
63
  `insights` 把所有 run 分成獨立區塊彙總:最終狀態、失敗原因(重試達上限、仲裁停止、agent 次數用完等)、用量(輸入/輸出/cache、強度占比、階段/agent,各 run 用量最高的任務)、各模型與步驟、步驟執行與失敗(與 `stats` 相同取 log 的檔頭檔尾,跨 run 不計總經過時間)、關卡重試原因。任務關卡與任務步驟不分 task id 合併計算。最後列出最多五則建議,規則由程式套門檻,不是再請 agent 分析。合計 token 不是主指標。覆蓋不足時會先警告占比可能失真。各分組是同一批呼叫的不同切片,不要跨組相加。舊 run 沒有重試或失敗原因紀錄,不會回填。單一 run 的全量明細仍用 `status <id>` 與 `stats <id>`。
61
64
 
62
- 執行紀錄在 `.agentflowctl/runs/<id>/`,工作分支在 `.agentflowctl/worktrees/<id>/`。不再需要某次 run 時,可用 `agentflowctl clean <id>` 清除 worktree 與紀錄;`flow/<id>` 分支會保留。
65
+ 執行紀錄在 `.agentflowctl/runs/<id>/`,工作分支在 `.agentflowctl/worktrees/<id>/`。不再需要某次 run 時,可用 `agentflowctl clean <id>` 清除 worktree 與紀錄;`agentflowctl clean --all` 一次清除所有 done、failed 的 run,以及沒有紀錄的 worktree(進行中、暫停、等待核准的不動)。`flow/<id>` 分支會保留。
63
66
 
64
67
  ## 執行停下來時怎麼做
65
68
 
@@ -69,8 +72,8 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
69
72
  | --- | --- |
70
73
  | 按 Ctrl-C,或終端機意外關閉 | 執行 `agentflowctl resume <id>`;沒有結束紀錄的步驟會重跑 |
71
74
  | `awaiting_approval`:計畫等你確認 | 閱讀 `.agentflowctl/worktrees/<id>/.flow/plan.md`,確認後執行 `agentflowctl approve <id>` |
72
- | `paused`:agent 額度用完 | 等額度恢復後執行 `agentflowctl resume <id>`;審查步驟不會換 agent 代審 |
73
- | `paused`:仲裁沒有產生有效裁決 | 依 `status` 的原因查看 log;若有 `.flow/plan-arbiter.json`,也檢查其內容,處理後執行 `agentflowctl resume <id>`。resume 會直接回到仲裁,不重跑計畫審查;暫停期間若改了計畫檔,或在 `flow.config.json` 把 `planArbiter` 關掉,就改成重新審查 |
75
+ | `paused`:agent 額度用完 | 等額度恢復後執行 `agentflowctl resume <id>`;審查步驟不會換 agent 代審。在仲裁途中暫停時,resume 直接回到仲裁,不重跑計畫審查;暫停期間若改了計畫檔,或在 `flow.config.json` 把 `planArbiter` 關掉,就改成重新審查 |
76
+ | `failed`:仲裁連續沒有產生有效裁決 | 仲裁者沒寫出 `.flow/plan-arbiter.json`、格式錯誤或交接無效時不會暫停,會把原因寫進 `.flow/feedback.md` 並自動重跑仲裁(不重跑計畫審查);無效的檔案移到 `.agentflowctl/runs/<id>/reviews/plan-arbiter-<輪>-<agent>-invalid.json`。連續達重試上限才失敗,查看 log 後執行 `agentflowctl resume <id>` 會再回到仲裁 |
74
77
  | `failed`:測試、檢查、審查或 agent 執行失敗 | 依 `status` 提示查看失敗的 log,處理原因後執行 `agentflowctl resume <id>`;失敗階段會重試 |
75
78
  | `failed`:已達 agent 執行次數上限 | 用 `agentflowctl resume <id> --max-agent-runs 100` 調高上限後接續,數字須大於已執行次數 |
76
79
 
@@ -97,6 +100,7 @@ agentflowctl resume f-xxxx
97
100
  | `run --cycle <名單>` | 指定這次參與的 agent,例如 `--cycle claude,codex`;優先於設定檔的 `cycle` |
98
101
  | `run --model-mode balanced\|adaptive` | 只覆蓋這次 run 的模型模式;`resume` 沿用建立時的模式 |
99
102
  | `run --base <分支>` | 指定起始分支;未設定時使用目前分支 |
103
+ | `run --max-attempts <次數>` | 覆蓋這次的重試上限(至少 3;預設取 `AGENTFLOWCTL_MAX_ATTEMPTS`,未設定為 5);失敗後可用 `resume <id> --max-attempts <次數>` 調高 |
100
104
  | `run --max-agent-runs <次數>` | 覆蓋這次的 `maxAgentRuns`;上限不夠時可用 `resume <id> --max-agent-runs <次數>` 調高 |
101
105
  | `-v` / `--verbose` | 執行時顯示 agent 文字、工具呼叫與專案指令,適用於 `run`、`resume`、`approve` |
102
106
 
@@ -105,6 +109,7 @@ agentflowctl resume f-xxxx
105
109
  ```bash
106
110
  agentflowctl run --req-file ./requirement.md --cycle claude,codex --max-agent-runs 80 --manual-plan
107
111
  agentflowctl resume f-xxxx --max-agent-runs 100
112
+ agentflowctl resume f-xxxx --max-attempts 8
108
113
  ```
109
114
 
110
115
  ### Agent 設定
@@ -189,7 +194,7 @@ Codex 另有幾點差異:
189
194
  | `tddSplit` | `true` | 有多位 agent 時,`true` 會把同一任務的測試與實作分給不同 agent |
190
195
  | `reviewQuorum` | `1` | 任務與最終程式碼審查需要幾位不同審查者核准 |
191
196
  | `planReviewQuorum` | `1` | 計畫需要幾位不同審查者核准 |
192
- | `planArbiter` | `true` | 計畫審查僵持時是否啟用仲裁 |
197
+ | `planArbiter` | `true` | 計畫審查僵持,或修訂一次後仍被要求修改時是否啟用仲裁 |
193
198
  | `planReviewLayers` | `{ "enabled": true, "minTasks": 7, "maxGroups": 5, "tasksPerGroup": 3 }` | 任務夠多時把計畫審查拆成索引與任務群;說明見表格下方 |
194
199
  | `tieBreak` | `"proceed"` | 兩位仲裁者意見分歧時,`"proceed"` 繼續、`"stop"` 停止 |
195
200
  | `maxAgentRuns` | `60` | 一次 run 最多執行幾次 agent;可用指令選項覆蓋 |
@@ -207,11 +212,23 @@ Codex 另有幾點差異:
207
212
 
208
213
  | 變數 | 預設 | 設定方式與用途 |
209
214
  | --- | --- | --- |
210
- | `AGENTFLOWCTL_MAX_ATTEMPTS` | `3` | 同一關連續失敗幾次後停止;例如 `AGENTFLOWCTL_MAX_ATTEMPTS=10 agentflowctl run --req "..."` |
215
+ | `AGENTFLOWCTL_MAX_ATTEMPTS` | `5` | 同一關連續失敗幾次後停止,至少 3(設得更小以 3 計);例如 `AGENTFLOWCTL_MAX_ATTEMPTS=10 agentflowctl run --req "..."` |
211
216
  | `AGENTFLOWCTL_VERBOSE` | 未開啟 | 設為 `1` 顯示詳細輸出,效果同 `-v` |
212
- | `AGENTFLOWCTL_MAX_TURNS` | `80` | 目前程式會讀取此值,但尚未用它限制 agent 執行 |
217
+ | `AGENTFLOWCTL_MAX_TURNS` | `200` | 目前程式會讀取此值,但尚未用它限制 agent 執行 |
218
+
219
+ 環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限(至少 3),單一 run 可用 `run`/`resume` 的 `--max-attempts` 覆蓋;計畫審查何時交付仲裁與它無關:意見沒有變化,或第 2 輪(修訂過一次)仍被要求修改時就交付,兩家 agent 時自動進入雙盲交叉仲裁,有第三方時由第三方單獨仲裁;`maxAgentRuns` 則是整次 run 的 agent 執行次數上限。修正成功、或計畫審查與程式碼審查整組完成一輪有效審查後,該關的失敗次數會歸零,所以上限只計算連續失敗。分層計畫審查時,同一輪裡只要有一次審查呼叫真的執行成功,計畫審查的失敗次數也會歸零;所以索引與各群輪流各失敗一次、每次重跑都有進展時,不會因累計達上限而失敗。
220
+
221
+ ## 發版(維護者)
222
+
223
+ 在 clone 下來的 repo 裡用 `pnpm release <patch|minor|major|x.y.z> [--dry-run]` 發版,只能在 `main` 執行:
224
+
225
+ ```bash
226
+ pnpm release patch --dry-run # 只檢查與驗證,不建立 Release
227
+ pnpm release minor # 確認後建立 v0.x+1.0 的 GitHub Release
228
+ pnpm release 1.0.0 # 指定版本,必須大於目前最新的 tag
229
+ ```
213
230
 
214
- 環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限;`maxAgentRuns` 則是整次 run 的 agent 執行次數上限。修正成功、或計畫審查與程式碼審查整組完成一輪有效審查後,該關的失敗次數會歸零,所以上限只計算連續失敗。分層計畫審查時,同一輪裡只要有一次審查呼叫真的執行成功,計畫審查的失敗次數也會歸零;所以索引與各群輪流各失敗一次、每次重跑都有進展時,不會因累計達上限而失敗。
231
+ 腳本會先檢查目前在 `main`、工作區乾淨、與 `origin/main` 同步、`gh` 已登入、新 tag 不存在,再跑 typecheck、test、build,列出自上個 tag 以來的 commit 並等你輸入 `y` 確認,然後用 `gh release create --generate-notes` 建立 Release。npm 由 Release 觸發的 `npm-publish.yml` 發布;版本號取自 tag,腳本不會改 `package.json`。
215
232
 
216
233
  ## 更多文件
217
234
 
package/dist/cli.js CHANGED
@@ -4,7 +4,7 @@ import { readFileSync } from "node:fs";
4
4
  import { join } from "node:path";
5
5
  import { stdin, stdout } from "node:process";
6
6
  import { createInterface } from "node:readline/promises";
7
- import { config } from "./config.js";
7
+ import { MIN_ATTEMPTS, config } from "./config.js";
8
8
  import { advance, loadRepoConfig } from "./engine.js";
9
9
  import { probeAgent, resolveAgent, runCommand } from "./runner.js";
10
10
  import { addWorktree, git } from "./git.js";
@@ -35,6 +35,12 @@ function printUsage(title, rows) {
35
35
  console.log(` ${key}: ${value};${c.runs} 次(未回報 ${c.unreportedRuns}、舊紀錄不明 ${c.legacyRuns})`);
36
36
  }
37
37
  }
38
+ function positiveInt(text, flag, min = 1) {
39
+ const n = Number(text);
40
+ if (!Number.isInteger(n) || n < min)
41
+ throw new Error(`${flag} 必須是不小於 ${min} 的整數:${text}`);
42
+ return n;
43
+ }
38
44
  function mustGetRun(id) {
39
45
  const run = getRun(id);
40
46
  if (!run)
@@ -121,6 +127,7 @@ program
121
127
  .option("--req-file <file>", "從檔案讀取需求")
122
128
  .option("--base <branch>", "基底分支(預設為目前的分支)")
123
129
  .option("--max-agent-runs <n>", "單一 run 最多執行幾次 agent(預設取 flow.config.json 的 maxAgentRuns)")
130
+ .option("--max-attempts <n>", "同一關連續失敗幾次後停止(預設取 AGENTFLOWCTL_MAX_ATTEMPTS,未設定為 5)")
124
131
  .option("--manual-plan", "計畫通過 AI 審查後,仍停下來等你確認", false)
125
132
  .option("--cycle <agents>", "參與的 agent,例如 claude,codex,gemini(順序不影響分工)")
126
133
  .option("--model-mode <mode>", "這次 run 的模型模式:balanced 或 adaptive")
@@ -154,6 +161,7 @@ program
154
161
  stage: "spec",
155
162
  autopilot: !opts.manualPlan,
156
163
  maxAgentRuns: opts.maxAgentRuns ? Number(opts.maxAgentRuns) : cfg.maxAgentRuns,
164
+ maxAttempts: opts.maxAttempts ? positiveInt(opts.maxAttempts, "--max-attempts", MIN_ATTEMPTS) : undefined,
157
165
  cycle,
158
166
  attempts: {},
159
167
  modelMode,
@@ -183,12 +191,15 @@ program
183
191
  .command("resume <id>")
184
192
  .description("從暫停、中斷或失敗的階段接續")
185
193
  .option("--max-agent-runs <n>", "調整 agent 執行次數上限")
194
+ .option("--max-attempts <n>", "調整同一關連續失敗的上限")
186
195
  .action(async (id, opts) => {
187
196
  let run = mustGetRun(id);
188
197
  if (run.modelMode === "adaptive")
189
198
  validateAdaptiveConfig(loadRepoConfig(), run.cycle);
190
199
  if (opts.maxAgentRuns)
191
200
  run = { ...run, maxAgentRuns: Number(opts.maxAgentRuns) };
201
+ if (opts.maxAttempts)
202
+ run = { ...run, maxAttempts: positiveInt(opts.maxAttempts, "--max-attempts", MIN_ATTEMPTS) };
192
203
  if (run.stage === "paused") {
193
204
  run = { ...run, stage: run.pausedStage ?? "spec", pausedStage: undefined, pauseReason: undefined };
194
205
  }
package/dist/config.js CHANGED
@@ -1,8 +1,10 @@
1
+ /** 重試上限的下限:低於這個值時,修正與審查來不及往返一輪 */
2
+ export const MIN_ATTEMPTS = 3;
1
3
  export const config = {
2
4
  /** 單一 Agent 執行最多幾輪工具迴圈,避免卡在迴圈裡 */
3
- maxTurns: Number(process.env.AGENTFLOWCTL_MAX_TURNS ?? 80),
4
- /** 同一個關卡連續失敗幾次後停止 */
5
- maxAttempts: Number(process.env.AGENTFLOWCTL_MAX_ATTEMPTS ?? 3),
5
+ maxTurns: Number(process.env.AGENTFLOWCTL_MAX_TURNS ?? 200),
6
+ /** 同一個關卡連續失敗幾次後停止;環境變數低於 3 時以 3 計 */
7
+ maxAttempts: Math.max(MIN_ATTEMPTS, Number(process.env.AGENTFLOWCTL_MAX_ATTEMPTS ?? 5) || 5),
6
8
  /** 終端機是否印出 agent 的文字、工具呼叫與專案指令;預設安靜,-v 或 AGENTFLOWCTL_VERBOSE=1 開啟 */
7
9
  verbose: process.env.AGENTFLOWCTL_VERBOSE === "1",
8
10
  };
package/dist/detect.js CHANGED
@@ -24,6 +24,16 @@ const CHECK_SCRIPTS = {
24
24
  test: ["test"],
25
25
  build: ["build"],
26
26
  };
27
+ /** 出現在依賴裡就代表專案有測試框架 */
28
+ const TEST_FRAMEWORKS = ["vitest", "jest", "mocha", "ava", "jasmine", "tap", "uvu", "@playwright/test", "cypress"];
29
+ function hasTestFramework(pkg) {
30
+ const deps = { ...pkg.dependencies, ...pkg.devDependencies };
31
+ if (TEST_FRAMEWORKS.some((name) => name in deps))
32
+ return true;
33
+ const script = pkg.scripts?.test;
34
+ // npm init 產生的佔位 script 不算
35
+ return typeof script === "string" && !/no test specified/i.test(script);
36
+ }
27
37
  function readPackageJson(root) {
28
38
  try {
29
39
  const pkg = JSON.parse(readFileSync(join(root, "package.json"), "utf8"));
@@ -56,24 +66,31 @@ export function detectProjectDefaults(root) {
56
66
  const script = CHECK_SCRIPTS[name]?.find((s) => typeof scripts[s] === "string");
57
67
  return { name, cmd: script ? `${manager} run ${script}` : withExec(cmd, manager) };
58
68
  });
59
- return { manager, source, install: INSTALL[manager], test: withExec(defaults.test, manager), checks };
69
+ return { manager, source, install: INSTALL[manager], test: withExec(defaults.test, manager), checks, testFramework: hasTestFramework(pkg) };
60
70
  }
71
+ /** 手動設定 test 指令就當作有測試框架 */
72
+ export const usesTestFramework = (raw, detected) => detected.testFramework || (!!raw && typeof raw === "object" && "test" in raw);
61
73
  /** 補上原始設定裡沒寫的 install、test、checks;不是物件就原樣回傳,交給 schema 報錯 */
62
74
  export function withProjectDefaults(raw, detected) {
63
75
  if (!raw || typeof raw !== "object" || Array.isArray(raw))
64
76
  return raw;
65
- const { install, test, checks } = detected;
77
+ const { install, test } = detected;
78
+ // 沒有測試框架也沒手動設定 test 時,預設的檢查不含 test
79
+ const checks = usesTestFramework(raw, detected) ? detected.checks : detected.checks.filter((c) => c.name !== "test");
66
80
  return { install, test, checks, ...raw };
67
81
  }
68
82
  /** run 開始時印出的說明:只列出這次用了偵測結果的欄位 */
69
83
  export function describeDetected(raw, detected) {
70
84
  const lines = [];
85
+ const framework = usesTestFramework(raw, detected);
71
86
  if (!("install" in raw))
72
87
  lines.push(` install:${detected.install}`);
73
- if (!("test" in raw))
88
+ if (!("test" in raw) && framework)
74
89
  lines.push(` test:${detected.test}`);
75
90
  if (!("checks" in raw))
76
- lines.push(...detected.checks.map((c) => ` checks.${c.name}:${c.cmd}`));
91
+ lines.push(...detected.checks.filter((c) => framework || c.name !== "test").map((c) => ` checks.${c.name}:${c.cmd}`));
92
+ if (!framework)
93
+ lines.push("ℹ️ 未偵測到測試框架,所有任務略過紅綠燈,也不跑 test 檢查(在 flow.config.json 設定 test 可改回來)");
77
94
  if (!lines.length)
78
95
  return [];
79
96
  const why = detected.source === "預設" ? "預設" : `依 ${detected.source}`;
package/dist/engine.js CHANGED
@@ -3,7 +3,7 @@ import { join } from "node:path";
3
3
  import { z } from "zod";
4
4
  import { config } from "./config.js";
5
5
  import { arbitrationDecision } from "./arbitration.js";
6
- import { detectProjectDefaults, withProjectDefaults } from "./detect.js";
6
+ import { detectProjectDefaults, usesTestFramework, withProjectDefaults } from "./detect.js";
7
7
  import { escapeXml, opinion, reviewIssue } from "./feedback.js";
8
8
  import { changedFiles, commitAll, discardChanges, git, headCommit, resetTo } from "./git.js";
9
9
  import { acceptHandoff, openActions, prepareHandoff, previewHandoff, readHandoff, recoverHandoff, reviewHandoffGate, validateHandoffResponse } from "./handoff.js";
@@ -125,18 +125,24 @@ function flowText(run, name) {
125
125
  function readFeedback(run) {
126
126
  return flowText(run, "feedback.md");
127
127
  }
128
+ /** 計畫審查第幾輪仍有人要求修改時交付仲裁;固定值,不受重試上限影響 */
129
+ const PLAN_ARBITRATION_ROUND = 2;
130
+ /** 這個 run 的重試上限:`--max-attempts` 存在 run 裡,沒設定才用環境變數 */
131
+ function attemptLimit(run) {
132
+ return run.maxAttempts ?? config.maxAttempts;
133
+ }
128
134
  /** 關卡未通過:寫入 feedback.md 給下一次嘗試參考,並把原因分類記進 retries.jsonl;超過上限就讓整個 run 失敗 */
129
135
  function retry(run, key, reason, backTo, category) {
130
136
  const n = (run.attempts[key] ?? 0) + 1;
131
137
  const attempts = { ...run.attempts, [key]: n };
132
- const final = n >= config.maxAttempts;
138
+ const final = n >= attemptLimit(run);
133
139
  mkdirSync(flowDir(run.id), { recursive: true });
134
140
  writeFileSync(flowFile(run, "feedback.md"), `# 前次嘗試未通過(第 ${n} 次)\n\n${reason}\n`);
135
141
  addRetry(run.id, { key, backTo, category, attempt: n, final }, run.updatedAt);
136
142
  if (final) {
137
143
  return { ...run, attempts, stage: "failed", failedStage: backTo, failureCategory: "retry_limit", failureReason: `${key} 連續失敗 ${n} 次:${tail(reason, 500)}` };
138
144
  }
139
- info(run, `⚠️ ${key} 未通過,重試(${n}/${config.maxAttempts})`);
145
+ info(run, `⚠️ ${key} 未通過,重試(${n}/${attemptLimit(run)})`);
140
146
  return { ...run, attempts, stage: backTo };
141
147
  }
142
148
  function succeed(run, key, next) {
@@ -158,6 +164,17 @@ export function loadRepoConfig() {
158
164
  throw new Error(r.error);
159
165
  return r.data;
160
166
  }
167
+ /** 專案有測試框架(偵測到或手動設定 test) */
168
+ function hasTestFramework() {
169
+ const root = projectRoot();
170
+ const p = join(root, "flow.config.json");
171
+ const r = existsSync(p) ? readJsonFile(p, z.unknown()) : undefined;
172
+ return usesTestFramework(r?.ok ? r.data : {}, detectProjectDefaults(root));
173
+ }
174
+ /** 這個任務要不要走紅綠燈:沒有測試框架一律不走,其餘依 planner 標記(沒標視為要走) */
175
+ function taskUsesTdd(task, framework) {
176
+ return framework && task.tdd !== false;
177
+ }
161
178
  function loadOrderedTasks(run) {
162
179
  const r = readJsonFile(flowFile(run, "tasks.ordered.json"), TaskList);
163
180
  if (!r.ok)
@@ -236,9 +253,12 @@ function acceptPlan(run, ordered) {
236
253
  function announceTasks(run, ordered) {
237
254
  const cfg = loadRepoConfig();
238
255
  info(run, `📋 共 ${ordered.length} 個任務:${ordered.map((t) => t.id).join(" → ")}`);
256
+ const framework = hasTestFramework();
239
257
  ordered.forEach((t, i) => {
240
258
  const a = taskAgents(run.cycle, i, cfg.tddSplit, run.id);
241
- info(run, ` ${t.id} 測試:${a.tests} 實作:${a.code} 審查:${a.review}`);
259
+ info(run, taskUsesTdd(t, framework)
260
+ ? ` ${t.id} 測試:${a.tests} 實作:${a.code} 審查:${a.review}`
261
+ : ` ${t.id} 略過 TDD 實作:${a.code} 審查:${a.review}`);
242
262
  });
243
263
  }
244
264
  /** 計畫定案後:預設直接開始實作;--manual-plan 時才停下來等人 */
@@ -346,7 +366,8 @@ async function concludePlanReview(run, cfg, round, calls) {
346
366
  const fingerprint = reviewFingerprint(issueLines);
347
367
  const stalled = existsSync(lastPath) && readFileSync(lastPath, "utf8") === fingerprint;
348
368
  writeFileSync(lastPath, fingerprint);
349
- const exhausted = round >= config.maxAttempts;
369
+ // 修訂過一次仍被要求修改就交付仲裁,不等到重試上限
370
+ const exhausted = round >= PLAN_ARBITRATION_ROUND;
350
371
  if ((stalled || exhausted) && cfg.planArbiter) {
351
372
  // 雙盲:帶有審查者名稱的 feedback.md 不留在 worktree,完整報告另存到 worktree 外
352
373
  rmSync(flowFile(run, "feedback.md"), { force: true });
@@ -587,8 +608,14 @@ async function arbitratePlan(run) {
587
608
  if (!result?.ok || handoffError) {
588
609
  const outputError = !r.ok ? `Agent 執行失敗:${r.summary}` : !result ? "未產生有效裁決" : result.ok ? "未產生有效裁決" : result.error;
589
610
  const reason = handoffError ?? outputError;
590
- info(run, ` ⏸️ ${arbiter} 未產生有效裁決:${reason}`);
591
- return { ...run, stage: "paused", pausedStage: "plan_review", pauseReason: `${arbiter} 仲裁未產生有效裁決:${reason}` };
611
+ const category = handoffError ? "handoff_invalid" : !r.ok ? "agent_error" : "format_invalid";
612
+ info(run, ` ✗ ${arbiter} 未產生有效裁決:${reason}`);
613
+ // 無效的裁決移到 worktree 外保存,供事後查看;plan-arbitration.json 還在,重試時直接回到仲裁
614
+ if (existsSync(flowFile(run, "plan-arbiter.json"))) {
615
+ mkdirSync(join(runDir(run.id), "reviews"), { recursive: true });
616
+ renameSync(flowFile(run, "plan-arbiter.json"), join(runDir(run.id), "reviews", `plan-arbiter-${arbitrationRound}-${arbiter}-invalid.json`));
617
+ }
618
+ return retry(run, "plan-arbitration-run", `上次仲裁未產生有效裁決:${reason}`, "plan_review", category);
592
619
  }
593
620
  mkdirSync(join(runDir(run.id), "reviews"), { recursive: true });
594
621
  renameSync(flowFile(run, "plan-arbiter.json"), join(runDir(run.id), "reviews", `plan-arbiter-${arbitrationRound}-${arbiter}.json`));
@@ -600,6 +627,8 @@ async function arbitratePlan(run) {
600
627
  }
601
628
  rmSync(flowFile(run, "dispute.md"), { force: true });
602
629
  rmSync(planArbitrationPath(run.id), { force: true });
630
+ run = { ...run, attempts: { ...run.attempts } };
631
+ delete run.attempts["plan-arbitration-run"];
603
632
  const approvals = verdicts.filter((v) => v.verdict === "approve").length;
604
633
  const unanimous = approvals === verdicts.length;
605
634
  const decision = arbitrationDecision(verdicts.map((v) => v.verdict), cfg.tieBreak);
@@ -655,6 +684,13 @@ async function implementStage(run) {
655
684
  if (run.taskPhase === "fix")
656
685
  return taskFixStep(run, task, progress);
657
686
  const agents = taskAgents(run.cycle, run.taskIndex, cfg.tddSplit, run.id);
687
+ const framework = hasTestFramework();
688
+ const tdd = taskUsesTdd(task, framework);
689
+ // ── 不走 TDD:略過紅燈,實作前的 HEAD 就是這個任務的起點 ──
690
+ if (!tdd && run.taskPhase === "tests") {
691
+ info(run, `⏭️ [${progress}] 略過 TDD(${framework ? "planner 標記不適合先寫測試" : "專案沒有測試框架"})`);
692
+ return { ...run, taskPhase: "code", taskBase: await headCommit(repo), testsCommit: undefined, lastTestsAuthor: undefined };
693
+ }
658
694
  // ── 紅燈:只寫測試,而且測試必須失敗 ──
659
695
  if (run.taskPhase === "tests") {
660
696
  const key = `${task.id}:tests`;
@@ -696,13 +732,16 @@ async function implementStage(run) {
696
732
  }
697
733
  // ── 綠燈:實作到測試通過,而且不可動測試 ──
698
734
  const key = `${task.id}:code`;
699
- const testsCommit = run.testsCommit;
735
+ // 走 TDD 時以測試 commit 為基準;不走 TDD 時以任務起點為基準
736
+ const testsCommit = tdd ? run.testsCommit : run.taskBase;
700
737
  if (!testsCommit)
701
- throw new Error("缺少 testsCommit,狀態不一致");
702
- info(run, `🛠️ [${progress}] 實作(${agents.code},測試由 ${run.lastTestsAuthor ?? agents.tests} 撰寫)`);
738
+ throw new Error(tdd ? "缺少 testsCommit,狀態不一致" : "缺少 taskBase,狀態不一致");
739
+ info(run, tdd ? `🛠️ [${progress}] 實作(${agents.code},測試由 ${run.lastTestsAuthor ?? agents.tests} 撰寫)` : `🛠️ [${progress}] 實作(${agents.code},不走 TDD)`);
703
740
  const redOutput = existsSync(flowFile(run, "red-output.txt")) ? readFileSync(flowFile(run, "red-output.txt"), "utf8") : "";
704
741
  const snap = snapshotPlan(run, LOCKED_FILES);
705
- const outcome = await agentStep(run, agents.code, `${task.id}-code`, renderPrompt("implement-code", { task: taskJson, acceptance: acceptanceJson, testCmd, redOutput: tail(redOutput, 3000) }), { kind: "write", reset: async () => { await resetTo(repo, testsCommit); restorePlan(run, snap); } });
742
+ const outcome = await agentStep(run, agents.code, `${task.id}-code`, tdd
743
+ ? renderPrompt("implement-code", { task: taskJson, acceptance: acceptanceJson, testCmd, redOutput: tail(redOutput, 3000) })
744
+ : renderPrompt("implement-direct", { task: taskJson, acceptance: acceptanceJson, testCmd: framework ? testCmd : "" }), { kind: "write", reset: async () => { await resetTo(repo, testsCommit); restorePlan(run, snap); } });
706
745
  const { r, agent: codeAuthor } = outcome;
707
746
  const tampered = restorePlan(run, snap);
708
747
  if (!r.ok)
@@ -711,13 +750,16 @@ async function implementStage(run) {
711
750
  await resetTo(repo, testsCommit);
712
751
  return retry(run, key, planTamperedMessage(tampered), "implement", "plan_tampered");
713
752
  }
714
- await commitAll(repo, `feat(${task.id}): ${task.title} [${codeAuthor}]`);
715
- const touched = (await changedFiles(repo, testsCommit, await headCommit(repo))).filter((f) => testRe.test(f));
753
+ const codeCommit = await commitAll(repo, `feat(${task.id}): ${task.title} [${codeAuthor}]`);
754
+ if (!tdd && !codeCommit)
755
+ return retry(run, key, "沒有任何檔案變更,這個任務必須完成實作。", "implement", "code_not_written");
756
+ const touched = tdd ? (await changedFiles(repo, testsCommit, await headCommit(repo))).filter((f) => testRe.test(f)) : [];
716
757
  if (touched.length) {
717
758
  await resetTo(repo, testsCommit);
718
759
  return retry(run, key, `實作階段不可修改測試檔,已還原你的變更:${touched.join(", ")}`, "implement", "tests_modified");
719
760
  }
720
- const green = await runCommand(target(run, `${task.id}-green`, CMD_AGENT), testCmd);
761
+ // 沒有測試框架時沒有東西可跑,後面的驗證與審查照常把關
762
+ const green = framework ? await runCommand(target(run, `${task.id}-green`, CMD_AGENT), testCmd) : { ok: true, output: "", seq: 0 };
721
763
  if (!green.ok) {
722
764
  info(run, ` ✗ 測試仍未通過${logHint(run, green.seq)}`);
723
765
  return retry(run, key, `測試仍未通過:\n\n\`\`\`\n${tail(green.output)}\n\`\`\``, "implement", "tests_not_green");
@@ -727,7 +769,7 @@ async function implementStage(run) {
727
769
  await resetTo(repo, testsCommit);
728
770
  return retry(run, key, handoffError, "implement", "handoff_invalid");
729
771
  }
730
- info(run, `🟢 [${progress}] 測試通過`);
772
+ info(run, framework ? `🟢 [${progress}] 測試通過` : `✔️ [${progress}] 實作完成`);
731
773
  return { ...succeed(run, key, "implement"), taskPhase: "review", lastWriter: codeAuthor };
732
774
  }
733
775
  // ── 任務審查:只看這個任務的變更與驗收條件 ──
package/dist/insights.js CHANGED
@@ -9,6 +9,7 @@ export const RETRY_LABEL = {
9
9
  arbitration_revise: "仲裁要求修訂",
10
10
  plan_tampered: "改動已鎖定的計畫檔",
11
11
  tests_not_written: "未寫測試",
12
+ code_not_written: "未實作",
12
13
  tests_not_red: "紅燈測試未失敗",
13
14
  tests_modified: "實作改了測試",
14
15
  tests_not_green: "測試仍未通過",
package/dist/schemas.js CHANGED
@@ -72,6 +72,8 @@ export const TaskItem = z.object({
72
72
  dependsOn: z.array(z.string()).default([]),
73
73
  acceptance: z.array(z.string()).min(1, "每個任務至少要對應一條驗收條件"),
74
74
  complexity: z.enum(["low", "medium", "high"]).optional(),
75
+ /** false=這個任務不適合先寫會失敗的測試(建置流程、設定、文件、純重構等),略過紅燈直接實作;沒寫視為 true */
76
+ tdd: z.boolean().optional(),
75
77
  });
76
78
  /** Agent 在 plan 階段產出的 .flow/tasks.json */
77
79
  export const TaskList = z.array(TaskItem).min(1);
@@ -181,6 +183,8 @@ export const FlowRun = z.object({
181
183
  autopilot: z.boolean(),
182
184
  /** 單一 run 最多執行幾次 agent */
183
185
  maxAgentRuns: z.number().int().positive(),
186
+ /** 這個 run 同一關連續失敗的上限;沒寫就用 AGENTFLOWCTL_MAX_ATTEMPTS(舊 state.json 沒有此欄位) */
187
+ maxAttempts: z.number().int().min(3).optional(),
184
188
  /** 暫停前所在的階段與原因(額度用完時) */
185
189
  pausedStage: Stage.optional(),
186
190
  pauseReason: z.string().optional(),
package/dist/store.js CHANGED
@@ -112,7 +112,7 @@ export function listSubstitutions(id) {
112
112
  }
113
113
  export const RetryCategories = [
114
114
  "agent_error", "missing_artifact", "format_invalid", "handoff_invalid", "open_handoff",
115
- "review_changes", "arbitration_revise", "plan_tampered", "tests_not_written", "tests_not_red", "tests_modified",
115
+ "review_changes", "arbitration_revise", "plan_tampered", "tests_not_written", "code_not_written", "tests_not_red", "tests_modified",
116
116
  "tests_not_green", "tests_deleted", "checks_failed",
117
117
  ];
118
118
  const retryPath = (id) => join(runDir(id), "retries.jsonl");
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "agentflowctl",
3
3
  "license": "MIT",
4
- "version": "0.14.0",
4
+ "version": "0.15.1",
5
5
  "description": "跨廠商 AI 開發 harness:Claude Code、Codex、Gemini 輪流實作、審查、修正",
6
6
  "keywords": [
7
7
  "ai",
@@ -41,7 +41,8 @@
41
41
  "build": "tsc",
42
42
  "typecheck": "tsc --noEmit",
43
43
  "test": "vitest run",
44
- "prepublishOnly": "pnpm run typecheck && pnpm test && pnpm run build"
44
+ "prepublishOnly": "pnpm run typecheck && pnpm test && pnpm run build",
45
+ "release": "bash scripts/release.sh"
45
46
  },
46
47
  "dependencies": {
47
48
  "commander": "^12.1.0",
@@ -0,0 +1,67 @@
1
+ <role>
2
+ 你是任務實作者,負責完成一個**不走 TDD** 的任務。這個任務不適合先寫會失敗的測試(例如建置流程、設定、文件、型別或純重構),或專案沒有測試框架,所以沒有紅燈測試可依循。你要依任務描述與驗收條件,用符合專案風格的最小改動完成它。
3
+ </role>
4
+
5
+ <context>
6
+ 目前的工作目錄就是專案(agentflowctl 為這次任務建立的專用 git worktree)。
7
+ </context>
8
+
9
+ <handoff>
10
+ 先閱讀 .flow/handoff-context.md,處理與本階段有關的待辦事項。完成時寫入 .flow/handoff-response.json;即使沒有事項也必須寫出空陣列:
11
+
12
+ ```json
13
+ { "newIssues": [], "dispositions": [] }
14
+ ```
15
+
16
+ 新增事項格式:{ "kind": "action 或 info", "summary": "具體問題", "evidence": "檔案位置或檢查證據", "targetStage": "plan 或 code" }。
17
+ 處置格式:{ "id": "既有事項 ID", "status": "proposed_resolved、resolved 或 accepted", "reason": "具體處理理由", "evidence": "檔案、commit 或檢查結果" }。只有「待處理事項」(action)可以處置:撰寫者只能用 proposed_resolved 提出修正;審查者可以用 resolved 或 accepted 結案。「參考資訊」(info)只供參考,不要放進 dispositions。重要疑慮必須放在這份檔案,不能只寫在回覆的 <concerns>。
18
+ </handoff>
19
+
20
+ <no_tdd>
21
+ 這個任務略過紅綠燈:沒有事先寫好的失敗測試,完成後由任務審查與專案的驗證指令(typecheck、lint、build 等)把關。
22
+ </no_tdd>
23
+
24
+ <task>
25
+ ```json
26
+ {{task}}
27
+ ```
28
+ </task>
29
+
30
+ <acceptance>
31
+ 這個任務負責的驗收條件:
32
+ ```json
33
+ {{acceptance}}
34
+ ```
35
+ </acceptance>
36
+
37
+ <inputs>
38
+ 先依上面的任務與驗收條件工作。只有資訊不足或互相矛盾時,再閱讀 .flow/spec.md、.flow/plan.md 的相關段落,並在交接中指出問題;不必通讀整份文件。
39
+ </inputs>
40
+
41
+ <steps>
42
+ 1. 若 .flow/feedback.md 存在,先閱讀,並依內容修正。
43
+ 2. 完成任務:最小改動、符合專案既有的程式風格與架構,並確保每一條驗收條件都能用具體的檔案或指令輸出證明。
44
+ 3. 自行執行能證明成果的指令(例如建置、型別檢查、執行腳本)並確認結果。
45
+ 4. 若專案有測試,外部流程會再執行 `{{testCmd}}` 確認既有測試沒有被破壞(空白代表專案沒有測試指令),不需要自行重跑全套測試。
46
+ </steps>
47
+
48
+ <constraints>
49
+ - 必須實際修改檔案;沒有任何變更會被視為失敗。
50
+ - 不要執行 git commit(權限設定已禁止)。
51
+ - **不可修改** .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json、.flow/tasks.ordered.json,修改會被自動還原並視為失敗。若認為規格或驗收條件有誤,請寫進 .flow/handoff-response.json 的 newIssues。
52
+ </constraints>
53
+
54
+ <reply_format>
55
+ 完成後,回覆的最後必須附上以下 XML 中繼資料(只附一次,標籤名稱不可更改):
56
+
57
+ ```xml
58
+ <result>
59
+ <status>done 或 blocked</status>
60
+ <summary>一兩句說明這次做了什麼;blocked 時說明卡在哪裡</summary>
61
+ <files_changed>
62
+ <file>每個新增或修改的檔案路徑各一行</file>
63
+ </files_changed>
64
+ <concerns>對需求、規格、計畫或測試的疑慮;沒有就留空</concerns>
65
+ </result>
66
+ ```
67
+ </reply_format>
@@ -24,6 +24,7 @@
24
24
  <inputs>
25
25
  - .flow/spec.md、.flow/acceptance.json、.flow/plan.md、.flow/tasks.json:目前的規格與計畫(審查意見的處理在 .flow/plan-replies.md;沒有這份檔表示尚未回應)
26
26
  - .flow/dispute.md:尚未被接受的審查意見
27
+ - .flow/feedback.md(可能不存在):上次仲裁輸出沒有通過程式檢查的原因,這次必須避免
27
28
  </inputs>
28
29
 
29
30
  <criteria>
@@ -47,6 +48,10 @@
47
48
  ```
48
49
 
49
50
  dispute.md 裡的每一條意見都要列一筆並說明你的判斷。
51
+ `status` 只能是以下三個值之一,不可自創其他值:
52
+ - `met`:這條意見不成立,或計畫已經處理好
53
+ - `not_met`:意見成立,計畫必須修改
54
+ - `partial`:意見部分成立,計畫仍需補強
50
55
  `verdict` 只能寫 `approve` 或 `changes_requested`。若計畫有會導致錯誤結果、遺漏需求或無法驗收的問題,請寫 `changes_requested`,不要寫 `reject`。
51
56
  </output_format>
52
57
 
@@ -30,7 +30,7 @@
30
30
  <review_focus>
31
31
  1. **需求覆蓋**:規格是否完整涵蓋原始需求?有沒有遺漏、誤解,或加入需求沒要求的範圍?
32
32
  2. **驗收條件**:每一條是否具體、可以用自動化測試驗證,而且只描述一個行為?把多個行為寫在同一條的,要求拆開。有沒有重要的邊界情況或錯誤處理沒被列入?
33
- 3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?
33
+ 3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試(標 `tdd: false` 的任務除外:核對它的改動內容確實不適合先寫失敗測試,例如建置流程、設定、文件、型別、純重構,並有寫明驗收方式;會改變程式行為卻標成 `false` 的,要求改回 `true`)?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?
34
34
  同時依 .flow/plan.md 的逐項理由及實際程式碼,獨立核對每個 task 的 `complexity`:分別看影響範圍、技術不確定性與失敗後果,取最高等級。`low` 須是沿用既有做法、侷限單一行為或模組且失敗可由局部測試發現;`medium` 包括多模組或介面協調、非典型邊界、相容性或狀態遷移風險;`high` 包括跨系統契約、架構或資料模型變更、未知的關鍵技術路徑,或資料遺失、權限、難以回復的風險。不要只憑檔案數、程式碼行數或驗收條件數判定。
35
35
  理由缺漏、與程式碼不符,或高低估會影響選模時,要求修正;在 `note` 指出 task ID、具體證據、建議等級及須修改的 .flow/plan.md/.flow/tasks.json 部分。不要為缺少高價值證據的細微措辭差異要求修改。
36
36
  4. **技術方向**:是否符合專案既有的架構與慣例?有沒有更簡單的做法,或明顯的風險?
package/prompts/plan.md CHANGED
@@ -39,6 +39,7 @@
39
39
  "title": "建立表單驗證 schema",
40
40
  "description": "具體要做什麼、要動哪些檔案、測試要驗證什麼行為",
41
41
  "complexity": "low",
42
+ "tdd": true,
42
43
  "dependsOn": [],
43
44
  "acceptance": ["AC-1"]
44
45
  }
@@ -51,7 +52,8 @@
51
52
  - 每個任務是一個可獨立測試的垂直切片,小到一次 TDD 循環就能完成;只動少數幾個檔案,測試只驗證一兩個行為。
52
53
  - `title` 用一句話說出這件事;需要用「並且」「以及」串起來的,就是兩個任務。
53
54
  - `description` 寫清楚要動哪些檔案(寫含目錄的路徑,例如 `src/form.ts`,不要只寫檔名)、測試要驗證哪個行為,以及這個任務不做什麼。計畫審查會依這些路徑把任務分群。
54
- - 每個任務都必須能寫出「在實作前會失敗」的測試;純設定或重構類工作請併入相關任務。
55
+ - 依改動內容標記 `tdd`:會改變程式行為、能寫出「在實作前會失敗」的測試的任務標 `true`(預設);改動內容不適合先寫失敗測試的任務標 `false`,這類任務會略過紅燈直接實作,改由任務審查與驗證指令把關。適合標 `false` 的例子:建置流程與打包設定(build、CI、bundler、tsconfig)、依賴與版本設定、文件與 prompt 文字、樣式與靜態資源、型別宣告、不改變行為的重構與搬移檔案。能併入相關行為任務的設定或重構,仍請併入,不要獨立成任務。標 `false` 時要在 .flow/plan.md 該 task 的節裡寫明理由,以及這個任務要怎麼驗收(例如「`npm run build` 通過」)。
56
+ - 專案沒有測試框架時,所有任務都會略過 TDD(程式會強制),此時仍請照實標記 `tdd`,並在 `description` 寫清楚驗收方式。
55
57
  - 每個任務先檢查預計修改的程式碼,再依「影響範圍、技術不確定性、失敗後果」三個面向判定 `complexity`,取其中最高的等級;不要只憑檔案數、程式碼行數或驗收條件數判定。
56
58
  - `low`:沿用現有做法,變更侷限在單一行為或模組,失敗容易由局部測試發現且不影響既有資料或對外契約。
57
59
  - `medium`:需要協調多個模組或既有介面、處理非典型邊界,或有相容性與狀態遷移風險,但可依已知做法實作與驗證。