agentflowctl 0.13.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,6 +2,8 @@
2
2
 
3
3
  讓 Claude Code、Codex、Gemini CLI 等 agent 在同一個專案裡分工:整理需求、規劃、寫測試與程式、交叉審查,最後建立 PR。agentflowctl 負責推進流程,並用檔案、測試和檢查結果決定能否進到下一步。
4
4
 
5
+ 目前內建支援三種 LLM CLI:**Claude Code、Codex、Gemini CLI**。未來可擴充自定義 LLM adapter,讓其他模型供應商加入流程。目前若要串接其他 CLI,可使用 `command` adapter,自行提供執行命令;需要模型驗證時,也須提供探測命令。
6
+
5
7
  每次執行都會建立獨立的 git worktree 與 `flow/<id>` 分支,不會直接修改你目前的工作目錄。兩個 agent 就能運作;若只有一個,也能執行,但無法做到跨 agent 審查。
6
8
 
7
9
  ## 開始使用
@@ -47,6 +49,7 @@ agentflowctl status f-xxxx # 看進度、結果與下一步
47
49
  agentflowctl logs f-xxxx # 列出各步驟的 log
48
50
  agentflowctl logs f-xxxx --latest # 看最新一份 log
49
51
  agentflowctl stats f-xxxx # 各步驟耗時、執行與失敗次數
52
+ agentflowctl insights # 這個專案所有 run 的結果、失敗原因、用量、步驟失敗、重試原因與改善建議
50
53
  agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
51
54
  ```
52
55
 
@@ -54,6 +57,8 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
54
57
 
55
58
  `stats` 依 log 的開始與結束時間統計每個步驟的執行次數、失敗次數、總耗時與最長一次,並分開列出 agent 與專案指令(install、測試、checks)各占多少時間,最耗時的步驟排在最前面。沒有結束紀錄的 log 列為未完成,不計入耗時;總經過時間包含暫停與等待核准。
56
59
 
60
+ `insights` 把所有 run 分成獨立區塊彙總:最終狀態、失敗原因(重試達上限、仲裁停止、agent 次數用完等)、用量(輸入/輸出/cache、強度占比、階段/agent,各 run 用量最高的任務)、各模型與步驟、步驟執行與失敗(與 `stats` 相同取 log 的檔頭檔尾,跨 run 不計總經過時間)、關卡重試原因。任務關卡與任務步驟不分 task id 合併計算。最後列出最多五則建議,規則由程式套門檻,不是再請 agent 分析。合計 token 不是主指標。覆蓋不足時會先警告占比可能失真。各分組是同一批呼叫的不同切片,不要跨組相加。舊 run 沒有重試或失敗原因紀錄,不會回填。單一 run 的全量明細仍用 `status <id>` 與 `stats <id>`。
61
+
57
62
  執行紀錄在 `.agentflowctl/runs/<id>/`,工作分支在 `.agentflowctl/worktrees/<id>/`。不再需要某次 run 時,可用 `agentflowctl clean <id>` 清除 worktree 與紀錄;`flow/<id>` 分支會保留。
58
63
 
59
64
  ## 執行停下來時怎麼做
@@ -65,7 +70,7 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
65
70
  | 按 Ctrl-C,或終端機意外關閉 | 執行 `agentflowctl resume <id>`;沒有結束紀錄的步驟會重跑 |
66
71
  | `awaiting_approval`:計畫等你確認 | 閱讀 `.agentflowctl/worktrees/<id>/.flow/plan.md`,確認後執行 `agentflowctl approve <id>` |
67
72
  | `paused`:agent 額度用完 | 等額度恢復後執行 `agentflowctl resume <id>`;審查步驟不會換 agent 代審 |
68
- | `paused`:仲裁沒有產生有效裁決 | 依 `status` 的原因查看 log;若有 `.flow/plan-arbiter.json`,也檢查其內容,處理後執行 `agentflowctl resume <id>` |
73
+ | `paused`:仲裁沒有產生有效裁決 | 依 `status` 的原因查看 log;若有 `.flow/plan-arbiter.json`,也檢查其內容,處理後執行 `agentflowctl resume <id>`。resume 會直接回到仲裁,不重跑計畫審查;暫停期間若改了計畫檔,或在 `flow.config.json` 把 `planArbiter` 關掉,就改成重新審查 |
69
74
  | `failed`:測試、檢查、審查或 agent 執行失敗 | 依 `status` 提示查看失敗的 log,處理原因後執行 `agentflowctl resume <id>`;失敗階段會重試 |
70
75
  | `failed`:已達 agent 執行次數上限 | 用 `agentflowctl resume <id> --max-agent-runs 100` 調高上限後接續,數字須大於已執行次數 |
71
76
 
@@ -130,13 +135,33 @@ agentflowctl model mode adaptive
130
135
  agentflowctl run --req-file ./requirement.md
131
136
  ```
132
137
 
133
- 把 `MODEL_NAME` 換成該 CLI 目前可呼叫的別名或完整 ID。`model add` 會用目前登入的帳號送出短請求,可能耗用少量 token;成功才寫入設定。需要重驗時執行 `model check`。用 `model set claude MODEL_NAME --strength medium` 改強度、`model remove claude MODEL_NAME` 移除模型,或用 `model stage taskReview high` 調整階段最低強度;`model stage` 不帶強度時列出各階段實際生效的強度。終端機每次呼叫會顯示送給 CLI 的模型名稱,Claude Code 與 Gemini CLI 回報的實際模型不同時也會顯示;`status <id>` 會按階段、任務、模型與步驟顯示用量。`run --model-mode balanced` 可暫時回到原設定。
138
+ 把 `MODEL_NAME` 換成該 CLI 目前可呼叫的別名或完整 ID。`model add` 會用目前登入的帳號送出短請求,可能耗用少量 token;成功才寫入設定。需要重驗時執行 `model check`。用 `model set claude MODEL_NAME --strength medium` 改強度、`model remove claude MODEL_NAME` 移除模型,或用 `model stage taskReview high` 調整階段最低強度;`model stage` 不帶強度時列出各階段實際生效的強度。終端機每次呼叫會顯示送給 CLI 的模型名稱;CLI 回報的實際模型不同時也會顯示(Claude Code 與 Gemini CLI 每次回報,Codex 只在模型被改派時回報);`status <id>` 會按階段、任務、模型與步驟顯示用量。`run --model-mode balanced` 可暫時回到原設定。
134
139
 
135
140
  `status <id>` 的用量以每次 LLM 呼叫為一筆,失敗、額度用完及代打也會計入呼叫次數。只有 CLI 同時回報輸入與輸出 token,才把兩者納入合計與模型強度占比;明確回報的 0 仍算已回報。缺少任一數字列為「未回報」;舊紀錄無法分辨真實 0 與預設補值,列為「舊紀錄不明」,原始數字只供查閱。各 agent、階段、任務、模型與步驟、模型強度是同一批呼叫的不同分組,不應跨組相加。輸入 token 一律包含 cache 讀取與寫入:Claude Code 回報的 `input_tokens` 不含 cache,agentflowctl 會把 cache 讀寫加回去;Codex 的 `input_tokens` 本來就包含 cache。有回報 cache 時,`status` 與 `logs` 會另外標出其中讀取與寫入 cache 各多少。`model add/check` 的探測請求可能耗用 token,但不屬於 run,因此不在 `status` 內。
136
141
 
137
142
  計畫 agent 會查閱相關程式碼,依影響範圍、技術不確定性與失敗後果為每個任務標註 `low`/`medium`/`high` 難度,取三者中最高等級,並在計畫中寫出依據;計畫審查會逐項核對。自動選模先遵守角色分配,再取階段強度與任務難度中較高者;失敗重試會提高強度。若分配到的 agent 沒有足夠強度的模型,會選它最強的模型並提示。這些強度是你對模型能力的設定,不由 CLI 自動評分。
138
143
 
139
- 模型探測必須能禁止工具。這版 Codex CLI 沒有可確認的無工具探測參數,因此 `model add` 無法驗證並登記 Codex 模型;手動寫入 `models` 仍可執行,但不代表已驗證可用。`balanced` 仍可照原方式使用。自訂 `command` adapter 另需提供會實際呼叫模型的探測命令;設定方式與限制見[詳細參考](docs/reference.md#模型設定與自動選模)。
144
+ #### 模型驗證
145
+
146
+ Claude Code、Codex 與 Gemini CLI 都使用專用探測流程:以 `--model` 指定模型,在獨立暫存目錄要求「只回答 OK」,不沿用正常工作時的 `extraArgs`。
147
+
148
+ | CLI | 探測時的工具限制 |
149
+ |---|---|
150
+ | Claude Code | `--tools ""` 停用內建工具,`--strict-mcp-config` 不載入外部 MCP,`--disable-slash-commands` 停用 slash commands |
151
+ | Codex | 唯讀沙箱與 `approval_policy="never"` 禁止寫入及權限升級;停用 shell、外部工具與 hooks,忽略使用者設定及 rules。仍可能有模型內建檔案工具,由唯讀沙箱限制 |
152
+ | Gemini CLI | admin/user policy 禁止所有工具(合法優先級 `999`),停用 extensions、MCP 與 hooks;政策載入問題使驗證失敗 |
153
+
154
+ 三家共用相同的通過條件:CLI 結束碼為 0、有非空文字與成功完成事件,且沒有工具嘗試或失敗事件。即使後來回覆成功,先前的失敗也不會被忽略。
155
+
156
+ 探測上限為 30 秒,結束後清除暫存目錄;macOS/Linux 逾時或按 Ctrl-C 中斷時會停止整個程序群組,包含 CLI 啟動的子程序。`model add` 成功才新增模型,失敗保持設定原狀;`model check` 只重驗、不修改設定。CLI 不支援探測參數(版本過舊或參數已改名)時直接失敗並提示更新 CLI,不會改用更寬鬆的權限重試。
157
+
158
+ Codex 另有幾點差異:
159
+ - 只寫在 stderr 的工具拒絕(例如唯讀沙箱擋下 patch)同樣算失敗。
160
+ - Codex 的 `error` item 是非致命通知,例如找不到模型 metadata 改用預設值、設定警告、棄用提示,不影響探測與 run 結果;`logs` 以 ⚠️ 顯示,不列入錯誤段落。真正的失敗是 `turn.failed` 與頂層 `error` 事件。
161
+ - Codex 只在模型被改派時回報實際模型,其餘情況只確認指定名稱可呼叫,不推測別名對應。
162
+ - `logs` 會顯示 Codex 每回合的完成事件,跨回合的 shell 指令分開整理。
163
+
164
+ 自訂 `command` adapter 另需提供會實際呼叫模型的探測命令;設定方式與各 CLI 已查核版本見[詳細參考](docs/reference.md#模型設定與自動選模)。
140
165
 
141
166
  ### 專案設定
142
167
 
@@ -165,12 +190,17 @@ agentflowctl run --req-file ./requirement.md
165
190
  | `reviewQuorum` | `1` | 任務與最終程式碼審查需要幾位不同審查者核准 |
166
191
  | `planReviewQuorum` | `1` | 計畫需要幾位不同審查者核准 |
167
192
  | `planArbiter` | `true` | 計畫審查僵持時是否啟用仲裁 |
193
+ | `planReviewLayers` | `{ "enabled": true, "minTasks": 7, "maxGroups": 5, "tasksPerGroup": 3 }` | 任務夠多時把計畫審查拆成索引與任務群;說明見表格下方 |
168
194
  | `tieBreak` | `"proceed"` | 兩位仲裁者意見分歧時,`"proceed"` 繼續、`"stop"` 停止 |
169
195
  | `maxAgentRuns` | `60` | 一次 run 最多執行幾次 agent;可用指令選項覆蓋 |
170
196
  | `install`、`test` | 依專案偵測 | 寫成指令字串,例如 `"install": "pnpm install"` |
171
197
  | `checks` | 依專案偵測 | 檢查清單,例如 `[{ "name": "test", "cmd": "pnpm test" }]`;提供時會取代整份預設清單 |
172
198
  | `testPattern` | 常見的 `.test.`、`.spec.` 檔名 | 辨識測試檔的正規表示式字串;非標準檔名時調整 |
173
199
 
200
+ 計畫審查會依任務規模選做法。同時符合下列條件時,每輪先做一次索引審查,再只審查有變動的任務群:任務達到 `planReviewLayers.minTasks` 個;依 description 寫的檔案路徑能分成至少兩群,而且最大一群不超過三分之二;`plan.md` 每個任務都有 `## T-<數字>` 標題。索引審查讀規格、全部任務描述、驗收條件與整體做法,人數是 `planReviewQuorum`。群數最多 `maxGroups`,也不超過任務數除以 `tasksPerGroup`;每群一位審查者,含 `high` 任務的群改由 `planReviewQuorum` 位審查。改了 `plan.md` 的整體做法時所有群都重審;某一次審查失敗時只重跑還沒完成的部分。已達門檻卻不符其他條件時,終端機會印出原因並改由審查者讀完整份規格與計畫。`"planReviewLayers": { "enabled": false }` 可以關閉,`doctor` 會顯示目前的設定。
201
+
202
+ 審查意見的處理寫在 `.flow/plan-replies.md`,每輪覆寫,不寫進 `plan.md` 文末。下一輪索引會看到整份回應;任務群只看到自己的 `## T-<數字>` 節。
203
+
174
204
  `install`、`test`、`checks` 未設定時,會依 `packageManager`、lockfile 和 `package.json` scripts 偵測。完整範例見 [examples/flow.config.json](examples/flow.config.json)。專案設定每一步都會重新讀取,但已建立 run 的參與 agent 與執行次數上限會沿用建立時的值;要調高後者請用 `resume --max-agent-runs`。
175
205
 
176
206
  ### 環境變數
@@ -181,7 +211,7 @@ agentflowctl run --req-file ./requirement.md
181
211
  | `AGENTFLOWCTL_VERBOSE` | 未開啟 | 設為 `1` 顯示詳細輸出,效果同 `-v` |
182
212
  | `AGENTFLOWCTL_MAX_TURNS` | `80` | 目前程式會讀取此值,但尚未用它限制 agent 執行 |
183
213
 
184
- 環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限;`maxAgentRuns` 則是整次 run 的 agent 執行次數上限。修正成功、或計畫審查與程式碼審查整組完成一輪有效審查後,該關的失敗次數會歸零,所以上限只計算連續失敗。
214
+ 環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限;`maxAgentRuns` 則是整次 run 的 agent 執行次數上限。修正成功、或計畫審查與程式碼審查整組完成一輪有效審查後,該關的失敗次數會歸零,所以上限只計算連續失敗。分層計畫審查時,同一輪裡只要有一次審查呼叫真的執行成功,計畫審查的失敗次數也會歸零;所以索引與各群輪流各失敗一次、每次重跑都有進展時,不會因累計達上限而失敗。
185
215
 
186
216
  ## 更多文件
187
217
 
@@ -1,4 +1,49 @@
1
1
  import { num, str, tryJson } from "./types.js";
2
+ /** 不是工具呼叫的 item;其餘 item(shell、MCP、搜尋、todo……)都當成工具 */
3
+ const NON_TOOL_ITEMS = new Set(["agent_message", "reasoning", "error"]);
4
+ function parse(line) {
5
+ const ev = tryJson(line);
6
+ if (!ev)
7
+ return [];
8
+ const out = [];
9
+ if (ev.type === "item.completed") {
10
+ const item = (ev.item ?? {});
11
+ if (item.type === "agent_message" && str(item.text)?.trim())
12
+ out.push({ kind: "text", text: str(item.text) });
13
+ else if (item.type === "command_execution")
14
+ out.push({ kind: "tool", name: "shell", detail: str(item.command) });
15
+ else if (item.type === "file_change") {
16
+ const changes = Array.isArray(item.changes) ? item.changes : [];
17
+ const paths = changes.map((c) => str(c.path)).filter(Boolean);
18
+ out.push({ kind: "tool", name: "edit", detail: paths.length ? paths.join(", ") : undefined });
19
+ }
20
+ else if (item.type === "error") {
21
+ // error item 是非致命通知(設定警告、棄用提示、模型改派);真正的失敗走 turn.failed 或頂層 error
22
+ const message = str(item.message) ?? "Codex 回報錯誤";
23
+ const rerouted = /^model rerouted: \S+ -> (\S+)/.exec(message)?.[1];
24
+ if (rerouted)
25
+ out.push({ kind: "model", id: rerouted });
26
+ out.push({ kind: "warning", message });
27
+ }
28
+ else if (typeof item.type === "string" && !NON_TOOL_ITEMS.has(item.type)) {
29
+ out.push({ kind: "tool", name: item.type, detail: str(item.tool) ?? str(item.query) });
30
+ }
31
+ }
32
+ else if (ev.type === "turn.completed") {
33
+ const usage = (ev.usage ?? {});
34
+ if (num(usage.input_tokens) !== undefined || num(usage.output_tokens) !== undefined) {
35
+ const cacheRead = num(usage.cached_input_tokens);
36
+ out.push({ kind: "usage", inputTokens: num(usage.input_tokens), outputTokens: num(usage.output_tokens),
37
+ ...(cacheRead !== undefined && { cacheReadTokens: cacheRead }) });
38
+ }
39
+ out.push({ kind: "done", ok: true });
40
+ }
41
+ else if (ev.type === "turn.failed" || ev.type === "error") {
42
+ const err = (ev.error ?? {});
43
+ out.push({ kind: "done", ok: false, summary: str(err.message) ?? str(ev.message) ?? "Codex 執行失敗" });
44
+ }
45
+ return out;
46
+ }
2
47
  /**
3
48
  * OpenAI Codex CLI:`codex exec --json`,prompt 由 stdin 傳入(`-`)。
4
49
  * workspace-write 沙箱只允許修改工作目錄,且預設不能連網,
@@ -6,41 +51,36 @@ import { num, str, tryJson } from "./types.js";
6
51
  */
7
52
  export const codex = {
8
53
  probe: () => ({ cmd: "codex", args: ["--version"] }),
9
- invoke: (o) => ({
54
+ // Codex 沒有 --tools 空清單;停用可執行工具,剩餘檔案工具由唯讀沙箱限制。
55
+ invokeModelProbe: (model, cwd) => ({
10
56
  cmd: "codex",
11
- args: ["exec", "--json", "--sandbox", "workspace-write", "-C", o.cwd, ...(o.model ? ["-m", o.model] : []), ...o.extraArgs, "-"],
12
- input: o.prompt,
57
+ args: ["exec", "--json", "--sandbox", "read-only", "--skip-git-repo-check", "--ephemeral",
58
+ "--ignore-user-config", "--ignore-rules", "--strict-config", "-C", cwd, "-m", model,
59
+ "-c", 'approval_policy="never"', "-c", 'web_search="disabled"', "-c", "mcp_servers={}",
60
+ "-c", "project_doc_max_bytes=0",
61
+ "-c", "tools.experimental_request_user_input.enabled=false", "-c", "tools.update_plan.enabled=false",
62
+ ...["shell_tool", "hooks", "apps", "plugins", "tool_suggest", "multi_agent", "multi_agent_v2",
63
+ "browser_use", "computer_use", "image_generation", "view_image", "code_mode", "code_mode_host",
64
+ "goals", "memories", "sleep_tool", "unbounded_connection_retries"].flatMap((feature) => ["--disable", feature]), "-"],
65
+ input: "只回答 OK,不要呼叫任何工具。",
13
66
  }),
14
- parse(line) {
67
+ // 探測也拒絕尚未完成的工具;一般 log 仍在 item.completed 時顯示結果。
68
+ parseModelProbe(line) {
69
+ const out = parse(line);
15
70
  const ev = tryJson(line);
16
- if (!ev)
17
- return [];
18
- const out = [];
19
- if (ev.type === "item.completed") {
20
- const item = (ev.item ?? {});
21
- if (item.type === "agent_message" && str(item.text)?.trim())
22
- out.push({ kind: "text", text: str(item.text) });
23
- else if (item.type === "command_execution")
24
- out.push({ kind: "tool", name: "shell", detail: str(item.command) });
25
- else if (item.type === "file_change") {
26
- const changes = Array.isArray(item.changes) ? item.changes : [];
27
- const paths = changes.map((c) => str(c.path)).filter(Boolean);
28
- out.push({ kind: "tool", name: "edit", detail: paths.length ? paths.join(", ") : undefined });
29
- }
30
- }
31
- else if (ev.type === "turn.completed") {
32
- const usage = (ev.usage ?? {});
33
- if (num(usage.input_tokens) !== undefined || num(usage.output_tokens) !== undefined) {
34
- const cacheRead = num(usage.cached_input_tokens);
35
- out.push({ kind: "usage", inputTokens: num(usage.input_tokens), outputTokens: num(usage.output_tokens),
36
- ...(cacheRead !== undefined && { cacheReadTokens: cacheRead }) });
37
- }
38
- }
39
- else if (ev.type === "turn.failed" || ev.type === "error") {
40
- const err = (ev.error ?? {});
41
- out.push({ kind: "done", ok: false, summary: str(err.message) ?? str(ev.message) ?? "Codex 執行失敗" });
71
+ const item = ev?.item;
72
+ if (ev?.type === "item.started" && typeof item?.type === "string" && !NON_TOOL_ITEMS.has(item.type)) {
73
+ out.push({ kind: "tool", name: item.type });
42
74
  }
43
75
  return out;
44
76
  },
77
+ // 拒絕 apply_patch 時可能只寫 stderr,沒有 file_change 事件。
78
+ modelProbeFailure: (stderr) => /patch rejected|tools::router.*error=/i.test(stderr) ? `模型檢查期間嘗試呼叫工具:${stderr.trim()}` : undefined,
79
+ invoke: (o) => ({
80
+ cmd: "codex",
81
+ args: ["exec", "--json", "--sandbox", "workspace-write", "-C", o.cwd, ...(o.model ? ["-m", o.model] : []), ...o.extraArgs, "-"],
82
+ input: o.prompt,
83
+ }),
84
+ parse,
45
85
  };
46
86
  //# sourceMappingURL=codex.js.map
@@ -1,4 +1,4 @@
1
- import { writeFileSync } from "node:fs";
1
+ import { mkdirSync, writeFileSync } from "node:fs";
2
2
  import { join } from "node:path";
3
3
  import { num, str, toolDetail, tryJson } from "./types.js";
4
4
  /**
@@ -11,11 +11,17 @@ export const gemini = {
11
11
  probe: () => ({ cmd: "gemini", args: ["--version"] }),
12
12
  invokeModelProbe: (model, cwd) => {
13
13
  const policy = join(cwd, "deny-tools.toml");
14
- writeFileSync(policy, '[[rule]]\ntoolName = "*"\ndecision = "deny"\npriority = 10000\n');
14
+ mkdirSync(join(cwd, ".gemini"), { recursive: true });
15
+ writeFileSync(join(cwd, ".gemini", "settings.json"), JSON.stringify({ hooksConfig: { enabled: false } }));
16
+ // priority 是 tier 內的 0–999;不能用 10000 嘗試跨 tier。
17
+ writeFileSync(policy, '[[rule]]\ntoolName = "*"\ndecision = "deny"\npriority = 999\n');
15
18
  return { cmd: "gemini", args: ["-p", "只回答 OK", "--model", model, "--output-format", "stream-json",
16
- "--approval-mode", "default", "--extensions", "none", "--policy", policy],
19
+ "--approval-mode", "default", "--extensions", "none", "--allowed-mcp-server-names", "",
20
+ "--admin-policy", policy, "--policy", policy],
17
21
  env: { GEMINI_CLI_TRUST_WORKSPACE: "true" } };
18
22
  },
23
+ modelProbeFailure: (stderr) => /ignoring --admin-policy|policy file (error|warning)|error loading policy|invalid policy|failed to load.*polic/i.test(stderr)
24
+ ? `CLI 未完整載入模型探測的工具限制:${stderr.trim()}` : undefined,
19
25
  invoke: (o) => ({
20
26
  cmd: "gemini",
21
27
  args: ["-p", o.prompt, "--output-format", "stream-json", "--approval-mode", "yolo", ...(o.model ? ["-m", o.model] : []), ...o.extraArgs],
@@ -41,7 +47,8 @@ export const gemini = {
41
47
  const outputTokens = num(stats.output_tokens) ?? num(stats.outputTokens);
42
48
  if (inputTokens !== undefined || outputTokens !== undefined)
43
49
  out.push({ kind: "usage", inputTokens, outputTokens });
44
- out.push({ kind: "done", ok: ev.status !== "error", summary: str(ev.response) });
50
+ const error = (ev.error ?? {});
51
+ out.push({ kind: "done", ok: ev.status !== "error", summary: str(error.message) ?? str(ev.response) });
45
52
  }
46
53
  else if (ev.type === "error") {
47
54
  out.push({ kind: "done", ok: false, summary: str(ev.message) ?? "Gemini 執行失敗" });
package/dist/cli.js CHANGED
@@ -13,9 +13,11 @@ import { describeDetected, detectProjectDefaults } from "./detect.js";
13
13
  import { CMD_AGENT, listLogs, localTime, logMark, nextLogFile, renderLog } from "./logs.js";
14
14
  import { flowDir, logDir, projectRoot, worktreeDir } from "./paths.js";
15
15
  import { ModelStage, ModelStrength, TaskList } from "./schemas.js";
16
+ import { computeInsights, failureLabel, retryLabel } from "./insights.js";
17
+ import { computeUsageInsights } from "./usageInsights.js";
16
18
  import { computeStats, formatDuration } from "./stats.js";
17
- import { agentRuns, getRun, listRuns, listSubstitutions, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
18
- import { readJsonFile } from "./util.js";
19
+ import { agentRuns, getRun, listRetries, listRuns, listSubstitutions, listUsage, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
20
+ import { padDisplay, readJsonFile } from "./util.js";
19
21
  import { openActions, readHandoff } from "./handoff.js";
20
22
  import { stopReport } from "./stopReport.js";
21
23
  import { runSetup, SETUP_ADAPTERS } from "./setup.js";
@@ -191,7 +193,7 @@ program
191
193
  run = { ...run, stage: run.pausedStage ?? "spec", pausedStage: undefined, pauseReason: undefined };
192
194
  }
193
195
  if (run.stage === "failed") {
194
- run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, modelRetryAttempts: {}, failedStage: undefined, failureReason: undefined };
196
+ run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, modelRetryAttempts: {}, failedStage: undefined, failureReason: undefined, failureCategory: undefined };
195
197
  }
196
198
  await drive(saveRun(run));
197
199
  });
@@ -200,7 +202,7 @@ program
200
202
  .description("標記 run 為失敗(執行中的 run 請直接在該終端機按 Ctrl-C)")
201
203
  .action((id) => {
202
204
  const run = mustGetRun(id);
203
- saveRun({ ...run, stage: "failed", failedStage: run.stage, failureReason: "使用者取消" });
205
+ saveRun({ ...run, stage: "failed", failedStage: run.stage, failureCategory: "cancelled", failureReason: "使用者取消" });
204
206
  console.log(`已取消 ${id}`);
205
207
  });
206
208
  program
@@ -266,6 +268,13 @@ program
266
268
  for (const sub of subs)
267
269
  console.log(` ${sub.at.slice(0, 16)} ${sub.step.padEnd(14)} ${sub.planned} → ${sub.actual}${sub.note ? `(${sub.note})` : ""}`);
268
270
  }
271
+ const retries = listRetries(id);
272
+ if (retries.length) {
273
+ console.log("\n重試紀錄");
274
+ for (const r of retries) {
275
+ console.log(` ${r.key.padEnd(19)} ${retryLabel(r.category)}${r.final ? "(達上限)" : `(第 ${r.attempt} 次)`}`);
276
+ }
277
+ }
269
278
  const tasks = readJsonFile(join(flowDir(id), "tasks.ordered.json"), TaskList);
270
279
  if (!tasks.ok)
271
280
  return;
@@ -502,6 +511,8 @@ async function doctor() {
502
511
  console.log(`\n單一 run 的 agent 執行上限:${cfg.maxAgentRuns} 次`);
503
512
  console.log(`修正策略:${cfg.fixStrategy} 測試與實作分開:${cfg.tddSplit ? "是" : "否"}`);
504
513
  console.log(`程式碼審查人數:${cfg.reviewQuorum} 計畫審查人數:${cfg.planReviewQuorum} 計畫仲裁:${cfg.planArbiter ? "開啟" : "關閉"}`);
514
+ const layers = cfg.planReviewLayers;
515
+ console.log(`計畫分層審查:${layers.enabled ? `任務達 ${layers.minTasks} 個時開啟,最多 ${layers.maxGroups} 群,每群平均至少 ${layers.tasksPerGroup} 個任務` : "關閉"}`);
505
516
  }
506
517
  program.command("doctor").description("檢查可用的 agent CLI 與目前參與的 agent").action(doctor);
507
518
  program
@@ -560,6 +571,105 @@ program
560
571
  }
561
572
  console.log(`\n次數多或失敗多的步驟可用 agentflowctl logs ${id} 找出編號查看原因;token 用量見 agentflowctl status ${id}`);
562
573
  });
574
+ program
575
+ .command("insights")
576
+ .description("彙總這個專案所有 run 的結果、失敗原因、用量、步驟失敗與重試原因,並給出改善建議")
577
+ .action(() => {
578
+ const runs = listRuns();
579
+ if (!runs.length)
580
+ return console.log("還沒有 run");
581
+ const rows = runs.map((r) => ({
582
+ id: r.id,
583
+ stage: r.stage,
584
+ failureCategory: r.failureCategory,
585
+ retries: listRetries(r.id),
586
+ usage: listUsage(r.id),
587
+ substitutions: listSubstitutions(r.id).length,
588
+ stats: computeStats(listLogs(logDir(r.id))),
589
+ }));
590
+ const insights = computeInsights(rows);
591
+ const usage = computeUsageInsights(rows);
592
+ const o = insights.byOutcome;
593
+ console.log(`專案彙總(${insights.runCount} 個 run)`);
594
+ console.log(` 完成 ${o.done} 失敗 ${o.failed} 暫停 ${o.paused} 等待核准 ${o.awaiting_approval} 進行中 ${o.active}`);
595
+ if (insights.byFailure.length) {
596
+ console.log("\n失敗原因");
597
+ for (const row of insights.byFailure)
598
+ console.log(` ${padDisplay(failureLabel(row.category), 20)} ${String(row.count).padStart(4)} 個 run`);
599
+ }
600
+ const t = usage.total;
601
+ console.log("\n用量");
602
+ if (t.runs) {
603
+ const value = t.reportedRuns
604
+ ? `輸入 ${t.inputTokens}${cacheNote(t)}、輸出 ${t.outputTokens}、合計 ${t.tokens} tokens`
605
+ : "未回報或回報狀態不明";
606
+ console.log(` ${value};${t.runs} 次(未回報 ${t.unreportedRuns}、舊紀錄不明 ${t.legacyRuns})`);
607
+ }
608
+ else
609
+ console.log(" 還沒有用量紀錄");
610
+ const byStrength = usage.byStrength;
611
+ const reportedTotal = Object.values(byStrength).reduce((sum, entry) => sum + entry.tokens, 0);
612
+ if (Object.keys(byStrength).length) {
613
+ console.log("\n模型強度用量(占比只計入明確回報)");
614
+ for (const strength of ["low", "medium", "high", "未知"]) {
615
+ const entry = byStrength[strength];
616
+ if (!entry)
617
+ continue;
618
+ const share = reportedTotal ? `${(entry.tokens / reportedTotal * 100).toFixed(1)}%` : "無法計算";
619
+ console.log(` ${strength}: ${entry.tokens} tokens,占比 ${share},呼叫 ${entry.runs} 次(未回報 ${entry.unreportedRuns}、舊紀錄不明 ${entry.legacyRuns})`);
620
+ }
621
+ console.log(` 高強度呼叫:${byStrength.high?.runs ?? 0} 次`);
622
+ }
623
+ const byTokens = (record, n) => Object.entries(record).sort((a, b) => b[1].tokens - a[1].tokens).slice(0, n);
624
+ printUsage("用量最高的階段", byTokens(usage.byStage, 8));
625
+ printUsage("各 agent 用量", byTokens(usage.byAgent, 8));
626
+ printUsage("各模型與步驟用量(任務步驟不分 task id,只加總明確回報)", byTokens(usage.byModelStage, 10));
627
+ const ss = usage.stepStats;
628
+ if (ss.steps.length) {
629
+ const totalMs = ss.agentMs + ss.cmdMs;
630
+ const share = (ms) => (totalMs ? `${(ms / totalMs * 100).toFixed(0)}%` : "-");
631
+ console.log("\n步驟執行與失敗(跨 run,任務步驟不分 task id,依失敗次數排序,不含總經過時間)");
632
+ console.log(` agent ${formatDuration(ss.agentMs).padStart(7)} ${share(ss.agentMs)}`);
633
+ console.log(` 專案指令 ${formatDuration(ss.cmdMs).padStart(7)} ${share(ss.cmdMs)}`);
634
+ if (ss.unfinished)
635
+ console.log(` 未完成 ${ss.unfinished} 份(沒有結束紀錄,不計入耗時)`);
636
+ console.log("\n 步驟 類型 次數 失敗 總耗時 最長");
637
+ for (const s of ss.steps.slice(0, 10)) {
638
+ console.log(` ${s.step.padEnd(19)} ${s.kind === "cmd" ? "指令 " : "agent"} ${String(s.runs).padStart(4)} ${String(s.failed).padStart(4)} ${formatDuration(s.totalMs).padStart(7)} ${formatDuration(s.maxMs).padStart(7)}${s.unfinished ? ` (未完成 ${s.unfinished})` : ""}`);
639
+ }
640
+ }
641
+ if (!insights.byCategory.length)
642
+ console.log("\n還沒有重試紀錄(舊 run 不會回填)");
643
+ else {
644
+ console.log("\n重試原因");
645
+ for (const row of insights.byCategory) {
646
+ const finals = row.finals ? `,其中 ${row.finals} 次達上限` : "";
647
+ console.log(` ${padDisplay(retryLabel(row.category), 20)} ${String(row.count).padStart(4)} 次${finals}`);
648
+ }
649
+ console.log("\n最常重試的關卡(任務關卡不分 task id 合併計算)");
650
+ for (const row of insights.byGate.slice(0, 10)) {
651
+ console.log(` ${padDisplay(row.gate, 19)} ${String(row.count).padStart(4)} 次`);
652
+ }
653
+ }
654
+ console.log("\n建議");
655
+ if (!usage.findings.length)
656
+ console.log(" 還沒有足夠訊號");
657
+ else
658
+ usage.findings.forEach((f, i) => {
659
+ console.log(` ${i + 1}. ${f.title}`);
660
+ console.log(` ${f.detail}`);
661
+ });
662
+ const usageById = new Map(usage.runs.map((r) => [r.id, r]));
663
+ console.log("\n各 run");
664
+ for (const r of insights.runs) {
665
+ const cat = r.topCategory ? ` 最多 ${retryLabel(r.topCategory)}` : "";
666
+ const u = usageById.get(r.id);
667
+ const tokens = !u?.calls ? "" : u.reportedCalls ? ` ${String(u.tokens).padStart(9)} tokens` : ` ${" ".repeat(3)}未回報${" ".repeat(7)}`;
668
+ const task = u?.topTask && u.topTask.tasks >= 2 ? ` 最耗任務 ${u.topTask.task}(${(u.topTask.share * 100).toFixed(0)}%)` : "";
669
+ console.log(` ${r.id} ${r.stage.padEnd(17)} 重試 ${String(r.retries).padStart(3)} 次${tokens}${task}${cat}`);
670
+ }
671
+ console.log("\n建議對應可改的 prompt、model stage 或關卡;各分組是同一批呼叫的不同切片,不要跨組相加。單一 run 用 agentflowctl status <id>、stats <id> 與 logs <id>");
672
+ });
563
673
  program.parseAsync().catch((err) => {
564
674
  console.error(`錯誤:${err.message}`);
565
675
  process.exit(1);