agentflowctl 0.13.1 → 0.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +19 -9
- package/dist/cli.js +126 -5
- package/dist/config.js +5 -3
- package/dist/detect.js +21 -4
- package/dist/engine.js +347 -105
- package/dist/insights.js +86 -0
- package/dist/modelSelection.js +19 -11
- package/dist/paths.js +4 -0
- package/dist/planReview.js +333 -0
- package/dist/schemas.js +18 -0
- package/dist/stats.js +26 -0
- package/dist/store.js +62 -17
- package/dist/usageInsights.js +156 -0
- package/dist/util.js +12 -0
- package/examples/flow.config.json +1 -0
- package/package.json +1 -1
- package/prompts/implement-direct.md +67 -0
- package/prompts/plan-arbiter.md +6 -1
- package/prompts/plan-fix.md +5 -2
- package/prompts/plan-review-group.md +92 -0
- package/prompts/plan-review-index.md +93 -0
- package/prompts/plan-review.md +2 -1
- package/prompts/plan.md +6 -4
package/README.md
CHANGED
|
@@ -31,13 +31,13 @@ npx agentflowctl run --req-file ./requirement.md
|
|
|
31
31
|
## 執行時會發生什麼
|
|
32
32
|
|
|
33
33
|
1. agent 整理需求與驗收條件,接著寫計畫,交給其他 agent 審查。
|
|
34
|
-
2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent
|
|
34
|
+
2. 依計畫逐個任務寫出會失敗的測試,再由另一位 agent 實作到測試通過;每個任務都會經過審查與驗證。計畫 agent 會依改動內容在任務標記 `tdd`:建置流程、設定、文件、型別、純重構這類不適合先寫失敗測試的任務,會略過紅綠燈直接實作,改由任務審查與驗證把關。專案沒有測試框架時,所有任務都略過紅綠燈,也不跑 `test` 檢查。
|
|
35
35
|
3. 全部任務完成後,再執行專案檢查與整體程式碼審查。未通過的項目會交回修正。
|
|
36
36
|
4. 有 `origin` 時會推送分支;若 `gh` 可用,會嘗試建立 PR。沒有 `origin` 時,完成的分支留在本機。
|
|
37
37
|
|
|
38
38
|
流程預設會自動往下走。想在計畫通過審查後親自確認,可加 `--manual-plan`;確認後執行 `agentflowctl approve <id>`。
|
|
39
39
|
|
|
40
|
-
agentflowctl 會依專案的 `packageManager`、lockfile 與 `package.json` scripts 選擇安裝、測試及檢查指令。第一次執行時,請留意終端機印出的偵測結果;需要調整可在 `flow.config.json` 指定 `install`、`test` 或 `checks`。
|
|
40
|
+
agentflowctl 會依專案的 `packageManager`、lockfile 與 `package.json` scripts 選擇安裝、測試及檢查指令。第一次執行時,請留意終端機印出的偵測結果;需要調整可在 `flow.config.json` 指定 `install`、`test` 或 `checks`。`package.json` 的依賴或 `test` script 看不出測試框架(且沒有手動設定 `test`)時,終端機會提示「未偵測到測試框架」,並略過紅綠燈;要改回來,在 `flow.config.json` 設定 `test`。
|
|
41
41
|
|
|
42
42
|
## 查看進度
|
|
43
43
|
|
|
@@ -49,6 +49,7 @@ agentflowctl status f-xxxx # 看進度、結果與下一步
|
|
|
49
49
|
agentflowctl logs f-xxxx # 列出各步驟的 log
|
|
50
50
|
agentflowctl logs f-xxxx --latest # 看最新一份 log
|
|
51
51
|
agentflowctl stats f-xxxx # 各步驟耗時、執行與失敗次數
|
|
52
|
+
agentflowctl insights # 這個專案所有 run 的結果、失敗原因、用量、步驟失敗、重試原因與改善建議
|
|
52
53
|
agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
|
|
53
54
|
```
|
|
54
55
|
|
|
@@ -56,7 +57,9 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
|
|
|
56
57
|
|
|
57
58
|
`stats` 依 log 的開始與結束時間統計每個步驟的執行次數、失敗次數、總耗時與最長一次,並分開列出 agent 與專案指令(install、測試、checks)各占多少時間,最耗時的步驟排在最前面。沒有結束紀錄的 log 列為未完成,不計入耗時;總經過時間包含暫停與等待核准。
|
|
58
59
|
|
|
59
|
-
|
|
60
|
+
`insights` 把所有 run 分成獨立區塊彙總:最終狀態、失敗原因(重試達上限、仲裁停止、agent 次數用完等)、用量(輸入/輸出/cache、強度占比、階段/agent,各 run 用量最高的任務)、各模型與步驟、步驟執行與失敗(與 `stats` 相同取 log 的檔頭檔尾,跨 run 不計總經過時間)、關卡重試原因。任務關卡與任務步驟不分 task id 合併計算。最後列出最多五則建議,規則由程式套門檻,不是再請 agent 分析。合計 token 不是主指標。覆蓋不足時會先警告占比可能失真。各分組是同一批呼叫的不同切片,不要跨組相加。舊 run 沒有重試或失敗原因紀錄,不會回填。單一 run 的全量明細仍用 `status <id>` 與 `stats <id>`。
|
|
61
|
+
|
|
62
|
+
執行紀錄在 `.agentflowctl/runs/<id>/`,工作分支在 `.agentflowctl/worktrees/<id>/`。不再需要某次 run 時,可用 `agentflowctl clean <id>` 清除 worktree 與紀錄;`agentflowctl clean --all` 一次清除所有 done、failed 的 run,以及沒有紀錄的 worktree(進行中、暫停、等待核准的不動)。`flow/<id>` 分支會保留。
|
|
60
63
|
|
|
61
64
|
## 執行停下來時怎麼做
|
|
62
65
|
|
|
@@ -66,8 +69,8 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
|
|
|
66
69
|
| --- | --- |
|
|
67
70
|
| 按 Ctrl-C,或終端機意外關閉 | 執行 `agentflowctl resume <id>`;沒有結束紀錄的步驟會重跑 |
|
|
68
71
|
| `awaiting_approval`:計畫等你確認 | 閱讀 `.agentflowctl/worktrees/<id>/.flow/plan.md`,確認後執行 `agentflowctl approve <id>` |
|
|
69
|
-
| `paused`:agent 額度用完 | 等額度恢復後執行 `agentflowctl resume <id>`;審查步驟不會換 agent
|
|
70
|
-
| `
|
|
72
|
+
| `paused`:agent 額度用完 | 等額度恢復後執行 `agentflowctl resume <id>`;審查步驟不會換 agent 代審。在仲裁途中暫停時,resume 直接回到仲裁,不重跑計畫審查;暫停期間若改了計畫檔,或在 `flow.config.json` 把 `planArbiter` 關掉,就改成重新審查 |
|
|
73
|
+
| `failed`:仲裁連續沒有產生有效裁決 | 仲裁者沒寫出 `.flow/plan-arbiter.json`、格式錯誤或交接無效時不會暫停,會把原因寫進 `.flow/feedback.md` 並自動重跑仲裁(不重跑計畫審查);無效的檔案移到 `.agentflowctl/runs/<id>/reviews/plan-arbiter-<輪>-<agent>-invalid.json`。連續達重試上限才失敗,查看 log 後執行 `agentflowctl resume <id>` 會再回到仲裁 |
|
|
71
74
|
| `failed`:測試、檢查、審查或 agent 執行失敗 | 依 `status` 提示查看失敗的 log,處理原因後執行 `agentflowctl resume <id>`;失敗階段會重試 |
|
|
72
75
|
| `failed`:已達 agent 執行次數上限 | 用 `agentflowctl resume <id> --max-agent-runs 100` 調高上限後接續,數字須大於已執行次數 |
|
|
73
76
|
|
|
@@ -94,6 +97,7 @@ agentflowctl resume f-xxxx
|
|
|
94
97
|
| `run --cycle <名單>` | 指定這次參與的 agent,例如 `--cycle claude,codex`;優先於設定檔的 `cycle` |
|
|
95
98
|
| `run --model-mode balanced\|adaptive` | 只覆蓋這次 run 的模型模式;`resume` 沿用建立時的模式 |
|
|
96
99
|
| `run --base <分支>` | 指定起始分支;未設定時使用目前分支 |
|
|
100
|
+
| `run --max-attempts <次數>` | 覆蓋這次的重試上限(至少 3;預設取 `AGENTFLOWCTL_MAX_ATTEMPTS`,未設定為 5);失敗後可用 `resume <id> --max-attempts <次數>` 調高 |
|
|
97
101
|
| `run --max-agent-runs <次數>` | 覆蓋這次的 `maxAgentRuns`;上限不夠時可用 `resume <id> --max-agent-runs <次數>` 調高 |
|
|
98
102
|
| `-v` / `--verbose` | 執行時顯示 agent 文字、工具呼叫與專案指令,適用於 `run`、`resume`、`approve` |
|
|
99
103
|
|
|
@@ -102,6 +106,7 @@ agentflowctl resume f-xxxx
|
|
|
102
106
|
```bash
|
|
103
107
|
agentflowctl run --req-file ./requirement.md --cycle claude,codex --max-agent-runs 80 --manual-plan
|
|
104
108
|
agentflowctl resume f-xxxx --max-agent-runs 100
|
|
109
|
+
agentflowctl resume f-xxxx --max-attempts 8
|
|
105
110
|
```
|
|
106
111
|
|
|
107
112
|
### Agent 設定
|
|
@@ -186,24 +191,29 @@ Codex 另有幾點差異:
|
|
|
186
191
|
| `tddSplit` | `true` | 有多位 agent 時,`true` 會把同一任務的測試與實作分給不同 agent |
|
|
187
192
|
| `reviewQuorum` | `1` | 任務與最終程式碼審查需要幾位不同審查者核准 |
|
|
188
193
|
| `planReviewQuorum` | `1` | 計畫需要幾位不同審查者核准 |
|
|
189
|
-
| `planArbiter` | `true` |
|
|
194
|
+
| `planArbiter` | `true` | 計畫審查僵持,或修訂一次後仍被要求修改時是否啟用仲裁 |
|
|
195
|
+
| `planReviewLayers` | `{ "enabled": true, "minTasks": 7, "maxGroups": 5, "tasksPerGroup": 3 }` | 任務夠多時把計畫審查拆成索引與任務群;說明見表格下方 |
|
|
190
196
|
| `tieBreak` | `"proceed"` | 兩位仲裁者意見分歧時,`"proceed"` 繼續、`"stop"` 停止 |
|
|
191
197
|
| `maxAgentRuns` | `60` | 一次 run 最多執行幾次 agent;可用指令選項覆蓋 |
|
|
192
198
|
| `install`、`test` | 依專案偵測 | 寫成指令字串,例如 `"install": "pnpm install"` |
|
|
193
199
|
| `checks` | 依專案偵測 | 檢查清單,例如 `[{ "name": "test", "cmd": "pnpm test" }]`;提供時會取代整份預設清單 |
|
|
194
200
|
| `testPattern` | 常見的 `.test.`、`.spec.` 檔名 | 辨識測試檔的正規表示式字串;非標準檔名時調整 |
|
|
195
201
|
|
|
202
|
+
計畫審查會依任務規模選做法。同時符合下列條件時,每輪先做一次索引審查,再只審查有變動的任務群:任務達到 `planReviewLayers.minTasks` 個;依 description 寫的檔案路徑能分成至少兩群,而且最大一群不超過三分之二;`plan.md` 每個任務都有 `## T-<數字>` 標題。索引審查讀規格、全部任務描述、驗收條件與整體做法,人數是 `planReviewQuorum`。群數最多 `maxGroups`,也不超過任務數除以 `tasksPerGroup`;每群一位審查者,含 `high` 任務的群改由 `planReviewQuorum` 位審查。改了 `plan.md` 的整體做法時所有群都重審;某一次審查失敗時只重跑還沒完成的部分。已達門檻卻不符其他條件時,終端機會印出原因並改由審查者讀完整份規格與計畫。`"planReviewLayers": { "enabled": false }` 可以關閉,`doctor` 會顯示目前的設定。
|
|
203
|
+
|
|
204
|
+
審查意見的處理寫在 `.flow/plan-replies.md`,每輪覆寫,不寫進 `plan.md` 文末。下一輪索引會看到整份回應;任務群只看到自己的 `## T-<數字>` 節。
|
|
205
|
+
|
|
196
206
|
`install`、`test`、`checks` 未設定時,會依 `packageManager`、lockfile 和 `package.json` scripts 偵測。完整範例見 [examples/flow.config.json](examples/flow.config.json)。專案設定每一步都會重新讀取,但已建立 run 的參與 agent 與執行次數上限會沿用建立時的值;要調高後者請用 `resume --max-agent-runs`。
|
|
197
207
|
|
|
198
208
|
### 環境變數
|
|
199
209
|
|
|
200
210
|
| 變數 | 預設 | 設定方式與用途 |
|
|
201
211
|
| --- | --- | --- |
|
|
202
|
-
| `AGENTFLOWCTL_MAX_ATTEMPTS` | `
|
|
212
|
+
| `AGENTFLOWCTL_MAX_ATTEMPTS` | `5` | 同一關連續失敗幾次後停止,至少 3(設得更小以 3 計);例如 `AGENTFLOWCTL_MAX_ATTEMPTS=10 agentflowctl run --req "..."` |
|
|
203
213
|
| `AGENTFLOWCTL_VERBOSE` | 未開啟 | 設為 `1` 顯示詳細輸出,效果同 `-v` |
|
|
204
|
-
| `AGENTFLOWCTL_MAX_TURNS` | `
|
|
214
|
+
| `AGENTFLOWCTL_MAX_TURNS` | `200` | 目前程式會讀取此值,但尚未用它限制 agent 執行 |
|
|
205
215
|
|
|
206
|
-
環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS`
|
|
216
|
+
環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限(至少 3),單一 run 可用 `run`/`resume` 的 `--max-attempts` 覆蓋;計畫審查何時交付仲裁與它無關:意見沒有變化,或第 2 輪(修訂過一次)仍被要求修改時就交付,兩家 agent 時自動進入雙盲交叉仲裁,有第三方時由第三方單獨仲裁;`maxAgentRuns` 則是整次 run 的 agent 執行次數上限。修正成功、或計畫審查與程式碼審查整組完成一輪有效審查後,該關的失敗次數會歸零,所以上限只計算連續失敗。分層計畫審查時,同一輪裡只要有一次審查呼叫真的執行成功,計畫審查的失敗次數也會歸零;所以索引與各群輪流各失敗一次、每次重跑都有進展時,不會因累計達上限而失敗。
|
|
207
217
|
|
|
208
218
|
## 更多文件
|
|
209
219
|
|
package/dist/cli.js
CHANGED
|
@@ -4,7 +4,7 @@ import { readFileSync } from "node:fs";
|
|
|
4
4
|
import { join } from "node:path";
|
|
5
5
|
import { stdin, stdout } from "node:process";
|
|
6
6
|
import { createInterface } from "node:readline/promises";
|
|
7
|
-
import { config } from "./config.js";
|
|
7
|
+
import { MIN_ATTEMPTS, config } from "./config.js";
|
|
8
8
|
import { advance, loadRepoConfig } from "./engine.js";
|
|
9
9
|
import { probeAgent, resolveAgent, runCommand } from "./runner.js";
|
|
10
10
|
import { addWorktree, git } from "./git.js";
|
|
@@ -13,9 +13,11 @@ import { describeDetected, detectProjectDefaults } from "./detect.js";
|
|
|
13
13
|
import { CMD_AGENT, listLogs, localTime, logMark, nextLogFile, renderLog } from "./logs.js";
|
|
14
14
|
import { flowDir, logDir, projectRoot, worktreeDir } from "./paths.js";
|
|
15
15
|
import { ModelStage, ModelStrength, TaskList } from "./schemas.js";
|
|
16
|
+
import { computeInsights, failureLabel, retryLabel } from "./insights.js";
|
|
17
|
+
import { computeUsageInsights } from "./usageInsights.js";
|
|
16
18
|
import { computeStats, formatDuration } from "./stats.js";
|
|
17
|
-
import { agentRuns, getRun, listRuns, listSubstitutions, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
|
|
18
|
-
import { readJsonFile } from "./util.js";
|
|
19
|
+
import { agentRuns, getRun, listRetries, listRuns, listSubstitutions, listUsage, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
|
|
20
|
+
import { padDisplay, readJsonFile } from "./util.js";
|
|
19
21
|
import { openActions, readHandoff } from "./handoff.js";
|
|
20
22
|
import { stopReport } from "./stopReport.js";
|
|
21
23
|
import { runSetup, SETUP_ADAPTERS } from "./setup.js";
|
|
@@ -33,6 +35,12 @@ function printUsage(title, rows) {
|
|
|
33
35
|
console.log(` ${key}: ${value};${c.runs} 次(未回報 ${c.unreportedRuns}、舊紀錄不明 ${c.legacyRuns})`);
|
|
34
36
|
}
|
|
35
37
|
}
|
|
38
|
+
function positiveInt(text, flag, min = 1) {
|
|
39
|
+
const n = Number(text);
|
|
40
|
+
if (!Number.isInteger(n) || n < min)
|
|
41
|
+
throw new Error(`${flag} 必須是不小於 ${min} 的整數:${text}`);
|
|
42
|
+
return n;
|
|
43
|
+
}
|
|
36
44
|
function mustGetRun(id) {
|
|
37
45
|
const run = getRun(id);
|
|
38
46
|
if (!run)
|
|
@@ -119,6 +127,7 @@ program
|
|
|
119
127
|
.option("--req-file <file>", "從檔案讀取需求")
|
|
120
128
|
.option("--base <branch>", "基底分支(預設為目前的分支)")
|
|
121
129
|
.option("--max-agent-runs <n>", "單一 run 最多執行幾次 agent(預設取 flow.config.json 的 maxAgentRuns)")
|
|
130
|
+
.option("--max-attempts <n>", "同一關連續失敗幾次後停止(預設取 AGENTFLOWCTL_MAX_ATTEMPTS,未設定為 5)")
|
|
122
131
|
.option("--manual-plan", "計畫通過 AI 審查後,仍停下來等你確認", false)
|
|
123
132
|
.option("--cycle <agents>", "參與的 agent,例如 claude,codex,gemini(順序不影響分工)")
|
|
124
133
|
.option("--model-mode <mode>", "這次 run 的模型模式:balanced 或 adaptive")
|
|
@@ -152,6 +161,7 @@ program
|
|
|
152
161
|
stage: "spec",
|
|
153
162
|
autopilot: !opts.manualPlan,
|
|
154
163
|
maxAgentRuns: opts.maxAgentRuns ? Number(opts.maxAgentRuns) : cfg.maxAgentRuns,
|
|
164
|
+
maxAttempts: opts.maxAttempts ? positiveInt(opts.maxAttempts, "--max-attempts", MIN_ATTEMPTS) : undefined,
|
|
155
165
|
cycle,
|
|
156
166
|
attempts: {},
|
|
157
167
|
modelMode,
|
|
@@ -181,17 +191,20 @@ program
|
|
|
181
191
|
.command("resume <id>")
|
|
182
192
|
.description("從暫停、中斷或失敗的階段接續")
|
|
183
193
|
.option("--max-agent-runs <n>", "調整 agent 執行次數上限")
|
|
194
|
+
.option("--max-attempts <n>", "調整同一關連續失敗的上限")
|
|
184
195
|
.action(async (id, opts) => {
|
|
185
196
|
let run = mustGetRun(id);
|
|
186
197
|
if (run.modelMode === "adaptive")
|
|
187
198
|
validateAdaptiveConfig(loadRepoConfig(), run.cycle);
|
|
188
199
|
if (opts.maxAgentRuns)
|
|
189
200
|
run = { ...run, maxAgentRuns: Number(opts.maxAgentRuns) };
|
|
201
|
+
if (opts.maxAttempts)
|
|
202
|
+
run = { ...run, maxAttempts: positiveInt(opts.maxAttempts, "--max-attempts", MIN_ATTEMPTS) };
|
|
190
203
|
if (run.stage === "paused") {
|
|
191
204
|
run = { ...run, stage: run.pausedStage ?? "spec", pausedStage: undefined, pauseReason: undefined };
|
|
192
205
|
}
|
|
193
206
|
if (run.stage === "failed") {
|
|
194
|
-
run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, modelRetryAttempts: {}, failedStage: undefined, failureReason: undefined };
|
|
207
|
+
run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, modelRetryAttempts: {}, failedStage: undefined, failureReason: undefined, failureCategory: undefined };
|
|
195
208
|
}
|
|
196
209
|
await drive(saveRun(run));
|
|
197
210
|
});
|
|
@@ -200,7 +213,7 @@ program
|
|
|
200
213
|
.description("標記 run 為失敗(執行中的 run 請直接在該終端機按 Ctrl-C)")
|
|
201
214
|
.action((id) => {
|
|
202
215
|
const run = mustGetRun(id);
|
|
203
|
-
saveRun({ ...run, stage: "failed", failedStage: run.stage, failureReason: "使用者取消" });
|
|
216
|
+
saveRun({ ...run, stage: "failed", failedStage: run.stage, failureCategory: "cancelled", failureReason: "使用者取消" });
|
|
204
217
|
console.log(`已取消 ${id}`);
|
|
205
218
|
});
|
|
206
219
|
program
|
|
@@ -266,6 +279,13 @@ program
|
|
|
266
279
|
for (const sub of subs)
|
|
267
280
|
console.log(` ${sub.at.slice(0, 16)} ${sub.step.padEnd(14)} ${sub.planned} → ${sub.actual}${sub.note ? `(${sub.note})` : ""}`);
|
|
268
281
|
}
|
|
282
|
+
const retries = listRetries(id);
|
|
283
|
+
if (retries.length) {
|
|
284
|
+
console.log("\n重試紀錄");
|
|
285
|
+
for (const r of retries) {
|
|
286
|
+
console.log(` ${r.key.padEnd(19)} ${retryLabel(r.category)}${r.final ? "(達上限)" : `(第 ${r.attempt} 次)`}`);
|
|
287
|
+
}
|
|
288
|
+
}
|
|
269
289
|
const tasks = readJsonFile(join(flowDir(id), "tasks.ordered.json"), TaskList);
|
|
270
290
|
if (!tasks.ok)
|
|
271
291
|
return;
|
|
@@ -502,6 +522,8 @@ async function doctor() {
|
|
|
502
522
|
console.log(`\n單一 run 的 agent 執行上限:${cfg.maxAgentRuns} 次`);
|
|
503
523
|
console.log(`修正策略:${cfg.fixStrategy} 測試與實作分開:${cfg.tddSplit ? "是" : "否"}`);
|
|
504
524
|
console.log(`程式碼審查人數:${cfg.reviewQuorum} 計畫審查人數:${cfg.planReviewQuorum} 計畫仲裁:${cfg.planArbiter ? "開啟" : "關閉"}`);
|
|
525
|
+
const layers = cfg.planReviewLayers;
|
|
526
|
+
console.log(`計畫分層審查:${layers.enabled ? `任務達 ${layers.minTasks} 個時開啟,最多 ${layers.maxGroups} 群,每群平均至少 ${layers.tasksPerGroup} 個任務` : "關閉"}`);
|
|
505
527
|
}
|
|
506
528
|
program.command("doctor").description("檢查可用的 agent CLI 與目前參與的 agent").action(doctor);
|
|
507
529
|
program
|
|
@@ -560,6 +582,105 @@ program
|
|
|
560
582
|
}
|
|
561
583
|
console.log(`\n次數多或失敗多的步驟可用 agentflowctl logs ${id} 找出編號查看原因;token 用量見 agentflowctl status ${id}`);
|
|
562
584
|
});
|
|
585
|
+
program
|
|
586
|
+
.command("insights")
|
|
587
|
+
.description("彙總這個專案所有 run 的結果、失敗原因、用量、步驟失敗與重試原因,並給出改善建議")
|
|
588
|
+
.action(() => {
|
|
589
|
+
const runs = listRuns();
|
|
590
|
+
if (!runs.length)
|
|
591
|
+
return console.log("還沒有 run");
|
|
592
|
+
const rows = runs.map((r) => ({
|
|
593
|
+
id: r.id,
|
|
594
|
+
stage: r.stage,
|
|
595
|
+
failureCategory: r.failureCategory,
|
|
596
|
+
retries: listRetries(r.id),
|
|
597
|
+
usage: listUsage(r.id),
|
|
598
|
+
substitutions: listSubstitutions(r.id).length,
|
|
599
|
+
stats: computeStats(listLogs(logDir(r.id))),
|
|
600
|
+
}));
|
|
601
|
+
const insights = computeInsights(rows);
|
|
602
|
+
const usage = computeUsageInsights(rows);
|
|
603
|
+
const o = insights.byOutcome;
|
|
604
|
+
console.log(`專案彙總(${insights.runCount} 個 run)`);
|
|
605
|
+
console.log(` 完成 ${o.done} 失敗 ${o.failed} 暫停 ${o.paused} 等待核准 ${o.awaiting_approval} 進行中 ${o.active}`);
|
|
606
|
+
if (insights.byFailure.length) {
|
|
607
|
+
console.log("\n失敗原因");
|
|
608
|
+
for (const row of insights.byFailure)
|
|
609
|
+
console.log(` ${padDisplay(failureLabel(row.category), 20)} ${String(row.count).padStart(4)} 個 run`);
|
|
610
|
+
}
|
|
611
|
+
const t = usage.total;
|
|
612
|
+
console.log("\n用量");
|
|
613
|
+
if (t.runs) {
|
|
614
|
+
const value = t.reportedRuns
|
|
615
|
+
? `輸入 ${t.inputTokens}${cacheNote(t)}、輸出 ${t.outputTokens}、合計 ${t.tokens} tokens`
|
|
616
|
+
: "未回報或回報狀態不明";
|
|
617
|
+
console.log(` ${value};${t.runs} 次(未回報 ${t.unreportedRuns}、舊紀錄不明 ${t.legacyRuns})`);
|
|
618
|
+
}
|
|
619
|
+
else
|
|
620
|
+
console.log(" 還沒有用量紀錄");
|
|
621
|
+
const byStrength = usage.byStrength;
|
|
622
|
+
const reportedTotal = Object.values(byStrength).reduce((sum, entry) => sum + entry.tokens, 0);
|
|
623
|
+
if (Object.keys(byStrength).length) {
|
|
624
|
+
console.log("\n模型強度用量(占比只計入明確回報)");
|
|
625
|
+
for (const strength of ["low", "medium", "high", "未知"]) {
|
|
626
|
+
const entry = byStrength[strength];
|
|
627
|
+
if (!entry)
|
|
628
|
+
continue;
|
|
629
|
+
const share = reportedTotal ? `${(entry.tokens / reportedTotal * 100).toFixed(1)}%` : "無法計算";
|
|
630
|
+
console.log(` ${strength}: ${entry.tokens} tokens,占比 ${share},呼叫 ${entry.runs} 次(未回報 ${entry.unreportedRuns}、舊紀錄不明 ${entry.legacyRuns})`);
|
|
631
|
+
}
|
|
632
|
+
console.log(` 高強度呼叫:${byStrength.high?.runs ?? 0} 次`);
|
|
633
|
+
}
|
|
634
|
+
const byTokens = (record, n) => Object.entries(record).sort((a, b) => b[1].tokens - a[1].tokens).slice(0, n);
|
|
635
|
+
printUsage("用量最高的階段", byTokens(usage.byStage, 8));
|
|
636
|
+
printUsage("各 agent 用量", byTokens(usage.byAgent, 8));
|
|
637
|
+
printUsage("各模型與步驟用量(任務步驟不分 task id,只加總明確回報)", byTokens(usage.byModelStage, 10));
|
|
638
|
+
const ss = usage.stepStats;
|
|
639
|
+
if (ss.steps.length) {
|
|
640
|
+
const totalMs = ss.agentMs + ss.cmdMs;
|
|
641
|
+
const share = (ms) => (totalMs ? `${(ms / totalMs * 100).toFixed(0)}%` : "-");
|
|
642
|
+
console.log("\n步驟執行與失敗(跨 run,任務步驟不分 task id,依失敗次數排序,不含總經過時間)");
|
|
643
|
+
console.log(` agent ${formatDuration(ss.agentMs).padStart(7)} ${share(ss.agentMs)}`);
|
|
644
|
+
console.log(` 專案指令 ${formatDuration(ss.cmdMs).padStart(7)} ${share(ss.cmdMs)}`);
|
|
645
|
+
if (ss.unfinished)
|
|
646
|
+
console.log(` 未完成 ${ss.unfinished} 份(沒有結束紀錄,不計入耗時)`);
|
|
647
|
+
console.log("\n 步驟 類型 次數 失敗 總耗時 最長");
|
|
648
|
+
for (const s of ss.steps.slice(0, 10)) {
|
|
649
|
+
console.log(` ${s.step.padEnd(19)} ${s.kind === "cmd" ? "指令 " : "agent"} ${String(s.runs).padStart(4)} ${String(s.failed).padStart(4)} ${formatDuration(s.totalMs).padStart(7)} ${formatDuration(s.maxMs).padStart(7)}${s.unfinished ? ` (未完成 ${s.unfinished})` : ""}`);
|
|
650
|
+
}
|
|
651
|
+
}
|
|
652
|
+
if (!insights.byCategory.length)
|
|
653
|
+
console.log("\n還沒有重試紀錄(舊 run 不會回填)");
|
|
654
|
+
else {
|
|
655
|
+
console.log("\n重試原因");
|
|
656
|
+
for (const row of insights.byCategory) {
|
|
657
|
+
const finals = row.finals ? `,其中 ${row.finals} 次達上限` : "";
|
|
658
|
+
console.log(` ${padDisplay(retryLabel(row.category), 20)} ${String(row.count).padStart(4)} 次${finals}`);
|
|
659
|
+
}
|
|
660
|
+
console.log("\n最常重試的關卡(任務關卡不分 task id 合併計算)");
|
|
661
|
+
for (const row of insights.byGate.slice(0, 10)) {
|
|
662
|
+
console.log(` ${padDisplay(row.gate, 19)} ${String(row.count).padStart(4)} 次`);
|
|
663
|
+
}
|
|
664
|
+
}
|
|
665
|
+
console.log("\n建議");
|
|
666
|
+
if (!usage.findings.length)
|
|
667
|
+
console.log(" 還沒有足夠訊號");
|
|
668
|
+
else
|
|
669
|
+
usage.findings.forEach((f, i) => {
|
|
670
|
+
console.log(` ${i + 1}. ${f.title}`);
|
|
671
|
+
console.log(` ${f.detail}`);
|
|
672
|
+
});
|
|
673
|
+
const usageById = new Map(usage.runs.map((r) => [r.id, r]));
|
|
674
|
+
console.log("\n各 run");
|
|
675
|
+
for (const r of insights.runs) {
|
|
676
|
+
const cat = r.topCategory ? ` 最多 ${retryLabel(r.topCategory)}` : "";
|
|
677
|
+
const u = usageById.get(r.id);
|
|
678
|
+
const tokens = !u?.calls ? "" : u.reportedCalls ? ` ${String(u.tokens).padStart(9)} tokens` : ` ${" ".repeat(3)}未回報${" ".repeat(7)}`;
|
|
679
|
+
const task = u?.topTask && u.topTask.tasks >= 2 ? ` 最耗任務 ${u.topTask.task}(${(u.topTask.share * 100).toFixed(0)}%)` : "";
|
|
680
|
+
console.log(` ${r.id} ${r.stage.padEnd(17)} 重試 ${String(r.retries).padStart(3)} 次${tokens}${task}${cat}`);
|
|
681
|
+
}
|
|
682
|
+
console.log("\n建議對應可改的 prompt、model stage 或關卡;各分組是同一批呼叫的不同切片,不要跨組相加。單一 run 用 agentflowctl status <id>、stats <id> 與 logs <id>");
|
|
683
|
+
});
|
|
563
684
|
program.parseAsync().catch((err) => {
|
|
564
685
|
console.error(`錯誤:${err.message}`);
|
|
565
686
|
process.exit(1);
|
package/dist/config.js
CHANGED
|
@@ -1,8 +1,10 @@
|
|
|
1
|
+
/** 重試上限的下限:低於這個值時,修正與審查來不及往返一輪 */
|
|
2
|
+
export const MIN_ATTEMPTS = 3;
|
|
1
3
|
export const config = {
|
|
2
4
|
/** 單一 Agent 執行最多幾輪工具迴圈,避免卡在迴圈裡 */
|
|
3
|
-
maxTurns: Number(process.env.AGENTFLOWCTL_MAX_TURNS ??
|
|
4
|
-
/**
|
|
5
|
-
maxAttempts: Number(process.env.AGENTFLOWCTL_MAX_ATTEMPTS ??
|
|
5
|
+
maxTurns: Number(process.env.AGENTFLOWCTL_MAX_TURNS ?? 200),
|
|
6
|
+
/** 同一個關卡連續失敗幾次後停止;環境變數低於 3 時以 3 計 */
|
|
7
|
+
maxAttempts: Math.max(MIN_ATTEMPTS, Number(process.env.AGENTFLOWCTL_MAX_ATTEMPTS ?? 5) || 5),
|
|
6
8
|
/** 終端機是否印出 agent 的文字、工具呼叫與專案指令;預設安靜,-v 或 AGENTFLOWCTL_VERBOSE=1 開啟 */
|
|
7
9
|
verbose: process.env.AGENTFLOWCTL_VERBOSE === "1",
|
|
8
10
|
};
|
package/dist/detect.js
CHANGED
|
@@ -24,6 +24,16 @@ const CHECK_SCRIPTS = {
|
|
|
24
24
|
test: ["test"],
|
|
25
25
|
build: ["build"],
|
|
26
26
|
};
|
|
27
|
+
/** 出現在依賴裡就代表專案有測試框架 */
|
|
28
|
+
const TEST_FRAMEWORKS = ["vitest", "jest", "mocha", "ava", "jasmine", "tap", "uvu", "@playwright/test", "cypress"];
|
|
29
|
+
function hasTestFramework(pkg) {
|
|
30
|
+
const deps = { ...pkg.dependencies, ...pkg.devDependencies };
|
|
31
|
+
if (TEST_FRAMEWORKS.some((name) => name in deps))
|
|
32
|
+
return true;
|
|
33
|
+
const script = pkg.scripts?.test;
|
|
34
|
+
// npm init 產生的佔位 script 不算
|
|
35
|
+
return typeof script === "string" && !/no test specified/i.test(script);
|
|
36
|
+
}
|
|
27
37
|
function readPackageJson(root) {
|
|
28
38
|
try {
|
|
29
39
|
const pkg = JSON.parse(readFileSync(join(root, "package.json"), "utf8"));
|
|
@@ -56,24 +66,31 @@ export function detectProjectDefaults(root) {
|
|
|
56
66
|
const script = CHECK_SCRIPTS[name]?.find((s) => typeof scripts[s] === "string");
|
|
57
67
|
return { name, cmd: script ? `${manager} run ${script}` : withExec(cmd, manager) };
|
|
58
68
|
});
|
|
59
|
-
return { manager, source, install: INSTALL[manager], test: withExec(defaults.test, manager), checks };
|
|
69
|
+
return { manager, source, install: INSTALL[manager], test: withExec(defaults.test, manager), checks, testFramework: hasTestFramework(pkg) };
|
|
60
70
|
}
|
|
71
|
+
/** 手動設定 test 指令就當作有測試框架 */
|
|
72
|
+
export const usesTestFramework = (raw, detected) => detected.testFramework || (!!raw && typeof raw === "object" && "test" in raw);
|
|
61
73
|
/** 補上原始設定裡沒寫的 install、test、checks;不是物件就原樣回傳,交給 schema 報錯 */
|
|
62
74
|
export function withProjectDefaults(raw, detected) {
|
|
63
75
|
if (!raw || typeof raw !== "object" || Array.isArray(raw))
|
|
64
76
|
return raw;
|
|
65
|
-
const { install, test
|
|
77
|
+
const { install, test } = detected;
|
|
78
|
+
// 沒有測試框架也沒手動設定 test 時,預設的檢查不含 test
|
|
79
|
+
const checks = usesTestFramework(raw, detected) ? detected.checks : detected.checks.filter((c) => c.name !== "test");
|
|
66
80
|
return { install, test, checks, ...raw };
|
|
67
81
|
}
|
|
68
82
|
/** run 開始時印出的說明:只列出這次用了偵測結果的欄位 */
|
|
69
83
|
export function describeDetected(raw, detected) {
|
|
70
84
|
const lines = [];
|
|
85
|
+
const framework = usesTestFramework(raw, detected);
|
|
71
86
|
if (!("install" in raw))
|
|
72
87
|
lines.push(` install:${detected.install}`);
|
|
73
|
-
if (!("test" in raw))
|
|
88
|
+
if (!("test" in raw) && framework)
|
|
74
89
|
lines.push(` test:${detected.test}`);
|
|
75
90
|
if (!("checks" in raw))
|
|
76
|
-
lines.push(...detected.checks.map((c) => ` checks.${c.name}:${c.cmd}`));
|
|
91
|
+
lines.push(...detected.checks.filter((c) => framework || c.name !== "test").map((c) => ` checks.${c.name}:${c.cmd}`));
|
|
92
|
+
if (!framework)
|
|
93
|
+
lines.push("ℹ️ 未偵測到測試框架,所有任務略過紅綠燈,也不跑 test 檢查(在 flow.config.json 設定 test 可改回來)");
|
|
77
94
|
if (!lines.length)
|
|
78
95
|
return [];
|
|
79
96
|
const why = detected.source === "預設" ? "預設" : `依 ${detected.source}`;
|