agentflowctl 0.13.1 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +10 -2
- package/dist/cli.js +114 -4
- package/dist/engine.js +291 -91
- package/dist/insights.js +85 -0
- package/dist/modelSelection.js +19 -11
- package/dist/paths.js +4 -0
- package/dist/planReview.js +333 -0
- package/dist/schemas.js +14 -0
- package/dist/stats.js +26 -0
- package/dist/store.js +62 -17
- package/dist/usageInsights.js +156 -0
- package/dist/util.js +12 -0
- package/examples/flow.config.json +1 -0
- package/package.json +1 -1
- package/prompts/plan-arbiter.md +1 -1
- package/prompts/plan-fix.md +5 -2
- package/prompts/plan-review-group.md +92 -0
- package/prompts/plan-review-index.md +93 -0
- package/prompts/plan-review.md +1 -0
- package/prompts/plan.md +3 -3
package/README.md
CHANGED
|
@@ -49,6 +49,7 @@ agentflowctl status f-xxxx # 看進度、結果與下一步
|
|
|
49
49
|
agentflowctl logs f-xxxx # 列出各步驟的 log
|
|
50
50
|
agentflowctl logs f-xxxx --latest # 看最新一份 log
|
|
51
51
|
agentflowctl stats f-xxxx # 各步驟耗時、執行與失敗次數
|
|
52
|
+
agentflowctl insights # 這個專案所有 run 的結果、失敗原因、用量、步驟失敗、重試原因與改善建議
|
|
52
53
|
agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
|
|
53
54
|
```
|
|
54
55
|
|
|
@@ -56,6 +57,8 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
|
|
|
56
57
|
|
|
57
58
|
`stats` 依 log 的開始與結束時間統計每個步驟的執行次數、失敗次數、總耗時與最長一次,並分開列出 agent 與專案指令(install、測試、checks)各占多少時間,最耗時的步驟排在最前面。沒有結束紀錄的 log 列為未完成,不計入耗時;總經過時間包含暫停與等待核准。
|
|
58
59
|
|
|
60
|
+
`insights` 把所有 run 分成獨立區塊彙總:最終狀態、失敗原因(重試達上限、仲裁停止、agent 次數用完等)、用量(輸入/輸出/cache、強度占比、階段/agent,各 run 用量最高的任務)、各模型與步驟、步驟執行與失敗(與 `stats` 相同取 log 的檔頭檔尾,跨 run 不計總經過時間)、關卡重試原因。任務關卡與任務步驟不分 task id 合併計算。最後列出最多五則建議,規則由程式套門檻,不是再請 agent 分析。合計 token 不是主指標。覆蓋不足時會先警告占比可能失真。各分組是同一批呼叫的不同切片,不要跨組相加。舊 run 沒有重試或失敗原因紀錄,不會回填。單一 run 的全量明細仍用 `status <id>` 與 `stats <id>`。
|
|
61
|
+
|
|
59
62
|
執行紀錄在 `.agentflowctl/runs/<id>/`,工作分支在 `.agentflowctl/worktrees/<id>/`。不再需要某次 run 時,可用 `agentflowctl clean <id>` 清除 worktree 與紀錄;`flow/<id>` 分支會保留。
|
|
60
63
|
|
|
61
64
|
## 執行停下來時怎麼做
|
|
@@ -67,7 +70,7 @@ agentflowctl resume f-xxxx # 從暫停、中斷或失敗處接續
|
|
|
67
70
|
| 按 Ctrl-C,或終端機意外關閉 | 執行 `agentflowctl resume <id>`;沒有結束紀錄的步驟會重跑 |
|
|
68
71
|
| `awaiting_approval`:計畫等你確認 | 閱讀 `.agentflowctl/worktrees/<id>/.flow/plan.md`,確認後執行 `agentflowctl approve <id>` |
|
|
69
72
|
| `paused`:agent 額度用完 | 等額度恢復後執行 `agentflowctl resume <id>`;審查步驟不會換 agent 代審 |
|
|
70
|
-
| `paused`:仲裁沒有產生有效裁決 | 依 `status` 的原因查看 log;若有 `.flow/plan-arbiter.json`,也檢查其內容,處理後執行 `agentflowctl resume <id
|
|
73
|
+
| `paused`:仲裁沒有產生有效裁決 | 依 `status` 的原因查看 log;若有 `.flow/plan-arbiter.json`,也檢查其內容,處理後執行 `agentflowctl resume <id>`。resume 會直接回到仲裁,不重跑計畫審查;暫停期間若改了計畫檔,或在 `flow.config.json` 把 `planArbiter` 關掉,就改成重新審查 |
|
|
71
74
|
| `failed`:測試、檢查、審查或 agent 執行失敗 | 依 `status` 提示查看失敗的 log,處理原因後執行 `agentflowctl resume <id>`;失敗階段會重試 |
|
|
72
75
|
| `failed`:已達 agent 執行次數上限 | 用 `agentflowctl resume <id> --max-agent-runs 100` 調高上限後接續,數字須大於已執行次數 |
|
|
73
76
|
|
|
@@ -187,12 +190,17 @@ Codex 另有幾點差異:
|
|
|
187
190
|
| `reviewQuorum` | `1` | 任務與最終程式碼審查需要幾位不同審查者核准 |
|
|
188
191
|
| `planReviewQuorum` | `1` | 計畫需要幾位不同審查者核准 |
|
|
189
192
|
| `planArbiter` | `true` | 計畫審查僵持時是否啟用仲裁 |
|
|
193
|
+
| `planReviewLayers` | `{ "enabled": true, "minTasks": 7, "maxGroups": 5, "tasksPerGroup": 3 }` | 任務夠多時把計畫審查拆成索引與任務群;說明見表格下方 |
|
|
190
194
|
| `tieBreak` | `"proceed"` | 兩位仲裁者意見分歧時,`"proceed"` 繼續、`"stop"` 停止 |
|
|
191
195
|
| `maxAgentRuns` | `60` | 一次 run 最多執行幾次 agent;可用指令選項覆蓋 |
|
|
192
196
|
| `install`、`test` | 依專案偵測 | 寫成指令字串,例如 `"install": "pnpm install"` |
|
|
193
197
|
| `checks` | 依專案偵測 | 檢查清單,例如 `[{ "name": "test", "cmd": "pnpm test" }]`;提供時會取代整份預設清單 |
|
|
194
198
|
| `testPattern` | 常見的 `.test.`、`.spec.` 檔名 | 辨識測試檔的正規表示式字串;非標準檔名時調整 |
|
|
195
199
|
|
|
200
|
+
計畫審查會依任務規模選做法。同時符合下列條件時,每輪先做一次索引審查,再只審查有變動的任務群:任務達到 `planReviewLayers.minTasks` 個;依 description 寫的檔案路徑能分成至少兩群,而且最大一群不超過三分之二;`plan.md` 每個任務都有 `## T-<數字>` 標題。索引審查讀規格、全部任務描述、驗收條件與整體做法,人數是 `planReviewQuorum`。群數最多 `maxGroups`,也不超過任務數除以 `tasksPerGroup`;每群一位審查者,含 `high` 任務的群改由 `planReviewQuorum` 位審查。改了 `plan.md` 的整體做法時所有群都重審;某一次審查失敗時只重跑還沒完成的部分。已達門檻卻不符其他條件時,終端機會印出原因並改由審查者讀完整份規格與計畫。`"planReviewLayers": { "enabled": false }` 可以關閉,`doctor` 會顯示目前的設定。
|
|
201
|
+
|
|
202
|
+
審查意見的處理寫在 `.flow/plan-replies.md`,每輪覆寫,不寫進 `plan.md` 文末。下一輪索引會看到整份回應;任務群只看到自己的 `## T-<數字>` 節。
|
|
203
|
+
|
|
196
204
|
`install`、`test`、`checks` 未設定時,會依 `packageManager`、lockfile 和 `package.json` scripts 偵測。完整範例見 [examples/flow.config.json](examples/flow.config.json)。專案設定每一步都會重新讀取,但已建立 run 的參與 agent 與執行次數上限會沿用建立時的值;要調高後者請用 `resume --max-agent-runs`。
|
|
197
205
|
|
|
198
206
|
### 環境變數
|
|
@@ -203,7 +211,7 @@ Codex 另有幾點差異:
|
|
|
203
211
|
| `AGENTFLOWCTL_VERBOSE` | 未開啟 | 設為 `1` 顯示詳細輸出,效果同 `-v` |
|
|
204
212
|
| `AGENTFLOWCTL_MAX_TURNS` | `80` | 目前程式會讀取此值,但尚未用它限制 agent 執行 |
|
|
205
213
|
|
|
206
|
-
環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限;`maxAgentRuns` 則是整次 run 的 agent
|
|
214
|
+
環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限;`maxAgentRuns` 則是整次 run 的 agent 執行次數上限。修正成功、或計畫審查與程式碼審查整組完成一輪有效審查後,該關的失敗次數會歸零,所以上限只計算連續失敗。分層計畫審查時,同一輪裡只要有一次審查呼叫真的執行成功,計畫審查的失敗次數也會歸零;所以索引與各群輪流各失敗一次、每次重跑都有進展時,不會因累計達上限而失敗。
|
|
207
215
|
|
|
208
216
|
## 更多文件
|
|
209
217
|
|
package/dist/cli.js
CHANGED
|
@@ -13,9 +13,11 @@ import { describeDetected, detectProjectDefaults } from "./detect.js";
|
|
|
13
13
|
import { CMD_AGENT, listLogs, localTime, logMark, nextLogFile, renderLog } from "./logs.js";
|
|
14
14
|
import { flowDir, logDir, projectRoot, worktreeDir } from "./paths.js";
|
|
15
15
|
import { ModelStage, ModelStrength, TaskList } from "./schemas.js";
|
|
16
|
+
import { computeInsights, failureLabel, retryLabel } from "./insights.js";
|
|
17
|
+
import { computeUsageInsights } from "./usageInsights.js";
|
|
16
18
|
import { computeStats, formatDuration } from "./stats.js";
|
|
17
|
-
import { agentRuns, getRun, listRuns, listSubstitutions, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
|
|
18
|
-
import { readJsonFile } from "./util.js";
|
|
19
|
+
import { agentRuns, getRun, listRetries, listRuns, listSubstitutions, listUsage, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
|
|
20
|
+
import { padDisplay, readJsonFile } from "./util.js";
|
|
19
21
|
import { openActions, readHandoff } from "./handoff.js";
|
|
20
22
|
import { stopReport } from "./stopReport.js";
|
|
21
23
|
import { runSetup, SETUP_ADAPTERS } from "./setup.js";
|
|
@@ -191,7 +193,7 @@ program
|
|
|
191
193
|
run = { ...run, stage: run.pausedStage ?? "spec", pausedStage: undefined, pauseReason: undefined };
|
|
192
194
|
}
|
|
193
195
|
if (run.stage === "failed") {
|
|
194
|
-
run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, modelRetryAttempts: {}, failedStage: undefined, failureReason: undefined };
|
|
196
|
+
run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, modelRetryAttempts: {}, failedStage: undefined, failureReason: undefined, failureCategory: undefined };
|
|
195
197
|
}
|
|
196
198
|
await drive(saveRun(run));
|
|
197
199
|
});
|
|
@@ -200,7 +202,7 @@ program
|
|
|
200
202
|
.description("標記 run 為失敗(執行中的 run 請直接在該終端機按 Ctrl-C)")
|
|
201
203
|
.action((id) => {
|
|
202
204
|
const run = mustGetRun(id);
|
|
203
|
-
saveRun({ ...run, stage: "failed", failedStage: run.stage, failureReason: "使用者取消" });
|
|
205
|
+
saveRun({ ...run, stage: "failed", failedStage: run.stage, failureCategory: "cancelled", failureReason: "使用者取消" });
|
|
204
206
|
console.log(`已取消 ${id}`);
|
|
205
207
|
});
|
|
206
208
|
program
|
|
@@ -266,6 +268,13 @@ program
|
|
|
266
268
|
for (const sub of subs)
|
|
267
269
|
console.log(` ${sub.at.slice(0, 16)} ${sub.step.padEnd(14)} ${sub.planned} → ${sub.actual}${sub.note ? `(${sub.note})` : ""}`);
|
|
268
270
|
}
|
|
271
|
+
const retries = listRetries(id);
|
|
272
|
+
if (retries.length) {
|
|
273
|
+
console.log("\n重試紀錄");
|
|
274
|
+
for (const r of retries) {
|
|
275
|
+
console.log(` ${r.key.padEnd(19)} ${retryLabel(r.category)}${r.final ? "(達上限)" : `(第 ${r.attempt} 次)`}`);
|
|
276
|
+
}
|
|
277
|
+
}
|
|
269
278
|
const tasks = readJsonFile(join(flowDir(id), "tasks.ordered.json"), TaskList);
|
|
270
279
|
if (!tasks.ok)
|
|
271
280
|
return;
|
|
@@ -502,6 +511,8 @@ async function doctor() {
|
|
|
502
511
|
console.log(`\n單一 run 的 agent 執行上限:${cfg.maxAgentRuns} 次`);
|
|
503
512
|
console.log(`修正策略:${cfg.fixStrategy} 測試與實作分開:${cfg.tddSplit ? "是" : "否"}`);
|
|
504
513
|
console.log(`程式碼審查人數:${cfg.reviewQuorum} 計畫審查人數:${cfg.planReviewQuorum} 計畫仲裁:${cfg.planArbiter ? "開啟" : "關閉"}`);
|
|
514
|
+
const layers = cfg.planReviewLayers;
|
|
515
|
+
console.log(`計畫分層審查:${layers.enabled ? `任務達 ${layers.minTasks} 個時開啟,最多 ${layers.maxGroups} 群,每群平均至少 ${layers.tasksPerGroup} 個任務` : "關閉"}`);
|
|
505
516
|
}
|
|
506
517
|
program.command("doctor").description("檢查可用的 agent CLI 與目前參與的 agent").action(doctor);
|
|
507
518
|
program
|
|
@@ -560,6 +571,105 @@ program
|
|
|
560
571
|
}
|
|
561
572
|
console.log(`\n次數多或失敗多的步驟可用 agentflowctl logs ${id} 找出編號查看原因;token 用量見 agentflowctl status ${id}`);
|
|
562
573
|
});
|
|
574
|
+
program
|
|
575
|
+
.command("insights")
|
|
576
|
+
.description("彙總這個專案所有 run 的結果、失敗原因、用量、步驟失敗與重試原因,並給出改善建議")
|
|
577
|
+
.action(() => {
|
|
578
|
+
const runs = listRuns();
|
|
579
|
+
if (!runs.length)
|
|
580
|
+
return console.log("還沒有 run");
|
|
581
|
+
const rows = runs.map((r) => ({
|
|
582
|
+
id: r.id,
|
|
583
|
+
stage: r.stage,
|
|
584
|
+
failureCategory: r.failureCategory,
|
|
585
|
+
retries: listRetries(r.id),
|
|
586
|
+
usage: listUsage(r.id),
|
|
587
|
+
substitutions: listSubstitutions(r.id).length,
|
|
588
|
+
stats: computeStats(listLogs(logDir(r.id))),
|
|
589
|
+
}));
|
|
590
|
+
const insights = computeInsights(rows);
|
|
591
|
+
const usage = computeUsageInsights(rows);
|
|
592
|
+
const o = insights.byOutcome;
|
|
593
|
+
console.log(`專案彙總(${insights.runCount} 個 run)`);
|
|
594
|
+
console.log(` 完成 ${o.done} 失敗 ${o.failed} 暫停 ${o.paused} 等待核准 ${o.awaiting_approval} 進行中 ${o.active}`);
|
|
595
|
+
if (insights.byFailure.length) {
|
|
596
|
+
console.log("\n失敗原因");
|
|
597
|
+
for (const row of insights.byFailure)
|
|
598
|
+
console.log(` ${padDisplay(failureLabel(row.category), 20)} ${String(row.count).padStart(4)} 個 run`);
|
|
599
|
+
}
|
|
600
|
+
const t = usage.total;
|
|
601
|
+
console.log("\n用量");
|
|
602
|
+
if (t.runs) {
|
|
603
|
+
const value = t.reportedRuns
|
|
604
|
+
? `輸入 ${t.inputTokens}${cacheNote(t)}、輸出 ${t.outputTokens}、合計 ${t.tokens} tokens`
|
|
605
|
+
: "未回報或回報狀態不明";
|
|
606
|
+
console.log(` ${value};${t.runs} 次(未回報 ${t.unreportedRuns}、舊紀錄不明 ${t.legacyRuns})`);
|
|
607
|
+
}
|
|
608
|
+
else
|
|
609
|
+
console.log(" 還沒有用量紀錄");
|
|
610
|
+
const byStrength = usage.byStrength;
|
|
611
|
+
const reportedTotal = Object.values(byStrength).reduce((sum, entry) => sum + entry.tokens, 0);
|
|
612
|
+
if (Object.keys(byStrength).length) {
|
|
613
|
+
console.log("\n模型強度用量(占比只計入明確回報)");
|
|
614
|
+
for (const strength of ["low", "medium", "high", "未知"]) {
|
|
615
|
+
const entry = byStrength[strength];
|
|
616
|
+
if (!entry)
|
|
617
|
+
continue;
|
|
618
|
+
const share = reportedTotal ? `${(entry.tokens / reportedTotal * 100).toFixed(1)}%` : "無法計算";
|
|
619
|
+
console.log(` ${strength}: ${entry.tokens} tokens,占比 ${share},呼叫 ${entry.runs} 次(未回報 ${entry.unreportedRuns}、舊紀錄不明 ${entry.legacyRuns})`);
|
|
620
|
+
}
|
|
621
|
+
console.log(` 高強度呼叫:${byStrength.high?.runs ?? 0} 次`);
|
|
622
|
+
}
|
|
623
|
+
const byTokens = (record, n) => Object.entries(record).sort((a, b) => b[1].tokens - a[1].tokens).slice(0, n);
|
|
624
|
+
printUsage("用量最高的階段", byTokens(usage.byStage, 8));
|
|
625
|
+
printUsage("各 agent 用量", byTokens(usage.byAgent, 8));
|
|
626
|
+
printUsage("各模型與步驟用量(任務步驟不分 task id,只加總明確回報)", byTokens(usage.byModelStage, 10));
|
|
627
|
+
const ss = usage.stepStats;
|
|
628
|
+
if (ss.steps.length) {
|
|
629
|
+
const totalMs = ss.agentMs + ss.cmdMs;
|
|
630
|
+
const share = (ms) => (totalMs ? `${(ms / totalMs * 100).toFixed(0)}%` : "-");
|
|
631
|
+
console.log("\n步驟執行與失敗(跨 run,任務步驟不分 task id,依失敗次數排序,不含總經過時間)");
|
|
632
|
+
console.log(` agent ${formatDuration(ss.agentMs).padStart(7)} ${share(ss.agentMs)}`);
|
|
633
|
+
console.log(` 專案指令 ${formatDuration(ss.cmdMs).padStart(7)} ${share(ss.cmdMs)}`);
|
|
634
|
+
if (ss.unfinished)
|
|
635
|
+
console.log(` 未完成 ${ss.unfinished} 份(沒有結束紀錄,不計入耗時)`);
|
|
636
|
+
console.log("\n 步驟 類型 次數 失敗 總耗時 最長");
|
|
637
|
+
for (const s of ss.steps.slice(0, 10)) {
|
|
638
|
+
console.log(` ${s.step.padEnd(19)} ${s.kind === "cmd" ? "指令 " : "agent"} ${String(s.runs).padStart(4)} ${String(s.failed).padStart(4)} ${formatDuration(s.totalMs).padStart(7)} ${formatDuration(s.maxMs).padStart(7)}${s.unfinished ? ` (未完成 ${s.unfinished})` : ""}`);
|
|
639
|
+
}
|
|
640
|
+
}
|
|
641
|
+
if (!insights.byCategory.length)
|
|
642
|
+
console.log("\n還沒有重試紀錄(舊 run 不會回填)");
|
|
643
|
+
else {
|
|
644
|
+
console.log("\n重試原因");
|
|
645
|
+
for (const row of insights.byCategory) {
|
|
646
|
+
const finals = row.finals ? `,其中 ${row.finals} 次達上限` : "";
|
|
647
|
+
console.log(` ${padDisplay(retryLabel(row.category), 20)} ${String(row.count).padStart(4)} 次${finals}`);
|
|
648
|
+
}
|
|
649
|
+
console.log("\n最常重試的關卡(任務關卡不分 task id 合併計算)");
|
|
650
|
+
for (const row of insights.byGate.slice(0, 10)) {
|
|
651
|
+
console.log(` ${padDisplay(row.gate, 19)} ${String(row.count).padStart(4)} 次`);
|
|
652
|
+
}
|
|
653
|
+
}
|
|
654
|
+
console.log("\n建議");
|
|
655
|
+
if (!usage.findings.length)
|
|
656
|
+
console.log(" 還沒有足夠訊號");
|
|
657
|
+
else
|
|
658
|
+
usage.findings.forEach((f, i) => {
|
|
659
|
+
console.log(` ${i + 1}. ${f.title}`);
|
|
660
|
+
console.log(` ${f.detail}`);
|
|
661
|
+
});
|
|
662
|
+
const usageById = new Map(usage.runs.map((r) => [r.id, r]));
|
|
663
|
+
console.log("\n各 run");
|
|
664
|
+
for (const r of insights.runs) {
|
|
665
|
+
const cat = r.topCategory ? ` 最多 ${retryLabel(r.topCategory)}` : "";
|
|
666
|
+
const u = usageById.get(r.id);
|
|
667
|
+
const tokens = !u?.calls ? "" : u.reportedCalls ? ` ${String(u.tokens).padStart(9)} tokens` : ` ${" ".repeat(3)}未回報${" ".repeat(7)}`;
|
|
668
|
+
const task = u?.topTask && u.topTask.tasks >= 2 ? ` 最耗任務 ${u.topTask.task}(${(u.topTask.share * 100).toFixed(0)}%)` : "";
|
|
669
|
+
console.log(` ${r.id} ${r.stage.padEnd(17)} 重試 ${String(r.retries).padStart(3)} 次${tokens}${task}${cat}`);
|
|
670
|
+
}
|
|
671
|
+
console.log("\n建議對應可改的 prompt、model stage 或關卡;各分組是同一批呼叫的不同切片,不要跨組相加。單一 run 用 agentflowctl status <id>、stats <id> 與 logs <id>");
|
|
672
|
+
});
|
|
563
673
|
program.parseAsync().catch((err) => {
|
|
564
674
|
console.error(`錯誤:${err.message}`);
|
|
565
675
|
process.exit(1);
|