agentflowctl 0.11.0 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +26 -1
- package/dist/agentConfig.js +4 -3
- package/dist/agents/claude.js +12 -6
- package/dist/agents/codex.js +3 -1
- package/dist/agents/command.js +4 -2
- package/dist/agents/gemini.js +17 -6
- package/dist/cli.js +141 -7
- package/dist/engine.js +35 -14
- package/dist/modelConfig.js +74 -0
- package/dist/modelProbe.js +74 -0
- package/dist/modelSelection.js +107 -0
- package/dist/proc.js +7 -2
- package/dist/runner.js +17 -4
- package/dist/schemas.js +17 -0
- package/dist/store.js +30 -10
- package/dist/tasks.js +7 -0
- package/package.json +1 -1
- package/prompts/plan-fix.md +1 -0
- package/prompts/plan-review.md +2 -0
- package/prompts/plan.md +7 -1
package/README.md
CHANGED
|
@@ -87,6 +87,7 @@ agentflowctl resume f-xxxx
|
|
|
87
87
|
| `run --req "..."` / `--req-file <檔案>` | 二選一,直接輸入需求或讀取檔案 |
|
|
88
88
|
| `run --manual-plan` | 計畫通過審查後等待你確認,再用 `approve <id>` 繼續 |
|
|
89
89
|
| `run --cycle <名單>` | 指定這次參與的 agent,例如 `--cycle claude,codex`;優先於設定檔的 `cycle` |
|
|
90
|
+
| `run --model-mode balanced\|adaptive` | 只覆蓋這次 run 的模型模式;`resume` 沿用建立時的模式 |
|
|
90
91
|
| `run --base <分支>` | 指定起始分支;未設定時使用目前分支 |
|
|
91
92
|
| `run --max-agent-runs <次數>` | 覆蓋這次的 `maxAgentRuns`;上限不夠時可用 `resume <id> --max-agent-runs <次數>` 調高 |
|
|
92
93
|
| `-v` / `--verbose` | 執行時顯示 agent 文字、工具呼叫與專案指令,適用於 `run`、`resume`、`approve` |
|
|
@@ -112,6 +113,28 @@ agentflowctl agent cycle claude,codex
|
|
|
112
113
|
|
|
113
114
|
`agent add` 的 `--adapter` 可填 `claude`、`codex`、`gemini` 或 `command`。`--model` 指定個別 agent 的模型;`--extra-arg=--參數` 可重複使用,傳給該 CLI。使用 `command` adapter 時,把指令寫在 `--` 後,例如 `agentflowctl agent add aider --adapter command -- aider --message {prompt}`。`agent remove <名稱>` 會移除設定與參與名單;`agent cycle` 不帶名單則顯示目前參與者。
|
|
114
115
|
|
|
116
|
+
`model add/set/remove` 只修改指定 agent 的模型清單;`model remove` 移除最後一個模型時,會檢查參與的 agent 是否仍有模型,不論目前使用哪種模型模式。同一 adapter 的 `agent set` 會保留清單;換 adapter 時會清掉舊 adapter 的模型設定。`agent setup` 遇到同名 agent 會先詢問是否覆寫。
|
|
117
|
+
|
|
118
|
+
### 依階段與任務難度選模型
|
|
119
|
+
|
|
120
|
+
預設是 `balanced`:沿用每個 agent 的 `model`、`defaultModels` 或 CLI 預設。想啟用自動選模,先為**每個參與的 agent** 登記可用模型與強度,再切換模式:
|
|
121
|
+
|
|
122
|
+
```bash
|
|
123
|
+
agentflowctl model add claude MODEL_NAME --strength low
|
|
124
|
+
agentflowctl model add claude ANOTHER_MODEL --strength high
|
|
125
|
+
agentflowctl model list
|
|
126
|
+
agentflowctl model mode adaptive
|
|
127
|
+
agentflowctl run --req-file ./requirement.md
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
把 `MODEL_NAME` 換成該 CLI 目前可呼叫的別名或完整 ID。`model add` 會用目前登入的帳號送出短請求,可能耗用少量 token;成功才寫入設定。需要重驗時執行 `model check`。用 `model set claude MODEL_NAME --strength medium` 改強度、`model remove claude MODEL_NAME` 移除模型,或用 `model stage taskReview high` 調整階段最低強度;`model stage` 不帶強度時列出各階段實際生效的強度。終端機每次呼叫會顯示送給 CLI 的模型名稱,Claude Code 與 Gemini CLI 回報的實際模型不同時也會顯示;`status <id>` 會按階段、任務、模型與步驟顯示用量。`run --model-mode balanced` 可暫時回到原設定。
|
|
131
|
+
|
|
132
|
+
`status <id>` 的用量以每次 LLM 呼叫為一筆,失敗、額度用完及代打也會計入呼叫次數。只有 CLI 同時回報輸入與輸出 token,才把兩者納入合計與模型強度占比;明確回報的 0 仍算已回報。缺少任一數字列為「未回報」;舊紀錄無法分辨真實 0 與預設補值,列為「舊紀錄不明」,原始數字只供查閱。各 agent、階段、任務、模型與步驟、模型強度是同一批呼叫的不同分組,不應跨組相加。`model add/check` 的探測請求可能耗用 token,但不屬於 run,因此不在 `status` 內。
|
|
133
|
+
|
|
134
|
+
計畫 agent 會查閱相關程式碼,依影響範圍、技術不確定性與失敗後果為每個任務標註 `low`/`medium`/`high` 難度,取三者中最高等級,並在計畫中寫出依據;計畫審查會逐項核對。自動選模先遵守角色分配,再取階段強度與任務難度中較高者;失敗重試會提高強度。若分配到的 agent 沒有足夠強度的模型,會選它最強的模型並提示。這些強度是你對模型能力的設定,不由 CLI 自動評分。
|
|
135
|
+
|
|
136
|
+
模型探測必須能禁止工具。這版 Codex CLI 沒有可確認的無工具探測參數,因此 `model add` 無法驗證並登記 Codex 模型;手動寫入 `models` 仍可執行,但不代表已驗證可用。`balanced` 仍可照原方式使用。自訂 `command` adapter 另需提供會實際呼叫模型的探測命令;設定方式與限制見[詳細參考](docs/reference.md#模型設定與自動選模)。
|
|
137
|
+
|
|
115
138
|
### 專案設定
|
|
116
139
|
|
|
117
140
|
你也可以直接編輯 `flow.config.json`。這是可用的最小範例;沒有寫的欄位會使用預設值:
|
|
@@ -132,6 +155,8 @@ agentflowctl agent cycle claude,codex
|
|
|
132
155
|
| `agents` | `{}` | 以名稱為 key 定義 agent;每個都要有 `adapter`,可加 `model`、`extraArgs`;`command` adapter 另需 `command` 指令陣列 |
|
|
133
156
|
| `cycle` | 自動偵測 | 填 agent 名稱陣列,例如 `["claude", "codex"]`;未填時使用已設定且可執行的 agent;順序不決定角色 |
|
|
134
157
|
| `defaultModels` | `{}` | 依 adapter 設預設模型,例如 `{ "claude": "模型名稱" }`;個別 agent 的 `model` 優先 |
|
|
158
|
+
| `modelSelection` | `balanced` | `mode` 可為 `balanced` 或 `adaptive`;`stageStrength` 可覆蓋各 LLM 階段強度 |
|
|
159
|
+
| `agents.<名稱>.models` | 無 | 自動選模時使用;每筆有 `name` 與 `strength`,建議用 `model add` 設定並實際驗證 |
|
|
135
160
|
| `fixStrategy` | `"ring"` | `"ring"` 由審查者以外的 agent 修正;`"author"` 交回最後作者 |
|
|
136
161
|
| `tddSplit` | `true` | 有多位 agent 時,`true` 會把同一任務的測試與實作分給不同 agent |
|
|
137
162
|
| `reviewQuorum` | `1` | 任務與最終程式碼審查需要幾位不同審查者核准 |
|
|
@@ -153,7 +178,7 @@ agentflowctl agent cycle claude,codex
|
|
|
153
178
|
| `AGENTFLOWCTL_VERBOSE` | 未開啟 | 設為 `1` 顯示詳細輸出,效果同 `-v` |
|
|
154
179
|
| `AGENTFLOWCTL_MAX_TURNS` | `80` | 目前程式會讀取此值,但尚未用它限制 agent 執行 |
|
|
155
180
|
|
|
156
|
-
環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限;`maxAgentRuns` 則是整次 run 的 agent
|
|
181
|
+
環境變數對新啟動的 agentflowctl 程序生效。`AGENTFLOWCTL_MAX_ATTEMPTS` 是單一關卡的重試上限;`maxAgentRuns` 則是整次 run 的 agent 執行次數上限。修正成功、或計畫審查與程式碼審查整組完成一輪有效審查後,該關的失敗次數會歸零,所以上限只計算連續失敗。
|
|
157
182
|
|
|
158
183
|
## 更多文件
|
|
159
184
|
|
package/dist/agentConfig.js
CHANGED
|
@@ -2,7 +2,7 @@ import { existsSync, readFileSync, writeFileSync } from "node:fs";
|
|
|
2
2
|
import { z } from "zod";
|
|
3
3
|
import { AgentDef, RepoConfig } from "./schemas.js";
|
|
4
4
|
/** 這些欄位的意義取決於 adapter,換 adapter 時要清掉 */
|
|
5
|
-
const ADAPTER_FIELDS = ["model", "extraArgs", "command"];
|
|
5
|
+
const ADAPTER_FIELDS = ["model", "models", "modelProbe", "extraArgs", "command"];
|
|
6
6
|
const agentsOf = (cfg) => ({ ...(cfg.agents ?? {}) });
|
|
7
7
|
const isDefined = (cfg, name) => name in agentsOf(cfg);
|
|
8
8
|
/** 只留下有值的欄位,驗證後回傳 */
|
|
@@ -33,13 +33,14 @@ export function setAgent(cfg, name, patch) {
|
|
|
33
33
|
if (!isDefined(cfg, name))
|
|
34
34
|
throw new Error(`未定義的 agent:${name}`);
|
|
35
35
|
if (Object.values(patch).every((v) => v === undefined)) {
|
|
36
|
-
throw new Error("沒有要修改的欄位(--adapter、--model、--extra-arg 或 -- <command>)");
|
|
36
|
+
throw new Error("沒有要修改的欄位(--adapter、--model、--model-probe-arg、--extra-arg 或 -- <command>)");
|
|
37
37
|
}
|
|
38
38
|
const agents = agentsOf(cfg);
|
|
39
39
|
const base = { ...agents[name] };
|
|
40
40
|
const changes = [];
|
|
41
41
|
if (patch.adapter !== undefined && patch.adapter !== base.adapter) {
|
|
42
|
-
const
|
|
42
|
+
const given = patch;
|
|
43
|
+
const cleared = ADAPTER_FIELDS.filter((f) => base[f] !== undefined && given[f] === undefined);
|
|
43
44
|
for (const f of cleared)
|
|
44
45
|
delete base[f];
|
|
45
46
|
if (cleared.length)
|
package/dist/agents/claude.js
CHANGED
|
@@ -24,6 +24,11 @@ function settingsFile(o) {
|
|
|
24
24
|
}
|
|
25
25
|
export const claude = {
|
|
26
26
|
probe: () => ({ cmd: "claude", args: ["--version"] }),
|
|
27
|
+
invokeModelProbe: (model) => ({
|
|
28
|
+
cmd: "claude",
|
|
29
|
+
args: ["-p", "只回答 OK", "--model", model, "--output-format", "stream-json", "--verbose",
|
|
30
|
+
"--tools", "", "--strict-mcp-config", "--disable-slash-commands"],
|
|
31
|
+
}),
|
|
27
32
|
invoke: (o) => ({
|
|
28
33
|
cmd: "claude",
|
|
29
34
|
args: [
|
|
@@ -40,7 +45,10 @@ export const claude = {
|
|
|
40
45
|
if (!ev)
|
|
41
46
|
return [];
|
|
42
47
|
const out = [];
|
|
43
|
-
if (ev.type === "
|
|
48
|
+
if (ev.type === "system" && ev.subtype === "init" && str(ev.model)) {
|
|
49
|
+
out.push({ kind: "model", id: str(ev.model) });
|
|
50
|
+
}
|
|
51
|
+
else if (ev.type === "assistant") {
|
|
44
52
|
const content = ev.message?.content ?? [];
|
|
45
53
|
for (const b of content) {
|
|
46
54
|
if (b.type === "text" && str(b.text)?.trim())
|
|
@@ -51,11 +59,9 @@ export const claude = {
|
|
|
51
59
|
}
|
|
52
60
|
else if (ev.type === "result") {
|
|
53
61
|
const usage = (ev.usage ?? {});
|
|
54
|
-
|
|
55
|
-
kind: "usage",
|
|
56
|
-
|
|
57
|
-
outputTokens: num(usage.output_tokens),
|
|
58
|
-
});
|
|
62
|
+
if (num(usage.input_tokens) !== undefined || num(usage.output_tokens) !== undefined) {
|
|
63
|
+
out.push({ kind: "usage", inputTokens: num(usage.input_tokens), outputTokens: num(usage.output_tokens) });
|
|
64
|
+
}
|
|
59
65
|
out.push({ kind: "done", ok: ev.is_error !== true, summary: str(ev.result) });
|
|
60
66
|
}
|
|
61
67
|
return out;
|
package/dist/agents/codex.js
CHANGED
|
@@ -30,7 +30,9 @@ export const codex = {
|
|
|
30
30
|
}
|
|
31
31
|
else if (ev.type === "turn.completed") {
|
|
32
32
|
const usage = (ev.usage ?? {});
|
|
33
|
-
|
|
33
|
+
if (num(usage.input_tokens) !== undefined || num(usage.output_tokens) !== undefined) {
|
|
34
|
+
out.push({ kind: "usage", inputTokens: num(usage.input_tokens), outputTokens: num(usage.output_tokens) });
|
|
35
|
+
}
|
|
34
36
|
}
|
|
35
37
|
else if (ev.type === "turn.failed" || ev.type === "error") {
|
|
36
38
|
const err = (ev.error ?? {});
|
package/dist/agents/command.js
CHANGED
|
@@ -10,9 +10,11 @@ export const command = {
|
|
|
10
10
|
if (!cmd)
|
|
11
11
|
throw new Error("command adapter 需要設定 command 陣列");
|
|
12
12
|
const hasPlaceholder = rest.some((a) => a.includes("{prompt}"));
|
|
13
|
+
if (o.command?.some((a) => a.includes("{model}")) && !o.model)
|
|
14
|
+
throw new Error("command 含 {model},但沒有設定 model");
|
|
13
15
|
return {
|
|
14
|
-
cmd,
|
|
15
|
-
args: [...rest.map((a) => a.replaceAll("{prompt}", o.prompt)), ...o.extraArgs],
|
|
16
|
+
cmd: cmd.replaceAll("{model}", o.model ?? ""),
|
|
17
|
+
args: [...rest.map((a) => a.replaceAll("{prompt}", o.prompt).replaceAll("{model}", o.model ?? "")), ...o.extraArgs],
|
|
16
18
|
input: hasPlaceholder ? undefined : o.prompt,
|
|
17
19
|
};
|
|
18
20
|
},
|
package/dist/agents/gemini.js
CHANGED
|
@@ -1,3 +1,5 @@
|
|
|
1
|
+
import { writeFileSync } from "node:fs";
|
|
2
|
+
import { join } from "node:path";
|
|
1
3
|
import { num, str, toolDetail, tryJson } from "./types.js";
|
|
2
4
|
/**
|
|
3
5
|
* Google Gemini CLI:`gemini -p ... --output-format stream-json`。
|
|
@@ -7,6 +9,13 @@ import { num, str, toolDetail, tryJson } from "./types.js";
|
|
|
7
9
|
*/
|
|
8
10
|
export const gemini = {
|
|
9
11
|
probe: () => ({ cmd: "gemini", args: ["--version"] }),
|
|
12
|
+
invokeModelProbe: (model, cwd) => {
|
|
13
|
+
const policy = join(cwd, "deny-tools.toml");
|
|
14
|
+
writeFileSync(policy, '[[rule]]\ntoolName = "*"\ndecision = "deny"\npriority = 10000\n');
|
|
15
|
+
return { cmd: "gemini", args: ["-p", "只回答 OK", "--model", model, "--output-format", "stream-json",
|
|
16
|
+
"--approval-mode", "default", "--extensions", "none", "--policy", policy],
|
|
17
|
+
env: { GEMINI_CLI_TRUST_WORKSPACE: "true" } };
|
|
18
|
+
},
|
|
10
19
|
invoke: (o) => ({
|
|
11
20
|
cmd: "gemini",
|
|
12
21
|
args: ["-p", o.prompt, "--output-format", "stream-json", "--approval-mode", "yolo", ...(o.model ? ["-m", o.model] : []), ...o.extraArgs],
|
|
@@ -17,7 +26,10 @@ export const gemini = {
|
|
|
17
26
|
if (!ev)
|
|
18
27
|
return [];
|
|
19
28
|
const out = [];
|
|
20
|
-
if (ev.type === "
|
|
29
|
+
if (ev.type === "init" && str(ev.model)) {
|
|
30
|
+
out.push({ kind: "model", id: str(ev.model) });
|
|
31
|
+
}
|
|
32
|
+
else if (ev.type === "message" && ev.role === "assistant" && str(ev.content)?.trim()) {
|
|
21
33
|
out.push({ kind: "text", text: str(ev.content) });
|
|
22
34
|
}
|
|
23
35
|
else if (ev.type === "tool_use") {
|
|
@@ -25,11 +37,10 @@ export const gemini = {
|
|
|
25
37
|
}
|
|
26
38
|
else if (ev.type === "result") {
|
|
27
39
|
const stats = (ev.stats ?? {});
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
});
|
|
40
|
+
const inputTokens = num(stats.input_tokens) ?? num(stats.inputTokens);
|
|
41
|
+
const outputTokens = num(stats.output_tokens) ?? num(stats.outputTokens);
|
|
42
|
+
if (inputTokens !== undefined || outputTokens !== undefined)
|
|
43
|
+
out.push({ kind: "usage", inputTokens, outputTokens });
|
|
33
44
|
out.push({ kind: "done", ok: ev.status !== "error", summary: str(ev.response) });
|
|
34
45
|
}
|
|
35
46
|
else if (ev.type === "error") {
|
package/dist/cli.js
CHANGED
|
@@ -12,13 +12,25 @@ import { cleanableRuns, cleanRun } from "./cleanup.js";
|
|
|
12
12
|
import { describeDetected, detectProjectDefaults } from "./detect.js";
|
|
13
13
|
import { CMD_AGENT, listLogs, localTime, logMark, nextLogFile, renderLog } from "./logs.js";
|
|
14
14
|
import { flowDir, logDir, projectRoot, worktreeDir } from "./paths.js";
|
|
15
|
-
import { TaskList } from "./schemas.js";
|
|
16
|
-
import { agentRuns, getRun, listRuns, listSubstitutions, saveRun, usageByAgent } from "./store.js";
|
|
15
|
+
import { ModelStage, ModelStrength, TaskList } from "./schemas.js";
|
|
16
|
+
import { agentRuns, getRun, listRuns, listSubstitutions, saveRun, usageByAgent, usageByModelStage, usageByStage, usageByStrength, usageByTask } from "./store.js";
|
|
17
17
|
import { readJsonFile } from "./util.js";
|
|
18
18
|
import { openActions, readHandoff } from "./handoff.js";
|
|
19
19
|
import { stopReport } from "./stopReport.js";
|
|
20
20
|
import { runSetup, SETUP_ADAPTERS } from "./setup.js";
|
|
21
21
|
import { addAgent, readRawConfig, removeAgent, setAgent, setCycle, writeRawConfig } from "./agentConfig.js";
|
|
22
|
+
import { addModel, removeModel, setModelMode, setModelStrength, setStageStrength } from "./modelConfig.js";
|
|
23
|
+
import { DEFAULT_STAGE_STRENGTH, effectiveStageStrengths, validateAdaptiveConfig } from "./modelSelection.js";
|
|
24
|
+
import { probeModel } from "./modelProbe.js";
|
|
25
|
+
function printUsage(title, rows) {
|
|
26
|
+
if (!rows.length)
|
|
27
|
+
return;
|
|
28
|
+
console.log(`\n${title}`);
|
|
29
|
+
for (const [key, c] of rows) {
|
|
30
|
+
const value = c.reportedRuns ? `輸入 ${c.inputTokens}、輸出 ${c.outputTokens}、合計 ${c.tokens} tokens` : "未回報或回報狀態不明";
|
|
31
|
+
console.log(` ${key}: ${value};${c.runs} 次(未回報 ${c.unreportedRuns}、舊紀錄不明 ${c.legacyRuns})`);
|
|
32
|
+
}
|
|
33
|
+
}
|
|
22
34
|
function mustGetRun(id) {
|
|
23
35
|
const run = getRun(id);
|
|
24
36
|
if (!run)
|
|
@@ -107,6 +119,7 @@ program
|
|
|
107
119
|
.option("--max-agent-runs <n>", "單一 run 最多執行幾次 agent(預設取 flow.config.json 的 maxAgentRuns)")
|
|
108
120
|
.option("--manual-plan", "計畫通過 AI 審查後,仍停下來等你確認", false)
|
|
109
121
|
.option("--cycle <agents>", "參與的 agent,例如 claude,codex,gemini(順序不影響分工)")
|
|
122
|
+
.option("--model-mode <mode>", "這次 run 的模型模式:balanced 或 adaptive")
|
|
110
123
|
.action(async (opts) => {
|
|
111
124
|
const requirement = opts.reqFile ? readFileSync(opts.reqFile, "utf8") : opts.req;
|
|
112
125
|
if (!requirement?.trim())
|
|
@@ -116,12 +129,17 @@ program
|
|
|
116
129
|
if (!base)
|
|
117
130
|
throw new Error("目前不在任何分支上,請用 --base 指定基底分支");
|
|
118
131
|
const cycle = await resolveCycle(opts.cycle);
|
|
132
|
+
const cfg = loadRepoConfig();
|
|
133
|
+
const modelMode = opts.modelMode ?? cfg.modelSelection.mode;
|
|
134
|
+
if (modelMode !== "balanced" && modelMode !== "adaptive")
|
|
135
|
+
throw new Error(`未知的模型模式:${modelMode}`);
|
|
136
|
+
if (modelMode === "adaptive")
|
|
137
|
+
validateAdaptiveConfig(cfg, cycle);
|
|
119
138
|
const id = `f-${Date.now().toString(36)}`;
|
|
120
139
|
const branch = `flow/${id}`;
|
|
121
140
|
console.log(`[${id}] 🌿 從 ${base} 建立 worktree(分支 ${branch})`);
|
|
122
141
|
await addWorktree(root, worktreeDir(id), base, branch);
|
|
123
142
|
console.log(`[${id}] 🤝 參與的 agent:${cycle.join("、")}(角色隨機分配)`);
|
|
124
|
-
const cfg = loadRepoConfig();
|
|
125
143
|
const now = new Date().toISOString();
|
|
126
144
|
// worktree 一建好就寫入紀錄:之後在任何地方中斷,都能用 resume 接續或用 clean 清掉
|
|
127
145
|
const run = saveRun({
|
|
@@ -134,6 +152,7 @@ program
|
|
|
134
152
|
maxAgentRuns: opts.maxAgentRuns ? Number(opts.maxAgentRuns) : cfg.maxAgentRuns,
|
|
135
153
|
cycle,
|
|
136
154
|
attempts: {},
|
|
155
|
+
modelMode,
|
|
137
156
|
taskIndex: 0,
|
|
138
157
|
taskPhase: "tests",
|
|
139
158
|
createdAt: now,
|
|
@@ -162,13 +181,15 @@ program
|
|
|
162
181
|
.option("--max-agent-runs <n>", "調整 agent 執行次數上限")
|
|
163
182
|
.action(async (id, opts) => {
|
|
164
183
|
let run = mustGetRun(id);
|
|
184
|
+
if (run.modelMode === "adaptive")
|
|
185
|
+
validateAdaptiveConfig(loadRepoConfig(), run.cycle);
|
|
165
186
|
if (opts.maxAgentRuns)
|
|
166
187
|
run = { ...run, maxAgentRuns: Number(opts.maxAgentRuns) };
|
|
167
188
|
if (run.stage === "paused") {
|
|
168
189
|
run = { ...run, stage: run.pausedStage ?? "spec", pausedStage: undefined, pauseReason: undefined };
|
|
169
190
|
}
|
|
170
191
|
if (run.stage === "failed") {
|
|
171
|
-
run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, failedStage: undefined, failureReason: undefined };
|
|
192
|
+
run = { ...run, stage: run.failedStage ?? "spec", attempts: {}, modelRetryAttempts: {}, failedStage: undefined, failureReason: undefined };
|
|
172
193
|
}
|
|
173
194
|
await drive(saveRun(run));
|
|
174
195
|
});
|
|
@@ -215,8 +236,27 @@ program
|
|
|
215
236
|
if (Object.keys(byAgent).length) {
|
|
216
237
|
console.log("\n各 agent 用量");
|
|
217
238
|
for (const [agent, c] of Object.entries(byAgent)) {
|
|
218
|
-
console.log(` ${agent.padEnd(10)} ${String(c.runs).padStart(3)} 次 ${String(c.tokens).padStart(9)} tokens`);
|
|
239
|
+
console.log(` ${agent.padEnd(10)} ${String(c.runs).padStart(3)} 次 ${String(c.tokens).padStart(9)} 已回報 tokens(未回報 ${c.unreportedRuns}、舊紀錄不明 ${c.legacyRuns}${c.legacyTokens ? `,原始數字 ${c.legacyTokens} tokens` : ""})`);
|
|
240
|
+
}
|
|
241
|
+
}
|
|
242
|
+
const byStage = usageByStage(id);
|
|
243
|
+
const stageOrder = [...Object.keys(DEFAULT_STAGE_STRENGTH), "其他"];
|
|
244
|
+
printUsage("各階段用量(同一步驟的所有任務合計)", Object.entries(byStage).sort(([a], [b]) => stageOrder.indexOf(a) - stageOrder.indexOf(b)));
|
|
245
|
+
const taskNo = (key) => /^T-(\d+)$/.exec(key) ? Number(key.slice(2)) : Infinity;
|
|
246
|
+
printUsage("各任務用量(寫測試、實作、任務審查、任務修正)", Object.entries(usageByTask(id)).sort(([a], [b]) => taskNo(a) - taskNo(b)));
|
|
247
|
+
printUsage("各模型與步驟用量(只加總明確回報)", Object.entries(usageByModelStage(id)));
|
|
248
|
+
const byStrength = usageByStrength(id);
|
|
249
|
+
const reportedTotal = Object.values(byStrength).reduce((sum, entry) => sum + entry.tokens, 0);
|
|
250
|
+
if (Object.keys(byStrength).length) {
|
|
251
|
+
console.log("\n模型強度用量(占比只計入明確回報)");
|
|
252
|
+
for (const strength of ["low", "medium", "high", "未知"]) {
|
|
253
|
+
const entry = byStrength[strength];
|
|
254
|
+
if (!entry)
|
|
255
|
+
continue;
|
|
256
|
+
const share = reportedTotal ? `${(entry.tokens / reportedTotal * 100).toFixed(1)}%` : "無法計算";
|
|
257
|
+
console.log(` ${strength}: ${entry.tokens} tokens,占比 ${share},呼叫 ${entry.runs} 次(未回報 ${entry.unreportedRuns}、舊紀錄不明 ${entry.legacyRuns})`);
|
|
219
258
|
}
|
|
259
|
+
console.log(` 高強度呼叫:${byStrength.high?.runs ?? 0} 次`);
|
|
220
260
|
}
|
|
221
261
|
const subs = listSubstitutions(id);
|
|
222
262
|
if (subs.length) {
|
|
@@ -237,6 +277,18 @@ program
|
|
|
237
277
|
// ───────────── agent 管理:讀寫 flow.config.json 的 agents 與 cycle ─────────────
|
|
238
278
|
const configPath = () => join(projectRoot(), "flow.config.json");
|
|
239
279
|
const collect = (value, prev = []) => [...prev, value];
|
|
280
|
+
const parseStrength = (value) => {
|
|
281
|
+
const result = ModelStrength.safeParse(value);
|
|
282
|
+
if (!result.success)
|
|
283
|
+
throw new Error(`未知的模型強度:${value}(請使用 low、medium 或 high)`);
|
|
284
|
+
return result.data;
|
|
285
|
+
};
|
|
286
|
+
const parseModelStage = (value) => {
|
|
287
|
+
const result = ModelStage.safeParse(value);
|
|
288
|
+
if (!result.success)
|
|
289
|
+
throw new Error(`未知的 LLM 階段:${value}`);
|
|
290
|
+
return result.data;
|
|
291
|
+
};
|
|
240
292
|
function applyEdit(edit, done) {
|
|
241
293
|
const before = readRawConfig(configPath());
|
|
242
294
|
const { cfg, changes } = edit(before);
|
|
@@ -279,8 +331,9 @@ agent
|
|
|
279
331
|
.requiredOption("--adapter <adapter>", "claude、codex、gemini 或 command")
|
|
280
332
|
.option("--model <model>", "模型名稱")
|
|
281
333
|
.option("--extra-arg <arg>", "額外參數,可重複;以 - 開頭時寫成 --extra-arg=--sandbox", collect)
|
|
334
|
+
.option("--model-probe-arg <arg>", "自訂 command 的模型探測命令參數,可重複", collect)
|
|
282
335
|
.action((name, command, opts) => {
|
|
283
|
-
applyEdit((cfg) => addAgent(cfg, name, { adapter: opts.adapter, model: opts.model, extraArgs: opts.extraArg, command: command.length ? command : undefined }), `已新增 ${name};要讓它參與請用 agent cycle`);
|
|
336
|
+
applyEdit((cfg) => addAgent(cfg, name, { adapter: opts.adapter, model: opts.model, extraArgs: opts.extraArg, modelProbe: opts.modelProbeArg, command: command.length ? command : undefined }), `已新增 ${name};要讓它參與請用 agent cycle`);
|
|
284
337
|
});
|
|
285
338
|
agent
|
|
286
339
|
.command("set <name> [command...]")
|
|
@@ -288,8 +341,9 @@ agent
|
|
|
288
341
|
.option("--adapter <adapter>", "claude、codex、gemini 或 command")
|
|
289
342
|
.option("--model <model>", "模型名稱")
|
|
290
343
|
.option("--extra-arg <arg>", "額外參數,可重複,會整個取代原本的設定", collect)
|
|
344
|
+
.option("--model-probe-arg <arg>", "自訂 command 的模型探測命令參數,可重複,會整個取代", collect)
|
|
291
345
|
.action((name, command, opts) => {
|
|
292
|
-
applyEdit((cfg) => setAgent(cfg, name, { adapter: opts.adapter, model: opts.model, extraArgs: opts.extraArg, command: command.length ? command : undefined }), `已更新 ${name}`);
|
|
346
|
+
applyEdit((cfg) => setAgent(cfg, name, { adapter: opts.adapter, model: opts.model, extraArgs: opts.extraArg, modelProbe: opts.modelProbeArg, command: command.length ? command : undefined }), `已更新 ${name}`);
|
|
293
347
|
});
|
|
294
348
|
agent
|
|
295
349
|
.command("remove <name>")
|
|
@@ -346,6 +400,86 @@ agent
|
|
|
346
400
|
console.log("");
|
|
347
401
|
await doctor();
|
|
348
402
|
});
|
|
403
|
+
// ───────────── 模型清單與各階段強度 ─────────────
|
|
404
|
+
const model = program.command("model").description("設定模型強度並檢查目前帳號能否呼叫模型");
|
|
405
|
+
model.command("add <agent> <name>")
|
|
406
|
+
.requiredOption("--strength <strength>", "low、medium 或 high")
|
|
407
|
+
.action(async (agent, name, opts) => {
|
|
408
|
+
const strength = parseStrength(opts.strength);
|
|
409
|
+
const before = readRawConfig(configPath());
|
|
410
|
+
const cfg = loadRepoConfig();
|
|
411
|
+
const def = cfg.agents[agent];
|
|
412
|
+
if (!def)
|
|
413
|
+
throw new Error(`未定義的 agent:${agent}`);
|
|
414
|
+
const next = addModel(before, agent, name, strength);
|
|
415
|
+
console.log("模型檢查會送出最短請求,可能耗用少量 token。");
|
|
416
|
+
const checked = await probeModel(def, name);
|
|
417
|
+
if (checked.status !== "ok")
|
|
418
|
+
throw new Error(`${agent} 的模型 ${name} 未加入:${checked.reason}`);
|
|
419
|
+
writeRawConfig(configPath(), next);
|
|
420
|
+
console.log(`✅ 已新增 ${agent} 的模型 ${name}(${strength})${checked.resolvedModel ? ` → ${checked.resolvedModel}` : ""}${def.adapter === "command" ? "(由自訂探測命令回報)" : ""};探測可能耗用少量 token`);
|
|
421
|
+
});
|
|
422
|
+
model.command("set <agent> <name>")
|
|
423
|
+
.requiredOption("--strength <strength>", "low、medium 或 high")
|
|
424
|
+
.action((agent, name, opts) => {
|
|
425
|
+
writeRawConfig(configPath(), setModelStrength(readRawConfig(configPath()), agent, name, parseStrength(opts.strength)));
|
|
426
|
+
console.log(`✅ 已將 ${agent} 的模型 ${name} 強度設為 ${opts.strength};此指令不重驗可用性`);
|
|
427
|
+
});
|
|
428
|
+
model.command("remove <agent> <name>").action((agent, name) => {
|
|
429
|
+
writeRawConfig(configPath(), removeModel(readRawConfig(configPath()), agent, name));
|
|
430
|
+
console.log(`✅ 已移除 ${agent} 的模型 ${name}`);
|
|
431
|
+
});
|
|
432
|
+
model.command("list [agent]").action((agent) => {
|
|
433
|
+
const cfg = loadRepoConfig();
|
|
434
|
+
const names = agent ? [agent] : Object.keys(cfg.agents);
|
|
435
|
+
for (const name of names) {
|
|
436
|
+
const def = cfg.agents[name];
|
|
437
|
+
if (!def)
|
|
438
|
+
throw new Error(`未定義的 agent:${name}`);
|
|
439
|
+
console.log(`${name}(${def.adapter})`);
|
|
440
|
+
for (const item of def.models ?? [])
|
|
441
|
+
console.log(` ${item.name} ${item.strength}`);
|
|
442
|
+
if (!def.models?.length)
|
|
443
|
+
console.log(" 尚未設定模型");
|
|
444
|
+
}
|
|
445
|
+
console.log("清單只顯示設定內容;要確認當前帳號是否可用,請執行 model check。");
|
|
446
|
+
});
|
|
447
|
+
model.command("check [agent]").action(async (agent) => {
|
|
448
|
+
const cfg = loadRepoConfig();
|
|
449
|
+
const names = agent ? [agent] : Object.keys(cfg.agents);
|
|
450
|
+
console.log("模型重驗會逐一送出短請求,可能耗用少量 token。");
|
|
451
|
+
for (const name of names) {
|
|
452
|
+
const def = cfg.agents[name];
|
|
453
|
+
if (!def)
|
|
454
|
+
throw new Error(`未定義的 agent:${name}`);
|
|
455
|
+
for (const item of def.models ?? []) {
|
|
456
|
+
const checked = await probeModel(def, item.name);
|
|
457
|
+
const at = new Date().toISOString();
|
|
458
|
+
const mark = checked.status === "ok" ? "✅" : checked.status === "unverifiable" ? "⚪" : "❌";
|
|
459
|
+
const result = checked.status === "ok" ? `可呼叫${checked.resolvedModel ? ` → ${checked.resolvedModel}` : ""}${def.adapter === "command" ? "(由自訂探測命令回報)" : ""}` : checked.reason;
|
|
460
|
+
console.log(`${mark} ${at} ${name} ${item.name}: ${result}`);
|
|
461
|
+
}
|
|
462
|
+
}
|
|
463
|
+
});
|
|
464
|
+
model.command("mode [mode]").action((mode) => {
|
|
465
|
+
if (!mode)
|
|
466
|
+
return console.log(`模型模式:${loadRepoConfig().modelSelection.mode}`);
|
|
467
|
+
if (mode !== "balanced" && mode !== "adaptive")
|
|
468
|
+
throw new Error(`未知的模型模式:${mode}`);
|
|
469
|
+
writeRawConfig(configPath(), setModelMode(readRawConfig(configPath()), mode));
|
|
470
|
+
console.log(`✅ 模型模式已設為 ${mode}`);
|
|
471
|
+
});
|
|
472
|
+
model.command("stage [stage] [strength]").action((stage, strength) => {
|
|
473
|
+
if (!stage || !strength) {
|
|
474
|
+
const all = effectiveStageStrengths(loadRepoConfig());
|
|
475
|
+
const stages = stage ? [parseModelStage(stage)] : Object.keys(all);
|
|
476
|
+
for (const s of stages)
|
|
477
|
+
console.log(`${s}: ${all[s].strength}${all[s].custom ? "(自訂)" : "(預設)"}`);
|
|
478
|
+
return;
|
|
479
|
+
}
|
|
480
|
+
writeRawConfig(configPath(), setStageStrength(readRawConfig(configPath()), parseModelStage(stage), parseStrength(strength)));
|
|
481
|
+
console.log(`✅ ${stage} 的最低強度已設為 ${strength}`);
|
|
482
|
+
});
|
|
349
483
|
/** 檢查設定的 agent 是否已安裝,印出參與的 agent 與主要設定 */
|
|
350
484
|
async function doctor() {
|
|
351
485
|
const cfg = loadRepoConfig();
|
package/dist/engine.js
CHANGED
|
@@ -14,7 +14,8 @@ import { arbiterPanel, availableAgent, fixAgent, planAgent, planFixAgent, review
|
|
|
14
14
|
import { resolveAgent, runAgent, runCommand } from "./runner.js";
|
|
15
15
|
import { AcceptanceList, ArbiterResult, ConsistentReviewResult, RepoConfig, TaskList, } from "./schemas.js";
|
|
16
16
|
import { addSubstitution, addUsage, agentRuns, saveRun } from "./store.js";
|
|
17
|
-
import {
|
|
17
|
+
import { clearModelReviewFailure, clearModelReviewStage, recordModelReviewFailure, selectModel } from "./modelSelection.js";
|
|
18
|
+
import { orderTasks, taskAcceptance, validateTaskComplexity } from "./tasks.js";
|
|
18
19
|
import { readJsonFile, renderPrompt, tail } from "./util.js";
|
|
19
20
|
// ───────────────────────── 共用工具 ─────────────────────────
|
|
20
21
|
const info = (run, msg) => console.log(`[${run.id}] ${msg}`);
|
|
@@ -60,10 +61,16 @@ async function agentStep(run, planned, step, prompt, mode) {
|
|
|
60
61
|
info(run, `🔁 ${agent} 額度已用完,${step} 由 ${sub} 代打${note ? `(注意:${note})` : ""}`);
|
|
61
62
|
agent = sub;
|
|
62
63
|
}
|
|
64
|
+
const selected = selectModel(run, cfg, agent, step, /^T-\d+-/.test(step) ? loadOrderedTasks(run)[run.taskIndex]?.complexity : undefined, mode.kind === "review" ? agent : undefined);
|
|
65
|
+
info(run, `🤖 ${step}:${agent} 使用 ${selected.name ?? "CLI 預設(名稱未知)"}${selected.insufficient ? `(低於目標 ${selected.targetStrength})` : ""}`);
|
|
63
66
|
const callKey = handoffKey(run, step, mode.slot ?? 0, agent);
|
|
64
67
|
prepareHandoff(run.id, callKey, handoffTarget(run), mode.blind ?? false);
|
|
65
|
-
const r = await runAgent(agent, resolveAgent(cfg, agent), target(run, step, agent), prompt);
|
|
66
|
-
|
|
68
|
+
const r = await runAgent(agent, { ...resolveAgent(cfg, agent), model: selected.name }, { ...target(run, step, agent), strength: selected.strength, targetStrength: selected.targetStrength }, prompt);
|
|
69
|
+
if (r.resolvedModel && r.resolvedModel !== selected.name)
|
|
70
|
+
info(run, ` ↳ CLI 回報實際模型:${r.resolvedModel}`);
|
|
71
|
+
addUsage(run.id, { stage: step, agent, model: selected.name, resolvedModel: r.resolvedModel,
|
|
72
|
+
strength: selected.strength, targetStrength: selected.targetStrength, usageReported: r.usageReported,
|
|
73
|
+
inputTokens: r.inputTokens, outputTokens: r.outputTokens });
|
|
67
74
|
if (!r.quotaExhausted) {
|
|
68
75
|
reportMeta(run, agent, r);
|
|
69
76
|
return { r, agent, step, callKey };
|
|
@@ -205,6 +212,9 @@ function validatePlan(run) {
|
|
|
205
212
|
const tasks = readJsonFile(flowFile(run, "tasks.json"), TaskList);
|
|
206
213
|
if (!tasks.ok)
|
|
207
214
|
return tasks.error;
|
|
215
|
+
const complexityError = validateTaskComplexity(tasks.data, run.modelMode ?? "balanced");
|
|
216
|
+
if (complexityError)
|
|
217
|
+
return complexityError;
|
|
208
218
|
return orderTasks(tasks.data, new Set(ids));
|
|
209
219
|
}
|
|
210
220
|
function acceptPlan(run, ordered) {
|
|
@@ -272,13 +282,14 @@ async function planReviewStage(run) {
|
|
|
272
282
|
if (tampered.length)
|
|
273
283
|
info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
|
|
274
284
|
if (!r.ok)
|
|
275
|
-
return retry(run, "plan-review-run", `Agent 執行失敗:${r.summary}`, "plan_review");
|
|
285
|
+
return retry(recordModelReviewFailure(run, "plan-review", reviewer), "plan-review-run", `Agent 執行失敗:${r.summary}`, "plan_review");
|
|
276
286
|
const review = readJsonFile(flowFile(run, "plan-review.json"), ConsistentReviewResult);
|
|
277
287
|
if (!review.ok)
|
|
278
|
-
return retry(run, "plan-review-run", review.error, "plan_review");
|
|
288
|
+
return retry(recordModelReviewFailure(run, "plan-review", reviewer), "plan-review-run", review.error, "plan_review");
|
|
279
289
|
const handoffError = finishHandoff(run, outcome, "reviewer", { target: "plan", verdict: review.data.verdict });
|
|
280
290
|
if (handoffError)
|
|
281
|
-
return retry(run, "plan-review-run", handoffError, "plan_review");
|
|
291
|
+
return retry(recordModelReviewFailure(run, "plan-review", reviewer), "plan-review-run", handoffError, "plan_review");
|
|
292
|
+
run = clearModelReviewFailure(run, "plan-review", reviewer);
|
|
282
293
|
// 審查紀錄移到 worktree 外面:之後的仲裁者看不到是哪一家提的意見
|
|
283
294
|
mkdirSync(join(runDir(run.id), "reviews"), { recursive: true });
|
|
284
295
|
renameSync(flowFile(run, "plan-review.json"), join(runDir(run.id), "reviews", `plan-review-${round}-${reviewer}.json`));
|
|
@@ -294,6 +305,10 @@ async function planReviewStage(run) {
|
|
|
294
305
|
issueLines.push(...lines);
|
|
295
306
|
issues.push(opinion(reviewer, lines));
|
|
296
307
|
}
|
|
308
|
+
run = clearModelReviewStage(run, "plan-review");
|
|
309
|
+
const attempts = { ...run.attempts };
|
|
310
|
+
delete attempts["plan-review-run"];
|
|
311
|
+
run = { ...run, attempts };
|
|
297
312
|
if (!firstObjector)
|
|
298
313
|
return planSettled(run, "plan-review");
|
|
299
314
|
const report = `計畫審查要求修改:\n\n${issues.join("\n\n")}`;
|
|
@@ -539,7 +554,7 @@ async function taskReviewStep(run, task, progress, taskJson, acceptanceJson) {
|
|
|
539
554
|
});
|
|
540
555
|
if ("run" in result)
|
|
541
556
|
return result.run;
|
|
542
|
-
const reviewed = succeed(
|
|
557
|
+
const reviewed = succeed(result.state, `${task.id}:review-run`, "implement");
|
|
543
558
|
if (!result.objector)
|
|
544
559
|
return { ...succeed(reviewed, key, "implement"), taskPhase: "verify" };
|
|
545
560
|
return {
|
|
@@ -665,7 +680,7 @@ async function fixStage(run) {
|
|
|
665
680
|
if ("run" in result)
|
|
666
681
|
return result.run;
|
|
667
682
|
// 修正者成為新的作者,下一輪審查會換成別人
|
|
668
|
-
return { ...
|
|
683
|
+
return { ...succeed(run, "fix", "verify"), lastWriter: result.agent };
|
|
669
684
|
}
|
|
670
685
|
/**
|
|
671
686
|
* 由作者以外的審查小組審查 base 之後的變更;回傳第一位要求修改的審查者與所有意見,
|
|
@@ -694,14 +709,16 @@ async function codeReview(run, opts) {
|
|
|
694
709
|
if (tampered.length)
|
|
695
710
|
info(run, ` ↩️ 已還原審查者修改的檔案:${tampered.join(", ")}`);
|
|
696
711
|
if (!r.ok)
|
|
697
|
-
return { run: retry(run, opts.runKey, `Agent 執行失敗:${r.summary}`, opts.backTo) };
|
|
712
|
+
return { run: retry(opts.step === "review" ? recordModelReviewFailure(run, "review", reviewer) : run, opts.runKey, `Agent 執行失敗:${r.summary}`, opts.backTo) };
|
|
698
713
|
const review = readJsonFile(flowFile(run, "review.json"), ConsistentReviewResult);
|
|
699
714
|
if (!review.ok)
|
|
700
|
-
return { run: retry(run, opts.runKey, review.error, opts.backTo) };
|
|
715
|
+
return { run: retry(opts.step === "review" ? recordModelReviewFailure(run, "review", reviewer) : run, opts.runKey, review.error, opts.backTo) };
|
|
701
716
|
const gate = opts.gate ? { target: "code", verdict: review.data.verdict } : undefined;
|
|
702
717
|
const handoffError = finishHandoff(run, outcome, "reviewer", gate);
|
|
703
718
|
if (handoffError)
|
|
704
|
-
return { run: retry(run, opts.runKey, handoffError, opts.backTo) };
|
|
719
|
+
return { run: retry(opts.step === "review" ? recordModelReviewFailure(run, "review", reviewer) : run, opts.runKey, handoffError, opts.backTo) };
|
|
720
|
+
if (opts.step === "review")
|
|
721
|
+
run = clearModelReviewFailure(run, "review", reviewer);
|
|
705
722
|
renameSync(flowFile(run, "review.json"), flowFile(run, opts.saveAs(reviewer)));
|
|
706
723
|
if (review.data.verdict === "approve") {
|
|
707
724
|
info(run, ` ✓ ${reviewer} 核准`);
|
|
@@ -713,7 +730,11 @@ async function codeReview(run, opts) {
|
|
|
713
730
|
.filter((i) => i.status !== "met")
|
|
714
731
|
.map((i) => reviewIssue(i.criterion, i.status, i.note))));
|
|
715
732
|
}
|
|
716
|
-
|
|
733
|
+
if (opts.step === "review")
|
|
734
|
+
run = clearModelReviewStage(run, "review");
|
|
735
|
+
const attempts = { ...run.attempts };
|
|
736
|
+
delete attempts[opts.runKey];
|
|
737
|
+
return { state: { ...run, attempts }, objector, issues };
|
|
717
738
|
}
|
|
718
739
|
async function reviewStage(run) {
|
|
719
740
|
const result = await codeReview(run, {
|
|
@@ -730,9 +751,9 @@ async function reviewStage(run) {
|
|
|
730
751
|
if ("run" in result)
|
|
731
752
|
return result.run;
|
|
732
753
|
if (!result.objector)
|
|
733
|
-
return succeed(
|
|
754
|
+
return succeed(result.state, "review", "pr");
|
|
734
755
|
return {
|
|
735
|
-
...retry(
|
|
756
|
+
...retry(result.state, "review", `程式碼審查要求修改:\n\n${result.issues.join("\n\n")}`, "fix"),
|
|
736
757
|
fixSource: "review",
|
|
737
758
|
lastReviewer: result.objector,
|
|
738
759
|
};
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
import { ModelStage, ModelStrength, RepoConfig } from "./schemas.js";
|
|
2
|
+
const agentsOf = (cfg) => ({ ...(cfg.agents ?? {}) });
|
|
3
|
+
const modelsOf = (agent) => [...(agent.models ?? [])];
|
|
4
|
+
function withAgent(cfg, agent, edit) {
|
|
5
|
+
const agents = agentsOf(cfg);
|
|
6
|
+
if (!agents[agent])
|
|
7
|
+
throw new Error(`未定義的 agent:${agent}`);
|
|
8
|
+
agents[agent] = edit({ ...agents[agent] });
|
|
9
|
+
const next = { ...cfg, agents };
|
|
10
|
+
RepoConfig.parse(next);
|
|
11
|
+
return next;
|
|
12
|
+
}
|
|
13
|
+
export function addModel(cfg, agent, name, strength) {
|
|
14
|
+
ModelStrength.parse(strength);
|
|
15
|
+
if (!name.trim())
|
|
16
|
+
throw new Error("模型名稱不可為空");
|
|
17
|
+
if (name !== name.trim())
|
|
18
|
+
throw new Error("模型名稱前後不可有空白");
|
|
19
|
+
return withAgent(cfg, agent, (def) => {
|
|
20
|
+
const models = modelsOf(def);
|
|
21
|
+
if (models.some((m) => m.name === name))
|
|
22
|
+
throw new Error(`agent ${agent} 的模型名稱 ${name} 重複`);
|
|
23
|
+
return { ...def, models: [...models, { name, strength }] };
|
|
24
|
+
});
|
|
25
|
+
}
|
|
26
|
+
export function setModelStrength(cfg, agent, name, strength) {
|
|
27
|
+
ModelStrength.parse(strength);
|
|
28
|
+
return withAgent(cfg, agent, (def) => {
|
|
29
|
+
const models = modelsOf(def);
|
|
30
|
+
if (!models.some((m) => m.name === name))
|
|
31
|
+
throw new Error(`agent ${agent} 沒有模型 ${name}`);
|
|
32
|
+
return { ...def, models: models.map((m) => m.name === name ? { ...m, strength } : m) };
|
|
33
|
+
});
|
|
34
|
+
}
|
|
35
|
+
/** 有 cycle 時看 cycle,否則看全部 agent;只做設定層的保守檢查。 */
|
|
36
|
+
function assertModelsPresent(cfg) {
|
|
37
|
+
const agents = agentsOf(cfg);
|
|
38
|
+
const affected = cfg.cycle ?? Object.keys(agents);
|
|
39
|
+
const missing = affected.filter((name) => !agents[name] || !modelsOf(agents[name]).length);
|
|
40
|
+
if (missing.length)
|
|
41
|
+
throw new Error(`以下 agent 缺少模型:${missing.join("、")}`);
|
|
42
|
+
}
|
|
43
|
+
export function removeModel(cfg, agent, name) {
|
|
44
|
+
const next = withAgent(cfg, agent, (def) => {
|
|
45
|
+
const models = modelsOf(def);
|
|
46
|
+
if (!models.some((m) => m.name === name))
|
|
47
|
+
throw new Error(`agent ${agent} 沒有模型 ${name}`);
|
|
48
|
+
return { ...def, models: models.filter((m) => m.name !== name) };
|
|
49
|
+
});
|
|
50
|
+
const lastModelRemoved = !modelsOf(agentsOf(next)[agent]).length;
|
|
51
|
+
if (lastModelRemoved || (cfg.modelSelection ?? {}).mode === "adaptive")
|
|
52
|
+
assertModelsPresent(next);
|
|
53
|
+
return next;
|
|
54
|
+
}
|
|
55
|
+
export function setModelMode(cfg, mode) {
|
|
56
|
+
if (mode !== "balanced" && mode !== "adaptive")
|
|
57
|
+
throw new Error(`未知的模型模式:${mode}`);
|
|
58
|
+
if (mode === "adaptive")
|
|
59
|
+
assertModelsPresent(cfg);
|
|
60
|
+
const next = { ...cfg, modelSelection: { ...(cfg.modelSelection ?? {}), mode } };
|
|
61
|
+
RepoConfig.parse(next);
|
|
62
|
+
return next;
|
|
63
|
+
}
|
|
64
|
+
export function setStageStrength(cfg, stage, strength) {
|
|
65
|
+
const parsed = ModelStage.safeParse(stage);
|
|
66
|
+
if (!parsed.success)
|
|
67
|
+
throw new Error(`未知的 LLM 階段:${stage}`);
|
|
68
|
+
ModelStrength.parse(strength);
|
|
69
|
+
const selection = (cfg.modelSelection ?? {});
|
|
70
|
+
const next = { ...cfg, modelSelection: { ...selection, stageStrength: { ...(selection.stageStrength ?? {}), [stage]: strength } } };
|
|
71
|
+
RepoConfig.parse(next);
|
|
72
|
+
return next;
|
|
73
|
+
}
|
|
74
|
+
//# sourceMappingURL=modelConfig.js.map
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
import { mkdtempSync, rmSync } from "node:fs";
|
|
2
|
+
import { tmpdir } from "node:os";
|
|
3
|
+
import { join } from "node:path";
|
|
4
|
+
import { ADAPTERS } from "./agents/index.js";
|
|
5
|
+
import { exec } from "./proc.js";
|
|
6
|
+
export function probeFailureReason(message) {
|
|
7
|
+
if (/\b429\b|quota|rate.?limit|usage limit/i.test(message))
|
|
8
|
+
return `額度或速率限制:${message}`;
|
|
9
|
+
if (/invalid model|model.*(not found|unknown|unavailable)|unknown model/i.test(message))
|
|
10
|
+
return `模型名稱無效或目前不可用:${message}`;
|
|
11
|
+
if (/auth|unauthori[sz]ed|forbidden|\b401\b|\b403\b/i.test(message))
|
|
12
|
+
return `認證或權限失敗:${message}`;
|
|
13
|
+
if (/network|connect|timed? out|dns|enotfound/i.test(message))
|
|
14
|
+
return `網路連線失敗:${message}`;
|
|
15
|
+
return message;
|
|
16
|
+
}
|
|
17
|
+
/** 在獨立暫存目錄對指定模型送出最短請求;安全能力不足時直接拒絕。 */
|
|
18
|
+
export async function probeModel(def, model, timeoutMs = 30000) {
|
|
19
|
+
const dir = mkdtempSync(join(tmpdir(), "agentflowctl-model-"));
|
|
20
|
+
try {
|
|
21
|
+
const adapter = ADAPTERS[def.adapter];
|
|
22
|
+
let inv;
|
|
23
|
+
if (def.adapter === "command") {
|
|
24
|
+
if (!def.modelProbe?.length || !def.modelProbe.some((arg) => arg.includes("{model}"))) {
|
|
25
|
+
return { status: "unverifiable", reason: "command agent 缺少含 {model} 的 modelProbe" };
|
|
26
|
+
}
|
|
27
|
+
const [cmd, ...rest] = def.modelProbe;
|
|
28
|
+
inv = { cmd: cmd.replaceAll("{model}", model), args: rest.map((arg) => arg.replaceAll("{model}", model)) };
|
|
29
|
+
}
|
|
30
|
+
else {
|
|
31
|
+
if (!adapter.invokeModelProbe)
|
|
32
|
+
return { status: "unverifiable", reason: `${def.adapter} CLI 無法安全停用工具,因此不能驗證模型` };
|
|
33
|
+
inv = adapter.invokeModelProbe(model, dir);
|
|
34
|
+
}
|
|
35
|
+
const events = [];
|
|
36
|
+
const result = await exec(inv.cmd, inv.args, {
|
|
37
|
+
cwd: dir, env: inv.env, input: inv.input, timeoutMs,
|
|
38
|
+
onStdoutLine: (line) => { if (def.adapter !== "command")
|
|
39
|
+
events.push(...adapter.parse(line)); },
|
|
40
|
+
});
|
|
41
|
+
if (result.code !== 0)
|
|
42
|
+
return { status: "failed", reason: result.code === 124 ? "模型檢查逾時" : probeFailureReason(result.stderr.trim() || result.stdout.trim() || `結束碼 ${result.code}`) };
|
|
43
|
+
if (def.adapter === "command") {
|
|
44
|
+
try {
|
|
45
|
+
const data = JSON.parse(result.stdout.trim());
|
|
46
|
+
if (data.requestedModel !== model || typeof data.resolvedModel !== "string" || !data.resolvedModel.trim()) {
|
|
47
|
+
return { status: "failed", reason: "自訂探測命令回報的模型名稱不符" };
|
|
48
|
+
}
|
|
49
|
+
return { status: "ok", resolvedModel: data.resolvedModel };
|
|
50
|
+
}
|
|
51
|
+
catch {
|
|
52
|
+
return { status: "failed", reason: "自訂探測命令未輸出有效 JSON" };
|
|
53
|
+
}
|
|
54
|
+
}
|
|
55
|
+
if (events.some((e) => e.kind === "tool"))
|
|
56
|
+
return { status: "failed", reason: "模型檢查期間出現工具呼叫,未通過無工具驗證" };
|
|
57
|
+
const done = events.filter((e) => e.kind === "done").at(-1);
|
|
58
|
+
if (done?.kind !== "done")
|
|
59
|
+
return { status: "failed", reason: "CLI 沒有回報完成" };
|
|
60
|
+
if (!done.ok)
|
|
61
|
+
return { status: "failed", reason: done.summary ?? "CLI 回報失敗" };
|
|
62
|
+
if (!events.some((e) => e.kind === "text" && e.text.trim()))
|
|
63
|
+
return { status: "failed", reason: "CLI 沒有回傳文字內容" };
|
|
64
|
+
const resolved = events.find((e) => e.kind === "model");
|
|
65
|
+
return { status: "ok", resolvedModel: resolved?.kind === "model" ? resolved.id : undefined };
|
|
66
|
+
}
|
|
67
|
+
catch (error) {
|
|
68
|
+
return { status: "failed", reason: error.message };
|
|
69
|
+
}
|
|
70
|
+
finally {
|
|
71
|
+
rmSync(dir, { recursive: true, force: true });
|
|
72
|
+
}
|
|
73
|
+
}
|
|
74
|
+
//# sourceMappingURL=modelProbe.js.map
|
|
@@ -0,0 +1,107 @@
|
|
|
1
|
+
const LEVEL = { low: 0, medium: 1, high: 2 };
|
|
2
|
+
const STRENGTHS = ["low", "medium", "high"];
|
|
3
|
+
export const DEFAULT_STAGE_STRENGTH = {
|
|
4
|
+
spec: "medium", plan: "high", planReview: "high", planFix: "medium", planArbiter: "high",
|
|
5
|
+
taskTests: "low", taskCode: "low", taskReview: "medium", taskFix: "low", fix: "medium", review: "high",
|
|
6
|
+
};
|
|
7
|
+
export function effectiveStageStrengths(cfg) {
|
|
8
|
+
const custom = cfg.modelSelection.stageStrength;
|
|
9
|
+
return Object.fromEntries(Object.keys(DEFAULT_STAGE_STRENGTH).map((stage) => [stage, { strength: custom[stage] ?? DEFAULT_STAGE_STRENGTH[stage], custom: custom[stage] !== undefined }]));
|
|
10
|
+
}
|
|
11
|
+
/** 小組審查失敗只影響該審查者的下一次選模。 */
|
|
12
|
+
export function recordModelReviewFailure(run, step, reviewer) {
|
|
13
|
+
const key = `${step}:${reviewer}`;
|
|
14
|
+
return { ...run, modelRetryAttempts: { ...run.modelRetryAttempts, [key]: (run.modelRetryAttempts?.[key] ?? 0) + 1 } };
|
|
15
|
+
}
|
|
16
|
+
export function clearModelReviewFailure(run, step, reviewer) {
|
|
17
|
+
const attempts = { ...run.modelRetryAttempts };
|
|
18
|
+
delete attempts[`${step}:${reviewer}`];
|
|
19
|
+
return { ...run, modelRetryAttempts: attempts };
|
|
20
|
+
}
|
|
21
|
+
export function clearModelReviewStage(run, step) {
|
|
22
|
+
const attempts = Object.fromEntries(Object.entries(run.modelRetryAttempts ?? {}).filter(([key]) => !key.startsWith(`${step}:`)));
|
|
23
|
+
return { ...run, modelRetryAttempts: attempts };
|
|
24
|
+
}
|
|
25
|
+
/** 把步驟名稱(例如 `T-1-code`、`plan-review`)對應到階段鍵;不認得時回傳 undefined。 */
|
|
26
|
+
export function stageOfStep(step) {
|
|
27
|
+
const task = /^T-\d+-(tests|code|review|fix)$/.exec(step);
|
|
28
|
+
if (task)
|
|
29
|
+
return { tests: "taskTests", code: "taskCode", review: "taskReview", fix: "taskFix" }[task[1]];
|
|
30
|
+
const stage = {
|
|
31
|
+
spec: "spec", plan: "plan", "plan-review": "planReview", "plan-fix": "planFix",
|
|
32
|
+
"plan-arbiter": "planArbiter", fix: "fix", review: "review",
|
|
33
|
+
};
|
|
34
|
+
return stage[step];
|
|
35
|
+
}
|
|
36
|
+
function stepStage(step) {
|
|
37
|
+
const found = stageOfStep(step);
|
|
38
|
+
if (!found)
|
|
39
|
+
throw new Error(`未知的 LLM 步驟:${step}`);
|
|
40
|
+
return found;
|
|
41
|
+
}
|
|
42
|
+
function escalation(run, stage, step, reviewer) {
|
|
43
|
+
const a = run.attempts;
|
|
44
|
+
if (stage === "planReview" || stage === "review")
|
|
45
|
+
return run.modelRetryAttempts?.[`${step}:${reviewer ?? ""}`] ?? 0;
|
|
46
|
+
if (stage === "planArbiter")
|
|
47
|
+
return 0;
|
|
48
|
+
if (stage === "planFix")
|
|
49
|
+
return (a["plan-fix"] ?? 0) + Math.max((a["plan-review"] ?? 0) + (a["plan-handoff"] ?? 0) - 1, 0);
|
|
50
|
+
if (stage === "taskFix") {
|
|
51
|
+
const id = step.slice(0, -4);
|
|
52
|
+
return (a[`${id}:fix`] ?? 0) + Math.max((a[`${id}:review`] ?? 0) + (a[`${id}:verify`] ?? 0) - 1, 0);
|
|
53
|
+
}
|
|
54
|
+
if (stage === "fix")
|
|
55
|
+
return (a.fix ?? 0) + Math.max((a.verify ?? 0) + (a.review ?? 0) - 1, 0);
|
|
56
|
+
if (stage === "taskTests" || stage === "taskCode" || stage === "taskReview") {
|
|
57
|
+
const id = step.slice(0, step.lastIndexOf("-"));
|
|
58
|
+
const suffix = step.slice(step.lastIndexOf("-") + 1);
|
|
59
|
+
return a[`${id}:${suffix === "review" ? "review-run" : suffix}`] ?? 0;
|
|
60
|
+
}
|
|
61
|
+
return a[step] ?? 0;
|
|
62
|
+
}
|
|
63
|
+
/** 每次呼叫前依目前設定與 run 狀態選擇模型;不修改任何狀態。 */
|
|
64
|
+
export function selectModel(run, cfg, agent, step, complexity, reviewer) {
|
|
65
|
+
const def = cfg.agents[agent];
|
|
66
|
+
if (!def)
|
|
67
|
+
throw new Error(`未定義的 agent:${agent}`);
|
|
68
|
+
const mode = run.modelMode ?? "balanced";
|
|
69
|
+
if (mode === "balanced") {
|
|
70
|
+
const fallback = def.adapter === "command" ? undefined : cfg.defaultModels[def.adapter];
|
|
71
|
+
return { mode, name: def.model ?? fallback, insufficient: false };
|
|
72
|
+
}
|
|
73
|
+
if (!def.models?.length)
|
|
74
|
+
throw new Error(`agent ${agent} 沒有設定 models`);
|
|
75
|
+
const stage = stepStage(step);
|
|
76
|
+
const floor = effectiveStageStrengths(cfg)[stage].strength;
|
|
77
|
+
const baseline = stage.startsWith("task") ? Math.max(LEVEL[floor], LEVEL[complexity ?? "medium"]) : LEVEL[floor];
|
|
78
|
+
const targetStrength = STRENGTHS[Math.min(2, baseline + escalation(run, stage, step, reviewer))];
|
|
79
|
+
const candidates = def.models.filter((m) => LEVEL[m.strength] >= LEVEL[targetStrength]);
|
|
80
|
+
const min = Math.min(...candidates.map((m) => LEVEL[m.strength]));
|
|
81
|
+
const chosen = candidates.find((m) => LEVEL[m.strength] === min)
|
|
82
|
+
?? def.models.reduce((best, current) => LEVEL[current.strength] > LEVEL[best.strength] ? current : best);
|
|
83
|
+
return { mode, name: chosen.name, strength: chosen.strength, targetStrength, insufficient: LEVEL[chosen.strength] < LEVEL[targetStrength] };
|
|
84
|
+
}
|
|
85
|
+
/** 啟用 adaptive 前檢查本次實際參與的 agent。 */
|
|
86
|
+
export function validateAdaptiveConfig(cfg, cycle) {
|
|
87
|
+
for (const name of cycle) {
|
|
88
|
+
const def = cfg.agents[name];
|
|
89
|
+
if (!def)
|
|
90
|
+
throw new Error(`未定義的 agent:${name}`);
|
|
91
|
+
if (!def.models?.length)
|
|
92
|
+
throw new Error(`agent ${name} 的 models 至少要有一個模型`);
|
|
93
|
+
const names = def.models.map((m) => m.name);
|
|
94
|
+
if (new Set(names).size !== names.length)
|
|
95
|
+
throw new Error(`agent ${name} 的 models 有重複名稱`);
|
|
96
|
+
if (def.adapter === "command") {
|
|
97
|
+
if (!def.command?.some((arg) => arg.includes("{model}")))
|
|
98
|
+
throw new Error(`agent ${name} 的 command 缺少 {model}`);
|
|
99
|
+
if (!def.modelProbe?.length || !def.modelProbe.some((arg) => arg.includes("{model}")))
|
|
100
|
+
throw new Error(`agent ${name} 的 modelProbe 缺少 {model}`);
|
|
101
|
+
}
|
|
102
|
+
else if (def.extraArgs.some((arg) => arg === "--model" || arg === "-m" || arg.startsWith("--model="))) {
|
|
103
|
+
throw new Error(`agent ${name} 的 extraArgs 與 adaptive 模型參數衝突`);
|
|
104
|
+
}
|
|
105
|
+
}
|
|
106
|
+
}
|
|
107
|
+
//# sourceMappingURL=modelSelection.js.map
|
package/dist/proc.js
CHANGED
|
@@ -6,6 +6,8 @@ export function exec(cmd, args, opts = {}) {
|
|
|
6
6
|
env: { ...process.env, ...opts.env },
|
|
7
7
|
shell: opts.shell ?? false,
|
|
8
8
|
});
|
|
9
|
+
let timedOut = false;
|
|
10
|
+
const timer = opts.timeoutMs ? setTimeout(() => { timedOut = true; child.kill("SIGKILL"); }, opts.timeoutMs) : undefined;
|
|
9
11
|
let stdout = "";
|
|
10
12
|
let stderr = "";
|
|
11
13
|
let pending = "";
|
|
@@ -22,11 +24,14 @@ export function exec(cmd, args, opts = {}) {
|
|
|
22
24
|
child.stderr.on("data", (d) => {
|
|
23
25
|
stderr += d.toString();
|
|
24
26
|
});
|
|
25
|
-
child.on("error",
|
|
27
|
+
child.on("error", (error) => { if (timer)
|
|
28
|
+
clearTimeout(timer); reject(error); });
|
|
26
29
|
child.on("close", (code) => {
|
|
30
|
+
if (timer)
|
|
31
|
+
clearTimeout(timer);
|
|
27
32
|
if (opts.onStdoutLine && pending)
|
|
28
33
|
opts.onStdoutLine(pending);
|
|
29
|
-
resolve({ code: code ?? 1, stdout, stderr });
|
|
34
|
+
resolve({ code: timedOut ? 124 : code ?? 1, stdout, stderr: timedOut ? `${stderr}\n執行逾時` : stderr });
|
|
30
35
|
});
|
|
31
36
|
child.stdin.on("error", () => { }); // 子程序不讀 stdin 時忽略 EPIPE
|
|
32
37
|
child.stdin.end(opts.input ?? "");
|
package/dist/runner.js
CHANGED
|
@@ -63,11 +63,14 @@ export async function runAgent(name, def, t, prompt) {
|
|
|
63
63
|
projectRoot: projectRoot(),
|
|
64
64
|
command: def.command,
|
|
65
65
|
});
|
|
66
|
-
appendLog(t.logFile, headerLine({ stage: t.stage, step: t.step, agent: name, adapter: def.adapter, startedAt: new Date().toISOString() }));
|
|
66
|
+
appendLog(t.logFile, headerLine({ stage: t.stage, step: t.step, agent: name, adapter: def.adapter, model: def.model, strength: t.strength, targetStrength: t.targetStrength, startedAt: new Date().toISOString() }));
|
|
67
67
|
let done;
|
|
68
68
|
let lastText = "";
|
|
69
69
|
let inputTokens = 0;
|
|
70
70
|
let outputTokens = 0;
|
|
71
|
+
let inputReported = false;
|
|
72
|
+
let outputReported = false;
|
|
73
|
+
let resolvedModel;
|
|
71
74
|
const r = await exec(inv.cmd, inv.args, {
|
|
72
75
|
cwd: t.cwd,
|
|
73
76
|
env: inv.env,
|
|
@@ -87,8 +90,17 @@ export async function runAgent(name, def, t, prompt) {
|
|
|
87
90
|
console.log(formatToolLine(name, ev));
|
|
88
91
|
}
|
|
89
92
|
else if (ev.kind === "usage") {
|
|
90
|
-
|
|
91
|
-
|
|
93
|
+
if (ev.inputTokens !== undefined) {
|
|
94
|
+
inputReported = true;
|
|
95
|
+
inputTokens += ev.inputTokens;
|
|
96
|
+
}
|
|
97
|
+
if (ev.outputTokens !== undefined) {
|
|
98
|
+
outputReported = true;
|
|
99
|
+
outputTokens += ev.outputTokens;
|
|
100
|
+
}
|
|
101
|
+
}
|
|
102
|
+
else if (ev.kind === "model") {
|
|
103
|
+
resolvedModel = ev.id;
|
|
92
104
|
}
|
|
93
105
|
else if (ev.kind === "done") {
|
|
94
106
|
done = { ok: ev.ok, summary: ev.summary };
|
|
@@ -105,7 +117,8 @@ export async function runAgent(name, def, t, prompt) {
|
|
|
105
117
|
const summary = done?.summary || lastText || tail(r.stdout, 2000) || tail(r.stderr, 2000);
|
|
106
118
|
const quotaExhausted = !ok && isQuotaError(`${summary}\n${done?.summary ?? ""}\n${r.stderr}\n${tail(r.stdout, 4000)}`);
|
|
107
119
|
const meta = parseResultMeta(summary) ?? parseResultMeta(lastText);
|
|
108
|
-
return { ok, quotaExhausted, summary, meta,
|
|
120
|
+
return { ok, quotaExhausted, summary, meta, usageReported: inputReported && outputReported,
|
|
121
|
+
inputTokens: inputReported ? inputTokens : undefined, outputTokens: outputReported ? outputTokens : undefined, resolvedModel };
|
|
109
122
|
}
|
|
110
123
|
/**
|
|
111
124
|
* 在 worktree 內執行專案指令(安裝、測試、建置)。
|
package/dist/schemas.js
CHANGED
|
@@ -71,6 +71,7 @@ export const TaskItem = z.object({
|
|
|
71
71
|
description: z.string().min(1),
|
|
72
72
|
dependsOn: z.array(z.string()).default([]),
|
|
73
73
|
acceptance: z.array(z.string()).min(1, "每個任務至少要對應一條驗收條件"),
|
|
74
|
+
complexity: z.enum(["low", "medium", "high"]).optional(),
|
|
74
75
|
});
|
|
75
76
|
/** Agent 在 plan 階段產出的 .flow/tasks.json */
|
|
76
77
|
export const TaskList = z.array(TaskItem).min(1);
|
|
@@ -97,12 +98,17 @@ export const ArbiterResult = ReviewResult.extend({
|
|
|
97
98
|
.transform((verdict) => verdict === "reject" ? "changes_requested" : verdict),
|
|
98
99
|
});
|
|
99
100
|
/** 一個 agent 的定義;名稱(agents 的 key)用在 cycle 裡 */
|
|
101
|
+
export const ModelStrength = z.enum(["low", "medium", "high"]);
|
|
102
|
+
export const ModelStage = z.enum(["spec", "plan", "planReview", "planFix", "planArbiter", "taskTests", "taskCode", "taskReview", "taskFix", "fix", "review"]);
|
|
103
|
+
export const ModelEntry = z.object({ name: z.string().trim().min(1), strength: ModelStrength });
|
|
100
104
|
export const AgentDef = z.object({
|
|
101
105
|
adapter: z.enum(["claude", "codex", "gemini", "command"]),
|
|
102
106
|
model: z.string().optional(),
|
|
107
|
+
models: z.array(ModelEntry).optional(),
|
|
103
108
|
extraArgs: z.array(z.string()).default([]),
|
|
104
109
|
/** 只有 command adapter 使用,`{prompt}` 會被替換成 prompt */
|
|
105
110
|
command: z.array(z.string()).optional(),
|
|
111
|
+
modelProbe: z.array(z.string()).optional(),
|
|
106
112
|
});
|
|
107
113
|
/** 目標專案可選的 flow.config.json,預設值對應 Vite + TypeScript + Vitest 專案 */
|
|
108
114
|
export const RepoConfig = z.object({
|
|
@@ -114,6 +120,15 @@ export const RepoConfig = z.object({
|
|
|
114
120
|
codex: z.string().trim().min(1).optional(),
|
|
115
121
|
gemini: z.string().trim().min(1).optional(),
|
|
116
122
|
}).default({}),
|
|
123
|
+
modelSelection: z.object({
|
|
124
|
+
mode: z.enum(["balanced", "adaptive"]).default("balanced"),
|
|
125
|
+
stageStrength: z.strictObject({
|
|
126
|
+
spec: ModelStrength.optional(), plan: ModelStrength.optional(), planReview: ModelStrength.optional(),
|
|
127
|
+
planFix: ModelStrength.optional(), planArbiter: ModelStrength.optional(), taskTests: ModelStrength.optional(),
|
|
128
|
+
taskCode: ModelStrength.optional(), taskReview: ModelStrength.optional(), taskFix: ModelStrength.optional(),
|
|
129
|
+
fix: ModelStrength.optional(), review: ModelStrength.optional(),
|
|
130
|
+
}).default({}),
|
|
131
|
+
}).default({ mode: "balanced", stageStrength: {} }),
|
|
117
132
|
/** 參與的 agent(順序不影響分工);未設定時取 agents 裡已安裝的 CLI */
|
|
118
133
|
cycle: z.array(z.string()).min(1).optional(),
|
|
119
134
|
/** review 後的修正由誰做:ring=輪到下一位;author=最後寫程式的 agent */
|
|
@@ -171,6 +186,8 @@ export const FlowRun = z.object({
|
|
|
171
186
|
fixSource: z.enum(["verify", "review"]).optional(),
|
|
172
187
|
/** 各關卡的連續失敗次數 */
|
|
173
188
|
attempts: z.record(z.string(), z.number()),
|
|
189
|
+
modelMode: z.enum(["balanced", "adaptive"]).optional(),
|
|
190
|
+
modelRetryAttempts: z.record(z.string(), z.number().int().nonnegative()).optional(),
|
|
174
191
|
taskIndex: z.number().int().nonnegative(),
|
|
175
192
|
/** 目前任務進行到哪一步:寫測試 → 實作 → 審查 → 驗證,審查或驗證未通過時進入修正 */
|
|
176
193
|
taskPhase: z.enum(["tests", "code", "review", "verify", "fix"]),
|
package/dist/store.js
CHANGED
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
import { appendFileSync, existsSync, mkdirSync, readdirSync, readFileSync, renameSync, writeFileSync } from "node:fs";
|
|
2
2
|
import { dirname, join } from "node:path";
|
|
3
|
+
import { stageOfStep } from "./modelSelection.js";
|
|
3
4
|
import { agentflowctlDir, runDir } from "./paths.js";
|
|
4
5
|
import { FlowRun } from "./schemas.js";
|
|
5
6
|
const statePath = (id) => join(runDir(id), "state.json");
|
|
@@ -34,21 +35,40 @@ export function addUsage(id, entry) {
|
|
|
34
35
|
mkdirSync(runDir(id), { recursive: true });
|
|
35
36
|
appendFileSync(usagePath(id), `${JSON.stringify({ at: new Date().toISOString(), ...entry })}\n`);
|
|
36
37
|
}
|
|
37
|
-
|
|
38
|
-
export function
|
|
38
|
+
const emptySummary = () => ({ tokens: 0, inputTokens: 0, outputTokens: 0, runs: 0, reportedRuns: 0, unreportedRuns: 0, legacyRuns: 0, legacyTokens: 0 });
|
|
39
|
+
export function listUsage(id) {
|
|
39
40
|
const p = usagePath(id);
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
41
|
+
return existsSync(p) ? readFileSync(p, "utf8").split("\n").filter(Boolean).map((line) => JSON.parse(line)) : [];
|
|
42
|
+
}
|
|
43
|
+
function addSummary(acc, e) {
|
|
44
|
+
acc.runs += 1;
|
|
45
|
+
if (e.usageReported === true && typeof e.inputTokens === "number" && typeof e.outputTokens === "number") {
|
|
46
|
+
acc.reportedRuns += 1;
|
|
47
|
+
acc.inputTokens += e.inputTokens ?? 0;
|
|
48
|
+
acc.outputTokens += e.outputTokens ?? 0;
|
|
47
49
|
acc.tokens += (e.inputTokens ?? 0) + (e.outputTokens ?? 0);
|
|
48
|
-
acc.runs += 1;
|
|
49
50
|
}
|
|
51
|
+
else if (e.usageReported === false || e.usageReported === true)
|
|
52
|
+
acc.unreportedRuns += 1;
|
|
53
|
+
else {
|
|
54
|
+
acc.legacyRuns += 1;
|
|
55
|
+
acc.legacyTokens += (e.inputTokens ?? 0) + (e.outputTokens ?? 0);
|
|
56
|
+
}
|
|
57
|
+
}
|
|
58
|
+
/** 只加總明確回報的 token,舊資料的 0 不視為已回報。 */
|
|
59
|
+
function groupUsage(id, keyOf) {
|
|
60
|
+
const out = {};
|
|
61
|
+
for (const e of listUsage(id))
|
|
62
|
+
addSummary(out[keyOf(e)] ??= emptySummary(), e);
|
|
50
63
|
return out;
|
|
51
64
|
}
|
|
65
|
+
export const usageByAgent = (id) => groupUsage(id, (e) => e.agent);
|
|
66
|
+
export const usageByModelStage = (id) => groupUsage(id, (e) => `${e.model ?? "CLI 預設(名稱未知)"} / ${e.stage}`);
|
|
67
|
+
export const usageByStrength = (id) => groupUsage(id, (e) => e.strength ?? "未知");
|
|
68
|
+
/** 依階段鍵加總(所有任務的同一步合在一起);認不得的步驟歸為「其他」。 */
|
|
69
|
+
export const usageByStage = (id) => groupUsage(id, (e) => stageOfStep(e.stage) ?? "其他");
|
|
70
|
+
/** 依任務加總寫測試、實作、任務審查與任務修正;規格、計畫與整體階段歸為「非任務步驟」。 */
|
|
71
|
+
export const usageByTask = (id) => groupUsage(id, (e) => /^(T-\d+)-/.exec(e.stage)?.[1] ?? "非任務步驟");
|
|
52
72
|
/** 這個 run 已執行 agent 的次數(每次執行都會記一筆用量) */
|
|
53
73
|
export function agentRuns(id) {
|
|
54
74
|
const p = usagePath(id);
|
package/dist/tasks.js
CHANGED
|
@@ -1,5 +1,12 @@
|
|
|
1
1
|
/** 一個任務最多做兩件事:對應的驗收條件超過這個數量就要再拆 */
|
|
2
2
|
export const MAX_TASK_ACCEPTANCE = 2;
|
|
3
|
+
/** adaptive 的新計畫必須明確標註難度;舊 run 仍可讀取缺少欄位的 task。 */
|
|
4
|
+
export function validateTaskComplexity(tasks, mode) {
|
|
5
|
+
if (mode === "balanced")
|
|
6
|
+
return undefined;
|
|
7
|
+
const missing = tasks.filter((task) => !task.complexity).map((task) => task.id);
|
|
8
|
+
return missing.length ? `任務缺少 complexity:${missing.join("、")}` : undefined;
|
|
9
|
+
}
|
|
3
10
|
/** 只取目前任務負責的驗收條件,避免每次實作都重讀整份清單。 */
|
|
4
11
|
export function taskAcceptance(task, acceptance) {
|
|
5
12
|
const byId = new Map(acceptance.map((item) => [item.id, item]));
|
package/package.json
CHANGED
package/prompts/plan-fix.md
CHANGED
|
@@ -32,6 +32,7 @@
|
|
|
32
32
|
- 每一條驗收條件都至少要有一個任務負責;`dependsOn` 不可有循環。
|
|
33
33
|
- 一個任務只做一件事,最多兩件:`acceptance` 最多列兩條驗收條件;驗收條件一條只描述一個行為。修改時若任務變大,請拆開,不要合併。
|
|
34
34
|
- 測試檔名必須符合正規表示式 `{{testPattern}}`。
|
|
35
|
+
- 修改 task 時保留或補上 `complexity`(`low`、`medium`、`high`)。依影響範圍、技術不確定性與失敗後果重新判定,取最高等級;同步更新 .flow/plan.md 中該 task 的逐項證據與最終等級。若不同意審查者建議的等級,在「## 審查回應」引用具體程式碼或測試依據。
|
|
35
36
|
</output_format>
|
|
36
37
|
|
|
37
38
|
<constraints>
|
package/prompts/plan-review.md
CHANGED
|
@@ -30,6 +30,8 @@
|
|
|
30
30
|
1. **需求覆蓋**:規格是否完整涵蓋原始需求?有沒有遺漏、誤解,或加入需求沒要求的範圍?
|
|
31
31
|
2. **驗收條件**:每一條是否具體、可以用自動化測試驗證,而且只描述一個行為?把多個行為寫在同一條的,要求拆開。有沒有重要的邊界情況或錯誤處理沒被列入?
|
|
32
32
|
3. **任務拆解**:每個任務是否只做一件事(最多兩件),小到一次 TDD 循環就能完成,而且能寫出「實作前會失敗」的測試?任務太大、一次要動很多檔案或驗證很多行為的,要求拆成更小的任務。相依順序是否合理?
|
|
33
|
+
同時依 .flow/plan.md 的逐項理由及實際程式碼,獨立核對每個 task 的 `complexity`:分別看影響範圍、技術不確定性與失敗後果,取最高等級。`low` 須是沿用既有做法、侷限單一行為或模組且失敗可由局部測試發現;`medium` 包括多模組或介面協調、非典型邊界、相容性或狀態遷移風險;`high` 包括跨系統契約、架構或資料模型變更、未知的關鍵技術路徑,或資料遺失、權限、難以回復的風險。不要只憑檔案數、程式碼行數或驗收條件數判定。
|
|
34
|
+
理由缺漏、與程式碼不符,或高低估會影響選模時,要求修正;在 `note` 指出 task ID、具體證據、建議等級及須修改的 .flow/plan.md/.flow/tasks.json 部分。不要為缺少高價值證據的細微措辭差異要求修改。
|
|
33
35
|
4. **技術方向**:是否符合專案既有的架構與慣例?有沒有更簡單的做法,或明顯的風險?
|
|
34
36
|
|
|
35
37
|
措辭、格式這類不影響實作結果的小問題,不需要要求修改。
|
package/prompts/plan.md
CHANGED
|
@@ -25,7 +25,7 @@
|
|
|
25
25
|
<steps>
|
|
26
26
|
1. 若 .flow/feedback.md 存在,先閱讀,並依內容修正前次的產出。
|
|
27
27
|
2. 閱讀規格與相關程式碼。
|
|
28
|
-
3. 撰寫 .flow/plan.md
|
|
28
|
+
3. 撰寫 .flow/plan.md:整體實作方式、要新增或修改的模組、任務順序的理由;逐一記錄 task 的難度判定依據。
|
|
29
29
|
4. 撰寫 .flow/tasks.json。
|
|
30
30
|
</steps>
|
|
31
31
|
|
|
@@ -38,6 +38,7 @@
|
|
|
38
38
|
"id": "T-1",
|
|
39
39
|
"title": "建立表單驗證 schema",
|
|
40
40
|
"description": "具體要做什麼、要動哪些檔案、測試要驗證什麼行為",
|
|
41
|
+
"complexity": "low",
|
|
41
42
|
"dependsOn": [],
|
|
42
43
|
"acceptance": ["AC-1"]
|
|
43
44
|
}
|
|
@@ -51,6 +52,11 @@
|
|
|
51
52
|
- `title` 用一句話說出這件事;需要用「並且」「以及」串起來的,就是兩個任務。
|
|
52
53
|
- `description` 寫清楚要動哪些檔案、測試要驗證哪個行為,以及這個任務不做什麼。
|
|
53
54
|
- 每個任務都必須能寫出「在實作前會失敗」的測試;純設定或重構類工作請併入相關任務。
|
|
55
|
+
- 每個任務先檢查預計修改的程式碼,再依「影響範圍、技術不確定性、失敗後果」三個面向判定 `complexity`,取其中最高的等級;不要只憑檔案數、程式碼行數或驗收條件數判定。
|
|
56
|
+
- `low`:沿用現有做法,變更侷限在單一行為或模組,失敗容易由局部測試發現且不影響既有資料或對外契約。
|
|
57
|
+
- `medium`:需要協調多個模組或既有介面、處理非典型邊界,或有相容性與狀態遷移風險,但可依已知做法實作與驗證。
|
|
58
|
+
- `high`:涉及跨系統契約、架構或資料模型變更;關鍵技術路徑尚不確定;或失敗可能造成資料遺失、權限問題或難以回復的影響。任一面向符合就標 `high`。
|
|
59
|
+
- 在 .flow/plan.md 逐一列出 task ID、三個面向的具體證據與最終等級;缺少證據時先查閱相關程式碼,不要一律標 `low` 或憑猜測調高。
|
|
54
60
|
- 測試檔名必須符合正規表示式 `{{testPattern}}`。
|
|
55
61
|
- 每一條驗收條件都至少要有一個任務負責;`dependsOn` 不可有循環。
|
|
56
62
|
</guidelines>
|