deepseek-foreman 0.2.2 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.en.md
CHANGED
|
@@ -140,6 +140,10 @@ Also: **hand-editing a bundle's patch file does not notify the running Host** (c
|
|
|
140
140
|
|
|
141
141
|
## Field data
|
|
142
142
|
|
|
143
|
+

|
|
144
|
+
|
|
145
|
+
> Per-item numbers, model-role logic and cost caveats: **[Quantified metrics](docs/metrics-2026-10-01.md)** (with data sources and how to reproduce).
|
|
146
|
+
|
|
143
147
|
The full sprint ran in one night (2026-09-30 23:00 – 2026-10-01 12:00); the numbers below come from that run's tickets and receipts, with the field report in [docs/dogfooding.md](docs/dogfooding.md).
|
|
144
148
|
|
|
145
149
|
| Item | Value |
|
|
@@ -172,6 +176,16 @@ The three pieces cover one segment each and add up:
|
|
|
172
176
|
|
|
173
177
|
**Honest accounting**: the overall numbers from the 19 tickets are in [Field data](#field-data) above; per-plugin isolated quantification **has had no A/B control, so no numbers are given** — only mechanisms: build dominates token spend → cheap workers; dialogue and judgement stay terse → persona; quality holds → cross-vendor review + item-by-item verification. The three mechanisms each own one segment, and they stack.
|
|
174
178
|
|
|
179
|
+
## Ticket tiers (effort scaling)
|
|
180
|
+
|
|
181
|
+
| Tier | Flow | Dispatch effort |
|
|
182
|
+
|---|---|---|
|
|
183
|
+
| `trivial` (≤10 lines/docs) | skip cross-vendor review, Lead accepts | low |
|
|
184
|
+
| `normal` (default) | full loop + cheap cross-vendor review | per workers.md |
|
|
185
|
+
| `critical` (data/release/security) | double review + physical read-only | high; escalate after 2 failures |
|
|
186
|
+
|
|
187
|
+
Every receipt logs a cost ledger (build token/effort/wall-clock); budget caps halt unattended runs. Model-role logic and measured data: [Quantified metrics](docs/metrics-2026-10-01.md).
|
|
188
|
+
|
|
175
189
|
## Status
|
|
176
190
|
|
|
177
191
|
Installed into dsh desktop 0.2.0-rc.2 (Intel iMac) and verified live: cold start is normal, `pick_route` is callable by the model, the vision constraint really blocks and returns a fallback; the final `subagent_readonly` state is live-verified too (read-only tool set enforced, delegation tools blocked).
|
package/README.md
CHANGED
|
@@ -138,6 +138,10 @@ bundle 的 `cordis.patch.yml` 里**新增插件行要包在 `insert:` 下**:
|
|
|
138
138
|
|
|
139
139
|
## 实测数据
|
|
140
140
|
|
|
141
|
+

|
|
142
|
+
|
|
143
|
+
> 逐项数字、模型分工逻辑、成本口径见 **[可量化测试数据](docs/metrics-2026-10-01.md)**(含图表数据源与复核方法)。
|
|
144
|
+
|
|
141
145
|
一晚(2026-09-30 23:00 – 2026-10-01 12:00)跑完整轮冲刺的记录,数字来自当轮工单与回执,实证见 [docs/dogfooding.md](docs/dogfooding.md)。
|
|
142
146
|
|
|
143
147
|
| 项 | 内容 |
|
|
@@ -170,6 +174,16 @@ bundle 的 `cordis.patch.yml` 里**新增插件行要包在 `insert:` 下**:
|
|
|
170
174
|
|
|
171
175
|
**口径(诚实)**:19 张工单的整体数据见上文[实测数据](#实测数据);逐插件的孤立量化贡献**没做 A/B 对照,不给数字**,只讲机制——token 大头在实现 → 便宜工人;对话与判断从简 → persona 精简;质量不降 → 异族审查 + 逐条核实。三个机制各管一段,叠加生效。
|
|
172
176
|
|
|
177
|
+
## 工单分级(effort scaling)
|
|
178
|
+
|
|
179
|
+
| 级别 | 走什么流程 | 派单 effort |
|
|
180
|
+
|---|---|---|
|
|
181
|
+
| `trivial`(≤10 行/纯文档) | 跳过异族审查,Lead 验收收尾 | low |
|
|
182
|
+
| `normal`(默认) | 完整闭环 + 异族便宜审查 | 按 workers.md |
|
|
183
|
+
| `critical`(数据/发布/安全) | 双审查 + 物理只读审查 | high;失败 2 次升档重派 |
|
|
184
|
+
|
|
185
|
+
成本台账随每张回执记录(施工 token/effort/wall-clock),`_receipts/progress.md` 周汇总;无人托管有预算闸(到限即停)。分工逻辑与实测数据见 [可量化测试数据](docs/metrics-2026-10-01.md)。
|
|
186
|
+
|
|
173
187
|
## 状态
|
|
174
188
|
|
|
175
189
|
已装进 dsh 桌面版 0.2.0-rc.2(Intel iMac)并 live 验证:冷启动正常,`pick_route` 可被模型调用,视觉约束会真的拦截并给 fallback;`subagent_readonly` 最终态同样 live 实测通过(只读工具集生效、委派工具被拦死)。
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "deepseek-foreman",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.1",
|
|
4
4
|
"description": "Foreman for DeepSeek Harness: your best model leads, cheaper models build, a rival vendor reviews. Role-to-route adjudication with peak-window, vision, output-size and cross-vendor-review constraints.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "lib/index.js",
|
|
@@ -67,9 +67,16 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
67
67
|
- **工单头写 `worker-route: <provider>/<model>`**,派单时照它传参。
|
|
68
68
|
- **碰线上服务的单,停机要短**:备份、大文件传输、长测试都在停服务之前做完;从停到拉起之间只切文件和做必要校验;所有远程连接设超时。
|
|
69
69
|
- **写两张,派一张。** `open/` 只放现在能跑的;下一张往往取决于上一张回执里的存疑项。
|
|
70
|
+
- **工单分级(effort scaling)**,工单头写 `difficulty:`:
|
|
71
|
+
- `trivial`(纯文档/≤10 行改动):跳过异族审查,Lead 验收即收尾;派单 effort 用 low
|
|
72
|
+
- `normal`(默认):完整闭环(派单→验收→异族审查→核实);effort 按 workers.md
|
|
73
|
+
- `critical`(碰数据/发布/安全/架构):双审查(两家各审一轮);effort high
|
|
74
|
+
- 失败 2 次升级:同角色 effort 升 max 或换高档施工模型重派,再败才挪 blocked
|
|
70
75
|
|
|
71
76
|
## 三、派单
|
|
72
77
|
|
|
78
|
+
**派单前必须调 `pick_route`**(把工单的 worker-route 传给它):拿到的 route 原样传给 subagent;被拒(高峰/视觉/输出上限/同族审查)就照它给的 fallback/alternatives 换,**不允许绕过硬约束自己挑模型**。
|
|
79
|
+
|
|
73
80
|
一次 `subagent` 调用 = 一张单。参数:
|
|
74
81
|
|
|
75
82
|
- `prompt`:照下面这段,**只给绝对路径**
|
|
@@ -102,6 +109,8 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
102
109
|
|
|
103
110
|
**回执是说法,不是证据。**
|
|
104
111
|
|
|
112
|
+
**回执阅读纪律(省 token)**:回执的「1. 做了什么」「4. 成本台账」「5. 风险与存疑」是结论层——Lead 默认只读这三段;「3. 验收证据」原文段只在抽查或存疑指向它时才读。验收输出超过 50 行让工人落盘到 _receipts/logs/,回执只贴命令+首尾 10 行。
|
|
113
|
+
|
|
105
114
|
1. **每条验收命令自己重跑一遍**,和回执里贴的对比。
|
|
106
115
|
2. **看改动范围**:用 git 就看 `git status` 和 `git diff`;不用 git 就逐个读回执列出的文件。是不是只动了工单允许的?有没有删掉不该动的?
|
|
107
116
|
3. **有实物就看实物**:打开页面、调接口、截图。
|
|
@@ -142,7 +151,7 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
142
151
|
**走之前**:本来会中途打断他的事一次问完,一次一件。他人已经走了就**不要停在问题上等**——一律按最保守的理解,他没明确允许的都算「绝不能自己做」。
|
|
143
152
|
|
|
144
153
|
1. 拆成工单,顺序写进 `_tickets/queue.md`。只有现在能跑的单放 `open/`;轮到下一张时再写进去。
|
|
145
|
-
2. 写 `_tickets/handoff.md
|
|
154
|
+
2. 写 `_tickets/handoff.md`:目标;工单顺序;**预算闸(写死数字)**:单张施工 token 上限(默认 500 万)、累计上限、Lead 轮数上限(默认 200 轮);哪些可以不问直接做;哪些绝不能自己做;几点停。**到任一上限立即停派,写报告等用户**。
|
|
146
155
|
3. 告诉他什么会让托管中断,请他走前处理:**电脑不能休眠**;权限确认不能没人点;**模型套餐有用量上限**,到顶就停,已派出去的任务会跑完,重置后要他叫你继续。
|
|
147
156
|
|
|
148
157
|
**他不在时**:
|
|
@@ -150,7 +159,7 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
150
159
|
- 按队列往下做,不等人说「继续」:派单 → 验收 → 审查 → 核实 → 修复单 → 下一张。
|
|
151
160
|
- 每个在跑的子任务都要有等待手段(`list_agents` 看状态,或后台轮询回执)。
|
|
152
161
|
- 遇到 `handoff.md` 没写到的决定:**不要猜,也不要全停**。记进 `_tickets/decisions.md`(问题、选项、你的建议、哪张单在等),那张单挪 `blocked/`,接着做别的。
|
|
153
|
-
- 每收尾一张单,在 `_receipts/progress.md`
|
|
162
|
+
- 每收尾一张单,在 `_receipts/progress.md` 加一行:`时间 | 工单 | 结果 | 施工token(万) | effort | Lead轮数 | 备注`。每周汇总一次给用户(施工总token、Lead轮数、重做率、审查命中率)。
|
|
154
163
|
- 会话随时可能中断。每张单做完就更新 `queue.md` 和 `progress.md`;会话开始时若 `handoff.md` 里还有活,先读这两个文件再继续。
|
|
155
164
|
- **不扩大范围。** 新想法写进 `queue.md` 当提议,不开工。
|
|
156
165
|
- 停下条件:队列做完;只剩卡住的单;到预算或停止时间;**同一张单失败两次**(挪 `blocked/` 写明原因)。
|