deepseek-foreman 0.2.1 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.en.md
CHANGED
|
@@ -4,10 +4,16 @@ English | [中文](README.md)
|
|
|
4
4
|
|
|
5
5
|
# deepseek-foreman
|
|
6
6
|
|
|
7
|
-
Write work as tickets and dispatch them to **cheaper models**; the Lead only breaks down the work, dispatches it, re-runs the acceptance commands personally,
|
|
7
|
+
**A strong model as foreman, cheap models doing the work, another vendor reviewing — multi-model teamwork that is cheaper and better.** Write work as tickets and dispatch them to **cheaper models**; the Lead only breaks down the work, dispatches it, re-runs the acceptance commands personally, and verifies every finding one by one. When nobody is around, it keeps working through the queue.
|
|
8
8
|
|
|
9
9
|
Built for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (`dsh`).
|
|
10
10
|
|
|
11
|
+
**Beginners: three steps** (no cordis / YAML / CLI experience needed):
|
|
12
|
+
|
|
13
|
+
1. Install dsh with routes for ≥2 different vendors;
|
|
14
|
+
2. add `deepseek-foreman` from the plugin page;
|
|
15
|
+
3. open a new session and say "**help me set up foreman — I have models from \<two vendors\>**".
|
|
16
|
+
|
|
11
17
|
> **Version 0.2.0 of this project was itself iterated to completion using this very system** — 19 tickets overnight; see [docs/dogfooding.md](docs/dogfooding.md).
|
|
12
18
|
|
|
13
19
|
## Origin
|
|
@@ -134,6 +140,10 @@ Also: **hand-editing a bundle's patch file does not notify the running Host** (c
|
|
|
134
140
|
|
|
135
141
|
## Field data
|
|
136
142
|
|
|
143
|
+

|
|
144
|
+
|
|
145
|
+
> Per-item numbers, model-role logic and cost caveats: **[Quantified metrics](docs/metrics-2026-10-01.md)** (with data sources and how to reproduce).
|
|
146
|
+
|
|
137
147
|
The full sprint ran in one night (2026-09-30 23:00 – 2026-10-01 12:00); the numbers below come from that run's tickets and receipts, with the field report in [docs/dogfooding.md](docs/dogfooding.md).
|
|
138
148
|
|
|
139
149
|
| Item | Value |
|
|
@@ -150,6 +160,22 @@ The full sprint ran in one night (2026-09-30 23:00 – 2026-10-01 12:00); the nu
|
|
|
150
160
|
|
|
151
161
|
**Token spend** (source: screenshot of the dsh subagent panel; measurement base = cumulative tokens of that build session): build tickets ran 330K–590K tokens each, verification calls 6K–16K tokens each. All build tokens were spent on cheap models — the Lead's context only holds tickets, receipts and review comments, never source code; that is the cost-saving mechanism. No strict A/B experiment was run, so no percentage claims are made; readers can check every line item in the dsh subagent panel themselves.
|
|
152
162
|
|
|
163
|
+
Per-plugin contribution and the two companion pieces are covered under [Works with](#works-with) below.
|
|
164
|
+
|
|
165
|
+
## Works with
|
|
166
|
+
|
|
167
|
+
The three pieces cover one segment each and add up:
|
|
168
|
+
|
|
169
|
+
| Piece | What it does | How to get it |
|
|
170
|
+
|---|---|---|
|
|
171
|
+
| **deepseek-foreman (this package)** | the strong model foremen: dispatch, personally re-run acceptance, verify item by item; cheap models build; another vendor reviews read-only. `pick_route`'s four hard constraints + physical read-only `subagent_readonly` | **included in this package** (takes effect on install) |
|
|
172
|
+
| **Concise-output persona** ([persona.example.md](persona.example.md)) | ponytail-style output discipline: check yourself level by level before acting (don't do it if it can be skipped, reuse before writing new, one line before ten), conclusion first, one-or-two-sentence reports | **included, opt-in**: select-all copy the template → the `personaPrefix` field of the system-prompt plugin in dsh settings → save |
|
|
173
|
+
| **Behaviour-constraint skill** (e.g. superpowers) | guard rails for the agent's behaviour: TDD, acceptance-first rules | **bring your own, optional** |
|
|
174
|
+
|
|
175
|
+
**Mechanism**: the persona hangs on the `personaPrefix` field of dsh's system-prompt plugin (deployment layer) and **child sessions inherit it automatically** (verified: both the `dsh-subagent` source and actual child-session behaviour) — workers receive deployment-level discipline, so it need not be restated per ticket.
|
|
176
|
+
|
|
177
|
+
**Honest accounting**: the overall numbers from the 19 tickets are in [Field data](#field-data) above; per-plugin isolated quantification **has had no A/B control, so no numbers are given** — only mechanisms: build dominates token spend → cheap workers; dialogue and judgement stay terse → persona; quality holds → cross-vendor review + item-by-item verification. The three mechanisms each own one segment, and they stack.
|
|
178
|
+
|
|
153
179
|
## Status
|
|
154
180
|
|
|
155
181
|
Installed into dsh desktop 0.2.0-rc.2 (Intel iMac) and verified live: cold start is normal, `pick_route` is callable by the model, the vision constraint really blocks and returns a fallback; the final `subagent_readonly` state is live-verified too (read-only tool set enforced, delegation tools blocked).
|
package/README.md
CHANGED
|
@@ -4,10 +4,16 @@
|
|
|
4
4
|
|
|
5
5
|
# deepseek-foreman
|
|
6
6
|
|
|
7
|
-
把活写成工单派给**更便宜的模型**去做,Lead
|
|
7
|
+
**让强模型当工头、便宜模型干活、别家厂商审查——多模型协作,又省又好。** 把活写成工单派给**更便宜的模型**去做,Lead 只负责拆活、派活、亲自重跑验收命令、逐条核实。人不在的时候按队列接着干。
|
|
8
8
|
|
|
9
9
|
面向 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(`dsh`)。
|
|
10
10
|
|
|
11
|
+
**小白三步**(没碰过 cordis / YAML / CLI 也能上手):
|
|
12
|
+
|
|
13
|
+
1. 装 dsh,配好 ≥2 家厂商的模型;
|
|
14
|
+
2. 插件页点装 `deepseek-foreman`;
|
|
15
|
+
3. 开新会话,说一句「**帮我配 foreman,我有 \<某两家\> 的模型**」。
|
|
16
|
+
|
|
11
17
|
> **本项目的 0.2.0 就是用这套系统自己迭代完成的**——19 张工单 overnight,详见 [docs/dogfooding.md](docs/dogfooding.md)。
|
|
12
18
|
|
|
13
19
|
## 来源
|
|
@@ -132,6 +138,10 @@ bundle 的 `cordis.patch.yml` 里**新增插件行要包在 `insert:` 下**:
|
|
|
132
138
|
|
|
133
139
|
## 实测数据
|
|
134
140
|
|
|
141
|
+

|
|
142
|
+
|
|
143
|
+
> 逐项数字、模型分工逻辑、成本口径见 **[可量化测试数据](docs/metrics-2026-10-01.md)**(含图表数据源与复核方法)。
|
|
144
|
+
|
|
135
145
|
一晚(2026-09-30 23:00 – 2026-10-01 12:00)跑完整轮冲刺的记录,数字来自当轮工单与回执,实证见 [docs/dogfooding.md](docs/dogfooding.md)。
|
|
136
146
|
|
|
137
147
|
| 项 | 内容 |
|
|
@@ -148,6 +158,22 @@ bundle 的 `cordis.patch.yml` 里**新增插件行要包在 `insert:` 下**:
|
|
|
148
158
|
|
|
149
159
|
**token 开销**(来源:dsh 子代理面板截图;口径=该次施工会话累计 token):施工类工单 33 万–59 万 token/单,验证类调用 0.6 万–1.6 万 token/次。施工 token 全部发生在便宜模型上,Lead 上下文只放工单/回执/审查意见、不读代码——这就是省钱机制。未做严格 A/B 对照实验,不给百分比承诺;读者可在 dsh 子代理面板自行核对每条开销。
|
|
150
160
|
|
|
161
|
+
逐插件的贡献口径与另外两件搭配件,见下面的[联合使用效果](#联合使用效果)一节。
|
|
162
|
+
|
|
163
|
+
## 联合使用效果
|
|
164
|
+
|
|
165
|
+
三件套合在一起用,各管一段、叠加生效:
|
|
166
|
+
|
|
167
|
+
| 件 | 干什么 | 怎么拿到 |
|
|
168
|
+
|---|---|---|
|
|
169
|
+
| **deepseek-foreman(本包)** | 强模型当工头:派单、亲自重跑验收、逐条核实;便宜模型施工;别家厂商只读审查。`pick_route` 四条硬约束 + `subagent_readonly` 物理只读 | **本包含**(装包即生效) |
|
|
170
|
+
| **精简输出 persona**([persona.example.md](persona.example.md)) | ponytail 式输出纪律:动手前逐级自检(能不做就不做、能复用不新写、能一行不十行)、输出先结论、汇报一两句 | **本包含,可选启用**:全选复制模板 → dsh 设置里 system-prompt 插件的 `personaPrefix` → 保存 |
|
|
171
|
+
| **行为约束类 skill**(如 superpowers) | 给 agent 的行为装护栏:TDD、验收先行那套规矩 | **自备,可选** |
|
|
172
|
+
|
|
173
|
+
**机制**:persona 挂在 dsh 的 system-prompt 插件 `personaPrefix`(deployment 层),**子会话自动继承**(已实证:dsh-subagent 源码 + 子会话实际行为双重印证)——工人拿到的是部署级纪律,每张工单不用按次重述。
|
|
174
|
+
|
|
175
|
+
**口径(诚实)**:19 张工单的整体数据见上文[实测数据](#实测数据);逐插件的孤立量化贡献**没做 A/B 对照,不给数字**,只讲机制——token 大头在实现 → 便宜工人;对话与判断从简 → persona 精简;质量不降 → 异族审查 + 逐条核实。三个机制各管一段,叠加生效。
|
|
176
|
+
|
|
151
177
|
## 状态
|
|
152
178
|
|
|
153
179
|
已装进 dsh 桌面版 0.2.0-rc.2(Intel iMac)并 live 验证:冷启动正常,`pick_route` 可被模型调用,视觉约束会真的拦截并给 fallback;`subagent_readonly` 最终态同样 live 实测通过(只读工具集生效、委派工具被拦死)。
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "deepseek-foreman",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.0",
|
|
4
4
|
"description": "Foreman for DeepSeek Harness: your best model leads, cheaper models build, a rival vendor reviews. Role-to-route adjudication with peak-window, vision, output-size and cross-vendor-review constraints.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "lib/index.js",
|
|
@@ -18,7 +18,8 @@
|
|
|
18
18
|
"lib/index.d.ts",
|
|
19
19
|
"cordis.patch.yml",
|
|
20
20
|
"roles.example.yml",
|
|
21
|
-
"skill/"
|
|
21
|
+
"skill/",
|
|
22
|
+
"persona.example.md"
|
|
22
23
|
],
|
|
23
24
|
"scripts": {
|
|
24
25
|
"build": "tsc -p tsconfig.json"
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
<!--
|
|
2
|
+
用法(可选,但推荐——工人的输出纪律全靠它):
|
|
3
|
+
1. 全选复制下面的规则正文;
|
|
4
|
+
2. 打开 dsh 设置 → system-prompt 插件 → personaPrefix 字段,粘贴进去;
|
|
5
|
+
3. 保存即生效。
|
|
6
|
+
该字段在 deployment 层,子会话自动继承:派出去的工人自动带上这套纪律,
|
|
7
|
+
每张工单不用重述。不装 ponytail / i-have-adhd 之类的插件也能拿到精简输出。
|
|
8
|
+
不装也能用本包,只是工人输出会更啰嗦。
|
|
9
|
+
-->
|
|
10
|
+
|
|
11
|
+
# 精简输出 persona 模板
|
|
12
|
+
|
|
13
|
+
以下 9 条即提示词正文,按需删改:
|
|
14
|
+
|
|
15
|
+
1. **先结论,再理由。** 回答开头就给答案,一句话说得清的不写三段。
|
|
16
|
+
2. **动手前逐级自检**:能不做就不做、能复用不新写、能一行不写十行。
|
|
17
|
+
3. **只做被要求的事**:不加没被要求的功能、抽象、配置项、注释、依赖、重构。
|
|
18
|
+
4. **汇报一两句**:做完只说改了哪个文件、验收结果;不写逐行流水账。
|
|
19
|
+
5. **证据不压缩**:验收命令与输出原样贴出,不为省篇幅删改、不摘抄。
|
|
20
|
+
6. **不编**:不编数字、不编命令输出;没验证的明说没验证。
|
|
21
|
+
7. **范围之外先问**:需求不明就问一句,不猜、不擅自扩范围。
|
|
22
|
+
8. **先证据后结论**:跑过验收命令、见到输出,才可以说「完成 / 通过」。
|
|
23
|
+
9. **默认中文汇报**(对方要求别的语言除外)。
|
|
@@ -67,9 +67,16 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
67
67
|
- **工单头写 `worker-route: <provider>/<model>`**,派单时照它传参。
|
|
68
68
|
- **碰线上服务的单,停机要短**:备份、大文件传输、长测试都在停服务之前做完;从停到拉起之间只切文件和做必要校验;所有远程连接设超时。
|
|
69
69
|
- **写两张,派一张。** `open/` 只放现在能跑的;下一张往往取决于上一张回执里的存疑项。
|
|
70
|
+
- **工单分级(effort scaling)**,工单头写 `difficulty:`:
|
|
71
|
+
- `trivial`(纯文档/≤10 行改动):跳过异族审查,Lead 验收即收尾;派单 effort 用 low
|
|
72
|
+
- `normal`(默认):完整闭环(派单→验收→异族审查→核实);effort 按 workers.md
|
|
73
|
+
- `critical`(碰数据/发布/安全/架构):双审查(两家各审一轮);effort high
|
|
74
|
+
- 失败 2 次升级:同角色 effort 升 max 或换高档施工模型重派,再败才挪 blocked
|
|
70
75
|
|
|
71
76
|
## 三、派单
|
|
72
77
|
|
|
78
|
+
**派单前必须调 `pick_route`**(把工单的 worker-route 传给它):拿到的 route 原样传给 subagent;被拒(高峰/视觉/输出上限/同族审查)就照它给的 fallback/alternatives 换,**不允许绕过硬约束自己挑模型**。
|
|
79
|
+
|
|
73
80
|
一次 `subagent` 调用 = 一张单。参数:
|
|
74
81
|
|
|
75
82
|
- `prompt`:照下面这段,**只给绝对路径**
|
|
@@ -102,6 +109,8 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
102
109
|
|
|
103
110
|
**回执是说法,不是证据。**
|
|
104
111
|
|
|
112
|
+
**回执阅读纪律(省 token)**:回执的「1. 做了什么」「4. 成本台账」「5. 风险与存疑」是结论层——Lead 默认只读这三段;「3. 验收证据」原文段只在抽查或存疑指向它时才读。验收输出超过 50 行让工人落盘到 _receipts/logs/,回执只贴命令+首尾 10 行。
|
|
113
|
+
|
|
105
114
|
1. **每条验收命令自己重跑一遍**,和回执里贴的对比。
|
|
106
115
|
2. **看改动范围**:用 git 就看 `git status` 和 `git diff`;不用 git 就逐个读回执列出的文件。是不是只动了工单允许的?有没有删掉不该动的?
|
|
107
116
|
3. **有实物就看实物**:打开页面、调接口、截图。
|
|
@@ -142,7 +151,7 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
142
151
|
**走之前**:本来会中途打断他的事一次问完,一次一件。他人已经走了就**不要停在问题上等**——一律按最保守的理解,他没明确允许的都算「绝不能自己做」。
|
|
143
152
|
|
|
144
153
|
1. 拆成工单,顺序写进 `_tickets/queue.md`。只有现在能跑的单放 `open/`;轮到下一张时再写进去。
|
|
145
|
-
2. 写 `_tickets/handoff.md
|
|
154
|
+
2. 写 `_tickets/handoff.md`:目标;工单顺序;**预算闸(写死数字)**:单张施工 token 上限(默认 500 万)、累计上限、Lead 轮数上限(默认 200 轮);哪些可以不问直接做;哪些绝不能自己做;几点停。**到任一上限立即停派,写报告等用户**。
|
|
146
155
|
3. 告诉他什么会让托管中断,请他走前处理:**电脑不能休眠**;权限确认不能没人点;**模型套餐有用量上限**,到顶就停,已派出去的任务会跑完,重置后要他叫你继续。
|
|
147
156
|
|
|
148
157
|
**他不在时**:
|
|
@@ -150,7 +159,7 @@ _receipts/ 回执、审查报告、进度、报告
|
|
|
150
159
|
- 按队列往下做,不等人说「继续」:派单 → 验收 → 审查 → 核实 → 修复单 → 下一张。
|
|
151
160
|
- 每个在跑的子任务都要有等待手段(`list_agents` 看状态,或后台轮询回执)。
|
|
152
161
|
- 遇到 `handoff.md` 没写到的决定:**不要猜,也不要全停**。记进 `_tickets/decisions.md`(问题、选项、你的建议、哪张单在等),那张单挪 `blocked/`,接着做别的。
|
|
153
|
-
- 每收尾一张单,在 `_receipts/progress.md`
|
|
162
|
+
- 每收尾一张单,在 `_receipts/progress.md` 加一行:`时间 | 工单 | 结果 | 施工token(万) | effort | Lead轮数 | 备注`。每周汇总一次给用户(施工总token、Lead轮数、重做率、审查命中率)。
|
|
154
163
|
- 会话随时可能中断。每张单做完就更新 `queue.md` 和 `progress.md`;会话开始时若 `handoff.md` 里还有活,先读这两个文件再继续。
|
|
155
164
|
- **不扩大范围。** 新想法写进 `queue.md` 当提议,不开工。
|
|
156
165
|
- 停下条件:队列做完;只剩卡住的单;到预算或停止时间;**同一张单失败两次**(挪 `blocked/` 写明原因)。
|