deepseek-foreman 0.2.0 → 0.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.en.md +43 -1
- package/README.md +43 -1
- package/package.json +3 -2
- package/persona.example.md +23 -0
package/README.en.md
CHANGED
|
@@ -4,10 +4,18 @@ English | [中文](README.md)
|
|
|
4
4
|
|
|
5
5
|
# deepseek-foreman
|
|
6
6
|
|
|
7
|
-
Write work as tickets and dispatch them to **cheaper models**; the Lead only breaks down the work, dispatches it, re-runs the acceptance commands personally,
|
|
7
|
+
**A strong model as foreman, cheap models doing the work, another vendor reviewing — multi-model teamwork that is cheaper and better.** Write work as tickets and dispatch them to **cheaper models**; the Lead only breaks down the work, dispatches it, re-runs the acceptance commands personally, and verifies every finding one by one. When nobody is around, it keeps working through the queue.
|
|
8
8
|
|
|
9
9
|
Built for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (`dsh`).
|
|
10
10
|
|
|
11
|
+
**Beginners: three steps** (no cordis / YAML / CLI experience needed):
|
|
12
|
+
|
|
13
|
+
1. Install dsh with routes for ≥2 different vendors;
|
|
14
|
+
2. add `deepseek-foreman` from the plugin page;
|
|
15
|
+
3. open a new session and say "**help me set up foreman — I have models from \<two vendors\>**".
|
|
16
|
+
|
|
17
|
+
> **Version 0.2.0 of this project was itself iterated to completion using this very system** — 19 tickets overnight; see [docs/dogfooding.md](docs/dogfooding.md).
|
|
18
|
+
|
|
11
19
|
## Origin
|
|
12
20
|
|
|
13
21
|
The design and the ticket/receipt contract are ported from [yanauto/opus-manager](https://github.com/yanauto/opus-manager) (MIT, Copyright (c) 2026 yanauto). Upstream is a **Claude Code skill**; this project is its **dsh port**. Upstream validated the flow over 8 weeks, 13 repositories and 360 tickets; this project keeps its directory contract, its acceptance criteria, and the principle that "a receipt is a claim, not evidence".
|
|
@@ -130,6 +138,40 @@ A bare entry `- id: ... / name: ...` is interpreted as **overriding an existing
|
|
|
130
138
|
|
|
131
139
|
Also: **hand-editing a bundle's patch file does not notify the running Host** (changes made outside the manager "announce nothing"). After editing, toggle the switch off and on again in the plugin page, or restart the app, for that layer to be reapplied.
|
|
132
140
|
|
|
141
|
+
## Field data
|
|
142
|
+
|
|
143
|
+
The full sprint ran in one night (2026-09-30 23:00 – 2026-10-01 12:00); the numbers below come from that run's tickets and receipts, with the field report in [docs/dogfooding.md](docs/dogfooding.md).
|
|
144
|
+
|
|
145
|
+
| Item | Value |
|
|
146
|
+
|---|---|
|
|
147
|
+
| Environment | dsh desktop 0.2.0-rc.2, macOS (Intel iMac), project on an ExFAT volume |
|
|
148
|
+
| Lead | Kimi K3 (break down the work, dispatch, personally re-run acceptance, verify item by item) |
|
|
149
|
+
| Builders | DeepSeek-V4.1-Flash (T001–T108, National Day half-price window), MiMo-V2.6-Flash (from T201, mainstay) |
|
|
150
|
+
| Cross-vendor review | Kimi K3; at wrap-up MiMo-V2.6-Pro was added to review K3's output |
|
|
151
|
+
|
|
152
|
+
- All 19 tickets completed the full loop: dispatch → Lead re-runs acceptance → cross-vendor review → item-by-item verification → fix.
|
|
153
|
+
- Self-checks grew from 20 to 110, all green; npm 0.2.0 published + GitHub made public + CI green on its first run.
|
|
154
|
+
- 20+ cross-vendor review findings in total: about 19 confirmed and all fixed, 1 rejected by the Lead after verification; including one high-severity finding (private route names nearly shipped open source).
|
|
155
|
+
- 2 "109 green offline, only blows up live" incidents were caught by the process and fixed.
|
|
156
|
+
|
|
157
|
+
**Token spend** (source: screenshot of the dsh subagent panel; measurement base = cumulative tokens of that build session): build tickets ran 330K–590K tokens each, verification calls 6K–16K tokens each. All build tokens were spent on cheap models — the Lead's context only holds tickets, receipts and review comments, never source code; that is the cost-saving mechanism. No strict A/B experiment was run, so no percentage claims are made; readers can check every line item in the dsh subagent panel themselves.
|
|
158
|
+
|
|
159
|
+
Per-plugin contribution and the two companion pieces are covered under [Works with](#works-with) below.
|
|
160
|
+
|
|
161
|
+
## Works with
|
|
162
|
+
|
|
163
|
+
The three pieces cover one segment each and add up:
|
|
164
|
+
|
|
165
|
+
| Piece | What it does | How to get it |
|
|
166
|
+
|---|---|---|
|
|
167
|
+
| **deepseek-foreman (this package)** | the strong model foremen: dispatch, personally re-run acceptance, verify item by item; cheap models build; another vendor reviews read-only. `pick_route`'s four hard constraints + physical read-only `subagent_readonly` | **included in this package** (takes effect on install) |
|
|
168
|
+
| **Concise-output persona** ([persona.example.md](persona.example.md)) | ponytail-style output discipline: check yourself level by level before acting (don't do it if it can be skipped, reuse before writing new, one line before ten), conclusion first, one-or-two-sentence reports | **included, opt-in**: select-all copy the template → the `personaPrefix` field of the system-prompt plugin in dsh settings → save |
|
|
169
|
+
| **Behaviour-constraint skill** (e.g. superpowers) | guard rails for the agent's behaviour: TDD, acceptance-first rules | **bring your own, optional** |
|
|
170
|
+
|
|
171
|
+
**Mechanism**: the persona hangs on the `personaPrefix` field of dsh's system-prompt plugin (deployment layer) and **child sessions inherit it automatically** (verified: both the `dsh-subagent` source and actual child-session behaviour) — workers receive deployment-level discipline, so it need not be restated per ticket.
|
|
172
|
+
|
|
173
|
+
**Honest accounting**: the overall numbers from the 19 tickets are in [Field data](#field-data) above; per-plugin isolated quantification **has had no A/B control, so no numbers are given** — only mechanisms: build dominates token spend → cheap workers; dialogue and judgement stay terse → persona; quality holds → cross-vendor review + item-by-item verification. The three mechanisms each own one segment, and they stack.
|
|
174
|
+
|
|
133
175
|
## Status
|
|
134
176
|
|
|
135
177
|
Installed into dsh desktop 0.2.0-rc.2 (Intel iMac) and verified live: cold start is normal, `pick_route` is callable by the model, the vision constraint really blocks and returns a fallback; the final `subagent_readonly` state is live-verified too (read-only tool set enforced, delegation tools blocked).
|
package/README.md
CHANGED
|
@@ -4,10 +4,18 @@
|
|
|
4
4
|
|
|
5
5
|
# deepseek-foreman
|
|
6
6
|
|
|
7
|
-
把活写成工单派给**更便宜的模型**去做,Lead
|
|
7
|
+
**让强模型当工头、便宜模型干活、别家厂商审查——多模型协作,又省又好。** 把活写成工单派给**更便宜的模型**去做,Lead 只负责拆活、派活、亲自重跑验收命令、逐条核实。人不在的时候按队列接着干。
|
|
8
8
|
|
|
9
9
|
面向 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(`dsh`)。
|
|
10
10
|
|
|
11
|
+
**小白三步**(没碰过 cordis / YAML / CLI 也能上手):
|
|
12
|
+
|
|
13
|
+
1. 装 dsh,配好 ≥2 家厂商的模型;
|
|
14
|
+
2. 插件页点装 `deepseek-foreman`;
|
|
15
|
+
3. 开新会话,说一句「**帮我配 foreman,我有 \<某两家\> 的模型**」。
|
|
16
|
+
|
|
17
|
+
> **本项目的 0.2.0 就是用这套系统自己迭代完成的**——19 张工单 overnight,详见 [docs/dogfooding.md](docs/dogfooding.md)。
|
|
18
|
+
|
|
11
19
|
## 来源
|
|
12
20
|
|
|
13
21
|
设计与工单/回执契约移植自 [yanauto/opus-manager](https://github.com/yanauto/opus-manager)(MIT,Copyright (c) 2026 yanauto)。上游是一个 **Claude Code skill**;本项目是它的 **dsh 移植版**。上游用 8 周、13 个仓库、360 张工单验证了这套流程,本项目沿用它的目录契约、验收标准和"回执是说法不是证据"原则。
|
|
@@ -128,6 +136,40 @@ bundle 的 `cordis.patch.yml` 里**新增插件行要包在 `insert:` 下**:
|
|
|
128
136
|
|
|
129
137
|
另外:**手改 bundle 的 patch 文件不会通知运行中的 Host**(管理器之外的改动"announces nothing")。改完要在插件页把开关关再开,或重启 app,才会重新应用这一层。
|
|
130
138
|
|
|
139
|
+
## 实测数据
|
|
140
|
+
|
|
141
|
+
一晚(2026-09-30 23:00 – 2026-10-01 12:00)跑完整轮冲刺的记录,数字来自当轮工单与回执,实证见 [docs/dogfooding.md](docs/dogfooding.md)。
|
|
142
|
+
|
|
143
|
+
| 项 | 内容 |
|
|
144
|
+
|---|---|
|
|
145
|
+
| 环境 | dsh 桌面版 0.2.0-rc.2,macOS(Intel iMac),项目在 ExFAT 卷上 |
|
|
146
|
+
| Lead | Kimi K3(拆活、派单、亲自重跑验收、逐条核实) |
|
|
147
|
+
| 施工 | DeepSeek-V4.1-Flash(T001–T108,国庆半价窗口)、MiMo-V2.6-Flash(T201 起,主力) |
|
|
148
|
+
| 异族审查 | Kimi K3;收尾新增 MiMo-V2.6-Pro 审 K3 产物 |
|
|
149
|
+
|
|
150
|
+
- 19 张工单全部走完「派单 → Lead 重跑验收 → 异族审查 → 逐条核实 → 修复」闭环。
|
|
151
|
+
- 自检从 20 项长到 110 项全绿;npm 发布 0.2.0 + GitHub 公开 + CI 首跑即绿。
|
|
152
|
+
- 异族审查累计 20+ 条发现:约 19 条成立全部修复、1 条不成立被 Lead 驳回;含 1 条高危(私有路由名差点开源出去)。
|
|
153
|
+
- 2 起「离线 109 项全绿、live 才炸」的事故被流程抓住并修复。
|
|
154
|
+
|
|
155
|
+
**token 开销**(来源:dsh 子代理面板截图;口径=该次施工会话累计 token):施工类工单 33 万–59 万 token/单,验证类调用 0.6 万–1.6 万 token/次。施工 token 全部发生在便宜模型上,Lead 上下文只放工单/回执/审查意见、不读代码——这就是省钱机制。未做严格 A/B 对照实验,不给百分比承诺;读者可在 dsh 子代理面板自行核对每条开销。
|
|
156
|
+
|
|
157
|
+
逐插件的贡献口径与另外两件搭配件,见下面的[联合使用效果](#联合使用效果)一节。
|
|
158
|
+
|
|
159
|
+
## 联合使用效果
|
|
160
|
+
|
|
161
|
+
三件套合在一起用,各管一段、叠加生效:
|
|
162
|
+
|
|
163
|
+
| 件 | 干什么 | 怎么拿到 |
|
|
164
|
+
|---|---|---|
|
|
165
|
+
| **deepseek-foreman(本包)** | 强模型当工头:派单、亲自重跑验收、逐条核实;便宜模型施工;别家厂商只读审查。`pick_route` 四条硬约束 + `subagent_readonly` 物理只读 | **本包含**(装包即生效) |
|
|
166
|
+
| **精简输出 persona**([persona.example.md](persona.example.md)) | ponytail 式输出纪律:动手前逐级自检(能不做就不做、能复用不新写、能一行不十行)、输出先结论、汇报一两句 | **本包含,可选启用**:全选复制模板 → dsh 设置里 system-prompt 插件的 `personaPrefix` → 保存 |
|
|
167
|
+
| **行为约束类 skill**(如 superpowers) | 给 agent 的行为装护栏:TDD、验收先行那套规矩 | **自备,可选** |
|
|
168
|
+
|
|
169
|
+
**机制**:persona 挂在 dsh 的 system-prompt 插件 `personaPrefix`(deployment 层),**子会话自动继承**(已实证:dsh-subagent 源码 + 子会话实际行为双重印证)——工人拿到的是部署级纪律,每张工单不用按次重述。
|
|
170
|
+
|
|
171
|
+
**口径(诚实)**:19 张工单的整体数据见上文[实测数据](#实测数据);逐插件的孤立量化贡献**没做 A/B 对照,不给数字**,只讲机制——token 大头在实现 → 便宜工人;对话与判断从简 → persona 精简;质量不降 → 异族审查 + 逐条核实。三个机制各管一段,叠加生效。
|
|
172
|
+
|
|
131
173
|
## 状态
|
|
132
174
|
|
|
133
175
|
已装进 dsh 桌面版 0.2.0-rc.2(Intel iMac)并 live 验证:冷启动正常,`pick_route` 可被模型调用,视觉约束会真的拦截并给 fallback;`subagent_readonly` 最终态同样 live 实测通过(只读工具集生效、委派工具被拦死)。
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "deepseek-foreman",
|
|
3
|
-
"version": "0.2.
|
|
3
|
+
"version": "0.2.2",
|
|
4
4
|
"description": "Foreman for DeepSeek Harness: your best model leads, cheaper models build, a rival vendor reviews. Role-to-route adjudication with peak-window, vision, output-size and cross-vendor-review constraints.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "lib/index.js",
|
|
@@ -18,7 +18,8 @@
|
|
|
18
18
|
"lib/index.d.ts",
|
|
19
19
|
"cordis.patch.yml",
|
|
20
20
|
"roles.example.yml",
|
|
21
|
-
"skill/"
|
|
21
|
+
"skill/",
|
|
22
|
+
"persona.example.md"
|
|
22
23
|
],
|
|
23
24
|
"scripts": {
|
|
24
25
|
"build": "tsc -p tsconfig.json"
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
<!--
|
|
2
|
+
用法(可选,但推荐——工人的输出纪律全靠它):
|
|
3
|
+
1. 全选复制下面的规则正文;
|
|
4
|
+
2. 打开 dsh 设置 → system-prompt 插件 → personaPrefix 字段,粘贴进去;
|
|
5
|
+
3. 保存即生效。
|
|
6
|
+
该字段在 deployment 层,子会话自动继承:派出去的工人自动带上这套纪律,
|
|
7
|
+
每张工单不用重述。不装 ponytail / i-have-adhd 之类的插件也能拿到精简输出。
|
|
8
|
+
不装也能用本包,只是工人输出会更啰嗦。
|
|
9
|
+
-->
|
|
10
|
+
|
|
11
|
+
# 精简输出 persona 模板
|
|
12
|
+
|
|
13
|
+
以下 9 条即提示词正文,按需删改:
|
|
14
|
+
|
|
15
|
+
1. **先结论,再理由。** 回答开头就给答案,一句话说得清的不写三段。
|
|
16
|
+
2. **动手前逐级自检**:能不做就不做、能复用不新写、能一行不写十行。
|
|
17
|
+
3. **只做被要求的事**:不加没被要求的功能、抽象、配置项、注释、依赖、重构。
|
|
18
|
+
4. **汇报一两句**:做完只说改了哪个文件、验收结果;不写逐行流水账。
|
|
19
|
+
5. **证据不压缩**:验收命令与输出原样贴出,不为省篇幅删改、不摘抄。
|
|
20
|
+
6. **不编**:不编数字、不编命令输出;没验证的明说没验证。
|
|
21
|
+
7. **范围之外先问**:需求不明就问一句,不猜、不擅自扩范围。
|
|
22
|
+
8. **先证据后结论**:跑过验收命令、见到输出,才可以说「完成 / 通过」。
|
|
23
|
+
9. **默认中文汇报**(对方要求别的语言除外)。
|