deepseek-foreman 0.4.1 → 0.4.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,205 +1,65 @@
1
- [English](README.en.md) | 中文
2
-
3
- > 一个 DeepSeek Harness 社区插件,与 DeepSeek 官方无隶属关系。
4
-
5
1
  # deepseek-foreman
6
2
 
7
- **让强模型当工头、便宜模型干活、别家厂商审查——多模型协作,又省又好。** 把活写成工单派给**更便宜的模型**去做,Lead 只负责拆活、派活、亲自重跑验收命令、逐条核实。人不在的时候按队列接着干。
8
-
9
- 面向 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(`dsh`)。
10
-
11
- **小白三步**(没碰过 cordis / YAML / CLI 也能上手,全程有 setup 向导兜底):
12
-
13
- 1. 装 dsh,配好 ≥2 家厂商的模型;
14
- 2. 插件页点装 `deepseek-foreman`;
15
- 3. 开新会话,说一句「**帮我配 foreman,我有 \<某两家\> 的模型**」。
16
-
17
- > **本项目的 0.2.0 就是用这套系统自己迭代完成的**——19 张工单 overnight,详见 [docs/dogfooding.md](docs/dogfooding.md)。
18
-
19
- ## 来源
20
-
21
- 设计与工单/回执契约移植自 [yanauto/opus-manager](https://github.com/yanauto/opus-manager)(MIT,Copyright (c) 2026 yanauto)。上游是一个 **Claude Code skill**;本项目是它的 **dsh 移植版**。上游用 8 周、13 个仓库、360 张工单验证了这套流程,本项目沿用它的目录契约、验收标准和"回执是说法不是证据"原则。
22
-
23
- ## 致谢
24
-
25
- 感谢 [yanauto/opus-manager](https://github.com/yanauto/opus-manager):工单/回执契约,以及"验收亲自重跑、换一家厂商只读审查、发现逐条核实"这套打法,都是上游用 8 周、13 个仓库、360 张工单实打实跑出来的。本项目只是它的 **dsh 移植版**——目录契约照搬、原则照搬,连"回执是说法不是证据"这句话也照搬。上游与本项目同以 **MIT** 发布,双份版权声明见 [LICENSE](LICENSE)。
26
-
27
- ## 为什么要移植到 dsh
28
-
29
- 上游必须靠外部命令行工具(`pi`、`cursor-agent`、`codex`、`agy`)当工人,因为 **Claude Code 没有"按次给子任务换模型"的入口**——它的模型静态写在 subagent 定义里,外加一个"所有子代理用同一个模型"的全局开关。
30
-
31
- `dsh` 的 `subagent` 工具在**调用那一刻**接受 `provider` / `model` / `reasoning_effort`,并受会话级 `allowedModels` 白名单约束。所以"经理一家、施工一家、审查一家"在 dsh 里是**同一进程内的三次工具调用**,不需要任何外部 CLI。
32
-
33
- ## 与上游的差异
34
-
35
- | | opus-manager | 本项目 |
36
- |---|---|---|
37
- | 宿主 | Claude Code skill | dsh skill |
38
- | 工人 | 外部 CLI 进程(pi / cursor-agent / codex / agy) | `subagent` 工具 + 按次传 `provider`/`model` |
39
- | 首次配置 | 扫描本机装了哪些 CLI、读各自 `--help`、写 `dispatch.sh` / `review.sh` | 读 `allowedModels` 白名单 + `list_subagent_models` 核对,**不写派单脚本** |
40
- | 工人脱离会话 | `nohup` / `Start-Process` | `run_in_background: true` 的 continuable 子会话 |
41
- | 只读审查 | 审查命令只开 `read,grep,find,ls` | 首选 `subagent_readonly` 物理只读(运行时只有 `read`/`grep`/`glob`);会话里没有该工具时降级为审查提示词写明「不改文件,只跑只读命令」+ Lead 审查前后各看一次 `git status` |
42
- | 上下文隔离 | 靠独立进程 | 靠独立 Session(子会话工作不进父对话) |
43
- | 省钱规则 | 一张单一个新会话 | 同上,**外加**:同模型继续干用 `subagent_fork`(保住前缀 KV cache),只有必须换模型才用 `subagent`;子会话继承部署 persona,工人纪律不用按次重述 |
44
- | 无人托管 | `queue.md` + 后台等待 | 同上,可叠 dsh 的 `goal` |
45
-
46
- 工单/回执模板、`_tickets/` 目录契约、验收与核实流程**保持一致**——这部分和宿主无关,是上游最有价值的东西。
47
-
48
- ## 安装
49
-
50
- 从装包到派出第一张工单 ≤ 10 分钟,0–6 步照做即可。逐步细节与故障速查见详版 [docs/install-flow.md](docs/install-flow.md)。
51
-
52
- 0. **dsh 桌面版**,已配好 ≥2 家厂商的 LLM 路由(API key 各家的)。
53
- 1. **`allowedModels` 白名单**(`@deepseek-ai/dsh-tool-subagent/model-selection-settings`):把允许子任务使用的 `provider/model` 列进去,**至少两家不同厂商**(异族审查的硬要求)。改完白名单要开新会话——它是会话快照,详见[前置条件](#前置条件)。
54
- 2. **装包**:dsh 插件页安装 `deepseek-foreman`,安装即生效的**四件套**——
55
- - 挂载 `pick_route`(路由裁决:高峰时段锁、视觉、输出上限、异族审查四条硬约束,见下文「插件」一节);
56
- - 挂载 `subagent_readonly`(只读审查实例:经它派出的子会话在运行时只有 `read` / `grep` / `glob` 三个读工具);
57
- - 自动把工单 skill 装进 `~/.dsh/skills/deepseek-foreman`:先建软链(改仓库即生效),软链被系统拒绝(如 Windows 权限)自动退化为递归拷贝;已装过(软链或目录)一律不动,不会覆盖;两步都失败也不抛错、不影响 dsh 启动,失败原因记进 `setup`(见第 5 步);插件配置 `installSkill: false` 可关闭自动安装;
58
- - 发现没有角色表 → 自动在 `~/.dsh/foreman.roles.yml` 铺一份带中文注释的模板(内容就是包里的 [roles.example.yml](roles.example.yml)),插件进入「未配置」引导态——不报错、不影响 dsh 启动。
59
-
60
- dsh 的 `skill-filesystem` 默认扫 `~/.dsh/skills`(`user-dsh` 根)和 `~/.agents/skills`(`user-agents` 根)。装在这里只给 dsh 用,不污染四工具共享的 `~/.agents/skills`。
61
- 3. **开新会话**:让白名单快照覆盖到新装的实例。
62
- 4. **编辑 `~/.dsh/foreman.roles.yml`**:把模板里「组合 A」(单一厂商全家桶)或「组合 B」(多厂商混合)其中一组的注释解开(**只解一组**),`provider` / `model` 换成第 1 步白名单里已有的路由。**保存即生效**——`pick_route` 每次调用查 mtime 热更新,不用重启;字段逐条有注释,详见下文「配置角色表」一节。
63
- 5. **自检**:对 dsh 说「**调 pick_route 看看 setup**」。`pick_route` 不带 role 即自检模式,返回的 `setup` 一眼看到还缺什么:
64
-
65
- | 字段 | 内容 |
66
- |---|---|
67
- | `skill` | skill 安装结果:`linked` / `copied` / `exists` / `disabled` / `failed: ...` |
68
- | `rolesFile` | 角色表路径、状态(`ok` / `unconfigured` / `error`)、角色数与错误摘要 |
69
- | `allowlist` | 扫各 profile 的 `cordis.patch.yml` 拿到的 `allowedModels` 白名单,与角色表对账;`unmatchedRoles` 列出不在白名单的角色 |
70
- | `hints` | 有问题时的一句人话指引(如「角色 daily-code 的路由不在 allowedModels,把它加进白名单后开新会话」) |
71
-
72
- 6. **走工单**:对 dsh 说「**走工单:把 xxx 项目里的 yyy 做了**」。之后 SOP 自动运转:写工单 → `pick_route` 选路由 → `subagent` 派单 → Lead 重跑验收 → `subagent_readonly` 派异族只读审查 → 逐条核实 → 收尾。用户不在时说「我走了你接着干」进入无人托管——九个环节与队列规矩见 [SKILL.md](skill/deepseek-foreman/SKILL.md) 的「九、无人托管」;派单后出岔子按 [docs/install-flow.md](docs/install-flow.md) 的「故障速查」查。
73
-
74
- ## 配置角色表(`~/.dsh/foreman.roles.yml`)
75
-
76
- 角色表不在 cordis config 里,而在一个外部 YAML 文件,默认 `~/.dsh/foreman.roles.yml`(可用插件的 `rolesFile` 字段换位置)。
77
-
78
- 装包时若发现该文件不存在,插件会**自动铺一份带中文注释的模板**(内容就是包里的 [roles.example.yml](roles.example.yml))并进入「未配置」引导态——`pick_route` 全部返回 `ok:false`,reason 写明文件位置和下一步(「按注释填好 provider/model,保存即生效」),插件本身不报错、不影响 dsh 启动。
79
-
80
- 要改的就是安装第 4 步这一件事:把模板里「组合 A」(单一厂商全家桶)或「组合 B」(多厂商混合)其中一组的注释解开(**只解一组**,同时解开会出现两个顶层 `roles:` 键),再把 `provider` / `model` 换成你自己白名单里已有的路由。改白名单要开新会话(见[前置条件](#前置条件));角色表文件不受这条限制,保存即生效。
81
-
82
- ```bash
83
- $EDITOR ~/.dsh/foreman.roles.yml # 填好 provider/model,保存即生效
84
- ```
85
-
86
- **改角色表不用重启**:`pick_route` 每次调用都查一次该文件的 mtime,变了就重读 + 重校验。文件写坏也不会把插件带崩:保留上一份可用角色表、本次调用返回 `ok:false` 并说明错在哪,改好保存后下一次调用自动恢复。
3
+ [![CI](https://github.com/biantao1108/deepseek-foreman/actions/workflows/ci.yml/badge.svg)](https://github.com/biantao1108/deepseek-foreman/actions)
4
+ [![npm](https://img.shields.io/npm/v/deepseek-foreman)](https://www.npmjs.com/package/deepseek-foreman)
5
+ [![license](https://img.shields.io/badge/license-MIT-green)](LICENSE)
6
+ **English** | [中文](README.zh-CN.md)
87
7
 
88
- ## 前置条件
8
+ > **A strong model as foreman. Cheap models do the work. Another vendor reviews.**
9
+ > Multi-model teamwork for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) — cheaper *and* better.
10
+ > A community plugin, not affiliated with DeepSeek.
89
11
 
90
- `dsh` 桌面版或 CLI,且 `cordis.patch.yml` 里配好:
12
+ ![Metrics dashboard](docs/metrics-2026-10-01.png)
91
13
 
92
- - `@deepseek-ai/dsh-tool-subagent/model-selection-settings` → `enabled: true` + `allowedModels` 至少两条**不同厂商**的路由
93
- - `standard` preset(其 `subagent` 工具实例带 `modelSelectionSettings: true`)
14
+ ## Why
94
15
 
95
- 两者缺一,`subagent` 就不会暴露 `provider`/`model` 入参,本 skill 的派单步骤无法执行。
96
-
97
- **白名单是会话快照**:`allowedModels` 在**新顶层会话创建时取一次**,之后改它不影响已经在跑的会话——改完白名单必须**开新会话**才生效。角色表文件(`~/.dsh/foreman.roles.yml`)不受这条限制,它每次调用都重读,改完保存即时生效。
98
-
99
- ## 插件(v1:路由裁决)
100
-
101
- `src/index.ts` 是 Cordis 插件,注册一个模型可见工具 `pick_route`。它只管 skill 管不了的三件硬约束——这三条写在 markdown 里只能靠模型自觉,写在代码里才能拒:
16
+ | Problem | What foreman does |
17
+ |---|---|
18
+ | Strong models burn quota writing code | Your Lead only **plans, dispatches, verifies** — implementation runs on flash-tier models |
19
+ | Same model writes & reviews its own code | **Cross-vendor review is enforced in runtime**: same vendor is refused, read-only reviewers literally cannot write |
20
+ | Rules in prompts get ignored | Four hard constraints live in code (`pick_route`): peak-hour lock, vision, output cap, cross-vendor |
21
+ | Config editing is scary | One YAML file, hot-reloaded; setup wizard + doctor self-check guide you |
102
22
 
103
- | 约束 | 配置字段 | 拒的时候 |
104
- |---|---|---|
105
- | **高峰时段锁** | `peakWindows` / `peakDays` | 工作日 09:00–18:00 派给 deepseek → 拒,并自动给 `fallback` 角色 |
106
- | **视觉能力** | `vision` | 带截图的活派给纯文本角色 → 拒 |
107
- | **输出上限** | `maxOutputTokens` | 整篇长产出派给上限 <100K 的角色 → 拒 |
108
- | **异族审查** | `vendor` + 调用时传 `review_for` | 审查者和写代码者同厂商 → 拒,并列出别家候选 |
23
+ **Measured** (23 tickets, one overnight run): 93%+ of tokens on cheap models, 24 review findings → 22 fixed, 100% ticket closure — see [metrics](docs/metrics-2026-10-01.md).
109
24
 
110
- 派单本身仍走 dsh 原生 `subagent`(它已支持按次 `provider`/`model`/`reasoning_effort`),工单契约仍是文件。**插件不做的事**:不重造任务板、不接管派单、不碰持久化。
25
+ ## Quick start
111
26
 
112
- 角色表的装法见上面「配置角色表」一节。也仍然兼容老接法:在 cordis config 里直接写 `roles: [...]`,非空时优先,`rolesFile` 被忽略(写法见 [example.cordis.yml](example.cordis.yml);出厂 bundle 的 [cordis.patch.yml](cordis.patch.yml) 里不再带真实角色表)。
27
+ 1. Install dsh, connect **≥2 model vendors**
28
+ 2. Install `deepseek-foreman` in the plugins page
29
+ 3. New session → say **"Set up foreman, I have <Vendor A> and <Vendor B>"**
113
30
 
114
- 本包在 bundle patch 里还挂载了 **`subagent_readonly`**(只读审查实例):经这个实例派出的子会话在**运行时**只有 `read` / `grep` / `glob` 三个读工具——写工具从提示里消失,执行也会被拒,只读审查从「提示词自觉」升级为运行时强制(原生 `subagent` 本身仍没有只读过滤参数,见上表)。只读审查的子会话跑**会话默认路由(通常=经理位模型)**——异族审查覆盖非经理模型写的活(bundle 层不能开按次换模型:standing 挂载需要 preset scope,见 dsh-tool-subagent 源码;k3 自己写的活仍是已知空洞);委派工具(`subagent` / `subagent_fork` / `subagent_readonly` / `workflow`)已被 `toolFilter.deny` 堵住,只读子会话派不出不受限子代理。
31
+ That's it. The wizard (`npx deepseek-foreman-setup`) checks your allowlist, the plugin auto-installs the ticket SOP, and `pick_route` self-checks everything with plain-language hints. Full walkthrough: [docs/install-flow.md](docs/install-flow.md).
115
32
 
116
- 注意:这个工具**能否出现在会话里取决于 dsh 的 preset 层**,本包只保证 bundle patch 这一层写对。若会话里没有 `subagent_readonly`,就按 [SKILL.md](skill/deepseek-foreman/SKILL.md) 的降级方案执行——审查提示词写明「不要改任何文件,只跑只读命令」,Lead 在审查前后各看一次 `git status` 核实。
33
+ ## How it works
117
34
 
118
- ```bash
119
- npm install && npm run build && node test/smoke.mjs # 110 项自检
120
35
  ```
121
-
122
- `Lead` 角色应与 `agent-default-model`(会话实际跑的模型)一致,否则"大脑"名不副实。
123
-
124
- ## 踩过的坑:bundle patch 必须用 `insert:`
125
-
126
- bundle 的 `cordis.patch.yml` 里**新增插件行要包在 `insert:` 下**:
127
-
128
- ```yaml
129
- - insert:
130
- - id: foreman
131
- name: deepseek-foreman
132
- config: { ... }
36
+ you ──► Lead (K3 / Opus-class) plans tickets, re-runs acceptance, verifies findings
37
+ │
38
+ ├──► worker (flash-tier) implements in its own session, 93% of tokens
39
+ │
40
+ └──► reviewer (other vendor) read-only code review → Lead checks each finding
133
41
  ```
134
42
 
135
- 写成裸条目 `- id: ... / name: ...` 会被解释为**按 id 覆盖已存在的行**;组合里没有这一行时**静默不生效**——不报错、不告警,插件页显示「这个插件包不包含任何组件」,模型侧 `NO_TOOL`。官方规范见 [Package and install a plugin](https://github.com/deepseek-ai/deepseek-harness/blob/main/docs/user/develop/basic/publish.md)。
136
-
137
- 另外:**手改 bundle 的 patch 文件不会通知运行中的 Host**(管理器之外的改动"announces nothing")。改完要在插件页把开关关再开,或重启 app,才会重新应用这一层。
138
-
139
- ## 实测数据
140
-
141
- ![效果看板](docs/metrics-2026-10-01.png)
142
-
143
- > 逐项数字、模型分工逻辑、成本口径见 **[可量化测试数据](docs/metrics-2026-10-01.md)**(含图表数据源与复核方法)。
144
-
145
- 一晚(2026-09-30 23:00 – 2026-10-01 12:00)跑完整轮冲刺的记录,数字来自当轮工单与回执,实证见 [docs/dogfooding.md](docs/dogfooding.md)。
146
-
147
- | 项 | 内容 |
148
- |---|---|
149
- | 环境 | dsh 桌面版 0.2.0-rc.2,macOS(Intel iMac),项目在 ExFAT 卷上 |
150
- | Lead | Kimi K3(拆活、派单、亲自重跑验收、逐条核实) |
151
- | 施工 | DeepSeek-V4.1-Flash(T001–T108,国庆半价窗口)、MiMo-V2.6-Flash(T201 起,主力) |
152
- | 异族审查 | Kimi K3;收尾新增 MiMo-V2.6-Pro 审 K3 产物 |
153
-
154
- - 19 张工单全部走完「派单 → Lead 重跑验收 → 异族审查 → 逐条核实 → 修复」闭环。
155
- - 自检从 20 项长到 110 项全绿;npm 发布 0.2.0 + GitHub 公开 + CI 首跑即绿。
156
- - 异族审查累计 20+ 条发现:约 19 条成立全部修复、1 条不成立被 Lead 驳回;含 1 条高危(私有路由名差点开源出去)。
157
- - 2 起「离线 109 项全绿、live 才炸」的事故被流程抓住并修复。
158
-
159
- **token 开销**(来源:dsh 子代理面板截图;口径=该次施工会话累计 token):施工类工单 33 万–59 万 token/单,验证类调用 0.6 万–1.6 万 token/次。施工 token 全部发生在便宜模型上,Lead 上下文只放工单/回执/审查意见、不读代码——这就是省钱机制。未做严格 A/B 对照实验,不给百分比承诺;读者可在 dsh 子代理面板自行核对每条开销。
160
-
161
- 逐插件的贡献口径与另外两件搭配件,见下面的[联合使用效果](#联合使用效果)一节。
43
+ Everything is plain Markdown in your repo: tickets, receipts, review reports. **A receipt is a claim, not evidence** — the Lead re-runs every acceptance command, and `accept_check` fingerprints the receipt against git HEAD.
162
44
 
163
- ## 联合使用效果
45
+ ## Ticket tiers (effort scaling)
164
46
 
165
- 三件套合在一起用,各管一段、叠加生效:
166
-
167
- | 件 | 干什么 | 怎么拿到 |
47
+ | Tier | Flow | Effort |
168
48
  |---|---|---|
169
- | **deepseek-foreman(本包)** | 强模型当工头:派单、亲自重跑验收、逐条核实;便宜模型施工;别家厂商只读审查。`pick_route` 四条硬约束 + `subagent_readonly` 物理只读 | **本包含**(装包即生效) |
170
- | **精简输出 persona**([persona.example.md](persona.example.md)) | ponytail 式输出纪律:动手前逐级自检(能不做就不做、能复用不新写、能一行不十行)、输出先结论、汇报一两句 | **本包含,可选启用**:全选复制模板 → dsh 设置里 system-prompt 插件的 `personaPrefix` → 保存 |
171
- | **行为约束类 skill**(如 superpowers) | 给 agent 的行为装护栏:TDD、验收先行那套规矩 | **自备,可选** |
172
-
173
- **机制**:persona 挂在 dsh 的 system-prompt 插件 `personaPrefix`(deployment 层),**子会话自动继承**(已实证:dsh-subagent 源码 + 子会话实际行为双重印证)——工人拿到的是部署级纪律,每张工单不用按次重述。
174
-
175
- **口径(诚实)**:19 张工单的整体数据见上文[实测数据](#实测数据);逐插件的孤立量化贡献**没做 A/B 对照,不给数字**,只讲机制——token 大头在实现 → 便宜工人;对话与判断从简 → persona 精简;质量不降 → 异族审查 + 逐条核实。三个机制各管一段,叠加生效。
176
-
177
- ## 工单分级(effort scaling)
178
-
179
- | 级别 | 走什么流程 | 派单 effort |
180
- |---|---|---|
181
- | `trivial`(≤10 行/纯文档) | 跳过异族审查,Lead 验收收尾 | low |
182
- | `normal`(默认) | 完整闭环 + 异族便宜审查 | 按 workers.md |
183
- | `critical`(数据/发布/安全) | 双审查 + 物理只读审查 | high;失败 2 次升档重派 |
184
-
185
- 成本台账随每张回执记录(施工 token/effort/wall-clock),`_receipts/progress.md` 周汇总;无人托管有预算闸(到限即停)。分工逻辑与实测数据见 [可量化测试数据](docs/metrics-2026-10-01.md)。
186
-
187
- ## 首次配置向导(R15)
188
-
189
- ```bash
190
- npx deepseek-foreman-setup # 核对白名单、指路角色表、给试跑话术
191
- ```
192
-
193
- 故障自助见 [故障速查](docs/install-flow.md#故障速查10-条与-doctor-hints-对应)(10 条,与 doctor hints 对应)。
49
+ | `trivial` (≤10 lines / docs) | no review, Lead accepts | low |
50
+ | `normal` (default) | full loop + cheap cross-vendor review | per config |
51
+ | `critical` (data / release / security) | double review + physical read-only | high; escalate after 2 failures |
194
52
 
195
- ## 状态
53
+ Cost ledger per receipt, budget caps for unattended runs. Model roles & field data: [metrics](docs/metrics-2026-10-01.md).
196
54
 
197
- 已装进 dsh 桌面版 0.2.0-rc.2(Intel iMac)并 live 验证:冷启动正常,`pick_route` 可被模型调用,视觉约束会真的拦截并给 fallback;`subagent_readonly` 最终态同样 live 实测通过(只读工具集生效、委派工具被拦死)。
55
+ ## Acknowledgements
198
56
 
199
- 工单 SOP(`skill/`)已在真实任务上跑完整轮冲刺:2026-10-01 一晚把 T001–T207(含修复单)全部走完,`_tickets/` / `_receipts/` 契约在本项目立了起来,「派单 → 施工 → 异族只读审查 → 逐条核实 → 修复回单」整条链路闭环,自检从 20 项一路长到 110 项。
57
+ Ported from [yanauto/opus-manager](https://github.com/yanauto/opus-manager) (MIT, © 2026 yanauto) — the ticket/receipt contract and the "re-run acceptance, cross-vendor review, verify every finding" playbook were proven there over 8 weeks / 13 repos / 360 tickets. Dual copyright in [LICENSE](LICENSE).
200
58
 
201
- 实证记录见 [docs/dogfooding.md](docs/dogfooding.md)。
59
+ ## Status
202
60
 
203
- ## 许可
61
+ - ✅ Shipped on [npm](https://www.npmjs.com/package/deepseek-foreman), CI green, 114 self-checks
62
+ - ✅ Built dogfooded: v0.2–v0.4 were iterated through this very ticket system
63
+ - 📄 Details & troubleshooting: [docs/install-flow.md](docs/install-flow.md) · [field reports](docs/dogfooding.md)
204
64
 
205
- MIT。见 [LICENSE](LICENSE)(含上游与本项目两份版权声明)。
65
+ MIT licensed. Community project — issues & PRs welcome.
@@ -0,0 +1,63 @@
1
+ # deepseek-foreman
2
+
3
+ [![CI](https://github.com/biantao1108/deepseek-foreman/actions/workflows/ci.yml/badge.svg)](https://github.com/biantao1108/deepseek-foreman/actions)
4
+ [![npm](https://img.shields.io/npm/v/deepseek-foreman)](https://www.npmjs.com/package/deepseek-foreman)
5
+ [![license](https://img.shields.io/badge/license-MIT-green)](LICENSE)
6
+ [English](README.md) | **中文**
7
+
8
+ > **强模型当工头,便宜模型干活,别家厂商审查。**
9
+ > 面向 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) 的多模型协作插件——又省又好。
10
+ > 社区插件,与 DeepSeek 官方无隶属关系。
11
+
12
+ ![效果看板](docs/metrics-2026-10-01.png)
13
+
14
+ ## 为什么
15
+
16
+ | 痛点 | foreman 怎么解 |
17
+ |---|---|
18
+ | 贵模型把额度烧在写代码上 | Lead 只做**拆单、派单、验收、核实**,实现全在便宜模型 |
19
+ | 自己写的代码自己审 | **异族审查运行时强制**:同厂商被拒,只读审查员物理上写不了 |
20
+ | 提示词规矩会被无视 | 四条硬约束在代码里(`pick_route`):高峰锁/视觉/输出上限/异族 |
21
+ | 配置吓人 | 一个 YAML 文件、改完即生效;setup 向导 + doctor 自检说人话 |
22
+
23
+ **实测**(23 张工单,一晚):93%+ token 在便宜模型、审查 24 条发现 22 条修复、工单 100% 闭环——见[可量化数据](docs/metrics-2026-10-01.md)。
24
+
25
+ ## 三步上手
26
+
27
+ 1. 装 dsh,配好 **≥2 家厂商的模型**
28
+ 2. 插件页点装 `deepseek-foreman`
29
+ 3. 开新会话,说一句:**「帮我配 foreman,我有 \<厂商A\> 和 \<厂商B\> 的模型」**
30
+
31
+ 向导(`npx deepseek-foreman-setup`)核对白名单、插件自动装工单 SOP、`pick_route` 自检全程说人话。完整流程:[docs/install-flow.md](docs/install-flow.md)。
32
+
33
+ ## 工作原理
34
+
35
+ ```
36
+ 你 ──► Lead(K3 级) 拆单、亲自重跑验收、逐条核实
37
+ ├──► 施工(便宜档) 独立会话实现,93% 的 token 在这
38
+ └──► 审查(别家厂商) 只读找问题 → Lead 逐条核实
39
+ ```
40
+
41
+ 工单/回执/审查报告全是仓库里的普通 Markdown。**回执是说法,不是证据**——Lead 重跑每条验收命令,`accept_check` 把回执指纹与 git HEAD 机械核对。
42
+
43
+ ## 工单分级(effort scaling)
44
+
45
+ | 级别 | 流程 | 档位 |
46
+ |---|---|---|
47
+ | `trivial`(≤10 行/文档) | 跳审查,Lead 验收收尾 | low |
48
+ | `normal`(默认) | 完整闭环 + 异族便宜审查 | 按配置 |
49
+ | `critical`(数据/发布/安全) | 双审查 + 物理只读 | high;失败 2 次升档 |
50
+
51
+ 每张回执带成本台账,无人托管有预算闸。模型分工与实测:[可量化数据](docs/metrics-2026-10-01.md)。
52
+
53
+ ## 致谢
54
+
55
+ 移植自 [yanauto/opus-manager](https://github.com/yanauto/opus-manager)(MIT,© 2026 yanauto)——工单/回执契约与「重跑验收、异族审查、逐条核实」打法在上游经 8 周 13 仓库 360 单验证。双版权声明见 [LICENSE](LICENSE)。
56
+
57
+ ## 状态
58
+
59
+ - ✅ 已上 [npm](https://www.npmjs.com/package/deepseek-foreman),CI 绿,114 项自检
60
+ - ✅ 项目自身用这套工单系统迭代完成(v0.2–v0.4)
61
+ - 📄 详情与故障自助:[docs/install-flow.md](docs/install-flow.md) · [实证记录](docs/dogfooding.md)
62
+
63
+ MIT 协议。社区项目,欢迎 Issue 和 PR。
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "deepseek-foreman",
3
- "version": "0.4.1",
3
+ "version": "0.4.2",
4
4
  "description": "Foreman for DeepSeek Harness: your best model leads, cheaper models build, a rival vendor reviews. Role-to-route adjudication with peak-window, vision, output-size and cross-vendor-review constraints.",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",
package/README.en.md DELETED
@@ -1,199 +0,0 @@
1
- English | [中文](README.md)
2
-
3
- > A community plugin for DeepSeek Harness, not affiliated with DeepSeek.
4
-
5
- # deepseek-foreman
6
-
7
- **A strong model as foreman, cheap models doing the work, another vendor reviewing — multi-model teamwork that is cheaper and better.** Write work as tickets and dispatch them to **cheaper models**; the Lead only breaks down the work, dispatches it, re-runs the acceptance commands personally, and verifies every finding one by one. When nobody is around, it keeps working through the queue.
8
-
9
- Built for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (`dsh`).
10
-
11
- **Beginners: three steps** (no cordis / YAML / CLI experience needed):
12
-
13
- 1. Install dsh with routes for ≥2 different vendors;
14
- 2. add `deepseek-foreman` from the plugin page;
15
- 3. open a new session and say "**help me set up foreman — I have models from \<two vendors\>**".
16
-
17
- > **Version 0.2.0 of this project was itself iterated to completion using this very system** — 19 tickets overnight; see [docs/dogfooding.md](docs/dogfooding.md).
18
-
19
- ## Origin
20
-
21
- The design and the ticket/receipt contract are ported from [yanauto/opus-manager](https://github.com/yanauto/opus-manager) (MIT, Copyright (c) 2026 yanauto). Upstream is a **Claude Code skill**; this project is its **dsh port**. Upstream validated the flow over 8 weeks, 13 repositories and 360 tickets; this project keeps its directory contract, its acceptance criteria, and the principle that "a receipt is a claim, not evidence".
22
-
23
- ## Acknowledgements
24
-
25
- Our thanks to [yanauto/opus-manager](https://github.com/yanauto/opus-manager): the ticket/receipt contract and the whole discipline of "re-run acceptance yourself, hand the work to a different vendor for read-only review, verify every finding one by one" were proven upstream over 8 weeks, 13 repositories and 360 tickets. This project is only its **dsh port** — the directory contract is kept, the principles are kept, and so is the line "a receipt is a claim, not evidence". Both are released under **MIT**; the two copyright notices live in [LICENSE](LICENSE).
26
-
27
- ## Why port it to dsh
28
-
29
- Upstream has to use external CLI tools (`pi`, `cursor-agent`, `codex`, `agy`) as workers, because **Claude Code has no entry point for swapping models per sub-task** — the model is hard-coded in the subagent definition, plus a global switch meaning "all subagents use the same model".
30
-
31
- dsh's `subagent` tool accepts `provider` / `model` / `reasoning_effort` **at call time**, constrained by the session-level `allowedModels` allowlist. So "one vendor as manager, one as builder, one as reviewer" in dsh is **three tool calls inside the same process**, with no external CLI required.
32
-
33
- ## Differences from upstream
34
-
35
- | | opus-manager | this project |
36
- |---|---|---|
37
- | Host | Claude Code skill | dsh skill |
38
- | Workers | external CLI processes (pi / cursor-agent / codex / agy) | `subagent` tool + per-call `provider`/`model` |
39
- | First-time setup | scan which CLIs are installed, read each `--help`, write `dispatch.sh` / `review.sh` | read the `allowedModels` allowlist + cross-check with `list_subagent_models`; **no dispatch scripts written** |
40
- | Worker detached from session | `nohup` / `Start-Process` | a continuable child session with `run_in_background: true` |
41
- | Read-only review | the review command opens only `read,grep,find,ls` | `subagent_readonly` first: physical read-only (only `read`/`grep`/`glob` at runtime); when that tool is absent from the session, fall back to a review prompt stating "do not modify files; run read-only commands only" + the Lead checks `git status` before and after |
42
- | Context isolation | separate process | separate Session (child-session work does not enter the parent conversation) |
43
- | Cost rule | one new session per ticket | same, **plus**: keep the same model by continuing with `subagent_fork` (preserving the prefix KV cache); use `subagent` only when the model must change; children inherit the deployment persona, so worker discipline is not restated per dispatch |
44
- | Unattended management | `queue.md` + background wait | same, stackable with dsh's `goal` |
45
-
46
- The ticket/receipt templates, the `_tickets/` directory contract, and the acceptance and verification flow **stay the same** — that part is host-independent, and it is upstream's most valuable piece.
47
-
48
- ## Install
49
-
50
- From package to first dispatched ticket in ≤ 10 minutes — just follow steps 0–6. Step-by-step detail and a troubleshooting table live in the full guide: [docs/install-flow.md](docs/install-flow.md).
51
-
52
- 0. **dsh desktop**, with LLM routes for **≥2 different vendors** configured (an API key for each).
53
- 1. **The `allowedModels` allowlist** (`@deepseek-ai/dsh-tool-subagent/model-selection-settings`): list every `provider/model` subtasks may use, **at least two from different vendors** (a hard requirement for cross-vendor review). After editing the allowlist, open a new session — it is a session snapshot; see [Prerequisites](#prerequisites).
54
- 2. **Install the package**: add `deepseek-foreman` from the dsh plugin page; four things take effect immediately:
55
- - it mounts `pick_route` (route adjudication: peak-hour lock, vision, output ceiling, cross-vendor review — see "Plugin" below);
56
- - it mounts `subagent_readonly` (the read-only review instance: a child session dispatched through it has only the `read` / `grep` / `glob` read tools at runtime);
57
- - it auto-installs the ticket skill into `~/.dsh/skills/deepseek-foreman`: a symlink first (so repo edits stay live), falling back to a recursive copy where symlinks are refused (e.g. Windows privileges); an existing install (symlink or directory) is left untouched — nothing is ever overwritten; if both steps fail it throws nothing and never blocks dsh startup — the reason is recorded in `setup` (step 5); set the plugin's `installSkill: false` to turn auto-install off;
58
- - if no role table is found, it **lays down a template with Chinese comments** at `~/.dsh/foreman.roles.yml` (the packaged [roles.example.yml](roles.example.yml)) and the plugin enters an "unconfigured" guidance state — no errors, dsh startup is never blocked.
59
-
60
- dsh's `skill-filesystem` scans `~/.dsh/skills` (the `user-dsh` root) and `~/.agents/skills` (the `user-agents` root) by default. Installing here serves dsh only and does not pollute the four-tool shared `~/.agents/skills`.
61
- 3. **Open a new session**, so the allowlist snapshot covers the newly installed instance.
62
- 4. **Edit `~/.dsh/foreman.roles.yml`**: uncomment one of the "combination A" (single-vendor, full stack) / "combination B" (multi-vendor mix) groups in the template (**only one**), then replace `provider` / `model` with routes already present in your step-1 allowlist. **Saving takes effect immediately** — every `pick_route` call checks the file's mtime and re-reads on change, no restart; every field has its own comment; see the "Configure the role table" section below.
63
- 5. **Self-check**: tell dsh "**call pick_route and show me setup**". A `pick_route` call without a role is the self-check mode; its `setup` block shows at a glance what is still missing:
64
-
65
- | Field | Contents |
66
- |---|---|
67
- | `skill` | skill install outcome: `linked` / `copied` / `exists` / `disabled` / `failed: ...` |
68
- | `rolesFile` | role-table path, status (`ok` / `unconfigured` / `error`), role count and an error summary |
69
- | `allowlist` | the `allowedModels` pairs scanned from each profile's `cordis.patch.yml`, reconciled against the role table; `unmatchedRoles` lists roles not on the allowlist |
70
- | `hints` | a plain-language hint for each problem found (e.g. "route deepseek/xxx of role daily-code is not in allowedModels; add it to the allowlist, then open a new session") |
71
-
72
- 6. **Work by ticket**: tell dsh "**work this ticket: do yyy in the xxx project**". The SOP then runs itself: write the ticket → `pick_route` picks the route → `subagent` dispatches → the Lead re-runs acceptance → `subagent_readonly` dispatches a cross-vendor read-only review → verify item by item → wrap up. Say "I'm leaving, keep going" to enter unattended mode — the nine-step section and the queue rules live in [SKILL.md](skill/deepseek-foreman/SKILL.md) §9 (无人托管 / unattended management); if something goes wrong after dispatch, use the 故障速查 (troubleshooting) table in [docs/install-flow.md](docs/install-flow.md).
73
-
74
- Note: hints 目前为中文输出 / hints are currently emitted in Chinese.
75
-
76
- ## Configure the role table (`~/.dsh/foreman.roles.yml`)
77
-
78
- The role table does not live in cordis config; it lives in an external YAML file, default `~/.dsh/foreman.roles.yml` (change the location with the plugin's `rolesFile` field).
79
-
80
- On package install, if that file does not exist, the plugin **lays down a template with Chinese comments** (the packaged [roles.example.yml](roles.example.yml)) and enters an "unconfigured" guidance state — every `pick_route` returns `ok:false`, with a reason naming the file location and the next step ("fill in provider/model per the comments; saving takes effect immediately"). The plugin itself neither errors nor blocks dsh startup.
81
-
82
- The only thing to edit is step 4 of the install flow: uncomment one of the "combination A" (single-vendor, full stack) / "combination B" (multi-vendor mix) groups in the template (**only one**; uncommenting both produces two top-level `roles:` keys), then replace `provider` / `model` with routes already present in your own allowlist. Changing the allowlist requires a new session (see [Prerequisites](#prerequisites)); the role-table file is not subject to this rule — saving takes effect immediately.
83
-
84
- ```bash
85
- $EDITOR ~/.dsh/foreman.roles.yml # fill in provider/model; saving takes effect immediately
86
- ```
87
-
88
- **No restart needed to change the role table**: every `pick_route` call checks the file's mtime and re-reads and re-validates on change. A broken file will not crash the plugin either: the last good role table is kept, the call returns `ok:false` explaining what is wrong, and the next call after you fix it recovers automatically.
89
-
90
- ## Prerequisites
91
-
92
- dsh desktop or CLI, with `cordis.patch.yml` configured:
93
-
94
- - `@deepseek-ai/dsh-tool-subagent/model-selection-settings` → `enabled: true` + `allowedModels` with at least two routes from **different vendors**
95
- - the `standard` preset (whose `subagent` tool instance carries `modelSelectionSettings: true`)
96
-
97
- Without both, `subagent` will not expose `provider`/`model` parameters, and this skill's dispatch step cannot run.
98
-
99
- **The allowlist is a session snapshot**: `allowedModels` is read once **when a new top-level session is created**; changing it afterwards does not affect sessions already running — after editing the allowlist you must **open a new session**. The role-table file (`~/.dsh/foreman.roles.yml`) is not subject to this rule; it is re-read on every call, so saving takes effect immediately.
100
-
101
- ## Plugin (v1: route adjudication)
102
-
103
- `src/index.ts` is a Cordis plugin registering one model-visible tool, `pick_route`. It handles only the hard constraints the skill cannot — these could only rely on model self-discipline in markdown, but can be actually refused in code:
104
-
105
- | Constraint | Config field | When refused |
106
- |---|---|---|
107
- | **Peak-hour lock** | `peakWindows` / `peakDays` | Weekday 09:00–18:00 dispatched to deepseek → refused, with an automatic `fallback` role |
108
- | **Vision capability** | `vision` | a job with screenshots dispatched to a text-only role → refused |
109
- | **Output ceiling** | `maxOutputTokens` | a whole long deliverable dispatched to a role capped below 100K → refused |
110
- | **Cross-vendor review** | `vendor` + a call-time `review_for` | reviewer and code author from the same vendor → refused, with other-vendor candidates listed |
111
-
112
- Dispatch itself still goes through dsh's native `subagent` (which already supports per-call `provider`/`model`/`reasoning_effort`); the ticket contract is still file-based. **What the plugin does not do**: rebuild a task board, take over dispatch, or touch persistence.
113
-
114
- The role table is configured in the "Configure the role table" section above. The old wiring still works too: write `roles: [...]` directly in cordis config; when non-empty it wins and `rolesFile` is ignored (see [example.cordis.yml](example.cordis.yml); the shipped bundle's [cordis.patch.yml](cordis.patch.yml) no longer carries a real role table).
115
-
116
- The bundle patch of this package also mounts **`subagent_readonly`** (the read-only review instance): a child session dispatched through it has, **at runtime**, only the three read tools `read` / `grep` / `glob` — write tools disappear from the prompt and their execution is refused, so read-only review goes from "prompt-level self-discipline" to runtime enforcement (the native `subagent` itself still has no read-only filter parameter; see the table above). the child session of a read-only review runs **the session's default route (usually the lead's model)** — cross-vendor review covers work written by non-lead models (per-call model switching is impossible at the bundle layer: a standing mount needs a preset scope, see the `dsh-tool-subagent` source; work written by k3 itself remains a known gap), and the delegation tools (`subagent` / `subagent_fork` / `subagent_readonly` / `workflow`) are blocked by `toolFilter.deny` so a read-only child session cannot dispatch an unrestricted subagent.
117
-
118
- Note: whether this tool **appears in a session depends on dsh's preset layer** — this package only guarantees that its own bundle-patch layer is written correctly. If `subagent_readonly` is not present in the session, fall back to the scheme in [SKILL.md](skill/deepseek-foreman/SKILL.md): the review prompt states "do not modify any files; run read-only commands only", and the Lead checks `git status` once before and once after the review.
119
-
120
- ```bash
121
- npm install && npm run build && node test/smoke.mjs # 110 self-checks
122
- ```
123
-
124
- The `Lead` role should match `agent-default-model` (the model the session actually runs as); otherwise the "brain" is misnamed.
125
-
126
- ## Pitfall: bundle patches must use `insert:`
127
-
128
- In a bundle's `cordis.patch.yml`, **a new plugin row must be wrapped under `insert:`**:
129
-
130
- ```yaml
131
- - insert:
132
- - id: foreman
133
- name: deepseek-foreman
134
- config: { ... }
135
- ```
136
-
137
- A bare entry `- id: ... / name: ...` is interpreted as **overriding an existing row by id**; when the composition has no such row it **silently does nothing** — no error, no warning, the plugin page says "this plugin package contains no components", and the model side returns `NO_TOOL`. The official spec is [Package and install a plugin](https://github.com/deepseek-ai/deepseek-harness/blob/main/docs/user/develop/basic/publish.md).
138
-
139
- Also: **hand-editing a bundle's patch file does not notify the running Host** (changes made outside the manager "announce nothing"). After editing, toggle the switch off and on again in the plugin page, or restart the app, for that layer to be reapplied.
140
-
141
- ## Field data
142
-
143
- ![Metrics dashboard](docs/metrics-2026-10-01.png)
144
-
145
- > Per-item numbers, model-role logic and cost caveats: **[Quantified metrics](docs/metrics-2026-10-01.md)** (with data sources and how to reproduce).
146
-
147
- The full sprint ran in one night (2026-09-30 23:00 – 2026-10-01 12:00); the numbers below come from that run's tickets and receipts, with the field report in [docs/dogfooding.md](docs/dogfooding.md).
148
-
149
- | Item | Value |
150
- |---|---|
151
- | Environment | dsh desktop 0.2.0-rc.2, macOS (Intel iMac), project on an ExFAT volume |
152
- | Lead | Kimi K3 (break down the work, dispatch, personally re-run acceptance, verify item by item) |
153
- | Builders | DeepSeek-V4.1-Flash (T001–T108, National Day half-price window), MiMo-V2.6-Flash (from T201, mainstay) |
154
- | Cross-vendor review | Kimi K3; at wrap-up MiMo-V2.6-Pro was added to review K3's output |
155
-
156
- - All 19 tickets completed the full loop: dispatch → Lead re-runs acceptance → cross-vendor review → item-by-item verification → fix.
157
- - Self-checks grew from 20 to 110, all green; npm 0.2.0 published + GitHub made public + CI green on its first run.
158
- - 20+ cross-vendor review findings in total: about 19 confirmed and all fixed, 1 rejected by the Lead after verification; including one high-severity finding (private route names nearly shipped open source).
159
- - 2 "109 green offline, only blows up live" incidents were caught by the process and fixed.
160
-
161
- **Token spend** (source: screenshot of the dsh subagent panel; measurement base = cumulative tokens of that build session): build tickets ran 330K–590K tokens each, verification calls 6K–16K tokens each. All build tokens were spent on cheap models — the Lead's context only holds tickets, receipts and review comments, never source code; that is the cost-saving mechanism. No strict A/B experiment was run, so no percentage claims are made; readers can check every line item in the dsh subagent panel themselves.
162
-
163
- Per-plugin contribution and the two companion pieces are covered under [Works with](#works-with) below.
164
-
165
- ## Works with
166
-
167
- The three pieces cover one segment each and add up:
168
-
169
- | Piece | What it does | How to get it |
170
- |---|---|---|
171
- | **deepseek-foreman (this package)** | the strong model foremen: dispatch, personally re-run acceptance, verify item by item; cheap models build; another vendor reviews read-only. `pick_route`'s four hard constraints + physical read-only `subagent_readonly` | **included in this package** (takes effect on install) |
172
- | **Concise-output persona** ([persona.example.md](persona.example.md)) | ponytail-style output discipline: check yourself level by level before acting (don't do it if it can be skipped, reuse before writing new, one line before ten), conclusion first, one-or-two-sentence reports | **included, opt-in**: select-all copy the template → the `personaPrefix` field of the system-prompt plugin in dsh settings → save |
173
- | **Behaviour-constraint skill** (e.g. superpowers) | guard rails for the agent's behaviour: TDD, acceptance-first rules | **bring your own, optional** |
174
-
175
- **Mechanism**: the persona hangs on the `personaPrefix` field of dsh's system-prompt plugin (deployment layer) and **child sessions inherit it automatically** (verified: both the `dsh-subagent` source and actual child-session behaviour) — workers receive deployment-level discipline, so it need not be restated per ticket.
176
-
177
- **Honest accounting**: the overall numbers from the 19 tickets are in [Field data](#field-data) above; per-plugin isolated quantification **has had no A/B control, so no numbers are given** — only mechanisms: build dominates token spend → cheap workers; dialogue and judgement stay terse → persona; quality holds → cross-vendor review + item-by-item verification. The three mechanisms each own one segment, and they stack.
178
-
179
- ## Ticket tiers (effort scaling)
180
-
181
- | Tier | Flow | Dispatch effort |
182
- |---|---|---|
183
- | `trivial` (≤10 lines/docs) | skip cross-vendor review, Lead accepts | low |
184
- | `normal` (default) | full loop + cheap cross-vendor review | per workers.md |
185
- | `critical` (data/release/security) | double review + physical read-only | high; escalate after 2 failures |
186
-
187
- Every receipt logs a cost ledger (build token/effort/wall-clock); budget caps halt unattended runs. Model-role logic and measured data: [Quantified metrics](docs/metrics-2026-10-01.md).
188
-
189
- ## Status
190
-
191
- Installed into dsh desktop 0.2.0-rc.2 (Intel iMac) and verified live: cold start is normal, `pick_route` is callable by the model, the vision constraint really blocks and returns a fallback; the final `subagent_readonly` state is live-verified too (read-only tool set enforced, delegation tools blocked).
192
-
193
- The ticket SOP (`skill/`) has run a full sprint on real tasks: on 2026-10-01 it took T001–T207 (fix tickets included) through the whole night, the `_tickets/` / `_receipts/` contract is established here, the whole chain (dispatch → build → cross-vendor read-only review → item-by-item verification → fix receipt) is closed, and the self-checks grew from 20 to 110.
194
-
195
- Field report: [docs/dogfooding.md](docs/dogfooding.md).
196
-
197
- ## License
198
-
199
- MIT. See [LICENSE](LICENSE) (containing both the upstream and this project's copyright notices).