dsh-embedded-workbench 0.8.11 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.en-US.md CHANGED
@@ -8,7 +8,7 @@ Embedded C/C++ firmware toolbox — 4 agents, 8 skills covering FreeRTOS, ISR, N
8
8
  storage, Keil MDK (AC5/AC6), ARMCLANG, HardFault triage, state machines,
9
9
  architecture principles, LVGL patterns, and claim fact-checking.
10
10
 
11
- **Cross-platform** — works with Claude Code, Codex CLI, Cursor, Kimi CLI, OpenCode, and ZCode. Built on the [Agent Skills](https://agentskills.io) open standard.
11
+ **Cross-platform** — works with Claude Code, Codex CLI, Cursor, Kimi CLI, OpenCode, ZCode, and DeepSeek Harness (dsh). Built on the [Agent Skills](https://agentskills.io) open standard.
12
12
 
13
13
  ## Components
14
14
 
@@ -32,8 +32,9 @@ architecture principles, LVGL patterns, and claim fact-checking.
32
32
  | `c-cpp-dev` | Code generation, style, memory layout, refactoring for C/C++ |
33
33
  | `state-machine-design` | State models, retries, timeouts, transition gates, implementation patterns |
34
34
  | `hardfault-triage` | Processor exception triage — fault registers, stack frames, PC-to-source, root-cause classification |
35
+ | `fact-check` | Claim-check fallback: verifies API names, file paths, enum values, counts, and mechanism feasibility against the codebase; used by the Plan Verification Gate when logicprobe is not installed |
35
36
 
36
- `logicprobe` (design-doc & plan claim verification) was **split out into its own plugin** — see [Other Plugins Recommended](#other-plugins-recommended).
37
+ `logicprobe` (design-doc & plan claim verification) was **split out into its own plugin** — see [Other Plugins Recommended](#other-plugins-recommended). When it is not installed, the Plan Verification Gate falls back to this plugin's built-in `fact-check` skill; only behavioral/model claims degrade to manual confirmation.
37
38
 
38
39
  > The skill content is mostly distilled from the author's personal embedded/firmware engineering experience and code-cleanliness discipline, based on real-world pitfalls and engineering constraints.
39
40
 
@@ -84,7 +85,7 @@ Then enable in `~/.claude/settings.json`:
84
85
  Native dsh support ships as a cordis plugin bundle at the repository root (the root `package.json` declares `dsh.bundle`):
85
86
 
86
87
  - The skills are discovered as-is by dsh's `skill-filesystem` provider (Agent Skills open standard) — zero code.
87
- - The bundle injects the first-model-step gate (1% Rule / Red Flags / Plan Verification Gate) into the first model step of every agent session — the dsh-native counterpart of the Claude `SessionStart` hook. It also registers a model-visible catalog entry (`cordis_inspect`).
88
+ - The bundle folds a **trimmed** first-model-step gate (the Plan Verification Gate plus a context-budget rule) into the first model step of every agent session — the dsh-native counterpart of the Claude `SessionStart` hook — and always registers the model-visible catalog entry (`cordis_inspect`). For why the 1% Rule and the Red Flags table left the payload, see [Design trade-offs and feedback](#design-trade-offs-and-feedback).
88
89
  - The 4 custom agents are intentionally not ported — dsh's native subagent tooling covers parallel multi-agent work.
89
90
 
90
91
  Install (native bundle, recommended):
@@ -104,13 +105,32 @@ Restart the profile, then run `dsh --profile web --dump-config`: the `id: embedd
104
105
 
105
106
  ## Usage
106
107
 
107
- The plugin auto-injects a capability notification into the first model step with a skill table, 1% Rule, and Red Flags reinforcement. Skills are loaded on demand:
108
+ Skills load on demand and do not depend on injection:
108
109
 
109
- - Say "use Multi-Agent Workflow" or invoke `Skill("embedded-workbench")` for the full workflow system
110
+ - Invoke `Skill("embedded-workbench")` for the workflow and engineering policies — it picks a light or full path by risk, and does not force fixed stages
110
111
  - Domain skills activate automatically when their `Use when` description matches your task — NOT clauses prevent false triggers (e.g., formatting-only won't load c-cpp-dev)
111
112
  - The agent proactively suggests verification, adversarial probing, and parallel subagents when it detects state machines, behavioral claims, or multi-module tasks
112
113
  - No manual CLAUDE.md configuration required
113
114
 
115
+ The plugin also folds a **trimmed** gate (about 400 tokens) into the first model step, carrying just two things: the Plan Verification Gate, and a context-budget rule (never guess a readout you cannot see; when a large step shows no signal, ask the user to decide). Set `enabled: false` to drop it entirely.
116
+
117
+ ## Design trade-offs and feedback
118
+
119
+ This revision walks back an earlier decision on the evidence, and the reasoning is below — challenge it.
120
+
121
+ **Background.** We checked the official documentation for all 8 supported harnesses one by one (and read the source for Codex CLI). One assumption did not survive: **7 of the 8 expose no context-budget readout to the model at all** (Claude Code, Copilot CLI, Cursor, OpenCode, Kimi CLI, ZCode and dsh show token figures only in the user's interface; only Codex has a `get_context_remaining` tool, and it is off by default). Asking the model to judge "do I have budget to delegate?" therefore had nothing to stand on.
122
+
123
+ **So we changed two things.**
124
+
125
+ 1. **Cost is now managed by trimming, not by switching injection off.** The first-step gate carries only two things: the **Plan Verification Gate** (verify, or tell the user you did not) and a **context-budget rule** (never guess a readout; when a large step shows no signal, ask the user to decide). The payload went from ~1,400 tokens to ~400 (Claude side −72%, dsh side −58%, measured) and it is on by default.
126
+ 2. **The verification gate stays; the enforcement scaffolding goes.** The 1% Rule and the 9-row Red Flags table left the **injected payload** because they are enforcement, and reported experience shows capable models follow that kind of prompt pressure literally — producing rigid phases, unnecessary questions, and six or seven agents on a five-line task at 10–15× overhead (see [obra/superpowers#1120](https://github.com/obra/superpowers/issues/1120), [openai/codex#22005](https://github.com/openai/codex/issues/22005), [#20366](https://github.com/openai/codex/issues/20366)). The full table still lives in `Skill("embedded-workbench")`: the discipline is available on request rather than applied to everyone by default. Workflow selection likewise moved from a fixed agent chain to **risk-proportional** paths.
127
+
128
+ **Deliberately kept.** The Plan Verification Gate is intact (logicprobe → the built-in `fact-check` when it is not installed → tell the user if you used neither). Following [Superpowers Lite](https://github.com/BB-84C/superpowers-lite), safety, permission and **verification** gates are the kind to keep; process ceremony is the kind to scale back.
129
+
130
+ **Known uncertainty.** These budget interfaces change fast and we checked once, on 2026-09-25. Every cell a vendor does not document is marked `UNVERIFIED` in [`platform-tool-mapping.md`](skills/embedded-workbench/references/platform-tool-mapping.md) rather than filled in by analogy.
131
+
132
+ **Disagree?** These are judgement calls, not settled facts — especially "the Red Flags table leaves the payload" and "the gate is on by default". Open an [issue](https://github.com/AmethystLuna/embedded-workbench/issues) with the model tier, harness and counter-example you are working with; we would rather adjust on evidence.
133
+
114
134
  ## Codex CLI
115
135
 
116
136
  This plugin also supports OpenAI Codex CLI. Skills follow the Agent Skills standard and work identically across both platforms. Agents are provided in Codex TOML format under `.codex/agents/`.
@@ -188,7 +208,7 @@ Skills are invoked with `$skill-name`. ZCode also auto-discovers from `.claude/s
188
208
  ## Requirements
189
209
 
190
210
  - Claude Code v2.1+ / Codex CLI latest / Cursor 2.5+ / Kimi CLI latest / OpenCode latest / ZCode 3.0+
191
- - DeepSeek Harness (dsh): dev preview — verified on mainline 2026-08-14 (gate bundle loaded and injected in-session)
211
+ - DeepSeek Harness (dsh): dev preview — verified per release through 0.1.7-rc.2 (2026-09-25; install / mount / start / uninstall and session-log evidence in [DSH-COMPATIBILITY.md](DSH-COMPATIBILITY.md))
192
212
  - No external dependencies
193
213
 
194
214
  ## Configuration
@@ -197,10 +217,10 @@ In DeepSeek Harness, the bundle accepts a small configuration object:
197
217
 
198
218
  | Key | Type | Default | Description |
199
219
  |---|---|---|---|
200
- | `enabled` | boolean | `true` | Set to `false` to disable the session-start gate injection. |
220
+ | `enabled` | boolean | `true` | Set to `false` to drop the first-step gate injection entirely; skill registration is unaffected. |
201
221
  | `gateContent` | string | built-in gate text | Override the text injected into the first model step. |
202
222
 
203
- To change it, override the row by id in your profile's `cordis.patch.yml`:
223
+ To change it, override the row by id in your profile's `cordis.patch.yml` (the example below customises the gate text):
204
224
 
205
225
  ```yaml
206
226
  - insert:
@@ -257,7 +277,7 @@ To report a security vulnerability, do **not** open a public issue. Use the priv
257
277
 
258
278
  | Plugin | Description |
259
279
  |--------|-------------|
260
- | [logicprobe](https://github.com/AmethystLuna/logicprobe) | Claim-verification skill: checks every verifiable claim in design docs, architecture specs, and refactoring plans against the codebase, escalating behavioral claims to executable-model verification. Split out of this plugin; the Plan Verification Gate requires it. |
280
+ | [logicprobe](https://github.com/AmethystLuna/logicprobe) | Claim-verification skill: checks every verifiable claim in design docs, architecture specs, and refactoring plans against the codebase, escalating behavioral claims to executable-model verification. Split out of this plugin; the Plan Verification Gate prefers it and falls back to the built-in `fact-check` skill when it is not installed. Install with `claude plugin install logicprobe@logicprobe` (on dsh: `dsh plugin --profile <name> add dsh-logicprobe`). |
261
281
  | [superpowers](https://github.com/obra/superpowers) | The original agent discipline engine — skill loading enforcement, Red Flags, subagent-driven development. Many of this plugin's agent-compliance patterns (1% Rule, Red Flags, `<SUBAGENT-STOP>`, instruction priority) were adapted from Superpowers. |
262
282
 
263
283
  ## Acknowledgments
package/README.md CHANGED
@@ -5,7 +5,7 @@
5
5
  嵌入式 C/C++ 固件开发工具箱 — 4 个代理、8 个技能,覆盖 FreeRTOS、中断、NVM 存储、Keil
6
6
  MDK(AC5/AC6)、ARMCLANG、HardFault 分析、状态机、架构原则、LVGL 陷阱。
7
7
 
8
- **跨平台** — 支持 Claude Code、Codex CLI、Cursor、Kimi CLI、OpenCode、ZCode。基于 [Agent Skills](https://agentskills.io) 开放标准构建。
8
+ **跨平台** — 支持 Claude Code、Codex CLI、Cursor、Kimi CLI、OpenCode、ZCode、DeepSeek Harness (dsh)。基于 [Agent Skills](https://agentskills.io) 开放标准构建。
9
9
 
10
10
  ## 组件
11
11
 
@@ -29,8 +29,9 @@ MDK(AC5/AC6)、ARMCLANG、HardFault 分析、状态机、架构原则、LVGL
29
29
  | `c-cpp-dev` | C/C++ 代码生成、风格、内存布局、重构 |
30
30
  | `state-machine-design` | 状态模型、重试、超时、转换门控、实现模式 |
31
31
  | `hardfault-triage` | 处理器异常分类 — 故障寄存器、栈帧、PC 定位源码、根因分类 |
32
+ | `fact-check` | 声称核查回退:逐条对照代码库核实 API 名、文件路径、枚举值、数量与机制可行性;logicprobe 未安装时由 Plan Verification Gate 使用 |
32
33
 
33
- `logicprobe`(文档与计划声称核查技能)**已拆分为独立插件** — 见下方[其他插件推荐](#其他插件推荐)。
34
+ `logicprobe`(文档与计划声称核查技能)**已拆分为独立插件** — 见下方[其他插件推荐](#其他插件推荐)。未安装时,Plan Verification Gate 回退到本插件自带的 `fact-check` 技能,只有行为/模型类声称降级为人工确认。
34
35
 
35
36
  > 技能内容大多来自作者个人嵌入式/固件开发工作经验和代码洁癖,按实际工程踩坑与约束沉淀,而非泛泛的模型生成内容。
36
37
 
@@ -81,7 +82,7 @@ git clone https://github.com/AmethystLuna/embedded-workbench.git ~/.claude/plugi
81
82
  原生 dsh 支持以 cordis 插件 bundle 的形式提供,位于**仓库根**(根 `package.json` 声明了 `dsh.bundle`):
82
83
 
83
84
  - 技能遵循 Agent Skills 开放标准,被 dsh 的 `skill-filesystem` provider 原样发现——零代码。
84
- - bundle 将首步门禁(1% Rule / Red Flags / Plan Verification Gate)注入每个 agent 会话的第一个模型步骤——是 Claude `SessionStart` hook 在 dsh 的原生对应物,并注册了模型可见的目录条目(`cordis_inspect`)。
85
+ - bundle 把一段**精简后**的首步门禁(Plan Verification Gate + 上下文预算规则)注入每个 agent 会话的第一个模型步骤(`enabled`,默认开启)——是 Claude `SessionStart` hook 在 dsh 的原生对应物;模型可见的目录条目(`cordis_inspect`)始终注册。为什么载荷里不再有 1% Rule / Red Flags,见[设计取舍与反馈](#设计取舍与反馈)。
85
86
  - 4 个自定义 agent 有意不移植——dsh 原生 subagent 工具已覆盖并行多 agent 工作。
86
87
 
87
88
  安装(原生 bundle,推荐):
@@ -101,13 +102,32 @@ npx -p @deepseek-ai/dsh dsh plugin --profile web add dsh-embedded-workbench
101
102
 
102
103
  ## 使用
103
104
 
104
- 插件在会话首个模型步骤自动注入能力通知(含技能表、1% Rule、Red Flags 强化)。技能按需加载:
105
+ 技能按需加载,不依赖注入:
105
106
 
106
- - 说"用 Multi-Agent Workflow"或调用 `Skill("embedded-workbench")` 加载完整工作流系统
107
+ - 调用 `Skill("embedded-workbench")` 加载工作流与工程策略——技能内部按风险比例选择轻量或完整路径,不强制固定阶段
107
108
  - 领域技能在任务匹配其 `Use when` 描述时自动激活——NOT 子句防止误触发(如纯格式化不会加载 c-cpp-dev)
108
109
  - Agent 在检测到状态机、行为声称或多模块任务时,主动建议验证、对抗探测和并行子代理
109
110
  - 无需手动配置 CLAUDE.md
110
111
 
112
+ 插件还会在会话首个模型步骤注入一段**精简**门禁(约 400 token),只含两件事:Plan Verification Gate,以及上下文预算规则(看不到读数就不要猜;工作量大且无信号时问用户,由用户决断)。设 `enabled: false` 可完全关闭。
113
+
114
+ ## 设计取舍与反馈
115
+
116
+ 这一版做了一次**基于证据的回退**,理由写在下面,欢迎质疑。
117
+
118
+ **背景**:我们逐个核对了 8 个受支持 harness 的官方文档(Codex 还对照了源码),发现一个此前想当然的前提是错的——**8 家里有 7 家根本不向模型暴露任何上下文预算读数**(Claude Code、Copilot CLI、Cursor、OpenCode、Kimi CLI、ZCode、dsh 都只把 token 数给用户界面;只有 Codex 有一个 `get_context_remaining` 工具,且默认关闭)。也就是说,让模型"按预算自行判断要不要委派/切窗口",本身没有依据。
119
+
120
+ **因此分两步处理**:
121
+
122
+ 1. **不再用"关掉注入"控制成本,改为"瘦身"**。首步门禁现在只保留两块:**Plan Verification Gate**(未核查就必须告诉用户,不许静默跳过)和**上下文预算规则**(看不到读数就不要猜;工作量大且无信号时问用户)。载荷从约 1,400 token 降到约 400 token(Claude 侧 −72%,dsh 侧 −58%,实测值),并且默认开启。
123
+ 2. **验证门禁保留,强制执行脚手架退场**。原来的 1% Rule 与 9 行 Red Flags 表从**注入载荷**中移除:它们属于"强制纪律",而社区实证显示这类提示会被能力较强的模型字面执行,产生僵硬阶段、多余提问,以及五行任务拉起六七个 agent 的 10–15× 开销(见 [obra/superpowers#1120](https://github.com/obra/superpowers/issues/1120)、[openai/codex#22005](https://github.com/openai/codex/issues/22005)、[#20366](https://github.com/openai/codex/issues/20366))。完整表格仍保留在 `Skill("embedded-workbench")` 正文里——需要纪律时纪律还在,只是不再对所有人默认施压。工作流选择也从固定 agent 串场改为**按风险比例**。
124
+
125
+ **有意保留的**:Plan Verification Gate 的语义没有削弱(logicprobe → 未安装则内置 `fact-check` → 两者都没用就必须告知用户)。按 [Superpowers Lite](https://github.com/BB-84C/superpowers-lite) 的原则,安全、权限与**验证**门禁是应当保留的一类,按比例裁掉的应该是流程仪式。
126
+
127
+ **已知的不确定**:各 harness 的预算接口变动很快,我们只在 2026-09-25 核对过一次;[`platform-tool-mapping.md`](skills/embedded-workbench/references/platform-tool-mapping.md) 里凡厂商未公开的格子都明确标为 `UNVERIFIED`,没有靠类比填空。
128
+
129
+ **有不同意见?** 这些取舍(尤其"Red Flags 从载荷退场"和"门禁默认开")是可讨论的判断,不是定论。欢迎到 [Issues](https://github.com/AmethystLuna/embedded-workbench/issues) 提出——写清你用的模型档位、harness 和反例,我们倾向按证据调整。
130
+
111
131
  ## Codex CLI
112
132
 
113
133
  本插件同样支持 OpenAI Codex CLI。技能遵循 Agent Skills 标准,跨平台行为一致。代理以 Codex TOML 格式提供于 `.codex/agents/`。
@@ -185,7 +205,7 @@ cp -r embedded-workbench/skills/* .zcode/skills/
185
205
  ## 依赖
186
206
 
187
207
  - Claude Code v2.1+ / Codex CLI 最新版 / Cursor 2.5+ / Kimi CLI 最新版 / OpenCode 最新版 / ZCode 3.0+
188
- - DeepSeek Harness (dsh): dev preview — 已实测 mainline 2026-08-14(gate bundle 加载并注入会话成功)
208
+ - DeepSeek Harness (dsh): dev preview — 已逐版本实测至 0.1.7-rc.2(2026-09-25,install / mount / start / uninstall 与会话日志证据见 [DSH-COMPATIBILITY.md](DSH-COMPATIBILITY.md))
189
209
  - 无外部依赖
190
210
 
191
211
  ## 配置
@@ -194,10 +214,10 @@ cp -r embedded-workbench/skills/* .zcode/skills/
194
214
 
195
215
  | 键 | 类型 | 默认值 | 说明 |
196
216
  |---|---|---|---|
197
- | `enabled` | boolean | `true` | 设为 `false` 可关闭首步 Gate 注入。 |
217
+ | `enabled` | boolean | `true` | 设为 `false` 可完全关闭首步 Gate 注入;技能注册不受影响。 |
198
218
  | `gateContent` | string | 内置 gate 文本 | 覆盖注入到首轮模型上下文中的文本。 |
199
219
 
200
- 在 profile 的 `cordis.patch.yml` 中按 row id 覆盖:
220
+ 在 profile 的 `cordis.patch.yml` 中按 row id 覆盖(下面的例子自定义 Gate 文本):
201
221
 
202
222
  ```yaml
203
223
  - insert:
@@ -254,7 +274,7 @@ bash tests/skill-triggering/run-all.sh
254
274
 
255
275
  | 插件 | 简介 |
256
276
  |------|------|
257
- | [logicprobe](https://github.com/AmethystLuna/logicprobe) | 声称核查技能:逐条核验设计文档、架构规格、重构计划中的可验证声称与代码库是否一致,行为类声称升级为可执行模型验证。自本插件拆分;Plan Verification Gate 依赖它。 |
277
+ | [logicprobe](https://github.com/AmethystLuna/logicprobe) | 声称核查技能:逐条核验设计文档、架构规格、重构计划中的可验证声称与代码库是否一致,行为类声称升级为可执行模型验证。自本插件拆分;Plan Verification Gate 优先使用它,未安装时回退到内置 `fact-check` 技能。安装:`claude plugin install logicprobe@logicprobe`(dsh:`dsh plugin --profile <name> add dsh-logicprobe`)。 |
258
278
  | [superpowers](https://github.com/obra/superpowers) | 原始 agent 纪律引擎——技能加载强制、Red Flags、子代理驱动开发。本插件的多项 agent 合规模式(1% Rule、Red Flags、`<SUBAGENT-STOP>`、指令优先级)均借鉴自 Superpowers。 |
259
279
 
260
280
  ## 致谢
package/cordis.patch.yml CHANGED
@@ -3,6 +3,10 @@
3
3
  # Row ids are stable identity in the config tree; later layers (the user's
4
4
  # profile cordis.patch.yml, $DSH_HOME/cordis.patch.yml, --patch overlays)
5
5
  # override a row by id, replacing the whole `config` (no deep merge).
6
+ #
7
+ # The first-model-step gate is ON by default and deliberately small — it carries
8
+ # the Plan Verification Gate and the context-budget rule, not the 1% Rule /
9
+ # Red Flags enforcement scaffolding. Set `enabled: false` per profile to drop it.
6
10
 
7
11
  - insert:
8
12
  - id: embedded-workbench
package/lib/index.js CHANGED
@@ -1,12 +1,11 @@
1
1
  /**
2
2
  * embedded-workbench — DeepSeek Harness native plugin for the Embedded
3
- * Workbench toolbox. Injects the session-start gate text (1% Rule, Red
4
- * Flags, Plan Verification Gate, skills roster) into the first model step
5
- * of every agent session, mirroring the SessionStart hook the Claude Code
6
- * plugin installs. The 8 skills ship in this package's `skills/` directory
3
+ * Workbench toolbox. The 8 skills ship in this package's `skills/` directory
7
4
  * and are registered at apply time into dsh's `ctx.skills` registry through
8
5
  * the standard filesystem provider, so they appear in every session catalog
9
- * without a manual copy step.
6
+ * without a manual copy step. The plugin also folds a short gate text into the
7
+ * first model step of every agent session, mirroring the SessionStart hook the
8
+ * Claude Code plugin installs.
10
9
  *
11
10
  * Injection listens on agent/pre-step and appends the gate to the FIRST
12
11
  * model step that runs, once per session (guarded by the session's durable
@@ -17,11 +16,14 @@
17
16
  * reminders (skill catalog, AGENTS.md, gate plugins) simply defer this message
18
17
  * to the first step after their promotion, and the history guard re-injects it
19
18
  * there. The default gate text is the dsh-native adaptation of
20
- * `hooks/session-start-content.md`: behavior rules
21
- * (1% Rule / Red Flags / Plan Verification Gate) stay in sync, while
22
- * presentation is adapted to dsh's native skill catalog — no roster table
23
- * (the model sees skills in its catalog) and no install instructions (those
24
- * live in `.dsh/INSTALL.md`). Deployments override via Config.
19
+ * `hooks/session-start-content.md`: the behavior rules stay in sync (the Plan
20
+ * Verification Gate and the context-budget rule), while presentation is adapted
21
+ * to dsh's native skill catalog — no roster table (the model sees skills in its
22
+ * catalog) and no install instructions (those live in `.dsh/INSTALL.md`). The
23
+ * payload is deliberately small: it carries the verification gate and the
24
+ * budget rule, not the 1% Rule / Red Flags enforcement scaffolding, which
25
+ * measurably pushes capable models into rigid phases and unnecessary fan-out.
26
+ * Deployments override via Config.
25
27
  *
26
28
  * @module embedded-workbench-dsh
27
29
  */
@@ -38,32 +40,21 @@ export const inject = ['skills'];
38
40
  // lands on `<package>/skills` regardless of where the package was installed.
39
41
  const SKILLS_DIR = fileURLToPath(new URL('../skills', import.meta.url));
40
42
  const GATE_PLUGIN_ID = 'embedded-workbench';
43
+ /** Producer-owned message source kind declared in `MessageSourceMap` above. */
44
+ const GATE_SOURCE_KIND = 'plugin:embedded-workbench';
41
45
  const DEFAULT_GATE_CONTENT = `<EXTREMELY_IMPORTANT>
42
- Plugin embedded-workbench is active. You have embedded C/C++ firmware development skills — names and "Use when" triggers are in your skill catalog; load them with the skill tool. No custom agents in dsh: use the native subagent tooling for parallel work.
46
+ Plugin embedded-workbench is active: embedded C/C++ firmware development skills are in your skill catalog. Load the one whose "Use when" matches before substantial work, with the skill tool.
43
47
 
44
- **1% Rule**: If there is even a 1% chance a skill applies to your task, invoke it before responding. If the skill turns out to be wrong for the situation, discard it and move on. The cost of loading a skill is trivial compared to the cost of a preventable mistake.
48
+ **Plan Verification Gate**: before calling exit_plan_mode (or presenting a plan for approval), load the logicprobe skill — or the built-in fact-check skill when logicprobe is not installed — and append a "## Plan Verification" block to the plan. If you verify with neither, tell the user the plan is unverified before asking for approval; a silent skip is not an option. "This change is too small to check" and "I already read the code, the paths are right" are the two rationalizations this gate exists to catch.
45
49
 
46
- **Red Flags** — if you think any of these, STOP. You are rationalizing:
50
+ **Context budget**: no token meter is visible to you, so never guess one. Act on what you can see — a pruned or spilled tool result means stop pulling it in whole, and a compaction checkpoint means move durable state into files. When a large step (many sources, a long sweep, several independent areas) shows no such signal, ask the user what to spend context on rather than deciding silently. If nobody can answer, take the reversible option and say so. When you do delegate, prefer \`subagent_fork\` over \`subagent\` if the sub-agent needs context you already built — its summary still lands here.
47
51
 
48
- | You think | Reality |
49
- |-----------|---------|
50
- | "This is just a quick fix" | Quick fixes break things. A 3-line design check costs 30 seconds. |
51
- | "I already understand this code" | You are looking at one file. The blast radius may span 5 modules. |
52
- | "The skill is overkill for this" | Simple things become complex. Check for skills. |
53
- | "Let me explore the codebase first" | Skills tell you HOW to explore. Check first. |
54
- | "I can just read the file directly" | Skills have patterns and pitfalls you will not discover by reading. |
55
- | "I remember this skill content" | Skills evolve. Always load the current version. |
56
- | "I've explored enough, time to exit plan mode" | The exit_plan_mode tool is the verification gate. Have you loaded the logicprobe skill — or, if it is not installed, the built-in fallback fact-check skill? Every plan must pass this gate before exit. |
57
- | "This plan is too simple for logicprobe" | logicprobe auto-classifies depth; the fallback fact-check verifies every claim regardless. You don't decide. |
58
- | "I already read the code, I know the file paths are correct" | Load the logicprobe skill or the fallback fact-check skill, verify each claim, append the "## Plan Verification" block. |
59
-
60
- **Plan Verification Gate**: Before calling exit_plan_mode (or presenting a plan for approval), load the logicprobe skill (a separate plugin) — or, if it is missing from your skill catalog, load the built-in fallback fact-check skill for claim-by-claim verification, tell the user that behavioral/model claims degrade to manual confirmation, and recommend installing logicprobe. If neither is loaded, inform the user "此计划未经核查,是否需要我先做事实核查?" Silent skip is not an option.
61
-
62
- To load workflows and engineering policies: load the embedded-workbench skill.
63
-
64
- **Proactive features**: When you see state machines, protocol refactoring, behavioral claims ("always"/"never"), or multi-module tasks — suggest verification (logicprobe, or the built-in fact-check fallback if logicprobe is not installed), adversarial probing, or parallel subagents BEFORE the user asks. Most users do not know these exist.
52
+ To load the workflows and engineering policies behind these skills: load the embedded-workbench skill.
65
53
  </EXTREMELY_IMPORTANT>`;
66
54
  export const Config = z.object({
55
+ // On by default, but deliberately small: the payload carries the verification
56
+ // gate and the context-budget rule only, so leaving it on costs a few hundred
57
+ // tokens once per session rather than the ~900 of the previous payload.
67
58
  enabled: z.boolean().default(true),
68
59
  gateContent: z.string().default(DEFAULT_GATE_CONTENT),
69
60
  });
@@ -71,7 +62,7 @@ function gateMessage(text) {
71
62
  return createUserMessage({
72
63
  content: [{ type: 'text', text }],
73
64
  // `form` omitted — an undeclared context is the documented default.
74
- source: { kind: 'plugin', plugin: GATE_PLUGIN_ID },
65
+ source: { kind: GATE_SOURCE_KIND },
75
66
  });
76
67
  }
77
68
  function readSessionEvents(session) {
@@ -94,7 +85,7 @@ function inspectProvider(config) {
94
85
  return {
95
86
  manifest: {
96
87
  id: 'embedded-workbench',
97
- description: 'Session-start gate injection for the Embedded Workbench toolbox — folds the 1% Rule / Red Flags / Plan Verification Gate text into the first model step of every agent session.',
88
+ description: 'Session-start gate injection for the Embedded Workbench toolbox — folds the Plan Verification Gate and the context-budget rule into the first model step of every agent session.',
98
89
  methods: [
99
90
  {
100
91
  name: 'status',
@@ -203,6 +194,11 @@ function gateInHistory(session) {
203
194
  if (event.type !== 'user/message')
204
195
  return false;
205
196
  const source = event.data?.source;
206
- return source?.kind === 'plugin' && source.plugin === GATE_PLUGIN_ID;
197
+ if (source === undefined)
198
+ return false;
199
+ // The v4 producer-owned kind, plus the pre-v4 wrapper this bundle wrote
200
+ // before DSH 0.1.7-alpha.1 retired it.
201
+ return source.kind === GATE_SOURCE_KIND
202
+ || (source.kind === 'plugin' && source.plugin === GATE_PLUGIN_ID);
207
203
  });
208
204
  }
@@ -1,12 +1,11 @@
1
1
  /**
2
2
  * embedded-workbench — DeepSeek Harness native plugin for the Embedded
3
- * Workbench toolbox. Injects the session-start gate text (1% Rule, Red
4
- * Flags, Plan Verification Gate, skills roster) into the first model step
5
- * of every agent session, mirroring the SessionStart hook the Claude Code
6
- * plugin installs. The 8 skills ship in this package's `skills/` directory
3
+ * Workbench toolbox. The 8 skills ship in this package's `skills/` directory
7
4
  * and are registered at apply time into dsh's `ctx.skills` registry through
8
5
  * the standard filesystem provider, so they appear in every session catalog
9
- * without a manual copy step.
6
+ * without a manual copy step. The plugin also folds a short gate text into the
7
+ * first model step of every agent session, mirroring the SessionStart hook the
8
+ * Claude Code plugin installs.
10
9
  *
11
10
  * Injection listens on agent/pre-step and appends the gate to the FIRST
12
11
  * model step that runs, once per session (guarded by the session's durable
@@ -17,27 +16,38 @@
17
16
  * reminders (skill catalog, AGENTS.md, gate plugins) simply defer this message
18
17
  * to the first step after their promotion, and the history guard re-injects it
19
18
  * there. The default gate text is the dsh-native adaptation of
20
- * `hooks/session-start-content.md`: behavior rules
21
- * (1% Rule / Red Flags / Plan Verification Gate) stay in sync, while
22
- * presentation is adapted to dsh's native skill catalog — no roster table
23
- * (the model sees skills in its catalog) and no install instructions (those
24
- * live in `.dsh/INSTALL.md`). Deployments override via Config.
19
+ * `hooks/session-start-content.md`: the behavior rules stay in sync (the Plan
20
+ * Verification Gate and the context-budget rule), while presentation is adapted
21
+ * to dsh's native skill catalog — no roster table (the model sees skills in its
22
+ * catalog) and no install instructions (those live in `.dsh/INSTALL.md`). The
23
+ * payload is deliberately small: it carries the verification gate and the
24
+ * budget rule, not the 1% Rule / Red Flags enforcement scaffolding, which
25
+ * measurably pushes capable models into rigid phases and unnecessary fan-out.
26
+ * Deployments override via Config.
25
27
  *
26
28
  * @module embedded-workbench-dsh
27
29
  */
28
30
  import type { Context } from '@deepseek-ai/cordis';
29
31
  import z from '@deepseek-ai/schemastery';
32
+ import type { ContextFormed } from '@deepseek-ai/dsh-llm';
33
+ declare module '@deepseek-ai/dsh-llm' {
34
+ interface MessageSourceMap {
35
+ 'plugin:embedded-workbench': {
36
+ kind: 'plugin:embedded-workbench';
37
+ } & ContextFormed;
38
+ }
39
+ }
30
40
  export declare const name = "embedded-workbench";
31
41
  export declare const inject: string[];
32
42
  export interface Config {
33
43
  enabled: boolean;
34
44
  gateContent: string;
35
45
  }
36
- export declare const Config: z<Schemastery.ObjectS<{
37
- enabled: z<boolean, boolean>;
38
- gateContent: z<string, string>;
39
- }>, Schemastery.ObjectT<{
40
- enabled: z<boolean, boolean>;
41
- gateContent: z<string, string>;
42
- }>>;
46
+ export declare const Config: z<Schemastery.ObjectS<NoInfer<{
47
+ enabled: z<boolean, boolean, "defined">;
48
+ gateContent: z<string, string, "defined">;
49
+ }>>, Schemastery.ObjectT<NoInfer<{
50
+ enabled: z<boolean, boolean, "defined">;
51
+ gateContent: z<string, string, "defined">;
52
+ }>>, "plain">;
43
53
  export declare function apply(ctx: Context, config: Config): void;
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "dsh-embedded-workbench",
3
- "version": "0.8.11",
4
- "description": "Embedded C/C++ firmware development toolbox — 4 agents, 8 skills covering FreeRTOS, ISR, NVM storage, Keil MDK, ARMCLANG, HardFault, state machines, architecture, LVGL patterns, and claim fact-checking. Ships a native DeepSeek Harness (dsh) bundle that injects the session-start gate into the first model step.",
3
+ "version": "0.9.0",
4
+ "description": "Embedded C/C++ firmware development toolbox — 4 agents, 8 skills covering FreeRTOS, ISR, NVM storage, Keil MDK, ARMCLANG, HardFault, state machines, architecture, LVGL patterns, and claim fact-checking. Ships a native DeepSeek Harness (dsh) bundle that folds a short gate — the Plan Verification Gate plus a context-budget rule — into the first model step.",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",
7
7
  "types": "lib/types/index.d.ts",
@@ -32,7 +32,7 @@
32
32
  "patch": "./cordis.patch.yml"
33
33
  },
34
34
  "compatibility": {
35
- "dsh": "^0.1.0-rc.7 || ^0.1.1-rc.1 || ^0.1.2-alpha.2 || ^0.1.2-alpha.3 || ^0.1.2-alpha.4 || ^0.1.2-alpha.5 || ^0.1.2-rc.1 || ^0.1.3-alpha.1 || ^0.1.3-alpha.2 || ^0.1.5-alpha.1 || ^0.1.5-rc.1 || ^0.1.5-alpha.2 || ^0.1.5-rc.2 || ^0.1.6-alpha.1 || ^0.1.6-alpha.2",
35
+ "dsh": "^0.1.0-rc.7 || ^0.1.1-rc.1 || ^0.1.2-alpha.2 || ^0.1.2-alpha.3 || ^0.1.2-alpha.4 || ^0.1.2-alpha.5 || ^0.1.2-rc.1 || ^0.1.3-alpha.1 || ^0.1.3-alpha.2 || ^0.1.5-alpha.1 || ^0.1.5-rc.1 || ^0.1.5-alpha.2 || ^0.1.5-rc.2 || ^0.1.5-rc.3 || ^0.1.6-alpha.1 || ^0.1.6-alpha.2 || ^0.1.7-alpha.1 || ^0.1.7-alpha.2 || ^0.1.7-rc.1 || ^0.1.7-rc.2",
36
36
  "dshReleases": {
37
37
  "0.1.0-rc.7": "compatible",
38
38
  "0.1.0-rc.8": "compatible",
@@ -49,8 +49,13 @@
49
49
  "0.1.5-alpha.2": "compatible",
50
50
  "0.1.5-rc.1": "compatible",
51
51
  "0.1.5-rc.2": "compatible",
52
+ "0.1.5-rc.3": "compatible",
52
53
  "0.1.6-alpha.1": "compatible",
53
- "0.1.6-alpha.2": "compatible"
54
+ "0.1.6-alpha.2": "compatible",
55
+ "0.1.7-alpha.1": "compatible",
56
+ "0.1.7-alpha.2": "compatible",
57
+ "0.1.7-rc.1": "compatible",
58
+ "0.1.7-rc.2": "compatible"
54
59
  },
55
60
  "profiles": [
56
61
  "headless"
@@ -63,15 +68,15 @@
63
68
  "bump": "node scripts/bump-version.mjs"
64
69
  },
65
70
  "peerDependencies": {
66
- "@deepseek-ai/cordis": "^4.0.2",
71
+ "@deepseek-ai/cordis": "^4.0.4",
67
72
  "@deepseek-ai/dsh-agent": "^0.1.0-rc.6",
68
73
  "@deepseek-ai/dsh-llm": "^0.1.0-rc.6",
69
74
  "@deepseek-ai/dsh-session": "^0.1.0-rc.6",
70
75
  "@deepseek-ai/dsh-skill-filesystem": "^0.1.0-rc.8",
71
- "@deepseek-ai/schemastery": "^3.18.2"
76
+ "@deepseek-ai/schemastery": "^3.18.4"
72
77
  },
73
78
  "devDependencies": {
74
- "@deepseek-ai/cordis": "^4.0.2",
79
+ "@deepseek-ai/cordis": "^4.0.4",
75
80
  "@deepseek-ai/dsh-agent": "^0.1.0-rc.6",
76
81
  "@deepseek-ai/dsh-cordis-host-runner": "^0.1.0-rc.8",
77
82
  "@deepseek-ai/dsh-home-paths": "^0.1.0-rc.8",
@@ -81,8 +86,8 @@
81
86
  "@deepseek-ai/dsh-skill": "^0.1.0-rc.8",
82
87
  "@deepseek-ai/dsh-skill-filesystem": "^0.1.0-rc.8",
83
88
  "@deepseek-ai/dsh-timeout": "^0.0.1-rc.1",
84
- "@types/node": "^26.5.0",
85
- "@deepseek-ai/schemastery": "^3.18.2",
89
+ "@deepseek-ai/schemastery": "^3.18.4",
90
+ "@types/node": "^26.6.2",
86
91
  "typescript": "^7.0.2"
87
92
  },
88
93
  "author": {
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: embedded-workbench
3
- description: "Use when starting any non-trivial coding task — loads multi-agent workflows, engineering policies, and principles for embedded C/C++ firmware development. NOT for trivial single-line fixes, formatting-only changes, or read-only queries."
3
+ description: "Use when starting any non-trivial coding task — loads risk-proportional workflows, engineering policies, and principles for embedded C/C++ firmware development. NOT for trivial single-line fixes, formatting-only changes, or read-only queries."
4
4
  ---
5
5
 
6
6
  <SUBAGENT-STOP>
@@ -23,7 +23,7 @@ If a user's CLAUDE.md says "skip design review for hotfixes" and the workflow re
23
23
 
24
24
  ## Platform Adaptation
25
25
 
26
- This plugin's skills and agents use Claude Code tool names (`Read`, `Write`, `Edit`, `Bash`, `Skill()`). If you are NOT on Claude Code, load `references/platform-tool-mapping.md` for the tool name equivalents on your platform (Codex CLI, Cursor, Kimi CLI, OpenCode, ZCode, Copilot CLI).
26
+ This plugin's skills and agents use Claude Code tool names (`Read`, `Write`, `Edit`, `Bash`, `Skill()`). If you are NOT on Claude Code, load `references/platform-tool-mapping.md` for the tool name equivalents on your platform (Codex CLI, Cursor, Kimi CLI, OpenCode, ZCode, Copilot CLI, DeepSeek Harness). On DeepSeek Harness (dsh) the names are lowercase (`read`/`write`/`edit`/`glob`/`grep`/`pwsh`), skills load through the `skill` tool, the plan gate is `exit_plan_mode`, parallel work uses `subagent`/`subagent_fork`, and the 4 custom agents below are not ported.
27
27
 
28
28
  ## Red Flags
29
29
 
@@ -34,8 +34,8 @@ If you catch yourself thinking any of these, STOP — you are rationalizing:
34
34
  | "This is just a quick fix, I don't need a plan" | Quick fixes are the most likely to break something else. A 3-line design check costs 30 seconds. |
35
35
  | "I already understand the architecture" | You're looking at one file. The blast radius may span 5 modules you haven't read. |
36
36
  | "The worker can figure out the details" | The worker has NO context from previous calls. A vague plan = the worker guessing. |
37
- | "I'll review it myself, no need for quality-coordinator" | Self-review catches ~60% of issues. A second pair catches the other 40%. |
38
- | "This change is too small for a Detailed Change Plan" | If it touches more than one function, it needs a plan. Even single-function changes benefit from explicit invariants. |
37
+ | "I'll review it myself, no need for quality-coordinator" | Self-review is enough for a bounded, reversible edit. A second pass earns its cost when the blast radius is unclear or the change is hard to undo. |
38
+ | "This change is too small for a Detailed Change Plan" | Bounded, reversible, single-site edits do not need the ceremony. Crossing a module boundary, a public interface, or a non-obvious failure mode does. Name the invariants either way. |
39
39
  | "I've explored enough, time to exit plan mode" | ExitPlanMode is the verification gate. Have you loaded `Skill("logicprobe")` or, if it is not installed, the built-in fallback `Skill("fact-check")`? Every plan — simple or complex — must pass this gate before exit. |
40
40
  | "This plan is too simple for logicprobe" | logicprobe auto-classifies depth (LIGHTWEIGHT/STANDARD/ESCALATED); the fallback fact-check verifies every claim regardless. You don't decide whether verification is needed. |
41
41
  | "I already read the code, I know the file paths and API names are correct" | Organic verification leaves no audit trail. Load `Skill("logicprobe")` or the fallback `Skill("fact-check")`, verify each claim, append the `## Plan Verification` block. |
@@ -46,65 +46,82 @@ If you catch yourself thinking any of these, STOP — you are rationalizing:
46
46
 
47
47
  - Build context before acting: identify domain → load relevant skills → read key sources → analyze → edit.
48
48
  - **Facts first, code is truth**: verify every document claim (counts, API names, enum values) against the actual codebase with Grep. Design on verified facts, not assumptions.
49
- - Use `Agent(subagent_type: "Explore")` for broad searches instead of chaining Grep/Glob.
49
+ - For a broad search that would otherwise take many Grep/Glob rounds, a read-only `Explore` sub-agent keeps the noise out of this context — but two direct searches are cheaper than dispatching one.
50
50
  - Verify every change with `Bash` compilation or tests before reporting success. No verification = no claim of success.
51
51
  - Reference file locations with line numbers in all reports: `[path/to/file.c#L100-L110]`.
52
52
  - Write project memory to `<workspace>/.github/memory/`, update `MEMORY.md` index. Personal preferences only in `~/.claude/projects/.../memory/`.
53
53
 
54
54
  ---
55
55
 
56
- ## Workflows
56
+ ## Proportionality
57
+
58
+ Match the process to the risk in front of you, not to habit. Start on the lightest path that covers the risk, and escalate only when you actually cross a boundary.
59
+
60
+ | Task shape | Path |
61
+ | --- | --- |
62
+ | Question, read-only investigation, or a bounded reversible edit (typo, constant, one call site) | Do it directly — no plan header, no sub-agent |
63
+ | One module, known repro, local refactor | Lite |
64
+ | Cross-module, public interface, shared-state ownership, new state machine, non-obvious failure/recovery mode | Full |
65
+ | Platform layer, contracts, staged migration | Full + audit matrix |
66
+
67
+ You have crossed into a heavier path when the change spans more than one module, alters a public interface or state ownership, hides a failure mode you cannot describe, or two attempts have not converged. Scaling a bounded edit up costs more than the edit.
57
68
 
58
- Sub-agents are **stateless** — each `Agent()` call is a fresh process. Plan-then-Implement uses two separate spawns: the first produces a Plan, the orchestrator approves it, the second implements. Only worth it when the Plan is specific enough for mechanical execution.
69
+ ## Context Budget
59
70
 
60
- ### Lite Workflow — Single-file fix, small bug, local refactor
71
+ Complexity decides **how much process**; the budget decides **whether to split the window**. Split it for discovery whose detail you will not cite again — a suite, a log sweep, docs, several independent areas. Keep it for phases that share context (plan → implement → test), or when the change is quick and latency matters.
61
72
 
62
- **When**: single file/module, known repro, no cross-module boundaries. Trivial changes (typo, constant) — fix directly.
73
+ **You cannot see your budget**: most harnesses show a token count to the user's interface, not to you, so never guess one. Act on what you do see — a result truncated, pruned, or spilled to a file, or a compaction / checkpoint summary. When a large step shows no signal at all, **ask the user** what to spend context on rather than deciding silently; if nobody can answer (headless run), take the reversible option and say so.
63
74
 
64
- `execution-worker` → Plan → **suggest user run `design-reviewer` or `logicprobe` to verify Plan claims against codebase** → approve → `execution-worker` → implement + verify. Self-check. Uncertain → `quality-coordinator`.
75
+ Delegation is not free: the sub-agent spends its own tokens and its summary still lands here. When it needs context you already built, inherit it (dsh `subagent_fork`; Codex `fork_turns`, default `all`; Claude Code fork mode; Kimi `/btw`) instead of rebuilding it in a prompt. Per-harness markers and tool names: `references/platform-tool-mapping.md`.
76
+
77
+ Delegation is not free: the sub-agent spends its own tokens, its summary still lands in this context, and its window is sized by *its* model, not this one. If the returns would be verbose, ask before delegating. Prefer read-only delegation when the point is to discard detail. When the delegated work needs the context you have already built, inherit it (dsh `subagent_fork`; Codex `fork_turns`, which defaults to `all`; Claude Code forked subagent; Kimi `/btw`) rather than rebuilding it inside a prompt (dsh `subagent`, Claude Code fresh subagent).
78
+
79
+ ## Workflows
65
80
 
66
- Escalate to Multi-Agent when cross-module or two revisions don't converge.
81
+ Sub-agents are **optional equipment, not mandatory stages** — dispatch one when it buys something concrete (a frozen plan, context isolation, an independent review pass), never as ceremony. Each `Agent()` call is stateless: it sees only what its prompt contains. Choose isolation when the detail can be discarded, and inheritance when the sub-agent needs the context you already built (see Context Budget above).
67
82
 
68
- ### Multi-Agent Workflow — Multi-module tasks, full plan/review/closure cycle
83
+ ### Lite — one module, known repro
69
84
 
70
- **When**: cross-module, new interfaces, state ownership changes, needs design package.
85
+ Plan the change, make it, verify it. Self-review is sufficient for a bounded, reversible edit.
71
86
 
72
- `architecture-steward` → design → `design-reviewer` → fact-check → per slice: `execution-worker` (Plan → approve → Implement) → `quality-coordinator` → closure (normal/failure/recovery paths).
87
+ If the plan rests on claims about the codebase (API names, file paths, enum values, counts), verify those claims before implementing — inline for a small change, or with `logicprobe` / the built-in `fact-check` skill. Escalate to Full when the change turns out to cross a boundary, or when two revisions fail to converge.
73
88
 
74
- ### Framework Workflow — Platform layer, contracts, staged migration
89
+ ### Full — cross-module, new interfaces, state-ownership changes
75
90
 
76
- **When**: framework incubation, runtime path mounting, contract/sentinel/audit definition, old/new coexistence.
91
+ Design first (module boundaries, slice breakdown, recovery paths), then per slice: plan → approve → implement → verify → close. Add an independent review pass (`quality-coordinator`, or a `subagent` with a review prompt) when the change is hard to reverse or the blast radius is unclear; for a well-bounded slice, self-review plus verification evidence is enough.
77
92
 
78
- Same as Multi-Agent, plus: design package includes audit matrix + rollback triggers; each slice reports audit delta; quality-coordinator checks audit ledger consistency.
93
+ ### Framework — platform layer, contracts, staged migration
79
94
 
80
- ### Sub-Agent Reference
95
+ Full, plus an audit matrix and rollback triggers; each slice reports its audit delta.
81
96
 
82
- | Agent | Role |
83
- | ------- | ------ |
84
- | `architecture-steward` | Read-only planning: design packages, module boundaries, slice breakdown |
85
- | `design-reviewer` | Design doc fact-check: verifies claims against codebase before implementation |
86
- | `execution-worker` | Plan round → Detailed Change Plan. Implement round → edit + verify |
87
- | `quality-coordinator` | Implementation review: bugs, compliance, closure completeness |
97
+ ### Sub-Agents
98
+
99
+ | Agent | Role | When it earns its cost |
100
+ | ------- | ------ | ------ |
101
+ | `architecture-steward` | Read-only planning: design packages, module boundaries, slice breakdown | A design spanning modules you have not read |
102
+ | `design-reviewer` | Design doc fact-check: verifies claims against codebase before implementation | The plan rests on claims you have not verified |
103
+ | `execution-worker` | Plan round → Detailed Change Plan. Implement round → edit + verify | You want the plan frozen before edits, or a noisy investigation kept out of this context |
104
+ | `quality-coordinator` | Implementation review: bugs, compliance, closure completeness | The change is hard to reverse or the risk surface is wide |
105
+
106
+ These are roles, not gates: the main model can play any of them directly when that is cheaper.
88
107
 
89
108
  ---
90
109
 
91
110
  ## Plan Mode Integration
92
111
 
93
- Claude Code's built-in `EnterPlanMode` / `ExitPlanMode` maps to the **Plan phase** of the Lite and Multi-Agent workflows. Plan mode is a read-only exploration + plan-writing phase — it does NOT exempt you from embedded-workbench verification gates.
112
+ Claude Code's built-in `EnterPlanMode` / `ExitPlanMode` maps to the **plan phase** of the Lite and Full paths. Plan mode is a read-only exploration + plan-writing phase — it does NOT exempt you from embedded-workbench verification gates.
94
113
 
95
114
  ### Plan Verification Gate
96
115
 
97
- > **⚠️ logicprobe 已拆分为独立插件 / moved to a standalone plugin** (v0.6.0): the full verification skill (executable model checks, adversarial probing) now ships in its own plugin — <https://github.com/AmethystLuna/logicprobe>. Install it with `claude plugin install logicprobe@logicprobe` (or clone to `~/.claude/plugins/dev/logicprobe`). This plugin ships a built-in simplified fallback — `Skill("fact-check")` — for claim-by-claim verification when logicprobe is not installed; behavioral/model claims then degrade to manual confirmation.
116
+ > **⚠️ logicprobe 已拆分为独立插件 / moved to a standalone plugin** (v0.6.0): the full verification skill (executable model checks, adversarial probing) now ships in its own plugin — <https://github.com/AmethystLuna/logicprobe> (install commands are in the README's "Other Plugins"). This plugin ships a built-in simplified fallback — `Skill("fact-check")` — for claim-by-claim verification when logicprobe is not installed; behavioral/model claims then degrade to manual confirmation.
98
117
 
99
118
  **Before calling `ExitPlanMode`**, exactly one of the following must happen:
100
119
 
101
- 1. **Load `Skill("logicprobe")`** (standalone plugin — install separately if missing) — the skill classifies depth (LIGHTWEIGHT / STANDARD / ESCALATED), runs verification (including executable model checks), and appends a `## Plan Verification` summary block to the plan file.
102
- 2. **Load `Skill("fact-check")`** (built-in fallback, only when logicprobe is not installed) — verifies every verifiable claim against the codebase with evidence, appends a `## Plan Verification` block marked `fact-check (fallback)`, and tells the user that state-machine/behavioral claims degrade to manual confirmation — recommend installing logicprobe.
103
- 3. **Inform the user** — if you choose not to load either skill, you MUST say: *"此计划未经核查。是否需要我在审批前运行事实核查?(This plan has not been fact-verified. Would you like me to run verification before approving?)"* The user must have the option to request verification before approving.
104
-
105
- Silent skip is not an option. Either verify, or tell the user you didn't.
120
+ 1. **Load `Skill("logicprobe")`** (standalone plugin) — it classifies depth (LIGHTWEIGHT / STANDARD / ESCALATED), runs the verification including executable model checks, and appends a `## Plan Verification` summary block to the plan.
121
+ 2. **Load `Skill("fact-check")`** (built-in fallback, only when logicprobe is not installed) — verifies every verifiable claim against the codebase with evidence, appends a `## Plan Verification` block marked `fact-check (fallback)`, and tells the user that state-machine/behavioral claims degrade to manual confirmation.
122
+ 3. **Inform the user** — if you load neither, say: *"此计划未经核查。是否需要我在审批前运行事实核查?(This plan has not been fact-verified. Would you like me to run verification before approving?)"*
106
123
 
107
- Plan mode permits `Read`, `Glob`, `Grep`, and `Skill` calls — all verification executes within plan mode before exit.
124
+ Silent skip is not an option. Plan mode permits `Read`, `Glob`, `Grep`, and `Skill` calls, so all of this executes before exit.
108
125
 
109
126
  ---
110
127
 
@@ -113,20 +130,18 @@ Plan mode permits `Read`, `Glob`, `Grep`, and `Skill` calls — all verification
113
130
  <HARD-GATE>
114
131
  ### Approval Gate
115
132
 
116
- - Implementation-bearing slices MUST produce a Detailed Change Plan before editing. No exceptions.
117
- - Plan must include: objective, entry point, intended files, change shape, invariants, risks, validation, stop conditions.
118
- - If execution reveals facts that change scope/boundaries/acceptance/verification surface, pause and require re-approval.
119
- - Do NOT skip the plan phase because "the change is obvious" or "I've done this before."
133
+ - A slice that crosses a module boundary, changes a public interface or state ownership, or carries a non-obvious failure/recovery mode MUST have its plan written and approved before editing: objective, entry point, intended files, change shape, invariants, risks, validation, stop conditions.
134
+ - Bounded, reversible, single-site edits proceed without the ceremony — state the intent, make the change, verify the result.
135
+ - If execution reveals facts that change scope, boundaries, acceptance, or the verification surface, pause and re-approve.
120
136
  </HARD-GATE>
121
137
 
122
138
  <HARD-GATE>
123
139
  ### Closure Gate
124
140
 
125
- - Slice is NOT done until: implementation intent + verification evidence + residual risks are all explicit.
141
+ - Work is not done until implementation intent, verification evidence, and residual risks are all explicit.
126
142
  - Skipped checks MUST record a concrete reason. "Looks good" is not a reason.
127
- - For fault/recovery scenarios, MUST cover normal, failure, and recovery paths.
128
- - Documentation and memory updates MUST be completed or explicitly skipped with reason.
129
- - Do NOT call a slice closed if verification, documentation impact, or audit deltas are unclear.
143
+ - For fault/recovery scenarios, cover the normal, failure, and recovery paths.
144
+ - Documentation and memory updates are completed, or explicitly skipped with a reason.
130
145
  </HARD-GATE>
131
146
 
132
147
  ### Escalation Triggers
@@ -140,7 +155,7 @@ Plan mode permits `Read`, `Glob`, `Grep`, and `Skill` calls — all verification
140
155
  Sub-agents are **stateless with no implicit context inheritance** — each spawn only gets what's in its prompt:
141
156
 
142
157
  - **Explicit prompt construction**: put design conclusions, approved Plans, review findings directly in the prompt. Do NOT assume the agent "remembers" previous conversations.
143
- - **Plan is the key handoff artifact**: between Design → Plan round → Implement round, the Detailed Change Plan and review verdicts are the only bridge. Vague Plans = the next agent guessing.
158
+ - **Plan is the key handoff artifact**: when you do dispatch sub-agents, the Detailed Change Plan and review verdicts are the only bridge between Design → Plan round → Implement round. Vague Plans = the next agent guessing.
144
159
  - **Pass only what's needed**: Design phase doesn't need full source code. Implement phase doesn't need the full Audit Matrix.
145
160
  - **Memory for cross-session persistence**: rules, pitfalls, constraints that need to survive across sessions go in `<workspace>/.github/memory/`. In-session coordination stays in chat.
146
161
  - **Long content via path references**: if context is too large, write long content to workspace docs and put only the path in the prompt. Let the agent Read it.
@@ -179,7 +194,7 @@ When multiple skills could apply, use this order:
179
194
  "Add retry logic" → state-machine-design first, then c-cpp-dev for implementation.
180
195
  "Review this design" → logicprobe first (or the built-in fact-check fallback if logicprobe is not installed), then escalate findings to design-reviewer agent.
181
196
 
182
- **Cross-domain links**: load secondary skills ONLY when the primary skill's findings indicate they are needed. Don't pre-load. `hardfault-triage` ↔ `keil-mdk-build` (.map file bridge — load keil-mdk-build only if .map analysis is needed). `hardfault-triage` ↔ `debug-methodology` (root-cause analysis — load debug-methodology only if the fault cause is complex). `embedded-firmware-dev` ↔ `state-machine-design` (state transitions — load state-machine-design only if state logic is involved). `embedded-firmware-dev` ↔ `debug-methodology` (debugging process). `logicprobe` ↔ `design-reviewer` agent (design doc review, logic verification). `logicprobe` ↔ `state-machine-design` (behavioral claim probing). `logicprobe` ↔ `fact-check` (built-in fallback when the logicprobe plugin is not installed).
197
+ **Cross-domain links**: load a secondary skill only when the primary skill's findings call for it — don't pre-load. Each skill's own `Use when` and NOT clauses already tell you when it applies.
183
198
 
184
199
  ## Domain Skills
185
200
 
@@ -198,45 +213,15 @@ Design doc review, claim verification, logic primitive + adversarial probing →
198
213
 
199
214
  ## Templates & References
200
215
 
201
- This skill's `references/` directory contains document templates and platform references. Use `Read` with the skill's reference path to load the relevant file when needed:
202
-
203
- ### Platform
204
-
205
- - `platform-tool-mapping.md` — Claude Code → Codex/Cursor/Kimi/OpenCode/ZCode/Copilot tool name equivalents. **Load this immediately if you are NOT on Claude Code.**
206
-
207
- ### Workflow Templates
216
+ The `references/` directory holds the workflow document templates plus three notes. Read the file you need by name:
208
217
 
209
- - `detailed-change-plan.md` — Pre-edit implementation plan
210
- - `task-charter.md` — Task scope and slice roadmap
211
- - `iteration-notes.md` — Per-slice execution notes
212
- - `steward-memo.md` — Pre-execution architecture framing
213
- - `result-note.md` — Post-edit closure evidence
214
- - `final-qc.md` — Formal review verdict
215
- - `decision-log.md` — Approved decisions with rationale
216
- - `audit-ledger.md` — Recurring audit tracking (Framework Workflow)
217
- - `contract-matrix.md` — Contract-to-sentinel mapping (Framework Workflow)
218
- - `durable-requirement-notes.md` — Long-lived business invariants
218
+ - `platform-tool-mapping.md` — tool-name equivalents for Codex/Cursor/Kimi/OpenCode/ZCode/Copilot/dsh, and the per-harness context-budget table. **Read this immediately if you are not on Claude Code.**
219
+ - `proactive-suggestions.md` — ready-to-use phrasings for the suggestions below.
220
+ - `INDEX.md` — the working-memory index template.
221
+ - Workflow templates: `detailed-change-plan.md`, `task-charter.md`, `iteration-notes.md`, `steward-memo.md`, `result-note.md`, `final-qc.md`, `decision-log.md`, `audit-ledger.md`, `contract-matrix.md`, `durable-requirement-notes.md`.
219
222
 
220
223
  ---
221
224
 
222
225
  ## Proactive Suggestions
223
226
 
224
- When you observe any of these patterns in the user's task, **suggest the relevant feature before the user asks**. Most users don't know these capabilities exist.
225
-
226
- | Pattern You Observe | Suggest |
227
- |---------------------|--------|
228
- | User describes refactoring a state machine (splitting/merging states, changing transitions) | "Before you start, would you like me to run logic-primitive verification on the refactoring? I can extract the current state machine from code, compare it against your plan, and flag any regressions, deadlocks, or behavioral deltas before you change a single line." |
229
- | User describes a new state machine or protocol with ≥3 states | "I can run an adversarial verification on this design — 14 automated checks for deadlocks, unreachable states, race conditions, guard completeness, and invariant violations. Want me to do that before we implement?" |
230
- | User pastes or writes a state enum + switch-case dispatcher | "I notice a state machine here. Would you like me to model it and run completeness checks? I can find missing transitions, detect absorbing error loops, and verify that every state is reachable." |
231
- | User says "always" / "never" / "guaranteed" about behavior | "That's a behavioral invariant. I can model this and try to find a counter-example — the shortest event sequence that would violate 'X always happens before Y'. Want me to check?" |
232
- | User reviews a PR or diff that touches a state machine file | "This PR changes state machine logic. Would you like me to extract the before/after models and verify no regressions were introduced?" |
233
- | User debugs a crash or lockup in a stateful module | "This might be a state machine completeness issue. I can model the state machine from the code and check for deadlocks, unreachable states, or event ordering problems that could cause the lockup." |
234
- | Task would benefit from parallel execution (multiple independent modules, files, or dimensions) | "These are independent. I can dispatch parallel subagents to handle each module concurrently and synthesize the results. Want me to do that?" |
235
- | User writes a Detailed Change Plan without design review | "Before implementing, would you like the design-reviewer agent to fact-check this plan against the codebase? It catches API mismatches, missing modules, and mechanism feasibility issues before you write code." |
236
-
237
- ### Suggestion Rules
238
-
239
- - **Suggest once per task**, not repeatedly. If the user declines, don't push.
240
- - **Be specific about what the feature does** — don't just name-drop. Say "I can find deadlocks and missing transitions" not "I can run logicprobe." If logicprobe is not installed, offer the built-in fact-check skill: "I can check every claim in the plan against the codebase."
241
- - **Estimate cost**: for lightweight checks, say "this takes ~30 seconds." For Python harness runs, say "this will generate and run a verification script."
242
- - **Respect the user's decision**: if they decline, move on. The features are tools, not requirements.
227
+ Suggest a capability when it clearly applies and the user is unlikely to know it exists — state-machine verification for a state machine, a protocol, or an "always"/"never" claim; parallel sub-agents for genuinely independent modules; a design fact-check for a plan that had no review. Suggest **once per task**, say what the check finds rather than which tool runs it, and drop it if the user declines. Ready-to-use phrasings: `references/proactive-suggestions.md`.
@@ -1,31 +1,31 @@
1
1
  # Platform Tool Mapping
2
2
 
3
- This plugin's skills and agents are written with Claude Code tool names. This reference maps each tool to equivalents on other platforms. When an agent prompt says "use `Read`" but you are on Codex CLI, use the mapped tool instead.
3
+ This plugin's skills and agents are written with Claude Code tool names. This reference maps each tool to equivalents on the other supported platforms — Codex CLI, Cursor, Kimi CLI, OpenCode, ZCode, Copilot CLI, and DeepSeek Harness (dsh). When an agent prompt says "use `Read`" but you are on Codex CLI, use the mapped tool instead.
4
4
 
5
5
  ## Core File Tools
6
6
 
7
- | Claude Code | Codex CLI | Cursor | Kimi CLI | OpenCode | ZCode | Copilot CLI |
8
- |-------------|-----------|--------|----------|----------|-------|-------------|
9
- | `Read` | `read_file` | `read_file` | `read_file` | `read` | `read_file` | `read_file` |
10
- | `Write` | `write_file` | `write_to_file` | `write_file` | `write` | `write_file` | `write_file` |
11
- | `Edit` | `edit_file` | `replace_in_file` | `edit_file` | `edit` | `edit_file` | `edit_file` |
12
- | `Glob` | `search_file` | `search_file` | `glob` | `glob` | `search_file` | `search_file` |
13
- | `Grep` | `search_content` | `search_content` | `grep` | `grep` | `search_content` | `search_content` |
14
- | `Bash` | `run_shell` | `execute_command` | `execute_command` | `terminal` | `run_shell` | `run_command` |
7
+ | Claude Code | Codex CLI | Cursor | Kimi CLI | OpenCode | ZCode | Copilot CLI | DeepSeek Harness (dsh) |
8
+ |-------------|-----------|--------|----------|----------|-------|-------------|------------------------|
9
+ | `Read` | `read_file` | `read_file` | `read_file` | `read` | `read_file` | `view` | `read` |
10
+ | `Write` | `write_file` | `write_to_file` | `write_file` | `write` | `write_file` | `create` | `write` |
11
+ | `Edit` | `edit_file` | `replace_in_file` | `edit_file` | `edit` | `edit_file` | `edit` | `edit` |
12
+ | `Glob` | `search_file` | `search_file` | `glob` | `glob` | `search_file` | `glob` | `glob` |
13
+ | `Grep` | `search_content` | `search_content` | `grep` | `grep` | `search_content` | `grep` | `grep` |
14
+ | `Bash` | `run_shell` | `execute_command` | `execute_command` | `terminal` | `run_shell` | `bash` | `pwsh` / `bash` |
15
15
 
16
16
  ## Agent & Skill Tools
17
17
 
18
- | Claude Code | Codex CLI | Cursor | Kimi CLI | OpenCode | ZCode | Copilot CLI |
19
- |-------------|-----------|--------|----------|----------|-------|-------------|
20
- | `Skill("name")` | `$name` (auto) | `use_skill` | `/skill:name` | `$name` | `$name` | `skill("name")` |
21
- | `Agent` | `task` | `task` | `agent` | `task` | `agent` | `task` |
18
+ | Claude Code | Codex CLI | Cursor | Kimi CLI | OpenCode | ZCode | Copilot CLI | DeepSeek Harness (dsh) |
19
+ |-------------|-----------|--------|----------|----------|-------|-------------|------------------------|
20
+ | `Skill("name")` | `$name` (auto) | `use_skill` | `/skill:name` | `$name` | `$name` | `skill` tool, or `/skill-name` | `skill` tool (`name: "..."`) |
21
+ | `Agent` | `task` | `task` | `agent` | `task` | `agent` | `task` (`agent_type`) | `subagent` / `subagent_fork` |
22
22
 
23
23
  ## Web Tools
24
24
 
25
- | Claude Code | Codex CLI | Cursor | Kimi CLI | OpenCode | ZCode | Copilot CLI |
26
- |-------------|-----------|--------|----------|----------|-------|-------------|
27
- | `WebFetch` | `web_fetch` | `web_fetch` | `web_search` | `fetch` | `web_fetch` | `web_fetch` |
28
- | `WebSearch` | `web_search` | `web_search` | `web_search` | `search` | `web_search` | `web_search` |
25
+ | Claude Code | Codex CLI | Cursor | Kimi CLI | OpenCode | ZCode | Copilot CLI | DeepSeek Harness (dsh) |
26
+ |-------------|-----------|--------|----------|----------|-------|-------------|------------------------|
27
+ | `WebFetch` | `web_fetch` | `web_fetch` | `web_search` | `fetch` | `web_fetch` | `web_fetch` | `web_fetch` |
28
+ | `WebSearch` | `web_search` | `web_search` | `web_search` | `search` | `web_search` | — (no tool) | `web_search` |
29
29
 
30
30
  ## Platform-Specific Notes
31
31
 
@@ -62,13 +62,50 @@ This plugin's skills and agents are written with Claude Code tool names. This re
62
62
 
63
63
  ### Copilot CLI (GitHub Copilot)
64
64
 
65
- - `run_command` requires explicit approval for destructive operations.
66
- - Skill invocation differs from Claude Code; check Copilot CLI docs.
67
- - Copilot CLI 1.0.11+ supports Agent Skills.
65
+ - Copilot CLI has a plugin system: a `plugin.json` manifest plus `agents/NAME.agent.md`, `skills/NAME/SKILL.md`, hooks, and MCP servers, distributed through marketplaces (`copilot-plugins`, `awesome-copilot`). Install with `copilot plugin install NAME@MARKETPLACE`, or register this repository with `copilot plugin marketplace add AmethystLuna/embedded-workbench`.
66
+ - **This repository already works as a Copilot plugin**: Copilot's legacy manifest lookup checks `.plugin/plugin.json`, `plugin.json`, `.github/plugin/plugin.json`, then `.claude-plugin/plugin.json` — and it also reads `marketplace.json` from `.claude-plugin/`. Both files already exist here.
67
+ - Skills follow the Agent Skills open standard, so `skills/NAME/SKILL.md` loads as-is. Invoke one with `/skill-name` in a prompt, e.g. `Use the /debug-methodology skill to ...`; inspect them with `/skills list` or `copilot skill list`.
68
+ - Tool names: `view` (read), `create` (write), `edit`, `glob`, `grep`, `bash`, `web_fetch`, `task`, `skill`. The edit tool is `str_replace_editor` under the hood, with `view`/`create`/`edit` as its aliases.
69
+ - There is **no** `web_search` tool — use `web_fetch` against a search URL.
70
+ - `task` requires an `agent_type` of `explore`, `task`, `general-purpose`, or `code-review`. `explore` is read-only and the cheapest; prefer it for searches.
71
+ - Sub-agent status comes from `read_agent` / `list_agents`; async shell sessions use `bash` with `mode: "async"` plus `read_bash` / `write_bash` / `stop_bash` / `list_bash`.
72
+ - Agents in this plugin are **not** loaded by Copilot: it requires `agents/NAME.agent.md`, and this repository ships `agents/NAME.md`.
73
+
74
+ ### DeepSeek Harness (dsh)
75
+
76
+ - dsh is the one platform this plugin ships a **native bundle** for (root `package.json` declares `dsh.bundle`), so no skill-copy step is needed: `dsh plugin --profile <name> add dsh-embedded-workbench`.
77
+ - Tool names are lowercase: `read`, `write`, `edit`, `glob`, `grep`. Shell access is `pwsh` on Windows hosts and `bash` elsewhere.
78
+ - Skills load through the `skill` tool by name — there is no `Skill(...)` call syntax.
79
+ - Plan mode is the native `exit_plan_mode` tool, not `ExitPlanMode`.
80
+ - The 4 Claude Code sub-agents are intentionally **not** ported: use the native sub-agent tools for parallel work, and take the steward/reviewer roles directly.
81
+ - `subagent` starts a **fresh** context and returns only its result — reach for it when the detail can be discarded (read-only discovery, sweeping logs, running a suite). `subagent_fork` is **seeded with this conversation** — reach for it when the sub-task needs context you already built, instead of rebuilding that context inside a prompt.
82
+ - `Agent(subagent_type: "Explore")` has no separate equivalent — use `glob` / `grep` directly, or dispatch a read-only `subagent`.
83
+ - **Context-budget markers** you can actually see (no token meter is exposed to the model): a tool result rewritten with `[... tool result middle pruned ...]`, a spill notice naming the omitted bytes and the complete-result path, and the compaction checkpoint preamble (`This is an automatically generated checkpoint condensing an earlier span…`). Treat any of them as "stop pulling content in whole" — switch to file references, or delegate the reading.
84
+ - Install, verify, and config-override details live in `.dsh/INSTALL.md` at the repository root.
85
+
86
+ ## Context Budget Interfaces
87
+
88
+ What each harness actually exposes to the **model** about its context budget, checked 2026-09-25 against vendor documentation (and, for Codex CLI, the `openai/codex` source). "User-only" means the signal exists but never reaches the model. `UNVERIFIED` means the vendor does not document it — do not invent a marker string for that cell.
89
+
90
+ | Harness | Model sees a token/percent readout? | Compaction — trigger and what the model receives | Oversized tool output | Sub-agent context |
91
+ | --- | --- | --- | --- | --- |
92
+ | Claude Code | No — `/context` and the status line are user-facing (the status line even exposes `used_percentage`, but to the terminal, not the model) | Automatic near the limit (`/autocompact <100K–1M>`). Older tool outputs are cleared first, then history is summarized; afterwards an oversized re-read returns as a `Referenced file` path instead of content, and invoked skill bodies are re-injected capped at 5,000 tokens each / 25,000 total | Bash success output past ~30,000 characters becomes a file path plus a preview of up to 2,000 characters (a failure is truncated in place with **no** path); hook `additionalContext` over 10,000 characters is saved to a file and the model gets a preview plus the path; Read adds a `PARTIAL view` notice and Glob flags a 100-file cap | Fresh isolated window; **fork mode** (on by default in interactive sessions, off under `-p`/SDK) inherits the parent conversation. Note a skill's `context: fork` is *not* a conversation fork |
93
+ | Codex CLI | Yes, but feature-gated: a `get_context_remaining` tool (feature `token_budget`, **off by default**) answers "You have {N} tokens left in this context window", and a reminder is injected once remaining drops below the threshold. `/status` percentages are user-only | Default 90% of the window, hard cap 95%. The model receives the summary behind a fixed handoff preamble ("Another language model started to solve this problem…") | `truncation_policy` defaults to bytes/10000: the model sees `Warning: truncated output (original token count: N)` and `…N tokens truncated…`. Spill-to-file exists only for hooks (`Full hook output saved to: <path>`) | `spawn_agent` `fork_turns`: `"none"` / `"all"` / a turn count — **the default is `"all"`** |
94
+ | Copilot CLI | No — `/context`, `/usage`, and `footer.showContextWindow` are user-only | ~80% of the window in the background (defers to ~90% when static context already uses ≥75%), pausing at ~95%. The model gets the summary plus preserved user instructions and plan/todo state; each compaction writes a numbered checkpoint (`/session checkpoints`) | **Over 20 KiB** is saved to a temporary file and the model gets the path plus a preview (`COPILOT_LARGE_OUTPUT_THRESHOLD_BYTES`) | Fresh window; `contextTier: "inherit"` inherits the parent's tier. Repository custom instructions are **not** inherited unless the agent sets `include-custom-instructions: true` |
95
+ | Cursor | Nothing documented — the context ring and breakdown tray are UI | Automatic once the window is full; `/summarize` (alias `/compress`). The injected summary text is not published | Not documented (UNVERIFIED) | Fresh window only; `model: inherit` inherits the **model**, not the conversation |
96
+ | OpenCode | No | `compaction.auto` is on, keeping ~15,000 recent tokens (`keep.tokens`) with `buffer` at 10% of the limit. The model sees the summary **as past conversation** — there is no distinct notice, and later compactions update the same summary. Native provider compaction substitutes an opaque encrypted item | Long tool output is shortened beside the summary; no marker string documented | UNVERIFIED |
97
+ | Kimi CLI | No — `/usage` is a user command | Automatic compression as the conversation approaches the window limit, plus `/compact [hint]`. `/undo` cannot cross a compaction | Not documented (UNVERIFIED) | Isolated context per sub-agent (its own event stream); built-ins `coder`, `explore`, `plan`. `/btw` runs in a **forked** sub-agent that sees the conversation |
98
+ | ZCode | No — the usage-stats panel shows the context composition as UI | `/compact` is user-initiated; automatic compaction is not documented | Not documented (UNVERIFIED) | UNVERIFIED |
99
+ | DeepSeek Harness (dsh) | No — `ctx.tokenMeter` deliberately adds no model-visible surface | At `floor(min(W × 0.8, W − O − 65536))`. The model receives the checkpoint preamble plus `<compacted-summary>`; `/compact` is a user command | `[... tool result middle pruned ...]` for results over 8,192 code points (4,096 head / 1,024 tail) once compaction pressure qualifies, plus a spill notice naming the omitted bytes and the complete-result path (12,500 estimated tokens in `dsh-base`) | `subagent` starts fresh; `subagent_fork` is seeded with the conversation so far |
100
+
101
+ Two consequences worth remembering:
102
+
103
+ - **A budget readout reaching the model is the exception, not the rule.** Only Codex documents one, and it is off by default. Everywhere else, act on the truncation, spill, or compaction notice your harness actually shows.
104
+ - **With no signal, ask the user instead of guessing.** Say what you are about to consume, that this harness gives you no budget readout, and the options with their costs, then do what they choose — falling back to the reversible option only when no one can answer.
68
105
 
69
106
  ## Sub-Agent Platform Equivalents
70
107
 
71
- This plugin defines 4 sub-agents. On platforms without an `Agent` tool:
108
+ This plugin defines 4 sub-agents. On platforms without an `Agent` tool (including DeepSeek Harness):
72
109
 
73
110
  | Claude Code Agent | Alternative Approach |
74
111
  |-------------------|---------------------|
@@ -81,7 +118,7 @@ This plugin defines 4 sub-agents. On platforms without an `Agent` tool:
81
118
 
82
119
  Load this when:
83
120
 
84
- - The session is NOT running on Claude Code (check environment: `$CLAUDE_PLUGIN_ROOT`, `$CODEX_CLI`, `$CURSOR_PLUGIN_ROOT`, etc.)
121
+ - The session is NOT running on Claude Code (check environment: `$CLAUDE_PLUGIN_ROOT`, `$CODEX_CLI`, `$CURSOR_PLUGIN_ROOT`, etc.; a dsh session runs under `$DSH_HOME` and sees none of those)
85
122
  - An agent prompt references a tool you don't recognize
86
123
  - You need to translate a skill or workflow instruction to your platform
87
124
 
@@ -0,0 +1,27 @@
1
+ # Proactive Suggestion Phrasings
2
+
3
+ Ready-to-use wording for the suggestions the bootstrap skill describes. Suggest once per
4
+ task, say what the check actually finds rather than which tool runs it, and drop it if the
5
+ user declines.
6
+
7
+ | Pattern you observe | Suggest |
8
+ | --- | --- |
9
+ | Refactoring a state machine (splitting/merging states, changing transitions) | "Before you start, would you like me to run logic-primitive verification on the refactoring? I can extract the current state machine from code, compare it against your plan, and flag regressions, deadlocks, or behavioral deltas before you change a line." |
10
+ | A new state machine or protocol with ≥3 states | "I can run an adversarial verification on this design — 22 automated checks (8 structural S1–S8 plus 14 adversarial A1–A14) for deadlocks, unreachable states, race conditions, guard completeness, and invariant violations." |
11
+ | A state enum plus a switch-case dispatcher | "I notice a state machine here. I can model it and run completeness checks — missing transitions, absorbing error loops, unreachable states." |
12
+ | "always" / "never" / "guaranteed" about behavior | "That's a behavioral invariant. I can model it and look for a counter-example — the shortest event sequence that violates 'X always happens before Y'." |
13
+ | A PR or diff touching a state machine file | "I can extract the before/after models and verify that no regressions were introduced." |
14
+ | A crash or lockup in a stateful module | "This might be a state-machine completeness issue. I can model it and check for deadlocks, unreachable states, or event-ordering problems behind the lockup." |
15
+ | Multiple independent modules, files, or dimensions | "These are independent. I can dispatch parallel sub-agents per module and synthesize the results." |
16
+ | A Detailed Change Plan with no design review | "Would you like the plan fact-checked against the codebase first? It catches API mismatches, missing modules, and mechanism-feasibility gaps before you write code." |
17
+
18
+ ## Rules
19
+
20
+ - **Suggest once per task**, not repeatedly. If the user declines, don't push.
21
+ - **Say what the check finds**, not which tool runs it. If logicprobe is not installed,
22
+ offer the built-in `fact-check` skill instead — "I can check every claim in the plan
23
+ against the codebase" — and note that behavioral/model claims then degrade to manual
24
+ confirmation.
25
+ - **Estimate cost** when it matters: a lightweight check is seconds; a full model
26
+ verification pass generates and runs a script.
27
+ - **Respect the decision**: these are tools, not requirements.
package/src/index.ts CHANGED
@@ -1,12 +1,11 @@
1
1
  /**
2
2
  * embedded-workbench — DeepSeek Harness native plugin for the Embedded
3
- * Workbench toolbox. Injects the session-start gate text (1% Rule, Red
4
- * Flags, Plan Verification Gate, skills roster) into the first model step
5
- * of every agent session, mirroring the SessionStart hook the Claude Code
6
- * plugin installs. The 8 skills ship in this package's `skills/` directory
3
+ * Workbench toolbox. The 8 skills ship in this package's `skills/` directory
7
4
  * and are registered at apply time into dsh's `ctx.skills` registry through
8
5
  * the standard filesystem provider, so they appear in every session catalog
9
- * without a manual copy step.
6
+ * without a manual copy step. The plugin also folds a short gate text into the
7
+ * first model step of every agent session, mirroring the SessionStart hook the
8
+ * Claude Code plugin installs.
10
9
  *
11
10
  * Injection listens on agent/pre-step and appends the gate to the FIRST
12
11
  * model step that runs, once per session (guarded by the session's durable
@@ -17,11 +16,14 @@
17
16
  * reminders (skill catalog, AGENTS.md, gate plugins) simply defer this message
18
17
  * to the first step after their promotion, and the history guard re-injects it
19
18
  * there. The default gate text is the dsh-native adaptation of
20
- * `hooks/session-start-content.md`: behavior rules
21
- * (1% Rule / Red Flags / Plan Verification Gate) stay in sync, while
22
- * presentation is adapted to dsh's native skill catalog — no roster table
23
- * (the model sees skills in its catalog) and no install instructions (those
24
- * live in `.dsh/INSTALL.md`). Deployments override via Config.
19
+ * `hooks/session-start-content.md`: the behavior rules stay in sync (the Plan
20
+ * Verification Gate and the context-budget rule), while presentation is adapted
21
+ * to dsh's native skill catalog — no roster table (the model sees skills in its
22
+ * catalog) and no install instructions (those live in `.dsh/INSTALL.md`). The
23
+ * payload is deliberately small: it carries the verification gate and the
24
+ * budget rule, not the 1% Rule / Red Flags enforcement scaffolding, which
25
+ * measurably pushes capable models into rigid phases and unnecessary fan-out.
26
+ * Deployments override via Config.
25
27
  *
26
28
  * @module embedded-workbench-dsh
27
29
  */
@@ -30,10 +32,23 @@ import { fileURLToPath } from 'node:url'
30
32
  import type { Context } from '@deepseek-ai/cordis'
31
33
  import z from '@deepseek-ai/schemastery'
32
34
  import { createUserMessage } from '@deepseek-ai/dsh-llm'
35
+ import type { ContextFormed } from '@deepseek-ai/dsh-llm'
33
36
  import type { Session, UserMessage } from '@deepseek-ai/dsh-session'
34
37
  import type { HostCordisInspectProviderRegistration } from '@deepseek-ai/dsh-cordis-host-runner'
35
38
  import { FileSystemSkillProvider } from '@deepseek-ai/dsh-skill-filesystem'
36
39
 
40
+ // DSH 0.1.7-alpha.1 (session format v4) retires the shared
41
+ // `{ kind: 'plugin', plugin }` wrapper: native admission rejects it in every
42
+ // declared durable message slot, and the official v3-to-v4 migration rewrites
43
+ // those historical rows to `plugin:<name>`. Declaring the producer-owned kind
44
+ // here keeps the write path and the history guard on one identity, and still
45
+ // compiles against the earlier releases that only declare `plugin`.
46
+ declare module '@deepseek-ai/dsh-llm' {
47
+ interface MessageSourceMap {
48
+ 'plugin:embedded-workbench': { kind: 'plugin:embedded-workbench' } & ContextFormed
49
+ }
50
+ }
51
+
37
52
  export const name = 'embedded-workbench'
38
53
 
39
54
  // Skills are contributed through the registry service, which dsh-base always
@@ -47,30 +62,17 @@ const SKILLS_DIR = fileURLToPath(new URL('../skills', import.meta.url))
47
62
 
48
63
  const GATE_PLUGIN_ID = 'embedded-workbench'
49
64
 
50
- const DEFAULT_GATE_CONTENT = `<EXTREMELY_IMPORTANT>
51
- Plugin embedded-workbench is active. You have embedded C/C++ firmware development skills — names and "Use when" triggers are in your skill catalog; load them with the skill tool. No custom agents in dsh: use the native subagent tooling for parallel work.
65
+ /** Producer-owned message source kind declared in `MessageSourceMap` above. */
66
+ const GATE_SOURCE_KIND: 'plugin:embedded-workbench' = 'plugin:embedded-workbench'
52
67
 
53
- **1% Rule**: If there is even a 1% chance a skill applies to your task, invoke it before responding. If the skill turns out to be wrong for the situation, discard it and move on. The cost of loading a skill is trivial compared to the cost of a preventable mistake.
54
-
55
- **Red Flags** — if you think any of these, STOP. You are rationalizing:
56
-
57
- | You think | Reality |
58
- |-----------|---------|
59
- | "This is just a quick fix" | Quick fixes break things. A 3-line design check costs 30 seconds. |
60
- | "I already understand this code" | You are looking at one file. The blast radius may span 5 modules. |
61
- | "The skill is overkill for this" | Simple things become complex. Check for skills. |
62
- | "Let me explore the codebase first" | Skills tell you HOW to explore. Check first. |
63
- | "I can just read the file directly" | Skills have patterns and pitfalls you will not discover by reading. |
64
- | "I remember this skill content" | Skills evolve. Always load the current version. |
65
- | "I've explored enough, time to exit plan mode" | The exit_plan_mode tool is the verification gate. Have you loaded the logicprobe skill — or, if it is not installed, the built-in fallback fact-check skill? Every plan must pass this gate before exit. |
66
- | "This plan is too simple for logicprobe" | logicprobe auto-classifies depth; the fallback fact-check verifies every claim regardless. You don't decide. |
67
- | "I already read the code, I know the file paths are correct" | Load the logicprobe skill or the fallback fact-check skill, verify each claim, append the "## Plan Verification" block. |
68
+ const DEFAULT_GATE_CONTENT = `<EXTREMELY_IMPORTANT>
69
+ Plugin embedded-workbench is active: embedded C/C++ firmware development skills are in your skill catalog. Load the one whose "Use when" matches before substantial work, with the skill tool.
68
70
 
69
- **Plan Verification Gate**: Before calling exit_plan_mode (or presenting a plan for approval), load the logicprobe skill (a separate plugin) — or, if it is missing from your skill catalog, load the built-in fallback fact-check skill for claim-by-claim verification, tell the user that behavioral/model claims degrade to manual confirmation, and recommend installing logicprobe. If neither is loaded, inform the user "此计划未经核查,是否需要我先做事实核查?" Silent skip is not an option.
71
+ **Plan Verification Gate**: before calling exit_plan_mode (or presenting a plan for approval), load the logicprobe skill — or the built-in fact-check skill when logicprobe is not installed — and append a "## Plan Verification" block to the plan. If you verify with neither, tell the user the plan is unverified before asking for approval; a silent skip is not an option. "This change is too small to check" and "I already read the code, the paths are right" are the two rationalizations this gate exists to catch.
70
72
 
71
- To load workflows and engineering policies: load the embedded-workbench skill.
73
+ **Context budget**: no token meter is visible to you, so never guess one. Act on what you can see — a pruned or spilled tool result means stop pulling it in whole, and a compaction checkpoint means move durable state into files. When a large step (many sources, a long sweep, several independent areas) shows no such signal, ask the user what to spend context on rather than deciding silently. If nobody can answer, take the reversible option and say so. When you do delegate, prefer \`subagent_fork\` over \`subagent\` if the sub-agent needs context you already built — its summary still lands here.
72
74
 
73
- **Proactive features**: When you see state machines, protocol refactoring, behavioral claims ("always"/"never"), or multi-module tasks — suggest verification (logicprobe, or the built-in fact-check fallback if logicprobe is not installed), adversarial probing, or parallel subagents BEFORE the user asks. Most users do not know these exist.
75
+ To load the workflows and engineering policies behind these skills: load the embedded-workbench skill.
74
76
  </EXTREMELY_IMPORTANT>`
75
77
 
76
78
  export interface Config {
@@ -79,6 +81,9 @@ export interface Config {
79
81
  }
80
82
 
81
83
  export const Config = z.object({
84
+ // On by default, but deliberately small: the payload carries the verification
85
+ // gate and the context-budget rule only, so leaving it on costs a few hundred
86
+ // tokens once per session rather than the ~900 of the previous payload.
82
87
  enabled: z.boolean().default(true),
83
88
  gateContent: z.string().default(DEFAULT_GATE_CONTENT),
84
89
  })
@@ -87,7 +92,7 @@ function gateMessage(text: string): UserMessage {
87
92
  return createUserMessage({
88
93
  content: [{ type: 'text', text }],
89
94
  // `form` omitted — an undeclared context is the documented default.
90
- source: { kind: 'plugin', plugin: GATE_PLUGIN_ID },
95
+ source: { kind: GATE_SOURCE_KIND },
91
96
  })
92
97
  }
93
98
 
@@ -127,7 +132,7 @@ function inspectProvider(config: Config): HostCordisInspectProviderRegistration
127
132
  return {
128
133
  manifest: {
129
134
  id: 'embedded-workbench',
130
- description: 'Session-start gate injection for the Embedded Workbench toolbox — folds the 1% Rule / Red Flags / Plan Verification Gate text into the first model step of every agent session.',
135
+ description: 'Session-start gate injection for the Embedded Workbench toolbox — folds the Plan Verification Gate and the context-budget rule into the first model step of every agent session.',
131
136
  methods: [
132
137
  {
133
138
  name: 'status',
@@ -231,6 +236,10 @@ function gateInHistory(session: Session): boolean {
231
236
  return readSessionEvents(session).some((event) => {
232
237
  if (event.type !== 'user/message') return false
233
238
  const source = event.data?.source as { kind?: string; plugin?: string } | undefined
234
- return source?.kind === 'plugin' && source.plugin === GATE_PLUGIN_ID
239
+ if (source === undefined) return false
240
+ // The v4 producer-owned kind, plus the pre-v4 wrapper this bundle wrote
241
+ // before DSH 0.1.7-alpha.1 retired it.
242
+ return source.kind === GATE_SOURCE_KIND
243
+ || (source.kind === 'plugin' && source.plugin === GATE_PLUGIN_ID)
235
244
  })
236
245
  }