dsh-speak 1.8.0 → 1.8.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,458 +1,529 @@
1
- # dsh-speak 🔊 — Voice announcements for AI coding harnesses
2
-
3
- **English** · [中文](README.zh-CN.md)
4
-
5
- ![鲸鱼娘大喇叭](鲸鱼娘大喇叭.png)
6
-
7
- [![Awesome DSH Plugin](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
8
-
9
- [![npm version](https://img.shields.io/npm/v/dsh-speak)](https://www.npmjs.com/package/dsh-speak)
10
-
11
- Let your agent **tell you** when a long task is done — no more staring at the screen.
12
-
13
- dsh-speak reads the final assistant reply aloud through system speech synthesis —
14
- on Windows using natural voices (Windows 11 built-in, or
15
- [NaturalVoiceSAPIAdapter] on Windows 10) with graceful fallback to stock voices;
16
- on macOS using the built-in `say` (can follow a Siri natural voice). It was built
17
- for [DeepSeek Harness](https://github.com/deepseek-ai/dsh)
18
- and is structured so any harness can plug in.
19
-
20
- ## Features
21
-
22
- - **Automatic**: DSH web plugin watches the session event stream and announces the
23
- final reply (skips reasoning/tool-call narration, merges multi-step messages).
24
- - **Gets your attention**: announces approval requests (hears "需要你的审批" when
25
- the agent is waiting on you) and questions the agent asks via `ask_user_question`.
26
- - **Final-reply replay** (1.7.0): every final reply (turn tail) has a 🔊 button
27
- in its action bar — click to replay that message, click again to stop, click
28
- another to switch. Speech execution stays fully owned by the DSH host (keeps
29
- speaking even with the browser closed).
30
- - **Host speech queue** (1.7.0): only one native speech process runs at a time;
31
- queued items continue automatically. A WebSocket syncs the live state (which
32
- message is speaking, queue length) to the UI.
33
- - **Optional event announcements** (1.6.0): turn end, command done, goal changes,
34
- tool errors, and todo updates can each be announced, toggled independently
35
- (off by default).
36
- - **Visual configuration** (1.7.0): a dedicated Settings → dsh-speak settings
37
- page — every option (master switch, automatic speech, Markdown cleaning, code
38
- blocks, event toggles, fixed prompt, …) is editable from the Web UI, no
39
- hand-edited YAML.
40
- - **Master switch** (1.6.0): silence everything with one toggle.
41
- - **Bundle auto-registration** (1.3.0): declare the package in `dsh.profile.bundles`
42
- and the plugin registers itself via the bundled `cordis.patch.yml` — no manual
43
- patch entry needed.
44
- - **Best-effort**: never throws, never blocks the harness, never breaks a session.
45
- - **Natural voices**: Windows prefers natural voices — Windows 11 built-in packs,
46
- or voices registered via NaturalVoiceSAPIAdapter on Windows 10 (e.g. Xiaoxiao);
47
- macOS uses the system reading voice (Siri natural voices on recent macOS). Both
48
- fall back to any installed voice.
49
- - **Robust text cleaning**: strips markdown/URLs/emoji that make speech synthesis
50
- fail silently, and guards the adapter's per-utterance character ceiling.
51
- - **Portable engine**: any process can speak with one line:
52
- Windows `powershell -File speak.ps1 -Text "你好"` / macOS `./speak.sh -t "你好"`.
53
-
54
- ## How it works
55
-
56
- ```
57
- harness event (DSH session event / Claude Code Stop hook / anything)
58
- │
59
- ▼ adapters/… (harness-specific trigger: filter, throttle, cancel)
60
- ▼ engine/speak.ps1 / speak.sh (harness-agnostic: clean text → SAPI5 / say)
61
- ▼ 🔊 you hear the final reply
62
- ```
63
-
64
- The adapter turns harness-specific events into engine calls; the engine cleans the
65
- text and speaks it, fully decoupled from any harness. Full design:
66
- [docs/DESIGN.md](docs/DESIGN.md).
67
-
68
- ## Prerequisites
69
-
70
- Windows:
71
-
72
- - Windows 10 or 11, PowerShell (any recent version).
73
- - Natural voices:
74
- - **Windows 11 (21H2–23H2)**: natural voice packs are built into the system —
75
- no extra installation. Enable/switch them in *Settings → Accessibility →
76
- Narrator* or *Settings → Time & Language → Speech*.
77
- - **Windows 11 24H2/25H2**: natural voices moved to MSIX app packages, which
78
- `System.Speech` may not enumerate (falls back to a robotic stock voice). As
79
- on Windows 10, install
80
- [NaturalVoiceSAPIAdapter](https://github.com/gexgd0419/NaturalVoiceSAPIAdapter)
81
- to bridge them.
82
- - **Windows 10**: install
83
- [NaturalVoiceSAPIAdapter](https://github.com/gexgd0419/NaturalVoiceSAPIAdapter)
84
- and use its VoiceDownloader to download the natural voice pack(s) you want
85
- (Chinese or any other language).
86
- - Without natural voices, the engine falls back to a stock voice (e.g. Huihui).
87
-
88
- macOS:
89
-
90
- - macOS (Apple Silicon or Intel), built-in `say` command — **no extra software**.
91
- - Chinese voices: see the [macOS](#macos) section (incl. the Siri natural-voice
92
- picker and its pitfalls).
93
-
94
- DSH web app:
95
-
96
- - Tested against **DSH 0.1.5-rc.1**. Two host/client APIs changed after 0.1.1, both
97
- handled here (1.8.0):
98
- - `@deepseek-ai/dsh-settings` deleted the `installSettingsSection` /
99
- `settingsNamespace` helpers — the plugin now registers its namespace through
100
- the `settings` **service**. On those older releases the plugin aborted the
101
- host boot (`settingsNamespace is not a function`); a missing settings provider
102
- now just leaves the composed patch `config` in force.
103
- - the Session snapshot stopped carrying Conversation target data — the 🔊 button
104
- resolves the clicked message through the Chat target hook `useChat`.
105
-
106
- ## Install & quick start
107
-
108
- ### DSH — Option A: npm plugin (recommended)
109
-
110
- ```powershell
111
- # 1. install the plugin into your web profile (adds dsh-speak to
112
- # ~/.dsh/profiles/web/package.json dependencies)
113
- dsh plugin --profile web add dsh-speak
114
-
115
- # 2. register it in ~/.dsh/profiles/web/cordis.patch.yml
116
- # (for npm packages the bare package name is used — no file:/// URL needed):
117
- # - insert:
118
- # - id: speech-hook
119
- # name: 'dsh-speak'
120
-
121
- # 3. restart the DSH web app — replies are now announced automatically
122
- ```
123
-
124
- > **No pnpm?** `dsh plugin` forwards to pnpm, which is not installed on every
125
- > machine. The exact same install can be done with npm directly:
126
- >
127
- > ```powershell
128
- > npm install --prefix "$env:USERPROFILE\.dsh\profiles\web" dsh-speak
129
- > ```
130
- >
131
- > On macOS (bash):
132
- >
133
- > ```bash
134
- > npm install --prefix "$HOME/.dsh/profiles/web" dsh-speak
135
- > ```
136
-
137
- The engine ships inside the package (`node_modules/dsh-speak/engine/`), so no extra
138
- copying is needed.
139
-
140
- > **Let your agent do it?** Paste this repo URL
141
- > (`https://github.com/Alan2Z/dsh-speak`) into your DSH session and ask it to
142
- > install the plugin — your agent follows this very README. Approving the
143
- > out-of-workspace writes (`~/.dsh`) is all that's needed.
144
-
145
- ### DSH — Option B: file install (no npm needed)
146
-
147
- ```powershell
148
- # 1. clone
149
- git clone https://github.com/Alan2Z/dsh-speak.git
150
- cd dsh-speak
151
-
152
- # 2. one-command install: copies engine + plugin, registers in cordis.patch.yml
153
- powershell.exe -NoProfile -ExecutionPolicy Bypass -File adapters\dsh\install.ps1
154
-
155
- # 3. verify the engine speaks
156
- powershell -NoProfile -ExecutionPolicy Bypass -File "$env:USERPROFILE\.dsh\hooks\speak.ps1" -Text "你好,语音播报已就绪。"
157
-
158
- # 4. restart the DSH web app — replies are now announced automatically
159
- ```
160
-
161
- What the file installer did:
162
-
163
- | file | destination |
164
- | ---- | ----------- |
165
- | `engine/*.ps1` | `%USERPROFILE%\.dsh\hooks\` |
166
- | `adapters/dsh/speech-hook.js` | `%USERPROFILE%\.dsh\profiles\web\plugins\` |
167
- | registration entry | appended to `%USERPROFILE%\.dsh\profiles\web\cordis.patch.yml` (backed up first) |
168
-
169
- ### macOS
170
-
171
- The same adapter runs natively on macOS — the plugin auto-detects the platform and
172
- calls `engine/speak.sh` (the built-in `say` command) instead of `speak.ps1`.
173
- **Since 1.2.0 the macOS engine ships in the npm package** — no extra software.
174
-
175
- ```bash
176
- # 1. install into your web profile (no pnpm needed — only `dsh plugin` requires it)
177
- npm install --prefix "$HOME/.dsh/profiles/web" dsh-speak
178
-
179
- # 2. register in ~/.dsh/profiles/web/cordis.patch.yml (bare package name — no file:/// URL):
180
- # - insert:
181
- # - id: speech-hook
182
- # name: 'dsh-speak'
183
-
184
- # 3. no restart needed — the patch watcher hot-reloads; replies are announced
185
- # after the throttle (~1.5 s); tool-calling replies are announced at turn end
186
- ```
187
-
188
- > With pnpm installed, `dsh plugin --profile web add dsh-speak` works identically.
189
-
190
- #### Voices (important — two pitfalls)
191
-
192
- - By default the engine follows the **system reading voice** (*Settings →
193
- Accessibility → Spoken Content → System Voice*). On **macOS 26** that picker has
194
- an **ⓘ circle icon** next to it — click it for the full voice list; the plain
195
- dropdown does **not** contain the Siri natural voices. Pick e.g. "普通话 Siri
196
- 声音1(男声)" there.
197
- - **Siri voice** (*Settings → Siri → Voice*) and the system reading voice are
198
- **two independent settings**; Siri voices are not exposed to `say -v '?'` and
199
- cannot be selected by name — they only work as the system default.
200
- - ⚠️ **Pitfall 1 (reproduced)**: opening the "Spoken Content / Siri Voice" settings
201
- pane — **even without changing anything** — drifts/resets the system voice to the
202
- classic "婷婷 (Tingting)". If the voice suddenly changes, re-pick it via the ⓘ
203
- entry.
204
- - ⚠️ **Pitfall 2**: the log lives at `$TMPDIR/dsh-speech-hook.log`
205
- (`os.tmpdir()` — **not** `/tmp`).
206
- - Use `-v Eddy|Flo|Tingting` to force a specific voice (`say -v '?'` lists them).
207
- - `say` has no volume flag — volume follows the system output volume.
208
-
209
- #### Test the engine alone (no DSH needed)
210
-
211
- ```bash
212
- curl -sfL -o ~/speak.sh "https://cdn.jsdelivr.net/gh/Alan2Z/dsh-speak@main/engine/speak.sh"
213
- chmod +x ~/speak.sh
214
- ~/speak.sh -t "你好,Mac 版语音播报测试"
215
- ~/speak.sh -t "测试" -v Eddy -r 200 # explicit voice + rate
216
- ```
217
-
218
- ### Claude Code
219
-
220
- Register the Stop hook in `~/.claude/settings.json`:
221
-
222
- ```json
223
- {
224
- "hooks": {
225
- "Stop": [
226
- {
227
- "hooks": [
228
- {
229
- "type": "command",
230
- "command": "powershell.exe -NoProfile -ExecutionPolicy Bypass -File C:\\path\\to\\dsh-speak\\adapters\\claude-code\\stop-hook.ps1"
231
- }
232
- ]
233
- }
234
- ]
235
- }
236
- }
237
- ```
238
-
239
- ### Any other harness
240
-
241
- Call the engine directly from your agent / wrapper / script:
242
-
243
- ```powershell
244
- # announce a one-liner
245
- powershell -NoProfile -ExecutionPolicy Bypass -File engine\speak.ps1 -Text "构建完成"
246
-
247
- # announce a long summary (blocking, returns when done)
248
- powershell -NoProfile -ExecutionPolicy Bypass -File engine\speech-summary.ps1 -Text "…"
249
-
250
- # ask for user attention (blocking, for prompts/approvals)
251
- powershell -NoProfile -ExecutionPolicy Bypass -File engine\speech-prompt.ps1 -Text "请做出选择"
252
- ```
253
-
254
- ## Configuration
255
-
256
- ### Engine parameters
257
-
258
- See [docs/DESIGN.md §5 configuration reference](docs/DESIGN.md#5-configuration-reference):
259
-
260
- ```powershell
261
- speak.ps1 -Text "…" -Volume 50 -Rate 1 -MaxChars 300 -LongTextMessage "本次播报内容较长,请自行阅读。"
262
- ```
263
-
264
- ### DSH plugin config
265
-
266
- **Either way works, and they stay in sync** (both write the same settings
267
- document):
268
-
269
- 1. **Web UI (1.7.0, recommended)**: a dedicated Settings → dsh-speak settings
270
- page. Every option is editable and saved there (visible in `dsh --dump-config`,
271
- per-profile, survives npm updates).
272
- 2. **Profile patch `config` block** (equivalent):
273
-
274
- ```yaml
275
- # ~/.dsh/profiles/web/cordis.patch.yml
276
- - insert:
277
- - id: speech-hook
278
- name: 'dsh-speak'
279
- config:
280
- enabled: true # master switch: false silences everything
281
- automaticSpeech: true # auto-speak final replies
282
- queueAllMessages: false # true = enqueue every assistant message as it arrives
283
- replayFullRead: false # true = manual replay skips the long-text truncation, reads everything
284
- cleanMarkdownFormatting: true # convert Markdown to natural speech
285
- readInlineCode: true # read inline code without backticks
286
- codeBlocks: smart # all | smart | replace (fenced code blocks)
287
- codeBlockMaxChars: 300 # smart-mode code block character limit
288
- codeBlockReplacementText: 'You can see the code in our history.' # replace-mode text
289
- throttleMs: 1500 # merge delay before announcing (ms)
290
- engine: '' # engine path override; '' = auto-resolve
291
- announceApprovals: true # speak approval requests
292
- announceQuestions: true # speak ask_user_question content
293
- stripApprovalPrefix: true # strip "escalate sandbox to ...: " prefix
294
- questionGapMs: 2000 # pause between multiple question announcements (ms)
295
- longTextMode: message # message | heading (speak largest md heading)
296
- longTextMessage: '本次播报内容较长,请自行阅读。' # fixed prompt for message mode
297
- maxChars: 300 # per-utterance ceiling (macOS default 0 = unlimited)
298
- volume: 50 # Windows only
299
- rate: 0 # 0 = engine default (Windows SAPI scale / macOS wpm)
300
- # —— optional event announcements (1.6.0, all off by default) ——
301
- announceTurnEnd: false # turn/end — "第 N 轮对话完成"
302
- announceCommandDone: false # command/done — command finished/failed
303
- announceGoalChange: false # goal/change — goal created/updated/completed
304
- announceToolErrors: false # tool/result error — announce (english dropped)
305
- announceTodoWrite: false # todo/write — todo list updated
306
- ```
307
-
308
- > Resolution order: schema default → patch `config` → UI user settings. Fields
309
- > written in YAML show up in the UI too. Platform note: `maxChars` defaults to
310
- > 0 on macOS (`say` has no ceiling) and 300 on Windows (SAPI safe limit).
311
-
312
- #### Option reference
313
-
314
- | option | default | effect |
315
- | ------ | ------- | ------ |
316
- | `enabled` | `true` | **master switch**: when off, nothing is ever announced (final reply / approvals / questions / optional events / replay) |
317
- | `automaticSpeech` | `true` | auto-speak final replies; manual replay always remains available |
318
- | `queueAllMessages` | `false` | `true` enqueues every assistant message as it arrives (intermediate messages spoken too, FIFO); default only speaks the throttled final reply |
319
- | `replayFullRead` | `false` | `true` makes manual replay skip the long-text heading truncation (`longTextMode: heading`) and read everything in chunks |
320
- | `cleanMarkdownFormatting` | `true` | converts Markdown into natural speech text (link labels kept, URLs/heading/emphasis cleaned) |
321
- | `readInlineCode` | `true` | read inline code without backtick markers |
322
- | `codeBlocks` | `smart` | fenced code blocks: `all` read / `smart` (read when ≤ `codeBlockMaxChars`) / `replace` with the replacement text |
323
- | `codeBlockMaxChars` | `300` | code block character limit for `smart` mode |
324
- | `codeBlockReplacementText` | `You can see the code in our history.` | replacement spoken in `replace` mode (or over-limit `smart`) |
325
- | `throttleMs` | `1500` | how long a reply's text waits before being announced (merges multi-step messages) |
326
- | `engine` | `''` | explicit engine script path; `''` auto-resolves: `<package>/engine/<platform>` → `~/.dsh/hooks/<platform>` |
327
- | `announceApprovals` | `true` | announce `approval/asked` events (reason, or the fixed prompt) |
328
- | `announceQuestions` | `true` | announce `ask_user_question`: each question spoken separately with a "问题N" prefix (when several) and "选项N" prefixes matching the UI numbering; a `questionGapMs` pause between questions |
329
- | `questionGapMs` | `2000` | pause between multiple question announcements (ms); 0 = no pause |
330
- | `stripApprovalPrefix` | `true` | strip the fixed English template prefix (`escalate sandbox to danger-full-access: `) from approval reasons, keeping the human explanation |
331
- | `longTextMode` | `message` | `message` = fixed prompt for over-long text; `heading` = speak the largest markdown heading instead (see below) |
332
- | `longTextMessage` | `本次播报内容较长,请自行阅读。` | the fixed prompt spoken for over-long text in `message` mode (editable in the UI) |
333
- | `maxChars` | platform | per-utterance ceiling. **macOS default 0 (`say` has no ceiling); Windows default 300** (SAPI fails silently beyond ~375-470) |
334
- | `volume` | `50` | Windows only (0-100); macOS volume follows the system |
335
- | `rate` | `0` | speech rate: Windows SAPI scale (-10 to 10, 0 = normal; try 1-3 for faster); macOS words-per-minute (default 175, 200 is a bit faster) |
336
- | `announceTurnEnd` | `false` | announce "第 N 轮对话完成/中断/异常结束" on turn end (`turn/end`) |
337
- | `announceCommandDone` | `false` | announce when a command finishes or fails (`command/done`) |
338
- | `announceGoalChange` | `false` | announce goal created/updated/completed/paused/resumed (`goal/change`, objective head) |
339
- | `announceToolErrors` | `false` | announce "工具调用出错" when a tool call returns an error: `tool/result` carrying `error` (structured failure identity) or a result block with `isError === true`. A **non-zero shell exit does NOT count** — pwsh/bash report `exit code: N` as result data by design, so only infrastructure failures (spawn errors, aborts) and structured tool failures (e.g. fs) set `isError` (English details / technical codes dropped, Chinese details kept) |
340
- | `announceTodoWrite` | `false` | announce "待办已更新:n/m 完成" when the agent updates its todos (`todo/write`) |
341
-
342
- #### Long-text modes
343
-
344
- When cleaned text exceeds `maxChars`:
345
-
346
- - **`message`** (default): speak `longTextMessage` (`本次播报内容较长,请自行阅读。`,
347
- editable in the UI or YAML).
348
- - **`heading`**: pick the *largest* markdown heading in the raw text — fewest `#`
349
- wins, tie → first. When there is **no heading at all**, speak a coherent opening
350
- instead of just the first line: the leading `maxChars` window, trimmed back to
351
- its last sentence end, and kept whole when that would drop more than half the
352
- window. Sentence ends are recognised bilingually: full-width `。!?;` and `…`
353
- always count, while half-width `.!?;` only count when followed by whitespace, a
354
- closing quote/bracket, or (for the very last character) one read past the window —
355
- so an English `period + space` at the edge still lands, but a decimal point such
356
- as `Version 0.1.` does not. (Before 1.8.0 this fallback spoke the first non-empty
357
- line only, which sounded like the narration stopped after line 1.) The chosen
358
- candidate is still cleaned and subject to the `maxChars` ceiling, falling back to
359
- the message if it is itself too long.
360
-
361
- Full architecture and design rationale: [docs/DESIGN.md](docs/DESIGN.md).
362
-
363
- ## Customizing (survives npm updates)
364
-
365
- You can tune behavior without forking, and your changes **survive `npm update`**:
366
-
367
- 1. **Copy the engine out and edit it** (recommended — this is where defaults live: volume,
368
- rate, `MaxChars`, `LongTextMessage`, voice logic):
369
-
370
- ```powershell
371
- # Windows
372
- Copy-Item "$env:USERPROFILE\.dsh\profiles\web\node_modules\dsh-speak\engine\speak.ps1" "$env:USERPROFILE\.dsh\hooks\my-speak.ps1"
373
- # macOS
374
- cp ~/.dsh/profiles/web/node_modules/dsh-speak/engine/speak.sh ~/.dsh/hooks/my-speak.sh
375
- ```
376
-
377
- Then point the plugin at your copy in the `config` block:
378
-
379
- > **Windows: keep the file's UTF-8 BOM.** `speak.ps1` is a UTF-8 script and Windows
380
- > PowerShell 5.1 only knows that from the 3-byte BOM (`EF BB BF`) at the start; an
381
- > editor that saves it without one makes the system ANSI code page decode it instead,
382
- > and Chinese text inside the script turns to mojibake — the symptom is **silence or
383
- > wrong trimming, with no error**. The shipped script keeps all of its *logic* ASCII-only
384
- > for that reason, so a lost BOM only garbles the Chinese comments and the default
385
- > prompt. After editing, check with
386
- > `Get-Content -Encoding Byte -TotalCount 3 your-speak.ps1` (expect `239 187 191`), or
387
- > run `node scripts/test-engine-static.js`.
388
-
389
- ```yaml
390
- - insert:
391
- - id: speech-hook
392
- name: 'dsh-speak'
393
- config:
394
- engine: 'C:/Users/<you>/.dsh/hooks/my-speak.ps1' # or ~/.dsh/hooks/my-speak.sh on macOS
395
- ```
396
-
397
- The plugin resolves the engine as `config.engine` → package engine → `~/.dsh/hooks/`,
398
- so your copy wins. `npm update` only touches the package — your engine stays.
399
-
400
- 2. **Edit the file inside `node_modules`** — works, but the next `npm update` overwrites it.
401
-
402
- 3. **Fork the repo** — full control, publish your own package if you want.
403
-
404
- ## Troubleshooting
405
-
406
- | symptom | cause | fix |
407
- | ------- | ----- | --- |
408
- | No sound at all, no error | no natural voice enabled/installed | Win11: enable a natural voice in *Settings → Narrator / Speech*; Win10: install NaturalVoiceSAPIAdapter + a voice pack. Test `speak.ps1` directly |
409
- | Long replies never spoken | adapter per-`Speak` character ceiling | already guarded at 300 chars — lower `-MaxChars` if needed |
410
- | Narration stops after the first line | with `longTextMode: heading`, text over `maxChars` and no markdown heading made the engine speak only the first non-empty line (pre-1.8.0) | fixed in 1.8.0 (speaks a coherent opening instead); to change the policy use `message` mode or raise `maxChars` |
411
- | `工具调用出错:Error: cannot read …` spoken | the "is this Chinese?" detail filter only checked for the presence of a CJK character, so a Chinese directory name inside an English error passed it (1.8.0 regression) | fixed in 1.8.0 — the detail now needs more Chinese characters than Latin letters |
412
- | Emoji-heavy text silent | SAPI fails silently on emoji | already stripped by the engine |
413
- | Plugin not loading | raw Windows path as plugin name | use the `file:///C:/…` URL form (installer does this) |
414
- | macOS: voice suddenly became "婷婷" | opening the "Spoken Content / Siri Voice" pane drifted the system voice | re-pick via Settings → Accessibility → Spoken Content → System Voice → ⓘ entry |
415
- | macOS: no log at `/tmp` | `os.tmpdir()` is `/var/folders/.../T`, not `/tmp` | log is at `$TMPDIR/dsh-speech-hook.log` |
416
-
417
- Plugin diagnostics: Windows `%TEMP%\dsh-speech-hook.log`; macOS `$TMPDIR/dsh-speech-hook.log`
418
-
419
- ## Repository layout
420
-
421
- ```
422
- engine/ harness-agnostic speech engine (PowerShell + SAPI5 / bash + say)
423
- speak.ps1 / speak.sh clean + speak (the only seam any adapter needs)
424
- speech-prompt.ps1 blocking short announcement
425
- speech-summary.ps1 blocking reply-summary announcement
426
- adapters/
427
- dsh/ DSH web plugin + one-command installer
428
- speech-hook.js session-event trigger (throttle/cancel + optional events + FIFO speech queue + WebSocket + settings registration)
429
- install.ps1 copies + registers + backs up
430
- claude-code/
431
- stop-hook.ps1 Claude Code Stop hook trigger
432
- client/
433
- client.js DSH browser bundle: turn-tail Speak/Stop button + Settings → dsh-speak settings page
434
- docs/
435
- DESIGN.md full design rationale, pitfalls, extension guide
436
- scripts/ tests + manual dev helpers (not shipped in the npm package)
437
- test-engine-static.js engine invariants: .ps1 BOM + PowerShell parse, .sh LF (also run by prepublishOnly)
438
- test-engine-longtext.js long-text guard contract for BOTH engines (speak.ps1 -DryRun / speak.sh's perl)
439
- test-speech-hook.js host plugin: event triggers, queue, tool-error detail filter
440
- test-client-bundle.js browser bundle: slot registration + component rendering
441
- test-settings-integration.js settings-service wiring + removed-API guard
442
- session-log-dump.js read a DSH session log (manual: what text reached the engine)
443
- settings-ui-check.py Playwright UI check (manual: needs a running, authenticated dsh)
444
- dsh-events-check.py Playwright disclosure check (manual)
445
- ```
446
-
447
- ## Writing a new adapter
448
-
449
- Three reference patterns exist: **event-stream** (DSH), **stop-hook** (Claude Code),
450
- **agent-called** (`speech-summary.ps1` from a shell). In every case the adapter only
451
- needs to: capture the *final reply text* → invoke the engine. See
452
- [docs/DESIGN.md §7](docs/DESIGN.md#7-extending).
453
-
454
- ## License
455
-
456
- MIT — see [LICENSE](LICENSE).
457
-
458
- [NaturalVoiceSAPIAdapter]: https://github.com/gexgd0419/NaturalVoiceSAPIAdapter
1
+ # dsh-speak 🔊 — Voice announcements for AI coding harnesses
2
+
3
+ **English** · [中文](README.zh-CN.md)
4
+
5
+ ![鲸鱼娘大喇叭](鲸鱼娘大喇叭.png)
6
+
7
+ [![Awesome DSH Plugin](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
8
+
9
+ [![npm version](https://img.shields.io/npm/v/dsh-speak)](https://www.npmjs.com/package/dsh-speak)
10
+
11
+ Let your agent **tell you** when a long task is done — no more staring at the screen.
12
+
13
+ dsh-speak reads the final assistant reply aloud through system speech synthesis —
14
+ on Windows using natural voices (Windows 11 built-in, or
15
+ [NaturalVoiceSAPIAdapter] on Windows 10) with graceful fallback to stock voices;
16
+ on macOS using the built-in `say` (can follow a Siri natural voice). It was built
17
+ for [DeepSeek Harness](https://github.com/deepseek-ai/dsh)
18
+ and is structured so any harness can plug in.
19
+
20
+ ## Features
21
+
22
+ - **Automatic**: DSH web plugin watches the session event stream and announces the
23
+ final reply (skips reasoning/tool-call narration, merges multi-step messages).
24
+ - **Gets your attention**: announces approval requests (hears "需要你的审批" when
25
+ the agent is waiting on you) and questions the agent asks via `ask_user_question`.
26
+ - **Final-reply replay** (1.7.0): every final reply (turn tail) has a 🔊 button
27
+ in its action bar — click to replay that message, click again to stop, click
28
+ another to switch. Speech execution stays fully owned by the DSH host (keeps
29
+ speaking even with the browser closed).
30
+ - **Host speech queue** (1.7.0): only one native speech process runs at a time;
31
+ queued items continue automatically. A WebSocket syncs the live state (which
32
+ message is speaking, queue length) to the UI.
33
+ - **Optional event announcements** (1.6.0): turn end, command done, goal changes,
34
+ tool errors, and todo updates can each be announced, toggled independently
35
+ (off by default).
36
+ - **Visual configuration** (1.7.0): a dedicated Settings → dsh-speak settings
37
+ page — every option (master switch, automatic speech, Markdown cleaning, code
38
+ blocks, event toggles, fixed prompt, …) is editable from the Web UI, no
39
+ hand-edited YAML.
40
+ - **Master switch** (1.6.0): silence everything with one toggle.
41
+ - **Bundle auto-registration** (1.3.0): declare the package in `dsh.profile.bundles`
42
+ and the plugin registers itself via the bundled `cordis.patch.yml` — no manual
43
+ patch entry needed.
44
+ - **Best-effort**: never throws, never blocks the harness, never breaks a session.
45
+ - **Natural voices**: Windows prefers natural voices — Windows 11 built-in packs,
46
+ or voices registered via NaturalVoiceSAPIAdapter on Windows 10 (e.g. Xiaoxiao);
47
+ macOS uses the system reading voice (Siri natural voices on recent macOS). Both
48
+ fall back to any installed voice.
49
+ - **Robust text cleaning**: strips markdown/URLs/emoji that make speech synthesis
50
+ fail silently, and guards the adapter's per-utterance character ceiling.
51
+ - **Portable engine**: any process can speak with one line:
52
+ Windows `powershell -File speak.ps1 -Text "你好"` / macOS `./speak.sh -t "你好"`.
53
+
54
+ ## How it works
55
+
56
+ ```
57
+ harness event (DSH session event / Claude Code Stop hook / anything)
58
+ │
59
+ ▼ adapters/… (harness-specific trigger: filter, throttle, cancel)
60
+ ▼ engine/speak.ps1 / speak.sh (harness-agnostic: clean text → SAPI5 / say)
61
+ ▼ 🔊 you hear the final reply
62
+ ```
63
+
64
+ The adapter turns harness-specific events into engine calls; the engine cleans the
65
+ text and speaks it, fully decoupled from any harness. Full design:
66
+ [docs/DESIGN.md](docs/DESIGN.md).
67
+
68
+ ## Prerequisites
69
+
70
+ Windows:
71
+
72
+ - Windows 10 or 11, PowerShell (any recent version).
73
+ - Natural voices:
74
+ - **Windows 11 (21H2–23H2)**: natural voice packs are built into the system —
75
+ no extra installation. Enable/switch them in *Settings → Accessibility →
76
+ Narrator* or *Settings → Time & Language → Speech*.
77
+ - **Windows 11 24H2/25H2**: natural voices moved to MSIX app packages, which
78
+ `System.Speech` may not enumerate (falls back to a robotic stock voice). As
79
+ on Windows 10, install
80
+ [NaturalVoiceSAPIAdapter](https://github.com/gexgd0419/NaturalVoiceSAPIAdapter)
81
+ to bridge them.
82
+ - **Windows 10**: install
83
+ [NaturalVoiceSAPIAdapter](https://github.com/gexgd0419/NaturalVoiceSAPIAdapter)
84
+ and use its VoiceDownloader to download the natural voice pack(s) you want
85
+ (Chinese or any other language).
86
+ - Without natural voices, the engine falls back to a stock voice (e.g. Huihui).
87
+
88
+ macOS:
89
+
90
+ - macOS (Apple Silicon or Intel), built-in `say` command — **no extra software**.
91
+ - Chinese voices: see the [macOS](#macos) section (incl. the Siri natural-voice
92
+ picker and its pitfalls).
93
+
94
+ DSH web app:
95
+
96
+ - Tested against **DSH 0.1.7-rc.2**. Host/client APIs changed twice since 0.1.1,
97
+ both handled here:
98
+ - **0.1.7 replaced the settings provider API with Config projection**: the
99
+ settings service now reads each active Loader entry's own exported `Config`
100
+ schema and projects its *volatile* fields into the settings UI
101
+ (`ctx.settings.describe()` on the host, `ctx.configForms` in the browser).
102
+ `settings.register(namespace, schema, { base })` — the 1.6.0–1.8.x wiring —
103
+ is gone, and the browser service `settingsScope` was replaced by
104
+ `configForms`. This plugin therefore exports its schema as `Config` and the
105
+ settings namespace is its **entry id** (`dsh-speak`; a leftover
106
+ `speech-hook` row from 1.8.x is still bound, see below).
107
+ - **0.1.2** stopped putting Conversation target data into the Session snapshot
108
+ (the 🔊 button resolves the clicked message through the Chat target hook
109
+ `useChat`), and deleted `@deepseek-ai/dsh-settings`'s `installSettingsSection`
110
+ / `settingsNamespace` helpers.
111
+ - The host floor lives where dsh-market reads it: `engines.dsh` in `package.json`
112
+ (`>=0.1.7-rc.2`). The catalog card and its "compatible with current DSH" filter
113
+ read exactly that field, so the floor moves only after a release has been
114
+ verified against the new host.
115
+ - 1.8.2 requires the settings model introduced in 0.1.7: on an older host the
116
+ browser half would wait forever for the `configForms` service. Use 1.8.1 for
117
+ DSH ≤ 0.1.6.
118
+
119
+ ## Install & quick start
120
+
121
+ ### DSH — Option A: npm plugin (recommended)
122
+
123
+ ```powershell
124
+ # 1. install the plugin into your web profile (adds dsh-speak to
125
+ # ~/.dsh/profiles/web/package.json dependencies AND to dsh.profile.bundles)
126
+ dsh plugin --profile web add dsh-speak
127
+
128
+ # 2. that is all — the package ships its own bundle patch, which registers the
129
+ # `dsh-speak` entry. Do NOT also hand-write an `insert:` row for it:
130
+ # registering the same entry twice makes the id non-unique, and DSH's config
131
+ # editor then rejects every settings write with
132
+ # `settings/rejected: Configuration for "dsh-speak" is overridden by a home
133
+ # patch or command-line overlay` (speech keeps working, the settings page
134
+ # silently bounces).
135
+ #
136
+ # To pin options in YAML anyway, add ONLY this top-level row (the settings page
137
+ # edits it in place; `config` on a top-level row is the shape the editor
138
+ # supports — see "DSH plugin config"):
139
+ # - id: dsh-speak
140
+ # name: 'dsh-speak'
141
+ # config: {}
142
+ #
143
+ # The entry id IS the settings namespace on DSH >= 0.1.7: your options are
144
+ # stored under that key in this same file. A row left over from 1.8.x
145
+ # (id: speech-hook) keeps working — the browser half binds either id.
146
+
147
+ # 3. restart the DSH web app — replies are now announced automatically
148
+ ```
149
+
150
+ > Installing from the in-app marketplace is the same path: it adds the package to
151
+ > `dsh.profile.bundles`. Only the manual installs below need hand-written rows.
152
+ > **One registration path per profile** — never both.
153
+
154
+ > **No pnpm?** `dsh plugin` forwards to pnpm, which is not installed on every
155
+ > machine. The exact same install can be done with npm directly:
156
+ >
157
+ > ```powershell
158
+ > npm install --prefix "$env:USERPROFILE\.dsh\profiles\web" dsh-speak
159
+ > ```
160
+ >
161
+ > On macOS (bash):
162
+ >
163
+ > ```bash
164
+ > npm install --prefix "$HOME/.dsh/profiles/web" dsh-speak
165
+ > ```
166
+ >
167
+ > An `npm install` alone does **not** register the plugin: add either the bundle
168
+ > name to `~/.dsh/profiles/web/package.json` → `dsh.profile.bundles`, or the two
169
+ > rows from the file-install section — not both.
170
+
171
+ The engine ships inside the package (`node_modules/dsh-speak/engine/`), so no extra
172
+ copying is needed.
173
+
174
+ > **Let your agent do it?** Paste this repo URL
175
+ > (`https://github.com/Alan2Z/dsh-speak`) into your DSH session and ask it to
176
+ > install the plugin — your agent follows this very README. Approving the
177
+ > out-of-workspace writes (`~/.dsh`) is all that's needed.
178
+
179
+ ### DSH — Option B: file install (no npm needed)
180
+
181
+ ```powershell
182
+ # 1. clone
183
+ git clone https://github.com/Alan2Z/dsh-speak.git
184
+ cd dsh-speak
185
+
186
+ # 2. one-command install: copies engine + plugin, registers in cordis.patch.yml
187
+ powershell.exe -NoProfile -ExecutionPolicy Bypass -File adapters\dsh\install.ps1
188
+
189
+ # 3. verify the engine speaks
190
+ powershell -NoProfile -ExecutionPolicy Bypass -File "$env:USERPROFILE\.dsh\hooks\speak.ps1" -Text "你好,语音播报已就绪。"
191
+
192
+ # 4. restart the DSH web app — replies are now announced automatically
193
+ ```
194
+
195
+ What the file installer did:
196
+
197
+ | file | destination |
198
+ | ---- | ----------- |
199
+ | `engine/*.ps1` | `%USERPROFILE%\.dsh\hooks\` |
200
+ | `adapters/dsh/speech-hook.js` | `%USERPROFILE%\.dsh\profiles\web\plugins\` |
201
+ | registration entry | appended to `%USERPROFILE%\.dsh\profiles\web\cordis.patch.yml` (backed up first) |
202
+
203
+ ### macOS
204
+
205
+ The same adapter runs natively on macOS — the plugin auto-detects the platform and
206
+ calls `engine/speak.sh` (the built-in `say` command) instead of `speak.ps1`.
207
+ **Since 1.2.0 the macOS engine ships in the npm package** — no extra software.
208
+
209
+ ```bash
210
+ # 1. install into your web profile (no pnpm needed — only `dsh plugin` requires it)
211
+ npm install --prefix "$HOME/.dsh/profiles/web" dsh-speak
212
+
213
+ # 2. register in ~/.dsh/profiles/web/cordis.patch.yml (bare package name — no file:/// URL):
214
+ # - insert:
215
+ # - id: dsh-speak
216
+ # name: 'dsh-speak'
217
+ # - id: dsh-speak
218
+ # name: 'dsh-speak'
219
+ # config: {}
220
+
221
+ # 3. no restart needed — the patch watcher hot-reloads; replies are announced
222
+ # after the throttle (~1.5 s); tool-calling replies are announced at turn end
223
+ ```
224
+
225
+ > With pnpm installed, `dsh plugin --profile web add dsh-speak` works identically.
226
+
227
+ #### Voices (important — two pitfalls)
228
+
229
+ - By default the engine follows the **system reading voice** (*Settings →
230
+ Accessibility → Spoken Content → System Voice*). On **macOS 26** that picker has
231
+ an **ⓘ circle icon** next to it — click it for the full voice list; the plain
232
+ dropdown does **not** contain the Siri natural voices. Pick e.g. "普通话 Siri
233
+ 声音1(男声)" there.
234
+ - **Siri voice** (*Settings → Siri → Voice*) and the system reading voice are
235
+ **two independent settings**; Siri voices are not exposed to `say -v '?'` and
236
+ cannot be selected by name — they only work as the system default.
237
+ - ⚠️ **Pitfall 1 (reproduced)**: opening the "Spoken Content / Siri Voice" settings
238
+ pane — **even without changing anything** — drifts/resets the system voice to the
239
+ classic "婷婷 (Tingting)". If the voice suddenly changes, re-pick it via the ⓘ
240
+ entry.
241
+ - ⚠️ **Pitfall 2**: the log lives at `$TMPDIR/dsh-speech-hook.log`
242
+ (`os.tmpdir()` — **not** `/tmp`).
243
+ - Use `-v Eddy|Flo|Tingting` to force a specific voice (`say -v '?'` lists them).
244
+ - `say` has no volume flag — volume follows the system output volume.
245
+
246
+ #### Test the engine alone (no DSH needed)
247
+
248
+ ```bash
249
+ curl -sfL -o ~/speak.sh "https://cdn.jsdelivr.net/gh/Alan2Z/dsh-speak@main/engine/speak.sh"
250
+ chmod +x ~/speak.sh
251
+ ~/speak.sh -t "你好,Mac 版语音播报测试"
252
+ ~/speak.sh -t "测试" -v Eddy -r 200 # explicit voice + rate
253
+ ```
254
+
255
+ ### Claude Code
256
+
257
+ Register the Stop hook in `~/.claude/settings.json`:
258
+
259
+ ```json
260
+ {
261
+ "hooks": {
262
+ "Stop": [
263
+ {
264
+ "hooks": [
265
+ {
266
+ "type": "command",
267
+ "command": "powershell.exe -NoProfile -ExecutionPolicy Bypass -File C:\\path\\to\\dsh-speak\\adapters\\claude-code\\stop-hook.ps1"
268
+ }
269
+ ]
270
+ }
271
+ ]
272
+ }
273
+ }
274
+ ```
275
+
276
+ ### Any other harness
277
+
278
+ Call the engine directly from your agent / wrapper / script:
279
+
280
+ ```powershell
281
+ # announce a one-liner
282
+ powershell -NoProfile -ExecutionPolicy Bypass -File engine\speak.ps1 -Text "构建完成"
283
+
284
+ # announce a long summary (blocking, returns when done)
285
+ powershell -NoProfile -ExecutionPolicy Bypass -File engine\speech-summary.ps1 -Text "…"
286
+
287
+ # ask for user attention (blocking, for prompts/approvals)
288
+ powershell -NoProfile -ExecutionPolicy Bypass -File engine\speech-prompt.ps1 -Text "请做出选择"
289
+ ```
290
+
291
+ ## Configuration
292
+
293
+ ### Engine parameters
294
+
295
+ See [docs/DESIGN.md §5 configuration reference](docs/DESIGN.md#5-configuration-reference):
296
+
297
+ ```powershell
298
+ speak.ps1 -Text "…" -Volume 50 -Rate 1 -MaxChars 300 -LongTextMessage "本次播报内容较长,请自行阅读。"
299
+ ```
300
+
301
+ ### DSH plugin config
302
+
303
+ **Either way works, and they stay in sync** (both write the same profile patch
304
+ document — the settings service persists UI edits into the row it read them from):
305
+
306
+ 1. **Web UI (1.7.0, recommended)**: a dedicated Settings → dsh-speak settings
307
+ page. Every option is editable and saved there (visible in `dsh --dump-config`,
308
+ per-profile, survives npm updates).
309
+ 2. **Profile patch `config` block** (equivalent):
310
+
311
+ ```yaml
312
+ # ~/.dsh/profiles/web/cordis.patch.yml
313
+ # TWO rows, and the shape matters: `insert` provides the entry, the TOP-LEVEL row
314
+ # carries the config the settings page edits. DSH's config editor rewrites a
315
+ # config in place only on a top-level row; a config nested inside the insert row
316
+ # (what 1.8.x profiles and the 1.8.2 notes first showed) makes the UI accept an
317
+ # edit, apply it live, and then silently roll it back on disk — the value reverts
318
+ # on the next boot.
319
+ - insert:
320
+ - id: dsh-speak # the entry id IS the settings namespace (0.1.7+)
321
+ name: 'dsh-speak'
322
+ - id: dsh-speak
323
+ name: 'dsh-speak'
324
+ config:
325
+ enabled: true # master switch: false silences everything
326
+ automaticSpeech: true # auto-speak final replies
327
+ queueAllMessages: false # true = enqueue every assistant message as it arrives
328
+ replayFullRead: false # true = manual replay skips the long-text truncation, reads everything
329
+ cleanMarkdownFormatting: true # convert Markdown to natural speech
330
+ readInlineCode: true # read inline code without backticks
331
+ codeBlocks: smart # all | smart | replace (fenced code blocks)
332
+ codeBlockMaxChars: 300 # smart-mode code block character limit
333
+ codeBlockReplacementText: 'You can see the code in our history.' # replace-mode text
334
+ throttleMs: 1500 # merge delay before announcing (ms)
335
+ engine: '' # engine path override; '' = auto-resolve
336
+ announceApprovals: true # speak approval requests
337
+ announceQuestions: true # speak ask_user_question content
338
+ stripApprovalPrefix: true # strip "escalate sandbox to ...: " prefix
339
+ questionGapMs: 2000 # pause between multiple question announcements (ms)
340
+ longTextMode: message # message | heading (speak largest md heading)
341
+ longTextMessage: '本次播报内容较长,请自行阅读。' # fixed prompt for message mode
342
+ maxChars: 300 # per-utterance ceiling (macOS default 0 = unlimited)
343
+ volume: 50 # Windows only
344
+ rate: 0 # 0 = engine default (Windows SAPI scale / macOS wpm)
345
+ # —— optional event announcements (1.6.0, all off by default) ——
346
+ announceTurnEnd: false # turn/end — "第 N 轮对话完成"
347
+ announceCommandDone: false # command/done — command finished/failed
348
+ announceGoalChange: false # goal/change — goal created/updated/completed
349
+ announceToolErrors: false # tool/result error — announce (english dropped)
350
+ announceTodoWrite: false # todo/write — todo list updated
351
+ ```
352
+
353
+ > Resolution order: schema default → patch `config` → UI user settings. Fields
354
+ > written in YAML show up in the UI too. Platform note: `maxChars` defaults to
355
+ > 0 on macOS (`say` has no ceiling) and 300 on Windows (SAPI safe limit).
356
+ >
357
+ > Values outside the ranges in the table (`volume` 0-100, Windows `rate` -10..10)
358
+ > are clamped before they reach the engine — SAPI throws on an out-of-range
359
+ > `Volume`/`Rate`, which surfaces as silence and nothing else. A clamped value is
360
+ > logged as `settings 值超出范围,已钳制: <field> <given> -> <used>`. An explicit
361
+ > `0` is always honored (`volume: 0` silences, `maxChars: 0` means unlimited,
362
+ > `throttleMs: 0` announces without merging).
363
+
364
+ > **Keep the config on the top-level row.** It is not a style preference: with the
365
+ > `config:` block nested inside the `insert` row, the settings page still renders
366
+ > and the running plugin still obeys an edit, but DSH's config editor cannot rewrite
367
+ > that row — it appends a new top-level row and then rolls the write back, so the
368
+ > value silently reverts the next time dsh starts. Check with
369
+ > `python scripts/settings-ui-check.py`, which asserts the write lands in the patch.
370
+
371
+ > **Upgrading from 1.8.x:** edit the options through the settings page, or move
372
+ > your old `config:` block onto the new top-level row (see the shape above). Before
373
+ > 0.1.7 the options lived in a `dsh-speak:` section of `~/.dsh/settings.yaml`; that
374
+ > file was migrated to `settings.yaml.imported` by DSH, and a section whose name
375
+ > matched no Loader entry (1.8.x registered the namespace in code, so `dsh-speak:`
376
+ > matched nothing) was left behind. If you had custom values, copy them into that
377
+ > `config:` block — the entry id (`dsh-speak`) is now the key DSH looks for.
378
+
379
+ #### Option reference
380
+
381
+ | option | default | effect |
382
+ | ------ | ------- | ------ |
383
+ | `enabled` | `true` | **master switch**: when off, nothing is ever announced (final reply / approvals / questions / optional events / replay) |
384
+ | `automaticSpeech` | `true` | auto-speak final replies; manual replay always remains available |
385
+ | `queueAllMessages` | `false` | `true` enqueues every assistant message as it arrives (intermediate messages spoken too, FIFO); default only speaks the throttled final reply |
386
+ | `replayFullRead` | `false` | `true` makes manual replay skip the long-text heading truncation (`longTextMode: heading`) and read everything in chunks |
387
+ | `cleanMarkdownFormatting` | `true` | converts Markdown into natural speech text (link labels kept, URLs/heading/emphasis cleaned) |
388
+ | `readInlineCode` | `true` | read inline code without backtick markers |
389
+ | `codeBlocks` | `smart` | fenced code blocks: `all` read / `smart` (read when ≤ `codeBlockMaxChars`) / `replace` with the replacement text |
390
+ | `codeBlockMaxChars` | `300` | code block character limit for `smart` mode |
391
+ | `codeBlockReplacementText` | `You can see the code in our history.` | replacement spoken in `replace` mode (or over-limit `smart`) |
392
+ | `throttleMs` | `1500` | how long a reply's text waits before being announced (merges multi-step messages) |
393
+ | `engine` | `''` | explicit engine script path; `''` auto-resolves: `<package>/engine/<platform>` → `~/.dsh/hooks/<platform>` |
394
+ | `announceApprovals` | `true` | announce `approval/asked` events (reason, or the fixed prompt) |
395
+ | `announceQuestions` | `true` | announce `ask_user_question`: each question spoken separately with a "问题N" prefix (when several) and "选项N" prefixes matching the UI numbering; a `questionGapMs` pause between questions |
396
+ | `questionGapMs` | `2000` | pause between multiple question announcements (ms); 0 = no pause |
397
+ | `stripApprovalPrefix` | `true` | strip the fixed English template prefix (`escalate sandbox to danger-full-access: `) from approval reasons, keeping the human explanation |
398
+ | `longTextMode` | `message` | `message` = fixed prompt for over-long text; `heading` = speak the largest markdown heading instead (see below) |
399
+ | `longTextMessage` | `本次播报内容较长,请自行阅读。` | the fixed prompt spoken for over-long text in `message` mode (editable in the UI) |
400
+ | `maxChars` | platform | per-utterance ceiling. **macOS default 0 (`say` has no ceiling); Windows default 300** (SAPI fails silently beyond ~375-470) |
401
+ | `volume` | `50` | Windows only (0-100); macOS volume follows the system |
402
+ | `rate` | `0` | speech rate: Windows SAPI scale (-10 to 10, 0 = normal; try 1-3 for faster); macOS words-per-minute (default 175, 200 is a bit faster) |
403
+ | `announceTurnEnd` | `false` | announce "第 N 轮对话完成/中断/异常结束" on turn end (`turn/end`) |
404
+ | `announceCommandDone` | `false` | announce when a command finishes or fails (`command/done`) |
405
+ | `announceGoalChange` | `false` | announce goal created/updated/completed/paused/resumed (`goal/change`, objective head) |
406
+ | `announceToolErrors` | `false` | announce "工具调用出错" when a tool call returns an error: `tool/result` carrying `error` (structured failure identity) or a result block with `isError === true`. A **non-zero shell exit does NOT count** — pwsh/bash report `exit code: N` as result data by design, so only infrastructure failures (spawn errors, aborts) and structured tool failures (e.g. fs) set `isError` (English details / technical codes dropped, Chinese details kept) |
407
+ | `announceTodoWrite` | `false` | announce "待办已更新:n/m 完成" when the agent updates its todos (`todo/write`) |
408
+
409
+ #### Long-text modes
410
+
411
+ When cleaned text exceeds `maxChars`:
412
+
413
+ - **`message`** (default): speak `longTextMessage` (`本次播报内容较长,请自行阅读。`,
414
+ editable in the UI or YAML).
415
+ - **`heading`**: pick the *largest* markdown heading in the raw text — fewest `#`
416
+ wins, tie → first. When there is **no heading at all**, speak a coherent opening
417
+ instead of just the first line: the leading `maxChars` window, trimmed back to
418
+ its last sentence end, and kept whole when that would drop more than half the
419
+ window. Sentence ends are recognised bilingually: full-width `。!?;` and `…`
420
+ always count, while half-width `.!?;` only count when followed by whitespace, a
421
+ closing quote/bracket, or (for the very last character) one read past the window —
422
+ so an English `period + space` at the edge still lands, but a decimal point such
423
+ as `Version 0.1.` does not. (Before 1.8.0 this fallback spoke the first non-empty
424
+ line only, which sounded like the narration stopped after line 1.) The chosen
425
+ candidate is still cleaned and subject to the `maxChars` ceiling, falling back to
426
+ the message if it is itself too long.
427
+
428
+ Full architecture and design rationale: [docs/DESIGN.md](docs/DESIGN.md).
429
+
430
+ ## Customizing (survives npm updates)
431
+
432
+ You can tune behavior without forking, and your changes **survive `npm update`**:
433
+
434
+ 1. **Copy the engine out and edit it** (recommended — this is where defaults live: volume,
435
+ rate, `MaxChars`, `LongTextMessage`, voice logic):
436
+
437
+ ```powershell
438
+ # Windows
439
+ Copy-Item "$env:USERPROFILE\.dsh\profiles\web\node_modules\dsh-speak\engine\speak.ps1" "$env:USERPROFILE\.dsh\hooks\my-speak.ps1"
440
+ # macOS
441
+ cp ~/.dsh/profiles/web/node_modules/dsh-speak/engine/speak.sh ~/.dsh/hooks/my-speak.sh
442
+ ```
443
+
444
+ Then point the plugin at your copy in the `config` block:
445
+
446
+ > **Windows: keep the file's UTF-8 BOM.** `speak.ps1` is a UTF-8 script and Windows
447
+ > PowerShell 5.1 only knows that from the 3-byte BOM (`EF BB BF`) at the start; an
448
+ > editor that saves it without one makes the system ANSI code page decode it instead,
449
+ > and Chinese text inside the script turns to mojibake — the symptom is **silence or
450
+ > wrong trimming, with no error**. The shipped script keeps all of its *logic* ASCII-only
451
+ > for that reason, so a lost BOM only garbles the Chinese comments and the default
452
+ > prompt. After editing, check with
453
+ > `Get-Content -Encoding Byte -TotalCount 3 your-speak.ps1` (expect `239 187 191`), or
454
+ > run `node scripts/test-engine-static.js`.
455
+
456
+ ```yaml
457
+ # edit the top-level row (NOT the insert row — see "DSH plugin config")
458
+ - id: dsh-speak
459
+ name: 'dsh-speak'
460
+ config:
461
+ engine: 'C:/Users/<you>/.dsh/hooks/my-speak.ps1' # or ~/.dsh/hooks/my-speak.sh on macOS
462
+ ```
463
+
464
+ The plugin resolves the engine as `config.engine` → package engine → `~/.dsh/hooks/`,
465
+ so your copy wins. `npm update` only touches the package — your engine stays.
466
+
467
+ 2. **Edit the file inside `node_modules`** — works, but the next `npm update` overwrites it.
468
+
469
+ 3. **Fork the repo** — full control, publish your own package if you want.
470
+
471
+ ## Troubleshooting
472
+
473
+ | symptom | cause | fix |
474
+ | ------- | ----- | --- |
475
+ | No sound at all, no error | no natural voice enabled/installed | Win11: enable a natural voice in *Settings → Narrator / Speech*; Win10: install NaturalVoiceSAPIAdapter + a voice pack. Test `speak.ps1` directly |
476
+ | Long replies never spoken | adapter per-`Speak` character ceiling | already guarded at 300 chars — lower `-MaxChars` if needed |
477
+ | Narration stops after the first line | with `longTextMode: heading`, text over `maxChars` and no markdown heading made the engine speak only the first non-empty line (pre-1.8.0) | fixed in 1.8.0 (speaks a coherent opening instead); to change the policy use `message` mode or raise `maxChars` |
478
+ | `工具调用出错:Error: cannot read …` spoken | the "is this Chinese?" detail filter only checked for the presence of a CJK character, so a Chinese directory name inside an English error passed it (1.8.0 regression) | fixed in 1.8.0 — the detail now needs more Chinese characters than Latin letters |
479
+ | Emoji-heavy text silent | SAPI fails silently on emoji | already stripped by the engine |
480
+ | Plugin not loading | raw Windows path as plugin name | use the `file:///C:/…` URL form (installer does this) |
481
+ | Settings page missing after upgrade to 1.8.2 | profile patch row `disabled: true`, or the entry id is neither `dsh-speak` nor the legacy `speech-hook` | enable the row and use one of those two ids as its `id` |
482
+ | Settings page renders but every change bounces back | DSH < 0.1.7 (the `configForms` service replaced `settingsScope`) | update DSH, or stay on dsh-speak 1.8.1 |
483
+ | Every change bounces back on DSH 0.1.7+, and the plugin log says `已有实例在运行` | the entry is registered **twice** — e.g. the package is in `dsh.profile.bundles` *and* a hand-written `insert:` row exists. The id stops being unique, and the config editor rejects each write (`settings/rejected: Configuration for "dsh-speak" is overridden by a home patch or command-line overlay`); speech keeps working, which is what makes it confusing | keep exactly one registration path (see Option A), then restart |
484
+ | A setting applies immediately but is back to the old value after a restart | the `config:` block sits INSIDE the profile patch's `insert:` row — DSH's config editor rewrites config in place only on a top-level row, and silently rolls the nested write back (it still answers `ok: true`) | split it into the two-row shape from [DSH plugin config](#dsh-plugin-config); `python scripts/settings-ui-check.py` asserts the write lands |
485
+ | macOS: voice suddenly became "婷婷" | opening the "Spoken Content / Siri Voice" pane drifted the system voice | re-pick via Settings → Accessibility → Spoken Content → System Voice → ⓘ entry |
486
+ | macOS: no log at `/tmp` | `os.tmpdir()` is `/var/folders/.../T`, not `/tmp` | log is at `$TMPDIR/dsh-speech-hook.log` |
487
+
488
+ Plugin diagnostics: Windows `%TEMP%\dsh-speech-hook.log`; macOS `$TMPDIR/dsh-speech-hook.log`
489
+
490
+ ## Repository layout
491
+
492
+ ```
493
+ engine/ harness-agnostic speech engine (PowerShell + SAPI5 / bash + say)
494
+ speak.ps1 / speak.sh clean + speak (the only seam any adapter needs)
495
+ speech-prompt.ps1 blocking short announcement
496
+ speech-summary.ps1 blocking reply-summary announcement
497
+ adapters/
498
+ dsh/ DSH web plugin + one-command installer
499
+ speech-hook.js session-event trigger (throttle/cancel + optional events + FIFO speech queue + WebSocket + Config-projected settings form)
500
+ install.ps1 copies + registers + backs up
501
+ claude-code/
502
+ stop-hook.ps1 Claude Code Stop hook trigger
503
+ client/
504
+ client.js DSH browser bundle: turn-tail Speak/Stop button + Settings → dsh-speak settings page
505
+ docs/
506
+ DESIGN.md full design rationale, pitfalls, extension guide
507
+ scripts/ tests + manual dev helpers (not shipped in the npm package)
508
+ test-engine-static.js engine invariants: .ps1 BOM + PowerShell parse, .sh LF (also run by prepublishOnly)
509
+ test-engine-longtext.js long-text guard contract for BOTH engines (speak.ps1 -DryRun / speak.sh's perl)
510
+ test-speech-hook.js host plugin: event triggers, queue, tool-error detail filter
511
+ test-client-bundle.js browser bundle: slot registration + component rendering
512
+ test-settings-integration.js Config-projection wiring (volatile refs + live writes) + removed-API guard
513
+ session-log-dump.js read a DSH session log (manual: what text reached the engine)
514
+ settings-ui-check.py Playwright UI check (manual: needs a running, authenticated dsh)
515
+ dsh-events-check.py Playwright disclosure check (manual)
516
+ ```
517
+
518
+ ## Writing a new adapter
519
+
520
+ Three reference patterns exist: **event-stream** (DSH), **stop-hook** (Claude Code),
521
+ **agent-called** (`speech-summary.ps1` from a shell). In every case the adapter only
522
+ needs to: capture the *final reply text* → invoke the engine. See
523
+ [docs/DESIGN.md §7](docs/DESIGN.md#7-extending).
524
+
525
+ ## License
526
+
527
+ MIT — see [LICENSE](LICENSE).
528
+
529
+ [NaturalVoiceSAPIAdapter]: https://github.com/gexgd0419/NaturalVoiceSAPIAdapter