dsh-live-voice 0.2.0 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,23 @@
2
2
 
3
3
  All notable changes to DSH Live Voice are documented in this file.
4
4
 
5
+ ## [0.2.1] - 2026-09-19
6
+
7
+ ### Documentation
8
+
9
+ - Reframe the README around the local-first, hands-free experience, including continuous conversations, spoken structured questions, voice commands, and behavior across chat navigation.
10
+ - Expand package and README discovery keywords for hands-free voice, conversational AI, spoken prompts, Qwen3 speech engines, and Apple Silicon.
11
+
12
+ ### Changes
13
+
14
+ - Update existing development dependencies.
15
+
16
+ ### Bug Fixes
17
+
18
+ - Persist microphone enabled or muted state as a global preference and propagate settings changes to every active voice controller. New controllers now inherit the saved microphone state, and capability information is refreshed after synchronized settings changes.
19
+ - Keep microphone, conversation, and Speak controls available when DSH omits the legacy `uiSession.pendingInteractions` store. Structured-question voice handling remains inactive when that optional store is unavailable.
20
+ - Add regression coverage for global microphone-state normalization and for mounting voice controls without the legacy pending-interaction store.
21
+
5
22
  ## [0.2.0] - Unreleased
6
23
 
7
24
  ### Release Changes
@@ -68,7 +85,8 @@ This release expands conversation control, audio routing, speech filtering, and
68
85
  - Support browser speech synthesis and native macOS `say` output.
69
86
  - Add Live Voice controls, Whisper settings, build and browser-preview scripts, and an initial test suite.
70
87
 
71
- [0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...HEAD
88
+ [0.2.1]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.0...HEAD
89
+ [0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...v0.2.0
72
90
  [0.1.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.2...v0.1.0
73
91
  [0.0.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.1-alpha.1...v0.0.2
74
92
  [0.0.1-alpha.1]: https://github.com/victorwads/dsh-live-voice/tree/v0.0.1-alpha.1
package/README.md CHANGED
@@ -4,9 +4,9 @@
4
4
  [![dsh.pub registry status](https://dsh.pub/api/badges/victorwads/dsh-live-voice.svg)](https://dsh.pub/en/plugins/dsh-live-voice/)
5
5
  [![license](https://img.shields.io/badge/license-GPL--3.0--only-blue)](LICENSE)
6
6
 
7
- **Local-first speech recognition, voice output, and continuous voice conversations for DSH.**
7
+ **Local-first, hands-free voice conversations for DSH — speak, listen, answer prompts, and keep working without touching the computer.**
8
8
 
9
- DSH Live Voice coordinates the microphone, composer, assistant messages, and speech output in one plugin — without letting listening and speaking compete with each other.
9
+ DSH Live Voice coordinates the microphone, composer, assistant messages, structured questions, and speech output in one plugin — without letting listening and speaking compete with each other. Once voice conversation mode is running, DSH can narrate responses and questions, capture your spoken answers, and continue the conversation while your hands stay free.
10
10
 
11
11
  ## Install
12
12
 
@@ -29,7 +29,10 @@ A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [ds
29
29
  | | Capability |
30
30
  |---|---|
31
31
  | 🎙️ | Voice typing directly into the DSH composer |
32
- | 💬 | Continuous voice conversations with automatic assistant speech |
32
+ | 👐 | Hands-free conversations: speak, hear responses, and continue without touching the computer |
33
+ | ❓ | Spoken DSH structured questions with automatic capture and submission of your answer |
34
+ | 💬 | Continuous voice conversations that stay active while you navigate between chats |
35
+ | 🗣️ | Configurable voice commands for sending, queueing, clearing, muting, resuming, stopping speech, and ending a conversation |
33
36
  | 🧠 | Browser SpeechRecognition, local loopback whisper.cpp, or Qwen3-ASR on Apple MLX |
34
37
  | 🔊 | Browser speech synthesis, native macOS `say`, or Qwen3-TTS on Apple MLX |
35
38
  | ⏱️ | Manual or automatic sending after configurable silence |
@@ -40,6 +43,18 @@ A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [ds
40
43
  | 🔈 | Play individual assistant messages on demand |
41
44
  | 🏠 | Whisper audio reaches the local server only through the authenticated DSH host |
42
45
 
46
+ ## Hands-free experience
47
+
48
+ Start voice conversation mode and choose an automatic delivery mode to keep a conversation moving without returning to the keyboard. DSH Live Voice can:
49
+
50
+ 1. Listen for your next message and deliver it after the configured silence period.
51
+ 2. Read assistant responses aloud as they arrive.
52
+ 3. Narrate DSH structured questions, listen for your next spoken response, and submit it as a custom answer.
53
+ 4. Keep voice conversation mode active when you move between chats until you explicitly end it.
54
+ 5. Accept configurable spoken commands for common conversation controls.
55
+
56
+ The hands-free experience coordinates speech input and output locally when you select local engines. The DSH language model itself may still be remote.
57
+
43
58
  ## Conversation flow
44
59
 
45
60
  The assistant never starts automatic playback while you are speaking. If a response is already waiting — including another assistant message — it waits until you finish and the configured continuous-silence delay has passed.
@@ -108,4 +123,4 @@ When a compatible Qwen3 speech API is already running on the DSH host, choose **
108
123
 
109
124
  ## Keywords
110
125
 
111
- `dsh`, `dsh-plugin`, `deepseek-harness`, `local-first`, `local-voice`, `voice-assistant`, `voice-conversation`, `continuous-conversation`, `voice-dictation`, `speech-to-text`, `text-to-speech`, `speech-recognition`, `speech-synthesis`, `stt`, `tts`, `whisper`, `whisper-cpp`, `web-speech-api`, `macos-say`, `turn-taking`, `voice-interruption`, `silence-detection`
126
+ `dsh`, `dsh-plugin`, `deepseek-harness`, `local-first`, `local-voice`, `hands-free`, `hands-free-ai`, `hands-free-assistant`, `voice-control`, `voice-commands`, `voice-assistant`, `voice-conversation`, `conversational-ai`, `continuous-conversation`, `voice-dictation`, `spoken-prompts`, `speech-to-text`, `text-to-speech`, `speech-recognition`, `speech-synthesis`, `stt`, `tts`, `whisper`, `whisper-cpp`, `qwen3-asr`, `qwen3-tts`, `mlx`, `apple-silicon`, `web-speech-api`, `macos-say`, `turn-taking`, `voice-interruption`, `silence-detection`
package/lib/client.js CHANGED
@@ -114,6 +114,7 @@ var defaultSettings = Object.freeze({
114
114
  recognitionProcessLocally: true,
115
115
  recognitionAutoInstall: true,
116
116
  voiceDetectionPreset: "natural",
117
+ microphoneEnabled: true,
117
118
  announceAssistantMessages: true,
118
119
  interruptSpeechOnUserMessage: false,
119
120
  sendingMode: "manual",
@@ -149,6 +150,7 @@ function normalizeSettings(value) {
149
150
  recognitionProcessLocally: typeof source.recognitionProcessLocally === "boolean" ? source.recognitionProcessLocally : defaultSettings.recognitionProcessLocally,
150
151
  recognitionAutoInstall: typeof source.recognitionAutoInstall === "boolean" ? source.recognitionAutoInstall : defaultSettings.recognitionAutoInstall,
151
152
  voiceDetectionPreset: Object.hasOwn(voiceDetectionPresets, source.voiceDetectionPreset) ? source.voiceDetectionPreset : defaultSettings.voiceDetectionPreset,
153
+ microphoneEnabled: typeof source.microphoneEnabled === "boolean" ? source.microphoneEnabled : defaultSettings.microphoneEnabled,
152
154
  announceAssistantMessages: typeof source.announceAssistantMessages === "boolean" ? source.announceAssistantMessages : defaultSettings.announceAssistantMessages,
153
155
  interruptSpeechOnUserMessage: typeof source.interruptSpeechOnUserMessage === "boolean" ? source.interruptSpeechOnUserMessage : defaultSettings.interruptSpeechOnUserMessage,
154
156
  sendingMode: source.sendingMode === "automatic" ? "queue" : ["manual", "queue", "steer"].includes(source.sendingMode) ? source.sendingMode : defaultSettings.sendingMode,
@@ -271,7 +273,7 @@ var VoiceCoordinator = class {
271
273
  conversation: false,
272
274
  listening: false,
273
275
  recognizing: false,
274
- muted: false,
276
+ muted: normalizeSettings(settings).microphoneEnabled === false,
275
277
  speaking: false,
276
278
  paused: false,
277
279
  starting: false,
@@ -338,7 +340,12 @@ var VoiceCoordinator = class {
338
340
  autoInstallLocalPack: settings.recognitionAutoInstall,
339
341
  voiceDetectionPreset: settings.voiceDetectionPreset
340
342
  });
341
- this.patch({ settings, error: null });
343
+ this.patch({
344
+ settings,
345
+ muted: settings.microphoneEnabled === false,
346
+ recognizing: settings.microphoneEnabled === false ? false : this.snapshot.recognizing,
347
+ error: null
348
+ });
342
349
  if (settings.sendingMode === "manual") this.cancelAutoSend();
343
350
  if (Object.hasOwn(next, "announceAssistantMessages") && !settings.announceAssistantMessages) {
344
351
  this.queue = [];
@@ -401,16 +408,17 @@ var VoiceCoordinator = class {
401
408
  return this.startListening(true);
402
409
  }
403
410
  muteListening() {
404
- if (!this.snapshot.settings.voiceCommandsEnabled) return this.stopListening();
405
411
  this.cancelAutoSend();
406
412
  this.composer.setDraft(this.transcript.update(this.composer.getDraft(), "", true));
407
413
  this.transcript.reset();
408
- this.patch({ muted: true, recognizing: false });
414
+ this.updateSettings({ microphoneEnabled: false });
415
+ if (!this.snapshot.settings.voiceCommandsEnabled) return this.stopListening();
409
416
  }
410
417
  resumeListeningInput() {
418
+ this.transcript.reset();
419
+ this.updateSettings({ microphoneEnabled: true });
411
420
  if (this.snapshot.listening || this.snapshot.starting) {
412
- this.transcript.reset();
413
- this.patch({ muted: false, recognizing: false });
421
+ this.patch({ recognizing: false });
414
422
  return;
415
423
  }
416
424
  return this.startListening(this.snapshot.conversation);
@@ -3244,7 +3252,7 @@ function apply(ctx) {
3244
3252
  });
3245
3253
  const pendingInteractions = ctx.uiSession.pendingInteractions;
3246
3254
  const refreshPendingQuestion = (baseline = false) => {
3247
- if (disposed || entry.closed) return;
3255
+ if (disposed || entry.closed || !pendingInteractions) return;
3248
3256
  const interaction = pendingInteractions.getSnapshot().get(sessionId);
3249
3257
  const key2 = interaction?.kind === "question" ? interaction.key : null;
3250
3258
  if (!key2 || entry.questionCapture && entry.questionCapture.interaction.key !== key2) {
@@ -3264,22 +3272,28 @@ function apply(ctx) {
3264
3272
  run(controller, controller.speak(text, key2));
3265
3273
  }
3266
3274
  };
3267
- refreshPendingQuestion(true);
3268
- entry.unsubscribePendingQuestion = pendingInteractions.subscribe(refreshPendingQuestion);
3275
+ if (pendingInteractions) {
3276
+ refreshPendingQuestion(true);
3277
+ entry.unsubscribePendingQuestion = pendingInteractions.subscribe(refreshPendingQuestion);
3278
+ }
3269
3279
  const update = controller.updateSettings.bind(controller);
3270
- controller.updateSettings = (next) => {
3271
- if (disposed || entry.closed) return;
3280
+ entry.applySettings = (next) => {
3272
3281
  update(next);
3273
3282
  engineBrowser.lang = controller.getSnapshot().settings.lang;
3274
3283
  engineQwen.lang = controller.getSnapshot().settings.lang;
3284
+ run(controller, controller.refreshCapabilities());
3285
+ };
3286
+ controller.updateSettings = (next) => {
3287
+ if (disposed || entry.closed) return;
3288
+ entry.applySettings(next);
3289
+ const settings2 = controller.getSnapshot().settings;
3275
3290
  try {
3276
- localStorage.setItem(
3277
- "dsh-live-voice.settings",
3278
- JSON.stringify(controller.getSnapshot().settings)
3279
- );
3291
+ localStorage.setItem("dsh-live-voice.settings", JSON.stringify(settings2));
3280
3292
  } catch {
3281
3293
  }
3282
- run(controller, controller.refreshCapabilities());
3294
+ for (const other of controllers.values()) {
3295
+ if (other !== entry && !other.closed) other.applySettings(settings2);
3296
+ }
3283
3297
  };
3284
3298
  for (const method of ["stopListening", "cancelDictation", "endConversation", "stopSpeech"]) {
3285
3299
  const original = controller[method].bind(controller);
@@ -3419,7 +3433,7 @@ function apply(ctx) {
3419
3433
  if (entry.closed) continue;
3420
3434
  if (entry.controller.getSnapshot().settings.recognitionEngine !== settings2.recognitionEngine)
3421
3435
  entry.controller.replaceRecognition(recognitionFor(settings2, entry.controller.meter));
3422
- entry.controller.updateSettings(settings2);
3436
+ entry.applySettings(settings2);
3423
3437
  run(entry.controller, entry.controller.refreshCapabilities());
3424
3438
  }
3425
3439
  };
package/lib/server.js CHANGED
@@ -460,6 +460,7 @@ var defaultSettings = Object.freeze({
460
460
  recognitionProcessLocally: true,
461
461
  recognitionAutoInstall: true,
462
462
  voiceDetectionPreset: "natural",
463
+ microphoneEnabled: true,
463
464
  announceAssistantMessages: true,
464
465
  interruptSpeechOnUserMessage: false,
465
466
  sendingMode: "manual",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dsh-live-voice",
3
- "version": "0.2.0",
3
+ "version": "0.2.1",
4
4
  "type": "module",
5
5
  "scripts": {
6
6
  "typecheck": "tsc --noEmit",
@@ -20,10 +20,17 @@
20
20
  "deepseek-harness",
21
21
  "local-first",
22
22
  "local-voice",
23
+ "hands-free",
24
+ "hands-free-ai",
25
+ "hands-free-assistant",
26
+ "voice-control",
27
+ "voice-commands",
23
28
  "voice-assistant",
24
29
  "voice-conversation",
30
+ "conversational-ai",
25
31
  "continuous-conversation",
26
32
  "voice-dictation",
33
+ "spoken-prompts",
27
34
  "speech-to-text",
28
35
  "text-to-speech",
29
36
  "speech-recognition",
@@ -68,16 +75,16 @@
68
75
  "url": "git+https://github.com/victorwads/dsh-live-voice.git"
69
76
  },
70
77
  "devDependencies": {
71
- "@types/node": "^26.5.1",
72
- "@types/react": "^19.3.0",
73
- "@types/react-dom": "^19.3.0",
74
- "esbuild": "^0.28.2",
75
- "jsdom": "^30.0.1",
76
- "playwright-core": "^1.63.0",
77
- "prettier": "3.9.6",
78
- "react": "^19.3.0",
79
- "react-dom": "^19.3.0",
80
- "typescript": "^7.0.2"
78
+ "@types/node": "26.6.2",
79
+ "@types/react": "19.3.0",
80
+ "@types/react-dom": "19.3.0",
81
+ "esbuild": "0.28.2",
82
+ "jsdom": "30.1.0",
83
+ "playwright-core": "1.63.0",
84
+ "prettier": "3.9.8",
85
+ "react": "19.3.0",
86
+ "react-dom": "19.3.0",
87
+ "typescript": "7.0.2"
81
88
  },
82
89
  "main": "./lib/server.js",
83
90
  "exports": {
@@ -172,9 +172,11 @@ export function apply(ctx) {
172
172
  refresh();
173
173
  for (const listener of entry.chatListeners) listener();
174
174
  });
175
+ // DSH 0.1.6 may omit this legacy store. Keep the remaining voice controls
176
+ // available while structured-question integration is unavailable.
175
177
  const pendingInteractions = ctx.uiSession.pendingInteractions;
176
178
  const refreshPendingQuestion = (baseline = false) => {
177
- if (disposed || entry.closed) return;
179
+ if (disposed || entry.closed || !pendingInteractions) return;
178
180
  const interaction = pendingInteractions.getSnapshot().get(sessionId);
179
181
  const key = interaction?.kind === 'question' ? interaction.key : null;
180
182
  if (!key || (entry.questionCapture && entry.questionCapture.interaction.key !== key)) {
@@ -194,21 +196,27 @@ export function apply(ctx) {
194
196
  run(controller, controller.speak(text, key));
195
197
  }
196
198
  };
197
- refreshPendingQuestion(true);
198
- entry.unsubscribePendingQuestion = pendingInteractions.subscribe(refreshPendingQuestion);
199
+ if (pendingInteractions) {
200
+ refreshPendingQuestion(true);
201
+ entry.unsubscribePendingQuestion = pendingInteractions.subscribe(refreshPendingQuestion);
202
+ }
199
203
  const update = controller.updateSettings.bind(controller);
200
- controller.updateSettings = (next) => {
201
- if (disposed || entry.closed) return;
204
+ entry.applySettings = (next) => {
202
205
  update(next);
203
206
  engineBrowser.lang = controller.getSnapshot().settings.lang;
204
207
  engineQwen.lang = controller.getSnapshot().settings.lang;
208
+ run(controller, controller.refreshCapabilities());
209
+ };
210
+ controller.updateSettings = (next) => {
211
+ if (disposed || entry.closed) return;
212
+ entry.applySettings(next);
213
+ const settings = controller.getSnapshot().settings;
205
214
  try {
206
- localStorage.setItem(
207
- 'dsh-live-voice.settings',
208
- JSON.stringify(controller.getSnapshot().settings),
209
- );
215
+ localStorage.setItem('dsh-live-voice.settings', JSON.stringify(settings));
210
216
  } catch {}
211
- run(controller, controller.refreshCapabilities());
217
+ for (const other of controllers.values()) {
218
+ if (other !== entry && !other.closed) other.applySettings(settings);
219
+ }
212
220
  };
213
221
  // Cancel only this entry's queued acquisition. Global cancellation here would
214
222
  // invalidate a newer session while its predecessor is being unmounted.
@@ -364,7 +372,7 @@ export function apply(ctx) {
364
372
  entry.controller.getSnapshot().settings.recognitionEngine !== settings.recognitionEngine
365
373
  )
366
374
  entry.controller.replaceRecognition(recognitionFor(settings, entry.controller.meter));
367
- entry.controller.updateSettings(settings);
375
+ entry.applySettings(settings);
368
376
  run(entry.controller, entry.controller.refreshCapabilities());
369
377
  }
370
378
  };
@@ -46,7 +46,7 @@ export class VoiceCoordinator {
46
46
  conversation: false,
47
47
  listening: false,
48
48
  recognizing: false,
49
- muted: false,
49
+ muted: normalizeSettings(settings).microphoneEnabled === false,
50
50
  speaking: false,
51
51
  paused: false,
52
52
  starting: false,
@@ -116,7 +116,12 @@ export class VoiceCoordinator {
116
116
  autoInstallLocalPack: settings.recognitionAutoInstall,
117
117
  voiceDetectionPreset: settings.voiceDetectionPreset,
118
118
  });
119
- this.patch({ settings, error: null });
119
+ this.patch({
120
+ settings,
121
+ muted: settings.microphoneEnabled === false,
122
+ recognizing: settings.microphoneEnabled === false ? false : this.snapshot.recognizing,
123
+ error: null,
124
+ });
120
125
  if (settings.sendingMode === 'manual') this.cancelAutoSend();
121
126
  if (Object.hasOwn(next, 'announceAssistantMessages') && !settings.announceAssistantMessages) {
122
127
  this.queue = [];
@@ -185,16 +190,17 @@ export class VoiceCoordinator {
185
190
  return this.startListening(true);
186
191
  }
187
192
  muteListening() {
188
- if (!this.snapshot.settings.voiceCommandsEnabled) return this.stopListening();
189
193
  this.cancelAutoSend();
190
194
  this.composer.setDraft(this.transcript.update(this.composer.getDraft(), '', true));
191
195
  this.transcript.reset();
192
- this.patch({ muted: true, recognizing: false });
196
+ this.updateSettings({ microphoneEnabled: false });
197
+ if (!this.snapshot.settings.voiceCommandsEnabled) return this.stopListening();
193
198
  }
194
199
  resumeListeningInput() {
200
+ this.transcript.reset();
201
+ this.updateSettings({ microphoneEnabled: true });
195
202
  if (this.snapshot.listening || this.snapshot.starting) {
196
- this.transcript.reset();
197
- this.patch({ muted: false, recognizing: false });
203
+ this.patch({ recognizing: false });
198
204
  return;
199
205
  }
200
206
  return this.startListening(this.snapshot.conversation);
@@ -36,6 +36,7 @@ export const defaultSettings = Object.freeze({
36
36
  recognitionProcessLocally: true,
37
37
  recognitionAutoInstall: true,
38
38
  voiceDetectionPreset: 'natural',
39
+ microphoneEnabled: true,
39
40
  announceAssistantMessages: true,
40
41
  interruptSpeechOnUserMessage: false,
41
42
  sendingMode: 'manual',
@@ -93,6 +94,10 @@ export function normalizeSettings(value) {
93
94
  voiceDetectionPreset: Object.hasOwn(voiceDetectionPresets, source.voiceDetectionPreset)
94
95
  ? source.voiceDetectionPreset
95
96
  : defaultSettings.voiceDetectionPreset,
97
+ microphoneEnabled:
98
+ typeof source.microphoneEnabled === 'boolean'
99
+ ? source.microphoneEnabled
100
+ : defaultSettings.microphoneEnabled,
96
101
  announceAssistantMessages:
97
102
  typeof source.announceAssistantMessages === 'boolean'
98
103
  ? source.announceAssistantMessages