dsh-live-voice 0.2.2 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -1
- package/DEVELOPMENT.md +13 -0
- package/LICENSE +201 -674
- package/PLAN.md +28 -0
- package/README.md +65 -110
- package/docs/ARCHITECTURE.md +40 -0
- package/docs/CHOOSING-AN-ENGINE.md +120 -0
- package/docs/CONFIGURATION.md +149 -0
- package/docs/REVIEW.md +35 -0
- package/docs/VOICE-LIFECYCLE.md +59 -0
- package/lib/client.js +3652 -1576
- package/lib/server.js +641 -453
- package/package.json +17 -6
- package/scripts/build.ts +16 -4
- package/scripts/preview-ui.ts +1 -1
- package/src/{client/index.ts → app/client/apply.tsx} +152 -59
- package/src/app/client/i18n/DshLanguageBoundary.tsx +42 -0
- package/src/app/client/i18n/catalogs/base.ts +231 -0
- package/src/app/client/i18n/catalogs/en.ts +258 -0
- package/src/app/client/i18n/catalogs/es.ts +279 -0
- package/src/app/client/i18n/catalogs/fr.ts +281 -0
- package/src/app/client/i18n/catalogs/hi.ts +263 -0
- package/src/app/client/i18n/catalogs/index.ts +18 -0
- package/src/app/client/i18n/catalogs/pt-BR.ts +270 -0
- package/src/app/client/i18n/catalogs/zh.ts +248 -0
- package/src/app/client/i18n/index.ts +4 -0
- package/src/app/client/i18n/registerDshLocales.ts +43 -0
- package/src/app/client/i18n/runtime.tsx +58 -0
- package/src/app/client/index.ts +3 -0
- package/src/app/client/registerSlots.tsx +30 -0
- package/src/app/client/slotDefinitions.ts +32 -0
- package/src/app/server/apply.ts +479 -0
- package/src/app/server/index.ts +2 -0
- package/src/app/server/registerRoutes.ts +11 -0
- package/src/modules/conversation/components/ConversationControls.tsx +40 -0
- package/src/modules/conversation/components/ConversationStatusBar.tsx +155 -0
- package/src/modules/conversation/components/DeliveryModeButton.tsx +43 -0
- package/src/modules/conversation/components/MicrophoneButton.tsx +6 -0
- package/src/modules/conversation/components/PlaybackControls.tsx +41 -0
- package/src/modules/conversation/components/SpeakButton.tsx +30 -0
- package/src/modules/conversation/components/Waveform.tsx +68 -0
- package/src/modules/conversation/components/conversationStatus.ts +17 -0
- package/src/modules/conversation/components/createConversationComponents.tsx +505 -0
- package/src/modules/conversation/components/index.ts +8 -0
- package/src/modules/conversation/hooks/index.ts +2 -0
- package/src/modules/conversation/hooks/useConversationActions.ts +27 -0
- package/src/modules/conversation/hooks/useConversationController.ts +15 -0
- package/src/modules/conversation/index.ts +3 -0
- package/src/{client → modules/conversation/models}/chat.ts +2 -1
- package/src/modules/conversation/models/index.ts +1 -0
- package/src/{core → modules/core}/coordinator.ts +133 -7
- package/src/modules/core/index.ts +6 -0
- package/src/{engines/qwen-http-host.ts → modules/core/qwen/QwenHttpHost.ts} +2 -2
- package/src/modules/core/qwen/QwenSettings.tsx +148 -0
- package/src/{core → modules/core}/settings.ts +23 -0
- package/src/modules/recognition/components/RecognitionCapabilityStatus.tsx +19 -0
- package/src/{engines/recognition/qwen-http.ts → modules/recognition/engines/qwen/QwenRecognitionEngine.ts} +1 -1
- package/src/{engines/recognition/whisper-http.ts → modules/recognition/engines/whisper/WhisperRecognitionEngine.ts} +1 -1
- package/src/modules/recognition/engines/whisper/WhisperSettings.tsx +150 -0
- package/src/modules/recognition/index.ts +4 -0
- package/src/modules/recognition/qwen/QwenRecognitionSettings.tsx +4 -0
- package/src/modules/settings/components/LiveVoiceSettings.tsx +72 -0
- package/src/modules/settings/components/SettingsHeader.tsx +49 -0
- package/src/modules/settings/components/VersionBadges.tsx +39 -0
- package/src/modules/settings/components/createLiveVoiceSettings.tsx +8 -0
- package/src/modules/settings/components/index.ts +4 -0
- package/src/modules/settings/hooks/index.ts +3 -0
- package/src/modules/settings/hooks/useAudioDevices.ts +38 -0
- package/src/modules/settings/hooks/useLiveVoiceSettings.ts +4 -0
- package/src/modules/settings/hooks/useReleaseStatus.ts +17 -0
- package/src/modules/settings/index.ts +4 -0
- package/src/modules/settings/models/index.ts +1 -0
- package/src/modules/settings/models/settingsStorage.ts +14 -0
- package/src/modules/settings/sections/conversation/ConversationDelaySettings.tsx +24 -0
- package/src/modules/settings/sections/conversation/ConversationSettingsSection.tsx +23 -0
- package/src/modules/settings/sections/conversation/DeliverySettings.tsx +38 -0
- package/src/modules/settings/sections/conversation/HoldToTalkSettings.tsx +20 -0
- package/src/modules/settings/sections/conversation/VoiceModeSettings.tsx +31 -0
- package/src/modules/settings/sections/conversation/index.ts +5 -0
- package/src/modules/settings/sections/index.ts +3 -0
- package/src/modules/settings/sections/recognition/RecognitionEngineSettings.tsx +109 -0
- package/src/modules/settings/sections/recognition/RecognitionFilterSettings.tsx +33 -0
- package/src/modules/settings/sections/recognition/RecognitionSettingsSection.tsx +20 -0
- package/src/modules/settings/sections/recognition/RecognitionStatus.tsx +19 -0
- package/src/modules/settings/sections/recognition/SilenceDetectionSettings.tsx +58 -0
- package/src/modules/settings/sections/recognition/VoiceCommandSettings.tsx +41 -0
- package/src/modules/settings/sections/recognition/index.ts +5 -0
- package/src/modules/settings/sections/speak/OutputFilterSettings.tsx +39 -0
- package/src/modules/settings/sections/speak/PlaybackPolicySettings.tsx +26 -0
- package/src/modules/settings/sections/speak/SpeakSettingsSection.tsx +16 -0
- package/src/modules/settings/sections/speak/SpeechAdvancedSettings.tsx +81 -0
- package/src/modules/settings/sections/speak/SpeechEngineSettings.tsx +80 -0
- package/src/modules/settings/sections/speak/index.ts +4 -0
- package/src/modules/settings/services/releases.ts +117 -0
- package/src/modules/speak/engines/audio/HostAudioEngine.ts +167 -0
- package/src/modules/speak/engines/audio/M4aAacTranscoder.ts +58 -0
- package/src/modules/speak/engines/qwen/QwenSpeakingEngine.ts +14 -0
- package/src/{engines/speaking/say.ts → modules/speak/engines/say/SaySpeakingEngine.ts} +26 -4
- package/src/modules/speak/index.ts +5 -0
- package/src/modules/speak/qwen/QwenSpeakingSettings.tsx +4 -0
- package/src/modules/speak/services/speechQueue.ts +2 -0
- package/src/server.ts +9 -365
- package/src/shared/design-system/buttons/IconButton.tsx +33 -0
- package/src/shared/design-system/buttons/PillButton.tsx +11 -0
- package/src/shared/design-system/buttons/ToggleButton.tsx +6 -0
- package/src/shared/design-system/buttons/index.ts +3 -0
- package/src/shared/design-system/feedback/ErrorMessage.tsx +27 -0
- package/src/shared/design-system/feedback/StatusBadge.tsx +9 -0
- package/src/shared/design-system/feedback/StatusMessage.tsx +4 -0
- package/src/shared/design-system/feedback/index.ts +3 -0
- package/src/shared/design-system/forms/CheckboxField.tsx +20 -0
- package/src/shared/design-system/forms/NumberField.tsx +12 -0
- package/src/shared/design-system/forms/SelectField.tsx +20 -0
- package/src/shared/design-system/forms/TextAreaField.tsx +12 -0
- package/src/shared/design-system/forms/TextField.tsx +12 -0
- package/src/shared/design-system/forms/index.ts +5 -0
- package/src/shared/design-system/icons/Icon.tsx +20 -0
- package/src/shared/design-system/icons/icons.ts +16 -0
- package/src/shared/design-system/icons/index.ts +2 -0
- package/src/shared/design-system/index.ts +5 -0
- package/src/shared/design-system/layout/SettingsCard.tsx +9 -0
- package/src/shared/design-system/layout/SettingsSection.tsx +9 -0
- package/src/shared/design-system/layout/SettingsSubcard.tsx +14 -0
- package/src/shared/design-system/layout/SettingsTabs.tsx +54 -0
- package/src/shared/design-system/layout/index.ts +4 -0
- package/src/{client/styles.ts → styles/index.ts} +3 -1
- package/src/client/components.ts +0 -1080
- package/src/client/qwen-settings.ts +0 -144
- package/src/client/whisper-settings.ts +0 -147
- package/src/engines/speaking/qwen-http.ts +0 -124
- /package/src/{core → modules/core}/filters.ts +0 -0
- /package/src/{core → modules/core}/microphone.ts +0 -0
- /package/src/{core → modules/core}/ownership.ts +0 -0
- /package/src/{core → modules/core}/transcript.ts +0 -0
- /package/src/{engines/recognition/browser.ts → modules/recognition/engines/browser/BrowserRecognitionEngine.ts} +0 -0
- /package/src/{engines/recognition/whisper-http-host.ts → modules/recognition/engines/whisper/whisperRecognitionHost.ts} +0 -0
- /package/src/{engines/speaking/browser.ts → modules/speak/engines/browser/BrowserSpeakingEngine.ts} +0 -0
- /package/src/{engines/speaking/say-client.ts → modules/speak/engines/say/sayClient.ts} +0 -0
package/PLAN.md
CHANGED
|
@@ -186,6 +186,34 @@ Pausing playback need not stop text generation. If the user introduces a new req
|
|
|
186
186
|
|
|
187
187
|
For spoken streaming, collect suitable text segments, synthesize them, and play them in order. Cancellation must invalidate pending work so delayed results from an old response cannot start playing later. Sentence boundaries, latency, buffering limits, and error recovery need evaluation.
|
|
188
188
|
|
|
189
|
+
### Dual-audio meeting transcription (planned)
|
|
190
|
+
|
|
191
|
+
Add an opt-in meeting mode with two independent audio-input pipelines. The primary pipeline remains the user’s normal microphone and labels finalized transcription with **“Me:”**. A secondary pipeline captures meeting participants from a user-selected shared-screen/system-audio stream, rather than assuming a second physical microphone, and labels finalized transcription with **“Them:”**. Each pipeline needs independent capture permission, meter/waveform, recognition work, segmentation, pending/error state, and cancellation; it must be possible to run either one alone or both together.
|
|
192
|
+
|
|
193
|
+
Meeting mode appends final segments from both pipelines to the current composer in arrival order, preserving the speaker labels so the editable draft becomes a live meeting transcript: “Me: …”, “Them: …”, and so on. It must not use voice commands or automatically send messages. The user remains free to edit the composer and manually submit any selected portion as a question to the agent, retaining the preceding live transcript as context for code reviews or other meeting discussion.
|
|
194
|
+
|
|
195
|
+
The design must account for browser screen/system-audio capture support, permission and device-selection UX, concurrent transcription ordering, overlap/echo between microphone and shared audio, and cleanup when either capture ends. It must never imply that system-audio capture is universally supported or silently capture meeting audio.
|
|
196
|
+
|
|
197
|
+
#### Expected structural change
|
|
198
|
+
|
|
199
|
+
This is a substantial input-flow refactor, not a complete plugin rewrite. The current session coordinator owns one microphone meter, one recognition instance, and one incremental composer transcript. Refactor it so the coordinator continues to own the conversation, composer integration, and shared policies, while reusable input pipelines own capture, meter/waveform, recognition, segmentation, pending work, errors, cancellation, and final-segment delivery.
|
|
200
|
+
|
|
201
|
+
The target shape is one VoiceCoordinator with an independent primary microphone pipeline and an optional shared-system-audio pipeline. The existing microphone adapter must accept an externally supplied MediaStream as well as getUserMedia(), because system audio originates through explicit getDisplayMedia() sharing. Composer delivery must serialize concurrent final segments and use an explicit ordering policy, likely segment-end time plus a small buffer, because recognition completion order can differ from speaking order.
|
|
202
|
+
|
|
203
|
+
The shared-audio pipeline must remain policy-isolated: it does not execute voice commands, trigger automatic delivery, or participate in microphone/TTS interruption as though meeting participants were the user. Initial implementation need not change the TTS queue; the meeting transcript stays manually editable and manually submitted. The planned backend-owned automatic-speech queue is complementary and will later make cross-chat output state independent from mounted frontend sessions.
|
|
204
|
+
|
|
205
|
+
### Settings changes should not restart voice (planned bug fix)
|
|
206
|
+
|
|
207
|
+
Changing a Live Voice preference currently stops active voice resources and makes the plugin behave as if it were restarting, even for values that can be applied live. Settings updates must be classified by effect: presentation and policy values should update immediately without interrupting capture, playback, queued speech, or the current conversation; engine, device, and other resource-boundary changes may require a targeted replacement only when necessary. The interface must state when a particular change will take effect on the next operation rather than silently resetting the whole plugin.
|
|
208
|
+
|
|
209
|
+
### Backend-owned automatic speech queue (planned)
|
|
210
|
+
|
|
211
|
+
The browser currently observes visible chat nodes and decides which assistant text enters its automatic-speech queue. This makes queue state dependent on the focused conversation and mounted frontend: switching chats can discard or duplicate state, and opening an older conversation can incorrectly announce historical assistant messages as if they were new.
|
|
212
|
+
|
|
213
|
+
Move automatic-speech admission, streaming-segment tracking, ordering, cancellation, and queue state to the DSH host. The host must retain per-chat/turn provenance and a durable high-water mark for consumed assistant text, then dispatch synthesized audio or playback work to the appropriate connected browser client. A client must be able to receive an announcement for a non-focused chat with clear contextual phrasing, for example identifying the chat before its message.
|
|
214
|
+
|
|
215
|
+
When several chats produce assistant output, the host must serialize announcements according to an explicit cross-chat policy: finish the active announcement, group pending segments by chat where appropriate, and announce one chat at a time rather than interleaving unrelated messages. The design must define how multiple browser clients are selected, how inactive clients are handled, and whether one client owns playback at a time. It must preserve correct queue state across navigation, avoid creating new work from historical chat content when an old chat is opened, and invalidate queued/in-flight speech on cancellation or an obsolete turn.
|
|
216
|
+
|
|
189
217
|
### Engine selection and interface
|
|
190
218
|
|
|
191
219
|
Provide one place to select microphone, output device, conversation mode, STT engine, TTS engine, and shortcuts. Keep provider-specific complexity behind clear options without hiding important costs, permissions, or data transmission.
|
package/README.md
CHANGED
|
@@ -1,137 +1,92 @@
|
|
|
1
1
|
# DSH Live Voice
|
|
2
2
|
|
|
3
3
|
[](https://www.npmjs.com/package/dsh-live-voice)
|
|
4
|
-
[](https://github.com/deepseek-ai/deepseek-harness/releases/tag/v0.1.6-alpha.2)
|
|
5
|
+
[](LICENSE)
|
|
6
6
|
|
|
7
|
-
**A local-first, hands-free voice assistant plugin for DeepSeek Harness (DSH)
|
|
7
|
+
**A local-first, hands-free voice assistant plugin for DeepSeek Harness (DSH).**
|
|
8
|
+
*Built in Brazil 🇧🇷 and tested daily with Brazilian Portuguese on macOS.*
|
|
8
9
|
|
|
9
|
-
|
|
10
|
+
Speak, listen, answer prompts, and code without touching your keyboard. DSH Live Voice coordinates speech-to-text (STT) and text-to-speech (TTS) into a single, conflict-free conversational flow.
|
|
10
11
|
|
|
11
|
-
|
|
12
|
+
---
|
|
12
13
|
|
|
13
|
-
|
|
14
|
+
## ⚡ Quick Start
|
|
14
15
|
|
|
15
|
-
|
|
16
|
+
Install with `npx dsh` into your DSH web profile:
|
|
16
17
|
|
|
17
|
-
|
|
18
|
+
```sh
|
|
19
|
+
npx dsh plugin add --profile web dsh-live-voice
|
|
20
|
+
```
|
|
18
21
|
|
|
19
|
-
|
|
22
|
+
Open **DSH Settings → Live Voice** after your next DSH startup.
|
|
20
23
|
|
|
21
|
-
|
|
24
|
+
---
|
|
22
25
|
|
|
23
|
-
|
|
26
|
+
## 🌍 Interface Languages
|
|
24
27
|
|
|
25
|
-
|
|
26
|
-
npx dshpub add victorwads/dsh-live-voice --profile web
|
|
27
|
-
```
|
|
28
|
+
DSH Live Voice’s plugin interface is translated into the following languages. This refers to the visible plugin UI—not speech-recognition or text-to-speech language support.
|
|
28
29
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
|
51
|
-
|
|
|
52
|
-
|
|
|
53
|
-
|
|
|
54
|
-
|
|
|
55
|
-
|
|
|
56
|
-
|
|
57
|
-
## Hands-free experience
|
|
58
|
-
|
|
59
|
-
Start voice conversation mode and choose an automatic delivery mode to keep a conversation moving without returning to the keyboard. DSH Live Voice can:
|
|
60
|
-
|
|
61
|
-
1. Listen for your next message and deliver it after the configured silence period.
|
|
62
|
-
2. Read assistant responses aloud as they arrive.
|
|
63
|
-
3. Narrate DSH structured questions, listen for your next spoken response, and submit it as a custom answer.
|
|
64
|
-
4. Keep voice conversation mode active when you move between chats until you explicitly end it.
|
|
65
|
-
5. Accept configurable spoken commands for common conversation controls.
|
|
66
|
-
6. Start temporary push-to-talk from anywhere on the page by holding Control, even when the voice bar is off; release to finish queued transcription and automatic delivery, or press Escape to cancel.
|
|
67
|
-
|
|
68
|
-
The hands-free experience coordinates speech input and output locally when you select local engines. The DSH language model itself may still be remote.
|
|
69
|
-
|
|
70
|
-
## Conversation flow
|
|
71
|
-
|
|
72
|
-
The assistant never starts automatic playback while you are speaking. If a response is already waiting — including another assistant message — it waits until you finish and the configured continuous-silence delay has passed.
|
|
73
|
-
|
|
74
|
-
### Speakers — gated listening (default)
|
|
75
|
-
|
|
76
|
-
Listening and playback take turns so the assistant does not hear its own voice.
|
|
77
|
-
|
|
78
|
-
```mermaid
|
|
79
|
-
sequenceDiagram
|
|
80
|
-
participant Interface
|
|
81
|
-
actor Você
|
|
82
|
-
actor Assistente
|
|
83
|
-
|
|
84
|
-
Note over Interface,Assistente: Aguardando você falar…
|
|
85
|
-
activate Você
|
|
86
|
-
Você->>Assistente: Começa a falar
|
|
87
|
-
Assistente-->>Você: Escuta enquanto você fala
|
|
88
|
-
Note over Interface,Assistente: Você parou de falar
|
|
89
|
-
opt Envio manual
|
|
90
|
-
Você->>Interface: Revisa a mensagem reconhecida
|
|
91
|
-
Interface-->>Você: Envia quando estiver pronto
|
|
92
|
-
end
|
|
93
|
-
opt Envio automático
|
|
94
|
-
Note over Você,Assistente: Envia após a contagem de silêncio
|
|
95
|
-
end
|
|
96
|
-
deactivate Você
|
|
97
|
-
Você->>Assistente: Entrega sua mensagem
|
|
98
|
-
Assistente-->>Você: Resposta pronta — aguarda silêncio contínuo
|
|
99
|
-
Assistente-->>Você: Para de escutar
|
|
100
|
-
activate Assistente
|
|
101
|
-
Assistente->>Você: Fala a resposta em voz alta
|
|
102
|
-
deactivate Assistente
|
|
103
|
-
Note over Interface,Assistente: Escutando novamente — aguardando você falar…
|
|
104
|
-
```
|
|
30
|
+
- 🇺🇸 **English**
|
|
31
|
+
- 🇧🇷 **Portuguese (Brazil)**
|
|
32
|
+
- 🇪🇸 **Spanish**
|
|
33
|
+
- 🇫🇷 **French**
|
|
34
|
+
- 🇮🇳 **Hindi**
|
|
35
|
+
- 🇨🇳 **Chinese**
|
|
36
|
+
|
|
37
|
+
---
|
|
38
|
+
|
|
39
|
+
## 📚 Documentation
|
|
40
|
+
|
|
41
|
+
Detailed guides for deep-diving into engines and configurations:
|
|
42
|
+
|
|
43
|
+
- ⚙️ **[Configuration & Conversation Flow Guide](docs/CONFIGURATION.md)** — Settings overview, speaker vs. headphone modes, sequence diagrams, silence delays, and external engine setup.
|
|
44
|
+
- 🧠 **[Choosing a Speech Recognition Engine](docs/CHOOSING-AN-ENGINE.md)** — Comparison between Browser STT, Qwen3 ASR, and Whisper, with RAM footprints and OS compatibility.
|
|
45
|
+
- 📖 **[The Story Behind the Project](HISTORY.md)** — Why this project was built and the human story behind coordinating voice.
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
49
|
+
## 🎯 Which Speech Engine Should I Use?
|
|
50
|
+
|
|
51
|
+
| Scenario | Recommendation | RAM | Why |
|
|
52
|
+
| --- | --- | --- | --- |
|
|
53
|
+
| 🇧🇷 **Portuguese on macOS** | **Qwen3 ASR (HTTP API)** | ~3 GB | Best accuracy in daily maintainer use. Whisper is second choice. |
|
|
54
|
+
| 🇺🇸 **English on macOS** | **Browser SpeechRecognition** | ~0 GB | Built-in macOS/browser API. Fast, zero extra RAM. |
|
|
55
|
+
| 🪟 **Windows** | **Qwen3 ASR** or **Whisper HTTP** | ~2–3 GB | Recommended starting point; Windows browser STT varies. |
|
|
56
|
+
| 🌐 **Multilingual / Other** | **Whisper HTTP (auto)** | ~2 GB | Automatic language detection across dozens of languages. |
|
|
105
57
|
|
|
106
|
-
|
|
58
|
+
👉 *For model requirements and server setup, see [Choosing a Speech Engine](docs/CHOOSING-AN-ENGINE.md).*
|
|
107
59
|
|
|
108
|
-
|
|
60
|
+
---
|
|
109
61
|
|
|
110
|
-
|
|
62
|
+
## ✨ Features at a Glance
|
|
111
63
|
|
|
112
|
-
- **
|
|
113
|
-
- **
|
|
114
|
-
- **
|
|
115
|
-
- **
|
|
116
|
-
- **
|
|
64
|
+
- 🎙️ **Voice Typing:** Speak directly into the DSH composer with live interim transcription.
|
|
65
|
+
- 👐 **Hands-Free Conversation:** Continuous dialogue that stays active across chat sessions.
|
|
66
|
+
- ❓ **Spoken Structured Questions:** Narrates DSH prompt questions and submits your spoken answer.
|
|
67
|
+
- ⌨️ **Hold-to-Talk (Push-to-Talk):** Hold `Control` anywhere on the page to speak; release to send.
|
|
68
|
+
- 🎧 **Acoustic Mode Isolation:** Gated listening for speakers (no echo) and open-mic interruption for headphones.
|
|
69
|
+
- 🗣️ **Spoken Commands:** Control the chat using phrases like *"send"*, *"mute"*, *"clear"*, and *"stop speaking"*.
|
|
70
|
+
- 🧹 **Smart Code Filtering:** Automatically skips or summarizes large code blocks instead of reading syntax out loud.
|
|
71
|
+
- 🏠 **Local-First & Private:** Audio runs locally on your machine (via Browser APIs, Apple MLX, or whisper.cpp); no external voice telemetry.
|
|
72
|
+
- 🌐 **Remote-Ready Host Audio:** Qwen and macOS Say synthesize on the DSH host, then DSH delivers compact audio to your browser—so playback works over remote and LAN connections.
|
|
117
73
|
|
|
118
|
-
|
|
74
|
+
### 🧑💻 Coming Soon: Meeting Mode
|
|
119
75
|
|
|
120
|
-
|
|
121
|
-
- **Speech output:** browser/device audio, native macOS `say`, or host-local Qwen3-TTS with WAV playback in the browser.
|
|
122
|
-
- **Whisper transport:** complete WAV utterances through authenticated same-origin DSH routes.
|
|
123
|
-
- **Privacy:** raw audio and transcripts are not logged by default.
|
|
76
|
+
**Meeting Mode** is a planned differentiator for collaborative coding conversations. It will keep two independent live transcription streams in the DSH composer: your microphone as **“Me:”**, and meeting participants from an explicitly shared screen/system-audio stream as **“Them:”**. This creates an editable, real-time record of a code review or technical discussion, so you can manually ask DSH a question with the meeting context already in the composer.
|
|
124
77
|
|
|
125
|
-
|
|
78
|
+
It will never automatically send the transcript or use meeting audio for voice commands. Sharing system audio will always require explicit browser permission and depends on browser and operating-system support.
|
|
126
79
|
|
|
127
|
-
|
|
80
|
+
---
|
|
128
81
|
|
|
129
|
-
|
|
82
|
+
## 🤝 Acknowledgments & Community
|
|
130
83
|
|
|
131
|
-
|
|
84
|
+
Listening and speaking should work together. A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [dsh-voice](https://github.com/GooDAnDReaDY/dsh-voice) and [Alan2Z](https://github.com/Alan2Z) for [dsh-speak](https://github.com/Alan2Z/dsh-speak), which inspired this unified coordinator. [Read the full story](HISTORY.md).
|
|
132
85
|
|
|
133
|
-
|
|
86
|
+
I use DSH Live Voice for at least 8 hours every day. Feedback, ideas, and contributions are welcome:
|
|
87
|
+
- [Open an Issue](https://github.com/victorwads/dsh-live-voice/issues)
|
|
88
|
+
- [Send a Pull Request](https://github.com/victorwads/dsh-live-voice/pulls)
|
|
134
89
|
|
|
135
|
-
|
|
90
|
+
---
|
|
136
91
|
|
|
137
|
-
|
|
92
|
+
### 📄 License [Apache-2.0](LICENSE)
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# Architecture
|
|
2
|
+
|
|
3
|
+
DSH Live Voice uses application composition roots, domain modules, and a domain-independent design system. There are no legacy client, core, or engine source trees.
|
|
4
|
+
|
|
5
|
+
## Dependency direction
|
|
6
|
+
|
|
7
|
+
1. `src/app` composes DSH services, slots, routes, modules, styles, and client language boundaries.
|
|
8
|
+
2. `src/modules` owns product behavior, feature UI, policies, engines, provider hosts, services, and models.
|
|
9
|
+
3. `src/shared` owns reusable domain-independent React primitives.
|
|
10
|
+
4. App code may import modules and shared code. Modules may import shared code. Shared code never imports modules.
|
|
11
|
+
5. Feature modules use the intentionally centralized client i18n API, but never register DSH slots or host routes.
|
|
12
|
+
6. Browser and host dependency graphs remain separate. Host-only Node imports must never enter the client bundle.
|
|
13
|
+
|
|
14
|
+
## Composition roots
|
|
15
|
+
|
|
16
|
+
- Client: `src/app/client/apply.tsx`.
|
|
17
|
+
- Server: `src/app/server/apply.ts`.
|
|
18
|
+
- Slot definitions and registration: `src/app/client/slotDefinitions.ts` and `registerSlots.tsx`.
|
|
19
|
+
- Styles: `src/styles/index.ts`.
|
|
20
|
+
|
|
21
|
+
## Internationalization
|
|
22
|
+
|
|
23
|
+
The only client language tree is `src/app/client/i18n`. It owns typed catalogs, the React runtime, DSH catalog registration, and synchronization with `ctx.locale`. Every DSH slot component is wrapped by the application language boundary. UI locale, recognition language, synthesis language, and user-defined command phrases are independent values.
|
|
24
|
+
|
|
25
|
+
## Modules
|
|
26
|
+
|
|
27
|
+
- `modules/core`: shared voice-domain coordination, settings normalization, ownership, transcript handling, microphone policy, filters, and shared Qwen configuration.
|
|
28
|
+
- `modules/conversation`: composer/chat UI, conversation components, chat models, and session-facing behavior.
|
|
29
|
+
- `modules/settings`: modular Settings shell, hooks, services, and Conversation/Recognition/Speech sections.
|
|
30
|
+
- `modules/recognition/engines/{browser,qwen,whisper}`: recognition adapters and provider-specific host/UI code.
|
|
31
|
+
- `modules/speak/engines/{browser,qwen,say,audio}`: speech adapters, host audio, transcoding, and provider clients.
|
|
32
|
+
|
|
33
|
+
## Stable contracts
|
|
34
|
+
|
|
35
|
+
- Preserve DSH slot names, IDs, order, and injected services.
|
|
36
|
+
- Preserve authenticated same-origin API routes.
|
|
37
|
+
- Preserve persisted key `dsh-live-voice.settings` and normalization.
|
|
38
|
+
- Preserve pause, resume, cancel, ownership, interruption, and stale-result semantics.
|
|
39
|
+
- Do not log raw audio or transcripts by default.
|
|
40
|
+
- Treat lifecycle behavior in [VOICE-LIFECYCLE.md](VOICE-LIFECYCLE.md) as the natural-language behavioral contract.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# Choosing a Speech Recognition Engine
|
|
2
|
+
|
|
3
|
+
DSH Live Voice brings hands-free voice conversations to DeepSeek Harness (DSH). Because speech recognition (STT) runs local-first, choosing the right engine depends on your **operating system**, **language**, and **available RAM**.
|
|
4
|
+
|
|
5
|
+
This guide summarizes practical recommendations based on real-world daily use, observed resource footprints, and platform compatibility.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Quick Recommendations
|
|
10
|
+
|
|
11
|
+
| Language | Platform | Recommended Engine | RAM Footprint | Notes |
|
|
12
|
+
| --- | --- | --- | --- | --- |
|
|
13
|
+
| **Portuguese (Brazil)** | macOS | **Qwen3 ASR (HTTP API)** | ~3 GB | **Best accuracy.** Maintainer's daily driver. Whisper is second choice. |
|
|
14
|
+
| **English (US)** | macOS | **Browser SpeechRecognition** | ~0 GB | Built-in macOS/browser API. Fast and lightweight. |
|
|
15
|
+
| **English / Portuguese** | Windows | **Qwen3 ASR** or **Whisper HTTP** | ~2–3 GB | Recommended starting point. Windows browser STT is inconsistent. |
|
|
16
|
+
| **Chinese (Mandarin)** | Any | **Qwen3 ASR** or **Whisper HTTP** | ~2–3 GB | Strong multilingual support in underlying models. |
|
|
17
|
+
| **Multilingual / Other** | Any | **Whisper HTTP (auto)** | ~2 GB | Automatic language detection (`auto`). |
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Engine Breakdown
|
|
22
|
+
|
|
23
|
+
### 1. Browser SpeechRecognition (Web Speech API)
|
|
24
|
+
*Default option in DSH Live Voice settings.*
|
|
25
|
+
|
|
26
|
+
- **How it works:** Uses the browser's native `SpeechRecognition` interface. On macOS/Chrome/Safari, it can leverage system dictation and locally downloaded language packs.
|
|
27
|
+
- **RAM usage:** Negligible (~0 GB extra), since it runs within the browser and operating system frameworks.
|
|
28
|
+
- **Strengths:**
|
|
29
|
+
- Zero setup required: no local server or model weights to download.
|
|
30
|
+
- Extremely responsive and lightweight.
|
|
31
|
+
- Outstanding recognition for **English** on macOS.
|
|
32
|
+
- **Trade-offs:**
|
|
33
|
+
- Availability and accuracy depend heavily on your browser vendor and operating system.
|
|
34
|
+
- When *Process recognition locally on this device* is checked, your browser must have the language pack installed (e.g. via browser settings or the auto-install prompt).
|
|
35
|
+
- Inconsistent across different platforms (especially Windows).
|
|
36
|
+
|
|
37
|
+
### 2. Qwen3 ASR (HTTP API)
|
|
38
|
+
*Local-first inference via host-local HTTP server (e.g. Apple MLX or OminiX-API).*
|
|
39
|
+
|
|
40
|
+
- **How it works:** Transcribes bounded audio utterances by posting 16 kHz mono WAV to an OpenAI-compatible `/v1/audio/transcriptions` endpoint running on your machine.
|
|
41
|
+
- **RAM usage:** ~3 GB RAM observed during typical local MLX inference.
|
|
42
|
+
- **Strengths:**
|
|
43
|
+
- **Highest accuracy for Brazilian Portuguese** in real-world testing.
|
|
44
|
+
- Native support for Portuguese, English, and Chinese.
|
|
45
|
+
- Decoupled from browser quirks and system dictation bugs.
|
|
46
|
+
- **Trade-offs:**
|
|
47
|
+
- Requires running a separate local HTTP process (e.g., via Python or MLX).
|
|
48
|
+
- Uses ~3 GB of memory.
|
|
49
|
+
|
|
50
|
+
### 3. Whisper (HTTP API)
|
|
51
|
+
*Local-first inference via whisper.cpp or compatible server.*
|
|
52
|
+
|
|
53
|
+
- **How it works:** Posts complete audio chunks to a host-local loopback server (e.g., `http://127.0.0.1:8080/inference` or standard `/v1/audio/transcriptions`).
|
|
54
|
+
- **RAM usage:** ~2 GB RAM observed in typical setups (e.g., `small` or quantized models).
|
|
55
|
+
- **Strengths:**
|
|
56
|
+
- Strong, reliable multilingual recognition across dozens of languages.
|
|
57
|
+
- Supports **Automatic — detect language** (`auto`), switching seamlessly without manual reconfiguration.
|
|
58
|
+
- Highly optimized C/C++ runtimes via `whisper.cpp`.
|
|
59
|
+
- **Trade-offs:**
|
|
60
|
+
- Requires running a separate local Whisper server.
|
|
61
|
+
- Transcription is chunk-based rather than real-time streaming.
|
|
62
|
+
|
|
63
|
+
---
|
|
64
|
+
|
|
65
|
+
## Platform & OS Guidance
|
|
66
|
+
|
|
67
|
+
### macOS (Apple Silicon / Intel)
|
|
68
|
+
macOS is the primary development and daily-testing platform for DSH Live Voice:
|
|
69
|
+
- **For Brazilian Portuguese:** Run **Qwen3 ASR** on Apple MLX. It delivers the most natural transcription with punctuation and handles conversational Portuguese with high precision.
|
|
70
|
+
- **For English:** Use **Browser SpeechRecognition**. It uses the native macOS speech synthesis and dictation stack, keeping your RAM free for local LLMs or developer tools.
|
|
71
|
+
- **Whisper:** A great fallback if you prefer a single model for multiple languages.
|
|
72
|
+
|
|
73
|
+
### Windows
|
|
74
|
+
- **Browser SpeechRecognition:** Often poorly implemented or inconsistent across Chromium builds on Windows, sometimes requiring external cloud endpoints or failing capability checks.
|
|
75
|
+
- **Recommendation:** Use **Qwen3 ASR** or **Whisper HTTP** running locally (e.g. via WSL, Docker, or native Windows binaries).
|
|
76
|
+
- *Validation notice:* The maintainer primarily develops and tests on macOS. Windows users are encouraged to test and submit feedback.
|
|
77
|
+
|
|
78
|
+
---
|
|
79
|
+
|
|
80
|
+
## Resource & Memory Comparison
|
|
81
|
+
|
|
82
|
+
| Metric | Browser SpeechRecognition | Whisper HTTP | Qwen3 ASR |
|
|
83
|
+
| --- | --- | --- | --- |
|
|
84
|
+
| **Additional Process** | None | Separate HTTP server | Separate HTTP server |
|
|
85
|
+
| **Typical RAM Usage** | Minimal (~0 GB) | ~2 GB | ~3 GB |
|
|
86
|
+
| **Setup Complexity** | Zero configuration | Moderate (run server binary) | Moderate (run MLX/Python server) |
|
|
87
|
+
| **Language Detection** | Specific locale only | Manual or `auto` | Manual locale mapping |
|
|
88
|
+
| **Local Privacy** | Configurable (local pack or cloud) | 100% local loopback | 100% local loopback |
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
## Language Support Details
|
|
93
|
+
|
|
94
|
+
1. **Portuguese (Brazil - `pt-BR`):**
|
|
95
|
+
- First-class citizen in DSH Live Voice.
|
|
96
|
+
- Mapped directly as `portuguese` for Qwen3 and `pt` for Whisper.
|
|
97
|
+
2. **English (United States - `en-US`):**
|
|
98
|
+
- Full native support across all three engines.
|
|
99
|
+
3. **Chinese:**
|
|
100
|
+
- Supported natively by Qwen3 and Whisper backends.
|
|
101
|
+
4. **Other Languages:**
|
|
102
|
+
- Select **Whisper HTTP API** with language set to **Automatic — detect language** to speak in Spanish, French, German, Japanese, and more.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## Speech Output (TTS) Note
|
|
107
|
+
|
|
108
|
+
Speech recognition (hearing you) and speech output (speaking back) are configured independently:
|
|
109
|
+
- **Browser Speech:** Uses local system voices installed in your OS.
|
|
110
|
+
- **macOS `say`:** Native host CLI synthesis. The host renders temporary WAV audio, transcodes it to compact AAC/M4A, removes temporary files, and the browser controls playback.
|
|
111
|
+
- **Qwen3 TTS:** High-quality neural synthesis running on the DSH host via MLX; its internal WAV response receives the same AAC/M4A transport conversion.
|
|
112
|
+
|
|
113
|
+
For both host engines, Live Voice preserves segment order and prepares at most three upcoming segments to reduce gaps. Stop, engine changes, and session teardown cancel or discard obsolete preparation. AAC in an M4A container is the shared browser transport; macOS `afconvert` provides the local conversion without a package dependency.
|
|
114
|
+
|
|
115
|
+
---
|
|
116
|
+
|
|
117
|
+
## Help Us Improve
|
|
118
|
+
|
|
119
|
+
Have you tested DSH Live Voice on Windows, Linux, or in other languages?
|
|
120
|
+
Please [open an issue](https://github.com/victorwads/dsh-live-voice/issues) or submit a pull request with your system specifications, engine configuration, and experience!
|
|
@@ -0,0 +1,149 @@
|
|
|
1
|
+
# Configuring DSH Live Voice & Conversation Flow
|
|
2
|
+
|
|
3
|
+
DSH Live Voice integrates directly into DeepSeek Harness settings under **DSH Settings → Live Voice**.
|
|
4
|
+
|
|
5
|
+
This guide covers all conversation modes, acoustic settings, silence thresholds, audio device routing, voice commands, and external engine configurations.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Acoustic Modes & Turn-Taking
|
|
10
|
+
|
|
11
|
+
How listening and speaking coordinate depends on whether you use speakers or headphones:
|
|
12
|
+
|
|
13
|
+
### Speakers — Gated Listening (Default)
|
|
14
|
+
When using speakers, microphone input is automatically gated while the assistant is speaking so the assistant never hears its own voice or transcribes its own output.
|
|
15
|
+
|
|
16
|
+
```mermaid
|
|
17
|
+
sequenceDiagram
|
|
18
|
+
participant Interface
|
|
19
|
+
actor You as User
|
|
20
|
+
actor Assistant
|
|
21
|
+
|
|
22
|
+
Note over Interface,Assistant: Waiting for you to speak…
|
|
23
|
+
activate You
|
|
24
|
+
You->>Assistant: Starts speaking
|
|
25
|
+
Assistant-->>You: Listens to speech
|
|
26
|
+
Note over Interface,Assistant: You pause / stop speaking
|
|
27
|
+
opt Manual Send
|
|
28
|
+
You->>Interface: Review transcribed draft
|
|
29
|
+
Interface-->>You: Send when ready
|
|
30
|
+
end
|
|
31
|
+
opt Automatic Send
|
|
32
|
+
Note over You,Assistant: Sends after silence countdown
|
|
33
|
+
end
|
|
34
|
+
deactivate You
|
|
35
|
+
You->>Assistant: Message delivered
|
|
36
|
+
Assistant-->>You: Response ready — waits for stable silence
|
|
37
|
+
Assistant-->>You: Pauses listening
|
|
38
|
+
activate Assistant
|
|
39
|
+
Assistant->>You: Speaks response aloud
|
|
40
|
+
deactivate Assistant
|
|
41
|
+
Note over Interface,Assistant: Listening resumed — ready for next turn
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
- **Interrupting on speakers:** Click **Take microphone** or use the push-to-talk shortcut to take your turn immediately and pause the assistant.
|
|
45
|
+
|
|
46
|
+
### Headphones — Open Microphone
|
|
47
|
+
With headphones, the microphone remains open during playback:
|
|
48
|
+
- **Speech-detected pause:** When you begin speaking, the assistant automatically pauses.
|
|
49
|
+
- **Debounced interruption:** Requires at least one recognized word to prevent accidental pauses from coughs or background noise.
|
|
50
|
+
- **Explicit resumption:** Silence alone does not automatically resume obsolete playback.
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## 2. Conversation Settings
|
|
55
|
+
|
|
56
|
+
| Setting | Default | Description |
|
|
57
|
+
| --- | --- | --- |
|
|
58
|
+
| **Sending Mode** | `Manual` | `Manual` (review before sending), `Automatic` (sends after silence countdown), or `Steer` (send directly to running agent). |
|
|
59
|
+
| **Auto-Send Delay** | `4 seconds` | Silence countdown before automatically submitting your draft (configurable from 2 to 10 seconds). |
|
|
60
|
+
| **Assistant Response Delay** | `3 seconds` | Continuous silence required after you stop speaking before the assistant begins automatic audio playback. |
|
|
61
|
+
| **Automatic Assistant Speech** | `Enabled` | Automatically reads new assistant responses aloud as they stream in. |
|
|
62
|
+
| **Stop Speech on Send** | `Disabled` | Immediately stops current assistant playback whenever you submit a new message. |
|
|
63
|
+
| **Hold-to-Talk (Push-to-Talk)** | `Enabled` | Hold `Control` anywhere on the page to speak; release to transcribe and auto-send, or press `Escape` to cancel. |
|
|
64
|
+
| **Hands-Free Structured Questions** | Automatic | Narrates DSH structured questions, captures your spoken response, and submits it as a custom answer. |
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
## 3. Speech Recognition (Input)
|
|
69
|
+
|
|
70
|
+
Configure your microphone and speech-to-text (STT) engine:
|
|
71
|
+
|
|
72
|
+
- **Recognition Engine:**
|
|
73
|
+
- **Browser SpeechRecognition:** Uses the browser/system dictation engine.
|
|
74
|
+
- **Qwen3 ASR — HTTP API:** Connects to a host-local Qwen3 model server.
|
|
75
|
+
- **Whisper — HTTP API:** Connects to a host-local whisper.cpp server.
|
|
76
|
+
- **Input Device:** Explicitly select a microphone (USB mic, headset, or built-in).
|
|
77
|
+
- **Recognition Language:** Choose **Português (Brasil)**, **English (United States)**, or **Automatic — detect language** (available on Qwen3 and Whisper).
|
|
78
|
+
- **Local Processing:**
|
|
79
|
+
- *Process recognition locally on this device:* Ensures speech is not sent to external browser vendor cloud servers.
|
|
80
|
+
- *Automatically install language pack:* Allows the browser to download offline language packs on demand.
|
|
81
|
+
- **Silence Detection (VAD) Profiles:**
|
|
82
|
+
- *Short (900 ms):* Fast turnaround for concise commands.
|
|
83
|
+
- *Natural (1500 ms - default):* Balanced for conversational cadence.
|
|
84
|
+
- *Long (2200 ms):* Accommodates thinking pauses during complex prompts.
|
|
85
|
+
- **Maximum Utterance Limit:** Split continuous speech after 10 to 300 seconds (default: 60s) to keep transcription chunks manageable.
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## 4. Speech Output (Speaking)
|
|
90
|
+
|
|
91
|
+
Configure how assistant messages are read aloud:
|
|
92
|
+
|
|
93
|
+
- **Speech Engine:**
|
|
94
|
+
- **Browser speech:** Plays audio through the browser device using local system voices.
|
|
95
|
+
- **macOS `say`:** Host-local synthesis to temporary WAV, converted to compact AAC/M4A for browser playback; temporary files are removed after conversion.
|
|
96
|
+
- **Qwen3 TTS:** Neural voice synthesis with WAV streaming playback in the browser.
|
|
97
|
+
- **Output Device:** Route playback to specific speakers or headphones.
|
|
98
|
+
- **Speech Rate:** Adjust reading speed from 0.1x to 3.0x (1.0 is normal).
|
|
99
|
+
- **Code Block Filtering:**
|
|
100
|
+
- Automatically filter out lengthy Markdown code blocks so the assistant doesn't recite syntax line by line.
|
|
101
|
+
- Read code blocks up to a configured line count (e.g. 5 lines).
|
|
102
|
+
- Replace larger code blocks with a custom spoken notice (e.g. *"Look at the code in our conversation"*).
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## 5. Spoken Voice Commands
|
|
107
|
+
|
|
108
|
+
When voice commands are enabled, you can speak exact trigger phrases to control the conversation hands-free:
|
|
109
|
+
|
|
110
|
+
| Action | Default Trigger Phrases |
|
|
111
|
+
| --- | --- |
|
|
112
|
+
| **Send Message** | `send`, `send message` |
|
|
113
|
+
| **Queue Message** | `queue`, `queue message` |
|
|
114
|
+
| **End Conversation** | `end`, `end conversation` |
|
|
115
|
+
| **Mute Microphone** | `mute`, `stop listening` |
|
|
116
|
+
| **Resume Microphone** | `resume`, `start listening` |
|
|
117
|
+
| **Stop Speech** | `stop talking`, `stop speaking`, `shut up` |
|
|
118
|
+
| **Clear Input** | `clear all`, `clear message` |
|
|
119
|
+
|
|
120
|
+
*Phrases are configurable in Settings → Live Voice → Commands, supporting comma-separated alternatives.*
|
|
121
|
+
|
|
122
|
+
---
|
|
123
|
+
|
|
124
|
+
## 6. External Host Engine Setup
|
|
125
|
+
|
|
126
|
+
### Qwen3 HTTP Service
|
|
127
|
+
When using Qwen3 ASR or TTS:
|
|
128
|
+
- Runs as an external local process (e.g. via Apple MLX, OminiX-API, or Python).
|
|
129
|
+
- Configured via base URL (default: `http://127.0.0.1:8080/`).
|
|
130
|
+
- Exposes standard OpenAI-compatible endpoints:
|
|
131
|
+
- `GET /health`
|
|
132
|
+
- `POST /v1/audio/transcriptions`
|
|
133
|
+
- `POST /v1/audio/speech`
|
|
134
|
+
- Credentials and host settings are stored with owner-only permissions in `~/.dsh/dsh-live-voice-qwen.json`.
|
|
135
|
+
|
|
136
|
+
### Whisper HTTP Service
|
|
137
|
+
When using Whisper:
|
|
138
|
+
- Connects to `whisper.cpp` server or compatible loopback HTTP daemon.
|
|
139
|
+
- Configured via inference endpoint URL (default: `http://127.0.0.1:8080/inference`).
|
|
140
|
+
- Audio is validated as bounded 16 kHz mono PCM16 WAV and passed strictly through authenticated same-origin DSH routes.
|
|
141
|
+
- Configuration is stored with owner-only permissions in `~/.dsh/dsh-live-voice-whisper.json`.
|
|
142
|
+
|
|
143
|
+
---
|
|
144
|
+
|
|
145
|
+
## 7. Privacy & Local Architecture
|
|
146
|
+
|
|
147
|
+
- **No Remote Telemetry:** Transcripts and raw microphone audio are never logged or phoned home.
|
|
148
|
+
- **Strict Loopback:** External engine bridges only accept unauthenticated loopback addresses (`127.0.0.1`, `::1`, `localhost`), preventing external network leakage.
|
|
149
|
+
- **Browser vs. Host:** Speech processing runs on your own hardware. Keep in mind that the DeepSeek Harness language model itself may be remote depending on your DSH setup.
|
package/docs/REVIEW.md
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Architecture Review Guide
|
|
2
|
+
|
|
3
|
+
## Authoritative boundaries
|
|
4
|
+
|
|
5
|
+
- Client root: `src/app/client/apply.tsx`.
|
|
6
|
+
- Server root: `src/app/server/apply.ts`.
|
|
7
|
+
- Client i18n: `src/app/client/i18n`.
|
|
8
|
+
- Conversation UI: `src/modules/conversation`.
|
|
9
|
+
- Settings runtime: `src/modules/settings/components/LiveVoiceSettings.tsx`.
|
|
10
|
+
- Recognition engines: `src/modules/recognition/engines`.
|
|
11
|
+
- Speech engines: `src/modules/speak/engines`.
|
|
12
|
+
- Shared UI primitives: `src/shared/design-system`.
|
|
13
|
+
|
|
14
|
+
No compatibility implementation remains under `src/client`, `src/core`, or `src/engines`; those trees must not be recreated. Barrels expose implementations physically owned by their modules rather than forwarding to legacy trees.
|
|
15
|
+
|
|
16
|
+
## Review rules
|
|
17
|
+
|
|
18
|
+
1. App composes modules; modules may import shared code; shared code never imports modules.
|
|
19
|
+
2. The centralized client i18n API is the sole intentional module-to-app dependency.
|
|
20
|
+
3. Every registered slot is wrapped by the DSH language boundary.
|
|
21
|
+
4. Provider adapters remain separate from conversation policy.
|
|
22
|
+
5. Browser bundles contain no Node built-ins.
|
|
23
|
+
6. Storage keys, routes, slot IDs, and established CSS classes remain stable.
|
|
24
|
+
|
|
25
|
+
## High-risk manual checks
|
|
26
|
+
|
|
27
|
+
- Start, stop, cancel, mute, and resume input.
|
|
28
|
+
- Start a conversation, switch chats, and verify transient mute does not leak into a new session.
|
|
29
|
+
- Test Speakers and Headphones turn-taking.
|
|
30
|
+
- Test pause, resume, next segment, and stop-all playback.
|
|
31
|
+
- Exercise queue and steer delivery, pending questions, devices, permissions, and provider settings.
|
|
32
|
+
- Change DSH UI language without changing STT/TTS selections.
|
|
33
|
+
- Test authenticated routes with real providers when available.
|
|
34
|
+
|
|
35
|
+
The automated suite covers cancellation, stale results, engines, routes, normalization, locale completeness, lifecycle, slots, modular UI, and design-system contracts. Physical devices and the authenticated GUI still require manual verification.
|