dsh-live-voice 0.3.2 β 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +15 -8
- package/lib/client.js +5959 -3222
- package/lib/server.js +29 -6
- package/package.json +9 -30
- package/src/app/AGENTS.md +1 -1
- package/src/app/ARCHITECTURE.md +1 -1
- package/src/app/client/apply.tsx +140 -16
- package/src/app/client/composerSelection.ts +92 -0
- package/src/app/client/i18n/catalogs/base.ts +33 -0
- package/src/app/client/i18n/catalogs/en.ts +36 -0
- package/src/app/client/i18n/catalogs/es.ts +36 -0
- package/src/app/client/i18n/catalogs/fr.ts +37 -0
- package/src/app/client/i18n/catalogs/hi.ts +35 -0
- package/src/app/client/i18n/catalogs/pt-BR.ts +37 -0
- package/src/app/client/i18n/catalogs/zh.ts +35 -0
- package/src/modules/AGENTS.md +1 -1
- package/src/modules/conversation/ARCHITECTURE.md +4 -2
- package/src/modules/conversation/components/AutoPlaybackToggle.tsx +31 -0
- package/src/modules/conversation/components/ConversationControls.tsx +23 -8
- package/src/modules/conversation/components/ConversationStatusBar.tsx +72 -94
- package/src/modules/conversation/components/DeliveryModeButton.tsx +3 -3
- package/src/modules/conversation/components/MeetingControls.tsx +64 -0
- package/src/modules/conversation/components/MicrophoneButton.tsx +6 -1
- package/src/modules/conversation/components/PlaybackControls.tsx +20 -5
- package/src/modules/conversation/components/RecognitionBar.tsx +14 -0
- package/src/modules/conversation/components/ScrollingSpeechCaption.tsx +100 -0
- package/src/modules/conversation/components/SpeechStatusBar.tsx +99 -0
- package/src/modules/conversation/components/Waveform.tsx +8 -3
- package/src/modules/conversation/components/conversationStatus.ts +0 -2
- package/src/modules/conversation/components/index.ts +1 -0
- package/src/modules/conversation/components/speechCaption.ts +27 -0
- package/src/modules/conversation/models/chat.ts +14 -2
- package/src/modules/conversation/models/meeting.ts +140 -0
- package/src/modules/conversation/models/meetingTranscript.ts +56 -0
- package/src/modules/core/ARCHITECTURE.md +1 -1
- package/src/modules/core/coordinator.ts +204 -51
- package/src/modules/core/developerExtension.ts +26 -0
- package/src/modules/core/diagnostics.ts +150 -0
- package/src/modules/core/filters.ts +20 -15
- package/src/modules/core/markdownSpeech.ts +154 -0
- package/src/modules/core/settings.ts +17 -13
- package/src/modules/core/sharedAudio.ts +82 -0
- package/src/modules/core/speechDefaults.ts +21 -0
- package/src/modules/recognition/engines/browser/BrowserRecognitionEngine.ts +8 -1
- package/src/modules/recognition/engines/whisper/WhisperRecognitionEngine.ts +10 -3
- package/src/modules/settings/components/LiveVoiceSettings.tsx +17 -3
- package/src/modules/settings/sections/conversation/ConversationDelaySettings.tsx +8 -3
- package/src/modules/settings/sections/conversation/DeliverySettings.tsx +25 -2
- package/src/modules/speak/engines/audio/HostAudioEngine.ts +19 -2
- package/src/modules/speak/engines/browser/BrowserSpeakingEngine.ts +33 -2
- package/src/shared/design-system/buttons/ToggleButton.tsx +8 -1
- package/src/shared/design-system/icons/icons.ts +4 -0
- package/src/styles/index.ts +11 -3
- package/CHANGELOG.md +0 -240
- package/DEVELOPMENT.md +0 -83
- package/PLAN.md +0 -296
- package/docs/ARCHITECTURE.md +0 -33
- package/docs/CHOOSING-AN-ENGINE.md +0 -120
- package/docs/CONFIGURATION.md +0 -163
- package/docs/REVIEW.md +0 -35
- package/docs/VOICE-LIFECYCLE.md +0 -63
- package/scripts/build.ts +0 -62
- package/scripts/check-dist.mjs +0 -38
- package/scripts/preview-ui.ts +0 -31
- package/scripts/probe-browser.ts +0 -46
package/README.md
CHANGED
|
@@ -1,10 +1,12 @@
|
|
|
1
1
|
# DSH Live Voice
|
|
2
2
|
|
|
3
|
+
Release history: [repository changelog](https://github.com/victorwads/dsh-live-voice/blob/main/CHANGELOG.md).
|
|
4
|
+
|
|
3
5
|
[](https://www.npmjs.com/package/dsh-live-voice)
|
|
4
|
-
[](https://github.com/deepseek-ai/deepseek-harness/releases/tag/dsh-v0.2.1-alpha.1)
|
|
5
7
|
[](LICENSE)
|
|
6
8
|
|
|
7
|
-
**Tested and working with DeepSeek Harness v0.2.
|
|
9
|
+
**Tested and working with DeepSeek Harness v0.2.1-alpha.1.** This is the current tested DSH version, recorded in the `dshTestedVersion` field in [package.json](package.json).
|
|
8
10
|
|
|
9
11
|
**A local-first, hands-free voice assistant plugin for DeepSeek Harness (DSH).**
|
|
10
12
|
*Built in Brazil π§π· and tested daily with Brazilian Portuguese on macOS.*
|
|
@@ -52,7 +54,7 @@ Detailed guides for deep-diving into engines and configurations:
|
|
|
52
54
|
|
|
53
55
|
| Scenario | Recommendation | RAM | Why |
|
|
54
56
|
| --- | --- | --- | --- |
|
|
55
|
-
| π§π· **Portuguese on macOS** | **Qwen3 ASR (HTTP API)** | ~
|
|
57
|
+
| π§π· **Portuguese on macOS** | **Qwen3 ASR (HTTP API)** | ~1.5 GB | Best accuracy in daily maintainer use; current RAM usage reported by the maintainer. Whisper is second choice. |
|
|
56
58
|
| πΊπΈ **English on macOS** | **Browser SpeechRecognition** | ~0 GB | Built-in macOS/browser API. Fast, zero extra RAM. |
|
|
57
59
|
| πͺ **Windows** | **Qwen3 ASR** or **Whisper HTTP** | ~2β3 GB | Recommended starting point; Windows browser STT varies. |
|
|
58
60
|
| π **Multilingual / Other** | **Whisper HTTP (auto)** | ~2 GB | Automatic language detection across dozens of languages. |
|
|
@@ -66,18 +68,23 @@ Detailed guides for deep-diving into engines and configurations:
|
|
|
66
68
|
- ποΈ **Voice Typing:** Append final recognized speech to the end of the DSH composer without replacing manual edits.
|
|
67
69
|
- π **Hands-Free Conversation:** Continuous dialogue that stays active across chat sessions.
|
|
68
70
|
- β **Spoken Structured Questions:** Narrates DSH prompt questions and submits your spoken answer.
|
|
69
|
-
- β¨οΈ **Hold-to-Talk (Push-to-Talk):**
|
|
71
|
+
- β¨οΈ **Optional Hold-to-Talk (Push-to-Talk):** Enable it in Settings, then hold `Control` anywhere on the page to speak; release to queue the message after the configured send delay. Disabled by default.
|
|
70
72
|
- π§ **Acoustic Mode Isolation:** Gated listening for speakers (no echo) and open-mic interruption for headphones.
|
|
71
73
|
- π£οΈ **Spoken Commands:** Control the chat using phrases like *"send"*, *"mute"*, *"clear"*, and *"stop speaking"*.
|
|
72
|
-
-
|
|
74
|
+
- β **Dedicated Speech Bar:** Smoothly scrolling approximate captions, a moving highlight, previous/next navigation, pause/resume, stop, and a current/total segment counter. Captions show the text sent to the speech engine, without claiming word-level alignment.
|
|
75
|
+
- π§Ή **Markdown-Aware Speech:** Remove formatting, announce links without reading full URLs, shorten file paths and line references, read checkbox states and table rows, and replace long code blocks with a localized notice. Preserve custom notices.
|
|
76
|
+
- β‘ **Responsive Conversation Settings:** Automatic-send delays of 600 ms, 800 ms, or 1β6 seconds (4 seconds by default), plus an assistant response delay of zero to 4 seconds (no delay by default). Recognition and synthesis still contribute to overall latency.
|
|
77
|
+
- π οΈ **Optional Live Voice Debugger:** Install `dsh-live-voice-debugger` separately to inspect runtime state and queues from the Developer tab in Live Voice Settings. The main plugin works without it.
|
|
73
78
|
- π **Local-First & Private:** Audio runs locally on your machine (via Browser APIs, Apple MLX, or whisper.cpp); no external voice telemetry.
|
|
74
79
|
- π **Remote-Ready Host Audio:** Qwen and macOS Say synthesize on the DSH host, then DSH delivers compact audio to your browserβso playback works over remote and LAN connections.
|
|
75
80
|
|
|
76
|
-
### π§βπ»
|
|
81
|
+
### π§βπ» Meeting Mode β Shared Audio Alongside Normal Voice
|
|
82
|
+
|
|
83
|
+
Select **Qwen HTTP** or **Whisper HTTP** recognition, then click the shared-audio icon beside the composer microphone. The first click opens the browser sharing dialog directly; enable audio in that dialog. The shared source gets its own recognition bar, using the same component as normal microphone recognition, without replacing microphone controls or assistant speech playback. Click the shared-audio toggle again to stop only that source. When all three bars are visible, speech/live captions come first, shared audio second, and microphone recognition last, separated by 2px. The microphone bar has one listening/ignoring toggle on the left; queue position appears only in the speech bar, while the automatic-speech toggle remains available in the composer between shared audio and microphone.
|
|
77
84
|
|
|
78
|
-
|
|
85
|
+
Normal microphone behavior, voice commands, delivery settings, and speech output remain available. The composer microphone remains a toggle while capturing, and either source can be stopped independently. The shared-audio bar also has a **Timestamp** clock toggle, off by default: new `Me:`/`Them:` blocks can include `[YYYY/MM/DD HH:MM:SS]` using the local machine time at chunk onset, not recognition completion. Consecutive chunks from the same source retain the existing block, and toggling timestamps never rewrites earlier text. With both sources capturing, source changes introduce **βMe:β** or **βThem:β**; consecutive chunks continue on new lines without repeating labels. Shared audio appends final transcripts but never triggers voice commands or automatic sending itself. Microphone transcripts retain the configured sending behavior.
|
|
79
86
|
|
|
80
|
-
|
|
87
|
+
Browser/OS support and the chosen sharing surface determine whether audio is available. Headphones are recommended to avoid unintended feedback when letting the agent speak into a shared meeting or tab. Automated capture, composer and UI tests do not replace real microphone/shared-audio validation.
|
|
81
88
|
|
|
82
89
|
---
|
|
83
90
|
|