dsh-live-voice 0.1.0 → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +92 -0
- package/README.md +19 -4
- package/lib/client.js +790 -215
- package/lib/server.js +16 -0
- package/package.json +19 -11
- package/scripts/build.ts +1 -1
- package/src/client/chat.ts +7 -0
- package/src/client/components.ts +453 -190
- package/src/client/index.ts +127 -19
- package/src/client/styles.ts +11 -7
- package/src/core/coordinator.ts +205 -26
- package/src/core/filters.ts +63 -0
- package/src/core/microphone.ts +7 -1
- package/src/core/ownership.ts +1 -1
- package/src/core/settings.ts +107 -3
- package/src/engines/speaking/qwen-http.ts +6 -1
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to DSH Live Voice are documented in this file.
|
|
4
|
+
|
|
5
|
+
## [0.2.1] - 2026-09-19
|
|
6
|
+
|
|
7
|
+
### Documentation
|
|
8
|
+
|
|
9
|
+
- Reframe the README around the local-first, hands-free experience, including continuous conversations, spoken structured questions, voice commands, and behavior across chat navigation.
|
|
10
|
+
- Expand package and README discovery keywords for hands-free voice, conversational AI, spoken prompts, Qwen3 speech engines, and Apple Silicon.
|
|
11
|
+
|
|
12
|
+
### Changes
|
|
13
|
+
|
|
14
|
+
- Update existing development dependencies.
|
|
15
|
+
|
|
16
|
+
### Bug Fixes
|
|
17
|
+
|
|
18
|
+
- Persist microphone enabled or muted state as a global preference and propagate settings changes to every active voice controller. New controllers now inherit the saved microphone state, and capability information is refreshed after synchronized settings changes.
|
|
19
|
+
- Keep microphone, conversation, and Speak controls available when DSH omits the legacy `uiSession.pendingInteractions` store. Structured-question voice handling remains inactive when that optional store is unavailable.
|
|
20
|
+
- Add regression coverage for global microphone-state normalization and for mounting voice controls without the legacy pending-interaction store.
|
|
21
|
+
|
|
22
|
+
## [0.2.0] - Unreleased
|
|
23
|
+
|
|
24
|
+
### Release Changes
|
|
25
|
+
|
|
26
|
+
This release expands conversation control, audio routing, speech filtering, and settings organization.
|
|
27
|
+
|
|
28
|
+
### Features
|
|
29
|
+
|
|
30
|
+
- Add hands-free DSH structured questions in voice conversation mode: narrate each prompt, capture the next spoken response as a custom answer, submit it automatically, and keep the Live Voice status bar visible over the question panel.
|
|
31
|
+
- Add audio input and output device preferences.
|
|
32
|
+
- Add a three-state delivery mode for manual review, queued delivery, and immediate steering.
|
|
33
|
+
- Add configurable speech-input filtering, including minimum-word filtering for final recognition results.
|
|
34
|
+
- Add configurable speech-output filtering to omit code blocks from spoken responses.
|
|
35
|
+
- Add exact voice commands for ending a conversation, muting or resuming listening, stopping speech, clearing input, and sending or queuing recognized text. Commands support multiple comma-separated phrases and normalize punctuation, case, and accents only while matching.
|
|
36
|
+
- Add a remaining-speech segment count to the automatic speech control; it remains expanded while speech is queued.
|
|
37
|
+
- Use the same newline-only segmentation and queue for automatic streaming speech and manual Speak actions.
|
|
38
|
+
- Add tabbed Live Voice settings to organize speech, conversation, commands, and advanced options.
|
|
39
|
+
|
|
40
|
+
### Changes
|
|
41
|
+
|
|
42
|
+
- Improve the Live Voice controls and settings interface for delivery, device, filtering, command, microphone-input, and speech-queue states.
|
|
43
|
+
- Expand coordinator, settings, component, and filter test coverage.
|
|
44
|
+
- Update compiled client and server bundles for the new functionality.
|
|
45
|
+
- Include this changelog in the published package.
|
|
46
|
+
|
|
47
|
+
### Bug Fixes
|
|
48
|
+
|
|
49
|
+
- Keep voice conversation mode active across chat navigation. When the current composer is replaced, the newly mounted chat automatically resumes voice mode; only an explicit End voice conversation action disables it. Composer drafts remain isolated per conversation, while internal controller disposal and hardware handoffs no longer count as user-requested conversation termination.
|
|
50
|
+
- Add lifecycle regression coverage for switching chats while voice mode is active and for preserving an explicit end across subsequent chats.
|
|
51
|
+
- Require at least one recognized word before headphone-mode microphone activity may pause assistant speech; audio activity alone no longer pauses playback.
|
|
52
|
+
- Debounce headphone-mode interruptions to reduce false pauses from short recognition events.
|
|
53
|
+
- Automatically resume speech when an interruption candidate ends without becoming valid user speech.
|
|
54
|
+
- Preserve manual pause behavior separately from automatic interruption handling.
|
|
55
|
+
|
|
56
|
+
## [0.1.0] - 2026-09-15
|
|
57
|
+
|
|
58
|
+
### Bug Fixes
|
|
59
|
+
|
|
60
|
+
- Reliably reset browser speech recognition after it ends or encounters an error.
|
|
61
|
+
- Allow a custom HTTP or HTTPS base URL for the host-local Qwen3 speech service.
|
|
62
|
+
|
|
63
|
+
## [0.0.2] - 2026-09-15
|
|
64
|
+
|
|
65
|
+
### Features
|
|
66
|
+
|
|
67
|
+
- Add host-local Apple MLX support for Qwen3-ASR and Qwen3-TTS.
|
|
68
|
+
- Add Qwen TTS voice selection in Live Voice settings.
|
|
69
|
+
- Expand Live Voice controls and their integration coverage.
|
|
70
|
+
- Add continuous integration, a pre-push hook, formatting configuration, and distribution-artifact validation.
|
|
71
|
+
|
|
72
|
+
### Changes
|
|
73
|
+
|
|
74
|
+
- Update the conversation coordinator, microphone, recognition, synthesis, Whisper, and native macOS `say` integrations for the new engines and controls.
|
|
75
|
+
- Include compiled client and server bundles in the published distribution.
|
|
76
|
+
|
|
77
|
+
## [0.0.1-alpha.1] - 2026-09-15
|
|
78
|
+
|
|
79
|
+
### Features
|
|
80
|
+
|
|
81
|
+
- First functional release of the DSH Live Voice plugin.
|
|
82
|
+
- Coordinate microphone capture, speech recognition, assistant-message delivery, and speech playback.
|
|
83
|
+
- Add voice typing in the DSH composer and continuous voice conversations.
|
|
84
|
+
- Support browser SpeechRecognition and authenticated loopback whisper.cpp HTTP recognition.
|
|
85
|
+
- Support browser speech synthesis and native macOS `say` output.
|
|
86
|
+
- Add Live Voice controls, Whisper settings, build and browser-preview scripts, and an initial test suite.
|
|
87
|
+
|
|
88
|
+
[0.2.1]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.0...HEAD
|
|
89
|
+
[0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...v0.2.0
|
|
90
|
+
[0.1.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.2...v0.1.0
|
|
91
|
+
[0.0.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.1-alpha.1...v0.0.2
|
|
92
|
+
[0.0.1-alpha.1]: https://github.com/victorwads/dsh-live-voice/tree/v0.0.1-alpha.1
|
package/README.md
CHANGED
|
@@ -4,9 +4,9 @@
|
|
|
4
4
|
[](https://dsh.pub/en/plugins/dsh-live-voice/)
|
|
5
5
|
[](LICENSE)
|
|
6
6
|
|
|
7
|
-
**Local-first
|
|
7
|
+
**Local-first, hands-free voice conversations for DSH — speak, listen, answer prompts, and keep working without touching the computer.**
|
|
8
8
|
|
|
9
|
-
DSH Live Voice coordinates the microphone, composer, assistant messages, and speech output in one plugin — without letting listening and speaking compete with each other.
|
|
9
|
+
DSH Live Voice coordinates the microphone, composer, assistant messages, structured questions, and speech output in one plugin — without letting listening and speaking compete with each other. Once voice conversation mode is running, DSH can narrate responses and questions, capture your spoken answers, and continue the conversation while your hands stay free.
|
|
10
10
|
|
|
11
11
|
## Install
|
|
12
12
|
|
|
@@ -29,7 +29,10 @@ A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [ds
|
|
|
29
29
|
| | Capability |
|
|
30
30
|
|---|---|
|
|
31
31
|
| 🎙️ | Voice typing directly into the DSH composer |
|
|
32
|
-
|
|
|
32
|
+
| 👐 | Hands-free conversations: speak, hear responses, and continue without touching the computer |
|
|
33
|
+
| ❓ | Spoken DSH structured questions with automatic capture and submission of your answer |
|
|
34
|
+
| 💬 | Continuous voice conversations that stay active while you navigate between chats |
|
|
35
|
+
| 🗣️ | Configurable voice commands for sending, queueing, clearing, muting, resuming, stopping speech, and ending a conversation |
|
|
33
36
|
| 🧠 | Browser SpeechRecognition, local loopback whisper.cpp, or Qwen3-ASR on Apple MLX |
|
|
34
37
|
| 🔊 | Browser speech synthesis, native macOS `say`, or Qwen3-TTS on Apple MLX |
|
|
35
38
|
| ⏱️ | Manual or automatic sending after configurable silence |
|
|
@@ -40,6 +43,18 @@ A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [ds
|
|
|
40
43
|
| 🔈 | Play individual assistant messages on demand |
|
|
41
44
|
| 🏠 | Whisper audio reaches the local server only through the authenticated DSH host |
|
|
42
45
|
|
|
46
|
+
## Hands-free experience
|
|
47
|
+
|
|
48
|
+
Start voice conversation mode and choose an automatic delivery mode to keep a conversation moving without returning to the keyboard. DSH Live Voice can:
|
|
49
|
+
|
|
50
|
+
1. Listen for your next message and deliver it after the configured silence period.
|
|
51
|
+
2. Read assistant responses aloud as they arrive.
|
|
52
|
+
3. Narrate DSH structured questions, listen for your next spoken response, and submit it as a custom answer.
|
|
53
|
+
4. Keep voice conversation mode active when you move between chats until you explicitly end it.
|
|
54
|
+
5. Accept configurable spoken commands for common conversation controls.
|
|
55
|
+
|
|
56
|
+
The hands-free experience coordinates speech input and output locally when you select local engines. The DSH language model itself may still be remote.
|
|
57
|
+
|
|
43
58
|
## Conversation flow
|
|
44
59
|
|
|
45
60
|
The assistant never starts automatic playback while you are speaking. If a response is already waiting — including another assistant message — it waits until you finish and the configured continuous-silence delay has passed.
|
|
@@ -108,4 +123,4 @@ When a compatible Qwen3 speech API is already running on the DSH host, choose **
|
|
|
108
123
|
|
|
109
124
|
## Keywords
|
|
110
125
|
|
|
111
|
-
`dsh`, `dsh-plugin`, `deepseek-harness`, `local-first`, `local-voice`, `voice-assistant`, `voice-conversation`, `continuous-conversation`, `voice-dictation`, `speech-to-text`, `text-to-speech`, `speech-recognition`, `speech-synthesis`, `stt`, `tts`, `whisper`, `whisper-cpp`, `web-speech-api`, `macos-say`, `turn-taking`, `voice-interruption`, `silence-detection`
|
|
126
|
+
`dsh`, `dsh-plugin`, `deepseek-harness`, `local-first`, `local-voice`, `hands-free`, `hands-free-ai`, `hands-free-assistant`, `voice-control`, `voice-commands`, `voice-assistant`, `voice-conversation`, `conversational-ai`, `continuous-conversation`, `voice-dictation`, `spoken-prompts`, `speech-to-text`, `text-to-speech`, `speech-recognition`, `speech-synthesis`, `stt`, `tts`, `whisper`, `whisper-cpp`, `qwen3-asr`, `qwen3-tts`, `mlx`, `apple-silicon`, `web-speech-api`, `macos-say`, `turn-taking`, `voice-interruption`, `silence-detection`
|