dsh-live-voice 0.2.0 → 0.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,44 @@
2
2
 
3
3
  All notable changes to DSH Live Voice are documented in this file.
4
4
 
5
+ ## [0.2.2] - 2026-09-19
6
+
7
+ ### Features
8
+
9
+ - Add an optional global hold-Control push-to-talk gesture that works while a composer is mounted even when regular voice mode is off. Releasing Control flushes queued transcription, waits the configured automatic-send delay, queues the completed draft once, and closes capture; Escape cancels the gesture.
10
+
11
+ ### Changes
12
+
13
+ - Increase the maximum continuous-speech transcription chunk from 20 to 60 seconds by default for Whisper HTTP and Qwen HTTP recognition. Add a Speech recognition setting that lets users configure the limit from 10 to 300 seconds; uninterrupted speech is split and sent for transcription only after the selected duration.
14
+
15
+ ### Bug Fixes
16
+
17
+ - Serialize Whisper HTTP and Qwen HTTP transcription segments through a per-session FIFO queue, preserving capture order while the microphone continues recording.
18
+ - Defer the automatic-send countdown while transcription segments are queued or a request is active.
19
+ - Discard queued segments and abort the active request when recognition stops, and clear the coordinator’s pending-transcription count.
20
+ - Add regression tests for serialized requests, ordered results, cancellation of queued work, and automatic-send gating until the transcription queue drains.
21
+
22
+ ### Compatibility
23
+
24
+ - Record that the latest DSH version tested with this plugin is **0.1.6-alpha.2**.
25
+
26
+ ## [0.2.1] - 2026-09-19
27
+
28
+ ### Documentation
29
+
30
+ - Reframe the README around the local-first, hands-free experience, including continuous conversations, spoken structured questions, voice commands, and behavior across chat navigation.
31
+ - Expand package and README discovery keywords for hands-free voice, conversational AI, spoken prompts, Qwen3 speech engines, and Apple Silicon.
32
+
33
+ ### Changes
34
+
35
+ - Update existing development dependencies.
36
+
37
+ ### Bug Fixes
38
+
39
+ - Persist microphone enabled or muted state as a global preference and propagate settings changes to every active voice controller. New controllers now inherit the saved microphone state, and capability information is refreshed after synchronized settings changes.
40
+ - Keep microphone, conversation, and Speak controls available when DSH omits the legacy `uiSession.pendingInteractions` store. Structured-question voice handling remains inactive when that optional store is unavailable.
41
+ - Add regression coverage for global microphone-state normalization and for mounting voice controls without the legacy pending-interaction store.
42
+
5
43
  ## [0.2.0] - Unreleased
6
44
 
7
45
  ### Release Changes
@@ -68,7 +106,9 @@ This release expands conversation control, audio routing, speech filtering, and
68
106
  - Support browser speech synthesis and native macOS `say` output.
69
107
  - Add Live Voice controls, Whisper settings, build and browser-preview scripts, and an initial test suite.
70
108
 
71
- [0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...HEAD
109
+ [0.2.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.1...HEAD
110
+ [0.2.1]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.0...v0.2.1
111
+ [0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...v0.2.0
72
112
  [0.1.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.2...v0.1.0
73
113
  [0.0.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.1-alpha.1...v0.0.2
74
114
  [0.0.1-alpha.1]: https://github.com/victorwads/dsh-live-voice/tree/v0.0.1-alpha.1
package/README.md CHANGED
@@ -4,9 +4,19 @@
4
4
  [![dsh.pub registry status](https://dsh.pub/api/badges/victorwads/dsh-live-voice.svg)](https://dsh.pub/en/plugins/dsh-live-voice/)
5
5
  [![license](https://img.shields.io/badge/license-GPL--3.0--only-blue)](LICENSE)
6
6
 
7
- **Local-first speech recognition, voice output, and continuous voice conversations for DSH.**
7
+ **A local-first, hands-free voice assistant plugin for DeepSeek Harness (DSH). Speak, listen, answer prompts, and keep working without touching the computer.**
8
8
 
9
- DSH Live Voice coordinates the microphone, composer, assistant messages, and speech output in one plugin — without letting listening and speaking compete with each other.
9
+ Looking for a DeepSeek Harness voice plugin, DSH microphone plugin, speech-to-text, text-to-speech, or hands-free AI assistant? DSH Live Voice brings those capabilities together in one coordinated plugin.
10
+
11
+ It coordinates the microphone, composer, assistant messages, structured questions, and speech output without letting listening and speaking compete. Once voice conversation mode is running, DSH can narrate responses and questions, capture your spoken answers, and continue the conversation while your hands stay free.
12
+
13
+ [![DSH tested](https://img.shields.io/badge/DSH%20tested-0.1.6--alpha.2-5c5cff?logo=deepseek&logoColor=white)](https://github.com/deepseek-ai/deepseek-harness/releases/tag/v0.1.6-alpha.2)
14
+
15
+ ## Built for daily use, maintained with you
16
+
17
+ I use DSH Live Voice for at least eight hours a day. For me, it is the best voice plugin for DSH — and I am committed to making it better through real, everyday use.
18
+
19
+ Have a bug to report, a feature you need, or a pull request to share? [Open an issue](https://github.com/victorwads/dsh-live-voice/issues) or [send a PR](https://github.com/victorwads/dsh-live-voice/pulls). I aim to respond quickly, review contributions promptly, and turn useful suggestions into features. Not every request will be implemented, but I welcome the conversation and will consider what you bring.
10
20
 
11
21
  ## Install
12
22
 
@@ -26,19 +36,36 @@ A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [ds
26
36
 
27
37
  ## Features
28
38
 
29
- | | Capability |
30
- |---|---|
31
- | 🎙️ | Voice typing directly into the DSH composer |
32
- | 💬 | Continuous voice conversations with automatic assistant speech |
33
- | 🧠 | Browser SpeechRecognition, local loopback whisper.cpp, or Qwen3-ASR on Apple MLX |
34
- | 🔊 | Browser speech synthesis, native macOS `say`, or Qwen3-TTS on Apple MLX |
35
- | ⏱️ | Manual or automatic sending after configurable silence |
36
- | 🫁 | Stable-silence delay prevents breathing pauses from starting assistant speech |
37
- | 🎧 | Open-microphone mode for headphones |
38
- | 🔒 | Gated microphone mode for speakers |
39
- | ✋ | Pause, resume, stop, and manual interruption controls |
40
- | 🔈 | Play individual assistant messages on demand |
41
- | 🏠 | Whisper audio reaches the local server only through the authenticated DSH host |
39
+ | | Capability |
40
+ | --- | ------------------------------------------------------------------------------------------------------------------------- |
41
+ | 🎙️ | Voice typing directly into the DSH composer |
42
+ | 👐 | Hands-free conversations: speak, hear responses, and continue without touching the computer |
43
+ | ❓ | Spoken DSH structured questions with automatic capture and submission of your answer |
44
+ | 💬 | Continuous voice conversations that stay active while you navigate between chats |
45
+ | 🗣️ | Configurable voice commands for sending, queueing, clearing, muting, resuming, stopping speech, and ending a conversation |
46
+ | ⌨️ | Optional hold-Control push-to-talk anywhere on the page while a composer is open |
47
+ | 🧠 | Browser SpeechRecognition, local loopback whisper.cpp, or Qwen3-ASR on Apple MLX |
48
+ | 🔊 | Browser speech synthesis, native macOS `say`, or Qwen3-TTS on Apple MLX |
49
+ | ⏱️ | Manual or automatic sending after configurable silence |
50
+ | 🫁 | Stable-silence delay prevents breathing pauses from starting assistant speech |
51
+ | 🎧 | Open-microphone mode for headphones |
52
+ | 🔒 | Gated microphone mode for speakers |
53
+ | ✋ | Pause, resume, stop, and manual interruption controls |
54
+ | 🔈 | Play individual assistant messages on demand |
55
+ | 🏠 | Whisper audio reaches the local server only through the authenticated DSH host |
56
+
57
+ ## Hands-free experience
58
+
59
+ Start voice conversation mode and choose an automatic delivery mode to keep a conversation moving without returning to the keyboard. DSH Live Voice can:
60
+
61
+ 1. Listen for your next message and deliver it after the configured silence period.
62
+ 2. Read assistant responses aloud as they arrive.
63
+ 3. Narrate DSH structured questions, listen for your next spoken response, and submit it as a custom answer.
64
+ 4. Keep voice conversation mode active when you move between chats until you explicitly end it.
65
+ 5. Accept configurable spoken commands for common conversation controls.
66
+ 6. Start temporary push-to-talk from anywhere on the page by holding Control, even when the voice bar is off; release to finish queued transcription and automatic delivery, or press Escape to cancel.
67
+
68
+ The hands-free experience coordinates speech input and output locally when you select local engines. The DSH language model itself may still be remote.
42
69
 
43
70
  ## Conversation flow
44
71
 
@@ -80,7 +107,6 @@ sequenceDiagram
80
107
 
81
108
  The microphone remains open during playback, allowing your voice to pause the assistant.
82
109
 
83
-
84
110
  ### Other conversation settings
85
111
 
86
112
  - **Sending mode:** review and send manually by default, or send automatically after a configurable silence countdown.
@@ -108,4 +134,4 @@ When a compatible Qwen3 speech API is already running on the DSH host, choose **
108
134
 
109
135
  ## Keywords
110
136
 
111
- `dsh`, `dsh-plugin`, `deepseek-harness`, `local-first`, `local-voice`, `voice-assistant`, `voice-conversation`, `continuous-conversation`, `voice-dictation`, `speech-to-text`, `text-to-speech`, `speech-recognition`, `speech-synthesis`, `stt`, `tts`, `whisper`, `whisper-cpp`, `web-speech-api`, `macos-say`, `turn-taking`, `voice-interruption`, `silence-detection`
137
+ `dsh`, `dsh-plugin`, `deepseek-harness`, `local-first`, `local-voice`, `hands-free`, `hands-free-ai`, `hands-free-assistant`, `voice-control`, `voice-commands`, `voice-assistant`, `voice-conversation`, `conversational-ai`, `continuous-conversation`, `voice-dictation`, `spoken-prompts`, `speech-to-text`, `text-to-speech`, `speech-recognition`, `speech-synthesis`, `stt`, `tts`, `whisper`, `whisper-cpp`, `qwen3-asr`, `qwen3-tts`, `mlx`, `apple-silicon`, `web-speech-api`, `macos-say`, `turn-taking`, `voice-interruption`, `silence-detection`