dsh-live-voice 0.2.1 → 0.2.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,46 @@
2
2
 
3
3
  All notable changes to DSH Live Voice are documented in this file.
4
4
 
5
+ ## [0.2.3] - 2026-09-21
6
+
7
+ ### Features
8
+
9
+ - Add an update notification in Live Voice settings when a newer GitHub release is available. The check runs at most once every 24 hours and the action opens the matching release.
10
+ - Show the installed Live Voice version and the tested DSH version in the settings header.
11
+ - Show Browser WebGPU Inference, sherpa-onnx Streaming, NVIDIA Parakeet, and Voxtral Realtime as disabled upcoming recognition engines.
12
+
13
+ ### Changes
14
+
15
+ - Rename recognition-engine options to describe the API they use: Browser SpeechRecognition, Qwen3 ASR HTTP API, and Whisper HTTP API.
16
+ - Show the expected HTTP URL or transcription route for the selected recognition engine.
17
+
18
+ ### Bug Fixes
19
+
20
+ - Avoid repeated GitHub requests when the release API is unavailable by caching failed checks for the same 24-hour interval.
21
+ - Handle release tags with a leading `v` and prerelease identifiers correctly when deciding whether an update is newer.
22
+ - Ignore unexpected release URLs and fall back to the repository releases page.
23
+
24
+ ## [0.2.2] - 2026-09-19
25
+
26
+ ### Features
27
+
28
+ - Add an optional global hold-Control push-to-talk gesture that works while a composer is mounted even when regular voice mode is off. Releasing Control flushes queued transcription, waits the configured automatic-send delay, queues the completed draft once, and closes capture; Escape cancels the gesture.
29
+
30
+ ### Changes
31
+
32
+ - Increase the maximum continuous-speech transcription chunk from 20 to 60 seconds by default for Whisper HTTP and Qwen HTTP recognition. Add a Speech recognition setting that lets users configure the limit from 10 to 300 seconds; uninterrupted speech is split and sent for transcription only after the selected duration.
33
+
34
+ ### Bug Fixes
35
+
36
+ - Serialize Whisper HTTP and Qwen HTTP transcription segments through a per-session FIFO queue, preserving capture order while the microphone continues recording.
37
+ - Defer the automatic-send countdown while transcription segments are queued or a request is active.
38
+ - Discard queued segments and abort the active request when recognition stops, and clear the coordinator’s pending-transcription count.
39
+ - Add regression tests for serialized requests, ordered results, cancellation of queued work, and automatic-send gating until the transcription queue drains.
40
+
41
+ ### Compatibility
42
+
43
+ - Record that the latest DSH version tested with this plugin is **0.1.6-alpha.2**.
44
+
5
45
  ## [0.2.1] - 2026-09-19
6
46
 
7
47
  ### Documentation
@@ -85,7 +125,9 @@ This release expands conversation control, audio routing, speech filtering, and
85
125
  - Support browser speech synthesis and native macOS `say` output.
86
126
  - Add Live Voice controls, Whisper settings, build and browser-preview scripts, and an initial test suite.
87
127
 
88
- [0.2.1]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.0...HEAD
128
+ [0.2.3]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.2...HEAD
129
+ [0.2.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.1...v0.2.2
130
+ [0.2.1]: https://github.com/victorwads/dsh-live-voice/compare/v0.2.0...v0.2.1
89
131
  [0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...v0.2.0
90
132
  [0.1.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.2...v0.1.0
91
133
  [0.0.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.1-alpha.1...v0.0.2
package/README.md CHANGED
@@ -1,12 +1,19 @@
1
1
  # DSH Live Voice
2
2
 
3
3
  [![npm version](https://img.shields.io/npm/v/dsh-live-voice?logo=npm&label=npm&color=brightgreen)](https://www.npmjs.com/package/dsh-live-voice)
4
- [![dsh.pub registry status](https://dsh.pub/api/badges/victorwads/dsh-live-voice.svg)](https://dsh.pub/en/plugins/dsh-live-voice/)
5
- [![license](https://img.shields.io/badge/license-GPL--3.0--only-blue)](LICENSE)
4
+ [![DSH](https://img.shields.io/badge/DSH-v0.1.6--alpha.2-5c5cff?logo=deepseek&logoColor=white)](https://github.com/deepseek-ai/deepseek-harness/releases/tag/v0.1.6-alpha.2)
6
5
 
7
- **Local-first, hands-free voice conversations for DSH — speak, listen, answer prompts, and keep working without touching the computer.**
6
+ **A local-first, hands-free voice assistant plugin for DeepSeek Harness (DSH). Speak, listen, answer prompts, and keep working without touching the computer.**
8
7
 
9
- DSH Live Voice coordinates the microphone, composer, assistant messages, structured questions, and speech output in one plugin — without letting listening and speaking compete with each other. Once voice conversation mode is running, DSH can narrate responses and questions, capture your spoken answers, and continue the conversation while your hands stay free.
8
+ Looking for a DeepSeek Harness voice plugin, DSH microphone plugin, speech-to-text, text-to-speech, or hands-free AI assistant? DSH Live Voice brings those capabilities together in one coordinated plugin.
9
+
10
+ It coordinates the microphone, composer, assistant messages, structured questions, and speech output without letting listening and speaking compete. Once voice conversation mode is running, DSH can narrate responses and questions, capture your spoken answers, and continue the conversation while your hands stay free.
11
+
12
+ ## Built for daily use, maintained with you
13
+
14
+ I use DSH Live Voice for at least eight hours a day. For me, it is the best voice plugin for DSH — and I am committed to making it better through real, everyday use.
15
+
16
+ Have a bug to report, a feature you need, or a pull request to share? [Open an issue](https://github.com/victorwads/dsh-live-voice/issues) or [send a PR](https://github.com/victorwads/dsh-live-voice/pulls). I aim to respond quickly, review contributions promptly, and turn useful suggestions into features. Not every request will be implemented, but I welcome the conversation and will consider what you bring.
10
17
 
11
18
  ## Install
12
19
 
@@ -26,22 +33,23 @@ A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [ds
26
33
 
27
34
  ## Features
28
35
 
29
- | | Capability |
30
- |---|---|
31
- | 🎙️ | Voice typing directly into the DSH composer |
32
- | 👐 | Hands-free conversations: speak, hear responses, and continue without touching the computer |
33
- | ❓ | Spoken DSH structured questions with automatic capture and submission of your answer |
34
- | 💬 | Continuous voice conversations that stay active while you navigate between chats |
35
- | 🗣️ | Configurable voice commands for sending, queueing, clearing, muting, resuming, stopping speech, and ending a conversation |
36
- | 🧠 | Browser SpeechRecognition, local loopback whisper.cpp, or Qwen3-ASR on Apple MLX |
37
- | 🔊 | Browser speech synthesis, native macOS `say`, or Qwen3-TTS on Apple MLX |
38
- | ⏱️ | Manual or automatic sending after configurable silence |
39
- | 🫁 | Stable-silence delay prevents breathing pauses from starting assistant speech |
40
- | 🎧 | Open-microphone mode for headphones |
41
- | 🔒 | Gated microphone mode for speakers |
42
- | ✋ | Pause, resume, stop, and manual interruption controls |
43
- | 🔈 | Play individual assistant messages on demand |
44
- | 🏠 | Whisper audio reaches the local server only through the authenticated DSH host |
36
+ | | Capability |
37
+ | --- | ------------------------------------------------------------------------------------------------------------------------- |
38
+ | 🎙️ | Voice typing directly into the DSH composer |
39
+ | 👐 | Hands-free conversations: speak, hear responses, and continue without touching the computer |
40
+ | ❓ | Spoken DSH structured questions with automatic capture and submission of your answer |
41
+ | 💬 | Continuous voice conversations that stay active while you navigate between chats |
42
+ | 🗣️ | Configurable voice commands for sending, queueing, clearing, muting, resuming, stopping speech, and ending a conversation |
43
+ | ⌨️ | Optional hold-Control push-to-talk anywhere on the page while a composer is open |
44
+ | 🧠 | Browser SpeechRecognition, local loopback whisper.cpp, or Qwen3-ASR on Apple MLX |
45
+ | 🔊 | Browser speech synthesis, native macOS `say`, or Qwen3-TTS on Apple MLX |
46
+ | ⏱️ | Manual or automatic sending after configurable silence |
47
+ | 🫁 | Stable-silence delay prevents breathing pauses from starting assistant speech |
48
+ | 🎧 | Open-microphone mode for headphones |
49
+ | 🔒 | Gated microphone mode for speakers |
50
+ | ✋ | Pause, resume, stop, and manual interruption controls |
51
+ | 🔈 | Play individual assistant messages on demand |
52
+ | 🏠 | Whisper audio reaches the local server only through the authenticated DSH host |
45
53
 
46
54
  ## Hands-free experience
47
55
 
@@ -52,6 +60,7 @@ Start voice conversation mode and choose an automatic delivery mode to keep a co
52
60
  3. Narrate DSH structured questions, listen for your next spoken response, and submit it as a custom answer.
53
61
  4. Keep voice conversation mode active when you move between chats until you explicitly end it.
54
62
  5. Accept configurable spoken commands for common conversation controls.
63
+ 6. Start temporary push-to-talk from anywhere on the page by holding Control, even when the voice bar is off; release to finish queued transcription and automatic delivery, or press Escape to cancel.
55
64
 
56
65
  The hands-free experience coordinates speech input and output locally when you select local engines. The DSH language model itself may still be remote.
57
66
 
@@ -95,7 +104,6 @@ sequenceDiagram
95
104
 
96
105
  The microphone remains open during playback, allowing your voice to pause the assistant.
97
106
 
98
-
99
107
  ### Other conversation settings
100
108
 
101
109
  - **Sending mode:** review and send manually by default, or send automatically after a configurable silence countdown.