dsh-live-voice 0.0.2 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,74 @@
1
+ # Changelog
2
+
3
+ All notable changes to DSH Live Voice are documented in this file.
4
+
5
+ ## [0.2.0] - Unreleased
6
+
7
+ ### Release Changes
8
+
9
+ This release expands conversation control, audio routing, speech filtering, and settings organization.
10
+
11
+ ### Features
12
+
13
+ - Add hands-free DSH structured questions in voice conversation mode: narrate each prompt, capture the next spoken response as a custom answer, submit it automatically, and keep the Live Voice status bar visible over the question panel.
14
+ - Add audio input and output device preferences.
15
+ - Add a three-state delivery mode for manual review, queued delivery, and immediate steering.
16
+ - Add configurable speech-input filtering, including minimum-word filtering for final recognition results.
17
+ - Add configurable speech-output filtering to omit code blocks from spoken responses.
18
+ - Add exact voice commands for ending a conversation, muting or resuming listening, stopping speech, clearing input, and sending or queuing recognized text. Commands support multiple comma-separated phrases and normalize punctuation, case, and accents only while matching.
19
+ - Add a remaining-speech segment count to the automatic speech control; it remains expanded while speech is queued.
20
+ - Use the same newline-only segmentation and queue for automatic streaming speech and manual Speak actions.
21
+ - Add tabbed Live Voice settings to organize speech, conversation, commands, and advanced options.
22
+
23
+ ### Changes
24
+
25
+ - Improve the Live Voice controls and settings interface for delivery, device, filtering, command, microphone-input, and speech-queue states.
26
+ - Expand coordinator, settings, component, and filter test coverage.
27
+ - Update compiled client and server bundles for the new functionality.
28
+ - Include this changelog in the published package.
29
+
30
+ ### Bug Fixes
31
+
32
+ - Keep voice conversation mode active across chat navigation. When the current composer is replaced, the newly mounted chat automatically resumes voice mode; only an explicit End voice conversation action disables it. Composer drafts remain isolated per conversation, while internal controller disposal and hardware handoffs no longer count as user-requested conversation termination.
33
+ - Add lifecycle regression coverage for switching chats while voice mode is active and for preserving an explicit end across subsequent chats.
34
+ - Require at least one recognized word before headphone-mode microphone activity may pause assistant speech; audio activity alone no longer pauses playback.
35
+ - Debounce headphone-mode interruptions to reduce false pauses from short recognition events.
36
+ - Automatically resume speech when an interruption candidate ends without becoming valid user speech.
37
+ - Preserve manual pause behavior separately from automatic interruption handling.
38
+
39
+ ## [0.1.0] - 2026-09-15
40
+
41
+ ### Bug Fixes
42
+
43
+ - Reliably reset browser speech recognition after it ends or encounters an error.
44
+ - Allow a custom HTTP or HTTPS base URL for the host-local Qwen3 speech service.
45
+
46
+ ## [0.0.2] - 2026-09-15
47
+
48
+ ### Features
49
+
50
+ - Add host-local Apple MLX support for Qwen3-ASR and Qwen3-TTS.
51
+ - Add Qwen TTS voice selection in Live Voice settings.
52
+ - Expand Live Voice controls and their integration coverage.
53
+ - Add continuous integration, a pre-push hook, formatting configuration, and distribution-artifact validation.
54
+
55
+ ### Changes
56
+
57
+ - Update the conversation coordinator, microphone, recognition, synthesis, Whisper, and native macOS `say` integrations for the new engines and controls.
58
+ - Include compiled client and server bundles in the published distribution.
59
+
60
+ ## [0.0.1-alpha.1] - 2026-09-15
61
+
62
+ ### Features
63
+
64
+ - First functional release of the DSH Live Voice plugin.
65
+ - Coordinate microphone capture, speech recognition, assistant-message delivery, and speech playback.
66
+ - Add voice typing in the DSH composer and continuous voice conversations.
67
+ - Support browser SpeechRecognition and authenticated loopback whisper.cpp HTTP recognition.
68
+ - Support browser speech synthesis and native macOS `say` output.
69
+ - Add Live Voice controls, Whisper settings, build and browser-preview scripts, and an initial test suite.
70
+
71
+ [0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...HEAD
72
+ [0.1.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.2...v0.1.0
73
+ [0.0.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.1-alpha.1...v0.0.2
74
+ [0.0.1-alpha.1]: https://github.com/victorwads/dsh-live-voice/tree/v0.0.1-alpha.1
package/README.md CHANGED
@@ -100,7 +100,7 @@ Speech processing can run locally, but the DSH language model may still be remot
100
100
 
101
101
  ## Qwen3 HTTP engine
102
102
 
103
- When a compatible Qwen3 speech API is already running on the DSH host, choose **Qwen3 ASR — local MLX server** under Speech recognition and **Qwen3 TTS — local MLX server** under Speech output. Configure its loopback base URL in **DSH Settings → Live Voice**. The plugin supports the OminiX-API contract and standard OpenAI-style speech endpoints at `GET /health`, `POST /v1/audio/transcriptions`, and `POST /v1/audio/speech`; it does not install, start, stop, or manage that external service or its model weights.
103
+ When a compatible Qwen3 speech API is already running on the DSH host, choose **Qwen3 ASR — local MLX server** under Speech recognition and **Qwen3 TTS — local MLX server** under Speech output. Configure its base URL in **DSH Settings → Live Voice**. The Qwen server may use any HTTP or HTTPS base URL reachable from the DSH host. The plugin supports the OminiX-API contract and standard OpenAI-style speech endpoints at `GET /health`, `POST /v1/audio/transcriptions`, and `POST /v1/audio/speech`; it does not install, start, stop, or manage that external service or its model weights.
104
104
 
105
105
  ## License
106
106