dsh-live-voice 0.0.2 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +74 -0
- package/README.md +1 -1
- package/lib/client.js +795 -207
- package/lib/server.js +18 -3
- package/package.json +2 -1
- package/scripts/build.ts +1 -1
- package/src/client/chat.ts +7 -0
- package/src/client/components.ts +453 -190
- package/src/client/index.ts +111 -11
- package/src/client/qwen-settings.ts +1 -1
- package/src/client/styles.ts +11 -7
- package/src/core/coordinator.ts +209 -26
- package/src/core/filters.ts +63 -0
- package/src/core/microphone.ts +7 -1
- package/src/core/ownership.ts +1 -1
- package/src/core/settings.ts +102 -3
- package/src/engines/qwen-http-host.ts +3 -9
- package/src/engines/recognition/browser.ts +23 -0
- package/src/engines/speaking/qwen-http.ts +6 -1
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to DSH Live Voice are documented in this file.
|
|
4
|
+
|
|
5
|
+
## [0.2.0] - Unreleased
|
|
6
|
+
|
|
7
|
+
### Release Changes
|
|
8
|
+
|
|
9
|
+
This release expands conversation control, audio routing, speech filtering, and settings organization.
|
|
10
|
+
|
|
11
|
+
### Features
|
|
12
|
+
|
|
13
|
+
- Add hands-free DSH structured questions in voice conversation mode: narrate each prompt, capture the next spoken response as a custom answer, submit it automatically, and keep the Live Voice status bar visible over the question panel.
|
|
14
|
+
- Add audio input and output device preferences.
|
|
15
|
+
- Add a three-state delivery mode for manual review, queued delivery, and immediate steering.
|
|
16
|
+
- Add configurable speech-input filtering, including minimum-word filtering for final recognition results.
|
|
17
|
+
- Add configurable speech-output filtering to omit code blocks from spoken responses.
|
|
18
|
+
- Add exact voice commands for ending a conversation, muting or resuming listening, stopping speech, clearing input, and sending or queuing recognized text. Commands support multiple comma-separated phrases and normalize punctuation, case, and accents only while matching.
|
|
19
|
+
- Add a remaining-speech segment count to the automatic speech control; it remains expanded while speech is queued.
|
|
20
|
+
- Use the same newline-only segmentation and queue for automatic streaming speech and manual Speak actions.
|
|
21
|
+
- Add tabbed Live Voice settings to organize speech, conversation, commands, and advanced options.
|
|
22
|
+
|
|
23
|
+
### Changes
|
|
24
|
+
|
|
25
|
+
- Improve the Live Voice controls and settings interface for delivery, device, filtering, command, microphone-input, and speech-queue states.
|
|
26
|
+
- Expand coordinator, settings, component, and filter test coverage.
|
|
27
|
+
- Update compiled client and server bundles for the new functionality.
|
|
28
|
+
- Include this changelog in the published package.
|
|
29
|
+
|
|
30
|
+
### Bug Fixes
|
|
31
|
+
|
|
32
|
+
- Keep voice conversation mode active across chat navigation. When the current composer is replaced, the newly mounted chat automatically resumes voice mode; only an explicit End voice conversation action disables it. Composer drafts remain isolated per conversation, while internal controller disposal and hardware handoffs no longer count as user-requested conversation termination.
|
|
33
|
+
- Add lifecycle regression coverage for switching chats while voice mode is active and for preserving an explicit end across subsequent chats.
|
|
34
|
+
- Require at least one recognized word before headphone-mode microphone activity may pause assistant speech; audio activity alone no longer pauses playback.
|
|
35
|
+
- Debounce headphone-mode interruptions to reduce false pauses from short recognition events.
|
|
36
|
+
- Automatically resume speech when an interruption candidate ends without becoming valid user speech.
|
|
37
|
+
- Preserve manual pause behavior separately from automatic interruption handling.
|
|
38
|
+
|
|
39
|
+
## [0.1.0] - 2026-09-15
|
|
40
|
+
|
|
41
|
+
### Bug Fixes
|
|
42
|
+
|
|
43
|
+
- Reliably reset browser speech recognition after it ends or encounters an error.
|
|
44
|
+
- Allow a custom HTTP or HTTPS base URL for the host-local Qwen3 speech service.
|
|
45
|
+
|
|
46
|
+
## [0.0.2] - 2026-09-15
|
|
47
|
+
|
|
48
|
+
### Features
|
|
49
|
+
|
|
50
|
+
- Add host-local Apple MLX support for Qwen3-ASR and Qwen3-TTS.
|
|
51
|
+
- Add Qwen TTS voice selection in Live Voice settings.
|
|
52
|
+
- Expand Live Voice controls and their integration coverage.
|
|
53
|
+
- Add continuous integration, a pre-push hook, formatting configuration, and distribution-artifact validation.
|
|
54
|
+
|
|
55
|
+
### Changes
|
|
56
|
+
|
|
57
|
+
- Update the conversation coordinator, microphone, recognition, synthesis, Whisper, and native macOS `say` integrations for the new engines and controls.
|
|
58
|
+
- Include compiled client and server bundles in the published distribution.
|
|
59
|
+
|
|
60
|
+
## [0.0.1-alpha.1] - 2026-09-15
|
|
61
|
+
|
|
62
|
+
### Features
|
|
63
|
+
|
|
64
|
+
- First functional release of the DSH Live Voice plugin.
|
|
65
|
+
- Coordinate microphone capture, speech recognition, assistant-message delivery, and speech playback.
|
|
66
|
+
- Add voice typing in the DSH composer and continuous voice conversations.
|
|
67
|
+
- Support browser SpeechRecognition and authenticated loopback whisper.cpp HTTP recognition.
|
|
68
|
+
- Support browser speech synthesis and native macOS `say` output.
|
|
69
|
+
- Add Live Voice controls, Whisper settings, build and browser-preview scripts, and an initial test suite.
|
|
70
|
+
|
|
71
|
+
[0.2.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.1.0...HEAD
|
|
72
|
+
[0.1.0]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.2...v0.1.0
|
|
73
|
+
[0.0.2]: https://github.com/victorwads/dsh-live-voice/compare/v0.0.1-alpha.1...v0.0.2
|
|
74
|
+
[0.0.1-alpha.1]: https://github.com/victorwads/dsh-live-voice/tree/v0.0.1-alpha.1
|
package/README.md
CHANGED
|
@@ -100,7 +100,7 @@ Speech processing can run locally, but the DSH language model may still be remot
|
|
|
100
100
|
|
|
101
101
|
## Qwen3 HTTP engine
|
|
102
102
|
|
|
103
|
-
When a compatible Qwen3 speech API is already running on the DSH host, choose **Qwen3 ASR — local MLX server** under Speech recognition and **Qwen3 TTS — local MLX server** under Speech output. Configure its
|
|
103
|
+
When a compatible Qwen3 speech API is already running on the DSH host, choose **Qwen3 ASR — local MLX server** under Speech recognition and **Qwen3 TTS — local MLX server** under Speech output. Configure its base URL in **DSH Settings → Live Voice**. The Qwen server may use any HTTP or HTTPS base URL reachable from the DSH host. The plugin supports the OminiX-API contract and standard OpenAI-style speech endpoints at `GET /health`, `POST /v1/audio/transcriptions`, and `POST /v1/audio/speech`; it does not install, start, stop, or manage that external service or its model weights.
|
|
104
104
|
|
|
105
105
|
## License
|
|
106
106
|
|