dsh-live-voice 0.0.1-alpha.1 → 0.0.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/DEVELOPMENT.md +16 -0
- package/PLAN.md +4 -0
- package/README.md +24 -5
- package/lib/client.js +1394 -254
- package/lib/server.js +580 -69
- package/package.json +22 -3
- package/scripts/build.ts +44 -8
- package/scripts/check-dist.mjs +38 -0
- package/scripts/preview-ui.ts +30 -7
- package/scripts/probe-browser.ts +44 -12
- package/src/client/chat.ts +32 -11
- package/src/client/components.ts +656 -88
- package/src/client/index.ts +362 -180
- package/src/client/qwen-settings.ts +144 -0
- package/src/client/styles.ts +6 -0
- package/src/client/whisper-settings.ts +142 -37
- package/src/core/coordinator.ts +402 -149
- package/src/core/microphone.ts +90 -23
- package/src/core/ownership.ts +24 -11
- package/src/core/settings.ts +107 -23
- package/src/core/transcript.ts +14 -5
- package/src/engines/qwen-http-host.ts +246 -0
- package/src/engines/recognition/browser.ts +167 -44
- package/src/engines/recognition/qwen-http.ts +36 -0
- package/src/engines/recognition/whisper-http-host.ts +198 -52
- package/src/engines/recognition/whisper-http.ts +186 -20
- package/src/engines/speaking/browser.ts +76 -20
- package/src/engines/speaking/qwen-http.ts +119 -0
- package/src/engines/speaking/say-client.ts +40 -12
- package/src/engines/speaking/say.ts +119 -40
- package/src/server.ts +295 -47
package/DEVELOPMENT.md
CHANGED
|
@@ -9,6 +9,22 @@ npm run build
|
|
|
9
9
|
|
|
10
10
|
Implementation, test, and build-script source is TypeScript (`.ts`). `npm run typecheck` compiles the TypeScript project; `npm test` typechecks, builds the browser and host bundles, transpiles the TypeScript tests, then runs them. Tests exercise engines, coordination, transcript edits, RPC, configuration, and UI contracts, but do not record the microphone. `npm run dev` watches the browser and host bundles only; it is not a replacement DSH server.
|
|
11
11
|
|
|
12
|
+
## Qwen3 HTTP engine contract
|
|
13
|
+
|
|
14
|
+
The plugin can connect to a separately managed Qwen3 speech service on an unauthenticated loopback HTTP base URL. The service lifecycle and weights are deliberately outside this repository. It must expose `GET /health`, `POST /v1/audio/transcriptions`, and `POST /v1/audio/speech`. Standard OpenAI multipart transcription and OminiX-API's JSON/base64 transcription contract are detected automatically. Host configuration is stored at `~/.dsh/dsh-live-voice-qwen.json` with owner-only permissions and rejects non-loopback URLs.
|
|
15
|
+
|
|
16
|
+
The plugin's browser never calls the speech service directly. It sends WAV/text through authenticated same-origin DSH routes; the host validates input and then calls loopback. Qwen TTS returns a WAV that is played on the browser device. Qwen STT is utterance-based: the plugin's VAD uses Natural (1500 ms) by default, sends mono 16 kHz PCM16 WAV, and maps `pt-BR` to Portuguese.
|
|
17
|
+
|
|
18
|
+
## Repository checks
|
|
19
|
+
|
|
20
|
+
Enable the repository's versioned Git hooks once per checkout:
|
|
21
|
+
|
|
22
|
+
```sh
|
|
23
|
+
npm run setup:hooks
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
The pre-push hook runs `npm run check:dist`. It rebuilds `lib/client.js` and `lib/server.js`, requires both files to be tracked, and blocks the push when the committed runtime bundles do not match the current source. The same guard runs in GitHub Actions on every push and pull request, so drift is still reported when a local hook is missing or bypassed. Configure the CI job as a required branch check if direct pushes must be rejected rather than reported after they arrive.
|
|
27
|
+
|
|
12
28
|
## Current local installation
|
|
13
29
|
|
|
14
30
|
Install through the official CLI only:
|
package/PLAN.md
CHANGED
|
@@ -2,6 +2,10 @@
|
|
|
2
2
|
|
|
3
3
|
This document describes the product plan and local implementation progress. The published npm version is a documentation placeholder; the working tree now contains an initial plugin undergoing integration validation. Features below describe intended behavior unless verified in the progress section.
|
|
4
4
|
|
|
5
|
+
## Qwen3 local engine
|
|
6
|
+
|
|
7
|
+
Qwen3 is available as independent recognition and speaking selections backed by a separately managed HTTP process on the DSH host. Runtime installation, weights, and service lifecycle stay outside the repository. The authenticated host bridge accepts only a loopback base URL, validates bounded mono PCM16 WAV before ASR forwarding, maps `pt-BR` to Portuguese, forwards TTS text without logging it, and returns WAV audio for browser playback. Browser cancellation aborts pending host/model requests and stale synthesized audio cannot begin playback after cancellation. Automated plugin coverage includes configuration normalization, loopback validation, HTTP payloads, WAV transport, browser playback, and route cleanup. Native model and endpoint acceptance must be reported separately from these tests.
|
|
8
|
+
|
|
5
9
|
## TypeScript migration
|
|
6
10
|
|
|
7
11
|
The implementation, tests, and developer scripts now use `.ts` source files. `tsconfig.json` centralizes compiler settings and `npm run typecheck` is part of the build path. The build generates a bundled browser client (`lib/client.js`), an ESM host bundle (`lib/server.js`), and temporary transpiled test artifacts that are ignored by Git. This is a source-language migration; the DSH runtime still receives JavaScript bundles. The compiler setup is transitional: current converted legacy files use `@ts-nocheck`, so this is not yet a claim that every implementation boundary has complete static typing.
|
package/README.md
CHANGED
|
@@ -1,12 +1,23 @@
|
|
|
1
1
|
# DSH Live Voice
|
|
2
2
|
|
|
3
|
-
[](https://www.npmjs.com/package/dsh-live-voice)
|
|
3
|
+
[](https://www.npmjs.com/package/dsh-live-voice)
|
|
4
|
+
[](https://dsh.pub/en/plugins/dsh-live-voice/)
|
|
4
5
|
[](LICENSE)
|
|
5
6
|
|
|
6
7
|
**Local-first speech recognition, voice output, and continuous voice conversations for DSH.**
|
|
7
8
|
|
|
8
9
|
DSH Live Voice coordinates the microphone, composer, assistant messages, and speech output in one plugin — without letting listening and speaking compete with each other.
|
|
9
10
|
|
|
11
|
+
## Install
|
|
12
|
+
|
|
13
|
+
Install the public repository through dsh.pub into your DSH web profile:
|
|
14
|
+
|
|
15
|
+
```sh
|
|
16
|
+
npx dshpub add victorwads/dsh-live-voice --profile web
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
The installer resolves the public repository to an exact commit before adding the bundle. Open **DSH Settings → Live Voice** after the next normal DSH startup.
|
|
20
|
+
|
|
10
21
|
## Why this project exists and Acknowledgments
|
|
11
22
|
|
|
12
23
|
Listening and speaking should work together, so you can interrupt and be heard without the assistant’s voice getting in the way. [Read the story behind the project](HISTORY.md).
|
|
@@ -19,8 +30,8 @@ A heartfelt thank you to [GooDAnDReaDY](https://github.com/GooDAnDReaDY) for [ds
|
|
|
19
30
|
|---|---|
|
|
20
31
|
| 🎙️ | Voice typing directly into the DSH composer |
|
|
21
32
|
| 💬 | Continuous voice conversations with automatic assistant speech |
|
|
22
|
-
| 🧠 | Browser SpeechRecognition
|
|
23
|
-
| 🔊 | Browser speech synthesis
|
|
33
|
+
| 🧠 | Browser SpeechRecognition, local loopback whisper.cpp, or Qwen3-ASR on Apple MLX |
|
|
34
|
+
| 🔊 | Browser speech synthesis, native macOS `say`, or Qwen3-TTS on Apple MLX |
|
|
24
35
|
| ⏱️ | Manual or automatic sending after configurable silence |
|
|
25
36
|
| 🫁 | Stable-silence delay prevents breathing pauses from starting assistant speech |
|
|
26
37
|
| 🎧 | Open-microphone mode for headphones |
|
|
@@ -80,13 +91,21 @@ The microphone remains open during playback, allowing your voice to pause the as
|
|
|
80
91
|
|
|
81
92
|
## Local-first architecture
|
|
82
93
|
|
|
83
|
-
- **Recognition:** Browser SpeechRecognition
|
|
84
|
-
- **Speech output:** browser/device audio
|
|
94
|
+
- **Recognition:** Browser SpeechRecognition, loopback whisper.cpp HTTP, or Qwen3-ASR through a host-local Apple MLX server.
|
|
95
|
+
- **Speech output:** browser/device audio, native macOS `say`, or host-local Qwen3-TTS with WAV playback in the browser.
|
|
85
96
|
- **Whisper transport:** complete WAV utterances through authenticated same-origin DSH routes.
|
|
86
97
|
- **Privacy:** raw audio and transcripts are not logged by default.
|
|
87
98
|
|
|
88
99
|
Speech processing can run locally, but the DSH language model may still be remote.
|
|
89
100
|
|
|
101
|
+
## Qwen3 HTTP engine
|
|
102
|
+
|
|
103
|
+
When a compatible Qwen3 speech API is already running on the DSH host, choose **Qwen3 ASR — local MLX server** under Speech recognition and **Qwen3 TTS — local MLX server** under Speech output. Configure its loopback base URL in **DSH Settings → Live Voice**. The plugin supports the OminiX-API contract and standard OpenAI-style speech endpoints at `GET /health`, `POST /v1/audio/transcriptions`, and `POST /v1/audio/speech`; it does not install, start, stop, or manage that external service or its model weights.
|
|
104
|
+
|
|
90
105
|
## License
|
|
91
106
|
|
|
92
107
|
[GPL-3.0-only](LICENSE). Commercial use and redistribution are allowed subject to the GPL. Third-party speech engines and models may have separate licenses.
|
|
108
|
+
|
|
109
|
+
## Keywords
|
|
110
|
+
|
|
111
|
+
`dsh`, `dsh-plugin`, `deepseek-harness`, `local-first`, `local-voice`, `voice-assistant`, `voice-conversation`, `continuous-conversation`, `voice-dictation`, `speech-to-text`, `text-to-speech`, `speech-recognition`, `speech-synthesis`, `stt`, `tts`, `whisper`, `whisper-cpp`, `web-speech-api`, `macos-say`, `turn-taking`, `voice-interruption`, `silence-detection`
|