dsh-live-voice 0.3.2 β†’ 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (65) hide show
  1. package/README.md +15 -8
  2. package/lib/client.js +5959 -3222
  3. package/lib/server.js +29 -6
  4. package/package.json +9 -30
  5. package/src/app/AGENTS.md +1 -1
  6. package/src/app/ARCHITECTURE.md +1 -1
  7. package/src/app/client/apply.tsx +140 -16
  8. package/src/app/client/composerSelection.ts +92 -0
  9. package/src/app/client/i18n/catalogs/base.ts +33 -0
  10. package/src/app/client/i18n/catalogs/en.ts +36 -0
  11. package/src/app/client/i18n/catalogs/es.ts +36 -0
  12. package/src/app/client/i18n/catalogs/fr.ts +37 -0
  13. package/src/app/client/i18n/catalogs/hi.ts +35 -0
  14. package/src/app/client/i18n/catalogs/pt-BR.ts +37 -0
  15. package/src/app/client/i18n/catalogs/zh.ts +35 -0
  16. package/src/modules/AGENTS.md +1 -1
  17. package/src/modules/conversation/ARCHITECTURE.md +4 -2
  18. package/src/modules/conversation/components/AutoPlaybackToggle.tsx +31 -0
  19. package/src/modules/conversation/components/ConversationControls.tsx +23 -8
  20. package/src/modules/conversation/components/ConversationStatusBar.tsx +72 -94
  21. package/src/modules/conversation/components/DeliveryModeButton.tsx +3 -3
  22. package/src/modules/conversation/components/MeetingControls.tsx +64 -0
  23. package/src/modules/conversation/components/MicrophoneButton.tsx +6 -1
  24. package/src/modules/conversation/components/PlaybackControls.tsx +20 -5
  25. package/src/modules/conversation/components/RecognitionBar.tsx +14 -0
  26. package/src/modules/conversation/components/ScrollingSpeechCaption.tsx +100 -0
  27. package/src/modules/conversation/components/SpeechStatusBar.tsx +99 -0
  28. package/src/modules/conversation/components/Waveform.tsx +8 -3
  29. package/src/modules/conversation/components/conversationStatus.ts +0 -2
  30. package/src/modules/conversation/components/index.ts +1 -0
  31. package/src/modules/conversation/components/speechCaption.ts +27 -0
  32. package/src/modules/conversation/models/chat.ts +14 -2
  33. package/src/modules/conversation/models/meeting.ts +140 -0
  34. package/src/modules/conversation/models/meetingTranscript.ts +56 -0
  35. package/src/modules/core/ARCHITECTURE.md +1 -1
  36. package/src/modules/core/coordinator.ts +204 -51
  37. package/src/modules/core/developerExtension.ts +26 -0
  38. package/src/modules/core/diagnostics.ts +150 -0
  39. package/src/modules/core/filters.ts +20 -15
  40. package/src/modules/core/markdownSpeech.ts +154 -0
  41. package/src/modules/core/settings.ts +17 -13
  42. package/src/modules/core/sharedAudio.ts +82 -0
  43. package/src/modules/core/speechDefaults.ts +21 -0
  44. package/src/modules/recognition/engines/browser/BrowserRecognitionEngine.ts +8 -1
  45. package/src/modules/recognition/engines/whisper/WhisperRecognitionEngine.ts +10 -3
  46. package/src/modules/settings/components/LiveVoiceSettings.tsx +17 -3
  47. package/src/modules/settings/sections/conversation/ConversationDelaySettings.tsx +8 -3
  48. package/src/modules/settings/sections/conversation/DeliverySettings.tsx +25 -2
  49. package/src/modules/speak/engines/audio/HostAudioEngine.ts +19 -2
  50. package/src/modules/speak/engines/browser/BrowserSpeakingEngine.ts +33 -2
  51. package/src/shared/design-system/buttons/ToggleButton.tsx +8 -1
  52. package/src/shared/design-system/icons/icons.ts +4 -0
  53. package/src/styles/index.ts +11 -3
  54. package/CHANGELOG.md +0 -240
  55. package/DEVELOPMENT.md +0 -83
  56. package/PLAN.md +0 -296
  57. package/docs/ARCHITECTURE.md +0 -33
  58. package/docs/CHOOSING-AN-ENGINE.md +0 -120
  59. package/docs/CONFIGURATION.md +0 -163
  60. package/docs/REVIEW.md +0 -35
  61. package/docs/VOICE-LIFECYCLE.md +0 -63
  62. package/scripts/build.ts +0 -62
  63. package/scripts/check-dist.mjs +0 -38
  64. package/scripts/preview-ui.ts +0 -31
  65. package/scripts/probe-browser.ts +0 -46
package/README.md CHANGED
@@ -1,10 +1,12 @@
1
1
  # DSH Live Voice
2
2
 
3
+ Release history: [repository changelog](https://github.com/victorwads/dsh-live-voice/blob/main/CHANGELOG.md).
4
+
3
5
  [![npm version](https://img.shields.io/npm/v/dsh-live-voice?logo=npm&label=npm&color=brightgreen)](https://www.npmjs.com/package/dsh-live-voice)
4
- [![Tested DSH](https://img.shields.io/badge/Tested_DSH-v0.2.0--rc.2-5c5cff?logo=deepseek&logoColor=white)](https://github.com/deepseek-ai/deepseek-harness/releases/tag/dsh-v0.2.0-rc.2)
6
+ [![Tested DSH](https://img.shields.io/badge/Tested_DSH-v0.2.1--alpha.1-5c5cff?logo=deepseek&logoColor=white)](https://github.com/deepseek-ai/deepseek-harness/releases/tag/dsh-v0.2.1-alpha.1)
5
7
  [![License: Apache-2.0](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](LICENSE)
6
8
 
7
- **Tested and working with DeepSeek Harness v0.2.0-rc.2.** This is the current tested DSH version, recorded in the `dshTestedVersion` field in [package.json](package.json).
9
+ **Tested and working with DeepSeek Harness v0.2.1-alpha.1.** This is the current tested DSH version, recorded in the `dshTestedVersion` field in [package.json](package.json).
8
10
 
9
11
  **A local-first, hands-free voice assistant plugin for DeepSeek Harness (DSH).**
10
12
  *Built in Brazil πŸ‡§πŸ‡· and tested daily with Brazilian Portuguese on macOS.*
@@ -52,7 +54,7 @@ Detailed guides for deep-diving into engines and configurations:
52
54
 
53
55
  | Scenario | Recommendation | RAM | Why |
54
56
  | --- | --- | --- | --- |
55
- | πŸ‡§πŸ‡· **Portuguese on macOS** | **Qwen3 ASR (HTTP API)** | ~3 GB | Best accuracy in daily maintainer use. Whisper is second choice. |
57
+ | πŸ‡§πŸ‡· **Portuguese on macOS** | **Qwen3 ASR (HTTP API)** | ~1.5 GB | Best accuracy in daily maintainer use; current RAM usage reported by the maintainer. Whisper is second choice. |
56
58
  | πŸ‡ΊπŸ‡Έ **English on macOS** | **Browser SpeechRecognition** | ~0 GB | Built-in macOS/browser API. Fast, zero extra RAM. |
57
59
  | πŸͺŸ **Windows** | **Qwen3 ASR** or **Whisper HTTP** | ~2–3 GB | Recommended starting point; Windows browser STT varies. |
58
60
  | 🌐 **Multilingual / Other** | **Whisper HTTP (auto)** | ~2 GB | Automatic language detection across dozens of languages. |
@@ -66,18 +68,23 @@ Detailed guides for deep-diving into engines and configurations:
66
68
  - πŸŽ™οΈ **Voice Typing:** Append final recognized speech to the end of the DSH composer without replacing manual edits.
67
69
  - πŸ‘ **Hands-Free Conversation:** Continuous dialogue that stays active across chat sessions.
68
70
  - ❓ **Spoken Structured Questions:** Narrates DSH prompt questions and submits your spoken answer.
69
- - ⌨️ **Hold-to-Talk (Push-to-Talk):** Hold `Control` anywhere on the page to speak; release to send.
71
+ - ⌨️ **Optional Hold-to-Talk (Push-to-Talk):** Enable it in Settings, then hold `Control` anywhere on the page to speak; release to queue the message after the configured send delay. Disabled by default.
70
72
  - 🎧 **Acoustic Mode Isolation:** Gated listening for speakers (no echo) and open-mic interruption for headphones.
71
73
  - πŸ—£οΈ **Spoken Commands:** Control the chat using phrases like *"send"*, *"mute"*, *"clear"*, and *"stop speaking"*.
72
- - 🧹 **Smart Code Filtering:** Automatically skips or summarizes large code blocks instead of reading syntax out loud.
74
+ - ⭐ **Dedicated Speech Bar:** Smoothly scrolling approximate captions, a moving highlight, previous/next navigation, pause/resume, stop, and a current/total segment counter. Captions show the text sent to the speech engine, without claiming word-level alignment.
75
+ - 🧹 **Markdown-Aware Speech:** Remove formatting, announce links without reading full URLs, shorten file paths and line references, read checkbox states and table rows, and replace long code blocks with a localized notice. Preserve custom notices.
76
+ - ⚑ **Responsive Conversation Settings:** Automatic-send delays of 600 ms, 800 ms, or 1–6 seconds (4 seconds by default), plus an assistant response delay of zero to 4 seconds (no delay by default). Recognition and synthesis still contribute to overall latency.
77
+ - πŸ› οΈ **Optional Live Voice Debugger:** Install `dsh-live-voice-debugger` separately to inspect runtime state and queues from the Developer tab in Live Voice Settings. The main plugin works without it.
73
78
  - 🏠 **Local-First & Private:** Audio runs locally on your machine (via Browser APIs, Apple MLX, or whisper.cpp); no external voice telemetry.
74
79
  - 🌐 **Remote-Ready Host Audio:** Qwen and macOS Say synthesize on the DSH host, then DSH delivers compact audio to your browserβ€”so playback works over remote and LAN connections.
75
80
 
76
- ### πŸ§‘β€πŸ’» Coming Soon: Meeting Mode
81
+ ### πŸ§‘β€πŸ’» Meeting Mode β€” Shared Audio Alongside Normal Voice
82
+
83
+ Select **Qwen HTTP** or **Whisper HTTP** recognition, then click the shared-audio icon beside the composer microphone. The first click opens the browser sharing dialog directly; enable audio in that dialog. The shared source gets its own recognition bar, using the same component as normal microphone recognition, without replacing microphone controls or assistant speech playback. Click the shared-audio toggle again to stop only that source. When all three bars are visible, speech/live captions come first, shared audio second, and microphone recognition last, separated by 2px. The microphone bar has one listening/ignoring toggle on the left; queue position appears only in the speech bar, while the automatic-speech toggle remains available in the composer between shared audio and microphone.
77
84
 
78
- **Meeting Mode** is a planned differentiator for collaborative coding conversations. It will keep two independent live transcription streams in the DSH composer: your microphone as **β€œMe:”**, and meeting participants from an explicitly shared screen/system-audio stream as **β€œThem:”**. This creates an editable, real-time record of a code review or technical discussion, so you can manually ask DSH a question with the meeting context already in the composer.
85
+ Normal microphone behavior, voice commands, delivery settings, and speech output remain available. The composer microphone remains a toggle while capturing, and either source can be stopped independently. The shared-audio bar also has a **Timestamp** clock toggle, off by default: new `Me:`/`Them:` blocks can include `[YYYY/MM/DD HH:MM:SS]` using the local machine time at chunk onset, not recognition completion. Consecutive chunks from the same source retain the existing block, and toggling timestamps never rewrites earlier text. With both sources capturing, source changes introduce **β€œMe:”** or **β€œThem:”**; consecutive chunks continue on new lines without repeating labels. Shared audio appends final transcripts but never triggers voice commands or automatic sending itself. Microphone transcripts retain the configured sending behavior.
79
86
 
80
- It will never automatically send the transcript or use meeting audio for voice commands. Sharing system audio will always require explicit browser permission and depends on browser and operating-system support.
87
+ Browser/OS support and the chosen sharing surface determine whether audio is available. Headphones are recommended to avoid unintended feedback when letting the agent speak into a shared meeting or tab. Automated capture, composer and UI tests do not replace real microphone/shared-audio validation.
81
88
 
82
89
  ---
83
90