@bojackduy/opencode-voice 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,729 @@
1
+ [![CI](https://github.com/bojackduy/opencode-voice/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/bojackduy/opencode-voice/actions/workflows/ci.yml)
2
+ [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)
3
+ [![npm](https://img.shields.io/npm/v/@bojackduy/opencode-voice)](https://www.npmjs.com/package/@bojackduy/opencode-voice)
4
+ [![Downloads](https://img.shields.io/npm/dm/@bojackduy/opencode-voice)](https://www.npmjs.com/package/@bojackduy/opencode-voice)
5
+
6
+ # opencode-voice
7
+
8
+ Speech-to-text and text-to-speech plugin for [OpenCode](https://opencode.ai/).
9
+
10
+ Record voice prompts with local whisper transcription, hear assistant responses
11
+ spoken aloud via Piper TTS. Both directions use an LLM to normalize text for
12
+ natural speech (fixing homophones, splitting camelCase identifiers, summarizing
13
+ code-heavy responses, etc.).
14
+
15
+ ## Install
16
+
17
+ Add to your `tui.json` (create at `~/.config/opencode/tui.json` if it doesn't
18
+ exist). You must configure at least `endpoint` and `model`:
19
+
20
+ > [!NOTE]
21
+ > **Clobbering default keybinds.** This plugin uses `ctrl+r` for voice
22
+ > recording, but OpenCode assigns it to session rename by default. Session
23
+ > rename is not used frequently and is still accessible via `/rename`, so we
24
+ > clobber the factory default to let the plugin use `ctrl+r` properly. See
25
+ > the `keybinds` section in the config below.
26
+
27
+ ```json
28
+ {
29
+ "$schema": "https://opencode.ai/tui.json",
30
+ "keybinds": {
31
+ "session_rename": "none"
32
+ },
33
+ "plugin": [
34
+ [
35
+ "@bojackduy/opencode-voice",
36
+ {
37
+ "endpoint": "https://api.anthropic.com/v1",
38
+ "model": "claude-haiku-4-5",
39
+ "apiKeyEnv": "ANTHROPIC_API_KEY"
40
+ }
41
+ ]
42
+ ]
43
+ }
44
+ ```
45
+
46
+ ### Refresh cached plugin after updates
47
+
48
+ If OpenCode keeps using an older published version of the plugin after an
49
+ update, clear the cached package and restart OpenCode:
50
+
51
+ ```bash
52
+ rm -rf ~/.cache/opencode/packages/@bojackduy/
53
+ ```
54
+
55
+ ## Prerequisites
56
+
57
+ ### Speech-to-text
58
+
59
+ The plugin uses [whisper.cpp](https://github.com/ggml-org/whisper.cpp) via a
60
+ `whisper-cli` binary and `sox` for microphone capture. Follow the subsection
61
+ for your OS to install the binary and verify your microphone, then run the
62
+ shared **Download model & smoke test** step at the end.
63
+
64
+ #### macOS
65
+
66
+ Install the `whisper-cpp` bottle (ships a `whisper-cli` with Metal enabled on
67
+ Apple Silicon) and `sox`:
68
+
69
+ ```bash
70
+ brew install whisper-cpp sox
71
+ ```
72
+
73
+ Verify your microphone by recording a 3-second clip and playing it back. The
74
+ first `sox -d` invocation triggers a macOS microphone permission prompt —
75
+ grant it in **System Settings → Privacy & Security → Microphone**, then rerun.
76
+ Remove the temp file once you've heard yourself clearly:
77
+
78
+ ```bash
79
+ sox -d /tmp/mic-check.wav trim 0 3 # speak for 3 seconds
80
+ play /tmp/mic-check.wav # you should hear yourself
81
+ rm /tmp/mic-check.wav # delete after verification
82
+ ```
83
+
84
+ #### Linux (including WSL2)
85
+
86
+ Install `sox` with its PulseAudio driver (a separate package on Debian/Ubuntu),
87
+ the PulseAudio tools so the plugin can enumerate input devices via `pactl`,
88
+ and the build tools for whisper.cpp:
89
+
90
+ ```bash
91
+ sudo apt install sox libsox-fmt-pulse pulseaudio-utils build-essential cmake
92
+ ```
93
+
94
+ On WSL2, make sure [WSLg](https://learn.microsoft.com/windows/wsl/tutorials/gui-apps)
95
+ is running — it bridges the Windows microphone into WSL as a PulseAudio source
96
+ (typically named `RDPSource`), which you can then pick with `/stt-mic`.
97
+
98
+ **WSL2 audio troubleshooting.** There is no `/dev/snd` in WSL2 — that is
99
+ normal. Audio goes through WSLg's PulseAudio server at `/mnt/wslg/PulseServer`,
100
+ so ALSA-only tools like `arecord -l` will never list a device. If `/stt-mic`
101
+ finds no devices or `pactl info` fails with `Connection refused`, WSLg's
102
+ PulseAudio is stuck; fix it from Windows PowerShell:
103
+
104
+ ```powershell
105
+ wsl --shutdown # then reopen Ubuntu (closes all WSL sessions)
106
+ ```
107
+
108
+ If the source list is still empty after a restart, check Windows
109
+ **Settings → Privacy & security → Microphone** and enable both "Microphone
110
+ access" and "Let desktop apps access your microphone" (WSLg captures audio via
111
+ a desktop RDP client), then run `wsl --update` for the latest WSLg.
112
+
113
+ Verify your microphone by recording a 3-second clip and playing it back.
114
+ Remove the temp file once you've heard yourself clearly; skip building
115
+ whisper.cpp until this works, otherwise `/stt-mic` will have nothing to select:
116
+
117
+ ```bash
118
+ sox -d /tmp/mic-check.wav trim 0 3 # speak for 3 seconds
119
+ play /tmp/mic-check.wav # you should hear yourself
120
+ rm /tmp/mic-check.wav # delete after verification
121
+ ```
122
+
123
+ `whisper-cli` is not packaged for Linux, so build whisper.cpp from source.
124
+ Pick **one** of the two builds below.
125
+
126
+ **CPU build** — works on any machine, adequate for `tiny`/`base`/`small`
127
+ models:
128
+
129
+ ```bash
130
+ git clone https://github.com/ggml-org/whisper.cpp ~/opt/whisper.cpp
131
+ cmake -B ~/opt/whisper.cpp/build -S ~/opt/whisper.cpp \
132
+ -DCMAKE_BUILD_TYPE=Release -DWHISPER_BUILD_TESTS=OFF
133
+ cmake --build ~/opt/whisper.cpp/build -j --target whisper-cli
134
+ sudo ln -sf ~/opt/whisper.cpp/build/bin/whisper-cli /usr/local/bin/whisper-cli
135
+ ```
136
+
137
+ **CUDA build** — NVIDIA GPU, ~100× faster encode for `medium`/`large` models.
138
+ Check your GPU with `nvidia-smi` and your toolkit with `nvcc --version`, then
139
+ pick the arch code from the table:
140
+
141
+ | GPU family | Arch | `CMAKE_CUDA_ARCHITECTURES` | Min. CUDA |
142
+ | ------------- | --------- | -------------------------- | --------- |
143
+ | RTX 20 / T4 | Turing | `75` | 10.0 |
144
+ | RTX 30 / A100 | Ampere | `86` | 11.0 |
145
+ | RTX 40 / L40 | Ada | `89` | 11.8 |
146
+ | H100 | Hopper | `90` | 12.0 |
147
+ | RTX 50 / B100 | Blackwell | `120` | 13.0 |
148
+
149
+ ```bash
150
+ git clone https://github.com/ggml-org/whisper.cpp ~/opt/whisper.cpp
151
+ cmake -B ~/opt/whisper.cpp/build -S ~/opt/whisper.cpp \
152
+ -DCMAKE_BUILD_TYPE=Release \
153
+ -DGGML_CUDA=ON \
154
+ -DCMAKE_CUDA_ARCHITECTURES=89 \
155
+ -DWHISPER_BUILD_TESTS=OFF
156
+ cmake --build ~/opt/whisper.cpp/build -j --target whisper-cli
157
+ sudo ln -sf ~/opt/whisper.cpp/build/bin/whisper-cli /usr/local/bin/whisper-cli
158
+ ```
159
+
160
+ If you have multiple CUDA toolkits installed (e.g. Blackwell requires CUDA 13
161
+ while the default `nvcc` is 12), also pass `-DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.3/bin/nvcc`
162
+ to point at the matching `nvcc`. CUDA runtime libraries are resolved via
163
+ ldconfig; no `LD_LIBRARY_PATH` is needed.
164
+
165
+ At runtime the plugin records through sox's `pulseaudio` driver when `pactl`
166
+ is available, and falls back to sox's default device otherwise.
167
+
168
+ #### Download model & smoke test
169
+
170
+ Download a whisper model to `~/.local/share/whisper-cpp/` (same path on both
171
+ OSes):
172
+
173
+ ```bash
174
+ mkdir -p ~/.local/share/whisper-cpp
175
+ curl -L -o ~/.local/share/whisper-cpp/ggml-large-v3-turbo-q5_0.bin \
176
+ https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3-turbo-q5_0.bin
177
+ ```
178
+
179
+ Smoke-test the install by transcribing a short recording:
180
+
181
+ ```bash
182
+ sox -d /tmp/smoke.wav trim 0 4 # say something for 4 seconds
183
+ whisper-cli -m ~/.local/share/whisper-cpp/ggml-large-v3-turbo-q5_0.bin \
184
+ -f /tmp/smoke.wav -l auto -nt
185
+ rm /tmp/smoke.wav
186
+ ```
187
+
188
+ Check the first `system_info:` line in the output to confirm the expected
189
+ backend is active:
190
+
191
+ | Install | Expect |
192
+ | ------------------------------ | ----------------------- |
193
+ | macOS Homebrew (Apple Silicon) | `METAL = 1` |
194
+ | Linux CUDA build | `CUDA : ARCHS = <n>` |
195
+ | CPU-only | `METAL = 0` / no `CUDA` |
196
+
197
+ Reference `encode time` on a 4-second clip: CPU `medium` ≈ 15–30 s; CUDA
198
+ `medium` ≈ 100–200 ms; CUDA `large-v3-turbo` ≈ 100–300 ms. Apple Silicon
199
+ Metal timings are hardware-dependent but typically sub-second. If your GPU
200
+ build shows CPU-level timings, the GPU backend failed to load — on Linux,
201
+ re-check `nvidia-smi` and rebuild with the arch code from the table above.
202
+
203
+ ### Text-to-speech
204
+
205
+ Install [Piper](https://github.com/rhasspy/piper):
206
+
207
+ ```bash
208
+ uv tool install piper-tts
209
+ ```
210
+
211
+ Or with pip:
212
+
213
+ ```bash
214
+ pip install piper-tts
215
+ ```
216
+
217
+ The plugin looks for `piper` on your `PATH` (`~/.local/bin` is typically on `PATH`).
218
+
219
+ Download voice models to `~/.local/share/piper-voices/` (English + Vietnamese
220
+ for mixed-language replies):
221
+
222
+ ```bash
223
+ mkdir -p ~/.local/share/piper-voices
224
+ curl -L -o ~/.local/share/piper-voices/en_US-ryan-high.onnx \
225
+ https://huggingface.co/rhasspy/piper-voices/resolve/main/en/en_US/ryan/high/en_US-ryan-high.onnx
226
+ curl -L -o ~/.local/share/piper-voices/en_US-ryan-high.onnx.json \
227
+ https://huggingface.co/rhasspy/piper-voices/resolve/main/en/en_US/ryan/high/en_US-ryan-high.onnx.json
228
+ curl -L -o ~/.local/share/piper-voices/vi_VN-vais1000-medium.onnx \
229
+ https://huggingface.co/rhasspy/piper-voices/resolve/main/vi/vi_VN/vais1000/medium/vi_VN-vais1000-medium.onnx
230
+ curl -L -o ~/.local/share/piper-voices/vi_VN-vais1000-medium.onnx.json \
231
+ https://huggingface.co/rhasspy/piper-voices/resolve/main/vi/vi_VN/vais1000/medium/vi_VN-vais1000-medium.onnx.json
232
+ ```
233
+
234
+ ### Vietnamese/English voice selection
235
+
236
+ Each call to speak (one full reply, or one sentence during conversation
237
+ streaming) picks a single voice for the whole thing - no mid-utterance
238
+ switching, since splicing separately synthesized clips sounds jarring.
239
+ The language is decided by majority vote over words: text is routed to the
240
+ Vietnamese voice only when a meaningful share of its words carry Vietnamese
241
+ diacritics, so a single stray Vietnamese word (e.g. a name) in an otherwise
242
+ English reply no longer flips the whole thing to the Vietnamese voice, and
243
+ vice versa. Toneless Vietnamese (no diacritics) still reads as English -
244
+ that is not distinguishable from English by this heuristic.
245
+
246
+ ### LLM endpoint
247
+
248
+ An OpenAI-compatible LLM endpoint is required for text normalization. For
249
+ speech-to-text it cleans up whisper output (punctuation, filler words, software
250
+ engineering homophones). For text-to-speech it converts markdown into natural
251
+ spoken text.
252
+
253
+ Configure your endpoint in `tui.json` via plugin options. Any OpenAI-compatible
254
+ endpoint works (Anthropic, OpenAI, Ollama, vLLM, LM Studio, etc.). The `apiKeyEnv`
255
+ option is optional - omit it for unauthenticated endpoints like Ollama.
256
+
257
+ ```json
258
+ {
259
+ "plugin": [
260
+ [
261
+ "@bojackduy/opencode-voice",
262
+ {
263
+ "endpoint": "https://api.anthropic.com/v1",
264
+ "model": "claude-haiku-4-5",
265
+ "apiKeyEnv": "ANTHROPIC_API_KEY"
266
+ }
267
+ ]
268
+ ]
269
+ }
270
+ ```
271
+
272
+ For unauthenticated local endpoints (e.g. Ollama):
273
+
274
+ ```json
275
+ {
276
+ "plugin": [
277
+ [
278
+ "@bojackduy/opencode-voice",
279
+ {
280
+ "endpoint": "http://localhost:11434/v1",
281
+ "model": "llama3.2"
282
+ }
283
+ ]
284
+ ]
285
+ }
286
+ ```
287
+
288
+ - `endpoint` _(required)_ - OpenAI-compatible base URL
289
+ - `model` _(required)_ - model name sent to `/chat/completions`
290
+ - `apiKeyEnv` _(optional)_ - environment variable containing the API key
291
+ - `maxTokens` _(optional)_ - maximum completion tokens for normalization calls
292
+ - `reasoningEffort` _(optional)_ - reasoning level for models that support it
293
+ - `chatTemplateKwargs` _(optional)_ - extra keyword arguments passed to the model's chat template (e.g. `{"enable_thinking": false}` for Qwen models to disable chain-of-thought)
294
+ - `retries` _(optional)_ - number of retry attempts for transient LLM failures
295
+ - `tmpDir` _(optional)_ - directory used for the temporary STT recording file (default `/tmp`)
296
+ - `sttLanguage` _(optional)_ - spoken language passed to local `whisper-cli -l` (default `auto`; any whisper.cpp language code, e.g. `en`, `zh`). Can be changed at runtime via `/stt-language`
297
+ - `trimSilence` _(optional)_ - whether to remove leading silence from recordings (default `true`). Set to `false` if your recordings are missing the first word or syllable
298
+
299
+ ### Logging
300
+
301
+ The plugin writes diagnostics through OpenCode's structured app logger. If this plugin is not working with your setup, check the OpenCode log file and, optionally, enable debug mode. See the [OpenCode Docs](https://opencode.ai/docs/troubleshooting/#logs) for details.
302
+
303
+ Routine plugin diagnostics use `debug`; recoverable issues use `warn`; failed
304
+ child processes, API calls, or unexpected exceptions use `error`.
305
+
306
+ ### STT API transcription (optional)
307
+
308
+ Instead of local `whisper-cli`, you can use an OpenAI-compatible speech-to-text
309
+ API (e.g. serving a Whisper model). This is useful when you want to run the
310
+ plugin on a machine without whisper-cpp installed.
311
+
312
+ ```json
313
+ {
314
+ "plugin": [
315
+ [
316
+ "@bojackduy/opencode-voice",
317
+ {
318
+ "sttEndpoint": "http://127.0.0.1:8000/v1",
319
+ "sttModel": "whisper-large-v3-turbo",
320
+ "sttApiKeyEnv": "MY_STT_API_KEY"
321
+ }
322
+ ]
323
+ ]
324
+ }
325
+ ```
326
+
327
+ - `sttEndpoint` _(optional)_ - OpenAI-compatible base URL with `/audio/transcriptions` support
328
+ - `sttModel` _(optional)_ - whisper model name to pass to the API (default: `whisper-large-v3-turbo`). Can be changed at runtime via `/stt-model`, which fetches available whisper models from the endpoint's `/models` listing
329
+ - `sttApiKeyEnv` _(optional)_ - environment variable containing the API key
330
+
331
+ OpenRouter note: when `sttEndpoint` points at `https://openrouter.ai/api/v1`, the plugin automatically uses OpenRouter's JSON/base64 transcription request format instead of multipart upload.
332
+
333
+ ### Voice LLM selection
334
+
335
+ Instead of hand-editing `endpoint`/`model` in `tui.json`, run `/voice-model`
336
+ (palette group `opencode-voice`) to pick from OpenCode's own provider catalog
337
+ — the same providers you connected in OpenCode, no file editing:
338
+
339
+ - The selection (`voice.providerID` + `voice.modelID` in kv) fills in whatever
340
+ `endpoint`/`model`/`apiKeyEnv` the plugin options leave out. Explicit
341
+ options always win; `/voice-model-clear` goes back to options-only.
342
+ - Only env-key providers work (key read live from the provider's env list,
343
+ e.g. `OPENROUTER_API_KEY`). OAuth/Console-managed credentials are invisible
344
+ to plugins — for those, keep explicit `endpoint` options.
345
+ - The picker warns when no key is exported yet.
346
+
347
+ ### Custom prompts
348
+
349
+ The LLM system prompts used for normalization can be fully replaced by pointing
350
+ to your own prompt files. This lets you fine-tune how transcriptions are cleaned
351
+ up or how responses are spoken.
352
+
353
+ ```json
354
+ {
355
+ "plugin": [
356
+ [
357
+ "@bojackduy/opencode-voice",
358
+ {
359
+ "sttPrompt": "~/.config/opencode/stt-prompt.md",
360
+ "ttsAutoPrompt": "~/.config/opencode/tts-auto-prompt.md",
361
+ "ttsManualPrompt": "~/.config/opencode/tts-manual-prompt.md"
362
+ }
363
+ ]
364
+ ]
365
+ }
366
+ ```
367
+
368
+ - `sttContextMessages` _(optional)_ - recent turns sent as knowledge with STT normalization (default `8`)
369
+ - `sttContextChars` _(optional)_ - max chars of that context (default `3000`, tail kept)
370
+ - `sttNormalizeTimeoutMs` _(optional)_ - worst-case budget per normalize call, raw transcript used on timeout (default `15000`)
371
+ - `sttNormalizeMode` _(optional)_ - `"interpretive"` (default: fix misheard words like "they face" → "database" using conversation context + workflow vocabulary) or `"strict"` (transcribe exactly, old behavior). A custom `sttPrompt` file overrides both.
372
+ - `sttAutoSubmit` _(optional)_ - one-shot `/stt-record` submits immediately instead of appending for edit (default `false`; conversation mode always submits)
373
+ - `sttMode` _(optional)_ - `"batch"` (default: record-then-transcribe as today) or `"streaming"` (live local dictation, see below)
374
+ - `sttStreamWindowMs` _(optional)_ - rolling audio window per live transcription (default `10000`)
375
+ - `sttStreamStepMs` _(optional)_ - live transcription cadence in ms (default `1000`)
376
+ - `sttPrompt` _(optional)_ - system prompt for cleaning up whisper transcriptions
377
+ - `ttsAutoPrompt` _(optional)_ - system prompt for auto-speaking assistant responses
378
+ - `ttsManualPrompt` _(optional)_ - system prompt for manually reading responses aloud
379
+
380
+ If a path is not set, the built-in default prompt is used.
381
+
382
+ ## Commands
383
+
384
+ ### Speech-to-text
385
+
386
+ | Command | Keybind | Description |
387
+ | --------------- | ---------- | -------------------------------------- |
388
+ | `/stt-record` | `ctrl+r` | Start/stop recording + transcribe |
389
+ | `/stt-submit` | `leader+r` | Stop recording, transcribe, and submit |
390
+ | `/stt-stop` | | Cancel recording |
391
+ | `/stt-model` | | Select whisper model |
392
+ | `/stt-language` | | Select transcription language |
393
+ | `/stt-mic` | | Select microphone |
394
+
395
+ `/stt-mic` lists CoreAudio input devices on macOS, and PulseAudio sources on
396
+ Linux (via `pactl`, monitor sources excluded). On systems without a supported
397
+ device listing, "System default" uses sox's default device (`sox -d`).
398
+
399
+ `/stt-language` offers a curated list of common languages (plus auto-detect)
400
+ and only affects local `whisper-cli` transcription, not the STT API. Languages
401
+ outside the list can be set via the `sttLanguage` plugin option.
402
+
403
+ In `streaming` mode (`sttMode: "streaming"`), `/stt-record` toggles
404
+ start/finalize of a live dictation session instead of one-shot recording
405
+ (see "Streaming dictation" below); `/stt-submit` finalizes and submits via
406
+ the dictated field only; `/stt-stop` cancels and removes only dictated text.
407
+
408
+ ### Text-to-speech
409
+
410
+ The `leader` key in OpenCode is `ctrl+x`. So `leader+s` means press `ctrl+x`
411
+ then `s`.
412
+
413
+ | Command | Keybind | Description |
414
+ | ------------ | ---------- | ------------------------ |
415
+ | `/tts-speak` | `leader+s` | Read last response aloud |
416
+ | `/tts-mode` | | Toggle auto TTS on/off |
417
+ | `/tts-stop` | `escape` | Stop playback |
418
+ | `/tts-voice` | | Select TTS voice |
419
+
420
+ ### Voice conversation
421
+
422
+ | Command | Keybind | Description |
423
+ | -------------------------- | ---------- | ----------------------------------------- |
424
+ | `/voice-conversation` | `leader+v` | Toggle hands-free voice conversation mode |
425
+ | `/voice-conversation-stop` | | Exit voice conversation mode |
426
+
427
+ One key drives the whole loop - its meaning follows the toast on screen:
428
+
429
+ ```
430
+ record -> transcribe -> normalize -> submit -> wait reply -> speak -> record ...
431
+ ```
432
+
433
+ | Toast shows | Pressing the key does |
434
+ | ------------------------- | ------------------------------------ |
435
+ | ● Recording | Finish the turn and submit |
436
+ | Speaking... | Pause speech (press again to record) |
437
+ | Paused | Record again |
438
+ | Waiting / Transcribing... | Exit the mode |
439
+
440
+ Empty or failed turns pause instead of re-recording, so the key never
441
+ surprises. Saying only a stop phrase (`stop`, `stop stop`, `dừng lại đi`,
442
+ ...) ends the mode without submitting - full sentences mentioning stop still
443
+ submit normally. `/tts-stop` pauses a speaking reply; `/voice-conversation-stop`
444
+ exits from anywhere. While the mode is on, the plain `/stt-record` keys act
445
+ as the conversation key and auto TTS stays silent (the loop speaks the reply
446
+ itself).
447
+
448
+ While waiting, assistant text is spoken sentence by sentence as it streams in
449
+ (local cleanup, no LLM), so the answer starts before the agent turn finishes.
450
+ Code-like sentences are skipped; replies with nothing streamable fall back to
451
+ the full narrated speak. Pressing the key mid-stream exits the mode.
452
+
453
+ - `ttsNormalizeMode` _(optional)_ - `"llm"` (default, polished narration) or
454
+ `"local"` (instant deterministic cleanup, no network). Streaming speech
455
+ always uses the local path.
456
+
457
+ Options: `conversationMaxTurns` (default `50`), `conversationTimeoutMs`
458
+ (default `300000`), `conversationRestartDelayMs` (default `350`),
459
+ `conversationStopPhrases` (custom stop-phrase list). Keybind override:
460
+ `"keybinds": { "voice.conversation": "none" }`.
461
+
462
+ ### Live voice notes
463
+
464
+ Continuous meeting/lecture transcription: the mic stays on and keeps
465
+ recording while a background lane transcribes and cleans up what was already
466
+ said, so a 30-minute meeting never blocks on waiting for whisper or the LLM.
467
+
468
+ | Command | Description |
469
+ | --------------------- | --------------------------------------------------------- |
470
+ | `/voice-notes-start` | Start continuous recording with background transcription |
471
+ | `/voice-notes-stop` | Stop, flush the last bit of audio, and wait for the save |
472
+ | `/voice-notes-cancel` | Stop capturing immediately; finishes saving in background |
473
+ | `/voice-notes-status` | Show elapsed time, chunks written/queued, and write lag |
474
+
475
+ No default keybind (palette/slash only, to avoid clashing with one-shot STT
476
+ or conversation mode - only one of the three can be active at a time).
477
+
478
+ Output goes to `voice-notes/<date>-<time>-notes.md` in your workspace (a
479
+ custom title becomes the file name via `notesTitle`, and `notesDir` moves the
480
+ folder). Two files are written per session:
481
+
482
+ - `<name>.md` - the readable, cleaned transcript with `[HH:MM:SS]` timestamps
483
+ - `<name>.raw.jsonl` - one JSON line per chunk with the raw whisper text,
484
+ normalized text, and timings, for recovery if a normalization pass
485
+ hallucinated
486
+
487
+ Recording is split into chunks on natural pauses (silence-aware), with a
488
+ forced cut every 20 seconds during continuous speech so processing never
489
+ falls more than ~20s behind. A forced cut carries a short audio overlap into
490
+ the next chunk so a word is never fully lost mid-cut; the writer removes the
491
+ duplicated words from the merged transcript automatically.
492
+
493
+ For local whisper transcription, live notes starts a persistent
494
+ `whisper-server` process (loads the model once) instead of the one-shot
495
+ `whisper-cli` path (which reloads the model - and would fall behind - on
496
+ every chunk). If `whisper-server` isn't installed or fails to start, it
497
+ falls back to per-chunk `whisper-cli` automatically. The `sttEndpoint` API
498
+ option, model, and language are reused from one-shot STT settings.
499
+
500
+ Multi-TUI: each OpenCode window runs its own server — the first takes
501
+ 127.0.0.1:8090, the next window takes the next free port up (8091, ...).
502
+ No configuration needed; each server loads its own model copy, so allow
503
+ roughly 2 GB RAM per window.
504
+
505
+ Far-field voices (a professor meters from the mic) are enhanced before
506
+ whisper hears them: each chunk is measured and gained up to a healthy speech
507
+ level with `sox` (`highpass 80` for room rumble + adaptive `gain -l` with the
508
+ limiter on, so quiet speech gets louder and loud speech is untouched). The
509
+ per-chunk levels land in the JSONL sidecar (`audio.rmsBefore/rmsAfter/gainDb`)
510
+ so you can see what the room actually sounded like. Silence detection also
511
+ learns the room: the first ~1.5s of mic audio sets the speech/silence
512
+ threshold instead of assuming one value fits every room.
513
+
514
+ Whisper hallucinates YouTube outros on quiet audio ("Hãy subscribe cho kênh
515
+ ...", "thank you for watching") - those are filtered by phrase (English +
516
+ Vietnamese), and any transcript that repeats verbatim 3x in a row is treated
517
+ as a repeat hallucination. Filtered chunks stay in the JSONL sidecar but never
518
+ reach the readable transcript.
519
+
520
+ Classroom tip: the LLM cleanup pass is slow over a network gateway and adds
521
+ lag per chunk - for lectures set `notesNormalize: false` (raw whisper text,
522
+ still enhanced + filtered) and clean up afterwards.
523
+
524
+ Options:
525
+
526
+ - `notesDir` _(optional, default `"voice-notes"`)_ - output folder, relative
527
+ to the workspace unless absolute
528
+ - `notesTitle` _(optional)_ - included in the session file name
529
+ - `notesChunkMaxSeconds` _(optional, default `20`)_ - forced cut duration
530
+ - `notesMinChunkSeconds` _(optional, default `3`)_ - minimum audio before a
531
+ natural (silence-triggered) cut is allowed
532
+ - `notesSilenceMs` _(optional, default `700`)_ - pause length that closes a
533
+ chunk naturally
534
+ - `notesOverlapMs` _(optional, default `400`)_ - audio overlap carried across
535
+ a forced cut
536
+ - `notesNormalize` _(optional, default `true`)_ - set `false` to skip the LLM
537
+ cleanup pass and write raw whisper text only
538
+ - `notesKeepAudio` _(optional, default `false`)_ - keep each transcribed
539
+ chunk's WAV file under `<name>-audio/` instead of deleting it
540
+ - `notesUseWhisperServer` _(optional, default `true`)_ - set `false` to force
541
+ per-chunk `whisper-cli` even when `whisper-server` is available
542
+ - `notesEnhance` _(optional, default `true`)_ - voice-only preprocessing
543
+ (adaptive gain + rumble filter) before whisper; set `false` if the mic is
544
+ already close/loud
545
+ - `notesEnhanceTargetRms` _(optional, default `0.1`)_ - speech level chunks
546
+ are gained up to
547
+ - `notesEnhanceMaxGainDb` _(optional, default `24`)_ - gain ceiling so
548
+ near-silence never becomes amplified noise
549
+ - `notesSilenceRms` _(optional)_ - explicit speech/silence RMS threshold,
550
+ disables room auto-calibration when set
551
+ - `notesCalibrationMs` _(optional, default `1500`)_ - mic audio used to learn
552
+ the room noise floor at session start; set `0` to disable
553
+
554
+ Not in v1: speaker diarization, automatic summaries/action items, and
555
+ uploading the notes anywhere - the Markdown file stays local.
556
+
557
+ ### Streaming dictation
558
+
559
+ Live local dictation for short prompts: the mic stays on, partial
560
+ transcriptions stream in as you speak, and one press of `/stt-record`
561
+ finalizes the text into the prompt box. Set `"sttMode": "streaming"` in the
562
+ plugin options (default `"batch"` keeps the record-then-transcribe flow
563
+ exactly as before).
564
+
565
+ How it works: a rolling audio window (default 10 s, every 1 s) is
566
+ transcribed by a persistent local `whisper-server` (model loaded once, no
567
+ per-utterance reload) and folded into stable + tentative text. No LLM is
568
+ involved anywhere in the streaming path — raw whisper text only — so it
569
+ costs zero quota and works fully offline once the model is downloaded.
570
+
571
+ What you see: a sticky status toast stays up for the whole session so
572
+ there is never a silent gap — `Loading speech model…` during cold start,
573
+ `● Streaming dictation — listening` with an elapsed timer while live
574
+ (plus a `(listening…)` note when inference runs slow), and `Finalizing…`
575
+ between the second keypress and the insert. The OpenCode renderer only
576
+ exposes `insertText`/`submit`, with no live range-replacement API, so
577
+ partials render as preview toasts (`🎙 ...`) alongside (not instead of)
578
+ the sticky status, and the full text is inserted EXACTLY ONCE at
579
+ finalize. If no editable field is focused, the transcript is kept via
580
+ toast instead of being redirected into the chat. `/stt-submit` submits
581
+ only through the dictated field's own `submit()` — a field that cannot
582
+ submit keeps its text un-submitted rather than falling through to the
583
+ primary chat. `/stt-stop` cancels and removes only plugin-dictated
584
+ text, never your typed prefix/suffix.
585
+
586
+ Warm server: the whisper model stays loaded across dictations (finalize
587
+ and cancel keep the lease); it reloads only on model/language change,
588
+ plugin unload, or a proven server fault. If the server fails mid-stream,
589
+ the error names the cause and the recorded audio is kept for a later
590
+ batch run. Streaming refuses while batch recording, live notes, or voice
591
+ conversation is active (and vice versa). Like live notes, streaming takes
592
+ the next free port when 8090 is owned by another window (see above).
593
+
594
+ Classroom/batch guidance: streaming is tuned for short dictations (prompt
595
+ length, seconds to a few minutes). For lectures and meetings, use live
596
+ voice notes instead — its chunked pipeline, overlap handling, and raw
597
+ JSONL sidecar are built for 30-minute sessions, while streaming keeps at
598
+ most ~5 minutes of bounded spool audio per session.
599
+
600
+ Options:
601
+
602
+ - `sttMode` _(optional, default `"batch"`)_ - `"streaming"` selects this mode
603
+ - `sttStreamWindowMs` _(optional, default `10000`)_ - rolling audio window
604
+ - `sttStreamStepMs` _(optional, default `1000`)_ - live transcription cadence
605
+ - `sttAutoSubmit` applies after streaming finalize (submit via the dictated
606
+ field only, as above)
607
+
608
+ Measured latencies (Apple M3 Pro, 36 GB, whisper.cpp 1.9.4 with Metal,
609
+ `ggml-large-v3-turbo-q5_0.bin`, scripted speech through the real
610
+ `/inference` path — goals from the plan, measured here, not promises):
611
+
612
+ | Measure | Measured |
613
+ | ------------------------------------------------------------------------ | --------------------------------- |
614
+ | Cold model-load to server ready | ~1.0–1.5 s |
615
+ | First inference, cold server, 10 s window | ~0.9 s (≈1.9 s total after spawn) |
616
+ | Warm 10 s-window transcribe (plain) | ~0.9 s |
617
+ | Warm 10 s-window transcribe (verbose format, with segments; opt-in only) | ~1.7 s |
618
+ | Stop-to-final, 2 s tail | ~0.8 s |
619
+
620
+ Honest limitations: ticks default to the plain format (~0.9 s per 10 s
621
+ window on the hardware above, no segments fetched — the stability
622
+ tracker is text-anchored and never reads segments), so partials refresh
623
+ at roughly that rate, not every cadence tick — the scheduler coalesces
624
+ while inference is busy, so partials are flicker-free but not instant.
625
+ Only one smaller-model data point exists (none: no smaller model file is
626
+ installed locally, and nothing was downloaded for measurement). E2E
627
+ coverage is mic-less/TUI-less by necessity: session flows are proven
628
+ through the real command paths with faked capture/renderer
629
+ (`test/streaming-e2e.test.js`, `test/streaming-gaps.test.js`), while the
630
+ numbers above come from the real server path with scripted audio — no
631
+ live-mic session was measured.
632
+
633
+ ## How it works
634
+
635
+ ### STT pipeline
636
+
637
+ 1. `sox` records audio from your microphone (CoreAudio on macOS, PulseAudio on
638
+ Linux when `pactl` is available, sox default device otherwise)
639
+ 2. `whisper-cli` transcribes locally using a ggml model, or an OpenAI-compatible
640
+ API endpoint if `sttEndpoint` is configured
641
+ 3. LLM normalizes the transcription: fixes punctuation, removes filler words,
642
+ corrects software engineering homophones ("Jason" to "JSON", "bullion" to
643
+ "boolean", etc.)
644
+ 4. Cleaned text is appended to the OpenCode prompt, or submitted immediately
645
+ when `/stt-submit` is used. If normalization fails (e.g. LLM endpoint
646
+ unreachable), the raw transcription is used as a fallback so you never lose
647
+ your input
648
+
649
+ ### TTS pipeline
650
+
651
+ 1. When the assistant finishes responding (or on manual trigger), the response
652
+ text is sent to the LLM for speech normalization
653
+ 2. The LLM decides how to handle it: narrate simple answers, summarize
654
+ code-heavy responses, or briefly notify for confirmations
655
+ 3. Piper synthesizes speech locally, piped through sox for playback
656
+
657
+ ### Auto TTS
658
+
659
+ When enabled (`/tts-mode`), the plugin automatically speaks:
660
+
661
+ - Assistant responses when a session goes idle after work
662
+ - Permission requests
663
+ - Questions that need your answer
664
+
665
+ ## Contributing
666
+
667
+ opencode-voice is open to contributions and ideas!
668
+
669
+ ### Issue conventions
670
+
671
+ **Format:** `type: brief description`
672
+
673
+ - `feat:` new features or functionality
674
+ - `fix:` bug fixes
675
+ - `enhance:` improvements to existing features
676
+ - `chore:` maintenance tasks, dependencies, cleanup
677
+ - `docs:` documentation updates
678
+ - `build:` build system, CI/CD changes
679
+
680
+ ### Development
681
+
682
+ ```bash
683
+ npm run check # lint + fmt
684
+ npm run lint # oxlint
685
+ npm run fmt # oxfmt --check
686
+ npm run fmt:fix # oxfmt --write
687
+ ```
688
+
689
+ ### Test local plugin in OpenCode
690
+
691
+ To test unpublished changes in the OpenCode TUI, point `~/.config/opencode/tui.json`
692
+ at the local repo path, not the npm package name:
693
+
694
+ ```json
695
+ {
696
+ "$schema": "https://opencode.ai/tui.json",
697
+ "plugin": ["/Users/your-user/opencode-voice"]
698
+ }
699
+ ```
700
+
701
+ ### Optional macOS Hammerspoon integration
702
+
703
+ If you use macOS, [Hammerspoon](https://www.hammerspoon.org/), and
704
+ [Ghostty](https://ghostty.org/), see
705
+ [`examples/hammerspoon/ghostty-fn.lua`](examples/hammerspoon/ghostty-fn.lua)
706
+ for an optional global `Fn` key setup.
707
+
708
+ Behavior:
709
+
710
+ - Press `Fn` to send `ctrl+r` and start recording.
711
+ - Hold `Fn` for at least 0.5 seconds and release to send `leader+r`, which
712
+ stops recording, normalizes, and submits the prompt.
713
+
714
+ Notes:
715
+
716
+ - It assumes OpenCode is using the default leader key, `ctrl+x`.
717
+ - It assumes OpenCode is running in Ghostty terminal `1`.
718
+ - It is best used as a push-to-talk flow: hold `Fn` while speaking, then
719
+ release to submit.
720
+ - Adjust `APP_NAME`, `TARGET_TERMINAL`, and `LONG_PRESS_THRESHOLD_SECONDS` to
721
+ fit your setup.
722
+
723
+ ### Release process
724
+
725
+ Manual releases via opencode; see [RELEASE_PROCESS.md](RELEASE_PROCESS.md).
726
+
727
+ ## License
728
+
729
+ This project is licensed under the [MIT License](LICENSE).