@bojackduy/opencode-voice 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -419,36 +419,47 @@ then `s`.
419
419
 
420
420
  ### Voice conversation
421
421
 
422
- | Command | Keybind | Description |
423
- | -------------------------- | ---------- | ----------------------------------------- |
424
- | `/voice-conversation` | `leader+v` | Toggle hands-free voice conversation mode |
425
- | `/voice-conversation-stop` | | Exit voice conversation mode |
422
+ | Command | Keybind | Description |
423
+ | -------------------------- | ---------- | ------------------------------------------- |
424
+ | `/voice-conversation` | `leader+v` | Toggle push-to-talk voice conversation mode |
425
+ | `/voice-conversation-stop` | | Exit voice conversation mode |
426
426
 
427
- One key drives the whole loop - its meaning follows the toast on screen:
427
+ One key drives the whole loop. It always means the same thing - "I want the
428
+ floor" - and what it does follows the toast on screen:
428
429
 
429
430
  ```
430
- record -> transcribe -> normalize -> submit -> wait reply -> speak -> record ...
431
+ press -> record -> press -> transcribe -> normalize -> submit -> speak -> press ...
431
432
  ```
432
433
 
433
- | Toast shows | Pressing the key does |
434
- | ------------------------- | ------------------------------------ |
435
- | ● Recording | Finish the turn and submit |
436
- | Speaking... | Pause speech (press again to record) |
437
- | Paused | Record again |
438
- | Waiting / Transcribing... | Exit the mode |
439
-
440
- Empty or failed turns pause instead of re-recording, so the key never
441
- surprises. Saying only a stop phrase (`stop`, `stop stop`, `dừng lại đi`,
442
- ...) ends the mode without submitting - full sentences mentioning stop still
443
- submit normally. `/tts-stop` pauses a speaking reply; `/voice-conversation-stop`
444
- exits from anywhere. While the mode is on, the plain `/stt-record` keys act
445
- as the conversation key and auto TTS stays silent (the loop speaks the reply
446
- itself).
434
+ | Toast shows | Pressing the key does |
435
+ | ------------------------ | ----------------------------------------- |
436
+ | Press `leader+v` to talk | Open the mic (this is the resting state) |
437
+ | ● Recording | Finish the turn and submit |
438
+ | Working... / Speaking | Barge in: cut audio, drop reply, open mic |
439
+ | Answer on screen | Same barge-in, while the agent resumes |
440
+ | Transcribing... | Report busy - too short to interrupt |
441
+
442
+ The loop never reopens the mic by itself: after a reply is spoken the mode rests
443
+ paused, so it cannot record the room while nobody is talking. Barge-in drops
444
+ the pending reply on purpose - pressing means talking now, not hearing the
445
+ rest.
446
+
447
+ Answering a permission or a question is the one thing the key cannot do,
448
+ because that answer is given on screen. The mode announces the gate, keeps the
449
+ mic shut, then speaks the agent's answer once the agent resumes.
450
+
451
+ Exit paths: a stop phrase (`stop`, `stop stop`, `dừng lại đi`, ...),
452
+ `/voice-conversation-stop`, or the global `/voice-cancel` (`leader+.`). Only a
453
+ bare stop phrase ends the mode - full sentences mentioning stop submit
454
+ normally. Empty or failed turns pause instead of re-recording, so the key never
455
+ surprises. `/tts-stop` pauses a speaking reply and rests paused. While the mode
456
+ is on, the plain `/stt-record` keys act as the conversation key and auto TTS
457
+ stays silent (the loop speaks the reply itself).
447
458
 
448
459
  While waiting, assistant text is spoken sentence by sentence as it streams in
449
460
  (local cleanup, no LLM), so the answer starts before the agent turn finishes.
450
461
  Code-like sentences are skipped; replies with nothing streamable fall back to
451
- the full narrated speak. Pressing the key mid-stream exits the mode.
462
+ the full narrated speak.
452
463
 
453
464
  - `ttsNormalizeMode` _(optional)_ - `"llm"` (default, polished narration) or
454
465
  `"local"` (instant deterministic cleanup, no network). Streaming speech
@@ -554,6 +565,19 @@ Options:
554
565
  Not in v1: speaker diarization, automatic summaries/action items, and
555
566
  uploading the notes anywhere - the Markdown file stays local.
556
567
 
568
+ ### Global cancel
569
+
570
+ | Command | Keybind | Description |
571
+ | --------------- | ---------- | ---------------------------------- |
572
+ | `/voice-cancel` | `leader+.` | Stop any in-flight voice operation |
573
+
574
+ One keypress cancels whatever is running: batch recording (or its
575
+ transcribe/normalize pipeline), streaming dictation, TTS playback (or a
576
+ pending speech handoff), the voice-conversation loop, and live-notes capture
577
+ (live notes finishes saving in the background, like `/voice-notes-cancel`).
578
+ Idle invocations toast `Nothing to cancel`. Keybind override:
579
+ `"keybinds": { "voice.cancel": "none" }`.
580
+
557
581
  ### Streaming dictation
558
582
 
559
583
  Live local dictation for short prompts: the mic stays on, partial
@@ -654,6 +678,10 @@ live-mic session was measured.
654
678
  code-heavy responses, or briefly notify for confirmations
655
679
  3. Piper synthesizes speech locally, piped through sox for playback
656
680
 
681
+ The LLM is only a polish layer: when the configured model is unavailable or out
682
+ of quota, TTS speaks the local cleanup automatically instead of going silent.
683
+ Set `ttsNormalizeMode: "local"` to skip narration entirely.
684
+
657
685
  ### Auto TTS
658
686
 
659
687
  When enabled (`/tts-mode`), the plugin automatically speaks:
package/index.js CHANGED
@@ -32,6 +32,7 @@
32
32
  // /voice-notes-stop - stop live notes, flush, and save (palette-only)
33
33
  // /voice-notes-cancel - stop live notes immediately, save in background (palette-only)
34
34
  // /voice-notes-status - show live-notes recording status (palette-only)
35
+ // /voice-cancel - cancel everything in flight (default: <leader>.)
35
36
  // All also palette-accessible via Ctrl+P or /slash. Override via plugin options `keybinds`:
36
37
  // { "keybinds": { "stt.record": "ctrl+r", "tts.speak-last": "none", "voice.conversation": "none" } }
37
38
  // Weird keys [ ] ; were chosen because opencode doesn't use them and shift variants were ignored in terminals.
@@ -43,6 +44,7 @@ import { registerSTT } from "./lib/stt.js";
43
44
  import { registerTTS } from "./lib/tts.js";
44
45
  import { registerConversation } from "./lib/conversation.js";
45
46
  import { registerLiveNotes } from "./lib/live-notes.js";
47
+ import { registerVoiceCancel } from "./lib/cancel.js";
46
48
  import { registerVoiceModel, resolveVoiceProviderModel } from "./lib/voice-model.js";
47
49
  import { createClient } from "./lib/llm-client.js";
48
50
  import { createLogger } from "./lib/logger.js";
@@ -112,6 +114,12 @@ export default {
112
114
  });
113
115
  shared.liveNotes = liveNotes.controller;
114
116
  const voiceModel = registerVoiceModel(api, kv, options, logger);
117
+ const cancel = registerVoiceCancel(api, options, logger, {
118
+ stt: stt.controller,
119
+ tts: tts.controller,
120
+ conversation: shared.conversation,
121
+ liveNotes: shared.liveNotes,
122
+ });
115
123
 
116
124
  api.command.register(() => [
117
125
  ...stt.commands,
@@ -119,6 +127,7 @@ export default {
119
127
  ...conversation.commands,
120
128
  ...liveNotes.commands,
121
129
  ...voiceModel.commands,
130
+ ...cancel.commands,
122
131
  ]);
123
132
  },
124
133
  };
package/lib/cancel.js ADDED
@@ -0,0 +1,106 @@
1
+ // Voice: cancel everything - a global panic button for in-flight voice work.
2
+ //
3
+ // One keypress stops whatever is running: batch STT recording (or the
4
+ // transcribe/normalize pipeline behind it), streaming dictation, TTS
5
+ // playback (or a pending normalize handoff), the voice-conversation loop,
6
+ // and live-notes capture (background-save semantics, like /voice-notes-cancel).
7
+ // Idle invocations never throw and toast "Nothing to cancel".
8
+
9
+ import { clearProcessingToast, clearStreamingToast } from "./stt.js";
10
+
11
+ export function registerVoiceCancel(api, opts, logger, { stt, tts, conversation, liveNotes }) {
12
+ function toast(message, variant = "info") {
13
+ api.ui.toast({ message, variant, duration: 3000 });
14
+ }
15
+
16
+ function guarded(fn) {
17
+ try {
18
+ return fn();
19
+ } catch {
20
+ return undefined;
21
+ }
22
+ }
23
+
24
+ async function cancelEverything() {
25
+ const cancelled = [];
26
+ try {
27
+ // TTS playback: stop() also invalidates a normalize handoff still in
28
+ // flight via the speech generation guard.
29
+ const wasSpeaking = guarded(() => tts?.isSpeaking?.()) === true;
30
+ if (wasSpeaking) {
31
+ guarded(() => tts?.stop?.());
32
+ cancelled.push("TTS");
33
+ }
34
+
35
+ // Streaming dictation owns the mic; batch recording otherwise. A
36
+ // transcribe/normalize pipeline with no active recording is still
37
+ // cancelled via the same path (it bumps the pipeline generation so
38
+ // the stale transcript never lands).
39
+ const streaming = guarded(() => stt?.isStreaming?.()) === true;
40
+ if (streaming) {
41
+ await guarded(() => stt?.cancelStreaming?.());
42
+ cancelled.push("streaming");
43
+ } else {
44
+ const busy =
45
+ guarded(() => stt?.isRecording?.()) === true ||
46
+ guarded(() => stt?.isProcessing?.()) === true;
47
+ if (busy) {
48
+ guarded(() => stt?.cancel?.());
49
+ cancelled.push("recording");
50
+ }
51
+ }
52
+
53
+ if (guarded(() => conversation?.isActive?.()) === true) {
54
+ guarded(() => conversation?.stop?.("cancelled"));
55
+ cancelled.push("conversation");
56
+ }
57
+
58
+ if (guarded(() => liveNotes?.isActive?.()) === true) {
59
+ guarded(() => liveNotes?.cancel?.());
60
+ cancelled.push("live notes");
61
+ }
62
+
63
+ guarded(() => clearProcessingToast());
64
+ guarded(() => clearStreamingToast());
65
+
66
+ if (cancelled.length > 0) {
67
+ toast(`Cancelled: ${cancelled.join(", ")}`, "success");
68
+ } else {
69
+ toast("Nothing to cancel");
70
+ }
71
+ logger?.log?.("VOICE", `Cancel everything: ${cancelled.join(", ") || "idle"}`, "debug");
72
+ } catch (err) {
73
+ logger?.log?.("VOICE", `Cancel everything failed: ${err?.message || err}`, "error");
74
+ }
75
+ }
76
+
77
+ const DEFAULT_KEYBINDS = {
78
+ "voice.cancel": "<leader>.",
79
+ };
80
+ function kb(value) {
81
+ const overrides = opts?.keybinds;
82
+ if (!overrides || typeof overrides !== "object" || Array.isArray(overrides)) {
83
+ return DEFAULT_KEYBINDS[value];
84
+ }
85
+ if (!Object.prototype.hasOwnProperty.call(overrides, value)) return DEFAULT_KEYBINDS[value];
86
+ const v = overrides[value];
87
+ if (!v || v === "none") return undefined;
88
+ return v;
89
+ }
90
+
91
+ const commands = [
92
+ {
93
+ title: "Voice: cancel everything",
94
+ value: "voice.cancel",
95
+ category: "opencode-voice",
96
+ description: "Stop any in-flight voice operation (recording, dictation, speech, loop, notes)",
97
+ ...(kb("voice.cancel") ? { keybind: kb("voice.cancel") } : {}),
98
+ slash: { name: "voice-cancel" },
99
+ onSelect() {
100
+ void cancelEverything();
101
+ },
102
+ },
103
+ ];
104
+
105
+ return { commands };
106
+ }
@@ -1,18 +1,27 @@
1
- // Voice conversation mode: hands-free talk loop with OpenCode.
1
+ // Voice conversation mode: push-to-talk turns with barge-in.
2
2
  //
3
- // One key drives the whole loop; its meaning depends on the state shown in
3
+ // One key drives the whole loop and its meaning is always the same: the user
4
+ // wants the floor. What the key actually does depends on the state shown in
4
5
  // the toast:
5
6
  //
6
- // record -> transcribe -> normalize -> submit -> wait reply -> speak -> record ...
7
+ // press -> record -> press -> transcribe -> normalize -> submit
8
+ // -> wait reply -> speak -> press -> record ...
7
9
  //
10
+ // - paused + key: open the mic (the resting state is the user's turn)
8
11
  // - recording + key: finish the turn and submit
9
- // - speaking + key: pause speech (press again to record the next turn)
10
- // - paused + key: record again
11
- // - waiting/processing + key: exit the mode
12
+ // - speaking/waiting + key: barge in - cut the audio, drop the pending reply,
13
+ // open the mic (one press, never a stop-then-record dance)
14
+ // - processing + key: transcription is too short to interrupt, report busy
12
15
  //
13
- // Saying a stop phrase ("stop", "dừng lại", ...) ends the mode without
14
- // submitting. Empty or failed turns pause instead of auto-recording, so the
15
- // key never surprises.
16
+ // The loop never reopens the mic by itself: after a reply is spoken the mode
17
+ // rests paused, so it can never record the room while nobody is talking. That
18
+ // auto-record is also why a question/permission gate used to eat the answer -
19
+ // the mic opened, the user's muttering was submitted, and the real reply was
20
+ // already torn down. Gates now announce and wait instead.
21
+ //
22
+ // Exit paths: a stop phrase ("stop", "dừng lại", ...), /voice-conversation-stop,
23
+ // or the global /voice-cancel. Empty or failed turns pause instead of
24
+ // auto-recording, so the key never surprises.
16
25
  //
17
26
  // Latency: while waiting, assistant text deltas are spoken sentence by
18
27
  // sentence (local cleanup, no LLM), so the answer starts before the turn
@@ -151,7 +160,34 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
151
160
  phase = "idle";
152
161
  stop("busy");
153
162
  toast("STT busy, conversation off", "warning");
163
+ return;
154
164
  }
165
+ toast(`Listening - press ${keyLabel()} to send`);
166
+ }
167
+
168
+ // The resting state: mic closed, one press to talk. Entered by start() and
169
+ // after a reply is spoken, never by the loop auto-recording.
170
+
171
+ function enterPaused() {
172
+ phase = "paused";
173
+ toast(`Press ${keyLabel()} to talk`);
174
+ }
175
+
176
+ // One press takes the floor: cut whatever audio is queued or playing, drop
177
+ // the pending reply, and open the mic. The mode stays on - a press always
178
+ // means "I want to talk now", so it must never exit silently.
179
+
180
+ function bargeIn() {
181
+ logger?.log("VOICE", "Conversation barge-in", "debug");
182
+ discardStreamAudio();
183
+ clearProcessingToast();
184
+ // Settle the pending reply wait as "stopped" so the abandoned answer is
185
+ // never flushed into the mic we are about to open.
186
+ if (waitCancel) {
187
+ waitCancel("stopped");
188
+ waitCancel = null;
189
+ }
190
+ beginTurn();
155
191
  }
156
192
 
157
193
  function endReplyStream() {
@@ -221,7 +257,7 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
221
257
  const { sentences, rest } = splitSpokenSentences(streamBuffer);
222
258
  const ready = final && rest.trim() ? [...sentences, rest] : sentences;
223
259
  if (ready.length > 0 && streamSpokenChars === 0) {
224
- showProcessingToast("Speaking...");
260
+ showProcessingToast(`Speaking - press ${keyLabel()} to interrupt`);
225
261
  }
226
262
  for (const sentence of ready) {
227
263
  if (!isSpeakableSentence(sentence)) continue;
@@ -234,9 +270,11 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
234
270
  }
235
271
 
236
272
  // Wait for the submitted turn to finish: idle, a permission/question gate,
237
- // or timeout. Resolves early when stop() cancels the wait.
273
+ // or timeout. Resolves early when stop() cancels the wait. gates:false drops
274
+ // the permission/question subscriptions, so a second wait survives the gate
275
+ // the user is answering on screen and only ends on idle/timeout/stop.
238
276
 
239
- function waitForReply(sessionID) {
277
+ function waitForReply(sessionID, { gates = true } = {}) {
240
278
  let settled = false;
241
279
  let unsubs = [];
242
280
  let timer = null;
@@ -276,12 +314,16 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
276
314
  finish("idle");
277
315
  }
278
316
  }),
279
- api.event.on("permission.asked", (event) => {
280
- if (forSession(event.properties)) finish("permission");
281
- }),
282
- api.event.on("question.asked", (event) => {
283
- if (forSession(event.properties)) finish("question");
284
- }),
317
+ ...(gates
318
+ ? [
319
+ api.event.on("permission.asked", (event) => {
320
+ if (forSession(event.properties)) finish("permission");
321
+ }),
322
+ api.event.on("question.asked", (event) => {
323
+ if (forSession(event.properties)) finish("question");
324
+ }),
325
+ ]
326
+ : []),
285
327
  ];
286
328
  timer = setTimeout(() => finish("timeout"), replyTimeoutMs);
287
329
  });
@@ -293,6 +335,7 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
293
335
  if (!active || phase !== "recording") return;
294
336
  const myGen = generation;
295
337
  phase = "processing";
338
+ showProcessingToast("Transcribing...");
296
339
 
297
340
  const res = await stt.transcribeTurn();
298
341
  if (!active || myGen !== generation) return;
@@ -304,12 +347,12 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
304
347
  stop("error");
305
348
  return;
306
349
  }
307
- pauseForRetry("Paused - press the key to retry");
350
+ pauseForRetry(`Paused - press ${keyLabel()} to retry`);
308
351
  return;
309
352
  }
310
353
  if (!res.text) {
311
354
  consecErrors = 0;
312
- pauseForRetry("No speech heard - press the key to try again");
355
+ pauseForRetry(`No speech heard - press ${keyLabel()} to try again`);
313
356
  return;
314
357
  }
315
358
  consecErrors = 0;
@@ -331,7 +374,7 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
331
374
  phase = "waiting";
332
375
  const sessionID = currentSessionID();
333
376
  const turnStart = Date.now();
334
- showProcessingToast("Waiting for reply...");
377
+ showProcessingToast(`Working - press ${keyLabel()} to talk`);
335
378
  startReplyStream(sessionID, turnStart);
336
379
  const tWait = Date.now();
337
380
  const outcome = await waitForReply(sessionID);
@@ -351,46 +394,78 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
351
394
  return;
352
395
  }
353
396
  if (outcome === "permission" || outcome === "question") {
354
- discardStreamAudio();
355
- phase = "speaking";
356
- showProcessingToast("Speaking...");
357
- await tts.speak(
358
- outcome === "permission"
359
- ? "Permission requested. Please check your screen."
360
- : "A question needs your answer. Please check your screen.",
361
- );
362
- clearProcessingToast();
363
- if (!active || myGen !== generation || phase !== "speaking") return;
364
- beginTurn();
397
+ await answerGateThenSpeak(myGen, sessionID, turnStart, outcome);
365
398
  return;
366
399
  }
367
400
  if (outcome === "stopped") return;
368
401
 
369
- // Streamed path: sentences already queued; flush the tail and let the
370
- // audio finish before listening again. Fall back to the full
371
- // LLM-narrated speak only when nothing streamable arrived.
402
+ await speakReply(myGen);
403
+ }
404
+
405
+ // A permission/question gate is answered ON SCREEN, so the mic must stay
406
+ // shut. Announce the gate, then re-arm the reply stream and wait for the
407
+ // agent to resume so its post-answer reply is spoken aloud - before this,
408
+ // endReplyStream() had already run and the answer was dropped in silence.
372
409
 
410
+ async function answerGateThenSpeak(myGen, sessionID, turnStart, outcome) {
411
+ discardStreamAudio();
412
+ phase = "speaking";
413
+ showProcessingToast("Speaking...");
414
+ await tts.speak(
415
+ outcome === "permission"
416
+ ? "Permission requested. Please check your screen."
417
+ : "A question needs your answer. Please check your screen.",
418
+ );
419
+ clearProcessingToast();
420
+ if (!active || myGen !== generation || phase !== "speaking") return;
421
+
422
+ // Still "waiting": the user answers on screen, and a press here is an
423
+ // ordinary barge-in (drop the pending reply, open the mic).
424
+ phase = "waiting";
425
+ showProcessingToast(`Answer on screen - press ${keyLabel()} to talk`);
426
+ startReplyStream(sessionID, turnStart);
427
+ const resumed = await waitForReply(sessionID, { gates: false });
428
+ if (!active || myGen !== generation) return;
429
+ if (resumed === "stopped") return;
430
+ endReplyStream();
431
+ if (resumed === "timeout") {
432
+ discardStreamAudio();
433
+ stop("timeout");
434
+ toast("No reply in time, conversation off", "warning");
435
+ return;
436
+ }
437
+ await speakReply(myGen);
438
+ }
439
+
440
+ // Speak whatever this turn produced, then rest paused with the mic OFF.
441
+ // Streamed sentences are already queued, so only the tail needs flushing;
442
+ // the full LLM-narrated speak is the fallback when nothing streamable
443
+ // arrived. Returns early on any generation/phase change (barge-in, stop).
444
+
445
+ async function speakReply(myGen) {
446
+ phase = "speaking";
447
+ showProcessingToast(`Speaking - press ${keyLabel()} to interrupt`);
373
448
  if (streamSpokenChars > 0) {
374
- phase = "speaking";
375
449
  flushStream(true);
376
450
  await speakTail;
377
451
  clearProcessingToast();
378
452
  if (!active || myGen !== generation || phase !== "speaking") return;
379
453
  await delay(restartDelayMs);
380
454
  if (!active || myGen !== generation || phase !== "speaking") return;
381
- beginTurn();
455
+ enterPaused();
382
456
  return;
383
457
  }
384
-
385
- phase = "speaking";
386
458
  const spoken = await tts.speakAssistantTurn();
387
459
  if (!active || myGen !== generation || phase !== "speaking") return;
460
+ // Drop the sticky speaking status before the paused toast, or it keeps
461
+ // re-announcing "interrupt" once the mic is closed again.
462
+ clearProcessingToast();
388
463
  if (!spoken.spoken) {
389
- toast("No reply to speak, listening again", "warning");
464
+ toast("No reply to speak", "warning");
390
465
  }
391
466
  await delay(restartDelayMs);
392
467
  if (!active || myGen !== generation || phase !== "speaking") return;
393
- beginTurn();
468
+ enterPaused();
394
469
  }
395
470
 
396
471
  // Pause instead of auto-recording so an empty or failed turn never feels
@@ -402,14 +477,15 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
402
477
  if (message) toast(message, "warning");
403
478
  }
404
479
 
405
- // Stop speech and hold the mode. The next keypress records again.
480
+ // Stop speech and hold the mode, mic still closed. The next keypress records
481
+ // again. Only the TTS stop key lands here; the conversation key barges in.
406
482
 
407
483
  function pauseSpeech() {
408
484
  if (!active || phase !== "speaking") return;
409
485
  logger?.log("VOICE", "Conversation speech paused", "debug");
410
486
  tts.stop();
411
487
  phase = "paused";
412
- toast("Paused - press the key to speak");
488
+ toast(`Speech paused - press ${keyLabel()} to talk`);
413
489
  }
414
490
 
415
491
  // Called by the TTS stop command while the mode is on. Returns true when the
@@ -431,7 +507,10 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
431
507
  return false;
432
508
  }
433
509
 
434
- function onKey(source) {
510
+ // One press, one meaning: "I want the floor". The source key (record,
511
+ // submit, toggle) no longer changes the outcome, and a press never exits.
512
+
513
+ function onKey() {
435
514
  if (!active) {
436
515
  start();
437
516
  return;
@@ -440,21 +519,16 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
440
519
  finishTurnFlow();
441
520
  return;
442
521
  }
443
- if (phase === "speaking") {
444
- pauseSpeech();
522
+ if (phase === "speaking" || phase === "waiting") {
523
+ bargeIn();
445
524
  return;
446
525
  }
447
526
  if (phase === "paused") {
448
527
  beginTurn();
449
528
  return;
450
529
  }
451
- // processing | waiting | idle
452
- if (source === "toggle") {
453
- stop("cancelled");
454
- toast("Conversation off");
455
- } else {
456
- toast("Conversation busy, please wait...", "warning");
457
- }
530
+ // processing | idle: the transcription in flight is too short to interrupt.
531
+ toast("Conversation busy, please wait...", "warning");
458
532
  }
459
533
 
460
534
  function start() {
@@ -471,12 +545,12 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
471
545
  generation += 1;
472
546
  turn = 0;
473
547
  consecErrors = 0;
474
- phase = "idle";
548
+ phase = "paused";
475
549
  tts.setConversationActive(true);
476
- stt.setStopHint("<leader>v");
550
+ // Recording toasts must name the key the user actually has bound.
551
+ stt.setStopHint(kb("voice.conversation"));
477
552
  logger?.log("VOICE", "Conversation started", "debug");
478
- toast("Conversation on - one key: send, pause speech, resume");
479
- beginTurn();
553
+ toast(`Conversation on - press ${keyLabel()} to talk`);
480
554
  }
481
555
 
482
556
  api.lifecycle?.onDispose?.(() => stop("dispose"));
@@ -495,16 +569,24 @@ export function registerConversation(api, opts, logger, { stt, tts, isLiveNotesA
495
569
  return v;
496
570
  }
497
571
 
572
+ // Every toast names the key that acts, resolved from the user's keybinds
573
+ // instead of the default literal - a rebound key must never be announced as
574
+ // something the user does not have.
575
+ function keyLabel() {
576
+ return kb("voice.conversation") || "the key";
577
+ }
578
+
498
579
  const commands = [
499
580
  {
500
581
  title: "Voice: conversation mode",
501
582
  value: "voice.conversation",
502
583
  category: "opencode-voice",
503
- description: "Toggle voice conversation (one key: send, pause speech, resume, exit)",
584
+ description:
585
+ "Toggle voice conversation (push-to-talk: press to record, press again to send, press while it speaks to interrupt)",
504
586
  ...(kb("voice.conversation") ? { keybind: kb("voice.conversation") } : {}),
505
587
  slash: { name: "voice-conversation" },
506
588
  onSelect() {
507
- onKey("toggle");
589
+ onKey();
508
590
  },
509
591
  },
510
592
  {
package/lib/live-notes.js CHANGED
@@ -618,5 +618,5 @@ export function registerLiveNotes(api, kv, complete, opts, logger, deps = {}) {
618
618
  },
619
619
  ];
620
620
 
621
- return { commands, controller: { isActive } };
621
+ return { commands, controller: { isActive, stop: stopCommand, cancel: cancelCommand } };
622
622
  }
package/lib/stt.js CHANGED
@@ -271,6 +271,10 @@ let soxProc = null;
271
271
  let soxStderr = "";
272
272
  let recording = false;
273
273
  let processing = false;
274
+ // Invalidated by cancelRecording/discardTurn: a cancel landing between
275
+ // LLM normalization and transcript-insert must prevent the stale text from
276
+ // landing afterwards. Checked in doTranscribePipeline before any append.
277
+ let pipelineGeneration = 0;
274
278
 
275
279
  // ---- Conversation-mode hooks (wired by index.js) ----
276
280
  // When voice conversation is active, the plain record/submit keys defer to the
@@ -1516,6 +1520,7 @@ async function appendTurnText(client, renderer, toast, text, logger) {
1516
1520
  // Clears the sticky toast left up by transcribeTurn and resets the flags.
1517
1521
 
1518
1522
  function discardTurn() {
1523
+ pipelineGeneration += 1;
1519
1524
  clearProcessingToast();
1520
1525
  processing = false;
1521
1526
  recording = false;
@@ -1524,6 +1529,7 @@ function discardTurn() {
1524
1529
  // Cancel an in-progress recording without transcribing.
1525
1530
 
1526
1531
  function cancelRecording(logger) {
1532
+ pipelineGeneration += 1;
1527
1533
  recording = false;
1528
1534
  clearRecordingToast();
1529
1535
  forceKillSox(logger);
@@ -1541,7 +1547,16 @@ async function doTranscribePipeline(
1541
1547
  renderer,
1542
1548
  ) {
1543
1549
  logger?.log("STT", `Pipeline started submit=${submit}`, "debug");
1550
+ const myGen = pipelineGeneration;
1544
1551
  const turn = await transcribeTurn(kv, complete, client, api, toast, systemPrompt, logger);
1552
+ // A cancel during transcribe/normalize invalidates this turn: drop the
1553
+ // stale text instead of appending it after the user cancelled.
1554
+ if (myGen !== pipelineGeneration) {
1555
+ clearProcessingToast();
1556
+ processing = false;
1557
+ recording = false;
1558
+ return;
1559
+ }
1545
1560
  if (!turn.text) return;
1546
1561
  if (submit) {
1547
1562
  await submitTurnText(client, renderer, toast, turn.text, logger);
package/lib/tts.js CHANGED
@@ -66,6 +66,44 @@ export function localSpeechCleanup(text) {
66
66
  return out.replace(/\s+/g, " ").trim();
67
67
  }
68
68
 
69
+ // ---- Speech must never depend on the LLM ----
70
+ // The narrator is a cosmetic polish layer on top of local Piper, which is
71
+ // always available. So when the narrator fails for any reason (quota
72
+ // exhausted, bad auth, unreachable endpoint, missing model, non-2xx) we speak
73
+ // the deterministic local cleanup instead of going silent - this is the one
74
+ // place the decision lives, so every caller benefits.
75
+ //
76
+ // Pure so it is testable without a TUI, a model, or piper.
77
+
78
+ export function resolveSpeechText(rawText, llmResult) {
79
+ if (llmResult?.text) return { text: llmResult.text };
80
+ return { text: localSpeechCleanup(rawText), error: llmResult?.error, fellBack: true };
81
+ }
82
+
83
+ // The llm-client retries with backoff (normalizeRetries in llm-client.js), so a
84
+ // model that is merely out of quota stalls for many seconds before admitting
85
+ // failure - waiting that long for a polish layer is worse than unpolished
86
+ // speech, so bound the wait. The timer is unref'd (same trick as
87
+ // wallTimeout in streaming-stt.js) and cleared on the happy path, so tests and
88
+ // short-lived processes never wait out the full bound.
89
+ const NORMALIZE_FALLBACK_TIMEOUT_MS = 12000;
90
+
91
+ async function normalizeWithinBound(promise) {
92
+ let timer = null;
93
+ const bound = new Promise((resolve) => {
94
+ timer = setTimeout(() => {
95
+ timer = null;
96
+ resolve({ text: null, error: "LLM normalization timed out" });
97
+ }, NORMALIZE_FALLBACK_TIMEOUT_MS);
98
+ timer.unref?.();
99
+ });
100
+ try {
101
+ return await Promise.race([promise, bound]);
102
+ } finally {
103
+ if (timer) clearTimeout(timer);
104
+ }
105
+ }
106
+
69
107
  // Split streamed text into complete spoken sentences. Returns what is ready
70
108
  // plus the trailing incomplete fragment to keep buffering. Sentences that
71
109
  // look like code dumps (overlong, backticks) are skipped by the caller.
@@ -281,6 +319,9 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
281
319
 
282
320
  let piperProc = null;
283
321
  let playProc = null;
322
+ // Invalidated by stopSpeech: a cancel landing between LLM normalization
323
+ // and playback must prevent the stale speech from starting afterwards.
324
+ let speechGen = 0;
284
325
 
285
326
  function killProcs() {
286
327
  if (piperProc) {
@@ -387,6 +428,7 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
387
428
  }
388
429
 
389
430
  function speak(text) {
431
+ const myGen = speechGen;
390
432
  if (!text) return Promise.resolve();
391
433
  const line = text.replace(/\n/g, " ").trim();
392
434
  if (!line) return Promise.resolve();
@@ -397,6 +439,8 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
397
439
  // switching.
398
440
  const voiceModel = resolveVoiceOrWarn(line);
399
441
  if (!voiceModel) return Promise.resolve();
442
+ // A cancel (stopSpeech) during normalization invalidates this speech.
443
+ if (myGen !== speechGen) return Promise.resolve();
400
444
  logger?.log?.("TTS", `Speak requested chars=${line.length} voice=${voiceModel}`, "debug");
401
445
 
402
446
  return new Promise((resolve) => {
@@ -420,6 +464,9 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
420
464
  }
421
465
 
422
466
  function stopSpeech() {
467
+ // Always invalidate: a pending normalize handoff must not speak after
468
+ // the cancel even when nothing is currently playing.
469
+ speechGen += 1;
423
470
  const wasPlaying = piperProc !== null || playProc !== null;
424
471
  killProcs();
425
472
  return wasPlaying;
@@ -458,15 +505,17 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
458
505
  lastSpokenMessageID = result.lastMessageID;
459
506
 
460
507
  showProcessingToast("Normalizing response...");
508
+ const genBeforeNormalize = speechGen;
461
509
  const llmResult = await normalizeOrLocal(result.text, systemAuto, 4096);
462
- if (!llmResult.text) {
510
+ if (genBeforeNormalize !== speechGen) {
511
+ clearProcessingToast();
512
+ return;
513
+ }
514
+ if (!announceNormalize("Auto", llmResult)) {
463
515
  clearProcessingToast();
464
- logger?.log?.("TTS", `Auto normalization failed: ${llmResult.error}`, "warn");
465
- toast(`TTS normalization failed: ${llmResult.error}`, "warning");
466
516
  return;
467
517
  }
468
518
 
469
- logger?.log?.("TTS", `Auto normalization succeeded chars=${llmResult.text.length}`, "debug");
470
519
  clearProcessingToast();
471
520
  await speakWithSessionPrefix(sessionID, llmResult.text, "Ready for your input.");
472
521
  });
@@ -499,15 +548,17 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
499
548
  }
500
549
 
501
550
  showProcessingToast("Normalizing response...");
551
+ const manualGen = speechGen;
502
552
  const llmResult = await normalizeOrLocal(result.text, systemManual, 4096);
503
- if (!llmResult.text) {
553
+ if (manualGen !== speechGen) {
554
+ clearProcessingToast();
555
+ return;
556
+ }
557
+ if (!announceNormalize("Manual", llmResult)) {
504
558
  clearProcessingToast();
505
- logger?.log?.("TTS", `Manual normalization failed: ${llmResult.error}`, "warn");
506
- toast(`TTS normalization failed: ${llmResult.error}`, "warning");
507
559
  return;
508
560
  }
509
561
 
510
- logger?.log?.("TTS", `Manual normalization succeeded chars=${llmResult.text.length}`, "debug");
511
562
  updateProcessingToast("Speaking...");
512
563
  await speak(llmResult.text);
513
564
  clearProcessingToast();
@@ -522,7 +573,31 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
522
573
 
523
574
  async function normalizeOrLocal(text, systemPrompt, maxTokens) {
524
575
  if (ttsLocal) return { text: localSpeechCleanup(text) };
525
- return normalizeForSpeech(text, systemPrompt, maxTokens);
576
+ const llmResult = await normalizeWithinBound(normalizeForSpeech(text, systemPrompt, maxTokens));
577
+ return resolveSpeechText(text, llmResult);
578
+ }
579
+
580
+ // Shared tail for all three normalize call sites: one informational toast
581
+ // when we are speaking the local fallback, and - when the cleanup left
582
+ // nothing speakable (a pure code dump) - a quiet debug log, because nothing
583
+ // failed there, there was just nothing to read. Returns whether to speak.
584
+
585
+ function announceNormalize(label, result) {
586
+ if (!result.text) {
587
+ logger?.log?.("TTS", `${label}: nothing speakable (${result.error})`, "debug");
588
+ return false;
589
+ }
590
+ if (result.fellBack) {
591
+ logger?.log?.(
592
+ "TTS",
593
+ `${label} narration unavailable (${result.error}); speaking local cleanup`,
594
+ "warn",
595
+ );
596
+ toast("Model unavailable - speaking text as-is");
597
+ return true;
598
+ }
599
+ logger?.log?.("TTS", `${label} normalization succeeded chars=${result.text.length}`, "debug");
600
+ return true;
526
601
  }
527
602
 
528
603
  async function speakAssistantTurn() {
@@ -536,15 +611,18 @@ export function registerTTS(api, kv, complete, prompts, opts, logger, deps = {})
536
611
 
537
612
  showProcessingToast("Normalizing response...");
538
613
  const tNormalize = Date.now();
614
+ const turnGen = speechGen;
539
615
  // Auto replies are narrated/summarized, so a tight cap is safe here (the
540
616
  // manual read-aloud keeps the full 4096).
541
617
  const llmResult = await normalizeOrLocal(result.text, systemAuto, 2048);
542
618
  const normalizeMs = Date.now() - tNormalize;
543
- if (!llmResult.text) {
619
+ if (turnGen !== speechGen) {
620
+ clearProcessingToast();
621
+ return { spoken: false };
622
+ }
623
+ if (!announceNormalize("Conversation", llmResult)) {
544
624
  clearProcessingToast();
545
- logger?.log?.("TTS", `Conversation normalization failed: ${llmResult.error}`, "warn");
546
- toast(`TTS normalization failed: ${llmResult.error}`, "warning");
547
- return { spoken: false, error: llmResult.error };
625
+ return { spoken: false };
548
626
  }
549
627
 
550
628
  updateProcessingToast("Speaking...");
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@bojackduy/opencode-voice",
3
- "version": "0.7.0",
3
+ "version": "0.9.0",
4
4
  "description": "Speech-to-text and text-to-speech for OpenCode. Record voice prompts with whisper-cpp, hear responses via Piper TTS, with LLM normalization through any OpenAI-compatible endpoint.",
5
5
  "keywords": [
6
6
  "opencode",