mac-voice-mcp 0.2.0 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,30 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
5
5
 
6
6
  ## [Unreleased]
7
7
 
8
+ ## [0.3.1] - 2026-09-29
9
+
10
+ ### Fixed
11
+ - **Recording broke on Macs without SoX when Bluetooth headphones were connected** ([#6](https://github.com/jeet0007/mac-voice-mcp/issues/6)). Every turn failed with "pure digital silence".
12
+ - The ffmpeg fallback now records from the input chosen in System Settings (`:default`) instead of device number 0. `VOICE_MCP_FFMPEG_DEVICE` still overrides it.
13
+ - `voice_setup` treats ffmpeg as a fallback, recommends SoX, and installs it with your OK.
14
+ - Silence recorded through ffmpeg now points at the input device and SoX, not only at mic permissions.
15
+ - **`voice_setup` notices tools installed since the last check**, such as `brew install sox` run in a terminal, without restarting the app.
16
+
17
+ ## [0.3.0] - 2026-09-28
18
+
19
+ ### Added
20
+ - **Voice conversations stay in voice (Claude Code plugin).** Once you've answered out loud, a plugin hook sends Claude back to reply by voice if it tries to answer in text. It does this at most once per spoken answer, so it can never loop. Voice mode ends when you type, when you don't answer, or when Claude says goodbye with `listen: false`. It's tracked per session and expires after 30 idle minutes. `VOICE_MCP_STAY_IN_VOICE=0` turns the hook off.
21
+ - **`speak_and_listen` with `listen: false`** speaks without opening the microphone, for one-way announcements or a goodbye. It needs only text-to-speech, not the recorder, whisper or a model.
22
+ - **One voice turn at a time across the whole Mac.** Claude Code windows, Claude Desktop and Cursor take turns at the speaker and mic, one turn at a time, so they never talk over each other. A waiting turn reports "waiting for another voice session" and gives up after `VOICE_MCP_LOCK_WAIT_SECONDS` (120) with a clear message. A lock left by a crashed session is taken over. If the cache folder isn't writable, turns go ahead without the lock.
23
+ - **README: an Upgrading section and a Voice mode section, plus a note on testing the plugin from a checkout.**
24
+
25
+ ### Changed
26
+ - **Claude waits longer for you to start talking: 15 seconds, up from 8** (`VOICE_MCP_START_TIMEOUT_SECONDS`), so reading or thinking doesn't end the conversation.
27
+ - **When you don't answer, Claude asks once more, then pauses cleanly.** The "no speech" result now tells Claude that the mic is off, so it tells you to type to carry on rather than to speak.
28
+
29
+ ### Removed
30
+ - The unused `packaging/` folder. `.github/` is the only copy of the workflows.
31
+
8
32
  ## [0.2.0] - 2026-09-28
9
33
 
10
34
  ### Added
package/README.md CHANGED
@@ -51,7 +51,7 @@ The server has **two tools and two prompts**:
51
51
  | `/mcp__voice-mcp__setup` | A guided setup: it checks, asks you, installs, then runs a spoken test. |
52
52
  | `/mcp__voice-mcp__voice_mode` | A hands-free session where Claude checks in by voice at natural points. |
53
53
 
54
- **Listening works like a conversation.** A soft chime plays when the mic opens. The server waits for you to start talking and hands back to Claude about a second after you stop. It adjusts to background noise, doesn't cut you off at pauses mid-sentence, and ignores coughs and clicks. If you say nothing for 8 seconds, Claude gets "no speech", which it is told never to treat as a yes.
54
+ **Listening works like a conversation.** A soft chime plays when the mic opens. The server waits for you to start talking and hands back to Claude about a second after you stop. It adjusts to background noise, doesn't cut you off at pauses mid-sentence, and ignores coughs and clicks. If you say nothing for 15 seconds, Claude gets "no speech", which it is told never to treat as a yes.
55
55
 
56
56
  **Replies come back fast.** whisper.cpp's server keeps the speech model loaded between turns, and the model starts loading while Claude is still talking. You don't wait for a model load on each reply. After 15 idle minutes the server shuts down to free memory. It is stopped automatically even if the MCP server crashes.
57
57
 
@@ -76,7 +76,7 @@ You need **Node.js 22 or newer**. Whichever way you install, run setup once afte
76
76
  /plugin marketplace add jeet0007/mac-voice-mcp
77
77
  /plugin install mac-voice-mcp@mac-voice-mcp
78
78
  ```
79
- The plugin adds `/mac-voice-mcp:setup` and `/mac-voice-mcp:talk`, plus a skill that teaches Claude how to use voice well and fix common problems. Each plugin version runs the matching npm release.
79
+ The plugin adds `/mac-voice-mcp:setup` and `/mac-voice-mcp:talk`, a skill that teaches Claude how to use voice well and fix common problems, and a hook that keeps a voice conversation in voice (see [Voice mode](#voice-mode)). Each plugin version runs the matching npm release.
80
80
  - **The official MCP Registry.** It's listed as [`io.github.jeet0007/mac-voice-mcp`](https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.jeet0007/mac-voice-mcp). Apps and directories that read the registry pick it up from there. In VS Code, open the Extensions view (⇧⌘X), search `@mcp mac-voice`, and click **Install**. Smithery, Glama, PulseMCP and mcp.so copy the registry, so it shows up there too.
81
81
 
82
82
  ### By hand
@@ -123,7 +123,7 @@ Setup checks what's already there before it changes anything:
123
123
  | Needed | Provided by | If it's missing |
124
124
  |---|---|---|
125
125
  | Voice | macOS `say`, using the most natural voice installed | Nothing to do. For a far better voice, add a free Premium one (see below). |
126
- | Microphone capture | SoX (`rec`) | `brew install sox` |
126
+ | Microphone capture | SoX (`rec`). ffmpeg works as a fallback, but setup recommends SoX | `brew install sox` |
127
127
  | Speech-to-text | whisper.cpp (`whisper-cli` and `whisper-server`, Metal-accelerated) | `brew install whisper-cpp` |
128
128
  | Speech model | `base.en`, ~140 MB | Downloaded once to `~/.cache/mac-voice-mcp/models/` |
129
129
 
@@ -160,6 +160,14 @@ One command does everything above: it builds the project, runs setup (asking bef
160
160
  git clone https://github.com/jeet0007/mac-voice-mcp && bash mac-voice-mcp/install.sh
161
161
  ```
162
162
 
163
+ ## Upgrading
164
+
165
+ - **Claude Code plugin:** run `/plugin marketplace update mac-voice-mcp`, then open `/plugin`, choose mac-voice-mcp under your installed plugins, and update it. Restart Claude Code. If there's no update option, uninstall and reinstall it.
166
+ - **Everything installed with `mac-voice-mcp@latest`** (Claude Desktop, Cursor, VS Code, `claude mcp add`): restart the app. npx fetches the new release when the server starts.
167
+ - **Configs without `@latest`:** change `mac-voice-mcp` to `mac-voice-mcp@latest` in the config, then restart the app. Otherwise npx keeps running the version it cached first.
168
+
169
+ Check which version you'd get with `npx -y mac-voice-mcp@latest --version`, and see what changed in the [changelog](CHANGELOG.md). Your model, voice and settings carry over.
170
+
163
171
  ## Using it
164
172
 
165
173
  - **"Work on X and check in with me by voice when you need a decision."** Claude works quietly and only speaks at decision points.
@@ -167,6 +175,18 @@ git clone https://github.com/jeet0007/mac-voice-mcp && bash mac-voice-mcp/instal
167
175
  - **"Read me a 20-second summary of this PR and ask if I should approve it."** Use this for one-off briefings.
168
176
  - **Just talk after the chime.** You don't need to hurry or fill silence. If you're still talking at the 30-second safety cap (`listen_seconds`), Claude is told your reply may be cut off and asks you to continue.
169
177
 
178
+ ### Voice mode
179
+
180
+ Once you answer out loud, you're in a voice conversation. Claude replies by voice, not in text, until one of these happens:
181
+
182
+ - **You type something.** You're back at the keyboard.
183
+ - **You say you're done.** Claude says a short goodbye without opening the mic (`speak_and_listen` with `listen: false`).
184
+ - **You don't answer.** After 15 seconds of silence Claude asks once more. If you still don't answer, it pauses and summarizes on screen, and the mic stays off. Type anything, or run `/mac-voice-mcp:talk`, to pick up again.
185
+
186
+ In Claude Code, the plugin enforces this with a hook. If Claude tries to answer in text mid-conversation, the hook sends it back once to answer by voice. If it stops again, the hook lets it. To turn the hook off, add `"VOICE_MCP_STAY_IN_VOICE": "0"` to the `env` block in `~/.claude/settings.json`. Other apps rely on the instructions alone.
187
+
188
+ **Several sessions, one mic.** Every mac-voice-mcp on your Mac takes turns: Claude Code windows, Claude Desktop and Cursor. While one is speaking or listening, the others wait for that turn to finish (shown as "waiting for another voice session"). They give up after 2 minutes with a message. If a session crashes, the next one takes the mic over.
189
+
170
190
  ## Getting Claude to sound natural
171
191
 
172
192
  Guidance reaches Claude through several channels, because each client shows different ones:
@@ -177,6 +197,7 @@ Guidance reaches Claude through several channels, because each client shows diff
177
197
  | Server instructions (when to use voice, how to handle replies, setup) | Claude Code (it reads up to 2 KB) |
178
198
  | The `voice_mode` and `setup` prompts | Claude Code (as slash commands), Claude Desktop, Cursor |
179
199
  | Server-side rewrite plus a `voice-mcp note` back to Claude | Always on |
200
+ | The plugin's `/mac-voice-mcp:talk`, `voice-help` skill and stay-in-voice hook | Claude Code, with the plugin |
180
201
 
181
202
  The rules Claude is given:
182
203
 
@@ -204,16 +225,18 @@ Everything is optional. Set these in your client config's `"env": { … }` block
204
225
  | `VOICE_MCP_RATE` | system rate | Words per minute, e.g. `200`. |
205
226
  | `VOICE_MCP_MAX_SPEAK_WORDS` | `120` | Longer text is cut at a sentence boundary ("the rest is on screen"). |
206
227
  | `VOICE_MCP_CHIME` | `1` | Set to `0` to turn off the mic open/close sounds. |
228
+ | `VOICE_MCP_LOCK_WAIT_SECONDS` | `120` | How long a turn waits while another session on this Mac is using the mic. |
207
229
 
208
230
  **Listening**
209
231
 
210
232
  | Variable | Default | |
211
233
  |---|---|---|
212
234
  | `VOICE_MCP_END_SILENCE_MS` | `1200` | How long a pause ends your turn. Use `1800` if it cuts you off while you think, `800` for snappier replies. |
213
- | `VOICE_MCP_START_TIMEOUT_SECONDS` | `8` | How long to wait for you to start talking. |
235
+ | `VOICE_MCP_START_TIMEOUT_SECONDS` | `15` | How long to wait for you to start talking. |
214
236
  | `VOICE_MCP_SPEECH_MARGIN_DB` | `12` | How much louder than room noise counts as speech. Raise it in noisy rooms. |
215
237
  | `VOICE_MCP_MIN_SPEECH_DB` | `-48` | The quietest level that ever counts as speech (dBFS). |
216
- | `VOICE_MCP_RECORDER` | `auto` | `sox` or `ffmpeg` (`ffmpeg` is macOS only). |
238
+ | `VOICE_MCP_RECORDER` | `auto` | `sox` or `ffmpeg` (`ffmpeg` is macOS only). `auto` uses SoX, falling back to ffmpeg. |
239
+ | `VOICE_MCP_FFMPEG_DEVICE` | `:default` | Which input ffmpeg records from. `:default` follows System Settings; `:1` picks device 1 (list them with `ffmpeg -f avfoundation -list_devices true -i ""`). |
217
240
 
218
241
  **Speech-to-text**
219
242
 
@@ -230,6 +253,12 @@ Everything is optional. Set these in your client config's `"env": { … }` block
230
253
  | `VOICE_MCP_CACHE_DIR` | `~/.cache/mac-voice-mcp` | Where models are downloaded or symlinked. |
231
254
  | `VOICE_MCP_DEBUG` | `0` | Verbose logs with per-turn timings, written to stderr. |
232
255
 
256
+ **Claude Code plugin hook.** Set this in the `env` block of `~/.claude/settings.json`, not in the server's config:
257
+
258
+ | Variable | Default | |
259
+ |---|---|---|
260
+ | `VOICE_MCP_STAY_IN_VOICE` | `1` | Set to `0` to stop the plugin's hook from sending Claude back to answer by voice. |
261
+
233
262
  **Models** (whisper.cpp names):
234
263
 
235
264
  | Model | Size | Good for |
@@ -244,7 +273,7 @@ Everything is optional. Set these in your client config's `"env": { … }` block
244
273
  | Symptom | Fix |
245
274
  |---|---|
246
275
  | "voice-mcp is not set up yet" | Ask Claude to *set up voice*, or run `npx -y mac-voice-mcp@latest setup`. |
247
- | "microphone returned pure digital silence" | macOS is blocking the mic for the host app. Go to **System Settings → Privacy & Security → Microphone**, enable Claude / Cursor / your terminal, then restart that app. |
276
+ | "microphone returned pure digital silence" | macOS is blocking the mic for the host app. Go to **System Settings → Privacy & Security → Microphone**, enable Claude / Cursor / your terminal, then restart that app. If the message says it's recording with ffmpeg, the input device is the likelier cause: run `brew install sox`. |
248
277
  | No permission prompt ever appears | Run `tccutil reset Microphone <bundle id>` and restart the app. Running `test` in Terminal only gives permission to Terminal, not to Claude Desktop. |
249
278
  | It cuts me off while I'm thinking | Set `VOICE_MCP_END_SILENCE_MS=1800` (or up to `2500`). |
250
279
  | It never stops listening | The room is too noisy for the defaults. Set `VOICE_MCP_SPEECH_MARGIN_DB=18`, or use a headset. |
@@ -253,6 +282,8 @@ Everything is optional. Set these in your client config's `"env": { … }` block
253
282
  | It garbles names or jargon | Set `VOICE_MCP_WHISPER_PROMPT="Priya, Postgres, Kubernetes"`, or switch to `small.en`. |
254
283
  | The voice sounds robotic | Download a Premium voice (see [Get a better voice](#then-set-up-and-allow-the-mic)). It's used automatically. |
255
284
  | It asks for approval every turn | Add the tool to Claude Code's allow list: see [Allow voice turns without prompts](#allow-voice-turns-without-prompts). |
285
+ | "Another voice session on this Mac…" | Another Claude window, Claude Desktop or Cursor held the speaker and mic for over 2 minutes, which means one very long turn. End that conversation, then try again. |
286
+ | Claude keeps answering by voice after I'm done | Type anything, or say "stop voice mode". To switch the plugin's hook off entirely, see [Voice mode](#voice-mode). |
256
287
  | Turns feel slow | Each result ends with a timing line, e.g. `spoke 3.1 s · listened 4.0 s · transcribed 0.3 s`. Ask Claude what it says. Transcribing should take well under a second. |
257
288
 
258
289
  ## Known limitations
@@ -260,7 +291,7 @@ Everything is optional. Set these in your client config's `"env": { … }` block
260
291
  - **You can't interrupt it.** It finishes speaking, then listens. Barge-in would mean listening while the speakers play, which needs headphones or echo cancellation.
261
292
  - **It's macOS-first.** Linux works with SoX and espeak-ng. Windows is untested.
262
293
  - **Turn-taking is based on loudness, not a speech model.** It adapts to background noise, but very noisy rooms, music or TV can confuse it. A headset helps, and so do the listening settings above.
263
- - **One conversation at a time.** There's one speaker and one microphone, so calls are queued.
294
+ - **One conversation at a time.** There's one speaker and one microphone, so turns from every session on the Mac are queued.
264
295
 
265
296
  ## Privacy and safety
266
297
 
@@ -297,11 +328,14 @@ npm run inspect # MCP Inspector
297
328
  | `model.ts` | Finding, symlinking or downloading the model |
298
329
  | `stt.ts` | The warm `whisper-server` with orphan guard, and the `whisper-cli` fallback |
299
330
  | `setup.ts` | Requirement checks and consent-based background installs |
331
+ | `lock.ts` | One voice turn at a time across every session on the Mac |
300
332
  | `voice.ts`, `server.ts`, `index.ts` | The round trip, the MCP tools and prompts, and the CLI |
301
- | `skills/` | The Claude Code plugin's `/mac-voice-mcp:talk` and `:setup` commands and the `voice-help` skill |
333
+ | `skills/`, `hooks/` | The Claude Code plugin's `/mac-voice-mcp:talk` and `:setup` commands, the `voice-help` skill, and the stay-in-voice hook |
302
334
 
303
335
  The package installs two commands: `mac-voice-mcp` (the one `npx -y mac-voice-mcp` runs), and `voice-mcp`.
304
336
 
337
+ **Testing the plugin: don't start Claude Code inside this repo.** In this folder, `npx mac-voice-mcp@<this version>` finds the checkout itself instead of downloading the package, can't run it, and the server fails with `CONNECTION_CLOSED`. Start `claude` in any other folder. To try unreleased plugin files (skills, hooks), swap the marketplace to your checkout: run `/plugin marketplace remove mac-voice-mcp`, then `/plugin marketplace add /path/to/checkout`, and install at user scope. The server still comes from npm, so test server changes with `npm test` and `npm run test:voice`.
338
+
305
339
  ### Releasing
306
340
 
307
341
  - **First release:** `bash publish.sh`. It asks before each public step and uses your own GitHub and npm logins. It creates the GitHub repo, publishes to npm, and lists the server in the [official MCP Registry](https://registry.modelcontextprotocol.io), which Smithery, Glama, PulseMCP and mcp.so pick up from.
package/dist/config.js CHANGED
@@ -31,6 +31,10 @@ const modelName = (process.env.VOICE_MCP_WHISPER_MODEL ?? "base.en").trim();
31
31
  const CACHE_ROOT = process.env.VOICE_MCP_CACHE_DIR?.trim() ||
32
32
  path.join(process.env.XDG_CACHE_HOME || path.join(os.homedir(), ".cache"), "mac-voice-mcp");
33
33
  export const CONFIG = {
34
+ /** Models, the mic lock and hook state live here. */
35
+ cacheDir: CACHE_ROOT,
36
+ /** How long a voice turn waits for another session to finish with the mic. */
37
+ lockWaitSeconds: Math.max(1, envNum("VOICE_MCP_LOCK_WAIT_SECONDS", 120)),
34
38
  // --- Speaking
35
39
  /**
36
40
  * macOS voice name, e.g. "Ava (Premium)" (`say -v '?'` lists them). Unset: the most natural
@@ -47,15 +51,15 @@ export const CONFIG = {
47
51
  /** End of turn: stop listening after this much silence once the user has spoken. */
48
52
  endSilenceMs: Math.max(300, envNum("VOICE_MCP_END_SILENCE_MS", 1200)),
49
53
  /** Give up if the user hasn't started talking within this many seconds. */
50
- startTimeoutSeconds: Math.max(1, envNum("VOICE_MCP_START_TIMEOUT_SECONDS", 8)),
54
+ startTimeoutSeconds: Math.max(1, envNum("VOICE_MCP_START_TIMEOUT_SECONDS", 15)),
51
55
  /** Speech must be this many dB above the room's background noise. Lower = more sensitive. */
52
56
  speechMarginDb: envNum("VOICE_MCP_SPEECH_MARGIN_DB", 12),
53
57
  /** …and never quieter than this absolute level (dBFS). */
54
58
  minSpeechDb: envNum("VOICE_MCP_MIN_SPEECH_DB", -48),
55
59
  /** Recorder: auto (SoX, else ffmpeg) | sox | ffmpeg. */
56
60
  recorder: (process.env.VOICE_MCP_RECORDER?.trim().toLowerCase() || "auto"),
57
- /** ffmpeg avfoundation audio input (macOS fallback recorder). */
58
- ffmpegDevice: process.env.VOICE_MCP_FFMPEG_DEVICE?.trim() || ":0",
61
+ /** ffmpeg avfoundation audio input (macOS fallback recorder). ":default" follows the input chosen in System Settings. */
62
+ ffmpegDevice: process.env.VOICE_MCP_FFMPEG_DEVICE?.trim() || ":default",
59
63
  // --- Speech-to-text
60
64
  /** whisper.cpp model name (tiny.en, base.en, small.en, large-v3-turbo-q5_0, ...). */
61
65
  modelName,
package/dist/index.js CHANGED
@@ -84,7 +84,10 @@ async function runSetupCli(args) {
84
84
  async function runTestCli(text) {
85
85
  const prompt = text || "Voice bridge test. Say something after the chime, and I'll print what I heard.";
86
86
  try {
87
- const result = await speakAndListen(prompt, envNum("VOICE_MCP_TEST_SECONDS", DEFAULT_LISTEN_SECONDS), undefined, (p) => console.error(` … ${p}`));
87
+ const result = await speakAndListen(prompt, {
88
+ listenSeconds: envNum("VOICE_MCP_TEST_SECONDS", DEFAULT_LISTEN_SECONDS),
89
+ onPhase: (p) => console.error(` … ${p}`),
90
+ });
88
91
  process.stdout.write(result.text + "\n");
89
92
  for (const note of result.notes)
90
93
  console.error(note);
package/dist/lock.js ADDED
@@ -0,0 +1,162 @@
1
+ /**
2
+ * One voice turn at a time across every voice-mcp on this Mac — several Claude Code windows,
3
+ * Claude Desktop and Cursor all share one speaker and one microphone.
4
+ *
5
+ * The lock is a file created atomically (O_EXCL) holding the owner's pid and start time.
6
+ * A lock whose owner has died (or that is implausibly old) is taken over, so a crash never
7
+ * blocks the mic for good. If the lock can't be used at all (read-only or missing cache
8
+ * folder), turns go ahead without it rather than failing.
9
+ */
10
+ import { closeSync, linkSync, mkdirSync, openSync, readFileSync, renameSync, statSync, unlinkSync, writeSync } from "node:fs";
11
+ import path from "node:path";
12
+ import { CONFIG, debug } from "./config.js";
13
+ import { CancelledError, sleep } from "./proc.js";
14
+ /** Longer than any real turn (speaking, up to two minutes of listening, transcribing). */
15
+ const STALE_AFTER_MS = 10 * 60_000;
16
+ /** A lock file that's still unreadable after this long was never going to be written. */
17
+ const UNREADABLE_GRACE_MS = 5_000;
18
+ const POLL_MS = 250;
19
+ export class MicBusyError extends Error {
20
+ name = "MicBusyError";
21
+ }
22
+ export const lockFile = () => path.join(CONFIG.cacheDir, "mic.lock");
23
+ function readOwner(file) {
24
+ try {
25
+ const o = JSON.parse(readFileSync(file, "utf8"));
26
+ return Number.isInteger(o.pid) && Number.isFinite(o.since) ? o : null;
27
+ }
28
+ catch {
29
+ return null;
30
+ }
31
+ }
32
+ function isAlive(pid) {
33
+ try {
34
+ process.kill(pid, 0);
35
+ return true;
36
+ }
37
+ catch (e) {
38
+ return e.code === "EPERM"; // exists, owned by someone else
39
+ }
40
+ }
41
+ /** The lock file's identity (inode) if it's stale, otherwise null. */
42
+ function staleIdentity(file, now) {
43
+ let ino;
44
+ let mtimeMs;
45
+ try {
46
+ ({ ino, mtimeMs } = statSync(file));
47
+ }
48
+ catch {
49
+ return null; // gone already; the next attempt will create it
50
+ }
51
+ const owner = readOwner(file);
52
+ const stale = owner ? !isAlive(owner.pid) || now - owner.since > STALE_AFTER_MS : now - mtimeMs > UNREADABLE_GRACE_MS;
53
+ return stale ? ino : null;
54
+ }
55
+ /**
56
+ * Remove a stale lock without ever removing a fresh one someone else just created: move it aside
57
+ * atomically, check it's the same file we judged stale, and put it back if it isn't.
58
+ */
59
+ function takeOver(file, ino) {
60
+ const aside = `${file}.${process.pid}.${Date.now()}.stale`;
61
+ try {
62
+ renameSync(file, aside);
63
+ }
64
+ catch {
65
+ return; // another session got there first
66
+ }
67
+ try {
68
+ if (statSync(aside).ino !== ino) {
69
+ try {
70
+ linkSync(aside, file); // it was a live lock after all: restore it (fails harmlessly if one exists)
71
+ }
72
+ catch {
73
+ /* a newer lock is already in place */
74
+ }
75
+ }
76
+ }
77
+ finally {
78
+ try {
79
+ unlinkSync(aside);
80
+ }
81
+ catch {
82
+ /* ignore */
83
+ }
84
+ }
85
+ }
86
+ /** Remove the file only if it's still ours. */
87
+ function releaseIfOurs(file, me) {
88
+ const now = readOwner(file);
89
+ if (!now || now.pid !== me.pid || now.since !== me.since)
90
+ return;
91
+ try {
92
+ unlinkSync(file);
93
+ }
94
+ catch {
95
+ /* already gone */
96
+ }
97
+ }
98
+ const UNUSABLE = new Set(["EACCES", "EPERM", "EROFS", "ENOENT", "ENOTDIR"]);
99
+ const noLock = () => { };
100
+ /** Wait for the mic, take it, and return the function that gives it back. */
101
+ export async function acquireMicLock(opts = {}) {
102
+ const file = lockFile();
103
+ const waitMs = opts.waitMs ?? CONFIG.lockWaitSeconds * 1000;
104
+ const deadline = Date.now() + waitMs;
105
+ let waiting = false;
106
+ try {
107
+ mkdirSync(path.dirname(file), { recursive: true });
108
+ }
109
+ catch (e) {
110
+ debug("mic lock unavailable, continuing without it:", e.message);
111
+ return noLock;
112
+ }
113
+ for (;;) {
114
+ if (opts.signal?.aborted)
115
+ throw new CancelledError();
116
+ const me = { pid: process.pid, since: Date.now() };
117
+ try {
118
+ const fd = openSync(file, "wx");
119
+ try {
120
+ writeSync(fd, JSON.stringify(me));
121
+ }
122
+ finally {
123
+ closeSync(fd);
124
+ }
125
+ let released = false;
126
+ const release = () => {
127
+ if (released)
128
+ return;
129
+ released = true;
130
+ process.off("exit", release);
131
+ releaseIfOurs(file, me);
132
+ };
133
+ process.on("exit", release); // a normal shutdown mid-turn still frees the mic
134
+ return release;
135
+ }
136
+ catch (e) {
137
+ const code = e.code ?? "";
138
+ if (UNUSABLE.has(code)) {
139
+ debug("mic lock unavailable, continuing without it:", e.message);
140
+ return noLock;
141
+ }
142
+ if (code !== "EEXIST")
143
+ throw e;
144
+ }
145
+ const ino = staleIdentity(file, Date.now());
146
+ if (ino !== null) {
147
+ debug("taking over a stale mic lock");
148
+ takeOver(file, ino);
149
+ continue;
150
+ }
151
+ if (!waiting) {
152
+ waiting = true;
153
+ opts.onWait?.();
154
+ }
155
+ if (Date.now() > deadline) {
156
+ throw new MicBusyError("Another voice session on this Mac (another Claude window, Claude Desktop or Cursor) has been using the " +
157
+ `speaker and microphone for over ${Math.round(waitMs / 1000)} seconds. ` +
158
+ "Tell the user on screen, and try again once its current turn is over.");
159
+ }
160
+ await sleep(POLL_MS);
161
+ }
162
+ }
package/dist/server.js CHANGED
@@ -37,12 +37,23 @@ export function createServer() {
37
37
  .optional()
38
38
  .describe(`Upper limit on how long to listen, in seconds (default ${DEFAULT_LISTEN_SECONDS}, max ${MAX_LISTEN_SECONDS}). ` +
39
39
  "Listening already stops when the user finishes talking, so you rarely need this."),
40
+ listen: z
41
+ .boolean()
42
+ .optional()
43
+ .describe("Default true. false: only speak, without opening the microphone — for a one-way announcement, " +
44
+ "or to say goodbye when the user ends voice mode."),
40
45
  },
41
46
  annotations: { readOnlyHint: false, destructiveHint: false, idempotentHint: false, openWorldHint: false },
42
- }, async ({ text_to_speak, listen_seconds }, extra) => {
47
+ }, async ({ text_to_speak, listen_seconds, listen }, extra) => {
43
48
  const progress = progressReporter(extra);
49
+ const phaseMessage = (phase) => phase === "waiting" ? "waiting for another voice session on this Mac to finish" : phase;
44
50
  try {
45
- const result = await exclusive(() => speakAndListen(text_to_speak, listen_seconds ?? DEFAULT_LISTEN_SECONDS, extra.signal, (phase) => progress?.(phase)));
51
+ const result = await exclusive(() => speakAndListen(text_to_speak, {
52
+ listenSeconds: listen_seconds ?? DEFAULT_LISTEN_SECONDS,
53
+ listen,
54
+ signal: extra.signal,
55
+ onPhase: (phase) => progress?.(phaseMessage(phase)),
56
+ }));
46
57
  return {
47
58
  content: [{ type: "text", text: result.text }, ...result.notes.map((note) => ({ type: "text", text: `\n\n${note}` }))],
48
59
  isError: !result.ok,
package/dist/setup.js CHANGED
@@ -76,9 +76,21 @@ async function checkRequirements() {
76
76
  : { label: "Voice", status: "ok", detail: choice.label });
77
77
  }
78
78
  const rec = findRecorder();
79
- checks.push(rec
80
- ? { label: "Recorder", status: "ok", detail: describeRecorder(rec) }
81
- : { label: "Recorder", status: brewing ? "installing" : "missing", detail: "SoX is not installed", brew: IS_WIN ? undefined : "sox" });
79
+ if (!rec) {
80
+ checks.push({ label: "Recorder", status: brewing ? "installing" : "missing", detail: "SoX is not installed", brew: IS_WIN ? undefined : "sox" });
81
+ }
82
+ else if (rec.kind === "ffmpeg" && CONFIG.recorder === "auto") {
83
+ // Works, but it's the fragile path: recommend SoX, which is what the turn-taking is tuned on.
84
+ checks.push({
85
+ label: "Recorder",
86
+ status: brewing ? "installing" : "optional",
87
+ detail: `${describeRecorder(rec)}, as a fallback. SoX is recommended: it's more reliable with Bluetooth headsets and other audio devices`,
88
+ brew: "sox",
89
+ });
90
+ }
91
+ else {
92
+ checks.push({ label: "Recorder", status: "ok", detail: describeRecorder(rec) });
93
+ }
82
94
  const cli = findWhisperCli();
83
95
  const server = findWhisperServer();
84
96
  if (cli || server) {
@@ -131,9 +143,11 @@ async function checkRequirements() {
131
143
  * @param onProgress optional heartbeat while waiting (used for MCP progress notifications).
132
144
  */
133
145
  export async function runSetupFlow(install, onProgress) {
146
+ resetWhichCache(); // see anything installed since the last check (e.g. `brew install sox` in a terminal)
134
147
  let checks = await checkRequirements();
135
148
  if (install) {
136
- const formulae = [...new Set(checks.filter((c) => c.status === "missing" && c.brew).map((c) => c.brew))];
149
+ // Missing pieces, plus recommended ones (SoX when only ffmpeg is there): the user agreed to install.
150
+ const formulae = [...new Set(checks.filter((c) => (c.status === "missing" || c.status === "optional") && c.brew).map((c) => c.brew))];
137
151
  const brew = findBrew();
138
152
  if (formulae.length && brew && !running(brewJob))
139
153
  brewJob = startBrewInstall(brew, formulae);
@@ -157,7 +171,7 @@ export async function runSetupFlow(install, onProgress) {
157
171
  const lines = [
158
172
  `voice-mcp setup — ${ready ? "READY" : installing ? "INSTALLING" : "NOT READY"}`,
159
173
  "",
160
- ...checks.map((c) => `${icon[c.status]} ${c.label}: ${c.detail}${c.brew && c.status === "missing" ? ` → brew install ${c.brew}` : ""}`),
174
+ ...checks.map((c) => `${icon[c.status]} ${c.label}: ${c.detail}${c.brew && c.status !== "ok" && c.status !== "installing" ? ` → brew install ${c.brew}` : ""}`),
161
175
  IS_MAC ? "• Microphone: macOS asks for permission the first time speak_and_listen listens — click Allow." : "",
162
176
  ];
163
177
  if (done.length)
@@ -169,6 +183,11 @@ export async function runSetupFlow(install, onProgress) {
169
183
  lines.push("");
170
184
  if (ready) {
171
185
  lines.push("Next: everything is in place — speak_and_listen is ready to use.");
186
+ const recommended = [...new Set(checks.filter((c) => c.status === "optional" && c.brew).map((c) => c.brew))];
187
+ if (recommended.length && !install && findBrew()) {
188
+ lines.push(`Recommended: ask the user whether to brew install ${recommended.join(" ")} (see • above). ` +
189
+ "Only if they agree, call voice_setup with install=true.");
190
+ }
172
191
  if (checks.some((c) => c.label === "Voice" && c.status === "optional")) {
173
192
  lines.push("Mention the optional voice tip (•) to the user once — it makes the voice sound far more natural. Nothing else to do.");
174
193
  }
package/dist/texts.js CHANGED
@@ -30,8 +30,8 @@ export const SERVER_INSTRUCTIONS = [
30
30
  '- "(No speech detected …)" means no answer: never treat silence as consent. Ask once more or',
31
31
  " continue with safe work and report on screen.",
32
32
  "- Once the user is talking with you by voice, stay in voice: answer every turn with speak_and_listen,",
33
- " not a text reply, until they say stop or start typing.",
34
- "- Listening ends on its own when the user stops talking; no need to set listen_seconds.",
33
+ " not a text reply, until they say stop (say goodbye with listen: false) or start typing.",
34
+ "- Listening ends when the user stops talking; listen_seconds is rarely needed.",
35
35
  '- If the result includes a "voice-mcp note", follow it.',
36
36
  "",
37
37
  "Setup: if speak_and_listen says voice-mcp is not set up, call voice_setup (check only), tell the user",
@@ -49,7 +49,9 @@ export const SPEAK_TOOL_DESCRIPTION = [
49
49
  "",
50
50
  "Once the user is talking with you by voice, keep the conversation in voice: answer each transcript with",
51
51
  "another speak_and_listen call (not a text reply) until they say stop or start typing.",
52
- 'A reply of "(No speech detected …)" means the user did not answer — never treat it as consent.',
52
+ "To end voice mode, or for a one-way announcement, pass listen: false: it speaks without opening the mic.",
53
+ 'A reply of "(No speech detected …)" means the user did not answer — never treat it as consent. Ask once more;',
54
+ "if still nothing, stop and say on screen that voice mode is paused and they can type to carry on (not speak: the mic is off).",
53
55
  'The "voice-mcp timing" line at the end is diagnostics: ignore it unless the user asks why things feel slow.',
54
56
  "If it reports that voice-mcp is not set up, call voice_setup.",
55
57
  ].join("\n");
@@ -81,9 +83,11 @@ export const VOICE_MODE_PROMPT = (task) => [
81
83
  " Don't narrate every small action.",
82
84
  "- If a transcript is unclear, ask again by voice.",
83
85
  "- Confirm by voice before anything destructive, irreversible, or that costs money.",
84
- "- If I don't answer, don't assume yes: carry on with safe work or pause, and summarize on screen.",
86
+ "- If I don't answer, don't assume yes: ask once more out loud. If there's still nothing, pause and summarize",
87
+ " on screen, and tell me to type anything to pick up again (the mic is off, so don't tell me to speak).",
85
88
  "- Keep writing full details (code, diffs, links) on screen as usual; the voice line is the headline.",
86
- '- Stop using voice when I say "stop voice mode", "I\'m back", or start typing again.',
89
+ '- Stop using voice when I say "stop voice mode", "I\'m back", or start typing again. To end it, say a short',
90
+ " goodbye with speak_and_listen and listen: false (it speaks without opening the mic).",
87
91
  "",
88
92
  SPEECH_RULES,
89
93
  "",
package/dist/voice.js CHANGED
@@ -4,8 +4,9 @@ import os from "node:os";
4
4
  import path from "node:path";
5
5
  import { chime, findRecorder, listenForTurn, MIC_PERMISSION_HINT, RECORDER_MISSING, speak } from "./audio.js";
6
6
  import { CONFIG, debug, DEFAULT_LISTEN_SECONDS, MAX_LISTEN_SECONDS, MAX_SPEAK_CHARS } from "./config.js";
7
+ import { acquireMicLock, MicBusyError } from "./lock.js";
7
8
  import { ensureModel } from "./model.js";
8
- import { SetupError } from "./proc.js";
9
+ import { resetWhichCache, SetupError } from "./proc.js";
9
10
  import { prepareSpeech } from "./speech-text.js";
10
11
  import { findWhisperCli, findWhisperServer, prewarm, STT_MISSING, transcribe } from "./stt.js";
11
12
  /** Everything speak_and_listen needs, checked before saying a word. Never installs anything. */
@@ -16,11 +17,33 @@ async function preflight() {
16
17
  throw new SetupError(STT_MISSING);
17
18
  return { model: await ensureModel({ download: false }) };
18
19
  }
19
- export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase) {
20
- const seconds = Math.min(MAX_LISTEN_SECONDS, Math.max(1, Number.isFinite(listenSeconds) ? listenSeconds : DEFAULT_LISTEN_SECONDS));
21
- const { model } = await preflight();
22
- // Load the model into the warm server while we talk, so it's ready when the user finishes.
23
- prewarm(model);
20
+ export async function speakAndListen(textToSpeak, opts = {}) {
21
+ const { signal, onPhase } = opts;
22
+ const listen = opts.listen !== false;
23
+ const requested = opts.listenSeconds ?? DEFAULT_LISTEN_SECONDS;
24
+ const seconds = Math.min(MAX_LISTEN_SECONDS, Math.max(1, Number.isFinite(requested) ? requested : DEFAULT_LISTEN_SECONDS));
25
+ const model = listen ? (await preflight()).model : null;
26
+ // Load the model into the warm server now, so it's ready when the user finishes (even if we wait below).
27
+ if (model)
28
+ prewarm(model);
29
+ // Wait for any other voice session on this Mac to finish with the speaker and mic.
30
+ let release;
31
+ try {
32
+ release = await acquireMicLock({ signal, onWait: () => onPhase?.("waiting") });
33
+ }
34
+ catch (err) {
35
+ if (err instanceof MicBusyError)
36
+ return { ok: false, text: err.message, notes: [] };
37
+ throw err;
38
+ }
39
+ try {
40
+ return await turn(textToSpeak, seconds, model, signal, onPhase);
41
+ }
42
+ finally {
43
+ release();
44
+ }
45
+ }
46
+ async function turn(textToSpeak, seconds, model, signal, onPhase) {
24
47
  const speech = prepareSpeech(textToSpeak, { maxWords: CONFIG.maxSpeakWords, maxChars: MAX_SPEAK_CHARS });
25
48
  const notes = [...speech.notes];
26
49
  const tmpDir = await mkdtemp(path.join(os.tmpdir(), "voice-mcp-"));
@@ -30,6 +53,10 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
30
53
  onPhase?.("speaking");
31
54
  await speak(speech.text, signal);
32
55
  const spoke = Date.now() - t0;
56
+ if (!model) {
57
+ const text = "(Spoken. The microphone was not opened, because listen was false.)";
58
+ return { ok: true, text, notes: [...notes, `voice-mcp timing: spoke ${secs(spoke)}`] };
59
+ }
33
60
  await chime("start");
34
61
  onPhase?.("listening");
35
62
  const tListen = Date.now();
@@ -39,14 +66,23 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
39
66
  debug("listen:", heard);
40
67
  const timing = (transcribedMs) => timingNote({ spokeMs: spoke, listenedMs: t1 - tListen, talkedSeconds: heard.speechSeconds, transcribedMs });
41
68
  if (heard.digitalSilence) {
69
+ resetWhichCache(); // if the user fixes it by installing SoX, the next turn picks it up
70
+ const viaFfmpeg = findRecorder()?.kind === "ffmpeg";
42
71
  return {
43
72
  ok: false,
44
- text: `The microphone returned pure digital silence, which usually means microphone access is blocked. ${MIC_PERMISSION_HINT}`,
73
+ text: viaFfmpeg
74
+ ? "The microphone returned pure digital silence. voice-mcp is recording with ffmpeg because SoX isn't installed, " +
75
+ "and ffmpeg may be using a silent or wrong input device (for example after connecting Bluetooth headphones). " +
76
+ "Suggest `brew install sox` (or call voice_setup), then try again. If that doesn't help: " +
77
+ MIC_PERMISSION_HINT
78
+ : `The microphone returned pure digital silence, which usually means microphone access is blocked. ${MIC_PERMISSION_HINT}`,
45
79
  notes,
46
80
  };
47
81
  }
48
82
  const waited = Math.min(CONFIG.startTimeoutSeconds, seconds);
49
- const noSpeech = `(No speech detected — the user did not reply within ${waited} seconds.)`;
83
+ const noSpeech = `(No speech detected — the user did not reply within ${waited} seconds. The microphone is now off. ` +
84
+ "Ask once more out loud. If there's still no answer, stop and say on screen that voice mode is paused " +
85
+ "and they can type anything to carry on — speaking won't work until you call speak_and_listen again.)";
50
86
  if (heard.reason === "no-speech")
51
87
  return { ok: true, text: noSpeech, notes: [...notes, timing()] };
52
88
  onPhase?.("transcribing");
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "mac-voice-mcp",
3
3
  "mcpName": "io.github.jeet0007/mac-voice-mcp",
4
- "version": "0.2.0",
4
+ "version": "0.3.1",
5
5
  "description": "Talk with Claude out loud on your Mac: speaks with macOS `say`, listens for one natural conversational turn, and transcribes on-device with whisper.cpp. An MCP server.",
6
6
  "type": "module",
7
7
  "bin": {