mac-voice-mcp 0.1.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,18 +1,60 @@
1
1
  # Changelog
2
2
 
3
- ## 0.1.1
3
+ All notable changes to this project are documented here.
4
+ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and the project uses [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
5
+
6
+ ## [Unreleased]
7
+
8
+ ## [0.2.0] - 2026-09-28
9
+
10
+ ### Added
11
+ - **A more natural voice, automatically.** On macOS, speech uses the best Premium or Enhanced voice installed for the spoken language, preferring your Mac's region, and otherwise the system voice. `voice_setup` shows which voice is in use and, if there's no Premium voice, how to download one for free in System Settings. `VOICE_MCP_VOICE=default` keeps the system voice.
12
+ - **A timing line on every turn.** Each `speak_and_listen` result ends with a line like `voice-mcp timing: spoke 3.1 s · listened 4.0 s (user talked 2.2 s) · transcribed 0.3 s`, which makes a slow turn easy to explain. The model is told to ignore it unless you ask.
13
+ - **Claude Code plugin commands and a skill:**
14
+ - `/mac-voice-mcp:talk [task]` starts a voice conversation. It pre-approves voice turns while it runs.
15
+ - `/mac-voice-mcp:setup` checks, installs with your OK, and tests.
16
+ - The `voice-help` skill gives Claude fuller usage guidance and a troubleshooting table.
17
+ - **README: "Allow voice turns without prompts".** The `permissions.allow` entries that stop Claude Code from asking for approval on every voice turn.
18
+ - **Releases:**
19
+ - `npm version patch|minor|major` now also updates `server.json` and the plugin.
20
+ - Each tag creates a GitHub Release from this changelog.
21
+ - Tests fail if any version, the plugin's pinned npm version or the changelog entry is out of sync.
22
+
23
+ ### Changed
24
+ - **The plugin runs the exact npm release it ships with** (`mac-voice-mcp@0.2.0`), instead of whatever version npx cached first.
25
+ - **The README's install configs use `mac-voice-mcp@latest`,** so npx picks up new releases. This covers Claude Desktop, Cursor, VS Code and `claude mcp add`.
26
+ - **The publish workflow is safe to re-run.**
27
+ - It skips npm and the MCP Registry when the version is already there.
28
+ - It retries the registry while npm catches up.
29
+ - It checks out without persisting credentials.
30
+
31
+ ## [0.1.1] - 2026-09-23
32
+
33
+ ### Added
34
+ - **README:** one-click install buttons for Cursor and VS Code, and install steps for the Claude Code plugin marketplace and the official MCP Registry.
35
+ - **Releases publish through npm trusted publishing, with provenance.**
4
36
 
5
37
  ### Changed
6
38
  - **Node.js 22 or newer is now required.** Node 18 and 20 no longer receive security fixes. CI tests on Node 22, 24 and 26.
7
39
 
8
40
  ### Fixed
9
- - **"It's on screen" when it isn't** ([#3](https://github.com/jeet0007/mac-voice-mcp/issues/3)). When the spoken text sends you to look at something, the model gets a reminder that only its own message text reaches your screen, not the speech and not Bash/tool output. The speaking rules now say the same. Thanks to Copilot for the first version ([#4](https://github.com/jeet0007/mac-voice-mcp/pull/4)). The trigger phrases were then narrowed so that ordinary sentences like "above thirty degrees" or "see the doctor" don't set it off.
41
+ - **"It's on screen" when it isn't** ([#3](https://github.com/jeet0007/mac-voice-mcp/issues/3)).
42
+ - When the spoken text sends you to look at something, the model is reminded that only its own message text reaches your screen, not the speech and not Bash/tool output. The speaking rules now say the same.
43
+ - Thanks to Copilot for the first version ([#4](https://github.com/jeet0007/mac-voice-mcp/pull/4)).
44
+ - The trigger phrases were narrowed so that ordinary sentences like "above thirty degrees" or "see the doctor" don't set it off.
45
+
46
+ ### Security
47
+ - TruffleHog secret scanning, CodeQL, dependency review, `npm audit` in CI, and Dependabot.
10
48
 
11
- ### Docs and tooling
12
- - README: one-click install buttons for Cursor and VS Code, and install steps for the Claude Code plugin marketplace and the official MCP Registry.
13
- - Security: TruffleHog secret scanning, CodeQL, dependency review, `npm audit` in CI, and Dependabot.
14
- - Releases publish through npm trusted publishing, with provenance.
49
+ ## [0.1.0] - 2026-09-23
15
50
 
16
- ## 0.1.0
51
+ ### Added
52
+ - First release:
53
+ - `speak_and_listen`: natural turn-taking, with on-device whisper.cpp and a warm server.
54
+ - `voice_setup`: checks first, installs only what's missing, and only after you agree.
55
+ - The `setup` and `voice_mode` prompts.
17
56
 
18
- First release: `speak_and_listen` (natural turn-taking, on-device whisper.cpp with a warm server), `voice_setup` (check first, install only what's missing, only after you agree), and the `setup` and `voice_mode` prompts.
57
+ [Unreleased]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.2.0...HEAD
58
+ [0.2.0]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.1.1...v0.2.0
59
+ [0.1.1]: https://github.com/jeet0007/mac-voice-mcp/compare/119a589...v0.1.1
60
+ [0.1.0]: https://github.com/jeet0007/mac-voice-mcp/tree/119a589
package/README.md CHANGED
@@ -64,9 +64,9 @@ You need **Node.js 22 or newer**. Whichever way you install, run setup once afte
64
64
  ### One click
65
65
 
66
66
  <p>
67
- <a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIm1hYy12b2ljZS1tY3AiXX0%3D"><img alt="Add to Cursor" src="https://cursor.com/deeplink/mcp-install-dark.svg" height="32"></a>
68
- <a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
69
- <a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%22%5D%7D&quality=insiders"><img alt="Install in VS Code Insiders" src="https://img.shields.io/badge/VS_Code_Insiders-Install_Server-24bfa5?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
67
+ <a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIm1hYy12b2ljZS1tY3BAbGF0ZXN0Il19"><img alt="Add to Cursor" src="https://cursor.com/deeplink/mcp-install-dark.svg" height="32"></a>
68
+ <a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
69
+ <a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D&quality=insiders"><img alt="Install in VS Code Insiders" src="https://img.shields.io/badge/VS_Code_Insiders-Install_Server-24bfa5?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
70
70
  </p>
71
71
 
72
72
  ### From a marketplace
@@ -76,6 +76,7 @@ You need **Node.js 22 or newer**. Whichever way you install, run setup once afte
76
76
  /plugin marketplace add jeet0007/mac-voice-mcp
77
77
  /plugin install mac-voice-mcp@mac-voice-mcp
78
78
  ```
79
+ The plugin adds `/mac-voice-mcp:setup` and `/mac-voice-mcp:talk`, plus a skill that teaches Claude how to use voice well and fix common problems. Each plugin version runs the matching npm release.
79
80
  - **The official MCP Registry.** It's listed as [`io.github.jeet0007/mac-voice-mcp`](https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.jeet0007/mac-voice-mcp). Apps and directories that read the registry pick it up from there. In VS Code, open the Extensions view (⇧⌘X), search `@mcp mac-voice`, and click **Install**. Smithery, Glama, PulseMCP and mcp.so copy the registry, so it shows up there too.
80
81
 
81
82
  ### By hand
@@ -83,7 +84,7 @@ You need **Node.js 22 or newer**. Whichever way you install, run setup once afte
83
84
  **Claude Code**
84
85
 
85
86
  ```bash
86
- claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp
87
+ claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp@latest
87
88
  ```
88
89
 
89
90
  **Claude Desktop.** Add this to `~/Library/Application Support/Claude/claude_desktop_config.json`, then quit (⌘Q) and reopen the app:
@@ -93,12 +94,14 @@ claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp
93
94
  "mcpServers": {
94
95
  "voice-mcp": {
95
96
  "command": "npx",
96
- "args": ["-y", "mac-voice-mcp"]
97
+ "args": ["-y", "mac-voice-mcp@latest"]
97
98
  }
98
99
  }
99
100
  }
100
101
  ```
101
102
 
103
+ `@latest` makes npx check for a new release each time the app starts. Without it, npx keeps running whichever version it cached first.
104
+
102
105
  If you get `spawn npx ENOENT`, use the full path from `which npx`, e.g. `"command": "/opt/homebrew/bin/npx"`.
103
106
 
104
107
  **Cursor.** Add the same `mcpServers` block to `~/.cursor/mcp.json`.
@@ -106,20 +109,20 @@ If you get `spawn npx ENOENT`, use the full path from `which npx`, e.g. `"comman
106
109
  **VS Code.** Run **MCP: Add Server** from the Command Palette, or:
107
110
 
108
111
  ```bash
109
- code --add-mcp '{"name":"voice-mcp","command":"npx","args":["-y","mac-voice-mcp"]}'
112
+ code --add-mcp '{"name":"voice-mcp","command":"npx","args":["-y","mac-voice-mcp@latest"]}'
110
113
  ```
111
114
 
112
- **Any other MCP client.** Run `npx -y mac-voice-mcp` as a stdio server.
115
+ **Any other MCP client.** Run `npx -y mac-voice-mcp@latest` as a stdio server.
113
116
 
114
117
  ### Then: set up and allow the mic
115
118
 
116
- **Run setup once.** Ask Claude to *"set up voice"*. In Claude Code you can also run `/mcp__voice-mcp__setup`, or from a terminal run `npx -y mac-voice-mcp setup`.
119
+ **Run setup once.** Ask Claude to *"set up voice"*. In Claude Code you can also run `/mac-voice-mcp:setup` (plugin) or `/mcp__voice-mcp__setup` (added by hand), or from a terminal run `npx -y mac-voice-mcp@latest setup`.
117
120
 
118
121
  Setup checks what's already there before it changes anything:
119
122
 
120
123
  | Needed | Provided by | If it's missing |
121
124
  |---|---|---|
122
- | Voice | macOS `say` | Nothing to do. It's part of macOS. |
125
+ | Voice | macOS `say`, using the most natural voice installed | Nothing to do. For a far better voice, add a free Premium one (see below). |
123
126
  | Microphone capture | SoX (`rec`) | `brew install sox` |
124
127
  | Speech-to-text | whisper.cpp (`whisper-cli` and `whisper-server`, Metal-accelerated) | `brew install whisper-cpp` |
125
128
  | Speech model | `base.en`, ~140 MB | Downloaded once to `~/.cache/mac-voice-mcp/models/` |
@@ -130,6 +133,25 @@ Setup checks what's already there before it changes anything:
130
133
 
131
134
  **Allow the microphone.** The first time Claude listens, macOS asks whether Claude (or Cursor, or your terminal) can use the microphone. Click Allow.
132
135
 
136
+ **Get a better voice (recommended).** macOS includes free Premium voices that sound far more natural than the default. Open **System Settings → Accessibility → Spoken Content → System Voice → Manage Voices…**, and download one, for example English → *Ava (Premium)* or *Zoe (Premium)*. The next voice turn uses it automatically. To choose a specific voice, or keep the system voice, see `VOICE_MCP_VOICE` under [Configuration](#configuration).
137
+
138
+ ### Allow voice turns without prompts
139
+
140
+ By default, Claude Code asks for approval every time Claude wants to speak, which breaks the flow of a conversation. To allow voice turns, add the tool to the `permissions.allow` list in `~/.claude/settings.json`. Use the name that matches how you installed it:
141
+
142
+ ```json
143
+ {
144
+ "permissions": {
145
+ "allow": [
146
+ "mcp__plugin_mac-voice-mcp_voice-mcp__speak_and_listen",
147
+ "mcp__voice-mcp__speak_and_listen"
148
+ ]
149
+ }
150
+ }
151
+ ```
152
+
153
+ The first name is for the plugin, the second for `claude mcp add voice-mcp …`. Leave `voice_setup` out, so installs still ask you first. While `/mac-voice-mcp:talk` runs, voice turns are already allowed.
154
+
133
155
  ### Installing from a clone
134
156
 
135
157
  One command does everything above: it builds the project, runs setup (asking before installing anything), adds voice-mcp to Claude Desktop (backing up your config first) and to Claude Code, and offers a spoken test. It's safe to re-run, because each step checks first and skips anything already done.
@@ -141,7 +163,7 @@ git clone https://github.com/jeet0007/mac-voice-mcp && bash mac-voice-mcp/instal
141
163
  ## Using it
142
164
 
143
165
  - **"Work on X and check in with me by voice when you need a decision."** Claude works quietly and only speaks at decision points.
144
- - **`/mcp__voice-mcp__voice_mode fix the flaky login test`.** Claude reads its plan back to you, then checks in at each checkpoint. Say "stop voice mode" or "I'm back" to end it.
166
+ - **`/mac-voice-mcp:talk fix the flaky login test`** (plugin), or **`/mcp__voice-mcp__voice_mode fix the flaky login test`** (added by hand). Claude reads its plan back to you, then checks in at each checkpoint. Say "stop voice mode" or "I'm back" to end it.
145
167
  - **"Read me a 20-second summary of this PR and ask if I should approve it."** Use this for one-off briefings.
146
168
  - **Just talk after the chime.** You don't need to hurry or fill silence. If you're still talking at the 30-second safety cap (`listen_seconds`), Claude is told your reply may be cut off and asks you to continue.
147
169
 
@@ -178,7 +200,7 @@ Everything is optional. Set these in your client config's `"env": { … }` block
178
200
 
179
201
  | Variable | Default | |
180
202
  |---|---|---|
181
- | `VOICE_MCP_VOICE` | system voice | macOS voice, e.g. `Samantha`, `Daniel`, `Kanya`. List them with `say -v '?'`. |
203
+ | `VOICE_MCP_VOICE` | most natural installed | Unset: the best Premium or Enhanced voice installed for the language, else the system voice. Set a name, e.g. `Ava (Premium)`, `Daniel`, `Kanya` (list them with `say -v '?'`), or `default` to always use the system voice. |
182
204
  | `VOICE_MCP_RATE` | system rate | Words per minute, e.g. `200`. |
183
205
  | `VOICE_MCP_MAX_SPEAK_WORDS` | `120` | Longer text is cut at a sentence boundary ("the rest is on screen"). |
184
206
  | `VOICE_MCP_CHIME` | `1` | Set to `0` to turn off the mic open/close sounds. |
@@ -221,7 +243,7 @@ Everything is optional. Set these in your client config's `"env": { … }` block
221
243
 
222
244
  | Symptom | Fix |
223
245
  |---|---|
224
- | "voice-mcp is not set up yet" | Ask Claude to *set up voice*, or run `npx -y mac-voice-mcp setup`. |
246
+ | "voice-mcp is not set up yet" | Ask Claude to *set up voice*, or run `npx -y mac-voice-mcp@latest setup`. |
225
247
  | "microphone returned pure digital silence" | macOS is blocking the mic for the host app. Go to **System Settings → Privacy & Security → Microphone**, enable Claude / Cursor / your terminal, then restart that app. |
226
248
  | No permission prompt ever appears | Run `tccutil reset Microphone <bundle id>` and restart the app. Running `test` in Terminal only gives permission to Terminal, not to Claude Desktop. |
227
249
  | It cuts me off while I'm thinking | Set `VOICE_MCP_END_SILENCE_MS=1800` (or up to `2500`). |
@@ -229,6 +251,9 @@ Everything is optional. Set these in your client config's `"env": { … }` block
229
251
  | It hears its own voice | Use headphones, or turn the speaker volume down. It only listens after it finishes speaking, but echo can linger. |
230
252
  | "Homebrew: not installed" | Install it from [brew.sh](https://brew.sh). It needs your password, so it can't run from Claude. Then run setup again. |
231
253
  | It garbles names or jargon | Set `VOICE_MCP_WHISPER_PROMPT="Priya, Postgres, Kubernetes"`, or switch to `small.en`. |
254
+ | The voice sounds robotic | Download a Premium voice (see [Get a better voice](#then-set-up-and-allow-the-mic)). It's used automatically. |
255
+ | It asks for approval every turn | Add the tool to Claude Code's allow list: see [Allow voice turns without prompts](#allow-voice-turns-without-prompts). |
256
+ | Turns feel slow | Each result ends with a timing line, e.g. `spoke 3.1 s · listened 4.0 s · transcribed 0.3 s`. Ask Claude what it says. Transcribing should take well under a second. |
232
257
 
233
258
  ## Known limitations
234
259
 
@@ -256,7 +281,7 @@ To report a vulnerability, see [SECURITY.md](SECURITY.md).
256
281
 
257
282
  ```bash
258
283
  npm install
259
- npm test # build + 27 tests: unit tests and end-to-end tests over MCP with stub binaries
284
+ npm test # build + unit tests, end-to-end tests over MCP with stub binaries, and release-metadata checks
260
285
  npm run audit # known-vulnerability and signature checks on dependencies
261
286
  npm run setup # check what's installed; offers to install what's missing
262
287
  npm run test:voice # one real speak → listen → transcribe turn
@@ -273,13 +298,19 @@ npm run inspect # MCP Inspector
273
298
  | `stt.ts` | The warm `whisper-server` with orphan guard, and the `whisper-cli` fallback |
274
299
  | `setup.ts` | Requirement checks and consent-based background installs |
275
300
  | `voice.ts`, `server.ts`, `index.ts` | The round trip, the MCP tools and prompts, and the CLI |
301
+ | `skills/` | The Claude Code plugin's `/mac-voice-mcp:talk` and `:setup` commands and the `voice-help` skill |
276
302
 
277
303
  The package installs two commands: `mac-voice-mcp` (the one `npx -y mac-voice-mcp` runs), and `voice-mcp`.
278
304
 
279
305
  ### Releasing
280
306
 
281
307
  - **First release:** `bash publish.sh`. It asks before each public step and uses your own GitHub and npm logins. It creates the GitHub repo, publishes to npm, and lists the server in the [official MCP Registry](https://registry.modelcontextprotocol.io), which Smithery, Glama, PulseMCP and mcp.so pick up from.
282
- - **Later releases:** bump `version` in `package.json` and `server.json`, then push a `v<version>` tag. The Publish workflow tests the build, checks that the tag matches both versions, and publishes to npm and the MCP Registry. It needs no secrets: both logins use GitHub's OIDC identity, once you've set the package's *Trusted Publisher* on npmjs.com (`publish.sh` prints the steps).
308
+ - **Later releases:**
309
+ 1. Add a section to `CHANGELOG.md` for the new version, in [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) format, dated today.
310
+ 2. Run `npm version patch` (bug fixes), `npm version minor` (new features) or `npm version major` (breaking changes), per [semver](https://semver.org). It updates `package.json`, `package-lock.json`, `server.json` and the plugin (including the npm version the plugin runs), commits, and tags `v<version>`.
311
+ 3. Run `git push --follow-tags`.
312
+
313
+ The Publish workflow then runs the tests, which also check that every version and the changelog entry match. It publishes to npm with provenance, lists the release in the MCP Registry, and creates a GitHub Release from the changelog section. It needs no secrets: npm and the registry both use GitHub's OIDC identity, once you've set the package's *Trusted Publisher* on npmjs.com (`publish.sh` prints the steps). If a step fails, fix the cause and use **Re-run failed jobs**. Steps that already finished are skipped.
283
314
 
284
315
  ## License
285
316
 
package/dist/audio.js CHANGED
@@ -15,6 +15,62 @@ export function findTts() {
15
15
  return "powershell.exe";
16
16
  return which("espeak-ng") ?? which("espeak") ?? which("spd-say");
17
17
  }
18
+ /** Parse `say -v '?'` lines such as "Ava (Premium) en_US # Hello! My name is Ava." */
19
+ export function parseSayVoices(output) {
20
+ const voices = [];
21
+ for (const line of output.split("\n")) {
22
+ const m = /^(.+?)\s+([a-z]{2,3}_[A-Za-z0-9]+)\s+#/.exec(line.trimEnd());
23
+ if (!m)
24
+ continue;
25
+ const name = m[1].trim();
26
+ const quality = /\(Premium\)/i.test(name) ? 3 : /\(Enhanced\)/i.test(name) ? 2 : 1;
27
+ voices.push({ name, locale: m[2], quality });
28
+ }
29
+ return voices;
30
+ }
31
+ /**
32
+ * The most natural installed voice for the language: Premium first, then Enhanced, preferring
33
+ * the system's region (en_US over en_GB on a US Mac). null = none installed; keep the system voice.
34
+ */
35
+ export function pickVoice(voices, language, systemLocale) {
36
+ const lang = (language && language !== "auto" ? language : systemLocale).toLowerCase().split(/[-_]/)[0] ?? "en";
37
+ const region = (systemLocale.split(/[-_]/)[1] ?? "").toUpperCase();
38
+ const inRegion = (v) => (region && v.locale.toUpperCase().endsWith(`_${region}`) ? 1 : 0);
39
+ const candidates = voices.filter((v) => v.quality >= 2 && v.locale.toLowerCase().startsWith(`${lang}_`));
40
+ candidates.sort((a, b) => b.quality - a.quality || inRegion(b) - inRegion(a) || a.name.localeCompare(b.name));
41
+ return candidates[0] ?? null;
42
+ }
43
+ export const VOICE_UPGRADE_HINT = "for a much more natural voice, download a Premium one: System Settings → Accessibility → Spoken Content → " +
44
+ "System Voice → Manage Voices… (for example English → Ava (Premium) or Zoe (Premium)). It's used automatically — no restart needed.";
45
+ let voiceChoice;
46
+ /** Forget the choice, so a voice downloaded meanwhile is picked up (voice_setup calls this). */
47
+ export function resetVoiceChoice() {
48
+ voiceChoice = undefined;
49
+ }
50
+ export function chooseVoice() {
51
+ voiceChoice ??= (async () => {
52
+ if (CONFIG.voice) {
53
+ return CONFIG.voice.toLowerCase() === "default"
54
+ ? { label: "the system voice (VOICE_MCP_VOICE=default)", canUpgrade: false }
55
+ : { voice: CONFIG.voice, label: `${CONFIG.voice} (set by VOICE_MCP_VOICE)`, canUpgrade: false };
56
+ }
57
+ const bin = IS_MAC ? findTts() : null;
58
+ if (!bin)
59
+ return { label: "the system voice", canUpgrade: false };
60
+ try {
61
+ const r = await run(bin, ["-v", "?"], { timeoutMs: 10_000 });
62
+ const locale = Intl.DateTimeFormat().resolvedOptions().locale;
63
+ const best = pickVoice(parseSayVoices(r.stdout), CONFIG.language, locale);
64
+ if (best)
65
+ return { voice: best.name, label: `${best.name} — the most natural voice installed`, canUpgrade: false };
66
+ }
67
+ catch {
68
+ /* fall back to the system voice */
69
+ }
70
+ return { label: "the system voice", canUpgrade: true };
71
+ })();
72
+ return voiceChoice;
73
+ }
18
74
  export async function speak(text, signal) {
19
75
  if (!text)
20
76
  return;
@@ -24,8 +80,9 @@ export async function speak(text, signal) {
24
80
  let result;
25
81
  if (IS_MAC) {
26
82
  const args = [];
27
- if (CONFIG.voice)
28
- args.push("-v", CONFIG.voice);
83
+ const { voice } = await chooseVoice();
84
+ if (voice)
85
+ args.push("-v", voice);
29
86
  if (CONFIG.rate)
30
87
  args.push("-r", CONFIG.rate);
31
88
  args.push("-f", "-"); // read from stdin: no argv length limits, no flag injection
package/dist/config.js CHANGED
@@ -32,7 +32,10 @@ const CACHE_ROOT = process.env.VOICE_MCP_CACHE_DIR?.trim() ||
32
32
  path.join(process.env.XDG_CACHE_HOME || path.join(os.homedir(), ".cache"), "mac-voice-mcp");
33
33
  export const CONFIG = {
34
34
  // --- Speaking
35
- /** macOS voice name, e.g. "Samantha" (`say -v '?'` lists them). */
35
+ /**
36
+ * macOS voice name, e.g. "Ava (Premium)" (`say -v '?'` lists them). Unset: the most natural
37
+ * installed voice for the language (Premium, then Enhanced), else the system voice. "default": always the system voice.
38
+ */
36
39
  voice: process.env.VOICE_MCP_VOICE?.trim() || undefined,
37
40
  /** Speech rate in words per minute (macOS `say -r`). */
38
41
  rate: process.env.VOICE_MCP_RATE?.trim() || undefined,
package/dist/setup.js CHANGED
@@ -7,7 +7,7 @@
7
7
  * the result. A slow `brew install` can never time the client out.
8
8
  */
9
9
  import { statSync } from "node:fs";
10
- import { describeRecorder, findRecorder, findTts } from "./audio.js";
10
+ import { chooseVoice, describeRecorder, findRecorder, findTts, resetVoiceChoice, VOICE_UPGRADE_HINT } from "./audio.js";
11
11
  import { CONFIG, IS_MAC, IS_WIN, log } from "./config.js";
12
12
  import { ensureModel, isModelDownloading, locateModel, modelSource, MODEL_SIZES_MB } from "./model.js";
13
13
  import { isExecutable, resetWhichCache, run, sleep, tail, which } from "./proc.js";
@@ -68,6 +68,13 @@ async function checkRequirements() {
68
68
  checks.push(tts
69
69
  ? { label: "Text-to-speech", status: "ok", detail: IS_MAC ? "macOS `say` (built in)" : tts }
70
70
  : { label: "Text-to-speech", status: "missing", detail: "no engine found — install espeak-ng (`sudo apt install espeak-ng`)" });
71
+ if (tts && IS_MAC) {
72
+ resetVoiceChoice(); // pick up a voice downloaded since the last check
73
+ const choice = await chooseVoice();
74
+ checks.push(choice.canUpgrade
75
+ ? { label: "Voice", status: "optional", detail: `${choice.label} — ${VOICE_UPGRADE_HINT}` }
76
+ : { label: "Voice", status: "ok", detail: choice.label });
77
+ }
71
78
  const rec = findRecorder();
72
79
  checks.push(rec
73
80
  ? { label: "Recorder", status: "ok", detail: describeRecorder(rec) }
@@ -162,6 +169,9 @@ export async function runSetupFlow(install, onProgress) {
162
169
  lines.push("");
163
170
  if (ready) {
164
171
  lines.push("Next: everything is in place — speak_and_listen is ready to use.");
172
+ if (checks.some((c) => c.label === "Voice" && c.status === "optional")) {
173
+ lines.push("Mention the optional voice tip (•) to the user once — it makes the voice sound far more natural. Nothing else to do.");
174
+ }
165
175
  }
166
176
  else if (installing) {
167
177
  lines.push("Next: installation is still running in the background. Tell the user, wait about a minute, then call voice_setup again (install=false) to check.");
package/dist/texts.js CHANGED
@@ -50,6 +50,7 @@ export const SPEAK_TOOL_DESCRIPTION = [
50
50
  "Once the user is talking with you by voice, keep the conversation in voice: answer each transcript with",
51
51
  "another speak_and_listen call (not a text reply) until they say stop or start typing.",
52
52
  'A reply of "(No speech detected …)" means the user did not answer — never treat it as consent.',
53
+ 'The "voice-mcp timing" line at the end is diagnostics: ignore it unless the user asks why things feel slow.',
53
54
  "If it reports that voice-mcp is not set up, call voice_setup.",
54
55
  ].join("\n");
55
56
  export const SETUP_TOOL_DESCRIPTION = [
package/dist/voice.js CHANGED
@@ -29,12 +29,15 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
29
29
  const t0 = Date.now();
30
30
  onPhase?.("speaking");
31
31
  await speak(speech.text, signal);
32
+ const spoke = Date.now() - t0;
32
33
  await chime("start");
33
34
  onPhase?.("listening");
35
+ const tListen = Date.now();
34
36
  const heard = await listenForTurn(seconds, wav, signal);
35
37
  void chime("stop");
36
38
  const t1 = Date.now();
37
39
  debug("listen:", heard);
40
+ const timing = (transcribedMs) => timingNote({ spokeMs: spoke, listenedMs: t1 - tListen, talkedSeconds: heard.speechSeconds, transcribedMs });
38
41
  if (heard.digitalSilence) {
39
42
  return {
40
43
  ok: false,
@@ -45,12 +48,12 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
45
48
  const waited = Math.min(CONFIG.startTimeoutSeconds, seconds);
46
49
  const noSpeech = `(No speech detected — the user did not reply within ${waited} seconds.)`;
47
50
  if (heard.reason === "no-speech")
48
- return { ok: true, text: noSpeech, notes };
51
+ return { ok: true, text: noSpeech, notes: [...notes, timing()] };
49
52
  onPhase?.("transcribing");
50
53
  const transcript = await transcribe(wav, model, signal);
51
- debug(`timing: speak+listen ${t1 - t0} ms, transcribe ${Date.now() - t1} ms`);
54
+ const transcribed = Date.now() - t1;
52
55
  if (!transcript)
53
- return { ok: true, text: noSpeech, notes };
56
+ return { ok: true, text: noSpeech, notes: [...notes, timing(transcribed)] };
54
57
  if (heard.reason === "max-duration") {
55
58
  const l = heard.levels;
56
59
  const f = (n) => (Number.isFinite(n) ? n.toFixed(0) : "?");
@@ -59,12 +62,23 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
59
62
  `(Audio levels, for tuning: background ${f(l.floorDb)} dB, voice ${f(l.speechDb)} dB, ` +
60
63
  `last second ${f(l.recentDb)} ± ${l.recentSpreadDb.toFixed(1)} dB.)`);
61
64
  }
62
- return { ok: true, text: transcript, notes };
65
+ return { ok: true, text: transcript, notes: [...notes, timing(transcribed)] };
63
66
  }
64
67
  finally {
65
68
  await rm(tmpDir, { recursive: true, force: true });
66
69
  }
67
70
  }
71
+ const secs = (ms) => `${(ms / 1000).toFixed(1)} s`;
72
+ /** One compact line per turn, so "why did that feel slow?" has an answer. */
73
+ export function timingNote(t) {
74
+ const parts = [
75
+ `spoke ${secs(t.spokeMs)}`,
76
+ `listened ${secs(t.listenedMs)}` + (t.talkedSeconds > 0 ? ` (user talked ${t.talkedSeconds.toFixed(1)} s)` : " (no speech)"),
77
+ ];
78
+ if (t.transcribedMs !== undefined)
79
+ parts.push(`transcribed ${secs(t.transcribedMs)}`);
80
+ return `voice-mcp timing: ${parts.join(" · ")}`;
81
+ }
68
82
  /** Serialize calls: there is one speaker and one microphone. */
69
83
  let queue = Promise.resolve();
70
84
  export function exclusive(fn) {
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "mac-voice-mcp",
3
3
  "mcpName": "io.github.jeet0007/mac-voice-mcp",
4
- "version": "0.1.1",
4
+ "version": "0.2.0",
5
5
  "description": "Talk with Claude out loud on your Mac: speaks with macOS `say`, listens for one natural conversational turn, and transcribes on-device with whisper.cpp. An MCP server.",
6
6
  "type": "module",
7
7
  "bin": {
@@ -26,7 +26,9 @@
26
26
  "inspect": "npx -y @modelcontextprotocol/inspector node dist/index.js",
27
27
  "prepare": "npm run build",
28
28
  "prepublishOnly": "npm test",
29
- "audit": "npm audit --omit=dev --audit-level=high && npm audit signatures"
29
+ "audit": "npm audit --omit=dev --audit-level=high && npm audit signatures",
30
+ "preversion": "npm test",
31
+ "version": "node scripts/sync-version.mjs && git add server.json .claude-plugin/plugin.json"
30
32
  },
31
33
  "dependencies": {
32
34
  "@modelcontextprotocol/sdk": "^1.30.0",