mac-voice-mcp 0.1.1 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +65 -8
- package/README.md +81 -17
- package/dist/audio.js +59 -2
- package/dist/config.js +9 -2
- package/dist/index.js +4 -1
- package/dist/lock.js +162 -0
- package/dist/server.js +13 -2
- package/dist/setup.js +11 -1
- package/dist/texts.js +10 -5
- package/dist/voice.js +53 -10
- package/package.json +4 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,18 +1,75 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
All notable changes to this project are documented here.
|
|
4
|
+
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and the project uses [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
5
|
+
|
|
6
|
+
## [Unreleased]
|
|
7
|
+
|
|
8
|
+
## [0.3.0] - 2026-09-28
|
|
9
|
+
|
|
10
|
+
### Added
|
|
11
|
+
- **Voice conversations stay in voice (Claude Code plugin).** Once you've answered out loud, a plugin hook sends Claude back to reply by voice if it tries to answer in text. It does this at most once per spoken answer, so it can never loop. Voice mode ends when you type, when you don't answer, or when Claude says goodbye with `listen: false`. It's tracked per session and expires after 30 idle minutes. `VOICE_MCP_STAY_IN_VOICE=0` turns the hook off.
|
|
12
|
+
- **`speak_and_listen` with `listen: false`** speaks without opening the microphone, for one-way announcements or a goodbye. It needs only text-to-speech, not the recorder, whisper or a model.
|
|
13
|
+
- **One voice turn at a time across the whole Mac.** Claude Code windows, Claude Desktop and Cursor take turns at the speaker and mic, one turn at a time, so they never talk over each other. A waiting turn reports "waiting for another voice session" and gives up after `VOICE_MCP_LOCK_WAIT_SECONDS` (120) with a clear message. A lock left by a crashed session is taken over. If the cache folder isn't writable, turns go ahead without the lock.
|
|
14
|
+
- **README: an Upgrading section and a Voice mode section, plus a note on testing the plugin from a checkout.**
|
|
15
|
+
|
|
16
|
+
### Changed
|
|
17
|
+
- **Claude waits longer for you to start talking: 15 seconds, up from 8** (`VOICE_MCP_START_TIMEOUT_SECONDS`), so reading or thinking doesn't end the conversation.
|
|
18
|
+
- **When you don't answer, Claude asks once more, then pauses cleanly.** The "no speech" result now tells Claude that the mic is off, so it tells you to type to carry on rather than to speak.
|
|
19
|
+
|
|
20
|
+
### Removed
|
|
21
|
+
- The unused `packaging/` folder. `.github/` is the only copy of the workflows.
|
|
22
|
+
|
|
23
|
+
## [0.2.0] - 2026-09-28
|
|
24
|
+
|
|
25
|
+
### Added
|
|
26
|
+
- **A more natural voice, automatically.** On macOS, speech uses the best Premium or Enhanced voice installed for the spoken language, preferring your Mac's region, and otherwise the system voice. `voice_setup` shows which voice is in use and, if there's no Premium voice, how to download one for free in System Settings. `VOICE_MCP_VOICE=default` keeps the system voice.
|
|
27
|
+
- **A timing line on every turn.** Each `speak_and_listen` result ends with a line like `voice-mcp timing: spoke 3.1 s · listened 4.0 s (user talked 2.2 s) · transcribed 0.3 s`, which makes a slow turn easy to explain. The model is told to ignore it unless you ask.
|
|
28
|
+
- **Claude Code plugin commands and a skill:**
|
|
29
|
+
- `/mac-voice-mcp:talk [task]` starts a voice conversation. It pre-approves voice turns while it runs.
|
|
30
|
+
- `/mac-voice-mcp:setup` checks, installs with your OK, and tests.
|
|
31
|
+
- The `voice-help` skill gives Claude fuller usage guidance and a troubleshooting table.
|
|
32
|
+
- **README: "Allow voice turns without prompts".** The `permissions.allow` entries that stop Claude Code from asking for approval on every voice turn.
|
|
33
|
+
- **Releases:**
|
|
34
|
+
- `npm version patch|minor|major` now also updates `server.json` and the plugin.
|
|
35
|
+
- Each tag creates a GitHub Release from this changelog.
|
|
36
|
+
- Tests fail if any version, the plugin's pinned npm version or the changelog entry is out of sync.
|
|
37
|
+
|
|
38
|
+
### Changed
|
|
39
|
+
- **The plugin runs the exact npm release it ships with** (`mac-voice-mcp@0.2.0`), instead of whatever version npx cached first.
|
|
40
|
+
- **The README's install configs use `mac-voice-mcp@latest`,** so npx picks up new releases. This covers Claude Desktop, Cursor, VS Code and `claude mcp add`.
|
|
41
|
+
- **The publish workflow is safe to re-run.**
|
|
42
|
+
- It skips npm and the MCP Registry when the version is already there.
|
|
43
|
+
- It retries the registry while npm catches up.
|
|
44
|
+
- It checks out without persisting credentials.
|
|
45
|
+
|
|
46
|
+
## [0.1.1] - 2026-09-23
|
|
47
|
+
|
|
48
|
+
### Added
|
|
49
|
+
- **README:** one-click install buttons for Cursor and VS Code, and install steps for the Claude Code plugin marketplace and the official MCP Registry.
|
|
50
|
+
- **Releases publish through npm trusted publishing, with provenance.**
|
|
4
51
|
|
|
5
52
|
### Changed
|
|
6
53
|
- **Node.js 22 or newer is now required.** Node 18 and 20 no longer receive security fixes. CI tests on Node 22, 24 and 26.
|
|
7
54
|
|
|
8
55
|
### Fixed
|
|
9
|
-
- **"It's on screen" when it isn't** ([#3](https://github.com/jeet0007/mac-voice-mcp/issues/3)).
|
|
56
|
+
- **"It's on screen" when it isn't** ([#3](https://github.com/jeet0007/mac-voice-mcp/issues/3)).
|
|
57
|
+
- When the spoken text sends you to look at something, the model is reminded that only its own message text reaches your screen, not the speech and not Bash/tool output. The speaking rules now say the same.
|
|
58
|
+
- Thanks to Copilot for the first version ([#4](https://github.com/jeet0007/mac-voice-mcp/pull/4)).
|
|
59
|
+
- The trigger phrases were narrowed so that ordinary sentences like "above thirty degrees" or "see the doctor" don't set it off.
|
|
60
|
+
|
|
61
|
+
### Security
|
|
62
|
+
- TruffleHog secret scanning, CodeQL, dependency review, `npm audit` in CI, and Dependabot.
|
|
10
63
|
|
|
11
|
-
|
|
12
|
-
- README: one-click install buttons for Cursor and VS Code, and install steps for the Claude Code plugin marketplace and the official MCP Registry.
|
|
13
|
-
- Security: TruffleHog secret scanning, CodeQL, dependency review, `npm audit` in CI, and Dependabot.
|
|
14
|
-
- Releases publish through npm trusted publishing, with provenance.
|
|
64
|
+
## [0.1.0] - 2026-09-23
|
|
15
65
|
|
|
16
|
-
|
|
66
|
+
### Added
|
|
67
|
+
- First release:
|
|
68
|
+
- `speak_and_listen`: natural turn-taking, with on-device whisper.cpp and a warm server.
|
|
69
|
+
- `voice_setup`: checks first, installs only what's missing, and only after you agree.
|
|
70
|
+
- The `setup` and `voice_mode` prompts.
|
|
17
71
|
|
|
18
|
-
|
|
72
|
+
[Unreleased]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.2.0...HEAD
|
|
73
|
+
[0.2.0]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.1.1...v0.2.0
|
|
74
|
+
[0.1.1]: https://github.com/jeet0007/mac-voice-mcp/compare/119a589...v0.1.1
|
|
75
|
+
[0.1.0]: https://github.com/jeet0007/mac-voice-mcp/tree/119a589
|
package/README.md
CHANGED
|
@@ -51,7 +51,7 @@ The server has **two tools and two prompts**:
|
|
|
51
51
|
| `/mcp__voice-mcp__setup` | A guided setup: it checks, asks you, installs, then runs a spoken test. |
|
|
52
52
|
| `/mcp__voice-mcp__voice_mode` | A hands-free session where Claude checks in by voice at natural points. |
|
|
53
53
|
|
|
54
|
-
**Listening works like a conversation.** A soft chime plays when the mic opens. The server waits for you to start talking and hands back to Claude about a second after you stop. It adjusts to background noise, doesn't cut you off at pauses mid-sentence, and ignores coughs and clicks. If you say nothing for
|
|
54
|
+
**Listening works like a conversation.** A soft chime plays when the mic opens. The server waits for you to start talking and hands back to Claude about a second after you stop. It adjusts to background noise, doesn't cut you off at pauses mid-sentence, and ignores coughs and clicks. If you say nothing for 15 seconds, Claude gets "no speech", which it is told never to treat as a yes.
|
|
55
55
|
|
|
56
56
|
**Replies come back fast.** whisper.cpp's server keeps the speech model loaded between turns, and the model starts loading while Claude is still talking. You don't wait for a model load on each reply. After 15 idle minutes the server shuts down to free memory. It is stopped automatically even if the MCP server crashes.
|
|
57
57
|
|
|
@@ -64,9 +64,9 @@ You need **Node.js 22 or newer**. Whichever way you install, run setup once afte
|
|
|
64
64
|
### One click
|
|
65
65
|
|
|
66
66
|
<p>
|
|
67
|
-
<a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=
|
|
68
|
-
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
69
|
-
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%22%5D%7D&quality=insiders"><img alt="Install in VS Code Insiders" src="https://img.shields.io/badge/VS_Code_Insiders-Install_Server-24bfa5?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
67
|
+
<a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIm1hYy12b2ljZS1tY3BAbGF0ZXN0Il19"><img alt="Add to Cursor" src="https://cursor.com/deeplink/mcp-install-dark.svg" height="32"></a>
|
|
68
|
+
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
69
|
+
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D&quality=insiders"><img alt="Install in VS Code Insiders" src="https://img.shields.io/badge/VS_Code_Insiders-Install_Server-24bfa5?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
70
70
|
</p>
|
|
71
71
|
|
|
72
72
|
### From a marketplace
|
|
@@ -76,6 +76,7 @@ You need **Node.js 22 or newer**. Whichever way you install, run setup once afte
|
|
|
76
76
|
/plugin marketplace add jeet0007/mac-voice-mcp
|
|
77
77
|
/plugin install mac-voice-mcp@mac-voice-mcp
|
|
78
78
|
```
|
|
79
|
+
The plugin adds `/mac-voice-mcp:setup` and `/mac-voice-mcp:talk`, a skill that teaches Claude how to use voice well and fix common problems, and a hook that keeps a voice conversation in voice (see [Voice mode](#voice-mode)). Each plugin version runs the matching npm release.
|
|
79
80
|
- **The official MCP Registry.** It's listed as [`io.github.jeet0007/mac-voice-mcp`](https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.jeet0007/mac-voice-mcp). Apps and directories that read the registry pick it up from there. In VS Code, open the Extensions view (⇧⌘X), search `@mcp mac-voice`, and click **Install**. Smithery, Glama, PulseMCP and mcp.so copy the registry, so it shows up there too.
|
|
80
81
|
|
|
81
82
|
### By hand
|
|
@@ -83,7 +84,7 @@ You need **Node.js 22 or newer**. Whichever way you install, run setup once afte
|
|
|
83
84
|
**Claude Code**
|
|
84
85
|
|
|
85
86
|
```bash
|
|
86
|
-
claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp
|
|
87
|
+
claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp@latest
|
|
87
88
|
```
|
|
88
89
|
|
|
89
90
|
**Claude Desktop.** Add this to `~/Library/Application Support/Claude/claude_desktop_config.json`, then quit (⌘Q) and reopen the app:
|
|
@@ -93,12 +94,14 @@ claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp
|
|
|
93
94
|
"mcpServers": {
|
|
94
95
|
"voice-mcp": {
|
|
95
96
|
"command": "npx",
|
|
96
|
-
"args": ["-y", "mac-voice-mcp"]
|
|
97
|
+
"args": ["-y", "mac-voice-mcp@latest"]
|
|
97
98
|
}
|
|
98
99
|
}
|
|
99
100
|
}
|
|
100
101
|
```
|
|
101
102
|
|
|
103
|
+
`@latest` makes npx check for a new release each time the app starts. Without it, npx keeps running whichever version it cached first.
|
|
104
|
+
|
|
102
105
|
If you get `spawn npx ENOENT`, use the full path from `which npx`, e.g. `"command": "/opt/homebrew/bin/npx"`.
|
|
103
106
|
|
|
104
107
|
**Cursor.** Add the same `mcpServers` block to `~/.cursor/mcp.json`.
|
|
@@ -106,20 +109,20 @@ If you get `spawn npx ENOENT`, use the full path from `which npx`, e.g. `"comman
|
|
|
106
109
|
**VS Code.** Run **MCP: Add Server** from the Command Palette, or:
|
|
107
110
|
|
|
108
111
|
```bash
|
|
109
|
-
code --add-mcp '{"name":"voice-mcp","command":"npx","args":["-y","mac-voice-mcp"]}'
|
|
112
|
+
code --add-mcp '{"name":"voice-mcp","command":"npx","args":["-y","mac-voice-mcp@latest"]}'
|
|
110
113
|
```
|
|
111
114
|
|
|
112
|
-
**Any other MCP client.** Run `npx -y mac-voice-mcp` as a stdio server.
|
|
115
|
+
**Any other MCP client.** Run `npx -y mac-voice-mcp@latest` as a stdio server.
|
|
113
116
|
|
|
114
117
|
### Then: set up and allow the mic
|
|
115
118
|
|
|
116
|
-
**Run setup once.** Ask Claude to *"set up voice"*. In Claude Code you can also run `/mcp__voice-mcp__setup
|
|
119
|
+
**Run setup once.** Ask Claude to *"set up voice"*. In Claude Code you can also run `/mac-voice-mcp:setup` (plugin) or `/mcp__voice-mcp__setup` (added by hand), or from a terminal run `npx -y mac-voice-mcp@latest setup`.
|
|
117
120
|
|
|
118
121
|
Setup checks what's already there before it changes anything:
|
|
119
122
|
|
|
120
123
|
| Needed | Provided by | If it's missing |
|
|
121
124
|
|---|---|---|
|
|
122
|
-
| Voice | macOS `say
|
|
125
|
+
| Voice | macOS `say`, using the most natural voice installed | Nothing to do. For a far better voice, add a free Premium one (see below). |
|
|
123
126
|
| Microphone capture | SoX (`rec`) | `brew install sox` |
|
|
124
127
|
| Speech-to-text | whisper.cpp (`whisper-cli` and `whisper-server`, Metal-accelerated) | `brew install whisper-cpp` |
|
|
125
128
|
| Speech model | `base.en`, ~140 MB | Downloaded once to `~/.cache/mac-voice-mcp/models/` |
|
|
@@ -130,6 +133,25 @@ Setup checks what's already there before it changes anything:
|
|
|
130
133
|
|
|
131
134
|
**Allow the microphone.** The first time Claude listens, macOS asks whether Claude (or Cursor, or your terminal) can use the microphone. Click Allow.
|
|
132
135
|
|
|
136
|
+
**Get a better voice (recommended).** macOS includes free Premium voices that sound far more natural than the default. Open **System Settings → Accessibility → Spoken Content → System Voice → Manage Voices…**, and download one, for example English → *Ava (Premium)* or *Zoe (Premium)*. The next voice turn uses it automatically. To choose a specific voice, or keep the system voice, see `VOICE_MCP_VOICE` under [Configuration](#configuration).
|
|
137
|
+
|
|
138
|
+
### Allow voice turns without prompts
|
|
139
|
+
|
|
140
|
+
By default, Claude Code asks for approval every time Claude wants to speak, which breaks the flow of a conversation. To allow voice turns, add the tool to the `permissions.allow` list in `~/.claude/settings.json`. Use the name that matches how you installed it:
|
|
141
|
+
|
|
142
|
+
```json
|
|
143
|
+
{
|
|
144
|
+
"permissions": {
|
|
145
|
+
"allow": [
|
|
146
|
+
"mcp__plugin_mac-voice-mcp_voice-mcp__speak_and_listen",
|
|
147
|
+
"mcp__voice-mcp__speak_and_listen"
|
|
148
|
+
]
|
|
149
|
+
}
|
|
150
|
+
}
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
The first name is for the plugin, the second for `claude mcp add voice-mcp …`. Leave `voice_setup` out, so installs still ask you first. While `/mac-voice-mcp:talk` runs, voice turns are already allowed.
|
|
154
|
+
|
|
133
155
|
### Installing from a clone
|
|
134
156
|
|
|
135
157
|
One command does everything above: it builds the project, runs setup (asking before installing anything), adds voice-mcp to Claude Desktop (backing up your config first) and to Claude Code, and offers a spoken test. It's safe to re-run, because each step checks first and skips anything already done.
|
|
@@ -138,13 +160,33 @@ One command does everything above: it builds the project, runs setup (asking bef
|
|
|
138
160
|
git clone https://github.com/jeet0007/mac-voice-mcp && bash mac-voice-mcp/install.sh
|
|
139
161
|
```
|
|
140
162
|
|
|
163
|
+
## Upgrading
|
|
164
|
+
|
|
165
|
+
- **Claude Code plugin:** run `/plugin marketplace update mac-voice-mcp`, then open `/plugin`, choose mac-voice-mcp under your installed plugins, and update it. Restart Claude Code. If there's no update option, uninstall and reinstall it.
|
|
166
|
+
- **Everything installed with `mac-voice-mcp@latest`** (Claude Desktop, Cursor, VS Code, `claude mcp add`): restart the app. npx fetches the new release when the server starts.
|
|
167
|
+
- **Configs without `@latest`:** change `mac-voice-mcp` to `mac-voice-mcp@latest` in the config, then restart the app. Otherwise npx keeps running the version it cached first.
|
|
168
|
+
|
|
169
|
+
Check which version you'd get with `npx -y mac-voice-mcp@latest --version`, and see what changed in the [changelog](CHANGELOG.md). Your model, voice and settings carry over.
|
|
170
|
+
|
|
141
171
|
## Using it
|
|
142
172
|
|
|
143
173
|
- **"Work on X and check in with me by voice when you need a decision."** Claude works quietly and only speaks at decision points.
|
|
144
|
-
- **`/mcp__voice-mcp__voice_mode fix the flaky login test
|
|
174
|
+
- **`/mac-voice-mcp:talk fix the flaky login test`** (plugin), or **`/mcp__voice-mcp__voice_mode fix the flaky login test`** (added by hand). Claude reads its plan back to you, then checks in at each checkpoint. Say "stop voice mode" or "I'm back" to end it.
|
|
145
175
|
- **"Read me a 20-second summary of this PR and ask if I should approve it."** Use this for one-off briefings.
|
|
146
176
|
- **Just talk after the chime.** You don't need to hurry or fill silence. If you're still talking at the 30-second safety cap (`listen_seconds`), Claude is told your reply may be cut off and asks you to continue.
|
|
147
177
|
|
|
178
|
+
### Voice mode
|
|
179
|
+
|
|
180
|
+
Once you answer out loud, you're in a voice conversation. Claude replies by voice, not in text, until one of these happens:
|
|
181
|
+
|
|
182
|
+
- **You type something.** You're back at the keyboard.
|
|
183
|
+
- **You say you're done.** Claude says a short goodbye without opening the mic (`speak_and_listen` with `listen: false`).
|
|
184
|
+
- **You don't answer.** After 15 seconds of silence Claude asks once more. If you still don't answer, it pauses and summarizes on screen, and the mic stays off. Type anything, or run `/mac-voice-mcp:talk`, to pick up again.
|
|
185
|
+
|
|
186
|
+
In Claude Code, the plugin enforces this with a hook. If Claude tries to answer in text mid-conversation, the hook sends it back once to answer by voice. If it stops again, the hook lets it. To turn the hook off, add `"VOICE_MCP_STAY_IN_VOICE": "0"` to the `env` block in `~/.claude/settings.json`. Other apps rely on the instructions alone.
|
|
187
|
+
|
|
188
|
+
**Several sessions, one mic.** Every mac-voice-mcp on your Mac takes turns: Claude Code windows, Claude Desktop and Cursor. While one is speaking or listening, the others wait for that turn to finish (shown as "waiting for another voice session"). They give up after 2 minutes with a message. If a session crashes, the next one takes the mic over.
|
|
189
|
+
|
|
148
190
|
## Getting Claude to sound natural
|
|
149
191
|
|
|
150
192
|
Guidance reaches Claude through several channels, because each client shows different ones:
|
|
@@ -155,6 +197,7 @@ Guidance reaches Claude through several channels, because each client shows diff
|
|
|
155
197
|
| Server instructions (when to use voice, how to handle replies, setup) | Claude Code (it reads up to 2 KB) |
|
|
156
198
|
| The `voice_mode` and `setup` prompts | Claude Code (as slash commands), Claude Desktop, Cursor |
|
|
157
199
|
| Server-side rewrite plus a `voice-mcp note` back to Claude | Always on |
|
|
200
|
+
| The plugin's `/mac-voice-mcp:talk`, `voice-help` skill and stay-in-voice hook | Claude Code, with the plugin |
|
|
158
201
|
|
|
159
202
|
The rules Claude is given:
|
|
160
203
|
|
|
@@ -178,17 +221,18 @@ Everything is optional. Set these in your client config's `"env": { … }` block
|
|
|
178
221
|
|
|
179
222
|
| Variable | Default | |
|
|
180
223
|
|---|---|---|
|
|
181
|
-
| `VOICE_MCP_VOICE` |
|
|
224
|
+
| `VOICE_MCP_VOICE` | most natural installed | Unset: the best Premium or Enhanced voice installed for the language, else the system voice. Set a name, e.g. `Ava (Premium)`, `Daniel`, `Kanya` (list them with `say -v '?'`), or `default` to always use the system voice. |
|
|
182
225
|
| `VOICE_MCP_RATE` | system rate | Words per minute, e.g. `200`. |
|
|
183
226
|
| `VOICE_MCP_MAX_SPEAK_WORDS` | `120` | Longer text is cut at a sentence boundary ("the rest is on screen"). |
|
|
184
227
|
| `VOICE_MCP_CHIME` | `1` | Set to `0` to turn off the mic open/close sounds. |
|
|
228
|
+
| `VOICE_MCP_LOCK_WAIT_SECONDS` | `120` | How long a turn waits while another session on this Mac is using the mic. |
|
|
185
229
|
|
|
186
230
|
**Listening**
|
|
187
231
|
|
|
188
232
|
| Variable | Default | |
|
|
189
233
|
|---|---|---|
|
|
190
234
|
| `VOICE_MCP_END_SILENCE_MS` | `1200` | How long a pause ends your turn. Use `1800` if it cuts you off while you think, `800` for snappier replies. |
|
|
191
|
-
| `VOICE_MCP_START_TIMEOUT_SECONDS` | `
|
|
235
|
+
| `VOICE_MCP_START_TIMEOUT_SECONDS` | `15` | How long to wait for you to start talking. |
|
|
192
236
|
| `VOICE_MCP_SPEECH_MARGIN_DB` | `12` | How much louder than room noise counts as speech. Raise it in noisy rooms. |
|
|
193
237
|
| `VOICE_MCP_MIN_SPEECH_DB` | `-48` | The quietest level that ever counts as speech (dBFS). |
|
|
194
238
|
| `VOICE_MCP_RECORDER` | `auto` | `sox` or `ffmpeg` (`ffmpeg` is macOS only). |
|
|
@@ -208,6 +252,12 @@ Everything is optional. Set these in your client config's `"env": { … }` block
|
|
|
208
252
|
| `VOICE_MCP_CACHE_DIR` | `~/.cache/mac-voice-mcp` | Where models are downloaded or symlinked. |
|
|
209
253
|
| `VOICE_MCP_DEBUG` | `0` | Verbose logs with per-turn timings, written to stderr. |
|
|
210
254
|
|
|
255
|
+
**Claude Code plugin hook.** Set this in the `env` block of `~/.claude/settings.json`, not in the server's config:
|
|
256
|
+
|
|
257
|
+
| Variable | Default | |
|
|
258
|
+
|---|---|---|
|
|
259
|
+
| `VOICE_MCP_STAY_IN_VOICE` | `1` | Set to `0` to stop the plugin's hook from sending Claude back to answer by voice. |
|
|
260
|
+
|
|
211
261
|
**Models** (whisper.cpp names):
|
|
212
262
|
|
|
213
263
|
| Model | Size | Good for |
|
|
@@ -221,7 +271,7 @@ Everything is optional. Set these in your client config's `"env": { … }` block
|
|
|
221
271
|
|
|
222
272
|
| Symptom | Fix |
|
|
223
273
|
|---|---|
|
|
224
|
-
| "voice-mcp is not set up yet" | Ask Claude to *set up voice*, or run `npx -y mac-voice-mcp setup`. |
|
|
274
|
+
| "voice-mcp is not set up yet" | Ask Claude to *set up voice*, or run `npx -y mac-voice-mcp@latest setup`. |
|
|
225
275
|
| "microphone returned pure digital silence" | macOS is blocking the mic for the host app. Go to **System Settings → Privacy & Security → Microphone**, enable Claude / Cursor / your terminal, then restart that app. |
|
|
226
276
|
| No permission prompt ever appears | Run `tccutil reset Microphone <bundle id>` and restart the app. Running `test` in Terminal only gives permission to Terminal, not to Claude Desktop. |
|
|
227
277
|
| It cuts me off while I'm thinking | Set `VOICE_MCP_END_SILENCE_MS=1800` (or up to `2500`). |
|
|
@@ -229,13 +279,18 @@ Everything is optional. Set these in your client config's `"env": { … }` block
|
|
|
229
279
|
| It hears its own voice | Use headphones, or turn the speaker volume down. It only listens after it finishes speaking, but echo can linger. |
|
|
230
280
|
| "Homebrew: not installed" | Install it from [brew.sh](https://brew.sh). It needs your password, so it can't run from Claude. Then run setup again. |
|
|
231
281
|
| It garbles names or jargon | Set `VOICE_MCP_WHISPER_PROMPT="Priya, Postgres, Kubernetes"`, or switch to `small.en`. |
|
|
282
|
+
| The voice sounds robotic | Download a Premium voice (see [Get a better voice](#then-set-up-and-allow-the-mic)). It's used automatically. |
|
|
283
|
+
| It asks for approval every turn | Add the tool to Claude Code's allow list: see [Allow voice turns without prompts](#allow-voice-turns-without-prompts). |
|
|
284
|
+
| "Another voice session on this Mac…" | Another Claude window, Claude Desktop or Cursor held the speaker and mic for over 2 minutes, which means one very long turn. End that conversation, then try again. |
|
|
285
|
+
| Claude keeps answering by voice after I'm done | Type anything, or say "stop voice mode". To switch the plugin's hook off entirely, see [Voice mode](#voice-mode). |
|
|
286
|
+
| Turns feel slow | Each result ends with a timing line, e.g. `spoke 3.1 s · listened 4.0 s · transcribed 0.3 s`. Ask Claude what it says. Transcribing should take well under a second. |
|
|
232
287
|
|
|
233
288
|
## Known limitations
|
|
234
289
|
|
|
235
290
|
- **You can't interrupt it.** It finishes speaking, then listens. Barge-in would mean listening while the speakers play, which needs headphones or echo cancellation.
|
|
236
291
|
- **It's macOS-first.** Linux works with SoX and espeak-ng. Windows is untested.
|
|
237
292
|
- **Turn-taking is based on loudness, not a speech model.** It adapts to background noise, but very noisy rooms, music or TV can confuse it. A headset helps, and so do the listening settings above.
|
|
238
|
-
- **One conversation at a time.** There's one speaker and one microphone, so
|
|
293
|
+
- **One conversation at a time.** There's one speaker and one microphone, so turns from every session on the Mac are queued.
|
|
239
294
|
|
|
240
295
|
## Privacy and safety
|
|
241
296
|
|
|
@@ -256,7 +311,7 @@ To report a vulnerability, see [SECURITY.md](SECURITY.md).
|
|
|
256
311
|
|
|
257
312
|
```bash
|
|
258
313
|
npm install
|
|
259
|
-
npm test # build +
|
|
314
|
+
npm test # build + unit tests, end-to-end tests over MCP with stub binaries, and release-metadata checks
|
|
260
315
|
npm run audit # known-vulnerability and signature checks on dependencies
|
|
261
316
|
npm run setup # check what's installed; offers to install what's missing
|
|
262
317
|
npm run test:voice # one real speak → listen → transcribe turn
|
|
@@ -272,14 +327,23 @@ npm run inspect # MCP Inspector
|
|
|
272
327
|
| `model.ts` | Finding, symlinking or downloading the model |
|
|
273
328
|
| `stt.ts` | The warm `whisper-server` with orphan guard, and the `whisper-cli` fallback |
|
|
274
329
|
| `setup.ts` | Requirement checks and consent-based background installs |
|
|
330
|
+
| `lock.ts` | One voice turn at a time across every session on the Mac |
|
|
275
331
|
| `voice.ts`, `server.ts`, `index.ts` | The round trip, the MCP tools and prompts, and the CLI |
|
|
332
|
+
| `skills/`, `hooks/` | The Claude Code plugin's `/mac-voice-mcp:talk` and `:setup` commands, the `voice-help` skill, and the stay-in-voice hook |
|
|
276
333
|
|
|
277
334
|
The package installs two commands: `mac-voice-mcp` (the one `npx -y mac-voice-mcp` runs), and `voice-mcp`.
|
|
278
335
|
|
|
336
|
+
**Testing the plugin: don't start Claude Code inside this repo.** In this folder, `npx mac-voice-mcp@<this version>` finds the checkout itself instead of downloading the package, can't run it, and the server fails with `CONNECTION_CLOSED`. Start `claude` in any other folder. To try unreleased plugin files (skills, hooks), swap the marketplace to your checkout: run `/plugin marketplace remove mac-voice-mcp`, then `/plugin marketplace add /path/to/checkout`, and install at user scope. The server still comes from npm, so test server changes with `npm test` and `npm run test:voice`.
|
|
337
|
+
|
|
279
338
|
### Releasing
|
|
280
339
|
|
|
281
340
|
- **First release:** `bash publish.sh`. It asks before each public step and uses your own GitHub and npm logins. It creates the GitHub repo, publishes to npm, and lists the server in the [official MCP Registry](https://registry.modelcontextprotocol.io), which Smithery, Glama, PulseMCP and mcp.so pick up from.
|
|
282
|
-
- **Later releases:**
|
|
341
|
+
- **Later releases:**
|
|
342
|
+
1. Add a section to `CHANGELOG.md` for the new version, in [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) format, dated today.
|
|
343
|
+
2. Run `npm version patch` (bug fixes), `npm version minor` (new features) or `npm version major` (breaking changes), per [semver](https://semver.org). It updates `package.json`, `package-lock.json`, `server.json` and the plugin (including the npm version the plugin runs), commits, and tags `v<version>`.
|
|
344
|
+
3. Run `git push --follow-tags`.
|
|
345
|
+
|
|
346
|
+
The Publish workflow then runs the tests, which also check that every version and the changelog entry match. It publishes to npm with provenance, lists the release in the MCP Registry, and creates a GitHub Release from the changelog section. It needs no secrets: npm and the registry both use GitHub's OIDC identity, once you've set the package's *Trusted Publisher* on npmjs.com (`publish.sh` prints the steps). If a step fails, fix the cause and use **Re-run failed jobs**. Steps that already finished are skipped.
|
|
283
347
|
|
|
284
348
|
## License
|
|
285
349
|
|
package/dist/audio.js
CHANGED
|
@@ -15,6 +15,62 @@ export function findTts() {
|
|
|
15
15
|
return "powershell.exe";
|
|
16
16
|
return which("espeak-ng") ?? which("espeak") ?? which("spd-say");
|
|
17
17
|
}
|
|
18
|
+
/** Parse `say -v '?'` lines such as "Ava (Premium) en_US # Hello! My name is Ava." */
|
|
19
|
+
export function parseSayVoices(output) {
|
|
20
|
+
const voices = [];
|
|
21
|
+
for (const line of output.split("\n")) {
|
|
22
|
+
const m = /^(.+?)\s+([a-z]{2,3}_[A-Za-z0-9]+)\s+#/.exec(line.trimEnd());
|
|
23
|
+
if (!m)
|
|
24
|
+
continue;
|
|
25
|
+
const name = m[1].trim();
|
|
26
|
+
const quality = /\(Premium\)/i.test(name) ? 3 : /\(Enhanced\)/i.test(name) ? 2 : 1;
|
|
27
|
+
voices.push({ name, locale: m[2], quality });
|
|
28
|
+
}
|
|
29
|
+
return voices;
|
|
30
|
+
}
|
|
31
|
+
/**
|
|
32
|
+
* The most natural installed voice for the language: Premium first, then Enhanced, preferring
|
|
33
|
+
* the system's region (en_US over en_GB on a US Mac). null = none installed; keep the system voice.
|
|
34
|
+
*/
|
|
35
|
+
export function pickVoice(voices, language, systemLocale) {
|
|
36
|
+
const lang = (language && language !== "auto" ? language : systemLocale).toLowerCase().split(/[-_]/)[0] ?? "en";
|
|
37
|
+
const region = (systemLocale.split(/[-_]/)[1] ?? "").toUpperCase();
|
|
38
|
+
const inRegion = (v) => (region && v.locale.toUpperCase().endsWith(`_${region}`) ? 1 : 0);
|
|
39
|
+
const candidates = voices.filter((v) => v.quality >= 2 && v.locale.toLowerCase().startsWith(`${lang}_`));
|
|
40
|
+
candidates.sort((a, b) => b.quality - a.quality || inRegion(b) - inRegion(a) || a.name.localeCompare(b.name));
|
|
41
|
+
return candidates[0] ?? null;
|
|
42
|
+
}
|
|
43
|
+
export const VOICE_UPGRADE_HINT = "for a much more natural voice, download a Premium one: System Settings → Accessibility → Spoken Content → " +
|
|
44
|
+
"System Voice → Manage Voices… (for example English → Ava (Premium) or Zoe (Premium)). It's used automatically — no restart needed.";
|
|
45
|
+
let voiceChoice;
|
|
46
|
+
/** Forget the choice, so a voice downloaded meanwhile is picked up (voice_setup calls this). */
|
|
47
|
+
export function resetVoiceChoice() {
|
|
48
|
+
voiceChoice = undefined;
|
|
49
|
+
}
|
|
50
|
+
export function chooseVoice() {
|
|
51
|
+
voiceChoice ??= (async () => {
|
|
52
|
+
if (CONFIG.voice) {
|
|
53
|
+
return CONFIG.voice.toLowerCase() === "default"
|
|
54
|
+
? { label: "the system voice (VOICE_MCP_VOICE=default)", canUpgrade: false }
|
|
55
|
+
: { voice: CONFIG.voice, label: `${CONFIG.voice} (set by VOICE_MCP_VOICE)`, canUpgrade: false };
|
|
56
|
+
}
|
|
57
|
+
const bin = IS_MAC ? findTts() : null;
|
|
58
|
+
if (!bin)
|
|
59
|
+
return { label: "the system voice", canUpgrade: false };
|
|
60
|
+
try {
|
|
61
|
+
const r = await run(bin, ["-v", "?"], { timeoutMs: 10_000 });
|
|
62
|
+
const locale = Intl.DateTimeFormat().resolvedOptions().locale;
|
|
63
|
+
const best = pickVoice(parseSayVoices(r.stdout), CONFIG.language, locale);
|
|
64
|
+
if (best)
|
|
65
|
+
return { voice: best.name, label: `${best.name} — the most natural voice installed`, canUpgrade: false };
|
|
66
|
+
}
|
|
67
|
+
catch {
|
|
68
|
+
/* fall back to the system voice */
|
|
69
|
+
}
|
|
70
|
+
return { label: "the system voice", canUpgrade: true };
|
|
71
|
+
})();
|
|
72
|
+
return voiceChoice;
|
|
73
|
+
}
|
|
18
74
|
export async function speak(text, signal) {
|
|
19
75
|
if (!text)
|
|
20
76
|
return;
|
|
@@ -24,8 +80,9 @@ export async function speak(text, signal) {
|
|
|
24
80
|
let result;
|
|
25
81
|
if (IS_MAC) {
|
|
26
82
|
const args = [];
|
|
27
|
-
|
|
28
|
-
|
|
83
|
+
const { voice } = await chooseVoice();
|
|
84
|
+
if (voice)
|
|
85
|
+
args.push("-v", voice);
|
|
29
86
|
if (CONFIG.rate)
|
|
30
87
|
args.push("-r", CONFIG.rate);
|
|
31
88
|
args.push("-f", "-"); // read from stdin: no argv length limits, no flag injection
|
package/dist/config.js
CHANGED
|
@@ -31,8 +31,15 @@ const modelName = (process.env.VOICE_MCP_WHISPER_MODEL ?? "base.en").trim();
|
|
|
31
31
|
const CACHE_ROOT = process.env.VOICE_MCP_CACHE_DIR?.trim() ||
|
|
32
32
|
path.join(process.env.XDG_CACHE_HOME || path.join(os.homedir(), ".cache"), "mac-voice-mcp");
|
|
33
33
|
export const CONFIG = {
|
|
34
|
+
/** Models, the mic lock and hook state live here. */
|
|
35
|
+
cacheDir: CACHE_ROOT,
|
|
36
|
+
/** How long a voice turn waits for another session to finish with the mic. */
|
|
37
|
+
lockWaitSeconds: Math.max(1, envNum("VOICE_MCP_LOCK_WAIT_SECONDS", 120)),
|
|
34
38
|
// --- Speaking
|
|
35
|
-
/**
|
|
39
|
+
/**
|
|
40
|
+
* macOS voice name, e.g. "Ava (Premium)" (`say -v '?'` lists them). Unset: the most natural
|
|
41
|
+
* installed voice for the language (Premium, then Enhanced), else the system voice. "default": always the system voice.
|
|
42
|
+
*/
|
|
36
43
|
voice: process.env.VOICE_MCP_VOICE?.trim() || undefined,
|
|
37
44
|
/** Speech rate in words per minute (macOS `say -r`). */
|
|
38
45
|
rate: process.env.VOICE_MCP_RATE?.trim() || undefined,
|
|
@@ -44,7 +51,7 @@ export const CONFIG = {
|
|
|
44
51
|
/** End of turn: stop listening after this much silence once the user has spoken. */
|
|
45
52
|
endSilenceMs: Math.max(300, envNum("VOICE_MCP_END_SILENCE_MS", 1200)),
|
|
46
53
|
/** Give up if the user hasn't started talking within this many seconds. */
|
|
47
|
-
startTimeoutSeconds: Math.max(1, envNum("VOICE_MCP_START_TIMEOUT_SECONDS",
|
|
54
|
+
startTimeoutSeconds: Math.max(1, envNum("VOICE_MCP_START_TIMEOUT_SECONDS", 15)),
|
|
48
55
|
/** Speech must be this many dB above the room's background noise. Lower = more sensitive. */
|
|
49
56
|
speechMarginDb: envNum("VOICE_MCP_SPEECH_MARGIN_DB", 12),
|
|
50
57
|
/** …and never quieter than this absolute level (dBFS). */
|
package/dist/index.js
CHANGED
|
@@ -84,7 +84,10 @@ async function runSetupCli(args) {
|
|
|
84
84
|
async function runTestCli(text) {
|
|
85
85
|
const prompt = text || "Voice bridge test. Say something after the chime, and I'll print what I heard.";
|
|
86
86
|
try {
|
|
87
|
-
const result = await speakAndListen(prompt,
|
|
87
|
+
const result = await speakAndListen(prompt, {
|
|
88
|
+
listenSeconds: envNum("VOICE_MCP_TEST_SECONDS", DEFAULT_LISTEN_SECONDS),
|
|
89
|
+
onPhase: (p) => console.error(` … ${p}`),
|
|
90
|
+
});
|
|
88
91
|
process.stdout.write(result.text + "\n");
|
|
89
92
|
for (const note of result.notes)
|
|
90
93
|
console.error(note);
|
package/dist/lock.js
ADDED
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* One voice turn at a time across every voice-mcp on this Mac — several Claude Code windows,
|
|
3
|
+
* Claude Desktop and Cursor all share one speaker and one microphone.
|
|
4
|
+
*
|
|
5
|
+
* The lock is a file created atomically (O_EXCL) holding the owner's pid and start time.
|
|
6
|
+
* A lock whose owner has died (or that is implausibly old) is taken over, so a crash never
|
|
7
|
+
* blocks the mic for good. If the lock can't be used at all (read-only or missing cache
|
|
8
|
+
* folder), turns go ahead without it rather than failing.
|
|
9
|
+
*/
|
|
10
|
+
import { closeSync, linkSync, mkdirSync, openSync, readFileSync, renameSync, statSync, unlinkSync, writeSync } from "node:fs";
|
|
11
|
+
import path from "node:path";
|
|
12
|
+
import { CONFIG, debug } from "./config.js";
|
|
13
|
+
import { CancelledError, sleep } from "./proc.js";
|
|
14
|
+
/** Longer than any real turn (speaking, up to two minutes of listening, transcribing). */
|
|
15
|
+
const STALE_AFTER_MS = 10 * 60_000;
|
|
16
|
+
/** A lock file that's still unreadable after this long was never going to be written. */
|
|
17
|
+
const UNREADABLE_GRACE_MS = 5_000;
|
|
18
|
+
const POLL_MS = 250;
|
|
19
|
+
export class MicBusyError extends Error {
|
|
20
|
+
name = "MicBusyError";
|
|
21
|
+
}
|
|
22
|
+
export const lockFile = () => path.join(CONFIG.cacheDir, "mic.lock");
|
|
23
|
+
function readOwner(file) {
|
|
24
|
+
try {
|
|
25
|
+
const o = JSON.parse(readFileSync(file, "utf8"));
|
|
26
|
+
return Number.isInteger(o.pid) && Number.isFinite(o.since) ? o : null;
|
|
27
|
+
}
|
|
28
|
+
catch {
|
|
29
|
+
return null;
|
|
30
|
+
}
|
|
31
|
+
}
|
|
32
|
+
function isAlive(pid) {
|
|
33
|
+
try {
|
|
34
|
+
process.kill(pid, 0);
|
|
35
|
+
return true;
|
|
36
|
+
}
|
|
37
|
+
catch (e) {
|
|
38
|
+
return e.code === "EPERM"; // exists, owned by someone else
|
|
39
|
+
}
|
|
40
|
+
}
|
|
41
|
+
/** The lock file's identity (inode) if it's stale, otherwise null. */
|
|
42
|
+
function staleIdentity(file, now) {
|
|
43
|
+
let ino;
|
|
44
|
+
let mtimeMs;
|
|
45
|
+
try {
|
|
46
|
+
({ ino, mtimeMs } = statSync(file));
|
|
47
|
+
}
|
|
48
|
+
catch {
|
|
49
|
+
return null; // gone already; the next attempt will create it
|
|
50
|
+
}
|
|
51
|
+
const owner = readOwner(file);
|
|
52
|
+
const stale = owner ? !isAlive(owner.pid) || now - owner.since > STALE_AFTER_MS : now - mtimeMs > UNREADABLE_GRACE_MS;
|
|
53
|
+
return stale ? ino : null;
|
|
54
|
+
}
|
|
55
|
+
/**
|
|
56
|
+
* Remove a stale lock without ever removing a fresh one someone else just created: move it aside
|
|
57
|
+
* atomically, check it's the same file we judged stale, and put it back if it isn't.
|
|
58
|
+
*/
|
|
59
|
+
function takeOver(file, ino) {
|
|
60
|
+
const aside = `${file}.${process.pid}.${Date.now()}.stale`;
|
|
61
|
+
try {
|
|
62
|
+
renameSync(file, aside);
|
|
63
|
+
}
|
|
64
|
+
catch {
|
|
65
|
+
return; // another session got there first
|
|
66
|
+
}
|
|
67
|
+
try {
|
|
68
|
+
if (statSync(aside).ino !== ino) {
|
|
69
|
+
try {
|
|
70
|
+
linkSync(aside, file); // it was a live lock after all: restore it (fails harmlessly if one exists)
|
|
71
|
+
}
|
|
72
|
+
catch {
|
|
73
|
+
/* a newer lock is already in place */
|
|
74
|
+
}
|
|
75
|
+
}
|
|
76
|
+
}
|
|
77
|
+
finally {
|
|
78
|
+
try {
|
|
79
|
+
unlinkSync(aside);
|
|
80
|
+
}
|
|
81
|
+
catch {
|
|
82
|
+
/* ignore */
|
|
83
|
+
}
|
|
84
|
+
}
|
|
85
|
+
}
|
|
86
|
+
/** Remove the file only if it's still ours. */
|
|
87
|
+
function releaseIfOurs(file, me) {
|
|
88
|
+
const now = readOwner(file);
|
|
89
|
+
if (!now || now.pid !== me.pid || now.since !== me.since)
|
|
90
|
+
return;
|
|
91
|
+
try {
|
|
92
|
+
unlinkSync(file);
|
|
93
|
+
}
|
|
94
|
+
catch {
|
|
95
|
+
/* already gone */
|
|
96
|
+
}
|
|
97
|
+
}
|
|
98
|
+
const UNUSABLE = new Set(["EACCES", "EPERM", "EROFS", "ENOENT", "ENOTDIR"]);
|
|
99
|
+
const noLock = () => { };
|
|
100
|
+
/** Wait for the mic, take it, and return the function that gives it back. */
|
|
101
|
+
export async function acquireMicLock(opts = {}) {
|
|
102
|
+
const file = lockFile();
|
|
103
|
+
const waitMs = opts.waitMs ?? CONFIG.lockWaitSeconds * 1000;
|
|
104
|
+
const deadline = Date.now() + waitMs;
|
|
105
|
+
let waiting = false;
|
|
106
|
+
try {
|
|
107
|
+
mkdirSync(path.dirname(file), { recursive: true });
|
|
108
|
+
}
|
|
109
|
+
catch (e) {
|
|
110
|
+
debug("mic lock unavailable, continuing without it:", e.message);
|
|
111
|
+
return noLock;
|
|
112
|
+
}
|
|
113
|
+
for (;;) {
|
|
114
|
+
if (opts.signal?.aborted)
|
|
115
|
+
throw new CancelledError();
|
|
116
|
+
const me = { pid: process.pid, since: Date.now() };
|
|
117
|
+
try {
|
|
118
|
+
const fd = openSync(file, "wx");
|
|
119
|
+
try {
|
|
120
|
+
writeSync(fd, JSON.stringify(me));
|
|
121
|
+
}
|
|
122
|
+
finally {
|
|
123
|
+
closeSync(fd);
|
|
124
|
+
}
|
|
125
|
+
let released = false;
|
|
126
|
+
const release = () => {
|
|
127
|
+
if (released)
|
|
128
|
+
return;
|
|
129
|
+
released = true;
|
|
130
|
+
process.off("exit", release);
|
|
131
|
+
releaseIfOurs(file, me);
|
|
132
|
+
};
|
|
133
|
+
process.on("exit", release); // a normal shutdown mid-turn still frees the mic
|
|
134
|
+
return release;
|
|
135
|
+
}
|
|
136
|
+
catch (e) {
|
|
137
|
+
const code = e.code ?? "";
|
|
138
|
+
if (UNUSABLE.has(code)) {
|
|
139
|
+
debug("mic lock unavailable, continuing without it:", e.message);
|
|
140
|
+
return noLock;
|
|
141
|
+
}
|
|
142
|
+
if (code !== "EEXIST")
|
|
143
|
+
throw e;
|
|
144
|
+
}
|
|
145
|
+
const ino = staleIdentity(file, Date.now());
|
|
146
|
+
if (ino !== null) {
|
|
147
|
+
debug("taking over a stale mic lock");
|
|
148
|
+
takeOver(file, ino);
|
|
149
|
+
continue;
|
|
150
|
+
}
|
|
151
|
+
if (!waiting) {
|
|
152
|
+
waiting = true;
|
|
153
|
+
opts.onWait?.();
|
|
154
|
+
}
|
|
155
|
+
if (Date.now() > deadline) {
|
|
156
|
+
throw new MicBusyError("Another voice session on this Mac (another Claude window, Claude Desktop or Cursor) has been using the " +
|
|
157
|
+
`speaker and microphone for over ${Math.round(waitMs / 1000)} seconds. ` +
|
|
158
|
+
"Tell the user on screen, and try again once its current turn is over.");
|
|
159
|
+
}
|
|
160
|
+
await sleep(POLL_MS);
|
|
161
|
+
}
|
|
162
|
+
}
|
package/dist/server.js
CHANGED
|
@@ -37,12 +37,23 @@ export function createServer() {
|
|
|
37
37
|
.optional()
|
|
38
38
|
.describe(`Upper limit on how long to listen, in seconds (default ${DEFAULT_LISTEN_SECONDS}, max ${MAX_LISTEN_SECONDS}). ` +
|
|
39
39
|
"Listening already stops when the user finishes talking, so you rarely need this."),
|
|
40
|
+
listen: z
|
|
41
|
+
.boolean()
|
|
42
|
+
.optional()
|
|
43
|
+
.describe("Default true. false: only speak, without opening the microphone — for a one-way announcement, " +
|
|
44
|
+
"or to say goodbye when the user ends voice mode."),
|
|
40
45
|
},
|
|
41
46
|
annotations: { readOnlyHint: false, destructiveHint: false, idempotentHint: false, openWorldHint: false },
|
|
42
|
-
}, async ({ text_to_speak, listen_seconds }, extra) => {
|
|
47
|
+
}, async ({ text_to_speak, listen_seconds, listen }, extra) => {
|
|
43
48
|
const progress = progressReporter(extra);
|
|
49
|
+
const phaseMessage = (phase) => phase === "waiting" ? "waiting for another voice session on this Mac to finish" : phase;
|
|
44
50
|
try {
|
|
45
|
-
const result = await exclusive(() => speakAndListen(text_to_speak,
|
|
51
|
+
const result = await exclusive(() => speakAndListen(text_to_speak, {
|
|
52
|
+
listenSeconds: listen_seconds ?? DEFAULT_LISTEN_SECONDS,
|
|
53
|
+
listen,
|
|
54
|
+
signal: extra.signal,
|
|
55
|
+
onPhase: (phase) => progress?.(phaseMessage(phase)),
|
|
56
|
+
}));
|
|
46
57
|
return {
|
|
47
58
|
content: [{ type: "text", text: result.text }, ...result.notes.map((note) => ({ type: "text", text: `\n\n${note}` }))],
|
|
48
59
|
isError: !result.ok,
|
package/dist/setup.js
CHANGED
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
* the result. A slow `brew install` can never time the client out.
|
|
8
8
|
*/
|
|
9
9
|
import { statSync } from "node:fs";
|
|
10
|
-
import { describeRecorder, findRecorder, findTts } from "./audio.js";
|
|
10
|
+
import { chooseVoice, describeRecorder, findRecorder, findTts, resetVoiceChoice, VOICE_UPGRADE_HINT } from "./audio.js";
|
|
11
11
|
import { CONFIG, IS_MAC, IS_WIN, log } from "./config.js";
|
|
12
12
|
import { ensureModel, isModelDownloading, locateModel, modelSource, MODEL_SIZES_MB } from "./model.js";
|
|
13
13
|
import { isExecutable, resetWhichCache, run, sleep, tail, which } from "./proc.js";
|
|
@@ -68,6 +68,13 @@ async function checkRequirements() {
|
|
|
68
68
|
checks.push(tts
|
|
69
69
|
? { label: "Text-to-speech", status: "ok", detail: IS_MAC ? "macOS `say` (built in)" : tts }
|
|
70
70
|
: { label: "Text-to-speech", status: "missing", detail: "no engine found — install espeak-ng (`sudo apt install espeak-ng`)" });
|
|
71
|
+
if (tts && IS_MAC) {
|
|
72
|
+
resetVoiceChoice(); // pick up a voice downloaded since the last check
|
|
73
|
+
const choice = await chooseVoice();
|
|
74
|
+
checks.push(choice.canUpgrade
|
|
75
|
+
? { label: "Voice", status: "optional", detail: `${choice.label} — ${VOICE_UPGRADE_HINT}` }
|
|
76
|
+
: { label: "Voice", status: "ok", detail: choice.label });
|
|
77
|
+
}
|
|
71
78
|
const rec = findRecorder();
|
|
72
79
|
checks.push(rec
|
|
73
80
|
? { label: "Recorder", status: "ok", detail: describeRecorder(rec) }
|
|
@@ -162,6 +169,9 @@ export async function runSetupFlow(install, onProgress) {
|
|
|
162
169
|
lines.push("");
|
|
163
170
|
if (ready) {
|
|
164
171
|
lines.push("Next: everything is in place — speak_and_listen is ready to use.");
|
|
172
|
+
if (checks.some((c) => c.label === "Voice" && c.status === "optional")) {
|
|
173
|
+
lines.push("Mention the optional voice tip (•) to the user once — it makes the voice sound far more natural. Nothing else to do.");
|
|
174
|
+
}
|
|
165
175
|
}
|
|
166
176
|
else if (installing) {
|
|
167
177
|
lines.push("Next: installation is still running in the background. Tell the user, wait about a minute, then call voice_setup again (install=false) to check.");
|
package/dist/texts.js
CHANGED
|
@@ -30,8 +30,8 @@ export const SERVER_INSTRUCTIONS = [
|
|
|
30
30
|
'- "(No speech detected …)" means no answer: never treat silence as consent. Ask once more or',
|
|
31
31
|
" continue with safe work and report on screen.",
|
|
32
32
|
"- Once the user is talking with you by voice, stay in voice: answer every turn with speak_and_listen,",
|
|
33
|
-
" not a text reply, until they say stop or start typing.",
|
|
34
|
-
"- Listening ends
|
|
33
|
+
" not a text reply, until they say stop (say goodbye with listen: false) or start typing.",
|
|
34
|
+
"- Listening ends when the user stops talking; listen_seconds is rarely needed.",
|
|
35
35
|
'- If the result includes a "voice-mcp note", follow it.',
|
|
36
36
|
"",
|
|
37
37
|
"Setup: if speak_and_listen says voice-mcp is not set up, call voice_setup (check only), tell the user",
|
|
@@ -49,7 +49,10 @@ export const SPEAK_TOOL_DESCRIPTION = [
|
|
|
49
49
|
"",
|
|
50
50
|
"Once the user is talking with you by voice, keep the conversation in voice: answer each transcript with",
|
|
51
51
|
"another speak_and_listen call (not a text reply) until they say stop or start typing.",
|
|
52
|
-
|
|
52
|
+
"To end voice mode, or for a one-way announcement, pass listen: false: it speaks without opening the mic.",
|
|
53
|
+
'A reply of "(No speech detected …)" means the user did not answer — never treat it as consent. Ask once more;',
|
|
54
|
+
"if still nothing, stop and say on screen that voice mode is paused and they can type to carry on (not speak: the mic is off).",
|
|
55
|
+
'The "voice-mcp timing" line at the end is diagnostics: ignore it unless the user asks why things feel slow.',
|
|
53
56
|
"If it reports that voice-mcp is not set up, call voice_setup.",
|
|
54
57
|
].join("\n");
|
|
55
58
|
export const SETUP_TOOL_DESCRIPTION = [
|
|
@@ -80,9 +83,11 @@ export const VOICE_MODE_PROMPT = (task) => [
|
|
|
80
83
|
" Don't narrate every small action.",
|
|
81
84
|
"- If a transcript is unclear, ask again by voice.",
|
|
82
85
|
"- Confirm by voice before anything destructive, irreversible, or that costs money.",
|
|
83
|
-
"- If I don't answer, don't assume yes:
|
|
86
|
+
"- If I don't answer, don't assume yes: ask once more out loud. If there's still nothing, pause and summarize",
|
|
87
|
+
" on screen, and tell me to type anything to pick up again (the mic is off, so don't tell me to speak).",
|
|
84
88
|
"- Keep writing full details (code, diffs, links) on screen as usual; the voice line is the headline.",
|
|
85
|
-
'- Stop using voice when I say "stop voice mode", "I\'m back", or start typing again.',
|
|
89
|
+
'- Stop using voice when I say "stop voice mode", "I\'m back", or start typing again. To end it, say a short',
|
|
90
|
+
" goodbye with speak_and_listen and listen: false (it speaks without opening the mic).",
|
|
86
91
|
"",
|
|
87
92
|
SPEECH_RULES,
|
|
88
93
|
"",
|
package/dist/voice.js
CHANGED
|
@@ -4,6 +4,7 @@ import os from "node:os";
|
|
|
4
4
|
import path from "node:path";
|
|
5
5
|
import { chime, findRecorder, listenForTurn, MIC_PERMISSION_HINT, RECORDER_MISSING, speak } from "./audio.js";
|
|
6
6
|
import { CONFIG, debug, DEFAULT_LISTEN_SECONDS, MAX_LISTEN_SECONDS, MAX_SPEAK_CHARS } from "./config.js";
|
|
7
|
+
import { acquireMicLock, MicBusyError } from "./lock.js";
|
|
7
8
|
import { ensureModel } from "./model.js";
|
|
8
9
|
import { SetupError } from "./proc.js";
|
|
9
10
|
import { prepareSpeech } from "./speech-text.js";
|
|
@@ -16,11 +17,33 @@ async function preflight() {
|
|
|
16
17
|
throw new SetupError(STT_MISSING);
|
|
17
18
|
return { model: await ensureModel({ download: false }) };
|
|
18
19
|
}
|
|
19
|
-
export async function speakAndListen(textToSpeak,
|
|
20
|
-
const
|
|
21
|
-
const
|
|
22
|
-
|
|
23
|
-
|
|
20
|
+
export async function speakAndListen(textToSpeak, opts = {}) {
|
|
21
|
+
const { signal, onPhase } = opts;
|
|
22
|
+
const listen = opts.listen !== false;
|
|
23
|
+
const requested = opts.listenSeconds ?? DEFAULT_LISTEN_SECONDS;
|
|
24
|
+
const seconds = Math.min(MAX_LISTEN_SECONDS, Math.max(1, Number.isFinite(requested) ? requested : DEFAULT_LISTEN_SECONDS));
|
|
25
|
+
const model = listen ? (await preflight()).model : null;
|
|
26
|
+
// Load the model into the warm server now, so it's ready when the user finishes (even if we wait below).
|
|
27
|
+
if (model)
|
|
28
|
+
prewarm(model);
|
|
29
|
+
// Wait for any other voice session on this Mac to finish with the speaker and mic.
|
|
30
|
+
let release;
|
|
31
|
+
try {
|
|
32
|
+
release = await acquireMicLock({ signal, onWait: () => onPhase?.("waiting") });
|
|
33
|
+
}
|
|
34
|
+
catch (err) {
|
|
35
|
+
if (err instanceof MicBusyError)
|
|
36
|
+
return { ok: false, text: err.message, notes: [] };
|
|
37
|
+
throw err;
|
|
38
|
+
}
|
|
39
|
+
try {
|
|
40
|
+
return await turn(textToSpeak, seconds, model, signal, onPhase);
|
|
41
|
+
}
|
|
42
|
+
finally {
|
|
43
|
+
release();
|
|
44
|
+
}
|
|
45
|
+
}
|
|
46
|
+
async function turn(textToSpeak, seconds, model, signal, onPhase) {
|
|
24
47
|
const speech = prepareSpeech(textToSpeak, { maxWords: CONFIG.maxSpeakWords, maxChars: MAX_SPEAK_CHARS });
|
|
25
48
|
const notes = [...speech.notes];
|
|
26
49
|
const tmpDir = await mkdtemp(path.join(os.tmpdir(), "voice-mcp-"));
|
|
@@ -29,12 +52,19 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
|
|
|
29
52
|
const t0 = Date.now();
|
|
30
53
|
onPhase?.("speaking");
|
|
31
54
|
await speak(speech.text, signal);
|
|
55
|
+
const spoke = Date.now() - t0;
|
|
56
|
+
if (!model) {
|
|
57
|
+
const text = "(Spoken. The microphone was not opened, because listen was false.)";
|
|
58
|
+
return { ok: true, text, notes: [...notes, `voice-mcp timing: spoke ${secs(spoke)}`] };
|
|
59
|
+
}
|
|
32
60
|
await chime("start");
|
|
33
61
|
onPhase?.("listening");
|
|
62
|
+
const tListen = Date.now();
|
|
34
63
|
const heard = await listenForTurn(seconds, wav, signal);
|
|
35
64
|
void chime("stop");
|
|
36
65
|
const t1 = Date.now();
|
|
37
66
|
debug("listen:", heard);
|
|
67
|
+
const timing = (transcribedMs) => timingNote({ spokeMs: spoke, listenedMs: t1 - tListen, talkedSeconds: heard.speechSeconds, transcribedMs });
|
|
38
68
|
if (heard.digitalSilence) {
|
|
39
69
|
return {
|
|
40
70
|
ok: false,
|
|
@@ -43,14 +73,16 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
|
|
|
43
73
|
};
|
|
44
74
|
}
|
|
45
75
|
const waited = Math.min(CONFIG.startTimeoutSeconds, seconds);
|
|
46
|
-
const noSpeech = `(No speech detected — the user did not reply within ${waited} seconds.
|
|
76
|
+
const noSpeech = `(No speech detected — the user did not reply within ${waited} seconds. The microphone is now off. ` +
|
|
77
|
+
"Ask once more out loud. If there's still no answer, stop and say on screen that voice mode is paused " +
|
|
78
|
+
"and they can type anything to carry on — speaking won't work until you call speak_and_listen again.)";
|
|
47
79
|
if (heard.reason === "no-speech")
|
|
48
|
-
return { ok: true, text: noSpeech, notes };
|
|
80
|
+
return { ok: true, text: noSpeech, notes: [...notes, timing()] };
|
|
49
81
|
onPhase?.("transcribing");
|
|
50
82
|
const transcript = await transcribe(wav, model, signal);
|
|
51
|
-
|
|
83
|
+
const transcribed = Date.now() - t1;
|
|
52
84
|
if (!transcript)
|
|
53
|
-
return { ok: true, text: noSpeech, notes };
|
|
85
|
+
return { ok: true, text: noSpeech, notes: [...notes, timing(transcribed)] };
|
|
54
86
|
if (heard.reason === "max-duration") {
|
|
55
87
|
const l = heard.levels;
|
|
56
88
|
const f = (n) => (Number.isFinite(n) ? n.toFixed(0) : "?");
|
|
@@ -59,12 +91,23 @@ export async function speakAndListen(textToSpeak, listenSeconds, signal, onPhase
|
|
|
59
91
|
`(Audio levels, for tuning: background ${f(l.floorDb)} dB, voice ${f(l.speechDb)} dB, ` +
|
|
60
92
|
`last second ${f(l.recentDb)} ± ${l.recentSpreadDb.toFixed(1)} dB.)`);
|
|
61
93
|
}
|
|
62
|
-
return { ok: true, text: transcript, notes };
|
|
94
|
+
return { ok: true, text: transcript, notes: [...notes, timing(transcribed)] };
|
|
63
95
|
}
|
|
64
96
|
finally {
|
|
65
97
|
await rm(tmpDir, { recursive: true, force: true });
|
|
66
98
|
}
|
|
67
99
|
}
|
|
100
|
+
const secs = (ms) => `${(ms / 1000).toFixed(1)} s`;
|
|
101
|
+
/** One compact line per turn, so "why did that feel slow?" has an answer. */
|
|
102
|
+
export function timingNote(t) {
|
|
103
|
+
const parts = [
|
|
104
|
+
`spoke ${secs(t.spokeMs)}`,
|
|
105
|
+
`listened ${secs(t.listenedMs)}` + (t.talkedSeconds > 0 ? ` (user talked ${t.talkedSeconds.toFixed(1)} s)` : " (no speech)"),
|
|
106
|
+
];
|
|
107
|
+
if (t.transcribedMs !== undefined)
|
|
108
|
+
parts.push(`transcribed ${secs(t.transcribedMs)}`);
|
|
109
|
+
return `voice-mcp timing: ${parts.join(" · ")}`;
|
|
110
|
+
}
|
|
68
111
|
/** Serialize calls: there is one speaker and one microphone. */
|
|
69
112
|
let queue = Promise.resolve();
|
|
70
113
|
export function exclusive(fn) {
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "mac-voice-mcp",
|
|
3
3
|
"mcpName": "io.github.jeet0007/mac-voice-mcp",
|
|
4
|
-
"version": "0.
|
|
4
|
+
"version": "0.3.0",
|
|
5
5
|
"description": "Talk with Claude out loud on your Mac: speaks with macOS `say`, listens for one natural conversational turn, and transcribes on-device with whisper.cpp. An MCP server.",
|
|
6
6
|
"type": "module",
|
|
7
7
|
"bin": {
|
|
@@ -26,7 +26,9 @@
|
|
|
26
26
|
"inspect": "npx -y @modelcontextprotocol/inspector node dist/index.js",
|
|
27
27
|
"prepare": "npm run build",
|
|
28
28
|
"prepublishOnly": "npm test",
|
|
29
|
-
"audit": "npm audit --omit=dev --audit-level=high && npm audit signatures"
|
|
29
|
+
"audit": "npm audit --omit=dev --audit-level=high && npm audit signatures",
|
|
30
|
+
"preversion": "npm test",
|
|
31
|
+
"version": "node scripts/sync-version.mjs && git add server.json .claude-plugin/plugin.json"
|
|
30
32
|
},
|
|
31
33
|
"dependencies": {
|
|
32
34
|
"@modelcontextprotocol/sdk": "^1.30.0",
|