mac-voice-mcp 0.3.1 → 0.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +38 -1
- package/README.md +72 -264
- package/dist/audio.js +133 -4
- package/dist/config.js +35 -3
- package/dist/doctor.js +185 -0
- package/dist/endpointer.js +14 -0
- package/dist/index.js +38 -6
- package/dist/kokoro-text.js +105 -0
- package/dist/kokoro-worker.js +103 -0
- package/dist/kokoro.js +449 -0
- package/dist/pcm.js +49 -0
- package/dist/proc.js +1 -0
- package/dist/server.js +9 -2
- package/dist/setup.js +58 -13
- package/dist/texts.js +2 -0
- package/dist/voice.js +23 -9
- package/dist/wer.js +41 -0
- package/package.json +4 -2
package/CHANGELOG.md
CHANGED
|
@@ -5,6 +5,39 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
|
|
|
5
5
|
|
|
6
6
|
## [Unreleased]
|
|
7
7
|
|
|
8
|
+
## [0.4.1] - 2026-10-10
|
|
9
|
+
|
|
10
|
+
### Added
|
|
11
|
+
- **README: a Disclaimer, a Third-party software table and a trademark note.** They spell out what the tool does on your Mac (mic, installs with your consent), that speech recognition can mishear, the license of everything it uses but doesn't bundle, and that the project isn't affiliated with Apple or Anthropic.
|
|
12
|
+
|
|
13
|
+
### Changed
|
|
14
|
+
- **whisper hears developer talk better.** In English it now gets a one-sentence hint that the conversation is a developer talking to a coding assistant. On 60 test clips read by three voices, word errors fell from 3.4% to 2.3% on everyday sentences and from 3.8% to 3.1% on developer jargon ("rebase", "backend"), and nothing was invented in silence. `VOICE_MCP_WHISPER_PROMPT` replaces the hint with your own, and `none` turns it off.
|
|
15
|
+
- **A shorter README, in the style of popular MCP servers:** what it's for, features, a four-step quick start and the most common fixes. The full guides moved to [`docs/`](docs/): install, voices, usage, configuration, troubleshooting, privacy and security, and development.
|
|
16
|
+
|
|
17
|
+
### Fixed
|
|
18
|
+
- **The first words of a reply were often lost.** The "mic open" chime played before the recorder had actually started, and opening a mic takes a moment (longer with Bluetooth), so anyone who answered right at the chime lost a word or two. In a test with 20 real recordings, 12 began mid-word, and the cut-off starts led whisper to mishear or invent the rest ("Open a draft PR…" came back as "You can find the link in the description below."). Now the mic opens first and the chime plays once audio is flowing; the chime's own sound is ignored. Re-recorded with the fix, 2 of 20 began mid-word, and whisper's word errors on that voice fell from about 20% to 5% (and from 28% to 12% with background noise).
|
|
19
|
+
|
|
20
|
+
### Security
|
|
21
|
+
- **MCP SDK 1.32.1** ([GHSA-6qxp-vccf-f47h](https://github.com/advisories/GHSA-6qxp-vccf-f47h)). The advisory is in the SDK's OAuth client, which this server never uses (it only talks over stdio), but `npm audit` flagged every install.
|
|
22
|
+
|
|
23
|
+
## [0.4.0] - 2026-10-06
|
|
24
|
+
|
|
25
|
+
### Added
|
|
26
|
+
- **An optional natural voice: Kokoro.** [Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) is a small neural voice that runs on your Mac and sounds much closer to a person than `say`.
|
|
27
|
+
- **Install it only if you want it:** `voice_setup` with `kokoro=true` (Claude asks first), or `npx mac-voice-mcp setup --kokoro`. It installs `kokoro-js@1.2.1` with npm into `~/.cache/mac-voice-mcp/kokoro/` and downloads the model once, about 1 GB on disk. Nothing is bundled, and speaking never downloads anything.
|
|
28
|
+
- **Once installed it's used automatically**, with the `af_heart` voice. It runs in its own warm process and speaks sentence by sentence, so the first sentence plays while the rest is generated. The timing line shows how soon the first sound came.
|
|
29
|
+
- **It can't leave you without a voice.** If Kokoro fails, the built-in voice says whatever hadn't been said yet, and Claude is told once.
|
|
30
|
+
- **Settings:** `VOICE_MCP_TTS` (`auto`, `say` or `kokoro`), `VOICE_MCP_KOKORO_VOICE`, `VOICE_MCP_KOKORO_SPEED`, `VOICE_MCP_KOKORO_DTYPE` and `VOICE_MCP_KOKORO_DIR`.
|
|
31
|
+
- **`doctor` and the CI round trip cover Kokoro too:** whisper must understand it, and the first sound must come quickly.
|
|
32
|
+
- **`doctor`: an objective self-check** (`npx mac-voice-mcp doctor`). It checks the whole pipeline on your Mac and marks each stage PASS, WARN or FAIL against fixed limits:
|
|
33
|
+
- **Setup:** everything is installed.
|
|
34
|
+
- **Speech → text:** the voice speaks a known sentence into a file, and whisper must get the words right.
|
|
35
|
+
- **Speaker → mic:** the same sentence is played aloud and recorded. It catches mic permission problems, the wrong input device and headphones.
|
|
36
|
+
|
|
37
|
+
Reports are saved as JSON under `~/.cache/mac-voice-mcp/doctor/`. Use `--no-loopback` to skip the speakers, and `--json` for machine-readable output.
|
|
38
|
+
- **A "Speech round trip" CI job on a real Mac.** It installs everything with our own `setup --install`, runs `doctor`, then speaks known sentences (everyday and developer jargon) with the real voice and transcribes them with the real whisper.cpp. The build fails on too many wrong words or slow transcription.
|
|
39
|
+
- **`npm run dev:plugin`** loads this checkout into Claude Code as a separate "mac-voice-mcp-dev" plugin. It runs the local build with `node`, with no npx and no clash with the installed plugin.
|
|
40
|
+
|
|
8
41
|
## [0.3.1] - 2026-09-29
|
|
9
42
|
|
|
10
43
|
### Fixed
|
|
@@ -78,7 +111,11 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
|
|
|
78
111
|
- `voice_setup`: checks first, installs only what's missing, and only after you agree.
|
|
79
112
|
- The `setup` and `voice_mode` prompts.
|
|
80
113
|
|
|
81
|
-
[Unreleased]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.
|
|
114
|
+
[Unreleased]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.4.1...HEAD
|
|
115
|
+
[0.4.1]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.4.0...v0.4.1
|
|
116
|
+
[0.4.0]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.3.1...v0.4.0
|
|
117
|
+
[0.3.1]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.3.0...v0.3.1
|
|
118
|
+
[0.3.0]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.2.0...v0.3.0
|
|
82
119
|
[0.2.0]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.1.1...v0.2.0
|
|
83
120
|
[0.1.1]: https://github.com/jeet0007/mac-voice-mcp/compare/119a589...v0.1.1
|
|
84
121
|
[0.1.0]: https://github.com/jeet0007/mac-voice-mcp/tree/119a589
|
package/README.md
CHANGED
|
@@ -24,7 +24,7 @@
|
|
|
24
24
|
|
|
25
25
|
---
|
|
26
26
|
|
|
27
|
-
Claude
|
|
27
|
+
Claude speaks through your Mac's speakers, listens to your answer the way a person would, and gets back what you said as text. Speech recognition runs on your Mac, so no audio leaves your computer.
|
|
28
28
|
|
|
29
29
|
```
|
|
30
30
|
Claude ──speak_and_listen("Tests pass. Open the PR?")──▶ 🔊 "Tests pass. Open the PR?"
|
|
@@ -32,112 +32,47 @@ Claude ──speak_and_listen("Tests pass. Open the PR?")──▶ 🔊 "Tests
|
|
|
32
32
|
Claude ◀──────────────── "Yes, and tag Priya." ───────── whisper.cpp on your Mac
|
|
33
33
|
```
|
|
34
34
|
|
|
35
|
-
|
|
35
|
+
**Perfect for:**
|
|
36
36
|
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
> It's an independent project, not made or endorsed by Anthropic. It works with any MCP client, including Claude Desktop, Claude Code and Cursor.
|
|
42
|
-
|
|
43
|
-
## How it works
|
|
44
|
-
|
|
45
|
-
The server has **two tools and two prompts**:
|
|
46
|
-
|
|
47
|
-
| | What it does |
|
|
48
|
-
|---|---|
|
|
49
|
-
| `speak_and_listen` | Speaks a short message, listens for one conversational turn and returns the transcript. |
|
|
50
|
-
| `voice_setup` | Checks what's installed. After you agree, it installs only what's missing. |
|
|
51
|
-
| `/mcp__voice-mcp__setup` | A guided setup: it checks, asks you, installs, then runs a spoken test. |
|
|
52
|
-
| `/mcp__voice-mcp__voice_mode` | A hands-free session where Claude checks in by voice at natural points. |
|
|
37
|
+
- Long tasks: Claude works quietly and checks in out loud when it needs a decision, so you can step away from the screen.
|
|
38
|
+
- Quick reviews: a 20-second summary of a pull request, then "approve it?"
|
|
39
|
+
- Resting your eyes, or your wrists, after a long day at the keyboard.
|
|
40
|
+
- Thinking out loud: talking a problem through instead of typing it.
|
|
53
41
|
|
|
54
|
-
|
|
42
|
+
## Features
|
|
55
43
|
|
|
56
|
-
|
|
44
|
+
- 🔒 **On-device.** whisper.cpp transcribes on your Mac. Recordings are deleted after each turn.
|
|
45
|
+
- 💬 **Natural turn-taking.** A soft chime, then it waits for you to start, and hands back about a second after you stop. No push-to-talk.
|
|
46
|
+
- ⚡ **Fast replies.** A warm whisper.cpp server keeps the model loaded, so transcribing takes a fraction of a second on Apple Silicon.
|
|
47
|
+
- 🗣️ **A natural voice, if you want one.** Uses your best macOS voice, or the optional [Kokoro](docs/voices.md#the-kokoro-voice-optional) neural voice.
|
|
48
|
+
- ✋ **Asks before installing anything.** Setup checks what you already have and reuses it.
|
|
49
|
+
- 🪟 **One mic, many windows.** Claude Code, Claude Desktop and Cursor take turns instead of talking over each other.
|
|
50
|
+
- 📏 **Measured, not guessed.** `doctor` checks the whole pipeline on your Mac and scores it PASS, WARN or FAIL.
|
|
57
51
|
|
|
58
|
-
|
|
52
|
+
## Quick start
|
|
59
53
|
|
|
60
|
-
|
|
54
|
+
You need a Mac (Apple Silicon recommended) and **Node.js 22 or newer**.
|
|
61
55
|
|
|
62
|
-
|
|
56
|
+
**1. Install.**
|
|
63
57
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
<p>
|
|
67
|
-
<a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIm1hYy12b2ljZS1tY3BAbGF0ZXN0Il19"><img alt="Add to Cursor" src="https://cursor.com/deeplink/mcp-install-dark.svg" height="32"></a>
|
|
68
|
-
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
69
|
-
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D&quality=insiders"><img alt="Install in VS Code Insiders" src="https://img.shields.io/badge/VS_Code_Insiders-Install_Server-24bfa5?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
70
|
-
</p>
|
|
71
|
-
|
|
72
|
-
### From a marketplace
|
|
73
|
-
|
|
74
|
-
- **Claude Code plugin marketplace.** This repo is its own marketplace:
|
|
58
|
+
- **Claude Code** (recommended). Install the plugin:
|
|
75
59
|
```
|
|
76
60
|
/plugin marketplace add jeet0007/mac-voice-mcp
|
|
77
61
|
/plugin install mac-voice-mcp@mac-voice-mcp
|
|
78
62
|
```
|
|
79
|
-
|
|
80
|
-
- **The official MCP Registry.** It's listed as [`io.github.jeet0007/mac-voice-mcp`](https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.jeet0007/mac-voice-mcp). Apps and directories that read the registry pick it up from there. In VS Code, open the Extensions view (⇧⌘X), search `@mcp mac-voice`, and click **Install**. Smithery, Glama, PulseMCP and mcp.so copy the registry, so it shows up there too.
|
|
81
|
-
|
|
82
|
-
### By hand
|
|
83
|
-
|
|
84
|
-
**Claude Code**
|
|
85
|
-
|
|
86
|
-
```bash
|
|
87
|
-
claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp@latest
|
|
88
|
-
```
|
|
89
|
-
|
|
90
|
-
**Claude Desktop.** Add this to `~/Library/Application Support/Claude/claude_desktop_config.json`, then quit (⌘Q) and reopen the app:
|
|
63
|
+
- **Cursor or VS Code:** <a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIm1hYy12b2ljZS1tY3BAbGF0ZXN0Il19"><img alt="Add to Cursor" src="https://cursor.com/deeplink/mcp-install-dark.svg" height="24"></a> <a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="24"></a>
|
|
91
64
|
|
|
92
|
-
|
|
93
|
-
{
|
|
94
|
-
"mcpServers": {
|
|
95
|
-
"voice-mcp": {
|
|
96
|
-
"command": "npx",
|
|
97
|
-
"args": ["-y", "mac-voice-mcp@latest"]
|
|
98
|
-
}
|
|
99
|
-
}
|
|
100
|
-
}
|
|
101
|
-
```
|
|
65
|
+
- **Claude Desktop and other apps:** add `npx -y mac-voice-mcp@latest` as a stdio MCP server. See [Install](docs/install.md#by-hand) for each app.
|
|
102
66
|
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
If you get `spawn npx ENOENT`, use the full path from `which npx`, e.g. `"command": "/opt/homebrew/bin/npx"`.
|
|
106
|
-
|
|
107
|
-
**Cursor.** Add the same `mcpServers` block to `~/.cursor/mcp.json`.
|
|
108
|
-
|
|
109
|
-
**VS Code.** Run **MCP: Add Server** from the Command Palette, or:
|
|
110
|
-
|
|
111
|
-
```bash
|
|
112
|
-
code --add-mcp '{"name":"voice-mcp","command":"npx","args":["-y","mac-voice-mcp@latest"]}'
|
|
113
|
-
```
|
|
67
|
+
**2. Set up.** Ask Claude to *"set up voice"* (or run `/mac-voice-mcp:setup`). It checks for SoX, whisper.cpp and a speech model, shows you what's missing, and installs it only after you say yes.
|
|
114
68
|
|
|
115
|
-
**
|
|
69
|
+
**3. Allow the microphone.** The first time Claude listens, macOS asks. Click **Allow**.
|
|
116
70
|
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
**Run setup once.** Ask Claude to *"set up voice"*. In Claude Code you can also run `/mac-voice-mcp:setup` (plugin) or `/mcp__voice-mcp__setup` (added by hand), or from a terminal run `npx -y mac-voice-mcp@latest setup`.
|
|
120
|
-
|
|
121
|
-
Setup checks what's already there before it changes anything:
|
|
122
|
-
|
|
123
|
-
| Needed | Provided by | If it's missing |
|
|
124
|
-
|---|---|---|
|
|
125
|
-
| Voice | macOS `say`, using the most natural voice installed | Nothing to do. For a far better voice, add a free Premium one (see below). |
|
|
126
|
-
| Microphone capture | SoX (`rec`). ffmpeg works as a fallback, but setup recommends SoX | `brew install sox` |
|
|
127
|
-
| Speech-to-text | whisper.cpp (`whisper-cli` and `whisper-server`, Metal-accelerated) | `brew install whisper-cpp` |
|
|
128
|
-
| Speech model | `base.en`, ~140 MB | Downloaded once to `~/.cache/mac-voice-mcp/models/` |
|
|
129
|
-
|
|
130
|
-
- **Nothing is redone.** Tools already on your PATH are used as they are. If the model is already somewhere on disk (a whisper.cpp checkout, Homebrew's share folder, another tool's cache, or anything Spotlight can find), it's **symlinked**, not downloaded again. `brew install` runs only for the missing formulae.
|
|
131
|
-
- **Nothing happens without your OK.** Claude calls `voice_setup` to check first, shows you the checklist, and asks before calling it with `install=true`.
|
|
132
|
-
- **Slow installs don't time out.** If `brew install whisper-cpp` takes a while, setup reports INSTALLING. The install carries on in the background, and the next check picks up the result.
|
|
133
|
-
|
|
134
|
-
**Allow the microphone.** The first time Claude listens, macOS asks whether Claude (or Cursor, or your terminal) can use the microphone. Click Allow.
|
|
135
|
-
|
|
136
|
-
**Get a better voice (recommended).** macOS includes free Premium voices that sound far more natural than the default. Open **System Settings → Accessibility → Spoken Content → System Voice → Manage Voices…**, and download one, for example English → *Ava (Premium)* or *Zoe (Premium)*. The next voice turn uses it automatically. To choose a specific voice, or keep the system voice, see `VOICE_MCP_VOICE` under [Configuration](#configuration).
|
|
71
|
+
**4. Talk.** Run `/mac-voice-mcp:talk fix the flaky login test`, or ask Claude to *"check in with me by voice when you need a decision"*. Reply after the chime.
|
|
137
72
|
|
|
138
73
|
### Allow voice turns without prompts
|
|
139
74
|
|
|
140
|
-
|
|
75
|
+
Claude Code asks for approval every time Claude wants to speak, which breaks the flow of a conversation. To allow voice turns, add the tool to `permissions.allow` in `~/.claude/settings.json`:
|
|
141
76
|
|
|
142
77
|
```json
|
|
143
78
|
{
|
|
@@ -150,202 +85,75 @@ By default, Claude Code asks for approval every time Claude wants to speak, whic
|
|
|
150
85
|
}
|
|
151
86
|
```
|
|
152
87
|
|
|
153
|
-
The first name is for the plugin, the second for `claude mcp add voice-mcp …`. Leave `voice_setup` out, so installs still ask you first.
|
|
154
|
-
|
|
155
|
-
### Installing from a clone
|
|
156
|
-
|
|
157
|
-
One command does everything above: it builds the project, runs setup (asking before installing anything), adds voice-mcp to Claude Desktop (backing up your config first) and to Claude Code, and offers a spoken test. It's safe to re-run, because each step checks first and skips anything already done.
|
|
88
|
+
The first name is for the plugin, the second for `claude mcp add voice-mcp …`. Leave `voice_setup` out, so installs still ask you first.
|
|
158
89
|
|
|
159
|
-
|
|
160
|
-
git clone https://github.com/jeet0007/mac-voice-mcp && bash mac-voice-mcp/install.sh
|
|
161
|
-
```
|
|
162
|
-
|
|
163
|
-
## Upgrading
|
|
164
|
-
|
|
165
|
-
- **Claude Code plugin:** run `/plugin marketplace update mac-voice-mcp`, then open `/plugin`, choose mac-voice-mcp under your installed plugins, and update it. Restart Claude Code. If there's no update option, uninstall and reinstall it.
|
|
166
|
-
- **Everything installed with `mac-voice-mcp@latest`** (Claude Desktop, Cursor, VS Code, `claude mcp add`): restart the app. npx fetches the new release when the server starts.
|
|
167
|
-
- **Configs without `@latest`:** change `mac-voice-mcp` to `mac-voice-mcp@latest` in the config, then restart the app. Otherwise npx keeps running the version it cached first.
|
|
168
|
-
|
|
169
|
-
Check which version you'd get with `npx -y mac-voice-mcp@latest --version`, and see what changed in the [changelog](CHANGELOG.md). Your model, voice and settings carry over.
|
|
170
|
-
|
|
171
|
-
## Using it
|
|
172
|
-
|
|
173
|
-
- **"Work on X and check in with me by voice when you need a decision."** Claude works quietly and only speaks at decision points.
|
|
174
|
-
- **`/mac-voice-mcp:talk fix the flaky login test`** (plugin), or **`/mcp__voice-mcp__voice_mode fix the flaky login test`** (added by hand). Claude reads its plan back to you, then checks in at each checkpoint. Say "stop voice mode" or "I'm back" to end it.
|
|
175
|
-
- **"Read me a 20-second summary of this PR and ask if I should approve it."** Use this for one-off briefings.
|
|
176
|
-
- **Just talk after the chime.** You don't need to hurry or fill silence. If you're still talking at the 30-second safety cap (`listen_seconds`), Claude is told your reply may be cut off and asks you to continue.
|
|
177
|
-
|
|
178
|
-
### Voice mode
|
|
179
|
-
|
|
180
|
-
Once you answer out loud, you're in a voice conversation. Claude replies by voice, not in text, until one of these happens:
|
|
90
|
+
## How it works
|
|
181
91
|
|
|
182
|
-
|
|
183
|
-
- **You say you're done.** Claude says a short goodbye without opening the mic (`speak_and_listen` with `listen: false`).
|
|
184
|
-
- **You don't answer.** After 15 seconds of silence Claude asks once more. If you still don't answer, it pauses and summarizes on screen, and the mic stays off. Type anything, or run `/mac-voice-mcp:talk`, to pick up again.
|
|
92
|
+
The server has two tools and two prompts:
|
|
185
93
|
|
|
186
|
-
|
|
94
|
+
| | What it does |
|
|
95
|
+
|---|---|
|
|
96
|
+
| `speak_and_listen` | Speaks a short message, listens for one conversational turn and returns the transcript. |
|
|
97
|
+
| `voice_setup` | Checks what's installed. After you agree, it installs only what's missing. |
|
|
98
|
+
| `setup` prompt | A guided setup: it checks, asks you, installs, then runs a spoken test. |
|
|
99
|
+
| `voice_mode` prompt | A hands-free session where Claude checks in by voice at natural points. |
|
|
187
100
|
|
|
188
|
-
|
|
101
|
+
Claude is told how to write for the ear: short sentences, no markdown, file names instead of paths. If code or links still get through, the server rewrites them before speaking. In Claude Code, the plugin also keeps a voice conversation in voice until you type. More in [Using it](docs/usage.md).
|
|
189
102
|
|
|
190
|
-
##
|
|
103
|
+
## Troubleshooting
|
|
191
104
|
|
|
192
|
-
|
|
105
|
+
Run `npx -y mac-voice-mcp@latest doctor` first. It tests setup, the voice and transcription, and your speakers and mic, and tells you which part fails.
|
|
193
106
|
|
|
194
|
-
|
|
|
107
|
+
| Symptom | Fix |
|
|
195
108
|
|---|---|
|
|
196
|
-
|
|
|
197
|
-
|
|
|
198
|
-
| The
|
|
199
|
-
|
|
|
200
|
-
| The plugin's `/mac-voice-mcp:talk`, `voice-help` skill and stay-in-voice hook | Claude Code, with the plugin |
|
|
201
|
-
|
|
202
|
-
The rules Claude is given:
|
|
203
|
-
|
|
204
|
-
```
|
|
205
|
-
- 1–3 short sentences, under ~40 words; lead with the outcome, then one question.
|
|
206
|
-
- Plain words only: no markdown, bullets, emoji, code, file paths, URLs, stack traces or tables.
|
|
207
|
-
- Describe code instead of reading it, say file names not paths, round numbers, spell out symbols.
|
|
208
|
-
- Ask one question at a time, answerable in a few words.
|
|
209
|
-
- Put the details (diffs, logs, links) in the on-screen reply, and say so out loud.
|
|
210
|
-
```
|
|
109
|
+
| "microphone returned pure digital silence" | Allow the mic for the app (Claude, Cursor or your terminal) in **System Settings → Privacy & Security → Microphone**, then restart it. |
|
|
110
|
+
| It cuts me off while I'm thinking | Set `VOICE_MCP_END_SILENCE_MS=1800`. |
|
|
111
|
+
| The voice sounds robotic | Download a Premium macOS voice or install Kokoro. See [Voices](docs/voices.md). |
|
|
112
|
+
| It asks for approval every turn | See [Allow voice turns without prompts](#allow-voice-turns-without-prompts). |
|
|
211
113
|
|
|
212
|
-
|
|
114
|
+
More in [Troubleshooting](docs/troubleshooting.md).
|
|
213
115
|
|
|
214
|
-
|
|
116
|
+
## Documentation
|
|
215
117
|
|
|
216
|
-
|
|
118
|
+
- [Install](docs/install.md): every app, the MCP Registry, installing from a clone, upgrading.
|
|
119
|
+
- [Voices](docs/voices.md): macOS Premium voices and the optional Kokoro voice.
|
|
120
|
+
- [Using it](docs/usage.md): voice mode, several sessions, getting Claude to sound natural.
|
|
121
|
+
- [Configuration](docs/configuration.md): every setting, and the speech models.
|
|
122
|
+
- [Troubleshooting](docs/troubleshooting.md): common problems and known limitations.
|
|
123
|
+
- [Privacy and security](docs/privacy-security.md): what stays on your Mac, and how the project is secured.
|
|
124
|
+
- [Development](docs/development.md): building, testing and releasing.
|
|
125
|
+
- [Changelog](CHANGELOG.md)
|
|
217
126
|
|
|
218
|
-
|
|
127
|
+
### 🤖 Vibe-coded
|
|
219
128
|
|
|
220
|
-
|
|
129
|
+
> This project was designed and written with Claude, in conversation. A human (me) steered it, tested it on a real Mac and checked the test suite, but most of the code was written by AI. Read the [Disclaimer](#disclaimer) before you rely on it.
|
|
221
130
|
|
|
222
|
-
|
|
223
|
-
|---|---|---|
|
|
224
|
-
| `VOICE_MCP_VOICE` | most natural installed | Unset: the best Premium or Enhanced voice installed for the language, else the system voice. Set a name, e.g. `Ava (Premium)`, `Daniel`, `Kanya` (list them with `say -v '?'`), or `default` to always use the system voice. |
|
|
225
|
-
| `VOICE_MCP_RATE` | system rate | Words per minute, e.g. `200`. |
|
|
226
|
-
| `VOICE_MCP_MAX_SPEAK_WORDS` | `120` | Longer text is cut at a sentence boundary ("the rest is on screen"). |
|
|
227
|
-
| `VOICE_MCP_CHIME` | `1` | Set to `0` to turn off the mic open/close sounds. |
|
|
228
|
-
| `VOICE_MCP_LOCK_WAIT_SECONDS` | `120` | How long a turn waits while another session on this Mac is using the mic. |
|
|
131
|
+
## Disclaimer
|
|
229
132
|
|
|
230
|
-
**
|
|
133
|
+
mac-voice-mcp is a free, open-source personal project. It's provided **as is, with no warranty of any kind**, and the authors aren't liable for any damage or loss from using it. See the [LICENSE](LICENSE) for the exact terms. In plain words:
|
|
231
134
|
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
236
|
-
| `VOICE_MCP_SPEECH_MARGIN_DB` | `12` | How much louder than room noise counts as speech. Raise it in noisy rooms. |
|
|
237
|
-
| `VOICE_MCP_MIN_SPEECH_DB` | `-48` | The quietest level that ever counts as speech (dBFS). |
|
|
238
|
-
| `VOICE_MCP_RECORDER` | `auto` | `sox` or `ffmpeg` (`ffmpeg` is macOS only). `auto` uses SoX, falling back to ffmpeg. |
|
|
239
|
-
| `VOICE_MCP_FFMPEG_DEVICE` | `:default` | Which input ffmpeg records from. `:default` follows System Settings; `:1` picks device 1 (list them with `ffmpeg -f avfoundation -list_devices true -i ""`). |
|
|
135
|
+
- **It turns on your microphone** whenever an AI assistant calls `speak_and_listen`, and plays sound through your speakers.
|
|
136
|
+
- **Speech recognition makes mistakes.** An assistant may act on a misheard word. Check what it heard before you let it do anything you can't undo, like deleting files, pushing code or sending messages.
|
|
137
|
+
- **It installs software on your Mac**, but only after you agree: Homebrew packages, and kokoro-js with npm if you choose the Kokoro voice. Those are separate projects, under their own licenses (below).
|
|
138
|
+
- **There's no guaranteed support.** Issues and pull requests are welcome, and fixed on a best-effort basis.
|
|
240
139
|
|
|
241
|
-
**
|
|
140
|
+
**Trademarks.** Apple, Mac, macOS and Siri are trademarks of Apple Inc. Claude is a trademark of Anthropic. mac-voice-mcp is an independent project, not affiliated with, sponsored or endorsed by Apple or Anthropic.
|
|
242
141
|
|
|
243
|
-
|
|
244
|
-
|---|---|---|
|
|
245
|
-
| `VOICE_MCP_WHISPER_MODEL` | `base.en` | Which model to use (see the table below). |
|
|
246
|
-
| `VOICE_MCP_LANGUAGE` | `en` for `*.en` models, otherwise `auto` | `en`, `th`, `ja`, `de`, … |
|
|
247
|
-
| `VOICE_MCP_WHISPER_PROMPT` | — | Words to bias toward: names, product terms, jargon. |
|
|
248
|
-
| `VOICE_MCP_WHISPER_MODEL_PATH` | — | Use this exact `ggml-*.bin` file. |
|
|
249
|
-
| `VOICE_MCP_MODEL_SEARCH_PATHS` | — | Extra folders to check for an existing model (`:`-separated). |
|
|
250
|
-
| `VOICE_MCP_WHISPER_SERVER` | `1` | Set to `0` to always use `whisper-cli`, with no warm server. |
|
|
251
|
-
| `VOICE_MCP_SERVER_IDLE_MINUTES` | `15` | How long the warm server stays up without use. |
|
|
252
|
-
| `VOICE_MCP_THREADS` | min(8, cores) | Number of whisper.cpp threads. |
|
|
253
|
-
| `VOICE_MCP_CACHE_DIR` | `~/.cache/mac-voice-mcp` | Where models are downloaded or symlinked. |
|
|
254
|
-
| `VOICE_MCP_DEBUG` | `0` | Verbose logs with per-turn timings, written to stderr. |
|
|
255
|
-
|
|
256
|
-
**Claude Code plugin hook.** Set this in the `env` block of `~/.claude/settings.json`, not in the server's config:
|
|
257
|
-
|
|
258
|
-
| Variable | Default | |
|
|
259
|
-
|---|---|---|
|
|
260
|
-
| `VOICE_MCP_STAY_IN_VOICE` | `1` | Set to `0` to stop the plugin's hook from sending Claude back to answer by voice. |
|
|
142
|
+
### Third-party software
|
|
261
143
|
|
|
262
|
-
|
|
144
|
+
mac-voice-mcp doesn't bundle any of these. It uses them when they're on your Mac, and installs them only with your consent. Each one comes under its own license.
|
|
263
145
|
|
|
264
|
-
|
|
|
146
|
+
| Software | Used for | License |
|
|
265
147
|
|---|---|---|
|
|
266
|
-
|
|
|
267
|
-
|
|
|
268
|
-
|
|
|
269
|
-
| `
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
|
274
|
-
|---|---|
|
|
275
|
-
| "voice-mcp is not set up yet" | Ask Claude to *set up voice*, or run `npx -y mac-voice-mcp@latest setup`. |
|
|
276
|
-
| "microphone returned pure digital silence" | macOS is blocking the mic for the host app. Go to **System Settings → Privacy & Security → Microphone**, enable Claude / Cursor / your terminal, then restart that app. If the message says it's recording with ffmpeg, the input device is the likelier cause: run `brew install sox`. |
|
|
277
|
-
| No permission prompt ever appears | Run `tccutil reset Microphone <bundle id>` and restart the app. Running `test` in Terminal only gives permission to Terminal, not to Claude Desktop. |
|
|
278
|
-
| It cuts me off while I'm thinking | Set `VOICE_MCP_END_SILENCE_MS=1800` (or up to `2500`). |
|
|
279
|
-
| It never stops listening | The room is too noisy for the defaults. Set `VOICE_MCP_SPEECH_MARGIN_DB=18`, or use a headset. |
|
|
280
|
-
| It hears its own voice | Use headphones, or turn the speaker volume down. It only listens after it finishes speaking, but echo can linger. |
|
|
281
|
-
| "Homebrew: not installed" | Install it from [brew.sh](https://brew.sh). It needs your password, so it can't run from Claude. Then run setup again. |
|
|
282
|
-
| It garbles names or jargon | Set `VOICE_MCP_WHISPER_PROMPT="Priya, Postgres, Kubernetes"`, or switch to `small.en`. |
|
|
283
|
-
| The voice sounds robotic | Download a Premium voice (see [Get a better voice](#then-set-up-and-allow-the-mic)). It's used automatically. |
|
|
284
|
-
| It asks for approval every turn | Add the tool to Claude Code's allow list: see [Allow voice turns without prompts](#allow-voice-turns-without-prompts). |
|
|
285
|
-
| "Another voice session on this Mac…" | Another Claude window, Claude Desktop or Cursor held the speaker and mic for over 2 minutes, which means one very long turn. End that conversation, then try again. |
|
|
286
|
-
| Claude keeps answering by voice after I'm done | Type anything, or say "stop voice mode". To switch the plugin's hook off entirely, see [Voice mode](#voice-mode). |
|
|
287
|
-
| Turns feel slow | Each result ends with a timing line, e.g. `spoke 3.1 s · listened 4.0 s · transcribed 0.3 s`. Ask Claude what it says. Transcribing should take well under a second. |
|
|
288
|
-
|
|
289
|
-
## Known limitations
|
|
290
|
-
|
|
291
|
-
- **You can't interrupt it.** It finishes speaking, then listens. Barge-in would mean listening while the speakers play, which needs headphones or echo cancellation.
|
|
292
|
-
- **It's macOS-first.** Linux works with SoX and espeak-ng. Windows is untested.
|
|
293
|
-
- **Turn-taking is based on loudness, not a speech model.** It adapts to background noise, but very noisy rooms, music or TV can confuse it. A headset helps, and so do the listening settings above.
|
|
294
|
-
- **One conversation at a time.** There's one speaker and one microphone, so turns from every session on the Mac are queued.
|
|
295
|
-
|
|
296
|
-
## Privacy and safety
|
|
297
|
-
|
|
298
|
-
- **Audio stays on your machine.** Recordings go to a temporary file that's deleted after each turn. The only network use is the one-time model download from Hugging Face.
|
|
299
|
-
- **The warm whisper server is local only.** It listens on `127.0.0.1` on a random port, and stops when idle or when this server exits.
|
|
300
|
-
- **Setup can only install known packages.** Its install list is fixed in the code (`sox`, `whisper-cpp`), so nothing Claude says can make it install anything else. It never uninstalls or modifies other software.
|
|
301
|
-
|
|
302
|
-
## Security
|
|
303
|
-
|
|
304
|
-
- **Secrets:** every push and pull request is scanned for leaked secrets with [TruffleHog](https://github.com/trufflesecurity/trufflehog), and the whole history is scanned before the first push. GitHub secret scanning with push protection is also on.
|
|
305
|
-
- **Dependencies:** [Dependabot](https://docs.github.com/code-security/dependabot) opens weekly update pull requests. CI fails on high-severity advisories (`npm audit`), and dependency review blocks pull requests that add vulnerable packages.
|
|
306
|
-
- **Code:** [CodeQL](https://codeql.github.com) runs with the `security-extended` queries.
|
|
307
|
-
- **Releases:** releases publish through [npm trusted publishing](https://docs.npmjs.com/trusted-publishers/), with no long-lived npm token and a signed provenance attestation for every version.
|
|
308
|
-
|
|
309
|
-
To report a vulnerability, see [SECURITY.md](SECURITY.md).
|
|
310
|
-
|
|
311
|
-
## Development
|
|
312
|
-
|
|
313
|
-
```bash
|
|
314
|
-
npm install
|
|
315
|
-
npm test # build + unit tests, end-to-end tests over MCP with stub binaries, and release-metadata checks
|
|
316
|
-
npm run audit # known-vulnerability and signature checks on dependencies
|
|
317
|
-
npm run setup # check what's installed; offers to install what's missing
|
|
318
|
-
npm run test:voice # one real speak → listen → transcribe turn
|
|
319
|
-
npm run inspect # MCP Inspector
|
|
320
|
-
```
|
|
321
|
-
|
|
322
|
-
| Module | Responsibility |
|
|
323
|
-
|---|---|
|
|
324
|
-
| `config.ts` | Environment settings, logging, the PATH fix-up for GUI apps |
|
|
325
|
-
| `speech-text.ts` | Rewriting screen text for speech, cleaning up transcripts (pure, unit-tested) |
|
|
326
|
-
| `endpointer.ts` | Turn-taking voice-activity detection (pure, unit-tested) |
|
|
327
|
-
| `audio.ts` | Text-to-speech, chimes, streaming mic capture |
|
|
328
|
-
| `model.ts` | Finding, symlinking or downloading the model |
|
|
329
|
-
| `stt.ts` | The warm `whisper-server` with orphan guard, and the `whisper-cli` fallback |
|
|
330
|
-
| `setup.ts` | Requirement checks and consent-based background installs |
|
|
331
|
-
| `lock.ts` | One voice turn at a time across every session on the Mac |
|
|
332
|
-
| `voice.ts`, `server.ts`, `index.ts` | The round trip, the MCP tools and prompts, and the CLI |
|
|
333
|
-
| `skills/`, `hooks/` | The Claude Code plugin's `/mac-voice-mcp:talk` and `:setup` commands, the `voice-help` skill, and the stay-in-voice hook |
|
|
334
|
-
|
|
335
|
-
The package installs two commands: `mac-voice-mcp` (the one `npx -y mac-voice-mcp` runs), and `voice-mcp`.
|
|
336
|
-
|
|
337
|
-
**Testing the plugin: don't start Claude Code inside this repo.** In this folder, `npx mac-voice-mcp@<this version>` finds the checkout itself instead of downloading the package, can't run it, and the server fails with `CONNECTION_CLOSED`. Start `claude` in any other folder. To try unreleased plugin files (skills, hooks), swap the marketplace to your checkout: run `/plugin marketplace remove mac-voice-mcp`, then `/plugin marketplace add /path/to/checkout`, and install at user scope. The server still comes from npm, so test server changes with `npm test` and `npm run test:voice`.
|
|
338
|
-
|
|
339
|
-
### Releasing
|
|
340
|
-
|
|
341
|
-
- **First release:** `bash publish.sh`. It asks before each public step and uses your own GitHub and npm logins. It creates the GitHub repo, publishes to npm, and lists the server in the [official MCP Registry](https://registry.modelcontextprotocol.io), which Smithery, Glama, PulseMCP and mcp.so pick up from.
|
|
342
|
-
- **Later releases:**
|
|
343
|
-
1. Add a section to `CHANGELOG.md` for the new version, in [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) format, dated today.
|
|
344
|
-
2. Run `npm version patch` (bug fixes), `npm version minor` (new features) or `npm version major` (breaking changes), per [semver](https://semver.org). It updates `package.json`, `package-lock.json`, `server.json` and the plugin (including the npm version the plugin runs), commits, and tags `v<version>`.
|
|
345
|
-
3. Run `git push --follow-tags`.
|
|
346
|
-
|
|
347
|
-
The Publish workflow then runs the tests, which also check that every version and the changelog entry match. It publishes to npm with provenance, lists the release in the MCP Registry, and creates a GitHub Release from the changelog section. It needs no secrets: npm and the registry both use GitHub's OIDC identity, once you've set the package's *Trusted Publisher* on npmjs.com (`publish.sh` prints the steps). If a step fails, fix the cause and use **Re-run failed jobs**. Steps that already finished are skipped.
|
|
148
|
+
| [SoX](https://sourceforge.net/projects/sox/) | Recording from the mic, Kokoro playback | GPL-2.0 |
|
|
149
|
+
| [whisper.cpp](https://github.com/ggml-org/whisper.cpp) | Speech-to-text | MIT |
|
|
150
|
+
| [Whisper models](https://huggingface.co/ggerganov/whisper.cpp) (OpenAI) | Speech-to-text | MIT |
|
|
151
|
+
| macOS voices (`say`) | The built-in voice | Apple's macOS license |
|
|
152
|
+
| [kokoro-js](https://github.com/hexgrad/kokoro) (optional) | The Kokoro voice | Apache-2.0 |
|
|
153
|
+
| [Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) model (optional) | The Kokoro voice | Apache-2.0 |
|
|
154
|
+
| [espeak-ng](https://github.com/espeak-ng/espeak-ng) (optional, inside kokoro-js) | Turning text into sounds for Kokoro | GPL-3.0 |
|
|
155
|
+
| [ffmpeg](https://ffmpeg.org) (fallback, never installed by setup) | Recording, if SoX is missing | LGPL-2.1 / GPL |
|
|
348
156
|
|
|
349
157
|
## License
|
|
350
158
|
|
|
351
|
-
MIT
|
|
159
|
+
[MIT](LICENSE) © Jeet
|