mac-voice-mcp 0.4.0 → 0.4.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +28 -1
- package/README.md +72 -290
- package/dist/audio.js +51 -16
- package/dist/config.js +20 -3
- package/dist/endpointer.js +14 -0
- package/dist/index.js +1 -1
- package/dist/pcm.js +17 -0
- package/dist/speech-text.js +29 -3
- package/dist/voice.js +10 -4
- package/package.json +2 -2
package/CHANGELOG.md
CHANGED
|
@@ -5,6 +5,31 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
|
|
|
5
5
|
|
|
6
6
|
## [Unreleased]
|
|
7
7
|
|
|
8
|
+
## [0.4.2] - 2026-10-11
|
|
9
|
+
|
|
10
|
+
### Changed
|
|
11
|
+
- **Docs and the voice-help skill now suggest `small.en` for noisy rooms,** with measured numbers: on real recordings with background noise it got 8% of words wrong instead of 12% with `base.en`, the same in a quiet room, at about 0.3 s per reply. `base.en` stays the default, so nobody gets a 466 MB download they didn't ask for.
|
|
12
|
+
|
|
13
|
+
### Fixed
|
|
14
|
+
- **`doctor` explains a voice that produces no audio.** Sometimes macOS `say` finishes without error but writes almost no sound (0.1 s for a whole sentence), for example when a voice's speech data is missing. `doctor` used to report that as "0% of words right". Now it tries once more, then says which voice wrote no audio and how to pick another (`VOICE_MCP_VOICE`). This was also behind the speech round trip failing on some CI runs, which now use a fixed classic voice.
|
|
15
|
+
- **A cough or a fan no longer counts as your answer.** For a sound with no words, whisper writes a note like `[gunshot]` (a cough), `[APPLAUSE]` (typing) or `[sound of running]` (a fan), and those notes were passed to Claude as if you'd said them. Now every sound note is dropped, and Claude is told it heard a sound but no words. On 10 noise-only clips (cough, typing, fan, hum, music, a door, breathing, silence), all 10 now come back as no speech.
|
|
16
|
+
- **Sentences whisper invents from videos are dropped**, such as "You can find the link in the description below." or "Thanks for watching!". Only whole sentences that nobody says to a coding assistant; "Thank you." and "Bye." stay.
|
|
17
|
+
|
|
18
|
+
## [0.4.1] - 2026-10-10
|
|
19
|
+
|
|
20
|
+
### Added
|
|
21
|
+
- **README: a Disclaimer, a Third-party software table and a trademark note.** They spell out what the tool does on your Mac (mic, installs with your consent), that speech recognition can mishear, the license of everything it uses but doesn't bundle, and that the project isn't affiliated with Apple or Anthropic.
|
|
22
|
+
|
|
23
|
+
### Changed
|
|
24
|
+
- **whisper hears developer talk better.** In English it now gets a one-sentence hint that the conversation is a developer talking to a coding assistant. On 60 test clips read by three voices, word errors fell from 3.4% to 2.3% on everyday sentences and from 3.8% to 3.1% on developer jargon ("rebase", "backend"), and nothing was invented in silence. `VOICE_MCP_WHISPER_PROMPT` replaces the hint with your own, and `none` turns it off.
|
|
25
|
+
- **A shorter README, in the style of popular MCP servers:** what it's for, features, a four-step quick start and the most common fixes. The full guides moved to [`docs/`](docs/): install, voices, usage, configuration, troubleshooting, privacy and security, and development.
|
|
26
|
+
|
|
27
|
+
### Fixed
|
|
28
|
+
- **The first words of a reply were often lost.** The "mic open" chime played before the recorder had actually started, and opening a mic takes a moment (longer with Bluetooth), so anyone who answered right at the chime lost a word or two. In a test with 20 real recordings, 12 began mid-word, and the cut-off starts led whisper to mishear or invent the rest ("Open a draft PR…" came back as "You can find the link in the description below."). Now the mic opens first and the chime plays once audio is flowing; the chime's own sound is ignored. Re-recorded with the fix, 2 of 20 began mid-word, and whisper's word errors on that voice fell from about 20% to 5% (and from 28% to 12% with background noise).
|
|
29
|
+
|
|
30
|
+
### Security
|
|
31
|
+
- **MCP SDK 1.32.1** ([GHSA-6qxp-vccf-f47h](https://github.com/advisories/GHSA-6qxp-vccf-f47h)). The advisory is in the SDK's OAuth client, which this server never uses (it only talks over stdio), but `npm audit` flagged every install.
|
|
32
|
+
|
|
8
33
|
## [0.4.0] - 2026-10-06
|
|
9
34
|
|
|
10
35
|
### Added
|
|
@@ -96,7 +121,9 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
|
|
|
96
121
|
- `voice_setup`: checks first, installs only what's missing, and only after you agree.
|
|
97
122
|
- The `setup` and `voice_mode` prompts.
|
|
98
123
|
|
|
99
|
-
[Unreleased]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.4.
|
|
124
|
+
[Unreleased]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.4.2...HEAD
|
|
125
|
+
[0.4.2]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.4.1...v0.4.2
|
|
126
|
+
[0.4.1]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.4.0...v0.4.1
|
|
100
127
|
[0.4.0]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.3.1...v0.4.0
|
|
101
128
|
[0.3.1]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.3.0...v0.3.1
|
|
102
129
|
[0.3.0]: https://github.com/jeet0007/mac-voice-mcp/compare/v0.2.0...v0.3.0
|
package/README.md
CHANGED
|
@@ -24,7 +24,7 @@
|
|
|
24
24
|
|
|
25
25
|
---
|
|
26
26
|
|
|
27
|
-
Claude
|
|
27
|
+
Claude speaks through your Mac's speakers, listens to your answer the way a person would, and gets back what you said as text. Speech recognition runs on your Mac, so no audio leaves your computer.
|
|
28
28
|
|
|
29
29
|
```
|
|
30
30
|
Claude ──speak_and_listen("Tests pass. Open the PR?")──▶ 🔊 "Tests pass. Open the PR?"
|
|
@@ -32,120 +32,47 @@ Claude ──speak_and_listen("Tests pass. Open the PR?")──▶ 🔊 "Tests
|
|
|
32
32
|
Claude ◀──────────────── "Yes, and tag Priya." ───────── whisper.cpp on your Mac
|
|
33
33
|
```
|
|
34
34
|
|
|
35
|
-
|
|
35
|
+
**Perfect for:**
|
|
36
36
|
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
> It's an independent project, not made or endorsed by Anthropic. It works with any MCP client, including Claude Desktop, Claude Code and Cursor.
|
|
42
|
-
|
|
43
|
-
## How it works
|
|
44
|
-
|
|
45
|
-
The server has **two tools and two prompts**:
|
|
46
|
-
|
|
47
|
-
| | What it does |
|
|
48
|
-
|---|---|
|
|
49
|
-
| `speak_and_listen` | Speaks a short message, listens for one conversational turn and returns the transcript. |
|
|
50
|
-
| `voice_setup` | Checks what's installed. After you agree, it installs only what's missing. |
|
|
51
|
-
| `/mcp__voice-mcp__setup` | A guided setup: it checks, asks you, installs, then runs a spoken test. |
|
|
52
|
-
| `/mcp__voice-mcp__voice_mode` | A hands-free session where Claude checks in by voice at natural points. |
|
|
37
|
+
- Long tasks: Claude works quietly and checks in out loud when it needs a decision, so you can step away from the screen.
|
|
38
|
+
- Quick reviews: a 20-second summary of a pull request, then "approve it?"
|
|
39
|
+
- Resting your eyes, or your wrists, after a long day at the keyboard.
|
|
40
|
+
- Thinking out loud: talking a problem through instead of typing it.
|
|
53
41
|
|
|
54
|
-
|
|
42
|
+
## Features
|
|
55
43
|
|
|
56
|
-
|
|
44
|
+
- 🔒 **On-device.** whisper.cpp transcribes on your Mac. Recordings are deleted after each turn.
|
|
45
|
+
- 💬 **Natural turn-taking.** A soft chime, then it waits for you to start, and hands back about a second after you stop. No push-to-talk.
|
|
46
|
+
- ⚡ **Fast replies.** A warm whisper.cpp server keeps the model loaded, so transcribing takes a fraction of a second on Apple Silicon.
|
|
47
|
+
- 🗣️ **A natural voice, if you want one.** Uses your best macOS voice, or the optional [Kokoro](docs/voices.md#the-kokoro-voice-optional) neural voice.
|
|
48
|
+
- ✋ **Asks before installing anything.** Setup checks what you already have and reuses it.
|
|
49
|
+
- 🪟 **One mic, many windows.** Claude Code, Claude Desktop and Cursor take turns instead of talking over each other.
|
|
50
|
+
- 📏 **Measured, not guessed.** `doctor` checks the whole pipeline on your Mac and scores it PASS, WARN or FAIL.
|
|
57
51
|
|
|
58
|
-
|
|
52
|
+
## Quick start
|
|
59
53
|
|
|
60
|
-
|
|
54
|
+
You need a Mac (Apple Silicon recommended) and **Node.js 22 or newer**.
|
|
61
55
|
|
|
62
|
-
|
|
56
|
+
**1. Install.**
|
|
63
57
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
<p>
|
|
67
|
-
<a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIm1hYy12b2ljZS1tY3BAbGF0ZXN0Il19"><img alt="Add to Cursor" src="https://cursor.com/deeplink/mcp-install-dark.svg" height="32"></a>
|
|
68
|
-
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
69
|
-
<a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D&quality=insiders"><img alt="Install in VS Code Insiders" src="https://img.shields.io/badge/VS_Code_Insiders-Install_Server-24bfa5?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="32"></a>
|
|
70
|
-
</p>
|
|
71
|
-
|
|
72
|
-
### From a marketplace
|
|
73
|
-
|
|
74
|
-
- **Claude Code plugin marketplace.** This repo is its own marketplace:
|
|
58
|
+
- **Claude Code** (recommended). Install the plugin:
|
|
75
59
|
```
|
|
76
60
|
/plugin marketplace add jeet0007/mac-voice-mcp
|
|
77
61
|
/plugin install mac-voice-mcp@mac-voice-mcp
|
|
78
62
|
```
|
|
79
|
-
|
|
80
|
-
- **The official MCP Registry.** It's listed as [`io.github.jeet0007/mac-voice-mcp`](https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.jeet0007/mac-voice-mcp). Apps and directories that read the registry pick it up from there. In VS Code, open the Extensions view (⇧⌘X), search `@mcp mac-voice`, and click **Install**. Smithery, Glama, PulseMCP and mcp.so copy the registry, so it shows up there too.
|
|
63
|
+
- **Cursor or VS Code:** <a href="https://cursor.com/en/install-mcp?name=voice-mcp&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIm1hYy12b2ljZS1tY3BAbGF0ZXN0Il19"><img alt="Add to Cursor" src="https://cursor.com/deeplink/mcp-install-dark.svg" height="24"></a> <a href="https://insiders.vscode.dev/redirect/mcp/install?name=voice-mcp&config=%7B%22type%22%3A%22stdio%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22mac-voice-mcp%40latest%22%5D%7D"><img alt="Install in VS Code" src="https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white" height="24"></a>
|
|
81
64
|
|
|
82
|
-
|
|
65
|
+
- **Claude Desktop and other apps:** add `npx -y mac-voice-mcp@latest` as a stdio MCP server. See [Install](docs/install.md#by-hand) for each app.
|
|
83
66
|
|
|
84
|
-
**Claude
|
|
67
|
+
**2. Set up.** Ask Claude to *"set up voice"* (or run `/mac-voice-mcp:setup`). It checks for SoX, whisper.cpp and a speech model, shows you what's missing, and installs it only after you say yes.
|
|
85
68
|
|
|
86
|
-
|
|
87
|
-
claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp@latest
|
|
88
|
-
```
|
|
69
|
+
**3. Allow the microphone.** The first time Claude listens, macOS asks. Click **Allow**.
|
|
89
70
|
|
|
90
|
-
**
|
|
91
|
-
|
|
92
|
-
```json
|
|
93
|
-
{
|
|
94
|
-
"mcpServers": {
|
|
95
|
-
"voice-mcp": {
|
|
96
|
-
"command": "npx",
|
|
97
|
-
"args": ["-y", "mac-voice-mcp@latest"]
|
|
98
|
-
}
|
|
99
|
-
}
|
|
100
|
-
}
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
`@latest` makes npx check for a new release each time the app starts. Without it, npx keeps running whichever version it cached first.
|
|
104
|
-
|
|
105
|
-
If you get `spawn npx ENOENT`, use the full path from `which npx`, e.g. `"command": "/opt/homebrew/bin/npx"`.
|
|
106
|
-
|
|
107
|
-
**Cursor.** Add the same `mcpServers` block to `~/.cursor/mcp.json`.
|
|
108
|
-
|
|
109
|
-
**VS Code.** Run **MCP: Add Server** from the Command Palette, or:
|
|
110
|
-
|
|
111
|
-
```bash
|
|
112
|
-
code --add-mcp '{"name":"voice-mcp","command":"npx","args":["-y","mac-voice-mcp@latest"]}'
|
|
113
|
-
```
|
|
114
|
-
|
|
115
|
-
**Any other MCP client.** Run `npx -y mac-voice-mcp@latest` as a stdio server.
|
|
116
|
-
|
|
117
|
-
### Then: set up and allow the mic
|
|
118
|
-
|
|
119
|
-
**Run setup once.** Ask Claude to *"set up voice"*. In Claude Code you can also run `/mac-voice-mcp:setup` (plugin) or `/mcp__voice-mcp__setup` (added by hand), or from a terminal run `npx -y mac-voice-mcp@latest setup`.
|
|
120
|
-
|
|
121
|
-
Setup checks what's already there before it changes anything:
|
|
122
|
-
|
|
123
|
-
| Needed | Provided by | If it's missing |
|
|
124
|
-
|---|---|---|
|
|
125
|
-
| Voice | macOS `say`, using the most natural voice installed | Nothing to do. For a far better voice, add a free Premium one (see below). |
|
|
126
|
-
| Microphone capture | SoX (`rec`). ffmpeg works as a fallback, but setup recommends SoX | `brew install sox` |
|
|
127
|
-
| Speech-to-text | whisper.cpp (`whisper-cli` and `whisper-server`, Metal-accelerated) | `brew install whisper-cpp` |
|
|
128
|
-
| Speech model | `base.en`, ~140 MB | Downloaded once to `~/.cache/mac-voice-mcp/models/` |
|
|
129
|
-
|
|
130
|
-
- **Nothing is redone.** Tools already on your PATH are used as they are. If the model is already somewhere on disk (a whisper.cpp checkout, Homebrew's share folder, another tool's cache, or anything Spotlight can find), it's **symlinked**, not downloaded again. `brew install` runs only for the missing formulae.
|
|
131
|
-
- **Nothing happens without your OK.** Claude calls `voice_setup` to check first, shows you the checklist, and asks before calling it with `install=true`.
|
|
132
|
-
- **Slow installs don't time out.** If `brew install whisper-cpp` takes a while, setup reports INSTALLING. The install carries on in the background, and the next check picks up the result.
|
|
133
|
-
|
|
134
|
-
**Allow the microphone.** The first time Claude listens, macOS asks whether Claude (or Cursor, or your terminal) can use the microphone. Click Allow.
|
|
135
|
-
|
|
136
|
-
**Get a better voice (recommended).** macOS includes free Premium voices that sound far more natural than the default. Open **System Settings → Accessibility → Spoken Content → System Voice → Manage Voices…**, and download one, for example English → *Ava (Premium)* or *Zoe (Premium)*. The next voice turn uses it automatically. To choose a specific voice, or keep the system voice, see `VOICE_MCP_VOICE` under [Configuration](#configuration).
|
|
137
|
-
|
|
138
|
-
**Or try the Kokoro voice (optional).** [Kokoro](https://huggingface.co/hexgrad/Kokoro-82M) is a small neural voice that sounds close to a person and runs entirely on your Mac. Ask Claude to *"install the Kokoro voice"*, or run `npx -y mac-voice-mcp@latest setup --kokoro`.
|
|
139
|
-
|
|
140
|
-
- **What it installs:** [kokoro-js](https://github.com/hexgrad/kokoro) with npm into `~/.cache/mac-voice-mcp/kokoro/`, and the model from Hugging Face, once. That's about 1 GB on disk, and nothing is bundled with this package.
|
|
141
|
-
- **How it's used:** once it's installed, every turn uses it, starting with `af_heart`. Pick another voice with `VOICE_MCP_KOKORO_VOICE` (see [Configuration](#configuration)). It speaks sentence by sentence, so it starts talking before the whole reply is generated. Each turn's timing line shows how soon the first sound came.
|
|
142
|
-
- **It can't leave you without a voice.** If Kokoro fails for any reason, the built-in voice takes over mid-sentence, and Claude tells you once.
|
|
143
|
-
- **To stop using it:** set `VOICE_MCP_TTS=say`, or delete `~/.cache/mac-voice-mcp/kokoro/`.
|
|
144
|
-
- **Limits:** English only (American and British voices). It speaks with its own pronunciation rules, which handle common developer words (JSON, `index.ts`, version numbers).
|
|
71
|
+
**4. Talk.** Run `/mac-voice-mcp:talk fix the flaky login test`, or ask Claude to *"check in with me by voice when you need a decision"*. Reply after the chime.
|
|
145
72
|
|
|
146
73
|
### Allow voice turns without prompts
|
|
147
74
|
|
|
148
|
-
|
|
75
|
+
Claude Code asks for approval every time Claude wants to speak, which breaks the flow of a conversation. To allow voice turns, add the tool to `permissions.allow` in `~/.claude/settings.json`:
|
|
149
76
|
|
|
150
77
|
```json
|
|
151
78
|
{
|
|
@@ -158,220 +85,75 @@ By default, Claude Code asks for approval every time Claude wants to speak, whic
|
|
|
158
85
|
}
|
|
159
86
|
```
|
|
160
87
|
|
|
161
|
-
The first name is for the plugin, the second for `claude mcp add voice-mcp …`. Leave `voice_setup` out, so installs still ask you first.
|
|
162
|
-
|
|
163
|
-
### Installing from a clone
|
|
88
|
+
The first name is for the plugin, the second for `claude mcp add voice-mcp …`. Leave `voice_setup` out, so installs still ask you first.
|
|
164
89
|
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
```bash
|
|
168
|
-
git clone https://github.com/jeet0007/mac-voice-mcp && bash mac-voice-mcp/install.sh
|
|
169
|
-
```
|
|
170
|
-
|
|
171
|
-
## Upgrading
|
|
172
|
-
|
|
173
|
-
- **Claude Code plugin:** run `/plugin marketplace update mac-voice-mcp`, then open `/plugin`, choose mac-voice-mcp under your installed plugins, and update it. Restart Claude Code. If there's no update option, uninstall and reinstall it.
|
|
174
|
-
- **Everything installed with `mac-voice-mcp@latest`** (Claude Desktop, Cursor, VS Code, `claude mcp add`): restart the app. npx fetches the new release when the server starts.
|
|
175
|
-
- **Configs without `@latest`:** change `mac-voice-mcp` to `mac-voice-mcp@latest` in the config, then restart the app. Otherwise npx keeps running the version it cached first.
|
|
176
|
-
|
|
177
|
-
Check which version you'd get with `npx -y mac-voice-mcp@latest --version`, and see what changed in the [changelog](CHANGELOG.md). Your model, voice and settings carry over.
|
|
178
|
-
|
|
179
|
-
## Using it
|
|
180
|
-
|
|
181
|
-
- **"Work on X and check in with me by voice when you need a decision."** Claude works quietly and only speaks at decision points.
|
|
182
|
-
- **`/mac-voice-mcp:talk fix the flaky login test`** (plugin), or **`/mcp__voice-mcp__voice_mode fix the flaky login test`** (added by hand). Claude reads its plan back to you, then checks in at each checkpoint. Say "stop voice mode" or "I'm back" to end it.
|
|
183
|
-
- **"Read me a 20-second summary of this PR and ask if I should approve it."** Use this for one-off briefings.
|
|
184
|
-
- **Just talk after the chime.** You don't need to hurry or fill silence. If you're still talking at the 30-second safety cap (`listen_seconds`), Claude is told your reply may be cut off and asks you to continue.
|
|
185
|
-
|
|
186
|
-
### Voice mode
|
|
187
|
-
|
|
188
|
-
Once you answer out loud, you're in a voice conversation. Claude replies by voice, not in text, until one of these happens:
|
|
90
|
+
## How it works
|
|
189
91
|
|
|
190
|
-
|
|
191
|
-
- **You say you're done.** Claude says a short goodbye without opening the mic (`speak_and_listen` with `listen: false`).
|
|
192
|
-
- **You don't answer.** After 15 seconds of silence Claude asks once more. If you still don't answer, it pauses and summarizes on screen, and the mic stays off. Type anything, or run `/mac-voice-mcp:talk`, to pick up again.
|
|
92
|
+
The server has two tools and two prompts:
|
|
193
93
|
|
|
194
|
-
|
|
94
|
+
| | What it does |
|
|
95
|
+
|---|---|
|
|
96
|
+
| `speak_and_listen` | Speaks a short message, listens for one conversational turn and returns the transcript. |
|
|
97
|
+
| `voice_setup` | Checks what's installed. After you agree, it installs only what's missing. |
|
|
98
|
+
| `setup` prompt | A guided setup: it checks, asks you, installs, then runs a spoken test. |
|
|
99
|
+
| `voice_mode` prompt | A hands-free session where Claude checks in by voice at natural points. |
|
|
195
100
|
|
|
196
|
-
|
|
101
|
+
Claude is told how to write for the ear: short sentences, no markdown, file names instead of paths. If code or links still get through, the server rewrites them before speaking. In Claude Code, the plugin also keeps a voice conversation in voice until you type. More in [Using it](docs/usage.md).
|
|
197
102
|
|
|
198
|
-
##
|
|
103
|
+
## Troubleshooting
|
|
199
104
|
|
|
200
|
-
|
|
105
|
+
Run `npx -y mac-voice-mcp@latest doctor` first. It tests setup, the voice and transcription, and your speakers and mic, and tells you which part fails.
|
|
201
106
|
|
|
202
|
-
|
|
|
107
|
+
| Symptom | Fix |
|
|
203
108
|
|---|---|
|
|
204
|
-
|
|
|
205
|
-
|
|
|
206
|
-
| The
|
|
207
|
-
|
|
|
208
|
-
| The plugin's `/mac-voice-mcp:talk`, `voice-help` skill and stay-in-voice hook | Claude Code, with the plugin |
|
|
109
|
+
| "microphone returned pure digital silence" | Allow the mic for the app (Claude, Cursor or your terminal) in **System Settings → Privacy & Security → Microphone**, then restart it. |
|
|
110
|
+
| It cuts me off while I'm thinking | Set `VOICE_MCP_END_SILENCE_MS=1800`. |
|
|
111
|
+
| The voice sounds robotic | Download a Premium macOS voice or install Kokoro. See [Voices](docs/voices.md). |
|
|
112
|
+
| It asks for approval every turn | See [Allow voice turns without prompts](#allow-voice-turns-without-prompts). |
|
|
209
113
|
|
|
210
|
-
|
|
114
|
+
More in [Troubleshooting](docs/troubleshooting.md).
|
|
211
115
|
|
|
212
|
-
|
|
213
|
-
- 1–3 short sentences, under ~40 words; lead with the outcome, then one question.
|
|
214
|
-
- Plain words only: no markdown, bullets, emoji, code, file paths, URLs, stack traces or tables.
|
|
215
|
-
- Describe code instead of reading it, say file names not paths, round numbers, spell out symbols.
|
|
216
|
-
- Ask one question at a time, answerable in a few words.
|
|
217
|
-
- Put the details (diffs, logs, links) in the on-screen reply, and say so out loud.
|
|
218
|
-
```
|
|
116
|
+
## Documentation
|
|
219
117
|
|
|
220
|
-
|
|
118
|
+
- [Install](docs/install.md): every app, the MCP Registry, installing from a clone, upgrading.
|
|
119
|
+
- [Voices](docs/voices.md): macOS Premium voices and the optional Kokoro voice.
|
|
120
|
+
- [Using it](docs/usage.md): voice mode, several sessions, getting Claude to sound natural.
|
|
121
|
+
- [Configuration](docs/configuration.md): every setting, and the speech models.
|
|
122
|
+
- [Troubleshooting](docs/troubleshooting.md): common problems and known limitations.
|
|
123
|
+
- [Privacy and security](docs/privacy-security.md): what stays on your Mac, and how the project is secured.
|
|
124
|
+
- [Development](docs/development.md): building, testing and releasing.
|
|
125
|
+
- [Changelog](CHANGELOG.md)
|
|
221
126
|
|
|
222
|
-
|
|
127
|
+
### 🤖 Vibe-coded
|
|
223
128
|
|
|
224
|
-
|
|
129
|
+
> This project was designed and written with Claude, in conversation. A human (me) steered it, tested it on a real Mac and checked the test suite, but most of the code was written by AI. Read the [Disclaimer](#disclaimer) before you rely on it.
|
|
225
130
|
|
|
226
|
-
|
|
131
|
+
## Disclaimer
|
|
227
132
|
|
|
228
|
-
**
|
|
133
|
+
mac-voice-mcp is a free, open-source personal project. It's provided **as is, with no warranty of any kind**, and the authors aren't liable for any damage or loss from using it. See the [LICENSE](LICENSE) for the exact terms. In plain words:
|
|
229
134
|
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
| `VOICE_MCP_TTS` | `auto` | `auto`: the [Kokoro voice](#then-set-up-and-allow-the-mic) once it's installed, else the built-in voice. `say`: always the built-in voice. `kokoro`: Kokoro, and setup offers to install it. |
|
|
235
|
-
| `VOICE_MCP_KOKORO_VOICE` | `af_heart` | Kokoro voice. American English starts with `a`, British with `b`, e.g. `af_bella`, `am_michael`, `bf_emma`, `bm_george`. |
|
|
236
|
-
| `VOICE_MCP_KOKORO_SPEED` | `1` | Kokoro speaking speed, `0.5` to `2`. |
|
|
237
|
-
| `VOICE_MCP_KOKORO_DTYPE` | `fp32` | `q8`: a smaller model (~90 MB instead of ~330 MB) that's about half as fast. |
|
|
238
|
-
| `VOICE_MCP_KOKORO_DIR` | `~/.cache/mac-voice-mcp/kokoro` | Where the Kokoro voice is installed. |
|
|
239
|
-
| `VOICE_MCP_MAX_SPEAK_WORDS` | `120` | Longer text is cut at a sentence boundary ("the rest is on screen"). |
|
|
240
|
-
| `VOICE_MCP_CHIME` | `1` | Set to `0` to turn off the mic open/close sounds. |
|
|
241
|
-
| `VOICE_MCP_LOCK_WAIT_SECONDS` | `120` | How long a turn waits while another session on this Mac is using the mic. |
|
|
242
|
-
|
|
243
|
-
**Listening**
|
|
244
|
-
|
|
245
|
-
| Variable | Default | |
|
|
246
|
-
|---|---|---|
|
|
247
|
-
| `VOICE_MCP_END_SILENCE_MS` | `1200` | How long a pause ends your turn. Use `1800` if it cuts you off while you think, `800` for snappier replies. |
|
|
248
|
-
| `VOICE_MCP_START_TIMEOUT_SECONDS` | `15` | How long to wait for you to start talking. |
|
|
249
|
-
| `VOICE_MCP_SPEECH_MARGIN_DB` | `12` | How much louder than room noise counts as speech. Raise it in noisy rooms. |
|
|
250
|
-
| `VOICE_MCP_MIN_SPEECH_DB` | `-48` | The quietest level that ever counts as speech (dBFS). |
|
|
251
|
-
| `VOICE_MCP_RECORDER` | `auto` | `sox` or `ffmpeg` (`ffmpeg` is macOS only). `auto` uses SoX, falling back to ffmpeg. |
|
|
252
|
-
| `VOICE_MCP_FFMPEG_DEVICE` | `:default` | Which input ffmpeg records from. `:default` follows System Settings; `:1` picks device 1 (list them with `ffmpeg -f avfoundation -list_devices true -i ""`). |
|
|
135
|
+
- **It turns on your microphone** whenever an AI assistant calls `speak_and_listen`, and plays sound through your speakers.
|
|
136
|
+
- **Speech recognition makes mistakes.** An assistant may act on a misheard word. Check what it heard before you let it do anything you can't undo, like deleting files, pushing code or sending messages.
|
|
137
|
+
- **It installs software on your Mac**, but only after you agree: Homebrew packages, and kokoro-js with npm if you choose the Kokoro voice. Those are separate projects, under their own licenses (below).
|
|
138
|
+
- **There's no guaranteed support.** Issues and pull requests are welcome, and fixed on a best-effort basis.
|
|
253
139
|
|
|
254
|
-
**
|
|
140
|
+
**Trademarks.** Apple, Mac, macOS and Siri are trademarks of Apple Inc. Claude is a trademark of Anthropic. mac-voice-mcp is an independent project, not affiliated with, sponsored or endorsed by Apple or Anthropic.
|
|
255
141
|
|
|
256
|
-
|
|
257
|
-
|---|---|---|
|
|
258
|
-
| `VOICE_MCP_WHISPER_MODEL` | `base.en` | Which model to use (see the table below). |
|
|
259
|
-
| `VOICE_MCP_LANGUAGE` | `en` for `*.en` models, otherwise `auto` | `en`, `th`, `ja`, `de`, … |
|
|
260
|
-
| `VOICE_MCP_WHISPER_PROMPT` | — | Words to bias toward: names, product terms, jargon. |
|
|
261
|
-
| `VOICE_MCP_WHISPER_MODEL_PATH` | — | Use this exact `ggml-*.bin` file. |
|
|
262
|
-
| `VOICE_MCP_MODEL_SEARCH_PATHS` | — | Extra folders to check for an existing model (`:`-separated). |
|
|
263
|
-
| `VOICE_MCP_WHISPER_SERVER` | `1` | Set to `0` to always use `whisper-cli`, with no warm server. |
|
|
264
|
-
| `VOICE_MCP_SERVER_IDLE_MINUTES` | `15` | How long the warm server stays up without use. |
|
|
265
|
-
| `VOICE_MCP_THREADS` | min(8, cores) | Number of whisper.cpp threads. |
|
|
266
|
-
| `VOICE_MCP_CACHE_DIR` | `~/.cache/mac-voice-mcp` | Where models are downloaded or symlinked. |
|
|
267
|
-
| `VOICE_MCP_DEBUG` | `0` | Verbose logs with per-turn timings, written to stderr. |
|
|
268
|
-
|
|
269
|
-
**Claude Code plugin hook.** Set this in the `env` block of `~/.claude/settings.json`, not in the server's config:
|
|
270
|
-
|
|
271
|
-
| Variable | Default | |
|
|
272
|
-
|---|---|---|
|
|
273
|
-
| `VOICE_MCP_STAY_IN_VOICE` | `1` | Set to `0` to stop the plugin's hook from sending Claude back to answer by voice. |
|
|
142
|
+
### Third-party software
|
|
274
143
|
|
|
275
|
-
|
|
144
|
+
mac-voice-mcp doesn't bundle any of these. It uses them when they're on your Mac, and installs them only with your consent. Each one comes under its own license.
|
|
276
145
|
|
|
277
|
-
|
|
|
146
|
+
| Software | Used for | License |
|
|
278
147
|
|---|---|---|
|
|
279
|
-
|
|
|
280
|
-
|
|
|
281
|
-
|
|
|
282
|
-
| `
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
|
287
|
-
|---|---|
|
|
288
|
-
| Not sure what's wrong | Run `npx -y mac-voice-mcp@latest doctor`. It checks setup, the voice and transcription, and your speakers and mic, and says which part fails. |
|
|
289
|
-
| "voice-mcp is not set up yet" | Ask Claude to *set up voice*, or run `npx -y mac-voice-mcp@latest setup`. |
|
|
290
|
-
| "microphone returned pure digital silence" | macOS is blocking the mic for the host app. Go to **System Settings → Privacy & Security → Microphone**, enable Claude / Cursor / your terminal, then restart that app. If the message says it's recording with ffmpeg, the input device is the likelier cause: run `brew install sox`. |
|
|
291
|
-
| No permission prompt ever appears | Run `tccutil reset Microphone <bundle id>` and restart the app. Running `test` in Terminal only gives permission to Terminal, not to Claude Desktop. |
|
|
292
|
-
| It cuts me off while I'm thinking | Set `VOICE_MCP_END_SILENCE_MS=1800` (or up to `2500`). |
|
|
293
|
-
| It never stops listening | The room is too noisy for the defaults. Set `VOICE_MCP_SPEECH_MARGIN_DB=18`, or use a headset. |
|
|
294
|
-
| It hears its own voice | Use headphones, or turn the speaker volume down. It only listens after it finishes speaking, but echo can linger. |
|
|
295
|
-
| "Homebrew: not installed" | Install it from [brew.sh](https://brew.sh). It needs your password, so it can't run from Claude. Then run setup again. |
|
|
296
|
-
| It garbles names or jargon | Set `VOICE_MCP_WHISPER_PROMPT="Priya, Postgres, Kubernetes"`, or switch to `small.en`. |
|
|
297
|
-
| The voice sounds robotic | Download a Premium voice, or install the Kokoro voice (see [Get a better voice](#then-set-up-and-allow-the-mic)). Either is used automatically. |
|
|
298
|
-
| "The Kokoro voice didn't work" | The built-in voice spoke instead. Ask Claude to *check voice setup*, or run `npx -y mac-voice-mcp@latest setup`. To reinstall it, delete `~/.cache/mac-voice-mcp/kokoro/` and run `setup --kokoro`. |
|
|
299
|
-
| It asks for approval every turn | Add the tool to Claude Code's allow list: see [Allow voice turns without prompts](#allow-voice-turns-without-prompts). |
|
|
300
|
-
| "Another voice session on this Mac…" | Another Claude window, Claude Desktop or Cursor held the speaker and mic for over 2 minutes, which means one very long turn. End that conversation, then try again. |
|
|
301
|
-
| Claude keeps answering by voice after I'm done | Type anything, or say "stop voice mode". To switch the plugin's hook off entirely, see [Voice mode](#voice-mode). |
|
|
302
|
-
| Turns feel slow | Each result ends with a timing line, e.g. `spoke 3.1 s · listened 4.0 s · transcribed 0.3 s`. Ask Claude what it says. Transcribing should take well under a second. |
|
|
303
|
-
|
|
304
|
-
## Known limitations
|
|
305
|
-
|
|
306
|
-
- **You can't interrupt it.** It finishes speaking, then listens. Barge-in would mean listening while the speakers play, which needs headphones or echo cancellation.
|
|
307
|
-
- **It's macOS-first.** Linux works with SoX and espeak-ng. Windows is untested.
|
|
308
|
-
- **Turn-taking is based on loudness, not a speech model.** It adapts to background noise, but very noisy rooms, music or TV can confuse it. A headset helps, and so do the listening settings above.
|
|
309
|
-
- **One conversation at a time.** There's one speaker and one microphone, so turns from every session on the Mac are queued.
|
|
310
|
-
|
|
311
|
-
## Privacy and safety
|
|
312
|
-
|
|
313
|
-
- **Audio stays on your machine.** Recordings go to a temporary file that's deleted after each turn. The only network use is the one-time model download from Hugging Face, plus npm and Hugging Face once more if you install the Kokoro voice. Speaking never downloads anything.
|
|
314
|
-
- **The warm whisper server is local only.** It listens on `127.0.0.1` on a random port, and stops when idle or when this server exits.
|
|
315
|
-
- **Setup can only install known packages.** Its install list is fixed in the code (`sox`, `whisper-cpp`, and `kokoro-js@1.2.1` for the optional voice), so nothing Claude says can make it install anything else. It never uninstalls or modifies other software.
|
|
316
|
-
- **Licenses of the optional Kokoro voice.** The Kokoro model and kokoro-js are Apache-2.0. kokoro-js turns text into sounds with a WebAssembly build of espeak-ng (GPL-3.0), through the `phonemizer` package. None of this is part of mac-voice-mcp (MIT). It's installed on your Mac only if you ask for it.
|
|
317
|
-
|
|
318
|
-
## Security
|
|
319
|
-
|
|
320
|
-
- **Secrets:** every push and pull request is scanned for leaked secrets with [TruffleHog](https://github.com/trufflesecurity/trufflehog), and the whole history is scanned before the first push. GitHub secret scanning with push protection is also on.
|
|
321
|
-
- **Dependencies:** [Dependabot](https://docs.github.com/code-security/dependabot) opens weekly update pull requests. CI fails on high-severity advisories (`npm audit`), and dependency review blocks pull requests that add vulnerable packages.
|
|
322
|
-
- **Code:** [CodeQL](https://codeql.github.com) runs with the `security-extended` queries.
|
|
323
|
-
- **Releases:** releases publish through [npm trusted publishing](https://docs.npmjs.com/trusted-publishers/), with no long-lived npm token and a signed provenance attestation for every version.
|
|
324
|
-
|
|
325
|
-
To report a vulnerability, see [SECURITY.md](SECURITY.md).
|
|
326
|
-
|
|
327
|
-
## Development
|
|
328
|
-
|
|
329
|
-
```bash
|
|
330
|
-
npm install
|
|
331
|
-
npm test # build + unit tests, end-to-end tests over MCP with stub binaries, and release-metadata checks
|
|
332
|
-
npm run audit # known-vulnerability and signature checks on dependencies
|
|
333
|
-
npm run setup # check what's installed; offers to install what's missing
|
|
334
|
-
npm run test:voice # one real speak → listen → transcribe turn
|
|
335
|
-
npm run doctor # objective self-check on this Mac: PASS / WARN / FAIL per stage, JSON report
|
|
336
|
-
npm run dev:plugin # try this checkout in Claude Code as the "mac-voice-mcp-dev" plugin
|
|
337
|
-
npm run inspect # MCP Inspector
|
|
338
|
-
```
|
|
339
|
-
|
|
340
|
-
| Module | Responsibility |
|
|
341
|
-
|---|---|
|
|
342
|
-
| `config.ts` | Environment settings, logging, the PATH fix-up for GUI apps |
|
|
343
|
-
| `speech-text.ts` | Rewriting screen text for speech, cleaning up transcripts (pure, unit-tested) |
|
|
344
|
-
| `endpointer.ts` | Turn-taking voice-activity detection (pure, unit-tested) |
|
|
345
|
-
| `audio.ts` | Text-to-speech, chimes, streaming mic capture |
|
|
346
|
-
| `kokoro.ts`, `kokoro-worker.ts` | The optional Kokoro voice: install, a warm worker process, sentence-by-sentence playback, fallback |
|
|
347
|
-
| `kokoro-text.ts`, `pcm.ts` | Pronunciation fixes and sentence chunks for Kokoro, PCM conversion (pure, unit-tested) |
|
|
348
|
-
| `model.ts` | Finding, symlinking or downloading the model |
|
|
349
|
-
| `stt.ts` | The warm `whisper-server` with orphan guard, and the `whisper-cli` fallback |
|
|
350
|
-
| `setup.ts` | Requirement checks and consent-based background installs |
|
|
351
|
-
| `lock.ts` | One voice turn at a time across every session on the Mac |
|
|
352
|
-
| `voice.ts`, `server.ts`, `index.ts` | The round trip, the MCP tools and prompts, and the CLI |
|
|
353
|
-
| `skills/`, `hooks/` | The Claude Code plugin's `/mac-voice-mcp:talk` and `:setup` commands, the `voice-help` skill, and the stay-in-voice hook |
|
|
354
|
-
|
|
355
|
-
The package installs two commands: `mac-voice-mcp` (the one `npx -y mac-voice-mcp` runs), and `voice-mcp`.
|
|
356
|
-
|
|
357
|
-
**Testing changes.** Three layers; none of them depend on anyone's ears.
|
|
358
|
-
|
|
359
|
-
1. **`npm test`** runs everywhere, with stub binaries. It covers the MCP tools, turn-taking, setup, the mic lock, the plugin hook and the release metadata.
|
|
360
|
-
2. **The "Speech round trip" CI job** runs on a real Mac. Our own `setup --install` puts in SoX, whisper.cpp and the model. Then `doctor --no-loopback` runs, and `test/roundtrip.test.mjs` speaks known sentences with the real voice and transcribes them with the real whisper.cpp. The build fails if too many words come back wrong or transcription is too slow. The same job then installs the Kokoro voice with `setup --kokoro` and repeats the round trip with it, which also checks how soon Kokoro starts talking. GitHub's Macs have only basic voices and no GPU for whisper.cpp, so the limits there are looser: it guards against breakage, and doctor on a real Mac is the quality bar. The numbers are kept as a build artifact.
|
|
361
|
-
3. **`npm run doctor`** on your own Mac adds the one thing CI can't test, your speakers and microphone. It plays a sentence through the speakers, records it and transcribes it. Each stage gets PASS, WARN or FAIL against fixed limits, and the report is saved under `~/.cache/mac-voice-mcp/doctor/`.
|
|
362
|
-
|
|
363
|
-
**Trying the plugin from a checkout:** run `npm run dev:plugin`, then the `claude --plugin-dir …` command it prints. This loads a throwaway plugin, "mac-voice-mcp-dev", that runs this checkout's build with `node`. Its own name means it doesn't clash with an installed mac-voice-mcp. It also works inside this repo, where `npx mac-voice-mcp@<this version>` would find the checkout instead of the package and fail with `CONNECTION_CLOSED`.
|
|
364
|
-
|
|
365
|
-
### Releasing
|
|
366
|
-
|
|
367
|
-
- **First release:** `bash publish.sh`. It asks before each public step and uses your own GitHub and npm logins. It creates the GitHub repo, publishes to npm, and lists the server in the [official MCP Registry](https://registry.modelcontextprotocol.io), which Smithery, Glama, PulseMCP and mcp.so pick up from.
|
|
368
|
-
- **Later releases:**
|
|
369
|
-
1. Add a section to `CHANGELOG.md` for the new version, in [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) format, dated today.
|
|
370
|
-
2. Run `npm version patch` (bug fixes), `npm version minor` (new features) or `npm version major` (breaking changes), per [semver](https://semver.org). It updates `package.json`, `package-lock.json`, `server.json` and the plugin (including the npm version the plugin runs), commits, and tags `v<version>`.
|
|
371
|
-
3. Run `git push --follow-tags`.
|
|
372
|
-
|
|
373
|
-
The Publish workflow then runs the tests, which also check that every version and the changelog entry match. It publishes to npm with provenance, lists the release in the MCP Registry, and creates a GitHub Release from the changelog section. It needs no secrets: npm and the registry both use GitHub's OIDC identity, once you've set the package's *Trusted Publisher* on npmjs.com (`publish.sh` prints the steps). If a step fails, fix the cause and use **Re-run failed jobs**. Steps that already finished are skipped.
|
|
148
|
+
| [SoX](https://sourceforge.net/projects/sox/) | Recording from the mic, Kokoro playback | GPL-2.0 |
|
|
149
|
+
| [whisper.cpp](https://github.com/ggml-org/whisper.cpp) | Speech-to-text | MIT |
|
|
150
|
+
| [Whisper models](https://huggingface.co/ggerganov/whisper.cpp) (OpenAI) | Speech-to-text | MIT |
|
|
151
|
+
| macOS voices (`say`) | The built-in voice | Apple's macOS license |
|
|
152
|
+
| [kokoro-js](https://github.com/hexgrad/kokoro) (optional) | The Kokoro voice | Apache-2.0 |
|
|
153
|
+
| [Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) model (optional) | The Kokoro voice | Apache-2.0 |
|
|
154
|
+
| [espeak-ng](https://github.com/espeak-ng/espeak-ng) (optional, inside kokoro-js) | Turning text into sounds for Kokoro | GPL-3.0 |
|
|
155
|
+
| [ffmpeg](https://ffmpeg.org) (fallback, never installed by setup) | Recording, if SoX is missing | LGPL-2.1 / GPL |
|
|
374
156
|
|
|
375
157
|
## License
|
|
376
158
|
|
|
377
|
-
MIT
|
|
159
|
+
[MIT](LICENSE) © Jeet
|
package/dist/audio.js
CHANGED
|
@@ -1,10 +1,11 @@
|
|
|
1
1
|
/** Speaking (native TTS), the mic chimes, and listening for one conversational turn. */
|
|
2
2
|
import { spawn } from "node:child_process";
|
|
3
3
|
import { existsSync } from "node:fs";
|
|
4
|
-
import { writeFile } from "node:fs/promises";
|
|
4
|
+
import { readFile, writeFile } from "node:fs/promises";
|
|
5
5
|
import { CONFIG, debug, IS_MAC, IS_WIN, log } from "./config.js";
|
|
6
|
-
import { Endpointer, FRAME_BYTES, wavHeader } from "./endpointer.js";
|
|
6
|
+
import { Endpointer, FRAME_BYTES, FRAME_MS, wavHeader } from "./endpointer.js";
|
|
7
7
|
import { KokoroError, kokoroEnabled, speakKokoro, synthesizeKokoroToFile } from "./kokoro.js";
|
|
8
|
+
import { wavSeconds } from "./pcm.js";
|
|
8
9
|
import { activeChildren, CancelledError, run, SetupError, tail, which } from "./proc.js";
|
|
9
10
|
// ---------------------------------------------------------------------------
|
|
10
11
|
// Text-to-speech
|
|
@@ -138,20 +139,33 @@ export async function synthesizeToFile(text, outFile) {
|
|
|
138
139
|
const bin = findTts();
|
|
139
140
|
if (!bin)
|
|
140
141
|
throw new SetupError("No text-to-speech engine found.");
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
142
|
+
const { voice } = IS_MAC ? await chooseVoice() : { voice: undefined };
|
|
143
|
+
const once = async () => {
|
|
144
|
+
let result;
|
|
145
|
+
if (IS_MAC) {
|
|
146
|
+
const args = [...(voice ? ["-v", voice] : []), "-o", outFile, "--file-format=WAVE", "--data-format=LEI16@16000", "-f", "-"];
|
|
147
|
+
result = await run(bin, args, { input: text, timeoutMs: 60_000 });
|
|
148
|
+
}
|
|
149
|
+
else if (/espeak/.test(bin)) {
|
|
150
|
+
result = await run(bin, ["-w", outFile, "--stdin"], { input: text, timeoutMs: 60_000 });
|
|
151
|
+
}
|
|
152
|
+
else {
|
|
153
|
+
throw new SetupError("Writing speech to a file needs macOS `say` or espeak-ng.");
|
|
154
|
+
}
|
|
155
|
+
if (result.code !== 0)
|
|
156
|
+
throw new Error(`Text-to-speech to file failed (exit ${result.code}): ${tail(result.stderr) || "no output"}`);
|
|
157
|
+
return wavSeconds(await readFile(outFile));
|
|
158
|
+
};
|
|
159
|
+
// Sometimes `say` exits fine but writes (almost) no audio, e.g. when the voice's speech data
|
|
160
|
+
// isn't available. Try once more, then say so plainly instead of handing silence to whisper.
|
|
161
|
+
const expected = text.trim().split(/\s+/).length > 2;
|
|
162
|
+
let seconds = await once();
|
|
163
|
+
if (expected && seconds < 0.3)
|
|
164
|
+
seconds = await once();
|
|
165
|
+
if (expected && seconds < 0.3) {
|
|
166
|
+
throw new Error(`Text-to-speech wrote only ${seconds.toFixed(2)} s of audio for "${text.slice(0, 40)}…" with ${voice ? `the voice "${voice}"` : "the system voice"}. ` +
|
|
167
|
+
"That voice may be missing its speech data: pick another with VOICE_MCP_VOICE (list them with `say -v '?'`).");
|
|
152
168
|
}
|
|
153
|
-
if (result.code !== 0)
|
|
154
|
-
throw new Error(`Text-to-speech to file failed (exit ${result.code}): ${tail(result.stderr) || "no output"}`);
|
|
155
169
|
return { engine: "built-in" };
|
|
156
170
|
}
|
|
157
171
|
/** A short, quiet cue that the mic just opened ("start") or closed ("stop"). macOS only. */
|
|
@@ -260,13 +274,19 @@ export async function recordForSeconds(seconds, outFile, onStarted) {
|
|
|
260
274
|
}
|
|
261
275
|
return { digitalSilence: !anyNonZero, peakDb: peak > 0 ? 10 * Math.log10(peak / 32768 ** 2) : -Infinity };
|
|
262
276
|
}
|
|
277
|
+
/** Audio the mic must have delivered before we tell the user it's open: the recorder is running, and the room's level is known. */
|
|
278
|
+
export const MIC_WARMUP_MS = 250;
|
|
263
279
|
/**
|
|
264
280
|
* Listen for one conversational turn and write it to `outFile` (16 kHz mono WAV).
|
|
265
281
|
* Streams raw PCM from the recorder, runs the endpointer on it live, and stops the
|
|
266
282
|
* recorder the moment the user finishes. Leading and trailing silence are trimmed
|
|
267
283
|
* (keeping a little padding) so whisper gets just the utterance.
|
|
284
|
+
*
|
|
285
|
+
* The recorder starts first and `onListening` (the chime) comes only once audio is flowing, so
|
|
286
|
+
* someone who starts talking right at the chime is never cut off: opening a mic takes a moment,
|
|
287
|
+
* longer with Bluetooth.
|
|
268
288
|
*/
|
|
269
|
-
export async function listenForTurn(maxSeconds, outFile, signal) {
|
|
289
|
+
export async function listenForTurn(maxSeconds, outFile, signal, opts = {}) {
|
|
270
290
|
const recorder = findRecorder();
|
|
271
291
|
if (!recorder)
|
|
272
292
|
throw new SetupError(RECORDER_MISSING);
|
|
@@ -281,6 +301,19 @@ export async function listenForTurn(maxSeconds, outFile, signal) {
|
|
|
281
301
|
let pending = Buffer.alloc(0);
|
|
282
302
|
let reason = null;
|
|
283
303
|
let stderr = "";
|
|
304
|
+
const warmupFrames = Math.ceil(MIC_WARMUP_MS / FRAME_MS);
|
|
305
|
+
let announced = false;
|
|
306
|
+
const announce = () => {
|
|
307
|
+
announced = true;
|
|
308
|
+
if (!opts.onListening)
|
|
309
|
+
return;
|
|
310
|
+
endpointer.setIgnoring(true);
|
|
311
|
+
Promise.resolve()
|
|
312
|
+
.then(opts.onListening)
|
|
313
|
+
.catch(() => { })
|
|
314
|
+
// A short tail: the chime's echo fades out of the mic a moment after it ends.
|
|
315
|
+
.finally(() => setTimeout(() => endpointer.setIgnoring(false), 80));
|
|
316
|
+
};
|
|
284
317
|
const args = recorderArgs(recorder);
|
|
285
318
|
debug("exec:", recorder.bin, args.join(" "));
|
|
286
319
|
const child = spawn(recorder.bin, args, { stdio: ["ignore", "pipe", "pipe"], windowsHide: true });
|
|
@@ -308,6 +341,8 @@ export async function listenForTurn(maxSeconds, outFile, signal) {
|
|
|
308
341
|
off += FRAME_BYTES;
|
|
309
342
|
frames.push(frame);
|
|
310
343
|
reason = endpointer.push(frame);
|
|
344
|
+
if (!announced && frames.length >= warmupFrames)
|
|
345
|
+
announce();
|
|
311
346
|
}
|
|
312
347
|
pending = pending.subarray(off);
|
|
313
348
|
if (reason)
|
package/dist/config.js
CHANGED
|
@@ -27,6 +27,23 @@ export function envBool(name, fallback) {
|
|
|
27
27
|
return !["0", "false", "no", "off"].includes(raw);
|
|
28
28
|
}
|
|
29
29
|
const modelName = (process.env.VOICE_MCP_WHISPER_MODEL ?? "base.en").trim();
|
|
30
|
+
const language = process.env.VOICE_MCP_LANGUAGE?.trim() || (modelName.endsWith(".en") ? "en" : "auto");
|
|
31
|
+
/**
|
|
32
|
+
* The default hint for English: telling whisper what the conversation is about makes it hear
|
|
33
|
+
* developer words better ("rebase", "backend", "passed"). Measured: 60 clips, 3 voices, word errors
|
|
34
|
+
* 3.4% → 2.3% on everyday sentences and 3.8% → 3.1% on jargon, with no words invented in silence.
|
|
35
|
+
* A plain word list did worse than no hint, so it's a sentence.
|
|
36
|
+
*/
|
|
37
|
+
export const DEVELOPER_PROMPT = "A software developer talks to a coding assistant about the repo: the build, tests, commits, branches, pull requests, npm, JSON, TypeScript, the API and CI.";
|
|
38
|
+
/** VOICE_MCP_WHISPER_PROMPT: unset = the developer hint (English only), "none" = no hint, anything else = that text. */
|
|
39
|
+
function whisperPrompt() {
|
|
40
|
+
const raw = process.env.VOICE_MCP_WHISPER_PROMPT?.trim();
|
|
41
|
+
if (raw && ["none", "off", "0", "false"].includes(raw.toLowerCase()))
|
|
42
|
+
return undefined;
|
|
43
|
+
if (raw)
|
|
44
|
+
return raw;
|
|
45
|
+
return language === "en" ? DEVELOPER_PROMPT : undefined;
|
|
46
|
+
}
|
|
30
47
|
/** Everything this server ever downloads or links lives here, shared by every version and every client. */
|
|
31
48
|
const CACHE_ROOT = process.env.VOICE_MCP_CACHE_DIR?.trim() ||
|
|
32
49
|
path.join(process.env.XDG_CACHE_HOME || path.join(os.homedir(), ".cache"), "mac-voice-mcp");
|
|
@@ -96,9 +113,9 @@ export const CONFIG = {
|
|
|
96
113
|
/** Stop the warm whisper-server after this many idle minutes to free memory. */
|
|
97
114
|
serverIdleMinutes: Math.max(1, envNum("VOICE_MCP_SERVER_IDLE_MINUTES", 15)),
|
|
98
115
|
/** Spoken language code ("en", "th", "de", ...) or "auto". Defaults to "en" for *.en models, else "auto". */
|
|
99
|
-
language
|
|
100
|
-
/**
|
|
101
|
-
prompt:
|
|
116
|
+
language,
|
|
117
|
+
/** Initial prompt that biases vocabulary: the developer hint by default (English), or your own (names, jargon). */
|
|
118
|
+
prompt: whisperPrompt(),
|
|
102
119
|
threads: Math.max(1, Math.floor(envNum("VOICE_MCP_THREADS", Math.min(8, os.cpus().length || 4)))),
|
|
103
120
|
debug: envBool("VOICE_MCP_DEBUG", false),
|
|
104
121
|
};
|
package/dist/endpointer.js
CHANGED
|
@@ -6,6 +6,7 @@
|
|
|
6
6
|
* - once speaking, ends the turn after `endSilenceMs` of quiet
|
|
7
7
|
* - ignores blips shorter than 300 ms (a cough, a click) and keeps listening
|
|
8
8
|
* - hard stop at `maxMs`
|
|
9
|
+
* - while `ignoring` (our own "mic open" chime is playing), frames are neither speech nor background
|
|
9
10
|
*/
|
|
10
11
|
export const SAMPLE_RATE = 16000;
|
|
11
12
|
export const FRAME_MS = 30;
|
|
@@ -46,6 +47,8 @@ export class Endpointer {
|
|
|
46
47
|
lastVoicedFrame = -1;
|
|
47
48
|
/** False while every sample so far has been exactly zero (a blocked mic). */
|
|
48
49
|
anyNonZero = false;
|
|
50
|
+
/** True while our own chime plays: its sound reaches the mic, but it isn't the user. */
|
|
51
|
+
ignoring = false;
|
|
49
52
|
constructor(opts) {
|
|
50
53
|
this.opts = opts;
|
|
51
54
|
this.maxFrames = Math.ceil(opts.maxMs / FRAME_MS);
|
|
@@ -53,6 +56,12 @@ export class Endpointer {
|
|
|
53
56
|
this.endSilenceFrames = Math.ceil(opts.endSilenceMs / FRAME_MS);
|
|
54
57
|
this.historySize = Math.max(FLOOR_WINDOW_FRAMES, this.endSilenceFrames, Math.ceil(1000 / FRAME_MS));
|
|
55
58
|
}
|
|
59
|
+
/** Stop (true) or resume (false) judging frames, e.g. while the "mic open" chime plays. Time still counts. */
|
|
60
|
+
setIgnoring(on) {
|
|
61
|
+
this.ignoring = on;
|
|
62
|
+
if (on && !this.started)
|
|
63
|
+
this.loudRun = 0;
|
|
64
|
+
}
|
|
56
65
|
/** Feed one 30 ms frame of 16-bit little-endian PCM; returns a reason once the turn is over. */
|
|
57
66
|
push(frame) {
|
|
58
67
|
const view = new DataView(frame.buffer, frame.byteOffset, frame.byteLength);
|
|
@@ -67,6 +76,8 @@ export class Endpointer {
|
|
|
67
76
|
const rms = Math.sqrt(sumSq / Math.max(1, n)) / 32768;
|
|
68
77
|
const db = Math.max(-100, 20 * Math.log10(rms + 1e-9));
|
|
69
78
|
const idx = this.frames++;
|
|
79
|
+
if (this.ignoring)
|
|
80
|
+
return this.timeLimit();
|
|
70
81
|
if (idx >= SETTLE_FRAMES) {
|
|
71
82
|
this.history.push(db);
|
|
72
83
|
if (this.history.length > this.historySize)
|
|
@@ -129,6 +140,9 @@ export class Endpointer {
|
|
|
129
140
|
}
|
|
130
141
|
}
|
|
131
142
|
}
|
|
143
|
+
return this.timeLimit();
|
|
144
|
+
}
|
|
145
|
+
timeLimit() {
|
|
132
146
|
if (this.frames >= this.maxFrames)
|
|
133
147
|
return this.everStarted ? "max-duration" : "no-speech";
|
|
134
148
|
if (!this.everStarted && this.frames >= this.startTimeoutFrames)
|
package/dist/index.js
CHANGED
|
@@ -63,7 +63,7 @@ Environment variables (all optional):
|
|
|
63
63
|
VOICE_MCP_WHISPER_MODEL_PATH use this ggml model file
|
|
64
64
|
VOICE_MCP_MODEL_SEARCH_PATHS extra folders to look in for an existing model (":"-separated)
|
|
65
65
|
VOICE_MCP_LANGUAGE spoken language: en, th, de, … or auto
|
|
66
|
-
VOICE_MCP_WHISPER_PROMPT
|
|
66
|
+
VOICE_MCP_WHISPER_PROMPT a sentence about the topic, to hear its words better (default: a developer hint; none = off)
|
|
67
67
|
VOICE_MCP_VOICE / _RATE macOS say voice and words-per-minute
|
|
68
68
|
VOICE_MCP_TTS auto (Kokoro once installed) | say | kokoro
|
|
69
69
|
VOICE_MCP_KOKORO_VOICE Kokoro voice (default af_heart; e.g. af_bella, am_michael, bf_emma)
|
package/dist/pcm.js
CHANGED
|
@@ -29,6 +29,23 @@ export function resamplePcm16(pcm, fromRate, toRate) {
|
|
|
29
29
|
}
|
|
30
30
|
return out;
|
|
31
31
|
}
|
|
32
|
+
/** Seconds of audio in a 16-bit mono WAV file's data chunk (`say` may add chunks before it). 0 if there is none. */
|
|
33
|
+
export function wavSeconds(wav) {
|
|
34
|
+
if (wav.length < 12 || wav.toString("ascii", 0, 4) !== "RIFF")
|
|
35
|
+
return 0;
|
|
36
|
+
let rate = 16_000;
|
|
37
|
+
let off = 12;
|
|
38
|
+
while (off + 8 <= wav.length) {
|
|
39
|
+
const id = wav.toString("ascii", off, off + 4);
|
|
40
|
+
const size = wav.readUInt32LE(off + 4);
|
|
41
|
+
if (id === "fmt " && off + 16 <= wav.length)
|
|
42
|
+
rate = wav.readUInt32LE(off + 12) || rate;
|
|
43
|
+
if (id === "data")
|
|
44
|
+
return Math.min(size, wav.length - off - 8) / 2 / rate;
|
|
45
|
+
off += 8 + size + (size % 2);
|
|
46
|
+
}
|
|
47
|
+
return 0;
|
|
48
|
+
}
|
|
32
49
|
/** A 44-byte WAV header for 16-bit mono PCM. */
|
|
33
50
|
export function wavHeaderFor(dataBytes, sampleRate) {
|
|
34
51
|
const h = Buffer.alloc(44);
|
package/dist/speech-text.js
CHANGED
|
@@ -112,13 +112,39 @@ export function prepareSpeech(input, opts = {}) {
|
|
|
112
112
|
return { text: s, notes: noteList };
|
|
113
113
|
}
|
|
114
114
|
/** Remove timestamps and non-speech markers like [BLANK_AUDIO] from whisper output. */
|
|
115
|
+
/**
|
|
116
|
+
* whisper's notes about sounds rather than words: [BLANK_AUDIO], [MUSIC PLAYING], [gunshot] (for a
|
|
117
|
+
* cough), [APPLAUSE] (for typing), (sound of running) (for a fan), *laughs*, ♪. Never something the user said.
|
|
118
|
+
*/
|
|
119
|
+
const SOUND_NOTES = /\[[^\]]*\]|\([^)]*\)|\*[^*\n]+\*|[♪♫]+/g;
|
|
120
|
+
/**
|
|
121
|
+
* Sentences whisper invents from the videos it learned from, typically on noise or a clipped start.
|
|
122
|
+
* Only whole sentences that nobody says to a coding assistant; "Thank you." and "Bye." are real answers.
|
|
123
|
+
*/
|
|
124
|
+
const STOCK_SENTENCES = [
|
|
125
|
+
/^(?:thanks|thank you)(?: (?:so|very) much)? for (?:watching|listening)(?: and see you next time)?$/,
|
|
126
|
+
/^(?:please |don'?t forget to )?(?:like and )?subscribe(?: to (?:my|our|the) channel)?$/,
|
|
127
|
+
/^(?:you can )?find the links? in the description(?: below)?$/,
|
|
128
|
+
/^see you in the next (?:video|episode)$/,
|
|
129
|
+
/^(?:subtitles|captions|transcription|translated|transcribed)(?: by| provided by)? .*$/,
|
|
130
|
+
];
|
|
131
|
+
function isStockSentence(sentence) {
|
|
132
|
+
const s = sentence.toLowerCase().replace(/[.!?,"“”]/g, "").replace(/\s+/g, " ").trim();
|
|
133
|
+
return STOCK_SENTENCES.some((re) => re.test(s));
|
|
134
|
+
}
|
|
135
|
+
/** whisper's raw output → just the words: no timestamps, sound notes or invented stock sentences. */
|
|
115
136
|
export function cleanTranscript(raw) {
|
|
116
|
-
|
|
137
|
+
const text = raw
|
|
117
138
|
.split("\n")
|
|
118
139
|
.map((l) => l.replace(/^\s*\[[\d:.\s\->]+\]\s*/, "")) // stray timestamps
|
|
119
140
|
.join(" ")
|
|
120
|
-
.replace(
|
|
121
|
-
.replace(/\
|
|
141
|
+
.replace(SOUND_NOTES, " ")
|
|
142
|
+
.replace(/\s+/g, " ")
|
|
143
|
+
.trim();
|
|
144
|
+
const sentences = text.match(/.+?(?:[.!?]+(?=\s|$)|$)/g) ?? []; // a "." inside "1.2.3" or "Amara.org" doesn't end a sentence
|
|
145
|
+
return sentences
|
|
146
|
+
.filter((s) => !isStockSentence(s))
|
|
147
|
+
.join("")
|
|
122
148
|
.replace(/\s+/g, " ")
|
|
123
149
|
.trim();
|
|
124
150
|
}
|
package/dist/voice.js
CHANGED
|
@@ -60,10 +60,10 @@ async function turn(textToSpeak, seconds, model, signal, onPhase) {
|
|
|
60
60
|
const text = "(Spoken. The microphone was not opened, because listen was false.)";
|
|
61
61
|
return { ok: true, text, notes: [...notes, timingNote({ spokeMs: spoke, spoken })] };
|
|
62
62
|
}
|
|
63
|
-
await chime("start");
|
|
64
63
|
onPhase?.("listening");
|
|
65
64
|
const tListen = Date.now();
|
|
66
|
-
|
|
65
|
+
// The chime plays once the mic is really recording, so the first word is never lost.
|
|
66
|
+
const heard = await listenForTurn(seconds, wav, signal, { onListening: () => chime("start") });
|
|
67
67
|
void chime("stop");
|
|
68
68
|
const t1 = Date.now();
|
|
69
69
|
debug("listen:", heard);
|
|
@@ -91,8 +91,14 @@ async function turn(textToSpeak, seconds, model, signal, onPhase) {
|
|
|
91
91
|
onPhase?.("transcribing");
|
|
92
92
|
const transcript = await transcribe(wav, model, signal);
|
|
93
93
|
const transcribed = Date.now() - t1;
|
|
94
|
-
if (!transcript)
|
|
95
|
-
|
|
94
|
+
if (!transcript) {
|
|
95
|
+
// Something was loud enough to sound like speech (a cough, typing, a door), but whisper found no words in it.
|
|
96
|
+
// Starts with "(No speech detected" so the plugin's voice-mode hook knows the user wasn't heard.
|
|
97
|
+
const noWords = "(No speech detected — the microphone picked up a sound but no words, probably a cough or background noise. The microphone is now off. " +
|
|
98
|
+
"Ask once more out loud. If there's still no answer, stop and say on screen that voice mode is paused " +
|
|
99
|
+
"and they can type anything to carry on — speaking won't work until you call speak_and_listen again.)";
|
|
100
|
+
return { ok: true, text: noWords, notes: [...notes, timing(transcribed)] };
|
|
101
|
+
}
|
|
96
102
|
if (heard.reason === "max-duration") {
|
|
97
103
|
const l = heard.levels;
|
|
98
104
|
const f = (n) => (Number.isFinite(n) ? n.toFixed(0) : "?");
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "mac-voice-mcp",
|
|
3
3
|
"mcpName": "io.github.jeet0007/mac-voice-mcp",
|
|
4
|
-
"version": "0.4.
|
|
4
|
+
"version": "0.4.2",
|
|
5
5
|
"description": "Talk with Claude out loud on your Mac: speaks with macOS `say`, listens for one natural conversational turn, and transcribes on-device with whisper.cpp. An MCP server.",
|
|
6
6
|
"type": "module",
|
|
7
7
|
"bin": {
|
|
@@ -33,7 +33,7 @@
|
|
|
33
33
|
"version": "node scripts/sync-version.mjs && git add server.json .claude-plugin/plugin.json"
|
|
34
34
|
},
|
|
35
35
|
"dependencies": {
|
|
36
|
-
"@modelcontextprotocol/sdk": "^1.
|
|
36
|
+
"@modelcontextprotocol/sdk": "^1.32.1",
|
|
37
37
|
"zod": "^4.1.0"
|
|
38
38
|
},
|
|
39
39
|
"devDependencies": {
|