whisper-windows-mcp 1.1.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/FUNDING.yml ADDED
@@ -0,0 +1,3 @@
1
+ github: eviscerations
2
+ ko_fi:
3
+ patreon:
package/README.md CHANGED
@@ -1,9 +1,9 @@
1
1
  # whisper-windows-mcp
2
2
 
3
- A Windows-native MCP (Model Context Protocol) server that lets Claude Desktop transcribe audio files locally using [whisper.cpp](https://github.com/ggerganov/whisper.cpp) — no internet connection required, no data sent to the cloud.
3
+ A Windows-native MCP (Model Context Protocol) server that lets Claude Desktop transcribe audio files locally using [whisper.cpp](https://github.com/ggml-org/whisper.cpp) — no internet connection required, no data sent to the cloud.
4
4
 
5
5
  > **Why does this exist?**
6
- > The popular `whisper-mcp` package was built for macOS and assumes a Unix environment. It does not work on Windows. This package was written specifically for Windows users who want the same local transcription functionality in Claude Desktop.
6
+ > The popular `whisper-mcp` package was built for macOS and assumes a Unix environment. It does not work on Windows. This package was written specifically for Windows users who want the same functionality.
7
7
 
8
8
  ---
9
9
 
@@ -12,8 +12,9 @@ A Windows-native MCP (Model Context Protocol) server that lets Claude Desktop tr
12
12
  Once installed, you can say things like this directly in Claude Desktop:
13
13
 
14
14
  - *"Transcribe C:\Users\Me\Downloads\meeting.mp3"*
15
- - *"Transcribe this recording and summarise the key points"*
15
+ - *"Transcribe this recording and give me a summary"*
16
16
  - *"Transcribe with timestamps so I can find specific moments"*
17
+ - *"Generate subtitles for this video"*
17
18
 
18
19
  Everything runs on your own machine. No audio ever leaves your computer.
19
20
 
@@ -21,61 +22,93 @@ Everything runs on your own machine. No audio ever leaves your computer.
21
22
 
22
23
  ## Requirements
23
24
 
24
- You need the following installed before proceeding. Each one is free.
25
+ Before installing this package, you need three things set up on your Windows machine:
25
26
 
26
- | Requirement | Purpose |
27
- |---|---|
28
- | [Node.js 18+](https://nodejs.org/en/download) | Runs the MCP server |
29
- | [whisper.cpp](https://github.com/ggerganov/whisper.cpp/releases/latest) | The transcription engine |
30
- | A Whisper model file | The AI model (downloaded in Step 2) |
27
+ 1. **Node.js 18 or later** — [download from nodejs.org](https://nodejs.org)
28
+ 2. **whisper.cpp binaries** — the actual transcription engine (see Step 1 below)
29
+ 3. **A Whisper model file** — the AI model that does the transcription (see Step 2 below)
31
30
 
32
31
  ---
33
32
 
34
- ## Step 1 — Install whisper.cpp
33
+ ## Step 1 — Install whisper.cpp binaries
34
+
35
+ ### Option A — Pre-built Vulkan release (recommended)
36
+
37
+ Download `whisper-vulkan-win-x64.zip` from the [releases page](https://github.com/eviscerations/whisper-windows-mcp/releases).
38
+
39
+ This is a custom-compiled build with **Vulkan GPU acceleration** enabled. It works with AMD, NVIDIA, and Intel GPUs on Windows — no vendor-specific SDK required.
40
+
41
+ Extract the zip to `C:\whisper\Release\`. You should end up with these files:
42
+
43
+ ```
44
+ C:\whisper\Release\whisper-cli.exe
45
+ C:\whisper\Release\ggml-vulkan.dll
46
+ C:\whisper\Release\ggml.dll
47
+ C:\whisper\Release\ggml-base.dll
48
+ C:\whisper\Release\ggml-cpu.dll
49
+ C:\whisper\Release\whisper.dll
50
+ ```
51
+
52
+ **GPU acceleration is automatic** — if a supported GPU is present, whisper.cpp will use it. No additional configuration needed.
53
+
54
+ ### Option B — Build from source
55
+
56
+ If you prefer to compile your own binary (advanced users):
35
57
 
36
- 1. Go to the [whisper.cpp latest release](https://github.com/ggerganov/whisper.cpp/releases/latest)
37
- 2. Download the file named **`whisper-bin-x64.zip`** (look for `win` and `x64` in the filename)
38
- 3. Extract the ZIP and move the contents to **`C:\whisper\Release\`** — create this folder if it doesn't exist
58
+ **Prerequisites:** Git, CMake, Visual Studio Build Tools 2022+ with "Desktop development with C++", Vulkan SDK from [lunarg.com](https://vulkan.lunarg.com/sdk/home#windows).
39
59
 
40
- ✅ You should now have **`C:\whisper\Release\whisper-cli.exe`**
60
+ ```
61
+ git clone https://github.com/ggml-org/whisper.cpp
62
+ cd whisper.cpp
63
+ cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
64
+ cmake --build build --config Release --target whisper-cli
65
+ ```
66
+
67
+ Copy the resulting binaries from `build\bin\Release\` to `C:\whisper\Release\`.
41
68
 
42
- > **Why this path?** You can install whisper.cpp anywhere, but `C:\whisper\Release\` matches the default config below and means less to edit later.
69
+ > **Note:** The default whisper.cpp Windows release on GitHub does not include a Vulkan build. You must either use the pre-built release above or compile from source with `-DGGML_VULKAN=ON`.
43
70
 
44
71
  ---
45
72
 
46
73
  ## Step 2 — Download a Whisper model
47
74
 
48
- The model is the AI that does the actual transcription. Click a link below to download directly:
75
+ Models are downloaded from Hugging Face. Choose one based on your needs:
49
76
 
50
- | Model | Download | Size | Speed | Best for |
77
+ | Model | File size | Speed | Accuracy | Recommended for |
51
78
  |---|---|---|---|---|
52
- | tiny.en | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-tiny.en.bin) | 75 MB | Very fast | Quick tests |
53
- | **base.en** | [**Download**](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en.bin) | 142 MB | Fast | **Recommended starting point** |
54
- | small.en | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.en.bin) | 466 MB | Moderate | Better accuracy |
55
- | medium.en | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-medium.en.bin) | 1.5 GB | Slow | High accuracy |
56
- | large-v3 | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3.bin) | 2.9 GB | Very slow | Maximum accuracy |
79
+ | `ggml-tiny.en.bin` | 75 MB | Very fast | Basic | Quick tests |
80
+ | `ggml-base.en.bin` | 142 MB | Fast | Good | Everyday use |
81
+ | `ggml-small.en.bin` | 466 MB | Moderate | Better | Important recordings |
82
+ | `ggml-medium.en.bin` | 1.5 GB | Fast on GPU | Very good | Best quality |
83
+ | `ggml-large-v3.bin` | 2.9 GB | Fast on GPU | Excellent | Maximum accuracy |
57
84
 
58
- Save the downloaded `.bin` file to **`C:\whisper\models\`** — create this folder if it doesn't exist.
85
+ **For most people, `base.en` or `small.en` is the best starting point.** With GPU acceleration, `medium.en` and `large-v3` become practical for everyday use.
59
86
 
60
- ✅ You should now have something like **`C:\whisper\models\ggml-base.en.bin`**
87
+ Download your chosen model from:
61
88
 
62
- ---
89
+ ```
90
+ https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en.bin
91
+ ```
63
92
 
64
- ## Step 3 — Install Node.js
93
+ Save it to `C:\whisper\models\` — create that folder if it doesn't exist.
65
94
 
66
- If you don't already have Node.js:
95
+ ---
67
96
 
68
- 1. Go to [nodejs.org](https://nodejs.org/en/download) and download the **Windows Installer (.msi)** — choose the LTS version
69
- 2. Run the installer and accept all defaults
97
+ ## Step 3 — Install this MCP server
70
98
 
71
- ✅ To verify, open Command Prompt and run `node --version` — you should see something like `v20.x.x`
99
+ ```
100
+ npm install -g whisper-windows-mcp
101
+ ```
102
+
103
+ Or use `npx` directly in your config (see Step 4).
72
104
 
73
105
  ---
74
106
 
75
107
  ## Step 4 — Configure Claude Desktop
76
108
 
77
- 1. Open Claude Desktop → **Settings → Developer → Edit Config**
78
- 2. Add the following (or merge the `mcpServers` block if you already have other servers):
109
+ 1. Open Claude Desktop
110
+ 2. Go to **Settings → Developer → Edit Config**
111
+ 3. Add the whisper-windows-mcp entry:
79
112
 
80
113
  ```json
81
114
  {
@@ -92,11 +125,10 @@ If you don't already have Node.js:
92
125
  }
93
126
  ```
94
127
 
95
- > ⚠️ **Path format:** In the JSON config, all backslashes must be doubled (`\\`). This is a JSON requirement. When typing paths into Claude in chat, use normal single backslashes.
128
+ > **Important:** If your `claude_desktop_config.json` already has other content, add the `"mcpServers"` block inside the existing `{}` — don't replace the whole file.
96
129
 
97
- 3. If you downloaded a different model, update `ggml-base.en.bin` to match your filename
98
- 4. Save the file, fully quit Claude Desktop, and reopen it
99
- 5. Go to **Settings → Developer** — you should see **whisper** with a green **running** badge
130
+ 4. Save the file and fully restart Claude Desktop.
131
+ 5. Go to **Settings → Developer** — you should see **whisper** listed with a green **running** badge.
100
132
 
101
133
  ---
102
134
 
@@ -106,65 +138,40 @@ In Claude Desktop, type:
106
138
 
107
139
  > *"Can you check your whisper config?"*
108
140
 
109
- Claude will verify that `whisper-cli.exe` and your model file are both found. Then try:
141
+ Claude will use the `check_config` tool to verify everything is set up correctly. Then try a transcription:
110
142
 
111
143
  > *"Please transcribe C:\Users\YourName\Downloads\recording.mp3"*
112
144
 
113
145
  ---
114
146
 
115
- ## Converting video files to audio
116
-
117
- Whisper processes audio. If you have a video file (MP4, MKV, etc.) you may want to extract the audio first — audio-only files are much smaller and faster to process.
147
+ ## GPU acceleration
118
148
 
119
- > **Tip:** whisper.cpp may handle MP4 files directly if FFmpeg is installed. Try transcribing an MP4 first before converting.
149
+ The pre-built Vulkan release enables GPU acceleration automatically. No flags or configuration required — whisper.cpp detects your GPU at startup and uses it if available.
120
150
 
121
- **Using VLC Media Player** (free, recommended for beginners):
151
+ **Confirmed working:** AMD Radeon RX Vega 56 (GCN 5th gen), and any GPU with Vulkan 1.0+ support.
122
152
 
123
- 1. Download [VLC](https://www.videolan.org/vlc/) if you don't have it
124
- 2. Open VLC → **Media → Convert / Save**
125
- 3. Click **Add**, select your video, then click **Convert / Save**
126
- 4. Under **Profile**, choose **Audio - MP3**
127
- 5. Set a destination filename and click **Start**
153
+ **Performance comparison with medium.en model:**
128
154
 
129
- A 1-hour MP4 that might be 2–4 GB typically becomes a 50–100 MB MP3.
155
+ | Hardware | ~5 min audio file |
156
+ |---|---|
157
+ | CPU only (Ryzen 7 2700x) | ~8–12 minutes |
158
+ | GPU (Vega 56 via Vulkan) | ~20–40 seconds |
130
159
 
131
- **Using FFmpeg** (command line, for advanced users):
132
- ```
133
- ffmpeg -i "C:\path\to\video.mp4" -vn -ac 1 -ar 16000 "C:\path\to\output.wav"
134
- ```
160
+ GPU utilization during transcription is typically 15–30% — efficient bursts, not sustained load.
135
161
 
136
162
  ---
137
163
 
138
164
  ## Output formats
139
165
 
140
- | Format | What you get | Ask Claude... |
141
- |---|---|---|
142
- | `text` (default) | Plain transcript, no timestamps | *"Transcribe this file"* |
143
- | `timestamps` | Transcript with `[00:00:00 --> 00:00:05]` time codes | *"Transcribe with timestamps"* |
144
- | `json` | Structured data | *"Transcribe as JSON"* |
145
-
146
- ---
147
-
148
- ## Transcription speed
149
-
150
- Whisper runs on CPU by default. Rough estimates for a 1-hour recording:
151
-
152
- | Model | Approximate time (CPU) |
153
- |---|---|
154
- | tiny.en | 5–10 minutes |
155
- | base.en | 10–20 minutes |
156
- | small.en | 20–35 minutes |
157
- | medium.en | 35–60 minutes |
158
- | large-v3 | 60–120 minutes |
159
-
160
- > **GPU acceleration** for AMD (ROCm) and NVIDIA (CUDA) on Windows is planned for a future update.
166
+ - **text** (default) — plain transcript
167
+ - **timestamps** — transcript with `[00:00:00 --> 00:00:05]` time codes
168
+ - **json** — structured output
169
+ - **srt** — subtitle file saved next to the source file
161
170
 
162
171
  ---
163
172
 
164
173
  ## Full config example
165
174
 
166
- If you have other MCP servers already configured:
167
-
168
175
  ```json
169
176
  {
170
177
  "preferences": {
@@ -176,26 +183,25 @@ If you have other MCP servers already configured:
176
183
  "args": ["-y", "whisper-windows-mcp"],
177
184
  "env": {
178
185
  "WHISPER_CLI_PATH": "C:\\whisper\\Release\\whisper-cli.exe",
179
- "WHISPER_MODEL": "C:\\whisper\\models\\ggml-base.en.bin"
186
+ "WHISPER_MODEL": "C:\\whisper\\models\\ggml-medium.en.bin"
180
187
  }
181
188
  }
182
189
  }
183
190
  }
184
191
  ```
185
192
 
186
- Config file location:
187
- ```
188
- C:\Users\YourUsername\AppData\Roaming\Claude\claude_desktop_config.json
189
- ```
190
-
191
- > The `AppData` folder is hidden by default. To show it: File Explorer → **View → Show → Hidden items**
193
+ Config file location: `C:\Users\YourUsername\AppData\Roaming\Claude\claude_desktop_config.json`
192
194
 
193
195
  ---
194
196
 
195
- ## Tested on
197
+ ## Optional environment variables
196
198
 
197
- - Windows 10 Pro (10.0.19045)
198
- - Windows 11 — untested, feedback welcome via [Issues](../../issues)
199
+ | Variable | Description |
200
+ |---|---|
201
+ | `WHISPER_CLI_PATH` | Path to whisper-cli.exe (required) |
202
+ | `WHISPER_MODEL` | Path to model .bin file (required) |
203
+ | `WHISPER_THREADS` | CPU thread count override |
204
+ | `FFMPEG_PATH` | Path to ffmpeg if not in system PATH |
199
205
 
200
206
  ---
201
207
 
@@ -204,23 +210,11 @@ C:\Users\YourUsername\AppData\Roaming\Claude\claude_desktop_config.json
204
210
  See [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for detailed solutions.
205
211
 
206
212
  Quick checklist:
207
-
208
- - [ ] Config paths use **double backslashes** (`C:\\whisper\\...`)
209
- - [ ] `whisper-cli.exe` exists at the path specified
210
- - [ ] The model `.bin` file exists at the path specified
211
- - [ ] Claude Desktop was **fully restarted** after editing the config
212
- - [ ] Whisper shows **running** in Settings → Developer
213
-
214
- ---
215
-
216
- ## Roadmap
217
-
218
- - [ ] SRT subtitle output
219
- - [ ] Direct MP4/video file support via FFmpeg
220
- - [ ] Translation to English from other languages
221
- - [ ] AMD GPU acceleration (ROCm)
222
- - [ ] NVIDIA GPU acceleration (CUDA)
223
- - [ ] Speaker diarization (automatic speaker identification)
213
+ - Paths in config use **double backslashes** (`C:\\whisper\\...`)
214
+ - `whisper-cli.exe` exists at the path specified
215
+ - The model `.bin` file exists at the path specified
216
+ - Claude Desktop was fully restarted after editing config
217
+ - The whisper server shows **running** in Settings → Developer
224
218
 
225
219
  ---
226
220
 
@@ -232,4 +226,4 @@ MIT — free to use, modify, and distribute.
232
226
 
233
227
  ## Contributing
234
228
 
235
- Pull requests welcome. GPU acceleration solutions for AMD or NVIDIA especially appreciated. Windows 11 feedback welcome via Issues.
229
+ Pull requests welcome. See [ROADMAP.md](ROADMAP.md) for planned features.
package/ROADMAP.md ADDED
@@ -0,0 +1,129 @@
1
+ # whisper-windows-mcp — Roadmap
2
+
3
+ Current version is v1.4.0. GPU acceleration via Vulkan is now working. This document tracks what has been completed and what remains.
4
+
5
+ ---
6
+
7
+ ## Completed
8
+
9
+ ### ✅ Priority 1 — GPU Acceleration (v1.4.0)
10
+
11
+ Compiled whisper.cpp from source with `-DGGML_VULKAN=ON` using Visual Studio Build Tools 2026 and Vulkan SDK 1.4.341.1. Pre-built Vulkan binaries are now distributed as a release asset (`whisper-vulkan-win-x64.zip`).
12
+
13
+ **Results:** AMD Radeon RX Vega 56 at ~16% GPU utilization, 1.2GB VRAM, CPU at ~15% during transcription. A ~5 minute file that previously took 8–12 minutes on CPU now completes in 20–40 seconds.
14
+
15
+ The official whisper.cpp Windows releases do not include a Vulkan build ([issue #3673](https://github.com/ggml-org/whisper.cpp/issues/3673)). The pre-built release in this repo fills that gap for AMD and Intel GPU users.
16
+
17
+ ### ✅ Priority 2 — Process Lock (v1.3.1)
18
+
19
+ Added `isWhisperRunning()` check using `tasklist /FI` before any transcription spawn. If `whisper-cli.exe` is already running, returns a clear error with Task Manager instructions rather than spawning a second competing process.
20
+
21
+ ---
22
+
23
+ ## Known Issues (Remaining)
24
+
25
+ ### 3. 4-Minute Claude MCP Timeout
26
+ The Claude web client cuts MCP connections after ~4 minutes. Whisper continues running in the background after the timeout fires, but Claude can't confirm completion. With GPU acceleration this is less frequently hit, but still a concern for very long files or large models.
27
+
28
+ ### 5. No Progress Visibility
29
+ The user has no indicator that transcription is happening or how far along it is. whisper-cli.exe outputs segment timestamps to stderr as it processes — the MCP should pipe and expose these.
30
+
31
+ ### 6. Background Batch Non-Functional
32
+ A background batch mode was attempted in a previous version but stripped due to no visible feedback. Needs a proper detached process architecture with job state files.
33
+
34
+ ### 7. No File Pre-Analysis
35
+ No way to know file duration or size before processing starts.
36
+
37
+ ---
38
+
39
+ ## Roadmap
40
+
41
+ ### Priority 3 — File Pre-Analysis Tool
42
+
43
+ New tool `analyze_media` using FFprobe:
44
+
45
+ ```
46
+ ffprobe -v quiet -print_format json -show_format -show_streams <file>
47
+ ```
48
+
49
+ Returns: duration, file size, codec, bitrate, estimated transcription time (CPU and GPU).
50
+
51
+ ---
52
+
53
+ ### Priority 4 — Progress Visibility
54
+
55
+ whisper-cli.exe outputs segment timestamps to stderr (e.g. `[00:01:30 --> 00:01:35]`). The MCP should:
56
+
57
+ 1. Pipe stderr to a log file during processing
58
+ 2. Expose a `check_progress` tool returning last timestamp, percentage complete, estimated time remaining, and whether the process is still running
59
+
60
+ ---
61
+
62
+ ### Priority 5 — Timeout Workaround (Detached Process Architecture)
63
+
64
+ Rearchitect transcription to use fully detached background processes:
65
+
66
+ 1. `transcribe_audio` → spawn whisper as detached process → write `job.json` → return immediately with job ID
67
+ 2. `check_progress` → read `job.json` → check PID → read log → return status + percentage
68
+ 3. When complete → read and return transcript
69
+
70
+ This eliminates the 4-minute timeout problem entirely and unlocks Priority 6.
71
+
72
+ ---
73
+
74
+ ### Priority 6 — Sequential Batch with Validation
75
+
76
+ Rebuild batch mode on top of Priority 5:
77
+
78
+ 1. `transcribe_batch` → runs `analyze_media` on folder → sorts by duration → processes one file at a time using detached process
79
+ 2. After each file: validate .txt exists, is non-empty, line count proportional to duration
80
+ 3. Flag suspect outputs for re-run
81
+ 4. `check_batch_progress` returns: files done, remaining, current file, ETA, any failed files
82
+
83
+ ---
84
+
85
+ ### Priority 7 — Multi-Language Support and Translation
86
+
87
+ Expose `--language` and `--translate` flags properly:
88
+
89
+ - `language` parameter: auto-detect (default) or specify (`ja`, `es`, `de`, etc.)
90
+ - `translate_to_english`: boolean — uses whisper's built-in translation model
91
+ - Dual output: two whisper passes, two output files
92
+
93
+ ---
94
+
95
+ ### Priority 8 — Filename-Based References Throughout
96
+
97
+ All tool outputs must reference the full source filename at all times. Never use positional indices as the primary identifier.
98
+
99
+ ---
100
+
101
+ ### Priority 9 — System Diagnostics Tool
102
+
103
+ New tool `check_system`:
104
+
105
+ - GPU vendor and model (via `wmic path win32_VideoController`)
106
+ - Whether `ggml-vulkan.dll` is present alongside `whisper-cli.exe`
107
+ - Recommended model size for available VRAM
108
+ - Estimated throughput based on hardware profile
109
+ - Actionable guidance if GPU binary is missing
110
+
111
+ ---
112
+
113
+ ## Design Principles
114
+
115
+ **Minimize Claude API usage.** The entire transcription workflow should require fewer than 20 Claude interactions for a 60-file batch.
116
+
117
+ **One whisper instance at all times.** Never spawn a second process while one is running.
118
+
119
+ **Local-first, private by default.** No audio leaves the machine. No cloud APIs required.
120
+
121
+ **Works for free-tier users.** Courtroom transcription, documentary research, foreign film subtitling — the tool should serve people who can't afford cloud transcription services.
122
+
123
+ ---
124
+
125
+ ## Contributing
126
+
127
+ Pull requests welcome for any of the above priorities. Check existing issues before starting work.
128
+
129
+ If you've tested GPU acceleration on hardware not listed above, please open an issue with your results — GPU model, VRAM, model size, and observed throughput.