whisper-windows-mcp 2.2.1 → 2.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,393 +1,393 @@
1
- # whisper-windows-mcp
2
-
3
- A Windows-native MCP (Model Context Protocol) server that lets Claude Desktop transcribe audio and video files locally using [whisper.cpp](https://github.com/ggml-org/whisper.cpp) — with GPU acceleration, multilingual support, and batch processing. All transcription runs locally — no audio, video, or file paths ever leave your machine.
4
-
5
- > **Why does this exist?**
6
- > The popular `whisper-mcp` package was built for macOS and assumes a Unix environment. It does not work on Windows. This package was written specifically for Windows users who want local AI transcription integrated with Claude Desktop.
7
-
8
- ---
9
-
10
- ## What you can do with it
11
-
12
- Once installed, you can say things like this directly in Claude Desktop:
13
-
14
- - *"Transcribe C:\Users\Me\Downloads\meeting.mp3"*
15
- - *"Transcribe this folder of recordings and save each as a text file"*
16
- - *"Generate Japanese and English subtitles for this video"*
17
- - *"Start a batch transcription of everything in this folder"*
18
- - *"How long will it take to transcribe these files?"*
19
- - *"Check if GPU acceleration is working"*
20
-
21
- ---
22
-
23
- ## Requirements
24
-
25
- 1. **Node.js 18 or later** — [nodejs.org](https://nodejs.org)
26
- 2. **whisper.cpp binaries with Vulkan GPU support** — see Step 1
27
- 3. **A Whisper model file** — see Step 2
28
- 4. **FFmpeg** — required for video files and non-WAV/MP3 audio
29
-
30
- ---
31
-
32
- ## Step 1 — Install whisper.cpp binaries
33
-
34
- ### Option A — Pre-built Vulkan release (recommended)
35
-
36
- Download `whisper-vulkan-win-x64.zip` from the [releases page](https://github.com/eviscerations/whisper-windows-mcp/releases/tag/v1.4.0).
37
-
38
- This is a custom-compiled build with **Vulkan GPU acceleration** enabled. Works with AMD, NVIDIA, and Intel GPUs — no vendor-specific SDK required.
39
-
40
- Extract to `C:\whisper\Release\`. You should end up with:
41
-
42
- ```
43
- C:\whisper\Release\whisper-cli.exe
44
- C:\whisper\Release\ggml-vulkan.dll
45
- C:\whisper\Release\ggml.dll
46
- C:\whisper\Release\ggml-base.dll
47
- C:\whisper\Release\ggml-cpu.dll
48
- C:\whisper\Release\whisper.dll
49
- ```
50
-
51
- GPU acceleration is automatic — no additional configuration needed.
52
-
53
- ### Option B — Build from source
54
-
55
- Requires: Git, CMake, Visual Studio Build Tools 2022+ with "Desktop development with C++", Vulkan SDK from [lunarg.com](https://vulkan.lunarg.com/sdk/home#windows).
56
-
57
- ```
58
- git clone https://github.com/ggml-org/whisper.cpp
59
- cd whisper.cpp
60
- cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
61
- cmake --build build --config Release --target whisper-cli
62
- ```
63
-
64
- Copy the binaries from `build\bin\Release\` to `C:\whisper\Release\`.
65
-
66
- > **Note:** The official whisper.cpp Windows releases on GitHub do not include a Vulkan build. You must use the pre-built release above or compile from source with `-DGGML_VULKAN=ON`.
67
-
68
- ---
69
-
70
- ## Step 2 — Download a Whisper model
71
-
72
- | Model | Size | Speed | Accuracy | Best for |
73
- |---|---|---|---|---|
74
- | `ggml-tiny.en.bin` | 75 MB | Very fast | Basic | Quick tests |
75
- | `ggml-base.en.bin` | 142 MB | Fast | Good | Everyday English |
76
- | `ggml-small.en.bin` | 466 MB | Moderate | Better | Important recordings |
77
- | `ggml-medium.en.bin` | 1.5 GB | Fast on GPU | Very good | Best quality English |
78
- | `ggml-large-v3-turbo.bin` | 1.6 GB | Fast on GPU | Excellent | **Recommended for English GPU batch work — ~6x faster than large-v3 with minimal accuracy loss** |
79
- | `ggml-large-v3.bin` | 2.9 GB | Fast on GPU | Excellent | Multilingual, maximum accuracy |
80
- | `ggml-medium.en-q5_0.bin` | 514 MB | Fast | Very good | **Best CPU-only English option — high accuracy at low memory** |
81
- | `ggml-large-v3-turbo-q5_0.bin` | 547 MB | Fast | Excellent | **Best CPU-only multilingual option** |
82
- | `ggml-large-v3-q5_0.bin` | 1.1 GB | Moderate on CPU | Excellent | Multilingual, CPU-friendly |
83
-
84
- Use `download_model` in Claude Desktop to install any of these directly. For **English-only** use: `large-v3-turbo` (GPU) or `medium.en-q5_0` (CPU) are the best starting points. For **multilingual** use: `large-v3-turbo` or `large-v3-turbo-q5_0` (CPU). English-only models (`*.en.bin`) output `[FOREIGN]` on non-English audio and cannot be used for other languages.
85
-
86
- ---
87
-
88
- ## Step 3 — Install FFmpeg
89
-
90
- FFmpeg is required for video files and non-native audio formats.
91
-
92
- Install via winget:
93
- ```
94
- winget install ffmpeg
95
- ```
96
-
97
- Or download from [ffmpeg.org](https://ffmpeg.org/download.html) and add to your PATH.
98
-
99
- Verify:
100
- ```
101
- ffmpeg -version
102
- ```
103
-
104
- ---
105
-
106
- ## Step 4 — Install this MCP server
107
-
108
- ```
109
- npm install -g whisper-windows-mcp
110
- ```
111
-
112
- ---
113
-
114
- ## Step 5 — Configure Claude Desktop
115
-
116
- Open Claude Desktop → Settings → Developer → Edit Config.
117
-
118
- Add the `whisper` entry:
119
-
120
- ```json
121
- {
122
- "mcpServers": {
123
- "whisper": {
124
- "command": "npx",
125
- "args": ["-y", "whisper-windows-mcp"],
126
- "env": {
127
- "WHISPER_CLI_PATH": "C:\\whisper\\Release\\whisper-cli.exe",
128
- "WHISPER_MODEL": "C:\\whisper\\models\\ggml-medium.en.bin"
129
- }
130
- }
131
- }
132
- }
133
- ```
134
-
135
- Config file location: `C:\Users\YourName\AppData\Roaming\Claude\claude_desktop_config.json`
136
-
137
- > Use **double backslashes** in all paths.
138
-
139
- Save and **fully restart** Claude Desktop. You should see **whisper** listed with a green running badge in Settings → Developer.
140
-
141
- ---
142
-
143
- ## Step 6 — Verify your setup
144
-
145
- In Claude Desktop, ask:
146
-
147
- > *"Check your whisper config"*
148
-
149
- Then:
150
-
151
- > *"Check your system hardware"*
152
-
153
- This confirms your GPU is detected and Vulkan acceleration is active.
154
-
155
- ---
156
-
157
- ## Available tools
158
-
159
- ### `transcribe_audio`
160
- Transcribe a single file. Supports blocking (default) or background mode for long files.
161
-
162
- | Parameter | Description |
163
- |---|---|
164
- | `file_path` | Absolute path to the file (required) |
165
- | `language` | Language code (`en`, `ja`, `es`, etc.) or `auto` to detect. Default: `en` |
166
- | `output_format` | `text` (default), `timestamps`, `json`, or `srt` |
167
- | `save_to_file` | Save transcript as .txt next to the source file |
168
- | `background` | Run as detached job — returns a job ID immediately. Use `check_progress` to monitor. Recommended for files over 10 minutes. |
169
- | `threads` | CPU thread override |
170
- | `temperature` | Sampling temperature 0.0–1.0. Default 0.0 (deterministic). Higher values reduce hallucination on noisy audio. |
171
- | `prompt` | Prior context string — improves accuracy for domain-specific vocabulary or speaker names. Example: `"Names: Keemstar, DramaAlert."` |
172
- | `condition_on_prev_text` | Re-enable context conditioning between segments. Default false. |
173
- | `beam_size` | Beam search width. Higher = more accurate, slower. Default 5. |
174
- | `best_of` | Candidate sequences evaluated. Default 5. |
175
- | `gpu_device` | GPU device index for multi-GPU systems. Default 0. |
176
- | `processors` | Parallel processor count. Default 1. |
177
- | `word_timestamps` | One word per timestamped segment. Useful for clip alignment. |
178
- | `max_segment_length` | Max segment length in characters. |
179
- | `diarize` | Stereo speaker diarization — requires stereo audio with speakers on separate channels. |
180
- | `vad_model` | Path to Silero VAD model .bin. Strips silence before transcription — reduces hallucinations on noisy files. |
181
- | `offset_t` | Start offset in milliseconds. |
182
- | `duration` | Process duration in milliseconds from offset. |
183
-
184
- ---
185
-
186
- ### `check_progress`
187
- Monitor a background transcription job started with `transcribe_audio` (background=true).
188
-
189
- Returns elapsed time, last processed timestamp, percentage, and the full transcript when complete.
190
-
191
- | Parameter | Description |
192
- |---|---|
193
- | `job_id` | Job ID returned by `transcribe_audio` |
194
-
195
- ---
196
-
197
- ### `start_batch`
198
- Automated sequential batch transcription of all untranscribed files in a folder. Sorts by duration (shortest first), processes one at a time as background jobs, validates each output.
199
-
200
- | Parameter | Description |
201
- |---|---|
202
- | `folder_path` | Path to folder (required) |
203
- | `language` | Language code. Default: `en` |
204
- | `threads` | CPU thread override |
205
-
206
- ---
207
-
208
- ### `check_batch_progress`
209
- Monitor a running batch. Automatically advances to the next file when the current one finishes. Returns overall progress, current file with timestamp, ETA, and any failed files.
210
-
211
- | Parameter | Description |
212
- |---|---|
213
- | `batch_id` | Batch ID returned by `start_batch` |
214
-
215
- ---
216
-
217
- ### `transcribe_batch` (interactive)
218
- Process files one at a time with a preview and confirmation before each. Useful when you want to review as you go.
219
-
220
- | Parameter | Description |
221
- |---|---|
222
- | `folder_path` | Path to folder (required) |
223
- | `file_index` | Which file to process (1-based). Omit to list files first. |
224
- | `language` | Language code. Default: `en` |
225
- | `recursive` | Include subfolders |
226
-
227
- ---
228
-
229
- ### `generate_subtitles`
230
- Generate SRT subtitle files. Supports automatic language detection and English translation output.
231
-
232
- | Parameter | Description |
233
- |---|---|
234
- | `file_path` | Path to file (required) |
235
- | `language` | Language code or `auto` to detect. Default: `en` |
236
- | `translate_to_english` | Also generate an English translation `.en.srt`. Only applies when source is not English. |
237
- | `threads` | CPU thread override |
238
-
239
- When both native and translation are requested, two files are saved next to the source:
240
- - `filename.ja.srt` — original language
241
- - `filename.en.srt` — English translation
242
-
243
- > Whisper's built-in translation only translates **to English**. For other target languages, translate the .srt file contents separately.
244
-
245
- ---
246
-
247
- ### `analyze_media`
248
- Analyze files before committing to transcription. Returns duration, size, codec, and estimated transcription time on CPU and GPU. For folders, shows all files in a sortable table with transcription status.
249
-
250
- | Parameter | Description |
251
- |---|---|
252
- | `path` | Path to a single file or folder (required) |
253
- | `sort_by` | For folders: `duration` (default), `name`, or `size` |
254
-
255
- ---
256
-
257
- ### `check_config`
258
- Verify whisper-cli.exe, the model file, and FFmpeg are all accessible. Run this first if anything is failing.
259
-
260
- ---
261
-
262
- ### `list_models`
263
- List all Whisper model files installed in your models directory. Shows filename, size, whether it is currently active, quantization status, and recommended use case. No network calls — reads local filesystem only.
264
-
265
- ---
266
-
267
- ### `download_model`
268
- Download a Whisper model directly from Hugging Face into your models directory. Accepts a model name (e.g. `large-v3-turbo`, `medium.en-q5_0`) and handles the download automatically. Only downloads from trusted Hugging Face namespaces. After downloading, use `switch_model` to activate it.
269
-
270
- | Parameter | Description |
271
- |---|---|
272
- | `model_name` | Model name to download, e.g. `large-v3-turbo`, `large-v3-turbo-q5_0`, `medium.en-q5_0` |
273
-
274
- ---
275
-
276
- ### `switch_model`
277
- Switch the active Whisper model for the current session without restarting Claude Desktop. Change is session-scoped — does not persist after restart. To make permanent, update `WHISPER_MODEL` in your config.
278
-
279
- | Parameter | Description |
280
- |---|---|
281
- | `model_name` | Model filename (e.g. `ggml-large-v3-turbo.bin`) or full path. Must be a `.bin` file in the configured models directory. |
282
-
283
- ---
284
-
285
- ### `check_system`
286
- Detect GPU hardware and verify Vulkan acceleration is available. Reports GPU name, VRAM, whether `ggml-vulkan.dll` is present, and recommends the best model size for your hardware.
287
-
288
- ---
289
-
290
- ## Supported formats
291
-
292
- | Type | Formats |
293
- |---|---|
294
- | Native (no conversion) | `mp3`, `wav` |
295
- | Video (auto-converted via FFmpeg) | `mp4`, `mkv`, `avi`, `mov`, `webm`, `flv`, `wmv`, `m4v`, `ts`, `3gp` |
296
- | Audio (auto-converted via FFmpeg) | `m4a`, `ogg`, `flac` |
297
-
298
- ---
299
-
300
- ## GPU acceleration
301
-
302
- The pre-built Vulkan release enables GPU acceleration automatically. Tested on AMD Radeon RX Vega 56 (GCN 5th gen). Any GPU with Vulkan 1.0+ support should work, including NVIDIA and Intel Arc.
303
-
304
- **Performance comparison (medium.en model, ~5 minute audio file):**
305
-
306
- | Hardware | Time |
307
- |---|---|
308
- | CPU only (Ryzen 7 2700x, 8 threads) | 8–12 minutes |
309
- | GPU (Vega 56 via Vulkan) | 20–40 seconds |
310
-
311
- GPU utilization during transcription is typically 15–20%, dropping back to idle between files. CPU stays around 15%.
312
-
313
- ---
314
-
315
- ## Multilingual support
316
-
317
- Whisper can auto-detect the spoken language and transcribe in that language. The built-in translation model translates **to English only**.
318
-
319
- For best multilingual accuracy, use the `large-v3` model. English-specific models (`*.en.bin`) cannot detect or transcribe other languages.
320
-
321
- **Example — foreign language video with subtitles:**
322
- 1. Ask Claude to generate subtitles with `language=auto` and `translate_to_english=true`
323
- 2. Whisper detects the language and generates a native-language SRT
324
- 3. A second pass generates an English translation SRT
325
- 4. Load either file in VLC via Subtitle → Add Subtitle File
326
-
327
- ---
328
-
329
- ## Designed for free-tier users
330
-
331
- This tool is built to minimize Claude API interactions. The entire transcription workflow — scan, analyze, queue, run, validate — is designed to require as few Claude interactions as possible. Heavy lifting is done locally on your machine.
332
-
333
- ---
334
-
335
- ## Optional environment variables
336
-
337
- | Variable | Description |
338
- |---|---|
339
- | `WHISPER_CLI_PATH` | Path to whisper-cli.exe (required) |
340
- | `WHISPER_MODEL` | Path to model .bin file (required) |
341
- | `WHISPER_THREADS` | CPU thread count override |
342
- | `FFMPEG_PATH` | Path to ffmpeg if not in system PATH |
343
- | `WHISPER_PRIVACY_MODE` | **Planned.** When set to `true`, tool responses return metadata only — no transcript text is returned to Claude's API. For regulated or confidential content. See [PRIVACY.md](PRIVACY.md). |
344
-
345
- ---
346
-
347
- ## Troubleshooting
348
-
349
- See [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for detailed solutions. See [PRIVACY.md](PRIVACY.md) for compliance guidance if you handle regulated content.
350
-
351
- Quick checklist:
352
- - Paths in config use **double backslashes** (`C:\\whisper\\...`)
353
- - `whisper-cli.exe` exists at the configured path
354
- - Model `.bin` file exists at the configured path
355
- - FFmpeg is installed and in PATH (`ffmpeg -version` works)
356
- - Claude Desktop was fully restarted after editing config
357
- - Whisper shows **running** in Settings → Developer
358
-
359
- ---
360
-
361
- ## Security and Privacy
362
-
363
- whisper-windows-mcp is designed with security as a core principle.
364
-
365
- **Audio never leaves your machine.** No audio or video files, no file paths, and no telemetry are ever transmitted to any server. No cloud APIs are required for core functionality.
366
-
367
- **Transcript text and the API boundary.** When a tool response includes transcript text, that text is processed by Claude's API — it leaves your local machine. For most users (public content, podcasts, streaming recordings) this is expected behavior. If you handle medical, legal, financial, or other regulated recordings, see [PRIVACY.md](PRIVACY.md) for compliance guidance and configuration options.
368
-
369
- A `WHISPER_PRIVACY_MODE` environment variable is planned that will restrict all tool responses to metadata only (filename, duration, word count) — no transcript text will be returned to Claude. This is the correct configuration for regulated or confidential content.
370
-
371
- **Input validation.** All file paths are validated before use — UNC paths (`\\server\share`) and directory traversal sequences (`..`) are rejected. Files over 10 GB are rejected to prevent resource exhaustion.
372
-
373
- **Transcript injection awareness.** Audio files can contain spoken content that, when transcribed, resembles instructions. Claude's built-in defenses handle this, but it is worth knowing that transcript content is treated as data — never as instructions — by the MCP server itself.
374
-
375
- **Model downloads are restricted.** The `download_model` tool only downloads from two trusted Hugging Face namespaces (`ggerganov/whisper.cpp` and `ggml-org`). Arbitrary URLs are rejected. Redirects are validated against an allowlist before following.
376
-
377
- **Model switching is sandboxed.** `switch_model` only accepts `.bin` files within the configured models directory. Paths outside that directory are rejected.
378
-
379
- **No new network dependencies.** Model downloads use Node.js built-in `https` — no external HTTP libraries are added to the package.
380
-
381
- ---
382
-
383
- ## License
384
-
385
- MIT
386
-
387
- ---
388
-
389
- ## Contributing
390
-
391
- Pull requests welcome. See [ROADMAP.md](ROADMAP.md) for planned features.
392
-
393
- If you've tested GPU acceleration on hardware not listed above, please open an issue with your results — GPU model, VRAM, model size, and observed throughput.
1
+ # whisper-windows-mcp
2
+
3
+ A Windows-native MCP (Model Context Protocol) server that lets Claude Desktop transcribe audio and video files locally using [whisper.cpp](https://github.com/ggml-org/whisper.cpp) — with GPU acceleration, multilingual support, and batch processing. All transcription runs locally — no audio, video, or file paths ever leave your machine.
4
+
5
+ > **Why does this exist?**
6
+ > The popular `whisper-mcp` package was built for macOS and assumes a Unix environment. It does not work on Windows. This package was written specifically for Windows users who want local AI transcription integrated with Claude Desktop.
7
+
8
+ ---
9
+
10
+ ## What you can do with it
11
+
12
+ Once installed, you can say things like this directly in Claude Desktop:
13
+
14
+ - *"Transcribe C:\Users\Me\Downloads\meeting.mp3"*
15
+ - *"Transcribe this folder of recordings and save each as a text file"*
16
+ - *"Generate Japanese and English subtitles for this video"*
17
+ - *"Start a batch transcription of everything in this folder"*
18
+ - *"How long will it take to transcribe these files?"*
19
+ - *"Check if GPU acceleration is working"*
20
+
21
+ ---
22
+
23
+ ## Requirements
24
+
25
+ 1. **Node.js 18 or later** — [nodejs.org](https://nodejs.org)
26
+ 2. **whisper.cpp binaries with Vulkan GPU support** — see Step 1
27
+ 3. **A Whisper model file** — see Step 2
28
+ 4. **FFmpeg** — required for video files and non-WAV/MP3 audio
29
+
30
+ ---
31
+
32
+ ## Step 1 — Install whisper.cpp binaries
33
+
34
+ ### Option A — Pre-built Vulkan release (recommended)
35
+
36
+ Download `whisper-vulkan-win-x64.zip` from the [releases page](https://github.com/eviscerations/whisper-windows-mcp/releases/tag/v1.4.0).
37
+
38
+ This is a custom-compiled build with **Vulkan GPU acceleration** enabled. Works with AMD, NVIDIA, and Intel GPUs — no vendor-specific SDK required.
39
+
40
+ Extract to `C:\whisper\Release\`. You should end up with:
41
+
42
+ ```
43
+ C:\whisper\Release\whisper-cli.exe
44
+ C:\whisper\Release\ggml-vulkan.dll
45
+ C:\whisper\Release\ggml.dll
46
+ C:\whisper\Release\ggml-base.dll
47
+ C:\whisper\Release\ggml-cpu.dll
48
+ C:\whisper\Release\whisper.dll
49
+ ```
50
+
51
+ GPU acceleration is automatic — no additional configuration needed.
52
+
53
+ ### Option B — Build from source
54
+
55
+ Requires: Git, CMake, Visual Studio Build Tools 2022+ with "Desktop development with C++", Vulkan SDK from [lunarg.com](https://vulkan.lunarg.com/sdk/home#windows).
56
+
57
+ ```
58
+ git clone https://github.com/ggml-org/whisper.cpp
59
+ cd whisper.cpp
60
+ cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
61
+ cmake --build build --config Release --target whisper-cli
62
+ ```
63
+
64
+ Copy the binaries from `build\bin\Release\` to `C:\whisper\Release\`.
65
+
66
+ > **Note:** The official whisper.cpp Windows releases on GitHub do not include a Vulkan build. You must use the pre-built release above or compile from source with `-DGGML_VULKAN=ON`.
67
+
68
+ ---
69
+
70
+ ## Step 2 — Download a Whisper model
71
+
72
+ | Model | Size | Speed | Accuracy | Best for |
73
+ |---|---|---|---|---|
74
+ | `ggml-tiny.en.bin` | 75 MB | Very fast | Basic | Quick tests |
75
+ | `ggml-base.en.bin` | 142 MB | Fast | Good | Everyday English |
76
+ | `ggml-small.en.bin` | 466 MB | Moderate | Better | Important recordings |
77
+ | `ggml-medium.en.bin` | 1.5 GB | Fast on GPU | Very good | Best quality English |
78
+ | `ggml-large-v3-turbo.bin` | 1.6 GB | Fast on GPU | Excellent | **Recommended for English GPU batch work — ~6x faster than large-v3 with minimal accuracy loss** |
79
+ | `ggml-large-v3.bin` | 2.9 GB | Fast on GPU | Excellent | Multilingual, maximum accuracy |
80
+ | `ggml-medium.en-q5_0.bin` | 514 MB | Fast | Very good | **Best CPU-only English option — high accuracy at low memory** |
81
+ | `ggml-large-v3-turbo-q5_0.bin` | 547 MB | Fast | Excellent | **Best CPU-only multilingual option** |
82
+ | `ggml-large-v3-q5_0.bin` | 1.1 GB | Moderate on CPU | Excellent | Multilingual, CPU-friendly |
83
+
84
+ Use `download_model` in Claude Desktop to install any of these directly. For **English-only** use: `large-v3-turbo` (GPU) or `medium.en-q5_0` (CPU) are the best starting points. For **multilingual** use: `large-v3-turbo` or `large-v3-turbo-q5_0` (CPU). English-only models (`*.en.bin`) output `[FOREIGN]` on non-English audio and cannot be used for other languages.
85
+
86
+ ---
87
+
88
+ ## Step 3 — Install FFmpeg
89
+
90
+ FFmpeg is required for video files and non-native audio formats.
91
+
92
+ Install via winget:
93
+ ```
94
+ winget install ffmpeg
95
+ ```
96
+
97
+ Or download from [ffmpeg.org](https://ffmpeg.org/download.html) and add to your PATH.
98
+
99
+ Verify:
100
+ ```
101
+ ffmpeg -version
102
+ ```
103
+
104
+ ---
105
+
106
+ ## Step 4 — Install this MCP server
107
+
108
+ ```
109
+ npm install -g whisper-windows-mcp
110
+ ```
111
+
112
+ ---
113
+
114
+ ## Step 5 — Configure Claude Desktop
115
+
116
+ Open Claude Desktop → Settings → Developer → Edit Config.
117
+
118
+ Add the `whisper` entry:
119
+
120
+ ```json
121
+ {
122
+ "mcpServers": {
123
+ "whisper": {
124
+ "command": "npx",
125
+ "args": ["-y", "whisper-windows-mcp"],
126
+ "env": {
127
+ "WHISPER_CLI_PATH": "C:\\whisper\\Release\\whisper-cli.exe",
128
+ "WHISPER_MODEL": "C:\\whisper\\models\\ggml-medium.en.bin"
129
+ }
130
+ }
131
+ }
132
+ }
133
+ ```
134
+
135
+ Config file location: `C:\Users\YourName\AppData\Roaming\Claude\claude_desktop_config.json`
136
+
137
+ > Use **double backslashes** in all paths.
138
+
139
+ Save and **fully restart** Claude Desktop. You should see **whisper** listed with a green running badge in Settings → Developer.
140
+
141
+ ---
142
+
143
+ ## Step 6 — Verify your setup
144
+
145
+ In Claude Desktop, ask:
146
+
147
+ > *"Check your whisper config"*
148
+
149
+ Then:
150
+
151
+ > *"Check your system hardware"*
152
+
153
+ This confirms your GPU is detected and Vulkan acceleration is active.
154
+
155
+ ---
156
+
157
+ ## Available tools
158
+
159
+ ### `transcribe_audio`
160
+ Transcribe a single file. Supports blocking (default) or background mode for long files.
161
+
162
+ | Parameter | Description |
163
+ |---|---|
164
+ | `file_path` | Absolute path to the file (required) |
165
+ | `language` | Language code (`en`, `ja`, `es`, etc.) or `auto` to detect. Default: `en` |
166
+ | `output_format` | `text` (default), `timestamps`, `json`, or `srt` |
167
+ | `save_to_file` | Save transcript as .txt next to the source file |
168
+ | `background` | Run as detached job — returns a job ID immediately. Use `check_progress` to monitor. Recommended for files over 10 minutes. |
169
+ | `threads` | CPU thread override |
170
+ | `temperature` | Sampling temperature 0.0–1.0. Default 0.0 (deterministic). Higher values reduce hallucination on noisy audio. |
171
+ | `prompt` | Prior context string — improves accuracy for domain-specific vocabulary or speaker names. Example: `"Names: Keemstar, DramaAlert."` |
172
+ | `condition_on_prev_text` | Re-enable context conditioning between segments. Default false. |
173
+ | `beam_size` | Beam search width. Higher = more accurate, slower. Default 5. |
174
+ | `best_of` | Candidate sequences evaluated. Default 5. |
175
+ | `gpu_device` | GPU device index for multi-GPU systems. Default 0. |
176
+ | `processors` | Parallel processor count. Default 1. |
177
+ | `word_timestamps` | One word per timestamped segment. Useful for clip alignment. |
178
+ | `max_segment_length` | Max segment length in characters. |
179
+ | `diarize` | Stereo speaker diarization — requires stereo audio with speakers on separate channels. |
180
+ | `vad_model` | Path to Silero VAD model .bin. Strips silence before transcription — reduces hallucinations on noisy files. |
181
+ | `offset_t` | Start offset in milliseconds. |
182
+ | `duration` | Process duration in milliseconds from offset. |
183
+
184
+ ---
185
+
186
+ ### `check_progress`
187
+ Monitor a background transcription job started with `transcribe_audio` (background=true).
188
+
189
+ Returns elapsed time, last processed timestamp, percentage, and the full transcript when complete.
190
+
191
+ | Parameter | Description |
192
+ |---|---|
193
+ | `job_id` | Job ID returned by `transcribe_audio` |
194
+
195
+ ---
196
+
197
+ ### `start_batch`
198
+ Automated sequential batch transcription of all untranscribed files in a folder. Sorts by duration (shortest first), processes one at a time as background jobs, validates each output.
199
+
200
+ | Parameter | Description |
201
+ |---|---|
202
+ | `folder_path` | Path to folder (required) |
203
+ | `language` | Language code. Default: `en` |
204
+ | `threads` | CPU thread override |
205
+
206
+ ---
207
+
208
+ ### `check_batch_progress`
209
+ Monitor a running batch. Automatically advances to the next file when the current one finishes. Returns overall progress, current file with timestamp, ETA, and any failed files.
210
+
211
+ | Parameter | Description |
212
+ |---|---|
213
+ | `batch_id` | Batch ID returned by `start_batch` |
214
+
215
+ ---
216
+
217
+ ### `transcribe_batch` (interactive)
218
+ Process files one at a time with a preview and confirmation before each. Useful when you want to review as you go.
219
+
220
+ | Parameter | Description |
221
+ |---|---|
222
+ | `folder_path` | Path to folder (required) |
223
+ | `file_index` | Which file to process (1-based). Omit to list files first. |
224
+ | `language` | Language code. Default: `en` |
225
+ | `recursive` | Include subfolders |
226
+
227
+ ---
228
+
229
+ ### `generate_subtitles`
230
+ Generate SRT subtitle files. Supports automatic language detection and English translation output.
231
+
232
+ | Parameter | Description |
233
+ |---|---|
234
+ | `file_path` | Path to file (required) |
235
+ | `language` | Language code or `auto` to detect. Default: `en` |
236
+ | `translate_to_english` | Also generate an English translation `.en.srt`. Only applies when source is not English. |
237
+ | `threads` | CPU thread override |
238
+
239
+ When both native and translation are requested, two files are saved next to the source:
240
+ - `filename.ja.srt` — original language
241
+ - `filename.en.srt` — English translation
242
+
243
+ > Whisper's built-in translation only translates **to English**. For other target languages, translate the .srt file contents separately.
244
+
245
+ ---
246
+
247
+ ### `analyze_media`
248
+ Analyze files before committing to transcription. Returns duration, size, codec, and estimated transcription time on CPU and GPU. For folders, shows all files in a sortable table with transcription status.
249
+
250
+ | Parameter | Description |
251
+ |---|---|
252
+ | `path` | Path to a single file or folder (required) |
253
+ | `sort_by` | For folders: `duration` (default), `name`, or `size` |
254
+
255
+ ---
256
+
257
+ ### `check_config`
258
+ Verify whisper-cli.exe, the model file, and FFmpeg are all accessible. Run this first if anything is failing.
259
+
260
+ ---
261
+
262
+ ### `list_models`
263
+ List all Whisper model files installed in your models directory. Shows filename, size, whether it is currently active, quantization status, and recommended use case. No network calls — reads local filesystem only.
264
+
265
+ ---
266
+
267
+ ### `download_model`
268
+ Download a Whisper model directly from Hugging Face into your models directory. Accepts a model name (e.g. `large-v3-turbo`, `medium.en-q5_0`) and handles the download automatically. Only downloads from trusted Hugging Face namespaces. After downloading, use `switch_model` to activate it.
269
+
270
+ | Parameter | Description |
271
+ |---|---|
272
+ | `model_name` | Model name to download, e.g. `large-v3-turbo`, `large-v3-turbo-q5_0`, `medium.en-q5_0` |
273
+
274
+ ---
275
+
276
+ ### `switch_model`
277
+ Switch the active Whisper model for the current session without restarting Claude Desktop. Change is session-scoped — does not persist after restart. To make permanent, update `WHISPER_MODEL` in your config.
278
+
279
+ | Parameter | Description |
280
+ |---|---|
281
+ | `model_name` | Model filename (e.g. `ggml-large-v3-turbo.bin`) or full path. Must be a `.bin` file in the configured models directory. |
282
+
283
+ ---
284
+
285
+ ### `check_system`
286
+ Detect GPU hardware and verify Vulkan acceleration is available. Reports GPU name, VRAM, whether `ggml-vulkan.dll` is present, and recommends the best model size for your hardware.
287
+
288
+ ---
289
+
290
+ ## Supported formats
291
+
292
+ | Type | Formats |
293
+ |---|---|
294
+ | Native (no conversion) | `mp3`, `wav` |
295
+ | Video (auto-converted via FFmpeg) | `mp4`, `mkv`, `avi`, `mov`, `webm`, `flv`, `wmv`, `m4v`, `ts`, `3gp` |
296
+ | Audio (auto-converted via FFmpeg) | `m4a`, `ogg`, `flac` |
297
+
298
+ ---
299
+
300
+ ## GPU acceleration
301
+
302
+ The pre-built Vulkan release enables GPU acceleration automatically. Tested on AMD Radeon RX Vega 56 (GCN 5th gen). Any GPU with Vulkan 1.0+ support should work, including NVIDIA and Intel Arc.
303
+
304
+ **Performance comparison (medium.en model, ~5 minute audio file):**
305
+
306
+ | Hardware | Time |
307
+ |---|---|
308
+ | CPU only (Ryzen 7 2700x, 8 threads) | 8–12 minutes |
309
+ | GPU (Vega 56 via Vulkan) | 20–40 seconds |
310
+
311
+ GPU utilization during transcription is typically 15–20%, dropping back to idle between files. CPU stays around 15%.
312
+
313
+ ---
314
+
315
+ ## Multilingual support
316
+
317
+ Whisper can auto-detect the spoken language and transcribe in that language. The built-in translation model translates **to English only**.
318
+
319
+ For best multilingual accuracy, use the `large-v3` model. English-specific models (`*.en.bin`) cannot detect or transcribe other languages.
320
+
321
+ **Example — foreign language video with subtitles:**
322
+ 1. Ask Claude to generate subtitles with `language=auto` and `translate_to_english=true`
323
+ 2. Whisper detects the language and generates a native-language SRT
324
+ 3. A second pass generates an English translation SRT
325
+ 4. Load either file in VLC via Subtitle → Add Subtitle File
326
+
327
+ ---
328
+
329
+ ## Designed for free-tier users
330
+
331
+ This tool is built to minimize Claude API interactions. The entire transcription workflow — scan, analyze, queue, run, validate — is designed to require as few Claude interactions as possible. Heavy lifting is done locally on your machine.
332
+
333
+ ---
334
+
335
+ ## Optional environment variables
336
+
337
+ | Variable | Description |
338
+ |---|---|
339
+ | `WHISPER_CLI_PATH` | Path to whisper-cli.exe (required) |
340
+ | `WHISPER_MODEL` | Path to model .bin file (required) |
341
+ | `WHISPER_THREADS` | CPU thread count override |
342
+ | `FFMPEG_PATH` | Path to ffmpeg if not in system PATH |
343
+ | `WHISPER_PRIVACY_MODE` | **Planned.** When set to `true`, tool responses return metadata only — no transcript text is returned to Claude's API. For regulated or confidential content. See [PRIVACY.md](PRIVACY.md). |
344
+
345
+ ---
346
+
347
+ ## Troubleshooting
348
+
349
+ See [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for detailed solutions. See [PRIVACY.md](PRIVACY.md) for compliance guidance if you handle regulated content.
350
+
351
+ Quick checklist:
352
+ - Paths in config use **double backslashes** (`C:\\whisper\\...`)
353
+ - `whisper-cli.exe` exists at the configured path
354
+ - Model `.bin` file exists at the configured path
355
+ - FFmpeg is installed and in PATH (`ffmpeg -version` works)
356
+ - Claude Desktop was fully restarted after editing config
357
+ - Whisper shows **running** in Settings → Developer
358
+
359
+ ---
360
+
361
+ ## Security and Privacy
362
+
363
+ whisper-windows-mcp is designed with security as a core principle.
364
+
365
+ **Audio never leaves your machine.** No audio or video files, no file paths, and no telemetry are ever transmitted to any server. No cloud APIs are required for core functionality.
366
+
367
+ **Transcript text and the API boundary.** When a tool response includes transcript text, that text is processed by Claude's API — it leaves your local machine. For most users (public content, podcasts, streaming recordings) this is expected behavior. If you handle medical, legal, financial, or other regulated recordings, see [PRIVACY.md](PRIVACY.md) for compliance guidance and configuration options.
368
+
369
+ A `WHISPER_PRIVACY_MODE` environment variable is planned that will restrict all tool responses to metadata only (filename, duration, word count) — no transcript text will be returned to Claude. This is the correct configuration for regulated or confidential content.
370
+
371
+ **Input validation.** All file paths are validated before use — UNC paths (`\\server\share`) and directory traversal sequences (`..`) are rejected. Files over 10 GB are rejected to prevent resource exhaustion.
372
+
373
+ **Transcript injection awareness.** Audio files can contain spoken content that, when transcribed, resembles instructions. Claude's built-in defenses handle this, but it is worth knowing that transcript content is treated as data — never as instructions — by the MCP server itself.
374
+
375
+ **Model downloads are restricted.** The `download_model` tool only downloads from two trusted Hugging Face namespaces (`ggerganov/whisper.cpp` and `ggml-org`). Arbitrary URLs are rejected. Redirects are validated against an allowlist before following.
376
+
377
+ **Model switching is sandboxed.** `switch_model` only accepts `.bin` files within the configured models directory. Paths outside that directory are rejected.
378
+
379
+ **No new network dependencies.** Model downloads use Node.js built-in `https` — no external HTTP libraries are added to the package.
380
+
381
+ ---
382
+
383
+ ## License
384
+
385
+ **Non-commercial use:** MIT — free for personal, educational, and non-commercial use. See [LICENSE](LICENSE).
386
+
387
+ **Commercial use:** A separate commercial license is required for any business, professional, or revenue-generating use. See [LICENSE-COMMERCIAL.md](LICENSE-COMMERCIAL.md) for terms and contact information.
388
+
389
+ ## Contributing
390
+
391
+ Pull requests welcome. See [ROADMAP.md](ROADMAP.md) for planned features.
392
+
393
+ If you've tested GPU acceleration on hardware not listed above, please open an issue with your results — GPU model, VRAM, model size, and observed throughput.