whisper-windows-mcp 1.3.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,290 +1,229 @@
1
- # whisper-windows-mcp
2
-
3
- A Windows-native MCP (Model Context Protocol) server that lets Claude Desktop transcribe audio and video files locally using [whisper.cpp](https://github.com/ggerganov/whisper.cpp) — no internet connection required, no data sent to the cloud.
4
-
5
- > **Why does this exist?**
6
- > The popular `whisper-mcp` package was built for macOS and assumes a Unix environment. It does not work on Windows. This package was written specifically for Windows users who want the same local transcription functionality in Claude Desktop.
7
-
8
- ---
9
-
10
- ## What you can do with it
11
-
12
- Once installed, you can say things like this directly in Claude Desktop:
13
-
14
- - *"Transcribe C:\Users\Me\Downloads\meeting.mp3"*
15
- - *"Transcribe C:\Users\Me\Videos\interview.mp4"* — video files work directly, no conversion needed
16
- - *"Transcribe this recording and summarise the key points"*
17
- - *"Transcribe with timestamps so I can find specific moments"*
18
- - *"Generate subtitles for C:\Users\Me\Videos\lecture.mp4"*
19
- - *"Transcribe all files in C:\Users\Me\Videos\clips\ one at a time"*
20
-
21
- Everything runs on your own machine. No audio ever leaves your computer.
22
-
23
- ---
24
-
25
- ## Requirements
26
-
27
- | Requirement | Purpose | Download |
28
- |---|---|---|
29
- | Node.js 18+ | Runs the MCP server | [nodejs.org](https://nodejs.org/en/download) |
30
- | whisper.cpp | The transcription engine | [Latest release](https://github.com/ggerganov/whisper.cpp/releases/latest) |
31
- | A Whisper model file | The AI model | See Step 2 below |
32
- | FFmpeg | Video file support | [ffmpeg.org](https://ffmpeg.org/download.html) |
33
-
34
- > FFmpeg is optional if you only use MP3/WAV files, but required for MP4, MKV, AVI, MOV and other video formats.
35
-
36
- ---
37
-
38
- ## Step 1 — Install whisper.cpp
39
-
40
- 1. Go to the [whisper.cpp latest release](https://github.com/ggerganov/whisper.cpp/releases/latest)
41
- 2. Download the file named **`whisper-bin-x64.zip`** (look for `win` and `x64` in the filename)
42
- 3. Extract the ZIP and move the contents to **`C:\whisper\Release\`** — create this folder if it doesn't exist
43
-
44
- ✅ You should now have **`C:\whisper\Release\whisper-cli.exe`**
45
-
46
- > **Why this path?** You can install whisper.cpp anywhere, but `C:\whisper\Release\` matches the default config below and means less to edit later.
47
-
48
- ---
49
-
50
- ## Step 2 — Download a Whisper model
51
-
52
- Click a link below to download directly:
53
-
54
- | Model | Download | Size | Speed | Best for |
55
- |---|---|---|---|---|
56
- | tiny.en | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-tiny.en.bin) | 75 MB | Very fast | Quick tests |
57
- | **base.en** | [**Download**](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en.bin) | 142 MB | Fast | **Recommended starting point** |
58
- | small.en | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.en.bin) | 466 MB | Moderate | Better accuracy |
59
- | medium.en | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-medium.en.bin) | 1.5 GB | Slow | High accuracy |
60
- | large-v3 | [Download](https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3.bin) | 2.9 GB | Very slow | Maximum accuracy |
61
-
62
- Save the downloaded `.bin` file to **`C:\whisper\models\`** — create this folder if it doesn't exist.
63
-
64
- ✅ You should now have something like **`C:\whisper\models\ggml-base.en.bin`**
65
-
66
- ---
67
-
68
- ## Step 3 — Install Node.js
69
-
70
- If you don't already have Node.js:
71
-
72
- 1. Go to [nodejs.org](https://nodejs.org/en/download) and download the **Windows Installer (.msi)** — choose the LTS version
73
- 2. Run the installer and accept all defaults
74
-
75
- ✅ Verify: open Command Prompt and run `node --version` — you should see something like `v20.x.x`
76
-
77
- ---
78
-
79
- ## Step 4 — Install FFmpeg (recommended)
80
-
81
- FFmpeg is required for video files (MP4, MKV, AVI, MOV, etc.).
82
-
83
- 1. Go to [ffmpeg.org/download.html](https://ffmpeg.org/download.html) and download a Windows build
84
- 2. Extract and move the `bin` folder contents (or the whole folder) somewhere permanent, e.g. `C:\ffmpeg\bin\`
85
- 3. Add `C:\ffmpeg\bin` to your system PATH:
86
- - Press **Win + S** → search **Environment Variables** → open it
87
- - Under **User Variables**, select **Path** → **Edit** → **New**
88
- - Add `C:\ffmpeg\bin` → click OK on all dialogs
89
-
90
- ✅ Verify: open a new Command Prompt and run `ffmpeg -version`
91
-
92
- ---
93
-
94
- ## Step 5 — Configure Claude Desktop
95
-
96
- 1. Open Claude Desktop → **Settings → Developer → Edit Config**
97
- 2. Add the following (or merge the `mcpServers` block if you have other servers):
98
-
99
- ```json
100
- {
101
- "mcpServers": {
102
- "whisper": {
103
- "command": "npx",
104
- "args": ["-y", "whisper-windows-mcp"],
105
- "env": {
106
- "WHISPER_CLI_PATH": "C:\\whisper\\Release\\whisper-cli.exe",
107
- "WHISPER_MODEL": "C:\\whisper\\models\\ggml-base.en.bin"
108
- }
109
- }
110
- }
111
- }
112
- ```
113
-
114
- > ⚠️ **Path format:** In the JSON config, all backslashes must be doubled (`\\`). This is a JSON requirement. When typing paths into Claude in chat, use normal single backslashes.
115
-
116
- 3. If you downloaded a different model, update `ggml-base.en.bin` to match your filename
117
- 4. Save the file, fully quit Claude Desktop, and reopen it
118
- 5. Go to **Settings → Developer** — you should see **whisper** with a green **running** badge
119
-
120
- ---
121
-
122
- ## Step 6 — Test it
123
-
124
- In Claude Desktop, type:
125
-
126
- > *"Can you check your whisper config?"*
127
-
128
- Claude will verify that everything is found correctly. Then try:
129
-
130
- > *"Please transcribe C:\Users\YourName\Downloads\recording.mp3"*
131
-
132
- ---
133
-
134
- ## Available tools
135
-
136
- | Tool | What it does |
137
- |---|---|
138
- | `transcribe_audio` | Transcribe a single file — audio or video |
139
- | `transcribe_batch` | Transcribe all files in a folder, one at a time with preview |
140
- | `generate_subtitles` | Generate an `.srt` subtitle file next to the source |
141
- | `check_config` | Verify all paths and FFmpeg are working |
142
-
143
- ---
144
-
145
- ## Output formats
146
-
147
- | Format | What you get | Ask Claude... |
148
- |---|---|---|
149
- | `text` (default) | Plain transcript, no timestamps | *"Transcribe this file"* |
150
- | `timestamps` | Transcript with `[00:00:00 --> 00:00:05]` time codes | *"Transcribe with timestamps"* |
151
- | `json` | Structured data | *"Transcribe as JSON"* |
152
- | `srt` | Subtitle file saved next to source | *"Generate subtitles for..."* |
153
-
154
- ---
155
-
156
- ## Supported file formats
157
-
158
- **Audio (native):** MP3, WAV
159
-
160
- **Audio (via FFmpeg):** M4A, FLAC, OGG
161
-
162
- **Video (via FFmpeg):** MP4, MKV, AVI, MOV, WebM, FLV, WMV, M4V
163
-
164
- > H.264 and H.265/HEVC video both work. FFmpeg must be installed and in your PATH for any video or non-MP3 audio format.
165
-
166
- ---
167
-
168
- ## Batch transcription
169
-
170
- To transcribe multiple files in a folder interactively:
171
-
172
- > *"Transcribe all files in C:\Users\Me\Videos\clips\"*
173
-
174
- Claude will list all detected files (with checkmarks on already-completed ones), then process them one at a time, showing you a preview of each transcript before moving to the next.
175
-
176
- **For large unattended overnight batches**, use whisper-cli directly from the command line — see [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for the syntax. This is more reliable than running through Claude for very large jobs.
177
-
178
- ---
179
-
180
- ## Transcription speed
181
-
182
- Whisper runs on CPU by default. Rough estimates for a 1-hour recording:
183
-
184
- | Model | Approximate time (CPU) |
185
- |---|---|
186
- | tiny.en | 5–10 minutes |
187
- | base.en | 10–20 minutes |
188
- | small.en | 20–35 minutes |
189
- | medium.en | 35–60 minutes |
190
- | large-v3 | 60–120 minutes |
191
-
192
- You can increase thread count by adding `"WHISPER_THREADS": "12"` to the `env` block in your config (replace `12` with however many threads you want — up to your CPU's logical core count).
193
-
194
- > **GPU acceleration** for AMD (ROCm) and NVIDIA (CUDA) on Windows is planned for a future update.
195
-
196
- ---
197
-
198
- ## Converting video to audio (optional)
199
-
200
- whisper-windows-mcp handles video files automatically via FFmpeg, so manual conversion is no longer required. However, if you want a smaller audio-only file for any reason, VLC makes it easy:
201
-
202
- 1. Open VLC → **Media → Convert / Save**
203
- 2. Click **Add**, select your video, then **Convert / Save**
204
- 3. Under **Profile**, choose **Audio - MP3**
205
- 4. Set a destination filename and click **Start**
206
-
207
- ---
208
-
209
- ## Full config example
210
-
211
- ```json
212
- {
213
- "preferences": {
214
- "coworkWebSearchEnabled": true
215
- },
216
- "mcpServers": {
217
- "whisper": {
218
- "command": "npx",
219
- "args": ["-y", "whisper-windows-mcp"],
220
- "env": {
221
- "WHISPER_CLI_PATH": "C:\\whisper\\Release\\whisper-cli.exe",
222
- "WHISPER_MODEL": "C:\\whisper\\models\\ggml-base.en.bin",
223
- "WHISPER_THREADS": "8",
224
- "FFMPEG_PATH": "ffmpeg"
225
- }
226
- }
227
- }
228
- }
229
- ```
230
-
231
- Config file location:
232
- ```
233
- C:\Users\YourUsername\AppData\Roaming\Claude\claude_desktop_config.json
234
- ```
235
-
236
- > The `AppData` folder is hidden by default. To show it: File Explorer → **View → Show → Hidden items**
237
-
238
- ---
239
-
240
- ## Tested on
241
-
242
- - Windows 10 Pro (10.0.19045)
243
- - Windows 11 — untested, feedback welcome via [Issues](../../issues)
244
-
245
- ---
246
-
247
- ## Troubleshooting
248
-
249
- See [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for detailed solutions including how to run large overnight batch jobs from the command line.
250
-
251
- Quick checklist:
252
-
253
- - [ ] Config paths use **double backslashes** (`C:\\whisper\\...`)
254
- - [ ] `whisper-cli.exe` exists at the path specified
255
- - [ ] The model `.bin` file exists at the path specified
256
- - [ ] Claude Desktop was **fully restarted** after editing the config
257
- - [ ] Whisper shows **running** in Settings → Developer
258
- - [ ] FFmpeg is in PATH if using video files
259
-
260
- ---
261
-
262
- ## Roadmap
263
-
264
- - [ ] AMD GPU acceleration (ROCm)
265
- - [ ] NVIDIA GPU acceleration (CUDA)
266
- - [ ] Speaker diarization (automatic A/B speaker identification)
267
- - [ ] Translation to English from other languages
268
- - [ ] Unattended background batch processing
269
-
270
- ---
271
-
272
- ## Support this project
273
-
274
- If this tool saved you time and you'd like to support continued development:
275
-
276
- - ⭐ **Star this repo** — it helps others find it
277
- - 💬 **Open an issue** if you find a bug or have a feature request
278
- - 💖 **Sponsor** — [GitHub Sponsors](https://github.com/sponsors/eviscerations) | [Ko-fi](https://ko-fi.com) | [Patreon](https://patreon.com)
279
-
280
- ---
281
-
282
- ## License
283
-
284
- MIT — free to use, modify, and distribute.
285
-
286
- ---
287
-
288
- ## Contributing
289
-
290
- Pull requests welcome. GPU acceleration for AMD or NVIDIA especially appreciated. Windows 11 feedback welcome via Issues.
1
+ # whisper-windows-mcp
2
+
3
+ A Windows-native MCP (Model Context Protocol) server that lets Claude Desktop transcribe audio files locally using [whisper.cpp](https://github.com/ggml-org/whisper.cpp) — no internet connection required, no data sent to the cloud.
4
+
5
+ > **Why does this exist?**
6
+ > The popular `whisper-mcp` package was built for macOS and assumes a Unix environment. It does not work on Windows. This package was written specifically for Windows users who want the same functionality.
7
+
8
+ ---
9
+
10
+ ## What you can do with it
11
+
12
+ Once installed, you can say things like this directly in Claude Desktop:
13
+
14
+ - *"Transcribe C:\Users\Me\Downloads\meeting.mp3"*
15
+ - *"Transcribe this recording and give me a summary"*
16
+ - *"Transcribe with timestamps so I can find specific moments"*
17
+ - *"Generate subtitles for this video"*
18
+
19
+ Everything runs on your own machine. No audio ever leaves your computer.
20
+
21
+ ---
22
+
23
+ ## Requirements
24
+
25
+ Before installing this package, you need three things set up on your Windows machine:
26
+
27
+ 1. **Node.js 18 or later** — [download from nodejs.org](https://nodejs.org)
28
+ 2. **whisper.cpp binaries** — the actual transcription engine (see Step 1 below)
29
+ 3. **A Whisper model file** — the AI model that does the transcription (see Step 2 below)
30
+
31
+ ---
32
+
33
+ ## Step 1 — Install whisper.cpp binaries
34
+
35
+ ### Option A — Pre-built Vulkan release (recommended)
36
+
37
+ Download `whisper-vulkan-win-x64.zip` from the [releases page](https://github.com/eviscerations/whisper-windows-mcp/releases).
38
+
39
+ This is a custom-compiled build with **Vulkan GPU acceleration** enabled. It works with AMD, NVIDIA, and Intel GPUs on Windows — no vendor-specific SDK required.
40
+
41
+ Extract the zip to `C:\whisper\Release\`. You should end up with these files:
42
+
43
+ ```
44
+ C:\whisper\Release\whisper-cli.exe
45
+ C:\whisper\Release\ggml-vulkan.dll
46
+ C:\whisper\Release\ggml.dll
47
+ C:\whisper\Release\ggml-base.dll
48
+ C:\whisper\Release\ggml-cpu.dll
49
+ C:\whisper\Release\whisper.dll
50
+ ```
51
+
52
+ **GPU acceleration is automatic** — if a supported GPU is present, whisper.cpp will use it. No additional configuration needed.
53
+
54
+ ### Option B — Build from source
55
+
56
+ If you prefer to compile your own binary (advanced users):
57
+
58
+ **Prerequisites:** Git, CMake, Visual Studio Build Tools 2022+ with "Desktop development with C++", Vulkan SDK from [lunarg.com](https://vulkan.lunarg.com/sdk/home#windows).
59
+
60
+ ```
61
+ git clone https://github.com/ggml-org/whisper.cpp
62
+ cd whisper.cpp
63
+ cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
64
+ cmake --build build --config Release --target whisper-cli
65
+ ```
66
+
67
+ Copy the resulting binaries from `build\bin\Release\` to `C:\whisper\Release\`.
68
+
69
+ > **Note:** The default whisper.cpp Windows release on GitHub does not include a Vulkan build. You must either use the pre-built release above or compile from source with `-DGGML_VULKAN=ON`.
70
+
71
+ ---
72
+
73
+ ## Step 2 — Download a Whisper model
74
+
75
+ Models are downloaded from Hugging Face. Choose one based on your needs:
76
+
77
+ | Model | File size | Speed | Accuracy | Recommended for |
78
+ |---|---|---|---|---|
79
+ | `ggml-tiny.en.bin` | 75 MB | Very fast | Basic | Quick tests |
80
+ | `ggml-base.en.bin` | 142 MB | Fast | Good | Everyday use |
81
+ | `ggml-small.en.bin` | 466 MB | Moderate | Better | Important recordings |
82
+ | `ggml-medium.en.bin` | 1.5 GB | Fast on GPU | Very good | Best quality |
83
+ | `ggml-large-v3.bin` | 2.9 GB | Fast on GPU | Excellent | Maximum accuracy |
84
+
85
+ **For most people, `base.en` or `small.en` is the best starting point.** With GPU acceleration, `medium.en` and `large-v3` become practical for everyday use.
86
+
87
+ Download your chosen model from:
88
+
89
+ ```
90
+ https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en.bin
91
+ ```
92
+
93
+ Save it to `C:\whisper\models\` — create that folder if it doesn't exist.
94
+
95
+ ---
96
+
97
+ ## Step 3 — Install this MCP server
98
+
99
+ ```
100
+ npm install -g whisper-windows-mcp
101
+ ```
102
+
103
+ Or use `npx` directly in your config (see Step 4).
104
+
105
+ ---
106
+
107
+ ## Step 4 — Configure Claude Desktop
108
+
109
+ 1. Open Claude Desktop
110
+ 2. Go to **Settings → Developer → Edit Config**
111
+ 3. Add the whisper-windows-mcp entry:
112
+
113
+ ```json
114
+ {
115
+ "mcpServers": {
116
+ "whisper": {
117
+ "command": "npx",
118
+ "args": ["-y", "whisper-windows-mcp"],
119
+ "env": {
120
+ "WHISPER_CLI_PATH": "C:\\whisper\\Release\\whisper-cli.exe",
121
+ "WHISPER_MODEL": "C:\\whisper\\models\\ggml-base.en.bin"
122
+ }
123
+ }
124
+ }
125
+ }
126
+ ```
127
+
128
+ > **Important:** If your `claude_desktop_config.json` already has other content, add the `"mcpServers"` block inside the existing `{}` — don't replace the whole file.
129
+
130
+ 4. Save the file and fully restart Claude Desktop.
131
+ 5. Go to **Settings → Developer** — you should see **whisper** listed with a green **running** badge.
132
+
133
+ ---
134
+
135
+ ## Step 5 — Test it
136
+
137
+ In Claude Desktop, type:
138
+
139
+ > *"Can you check your whisper config?"*
140
+
141
+ Claude will use the `check_config` tool to verify everything is set up correctly. Then try a transcription:
142
+
143
+ > *"Please transcribe C:\Users\YourName\Downloads\recording.mp3"*
144
+
145
+ ---
146
+
147
+ ## GPU acceleration
148
+
149
+ The pre-built Vulkan release enables GPU acceleration automatically. No flags or configuration required — whisper.cpp detects your GPU at startup and uses it if available.
150
+
151
+ **Confirmed working:** AMD Radeon RX Vega 56 (GCN 5th gen), and any GPU with Vulkan 1.0+ support.
152
+
153
+ **Performance comparison with medium.en model:**
154
+
155
+ | Hardware | ~5 min audio file |
156
+ |---|---|
157
+ | CPU only (Ryzen 7 2700x) | ~8–12 minutes |
158
+ | GPU (Vega 56 via Vulkan) | ~20–40 seconds |
159
+
160
+ GPU utilization during transcription is typically 15–30% — efficient bursts, not sustained load.
161
+
162
+ ---
163
+
164
+ ## Output formats
165
+
166
+ - **text** (default) — plain transcript
167
+ - **timestamps** — transcript with `[00:00:00 --> 00:00:05]` time codes
168
+ - **json** — structured output
169
+ - **srt** — subtitle file saved next to the source file
170
+
171
+ ---
172
+
173
+ ## Full config example
174
+
175
+ ```json
176
+ {
177
+ "preferences": {
178
+ "coworkWebSearchEnabled": true
179
+ },
180
+ "mcpServers": {
181
+ "whisper": {
182
+ "command": "npx",
183
+ "args": ["-y", "whisper-windows-mcp"],
184
+ "env": {
185
+ "WHISPER_CLI_PATH": "C:\\whisper\\Release\\whisper-cli.exe",
186
+ "WHISPER_MODEL": "C:\\whisper\\models\\ggml-medium.en.bin"
187
+ }
188
+ }
189
+ }
190
+ }
191
+ ```
192
+
193
+ Config file location: `C:\Users\YourUsername\AppData\Roaming\Claude\claude_desktop_config.json`
194
+
195
+ ---
196
+
197
+ ## Optional environment variables
198
+
199
+ | Variable | Description |
200
+ |---|---|
201
+ | `WHISPER_CLI_PATH` | Path to whisper-cli.exe (required) |
202
+ | `WHISPER_MODEL` | Path to model .bin file (required) |
203
+ | `WHISPER_THREADS` | CPU thread count override |
204
+ | `FFMPEG_PATH` | Path to ffmpeg if not in system PATH |
205
+
206
+ ---
207
+
208
+ ## Troubleshooting
209
+
210
+ See [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for detailed solutions.
211
+
212
+ Quick checklist:
213
+ - Paths in config use **double backslashes** (`C:\\whisper\\...`)
214
+ - `whisper-cli.exe` exists at the path specified
215
+ - The model `.bin` file exists at the path specified
216
+ - Claude Desktop was fully restarted after editing config
217
+ - The whisper server shows **running** in Settings → Developer
218
+
219
+ ---
220
+
221
+ ## License
222
+
223
+ MIT — free to use, modify, and distribute.
224
+
225
+ ---
226
+
227
+ ## Contributing
228
+
229
+ Pull requests welcome. See [ROADMAP.md](ROADMAP.md) for planned features.
package/ROADMAP.md ADDED
@@ -0,0 +1,129 @@
1
+ # whisper-windows-mcp — Roadmap
2
+
3
+ Current version is v1.4.0. GPU acceleration via Vulkan is now working. This document tracks what has been completed and what remains.
4
+
5
+ ---
6
+
7
+ ## Completed
8
+
9
+ ### ✅ Priority 1 — GPU Acceleration (v1.4.0)
10
+
11
+ Compiled whisper.cpp from source with `-DGGML_VULKAN=ON` using Visual Studio Build Tools 2026 and Vulkan SDK 1.4.341.1. Pre-built Vulkan binaries are now distributed as a release asset (`whisper-vulkan-win-x64.zip`).
12
+
13
+ **Results:** AMD Radeon RX Vega 56 at ~16% GPU utilization, 1.2GB VRAM, CPU at ~15% during transcription. A ~5 minute file that previously took 8–12 minutes on CPU now completes in 20–40 seconds.
14
+
15
+ The official whisper.cpp Windows releases do not include a Vulkan build ([issue #3673](https://github.com/ggml-org/whisper.cpp/issues/3673)). The pre-built release in this repo fills that gap for AMD and Intel GPU users.
16
+
17
+ ### ✅ Priority 2 — Process Lock (v1.3.1)
18
+
19
+ Added `isWhisperRunning()` check using `tasklist /FI` before any transcription spawn. If `whisper-cli.exe` is already running, returns a clear error with Task Manager instructions rather than spawning a second competing process.
20
+
21
+ ---
22
+
23
+ ## Known Issues (Remaining)
24
+
25
+ ### 3. 4-Minute Claude MCP Timeout
26
+ The Claude web client cuts MCP connections after ~4 minutes. Whisper continues running in the background after the timeout fires, but Claude can't confirm completion. With GPU acceleration this is less frequently hit, but still a concern for very long files or large models.
27
+
28
+ ### 5. No Progress Visibility
29
+ The user has no indicator that transcription is happening or how far along it is. whisper-cli.exe outputs segment timestamps to stderr as it processes — the MCP should pipe and expose these.
30
+
31
+ ### 6. Background Batch Non-Functional
32
+ A background batch mode was attempted in a previous version but stripped due to no visible feedback. Needs a proper detached process architecture with job state files.
33
+
34
+ ### 7. No File Pre-Analysis
35
+ No way to know file duration or size before processing starts.
36
+
37
+ ---
38
+
39
+ ## Roadmap
40
+
41
+ ### Priority 3 — File Pre-Analysis Tool
42
+
43
+ New tool `analyze_media` using FFprobe:
44
+
45
+ ```
46
+ ffprobe -v quiet -print_format json -show_format -show_streams <file>
47
+ ```
48
+
49
+ Returns: duration, file size, codec, bitrate, estimated transcription time (CPU and GPU).
50
+
51
+ ---
52
+
53
+ ### Priority 4 — Progress Visibility
54
+
55
+ whisper-cli.exe outputs segment timestamps to stderr (e.g. `[00:01:30 --> 00:01:35]`). The MCP should:
56
+
57
+ 1. Pipe stderr to a log file during processing
58
+ 2. Expose a `check_progress` tool returning last timestamp, percentage complete, estimated time remaining, and whether the process is still running
59
+
60
+ ---
61
+
62
+ ### Priority 5 — Timeout Workaround (Detached Process Architecture)
63
+
64
+ Rearchitect transcription to use fully detached background processes:
65
+
66
+ 1. `transcribe_audio` → spawn whisper as detached process → write `job.json` → return immediately with job ID
67
+ 2. `check_progress` → read `job.json` → check PID → read log → return status + percentage
68
+ 3. When complete → read and return transcript
69
+
70
+ This eliminates the 4-minute timeout problem entirely and unlocks Priority 6.
71
+
72
+ ---
73
+
74
+ ### Priority 6 — Sequential Batch with Validation
75
+
76
+ Rebuild batch mode on top of Priority 5:
77
+
78
+ 1. `transcribe_batch` → runs `analyze_media` on folder → sorts by duration → processes one file at a time using detached process
79
+ 2. After each file: validate .txt exists, is non-empty, line count proportional to duration
80
+ 3. Flag suspect outputs for re-run
81
+ 4. `check_batch_progress` returns: files done, remaining, current file, ETA, any failed files
82
+
83
+ ---
84
+
85
+ ### Priority 7 — Multi-Language Support and Translation
86
+
87
+ Expose `--language` and `--translate` flags properly:
88
+
89
+ - `language` parameter: auto-detect (default) or specify (`ja`, `es`, `de`, etc.)
90
+ - `translate_to_english`: boolean — uses whisper's built-in translation model
91
+ - Dual output: two whisper passes, two output files
92
+
93
+ ---
94
+
95
+ ### Priority 8 — Filename-Based References Throughout
96
+
97
+ All tool outputs must reference the full source filename at all times. Never use positional indices as the primary identifier.
98
+
99
+ ---
100
+
101
+ ### Priority 9 — System Diagnostics Tool
102
+
103
+ New tool `check_system`:
104
+
105
+ - GPU vendor and model (via `wmic path win32_VideoController`)
106
+ - Whether `ggml-vulkan.dll` is present alongside `whisper-cli.exe`
107
+ - Recommended model size for available VRAM
108
+ - Estimated throughput based on hardware profile
109
+ - Actionable guidance if GPU binary is missing
110
+
111
+ ---
112
+
113
+ ## Design Principles
114
+
115
+ **Minimize Claude API usage.** The entire transcription workflow should require fewer than 20 Claude interactions for a 60-file batch.
116
+
117
+ **One whisper instance at all times.** Never spawn a second process while one is running.
118
+
119
+ **Local-first, private by default.** No audio leaves the machine. No cloud APIs required.
120
+
121
+ **Works for free-tier users.** Courtroom transcription, documentary research, foreign film subtitling — the tool should serve people who can't afford cloud transcription services.
122
+
123
+ ---
124
+
125
+ ## Contributing
126
+
127
+ Pull requests welcome for any of the above priorities. Check existing issues before starting work.
128
+
129
+ If you've tested GPU acceleration on hardware not listed above, please open an issue with your results — GPU model, VRAM, model size, and observed throughput.