vocalize-cli 0.9.1__tar.gz → 0.10.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/CHANGELOG.md +157 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/PKG-INFO +239 -7
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/README.md +238 -6
- vocalize_cli-0.10.1/docs/dictation.md +531 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/docs/installation.md +79 -8
- vocalize_cli-0.10.1/docs/next-features-analysis.md +133 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/choreography.md +68 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/decisions.md +445 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/design.md +210 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/plan.md +199 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/review-0.10.0.md +58 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-1-status/project-plan.md +39 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-1-status/report.md +125 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-1-status/validate-exit.sh +110 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-10-release-0-11-0/project-plan.md +46 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-10-release-0-11-0/validate-exit.sh +113 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-2-stt-runtime/project-plan.md +50 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-2-stt-runtime/report.md +191 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-2-stt-runtime/validate-exit.sh +114 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-3-recorder/project-plan.md +48 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-3-recorder/report.md +80 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-3-recorder/validate-exit.sh +115 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-4-dictation/project-plan.md +55 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-4-dictation/report.md +106 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-4-dictation/validate-exit.sh +119 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-5-resume/project-plan.md +46 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-5-resume/report.md +94 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-5-resume/validate-exit.sh +113 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-6-release-0-10-0/project-plan.md +50 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-6-release-0-10-0/report.md +72 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-6-release-0-10-0/validate-exit.sh +115 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-7-portal-read/project-plan.md +39 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-7-portal-read/validate-exit.sh +111 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-8-portal-write/project-plan.md +42 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-8-portal-write/validate-exit.sh +113 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-9-portal-page/project-plan.md +44 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-9-portal-page/validate-exit.sh +112 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/spike-2026-09-01.md +71 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/split-assessment.md +47 -0
- vocalize_cli-0.10.1/docs/plans/2026-09-next-features/verification.md +103 -0
- vocalize_cli-0.10.1/docs/research/2026-09-01-config-portal-design.md +58 -0
- vocalize_cli-0.10.1/docs/research/2026-09-01-dictation-design.md +158 -0
- vocalize_cli-0.10.1/docs/research/2026-09-01-voicebox-findings.md +59 -0
- vocalize_cli-0.10.1/docs/roadmap.md +24 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/claude_stop_hook.py +38 -3
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/install_quick_action.py +1 -0
- vocalize_cli-0.10.1/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Info.plist +26 -0
- vocalize_cli-0.10.1/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Resources/document.wflow +206 -0
- vocalize_cli-0.10.1/tests/conftest.py +151 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_audio.py +297 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_chain.py +63 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_claude_stop_hook.py +61 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_cli.py +341 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_config.py +153 -0
- vocalize_cli-0.10.1/tests/test_dictate.py +2335 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_install_quick_action.py +81 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_kokoro_manifest.py +19 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_kokoro_provider.py +8 -2
- vocalize_cli-0.10.1/tests/test_listen_check.py +591 -0
- vocalize_cli-0.10.1/tests/test_local_install.py +1139 -0
- vocalize_cli-0.10.1/tests/test_readiness.py +608 -0
- vocalize_cli-0.10.1/tests/test_recorder_build.py +546 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_speak_options.py +13 -0
- vocalize_cli-0.10.1/tests/test_uv_path.py +62 -0
- vocalize_cli-0.10.1/tests/test_whisper_manifest.py +165 -0
- vocalize_cli-0.10.1/tests/test_whisper_worker.py +292 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/__init__.py +1 -1
- vocalize_cli-0.10.1/vocalize/audio.py +563 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/chain.py +43 -1
- vocalize_cli-0.10.1/vocalize/cli.py +1652 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/config.py +142 -1
- vocalize_cli-0.10.1/vocalize/dictate.py +1335 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/exceptions.py +21 -1
- vocalize_cli-0.10.1/vocalize/interrupted.py +323 -0
- vocalize_cli-0.10.1/vocalize/local/__init__.py +33 -0
- vocalize_cli-0.10.1/vocalize/local/install.py +598 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/local/kokoro_manifest.py +19 -0
- vocalize_cli-0.10.1/vocalize/local/whisper_manifest.py +153 -0
- vocalize_cli-0.10.1/vocalize/local/whisper_worker.py +166 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/kokoro.py +6 -38
- vocalize_cli-0.10.1/vocalize/readiness.py +380 -0
- vocalize_cli-0.10.1/vocalize/recorder/Info.plist.in +40 -0
- vocalize_cli-0.10.1/vocalize/recorder/Recorder.entitlements +8 -0
- vocalize_cli-0.10.1/vocalize/recorder/VocalizeRecorder.swift +527 -0
- vocalize_cli-0.9.1/tests/conftest.py +0 -78
- vocalize_cli-0.9.1/tests/test_local_install.py +0 -548
- vocalize_cli-0.9.1/vocalize/audio.py +0 -300
- vocalize_cli-0.9.1/vocalize/cli.py +0 -859
- vocalize_cli-0.9.1/vocalize/local/__init__.py +0 -7
- vocalize_cli-0.9.1/vocalize/local/install.py +0 -217
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/.env.example +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/.github/workflows/ci.yml +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/.gitignore +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/LICENSE +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/docs/provider-credentials.md +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/install_hook.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Info.plist +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Resources/document.wflow +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Info.plist +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Resources/document.wflow +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Info.plist +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Resources/document.wflow +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/speak_options.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/speak_url_gate.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/pyproject.toml +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_auth.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_cache.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_clipboard.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_elevenlabs_provider.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_exceptions.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_google_provider.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_http.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_install_hook.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_kokoro_worker.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_ledger.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_openai_provider.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_polly_provider.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_preprocess.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_providers_registry.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_say_provider.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_speak_url_gate.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_tts.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_wizard.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/__main__.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/auth.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/cache.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/clipboard.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/ledger.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/local/kokoro_worker.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/preprocess.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/__init__.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/_http.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/elevenlabs.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/google.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/openai.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/polly.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/say.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/tts.py +0 -0
- {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/wizard.py +0 -0
|
@@ -3,6 +3,163 @@
|
|
|
3
3
|
All notable changes to this project are documented here. Format follows
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
5
5
|
|
|
6
|
+
## 0.10.1 - 2026-09-02
|
|
7
|
+
|
|
8
|
+
Three fixes found in the first owner-present run of 0.10.0's dictation.
|
|
9
|
+
Together they meant no hotkey dictation could succeed on 0.10.0; upgrade.
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- **No dictation could ever start on a fresh install.** The recorder was
|
|
14
|
+
signed with the hardened runtime but without the
|
|
15
|
+
`com.apple.security.device.audio-input` entitlement, so macOS refused the
|
|
16
|
+
microphone on the spot — no permission dialog, status stuck at
|
|
17
|
+
`notDetermined` — and every first press ended in "The recorder did not
|
|
18
|
+
start". The bundle is now signed with
|
|
19
|
+
`vocalize/recorder/Recorder.entitlements`, and the entitlements are part
|
|
20
|
+
of the recorder's fingerprint, so `vocalize local install --stt` rebuilds
|
|
21
|
+
the bundle once (and, as with any rebuild, macOS asks for the microphone
|
|
22
|
+
again — it never actually asked before).
|
|
23
|
+
- **Every hotkey dictation ended in "Dictation failed" on a machine whose
|
|
24
|
+
`uv` came from Homebrew.** A Services environment has a bare PATH, and
|
|
25
|
+
`uv_path()` looked only there and in `~/.local/bin`; the same dictation
|
|
26
|
+
worked from a terminal. `/opt/homebrew/bin/uv` and `/usr/local/bin/uv`
|
|
27
|
+
are now tried too (this also covers Kokoro from a Quick Action).
|
|
28
|
+
- **Holding the dictation hotkey down turned into a cancel-and-restart
|
|
29
|
+
loop.** macOS re-fires a Service shortcut at the key-repeat rate, and
|
|
30
|
+
every repeat landed as a second press. Presses within half a second of
|
|
31
|
+
the previous one are now ignored as the same press; a deliberate cancel
|
|
32
|
+
is "press, a beat, press" inside the two-second window, as before.
|
|
33
|
+
- `hooks/claude_stop_hook.py --latest`, run from inside a Claude Code turn
|
|
34
|
+
(which is how `/speak` runs it), spoke the agent's own status line —
|
|
35
|
+
"Checking settings." — instead of the response the user asked to hear.
|
|
36
|
+
It now skips the turn in progress, back past the `/speak` message
|
|
37
|
+
itself, and speaks the response before it. From a plain terminal, where
|
|
38
|
+
no turn is in progress, `--latest` still speaks the newest response; the
|
|
39
|
+
hook tells the two apart by the `CLAUDECODE` variable Claude Code sets
|
|
40
|
+
in its shell. The Stop-hook path is unchanged.
|
|
41
|
+
|
|
42
|
+
## 0.10.0 - 2026-09-02
|
|
43
|
+
|
|
44
|
+
### Added
|
|
45
|
+
|
|
46
|
+
- **Local dictation.** Press a hotkey, speak, press it again, and the
|
|
47
|
+
transcript is on your clipboard — speech to text, entirely on-device via
|
|
48
|
+
[whisper.cpp](https://github.com/ggerganov/whisper.cpp)
|
|
49
|
+
(`pywhispercpp`). Nothing about a dictation leaves the machine unless
|
|
50
|
+
`[stt] cleanup` is turned on, and even then only the transcript is sent
|
|
51
|
+
(to `claude -p`, tools denied), never the audio. See
|
|
52
|
+
[docs/dictation.md](docs/dictation.md) for the full guide.
|
|
53
|
+
- `vocalize listen` (`--toggle`, `--cancel`, `--check`, `--list-devices`,
|
|
54
|
+
`--wav FILE`, `--cleanup`, `--max-seconds`) and `vocalize dictate` (an
|
|
55
|
+
alias for `listen --toggle`, under the name the hotkey uses).
|
|
56
|
+
- New Quick Action, **"Dictate with Vocalize"** — a no-input Service for
|
|
57
|
+
the dictation hotkey (⌃⌥⌘D suggested), installed by the existing
|
|
58
|
+
`hooks/install_quick_action.py` alongside the other three.
|
|
59
|
+
- `vocalize local install --stt [--model base.en|small.en|large-v3-turbo-q5_0]`
|
|
60
|
+
and `vocalize local uninstall --stt` — opt-in download-and-verify of a
|
|
61
|
+
whisper.cpp model, plus build-and-sign of **Vocalize Recorder**, the
|
|
62
|
+
small `.app` bundle that holds the microphone permission (macOS only
|
|
63
|
+
grants that to something with an identity). Nothing is downloaded or
|
|
64
|
+
compiled until you run `install --stt`, mirroring Kokoro's opt-in
|
|
65
|
+
install; the one-time Metal shader warm-up (~8s) is paid here, never
|
|
66
|
+
during a dictation.
|
|
67
|
+
- New `[stt]` config table — `model`, `language`, `input_device`,
|
|
68
|
+
`cleanup`, `paste` (reserved, not implemented yet), `max_seconds`,
|
|
69
|
+
`sounds` — validated on the way in the same way `[providers.*]` is, and
|
|
70
|
+
printed by `vocalize settings` as `stt.*` lines.
|
|
71
|
+
- `vocalize status` — a one-screen readiness check across every provider
|
|
72
|
+
in your chain, plus four dictation rows (`stt model`, `recorder`,
|
|
73
|
+
`microphone`, `input device`) once dictation has been set up at all.
|
|
74
|
+
`--json` prints the same rows as a list; exit 0 when everything is `ok`,
|
|
75
|
+
1 otherwise.
|
|
76
|
+
- `vocalize resume [--forget]` — continue (or discard) a text-to-speech
|
|
77
|
+
read that a dictation interrupted. Starting a dictation stops any read
|
|
78
|
+
in progress, but vocalize now remembers exactly where it stopped and
|
|
79
|
+
offers to continue once the transcript has landed (a macOS dialog,
|
|
80
|
+
default Continue, 15s to answer); the record lives at
|
|
81
|
+
`~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour.
|
|
82
|
+
- New [docs/dictation.md](docs/dictation.md): install, the hotkey, every
|
|
83
|
+
`[stt]` key, `vocalize status`'s dictation rows, `resume`, and
|
|
84
|
+
troubleshooting keyed on `vocalize listen --check`'s exact messages and
|
|
85
|
+
exit codes.
|
|
86
|
+
|
|
87
|
+
### Changed
|
|
88
|
+
|
|
89
|
+
- `vocalize/local/install.py` generalized to support more than one local
|
|
90
|
+
runtime's manifest and model files (previously hard-coded to Kokoro's).
|
|
91
|
+
Kokoro's own install, stamp, and `local status` output are unchanged
|
|
92
|
+
byte-for-byte; the whisper runtime downloads and stamps only the single
|
|
93
|
+
model you selected, not all three.
|
|
94
|
+
- `vocalize listen --check` measures the microphone permission by
|
|
95
|
+
launching Vocalize Recorder the same way a real dictation does
|
|
96
|
+
(through LaunchServices), not by exec'ing its binary directly — macOS
|
|
97
|
+
attributes a TCC grant to the *responsible* process, and exec'ing the
|
|
98
|
+
binary as a child of your shell reported the terminal's own grant
|
|
99
|
+
instead. Exit codes: 0 authorized-and-ready, 2 denied, 3 no usable
|
|
100
|
+
input device, 5 not asked yet (macOS `notDetermined`) — matching the
|
|
101
|
+
recorder's own contract — plus a new exit 1 meaning "vocalize's own
|
|
102
|
+
local install isn't finished" (not built, no model on disk, or the
|
|
103
|
+
recorder never reported back), which is a setup problem, not a
|
|
104
|
+
permission one.
|
|
105
|
+
- `audio.stop_playback()` gained a `remember=` flag. A dictation's stop
|
|
106
|
+
passes it, leaving a marker so the process that was playing can record
|
|
107
|
+
where it stopped — this is what makes `vocalize resume` possible. A
|
|
108
|
+
plain `vocalize stop` records nothing, as before.
|
|
109
|
+
- **A stop now silences every read already in flight**, not only the
|
|
110
|
+
player it kills. Playback is serialized machine-wide, so stopping one
|
|
111
|
+
read used to let the next queued one start speaking immediately — into
|
|
112
|
+
the microphone a dictation had just opened. A read *started* after the
|
|
113
|
+
stop is unaffected.
|
|
114
|
+
- Dictated text reaches the clipboard as a single line. Newlines are
|
|
115
|
+
collapsed there so a paste into a terminal cannot run as several
|
|
116
|
+
commands; `vocalize listen`'s stdout keeps them.
|
|
117
|
+
- `vocalize listen --check` now measures the input device configured in
|
|
118
|
+
`[stt] input_device` rather than the system default, and records what it
|
|
119
|
+
saw with a timestamp — so `vocalize status` says how old that
|
|
120
|
+
"authorized" verdict is instead of implying it is current.
|
|
121
|
+
|
|
122
|
+
### Fixed
|
|
123
|
+
|
|
124
|
+
- `vocalize local install --stt`, re-run against a model that already
|
|
125
|
+
verified, now re-warms the runtime instead of reporting "already
|
|
126
|
+
installed" and stopping — a machine where only the runtime failed to
|
|
127
|
+
start (no Metal, a build hiccup) previously had no way to retry that
|
|
128
|
+
short of a full uninstall and 465 MB re-download.
|
|
129
|
+
- `vocalize local status` reports every installed speech-to-text model,
|
|
130
|
+
not just the default — installing a non-default model with `--model`
|
|
131
|
+
no longer looks unfinished.
|
|
132
|
+
- `vocalize local uninstall --stt` no longer crashes on a symlinked model
|
|
133
|
+
directory or recorder bundle; it reports the symlink and leaves it for
|
|
134
|
+
you to remove.
|
|
135
|
+
- `vocalize status` no longer raises on an unrecognized `VOCALIZE_CHAIN`;
|
|
136
|
+
like any other misconfiguration, it degrades to one failing row instead
|
|
137
|
+
of crashing the command. A probe that raises is reported by exception
|
|
138
|
+
type only — never its message, which could otherwise echo
|
|
139
|
+
credential-shaped text onto the screen.
|
|
140
|
+
- A read stopped by a dictation while a streaming provider's next chunk
|
|
141
|
+
was still rendering (nothing audible playing at that exact instant)
|
|
142
|
+
used to lose the rest of the read with no way to get it back; it's now
|
|
143
|
+
recorded and resumable like any other interruption. The same now holds
|
|
144
|
+
for a read still being synthesized (no player exists yet) and for a
|
|
145
|
+
plain `vocalize stop` landing in that gap.
|
|
146
|
+
- The first dictation on a fresh install no longer fails while macOS is
|
|
147
|
+
asking for the microphone. The permission dialog can sit on screen for
|
|
148
|
+
minutes; the press now waits for your answer and starts recording when
|
|
149
|
+
you click Allow, instead of giving up after five seconds and reporting
|
|
150
|
+
a failure that had not happened.
|
|
151
|
+
- `vocalize resume` continues the read in the voice, model, speed and
|
|
152
|
+
chunk size it was stopped in. It previously fell back to the config
|
|
153
|
+
defaults, which also missed the audio cache and re-synthesized (and
|
|
154
|
+
re-charged for) the whole remainder.
|
|
155
|
+
- A ten-minute dictation is no longer mistaken for a crashed one. The
|
|
156
|
+
claim a stop puts on a take is now aged from its own progress rather
|
|
157
|
+
than from when recording began, so a long take or a stop queued behind
|
|
158
|
+
a long read cannot be reaped mid-transcription.
|
|
159
|
+
- `~/.cache/vocalize` and `~/.cache/vocalize/bin` are tightened to 0700
|
|
160
|
+
even when they already existed. The files inside were always 0600, but
|
|
161
|
+
the directory listing said whether a dictation was in progress.
|
|
162
|
+
|
|
6
163
|
## 0.9.1 - 2026-09-01
|
|
7
164
|
|
|
8
165
|
### Fixed
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: vocalize-cli
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.10.1
|
|
4
4
|
Summary: A CLI that turns text, markdown, or piped stdin into speech via the ElevenLabs API, with markdown-table-aware preprocessing.
|
|
5
5
|
Project-URL: Homepage, https://github.com/matthager12-collab/vocalize
|
|
6
6
|
Project-URL: Repository, https://github.com/matthager12-collab/vocalize
|
|
@@ -289,6 +289,12 @@ in the current directory, then the OS keychain. `vocalize auth login` sets
|
|
|
289
289
|
up the keychain entry; `vocalize auth status` shows which of those sources
|
|
290
290
|
is currently supplying the key.
|
|
291
291
|
|
|
292
|
+
A `[stt]` table configures dictation the same way `[providers.<name>]`
|
|
293
|
+
configures a TTS provider — see
|
|
294
|
+
[Configuration: the `[stt]` table](#configuration-the-stt-table) under
|
|
295
|
+
[Dictation](#dictation-speech-to-text) below for every key and its
|
|
296
|
+
allowlist.
|
|
297
|
+
|
|
292
298
|
## Providers and fallback
|
|
293
299
|
|
|
294
300
|
vocalize tries providers in order — a **chain** — until one speaks. The
|
|
@@ -417,9 +423,201 @@ instead of waiting for the whole thing to render. Measured on this Mac
|
|
|
417
423
|
rendering. `vocalize stop`, run from any terminal, halts a Kokoro read
|
|
418
424
|
mid-sentence the same as any other provider.
|
|
419
425
|
|
|
426
|
+
## Dictation (speech to text)
|
|
427
|
+
|
|
428
|
+
Everything above turns text into speech. Dictation runs the other way:
|
|
429
|
+
press a hotkey, speak, press it again, and the words land on your
|
|
430
|
+
clipboard. It's [whisper.cpp](https://github.com/ggerganov/whisper.cpp) via
|
|
431
|
+
[`pywhispercpp`](https://github.com/absadiki/pywhispercpp), entirely
|
|
432
|
+
on-device — nothing you say leaves the Mac unless you turn on `--cleanup`
|
|
433
|
+
(below), and even then only the *transcript* goes anywhere, never the audio.
|
|
434
|
+
|
|
435
|
+
### Install
|
|
436
|
+
|
|
437
|
+
```bash
|
|
438
|
+
vocalize local install --stt
|
|
439
|
+
```
|
|
440
|
+
|
|
441
|
+
A separate opt-in from Kokoro's `vocalize local install` — nothing here is
|
|
442
|
+
downloaded or built until you run this. It:
|
|
443
|
+
|
|
444
|
+
1. Downloads one whisper.cpp model (`small.en` by default, ~465 MB) from a
|
|
445
|
+
pinned Hugging Face revision, verified against a pinned sha256 before
|
|
446
|
+
it's kept.
|
|
447
|
+
2. Compiles and ad-hoc signs a small Swift recorder bundle, **Vocalize
|
|
448
|
+
Recorder** — the thing that actually owns the microphone permission;
|
|
449
|
+
macOS won't grant that to a bare command-line tool.
|
|
450
|
+
3. Warms the runtime, paying a one-time ~8-second Metal shader compile
|
|
451
|
+
right here, so no dictation ever stalls on it later.
|
|
452
|
+
|
|
453
|
+
The first real dictation prompts for microphone access naming **"Vocalize
|
|
454
|
+
Recorder"** — approve it once, like any other app's first-run permission
|
|
455
|
+
prompt. Pick a different model with `--model large-v3-turbo-q5_0` (~547 MB,
|
|
456
|
+
more accurate) or `--model base.en` (~141 MB, fastest, least accurate).
|
|
457
|
+
|
|
458
|
+
### The hotkey
|
|
459
|
+
|
|
460
|
+
```bash
|
|
461
|
+
python3 hooks/install_quick_action.py
|
|
462
|
+
```
|
|
463
|
+
|
|
464
|
+
then assign a shortcut under **System Settings › Keyboard › Keyboard
|
|
465
|
+
Shortcuts › Services › Text › "Dictate with Vocalize"** — ⌃⌥⌘D is free by
|
|
466
|
+
default and a sensible pick. `vocalize dictate` is the same command from a
|
|
467
|
+
terminal, if you'd rather trigger it that way.
|
|
468
|
+
|
|
469
|
+
### How a dictation works
|
|
470
|
+
|
|
471
|
+
- **Press** the hotkey — a Tink plays, the recorder starts, and any read
|
|
472
|
+
currently playing is stopped first (dictation and playback never
|
|
473
|
+
overlap; vocalize remembers where the read was cut off — see
|
|
474
|
+
[Continuing an interrupted read](#continuing-an-interrupted-read)).
|
|
475
|
+
- **Speak.**
|
|
476
|
+
- **Press again** — a Pop plays, recording stops, and the audio is
|
|
477
|
+
transcribed on-device. If anything was heard, a Glass plays and the
|
|
478
|
+
transcript is on your clipboard; nothing is typed for you automatically.
|
|
479
|
+
- **Nothing heard** (silence, or a microphone that isn't actually picking
|
|
480
|
+
anything up) ends the dictation quietly — no clipboard write.
|
|
481
|
+
- **Cancel** with a second press *within two seconds* of the first (but
|
|
482
|
+
not within half a second — that's a held key, and it's ignored), or at
|
|
483
|
+
any point with `vocalize listen --cancel` — the audio is discarded, never
|
|
484
|
+
transcribed.
|
|
485
|
+
- **A third press while transcribing is refused**: a Pop, and "Still
|
|
486
|
+
transcribing the last dictation." Wait for the clipboard notification, or
|
|
487
|
+
`--cancel`, before dictating again.
|
|
488
|
+
|
|
489
|
+
### `vocalize listen`
|
|
490
|
+
|
|
491
|
+
`vocalize dictate` (the hotkey's command) is `vocalize listen --toggle`
|
|
492
|
+
under another name. `listen` is the general primitive:
|
|
493
|
+
|
|
494
|
+
```bash
|
|
495
|
+
vocalize listen # record until Enter/Ctrl-C, print to stdout
|
|
496
|
+
vocalize listen --toggle # start, or stop and copy to the clipboard
|
|
497
|
+
vocalize listen --cancel # discard whatever is in progress
|
|
498
|
+
vocalize listen --wav clip.wav # transcribe a file you already have
|
|
499
|
+
vocalize listen --check # microphone + install readiness
|
|
500
|
+
vocalize listen --list-devices # input device names for [stt] input_device
|
|
501
|
+
vocalize listen --max-seconds 30 # cap this one recording
|
|
502
|
+
```
|
|
503
|
+
|
|
504
|
+
`--wav` is trusted input — the file has to be 16 kHz mono 16-bit WAV
|
|
505
|
+
(exactly what the recorder, and `say --data-format=LEI16@16000`, produce);
|
|
506
|
+
a malformed file gets a plain error naming the format, not a crash.
|
|
507
|
+
`--cleanup` tidies the transcript with Claude before it's delivered; it
|
|
508
|
+
applies to a live recording (`--toggle`/`dictate`, or plain `listen`) and
|
|
509
|
+
has no effect on `--wav`, which transcribes literally. `--max-seconds`
|
|
510
|
+
overrides `[stt] max_seconds` for one invocation.
|
|
511
|
+
|
|
512
|
+
### Configuration: the `[stt]` table
|
|
513
|
+
|
|
514
|
+
```toml
|
|
515
|
+
[stt]
|
|
516
|
+
model = "small.en" # base.en | small.en | large-v3-turbo-q5_0
|
|
517
|
+
language = "en" # a whisper.cpp language code
|
|
518
|
+
input_device = "" # "" = system default; else an exact name from --list-devices
|
|
519
|
+
cleanup = false # send the transcript (never audio) to Claude first
|
|
520
|
+
max_seconds = 120 # 1-600; the recorder self-stops here, dictate backstops it
|
|
521
|
+
sounds = true # the Tink/Pop/Glass feedback sounds
|
|
522
|
+
```
|
|
523
|
+
|
|
524
|
+
| Key | Allowed values | Default |
|
|
525
|
+
|---|---|---|
|
|
526
|
+
| `model` | `base.en`, `small.en`, `large-v3-turbo-q5_0` | `small.en` |
|
|
527
|
+
| `language` | a whisper.cpp language code (`en`, `es`, `fr`, …); an `.en` model must stay `en` | `en` |
|
|
528
|
+
| `input_device` | `""` (system default) or an exact name from `vocalize listen --list-devices`; ≤ 128 characters, printable, can't start with `-` | `""` |
|
|
529
|
+
| `cleanup` | `true` / `false` | `false` |
|
|
530
|
+
| `paste` | reserved — not implemented in 0.10.0 | `false` |
|
|
531
|
+
| `max_seconds` | integer, 1–600 | `120` |
|
|
532
|
+
| `sounds` | `true` / `false` | `true` |
|
|
533
|
+
|
|
534
|
+
An unknown key warns on stderr; a bad value is a `ConfigError` naming it —
|
|
535
|
+
every one of these becomes a subprocess argument eventually, so nothing
|
|
536
|
+
here is trusted on the way in.
|
|
537
|
+
|
|
538
|
+
**The input-device gotcha:** a pair of Bluetooth earbuds that are *paired*
|
|
539
|
+
but not actually in your ears still shows up as the default input device —
|
|
540
|
+
and delivers digital silence. If dictation keeps saying nothing was heard,
|
|
541
|
+
run `vocalize listen --list-devices`, copy the real microphone's name, and
|
|
542
|
+
set it:
|
|
543
|
+
|
|
544
|
+
```toml
|
|
545
|
+
[stt]
|
|
546
|
+
input_device = "MacBook Pro Microphone"
|
|
547
|
+
```
|
|
548
|
+
|
|
549
|
+
### `vocalize status`
|
|
550
|
+
|
|
551
|
+
```bash
|
|
552
|
+
vocalize status
|
|
553
|
+
```
|
|
554
|
+
|
|
555
|
+
prints one row per provider in your chain, plus four dictation rows once
|
|
556
|
+
dictation has been set up at all — an `[stt]` table in your config, a
|
|
557
|
+
built recorder, or a model on disk (a machine that never opted in doesn't
|
|
558
|
+
get four permanent red rows for a feature nobody asked for):
|
|
559
|
+
|
|
560
|
+
| Row | Reports |
|
|
561
|
+
|---|---|
|
|
562
|
+
| `stt model` | whether a whisper.cpp model is on disk |
|
|
563
|
+
| `recorder` | whether Vocalize Recorder is built |
|
|
564
|
+
| `microphone` | authorized / denied / not asked yet — from the last `listen --check`, never by launching the recorder itself |
|
|
565
|
+
| `input device` | whether the configured (or default) input device is actually present |
|
|
566
|
+
|
|
567
|
+
`--json` prints the same rows as a list. Exit code is 0 when every row —
|
|
568
|
+
providers and dictation both — is `ok`, 1 otherwise, so it composes with
|
|
569
|
+
`&&` in a script.
|
|
570
|
+
|
|
571
|
+
### Continuing an interrupted read
|
|
572
|
+
|
|
573
|
+
Starting a dictation stops any read in progress, but doesn't throw the rest
|
|
574
|
+
of it away (DEC-003). The process that was speaking remembers exactly
|
|
575
|
+
where it stopped, and once your transcript has landed, vocalize asks:
|
|
576
|
+
**"Continue the read you interrupted?"** — answer, or let the dialog give
|
|
577
|
+
up after 15 seconds (counted as no). From a terminal, the same thing is:
|
|
578
|
+
|
|
579
|
+
```bash
|
|
580
|
+
vocalize resume # continue where the last read left off
|
|
581
|
+
vocalize resume --forget # discard it instead
|
|
582
|
+
```
|
|
583
|
+
|
|
584
|
+
The record — one piece of audio and the text not yet spoken — lives at
|
|
585
|
+
`~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour, and is
|
|
586
|
+
deleted the moment you resume, decline, or `--forget` it. This is the
|
|
587
|
+
interrupted *read*, not dictation audio or a transcript — see Privacy,
|
|
588
|
+
below.
|
|
589
|
+
|
|
590
|
+
### Uninstall
|
|
591
|
+
|
|
592
|
+
```bash
|
|
593
|
+
vocalize local uninstall --stt
|
|
594
|
+
```
|
|
595
|
+
|
|
596
|
+
Removes the downloaded model file and the recorder bundle. The microphone
|
|
597
|
+
permission grant itself stays in System Settings — remove it there
|
|
598
|
+
(Privacy & Security › Microphone) if you want that gone too.
|
|
599
|
+
|
|
600
|
+
### Privacy
|
|
601
|
+
|
|
602
|
+
vocalize writes no transcript to disk and shows none in a notification —
|
|
603
|
+
it's held in memory and goes to your clipboard (or stdout, for
|
|
604
|
+
`vocalize listen`) and nowhere else. Recorded audio lives only in a
|
|
605
|
+
private temporary directory deleted the moment a dictation ends, one way
|
|
606
|
+
or another (a sweep clears anything a hard kill leaves behind after 24
|
|
607
|
+
hours). The **only** thing that ever leaves the machine is the transcript,
|
|
608
|
+
and only when you turn on `--cleanup`: it's sent to `claude -p` with every
|
|
609
|
+
tool denied, purely to fix punctuation and casing. The audio itself is
|
|
610
|
+
never sent anywhere.
|
|
611
|
+
|
|
612
|
+
`--cleanup` has one consequence worth knowing before you turn it on:
|
|
613
|
+
Claude Code logs the prompt and stdin of every print-mode run, so the
|
|
614
|
+
transcript is written in plaintext to `~/.claude/projects/…`. vocalize
|
|
615
|
+
cannot suppress that. It is off by default. Full accounting in
|
|
616
|
+
[docs/dictation.md](docs/dictation.md#privacy).
|
|
617
|
+
|
|
420
618
|
## macOS Quick Actions (highlight → speak)
|
|
421
619
|
|
|
422
|
-
|
|
620
|
+
Four Services let you use vocalize from any app without a terminal:
|
|
423
621
|
|
|
424
622
|
- **Speak with Vocalize** — highlight text anywhere, right-click →
|
|
425
623
|
Services → Speak with Vocalize. It stops whatever was already playing
|
|
@@ -436,8 +634,11 @@ Two Services let you use vocalize from any app without a terminal:
|
|
|
436
634
|
(`~/.claude/plans/`) aloud on demand. Made for the plan-approval moment:
|
|
437
635
|
the proposal card is up, you press your shortcut, hear the plan, then
|
|
438
636
|
accept or reject. Nothing reads unless you trigger it.
|
|
637
|
+
- **Dictate with Vocalize** — the dictation hotkey (see
|
|
638
|
+
[Dictation](#dictation-speech-to-text) above). Takes no input and shows
|
|
639
|
+
no window; press it, speak, press it again.
|
|
439
640
|
|
|
440
|
-
Install
|
|
641
|
+
Install all four:
|
|
441
642
|
|
|
442
643
|
```bash
|
|
443
644
|
python3 hooks/install_quick_action.py
|
|
@@ -605,14 +806,26 @@ vocalize/
|
|
|
605
806
|
chain.py # tries each provider in the chain in turn until one speaks
|
|
606
807
|
ledger.py # ~/.cache/vocalize/usage.json — local monthly budget tracking
|
|
607
808
|
providers/ # elevenlabs.py, openai.py, google.py, polly.py, say.py, kokoro.py
|
|
608
|
-
local/ #
|
|
809
|
+
local/ # opt-in download/verify + the uv-run worker scripts —
|
|
810
|
+
# kokoro_manifest.py/kokoro_worker.py (TTS) and
|
|
811
|
+
# whisper_manifest.py/whisper_worker.py (dictation)
|
|
812
|
+
recorder/ # VocalizeRecorder.swift + Info.plist.in — the ad-hoc-signed
|
|
813
|
+
# .app bundle that owns the microphone permission
|
|
609
814
|
audio.py # save to disk + play via the OS's native player
|
|
610
815
|
# (afplay / mpg123 / ffplay / PowerShell, whichever exists)
|
|
816
|
+
dictate.py # the hotkey's toggle state machine: record, transcribe,
|
|
817
|
+
# clipboard — nothing here ever becomes a file or a log line
|
|
818
|
+
interrupted.py # the record of a read a dictation interrupted, and the
|
|
819
|
+
# slice `vocalize resume` plays to continue it
|
|
820
|
+
readiness.py # vocalize status's per-provider + per-dictation-row probes,
|
|
821
|
+
# each on a timed daemon thread so a wedged keychain or
|
|
822
|
+
# microphone check can never hang the command
|
|
611
823
|
cli.py # click-based CLI wiring the above together
|
|
612
824
|
hooks/
|
|
613
825
|
claude_stop_hook.py # Claude Code Stop hook -> calls the vocalize CLI
|
|
614
826
|
install_hook.py # safely merges the hook into ~/.claude/settings.json
|
|
615
|
-
tests/ # pytest, all mocked — no API key
|
|
827
|
+
tests/ # pytest, all mocked — no API key, network, or
|
|
828
|
+
# microphone needed to run these
|
|
616
829
|
```
|
|
617
830
|
|
|
618
831
|
## Testing
|
|
@@ -622,8 +835,11 @@ pip install -e ".[dev]"
|
|
|
622
835
|
pytest
|
|
623
836
|
```
|
|
624
837
|
|
|
625
|
-
|
|
626
|
-
`tts.py`, so tests pass in a fake client instead of hitting the real
|
|
838
|
+
Over 1,180 tests, all offline: the ElevenLabs client is dependency-injected
|
|
839
|
+
into `tts.py`, so tests pass in a fake client instead of hitting the real
|
|
840
|
+
API, and dictation's tests fake the recorder (a tiny shell script honoring
|
|
841
|
+
a stop file) and the whisper worker instead of touching a microphone or a
|
|
842
|
+
model.
|
|
627
843
|
|
|
628
844
|
## Known limitations
|
|
629
845
|
|
|
@@ -676,6 +892,22 @@ All tests run offline: the ElevenLabs client is dependency-injected into
|
|
|
676
892
|
confidential information"); until you click Always Allow, every command
|
|
677
893
|
that needs that key waits. Click it once per Python binary, or supply the
|
|
678
894
|
key through its environment variable instead.
|
|
895
|
+
- **Dictation is macOS only.** It depends on `AVFoundation`, LaunchServices,
|
|
896
|
+
and a Swift-compiled `.app` bundle for the microphone permission — there's
|
|
897
|
+
no equivalent path on Linux or Windows.
|
|
898
|
+
- **Rebuilding the recorder means re-granting the microphone.** Vocalize
|
|
899
|
+
Recorder's ad-hoc code signature is what macOS ties the permission grant
|
|
900
|
+
to; when its Swift source changes (a vocalize upgrade that touches it),
|
|
901
|
+
`vocalize local install --stt` rebuilds the bundle and warns you to
|
|
902
|
+
re-approve it in System Settings › Privacy & Security › Microphone. An
|
|
903
|
+
install that doesn't change the source never re-signs, so this isn't
|
|
904
|
+
every upgrade — only ones that touch the recorder.
|
|
905
|
+
- **`small.en` mishears jargon.** The default model does fine on ordinary
|
|
906
|
+
speech but can mangle project-specific words (`pyproject`, a function
|
|
907
|
+
name) — pick `large-v3-turbo-q5_0` for better accuracy, or turn on
|
|
908
|
+
`[stt] cleanup` so Claude fixes obvious transcription noise before it
|
|
909
|
+
reaches your clipboard (it still can't guess a word it never heard
|
|
910
|
+
correctly).
|
|
679
911
|
|
|
680
912
|
## License
|
|
681
913
|
|