vocalize-cli 0.9.0__tar.gz → 0.10.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/CHANGELOG.md +141 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/PKG-INFO +238 -7
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/README.md +237 -6
- vocalize_cli-0.10.0/docs/dictation.md +528 -0
- vocalize_cli-0.10.0/docs/installation.md +470 -0
- vocalize_cli-0.10.0/docs/next-features-analysis.md +133 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/choreography.md +68 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/decisions.md +445 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/design.md +210 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/plan.md +199 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/review-0.10.0.md +58 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-1-status/project-plan.md +39 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-1-status/report.md +125 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-1-status/validate-exit.sh +110 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-10-release-0-11-0/project-plan.md +46 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-10-release-0-11-0/validate-exit.sh +113 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-2-stt-runtime/project-plan.md +50 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-2-stt-runtime/report.md +191 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-2-stt-runtime/validate-exit.sh +114 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-3-recorder/project-plan.md +48 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-3-recorder/report.md +80 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-3-recorder/validate-exit.sh +115 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-4-dictation/project-plan.md +55 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-4-dictation/report.md +106 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-4-dictation/validate-exit.sh +119 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-5-resume/project-plan.md +46 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-5-resume/report.md +94 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-5-resume/validate-exit.sh +113 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-6-release-0-10-0/project-plan.md +50 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-6-release-0-10-0/report.md +71 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-6-release-0-10-0/validate-exit.sh +115 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-7-portal-read/project-plan.md +39 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-7-portal-read/validate-exit.sh +111 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-8-portal-write/project-plan.md +42 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-8-portal-write/validate-exit.sh +113 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-9-portal-page/project-plan.md +44 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-9-portal-page/validate-exit.sh +112 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/spike-2026-09-01.md +71 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/split-assessment.md +47 -0
- vocalize_cli-0.10.0/docs/plans/2026-09-next-features/verification.md +103 -0
- vocalize_cli-0.10.0/docs/research/2026-09-01-config-portal-design.md +58 -0
- vocalize_cli-0.10.0/docs/research/2026-09-01-dictation-design.md +158 -0
- vocalize_cli-0.10.0/docs/research/2026-09-01-voicebox-findings.md +59 -0
- vocalize_cli-0.10.0/docs/roadmap.md +22 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/install_quick_action.py +1 -0
- vocalize_cli-0.10.0/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Info.plist +26 -0
- vocalize_cli-0.10.0/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Resources/document.wflow +206 -0
- vocalize_cli-0.10.0/tests/conftest.py +151 -0
- vocalize_cli-0.10.0/tests/test_audio.py +731 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_chain.py +63 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_cli.py +341 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_config.py +153 -0
- vocalize_cli-0.10.0/tests/test_dictate.py +2264 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_install_quick_action.py +81 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_kokoro_manifest.py +19 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_kokoro_provider.py +6 -2
- vocalize_cli-0.10.0/tests/test_listen_check.py +591 -0
- vocalize_cli-0.10.0/tests/test_local_install.py +1139 -0
- vocalize_cli-0.10.0/tests/test_readiness.py +608 -0
- vocalize_cli-0.10.0/tests/test_recorder_build.py +510 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_speak_options.py +13 -0
- vocalize_cli-0.10.0/tests/test_whisper_manifest.py +165 -0
- vocalize_cli-0.10.0/tests/test_whisper_worker.py +292 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/__init__.py +1 -1
- vocalize_cli-0.10.0/vocalize/audio.py +563 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/chain.py +43 -1
- vocalize_cli-0.10.0/vocalize/cli.py +1652 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/config.py +142 -1
- vocalize_cli-0.10.0/vocalize/dictate.py +1292 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/exceptions.py +21 -1
- vocalize_cli-0.10.0/vocalize/interrupted.py +323 -0
- vocalize_cli-0.10.0/vocalize/local/__init__.py +25 -0
- vocalize_cli-0.10.0/vocalize/local/install.py +588 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/local/kokoro_manifest.py +19 -0
- vocalize_cli-0.10.0/vocalize/local/whisper_manifest.py +153 -0
- vocalize_cli-0.10.0/vocalize/local/whisper_worker.py +166 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/kokoro.py +6 -38
- vocalize_cli-0.10.0/vocalize/readiness.py +380 -0
- vocalize_cli-0.10.0/vocalize/recorder/Info.plist.in +40 -0
- vocalize_cli-0.10.0/vocalize/recorder/VocalizeRecorder.swift +527 -0
- vocalize_cli-0.9.0/tests/conftest.py +0 -67
- vocalize_cli-0.9.0/tests/test_audio.py +0 -375
- vocalize_cli-0.9.0/tests/test_local_install.py +0 -548
- vocalize_cli-0.9.0/vocalize/audio.py +0 -246
- vocalize_cli-0.9.0/vocalize/cli.py +0 -859
- vocalize_cli-0.9.0/vocalize/local/__init__.py +0 -7
- vocalize_cli-0.9.0/vocalize/local/install.py +0 -217
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/.env.example +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/.github/workflows/ci.yml +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/.gitignore +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/LICENSE +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/docs/provider-credentials.md +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/claude_stop_hook.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/install_hook.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Info.plist +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Resources/document.wflow +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Info.plist +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Resources/document.wflow +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Info.plist +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Resources/document.wflow +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/speak_options.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/speak_url_gate.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/pyproject.toml +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_auth.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_cache.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_claude_stop_hook.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_clipboard.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_elevenlabs_provider.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_exceptions.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_google_provider.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_http.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_install_hook.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_kokoro_worker.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_ledger.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_openai_provider.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_polly_provider.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_preprocess.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_providers_registry.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_say_provider.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_speak_url_gate.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_tts.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_wizard.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/__main__.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/auth.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/cache.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/clipboard.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/ledger.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/local/kokoro_worker.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/preprocess.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/__init__.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/_http.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/elevenlabs.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/google.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/openai.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/polly.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/say.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/tts.py +0 -0
- {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/wizard.py +0 -0
|
@@ -3,6 +3,147 @@
|
|
|
3
3
|
All notable changes to this project are documented here. Format follows
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
5
5
|
|
|
6
|
+
## 0.10.0 - 2026-09-02
|
|
7
|
+
|
|
8
|
+
### Added
|
|
9
|
+
|
|
10
|
+
- **Local dictation.** Press a hotkey, speak, press it again, and the
|
|
11
|
+
transcript is on your clipboard — speech to text, entirely on-device via
|
|
12
|
+
[whisper.cpp](https://github.com/ggerganov/whisper.cpp)
|
|
13
|
+
(`pywhispercpp`). Nothing about a dictation leaves the machine unless
|
|
14
|
+
`[stt] cleanup` is turned on, and even then only the transcript is sent
|
|
15
|
+
(to `claude -p`, tools denied), never the audio. See
|
|
16
|
+
[docs/dictation.md](docs/dictation.md) for the full guide.
|
|
17
|
+
- `vocalize listen` (`--toggle`, `--cancel`, `--check`, `--list-devices`,
|
|
18
|
+
`--wav FILE`, `--cleanup`, `--max-seconds`) and `vocalize dictate` (an
|
|
19
|
+
alias for `listen --toggle`, under the name the hotkey uses).
|
|
20
|
+
- New Quick Action, **"Dictate with Vocalize"** — a no-input Service for
|
|
21
|
+
the dictation hotkey (⌃⌥⌘D suggested), installed by the existing
|
|
22
|
+
`hooks/install_quick_action.py` alongside the other three.
|
|
23
|
+
- `vocalize local install --stt [--model base.en|small.en|large-v3-turbo-q5_0]`
|
|
24
|
+
and `vocalize local uninstall --stt` — opt-in download-and-verify of a
|
|
25
|
+
whisper.cpp model, plus build-and-sign of **Vocalize Recorder**, the
|
|
26
|
+
small `.app` bundle that holds the microphone permission (macOS only
|
|
27
|
+
grants that to something with an identity). Nothing is downloaded or
|
|
28
|
+
compiled until you run `install --stt`, mirroring Kokoro's opt-in
|
|
29
|
+
install; the one-time Metal shader warm-up (~8s) is paid here, never
|
|
30
|
+
during a dictation.
|
|
31
|
+
- New `[stt]` config table — `model`, `language`, `input_device`,
|
|
32
|
+
`cleanup`, `paste` (reserved, not implemented yet), `max_seconds`,
|
|
33
|
+
`sounds` — validated on the way in the same way `[providers.*]` is, and
|
|
34
|
+
printed by `vocalize settings` as `stt.*` lines.
|
|
35
|
+
- `vocalize status` — a one-screen readiness check across every provider
|
|
36
|
+
in your chain, plus four dictation rows (`stt model`, `recorder`,
|
|
37
|
+
`microphone`, `input device`) once dictation has been set up at all.
|
|
38
|
+
`--json` prints the same rows as a list; exit 0 when everything is `ok`,
|
|
39
|
+
1 otherwise.
|
|
40
|
+
- `vocalize resume [--forget]` — continue (or discard) a text-to-speech
|
|
41
|
+
read that a dictation interrupted. Starting a dictation stops any read
|
|
42
|
+
in progress, but vocalize now remembers exactly where it stopped and
|
|
43
|
+
offers to continue once the transcript has landed (a macOS dialog,
|
|
44
|
+
default Continue, 15s to answer); the record lives at
|
|
45
|
+
`~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour.
|
|
46
|
+
- New [docs/dictation.md](docs/dictation.md): install, the hotkey, every
|
|
47
|
+
`[stt]` key, `vocalize status`'s dictation rows, `resume`, and
|
|
48
|
+
troubleshooting keyed on `vocalize listen --check`'s exact messages and
|
|
49
|
+
exit codes.
|
|
50
|
+
|
|
51
|
+
### Changed
|
|
52
|
+
|
|
53
|
+
- `vocalize/local/install.py` generalized to support more than one local
|
|
54
|
+
runtime's manifest and model files (previously hard-coded to Kokoro's).
|
|
55
|
+
Kokoro's own install, stamp, and `local status` output are unchanged
|
|
56
|
+
byte-for-byte; the whisper runtime downloads and stamps only the single
|
|
57
|
+
model you selected, not all three.
|
|
58
|
+
- `vocalize listen --check` measures the microphone permission by
|
|
59
|
+
launching Vocalize Recorder the same way a real dictation does
|
|
60
|
+
(through LaunchServices), not by exec'ing its binary directly — macOS
|
|
61
|
+
attributes a TCC grant to the *responsible* process, and exec'ing the
|
|
62
|
+
binary as a child of your shell reported the terminal's own grant
|
|
63
|
+
instead. Exit codes: 0 authorized-and-ready, 2 denied, 3 no usable
|
|
64
|
+
input device, 5 not asked yet (macOS `notDetermined`) — matching the
|
|
65
|
+
recorder's own contract — plus a new exit 1 meaning "vocalize's own
|
|
66
|
+
local install isn't finished" (not built, no model on disk, or the
|
|
67
|
+
recorder never reported back), which is a setup problem, not a
|
|
68
|
+
permission one.
|
|
69
|
+
- `audio.stop_playback()` gained a `remember=` flag. A dictation's stop
|
|
70
|
+
passes it, leaving a marker so the process that was playing can record
|
|
71
|
+
where it stopped — this is what makes `vocalize resume` possible. A
|
|
72
|
+
plain `vocalize stop` records nothing, as before.
|
|
73
|
+
- **A stop now silences every read already in flight**, not only the
|
|
74
|
+
player it kills. Playback is serialized machine-wide, so stopping one
|
|
75
|
+
read used to let the next queued one start speaking immediately — into
|
|
76
|
+
the microphone a dictation had just opened. A read *started* after the
|
|
77
|
+
stop is unaffected.
|
|
78
|
+
- Dictated text reaches the clipboard as a single line. Newlines are
|
|
79
|
+
collapsed there so a paste into a terminal cannot run as several
|
|
80
|
+
commands; `vocalize listen`'s stdout keeps them.
|
|
81
|
+
- `vocalize listen --check` now measures the input device configured in
|
|
82
|
+
`[stt] input_device` rather than the system default, and records what it
|
|
83
|
+
saw with a timestamp — so `vocalize status` says how old that
|
|
84
|
+
"authorized" verdict is instead of implying it is current.
|
|
85
|
+
|
|
86
|
+
### Fixed
|
|
87
|
+
|
|
88
|
+
- `vocalize local install --stt`, re-run against a model that already
|
|
89
|
+
verified, now re-warms the runtime instead of reporting "already
|
|
90
|
+
installed" and stopping — a machine where only the runtime failed to
|
|
91
|
+
start (no Metal, a build hiccup) previously had no way to retry that
|
|
92
|
+
short of a full uninstall and 465 MB re-download.
|
|
93
|
+
- `vocalize local status` reports every installed speech-to-text model,
|
|
94
|
+
not just the default — installing a non-default model with `--model`
|
|
95
|
+
no longer looks unfinished.
|
|
96
|
+
- `vocalize local uninstall --stt` no longer crashes on a symlinked model
|
|
97
|
+
directory or recorder bundle; it reports the symlink and leaves it for
|
|
98
|
+
you to remove.
|
|
99
|
+
- `vocalize status` no longer raises on an unrecognized `VOCALIZE_CHAIN`;
|
|
100
|
+
like any other misconfiguration, it degrades to one failing row instead
|
|
101
|
+
of crashing the command. A probe that raises is reported by exception
|
|
102
|
+
type only — never its message, which could otherwise echo
|
|
103
|
+
credential-shaped text onto the screen.
|
|
104
|
+
- A read stopped by a dictation while a streaming provider's next chunk
|
|
105
|
+
was still rendering (nothing audible playing at that exact instant)
|
|
106
|
+
used to lose the rest of the read with no way to get it back; it's now
|
|
107
|
+
recorded and resumable like any other interruption. The same now holds
|
|
108
|
+
for a read still being synthesized (no player exists yet) and for a
|
|
109
|
+
plain `vocalize stop` landing in that gap.
|
|
110
|
+
- The first dictation on a fresh install no longer fails while macOS is
|
|
111
|
+
asking for the microphone. The permission dialog can sit on screen for
|
|
112
|
+
minutes; the press now waits for your answer and starts recording when
|
|
113
|
+
you click Allow, instead of giving up after five seconds and reporting
|
|
114
|
+
a failure that had not happened.
|
|
115
|
+
- `vocalize resume` continues the read in the voice, model, speed and
|
|
116
|
+
chunk size it was stopped in. It previously fell back to the config
|
|
117
|
+
defaults, which also missed the audio cache and re-synthesized (and
|
|
118
|
+
re-charged for) the whole remainder.
|
|
119
|
+
- A ten-minute dictation is no longer mistaken for a crashed one. The
|
|
120
|
+
claim a stop puts on a take is now aged from its own progress rather
|
|
121
|
+
than from when recording began, so a long take or a stop queued behind
|
|
122
|
+
a long read cannot be reaped mid-transcription.
|
|
123
|
+
- `~/.cache/vocalize` and `~/.cache/vocalize/bin` are tightened to 0700
|
|
124
|
+
even when they already existed. The files inside were always 0600, but
|
|
125
|
+
the directory listing said whether a dictation was in progress.
|
|
126
|
+
|
|
127
|
+
## 0.9.1 - 2026-09-01
|
|
128
|
+
|
|
129
|
+
### Fixed
|
|
130
|
+
|
|
131
|
+
- Concurrent invocations no longer talk over each other. Playback is now
|
|
132
|
+
serialized machine-wide on an exclusive file lock
|
|
133
|
+
(`~/.cache/vocalize/play.lock`): a read that arrives while another is
|
|
134
|
+
playing queues and starts the moment the first one ends. Only the audible
|
|
135
|
+
part is serialized — synthesis still runs concurrently — and the lock
|
|
136
|
+
dies with its process, so a killed or timed-out waiter can never leave a
|
|
137
|
+
stale lock behind. Chunked reads hold the slot for the whole sequence, so
|
|
138
|
+
pieces of two reads never interleave. On platforms without `fcntl`
|
|
139
|
+
(Windows), the lock is skipped and the old overlapping behavior remains.
|
|
140
|
+
|
|
141
|
+
### Changed
|
|
142
|
+
|
|
143
|
+
- `vocalize stop` semantics with a queue: stopping kills the *current*
|
|
144
|
+
player; the next queued read (if any) then begins. Run `stop` again to
|
|
145
|
+
silence that one too.
|
|
146
|
+
|
|
6
147
|
## 0.9.0 - 2026-09-01
|
|
7
148
|
|
|
8
149
|
### Added
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: vocalize-cli
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.10.0
|
|
4
4
|
Summary: A CLI that turns text, markdown, or piped stdin into speech via the ElevenLabs API, with markdown-table-aware preprocessing.
|
|
5
5
|
Project-URL: Homepage, https://github.com/matthager12-collab/vocalize
|
|
6
6
|
Project-URL: Repository, https://github.com/matthager12-collab/vocalize
|
|
@@ -289,6 +289,12 @@ in the current directory, then the OS keychain. `vocalize auth login` sets
|
|
|
289
289
|
up the keychain entry; `vocalize auth status` shows which of those sources
|
|
290
290
|
is currently supplying the key.
|
|
291
291
|
|
|
292
|
+
A `[stt]` table configures dictation the same way `[providers.<name>]`
|
|
293
|
+
configures a TTS provider — see
|
|
294
|
+
[Configuration: the `[stt]` table](#configuration-the-stt-table) under
|
|
295
|
+
[Dictation](#dictation-speech-to-text) below for every key and its
|
|
296
|
+
allowlist.
|
|
297
|
+
|
|
292
298
|
## Providers and fallback
|
|
293
299
|
|
|
294
300
|
vocalize tries providers in order — a **chain** — until one speaks. The
|
|
@@ -417,9 +423,200 @@ instead of waiting for the whole thing to render. Measured on this Mac
|
|
|
417
423
|
rendering. `vocalize stop`, run from any terminal, halts a Kokoro read
|
|
418
424
|
mid-sentence the same as any other provider.
|
|
419
425
|
|
|
426
|
+
## Dictation (speech to text)
|
|
427
|
+
|
|
428
|
+
Everything above turns text into speech. Dictation runs the other way:
|
|
429
|
+
press a hotkey, speak, press it again, and the words land on your
|
|
430
|
+
clipboard. It's [whisper.cpp](https://github.com/ggerganov/whisper.cpp) via
|
|
431
|
+
[`pywhispercpp`](https://github.com/absadiki/pywhispercpp), entirely
|
|
432
|
+
on-device — nothing you say leaves the Mac unless you turn on `--cleanup`
|
|
433
|
+
(below), and even then only the *transcript* goes anywhere, never the audio.
|
|
434
|
+
|
|
435
|
+
### Install
|
|
436
|
+
|
|
437
|
+
```bash
|
|
438
|
+
vocalize local install --stt
|
|
439
|
+
```
|
|
440
|
+
|
|
441
|
+
A separate opt-in from Kokoro's `vocalize local install` — nothing here is
|
|
442
|
+
downloaded or built until you run this. It:
|
|
443
|
+
|
|
444
|
+
1. Downloads one whisper.cpp model (`small.en` by default, ~465 MB) from a
|
|
445
|
+
pinned Hugging Face revision, verified against a pinned sha256 before
|
|
446
|
+
it's kept.
|
|
447
|
+
2. Compiles and ad-hoc signs a small Swift recorder bundle, **Vocalize
|
|
448
|
+
Recorder** — the thing that actually owns the microphone permission;
|
|
449
|
+
macOS won't grant that to a bare command-line tool.
|
|
450
|
+
3. Warms the runtime, paying a one-time ~8-second Metal shader compile
|
|
451
|
+
right here, so no dictation ever stalls on it later.
|
|
452
|
+
|
|
453
|
+
The first real dictation prompts for microphone access naming **"Vocalize
|
|
454
|
+
Recorder"** — approve it once, like any other app's first-run permission
|
|
455
|
+
prompt. Pick a different model with `--model large-v3-turbo-q5_0` (~547 MB,
|
|
456
|
+
more accurate) or `--model base.en` (~141 MB, fastest, least accurate).
|
|
457
|
+
|
|
458
|
+
### The hotkey
|
|
459
|
+
|
|
460
|
+
```bash
|
|
461
|
+
python3 hooks/install_quick_action.py
|
|
462
|
+
```
|
|
463
|
+
|
|
464
|
+
then assign a shortcut under **System Settings › Keyboard › Keyboard
|
|
465
|
+
Shortcuts › Services › Text › "Dictate with Vocalize"** — ⌃⌥⌘D is free by
|
|
466
|
+
default and a sensible pick. `vocalize dictate` is the same command from a
|
|
467
|
+
terminal, if you'd rather trigger it that way.
|
|
468
|
+
|
|
469
|
+
### How a dictation works
|
|
470
|
+
|
|
471
|
+
- **Press** the hotkey — a Tink plays, the recorder starts, and any read
|
|
472
|
+
currently playing is stopped first (dictation and playback never
|
|
473
|
+
overlap; vocalize remembers where the read was cut off — see
|
|
474
|
+
[Continuing an interrupted read](#continuing-an-interrupted-read)).
|
|
475
|
+
- **Speak.**
|
|
476
|
+
- **Press again** — a Pop plays, recording stops, and the audio is
|
|
477
|
+
transcribed on-device. If anything was heard, a Glass plays and the
|
|
478
|
+
transcript is on your clipboard; nothing is typed for you automatically.
|
|
479
|
+
- **Nothing heard** (silence, or a microphone that isn't actually picking
|
|
480
|
+
anything up) ends the dictation quietly — no clipboard write.
|
|
481
|
+
- **Cancel** with a second press *within two seconds* of the first, or at
|
|
482
|
+
any point with `vocalize listen --cancel` — the audio is discarded, never
|
|
483
|
+
transcribed.
|
|
484
|
+
- **A third press while transcribing is refused**: a Pop, and "Still
|
|
485
|
+
transcribing the last dictation." Wait for the clipboard notification, or
|
|
486
|
+
`--cancel`, before dictating again.
|
|
487
|
+
|
|
488
|
+
### `vocalize listen`
|
|
489
|
+
|
|
490
|
+
`vocalize dictate` (the hotkey's command) is `vocalize listen --toggle`
|
|
491
|
+
under another name. `listen` is the general primitive:
|
|
492
|
+
|
|
493
|
+
```bash
|
|
494
|
+
vocalize listen # record until Enter/Ctrl-C, print to stdout
|
|
495
|
+
vocalize listen --toggle # start, or stop and copy to the clipboard
|
|
496
|
+
vocalize listen --cancel # discard whatever is in progress
|
|
497
|
+
vocalize listen --wav clip.wav # transcribe a file you already have
|
|
498
|
+
vocalize listen --check # microphone + install readiness
|
|
499
|
+
vocalize listen --list-devices # input device names for [stt] input_device
|
|
500
|
+
vocalize listen --max-seconds 30 # cap this one recording
|
|
501
|
+
```
|
|
502
|
+
|
|
503
|
+
`--wav` is trusted input — the file has to be 16 kHz mono 16-bit WAV
|
|
504
|
+
(exactly what the recorder, and `say --data-format=LEI16@16000`, produce);
|
|
505
|
+
a malformed file gets a plain error naming the format, not a crash.
|
|
506
|
+
`--cleanup` tidies the transcript with Claude before it's delivered; it
|
|
507
|
+
applies to a live recording (`--toggle`/`dictate`, or plain `listen`) and
|
|
508
|
+
has no effect on `--wav`, which transcribes literally. `--max-seconds`
|
|
509
|
+
overrides `[stt] max_seconds` for one invocation.
|
|
510
|
+
|
|
511
|
+
### Configuration: the `[stt]` table
|
|
512
|
+
|
|
513
|
+
```toml
|
|
514
|
+
[stt]
|
|
515
|
+
model = "small.en" # base.en | small.en | large-v3-turbo-q5_0
|
|
516
|
+
language = "en" # a whisper.cpp language code
|
|
517
|
+
input_device = "" # "" = system default; else an exact name from --list-devices
|
|
518
|
+
cleanup = false # send the transcript (never audio) to Claude first
|
|
519
|
+
max_seconds = 120 # 1-600; the recorder self-stops here, dictate backstops it
|
|
520
|
+
sounds = true # the Tink/Pop/Glass feedback sounds
|
|
521
|
+
```
|
|
522
|
+
|
|
523
|
+
| Key | Allowed values | Default |
|
|
524
|
+
|---|---|---|
|
|
525
|
+
| `model` | `base.en`, `small.en`, `large-v3-turbo-q5_0` | `small.en` |
|
|
526
|
+
| `language` | a whisper.cpp language code (`en`, `es`, `fr`, …); an `.en` model must stay `en` | `en` |
|
|
527
|
+
| `input_device` | `""` (system default) or an exact name from `vocalize listen --list-devices`; ≤ 128 characters, printable, can't start with `-` | `""` |
|
|
528
|
+
| `cleanup` | `true` / `false` | `false` |
|
|
529
|
+
| `paste` | reserved — not implemented in 0.10.0 | `false` |
|
|
530
|
+
| `max_seconds` | integer, 1–600 | `120` |
|
|
531
|
+
| `sounds` | `true` / `false` | `true` |
|
|
532
|
+
|
|
533
|
+
An unknown key warns on stderr; a bad value is a `ConfigError` naming it —
|
|
534
|
+
every one of these becomes a subprocess argument eventually, so nothing
|
|
535
|
+
here is trusted on the way in.
|
|
536
|
+
|
|
537
|
+
**The input-device gotcha:** a pair of Bluetooth earbuds that are *paired*
|
|
538
|
+
but not actually in your ears still shows up as the default input device —
|
|
539
|
+
and delivers digital silence. If dictation keeps saying nothing was heard,
|
|
540
|
+
run `vocalize listen --list-devices`, copy the real microphone's name, and
|
|
541
|
+
set it:
|
|
542
|
+
|
|
543
|
+
```toml
|
|
544
|
+
[stt]
|
|
545
|
+
input_device = "MacBook Pro Microphone"
|
|
546
|
+
```
|
|
547
|
+
|
|
548
|
+
### `vocalize status`
|
|
549
|
+
|
|
550
|
+
```bash
|
|
551
|
+
vocalize status
|
|
552
|
+
```
|
|
553
|
+
|
|
554
|
+
prints one row per provider in your chain, plus four dictation rows once
|
|
555
|
+
dictation has been set up at all — an `[stt]` table in your config, a
|
|
556
|
+
built recorder, or a model on disk (a machine that never opted in doesn't
|
|
557
|
+
get four permanent red rows for a feature nobody asked for):
|
|
558
|
+
|
|
559
|
+
| Row | Reports |
|
|
560
|
+
|---|---|
|
|
561
|
+
| `stt model` | whether a whisper.cpp model is on disk |
|
|
562
|
+
| `recorder` | whether Vocalize Recorder is built |
|
|
563
|
+
| `microphone` | authorized / denied / not asked yet — from the last `listen --check`, never by launching the recorder itself |
|
|
564
|
+
| `input device` | whether the configured (or default) input device is actually present |
|
|
565
|
+
|
|
566
|
+
`--json` prints the same rows as a list. Exit code is 0 when every row —
|
|
567
|
+
providers and dictation both — is `ok`, 1 otherwise, so it composes with
|
|
568
|
+
`&&` in a script.
|
|
569
|
+
|
|
570
|
+
### Continuing an interrupted read
|
|
571
|
+
|
|
572
|
+
Starting a dictation stops any read in progress, but doesn't throw the rest
|
|
573
|
+
of it away (DEC-003). The process that was speaking remembers exactly
|
|
574
|
+
where it stopped, and once your transcript has landed, vocalize asks:
|
|
575
|
+
**"Continue the read you interrupted?"** — answer, or let the dialog give
|
|
576
|
+
up after 15 seconds (counted as no). From a terminal, the same thing is:
|
|
577
|
+
|
|
578
|
+
```bash
|
|
579
|
+
vocalize resume # continue where the last read left off
|
|
580
|
+
vocalize resume --forget # discard it instead
|
|
581
|
+
```
|
|
582
|
+
|
|
583
|
+
The record — one piece of audio and the text not yet spoken — lives at
|
|
584
|
+
`~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour, and is
|
|
585
|
+
deleted the moment you resume, decline, or `--forget` it. This is the
|
|
586
|
+
interrupted *read*, not dictation audio or a transcript — see Privacy,
|
|
587
|
+
below.
|
|
588
|
+
|
|
589
|
+
### Uninstall
|
|
590
|
+
|
|
591
|
+
```bash
|
|
592
|
+
vocalize local uninstall --stt
|
|
593
|
+
```
|
|
594
|
+
|
|
595
|
+
Removes the downloaded model file and the recorder bundle. The microphone
|
|
596
|
+
permission grant itself stays in System Settings — remove it there
|
|
597
|
+
(Privacy & Security › Microphone) if you want that gone too.
|
|
598
|
+
|
|
599
|
+
### Privacy
|
|
600
|
+
|
|
601
|
+
vocalize writes no transcript to disk and shows none in a notification —
|
|
602
|
+
it's held in memory and goes to your clipboard (or stdout, for
|
|
603
|
+
`vocalize listen`) and nowhere else. Recorded audio lives only in a
|
|
604
|
+
private temporary directory deleted the moment a dictation ends, one way
|
|
605
|
+
or another (a sweep clears anything a hard kill leaves behind after 24
|
|
606
|
+
hours). The **only** thing that ever leaves the machine is the transcript,
|
|
607
|
+
and only when you turn on `--cleanup`: it's sent to `claude -p` with every
|
|
608
|
+
tool denied, purely to fix punctuation and casing. The audio itself is
|
|
609
|
+
never sent anywhere.
|
|
610
|
+
|
|
611
|
+
`--cleanup` has one consequence worth knowing before you turn it on:
|
|
612
|
+
Claude Code logs the prompt and stdin of every print-mode run, so the
|
|
613
|
+
transcript is written in plaintext to `~/.claude/projects/…`. vocalize
|
|
614
|
+
cannot suppress that. It is off by default. Full accounting in
|
|
615
|
+
[docs/dictation.md](docs/dictation.md#privacy).
|
|
616
|
+
|
|
420
617
|
## macOS Quick Actions (highlight → speak)
|
|
421
618
|
|
|
422
|
-
|
|
619
|
+
Four Services let you use vocalize from any app without a terminal:
|
|
423
620
|
|
|
424
621
|
- **Speak with Vocalize** — highlight text anywhere, right-click →
|
|
425
622
|
Services → Speak with Vocalize. It stops whatever was already playing
|
|
@@ -436,8 +633,11 @@ Two Services let you use vocalize from any app without a terminal:
|
|
|
436
633
|
(`~/.claude/plans/`) aloud on demand. Made for the plan-approval moment:
|
|
437
634
|
the proposal card is up, you press your shortcut, hear the plan, then
|
|
438
635
|
accept or reject. Nothing reads unless you trigger it.
|
|
636
|
+
- **Dictate with Vocalize** — the dictation hotkey (see
|
|
637
|
+
[Dictation](#dictation-speech-to-text) above). Takes no input and shows
|
|
638
|
+
no window; press it, speak, press it again.
|
|
439
639
|
|
|
440
|
-
Install
|
|
640
|
+
Install all four:
|
|
441
641
|
|
|
442
642
|
```bash
|
|
443
643
|
python3 hooks/install_quick_action.py
|
|
@@ -605,14 +805,26 @@ vocalize/
|
|
|
605
805
|
chain.py # tries each provider in the chain in turn until one speaks
|
|
606
806
|
ledger.py # ~/.cache/vocalize/usage.json — local monthly budget tracking
|
|
607
807
|
providers/ # elevenlabs.py, openai.py, google.py, polly.py, say.py, kokoro.py
|
|
608
|
-
local/ #
|
|
808
|
+
local/ # opt-in download/verify + the uv-run worker scripts —
|
|
809
|
+
# kokoro_manifest.py/kokoro_worker.py (TTS) and
|
|
810
|
+
# whisper_manifest.py/whisper_worker.py (dictation)
|
|
811
|
+
recorder/ # VocalizeRecorder.swift + Info.plist.in — the ad-hoc-signed
|
|
812
|
+
# .app bundle that owns the microphone permission
|
|
609
813
|
audio.py # save to disk + play via the OS's native player
|
|
610
814
|
# (afplay / mpg123 / ffplay / PowerShell, whichever exists)
|
|
815
|
+
dictate.py # the hotkey's toggle state machine: record, transcribe,
|
|
816
|
+
# clipboard — nothing here ever becomes a file or a log line
|
|
817
|
+
interrupted.py # the record of a read a dictation interrupted, and the
|
|
818
|
+
# slice `vocalize resume` plays to continue it
|
|
819
|
+
readiness.py # vocalize status's per-provider + per-dictation-row probes,
|
|
820
|
+
# each on a timed daemon thread so a wedged keychain or
|
|
821
|
+
# microphone check can never hang the command
|
|
611
822
|
cli.py # click-based CLI wiring the above together
|
|
612
823
|
hooks/
|
|
613
824
|
claude_stop_hook.py # Claude Code Stop hook -> calls the vocalize CLI
|
|
614
825
|
install_hook.py # safely merges the hook into ~/.claude/settings.json
|
|
615
|
-
tests/ # pytest, all mocked — no API key
|
|
826
|
+
tests/ # pytest, all mocked — no API key, network, or
|
|
827
|
+
# microphone needed to run these
|
|
616
828
|
```
|
|
617
829
|
|
|
618
830
|
## Testing
|
|
@@ -622,8 +834,11 @@ pip install -e ".[dev]"
|
|
|
622
834
|
pytest
|
|
623
835
|
```
|
|
624
836
|
|
|
625
|
-
|
|
626
|
-
`tts.py`, so tests pass in a fake client instead of hitting the real
|
|
837
|
+
Over 1,180 tests, all offline: the ElevenLabs client is dependency-injected
|
|
838
|
+
into `tts.py`, so tests pass in a fake client instead of hitting the real
|
|
839
|
+
API, and dictation's tests fake the recorder (a tiny shell script honoring
|
|
840
|
+
a stop file) and the whisper worker instead of touching a microphone or a
|
|
841
|
+
model.
|
|
627
842
|
|
|
628
843
|
## Known limitations
|
|
629
844
|
|
|
@@ -676,6 +891,22 @@ All tests run offline: the ElevenLabs client is dependency-injected into
|
|
|
676
891
|
confidential information"); until you click Always Allow, every command
|
|
677
892
|
that needs that key waits. Click it once per Python binary, or supply the
|
|
678
893
|
key through its environment variable instead.
|
|
894
|
+
- **Dictation is macOS only.** It depends on `AVFoundation`, LaunchServices,
|
|
895
|
+
and a Swift-compiled `.app` bundle for the microphone permission — there's
|
|
896
|
+
no equivalent path on Linux or Windows.
|
|
897
|
+
- **Rebuilding the recorder means re-granting the microphone.** Vocalize
|
|
898
|
+
Recorder's ad-hoc code signature is what macOS ties the permission grant
|
|
899
|
+
to; when its Swift source changes (a vocalize upgrade that touches it),
|
|
900
|
+
`vocalize local install --stt` rebuilds the bundle and warns you to
|
|
901
|
+
re-approve it in System Settings › Privacy & Security › Microphone. An
|
|
902
|
+
install that doesn't change the source never re-signs, so this isn't
|
|
903
|
+
every upgrade — only ones that touch the recorder.
|
|
904
|
+
- **`small.en` mishears jargon.** The default model does fine on ordinary
|
|
905
|
+
speech but can mangle project-specific words (`pyproject`, a function
|
|
906
|
+
name) — pick `large-v3-turbo-q5_0` for better accuracy, or turn on
|
|
907
|
+
`[stt] cleanup` so Claude fixes obvious transcription noise before it
|
|
908
|
+
reaches your clipboard (it still can't guess a word it never heard
|
|
909
|
+
correctly).
|
|
679
910
|
|
|
680
911
|
## License
|
|
681
912
|
|