vocalize-cli 0.9.1__tar.gz → 0.10.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (139) hide show
  1. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/CHANGELOG.md +157 -0
  2. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/PKG-INFO +239 -7
  3. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/README.md +238 -6
  4. vocalize_cli-0.10.1/docs/dictation.md +531 -0
  5. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/docs/installation.md +79 -8
  6. vocalize_cli-0.10.1/docs/next-features-analysis.md +133 -0
  7. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/choreography.md +68 -0
  8. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/decisions.md +445 -0
  9. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/design.md +210 -0
  10. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/plan.md +199 -0
  11. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/review-0.10.0.md +58 -0
  12. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-1-status/project-plan.md +39 -0
  13. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-1-status/report.md +125 -0
  14. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-1-status/validate-exit.sh +110 -0
  15. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-10-release-0-11-0/project-plan.md +46 -0
  16. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-10-release-0-11-0/validate-exit.sh +113 -0
  17. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-2-stt-runtime/project-plan.md +50 -0
  18. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-2-stt-runtime/report.md +191 -0
  19. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-2-stt-runtime/validate-exit.sh +114 -0
  20. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-3-recorder/project-plan.md +48 -0
  21. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-3-recorder/report.md +80 -0
  22. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-3-recorder/validate-exit.sh +115 -0
  23. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-4-dictation/project-plan.md +55 -0
  24. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-4-dictation/report.md +106 -0
  25. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-4-dictation/validate-exit.sh +119 -0
  26. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-5-resume/project-plan.md +46 -0
  27. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-5-resume/report.md +94 -0
  28. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-5-resume/validate-exit.sh +113 -0
  29. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-6-release-0-10-0/project-plan.md +50 -0
  30. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-6-release-0-10-0/report.md +72 -0
  31. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-6-release-0-10-0/validate-exit.sh +115 -0
  32. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-7-portal-read/project-plan.md +39 -0
  33. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-7-portal-read/validate-exit.sh +111 -0
  34. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-8-portal-write/project-plan.md +42 -0
  35. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-8-portal-write/validate-exit.sh +113 -0
  36. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-9-portal-page/project-plan.md +44 -0
  37. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/run-9-portal-page/validate-exit.sh +112 -0
  38. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/spike-2026-09-01.md +71 -0
  39. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/split-assessment.md +47 -0
  40. vocalize_cli-0.10.1/docs/plans/2026-09-next-features/verification.md +103 -0
  41. vocalize_cli-0.10.1/docs/research/2026-09-01-config-portal-design.md +58 -0
  42. vocalize_cli-0.10.1/docs/research/2026-09-01-dictation-design.md +158 -0
  43. vocalize_cli-0.10.1/docs/research/2026-09-01-voicebox-findings.md +59 -0
  44. vocalize_cli-0.10.1/docs/roadmap.md +24 -0
  45. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/claude_stop_hook.py +38 -3
  46. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/install_quick_action.py +1 -0
  47. vocalize_cli-0.10.1/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Info.plist +26 -0
  48. vocalize_cli-0.10.1/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Resources/document.wflow +206 -0
  49. vocalize_cli-0.10.1/tests/conftest.py +151 -0
  50. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_audio.py +297 -0
  51. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_chain.py +63 -0
  52. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_claude_stop_hook.py +61 -0
  53. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_cli.py +341 -0
  54. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_config.py +153 -0
  55. vocalize_cli-0.10.1/tests/test_dictate.py +2335 -0
  56. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_install_quick_action.py +81 -0
  57. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_kokoro_manifest.py +19 -0
  58. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_kokoro_provider.py +8 -2
  59. vocalize_cli-0.10.1/tests/test_listen_check.py +591 -0
  60. vocalize_cli-0.10.1/tests/test_local_install.py +1139 -0
  61. vocalize_cli-0.10.1/tests/test_readiness.py +608 -0
  62. vocalize_cli-0.10.1/tests/test_recorder_build.py +546 -0
  63. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_speak_options.py +13 -0
  64. vocalize_cli-0.10.1/tests/test_uv_path.py +62 -0
  65. vocalize_cli-0.10.1/tests/test_whisper_manifest.py +165 -0
  66. vocalize_cli-0.10.1/tests/test_whisper_worker.py +292 -0
  67. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/__init__.py +1 -1
  68. vocalize_cli-0.10.1/vocalize/audio.py +563 -0
  69. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/chain.py +43 -1
  70. vocalize_cli-0.10.1/vocalize/cli.py +1652 -0
  71. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/config.py +142 -1
  72. vocalize_cli-0.10.1/vocalize/dictate.py +1335 -0
  73. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/exceptions.py +21 -1
  74. vocalize_cli-0.10.1/vocalize/interrupted.py +323 -0
  75. vocalize_cli-0.10.1/vocalize/local/__init__.py +33 -0
  76. vocalize_cli-0.10.1/vocalize/local/install.py +598 -0
  77. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/local/kokoro_manifest.py +19 -0
  78. vocalize_cli-0.10.1/vocalize/local/whisper_manifest.py +153 -0
  79. vocalize_cli-0.10.1/vocalize/local/whisper_worker.py +166 -0
  80. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/kokoro.py +6 -38
  81. vocalize_cli-0.10.1/vocalize/readiness.py +380 -0
  82. vocalize_cli-0.10.1/vocalize/recorder/Info.plist.in +40 -0
  83. vocalize_cli-0.10.1/vocalize/recorder/Recorder.entitlements +8 -0
  84. vocalize_cli-0.10.1/vocalize/recorder/VocalizeRecorder.swift +527 -0
  85. vocalize_cli-0.9.1/tests/conftest.py +0 -78
  86. vocalize_cli-0.9.1/tests/test_local_install.py +0 -548
  87. vocalize_cli-0.9.1/vocalize/audio.py +0 -300
  88. vocalize_cli-0.9.1/vocalize/cli.py +0 -859
  89. vocalize_cli-0.9.1/vocalize/local/__init__.py +0 -7
  90. vocalize_cli-0.9.1/vocalize/local/install.py +0 -217
  91. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/.env.example +0 -0
  92. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/.github/workflows/ci.yml +0 -0
  93. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/.gitignore +0 -0
  94. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/LICENSE +0 -0
  95. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/docs/provider-credentials.md +0 -0
  96. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/install_hook.py +0 -0
  97. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Info.plist +0 -0
  98. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Resources/document.wflow +0 -0
  99. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Info.plist +0 -0
  100. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Resources/document.wflow +0 -0
  101. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Info.plist +0 -0
  102. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Resources/document.wflow +0 -0
  103. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/speak_options.py +0 -0
  104. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/hooks/speak_url_gate.py +0 -0
  105. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/pyproject.toml +0 -0
  106. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_auth.py +0 -0
  107. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_cache.py +0 -0
  108. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_clipboard.py +0 -0
  109. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_elevenlabs_provider.py +0 -0
  110. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_exceptions.py +0 -0
  111. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_google_provider.py +0 -0
  112. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_http.py +0 -0
  113. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_install_hook.py +0 -0
  114. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_kokoro_worker.py +0 -0
  115. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_ledger.py +0 -0
  116. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_openai_provider.py +0 -0
  117. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_polly_provider.py +0 -0
  118. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_preprocess.py +0 -0
  119. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_providers_registry.py +0 -0
  120. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_say_provider.py +0 -0
  121. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_speak_url_gate.py +0 -0
  122. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_tts.py +0 -0
  123. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/tests/test_wizard.py +0 -0
  124. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/__main__.py +0 -0
  125. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/auth.py +0 -0
  126. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/cache.py +0 -0
  127. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/clipboard.py +0 -0
  128. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/ledger.py +0 -0
  129. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/local/kokoro_worker.py +0 -0
  130. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/preprocess.py +0 -0
  131. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/__init__.py +0 -0
  132. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/_http.py +0 -0
  133. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/elevenlabs.py +0 -0
  134. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/google.py +0 -0
  135. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/openai.py +0 -0
  136. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/polly.py +0 -0
  137. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/providers/say.py +0 -0
  138. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/tts.py +0 -0
  139. {vocalize_cli-0.9.1 → vocalize_cli-0.10.1}/vocalize/wizard.py +0 -0
@@ -3,6 +3,163 @@
3
3
  All notable changes to this project are documented here. Format follows
4
4
  [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
5
5
 
6
+ ## 0.10.1 - 2026-09-02
7
+
8
+ Three fixes found in the first owner-present run of 0.10.0's dictation.
9
+ Together they meant no hotkey dictation could succeed on 0.10.0; upgrade.
10
+
11
+ ### Fixed
12
+
13
+ - **No dictation could ever start on a fresh install.** The recorder was
14
+ signed with the hardened runtime but without the
15
+ `com.apple.security.device.audio-input` entitlement, so macOS refused the
16
+ microphone on the spot — no permission dialog, status stuck at
17
+ `notDetermined` — and every first press ended in "The recorder did not
18
+ start". The bundle is now signed with
19
+ `vocalize/recorder/Recorder.entitlements`, and the entitlements are part
20
+ of the recorder's fingerprint, so `vocalize local install --stt` rebuilds
21
+ the bundle once (and, as with any rebuild, macOS asks for the microphone
22
+ again — it never actually asked before).
23
+ - **Every hotkey dictation ended in "Dictation failed" on a machine whose
24
+ `uv` came from Homebrew.** A Services environment has a bare PATH, and
25
+ `uv_path()` looked only there and in `~/.local/bin`; the same dictation
26
+ worked from a terminal. `/opt/homebrew/bin/uv` and `/usr/local/bin/uv`
27
+ are now tried too (this also covers Kokoro from a Quick Action).
28
+ - **Holding the dictation hotkey down turned into a cancel-and-restart
29
+ loop.** macOS re-fires a Service shortcut at the key-repeat rate, and
30
+ every repeat landed as a second press. Presses within half a second of
31
+ the previous one are now ignored as the same press; a deliberate cancel
32
+ is "press, a beat, press" inside the two-second window, as before.
33
+ - `hooks/claude_stop_hook.py --latest`, run from inside a Claude Code turn
34
+ (which is how `/speak` runs it), spoke the agent's own status line —
35
+ "Checking settings." — instead of the response the user asked to hear.
36
+ It now skips the turn in progress, back past the `/speak` message
37
+ itself, and speaks the response before it. From a plain terminal, where
38
+ no turn is in progress, `--latest` still speaks the newest response; the
39
+ hook tells the two apart by the `CLAUDECODE` variable Claude Code sets
40
+ in its shell. The Stop-hook path is unchanged.
41
+
42
+ ## 0.10.0 - 2026-09-02
43
+
44
+ ### Added
45
+
46
+ - **Local dictation.** Press a hotkey, speak, press it again, and the
47
+ transcript is on your clipboard — speech to text, entirely on-device via
48
+ [whisper.cpp](https://github.com/ggerganov/whisper.cpp)
49
+ (`pywhispercpp`). Nothing about a dictation leaves the machine unless
50
+ `[stt] cleanup` is turned on, and even then only the transcript is sent
51
+ (to `claude -p`, tools denied), never the audio. See
52
+ [docs/dictation.md](docs/dictation.md) for the full guide.
53
+ - `vocalize listen` (`--toggle`, `--cancel`, `--check`, `--list-devices`,
54
+ `--wav FILE`, `--cleanup`, `--max-seconds`) and `vocalize dictate` (an
55
+ alias for `listen --toggle`, under the name the hotkey uses).
56
+ - New Quick Action, **"Dictate with Vocalize"** — a no-input Service for
57
+ the dictation hotkey (⌃⌥⌘D suggested), installed by the existing
58
+ `hooks/install_quick_action.py` alongside the other three.
59
+ - `vocalize local install --stt [--model base.en|small.en|large-v3-turbo-q5_0]`
60
+ and `vocalize local uninstall --stt` — opt-in download-and-verify of a
61
+ whisper.cpp model, plus build-and-sign of **Vocalize Recorder**, the
62
+ small `.app` bundle that holds the microphone permission (macOS only
63
+ grants that to something with an identity). Nothing is downloaded or
64
+ compiled until you run `install --stt`, mirroring Kokoro's opt-in
65
+ install; the one-time Metal shader warm-up (~8s) is paid here, never
66
+ during a dictation.
67
+ - New `[stt]` config table — `model`, `language`, `input_device`,
68
+ `cleanup`, `paste` (reserved, not implemented yet), `max_seconds`,
69
+ `sounds` — validated on the way in the same way `[providers.*]` is, and
70
+ printed by `vocalize settings` as `stt.*` lines.
71
+ - `vocalize status` — a one-screen readiness check across every provider
72
+ in your chain, plus four dictation rows (`stt model`, `recorder`,
73
+ `microphone`, `input device`) once dictation has been set up at all.
74
+ `--json` prints the same rows as a list; exit 0 when everything is `ok`,
75
+ 1 otherwise.
76
+ - `vocalize resume [--forget]` — continue (or discard) a text-to-speech
77
+ read that a dictation interrupted. Starting a dictation stops any read
78
+ in progress, but vocalize now remembers exactly where it stopped and
79
+ offers to continue once the transcript has landed (a macOS dialog,
80
+ default Continue, 15s to answer); the record lives at
81
+ `~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour.
82
+ - New [docs/dictation.md](docs/dictation.md): install, the hotkey, every
83
+ `[stt]` key, `vocalize status`'s dictation rows, `resume`, and
84
+ troubleshooting keyed on `vocalize listen --check`'s exact messages and
85
+ exit codes.
86
+
87
+ ### Changed
88
+
89
+ - `vocalize/local/install.py` generalized to support more than one local
90
+ runtime's manifest and model files (previously hard-coded to Kokoro's).
91
+ Kokoro's own install, stamp, and `local status` output are unchanged
92
+ byte-for-byte; the whisper runtime downloads and stamps only the single
93
+ model you selected, not all three.
94
+ - `vocalize listen --check` measures the microphone permission by
95
+ launching Vocalize Recorder the same way a real dictation does
96
+ (through LaunchServices), not by exec'ing its binary directly — macOS
97
+ attributes a TCC grant to the *responsible* process, and exec'ing the
98
+ binary as a child of your shell reported the terminal's own grant
99
+ instead. Exit codes: 0 authorized-and-ready, 2 denied, 3 no usable
100
+ input device, 5 not asked yet (macOS `notDetermined`) — matching the
101
+ recorder's own contract — plus a new exit 1 meaning "vocalize's own
102
+ local install isn't finished" (not built, no model on disk, or the
103
+ recorder never reported back), which is a setup problem, not a
104
+ permission one.
105
+ - `audio.stop_playback()` gained a `remember=` flag. A dictation's stop
106
+ passes it, leaving a marker so the process that was playing can record
107
+ where it stopped — this is what makes `vocalize resume` possible. A
108
+ plain `vocalize stop` records nothing, as before.
109
+ - **A stop now silences every read already in flight**, not only the
110
+ player it kills. Playback is serialized machine-wide, so stopping one
111
+ read used to let the next queued one start speaking immediately — into
112
+ the microphone a dictation had just opened. A read *started* after the
113
+ stop is unaffected.
114
+ - Dictated text reaches the clipboard as a single line. Newlines are
115
+ collapsed there so a paste into a terminal cannot run as several
116
+ commands; `vocalize listen`'s stdout keeps them.
117
+ - `vocalize listen --check` now measures the input device configured in
118
+ `[stt] input_device` rather than the system default, and records what it
119
+ saw with a timestamp — so `vocalize status` says how old that
120
+ "authorized" verdict is instead of implying it is current.
121
+
122
+ ### Fixed
123
+
124
+ - `vocalize local install --stt`, re-run against a model that already
125
+ verified, now re-warms the runtime instead of reporting "already
126
+ installed" and stopping — a machine where only the runtime failed to
127
+ start (no Metal, a build hiccup) previously had no way to retry that
128
+ short of a full uninstall and 465 MB re-download.
129
+ - `vocalize local status` reports every installed speech-to-text model,
130
+ not just the default — installing a non-default model with `--model`
131
+ no longer looks unfinished.
132
+ - `vocalize local uninstall --stt` no longer crashes on a symlinked model
133
+ directory or recorder bundle; it reports the symlink and leaves it for
134
+ you to remove.
135
+ - `vocalize status` no longer raises on an unrecognized `VOCALIZE_CHAIN`;
136
+ like any other misconfiguration, it degrades to one failing row instead
137
+ of crashing the command. A probe that raises is reported by exception
138
+ type only — never its message, which could otherwise echo
139
+ credential-shaped text onto the screen.
140
+ - A read stopped by a dictation while a streaming provider's next chunk
141
+ was still rendering (nothing audible playing at that exact instant)
142
+ used to lose the rest of the read with no way to get it back; it's now
143
+ recorded and resumable like any other interruption. The same now holds
144
+ for a read still being synthesized (no player exists yet) and for a
145
+ plain `vocalize stop` landing in that gap.
146
+ - The first dictation on a fresh install no longer fails while macOS is
147
+ asking for the microphone. The permission dialog can sit on screen for
148
+ minutes; the press now waits for your answer and starts recording when
149
+ you click Allow, instead of giving up after five seconds and reporting
150
+ a failure that had not happened.
151
+ - `vocalize resume` continues the read in the voice, model, speed and
152
+ chunk size it was stopped in. It previously fell back to the config
153
+ defaults, which also missed the audio cache and re-synthesized (and
154
+ re-charged for) the whole remainder.
155
+ - A ten-minute dictation is no longer mistaken for a crashed one. The
156
+ claim a stop puts on a take is now aged from its own progress rather
157
+ than from when recording began, so a long take or a stop queued behind
158
+ a long read cannot be reaped mid-transcription.
159
+ - `~/.cache/vocalize` and `~/.cache/vocalize/bin` are tightened to 0700
160
+ even when they already existed. The files inside were always 0600, but
161
+ the directory listing said whether a dictation was in progress.
162
+
6
163
  ## 0.9.1 - 2026-09-01
7
164
 
8
165
  ### Fixed
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: vocalize-cli
3
- Version: 0.9.1
3
+ Version: 0.10.1
4
4
  Summary: A CLI that turns text, markdown, or piped stdin into speech via the ElevenLabs API, with markdown-table-aware preprocessing.
5
5
  Project-URL: Homepage, https://github.com/matthager12-collab/vocalize
6
6
  Project-URL: Repository, https://github.com/matthager12-collab/vocalize
@@ -289,6 +289,12 @@ in the current directory, then the OS keychain. `vocalize auth login` sets
289
289
  up the keychain entry; `vocalize auth status` shows which of those sources
290
290
  is currently supplying the key.
291
291
 
292
+ A `[stt]` table configures dictation the same way `[providers.<name>]`
293
+ configures a TTS provider — see
294
+ [Configuration: the `[stt]` table](#configuration-the-stt-table) under
295
+ [Dictation](#dictation-speech-to-text) below for every key and its
296
+ allowlist.
297
+
292
298
  ## Providers and fallback
293
299
 
294
300
  vocalize tries providers in order — a **chain** — until one speaks. The
@@ -417,9 +423,201 @@ instead of waiting for the whole thing to render. Measured on this Mac
417
423
  rendering. `vocalize stop`, run from any terminal, halts a Kokoro read
418
424
  mid-sentence the same as any other provider.
419
425
 
426
+ ## Dictation (speech to text)
427
+
428
+ Everything above turns text into speech. Dictation runs the other way:
429
+ press a hotkey, speak, press it again, and the words land on your
430
+ clipboard. It's [whisper.cpp](https://github.com/ggerganov/whisper.cpp) via
431
+ [`pywhispercpp`](https://github.com/absadiki/pywhispercpp), entirely
432
+ on-device — nothing you say leaves the Mac unless you turn on `--cleanup`
433
+ (below), and even then only the *transcript* goes anywhere, never the audio.
434
+
435
+ ### Install
436
+
437
+ ```bash
438
+ vocalize local install --stt
439
+ ```
440
+
441
+ A separate opt-in from Kokoro's `vocalize local install` — nothing here is
442
+ downloaded or built until you run this. It:
443
+
444
+ 1. Downloads one whisper.cpp model (`small.en` by default, ~465 MB) from a
445
+ pinned Hugging Face revision, verified against a pinned sha256 before
446
+ it's kept.
447
+ 2. Compiles and ad-hoc signs a small Swift recorder bundle, **Vocalize
448
+ Recorder** — the thing that actually owns the microphone permission;
449
+ macOS won't grant that to a bare command-line tool.
450
+ 3. Warms the runtime, paying a one-time ~8-second Metal shader compile
451
+ right here, so no dictation ever stalls on it later.
452
+
453
+ The first real dictation prompts for microphone access naming **"Vocalize
454
+ Recorder"** — approve it once, like any other app's first-run permission
455
+ prompt. Pick a different model with `--model large-v3-turbo-q5_0` (~547 MB,
456
+ more accurate) or `--model base.en` (~141 MB, fastest, least accurate).
457
+
458
+ ### The hotkey
459
+
460
+ ```bash
461
+ python3 hooks/install_quick_action.py
462
+ ```
463
+
464
+ then assign a shortcut under **System Settings › Keyboard › Keyboard
465
+ Shortcuts › Services › Text › "Dictate with Vocalize"** — ⌃⌥⌘D is free by
466
+ default and a sensible pick. `vocalize dictate` is the same command from a
467
+ terminal, if you'd rather trigger it that way.
468
+
469
+ ### How a dictation works
470
+
471
+ - **Press** the hotkey — a Tink plays, the recorder starts, and any read
472
+ currently playing is stopped first (dictation and playback never
473
+ overlap; vocalize remembers where the read was cut off — see
474
+ [Continuing an interrupted read](#continuing-an-interrupted-read)).
475
+ - **Speak.**
476
+ - **Press again** — a Pop plays, recording stops, and the audio is
477
+ transcribed on-device. If anything was heard, a Glass plays and the
478
+ transcript is on your clipboard; nothing is typed for you automatically.
479
+ - **Nothing heard** (silence, or a microphone that isn't actually picking
480
+ anything up) ends the dictation quietly — no clipboard write.
481
+ - **Cancel** with a second press *within two seconds* of the first (but
482
+ not within half a second — that's a held key, and it's ignored), or at
483
+ any point with `vocalize listen --cancel` — the audio is discarded, never
484
+ transcribed.
485
+ - **A third press while transcribing is refused**: a Pop, and "Still
486
+ transcribing the last dictation." Wait for the clipboard notification, or
487
+ `--cancel`, before dictating again.
488
+
489
+ ### `vocalize listen`
490
+
491
+ `vocalize dictate` (the hotkey's command) is `vocalize listen --toggle`
492
+ under another name. `listen` is the general primitive:
493
+
494
+ ```bash
495
+ vocalize listen # record until Enter/Ctrl-C, print to stdout
496
+ vocalize listen --toggle # start, or stop and copy to the clipboard
497
+ vocalize listen --cancel # discard whatever is in progress
498
+ vocalize listen --wav clip.wav # transcribe a file you already have
499
+ vocalize listen --check # microphone + install readiness
500
+ vocalize listen --list-devices # input device names for [stt] input_device
501
+ vocalize listen --max-seconds 30 # cap this one recording
502
+ ```
503
+
504
+ `--wav` is trusted input — the file has to be 16 kHz mono 16-bit WAV
505
+ (exactly what the recorder, and `say --data-format=LEI16@16000`, produce);
506
+ a malformed file gets a plain error naming the format, not a crash.
507
+ `--cleanup` tidies the transcript with Claude before it's delivered; it
508
+ applies to a live recording (`--toggle`/`dictate`, or plain `listen`) and
509
+ has no effect on `--wav`, which transcribes literally. `--max-seconds`
510
+ overrides `[stt] max_seconds` for one invocation.
511
+
512
+ ### Configuration: the `[stt]` table
513
+
514
+ ```toml
515
+ [stt]
516
+ model = "small.en" # base.en | small.en | large-v3-turbo-q5_0
517
+ language = "en" # a whisper.cpp language code
518
+ input_device = "" # "" = system default; else an exact name from --list-devices
519
+ cleanup = false # send the transcript (never audio) to Claude first
520
+ max_seconds = 120 # 1-600; the recorder self-stops here, dictate backstops it
521
+ sounds = true # the Tink/Pop/Glass feedback sounds
522
+ ```
523
+
524
+ | Key | Allowed values | Default |
525
+ |---|---|---|
526
+ | `model` | `base.en`, `small.en`, `large-v3-turbo-q5_0` | `small.en` |
527
+ | `language` | a whisper.cpp language code (`en`, `es`, `fr`, …); an `.en` model must stay `en` | `en` |
528
+ | `input_device` | `""` (system default) or an exact name from `vocalize listen --list-devices`; ≤ 128 characters, printable, can't start with `-` | `""` |
529
+ | `cleanup` | `true` / `false` | `false` |
530
+ | `paste` | reserved — not implemented in 0.10.0 | `false` |
531
+ | `max_seconds` | integer, 1–600 | `120` |
532
+ | `sounds` | `true` / `false` | `true` |
533
+
534
+ An unknown key warns on stderr; a bad value is a `ConfigError` naming it —
535
+ every one of these becomes a subprocess argument eventually, so nothing
536
+ here is trusted on the way in.
537
+
538
+ **The input-device gotcha:** a pair of Bluetooth earbuds that are *paired*
539
+ but not actually in your ears still shows up as the default input device —
540
+ and delivers digital silence. If dictation keeps saying nothing was heard,
541
+ run `vocalize listen --list-devices`, copy the real microphone's name, and
542
+ set it:
543
+
544
+ ```toml
545
+ [stt]
546
+ input_device = "MacBook Pro Microphone"
547
+ ```
548
+
549
+ ### `vocalize status`
550
+
551
+ ```bash
552
+ vocalize status
553
+ ```
554
+
555
+ prints one row per provider in your chain, plus four dictation rows once
556
+ dictation has been set up at all — an `[stt]` table in your config, a
557
+ built recorder, or a model on disk (a machine that never opted in doesn't
558
+ get four permanent red rows for a feature nobody asked for):
559
+
560
+ | Row | Reports |
561
+ |---|---|
562
+ | `stt model` | whether a whisper.cpp model is on disk |
563
+ | `recorder` | whether Vocalize Recorder is built |
564
+ | `microphone` | authorized / denied / not asked yet — from the last `listen --check`, never by launching the recorder itself |
565
+ | `input device` | whether the configured (or default) input device is actually present |
566
+
567
+ `--json` prints the same rows as a list. Exit code is 0 when every row —
568
+ providers and dictation both — is `ok`, 1 otherwise, so it composes with
569
+ `&&` in a script.
570
+
571
+ ### Continuing an interrupted read
572
+
573
+ Starting a dictation stops any read in progress, but doesn't throw the rest
574
+ of it away (DEC-003). The process that was speaking remembers exactly
575
+ where it stopped, and once your transcript has landed, vocalize asks:
576
+ **"Continue the read you interrupted?"** — answer, or let the dialog give
577
+ up after 15 seconds (counted as no). From a terminal, the same thing is:
578
+
579
+ ```bash
580
+ vocalize resume # continue where the last read left off
581
+ vocalize resume --forget # discard it instead
582
+ ```
583
+
584
+ The record — one piece of audio and the text not yet spoken — lives at
585
+ `~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour, and is
586
+ deleted the moment you resume, decline, or `--forget` it. This is the
587
+ interrupted *read*, not dictation audio or a transcript — see Privacy,
588
+ below.
589
+
590
+ ### Uninstall
591
+
592
+ ```bash
593
+ vocalize local uninstall --stt
594
+ ```
595
+
596
+ Removes the downloaded model file and the recorder bundle. The microphone
597
+ permission grant itself stays in System Settings — remove it there
598
+ (Privacy & Security › Microphone) if you want that gone too.
599
+
600
+ ### Privacy
601
+
602
+ vocalize writes no transcript to disk and shows none in a notification —
603
+ it's held in memory and goes to your clipboard (or stdout, for
604
+ `vocalize listen`) and nowhere else. Recorded audio lives only in a
605
+ private temporary directory deleted the moment a dictation ends, one way
606
+ or another (a sweep clears anything a hard kill leaves behind after 24
607
+ hours). The **only** thing that ever leaves the machine is the transcript,
608
+ and only when you turn on `--cleanup`: it's sent to `claude -p` with every
609
+ tool denied, purely to fix punctuation and casing. The audio itself is
610
+ never sent anywhere.
611
+
612
+ `--cleanup` has one consequence worth knowing before you turn it on:
613
+ Claude Code logs the prompt and stdin of every print-mode run, so the
614
+ transcript is written in plaintext to `~/.claude/projects/…`. vocalize
615
+ cannot suppress that. It is off by default. Full accounting in
616
+ [docs/dictation.md](docs/dictation.md#privacy).
617
+
420
618
  ## macOS Quick Actions (highlight → speak)
421
619
 
422
- Two Services let you use vocalize from any app without a terminal:
620
+ Four Services let you use vocalize from any app without a terminal:
423
621
 
424
622
  - **Speak with Vocalize** — highlight text anywhere, right-click →
425
623
  Services → Speak with Vocalize. It stops whatever was already playing
@@ -436,8 +634,11 @@ Two Services let you use vocalize from any app without a terminal:
436
634
  (`~/.claude/plans/`) aloud on demand. Made for the plan-approval moment:
437
635
  the proposal card is up, you press your shortcut, hear the plan, then
438
636
  accept or reject. Nothing reads unless you trigger it.
637
+ - **Dictate with Vocalize** — the dictation hotkey (see
638
+ [Dictation](#dictation-speech-to-text) above). Takes no input and shows
639
+ no window; press it, speak, press it again.
439
640
 
440
- Install both:
641
+ Install all four:
441
642
 
442
643
  ```bash
443
644
  python3 hooks/install_quick_action.py
@@ -605,14 +806,26 @@ vocalize/
605
806
  chain.py # tries each provider in the chain in turn until one speaks
606
807
  ledger.py # ~/.cache/vocalize/usage.json — local monthly budget tracking
607
808
  providers/ # elevenlabs.py, openai.py, google.py, polly.py, say.py, kokoro.py
608
- local/ # Kokoro's opt-in download/verify + the uv-run worker script
809
+ local/ # opt-in download/verify + the uv-run worker scripts —
810
+ # kokoro_manifest.py/kokoro_worker.py (TTS) and
811
+ # whisper_manifest.py/whisper_worker.py (dictation)
812
+ recorder/ # VocalizeRecorder.swift + Info.plist.in — the ad-hoc-signed
813
+ # .app bundle that owns the microphone permission
609
814
  audio.py # save to disk + play via the OS's native player
610
815
  # (afplay / mpg123 / ffplay / PowerShell, whichever exists)
816
+ dictate.py # the hotkey's toggle state machine: record, transcribe,
817
+ # clipboard — nothing here ever becomes a file or a log line
818
+ interrupted.py # the record of a read a dictation interrupted, and the
819
+ # slice `vocalize resume` plays to continue it
820
+ readiness.py # vocalize status's per-provider + per-dictation-row probes,
821
+ # each on a timed daemon thread so a wedged keychain or
822
+ # microphone check can never hang the command
611
823
  cli.py # click-based CLI wiring the above together
612
824
  hooks/
613
825
  claude_stop_hook.py # Claude Code Stop hook -> calls the vocalize CLI
614
826
  install_hook.py # safely merges the hook into ~/.claude/settings.json
615
- tests/ # pytest, all mocked — no API key needed to run these
827
+ tests/ # pytest, all mocked — no API key, network, or
828
+ # microphone needed to run these
616
829
  ```
617
830
 
618
831
  ## Testing
@@ -622,8 +835,11 @@ pip install -e ".[dev]"
622
835
  pytest
623
836
  ```
624
837
 
625
- All tests run offline: the ElevenLabs client is dependency-injected into
626
- `tts.py`, so tests pass in a fake client instead of hitting the real API.
838
+ Over 1,180 tests, all offline: the ElevenLabs client is dependency-injected
839
+ into `tts.py`, so tests pass in a fake client instead of hitting the real
840
+ API, and dictation's tests fake the recorder (a tiny shell script honoring
841
+ a stop file) and the whisper worker instead of touching a microphone or a
842
+ model.
627
843
 
628
844
  ## Known limitations
629
845
 
@@ -676,6 +892,22 @@ All tests run offline: the ElevenLabs client is dependency-injected into
676
892
  confidential information"); until you click Always Allow, every command
677
893
  that needs that key waits. Click it once per Python binary, or supply the
678
894
  key through its environment variable instead.
895
+ - **Dictation is macOS only.** It depends on `AVFoundation`, LaunchServices,
896
+ and a Swift-compiled `.app` bundle for the microphone permission — there's
897
+ no equivalent path on Linux or Windows.
898
+ - **Rebuilding the recorder means re-granting the microphone.** Vocalize
899
+ Recorder's ad-hoc code signature is what macOS ties the permission grant
900
+ to; when its Swift source changes (a vocalize upgrade that touches it),
901
+ `vocalize local install --stt` rebuilds the bundle and warns you to
902
+ re-approve it in System Settings › Privacy & Security › Microphone. An
903
+ install that doesn't change the source never re-signs, so this isn't
904
+ every upgrade — only ones that touch the recorder.
905
+ - **`small.en` mishears jargon.** The default model does fine on ordinary
906
+ speech but can mangle project-specific words (`pyproject`, a function
907
+ name) — pick `large-v3-turbo-q5_0` for better accuracy, or turn on
908
+ `[stt] cleanup` so Claude fixes obvious transcription noise before it
909
+ reaches your clipboard (it still can't guess a word it never heard
910
+ correctly).
679
911
 
680
912
  ## License
681
913