vocalize-cli 0.9.0__tar.gz → 0.10.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (138) hide show
  1. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/CHANGELOG.md +141 -0
  2. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/PKG-INFO +238 -7
  3. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/README.md +237 -6
  4. vocalize_cli-0.10.0/docs/dictation.md +528 -0
  5. vocalize_cli-0.10.0/docs/installation.md +470 -0
  6. vocalize_cli-0.10.0/docs/next-features-analysis.md +133 -0
  7. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/choreography.md +68 -0
  8. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/decisions.md +445 -0
  9. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/design.md +210 -0
  10. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/plan.md +199 -0
  11. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/review-0.10.0.md +58 -0
  12. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-1-status/project-plan.md +39 -0
  13. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-1-status/report.md +125 -0
  14. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-1-status/validate-exit.sh +110 -0
  15. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-10-release-0-11-0/project-plan.md +46 -0
  16. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-10-release-0-11-0/validate-exit.sh +113 -0
  17. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-2-stt-runtime/project-plan.md +50 -0
  18. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-2-stt-runtime/report.md +191 -0
  19. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-2-stt-runtime/validate-exit.sh +114 -0
  20. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-3-recorder/project-plan.md +48 -0
  21. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-3-recorder/report.md +80 -0
  22. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-3-recorder/validate-exit.sh +115 -0
  23. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-4-dictation/project-plan.md +55 -0
  24. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-4-dictation/report.md +106 -0
  25. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-4-dictation/validate-exit.sh +119 -0
  26. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-5-resume/project-plan.md +46 -0
  27. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-5-resume/report.md +94 -0
  28. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-5-resume/validate-exit.sh +113 -0
  29. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-6-release-0-10-0/project-plan.md +50 -0
  30. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-6-release-0-10-0/report.md +71 -0
  31. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-6-release-0-10-0/validate-exit.sh +115 -0
  32. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-7-portal-read/project-plan.md +39 -0
  33. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-7-portal-read/validate-exit.sh +111 -0
  34. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-8-portal-write/project-plan.md +42 -0
  35. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-8-portal-write/validate-exit.sh +113 -0
  36. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-9-portal-page/project-plan.md +44 -0
  37. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/run-9-portal-page/validate-exit.sh +112 -0
  38. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/spike-2026-09-01.md +71 -0
  39. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/split-assessment.md +47 -0
  40. vocalize_cli-0.10.0/docs/plans/2026-09-next-features/verification.md +103 -0
  41. vocalize_cli-0.10.0/docs/research/2026-09-01-config-portal-design.md +58 -0
  42. vocalize_cli-0.10.0/docs/research/2026-09-01-dictation-design.md +158 -0
  43. vocalize_cli-0.10.0/docs/research/2026-09-01-voicebox-findings.md +59 -0
  44. vocalize_cli-0.10.0/docs/roadmap.md +22 -0
  45. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/install_quick_action.py +1 -0
  46. vocalize_cli-0.10.0/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Info.plist +26 -0
  47. vocalize_cli-0.10.0/hooks/quick_actions/Dictate with Vocalize.workflow/Contents/Resources/document.wflow +206 -0
  48. vocalize_cli-0.10.0/tests/conftest.py +151 -0
  49. vocalize_cli-0.10.0/tests/test_audio.py +731 -0
  50. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_chain.py +63 -0
  51. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_cli.py +341 -0
  52. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_config.py +153 -0
  53. vocalize_cli-0.10.0/tests/test_dictate.py +2264 -0
  54. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_install_quick_action.py +81 -0
  55. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_kokoro_manifest.py +19 -0
  56. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_kokoro_provider.py +6 -2
  57. vocalize_cli-0.10.0/tests/test_listen_check.py +591 -0
  58. vocalize_cli-0.10.0/tests/test_local_install.py +1139 -0
  59. vocalize_cli-0.10.0/tests/test_readiness.py +608 -0
  60. vocalize_cli-0.10.0/tests/test_recorder_build.py +510 -0
  61. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_speak_options.py +13 -0
  62. vocalize_cli-0.10.0/tests/test_whisper_manifest.py +165 -0
  63. vocalize_cli-0.10.0/tests/test_whisper_worker.py +292 -0
  64. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/__init__.py +1 -1
  65. vocalize_cli-0.10.0/vocalize/audio.py +563 -0
  66. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/chain.py +43 -1
  67. vocalize_cli-0.10.0/vocalize/cli.py +1652 -0
  68. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/config.py +142 -1
  69. vocalize_cli-0.10.0/vocalize/dictate.py +1292 -0
  70. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/exceptions.py +21 -1
  71. vocalize_cli-0.10.0/vocalize/interrupted.py +323 -0
  72. vocalize_cli-0.10.0/vocalize/local/__init__.py +25 -0
  73. vocalize_cli-0.10.0/vocalize/local/install.py +588 -0
  74. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/local/kokoro_manifest.py +19 -0
  75. vocalize_cli-0.10.0/vocalize/local/whisper_manifest.py +153 -0
  76. vocalize_cli-0.10.0/vocalize/local/whisper_worker.py +166 -0
  77. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/kokoro.py +6 -38
  78. vocalize_cli-0.10.0/vocalize/readiness.py +380 -0
  79. vocalize_cli-0.10.0/vocalize/recorder/Info.plist.in +40 -0
  80. vocalize_cli-0.10.0/vocalize/recorder/VocalizeRecorder.swift +527 -0
  81. vocalize_cli-0.9.0/tests/conftest.py +0 -67
  82. vocalize_cli-0.9.0/tests/test_audio.py +0 -375
  83. vocalize_cli-0.9.0/tests/test_local_install.py +0 -548
  84. vocalize_cli-0.9.0/vocalize/audio.py +0 -246
  85. vocalize_cli-0.9.0/vocalize/cli.py +0 -859
  86. vocalize_cli-0.9.0/vocalize/local/__init__.py +0 -7
  87. vocalize_cli-0.9.0/vocalize/local/install.py +0 -217
  88. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/.env.example +0 -0
  89. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/.github/workflows/ci.yml +0 -0
  90. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/.gitignore +0 -0
  91. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/LICENSE +0 -0
  92. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/docs/provider-credentials.md +0 -0
  93. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/claude_stop_hook.py +0 -0
  94. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/install_hook.py +0 -0
  95. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Info.plist +0 -0
  96. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak Latest Plan.workflow/Contents/Resources/document.wflow +0 -0
  97. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Info.plist +0 -0
  98. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Speak with Vocalize.workflow/Contents/Resources/document.wflow +0 -0
  99. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Info.plist +0 -0
  100. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/quick_actions/Stop Vocalize.workflow/Contents/Resources/document.wflow +0 -0
  101. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/speak_options.py +0 -0
  102. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/hooks/speak_url_gate.py +0 -0
  103. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/pyproject.toml +0 -0
  104. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_auth.py +0 -0
  105. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_cache.py +0 -0
  106. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_claude_stop_hook.py +0 -0
  107. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_clipboard.py +0 -0
  108. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_elevenlabs_provider.py +0 -0
  109. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_exceptions.py +0 -0
  110. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_google_provider.py +0 -0
  111. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_http.py +0 -0
  112. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_install_hook.py +0 -0
  113. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_kokoro_worker.py +0 -0
  114. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_ledger.py +0 -0
  115. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_openai_provider.py +0 -0
  116. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_polly_provider.py +0 -0
  117. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_preprocess.py +0 -0
  118. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_providers_registry.py +0 -0
  119. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_say_provider.py +0 -0
  120. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_speak_url_gate.py +0 -0
  121. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_tts.py +0 -0
  122. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/tests/test_wizard.py +0 -0
  123. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/__main__.py +0 -0
  124. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/auth.py +0 -0
  125. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/cache.py +0 -0
  126. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/clipboard.py +0 -0
  127. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/ledger.py +0 -0
  128. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/local/kokoro_worker.py +0 -0
  129. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/preprocess.py +0 -0
  130. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/__init__.py +0 -0
  131. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/_http.py +0 -0
  132. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/elevenlabs.py +0 -0
  133. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/google.py +0 -0
  134. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/openai.py +0 -0
  135. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/polly.py +0 -0
  136. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/providers/say.py +0 -0
  137. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/tts.py +0 -0
  138. {vocalize_cli-0.9.0 → vocalize_cli-0.10.0}/vocalize/wizard.py +0 -0
@@ -3,6 +3,147 @@
3
3
  All notable changes to this project are documented here. Format follows
4
4
  [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
5
5
 
6
+ ## 0.10.0 - 2026-09-02
7
+
8
+ ### Added
9
+
10
+ - **Local dictation.** Press a hotkey, speak, press it again, and the
11
+ transcript is on your clipboard — speech to text, entirely on-device via
12
+ [whisper.cpp](https://github.com/ggerganov/whisper.cpp)
13
+ (`pywhispercpp`). Nothing about a dictation leaves the machine unless
14
+ `[stt] cleanup` is turned on, and even then only the transcript is sent
15
+ (to `claude -p`, tools denied), never the audio. See
16
+ [docs/dictation.md](docs/dictation.md) for the full guide.
17
+ - `vocalize listen` (`--toggle`, `--cancel`, `--check`, `--list-devices`,
18
+ `--wav FILE`, `--cleanup`, `--max-seconds`) and `vocalize dictate` (an
19
+ alias for `listen --toggle`, under the name the hotkey uses).
20
+ - New Quick Action, **"Dictate with Vocalize"** — a no-input Service for
21
+ the dictation hotkey (⌃⌥⌘D suggested), installed by the existing
22
+ `hooks/install_quick_action.py` alongside the other three.
23
+ - `vocalize local install --stt [--model base.en|small.en|large-v3-turbo-q5_0]`
24
+ and `vocalize local uninstall --stt` — opt-in download-and-verify of a
25
+ whisper.cpp model, plus build-and-sign of **Vocalize Recorder**, the
26
+ small `.app` bundle that holds the microphone permission (macOS only
27
+ grants that to something with an identity). Nothing is downloaded or
28
+ compiled until you run `install --stt`, mirroring Kokoro's opt-in
29
+ install; the one-time Metal shader warm-up (~8s) is paid here, never
30
+ during a dictation.
31
+ - New `[stt]` config table — `model`, `language`, `input_device`,
32
+ `cleanup`, `paste` (reserved, not implemented yet), `max_seconds`,
33
+ `sounds` — validated on the way in the same way `[providers.*]` is, and
34
+ printed by `vocalize settings` as `stt.*` lines.
35
+ - `vocalize status` — a one-screen readiness check across every provider
36
+ in your chain, plus four dictation rows (`stt model`, `recorder`,
37
+ `microphone`, `input device`) once dictation has been set up at all.
38
+ `--json` prints the same rows as a list; exit 0 when everything is `ok`,
39
+ 1 otherwise.
40
+ - `vocalize resume [--forget]` — continue (or discard) a text-to-speech
41
+ read that a dictation interrupted. Starting a dictation stops any read
42
+ in progress, but vocalize now remembers exactly where it stopped and
43
+ offers to continue once the transcript has landed (a macOS dialog,
44
+ default Continue, 15s to answer); the record lives at
45
+ `~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour.
46
+ - New [docs/dictation.md](docs/dictation.md): install, the hotkey, every
47
+ `[stt]` key, `vocalize status`'s dictation rows, `resume`, and
48
+ troubleshooting keyed on `vocalize listen --check`'s exact messages and
49
+ exit codes.
50
+
51
+ ### Changed
52
+
53
+ - `vocalize/local/install.py` generalized to support more than one local
54
+ runtime's manifest and model files (previously hard-coded to Kokoro's).
55
+ Kokoro's own install, stamp, and `local status` output are unchanged
56
+ byte-for-byte; the whisper runtime downloads and stamps only the single
57
+ model you selected, not all three.
58
+ - `vocalize listen --check` measures the microphone permission by
59
+ launching Vocalize Recorder the same way a real dictation does
60
+ (through LaunchServices), not by exec'ing its binary directly — macOS
61
+ attributes a TCC grant to the *responsible* process, and exec'ing the
62
+ binary as a child of your shell reported the terminal's own grant
63
+ instead. Exit codes: 0 authorized-and-ready, 2 denied, 3 no usable
64
+ input device, 5 not asked yet (macOS `notDetermined`) — matching the
65
+ recorder's own contract — plus a new exit 1 meaning "vocalize's own
66
+ local install isn't finished" (not built, no model on disk, or the
67
+ recorder never reported back), which is a setup problem, not a
68
+ permission one.
69
+ - `audio.stop_playback()` gained a `remember=` flag. A dictation's stop
70
+ passes it, leaving a marker so the process that was playing can record
71
+ where it stopped — this is what makes `vocalize resume` possible. A
72
+ plain `vocalize stop` records nothing, as before.
73
+ - **A stop now silences every read already in flight**, not only the
74
+ player it kills. Playback is serialized machine-wide, so stopping one
75
+ read used to let the next queued one start speaking immediately — into
76
+ the microphone a dictation had just opened. A read *started* after the
77
+ stop is unaffected.
78
+ - Dictated text reaches the clipboard as a single line. Newlines are
79
+ collapsed there so a paste into a terminal cannot run as several
80
+ commands; `vocalize listen`'s stdout keeps them.
81
+ - `vocalize listen --check` now measures the input device configured in
82
+ `[stt] input_device` rather than the system default, and records what it
83
+ saw with a timestamp — so `vocalize status` says how old that
84
+ "authorized" verdict is instead of implying it is current.
85
+
86
+ ### Fixed
87
+
88
+ - `vocalize local install --stt`, re-run against a model that already
89
+ verified, now re-warms the runtime instead of reporting "already
90
+ installed" and stopping — a machine where only the runtime failed to
91
+ start (no Metal, a build hiccup) previously had no way to retry that
92
+ short of a full uninstall and 465 MB re-download.
93
+ - `vocalize local status` reports every installed speech-to-text model,
94
+ not just the default — installing a non-default model with `--model`
95
+ no longer looks unfinished.
96
+ - `vocalize local uninstall --stt` no longer crashes on a symlinked model
97
+ directory or recorder bundle; it reports the symlink and leaves it for
98
+ you to remove.
99
+ - `vocalize status` no longer raises on an unrecognized `VOCALIZE_CHAIN`;
100
+ like any other misconfiguration, it degrades to one failing row instead
101
+ of crashing the command. A probe that raises is reported by exception
102
+ type only — never its message, which could otherwise echo
103
+ credential-shaped text onto the screen.
104
+ - A read stopped by a dictation while a streaming provider's next chunk
105
+ was still rendering (nothing audible playing at that exact instant)
106
+ used to lose the rest of the read with no way to get it back; it's now
107
+ recorded and resumable like any other interruption. The same now holds
108
+ for a read still being synthesized (no player exists yet) and for a
109
+ plain `vocalize stop` landing in that gap.
110
+ - The first dictation on a fresh install no longer fails while macOS is
111
+ asking for the microphone. The permission dialog can sit on screen for
112
+ minutes; the press now waits for your answer and starts recording when
113
+ you click Allow, instead of giving up after five seconds and reporting
114
+ a failure that had not happened.
115
+ - `vocalize resume` continues the read in the voice, model, speed and
116
+ chunk size it was stopped in. It previously fell back to the config
117
+ defaults, which also missed the audio cache and re-synthesized (and
118
+ re-charged for) the whole remainder.
119
+ - A ten-minute dictation is no longer mistaken for a crashed one. The
120
+ claim a stop puts on a take is now aged from its own progress rather
121
+ than from when recording began, so a long take or a stop queued behind
122
+ a long read cannot be reaped mid-transcription.
123
+ - `~/.cache/vocalize` and `~/.cache/vocalize/bin` are tightened to 0700
124
+ even when they already existed. The files inside were always 0600, but
125
+ the directory listing said whether a dictation was in progress.
126
+
127
+ ## 0.9.1 - 2026-09-01
128
+
129
+ ### Fixed
130
+
131
+ - Concurrent invocations no longer talk over each other. Playback is now
132
+ serialized machine-wide on an exclusive file lock
133
+ (`~/.cache/vocalize/play.lock`): a read that arrives while another is
134
+ playing queues and starts the moment the first one ends. Only the audible
135
+ part is serialized — synthesis still runs concurrently — and the lock
136
+ dies with its process, so a killed or timed-out waiter can never leave a
137
+ stale lock behind. Chunked reads hold the slot for the whole sequence, so
138
+ pieces of two reads never interleave. On platforms without `fcntl`
139
+ (Windows), the lock is skipped and the old overlapping behavior remains.
140
+
141
+ ### Changed
142
+
143
+ - `vocalize stop` semantics with a queue: stopping kills the *current*
144
+ player; the next queued read (if any) then begins. Run `stop` again to
145
+ silence that one too.
146
+
6
147
  ## 0.9.0 - 2026-09-01
7
148
 
8
149
  ### Added
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: vocalize-cli
3
- Version: 0.9.0
3
+ Version: 0.10.0
4
4
  Summary: A CLI that turns text, markdown, or piped stdin into speech via the ElevenLabs API, with markdown-table-aware preprocessing.
5
5
  Project-URL: Homepage, https://github.com/matthager12-collab/vocalize
6
6
  Project-URL: Repository, https://github.com/matthager12-collab/vocalize
@@ -289,6 +289,12 @@ in the current directory, then the OS keychain. `vocalize auth login` sets
289
289
  up the keychain entry; `vocalize auth status` shows which of those sources
290
290
  is currently supplying the key.
291
291
 
292
+ A `[stt]` table configures dictation the same way `[providers.<name>]`
293
+ configures a TTS provider — see
294
+ [Configuration: the `[stt]` table](#configuration-the-stt-table) under
295
+ [Dictation](#dictation-speech-to-text) below for every key and its
296
+ allowlist.
297
+
292
298
  ## Providers and fallback
293
299
 
294
300
  vocalize tries providers in order — a **chain** — until one speaks. The
@@ -417,9 +423,200 @@ instead of waiting for the whole thing to render. Measured on this Mac
417
423
  rendering. `vocalize stop`, run from any terminal, halts a Kokoro read
418
424
  mid-sentence the same as any other provider.
419
425
 
426
+ ## Dictation (speech to text)
427
+
428
+ Everything above turns text into speech. Dictation runs the other way:
429
+ press a hotkey, speak, press it again, and the words land on your
430
+ clipboard. It's [whisper.cpp](https://github.com/ggerganov/whisper.cpp) via
431
+ [`pywhispercpp`](https://github.com/absadiki/pywhispercpp), entirely
432
+ on-device — nothing you say leaves the Mac unless you turn on `--cleanup`
433
+ (below), and even then only the *transcript* goes anywhere, never the audio.
434
+
435
+ ### Install
436
+
437
+ ```bash
438
+ vocalize local install --stt
439
+ ```
440
+
441
+ A separate opt-in from Kokoro's `vocalize local install` — nothing here is
442
+ downloaded or built until you run this. It:
443
+
444
+ 1. Downloads one whisper.cpp model (`small.en` by default, ~465 MB) from a
445
+ pinned Hugging Face revision, verified against a pinned sha256 before
446
+ it's kept.
447
+ 2. Compiles and ad-hoc signs a small Swift recorder bundle, **Vocalize
448
+ Recorder** — the thing that actually owns the microphone permission;
449
+ macOS won't grant that to a bare command-line tool.
450
+ 3. Warms the runtime, paying a one-time ~8-second Metal shader compile
451
+ right here, so no dictation ever stalls on it later.
452
+
453
+ The first real dictation prompts for microphone access naming **"Vocalize
454
+ Recorder"** — approve it once, like any other app's first-run permission
455
+ prompt. Pick a different model with `--model large-v3-turbo-q5_0` (~547 MB,
456
+ more accurate) or `--model base.en` (~141 MB, fastest, least accurate).
457
+
458
+ ### The hotkey
459
+
460
+ ```bash
461
+ python3 hooks/install_quick_action.py
462
+ ```
463
+
464
+ then assign a shortcut under **System Settings › Keyboard › Keyboard
465
+ Shortcuts › Services › Text › "Dictate with Vocalize"** — ⌃⌥⌘D is free by
466
+ default and a sensible pick. `vocalize dictate` is the same command from a
467
+ terminal, if you'd rather trigger it that way.
468
+
469
+ ### How a dictation works
470
+
471
+ - **Press** the hotkey — a Tink plays, the recorder starts, and any read
472
+ currently playing is stopped first (dictation and playback never
473
+ overlap; vocalize remembers where the read was cut off — see
474
+ [Continuing an interrupted read](#continuing-an-interrupted-read)).
475
+ - **Speak.**
476
+ - **Press again** — a Pop plays, recording stops, and the audio is
477
+ transcribed on-device. If anything was heard, a Glass plays and the
478
+ transcript is on your clipboard; nothing is typed for you automatically.
479
+ - **Nothing heard** (silence, or a microphone that isn't actually picking
480
+ anything up) ends the dictation quietly — no clipboard write.
481
+ - **Cancel** with a second press *within two seconds* of the first, or at
482
+ any point with `vocalize listen --cancel` — the audio is discarded, never
483
+ transcribed.
484
+ - **A third press while transcribing is refused**: a Pop, and "Still
485
+ transcribing the last dictation." Wait for the clipboard notification, or
486
+ `--cancel`, before dictating again.
487
+
488
+ ### `vocalize listen`
489
+
490
+ `vocalize dictate` (the hotkey's command) is `vocalize listen --toggle`
491
+ under another name. `listen` is the general primitive:
492
+
493
+ ```bash
494
+ vocalize listen # record until Enter/Ctrl-C, print to stdout
495
+ vocalize listen --toggle # start, or stop and copy to the clipboard
496
+ vocalize listen --cancel # discard whatever is in progress
497
+ vocalize listen --wav clip.wav # transcribe a file you already have
498
+ vocalize listen --check # microphone + install readiness
499
+ vocalize listen --list-devices # input device names for [stt] input_device
500
+ vocalize listen --max-seconds 30 # cap this one recording
501
+ ```
502
+
503
+ `--wav` is trusted input — the file has to be 16 kHz mono 16-bit WAV
504
+ (exactly what the recorder, and `say --data-format=LEI16@16000`, produce);
505
+ a malformed file gets a plain error naming the format, not a crash.
506
+ `--cleanup` tidies the transcript with Claude before it's delivered; it
507
+ applies to a live recording (`--toggle`/`dictate`, or plain `listen`) and
508
+ has no effect on `--wav`, which transcribes literally. `--max-seconds`
509
+ overrides `[stt] max_seconds` for one invocation.
510
+
511
+ ### Configuration: the `[stt]` table
512
+
513
+ ```toml
514
+ [stt]
515
+ model = "small.en" # base.en | small.en | large-v3-turbo-q5_0
516
+ language = "en" # a whisper.cpp language code
517
+ input_device = "" # "" = system default; else an exact name from --list-devices
518
+ cleanup = false # send the transcript (never audio) to Claude first
519
+ max_seconds = 120 # 1-600; the recorder self-stops here, dictate backstops it
520
+ sounds = true # the Tink/Pop/Glass feedback sounds
521
+ ```
522
+
523
+ | Key | Allowed values | Default |
524
+ |---|---|---|
525
+ | `model` | `base.en`, `small.en`, `large-v3-turbo-q5_0` | `small.en` |
526
+ | `language` | a whisper.cpp language code (`en`, `es`, `fr`, …); an `.en` model must stay `en` | `en` |
527
+ | `input_device` | `""` (system default) or an exact name from `vocalize listen --list-devices`; ≤ 128 characters, printable, can't start with `-` | `""` |
528
+ | `cleanup` | `true` / `false` | `false` |
529
+ | `paste` | reserved — not implemented in 0.10.0 | `false` |
530
+ | `max_seconds` | integer, 1–600 | `120` |
531
+ | `sounds` | `true` / `false` | `true` |
532
+
533
+ An unknown key warns on stderr; a bad value is a `ConfigError` naming it —
534
+ every one of these becomes a subprocess argument eventually, so nothing
535
+ here is trusted on the way in.
536
+
537
+ **The input-device gotcha:** a pair of Bluetooth earbuds that are *paired*
538
+ but not actually in your ears still shows up as the default input device —
539
+ and delivers digital silence. If dictation keeps saying nothing was heard,
540
+ run `vocalize listen --list-devices`, copy the real microphone's name, and
541
+ set it:
542
+
543
+ ```toml
544
+ [stt]
545
+ input_device = "MacBook Pro Microphone"
546
+ ```
547
+
548
+ ### `vocalize status`
549
+
550
+ ```bash
551
+ vocalize status
552
+ ```
553
+
554
+ prints one row per provider in your chain, plus four dictation rows once
555
+ dictation has been set up at all — an `[stt]` table in your config, a
556
+ built recorder, or a model on disk (a machine that never opted in doesn't
557
+ get four permanent red rows for a feature nobody asked for):
558
+
559
+ | Row | Reports |
560
+ |---|---|
561
+ | `stt model` | whether a whisper.cpp model is on disk |
562
+ | `recorder` | whether Vocalize Recorder is built |
563
+ | `microphone` | authorized / denied / not asked yet — from the last `listen --check`, never by launching the recorder itself |
564
+ | `input device` | whether the configured (or default) input device is actually present |
565
+
566
+ `--json` prints the same rows as a list. Exit code is 0 when every row —
567
+ providers and dictation both — is `ok`, 1 otherwise, so it composes with
568
+ `&&` in a script.
569
+
570
+ ### Continuing an interrupted read
571
+
572
+ Starting a dictation stops any read in progress, but doesn't throw the rest
573
+ of it away (DEC-003). The process that was speaking remembers exactly
574
+ where it stopped, and once your transcript has landed, vocalize asks:
575
+ **"Continue the read you interrupted?"** — answer, or let the dialog give
576
+ up after 15 seconds (counted as no). From a terminal, the same thing is:
577
+
578
+ ```bash
579
+ vocalize resume # continue where the last read left off
580
+ vocalize resume --forget # discard it instead
581
+ ```
582
+
583
+ The record — one piece of audio and the text not yet spoken — lives at
584
+ `~/.cache/vocalize/interrupted.*`, mode 0600, for at most an hour, and is
585
+ deleted the moment you resume, decline, or `--forget` it. This is the
586
+ interrupted *read*, not dictation audio or a transcript — see Privacy,
587
+ below.
588
+
589
+ ### Uninstall
590
+
591
+ ```bash
592
+ vocalize local uninstall --stt
593
+ ```
594
+
595
+ Removes the downloaded model file and the recorder bundle. The microphone
596
+ permission grant itself stays in System Settings — remove it there
597
+ (Privacy & Security › Microphone) if you want that gone too.
598
+
599
+ ### Privacy
600
+
601
+ vocalize writes no transcript to disk and shows none in a notification —
602
+ it's held in memory and goes to your clipboard (or stdout, for
603
+ `vocalize listen`) and nowhere else. Recorded audio lives only in a
604
+ private temporary directory deleted the moment a dictation ends, one way
605
+ or another (a sweep clears anything a hard kill leaves behind after 24
606
+ hours). The **only** thing that ever leaves the machine is the transcript,
607
+ and only when you turn on `--cleanup`: it's sent to `claude -p` with every
608
+ tool denied, purely to fix punctuation and casing. The audio itself is
609
+ never sent anywhere.
610
+
611
+ `--cleanup` has one consequence worth knowing before you turn it on:
612
+ Claude Code logs the prompt and stdin of every print-mode run, so the
613
+ transcript is written in plaintext to `~/.claude/projects/…`. vocalize
614
+ cannot suppress that. It is off by default. Full accounting in
615
+ [docs/dictation.md](docs/dictation.md#privacy).
616
+
420
617
  ## macOS Quick Actions (highlight → speak)
421
618
 
422
- Two Services let you use vocalize from any app without a terminal:
619
+ Four Services let you use vocalize from any app without a terminal:
423
620
 
424
621
  - **Speak with Vocalize** — highlight text anywhere, right-click →
425
622
  Services → Speak with Vocalize. It stops whatever was already playing
@@ -436,8 +633,11 @@ Two Services let you use vocalize from any app without a terminal:
436
633
  (`~/.claude/plans/`) aloud on demand. Made for the plan-approval moment:
437
634
  the proposal card is up, you press your shortcut, hear the plan, then
438
635
  accept or reject. Nothing reads unless you trigger it.
636
+ - **Dictate with Vocalize** — the dictation hotkey (see
637
+ [Dictation](#dictation-speech-to-text) above). Takes no input and shows
638
+ no window; press it, speak, press it again.
439
639
 
440
- Install both:
640
+ Install all four:
441
641
 
442
642
  ```bash
443
643
  python3 hooks/install_quick_action.py
@@ -605,14 +805,26 @@ vocalize/
605
805
  chain.py # tries each provider in the chain in turn until one speaks
606
806
  ledger.py # ~/.cache/vocalize/usage.json — local monthly budget tracking
607
807
  providers/ # elevenlabs.py, openai.py, google.py, polly.py, say.py, kokoro.py
608
- local/ # Kokoro's opt-in download/verify + the uv-run worker script
808
+ local/ # opt-in download/verify + the uv-run worker scripts —
809
+ # kokoro_manifest.py/kokoro_worker.py (TTS) and
810
+ # whisper_manifest.py/whisper_worker.py (dictation)
811
+ recorder/ # VocalizeRecorder.swift + Info.plist.in — the ad-hoc-signed
812
+ # .app bundle that owns the microphone permission
609
813
  audio.py # save to disk + play via the OS's native player
610
814
  # (afplay / mpg123 / ffplay / PowerShell, whichever exists)
815
+ dictate.py # the hotkey's toggle state machine: record, transcribe,
816
+ # clipboard — nothing here ever becomes a file or a log line
817
+ interrupted.py # the record of a read a dictation interrupted, and the
818
+ # slice `vocalize resume` plays to continue it
819
+ readiness.py # vocalize status's per-provider + per-dictation-row probes,
820
+ # each on a timed daemon thread so a wedged keychain or
821
+ # microphone check can never hang the command
611
822
  cli.py # click-based CLI wiring the above together
612
823
  hooks/
613
824
  claude_stop_hook.py # Claude Code Stop hook -> calls the vocalize CLI
614
825
  install_hook.py # safely merges the hook into ~/.claude/settings.json
615
- tests/ # pytest, all mocked — no API key needed to run these
826
+ tests/ # pytest, all mocked — no API key, network, or
827
+ # microphone needed to run these
616
828
  ```
617
829
 
618
830
  ## Testing
@@ -622,8 +834,11 @@ pip install -e ".[dev]"
622
834
  pytest
623
835
  ```
624
836
 
625
- All tests run offline: the ElevenLabs client is dependency-injected into
626
- `tts.py`, so tests pass in a fake client instead of hitting the real API.
837
+ Over 1,180 tests, all offline: the ElevenLabs client is dependency-injected
838
+ into `tts.py`, so tests pass in a fake client instead of hitting the real
839
+ API, and dictation's tests fake the recorder (a tiny shell script honoring
840
+ a stop file) and the whisper worker instead of touching a microphone or a
841
+ model.
627
842
 
628
843
  ## Known limitations
629
844
 
@@ -676,6 +891,22 @@ All tests run offline: the ElevenLabs client is dependency-injected into
676
891
  confidential information"); until you click Always Allow, every command
677
892
  that needs that key waits. Click it once per Python binary, or supply the
678
893
  key through its environment variable instead.
894
+ - **Dictation is macOS only.** It depends on `AVFoundation`, LaunchServices,
895
+ and a Swift-compiled `.app` bundle for the microphone permission — there's
896
+ no equivalent path on Linux or Windows.
897
+ - **Rebuilding the recorder means re-granting the microphone.** Vocalize
898
+ Recorder's ad-hoc code signature is what macOS ties the permission grant
899
+ to; when its Swift source changes (a vocalize upgrade that touches it),
900
+ `vocalize local install --stt` rebuilds the bundle and warns you to
901
+ re-approve it in System Settings › Privacy & Security › Microphone. An
902
+ install that doesn't change the source never re-signs, so this isn't
903
+ every upgrade — only ones that touch the recorder.
904
+ - **`small.en` mishears jargon.** The default model does fine on ordinary
905
+ speech but can mangle project-specific words (`pyproject`, a function
906
+ name) — pick `large-v3-turbo-q5_0` for better accuracy, or turn on
907
+ `[stt] cleanup` so Claude fixes obvious transcription noise before it
908
+ reaches your clipboard (it still can't guess a word it never heard
909
+ correctly).
679
910
 
680
911
  ## License
681
912