sqs-dsh-better-input 0.1.0 → 0.2.0-rc.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +49 -16
- package/lib/client/MicrophoneButton.d.ts +18 -7
- package/lib/client/VoiceRecognitionBar.d.ts +11 -2
- package/lib/client/audio-capture.d.ts +76 -0
- package/lib/client/native-speech.d.ts +158 -0
- package/lib/client/settings-controller.d.ts +21 -1
- package/lib/client/strings.d.ts +19 -2
- package/lib/client/styles.d.ts +17 -0
- package/lib/client/voice-session.d.ts +32 -0
- package/lib/client.js +1272 -226
- package/lib/client.js.map +1 -1
- package/lib/config.d.ts +136 -9
- package/lib/index.d.ts +10 -4
- package/lib/index.js +280 -10
- package/lib/index.js.map +1 -1
- package/lib/remote-contract.d.ts +55 -1
- package/lib/remote.d.ts +7 -1
- package/lib/remote.js +45 -1
- package/lib/remote.js.map +1 -1
- package/lib/settings/store.d.ts +8 -1
- package/lib/speech/service.d.ts +51 -0
- package/lib/speech/wave.d.ts +34 -0
- package/lib/typert.d.ts +3 -2
- package/lib/typert.js +123 -2
- package/lib/typert.js.map +1 -1
- package/package.json +27 -24
- package/lib/client/web-speech.d.ts +0 -35
package/README.md
CHANGED
|
@@ -22,16 +22,33 @@
|
|
|
22
22
|
|
|
23
23
|
## ⚠️ 关于这个版本(二次开发说明)
|
|
24
24
|
|
|
25
|
-
本仓库是 [`dsh-better-input`](https://github.com/DIAG5/dsh-better-input)(MIT)的**二次开发版本**,由 **sanqianshuang**
|
|
25
|
+
本仓库是 [`dsh-better-input`](https://github.com/DIAG5/dsh-better-input)(MIT)的**二次开发版本**,由 **sanqianshuang** 维护。当前版本 **`0.2.0-rc.2`** —— **与它所适配的 DSH 主线同号**,一眼即可判断该装哪一版(`0.1.0` 适配的是 DSH `0.1.7-rc.2`)。与原版的主要差别:
|
|
26
26
|
|
|
27
27
|
| 变更 | 说明 |
|
|
28
28
|
| --- | --- |
|
|
29
|
+
| 🔧 **修复「插件被 DSH 静默丢弃」(关键)** | `0.1.0` 声明的 peer 范围是 `>=0.1.7-rc.2 <0.2.0-0`,而 `0.2.0-rc.2` **不满足** `<0.2.0-0`(`0.2.0-0` 是 `0.2.0` 的最小 prerelease)。DSH 的兼容性预检因此把整个 bundle 降级丢弃——插件完全没加载,Host 服务、麦克风按钮、设置页全部消失,而 `npm run build` / `tsc` / 各构建守卫**全部照常通过**。现收紧为 `>=0.2.0-rc.2 <0.3.0-0`,并新增 `npm run check:peers` 守卫(以 DSH 自身的兼容性判定为准)防复发。 |
|
|
30
|
+
| 🎙️ **语音识别改用 DSH 自带的本地识别器** | 不再使用浏览器 Web Speech API(Chrome/Edge 上那其实是**云端识别**,音频会上行给浏览器厂商)。改为把 16 kHz 单声道 PCM16 WAV 交给 dsh 的语音服务,由本地 SenseVoice 在**本机 CPU** 上转写:**离线可用、音频不出机器、无需厂商凭据**。语种随之收窄为 自动 / 中文 / 英语 / 粤语 / 日语 / 韩语。详见 [CHANGELOG](./CHANGELOG.md) `[0.2.0-rc.2]`。 |
|
|
31
|
+
| 🔀 **保留「边说边出字」** | dsh 的语音服务只有整段转写(没有流式接口),所以流式由本插件自己做:录制时按静音优先、默认每 3 秒切段转写并流入输入框;**停止后用整段录音重转一次作为最终稿**再润色 —— 分段只影响预览观感,不会污染最终文本。 |
|
|
32
|
+
| ⏱️ **识别条重做:贴齐输入框 + 计时 + 静音自动停止** | 原识别条是 `conversation.input.dock` 槽位里唯一没有自己做宽度约束的条目,会被 flex 父级拉满整列,**比输入框宽出一大截**;现在按 dsh 自己的 dock 约定(`--dsh-composer-card-max-width` / `--dsh-composer-side-clearance`)居中贴齐输入框。同时条上新增**已监听时长**(`mm:ss`)、**实时音量表**与**静音倒计时**:说话后静音满 10 秒(可调、可关)自动停止并转写。 |
|
|
29
33
|
| 🗑️ **移除「文件输入 / 文件转 Markdown / OCR」整块** | DSH `0.1.7` 已原生支持文件拖拽上传、上传进度、文件卡片与重试(`conversation.input.attachments` / `DropOverlay` / `FileCard`),插件再挂一套 📎 文件面板与 `@` 引用芯片属于重复实现,故整体删除。 |
|
|
30
|
-
| 🔧 **修复 Typert
|
|
31
|
-
| 🔧 **修复设置存储被移除的 API** | `0.1.7` 删除了 `settings.register(namespace, schema)
|
|
32
|
-
| 🔧 **对齐依赖版本** | `@deepseek-ai/cordis` → `~4.0.4`、`@deepseek-ai/schemastery` → `~3.18.4
|
|
34
|
+
| 🔧 **修复 Typert 兼容性** | `0.1.7` 收紧了 Typert 边界:`mode: 'strict'` 的 codec 必须带 **`create()` 工厂**;旧版把 schema 直接放在 `schema` 属性上,会被 `requireStrictCodec` 拒绝(报 `strict codec has no create() factory`),导致整个 entry 无法激活、插件根本不加载。现已全部改为工厂形式。 |
|
|
35
|
+
| 🔧 **修复设置存储被移除的 API** | `0.1.7` 删除了 `settings.register(namespace, schema)`(`0.2.0` 仍无此 API),改为按 profile entry 的 `SettingsForms`(带 revision CAS)。本插件的设置改为**自持 JSON 文件**,不再依赖该 API,因而不会被后续版本反复打破。 |
|
|
36
|
+
| 🔧 **对齐依赖版本** | `@deepseek-ai/cordis` → `~4.0.4`、`@deepseek-ai/schemastery` → `~3.18.4`;全部 `@deepseek-ai/dsh-*` 类型包钉到 `0.2.0-rc.2`,peer 范围 `>=0.2.0-rc.2 <0.3.0-0`。 |
|
|
33
37
|
| 🏷️ **身份与命名** | 包名 `sqs-dsh-better-input`(npm 同名)、作者 `sanqianshuang`;设置与模板数据目录为 `~/.dsh/sqs-dsh-better-input/`(与原版互不干扰,可并存安装)。 |
|
|
34
38
|
|
|
39
|
+
> **从 `0.1.0` 升级**:设置项、数据目录与用户可见行为**向下兼容**,不需要迁移任何数据;本版新增的设置键(`streamingPreview` / `segmentSeconds` / `autoStopSeconds`)在读取旧 `settings.json` 时自动补齐,越界值读取即修复。
|
|
40
|
+
|
|
41
|
+
> **与 DSH 自带的语音输入的关系(重要)**:`0.2.0` 增加了一个**可选且默认关闭**的 bundle
|
|
42
|
+
> `@deepseek-ai/dsh-experimental-voice-input-bundle`,它带来 DSH 自己的麦克风
|
|
43
|
+
> (`conversation.input.activity` 槽位,位于模型选择器与发送按钮之间)与**本地 SenseVoice 识别器**。
|
|
44
|
+
>
|
|
45
|
+
> 本插件的麦克风在 `conversation.input.right`(模型选择器左侧),两者**槽位不冲突、可同时开启**;
|
|
46
|
+
> 但**功能上重叠**:本版起本插件也用同一个本地识别器。区别在于本插件额外提供
|
|
47
|
+
> **分段流式上屏**、识别后的 **AI 润色**、**提示词优化**与**模板库** —— DSH 自带的那个转写完就直接进草稿,没有任何后处理。
|
|
48
|
+
>
|
|
49
|
+
> 因此:**本插件必须依赖该 bundle 提供的识别服务**(在插件管理页启用「语音输入」并完成约 239 MB 的模型准备)。
|
|
50
|
+
> 未启用时本插件其余功能完全正常,只是麦克风会提示"本地识别不可用"。
|
|
51
|
+
|
|
35
52
|
> 原版的 LICENSE 与版权声明已在 [LICENSE](./LICENSE) 中保留,符合 MIT 要求。
|
|
36
53
|
|
|
37
54
|
## ✨ 功能
|
|
@@ -40,7 +57,7 @@
|
|
|
40
57
|
<tr><th align="center" width="120">模块</th><th align="left">说明</th></tr>
|
|
41
58
|
<tr>
|
|
42
59
|
<td align="center">🎙️<br/><b>语音输入</b></td>
|
|
43
|
-
<td>点击麦克风,边说边转写,文字<strong
|
|
60
|
+
<td>点击麦克风,边说边转写,文字<strong>分段流式</strong>进入输入框。识别由 <strong>dsh 自带的本地识别器(SenseVoice)</strong> 在本机 CPU 完成 —— <strong>离线可用、音频不出机器、无需 API Key</strong>。停止后再用整段录音重转一次作为最终稿,然后才进入 AI 润色。</td>
|
|
44
61
|
</tr>
|
|
45
62
|
<tr>
|
|
46
63
|
<td align="center">🤖<br/><b>AI 润色</b></td>
|
|
@@ -59,8 +76,8 @@
|
|
|
59
76
|
<td>润色 / 优化进行中你手动改了草稿,结果<strong>不会覆盖</strong>你的编辑;失败保留原文。</td>
|
|
60
77
|
</tr>
|
|
61
78
|
<tr>
|
|
62
|
-
<td align="center">⏱️<br/><b
|
|
63
|
-
<td
|
|
79
|
+
<td align="center">⏱️<br/><b>计时与自动停止</b></td>
|
|
80
|
+
<td>输入框上方的识别条实时显示<strong>已监听时长</strong>与音量表;说话后一旦静音,条上出现 <strong>倒计时</strong>(默认 10 秒),归零即自动停止、转写、润色。单次录音另有 1–120 秒硬上限兜底(可自定义)。两者都在设置页可调。</td>
|
|
64
81
|
</tr>
|
|
65
82
|
<tr>
|
|
66
83
|
<td align="center">⚙️<br/><b>可视化设置页</b></td>
|
|
@@ -74,7 +91,7 @@
|
|
|
74
91
|
|
|
75
92
|
## 🚀 安装
|
|
76
93
|
|
|
77
|
-
前置:[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(`0.
|
|
94
|
+
前置:[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(`0.2.0-rc.2`)+ Node.js `^22.19.0 || >=24.0.0`,并在 profile 中启用 DSH 的语音 bundle(插件管理页的「语音输入」,首次使用需准备约 239 MB 模型)。浏览器只需支持 `getUserMedia` 与 Web Audio,不再需要 Chromium 专有的 Web Speech API。
|
|
78
95
|
|
|
79
96
|
### 方式 A:从 npm 安装(推荐)
|
|
80
97
|
|
|
@@ -115,7 +132,7 @@ npm run build
|
|
|
115
132
|
dsh plugin --profile web add "$PWD"
|
|
116
133
|
```
|
|
117
134
|
|
|
118
|
-
> `--legacy-peer-deps`:`0.
|
|
135
|
+
> `--legacy-peer-deps`:`0.2.0-rc.2` 的这批类型包(dev 依赖)之间存在 npm 难以自解的 peer 链(如 `dsh-api-remotes` → `dsh-scope`)。它们只用于编译期类型、不会被打进产物(客户端 bundle 已把 `@deepseek-ai/*` 全部外置),因此放宽 peer 校验是安全的。
|
|
119
136
|
|
|
120
137
|
### 备选:不装依赖,写进 preset 的 `cordis.yml`
|
|
121
138
|
|
|
@@ -133,10 +150,16 @@ dsh plugin --profile web add "$PWD"
|
|
|
133
150
|
|
|
134
151
|
1. 打开任意对话,点击输入框右侧的**麦克风按钮**
|
|
135
152
|
2. 开始说话,识别文字**实时流入**输入框
|
|
153
|
+
|
|
154
|
+
> 输入框上方会出现识别条:`正在聆听… 00:12` 是**已监听时长**,后面的波形是你的**实时音量**(能一眼看出麦克风是否收到声音)。
|
|
155
|
+
> 说完停下后,条上出现倒计时 `10 秒后自动停止`(设置里可改,也可设为 0 关闭);倒计时归零就自动停止并开始转写,不用再点一次按钮。你随时可以手动点**停止语音输入**提前结束。
|
|
156
|
+
|
|
136
157
|
3. 再点按钮(或识别条上的**停止**)结束
|
|
137
158
|
4. 检查、修改、发送
|
|
138
159
|
|
|
139
|
-
>
|
|
160
|
+
> 识别在**你的机器上**完成:浏览器只负责录音,转写由 dsh 的语音服务(本地 SenseVoice)完成,音频不上行到任何云端。
|
|
161
|
+
> 首次录音会唤醒识别进程(约 1 秒),之后更快。
|
|
162
|
+
> 若提示"本地识别不可用",去 设置 → **BetterInput** → 「本地识别器」查看状态并点「下载并准备模型」。
|
|
140
163
|
|
|
141
164
|
### 2. AI 润色
|
|
142
165
|
|
|
@@ -167,8 +190,12 @@ dsh plugin --profile web add "$PWD"
|
|
|
167
190
|
| 设置项 | 说明 |
|
|
168
191
|
| --- | --- |
|
|
169
192
|
| 界面语言 | 插件界面文案支持中文 / 英文,跟随 DSH 界面语言切换,即时生效 |
|
|
170
|
-
|
|
|
171
|
-
|
|
|
193
|
+
| 本地识别器 | 显示 dsh 语音服务的识别器与就绪状态;未准备时提供「下载并准备模型」 |
|
|
194
|
+
| 识别语言 | 自动检测 / 中文 / English / 粤语 / 日本語 / 한국어(旧版填写的 `zh-CN` 这类值会自动映射) |
|
|
195
|
+
| 单次录音上限 | 1–120 秒,默认 120,到点自动停止 |
|
|
196
|
+
| 静音自动停止 | 0 = 关闭;3–120 秒,默认 **10**。说话后静音满这么久即自动停止并转写(倒计时显示在识别条上) |
|
|
197
|
+
| 边录边出字(分段预览) | 开(默认)/ 关。关闭后只在停止时转写一次 |
|
|
198
|
+
| 分段长度 | 1–10 秒,默认 3。优先在静音处切分;越短上屏越快、越容易切断词 |
|
|
172
199
|
| AI 润色 | 开 / 关;开启后每次语音识别结束自动润色进草稿 |
|
|
173
200
|
| 润色模型 / 思考强度 / 自定义提示词 | 各自独立配置 |
|
|
174
201
|
| 优化模型 / 思考强度 / 自定义提示词 | 各自独立配置 |
|
|
@@ -180,9 +207,13 @@ dsh plugin --profile web add "$PWD"
|
|
|
180
207
|
|
|
181
208
|
## 🧩 兼容性
|
|
182
209
|
|
|
183
|
-
- DeepSeek Harness `0.
|
|
210
|
+
- DeepSeek Harness `0.2.0-rc.2`(peer 范围 `>=0.2.0-rc.2 <0.3.0-0`)
|
|
184
211
|
- Node.js `^22.19.0 || >=24.0.0`
|
|
185
|
-
-
|
|
212
|
+
- 任意支持 `getUserMedia` + Web Audio 的现代浏览器(Chrome / Edge / Firefox;不再需要 Chromium 专有的 Web Speech API)
|
|
213
|
+
- profile 中启用 DSH 语音 bundle,并完成本地识别模型准备(约 239 MB)
|
|
214
|
+
|
|
215
|
+
> peer 范围必须容纳正在运行的 dsh 版本,否则 DSH 会**静默丢弃整个插件**(详见
|
|
216
|
+
> [AGENTS.md](./AGENTS.md) 硬规则 1)。`npm run check:peers` 以 DSH 自身的判定逻辑校验这一点。
|
|
186
217
|
|
|
187
218
|
## 🛠️ 开发
|
|
188
219
|
|
|
@@ -190,6 +221,7 @@ dsh plugin --profile web add "$PWD"
|
|
|
190
221
|
npm install --legacy-peer-deps
|
|
191
222
|
npm run check # 类型检查
|
|
192
223
|
npm run build # 构建 lib/(Host ESM + 浏览器 bundle)
|
|
224
|
+
npm run verify # 构建 + 三条构建后守卫(bundle / typert / peers)—— 提交前请跑
|
|
193
225
|
```
|
|
194
226
|
|
|
195
227
|
改 Client 端:`npm run dev:watch` 后刷新 UI;改 Host 端:重启 `dsh web`。
|
|
@@ -200,10 +232,11 @@ npm run build # 构建 lib/(Host ESM + 浏览器 bundle)
|
|
|
200
232
|
- `src/identity.ts` — 包名 / 作者 / 仓库等身份信息的唯一来源
|
|
201
233
|
- `src/polish/service.ts` — `BetterInputPolishService`(Typert remote):设置读写、dsh 模型路由发现、LLM 润色与提示词优化、提示词模板存取
|
|
202
234
|
- `src/polish/prompts.ts` — 内置润色 / 优化提示词与输出契约守卫
|
|
203
|
-
- `src/settings/store.ts` — 插件自持设置的 JSON
|
|
235
|
+
- `src/settings/store.ts` — 插件自持设置的 JSON 存储(原子写入、损坏自愈、越界值读取即修复)
|
|
236
|
+
- `src/speech/` — Host 端语音入口:`service.ts`(`BetterInputSpeech`:状态 / 准备模型 / 转写)与 `wave.ts`(逐字段镜像 dsh 的 WAV 校验)
|
|
204
237
|
- `src/templates/` — 提示词模板的数据模型与宿主端 JSON 存储
|
|
205
238
|
- `src/about.ts` — 插件身份读取与 npm 版本检查
|
|
206
|
-
- `src/client/` — 浏览器端:麦克风 / 优化按钮(`conversation.input.right`)、识别条(`conversation.input.dock`)、设置页与模板管理(`settings.section`)、`/`
|
|
239
|
+
- `src/client/` — 浏览器端:麦克风 / 优化按钮(`conversation.input.right`)、识别条(`conversation.input.dock`)、设置页与模板管理(`settings.section`)、`/` 模板候选源;`audio-capture.ts`(采集 → 16 kHz WAV)、`native-speech.ts`(分段预览 + 终稿 + `SilenceWatch` 静音倒计时)、`styles.ts`(识别条样式,按 dsh 的 dock 宽度约定贴齐输入框)
|
|
207
240
|
- `src/typert.ts` / `src/remote.ts` — Client↔Host 类型化通信契约
|
|
208
241
|
|
|
209
242
|
## 📄 License
|
|
@@ -25,10 +25,15 @@ export type SettingsFace = {
|
|
|
25
25
|
readonly settings: BetterInputSettings;
|
|
26
26
|
};
|
|
27
27
|
/**
|
|
28
|
-
* The microphone button in the composer tool row. Click to start
|
|
29
|
-
* click again to stop.
|
|
30
|
-
*
|
|
31
|
-
*
|
|
28
|
+
* The microphone button in the composer tool row. Click to start recording,
|
|
29
|
+
* click again to stop.
|
|
30
|
+
*
|
|
31
|
+
* The recording is transcribed by dsh's own local recognizer
|
|
32
|
+
* (`remote.transcribeSpeech` → `ctx.speechToText`, the SenseVoice provider from
|
|
33
|
+
* the optional voice-input bundle) instead of the browser's Web Speech API.
|
|
34
|
+
* While recording, the text of each finished segment streams into the draft;
|
|
35
|
+
* when the user stops, one pass over the whole recording produces the
|
|
36
|
+
* authoritative transcript and AI polishing runs on that.
|
|
32
37
|
*/
|
|
33
38
|
export declare function MicrophoneButton({ useInput, inputActions, voiceSession, remote, useSettings, t }: InputZoneLikeProps): import("react").JSX.Element;
|
|
34
39
|
export interface PolishDraftOptions {
|
|
@@ -53,9 +58,15 @@ export interface PolishDraftOptions {
|
|
|
53
58
|
}
|
|
54
59
|
export declare function polishDraft(options: PolishDraftOptions): Promise<void>;
|
|
55
60
|
/**
|
|
56
|
-
* Only replace the draft when the user has not edited it since
|
|
57
|
-
*
|
|
58
|
-
*
|
|
61
|
+
* Only replace the draft when the user has not edited it since our own last
|
|
62
|
+
* write. Both the text we wrote last (`draftAtStop` — the finished transcript,
|
|
63
|
+
* or the last streamed preview segment) and the untouched base draft count as
|
|
64
|
+
* unchanged.
|
|
65
|
+
*
|
|
66
|
+
* This matters more since the preview became segmented: the text on screen
|
|
67
|
+
* while recording is a *preview*, and the authoritative transcript that arrives
|
|
68
|
+
* after stopping can differ from it. Comparing against the text this session
|
|
69
|
+
* wrote is what keeps a user's mid-recording edit from being overwritten.
|
|
59
70
|
*/
|
|
60
71
|
export declare function shouldApplyPolishResult(currentDraft: string, draftAtStop: string, baseDraft: string): boolean;
|
|
61
72
|
/** Append transcript to a base draft with one space separator. */
|
|
@@ -7,8 +7,17 @@ export type RecognitionBarProps = {
|
|
|
7
7
|
readonly t: Translate;
|
|
8
8
|
};
|
|
9
9
|
/**
|
|
10
|
-
* The recognition status bar above the composer.
|
|
11
|
-
*
|
|
10
|
+
* The recognition status bar above the composer.
|
|
11
|
+
*
|
|
12
|
+
* Shows the live state, a stopwatch for the current listening session, a level
|
|
13
|
+
* meter fed by the capture's own RMS, and — once the user has spoken and gone
|
|
14
|
+
* quiet — the countdown to the silence auto-stop. Renders nothing when idle.
|
|
15
|
+
*
|
|
16
|
+
* Timer and level come from `session.meter`, a store that only this component
|
|
17
|
+
* subscribes to: they change five times a second, which would otherwise
|
|
18
|
+
* re-render the microphone button (and every other session subscriber) too.
|
|
12
19
|
*/
|
|
13
20
|
export declare function VoiceRecognitionBar({ voiceSession, t }: RecognitionBarProps): import("react").JSX.Element | null;
|
|
21
|
+
/** `mm:ss`, the stopwatch shown next to the listening label. */
|
|
22
|
+
export declare function formatClock(seconds: number): string;
|
|
14
23
|
export {};
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Browser microphone capture for the native dsh recognizers.
|
|
3
|
+
*
|
|
4
|
+
* dsh's speech service accepts exactly one audio format: a canonical 16 kHz
|
|
5
|
+
* mono PCM16 WAV (`@deepseek-ai/dsh-experimental-speech-to-text/wave`). This
|
|
6
|
+
* module owns that contract on the browser side — capture, resample, encode —
|
|
7
|
+
* and keeps the recorded samples so a finished recording can still be sliced
|
|
8
|
+
* (preview segments) or re-encoded (the final authoritative pass) after the
|
|
9
|
+
* microphone itself has been released.
|
|
10
|
+
*
|
|
11
|
+
* The previous implementation used the Web Speech API, which streamed results
|
|
12
|
+
* out of the browser and never owned the audio. Recognising locally means we do
|
|
13
|
+
* own the audio, hence the explicit format handling here.
|
|
14
|
+
*/
|
|
15
|
+
/** Why capture could not start, for localized messaging by the caller. */
|
|
16
|
+
export type CaptureFailureKind = 'unavailable' | 'permission' | 'interrupted';
|
|
17
|
+
/** Capture failure whose `kind` the caller maps onto a localized string. */
|
|
18
|
+
export declare class CaptureError extends Error {
|
|
19
|
+
readonly kind: CaptureFailureKind;
|
|
20
|
+
constructor(kind: CaptureFailureKind, message?: string);
|
|
21
|
+
}
|
|
22
|
+
/**
|
|
23
|
+
* One microphone acquisition.
|
|
24
|
+
*
|
|
25
|
+
* Samples are retained at the rate the AudioContext actually runs at and are
|
|
26
|
+
* converted to 16 kHz on demand, so a browser that ignores the requested
|
|
27
|
+
* sample rate still produces a valid recording.
|
|
28
|
+
*/
|
|
29
|
+
export declare class MicrophoneCapture {
|
|
30
|
+
private stream;
|
|
31
|
+
private context;
|
|
32
|
+
private source;
|
|
33
|
+
private processor;
|
|
34
|
+
private sink;
|
|
35
|
+
private samples;
|
|
36
|
+
private frames;
|
|
37
|
+
private captureRate;
|
|
38
|
+
private released;
|
|
39
|
+
private disposed;
|
|
40
|
+
/** Acquire the microphone and start buffering samples. */
|
|
41
|
+
start(): Promise<void>;
|
|
42
|
+
/** Seconds captured so far, in capture-rate terms. */
|
|
43
|
+
secondsRecorded(): number;
|
|
44
|
+
/** Root-mean-square level of the trailing window, 0 when idle. */
|
|
45
|
+
level(windowSeconds?: number): number;
|
|
46
|
+
/** Largest absolute sample in a range, used to skip silent segments. */
|
|
47
|
+
peak(fromSeconds: number, toSeconds: number): number;
|
|
48
|
+
/** Copy a capture-time range as 16 kHz mono samples. */
|
|
49
|
+
slice(fromSeconds: number, toSeconds: number): Float32Array;
|
|
50
|
+
/**
|
|
51
|
+
* Release the microphone and the audio graph, keeping the samples.
|
|
52
|
+
*
|
|
53
|
+
* Called as soon as the user stops speaking so the browser indicator goes
|
|
54
|
+
* away while the final transcription is still running.
|
|
55
|
+
*/
|
|
56
|
+
release(): Promise<void>;
|
|
57
|
+
/** Release the microphone and drop the buffered samples. */
|
|
58
|
+
dispose(): Promise<void>;
|
|
59
|
+
private collect;
|
|
60
|
+
private ensure;
|
|
61
|
+
private range;
|
|
62
|
+
}
|
|
63
|
+
/** Whether this browser can capture at all (secure context + Web Audio). */
|
|
64
|
+
export declare function isCaptureSupported(): boolean;
|
|
65
|
+
/**
|
|
66
|
+
* Encode 16 kHz mono samples as the canonical PCM16 WAV the Host validates.
|
|
67
|
+
*
|
|
68
|
+
* Field for field: RIFF/WAVE/`fmt `/16/PCM=1/mono/16000/byteRate 32000/
|
|
69
|
+
* blockAlign 2/bits 16/`data`, RIFF size = length - 8, data size =
|
|
70
|
+
* length - 44, an even PCM length, and at least one sample.
|
|
71
|
+
*/
|
|
72
|
+
export declare function encodeWave(samples: Float32Array): Uint8Array;
|
|
73
|
+
/** Base64 for the JSON Remote carrier, chunked to avoid argument-count limits. */
|
|
74
|
+
export declare function audioBase64(bytes: Uint8Array): string;
|
|
75
|
+
/** Linear resampling; adequate for a fallback path off the 16 kHz happy path. */
|
|
76
|
+
export declare function resampleLinear(samples: Float32Array, fromRate: number, toRate: number): Float32Array;
|
|
@@ -0,0 +1,158 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Native speech session: dsh's local recognizers with segment-level streaming.
|
|
3
|
+
*
|
|
4
|
+
* dsh's speech service transcribes one complete recording — its own README says
|
|
5
|
+
* streaming recognition has no service method — so "text appears while you
|
|
6
|
+
* speak" is emulated here: the recording is cut into segments (preferably at
|
|
7
|
+
* silence boundaries), each segment is transcribed over the plugin's Remote and
|
|
8
|
+
* its text is appended to the draft, and when the user stops, the whole
|
|
9
|
+
* recording is transcribed once more to produce an authoritative transcript
|
|
10
|
+
* before polishing runs on it.
|
|
11
|
+
*
|
|
12
|
+
* The two-pass design is what keeps the streamed preview cheap: a mis-split
|
|
13
|
+
* segment can only make the *preview* worse, never the finished text.
|
|
14
|
+
*/
|
|
15
|
+
import { type BetterInputSettings } from '../config.js';
|
|
16
|
+
import { CaptureError } from './audio-capture.js';
|
|
17
|
+
/** How often the segment cutter inspects the recording. */
|
|
18
|
+
export declare const TICK_MS = 200;
|
|
19
|
+
export type NativeSpeechPhase = 'capturing' | 'finalizing';
|
|
20
|
+
/** What one {@link SilenceWatch} frame reported. */
|
|
21
|
+
export type SilenceReading = {
|
|
22
|
+
/** True once any frame has been loud enough — the countdown arms on this. */
|
|
23
|
+
readonly hasSpoken: boolean;
|
|
24
|
+
/** Seconds of continuous quiet since the last sound. */
|
|
25
|
+
readonly quietSeconds: number;
|
|
26
|
+
/** Seconds left before the auto-stop, `null` while no countdown runs. */
|
|
27
|
+
readonly remainingSeconds: number | null;
|
|
28
|
+
/** Whether the quiet window has fully elapsed. */
|
|
29
|
+
readonly elapsed: boolean;
|
|
30
|
+
};
|
|
31
|
+
/**
|
|
32
|
+
* The silence auto-stop timer: "speak, then stop talking, and the recording
|
|
33
|
+
* finishes by itself".
|
|
34
|
+
*
|
|
35
|
+
* Deliberately pure and free of Web Audio so the guard script can drive it:
|
|
36
|
+
* the whole user-visible contract is here — the countdown never runs before the
|
|
37
|
+
* first sound (`hasSpoken`), any new sound resets it to the full window, and it
|
|
38
|
+
* cannot fire at all when the window is `0` (disabled).
|
|
39
|
+
*/
|
|
40
|
+
export declare class SilenceWatch {
|
|
41
|
+
/** Quiet window in seconds; `0` disables the auto-stop. */
|
|
42
|
+
private readonly limitSeconds;
|
|
43
|
+
/** Frame length in seconds. */
|
|
44
|
+
private readonly tickSeconds;
|
|
45
|
+
private spoken;
|
|
46
|
+
private quiet;
|
|
47
|
+
constructor(
|
|
48
|
+
/** Quiet window in seconds; `0` disables the auto-stop. */
|
|
49
|
+
limitSeconds: number,
|
|
50
|
+
/** Frame length in seconds. */
|
|
51
|
+
tickSeconds: number);
|
|
52
|
+
observe(sounding: boolean): SilenceReading;
|
|
53
|
+
/** Seconds left in the window, `null` while counting is not allowed. */
|
|
54
|
+
remainingSeconds(): number | null;
|
|
55
|
+
}
|
|
56
|
+
/**
|
|
57
|
+
* One 200 ms telemetry frame for the status bar: how long the microphone has
|
|
58
|
+
* been open, the current input level (for the level meter) and — once speech
|
|
59
|
+
* has been heard — how long is left before the silence auto-stop.
|
|
60
|
+
*/
|
|
61
|
+
export type SpeechTick = {
|
|
62
|
+
readonly elapsedSeconds: number;
|
|
63
|
+
/** Root-mean-square of the trailing window, ~0 while silent. */
|
|
64
|
+
readonly level: number;
|
|
65
|
+
/** Configured silence window in seconds, `0` when the auto-stop is off. */
|
|
66
|
+
readonly autoStopSeconds: number;
|
|
67
|
+
/** Seconds left before the auto-stop, `null` when no countdown is running. */
|
|
68
|
+
readonly autoStopRemainingSeconds: number | null;
|
|
69
|
+
};
|
|
70
|
+
/** The shape this session needs from `remote.transcribeSpeech`. */
|
|
71
|
+
export type TranscribeOutcome = {
|
|
72
|
+
readonly ok: true;
|
|
73
|
+
readonly value: {
|
|
74
|
+
readonly text: string;
|
|
75
|
+
};
|
|
76
|
+
} | {
|
|
77
|
+
readonly ok: false;
|
|
78
|
+
readonly error: {
|
|
79
|
+
readonly message: string;
|
|
80
|
+
};
|
|
81
|
+
};
|
|
82
|
+
export type TranscribeCall = (audioBase64: string, language: string, signal: AbortSignal) => Promise<TranscribeOutcome>;
|
|
83
|
+
export type NativeSpeechSessionOptions = {
|
|
84
|
+
/** SenseVoice language hint; empty means automatic detection. */
|
|
85
|
+
language: string;
|
|
86
|
+
/** Cut a preview segment once this many seconds have accumulated. */
|
|
87
|
+
segmentSeconds: number;
|
|
88
|
+
/** Stream segment transcripts into the draft while the user is still speaking. */
|
|
89
|
+
streamingPreview: boolean;
|
|
90
|
+
/**
|
|
91
|
+
* Stop after this many seconds without speech once the user has spoken.
|
|
92
|
+
* `0` disables it; the hard recording ceiling still applies either way.
|
|
93
|
+
*/
|
|
94
|
+
autoStopSeconds: number;
|
|
95
|
+
/** Plugin Remote call that performs one transcription on the Host. */
|
|
96
|
+
transcribe: TranscribeCall;
|
|
97
|
+
/** Growing preview text, called after every completed segment. */
|
|
98
|
+
onPreview: (text: string) => void;
|
|
99
|
+
/** Authoritative transcript after the recording stopped. */
|
|
100
|
+
onEnd: (text: string) => void;
|
|
101
|
+
/** Capture and transcription failures; the session decides whether to continue. */
|
|
102
|
+
onError: (error: Error) => void;
|
|
103
|
+
/** Phase changes, for the surrounding status UI. */
|
|
104
|
+
onPhase?: (phase: NativeSpeechPhase) => void;
|
|
105
|
+
/** Timer/level telemetry, every {@link TICK_MS}, for the status bar. */
|
|
106
|
+
onTick?: (tick: SpeechTick) => void;
|
|
107
|
+
/** Called just before the session stops itself on the silence timeout. */
|
|
108
|
+
onAutoStop?: () => void;
|
|
109
|
+
};
|
|
110
|
+
/** Settings face this session reads; kept structural so tests need no defaults. */
|
|
111
|
+
export type SpeechSessionSettings = Pick<BetterInputSettings, 'language' | 'segmentSeconds' | 'streamingPreview' | 'autoStopSeconds'>;
|
|
112
|
+
export declare function sessionOptionsFor(settings: SpeechSessionSettings, transcribe: TranscribeCall, callbacks: Pick<NativeSpeechSessionOptions, 'onPreview' | 'onEnd' | 'onError' | 'onPhase' | 'onTick' | 'onAutoStop'>): NativeSpeechSessionOptions;
|
|
113
|
+
export declare class NativeSpeechSession {
|
|
114
|
+
private readonly options;
|
|
115
|
+
private readonly capture;
|
|
116
|
+
private readonly segments;
|
|
117
|
+
private timer;
|
|
118
|
+
/** Seconds already submitted as preview segments. */
|
|
119
|
+
private cursor;
|
|
120
|
+
private speechSeen;
|
|
121
|
+
/** Silence since the last cut, used by the segment cutter. */
|
|
122
|
+
private silenceSeconds;
|
|
123
|
+
/**
|
|
124
|
+
* The auto-stop timer, replaced on every `start()`.
|
|
125
|
+
*
|
|
126
|
+
* Kept separate from `silenceSeconds` on purpose: cutting a segment resets
|
|
127
|
+
* that one, and the countdown the user is watching must not jump back to the
|
|
128
|
+
* full window every time a preview segment is cut.
|
|
129
|
+
*/
|
|
130
|
+
private watch;
|
|
131
|
+
/** Serialises transcriptions: the native worker is serial anyway. */
|
|
132
|
+
private queue;
|
|
133
|
+
private inFlight;
|
|
134
|
+
private stopping;
|
|
135
|
+
private aborted;
|
|
136
|
+
private ended;
|
|
137
|
+
constructor(options: NativeSpeechSessionOptions);
|
|
138
|
+
get active(): boolean;
|
|
139
|
+
/** Acquire the microphone and begin cutting preview segments. */
|
|
140
|
+
start(): Promise<void>;
|
|
141
|
+
/** Finish the recording: flush the tail, then transcribe the whole thing. */
|
|
142
|
+
stop(): void;
|
|
143
|
+
/** Discard everything: no callbacks, no transcript, microphone released. */
|
|
144
|
+
abort(): void;
|
|
145
|
+
private tick;
|
|
146
|
+
/** One telemetry frame for the status bar (timer, level, countdown). */
|
|
147
|
+
private publishTick;
|
|
148
|
+
/** Submit `[cursor, endSeconds)` as one preview segment. */
|
|
149
|
+
private cutSegment;
|
|
150
|
+
private submit;
|
|
151
|
+
private finish;
|
|
152
|
+
private clearTimer;
|
|
153
|
+
}
|
|
154
|
+
/** Localized by the caller: capture failures carry a kind, others their message. */
|
|
155
|
+
export declare function captureFailureMessage(error: Error): {
|
|
156
|
+
kind: CaptureError['kind'] | '';
|
|
157
|
+
message: string;
|
|
158
|
+
};
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import { type BetterInputSettingsPatch, type BetterInputSettingsView, type PolishRoute, type ReasoningEffortInfo } from '../config.js';
|
|
2
|
-
import type { AboutInfoWire, UpdateCheckResultWire } from '../remote-contract.js';
|
|
2
|
+
import type { AboutInfoWire, SpeechStatusWire, UpdateCheckResultWire } from '../remote-contract.js';
|
|
3
3
|
import type { BetterInputRemote } from '../remote.js';
|
|
4
4
|
export type SettingsStatus = 'loading' | 'ready' | 'error';
|
|
5
5
|
export type SettingsSnapshot = {
|
|
@@ -31,6 +31,13 @@ export type UpdateSnapshot = {
|
|
|
31
31
|
readonly update: UpdateCheckResultWire | null;
|
|
32
32
|
readonly detail: string;
|
|
33
33
|
};
|
|
34
|
+
export type SpeechSnapshot = {
|
|
35
|
+
readonly status: 'loading' | 'ready' | 'error';
|
|
36
|
+
readonly view: SpeechStatusWire;
|
|
37
|
+
/** True while a prepare request is in flight, so the button can be disabled. */
|
|
38
|
+
readonly preparing: boolean;
|
|
39
|
+
readonly detail: string;
|
|
40
|
+
};
|
|
34
41
|
type Listener = () => void;
|
|
35
42
|
/**
|
|
36
43
|
* Settings read/write controller for the settings page and the microphone
|
|
@@ -46,6 +53,7 @@ export declare class SettingsController {
|
|
|
46
53
|
private effortsSnapshot;
|
|
47
54
|
private aboutSnapshot;
|
|
48
55
|
private updateSnapshot;
|
|
56
|
+
private speechSnapshot;
|
|
49
57
|
private readonly listeners;
|
|
50
58
|
private disposed;
|
|
51
59
|
constructor(remote: BetterInputRemote);
|
|
@@ -54,8 +62,19 @@ export declare class SettingsController {
|
|
|
54
62
|
readonly getEffortsSnapshot: () => EffortsSnapshot;
|
|
55
63
|
readonly getAboutSnapshot: () => AboutSnapshot;
|
|
56
64
|
readonly getUpdateSnapshot: () => UpdateSnapshot;
|
|
65
|
+
readonly getSpeechSnapshot: () => SpeechSnapshot;
|
|
57
66
|
readonly subscribe: (listener: Listener) => (() => void);
|
|
58
67
|
refreshSettings(): Promise<void>;
|
|
68
|
+
/**
|
|
69
|
+
* Read dsh's speech service status through the plugin's Remote.
|
|
70
|
+
*
|
|
71
|
+
* Never fails the page: an unavailable or provider-less speech service is a
|
|
72
|
+
* legitimate state (`available: false`), so the page explains it instead of
|
|
73
|
+
* showing a request error.
|
|
74
|
+
*/
|
|
75
|
+
refreshSpeechStatus(): Promise<void>;
|
|
76
|
+
/** Start (or join) dsh's provider-owned preparation task, then re-read status. */
|
|
77
|
+
prepareSpeech(providerId: string): Promise<boolean>;
|
|
59
78
|
refreshRoutes(): Promise<void>;
|
|
60
79
|
private readonly autoPopulateDefaultRoutesIfNeeded;
|
|
61
80
|
update(patch: BetterInputSettingsPatch): Promise<boolean>;
|
|
@@ -77,4 +96,5 @@ export declare function useRoutesSnapshot(controller: SettingsController): Route
|
|
|
77
96
|
export declare function useEffortsSnapshot(controller: SettingsController): EffortsSnapshot;
|
|
78
97
|
export declare function useAboutSnapshot(controller: SettingsController): AboutSnapshot;
|
|
79
98
|
export declare function useUpdateSnapshot(controller: SettingsController): UpdateSnapshot;
|
|
99
|
+
export declare function useSpeechSnapshot(controller: SettingsController): SpeechSnapshot;
|
|
80
100
|
export {};
|
package/lib/client/strings.d.ts
CHANGED
|
@@ -8,8 +8,14 @@ export type BetterInputStrings = {
|
|
|
8
8
|
voiceStart: string;
|
|
9
9
|
voiceStop: string;
|
|
10
10
|
voiceBusy: string;
|
|
11
|
-
|
|
11
|
+
voicePermissionDenied: string;
|
|
12
|
+
voiceCaptureUnavailable: string;
|
|
12
13
|
listening: string;
|
|
14
|
+
voiceCancel: string;
|
|
15
|
+
autoStopIn: string;
|
|
16
|
+
autoStopTitle: string;
|
|
17
|
+
autoStopLabel: string;
|
|
18
|
+
autoStopHint: string;
|
|
13
19
|
transcribing: string;
|
|
14
20
|
polishing: string;
|
|
15
21
|
voiceFailed: string;
|
|
@@ -21,9 +27,20 @@ export type BetterInputStrings = {
|
|
|
21
27
|
saveFailed: string;
|
|
22
28
|
languageLabel: string;
|
|
23
29
|
languageHint: string;
|
|
24
|
-
|
|
30
|
+
languageAuto: string;
|
|
31
|
+
languageCantonese: string;
|
|
32
|
+
speechStatusLabel: string;
|
|
33
|
+
speechStatusReady: string;
|
|
34
|
+
speechStatusUnavailable: string;
|
|
35
|
+
speechStatusPreparing: string;
|
|
36
|
+
speechPrepareButton: string;
|
|
37
|
+
speechPrepareBusy: string;
|
|
25
38
|
recordingLimitLabel: string;
|
|
26
39
|
recordingLimitHint: string;
|
|
40
|
+
streamingPreviewLabel: string;
|
|
41
|
+
streamingPreviewHint: string;
|
|
42
|
+
segmentSecondsLabel: string;
|
|
43
|
+
segmentSecondsHint: string;
|
|
27
44
|
polishLabel: string;
|
|
28
45
|
polishHint: string;
|
|
29
46
|
on: string;
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Plugin-owned stylesheet, injected once into `document.head`.
|
|
3
|
+
*
|
|
4
|
+
* The status bar lives in `conversation.input.dock`, the same slot dsh's own
|
|
5
|
+
* QueueDock occupies. A dock child is stretched to the **full width of the
|
|
6
|
+
* conversation column** by the flex parent, so it must constrain itself or it
|
|
7
|
+
* draws wider than the composer card above which it sits — dsh's QueueDock does
|
|
8
|
+
* exactly this:
|
|
9
|
+
*
|
|
10
|
+
* width: calc(100% - 2 * clearance - 2 * inset); max-width: card - 2 * inset
|
|
11
|
+
*
|
|
12
|
+
* We mirror the same two custom properties (`--dsh-composer-card-max-width`,
|
|
13
|
+
* `--dsh-composer-side-clearance`) without the extra dock inset, so the bar lines
|
|
14
|
+
* up edge to edge with the composer card in every layout, including the narrow
|
|
15
|
+
* `embedded` variant. Fallbacks keep it sane if a future dsh drops the names.
|
|
16
|
+
*/
|
|
17
|
+
export declare const PLUGIN_CSS = "\n.sqs-bi-bar {\n box-sizing: border-box;\n width: calc(100% - 2 * var(--dsh-composer-side-clearance, 16px));\n max-width: var(--dsh-composer-card-max-width, 952px);\n margin: 0 auto;\n flex: none;\n display: flex;\n align-items: center;\n gap: 8px;\n min-height: 32px;\n padding: 3px 8px;\n border-radius: var(--dsw-radius-lg, 12px);\n background: var(--dsw-alias-interactive-bg-hover, rgba(128, 128, 128, 0.1));\n color: var(--dsw-alias-label-secondary, inherit);\n font-size: 12px;\n line-height: 18px;\n position: relative;\n overflow: hidden;\n}\n.sqs-bi-bar:after {\n content: \"\";\n position: absolute;\n inset: 0;\n border: 0.5px solid var(--dsw-alias-border-l1, transparent);\n border-radius: inherit;\n pointer-events: none;\n}\n.sqs-bi-dot {\n flex: none;\n width: 8px;\n height: 8px;\n border-radius: 50%;\n background: var(--dsw-alias-state-error-secondary, #e5484d);\n animation: sqs-bi-pulse 1.2s ease-in-out infinite;\n}\n.sqs-bi-spin {\n flex: none;\n width: 11px;\n height: 11px;\n border: 1.5px solid currentColor;\n border-top-color: transparent;\n border-radius: 50%;\n opacity: 0.7;\n animation: sqs-bi-spin 0.9s linear infinite;\n}\n.sqs-bi-label {\n flex: none;\n max-width: 40%;\n overflow: hidden;\n text-overflow: ellipsis;\n white-space: nowrap;\n color: var(--dsw-alias-label-primary, inherit);\n font-weight: 500;\n}\n.sqs-bi-detail {\n flex: 0 1 auto;\n min-width: 0;\n overflow: hidden;\n text-overflow: ellipsis;\n white-space: nowrap;\n opacity: 0.75;\n}\n.sqs-bi-clock {\n flex: none;\n font-variant-numeric: tabular-nums;\n font-feature-settings: \"tnum\";\n padding: 0 6px;\n border-radius: 5px;\n background: var(--dsw-alias-bg-layer-1, rgba(128, 128, 128, 0.16));\n color: var(--dsw-alias-label-primary, inherit);\n}\n.sqs-bi-meter {\n flex: 0 1 108px;\n min-width: 0;\n display: flex;\n align-items: center;\n gap: 2px;\n height: 18px;\n overflow: hidden;\n opacity: 0.55;\n}\n.sqs-bi-meter i {\n flex: none;\n display: block;\n width: 3px;\n border-radius: 2px;\n background: currentColor;\n transition: height 0.12s linear;\n}\n.sqs-bi-space {\n flex: 1 1 auto;\n min-width: 4px;\n}\n.sqs-bi-count {\n flex: none;\n display: inline-flex;\n align-items: center;\n gap: 5px;\n white-space: nowrap;\n padding: 0 6px 0 4px;\n border-radius: 5px;\n background: var(--dsw-alias-bg-layer-1, rgba(128, 128, 128, 0.16));\n color: var(--dsw-alias-label-primary, inherit);\n}\n.sqs-bi-count[data-urgent=\"true\"] {\n color: var(--dsw-alias-state-error-secondary, #e5484d);\n}\n.sqs-bi-ring {\n flex: none;\n transform: rotate(-90deg);\n}\n.sqs-bi-ring circle {\n fill: none;\n stroke: currentColor;\n stroke-width: 2;\n}\n.sqs-bi-ring .sqs-bi-ring-track {\n opacity: 0.25;\n}\n.sqs-bi-btn {\n flex: none;\n appearance: none;\n border: 0;\n border-radius: 6px;\n padding: 3px 10px;\n font: inherit;\n line-height: 16px;\n white-space: nowrap;\n cursor: pointer;\n background: transparent;\n color: inherit;\n}\n.sqs-bi-btn[data-tone=\"ghost\"]:hover {\n background: var(--dsw-alias-bg-layer-1, rgba(128, 128, 128, 0.16));\n}\n.sqs-bi-btn[data-tone=\"danger\"] {\n background: var(--dsw-alias-state-error-secondary, #e5484d);\n color: #fff;\n}\n.sqs-bi-btn[data-tone=\"danger\"]:hover {\n filter: brightness(1.06);\n}\n.sqs-bi-bar[data-state=\"error\"] .sqs-bi-label,\n.sqs-bi-bar[data-state=\"polish-error\"] .sqs-bi-label {\n color: var(--dsw-alias-state-error-secondary, #e5484d);\n}\n@keyframes sqs-bi-pulse {\n 0%, 100% { opacity: 1; }\n 50% { opacity: 0.3; }\n}\n@keyframes sqs-bi-spin {\n to { transform: rotate(360deg); }\n}\n";
|
|
@@ -6,12 +6,42 @@ export type VoiceInputSessionSnapshot = {
|
|
|
6
6
|
type Listener = () => void;
|
|
7
7
|
type StopListener = () => void;
|
|
8
8
|
export declare const VOICE_ERROR_DISMISS_MS = 2600;
|
|
9
|
+
/**
|
|
10
|
+
* Live recording telemetry for the status bar: elapsed time, input level and
|
|
11
|
+
* the silence auto-stop countdown.
|
|
12
|
+
*
|
|
13
|
+
* Deliberately a **separate** store from the session state: it updates five
|
|
14
|
+
* times a second while recording, and the microphone button (which also
|
|
15
|
+
* subscribes to the session) has no use for it. `getSnapshot` returns a cached
|
|
16
|
+
* object so an unchanged frame does not re-render anything — required by
|
|
17
|
+
* `useSyncExternalStore`.
|
|
18
|
+
*/
|
|
19
|
+
export type VoiceMeterSnapshot = {
|
|
20
|
+
readonly elapsedSeconds: number;
|
|
21
|
+
readonly level: number;
|
|
22
|
+
/** Configured silence window in seconds, `0` when the auto-stop is off. */
|
|
23
|
+
readonly autoStopSeconds: number;
|
|
24
|
+
/** Seconds left before the auto-stop, `null` when no countdown is running. */
|
|
25
|
+
readonly autoStopRemainingSeconds: number | null;
|
|
26
|
+
};
|
|
27
|
+
export declare const EMPTY_VOICE_METER: VoiceMeterSnapshot;
|
|
28
|
+
export declare class VoiceMeter {
|
|
29
|
+
private snapshot;
|
|
30
|
+
private readonly listeners;
|
|
31
|
+
readonly getSnapshot: () => VoiceMeterSnapshot;
|
|
32
|
+
readonly subscribe: (listener: Listener) => (() => void);
|
|
33
|
+
publish(next: VoiceMeterSnapshot): void;
|
|
34
|
+
reset(): void;
|
|
35
|
+
dispose(): void;
|
|
36
|
+
}
|
|
9
37
|
/**
|
|
10
38
|
* Shared voice-input state for one session, written from scratch for
|
|
11
39
|
* sqs-dsh-better-input. The microphone button and the recognition bar both
|
|
12
40
|
* subscribe; the bar can request stop/cancel through the same instance.
|
|
13
41
|
*/
|
|
14
42
|
export declare class VoiceInputSession {
|
|
43
|
+
/** Recording telemetry, subscribed to only by the recognition bar. */
|
|
44
|
+
readonly meter: VoiceMeter;
|
|
15
45
|
private snapshot;
|
|
16
46
|
private readonly listeners;
|
|
17
47
|
private readonly stopListeners;
|
|
@@ -32,4 +62,6 @@ export declare class VoiceInputSession {
|
|
|
32
62
|
private emit;
|
|
33
63
|
}
|
|
34
64
|
export declare function useVoiceInputSession(session: VoiceInputSession): VoiceInputSessionSnapshot;
|
|
65
|
+
/** Subscribe to the live recording telemetry of one session. */
|
|
66
|
+
export declare function useVoiceMeter(session: VoiceInputSession): VoiceMeterSnapshot;
|
|
35
67
|
export {};
|