dsh-voice 0.3.4 → 0.3.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +21 -0
- package/README.en.md +23 -70
- package/README.md +22 -72
- package/docs/CHANGELOG.en.md +13 -0
- package/docs/USAGE.en.md +76 -0
- package/docs/USAGE.md +76 -0
- package/docs/VALIDATION.md +11 -0
- package/lib/stt.js +23 -4
- package/package.json +5 -3
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# 更新记录
|
|
2
|
+
|
|
3
|
+
[返回简介](README.md) · [使用说明](docs/USAGE.md) · [验证记录](docs/VALIDATION.md)
|
|
4
|
+
|
|
5
|
+
[历史英文记录](docs/CHANGELOG.en.md)
|
|
6
|
+
|
|
7
|
+
## 0.3.6 (2026-10-05)
|
|
8
|
+
|
|
9
|
+
- ASR 网络与响应体读取均受超时和取消控制;保留用户取消原因,空白或缺失识别结果明确报错,不伪造成功文本文件。
|
|
10
|
+
|
|
11
|
+
## 0.3.5 (2026-09-28)
|
|
12
|
+
|
|
13
|
+
- 更新官方 Harness 0.2.0-rc.1 的兼容声明和共同加载验证;运行时代码未变。验证范围见[验证记录](docs/VALIDATION.md)。
|
|
14
|
+
|
|
15
|
+
## 0.3.4 (2026-09-19)
|
|
16
|
+
|
|
17
|
+
- 修复代理路径 STT 上传 `[object FormData]`、25MB 校验在读入之后、输出目录不存在白烧一次合成、默认输出落宿主 cwd;`exec.signal` 全程透传、preview 改 3 路并发。测试 55 项。
|
|
18
|
+
|
|
19
|
+
## 更早的改动
|
|
20
|
+
|
|
21
|
+
完整历史可查阅 [GitHub 提交记录](https://github.com/STARDUSTLC666/dsh-voice/commits/master)。
|
package/README.en.md
CHANGED
|
@@ -1,93 +1,46 @@
|
|
|
1
|
-
[中文](README.md)
|
|
2
|
-
|
|
3
|
-
   
|
|
4
|
-
|
|
5
1
|
# dsh-voice
|
|
6
2
|
|
|
7
|
-
[
|
|
3
|
+
[中文](README.md)
|
|
8
4
|
|
|
9
|
-
|
|
5
|
+

|
|
10
6
|
|
|
11
|
-
|
|
12
|
-
- **voice_stt**: speech-to-text over any **OpenAI-compatible ASR endpoint** (Groq / OpenAI / custom)
|
|
13
|
-
- **voice_list**: voice catalog
|
|
14
|
-
- **voice_preview**: batch-generate short preview MP3s to pick a voice by ear
|
|
15
|
-
- **voice_health**: offline self-check of TTS voice / ASR key / endpoint / proxy
|
|
7
|
+
Generate speech from text or transcribe audio through a compatible API.
|
|
16
8
|
|
|
17
|
-
|
|
9
|
+
[](https://www.npmjs.com/package/dsh-voice) [](https://www.npmjs.com/package/dsh-voice)
|
|
18
10
|
|
|
19
|
-
|
|
11
|
+
## What it does
|
|
20
12
|
|
|
21
|
-
|
|
13
|
+
- Generate speech using the Edge online read-aloud service.
|
|
14
|
+
- List voices and generate batch voice previews.
|
|
15
|
+
- Transcribe audio using an OpenAI-compatible ASR endpoint.
|
|
22
16
|
|
|
23
|
-
|
|
24
|
-
dsh plugin --profile web add dsh-voice
|
|
25
|
-
```
|
|
17
|
+
## Install
|
|
26
18
|
|
|
27
|
-
|
|
19
|
+
In DSH Desktop, install `dsh-voice` from the Plugins panel. If the bundled dsh command is available:
|
|
28
20
|
|
|
29
21
|
```bash
|
|
30
|
-
dsh plugin --profile
|
|
22
|
+
dsh plugin --profile desktop add dsh-voice
|
|
31
23
|
```
|
|
32
24
|
|
|
33
|
-
|
|
34
|
-
|
|
25
|
+
For the web version, replace `desktop` with `web`. Restart DSH after installation.
|
|
35
26
|
|
|
36
|
-
##
|
|
27
|
+
## Start using it
|
|
37
28
|
|
|
38
|
-
|
|
29
|
+
Ask to turn text into a Chinese MP3 and preview two voices. After configuring ASR, you can also request audio transcription.
|
|
39
30
|
|
|
40
|
-
|
|
41
|
-
- id: voice
|
|
42
|
-
name: 'dsh-voice'
|
|
43
|
-
config:
|
|
44
|
-
asrEngine: groq # groq | openai | custom
|
|
45
|
-
asrModel: whisper-large-v3-turbo # Groq whisper model
|
|
46
|
-
# asrApiKey: gsk_... # prefer env var DSH_VOICE_ASR_KEY
|
|
47
|
-
ttsVoice: zh-CN-XiaoxiaoNeural # default voice
|
|
48
|
-
# proxyUrl: http://127.0.0.1:7890 # enable when the ASR endpoint needs a special proxy
|
|
49
|
-
```
|
|
31
|
+
## Requirements and configuration
|
|
50
32
|
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
| Tool | Purpose | Key parameters |
|
|
54
|
-
| :-- | :-- | :-- |
|
|
55
|
-
| `voice_tts` | Synthesize MP3 from text (free) | `text` required; `voice`/`rate`/`pitch`/`output` optional |
|
|
56
|
-
| `voice_stt` | Transcribe audio to text | `audio` required; `engine`/`model`/`language`/`prompt`/`output` optional |
|
|
57
|
-
| `voice_list` | Curated voice catalog | none |
|
|
58
|
-
| `voice_preview` | Batch-generate short preview MP3s | optional `voices` (≤8) / `text` / `outputDir` |
|
|
59
|
-
| `voice_health` | Offline config self-check | none |
|
|
60
|
-
|
|
61
|
-
### Examples
|
|
62
|
-
|
|
63
|
-
```text
|
|
64
|
-
voice_tts { text: hello world } # outputs voice_output.mp3
|
|
65
|
-
voice_tts { text: hello, voice: en-US-AriaNeural } # English female voice
|
|
66
|
-
voice_stt { audio: E:\audio\meeting.mp3, language: zh } # transcribe a recording
|
|
67
|
-
voice_list {}
|
|
68
|
-
voice_preview { voices: [zh-CN-XiaoxiaoNeural, en-US-AriaNeural] } # generate two preview samples
|
|
69
|
-
voice_health {} # self-check TTS / ASR / proxy config
|
|
70
|
-
```
|
|
33
|
+
Speech synthesis requires network access. Transcription needs an ASR endpoint and key. A plugin-specific proxy can be configured.
|
|
71
34
|
|
|
72
|
-
|
|
35
|
+
Detailed configuration, tool arguments and troubleshooting are in the [usage guide](docs/USAGE.en.md). For standalone development, follow the Node requirement in [package.json](package.json).
|
|
73
36
|
|
|
74
|
-
|
|
75
|
-
- **Zero API cost**: TTS is completely free; STT costs only whatever your ASR provider charges
|
|
76
|
-
- Up-front validation: text ≤ 5000 chars, audio ≤ 25MB; same-name outputs auto-suffixed
|
|
77
|
-
- Protocol aligned with current open-source edge-tts (7.x) — no reliance on the outdated token endpoint
|
|
37
|
+
## Documentation
|
|
78
38
|
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
pnpm test # build + offline unit tests with mocked TTS/ASR
|
|
84
|
-
pnpm test:integration # opt-in real edge-tts request; network and assertion errors fail
|
|
85
|
-
```
|
|
39
|
+
- [Usage and troubleshooting](docs/USAGE.en.md)
|
|
40
|
+
- [Changelog](CHANGELOG.md)
|
|
41
|
+
- [Validation scope and history](docs/VALIDATION.md)
|
|
42
|
+
- [Report a problem or suggest a feature](https://github.com/STARDUSTLC666/dsh-voice/issues)
|
|
86
43
|
|
|
87
44
|
## License
|
|
88
45
|
|
|
89
|
-
MIT
|
|
90
|
-
|
|
91
|
-
## Changelog
|
|
92
|
-
|
|
93
|
-
- **0.3.4 (2026-09-18)**: 修复代理路径 STT 上传 `[object FormData]`、25MB 校验在读入之后、输出目录不存在白烧一次合成、默认输出落宿主 cwd; `exec.signal` 全程透传、preview 改 3 路并发. 测试 55 项.
|
|
46
|
+
[MIT](LICENSE)
|
package/README.md
CHANGED
|
@@ -1,96 +1,46 @@
|
|
|
1
|
-
[English](README.en.md)
|
|
2
|
-
|
|
3
1
|
# dsh-voice
|
|
4
2
|
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
   
|
|
8
|
-
|
|
9
|
-
[](https://awesome-dsh-plugin.com)
|
|
3
|
+
[English](README.en.md)
|
|
10
4
|
|
|
5
|
+

|
|
11
6
|
|
|
12
|
-
|
|
7
|
+
把文字生成语音,或通过兼容接口把音频转为文字。
|
|
13
8
|
|
|
14
|
-
-
|
|
15
|
-
- **voice_stt**:语音转文字,走 **OpenAI 兼容 ASR 接口**(Groq / OpenAI / 自定义端点)
|
|
16
|
-
- **voice_list**:音色清单
|
|
17
|
-
- **voice_preview**:音色试听,用一段样例文本批量生成短 MP3,挑音色不用盲选
|
|
18
|
-
- **voice_health**:配置自检,体检 TTS 音色 / ASR 密钥 / 接口地址 / 代理(不联网)
|
|
9
|
+
[](https://www.npmjs.com/package/dsh-voice) [](https://www.npmjs.com/package/dsh-voice)
|
|
19
10
|
|
|
20
|
-
##
|
|
11
|
+
## 功能
|
|
21
12
|
|
|
22
|
-
|
|
13
|
+
- 使用 Edge 在线朗读服务生成语音。
|
|
14
|
+
- 列出音色并批量试听。
|
|
15
|
+
- 通过 OpenAI 兼容 ASR 接口转写音频。
|
|
23
16
|
|
|
24
17
|
## 安装
|
|
25
18
|
|
|
26
|
-
|
|
27
|
-
dsh plugin --profile web add dsh-voice
|
|
28
|
-
```
|
|
29
|
-
|
|
30
|
-
## 卸载
|
|
19
|
+
桌面版可在「插件」面板按包名 `dsh-voice` 安装。已配置 dsh 命令时也可使用:
|
|
31
20
|
|
|
32
21
|
```bash
|
|
33
|
-
dsh plugin --profile
|
|
22
|
+
dsh plugin --profile desktop add dsh-voice
|
|
34
23
|
```
|
|
35
24
|
|
|
36
|
-
|
|
25
|
+
网页版把命令中的 `desktop` 改为 `web`。安装后重启 DSH。
|
|
37
26
|
|
|
27
|
+
## 开始使用
|
|
38
28
|
|
|
39
|
-
|
|
29
|
+
可说:“把这段文字生成中文 MP3,并给我两个音色试听。”配置 ASR 后,也可要求转写音频文件。
|
|
40
30
|
|
|
41
|
-
|
|
31
|
+
## 依赖与配置
|
|
42
32
|
|
|
43
|
-
|
|
44
|
-
- id: voice
|
|
45
|
-
name: 'dsh-voice'
|
|
46
|
-
config:
|
|
47
|
-
asrEngine: groq # groq | openai | custom
|
|
48
|
-
asrModel: whisper-large-v3-turbo # groq 的 whisper 模型
|
|
49
|
-
# asrApiKey: gsk_... # 推荐改用环境变量 DSH_VOICE_ASR_KEY
|
|
50
|
-
ttsVoice: zh-CN-XiaoxiaoNeural # 默认音色
|
|
51
|
-
# proxyUrl: http://127.0.0.1:7890 # ASR 接口需要特殊代理时启用
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
## 工具一览
|
|
55
|
-
|
|
56
|
-
| 工具 | 作用 | 关键参数 |
|
|
57
|
-
| :-- | :-- | :-- |
|
|
58
|
-
| `voice_tts` | 文字合成 MP3(免费) | `text` 必填;`voice`/`rate`/`pitch`/`output` 可选 |
|
|
59
|
-
| `voice_stt` | 音频转文字 | `audio` 必填;`engine`/`model`/`language`/`prompt`/`output` 可选 |
|
|
60
|
-
| `voice_list` | 常用音色清单 | 无 |
|
|
61
|
-
| `voice_preview` | 音色试听:批量生成短样例 MP3 | `voices`(≤8 个)/ `text` / `outputDir` 可选 |
|
|
62
|
-
| `voice_health` | 配置自检(不联网) | 无 |
|
|
63
|
-
|
|
64
|
-
### 示例
|
|
65
|
-
|
|
66
|
-
```text
|
|
67
|
-
voice_tts { text: 今天的 AI 早报来了 } # 晓晓女声,输出 voice_output.mp3
|
|
68
|
-
voice_tts { text: hello, voice: en-US-AriaNeural } # 英文女声
|
|
69
|
-
voice_stt { audio: E:\audio\meeting.mp3, language: zh } # 转写会议录音
|
|
70
|
-
voice_list {}
|
|
71
|
-
voice_preview { voices: [zh-CN-XiaoxiaoNeural, en-US-AriaNeural] } # 生成两个试听样例
|
|
72
|
-
voice_health {} # 自检 TTS / ASR / 代理配置
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
## 硬核细节
|
|
33
|
+
语音合成需要网络连接;转写需配置对应 ASR 服务与密钥。代理可以单独设置。
|
|
76
34
|
|
|
77
|
-
|
|
78
|
-
- **零 API 成本**:TTS 完全免费;STT 只花你选的 ASR 接口的钱
|
|
79
|
-
- 文本 ≤5000 字符、音频 ≤25MB 前置校验;输出同名自动加序号
|
|
80
|
-
- 协议对齐开源 edge-tts 当前版本(7.x),不依赖过时的令牌端点
|
|
35
|
+
详细配置、工具参数与排错见[使用说明](docs/USAGE.md)。从源码独立开发时,Node 要求以 [package.json](package.json) 为准。
|
|
81
36
|
|
|
82
|
-
##
|
|
37
|
+
## 文档
|
|
83
38
|
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
```
|
|
39
|
+
- [使用与排错](docs/USAGE.md)
|
|
40
|
+
- [更新记录](CHANGELOG.md)
|
|
41
|
+
- [验证范围与历史记录](docs/VALIDATION.md)
|
|
42
|
+
- [问题反馈与功能建议](https://github.com/STARDUSTLC666/dsh-voice/issues)
|
|
89
43
|
|
|
90
44
|
## License
|
|
91
45
|
|
|
92
|
-
MIT
|
|
93
|
-
|
|
94
|
-
## 版本记录
|
|
95
|
-
|
|
96
|
-
- **0.3.4(2026-09-18)**:修复代理路径 STT 上传 `[object FormData]`、25MB 校验在读入之后、输出目录不存在白烧一次合成、默认输出落宿主 cwd;`exec.signal` 全程透传、preview 改 3 路并发。测试 55 项。
|
|
46
|
+
[MIT](LICENSE)
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# Historical release notes
|
|
2
|
+
|
|
3
|
+
[Current changelog](../CHANGELOG.md) · [Overview](../README.en.md)
|
|
4
|
+
|
|
5
|
+
These English notes preserve the earlier translations. The main changelog contains the consolidated version history.
|
|
6
|
+
|
|
7
|
+
## 0.3.6 (2026-10-05)
|
|
8
|
+
|
|
9
|
+
- Apply cancellation and timeouts to both ASR transport and response parsing. Preserve cancellation reasons and reject empty or missing transcripts.
|
|
10
|
+
|
|
11
|
+
## 0.3.4 (2026-09-19)
|
|
12
|
+
|
|
13
|
+
- 修复代理路径 STT 上传 `[object FormData]`、25MB 校验在读入之后、输出目录不存在白烧一次合成、默认输出落宿主 cwd; `exec.signal` 全程透传、preview 改 3 路并发. 测试 55 项.
|
package/docs/USAGE.en.md
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# dsh-voice usage guide
|
|
2
|
+
|
|
3
|
+
[Overview](../README.en.md) · [Changelog](../CHANGELOG.md) · [Validation](VALIDATION.md)
|
|
4
|
+
|
|
5
|
+
## Current improvements
|
|
6
|
+
|
|
7
|
+
Check the ASR endpoint, key and response format after a failure. Blank results are not saved as successful transcripts. Verify TTS and ASR service availability separately.
|
|
8
|
+
|
|
9
|
+
## Installation
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
dsh plugin --profile web add dsh-voice
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
## Uninstall
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
dsh plugin --profile web remove dsh-voice
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
Then restart the web service. To clean up fully, also remove the plugin entry from your profile `cordis.patch.yml` if you overrode it.
|
|
22
|
+
|
|
23
|
+
## Configuration
|
|
24
|
+
|
|
25
|
+
`voice_tts` works with zero config; `voice_stt` needs an ASR key:
|
|
26
|
+
|
|
27
|
+
```yaml
|
|
28
|
+
- id: voice
|
|
29
|
+
name: 'dsh-voice'
|
|
30
|
+
config:
|
|
31
|
+
asrEngine: groq # groq | openai | custom
|
|
32
|
+
asrModel: whisper-large-v3-turbo # Groq whisper model
|
|
33
|
+
# asrApiKey: gsk_... # prefer env var DSH_VOICE_ASR_KEY
|
|
34
|
+
ttsVoice: zh-CN-XiaoxiaoNeural # default voice
|
|
35
|
+
# proxyUrl: http://127.0.0.1:7890 # enable when the ASR endpoint needs a special proxy
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
## Tools
|
|
39
|
+
|
|
40
|
+
| Tool | Purpose | Key parameters |
|
|
41
|
+
| :-- | :-- | :-- |
|
|
42
|
+
| `voice_tts` | Synthesize MP3 from text (Edge online service) | `text` required; `voice`/`rate`/`pitch`/`output` optional |
|
|
43
|
+
| `voice_stt` | Transcribe audio to text | `audio` required; `engine`/`model`/`language`/`prompt`/`output` optional |
|
|
44
|
+
| `voice_list` | Curated voice catalog | none |
|
|
45
|
+
| `voice_preview` | Batch-generate short preview MP3s | optional `voices` (≤8) / `text` / `outputDir` |
|
|
46
|
+
| `voice_health` | Offline config self-check | none |
|
|
47
|
+
|
|
48
|
+
### Examples
|
|
49
|
+
|
|
50
|
+
```text
|
|
51
|
+
voice_tts { text: hello world } # outputs voice_output.mp3
|
|
52
|
+
voice_tts { text: hello, voice: en-US-AriaNeural } # English female voice
|
|
53
|
+
voice_stt { audio: E:\audio\meeting.mp3, language: zh } # transcribe a recording
|
|
54
|
+
voice_list {}
|
|
55
|
+
voice_preview { voices: [zh-CN-XiaoxiaoNeural, en-US-AriaNeural] } # generate two preview samples
|
|
56
|
+
voice_health {} # self-check TTS / ASR / proxy config
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## Implementation and file limits
|
|
60
|
+
|
|
61
|
+
- **Direct edge-tts protocol**: the Sec-MS-GEC token is **generated locally** with the official DRM algorithm (SHA256 of Windows file time + trusted client token, 5-minute windows); transport uses the `ws` library with permessage-deflate and an optional HTTP CONNECT proxy tunnel
|
|
62
|
+
- **API configuration**: TTS needs no separate API key. STT uses your configured ASR service, whose pricing and limits apply.
|
|
63
|
+
- Up-front validation: text ≤ 5000 chars, audio ≤ 25MB; same-name outputs auto-suffixed
|
|
64
|
+
- Protocol aligned with current open-source edge-tts (7.x) — no reliance on the outdated token endpoint
|
|
65
|
+
|
|
66
|
+
## Development
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
pnpm install
|
|
70
|
+
pnpm test # build + offline unit tests with mocked TTS/ASR
|
|
71
|
+
pnpm test:integration # opt-in real edge-tts request; network and assertion errors fail
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
## License
|
|
75
|
+
|
|
76
|
+
MIT
|
package/docs/USAGE.md
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# dsh-voice 使用说明
|
|
2
|
+
|
|
3
|
+
[返回简介](../README.md) · [更新记录](../CHANGELOG.md) · [验证记录](VALIDATION.md)
|
|
4
|
+
|
|
5
|
+
## 本次改进
|
|
6
|
+
|
|
7
|
+
转写失败时检查 ASR 地址、密钥与响应格式;空白识别结果不会作为成功转写保存。TTS 与 ASR 的外部服务可用性应分别核对。
|
|
8
|
+
|
|
9
|
+
## 安装
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
dsh plugin --profile web add dsh-voice
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
## 卸载
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
dsh plugin --profile web remove dsh-voice
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
卸载后重启 Web 服务。如需彻底清理,可再手动删除自己 profile `cordis.patch.yml` 中覆盖的插件行。
|
|
22
|
+
|
|
23
|
+
## 配置
|
|
24
|
+
|
|
25
|
+
`voice_tts` 零配置可用;`voice_stt` 需要 ASR 密钥:
|
|
26
|
+
|
|
27
|
+
```yaml
|
|
28
|
+
- id: voice
|
|
29
|
+
name: 'dsh-voice'
|
|
30
|
+
config:
|
|
31
|
+
asrEngine: groq # groq | openai | custom
|
|
32
|
+
asrModel: whisper-large-v3-turbo # groq 的 whisper 模型
|
|
33
|
+
# asrApiKey: gsk_... # 推荐改用环境变量 DSH_VOICE_ASR_KEY
|
|
34
|
+
ttsVoice: zh-CN-XiaoxiaoNeural # 默认音色
|
|
35
|
+
# proxyUrl: http://127.0.0.1:7890 # ASR 接口需要特殊代理时启用
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
## 工具一览
|
|
39
|
+
|
|
40
|
+
| 工具 | 作用 | 关键参数 |
|
|
41
|
+
| :-- | :-- | :-- |
|
|
42
|
+
| `voice_tts` | 文字合成 MP3(Edge 在线服务) | `text` 必填;`voice`/`rate`/`pitch`/`output` 可选 |
|
|
43
|
+
| `voice_stt` | 音频转文字 | `audio` 必填;`engine`/`model`/`language`/`prompt`/`output` 可选 |
|
|
44
|
+
| `voice_list` | 常用音色清单 | 无 |
|
|
45
|
+
| `voice_preview` | 音色试听:批量生成短样例 MP3 | `voices`(≤8 个)/ `text` / `outputDir` 可选 |
|
|
46
|
+
| `voice_health` | 配置自检(不联网) | 无 |
|
|
47
|
+
|
|
48
|
+
### 示例
|
|
49
|
+
|
|
50
|
+
```text
|
|
51
|
+
voice_tts { text: 今天的 AI 早报来了 } # 晓晓女声,输出 voice_output.mp3
|
|
52
|
+
voice_tts { text: hello, voice: en-US-AriaNeural } # 英文女声
|
|
53
|
+
voice_stt { audio: E:\audio\meeting.mp3, language: zh } # 转写会议录音
|
|
54
|
+
voice_list {}
|
|
55
|
+
voice_preview { voices: [zh-CN-XiaoxiaoNeural, en-US-AriaNeural] } # 生成两个试听样例
|
|
56
|
+
voice_health {} # 自检 TTS / ASR / 代理配置
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## 实现与文件限制
|
|
60
|
+
|
|
61
|
+
- **edge-tts 协议直连**:Sec-MS-GEC 令牌按官方 DRM 算法**本地生成**(SHA256(Windows 文件时间 + TrustedClientToken),5 分钟窗口),WS 传输用 `ws` 库 + permessage-deflate 压缩 + 可选 HTTP CONNECT 代理隧道
|
|
62
|
+
- **接口配置**:TTS 无需另填 API Key;STT 使用你配置的 ASR 服务,其计费与额度由该服务决定。
|
|
63
|
+
- 文本 ≤5000 字符、音频 ≤25MB 前置校验;输出同名自动加序号
|
|
64
|
+
- Edge TTS 协议适配保留在插件实现中,不使用旧令牌端点。
|
|
65
|
+
|
|
66
|
+
## 开发
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
pnpm install
|
|
70
|
+
pnpm test # 构建 + 离线单元测试(使用模拟的 TTS/ASR)
|
|
71
|
+
pnpm test:integration # 显式联网,调用真实 edge-tts;网络或断言失败均报错
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
## License
|
|
75
|
+
|
|
76
|
+
MIT
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
# dsh-voice 验证记录
|
|
2
|
+
|
|
3
|
+
本页整理原 README 的历史验证说明,保留当时的版本、日期与范围。自动测试、启动检查、浏览器操作和真实服务验收分别记录,不能相互替代。更详细的版本验收文件仍保留在仓库中。
|
|
4
|
+
|
|
5
|
+
## 原中文记录
|
|
6
|
+
|
|
7
|
+
验证宿主:官方源码构建的 Harness `0.2.0-rc.1`(commit `407e65c8`)+ Node `24.16.0`(2026-09-28)。55 项插件测试在隔离环境全部通过;同一个宿主里 18 个插件共同加载,注册 5 个工具,工具 schema 与健康检查契约通过。本轮未启用真实端口与外部服务。
|
|
8
|
+
|
|
9
|
+
## Original English record
|
|
10
|
+
|
|
11
|
+
Validation host: Harness `0.2.0-rc.1` built from official sources (commit `407e65c8`) with Node `24.16.0` on 2026-09-28. All 55 plugin tests pass in an isolated environment; all 18 plugins mount together in one host registering 5 tools, with tool schemas and health-check contracts passing. No live ports or external services were exercised in this round.
|
package/lib/stt.js
CHANGED
|
@@ -59,24 +59,43 @@ export async function transcribe(baseUrl, apiKey, options, fetchImpl = globalThi
|
|
|
59
59
|
headers: { authorization: 'Bearer ' + apiKey },
|
|
60
60
|
body: form,
|
|
61
61
|
signal: requestSignal,
|
|
62
|
-
}),
|
|
62
|
+
}), requestSignal);
|
|
63
63
|
}
|
|
64
64
|
catch (error) {
|
|
65
65
|
throwIfAborted(abortSignal);
|
|
66
|
+
if (timeoutSignal.aborted)
|
|
67
|
+
throw new Error('ASR 请求超时,请重试或缩短音频。');
|
|
66
68
|
throw new Error('ASR 请求失败:' + (error instanceof Error ? error.message : String(error)) + '。若接口需要特殊代理(梯子),请在 cordis.patch.yml 配置 proxyUrl 后重启。');
|
|
67
69
|
}
|
|
68
70
|
if (!response.ok) {
|
|
69
|
-
|
|
71
|
+
let body;
|
|
72
|
+
try {
|
|
73
|
+
body = (await raceAbort(response.text(), requestSignal)).slice(0, 300);
|
|
74
|
+
}
|
|
75
|
+
catch (error) {
|
|
76
|
+
throwIfAborted(abortSignal);
|
|
77
|
+
if (timeoutSignal.aborted)
|
|
78
|
+
throw new Error('ASR 响应读取超时,请重试或缩短音频。');
|
|
79
|
+
throw error;
|
|
80
|
+
}
|
|
70
81
|
throw new Error('ASR 失败:HTTP ' + response.status + '。' + body);
|
|
71
82
|
}
|
|
72
83
|
let json;
|
|
73
84
|
try {
|
|
74
|
-
json = await response.json();
|
|
85
|
+
json = await raceAbort(response.json(), requestSignal);
|
|
75
86
|
}
|
|
76
87
|
catch {
|
|
88
|
+
throwIfAborted(abortSignal);
|
|
89
|
+
if (timeoutSignal.aborted)
|
|
90
|
+
throw new Error('ASR 响应读取超时,请重试或缩短音频。');
|
|
77
91
|
throw new Error('ASR 响应不是合法 JSON。');
|
|
78
92
|
}
|
|
79
|
-
|
|
93
|
+
if (typeof json !== 'object' || json === null || typeof json.text !== 'string') {
|
|
94
|
+
throw new Error('ASR 响应缺少文本字段,请检查接口是否兼容 /audio/transcriptions。');
|
|
95
|
+
}
|
|
96
|
+
const text = json.text;
|
|
97
|
+
if (text.trim() === '')
|
|
98
|
+
throw new Error('未识别到语音,请检查音频是否包含清晰的人声。');
|
|
80
99
|
return { text, model: options.model };
|
|
81
100
|
}
|
|
82
101
|
/** 常见音频扩展名 → MIME。 */
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "dsh-voice",
|
|
3
|
-
"version": "0.3.
|
|
4
|
-
"description": "
|
|
3
|
+
"version": "0.3.6",
|
|
4
|
+
"description": "把文字生成语音,或通过兼容接口把音频转为文字。",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "lib/index.js",
|
|
7
7
|
"types": "lib/index.d.ts",
|
|
@@ -17,7 +17,9 @@
|
|
|
17
17
|
"lib",
|
|
18
18
|
"cordis.patch.yml",
|
|
19
19
|
"README.md",
|
|
20
|
-
"README.en.md"
|
|
20
|
+
"README.en.md",
|
|
21
|
+
"docs",
|
|
22
|
+
"CHANGELOG.md"
|
|
21
23
|
],
|
|
22
24
|
"scripts": {
|
|
23
25
|
"build": "tsc -p tsconfig.json",
|