@jerryliang122/openclaw-qqbot 1.0.8 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +15 -28
- package/README.zh.md +15 -28
- package/dist/index.cjs +264 -247
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +21 -7
- package/package.json +1 -1
- package/src/middleware/attachment.ts +30 -19
- package/src/openclaw-plugin-sdk.d.ts +38 -2
- package/src/tools/platform.ts +86 -72
- package/src/tools/remind.ts +53 -42
- package/src/tools/secret-input.ts +112 -80
- package/src/tools/tool-session.ts +74 -0
- package/src/types.ts +21 -7
- package/src/utils/stt.ts +49 -92
package/README.md
CHANGED
|
@@ -810,44 +810,31 @@ The plugin processes messages through a carefully ordered middleware chain:
|
|
|
810
810
|
|
|
811
811
|
#### STT (Speech-to-Text) — Transcribe Incoming Voice Messages
|
|
812
812
|
|
|
813
|
-
STT
|
|
814
|
-
|
|
815
|
-
| Priority | Config Path | Scope |
|
|
816
|
-
|----------|------------|-------|
|
|
817
|
-
| 1 (highest) | `channels.qqbot.stt` | Plugin-specific |
|
|
818
|
-
| 2 (fallback) | `tools.media.audio.models[0]` | Framework-level |
|
|
813
|
+
Transcription runs through the **framework audio-understanding pipeline** (`openclaw/plugin-sdk/media-understanding-runtime`), so STT credentials come from the framework-level config (same as the built-in Telegram channel) — the plugin no longer ships its own OpenAI-compatible HTTP call:
|
|
819
814
|
|
|
820
815
|
```json
|
|
821
816
|
{
|
|
822
|
-
"
|
|
823
|
-
"
|
|
824
|
-
"
|
|
825
|
-
"provider": "your-provider",
|
|
826
|
-
"model": "your-stt-model"
|
|
817
|
+
"tools": {
|
|
818
|
+
"media": {
|
|
819
|
+
"audio": {
|
|
820
|
+
"models": [{ "provider": "your-provider", "model": "your-stt-model" }]
|
|
827
821
|
}
|
|
828
822
|
}
|
|
829
823
|
}
|
|
830
824
|
}
|
|
831
825
|
```
|
|
832
826
|
|
|
833
|
-
|
|
834
|
-
- Set `enabled: false` to disable
|
|
835
|
-
- When configured, incoming voice messages are automatically converted (SILK→WAV) and transcribed
|
|
836
|
-
- `asrFallback` — platform ASR (`asr_refer_text`) participation switch. Unless explicitly set to `true`, QQ's built-in platform transcript is **discarded in all cases**: not used as a fallback when your STT fails or returns empty, and not used as the sole source when STT is not configured at all (voice messages then render as `[Voice message - transcription unavailable]`; the audio URL is still referenced via the `- Voice:` line). The flag is read from `channels.qqbot.stt.asrFallback` regardless of whether STT credentials resolve — `stt: { "asrFallback": true }` alone restores the legacy platform-transcript behavior:
|
|
827
|
+
Voice message handling order:
|
|
837
828
|
|
|
838
|
-
|
|
839
|
-
|
|
840
|
-
|
|
841
|
-
|
|
842
|
-
|
|
843
|
-
|
|
844
|
-
|
|
845
|
-
|
|
846
|
-
|
|
847
|
-
}
|
|
848
|
-
}
|
|
849
|
-
}
|
|
850
|
-
```
|
|
829
|
+
1. **Framework STT configured** (`tools.media.audio.models` non-empty) → voice is downloaded, converted (SILK→WAV), and transcribed through the framework pipeline. On failure or empty transcript, falls back to the platform transcript.
|
|
830
|
+
2. **Framework STT not configured** → the **QQ platform transcript (`asr_refer_text`)** — QQ auto-STTs voice messages and ships the text in the event JSON — is used directly as the sole source, no download and no external call needed.
|
|
831
|
+
3. Neither available → placeholder text (`[Voice message - transcription unavailable]`); the audio URL is still referenced via the `- Voice:` line.
|
|
832
|
+
|
|
833
|
+
Plugin-level behavior switches under `channels.qqbot.stt`:
|
|
834
|
+
|
|
835
|
+
- `enabled: false` — never call external STT; use only the platform transcript (or the placeholder)
|
|
836
|
+
- `asrFallback: false` — strict mode: discard the platform transcript in **all** cases (restores the pre-2026-10 behavior)
|
|
837
|
+
- `provider` / `baseUrl` / `apiKey` / `model` — **deprecated and ignored** (2026-10); a one-time migration notice is logged if still present. Move credentials to `tools.media.audio.models`.
|
|
851
838
|
|
|
852
839
|
#### TTS (Text-to-Speech) — Send Voice Messages
|
|
853
840
|
|
package/README.zh.md
CHANGED
|
@@ -631,44 +631,31 @@ openclaw message send --channel "qqbot" \
|
|
|
631
631
|
|
|
632
632
|
#### STT(语音转文字)— 自动转录用户发来的语音消息
|
|
633
633
|
|
|
634
|
-
STT
|
|
635
|
-
|
|
636
|
-
| 优先级 | 配置路径 | 作用域 |
|
|
637
|
-
|--------|----------|--------|
|
|
638
|
-
| 1(highest) | `channels.qqbot.stt` | 插件专属 |
|
|
639
|
-
| 2(fallback) | `tools.media.audio.models[0]` | 框架级 |
|
|
634
|
+
转录统一走**框架音频理解管线**(`openclaw/plugin-sdk/media-understanding-runtime`),STT 凭证只认框架级配置(与内置 Telegram 通道一致)——插件不再自带 OpenAI 兼容 HTTP 调用:
|
|
640
635
|
|
|
641
636
|
```json
|
|
642
637
|
{
|
|
643
|
-
"
|
|
644
|
-
"
|
|
645
|
-
"
|
|
646
|
-
"provider": "your-provider",
|
|
647
|
-
"model": "your-stt-model"
|
|
638
|
+
"tools": {
|
|
639
|
+
"media": {
|
|
640
|
+
"audio": {
|
|
641
|
+
"models": [{ "provider": "your-provider", "model": "your-stt-model" }]
|
|
648
642
|
}
|
|
649
643
|
}
|
|
650
644
|
}
|
|
651
645
|
}
|
|
652
646
|
```
|
|
653
647
|
|
|
654
|
-
|
|
655
|
-
- 设置 `enabled: false` 可禁用
|
|
656
|
-
- 配置后,用户发来的语音消息会自动转换(SILK→WAV)并转录为文字
|
|
657
|
-
- `asrFallback` — 平台转写(`asr_refer_text`)参与开关。**未显式设为 `true` 时,平台转写在所有场景下都被丢弃**:自有 STT 失败或返回空时不作兜底,STT 未配置时也不作为唯一来源(此时语音消息渲染为 `[Voice message - transcription unavailable]` 占位文本,音频 URL 仍通过 `- Voice:` 行引用)。该开关从 `channels.qqbot.stt.asrFallback` 读取,与 STT 凭证是否解析成功无关——仅写 `stt: { "asrFallback": true }` 即可恢复平台转写参与旧行为:
|
|
648
|
+
语音消息处理顺序:
|
|
658
649
|
|
|
659
|
-
|
|
660
|
-
|
|
661
|
-
|
|
662
|
-
|
|
663
|
-
|
|
664
|
-
|
|
665
|
-
|
|
666
|
-
|
|
667
|
-
|
|
668
|
-
}
|
|
669
|
-
}
|
|
670
|
-
}
|
|
671
|
-
```
|
|
650
|
+
1. **框架 STT 已配置**(`tools.media.audio.models` 非空)→ 下载语音、转换(SILK→WAV)后经框架管线转录;转录失败或为空时兜底用平台转写。
|
|
651
|
+
2. **框架 STT 未配置** → **直接采用 QQ 平台转写**(`asr_refer_text`,QQ 平台对语音消息自动 STT 并随事件 JSON 下发)作为唯一来源——无需下载、不发起任何外部调用。
|
|
652
|
+
3. 两者都不可用 → 占位文本(`[Voice message - transcription unavailable]`),音频 URL 仍通过 `- Voice:` 行引用。
|
|
653
|
+
|
|
654
|
+
`channels.qqbot.stt` 下仅保留行为开关:
|
|
655
|
+
|
|
656
|
+
- `enabled: false` — 不调用外部 STT,语音只用平台转写(或无转写时占位文本)
|
|
657
|
+
- `asrFallback: false` — 严格模式:所有场景丢弃平台转写(恢复 2026-10 之前的旧行为)
|
|
658
|
+
- `provider` / `baseUrl` / `apiKey` / `model` — **已废弃并被忽略**(2026-10);检测到仍配置时会打一次性迁移提示日志,请把凭证迁移到 `tools.media.audio.models`。
|
|
672
659
|
|
|
673
660
|
#### TTS(文字转语音)— 机器人发送语音消息
|
|
674
661
|
|