@jerryliang122/openclaw-qqbot 1.0.8 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -810,44 +810,31 @@ The plugin processes messages through a carefully ordered middleware chain:
810
810
 
811
811
  #### STT (Speech-to-Text) — Transcribe Incoming Voice Messages
812
812
 
813
- STT supports two-level configuration with priority fallback:
814
-
815
- | Priority | Config Path | Scope |
816
- |----------|------------|-------|
817
- | 1 (highest) | `channels.qqbot.stt` | Plugin-specific |
818
- | 2 (fallback) | `tools.media.audio.models[0]` | Framework-level |
813
+ Transcription runs through the **framework audio-understanding pipeline** (`openclaw/plugin-sdk/media-understanding-runtime`), so STT credentials come from the framework-level config (same as the built-in Telegram channel) — the plugin no longer ships its own OpenAI-compatible HTTP call:
819
814
 
820
815
  ```json
821
816
  {
822
- "channels": {
823
- "qqbot": {
824
- "stt": {
825
- "provider": "your-provider",
826
- "model": "your-stt-model"
817
+ "tools": {
818
+ "media": {
819
+ "audio": {
820
+ "models": [{ "provider": "your-provider", "model": "your-stt-model" }]
827
821
  }
828
822
  }
829
823
  }
830
824
  }
831
825
  ```
832
826
 
833
- - `provider` — references a key in `models.providers` to inherit `baseUrl` and `apiKey`
834
- - Set `enabled: false` to disable
835
- - When configured, incoming voice messages are automatically converted (SILK→WAV) and transcribed
836
- - `asrFallback` — platform ASR (`asr_refer_text`) participation switch. Unless explicitly set to `true`, QQ's built-in platform transcript is **discarded in all cases**: not used as a fallback when your STT fails or returns empty, and not used as the sole source when STT is not configured at all (voice messages then render as `[Voice message - transcription unavailable]`; the audio URL is still referenced via the `- Voice:` line). The flag is read from `channels.qqbot.stt.asrFallback` regardless of whether STT credentials resolve — `stt: { "asrFallback": true }` alone restores the legacy platform-transcript behavior:
827
+ Voice message handling order:
837
828
 
838
- ```json
839
- {
840
- "channels": {
841
- "qqbot": {
842
- "stt": {
843
- "provider": "your-provider",
844
- "model": "your-stt-model",
845
- "asrFallback": true
846
- }
847
- }
848
- }
849
- }
850
- ```
829
+ 1. **Framework STT configured** (`tools.media.audio.models` non-empty) → voice is downloaded, converted (SILK→WAV), and transcribed through the framework pipeline. On failure or empty transcript, falls back to the platform transcript.
830
+ 2. **Framework STT not configured** → the **QQ platform transcript (`asr_refer_text`)** — QQ auto-STTs voice messages and ships the text in the event JSON — is used directly as the sole source, no download and no external call needed.
831
+ 3. Neither available → placeholder text (`[Voice message - transcription unavailable]`); the audio URL is still referenced via the `- Voice:` line.
832
+
833
+ Plugin-level behavior switches under `channels.qqbot.stt`:
834
+
835
+ - `enabled: false` — never call external STT; use only the platform transcript (or the placeholder)
836
+ - `asrFallback: false` — strict mode: discard the platform transcript in **all** cases (restores the pre-2026-10 behavior)
837
+ - `provider` / `baseUrl` / `apiKey` / `model` — **deprecated and ignored** (2026-10); a one-time migration notice is logged if still present. Move credentials to `tools.media.audio.models`.
851
838
 
852
839
  #### TTS (Text-to-Speech) — Send Voice Messages
853
840
 
package/README.zh.md CHANGED
@@ -631,44 +631,31 @@ openclaw message send --channel "qqbot" \
631
631
 
632
632
  #### STT(语音转文字)— 自动转录用户发来的语音消息
633
633
 
634
- STT 支持两级配置,按优先级查找:
635
-
636
- | 优先级 | 配置路径 | 作用域 |
637
- |--------|----------|--------|
638
- | 1(highest) | `channels.qqbot.stt` | 插件专属 |
639
- | 2(fallback) | `tools.media.audio.models[0]` | 框架级 |
634
+ 转录统一走**框架音频理解管线**(`openclaw/plugin-sdk/media-understanding-runtime`),STT 凭证只认框架级配置(与内置 Telegram 通道一致)——插件不再自带 OpenAI 兼容 HTTP 调用:
640
635
 
641
636
  ```json
642
637
  {
643
- "channels": {
644
- "qqbot": {
645
- "stt": {
646
- "provider": "your-provider",
647
- "model": "your-stt-model"
638
+ "tools": {
639
+ "media": {
640
+ "audio": {
641
+ "models": [{ "provider": "your-provider", "model": "your-stt-model" }]
648
642
  }
649
643
  }
650
644
  }
651
645
  }
652
646
  ```
653
647
 
654
- - `provider` — 引用 `models.providers` 中的 key,自动继承 `baseUrl` 和 `apiKey`
655
- - 设置 `enabled: false` 可禁用
656
- - 配置后,用户发来的语音消息会自动转换(SILK→WAV)并转录为文字
657
- - `asrFallback` — 平台转写(`asr_refer_text`)参与开关。**未显式设为 `true` 时,平台转写在所有场景下都被丢弃**:自有 STT 失败或返回空时不作兜底,STT 未配置时也不作为唯一来源(此时语音消息渲染为 `[Voice message - transcription unavailable]` 占位文本,音频 URL 仍通过 `- Voice:` 行引用)。该开关从 `channels.qqbot.stt.asrFallback` 读取,与 STT 凭证是否解析成功无关——仅写 `stt: { "asrFallback": true }` 即可恢复平台转写参与旧行为:
648
+ 语音消息处理顺序:
658
649
 
659
- ```json
660
- {
661
- "channels": {
662
- "qqbot": {
663
- "stt": {
664
- "provider": "your-provider",
665
- "model": "your-stt-model",
666
- "asrFallback": true
667
- }
668
- }
669
- }
670
- }
671
- ```
650
+ 1. **框架 STT 已配置**(`tools.media.audio.models` 非空)→ 下载语音、转换(SILK→WAV)后经框架管线转录;转录失败或为空时兜底用平台转写。
651
+ 2. **框架 STT 未配置** → **直接采用 QQ 平台转写**(`asr_refer_text`,QQ 平台对语音消息自动 STT 并随事件 JSON 下发)作为唯一来源——无需下载、不发起任何外部调用。
652
+ 3. 两者都不可用 → 占位文本(`[Voice message - transcription unavailable]`),音频 URL 仍通过 `- Voice:` 行引用。
653
+
654
+ `channels.qqbot.stt` 下仅保留行为开关:
655
+
656
+ - `enabled: false` — 不调用外部 STT,语音只用平台转写(或无转写时占位文本)
657
+ - `asrFallback: false` — 严格模式:所有场景丢弃平台转写(恢复 2026-10 之前的旧行为)
658
+ - `provider` / `baseUrl` / `apiKey` / `model` — **已废弃并被忽略**(2026-10);检测到仍配置时会打一次性迁移提示日志,请把凭证迁移到 `tools.media.audio.models`。
672
659
 
673
660
  #### TTS(文字转语音)— 机器人发送语音消息
674
661