@jerryliang122/openclaw-qqbot 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -810,31 +810,36 @@ The plugin processes messages through a carefully ordered middleware chain:
810
810
 
811
811
  #### STT (Speech-to-Text) — Transcribe Incoming Voice Messages
812
812
 
813
- Transcription runs through the **framework audio-understanding pipeline** (`openclaw/plugin-sdk/media-understanding-runtime`), so STT credentials come from the framework-level config (same as the built-in Telegram channel) — the plugin no longer ships its own OpenAI-compatible HTTP call:
813
+ Transcription runs through the **framework audio-understanding pipeline** (`openclaw/plugin-sdk/media-understanding-runtime`), so STT config comes from the framework-level `tools.media.models` (same as the built-in Telegram channel) — the plugin no longer ships its own OpenAI-compatible HTTP call:
814
814
 
815
815
  ```json
816
816
  {
817
817
  "tools": {
818
818
  "media": {
819
- "audio": {
820
- "models": [{ "provider": "your-provider", "model": "your-stt-model" }]
821
- }
819
+ "models": [
820
+ {
821
+ "type": "cli",
822
+ "command": "your-asr-cli",
823
+ "args": ["-m", "your-model", "{{AttachmentPath}}"],
824
+ "maxBytes": 52428800,
825
+ "timeoutSeconds": 120,
826
+ "capabilities": ["audio"]
827
+ }
828
+ ]
822
829
  }
823
830
  }
824
831
  }
825
832
  ```
826
833
 
827
- Voice message handling order:
834
+ A provider-form entry also works (e.g. `{ "provider": "openai", "model": "whisper-1", "capabilities": ["audio"] }`). **`capabilities` must include `"audio"`** — the shared model list is selected by capability tags; untagged entries never take part in voice transcription.
828
835
 
829
- 1. **Framework STT configured** (`tools.media.audio.models` non-empty) → voice is downloaded, converted (SILK→WAV), and transcribed through the framework pipeline. On failure or empty transcript, falls back to the platform transcript.
830
- 2. **Framework STT not configured** → the **QQ platform transcript (`asr_refer_text`)** — QQ auto-STTs voice messages and ships the text in the event JSON — is used directly as the sole source, no download and no external call needed.
831
- 3. Neither available → placeholder text (`[Voice message - transcription unavailable]`); the audio URL is still referenced via the `- Voice:` line.
836
+ Voice message handling order (hardcoded since 2026-10; no plugin-level switches):
832
837
 
833
- Plugin-level behavior switches under `channels.qqbot.stt`:
838
+ 1. **Framework STT configured** (a `tools.media.models` entry whose `capabilities` includes `"audio"`) → voice is downloaded, converted (SILK→WAV), and transcribed through the framework pipeline. **The framework is trusted strictly** — on failure or empty transcript a placeholder (`[Voice message - transcription failed]`) is emitted; there is **no** fallback to the platform transcript.
839
+ 2. **Framework STT not configured** → the **QQ platform transcript (`asr_refer_text`)** — QQ auto-STTs voice messages and ships the text in the event JSON — is used directly as the sole source, no download and no external call needed.
840
+ 3. Not configured and no platform transcript → placeholder text (`[Voice message - transcription unavailable]`); the audio URL is still referenced via the `- Voice:` line.
834
841
 
835
- - `enabled: false` — never call external STT; use only the platform transcript (or the placeholder)
836
- - `asrFallback: false` — strict mode: discard the platform transcript in **all** cases (restores the pre-2026-10 behavior)
837
- - `provider` / `baseUrl` / `apiKey` / `model` — **deprecated and ignored** (2026-10); a one-time migration notice is logged if still present. Move credentials to `tools.media.audio.models`.
842
+ The whole `channels.qqbot.stt` block is **deprecated and ignored** (legacy credential keys plus the historical `enabled`/`asrFallback` switches, removed 2026-10-04); a one-time migration notice is logged if any key is present. STT is controlled solely by the framework config — remove the audio entry from `tools.media.models` (or set `tools.media.audio.enabled: false`) to disable framework transcription.
838
843
 
839
844
  #### TTS (Text-to-Speech) — Send Voice Messages
840
845
 
package/README.zh.md CHANGED
@@ -631,31 +631,36 @@ openclaw message send --channel "qqbot" \
631
631
 
632
632
  #### STT(语音转文字)— 自动转录用户发来的语音消息
633
633
 
634
- 转录统一走**框架音频理解管线**(`openclaw/plugin-sdk/media-understanding-runtime`),STT 凭证只认框架级配置(与内置 Telegram 通道一致)——插件不再自带 OpenAI 兼容 HTTP 调用:
634
+ 转录统一走**框架音频理解管线**(`openclaw/plugin-sdk/media-understanding-runtime`),STT 配置只认框架级 `tools.media.models`(与内置 Telegram 通道一致)——插件不再自带 OpenAI 兼容 HTTP 调用:
635
635
 
636
636
  ```json
637
637
  {
638
638
  "tools": {
639
639
  "media": {
640
- "audio": {
641
- "models": [{ "provider": "your-provider", "model": "your-stt-model" }]
642
- }
640
+ "models": [
641
+ {
642
+ "type": "cli",
643
+ "command": "your-asr-cli",
644
+ "args": ["-m", "your-model", "{{AttachmentPath}}"],
645
+ "maxBytes": 52428800,
646
+ "timeoutSeconds": 120,
647
+ "capabilities": ["audio"]
648
+ }
649
+ ]
643
650
  }
644
651
  }
645
652
  }
646
653
  ```
647
654
 
648
- 语音消息处理顺序:
655
+ 也支持 provider 形态(如 `{ "provider": "openai", "model": "whisper-1", "capabilities": ["audio"] }`)。**`capabilities` 必须包含 `"audio"`**——模型列表按能力标签选择,无标签条目不参与语音转录。
649
656
 
650
- 1. **框架 STT 已配置**(`tools.media.audio.models` 非空)→ 下载语音、转换(SILK→WAV)后经框架管线转录;转录失败或为空时兜底用平台转写。
651
- 2. **框架 STT 未配置** → **直接采用 QQ 平台转写**(`asr_refer_text`,QQ 平台对语音消息自动 STT 并随事件 JSON 下发)作为唯一来源——无需下载、不发起任何外部调用。
652
- 3. 两者都不可用 → 占位文本(`[Voice message - transcription unavailable]`),音频 URL 仍通过 `- Voice:` 行引用。
657
+ 语音消息处理顺序(2026-10 起硬编码,无插件级开关):
653
658
 
654
- `channels.qqbot.stt` 下仅保留行为开关:
659
+ 1. **框架 STT 已配置**(`tools.media.models` 存在 `capabilities` 含 `"audio"` 的条目)→ 下载语音、转换(SILK→WAV)后经框架管线转录;**严格信框架**——转录失败或为空时输出占位文本(`[Voice message - transcription failed]`),不回退平台转写。
660
+ 2. **框架 STT 未配置** → **直接采用 QQ 平台转写**(`asr_refer_text`,QQ 平台对语音消息自动 STT 并随事件 JSON 下发)作为唯一来源——无需下载、不发起任何外部调用。
661
+ 3. 未配置且无平台转写 → 占位文本(`[Voice message - transcription unavailable]`),音频 URL 仍通过 `- Voice:` 行引用。
655
662
 
656
- - `enabled: false` — 不调用外部 STT,语音只用平台转写(或无转写时占位文本)
657
- - `asrFallback: false` — 严格模式:所有场景丢弃平台转写(恢复 2026-10 之前的旧行为)
658
- - `provider` / `baseUrl` / `apiKey` / `model` — **已废弃并被忽略**(2026-10);检测到仍配置时会打一次性迁移提示日志,请把凭证迁移到 `tools.media.audio.models`。
663
+ `channels.qqbot.stt` 整块**已废弃并被忽略**(旧凭证键 + 历史 `enabled`/`asrFallback` 开关,2026-10-04 移除);检测到任何键时会打一次性迁移提示日志。STT 启停只由框架配置控制——删掉 `tools.media.models` 的 audio 条目、或设 `tools.media.audio.enabled: false` 即关闭框架转录。
659
664
 
660
665
  #### TTS(文字转语音)— 机器人发送语音消息
661
666
 
package/dist/index.cjs CHANGED
@@ -8259,7 +8259,7 @@ var init_outbound_service = __esm({
8259
8259
 
8260
8260
  // src/utils/pkg-version.ts
8261
8261
  function getPackageVersion() {
8262
- return true ? "2.0.0" : "unknown";
8262
+ return true ? "2.1.0" : "unknown";
8263
8263
  }
8264
8264
  function getOpenclawVersion(runtimeVersion) {
8265
8265
  return runtimeVersion && runtimeVersion !== "unknown" ? runtimeVersion : "unknown";
@@ -12906,33 +12906,32 @@ init_webhook_verify();
12906
12906
  // src/utils/stt.ts
12907
12907
  var path16 = __toESM(require("path"), 1);
12908
12908
  var import_media_understanding_runtime = require("openclaw/plugin-sdk/media-understanding-runtime");
12909
- function shouldUsePlatformAsr(cfg) {
12910
- const channels = asRecord(cfg.channels);
12911
- const qqbot = asRecord(channels?.qqbot);
12912
- return asRecord(qqbot?.stt)?.asrFallback !== false;
12913
- }
12914
12909
  function isFrameworkSttConfigured(cfg) {
12915
- const channels = asRecord(cfg.channels);
12916
- const qqbot = asRecord(channels?.qqbot);
12917
- if (asRecord(qqbot?.stt)?.enabled === false) {
12918
- return false;
12919
- }
12920
12910
  const tools = asRecord(cfg.tools);
12921
12911
  const media = asRecord(tools?.media);
12922
- const audio = asRecord(media?.audio);
12923
- if (!audio || audio.enabled === false) {
12912
+ if (asRecord(media?.audio)?.enabled === false) {
12924
12913
  return false;
12925
12914
  }
12926
- return Array.isArray(audio.models) && audio.models.length > 0;
12915
+ const models = media?.models;
12916
+ if (!Array.isArray(models)) {
12917
+ return false;
12918
+ }
12919
+ return models.some((entry) => {
12920
+ const capabilities = asRecord(entry)?.capabilities;
12921
+ return Array.isArray(capabilities) && capabilities.includes("audio");
12922
+ });
12927
12923
  }
12928
- function hasLegacySttCredentials(cfg) {
12924
+ function hasLegacySttConfig(cfg) {
12929
12925
  const channels = asRecord(cfg.channels);
12930
12926
  const qqbot = asRecord(channels?.qqbot);
12931
12927
  const stt = asRecord(qqbot?.stt);
12932
12928
  if (!stt) return false;
12933
- return ["provider", "baseUrl", "apiKey", "model"].some(
12934
- (key2) => typeof stt[key2] === "string" && stt[key2].trim().length > 0
12935
- );
12929
+ const legacyKeys = ["provider", "baseUrl", "apiKey", "model", "enabled", "asrFallback"];
12930
+ return legacyKeys.some((key2) => {
12931
+ const value = stt[key2];
12932
+ if (typeof value === "string") return value.trim().length > 0;
12933
+ return value != null;
12934
+ });
12936
12935
  }
12937
12936
  async function transcribeAudioViaFramework(audioPath, cfg) {
12938
12937
  const result = await (0, import_media_understanding_runtime.transcribeAudioFile)({
@@ -13023,9 +13022,8 @@ function attachmentProcessor(opts) {
13023
13022
  };
13024
13023
  }
13025
13024
  async function processAttachments(attachments, cfg, log5) {
13026
- const usePlatformAsr = shouldUsePlatformAsr(cfg);
13027
13025
  const sttConfigured = isFrameworkSttConfigured(cfg);
13028
- warnLegacySttCredentials(cfg, log5);
13026
+ warnLegacySttConfig(cfg, log5);
13029
13027
  const audioPolicy = resolveAudioPolicy(cfg);
13030
13028
  const imageUrls = [];
13031
13029
  const otherParts = [];
@@ -13043,7 +13041,7 @@ async function processAttachments(attachments, cfg, log5) {
13043
13041
  return { type: "image", localPath, url, contentType: att.content_type ?? "image/png", filename: att.filename };
13044
13042
  }
13045
13043
  if (isVoice) {
13046
- const transcript = await processVoiceAttachment(att, cfg, usePlatformAsr, sttConfigured, audioPolicy, log5);
13044
+ const transcript = await processVoiceAttachment(att, cfg, sttConfigured, audioPolicy, log5);
13047
13045
  return { type: "voice", transcript };
13048
13046
  }
13049
13047
  if (url) {
@@ -13117,22 +13115,18 @@ function kindFromContentType(contentType) {
13117
13115
  return "document";
13118
13116
  }
13119
13117
  var legacySttWarned = false;
13120
- function warnLegacySttCredentials(cfg, log5) {
13121
- if (!legacySttWarned && hasLegacySttCredentials(cfg)) {
13118
+ function warnLegacySttConfig(cfg, log5) {
13119
+ if (!legacySttWarned && hasLegacySttConfig(cfg)) {
13122
13120
  legacySttWarned = true;
13123
13121
  log5?.info(
13124
- "Voice: channels.qqbot.stt credentials (provider/baseUrl/apiKey/model) are deprecated and ignored; configure tools.media.audio.models instead \u2014 platform asr_refer_text is used when framework STT is absent"
13122
+ "Voice: channels.qqbot.stt is deprecated and ignored entirely (credentials + enabled/asrFallback); configure an audio-capable tools.media.models entry for STT \u2014 platform asr_refer_text is used only when framework STT is absent"
13125
13123
  );
13126
13124
  }
13127
13125
  }
13128
- async function processVoiceAttachment(att, cfg, usePlatformAsr, sttConfigured, audioPolicy, log5) {
13129
- const rawAsrText = att.asr_refer_text?.trim() || void 0;
13130
- const asrReferText = usePlatformAsr ? rawAsrText : void 0;
13126
+ async function processVoiceAttachment(att, cfg, sttConfigured, audioPolicy, log5) {
13131
13127
  const remoteUrl = normalizeUrl(att.voice_wav_url) || normalizeUrl(att.url) || void 0;
13132
13128
  if (!sttConfigured) {
13133
- if (!usePlatformAsr && rawAsrText) {
13134
- log5?.info(`Voice: framework STT not configured; platform asr_refer_text discarded (asrFallback: false)`);
13135
- }
13129
+ const asrReferText = att.asr_refer_text?.trim() || void 0;
13136
13130
  if (asrReferText) {
13137
13131
  log5?.debug?.(`Voice: using platform asr_refer_text (framework STT not configured)`);
13138
13132
  return { text: asrReferText, source: "asr", asrReferText, remoteUrl };
@@ -13140,7 +13134,6 @@ async function processVoiceAttachment(att, cfg, usePlatformAsr, sttConfigured, a
13140
13134
  return {
13141
13135
  text: "[Voice message - transcription unavailable]",
13142
13136
  source: "fallback",
13143
- asrReferText,
13144
13137
  remoteUrl
13145
13138
  };
13146
13139
  }
@@ -13184,23 +13177,18 @@ async function processVoiceAttachment(att, cfg, usePlatformAsr, sttConfigured, a
13184
13177
  const transcript = await transcribeAudioViaFramework(localPath, cfg);
13185
13178
  if (transcript) {
13186
13179
  log5?.debug?.(`Voice STT (framework): ${transcript.slice(0, 80)}...`);
13187
- return { text: transcript, source: "stt", duration, localPath, remoteUrl, asrReferText };
13180
+ return { text: transcript, source: "stt", duration, localPath, remoteUrl };
13188
13181
  }
13189
13182
  } catch (err) {
13190
13183
  log5?.error(`Voice STT (framework) failed: ${err instanceof Error ? err.message : String(err)}`);
13191
13184
  }
13192
13185
  }
13193
- if (asrReferText) {
13194
- log5?.debug?.(`Voice: falling back to platform asr_refer_text after framework STT failure`);
13195
- return { text: asrReferText, source: "asr", duration, localPath, remoteUrl, asrReferText };
13196
- }
13197
13186
  return {
13198
13187
  text: "[Voice message - transcription failed]",
13199
13188
  source: "fallback",
13200
13189
  duration,
13201
13190
  localPath,
13202
- remoteUrl,
13203
- asrReferText
13191
+ remoteUrl
13204
13192
  };
13205
13193
  }
13206
13194
  function resolveAudioPolicy(cfg) {
@@ -13322,7 +13310,7 @@ function buildDynamicCtx(processed, msg, quote) {
13322
13310
  lines.push(`- Voice: ${voiceRefs.join(", ")}`);
13323
13311
  }
13324
13312
  const asrTexts = unique(
13325
- transcripts.map((t) => t.source === "asr" ? t.text : t.asrReferText).filter(isNonEmpty)
13313
+ transcripts.filter((t) => t.source === "asr").map((t) => t.text).filter(isNonEmpty)
13326
13314
  );
13327
13315
  if (asrTexts.length > 0) {
13328
13316
  lines.push(`- ASR: ${asrTexts.join(" | ")}`);