@jerryliang122/openclaw-qqbot 2.0.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +17 -12
- package/README.zh.md +17 -12
- package/dist/index.cjs +27 -39
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +7 -19
- package/package.json +1 -1
- package/src/dispatch/body-assembler.ts +4 -2
- package/src/middleware/attachment.ts +14 -27
- package/src/types.ts +7 -19
- package/src/utils/stt.ts +42 -37
package/README.md
CHANGED
|
@@ -810,31 +810,36 @@ The plugin processes messages through a carefully ordered middleware chain:
|
|
|
810
810
|
|
|
811
811
|
#### STT (Speech-to-Text) — Transcribe Incoming Voice Messages
|
|
812
812
|
|
|
813
|
-
Transcription runs through the **framework audio-understanding pipeline** (`openclaw/plugin-sdk/media-understanding-runtime`), so STT
|
|
813
|
+
Transcription runs through the **framework audio-understanding pipeline** (`openclaw/plugin-sdk/media-understanding-runtime`), so STT config comes from the framework-level `tools.media.models` (same as the built-in Telegram channel) — the plugin no longer ships its own OpenAI-compatible HTTP call:
|
|
814
814
|
|
|
815
815
|
```json
|
|
816
816
|
{
|
|
817
817
|
"tools": {
|
|
818
818
|
"media": {
|
|
819
|
-
"
|
|
820
|
-
|
|
821
|
-
|
|
819
|
+
"models": [
|
|
820
|
+
{
|
|
821
|
+
"type": "cli",
|
|
822
|
+
"command": "your-asr-cli",
|
|
823
|
+
"args": ["-m", "your-model", "{{AttachmentPath}}"],
|
|
824
|
+
"maxBytes": 52428800,
|
|
825
|
+
"timeoutSeconds": 120,
|
|
826
|
+
"capabilities": ["audio"]
|
|
827
|
+
}
|
|
828
|
+
]
|
|
822
829
|
}
|
|
823
830
|
}
|
|
824
831
|
}
|
|
825
832
|
```
|
|
826
833
|
|
|
827
|
-
|
|
834
|
+
A provider-form entry also works (e.g. `{ "provider": "openai", "model": "whisper-1", "capabilities": ["audio"] }`). **`capabilities` must include `"audio"`** — the shared model list is selected by capability tags; untagged entries never take part in voice transcription.
|
|
828
835
|
|
|
829
|
-
|
|
830
|
-
2. **Framework STT not configured** → the **QQ platform transcript (`asr_refer_text`)** — QQ auto-STTs voice messages and ships the text in the event JSON — is used directly as the sole source, no download and no external call needed.
|
|
831
|
-
3. Neither available → placeholder text (`[Voice message - transcription unavailable]`); the audio URL is still referenced via the `- Voice:` line.
|
|
836
|
+
Voice message handling order (hardcoded since 2026-10; no plugin-level switches):
|
|
832
837
|
|
|
833
|
-
|
|
838
|
+
1. **Framework STT configured** (a `tools.media.models` entry whose `capabilities` includes `"audio"`) → voice is downloaded, converted (SILK→WAV), and transcribed through the framework pipeline. **The framework is trusted strictly** — on failure or empty transcript a placeholder (`[Voice message - transcription failed]`) is emitted; there is **no** fallback to the platform transcript.
|
|
839
|
+
2. **Framework STT not configured** → the **QQ platform transcript (`asr_refer_text`)** — QQ auto-STTs voice messages and ships the text in the event JSON — is used directly as the sole source, no download and no external call needed.
|
|
840
|
+
3. Not configured and no platform transcript → placeholder text (`[Voice message - transcription unavailable]`); the audio URL is still referenced via the `- Voice:` line.
|
|
834
841
|
|
|
835
|
-
|
|
836
|
-
- `asrFallback: false` — strict mode: discard the platform transcript in **all** cases (restores the pre-2026-10 behavior)
|
|
837
|
-
- `provider` / `baseUrl` / `apiKey` / `model` — **deprecated and ignored** (2026-10); a one-time migration notice is logged if still present. Move credentials to `tools.media.audio.models`.
|
|
842
|
+
The whole `channels.qqbot.stt` block is **deprecated and ignored** (legacy credential keys plus the historical `enabled`/`asrFallback` switches, removed 2026-10-04); a one-time migration notice is logged if any key is present. STT is controlled solely by the framework config — remove the audio entry from `tools.media.models` (or set `tools.media.audio.enabled: false`) to disable framework transcription.
|
|
838
843
|
|
|
839
844
|
#### TTS (Text-to-Speech) — Send Voice Messages
|
|
840
845
|
|
package/README.zh.md
CHANGED
|
@@ -631,31 +631,36 @@ openclaw message send --channel "qqbot" \
|
|
|
631
631
|
|
|
632
632
|
#### STT(语音转文字)— 自动转录用户发来的语音消息
|
|
633
633
|
|
|
634
|
-
转录统一走**框架音频理解管线**(`openclaw/plugin-sdk/media-understanding-runtime`),STT
|
|
634
|
+
转录统一走**框架音频理解管线**(`openclaw/plugin-sdk/media-understanding-runtime`),STT 配置只认框架级 `tools.media.models`(与内置 Telegram 通道一致)——插件不再自带 OpenAI 兼容 HTTP 调用:
|
|
635
635
|
|
|
636
636
|
```json
|
|
637
637
|
{
|
|
638
638
|
"tools": {
|
|
639
639
|
"media": {
|
|
640
|
-
"
|
|
641
|
-
|
|
642
|
-
|
|
640
|
+
"models": [
|
|
641
|
+
{
|
|
642
|
+
"type": "cli",
|
|
643
|
+
"command": "your-asr-cli",
|
|
644
|
+
"args": ["-m", "your-model", "{{AttachmentPath}}"],
|
|
645
|
+
"maxBytes": 52428800,
|
|
646
|
+
"timeoutSeconds": 120,
|
|
647
|
+
"capabilities": ["audio"]
|
|
648
|
+
}
|
|
649
|
+
]
|
|
643
650
|
}
|
|
644
651
|
}
|
|
645
652
|
}
|
|
646
653
|
```
|
|
647
654
|
|
|
648
|
-
|
|
655
|
+
也支持 provider 形态(如 `{ "provider": "openai", "model": "whisper-1", "capabilities": ["audio"] }`)。**`capabilities` 必须包含 `"audio"`**——模型列表按能力标签选择,无标签条目不参与语音转录。
|
|
649
656
|
|
|
650
|
-
|
|
651
|
-
2. **框架 STT 未配置** → **直接采用 QQ 平台转写**(`asr_refer_text`,QQ 平台对语音消息自动 STT 并随事件 JSON 下发)作为唯一来源——无需下载、不发起任何外部调用。
|
|
652
|
-
3. 两者都不可用 → 占位文本(`[Voice message - transcription unavailable]`),音频 URL 仍通过 `- Voice:` 行引用。
|
|
657
|
+
语音消息处理顺序(2026-10 起硬编码,无插件级开关):
|
|
653
658
|
|
|
654
|
-
|
|
659
|
+
1. **框架 STT 已配置**(`tools.media.models` 存在 `capabilities` 含 `"audio"` 的条目)→ 下载语音、转换(SILK→WAV)后经框架管线转录;**严格信框架**——转录失败或为空时输出占位文本(`[Voice message - transcription failed]`),不回退平台转写。
|
|
660
|
+
2. **框架 STT 未配置** → **直接采用 QQ 平台转写**(`asr_refer_text`,QQ 平台对语音消息自动 STT 并随事件 JSON 下发)作为唯一来源——无需下载、不发起任何外部调用。
|
|
661
|
+
3. 未配置且无平台转写 → 占位文本(`[Voice message - transcription unavailable]`),音频 URL 仍通过 `- Voice:` 行引用。
|
|
655
662
|
|
|
656
|
-
- `enabled: false`
|
|
657
|
-
- `asrFallback: false` — 严格模式:所有场景丢弃平台转写(恢复 2026-10 之前的旧行为)
|
|
658
|
-
- `provider` / `baseUrl` / `apiKey` / `model` — **已废弃并被忽略**(2026-10);检测到仍配置时会打一次性迁移提示日志,请把凭证迁移到 `tools.media.audio.models`。
|
|
663
|
+
`channels.qqbot.stt` 整块**已废弃并被忽略**(旧凭证键 + 历史 `enabled`/`asrFallback` 开关,2026-10-04 移除);检测到任何键时会打一次性迁移提示日志。STT 启停只由框架配置控制——删掉 `tools.media.models` 的 audio 条目、或设 `tools.media.audio.enabled: false` 即关闭框架转录。
|
|
659
664
|
|
|
660
665
|
#### TTS(文字转语音)— 机器人发送语音消息
|
|
661
666
|
|
package/dist/index.cjs
CHANGED
|
@@ -8259,7 +8259,7 @@ var init_outbound_service = __esm({
|
|
|
8259
8259
|
|
|
8260
8260
|
// src/utils/pkg-version.ts
|
|
8261
8261
|
function getPackageVersion() {
|
|
8262
|
-
return true ? "2.
|
|
8262
|
+
return true ? "2.1.0" : "unknown";
|
|
8263
8263
|
}
|
|
8264
8264
|
function getOpenclawVersion(runtimeVersion) {
|
|
8265
8265
|
return runtimeVersion && runtimeVersion !== "unknown" ? runtimeVersion : "unknown";
|
|
@@ -12906,33 +12906,32 @@ init_webhook_verify();
|
|
|
12906
12906
|
// src/utils/stt.ts
|
|
12907
12907
|
var path16 = __toESM(require("path"), 1);
|
|
12908
12908
|
var import_media_understanding_runtime = require("openclaw/plugin-sdk/media-understanding-runtime");
|
|
12909
|
-
function shouldUsePlatformAsr(cfg) {
|
|
12910
|
-
const channels = asRecord(cfg.channels);
|
|
12911
|
-
const qqbot = asRecord(channels?.qqbot);
|
|
12912
|
-
return asRecord(qqbot?.stt)?.asrFallback !== false;
|
|
12913
|
-
}
|
|
12914
12909
|
function isFrameworkSttConfigured(cfg) {
|
|
12915
|
-
const channels = asRecord(cfg.channels);
|
|
12916
|
-
const qqbot = asRecord(channels?.qqbot);
|
|
12917
|
-
if (asRecord(qqbot?.stt)?.enabled === false) {
|
|
12918
|
-
return false;
|
|
12919
|
-
}
|
|
12920
12910
|
const tools = asRecord(cfg.tools);
|
|
12921
12911
|
const media = asRecord(tools?.media);
|
|
12922
|
-
|
|
12923
|
-
if (!audio || audio.enabled === false) {
|
|
12912
|
+
if (asRecord(media?.audio)?.enabled === false) {
|
|
12924
12913
|
return false;
|
|
12925
12914
|
}
|
|
12926
|
-
|
|
12915
|
+
const models = media?.models;
|
|
12916
|
+
if (!Array.isArray(models)) {
|
|
12917
|
+
return false;
|
|
12918
|
+
}
|
|
12919
|
+
return models.some((entry) => {
|
|
12920
|
+
const capabilities = asRecord(entry)?.capabilities;
|
|
12921
|
+
return Array.isArray(capabilities) && capabilities.includes("audio");
|
|
12922
|
+
});
|
|
12927
12923
|
}
|
|
12928
|
-
function
|
|
12924
|
+
function hasLegacySttConfig(cfg) {
|
|
12929
12925
|
const channels = asRecord(cfg.channels);
|
|
12930
12926
|
const qqbot = asRecord(channels?.qqbot);
|
|
12931
12927
|
const stt = asRecord(qqbot?.stt);
|
|
12932
12928
|
if (!stt) return false;
|
|
12933
|
-
|
|
12934
|
-
|
|
12935
|
-
|
|
12929
|
+
const legacyKeys = ["provider", "baseUrl", "apiKey", "model", "enabled", "asrFallback"];
|
|
12930
|
+
return legacyKeys.some((key2) => {
|
|
12931
|
+
const value = stt[key2];
|
|
12932
|
+
if (typeof value === "string") return value.trim().length > 0;
|
|
12933
|
+
return value != null;
|
|
12934
|
+
});
|
|
12936
12935
|
}
|
|
12937
12936
|
async function transcribeAudioViaFramework(audioPath, cfg) {
|
|
12938
12937
|
const result = await (0, import_media_understanding_runtime.transcribeAudioFile)({
|
|
@@ -13023,9 +13022,8 @@ function attachmentProcessor(opts) {
|
|
|
13023
13022
|
};
|
|
13024
13023
|
}
|
|
13025
13024
|
async function processAttachments(attachments, cfg, log5) {
|
|
13026
|
-
const usePlatformAsr = shouldUsePlatformAsr(cfg);
|
|
13027
13025
|
const sttConfigured = isFrameworkSttConfigured(cfg);
|
|
13028
|
-
|
|
13026
|
+
warnLegacySttConfig(cfg, log5);
|
|
13029
13027
|
const audioPolicy = resolveAudioPolicy(cfg);
|
|
13030
13028
|
const imageUrls = [];
|
|
13031
13029
|
const otherParts = [];
|
|
@@ -13043,7 +13041,7 @@ async function processAttachments(attachments, cfg, log5) {
|
|
|
13043
13041
|
return { type: "image", localPath, url, contentType: att.content_type ?? "image/png", filename: att.filename };
|
|
13044
13042
|
}
|
|
13045
13043
|
if (isVoice) {
|
|
13046
|
-
const transcript = await processVoiceAttachment(att, cfg,
|
|
13044
|
+
const transcript = await processVoiceAttachment(att, cfg, sttConfigured, audioPolicy, log5);
|
|
13047
13045
|
return { type: "voice", transcript };
|
|
13048
13046
|
}
|
|
13049
13047
|
if (url) {
|
|
@@ -13117,22 +13115,18 @@ function kindFromContentType(contentType) {
|
|
|
13117
13115
|
return "document";
|
|
13118
13116
|
}
|
|
13119
13117
|
var legacySttWarned = false;
|
|
13120
|
-
function
|
|
13121
|
-
if (!legacySttWarned &&
|
|
13118
|
+
function warnLegacySttConfig(cfg, log5) {
|
|
13119
|
+
if (!legacySttWarned && hasLegacySttConfig(cfg)) {
|
|
13122
13120
|
legacySttWarned = true;
|
|
13123
13121
|
log5?.info(
|
|
13124
|
-
"Voice: channels.qqbot.stt
|
|
13122
|
+
"Voice: channels.qqbot.stt is deprecated and ignored entirely (credentials + enabled/asrFallback); configure an audio-capable tools.media.models entry for STT \u2014 platform asr_refer_text is used only when framework STT is absent"
|
|
13125
13123
|
);
|
|
13126
13124
|
}
|
|
13127
13125
|
}
|
|
13128
|
-
async function processVoiceAttachment(att, cfg,
|
|
13129
|
-
const rawAsrText = att.asr_refer_text?.trim() || void 0;
|
|
13130
|
-
const asrReferText = usePlatformAsr ? rawAsrText : void 0;
|
|
13126
|
+
async function processVoiceAttachment(att, cfg, sttConfigured, audioPolicy, log5) {
|
|
13131
13127
|
const remoteUrl = normalizeUrl(att.voice_wav_url) || normalizeUrl(att.url) || void 0;
|
|
13132
13128
|
if (!sttConfigured) {
|
|
13133
|
-
|
|
13134
|
-
log5?.info(`Voice: framework STT not configured; platform asr_refer_text discarded (asrFallback: false)`);
|
|
13135
|
-
}
|
|
13129
|
+
const asrReferText = att.asr_refer_text?.trim() || void 0;
|
|
13136
13130
|
if (asrReferText) {
|
|
13137
13131
|
log5?.debug?.(`Voice: using platform asr_refer_text (framework STT not configured)`);
|
|
13138
13132
|
return { text: asrReferText, source: "asr", asrReferText, remoteUrl };
|
|
@@ -13140,7 +13134,6 @@ async function processVoiceAttachment(att, cfg, usePlatformAsr, sttConfigured, a
|
|
|
13140
13134
|
return {
|
|
13141
13135
|
text: "[Voice message - transcription unavailable]",
|
|
13142
13136
|
source: "fallback",
|
|
13143
|
-
asrReferText,
|
|
13144
13137
|
remoteUrl
|
|
13145
13138
|
};
|
|
13146
13139
|
}
|
|
@@ -13184,23 +13177,18 @@ async function processVoiceAttachment(att, cfg, usePlatformAsr, sttConfigured, a
|
|
|
13184
13177
|
const transcript = await transcribeAudioViaFramework(localPath, cfg);
|
|
13185
13178
|
if (transcript) {
|
|
13186
13179
|
log5?.debug?.(`Voice STT (framework): ${transcript.slice(0, 80)}...`);
|
|
13187
|
-
return { text: transcript, source: "stt", duration, localPath, remoteUrl
|
|
13180
|
+
return { text: transcript, source: "stt", duration, localPath, remoteUrl };
|
|
13188
13181
|
}
|
|
13189
13182
|
} catch (err) {
|
|
13190
13183
|
log5?.error(`Voice STT (framework) failed: ${err instanceof Error ? err.message : String(err)}`);
|
|
13191
13184
|
}
|
|
13192
13185
|
}
|
|
13193
|
-
if (asrReferText) {
|
|
13194
|
-
log5?.debug?.(`Voice: falling back to platform asr_refer_text after framework STT failure`);
|
|
13195
|
-
return { text: asrReferText, source: "asr", duration, localPath, remoteUrl, asrReferText };
|
|
13196
|
-
}
|
|
13197
13186
|
return {
|
|
13198
13187
|
text: "[Voice message - transcription failed]",
|
|
13199
13188
|
source: "fallback",
|
|
13200
13189
|
duration,
|
|
13201
13190
|
localPath,
|
|
13202
|
-
remoteUrl
|
|
13203
|
-
asrReferText
|
|
13191
|
+
remoteUrl
|
|
13204
13192
|
};
|
|
13205
13193
|
}
|
|
13206
13194
|
function resolveAudioPolicy(cfg) {
|
|
@@ -13322,7 +13310,7 @@ function buildDynamicCtx(processed, msg, quote) {
|
|
|
13322
13310
|
lines.push(`- Voice: ${voiceRefs.join(", ")}`);
|
|
13323
13311
|
}
|
|
13324
13312
|
const asrTexts = unique(
|
|
13325
|
-
transcripts.
|
|
13313
|
+
transcripts.filter((t) => t.source === "asr").map((t) => t.text).filter(isNonEmpty)
|
|
13326
13314
|
);
|
|
13327
13315
|
if (asrTexts.length > 0) {
|
|
13328
13316
|
lines.push(`- ASR: ${asrTexts.join(" | ")}`);
|