dsh-audiogen 0.3.4 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "dsh-audiogen",
3
3
  "description": "AI audio generation plugin for the dsh web GUI: multi-vendor TTS/music/sound-effect channels (OpenAI-compatible, ElevenLabs, MiniMax, Stability AI and custom), per-channel model/voice catalogs, Agent tool and a sidebar AI 音频 panel.",
4
- "version": "0.3.4",
4
+ "version": "0.4.0",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",
7
7
  "exports": {
@@ -7,7 +7,13 @@
7
7
  - 帮助用户把音色或音效需求细化成可复用的描述,供后续 TTS/music/sfx 生成。
8
8
  - 输出不建议直接生成音频,而是给出音色参数建议和可用的模型/音色清单。
9
9
 
10
+ ## 面板/工具的音色设计(voice_design)
11
+ - 渠道支持:MiniMax(POST /v1/voice_design,prompt + preview_text)与 ElevenLabs(POST /v1/text-to-voice/design)。
12
+ - MiniMax 参数:`prompt`(音色描述)、`preview_text`(试听文本,任意长度)。
13
+ - ElevenLabs 参数:`voice_description`(音色描述);试听文本 `preview_text` 需 100-1000 字符,过短时自动 `auto_generate_text`;响应含 `previews[].audio_base_64` 与 `generated_voice_id`(可后续在 TTS 中复用该 voice_id)。
14
+ - 面板音色设计模式顶部可切换「厂商 / 渠道」。
15
+
10
16
  ## 流程
11
17
  1. 询问目标风格、音域、情绪、适用场景。
12
18
  2. 生成结构化的音色描述(如温暖复古合成器、未来感 UI 提示音)。
13
- 3. 在后续生成中复用该描述,必要时通过 `generate_audio` 试听。
19
+ 3. 在后续生成中复用该描述,必要时通过 `generate_audio`(mode=voice_design)试听。
@@ -37,3 +37,40 @@
37
37
  ## 常见错误
38
38
  - `lyrics-required`:MiniMax 音乐生成需要歌词,或开启纯音乐。
39
39
  - `HTTP 400` 且含 `2013`:上游参数不合法,检查 lyrics / audio_setting 枚举。
40
+
41
+ ## ElevenLabs Music(POST /v1/music,模型 music_v1 / music_v2)
42
+
43
+ | 字段 | 工具/面板参数 | 说明 |
44
+ | --- | --- | --- |
45
+ | model_id | model(music_v2 等) | 官方枚举:music_v1 / music_v2,默认 music_v1 |
46
+ | prompt | prompt | 音乐/歌词主题描述(不能与 composition_plan 同用;引擎用 prompt) |
47
+ | music_length_ms | duration(秒,自动×1000) | 3000ms - 600000ms(3s-600s),超出自动收敛区间 |
48
+ | lyrics_text | lyrics(歌词) | 选填歌词文本 |
49
+ | force_instrumental | is_instrumental(纯音乐) | true 保证无演唱(无词) |
50
+ | seed / generation_mode / finetune_* | — | 高级字段,暂未透出 |
51
+
52
+ > 响应为音频字节流(audio/*,常为 mp3)。请求同时携带 `xi-api-key` 与 `Authorization: Bearer`,以兼容 New API 类网关(官方站任一头即可)。
53
+
54
+ ## Stability Stable Audio(官方 v2beta,multipart/form-data)
55
+
56
+ 官方端点(文本到音频,TTS 描述 / 音乐 / 音效统一走该接口,不同模型参数不同):
57
+ - `POST /v2beta/audio/stable-audio/text-to-audio` → 模型 `stable-audio-3`(202 异步,随后轮询 GET /v2beta/audio/results/{id})
58
+ - `POST /v2beta/audio/stable-audio-2/text-to-audio` → 模型 `stable-audio-2` / `stable-audio-2.5`(同步返回音频)
59
+
60
+ | 字段 | 工具/面板参数 | 说明 |
61
+ | --- | --- | --- |
62
+ | prompt | prompt | 必填,描述性提示词(乐器/情绪/风格/体裁,≤10000 字符) |
63
+ | model | model | stable-audio-3 / stable-audio-2.5 / stable-audio-2 |
64
+ | duration | duration | 秒数:3 ≤380(默认 190);2/2.5 ≤190(默认 190) |
65
+ | seed | seed | 0-4294967294,默认 0=随机;同参数同 seed 可复现 |
66
+ | steps | steps | 采样步数:2 → 30-100(默认 50);2.5/3 → 4-8(默认 8) |
67
+ | cfg_scale | cfg_scale | 1-25:2 默认 7,2.5/3 默认 1;越高越贴提示词 |
68
+ | output_format | format | mp3 / wav |
69
+
70
+ > 引擎按模型自动收敛步数/时长区间;渠道 preset/apiUrl 含 `stability` 或模型名以 `stable-audio-` 开头即走官方协议(自定义渠道同样适用)。
71
+
72
+ ### 双通道(自动选择)
73
+ - **官方 v2beta**:apiUrl 为 `https://api.stability.ai`(含 `/v2beta`、`/v2beta/audio` 形态)→ multipart 原生端点(2/2.5 同步、3 异步轮询)。
74
+ - **OpenAI 兼容网关**:apiUrl 以 `/v1` 结尾或含 `/audio/speech`(如 New API)→ `POST {apiUrl}/audio/speech`,JSON:
75
+ `{ "model": "stable-audio-2.5", "input": "<prompt>", "output_format": "mp3", "duration": 30, "seed": 0, "steps": 8, "cfg_scale": 1 }`(网关把该模型映射到 Stable 上游)。
76
+ - 一方返回 `404 Invalid URL`(未路由)时自动换另一方重试;参数在两种通道均按模型收敛(duration/seed/steps/cfg_scale/output_format)。
@@ -2,15 +2,29 @@
2
2
 
3
3
  ## 触发
4
4
  - `/audio:sfx <描述>`
5
- - 用户说“生成一个音效 / 提示音 / 環境音”
5
+ - 用户说“生成一个音效 / 提示音 / 环境音”
6
6
 
7
7
  ## 参数
8
8
  - prompt: 必填,音效描述
9
- - model: 可选,已配置模型
10
- - duration: 可选,秒数
9
+ - model: 可选,已配置模型(ElevenLabs:eleven_text_to_sound_v2)
10
+ - duration: 可选,秒数(ElevenLabs 0.5-30,留空则自动)
11
+ - loop: 可选,ElevenLabs 无缝循环音效(仅 eleven_text_to_sound_v2)
12
+ - prompt_influence: 可选,ElevenLabs 提示词影响度 0-1(默认 0.3,越高越贴近提示词、越少随机)
11
13
  - format: 可选
12
14
 
13
15
  ## 流程
14
16
  1. 确认已配置音效生成渠道。
15
17
  2. 调用 `generate_audio`,mode=sfx。
16
18
  3. 将音频 URL 返回。
19
+
20
+ ## ElevenLabs Sound Generation(POST /v1/sound-generation)
21
+
22
+ | 字段 | 工具/面板参数 | 说明 |
23
+ | --- | --- | --- |
24
+ | text | prompt | 必填,转换为音效的文本/描述 |
25
+ | model_id | model | 官方枚举:eleven_text_to_sound_v2(默认) |
26
+ | duration_seconds | duration | 0.5-30 秒;留空由提示词推算最优时长 |
27
+ | loop | loop | 是否生成平滑循环音效(仅 eleven_text_to_sound_v2) |
28
+ | prompt_influence | prompt_influence | 0-1,默认 0.3;越高越贴提示词,越低越多样 |
29
+
30
+ > 响应为 audio/mpeg 二进制。请求同时携带 `xi-api-key` 与 `Authorization: Bearer`,兼容 New API 类网关。
@@ -9,14 +9,15 @@ import { defineTool } from '@deepseek-ai/dsh-tools'
9
9
  import { randomUUID } from 'node:crypto'
10
10
  import type { AudioChannel } from './audio-engine.ts'
11
11
  import { generateAudio, AudioGenError } from './audio-engine.ts'
12
- import { appendHistory, saveAudioFile } from './audio-store.ts'
13
- import type { AudioMode, GenerateAudioRequest } from './protocol.ts'
12
+ import { appendHistory, saveAudioFile, saveToLibrary, listLibrary } from './audio-store.ts'
13
+ import type { AudioMode, GenerateAudioRequest, LibraryType } from './protocol.ts'
14
14
 
15
15
  export interface AgentAudioToolConfig {
16
16
  enabled: boolean
17
17
  allowAgentAudioGeneration: boolean
18
18
  channels: AudioChannel[]
19
19
  defaultChannelId: string
20
+ autoSaveToLibrary: boolean
20
21
  }
21
22
 
22
23
  interface AgentAudioRef {
@@ -27,12 +28,19 @@ interface AgentAudioRef {
27
28
  voiceId?: string
28
29
  }
29
30
 
31
+ /** Internal: the persisted file name, needed for library copies. */
32
+ interface SavedAudioRef extends AgentAudioRef {
33
+ file: string
34
+ }
35
+
30
36
  interface AgentAudioResult {
31
37
  status: string
32
38
  message: string
33
39
  mode: AudioMode
34
40
  model: string
35
41
  audio: AgentAudioRef[]
42
+ /** Resource-library entry ids when the audio was saved to the library. */
43
+ resources?: string[]
36
44
  error?: string
37
45
  }
38
46
 
@@ -57,6 +65,7 @@ const resultSchema = {
57
65
  mode: { type: 'string', required: true, enum: ['tts', 'music', 'sfx', 'voice_design'] },
58
66
  model: { type: 'string', required: true },
59
67
  audio: { type: 'array', required: true, items: audioRefSchema },
68
+ resources: { type: 'array', items: { type: 'string' } },
60
69
  error: { type: 'string' },
61
70
  },
62
71
  } as const
@@ -96,11 +105,17 @@ function ensureConfigured(config: AgentAudioToolConfig): void {
96
105
  if (!usable) throw new AudioGenError('Audio API credentials are not configured. Open Settings > Plugins > AI Audio, add a channel and fill its API URL and API key.', 'audio-api-not-configured')
97
106
  }
98
107
 
108
+ /** Library type from the generation mode, with an explicit override. */
109
+ function libraryTypeOf(mode: AudioMode, override: unknown): LibraryType {
110
+ if (override === 'voice' || override === 'music' || override === 'sfx' || override === 'tts') return override
111
+ if (mode === 'voice_design') return 'voice'
112
+ return mode
113
+ }
114
+
99
115
  /** Register the Agent audio tool. */
100
116
  export function registerAgentAudioTools(ctx: Context, resolve: () => AgentAudioToolConfig): () => void {
101
117
  const disposer = ctx.tools.register(defineTool({
102
- name: 'generate_audio',
103
- description: 'Generate audio with the configured audio provider. Supports text-to-speech, music generation, sound effects and MiniMax voice design. The tool call waits for the upstream result and returns same-origin audio URLs; pass those URLs to the user for playback or download. If multiple models are configured, first ask the user which one to use or pass model explicitly.',
118
+ name: 'generate_audio', description: 'Generate audio with the configured audio provider. Supports text-to-speech, music generation, sound effects and voice design (MiniMax /v1/voice_design, ElevenLabs /v1/text-to-voice/design). The tool call waits for the upstream result and returns same-origin audio URLs; pass those URLs to the user for playback or download. If multiple models are configured, first ask the user which one to use or pass model explicitly.',
104
119
  parameters: {
105
120
  prompt: { type: 'string', required: true, description: 'For tts, the text to speak. For music/sfx, a descriptive prompt.' },
106
121
  mode: { type: 'string', enum: ['tts', 'music', 'sfx', 'voice_design'], description: 'Generation mode. Defaults to tts.' },
@@ -111,6 +126,11 @@ export function registerAgentAudioTools(ctx: Context, resolve: () => AgentAudioT
111
126
  duration: { type: 'number', description: 'Requested duration in seconds for music/sfx.' },
112
127
  lyrics: { type: 'string', description: 'Lyrics for music generation (MiniMax music-3.0/music-cover). Required unless is_instrumental is true. Split verses with an empty line.' },
113
128
  is_instrumental: { type: 'boolean', description: 'Generate purely instrumental music without vocals/lyrics (MiniMax is_instrumental). When true, lyrics may be omitted.' },
129
+ loop: { type: 'boolean', description: 'Create a seamlessly looping sound effect (ElevenLabs sound generation loop, only for eleven_text_to_sound_v2).' },
130
+ prompt_influence: { type: 'number', description: 'Sound effect prompt influence 0-1 (ElevenLabs prompt_influence, default 0.3): higher follows the prompt more closely, lower is more variable.' },
131
+ seed: { type: 'integer', description: 'Stable Audio random seed 0-4294967294 (default 0 = random); same seed yields reproducible audio.' },
132
+ steps: { type: 'integer', description: 'Stable Audio sampling steps, model-dependent: stable-audio-2 30-100, stable-audio-2.5/3 4-8 (out-of-range auto-clamped).' },
133
+ cfg_scale: { type: 'number', description: 'Stable Audio prompt adherence 1-25 (stable-audio-2 default 7, 2.5/3 default 1); higher follows the prompt more strictly.' },
114
134
  format: { type: 'string', description: 'Output format such as mp3 or wav. MiniMax music supports mp3/wav/pcm.' },
115
135
  // ---- MiniMax TTS only (ignored by other providers) ----
116
136
  emotion: { type: 'string', description: 'MiniMax TTS emotion, e.g. happy/sad/angry/nervous/fearful/bored (voice_setting.emotion).' },
@@ -150,6 +170,11 @@ export function registerAgentAudioTools(ctx: Context, resolve: () => AgentAudioT
150
170
  },
151
171
  description: 'MiniMax TTS dual-voice blend weights (timbre_weights).',
152
172
  },
173
+ // ---- resource library ----
174
+ save_to_library: { type: 'boolean', description: 'Save the generated audio into the local resource library after success. Also enabled globally by the "auto save to library" setting; pass false to skip a single run.' },
175
+ library_name: { type: 'string', description: 'Resource name in the library. Defaults to the prompt.' },
176
+ library_type: { type: 'string', enum: ['voice', 'music', 'sfx', 'tts'], description: 'Resource type in the library. Defaults to the generation mode (voice_design → voice).' },
177
+ library_tags: { type: 'array', items: { type: 'string' }, description: 'Tags for the library resource.' },
153
178
  },
154
179
  output: {
155
180
  schema: resultSchema,
@@ -199,6 +224,11 @@ export function registerAgentAudioTools(ctx: Context, resolve: () => AgentAudioT
199
224
  ...(typeof args.duration === 'number' ? { duration: args.duration } : {}),
200
225
  ...(typeof args.lyrics === 'string' && args.lyrics.trim() !== '' ? { lyrics: args.lyrics.trim() } : {}),
201
226
  ...(typeof args.is_instrumental === 'boolean' ? { isInstrumental: args.is_instrumental } : {}),
227
+ ...(typeof args.loop === 'boolean' ? { loop: args.loop } : {}),
228
+ ...(typeof args.prompt_influence === 'number' && Number.isFinite(args.prompt_influence) ? { promptInfluence: args.prompt_influence } : {}),
229
+ ...(typeof args.seed === 'number' && Number.isFinite(args.seed) ? { seed: args.seed } : {}),
230
+ ...(typeof args.steps === 'number' && Number.isFinite(args.steps) ? { steps: args.steps } : {}),
231
+ ...(typeof args.cfg_scale === 'number' && Number.isFinite(args.cfg_scale) ? { cfgScale: args.cfg_scale } : {}),
202
232
  ...(typeof args.format === 'string' && args.format.trim() !== '' ? { format: args.format.trim() } : {}),
203
233
  // ---- MiniMax TTS 专属字段 ----
204
234
  ...(typeof args.emotion === 'string' && args.emotion.trim() !== '' ? { emotion: args.emotion.trim() } : {}),
@@ -222,13 +252,22 @@ export function registerAgentAudioTools(ctx: Context, resolve: () => AgentAudioT
222
252
  try {
223
253
  const outputs = await generateAudio(picked.channel, request, exec.signal)
224
254
  const audio: AgentAudioRef[] = []
255
+ const saved: SavedAudioRef[] = []
225
256
  for (const [index, output] of outputs.entries()) {
226
- const saved = await saveAudioFile(output.data, output.mime, `generated-${index + 1}`)
257
+ const stored = await saveAudioFile(output.data, output.mime, `generated-${index + 1}`)
258
+ saved.push({
259
+ id: stored.id,
260
+ url: `/api/dsh-audiogen/audio/${encodeURIComponent(stored.file)}`,
261
+ file: stored.file,
262
+ mime: stored.mime,
263
+ bytes: stored.bytes,
264
+ ...(output.voiceId === undefined ? {} : { voiceId: output.voiceId }),
265
+ })
227
266
  audio.push({
228
- id: saved.id,
229
- url: `/api/dsh-audiogen/audio/${encodeURIComponent(saved.file)}`,
230
- mime: saved.mime,
231
- bytes: saved.bytes,
267
+ id: stored.id,
268
+ url: `/api/dsh-audiogen/audio/${encodeURIComponent(stored.file)}`,
269
+ mime: stored.mime,
270
+ bytes: stored.bytes,
232
271
  ...(output.voiceId === undefined ? {} : { voiceId: output.voiceId }),
233
272
  })
234
273
  }
@@ -244,25 +283,60 @@ export function registerAgentAudioTools(ctx: Context, resolve: () => AgentAudioT
244
283
  ...(request.duration === undefined ? {} : { duration: request.duration }),
245
284
  ...(request.format === undefined ? {} : { format: request.format }),
246
285
  audio: outputs.map((output, index) => ({
247
- id: audio[index]!.id,
286
+ id: saved[index]!.id,
287
+ file: saved[index]!.file,
248
288
  b64: Buffer.from(output.data).toString('base64'),
249
- mime: audio[index]!.mime,
250
- bytes: audio[index]!.bytes,
251
- url: audio[index]!.url,
289
+ mime: saved[index]!.mime,
290
+ bytes: saved[index]!.bytes,
291
+ url: saved[index]!.url,
252
292
  ...(output.voiceId === undefined ? {} : { voiceId: output.voiceId }),
253
293
  })),
254
294
  channelId: picked.channel.id,
255
295
  channel: picked.channel.name,
296
+ params: { ...request },
256
297
  })
257
298
  } catch {
258
299
  // History is best-effort and must not fail the agent tool.
259
300
  }
301
+ // ---- 资源库保存:显式参数优先;设置自动入库时可用 false 跳过 ----
302
+ const wantSave = args.save_to_library === true || (config.autoSaveToLibrary && args.save_to_library !== false)
303
+ let resources: string[] | undefined
304
+ if (wantSave) {
305
+ try {
306
+ const entry = await saveToLibrary({
307
+ audioFiles: saved.map(item => ({
308
+ id: item.id,
309
+ file: item.file,
310
+ mime: item.mime,
311
+ ...(item.voiceId === undefined ? {} : { voiceId: item.voiceId }),
312
+ })),
313
+ type: libraryTypeOf(request.mode, args.library_type),
314
+ ...(typeof args.library_name === 'string' && args.library_name.trim() !== '' ? { name: args.library_name.trim() } : {}),
315
+ ...(Array.isArray(args.library_tags) ? { tags: args.library_tags.filter((tag): tag is string => typeof tag === 'string' && tag.trim() !== '').map(tag => tag.trim()) } : {}),
316
+ provenance: {
317
+ mode: request.mode,
318
+ prompt: request.prompt,
319
+ channel: picked.channel.name,
320
+ channelId: picked.channel.id,
321
+ apiUrl: picked.channel.apiUrl,
322
+ model: picked.alias,
323
+ upstream: picked.upstream,
324
+ ...(request.voice === undefined ? {} : { voice: request.voice }),
325
+ params: { ...request },
326
+ },
327
+ })
328
+ resources = [entry.id]
329
+ } catch {
330
+ // library-save is best-effort; generation already succeeded.
331
+ }
332
+ }
260
333
  return {
261
334
  status: 'completed',
262
335
  message: 'Audio generation completed. The audio files can be played/downloaded from the returned URLs.',
263
336
  mode: request.mode,
264
337
  model: picked.alias,
265
338
  audio,
339
+ ...(resources === undefined ? {} : { resources }),
266
340
  }
267
341
  } catch (error) {
268
342
  if (exec.signal?.aborted === true) throw error
@@ -277,5 +351,76 @@ export function registerAgentAudioTools(ctx: Context, resolve: () => AgentAudioT
277
351
  }
278
352
  },
279
353
  }))
280
- return disposer
354
+
355
+ const searchDisposer = ctx.tools.register(defineTool({
356
+ name: 'search_audio_library',
357
+ description: 'Search curated audio resources in the local resource library (voice / music / sfx / tts). Returns matching resources with type, category, name, tags, full provenance (channel, model, voiceId, prompt) and same-origin audio URLs the user can play. Use it before generating to reuse an existing voice, music bed or sound effect instead of generating a new one.',
358
+ parameters: {
359
+ type: { type: 'string', enum: ['voice', 'music', 'sfx', 'tts'], description: 'Filter by resource type.' },
360
+ category: { type: 'string', description: 'Filter by category (voice: male/female/custom; tts: the speaking voice key).' },
361
+ keyword: { type: 'string', description: 'Search name, tags, prompt and model.' },
362
+ },
363
+ output: {
364
+ schema: {
365
+ type: 'object',
366
+ additionalProperties: false,
367
+ properties: {
368
+ status: { type: 'string', required: true, enum: ['ok'] },
369
+ count: { type: 'integer', required: true },
370
+ entries: {
371
+ type: 'array', required: true,
372
+ items: {
373
+ type: 'object',
374
+ additionalProperties: false,
375
+ properties: {
376
+ id: { type: 'string', required: true },
377
+ name: { type: 'string', required: true },
378
+ type: { type: 'string', required: true, enum: ['voice', 'music', 'sfx', 'tts'] },
379
+ category: { type: 'string' },
380
+ tags: { type: 'array', items: { type: 'string' }, required: true },
381
+ prompt: { type: 'string', required: true },
382
+ model: { type: 'string' },
383
+ channel: { type: 'string' },
384
+ voiceId: { type: 'string' },
385
+ urls: { type: 'array', items: { type: 'string' }, required: true },
386
+ },
387
+ },
388
+ },
389
+ },
390
+ } as const,
391
+ render: (_args, value) => [{ type: 'text', text: JSON.stringify(value) }],
392
+ },
393
+ isConcurrencySafe: () => true,
394
+ async execute(args) {
395
+ const keyword = typeof args.keyword === 'string' ? args.keyword.trim().toLowerCase() : ''
396
+ const wantedType = args.type === 'voice' || args.type === 'music' || args.type === 'sfx' || args.type === 'tts' ? args.type : undefined
397
+ const wantedCategory = typeof args.category === 'string' && args.category.trim() !== '' ? args.category.trim() : undefined
398
+ const all = await listLibrary()
399
+ const entries = all.filter(entry => {
400
+ if (wantedType !== undefined && entry.type !== wantedType) return false
401
+ if (wantedCategory !== undefined && (entry.category ?? '') !== wantedCategory) return false
402
+ if (keyword !== '') {
403
+ const haystack = [entry.name, ...entry.tags, entry.provenance.prompt, entry.provenance.model ?? '', entry.provenance.channel ?? ''].join(' ').toLowerCase()
404
+ if (!haystack.includes(keyword)) return false
405
+ }
406
+ return true
407
+ }).slice(0, 30).map(entry => ({
408
+ id: entry.id,
409
+ name: entry.name,
410
+ type: entry.type,
411
+ ...(entry.category === undefined ? {} : { category: entry.category }),
412
+ tags: entry.tags,
413
+ prompt: entry.provenance.prompt,
414
+ ...(entry.provenance.model === undefined ? {} : { model: entry.provenance.model }),
415
+ ...(entry.provenance.channel === undefined ? {} : { channel: entry.provenance.channel }),
416
+ ...(entry.provenance.voiceId === undefined ? {} : { voiceId: entry.provenance.voiceId }),
417
+ urls: entry.files.map(file => file.url),
418
+ }))
419
+ return { status: 'ok' as const, count: entries.length, entries }
420
+ },
421
+ }))
422
+ return () => {
423
+ disposer()
424
+ searchDisposer()
425
+ }
281
426
  }