dsh-audiogen 0.4.11 → 0.4.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # 🎧 dsh-audiogen
2
2
 
3
- **AI audio generation for DeepSeek Harness (DSH)** — turn your DSH web GUI into an audio studio: text-to-speech, music, sound effects and voice design, from a sidebar panel or straight from the Agent.
3
+ **AI audio generation for DeepSeek Harness (DSH)** — turn your DSH web GUI into an audio studio: text-to-speech, music, sound effects and voice design, from the sidebar panel or straight from the Agent.
4
4
 
5
5
  [English](README.md) | [简体中文](README.zh-CN.md)
6
6
 
@@ -9,29 +9,29 @@
9
9
  ![DSH plugin](https://img.shields.io/badge/DSH-plugin-brightgreen?style=flat-square)
10
10
  ![Node](https://img.shields.io/badge/node-%3E%3D20-blue?style=flat-square)
11
11
 
12
- ![Main panel](docs/images/panel-compare.png)
12
+ ![Main panel](docs/images/hero.png)
13
13
 
14
14
  ## ✨ Features
15
15
 
16
- - **Four generation modes**: 文本转语音 (TTS) · 音乐生成 (Music) · 音效生成 (Sound effects) · 音色设计 (Voice design)
16
+ - **Four generation modes**: text-to-speech, music, sound effects, and voice design
17
17
  - **Multi-vendor channels in one place**: OpenAI-compatible TTS, MiniMax, ElevenLabs, Stability AI, or any custom OpenAI-compatible / generic POST endpoint
18
- - **Per-channel model & voice catalogs** with one-click discovery, aliases, capability categories and per-model advanced fields (duration, seed, steps, cfg_scale, loop, prompt influence, …)
19
- - **Model comparison**: same prompt across 2–4 models with per-model parameter overrides — results grouped side by side
20
- - **✨ Prompt enhancement**: rewrite any rough idea into a ready-to-generate description with an LLM (choose any model from *Settings → Models*; falls back to the agent default model)
18
+ - **Per-channel model & voice catalogs** with one-click discovery, display aliases, capability categories, and per-model advanced fields (duration, seed, steps, cfg_scale, loop, prompt influence, …)
19
+ - **Model comparison**: run the same prompt across 2–4 models at once with per-model parameter overrides — results are grouped side by side
20
+ - **Prompt enhancement**: rewrite a rough idea into a ready-to-generate description with an LLM (pick any model from *Settings → Models*; falls back to the agent default model)
21
21
  - **History with one-click restore**: prompt, config, model set *and the original audio* come back into the panel — no regeneration, no extra cost
22
- - **Resource library**: auto-save generated audio (or opt-in per run), organized by type — voices / music / sfx / TTS — with search, tags, rename, category moves, and full provenance (channel, model, voiceId, prompt, params snapshot). Reuse a voice or music bed instead of regenerating
22
+ - **Resource library**: auto-save generated audio (or opt in per run), organized by type — voices / music / SFX / TTS — with search, tags, rename, category moves, and full provenance (channel, model, voice id, prompt, params snapshot). Reuse a voice or music bed instead of regenerating
23
23
  - **Agent tools**: `generate_audio` and `search_audio_library`, plus bundled session skills — the Agent can generate and find audio on demand
24
24
  - **Keys stay local**: API keys live in the local DSH settings document and generation is proxied by the local host; the browser and the Agent never touch plaintext credentials
25
25
 
26
26
  ## 📸 Screenshots
27
27
 
28
- | Generation panel | Model comparison |
28
+ | Generation panel | Resource library |
29
29
  | --- | --- |
30
- | ![Generation](docs/images/panel-music.png) | ![Compare](docs/images/panel-compare.png) |
30
+ | ![Generation](docs/images/hero.png) | ![Library](docs/images/library.png) |
31
31
 
32
- | History & restore | Channels settings |
32
+ | Library full provenance drawer | Channels settings |
33
33
  | --- | --- |
34
- | ![History](docs/images/history-panel.png) | ![Settings](docs/images/settings-channels.png) |
34
+ | ![Library detail](docs/images/library-detail.png) | ![Settings](docs/images/settings-channels.png) |
35
35
 
36
36
  | Channel editor (model catalog & auto capabilities) | LLM models (Settings → Models) |
37
37
  | --- | --- |
@@ -51,18 +51,18 @@ Local development install:
51
51
  dsh plugin --profile web add /path/to/dsh-audiogen
52
52
  ```
53
53
 
54
- Restart `dsh web` after install — the sidebar will show **AI 音频**.
54
+ Restart `dsh web` after install — the sidebar will show the **AI Audio** entry.
55
55
 
56
56
  ## 🚀 Quick start
57
57
 
58
- 1. Open **Settings → Plugins → AI 音频**
58
+ 1. Open **Settings → Plugins → AI Audio**
59
59
  2. Add a channel: pick a preset provider (+ Add provider) or a custom endpoint (+ Add custom provider)
60
60
  3. Fill in the API URL, API key, and the model/voice catalog (use *Fetch available models* to import them)
61
- 4. Save, then open the **AI 音频** sidebar panel:
62
- - choose a mode (TTS / Music / SFX / Voice design)
63
- - type your text or prompt (optional: ✨ 增强提示词)
64
- - pick a model — or tick 模型对比 for 2–4 models at once
65
- - hit **开始生成** and play the results, download them, or add them to the 资源库
61
+ 4. Save, then open the **AI Audio** sidebar panel:
62
+ - choose a mode (Speech / Music / Sound effects / Voice design)
63
+ - type your text or prompt (optional: ✨ Enhance prompt)
64
+ - pick a model — or tick **Model comparison** for 2–4 models at once
65
+ - press **Start generation** and play the results, download them, or add them to the resource library
66
66
 
67
67
  ## 🎛 Modes supported by each vendor
68
68
 
@@ -70,7 +70,7 @@ Restart `dsh web` after install — the sidebar will show **AI 音频**.
70
70
  | --- | --- | --- | --- | --- |
71
71
  | TTS | ✅ (8 voices) | ✅ (voices + streams) | — | ✅ |
72
72
  | Music | ✅ (`music-3.0` / `music-2.6` / `music-cover`) | ✅ (`music_v2`) | ✅ (`stable-audio-*`) | ✅ (generic POST) |
73
- | Sound effects | — | ✅ (`eleven_text_to_sound_v2`, loop / prompt_influence) | ✅ (`stable-audio-*` — same text-to-audio protocol, auto-detected in music + SFX) | ✅ (generic POST) |
73
+ | Sound effects | — | ✅ (`eleven_text_to_sound_v2`, loop / prompt influence) | ✅ (`stable-audio-*` — same text-to-audio protocol; auto-detected in both Music and SFX) | ✅ (generic POST) |
74
74
  | Voice design | ✅ (`/v1/voice_design`) | ✅ (`/v1/text-to-voice/design`) | — | — |
75
75
 
76
76
  ## 🤖 Agent usage
@@ -83,10 +83,10 @@ Restart `dsh web` after install — the sidebar will show **AI 音频**.
83
83
  Typical session commands (skills bundled with the plugin):
84
84
 
85
85
  ```text
86
- /audio:tts 用温暖的声音朗读这句话
87
- /audio:music 生成一段 30 秒的 Lo-fi 背景音乐
88
- /audio:sfx 生成一声科幻风格的 UI 提示音
89
- /audio:design 一个温暖复古的合成器音色
86
+ /audio:tts Read this sentence with a warm voice
87
+ /audio:music Generate a 30-second lo-fi background track
88
+ /audio:sfx Create a sci-fi UI cue
89
+ /audio:design Craft a warm retro synth voice
90
90
  ```
91
91
 
92
92
  ## 🔐 Security & data notes
package/README.zh-CN.md CHANGED
@@ -9,7 +9,7 @@
9
9
  ![DSH plugin](https://img.shields.io/badge/DSH-plugin-brightgreen?style=flat-square)
10
10
  ![Node](https://img.shields.io/badge/node-%3E%3D20-blue?style=flat-square)
11
11
 
12
- ![主面板](docs/images/panel-compare.png)
12
+ ![主面板](docs/images/hero.png)
13
13
 
14
14
  ## ✨ 功能特性
15
15
 
@@ -25,13 +25,13 @@
25
25
 
26
26
  ## 📸 截图
27
27
 
28
- | 生成面板 | 模型对比 |
28
+ | 生成面板 | 资源库 |
29
29
  | --- | --- |
30
- | ![生成](docs/images/panel-music.png) | ![对比](docs/images/panel-compare.png) |
30
+ | ![生成](docs/images/hero.png) | ![资源库](docs/images/library.png) |
31
31
 
32
- | 历史记录与恢复 | 渠道设置 |
32
+ | 资源库详情(完整溯源) | 渠道设置 |
33
33
  | --- | --- |
34
- | ![历史](docs/images/history-panel.png) | ![设置](docs/images/settings-channels.png) |
34
+ | ![资源库详情](docs/images/library-detail.png) | ![设置](docs/images/settings-channels.png) |
35
35
 
36
36
  | 渠道编辑(模型目录与自动能力识别) | LLM 模型(设置 → 模型) |
37
37
  | --- | --- |
Binary file
Binary file
Binary file
Binary file
Binary file
Binary file