eleven_rb 0.4.0 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +50 -0
- data/README.md +125 -4
- data/lib/eleven_rb/callbacks.rb +25 -1
- data/lib/eleven_rb/client.rb +8 -0
- data/lib/eleven_rb/configuration.rb +7 -3
- data/lib/eleven_rb/http/client.rb +32 -12
- data/lib/eleven_rb/model_capabilities.rb +103 -0
- data/lib/eleven_rb/objects/audio.rb +13 -4
- data/lib/eleven_rb/objects/cost_info.rb +8 -3
- data/lib/eleven_rb/objects/model.rb +10 -0
- data/lib/eleven_rb/objects/voice_settings.rb +33 -0
- data/lib/eleven_rb/resources/base.rb +18 -0
- data/lib/eleven_rb/resources/models.rb +37 -4
- data/lib/eleven_rb/resources/text_to_dialogue.rb +203 -0
- data/lib/eleven_rb/resources/text_to_speech.rb +150 -43
- data/lib/eleven_rb/version.rb +1 -1
- data/lib/eleven_rb.rb +3 -0
- metadata +7 -5
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 18424bceaba545d05f991cb1ae674fdc36572a7a9573ef769701deb35537b393
|
|
4
|
+
data.tar.gz: a5c2a3814295f5d19945b1731b639fe21ae0f0111ef51741fdb99bae5a978cf7
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: af3da66572e56a9173b5f427a3fa30212ebd686a5820b060ab5c81e5e41d60ae64a74e7956e0ddad4365d0f8e68ceefaa5572d97f485a0e424609c195f34be31
|
|
7
|
+
data.tar.gz: 3af1d07b75e5405724c50e87d3b89f88061d3540043716b38242acce0941f5152294148cd1d9109c20f05bbd05f5131e6b1225260bd373371b45061ad4bbca26
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,56 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [1.1.0] - 2026-09-30
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- Eleven v4 support (`eleven_v4`, `eleven_v4_turbo`) across Text-to-Speech and Text-to-Dialogue
|
|
15
|
+
- `ElevenRb::ModelCapabilities` — per model-family table of honoured voice settings, max text length, SSML `<break>` support, audio-tag support and continuity support (`for`, `supported_voice_settings`, `max_text_length`, `supports?`)
|
|
16
|
+
- `Objects::VoiceSettings.for_model(model_id, overrides)` → `[settings, dropped_keys]`, building only the settings a model honours
|
|
17
|
+
- `Configuration#strict_voice_settings` (default `false`): raise `ValidationError` instead of warning when a voice setting the model ignores is passed
|
|
18
|
+
- Optional TTS keywords on `generate`, `stream` and `generate_with_timestamps` (omitted from the body when nil): `language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`
|
|
19
|
+
- Response metadata: `Objects::Audio#request_id`, `#character_cost` (from the `request-id` / `character-cost` headers) and `#dropped_settings`; `on_audio_generated` also receives `request_id:` (nil for streams)
|
|
20
|
+
- `TextToSpeech#generate_with_timestamps` now also returns `normalized_alignment`, `request_id` and `character_cost`
|
|
21
|
+
- `TextToDialogue#generate_with_timestamps` (`POST /v1/text-to-dialogue/with-timestamps`) returning `audio`, `alignment`, `normalized_alignment`, `voice_segments` and `request_id`
|
|
22
|
+
- Text-to-Dialogue keywords `use_pvc_as_ivc`, `previous_text`, `future_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`; a logger warning above `RECOMMENDED_MAX_TEXT_LENGTH` (2,000 characters)
|
|
23
|
+
- `HTTP::Client#post(..., with_meta: true)` returning `{ body:, headers: }`, and `Resources::Base#post_with_meta` / `#post_binary_with_meta`
|
|
24
|
+
- `Models#find(model_id)`
|
|
25
|
+
- `Objects::Model#model_rates`, `#maximum_text_length_per_request`, `#requires_alpha_access`, `#supported_voice_settings`
|
|
26
|
+
- `CostInfo::COST_PER_1K_CHARS` entries for `eleven_v4` ($0.30), `eleven_v4_turbo` ($0.15) and `eleven_v3_conversational` ($0.15)
|
|
27
|
+
- `opus_*` output formats map to the `ogg` extension and `audio/ogg` content type
|
|
28
|
+
|
|
29
|
+
### Changed
|
|
30
|
+
|
|
31
|
+
- Voice settings are filtered per model: known keys a model ignores (`stability`, `similarity_boost`, `style`, `use_speaker_boost`, `speed` minus the model's supported set) are dropped from the request (logged, and listed on `audio.dropped_settings`); unknown voice-setting keys pass through untouched on every model, so a new API field is never swallowed. `eleven_v3` / `eleven_v4` now send only `stability` and `similarity_boost` by default; `eleven_multilingual_v2` requests are byte-identical to 1.0.0. Override keys are symbolized and nil values removed (a nil override is never reported as dropped)
|
|
32
|
+
- Text length is capped per model (`ModelCapabilities.max_text_length`): 5,000 for `eleven_v3`, 10,000 for `eleven_v4` and `eleven_multilingual_v2`, 30,000 / 40,000 for the flash and turbo models. `TextToSpeech::MAX_TEXT_LENGTH` stays defined but is no longer the cap
|
|
33
|
+
- `TextToDialogue::DEFAULT_MODEL` is now `eleven_v4`
|
|
34
|
+
- Text-to-Dialogue's hard text cap now follows the model (10,000 characters on `eleven_v4`, 5,000 on `eleven_v3`) instead of a flat 5,000; `TextToDialogue::MAX_TEXT_LENGTH` stays defined but is no longer the cap
|
|
35
|
+
- `Models#latest` returns `eleven_v4`, else `eleven_v3`, else the default model
|
|
36
|
+
- `TextToSpeech::OUTPUT_FORMATS` refreshed to the current list (documentation only, not validated)
|
|
37
|
+
- Callbacks that declare their keywords explicitly (no `**rest`) receive only the keywords they declare, so callbacks written for 1.0.0 keep working as new keywords are added
|
|
38
|
+
|
|
39
|
+
### Fixed
|
|
40
|
+
|
|
41
|
+
- `Models#list` (and everything built on it: `get`, `default`, `latest`, `multilingual`, `turbo`, `tts_capable`, `ids`, `TTSAdapter#list_models`) recursed until `SystemStackError`, because `Models#get(model_id)` shadowed `Resources::Base#get`
|
|
42
|
+
|
|
43
|
+
## [1.0.0] - 2026-03-10
|
|
44
|
+
|
|
45
|
+
### Added
|
|
46
|
+
|
|
47
|
+
- Text-to-Dialogue multi-speaker audio generation via `client.text_to_dialogue.generate` (`POST /v1/text-to-dialogue`)
|
|
48
|
+
- `Client#text_to_dialogue` resource with `dialogue` alias
|
|
49
|
+
- Multi-speaker input validation (max 10 unique voices, 5000 character limit)
|
|
50
|
+
- `eleven_v3` model added to `CostInfo::COST_PER_1K_CHARS` ($0.30/1K chars)
|
|
51
|
+
- `Models#latest` method returning the most capable model (`eleven_v3`)
|
|
52
|
+
- Audio tags support via v3 model (`[laughs]`, `[whispers]`, `[excited]`, etc.)
|
|
53
|
+
- `CostInfo` now accepts `character_count:` keyword as alternative to `text:`
|
|
54
|
+
- TTS generation with word-level timestamps via `client.tts.generate_with_timestamps`
|
|
55
|
+
|
|
56
|
+
### Changed
|
|
57
|
+
|
|
58
|
+
- `CostInfo#initialize` signature: `text:` is now optional when `character_count:` is provided (backwards-compatible)
|
|
59
|
+
|
|
10
60
|
## [0.4.0] - 2026-03-10
|
|
11
61
|
|
|
12
62
|
### Added
|
data/README.md
CHANGED
|
@@ -4,12 +4,14 @@
|
|
|
4
4
|
[](https://github.com/webventures/eleven_rb/actions/workflows/ci.yml)
|
|
5
5
|
[](https://opensource.org/licenses/MIT)
|
|
6
6
|
|
|
7
|
-
A Ruby client for the [ElevenLabs](https://try.elevenlabs.io/qyk2j8gumrjz) Text-to-Speech, Speech-to-Speech, Sound Effects, and Music API.
|
|
7
|
+
A Ruby client for the [ElevenLabs](https://try.elevenlabs.io/qyk2j8gumrjz) Text-to-Speech, Speech-to-Speech, Text-to-Dialogue, Sound Effects, and Music API.
|
|
8
8
|
|
|
9
9
|
## Features
|
|
10
10
|
|
|
11
11
|
- Text-to-Speech generation and streaming
|
|
12
12
|
- Speech-to-Speech voice conversion
|
|
13
|
+
- Text-to-Dialogue multi-speaker generation with audio tags
|
|
14
|
+
- Eleven v4 support, with per-model voice-settings filtering and text caps
|
|
13
15
|
- Sound effects generation from text descriptions
|
|
14
16
|
- Music generation from prompts or composition plans
|
|
15
17
|
- Voice management (list, get, create, update, delete)
|
|
@@ -73,7 +75,7 @@ audio.save_to_file("output.mp3")
|
|
|
73
75
|
audio = client.tts.generate(
|
|
74
76
|
"Hello world",
|
|
75
77
|
voice_id: "voice_id",
|
|
76
|
-
model_id: "
|
|
78
|
+
model_id: "eleven_v4", # Most expressive; audio tags; see "Eleven v4" below
|
|
77
79
|
voice_settings: {
|
|
78
80
|
stability: 0.5,
|
|
79
81
|
similarity_boost: 0.75
|
|
@@ -87,8 +89,72 @@ File.open("output.mp3", "wb") do |file|
|
|
|
87
89
|
file.write(chunk)
|
|
88
90
|
end
|
|
89
91
|
end
|
|
92
|
+
|
|
93
|
+
# Word-level timestamps
|
|
94
|
+
result = client.tts.generate_with_timestamps("Hello world", voice_id: "voice_id")
|
|
95
|
+
result[:audio] # => ElevenRb::Objects::Audio
|
|
96
|
+
result[:alignment] # => { "characters" => [...], "character_start_times_seconds" => [...], ... }
|
|
97
|
+
result[:normalized_alignment]
|
|
98
|
+
result[:request_id] # from the request-id response header
|
|
99
|
+
result[:character_cost] # Integer, from the character-cost response header
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Optional keywords on `generate`, `stream` and `generate_with_timestamps` (each is left out of the request when nil):
|
|
103
|
+
`language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`,
|
|
104
|
+
`next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`.
|
|
105
|
+
|
|
106
|
+
### Eleven v4
|
|
107
|
+
|
|
108
|
+
`eleven_v4` (and the cheaper, faster `eleven_v4_turbo`) take up to 10,000 characters per request.
|
|
109
|
+
|
|
110
|
+
```ruby
|
|
111
|
+
audio = client.tts.generate(
|
|
112
|
+
"[sighs] Right... let's try that one more time.",
|
|
113
|
+
voice_id: "voice_id",
|
|
114
|
+
model_id: "eleven_v4",
|
|
115
|
+
voice_settings: { stability: 0.5, similarity_boost: 0.75 },
|
|
116
|
+
seed: 42, # reproducible re-rolls
|
|
117
|
+
previous_text: "That did not go to plan.", # continuity with the take before
|
|
118
|
+
next_text: "Here we go." # ...and the one after
|
|
119
|
+
)
|
|
120
|
+
audio.request_id # pass as previous_request_ids: / next_request_ids: on neighbouring takes
|
|
121
|
+
audio.character_cost
|
|
122
|
+
audio.dropped_settings # => [] (voice settings the gem removed, see below)
|
|
90
123
|
```
|
|
91
124
|
|
|
125
|
+
- **Only `stability` and `similarity_boost` take effect.** The API accepts `speed`, `style` and
|
|
126
|
+
`use_speaker_boost` on v4 and silently ignores them, so the gem drops them from the request, logs a
|
|
127
|
+
warning through `logger`, and lists them on `audio.dropped_settings`. Set
|
|
128
|
+
`strict_voice_settings: true` on the client to raise `ElevenRb::Errors::ValidationError` instead.
|
|
129
|
+
Only those known keys are ever dropped: a voice-setting key the gem does not know passes through
|
|
130
|
+
unchanged on every model, so new API fields keep working.
|
|
131
|
+
Pace a v4 take with the text (ellipses, pause tags), not `speed`.
|
|
132
|
+
- **SSML `<break time="…"/>` tags are ignored** by v4 (no pause is produced). Use an audio tag or
|
|
133
|
+
punctuation for pauses.
|
|
134
|
+
- **Audio tags** go in square brackets inside the text: `[sighs]`, `[whispers]`, `[laughs]`,
|
|
135
|
+
`[short pause]`. Note that tags appear in the returned alignment like any other characters.
|
|
136
|
+
- **Continuity**: `previous_text` / `next_text` (or `previous_request_ids` / `next_request_ids`) tell the
|
|
137
|
+
model what surrounds a take, so stitched takes keep a consistent delivery.
|
|
138
|
+
- **`seed`** makes v4 generation reproducible, so a re-roll changes one thing at a time.
|
|
139
|
+
|
|
140
|
+
`ElevenRb::ModelCapabilities` answers these questions for any model id:
|
|
141
|
+
|
|
142
|
+
```ruby
|
|
143
|
+
ElevenRb::ModelCapabilities.supported_voice_settings("eleven_v4") # => [:stability, :similarity_boost]
|
|
144
|
+
ElevenRb::ModelCapabilities.max_text_length("eleven_v3") # => 5000
|
|
145
|
+
ElevenRb::ModelCapabilities.supports?("eleven_v4", :audio_tags) # => true
|
|
146
|
+
ElevenRb::ModelCapabilities.supports?("eleven_v4", :ssml_break) # => false
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
| Model family | Voice settings honoured | Max chars | Audio tags | SSML breaks |
|
|
150
|
+
|---|---|---|---|---|
|
|
151
|
+
| `eleven_v4`, `eleven_v4_turbo` | stability, similarity_boost | 10,000 | yes | no |
|
|
152
|
+
| `eleven_v3` | stability, similarity_boost | 5,000 | yes | no |
|
|
153
|
+
| `eleven_v3_conversational` | stability, similarity_boost, use_speaker_boost | 5,000 | yes | no |
|
|
154
|
+
| `eleven_flash_v2_5`, `eleven_turbo_v2_5` | all five (incl. style, speed) | 40,000 | no | yes |
|
|
155
|
+
| `eleven_flash_v2`, `eleven_turbo_v2` | all five | 30,000 | no | yes |
|
|
156
|
+
| `eleven_multilingual_v2` and others | all five | 10,000 | no | yes |
|
|
157
|
+
|
|
92
158
|
### Speech-to-Speech
|
|
93
159
|
|
|
94
160
|
```ruby
|
|
@@ -111,6 +177,54 @@ io = File.open("input.mp3", "rb")
|
|
|
111
177
|
audio = client.sts.convert(io, voice_id: "voice_id")
|
|
112
178
|
```
|
|
113
179
|
|
|
180
|
+
### Text-to-Dialogue
|
|
181
|
+
|
|
182
|
+
```ruby
|
|
183
|
+
# Generate multi-speaker dialogue
|
|
184
|
+
audio = client.text_to_dialogue.generate([
|
|
185
|
+
{ text: "[excited] Welcome to the show!", voice_id: "voice_abc" },
|
|
186
|
+
{ text: "[laughs] Thanks for having me.", voice_id: "voice_xyz" },
|
|
187
|
+
{ text: "So tell us about your project...", voice_id: "voice_abc" }
|
|
188
|
+
])
|
|
189
|
+
audio.save_to_file("dialogue.mp3")
|
|
190
|
+
|
|
191
|
+
# With options (the default model is eleven_v4)
|
|
192
|
+
audio = client.dialogue.generate(
|
|
193
|
+
inputs,
|
|
194
|
+
model_id: "eleven_v4",
|
|
195
|
+
language_code: "en",
|
|
196
|
+
settings: { stability: 0.5, similarity: 0.75 }, # sent to the API unchanged
|
|
197
|
+
seed: 42,
|
|
198
|
+
previous_text: "Earlier in the scene...",
|
|
199
|
+
output_format: "mp3_44100_192"
|
|
200
|
+
)
|
|
201
|
+
|
|
202
|
+
# With timestamps and per-speaker segments
|
|
203
|
+
result = client.dialogue.generate_with_timestamps(inputs)
|
|
204
|
+
result[:audio] # => ElevenRb::Objects::Audio
|
|
205
|
+
result[:alignment] # character timings
|
|
206
|
+
result[:voice_segments] # which voice speaks when
|
|
207
|
+
result[:request_id]
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
The hard text cap follows the model (10,000 characters for `eleven_v4`, 5,000 for `eleven_v3`); above
|
|
211
|
+
2,000 characters the gem logs a warning, since shorter dialogue requests give more reliable results.
|
|
212
|
+
|
|
213
|
+
### Audio Tags
|
|
214
|
+
|
|
215
|
+
The `eleven_v4` and `eleven_v3` models support inline audio tags, in square brackets, for expressive speech
|
|
216
|
+
(`ElevenRb::ModelCapabilities.supports?(model_id, :audio_tags)`):
|
|
217
|
+
|
|
218
|
+
```ruby
|
|
219
|
+
audio = client.tts.generate(
|
|
220
|
+
"[excited] Oh wow, this is AMAZING! [laughs] I can't believe it...",
|
|
221
|
+
voice_id: "voice_id",
|
|
222
|
+
model_id: "eleven_v4"
|
|
223
|
+
)
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
Supported tags include `[laughs]`, `[whispers]`, `[sighs]`, `[excited]`, `[sarcastic]`, `[curious]`, `[short pause]`, and more. Use CAPS for emphasis, `...` for pauses, and `—` for interruptions. See the [ElevenLabs audio tags documentation](https://elevenlabs.io/docs/guides/audio-tags) for the full list.
|
|
227
|
+
|
|
114
228
|
### Sound Effects
|
|
115
229
|
|
|
116
230
|
```ruby
|
|
@@ -252,7 +366,7 @@ client = ElevenRb::Client.new(
|
|
|
252
366
|
Sentry.capture_exception(error, extra: { path: path })
|
|
253
367
|
},
|
|
254
368
|
|
|
255
|
-
# Cost tracking
|
|
369
|
+
# Cost tracking (add request_id: to also receive the API's request id)
|
|
256
370
|
on_audio_generated: ->(audio:, voice_id:, text:, cost_info:) {
|
|
257
371
|
UsageRecord.create!(
|
|
258
372
|
characters: cost_info[:character_count],
|
|
@@ -274,6 +388,12 @@ client = ElevenRb::Client.new(
|
|
|
274
388
|
models = client.models.list
|
|
275
389
|
models.each { |m| puts "#{m.name} (#{m.model_id})" }
|
|
276
390
|
|
|
391
|
+
# Find one model
|
|
392
|
+
client.models.find("eleven_v4") # => ElevenRb::Objects::Model (get is an alias)
|
|
393
|
+
|
|
394
|
+
# Get the latest/most capable model: eleven_v4, else eleven_v3, else the default
|
|
395
|
+
client.models.latest.model_id # => "eleven_v4"
|
|
396
|
+
|
|
277
397
|
# Get multilingual models
|
|
278
398
|
client.models.multilingual
|
|
279
399
|
|
|
@@ -303,7 +423,8 @@ client = ElevenRb::Client.new(
|
|
|
303
423
|
open_timeout: 10, # Connection timeout
|
|
304
424
|
max_retries: 3, # Max retry attempts
|
|
305
425
|
retry_delay: 1.0, # Base delay between retries
|
|
306
|
-
logger: Rails.logger
|
|
426
|
+
logger: Rails.logger, # Optional logger (receives dropped-setting warnings)
|
|
427
|
+
strict_voice_settings: false # true: raise instead of dropping settings a model ignores
|
|
307
428
|
)
|
|
308
429
|
```
|
|
309
430
|
|
data/lib/eleven_rb/callbacks.rb
CHANGED
|
@@ -28,6 +28,10 @@ module ElevenRb
|
|
|
28
28
|
|
|
29
29
|
# Trigger a callback if it's configured
|
|
30
30
|
#
|
|
31
|
+
# A callback that names its keywords explicitly (no `**rest`) receives only
|
|
32
|
+
# the keywords it declares, so callbacks written before a keyword was added
|
|
33
|
+
# (e.g. `request_id:` on on_audio_generated in 1.1.0) keep working.
|
|
34
|
+
#
|
|
31
35
|
# @param callback_name [Symbol] the name of the callback
|
|
32
36
|
# @param kwargs [Hash] keyword arguments to pass to the callback
|
|
33
37
|
# @return [Object, nil] the return value of the callback, or nil
|
|
@@ -36,12 +40,32 @@ module ElevenRb
|
|
|
36
40
|
return unless callback.respond_to?(:call)
|
|
37
41
|
|
|
38
42
|
begin
|
|
39
|
-
callback.call(**kwargs)
|
|
43
|
+
callback.call(**accepted_callback_kwargs(callback, kwargs))
|
|
40
44
|
rescue StandardError => e
|
|
41
45
|
# Don't let callback errors break the main flow
|
|
42
46
|
warn "[ElevenRb] Callback error in #{callback_name}: #{e.message}"
|
|
43
47
|
nil
|
|
44
48
|
end
|
|
45
49
|
end
|
|
50
|
+
|
|
51
|
+
private
|
|
52
|
+
|
|
53
|
+
def accepted_callback_kwargs(callback, kwargs)
|
|
54
|
+
params = callback_parameters(callback)
|
|
55
|
+
return kwargs if params.nil? || params.any? { |type, _| type == :keyrest }
|
|
56
|
+
|
|
57
|
+
accepted = params.filter_map { |type, name| name if %i[key keyreq].include?(type) }
|
|
58
|
+
return kwargs if accepted.empty?
|
|
59
|
+
|
|
60
|
+
kwargs.slice(*accepted)
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
def callback_parameters(callback)
|
|
64
|
+
return callback.parameters if callback.respond_to?(:parameters)
|
|
65
|
+
|
|
66
|
+
callback.method(:call).parameters
|
|
67
|
+
rescue NameError
|
|
68
|
+
nil
|
|
69
|
+
end
|
|
46
70
|
end
|
|
47
71
|
end
|
data/lib/eleven_rb/client.rb
CHANGED
|
@@ -101,6 +101,14 @@ module ElevenRb
|
|
|
101
101
|
@music ||= Resources::Music.new(http_client)
|
|
102
102
|
end
|
|
103
103
|
|
|
104
|
+
# Text-to-dialogue resource
|
|
105
|
+
#
|
|
106
|
+
# @return [Resources::TextToDialogue]
|
|
107
|
+
def text_to_dialogue
|
|
108
|
+
@text_to_dialogue ||= Resources::TextToDialogue.new(http_client)
|
|
109
|
+
end
|
|
110
|
+
alias dialogue text_to_dialogue
|
|
111
|
+
|
|
104
112
|
# Voice slot manager
|
|
105
113
|
#
|
|
106
114
|
# @return [VoiceSlotManager]
|
|
@@ -21,12 +21,13 @@ module ElevenRb
|
|
|
21
21
|
open_timeout: 10,
|
|
22
22
|
max_retries: 3,
|
|
23
23
|
retry_delay: 1.0,
|
|
24
|
-
retry_statuses: [429, 500, 502, 503, 504].freeze
|
|
24
|
+
retry_statuses: [429, 500, 502, 503, 504].freeze,
|
|
25
|
+
strict_voice_settings: false
|
|
25
26
|
}.freeze
|
|
26
27
|
|
|
27
28
|
attr_accessor :api_key, :base_url, :timeout, :open_timeout,
|
|
28
29
|
:max_retries, :retry_delay, :retry_statuses,
|
|
29
|
-
:logger
|
|
30
|
+
:logger, :strict_voice_settings
|
|
30
31
|
|
|
31
32
|
# Initialize a new configuration
|
|
32
33
|
#
|
|
@@ -39,6 +40,8 @@ module ElevenRb
|
|
|
39
40
|
# @option options [Float] :retry_delay Base delay between retries in seconds (default: 1.0)
|
|
40
41
|
# @option options [Array<Integer>] :retry_statuses HTTP status codes to retry (default: [429, 500, 502, 503, 504])
|
|
41
42
|
# @option options [Logger] :logger Logger instance for debug output
|
|
43
|
+
# @option options [Boolean] :strict_voice_settings Raise instead of warn when a voice setting the
|
|
44
|
+
# model does not honour is passed (default: false)
|
|
42
45
|
# @option options [Proc] :on_request Callback before each request
|
|
43
46
|
# @option options [Proc] :on_response Callback after successful response
|
|
44
47
|
# @option options [Proc] :on_error Callback when an error occurs
|
|
@@ -87,7 +90,8 @@ module ElevenRb
|
|
|
87
90
|
timeout: timeout,
|
|
88
91
|
open_timeout: open_timeout,
|
|
89
92
|
max_retries: max_retries,
|
|
90
|
-
retry_delay: retry_delay
|
|
93
|
+
retry_delay: retry_delay,
|
|
94
|
+
strict_voice_settings: strict_voice_settings
|
|
91
95
|
}
|
|
92
96
|
end
|
|
93
97
|
end
|
|
@@ -32,9 +32,11 @@ module ElevenRb
|
|
|
32
32
|
# @param path [String] the API path
|
|
33
33
|
# @param body [Hash] request body
|
|
34
34
|
# @param response_type [Symbol] :json or :binary
|
|
35
|
-
# @
|
|
36
|
-
|
|
37
|
-
|
|
35
|
+
# @param with_meta [Boolean] when true, return `{ body:, headers: }` instead of the body alone
|
|
36
|
+
# (headers as a Hash of lower-cased String keys to String values)
|
|
37
|
+
# @return [Hash, Array, String] parsed response (or `{ body:, headers: }` with with_meta)
|
|
38
|
+
def post(path, body = {}, response_type: :json, with_meta: false)
|
|
39
|
+
request(:post, path, body: body, response_type: response_type, with_meta: with_meta)
|
|
38
40
|
end
|
|
39
41
|
|
|
40
42
|
# Make a DELETE request
|
|
@@ -68,7 +70,7 @@ module ElevenRb
|
|
|
68
70
|
private
|
|
69
71
|
|
|
70
72
|
def request(method, path, body: nil, params: nil, response_type: :json, multipart: false, stream: false,
|
|
71
|
-
attempt: 1, &block)
|
|
73
|
+
attempt: 1, with_meta: false, &block)
|
|
72
74
|
config.validate!
|
|
73
75
|
url = "#{config.base_url}#{path}"
|
|
74
76
|
start_time = Time.now
|
|
@@ -86,15 +88,12 @@ module ElevenRb
|
|
|
86
88
|
# Trigger after response callback
|
|
87
89
|
config.trigger(:on_response, method: method, path: path, response: response, duration: duration)
|
|
88
90
|
|
|
89
|
-
|
|
90
|
-
return response.body if response_type == :binary && response.success?
|
|
91
|
-
|
|
92
|
-
handle_response(response)
|
|
91
|
+
build_result(response, response_type, with_meta)
|
|
93
92
|
rescue Errors::RateLimitError => e
|
|
94
93
|
config.trigger(:on_rate_limit, retry_after: e.retry_after, error: e)
|
|
95
|
-
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
|
|
94
|
+
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
|
|
96
95
|
rescue Errors::ServerError => e
|
|
97
|
-
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
|
|
96
|
+
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
|
|
98
97
|
rescue Errors::Base => e
|
|
99
98
|
config.trigger(:on_error, error: e, method: method, path: path,
|
|
100
99
|
context: { body: sanitize_body_for_logging(body) })
|
|
@@ -107,6 +106,17 @@ module ElevenRb
|
|
|
107
106
|
end
|
|
108
107
|
end
|
|
109
108
|
|
|
109
|
+
def build_result(response, response_type, with_meta)
|
|
110
|
+
# Return binary data directly
|
|
111
|
+
result = if response_type == :binary && response.success?
|
|
112
|
+
response.body
|
|
113
|
+
else
|
|
114
|
+
handle_response(response)
|
|
115
|
+
end
|
|
116
|
+
|
|
117
|
+
with_meta ? { body: result, headers: normalize_headers(response) } : result
|
|
118
|
+
end
|
|
119
|
+
|
|
110
120
|
def execute_request(method, url, body, params, multipart, stream, &block)
|
|
111
121
|
options = build_options(body, params, multipart, stream, &block)
|
|
112
122
|
|
|
@@ -210,7 +220,7 @@ module ElevenRb
|
|
|
210
220
|
raise error_class.new(message, **error_kwargs)
|
|
211
221
|
end
|
|
212
222
|
|
|
213
|
-
def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, &block)
|
|
223
|
+
def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
|
|
214
224
|
raise error if attempt > config.max_retries || !config.retry_statuses.include?(error.http_status)
|
|
215
225
|
|
|
216
226
|
delay = if error.is_a?(Errors::RateLimitError) && error.retry_after
|
|
@@ -224,7 +234,17 @@ module ElevenRb
|
|
|
224
234
|
sleep(delay)
|
|
225
235
|
|
|
226
236
|
request(method, path,
|
|
227
|
-
body: body, params: params, response_type: response_type, multipart: multipart, stream: stream,
|
|
237
|
+
body: body, params: params, response_type: response_type, multipart: multipart, stream: stream,
|
|
238
|
+
attempt: attempt + 1, with_meta: with_meta, &block)
|
|
239
|
+
end
|
|
240
|
+
|
|
241
|
+
# Response headers as a plain Hash of lower-cased String keys to String values
|
|
242
|
+
# (HTTParty exposes multi-value arrays; the first value is kept)
|
|
243
|
+
def normalize_headers(response)
|
|
244
|
+
raw = response.headers.respond_to?(:to_hash) ? response.headers.to_hash : response.headers.to_h
|
|
245
|
+
raw.each_with_object({}) do |(key, value), out|
|
|
246
|
+
out[key.to_s.downcase] = value.is_a?(Array) ? value.first : value
|
|
247
|
+
end
|
|
228
248
|
end
|
|
229
249
|
|
|
230
250
|
def wrap_error(error)
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module ElevenRb
|
|
4
|
+
# What each ElevenLabs model family accepts, keyed by model-id family.
|
|
5
|
+
#
|
|
6
|
+
# The API silently ignores voice settings a model does not honour (Eleven v4
|
|
7
|
+
# accepts `speed`, `style` and `use_speaker_boost` and does nothing with them),
|
|
8
|
+
# so the gem uses this table to send only the settings that take effect, to
|
|
9
|
+
# cap text length per model, and to answer feature questions (audio tags,
|
|
10
|
+
# SSML `<break>` tags, request continuity via previous_text / next_text).
|
|
11
|
+
#
|
|
12
|
+
# @example
|
|
13
|
+
# ElevenRb::ModelCapabilities.supported_voice_settings('eleven_v4')
|
|
14
|
+
# # => [:stability, :similarity_boost]
|
|
15
|
+
# ElevenRb::ModelCapabilities.max_text_length('eleven_flash_v2_5') # => 40_000
|
|
16
|
+
# ElevenRb::ModelCapabilities.supports?('eleven_v4', :audio_tags) # => true
|
|
17
|
+
module ModelCapabilities
|
|
18
|
+
# Capability record for one model family
|
|
19
|
+
Capabilities = Struct.new(:voice_settings, :max_text_length, :ssml_break, :audio_tags, :continuity,
|
|
20
|
+
keyword_init: true) do
|
|
21
|
+
def ssml_break? = ssml_break
|
|
22
|
+
def audio_tags? = audio_tags
|
|
23
|
+
def continuity? = continuity
|
|
24
|
+
end
|
|
25
|
+
|
|
26
|
+
ALL_VOICE_SETTINGS = %i[stability similarity_boost style use_speaker_boost speed].freeze
|
|
27
|
+
|
|
28
|
+
def self.build(voice_settings:, max_text_length:, ssml_break:, audio_tags:, continuity:)
|
|
29
|
+
Capabilities.new(
|
|
30
|
+
voice_settings: voice_settings.freeze,
|
|
31
|
+
max_text_length: max_text_length,
|
|
32
|
+
ssml_break: ssml_break,
|
|
33
|
+
audio_tags: audio_tags,
|
|
34
|
+
continuity: continuity
|
|
35
|
+
).freeze
|
|
36
|
+
end
|
|
37
|
+
private_class_method :build
|
|
38
|
+
|
|
39
|
+
V4 = build(voice_settings: %i[stability similarity_boost], max_text_length: 10_000,
|
|
40
|
+
ssml_break: false, audio_tags: true, continuity: true)
|
|
41
|
+
V3_CONVERSATIONAL = build(voice_settings: %i[stability similarity_boost use_speaker_boost], max_text_length: 5_000,
|
|
42
|
+
ssml_break: false, audio_tags: true, continuity: true)
|
|
43
|
+
V3 = build(voice_settings: %i[stability similarity_boost], max_text_length: 5_000,
|
|
44
|
+
ssml_break: false, audio_tags: true, continuity: true)
|
|
45
|
+
V2_5_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 40_000,
|
|
46
|
+
ssml_break: true, audio_tags: false, continuity: true)
|
|
47
|
+
V2_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 30_000,
|
|
48
|
+
ssml_break: true, audio_tags: false, continuity: true)
|
|
49
|
+
DEFAULT = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 10_000,
|
|
50
|
+
ssml_break: true, audio_tags: false, continuity: true)
|
|
51
|
+
|
|
52
|
+
# Ordered [matcher, capabilities] pairs; the first match wins.
|
|
53
|
+
FAMILIES = [
|
|
54
|
+
[/\Aeleven_v4/, V4],
|
|
55
|
+
[/\Aeleven_v3_conversational/, V3_CONVERSATIONAL],
|
|
56
|
+
[/\Aeleven_v3/, V3],
|
|
57
|
+
[/\Aeleven_(flash|turbo)_v2_5/, V2_5_FAST],
|
|
58
|
+
[/\Aeleven_(flash|turbo)_v2/, V2_FAST]
|
|
59
|
+
].freeze
|
|
60
|
+
|
|
61
|
+
FEATURES = %i[audio_tags ssml_break continuity].freeze
|
|
62
|
+
|
|
63
|
+
module_function
|
|
64
|
+
|
|
65
|
+
# Capabilities for a model id (unknown ids get the permissive default)
|
|
66
|
+
#
|
|
67
|
+
# @param model_id [String, Symbol, nil]
|
|
68
|
+
# @return [Capabilities] frozen
|
|
69
|
+
def for(model_id)
|
|
70
|
+
id = model_id.to_s
|
|
71
|
+
FAMILIES.each { |matcher, caps| return caps if matcher.match?(id) }
|
|
72
|
+
DEFAULT
|
|
73
|
+
end
|
|
74
|
+
|
|
75
|
+
# Voice settings keys the model honours
|
|
76
|
+
#
|
|
77
|
+
# @param model_id [String]
|
|
78
|
+
# @return [Array<Symbol>]
|
|
79
|
+
def supported_voice_settings(model_id)
|
|
80
|
+
self.for(model_id).voice_settings
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
# Maximum characters per request for the model
|
|
84
|
+
#
|
|
85
|
+
# @param model_id [String]
|
|
86
|
+
# @return [Integer]
|
|
87
|
+
def max_text_length(model_id)
|
|
88
|
+
self.for(model_id).max_text_length
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
# Whether the model supports a feature
|
|
92
|
+
#
|
|
93
|
+
# @param model_id [String]
|
|
94
|
+
# @param feature [Symbol] :audio_tags, :ssml_break or :continuity
|
|
95
|
+
# @return [Boolean]
|
|
96
|
+
def supports?(model_id, feature)
|
|
97
|
+
feature = feature.to_sym
|
|
98
|
+
raise ArgumentError, "Unknown feature #{feature.inspect} (expected one of #{FEATURES.join(', ')})" unless FEATURES.include?(feature)
|
|
99
|
+
|
|
100
|
+
self.for(model_id).public_send(feature) ? true : false
|
|
101
|
+
end
|
|
102
|
+
end
|
|
103
|
+
end
|
|
@@ -6,7 +6,8 @@ module ElevenRb
|
|
|
6
6
|
module Objects
|
|
7
7
|
# Represents generated audio data
|
|
8
8
|
class Audio
|
|
9
|
-
attr_reader :data, :format, :voice_id, :text, :model_id
|
|
9
|
+
attr_reader :data, :format, :voice_id, :text, :model_id,
|
|
10
|
+
:request_id, :character_cost, :dropped_settings
|
|
10
11
|
|
|
11
12
|
# Initialize audio object
|
|
12
13
|
#
|
|
@@ -15,12 +16,20 @@ module ElevenRb
|
|
|
15
16
|
# @param voice_id [String] the voice ID used
|
|
16
17
|
# @param text [String] the text that was converted
|
|
17
18
|
# @param model_id [String, nil] the model ID used
|
|
18
|
-
|
|
19
|
+
# @param request_id [String, nil] the API's `request-id` response header (usable as
|
|
20
|
+
# previous_request_ids / next_request_ids on a later request)
|
|
21
|
+
# @param character_cost [Integer, nil] the API's `character-cost` response header
|
|
22
|
+
# @param dropped_settings [Array<Symbol>] voice settings the gem dropped because the model ignores them
|
|
23
|
+
def initialize(data:, format:, voice_id:, text:, model_id: nil, request_id: nil, character_cost: nil,
|
|
24
|
+
dropped_settings: [])
|
|
19
25
|
@data = data
|
|
20
26
|
@format = format
|
|
21
27
|
@voice_id = voice_id
|
|
22
28
|
@text = text
|
|
23
29
|
@model_id = model_id
|
|
30
|
+
@request_id = request_id
|
|
31
|
+
@character_cost = character_cost
|
|
32
|
+
@dropped_settings = Array(dropped_settings).dup.freeze
|
|
24
33
|
end
|
|
25
34
|
|
|
26
35
|
# Save audio to a file
|
|
@@ -62,7 +71,7 @@ module ElevenRb
|
|
|
62
71
|
'audio/mpeg'
|
|
63
72
|
when /pcm/
|
|
64
73
|
'audio/pcm'
|
|
65
|
-
when /ogg/
|
|
74
|
+
when /ogg|opus/
|
|
66
75
|
'audio/ogg'
|
|
67
76
|
when /wav/
|
|
68
77
|
'audio/wav'
|
|
@@ -82,7 +91,7 @@ module ElevenRb
|
|
|
82
91
|
'mp3'
|
|
83
92
|
when /pcm/
|
|
84
93
|
'pcm'
|
|
85
|
-
when /ogg/
|
|
94
|
+
when /ogg|opus/
|
|
86
95
|
'ogg'
|
|
87
96
|
when /wav/
|
|
88
97
|
'wav'
|
|
@@ -12,6 +12,10 @@ module ElevenRb
|
|
|
12
12
|
'eleven_monolingual_v1' => 0.30,
|
|
13
13
|
'eleven_multilingual_v1' => 0.30,
|
|
14
14
|
'eleven_multilingual_v2' => 0.30,
|
|
15
|
+
'eleven_v3' => 0.30,
|
|
16
|
+
'eleven_v3_conversational' => 0.15,
|
|
17
|
+
'eleven_v4' => 0.30,
|
|
18
|
+
'eleven_v4_turbo' => 0.15,
|
|
15
19
|
'eleven_turbo_v2' => 0.18,
|
|
16
20
|
'eleven_turbo_v2_5' => 0.18,
|
|
17
21
|
'eleven_english_sts_v2' => 0.30,
|
|
@@ -23,11 +27,12 @@ module ElevenRb
|
|
|
23
27
|
|
|
24
28
|
# Initialize cost info
|
|
25
29
|
#
|
|
26
|
-
# @param text [String] the text being converted
|
|
30
|
+
# @param text [String, nil] the text being converted
|
|
31
|
+
# @param character_count [Integer, nil] direct character count (alternative to text)
|
|
27
32
|
# @param voice_id [String] the voice ID
|
|
28
33
|
# @param model_id [String] the model ID
|
|
29
|
-
def initialize(
|
|
30
|
-
@character_count = text
|
|
34
|
+
def initialize(voice_id:, model_id:, text: nil, character_count: nil)
|
|
35
|
+
@character_count = character_count || text&.length || 0
|
|
31
36
|
@voice_id = voice_id
|
|
32
37
|
@model_id = model_id
|
|
33
38
|
end
|
|
@@ -18,6 +18,16 @@ module ElevenRb
|
|
|
18
18
|
attribute :max_characters_request_free_user
|
|
19
19
|
attribute :max_characters_request_subscribed_user
|
|
20
20
|
attribute :concurrency_group
|
|
21
|
+
attribute :model_rates
|
|
22
|
+
attribute :maximum_text_length_per_request
|
|
23
|
+
attribute :requires_alpha_access, type: :boolean
|
|
24
|
+
|
|
25
|
+
# Voice settings keys this model honours (from ModelCapabilities)
|
|
26
|
+
#
|
|
27
|
+
# @return [Array<Symbol>]
|
|
28
|
+
def supported_voice_settings
|
|
29
|
+
ModelCapabilities.supported_voice_settings(model_id)
|
|
30
|
+
end
|
|
21
31
|
|
|
22
32
|
# Check if this model supports a given language
|
|
23
33
|
#
|