eleven_rb 1.0.0 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +33 -0
- data/README.md +92 -11
- data/lib/eleven_rb/callbacks.rb +25 -1
- data/lib/eleven_rb/configuration.rb +7 -3
- data/lib/eleven_rb/http/client.rb +32 -12
- data/lib/eleven_rb/model_capabilities.rb +103 -0
- data/lib/eleven_rb/objects/audio.rb +13 -4
- data/lib/eleven_rb/objects/cost_info.rb +3 -0
- data/lib/eleven_rb/objects/model.rb +10 -0
- data/lib/eleven_rb/objects/voice_settings.rb +33 -0
- data/lib/eleven_rb/resources/base.rb +18 -0
- data/lib/eleven_rb/resources/models.rb +32 -6
- data/lib/eleven_rb/resources/text_to_dialogue.rb +117 -27
- data/lib/eleven_rb/resources/text_to_speech.rb +150 -43
- data/lib/eleven_rb/version.rb +1 -1
- data/lib/eleven_rb.rb +2 -0
- metadata +2 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 18424bceaba545d05f991cb1ae674fdc36572a7a9573ef769701deb35537b393
|
|
4
|
+
data.tar.gz: a5c2a3814295f5d19945b1731b639fe21ae0f0111ef51741fdb99bae5a978cf7
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: af3da66572e56a9173b5f427a3fa30212ebd686a5820b060ab5c81e5e41d60ae64a74e7956e0ddad4365d0f8e68ceefaa5572d97f485a0e424609c195f34be31
|
|
7
|
+
data.tar.gz: 3af1d07b75e5405724c50e87d3b89f88061d3540043716b38242acce0941f5152294148cd1d9109c20f05bbd05f5131e6b1225260bd373371b45061ad4bbca26
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,39 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [1.1.0] - 2026-09-30
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- Eleven v4 support (`eleven_v4`, `eleven_v4_turbo`) across Text-to-Speech and Text-to-Dialogue
|
|
15
|
+
- `ElevenRb::ModelCapabilities` — per model-family table of honoured voice settings, max text length, SSML `<break>` support, audio-tag support and continuity support (`for`, `supported_voice_settings`, `max_text_length`, `supports?`)
|
|
16
|
+
- `Objects::VoiceSettings.for_model(model_id, overrides)` → `[settings, dropped_keys]`, building only the settings a model honours
|
|
17
|
+
- `Configuration#strict_voice_settings` (default `false`): raise `ValidationError` instead of warning when a voice setting the model ignores is passed
|
|
18
|
+
- Optional TTS keywords on `generate`, `stream` and `generate_with_timestamps` (omitted from the body when nil): `language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`
|
|
19
|
+
- Response metadata: `Objects::Audio#request_id`, `#character_cost` (from the `request-id` / `character-cost` headers) and `#dropped_settings`; `on_audio_generated` also receives `request_id:` (nil for streams)
|
|
20
|
+
- `TextToSpeech#generate_with_timestamps` now also returns `normalized_alignment`, `request_id` and `character_cost`
|
|
21
|
+
- `TextToDialogue#generate_with_timestamps` (`POST /v1/text-to-dialogue/with-timestamps`) returning `audio`, `alignment`, `normalized_alignment`, `voice_segments` and `request_id`
|
|
22
|
+
- Text-to-Dialogue keywords `use_pvc_as_ivc`, `previous_text`, `future_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`; a logger warning above `RECOMMENDED_MAX_TEXT_LENGTH` (2,000 characters)
|
|
23
|
+
- `HTTP::Client#post(..., with_meta: true)` returning `{ body:, headers: }`, and `Resources::Base#post_with_meta` / `#post_binary_with_meta`
|
|
24
|
+
- `Models#find(model_id)`
|
|
25
|
+
- `Objects::Model#model_rates`, `#maximum_text_length_per_request`, `#requires_alpha_access`, `#supported_voice_settings`
|
|
26
|
+
- `CostInfo::COST_PER_1K_CHARS` entries for `eleven_v4` ($0.30), `eleven_v4_turbo` ($0.15) and `eleven_v3_conversational` ($0.15)
|
|
27
|
+
- `opus_*` output formats map to the `ogg` extension and `audio/ogg` content type
|
|
28
|
+
|
|
29
|
+
### Changed
|
|
30
|
+
|
|
31
|
+
- Voice settings are filtered per model: known keys a model ignores (`stability`, `similarity_boost`, `style`, `use_speaker_boost`, `speed` minus the model's supported set) are dropped from the request (logged, and listed on `audio.dropped_settings`); unknown voice-setting keys pass through untouched on every model, so a new API field is never swallowed. `eleven_v3` / `eleven_v4` now send only `stability` and `similarity_boost` by default; `eleven_multilingual_v2` requests are byte-identical to 1.0.0. Override keys are symbolized and nil values removed (a nil override is never reported as dropped)
|
|
32
|
+
- Text length is capped per model (`ModelCapabilities.max_text_length`): 5,000 for `eleven_v3`, 10,000 for `eleven_v4` and `eleven_multilingual_v2`, 30,000 / 40,000 for the flash and turbo models. `TextToSpeech::MAX_TEXT_LENGTH` stays defined but is no longer the cap
|
|
33
|
+
- `TextToDialogue::DEFAULT_MODEL` is now `eleven_v4`
|
|
34
|
+
- Text-to-Dialogue's hard text cap now follows the model (10,000 characters on `eleven_v4`, 5,000 on `eleven_v3`) instead of a flat 5,000; `TextToDialogue::MAX_TEXT_LENGTH` stays defined but is no longer the cap
|
|
35
|
+
- `Models#latest` returns `eleven_v4`, else `eleven_v3`, else the default model
|
|
36
|
+
- `TextToSpeech::OUTPUT_FORMATS` refreshed to the current list (documentation only, not validated)
|
|
37
|
+
- Callbacks that declare their keywords explicitly (no `**rest`) receive only the keywords they declare, so callbacks written for 1.0.0 keep working as new keywords are added
|
|
38
|
+
|
|
39
|
+
### Fixed
|
|
40
|
+
|
|
41
|
+
- `Models#list` (and everything built on it: `get`, `default`, `latest`, `multilingual`, `turbo`, `tts_capable`, `ids`, `TTSAdapter#list_models`) recursed until `SystemStackError`, because `Models#get(model_id)` shadowed `Resources::Base#get`
|
|
42
|
+
|
|
10
43
|
## [1.0.0] - 2026-03-10
|
|
11
44
|
|
|
12
45
|
### Added
|
data/README.md
CHANGED
|
@@ -11,6 +11,7 @@ A Ruby client for the [ElevenLabs](https://try.elevenlabs.io/qyk2j8gumrjz) Text-
|
|
|
11
11
|
- Text-to-Speech generation and streaming
|
|
12
12
|
- Speech-to-Speech voice conversion
|
|
13
13
|
- Text-to-Dialogue multi-speaker generation with audio tags
|
|
14
|
+
- Eleven v4 support, with per-model voice-settings filtering and text caps
|
|
14
15
|
- Sound effects generation from text descriptions
|
|
15
16
|
- Music generation from prompts or composition plans
|
|
16
17
|
- Voice management (list, get, create, update, delete)
|
|
@@ -74,7 +75,7 @@ audio.save_to_file("output.mp3")
|
|
|
74
75
|
audio = client.tts.generate(
|
|
75
76
|
"Hello world",
|
|
76
77
|
voice_id: "voice_id",
|
|
77
|
-
model_id: "
|
|
78
|
+
model_id: "eleven_v4", # Most expressive; audio tags; see "Eleven v4" below
|
|
78
79
|
voice_settings: {
|
|
79
80
|
stability: 0.5,
|
|
80
81
|
similarity_boost: 0.75
|
|
@@ -88,8 +89,72 @@ File.open("output.mp3", "wb") do |file|
|
|
|
88
89
|
file.write(chunk)
|
|
89
90
|
end
|
|
90
91
|
end
|
|
92
|
+
|
|
93
|
+
# Word-level timestamps
|
|
94
|
+
result = client.tts.generate_with_timestamps("Hello world", voice_id: "voice_id")
|
|
95
|
+
result[:audio] # => ElevenRb::Objects::Audio
|
|
96
|
+
result[:alignment] # => { "characters" => [...], "character_start_times_seconds" => [...], ... }
|
|
97
|
+
result[:normalized_alignment]
|
|
98
|
+
result[:request_id] # from the request-id response header
|
|
99
|
+
result[:character_cost] # Integer, from the character-cost response header
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Optional keywords on `generate`, `stream` and `generate_with_timestamps` (each is left out of the request when nil):
|
|
103
|
+
`language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`,
|
|
104
|
+
`next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`.
|
|
105
|
+
|
|
106
|
+
### Eleven v4
|
|
107
|
+
|
|
108
|
+
`eleven_v4` (and the cheaper, faster `eleven_v4_turbo`) take up to 10,000 characters per request.
|
|
109
|
+
|
|
110
|
+
```ruby
|
|
111
|
+
audio = client.tts.generate(
|
|
112
|
+
"[sighs] Right... let's try that one more time.",
|
|
113
|
+
voice_id: "voice_id",
|
|
114
|
+
model_id: "eleven_v4",
|
|
115
|
+
voice_settings: { stability: 0.5, similarity_boost: 0.75 },
|
|
116
|
+
seed: 42, # reproducible re-rolls
|
|
117
|
+
previous_text: "That did not go to plan.", # continuity with the take before
|
|
118
|
+
next_text: "Here we go." # ...and the one after
|
|
119
|
+
)
|
|
120
|
+
audio.request_id # pass as previous_request_ids: / next_request_ids: on neighbouring takes
|
|
121
|
+
audio.character_cost
|
|
122
|
+
audio.dropped_settings # => [] (voice settings the gem removed, see below)
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
- **Only `stability` and `similarity_boost` take effect.** The API accepts `speed`, `style` and
|
|
126
|
+
`use_speaker_boost` on v4 and silently ignores them, so the gem drops them from the request, logs a
|
|
127
|
+
warning through `logger`, and lists them on `audio.dropped_settings`. Set
|
|
128
|
+
`strict_voice_settings: true` on the client to raise `ElevenRb::Errors::ValidationError` instead.
|
|
129
|
+
Only those known keys are ever dropped: a voice-setting key the gem does not know passes through
|
|
130
|
+
unchanged on every model, so new API fields keep working.
|
|
131
|
+
Pace a v4 take with the text (ellipses, pause tags), not `speed`.
|
|
132
|
+
- **SSML `<break time="…"/>` tags are ignored** by v4 (no pause is produced). Use an audio tag or
|
|
133
|
+
punctuation for pauses.
|
|
134
|
+
- **Audio tags** go in square brackets inside the text: `[sighs]`, `[whispers]`, `[laughs]`,
|
|
135
|
+
`[short pause]`. Note that tags appear in the returned alignment like any other characters.
|
|
136
|
+
- **Continuity**: `previous_text` / `next_text` (or `previous_request_ids` / `next_request_ids`) tell the
|
|
137
|
+
model what surrounds a take, so stitched takes keep a consistent delivery.
|
|
138
|
+
- **`seed`** makes v4 generation reproducible, so a re-roll changes one thing at a time.
|
|
139
|
+
|
|
140
|
+
`ElevenRb::ModelCapabilities` answers these questions for any model id:
|
|
141
|
+
|
|
142
|
+
```ruby
|
|
143
|
+
ElevenRb::ModelCapabilities.supported_voice_settings("eleven_v4") # => [:stability, :similarity_boost]
|
|
144
|
+
ElevenRb::ModelCapabilities.max_text_length("eleven_v3") # => 5000
|
|
145
|
+
ElevenRb::ModelCapabilities.supports?("eleven_v4", :audio_tags) # => true
|
|
146
|
+
ElevenRb::ModelCapabilities.supports?("eleven_v4", :ssml_break) # => false
|
|
91
147
|
```
|
|
92
148
|
|
|
149
|
+
| Model family | Voice settings honoured | Max chars | Audio tags | SSML breaks |
|
|
150
|
+
|---|---|---|---|---|
|
|
151
|
+
| `eleven_v4`, `eleven_v4_turbo` | stability, similarity_boost | 10,000 | yes | no |
|
|
152
|
+
| `eleven_v3` | stability, similarity_boost | 5,000 | yes | no |
|
|
153
|
+
| `eleven_v3_conversational` | stability, similarity_boost, use_speaker_boost | 5,000 | yes | no |
|
|
154
|
+
| `eleven_flash_v2_5`, `eleven_turbo_v2_5` | all five (incl. style, speed) | 40,000 | no | yes |
|
|
155
|
+
| `eleven_flash_v2`, `eleven_turbo_v2` | all five | 30,000 | no | yes |
|
|
156
|
+
| `eleven_multilingual_v2` and others | all five | 10,000 | no | yes |
|
|
157
|
+
|
|
93
158
|
### Speech-to-Speech
|
|
94
159
|
|
|
95
160
|
```ruby
|
|
@@ -123,30 +188,42 @@ audio = client.text_to_dialogue.generate([
|
|
|
123
188
|
])
|
|
124
189
|
audio.save_to_file("dialogue.mp3")
|
|
125
190
|
|
|
126
|
-
# With options
|
|
191
|
+
# With options (the default model is eleven_v4)
|
|
127
192
|
audio = client.dialogue.generate(
|
|
128
193
|
inputs,
|
|
129
|
-
model_id: "
|
|
194
|
+
model_id: "eleven_v4",
|
|
130
195
|
language_code: "en",
|
|
131
|
-
settings: { stability: 0.5 },
|
|
196
|
+
settings: { stability: 0.5, similarity: 0.75 }, # sent to the API unchanged
|
|
132
197
|
seed: 42,
|
|
198
|
+
previous_text: "Earlier in the scene...",
|
|
133
199
|
output_format: "mp3_44100_192"
|
|
134
200
|
)
|
|
201
|
+
|
|
202
|
+
# With timestamps and per-speaker segments
|
|
203
|
+
result = client.dialogue.generate_with_timestamps(inputs)
|
|
204
|
+
result[:audio] # => ElevenRb::Objects::Audio
|
|
205
|
+
result[:alignment] # character timings
|
|
206
|
+
result[:voice_segments] # which voice speaks when
|
|
207
|
+
result[:request_id]
|
|
135
208
|
```
|
|
136
209
|
|
|
210
|
+
The hard text cap follows the model (10,000 characters for `eleven_v4`, 5,000 for `eleven_v3`); above
|
|
211
|
+
2,000 characters the gem logs a warning, since shorter dialogue requests give more reliable results.
|
|
212
|
+
|
|
137
213
|
### Audio Tags
|
|
138
214
|
|
|
139
|
-
The `eleven_v3`
|
|
215
|
+
The `eleven_v4` and `eleven_v3` models support inline audio tags, in square brackets, for expressive speech
|
|
216
|
+
(`ElevenRb::ModelCapabilities.supports?(model_id, :audio_tags)`):
|
|
140
217
|
|
|
141
218
|
```ruby
|
|
142
219
|
audio = client.tts.generate(
|
|
143
220
|
"[excited] Oh wow, this is AMAZING! [laughs] I can't believe it...",
|
|
144
221
|
voice_id: "voice_id",
|
|
145
|
-
model_id: "
|
|
222
|
+
model_id: "eleven_v4"
|
|
146
223
|
)
|
|
147
224
|
```
|
|
148
225
|
|
|
149
|
-
Supported tags include `[laughs]`, `[whispers]`, `[sighs]`, `[excited]`, `[sarcastic]`, `[curious]`, `[pause]`, and more. Use CAPS for emphasis, `...` for pauses, and `—` for interruptions. See the [ElevenLabs
|
|
226
|
+
Supported tags include `[laughs]`, `[whispers]`, `[sighs]`, `[excited]`, `[sarcastic]`, `[curious]`, `[short pause]`, and more. Use CAPS for emphasis, `...` for pauses, and `—` for interruptions. See the [ElevenLabs audio tags documentation](https://elevenlabs.io/docs/guides/audio-tags) for the full list.
|
|
150
227
|
|
|
151
228
|
### Sound Effects
|
|
152
229
|
|
|
@@ -289,7 +366,7 @@ client = ElevenRb::Client.new(
|
|
|
289
366
|
Sentry.capture_exception(error, extra: { path: path })
|
|
290
367
|
},
|
|
291
368
|
|
|
292
|
-
# Cost tracking
|
|
369
|
+
# Cost tracking (add request_id: to also receive the API's request id)
|
|
293
370
|
on_audio_generated: ->(audio:, voice_id:, text:, cost_info:) {
|
|
294
371
|
UsageRecord.create!(
|
|
295
372
|
characters: cost_info[:character_count],
|
|
@@ -311,8 +388,11 @@ client = ElevenRb::Client.new(
|
|
|
311
388
|
models = client.models.list
|
|
312
389
|
models.each { |m| puts "#{m.name} (#{m.model_id})" }
|
|
313
390
|
|
|
314
|
-
#
|
|
315
|
-
client.models.
|
|
391
|
+
# Find one model
|
|
392
|
+
client.models.find("eleven_v4") # => ElevenRb::Objects::Model (get is an alias)
|
|
393
|
+
|
|
394
|
+
# Get the latest/most capable model: eleven_v4, else eleven_v3, else the default
|
|
395
|
+
client.models.latest.model_id # => "eleven_v4"
|
|
316
396
|
|
|
317
397
|
# Get multilingual models
|
|
318
398
|
client.models.multilingual
|
|
@@ -343,7 +423,8 @@ client = ElevenRb::Client.new(
|
|
|
343
423
|
open_timeout: 10, # Connection timeout
|
|
344
424
|
max_retries: 3, # Max retry attempts
|
|
345
425
|
retry_delay: 1.0, # Base delay between retries
|
|
346
|
-
logger: Rails.logger
|
|
426
|
+
logger: Rails.logger, # Optional logger (receives dropped-setting warnings)
|
|
427
|
+
strict_voice_settings: false # true: raise instead of dropping settings a model ignores
|
|
347
428
|
)
|
|
348
429
|
```
|
|
349
430
|
|
data/lib/eleven_rb/callbacks.rb
CHANGED
|
@@ -28,6 +28,10 @@ module ElevenRb
|
|
|
28
28
|
|
|
29
29
|
# Trigger a callback if it's configured
|
|
30
30
|
#
|
|
31
|
+
# A callback that names its keywords explicitly (no `**rest`) receives only
|
|
32
|
+
# the keywords it declares, so callbacks written before a keyword was added
|
|
33
|
+
# (e.g. `request_id:` on on_audio_generated in 1.1.0) keep working.
|
|
34
|
+
#
|
|
31
35
|
# @param callback_name [Symbol] the name of the callback
|
|
32
36
|
# @param kwargs [Hash] keyword arguments to pass to the callback
|
|
33
37
|
# @return [Object, nil] the return value of the callback, or nil
|
|
@@ -36,12 +40,32 @@ module ElevenRb
|
|
|
36
40
|
return unless callback.respond_to?(:call)
|
|
37
41
|
|
|
38
42
|
begin
|
|
39
|
-
callback.call(**kwargs)
|
|
43
|
+
callback.call(**accepted_callback_kwargs(callback, kwargs))
|
|
40
44
|
rescue StandardError => e
|
|
41
45
|
# Don't let callback errors break the main flow
|
|
42
46
|
warn "[ElevenRb] Callback error in #{callback_name}: #{e.message}"
|
|
43
47
|
nil
|
|
44
48
|
end
|
|
45
49
|
end
|
|
50
|
+
|
|
51
|
+
private
|
|
52
|
+
|
|
53
|
+
def accepted_callback_kwargs(callback, kwargs)
|
|
54
|
+
params = callback_parameters(callback)
|
|
55
|
+
return kwargs if params.nil? || params.any? { |type, _| type == :keyrest }
|
|
56
|
+
|
|
57
|
+
accepted = params.filter_map { |type, name| name if %i[key keyreq].include?(type) }
|
|
58
|
+
return kwargs if accepted.empty?
|
|
59
|
+
|
|
60
|
+
kwargs.slice(*accepted)
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
def callback_parameters(callback)
|
|
64
|
+
return callback.parameters if callback.respond_to?(:parameters)
|
|
65
|
+
|
|
66
|
+
callback.method(:call).parameters
|
|
67
|
+
rescue NameError
|
|
68
|
+
nil
|
|
69
|
+
end
|
|
46
70
|
end
|
|
47
71
|
end
|
|
@@ -21,12 +21,13 @@ module ElevenRb
|
|
|
21
21
|
open_timeout: 10,
|
|
22
22
|
max_retries: 3,
|
|
23
23
|
retry_delay: 1.0,
|
|
24
|
-
retry_statuses: [429, 500, 502, 503, 504].freeze
|
|
24
|
+
retry_statuses: [429, 500, 502, 503, 504].freeze,
|
|
25
|
+
strict_voice_settings: false
|
|
25
26
|
}.freeze
|
|
26
27
|
|
|
27
28
|
attr_accessor :api_key, :base_url, :timeout, :open_timeout,
|
|
28
29
|
:max_retries, :retry_delay, :retry_statuses,
|
|
29
|
-
:logger
|
|
30
|
+
:logger, :strict_voice_settings
|
|
30
31
|
|
|
31
32
|
# Initialize a new configuration
|
|
32
33
|
#
|
|
@@ -39,6 +40,8 @@ module ElevenRb
|
|
|
39
40
|
# @option options [Float] :retry_delay Base delay between retries in seconds (default: 1.0)
|
|
40
41
|
# @option options [Array<Integer>] :retry_statuses HTTP status codes to retry (default: [429, 500, 502, 503, 504])
|
|
41
42
|
# @option options [Logger] :logger Logger instance for debug output
|
|
43
|
+
# @option options [Boolean] :strict_voice_settings Raise instead of warn when a voice setting the
|
|
44
|
+
# model does not honour is passed (default: false)
|
|
42
45
|
# @option options [Proc] :on_request Callback before each request
|
|
43
46
|
# @option options [Proc] :on_response Callback after successful response
|
|
44
47
|
# @option options [Proc] :on_error Callback when an error occurs
|
|
@@ -87,7 +90,8 @@ module ElevenRb
|
|
|
87
90
|
timeout: timeout,
|
|
88
91
|
open_timeout: open_timeout,
|
|
89
92
|
max_retries: max_retries,
|
|
90
|
-
retry_delay: retry_delay
|
|
93
|
+
retry_delay: retry_delay,
|
|
94
|
+
strict_voice_settings: strict_voice_settings
|
|
91
95
|
}
|
|
92
96
|
end
|
|
93
97
|
end
|
|
@@ -32,9 +32,11 @@ module ElevenRb
|
|
|
32
32
|
# @param path [String] the API path
|
|
33
33
|
# @param body [Hash] request body
|
|
34
34
|
# @param response_type [Symbol] :json or :binary
|
|
35
|
-
# @
|
|
36
|
-
|
|
37
|
-
|
|
35
|
+
# @param with_meta [Boolean] when true, return `{ body:, headers: }` instead of the body alone
|
|
36
|
+
# (headers as a Hash of lower-cased String keys to String values)
|
|
37
|
+
# @return [Hash, Array, String] parsed response (or `{ body:, headers: }` with with_meta)
|
|
38
|
+
def post(path, body = {}, response_type: :json, with_meta: false)
|
|
39
|
+
request(:post, path, body: body, response_type: response_type, with_meta: with_meta)
|
|
38
40
|
end
|
|
39
41
|
|
|
40
42
|
# Make a DELETE request
|
|
@@ -68,7 +70,7 @@ module ElevenRb
|
|
|
68
70
|
private
|
|
69
71
|
|
|
70
72
|
def request(method, path, body: nil, params: nil, response_type: :json, multipart: false, stream: false,
|
|
71
|
-
attempt: 1, &block)
|
|
73
|
+
attempt: 1, with_meta: false, &block)
|
|
72
74
|
config.validate!
|
|
73
75
|
url = "#{config.base_url}#{path}"
|
|
74
76
|
start_time = Time.now
|
|
@@ -86,15 +88,12 @@ module ElevenRb
|
|
|
86
88
|
# Trigger after response callback
|
|
87
89
|
config.trigger(:on_response, method: method, path: path, response: response, duration: duration)
|
|
88
90
|
|
|
89
|
-
|
|
90
|
-
return response.body if response_type == :binary && response.success?
|
|
91
|
-
|
|
92
|
-
handle_response(response)
|
|
91
|
+
build_result(response, response_type, with_meta)
|
|
93
92
|
rescue Errors::RateLimitError => e
|
|
94
93
|
config.trigger(:on_rate_limit, retry_after: e.retry_after, error: e)
|
|
95
|
-
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
|
|
94
|
+
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
|
|
96
95
|
rescue Errors::ServerError => e
|
|
97
|
-
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
|
|
96
|
+
handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
|
|
98
97
|
rescue Errors::Base => e
|
|
99
98
|
config.trigger(:on_error, error: e, method: method, path: path,
|
|
100
99
|
context: { body: sanitize_body_for_logging(body) })
|
|
@@ -107,6 +106,17 @@ module ElevenRb
|
|
|
107
106
|
end
|
|
108
107
|
end
|
|
109
108
|
|
|
109
|
+
def build_result(response, response_type, with_meta)
|
|
110
|
+
# Return binary data directly
|
|
111
|
+
result = if response_type == :binary && response.success?
|
|
112
|
+
response.body
|
|
113
|
+
else
|
|
114
|
+
handle_response(response)
|
|
115
|
+
end
|
|
116
|
+
|
|
117
|
+
with_meta ? { body: result, headers: normalize_headers(response) } : result
|
|
118
|
+
end
|
|
119
|
+
|
|
110
120
|
def execute_request(method, url, body, params, multipart, stream, &block)
|
|
111
121
|
options = build_options(body, params, multipart, stream, &block)
|
|
112
122
|
|
|
@@ -210,7 +220,7 @@ module ElevenRb
|
|
|
210
220
|
raise error_class.new(message, **error_kwargs)
|
|
211
221
|
end
|
|
212
222
|
|
|
213
|
-
def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, &block)
|
|
223
|
+
def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
|
|
214
224
|
raise error if attempt > config.max_retries || !config.retry_statuses.include?(error.http_status)
|
|
215
225
|
|
|
216
226
|
delay = if error.is_a?(Errors::RateLimitError) && error.retry_after
|
|
@@ -224,7 +234,17 @@ module ElevenRb
|
|
|
224
234
|
sleep(delay)
|
|
225
235
|
|
|
226
236
|
request(method, path,
|
|
227
|
-
body: body, params: params, response_type: response_type, multipart: multipart, stream: stream,
|
|
237
|
+
body: body, params: params, response_type: response_type, multipart: multipart, stream: stream,
|
|
238
|
+
attempt: attempt + 1, with_meta: with_meta, &block)
|
|
239
|
+
end
|
|
240
|
+
|
|
241
|
+
# Response headers as a plain Hash of lower-cased String keys to String values
|
|
242
|
+
# (HTTParty exposes multi-value arrays; the first value is kept)
|
|
243
|
+
def normalize_headers(response)
|
|
244
|
+
raw = response.headers.respond_to?(:to_hash) ? response.headers.to_hash : response.headers.to_h
|
|
245
|
+
raw.each_with_object({}) do |(key, value), out|
|
|
246
|
+
out[key.to_s.downcase] = value.is_a?(Array) ? value.first : value
|
|
247
|
+
end
|
|
228
248
|
end
|
|
229
249
|
|
|
230
250
|
def wrap_error(error)
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module ElevenRb
|
|
4
|
+
# What each ElevenLabs model family accepts, keyed by model-id family.
|
|
5
|
+
#
|
|
6
|
+
# The API silently ignores voice settings a model does not honour (Eleven v4
|
|
7
|
+
# accepts `speed`, `style` and `use_speaker_boost` and does nothing with them),
|
|
8
|
+
# so the gem uses this table to send only the settings that take effect, to
|
|
9
|
+
# cap text length per model, and to answer feature questions (audio tags,
|
|
10
|
+
# SSML `<break>` tags, request continuity via previous_text / next_text).
|
|
11
|
+
#
|
|
12
|
+
# @example
|
|
13
|
+
# ElevenRb::ModelCapabilities.supported_voice_settings('eleven_v4')
|
|
14
|
+
# # => [:stability, :similarity_boost]
|
|
15
|
+
# ElevenRb::ModelCapabilities.max_text_length('eleven_flash_v2_5') # => 40_000
|
|
16
|
+
# ElevenRb::ModelCapabilities.supports?('eleven_v4', :audio_tags) # => true
|
|
17
|
+
module ModelCapabilities
|
|
18
|
+
# Capability record for one model family
|
|
19
|
+
Capabilities = Struct.new(:voice_settings, :max_text_length, :ssml_break, :audio_tags, :continuity,
|
|
20
|
+
keyword_init: true) do
|
|
21
|
+
def ssml_break? = ssml_break
|
|
22
|
+
def audio_tags? = audio_tags
|
|
23
|
+
def continuity? = continuity
|
|
24
|
+
end
|
|
25
|
+
|
|
26
|
+
ALL_VOICE_SETTINGS = %i[stability similarity_boost style use_speaker_boost speed].freeze
|
|
27
|
+
|
|
28
|
+
def self.build(voice_settings:, max_text_length:, ssml_break:, audio_tags:, continuity:)
|
|
29
|
+
Capabilities.new(
|
|
30
|
+
voice_settings: voice_settings.freeze,
|
|
31
|
+
max_text_length: max_text_length,
|
|
32
|
+
ssml_break: ssml_break,
|
|
33
|
+
audio_tags: audio_tags,
|
|
34
|
+
continuity: continuity
|
|
35
|
+
).freeze
|
|
36
|
+
end
|
|
37
|
+
private_class_method :build
|
|
38
|
+
|
|
39
|
+
V4 = build(voice_settings: %i[stability similarity_boost], max_text_length: 10_000,
|
|
40
|
+
ssml_break: false, audio_tags: true, continuity: true)
|
|
41
|
+
V3_CONVERSATIONAL = build(voice_settings: %i[stability similarity_boost use_speaker_boost], max_text_length: 5_000,
|
|
42
|
+
ssml_break: false, audio_tags: true, continuity: true)
|
|
43
|
+
V3 = build(voice_settings: %i[stability similarity_boost], max_text_length: 5_000,
|
|
44
|
+
ssml_break: false, audio_tags: true, continuity: true)
|
|
45
|
+
V2_5_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 40_000,
|
|
46
|
+
ssml_break: true, audio_tags: false, continuity: true)
|
|
47
|
+
V2_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 30_000,
|
|
48
|
+
ssml_break: true, audio_tags: false, continuity: true)
|
|
49
|
+
DEFAULT = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 10_000,
|
|
50
|
+
ssml_break: true, audio_tags: false, continuity: true)
|
|
51
|
+
|
|
52
|
+
# Ordered [matcher, capabilities] pairs; the first match wins.
|
|
53
|
+
FAMILIES = [
|
|
54
|
+
[/\Aeleven_v4/, V4],
|
|
55
|
+
[/\Aeleven_v3_conversational/, V3_CONVERSATIONAL],
|
|
56
|
+
[/\Aeleven_v3/, V3],
|
|
57
|
+
[/\Aeleven_(flash|turbo)_v2_5/, V2_5_FAST],
|
|
58
|
+
[/\Aeleven_(flash|turbo)_v2/, V2_FAST]
|
|
59
|
+
].freeze
|
|
60
|
+
|
|
61
|
+
FEATURES = %i[audio_tags ssml_break continuity].freeze
|
|
62
|
+
|
|
63
|
+
module_function
|
|
64
|
+
|
|
65
|
+
# Capabilities for a model id (unknown ids get the permissive default)
|
|
66
|
+
#
|
|
67
|
+
# @param model_id [String, Symbol, nil]
|
|
68
|
+
# @return [Capabilities] frozen
|
|
69
|
+
def for(model_id)
|
|
70
|
+
id = model_id.to_s
|
|
71
|
+
FAMILIES.each { |matcher, caps| return caps if matcher.match?(id) }
|
|
72
|
+
DEFAULT
|
|
73
|
+
end
|
|
74
|
+
|
|
75
|
+
# Voice settings keys the model honours
|
|
76
|
+
#
|
|
77
|
+
# @param model_id [String]
|
|
78
|
+
# @return [Array<Symbol>]
|
|
79
|
+
def supported_voice_settings(model_id)
|
|
80
|
+
self.for(model_id).voice_settings
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
# Maximum characters per request for the model
|
|
84
|
+
#
|
|
85
|
+
# @param model_id [String]
|
|
86
|
+
# @return [Integer]
|
|
87
|
+
def max_text_length(model_id)
|
|
88
|
+
self.for(model_id).max_text_length
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
# Whether the model supports a feature
|
|
92
|
+
#
|
|
93
|
+
# @param model_id [String]
|
|
94
|
+
# @param feature [Symbol] :audio_tags, :ssml_break or :continuity
|
|
95
|
+
# @return [Boolean]
|
|
96
|
+
def supports?(model_id, feature)
|
|
97
|
+
feature = feature.to_sym
|
|
98
|
+
raise ArgumentError, "Unknown feature #{feature.inspect} (expected one of #{FEATURES.join(', ')})" unless FEATURES.include?(feature)
|
|
99
|
+
|
|
100
|
+
self.for(model_id).public_send(feature) ? true : false
|
|
101
|
+
end
|
|
102
|
+
end
|
|
103
|
+
end
|
|
@@ -6,7 +6,8 @@ module ElevenRb
|
|
|
6
6
|
module Objects
|
|
7
7
|
# Represents generated audio data
|
|
8
8
|
class Audio
|
|
9
|
-
attr_reader :data, :format, :voice_id, :text, :model_id
|
|
9
|
+
attr_reader :data, :format, :voice_id, :text, :model_id,
|
|
10
|
+
:request_id, :character_cost, :dropped_settings
|
|
10
11
|
|
|
11
12
|
# Initialize audio object
|
|
12
13
|
#
|
|
@@ -15,12 +16,20 @@ module ElevenRb
|
|
|
15
16
|
# @param voice_id [String] the voice ID used
|
|
16
17
|
# @param text [String] the text that was converted
|
|
17
18
|
# @param model_id [String, nil] the model ID used
|
|
18
|
-
|
|
19
|
+
# @param request_id [String, nil] the API's `request-id` response header (usable as
|
|
20
|
+
# previous_request_ids / next_request_ids on a later request)
|
|
21
|
+
# @param character_cost [Integer, nil] the API's `character-cost` response header
|
|
22
|
+
# @param dropped_settings [Array<Symbol>] voice settings the gem dropped because the model ignores them
|
|
23
|
+
def initialize(data:, format:, voice_id:, text:, model_id: nil, request_id: nil, character_cost: nil,
|
|
24
|
+
dropped_settings: [])
|
|
19
25
|
@data = data
|
|
20
26
|
@format = format
|
|
21
27
|
@voice_id = voice_id
|
|
22
28
|
@text = text
|
|
23
29
|
@model_id = model_id
|
|
30
|
+
@request_id = request_id
|
|
31
|
+
@character_cost = character_cost
|
|
32
|
+
@dropped_settings = Array(dropped_settings).dup.freeze
|
|
24
33
|
end
|
|
25
34
|
|
|
26
35
|
# Save audio to a file
|
|
@@ -62,7 +71,7 @@ module ElevenRb
|
|
|
62
71
|
'audio/mpeg'
|
|
63
72
|
when /pcm/
|
|
64
73
|
'audio/pcm'
|
|
65
|
-
when /ogg/
|
|
74
|
+
when /ogg|opus/
|
|
66
75
|
'audio/ogg'
|
|
67
76
|
when /wav/
|
|
68
77
|
'audio/wav'
|
|
@@ -82,7 +91,7 @@ module ElevenRb
|
|
|
82
91
|
'mp3'
|
|
83
92
|
when /pcm/
|
|
84
93
|
'pcm'
|
|
85
|
-
when /ogg/
|
|
94
|
+
when /ogg|opus/
|
|
86
95
|
'ogg'
|
|
87
96
|
when /wav/
|
|
88
97
|
'wav'
|
|
@@ -13,6 +13,9 @@ module ElevenRb
|
|
|
13
13
|
'eleven_multilingual_v1' => 0.30,
|
|
14
14
|
'eleven_multilingual_v2' => 0.30,
|
|
15
15
|
'eleven_v3' => 0.30,
|
|
16
|
+
'eleven_v3_conversational' => 0.15,
|
|
17
|
+
'eleven_v4' => 0.30,
|
|
18
|
+
'eleven_v4_turbo' => 0.15,
|
|
16
19
|
'eleven_turbo_v2' => 0.18,
|
|
17
20
|
'eleven_turbo_v2_5' => 0.18,
|
|
18
21
|
'eleven_english_sts_v2' => 0.30,
|
|
@@ -18,6 +18,16 @@ module ElevenRb
|
|
|
18
18
|
attribute :max_characters_request_free_user
|
|
19
19
|
attribute :max_characters_request_subscribed_user
|
|
20
20
|
attribute :concurrency_group
|
|
21
|
+
attribute :model_rates
|
|
22
|
+
attribute :maximum_text_length_per_request
|
|
23
|
+
attribute :requires_alpha_access, type: :boolean
|
|
24
|
+
|
|
25
|
+
# Voice settings keys this model honours (from ModelCapabilities)
|
|
26
|
+
#
|
|
27
|
+
# @return [Array<Symbol>]
|
|
28
|
+
def supported_voice_settings
|
|
29
|
+
ModelCapabilities.supported_voice_settings(model_id)
|
|
30
|
+
end
|
|
21
31
|
|
|
22
32
|
# Check if this model supports a given language
|
|
23
33
|
#
|
|
@@ -25,6 +25,39 @@ module ElevenRb
|
|
|
25
25
|
from_response(DEFAULTS.merge(overrides))
|
|
26
26
|
end
|
|
27
27
|
|
|
28
|
+
# Voice-setting keys the capability table knows about. Only these can be
|
|
29
|
+
# dropped for a model; any other key is passed through untouched so a
|
|
30
|
+
# future API field is never swallowed.
|
|
31
|
+
KNOWN_KEYS = ModelCapabilities::ALL_VOICE_SETTINGS
|
|
32
|
+
|
|
33
|
+
# Build the voice_settings hash a model actually honours
|
|
34
|
+
#
|
|
35
|
+
# Starts from DEFAULTS filtered to the model's supported keys and merges the
|
|
36
|
+
# overrides (keys symbolized). Known keys the model does not support are
|
|
37
|
+
# dropped; unknown keys pass through after the supported ones, in the
|
|
38
|
+
# caller's order; nil values are removed. Only non-nil override keys count
|
|
39
|
+
# as dropped: defaults the model does not take are filtered silently.
|
|
40
|
+
#
|
|
41
|
+
# @example
|
|
42
|
+
# VoiceSettings.for_model('eleven_multilingual_v2')
|
|
43
|
+
# # => [{ stability: 0.5, similarity_boost: 0.75, style: 0.0, use_speaker_boost: true }, []]
|
|
44
|
+
# VoiceSettings.for_model('eleven_v4', speed: 1.1)
|
|
45
|
+
# # => [{ stability: 0.5, similarity_boost: 0.75 }, [:speed]]
|
|
46
|
+
#
|
|
47
|
+
# @param model_id [String] the model ID
|
|
48
|
+
# @param overrides [Hash] caller settings (String or Symbol keys)
|
|
49
|
+
# @return [Array(Hash, Array<Symbol>)] the settings hash and the dropped override keys
|
|
50
|
+
def self.for_model(model_id, overrides = {})
|
|
51
|
+
supported = ModelCapabilities.supported_voice_settings(model_id)
|
|
52
|
+
requested = (overrides || {}).to_h.transform_keys(&:to_sym)
|
|
53
|
+
|
|
54
|
+
dropped = requested.compact.keys & (KNOWN_KEYS - supported)
|
|
55
|
+
unknown = requested.except(*KNOWN_KEYS)
|
|
56
|
+
settings = DEFAULTS.slice(*supported).merge(requested.slice(*supported)).merge(unknown).compact
|
|
57
|
+
|
|
58
|
+
[settings, dropped]
|
|
59
|
+
end
|
|
60
|
+
|
|
28
61
|
# Convert to hash suitable for API request
|
|
29
62
|
#
|
|
30
63
|
# @return [Hash]
|
|
@@ -51,6 +51,24 @@ module ElevenRb
|
|
|
51
51
|
http_client.post(path, body, response_type: :binary)
|
|
52
52
|
end
|
|
53
53
|
|
|
54
|
+
# Make a JSON POST request and also return the response headers
|
|
55
|
+
#
|
|
56
|
+
# @param path [String]
|
|
57
|
+
# @param body [Hash]
|
|
58
|
+
# @return [Hash] `{ body: Hash, headers: Hash<String, String> }`
|
|
59
|
+
def post_with_meta(path, body = {})
|
|
60
|
+
http_client.post(path, body, response_type: :json, with_meta: true)
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
# Make a binary POST request and also return the response headers
|
|
64
|
+
#
|
|
65
|
+
# @param path [String]
|
|
66
|
+
# @param body [Hash]
|
|
67
|
+
# @return [Hash] `{ body: String, headers: Hash<String, String> }`
|
|
68
|
+
def post_binary_with_meta(path, body = {})
|
|
69
|
+
http_client.post(path, body, response_type: :binary, with_meta: true)
|
|
70
|
+
end
|
|
71
|
+
|
|
54
72
|
# Make a streaming POST request
|
|
55
73
|
#
|
|
56
74
|
# @param path [String]
|
|
@@ -9,23 +9,36 @@ module ElevenRb
|
|
|
9
9
|
#
|
|
10
10
|
# @example Find multilingual models
|
|
11
11
|
# client.models.multilingual
|
|
12
|
+
#
|
|
13
|
+
# @example Find one model
|
|
14
|
+
# client.models.find('eleven_v4')
|
|
12
15
|
class Models < Base
|
|
13
16
|
# List all available models
|
|
14
17
|
#
|
|
15
18
|
# @return [Array<Objects::Model>]
|
|
16
19
|
def list
|
|
17
|
-
|
|
20
|
+
# Call the HTTP client directly: #get below is the model lookup (kept for
|
|
21
|
+
# compatibility) and shadows Base#get, which made this method recurse.
|
|
22
|
+
response = http_client.get('/models')
|
|
18
23
|
response.map { |m| Objects::Model.from_response(m) }
|
|
19
24
|
end
|
|
20
25
|
|
|
21
|
-
#
|
|
26
|
+
# Find a specific model by ID
|
|
22
27
|
#
|
|
23
28
|
# @param model_id [String] the model ID
|
|
24
29
|
# @return [Objects::Model, nil]
|
|
25
|
-
def
|
|
30
|
+
def find(model_id)
|
|
26
31
|
list.find { |m| m.model_id == model_id }
|
|
27
32
|
end
|
|
28
33
|
|
|
34
|
+
# Alias of {#find}, kept for backwards compatibility
|
|
35
|
+
#
|
|
36
|
+
# @param model_id [String] the model ID
|
|
37
|
+
# @return [Objects::Model, nil]
|
|
38
|
+
def get(model_id)
|
|
39
|
+
find(model_id)
|
|
40
|
+
end
|
|
41
|
+
|
|
29
42
|
# Get all multilingual models
|
|
30
43
|
#
|
|
31
44
|
# @return [Array<Objects::Model>]
|
|
@@ -51,14 +64,20 @@ module ElevenRb
|
|
|
51
64
|
#
|
|
52
65
|
# @return [Objects::Model, nil]
|
|
53
66
|
def default
|
|
54
|
-
|
|
67
|
+
default_from(list)
|
|
55
68
|
end
|
|
56
69
|
|
|
57
|
-
# Get the latest/most capable model
|
|
70
|
+
# Get the latest/most capable model available to the account:
|
|
71
|
+
# eleven_v4, else eleven_v3, else {#default} (one /models request)
|
|
58
72
|
#
|
|
59
73
|
# @return [Objects::Model, nil]
|
|
60
74
|
def latest
|
|
61
|
-
|
|
75
|
+
models = list
|
|
76
|
+
%w[eleven_v4 eleven_v3].each do |model_id|
|
|
77
|
+
model = models.find { |m| m.model_id == model_id }
|
|
78
|
+
return model if model
|
|
79
|
+
end
|
|
80
|
+
default_from(models)
|
|
62
81
|
end
|
|
63
82
|
|
|
64
83
|
# Get model IDs as array
|
|
@@ -67,6 +86,13 @@ module ElevenRb
|
|
|
67
86
|
def ids
|
|
68
87
|
list.map(&:model_id)
|
|
69
88
|
end
|
|
89
|
+
|
|
90
|
+
private
|
|
91
|
+
|
|
92
|
+
# eleven_multilingual_v2, else the first TTS-capable model, from an already-fetched list
|
|
93
|
+
def default_from(models)
|
|
94
|
+
models.find { |m| m.model_id == 'eleven_multilingual_v2' } || models.find(&:can_do_text_to_speech)
|
|
95
|
+
end
|
|
70
96
|
end
|
|
71
97
|
end
|
|
72
98
|
end
|
|
@@ -10,20 +10,51 @@ module ElevenRb
|
|
|
10
10
|
# { text: "[laughs] Thanks!", voice_id: "voice_xyz" }
|
|
11
11
|
# ])
|
|
12
12
|
# audio.save_to_file("dialogue.mp3")
|
|
13
|
+
#
|
|
14
|
+
# @example Dialogue with timestamps and per-speaker segments
|
|
15
|
+
# result = client.dialogue.generate_with_timestamps(inputs, seed: 7)
|
|
16
|
+
# result[:voice_segments] # => [{ "voice_id" => ..., "start_time_seconds" => ... }, ...]
|
|
13
17
|
class TextToDialogue < Base
|
|
14
|
-
DEFAULT_MODEL = '
|
|
18
|
+
DEFAULT_MODEL = 'eleven_v4'
|
|
15
19
|
MAX_VOICES_PER_REQUEST = 10
|
|
20
|
+
|
|
21
|
+
# Kept for compatibility. The hard cap now comes from
|
|
22
|
+
# ModelCapabilities.max_text_length(model_id) (5,000 for eleven_v3).
|
|
16
23
|
MAX_TEXT_LENGTH = 5000
|
|
17
24
|
|
|
25
|
+
# Above this many characters the API recommends splitting the dialogue;
|
|
26
|
+
# the gem logs a warning but still sends the request.
|
|
27
|
+
RECOMMENDED_MAX_TEXT_LENGTH = 2_000
|
|
28
|
+
|
|
29
|
+
# Optional request-body keys, in the order they are written to the body.
|
|
30
|
+
# Each is omitted from the body when nil.
|
|
31
|
+
OPTIONAL_BODY_KEYS = %i[
|
|
32
|
+
language_code
|
|
33
|
+
settings
|
|
34
|
+
seed
|
|
35
|
+
use_pvc_as_ivc
|
|
36
|
+
previous_text
|
|
37
|
+
future_text
|
|
38
|
+
previous_request_ids
|
|
39
|
+
next_request_ids
|
|
40
|
+
pronunciation_dictionary_locators
|
|
41
|
+
].freeze
|
|
42
|
+
|
|
18
43
|
# Generate dialogue audio from multiple speaker inputs
|
|
19
44
|
#
|
|
20
45
|
# @param inputs [Array<Hash>] Array of { text:, voice_id: } hashes
|
|
21
|
-
# @param model_id [String] Model to use (
|
|
46
|
+
# @param model_id [String] Model to use (default: eleven_v4)
|
|
22
47
|
# @param language_code [String, nil] ISO 639-1 language code
|
|
23
|
-
# @param settings [Hash, nil] Generation settings (stability
|
|
48
|
+
# @param settings [Hash, nil] Generation settings, sent unchanged (e.g. stability, similarity)
|
|
24
49
|
# @param seed [Integer, nil] Seed for reproducibility
|
|
25
50
|
# @param output_format [String] Audio output format
|
|
26
51
|
# @param apply_text_normalization [String] "auto", "on", or "off"
|
|
52
|
+
# @param use_pvc_as_ivc [Boolean, nil] use the IVC version of professional voices
|
|
53
|
+
# @param previous_text [String, nil] text that comes before this dialogue (continuity)
|
|
54
|
+
# @param future_text [String, nil] text that comes after this dialogue (continuity)
|
|
55
|
+
# @param previous_request_ids [Array<String>, nil] request IDs of preceding generations
|
|
56
|
+
# @param next_request_ids [Array<String>, nil] request IDs of following generations
|
|
57
|
+
# @param pronunciation_dictionary_locators [Array<Hash>, nil] pronunciation dictionaries to apply
|
|
27
58
|
# @return [Objects::Audio]
|
|
28
59
|
def generate(
|
|
29
60
|
inputs,
|
|
@@ -32,45 +63,92 @@ module ElevenRb
|
|
|
32
63
|
settings: nil,
|
|
33
64
|
seed: nil,
|
|
34
65
|
output_format: 'mp3_44100_128',
|
|
35
|
-
apply_text_normalization: 'auto'
|
|
66
|
+
apply_text_normalization: 'auto',
|
|
67
|
+
use_pvc_as_ivc: nil,
|
|
68
|
+
previous_text: nil,
|
|
69
|
+
future_text: nil,
|
|
70
|
+
previous_request_ids: nil,
|
|
71
|
+
next_request_ids: nil,
|
|
72
|
+
pronunciation_dictionary_locators: nil
|
|
36
73
|
)
|
|
37
|
-
validate_inputs!(inputs)
|
|
74
|
+
validate_inputs!(inputs, model_id)
|
|
38
75
|
|
|
39
|
-
body = build_request_body(inputs, model_id,
|
|
40
|
-
|
|
76
|
+
body = build_request_body(inputs, model_id, apply_text_normalization, optional_values(binding))
|
|
77
|
+
response = post_binary_with_meta("/text-to-dialogue?output_format=#{output_format}", body)
|
|
41
78
|
|
|
42
|
-
response
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
)
|
|
79
|
+
build_audio_response(response[:body], inputs, output_format, model_id,
|
|
80
|
+
request_id: response[:headers]['request-id'])
|
|
81
|
+
end
|
|
46
82
|
|
|
47
|
-
|
|
83
|
+
# Generate dialogue audio with character timestamps and per-voice segments
|
|
84
|
+
#
|
|
85
|
+
# Takes the same keywords as {#generate}.
|
|
86
|
+
#
|
|
87
|
+
# @param inputs [Array<Hash>] Array of { text:, voice_id: } hashes
|
|
88
|
+
# @return [Hash] `{ audio:, alignment:, normalized_alignment:, voice_segments:, request_id: }`
|
|
89
|
+
def generate_with_timestamps(
|
|
90
|
+
inputs,
|
|
91
|
+
model_id: DEFAULT_MODEL,
|
|
92
|
+
language_code: nil,
|
|
93
|
+
settings: nil,
|
|
94
|
+
seed: nil,
|
|
95
|
+
output_format: 'mp3_44100_128',
|
|
96
|
+
apply_text_normalization: 'auto',
|
|
97
|
+
use_pvc_as_ivc: nil,
|
|
98
|
+
previous_text: nil,
|
|
99
|
+
future_text: nil,
|
|
100
|
+
previous_request_ids: nil,
|
|
101
|
+
next_request_ids: nil,
|
|
102
|
+
pronunciation_dictionary_locators: nil
|
|
103
|
+
)
|
|
104
|
+
validate_inputs!(inputs, model_id)
|
|
105
|
+
|
|
106
|
+
body = build_request_body(inputs, model_id, apply_text_normalization, optional_values(binding))
|
|
107
|
+
result = post_with_meta("/text-to-dialogue/with-timestamps?output_format=#{output_format}", body)
|
|
108
|
+
response = result[:body]
|
|
109
|
+
request_id = result[:headers]['request-id']
|
|
110
|
+
|
|
111
|
+
audio_data = Base64.decode64(response['audio_base64']) if response['audio_base64']
|
|
112
|
+
audio = (build_audio_response(audio_data, inputs, output_format, model_id, request_id: request_id) if audio_data)
|
|
113
|
+
|
|
114
|
+
{
|
|
115
|
+
audio: audio,
|
|
116
|
+
alignment: response['alignment'],
|
|
117
|
+
normalized_alignment: response['normalized_alignment'],
|
|
118
|
+
voice_segments: response['voice_segments'],
|
|
119
|
+
request_id: request_id
|
|
120
|
+
}
|
|
48
121
|
end
|
|
49
122
|
|
|
50
123
|
private
|
|
51
124
|
|
|
52
|
-
|
|
53
|
-
|
|
125
|
+
# The optional keyword values of the calling method, keyed by OPTIONAL_BODY_KEYS
|
|
126
|
+
def optional_values(caller_binding)
|
|
127
|
+
OPTIONAL_BODY_KEYS.to_h { |key| [key, caller_binding.local_variable_get(key)] }
|
|
128
|
+
end
|
|
129
|
+
|
|
130
|
+
def build_request_body(inputs, model_id, apply_text_normalization, options)
|
|
54
131
|
body = {
|
|
55
132
|
inputs: inputs.map { |i| { text: i[:text], voice_id: i[:voice_id] } },
|
|
56
133
|
model_id: model_id,
|
|
57
134
|
apply_text_normalization: apply_text_normalization
|
|
58
135
|
}
|
|
59
136
|
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
137
|
+
OPTIONAL_BODY_KEYS.each do |key|
|
|
138
|
+
body[key] = options[key] unless options[key].nil?
|
|
139
|
+
end
|
|
63
140
|
body
|
|
64
141
|
end
|
|
65
142
|
|
|
66
|
-
def build_audio_response(
|
|
143
|
+
def build_audio_response(data, inputs, output_format, model_id, request_id: nil)
|
|
67
144
|
total_text = inputs.map { |i| i[:text] }.join("\n")
|
|
68
145
|
total_chars = inputs.sum { |i| i[:text].length }
|
|
69
146
|
primary_voice = inputs.first[:voice_id]
|
|
70
147
|
|
|
71
148
|
audio = Objects::Audio.new(
|
|
72
|
-
data:
|
|
73
|
-
voice_id: primary_voice, text: total_text, model_id: model_id
|
|
149
|
+
data: data, format: output_format,
|
|
150
|
+
voice_id: primary_voice, text: total_text, model_id: model_id,
|
|
151
|
+
request_id: request_id
|
|
74
152
|
)
|
|
75
153
|
|
|
76
154
|
cost_info = Objects::CostInfo.new(
|
|
@@ -80,13 +158,14 @@ module ElevenRb
|
|
|
80
158
|
http_client.config.trigger(
|
|
81
159
|
:on_audio_generated,
|
|
82
160
|
audio: audio, voice_id: primary_voice,
|
|
83
|
-
text: total_text, cost_info: cost_info.to_h
|
|
161
|
+
text: total_text, cost_info: cost_info.to_h,
|
|
162
|
+
request_id: request_id
|
|
84
163
|
)
|
|
85
164
|
|
|
86
165
|
audio
|
|
87
166
|
end
|
|
88
167
|
|
|
89
|
-
def validate_inputs!(inputs)
|
|
168
|
+
def validate_inputs!(inputs, model_id = DEFAULT_MODEL)
|
|
90
169
|
raise Errors::ValidationError, 'inputs must be a non-empty array' unless inputs.is_a?(Array) && !inputs.empty?
|
|
91
170
|
|
|
92
171
|
inputs.each_with_index do |input, i|
|
|
@@ -101,12 +180,23 @@ module ElevenRb
|
|
|
101
180
|
"(got #{unique_voices.length})"
|
|
102
181
|
end
|
|
103
182
|
|
|
104
|
-
|
|
105
|
-
|
|
183
|
+
validate_text_length!(inputs.sum { |i| i[:text].length }, model_id)
|
|
184
|
+
end
|
|
185
|
+
|
|
186
|
+
def validate_text_length!(total_chars, model_id)
|
|
187
|
+
max_length = ModelCapabilities.max_text_length(model_id)
|
|
188
|
+
if total_chars > max_length
|
|
189
|
+
raise Errors::ValidationError,
|
|
190
|
+
"Total text length #{total_chars} exceeds maximum " \
|
|
191
|
+
"#{max_length} characters for #{model_id}"
|
|
192
|
+
end
|
|
193
|
+
|
|
194
|
+
return unless total_chars > RECOMMENDED_MAX_TEXT_LENGTH
|
|
106
195
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
196
|
+
http_client.config.logger&.warn(
|
|
197
|
+
"[ElevenRb] text-to-dialogue text is #{total_chars} characters; " \
|
|
198
|
+
"#{RECOMMENDED_MAX_TEXT_LENGTH} or fewer per request is recommended"
|
|
199
|
+
)
|
|
110
200
|
end
|
|
111
201
|
end
|
|
112
202
|
end
|
|
@@ -12,18 +12,67 @@ module ElevenRb
|
|
|
12
12
|
# client.tts.stream("Hello world", voice_id: "voice_id") do |chunk|
|
|
13
13
|
# io.write(chunk)
|
|
14
14
|
# end
|
|
15
|
+
#
|
|
16
|
+
# @example Eleven v4 with continuity and a fixed seed
|
|
17
|
+
# audio = client.tts.generate(
|
|
18
|
+
# "[sighs] Right. Let's try that again.",
|
|
19
|
+
# voice_id: "voice_id",
|
|
20
|
+
# model_id: "eleven_v4",
|
|
21
|
+
# seed: 42,
|
|
22
|
+
# previous_text: "That did not go to plan."
|
|
23
|
+
# )
|
|
24
|
+
# audio.request_id # => "abc123" (from the request-id response header)
|
|
15
25
|
class TextToSpeech < Base
|
|
16
26
|
DEFAULT_MODEL = 'eleven_multilingual_v2'
|
|
27
|
+
|
|
28
|
+
# Kept for compatibility. The per-request cap now comes from
|
|
29
|
+
# ModelCapabilities.max_text_length(model_id) (5,000 for eleven_v3).
|
|
17
30
|
MAX_TEXT_LENGTH = 5000
|
|
18
31
|
|
|
32
|
+
# Output formats the API accepts (documentation only; not validated)
|
|
19
33
|
OUTPUT_FORMATS = %w[
|
|
34
|
+
mp3_22050_32
|
|
35
|
+
mp3_24000_48
|
|
36
|
+
mp3_44100_32
|
|
37
|
+
mp3_44100_64
|
|
38
|
+
mp3_44100_96
|
|
20
39
|
mp3_44100_128
|
|
21
40
|
mp3_44100_192
|
|
41
|
+
opus_48000_32
|
|
42
|
+
opus_48000_64
|
|
43
|
+
opus_48000_96
|
|
44
|
+
opus_48000_128
|
|
45
|
+
opus_48000_192
|
|
46
|
+
pcm_8000
|
|
22
47
|
pcm_16000
|
|
23
48
|
pcm_22050
|
|
24
49
|
pcm_24000
|
|
50
|
+
pcm_32000
|
|
25
51
|
pcm_44100
|
|
52
|
+
pcm_48000
|
|
53
|
+
wav_8000
|
|
54
|
+
wav_16000
|
|
55
|
+
wav_22050
|
|
56
|
+
wav_24000
|
|
57
|
+
wav_32000
|
|
58
|
+
wav_44100
|
|
59
|
+
wav_48000
|
|
26
60
|
ulaw_8000
|
|
61
|
+
alaw_8000
|
|
62
|
+
].freeze
|
|
63
|
+
|
|
64
|
+
# Optional request-body keys, in the order they are written to the body.
|
|
65
|
+
# Each is omitted from the body when nil.
|
|
66
|
+
OPTIONAL_BODY_KEYS = %i[
|
|
67
|
+
language_code
|
|
68
|
+
apply_text_normalization
|
|
69
|
+
seed
|
|
70
|
+
previous_text
|
|
71
|
+
next_text
|
|
72
|
+
previous_request_ids
|
|
73
|
+
next_request_ids
|
|
74
|
+
pronunciation_dictionary_locators
|
|
75
|
+
use_pvc_as_ivc
|
|
27
76
|
].freeze
|
|
28
77
|
|
|
29
78
|
# Generate audio from text
|
|
@@ -31,30 +80,39 @@ module ElevenRb
|
|
|
31
80
|
# @param text [String] the text to convert
|
|
32
81
|
# @param voice_id [String] the voice ID to use
|
|
33
82
|
# @param model_id [String] the model to use (default: eleven_multilingual_v2)
|
|
34
|
-
# @param voice_settings [Hash] voice settings overrides
|
|
83
|
+
# @param voice_settings [Hash] voice settings overrides (keys the model ignores are dropped)
|
|
35
84
|
# @param output_format [String] audio output format
|
|
36
|
-
# @
|
|
37
|
-
|
|
38
|
-
|
|
85
|
+
# @param language_code [String, nil] ISO 639-1 language code to enforce
|
|
86
|
+
# @param apply_text_normalization [String, nil] "auto", "on" or "off"
|
|
87
|
+
# @param seed [Integer, nil] seed for reproducible generation
|
|
88
|
+
# @param previous_text [String, nil] text that comes before this request (continuity)
|
|
89
|
+
# @param next_text [String, nil] text that comes after this request (continuity)
|
|
90
|
+
# @param previous_request_ids [Array<String>, nil] request IDs of preceding generations
|
|
91
|
+
# @param next_request_ids [Array<String>, nil] request IDs of following generations
|
|
92
|
+
# @param pronunciation_dictionary_locators [Array<Hash>, nil] `{ pronunciation_dictionary_id:, version_id: }`
|
|
93
|
+
# @param use_pvc_as_ivc [Boolean, nil] use the IVC version of a professional voice
|
|
94
|
+
# @return [Objects::Audio] with request_id, character_cost and dropped_settings
|
|
95
|
+
def generate(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {}, output_format: 'mp3_44100_128',
|
|
96
|
+
language_code: nil, apply_text_normalization: nil, seed: nil, previous_text: nil,
|
|
97
|
+
next_text: nil, previous_request_ids: nil, next_request_ids: nil,
|
|
98
|
+
pronunciation_dictionary_locators: nil, use_pvc_as_ivc: nil)
|
|
99
|
+
validate_text!(text, model_id)
|
|
39
100
|
validate_presence!(voice_id, 'voice_id')
|
|
40
|
-
|
|
41
|
-
settings = Objects::VoiceSettings::DEFAULTS.merge(voice_settings)
|
|
42
|
-
|
|
43
|
-
body = {
|
|
44
|
-
text: text,
|
|
45
|
-
model_id: model_id,
|
|
46
|
-
voice_settings: settings
|
|
47
|
-
}
|
|
101
|
+
body, dropped = build_body(text, model_id, voice_settings, optional_values(binding))
|
|
48
102
|
|
|
49
103
|
path = "/text-to-speech/#{voice_id}?output_format=#{output_format}"
|
|
50
|
-
response =
|
|
104
|
+
response = post_binary_with_meta(path, body)
|
|
105
|
+
headers = response[:headers]
|
|
51
106
|
|
|
52
107
|
audio = Objects::Audio.new(
|
|
53
|
-
data: response,
|
|
108
|
+
data: response[:body],
|
|
54
109
|
format: output_format,
|
|
55
110
|
voice_id: voice_id,
|
|
56
111
|
text: text,
|
|
57
|
-
model_id: model_id
|
|
112
|
+
model_id: model_id,
|
|
113
|
+
request_id: headers['request-id'],
|
|
114
|
+
character_cost: integer_header(headers, 'character-cost'),
|
|
115
|
+
dropped_settings: dropped
|
|
58
116
|
)
|
|
59
117
|
|
|
60
118
|
# Trigger cost tracking callback
|
|
@@ -64,7 +122,8 @@ module ElevenRb
|
|
|
64
122
|
audio: audio,
|
|
65
123
|
voice_id: voice_id,
|
|
66
124
|
text: text,
|
|
67
|
-
cost_info: cost_info.to_h
|
|
125
|
+
cost_info: cost_info.to_h,
|
|
126
|
+
request_id: audio.request_id
|
|
68
127
|
)
|
|
69
128
|
|
|
70
129
|
audio
|
|
@@ -72,6 +131,8 @@ module ElevenRb
|
|
|
72
131
|
|
|
73
132
|
# Stream audio from text
|
|
74
133
|
#
|
|
134
|
+
# Takes the same optional keywords as {#generate}.
|
|
135
|
+
#
|
|
75
136
|
# @param text [String] the text to convert
|
|
76
137
|
# @param voice_id [String] the voice ID to use
|
|
77
138
|
# @param model_id [String] the model to use
|
|
@@ -79,18 +140,15 @@ module ElevenRb
|
|
|
79
140
|
# @param output_format [String] audio output format
|
|
80
141
|
# @yield [String] each chunk of audio data
|
|
81
142
|
# @return [void]
|
|
82
|
-
def stream(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {}, output_format: 'mp3_44100_128',
|
|
83
|
-
|
|
143
|
+
def stream(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {}, output_format: 'mp3_44100_128',
|
|
144
|
+
language_code: nil, apply_text_normalization: nil, seed: nil, previous_text: nil,
|
|
145
|
+
next_text: nil, previous_request_ids: nil, next_request_ids: nil,
|
|
146
|
+
pronunciation_dictionary_locators: nil, use_pvc_as_ivc: nil, &block)
|
|
147
|
+
validate_text!(text, model_id)
|
|
84
148
|
validate_presence!(voice_id, 'voice_id')
|
|
85
149
|
raise ArgumentError, 'Block required for streaming' unless block_given?
|
|
86
150
|
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
body = {
|
|
90
|
-
text: text,
|
|
91
|
-
model_id: model_id,
|
|
92
|
-
voice_settings: settings
|
|
93
|
-
}
|
|
151
|
+
body, = build_body(text, model_id, voice_settings, optional_values(binding))
|
|
94
152
|
|
|
95
153
|
path = "/text-to-speech/#{voice_id}/stream?output_format=#{output_format}"
|
|
96
154
|
post_stream(path, body, &block)
|
|
@@ -102,33 +160,35 @@ module ElevenRb
|
|
|
102
160
|
audio: nil, # No audio object for streaming
|
|
103
161
|
voice_id: voice_id,
|
|
104
162
|
text: text,
|
|
105
|
-
cost_info: cost_info.to_h
|
|
163
|
+
cost_info: cost_info.to_h,
|
|
164
|
+
request_id: nil
|
|
106
165
|
)
|
|
107
166
|
end
|
|
108
167
|
|
|
109
168
|
# Generate audio with timestamps
|
|
110
169
|
#
|
|
170
|
+
# Takes the same optional keywords as {#generate}.
|
|
171
|
+
#
|
|
111
172
|
# @param text [String] the text to convert
|
|
112
173
|
# @param voice_id [String] the voice ID to use
|
|
113
174
|
# @param model_id [String] the model to use
|
|
114
175
|
# @param voice_settings [Hash] voice settings overrides
|
|
115
176
|
# @param output_format [String] audio output format
|
|
116
|
-
# @return [Hash]
|
|
177
|
+
# @return [Hash] `{ audio:, alignment:, normalized_alignment:, request_id:, character_cost: }`
|
|
117
178
|
def generate_with_timestamps(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {},
|
|
118
|
-
output_format: 'mp3_44100_128'
|
|
119
|
-
|
|
179
|
+
output_format: 'mp3_44100_128', language_code: nil,
|
|
180
|
+
apply_text_normalization: nil, seed: nil, previous_text: nil,
|
|
181
|
+
next_text: nil, previous_request_ids: nil, next_request_ids: nil,
|
|
182
|
+
pronunciation_dictionary_locators: nil, use_pvc_as_ivc: nil)
|
|
183
|
+
validate_text!(text, model_id)
|
|
120
184
|
validate_presence!(voice_id, 'voice_id')
|
|
121
|
-
|
|
122
|
-
settings = Objects::VoiceSettings::DEFAULTS.merge(voice_settings)
|
|
123
|
-
|
|
124
|
-
body = {
|
|
125
|
-
text: text,
|
|
126
|
-
model_id: model_id,
|
|
127
|
-
voice_settings: settings
|
|
128
|
-
}
|
|
185
|
+
body, dropped = build_body(text, model_id, voice_settings, optional_values(binding))
|
|
129
186
|
|
|
130
187
|
path = "/text-to-speech/#{voice_id}/with-timestamps?output_format=#{output_format}"
|
|
131
|
-
|
|
188
|
+
result = post_with_meta(path, body)
|
|
189
|
+
response = result[:body]
|
|
190
|
+
request_id = result[:headers]['request-id']
|
|
191
|
+
character_cost = integer_header(result[:headers], 'character-cost')
|
|
132
192
|
|
|
133
193
|
# Decode base64 audio
|
|
134
194
|
audio_data = Base64.decode64(response['audio_base64']) if response['audio_base64']
|
|
@@ -139,25 +199,72 @@ module ElevenRb
|
|
|
139
199
|
format: output_format,
|
|
140
200
|
voice_id: voice_id,
|
|
141
201
|
text: text,
|
|
142
|
-
model_id: model_id
|
|
202
|
+
model_id: model_id,
|
|
203
|
+
request_id: request_id,
|
|
204
|
+
character_cost: character_cost,
|
|
205
|
+
dropped_settings: dropped
|
|
143
206
|
)
|
|
144
207
|
end
|
|
145
208
|
|
|
146
209
|
{
|
|
147
210
|
audio: audio,
|
|
148
|
-
alignment: response['alignment']
|
|
211
|
+
alignment: response['alignment'],
|
|
212
|
+
normalized_alignment: response['normalized_alignment'],
|
|
213
|
+
request_id: request_id,
|
|
214
|
+
character_cost: character_cost
|
|
149
215
|
}
|
|
150
216
|
end
|
|
151
217
|
|
|
152
218
|
private
|
|
153
219
|
|
|
154
|
-
|
|
220
|
+
# Shared request body for generate / stream / generate_with_timestamps
|
|
221
|
+
#
|
|
222
|
+
# @return [Array(Hash, Array<Symbol>)] the body and the dropped voice-setting keys
|
|
223
|
+
def build_body(text, model_id, voice_settings, options)
|
|
224
|
+
settings, dropped = resolve_voice_settings(model_id, voice_settings)
|
|
225
|
+
|
|
226
|
+
body = { text: text, model_id: model_id, voice_settings: settings }
|
|
227
|
+
OPTIONAL_BODY_KEYS.each do |key|
|
|
228
|
+
body[key] = options[key] unless options[key].nil?
|
|
229
|
+
end
|
|
230
|
+
|
|
231
|
+
[body, dropped]
|
|
232
|
+
end
|
|
233
|
+
|
|
234
|
+
# The optional keyword values of the calling method, keyed by OPTIONAL_BODY_KEYS
|
|
235
|
+
def optional_values(caller_binding)
|
|
236
|
+
OPTIONAL_BODY_KEYS.to_h { |key| [key, caller_binding.local_variable_get(key)] }
|
|
237
|
+
end
|
|
238
|
+
|
|
239
|
+
def resolve_voice_settings(model_id, voice_settings)
|
|
240
|
+
settings, dropped = Objects::VoiceSettings.for_model(model_id, voice_settings)
|
|
241
|
+
return [settings, dropped] if dropped.empty?
|
|
242
|
+
|
|
243
|
+
message = "voice settings #{dropped.join(', ')} are not supported by #{model_id} " \
|
|
244
|
+
"(supported: #{ModelCapabilities.supported_voice_settings(model_id).join(', ')})"
|
|
245
|
+
raise Errors::ValidationError, message if http_client.config.strict_voice_settings
|
|
246
|
+
|
|
247
|
+
http_client.config.logger&.warn("[ElevenRb] #{message}; dropped from the request")
|
|
248
|
+
[settings, dropped]
|
|
249
|
+
end
|
|
250
|
+
|
|
251
|
+
def integer_header(headers, name)
|
|
252
|
+
value = headers[name]
|
|
253
|
+
return nil if value.nil? || value.to_s.strip.empty?
|
|
254
|
+
|
|
255
|
+
Integer(value.to_s.strip, 10)
|
|
256
|
+
rescue ArgumentError
|
|
257
|
+
nil
|
|
258
|
+
end
|
|
259
|
+
|
|
260
|
+
def validate_text!(text, model_id = DEFAULT_MODEL)
|
|
155
261
|
validate_presence!(text, 'text')
|
|
156
262
|
|
|
157
|
-
|
|
263
|
+
max_length = ModelCapabilities.max_text_length(model_id)
|
|
264
|
+
return unless text.length > max_length
|
|
158
265
|
|
|
159
266
|
raise Errors::ValidationError,
|
|
160
|
-
"text exceeds maximum length of #{
|
|
267
|
+
"text exceeds maximum length of #{max_length} characters for #{model_id} (got #{text.length})"
|
|
161
268
|
end
|
|
162
269
|
end
|
|
163
270
|
end
|
data/lib/eleven_rb/version.rb
CHANGED
data/lib/eleven_rb.rb
CHANGED
|
@@ -60,6 +60,7 @@ module ElevenRb
|
|
|
60
60
|
retry_delay: config.retry_delay,
|
|
61
61
|
retry_statuses: config.retry_statuses,
|
|
62
62
|
logger: config.logger,
|
|
63
|
+
strict_voice_settings: config.strict_voice_settings,
|
|
63
64
|
on_request: config.on_request,
|
|
64
65
|
on_response: config.on_response,
|
|
65
66
|
on_error: config.on_error,
|
|
@@ -79,6 +80,7 @@ require_relative 'eleven_rb/errors'
|
|
|
79
80
|
require_relative 'eleven_rb/callbacks'
|
|
80
81
|
require_relative 'eleven_rb/instrumentation'
|
|
81
82
|
require_relative 'eleven_rb/configuration'
|
|
83
|
+
require_relative 'eleven_rb/model_capabilities'
|
|
82
84
|
|
|
83
85
|
# HTTP layer
|
|
84
86
|
require_relative 'eleven_rb/http/client'
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: eleven_rb
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.
|
|
4
|
+
version: 1.1.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Web Ventures Ltd
|
|
@@ -144,6 +144,7 @@ files:
|
|
|
144
144
|
- lib/eleven_rb/errors.rb
|
|
145
145
|
- lib/eleven_rb/http/client.rb
|
|
146
146
|
- lib/eleven_rb/instrumentation.rb
|
|
147
|
+
- lib/eleven_rb/model_capabilities.rb
|
|
147
148
|
- lib/eleven_rb/objects/audio.rb
|
|
148
149
|
- lib/eleven_rb/objects/base.rb
|
|
149
150
|
- lib/eleven_rb/objects/cost_info.rb
|