eleven_rb 0.4.0 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: fa71eff851a0c6b80f139e801962bceaf1bb371f6e9a0cd325c47b2ee6f4994c
4
- data.tar.gz: aa640970faba75afe3cbacdc16958ee63c1f2ecfb009512cd9e8182117e16d29
3
+ metadata.gz: 18424bceaba545d05f991cb1ae674fdc36572a7a9573ef769701deb35537b393
4
+ data.tar.gz: a5c2a3814295f5d19945b1731b639fe21ae0f0111ef51741fdb99bae5a978cf7
5
5
  SHA512:
6
- metadata.gz: a537ba9de014afc366c348a71613f257b6380ae784979fb42cd55522610c85661e0e34907c770642fa268051cbb91b24110e00d5e8b7305d5c913a81e04705a6
7
- data.tar.gz: c1f4e236fb327b737b4e6346bf3a71bfec357d4f7f09a0e1eaf9c8889ef2251ff40b8eeb3be959e98b58e2a76ff6305e7391d64f0c33a9ce1fb38897b958f1fa
6
+ metadata.gz: af3da66572e56a9173b5f427a3fa30212ebd686a5820b060ab5c81e5e41d60ae64a74e7956e0ddad4365d0f8e68ceefaa5572d97f485a0e424609c195f34be31
7
+ data.tar.gz: 3af1d07b75e5405724c50e87d3b89f88061d3540043716b38242acce0941f5152294148cd1d9109c20f05bbd05f5131e6b1225260bd373371b45061ad4bbca26
data/CHANGELOG.md CHANGED
@@ -7,6 +7,56 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [1.1.0] - 2026-09-30
11
+
12
+ ### Added
13
+
14
+ - Eleven v4 support (`eleven_v4`, `eleven_v4_turbo`) across Text-to-Speech and Text-to-Dialogue
15
+ - `ElevenRb::ModelCapabilities` — per model-family table of honoured voice settings, max text length, SSML `<break>` support, audio-tag support and continuity support (`for`, `supported_voice_settings`, `max_text_length`, `supports?`)
16
+ - `Objects::VoiceSettings.for_model(model_id, overrides)` → `[settings, dropped_keys]`, building only the settings a model honours
17
+ - `Configuration#strict_voice_settings` (default `false`): raise `ValidationError` instead of warning when a voice setting the model ignores is passed
18
+ - Optional TTS keywords on `generate`, `stream` and `generate_with_timestamps` (omitted from the body when nil): `language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`
19
+ - Response metadata: `Objects::Audio#request_id`, `#character_cost` (from the `request-id` / `character-cost` headers) and `#dropped_settings`; `on_audio_generated` also receives `request_id:` (nil for streams)
20
+ - `TextToSpeech#generate_with_timestamps` now also returns `normalized_alignment`, `request_id` and `character_cost`
21
+ - `TextToDialogue#generate_with_timestamps` (`POST /v1/text-to-dialogue/with-timestamps`) returning `audio`, `alignment`, `normalized_alignment`, `voice_segments` and `request_id`
22
+ - Text-to-Dialogue keywords `use_pvc_as_ivc`, `previous_text`, `future_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`; a logger warning above `RECOMMENDED_MAX_TEXT_LENGTH` (2,000 characters)
23
+ - `HTTP::Client#post(..., with_meta: true)` returning `{ body:, headers: }`, and `Resources::Base#post_with_meta` / `#post_binary_with_meta`
24
+ - `Models#find(model_id)`
25
+ - `Objects::Model#model_rates`, `#maximum_text_length_per_request`, `#requires_alpha_access`, `#supported_voice_settings`
26
+ - `CostInfo::COST_PER_1K_CHARS` entries for `eleven_v4` ($0.30), `eleven_v4_turbo` ($0.15) and `eleven_v3_conversational` ($0.15)
27
+ - `opus_*` output formats map to the `ogg` extension and `audio/ogg` content type
28
+
29
+ ### Changed
30
+
31
+ - Voice settings are filtered per model: known keys a model ignores (`stability`, `similarity_boost`, `style`, `use_speaker_boost`, `speed` minus the model's supported set) are dropped from the request (logged, and listed on `audio.dropped_settings`); unknown voice-setting keys pass through untouched on every model, so a new API field is never swallowed. `eleven_v3` / `eleven_v4` now send only `stability` and `similarity_boost` by default; `eleven_multilingual_v2` requests are byte-identical to 1.0.0. Override keys are symbolized and nil values removed (a nil override is never reported as dropped)
32
+ - Text length is capped per model (`ModelCapabilities.max_text_length`): 5,000 for `eleven_v3`, 10,000 for `eleven_v4` and `eleven_multilingual_v2`, 30,000 / 40,000 for the flash and turbo models. `TextToSpeech::MAX_TEXT_LENGTH` stays defined but is no longer the cap
33
+ - `TextToDialogue::DEFAULT_MODEL` is now `eleven_v4`
34
+ - Text-to-Dialogue's hard text cap now follows the model (10,000 characters on `eleven_v4`, 5,000 on `eleven_v3`) instead of a flat 5,000; `TextToDialogue::MAX_TEXT_LENGTH` stays defined but is no longer the cap
35
+ - `Models#latest` returns `eleven_v4`, else `eleven_v3`, else the default model
36
+ - `TextToSpeech::OUTPUT_FORMATS` refreshed to the current list (documentation only, not validated)
37
+ - Callbacks that declare their keywords explicitly (no `**rest`) receive only the keywords they declare, so callbacks written for 1.0.0 keep working as new keywords are added
38
+
39
+ ### Fixed
40
+
41
+ - `Models#list` (and everything built on it: `get`, `default`, `latest`, `multilingual`, `turbo`, `tts_capable`, `ids`, `TTSAdapter#list_models`) recursed until `SystemStackError`, because `Models#get(model_id)` shadowed `Resources::Base#get`
42
+
43
+ ## [1.0.0] - 2026-03-10
44
+
45
+ ### Added
46
+
47
+ - Text-to-Dialogue multi-speaker audio generation via `client.text_to_dialogue.generate` (`POST /v1/text-to-dialogue`)
48
+ - `Client#text_to_dialogue` resource with `dialogue` alias
49
+ - Multi-speaker input validation (max 10 unique voices, 5000 character limit)
50
+ - `eleven_v3` model added to `CostInfo::COST_PER_1K_CHARS` ($0.30/1K chars)
51
+ - `Models#latest` method returning the most capable model (`eleven_v3`)
52
+ - Audio tags support via v3 model (`[laughs]`, `[whispers]`, `[excited]`, etc.)
53
+ - `CostInfo` now accepts `character_count:` keyword as alternative to `text:`
54
+ - TTS generation with word-level timestamps via `client.tts.generate_with_timestamps`
55
+
56
+ ### Changed
57
+
58
+ - `CostInfo#initialize` signature: `text:` is now optional when `character_count:` is provided (backwards-compatible)
59
+
10
60
  ## [0.4.0] - 2026-03-10
11
61
 
12
62
  ### Added
data/README.md CHANGED
@@ -4,12 +4,14 @@
4
4
  [![CI](https://github.com/webventures/eleven_rb/actions/workflows/ci.yml/badge.svg)](https://github.com/webventures/eleven_rb/actions/workflows/ci.yml)
5
5
  [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
6
6
 
7
- A Ruby client for the [ElevenLabs](https://try.elevenlabs.io/qyk2j8gumrjz) Text-to-Speech, Speech-to-Speech, Sound Effects, and Music API.
7
+ A Ruby client for the [ElevenLabs](https://try.elevenlabs.io/qyk2j8gumrjz) Text-to-Speech, Speech-to-Speech, Text-to-Dialogue, Sound Effects, and Music API.
8
8
 
9
9
  ## Features
10
10
 
11
11
  - Text-to-Speech generation and streaming
12
12
  - Speech-to-Speech voice conversion
13
+ - Text-to-Dialogue multi-speaker generation with audio tags
14
+ - Eleven v4 support, with per-model voice-settings filtering and text caps
13
15
  - Sound effects generation from text descriptions
14
16
  - Music generation from prompts or composition plans
15
17
  - Voice management (list, get, create, update, delete)
@@ -73,7 +75,7 @@ audio.save_to_file("output.mp3")
73
75
  audio = client.tts.generate(
74
76
  "Hello world",
75
77
  voice_id: "voice_id",
76
- model_id: "eleven_multilingual_v2",
78
+ model_id: "eleven_v4", # Most expressive; audio tags; see "Eleven v4" below
77
79
  voice_settings: {
78
80
  stability: 0.5,
79
81
  similarity_boost: 0.75
@@ -87,8 +89,72 @@ File.open("output.mp3", "wb") do |file|
87
89
  file.write(chunk)
88
90
  end
89
91
  end
92
+
93
+ # Word-level timestamps
94
+ result = client.tts.generate_with_timestamps("Hello world", voice_id: "voice_id")
95
+ result[:audio] # => ElevenRb::Objects::Audio
96
+ result[:alignment] # => { "characters" => [...], "character_start_times_seconds" => [...], ... }
97
+ result[:normalized_alignment]
98
+ result[:request_id] # from the request-id response header
99
+ result[:character_cost] # Integer, from the character-cost response header
100
+ ```
101
+
102
+ Optional keywords on `generate`, `stream` and `generate_with_timestamps` (each is left out of the request when nil):
103
+ `language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`,
104
+ `next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`.
105
+
106
+ ### Eleven v4
107
+
108
+ `eleven_v4` (and the cheaper, faster `eleven_v4_turbo`) take up to 10,000 characters per request.
109
+
110
+ ```ruby
111
+ audio = client.tts.generate(
112
+ "[sighs] Right... let's try that one more time.",
113
+ voice_id: "voice_id",
114
+ model_id: "eleven_v4",
115
+ voice_settings: { stability: 0.5, similarity_boost: 0.75 },
116
+ seed: 42, # reproducible re-rolls
117
+ previous_text: "That did not go to plan.", # continuity with the take before
118
+ next_text: "Here we go." # ...and the one after
119
+ )
120
+ audio.request_id # pass as previous_request_ids: / next_request_ids: on neighbouring takes
121
+ audio.character_cost
122
+ audio.dropped_settings # => [] (voice settings the gem removed, see below)
90
123
  ```
91
124
 
125
+ - **Only `stability` and `similarity_boost` take effect.** The API accepts `speed`, `style` and
126
+ `use_speaker_boost` on v4 and silently ignores them, so the gem drops them from the request, logs a
127
+ warning through `logger`, and lists them on `audio.dropped_settings`. Set
128
+ `strict_voice_settings: true` on the client to raise `ElevenRb::Errors::ValidationError` instead.
129
+ Only those known keys are ever dropped: a voice-setting key the gem does not know passes through
130
+ unchanged on every model, so new API fields keep working.
131
+ Pace a v4 take with the text (ellipses, pause tags), not `speed`.
132
+ - **SSML `<break time="…"/>` tags are ignored** by v4 (no pause is produced). Use an audio tag or
133
+ punctuation for pauses.
134
+ - **Audio tags** go in square brackets inside the text: `[sighs]`, `[whispers]`, `[laughs]`,
135
+ `[short pause]`. Note that tags appear in the returned alignment like any other characters.
136
+ - **Continuity**: `previous_text` / `next_text` (or `previous_request_ids` / `next_request_ids`) tell the
137
+ model what surrounds a take, so stitched takes keep a consistent delivery.
138
+ - **`seed`** makes v4 generation reproducible, so a re-roll changes one thing at a time.
139
+
140
+ `ElevenRb::ModelCapabilities` answers these questions for any model id:
141
+
142
+ ```ruby
143
+ ElevenRb::ModelCapabilities.supported_voice_settings("eleven_v4") # => [:stability, :similarity_boost]
144
+ ElevenRb::ModelCapabilities.max_text_length("eleven_v3") # => 5000
145
+ ElevenRb::ModelCapabilities.supports?("eleven_v4", :audio_tags) # => true
146
+ ElevenRb::ModelCapabilities.supports?("eleven_v4", :ssml_break) # => false
147
+ ```
148
+
149
+ | Model family | Voice settings honoured | Max chars | Audio tags | SSML breaks |
150
+ |---|---|---|---|---|
151
+ | `eleven_v4`, `eleven_v4_turbo` | stability, similarity_boost | 10,000 | yes | no |
152
+ | `eleven_v3` | stability, similarity_boost | 5,000 | yes | no |
153
+ | `eleven_v3_conversational` | stability, similarity_boost, use_speaker_boost | 5,000 | yes | no |
154
+ | `eleven_flash_v2_5`, `eleven_turbo_v2_5` | all five (incl. style, speed) | 40,000 | no | yes |
155
+ | `eleven_flash_v2`, `eleven_turbo_v2` | all five | 30,000 | no | yes |
156
+ | `eleven_multilingual_v2` and others | all five | 10,000 | no | yes |
157
+
92
158
  ### Speech-to-Speech
93
159
 
94
160
  ```ruby
@@ -111,6 +177,54 @@ io = File.open("input.mp3", "rb")
111
177
  audio = client.sts.convert(io, voice_id: "voice_id")
112
178
  ```
113
179
 
180
+ ### Text-to-Dialogue
181
+
182
+ ```ruby
183
+ # Generate multi-speaker dialogue
184
+ audio = client.text_to_dialogue.generate([
185
+ { text: "[excited] Welcome to the show!", voice_id: "voice_abc" },
186
+ { text: "[laughs] Thanks for having me.", voice_id: "voice_xyz" },
187
+ { text: "So tell us about your project...", voice_id: "voice_abc" }
188
+ ])
189
+ audio.save_to_file("dialogue.mp3")
190
+
191
+ # With options (the default model is eleven_v4)
192
+ audio = client.dialogue.generate(
193
+ inputs,
194
+ model_id: "eleven_v4",
195
+ language_code: "en",
196
+ settings: { stability: 0.5, similarity: 0.75 }, # sent to the API unchanged
197
+ seed: 42,
198
+ previous_text: "Earlier in the scene...",
199
+ output_format: "mp3_44100_192"
200
+ )
201
+
202
+ # With timestamps and per-speaker segments
203
+ result = client.dialogue.generate_with_timestamps(inputs)
204
+ result[:audio] # => ElevenRb::Objects::Audio
205
+ result[:alignment] # character timings
206
+ result[:voice_segments] # which voice speaks when
207
+ result[:request_id]
208
+ ```
209
+
210
+ The hard text cap follows the model (10,000 characters for `eleven_v4`, 5,000 for `eleven_v3`); above
211
+ 2,000 characters the gem logs a warning, since shorter dialogue requests give more reliable results.
212
+
213
+ ### Audio Tags
214
+
215
+ The `eleven_v4` and `eleven_v3` models support inline audio tags, in square brackets, for expressive speech
216
+ (`ElevenRb::ModelCapabilities.supports?(model_id, :audio_tags)`):
217
+
218
+ ```ruby
219
+ audio = client.tts.generate(
220
+ "[excited] Oh wow, this is AMAZING! [laughs] I can't believe it...",
221
+ voice_id: "voice_id",
222
+ model_id: "eleven_v4"
223
+ )
224
+ ```
225
+
226
+ Supported tags include `[laughs]`, `[whispers]`, `[sighs]`, `[excited]`, `[sarcastic]`, `[curious]`, `[short pause]`, and more. Use CAPS for emphasis, `...` for pauses, and `—` for interruptions. See the [ElevenLabs audio tags documentation](https://elevenlabs.io/docs/guides/audio-tags) for the full list.
227
+
114
228
  ### Sound Effects
115
229
 
116
230
  ```ruby
@@ -252,7 +366,7 @@ client = ElevenRb::Client.new(
252
366
  Sentry.capture_exception(error, extra: { path: path })
253
367
  },
254
368
 
255
- # Cost tracking
369
+ # Cost tracking (add request_id: to also receive the API's request id)
256
370
  on_audio_generated: ->(audio:, voice_id:, text:, cost_info:) {
257
371
  UsageRecord.create!(
258
372
  characters: cost_info[:character_count],
@@ -274,6 +388,12 @@ client = ElevenRb::Client.new(
274
388
  models = client.models.list
275
389
  models.each { |m| puts "#{m.name} (#{m.model_id})" }
276
390
 
391
+ # Find one model
392
+ client.models.find("eleven_v4") # => ElevenRb::Objects::Model (get is an alias)
393
+
394
+ # Get the latest/most capable model: eleven_v4, else eleven_v3, else the default
395
+ client.models.latest.model_id # => "eleven_v4"
396
+
277
397
  # Get multilingual models
278
398
  client.models.multilingual
279
399
 
@@ -303,7 +423,8 @@ client = ElevenRb::Client.new(
303
423
  open_timeout: 10, # Connection timeout
304
424
  max_retries: 3, # Max retry attempts
305
425
  retry_delay: 1.0, # Base delay between retries
306
- logger: Rails.logger # Optional logger
426
+ logger: Rails.logger, # Optional logger (receives dropped-setting warnings)
427
+ strict_voice_settings: false # true: raise instead of dropping settings a model ignores
307
428
  )
308
429
  ```
309
430
 
@@ -28,6 +28,10 @@ module ElevenRb
28
28
 
29
29
  # Trigger a callback if it's configured
30
30
  #
31
+ # A callback that names its keywords explicitly (no `**rest`) receives only
32
+ # the keywords it declares, so callbacks written before a keyword was added
33
+ # (e.g. `request_id:` on on_audio_generated in 1.1.0) keep working.
34
+ #
31
35
  # @param callback_name [Symbol] the name of the callback
32
36
  # @param kwargs [Hash] keyword arguments to pass to the callback
33
37
  # @return [Object, nil] the return value of the callback, or nil
@@ -36,12 +40,32 @@ module ElevenRb
36
40
  return unless callback.respond_to?(:call)
37
41
 
38
42
  begin
39
- callback.call(**kwargs)
43
+ callback.call(**accepted_callback_kwargs(callback, kwargs))
40
44
  rescue StandardError => e
41
45
  # Don't let callback errors break the main flow
42
46
  warn "[ElevenRb] Callback error in #{callback_name}: #{e.message}"
43
47
  nil
44
48
  end
45
49
  end
50
+
51
+ private
52
+
53
+ def accepted_callback_kwargs(callback, kwargs)
54
+ params = callback_parameters(callback)
55
+ return kwargs if params.nil? || params.any? { |type, _| type == :keyrest }
56
+
57
+ accepted = params.filter_map { |type, name| name if %i[key keyreq].include?(type) }
58
+ return kwargs if accepted.empty?
59
+
60
+ kwargs.slice(*accepted)
61
+ end
62
+
63
+ def callback_parameters(callback)
64
+ return callback.parameters if callback.respond_to?(:parameters)
65
+
66
+ callback.method(:call).parameters
67
+ rescue NameError
68
+ nil
69
+ end
46
70
  end
47
71
  end
@@ -101,6 +101,14 @@ module ElevenRb
101
101
  @music ||= Resources::Music.new(http_client)
102
102
  end
103
103
 
104
+ # Text-to-dialogue resource
105
+ #
106
+ # @return [Resources::TextToDialogue]
107
+ def text_to_dialogue
108
+ @text_to_dialogue ||= Resources::TextToDialogue.new(http_client)
109
+ end
110
+ alias dialogue text_to_dialogue
111
+
104
112
  # Voice slot manager
105
113
  #
106
114
  # @return [VoiceSlotManager]
@@ -21,12 +21,13 @@ module ElevenRb
21
21
  open_timeout: 10,
22
22
  max_retries: 3,
23
23
  retry_delay: 1.0,
24
- retry_statuses: [429, 500, 502, 503, 504].freeze
24
+ retry_statuses: [429, 500, 502, 503, 504].freeze,
25
+ strict_voice_settings: false
25
26
  }.freeze
26
27
 
27
28
  attr_accessor :api_key, :base_url, :timeout, :open_timeout,
28
29
  :max_retries, :retry_delay, :retry_statuses,
29
- :logger
30
+ :logger, :strict_voice_settings
30
31
 
31
32
  # Initialize a new configuration
32
33
  #
@@ -39,6 +40,8 @@ module ElevenRb
39
40
  # @option options [Float] :retry_delay Base delay between retries in seconds (default: 1.0)
40
41
  # @option options [Array<Integer>] :retry_statuses HTTP status codes to retry (default: [429, 500, 502, 503, 504])
41
42
  # @option options [Logger] :logger Logger instance for debug output
43
+ # @option options [Boolean] :strict_voice_settings Raise instead of warn when a voice setting the
44
+ # model does not honour is passed (default: false)
42
45
  # @option options [Proc] :on_request Callback before each request
43
46
  # @option options [Proc] :on_response Callback after successful response
44
47
  # @option options [Proc] :on_error Callback when an error occurs
@@ -87,7 +90,8 @@ module ElevenRb
87
90
  timeout: timeout,
88
91
  open_timeout: open_timeout,
89
92
  max_retries: max_retries,
90
- retry_delay: retry_delay
93
+ retry_delay: retry_delay,
94
+ strict_voice_settings: strict_voice_settings
91
95
  }
92
96
  end
93
97
  end
@@ -32,9 +32,11 @@ module ElevenRb
32
32
  # @param path [String] the API path
33
33
  # @param body [Hash] request body
34
34
  # @param response_type [Symbol] :json or :binary
35
- # @return [Hash, Array, String] parsed response
36
- def post(path, body = {}, response_type: :json)
37
- request(:post, path, body: body, response_type: response_type)
35
+ # @param with_meta [Boolean] when true, return `{ body:, headers: }` instead of the body alone
36
+ # (headers as a Hash of lower-cased String keys to String values)
37
+ # @return [Hash, Array, String] parsed response (or `{ body:, headers: }` with with_meta)
38
+ def post(path, body = {}, response_type: :json, with_meta: false)
39
+ request(:post, path, body: body, response_type: response_type, with_meta: with_meta)
38
40
  end
39
41
 
40
42
  # Make a DELETE request
@@ -68,7 +70,7 @@ module ElevenRb
68
70
  private
69
71
 
70
72
  def request(method, path, body: nil, params: nil, response_type: :json, multipart: false, stream: false,
71
- attempt: 1, &block)
73
+ attempt: 1, with_meta: false, &block)
72
74
  config.validate!
73
75
  url = "#{config.base_url}#{path}"
74
76
  start_time = Time.now
@@ -86,15 +88,12 @@ module ElevenRb
86
88
  # Trigger after response callback
87
89
  config.trigger(:on_response, method: method, path: path, response: response, duration: duration)
88
90
 
89
- # Return binary data directly
90
- return response.body if response_type == :binary && response.success?
91
-
92
- handle_response(response)
91
+ build_result(response, response_type, with_meta)
93
92
  rescue Errors::RateLimitError => e
94
93
  config.trigger(:on_rate_limit, retry_after: e.retry_after, error: e)
95
- handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
94
+ handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
96
95
  rescue Errors::ServerError => e
97
- handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
96
+ handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
98
97
  rescue Errors::Base => e
99
98
  config.trigger(:on_error, error: e, method: method, path: path,
100
99
  context: { body: sanitize_body_for_logging(body) })
@@ -107,6 +106,17 @@ module ElevenRb
107
106
  end
108
107
  end
109
108
 
109
+ def build_result(response, response_type, with_meta)
110
+ # Return binary data directly
111
+ result = if response_type == :binary && response.success?
112
+ response.body
113
+ else
114
+ handle_response(response)
115
+ end
116
+
117
+ with_meta ? { body: result, headers: normalize_headers(response) } : result
118
+ end
119
+
110
120
  def execute_request(method, url, body, params, multipart, stream, &block)
111
121
  options = build_options(body, params, multipart, stream, &block)
112
122
 
@@ -210,7 +220,7 @@ module ElevenRb
210
220
  raise error_class.new(message, **error_kwargs)
211
221
  end
212
222
 
213
- def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, &block)
223
+ def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
214
224
  raise error if attempt > config.max_retries || !config.retry_statuses.include?(error.http_status)
215
225
 
216
226
  delay = if error.is_a?(Errors::RateLimitError) && error.retry_after
@@ -224,7 +234,17 @@ module ElevenRb
224
234
  sleep(delay)
225
235
 
226
236
  request(method, path,
227
- body: body, params: params, response_type: response_type, multipart: multipart, stream: stream, attempt: attempt + 1, &block)
237
+ body: body, params: params, response_type: response_type, multipart: multipart, stream: stream,
238
+ attempt: attempt + 1, with_meta: with_meta, &block)
239
+ end
240
+
241
+ # Response headers as a plain Hash of lower-cased String keys to String values
242
+ # (HTTParty exposes multi-value arrays; the first value is kept)
243
+ def normalize_headers(response)
244
+ raw = response.headers.respond_to?(:to_hash) ? response.headers.to_hash : response.headers.to_h
245
+ raw.each_with_object({}) do |(key, value), out|
246
+ out[key.to_s.downcase] = value.is_a?(Array) ? value.first : value
247
+ end
228
248
  end
229
249
 
230
250
  def wrap_error(error)
@@ -0,0 +1,103 @@
1
+ # frozen_string_literal: true
2
+
3
+ module ElevenRb
4
+ # What each ElevenLabs model family accepts, keyed by model-id family.
5
+ #
6
+ # The API silently ignores voice settings a model does not honour (Eleven v4
7
+ # accepts `speed`, `style` and `use_speaker_boost` and does nothing with them),
8
+ # so the gem uses this table to send only the settings that take effect, to
9
+ # cap text length per model, and to answer feature questions (audio tags,
10
+ # SSML `<break>` tags, request continuity via previous_text / next_text).
11
+ #
12
+ # @example
13
+ # ElevenRb::ModelCapabilities.supported_voice_settings('eleven_v4')
14
+ # # => [:stability, :similarity_boost]
15
+ # ElevenRb::ModelCapabilities.max_text_length('eleven_flash_v2_5') # => 40_000
16
+ # ElevenRb::ModelCapabilities.supports?('eleven_v4', :audio_tags) # => true
17
+ module ModelCapabilities
18
+ # Capability record for one model family
19
+ Capabilities = Struct.new(:voice_settings, :max_text_length, :ssml_break, :audio_tags, :continuity,
20
+ keyword_init: true) do
21
+ def ssml_break? = ssml_break
22
+ def audio_tags? = audio_tags
23
+ def continuity? = continuity
24
+ end
25
+
26
+ ALL_VOICE_SETTINGS = %i[stability similarity_boost style use_speaker_boost speed].freeze
27
+
28
+ def self.build(voice_settings:, max_text_length:, ssml_break:, audio_tags:, continuity:)
29
+ Capabilities.new(
30
+ voice_settings: voice_settings.freeze,
31
+ max_text_length: max_text_length,
32
+ ssml_break: ssml_break,
33
+ audio_tags: audio_tags,
34
+ continuity: continuity
35
+ ).freeze
36
+ end
37
+ private_class_method :build
38
+
39
+ V4 = build(voice_settings: %i[stability similarity_boost], max_text_length: 10_000,
40
+ ssml_break: false, audio_tags: true, continuity: true)
41
+ V3_CONVERSATIONAL = build(voice_settings: %i[stability similarity_boost use_speaker_boost], max_text_length: 5_000,
42
+ ssml_break: false, audio_tags: true, continuity: true)
43
+ V3 = build(voice_settings: %i[stability similarity_boost], max_text_length: 5_000,
44
+ ssml_break: false, audio_tags: true, continuity: true)
45
+ V2_5_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 40_000,
46
+ ssml_break: true, audio_tags: false, continuity: true)
47
+ V2_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 30_000,
48
+ ssml_break: true, audio_tags: false, continuity: true)
49
+ DEFAULT = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 10_000,
50
+ ssml_break: true, audio_tags: false, continuity: true)
51
+
52
+ # Ordered [matcher, capabilities] pairs; the first match wins.
53
+ FAMILIES = [
54
+ [/\Aeleven_v4/, V4],
55
+ [/\Aeleven_v3_conversational/, V3_CONVERSATIONAL],
56
+ [/\Aeleven_v3/, V3],
57
+ [/\Aeleven_(flash|turbo)_v2_5/, V2_5_FAST],
58
+ [/\Aeleven_(flash|turbo)_v2/, V2_FAST]
59
+ ].freeze
60
+
61
+ FEATURES = %i[audio_tags ssml_break continuity].freeze
62
+
63
+ module_function
64
+
65
+ # Capabilities for a model id (unknown ids get the permissive default)
66
+ #
67
+ # @param model_id [String, Symbol, nil]
68
+ # @return [Capabilities] frozen
69
+ def for(model_id)
70
+ id = model_id.to_s
71
+ FAMILIES.each { |matcher, caps| return caps if matcher.match?(id) }
72
+ DEFAULT
73
+ end
74
+
75
+ # Voice settings keys the model honours
76
+ #
77
+ # @param model_id [String]
78
+ # @return [Array<Symbol>]
79
+ def supported_voice_settings(model_id)
80
+ self.for(model_id).voice_settings
81
+ end
82
+
83
+ # Maximum characters per request for the model
84
+ #
85
+ # @param model_id [String]
86
+ # @return [Integer]
87
+ def max_text_length(model_id)
88
+ self.for(model_id).max_text_length
89
+ end
90
+
91
+ # Whether the model supports a feature
92
+ #
93
+ # @param model_id [String]
94
+ # @param feature [Symbol] :audio_tags, :ssml_break or :continuity
95
+ # @return [Boolean]
96
+ def supports?(model_id, feature)
97
+ feature = feature.to_sym
98
+ raise ArgumentError, "Unknown feature #{feature.inspect} (expected one of #{FEATURES.join(', ')})" unless FEATURES.include?(feature)
99
+
100
+ self.for(model_id).public_send(feature) ? true : false
101
+ end
102
+ end
103
+ end
@@ -6,7 +6,8 @@ module ElevenRb
6
6
  module Objects
7
7
  # Represents generated audio data
8
8
  class Audio
9
- attr_reader :data, :format, :voice_id, :text, :model_id
9
+ attr_reader :data, :format, :voice_id, :text, :model_id,
10
+ :request_id, :character_cost, :dropped_settings
10
11
 
11
12
  # Initialize audio object
12
13
  #
@@ -15,12 +16,20 @@ module ElevenRb
15
16
  # @param voice_id [String] the voice ID used
16
17
  # @param text [String] the text that was converted
17
18
  # @param model_id [String, nil] the model ID used
18
- def initialize(data:, format:, voice_id:, text:, model_id: nil)
19
+ # @param request_id [String, nil] the API's `request-id` response header (usable as
20
+ # previous_request_ids / next_request_ids on a later request)
21
+ # @param character_cost [Integer, nil] the API's `character-cost` response header
22
+ # @param dropped_settings [Array<Symbol>] voice settings the gem dropped because the model ignores them
23
+ def initialize(data:, format:, voice_id:, text:, model_id: nil, request_id: nil, character_cost: nil,
24
+ dropped_settings: [])
19
25
  @data = data
20
26
  @format = format
21
27
  @voice_id = voice_id
22
28
  @text = text
23
29
  @model_id = model_id
30
+ @request_id = request_id
31
+ @character_cost = character_cost
32
+ @dropped_settings = Array(dropped_settings).dup.freeze
24
33
  end
25
34
 
26
35
  # Save audio to a file
@@ -62,7 +71,7 @@ module ElevenRb
62
71
  'audio/mpeg'
63
72
  when /pcm/
64
73
  'audio/pcm'
65
- when /ogg/
74
+ when /ogg|opus/
66
75
  'audio/ogg'
67
76
  when /wav/
68
77
  'audio/wav'
@@ -82,7 +91,7 @@ module ElevenRb
82
91
  'mp3'
83
92
  when /pcm/
84
93
  'pcm'
85
- when /ogg/
94
+ when /ogg|opus/
86
95
  'ogg'
87
96
  when /wav/
88
97
  'wav'
@@ -12,6 +12,10 @@ module ElevenRb
12
12
  'eleven_monolingual_v1' => 0.30,
13
13
  'eleven_multilingual_v1' => 0.30,
14
14
  'eleven_multilingual_v2' => 0.30,
15
+ 'eleven_v3' => 0.30,
16
+ 'eleven_v3_conversational' => 0.15,
17
+ 'eleven_v4' => 0.30,
18
+ 'eleven_v4_turbo' => 0.15,
15
19
  'eleven_turbo_v2' => 0.18,
16
20
  'eleven_turbo_v2_5' => 0.18,
17
21
  'eleven_english_sts_v2' => 0.30,
@@ -23,11 +27,12 @@ module ElevenRb
23
27
 
24
28
  # Initialize cost info
25
29
  #
26
- # @param text [String] the text being converted
30
+ # @param text [String, nil] the text being converted
31
+ # @param character_count [Integer, nil] direct character count (alternative to text)
27
32
  # @param voice_id [String] the voice ID
28
33
  # @param model_id [String] the model ID
29
- def initialize(text:, voice_id:, model_id:)
30
- @character_count = text.length
34
+ def initialize(voice_id:, model_id:, text: nil, character_count: nil)
35
+ @character_count = character_count || text&.length || 0
31
36
  @voice_id = voice_id
32
37
  @model_id = model_id
33
38
  end
@@ -18,6 +18,16 @@ module ElevenRb
18
18
  attribute :max_characters_request_free_user
19
19
  attribute :max_characters_request_subscribed_user
20
20
  attribute :concurrency_group
21
+ attribute :model_rates
22
+ attribute :maximum_text_length_per_request
23
+ attribute :requires_alpha_access, type: :boolean
24
+
25
+ # Voice settings keys this model honours (from ModelCapabilities)
26
+ #
27
+ # @return [Array<Symbol>]
28
+ def supported_voice_settings
29
+ ModelCapabilities.supported_voice_settings(model_id)
30
+ end
21
31
 
22
32
  # Check if this model supports a given language
23
33
  #