eleven_rb 1.0.0 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: ed711abcce18771ad13f10bcb29754605be61f7d02f7114f0e0b28b0dad4d556
4
- data.tar.gz: 146285726bc80b0c3eab0b307a7ec4b788a8f3465903992bb12fc2b34bc1694b
3
+ metadata.gz: 18424bceaba545d05f991cb1ae674fdc36572a7a9573ef769701deb35537b393
4
+ data.tar.gz: a5c2a3814295f5d19945b1731b639fe21ae0f0111ef51741fdb99bae5a978cf7
5
5
  SHA512:
6
- metadata.gz: 6bf8e216c83287bb099e4a6bbed4ef718329f361fb7dfb4c70bf122f2512c74916eb1540fe6a1dfd4ae01e0edc53edc05408a017946f504c09611a54a6c2370b
7
- data.tar.gz: 1839c52e3adf4efed58c410f08fa5c5e4818fde0964922e0752019b6606d726ed66a4f48a767b4e2110aa3b0cf98e7f4666eaef64552ec9f0f976378b1ef5094
6
+ metadata.gz: af3da66572e56a9173b5f427a3fa30212ebd686a5820b060ab5c81e5e41d60ae64a74e7956e0ddad4365d0f8e68ceefaa5572d97f485a0e424609c195f34be31
7
+ data.tar.gz: 3af1d07b75e5405724c50e87d3b89f88061d3540043716b38242acce0941f5152294148cd1d9109c20f05bbd05f5131e6b1225260bd373371b45061ad4bbca26
data/CHANGELOG.md CHANGED
@@ -7,6 +7,39 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [1.1.0] - 2026-09-30
11
+
12
+ ### Added
13
+
14
+ - Eleven v4 support (`eleven_v4`, `eleven_v4_turbo`) across Text-to-Speech and Text-to-Dialogue
15
+ - `ElevenRb::ModelCapabilities` — per model-family table of honoured voice settings, max text length, SSML `<break>` support, audio-tag support and continuity support (`for`, `supported_voice_settings`, `max_text_length`, `supports?`)
16
+ - `Objects::VoiceSettings.for_model(model_id, overrides)` → `[settings, dropped_keys]`, building only the settings a model honours
17
+ - `Configuration#strict_voice_settings` (default `false`): raise `ValidationError` instead of warning when a voice setting the model ignores is passed
18
+ - Optional TTS keywords on `generate`, `stream` and `generate_with_timestamps` (omitted from the body when nil): `language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`
19
+ - Response metadata: `Objects::Audio#request_id`, `#character_cost` (from the `request-id` / `character-cost` headers) and `#dropped_settings`; `on_audio_generated` also receives `request_id:` (nil for streams)
20
+ - `TextToSpeech#generate_with_timestamps` now also returns `normalized_alignment`, `request_id` and `character_cost`
21
+ - `TextToDialogue#generate_with_timestamps` (`POST /v1/text-to-dialogue/with-timestamps`) returning `audio`, `alignment`, `normalized_alignment`, `voice_segments` and `request_id`
22
+ - Text-to-Dialogue keywords `use_pvc_as_ivc`, `previous_text`, `future_text`, `previous_request_ids`, `next_request_ids`, `pronunciation_dictionary_locators`; a logger warning above `RECOMMENDED_MAX_TEXT_LENGTH` (2,000 characters)
23
+ - `HTTP::Client#post(..., with_meta: true)` returning `{ body:, headers: }`, and `Resources::Base#post_with_meta` / `#post_binary_with_meta`
24
+ - `Models#find(model_id)`
25
+ - `Objects::Model#model_rates`, `#maximum_text_length_per_request`, `#requires_alpha_access`, `#supported_voice_settings`
26
+ - `CostInfo::COST_PER_1K_CHARS` entries for `eleven_v4` ($0.30), `eleven_v4_turbo` ($0.15) and `eleven_v3_conversational` ($0.15)
27
+ - `opus_*` output formats map to the `ogg` extension and `audio/ogg` content type
28
+
29
+ ### Changed
30
+
31
+ - Voice settings are filtered per model: known keys a model ignores (`stability`, `similarity_boost`, `style`, `use_speaker_boost`, `speed` minus the model's supported set) are dropped from the request (logged, and listed on `audio.dropped_settings`); unknown voice-setting keys pass through untouched on every model, so a new API field is never swallowed. `eleven_v3` / `eleven_v4` now send only `stability` and `similarity_boost` by default; `eleven_multilingual_v2` requests are byte-identical to 1.0.0. Override keys are symbolized and nil values removed (a nil override is never reported as dropped)
32
+ - Text length is capped per model (`ModelCapabilities.max_text_length`): 5,000 for `eleven_v3`, 10,000 for `eleven_v4` and `eleven_multilingual_v2`, 30,000 / 40,000 for the flash and turbo models. `TextToSpeech::MAX_TEXT_LENGTH` stays defined but is no longer the cap
33
+ - `TextToDialogue::DEFAULT_MODEL` is now `eleven_v4`
34
+ - Text-to-Dialogue's hard text cap now follows the model (10,000 characters on `eleven_v4`, 5,000 on `eleven_v3`) instead of a flat 5,000; `TextToDialogue::MAX_TEXT_LENGTH` stays defined but is no longer the cap
35
+ - `Models#latest` returns `eleven_v4`, else `eleven_v3`, else the default model
36
+ - `TextToSpeech::OUTPUT_FORMATS` refreshed to the current list (documentation only, not validated)
37
+ - Callbacks that declare their keywords explicitly (no `**rest`) receive only the keywords they declare, so callbacks written for 1.0.0 keep working as new keywords are added
38
+
39
+ ### Fixed
40
+
41
+ - `Models#list` (and everything built on it: `get`, `default`, `latest`, `multilingual`, `turbo`, `tts_capable`, `ids`, `TTSAdapter#list_models`) recursed until `SystemStackError`, because `Models#get(model_id)` shadowed `Resources::Base#get`
42
+
10
43
  ## [1.0.0] - 2026-03-10
11
44
 
12
45
  ### Added
data/README.md CHANGED
@@ -11,6 +11,7 @@ A Ruby client for the [ElevenLabs](https://try.elevenlabs.io/qyk2j8gumrjz) Text-
11
11
  - Text-to-Speech generation and streaming
12
12
  - Speech-to-Speech voice conversion
13
13
  - Text-to-Dialogue multi-speaker generation with audio tags
14
+ - Eleven v4 support, with per-model voice-settings filtering and text caps
14
15
  - Sound effects generation from text descriptions
15
16
  - Music generation from prompts or composition plans
16
17
  - Voice management (list, get, create, update, delete)
@@ -74,7 +75,7 @@ audio.save_to_file("output.mp3")
74
75
  audio = client.tts.generate(
75
76
  "Hello world",
76
77
  voice_id: "voice_id",
77
- model_id: "eleven_v3", # Most expressive, 70+ languages, audio tags
78
+ model_id: "eleven_v4", # Most expressive; audio tags; see "Eleven v4" below
78
79
  voice_settings: {
79
80
  stability: 0.5,
80
81
  similarity_boost: 0.75
@@ -88,8 +89,72 @@ File.open("output.mp3", "wb") do |file|
88
89
  file.write(chunk)
89
90
  end
90
91
  end
92
+
93
+ # Word-level timestamps
94
+ result = client.tts.generate_with_timestamps("Hello world", voice_id: "voice_id")
95
+ result[:audio] # => ElevenRb::Objects::Audio
96
+ result[:alignment] # => { "characters" => [...], "character_start_times_seconds" => [...], ... }
97
+ result[:normalized_alignment]
98
+ result[:request_id] # from the request-id response header
99
+ result[:character_cost] # Integer, from the character-cost response header
100
+ ```
101
+
102
+ Optional keywords on `generate`, `stream` and `generate_with_timestamps` (each is left out of the request when nil):
103
+ `language_code`, `apply_text_normalization`, `seed`, `previous_text`, `next_text`, `previous_request_ids`,
104
+ `next_request_ids`, `pronunciation_dictionary_locators`, `use_pvc_as_ivc`.
105
+
106
+ ### Eleven v4
107
+
108
+ `eleven_v4` (and the cheaper, faster `eleven_v4_turbo`) take up to 10,000 characters per request.
109
+
110
+ ```ruby
111
+ audio = client.tts.generate(
112
+ "[sighs] Right... let's try that one more time.",
113
+ voice_id: "voice_id",
114
+ model_id: "eleven_v4",
115
+ voice_settings: { stability: 0.5, similarity_boost: 0.75 },
116
+ seed: 42, # reproducible re-rolls
117
+ previous_text: "That did not go to plan.", # continuity with the take before
118
+ next_text: "Here we go." # ...and the one after
119
+ )
120
+ audio.request_id # pass as previous_request_ids: / next_request_ids: on neighbouring takes
121
+ audio.character_cost
122
+ audio.dropped_settings # => [] (voice settings the gem removed, see below)
123
+ ```
124
+
125
+ - **Only `stability` and `similarity_boost` take effect.** The API accepts `speed`, `style` and
126
+ `use_speaker_boost` on v4 and silently ignores them, so the gem drops them from the request, logs a
127
+ warning through `logger`, and lists them on `audio.dropped_settings`. Set
128
+ `strict_voice_settings: true` on the client to raise `ElevenRb::Errors::ValidationError` instead.
129
+ Only those known keys are ever dropped: a voice-setting key the gem does not know passes through
130
+ unchanged on every model, so new API fields keep working.
131
+ Pace a v4 take with the text (ellipses, pause tags), not `speed`.
132
+ - **SSML `<break time="…"/>` tags are ignored** by v4 (no pause is produced). Use an audio tag or
133
+ punctuation for pauses.
134
+ - **Audio tags** go in square brackets inside the text: `[sighs]`, `[whispers]`, `[laughs]`,
135
+ `[short pause]`. Note that tags appear in the returned alignment like any other characters.
136
+ - **Continuity**: `previous_text` / `next_text` (or `previous_request_ids` / `next_request_ids`) tell the
137
+ model what surrounds a take, so stitched takes keep a consistent delivery.
138
+ - **`seed`** makes v4 generation reproducible, so a re-roll changes one thing at a time.
139
+
140
+ `ElevenRb::ModelCapabilities` answers these questions for any model id:
141
+
142
+ ```ruby
143
+ ElevenRb::ModelCapabilities.supported_voice_settings("eleven_v4") # => [:stability, :similarity_boost]
144
+ ElevenRb::ModelCapabilities.max_text_length("eleven_v3") # => 5000
145
+ ElevenRb::ModelCapabilities.supports?("eleven_v4", :audio_tags) # => true
146
+ ElevenRb::ModelCapabilities.supports?("eleven_v4", :ssml_break) # => false
91
147
  ```
92
148
 
149
+ | Model family | Voice settings honoured | Max chars | Audio tags | SSML breaks |
150
+ |---|---|---|---|---|
151
+ | `eleven_v4`, `eleven_v4_turbo` | stability, similarity_boost | 10,000 | yes | no |
152
+ | `eleven_v3` | stability, similarity_boost | 5,000 | yes | no |
153
+ | `eleven_v3_conversational` | stability, similarity_boost, use_speaker_boost | 5,000 | yes | no |
154
+ | `eleven_flash_v2_5`, `eleven_turbo_v2_5` | all five (incl. style, speed) | 40,000 | no | yes |
155
+ | `eleven_flash_v2`, `eleven_turbo_v2` | all five | 30,000 | no | yes |
156
+ | `eleven_multilingual_v2` and others | all five | 10,000 | no | yes |
157
+
93
158
  ### Speech-to-Speech
94
159
 
95
160
  ```ruby
@@ -123,30 +188,42 @@ audio = client.text_to_dialogue.generate([
123
188
  ])
124
189
  audio.save_to_file("dialogue.mp3")
125
190
 
126
- # With options
191
+ # With options (the default model is eleven_v4)
127
192
  audio = client.dialogue.generate(
128
193
  inputs,
129
- model_id: "eleven_v3",
194
+ model_id: "eleven_v4",
130
195
  language_code: "en",
131
- settings: { stability: 0.5 },
196
+ settings: { stability: 0.5, similarity: 0.75 }, # sent to the API unchanged
132
197
  seed: 42,
198
+ previous_text: "Earlier in the scene...",
133
199
  output_format: "mp3_44100_192"
134
200
  )
201
+
202
+ # With timestamps and per-speaker segments
203
+ result = client.dialogue.generate_with_timestamps(inputs)
204
+ result[:audio] # => ElevenRb::Objects::Audio
205
+ result[:alignment] # character timings
206
+ result[:voice_segments] # which voice speaks when
207
+ result[:request_id]
135
208
  ```
136
209
 
210
+ The hard text cap follows the model (10,000 characters for `eleven_v4`, 5,000 for `eleven_v3`); above
211
+ 2,000 characters the gem logs a warning, since shorter dialogue requests give more reliable results.
212
+
137
213
  ### Audio Tags
138
214
 
139
- The `eleven_v3` model supports inline audio tags for expressive speech:
215
+ The `eleven_v4` and `eleven_v3` models support inline audio tags, in square brackets, for expressive speech
216
+ (`ElevenRb::ModelCapabilities.supports?(model_id, :audio_tags)`):
140
217
 
141
218
  ```ruby
142
219
  audio = client.tts.generate(
143
220
  "[excited] Oh wow, this is AMAZING! [laughs] I can't believe it...",
144
221
  voice_id: "voice_id",
145
- model_id: "eleven_v3"
222
+ model_id: "eleven_v4"
146
223
  )
147
224
  ```
148
225
 
149
- Supported tags include `[laughs]`, `[whispers]`, `[sighs]`, `[excited]`, `[sarcastic]`, `[curious]`, `[pause]`, and more. Use CAPS for emphasis, `...` for pauses, and `—` for interruptions. See the [ElevenLabs v3 documentation](https://elevenlabs.io/docs/guides/audio-tags) for the full list.
226
+ Supported tags include `[laughs]`, `[whispers]`, `[sighs]`, `[excited]`, `[sarcastic]`, `[curious]`, `[short pause]`, and more. Use CAPS for emphasis, `...` for pauses, and `—` for interruptions. See the [ElevenLabs audio tags documentation](https://elevenlabs.io/docs/guides/audio-tags) for the full list.
150
227
 
151
228
  ### Sound Effects
152
229
 
@@ -289,7 +366,7 @@ client = ElevenRb::Client.new(
289
366
  Sentry.capture_exception(error, extra: { path: path })
290
367
  },
291
368
 
292
- # Cost tracking
369
+ # Cost tracking (add request_id: to also receive the API's request id)
293
370
  on_audio_generated: ->(audio:, voice_id:, text:, cost_info:) {
294
371
  UsageRecord.create!(
295
372
  characters: cost_info[:character_count],
@@ -311,8 +388,11 @@ client = ElevenRb::Client.new(
311
388
  models = client.models.list
312
389
  models.each { |m| puts "#{m.name} (#{m.model_id})" }
313
390
 
314
- # Get the latest/most capable model
315
- client.models.latest # => "eleven_v3"
391
+ # Find one model
392
+ client.models.find("eleven_v4") # => ElevenRb::Objects::Model (get is an alias)
393
+
394
+ # Get the latest/most capable model: eleven_v4, else eleven_v3, else the default
395
+ client.models.latest.model_id # => "eleven_v4"
316
396
 
317
397
  # Get multilingual models
318
398
  client.models.multilingual
@@ -343,7 +423,8 @@ client = ElevenRb::Client.new(
343
423
  open_timeout: 10, # Connection timeout
344
424
  max_retries: 3, # Max retry attempts
345
425
  retry_delay: 1.0, # Base delay between retries
346
- logger: Rails.logger # Optional logger
426
+ logger: Rails.logger, # Optional logger (receives dropped-setting warnings)
427
+ strict_voice_settings: false # true: raise instead of dropping settings a model ignores
347
428
  )
348
429
  ```
349
430
 
@@ -28,6 +28,10 @@ module ElevenRb
28
28
 
29
29
  # Trigger a callback if it's configured
30
30
  #
31
+ # A callback that names its keywords explicitly (no `**rest`) receives only
32
+ # the keywords it declares, so callbacks written before a keyword was added
33
+ # (e.g. `request_id:` on on_audio_generated in 1.1.0) keep working.
34
+ #
31
35
  # @param callback_name [Symbol] the name of the callback
32
36
  # @param kwargs [Hash] keyword arguments to pass to the callback
33
37
  # @return [Object, nil] the return value of the callback, or nil
@@ -36,12 +40,32 @@ module ElevenRb
36
40
  return unless callback.respond_to?(:call)
37
41
 
38
42
  begin
39
- callback.call(**kwargs)
43
+ callback.call(**accepted_callback_kwargs(callback, kwargs))
40
44
  rescue StandardError => e
41
45
  # Don't let callback errors break the main flow
42
46
  warn "[ElevenRb] Callback error in #{callback_name}: #{e.message}"
43
47
  nil
44
48
  end
45
49
  end
50
+
51
+ private
52
+
53
+ def accepted_callback_kwargs(callback, kwargs)
54
+ params = callback_parameters(callback)
55
+ return kwargs if params.nil? || params.any? { |type, _| type == :keyrest }
56
+
57
+ accepted = params.filter_map { |type, name| name if %i[key keyreq].include?(type) }
58
+ return kwargs if accepted.empty?
59
+
60
+ kwargs.slice(*accepted)
61
+ end
62
+
63
+ def callback_parameters(callback)
64
+ return callback.parameters if callback.respond_to?(:parameters)
65
+
66
+ callback.method(:call).parameters
67
+ rescue NameError
68
+ nil
69
+ end
46
70
  end
47
71
  end
@@ -21,12 +21,13 @@ module ElevenRb
21
21
  open_timeout: 10,
22
22
  max_retries: 3,
23
23
  retry_delay: 1.0,
24
- retry_statuses: [429, 500, 502, 503, 504].freeze
24
+ retry_statuses: [429, 500, 502, 503, 504].freeze,
25
+ strict_voice_settings: false
25
26
  }.freeze
26
27
 
27
28
  attr_accessor :api_key, :base_url, :timeout, :open_timeout,
28
29
  :max_retries, :retry_delay, :retry_statuses,
29
- :logger
30
+ :logger, :strict_voice_settings
30
31
 
31
32
  # Initialize a new configuration
32
33
  #
@@ -39,6 +40,8 @@ module ElevenRb
39
40
  # @option options [Float] :retry_delay Base delay between retries in seconds (default: 1.0)
40
41
  # @option options [Array<Integer>] :retry_statuses HTTP status codes to retry (default: [429, 500, 502, 503, 504])
41
42
  # @option options [Logger] :logger Logger instance for debug output
43
+ # @option options [Boolean] :strict_voice_settings Raise instead of warn when a voice setting the
44
+ # model does not honour is passed (default: false)
42
45
  # @option options [Proc] :on_request Callback before each request
43
46
  # @option options [Proc] :on_response Callback after successful response
44
47
  # @option options [Proc] :on_error Callback when an error occurs
@@ -87,7 +90,8 @@ module ElevenRb
87
90
  timeout: timeout,
88
91
  open_timeout: open_timeout,
89
92
  max_retries: max_retries,
90
- retry_delay: retry_delay
93
+ retry_delay: retry_delay,
94
+ strict_voice_settings: strict_voice_settings
91
95
  }
92
96
  end
93
97
  end
@@ -32,9 +32,11 @@ module ElevenRb
32
32
  # @param path [String] the API path
33
33
  # @param body [Hash] request body
34
34
  # @param response_type [Symbol] :json or :binary
35
- # @return [Hash, Array, String] parsed response
36
- def post(path, body = {}, response_type: :json)
37
- request(:post, path, body: body, response_type: response_type)
35
+ # @param with_meta [Boolean] when true, return `{ body:, headers: }` instead of the body alone
36
+ # (headers as a Hash of lower-cased String keys to String values)
37
+ # @return [Hash, Array, String] parsed response (or `{ body:, headers: }` with with_meta)
38
+ def post(path, body = {}, response_type: :json, with_meta: false)
39
+ request(:post, path, body: body, response_type: response_type, with_meta: with_meta)
38
40
  end
39
41
 
40
42
  # Make a DELETE request
@@ -68,7 +70,7 @@ module ElevenRb
68
70
  private
69
71
 
70
72
  def request(method, path, body: nil, params: nil, response_type: :json, multipart: false, stream: false,
71
- attempt: 1, &block)
73
+ attempt: 1, with_meta: false, &block)
72
74
  config.validate!
73
75
  url = "#{config.base_url}#{path}"
74
76
  start_time = Time.now
@@ -86,15 +88,12 @@ module ElevenRb
86
88
  # Trigger after response callback
87
89
  config.trigger(:on_response, method: method, path: path, response: response, duration: duration)
88
90
 
89
- # Return binary data directly
90
- return response.body if response_type == :binary && response.success?
91
-
92
- handle_response(response)
91
+ build_result(response, response_type, with_meta)
93
92
  rescue Errors::RateLimitError => e
94
93
  config.trigger(:on_rate_limit, retry_after: e.retry_after, error: e)
95
- handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
94
+ handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
96
95
  rescue Errors::ServerError => e
97
- handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, &block)
96
+ handle_retry(e, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
98
97
  rescue Errors::Base => e
99
98
  config.trigger(:on_error, error: e, method: method, path: path,
100
99
  context: { body: sanitize_body_for_logging(body) })
@@ -107,6 +106,17 @@ module ElevenRb
107
106
  end
108
107
  end
109
108
 
109
+ def build_result(response, response_type, with_meta)
110
+ # Return binary data directly
111
+ result = if response_type == :binary && response.success?
112
+ response.body
113
+ else
114
+ handle_response(response)
115
+ end
116
+
117
+ with_meta ? { body: result, headers: normalize_headers(response) } : result
118
+ end
119
+
110
120
  def execute_request(method, url, body, params, multipart, stream, &block)
111
121
  options = build_options(body, params, multipart, stream, &block)
112
122
 
@@ -210,7 +220,7 @@ module ElevenRb
210
220
  raise error_class.new(message, **error_kwargs)
211
221
  end
212
222
 
213
- def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, &block)
223
+ def handle_retry(error, method, path, body, params, response_type, multipart, stream, attempt, with_meta, &block)
214
224
  raise error if attempt > config.max_retries || !config.retry_statuses.include?(error.http_status)
215
225
 
216
226
  delay = if error.is_a?(Errors::RateLimitError) && error.retry_after
@@ -224,7 +234,17 @@ module ElevenRb
224
234
  sleep(delay)
225
235
 
226
236
  request(method, path,
227
- body: body, params: params, response_type: response_type, multipart: multipart, stream: stream, attempt: attempt + 1, &block)
237
+ body: body, params: params, response_type: response_type, multipart: multipart, stream: stream,
238
+ attempt: attempt + 1, with_meta: with_meta, &block)
239
+ end
240
+
241
+ # Response headers as a plain Hash of lower-cased String keys to String values
242
+ # (HTTParty exposes multi-value arrays; the first value is kept)
243
+ def normalize_headers(response)
244
+ raw = response.headers.respond_to?(:to_hash) ? response.headers.to_hash : response.headers.to_h
245
+ raw.each_with_object({}) do |(key, value), out|
246
+ out[key.to_s.downcase] = value.is_a?(Array) ? value.first : value
247
+ end
228
248
  end
229
249
 
230
250
  def wrap_error(error)
@@ -0,0 +1,103 @@
1
+ # frozen_string_literal: true
2
+
3
+ module ElevenRb
4
+ # What each ElevenLabs model family accepts, keyed by model-id family.
5
+ #
6
+ # The API silently ignores voice settings a model does not honour (Eleven v4
7
+ # accepts `speed`, `style` and `use_speaker_boost` and does nothing with them),
8
+ # so the gem uses this table to send only the settings that take effect, to
9
+ # cap text length per model, and to answer feature questions (audio tags,
10
+ # SSML `<break>` tags, request continuity via previous_text / next_text).
11
+ #
12
+ # @example
13
+ # ElevenRb::ModelCapabilities.supported_voice_settings('eleven_v4')
14
+ # # => [:stability, :similarity_boost]
15
+ # ElevenRb::ModelCapabilities.max_text_length('eleven_flash_v2_5') # => 40_000
16
+ # ElevenRb::ModelCapabilities.supports?('eleven_v4', :audio_tags) # => true
17
+ module ModelCapabilities
18
+ # Capability record for one model family
19
+ Capabilities = Struct.new(:voice_settings, :max_text_length, :ssml_break, :audio_tags, :continuity,
20
+ keyword_init: true) do
21
+ def ssml_break? = ssml_break
22
+ def audio_tags? = audio_tags
23
+ def continuity? = continuity
24
+ end
25
+
26
+ ALL_VOICE_SETTINGS = %i[stability similarity_boost style use_speaker_boost speed].freeze
27
+
28
+ def self.build(voice_settings:, max_text_length:, ssml_break:, audio_tags:, continuity:)
29
+ Capabilities.new(
30
+ voice_settings: voice_settings.freeze,
31
+ max_text_length: max_text_length,
32
+ ssml_break: ssml_break,
33
+ audio_tags: audio_tags,
34
+ continuity: continuity
35
+ ).freeze
36
+ end
37
+ private_class_method :build
38
+
39
+ V4 = build(voice_settings: %i[stability similarity_boost], max_text_length: 10_000,
40
+ ssml_break: false, audio_tags: true, continuity: true)
41
+ V3_CONVERSATIONAL = build(voice_settings: %i[stability similarity_boost use_speaker_boost], max_text_length: 5_000,
42
+ ssml_break: false, audio_tags: true, continuity: true)
43
+ V3 = build(voice_settings: %i[stability similarity_boost], max_text_length: 5_000,
44
+ ssml_break: false, audio_tags: true, continuity: true)
45
+ V2_5_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 40_000,
46
+ ssml_break: true, audio_tags: false, continuity: true)
47
+ V2_FAST = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 30_000,
48
+ ssml_break: true, audio_tags: false, continuity: true)
49
+ DEFAULT = build(voice_settings: ALL_VOICE_SETTINGS, max_text_length: 10_000,
50
+ ssml_break: true, audio_tags: false, continuity: true)
51
+
52
+ # Ordered [matcher, capabilities] pairs; the first match wins.
53
+ FAMILIES = [
54
+ [/\Aeleven_v4/, V4],
55
+ [/\Aeleven_v3_conversational/, V3_CONVERSATIONAL],
56
+ [/\Aeleven_v3/, V3],
57
+ [/\Aeleven_(flash|turbo)_v2_5/, V2_5_FAST],
58
+ [/\Aeleven_(flash|turbo)_v2/, V2_FAST]
59
+ ].freeze
60
+
61
+ FEATURES = %i[audio_tags ssml_break continuity].freeze
62
+
63
+ module_function
64
+
65
+ # Capabilities for a model id (unknown ids get the permissive default)
66
+ #
67
+ # @param model_id [String, Symbol, nil]
68
+ # @return [Capabilities] frozen
69
+ def for(model_id)
70
+ id = model_id.to_s
71
+ FAMILIES.each { |matcher, caps| return caps if matcher.match?(id) }
72
+ DEFAULT
73
+ end
74
+
75
+ # Voice settings keys the model honours
76
+ #
77
+ # @param model_id [String]
78
+ # @return [Array<Symbol>]
79
+ def supported_voice_settings(model_id)
80
+ self.for(model_id).voice_settings
81
+ end
82
+
83
+ # Maximum characters per request for the model
84
+ #
85
+ # @param model_id [String]
86
+ # @return [Integer]
87
+ def max_text_length(model_id)
88
+ self.for(model_id).max_text_length
89
+ end
90
+
91
+ # Whether the model supports a feature
92
+ #
93
+ # @param model_id [String]
94
+ # @param feature [Symbol] :audio_tags, :ssml_break or :continuity
95
+ # @return [Boolean]
96
+ def supports?(model_id, feature)
97
+ feature = feature.to_sym
98
+ raise ArgumentError, "Unknown feature #{feature.inspect} (expected one of #{FEATURES.join(', ')})" unless FEATURES.include?(feature)
99
+
100
+ self.for(model_id).public_send(feature) ? true : false
101
+ end
102
+ end
103
+ end
@@ -6,7 +6,8 @@ module ElevenRb
6
6
  module Objects
7
7
  # Represents generated audio data
8
8
  class Audio
9
- attr_reader :data, :format, :voice_id, :text, :model_id
9
+ attr_reader :data, :format, :voice_id, :text, :model_id,
10
+ :request_id, :character_cost, :dropped_settings
10
11
 
11
12
  # Initialize audio object
12
13
  #
@@ -15,12 +16,20 @@ module ElevenRb
15
16
  # @param voice_id [String] the voice ID used
16
17
  # @param text [String] the text that was converted
17
18
  # @param model_id [String, nil] the model ID used
18
- def initialize(data:, format:, voice_id:, text:, model_id: nil)
19
+ # @param request_id [String, nil] the API's `request-id` response header (usable as
20
+ # previous_request_ids / next_request_ids on a later request)
21
+ # @param character_cost [Integer, nil] the API's `character-cost` response header
22
+ # @param dropped_settings [Array<Symbol>] voice settings the gem dropped because the model ignores them
23
+ def initialize(data:, format:, voice_id:, text:, model_id: nil, request_id: nil, character_cost: nil,
24
+ dropped_settings: [])
19
25
  @data = data
20
26
  @format = format
21
27
  @voice_id = voice_id
22
28
  @text = text
23
29
  @model_id = model_id
30
+ @request_id = request_id
31
+ @character_cost = character_cost
32
+ @dropped_settings = Array(dropped_settings).dup.freeze
24
33
  end
25
34
 
26
35
  # Save audio to a file
@@ -62,7 +71,7 @@ module ElevenRb
62
71
  'audio/mpeg'
63
72
  when /pcm/
64
73
  'audio/pcm'
65
- when /ogg/
74
+ when /ogg|opus/
66
75
  'audio/ogg'
67
76
  when /wav/
68
77
  'audio/wav'
@@ -82,7 +91,7 @@ module ElevenRb
82
91
  'mp3'
83
92
  when /pcm/
84
93
  'pcm'
85
- when /ogg/
94
+ when /ogg|opus/
86
95
  'ogg'
87
96
  when /wav/
88
97
  'wav'
@@ -13,6 +13,9 @@ module ElevenRb
13
13
  'eleven_multilingual_v1' => 0.30,
14
14
  'eleven_multilingual_v2' => 0.30,
15
15
  'eleven_v3' => 0.30,
16
+ 'eleven_v3_conversational' => 0.15,
17
+ 'eleven_v4' => 0.30,
18
+ 'eleven_v4_turbo' => 0.15,
16
19
  'eleven_turbo_v2' => 0.18,
17
20
  'eleven_turbo_v2_5' => 0.18,
18
21
  'eleven_english_sts_v2' => 0.30,
@@ -18,6 +18,16 @@ module ElevenRb
18
18
  attribute :max_characters_request_free_user
19
19
  attribute :max_characters_request_subscribed_user
20
20
  attribute :concurrency_group
21
+ attribute :model_rates
22
+ attribute :maximum_text_length_per_request
23
+ attribute :requires_alpha_access, type: :boolean
24
+
25
+ # Voice settings keys this model honours (from ModelCapabilities)
26
+ #
27
+ # @return [Array<Symbol>]
28
+ def supported_voice_settings
29
+ ModelCapabilities.supported_voice_settings(model_id)
30
+ end
21
31
 
22
32
  # Check if this model supports a given language
23
33
  #
@@ -25,6 +25,39 @@ module ElevenRb
25
25
  from_response(DEFAULTS.merge(overrides))
26
26
  end
27
27
 
28
+ # Voice-setting keys the capability table knows about. Only these can be
29
+ # dropped for a model; any other key is passed through untouched so a
30
+ # future API field is never swallowed.
31
+ KNOWN_KEYS = ModelCapabilities::ALL_VOICE_SETTINGS
32
+
33
+ # Build the voice_settings hash a model actually honours
34
+ #
35
+ # Starts from DEFAULTS filtered to the model's supported keys and merges the
36
+ # overrides (keys symbolized). Known keys the model does not support are
37
+ # dropped; unknown keys pass through after the supported ones, in the
38
+ # caller's order; nil values are removed. Only non-nil override keys count
39
+ # as dropped: defaults the model does not take are filtered silently.
40
+ #
41
+ # @example
42
+ # VoiceSettings.for_model('eleven_multilingual_v2')
43
+ # # => [{ stability: 0.5, similarity_boost: 0.75, style: 0.0, use_speaker_boost: true }, []]
44
+ # VoiceSettings.for_model('eleven_v4', speed: 1.1)
45
+ # # => [{ stability: 0.5, similarity_boost: 0.75 }, [:speed]]
46
+ #
47
+ # @param model_id [String] the model ID
48
+ # @param overrides [Hash] caller settings (String or Symbol keys)
49
+ # @return [Array(Hash, Array<Symbol>)] the settings hash and the dropped override keys
50
+ def self.for_model(model_id, overrides = {})
51
+ supported = ModelCapabilities.supported_voice_settings(model_id)
52
+ requested = (overrides || {}).to_h.transform_keys(&:to_sym)
53
+
54
+ dropped = requested.compact.keys & (KNOWN_KEYS - supported)
55
+ unknown = requested.except(*KNOWN_KEYS)
56
+ settings = DEFAULTS.slice(*supported).merge(requested.slice(*supported)).merge(unknown).compact
57
+
58
+ [settings, dropped]
59
+ end
60
+
28
61
  # Convert to hash suitable for API request
29
62
  #
30
63
  # @return [Hash]
@@ -51,6 +51,24 @@ module ElevenRb
51
51
  http_client.post(path, body, response_type: :binary)
52
52
  end
53
53
 
54
+ # Make a JSON POST request and also return the response headers
55
+ #
56
+ # @param path [String]
57
+ # @param body [Hash]
58
+ # @return [Hash] `{ body: Hash, headers: Hash<String, String> }`
59
+ def post_with_meta(path, body = {})
60
+ http_client.post(path, body, response_type: :json, with_meta: true)
61
+ end
62
+
63
+ # Make a binary POST request and also return the response headers
64
+ #
65
+ # @param path [String]
66
+ # @param body [Hash]
67
+ # @return [Hash] `{ body: String, headers: Hash<String, String> }`
68
+ def post_binary_with_meta(path, body = {})
69
+ http_client.post(path, body, response_type: :binary, with_meta: true)
70
+ end
71
+
54
72
  # Make a streaming POST request
55
73
  #
56
74
  # @param path [String]
@@ -9,23 +9,36 @@ module ElevenRb
9
9
  #
10
10
  # @example Find multilingual models
11
11
  # client.models.multilingual
12
+ #
13
+ # @example Find one model
14
+ # client.models.find('eleven_v4')
12
15
  class Models < Base
13
16
  # List all available models
14
17
  #
15
18
  # @return [Array<Objects::Model>]
16
19
  def list
17
- response = get('/models')
20
+ # Call the HTTP client directly: #get below is the model lookup (kept for
21
+ # compatibility) and shadows Base#get, which made this method recurse.
22
+ response = http_client.get('/models')
18
23
  response.map { |m| Objects::Model.from_response(m) }
19
24
  end
20
25
 
21
- # Get a specific model by ID
26
+ # Find a specific model by ID
22
27
  #
23
28
  # @param model_id [String] the model ID
24
29
  # @return [Objects::Model, nil]
25
- def get(model_id)
30
+ def find(model_id)
26
31
  list.find { |m| m.model_id == model_id }
27
32
  end
28
33
 
34
+ # Alias of {#find}, kept for backwards compatibility
35
+ #
36
+ # @param model_id [String] the model ID
37
+ # @return [Objects::Model, nil]
38
+ def get(model_id)
39
+ find(model_id)
40
+ end
41
+
29
42
  # Get all multilingual models
30
43
  #
31
44
  # @return [Array<Objects::Model>]
@@ -51,14 +64,20 @@ module ElevenRb
51
64
  #
52
65
  # @return [Objects::Model, nil]
53
66
  def default
54
- get('eleven_multilingual_v2') || tts_capable.first
67
+ default_from(list)
55
68
  end
56
69
 
57
- # Get the latest/most capable model
70
+ # Get the latest/most capable model available to the account:
71
+ # eleven_v4, else eleven_v3, else {#default} (one /models request)
58
72
  #
59
73
  # @return [Objects::Model, nil]
60
74
  def latest
61
- get('eleven_v3') || default
75
+ models = list
76
+ %w[eleven_v4 eleven_v3].each do |model_id|
77
+ model = models.find { |m| m.model_id == model_id }
78
+ return model if model
79
+ end
80
+ default_from(models)
62
81
  end
63
82
 
64
83
  # Get model IDs as array
@@ -67,6 +86,13 @@ module ElevenRb
67
86
  def ids
68
87
  list.map(&:model_id)
69
88
  end
89
+
90
+ private
91
+
92
+ # eleven_multilingual_v2, else the first TTS-capable model, from an already-fetched list
93
+ def default_from(models)
94
+ models.find { |m| m.model_id == 'eleven_multilingual_v2' } || models.find(&:can_do_text_to_speech)
95
+ end
70
96
  end
71
97
  end
72
98
  end
@@ -10,20 +10,51 @@ module ElevenRb
10
10
  # { text: "[laughs] Thanks!", voice_id: "voice_xyz" }
11
11
  # ])
12
12
  # audio.save_to_file("dialogue.mp3")
13
+ #
14
+ # @example Dialogue with timestamps and per-speaker segments
15
+ # result = client.dialogue.generate_with_timestamps(inputs, seed: 7)
16
+ # result[:voice_segments] # => [{ "voice_id" => ..., "start_time_seconds" => ... }, ...]
13
17
  class TextToDialogue < Base
14
- DEFAULT_MODEL = 'eleven_v3'
18
+ DEFAULT_MODEL = 'eleven_v4'
15
19
  MAX_VOICES_PER_REQUEST = 10
20
+
21
+ # Kept for compatibility. The hard cap now comes from
22
+ # ModelCapabilities.max_text_length(model_id) (5,000 for eleven_v3).
16
23
  MAX_TEXT_LENGTH = 5000
17
24
 
25
+ # Above this many characters the API recommends splitting the dialogue;
26
+ # the gem logs a warning but still sends the request.
27
+ RECOMMENDED_MAX_TEXT_LENGTH = 2_000
28
+
29
+ # Optional request-body keys, in the order they are written to the body.
30
+ # Each is omitted from the body when nil.
31
+ OPTIONAL_BODY_KEYS = %i[
32
+ language_code
33
+ settings
34
+ seed
35
+ use_pvc_as_ivc
36
+ previous_text
37
+ future_text
38
+ previous_request_ids
39
+ next_request_ids
40
+ pronunciation_dictionary_locators
41
+ ].freeze
42
+
18
43
  # Generate dialogue audio from multiple speaker inputs
19
44
  #
20
45
  # @param inputs [Array<Hash>] Array of { text:, voice_id: } hashes
21
- # @param model_id [String] Model to use (only eleven_v3 supported)
46
+ # @param model_id [String] Model to use (default: eleven_v4)
22
47
  # @param language_code [String, nil] ISO 639-1 language code
23
- # @param settings [Hash, nil] Generation settings (stability: 0.0-1.0)
48
+ # @param settings [Hash, nil] Generation settings, sent unchanged (e.g. stability, similarity)
24
49
  # @param seed [Integer, nil] Seed for reproducibility
25
50
  # @param output_format [String] Audio output format
26
51
  # @param apply_text_normalization [String] "auto", "on", or "off"
52
+ # @param use_pvc_as_ivc [Boolean, nil] use the IVC version of professional voices
53
+ # @param previous_text [String, nil] text that comes before this dialogue (continuity)
54
+ # @param future_text [String, nil] text that comes after this dialogue (continuity)
55
+ # @param previous_request_ids [Array<String>, nil] request IDs of preceding generations
56
+ # @param next_request_ids [Array<String>, nil] request IDs of following generations
57
+ # @param pronunciation_dictionary_locators [Array<Hash>, nil] pronunciation dictionaries to apply
27
58
  # @return [Objects::Audio]
28
59
  def generate(
29
60
  inputs,
@@ -32,45 +63,92 @@ module ElevenRb
32
63
  settings: nil,
33
64
  seed: nil,
34
65
  output_format: 'mp3_44100_128',
35
- apply_text_normalization: 'auto'
66
+ apply_text_normalization: 'auto',
67
+ use_pvc_as_ivc: nil,
68
+ previous_text: nil,
69
+ future_text: nil,
70
+ previous_request_ids: nil,
71
+ next_request_ids: nil,
72
+ pronunciation_dictionary_locators: nil
36
73
  )
37
- validate_inputs!(inputs)
74
+ validate_inputs!(inputs, model_id)
38
75
 
39
- body = build_request_body(inputs, model_id, language_code, settings, seed,
40
- apply_text_normalization)
76
+ body = build_request_body(inputs, model_id, apply_text_normalization, optional_values(binding))
77
+ response = post_binary_with_meta("/text-to-dialogue?output_format=#{output_format}", body)
41
78
 
42
- response = post_binary(
43
- "/text-to-dialogue?output_format=#{output_format}",
44
- body
45
- )
79
+ build_audio_response(response[:body], inputs, output_format, model_id,
80
+ request_id: response[:headers]['request-id'])
81
+ end
46
82
 
47
- build_audio_response(response, inputs, output_format, model_id)
83
+ # Generate dialogue audio with character timestamps and per-voice segments
84
+ #
85
+ # Takes the same keywords as {#generate}.
86
+ #
87
+ # @param inputs [Array<Hash>] Array of { text:, voice_id: } hashes
88
+ # @return [Hash] `{ audio:, alignment:, normalized_alignment:, voice_segments:, request_id: }`
89
+ def generate_with_timestamps(
90
+ inputs,
91
+ model_id: DEFAULT_MODEL,
92
+ language_code: nil,
93
+ settings: nil,
94
+ seed: nil,
95
+ output_format: 'mp3_44100_128',
96
+ apply_text_normalization: 'auto',
97
+ use_pvc_as_ivc: nil,
98
+ previous_text: nil,
99
+ future_text: nil,
100
+ previous_request_ids: nil,
101
+ next_request_ids: nil,
102
+ pronunciation_dictionary_locators: nil
103
+ )
104
+ validate_inputs!(inputs, model_id)
105
+
106
+ body = build_request_body(inputs, model_id, apply_text_normalization, optional_values(binding))
107
+ result = post_with_meta("/text-to-dialogue/with-timestamps?output_format=#{output_format}", body)
108
+ response = result[:body]
109
+ request_id = result[:headers]['request-id']
110
+
111
+ audio_data = Base64.decode64(response['audio_base64']) if response['audio_base64']
112
+ audio = (build_audio_response(audio_data, inputs, output_format, model_id, request_id: request_id) if audio_data)
113
+
114
+ {
115
+ audio: audio,
116
+ alignment: response['alignment'],
117
+ normalized_alignment: response['normalized_alignment'],
118
+ voice_segments: response['voice_segments'],
119
+ request_id: request_id
120
+ }
48
121
  end
49
122
 
50
123
  private
51
124
 
52
- def build_request_body(inputs, model_id, language_code, settings, seed,
53
- apply_text_normalization)
125
+ # The optional keyword values of the calling method, keyed by OPTIONAL_BODY_KEYS
126
+ def optional_values(caller_binding)
127
+ OPTIONAL_BODY_KEYS.to_h { |key| [key, caller_binding.local_variable_get(key)] }
128
+ end
129
+
130
+ def build_request_body(inputs, model_id, apply_text_normalization, options)
54
131
  body = {
55
132
  inputs: inputs.map { |i| { text: i[:text], voice_id: i[:voice_id] } },
56
133
  model_id: model_id,
57
134
  apply_text_normalization: apply_text_normalization
58
135
  }
59
136
 
60
- body[:language_code] = language_code if language_code
61
- body[:settings] = settings if settings
62
- body[:seed] = seed if seed
137
+ OPTIONAL_BODY_KEYS.each do |key|
138
+ body[key] = options[key] unless options[key].nil?
139
+ end
63
140
  body
64
141
  end
65
142
 
66
- def build_audio_response(response, inputs, output_format, model_id)
143
+ def build_audio_response(data, inputs, output_format, model_id, request_id: nil)
67
144
  total_text = inputs.map { |i| i[:text] }.join("\n")
68
145
  total_chars = inputs.sum { |i| i[:text].length }
69
146
  primary_voice = inputs.first[:voice_id]
70
147
 
71
148
  audio = Objects::Audio.new(
72
- data: response, format: output_format,
73
- voice_id: primary_voice, text: total_text, model_id: model_id
149
+ data: data, format: output_format,
150
+ voice_id: primary_voice, text: total_text, model_id: model_id,
151
+ request_id: request_id
74
152
  )
75
153
 
76
154
  cost_info = Objects::CostInfo.new(
@@ -80,13 +158,14 @@ module ElevenRb
80
158
  http_client.config.trigger(
81
159
  :on_audio_generated,
82
160
  audio: audio, voice_id: primary_voice,
83
- text: total_text, cost_info: cost_info.to_h
161
+ text: total_text, cost_info: cost_info.to_h,
162
+ request_id: request_id
84
163
  )
85
164
 
86
165
  audio
87
166
  end
88
167
 
89
- def validate_inputs!(inputs)
168
+ def validate_inputs!(inputs, model_id = DEFAULT_MODEL)
90
169
  raise Errors::ValidationError, 'inputs must be a non-empty array' unless inputs.is_a?(Array) && !inputs.empty?
91
170
 
92
171
  inputs.each_with_index do |input, i|
@@ -101,12 +180,23 @@ module ElevenRb
101
180
  "(got #{unique_voices.length})"
102
181
  end
103
182
 
104
- total_chars = inputs.sum { |i| i[:text].length }
105
- return unless total_chars > MAX_TEXT_LENGTH
183
+ validate_text_length!(inputs.sum { |i| i[:text].length }, model_id)
184
+ end
185
+
186
+ def validate_text_length!(total_chars, model_id)
187
+ max_length = ModelCapabilities.max_text_length(model_id)
188
+ if total_chars > max_length
189
+ raise Errors::ValidationError,
190
+ "Total text length #{total_chars} exceeds maximum " \
191
+ "#{max_length} characters for #{model_id}"
192
+ end
193
+
194
+ return unless total_chars > RECOMMENDED_MAX_TEXT_LENGTH
106
195
 
107
- raise Errors::ValidationError,
108
- "Total text length #{total_chars} exceeds maximum " \
109
- "#{MAX_TEXT_LENGTH} characters"
196
+ http_client.config.logger&.warn(
197
+ "[ElevenRb] text-to-dialogue text is #{total_chars} characters; " \
198
+ "#{RECOMMENDED_MAX_TEXT_LENGTH} or fewer per request is recommended"
199
+ )
110
200
  end
111
201
  end
112
202
  end
@@ -12,18 +12,67 @@ module ElevenRb
12
12
  # client.tts.stream("Hello world", voice_id: "voice_id") do |chunk|
13
13
  # io.write(chunk)
14
14
  # end
15
+ #
16
+ # @example Eleven v4 with continuity and a fixed seed
17
+ # audio = client.tts.generate(
18
+ # "[sighs] Right. Let's try that again.",
19
+ # voice_id: "voice_id",
20
+ # model_id: "eleven_v4",
21
+ # seed: 42,
22
+ # previous_text: "That did not go to plan."
23
+ # )
24
+ # audio.request_id # => "abc123" (from the request-id response header)
15
25
  class TextToSpeech < Base
16
26
  DEFAULT_MODEL = 'eleven_multilingual_v2'
27
+
28
+ # Kept for compatibility. The per-request cap now comes from
29
+ # ModelCapabilities.max_text_length(model_id) (5,000 for eleven_v3).
17
30
  MAX_TEXT_LENGTH = 5000
18
31
 
32
+ # Output formats the API accepts (documentation only; not validated)
19
33
  OUTPUT_FORMATS = %w[
34
+ mp3_22050_32
35
+ mp3_24000_48
36
+ mp3_44100_32
37
+ mp3_44100_64
38
+ mp3_44100_96
20
39
  mp3_44100_128
21
40
  mp3_44100_192
41
+ opus_48000_32
42
+ opus_48000_64
43
+ opus_48000_96
44
+ opus_48000_128
45
+ opus_48000_192
46
+ pcm_8000
22
47
  pcm_16000
23
48
  pcm_22050
24
49
  pcm_24000
50
+ pcm_32000
25
51
  pcm_44100
52
+ pcm_48000
53
+ wav_8000
54
+ wav_16000
55
+ wav_22050
56
+ wav_24000
57
+ wav_32000
58
+ wav_44100
59
+ wav_48000
26
60
  ulaw_8000
61
+ alaw_8000
62
+ ].freeze
63
+
64
+ # Optional request-body keys, in the order they are written to the body.
65
+ # Each is omitted from the body when nil.
66
+ OPTIONAL_BODY_KEYS = %i[
67
+ language_code
68
+ apply_text_normalization
69
+ seed
70
+ previous_text
71
+ next_text
72
+ previous_request_ids
73
+ next_request_ids
74
+ pronunciation_dictionary_locators
75
+ use_pvc_as_ivc
27
76
  ].freeze
28
77
 
29
78
  # Generate audio from text
@@ -31,30 +80,39 @@ module ElevenRb
31
80
  # @param text [String] the text to convert
32
81
  # @param voice_id [String] the voice ID to use
33
82
  # @param model_id [String] the model to use (default: eleven_multilingual_v2)
34
- # @param voice_settings [Hash] voice settings overrides
83
+ # @param voice_settings [Hash] voice settings overrides (keys the model ignores are dropped)
35
84
  # @param output_format [String] audio output format
36
- # @return [Objects::Audio]
37
- def generate(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {}, output_format: 'mp3_44100_128')
38
- validate_text!(text)
85
+ # @param language_code [String, nil] ISO 639-1 language code to enforce
86
+ # @param apply_text_normalization [String, nil] "auto", "on" or "off"
87
+ # @param seed [Integer, nil] seed for reproducible generation
88
+ # @param previous_text [String, nil] text that comes before this request (continuity)
89
+ # @param next_text [String, nil] text that comes after this request (continuity)
90
+ # @param previous_request_ids [Array<String>, nil] request IDs of preceding generations
91
+ # @param next_request_ids [Array<String>, nil] request IDs of following generations
92
+ # @param pronunciation_dictionary_locators [Array<Hash>, nil] `{ pronunciation_dictionary_id:, version_id: }`
93
+ # @param use_pvc_as_ivc [Boolean, nil] use the IVC version of a professional voice
94
+ # @return [Objects::Audio] with request_id, character_cost and dropped_settings
95
+ def generate(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {}, output_format: 'mp3_44100_128',
96
+ language_code: nil, apply_text_normalization: nil, seed: nil, previous_text: nil,
97
+ next_text: nil, previous_request_ids: nil, next_request_ids: nil,
98
+ pronunciation_dictionary_locators: nil, use_pvc_as_ivc: nil)
99
+ validate_text!(text, model_id)
39
100
  validate_presence!(voice_id, 'voice_id')
40
-
41
- settings = Objects::VoiceSettings::DEFAULTS.merge(voice_settings)
42
-
43
- body = {
44
- text: text,
45
- model_id: model_id,
46
- voice_settings: settings
47
- }
101
+ body, dropped = build_body(text, model_id, voice_settings, optional_values(binding))
48
102
 
49
103
  path = "/text-to-speech/#{voice_id}?output_format=#{output_format}"
50
- response = post_binary(path, body)
104
+ response = post_binary_with_meta(path, body)
105
+ headers = response[:headers]
51
106
 
52
107
  audio = Objects::Audio.new(
53
- data: response,
108
+ data: response[:body],
54
109
  format: output_format,
55
110
  voice_id: voice_id,
56
111
  text: text,
57
- model_id: model_id
112
+ model_id: model_id,
113
+ request_id: headers['request-id'],
114
+ character_cost: integer_header(headers, 'character-cost'),
115
+ dropped_settings: dropped
58
116
  )
59
117
 
60
118
  # Trigger cost tracking callback
@@ -64,7 +122,8 @@ module ElevenRb
64
122
  audio: audio,
65
123
  voice_id: voice_id,
66
124
  text: text,
67
- cost_info: cost_info.to_h
125
+ cost_info: cost_info.to_h,
126
+ request_id: audio.request_id
68
127
  )
69
128
 
70
129
  audio
@@ -72,6 +131,8 @@ module ElevenRb
72
131
 
73
132
  # Stream audio from text
74
133
  #
134
+ # Takes the same optional keywords as {#generate}.
135
+ #
75
136
  # @param text [String] the text to convert
76
137
  # @param voice_id [String] the voice ID to use
77
138
  # @param model_id [String] the model to use
@@ -79,18 +140,15 @@ module ElevenRb
79
140
  # @param output_format [String] audio output format
80
141
  # @yield [String] each chunk of audio data
81
142
  # @return [void]
82
- def stream(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {}, output_format: 'mp3_44100_128', &block)
83
- validate_text!(text)
143
+ def stream(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {}, output_format: 'mp3_44100_128',
144
+ language_code: nil, apply_text_normalization: nil, seed: nil, previous_text: nil,
145
+ next_text: nil, previous_request_ids: nil, next_request_ids: nil,
146
+ pronunciation_dictionary_locators: nil, use_pvc_as_ivc: nil, &block)
147
+ validate_text!(text, model_id)
84
148
  validate_presence!(voice_id, 'voice_id')
85
149
  raise ArgumentError, 'Block required for streaming' unless block_given?
86
150
 
87
- settings = Objects::VoiceSettings::DEFAULTS.merge(voice_settings)
88
-
89
- body = {
90
- text: text,
91
- model_id: model_id,
92
- voice_settings: settings
93
- }
151
+ body, = build_body(text, model_id, voice_settings, optional_values(binding))
94
152
 
95
153
  path = "/text-to-speech/#{voice_id}/stream?output_format=#{output_format}"
96
154
  post_stream(path, body, &block)
@@ -102,33 +160,35 @@ module ElevenRb
102
160
  audio: nil, # No audio object for streaming
103
161
  voice_id: voice_id,
104
162
  text: text,
105
- cost_info: cost_info.to_h
163
+ cost_info: cost_info.to_h,
164
+ request_id: nil
106
165
  )
107
166
  end
108
167
 
109
168
  # Generate audio with timestamps
110
169
  #
170
+ # Takes the same optional keywords as {#generate}.
171
+ #
111
172
  # @param text [String] the text to convert
112
173
  # @param voice_id [String] the voice ID to use
113
174
  # @param model_id [String] the model to use
114
175
  # @param voice_settings [Hash] voice settings overrides
115
176
  # @param output_format [String] audio output format
116
- # @return [Hash] contains :audio and :alignment data
177
+ # @return [Hash] `{ audio:, alignment:, normalized_alignment:, request_id:, character_cost: }`
117
178
  def generate_with_timestamps(text, voice_id:, model_id: DEFAULT_MODEL, voice_settings: {},
118
- output_format: 'mp3_44100_128')
119
- validate_text!(text)
179
+ output_format: 'mp3_44100_128', language_code: nil,
180
+ apply_text_normalization: nil, seed: nil, previous_text: nil,
181
+ next_text: nil, previous_request_ids: nil, next_request_ids: nil,
182
+ pronunciation_dictionary_locators: nil, use_pvc_as_ivc: nil)
183
+ validate_text!(text, model_id)
120
184
  validate_presence!(voice_id, 'voice_id')
121
-
122
- settings = Objects::VoiceSettings::DEFAULTS.merge(voice_settings)
123
-
124
- body = {
125
- text: text,
126
- model_id: model_id,
127
- voice_settings: settings
128
- }
185
+ body, dropped = build_body(text, model_id, voice_settings, optional_values(binding))
129
186
 
130
187
  path = "/text-to-speech/#{voice_id}/with-timestamps?output_format=#{output_format}"
131
- response = post(path, body)
188
+ result = post_with_meta(path, body)
189
+ response = result[:body]
190
+ request_id = result[:headers]['request-id']
191
+ character_cost = integer_header(result[:headers], 'character-cost')
132
192
 
133
193
  # Decode base64 audio
134
194
  audio_data = Base64.decode64(response['audio_base64']) if response['audio_base64']
@@ -139,25 +199,72 @@ module ElevenRb
139
199
  format: output_format,
140
200
  voice_id: voice_id,
141
201
  text: text,
142
- model_id: model_id
202
+ model_id: model_id,
203
+ request_id: request_id,
204
+ character_cost: character_cost,
205
+ dropped_settings: dropped
143
206
  )
144
207
  end
145
208
 
146
209
  {
147
210
  audio: audio,
148
- alignment: response['alignment']
211
+ alignment: response['alignment'],
212
+ normalized_alignment: response['normalized_alignment'],
213
+ request_id: request_id,
214
+ character_cost: character_cost
149
215
  }
150
216
  end
151
217
 
152
218
  private
153
219
 
154
- def validate_text!(text)
220
+ # Shared request body for generate / stream / generate_with_timestamps
221
+ #
222
+ # @return [Array(Hash, Array<Symbol>)] the body and the dropped voice-setting keys
223
+ def build_body(text, model_id, voice_settings, options)
224
+ settings, dropped = resolve_voice_settings(model_id, voice_settings)
225
+
226
+ body = { text: text, model_id: model_id, voice_settings: settings }
227
+ OPTIONAL_BODY_KEYS.each do |key|
228
+ body[key] = options[key] unless options[key].nil?
229
+ end
230
+
231
+ [body, dropped]
232
+ end
233
+
234
+ # The optional keyword values of the calling method, keyed by OPTIONAL_BODY_KEYS
235
+ def optional_values(caller_binding)
236
+ OPTIONAL_BODY_KEYS.to_h { |key| [key, caller_binding.local_variable_get(key)] }
237
+ end
238
+
239
+ def resolve_voice_settings(model_id, voice_settings)
240
+ settings, dropped = Objects::VoiceSettings.for_model(model_id, voice_settings)
241
+ return [settings, dropped] if dropped.empty?
242
+
243
+ message = "voice settings #{dropped.join(', ')} are not supported by #{model_id} " \
244
+ "(supported: #{ModelCapabilities.supported_voice_settings(model_id).join(', ')})"
245
+ raise Errors::ValidationError, message if http_client.config.strict_voice_settings
246
+
247
+ http_client.config.logger&.warn("[ElevenRb] #{message}; dropped from the request")
248
+ [settings, dropped]
249
+ end
250
+
251
+ def integer_header(headers, name)
252
+ value = headers[name]
253
+ return nil if value.nil? || value.to_s.strip.empty?
254
+
255
+ Integer(value.to_s.strip, 10)
256
+ rescue ArgumentError
257
+ nil
258
+ end
259
+
260
+ def validate_text!(text, model_id = DEFAULT_MODEL)
155
261
  validate_presence!(text, 'text')
156
262
 
157
- return unless text.length > MAX_TEXT_LENGTH
263
+ max_length = ModelCapabilities.max_text_length(model_id)
264
+ return unless text.length > max_length
158
265
 
159
266
  raise Errors::ValidationError,
160
- "text exceeds maximum length of #{MAX_TEXT_LENGTH} characters (got #{text.length})"
267
+ "text exceeds maximum length of #{max_length} characters for #{model_id} (got #{text.length})"
161
268
  end
162
269
  end
163
270
  end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module ElevenRb
4
- VERSION = '1.0.0'
4
+ VERSION = '1.1.0'
5
5
  end
data/lib/eleven_rb.rb CHANGED
@@ -60,6 +60,7 @@ module ElevenRb
60
60
  retry_delay: config.retry_delay,
61
61
  retry_statuses: config.retry_statuses,
62
62
  logger: config.logger,
63
+ strict_voice_settings: config.strict_voice_settings,
63
64
  on_request: config.on_request,
64
65
  on_response: config.on_response,
65
66
  on_error: config.on_error,
@@ -79,6 +80,7 @@ require_relative 'eleven_rb/errors'
79
80
  require_relative 'eleven_rb/callbacks'
80
81
  require_relative 'eleven_rb/instrumentation'
81
82
  require_relative 'eleven_rb/configuration'
83
+ require_relative 'eleven_rb/model_capabilities'
82
84
 
83
85
  # HTTP layer
84
86
  require_relative 'eleven_rb/http/client'
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: eleven_rb
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.0.0
4
+ version: 1.1.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Web Ventures Ltd
@@ -144,6 +144,7 @@ files:
144
144
  - lib/eleven_rb/errors.rb
145
145
  - lib/eleven_rb/http/client.rb
146
146
  - lib/eleven_rb/instrumentation.rb
147
+ - lib/eleven_rb/model_capabilities.rb
147
148
  - lib/eleven_rb/objects/audio.rb
148
149
  - lib/eleven_rb/objects/base.rb
149
150
  - lib/eleven_rb/objects/cost_info.rb