ruby_llm-contract 1.1.2 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 7ae19c7459f4999409748c604e8b80aa756e5b2333863c8cc16b4070bf57dff9
4
- data.tar.gz: 3be220b3be5c0c313f4b3f5160da5416dc9be8c35629c0012d756e0d027aa646
3
+ metadata.gz: fead48c5516f8e0aa738d3ab1633ca2d07f7173aba6466a6d29e37634263b857
4
+ data.tar.gz: 4ae4c314cd6365c5f5de455f9dc5ddc0741abd26fa751804b22c5ffded0e6cb1
5
5
  SHA512:
6
- metadata.gz: 16bc0227706b8e1cc16f5d93c7771f82e7835c40c20888cec9163a9537d30f80a86f761c47fd10bd778df8f7023ce041bd0d574248950e11bdad5dd3ac4860da
7
- data.tar.gz: 4d4374cc2d65aa555f30656c2475bae261b0af1be0656b7449c22c38467b7f1e32f7ba159d215adc17254d0c624f7144f91087644f578b5cb2727fc8f9087c91
6
+ metadata.gz: 9b66df1e490e238fdacc84a728b874a008b80c8b44cdd882b444a84389bb77996ee6132efb6adaad79c91a04bee30eb7351241a4989ff1a4ed5df8d7a6cff619
7
+ data.tar.gz: a055fd3d06ee4fa9591fa2a6cfed6872142febd950fe3a455daa6e588f04e157a585701456f4815483ecc3accb803cf3bbcffa9308c6dfa5011c00f6a7143274
data/CHANGELOG.md CHANGED
@@ -1,5 +1,52 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.2.0 (2026-10-10)
4
+
5
+ Everything new is opt-in, except the temperature fix below.
6
+
7
+ ### Added
8
+
9
+ - **`token_count :exact`** on a step: `max_input` and `max_cost` measure the input with
10
+ the provider's count (RubyLLM's `chat.count_tokens`: OpenAI Responses, Anthropic,
11
+ Gemini, Vertex AI, Bedrock Converse) instead of the chars/4 heuristic, attachments
12
+ included, so `attachment_token_estimate` is not needed. One extra request per attempt.
13
+ A provider without the endpoint, a failed count or an adapter without `count_tokens`
14
+ refuses the call (`:limit_exceeded`) rather than falling back to the heuristic; the
15
+ `Test` adapter counts with the heuristic. What a provider counts beyond prompt and
16
+ attachments is its own (Gemini's count leaves the output schema out).
17
+ `Adapters::RubyLLM#count_tokens`.
18
+ - **`on_incomplete_output :refuse`** on a step: a response the provider stopped at a token
19
+ limit fails as `:output_truncated`, one stopped by a content filter as
20
+ `:content_filtered`, before validation, so it is never `:ok`. Both statuses are outside
21
+ the default `retry_on`. The default, `:accept`, keeps 1.1.2's behaviour.
22
+ - **`retry_on` takes exception classes**: `retry_on :validation_failed,
23
+ RubyLLM::RateLimitError` retries an adapter error only of that class or a subclass;
24
+ `:adapter_error` still retries any. An adapter error's trace (and attempt) has
25
+ `error_class`.
26
+ - **`provider_options`** on a step (inherited, `:default` resets) and per call
27
+ (`context: { provider_options: }`, merged over the step's), passed to RubyLLM's
28
+ `with_provider_options`. Keys the step controls raise `ArgumentError`.
29
+ - **`workflow_instrumentation`** in `RubyLLM::Contract.configure`: a pipeline runs as one
30
+ `RubyLLM.workflow` with a `workflow.step` per alias, a step on its own as a workflow
31
+ named after its class, so RubyLLM's events and OpenTelemetry spans carry the step.
32
+ - **`register_model` takes cache prices**: `cache_read_per_1m:` and `cache_write_per_1m:`,
33
+ both optional.
34
+ - Guides: retry by error class, cut-off responses, exact token counts, local and custom
35
+ models (Ollama), provider options, workflow instrumentation, and `RubyLLM::Judge` (2.1)
36
+ in the LLM-judge guide. The README lists the new options and answers how to run local
37
+ models.
38
+
39
+ ### Changed
40
+
41
+ - **A temperature is no longer sent to a model RubyLLM's registry marks as taking none**
42
+ (OpenAI's o-series and gpt-5 family). RubyLLM 1.x quietly sent 1.0 to those models;
43
+ 2.x sends the value as given, so a step with `temperature` that escalated to one failed
44
+ with an adapter error. The adapter now leaves it out and warns once per model; a model
45
+ the registry does not know still gets it.
46
+ - **`retry_on` raises on a condition it cannot match.** A status given as a String (or
47
+ any value that is neither a Symbol nor an exception class) never matched a result, so
48
+ it silently turned retries off; it now raises `ArgumentError` when the step is defined.
49
+
3
50
  ## 1.1.2 (2026-10-10)
4
51
 
5
52
  ### Changed - costs and gates may move
data/README.md CHANGED
@@ -131,6 +131,10 @@ Everything below is optional — the example above is a complete step. Reach for
131
131
  - **[A/B test prompts](docs/guide/eval_first.md)** — measure whether a new prompt is safe to ship before merging.
132
132
  - **[Budget caps](docs/guide/getting_started.md)** — refuse the request pre-flight when an estimate exceeds the limit.
133
133
  - **[Cost tracking](docs/guide/getting_started.md#what-a-traces-usage-and-cost-count)** - per-call cost with prompt-cache and per-provider prices; eval cost gates fail closed when a cost is unknown.
134
+ - **[Exact token counts](docs/guide/getting_started.md#counting-input-tokens-exactly)** - `token_count :exact` checks `max_input` / `max_cost` against the provider's own count instead of a heuristic.
135
+ - **[Cut-off responses](docs/guide/getting_started.md#responses-cut-off-by-a-token-limit)** - `on_incomplete_output :refuse` fails an answer the provider stopped at a token limit, even one that would validate.
136
+ - **[Retry by error class](docs/guide/getting_started.md#retrying-provider-errors-by-class)** - `retry_on RubyLLM::RateLimitError` retries a rate limit without retrying an exhausted balance.
137
+ - **[Provider options](docs/guide/getting_started.md#provider-options)** - send `service_tier`, `seed` and other provider settings from the step or per call.
134
138
  - **[Reasoning effort / thinking config](docs/guide/optimizing_retry_policy.md)** — Anthropic / OpenAI thinking configuration on the Step class.
135
139
 
136
140
  Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and per-step models.
@@ -194,6 +198,8 @@ what 1.0.0 itself was.
194
198
 
195
199
  **Costs went up, or an eval cost gate turned red, after upgrading to 1.1.2?** Earlier versions left prompt-cache tokens out of `input_tokens` and the cost, priced every model at its default provider, counted a missing price as $0, and treated a call without token counts as free. The [CHANGELOG](CHANGELOG.md) lists each change; `on_unknown_pricing: :warn` turns the unknown-cost refusal into a warning.
196
200
 
201
+ **Local models (Ollama)?** Point RubyLLM at the server (`c.ollama_api_base = "http://localhost:11434/v1"`) and run steps with `context: { provider: :ollama, model: "...", assume_model_exists: true }`. Local models have no price in RubyLLM's registry, so register one (`CostCalculator.register_model("gemma-fast:latest", input_per_1m: 0, output_per_1m: 0)`) or the eval cost gates treat their cost as unknown. Ollama has no token-counting endpoint, so leave `token_count` at `:estimate`. See [local and custom models](docs/guide/getting_started.md#local-and-custom-models).
202
+
197
203
  **Upgraded from pre-0.10.0 and getting `:limit_exceeded` with attachments?** Multimodal contracts with `max_cost`/`max_input` need `attachment_token_estimate`. See [multimodal input guide](docs/guide/multimodal_input.md#cost-attachment_token_estimate-is-required) for setup, fail-closed behaviour, and `on_unknown_attachment_size :warn` opt-out.
198
204
 
199
205
  ## License
@@ -65,6 +65,32 @@ result.trace[:attempts]
65
65
 
66
66
  If the whole chain exhausts, `result.status` is the status of the last attempt (`:validation_failed` or `:parse_error`) and `result.parsed_output` is the last attempt's output. The caller decides what to do — ship it anyway, fall back to a template, or raise.
67
67
 
68
+ ### Retrying provider errors by class
69
+
70
+ A provider error that RubyLLM's own retries could not get past ends as `:adapter_error`, with the exception class in `trace[:error_class]` (e.g. `"RubyLLM::RateLimitError"`). Adapter errors are not retried by default. List exception classes in `retry_on` to retry only those, typically together with `escalate` so the next attempt goes to another model:
71
+
72
+ ```ruby
73
+ retry_policy do
74
+ escalate "gpt-4.1-mini", "claude-haiku-4-5"
75
+ retry_on :validation_failed, :parse_error, RubyLLM::RateLimitError, RubyLLM::OverloadedError
76
+ end
77
+ ```
78
+
79
+ A class matches its subclasses; `:adapter_error` in the list retries every adapter error. As with statuses, the list replaces the defaults. Switching to a fallback model on transport errors inside one call is RubyLLM's `chat.with_fallbacks`; the contract's adapter does not call it, so fallbacks across models belong in `escalate`.
80
+
81
+ ### Responses cut off by a token limit
82
+
83
+ When a provider stops at a token limit or a content filter, `trace[:finish_reason]` says so (`:max_tokens`, `:content_filter`), and a failed result gets a last validation error naming it. A cut-off answer that still parses and validates stays `:ok` unless the step refuses it:
84
+
85
+ ```ruby
86
+ class SummarizeArticle < RubyLLM::Contract::Step::Base
87
+ max_output 400
88
+ on_incomplete_output :refuse # default :accept
89
+ end
90
+ ```
91
+
92
+ With `:refuse`, such a response fails before validation as `:output_truncated` (token limit) or `:content_filtered`, so `validate` blocks and observers never see it. Raw output, usage and cost stay on the result. Neither status is retried by default; `retry_on :output_truncated` opts in (useful when a later attempt has a larger `max_output`). Anthropic reports an exhausted context window as a token-limit stop too, so `:output_truncated` means "stopped on tokens", not necessarily "hit max_output".
93
+
68
94
  ### Per-attempt reasoning effort
69
95
 
70
96
  `models:` accepts config hashes as well as model-name strings, so a fallback can "try harder" (more reasoning) on retry, not just switch model:
@@ -191,6 +217,31 @@ When a cost does not cover the whole call, `result.trace.cost_unknown?` is true
191
217
 
192
218
  A custom adapter returning `Response.new(content:, usage:)` is priced from `usage` as before. To report its own cost, pass `cost:` - `nil` means unknown - with `cost_complete:` and `usage_complete:`.
193
219
 
220
+ ### Counting input tokens exactly
221
+
222
+ `max_input` and `max_cost` measure the input with a chars/4 heuristic (±30%). With `token_count :exact` they ask the provider instead:
223
+
224
+ ```ruby
225
+ class SummarizeArticle < RubyLLM::Contract::Step::Base
226
+ token_count :exact # default :estimate
227
+ max_input 20_000
228
+ end
229
+ ```
230
+
231
+ The count comes from RubyLLM's `chat.count_tokens` (OpenAI Responses, Anthropic, Gemini, Vertex AI, Bedrock Converse) and covers the prompt and attachments, so `attachment_token_estimate` is not needed. What else a provider counts is its own: OpenAI's and Anthropic's counts include the output schema, Gemini's leaves it out. It costs one extra request per attempt, and RubyLLM leaves `provider_options` out of it. Where no count can be had - a provider without the endpoint (Ollama, for one), a failed request, or an adapter without `count_tokens` - the call is refused as `:limit_exceeded` with the reason, never measured with the heuristic under the name "exact". The `Test` adapter counts with the heuristic, since it has no provider to ask. The output side of `max_cost` is still an estimate (`max_output`, or the input size without one).
232
+
233
+ ### Local and custom models
234
+
235
+ A local model (Ollama) or a fine-tuned one has no price in RubyLLM's registry, so its cost is unknown and the eval cost gates refuse. Register a price - zero for a model that costs nothing per token:
236
+
237
+ ```ruby
238
+ RubyLLM::Contract::CostCalculator.register_model("gemma-fast:latest", input_per_1m: 0, output_per_1m: 0)
239
+ RubyLLM::Contract::CostCalculator.register_model("ft:gpt-4o-custom",
240
+ input_per_1m: 3.0, output_per_1m: 6.0, cache_read_per_1m: 1.5, cache_write_per_1m: 3.75)
241
+ ```
242
+
243
+ `register_model` sets a price for a model id, not a model RubyLLM can call: the step still needs `provider:` (and `assume_model_exists: true` for an id RubyLLM does not list). Cache prices are optional; without them cache reads are charged at the input price and a call that wrote to the cache has no price. A price does not make up for missing counts: a provider that sends no token counts still leaves the cost unknown.
244
+
194
245
  ### Preflight cost estimates
195
246
 
196
247
  Check what a call is likely to cost before invoking it:
@@ -211,6 +262,24 @@ SummarizeArticle.estimate_eval_cost("regression",
211
262
 
212
263
  `estimate_cost` returns `nil` when pricing isn't registered. `estimate_eval_cost` silently treats unknown-pricing cases as `$0.00` and sums the rest — it does **not** fail closed the way `max_cost` does. Treat its output as a floor, not a guarantee; register pricing via `CostCalculator.register_model` before relying on it for budget decisions.
213
264
 
265
+ ## Provider options
266
+
267
+ Settings RubyLLM has no method for go to the provider as-is through `provider_options`, on the step or per call:
268
+
269
+ ```ruby
270
+ class SummarizeArticle < RubyLLM::Contract::Step::Base
271
+ provider_options service_tier: "flex"
272
+ end
273
+
274
+ SummarizeArticle.run(text, context: { provider_options: { service_tier: "priority", seed: 7 } })
275
+ ```
276
+
277
+ The call's keys win over the step's; a subclass replaces its parent's hash, and `provider_options :default` stops inheriting it. Keys the step itself sets (`model`, `temperature`, the max-token keys, the messages, and on a step with `output_schema` the structured-output keys `response_format` and `text`) raise `ArgumentError`, so a call cannot differ from what its limits and schema were checked against. Only top-level keys are checked. `compare_models` and `optimize_retry_policy` candidates do not carry provider options.
278
+
279
+ ## Temperature on reasoning models
280
+
281
+ RubyLLM 1.x sent temperature 1.0 to OpenAI reasoning models (o-series, gpt-5), which reject other values; RubyLLM 2.x sends what it is given. A step with `temperature` therefore leaves it out for a model RubyLLM's registry marks as taking none, and warns once per model - so `temperature 0` with `escalate "gpt-4.1-mini", "gpt-5-mini"` still works. A model the registry does not know gets the temperature as set.
282
+
214
283
  ## `output_schema` vs `with_schema`
215
284
 
216
285
  `with_schema` in `ruby_llm` tells the provider to force a specific JSON structure. `output_schema` in this gem does the same thing (calls `with_schema` under the hood) **plus** validates the response client-side. Cheaper models sometimes ignore schema constraints — `with_schema` is a request; `output_schema` is a request plus verification.
@@ -195,6 +195,20 @@ What to do operationally:
195
195
 
196
196
  The drop is your signal to refine the judge's prompt, not to lower the gate.
197
197
 
198
+ ## `RubyLLM::Judge` (RubyLLM 2.1)
199
+
200
+ RubyLLM 2.1 adds `RubyLLM::Judge` for typed questions - `probability`, `choice` and `score` - answered by **decision models**: OpenAI's judgment models, TypeSafe, or local decision models through Ollama. It does not run on ordinary chat models; those raise that they do not support judgments. Where you have such a model, a judge class can stand in for the judge Step in step 1 above:
201
+
202
+ ```ruby
203
+ class SummaryFaithfulness < RubyLLM::Judge
204
+ probability :faithful, "Is every claim in the summary supported by the article?"
205
+ end
206
+
207
+ SummaryFaithfulness.judge(article_and_summary).faithful.probability # => 0.93
208
+ ```
209
+
210
+ Calibrate it against human labels exactly as in step 2; the threshold is still yours to set. On an ordinary chat model, keep the judge as a Contract Step - the pattern this guide builds.
211
+
198
212
  ## When to reach for Tribunal instead
199
213
 
200
214
  [`ruby_llm-tribunal`](https://github.com/Alqemist-labs/ruby_llm-tribunal) ships an off-the-shelf catalog of common LLM-as-judge assertions (`assert_faithful`, `assert_hallucination`, `assert_refusal`, `assert_no_pii`, etc.) — a shortcut when your check matches one of those domain-general categories. The methodology in this guide still applies: calibrate the judge against your human-labeled production data **before** trusting Tribunal's `default_threshold = 0.8`, refine the prompt when it over-flags, watch for the anti-patterns above. See [Relation to Tribunal](relation_to_tribunal.md) for the full positioning — what each gem documents (and doesn't), a concrete decision tree on catalog-vs-custom-judge, and three working integration patterns.
@@ -125,6 +125,14 @@ end
125
125
 
126
126
  Trace inspection in an admin UI: `result.trace[:attempts]` gives you per-attempt model, status, cost, latency — render it in a partial to debug production failures without re-running.
127
127
 
128
+ RubyLLM 2.1 emits its own events and OpenTelemetry spans for every call. To group them by step, turn on workflow instrumentation in the initializer:
129
+
130
+ ```ruby
131
+ RubyLLM::Contract.configure { |c| c.workflow_instrumentation = true }
132
+ ```
133
+
134
+ A pipeline then runs as one `RubyLLM.workflow` named after its class, with each step as `workflow.step(alias)`; a step run on its own is a workflow named after its class. RubyLLM's events carry `workflow_name` and `workflow_step_name`, and a workflow your code opened around the call becomes the parent. Off by default. In a Rails app RubyLLM also records each call in its usage ledger when configured; that figure and `trace[:cost]` are the same RubyLLM price, unless `CostCalculator.register_model` overrides it for the model.
135
+
128
136
  ## 5. Testing — RSpec and Minitest
129
137
 
130
138
  Add to `spec/spec_helper.rb` (or `test_helper.rb`):
@@ -6,29 +6,59 @@ module RubyLLM
6
6
  module Contract
7
7
  module Adapters
8
8
  class RubyLLM < Base
9
+ # `with: nil` adds no attachment. The text is never nil (`fetch(:content, "")`),
10
+ # and RubyLLM raises only when both text and attachments are nil.
9
11
  def call(messages:, **options)
10
- system_contents, conversation = partition_messages(messages)
11
- conversation = fallback_conversation(system_contents, conversation)
12
-
13
- chat = build_chat(options, system_contents)
14
- add_history(chat, conversation[0..-2])
15
-
16
- # `with: nil` adds no attachment. The text is never nil (`fetch(:content, "")`),
17
- # and RubyLLM raises only when both text and attachments are nil.
18
- response = chat.ask(
19
- conversation.last&.fetch(:content, ""),
20
- with: options[:attachment]
21
- )
12
+ chat, text = prepared_chat(messages, options)
13
+ response = chat.ask(text, with: options[:attachment])
22
14
  build_response(response, options[:model])
23
15
  end
24
16
 
17
+ # Input tokens of the request `call` would send, counted by the
18
+ # provider (`chat.count_tokens`), attachments included. `ask` is
19
+ # `ask_later` plus `complete` in RubyLLM, so the staged message is the
20
+ # one `call` sends. RubyLLM leaves provider_options out of the count.
21
+ # Raises RubyLLM::Error where the provider has no counting endpoint.
22
+ def count_tokens(messages:, **options)
23
+ chat, text = prepared_chat(messages, options)
24
+ chat.ask_later(text, with: options[:attachment]).count_tokens
25
+ rescue KeyError => e
26
+ # RubyLLM reads OpenAI's count with `fetch`; a reply without the
27
+ # field is a failed count, not a bug in the caller.
28
+ raise ::RubyLLM::Error, "count response without input_tokens (#{e.message})"
29
+ end
30
+
25
31
  CHAT_OPTION_METHODS = {
26
- temperature: :with_temperature,
27
32
  schema: :with_schema
28
33
  }.freeze
29
34
 
35
+ @dropped_temperature_models = Set.new
36
+ @dropped_temperature_lock = Mutex.new
37
+
38
+ # Warns once per provider and model, across threads and subclasses
39
+ # (called on this class, whose state subclasses do not inherit).
40
+ def self.warn_dropped_temperature(model)
41
+ key = "#{model.provider}:#{model.id}"
42
+ first = @dropped_temperature_lock.synchronize { @dropped_temperature_models.add?(key) }
43
+ return unless first
44
+
45
+ warn "[ruby_llm-contract] #{model.id} (#{model.provider}) does not accept temperature " \
46
+ "according to RubyLLM's model registry; the step's temperature is not sent"
47
+ end
48
+
30
49
  private
31
50
 
51
+ # The configured chat with the history added, and the text of the last
52
+ # message, which sending and counting both stage themselves.
53
+ def prepared_chat(messages, options)
54
+ system_contents, conversation = partition_messages(messages)
55
+ conversation = fallback_conversation(system_contents, conversation)
56
+
57
+ chat = build_chat(options, system_contents)
58
+ add_history(chat, conversation[0..-2])
59
+ [chat, conversation.last&.fetch(:content, "")]
60
+ end
61
+
32
62
  # When prompt has only system/section/rule nodes and no user message,
33
63
  # pop the last system message and use it as the user ask.
34
64
  def fallback_conversation(system_contents, conversation)
@@ -56,6 +86,7 @@ module RubyLLM
56
86
  CHAT_OPTION_METHODS.each do |key, method_name|
57
87
  chat.public_send(method_name, options[key]) if options[key]
58
88
  end
89
+ apply_temperature(chat, options[:temperature]) unless options[:temperature].nil?
59
90
 
60
91
  # Resolve thinking config from BOTH sources, with `:reasoning_effort`
61
92
  # taking precedence over `:thinking[:effort]`. This is the per-attempt
@@ -69,6 +100,29 @@ module RubyLLM
69
100
  # replacement for this one passthrough. `reasoning_effort` is not
70
101
  # forwarded here — it goes through `with_thinking` above.
71
102
  chat.with_max_output_tokens(options[:max_tokens]) if options[:max_tokens]
103
+ apply_provider_options(chat, options)
104
+ end
105
+
106
+ def apply_provider_options(chat, options)
107
+ provider_options = options[:provider_options]
108
+ return if provider_options.nil? || provider_options.empty?
109
+
110
+ ProviderOptions.validate!(provider_options, schema: !options[:schema].nil?)
111
+ chat.with_provider_options(provider_options)
112
+ end
113
+
114
+ # RubyLLM 1.x set temperature to 1.0 for OpenAI reasoning models, which
115
+ # reject any other value; 2.x sends it as given. A step escalating from
116
+ # a sampling model to one of those would fail, so a temperature is left
117
+ # out where the registry says the model takes none. A model RubyLLM does
118
+ # not know (assume_model_exists) gets it as before.
119
+ def apply_temperature(chat, temperature)
120
+ model = chat.respond_to?(:model) ? chat.model : nil
121
+ if model.respond_to?(:metadata) && model.metadata.is_a?(Hash) && model.metadata[:temperature] == false
122
+ Adapters::RubyLLM.warn_dropped_temperature(model)
123
+ else
124
+ chat.with_temperature(temperature)
125
+ end
72
126
  end
73
127
 
74
128
  # Returns merged `{ effort:, budget: }` or nil. `options[:reasoning_effort]`
@@ -43,6 +43,12 @@ module RubyLLM
43
43
  end
44
44
  end
45
45
 
46
+ # There is no provider to ask, so `token_count :exact` steps measure
47
+ # their input with the same heuristic as `:estimate` under this adapter.
48
+ def count_tokens(messages:, **_options)
49
+ TokenEstimator.estimate(messages)
50
+ end
51
+
46
52
  def call(messages:, **_options) # rubocop:disable Lint/UnusedMethodArgument
47
53
  content = if @responses
48
54
  c = @responses[@index] || @responses.last
@@ -9,13 +9,17 @@ module RubyLLM
9
9
  #
10
10
  # Then configure contract-specific options:
11
11
  # RubyLLM::Contract.configure { |c| c.default_model = "gpt-4.1-mini" }
12
+ #
13
+ # `workflow_instrumentation = true` wraps steps and pipelines in
14
+ # `RubyLLM.workflow` (see WorkflowScope).
12
15
  class Configuration
13
- attr_accessor :default_adapter, :default_model, :logger
16
+ attr_accessor :default_adapter, :default_model, :logger, :workflow_instrumentation
14
17
 
15
18
  def initialize
16
19
  @default_adapter = nil
17
20
  @default_model = nil
18
21
  @logger = nil
22
+ @workflow_instrumentation = false
19
23
  end
20
24
  end
21
25
  end
@@ -29,7 +29,9 @@ module RubyLLM
29
29
  # which calls `calculate` per attempt and sums; not in this module).
30
30
  module CostCalculator
31
31
  # Simple struct for custom-registered model pricing
32
- RegisteredModel = Struct.new(:input_price_per_million, :output_price_per_million, keyword_init: true)
32
+ RegisteredModel = Struct.new(:input_price_per_million, :output_price_per_million,
33
+ :cache_read_price_per_million, :cache_write_price_per_million,
34
+ keyword_init: true)
33
35
 
34
36
  @custom_models = {}
35
37
 
@@ -40,13 +42,22 @@ module RubyLLM
40
42
  # CostCalculator.register_model("ft:gpt-4o-custom",
41
43
  # input_per_1m: 3.0, output_per_1m: 6.0)
42
44
  #
43
- def self.register_model(model_name, input_per_1m:, output_per_1m:)
45
+ # Prompt-cache prices are optional: without `cache_read_per_1m` cache
46
+ # reads are charged at the input price; without `cache_write_per_1m` a
47
+ # call that wrote to the cache has no price (writes can cost more than
48
+ # input).
49
+ def self.register_model(model_name, input_per_1m:, output_per_1m:, cache_read_per_1m: nil,
50
+ cache_write_per_1m: nil)
44
51
  validate_price!(:input_per_1m, input_per_1m)
45
52
  validate_price!(:output_per_1m, output_per_1m)
53
+ validate_price!(:cache_read_per_1m, cache_read_per_1m) unless cache_read_per_1m.nil?
54
+ validate_price!(:cache_write_per_1m, cache_write_per_1m) unless cache_write_per_1m.nil?
46
55
 
47
56
  @custom_models[model_name] = RegisteredModel.new(
48
57
  input_price_per_million: input_per_1m,
49
- output_price_per_million: output_per_1m
58
+ output_price_per_million: output_per_1m,
59
+ cache_read_price_per_million: cache_read_per_1m,
60
+ cache_write_price_per_million: cache_write_per_1m
50
61
  )
51
62
  end
52
63
 
@@ -113,18 +124,32 @@ module RubyLLM
113
124
  end
114
125
  end
115
126
 
116
- # Our registry knows input and output prices only. Cache reads are
127
+ # input_tokens includes the cache reads and writes, so each part is
128
+ # charged once, at its own price. Without a cache-read price, reads are
117
129
  # charged at the input price, which assumes they cost no more than
118
130
  # uncached input (true of every provider in RubyLLM's registry, not
119
131
  # guaranteed for a custom one). Cache writes can cost more than input
120
- # (Anthropic charges 1.25x-2x), so a call that wrote to the cache has
121
- # no price here rather than an undercount.
132
+ # (Anthropic charges 1.25x-2x), so without a cache-write price a call
133
+ # that wrote to the cache has no price rather than an undercount.
122
134
  def self.registered_cost(model_info, usage)
123
- return nil if (usage[:cache_write_tokens] || 0).positive?
135
+ cache_read = usage[:cache_read_tokens] || 0
136
+ cache_write = usage[:cache_write_tokens] || 0
137
+ read_price, write_price = registered_cache_prices(model_info)
138
+ return nil if cache_write.positive? && write_price.nil?
139
+
140
+ uncached = [(usage[:input_tokens] || 0) - cache_read - cache_write, 0].max
141
+ (token_cost(uncached, model_info.input_price_per_million) +
142
+ token_cost(cache_read, read_price) +
143
+ token_cost(cache_write, write_price.to_f) +
144
+ token_cost(usage[:output_tokens], model_info.output_price_per_million)).round(6)
145
+ end
124
146
 
125
- input_cost = token_cost(usage[:input_tokens], model_info.input_price_per_million)
126
- output_cost = token_cost(usage[:output_tokens], model_info.output_price_per_million)
127
- (input_cost + output_cost).round(6)
147
+ # [read price, write price]; a model without a read price reads at its
148
+ # input price. Duck-typed registry entries may lack both readers.
149
+ def self.registered_cache_prices(model_info)
150
+ read = model_info.cache_read_price_per_million if model_info.respond_to?(:cache_read_price_per_million)
151
+ write = model_info.cache_write_price_per_million if model_info.respond_to?(:cache_write_price_per_million)
152
+ [read || model_info.input_price_per_million, write]
128
153
  end
129
154
 
130
155
  # RubyLLM counts uncached input apart from cache reads and writes; our
@@ -174,7 +199,8 @@ module RubyLLM
174
199
  # to inspect model pricing before invoking `calculate` (e.g., to short-
175
200
  # circuit estimate when the model is unknown). Exposing it removes
176
201
  # `CostCalculator.send(:find_model)` workarounds at call sites.
177
- private_class_method :compute_cost, :registered_cost, :tokens_for, :token_cost, :validate_price!
202
+ private_class_method :compute_cost, :registered_cost, :registered_cache_prices, :tokens_for, :token_cost,
203
+ :validate_price!
178
204
  end
179
205
  end
180
206
  end
@@ -48,7 +48,9 @@ module RubyLLM
48
48
  end
49
49
 
50
50
  def run(input, context: {}, timeout_ms: nil)
51
- Runner.new(steps: steps, context: context, timeout_ms: timeout_ms, token_budget: token_budget).call(input)
51
+ WorkflowScope.workflow(name || "anonymous pipeline") do
52
+ Runner.new(steps: steps, context: context, timeout_ms: timeout_ms, token_budget: token_budget).call(input)
53
+ end
52
54
  end
53
55
 
54
56
  def test(input, responses: {}, timeout_ms: nil)
@@ -37,7 +37,9 @@ module RubyLLM
37
37
 
38
38
  def execute_step(step_def, execution)
39
39
  step_context = build_step_context(step_def)
40
- result = step_def[:step_class].run(execution.current_input, context: step_context)
40
+ result = WorkflowScope.step(step_def[:alias]) do
41
+ step_def[:step_class].run(execution.current_input, context: step_context)
42
+ end
41
43
 
42
44
  execution.record_step(step_def[:alias], result)
43
45
  end
@@ -0,0 +1,45 @@
1
+ # frozen_string_literal: true
2
+
3
+ module RubyLLM
4
+ module Contract
5
+ # Options in the provider's own request vocabulary, passed to
6
+ # `chat.with_provider_options` and merged into the request as-is. Keys the
7
+ # contract itself sets are refused: an override would make the call differ
8
+ # from what max_cost, max_output and the schema were checked against.
9
+ # Only top-level keys are checked; a provider that nests these settings
10
+ # (Gemini's generationConfig, for one) is not.
11
+ module ProviderOptions
12
+ RESERVED_KEYS = %w[model messages input instructions system temperature stream
13
+ max_tokens max_output_tokens max_completion_tokens].freeze
14
+ # Where OpenAI's protocols put the structured-output schema.
15
+ SCHEMA_KEYS = %w[response_format text].freeze
16
+
17
+ def self.validate!(options, schema: false)
18
+ raise ArgumentError, "provider_options must be a Hash, got #{options.class}" unless options.is_a?(Hash)
19
+
20
+ reserved = schema ? RESERVED_KEYS + SCHEMA_KEYS : RESERVED_KEYS
21
+ clashes = options.keys.map(&:to_s) & reserved
22
+ return options if clashes.empty?
23
+
24
+ raise ArgumentError, "provider_options cannot set #{clashes.join(", ")}: the step sets " \
25
+ "#{clashes.one? ? "it" : "them"} (use the step DSL or context instead)"
26
+ end
27
+
28
+ # A frozen deep copy, so a hash the caller keeps mutating cannot change
29
+ # a step definition shared across threads.
30
+ def self.frozen_copy(object)
31
+ case object
32
+ when Hash then object.to_h { |key, value| [key, frozen_copy(value)] }.freeze
33
+ when Array then object.map { |value| frozen_copy(value) }.freeze
34
+ else object.frozen? ? object : object.dup.freeze
35
+ end
36
+ end
37
+
38
+ # Context keys win over the class-level ones; nil when nothing is set.
39
+ def self.merge(class_level, context_level)
40
+ merged = (class_level || {}).merge(context_level || {})
41
+ merged.empty? ? nil : merged
42
+ end
43
+ end
44
+ end
45
+ end
@@ -30,7 +30,8 @@ module RubyLLM
30
30
  status: :adapter_error,
31
31
  raw_output: nil,
32
32
  parsed_output: nil,
33
- validation_errors: [e.message]
33
+ validation_errors: [e.message],
34
+ trace: Trace.new(error_class: e.class.name, error_type: e.class)
34
35
  )
35
36
  [result, 0]
36
37
  end
@@ -93,7 +93,8 @@ module RubyLLM
93
93
 
94
94
  # Forwarded to the adapter as-is. A key known but not forwarded would be
95
95
  # silently dropped, so the known list is built from this one.
96
- ADAPTER_CONTEXT_KEYS = %i[provider assume_model_exists max_tokens reasoning_effort attachment].freeze
96
+ ADAPTER_CONTEXT_KEYS = %i[provider assume_model_exists max_tokens reasoning_effort attachment
97
+ provider_options].freeze
97
98
  KNOWN_CONTEXT_KEYS = (%i[adapter model temperature retry_policy_override] + ADAPTER_CONTEXT_KEYS).freeze
98
99
 
99
100
  include Concerns::ContextHelpers
@@ -102,7 +103,7 @@ module RubyLLM
102
103
  context = safe_context(context)
103
104
  warn_unknown_context_keys(context)
104
105
 
105
- result = dispatch_run(input, context)
106
+ result = WorkflowScope.workflow(name || "anonymous step") { dispatch_run(input, context) }
106
107
  log_result(result)
107
108
  invoke_around_call(input, result)
108
109
  end
@@ -196,7 +197,7 @@ module RubyLLM
196
197
 
197
198
  def runtime_settings(context)
198
199
  policy = context.key?(:retry_policy_override) ? context[:retry_policy_override] : retry_policy
199
- extra = context.slice(*ADAPTER_CONTEXT_KEYS)
200
+ extra = context.slice(*ADAPTER_CONTEXT_KEYS).merge(merged_provider_options(context))
200
201
 
201
202
  # Always pass the class-level `thinking` config to the adapter when
202
203
  # set, so fields like `budget` survive a per-call `reasoning_effort`
@@ -221,6 +222,13 @@ module RubyLLM
221
222
  }
222
223
  end
223
224
 
225
+ # Class-level provider_options with the call's merged over them; a
226
+ # one-key hash so `merge` drops it entirely when neither is set.
227
+ def merged_provider_options(context)
228
+ merged = ProviderOptions.merge(provider_options, context[:provider_options])
229
+ merged ? { provider_options: merged } : {}
230
+ end
231
+
224
232
  def current_model_config
225
233
  policy = retry_policy
226
234
  if policy && policy.config_list.any?
@@ -275,7 +283,9 @@ module RubyLLM
275
283
  on_unknown_attachment_size: on_unknown_attachment_size,
276
284
  temperature: context_temperature || temperature,
277
285
  extra_options: extra_options,
278
- observers: class_observers
286
+ observers: class_observers,
287
+ on_incomplete_output: on_incomplete_output,
288
+ token_count: token_count
279
289
  )
280
290
  end
281
291
 
@@ -185,6 +185,46 @@ module RubyLLM
185
185
  inherited_value(:on_unknown_attachment_size) || UnknownPolicy::DEFAULT
186
186
  end
187
187
 
188
+ TOKEN_COUNT_MODES = %i[estimate exact].freeze
189
+ TOKEN_COUNT_DEFAULT = :estimate
190
+
191
+ # How max_input and max_cost measure the input before a call.
192
+ # `:estimate` (default) is the local chars/4 heuristic. `:exact` asks
193
+ # the provider (`chat.count_tokens`): one extra request per attempt,
194
+ # attachments included, no attachment_token_estimate needed. A provider
195
+ # or adapter that cannot count refuses the call (:limit_exceeded).
196
+ def token_count(mode = nil)
197
+ if mode
198
+ unless TOKEN_COUNT_MODES.include?(mode)
199
+ raise ArgumentError, "token_count must be one of #{TOKEN_COUNT_MODES.inspect}, got #{mode.inspect}"
200
+ end
201
+
202
+ return @token_count = mode
203
+ end
204
+
205
+ inherited_value(:token_count) || TOKEN_COUNT_DEFAULT
206
+ end
207
+
208
+ INCOMPLETE_OUTPUT_MODES = %i[accept refuse].freeze
209
+ INCOMPLETE_OUTPUT_DEFAULT = :accept
210
+
211
+ # `:refuse` fails a response the provider cut off at a token limit
212
+ # (:output_truncated) or stopped with a content filter
213
+ # (:content_filtered), before validation, so it is never :ok.
214
+ # `:accept` (default) validates it like any other response.
215
+ def on_incomplete_output(mode = nil)
216
+ if mode
217
+ unless INCOMPLETE_OUTPUT_MODES.include?(mode)
218
+ raise ArgumentError, "on_incomplete_output must be one of #{INCOMPLETE_OUTPUT_MODES.inspect}, " \
219
+ "got #{mode.inspect}"
220
+ end
221
+
222
+ return @on_incomplete_output = mode
223
+ end
224
+
225
+ inherited_value(:on_incomplete_output) || INCOMPLETE_OUTPUT_DEFAULT
226
+ end
227
+
188
228
  def model(name = nil)
189
229
  if name == :default
190
230
  @model = UNSET
@@ -215,6 +255,20 @@ module RubyLLM
215
255
  inherited_value_with_reset(:temperature)
216
256
  end
217
257
 
258
+ # Request options in the provider's vocabulary, e.g.
259
+ # `provider_options service_tier: "flex"`. Replaces an inherited hash;
260
+ # `provider_options :default` stops inheriting. `context: { provider_options: }`
261
+ # is merged over it per call.
262
+ def provider_options(options = nil)
263
+ if options == :default
264
+ @provider_options = UNSET
265
+ return nil
266
+ end
267
+ return @provider_options = ProviderOptions.frozen_copy(ProviderOptions.validate!(options)) if options
268
+
269
+ inherited_value_with_reset(:provider_options)
270
+ end
271
+
218
272
  def thinking(effort: nil, budget: nil)
219
273
  if effort == :default
220
274
  @thinking = UNSET
@@ -8,6 +8,7 @@ module RubyLLM
8
8
 
9
9
  def check_limits(messages)
10
10
  return nil unless max_input || max_cost
11
+ return check_counted_limits(messages) if exact_token_count?
11
12
 
12
13
  text_tokens = TokenEstimator.estimate(messages)
13
14
  attachment_tokens, attachment_error = resolve_attachment_tokens
@@ -23,6 +24,31 @@ module RubyLLM
23
24
  build_limit_result(messages, estimated, errors)
24
25
  end
25
26
 
27
+ # `token_count :exact`: the provider counts the input, attachments
28
+ # included. A count that cannot be had refuses the call rather than
29
+ # falling back to the heuristic under the name "exact".
30
+ def check_counted_limits(messages)
31
+ counted, count_error = count_input_tokens(messages)
32
+ return build_limit_result(messages, nil, [count_error], method: :exact) if count_error
33
+
34
+ errors = collect_limit_errors(counted, method: :exact)
35
+ errors.empty? ? nil : build_limit_result(messages, counted, errors, method: :exact)
36
+ end
37
+
38
+ def count_input_tokens(messages)
39
+ unless count_adapter.respond_to?(:count_tokens)
40
+ return [nil, "token_count :exact needs an adapter that counts tokens; " \
41
+ "#{count_adapter.class} does not"]
42
+ end
43
+
44
+ counted = count_adapter.count_tokens(messages: messages, **count_options)
45
+ return [counted, nil] if counted.is_a?(Integer) && !counted.negative?
46
+
47
+ [nil, "token_count :exact: no usable token count (got #{counted.inspect})"]
48
+ rescue ::RubyLLM::Error, ::Faraday::Error => e
49
+ [nil, "token_count :exact: the provider could not count the input tokens (#{e.message})"]
50
+ end
51
+
26
52
  # Fail-closed: when an attachment is passed via context but no
27
53
  # attachment_token_estimate is declared, the gem cannot bound vision/
28
54
  # PDF cost. Refuses with a clear error unless on_unknown_attachment_size
@@ -50,12 +76,17 @@ module RubyLLM
50
76
  [estimate, nil]
51
77
  end
52
78
 
53
- def collect_limit_errors(estimated)
79
+ def collect_limit_errors(estimated, method: :heuristic)
54
80
  errors = []
55
81
  if max_input && estimated > max_input
56
- errors << "Input token limit exceeded: estimated #{estimated} tokens (heuristic ±30%), max #{max_input}"
82
+ measured = if method == :exact
83
+ "#{estimated} tokens (counted by the provider)"
84
+ else
85
+ "estimated #{estimated} tokens (heuristic ±30%)"
86
+ end
87
+ errors << "Input token limit exceeded: #{measured}, max #{max_input}"
57
88
  end
58
- append_cost_error(estimated, errors) if max_cost
89
+ append_cost_error(estimated, errors, method) if max_cost
59
90
  errors
60
91
  end
61
92
 
@@ -66,7 +97,7 @@ module RubyLLM
66
97
  # for models expensive on completion side.
67
98
  DEFAULT_OUTPUT_RATIO = 1
68
99
 
69
- def append_cost_error(estimated, errors)
100
+ def append_cost_error(estimated, errors, method = :heuristic)
70
101
  estimated_output = effective_max_output || (estimated * DEFAULT_OUTPUT_RATIO)
71
102
  estimated_cost = CostCalculator.calculate(
72
103
  model_name: model_name,
@@ -77,8 +108,12 @@ module RubyLLM
77
108
  if estimated_cost.nil?
78
109
  handle_unknown_pricing(errors)
79
110
  elsif estimated_cost > max_cost
80
- errors << "Cost limit exceeded: estimated $#{format("%.6f", estimated_cost)} " \
81
- "(#{estimated} input + #{estimated_output} output tokens, heuristic ±30%), " \
111
+ tokens = if method == :exact
112
+ "#{estimated} input tokens counted by the provider + #{estimated_output} output"
113
+ else
114
+ "#{estimated} input + #{estimated_output} output tokens, heuristic ±30%"
115
+ end
116
+ errors << "Cost limit exceeded: estimated $#{format("%.6f", estimated_cost)} (#{tokens}), " \
82
117
  "max $#{format("%.6f", max_cost)}"
83
118
  end
84
119
  end
@@ -94,17 +129,16 @@ module RubyLLM
94
129
  end
95
130
  end
96
131
 
97
- def build_limit_result(messages, estimated, errors)
132
+ # No request was sent, so usage is zero; the measured input goes
133
+ # alongside, and is left out when it could not be measured.
134
+ def build_limit_result(messages, estimated, errors, method: :heuristic)
135
+ usage = { input_tokens: 0, output_tokens: 0, estimated_input_tokens: estimated, estimate_method: method }
98
136
  Result.new(
99
137
  status: :limit_exceeded,
100
138
  raw_output: nil,
101
139
  parsed_output: nil,
102
140
  validation_errors: errors,
103
- trace: Trace.new(
104
- messages: messages, model: model_name,
105
- usage: { input_tokens: 0, output_tokens: 0, estimated_input_tokens: estimated,
106
- estimate_method: :heuristic }
107
- )
141
+ trace: Trace.new(messages: messages, model: model_name, usage: usage.compact)
108
142
  )
109
143
  end
110
144
  end
@@ -12,13 +12,18 @@ module RubyLLM
12
12
  content_filter: "provider reported content filtering or refusal"
13
13
  }.freeze
14
14
 
15
- def initialize(contract_definition:, output_type:, output_schema:, model:, observers:, max_output: nil)
15
+ # Statuses for `on_incomplete_output :refuse`.
16
+ INCOMPLETE_STATUSES = { max_tokens: :output_truncated, content_filter: :content_filtered }.freeze
17
+
18
+ def initialize(contract_definition:, output_type:, output_schema:, model:, observers:, max_output: nil,
19
+ on_incomplete_output: Dsl::INCOMPLETE_OUTPUT_DEFAULT)
16
20
  @contract_definition = contract_definition
17
21
  @output_type = output_type
18
22
  @output_schema = output_schema
19
23
  @model = model
20
24
  @observers = observers
21
25
  @max_output = max_output
26
+ @on_incomplete_output = on_incomplete_output
22
27
  end
23
28
 
24
29
  def error_result(error_result:, messages:)
@@ -27,15 +32,18 @@ module RubyLLM
27
32
  raw_output: error_result.raw_output,
28
33
  parsed_output: error_result.parsed_output,
29
34
  validation_errors: error_result.validation_errors,
30
- trace: Trace.new(messages: messages, model: @model)
35
+ trace: Trace.new(messages: messages, model: @model, **error_fields(error_result.trace))
31
36
  )
32
37
  end
33
38
 
34
39
  def success_result(response:, messages:, latency_ms:, input:)
35
40
  raw_output = response.content
36
- validation_result = validate_output(raw_output, input)
37
41
  trace = Trace.new(messages: messages, model: @model, latency_ms: latency_ms, usage: response.usage,
38
42
  **response_accounting(response))
43
+ incomplete = incomplete_result(response, trace)
44
+ return incomplete if incomplete
45
+
46
+ validation_result = validate_output(raw_output, input)
39
47
 
40
48
  Result.new(
41
49
  status: validation_result[:status],
@@ -49,18 +57,43 @@ module RubyLLM
49
57
 
50
58
  private
51
59
 
60
+ def error_fields(trace)
61
+ return {} unless trace.respond_to?(:error_class) && trace.error_class
62
+
63
+ { error_class: trace.error_class, error_type: trace.error_type }
64
+ end
65
+
66
+ # Decided before validation, so neither validate blocks nor observers
67
+ # see a response the step refuses.
68
+ def incomplete_result(response, trace)
69
+ return unless @on_incomplete_output == :refuse
70
+
71
+ status = INCOMPLETE_STATUSES[finish_reason(response)]
72
+ return unless status
73
+
74
+ Result.new(status: status, raw_output: response.content, parsed_output: nil,
75
+ validation_errors: [finish_reason_error(finish_reason(response))], trace: trace)
76
+ end
77
+
78
+ def finish_reason(response)
79
+ response.respond_to?(:finish_reason) ? response.finish_reason : nil
80
+ end
81
+
82
+ def finish_reason_error(reason)
83
+ message = FINISH_REASON_ERRORS[reason]
84
+ return message unless message && reason == :max_tokens && @max_output
85
+
86
+ "#{message}; configured max_output: #{@max_output}"
87
+ end
88
+
52
89
  # Only a failed result says why the provider stopped; a result that
53
90
  # passed validation keeps finish_reason in its trace alone.
54
91
  def validation_errors(validation_result, response)
55
92
  errors = validation_result[:errors]
56
93
  return errors if validation_result[:status] == :ok
57
94
 
58
- reason = response.respond_to?(:finish_reason) ? response.finish_reason : nil
59
- message = FINISH_REASON_ERRORS[reason]
60
- return errors unless message
61
-
62
- message += "; configured max_output: #{@max_output}" if reason == :max_tokens && @max_output
63
- [*errors, message]
95
+ message = finish_reason_error(finish_reason(response))
96
+ message ? [*errors, message] : errors
64
97
  end
65
98
 
66
99
  # Custom adapters may return any object with `content` and `usage`.
@@ -7,7 +7,7 @@ module RubyLLM
7
7
  include Concerns::UsageAggregator
8
8
 
9
9
  # Copied from each attempt's trace into its attempts entry when present.
10
- ATTEMPT_TRACE_FIELDS = %i[usage latency_ms cost finish_reason].freeze
10
+ ATTEMPT_TRACE_FIELDS = %i[usage latency_ms cost finish_reason error_class].freeze
11
11
 
12
12
  private
13
13
 
@@ -4,7 +4,7 @@ module RubyLLM
4
4
  module Contract
5
5
  module Step
6
6
  class RetryPolicy
7
- attr_reader :max_attempts, :retryable_statuses
7
+ attr_reader :max_attempts, :retryable_statuses, :retryable_errors
8
8
 
9
9
  # Breaking (0.7.0): :adapter_error removed from defaults. ruby_llm's Faraday
10
10
  # middleware already retries transport errors (RateLimitError, ServerError,
@@ -18,6 +18,7 @@ module RubyLLM
18
18
  def initialize(models: nil, attempts: nil, retry_on: nil, &block)
19
19
  @configs = []
20
20
  @retryable_statuses = DEFAULT_RETRY_ON.dup
21
+ @retryable_errors = []
21
22
 
22
23
  if block
23
24
  @max_attempts = 1
@@ -49,12 +50,20 @@ module RubyLLM
49
50
  @configs
50
51
  end
51
52
 
52
- def retry_on(*statuses)
53
- @retryable_statuses = statuses.flatten
53
+ # Statuses, and exception classes for adapter errors:
54
+ # `retry_on :validation_failed, RubyLLM::RateLimitError` retries an
55
+ # adapter error only when it was a RateLimitError (or a subclass), while
56
+ # `:adapter_error` retries any. Replaces the defaults, as before.
57
+ def retry_on(*conditions)
58
+ @retryable_statuses, @retryable_errors = split_conditions(conditions.flatten)
54
59
  end
55
60
 
56
61
  def retryable?(result)
57
- retryable_statuses.include?(result.status)
62
+ return true if retryable_statuses.include?(result.status)
63
+ return false unless result.status == :adapter_error && retryable_errors.any?
64
+
65
+ error_type = result.trace.respond_to?(:error_type) ? result.trace.error_type : nil
66
+ !error_type.nil? && retryable_errors.any? { |klass| error_type <= klass }
58
67
  end
59
68
 
60
69
  def config_for_attempt(attempt, default_config)
@@ -78,7 +87,24 @@ module RubyLLM
78
87
  else
79
88
  @max_attempts = attempts || 1
80
89
  end
81
- @retryable_statuses = Array(retry_on).dup if retry_on
90
+ @retryable_statuses, @retryable_errors = split_conditions(Array(retry_on)) if retry_on
91
+ end
92
+
93
+ # Symbols are statuses; exception classes match adapter errors. Anything
94
+ # else (a String status among them) would never match, so it raises
95
+ # instead of quietly turning retries off.
96
+ def split_conditions(conditions)
97
+ invalid = conditions.reject { |condition| condition.is_a?(Symbol) || exception_class?(condition) }
98
+ unless invalid.empty?
99
+ raise ArgumentError,
100
+ "retry_on takes status symbols and exception classes, got #{invalid.map(&:inspect).join(", ")}"
101
+ end
102
+
103
+ conditions.partition { |condition| condition.is_a?(Symbol) }
104
+ end
105
+
106
+ def exception_class?(condition)
107
+ condition.is_a?(Class) && condition <= Exception
82
108
  end
83
109
 
84
110
  def normalize_config(entry)
@@ -60,7 +60,8 @@ module RubyLLM
60
60
  output_schema: @config.output_schema,
61
61
  model: @config.model,
62
62
  observers: @config.observers,
63
- max_output: @config.effective_max_output
63
+ max_output: @config.effective_max_output,
64
+ on_incomplete_output: @config.on_incomplete_output
64
65
  )
65
66
  end
66
67
 
@@ -84,6 +85,18 @@ module RubyLLM
84
85
  @config.attachment_token_estimate
85
86
  end
86
87
 
88
+ def exact_token_count?
89
+ @config.token_count == :exact
90
+ end
91
+
92
+ def count_adapter
93
+ @config.adapter
94
+ end
95
+
96
+ def count_options
97
+ @config.adapter_options
98
+ end
99
+
87
100
  def on_unknown_attachment_size
88
101
  @config.on_unknown_attachment_size
89
102
  end
@@ -19,7 +19,9 @@ module RubyLLM
19
19
  :on_unknown_attachment_size,
20
20
  :temperature,
21
21
  :extra_options,
22
- :observers
22
+ :observers,
23
+ :on_incomplete_output,
24
+ :token_count
23
25
  ) do
24
26
  # Factory with sensible defaults for optional fields. Lets callers
25
27
  # (Step::Base#run_once and tests) construct a RunnerConfig without
@@ -30,7 +32,8 @@ module RubyLLM
30
32
  output_schema: nil, max_output: nil,
31
33
  max_input: nil, max_cost: nil, on_unknown_pricing: UnknownPolicy::DEFAULT,
32
34
  attachment_token_estimate: nil, on_unknown_attachment_size: UnknownPolicy::DEFAULT,
33
- temperature: nil, extra_options: {}, observers: [])
35
+ temperature: nil, extra_options: {}, observers: [],
36
+ on_incomplete_output: Dsl::INCOMPLETE_OUTPUT_DEFAULT, token_count: Dsl::TOKEN_COUNT_DEFAULT)
34
37
  new(
35
38
  input_type: input_type, output_type: output_type,
36
39
  prompt_block: prompt_block, contract_definition: contract_definition,
@@ -41,7 +44,7 @@ module RubyLLM
41
44
  attachment_token_estimate: attachment_token_estimate,
42
45
  on_unknown_attachment_size: on_unknown_attachment_size,
43
46
  temperature: temperature, extra_options: extra_options,
44
- observers: observers
47
+ observers: observers, on_incomplete_output: on_incomplete_output, token_count: token_count
45
48
  )
46
49
  end
47
50
 
@@ -8,7 +8,7 @@ module RubyLLM
8
8
  include Concerns::DeepFreeze
9
9
 
10
10
  attr_reader :messages, :model, :latency_ms, :usage, :attempts, :cost,
11
- :usage_complete, :cost_complete, :finish_reason
11
+ :usage_complete, :cost_complete, :finish_reason, :error_class, :error_type
12
12
 
13
13
  # Marks `cost:` as not given, as distinct from an explicit nil. Only an
14
14
  # omitted cost is priced from the registry; nil from an adapter that
@@ -18,8 +18,12 @@ module RubyLLM
18
18
  # usage_complete / cost_complete are nil for traces that carry no
19
19
  # accounting state (Test adapter, older adapters, hand-built traces);
20
20
  # those keep the token-count rule in `unpriced?`.
21
+ # error_class names the exception behind an :adapter_error and is part
22
+ # of to_h; error_type is the class itself, kept in memory only so
23
+ # `retry_on SomeError` can match subclasses. A trace rebuilt from a
24
+ # hash has the name but not the class.
21
25
  def initialize(messages: nil, model: nil, latency_ms: nil, usage: nil, attempts: nil, cost: COST_UNSET,
22
- usage_complete: nil, cost_complete: nil, finish_reason: nil)
26
+ usage_complete: nil, cost_complete: nil, finish_reason: nil, error_class: nil, error_type: nil)
23
27
  @messages = deep_dup_freeze(messages)
24
28
  @model = model.frozen? ? model : model&.dup&.freeze
25
29
  @latency_ms = latency_ms
@@ -28,13 +32,15 @@ module RubyLLM
28
32
  @usage_complete = usage_complete
29
33
  @cost_complete = cost_complete
30
34
  @finish_reason = finish_reason
35
+ @error_class = error_class
36
+ @error_type = error_type
31
37
  @cost = resolve_cost(cost)
32
38
  @keep_nil_cost = keep_nil_cost?(cost)
33
39
  freeze
34
40
  end
35
41
 
36
42
  KNOWN_KEYS = %i[messages model latency_ms usage attempts cost
37
- usage_complete cost_complete finish_reason].freeze
43
+ usage_complete cost_complete finish_reason error_class].freeze
38
44
 
39
45
  def [](key)
40
46
  return nil unless KNOWN_KEYS.include?(key.to_sym)
@@ -57,7 +63,8 @@ module RubyLLM
57
63
  # Never re-prices: a retry merges a subtotal and an aggregate usage that
58
64
  # no single model's pricing describes.
59
65
  def merge(**overrides)
60
- self.class.new(**KNOWN_KEYS.to_h { |key| [key, overrides.fetch(key) { public_send(key) }] })
66
+ self.class.new(**KNOWN_KEYS.to_h { |key| [key, overrides.fetch(key) { public_send(key) }] },
67
+ error_type: overrides.fetch(:error_type, @error_type))
61
68
  end
62
69
 
63
70
  # True when a total built from this trace undercounts: the adapter said
@@ -87,7 +94,7 @@ module RubyLLM
87
94
  def to_h
88
95
  hash = { messages: @messages, model: @model, latency_ms: @latency_ms,
89
96
  usage: @usage, attempts: @attempts, cost: @cost, usage_complete: @usage_complete,
90
- cost_complete: @cost_complete, finish_reason: @finish_reason }.compact
97
+ cost_complete: @cost_complete, finish_reason: @finish_reason, error_class: @error_class }.compact
91
98
  hash[:cost] = nil if @keep_nil_cost
92
99
  hash
93
100
  end
@@ -2,6 +2,6 @@
2
2
 
3
3
  module RubyLLM
4
4
  module Contract
5
- VERSION = "1.1.2"
5
+ VERSION = "1.2.0"
6
6
  end
7
7
  end
@@ -0,0 +1,48 @@
1
+ # frozen_string_literal: true
2
+
3
+ module RubyLLM
4
+ module Contract
5
+ # Groups a step's or pipeline's RubyLLM calls under `RubyLLM.workflow`, so
6
+ # OpenTelemetry spans and instrumentation events carry the step's name.
7
+ # Off unless `RubyLLM::Contract.configure { |c| c.workflow_instrumentation = true }`.
8
+ #
9
+ # A pipeline opens one workflow and runs each step as `workflow.step(alias)`;
10
+ # a step run inside it does not open another. A step run on its own opens a
11
+ # workflow named after its class. A workflow the application opened itself
12
+ # becomes the parent, as RubyLLM links nested workflows.
13
+ module WorkflowScope
14
+ CURRENT_KEY = :ruby_llm_contract_workflow
15
+
16
+ def self.workflow(name, &block)
17
+ return yield unless enabled? && current.nil?
18
+
19
+ ::RubyLLM.workflow(name.to_s) { |workflow| within(workflow, &block) }
20
+ end
21
+
22
+ def self.step(name, &)
23
+ workflow = current
24
+ return yield unless enabled? && workflow
25
+
26
+ workflow.step(name.to_s, &)
27
+ end
28
+
29
+ def self.enabled?
30
+ Contract.configuration.workflow_instrumentation
31
+ end
32
+
33
+ # Fiber-local, so concurrent eval threads each get their own.
34
+ def self.current
35
+ Thread.current[CURRENT_KEY]
36
+ end
37
+
38
+ def self.within(workflow)
39
+ previous = Thread.current[CURRENT_KEY]
40
+ Thread.current[CURRENT_KEY] = workflow
41
+ yield
42
+ ensure
43
+ Thread.current[CURRENT_KEY] = previous
44
+ end
45
+ private_class_method :within
46
+ end
47
+ end
48
+ end
@@ -3,6 +3,8 @@
3
3
  require_relative "contract/version"
4
4
  require_relative "contract/errors"
5
5
  require_relative "contract/unknown_policy"
6
+ require_relative "contract/provider_options"
7
+ require_relative "contract/workflow_scope"
6
8
  require_relative "contract/types"
7
9
 
8
10
  module RubyLLM
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: ruby_llm-contract
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.1.2
4
+ version: 1.2.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Justyna
@@ -170,6 +170,7 @@ files:
170
170
  - lib/ruby_llm/contract/prompt/nodes/system_node.rb
171
171
  - lib/ruby_llm/contract/prompt/nodes/user_node.rb
172
172
  - lib/ruby_llm/contract/prompt/renderer.rb
173
+ - lib/ruby_llm/contract/provider_options.rb
173
174
  - lib/ruby_llm/contract/railtie.rb
174
175
  - lib/ruby_llm/contract/rake_task.rb
175
176
  - lib/ruby_llm/contract/rake_task/suite_gate.rb
@@ -195,6 +196,7 @@ files:
195
196
  - lib/ruby_llm/contract/types.rb
196
197
  - lib/ruby_llm/contract/unknown_policy.rb
197
198
  - lib/ruby_llm/contract/version.rb
199
+ - lib/ruby_llm/contract/workflow_scope.rb
198
200
  - ruby_llm-contract.gemspec
199
201
  homepage: https://github.com/justi/ruby_llm-contract
200
202
  licenses: