ruby_llm-contract 1.1.1 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +103 -0
  3. data/README.md +18 -3
  4. data/docs/guide/getting_started.md +86 -1
  5. data/docs/guide/llm_judge.md +14 -0
  6. data/docs/guide/rails_integration.md +8 -0
  7. data/docs/guide/relation_to_agent.md +8 -2
  8. data/lib/ruby_llm/contract/adapters/response.rb +28 -2
  9. data/lib/ruby_llm/contract/adapters/ruby_llm.rb +136 -28
  10. data/lib/ruby_llm/contract/adapters/test.rb +6 -0
  11. data/lib/ruby_llm/contract/concerns/usage_aggregator.rb +8 -5
  12. data/lib/ruby_llm/contract/configuration.rb +5 -1
  13. data/lib/ruby_llm/contract/cost_calculator.rb +98 -44
  14. data/lib/ruby_llm/contract/eval/eval_history.rb +2 -1
  15. data/lib/ruby_llm/contract/eval/model_comparison.rb +11 -1
  16. data/lib/ruby_llm/contract/eval/recommender.rb +10 -3
  17. data/lib/ruby_llm/contract/eval/report_stats.rb +8 -3
  18. data/lib/ruby_llm/contract/eval/report_storage.rb +8 -2
  19. data/lib/ruby_llm/contract/pipeline/base.rb +3 -1
  20. data/lib/ruby_llm/contract/pipeline/runner.rb +11 -1
  21. data/lib/ruby_llm/contract/pipeline/trace.rb +10 -0
  22. data/lib/ruby_llm/contract/provider_options.rb +45 -0
  23. data/lib/ruby_llm/contract/step/adapter_caller.rb +2 -1
  24. data/lib/ruby_llm/contract/step/base.rb +29 -12
  25. data/lib/ruby_llm/contract/step/dsl.rb +54 -0
  26. data/lib/ruby_llm/contract/step/limit_checker.rb +48 -13
  27. data/lib/ruby_llm/contract/step/result_builder.rb +66 -4
  28. data/lib/ruby_llm/contract/step/retry_executor.rb +35 -4
  29. data/lib/ruby_llm/contract/step/retry_policy.rb +31 -5
  30. data/lib/ruby_llm/contract/step/runner.rb +20 -1
  31. data/lib/ruby_llm/contract/step/runner_config.rb +6 -3
  32. data/lib/ruby_llm/contract/step/trace.rb +61 -20
  33. data/lib/ruby_llm/contract/token_estimator.rb +5 -3
  34. data/lib/ruby_llm/contract/version.rb +1 -1
  35. data/lib/ruby_llm/contract/workflow_scope.rb +48 -0
  36. data/lib/ruby_llm/contract.rb +2 -0
  37. metadata +3 -1
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 390af91f8200f1098440872d22515fa6232888d13d21a149c0b11f0bf3c15d48
4
- data.tar.gz: d39d8a0c2e8914d08e4d996e6a3eb4d2c53613304cabcc6e077af2d9963ad4f5
3
+ metadata.gz: fead48c5516f8e0aa738d3ab1633ca2d07f7173aba6466a6d29e37634263b857
4
+ data.tar.gz: 4ae4c314cd6365c5f5de455f9dc5ddc0741abd26fa751804b22c5ffded0e6cb1
5
5
  SHA512:
6
- metadata.gz: 81e03a20720536700acfe5bc840720a84675ca738e248fbdb490dac1dc63479c93924923d16cb28433329b53d24b796a42332736d29ec6f4d42a41240dfd4367
7
- data.tar.gz: a433acb982c831c7ee1e2db5675135c4a42762d75c78e53708673b8ca50f5f165fe2cba9a088828b95738f24f29afb01e2033f694e92c69182662a6307492aa4
6
+ metadata.gz: 9b66df1e490e238fdacc84a728b874a008b80c8b44cdd882b444a84389bb77996ee6132efb6adaad79c91a04bee30eb7351241a4989ff1a4ed5df8d7a6cff619
7
+ data.tar.gz: a055fd3d06ee4fa9591fa2a6cfed6872142febd950fe3a455daa6e588f04e157a585701456f4815483ecc3accb803cf3bbcffa9308c6dfa5011c00f6a7143274
data/CHANGELOG.md CHANGED
@@ -1,5 +1,108 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.2.0 (2026-10-10)
4
+
5
+ Everything new is opt-in, except the temperature fix below.
6
+
7
+ ### Added
8
+
9
+ - **`token_count :exact`** on a step: `max_input` and `max_cost` measure the input with
10
+ the provider's count (RubyLLM's `chat.count_tokens`: OpenAI Responses, Anthropic,
11
+ Gemini, Vertex AI, Bedrock Converse) instead of the chars/4 heuristic, attachments
12
+ included, so `attachment_token_estimate` is not needed. One extra request per attempt.
13
+ A provider without the endpoint, a failed count or an adapter without `count_tokens`
14
+ refuses the call (`:limit_exceeded`) rather than falling back to the heuristic; the
15
+ `Test` adapter counts with the heuristic. What a provider counts beyond prompt and
16
+ attachments is its own (Gemini's count leaves the output schema out).
17
+ `Adapters::RubyLLM#count_tokens`.
18
+ - **`on_incomplete_output :refuse`** on a step: a response the provider stopped at a token
19
+ limit fails as `:output_truncated`, one stopped by a content filter as
20
+ `:content_filtered`, before validation, so it is never `:ok`. Both statuses are outside
21
+ the default `retry_on`. The default, `:accept`, keeps 1.1.2's behaviour.
22
+ - **`retry_on` takes exception classes**: `retry_on :validation_failed,
23
+ RubyLLM::RateLimitError` retries an adapter error only of that class or a subclass;
24
+ `:adapter_error` still retries any. An adapter error's trace (and attempt) has
25
+ `error_class`.
26
+ - **`provider_options`** on a step (inherited, `:default` resets) and per call
27
+ (`context: { provider_options: }`, merged over the step's), passed to RubyLLM's
28
+ `with_provider_options`. Keys the step controls raise `ArgumentError`.
29
+ - **`workflow_instrumentation`** in `RubyLLM::Contract.configure`: a pipeline runs as one
30
+ `RubyLLM.workflow` with a `workflow.step` per alias, a step on its own as a workflow
31
+ named after its class, so RubyLLM's events and OpenTelemetry spans carry the step.
32
+ - **`register_model` takes cache prices**: `cache_read_per_1m:` and `cache_write_per_1m:`,
33
+ both optional.
34
+ - Guides: retry by error class, cut-off responses, exact token counts, local and custom
35
+ models (Ollama), provider options, workflow instrumentation, and `RubyLLM::Judge` (2.1)
36
+ in the LLM-judge guide. The README lists the new options and answers how to run local
37
+ models.
38
+
39
+ ### Changed
40
+
41
+ - **A temperature is no longer sent to a model RubyLLM's registry marks as taking none**
42
+ (OpenAI's o-series and gpt-5 family). RubyLLM 1.x quietly sent 1.0 to those models;
43
+ 2.x sends the value as given, so a step with `temperature` that escalated to one failed
44
+ with an adapter error. The adapter now leaves it out and warns once per model; a model
45
+ the registry does not know still gets it.
46
+ - **`retry_on` raises on a condition it cannot match.** A status given as a String (or
47
+ any value that is neither a Symbol nor an exception class) never matched a result, so
48
+ it silently turned retries off; it now raises `ArgumentError` when the step is defined.
49
+
50
+ ## 1.1.2 (2026-10-10)
51
+
52
+ ### Changed - costs and gates may move
53
+
54
+ - **Costs now count the prompt cache, and may go up.** RubyLLM 2.x reports cached prompt
55
+ tokens apart from `tokens.input`, and the adapter read only input and output: 1,000
56
+ uncached + 9,000 cached + 200 output tokens on gpt-4.1-mini came out at $0.00072
57
+ instead of $0.00162. `trace[:usage][:input_tokens]` is now the whole prompt and
58
+ `trace[:cost]` is the cost RubyLLM puts on the response: the price of the provider the
59
+ call went to, cache and long-context prices, the amount the provider reported when it
60
+ reports one, and RubyLLM's own retries of the request. Eval totals, history and
61
+ `compare_models` figures can rise against older baselines; baselines compare scores,
62
+ not costs, so they do not regress on this.
63
+ - **A call without token counts is no longer free.** When the provider sends no usage
64
+ (for the call or for any of RubyLLM's attempts at it), the counts still read 0 but
65
+ `trace[:usage_complete]` is `false` and, unless the provider reported what it billed,
66
+ the cost is unknown, so the eval cost gates refuse (opt out with
67
+ `on_unknown_pricing: :warn`).
68
+ - **A missing price is unknown, not $0.** A model priced for input but not output has no
69
+ cost instead of a partial one, so `max_cost` refuses it pre-flight. A call that wrote
70
+ to the prompt cache on a model without a cache-write price has an unknown cost
71
+ afterwards (pre-flight estimates do not predict cache writes).
72
+
73
+ ### Fixed
74
+
75
+ - `CostCalculator` prices RubyLLM's models through `Model#cost_for`, so pre-flight
76
+ estimates use cache and long-context prices, and the price of the provider the call
77
+ goes to (12 model ids carry different prices per provider). `register_model` keeps
78
+ overriding the price for its id; cache reads are charged at its input price, cache
79
+ writes have no price.
80
+ - `compare_models` no longer names a model with unpriced cases the cheapest
81
+ (`best_for`), gives it an infinite `cost_per_point`, and `recommend` neither ranks it
82
+ on cost nor promises savings against it.
83
+ - `single_shot_cost` reported a later attempt's cost when the first attempt was unpriced.
84
+ - The step log printed an unknown cost as `$0.000000`; it now prints `cost=unknown`.
85
+ - Docs no longer say RubyLLM has no evaluation framework: `relation_to_agent.md` now
86
+ places `RubyLLM::Evaluation` (2.1) next to `define_eval`. The README and the getting
87
+ started guide describe what a trace's usage and cost count, and the README's 1.x
88
+ promise now says new trace keys may be added.
89
+
90
+ ### Added
91
+
92
+ - `trace[:usage]` breakdown keys when reported: `cache_read_tokens`, `cache_write_tokens`,
93
+ `thinking_tokens` (parts of `input_tokens` / `output_tokens`, not extra tokens).
94
+ - `trace[:finish_reason]`, also on each retry attempt. A failed result whose response
95
+ stopped at a token limit or a content filter gets a last validation error saying so,
96
+ with the `max_output` the call ran with. Statuses and retries are unchanged.
97
+ - An adapter may report its own cost: `Adapters::Response.new(content:, usage:, cost:,
98
+ cost_complete:, usage_complete:, finish_reason:)`, where `cost: nil` means unknown.
99
+ A response without `cost:` is priced from `usage` as before.
100
+ - `provider:` on `CostCalculator.calculate`, `CostCalculator.find_model`,
101
+ `estimate_cost` and `estimate_eval_cost`.
102
+ - Baselines and history mark unpriced cases and runs with `cost_unknown: true`.
103
+ - A pipeline with a `token_budget` warns when a step's provider left out token counts;
104
+ `Pipeline::Trace#usage_complete?`.
105
+
3
106
  ## 1.1.1 (2026-10-10)
4
107
 
5
108
  ### Fixed
data/README.md CHANGED
@@ -91,6 +91,8 @@ result.trace[:attempts]
91
91
  # ]
92
92
  ```
93
93
 
94
+ `usage[:input_tokens]` is the whole prompt, prompt-cache tokens included (broken out as `cache_read_tokens` / `cache_write_tokens` when the provider reports them), and `cost` is what RubyLLM prices the response at: the provider's price, cache and long-context rates, or the amount the provider reported. A cost the gem cannot fully price is marked unknown rather than counted as $0 - see [what a trace's usage and cost count](docs/guide/getting_started.md#what-a-traces-usage-and-cost-count).
95
+
94
96
  If the response is malformed, the TL;DR overflows the card, or the takeaway count is off, the gem moves to the next step. This is model **escalation**, not a fallback list — each step is an independent config (`model`, `reasoning_effort`), so the retry policy spends more compute only when the cheaper one couldn't satisfy the contract.
95
97
 
96
98
  ### Add a CI gate in 6 lines
@@ -128,6 +130,11 @@ Everything below is optional — the example above is a complete step. Reach for
128
130
  - **[Find the cheapest viable fallback list](docs/guide/optimizing_retry_policy.md)** — empirically pick the cheapest model chain that still passes your evals.
129
131
  - **[A/B test prompts](docs/guide/eval_first.md)** — measure whether a new prompt is safe to ship before merging.
130
132
  - **[Budget caps](docs/guide/getting_started.md)** — refuse the request pre-flight when an estimate exceeds the limit.
133
+ - **[Cost tracking](docs/guide/getting_started.md#what-a-traces-usage-and-cost-count)** - per-call cost with prompt-cache and per-provider prices; eval cost gates fail closed when a cost is unknown.
134
+ - **[Exact token counts](docs/guide/getting_started.md#counting-input-tokens-exactly)** - `token_count :exact` checks `max_input` / `max_cost` against the provider's own count instead of a heuristic.
135
+ - **[Cut-off responses](docs/guide/getting_started.md#responses-cut-off-by-a-token-limit)** - `on_incomplete_output :refuse` fails an answer the provider stopped at a token limit, even one that would validate.
136
+ - **[Retry by error class](docs/guide/getting_started.md#retrying-provider-errors-by-class)** - `retry_on RubyLLM::RateLimitError` retries a rate limit without retrying an exhausted balance.
137
+ - **[Provider options](docs/guide/getting_started.md#provider-options)** - send `service_tier`, `seed` and other provider settings from the step or per call.
131
138
  - **[Reasoning effort / thinking config](docs/guide/optimizing_retry_policy.md)** — Anthropic / OpenAI thinking configuration on the Step class.
132
139
 
133
140
  Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and per-step models.
@@ -136,7 +143,7 @@ Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and
136
143
 
137
144
  ## Relation to `RubyLLM::Agent`
138
145
 
139
- `Step::Base` and `RubyLLM::Agent` (since RubyLLM 1.12) are **siblings** targeting the same niche: reusable, class-based prompts. Both call into `RubyLLM::Chat` directly — Step does not wrap Agent. Step adds the contract layer: `validate` (business invariants), `retry_policy escalate(...)` (model escalation on validation failure), `max_cost` pre-flight refusal, regression-eval framework, pipeline composition. **[Full feature mapping →](docs/guide/relation_to_agent.md)**
146
+ `Step::Base` and `RubyLLM::Agent` (since RubyLLM 1.12) are **siblings** targeting the same niche: reusable, class-based prompts. Both call into `RubyLLM::Chat` directly - Step does not wrap Agent. Step adds the contract layer: `validate` (business invariants), `retry_policy escalate(...)` (model escalation on validation failure), `max_cost` pre-flight refusal, regression-eval framework, pipeline composition. RubyLLM 2.1's `RubyLLM::Evaluation` grades answers; `define_eval` guards against regressions and cost in CI - the mapping covers both. **[Full feature mapping →](docs/guide/relation_to_agent.md)**
140
147
 
141
148
  ## Relation to `ruby_llm-tribunal`
142
149
 
@@ -153,7 +160,7 @@ Different layers, complementary. [`ruby_llm-tribunal`](https://github.com/Alqemi
153
160
  | [Why contracts?](docs/guide/why.md) | Recognise the four production failures the gem exists for |
154
161
  | [Relation to RubyLLM::Agent](docs/guide/relation_to_agent.md) | Sibling abstractions; what each adds; runtime call path; coexistence patterns |
155
162
  | [Relation to ruby_llm-tribunal](docs/guide/relation_to_tribunal.md) | Different layers (test framework vs runtime contract); visual flows; integration recipes |
156
- | [Getting Started](docs/guide/getting_started.md) | Walk the full feature set on one concrete step |
163
+ | [Getting Started](docs/guide/getting_started.md) | Walk the full feature set on one concrete step, including what usage and cost count |
157
164
  | [Rails integration](docs/guide/rails_integration.md) | Directory, initializer, jobs, logging, specs, CI gate — 7 FAQs for Rails devs |
158
165
  | [Adopt in an existing Rails app](docs/guide/migration.md) | Replace raw `LlmClient.call` with a contract, Before/After |
159
166
  | [Prevent silent prompt regressions](docs/guide/eval_first.md) | Evals, baselines, CI gates that block quality drift |
@@ -173,7 +180,11 @@ Stable since **1.0.0**. Semver tracked; breaking changes flagged in [CHANGELOG](
173
180
  What 1.0 promises not to break inside the 1.x line: the documented DSL
174
181
  (`prompt`, `output_schema`, `validate`, `retry_policy`, `max_cost`, `define_eval`,
175
182
  pipelines), the `Result` and `trace` shapes including `trace[:usage]`'s
176
- `{ input_tokens:, output_tokens: }` keys, and the adapter interface. A new major
183
+ `{ input_tokens:, output_tokens: }` keys, and the adapter interface. New keys may be
184
+ added to both (1.1.2 added `cache_read_tokens`, `cache_write_tokens` and
185
+ `thinking_tokens` to `trace[:usage]`, and `usage_complete`, `cost_complete` and
186
+ `finish_reason` to the trace), and a figure may change when it was wrong: since 1.1.2
187
+ `input_tokens` counts prompt-cache tokens. A new major
177
188
  version of a runtime dependency may require a new major version here — that is
178
189
  what 1.0.0 itself was.
179
190
 
@@ -185,6 +196,10 @@ what 1.0.0 itself was.
185
196
 
186
197
  **Where in a Rails app?** Default `app/contracts/`. The Railtie reloads `app/contracts/eval/` and `app/steps/eval/` in development; any autoloaded directory also works. See [Rails integration](docs/guide/rails_integration.md).
187
198
 
199
+ **Costs went up, or an eval cost gate turned red, after upgrading to 1.1.2?** Earlier versions left prompt-cache tokens out of `input_tokens` and the cost, priced every model at its default provider, counted a missing price as $0, and treated a call without token counts as free. The [CHANGELOG](CHANGELOG.md) lists each change; `on_unknown_pricing: :warn` turns the unknown-cost refusal into a warning.
200
+
201
+ **Local models (Ollama)?** Point RubyLLM at the server (`c.ollama_api_base = "http://localhost:11434/v1"`) and run steps with `context: { provider: :ollama, model: "...", assume_model_exists: true }`. Local models have no price in RubyLLM's registry, so register one (`CostCalculator.register_model("gemma-fast:latest", input_per_1m: 0, output_per_1m: 0)`) or the eval cost gates treat their cost as unknown. Ollama has no token-counting endpoint, so leave `token_count` at `:estimate`. See [local and custom models](docs/guide/getting_started.md#local-and-custom-models).
202
+
188
203
  **Upgraded from pre-0.10.0 and getting `:limit_exceeded` with attachments?** Multimodal contracts with `max_cost`/`max_input` need `attachment_token_estimate`. See [multimodal input guide](docs/guide/multimodal_input.md#cost-attachment_token_estimate-is-required) for setup, fail-closed behaviour, and `on_unknown_attachment_size :warn` opt-out.
189
204
 
190
205
  ## License
@@ -65,6 +65,32 @@ result.trace[:attempts]
65
65
 
66
66
  If the whole chain exhausts, `result.status` is the status of the last attempt (`:validation_failed` or `:parse_error`) and `result.parsed_output` is the last attempt's output. The caller decides what to do — ship it anyway, fall back to a template, or raise.
67
67
 
68
+ ### Retrying provider errors by class
69
+
70
+ A provider error that RubyLLM's own retries could not get past ends as `:adapter_error`, with the exception class in `trace[:error_class]` (e.g. `"RubyLLM::RateLimitError"`). Adapter errors are not retried by default. List exception classes in `retry_on` to retry only those, typically together with `escalate` so the next attempt goes to another model:
71
+
72
+ ```ruby
73
+ retry_policy do
74
+ escalate "gpt-4.1-mini", "claude-haiku-4-5"
75
+ retry_on :validation_failed, :parse_error, RubyLLM::RateLimitError, RubyLLM::OverloadedError
76
+ end
77
+ ```
78
+
79
+ A class matches its subclasses; `:adapter_error` in the list retries every adapter error. As with statuses, the list replaces the defaults. Switching to a fallback model on transport errors inside one call is RubyLLM's `chat.with_fallbacks`; the contract's adapter does not call it, so fallbacks across models belong in `escalate`.
80
+
81
+ ### Responses cut off by a token limit
82
+
83
+ When a provider stops at a token limit or a content filter, `trace[:finish_reason]` says so (`:max_tokens`, `:content_filter`), and a failed result gets a last validation error naming it. A cut-off answer that still parses and validates stays `:ok` unless the step refuses it:
84
+
85
+ ```ruby
86
+ class SummarizeArticle < RubyLLM::Contract::Step::Base
87
+ max_output 400
88
+ on_incomplete_output :refuse # default :accept
89
+ end
90
+ ```
91
+
92
+ With `:refuse`, such a response fails before validation as `:output_truncated` (token limit) or `:content_filtered`, so `validate` blocks and observers never see it. Raw output, usage and cost stay on the result. Neither status is retried by default; `retry_on :output_truncated` opts in (useful when a later attempt has a larger `max_output`). Anthropic reports an exhausted context window as a token-limit stop too, so `:output_truncated` means "stopped on tokens", not necessarily "hit max_output".
93
+
68
94
  ### Per-attempt reasoning effort
69
95
 
70
96
  `models:` accepts config hashes as well as model-name strings, so a fallback can "try harder" (more reasoning) on retry, not just switch model:
@@ -173,7 +199,48 @@ max_cost 0.01, on_unknown_pricing: :warn
173
199
 
174
200
  Default is `:refuse`. Use `:warn` only when you accept running without a cost ceiling (fine-tuned models you trust, private endpoints).
175
201
 
176
- The eval-level gates - `with_maximum_cost` on `pass_eval`, `maximum_cost` on the rake task, and `maximum_cost:` on `assert_eval_passes` - follow the same rule since 1.1.0. A report's total counts a case it cannot price as $0, so when any case ran on a model without pricing data the gate fails and names the cases instead of comparing an undercounted total. Opt out with `.on_unknown_pricing(:warn)`, `t.on_unknown_pricing = :warn` or `on_unknown_pricing: :warn`, which checks the budget against the priced cases and prints a warning. Offline runs (`sample_response`, a `Test` adapter without `usage:`) report zero tokens, cost $0 with or without pricing, and are never flagged. The check sees only reported tokens: an adapter that returns no token counts looks like a zero-token run.
202
+ The eval-level gates - `with_maximum_cost` on `pass_eval`, `maximum_cost` on the rake task, and `maximum_cost:` on `assert_eval_passes` - follow the same rule since 1.1.0. A report's total counts a case it cannot price as $0, so when any case ran on a model without pricing data the gate fails and names the cases instead of comparing an undercounted total. Opt out with `.on_unknown_pricing(:warn)`, `t.on_unknown_pricing = :warn` or `on_unknown_pricing: :warn`, which checks the budget against the priced cases and prints a warning. Offline runs (`sample_response`, a `Test` adapter without `usage:`) report zero tokens, cost $0 with or without pricing, and are never flagged. Since 1.1.2 a real call whose provider sent no token counts is flagged too, instead of passing as a free zero-token run.
203
+
204
+ ### What a trace's usage and cost count
205
+
206
+ `trace[:usage][:input_tokens]` is the whole prompt the model took in, including tokens served from or written to the provider's prompt cache. When the provider reports them, the trace adds a breakdown: `cache_read_tokens`, `cache_write_tokens` and `thinking_tokens`. These are parts of `input_tokens` / `output_tokens`, not extra tokens; reasoning tokens are already in `output_tokens` when the provider bills them as output, and a model with a separate reasoning price is charged that price for them.
207
+
208
+ ```ruby
209
+ result.trace[:usage]
210
+ # => { input_tokens: 10_000, output_tokens: 200, cache_read_tokens: 9_000 }
211
+ result.trace[:cost] # => 0.00162 (gpt-4.1-mini: cache reads at the cache price)
212
+ ```
213
+
214
+ With the RubyLLM adapter, `trace[:cost]` is the cost RubyLLM puts on the response (`response.cost.total`): the price of the provider the call went to, cache and long-context prices, and the amount the provider reported when it reports one (OpenRouter, for example). A price set with `CostCalculator.register_model` still overrides it for that model id, unless the provider reported what it billed; there, cache reads are charged at the input price and a call that wrote to the cache has no price, since cache writes can cost more than input.
215
+
216
+ When a cost does not cover the whole call, `result.trace.cost_unknown?` is true and the eval gates above refuse. That happens when pricing is missing for a token category the call used, or when the provider sent no token counts for the call or for one of RubyLLM's attempts at it (`trace[:usage_complete]` is then `false` and missing counts read as 0) and did not report what it billed either. On a retried step `trace[:cost]` is the sum of the attempts that could be priced, and `cost_unknown?` is true if any attempt could not. `max_cost` itself is a pre-flight check on an estimate; it does not inspect the finished call.
217
+
218
+ A custom adapter returning `Response.new(content:, usage:)` is priced from `usage` as before. To report its own cost, pass `cost:` - `nil` means unknown - with `cost_complete:` and `usage_complete:`.
219
+
220
+ ### Counting input tokens exactly
221
+
222
+ `max_input` and `max_cost` measure the input with a chars/4 heuristic (±30%). With `token_count :exact` they ask the provider instead:
223
+
224
+ ```ruby
225
+ class SummarizeArticle < RubyLLM::Contract::Step::Base
226
+ token_count :exact # default :estimate
227
+ max_input 20_000
228
+ end
229
+ ```
230
+
231
+ The count comes from RubyLLM's `chat.count_tokens` (OpenAI Responses, Anthropic, Gemini, Vertex AI, Bedrock Converse) and covers the prompt and attachments, so `attachment_token_estimate` is not needed. What else a provider counts is its own: OpenAI's and Anthropic's counts include the output schema, Gemini's leaves it out. It costs one extra request per attempt, and RubyLLM leaves `provider_options` out of it. Where no count can be had - a provider without the endpoint (Ollama, for one), a failed request, or an adapter without `count_tokens` - the call is refused as `:limit_exceeded` with the reason, never measured with the heuristic under the name "exact". The `Test` adapter counts with the heuristic, since it has no provider to ask. The output side of `max_cost` is still an estimate (`max_output`, or the input size without one).
232
+
233
+ ### Local and custom models
234
+
235
+ A local model (Ollama) or a fine-tuned one has no price in RubyLLM's registry, so its cost is unknown and the eval cost gates refuse. Register a price - zero for a model that costs nothing per token:
236
+
237
+ ```ruby
238
+ RubyLLM::Contract::CostCalculator.register_model("gemma-fast:latest", input_per_1m: 0, output_per_1m: 0)
239
+ RubyLLM::Contract::CostCalculator.register_model("ft:gpt-4o-custom",
240
+ input_per_1m: 3.0, output_per_1m: 6.0, cache_read_per_1m: 1.5, cache_write_per_1m: 3.75)
241
+ ```
242
+
243
+ `register_model` sets a price for a model id, not a model RubyLLM can call: the step still needs `provider:` (and `assume_model_exists: true` for an id RubyLLM does not list). Cache prices are optional; without them cache reads are charged at the input price and a call that wrote to the cache has no price. A price does not make up for missing counts: a provider that sends no token counts still leaves the cost unknown.
177
244
 
178
245
  ### Preflight cost estimates
179
246
 
@@ -195,6 +262,24 @@ SummarizeArticle.estimate_eval_cost("regression",
195
262
 
196
263
  `estimate_cost` returns `nil` when pricing isn't registered. `estimate_eval_cost` silently treats unknown-pricing cases as `$0.00` and sums the rest — it does **not** fail closed the way `max_cost` does. Treat its output as a floor, not a guarantee; register pricing via `CostCalculator.register_model` before relying on it for budget decisions.
197
264
 
265
+ ## Provider options
266
+
267
+ Settings RubyLLM has no method for go to the provider as-is through `provider_options`, on the step or per call:
268
+
269
+ ```ruby
270
+ class SummarizeArticle < RubyLLM::Contract::Step::Base
271
+ provider_options service_tier: "flex"
272
+ end
273
+
274
+ SummarizeArticle.run(text, context: { provider_options: { service_tier: "priority", seed: 7 } })
275
+ ```
276
+
277
+ The call's keys win over the step's; a subclass replaces its parent's hash, and `provider_options :default` stops inheriting it. Keys the step itself sets (`model`, `temperature`, the max-token keys, the messages, and on a step with `output_schema` the structured-output keys `response_format` and `text`) raise `ArgumentError`, so a call cannot differ from what its limits and schema were checked against. Only top-level keys are checked. `compare_models` and `optimize_retry_policy` candidates do not carry provider options.
278
+
279
+ ## Temperature on reasoning models
280
+
281
+ RubyLLM 1.x sent temperature 1.0 to OpenAI reasoning models (o-series, gpt-5), which reject other values; RubyLLM 2.x sends what it is given. A step with `temperature` therefore leaves it out for a model RubyLLM's registry marks as taking none, and warns once per model - so `temperature 0` with `escalate "gpt-4.1-mini", "gpt-5-mini"` still works. A model the registry does not know gets the temperature as set.
282
+
198
283
  ## `output_schema` vs `with_schema`
199
284
 
200
285
  `with_schema` in `ruby_llm` tells the provider to force a specific JSON structure. `output_schema` in this gem does the same thing (calls `with_schema` under the hood) **plus** validates the response client-side. Cheaper models sometimes ignore schema constraints — `with_schema` is a request; `output_schema` is a request plus verification.
@@ -195,6 +195,20 @@ What to do operationally:
195
195
 
196
196
  The drop is your signal to refine the judge's prompt, not to lower the gate.
197
197
 
198
+ ## `RubyLLM::Judge` (RubyLLM 2.1)
199
+
200
+ RubyLLM 2.1 adds `RubyLLM::Judge` for typed questions - `probability`, `choice` and `score` - answered by **decision models**: OpenAI's judgment models, TypeSafe, or local decision models through Ollama. It does not run on ordinary chat models; those raise that they do not support judgments. Where you have such a model, a judge class can stand in for the judge Step in step 1 above:
201
+
202
+ ```ruby
203
+ class SummaryFaithfulness < RubyLLM::Judge
204
+ probability :faithful, "Is every claim in the summary supported by the article?"
205
+ end
206
+
207
+ SummaryFaithfulness.judge(article_and_summary).faithful.probability # => 0.93
208
+ ```
209
+
210
+ Calibrate it against human labels exactly as in step 2; the threshold is still yours to set. On an ordinary chat model, keep the judge as a Contract Step - the pattern this guide builds.
211
+
198
212
  ## When to reach for Tribunal instead
199
213
 
200
214
  [`ruby_llm-tribunal`](https://github.com/Alqemist-labs/ruby_llm-tribunal) ships an off-the-shelf catalog of common LLM-as-judge assertions (`assert_faithful`, `assert_hallucination`, `assert_refusal`, `assert_no_pii`, etc.) — a shortcut when your check matches one of those domain-general categories. The methodology in this guide still applies: calibrate the judge against your human-labeled production data **before** trusting Tribunal's `default_threshold = 0.8`, refine the prompt when it over-flags, watch for the anti-patterns above. See [Relation to Tribunal](relation_to_tribunal.md) for the full positioning — what each gem documents (and doesn't), a concrete decision tree on catalog-vs-custom-judge, and three working integration patterns.
@@ -125,6 +125,14 @@ end
125
125
 
126
126
  Trace inspection in an admin UI: `result.trace[:attempts]` gives you per-attempt model, status, cost, latency — render it in a partial to debug production failures without re-running.
127
127
 
128
+ RubyLLM 2.1 emits its own events and OpenTelemetry spans for every call. To group them by step, turn on workflow instrumentation in the initializer:
129
+
130
+ ```ruby
131
+ RubyLLM::Contract.configure { |c| c.workflow_instrumentation = true }
132
+ ```
133
+
134
+ A pipeline then runs as one `RubyLLM.workflow` named after its class, with each step as `workflow.step(alias)`; a step run on its own is a workflow named after its class. RubyLLM's events carry `workflow_name` and `workflow_step_name`, and a workflow your code opened around the call becomes the parent. Off by default. In a Rails app RubyLLM also records each call in its usage ledger when configured; that figure and `trace[:cost]` are the same RubyLLM price, unless `CostCalculator.register_model` overrides it for the model.
135
+
128
136
  ## 5. Testing — RSpec and Minitest
129
137
 
130
138
  Add to `spec/spec_helper.rb` (or `test_helper.rb`):
@@ -12,8 +12,8 @@
12
12
  | `validate :rule do ... end` business invariants on output | only in `ruby_llm-contract` |
13
13
  | `retry_policy escalate(...)` model escalation on validation failure | only here (different from RubyLLM's network-level retry) |
14
14
  | `max_cost` / `max_input` / `max_output` pre-flight refusal | only here |
15
- | `define_eval` + baseline regression + `compare_models` + `optimize_retry_policy` | only here (RubyLLM does not ship an evaluation framework) |
16
- | Pipeline composition with `step SomeStep, as: :alias` | only here (RubyLLM intentionally leaves workflows as plain Ruby) |
15
+ | `define_eval` + baseline regression + `compare_models` + `optimize_retry_policy` | only here; RubyLLM 2.1's `RubyLLM::Evaluation` grades answers (see below) but keeps no baseline and has no cost gate |
16
+ | Pipeline composition with `step SomeStep, as: :alias` | only here; `RubyLLM.workflow` (2.1) groups calls for tracing, without typed step outputs, fail-fast or budgets |
17
17
  | `around_call`, named `observe` hooks with pass/fail recorded in trace | only here |
18
18
 
19
19
  ## Runtime relationship
@@ -32,6 +32,12 @@ Step.run(input)
32
32
 
33
33
  This may change in a future release if upstream APIs make a layered design natural. The decision is not committed; it depends on adopter signal.
34
34
 
35
+ ## Relation to `RubyLLM::Evaluation`
36
+
37
+ RubyLLM 2.1 ships `RubyLLM::Evaluation`: a dataset of cases run through your `perform` method, graded for correctness by an LLM or by your own criteria and Ruby assertions, with RSpec/Minitest integration and a report of verdicts, tokens and cost. Use it to answer "is this answer right?".
38
+
39
+ `define_eval` answers a different question: did this prompt or model get worse than last time, and does it stay within budget? It keeps baselines and flags regressions, compares prompts (`compare_with`) and models (`compare_models`, `recommend`), runs offline with `sample_response`, and gates CI on score and cost. The two can share a project; an Evaluation's `perform` can call `Step.run`.
40
+
35
41
  ## Coexistence on the same project
36
42
 
37
43
  The two abstractions can live in the same Rails (or non-Rails) project. Pick one per use case:
@@ -3,16 +3,42 @@
3
3
  module RubyLLM
4
4
  module Contract
5
5
  module Adapters
6
+ # `content` and `usage` are the whole interface an adapter has to fill.
7
+ # The rest is optional accounting: an adapter that leaves `cost` out gets
8
+ # its usage priced from the registry, as before; one that passes `cost`
9
+ # (nil included) owns the figure, and `cost_complete` / `usage_complete`
10
+ # say whether it and the token counts cover the whole call.
6
11
  class Response
7
12
  include Concerns::DeepFreeze
8
13
 
9
- attr_reader :content, :usage
14
+ attr_reader :content, :usage, :usage_complete, :cost_complete, :finish_reason
10
15
 
11
- def initialize(content:, usage: {})
16
+ def initialize(content:, usage: {}, cost: Step::Trace::COST_UNSET, usage_complete: nil,
17
+ cost_complete: nil, finish_reason: nil)
12
18
  @content = deep_dup_freeze(content)
13
19
  @usage = deep_dup_freeze(usage)
20
+ @cost = cost
21
+ @usage_complete = usage_complete
22
+ @cost_complete = cost_complete
23
+ @finish_reason = finish_reason
14
24
  freeze
15
25
  end
26
+
27
+ def cost
28
+ cost_provided? ? @cost : nil
29
+ end
30
+
31
+ def cost_provided?
32
+ !@cost.equal?(Step::Trace::COST_UNSET)
33
+ end
34
+
35
+ # Keyword arguments for Step::Trace.new. Omits what the adapter did not
36
+ # report, so a plain Response keeps the registry pricing path.
37
+ def trace_accounting
38
+ fields = { usage_complete: @usage_complete, cost_complete: @cost_complete,
39
+ finish_reason: @finish_reason }.compact
40
+ cost_provided? ? fields.merge(cost: @cost) : fields
41
+ end
16
42
  end
17
43
  end
18
44
  end
@@ -6,32 +6,59 @@ module RubyLLM
6
6
  module Contract
7
7
  module Adapters
8
8
  class RubyLLM < Base
9
+ # `with: nil` adds no attachment. The text is never nil (`fetch(:content, "")`),
10
+ # and RubyLLM raises only when both text and attachments are nil.
9
11
  def call(messages:, **options)
10
- system_contents, conversation = partition_messages(messages)
11
- conversation = fallback_conversation(system_contents, conversation)
12
-
13
- chat = build_chat(options, system_contents)
14
- add_history(chat, conversation[0..-2])
12
+ chat, text = prepared_chat(messages, options)
13
+ response = chat.ask(text, with: options[:attachment])
14
+ build_response(response, options[:model])
15
+ end
15
16
 
16
- # `with: nil` is a documented no-op in RubyLLM (verified against
17
- # 1.15.0: chat.rb:36-37 `build_content(message, nil)` -> content.rb:8-14
18
- # `Content.new(text, nil)` keeps text-only path when attachments
19
- # are empty; raise only fires when BOTH text and attachments are nil,
20
- # and we always pass a non-nil string thanks to `&.fetch(:content, "")`).
21
- response = chat.ask(
22
- conversation.last&.fetch(:content, ""),
23
- with: options[:attachment]
24
- )
25
- build_response(response)
17
+ # Input tokens of the request `call` would send, counted by the
18
+ # provider (`chat.count_tokens`), attachments included. `ask` is
19
+ # `ask_later` plus `complete` in RubyLLM, so the staged message is the
20
+ # one `call` sends. RubyLLM leaves provider_options out of the count.
21
+ # Raises RubyLLM::Error where the provider has no counting endpoint.
22
+ def count_tokens(messages:, **options)
23
+ chat, text = prepared_chat(messages, options)
24
+ chat.ask_later(text, with: options[:attachment]).count_tokens
25
+ rescue KeyError => e
26
+ # RubyLLM reads OpenAI's count with `fetch`; a reply without the
27
+ # field is a failed count, not a bug in the caller.
28
+ raise ::RubyLLM::Error, "count response without input_tokens (#{e.message})"
26
29
  end
27
30
 
28
31
  CHAT_OPTION_METHODS = {
29
- temperature: :with_temperature,
30
32
  schema: :with_schema
31
33
  }.freeze
32
34
 
35
+ @dropped_temperature_models = Set.new
36
+ @dropped_temperature_lock = Mutex.new
37
+
38
+ # Warns once per provider and model, across threads and subclasses
39
+ # (called on this class, whose state subclasses do not inherit).
40
+ def self.warn_dropped_temperature(model)
41
+ key = "#{model.provider}:#{model.id}"
42
+ first = @dropped_temperature_lock.synchronize { @dropped_temperature_models.add?(key) }
43
+ return unless first
44
+
45
+ warn "[ruby_llm-contract] #{model.id} (#{model.provider}) does not accept temperature " \
46
+ "according to RubyLLM's model registry; the step's temperature is not sent"
47
+ end
48
+
33
49
  private
34
50
 
51
+ # The configured chat with the history added, and the text of the last
52
+ # message, which sending and counting both stage themselves.
53
+ def prepared_chat(messages, options)
54
+ system_contents, conversation = partition_messages(messages)
55
+ conversation = fallback_conversation(system_contents, conversation)
56
+
57
+ chat = build_chat(options, system_contents)
58
+ add_history(chat, conversation[0..-2])
59
+ [chat, conversation.last&.fetch(:content, "")]
60
+ end
61
+
35
62
  # When prompt has only system/section/rule nodes and no user message,
36
63
  # pop the last system message and use it as the user ask.
37
64
  def fallback_conversation(system_contents, conversation)
@@ -59,13 +86,13 @@ module RubyLLM
59
86
  CHAT_OPTION_METHODS.each do |key, method_name|
60
87
  chat.public_send(method_name, options[key]) if options[key]
61
88
  end
89
+ apply_temperature(chat, options[:temperature]) unless options[:temperature].nil?
62
90
 
63
91
  # Resolve thinking config from BOTH sources, with `:reasoning_effort`
64
92
  # taking precedence over `:thinking[:effort]`. This is the per-attempt
65
93
  # override path used by `retry_policy { escalate({model:, reasoning_effort:}) }`
66
94
  # — the attempt-specific effort must win over the class-level default.
67
- # Forwarded provider-agnostically via `chat.with_thinking(**)` —
68
- # available since RubyLLM 1.12 (gemspec enforces this minimum).
95
+ # Forwarded provider-agnostically via `chat.with_thinking(**)`.
69
96
  thinking_config = resolve_thinking_config(options)
70
97
  chat.with_thinking(**thinking_config) if thinking_config
71
98
 
@@ -73,6 +100,29 @@ module RubyLLM
73
100
  # replacement for this one passthrough. `reasoning_effort` is not
74
101
  # forwarded here — it goes through `with_thinking` above.
75
102
  chat.with_max_output_tokens(options[:max_tokens]) if options[:max_tokens]
103
+ apply_provider_options(chat, options)
104
+ end
105
+
106
+ def apply_provider_options(chat, options)
107
+ provider_options = options[:provider_options]
108
+ return if provider_options.nil? || provider_options.empty?
109
+
110
+ ProviderOptions.validate!(provider_options, schema: !options[:schema].nil?)
111
+ chat.with_provider_options(provider_options)
112
+ end
113
+
114
+ # RubyLLM 1.x set temperature to 1.0 for OpenAI reasoning models, which
115
+ # reject any other value; 2.x sends it as given. A step escalating from
116
+ # a sampling model to one of those would fail, so a temperature is left
117
+ # out where the registry says the model takes none. A model RubyLLM does
118
+ # not know (assume_model_exists) gets it as before.
119
+ def apply_temperature(chat, temperature)
120
+ model = chat.respond_to?(:model) ? chat.model : nil
121
+ if model.respond_to?(:metadata) && model.metadata.is_a?(Hash) && model.metadata[:temperature] == false
122
+ Adapters::RubyLLM.warn_dropped_temperature(model)
123
+ else
124
+ chat.with_temperature(temperature)
125
+ end
76
126
  end
77
127
 
78
128
  # Returns merged `{ effort:, budget: }` or nil. `options[:reasoning_effort]`
@@ -84,26 +134,84 @@ module RubyLLM
84
134
  base.empty? ? nil : base
85
135
  end
86
136
 
87
- def build_response(response)
137
+ def build_response(response, model)
88
138
  content = response.content
89
139
  content = content.to_s unless content.is_a?(Hash) || content.is_a?(Array)
90
140
 
91
- # This is the ONLY place upstream token counts are read. RubyLLM 2.0
92
- # replaced `Message#input_tokens`/`#output_tokens` with a `Tokens` value
93
- # object. The `{ input_tokens:, output_tokens: }` shape below is this
94
- # gem's own public contract (Step::Trace#usage, cost calculation,
95
- # documented in the README) and deliberately does NOT change.
141
+ # This is the ONLY place upstream token counts are read. The
142
+ # `{ input_tokens:, output_tokens: }` shape is this gem's public
143
+ # contract (Step::Trace#usage, README) and does not change; RubyLLM's
144
+ # own counts go into it as described in `usage_from`.
96
145
  tokens = response.tokens
146
+ usage_complete = every_attempt_counted?(response, tokens)
147
+ reported = tokens&.reported_cost
148
+ usage = usage_from(tokens)
149
+ cost = call_cost(response, model, usage, reported)
97
150
 
98
151
  Response.new(
99
152
  content: content,
100
- usage: {
101
- input_tokens: tokens&.input || 0,
102
- output_tokens: tokens&.output || 0
103
- }
153
+ usage: usage,
154
+ cost: cost,
155
+ usage_complete: usage_complete,
156
+ # A cost RubyLLM priced from partial counts covers part of the call;
157
+ # one the provider reported covers all of it.
158
+ cost_complete: !cost.nil? && (usage_complete || !reported.nil?),
159
+ finish_reason: response.finish_reason
104
160
  )
105
161
  end
106
162
 
163
+ # `response.tokens` sums RubyLLM's attempts at the request, skipping
164
+ # counts an attempt never got, so the sum alone can look complete.
165
+ # RubyLLM records each attempt; one that failed after the provider may
166
+ # have billed it keeps nil counts, one never sent or refused gets zeros.
167
+ def every_attempt_counted?(response, tokens)
168
+ entries = response.respond_to?(:ruby_llm_usage_entries) ? Array(response.ruby_llm_usage_entries) : []
169
+ return entries.all? { |entry| reported_both_counts?(entry.tokens) } if entries.any?
170
+
171
+ reported_both_counts?(tokens)
172
+ end
173
+
174
+ # Zero is a reported count (a full cache hit leaves uncached input at 0);
175
+ # nil is a count the provider did not send.
176
+ def reported_both_counts?(tokens)
177
+ !tokens.nil? && !tokens.input.nil? && !tokens.output.nil?
178
+ end
179
+
180
+ # RubyLLM 2.x counts `input` without the prompt-cache tokens, which it
181
+ # reports apart as `cache_read` / `cache_write`. Our input_tokens is the
182
+ # whole prompt, so they are added back and also kept as a breakdown.
183
+ # `output` already includes thinking tokens when the provider bills them
184
+ # as output, so `thinking_tokens` is a breakdown, never added. A count
185
+ # the provider did not report is 0 here; `usage_complete` says so.
186
+ def usage_from(tokens)
187
+ return { input_tokens: 0, output_tokens: 0 } unless tokens
188
+
189
+ cache_read = tokens.cache_read.to_i
190
+ cache_write = tokens.cache_write.to_i
191
+ usage = { input_tokens: tokens.input.to_i + cache_read + cache_write, output_tokens: tokens.output.to_i }
192
+ usage[:cache_read_tokens] = cache_read if cache_read.positive?
193
+ usage[:cache_write_tokens] = cache_write if cache_write.positive?
194
+ usage[:thinking_tokens] = tokens.thinking.to_i if tokens.thinking.to_i.positive?
195
+ usage
196
+ end
197
+
198
+ # RubyLLM prices the response itself: per provider, with cache and
199
+ # long-context prices, the amount the provider reported when it did,
200
+ # and its own retries of the request. It returns nil when it cannot
201
+ # price every attempt. A price set with `register_model` still wins
202
+ # for its id, unless the provider reported what it billed; it prices
203
+ # the summed counts, so it is complete only when every attempt was
204
+ # counted (`cost_complete` in build_response).
205
+ def call_cost(response, model, usage, reported)
206
+ if reported.nil? && CostCalculator.registered?(model)
207
+ CostCalculator.calculate(model_name: model, usage: usage)
208
+ else
209
+ response.cost.total&.round(6)
210
+ end
211
+ rescue StandardError
212
+ nil
213
+ end
214
+
107
215
  def partition_messages(messages)
108
216
  system_contents = []
109
217
  conversation = []
@@ -43,6 +43,12 @@ module RubyLLM
43
43
  end
44
44
  end
45
45
 
46
+ # There is no provider to ask, so `token_count :exact` steps measure
47
+ # their input with the same heuristic as `:estimate` under this adapter.
48
+ def count_tokens(messages:, **_options)
49
+ TokenEstimator.estimate(messages)
50
+ end
51
+
46
52
  def call(messages:, **_options) # rubocop:disable Lint/UnusedMethodArgument
47
53
  content = if @responses
48
54
  c = @responses[@index] || @responses.last