ruby_llm-contract 1.1.1 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +103 -0
- data/README.md +18 -3
- data/docs/guide/getting_started.md +86 -1
- data/docs/guide/llm_judge.md +14 -0
- data/docs/guide/rails_integration.md +8 -0
- data/docs/guide/relation_to_agent.md +8 -2
- data/lib/ruby_llm/contract/adapters/response.rb +28 -2
- data/lib/ruby_llm/contract/adapters/ruby_llm.rb +136 -28
- data/lib/ruby_llm/contract/adapters/test.rb +6 -0
- data/lib/ruby_llm/contract/concerns/usage_aggregator.rb +8 -5
- data/lib/ruby_llm/contract/configuration.rb +5 -1
- data/lib/ruby_llm/contract/cost_calculator.rb +98 -44
- data/lib/ruby_llm/contract/eval/eval_history.rb +2 -1
- data/lib/ruby_llm/contract/eval/model_comparison.rb +11 -1
- data/lib/ruby_llm/contract/eval/recommender.rb +10 -3
- data/lib/ruby_llm/contract/eval/report_stats.rb +8 -3
- data/lib/ruby_llm/contract/eval/report_storage.rb +8 -2
- data/lib/ruby_llm/contract/pipeline/base.rb +3 -1
- data/lib/ruby_llm/contract/pipeline/runner.rb +11 -1
- data/lib/ruby_llm/contract/pipeline/trace.rb +10 -0
- data/lib/ruby_llm/contract/provider_options.rb +45 -0
- data/lib/ruby_llm/contract/step/adapter_caller.rb +2 -1
- data/lib/ruby_llm/contract/step/base.rb +29 -12
- data/lib/ruby_llm/contract/step/dsl.rb +54 -0
- data/lib/ruby_llm/contract/step/limit_checker.rb +48 -13
- data/lib/ruby_llm/contract/step/result_builder.rb +66 -4
- data/lib/ruby_llm/contract/step/retry_executor.rb +35 -4
- data/lib/ruby_llm/contract/step/retry_policy.rb +31 -5
- data/lib/ruby_llm/contract/step/runner.rb +20 -1
- data/lib/ruby_llm/contract/step/runner_config.rb +6 -3
- data/lib/ruby_llm/contract/step/trace.rb +61 -20
- data/lib/ruby_llm/contract/token_estimator.rb +5 -3
- data/lib/ruby_llm/contract/version.rb +1 -1
- data/lib/ruby_llm/contract/workflow_scope.rb +48 -0
- data/lib/ruby_llm/contract.rb +2 -0
- metadata +3 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: fead48c5516f8e0aa738d3ab1633ca2d07f7173aba6466a6d29e37634263b857
|
|
4
|
+
data.tar.gz: 4ae4c314cd6365c5f5de455f9dc5ddc0741abd26fa751804b22c5ffded0e6cb1
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 9b66df1e490e238fdacc84a728b874a008b80c8b44cdd882b444a84389bb77996ee6132efb6adaad79c91a04bee30eb7351241a4989ff1a4ed5df8d7a6cff619
|
|
7
|
+
data.tar.gz: a055fd3d06ee4fa9591fa2a6cfed6872142febd950fe3a455daa6e588f04e157a585701456f4815483ecc3accb803cf3bbcffa9308c6dfa5011c00f6a7143274
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,108 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 1.2.0 (2026-10-10)
|
|
4
|
+
|
|
5
|
+
Everything new is opt-in, except the temperature fix below.
|
|
6
|
+
|
|
7
|
+
### Added
|
|
8
|
+
|
|
9
|
+
- **`token_count :exact`** on a step: `max_input` and `max_cost` measure the input with
|
|
10
|
+
the provider's count (RubyLLM's `chat.count_tokens`: OpenAI Responses, Anthropic,
|
|
11
|
+
Gemini, Vertex AI, Bedrock Converse) instead of the chars/4 heuristic, attachments
|
|
12
|
+
included, so `attachment_token_estimate` is not needed. One extra request per attempt.
|
|
13
|
+
A provider without the endpoint, a failed count or an adapter without `count_tokens`
|
|
14
|
+
refuses the call (`:limit_exceeded`) rather than falling back to the heuristic; the
|
|
15
|
+
`Test` adapter counts with the heuristic. What a provider counts beyond prompt and
|
|
16
|
+
attachments is its own (Gemini's count leaves the output schema out).
|
|
17
|
+
`Adapters::RubyLLM#count_tokens`.
|
|
18
|
+
- **`on_incomplete_output :refuse`** on a step: a response the provider stopped at a token
|
|
19
|
+
limit fails as `:output_truncated`, one stopped by a content filter as
|
|
20
|
+
`:content_filtered`, before validation, so it is never `:ok`. Both statuses are outside
|
|
21
|
+
the default `retry_on`. The default, `:accept`, keeps 1.1.2's behaviour.
|
|
22
|
+
- **`retry_on` takes exception classes**: `retry_on :validation_failed,
|
|
23
|
+
RubyLLM::RateLimitError` retries an adapter error only of that class or a subclass;
|
|
24
|
+
`:adapter_error` still retries any. An adapter error's trace (and attempt) has
|
|
25
|
+
`error_class`.
|
|
26
|
+
- **`provider_options`** on a step (inherited, `:default` resets) and per call
|
|
27
|
+
(`context: { provider_options: }`, merged over the step's), passed to RubyLLM's
|
|
28
|
+
`with_provider_options`. Keys the step controls raise `ArgumentError`.
|
|
29
|
+
- **`workflow_instrumentation`** in `RubyLLM::Contract.configure`: a pipeline runs as one
|
|
30
|
+
`RubyLLM.workflow` with a `workflow.step` per alias, a step on its own as a workflow
|
|
31
|
+
named after its class, so RubyLLM's events and OpenTelemetry spans carry the step.
|
|
32
|
+
- **`register_model` takes cache prices**: `cache_read_per_1m:` and `cache_write_per_1m:`,
|
|
33
|
+
both optional.
|
|
34
|
+
- Guides: retry by error class, cut-off responses, exact token counts, local and custom
|
|
35
|
+
models (Ollama), provider options, workflow instrumentation, and `RubyLLM::Judge` (2.1)
|
|
36
|
+
in the LLM-judge guide. The README lists the new options and answers how to run local
|
|
37
|
+
models.
|
|
38
|
+
|
|
39
|
+
### Changed
|
|
40
|
+
|
|
41
|
+
- **A temperature is no longer sent to a model RubyLLM's registry marks as taking none**
|
|
42
|
+
(OpenAI's o-series and gpt-5 family). RubyLLM 1.x quietly sent 1.0 to those models;
|
|
43
|
+
2.x sends the value as given, so a step with `temperature` that escalated to one failed
|
|
44
|
+
with an adapter error. The adapter now leaves it out and warns once per model; a model
|
|
45
|
+
the registry does not know still gets it.
|
|
46
|
+
- **`retry_on` raises on a condition it cannot match.** A status given as a String (or
|
|
47
|
+
any value that is neither a Symbol nor an exception class) never matched a result, so
|
|
48
|
+
it silently turned retries off; it now raises `ArgumentError` when the step is defined.
|
|
49
|
+
|
|
50
|
+
## 1.1.2 (2026-10-10)
|
|
51
|
+
|
|
52
|
+
### Changed - costs and gates may move
|
|
53
|
+
|
|
54
|
+
- **Costs now count the prompt cache, and may go up.** RubyLLM 2.x reports cached prompt
|
|
55
|
+
tokens apart from `tokens.input`, and the adapter read only input and output: 1,000
|
|
56
|
+
uncached + 9,000 cached + 200 output tokens on gpt-4.1-mini came out at $0.00072
|
|
57
|
+
instead of $0.00162. `trace[:usage][:input_tokens]` is now the whole prompt and
|
|
58
|
+
`trace[:cost]` is the cost RubyLLM puts on the response: the price of the provider the
|
|
59
|
+
call went to, cache and long-context prices, the amount the provider reported when it
|
|
60
|
+
reports one, and RubyLLM's own retries of the request. Eval totals, history and
|
|
61
|
+
`compare_models` figures can rise against older baselines; baselines compare scores,
|
|
62
|
+
not costs, so they do not regress on this.
|
|
63
|
+
- **A call without token counts is no longer free.** When the provider sends no usage
|
|
64
|
+
(for the call or for any of RubyLLM's attempts at it), the counts still read 0 but
|
|
65
|
+
`trace[:usage_complete]` is `false` and, unless the provider reported what it billed,
|
|
66
|
+
the cost is unknown, so the eval cost gates refuse (opt out with
|
|
67
|
+
`on_unknown_pricing: :warn`).
|
|
68
|
+
- **A missing price is unknown, not $0.** A model priced for input but not output has no
|
|
69
|
+
cost instead of a partial one, so `max_cost` refuses it pre-flight. A call that wrote
|
|
70
|
+
to the prompt cache on a model without a cache-write price has an unknown cost
|
|
71
|
+
afterwards (pre-flight estimates do not predict cache writes).
|
|
72
|
+
|
|
73
|
+
### Fixed
|
|
74
|
+
|
|
75
|
+
- `CostCalculator` prices RubyLLM's models through `Model#cost_for`, so pre-flight
|
|
76
|
+
estimates use cache and long-context prices, and the price of the provider the call
|
|
77
|
+
goes to (12 model ids carry different prices per provider). `register_model` keeps
|
|
78
|
+
overriding the price for its id; cache reads are charged at its input price, cache
|
|
79
|
+
writes have no price.
|
|
80
|
+
- `compare_models` no longer names a model with unpriced cases the cheapest
|
|
81
|
+
(`best_for`), gives it an infinite `cost_per_point`, and `recommend` neither ranks it
|
|
82
|
+
on cost nor promises savings against it.
|
|
83
|
+
- `single_shot_cost` reported a later attempt's cost when the first attempt was unpriced.
|
|
84
|
+
- The step log printed an unknown cost as `$0.000000`; it now prints `cost=unknown`.
|
|
85
|
+
- Docs no longer say RubyLLM has no evaluation framework: `relation_to_agent.md` now
|
|
86
|
+
places `RubyLLM::Evaluation` (2.1) next to `define_eval`. The README and the getting
|
|
87
|
+
started guide describe what a trace's usage and cost count, and the README's 1.x
|
|
88
|
+
promise now says new trace keys may be added.
|
|
89
|
+
|
|
90
|
+
### Added
|
|
91
|
+
|
|
92
|
+
- `trace[:usage]` breakdown keys when reported: `cache_read_tokens`, `cache_write_tokens`,
|
|
93
|
+
`thinking_tokens` (parts of `input_tokens` / `output_tokens`, not extra tokens).
|
|
94
|
+
- `trace[:finish_reason]`, also on each retry attempt. A failed result whose response
|
|
95
|
+
stopped at a token limit or a content filter gets a last validation error saying so,
|
|
96
|
+
with the `max_output` the call ran with. Statuses and retries are unchanged.
|
|
97
|
+
- An adapter may report its own cost: `Adapters::Response.new(content:, usage:, cost:,
|
|
98
|
+
cost_complete:, usage_complete:, finish_reason:)`, where `cost: nil` means unknown.
|
|
99
|
+
A response without `cost:` is priced from `usage` as before.
|
|
100
|
+
- `provider:` on `CostCalculator.calculate`, `CostCalculator.find_model`,
|
|
101
|
+
`estimate_cost` and `estimate_eval_cost`.
|
|
102
|
+
- Baselines and history mark unpriced cases and runs with `cost_unknown: true`.
|
|
103
|
+
- A pipeline with a `token_budget` warns when a step's provider left out token counts;
|
|
104
|
+
`Pipeline::Trace#usage_complete?`.
|
|
105
|
+
|
|
3
106
|
## 1.1.1 (2026-10-10)
|
|
4
107
|
|
|
5
108
|
### Fixed
|
data/README.md
CHANGED
|
@@ -91,6 +91,8 @@ result.trace[:attempts]
|
|
|
91
91
|
# ]
|
|
92
92
|
```
|
|
93
93
|
|
|
94
|
+
`usage[:input_tokens]` is the whole prompt, prompt-cache tokens included (broken out as `cache_read_tokens` / `cache_write_tokens` when the provider reports them), and `cost` is what RubyLLM prices the response at: the provider's price, cache and long-context rates, or the amount the provider reported. A cost the gem cannot fully price is marked unknown rather than counted as $0 - see [what a trace's usage and cost count](docs/guide/getting_started.md#what-a-traces-usage-and-cost-count).
|
|
95
|
+
|
|
94
96
|
If the response is malformed, the TL;DR overflows the card, or the takeaway count is off, the gem moves to the next step. This is model **escalation**, not a fallback list — each step is an independent config (`model`, `reasoning_effort`), so the retry policy spends more compute only when the cheaper one couldn't satisfy the contract.
|
|
95
97
|
|
|
96
98
|
### Add a CI gate in 6 lines
|
|
@@ -128,6 +130,11 @@ Everything below is optional — the example above is a complete step. Reach for
|
|
|
128
130
|
- **[Find the cheapest viable fallback list](docs/guide/optimizing_retry_policy.md)** — empirically pick the cheapest model chain that still passes your evals.
|
|
129
131
|
- **[A/B test prompts](docs/guide/eval_first.md)** — measure whether a new prompt is safe to ship before merging.
|
|
130
132
|
- **[Budget caps](docs/guide/getting_started.md)** — refuse the request pre-flight when an estimate exceeds the limit.
|
|
133
|
+
- **[Cost tracking](docs/guide/getting_started.md#what-a-traces-usage-and-cost-count)** - per-call cost with prompt-cache and per-provider prices; eval cost gates fail closed when a cost is unknown.
|
|
134
|
+
- **[Exact token counts](docs/guide/getting_started.md#counting-input-tokens-exactly)** - `token_count :exact` checks `max_input` / `max_cost` against the provider's own count instead of a heuristic.
|
|
135
|
+
- **[Cut-off responses](docs/guide/getting_started.md#responses-cut-off-by-a-token-limit)** - `on_incomplete_output :refuse` fails an answer the provider stopped at a token limit, even one that would validate.
|
|
136
|
+
- **[Retry by error class](docs/guide/getting_started.md#retrying-provider-errors-by-class)** - `retry_on RubyLLM::RateLimitError` retries a rate limit without retrying an exhausted balance.
|
|
137
|
+
- **[Provider options](docs/guide/getting_started.md#provider-options)** - send `service_tier`, `seed` and other provider settings from the step or per call.
|
|
131
138
|
- **[Reasoning effort / thinking config](docs/guide/optimizing_retry_policy.md)** — Anthropic / OpenAI thinking configuration on the Step class.
|
|
132
139
|
|
|
133
140
|
Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and per-step models.
|
|
@@ -136,7 +143,7 @@ Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and
|
|
|
136
143
|
|
|
137
144
|
## Relation to `RubyLLM::Agent`
|
|
138
145
|
|
|
139
|
-
`Step::Base` and `RubyLLM::Agent` (since RubyLLM 1.12) are **siblings** targeting the same niche: reusable, class-based prompts. Both call into `RubyLLM::Chat` directly
|
|
146
|
+
`Step::Base` and `RubyLLM::Agent` (since RubyLLM 1.12) are **siblings** targeting the same niche: reusable, class-based prompts. Both call into `RubyLLM::Chat` directly - Step does not wrap Agent. Step adds the contract layer: `validate` (business invariants), `retry_policy escalate(...)` (model escalation on validation failure), `max_cost` pre-flight refusal, regression-eval framework, pipeline composition. RubyLLM 2.1's `RubyLLM::Evaluation` grades answers; `define_eval` guards against regressions and cost in CI - the mapping covers both. **[Full feature mapping →](docs/guide/relation_to_agent.md)**
|
|
140
147
|
|
|
141
148
|
## Relation to `ruby_llm-tribunal`
|
|
142
149
|
|
|
@@ -153,7 +160,7 @@ Different layers, complementary. [`ruby_llm-tribunal`](https://github.com/Alqemi
|
|
|
153
160
|
| [Why contracts?](docs/guide/why.md) | Recognise the four production failures the gem exists for |
|
|
154
161
|
| [Relation to RubyLLM::Agent](docs/guide/relation_to_agent.md) | Sibling abstractions; what each adds; runtime call path; coexistence patterns |
|
|
155
162
|
| [Relation to ruby_llm-tribunal](docs/guide/relation_to_tribunal.md) | Different layers (test framework vs runtime contract); visual flows; integration recipes |
|
|
156
|
-
| [Getting Started](docs/guide/getting_started.md) | Walk the full feature set on one concrete step |
|
|
163
|
+
| [Getting Started](docs/guide/getting_started.md) | Walk the full feature set on one concrete step, including what usage and cost count |
|
|
157
164
|
| [Rails integration](docs/guide/rails_integration.md) | Directory, initializer, jobs, logging, specs, CI gate — 7 FAQs for Rails devs |
|
|
158
165
|
| [Adopt in an existing Rails app](docs/guide/migration.md) | Replace raw `LlmClient.call` with a contract, Before/After |
|
|
159
166
|
| [Prevent silent prompt regressions](docs/guide/eval_first.md) | Evals, baselines, CI gates that block quality drift |
|
|
@@ -173,7 +180,11 @@ Stable since **1.0.0**. Semver tracked; breaking changes flagged in [CHANGELOG](
|
|
|
173
180
|
What 1.0 promises not to break inside the 1.x line: the documented DSL
|
|
174
181
|
(`prompt`, `output_schema`, `validate`, `retry_policy`, `max_cost`, `define_eval`,
|
|
175
182
|
pipelines), the `Result` and `trace` shapes including `trace[:usage]`'s
|
|
176
|
-
`{ input_tokens:, output_tokens: }` keys, and the adapter interface.
|
|
183
|
+
`{ input_tokens:, output_tokens: }` keys, and the adapter interface. New keys may be
|
|
184
|
+
added to both (1.1.2 added `cache_read_tokens`, `cache_write_tokens` and
|
|
185
|
+
`thinking_tokens` to `trace[:usage]`, and `usage_complete`, `cost_complete` and
|
|
186
|
+
`finish_reason` to the trace), and a figure may change when it was wrong: since 1.1.2
|
|
187
|
+
`input_tokens` counts prompt-cache tokens. A new major
|
|
177
188
|
version of a runtime dependency may require a new major version here — that is
|
|
178
189
|
what 1.0.0 itself was.
|
|
179
190
|
|
|
@@ -185,6 +196,10 @@ what 1.0.0 itself was.
|
|
|
185
196
|
|
|
186
197
|
**Where in a Rails app?** Default `app/contracts/`. The Railtie reloads `app/contracts/eval/` and `app/steps/eval/` in development; any autoloaded directory also works. See [Rails integration](docs/guide/rails_integration.md).
|
|
187
198
|
|
|
199
|
+
**Costs went up, or an eval cost gate turned red, after upgrading to 1.1.2?** Earlier versions left prompt-cache tokens out of `input_tokens` and the cost, priced every model at its default provider, counted a missing price as $0, and treated a call without token counts as free. The [CHANGELOG](CHANGELOG.md) lists each change; `on_unknown_pricing: :warn` turns the unknown-cost refusal into a warning.
|
|
200
|
+
|
|
201
|
+
**Local models (Ollama)?** Point RubyLLM at the server (`c.ollama_api_base = "http://localhost:11434/v1"`) and run steps with `context: { provider: :ollama, model: "...", assume_model_exists: true }`. Local models have no price in RubyLLM's registry, so register one (`CostCalculator.register_model("gemma-fast:latest", input_per_1m: 0, output_per_1m: 0)`) or the eval cost gates treat their cost as unknown. Ollama has no token-counting endpoint, so leave `token_count` at `:estimate`. See [local and custom models](docs/guide/getting_started.md#local-and-custom-models).
|
|
202
|
+
|
|
188
203
|
**Upgraded from pre-0.10.0 and getting `:limit_exceeded` with attachments?** Multimodal contracts with `max_cost`/`max_input` need `attachment_token_estimate`. See [multimodal input guide](docs/guide/multimodal_input.md#cost-attachment_token_estimate-is-required) for setup, fail-closed behaviour, and `on_unknown_attachment_size :warn` opt-out.
|
|
189
204
|
|
|
190
205
|
## License
|
|
@@ -65,6 +65,32 @@ result.trace[:attempts]
|
|
|
65
65
|
|
|
66
66
|
If the whole chain exhausts, `result.status` is the status of the last attempt (`:validation_failed` or `:parse_error`) and `result.parsed_output` is the last attempt's output. The caller decides what to do — ship it anyway, fall back to a template, or raise.
|
|
67
67
|
|
|
68
|
+
### Retrying provider errors by class
|
|
69
|
+
|
|
70
|
+
A provider error that RubyLLM's own retries could not get past ends as `:adapter_error`, with the exception class in `trace[:error_class]` (e.g. `"RubyLLM::RateLimitError"`). Adapter errors are not retried by default. List exception classes in `retry_on` to retry only those, typically together with `escalate` so the next attempt goes to another model:
|
|
71
|
+
|
|
72
|
+
```ruby
|
|
73
|
+
retry_policy do
|
|
74
|
+
escalate "gpt-4.1-mini", "claude-haiku-4-5"
|
|
75
|
+
retry_on :validation_failed, :parse_error, RubyLLM::RateLimitError, RubyLLM::OverloadedError
|
|
76
|
+
end
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
A class matches its subclasses; `:adapter_error` in the list retries every adapter error. As with statuses, the list replaces the defaults. Switching to a fallback model on transport errors inside one call is RubyLLM's `chat.with_fallbacks`; the contract's adapter does not call it, so fallbacks across models belong in `escalate`.
|
|
80
|
+
|
|
81
|
+
### Responses cut off by a token limit
|
|
82
|
+
|
|
83
|
+
When a provider stops at a token limit or a content filter, `trace[:finish_reason]` says so (`:max_tokens`, `:content_filter`), and a failed result gets a last validation error naming it. A cut-off answer that still parses and validates stays `:ok` unless the step refuses it:
|
|
84
|
+
|
|
85
|
+
```ruby
|
|
86
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
87
|
+
max_output 400
|
|
88
|
+
on_incomplete_output :refuse # default :accept
|
|
89
|
+
end
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
With `:refuse`, such a response fails before validation as `:output_truncated` (token limit) or `:content_filtered`, so `validate` blocks and observers never see it. Raw output, usage and cost stay on the result. Neither status is retried by default; `retry_on :output_truncated` opts in (useful when a later attempt has a larger `max_output`). Anthropic reports an exhausted context window as a token-limit stop too, so `:output_truncated` means "stopped on tokens", not necessarily "hit max_output".
|
|
93
|
+
|
|
68
94
|
### Per-attempt reasoning effort
|
|
69
95
|
|
|
70
96
|
`models:` accepts config hashes as well as model-name strings, so a fallback can "try harder" (more reasoning) on retry, not just switch model:
|
|
@@ -173,7 +199,48 @@ max_cost 0.01, on_unknown_pricing: :warn
|
|
|
173
199
|
|
|
174
200
|
Default is `:refuse`. Use `:warn` only when you accept running without a cost ceiling (fine-tuned models you trust, private endpoints).
|
|
175
201
|
|
|
176
|
-
The eval-level gates - `with_maximum_cost` on `pass_eval`, `maximum_cost` on the rake task, and `maximum_cost:` on `assert_eval_passes` - follow the same rule since 1.1.0. A report's total counts a case it cannot price as $0, so when any case ran on a model without pricing data the gate fails and names the cases instead of comparing an undercounted total. Opt out with `.on_unknown_pricing(:warn)`, `t.on_unknown_pricing = :warn` or `on_unknown_pricing: :warn`, which checks the budget against the priced cases and prints a warning. Offline runs (`sample_response`, a `Test` adapter without `usage:`) report zero tokens, cost $0 with or without pricing, and are never flagged.
|
|
202
|
+
The eval-level gates - `with_maximum_cost` on `pass_eval`, `maximum_cost` on the rake task, and `maximum_cost:` on `assert_eval_passes` - follow the same rule since 1.1.0. A report's total counts a case it cannot price as $0, so when any case ran on a model without pricing data the gate fails and names the cases instead of comparing an undercounted total. Opt out with `.on_unknown_pricing(:warn)`, `t.on_unknown_pricing = :warn` or `on_unknown_pricing: :warn`, which checks the budget against the priced cases and prints a warning. Offline runs (`sample_response`, a `Test` adapter without `usage:`) report zero tokens, cost $0 with or without pricing, and are never flagged. Since 1.1.2 a real call whose provider sent no token counts is flagged too, instead of passing as a free zero-token run.
|
|
203
|
+
|
|
204
|
+
### What a trace's usage and cost count
|
|
205
|
+
|
|
206
|
+
`trace[:usage][:input_tokens]` is the whole prompt the model took in, including tokens served from or written to the provider's prompt cache. When the provider reports them, the trace adds a breakdown: `cache_read_tokens`, `cache_write_tokens` and `thinking_tokens`. These are parts of `input_tokens` / `output_tokens`, not extra tokens; reasoning tokens are already in `output_tokens` when the provider bills them as output, and a model with a separate reasoning price is charged that price for them.
|
|
207
|
+
|
|
208
|
+
```ruby
|
|
209
|
+
result.trace[:usage]
|
|
210
|
+
# => { input_tokens: 10_000, output_tokens: 200, cache_read_tokens: 9_000 }
|
|
211
|
+
result.trace[:cost] # => 0.00162 (gpt-4.1-mini: cache reads at the cache price)
|
|
212
|
+
```
|
|
213
|
+
|
|
214
|
+
With the RubyLLM adapter, `trace[:cost]` is the cost RubyLLM puts on the response (`response.cost.total`): the price of the provider the call went to, cache and long-context prices, and the amount the provider reported when it reports one (OpenRouter, for example). A price set with `CostCalculator.register_model` still overrides it for that model id, unless the provider reported what it billed; there, cache reads are charged at the input price and a call that wrote to the cache has no price, since cache writes can cost more than input.
|
|
215
|
+
|
|
216
|
+
When a cost does not cover the whole call, `result.trace.cost_unknown?` is true and the eval gates above refuse. That happens when pricing is missing for a token category the call used, or when the provider sent no token counts for the call or for one of RubyLLM's attempts at it (`trace[:usage_complete]` is then `false` and missing counts read as 0) and did not report what it billed either. On a retried step `trace[:cost]` is the sum of the attempts that could be priced, and `cost_unknown?` is true if any attempt could not. `max_cost` itself is a pre-flight check on an estimate; it does not inspect the finished call.
|
|
217
|
+
|
|
218
|
+
A custom adapter returning `Response.new(content:, usage:)` is priced from `usage` as before. To report its own cost, pass `cost:` - `nil` means unknown - with `cost_complete:` and `usage_complete:`.
|
|
219
|
+
|
|
220
|
+
### Counting input tokens exactly
|
|
221
|
+
|
|
222
|
+
`max_input` and `max_cost` measure the input with a chars/4 heuristic (±30%). With `token_count :exact` they ask the provider instead:
|
|
223
|
+
|
|
224
|
+
```ruby
|
|
225
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
226
|
+
token_count :exact # default :estimate
|
|
227
|
+
max_input 20_000
|
|
228
|
+
end
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
The count comes from RubyLLM's `chat.count_tokens` (OpenAI Responses, Anthropic, Gemini, Vertex AI, Bedrock Converse) and covers the prompt and attachments, so `attachment_token_estimate` is not needed. What else a provider counts is its own: OpenAI's and Anthropic's counts include the output schema, Gemini's leaves it out. It costs one extra request per attempt, and RubyLLM leaves `provider_options` out of it. Where no count can be had - a provider without the endpoint (Ollama, for one), a failed request, or an adapter without `count_tokens` - the call is refused as `:limit_exceeded` with the reason, never measured with the heuristic under the name "exact". The `Test` adapter counts with the heuristic, since it has no provider to ask. The output side of `max_cost` is still an estimate (`max_output`, or the input size without one).
|
|
232
|
+
|
|
233
|
+
### Local and custom models
|
|
234
|
+
|
|
235
|
+
A local model (Ollama) or a fine-tuned one has no price in RubyLLM's registry, so its cost is unknown and the eval cost gates refuse. Register a price - zero for a model that costs nothing per token:
|
|
236
|
+
|
|
237
|
+
```ruby
|
|
238
|
+
RubyLLM::Contract::CostCalculator.register_model("gemma-fast:latest", input_per_1m: 0, output_per_1m: 0)
|
|
239
|
+
RubyLLM::Contract::CostCalculator.register_model("ft:gpt-4o-custom",
|
|
240
|
+
input_per_1m: 3.0, output_per_1m: 6.0, cache_read_per_1m: 1.5, cache_write_per_1m: 3.75)
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
`register_model` sets a price for a model id, not a model RubyLLM can call: the step still needs `provider:` (and `assume_model_exists: true` for an id RubyLLM does not list). Cache prices are optional; without them cache reads are charged at the input price and a call that wrote to the cache has no price. A price does not make up for missing counts: a provider that sends no token counts still leaves the cost unknown.
|
|
177
244
|
|
|
178
245
|
### Preflight cost estimates
|
|
179
246
|
|
|
@@ -195,6 +262,24 @@ SummarizeArticle.estimate_eval_cost("regression",
|
|
|
195
262
|
|
|
196
263
|
`estimate_cost` returns `nil` when pricing isn't registered. `estimate_eval_cost` silently treats unknown-pricing cases as `$0.00` and sums the rest — it does **not** fail closed the way `max_cost` does. Treat its output as a floor, not a guarantee; register pricing via `CostCalculator.register_model` before relying on it for budget decisions.
|
|
197
264
|
|
|
265
|
+
## Provider options
|
|
266
|
+
|
|
267
|
+
Settings RubyLLM has no method for go to the provider as-is through `provider_options`, on the step or per call:
|
|
268
|
+
|
|
269
|
+
```ruby
|
|
270
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
271
|
+
provider_options service_tier: "flex"
|
|
272
|
+
end
|
|
273
|
+
|
|
274
|
+
SummarizeArticle.run(text, context: { provider_options: { service_tier: "priority", seed: 7 } })
|
|
275
|
+
```
|
|
276
|
+
|
|
277
|
+
The call's keys win over the step's; a subclass replaces its parent's hash, and `provider_options :default` stops inheriting it. Keys the step itself sets (`model`, `temperature`, the max-token keys, the messages, and on a step with `output_schema` the structured-output keys `response_format` and `text`) raise `ArgumentError`, so a call cannot differ from what its limits and schema were checked against. Only top-level keys are checked. `compare_models` and `optimize_retry_policy` candidates do not carry provider options.
|
|
278
|
+
|
|
279
|
+
## Temperature on reasoning models
|
|
280
|
+
|
|
281
|
+
RubyLLM 1.x sent temperature 1.0 to OpenAI reasoning models (o-series, gpt-5), which reject other values; RubyLLM 2.x sends what it is given. A step with `temperature` therefore leaves it out for a model RubyLLM's registry marks as taking none, and warns once per model - so `temperature 0` with `escalate "gpt-4.1-mini", "gpt-5-mini"` still works. A model the registry does not know gets the temperature as set.
|
|
282
|
+
|
|
198
283
|
## `output_schema` vs `with_schema`
|
|
199
284
|
|
|
200
285
|
`with_schema` in `ruby_llm` tells the provider to force a specific JSON structure. `output_schema` in this gem does the same thing (calls `with_schema` under the hood) **plus** validates the response client-side. Cheaper models sometimes ignore schema constraints — `with_schema` is a request; `output_schema` is a request plus verification.
|
data/docs/guide/llm_judge.md
CHANGED
|
@@ -195,6 +195,20 @@ What to do operationally:
|
|
|
195
195
|
|
|
196
196
|
The drop is your signal to refine the judge's prompt, not to lower the gate.
|
|
197
197
|
|
|
198
|
+
## `RubyLLM::Judge` (RubyLLM 2.1)
|
|
199
|
+
|
|
200
|
+
RubyLLM 2.1 adds `RubyLLM::Judge` for typed questions - `probability`, `choice` and `score` - answered by **decision models**: OpenAI's judgment models, TypeSafe, or local decision models through Ollama. It does not run on ordinary chat models; those raise that they do not support judgments. Where you have such a model, a judge class can stand in for the judge Step in step 1 above:
|
|
201
|
+
|
|
202
|
+
```ruby
|
|
203
|
+
class SummaryFaithfulness < RubyLLM::Judge
|
|
204
|
+
probability :faithful, "Is every claim in the summary supported by the article?"
|
|
205
|
+
end
|
|
206
|
+
|
|
207
|
+
SummaryFaithfulness.judge(article_and_summary).faithful.probability # => 0.93
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
Calibrate it against human labels exactly as in step 2; the threshold is still yours to set. On an ordinary chat model, keep the judge as a Contract Step - the pattern this guide builds.
|
|
211
|
+
|
|
198
212
|
## When to reach for Tribunal instead
|
|
199
213
|
|
|
200
214
|
[`ruby_llm-tribunal`](https://github.com/Alqemist-labs/ruby_llm-tribunal) ships an off-the-shelf catalog of common LLM-as-judge assertions (`assert_faithful`, `assert_hallucination`, `assert_refusal`, `assert_no_pii`, etc.) — a shortcut when your check matches one of those domain-general categories. The methodology in this guide still applies: calibrate the judge against your human-labeled production data **before** trusting Tribunal's `default_threshold = 0.8`, refine the prompt when it over-flags, watch for the anti-patterns above. See [Relation to Tribunal](relation_to_tribunal.md) for the full positioning — what each gem documents (and doesn't), a concrete decision tree on catalog-vs-custom-judge, and three working integration patterns.
|
|
@@ -125,6 +125,14 @@ end
|
|
|
125
125
|
|
|
126
126
|
Trace inspection in an admin UI: `result.trace[:attempts]` gives you per-attempt model, status, cost, latency — render it in a partial to debug production failures without re-running.
|
|
127
127
|
|
|
128
|
+
RubyLLM 2.1 emits its own events and OpenTelemetry spans for every call. To group them by step, turn on workflow instrumentation in the initializer:
|
|
129
|
+
|
|
130
|
+
```ruby
|
|
131
|
+
RubyLLM::Contract.configure { |c| c.workflow_instrumentation = true }
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
A pipeline then runs as one `RubyLLM.workflow` named after its class, with each step as `workflow.step(alias)`; a step run on its own is a workflow named after its class. RubyLLM's events carry `workflow_name` and `workflow_step_name`, and a workflow your code opened around the call becomes the parent. Off by default. In a Rails app RubyLLM also records each call in its usage ledger when configured; that figure and `trace[:cost]` are the same RubyLLM price, unless `CostCalculator.register_model` overrides it for the model.
|
|
135
|
+
|
|
128
136
|
## 5. Testing — RSpec and Minitest
|
|
129
137
|
|
|
130
138
|
Add to `spec/spec_helper.rb` (or `test_helper.rb`):
|
|
@@ -12,8 +12,8 @@
|
|
|
12
12
|
| `validate :rule do ... end` business invariants on output | only in `ruby_llm-contract` |
|
|
13
13
|
| `retry_policy escalate(...)` model escalation on validation failure | only here (different from RubyLLM's network-level retry) |
|
|
14
14
|
| `max_cost` / `max_input` / `max_output` pre-flight refusal | only here |
|
|
15
|
-
| `define_eval` + baseline regression + `compare_models` + `optimize_retry_policy` | only here
|
|
16
|
-
| Pipeline composition with `step SomeStep, as: :alias` | only here
|
|
15
|
+
| `define_eval` + baseline regression + `compare_models` + `optimize_retry_policy` | only here; RubyLLM 2.1's `RubyLLM::Evaluation` grades answers (see below) but keeps no baseline and has no cost gate |
|
|
16
|
+
| Pipeline composition with `step SomeStep, as: :alias` | only here; `RubyLLM.workflow` (2.1) groups calls for tracing, without typed step outputs, fail-fast or budgets |
|
|
17
17
|
| `around_call`, named `observe` hooks with pass/fail recorded in trace | only here |
|
|
18
18
|
|
|
19
19
|
## Runtime relationship
|
|
@@ -32,6 +32,12 @@ Step.run(input)
|
|
|
32
32
|
|
|
33
33
|
This may change in a future release if upstream APIs make a layered design natural. The decision is not committed; it depends on adopter signal.
|
|
34
34
|
|
|
35
|
+
## Relation to `RubyLLM::Evaluation`
|
|
36
|
+
|
|
37
|
+
RubyLLM 2.1 ships `RubyLLM::Evaluation`: a dataset of cases run through your `perform` method, graded for correctness by an LLM or by your own criteria and Ruby assertions, with RSpec/Minitest integration and a report of verdicts, tokens and cost. Use it to answer "is this answer right?".
|
|
38
|
+
|
|
39
|
+
`define_eval` answers a different question: did this prompt or model get worse than last time, and does it stay within budget? It keeps baselines and flags regressions, compares prompts (`compare_with`) and models (`compare_models`, `recommend`), runs offline with `sample_response`, and gates CI on score and cost. The two can share a project; an Evaluation's `perform` can call `Step.run`.
|
|
40
|
+
|
|
35
41
|
## Coexistence on the same project
|
|
36
42
|
|
|
37
43
|
The two abstractions can live in the same Rails (or non-Rails) project. Pick one per use case:
|
|
@@ -3,16 +3,42 @@
|
|
|
3
3
|
module RubyLLM
|
|
4
4
|
module Contract
|
|
5
5
|
module Adapters
|
|
6
|
+
# `content` and `usage` are the whole interface an adapter has to fill.
|
|
7
|
+
# The rest is optional accounting: an adapter that leaves `cost` out gets
|
|
8
|
+
# its usage priced from the registry, as before; one that passes `cost`
|
|
9
|
+
# (nil included) owns the figure, and `cost_complete` / `usage_complete`
|
|
10
|
+
# say whether it and the token counts cover the whole call.
|
|
6
11
|
class Response
|
|
7
12
|
include Concerns::DeepFreeze
|
|
8
13
|
|
|
9
|
-
attr_reader :content, :usage
|
|
14
|
+
attr_reader :content, :usage, :usage_complete, :cost_complete, :finish_reason
|
|
10
15
|
|
|
11
|
-
def initialize(content:, usage: {}
|
|
16
|
+
def initialize(content:, usage: {}, cost: Step::Trace::COST_UNSET, usage_complete: nil,
|
|
17
|
+
cost_complete: nil, finish_reason: nil)
|
|
12
18
|
@content = deep_dup_freeze(content)
|
|
13
19
|
@usage = deep_dup_freeze(usage)
|
|
20
|
+
@cost = cost
|
|
21
|
+
@usage_complete = usage_complete
|
|
22
|
+
@cost_complete = cost_complete
|
|
23
|
+
@finish_reason = finish_reason
|
|
14
24
|
freeze
|
|
15
25
|
end
|
|
26
|
+
|
|
27
|
+
def cost
|
|
28
|
+
cost_provided? ? @cost : nil
|
|
29
|
+
end
|
|
30
|
+
|
|
31
|
+
def cost_provided?
|
|
32
|
+
!@cost.equal?(Step::Trace::COST_UNSET)
|
|
33
|
+
end
|
|
34
|
+
|
|
35
|
+
# Keyword arguments for Step::Trace.new. Omits what the adapter did not
|
|
36
|
+
# report, so a plain Response keeps the registry pricing path.
|
|
37
|
+
def trace_accounting
|
|
38
|
+
fields = { usage_complete: @usage_complete, cost_complete: @cost_complete,
|
|
39
|
+
finish_reason: @finish_reason }.compact
|
|
40
|
+
cost_provided? ? fields.merge(cost: @cost) : fields
|
|
41
|
+
end
|
|
16
42
|
end
|
|
17
43
|
end
|
|
18
44
|
end
|
|
@@ -6,32 +6,59 @@ module RubyLLM
|
|
|
6
6
|
module Contract
|
|
7
7
|
module Adapters
|
|
8
8
|
class RubyLLM < Base
|
|
9
|
+
# `with: nil` adds no attachment. The text is never nil (`fetch(:content, "")`),
|
|
10
|
+
# and RubyLLM raises only when both text and attachments are nil.
|
|
9
11
|
def call(messages:, **options)
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
add_history(chat, conversation[0..-2])
|
|
12
|
+
chat, text = prepared_chat(messages, options)
|
|
13
|
+
response = chat.ask(text, with: options[:attachment])
|
|
14
|
+
build_response(response, options[:model])
|
|
15
|
+
end
|
|
15
16
|
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
17
|
+
# Input tokens of the request `call` would send, counted by the
|
|
18
|
+
# provider (`chat.count_tokens`), attachments included. `ask` is
|
|
19
|
+
# `ask_later` plus `complete` in RubyLLM, so the staged message is the
|
|
20
|
+
# one `call` sends. RubyLLM leaves provider_options out of the count.
|
|
21
|
+
# Raises RubyLLM::Error where the provider has no counting endpoint.
|
|
22
|
+
def count_tokens(messages:, **options)
|
|
23
|
+
chat, text = prepared_chat(messages, options)
|
|
24
|
+
chat.ask_later(text, with: options[:attachment]).count_tokens
|
|
25
|
+
rescue KeyError => e
|
|
26
|
+
# RubyLLM reads OpenAI's count with `fetch`; a reply without the
|
|
27
|
+
# field is a failed count, not a bug in the caller.
|
|
28
|
+
raise ::RubyLLM::Error, "count response without input_tokens (#{e.message})"
|
|
26
29
|
end
|
|
27
30
|
|
|
28
31
|
CHAT_OPTION_METHODS = {
|
|
29
|
-
temperature: :with_temperature,
|
|
30
32
|
schema: :with_schema
|
|
31
33
|
}.freeze
|
|
32
34
|
|
|
35
|
+
@dropped_temperature_models = Set.new
|
|
36
|
+
@dropped_temperature_lock = Mutex.new
|
|
37
|
+
|
|
38
|
+
# Warns once per provider and model, across threads and subclasses
|
|
39
|
+
# (called on this class, whose state subclasses do not inherit).
|
|
40
|
+
def self.warn_dropped_temperature(model)
|
|
41
|
+
key = "#{model.provider}:#{model.id}"
|
|
42
|
+
first = @dropped_temperature_lock.synchronize { @dropped_temperature_models.add?(key) }
|
|
43
|
+
return unless first
|
|
44
|
+
|
|
45
|
+
warn "[ruby_llm-contract] #{model.id} (#{model.provider}) does not accept temperature " \
|
|
46
|
+
"according to RubyLLM's model registry; the step's temperature is not sent"
|
|
47
|
+
end
|
|
48
|
+
|
|
33
49
|
private
|
|
34
50
|
|
|
51
|
+
# The configured chat with the history added, and the text of the last
|
|
52
|
+
# message, which sending and counting both stage themselves.
|
|
53
|
+
def prepared_chat(messages, options)
|
|
54
|
+
system_contents, conversation = partition_messages(messages)
|
|
55
|
+
conversation = fallback_conversation(system_contents, conversation)
|
|
56
|
+
|
|
57
|
+
chat = build_chat(options, system_contents)
|
|
58
|
+
add_history(chat, conversation[0..-2])
|
|
59
|
+
[chat, conversation.last&.fetch(:content, "")]
|
|
60
|
+
end
|
|
61
|
+
|
|
35
62
|
# When prompt has only system/section/rule nodes and no user message,
|
|
36
63
|
# pop the last system message and use it as the user ask.
|
|
37
64
|
def fallback_conversation(system_contents, conversation)
|
|
@@ -59,13 +86,13 @@ module RubyLLM
|
|
|
59
86
|
CHAT_OPTION_METHODS.each do |key, method_name|
|
|
60
87
|
chat.public_send(method_name, options[key]) if options[key]
|
|
61
88
|
end
|
|
89
|
+
apply_temperature(chat, options[:temperature]) unless options[:temperature].nil?
|
|
62
90
|
|
|
63
91
|
# Resolve thinking config from BOTH sources, with `:reasoning_effort`
|
|
64
92
|
# taking precedence over `:thinking[:effort]`. This is the per-attempt
|
|
65
93
|
# override path used by `retry_policy { escalate({model:, reasoning_effort:}) }`
|
|
66
94
|
# — the attempt-specific effort must win over the class-level default.
|
|
67
|
-
# Forwarded provider-agnostically via `chat.with_thinking(**)
|
|
68
|
-
# available since RubyLLM 1.12 (gemspec enforces this minimum).
|
|
95
|
+
# Forwarded provider-agnostically via `chat.with_thinking(**)`.
|
|
69
96
|
thinking_config = resolve_thinking_config(options)
|
|
70
97
|
chat.with_thinking(**thinking_config) if thinking_config
|
|
71
98
|
|
|
@@ -73,6 +100,29 @@ module RubyLLM
|
|
|
73
100
|
# replacement for this one passthrough. `reasoning_effort` is not
|
|
74
101
|
# forwarded here — it goes through `with_thinking` above.
|
|
75
102
|
chat.with_max_output_tokens(options[:max_tokens]) if options[:max_tokens]
|
|
103
|
+
apply_provider_options(chat, options)
|
|
104
|
+
end
|
|
105
|
+
|
|
106
|
+
def apply_provider_options(chat, options)
|
|
107
|
+
provider_options = options[:provider_options]
|
|
108
|
+
return if provider_options.nil? || provider_options.empty?
|
|
109
|
+
|
|
110
|
+
ProviderOptions.validate!(provider_options, schema: !options[:schema].nil?)
|
|
111
|
+
chat.with_provider_options(provider_options)
|
|
112
|
+
end
|
|
113
|
+
|
|
114
|
+
# RubyLLM 1.x set temperature to 1.0 for OpenAI reasoning models, which
|
|
115
|
+
# reject any other value; 2.x sends it as given. A step escalating from
|
|
116
|
+
# a sampling model to one of those would fail, so a temperature is left
|
|
117
|
+
# out where the registry says the model takes none. A model RubyLLM does
|
|
118
|
+
# not know (assume_model_exists) gets it as before.
|
|
119
|
+
def apply_temperature(chat, temperature)
|
|
120
|
+
model = chat.respond_to?(:model) ? chat.model : nil
|
|
121
|
+
if model.respond_to?(:metadata) && model.metadata.is_a?(Hash) && model.metadata[:temperature] == false
|
|
122
|
+
Adapters::RubyLLM.warn_dropped_temperature(model)
|
|
123
|
+
else
|
|
124
|
+
chat.with_temperature(temperature)
|
|
125
|
+
end
|
|
76
126
|
end
|
|
77
127
|
|
|
78
128
|
# Returns merged `{ effort:, budget: }` or nil. `options[:reasoning_effort]`
|
|
@@ -84,26 +134,84 @@ module RubyLLM
|
|
|
84
134
|
base.empty? ? nil : base
|
|
85
135
|
end
|
|
86
136
|
|
|
87
|
-
def build_response(response)
|
|
137
|
+
def build_response(response, model)
|
|
88
138
|
content = response.content
|
|
89
139
|
content = content.to_s unless content.is_a?(Hash) || content.is_a?(Array)
|
|
90
140
|
|
|
91
|
-
# This is the ONLY place upstream token counts are read.
|
|
92
|
-
#
|
|
93
|
-
#
|
|
94
|
-
#
|
|
95
|
-
# documented in the README) and deliberately does NOT change.
|
|
141
|
+
# This is the ONLY place upstream token counts are read. The
|
|
142
|
+
# `{ input_tokens:, output_tokens: }` shape is this gem's public
|
|
143
|
+
# contract (Step::Trace#usage, README) and does not change; RubyLLM's
|
|
144
|
+
# own counts go into it as described in `usage_from`.
|
|
96
145
|
tokens = response.tokens
|
|
146
|
+
usage_complete = every_attempt_counted?(response, tokens)
|
|
147
|
+
reported = tokens&.reported_cost
|
|
148
|
+
usage = usage_from(tokens)
|
|
149
|
+
cost = call_cost(response, model, usage, reported)
|
|
97
150
|
|
|
98
151
|
Response.new(
|
|
99
152
|
content: content,
|
|
100
|
-
usage:
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
153
|
+
usage: usage,
|
|
154
|
+
cost: cost,
|
|
155
|
+
usage_complete: usage_complete,
|
|
156
|
+
# A cost RubyLLM priced from partial counts covers part of the call;
|
|
157
|
+
# one the provider reported covers all of it.
|
|
158
|
+
cost_complete: !cost.nil? && (usage_complete || !reported.nil?),
|
|
159
|
+
finish_reason: response.finish_reason
|
|
104
160
|
)
|
|
105
161
|
end
|
|
106
162
|
|
|
163
|
+
# `response.tokens` sums RubyLLM's attempts at the request, skipping
|
|
164
|
+
# counts an attempt never got, so the sum alone can look complete.
|
|
165
|
+
# RubyLLM records each attempt; one that failed after the provider may
|
|
166
|
+
# have billed it keeps nil counts, one never sent or refused gets zeros.
|
|
167
|
+
def every_attempt_counted?(response, tokens)
|
|
168
|
+
entries = response.respond_to?(:ruby_llm_usage_entries) ? Array(response.ruby_llm_usage_entries) : []
|
|
169
|
+
return entries.all? { |entry| reported_both_counts?(entry.tokens) } if entries.any?
|
|
170
|
+
|
|
171
|
+
reported_both_counts?(tokens)
|
|
172
|
+
end
|
|
173
|
+
|
|
174
|
+
# Zero is a reported count (a full cache hit leaves uncached input at 0);
|
|
175
|
+
# nil is a count the provider did not send.
|
|
176
|
+
def reported_both_counts?(tokens)
|
|
177
|
+
!tokens.nil? && !tokens.input.nil? && !tokens.output.nil?
|
|
178
|
+
end
|
|
179
|
+
|
|
180
|
+
# RubyLLM 2.x counts `input` without the prompt-cache tokens, which it
|
|
181
|
+
# reports apart as `cache_read` / `cache_write`. Our input_tokens is the
|
|
182
|
+
# whole prompt, so they are added back and also kept as a breakdown.
|
|
183
|
+
# `output` already includes thinking tokens when the provider bills them
|
|
184
|
+
# as output, so `thinking_tokens` is a breakdown, never added. A count
|
|
185
|
+
# the provider did not report is 0 here; `usage_complete` says so.
|
|
186
|
+
def usage_from(tokens)
|
|
187
|
+
return { input_tokens: 0, output_tokens: 0 } unless tokens
|
|
188
|
+
|
|
189
|
+
cache_read = tokens.cache_read.to_i
|
|
190
|
+
cache_write = tokens.cache_write.to_i
|
|
191
|
+
usage = { input_tokens: tokens.input.to_i + cache_read + cache_write, output_tokens: tokens.output.to_i }
|
|
192
|
+
usage[:cache_read_tokens] = cache_read if cache_read.positive?
|
|
193
|
+
usage[:cache_write_tokens] = cache_write if cache_write.positive?
|
|
194
|
+
usage[:thinking_tokens] = tokens.thinking.to_i if tokens.thinking.to_i.positive?
|
|
195
|
+
usage
|
|
196
|
+
end
|
|
197
|
+
|
|
198
|
+
# RubyLLM prices the response itself: per provider, with cache and
|
|
199
|
+
# long-context prices, the amount the provider reported when it did,
|
|
200
|
+
# and its own retries of the request. It returns nil when it cannot
|
|
201
|
+
# price every attempt. A price set with `register_model` still wins
|
|
202
|
+
# for its id, unless the provider reported what it billed; it prices
|
|
203
|
+
# the summed counts, so it is complete only when every attempt was
|
|
204
|
+
# counted (`cost_complete` in build_response).
|
|
205
|
+
def call_cost(response, model, usage, reported)
|
|
206
|
+
if reported.nil? && CostCalculator.registered?(model)
|
|
207
|
+
CostCalculator.calculate(model_name: model, usage: usage)
|
|
208
|
+
else
|
|
209
|
+
response.cost.total&.round(6)
|
|
210
|
+
end
|
|
211
|
+
rescue StandardError
|
|
212
|
+
nil
|
|
213
|
+
end
|
|
214
|
+
|
|
107
215
|
def partition_messages(messages)
|
|
108
216
|
system_contents = []
|
|
109
217
|
conversation = []
|
|
@@ -43,6 +43,12 @@ module RubyLLM
|
|
|
43
43
|
end
|
|
44
44
|
end
|
|
45
45
|
|
|
46
|
+
# There is no provider to ask, so `token_count :exact` steps measure
|
|
47
|
+
# their input with the same heuristic as `:estimate` under this adapter.
|
|
48
|
+
def count_tokens(messages:, **_options)
|
|
49
|
+
TokenEstimator.estimate(messages)
|
|
50
|
+
end
|
|
51
|
+
|
|
46
52
|
def call(messages:, **_options) # rubocop:disable Lint/UnusedMethodArgument
|
|
47
53
|
content = if @responses
|
|
48
54
|
c = @responses[@index] || @responses.last
|