ruby_llm-contract 0.8.0 → 0.10.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +79 -1
  3. data/Gemfile.lock +2 -2
  4. data/README.md +96 -37
  5. data/docs/architecture.md +50 -0
  6. data/docs/guide/best_practices.md +136 -0
  7. data/docs/guide/eval_first.md +192 -0
  8. data/docs/guide/getting_started.md +199 -0
  9. data/docs/guide/migration.md +185 -0
  10. data/docs/guide/multimodal_input.md +160 -0
  11. data/docs/guide/optimizing_retry_policy.md +131 -0
  12. data/docs/guide/output_schema.md +93 -0
  13. data/docs/guide/pipeline.md +154 -0
  14. data/docs/guide/prompt_ast.md +76 -0
  15. data/docs/guide/rails_integration.md +218 -0
  16. data/docs/guide/relation_to_agent.md +52 -0
  17. data/docs/guide/relation_to_tribunal.md +135 -0
  18. data/docs/guide/testing.md +282 -0
  19. data/docs/guide/why.md +103 -0
  20. data/lib/ruby_llm/contract/adapters/ruby_llm.rb +9 -1
  21. data/lib/ruby_llm/contract/concerns/eval_host.rb +6 -9
  22. data/lib/ruby_llm/contract/concerns/stub_helpers.rb +97 -0
  23. data/lib/ruby_llm/contract/contract/definition.rb +2 -0
  24. data/lib/ruby_llm/contract/cost_calculator.rb +11 -2
  25. data/lib/ruby_llm/contract/eval/recommender.rb +3 -1
  26. data/lib/ruby_llm/contract/eval/retry_optimizer.rb +16 -13
  27. data/lib/ruby_llm/contract/eval.rb +13 -0
  28. data/lib/ruby_llm/contract/minitest.rb +6 -108
  29. data/lib/ruby_llm/contract/pipeline/result.rb +1 -1
  30. data/lib/ruby_llm/contract/rake_task/suite_gate.rb +117 -0
  31. data/lib/ruby_llm/contract/rake_task.rb +30 -51
  32. data/lib/ruby_llm/contract/rspec/helpers.rb +9 -123
  33. data/lib/ruby_llm/contract/step/base.rb +56 -24
  34. data/lib/ruby_llm/contract/step/dsl.rb +91 -63
  35. data/lib/ruby_llm/contract/step/limit_checker.rb +34 -1
  36. data/lib/ruby_llm/contract/step/retry_executor.rb +6 -13
  37. data/lib/ruby_llm/contract/step/runner.rb +22 -20
  38. data/lib/ruby_llm/contract/step/runner_config.rb +26 -0
  39. data/lib/ruby_llm/contract/version.rb +1 -1
  40. data/lib/ruby_llm/contract.rb +1 -0
  41. data/ruby_llm-contract.gemspec +5 -1
  42. metadata +18 -4
  43. data/.rspec +0 -3
  44. data/.rubycritic.yml +0 -8
  45. data/.simplecov +0 -22
@@ -0,0 +1,282 @@
1
+ # Testing
2
+
3
+ > Read this when writing deterministic unit specs for contract steps. Skip if you only ever run live evals in CI and never stub LLM calls.
4
+
5
+ CI that hits a real LLM on every commit is slow and costly: tests take minutes instead of milliseconds, every rerun spends money, and flakes on provider hiccups block merges. The Test adapter + `stub_step` make LLM-backed specs run like ordinary unit tests — deterministic, offline, free. Live evals stay as an opt-in CI stage for quality gating (see [Eval-First](eval_first.md)).
6
+
7
+ How to write deterministic specs and matchers for steps built on `ruby_llm-contract`. Examples use `SummarizeArticle` (the flagship step from the [README](../../README.md)).
8
+
9
+ ## Test adapter
10
+
11
+ Ships deterministic specs with zero API calls. Accepts a String, Hash, or Array:
12
+
13
+ ```ruby
14
+ # String JSON
15
+ adapter = RubyLLM::Contract::Adapters::Test.new(
16
+ response: '{"tldr":"...","takeaways":["a","b","c"],"tone":"neutral"}'
17
+ )
18
+
19
+ # Hash — auto-converted to JSON
20
+ adapter = RubyLLM::Contract::Adapters::Test.new(
21
+ response: { tldr: "...", takeaways: %w[a b c], tone: "neutral" }
22
+ )
23
+
24
+ # Multiple sequential responses (one per call)
25
+ adapter = RubyLLM::Contract::Adapters::Test.new(
26
+ responses: [
27
+ { tldr: "...", takeaways: %w[a b c], tone: "neutral" },
28
+ { tldr: "...", takeaways: %w[x y z], tone: "analytical" }
29
+ ]
30
+ )
31
+
32
+ result = SummarizeArticle.run("article text", context: { adapter: adapter })
33
+ result.ok? # => true
34
+ ```
35
+
36
+ Multi-step pipeline testing with per-step named responses (using `ArticleCardPipeline` from [Pipeline](pipeline.md)):
37
+
38
+ ```ruby
39
+ result = ArticleCardPipeline.test("article text",
40
+ responses: {
41
+ summarize: { tldr: "...", takeaways: %w[a b c], tone: "analytical" },
42
+ tag: { tldr: "...", takeaways: %w[a b c], tone: "analytical",
43
+ hashtags: %w[#ruby #release] },
44
+ card: { headline: "Ruby 3.4 ships", summary: "...",
45
+ hashtags: %w[#ruby #release], sentiment_icon: "🧠" }
46
+ }
47
+ )
48
+ ```
49
+
50
+ ## Output keys are always symbols
51
+
52
+ Parsed output uses **symbol keys**, never strings:
53
+
54
+ ```ruby
55
+ result.parsed_output[:tldr] # => "..." ✓
56
+ result.parsed_output["tldr"] # => nil ✗
57
+ ```
58
+
59
+ The gem warns if a `validate` or `verify` block returns `nil` — usually a sign of string-key access on symbol-keyed data.
60
+
61
+ ## RSpec setup
62
+
63
+ In `spec_helper.rb`:
64
+
65
+ ```ruby
66
+ require "ruby_llm/contract/rspec"
67
+ ```
68
+
69
+ You get the `satisfy_contract` matcher, `pass_eval` matcher, and the `stub_step` helpers.
70
+
71
+ ## stub_step helpers
72
+
73
+ `stub_step` canned-responses a single step; other steps run normally.
74
+
75
+ ```ruby
76
+ RSpec.describe SummarizeArticle do
77
+ before do
78
+ stub_step(described_class,
79
+ response: { tldr: "...", takeaways: %w[a b c], tone: "neutral" })
80
+ end
81
+
82
+ it "satisfies its contract" do
83
+ result = described_class.run("article text")
84
+ expect(result).to satisfy_contract
85
+ end
86
+ end
87
+ ```
88
+
89
+ **Sequential responses:**
90
+
91
+ ```ruby
92
+ stub_step(described_class, responses: [
93
+ { tldr: "...", takeaways: %w[a b c], tone: "neutral" },
94
+ { tldr: "...", takeaways: %w[x y z], tone: "analytical" }
95
+ ])
96
+ ```
97
+
98
+ **Block form (auto-cleanup):**
99
+
100
+ ```ruby
101
+ stub_step(SummarizeArticle, response: { tldr: "...", takeaways: %w[a b c], tone: "neutral" }) do
102
+ result = SummarizeArticle.run("article text")
103
+ # stub is active here
104
+ end
105
+ # stub gone — original adapter restored
106
+ ```
107
+
108
+ **Multiple steps at once:**
109
+
110
+ ```ruby
111
+ stub_steps(
112
+ SummarizeArticle => { response: { tldr: "...", takeaways: %w[a b c], tone: "neutral" } },
113
+ GenerateHashtags => { response: { tldr: "...", takeaways: %w[a b c],
114
+ tone: "neutral",
115
+ hashtags: %w[#ruby #release] } }
116
+ ) do
117
+ result = ArticleCardPipeline.run("article text")
118
+ end
119
+ ```
120
+
121
+ **Global stub for all steps:**
122
+
123
+ ```ruby
124
+ stub_all_steps(response: { default: true })
125
+ ```
126
+
127
+ In RSpec, non-block stubs auto-clean after each example. In Minitest, `teardown` restores the original adapter (via `MinitestHelpers`).
128
+
129
+ ## Minitest
130
+
131
+ Require `ruby_llm/contract/minitest` in your `test_helper.rb`. You get the same `satisfy_contract` / `pass_eval` assertions and `stub_step` helper, adapted for Minitest syntax.
132
+
133
+ ## RSpec matchers
134
+
135
+ ```ruby
136
+ RSpec.describe SummarizeArticle do
137
+ before do
138
+ stub_step(described_class,
139
+ response: { tldr: "...", takeaways: %w[a b c], tone: "neutral" })
140
+ end
141
+
142
+ it "satisfies its contract" do
143
+ result = described_class.run("article text")
144
+ expect(result).to satisfy_contract
145
+ end
146
+
147
+ it "rejects invalid output" do
148
+ stub_step(described_class, response: { tldr: "x" * 300, takeaways: %w[a b c], tone: "neutral" })
149
+ result = described_class.run("article text")
150
+ expect(result).not_to satisfy_contract # TL;DR > 200 chars fails validate
151
+ end
152
+
153
+ it "passes its eval" do
154
+ expect(described_class).to pass_eval("smoke")
155
+ end
156
+ end
157
+ ```
158
+
159
+ `pass_eval` supports a matcher chain — full reference lives in [Getting Started](getting_started.md) under Evals and CI gates. Quick summary:
160
+
161
+ - `.with_context(model: "gpt-4.1-mini")` — pick model / pass adapter
162
+ - `.with_minimum_score(0.8)` — gate on average score
163
+ - `.with_maximum_cost(0.01)` — gate on total cost
164
+ - `.without_regressions` — block any previously-passing case that now fails (reads the baseline)
165
+ - `.compared_with(SummarizeArticleV1)` — A/B against another step; implies regression check
166
+
167
+ ## Offline vs online eval
168
+
169
+ Evals run in one of two modes depending on how they're defined and what context is passed:
170
+
171
+ | Has `sample_response`? | Context has adapter/model? | Mode | API calls |
172
+ |---|---|---|---|
173
+ | Yes | No | **Offline** — uses `sample_response` as canned answer | Zero |
174
+ | Yes | Yes | **Online** — ignores `sample_response`, calls real LLM | Real |
175
+ | No | Yes | **Online** — calls real LLM | Real |
176
+ | No | No | **Skipped** — returns `:skipped`, excluded from score | Zero |
177
+
178
+ Default is offline. To force online, pass adapter or model in context:
179
+
180
+ ```ruby
181
+ # Online — real LLM call
182
+ report = SummarizeArticle.run_eval("regression", context: { model: "gpt-4.1-nano" })
183
+
184
+ # Offline — uses sample_response
185
+ report = SummarizeArticle.run_eval("smoke")
186
+ ```
187
+
188
+ `compare_with` intentionally ignores `sample_response` because canned data would make both sides look identical. Always pass `model:` or an adapter to A/B.
189
+
190
+ ## Inspecting failures
191
+
192
+ `run_eval` returns a `Report`. Drill into per-case failures:
193
+
194
+ ```ruby
195
+ report = SummarizeArticle.run_eval("regression")
196
+
197
+ report.score # => 0.5
198
+ report.pass_rate # => "1/2"
199
+ report.total_cost # => 0.003
200
+
201
+ report.failures.each do |result|
202
+ puts result.name # => "critical review"
203
+ puts result.mismatches # => { tone: { expected: "negative", got: "analytical" } }
204
+ puts result.output # full parsed output hash
205
+ puts result.details # human-readable explanation
206
+ end
207
+ ```
208
+
209
+ `mismatches` is a hash of keys where expected and actual output diverge — pinpoints which field the model got wrong.
210
+
211
+ ## Soft observations
212
+
213
+ Log suspicious-but-not-invalid output without failing the contract:
214
+
215
+ ```ruby
216
+ class CompareArticles < RubyLLM::Contract::Step::Base
217
+ prompt "Score the article pair for relevance. Return JSON: {score_a: 1-10, score_b: 1-10}.\n\n{input}"
218
+
219
+ output_schema do
220
+ integer :score_a, minimum: 1, maximum: 10
221
+ integer :score_b, minimum: 1, maximum: 10
222
+ end
223
+
224
+ validate("scores in range") { |o, _| (1..10).cover?(o[:score_a]) && (1..10).cover?(o[:score_b]) }
225
+ observe("scores should differ") { |o, _| o[:score_a] != o[:score_b] }
226
+ end
227
+
228
+ adapter = RubyLLM::Contract::Adapters::Test.new(response: { score_a: 5, score_b: 5 })
229
+ result = CompareArticles.run("two identical-looking articles", context: { adapter: adapter })
230
+
231
+ result.ok? # => true (observe never fails the contract)
232
+ result.observations # => [{ description: "scores should differ", passed: false }]
233
+ ```
234
+
235
+ `observe` runs only after validation passes. Failed observations are logged via `RubyLLM::Contract.logger` — useful for "I want to know this happened without blocking the response".
236
+
237
+ > **Not the same as `Chat#on_end_message` / `on_tool_call`.** RubyLLM exposes anonymous, global callbacks attached to a chat instance — they receive the raw `Message` and have no inherent pass/fail concept. `Step.observe` is per-step, named with a description, runs against `parsed_output` (already JSON-parsed and schema-validated), and the pass/fail outcome is recorded in `result.observations`. The two can coexist: use Chat callbacks for transport/tool tracing, use `observe` for domain assertions you want captured in the step trace.
238
+
239
+ ## Asserting on `around_call`
240
+
241
+ `around_call` fires **once per run** with the final result (after retry fallback) and exceptions propagate. That makes it straightforward to test:
242
+
243
+ ```ruby
244
+ class LoggedSummarize < RubyLLM::Contract::Step::Base
245
+ prompt "Summarize: {input}"
246
+ output_schema { string :tldr }
247
+
248
+ around_call do |_step, input, result|
249
+ CallLog.record(model: result.trace.model, cost: result.trace.cost, input_size: input.length)
250
+ end
251
+ end
252
+
253
+ RSpec.describe LoggedSummarize do
254
+ it "logs once per run, with final model + total cost" do
255
+ adapter = RubyLLM::Contract::Adapters::Test.new(response: { tldr: "ok" })
256
+ expect(CallLog).to receive(:record).once.with(hash_including(:model, :cost, :input_size))
257
+
258
+ LoggedSummarize.run("article text", context: { adapter: adapter })
259
+ end
260
+ end
261
+ ```
262
+
263
+ The callback receives `(step, input, result)` — the same `Result` the caller sees. Not invoked per-attempt inside a `retry_policy` chain; if you need per-attempt visibility, read `result.trace[:attempts]` inside the block.
264
+
265
+ ## Baseline file format
266
+
267
+ Baselines are JSON files in `.eval_baselines/` — commit them to git:
268
+
269
+ ```
270
+ .eval_baselines/
271
+ SummarizeArticle/
272
+ regression_gpt-4_1-nano.json
273
+ regression_gpt-4_1-mini.json
274
+ ```
275
+
276
+ Each file contains dataset name, step name, score, and per-case results. No timestamps — re-saving an identical baseline produces no git diff. Baseline semantics (what counts as a regression, how `compare_with_baseline` works) are covered in [Getting Started](getting_started.md#evals-and-ci-gates) and [Eval-First](eval_first.md).
277
+
278
+ ## See also
279
+
280
+ - [Getting Started](getting_started.md) — `pass_eval` matcher chain, threshold gating, Rake task, baseline regressions.
281
+ - [Eval-First](eval_first.md) — `compare_with` prompt A/B workflow.
282
+ - [Pipeline](pipeline.md) — pipeline-level testing with named step responses.
data/docs/guide/why.md ADDED
@@ -0,0 +1,103 @@
1
+ # Why contracts?
2
+
3
+ > Read this if you're not sure whether `ruby_llm-contract` solves a problem you actually have. It's the fastest way to recognise the production failure modes the gem exists for.
4
+
5
+ LLMs return JSON that *looks* correct — valid shape, right types, right fields — while being silently wrong in ways that hurt users, burn budget, or break downstream code. Schema validation alone does not catch these. Contracts layer on business rules, retries, evals, and cost caps so the wrong output is caught at the boundary of your system instead of shipping to production.
6
+
7
+ Below are four failure modes teams actually hit. If one looks familiar, the gem is probably worth 30 minutes of your time.
8
+
9
+ ## Failure 1 — Schema-valid, logically wrong
10
+
11
+ A `SummarizeArticle` step produces `{ tldr: "...", takeaways: [...], tone: "analytical" }`. Schema passes. The TL;DR is 520 characters long and overflows the UI card. Or the takeaways are all variations of the same sentence. Or the article was a service-outage complaint and `tone` came back `"analytical"` instead of `"negative"` — so customer success's "critical feedback" filter never sees it.
12
+
13
+ JSON schema enforces **shape**. It cannot enforce *length fits the card*, *takeaways are distinct*, or *tone matches content*. Those are business rules, and without them you find out from a Slack thread or a support ticket.
14
+
15
+ ```ruby
16
+ validate("TL;DR fits the card") { |o, _| o[:tldr].length <= 200 }
17
+ validate("takeaways are unique") { |o, _| o[:takeaways] == o[:takeaways].uniq }
18
+ validate("negative tone requires concrete risk") do |o, _|
19
+ next true unless o[:tone] == "negative"
20
+ o[:takeaways].any? { |t| t.match?(/fail|break|crash|outage|risk/i) }
21
+ end
22
+ ```
23
+
24
+ Wrong output never reaches `Article.update!` — the contract refuses before it persists.
25
+
26
+ ## Failure 2 — Silent prompt regression
27
+
28
+ `SummarizeArticle` ships and works. Two weeks later, someone tweaks the system prompt to emphasise negative sentiment because customer success complained about missed complaints. The tweak fixes that case and silently breaks three neutral product-update articles that now get labelled `"negative"`. Nobody knows for a week.
29
+
30
+ Without evals, *every prompt change is a blind deploy.* Contracts invert this:
31
+
32
+ ```ruby
33
+ SummarizeArticle.define_eval("regression") do
34
+ add_case "outage complaint", input: "...", expected: { tone: "negative" }
35
+ add_case "neutral product update", input: "...", expected: { tone: "neutral" }
36
+ end
37
+
38
+ # In CI — blocks merge when a prompt tweak regresses any previously-passing case
39
+ expect(SummarizeArticle).to pass_eval("regression").without_regressions
40
+ ```
41
+
42
+ The "tweak helped one case, broke three" scenario is caught at PR review. No Slack-thread surprises.
43
+
44
+ ## Failure 3 — Sampling variance on fixed-temperature models
45
+
46
+ OpenAI's gpt-5 and o-series run with `temperature=1.0` server-side — you cannot lower it. That means the same prompt on the same model can produce different answers between calls. An outage complaint classified `tone: "negative"` on Monday may come back `tone: "positive"` on Tuesday, with no code change in between. Schema passes both. Your customer-success filter silently misroutes the Tuesday case.
47
+
48
+ A `validate` block that cross-checks fields against each other turns a one-in-N flaky output into a deterministic retry:
49
+
50
+ ```ruby
51
+ validate("tone matches severity keywords") do |o, _|
52
+ severity = /fail|crash|outage|broken|bug|error/i
53
+ flagged = o[:takeaways].any? { |t| t.match?(severity) }
54
+ next true unless flagged
55
+ %w[negative analytical].include?(o[:tone])
56
+ end
57
+
58
+ retry_policy models: %w[gpt-5-nano gpt-5-mini gpt-5]
59
+ ```
60
+
61
+ Nano misclassifies the tone on the first attempt → contract rejects → mini gets the call and returns a different sample. Variance absorbed; the user never sees the flaky run. Your logs show the retry rate and the cost delta.
62
+
63
+ **See it in 30 seconds:** `ruby examples/01_fallback_showcase.rb` — zero API keys required. The Test adapter simulates a tone/takeaways mismatch on the first attempt and a consistent sample on the retry, then prints the per-attempt trace.
64
+
65
+ `retry_policy` has three other shapes beyond cross-model escalation — same-model `attempts: 3` (absorbs sampling variance without paying for a stronger tier), `reasoning_effort` escalation (low → medium → high on one model), and cross-provider fallback (Ollama → Anthropic → OpenAI — local first because it costs nothing, hosted last because it is the most accurate). `examples/06_retry_variants.rb` runs all three through the Test adapter with the trace printed.
66
+
67
+ ## Failure 4 — Runaway cost and no fallback policy
68
+
69
+ Someone pastes a 40-page PDF into the endpoint that calls `SummarizeArticle`. The prompt expands to 80k tokens. Your provider bill jumps. Meanwhile, a separate team uses `gpt-4.1` for every single call because "quality matters" — even though 80% of their traffic is trivially handled by `gpt-4.1-nano` at 1/30th the cost. Neither situation is visible until the invoice arrives.
70
+
71
+ Contracts make cost a first-class concern:
72
+
73
+ ```ruby
74
+ max_input 2_000 # refuses before calling the API if tokens exceed budget
75
+ max_output 4_000
76
+ max_cost 0.01 # refuses if estimated cost exceeds cap
77
+ retry_policy models: %w[gpt-4.1-nano gpt-4.1-mini gpt-4.1] # cheap first, escalate only on failure
78
+ ```
79
+
80
+ The 40-page PDF returns `status: :limit_exceeded` — zero tokens spent. The 80/20 traffic pattern resolves on nano; only the hard 20% escalate. [`optimize_retry_policy`](optimizing_retry_policy.md) tells you empirically which fallback list is cheapest for *your* evals.
81
+
82
+ ## Also catches
83
+
84
+ - **Leaked prompt placeholders** — model echoes `{article}` or `{audience}` into the output because the template string wasn't interpolated. Validate string-equality check stops it before a user sees it.
85
+ - **Lazy models echoing input verbatim** — cheap model returns the article text as the "summary". 2-arity `validate("tldr shorter than input")` catches it.
86
+ - **Tone mislabel breaking downstream routing** — content says "negative", model labels it "analytical", a customer success filter misses it. Cross-validate catches the label/content drift.
87
+
88
+ ## Failure → contract mechanism
89
+
90
+ | Failure in production | Contract mechanism |
91
+ |---|---|
92
+ | Schema-valid but logically wrong output | `validate(...) { |o, i| ... }` with 2-arity for cross-checks |
93
+ | Silent prompt regression after a tweak | `define_eval` + `pass_eval(...).without_regressions` in CI |
94
+ | Sampling variance on fixed-temperature models (gpt-5 / o-series) | Cross-field `validate(...)` + `retry_policy models: [...]` |
95
+ | Runaway cost on pathological inputs | `max_input`, `max_output`, `max_cost` preflight |
96
+ | 80/20 traffic paying the premium model rate | `retry_policy` + `optimize_retry_policy` |
97
+ | Leaked placeholder / input echo / tone drift | `validate` with content and cross-input checks |
98
+
99
+ ## What next
100
+
101
+ - **If one of the failures above looks familiar** → [Getting Started](getting_started.md) walks through every feature in order on the same `SummarizeArticle` step.
102
+ - **If you're adopting in an existing Rails app** → [Migration](migration.md) shows Before/After for replacing a raw `LlmClient.new.call` service.
103
+ - **If you already ship LLM features and want to make them regression-safe** → [Eval-First](eval_first.md) is the workflow that prevents Failure 2 from ever happening again.
@@ -13,7 +13,15 @@ module RubyLLM
13
13
  chat = build_chat(options, system_contents)
14
14
  add_history(chat, conversation[0..-2])
15
15
 
16
- response = chat.ask(conversation.last&.fetch(:content, ""))
16
+ # `with: nil` is a documented no-op in RubyLLM (verified against
17
+ # 1.15.0: chat.rb:36-37 `build_content(message, nil)` -> content.rb:8-14
18
+ # `Content.new(text, nil)` keeps text-only path when attachments
19
+ # are empty; raise only fires when BOTH text and attachments are nil,
20
+ # and we always pass a non-nil string thanks to `&.fetch(:content, "")`).
21
+ response = chat.ask(
22
+ conversation.last&.fetch(:content, ""),
23
+ with: options[:attachment]
24
+ )
17
25
  build_response(response)
18
26
  end
19
27
 
@@ -190,16 +190,13 @@ module RubyLLM
190
190
  context.merge(adapter: sample_adapter)
191
191
  end
192
192
 
193
+ # `Class#subclasses` is available from Ruby 3.1; the gemspec requires
194
+ # `>= 3.2.0` so the legacy `ObjectSpace.each_object` fallback would be
195
+ # dead code on every supported runtime.
193
196
  def register_subclasses(klass)
194
- if klass.respond_to?(:subclasses)
195
- klass.subclasses.each do |sub|
196
- Contract.register_eval_host(sub)
197
- register_subclasses(sub)
198
- end
199
- else
200
- ObjectSpace.each_object(Class) do |sub|
201
- Contract.register_eval_host(sub) if sub < klass
202
- end
197
+ klass.subclasses.each do |sub|
198
+ Contract.register_eval_host(sub)
199
+ register_subclasses(sub)
203
200
  end
204
201
  end
205
202
 
@@ -0,0 +1,97 @@
1
+ # frozen_string_literal: true
2
+
3
+ module RubyLLM
4
+ module Contract
5
+ module Concerns
6
+ # Shared implementation of `stub_step`, `stub_steps`, and `stub_all_steps`.
7
+ # Included by both `RubyLLM::Contract::RSpec::Helpers` and
8
+ # `RubyLLM::Contract::MinitestHelpers` so the two test-framework
9
+ # adapters cannot drift on stub semantics (Codex DRY finding #1: the
10
+ # prior parallel implementations had already diverged on
11
+ # `normalize_test_response` — RSpec had it, Minitest didn't).
12
+ #
13
+ # Cleanup between examples is the responsibility of the host helper:
14
+ # - RSpec: `around(:each)` hook in `lib/ruby_llm/contract/rspec.rb`
15
+ # restores `step_adapter_overrides`.
16
+ # - Minitest: `teardown` in `MinitestHelpers` clears overrides and
17
+ # restores `default_adapter`.
18
+ module StubHelpers
19
+ # Stub a single step to return a canned response without API calls.
20
+ # Block form scopes the stub to the block; non-block form lives
21
+ # until the host's teardown/around hook fires.
22
+ def stub_step(step_class, response: nil, responses: nil, &block)
23
+ adapter = build_test_adapter(response: response, responses: responses)
24
+ overrides = RubyLLM::Contract.step_adapter_overrides
25
+
26
+ if block
27
+ previous = overrides[step_class]
28
+ overrides[step_class] = adapter
29
+ begin
30
+ yield
31
+ ensure
32
+ previous ? (overrides[step_class] = previous) : overrides.delete(step_class)
33
+ end
34
+ else
35
+ overrides[step_class] = adapter
36
+ end
37
+ end
38
+
39
+ # Stub multiple steps with different responses. Requires a block.
40
+ def stub_steps(stubs, &block)
41
+ raise ArgumentError, "stub_steps requires a block" unless block
42
+
43
+ overrides = RubyLLM::Contract.step_adapter_overrides
44
+ previous = {}
45
+
46
+ stubs.each do |step_class, opts|
47
+ opts = opts.transform_keys(&:to_sym)
48
+ previous[step_class] = overrides[step_class]
49
+ overrides[step_class] = build_test_adapter(**opts.slice(:response, :responses))
50
+ end
51
+
52
+ begin
53
+ yield
54
+ ensure
55
+ stubs.each_key do |step_class|
56
+ previous[step_class] ? (overrides[step_class] = previous[step_class]) : overrides.delete(step_class)
57
+ end
58
+ end
59
+ end
60
+
61
+ # Set a global test adapter for ALL steps. Block form restores the
62
+ # previous adapter on exit; non-block form persists until host cleanup.
63
+ def stub_all_steps(response: nil, responses: nil, &block)
64
+ adapter = build_test_adapter(response: response, responses: responses)
65
+
66
+ if block
67
+ previous = RubyLLM::Contract.configuration.default_adapter
68
+ begin
69
+ RubyLLM::Contract.configuration.default_adapter = adapter
70
+ yield
71
+ ensure
72
+ RubyLLM::Contract.configuration.default_adapter = previous
73
+ end
74
+ else
75
+ RubyLLM::Contract.configure { |c| c.default_adapter = adapter }
76
+ end
77
+ end
78
+
79
+ private
80
+
81
+ def build_test_adapter(response: nil, responses: nil)
82
+ if responses
83
+ Adapters::Test.new(responses: responses.map { |r| normalize_test_response(r) })
84
+ else
85
+ Adapters::Test.new(response: normalize_test_response(response))
86
+ end
87
+ end
88
+
89
+ # Hook for host frameworks to inject custom serialization (e.g.
90
+ # turning hashes into JSON strings). Default: identity.
91
+ def normalize_test_response(value)
92
+ value
93
+ end
94
+ end
95
+ end
96
+ end
97
+ end
@@ -18,6 +18,8 @@ module RubyLLM
18
18
  end
19
19
 
20
20
  def invariant(description, &block)
21
+ raise ArgumentError, "invariant description must be a non-empty string", caller if description.to_s.empty?
22
+
21
23
  @invariants << Invariant.new(description, block)
22
24
  end
23
25
  alias validate invariant
@@ -89,8 +89,13 @@ module RubyLLM
89
89
  (input_cost + output_cost).round(6)
90
90
  end
91
91
 
92
+ # Provider pricing is denominated per 1M tokens; divide here to get
93
+ # the dollar cost for the actual usage count. Named constant for
94
+ # consistency with how RubyLLM and provider docs express prices.
95
+ TOKENS_PER_MILLION = 1_000_000.0
96
+
92
97
  def self.token_cost(tokens, price_per_million)
93
- (tokens || 0) * (price_per_million || 0) / 1_000_000.0
98
+ (tokens || 0) * (price_per_million || 0) / TOKENS_PER_MILLION
94
99
  end
95
100
 
96
101
  def self.find_model(model_name)
@@ -111,7 +116,11 @@ module RubyLLM
111
116
  end
112
117
  end
113
118
 
114
- private_class_method :compute_cost, :token_cost, :find_model, :validate_price!
119
+ # `find_model` is intentionally public: `Step::Base#estimate_cost` needs
120
+ # to inspect model pricing before invoking `calculate` (e.g., to short-
121
+ # circuit estimate when the model is unknown). Exposing it removes
122
+ # `CostCalculator.send(:find_model)` workarounds at call sites.
123
+ private_class_method :compute_cost, :token_cost, :validate_price!
115
124
  end
116
125
  end
117
126
  end
@@ -4,7 +4,9 @@ module RubyLLM
4
4
  module Contract
5
5
  module Eval
6
6
  class Recommender
7
- def initialize(comparison:, min_score:, min_first_try_pass_rate: 0.8, current_config: nil)
7
+ def initialize(comparison:, min_score:,
8
+ min_first_try_pass_rate: DEFAULT_MIN_FIRST_TRY_PASS_RATE,
9
+ current_config: nil)
8
10
  @comparison = comparison
9
11
  @min_score = min_score
10
12
  @min_first_try_pass_rate = min_first_try_pass_rate
@@ -98,7 +98,9 @@ module RubyLLM
98
98
  end
99
99
  end
100
100
 
101
- def initialize(step:, candidates:, context: {}, min_score: 0.95, runs: 1, production_mode: nil)
101
+ def initialize(step:, candidates:, context: {},
102
+ min_score: DEFAULT_MIN_SCORE,
103
+ runs: 1, production_mode: nil)
102
104
  @step = step
103
105
  @candidates = candidates
104
106
  @context = context
@@ -113,10 +115,19 @@ module RubyLLM
113
115
 
114
116
  score_matrix = {}
115
117
  evals.each do |eval_name|
116
- comparison = with_retry_disabled do
117
- @step.compare_models(eval_name, candidates: @candidates, context: @context,
118
- runs: @runs, production_mode: @production_mode)
119
- end
118
+ # `retry_policy_override: nil` in context disables the step's
119
+ # class-level retry policy for this comparison run — see
120
+ # step/base.rb#runtime_settings, which honours the key when
121
+ # present (even when value is nil). Replaces the prior
122
+ # `define_singleton_method(:retry_policy)` mutation, which was
123
+ # not thread-safe across concurrent optimizer calls.
124
+ comparison = @step.compare_models(
125
+ eval_name,
126
+ candidates: @candidates,
127
+ context: @context.merge(retry_policy_override: nil),
128
+ runs: @runs,
129
+ production_mode: @production_mode
130
+ )
120
131
  score_matrix[eval_name] = extract_scores(comparison)
121
132
  end
122
133
 
@@ -203,14 +214,6 @@ module RubyLLM
203
214
  end
204
215
  end
205
216
 
206
- def with_retry_disabled(&block)
207
- original = @step.retry_policy if @step.respond_to?(:retry_policy)
208
- @step.define_singleton_method(:retry_policy) { nil }
209
- block.call
210
- ensure
211
- @step.define_singleton_method(:retry_policy) { original }
212
- end
213
-
214
217
  def empty_result(evals)
215
218
  Result.new(
216
219
  step_name: @step.name || @step.to_s,
@@ -33,3 +33,16 @@ require_relative "eval/eval_history"
33
33
  require_relative "eval/recommendation"
34
34
  require_relative "eval/recommender"
35
35
  require_relative "eval/retry_optimizer"
36
+
37
+ module RubyLLM
38
+ module Contract
39
+ module Eval
40
+ # Default thresholds shared across `recommend`, `optimize_retry_policy`,
41
+ # `Recommender`, and `RetryOptimizer`. Centralised so a single change in
42
+ # what "viable" means (e.g. tightening from 0.95 to 0.97) propagates
43
+ # everywhere instead of needing the same edit in 4 places.
44
+ DEFAULT_MIN_SCORE = 0.95
45
+ DEFAULT_MIN_FIRST_TRY_PASS_RATE = 0.8
46
+ end
47
+ end
48
+ end