ruby_llm-contract 0.10.6 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. checksums.yaml +4 -4
  2. data/.ruby-version +1 -0
  3. data/CHANGELOG.md +354 -211
  4. data/README.md +12 -2
  5. data/docs/guide/getting_started.md +3 -1
  6. data/docs/guide/llm_judge.md +1 -1
  7. data/docs/guide/multimodal_input.md +4 -3
  8. data/docs/guide/output_schema.md +2 -2
  9. data/docs/guide/testing.md +1 -1
  10. data/examples/README.md +1 -1
  11. data/lib/ruby_llm/contract/adapters/ruby_llm.rb +13 -9
  12. data/lib/ruby_llm/contract/concerns/context_helpers.rb +0 -2
  13. data/lib/ruby_llm/contract/concerns/deep_symbolize.rb +0 -1
  14. data/lib/ruby_llm/contract/contract/parser.rb +0 -3
  15. data/lib/ruby_llm/contract/contract/schema_validator/bound_rule.rb +0 -1
  16. data/lib/ruby_llm/contract/contract/schema_validator/enum_rule.rb +0 -1
  17. data/lib/ruby_llm/contract/contract/schema_validator/node.rb +0 -4
  18. data/lib/ruby_llm/contract/contract/schema_validator/scalar_rules.rb +0 -1
  19. data/lib/ruby_llm/contract/contract/schema_validator/type_rule.rb +0 -1
  20. data/lib/ruby_llm/contract/cost_calculator.rb +28 -2
  21. data/lib/ruby_llm/contract/eval/aggregated_report.rb +4 -0
  22. data/lib/ruby_llm/contract/eval/baseline_diff.rb +22 -3
  23. data/lib/ruby_llm/contract/eval/candidate_label.rb +31 -0
  24. data/lib/ruby_llm/contract/eval/case_executor.rb +2 -2
  25. data/lib/ruby_llm/contract/eval/case_result.rb +17 -3
  26. data/lib/ruby_llm/contract/eval/case_result_builder.rb +2 -1
  27. data/lib/ruby_llm/contract/eval/contract_detail_builder.rb +0 -2
  28. data/lib/ruby_llm/contract/eval/model_comparison.rb +3 -2
  29. data/lib/ruby_llm/contract/eval/pipeline_result_adapter.rb +1 -2
  30. data/lib/ruby_llm/contract/eval/prompt_diff_comparator.rb +0 -1
  31. data/lib/ruby_llm/contract/eval/prompt_diff_presenter.rb +0 -1
  32. data/lib/ruby_llm/contract/eval/prompt_diff_serializer.rb +0 -1
  33. data/lib/ruby_llm/contract/eval/report.rb +1 -1
  34. data/lib/ruby_llm/contract/eval/report_presenter.rb +0 -1
  35. data/lib/ruby_llm/contract/eval/report_stats.rb +5 -1
  36. data/lib/ruby_llm/contract/eval/report_storage.rb +0 -1
  37. data/lib/ruby_llm/contract/eval/retry_optimizer.rb +10 -20
  38. data/lib/ruby_llm/contract/eval/step_expectation_applier.rb +2 -1
  39. data/lib/ruby_llm/contract/eval/trait_evaluator.rb +0 -2
  40. data/lib/ruby_llm/contract/eval/unknown_cost_gate.rb +40 -0
  41. data/lib/ruby_llm/contract/eval.rb +2 -0
  42. data/lib/ruby_llm/contract/minitest.rb +6 -1
  43. data/lib/ruby_llm/contract/pipeline/runner.rb +0 -1
  44. data/lib/ruby_llm/contract/pipeline/trace.rb +7 -0
  45. data/lib/ruby_llm/contract/rake_task/suite_gate.rb +18 -21
  46. data/lib/ruby_llm/contract/rake_task.rb +9 -4
  47. data/lib/ruby_llm/contract/rspec/pass_eval.rb +80 -52
  48. data/lib/ruby_llm/contract/step/base.rb +12 -3
  49. data/lib/ruby_llm/contract/step/dsl.rb +2 -4
  50. data/lib/ruby_llm/contract/step/limit_checker.rb +0 -2
  51. data/lib/ruby_llm/contract/step/retry_executor.rb +0 -2
  52. data/lib/ruby_llm/contract/step/trace.rb +19 -0
  53. data/lib/ruby_llm/contract/token_estimator.rb +0 -2
  54. data/lib/ruby_llm/contract/version.rb +1 -1
  55. data/lib/ruby_llm/contract.rb +0 -4
  56. data/ruby_llm-contract.gemspec +6 -5
  57. metadata +9 -11
  58. data/.rubocop.yml +0 -58
  59. data/Gemfile +0 -13
  60. data/Gemfile.lock +0 -278
  61. data/Rakefile +0 -8
data/CHANGELOG.md CHANGED
@@ -1,13 +1,156 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.1.0 (2026-10-10)
4
+
5
+ ### Changed - may turn a green CI red
6
+
7
+ - **Eval cost gates now fail closed on unknown pricing.** `pass_eval(...).with_maximum_cost`,
8
+ `RakeTask#maximum_cost` and `assert_eval_passes(maximum_cost:)` compared the report's
9
+ `total_cost`, which counts a case it cannot price as $0. An eval running on a model
10
+ without pricing data (a local or fine-tuned model) therefore passed any budget. The
11
+ gates now fail and name the unpriced cases, matching what step-level `max_cost`
12
+ already did. Opt out with `.on_unknown_pricing(:warn)`, `t.on_unknown_pricing = :warn`
13
+ or `on_unknown_pricing: :warn`: the budget is then checked against the priced cases
14
+ and a warning lists the rest. Offline runs (`sample_response`, a `Test` adapter
15
+ without `usage:`) report zero tokens and are never flagged. Only gates with a
16
+ maximum cost set are affected.
17
+
18
+ ### Added
19
+
20
+ - `Report#unknown_cost_results` (also on `AggregatedReport`), `CaseResult#cost_unknown?`,
21
+ and `cost_unknown?` on step and pipeline traces. A retried step is judged per attempt,
22
+ so a priced subtotal no longer hides an unpriced attempt. `CaseResult#to_h` gains a
23
+ `cost_unknown: true` key only for flagged cases; baseline files are unchanged.
24
+
25
+ ### Fixed
26
+
27
+ - **A previously passing case that was skipped (no adapter) is reported as `SKIPPED`,
28
+ not `REGRESSED`.** The regression gate still fails on it - a skip is no evidence the
29
+ case still passes - but the report no longer showed an unchanged score next to a
30
+ "regression". Applies to `BaselineDiff#to_s` (rake task output) and the
31
+ `without_regressions` failure message; new `BaselineDiff#skipped_passing_cases`.
32
+ - A code comment in `CostCalculator` claimed an unreadable pricing shape refuses budgeted
33
+ calls "with no error"; the refusal does carry an error, it just names missing pricing
34
+ as the cause.
35
+
36
+ ## 1.0.0 (2026-10-10)
37
+
38
+ **Requires `ruby_llm ~> 2.0`.** This is the only breaking change: it is a
39
+ dependency requirement, not a change to this gem's API. Your call sites,
40
+ `output_schema` blocks, `trace[:usage]` keys and adapter interface are all
41
+ untouched - if you are on ruby_llm 2.x, upgrading should need no code edits.
42
+
43
+ 1.0 also states what is now stable: the documented DSL, the `Result`/`trace`
44
+ shapes including `trace[:usage]`'s `{ input_tokens:, output_tokens: }` keys, and
45
+ the adapter interface. A future major version of a runtime dependency may again
46
+ require a major version here - that is what this release is.
47
+
48
+ ### Added
49
+
50
+ - **CI, for the first time.** Matrix over Ruby 3.2 and 3.4 against both
51
+ ruby_llm 2.0.0 (pinned, so a regression at the major boundary stays
52
+ distinguishable) and the newest 2.x the gemspec allows. The job also runs every
53
+ offline example, builds the gem and fails on any `gem build` warning, and
54
+ installs the built gem to prove it `require`s. The seven live-key specs are
55
+ bounded explicitly: CI asserts the pending count is exactly 7 and that all
56
+ seven come from `spec/integration/cost_of_quality_real_spec.rb`, so a failing
57
+ spec cannot be quietly downgraded to pending.
58
+ - **`.rubocop_todo.yml`** so lint can be enforced in CI at zero offences while the
59
+ pre-existing backlog stays visible and shrinkable, rather than blocking the gate.
60
+ - **A boundary spec for the RubyLLM 2.x seam**
61
+ (`spec/ruby_llm/contract/adapters/ruby_llm_2x_boundary_spec.rb`). Every
62
+ upstream-version-specific detail is supposed to live in `Adapters::RubyLLM` and
63
+ `CostCalculator`; these tests pin that, including that `max_cost` still fails
64
+ closed if provider pricing ever becomes unreachable again.
65
+
66
+ ### Changed
67
+
68
+ - **`ruby_llm` requirement raised to `~> 2.0`** (from `~> 1.12`).
69
+ - **`ruby_llm-schema` replaced by `schematist ~> 1.1`.** `ruby_llm-schema` 1.0.0 is
70
+ 31 lines that emit a deprecation warning and alias `RubyLLM::Schema =
71
+ Schematist::Schema`; ruby_llm 2.x depends on schematist directly. The
72
+ `output_schema do ... end` DSL is unchanged - same object behind a different
73
+ constant.
74
+ - **Adapter reads token counts from `Message#tokens`.** RubyLLM 2.0 replaced
75
+ `Message#input_tokens`/`#output_tokens` with a `Tokens` value object. This is the
76
+ only place in the gem that touches upstream token counts, and the normalised
77
+ `{ input_tokens:, output_tokens: }` hash it publishes is deliberately unchanged.
78
+ - **`max_tokens` now forwards via `with_max_output_tokens`**; `Chat#with_params` no
79
+ longer exists in 2.0.
80
+
81
+ ### Fixed
82
+
83
+ - **`estimate_eval_cost` raised `TypeError` instead of flooring at $0.00.** With
84
+ ruby_llm 2.x, `CostCalculator#calculate` returns nil for two distinct misses -
85
+ model absent from the registry, and model present but with unreadable pricing -
86
+ and the summation guarded only the first. Measured against the real 2.0.0
87
+ registry, 669 of 1672 models take the second path, so any eval estimate touching
88
+ one of them crashed. `getting_started.md` documents this method as a floor, not
89
+ a fail-closed, and it behaves that way again.
90
+ - **The optimizer's candidate table could mangle a model name containing
91
+ brackets.** The short label was built by cutting the rendered label apart, so
92
+ `"custom (beta) (effort: high)"` came out as `"custom (beta@high)"` - the
93
+ model's own closing bracket was consumed. The notation now has one owner
94
+ (`Eval::CandidateLabel`, which both renders and parses), and the short form is
95
+ composed from parsed parts instead of string surgery.
96
+ - **A reworded skipped-case reason could have reported a phantom regression.**
97
+ `BaselineDiff` excludes skipped cases from the score denominator by matching a
98
+ string prefix that `CaseExecutor` writes, and the two sides are filtered
99
+ differently - the baseline side structurally, the current side only by that
100
+ string. Both ends now read one constant, and a baseline file written by 0.10.6
101
+ is covered by its own spec (the fixture was generated by installing 0.10.6, not
102
+ written by hand).
103
+ - **Cost calculation against ruby_llm 2.x.** 2.0 dropped the flat
104
+ `input_price_per_million` / `output_price_per_million` readers on `Model` for a
105
+ nested `pricing -> text_tokens -> standard` walk. Because `CostCalculator`
106
+ rescued `StandardError` into a nil cost, every priced model silently became
107
+ "unknown pricing" - and since `max_cost` fails closed on unknown pricing, every
108
+ budgeted call would have been refused. Pricing extraction is now explicit for
109
+ both shapes (upstream models and locally `register_model`-ed ones), and an
110
+ unrecognised shape returns nil rather than a swallowed exception.
111
+ - **A shipped doc example taught the wrong test double.**
112
+ `docs/guide/multimodal_input.md` showed `double(content:, input_tokens:,
113
+ output_tokens:)`, which a verified double rejects against ruby_llm 2.0.
114
+ - **`FINDING 7` in the audit specs asserted a constraint on `ruby_llm-schema`**, a
115
+ dependency this release removes. Generalised to "no runtime dependency is
116
+ declared without a version constraint", so the check cannot be retired by
117
+ renaming the gem it happened to be about.
118
+ - Stale version and dependency claims in `README.md`, `docs/guide/output_schema.md`
119
+ and `docs/guide/multimodal_input.md`.
120
+
121
+ ### Housekeeping (was staged as 0.11.0)
122
+
123
+ Result of a 20-pass systematic audit of the whole repository. No API changes and no behaviour change inside `Step`/`Pipeline`/`Eval`; the user-visible differences are in what the gem **packages**, what the RubyGems page **links**, and what the docs **claim**.
124
+
125
+ ### Changed
126
+
127
+ - **The published gem no longer ships development-only files.** Measured by building the gem and unpacking it: `.rubocop.yml`, `Gemfile`, `Gemfile.lock` and `Rakefile` were being packaged, contradicting the gemspec's own stated policy ("dev configs excluded so the published gem contains only what adopters actually need at runtime"). The shipped `Rakefile` was unusable anyway - it `require`s rspec, which is not a runtime dependency, and defines a task over `spec/`, which the same build excludes. Package contents: **137 → 133 files**; `lib/` (105) and `docs/` (16) unchanged.
128
+ - **`Gemfile.lock` is no longer tracked.** It was in the index despite being matched by `.gitignore:10`, and was repacked on every release. A library should not pin its adopters' dependency versions.
129
+ - **The RubyGems page will now show a "Source Code" link.** `metadata["homepage_uri"]` duplicated `spec.homepage` and, being listed first, suppressed `source_code_uri` in the UI - `gem build` warned about this on every build. Dropped the duplicate; `source_code_uri` stays, because `bundle info` and the RubyGems API read it as a distinct field.
130
+
131
+ ### Fixed
132
+
133
+ - **`.rubocop.yml` declared `AllCops` twice.** Psych applies last-key-wins, so `TargetRubyVersion: 3.2`, `NewCops: enable` and `SuggestExtensions: false` were silently discarded, and the second block's `Exclude` **replaced** RuboCop's defaults instead of adding to them. Effect: the linter scanned 353 files / 2423 offences, 147 of those files untracked under `tmp/` (they carried 2155 of the offences). Merging the blocks alone does not fix it - `inherit_mode: merge` is what restores the defaults. Now 206 files / 300 offences, all pre-existing formatting on real code. (Both counts measured with rubocop 1.88.0; `Gemfile.lock` is no longer tracked, so pin the tool when reproducing.)
134
+ - **`docs/guide/multimodal_input.md` linked to a file that could never resolve** - `../decisions/ADR-0022-*.md` points at `docs/decisions/` (no such directory; the ADR lives in `doc/decisions/`, singular) and that path is gitignored. Since `docs/guide/*` is packaged, this reproduced exactly the "links 404'd locally" defect that 0.10.2 set out to fix. Same dangling reference removed from the 0.9.0 CHANGELOG entry.
135
+ - **Two documented rake-task environment variables never existed.** `REASONING_EFFORT=` and `OLLAMA_API_BASE=` were listed in the 0.5.0 CHANGELOG entry, but `git log --all -S<name> -- lib/` is empty for both - they were never read at any point in the project's history, and ruby_llm's own configuration only seeds `RUBYLLM_DEBUG` / `RUBYLLM_STREAM_DEBUG` from the environment. Setting either was a silent no-op. Reasoning effort is selected with `CANDIDATES=model@effort`; Ollama's base URL goes through `RubyLLM.configure`.
136
+ - **README stated the wrong current version** (0.10.4 while `version.rb` was 0.10.6) - the one sentence that has to change every release.
137
+ - **`examples/README.md` claimed every example carries an "Expected output" section**; four of seven do, for three different reasons, so the sentence was the error rather than the files.
138
+ - **`docs/architecture.md` was reachable from nowhere** - the only one of 16 tracked docs with no inbound link. Added to the README index.
139
+
140
+ ### Removed
141
+
142
+ - Dead code with a full evidence battery behind each cut: `SchemaValidator::Node#numeric?` (zero readers; the one place needing the check inlines `value.is_a?(Numeric)`), an abandoned twin test helper whose sibling has 14 call sites, three unused `let`/helper definitions, four useless assignments, seven redundant `require`s in specs, and a `SimpleCov.start` that was a measured no-op because `require "simplecov"` already autoloads `.simplecov`.
143
+ - 31 WHAT comment blocks (51 lines) in `lib/` plus 20 comment lines in `spec/`, all restating the line or class below them - including all six `Extracted from ...` comments, five of them the "to reduce class length" boilerplate (refactor history git already records) and ten one-line class docstrings paraphrasing their own class name. WHY comments - invariants, external-behaviour notes, anti-duplication markers - were deliberately kept; so were the three evaluator banners that name a comparison semantic the class name cannot convey.
144
+ - Private-project residue: a spec block still labelled "reddit promo planner shape" although that project was removed as private in #25, and four dangling `(Batch N / TODO)` references to a `TODO.md` deleted in 0.10.1.
145
+
3
146
  ## 0.10.6 (2026-06-11)
4
147
 
5
- Docs accuracy patch: correct the positioning of `llm_judge.md` against `ruby_llm-tribunal` after a deeper audit of Tribunal's documented scope. The previous wording ("you are reinventing what `ruby_llm-tribunal` ships as a built-in catalog") read as if Tribunal made `llm_judge.md` redundant — incorrect. Tribunal's README ships an off-the-shelf implementation catalog (`assert_faithful`, `assert_hallucination`, `assert_refusal`, `assert_no_pii`, etc.) but does **not** document calibration workflow, per-claim breakdown, judge-prompt iteration, or judge-as-`evaluator:` integration — exactly the methodology `llm_judge.md` covers. The two are complementary layers, not alternatives. No code behaviour change.
148
+ Docs accuracy patch: correct the positioning of `llm_judge.md` against `ruby_llm-tribunal` after a deeper audit of Tribunal's documented scope. The previous wording ("you are reinventing what `ruby_llm-tribunal` ships as a built-in catalog") read as if Tribunal made `llm_judge.md` redundant - incorrect. Tribunal's README ships an off-the-shelf implementation catalog (`assert_faithful`, `assert_hallucination`, `assert_refusal`, `assert_no_pii`, etc.) but does **not** document calibration workflow, per-claim breakdown, judge-prompt iteration, or judge-as-`evaluator:` integration - exactly the methodology `llm_judge.md` covers. The two are complementary layers, not alternatives. No code behaviour change.
6
149
 
7
150
  ### Fixed
8
151
 
9
- - **`docs/guide/relation_to_tribunal.md`** — added the **"Tribunal's catalog vs Contract's `llm_judge.md` — concrete decision tree"** sub-section under "When to use which", giving adopters a sharp three-way fork: reach for Tribunal's catalog when the check is domain-general (faithfulness vs context, hallucination, refusal, PII, jailbreak, toxicity, bias); build a custom judge per `llm_judge.md` when the criterion is domain-specific, when the judge needs to live inside a `define_eval` regression gate as the `evaluator:` lambda, when per-claim sentence-level debug output is required, or when the judge prompt itself needs to be iterated against your data; use both for the same project at different lifecycle stages (Tribunal at spec-time, calibrated custom judge at CI merge gate). Added the **"What Tribunal documents — and what it doesn't"** sub-section with a five-row comparison table making explicit which methodology gaps `llm_judge.md` covers that Tribunal's README leaves to the adopter (calibration against humans, prompt iteration on over-flag, per-claim breakdown, evaluator-lambda integration, the six anti-patterns).
10
- - **`docs/guide/llm_judge.md`** — rewrote the closing "When to escalate to Tribunal's catalog" section as **"When to reach for Tribunal instead"**: shorter, accurate (Tribunal is a complementary catalog, not a replacement), points to `relation_to_tribunal.md` for the full decision tree and integration patterns. The previous wording implied that building any of the four standard judges (faithful / hallucination / refusal / PII) was "reinvention" — true for the **implementation** (Tribunal ships them), false for the **methodology** (Tribunal's README doesn't document calibration, anti-patterns, or per-claim breakdown). The methodology applies equally to Tribunal's built-ins, Tribunal's custom registered judges, and Contract `Step::Base` judges.
152
+ - **`docs/guide/relation_to_tribunal.md`** - added the **"Tribunal's catalog vs Contract's `llm_judge.md` - concrete decision tree"** sub-section under "When to use which", giving adopters a sharp three-way fork: reach for Tribunal's catalog when the check is domain-general (faithfulness vs context, hallucination, refusal, PII, jailbreak, toxicity, bias); build a custom judge per `llm_judge.md` when the criterion is domain-specific, when the judge needs to live inside a `define_eval` regression gate as the `evaluator:` lambda, when per-claim sentence-level debug output is required, or when the judge prompt itself needs to be iterated against your data; use both for the same project at different lifecycle stages (Tribunal at spec-time, calibrated custom judge at CI merge gate). Added the **"What Tribunal documents - and what it doesn't"** sub-section with a five-row comparison table making explicit which methodology gaps `llm_judge.md` covers that Tribunal's README leaves to the adopter (calibration against humans, prompt iteration on over-flag, per-claim breakdown, evaluator-lambda integration, the six anti-patterns).
153
+ - **`docs/guide/llm_judge.md`** - rewrote the closing "When to escalate to Tribunal's catalog" section as **"When to reach for Tribunal instead"**: shorter, accurate (Tribunal is a complementary catalog, not a replacement), points to `relation_to_tribunal.md` for the full decision tree and integration patterns. The previous wording implied that building any of the four standard judges (faithful / hallucination / refusal / PII) was "reinvention" - true for the **implementation** (Tribunal ships them), false for the **methodology** (Tribunal's README doesn't document calibration, anti-patterns, or per-claim breakdown). The methodology applies equally to Tribunal's built-ins, Tribunal's custom registered judges, and Contract `Step::Base` judges.
11
154
 
12
155
  ## 0.10.5 (2026-06-11)
13
156
 
@@ -15,21 +158,21 @@ Docs release: new `llm_judge.md` guide + comprehensive clarity audit across all
15
158
 
16
159
  ### Added
17
160
 
18
- - **`docs/guide/llm_judge.md`** — new guide for the LLM-as-judge pattern. Three-step workflow (build judge as a Contract Step → calibrate against humans on real production data → use as eval gate). Hook is the Cursor "Sam" April 2025 incident (support chatbot hallucinated company policy, schema valid, brand damage). Covers claims-breakdown alternative output schema, the "calibration is itself iterative" failure mode (over-flagging stylistic courtesy), six anti-patterns (stubbing the judge, calibrating on synthetic data, per-language regex, judge on request path, proxy label assertion, calibrating once and shipping), and an escalation pointer to `ruby_llm-tribunal`'s catalog. 206 lines. Linked from `eval_first.md` (oracle-validation rule) and `relation_to_tribunal.md` (composition recipe). Cited research: Eugene Yan on evals, Hamel Husain field guide, May 2025 LLM-hallucination court-case database, Klarna AI customer-service reversal.
161
+ - **`docs/guide/llm_judge.md`** - new guide for the LLM-as-judge pattern. Three-step workflow (build judge as a Contract Step → calibrate against humans on real production data → use as eval gate). Hook is the Cursor "Sam" April 2025 incident (support chatbot hallucinated company policy, schema valid, brand damage). Covers claims-breakdown alternative output schema, the "calibration is itself iterative" failure mode (over-flagging stylistic courtesy), six anti-patterns (stubbing the judge, calibrating on synthetic data, per-language regex, judge on request path, proxy label assertion, calibrating once and shipping), and an escalation pointer to `ruby_llm-tribunal`'s catalog. 206 lines. Linked from `eval_first.md` (oracle-validation rule) and `relation_to_tribunal.md` (composition recipe). Cited research: Eugene Yan on evals, Hamel Husain field guide, May 2025 LLM-hallucination court-case database, Klarna AI customer-service reversal.
19
162
 
20
163
  ### Fixed
21
164
 
22
- - **`docs/guide/migration.md` — `save_baseline = false` in CI Rakefile.** Pre-0.10.5 the migration template set `save_baseline = true`, which dirties the working checkout on every CI run and races the regression check — directly violating the project invariant ("`save_baseline: true` in CI dirties checkout and races the regression check — keep it false; refresh in a separate job"). Adopters copy-pasting the template hit non-deterministic CI failures + git status noise. The corrected template now sets `false` with an inline comment explaining the baseline-refresh-in-separate-workflow rationale.
23
- - **`docs/guide/relation_to_tribunal.md` — evaluator lambda arity corrected from 3 to 1.** The example previously used `->(output, _expected, _input)`, but `ProcEvaluator` only accepts arity 1 or 2 — adopters copy-pasting hit `ArgumentError: wrong number of arguments`. Now reads `->(output)`.
24
- - **`docs/guide/eval_first.md`** — added explicit definition of "partial match" (`expected:` is treated as a subset of `parsed_output`; extra output keys ignored, listed keys must equal), added definition of `adapter` (the layer that actually executes the LLM call), clarified `system`/`rule`/`example`/`validate` as prompt-shaping building blocks with a `prompt_ast.md` link.
25
- - **`docs/guide/getting_started.md`** — explained the `{X}` prompt template syntax (gem-specific, not ERB, not `String#%`), described `validate(...) { |o, _| ... }` arity convention, named the retry triggers `:validation_failed` / `:parse_error`, differentiated `default_input` from `add_case input:`, explained partial-match semantics.
26
- - **`README.md`** — three clarity edits to the SummarizeArticle example: `{input}` placeholder syntax explained, `validate` lambda args (`|o, _|`) explained, `retry_policy do escalate(...) end` block form aliased to `retry_policy models: %w[...]` shorthand (both forms share the same DSL — `models` is an alias for `escalate`). Plus the multimodal upgrade FAQ entry was compressed from 7 sentences to a 2-sentence SEO pointer to `multimodal_input.md` (where the full `attachment_token_estimate` setup, fail-closed behaviour, and `on_unknown_attachment_size :warn` opt-out already lived canonically) — reduces README cognitive load for the 90% of adopters not upgrading from pre-0.10.0.
27
- - **`docs/guide/testing.md`** — clarified pipeline test responses (`responses: { :summarize, :tag, :card }` keys must match `add_step :name, ...` identifiers), distinguished `validate` block (Step-level) from `verify` block (eval-case evaluator declared via `verify(name) { |output| ... }` inside `define_eval`), explained `stub_all_steps(response: { ... })` shape semantics.
28
- - **`docs/guide/best_practices.md`** — explained `rule` as a prompt DSL element distinct from `system` / `user` with a link to `prompt_ast.md`.
29
- - **`docs/guide/optimizing_retry_policy.md`** — explained `gpt-4.1-mini@low` CLI shorthand for `{model:, reasoning_effort:}`, defined `production_mode: { fallback: ... }` as an optional `compare_models` kwarg that reports effective cost (first-try + weighted fallback) instead of first-attempt only, compressed the "two orthogonal dimensions" callout from 7 concepts to 3 (DSL alias details moved to the dedicated `thinking` DSL note at the end).
30
- - **`docs/guide/output_schema.md`** — compressed the `RubyLLM::Agent.schema` boundary callout from 6 concepts to 3 + pointer (full Agent-vs-Step comparison lives canonically in `relation_to_agent.md`).
31
- - **`docs/guide/pipeline.md`** — documented that `Pipeline.run_eval` matches **only** the final step's output against `expected:` (assertions on intermediate-step outputs silently never match — gem-level invariant).
32
- - **`docs/guide/prompt_ast.md`** — explained `Types::Hash.schema(...)` typed-input declaration as distinct from plain `input_type Hash` (typed form raises `TypeError` on missing/wrong-type keys; plain form accepts any hash).
165
+ - **`docs/guide/migration.md` - `save_baseline = false` in CI Rakefile.** Pre-0.10.5 the migration template set `save_baseline = true`, which dirties the working checkout on every CI run and races the regression check - directly violating the project invariant ("`save_baseline: true` in CI dirties checkout and races the regression check - keep it false; refresh in a separate job"). Adopters copy-pasting the template hit non-deterministic CI failures + git status noise. The corrected template now sets `false` with an inline comment explaining the baseline-refresh-in-separate-workflow rationale.
166
+ - **`docs/guide/relation_to_tribunal.md` - evaluator lambda arity corrected from 3 to 1.** The example previously used `->(output, _expected, _input)`, but `ProcEvaluator` only accepts arity 1 or 2 - adopters copy-pasting hit `ArgumentError: wrong number of arguments`. Now reads `->(output)`.
167
+ - **`docs/guide/eval_first.md`** - added explicit definition of "partial match" (`expected:` is treated as a subset of `parsed_output`; extra output keys ignored, listed keys must equal), added definition of `adapter` (the layer that actually executes the LLM call), clarified `system`/`rule`/`example`/`validate` as prompt-shaping building blocks with a `prompt_ast.md` link.
168
+ - **`docs/guide/getting_started.md`** - explained the `{X}` prompt template syntax (gem-specific, not ERB, not `String#%`), described `validate(...) { |o, _| ... }` arity convention, named the retry triggers `:validation_failed` / `:parse_error`, differentiated `default_input` from `add_case input:`, explained partial-match semantics.
169
+ - **`README.md`** - three clarity edits to the SummarizeArticle example: `{input}` placeholder syntax explained, `validate` lambda args (`|o, _|`) explained, `retry_policy do escalate(...) end` block form aliased to `retry_policy models: %w[...]` shorthand (both forms share the same DSL - `models` is an alias for `escalate`). Plus the multimodal upgrade FAQ entry was compressed from 7 sentences to a 2-sentence SEO pointer to `multimodal_input.md` (where the full `attachment_token_estimate` setup, fail-closed behaviour, and `on_unknown_attachment_size :warn` opt-out already lived canonically) - reduces README cognitive load for the 90% of adopters not upgrading from pre-0.10.0.
170
+ - **`docs/guide/testing.md`** - clarified pipeline test responses (`responses: { :summarize, :tag, :card }` keys must match `add_step :name, ...` identifiers), distinguished `validate` block (Step-level) from `verify` block (eval-case evaluator declared via `verify(name) { |output| ... }` inside `define_eval`), explained `stub_all_steps(response: { ... })` shape semantics.
171
+ - **`docs/guide/best_practices.md`** - explained `rule` as a prompt DSL element distinct from `system` / `user` with a link to `prompt_ast.md`.
172
+ - **`docs/guide/optimizing_retry_policy.md`** - explained `gpt-4.1-mini@low` CLI shorthand for `{model:, reasoning_effort:}`, defined `production_mode: { fallback: ... }` as an optional `compare_models` kwarg that reports effective cost (first-try + weighted fallback) instead of first-attempt only, compressed the "two orthogonal dimensions" callout from 7 concepts to 3 (DSL alias details moved to the dedicated `thinking` DSL note at the end).
173
+ - **`docs/guide/output_schema.md`** - compressed the `RubyLLM::Agent.schema` boundary callout from 6 concepts to 3 + pointer (full Agent-vs-Step comparison lives canonically in `relation_to_agent.md`).
174
+ - **`docs/guide/pipeline.md`** - documented that `Pipeline.run_eval` matches **only** the final step's output against `expected:` (assertions on intermediate-step outputs silently never match - gem-level invariant).
175
+ - **`docs/guide/prompt_ast.md`** - explained `Types::Hash.schema(...)` typed-input declaration as distinct from plain `input_type Hash` (typed form raises `TypeError` on missing/wrong-type keys; plain form accepts any hash).
33
176
 
34
177
  ### Audit method
35
178
 
@@ -41,23 +184,23 @@ Patch release: the `ruby_llm_contract:optimize` rake task now auto-loads in Rail
41
184
 
42
185
  ### Fixed
43
186
 
44
- - **`ruby_llm_contract:optimize` no longer requires manual `require "ruby_llm/contract/rake_task"` in Rails apps.** Pre-0.10.4 the docs claimed the task was "included" with `RubyLLM::Contract::RakeTask`, but the railtie did not load the file — adopters running `bin/rails ruby_llm_contract:optimize` got `Unrecognized command` until they added the require to `Rakefile` or `lib/tasks/*.rake`. The railtie now uses the standard `rake_tasks { require "..." }` idiom, lazy-loading the file only when `rake` is invoked (no boot cost).
45
- - **Docs corrected** in `docs/guide/optimizing_retry_policy.md` — the "rake task is included" line now explicitly states it auto-loads on 0.10.4+ in Rails, and that non-Rails / older-Rails setups still need the explicit `require`.
187
+ - **`ruby_llm_contract:optimize` no longer requires manual `require "ruby_llm/contract/rake_task"` in Rails apps.** Pre-0.10.4 the docs claimed the task was "included" with `RubyLLM::Contract::RakeTask`, but the railtie did not load the file - adopters running `bin/rails ruby_llm_contract:optimize` got `Unrecognized command` until they added the require to `Rakefile` or `lib/tasks/*.rake`. The railtie now uses the standard `rake_tasks { require "..." }` idiom, lazy-loading the file only when `rake` is invoked (no boot cost).
188
+ - **Docs corrected** in `docs/guide/optimizing_retry_policy.md` - the "rake task is included" line now explicitly states it auto-loads on 0.10.4+ in Rails, and that non-Rails / older-Rails setups still need the explicit `require`.
46
189
 
47
190
  ## 0.10.3 (2026-06-10)
48
191
 
49
- Hot-fix release: the schema validator now correctly accepts `nil` on **required-but-nullable** fields. This unblocks OpenAI structured-output strict mode, where every property has to be in `required` and "nullable" is expressed as a `null` branch in `anyOf`/`oneOf` or as an array `type` — exactly the combination the prior validator wrongly rejected.
192
+ Hot-fix release: the schema validator now correctly accepts `nil` on **required-but-nullable** fields. This unblocks OpenAI structured-output strict mode, where every property has to be in `required` and "nullable" is expressed as a `null` branch in `anyOf`/`oneOf` or as an array `type` - exactly the combination the prior validator wrongly rejected.
50
193
 
51
194
  ### Fixed
52
195
 
53
- - **`SchemaValidator` no longer rejects legal `nil` on required-nullable fields.** Pre-0.10.3 (every published version 0.2.x–0.10.2) conflated "required" (must be present) with "non-nullable" (cannot be null) — JSON Schema treats those as orthogonal. The validator now honours all three idioms for nullability:
196
+ - **`SchemaValidator` no longer rejects legal `nil` on required-nullable fields.** Pre-0.10.3 (every published version 0.2.x–0.10.2) conflated "required" (must be present) with "non-nullable" (cannot be null) - JSON Schema treats those as orthogonal. The validator now honours all three idioms for nullability:
54
197
  - `type: ["string", "null"]` (array form)
55
198
  - `type: "null"` (degenerate scalar form)
56
199
  - `anyOf: [{type: "string"}, {type: "null"}]` and `oneOf` equivalents
57
200
 
58
201
  **Adopter impact:** if you bypassed the bug by setting `required: false` + disabling OpenAI `strict: true` on every nullable field, you can now restore `strict: true` and keep the field required (= OpenAI's standard nullable idiom). Non-nullable required fields still reject `nil` exactly as before; this is a strictly additive fix.
59
202
 
60
- Discovered via dogfooding in a production adopter using OpenAI strict structured output with 17 nullable fields (`"set the rest to null"` prompt). Smoking-gun mutation: remove `nullable_schema?` short-circuit from `validate_nil_field` — the three new spec cases in `spec/ruby_llm/contract/contract/schema_validator_spec.rb` fail.
203
+ Discovered via dogfooding in a production adopter using OpenAI strict structured output with 17 nullable fields (`"set the rest to null"` prompt). Smoking-gun mutation: remove `nullable_schema?` short-circuit from `validate_nil_field` - the three new spec cases in `spec/ruby_llm/contract/contract/schema_validator_spec.rb` fail.
61
204
 
62
205
  ## 0.10.2 (2026-06-10)
63
206
 
@@ -87,7 +230,7 @@ First published release since 0.8.0. Consolidates work originally tagged as 0.9.
87
230
 
88
231
  ### Breaking changes
89
232
 
90
- - **`validate(description, &block)` and `Definition#invariant(description, &block)` now raise `ArgumentError` when `description` is `nil` or empty.** Pre-0.10.0 the empty descriptor was silently accepted and produced `""` entries in `result.validation_errors`, making debugging impossible. Codex audit found zero production use sites across `lib/`, `examples/`, `README` — only the regression-marker test certifying the bug.
233
+ - **`validate(description, &block)` and `Definition#invariant(description, &block)` now raise `ArgumentError` when `description` is `nil` or empty.** Pre-0.10.0 the empty descriptor was silently accepted and produced `""` entries in `result.validation_errors`, making debugging impossible. Codex audit found zero production use sites across `lib/`, `examples/`, `README` - only the regression-marker test certifying the bug.
91
234
 
92
235
  ### Migration
93
236
 
@@ -103,9 +246,9 @@ validate("score in range 0-100") { |o| o[:score].between?(0, 100) }
103
246
 
104
247
  ### Added
105
248
 
106
- - **Multimodal input via `context: { attachment: ... }`** — pass a file/IO/URL through `Step.run(input, context: { attachment: path })`; the adapter forwards it to `RubyLLM::Chat#ask(content, with: attachment)`. RubyLLM normalises wire format per provider (Anthropic url/base64, OpenAI `image_url`/`file`, Gemini `inline_data`). Multi-attachment supported natively (`with: [pdf1, pdf2]` or `with: { images: [...], pdfs: [...] }`). See [multimodal input guide](docs/guide/multimodal_input.md) and [ADR-0022](doc/decisions/ADR-0022-v09-multimodal-input.md).
107
- - **`attachment_token_estimate(n)` class macro** — adopter-declared conservative estimate of attachment input tokens. Applied to BOTH runtime (`limit_checker`) and pre-flight (`estimate_cost`) — same source of truth, no estimate/runtime drift.
108
- - **`on_unknown_attachment_size(:refuse | :warn)` class macro** — mirrors `on_unknown_pricing` opt-out semantics. Defaults to `:refuse`. Never settable as global default — same invariant as `max_cost` fail-closed.
249
+ - **Multimodal input via `context: { attachment: ... }`** - pass a file/IO/URL through `Step.run(input, context: { attachment: path })`; the adapter forwards it to `RubyLLM::Chat#ask(content, with: attachment)`. RubyLLM normalises wire format per provider (Anthropic url/base64, OpenAI `image_url`/`file`, Gemini `inline_data`). Multi-attachment supported natively (`with: [pdf1, pdf2]` or `with: { images: [...], pdfs: [...] }`). See [multimodal input guide](docs/guide/multimodal_input.md) (rationale in the internal ADR-0022, not shipped).
250
+ - **`attachment_token_estimate(n)` class macro** - adopter-declared conservative estimate of attachment input tokens. Applied to BOTH runtime (`limit_checker`) and pre-flight (`estimate_cost`) - same source of truth, no estimate/runtime drift.
251
+ - **`on_unknown_attachment_size(:refuse | :warn)` class macro** - mirrors `on_unknown_pricing` opt-out semantics. Defaults to `:refuse`. Never settable as global default - same invariant as `max_cost` fail-closed.
109
252
 
110
253
  ### Behavioural change (READ BEFORE UPGRADING)
111
254
 
@@ -113,13 +256,13 @@ validate("score in range 0-100") { |o| o[:score].between?(0, 100) }
113
256
 
114
257
  ### Changed
115
258
 
116
- - **`run_eval` (no args) return shape pinned to `Hash<String, Report>` keyed by eval name.** Documents the existing contract used by `RubyLLM::Contract::RakeTask#collect_host_reports` and adopters. No runtime change vs 0.8.0 — only the spec assertion now locks the shape.
117
- - **`Parser.parse(text, strategy: :json)` first-bracket-wins boundary documented.** Extraction commits to the first balanced `{` or `[` structure and does NOT retry on later candidates. Empty `{}` followed by real JSON parses as the empty Hash; non-JSON `{braces}` before real JSON raises `ParseError`. No runtime change — this codifies long-standing behavior with explicit boundary tests.
259
+ - **`run_eval` (no args) return shape pinned to `Hash<String, Report>` keyed by eval name.** Documents the existing contract used by `RubyLLM::Contract::RakeTask#collect_host_reports` and adopters. No runtime change vs 0.8.0 - only the spec assertion now locks the shape.
260
+ - **`Parser.parse(text, strategy: :json)` first-bracket-wins boundary documented.** Extraction commits to the first balanced `{` or `[` structure and does NOT retry on later candidates. Empty `{}` followed by real JSON parses as the empty Hash; non-JSON `{braces}` before real JSON raises `ParseError`. No runtime change - this codifies long-standing behavior with explicit boundary tests.
118
261
 
119
262
  ### Fixed
120
263
 
121
264
  - **`with_retry_disabled` no longer mutates the step class's singleton method.** The optimizer now passes `retry_policy_override: nil` through `context:` to `compare_models`, which `Step::Base#runtime_settings` already honours. Removes a concurrency hazard where two parallel `optimize_retry_policy` calls on the same step would race on the singleton restore in `ensure`.
122
- - **`CostCalculator.find_model` exposed as a public class method.** Removes two `CostCalculator.send(:find_model, ...)` workarounds in `Step::Base#estimate_cost`. The `estimated_cost_for` helper is gone — `estimate_cost` now routes through the existing public `CostCalculator.calculate(model_name:, usage:)`.
265
+ - **`CostCalculator.find_model` exposed as a public class method.** Removes two `CostCalculator.send(:find_model, ...)` workarounds in `Step::Base#estimate_cost`. The `estimated_cost_for` helper is gone - `estimate_cost` now routes through the existing public `CostCalculator.calculate(model_name:, usage:)`.
123
266
  - **`stub_step` unified on a single storage path.** Both block and non-block forms now write to `RubyLLM::Contract.step_adapter_overrides` (thread-local). The `around(:each)` hook in `rspec.rb` handles cleanup between examples. Removes the prior `allow(step).to receive(:run)` branch.
124
267
 
125
268
  ### Internal
@@ -129,9 +272,9 @@ validate("score in range 0-100") { |o| o[:score].between?(0, 100) }
129
272
 
130
273
  ### Deferred (not in 0.10.x)
131
274
 
132
- - `add_history` multi-turn replay of prior attachments — single-turn multimodal supported; follow-up questions on the same document deferred to a later release.
133
- - Streaming + attachment — contract steps remain synchronous.
134
- - Provider-specific attachment size caps — surface only via `attachment_token_estimate` calibration; consult provider docs.
275
+ - `add_history` multi-turn replay of prior attachments - single-turn multimodal supported; follow-up questions on the same document deferred to a later release.
276
+ - Streaming + attachment - contract steps remain synchronous.
277
+ - Provider-specific attachment size caps - surface only via `attachment_token_estimate` calibration; consult provider docs.
135
278
 
136
279
  ### Tests
137
280
 
@@ -143,20 +286,20 @@ Narrative repositioning + small API additions. Internal architecture unchanged:
143
286
 
144
287
  ### Added
145
288
 
146
- - **`thinking(effort:, budget:)` class macro on `Step::Base`** — mirrors `RubyLLM::Agent.thinking` signature exactly. Stored as `{ effort:, budget: }` hash; reader returns the hash; supports `:default` reset semantics; superclass inheritance like `model`/`temperature`. The convenience alias `reasoning_effort(:low)` is implemented as `thinking(effort: :low)` — single normalized state, not separate ivar.
147
- - **Adapter wiring for `with_thinking`** — when `thinking` is set on the Step class, OR when `reasoning_effort:` is passed through context, OR when an attempt config in `retry_policy escalate(...)` carries `reasoning_effort:`, the RubyLLM adapter resolves the effective `{ effort:, budget: }` hash and forwards it via `chat.with_thinking(**)` — provider-agnostic (supports OpenAI `reasoning_effort` AND Anthropic extended-thinking budget). Precedence: per-attempt / context `reasoning_effort` overrides class-level `thinking[:effort]`; budget is taken from class-level `thinking[:budget]`. **Behavioural change vs 0.7.x**: `reasoning_effort` is now forwarded via `with_thinking` instead of `with_params`. Same wire-level OpenAI parameter; provider-agnostic Anthropic support is now automatic.
289
+ - **`thinking(effort:, budget:)` class macro on `Step::Base`** - mirrors `RubyLLM::Agent.thinking` signature exactly. Stored as `{ effort:, budget: }` hash; reader returns the hash; supports `:default` reset semantics; superclass inheritance like `model`/`temperature`. The convenience alias `reasoning_effort(:low)` is implemented as `thinking(effort: :low)` - single normalized state, not separate ivar.
290
+ - **Adapter wiring for `with_thinking`** - when `thinking` is set on the Step class, OR when `reasoning_effort:` is passed through context, OR when an attempt config in `retry_policy escalate(...)` carries `reasoning_effort:`, the RubyLLM adapter resolves the effective `{ effort:, budget: }` hash and forwards it via `chat.with_thinking(**)` - provider-agnostic (supports OpenAI `reasoning_effort` AND Anthropic extended-thinking budget). Precedence: per-attempt / context `reasoning_effort` overrides class-level `thinking[:effort]`; budget is taken from class-level `thinking[:budget]`. **Behavioural change vs 0.7.x**: `reasoning_effort` is now forwarded via `with_thinking` instead of `with_params`. Same wire-level OpenAI parameter; provider-agnostic Anthropic support is now automatic.
148
291
 
149
292
  ### Dependencies
150
293
 
151
- - **`ruby_llm` constraint bumped from `~> 1.0` to `~> 1.12`** — `Chat#with_thinking` is the canonical path for reasoning effort + extended thinking; it shipped in RubyLLM 1.12. Adopters on `ruby_llm < 1.12` need to bump RubyLLM before upgrading this gem to 0.8.0.
294
+ - **`ruby_llm` constraint bumped from `~> 1.0` to `~> 1.12`** - `Chat#with_thinking` is the canonical path for reasoning effort + extended thinking; it shipped in RubyLLM 1.12. Adopters on `ruby_llm < 1.12` need to bump RubyLLM before upgrading this gem to 0.8.0.
152
295
 
153
296
  ### Changed
154
297
 
155
- - **Tagline + README opening** — repositioned around "Contracts + Evals for RubyLLM". New "Relation to RubyLLM::Agent" section explicitly frames Step as a sibling abstraction (same niche as Agent, wider contract), not an alternative or foundation. README does not claim "Step uses Agent under the hood" — current call path is `Step → Runner → Adapters::RubyLLM → RubyLLM.chat` directly.
156
- - **`TokenEstimator` documented as heuristic** — module docstring expanded with explicit "±30% accuracy" framing. Refusal messages from `LimitChecker` now include `(heuristic ±30%)` suffix so adopters know the pre-flight number is estimated, not measured. RubyLLM 1.14 also has no pre-flight tokenizer; `RubyLLM::Tokens` is post-hoc only.
157
- - **`CostCalculator` repositioned in docs** — module narrative reframed from "cost calculator" to "fine-tune pricing registry + lookup with fallback chain". Math methods (`compute_cost`, `token_cost`, etc.) were already private; this release makes the docs match. Public API surface unchanged: `register_model`, `unregister_model`, `reset_custom_models!`, `calculate`.
158
- - **`output_schema` reframed in docs** — described as "wrapper around `RubyLLM::Schema` + client-side validation step", not a standalone feature. The schema language is identical to what `RubyLLM::Agent.schema` accepts; the difference is what wraps it.
159
- - **README retry framing** — `retry_policy escalate(...)` (model escalation on validation failure) is the marketed default. `retry_policy attempts: N` (same-model retry) stays in the API for backward compat and niche cases (subjective criteria, multi-step pipelines, weaker models) but is no longer marketed as a recommended default. Empirical basis: four small experiments across PDF quiz generation, GSM8K math (n=30 + n=120), and multi-constraint schedule generation found no useful lift for nano-class models on tasks with clear correctness criteria.
298
+ - **Tagline + README opening** - repositioned around "Contracts + Evals for RubyLLM". New "Relation to RubyLLM::Agent" section explicitly frames Step as a sibling abstraction (same niche as Agent, wider contract), not an alternative or foundation. README does not claim "Step uses Agent under the hood" - current call path is `Step → Runner → Adapters::RubyLLM → RubyLLM.chat` directly.
299
+ - **`TokenEstimator` documented as heuristic** - module docstring expanded with explicit "±30% accuracy" framing. Refusal messages from `LimitChecker` now include `(heuristic ±30%)` suffix so adopters know the pre-flight number is estimated, not measured. RubyLLM 1.14 also has no pre-flight tokenizer; `RubyLLM::Tokens` is post-hoc only.
300
+ - **`CostCalculator` repositioned in docs** - module narrative reframed from "cost calculator" to "fine-tune pricing registry + lookup with fallback chain". Math methods (`compute_cost`, `token_cost`, etc.) were already private; this release makes the docs match. Public API surface unchanged: `register_model`, `unregister_model`, `reset_custom_models!`, `calculate`.
301
+ - **`output_schema` reframed in docs** - described as "wrapper around `RubyLLM::Schema` + client-side validation step", not a standalone feature. The schema language is identical to what `RubyLLM::Agent.schema` accepts; the difference is what wraps it.
302
+ - **README retry framing** - `retry_policy escalate(...)` (model escalation on validation failure) is the marketed default. `retry_policy attempts: N` (same-model retry) stays in the API for backward compat and niche cases (subjective criteria, multi-step pipelines, weaker models) but is no longer marketed as a recommended default. Empirical basis: four small experiments across PDF quiz generation, GSM8K math (n=30 + n=120), and multi-constraint schedule generation found no useful lift for nano-class models on tasks with clear correctness criteria.
160
303
 
161
304
  ### Documentation
162
305
 
@@ -165,30 +308,30 @@ Narrative repositioning + small API additions. Internal architecture unchanged:
165
308
 
166
309
  ### Issues closed
167
310
 
168
- - **#11** (Optimizer is blind to same-model attempts) — closed after empirical experiments. `attempts: N` retry stays in API; not marketed as a default.
169
- - **#6** (Production cost reporting) — already implemented in 0.7.x; close confirmed.
311
+ - **#11** (Optimizer is blind to same-model attempts) - closed after empirical experiments. `attempts: N` retry stays in API; not marketed as a default.
312
+ - **#6** (Production cost reporting) - already implemented in 0.7.x; close confirmed.
170
313
 
171
314
  ### Not in this release (deferred)
172
315
 
173
316
  - `output_schema` Proc form for runtime-input-aware schemas (parity with `Agent.schema` Proc form). Additive, low-risk; deferred to 0.9 to keep 0.8 scope tight.
174
- - H4 (Step composing `RubyLLM::Agent` internally as config holder) — verified feasible but ROI insufficient for current adopter base; trigger-based revisit, no calendar commitment.
317
+ - H4 (Step composing `RubyLLM::Agent` internally as config holder) - verified feasible but ROI insufficient for current adopter base; trigger-based revisit, no calendar commitment.
175
318
 
176
319
  ## 0.7.3 (2026-04-24)
177
320
 
178
- Adoption-friction release. No runtime behavior changes — every delta is in `docs/`, `examples/`, or `spec/integration/` (plus the `version.rb` / Gemfile.lock bumps). Upgrading from 0.7.2 picks up the expanded guide set, the new runnable showcases, and one extra integration spec.
321
+ Adoption-friction release. No runtime behavior changes - every delta is in `docs/`, `examples/`, or `spec/integration/` (plus the `version.rb` / Gemfile.lock bumps). Upgrading from 0.7.2 picks up the expanded guide set, the new runnable showcases, and one extra integration spec.
179
322
 
180
323
  ### Documentation
181
324
 
182
- - **New guide: `docs/guide/why.md`** — four production failure modes the gem exists for (schema-valid logically wrong, silent prompt regression, sampling variance on fixed-temperature models, runaway cost). Opens from a concrete incident each time; designed for readers who have not yet felt the pain the gem relieves.
183
- - **New guide: `docs/guide/rails_integration.md`** — seven Rails-specific FAQs with runnable snippets: where step classes live (`app/contracts/`), initializer setup, background jobs, `around_call` observability, RSpec/Minitest stubs, error handling in controllers, CI gate wiring.
184
- - **README adoption-friction pass** — added a short "Do I need this?" block after Install, a reading-order hint (`README → why.md → getting_started.md`), and outcome-based labels in the docs index ("Prevent silent prompt regressions" instead of "Eval-First", etc.).
185
- - **TL;DR box at the top of every guide** — single-sentence orientation for readers who land via search; "Skip if" clause added where real confusion exists (`eval_first.md`, `testing.md`, `migration.md`).
186
- - **API coverage gaps closed** — `estimate_cost` / `estimate_eval_cost`, `max_cost on_unknown_pricing: :warn`, `run_eval(..., concurrency:)`, `around_call` testing patterns now documented in `getting_started.md`, `eval_first.md`, `testing.md`.
187
- - **Industry-standard terminology** — `temperature-locked` → `fixed-temperature`, `variance-induced` → `sampling variance`, `severity signals` → `severity keywords`, `takeaway drift` → `tone/takeaways mismatch`.
188
- - **`docs/architecture.md` refresh** — diagram now reflects the current class layout: added `Step::RetryPolicy`, `Pipeline::Result`, `Eval::AggregatedReport`, `Eval::BaselineDiff`, `Eval::PromptDiffComparator`, `Eval::EvalHistory`, `Eval::RetryOptimizer`, `OptimizeRakeTask`. Replaced the outdated `Eval::TraitEvaluator` entry with `Eval::ExpectationEvaluator`.
189
- - **Business framing added to guides** — every guide opens with a concrete production scenario or "why it matters" hook before the API reference.
325
+ - **New guide: `docs/guide/why.md`** - four production failure modes the gem exists for (schema-valid logically wrong, silent prompt regression, sampling variance on fixed-temperature models, runaway cost). Opens from a concrete incident each time; designed for readers who have not yet felt the pain the gem relieves.
326
+ - **New guide: `docs/guide/rails_integration.md`** - seven Rails-specific FAQs with runnable snippets: where step classes live (`app/contracts/`), initializer setup, background jobs, `around_call` observability, RSpec/Minitest stubs, error handling in controllers, CI gate wiring.
327
+ - **README adoption-friction pass** - added a short "Do I need this?" block after Install, a reading-order hint (`README → why.md → getting_started.md`), and outcome-based labels in the docs index ("Prevent silent prompt regressions" instead of "Eval-First", etc.).
328
+ - **TL;DR box at the top of every guide** - single-sentence orientation for readers who land via search; "Skip if" clause added where real confusion exists (`eval_first.md`, `testing.md`, `migration.md`).
329
+ - **API coverage gaps closed** - `estimate_cost` / `estimate_eval_cost`, `max_cost on_unknown_pricing: :warn`, `run_eval(..., concurrency:)`, `around_call` testing patterns now documented in `getting_started.md`, `eval_first.md`, `testing.md`.
330
+ - **Industry-standard terminology** - `temperature-locked` → `fixed-temperature`, `variance-induced` → `sampling variance`, `severity signals` → `severity keywords`, `takeaway drift` → `tone/takeaways mismatch`.
331
+ - **`docs/architecture.md` refresh** - diagram now reflects the current class layout: added `Step::RetryPolicy`, `Pipeline::Result`, `Eval::AggregatedReport`, `Eval::BaselineDiff`, `Eval::PromptDiffComparator`, `Eval::EvalHistory`, `Eval::RetryOptimizer`, `OptimizeRakeTask`. Replaced the outdated `Eval::TraitEvaluator` entry with `Eval::ExpectationEvaluator`.
332
+ - **Business framing added to guides** - every guide opens with a concrete production scenario or "why it matters" hook before the API reference.
190
333
 
191
- ### Examples — consolidated on `SummarizeArticle`, renumbered 00-06
334
+ ### Examples - consolidated on `SummarizeArticle`, renumbered 00-06
192
335
 
193
336
  The previous 12-file set mixed a private Reddit promo planner, customer support, meetings, keyword extraction, and translation. The new set is seven runnable files, each answering one adopter question on the README's `SummarizeArticle` case.
194
337
 
@@ -204,19 +347,19 @@ The previous 12-file set mixed a private Reddit promo planner, customer support,
204
347
 
205
348
  Every file carries an "Expected output" block in its header so readers see the result without running the script. The `docs/ideas/` directory is now fully untracked (already in `.gitignore`; one stray file removed from version control).
206
349
 
207
- ### Examples — bug fixes carried along
350
+ ### Examples - bug fixes carried along
208
351
 
209
- - **Schema pitfall fixed in 5 files** — `array :x do; string :y; ...; end` silently produces `items: string` and drops every declaration after the first, matching the documented pitfall in `spec/ruby_llm/contract/nested_schema_spec.rb:71`. Every affected array block is now wrapped in `object do...end`.
210
- - **`examples/05_eval_dataset.rb` (pre-renumber: `09_eval_dataset.rb`) `result[:passed]` → `result.passed?`** — the previous code called `[]` on an `Eval::CaseResult` and raised `NoMethodError` at runtime.
352
+ - **Schema pitfall fixed in 5 files** - `array :x do; string :y; ...; end` silently produces `items: string` and drops every declaration after the first, matching the documented pitfall in `spec/ruby_llm/contract/nested_schema_spec.rb:71`. Every affected array block is now wrapped in `object do...end`.
353
+ - **`examples/05_eval_dataset.rb` (pre-renumber: `09_eval_dataset.rb`) `result[:passed]` → `result.passed?`** - the previous code called `[]` on an `Eval::CaseResult` and raised `NoMethodError` at runtime.
211
354
 
212
355
  ### Testing
213
356
 
214
- - **New `spec/integration/pipeline_eval_spec.rb`** — three cases guaranteeing pipeline-level `run_eval` stays functional: happy path, final-step mismatch, and fail-fast propagation when an intermediate `validate` rejects. Closes the "09 STEP 5 pipeline evaluation" known issue flagged in the 0.7.2 release. The fail-fast case asserts `step_status == :validation_failed` and the validate's label in `details`, so a regression that short-circuits on schema instead of validate would fail loudly.
357
+ - **New `spec/integration/pipeline_eval_spec.rb`** - three cases guaranteeing pipeline-level `run_eval` stays functional: happy path, final-step mismatch, and fail-fast propagation when an intermediate `validate` rejects. Closes the "09 STEP 5 pipeline evaluation" known issue flagged in the 0.7.2 release. The fail-fast case asserts `step_status == :validation_failed` and the validate's label in `details`, so a regression that short-circuits on schema instead of validate would fail loudly.
215
358
 
216
359
  ### Deleted (private-project cleanup)
217
360
 
218
- - `examples/01_classify_threads.rb`, `02_generate_comment.rb`, `03_target_audience.rb`, `10_reddit_full_showcase.rb`, `spec/integration/reddit_pipeline_spec.rb` — Reddit Promo Planner was a separate private project; its examples do not belong in the gem's public repo.
219
- - `examples/02_output_schema.rb` — fully covered by `docs/guide/output_schema.md`; deleting avoids duplication.
361
+ - `examples/01_classify_threads.rb`, `02_generate_comment.rb`, `03_target_audience.rb`, `10_reddit_full_showcase.rb`, `spec/integration/reddit_pipeline_spec.rb` - Reddit Promo Planner was a separate private project; its examples do not belong in the gem's public repo.
362
+ - `examples/02_output_schema.rb` - fully covered by `docs/guide/output_schema.md`; deleting avoids duplication.
220
363
 
221
364
  ## 0.7.2 (2026-04-22)
222
365
 
@@ -226,24 +369,24 @@ Every file carries an "Expected output" block in its header so readers see the r
226
369
 
227
370
  ### Documentation
228
371
 
229
- - **`docs/guide/optimizing_retry_policy.md` rewritten** — 17.7k → 6.4k characters. Continues the `SummarizeArticle` narrative from README. Offline mode clearly positioned as wiring-check; real optimization runs via `LIVE=1 RUNS=3`. Output samples match actual `print_summary` format.
230
- - **`docs/guide/getting_started.md` rewritten** — 8.7k → 6.1k. Every example uses `SummarizeArticle`. Evals + CI gates section moved before Budget caps. Structured Prompts / Dynamic Prompts / "Already using ruby_llm?" / Reasoning effort sections removed; content delegated to `prompt_ast.md` and README.
231
- - **`docs/guide/eval_first.md` refined** — 6.3k → 5.0k. Switched to `SummarizeArticle` case. Team workflow section compressed with links back to `getting_started.md` for the matcher chain.
232
- - **`docs/guide/testing.md` refined** — 10.7k → 7.4k. Switched to `SummarizeArticle` case. Threshold gating / Rake task / baseline walkthrough / prompt A/B sections delegated back to `getting_started.md` and `eval_first.md`.
233
- - **`docs/guide/output_schema.md` DSL bug fix** — the Supported constraints table documented JSON Schema camelCase keys (`minLength`, `minItems`, `additionalProperties`) that are not valid DSL arguments. Every copy-paste from the previous table would have raised `ArgumentError`. Switched to snake_case (`min_length`, `min_items`, `additional_properties`) as the DSL actually expects; added a short note on the internal camelCase conversion.
234
- - **`docs/guide/best_practices.md`, `pipeline.md`, `migration.md` sanity pass** — terminology alignment (model escalation → model fallback where narrative; `escalate` DSL method unchanged) and `SummarizeArticle` case where the guide is not inherently multi-step.
372
+ - **`docs/guide/optimizing_retry_policy.md` rewritten** - 17.7k → 6.4k characters. Continues the `SummarizeArticle` narrative from README. Offline mode clearly positioned as wiring-check; real optimization runs via `LIVE=1 RUNS=3`. Output samples match actual `print_summary` format.
373
+ - **`docs/guide/getting_started.md` rewritten** - 8.7k → 6.1k. Every example uses `SummarizeArticle`. Evals + CI gates section moved before Budget caps. Structured Prompts / Dynamic Prompts / "Already using ruby_llm?" / Reasoning effort sections removed; content delegated to `prompt_ast.md` and README.
374
+ - **`docs/guide/eval_first.md` refined** - 6.3k → 5.0k. Switched to `SummarizeArticle` case. Team workflow section compressed with links back to `getting_started.md` for the matcher chain.
375
+ - **`docs/guide/testing.md` refined** - 10.7k → 7.4k. Switched to `SummarizeArticle` case. Threshold gating / Rake task / baseline walkthrough / prompt A/B sections delegated back to `getting_started.md` and `eval_first.md`.
376
+ - **`docs/guide/output_schema.md` DSL bug fix** - the Supported constraints table documented JSON Schema camelCase keys (`minLength`, `minItems`, `additionalProperties`) that are not valid DSL arguments. Every copy-paste from the previous table would have raised `ArgumentError`. Switched to snake_case (`min_length`, `min_items`, `additional_properties`) as the DSL actually expects; added a short note on the internal camelCase conversion.
377
+ - **`docs/guide/best_practices.md`, `pipeline.md`, `migration.md` sanity pass** - terminology alignment (model escalation → model fallback where narrative; `escalate` DSL method unchanged) and `SummarizeArticle` case where the guide is not inherently multi-step.
235
378
 
236
379
  ## 0.7.1 (2026-04-22)
237
380
 
238
381
  ### Changed (behavioral, follow-up to v0.7.0)
239
382
 
240
- - **`Step::Base#run_once` no longer swallows adapter-phase `ArgumentError` as `:input_error`.** The previous blanket `rescue ArgumentError` was there to convert DSL misconfiguration (e.g. missing `prompt`) into an `:input_error` Result. Side effect: programmer bugs in adapter code that raised `ArgumentError` (wrong arity, bad config argument) were silently coerced into `:input_error` and retried as if the user had given bad input. Now the rescue is narrowed to the Runner-construction phase only — DSL configuration errors still produce `:input_error` (the `prompt has not been set` case is regression-tested), but `ArgumentError` raised from adapter code during `Runner#call` propagates to the caller. Input-type validation failures continue to produce `:input_error` through `InputValidator`'s own scoped rescue, unchanged.
383
+ - **`Step::Base#run_once` no longer swallows adapter-phase `ArgumentError` as `:input_error`.** The previous blanket `rescue ArgumentError` was there to convert DSL misconfiguration (e.g. missing `prompt`) into an `:input_error` Result. Side effect: programmer bugs in adapter code that raised `ArgumentError` (wrong arity, bad config argument) were silently coerced into `:input_error` and retried as if the user had given bad input. Now the rescue is narrowed to the Runner-construction phase only - DSL configuration errors still produce `:input_error` (the `prompt has not been set` case is regression-tested), but `ArgumentError` raised from adapter code during `Runner#call` propagates to the caller. Input-type validation failures continue to produce `:input_error` through `InputValidator`'s own scoped rescue, unchanged.
241
384
 
242
385
  ## 0.7.0 (2026-04-21)
243
386
 
244
387
  ### Breaking changes
245
388
 
246
- - **`:adapter_error` removed from `DEFAULT_RETRY_ON`.** New default: `[:validation_failed, :parse_error]`. `ruby_llm` already retries transport errors (`RateLimitError`, `ServerError`, `ServiceUnavailableError`, `OverloadedError`, timeouts) at the Faraday layer, so the previous default re-ran the same model on errors the HTTP middleware already retried with backoff. To restore pre-0.7 behavior: `retry_on :validation_failed, :parse_error, :adapter_error`. Recommended pattern: pair `:adapter_error` with `escalate "model_a", "model_b"` — a different model/provider can bypass what transport retry could not.
389
+ - **`:adapter_error` removed from `DEFAULT_RETRY_ON`.** New default: `[:validation_failed, :parse_error]`. `ruby_llm` already retries transport errors (`RateLimitError`, `ServerError`, `ServiceUnavailableError`, `OverloadedError`, timeouts) at the Faraday layer, so the previous default re-ran the same model on errors the HTTP middleware already retried with backoff. To restore pre-0.7 behavior: `retry_on :validation_failed, :parse_error, :adapter_error`. Recommended pattern: pair `:adapter_error` with `escalate "model_a", "model_b"` - a different model/provider can bypass what transport retry could not.
247
390
  - **`AdapterCaller` narrows `rescue` from `StandardError` to `RubyLLM::Error` + `Faraday::Error`.** Provider errors and transport errors that escape ruby_llm's Faraday retry middleware (`Faraday::TimeoutError`, `Faraday::ConnectionFailed`) still produce `:adapter_error` as before. Programmer errors that are neither (`NoMethodError`, adapter code bugs) now propagate instead of being silently converted to `:adapter_error` and retried. **Known limitation:** adapter code raising `ArgumentError` is still coerced into `:input_error` by `Step::Base#run_once` (which rescues `ArgumentError` for input-type validation). Disambiguating adapter-ArgumentError vs input-validation-ArgumentError requires a `run_once` refactor and is tracked as a follow-up.
248
391
 
249
392
  ### Migration
@@ -270,43 +413,43 @@ end
270
413
 
271
414
  ### Features
272
415
 
273
- - **`production_mode:` on `compare_models` and `optimize_retry_policy`** — measures retry-aware, end-to-end cost per successful output. Pass `production_mode: { fallback: "gpt-5-mini" }` and each candidate runs with a runtime-injected `[candidate, fallback]` retry chain. The report exposes `escalation_rate`, `single_shot_cost`, and `effective_cost` so "the cheaper candidate" decision matches production cost rather than first-attempt cost.
274
- - **New Report metrics** — `escalation_rate`, `single_shot_cost`, `effective_cost`, `single_shot_latency_ms`, `effective_latency_ms`, `latency_percentiles` (p50/p95/max). `AggregatedReport` averages all of them across `runs:`.
275
- - **Extended `ModelComparison#table`** — when `production_mode:` is set, renders a `Chain` column (`candidate → fallback`) with `single-shot`, `escalation`, `effective cost`, `latency`, `score`. Edge case `candidate == fallback` renders as a single model and `—` in the escalation column, with retry injection skipped entirely so `effective == single-shot` by construction, not by coincidence.
276
- - **`context[:retry_policy_override]`** — new context key that nullifies or replaces class-level `retry_policy` for a single call. Used internally by production-mode injection; safe to use directly when you need a transient override that doesn't mutate the step class.
416
+ - **`production_mode:` on `compare_models` and `optimize_retry_policy`** - measures retry-aware, end-to-end cost per successful output. Pass `production_mode: { fallback: "gpt-5-mini" }` and each candidate runs with a runtime-injected `[candidate, fallback]` retry chain. The report exposes `escalation_rate`, `single_shot_cost`, and `effective_cost` so "the cheaper candidate" decision matches production cost rather than first-attempt cost.
417
+ - **New Report metrics** - `escalation_rate`, `single_shot_cost`, `effective_cost`, `single_shot_latency_ms`, `effective_latency_ms`, `latency_percentiles` (p50/p95/max). `AggregatedReport` averages all of them across `runs:`.
418
+ - **Extended `ModelComparison#table`** - when `production_mode:` is set, renders a `Chain` column (`candidate → fallback`) with `single-shot`, `escalation`, `effective cost`, `latency`, `score`. Edge case `candidate == fallback` renders as a single model and `—` in the escalation column, with retry injection skipped entirely so `effective == single-shot` by construction, not by coincidence.
419
+ - **`context[:retry_policy_override]`** - new context key that nullifies or replaces class-level `retry_policy` for a single call. Used internally by production-mode injection; safe to use directly when you need a transient override that doesn't mutate the step class.
277
420
 
278
421
  ### Scope
279
422
 
280
423
  - Single-fallback (2-tier) chains only. Multi-tier chains can be inspected post-hoc via `trace.attempts` but aren't summarized in the optimize table.
281
- - Costs with `runs: 3 + production_mode: { fallback: "gpt-5-mini" }` are ≈3× a single-shot eval plus the actual retry attempts — not 6×. Production-mode metrics come from a single pass.
282
- - **Step-only.** Calling `compare_models` with `production_mode:` on a `Pipeline::Base` subclass raises `ArgumentError` — retry injection is Step-level and pipeline-wide fallback semantics aren't defined yet. Benchmark individual steps.
424
+ - Costs with `runs: 3 + production_mode: { fallback: "gpt-5-mini" }` are ≈3× a single-shot eval plus the actual retry attempts - not 6×. Production-mode metrics come from a single pass.
425
+ - **Step-only.** Calling `compare_models` with `production_mode:` on a `Pipeline::Base` subclass raises `ArgumentError` - retry injection is Step-level and pipeline-wide fallback semantics aren't defined yet. Benchmark individual steps.
283
426
 
284
427
  ### Documentation
285
428
 
286
- - **Guide: [Production-mode cost measurement](docs/guide/optimizing_retry_policy.md#production-mode-cost-measurement)** — API, metric interpretation, 2-tier scope note.
429
+ - **Guide: [Production-mode cost measurement](docs/guide/optimizing_retry_policy.md#measure-effective-cost-before-shipping)** - API, metric interpretation, 2-tier scope note.
287
430
 
288
431
  ## 0.6.3 (2026-04-20)
289
432
 
290
433
  ### Features
291
434
 
292
- - **`runs:` parameter on `compare_models` and `optimize_retry_policy`** — runs each candidate N times per eval and aggregates the mean score, mean cost per run, and mean latency. Reduces sampling variance in live mode where LLM outputs are non-deterministic (gpt-5 family enforces `temperature=1.0` server-side, so a single unlucky sample can misclassify a viable candidate as "failing"). Default `runs: 1` — backward compatible.
293
- - **`RUNS=N` on `rake ruby_llm_contract:optimize`** — CLI flag for variance-aware optimization.
294
- - **`Eval::AggregatedReport`** — duck-type `Report` exposing `score` (mean), `score_min`/`score_max` (spread), `total_cost` (mean per run), `pass_rate` (clean-pass count x/N), and `clean_passes`.
295
- - **Guide: [Reducing variance with `runs:`](docs/guide/optimizing_retry_policy.md#reducing-variance-with-runs)** — when to use it and why.
435
+ - **`runs:` parameter on `compare_models` and `optimize_retry_policy`** - runs each candidate N times per eval and aggregates the mean score, mean cost per run, and mean latency. Reduces sampling variance in live mode where LLM outputs are non-deterministic (gpt-5 family enforces `temperature=1.0` server-side, so a single unlucky sample can misclassify a viable candidate as "failing"). Default `runs: 1` - backward compatible.
436
+ - **`RUNS=N` on `rake ruby_llm_contract:optimize`** - CLI flag for variance-aware optimization.
437
+ - **`Eval::AggregatedReport`** - duck-type `Report` exposing `score` (mean), `score_min`/`score_max` (spread), `total_cost` (mean per run), `pass_rate` (clean-pass count x/N), and `clean_passes`.
438
+ - **Guide: [Reducing variance with `runs:`](docs/guide/optimizing_retry_policy.md)** - when to use it and why.
296
439
 
297
440
  ## 0.6.2 (2026-04-18)
298
441
 
299
442
  ### Features
300
443
 
301
- - **`Step.optimize_retry_policy`** — runs `compare_models` on ALL evals for the step, builds a score matrix, identifies the constraining eval, and suggests a retry chain. Chain's last model always passes all evals (safe fallback).
302
- - **`rake ruby_llm_contract:optimize`** — one-command retry chain optimization. Prints score table, constraining eval, suggested chain, and copy-paste DSL.
303
- - **Offline by default** — `optimize` uses `sample_response` (zero API calls) unless `LIVE=1` or `PROVIDER=` is set.
304
- - **`EVAL_DIRS=` support** — non-Rails setups can specify eval file directories.
305
- - **Guide: [Optimizing retry_policy](docs/guide/optimizing_retry_policy.md)** — full procedure with prerequisites, troubleshooting, and real-world example.
444
+ - **`Step.optimize_retry_policy`** - runs `compare_models` on ALL evals for the step, builds a score matrix, identifies the constraining eval, and suggests a retry chain. Chain's last model always passes all evals (safe fallback).
445
+ - **`rake ruby_llm_contract:optimize`** - one-command retry chain optimization. Prints score table, constraining eval, suggested chain, and copy-paste DSL.
446
+ - **Offline by default** - `optimize` uses `sample_response` (zero API calls) unless `LIVE=1` or `PROVIDER=` is set.
447
+ - **`EVAL_DIRS=` support** - non-Rails setups can specify eval file directories.
448
+ - **Guide: [Optimizing retry_policy](docs/guide/optimizing_retry_policy.md)** - full procedure with prerequisites, troubleshooting, and real-world example.
306
449
 
307
450
  ### Fixes
308
451
 
309
- - Chain semantics aligned with `retry_executor` — retry fires on `validation_failed`/`parse_error`, not on low eval score. Disjoint eval coverage (A passes e1, B passes e2, neither passes both) correctly returns empty chain.
452
+ - Chain semantics aligned with `retry_executor` - retry fires on `validation_failed`/`parse_error`, not on low eval score. Disjoint eval coverage (A passes e1, B passes e2, neither passes both) correctly returns empty chain.
310
453
  - Removed ActiveSupport dependency from rake task (`.presence` → `.empty?`).
311
454
  - Added `require "set"` for non-Rails environments.
312
455
 
@@ -314,22 +457,22 @@ end
314
457
 
315
458
  ### Features
316
459
 
317
- - **Multi-provider operator tooling** — rake tasks support `PROVIDER=openai|anthropic|ollama`, `CANDIDATES=model@effort,...`, and `REASONING_EFFORT=low|medium|high`.
318
- - **`rake ruby_llm_contract:recommend`** — wraps `Step.recommend` with CLI interface, prints best config, retry chain, DSL, rationale, and savings.
319
- - **Ollama support** — `PROVIDER=ollama` with configurable `OLLAMA_API_BASE`.
460
+ - **Multi-provider operator tooling** - rake tasks support `PROVIDER=openai|anthropic|ollama` and `CANDIDATES=model@effort,...`.
461
+ - **`rake ruby_llm_contract:recommend`** - wraps `Step.recommend` with CLI interface, prints best config, retry chain, DSL, rationale, and savings.
462
+ - **Ollama support** - `PROVIDER=ollama` (base URL via `RubyLLM.configure { |c| c.ollama_api_base = ... }`).
320
463
 
321
464
  ## 0.6.0 (2026-04-12)
322
465
 
323
- "What should I do?" — model + configuration recommendation.
466
+ "What should I do?" - model + configuration recommendation.
324
467
 
325
468
  ### Features
326
469
 
327
- - **`Step.recommend`** — `ClassifyTicket.recommend("eval", candidates: [...], min_score: 0.95)` runs eval on all candidates and returns a `Recommendation` with optimal model, retry chain, rationale, savings vs current config, and `to_dsl` code output.
328
- - **Candidates as configurations** — `candidates:` accepts `{ model:, reasoning_effort: }` hashes, not just model name strings. `gpt-5-mini` with `reasoning_effort: "low"` is a different candidate than with `"high"`.
329
- - **`compare_models` extended** — new `candidates:` parameter alongside existing `models:` (backward compatible). Candidate labels include reasoning effort in output table.
330
- - **Per-attempt `reasoning_effort` in retry policies** — `escalate` accepts config hashes: `escalate({ model: "gpt-5-nano" }, { model: "gpt-5-mini", reasoning_effort: "high" })`. Each attempt gets its own reasoning_effort forwarded to the provider.
331
- - **`pass_rate_ratio`** — numeric float (0.0–1.0) on `Report` and `ReportStats`, complementing the string `pass_rate` (`"3/5"`).
332
- - **History entries enriched** — `save_history!` accepts `reasoning_effort:` and stores `model`, `reasoning_effort`, `pass_rate_ratio` in JSONL entries.
470
+ - **`Step.recommend`** - `ClassifyTicket.recommend("eval", candidates: [...], min_score: 0.95)` runs eval on all candidates and returns a `Recommendation` with optimal model, retry chain, rationale, savings vs current config, and `to_dsl` code output.
471
+ - **Candidates as configurations** - `candidates:` accepts `{ model:, reasoning_effort: }` hashes, not just model name strings. `gpt-5-mini` with `reasoning_effort: "low"` is a different candidate than with `"high"`.
472
+ - **`compare_models` extended** - new `candidates:` parameter alongside existing `models:` (backward compatible). Candidate labels include reasoning effort in output table.
473
+ - **Per-attempt `reasoning_effort` in retry policies** - `escalate` accepts config hashes: `escalate({ model: "gpt-5-nano" }, { model: "gpt-5-mini", reasoning_effort: "high" })`. Each attempt gets its own reasoning_effort forwarded to the provider.
474
+ - **`pass_rate_ratio`** - numeric float (0.0–1.0) on `Report` and `ReportStats`, complementing the string `pass_rate` (`"3/5"`).
475
+ - **History entries enriched** - `save_history!` accepts `reasoning_effort:` and stores `model`, `reasoning_effort`, `pass_rate_ratio` in JSONL entries.
333
476
 
334
477
  ### Game changer continuity
335
478
 
@@ -345,7 +488,7 @@ v0.6: "What should I do?" → recommend (actionable advice)
345
488
 
346
489
  ### Features
347
490
 
348
- - **`reasoning_effort` forwarded to provider** — `context: { reasoning_effort: "low" }` now passed through `with_params` to the LLM. Previously accepted as a known context key but silently ignored by the RubyLLM adapter.
491
+ - **`reasoning_effort` forwarded to provider** - `context: { reasoning_effort: "low" }` now passed through `with_params` to the LLM. Previously accepted as a known context key but silently ignored by the RubyLLM adapter.
349
492
 
350
493
  ## 0.5.0 (2026-03-25)
351
494
 
@@ -353,9 +496,9 @@ Data-Driven Prompt Engineering.
353
496
 
354
497
  ### Features
355
498
 
356
- - **`observe` DSL** — soft observations that log but never fail. `observe("scores differ") { |o| o[:a] != o[:b] }`. Results in `result.observations`. Logged via `Contract.logger` when they fail. Runs only when validation passes.
357
- - **`compare_with`** — prompt A/B testing. `StepV2.compare_with(StepV1, eval: "regression", model: "nano")` returns `PromptDiff` with `improvements`, `regressions`, `score_delta`, `safe_to_switch?`. Reuses `BaselineDiff` internally.
358
- - **RSpec `compared_with` chain** — `expect(StepV2).to pass_eval("x").compared_with(StepV1).without_regressions` blocks merge if new prompt regresses any case.
499
+ - **`observe` DSL** - soft observations that log but never fail. `observe("scores differ") { |o| o[:a] != o[:b] }`. Results in `result.observations`. Logged via `Contract.logger` when they fail. Runs only when validation passes.
500
+ - **`compare_with`** - prompt A/B testing. `StepV2.compare_with(StepV1, eval: "regression", model: "nano")` returns `PromptDiff` with `improvements`, `regressions`, `score_delta`, `safe_to_switch?`. Reuses `BaselineDiff` internally.
501
+ - **RSpec `compared_with` chain** - `expect(StepV2).to pass_eval("x").compared_with(StepV1).without_regressions` blocks merge if new prompt regresses any case.
359
502
 
360
503
  ### Game changer continuity
361
504
 
@@ -368,25 +511,25 @@ v0.5: "Which prompt is better?" → compare_with (A/B testing)
368
511
 
369
512
  ## 0.4.5 (2026-03-24)
370
513
 
371
- Audit hardening — 18 bugs fixed across 4 audit rounds.
514
+ Audit hardening - 18 bugs fixed across 4 audit rounds.
372
515
 
373
516
  ### Fixes
374
517
 
375
- - **RakeTask history before abort** — `track_history` now saves all reports (pass and fail) before gating, so failed runs appear in eval history.
376
- - **RSpec/Minitest stub scoping** — block form `stub_step` uses thread-local overrides with real cleanup. Non-block `stub_all_steps` auto-restored by RSpec `around(:each)` hook and Minitest `setup`/`teardown`.
377
- - **StepAdapterOverride** — handles `context: nil` and respects string key `"adapter"`. Moved to `contract.rb` so both test frameworks share one mechanism.
378
- - **max_cost fail closed output estimate** — preflight uses 1x input tokens as output estimate when `max_output` not set, preventing cost bypass for output-expensive models.
379
- - **reset_configuration! clears overrides** — `step_adapter_overrides` now cleared on reset.
380
- - **CostCalculator.register_model** — validates `Numeric`, `finite?`, non-negative. Rejects NaN, Infinity, strings, nil.
381
- - **Pipeline token_budget** — rejects negative and zero values (parity with `timeout_ms`).
382
- - **track_history model fallback** — uses step DSL `model`, then `default_model` when context has no model. Handles string key `"model"`.
383
- - **estimate_cost / estimate_eval_cost** — falls back to step DSL model when no explicit model arg given.
384
- - **stub_steps string keys** — both RSpec and Minitest normalize string-keyed options with `transform_keys(:to_sym)`.
385
- - **DSL `:default` reset** — `model(:default)`, `temperature(:default)`, `max_cost(:default)` reset inherited parent values.
518
+ - **RakeTask history before abort** - `track_history` now saves all reports (pass and fail) before gating, so failed runs appear in eval history.
519
+ - **RSpec/Minitest stub scoping** - block form `stub_step` uses thread-local overrides with real cleanup. Non-block `stub_all_steps` auto-restored by RSpec `around(:each)` hook and Minitest `setup`/`teardown`.
520
+ - **StepAdapterOverride** - handles `context: nil` and respects string key `"adapter"`. Moved to `contract.rb` so both test frameworks share one mechanism.
521
+ - **max_cost fail closed output estimate** - preflight uses 1x input tokens as output estimate when `max_output` not set, preventing cost bypass for output-expensive models.
522
+ - **reset_configuration! clears overrides** - `step_adapter_overrides` now cleared on reset.
523
+ - **CostCalculator.register_model** - validates `Numeric`, `finite?`, non-negative. Rejects NaN, Infinity, strings, nil.
524
+ - **Pipeline token_budget** - rejects negative and zero values (parity with `timeout_ms`).
525
+ - **track_history model fallback** - uses step DSL `model`, then `default_model` when context has no model. Handles string key `"model"`.
526
+ - **estimate_cost / estimate_eval_cost** - falls back to step DSL model when no explicit model arg given.
527
+ - **stub_steps string keys** - both RSpec and Minitest normalize string-keyed options with `transform_keys(:to_sym)`.
528
+ - **DSL `:default` reset** - `model(:default)`, `temperature(:default)`, `max_cost(:default)` reset inherited parent values.
386
529
 
387
530
  ## 0.4.4 (2026-03-24)
388
531
 
389
- - **`stub_steps` (plural)** — stub multiple steps with different responses in one block. No nesting needed. Works in RSpec and Minitest:
532
+ - **`stub_steps` (plural)** - stub multiple steps with different responses in one block. No nesting needed. Works in RSpec and Minitest:
390
533
  ```ruby
391
534
  stub_steps(
392
535
  ClassifyTicket => { response: { priority: "high" } },
@@ -400,33 +543,33 @@ Production feedback release.
400
543
 
401
544
  ### Features
402
545
 
403
- - **`stub_step` block form** — `stub_step(Step, response: x) { test }` auto-resets adapter after block. Works in RSpec and Minitest. Eliminates leaked test state.
404
- - **Minitest per-step routing** — `stub_step(StepA, ...)` now actually routes to StepA only (was setting global adapter, ignoring step class).
405
- - **`track_history` in RakeTask** — `t.track_history = true` auto-appends every eval run (pass and fail) to `.eval_history/`. Drift detection without manual `save_history!` calls.
406
- - **`max_cost` fail closed** — unknown model pricing now refuses the call instead of silently skipping. Set `on_unknown_pricing: :warn` for old behavior.
407
- - **`CostCalculator.register_model`** — register pricing for custom/fine-tuned models: `register_model("ft:gpt-4o", input_per_1m: 3.0, output_per_1m: 6.0)`.
546
+ - **`stub_step` block form** - `stub_step(Step, response: x) { test }` auto-resets adapter after block. Works in RSpec and Minitest. Eliminates leaked test state.
547
+ - **Minitest per-step routing** - `stub_step(StepA, ...)` now actually routes to StepA only (was setting global adapter, ignoring step class).
548
+ - **`track_history` in RakeTask** - `t.track_history = true` auto-appends every eval run (pass and fail) to `.eval_history/`. Drift detection without manual `save_history!` calls.
549
+ - **`max_cost` fail closed** - unknown model pricing now refuses the call instead of silently skipping. Set `on_unknown_pricing: :warn` for old behavior.
550
+ - **`CostCalculator.register_model`** - register pricing for custom/fine-tuned models: `register_model("ft:gpt-4o", input_per_1m: 3.0, output_per_1m: 6.0)`.
408
551
 
409
552
  ## 0.4.2 (2026-03-24)
410
553
 
411
- - **RakeTask lazy context** — `t.context` now accepts a Proc, resolved at task runtime (after `:environment`). Fixes adapter not being available at Rake load time in Rails apps.
554
+ - **RakeTask lazy context** - `t.context` now accepts a Proc, resolved at task runtime (after `:environment`). Fixes adapter not being available at Rake load time in Rails apps.
412
555
 
413
556
  ## 0.4.1 (2026-03-24)
414
557
 
415
- - **RakeTask `:environment` fix** — uses `defined?(::Rails)` instead of `Rake::Task.task_defined?(:environment)`. Works in Rails 8 without manual `Rake::Task.enhance`.
416
- - **Concurrent eval deterministic** — `clone_for_concurrency` protocol, `ContextHelpers` extracted.
417
- - **README** — added eval history, concurrency, quality tracking examples.
558
+ - **RakeTask `:environment` fix** - uses `defined?(::Rails)` instead of `Rake::Task.task_defined?(:environment)`. Works in Rails 8 without manual `Rake::Task.enhance`.
559
+ - **Concurrent eval deterministic** - `clone_for_concurrency` protocol, `ContextHelpers` extracted.
560
+ - **README** - added eval history, concurrency, quality tracking examples.
418
561
 
419
562
  ## 0.4.0 (2026-03-24)
420
563
 
421
- Observability & Scale — see what changed, run it fast, debug it easily.
564
+ Observability & Scale - see what changed, run it fast, debug it easily.
422
565
 
423
566
  ### Features
424
567
 
425
- - **Structured logging** — `Contract.configure { |c| c.logger = Rails.logger }`. Auto-logs model, status, latency, tokens, cost on every `step.run`.
426
- - **Batch eval concurrency** — `run_eval("regression", concurrency: 4)`. Parallel case execution via Concurrent::Future. 4x faster CI for large eval suites.
427
- - **Eval history & trending** — `report.save_history!` appends to JSONL. `report.eval_history` returns `EvalHistory` with `score_trend`, `drift?`, run-by-run scores.
428
- - **Pipeline per-step eval** — `add_case(..., step_expectations: { classify: { priority: "high" } })`. See which step in a pipeline regressed.
429
- - **Minitest support** — `assert_satisfies_contract`, `assert_eval_passes`, `stub_step` for Minitest users. `require "ruby_llm/contract/minitest"`.
568
+ - **Structured logging** - `Contract.configure { |c| c.logger = Rails.logger }`. Auto-logs model, status, latency, tokens, cost on every `step.run`.
569
+ - **Batch eval concurrency** - `run_eval("regression", concurrency: 4)`. Parallel case execution via Concurrent::Future. 4x faster CI for large eval suites.
570
+ - **Eval history & trending** - `report.save_history!` appends to JSONL. `report.eval_history` returns `EvalHistory` with `score_trend`, `drift?`, run-by-run scores.
571
+ - **Pipeline per-step eval** - `add_case(..., step_expectations: { classify: { priority: "high" } })`. See which step in a pipeline regressed.
572
+ - **Minitest support** - `assert_satisfies_contract`, `assert_eval_passes`, `stub_step` for Minitest users. `require "ruby_llm/contract/minitest"`.
430
573
 
431
574
  ### Game changer continuity
432
575
 
@@ -440,71 +583,71 @@ v0.4: "Show me the trend" → eval history (time series)
440
583
 
441
584
  ## 0.3.7 (2026-03-24)
442
585
 
443
- - **Trait missing key = error** — `expected_traits: { title: 0..5 }` on output `{}` now fails instead of silently passing.
444
- - **nil input in dynamic prompts** — `run(nil)` with `prompt { |input| ... }` correctly passes nil to block.
445
- - **Defensive sample pre-validation** — `sample_response` uses the same parser as runtime (handles code fences, BOM, prose around JSON).
446
- - **Baseline diff excludes skipped** — self-compare with skipped cases no longer shows artificial score delta.
447
- - **Zeitwerk eval/ ignore** — `eager_load_contract_dirs!` ignores `eval/` subdirs before eager load.
586
+ - **Trait missing key = error** - `expected_traits: { title: 0..5 }` on output `{}` now fails instead of silently passing.
587
+ - **nil input in dynamic prompts** - `run(nil)` with `prompt { |input| ... }` correctly passes nil to block.
588
+ - **Defensive sample pre-validation** - `sample_response` uses the same parser as runtime (handles code fences, BOM, prose around JSON).
589
+ - **Baseline diff excludes skipped** - self-compare with skipped cases no longer shows artificial score delta.
590
+ - **Zeitwerk eval/ ignore** - `eager_load_contract_dirs!` ignores `eval/` subdirs before eager load.
448
591
 
449
592
  ## 0.3.6 (2026-03-24)
450
593
 
451
- - **Recursive array/object validation** — nested arrays (`array of array of string`) validated recursively. Object items validated even without `:properties` (e.g. `additionalProperties: false`).
452
- - **Deep symbolize in sample pre-validation** — array samples with string keys (`[{"name" => "Alice"}]`) correctly symbolized before schema validation.
594
+ - **Recursive array/object validation** - nested arrays (`array of array of string`) validated recursively. Object items validated even without `:properties` (e.g. `additionalProperties: false`).
595
+ - **Deep symbolize in sample pre-validation** - array samples with string keys (`[{"name" => "Alice"}]`) correctly symbolized before schema validation.
453
596
 
454
597
  ## 0.3.5 (2026-03-24)
455
598
 
456
- - **String constraints in SchemaValidator** — `minLength`/`maxLength` enforced for root and nested strings.
457
- - **Array item validation** — scalar items (string, integer) validated against items schema type and constraints.
458
- - **Non-JSON sample_response fails fast** — `sample_response("hello")` with object schema raises ArgumentError at definition time instead of silently passing.
459
- - **`max_tokens` in KNOWN_CONTEXT_KEYS** — no more spurious "Unknown context keys" warning.
460
- - **Duplicate models deduplicated** — `compare_models(models: ["m", "m"])` runs model once.
599
+ - **String constraints in SchemaValidator** - `minLength`/`maxLength` enforced for root and nested strings.
600
+ - **Array item validation** - scalar items (string, integer) validated against items schema type and constraints.
601
+ - **Non-JSON sample_response fails fast** - `sample_response("hello")` with object schema raises ArgumentError at definition time instead of silently passing.
602
+ - **`max_tokens` in KNOWN_CONTEXT_KEYS** - no more spurious "Unknown context keys" warning.
603
+ - **Duplicate models deduplicated** - `compare_models(models: ["m", "m"])` runs model once.
461
604
 
462
605
  ## 0.3.4 (2026-03-24)
463
606
 
464
- - **SchemaValidator validates non-object roots** — boolean, integer, number, array root schemas now enforce type, min/max, enum, minItems/maxItems. Previously only object schemas were validated.
465
- - **Removed passing cases = regression** — `regressed?` returns true when baseline had passing cases that are now missing. Prevents gate bypass by deleting eval cases.
466
- - **JSON string sample_response fixed** — `sample_response('{"name":"Alice"}')` correctly parsed for pre-validation instead of double-encoding.
467
- - **`context[:max_tokens]` forwarded** — overrides step's `max_output` for adapter call AND budget precheck.
607
+ - **SchemaValidator validates non-object roots** - boolean, integer, number, array root schemas now enforce type, min/max, enum, minItems/maxItems. Previously only object schemas were validated.
608
+ - **Removed passing cases = regression** - `regressed?` returns true when baseline had passing cases that are now missing. Prevents gate bypass by deleting eval cases.
609
+ - **JSON string sample_response fixed** - `sample_response('{"name":"Alice"}')` correctly parsed for pre-validation instead of double-encoding.
610
+ - **`context[:max_tokens]` forwarded** - overrides step's `max_output` for adapter call AND budget precheck.
468
611
 
469
612
  ## 0.3.3 (2026-03-23)
470
613
 
471
- - **Skipped cases visible in regression diff** — baseline PASS → current SKIP now detected as regression by `without_regressions` and `fail_on_regression`.
472
- - **Skip only on missing adapter** — eval runner no longer masks evaluator errors as SKIP. Only "No adapter configured" triggers skip.
473
- - **Array/Hash sample pre-validation** — `sample_response([{...}])` correctly validated against schema instead of silently skipping.
474
- - **`assume_model_exists: false` forwarded** — boolean `false` no longer dropped by truthiness check in adapter options.
475
- - **Duplicate case names caught at definition** — `add_case`/`verify` with same name raises immediately, not at run time.
614
+ - **Skipped cases visible in regression diff** - baseline PASS → current SKIP now detected as regression by `without_regressions` and `fail_on_regression`.
615
+ - **Skip only on missing adapter** - eval runner no longer masks evaluator errors as SKIP. Only "No adapter configured" triggers skip.
616
+ - **Array/Hash sample pre-validation** - `sample_response([{...}])` correctly validated against schema instead of silently skipping.
617
+ - **`assume_model_exists: false` forwarded** - boolean `false` no longer dropped by truthiness check in adapter options.
618
+ - **Duplicate case names caught at definition** - `add_case`/`verify` with same name raises immediately, not at run time.
476
619
 
477
620
  ## 0.3.2 (2026-03-23)
478
621
 
479
- - **Array response preserved** — `Adapters::RubyLLM` no longer stringifies Array content. Steps with `output_type Array` work correctly.
480
- - **Falsy prompt input** — `run(false)` and `build_messages(false)` pass `false` to dynamic prompt blocks instead of falling back to `instance_eval`.
481
- - **`retry_on` flatten** — `retry_on([:a, :b])` no longer wraps in nested array.
482
- - **Builder reset** — `Prompt::Builder` resets nodes on each build (no accumulation on reuse).
483
- - **Pipeline false output** — `output: false` no longer shows "(no output)" in pretty_print.
622
+ - **Array response preserved** - `Adapters::RubyLLM` no longer stringifies Array content. Steps with `output_type Array` work correctly.
623
+ - **Falsy prompt input** - `run(false)` and `build_messages(false)` pass `false` to dynamic prompt blocks instead of falling back to `instance_eval`.
624
+ - **`retry_on` flatten** - `retry_on([:a, :b])` no longer wraps in nested array.
625
+ - **Builder reset** - `Prompt::Builder` resets nodes on each build (no accumulation on reuse).
626
+ - **Pipeline false output** - `output: false` no longer shows "(no output)" in pretty_print.
484
627
 
485
628
  ## 0.3.1 (2026-03-23)
486
629
 
487
630
  Fixes from persona_tool production deployment (4 services migrated).
488
631
 
489
- - **Proc/Lambda in `expected_traits`** — `expected_traits: { score: ->(v) { v > 3 } }` now works.
490
- - **Zeitwerk eager-load** — `load_evals!` eager-loads `app/contracts/` and `app/steps/` before loading eval files. Fixes uninitialized constant errors in Rake tasks.
491
- - **Falsy values** — `expected: false`, `input: false`, `sample_response(nil)` all handled correctly.
492
- - **Context key forwarding** — `provider:` and `assume_model_exists:` forwarded to adapter. `schema:` and `max_tokens:` are step-level only (no split-brain).
493
- - **Deep-freeze immutability** — constructors never mutate caller's data.
632
+ - **Proc/Lambda in `expected_traits`** - `expected_traits: { score: ->(v) { v > 3 } }` now works.
633
+ - **Zeitwerk eager-load** - `load_evals!` eager-loads `app/contracts/` and `app/steps/` before loading eval files. Fixes uninitialized constant errors in Rake tasks.
634
+ - **Falsy values** - `expected: false`, `input: false`, `sample_response(nil)` all handled correctly.
635
+ - **Context key forwarding** - `provider:` and `assume_model_exists:` forwarded to adapter. `schema:` and `max_tokens:` are step-level only (no split-brain).
636
+ - **Deep-freeze immutability** - constructors never mutate caller's data.
494
637
 
495
638
  ## 0.3.0 (2026-03-23)
496
639
 
497
- Baseline regression detection — know when quality drops before users do.
640
+ Baseline regression detection - know when quality drops before users do.
498
641
 
499
642
  ### Features
500
643
 
501
- - **`report.save_baseline!`** — serialize eval results to `.eval_baselines/` (JSON, git-tracked)
502
- - **`report.compare_with_baseline`** — returns `BaselineDiff` with regressions, improvements, score_delta, new/removed cases
503
- - **`diff.regressed?`** — true when any previously-passing case now fails
504
- - **`without_regressions` RSpec chain** — `expect(Step).to pass_eval("x").without_regressions`
505
- - **RakeTask `fail_on_regression`** — blocks CI when regressions detected
506
- - **RakeTask `save_baseline`** — auto-save after successful run
507
- - **Migration guide** — `docs/guide/migration.md` with 7 patterns for adopting the gem in existing Rails apps
644
+ - **`report.save_baseline!`** - serialize eval results to `.eval_baselines/` (JSON, git-tracked)
645
+ - **`report.compare_with_baseline`** - returns `BaselineDiff` with regressions, improvements, score_delta, new/removed cases
646
+ - **`diff.regressed?`** - true when any previously-passing case now fails
647
+ - **`without_regressions` RSpec chain** - `expect(Step).to pass_eval("x").without_regressions`
648
+ - **RakeTask `fail_on_regression`** - blocks CI when regressions detected
649
+ - **RakeTask `save_baseline`** - auto-save after successful run
650
+ - **Migration guide** - `docs/guide/migration.md` with 7 patterns for adopting the gem in existing Rails apps
508
651
 
509
652
  ### Stats
510
653
 
@@ -514,21 +657,21 @@ Baseline regression detection — know when quality drops before users do.
514
657
 
515
658
  Production hardening from senior Rails review panel.
516
659
 
517
- - **`around_call` propagates exceptions** — no longer silently swallows DB errors, timeouts, etc. User who wants swallowing can rescue in their block.
518
- - **Nil section content skipped** — `section "X", nil` no longer renders `"null"` to the LLM. Section is omitted entirely.
519
- - **Range support in `expected:`** — `expected: { score: 1..5 }` works in `add_case`. Previously only Regexp was supported.
520
- - **`Trace#dig`** — `trace.dig(:usage, :input_tokens)` works on both Step and Pipeline traces.
660
+ - **`around_call` propagates exceptions** - no longer silently swallows DB errors, timeouts, etc. User who wants swallowing can rescue in their block.
661
+ - **Nil section content skipped** - `section "X", nil` no longer renders `"null"` to the LLM. Section is omitted entirely.
662
+ - **Range support in `expected:`** - `expected: { score: 1..5 }` works in `add_case`. Previously only Regexp was supported.
663
+ - **`Trace#dig`** - `trace.dig(:usage, :input_tokens)` works on both Step and Pipeline traces.
521
664
 
522
665
  ## 0.2.2 (2026-03-23)
523
666
 
524
667
  Fixes from first real-world integration (persona_tool).
525
668
 
526
- - **`around_call` fires per-run** — not per-attempt. With retry_policy, callback fires once with final result. Signature: `around_call { |step, input, result| ... }`
527
- - **`Result#trace` always `Trace` object** — never bare Hash. `result.trace.model` works on success AND failure.
528
- - **`around_call` exception safe** — warns and returns result instead of crashing.
529
- - **`model` DSL** — `model "gpt-4o-mini"` per-step. Priority: context > step DSL > global config.
530
- - **Test adapter `raw_output` always String** — Hash/Array normalized to `.to_json`.
531
- - **`Trace#dig`** — `trace.dig(:usage, :input_tokens)` works.
669
+ - **`around_call` fires per-run** - not per-attempt. With retry_policy, callback fires once with final result. Signature: `around_call { |step, input, result| ... }`
670
+ - **`Result#trace` always `Trace` object** - never bare Hash. `result.trace.model` works on success AND failure.
671
+ - **`around_call` exception safe** - warns and returns result instead of crashing.
672
+ - **`model` DSL** - `model "gpt-4o-mini"` per-step. Priority: context > step DSL > global config.
673
+ - **Test adapter `raw_output` always String** - Hash/Array normalized to `.to_json`.
674
+ - **`Trace#dig`** - `trace.dig(:usage, :input_tokens)` works.
532
675
 
533
676
  ## 0.2.1 (2026-03-23)
534
677
 
@@ -536,18 +679,18 @@ Production DX improvements from first real-world integration (persona_tool).
536
679
 
537
680
  ### Features
538
681
 
539
- - **`temperature` DSL** — `temperature 0.3` in step definition, overridable via `context: { temperature: 0.7 }`. RubyLLM handles per-model normalization natively.
540
- - **`around_call` hook** — callback for logging, metrics, observability. Replaces need for custom middleware.
541
- - **`build_messages` public** — inspect rendered prompt without running the step.
542
- - **`stub_step` RSpec helper** — `stub_step(MyStep, response: { ... })` reduces test boilerplate. Auto-included via `require "ruby_llm/contract/rspec"`.
543
- - **`estimate_cost` / `estimate_eval_cost`** — predict spend before API calls.
682
+ - **`temperature` DSL** - `temperature 0.3` in step definition, overridable via `context: { temperature: 0.7 }`. RubyLLM handles per-model normalization natively.
683
+ - **`around_call` hook** - callback for logging, metrics, observability. Replaces need for custom middleware.
684
+ - **`build_messages` public** - inspect rendered prompt without running the step.
685
+ - **`stub_step` RSpec helper** - `stub_step(MyStep, response: { ... })` reduces test boilerplate. Auto-included via `require "ruby_llm/contract/rspec"`.
686
+ - **`estimate_cost` / `estimate_eval_cost`** - predict spend before API calls.
544
687
 
545
688
  ### Fixes
546
689
 
547
- - **Reload lifecycle** — `load_evals!` clears definitions before re-loading. Railtie hooks `config.to_prepare` for development reload. `define_eval` warns on duplicate name (suppressed during reload).
548
- - **Pipeline eval cost** — uses `Pipeline::Trace#total_cost` (all steps), not just last step.
549
- - **Adapter isolation** — `compare_models` and `run_all_own_evals` deep-dup context per run.
550
- - **Offline mode** — cases without adapter return `:skipped` instead of crashing. Skipped cases excluded from score.
690
+ - **Reload lifecycle** - `load_evals!` clears definitions before re-loading. Railtie hooks `config.to_prepare` for development reload. `define_eval` warns on duplicate name (suppressed during reload).
691
+ - **Pipeline eval cost** - uses `Pipeline::Trace#total_cost` (all steps), not just last step.
692
+ - **Adapter isolation** - `compare_models` and `run_all_own_evals` deep-dup context per run.
693
+ - **Offline mode** - cases without adapter return `:skipped` instead of crashing. Skipped cases excluded from score.
551
694
  - **`expected_traits`** reachable from `define_eval` DSL via `add_case`.
552
695
  - **`verify`** raises when both positional and `expect:` keyword provided.
553
696
  - **`best_for`** excludes zero-score models from recommendation.
@@ -577,28 +720,28 @@ Contracts for LLM quality. Know which model to use, what it costs, and when accu
577
720
 
578
721
  ### Features
579
722
 
580
- - **`add_case` in `define_eval`** — `add_case "billing", input: "...", expected: { priority: "high" }` with partial matching. Supports `expected_traits:` for regex/range matching.
581
- - **`CaseResult` value objects** — `result.name`, `result.passed?`, `result.output`, `result.expected`, `result.mismatches` (structured diff), `result.cost`, `result.duration_ms`.
582
- - **`report.failures`** — returns only failed cases. `report.skipped` counts skipped (offline) cases.
583
- - **Model comparison** — `Step.compare_models("eval", models: %w[nano mini full])` runs same eval across models. Returns table with score/cost/latency per model. `comparison.best_for(min_score: 0.95)` returns cheapest model meeting threshold.
584
- - **Cost tracking** — `report.total_cost`, `report.avg_latency_ms`, per-case `result.cost`. Pipeline eval uses total pipeline cost, not just last step.
585
- - **Cost prediction** — `Step.estimate_cost(input:, model:)` and `Step.estimate_eval_cost("eval", models: [...])` predict spend before API calls.
586
- - **CI gating** — `pass_eval("regression").with_minimum_score(0.8).with_maximum_cost(0.01)`. RakeTask with suite-level `minimum_score` and `maximum_cost`.
587
- - **`RubyLLM::Contract.run_all_evals`** — discovers all Steps/Pipelines with evals, runs them all. Includes inherited evals.
588
- - **`RubyLLM::Contract::RakeTask`** — `rake ruby_llm_contract:eval` with `minimum_score`, `maximum_cost`, `fail_on_empty`, `eval_dirs`.
589
- - **Rails Railtie** — auto-loads eval files via `config.after_initialize` + `config.to_prepare` (supports development reload).
590
- - **Offline mode** — cases without adapter return `:skipped` instead of crashing. Skipped cases excluded from score/passed.
591
- - **Safe `define_eval`** — warns on duplicate name; suppressed during reload.
723
+ - **`add_case` in `define_eval`** - `add_case "billing", input: "...", expected: { priority: "high" }` with partial matching. Supports `expected_traits:` for regex/range matching.
724
+ - **`CaseResult` value objects** - `result.name`, `result.passed?`, `result.output`, `result.expected`, `result.mismatches` (structured diff), `result.cost`, `result.duration_ms`.
725
+ - **`report.failures`** - returns only failed cases. `report.skipped` counts skipped (offline) cases.
726
+ - **Model comparison** - `Step.compare_models("eval", models: %w[nano mini full])` runs same eval across models. Returns table with score/cost/latency per model. `comparison.best_for(min_score: 0.95)` returns cheapest model meeting threshold.
727
+ - **Cost tracking** - `report.total_cost`, `report.avg_latency_ms`, per-case `result.cost`. Pipeline eval uses total pipeline cost, not just last step.
728
+ - **Cost prediction** - `Step.estimate_cost(input:, model:)` and `Step.estimate_eval_cost("eval", models: [...])` predict spend before API calls.
729
+ - **CI gating** - `pass_eval("regression").with_minimum_score(0.8).with_maximum_cost(0.01)`. RakeTask with suite-level `minimum_score` and `maximum_cost`.
730
+ - **`RubyLLM::Contract.run_all_evals`** - discovers all Steps/Pipelines with evals, runs them all. Includes inherited evals.
731
+ - **`RubyLLM::Contract::RakeTask`** - `rake ruby_llm_contract:eval` with `minimum_score`, `maximum_cost`, `fail_on_empty`, `eval_dirs`.
732
+ - **Rails Railtie** - auto-loads eval files via `config.after_initialize` + `config.to_prepare` (supports development reload).
733
+ - **Offline mode** - cases without adapter return `:skipped` instead of crashing. Skipped cases excluded from score/passed.
734
+ - **Safe `define_eval`** - warns on duplicate name; suppressed during reload.
592
735
 
593
736
  ### Fixes
594
737
 
595
- - **P1: Eval files not autoloaded by Rails** — Railtie uses `load` (not Zeitwerk). Hooks into reloader for dev.
596
- - **P2: report.results returns raw Hashes** — now returns `CaseResult` objects.
597
- - **P3: No way to run all evals at once** — `Contract.run_all_evals` + Rake task.
598
- - **P4: String vs symbol key mismatch** — warns when `validate` or `verify` proc returns nil.
599
- - **Pipeline eval cost** — uses `Pipeline::Trace#total_cost` (all steps), not just last step.
600
- - **Reload lifecycle** — `load_evals!` clears definitions before re-loading. Registry filters stale hosts.
601
- - **Adapter isolation** — `compare_models` and `run_all_own_evals` deep-dup context per run.
738
+ - **P1: Eval files not autoloaded by Rails** - Railtie uses `load` (not Zeitwerk). Hooks into reloader for dev.
739
+ - **P2: report.results returns raw Hashes** - now returns `CaseResult` objects.
740
+ - **P3: No way to run all evals at once** - `Contract.run_all_evals` + Rake task.
741
+ - **P4: String vs symbol key mismatch** - warns when `validate` or `verify` proc returns nil.
742
+ - **Pipeline eval cost** - uses `Pipeline::Trace#total_cost` (all steps), not just last step.
743
+ - **Reload lifecycle** - `load_evals!` clears definitions before re-loading. Registry filters stale hosts.
744
+ - **Adapter isolation** - `compare_models` and `run_all_own_evals` deep-dup context per run.
602
745
 
603
746
  ### Verified with real API
604
747
 
@@ -621,16 +764,16 @@ Initial release.
621
764
 
622
765
  ### Features
623
766
 
624
- - **Step abstraction** — `RubyLLM::Contract::Step::Base` with prompt DSL, typed input/output
625
- - **Output schema** — declarative structure via ruby_llm-schema, sent to provider for enforcement
626
- - **Validate** — business logic checks (1-arity and 2-arity with input cross-validation)
627
- - **Retry with model escalation** — start cheap, auto-escalate on contract failure or network error
628
- - **Preflight limits** — `max_input`, `max_cost`, `max_output` refuse before calling the LLM
629
- - **Pipeline** — multi-step composition with fail-fast, timeout, token budget
630
- - **Eval** — offline contract verification with `define_eval`, `run_eval`, zero-verify auto-case
631
- - **Adapters** — RubyLLM (production), Test (deterministic specs)
632
- - **RSpec matchers** — `satisfy_contract`, `pass_eval`
633
- - **Structured trace** — model, latency, tokens, cost, attempt log per step
767
+ - **Step abstraction** - `RubyLLM::Contract::Step::Base` with prompt DSL, typed input/output
768
+ - **Output schema** - declarative structure via ruby_llm-schema, sent to provider for enforcement
769
+ - **Validate** - business logic checks (1-arity and 2-arity with input cross-validation)
770
+ - **Retry with model escalation** - start cheap, auto-escalate on contract failure or network error
771
+ - **Preflight limits** - `max_input`, `max_cost`, `max_output` refuse before calling the LLM
772
+ - **Pipeline** - multi-step composition with fail-fast, timeout, token budget
773
+ - **Eval** - offline contract verification with `define_eval`, `run_eval`, zero-verify auto-case
774
+ - **Adapters** - RubyLLM (production), Test (deterministic specs)
775
+ - **RSpec matchers** - `satisfy_contract`, `pass_eval`
776
+ - **Structured trace** - model, latency, tokens, cost, attempt log per step
634
777
 
635
778
  ### Robustness
636
779