ruby_llm-contract 0.8.0 → 0.10.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +79 -1
- data/Gemfile.lock +2 -2
- data/README.md +96 -37
- data/docs/architecture.md +50 -0
- data/docs/guide/best_practices.md +136 -0
- data/docs/guide/eval_first.md +192 -0
- data/docs/guide/getting_started.md +199 -0
- data/docs/guide/migration.md +185 -0
- data/docs/guide/multimodal_input.md +160 -0
- data/docs/guide/optimizing_retry_policy.md +131 -0
- data/docs/guide/output_schema.md +93 -0
- data/docs/guide/pipeline.md +154 -0
- data/docs/guide/prompt_ast.md +76 -0
- data/docs/guide/rails_integration.md +218 -0
- data/docs/guide/relation_to_agent.md +52 -0
- data/docs/guide/relation_to_tribunal.md +135 -0
- data/docs/guide/testing.md +282 -0
- data/docs/guide/why.md +103 -0
- data/lib/ruby_llm/contract/adapters/ruby_llm.rb +9 -1
- data/lib/ruby_llm/contract/concerns/eval_host.rb +6 -9
- data/lib/ruby_llm/contract/concerns/stub_helpers.rb +97 -0
- data/lib/ruby_llm/contract/contract/definition.rb +2 -0
- data/lib/ruby_llm/contract/cost_calculator.rb +11 -2
- data/lib/ruby_llm/contract/eval/recommender.rb +3 -1
- data/lib/ruby_llm/contract/eval/retry_optimizer.rb +16 -13
- data/lib/ruby_llm/contract/eval.rb +13 -0
- data/lib/ruby_llm/contract/minitest.rb +6 -108
- data/lib/ruby_llm/contract/pipeline/result.rb +1 -1
- data/lib/ruby_llm/contract/rake_task/suite_gate.rb +117 -0
- data/lib/ruby_llm/contract/rake_task.rb +30 -51
- data/lib/ruby_llm/contract/rspec/helpers.rb +9 -123
- data/lib/ruby_llm/contract/step/base.rb +56 -24
- data/lib/ruby_llm/contract/step/dsl.rb +91 -63
- data/lib/ruby_llm/contract/step/limit_checker.rb +34 -1
- data/lib/ruby_llm/contract/step/retry_executor.rb +6 -13
- data/lib/ruby_llm/contract/step/runner.rb +22 -20
- data/lib/ruby_llm/contract/step/runner_config.rb +26 -0
- data/lib/ruby_llm/contract/version.rb +1 -1
- data/lib/ruby_llm/contract.rb +1 -0
- data/ruby_llm-contract.gemspec +5 -1
- metadata +18 -4
- data/.rspec +0 -3
- data/.rubycritic.yml +0 -8
- data/.simplecov +0 -22
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: ac8c19463285a3e2c0f9050e835125454b21d2b1500dcb5bad2dd4a114666bd0
|
|
4
|
+
data.tar.gz: 339c0609e3dccbf55da649c3470ed234c5abafa55c8649aee5d8dd41f9bd82ec
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: c04fea66393868fbdb369f5a75eaeb66a28172c8f9ddabd0a8c053fd55e5daa06174aeeb00c4bb7e9aabaf613bd1ef4ea354ff7d0f4b817f2ec41af1a4c61667
|
|
7
|
+
data.tar.gz: 71f502ebe497ede22df508d3b56396660fbd1eba109066885d5e905724bd82d30a0d46fddfe4f1c1e7efe47f95ea41b4f61341a4698c12d3f3a397cab2510068
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,83 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.10.2 (2026-06-10)
|
|
4
|
+
|
|
5
|
+
Patch release: ship the `docs/guide/` directory inside the gem so adopters and LLM integration agents can read the manuals locally (via `bundle show ruby_llm-contract` or `gem unpack`) without an internet round-trip to GitHub. No code behavior change.
|
|
6
|
+
|
|
7
|
+
### Added
|
|
8
|
+
|
|
9
|
+
- **`docs/guide/*` is now packaged with the gem** (14 files, ~120 KB). Previously the README's "See also" links pointed at `docs/guide/getting_started.md`, `docs/guide/optimizing_retry_policy.md`, etc., but those files were stripped from the published gem - LLM integration agents (Cursor, Claude Code, Copilot) reported "no documentation in the gem" because the links 404'd locally. `docs/ideas/` and `doc/decisions/` remain excluded.
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- **`models:` keyword form documented + pinned for hash configs.** `retry_policy models: ["gpt-5-nano", { model: "gpt-5-mini", reasoning_effort: "high" }]` is now covered by a spec and shown in [getting_started.md](docs/guide/getting_started.md). The block form (`escalate(...)`) and the keyword form share the same `@configs` storage; both forms accept config hashes.
|
|
14
|
+
- **`reasoning_effort` examples corrected to gpt-5 family.** Pre-0.10.2 docs and specs paired `reasoning_effort` with `gpt-4.1-*` model names, which is incorrect: `gpt-4.1` is not a reasoning model. Updated across `docs/guide/getting_started.md`, `docs/guide/optimizing_retry_policy.md`, `CHANGELOG.md`, and 7 spec files. Non-reasoning `gpt-4.1` examples (model fallback chains without `reasoning_effort`) are unchanged.
|
|
15
|
+
- **Version mentions corrected in README and multimodal guide.** README FAQ ("Upgraded to 0.9.0 - why?") now reads "Upgraded to 0.10.0 from 0.8.x"; `docs/guide/multimodal_input.md` references `0.10.0+` and `0.10.x` instead of `0.9.0`. 0.9.0 and 0.9.1 were tagged but never published to rubygems; adopters jump from 0.8.0 directly to 0.10.x.
|
|
16
|
+
|
|
17
|
+
## 0.10.1 (2026-06-01)
|
|
18
|
+
|
|
19
|
+
Patch release fixing gem packaging. 0.10.0 was yanked from rubygems.org due to the issue documented below; 0.10.1 is the recommended upgrade target. No code behavior change vs 0.10.0.
|
|
20
|
+
|
|
21
|
+
### Fixed
|
|
22
|
+
|
|
23
|
+
- **Gem no longer ships internal tracker / dev configs.** Excluded from `spec.files`: `TODO.md`, `.rspec`, `.rubycritic.yml`, `.simplecov`, and the `.revive/` directory. Pre-0.10.1 the published gem contained these files; adopters who already extracted 0.10.0 can safely delete them.
|
|
24
|
+
|
|
25
|
+
## 0.10.0 (2026-06-01)
|
|
26
|
+
|
|
27
|
+
First published release since 0.8.0. Consolidates work originally tagged as 0.9.0 (multimodal input) and 0.9.1 (internal quality refactor), neither of which was pushed to rubygems. Adopters upgrading from 0.8.0 should read the **Behavioural change** and **Breaking changes** sections below before installing.
|
|
28
|
+
|
|
29
|
+
### Breaking changes
|
|
30
|
+
|
|
31
|
+
- **`validate(description, &block)` and `Definition#invariant(description, &block)` now raise `ArgumentError` when `description` is `nil` or empty.** Pre-0.10.0 the empty descriptor was silently accepted and produced `""` entries in `result.validation_errors`, making debugging impossible. Codex audit found zero production use sites across `lib/`, `examples/`, `README` — only the regression-marker test certifying the bug.
|
|
32
|
+
|
|
33
|
+
### Migration
|
|
34
|
+
|
|
35
|
+
Ensure every `validate` / `invariant` call has a non-empty descriptor (this is already how every README example writes them):
|
|
36
|
+
|
|
37
|
+
```ruby
|
|
38
|
+
# Before (silently accepted, produced "" in validation_errors):
|
|
39
|
+
validate("") { |o| o[:score].between?(0, 100) }
|
|
40
|
+
|
|
41
|
+
# After (required):
|
|
42
|
+
validate("score in range 0-100") { |o| o[:score].between?(0, 100) }
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
### Added
|
|
46
|
+
|
|
47
|
+
- **Multimodal input via `context: { attachment: ... }`** — pass a file/IO/URL through `Step.run(input, context: { attachment: path })`; the adapter forwards it to `RubyLLM::Chat#ask(content, with: attachment)`. RubyLLM normalises wire format per provider (Anthropic url/base64, OpenAI `image_url`/`file`, Gemini `inline_data`). Multi-attachment supported natively (`with: [pdf1, pdf2]` or `with: { images: [...], pdfs: [...] }`). See [multimodal input guide](docs/guide/multimodal_input.md) and [ADR-0022](doc/decisions/ADR-0022-v09-multimodal-input.md).
|
|
48
|
+
- **`attachment_token_estimate(n)` class macro** — adopter-declared conservative estimate of attachment input tokens. Applied to BOTH runtime (`limit_checker`) and pre-flight (`estimate_cost`) — same source of truth, no estimate/runtime drift.
|
|
49
|
+
- **`on_unknown_attachment_size(:refuse | :warn)` class macro** — mirrors `on_unknown_pricing` opt-out semantics. Defaults to `:refuse`. Never settable as global default — same invariant as `max_cost` fail-closed.
|
|
50
|
+
|
|
51
|
+
### Behavioural change (READ BEFORE UPGRADING)
|
|
52
|
+
|
|
53
|
+
- **Contracts with `max_cost` or `max_input` AND `context[:attachment]` set AND no `attachment_token_estimate` declared → REFUSE with `:limit_exceeded`.** This is fail-closed semantics: the gem cannot bound vision/PDF token cost without an adopter-declared estimate. Opt out per-step with `on_unknown_attachment_size :warn`. Text-only contracts and contracts without `max_cost`/`max_input` are unaffected.
|
|
54
|
+
|
|
55
|
+
### Changed
|
|
56
|
+
|
|
57
|
+
- **`run_eval` (no args) return shape pinned to `Hash<String, Report>` keyed by eval name.** Documents the existing contract used by `RubyLLM::Contract::RakeTask#collect_host_reports` and adopters. No runtime change vs 0.8.0 — only the spec assertion now locks the shape.
|
|
58
|
+
- **`Parser.parse(text, strategy: :json)` first-bracket-wins boundary documented.** Extraction commits to the first balanced `{` or `[` structure and does NOT retry on later candidates. Empty `{}` followed by real JSON parses as the empty Hash; non-JSON `{braces}` before real JSON raises `ParseError`. No runtime change — this codifies long-standing behavior with explicit boundary tests.
|
|
59
|
+
|
|
60
|
+
### Fixed
|
|
61
|
+
|
|
62
|
+
- **`with_retry_disabled` no longer mutates the step class's singleton method.** The optimizer now passes `retry_policy_override: nil` through `context:` to `compare_models`, which `Step::Base#runtime_settings` already honours. Removes a concurrency hazard where two parallel `optimize_retry_policy` calls on the same step would race on the singleton restore in `ensure`.
|
|
63
|
+
- **`CostCalculator.find_model` exposed as a public class method.** Removes two `CostCalculator.send(:find_model, ...)` workarounds in `Step::Base#estimate_cost`. The `estimated_cost_for` helper is gone — `estimate_cost` now routes through the existing public `CostCalculator.calculate(model_name:, usage:)`.
|
|
64
|
+
- **`stub_step` unified on a single storage path.** Both block and non-block forms now write to `RubyLLM::Contract.step_adapter_overrides` (thread-local). The `around(:each)` hook in `rspec.rb` handles cleanup between examples. Removes the prior `allow(step).to receive(:run)` branch.
|
|
65
|
+
|
|
66
|
+
### Internal
|
|
67
|
+
|
|
68
|
+
- **Anti-facade audit complete: 89/89 spec files under per-test 17-mode walk** (Phase A: 26 specs, Phase C: 63 specs via parallel Codex fan-out). Net +30 strengthened tests against mutation-blind assertions, zero public API change beyond the breaking entry above.
|
|
69
|
+
- **Dead `ObjectSpace.each_object(Class)` fallback removed** in `concerns/eval_host.rb#register_subclasses`. The gemspec requires Ruby `>= 3.2.0`, so `Class#subclasses` (Ruby ≥ 3.1) is always available; the legacy fallback was unreachable code that would have iterated all loaded classes O(n) and was not thread-safe.
|
|
70
|
+
|
|
71
|
+
### Deferred (not in 0.10.x)
|
|
72
|
+
|
|
73
|
+
- `add_history` multi-turn replay of prior attachments — single-turn multimodal supported; follow-up questions on the same document deferred to a later release.
|
|
74
|
+
- Streaming + attachment — contract steps remain synchronous.
|
|
75
|
+
- Provider-specific attachment size caps — surface only via `attachment_token_estimate` calibration; consult provider docs.
|
|
76
|
+
|
|
77
|
+
### Tests
|
|
78
|
+
|
|
79
|
+
- Suite: 1401 examples / 0 failures / 7 pending (was 1346/0/8 at 0.8.0).
|
|
80
|
+
|
|
3
81
|
## 0.8.0 (2026-04-26)
|
|
4
82
|
|
|
5
83
|
Narrative repositioning + small API additions. Internal architecture unchanged: no `Step::Base` refactor, no breaking changes to existing DSL.
|
|
@@ -190,7 +268,7 @@ end
|
|
|
190
268
|
- **`Step.recommend`** — `ClassifyTicket.recommend("eval", candidates: [...], min_score: 0.95)` runs eval on all candidates and returns a `Recommendation` with optimal model, retry chain, rationale, savings vs current config, and `to_dsl` code output.
|
|
191
269
|
- **Candidates as configurations** — `candidates:` accepts `{ model:, reasoning_effort: }` hashes, not just model name strings. `gpt-5-mini` with `reasoning_effort: "low"` is a different candidate than with `"high"`.
|
|
192
270
|
- **`compare_models` extended** — new `candidates:` parameter alongside existing `models:` (backward compatible). Candidate labels include reasoning effort in output table.
|
|
193
|
-
- **Per-attempt `reasoning_effort` in retry policies** — `escalate` accepts config hashes: `escalate({ model: "gpt-
|
|
271
|
+
- **Per-attempt `reasoning_effort` in retry policies** — `escalate` accepts config hashes: `escalate({ model: "gpt-5-nano" }, { model: "gpt-5-mini", reasoning_effort: "high" })`. Each attempt gets its own reasoning_effort forwarded to the provider.
|
|
194
272
|
- **`pass_rate_ratio`** — numeric float (0.0–1.0) on `Report` and `ReportStats`, complementing the string `pass_rate` (`"3/5"`).
|
|
195
273
|
- **History entries enriched** — `save_history!` accepts `reasoning_effort:` and stores `model`, `reasoning_effort`, `pass_rate_ratio` in JSONL entries.
|
|
196
274
|
|
data/Gemfile.lock
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
PATH
|
|
2
2
|
remote: .
|
|
3
3
|
specs:
|
|
4
|
-
ruby_llm-contract (0.
|
|
4
|
+
ruby_llm-contract (0.10.2)
|
|
5
5
|
dry-types (~> 1.7)
|
|
6
6
|
ruby_llm (~> 1.12)
|
|
7
7
|
ruby_llm-schema (~> 0.3)
|
|
@@ -258,7 +258,7 @@ CHECKSUMS
|
|
|
258
258
|
rubocop-ast (1.49.1) sha256=4412f3ee70f6fe4546cc489548e0f6fcf76cafcfa80fa03af67098ffed755035
|
|
259
259
|
ruby-progressbar (1.13.0) sha256=80fc9c47a9b640d6834e0dc7b3c94c9df37f08cb072b7761e4a71e22cff29b33
|
|
260
260
|
ruby_llm (1.14.0) sha256=57c6f7034fc4a44504ea137d70f853b07824f1c1cdbe774ab3ab3522e7098deb
|
|
261
|
-
ruby_llm-contract (0.
|
|
261
|
+
ruby_llm-contract (0.10.2)
|
|
262
262
|
ruby_llm-schema (0.3.0) sha256=a591edc5ca1b7f0304f0e2261de61ba4b3bea17be09f5cf7558153adfda3dec6
|
|
263
263
|
ruby_parser (3.22.0) sha256=1eb4937cd9eb220aa2d194e352a24dba90aef00751e24c8dfffdb14000f15d23
|
|
264
264
|
rubycritic (4.12.0) sha256=024fed90fe656fa939f6ea80aab17569699ac3863d0b52fd72cb99892247abc8
|
data/README.md
CHANGED
|
@@ -2,9 +2,9 @@
|
|
|
2
2
|
|
|
3
3
|
**Contracts + Evals for [ruby_llm](https://github.com/crmne/ruby_llm).**
|
|
4
4
|
|
|
5
|
-
Your eval passed. Prod broke anyway? This gem wraps `RubyLLM::Chat` with input/output contracts, business-rule validation, retry with model escalation on validation failure, pre-flight cost ceilings, and
|
|
5
|
+
Your eval passed. Prod broke anyway? This gem wraps `RubyLLM::Chat` with input/output contracts, business-rule validation, retry with model escalation on validation failure, pre-flight cost ceilings, and a regression-eval framework — so a flaky cheap-model call escalates to a stronger model instead of shipping garbage to your user.
|
|
6
6
|
|
|
7
|
-
`ruby_llm` handles the HTTP side (rate limits, timeouts, streaming, tool calls, embeddings). This gem handles what the model *returned
|
|
7
|
+
`ruby_llm` handles the HTTP side (rate limits, timeouts, streaming, tool calls, embeddings). This gem handles what the model *returned* at **runtime**: schema validation, business rules, model escalation on failed validation, regression datasets that gate prompt/model changes in CI.
|
|
8
8
|
|
|
9
9
|
## Install
|
|
10
10
|
|
|
@@ -13,23 +13,24 @@ gem "ruby_llm-contract"
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
```ruby
|
|
16
|
-
RubyLLM.configure
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
Works with any `ruby_llm` provider (OpenAI, Anthropic, Gemini, etc).
|
|
21
|
-
|
|
22
|
-
## Do I need this?
|
|
16
|
+
RubyLLM.configure do |c|
|
|
17
|
+
c.openai_api_key = ENV["OPENAI_API_KEY"]
|
|
18
|
+
c.default_model = "gpt-4.1-mini" # used when a Step has no explicit model
|
|
19
|
+
end
|
|
23
20
|
|
|
24
|
-
|
|
21
|
+
# Required: boots the gem so `Step.run` knows how to talk to your LLM.
|
|
22
|
+
# Empty block is fine. Pass options here if you need them (e.g. `c.logger`).
|
|
23
|
+
RubyLLM::Contract.configure { }
|
|
24
|
+
```
|
|
25
25
|
|
|
26
|
-
|
|
26
|
+
Works with any `ruby_llm` provider (OpenAI, Anthropic, Gemini, etc). Requires `ruby_llm ~> 1.12` and Ruby ≥ 3.2.
|
|
27
27
|
|
|
28
28
|
## Example
|
|
29
29
|
|
|
30
30
|
A Rails app takes article text extracted from a user-submitted URL and wants to show a summary card: a short TL;DR, 3–5 key takeaways, and a tone label. The output has to fit the UI (TL;DR under 200 chars) and the schema has to be strict enough to render without conditionals.
|
|
31
31
|
|
|
32
32
|
```ruby
|
|
33
|
+
# app/contracts/summarize_article.rb
|
|
33
34
|
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
34
35
|
prompt <<~PROMPT
|
|
35
36
|
Summarize this article for a UI card. Return a short TL;DR,
|
|
@@ -45,48 +46,93 @@ class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
|
45
46
|
end
|
|
46
47
|
|
|
47
48
|
validate("TL;DR fits the card") { |o, _| o[:tldr].length <= 200 }
|
|
48
|
-
validate("takeaways are unique") { |o, _| o[:takeaways]
|
|
49
|
+
validate("takeaways are unique") { |o, _| o[:takeaways] == o[:takeaways].uniq }
|
|
49
50
|
|
|
50
|
-
|
|
51
|
+
# Cheapest first; last step adds a reasoning model with more thinking.
|
|
52
|
+
retry_policy do
|
|
53
|
+
escalate "gpt-4.1-nano",
|
|
54
|
+
"gpt-4.1-mini",
|
|
55
|
+
{ model: "gpt-5", reasoning_effort: "high" }
|
|
56
|
+
end
|
|
51
57
|
end
|
|
52
58
|
|
|
53
59
|
result = SummarizeArticle.run(article_text)
|
|
54
|
-
result.
|
|
55
|
-
result.
|
|
56
|
-
result.trace[:
|
|
60
|
+
result.status # => :ok (or :validation_failed if all steps fail)
|
|
61
|
+
result.parsed_output # => { tldr: "...", takeaways: [...], tone: "..." }
|
|
62
|
+
result.trace[:model] # => "gpt-4.1-mini" (winning step)
|
|
63
|
+
result.trace[:cost] # => 0.000520 (total across all attempts)
|
|
64
|
+
|
|
65
|
+
result.trace[:attempts]
|
|
66
|
+
# => [
|
|
67
|
+
# {
|
|
68
|
+
# attempt: 1,
|
|
69
|
+
# model: "gpt-4.1-nano",
|
|
70
|
+
# status: :validation_failed,
|
|
71
|
+
# usage: { input_tokens: 256, output_tokens: 84 },
|
|
72
|
+
# latency_ms: 45,
|
|
73
|
+
# cost: 0.000100
|
|
74
|
+
# },
|
|
75
|
+
# {
|
|
76
|
+
# attempt: 2,
|
|
77
|
+
# model: "gpt-4.1-mini",
|
|
78
|
+
# status: :ok,
|
|
79
|
+
# usage: { input_tokens: 256, output_tokens: 92 },
|
|
80
|
+
# latency_ms: 92,
|
|
81
|
+
# cost: 0.000420
|
|
82
|
+
# }
|
|
83
|
+
# ]
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
If the response is malformed, the TL;DR overflows the card, or the takeaway count is off, the gem moves to the next step. This is model **escalation**, not a fallback list — each step is an independent config (`model`, `reasoning_effort`), so the retry policy spends more compute only when the cheaper one couldn't satisfy the contract.
|
|
87
|
+
|
|
88
|
+
### Add a CI gate in 6 lines
|
|
89
|
+
|
|
90
|
+
The contract above already runs in production. The same `Step` doubles as the unit your regression eval runs against:
|
|
91
|
+
|
|
92
|
+
```ruby
|
|
93
|
+
SummarizeArticle.define_eval("regression") do
|
|
94
|
+
# `expected:` is a partial hash match — only listed keys check parsed_output.
|
|
95
|
+
add_case "neutral release",
|
|
96
|
+
input: "Ruby 3.4 shipped frozen string literals...",
|
|
97
|
+
expected: { tone: "analytical" }
|
|
98
|
+
add_case "outage post",
|
|
99
|
+
input: "Service was down for 4 hours...",
|
|
100
|
+
expected: { tone: "negative" }
|
|
101
|
+
end
|
|
102
|
+
|
|
103
|
+
# in CI (RSpec):
|
|
104
|
+
expect(SummarizeArticle).to pass_eval("regression").without_regressions
|
|
57
105
|
```
|
|
58
106
|
|
|
59
|
-
|
|
107
|
+
A bad prompt edit or model swap that drops accuracy on the frozen dataset → red CI, blocked merge. The first CI run records a baseline; subsequent runs compare against it. Every production miss should become the next `add_case`. See [Prevent silent prompt regressions](docs/guide/eval_first.md) for the full flywheel.
|
|
108
|
+
|
|
109
|
+
## Do I need this?
|
|
110
|
+
|
|
111
|
+
Use this if LLM output affects production behaviour, money, user trust, or downstream code. You probably don't need it if you have one low-risk prompt, manually inspect every result, or only generate best-effort prose.
|
|
60
112
|
|
|
61
|
-
|
|
113
|
+
Already using structured outputs from your provider? This gem adds business-rule validation, retry with model escalation, evals, regression gating, and test stubs on top of them — the layer that stops schema-valid-but-wrong output from reaching users. See [Why contracts?](docs/guide/why.md) for the four production failure modes the gem exists for.
|
|
62
114
|
|
|
63
115
|
## Most useful next
|
|
64
116
|
|
|
65
117
|
Everything below is optional — the example above is a complete step. Reach for these when one step isn't enough.
|
|
66
118
|
|
|
67
|
-
- **[CI regression gates](docs/guide/getting_started.md)** —
|
|
68
|
-
- **[Find the cheapest viable fallback list](docs/guide/optimizing_retry_policy.md)** —
|
|
69
|
-
- **[A/B test prompts](docs/guide/eval_first.md)** —
|
|
70
|
-
- **[Budget caps](docs/guide/getting_started.md)** —
|
|
71
|
-
- **[Reasoning effort / thinking config](docs/guide/optimizing_retry_policy.md)** —
|
|
119
|
+
- **[CI regression gates](docs/guide/getting_started.md)** — block CI when accuracy drops on a model update or prompt tweak.
|
|
120
|
+
- **[Find the cheapest viable fallback list](docs/guide/optimizing_retry_policy.md)** — empirically pick the cheapest model chain that still passes your evals.
|
|
121
|
+
- **[A/B test prompts](docs/guide/eval_first.md)** — measure whether a new prompt is safe to ship before merging.
|
|
122
|
+
- **[Budget caps](docs/guide/getting_started.md)** — refuse the request pre-flight when an estimate exceeds the limit.
|
|
123
|
+
- **[Reasoning effort / thinking config](docs/guide/optimizing_retry_policy.md)** — Anthropic / OpenAI thinking configuration on the Step class.
|
|
72
124
|
|
|
73
|
-
Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and
|
|
125
|
+
Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and per-step models.
|
|
74
126
|
|
|
75
127
|
## Relation to `RubyLLM::Agent`
|
|
76
128
|
|
|
77
|
-
`RubyLLM::Agent` (since RubyLLM 1.12)
|
|
129
|
+
`Step::Base` and `RubyLLM::Agent` (since RubyLLM 1.12) are **siblings** targeting the same niche: reusable, class-based prompts. Both call into `RubyLLM::Chat` directly — Step does not wrap Agent. Step adds the contract layer: `validate` (business invariants), `retry_policy escalate(...)` (model escalation on validation failure), `max_cost` pre-flight refusal, regression-eval framework, pipeline composition. **[Full feature mapping →](docs/guide/relation_to_agent.md)**
|
|
78
130
|
|
|
79
|
-
|
|
80
|
-
|---|---|
|
|
81
|
-
| `model`, `temperature`, `schema`, `instructions`, `tools`, `thinking` | covered by both — same idea, different DSL surface |
|
|
82
|
-
| `validate :rule do |out| ... end` business invariants | only here |
|
|
83
|
-
| `retry_policy escalate(...)` model escalation on validation failure | only here (different from RubyLLM's network-level retry) |
|
|
84
|
-
| `max_cost` / `max_input` / `max_output` pre-flight refusal | only here |
|
|
85
|
-
| `define_eval` + baseline regression + `compare_models` + `optimize_retry_policy` | only here (RubyLLM does not ship an eval framework) |
|
|
86
|
-
| Pipeline composition with `step SomeStep, as: :alias` | only here (RubyLLM intentionally leaves workflows as plain Ruby) |
|
|
87
|
-
| `around_call`, named `observe` hooks with pass/fail in trace | only here |
|
|
131
|
+
## Relation to `ruby_llm-tribunal`
|
|
88
132
|
|
|
89
|
-
|
|
133
|
+
Different layers, complementary. [`ruby_llm-tribunal`](https://github.com/Alqemist-labs/ruby_llm-tribunal) is a **test framework** that grades outputs **after they've reached your code**, typically in a spec. `ruby_llm-contract` is **runtime** — schema + `validate` rules gate the call **before the output reaches your code**, retry/escalate attempts to recover from failed outputs, `max_cost` refuses pre-flight. Our `define_eval` is *regression* (does this prompt/model still pass on a frozen dataset?), not *grading*.
|
|
134
|
+
|
|
135
|
+
**One-liner:** Tribunal answers *"is this output good?"* (fail → red test in CI). Contract answers *"what do we do when it isn't?"* (fail → retry/escalate, or fail closed). **[Visual flows + coexistence patterns →](docs/guide/relation_to_tribunal.md)**
|
|
90
136
|
|
|
91
137
|
## Docs
|
|
92
138
|
|
|
@@ -95,6 +141,8 @@ Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and
|
|
|
95
141
|
| Guide | What it does for your app |
|
|
96
142
|
|-------|---------------------------|
|
|
97
143
|
| [Why contracts?](docs/guide/why.md) | Recognise the four production failures the gem exists for |
|
|
144
|
+
| [Relation to RubyLLM::Agent](docs/guide/relation_to_agent.md) | Sibling abstractions; what each adds; runtime call path; coexistence patterns |
|
|
145
|
+
| [Relation to ruby_llm-tribunal](docs/guide/relation_to_tribunal.md) | Different layers (test framework vs runtime contract); visual flows; integration recipes |
|
|
98
146
|
| [Getting Started](docs/guide/getting_started.md) | Walk the full feature set on one concrete step |
|
|
99
147
|
| [Rails integration](docs/guide/rails_integration.md) | Directory, initializer, jobs, logging, specs, CI gate — 7 FAQs for Rails devs |
|
|
100
148
|
| [Adopt in an existing Rails app](docs/guide/migration.md) | Replace raw `LlmClient.call` with a contract, Before/After |
|
|
@@ -103,12 +151,23 @@ Also supports [multi-step pipelines](docs/guide/pipeline.md) with fail-fast and
|
|
|
103
151
|
| [Write validate rules that catch real bugs](docs/guide/best_practices.md) | Patterns for cross-input checks and content-quality rules |
|
|
104
152
|
| [Stub LLM calls in tests](docs/guide/testing.md) | Deterministic specs, RSpec + Minitest matchers |
|
|
105
153
|
| [Chain LLM calls into a pipeline](docs/guide/pipeline.md) | Multi-step with fail-fast and per-step models |
|
|
154
|
+
| [Multimodal input (PDF / image / audio)](docs/guide/multimodal_input.md) | Route attachments through the contract; `attachment_token_estimate`, fail-closed cost, calibration table |
|
|
106
155
|
| [Schema DSL reference](docs/guide/output_schema.md) | Every constraint, nested objects, pattern table |
|
|
107
156
|
| [Prompt DSL reference](docs/guide/prompt_ast.md) | `system` / `rule` / `section` / `example` / `user` nodes |
|
|
108
157
|
|
|
109
|
-
##
|
|
158
|
+
## Status & versioning
|
|
159
|
+
|
|
160
|
+
Pre-1.0 (currently **0.10.2**). Semver tracked; breaking changes flagged in [CHANGELOG](CHANGELOG.md). Pin `~> 0.10.2` until 1.0 ships.
|
|
161
|
+
|
|
162
|
+
## FAQ
|
|
163
|
+
|
|
164
|
+
**Thread-safe / Sidekiq?** Yes. Each `Step.run` builds an isolated `RubyLLM::Chat`; class-level state (`output_schema`, `validate`, `retry_policy`) is set up once at class load and read-only afterwards. Safe to run from concurrent jobs/threads.
|
|
165
|
+
|
|
166
|
+
**How do I stub `Step.run` in specs?** Include `RubyLLM::Contract::RSpec::Helpers` and use `stub_step(MyStep, response: { ... })`. The block form scopes the stub to one `it`. See [testing guide](docs/guide/testing.md).
|
|
167
|
+
|
|
168
|
+
**Where in a Rails app?** Default `app/contracts/`. The Railtie reloads `app/contracts/eval/` and `app/steps/eval/` in development; any autoloaded directory also works. See [Rails integration](docs/guide/rails_integration.md).
|
|
110
169
|
|
|
111
|
-
|
|
170
|
+
**Upgraded to 0.10.0 from 0.8.x and my contract started refusing — why?** 0.10.0 added multimodal input. If your contract has `max_cost` or `max_input` set AND now receives `context: { attachment: ... }`, you must declare `attachment_token_estimate(n)` (conservative input-token budget for the attachment) — otherwise the call fails closed with `:limit_exceeded`. The gem cannot bound vision/PDF cost without your estimate. Opt out per-step with `on_unknown_attachment_size :warn`. Text-only contracts are unaffected. See [multimodal input guide](docs/guide/multimodal_input.md).
|
|
112
171
|
|
|
113
172
|
## License
|
|
114
173
|
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
# Architecture
|
|
2
|
+
|
|
3
|
+
```
|
|
4
|
+
RubyLLM::Contract::Pipeline::Base # optional: compose steps
|
|
5
|
+
├── Pipeline::Runner # sequential execution, fail-fast, trace, timeout
|
|
6
|
+
└── Pipeline::Result # per-step outputs + aggregated trace
|
|
7
|
+
|
|
8
|
+
RubyLLM::Contract::Step::Base # single contracted step
|
|
9
|
+
├── Step::Dsl # DSL macros (prompt, validate, output_schema, etc.)
|
|
10
|
+
├── Step::RetryPolicy # attempts, models, reasoning_effort, retry_on
|
|
11
|
+
├── Step::RetryExecutor # retry loop driven by RetryPolicy
|
|
12
|
+
├── Step::LimitChecker # preflight cost / input / output checks
|
|
13
|
+
├── Step::Runner # runtime flow (single attempt)
|
|
14
|
+
├── Step::Result # status + outputs + errors + trace
|
|
15
|
+
├── Step::Trace # model, latency, tokens, cost, attempts
|
|
16
|
+
├── Prompt::AST # structured prompt (immutable)
|
|
17
|
+
│ ├── Prompt::Builder # DSL: system, rule, example, user, section
|
|
18
|
+
│ └── Prompt::Renderer # AST → messages array
|
|
19
|
+
├── Contract::Definition # parse strategy + validates
|
|
20
|
+
│ ├── Contract::Parser # :json / :text (auto-inferred from output type)
|
|
21
|
+
│ ├── Contract::Validator # runs parse + schema + validates + observations
|
|
22
|
+
│ └── Contract::SchemaValidator # JSON Schema validation (nested)
|
|
23
|
+
├── CostCalculator # per-step cost estimation from model pricing
|
|
24
|
+
├── TokenEstimator # input-token count estimation for limit checks
|
|
25
|
+
├── estimate_cost # single-call cost estimate (class method)
|
|
26
|
+
├── estimate_eval_cost # cost estimate for a full eval across models
|
|
27
|
+
└── Adapters::Base # provider interface
|
|
28
|
+
├── Adapters::RubyLLM # real LLM calls via ruby_llm (any provider)
|
|
29
|
+
└── Adapters::Test # canned responses for specs and examples
|
|
30
|
+
|
|
31
|
+
RubyLLM::Contract::Eval # quality measurement
|
|
32
|
+
├── Eval::EvalDefinition # define_eval DSL (verify, add_case, default_input, sample_response)
|
|
33
|
+
├── Eval::Dataset # test cases
|
|
34
|
+
├── Eval::Runner # sequential or concurrent execution
|
|
35
|
+
├── Eval::Report # score, pass_rate, per-case results
|
|
36
|
+
├── Eval::AggregatedReport # merged reports across models or runs
|
|
37
|
+
├── Eval::CaseResult # value object (name, passed?, output, expected, mismatches, cost)
|
|
38
|
+
├── Eval::ExpectationEvaluator # expected / expected_traits / evaluator proc
|
|
39
|
+
├── Eval::ModelComparison # compare_models result (table, best_for, candidate configs)
|
|
40
|
+
├── Eval::Recommender # model recommendation algorithm (candidates → optimal config)
|
|
41
|
+
├── Eval::Recommendation # recommendation result (best, retry_chain, savings, to_dsl)
|
|
42
|
+
├── Eval::RetryOptimizer # optimize_retry_policy result (per-eval breakdown, fallback list)
|
|
43
|
+
├── Eval::BaselineDiff # save_baseline! + without_regressions comparison
|
|
44
|
+
├── Eval::PromptDiffComparator # compare_with prompt A/B diff
|
|
45
|
+
└── Eval::EvalHistory # time-series view across saved reports
|
|
46
|
+
|
|
47
|
+
RubyLLM::Contract::RakeTask # rake ruby_llm_contract:eval
|
|
48
|
+
RubyLLM::Contract::OptimizeRakeTask # rake ruby_llm_contract:optimize
|
|
49
|
+
RubyLLM::Contract::Railtie # auto-loads eval files in Rails
|
|
50
|
+
```
|
|
@@ -0,0 +1,136 @@
|
|
|
1
|
+
# Best Practices
|
|
2
|
+
|
|
3
|
+
> Read this when writing your first `validate` rules and you want patterns that catch real bugs instead of restating schema.
|
|
4
|
+
|
|
5
|
+
Schema guarantees valid JSON structure. An LLM can still return structurally perfect JSON that is **semantically wrong**. Schema handles _shape_, validates handle _meaning_.
|
|
6
|
+
|
|
7
|
+
All examples extend the `SummarizeArticle` step from the [README](../../README.md).
|
|
8
|
+
|
|
9
|
+
## 1. Guard against empty / placeholder output
|
|
10
|
+
|
|
11
|
+
**Why it matters:** a cheap model that answers `{"tldr": "This article discusses...", "takeaways": ["This article discusses X", ...]}` passes the schema but renders a broken UI card that tells the user nothing. The validate catches it before `Article.update!` persists it.
|
|
12
|
+
|
|
13
|
+
```ruby
|
|
14
|
+
output_schema do
|
|
15
|
+
string :tldr
|
|
16
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
17
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
18
|
+
end
|
|
19
|
+
|
|
20
|
+
validate("tldr is substantive") do |o, _|
|
|
21
|
+
o[:tldr].to_s.split.length >= 5 # at least five words
|
|
22
|
+
end
|
|
23
|
+
|
|
24
|
+
validate("no boilerplate takeaways") do |o, _|
|
|
25
|
+
o[:takeaways].none? { |t| t.downcase.start_with?("this article") }
|
|
26
|
+
end
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
## 2. Cross-validate output against input
|
|
30
|
+
|
|
31
|
+
**Why it matters:** a lazy model will return the article text verbatim as the "summary", or invent takeaways about topics the article never mentions. The 2-arity form is how you catch answers that are internally consistent but unfaithful to the actual input.
|
|
32
|
+
|
|
33
|
+
`validate` blocks with 2-arity `|output, input|` compare the model's answer against what was asked:
|
|
34
|
+
|
|
35
|
+
```ruby
|
|
36
|
+
validate("tldr is shorter than the article") do |output, input|
|
|
37
|
+
output[:tldr].length < input.length / 2
|
|
38
|
+
end
|
|
39
|
+
|
|
40
|
+
validate("every takeaway appears, in spirit, in the article") do |output, input|
|
|
41
|
+
output[:takeaways].all? { |t|
|
|
42
|
+
# cheap keyword overlap heuristic
|
|
43
|
+
t.downcase.split.any? { |w| input.downcase.include?(w) && w.length > 4 }
|
|
44
|
+
}
|
|
45
|
+
end
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
## 3. Conditional logic schema can't express
|
|
49
|
+
|
|
50
|
+
**Why it matters:** customer success filters on `tone == "negative"` to route angry users to a human. If the model labels an outage complaint "negative" but the takeaways are all positive-sounding, the filter runs on a label that doesn't match the content — the routing breaks silently.
|
|
51
|
+
|
|
52
|
+
```ruby
|
|
53
|
+
validate("negative tone requires at least one concrete concern") do |output, _input|
|
|
54
|
+
next true unless output[:tone] == "negative"
|
|
55
|
+
output[:takeaways].any? { |t| t.match?(/fail|break|crash|outage|vulnerab|risk/i) }
|
|
56
|
+
end
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
A model that picks `tone: "negative"` but gives three upbeat takeaways fails this check. Schema can't catch it because each takeaway is, individually, a valid string.
|
|
60
|
+
|
|
61
|
+
## 4. Content quality
|
|
62
|
+
|
|
63
|
+
**Why it matters:** a TL;DR with `## Summary` leaks markdown into a plain-text card. A one-word takeaway ("Fast.") wastes a UI slot. A leaked `{article}` placeholder reveals the prompt template to end users. All pass schema; all embarrass you in front of customers.
|
|
64
|
+
|
|
65
|
+
```ruby
|
|
66
|
+
validate("no markdown headings in the TL;DR") do |o, _|
|
|
67
|
+
!o[:tldr].match?(/^\#{1,6}\s/)
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
validate("takeaways aren't single words") do |o, _|
|
|
71
|
+
o[:takeaways].all? { |t| t.split.length >= 3 }
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
validate("no template placeholders leaked") do |o, _|
|
|
75
|
+
!(o[:tldr] + o[:takeaways].join(" ")).include?("{")
|
|
76
|
+
end
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
## 5. Pipeline: preserve data between steps
|
|
80
|
+
|
|
81
|
+
In a pipeline, each step only sees the previous step's output. If a later step needs original article metadata, an intermediate step must carry it through. Suppose a pipeline `SummarizeArticle → GenerateHashtags`, where `GenerateHashtags` needs the `tone` from the summary:
|
|
82
|
+
|
|
83
|
+
```ruby
|
|
84
|
+
class GenerateHashtags < RubyLLM::Contract::Step::Base
|
|
85
|
+
input_type Hash
|
|
86
|
+
|
|
87
|
+
output_schema do
|
|
88
|
+
# Carry through the fields a downstream step (or the caller) might need
|
|
89
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
90
|
+
array :hashtags, of: :string, min_items: 2, max_items: 5
|
|
91
|
+
end
|
|
92
|
+
|
|
93
|
+
prompt do
|
|
94
|
+
rule "Preserve the tone label from the input unchanged."
|
|
95
|
+
user "Summary: {tldr}\nTone: {tone}\nTakeaways: {takeaways}"
|
|
96
|
+
end
|
|
97
|
+
|
|
98
|
+
validate("tone preserved") { |o, input| o[:tone] == input[:tone] }
|
|
99
|
+
end
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
The explicit `validate("tone preserved")` catches the case where the model silently rewrites the tone during a downstream transform.
|
|
103
|
+
|
|
104
|
+
## 6. Model fallback
|
|
105
|
+
|
|
106
|
+
**Why it matters:** 80% of production articles are short and simple — `gpt-4.1-nano` handles them for ~$0.0001. The remaining 20% are dense, critical, or multi-topic — those need `gpt-4.1-mini` or `gpt-4.1`. Paying `gpt-4.1` rates for every call when nano is enough for most is throwing money away. Contracts tell you when nano wasn't enough, so fallback is cost-aware, not hope-based.
|
|
107
|
+
|
|
108
|
+
Small models are cheap but hallucinate. Big models are accurate but expensive. Start cheap, fall back only when validates catch a failure:
|
|
109
|
+
|
|
110
|
+
```ruby
|
|
111
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
112
|
+
output_schema do
|
|
113
|
+
string :tldr
|
|
114
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
115
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
116
|
+
end
|
|
117
|
+
|
|
118
|
+
validate("TL;DR fits the card") { |o, _| o[:tldr].length <= 200 }
|
|
119
|
+
validate("takeaways are unique") { |o, _| o[:takeaways] == o[:takeaways].uniq }
|
|
120
|
+
|
|
121
|
+
retry_policy models: %w[gpt-4.1-nano gpt-4.1-mini gpt-4.1]
|
|
122
|
+
end
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
**Key insight:** without contracts, you can't do model fallback — you'd have no way to know if the cheap model's output is good enough. Validates are the quality gate that makes cost optimization possible. See [Optimizing retry_policy](optimizing_retry_policy.md) for how to find the cheapest viable fallback list for your step.
|
|
126
|
+
|
|
127
|
+
## Summary
|
|
128
|
+
|
|
129
|
+
| What to validate | Use |
|
|
130
|
+
|---|---|
|
|
131
|
+
| Field types, enums, ranges, required fields | `output_schema` |
|
|
132
|
+
| Output makes sense given the input | `validate` (2-arity `\|output, input\|`) |
|
|
133
|
+
| Conditional business rules | `validate` |
|
|
134
|
+
| Content quality (not empty, not template) | `validate` |
|
|
135
|
+
| Data preserved across pipeline steps | `validate` + schema carry-through |
|
|
136
|
+
| Cost optimization via model fallback | `retry_policy` + `validate` as quality gate |
|