ruby_llm-contract 0.8.0 → 0.10.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +79 -1
- data/Gemfile.lock +2 -2
- data/README.md +96 -37
- data/docs/architecture.md +50 -0
- data/docs/guide/best_practices.md +136 -0
- data/docs/guide/eval_first.md +192 -0
- data/docs/guide/getting_started.md +199 -0
- data/docs/guide/migration.md +185 -0
- data/docs/guide/multimodal_input.md +160 -0
- data/docs/guide/optimizing_retry_policy.md +131 -0
- data/docs/guide/output_schema.md +93 -0
- data/docs/guide/pipeline.md +154 -0
- data/docs/guide/prompt_ast.md +76 -0
- data/docs/guide/rails_integration.md +218 -0
- data/docs/guide/relation_to_agent.md +52 -0
- data/docs/guide/relation_to_tribunal.md +135 -0
- data/docs/guide/testing.md +282 -0
- data/docs/guide/why.md +103 -0
- data/lib/ruby_llm/contract/adapters/ruby_llm.rb +9 -1
- data/lib/ruby_llm/contract/concerns/eval_host.rb +6 -9
- data/lib/ruby_llm/contract/concerns/stub_helpers.rb +97 -0
- data/lib/ruby_llm/contract/contract/definition.rb +2 -0
- data/lib/ruby_llm/contract/cost_calculator.rb +11 -2
- data/lib/ruby_llm/contract/eval/recommender.rb +3 -1
- data/lib/ruby_llm/contract/eval/retry_optimizer.rb +16 -13
- data/lib/ruby_llm/contract/eval.rb +13 -0
- data/lib/ruby_llm/contract/minitest.rb +6 -108
- data/lib/ruby_llm/contract/pipeline/result.rb +1 -1
- data/lib/ruby_llm/contract/rake_task/suite_gate.rb +117 -0
- data/lib/ruby_llm/contract/rake_task.rb +30 -51
- data/lib/ruby_llm/contract/rspec/helpers.rb +9 -123
- data/lib/ruby_llm/contract/step/base.rb +56 -24
- data/lib/ruby_llm/contract/step/dsl.rb +91 -63
- data/lib/ruby_llm/contract/step/limit_checker.rb +34 -1
- data/lib/ruby_llm/contract/step/retry_executor.rb +6 -13
- data/lib/ruby_llm/contract/step/runner.rb +22 -20
- data/lib/ruby_llm/contract/step/runner_config.rb +26 -0
- data/lib/ruby_llm/contract/version.rb +1 -1
- data/lib/ruby_llm/contract.rb +1 -0
- data/ruby_llm-contract.gemspec +5 -1
- metadata +18 -4
- data/.rspec +0 -3
- data/.rubycritic.yml +0 -8
- data/.simplecov +0 -22
|
@@ -0,0 +1,192 @@
|
|
|
1
|
+
# Eval-First
|
|
2
|
+
|
|
3
|
+
> Read this when you need to prevent silent prompt regressions in CI. Skip if your LLM output is evaluated only by humans and never gated in an automated pipeline.
|
|
4
|
+
|
|
5
|
+
If you change prompts by feel, you ship regressions by feel.
|
|
6
|
+
|
|
7
|
+
Concrete scenario: `SummarizeArticle` has been running in production for two weeks. Customer success notices that complaints about service outages keep getting `tone: "analytical"` instead of `"negative"` — so their "critical feedback" filter silently misses angry users. Someone tweaks the system prompt to emphasise negative sentiment. It fixes the outage article but now three neutral product-update articles get misclassified as `"negative"`. You find out from a Slack thread.
|
|
8
|
+
|
|
9
|
+
That is the cost of prompt-by-feel. Evals are how you stop it.
|
|
10
|
+
|
|
11
|
+
`ruby_llm-contract` works best when you treat evals as the source of truth:
|
|
12
|
+
|
|
13
|
+
1. Capture real failures from production (the outage article, verbatim).
|
|
14
|
+
2. Turn them into eval cases (`add_case "service outage complaint"`).
|
|
15
|
+
3. Change the prompt.
|
|
16
|
+
4. Re-run the same eval — plus all previously-passing cases.
|
|
17
|
+
5. Merge only if the eval says quality improved or stayed safe on every case.
|
|
18
|
+
|
|
19
|
+
## Core rule
|
|
20
|
+
|
|
21
|
+
**Do not start with the prompt. Start with the eval.**
|
|
22
|
+
|
|
23
|
+
Using the `SummarizeArticle` step from the [README](../../README.md):
|
|
24
|
+
|
|
25
|
+
```ruby
|
|
26
|
+
SummarizeArticle.define_eval("regression") do
|
|
27
|
+
add_case "ruby release",
|
|
28
|
+
input: "Ruby 3.4 shipped with frozen string literals...",
|
|
29
|
+
expected: { tone: "analytical" } # partial match
|
|
30
|
+
|
|
31
|
+
add_case "critical review",
|
|
32
|
+
input: "Mesh networking hardware failed under load...",
|
|
33
|
+
expected: { tone: "negative" }
|
|
34
|
+
end
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Only after the eval exists, touch: `system`, `rule`, `example`, `validate`, prompt versions.
|
|
38
|
+
|
|
39
|
+
## Three eval kinds
|
|
40
|
+
|
|
41
|
+
### 1. `smoke` — wiring check, offline
|
|
42
|
+
|
|
43
|
+
```ruby
|
|
44
|
+
SummarizeArticle.define_eval("smoke") do
|
|
45
|
+
default_input "Ruby 3.4 shipped with frozen string literals..."
|
|
46
|
+
sample_response({
|
|
47
|
+
tldr: "...",
|
|
48
|
+
takeaways: ["point one", "point two", "point three"],
|
|
49
|
+
tone: "analytical"
|
|
50
|
+
})
|
|
51
|
+
end
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
`sample_response` returns canned data. Zero API calls. Verifies schema + validates parse and the step wiring is intact. **Not a quality signal.**
|
|
55
|
+
|
|
56
|
+
### 2. `regression` — real quality measurement
|
|
57
|
+
|
|
58
|
+
Represent real traffic and known failures. Good sources: production logs, bad completions, incidents, QA edge cases, cases a human had to correct.
|
|
59
|
+
|
|
60
|
+
Every production failure becomes `add_case`. That's the flywheel.
|
|
61
|
+
|
|
62
|
+
### 3. `ab` — prompt iteration
|
|
63
|
+
|
|
64
|
+
Compare two prompt versions on the same eval:
|
|
65
|
+
|
|
66
|
+
```ruby
|
|
67
|
+
diff = SummarizeArticleV2.compare_with(
|
|
68
|
+
SummarizeArticleV1,
|
|
69
|
+
eval: "regression",
|
|
70
|
+
model: "gpt-4.1-mini"
|
|
71
|
+
)
|
|
72
|
+
|
|
73
|
+
diff.safe_to_switch? # => true if no cases regressed
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
This is the cleanest eval-first move: same eval, same cases, two prompt versions, one answer.
|
|
77
|
+
|
|
78
|
+
## What counts as eval-first
|
|
79
|
+
|
|
80
|
+
**Good** — eval exists before the prompt change:
|
|
81
|
+
|
|
82
|
+
```ruby
|
|
83
|
+
SummarizeArticle.define_eval("regression") do
|
|
84
|
+
add_case "short article", input: "...", expected: { tone: "neutral" }
|
|
85
|
+
end
|
|
86
|
+
|
|
87
|
+
# Prompt iteration happens afterward
|
|
88
|
+
diff = SummarizeArticleV2.compare_with(
|
|
89
|
+
SummarizeArticleV1, eval: "regression", model: "gpt-4.1-mini"
|
|
90
|
+
)
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
**Bad**:
|
|
94
|
+
|
|
95
|
+
```ruby
|
|
96
|
+
# Tweak prompt for an hour
|
|
97
|
+
# Maybe add an example
|
|
98
|
+
# Maybe tighten a rule
|
|
99
|
+
# Then eyeball one or two responses
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
That's prompt guessing, not eval-first.
|
|
103
|
+
|
|
104
|
+
## `sample_response`: useful, but not the main thing
|
|
105
|
+
|
|
106
|
+
Good for: offline smoke tests, local development, testing evaluator wiring, checking schema + validate behavior with zero API calls.
|
|
107
|
+
|
|
108
|
+
Not enough for real prompt decisions. For those:
|
|
109
|
+
|
|
110
|
+
- `run_eval(..., context: { model: "..." })` with a real model, or pass an explicit adapter.
|
|
111
|
+
- `compare_with(...)` for prompt A/B.
|
|
112
|
+
|
|
113
|
+
`compare_with` intentionally ignores `sample_response` — canned data would make both sides look the same.
|
|
114
|
+
|
|
115
|
+
## Parallel eval runs
|
|
116
|
+
|
|
117
|
+
For larger datasets, `run_eval` accepts a `concurrency:` argument — cases run in parallel using a thread pool:
|
|
118
|
+
|
|
119
|
+
```ruby
|
|
120
|
+
report = SummarizeArticle.run_eval("regression",
|
|
121
|
+
context: { model: "gpt-4.1-mini" },
|
|
122
|
+
concurrency: 8)
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Same accepted by `compare_models` and `optimize_retry_policy`. Thread count is a ceiling — dataset order of results is preserved. Keep it low enough to respect the provider's rate limits.
|
|
126
|
+
|
|
127
|
+
## Budgeting an eval before you run it
|
|
128
|
+
|
|
129
|
+
`estimate_eval_cost` gives you a cost projection without calling the LLM:
|
|
130
|
+
|
|
131
|
+
```ruby
|
|
132
|
+
SummarizeArticle.estimate_eval_cost("regression",
|
|
133
|
+
models: %w[gpt-4.1-nano gpt-4.1-mini gpt-4.1])
|
|
134
|
+
# => { "gpt-4.1-nano" => 0.00041, "gpt-4.1-mini" => 0.0018, "gpt-4.1" => 0.0092 }
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
Use it in CI to decide which models are worth running regression on, or to cap worst-case spend per build.
|
|
138
|
+
|
|
139
|
+
## Team workflow
|
|
140
|
+
|
|
141
|
+
1. **Build one eval that matters** — 10–30 cases representing real mistakes and important business paths.
|
|
142
|
+
2. **Gate CI** — `pass_eval("regression").with_context(model: "...").with_minimum_score(0.8)`. See [Getting Started](getting_started.md) for the full matcher chain.
|
|
143
|
+
3. **Save a baseline** — `report.save_baseline!` makes quality drift visible.
|
|
144
|
+
4. **Change prompts only through comparison** — `pass_eval(...).compared_with(SummarizeArticleV1)` in CI so any regression blocks the merge.
|
|
145
|
+
5. **Feed production failures back** — every miss in prod → new `add_case`, then fix. The eval gets stronger over time.
|
|
146
|
+
|
|
147
|
+
## Few-shot examples fit naturally
|
|
148
|
+
|
|
149
|
+
Adding `example input: ..., output: ...` inside the prompt is still a prompt change. The eval-first way:
|
|
150
|
+
|
|
151
|
+
1. Add examples to the prompt.
|
|
152
|
+
2. Rerun the existing regression eval.
|
|
153
|
+
3. `compare_with` against the old prompt.
|
|
154
|
+
|
|
155
|
+
Few-shot isn't the proof. The eval is.
|
|
156
|
+
|
|
157
|
+
## Model selection comes after prompt stability
|
|
158
|
+
|
|
159
|
+
Don't optimize cost before you stabilize quality:
|
|
160
|
+
|
|
161
|
+
1. Build `regression`.
|
|
162
|
+
2. Improve the prompt with `compare_with`.
|
|
163
|
+
3. Lock quality in CI.
|
|
164
|
+
4. Then run `compare_models` (see [Optimizing retry_policy](optimizing_retry_policy.md)).
|
|
165
|
+
|
|
166
|
+
```ruby
|
|
167
|
+
comparison = SummarizeArticle.compare_models(
|
|
168
|
+
"regression",
|
|
169
|
+
candidates: [{ model: "gpt-4.1-nano" }, { model: "gpt-4.1-mini" }, { model: "gpt-4.1" }]
|
|
170
|
+
)
|
|
171
|
+
|
|
172
|
+
comparison.best_for(min_score: 0.95)
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
## Strong defaults for teams
|
|
176
|
+
|
|
177
|
+
- `smoke` uses `sample_response`.
|
|
178
|
+
- `regression` uses real model calls.
|
|
179
|
+
- Every prompt change uses `compare_with`.
|
|
180
|
+
- Every merge runs `pass_eval`.
|
|
181
|
+
- Every production failure becomes a new `add_case`.
|
|
182
|
+
|
|
183
|
+
## Short version
|
|
184
|
+
|
|
185
|
+
1. Write `define_eval` before touching the prompt.
|
|
186
|
+
2. Treat `sample_response` as smoke only.
|
|
187
|
+
3. Use `run_eval("name", context: { model: "..." })` for real quality measurement.
|
|
188
|
+
4. Use `compare_with` for every serious prompt change.
|
|
189
|
+
5. Gate merges with `pass_eval`.
|
|
190
|
+
6. Feed every production miss back into the dataset.
|
|
191
|
+
|
|
192
|
+
Prompts stop being vibes and start being engineering.
|
|
@@ -0,0 +1,199 @@
|
|
|
1
|
+
# Getting Started
|
|
2
|
+
|
|
3
|
+
> Read this to walk through every feature on one concrete step for the first time.
|
|
4
|
+
|
|
5
|
+
The README shows a minimal `SummarizeArticle` step. This guide walks through the features you reach for as production requirements grow: budget caps so runaway inputs don't drain your LLM provider budget, evals so you catch regressions in CI, and CI gating so a merge that lowers accuracy gets blocked.
|
|
6
|
+
|
|
7
|
+
## The walkthrough
|
|
8
|
+
|
|
9
|
+
Start with the README example, then add features one layer at a time. Each is optional — use what you need.
|
|
10
|
+
|
|
11
|
+
```ruby
|
|
12
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
13
|
+
# 1. Prompt (required)
|
|
14
|
+
prompt <<~PROMPT
|
|
15
|
+
Summarize this article for a UI card. Return a short TL;DR,
|
|
16
|
+
3 to 5 key takeaways, and a tone label.
|
|
17
|
+
|
|
18
|
+
{input}
|
|
19
|
+
PROMPT
|
|
20
|
+
|
|
21
|
+
# 2. Schema — sent to the provider via with_schema, validated client-side
|
|
22
|
+
output_schema do
|
|
23
|
+
string :tldr
|
|
24
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
25
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
26
|
+
end
|
|
27
|
+
|
|
28
|
+
# 3. Business rules — things JSON Schema cannot express
|
|
29
|
+
validate("TL;DR fits the card") { |o, _| o[:tldr].length <= 200 }
|
|
30
|
+
validate("takeaways are unique") { |o, _| o[:takeaways] == o[:takeaways].uniq }
|
|
31
|
+
|
|
32
|
+
# 4. Retry with model fallback on validation_failed / parse_error
|
|
33
|
+
retry_policy models: %w[gpt-4.1-nano gpt-4.1-mini gpt-4.1]
|
|
34
|
+
|
|
35
|
+
# 5. Refuse before calling the LLM if input is too large or estimated cost exceeds the cap
|
|
36
|
+
max_input 2_000
|
|
37
|
+
max_output 4_000
|
|
38
|
+
max_cost 0.01
|
|
39
|
+
end
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
## Validation and retry behavior
|
|
43
|
+
|
|
44
|
+
When the cheap model returns output that fails a `validate` block or can't be parsed, retry falls back to the next model in `models:` and tries again.
|
|
45
|
+
|
|
46
|
+
```ruby
|
|
47
|
+
result = SummarizeArticle.run(article_text)
|
|
48
|
+
|
|
49
|
+
result.status # => :ok
|
|
50
|
+
result.parsed_output # => { tldr: "...", takeaways: [...], tone: "analytical" }
|
|
51
|
+
result.trace[:model] # => "gpt-4.1-mini" (first model that passed)
|
|
52
|
+
result.trace[:cost] # => 0.00052 (sum of all attempts)
|
|
53
|
+
result.trace[:attempts]
|
|
54
|
+
# => [
|
|
55
|
+
# { attempt: 1, model: "gpt-4.1-nano", status: :validation_failed,
|
|
56
|
+
# cost: 0.00010, latency_ms: 45, ... },
|
|
57
|
+
# { attempt: 2, model: "gpt-4.1-mini", status: :ok,
|
|
58
|
+
# cost: 0.00042, latency_ms: 92, ... }
|
|
59
|
+
# ]
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
If the whole chain exhausts, `result.status` is the status of the last attempt (`:validation_failed` or `:parse_error`) and `result.parsed_output` is the last attempt's output. The caller decides what to do — ship it anyway, fall back to a template, or raise.
|
|
63
|
+
|
|
64
|
+
### Per-attempt reasoning effort
|
|
65
|
+
|
|
66
|
+
`models:` accepts config hashes as well as model-name strings, so a fallback can "try harder" (more reasoning) on retry, not just switch model:
|
|
67
|
+
|
|
68
|
+
```ruby
|
|
69
|
+
retry_policy models: [
|
|
70
|
+
"gpt-5-nano", # attempt 1: cheap + fast
|
|
71
|
+
{ model: "gpt-5-mini", reasoning_effort: "high" } # attempt 2: stronger + more reasoning
|
|
72
|
+
]
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
`reasoning_effort` is a gpt-5 family feature (gpt-4.1 is not a reasoning model). The per-attempt value is forwarded via `with_thinking` (provider-agnostic — OpenAI `reasoning_effort` and Anthropic extended-thinking budget both supported). Passing it alongside a non-reasoning model is forwarded unchanged to the provider, which will either ignore it or reject the request — the gem does not guard against this.
|
|
76
|
+
|
|
77
|
+
## Evals and CI gates
|
|
78
|
+
|
|
79
|
+
An eval is a named scenario you can run to verify the step still works. `sample_response` makes it offline — zero API calls — so CI can run it on every merge without burning budget.
|
|
80
|
+
|
|
81
|
+
```ruby
|
|
82
|
+
SummarizeArticle.define_eval("smoke") do
|
|
83
|
+
default_input <<~ARTICLE
|
|
84
|
+
Ruby 3.4 ships with frozen string literals on by default, measurable YJIT
|
|
85
|
+
speedups on Rails workloads, and tightened Warning.warn category filtering.
|
|
86
|
+
The release notes also mention several parser fixes and faster keyword
|
|
87
|
+
argument handling.
|
|
88
|
+
ARTICLE
|
|
89
|
+
|
|
90
|
+
sample_response({
|
|
91
|
+
tldr: "Ruby 3.4 brings frozen string literals by default, YJIT speedups, and parser fixes.",
|
|
92
|
+
takeaways: [
|
|
93
|
+
"Frozen string literals are the default",
|
|
94
|
+
"YJIT adds measurable speedups on Rails workloads",
|
|
95
|
+
"Warning.warn category filtering is tighter"
|
|
96
|
+
],
|
|
97
|
+
tone: "analytical"
|
|
98
|
+
})
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
report = SummarizeArticle.run_eval("smoke")
|
|
102
|
+
report.passed? # => true — schema + validates pass on the canned response
|
|
103
|
+
report.score # => 1.0
|
|
104
|
+
report.print_summary
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
For real regression testing, define cases with expected output (online — calls the LLM):
|
|
108
|
+
|
|
109
|
+
```ruby
|
|
110
|
+
SummarizeArticle.define_eval("regression") do
|
|
111
|
+
add_case "ruby release",
|
|
112
|
+
input: "Ruby 3.4 was released...",
|
|
113
|
+
expected: { tone: "analytical" } # partial match
|
|
114
|
+
|
|
115
|
+
add_case "critical review",
|
|
116
|
+
input: "The new mesh networking hardware failed under load...",
|
|
117
|
+
expected: { tone: "negative" }
|
|
118
|
+
end
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Gate CI on score and cost thresholds:
|
|
122
|
+
|
|
123
|
+
```ruby
|
|
124
|
+
# RSpec — blocks merge if accuracy drops or cost spikes
|
|
125
|
+
expect(SummarizeArticle).to pass_eval("regression")
|
|
126
|
+
.with_minimum_score(0.8)
|
|
127
|
+
.with_maximum_cost(0.01)
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
Save a baseline once, then block regressions automatically:
|
|
131
|
+
|
|
132
|
+
```ruby
|
|
133
|
+
report = SummarizeArticle.run_eval("regression")
|
|
134
|
+
report.save_baseline!
|
|
135
|
+
|
|
136
|
+
# In CI:
|
|
137
|
+
expect(SummarizeArticle).to pass_eval("regression").without_regressions
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
`without_regressions` fails the build only if a previously-passing case now fails — a new model version, a prompt tweak, or an upstream change that silently lowered quality.
|
|
141
|
+
|
|
142
|
+
## Budget caps
|
|
143
|
+
|
|
144
|
+
`max_input`, `max_output`, and `max_cost` are preflight checks — the LLM is never called if an estimate exceeds the limit. Zero tokens spent, zero cost.
|
|
145
|
+
|
|
146
|
+
```ruby
|
|
147
|
+
result = SummarizeArticle.run(giant_10mb_document)
|
|
148
|
+
result.status # => :limit_exceeded
|
|
149
|
+
result.validation_errors
|
|
150
|
+
# => ["Input token limit exceeded: estimated 32000 tokens (heuristic ±30%), max 2000"]
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
`max_cost` fails closed when the model's pricing isn't known — register custom or fine-tuned models explicitly:
|
|
154
|
+
|
|
155
|
+
```ruby
|
|
156
|
+
RubyLLM::Contract::CostCalculator.register_model("ft:gpt-4o-custom",
|
|
157
|
+
input_per_1m: 3.0, output_per_1m: 6.0)
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
Or opt into a soft warning instead of a refusal when pricing is missing:
|
|
161
|
+
|
|
162
|
+
```ruby
|
|
163
|
+
max_cost 0.01, on_unknown_pricing: :warn
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
Default is `:refuse`. Use `:warn` only when you accept running without a cost ceiling (fine-tuned models you trust, private endpoints).
|
|
167
|
+
|
|
168
|
+
### Preflight cost estimates
|
|
169
|
+
|
|
170
|
+
Check what a call is likely to cost before invoking it:
|
|
171
|
+
|
|
172
|
+
```ruby
|
|
173
|
+
SummarizeArticle.estimate_cost(input: article_text)
|
|
174
|
+
# => {
|
|
175
|
+
# model: "gpt-4.1-mini",
|
|
176
|
+
# input_tokens: 812, output_tokens_estimate: 4000,
|
|
177
|
+
# estimated_cost: 0.00243
|
|
178
|
+
# }
|
|
179
|
+
|
|
180
|
+
# Estimate what a full eval would cost across candidate models
|
|
181
|
+
SummarizeArticle.estimate_eval_cost("regression",
|
|
182
|
+
models: %w[gpt-4.1-nano gpt-4.1-mini gpt-4.1])
|
|
183
|
+
# => { "gpt-4.1-nano" => 0.00041, "gpt-4.1-mini" => 0.0018, "gpt-4.1" => 0.0092 }
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
`estimate_cost` returns `nil` when pricing isn't registered. `estimate_eval_cost` silently treats unknown-pricing cases as `$0.00` and sums the rest — it does **not** fail closed the way `max_cost` does. Treat its output as a floor, not a guarantee; register pricing via `CostCalculator.register_model` before relying on it for budget decisions.
|
|
187
|
+
|
|
188
|
+
## `output_schema` vs `with_schema`
|
|
189
|
+
|
|
190
|
+
`with_schema` in `ruby_llm` tells the provider to force a specific JSON structure. `output_schema` in this gem does the same thing (calls `with_schema` under the hood) **plus** validates the response client-side. Cheaper models sometimes ignore schema constraints — `with_schema` is a request; `output_schema` is a request plus verification.
|
|
191
|
+
|
|
192
|
+
## See also
|
|
193
|
+
|
|
194
|
+
- [Prompt AST](prompt_ast.md) — prompt DSL variants: `system`, `rule`, `section`, `example`, `user`, and dynamic prompts with `|input|`.
|
|
195
|
+
- [Eval-First](eval_first.md) — datasets, baselines, A/B gates, the workflow that makes the above evals useful.
|
|
196
|
+
- [Optimizing retry_policy](optimizing_retry_policy.md) — find the cheapest viable fallback list with `compare_models` and `optimize_retry_policy`.
|
|
197
|
+
- [Testing](testing.md) — test adapter, `stub_step`, full RSpec + Minitest matcher reference.
|
|
198
|
+
- [Output Schema](output_schema.md) — nested objects in arrays, constraints, pattern reference.
|
|
199
|
+
- [Rails integration](rails_integration.md) — where step classes live, initializer, jobs, logging, specs, CI gate.
|
|
@@ -0,0 +1,185 @@
|
|
|
1
|
+
# Migration Guide
|
|
2
|
+
|
|
3
|
+
> Read this when adopting the gem in an existing Rails app with a raw `LlmClient.call` service you want to replace. Skip if starting fresh — go to [Getting Started](getting_started.md) instead.
|
|
4
|
+
|
|
5
|
+
How to adopt `ruby_llm-contract` in an existing Rails app. Examples use `SummarizeArticle` — the flagship step from the [README](../../README.md) — but the pattern applies to any single-input / structured-output service.
|
|
6
|
+
|
|
7
|
+
## Step 1: Start with the simplest service
|
|
8
|
+
|
|
9
|
+
Pick the LLM service with: single input → JSON output → DB save. Don't start with parallel batches or complex pipelines.
|
|
10
|
+
|
|
11
|
+
## Step 2: Define the contract
|
|
12
|
+
|
|
13
|
+
**Before — raw HTTP:**
|
|
14
|
+
|
|
15
|
+
```ruby
|
|
16
|
+
class ArticleSummaryService
|
|
17
|
+
def call(article_text)
|
|
18
|
+
response = LlmClient.new(model: "gpt-4o-mini").call(prompt(article_text))
|
|
19
|
+
JSON.parse(response[:content], symbolize_names: true)
|
|
20
|
+
end
|
|
21
|
+
end
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
**After — contract:**
|
|
25
|
+
|
|
26
|
+
```ruby
|
|
27
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
28
|
+
model "gpt-4.1-mini"
|
|
29
|
+
|
|
30
|
+
prompt do
|
|
31
|
+
system "You summarize articles for a UI card."
|
|
32
|
+
rule "Return valid JSON only."
|
|
33
|
+
user "{input}"
|
|
34
|
+
end
|
|
35
|
+
|
|
36
|
+
output_schema do
|
|
37
|
+
string :tldr
|
|
38
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
39
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
validate("TL;DR fits the card") { |o, _| o[:tldr].length <= 200 }
|
|
43
|
+
retry_policy models: %w[gpt-4.1-mini gpt-4.1]
|
|
44
|
+
end
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Step 3: Replace the caller
|
|
48
|
+
|
|
49
|
+
```ruby
|
|
50
|
+
# Before
|
|
51
|
+
parsed = ArticleSummaryService.new.call(article_text)
|
|
52
|
+
Article.update!(summary: parsed["tldr"])
|
|
53
|
+
|
|
54
|
+
# After
|
|
55
|
+
result = SummarizeArticle.run(article_text)
|
|
56
|
+
if result.ok?
|
|
57
|
+
Article.update!(summary: result.parsed_output[:tldr])
|
|
58
|
+
else
|
|
59
|
+
Rails.logger.warn "Summary failed: #{result.status}"
|
|
60
|
+
end
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Step 4: Add logging via around_call
|
|
64
|
+
|
|
65
|
+
```ruby
|
|
66
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
67
|
+
# ... prompt, schema, validates ...
|
|
68
|
+
|
|
69
|
+
around_call do |step, input, result|
|
|
70
|
+
AiCallLog.create!(
|
|
71
|
+
ai_model: result.trace.model,
|
|
72
|
+
duration_ms: result.trace.latency_ms,
|
|
73
|
+
input_tokens: result.trace.usage&.dig(:input_tokens),
|
|
74
|
+
output_tokens: result.trace.usage&.dig(:output_tokens),
|
|
75
|
+
cost: result.trace.cost,
|
|
76
|
+
status: result.status.to_s
|
|
77
|
+
)
|
|
78
|
+
end
|
|
79
|
+
end
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
## Step 5: Add eval cases
|
|
83
|
+
|
|
84
|
+
Use real inputs from production logs:
|
|
85
|
+
|
|
86
|
+
```ruby
|
|
87
|
+
SummarizeArticle.define_eval("regression") do
|
|
88
|
+
add_case "short news",
|
|
89
|
+
input: "Ruby 3.4 ships with frozen string literals by default...",
|
|
90
|
+
expected: { tone: "analytical" }
|
|
91
|
+
|
|
92
|
+
add_case "critical review",
|
|
93
|
+
input: "The new mesh networking hardware failed under load...",
|
|
94
|
+
expected: { tone: "negative" }
|
|
95
|
+
end
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
## Step 6: Find the cheapest model
|
|
99
|
+
|
|
100
|
+
```ruby
|
|
101
|
+
comparison = SummarizeArticle.compare_models("regression",
|
|
102
|
+
candidates: [{ model: "gpt-4.1-nano" }, { model: "gpt-4.1-mini" }])
|
|
103
|
+
|
|
104
|
+
comparison.print_summary
|
|
105
|
+
comparison.best_for(min_score: 0.95) # => cheapest model at >= 95%
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
Full optimization workflow — multi-eval, fallback list, production-mode cost — in [Optimizing retry_policy](optimizing_retry_policy.md).
|
|
109
|
+
|
|
110
|
+
## Step 7: Add CI gate
|
|
111
|
+
|
|
112
|
+
```ruby
|
|
113
|
+
# Rakefile
|
|
114
|
+
require "ruby_llm/contract/rake_task"
|
|
115
|
+
RubyLLM::Contract::RakeTask.new do |t|
|
|
116
|
+
t.minimum_score = 0.8
|
|
117
|
+
t.maximum_cost = 0.05
|
|
118
|
+
t.fail_on_regression = true
|
|
119
|
+
t.save_baseline = true
|
|
120
|
+
end
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
**Rails apps:** if your adapter is configured in an initializer, use a Proc so context is resolved after Rails boots:
|
|
124
|
+
|
|
125
|
+
```ruby
|
|
126
|
+
RubyLLM::Contract::RakeTask.new do |t|
|
|
127
|
+
t.context = -> { { adapter: RubyLLM::Contract.configuration.default_adapter } }
|
|
128
|
+
t.minimum_score = 0.8
|
|
129
|
+
end
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
## Common patterns
|
|
133
|
+
|
|
134
|
+
| Old pattern | New pattern |
|
|
135
|
+
|---|---|
|
|
136
|
+
| `LlmClient.new(model:).call(prompt)` | `MyStep.run(input)` |
|
|
137
|
+
| `JSON.parse(response[:content])` | `result.parsed_output` |
|
|
138
|
+
| `begin; rescue; retry; end` | `retry_policy models: [...]` |
|
|
139
|
+
| `body[:temperature] = 0.7` | `temperature 0.7` |
|
|
140
|
+
| `AiCallLog.create(...)` | `around_call { \|s, i, r\| ... }` |
|
|
141
|
+
| `response_format: JsonSchema.build(...)` | `output_schema do...end` |
|
|
142
|
+
| `stub_request(:post, ...)` | `stub_step(MyStep, response: {...})` |
|
|
143
|
+
|
|
144
|
+
## Anti-patterns
|
|
145
|
+
|
|
146
|
+
- **Don't migrate markdown/text output services.** The gem is for structured JSON. Prose output gets no benefit from schema validation.
|
|
147
|
+
- **Don't put parallelism in the gem.** Thread management is your app's concern. The gem provides the contract; you call it however you want.
|
|
148
|
+
- **Don't migrate all services at once.** Start with one. Validate the pattern. Then migrate the next.
|
|
149
|
+
|
|
150
|
+
## Parallel batch generation
|
|
151
|
+
|
|
152
|
+
The gem handles single calls. You handle parallelism:
|
|
153
|
+
|
|
154
|
+
```ruby
|
|
155
|
+
class SummarizeBatch < RubyLLM::Contract::Step::Base
|
|
156
|
+
output_schema do
|
|
157
|
+
array :summaries do
|
|
158
|
+
object do
|
|
159
|
+
string :article_id
|
|
160
|
+
string :tldr
|
|
161
|
+
end
|
|
162
|
+
end
|
|
163
|
+
end
|
|
164
|
+
retry_policy models: %w[gpt-4.1-mini gpt-4.1]
|
|
165
|
+
end
|
|
166
|
+
|
|
167
|
+
# Your orchestrator
|
|
168
|
+
threads = 10.times.map do |i|
|
|
169
|
+
Thread.new { Rails.application.executor.wrap { SummarizeBatch.run(input(i)) } }
|
|
170
|
+
end
|
|
171
|
+
results = threads.map(&:value)
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
**Note:** in tests, `stub_step` overrides are thread-local. If your orchestrator spawns threads, propagate overrides manually:
|
|
175
|
+
|
|
176
|
+
```ruby
|
|
177
|
+
overrides = RubyLLM::Contract.step_adapter_overrides.dup
|
|
178
|
+
Thread.new { RubyLLM::Contract.step_adapter_overrides = overrides; SummarizeBatch.run(input) }
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
## See also
|
|
182
|
+
|
|
183
|
+
- [Getting Started](getting_started.md) — the full walkthrough of every feature `SummarizeArticle` uses.
|
|
184
|
+
- [Testing](testing.md) — `stub_step` reference for migrating your test adapter mocks.
|
|
185
|
+
- [Eval-First](eval_first.md) — how to build the "regression" eval from production logs.
|