ruby_llm-contract 0.10.1 → 0.10.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +30 -1
- data/Gemfile.lock +2 -2
- data/README.md +2 -2
- data/docs/architecture.md +50 -0
- data/docs/guide/best_practices.md +136 -0
- data/docs/guide/eval_first.md +192 -0
- data/docs/guide/getting_started.md +199 -0
- data/docs/guide/migration.md +185 -0
- data/docs/guide/multimodal_input.md +160 -0
- data/docs/guide/optimizing_retry_policy.md +131 -0
- data/docs/guide/output_schema.md +93 -0
- data/docs/guide/pipeline.md +154 -0
- data/docs/guide/prompt_ast.md +76 -0
- data/docs/guide/rails_integration.md +218 -0
- data/docs/guide/relation_to_agent.md +52 -0
- data/docs/guide/relation_to_tribunal.md +135 -0
- data/docs/guide/testing.md +282 -0
- data/docs/guide/why.md +103 -0
- data/lib/ruby_llm/contract/contract/schema_validator/object_rules.rb +24 -0
- data/lib/ruby_llm/contract/version.rb +1 -1
- data/ruby_llm-contract.gemspec +1 -1
- metadata +16 -1
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
# Output Schema
|
|
2
|
+
|
|
3
|
+
> Read this as a reference for the schema DSL — every constraint, nested objects, arrays of objects, the full pattern table.
|
|
4
|
+
|
|
5
|
+
Declare the expected output structure using [ruby_llm-schema](https://github.com/danielfriis/ruby_llm-schema) DSL. The schema serves **two purposes**:
|
|
6
|
+
|
|
7
|
+
1. **Output validation** — replaces type and shape checks (enums, ranges, required fields). One declaration instead of many.
|
|
8
|
+
2. **Provider-side request** — with the RubyLLM adapter, the schema is sent to the LLM provider via `chat.with_schema(...)`, asking the model to return JSON matching the shape. Cheaper models sometimes ignore the request, which is why client-side validation (point 1) still matters.
|
|
9
|
+
|
|
10
|
+
> **Same DSL `RubyLLM::Agent.schema` accepts.** `output_schema do ... end` here is a wrapper around `RubyLLM::Schema.create(&block)` plus a client-side validation step. `RubyLLM::Agent.schema` accepts the same block — choosing one over the other does not change the schema language. The difference: `Agent.schema` lets you pass a `Proc` evaluated in runtime context (dynamic per-call schema); `Step.output_schema` is eager-compiled at class load and additionally drives `output_type` inference and `Validator.validate`. Both can coexist.
|
|
11
|
+
|
|
12
|
+
All examples below extend the `SummarizeArticle` step from the [README](../../README.md).
|
|
13
|
+
|
|
14
|
+
## Schema replaces type and shape checks
|
|
15
|
+
|
|
16
|
+
```ruby
|
|
17
|
+
# WITHOUT schema — many validates:
|
|
18
|
+
validate("tldr must be a string") { |o| o[:tldr].is_a?(String) }
|
|
19
|
+
validate("takeaways must be an array") { |o| o[:takeaways].is_a?(Array) }
|
|
20
|
+
validate("takeaways 3 to 5") { |o| (3..5).cover?(o[:takeaways].size) }
|
|
21
|
+
ALLOWED_TONES = %w[neutral positive negative analytical].freeze
|
|
22
|
+
validate("tone must be an allowed label") { |o| ALLOWED_TONES.include?(o[:tone]) }
|
|
23
|
+
|
|
24
|
+
# WITH schema — one declaration:
|
|
25
|
+
output_schema do
|
|
26
|
+
string :tldr
|
|
27
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
28
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
29
|
+
end
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## Nested objects in arrays
|
|
33
|
+
|
|
34
|
+
Use `object do...end` inside `array` when you need more than a primitive per element. Concrete scenario: the UI card grows a "confidence bar" next to each takeaway so editors can see which points the model was sure about vs guessing. That requires `confidence` paired with `text`, not two parallel arrays that could desync. Nested objects make the pairing a schema invariant:
|
|
35
|
+
|
|
36
|
+
```ruby
|
|
37
|
+
output_schema do
|
|
38
|
+
string :tldr
|
|
39
|
+
array :takeaways, min_items: 3, max_items: 5 do
|
|
40
|
+
object do
|
|
41
|
+
string :text
|
|
42
|
+
number :confidence, minimum: 0.0, maximum: 1.0
|
|
43
|
+
end
|
|
44
|
+
end
|
|
45
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
46
|
+
end
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
## Schema pattern reference
|
|
50
|
+
|
|
51
|
+
| Your output looks like | Schema pattern | Example |
|
|
52
|
+
|---|---|---|
|
|
53
|
+
| `{"tldr": "...", "tone": "positive"}` | Flat fields | `string :tldr; string :tone, enum: [...]` |
|
|
54
|
+
| `{"takeaways": ["...", "..."]}` | Array of primitives | `array :takeaways, of: :string, min_items: 3, max_items: 5` |
|
|
55
|
+
| `{"takeaways": [{"text": "...", "confidence": 0.9}]}` | Array of objects | `array :takeaways do; object do; string :text; number :confidence; end; end` |
|
|
56
|
+
|
|
57
|
+
Without `object do...end`, `array :takeaways do; string :text; end` tells the provider "takeaways is an array of strings" — not objects. That's what you get back.
|
|
58
|
+
|
|
59
|
+
## Why schema alone is not enough
|
|
60
|
+
|
|
61
|
+
Schema validates **shape** — correct types, allowed values, field presence. But LLMs can return structurally valid JSON that is **logically wrong**. Validates catch what schema can't:
|
|
62
|
+
|
|
63
|
+
```ruby
|
|
64
|
+
output_schema do
|
|
65
|
+
string :tldr
|
|
66
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
67
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
# Schema allows any string for :tldr — but a 500-char "summary" breaks the UI card.
|
|
71
|
+
validate("TL;DR fits the card") { |o, _| o[:tldr].length <= 200 }
|
|
72
|
+
|
|
73
|
+
# Schema enforces 3–5 takeaways — but says nothing about them being distinct.
|
|
74
|
+
validate("takeaways are unique") { |o, _| o[:takeaways] == o[:takeaways].uniq }
|
|
75
|
+
|
|
76
|
+
# Schema can't express cross-field rules.
|
|
77
|
+
validate("critical tone requires at least one concrete risk") do |o, _|
|
|
78
|
+
next true unless o[:tone] == "negative"
|
|
79
|
+
o[:takeaways].any? { |t| t.match?(/fail|break|crash|outage|vulnerab/i) }
|
|
80
|
+
end
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
## Supported constraints
|
|
84
|
+
|
|
85
|
+
| Constraint | Types | Example |
|
|
86
|
+
|---|---|---|
|
|
87
|
+
| `enum` | string, integer | `string :tone, enum: %w[neutral positive negative analytical]` |
|
|
88
|
+
| `minimum` / `maximum` | number, integer | `number :confidence, minimum: 0.0, maximum: 1.0` |
|
|
89
|
+
| `min_length` / `max_length` | string | `string :tldr, min_length: 1, max_length: 200` |
|
|
90
|
+
| `min_items` / `max_items` | array | `array :takeaways, of: :string, min_items: 3, max_items: 5` |
|
|
91
|
+
| `additional_properties` | object | Set to `false` in the schema to reject extra keys |
|
|
92
|
+
|
|
93
|
+
Keyword args use Ruby snake_case (`min_length`, `min_items`). The DSL converts them internally to JSON Schema's camelCase (`minLength`, `minItems`) before sending the schema to the provider — you don't need to write camelCase in Ruby.
|
|
@@ -0,0 +1,154 @@
|
|
|
1
|
+
# Pipeline
|
|
2
|
+
|
|
3
|
+
> Read this when one step isn't enough — you need multi-step with fail-fast, automatic data threading, and per-step models.
|
|
4
|
+
|
|
5
|
+
Chain multiple steps with automatic data threading, fail-fast, per-step models, trace, and timeout.
|
|
6
|
+
|
|
7
|
+
A pipeline needs more than one step to be interesting. This guide grows the `SummarizeArticle` step from the [README](../../README.md) into a three-step content pipeline that tags and routes the summary to a UI card.
|
|
8
|
+
|
|
9
|
+
## Full example: article → summary → hashtags → card
|
|
10
|
+
|
|
11
|
+
```ruby
|
|
12
|
+
# Step 1 — the flagship step from README, unchanged.
|
|
13
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
14
|
+
prompt <<~PROMPT
|
|
15
|
+
Summarize this article for a UI card. Return a short TL;DR,
|
|
16
|
+
3 to 5 key takeaways, and a tone label.
|
|
17
|
+
|
|
18
|
+
{input}
|
|
19
|
+
PROMPT
|
|
20
|
+
|
|
21
|
+
output_schema do
|
|
22
|
+
string :tldr
|
|
23
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
24
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
25
|
+
end
|
|
26
|
+
|
|
27
|
+
validate("TL;DR fits the card") { |o, _| o[:tldr].length <= 200 }
|
|
28
|
+
end
|
|
29
|
+
|
|
30
|
+
# Step 2 — reads SummarizeArticle's output, produces hashtags suitable for social posts.
|
|
31
|
+
class GenerateHashtags < RubyLLM::Contract::Step::Base
|
|
32
|
+
input_type Hash
|
|
33
|
+
|
|
34
|
+
output_schema do
|
|
35
|
+
# Carry through the summary fields downstream consumers (and the next step) need.
|
|
36
|
+
string :tldr
|
|
37
|
+
array :takeaways, of: :string
|
|
38
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
39
|
+
# Add new field.
|
|
40
|
+
array :hashtags, of: :string, min_items: 2, max_items: 5
|
|
41
|
+
end
|
|
42
|
+
|
|
43
|
+
prompt do
|
|
44
|
+
rule "Preserve tldr / takeaways / tone exactly as given."
|
|
45
|
+
user "Article summary: {tldr}\nTone: {tone}\nGenerate 2 to 5 concise hashtags."
|
|
46
|
+
end
|
|
47
|
+
|
|
48
|
+
validate("tone preserved") { |o, input| o[:tone] == input[:tone] }
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
# Step 3 — final shape the UI card consumes.
|
|
52
|
+
class BuildArticleCard < RubyLLM::Contract::Step::Base
|
|
53
|
+
input_type Hash
|
|
54
|
+
|
|
55
|
+
output_schema do
|
|
56
|
+
string :headline
|
|
57
|
+
string :summary
|
|
58
|
+
array :hashtags, of: :string
|
|
59
|
+
string :sentiment_icon, enum: %w[😐 🙂 ⚠️ 🧠]
|
|
60
|
+
end
|
|
61
|
+
|
|
62
|
+
prompt do
|
|
63
|
+
rule "Headline <= 70 chars. Summary is the incoming tldr reprinted verbatim."
|
|
64
|
+
rule "Pick sentiment_icon from: 😐 neutral, 🙂 positive, ⚠️ negative, 🧠 analytical."
|
|
65
|
+
user "TL;DR: {tldr}\nTone: {tone}\nHashtags: {hashtags}"
|
|
66
|
+
end
|
|
67
|
+
|
|
68
|
+
validate("summary is the tldr verbatim") { |o, input| o[:summary] == input[:tldr] }
|
|
69
|
+
end
|
|
70
|
+
|
|
71
|
+
# Pipeline: summarize → hashtags → card
|
|
72
|
+
class ArticleCardPipeline < RubyLLM::Contract::Pipeline::Base
|
|
73
|
+
step SummarizeArticle, as: :summarize
|
|
74
|
+
step GenerateHashtags, as: :tag
|
|
75
|
+
step BuildArticleCard, as: :card
|
|
76
|
+
end
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
## Running and inspecting
|
|
80
|
+
|
|
81
|
+
```ruby
|
|
82
|
+
result = ArticleCardPipeline.run(article_text, context: { adapter: adapter })
|
|
83
|
+
result.ok? # => true
|
|
84
|
+
result.outputs_by_step[:summarize] # => { tldr: "...", takeaways: [...], tone: "analytical" }
|
|
85
|
+
result.outputs_by_step[:card] # => { headline: "...", summary: "...", ... }
|
|
86
|
+
result.trace.total_cost # => 0.000128 (all steps combined)
|
|
87
|
+
result.trace.total_latency_ms # => 2340
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
## Fail-fast behavior
|
|
91
|
+
|
|
92
|
+
When a step's schema, validate, or preflight check rejects the output, the pipeline stops there — downstream steps never run:
|
|
93
|
+
|
|
94
|
+
```ruby
|
|
95
|
+
# Summarize returns a TL;DR over 200 chars → the "TL;DR fits the card" validate fails
|
|
96
|
+
adapter = RubyLLM::Contract::Adapters::Test.new(response: {
|
|
97
|
+
tldr: "x" * 500,
|
|
98
|
+
takeaways: %w[one two three],
|
|
99
|
+
tone: "neutral"
|
|
100
|
+
})
|
|
101
|
+
|
|
102
|
+
result = ArticleCardPipeline.run("article text", context: { adapter: adapter })
|
|
103
|
+
result.failed? # => true
|
|
104
|
+
result.failed_step # => :summarize (validate rejected; retries exhausted)
|
|
105
|
+
# tag and card never run — no downstream tokens spent on garbage
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
## Per-step model override
|
|
109
|
+
|
|
110
|
+
```ruby
|
|
111
|
+
class ArticleCardPipeline < RubyLLM::Contract::Pipeline::Base
|
|
112
|
+
step SummarizeArticle, as: :summarize, model: "gpt-4.1-mini"
|
|
113
|
+
step GenerateHashtags, as: :tag, model: "gpt-4.1-nano"
|
|
114
|
+
step BuildArticleCard, as: :card, model: "gpt-4.1-nano"
|
|
115
|
+
end
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
## Timeout
|
|
119
|
+
|
|
120
|
+
```ruby
|
|
121
|
+
result = ArticleCardPipeline.run(article_text, timeout_ms: 30_000)
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
## Pipeline eval
|
|
125
|
+
|
|
126
|
+
```ruby
|
|
127
|
+
ArticleCardPipeline.define_eval("e2e") do
|
|
128
|
+
add_case "ruby 3.4 release",
|
|
129
|
+
input: "Ruby 3.4 ships with frozen string literals by default and better YJIT...",
|
|
130
|
+
expected: { sentiment_icon: "🧠" }
|
|
131
|
+
end
|
|
132
|
+
|
|
133
|
+
report = ArticleCardPipeline.run_eval("e2e", context: { model: "gpt-4.1-mini" })
|
|
134
|
+
report.print_summary
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
## Pretty print
|
|
138
|
+
|
|
139
|
+
```ruby
|
|
140
|
+
puts result
|
|
141
|
+
# Pipeline: ok 3 steps 1234ms 450+120 tokens trace=abc12345
|
|
142
|
+
|
|
143
|
+
result.pretty_print
|
|
144
|
+
# Full ASCII table with per-step outputs (Pipeline::Result)
|
|
145
|
+
|
|
146
|
+
# For eval reports, use print_summary instead:
|
|
147
|
+
report.print_summary
|
|
148
|
+
# Tabular pass/fail breakdown (Eval::Report)
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
## See also
|
|
152
|
+
|
|
153
|
+
- [Testing](testing.md) — `ArticleCardPipeline.test(..., responses: { summarize: ..., tag: ..., card: ... })` for pipeline-level spec adapters.
|
|
154
|
+
- [Optimizing retry_policy](optimizing_retry_policy.md) — `optimize_retry_policy` runs per-step; pipelines benchmark one step at a time.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Prompt AST
|
|
2
|
+
|
|
3
|
+
> Read this when your prompt has more than one shape (per-tenant, per-language, per-audience) and string concatenation is starting to drift.
|
|
4
|
+
|
|
5
|
+
Prompts are structured data, not strings. That matters the moment `SummarizeArticle` has to ship in more than one shape — a different audience per tenant, a different language per region, a different tone template for a B2B vs consumer card. Building those variants by string-concatenating a monolithic prompt leads to silent drift across environments. The AST gives you typed nodes (`system`, `rule`, `section`, `example`, `user`) that compose, diff, and snapshot-test cleanly.
|
|
6
|
+
|
|
7
|
+
Available node types:
|
|
8
|
+
|
|
9
|
+
```ruby
|
|
10
|
+
prompt do
|
|
11
|
+
system "You summarize articles for a UI card." # system message
|
|
12
|
+
rule "Return valid JSON only." # appended as separate system message
|
|
13
|
+
section "AUDIENCE", "Rails developers" # labeled system message: [AUDIENCE]\n...
|
|
14
|
+
example input: "Ruby 3.4 ships frozen strings...", # user/assistant few-shot pair
|
|
15
|
+
output: '{"tldr":"...","takeaways":[...],"tone":"analytical"}'
|
|
16
|
+
user "{input}" # user message with interpolation
|
|
17
|
+
end
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
Or just a plain string (wraps as a single user message):
|
|
21
|
+
|
|
22
|
+
```ruby
|
|
23
|
+
prompt "Summarize this article for a UI card. {input}"
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
The AST is immutable, diffable, and hashable. Useful for snapshot testing and auditing prompt changes.
|
|
27
|
+
|
|
28
|
+
## Hash inputs with variable interpolation
|
|
29
|
+
|
|
30
|
+
When input is a Hash, each key becomes a template variable. Concrete scenario: a multi-language newsletter product where the same article has to be summarised in Polish for EU subscribers, English for US, with different audiences per tier (Rails developers vs engineering managers). Hash inputs let one step cover all of these without forking the class:
|
|
31
|
+
|
|
32
|
+
```ruby
|
|
33
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
34
|
+
input_type RubyLLM::Contract::Types::Hash.schema(
|
|
35
|
+
article: RubyLLM::Contract::Types::String,
|
|
36
|
+
audience: RubyLLM::Contract::Types::String,
|
|
37
|
+
language: RubyLLM::Contract::Types::String
|
|
38
|
+
)
|
|
39
|
+
|
|
40
|
+
prompt do
|
|
41
|
+
system "You summarize articles for a UI card."
|
|
42
|
+
rule "Write the TL;DR and takeaways in {language}."
|
|
43
|
+
section "AUDIENCE", "{audience}"
|
|
44
|
+
user "{article}"
|
|
45
|
+
end
|
|
46
|
+
|
|
47
|
+
output_schema do
|
|
48
|
+
string :tldr
|
|
49
|
+
array :takeaways, of: :string, min_items: 3, max_items: 5
|
|
50
|
+
string :tone, enum: %w[neutral positive negative analytical]
|
|
51
|
+
end
|
|
52
|
+
end
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
Every `{key}` in a prompt node is pulled from the input hash at run time. Missing keys raise — making wire-up bugs loud, not silent.
|
|
56
|
+
|
|
57
|
+
> **Not the same as `RubyLLM::Agent.inputs`.** `Step.input_type` is a *runtime type check* on the positional argument passed to `run(input)` — it raises `TypeError` if the input violates the declared shape. `RubyLLM::Agent.inputs` is a list of *named template locals* injected into ERB instructions. They solve different problems (validation vs templating) and can coexist on the same project.
|
|
58
|
+
|
|
59
|
+
> **Not the same as `RubyLLM::Agent` ERB templates either.** The `prompt do ... end` DSL above builds a multi-role message list (`system` / `user` / `assistant` / `example` nodes) into a node-AST. `Agent.instructions :name` loads a single-string ERB file from `app/prompts/<agent_path>/<name>.txt.erb` for the system prompt only. Different output shape, different scope. Variable interpolation here uses `{key}` substitution; ERB uses full `<%= ruby %>`.
|
|
60
|
+
|
|
61
|
+
## Cross-validating output against input
|
|
62
|
+
|
|
63
|
+
Validate blocks support 2-arity `|output, input|` so you can check that the model's answer stays faithful to the request:
|
|
64
|
+
|
|
65
|
+
```ruby
|
|
66
|
+
validate("tldr is not just the article reprinted") do |output, input|
|
|
67
|
+
# Guard against lazy models that return the input verbatim.
|
|
68
|
+
output[:tldr].length < input[:article].length / 2
|
|
69
|
+
end
|
|
70
|
+
|
|
71
|
+
validate("no takeaway repeats the TL;DR") do |output, _input|
|
|
72
|
+
output[:takeaways].none? { |t| t == output[:tldr] }
|
|
73
|
+
end
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
The first example uses `input`; the second ignores it. Both are legal 2-arity signatures — Ruby accepts the unused `_input` parameter naming convention.
|
|
@@ -0,0 +1,218 @@
|
|
|
1
|
+
# Rails integration
|
|
2
|
+
|
|
3
|
+
> Read this when you've seen the `SummarizeArticle` example and want to know where contract steps fit in an actual Rails app — directory, initializer, jobs, logging, tests, CI. Skip if you're writing a non-Rails script.
|
|
4
|
+
|
|
5
|
+
Seven pre-emptive answers to the questions that come up first.
|
|
6
|
+
|
|
7
|
+
## 1. Where do step classes live?
|
|
8
|
+
|
|
9
|
+
**Recommended: `app/contracts/`.** The gem's Railtie auto-reloads eval files under `app/contracts/eval/` and `app/steps/eval/` in development, so picking `app/contracts/` aligns with the default reload paths.
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
app/contracts/summarize_article.rb # class SummarizeArticle
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Any autoloaded directory works (`app/llm_steps/`, `app/services/llm/`, etc.) — Rails 7/8 autoloading resolves them all, and the step class itself does not depend on the path. Pick the default if you have no stronger convention.
|
|
16
|
+
|
|
17
|
+
Keep evals in the same file as the step (`define_eval` block at the bottom of the class) — one source of truth per contract. If your evals grow too large for the class file, move them to `app/contracts/eval/summarize_article_eval.rb` — the Railtie reloads that directory explicitly in development.
|
|
18
|
+
|
|
19
|
+
## 2. Initializer configuration
|
|
20
|
+
|
|
21
|
+
```ruby
|
|
22
|
+
# config/initializers/ruby_llm_contract.rb
|
|
23
|
+
RubyLLM.configure do |c|
|
|
24
|
+
c.openai_api_key = ENV.fetch("OPENAI_API_KEY", nil)
|
|
25
|
+
c.anthropic_api_key = ENV.fetch("ANTHROPIC_API_KEY", nil)
|
|
26
|
+
end
|
|
27
|
+
|
|
28
|
+
RubyLLM::Contract.configure do |c|
|
|
29
|
+
c.default_model = Rails.env.production? ? "gpt-5-mini" : "gpt-5-nano"
|
|
30
|
+
c.default_adapter = RubyLLM::Contract::Adapters::RubyLLM.new
|
|
31
|
+
end
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
In specs, override the default adapter to `Adapters::Test` in `spec_helper.rb` (or use `stub_step` per-example — see §5).
|
|
35
|
+
|
|
36
|
+
Evals defined inside a step class (the recommended pattern) are picked up as soon as Rails autoloads the class — you do not need the eval file in any special directory. If you move evals into separate files under `app/contracts/eval/` or `app/steps/eval/`, the gem's Railtie reloads those two directories explicitly on each request in development; other directories follow standard Rails autoloading rules.
|
|
37
|
+
|
|
38
|
+
## 3. Background jobs — never call LLMs inline in a controller
|
|
39
|
+
|
|
40
|
+
LLM calls take 0.8–5 seconds and can fail. Wrap every step invocation in an ActiveJob:
|
|
41
|
+
|
|
42
|
+
```ruby
|
|
43
|
+
class SummarizeArticleJob < ApplicationJob
|
|
44
|
+
queue_as :llm
|
|
45
|
+
|
|
46
|
+
def perform(article_id)
|
|
47
|
+
article = Article.find(article_id)
|
|
48
|
+
result = SummarizeArticle.run(article.body)
|
|
49
|
+
|
|
50
|
+
if result.ok?
|
|
51
|
+
# parsed_output uses symbol keys in memory. jsonb/json columns round-trip
|
|
52
|
+
# as strings on reload, so either use deep_stringify_keys before write or
|
|
53
|
+
# access downstream with string keys — pick one convention and stick to it.
|
|
54
|
+
article.update!(summary: result.parsed_output.deep_stringify_keys)
|
|
55
|
+
else
|
|
56
|
+
article.update!(summary_error: result.validation_errors.join("; "))
|
|
57
|
+
end
|
|
58
|
+
end
|
|
59
|
+
end
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
`SummarizeArticleJob.perform_later(article.id)` returns in milliseconds; the controller stays responsive. If you use Sidekiq, pair `queue_as :llm` with a dedicated concurrency cap in `sidekiq.yml` so long-running LLM calls do not starve other job queues (mailers, webhooks, cleanups).
|
|
63
|
+
|
|
64
|
+
## 4. Logging and observability
|
|
65
|
+
|
|
66
|
+
`around_call` runs once per `run()` with the final `Result` (after all retries). Use it to write one row per LLM call:
|
|
67
|
+
|
|
68
|
+
```ruby
|
|
69
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
70
|
+
# ... prompt, schema, validates ...
|
|
71
|
+
|
|
72
|
+
around_call do |step, input, result|
|
|
73
|
+
AiCallLog.create!(
|
|
74
|
+
step: step.name,
|
|
75
|
+
model: result.trace[:model],
|
|
76
|
+
status: result.status.to_s,
|
|
77
|
+
latency_ms: result.trace[:latency_ms],
|
|
78
|
+
input_tokens: result.trace[:usage]&.dig(:input_tokens),
|
|
79
|
+
output_tokens: result.trace[:usage]&.dig(:output_tokens),
|
|
80
|
+
cost: result.trace[:cost],
|
|
81
|
+
validation_errors: result.validation_errors
|
|
82
|
+
)
|
|
83
|
+
end
|
|
84
|
+
end
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
The `AiCallLog` model assumed above is a thin audit record. One possible migration:
|
|
88
|
+
|
|
89
|
+
```ruby
|
|
90
|
+
# rails g model AiCallLog step:string model:string status:string ...
|
|
91
|
+
create_table :ai_call_logs do |t|
|
|
92
|
+
t.string :step, null: false
|
|
93
|
+
t.string :model
|
|
94
|
+
t.string :status, null: false
|
|
95
|
+
t.integer :latency_ms
|
|
96
|
+
t.integer :input_tokens
|
|
97
|
+
t.integer :output_tokens
|
|
98
|
+
t.decimal :cost, precision: 10, scale: 6
|
|
99
|
+
t.jsonb :validation_errors, default: []
|
|
100
|
+
t.timestamps
|
|
101
|
+
end
|
|
102
|
+
add_index :ai_call_logs, :step
|
|
103
|
+
add_index :ai_call_logs, :status
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
For Appsignal / Honeybadger / Datadog, emit an `ActiveSupport::Notifications` event from inside the same `around_call` and subscribe in an initializer:
|
|
107
|
+
|
|
108
|
+
```ruby
|
|
109
|
+
class SummarizeArticle < RubyLLM::Contract::Step::Base
|
|
110
|
+
# ... prompt, schema, validates ...
|
|
111
|
+
|
|
112
|
+
around_call do |step, _input, result|
|
|
113
|
+
ActiveSupport::Notifications.instrument(
|
|
114
|
+
"ruby_llm_contract.run",
|
|
115
|
+
step: step.name, model: result.trace[:model], status: result.status
|
|
116
|
+
)
|
|
117
|
+
end
|
|
118
|
+
end
|
|
119
|
+
|
|
120
|
+
# config/initializers/observability.rb
|
|
121
|
+
ActiveSupport::Notifications.subscribe("ruby_llm_contract.run") do |*, payload|
|
|
122
|
+
Appsignal.increment_counter("llm.run.#{payload[:status]}", 1, step: payload[:step])
|
|
123
|
+
end
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
Trace inspection in an admin UI: `result.trace[:attempts]` gives you per-attempt model, status, cost, latency — render it in a partial to debug production failures without re-running.
|
|
127
|
+
|
|
128
|
+
## 5. Testing — RSpec and Minitest
|
|
129
|
+
|
|
130
|
+
Add to `spec/spec_helper.rb` (or `test_helper.rb`):
|
|
131
|
+
|
|
132
|
+
```ruby
|
|
133
|
+
require "ruby_llm/contract/rspec" # or ruby_llm/contract/minitest
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
Then in specs:
|
|
137
|
+
|
|
138
|
+
```ruby
|
|
139
|
+
RSpec.describe ArticlesController do
|
|
140
|
+
it "saves the summary when the step passes" do
|
|
141
|
+
stub_step(SummarizeArticle, response: {
|
|
142
|
+
tldr: "...", takeaways: %w[a b c], tone: "analytical"
|
|
143
|
+
})
|
|
144
|
+
|
|
145
|
+
post :summarize, params: { id: article.id }
|
|
146
|
+
|
|
147
|
+
# NOTE: jsonb/json column round-trips as string keys on reload.
|
|
148
|
+
expect(article.reload.summary["tldr"]).to eq("...")
|
|
149
|
+
end
|
|
150
|
+
end
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
For the step itself, use the `satisfy_contract` and `pass_eval` matchers — details in the [Testing guide](testing.md).
|
|
154
|
+
|
|
155
|
+
## 6. Error handling in controllers
|
|
156
|
+
|
|
157
|
+
Never raise on `result.failed?` — that crashes the request. Branch instead:
|
|
158
|
+
|
|
159
|
+
```ruby
|
|
160
|
+
class ArticlesController < ApplicationController
|
|
161
|
+
def summarize
|
|
162
|
+
SummarizeArticleJob.perform_later(params[:id])
|
|
163
|
+
head :accepted
|
|
164
|
+
end
|
|
165
|
+
|
|
166
|
+
# For synchronous cases (admin tools, small content):
|
|
167
|
+
def preview
|
|
168
|
+
result = SummarizeArticle.run(@article.body)
|
|
169
|
+
|
|
170
|
+
if result.ok?
|
|
171
|
+
render json: result.parsed_output
|
|
172
|
+
else
|
|
173
|
+
Rails.logger.warn "[llm] #{SummarizeArticle.name} failed: #{result.status}"
|
|
174
|
+
render json: { error: "Could not summarize; try again shortly." }, status: :service_unavailable
|
|
175
|
+
end
|
|
176
|
+
end
|
|
177
|
+
end
|
|
178
|
+
```
|
|
179
|
+
|
|
180
|
+
When `retry_policy` exhausts and all models fail, `result.failed?` is true but `result.parsed_output` still contains the last attempt's output — useful for logging what the model *did* return before the validate rejected it.
|
|
181
|
+
|
|
182
|
+
## 7. CI gate — block regressions before merge
|
|
183
|
+
|
|
184
|
+
Add to your `Rakefile`:
|
|
185
|
+
|
|
186
|
+
```ruby
|
|
187
|
+
require "ruby_llm/contract/rake_task"
|
|
188
|
+
|
|
189
|
+
RubyLLM::Contract::RakeTask.new do |t|
|
|
190
|
+
t.minimum_score = 0.8
|
|
191
|
+
t.maximum_cost = 0.05
|
|
192
|
+
t.fail_on_regression = true
|
|
193
|
+
t.save_baseline = false # read-only in CI; refresh baselines manually
|
|
194
|
+
end
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
Then wire it in GitHub Actions:
|
|
198
|
+
|
|
199
|
+
```yaml
|
|
200
|
+
- name: LLM contract evals
|
|
201
|
+
env:
|
|
202
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
203
|
+
run: bundle exec rake ruby_llm_contract:eval
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
The job fails when a previously-passing eval case now fails, when the average score drops below the threshold, or when total cost exceeds the cap. That is the signal that blocks a prompt regression or an accidental model upgrade from shipping.
|
|
207
|
+
|
|
208
|
+
**Two practical notes:**
|
|
209
|
+
|
|
210
|
+
- **Live evals spend real money on every run** — provider tokens per case × number of cases × every merge. Keep the dataset small and targeted (5–15 high-value cases), use cheap models where quality allows, and rely on offline `sample_response` smoke tests in the bulk of CI runs. Live evals belong on merge-candidate branches and scheduled nightly runs, not on every commit.
|
|
211
|
+
- **Baselines are checkout-managed** — commit them to git under `.eval_baselines/`. Refresh them in a separate manual workflow (or locally + a dedicated PR) rather than from the merge gate, which would dirty the checkout and race with the regression check it is supposed to run.
|
|
212
|
+
|
|
213
|
+
## See also
|
|
214
|
+
|
|
215
|
+
- [Getting Started](getting_started.md) — the feature walkthrough the step above is built on
|
|
216
|
+
- [Migration](migration.md) — before/after for replacing a raw `LlmClient.new.call` service with a contract
|
|
217
|
+
- [Eval-First](eval_first.md) — the workflow behind the CI gate above
|
|
218
|
+
- [Testing](testing.md) — `satisfy_contract` and `pass_eval` matcher chains
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# Relation to `RubyLLM::Agent`
|
|
2
|
+
|
|
3
|
+
> Read this when you already use `RubyLLM::Agent` (or are about to) and want to understand where `ruby_llm-contract` sits in the same project.
|
|
4
|
+
|
|
5
|
+
`RubyLLM::Agent` shipped in RubyLLM 1.12. `Step::Base` from this gem and `Agent` target the **same niche**: reusable, class-based prompts. They are **siblings**, not foundation-and-roof.
|
|
6
|
+
|
|
7
|
+
## Feature mapping
|
|
8
|
+
|
|
9
|
+
| What you write | Where it lives |
|
|
10
|
+
|---|---|
|
|
11
|
+
| `model`, `temperature`, `schema`, `instructions`, `tools`, `thinking` | covered by both — same idea, different DSL surface |
|
|
12
|
+
| `validate :rule do ... end` business invariants on output | only in `ruby_llm-contract` |
|
|
13
|
+
| `retry_policy escalate(...)` model escalation on validation failure | only here (different from RubyLLM's network-level retry) |
|
|
14
|
+
| `max_cost` / `max_input` / `max_output` pre-flight refusal | only here |
|
|
15
|
+
| `define_eval` + baseline regression + `compare_models` + `optimize_retry_policy` | only here (RubyLLM does not ship an evaluation framework) |
|
|
16
|
+
| Pipeline composition with `step SomeStep, as: :alias` | only here (RubyLLM intentionally leaves workflows as plain Ruby) |
|
|
17
|
+
| `around_call`, named `observe` hooks with pass/fail recorded in trace | only here |
|
|
18
|
+
|
|
19
|
+
## Runtime relationship
|
|
20
|
+
|
|
21
|
+
`Step::Base` does **not** use `Agent` internally today. The actual call path is:
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
Step.run(input)
|
|
25
|
+
→ Runner
|
|
26
|
+
→ Adapters::RubyLLM
|
|
27
|
+
→ RubyLLM.chat(model:, ...)
|
|
28
|
+
→ ... .ask(prompt)
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
`Agent` is a sibling abstraction calling into `RubyLLM::Chat` through its own `apply_configuration` path. Both end up at `Chat`. They do not share the macro-storage layer.
|
|
32
|
+
|
|
33
|
+
This may change in a future release if upstream APIs make a layered design natural. The decision is not committed; it depends on adopter signal.
|
|
34
|
+
|
|
35
|
+
## Coexistence on the same project
|
|
36
|
+
|
|
37
|
+
The two abstractions can live in the same Rails (or non-Rails) project. Pick one per use case:
|
|
38
|
+
|
|
39
|
+
- **`RubyLLM::Agent`** when you want a reusable prompt with `model` + `instructions` + `schema` + `tools` and that is enough — no retry-on-validation-failure, no business invariants, no eval framework, no budget gating.
|
|
40
|
+
- **`ruby_llm-contract`'s `Step::Base`** when you need any of: invariants (`validate`), retry with model escalation on validation failure, pre-flight cost ceilings, an evaluation framework with baseline regression, or pipeline composition.
|
|
41
|
+
|
|
42
|
+
A common pattern: simple ad-hoc prompts as `Agent`, contracts on the LLM features that touch production behaviour or money as `Step`.
|
|
43
|
+
|
|
44
|
+
## On retry strategies
|
|
45
|
+
|
|
46
|
+
The retry-strategy framing in this gem favours `retry_policy escalate(model_2, ...)` (model escalation, addresses model bias) over same-model `retry_policy attempts: N` (variance retry).
|
|
47
|
+
|
|
48
|
+
This is grounded in empirical comparison across PDF quiz generation, GSM8K math (n=30 + n=120), and multi-constraint schedule generation: same-model retry produced no useful lift for nano-class models on tasks with clear correctness criteria. Model escalation did move the needle when same-model retry could not.
|
|
49
|
+
|
|
50
|
+
`attempts: N` stays in the gem API (backward compat + niche cases like subjective-criteria tasks, multi-step pipelines, weaker open-source models) but is not marketed as a default retry strategy.
|
|
51
|
+
|
|
52
|
+
See [Optimize retry policy](optimizing_retry_policy.md) for the empirical tooling.
|