squishling 0.2.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +83 -0
- data/README.md +85 -30
- data/docs/configuration.md +40 -16
- data/docs/failures.md +68 -8
- data/docs/harnesses.md +198 -0
- data/docs/measuring-tokens.md +148 -0
- data/docs/naming.md +43 -0
- data/docs/routing.md +36 -27
- data/docs/schemas.md +4 -3
- data/lib/squishling/appendices.rb +3 -3
- data/lib/squishling/class_methods.rb +60 -29
- data/lib/squishling/collisions.rb +54 -0
- data/lib/squishling/configuration.rb +19 -0
- data/lib/squishling/definition.rb +48 -23
- data/lib/squishling/errors.rb +22 -1
- data/lib/squishling/harness.rb +153 -0
- data/lib/squishling/invoker.rb +188 -133
- data/lib/squishling/judge.rb +142 -0
- data/lib/squishling/llm_client.rb +150 -0
- data/lib/squishling/model_path.rb +21 -7
- data/lib/squishling/output_check.rb +109 -0
- data/lib/squishling/params.rb +37 -1
- data/lib/squishling/router.rb +8 -2
- data/lib/squishling/source.rb +5 -5
- data/lib/squishling/squawk.rb +52 -0
- data/lib/squishling/version.rb +1 -1
- data/lib/squishling.rb +24 -7
- metadata +12 -3
data/docs/harnesses.md
ADDED
|
@@ -0,0 +1,198 @@
|
|
|
1
|
+
# Harnesses: escalation, squishsum, and ensemble
|
|
2
|
+
|
|
3
|
+
A squished call's **harness** decides how it uses its [escalation](configuration.md#models-and-escalation):
|
|
4
|
+
|
|
5
|
+
| Harness | What it does | Requests per call |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| `:escalation` (default) | Tries each attempt in order until an output passes the schema and `squish_validate`. | 1 or more |
|
|
8
|
+
| `:squishsum` | Asks the escalation's **first step** twice, concurrently. Identical outputs are accepted; different ones raise `Squishling::DisagreementError`. | 2 or more |
|
|
9
|
+
| `:judged_squishsum` | Like `:squishsum`, but when the two outputs differ, a **judge** picks one or rejects both. | 2 or more, plus 1 or more judge requests when they differ |
|
|
10
|
+
| `:ensemble` | Asks the escalation's **first and second steps** once each, concurrently, so two different models are compared. Identical outputs are accepted; different ones raise `DisagreementError`. | 2 or more |
|
|
11
|
+
| `:judged_ensemble` | Like `:ensemble`, but when the two outputs differ, a **judge** picks one or rejects both. | 2 or more, plus 1 or more judge requests when they differ |
|
|
12
|
+
|
|
13
|
+
Use a squishsum harness when a confidently wrong answer costs more than a second request: one sample can make a
|
|
14
|
+
mistake, but two independent samples rarely make the same one. Use an [ensemble](#ensembles) when the two samples
|
|
15
|
+
should come from different models: a model rarely disagrees with itself at low temperature, but a cheap and a
|
|
16
|
+
stronger model often disagree on exactly the cases the cheap one gets confidently wrong.
|
|
17
|
+
|
|
18
|
+
```ruby
|
|
19
|
+
class TicketTriager
|
|
20
|
+
include Squishling
|
|
21
|
+
|
|
22
|
+
squishling escalation: [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5"],
|
|
23
|
+
harness: :judged_squishsum # two Haiku samples; Sonnet judges when they differ
|
|
24
|
+
purpose "Assign a priority and a team."
|
|
25
|
+
output_schema do
|
|
26
|
+
string :priority, enum: %w[low medium high]
|
|
27
|
+
string :team
|
|
28
|
+
end
|
|
29
|
+
end
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## Declaring a harness
|
|
33
|
+
|
|
34
|
+
`harness:` takes a type, or a Hash with `type:` and its options:
|
|
35
|
+
|
|
36
|
+
```ruby
|
|
37
|
+
Squishling.configure { |c| c.default_harness = :squishsum } # every class without its own
|
|
38
|
+
|
|
39
|
+
squishling harness: :judged_squishsum # this class (and subclasses)
|
|
40
|
+
squish :triage, harness: { type: :squishsum, # one method
|
|
41
|
+
compare: ->(a, b, **) { a.priority == b.priority } }
|
|
42
|
+
squish!(harness: :escalation) # one call (see routing.md)
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
The first level that declares one wins: `squish!`, then the method, then the class, then
|
|
46
|
+
`Squishling.config.default_harness`, then `:escalation`. Options aren't merged across levels.
|
|
47
|
+
|
|
48
|
+
| Option | Harnesses | Description |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| `type:` | all | `:escalation`, `:squishsum`, `:judged_squishsum`, `:ensemble`, or `:judged_ensemble` (required in the Hash form). |
|
|
51
|
+
| `compare:` | squishsum and ensemble ones | `->(a, b, **inputs) { ... }`, evaluated against the instance with the two typed results and the method's inputs. Truthy means the samples agree. Default: the two outputs are exactly equal. |
|
|
52
|
+
| `judge:` | judged ones | The judge step (see [The judge](#the-judge)). Default: the escalation's next step after the sampled ones. |
|
|
53
|
+
| `judge_instructions:` | judged ones | A String, or a Proc evaluated against the instance, that replaces the [default judge prompt](#the-default-judge-prompt). |
|
|
54
|
+
|
|
55
|
+
## Samples
|
|
56
|
+
|
|
57
|
+
Under `:squishsum`, both samples run on the escalation's first step, with its model, provider, and params
|
|
58
|
+
(an [ensemble](#ensembles) runs them on its first and second steps):
|
|
59
|
+
|
|
60
|
+
- **They're independent.** Each sample has its own chat, and the two requests are sent concurrently on separate
|
|
61
|
+
threads. Only the provider requests run on those threads. Parsing, schema validation, `squish_validate`,
|
|
62
|
+
`compare:`, and the judge all run on the calling thread, so your Squishling callbacks never run concurrently.
|
|
63
|
+
RubyLLM instrumentation subscribers (e.g. `ActiveSupport::Notifications` listeners for `chat.ruby_llm`) do see
|
|
64
|
+
each sample's request on its worker thread, outside the Rails executor, so keep them thread-safe.
|
|
65
|
+
- **Each one is retried like an escalation step.** A sample with invalid output is re-asked in its own chat, up to
|
|
66
|
+
their step's `attempts:`, while a sample that already passed keeps its result. A failed request gets a fresh
|
|
67
|
+
chat. A sample still failing after its attempts fails the call with `InvalidOutputError` or `LLMError`, as
|
|
68
|
+
usual. Under `:squishsum`, later escalation steps are never used for samples (see [Ensembles](#ensembles)).
|
|
69
|
+
- **Agreement is exact by default.** Free-text fields seldom match word for word, so pass `compare:` to decide
|
|
70
|
+
which fields must match:
|
|
71
|
+
|
|
72
|
+
```ruby
|
|
73
|
+
squishling harness: { type: :squishsum, compare: ->(a, b, **) { a.priority == b.priority && a.team == b.team } }
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
When the samples agree, the first one is returned.
|
|
77
|
+
|
|
78
|
+
## Ensembles
|
|
79
|
+
|
|
80
|
+
`:ensemble` and `:judged_ensemble` sample two *different* steps instead of one step twice. Sample `a` is the
|
|
81
|
+
escalation's first step and sample `b` its second, each with its own model, provider, and params:
|
|
82
|
+
|
|
83
|
+
```ruby
|
|
84
|
+
class TicketTriager
|
|
85
|
+
include Squishling
|
|
86
|
+
|
|
87
|
+
squishling escalation: [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5", "claude-opus-5-5"],
|
|
88
|
+
harness: :judged_ensemble # a = Haiku, b = Sonnet, Opus judges when they differ
|
|
89
|
+
# purpose, output_schema, ...
|
|
90
|
+
end
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
- A step's `attempts:` are that sample's retries (here Haiku gets two attempts, Sonnet one), exactly as in
|
|
94
|
+
[Samples](#samples). Steps after the second are never sampled.
|
|
95
|
+
- Everything else matches squishsum: independence and concurrency, `compare:`, `DisagreementError`, and the
|
|
96
|
+
judge. `DisagreementError#models` names the model behind each sample.
|
|
97
|
+
- The judge defaults to the escalation's **third** step. With only two steps, a `:judged_ensemble` needs `judge:`.
|
|
98
|
+
An ensemble with fewer than two steps, or a judged one with neither a third step nor `judge:`, raises
|
|
99
|
+
`ConfigurationError` before any request.
|
|
100
|
+
- A global `default_harness = :ensemble` therefore needs every class to declare at least two steps, and a
|
|
101
|
+
`squish!(model: ...)` override (one step) can't be combined with an ensemble harness.
|
|
102
|
+
- Steps are told apart by value, so two adjacent steps with the same model, provider, and params count as one step.
|
|
103
|
+
|
|
104
|
+
## The judge
|
|
105
|
+
|
|
106
|
+
With a judged harness, two valid samples that disagree go to a judge. The judge picks candidate `a` or `b`,
|
|
107
|
+
whose typed result is returned (`squished?` is `true`), or it picks `neither`, which raises
|
|
108
|
+
`DisagreementError`.
|
|
109
|
+
|
|
110
|
+
The judge is:
|
|
111
|
+
|
|
112
|
+
1. the `judge:` step, if declared: a model name, or a Hash with `model:`, `provider:`, `params:`, and `attempts:`
|
|
113
|
+
(as in an escalation step, without `order:`), plus `type:`; or
|
|
114
|
+
2. the escalation's next step after the sampled ones (the second under `:judged_squishsum`, the third under
|
|
115
|
+
`:judged_ensemble`), with its attempts.
|
|
116
|
+
|
|
117
|
+
A judged harness with neither raises `ConfigurationError` before any sample is requested. A `judge:`
|
|
118
|
+
step uses only its own `provider:`; it doesn't inherit the class or method `provider:`, which belongs to their models
|
|
119
|
+
(a chat judge's params, though, do merge over the method's).
|
|
120
|
+
|
|
121
|
+
### Chat judges
|
|
122
|
+
|
|
123
|
+
By default the judge is a chat model, asked for a strict `{ "verdict": "a" | "b" | "neither", "reason": "..." }`.
|
|
124
|
+
It gets the same context the samples had, plus the samples themselves:
|
|
125
|
+
|
|
126
|
+
- **System prompt:** the judge prompt, then the operation's own purpose (including any `append_to_purpose`
|
|
127
|
+
sections).
|
|
128
|
+
- **User message:** JSON with the samples' `input` (`arguments` and `context`), the `output_schema`, and the
|
|
129
|
+
`candidates` `a` and `b`.
|
|
130
|
+
|
|
131
|
+
A verdict that doesn't match its schema is retried per the judge step's `attempts:`; a judge that never returns a
|
|
132
|
+
valid verdict raises `InvalidOutputError`. A chat judge's params are the method's
|
|
133
|
+
[generation params](configuration.md#generation-params) with its own `params:` merged over them.
|
|
134
|
+
|
|
135
|
+
### System One judgment models (Jev)
|
|
136
|
+
|
|
137
|
+
A judge can also be a System One decision model, such as TypeSafe's Jev, via
|
|
138
|
+
[RubyLLM judgments](https://rubyllm.com/judgments/) (RubyLLM 2.1+). Decision models return calibrated
|
|
139
|
+
probabilities rather than text, so they are much faster and cheaper than a chat judge.
|
|
140
|
+
|
|
141
|
+
```ruby
|
|
142
|
+
squishling harness: {
|
|
143
|
+
type: :judged_squishsum,
|
|
144
|
+
judge: { model: "jev-latest", type: :judgment, min_confidence: 0.8 }
|
|
145
|
+
}
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
Squishling asks one `choice` question (`a`, `b`, or `neither`), with the judge prompt as its instructions and the
|
|
149
|
+
operation's purpose, input, output schema, and candidates as the judgment input. A pick of `a` or `b` is
|
|
150
|
+
accepted only when its probability is at least `min_confidence:` (default `0.8`). Otherwise, as with `neither`,
|
|
151
|
+
`DisagreementError` is raised with the choice and its probability in its `reason`. A judgment judge sends only its own
|
|
152
|
+
`params:` (as RubyLLM `provider_options:`). `attempts:` retries a failed request. Configure the provider's
|
|
153
|
+
credentials in RubyLLM as usual.
|
|
154
|
+
|
|
155
|
+
### The default judge prompt
|
|
156
|
+
|
|
157
|
+
```text
|
|
158
|
+
You are an impartial judge. Two independent attempts at the same operation returned different outputs,
|
|
159
|
+
candidate "a" and candidate "b". Both already match the required output format, so judge only their
|
|
160
|
+
content. You are given the operation's purpose and its input. Decide which candidate correctly and
|
|
161
|
+
faithfully carries out the purpose for this input. If both do (for example, they differ only in
|
|
162
|
+
wording), choose either one, preferring "a". Choose "neither" only when both are wrong or you can't tell
|
|
163
|
+
whether either is correct. Never combine them or invent a third answer.
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
It is available as `Squishling::Harness::DEFAULT_JUDGE_INSTRUCTIONS`. Replace it with `judge_instructions:`:
|
|
167
|
+
|
|
168
|
+
```ruby
|
|
169
|
+
squishling harness: { type: :judged_squishsum,
|
|
170
|
+
judge_instructions: "Pick the candidate whose priority follows our SLA rules: ..." }
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
## When the harness fails
|
|
174
|
+
|
|
175
|
+
A disagreement raises `Squishling::DisagreementError`, a subclass of `InvalidOutputError`. It's raised when the
|
|
176
|
+
samples differ with no judge, or when the judge rejects both. Like any failure on the LLM path, it goes to
|
|
177
|
+
[`squish_fallback`](failures.md#fallbacks), which can still return one of the candidates:
|
|
178
|
+
|
|
179
|
+
```ruby
|
|
180
|
+
squish_fallback do |error, **|
|
|
181
|
+
raise error unless error.is_a?(Squishling::DisagreementError)
|
|
182
|
+
|
|
183
|
+
error.candidates.first # already typed and validated; squished? is true
|
|
184
|
+
end
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
| Attribute | Value |
|
|
188
|
+
|---|---|
|
|
189
|
+
| `candidates` | The two typed results, `[a, b]` |
|
|
190
|
+
| `verdict` | `nil` when there was no judge; `:neither` when the judge rejected both |
|
|
191
|
+
| `reason` | The judge's reason (or, for a judgment model, the choice and its probability). A chat judge's reason is model-written, so it is not part of `message` or the logs; a judgment model's choice and probability are |
|
|
192
|
+
| `raw` | Both candidates as Hashes, `[a.to_h, b.to_h]` |
|
|
193
|
+
| `attempts` | `nil` |
|
|
194
|
+
| `models` | The model behind each role: sample `a`, sample `b`, then the judge if one ran (one entry each, not one per attempt) |
|
|
195
|
+
|
|
196
|
+
Other failures are unchanged. A sample or judge that never produces valid output raises `InvalidOutputError`, a
|
|
197
|
+
failed request on the last attempt raises `LLMError`, and setup mistakes (including a provider 400 on either
|
|
198
|
+
sample) raise `ConfigurationError`, which is never sent to the fallback.
|
|
@@ -0,0 +1,148 @@
|
|
|
1
|
+
# Measuring token spend
|
|
2
|
+
|
|
3
|
+
A squished path costs tokens on every call. A Ruby path costs engineer time once, plus upkeep. Measuring the
|
|
4
|
+
first tells you when the second is worth paying.
|
|
5
|
+
|
|
6
|
+
## What Squishling gives you
|
|
7
|
+
|
|
8
|
+
- `result.squished?` says which path served a call, so you can count LLM calls per class.
|
|
9
|
+
- The [`squawk` hook](failures.md#observing-every-attempt-squawk) runs after every LLM attempt and receives the
|
|
10
|
+
method (`label`, e.g. `"ShipmentUpdate#call"`), the model and provider, which attempt it was, and the token usage
|
|
11
|
+
RubyLLM reported. That is enough to attribute spend to a squishling method without any other tooling.
|
|
12
|
+
- Cost in dollars comes from RubyLLM's model pricing. Multiply the reported tokens by your model's price, or read it
|
|
13
|
+
from RubyLLM directly (see [the instrumenter](#lower-level-rubyllms-instrumenter)).
|
|
14
|
+
|
|
15
|
+
## Deciding when to write the code
|
|
16
|
+
|
|
17
|
+
For each squished path, over a month:
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
token cost = calls × (input tokens × input price + output tokens × output price)
|
|
21
|
+
code cost = time to write it + upkeep (new cases, format changes, on-call)
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Write the Ruby when the token cost keeps exceeding the code cost, and narrow `squish_when` so only the inputs
|
|
25
|
+
Ruby can't handle still go to the LLM. See [Hardening a path](routing.md#hardening-a-path). If the token cost
|
|
26
|
+
stays small, the debt you haven't taken on is cheaper than the code.
|
|
27
|
+
|
|
28
|
+
## Tools
|
|
29
|
+
|
|
30
|
+
| Option | Good for | Setup |
|
|
31
|
+
|---|---|---|
|
|
32
|
+
| [Coolhand Labs](#coolhand-labs) | Picking up your squishling calls automatically and measuring their cost and accuracy | The `coolhand` gem and an initializer |
|
|
33
|
+
| [Custom: the `squawk` hook](#custom-the-squawk-hook) | Logging or metrics in your own stack, with no new dependency | A lambda and one config line |
|
|
34
|
+
| [OpenTelemetry](#opentelemetry) | Sending traces to Datadog, Langfuse, Arize, Braintrust, LangSmith, or any OTel backend | The OpenTelemetry SDK, an exporter, and one line |
|
|
35
|
+
| [LangSmith](#langsmith) | Browsing traces and token usage in LangSmith | Through OpenTelemetry |
|
|
36
|
+
|
|
37
|
+
### Coolhand Labs
|
|
38
|
+
|
|
39
|
+
Add and initialize the [`coolhand`](https://github.com/Coolhand-Labs/coolhand-ruby) gem. It intercepts the LLM
|
|
40
|
+
requests your app makes, so it picks up your squishling calls automatically and measures their cost, and their
|
|
41
|
+
accuracy through Coolhand's feedback tools. No changes to your squishling classes.
|
|
42
|
+
|
|
43
|
+
```ruby
|
|
44
|
+
# Gemfile
|
|
45
|
+
gem "coolhand"
|
|
46
|
+
|
|
47
|
+
# config/initializers/coolhand.rb
|
|
48
|
+
Coolhand.configure do |config|
|
|
49
|
+
config.api_key = ENV.fetch("COOLHAND_API_KEY")
|
|
50
|
+
end
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Get an API key at [coolhandlabs.com](https://coolhandlabs.com/). To keep the data in your own infrastructure, point
|
|
54
|
+
`config.base_url` at a Coolhand-compatible endpoint of your own.
|
|
55
|
+
|
|
56
|
+
### Custom: the `squawk` hook
|
|
57
|
+
|
|
58
|
+
`config.squawk` runs after every LLM attempt, accepted or rejected, and is handed the method's label, the model,
|
|
59
|
+
and the token usage RubyLLM reported. A lambda can take just the keywords it needs:
|
|
60
|
+
|
|
61
|
+
```ruby
|
|
62
|
+
Squishling.configure do |config|
|
|
63
|
+
config.squawk = lambda do |metadata:, **|
|
|
64
|
+
usage = metadata[:usage] || {} # nil when the provider call failed; counts it didn't report are absent
|
|
65
|
+
Rails.logger.info("squishling #{metadata[:label]} model=#{metadata[:model]} " \
|
|
66
|
+
"attempt=#{metadata[:attempt]}/#{metadata[:attempts]} " \
|
|
67
|
+
"in=#{usage[:input_tokens]} out=#{usage[:output_tokens]}")
|
|
68
|
+
end
|
|
69
|
+
end
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
Each line carries the `"Class#method"` that made the call, so grouping by label gives you tokens per squishling. A call
|
|
73
|
+
that stays in Ruby makes no LLM request and never reaches the hook, so a label that stops appearing is a path that has
|
|
74
|
+
been hardened. You can also set the hook per class (`squishling squawk: ...`) or per method. See
|
|
75
|
+
[Observing every attempt](failures.md#observing-every-attempt-squawk) for every keyword and the rules for exceptions.
|
|
76
|
+
|
|
77
|
+
### Lower level: RubyLLM's instrumenter
|
|
78
|
+
|
|
79
|
+
To meter every RubyLLM request in your app, not only Squishling's, use RubyLLM's own instrumentation. It reports a
|
|
80
|
+
`usage.ruby_llm` event for each provider request, with `model`, `provider`, `status`, `tokens` (`input`, `output`,
|
|
81
|
+
and more) and `cost`. Give it an object that responds to `instrument(name, payload)` and optionally takes a block. In
|
|
82
|
+
Rails, `ActiveSupport::Notifications` is used automatically.
|
|
83
|
+
|
|
84
|
+
```ruby
|
|
85
|
+
class TokenMeter
|
|
86
|
+
def instrument(name, payload = {})
|
|
87
|
+
if name == "usage.ruby_llm"
|
|
88
|
+
Rails.logger.info("llm model=#{payload[:model]} workflow=#{payload[:workflow_name]} " \
|
|
89
|
+
"in=#{payload[:tokens].input} out=#{payload[:tokens].output}")
|
|
90
|
+
end
|
|
91
|
+
block_given? ? yield : nil
|
|
92
|
+
end
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
RubyLLM.configure { |config| config.instrumenter = TokenMeter.new }
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
These events don't know which squishling made the request. To attribute the spend, wrap the call in a RubyLLM
|
|
99
|
+
workflow, and every event emitted inside the block carries its name, id and any `metadata:` you pass. (The
|
|
100
|
+
`squawk` hook above needs no wrapping.)
|
|
101
|
+
|
|
102
|
+
```ruby
|
|
103
|
+
RubyLLM.workflow("ShipmentUpdate", metadata: { carrier: "acme-freight" }) do
|
|
104
|
+
ShipmentUpdate.call(carrier: "acme-freight", payload: body)
|
|
105
|
+
end
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
### OpenTelemetry
|
|
109
|
+
|
|
110
|
+
RubyLLM 2.1+ ships its own OpenTelemetry tracer. Each request becomes a span with the model, request settings and
|
|
111
|
+
token usage (`gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`), following the OpenTelemetry GenAI
|
|
112
|
+
conventions. It does not record message text. Your app owns the SDK and exporters.
|
|
113
|
+
|
|
114
|
+
```ruby
|
|
115
|
+
# Gemfile
|
|
116
|
+
gem "opentelemetry-sdk"
|
|
117
|
+
gem "opentelemetry-exporter-otlp" # or your backend's exporter
|
|
118
|
+
|
|
119
|
+
# config/initializers/opentelemetry.rb
|
|
120
|
+
OpenTelemetry::SDK.configure # reads the standard OTEL_* environment variables
|
|
121
|
+
RubyLLM::OpenTelemetry.enable # starts tracing RubyLLM requests
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
Spans don't say which squishling made the request, so pair this with a `RubyLLM.workflow` (see
|
|
125
|
+
[above](#lower-level-rubyllms-instrumenter)) or group by model and prompt in your backend.
|
|
126
|
+
|
|
127
|
+
### LangSmith
|
|
128
|
+
|
|
129
|
+
LangSmith accepts OpenTelemetry traces, so the OpenTelemetry setup above covers it: point the OTLP exporter at
|
|
130
|
+
LangSmith. See LangSmith's
|
|
131
|
+
[OpenTelemetry docs](https://docs.smith.langchain.com/observability/how_to_guides/tracing/trace_with_opentelemetry)
|
|
132
|
+
for the endpoint and headers. A community Ruby gem, [`langsmith-sdk`](https://github.com/felipekb/langsmith-ruby-sdk),
|
|
133
|
+
also exists; it isn't published by LangChain.
|
|
134
|
+
|
|
135
|
+
## Privacy
|
|
136
|
+
|
|
137
|
+
Squishling only sends the model what you name: the method arguments, any `squish_context` values, and any source you
|
|
138
|
+
append. Tools that sit on the request path are a separate route out of your process, and they see both sides:
|
|
139
|
+
|
|
140
|
+
- **Request interceptors, such as `coolhand`,** record the prompts (your purpose text and those inputs) and the
|
|
141
|
+
model's responses. Coolhand sends them to coolhandlabs.com unless you point `config.base_url` at an endpoint of your
|
|
142
|
+
own.
|
|
143
|
+
- **A `squawk` hook that forwards `output:` or `metadata[:input]`** sends model output and your inputs wherever you
|
|
144
|
+
send them. That is the sanctioned route for model output; see
|
|
145
|
+
[Observing every attempt](failures.md#observing-every-attempt-squawk). The examples above log only metadata.
|
|
146
|
+
- **OpenTelemetry** with RubyLLM's tracer records usage and request settings, not message text.
|
|
147
|
+
|
|
148
|
+
Check what your tool stores before pointing it at inputs that contain personal data.
|
data/docs/naming.md
ADDED
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# Naming and collisions
|
|
2
|
+
|
|
3
|
+
`include Squishling` adds a small set of methods to your class. This page lists them and explains what happens
|
|
4
|
+
when one of those names is already taken.
|
|
5
|
+
|
|
6
|
+
## What Squishling adds
|
|
7
|
+
|
|
8
|
+
| Where | Names |
|
|
9
|
+
|---|---|
|
|
10
|
+
| Class-level DSL | `squishling`, `purpose`, `append_to_purpose`, `output_schema`, `squish`, `squish_when`, `squish_fallback`, `squish_validate`, `squish_context`, `squished_methods`, `call`, and the `squishling_*` readers |
|
|
11
|
+
| Instance | `squishling_result`, `squish!`, and `result` (an alias for `squishling_result`) |
|
|
12
|
+
|
|
13
|
+
Methods your class defines itself always win over these, whether you define them before or after the `include`.
|
|
14
|
+
Only methods your class *inherits* can be shadowed.
|
|
15
|
+
|
|
16
|
+
## Inherited collisions raise
|
|
17
|
+
|
|
18
|
+
If a parent class or an earlier-included module already defines a class-level DSL method, `squish!` or
|
|
19
|
+
`squishling_result`, `include Squishling` raises `Squishling::ConfigurationError` listing every clash and where
|
|
20
|
+
it comes from:
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
ConfigurationError: MyApp inherits methods that Squishling would override: MyApp.call (from Sinatra::Base's
|
|
24
|
+
class methods). Include Squishling in a plain Ruby class instead ...
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
This is deliberate. Squishling's modules sit ahead of inherited ones, so without the check
|
|
28
|
+
`Sinatra::Base.call(env)` (the Rack entry point), or any `call` or `purpose` your parent class defines, would be
|
|
29
|
+
replaced silently. Put Squishling in a plain Ruby class and have your framework object call it instead.
|
|
30
|
+
|
|
31
|
+
## `result` is skipped when taken
|
|
32
|
+
|
|
33
|
+
`result` is a convenience for `squishling_result`. If the class already has a `result` (its own or inherited), it is
|
|
34
|
+
left alone and Squishling does not add the alias. Use `squishling_result(...)` instead; it is always available.
|
|
35
|
+
|
|
36
|
+
**ActiveRecord caveat.** Column readers are generated lazily, so Squishling can't see a `result` column when you
|
|
37
|
+
include it. Its `result` would then shadow the column reader. On a model with a `result` column, call
|
|
38
|
+
`squishling_result` (or keep Squishling out of the model and use a plain class).
|
|
39
|
+
|
|
40
|
+
## Not a collision
|
|
41
|
+
|
|
42
|
+
ActiveSupport adds `String#squish` and `String#squish!`. Those live on `String`; Squishling's `squish!` lives on
|
|
43
|
+
your class, so they never meet. They only look alike when reading code in a Rails app.
|
data/docs/routing.md
CHANGED
|
@@ -35,24 +35,27 @@ end
|
|
|
35
35
|
def summarize(text) = raise NotImplementedError # elastic until someone writes it
|
|
36
36
|
```
|
|
37
37
|
|
|
38
|
-
When the Ruby implementation runs,
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
38
|
+
When the Ruby implementation runs, whatever it returns (a `Hash`, `nil`, a string, another schema's result, …)
|
|
39
|
+
is validated against the schema and turned into the same typed result the LLM path produces.
|
|
40
|
+
`squishling_result(...)` does the same explicitly; `result(...)` is a convenience alias, skipped when the class
|
|
41
|
+
already has a `result` (see [Naming and collisions](naming.md)). Invalid deterministic output raises
|
|
42
|
+
`Squishling::InvalidOutputError` too, so a hardened path can't silently drift from the contract.
|
|
42
43
|
|
|
43
44
|
A `NotImplementedError` raised anywhere inside the method, including from code it calls, also routes to the
|
|
44
45
|
LLM.
|
|
45
46
|
|
|
46
|
-
Each call is routed on its own, including a squished method that calls itself on smaller inputs
|
|
47
|
+
Each call is routed on its own, including a squished method that calls itself on smaller inputs (those
|
|
48
|
+
recursive calls must return schema-valid values too). A subclass
|
|
47
49
|
override that calls `super` is one call: it's routed once, at the subclass. A call to the method from its own
|
|
48
|
-
`squish_when` or `squish_fallback` (or
|
|
49
|
-
implementation, so a fallback can hand the input back to Ruby with `call(**inputs)`.
|
|
50
|
+
`squish_when` or `squish_fallback` (or a purpose proc) isn't routed again: it runs the Ruby
|
|
51
|
+
implementation, so a fallback can hand the input back to Ruby with `call(**inputs)`. These inner calls
|
|
52
|
+
return a `Hash` as the typed result and any other value unchanged; only the outermost return is validated in full.
|
|
50
53
|
|
|
51
54
|
## Hardening a path
|
|
52
55
|
|
|
53
56
|
This is the workflow from [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai):
|
|
54
57
|
|
|
55
|
-
1. Ship the class with
|
|
58
|
+
1. Ship the class with purpose and a schema but no implementation. Every call goes to the LLM.
|
|
56
59
|
2. Watch which inputs carry the volume. `squished?` on each result tells you which path served it.
|
|
57
60
|
3. Write Ruby for the high-volume cases and narrow `squish_when` so only the rest go to the LLM.
|
|
58
61
|
|
|
@@ -68,7 +71,7 @@ class TicketTriager
|
|
|
68
71
|
|
|
69
72
|
squish_context :customer_tier, :product # instance state sent alongside the arguments
|
|
70
73
|
|
|
71
|
-
squish :triage,
|
|
74
|
+
squish :triage, purpose: "Assign a priority and team.", model: "claude-haiku-4-5" do
|
|
72
75
|
string :priority, enum: %w[low med high]
|
|
73
76
|
string :team
|
|
74
77
|
end
|
|
@@ -83,10 +86,11 @@ end
|
|
|
83
86
|
|
|
84
87
|
`squish` accepts:
|
|
85
88
|
|
|
86
|
-
- `
|
|
87
|
-
- `
|
|
89
|
+
- `purpose:` (replaces the class's)
|
|
90
|
+
- `append_to_purpose:` (added to the class's, see [Appending to the purpose](#appending-to-the-purpose))
|
|
88
91
|
- `output_schema:` (or a schema block)
|
|
89
92
|
- `model:` or `escalation:`, `provider:`, and `params:` (generation params; see [Configuration](configuration.md))
|
|
93
|
+
- `harness:` (see [Harnesses](harnesses.md))
|
|
90
94
|
- `when:`, a predicate proc
|
|
91
95
|
- `validate:` and `fallback:` (see [Failure handling](failures.md))
|
|
92
96
|
|
|
@@ -103,8 +107,8 @@ from the method.
|
|
|
103
107
|
class InvoiceParser
|
|
104
108
|
include Squishling
|
|
105
109
|
|
|
106
|
-
|
|
107
|
-
|
|
110
|
+
purpose "Extract invoice fields from the client's raw data."
|
|
111
|
+
append_to_purpose "Here is the Ruby that parses well-formed invoices, for context on the logic and goals:",
|
|
108
112
|
self
|
|
109
113
|
output_schema do
|
|
110
114
|
string :invoice_number
|
|
@@ -115,7 +119,7 @@ class InvoiceParser
|
|
|
115
119
|
parsed = AcmeParser.parse(data)
|
|
116
120
|
result(invoice_number: parsed.id, total: parsed.sum)
|
|
117
121
|
rescue AcmeParser::ParseError => e
|
|
118
|
-
squish!(
|
|
122
|
+
squish!(append_to_purpose: "The Ruby parser above failed on this input; the error is in the context.",
|
|
119
123
|
context: { parse_error: e })
|
|
120
124
|
end
|
|
121
125
|
end
|
|
@@ -126,9 +130,10 @@ end
|
|
|
126
130
|
| Option | Effect |
|
|
127
131
|
|---|---|
|
|
128
132
|
| `context:` | A Hash sent under `"context"` with any `squish_context` values (a same-named key wins). Exceptions are sent as `{ "class", "message" }`, never their backtrace. |
|
|
129
|
-
| `
|
|
130
|
-
| `
|
|
133
|
+
| `append_to_purpose:` | Added to the declared sections; `false` (alone or first in an Array) drops them for this call |
|
|
134
|
+
| `purpose:` | Replaces the purpose |
|
|
131
135
|
| `model:` or `escalation:`, `provider:`, `params:` | E.g. send this call to a stronger model, or a whole [escalation](configuration.md#models-and-escalation), when Ruby fails. A `provider:` needs a `model:` or `escalation:`; `params:` merge key by key over the declared ones. |
|
|
136
|
+
| `harness:` | Replaces the declared [harness](harnesses.md), e.g. `harness: :judged_squishsum` to double-check a hand-off |
|
|
132
137
|
|
|
133
138
|
- **The output schema can't be overridden.** The call still returns the method's result type.
|
|
134
139
|
- **Failures** go through the normal LLM path: a declared `squish_fallback` is used, otherwise
|
|
@@ -148,9 +153,9 @@ end
|
|
|
148
153
|
or in a parent implementation reached through `super`. Calling it anywhere else, or from a `squish_fallback`
|
|
149
154
|
(which would loop), raises `Squishling::Error`.
|
|
150
155
|
|
|
151
|
-
## Appending to the
|
|
156
|
+
## Appending to the purpose
|
|
152
157
|
|
|
153
|
-
`
|
|
158
|
+
`append_to_purpose` adds sections to the system prompt after the purpose. It takes items, an Array of
|
|
154
159
|
them, or a block (treated as a Proc item). Each item is one of:
|
|
155
160
|
|
|
156
161
|
| Item | Sent as |
|
|
@@ -164,19 +169,19 @@ them, or a block (treated as a Proc item). Each item is one of:
|
|
|
164
169
|
class InvoiceParser
|
|
165
170
|
include Squishling
|
|
166
171
|
|
|
167
|
-
|
|
168
|
-
|
|
172
|
+
append_to_purpose "The Ruby that parses well-formed invoices:", self, AcmeParser
|
|
173
|
+
append_to_purpose -> { "This client's invoices are in #{currency}." }
|
|
169
174
|
|
|
170
|
-
squish :summarize,
|
|
175
|
+
squish :summarize, append_to_purpose: false do # no appendices for this method
|
|
171
176
|
string :summary
|
|
172
177
|
end
|
|
173
178
|
end
|
|
174
179
|
```
|
|
175
180
|
|
|
176
|
-
The same option goes in the class-wide call: `squishling
|
|
181
|
+
The same option goes in the class-wide call: `squishling append_to_purpose: ["...", self]`.
|
|
177
182
|
|
|
178
|
-
Sections are added down the chain: class, then subclass, then `squish :name,
|
|
179
|
-
`squish!(
|
|
183
|
+
Sections are added down the chain: class, then subclass, then `squish :name, append_to_purpose:`, then
|
|
184
|
+
`squish!(append_to_purpose:)`. `false` drops everything declared above it, so `[false, "Only this."]`
|
|
180
185
|
replaces it.
|
|
181
186
|
|
|
182
187
|
Source is read with Ruby's own parser (Prism) the first time it's needed and cached. Some limits:
|
|
@@ -195,8 +200,8 @@ Source is read with Ruby's own parser (Prism) the first time it's needed and cac
|
|
|
195
200
|
|
|
196
201
|
## What the LLM sees
|
|
197
202
|
|
|
198
|
-
- **System prompt:** your
|
|
199
|
-
`
|
|
203
|
+
- **System prompt:** your purpose (a String, or a Proc evaluated against the instance), then any
|
|
204
|
+
`append_to_purpose` sections, then a short note describing the input format.
|
|
200
205
|
- **User message:** JSON with the method's arguments, mapped to their parameter names. Any
|
|
201
206
|
`squish_context` values, and a `squish!` call's `context:`, go under `"context"`:
|
|
202
207
|
|
|
@@ -205,11 +210,15 @@ Source is read with Ruby's own parser (Prism) the first time it's needed and cac
|
|
|
205
210
|
"context": { "customer_tier": "enterprise", "product": "API" } }
|
|
206
211
|
```
|
|
207
212
|
|
|
213
|
+
- **Squishsum and ensemble harnesses:** both samples get exactly this input, and a [judge](harnesses.md#the-judge) gets
|
|
214
|
+
the same purpose and input plus both samples' outputs.
|
|
208
215
|
- **Retries and escalation:** another attempt of the same step gets the validation errors in the same conversation. A
|
|
209
216
|
later [escalation](configuration.md#models-and-escalation) step, which may be a different provider (say, local
|
|
210
217
|
Ollama, then hosted Anthropic Claude), gets the same JSON plus the previous model's rejected output and the
|
|
211
218
|
errors, including any messages your [`squish_validate`](failures.md#output-checks-squish_validate) check
|
|
212
|
-
returned.
|
|
219
|
+
returned. The rejected output is model-generated and is truncated to 4,000 characters; set
|
|
220
|
+
`forward_rejected: false` on a step to start it from the original input only, without the rejected output or
|
|
221
|
+
its errors. Don't put data in those messages that you wouldn't send as an argument.
|
|
213
222
|
|
|
214
223
|
`squish_context` names are read from a method of that name if there is one, otherwise from the instance
|
|
215
224
|
variable. Only context you name is sent. Instance variables are never dumped wholesale, so API clients,
|
data/docs/schemas.md
CHANGED
|
@@ -59,7 +59,8 @@ r.squished? # => true when it came from the LLM
|
|
|
59
59
|
```
|
|
60
60
|
|
|
61
61
|
On the deterministic path, return a `Hash` or build the result with `result(...)`. Either way it's validated
|
|
62
|
-
against the schema
|
|
62
|
+
against the schema, and so is anything else you return (`nil`, a string, another schema's result), which raises
|
|
63
|
+
`Squishling::InvalidOutputError` unless it matches.
|
|
63
64
|
|
|
64
65
|
## Optional vs. empty
|
|
65
66
|
|
|
@@ -84,7 +85,7 @@ end
|
|
|
84
85
|
branch as usual: `vitals` is a `Data` object or `nil`, and a list of objects is a list of `Data` objects. In a raw
|
|
85
86
|
JSON Schema, `type: ["array", "null"]` works too. A union with more than one non-null branch is ambiguous, so its
|
|
86
87
|
values come back as plain hashes. If you need the model to tell "none" apart from "not mentioned", say so in your
|
|
87
|
-
|
|
88
|
+
purpose.
|
|
88
89
|
|
|
89
90
|
## Contracts beyond the schema
|
|
90
91
|
|
|
@@ -110,7 +111,7 @@ Providers' strict modes don't support conditional keywords (`if`/`then`/`else` f
|
|
|
110
111
|
`dependentSchemas` from `dependent`), so Squishling keeps them out of the schema it sends and enforces them itself.
|
|
111
112
|
Output that breaks a rule is rejected like any other invalid output: the call moves on to the next attempt in its
|
|
112
113
|
[escalation](configuration.md#models-and-escalation) with the errors. The model doesn't see these rules
|
|
113
|
-
up front, so state them in your
|
|
114
|
+
up front, so state them in your purpose too.
|
|
114
115
|
|
|
115
116
|
For anything a schema can't express (sums, lookups against your data), use
|
|
116
117
|
[`squish_validate`](failures.md#output-checks-squish_validate).
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
module Squishling
|
|
4
|
-
#
|
|
4
|
+
# append_to_purpose: extra system-prompt sections placed after the purpose, layered
|
|
5
5
|
# class -> subclass -> method -> call. Each level adds to the levels above it; `false` drops them.
|
|
6
6
|
module Appendices
|
|
7
7
|
ITEM_TYPES = [String, Module, Method, UnboundMethod, Proc].freeze
|
|
@@ -29,7 +29,7 @@ module Squishling
|
|
|
29
29
|
|
|
30
30
|
values = receiver.instance_exec(&item)
|
|
31
31
|
(values.is_a?(Array) ? values : [values]).select(&:itself).map do |value|
|
|
32
|
-
check!(value, "#{label}
|
|
32
|
+
check!(value, "#{label} append_to_purpose proc", procs: false)
|
|
33
33
|
render_item(value)
|
|
34
34
|
end
|
|
35
35
|
end
|
|
@@ -43,7 +43,7 @@ module Squishling
|
|
|
43
43
|
allowed = procs ? ITEM_TYPES : ITEM_TYPES - [Proc]
|
|
44
44
|
return if allowed.any? { |type| item.is_a?(type) }
|
|
45
45
|
|
|
46
|
-
raise ConfigurationError, "#{label}:
|
|
46
|
+
raise ConfigurationError, "#{label}: append_to_purpose items must be Strings, classes or modules, " \
|
|
47
47
|
"#{procs ? 'methods, or procs' : 'or methods'} (got #{describe(item)})"
|
|
48
48
|
end
|
|
49
49
|
|