squishling 0.1.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +123 -0
- data/README.md +117 -35
- data/docs/configuration.md +113 -32
- data/docs/failures.md +111 -14
- data/docs/harnesses.md +198 -0
- data/docs/measuring-tokens.md +148 -0
- data/docs/naming.md +43 -0
- data/docs/routing.md +46 -31
- data/docs/schemas.md +32 -2
- data/lib/squishling/appendices.rb +3 -3
- data/lib/squishling/class_methods.rb +81 -31
- data/lib/squishling/collisions.rb +54 -0
- data/lib/squishling/configuration.rb +58 -9
- data/lib/squishling/definition.rb +78 -39
- data/lib/squishling/errors.rb +29 -5
- data/lib/squishling/harness.rb +153 -0
- data/lib/squishling/invoker.rb +223 -74
- data/lib/squishling/judge.rb +142 -0
- data/lib/squishling/llm_client.rb +150 -0
- data/lib/squishling/model_path.rb +129 -0
- data/lib/squishling/output_check.rb +109 -0
- data/lib/squishling/params.rb +39 -2
- data/lib/squishling/router.rb +13 -7
- data/lib/squishling/schema.rb +45 -3
- data/lib/squishling/source.rb +5 -5
- data/lib/squishling/squawk.rb +52 -0
- data/lib/squishling/version.rb +1 -1
- data/lib/squishling.rb +25 -5
- metadata +13 -3
data/docs/failures.md
CHANGED
|
@@ -2,34 +2,131 @@
|
|
|
2
2
|
|
|
3
3
|
## When the LLM fails
|
|
4
4
|
|
|
5
|
+
Every squished call works through its [escalation](configuration.md#models-and-escalation), one attempt at a time.
|
|
6
|
+
A failed attempt moves on to the next one; once the last fails, the error is raised (or handed to the
|
|
7
|
+
[fallback](#fallbacks)). A plain `model:` makes a single attempt.
|
|
8
|
+
|
|
5
9
|
| Failure | What Squishling does |
|
|
6
10
|
|---|---|
|
|
7
|
-
| Rate limit, 5xx, overload, timeout, connection error | RubyLLM retries at the HTTP level (`RubyLLM.config.max_retries`, default 3). If it still fails, raises `Squishling::LLMError
|
|
8
|
-
| Bad API key, unknown model, missing provider config | Raises `Squishling::ConfigurationError`. Never
|
|
9
|
-
| Provider rejects the request (400 Bad Request), e.g. an unsupported `temperature` or a schema it won't accept | Raises `ConfigurationError` with the provider's message and the params in use. Never
|
|
10
|
-
| Empty or `nil` response (e.g. a refusal or a max-tokens cutoff) |
|
|
11
|
-
| Malformed or truncated JSON |
|
|
12
|
-
| JSON that doesn't match the schema (wrong types, missing or extra keys, root not an object) |
|
|
11
|
+
| Rate limit, 5xx, overload, timeout, connection error | RubyLLM retries at the HTTP level first (`RubyLLM.config.max_retries`, default 3). If it still fails, moves to the next attempt in a fresh chat; on the last attempt, raises `Squishling::LLMError` with the original exception as its `cause`. |
|
|
12
|
+
| Bad API key, unknown model, missing provider config | Raises `Squishling::ConfigurationError`. Never escalated, never sent to the fallback. |
|
|
13
|
+
| Provider rejects the request (400 Bad Request), e.g. an unsupported `temperature` or a schema it won't accept | Raises `ConfigurationError` with the provider's message and the params in use. Never escalated, never sent to the fallback, so a setup mistake can't be silently covered up on every call. |
|
|
14
|
+
| Empty or `nil` response (e.g. a refusal or a max-tokens cutoff) | Moves to the next attempt, then raises `Squishling::InvalidOutputError` |
|
|
15
|
+
| Malformed or truncated JSON | Moves to the next attempt, then raises `InvalidOutputError`. JSON wrapped in a markdown code fence is accepted. |
|
|
16
|
+
| JSON that doesn't match the schema (wrong types, missing or extra keys, `null` in a non-`optional` field, a broken [conditional rule](schemas.md#contracts-beyond-the-schema), root not an object) | Moves to the next attempt with the validation errors, then raises `InvalidOutputError` |
|
|
17
|
+
| Schema-valid output that [`squish_validate`](#output-checks-squish_validate) rejects | Moves to the next attempt with your messages, then raises `InvalidOutputError` |
|
|
18
|
+
| With a [squishsum or ensemble harness](harnesses.md): two valid samples that differ, with no judge or a judge that rejects both | Raises `Squishling::DisagreementError` (an `InvalidOutputError`), carrying both typed `candidates` |
|
|
13
19
|
|
|
14
20
|
Squishling validates output itself with [json_schemer](https://github.com/davishmcclurg/json_schemer), because
|
|
15
|
-
RubyLLM doesn't, and some providers don't enforce strict mode.
|
|
16
|
-
|
|
17
|
-
|
|
21
|
+
RubyLLM doesn't, and some providers don't enforce strict mode. When the next attempt is on the same step (same
|
|
22
|
+
model, provider, params, and `forward_rejected:`), the re-ask happens in the same conversation, so the model sees
|
|
23
|
+
what it got wrong. A different step gets a fresh chat with the original input plus the rejected output (truncated to
|
|
24
|
+
4,000 characters) and its errors, so that output is sent to the next step's provider, which may not be the one that
|
|
25
|
+
produced it. Set `forward_rejected: false` on a step to leave both out. `InvalidOutputError` exposes `errors`, `raw`
|
|
26
|
+
(the last response), `attempts`, and `models` (the model tried on each attempt); for a
|
|
27
|
+
[`DisagreementError`](harnesses.md#when-the-harness-fails), `raw` holds both candidates, `attempts` is `nil`, and
|
|
28
|
+
`models` has one entry per role (sample a, sample b, then the judge). With a `logger` configured, every
|
|
29
|
+
escalation is logged as a warning.
|
|
30
|
+
|
|
31
|
+
### What ends up in errors and logs
|
|
32
|
+
|
|
33
|
+
Model output can echo sensitive input, so the raw response is kept in one place: `InvalidOutputError#raw` (plus the opt-in `squawk` hook below).
|
|
34
|
+
`InvalidOutputError#message`, `#errors`, and `config.logger` warnings are built from these:
|
|
35
|
+
|
|
36
|
+
- For unparseable JSON, only the position (`response was not valid JSON (at line 1 column 7)`), never the parser's
|
|
37
|
+
snippet of the response.
|
|
38
|
+
- JSON Schema errors, which name the schema path and the rule that failed. They never contain the rejected value, but
|
|
39
|
+
they do name an extra key the model added (`object property at /foo is a disallowed additional property`), and
|
|
40
|
+
that key name is chosen by the model.
|
|
41
|
+
- Your own `squish_validate` messages, verbatim. If you interpolate result values into them, those values reach
|
|
42
|
+
the message, the logger, and the next attempt's prompt.
|
|
43
|
+
- A fixed message when a response contains a value JSON can't represent (`1e400` parses to `Infinity`, or a string
|
|
44
|
+
with invalid UTF-8).
|
|
45
|
+
- At most the first 20 of these errors, plus a count of the rest, so a very large malformed response can't flood
|
|
46
|
+
the retry message or the log.
|
|
47
|
+
- For a [`DisagreementError`](harnesses.md#when-the-harness-fails), only what Squishling wrote: a `:judgment` judge's
|
|
48
|
+
choice and probability. A chat judge's `reason` is model-written, so it is available as `#reason` but never in the
|
|
49
|
+
message or logs.
|
|
50
|
+
|
|
51
|
+
The same applies to anything you log yourself, such as `error.message` in a [fallback](#fallbacks). Use `error.raw`
|
|
52
|
+
only where you are willing to store model output. To send the output of every attempt somewhere on purpose, use
|
|
53
|
+
[`squawk`](#observing-every-attempt-squawk).
|
|
54
|
+
|
|
55
|
+
## Observing every attempt (`squawk`)
|
|
56
|
+
|
|
57
|
+
`squawk` is an opt-in hook for sending what the model returned to an error tracker or tracing tool. It runs after
|
|
58
|
+
every LLM attempt, accepted or rejected, and does nothing unless you set it:
|
|
59
|
+
|
|
60
|
+
```ruby
|
|
61
|
+
Squishling.configure do |config|
|
|
62
|
+
config.squawk = lambda do |output:, metadata:, error:|
|
|
63
|
+
ObservabilitySolution.record(output, metadata, error) if error
|
|
64
|
+
end
|
|
65
|
+
end
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
| Keyword | Value |
|
|
69
|
+
|---|---|
|
|
70
|
+
| `output` | The raw response: a Hash or String, or `nil` when the provider call itself failed |
|
|
71
|
+
| `error` | `nil` for an accepted attempt. Otherwise the `InvalidOutputError` (with `errors` and `raw`) or `LLMError` that ended the attempt |
|
|
72
|
+
| `metadata` | `label` (`"Class#method"`), `attempt`, `attempts` (the escalation's length; under a harness, the sample's or judge's own attempts), `final` (the last attempt), `model`, `provider`, `params`, `input` (the JSON sent: arguments and named context only), `usage` (token counts, when RubyLLM reports them) |
|
|
73
|
+
|
|
74
|
+
- Set it globally with `config.squawk`, per class with `squishling squawk: ...`, or per method with
|
|
75
|
+
`squish :triage, squawk: ...`. The method's hook wins over the class's, which wins over the configured one;
|
|
76
|
+
`squawk: false` silences an inherited hook. Subclasses and `squish!` calls inherit it.
|
|
77
|
+
- Any object that responds to `call` works. The hook receives only the keywords it declares, unless it takes `**`,
|
|
78
|
+
so a lambda that wants just `error:` is fine and new metadata fields won't break it.
|
|
79
|
+
- It runs inline, so keep it quick. Exceptions it raises propagate unchanged and fail the call, like any of your own
|
|
80
|
+
code, so rescue inside the hook if an outage in your tracing tool shouldn't.
|
|
81
|
+
- `output` is the same object the result is built from, so treat it as read-only.
|
|
82
|
+
- Under a [harness](harnesses.md), it runs for every sample and judge attempt too, each with its own `input`.
|
|
83
|
+
- It isn't called for a `ConfigurationError`, for deterministic and fallback returns, or for an attempt where your own
|
|
84
|
+
`squish_validate` raises (that exception propagates first).
|
|
85
|
+
|
|
86
|
+
## Output checks (`squish_validate`)
|
|
87
|
+
|
|
88
|
+
The schema covers shape and types. For rules it can't express, such as cross-field arithmetic or a lookup against
|
|
89
|
+
your data, add a Ruby check. It runs only on LLM output that already matches the schema, receives the typed result
|
|
90
|
+
plus the method's inputs as keywords, and runs against the instance:
|
|
91
|
+
|
|
92
|
+
```ruby
|
|
93
|
+
class InvoiceParser
|
|
94
|
+
include Squishling
|
|
95
|
+
# ...
|
|
96
|
+
squish_validate do |result, **|
|
|
97
|
+
errors = []
|
|
98
|
+
errors << "total must equal the sum of line_items" unless result.total == result.line_items.sum(&:amount)
|
|
99
|
+
errors << "unknown currency" unless Currency.supported?(result.currency)
|
|
100
|
+
errors
|
|
101
|
+
end
|
|
102
|
+
end
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
- Return `nil`, `true`, `""`, or `[]` to accept the output.
|
|
106
|
+
- Return a message or an array of messages to reject it. The messages are sent to the next attempt and end up in
|
|
107
|
+
`InvalidOutputError#errors` if every attempt fails. `false` rejects with a generic message.
|
|
108
|
+
- A [dry-validation](https://dry-rb.org/gems/dry-validation/) result works too, if your app already uses it:
|
|
109
|
+
`squish_validate { |result, **| InvoiceContract.new.call(result.to_h) }`. Squishling doesn't depend on dry-rb.
|
|
110
|
+
- Per method: `squish :triage, validate: ->(result, **inputs) { ... }`. Subclasses inherit the class-level check.
|
|
111
|
+
- It doesn't run on deterministic or fallback returns: those are your own code.
|
|
112
|
+
- Exceptions raised inside it propagate unwrapped, like any of your own code.
|
|
18
113
|
|
|
19
114
|
## Errors
|
|
20
115
|
|
|
21
116
|
| Error | Raised when |
|
|
22
117
|
|---|---|
|
|
23
118
|
| `Squishling::Error` | Base class for everything below. Raised directly for misuse at call time, e.g. `result` or `squish!` called outside a squished method |
|
|
24
|
-
| `Squishling::ConfigurationError` | Missing
|
|
25
|
-
| `Squishling::InvalidOutputError` | LLM output still invalid
|
|
26
|
-
| `Squishling::
|
|
119
|
+
| `Squishling::ConfigurationError` | Missing purpose or schema, a non-strict schema, invalid or reserved params, an invalid `model:`/`escalation:` declaration, an invalid `append_to_purpose` item or unavailable source, bad credentials, an unknown model, a request the provider rejects (400) |
|
|
120
|
+
| `Squishling::InvalidOutputError` | LLM output still invalid (schema or `squish_validate`) after every attempt in the escalation, or a deterministic/fallback return that doesn't match the schema |
|
|
121
|
+
| `Squishling::DisagreementError` | A subclass of `InvalidOutputError`: a [squishsum or ensemble harness](harnesses.md#when-the-harness-fails)'s samples disagreed and no judge accepted either. `candidates`, `verdict`, and `reason` describe what happened |
|
|
122
|
+
| `Squishling::LLMError` | The provider call failed on the last attempt in the escalation, after RubyLLM's own retries, including context-length errors (`cause` holds the original) |
|
|
27
123
|
|
|
28
124
|
Errors raised by your own Ruby code are not wrapped.
|
|
29
125
|
|
|
30
126
|
## Fallbacks
|
|
31
127
|
|
|
32
|
-
Use `squish_fallback` to decide what happens when the LLM path fails with `InvalidOutputError`
|
|
128
|
+
Use `squish_fallback` to decide what happens when the LLM path fails with `InvalidOutputError` (including a
|
|
129
|
+
`DisagreementError`) or `LLMError`. It's the one fallback for every [harness](harnesses.md).
|
|
33
130
|
It receives the error plus the method's inputs as keywords, and runs against the instance:
|
|
34
131
|
|
|
35
132
|
```ruby
|
|
@@ -49,7 +146,7 @@ end
|
|
|
49
146
|
- Per method: `squish :triage, fallback: ->(error, **inputs) { ... }`.
|
|
50
147
|
- Subclasses inherit the class-level fallback.
|
|
51
148
|
- Calls handed to the LLM with `squish!` use the fallback too. A fallback can't call `squish!` itself; that
|
|
52
|
-
raises `Squishling::Error` rather than looping. See [
|
|
149
|
+
raises `Squishling::Error` rather than looping. See [Handing off to the LLM](routing.md#handing-off-to-the-llm-with-squish).
|
|
53
150
|
- Without a fallback, the error propagates.
|
|
54
151
|
|
|
55
152
|
Routing to Ruby code isn't automatic on failure. The predicate already chose the LLM for this input, so the
|
data/docs/harnesses.md
ADDED
|
@@ -0,0 +1,198 @@
|
|
|
1
|
+
# Harnesses: escalation, squishsum, and ensemble
|
|
2
|
+
|
|
3
|
+
A squished call's **harness** decides how it uses its [escalation](configuration.md#models-and-escalation):
|
|
4
|
+
|
|
5
|
+
| Harness | What it does | Requests per call |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| `:escalation` (default) | Tries each attempt in order until an output passes the schema and `squish_validate`. | 1 or more |
|
|
8
|
+
| `:squishsum` | Asks the escalation's **first step** twice, concurrently. Identical outputs are accepted; different ones raise `Squishling::DisagreementError`. | 2 or more |
|
|
9
|
+
| `:judged_squishsum` | Like `:squishsum`, but when the two outputs differ, a **judge** picks one or rejects both. | 2 or more, plus 1 or more judge requests when they differ |
|
|
10
|
+
| `:ensemble` | Asks the escalation's **first and second steps** once each, concurrently, so two different models are compared. Identical outputs are accepted; different ones raise `DisagreementError`. | 2 or more |
|
|
11
|
+
| `:judged_ensemble` | Like `:ensemble`, but when the two outputs differ, a **judge** picks one or rejects both. | 2 or more, plus 1 or more judge requests when they differ |
|
|
12
|
+
|
|
13
|
+
Use a squishsum harness when a confidently wrong answer costs more than a second request: one sample can make a
|
|
14
|
+
mistake, but two independent samples rarely make the same one. Use an [ensemble](#ensembles) when the two samples
|
|
15
|
+
should come from different models: a model rarely disagrees with itself at low temperature, but a cheap and a
|
|
16
|
+
stronger model often disagree on exactly the cases the cheap one gets confidently wrong.
|
|
17
|
+
|
|
18
|
+
```ruby
|
|
19
|
+
class TicketTriager
|
|
20
|
+
include Squishling
|
|
21
|
+
|
|
22
|
+
squishling escalation: [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5"],
|
|
23
|
+
harness: :judged_squishsum # two Haiku samples; Sonnet judges when they differ
|
|
24
|
+
purpose "Assign a priority and a team."
|
|
25
|
+
output_schema do
|
|
26
|
+
string :priority, enum: %w[low medium high]
|
|
27
|
+
string :team
|
|
28
|
+
end
|
|
29
|
+
end
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## Declaring a harness
|
|
33
|
+
|
|
34
|
+
`harness:` takes a type, or a Hash with `type:` and its options:
|
|
35
|
+
|
|
36
|
+
```ruby
|
|
37
|
+
Squishling.configure { |c| c.default_harness = :squishsum } # every class without its own
|
|
38
|
+
|
|
39
|
+
squishling harness: :judged_squishsum # this class (and subclasses)
|
|
40
|
+
squish :triage, harness: { type: :squishsum, # one method
|
|
41
|
+
compare: ->(a, b, **) { a.priority == b.priority } }
|
|
42
|
+
squish!(harness: :escalation) # one call (see routing.md)
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
The first level that declares one wins: `squish!`, then the method, then the class, then
|
|
46
|
+
`Squishling.config.default_harness`, then `:escalation`. Options aren't merged across levels.
|
|
47
|
+
|
|
48
|
+
| Option | Harnesses | Description |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| `type:` | all | `:escalation`, `:squishsum`, `:judged_squishsum`, `:ensemble`, or `:judged_ensemble` (required in the Hash form). |
|
|
51
|
+
| `compare:` | squishsum and ensemble ones | `->(a, b, **inputs) { ... }`, evaluated against the instance with the two typed results and the method's inputs. Truthy means the samples agree. Default: the two outputs are exactly equal. |
|
|
52
|
+
| `judge:` | judged ones | The judge step (see [The judge](#the-judge)). Default: the escalation's next step after the sampled ones. |
|
|
53
|
+
| `judge_instructions:` | judged ones | A String, or a Proc evaluated against the instance, that replaces the [default judge prompt](#the-default-judge-prompt). |
|
|
54
|
+
|
|
55
|
+
## Samples
|
|
56
|
+
|
|
57
|
+
Under `:squishsum`, both samples run on the escalation's first step, with its model, provider, and params
|
|
58
|
+
(an [ensemble](#ensembles) runs them on its first and second steps):
|
|
59
|
+
|
|
60
|
+
- **They're independent.** Each sample has its own chat, and the two requests are sent concurrently on separate
|
|
61
|
+
threads. Only the provider requests run on those threads. Parsing, schema validation, `squish_validate`,
|
|
62
|
+
`compare:`, and the judge all run on the calling thread, so your Squishling callbacks never run concurrently.
|
|
63
|
+
RubyLLM instrumentation subscribers (e.g. `ActiveSupport::Notifications` listeners for `chat.ruby_llm`) do see
|
|
64
|
+
each sample's request on its worker thread, outside the Rails executor, so keep them thread-safe.
|
|
65
|
+
- **Each one is retried like an escalation step.** A sample with invalid output is re-asked in its own chat, up to
|
|
66
|
+
their step's `attempts:`, while a sample that already passed keeps its result. A failed request gets a fresh
|
|
67
|
+
chat. A sample still failing after its attempts fails the call with `InvalidOutputError` or `LLMError`, as
|
|
68
|
+
usual. Under `:squishsum`, later escalation steps are never used for samples (see [Ensembles](#ensembles)).
|
|
69
|
+
- **Agreement is exact by default.** Free-text fields seldom match word for word, so pass `compare:` to decide
|
|
70
|
+
which fields must match:
|
|
71
|
+
|
|
72
|
+
```ruby
|
|
73
|
+
squishling harness: { type: :squishsum, compare: ->(a, b, **) { a.priority == b.priority && a.team == b.team } }
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
When the samples agree, the first one is returned.
|
|
77
|
+
|
|
78
|
+
## Ensembles
|
|
79
|
+
|
|
80
|
+
`:ensemble` and `:judged_ensemble` sample two *different* steps instead of one step twice. Sample `a` is the
|
|
81
|
+
escalation's first step and sample `b` its second, each with its own model, provider, and params:
|
|
82
|
+
|
|
83
|
+
```ruby
|
|
84
|
+
class TicketTriager
|
|
85
|
+
include Squishling
|
|
86
|
+
|
|
87
|
+
squishling escalation: [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5", "claude-opus-5-5"],
|
|
88
|
+
harness: :judged_ensemble # a = Haiku, b = Sonnet, Opus judges when they differ
|
|
89
|
+
# purpose, output_schema, ...
|
|
90
|
+
end
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
- A step's `attempts:` are that sample's retries (here Haiku gets two attempts, Sonnet one), exactly as in
|
|
94
|
+
[Samples](#samples). Steps after the second are never sampled.
|
|
95
|
+
- Everything else matches squishsum: independence and concurrency, `compare:`, `DisagreementError`, and the
|
|
96
|
+
judge. `DisagreementError#models` names the model behind each sample.
|
|
97
|
+
- The judge defaults to the escalation's **third** step. With only two steps, a `:judged_ensemble` needs `judge:`.
|
|
98
|
+
An ensemble with fewer than two steps, or a judged one with neither a third step nor `judge:`, raises
|
|
99
|
+
`ConfigurationError` before any request.
|
|
100
|
+
- A global `default_harness = :ensemble` therefore needs every class to declare at least two steps, and a
|
|
101
|
+
`squish!(model: ...)` override (one step) can't be combined with an ensemble harness.
|
|
102
|
+
- Steps are told apart by value, so two adjacent steps with the same model, provider, and params count as one step.
|
|
103
|
+
|
|
104
|
+
## The judge
|
|
105
|
+
|
|
106
|
+
With a judged harness, two valid samples that disagree go to a judge. The judge picks candidate `a` or `b`,
|
|
107
|
+
whose typed result is returned (`squished?` is `true`), or it picks `neither`, which raises
|
|
108
|
+
`DisagreementError`.
|
|
109
|
+
|
|
110
|
+
The judge is:
|
|
111
|
+
|
|
112
|
+
1. the `judge:` step, if declared: a model name, or a Hash with `model:`, `provider:`, `params:`, and `attempts:`
|
|
113
|
+
(as in an escalation step, without `order:`), plus `type:`; or
|
|
114
|
+
2. the escalation's next step after the sampled ones (the second under `:judged_squishsum`, the third under
|
|
115
|
+
`:judged_ensemble`), with its attempts.
|
|
116
|
+
|
|
117
|
+
A judged harness with neither raises `ConfigurationError` before any sample is requested. A `judge:`
|
|
118
|
+
step uses only its own `provider:`; it doesn't inherit the class or method `provider:`, which belongs to their models
|
|
119
|
+
(a chat judge's params, though, do merge over the method's).
|
|
120
|
+
|
|
121
|
+
### Chat judges
|
|
122
|
+
|
|
123
|
+
By default the judge is a chat model, asked for a strict `{ "verdict": "a" | "b" | "neither", "reason": "..." }`.
|
|
124
|
+
It gets the same context the samples had, plus the samples themselves:
|
|
125
|
+
|
|
126
|
+
- **System prompt:** the judge prompt, then the operation's own purpose (including any `append_to_purpose`
|
|
127
|
+
sections).
|
|
128
|
+
- **User message:** JSON with the samples' `input` (`arguments` and `context`), the `output_schema`, and the
|
|
129
|
+
`candidates` `a` and `b`.
|
|
130
|
+
|
|
131
|
+
A verdict that doesn't match its schema is retried per the judge step's `attempts:`; a judge that never returns a
|
|
132
|
+
valid verdict raises `InvalidOutputError`. A chat judge's params are the method's
|
|
133
|
+
[generation params](configuration.md#generation-params) with its own `params:` merged over them.
|
|
134
|
+
|
|
135
|
+
### System One judgment models (Jev)
|
|
136
|
+
|
|
137
|
+
A judge can also be a System One decision model, such as TypeSafe's Jev, via
|
|
138
|
+
[RubyLLM judgments](https://rubyllm.com/judgments/) (RubyLLM 2.1+). Decision models return calibrated
|
|
139
|
+
probabilities rather than text, so they are much faster and cheaper than a chat judge.
|
|
140
|
+
|
|
141
|
+
```ruby
|
|
142
|
+
squishling harness: {
|
|
143
|
+
type: :judged_squishsum,
|
|
144
|
+
judge: { model: "jev-latest", type: :judgment, min_confidence: 0.8 }
|
|
145
|
+
}
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
Squishling asks one `choice` question (`a`, `b`, or `neither`), with the judge prompt as its instructions and the
|
|
149
|
+
operation's purpose, input, output schema, and candidates as the judgment input. A pick of `a` or `b` is
|
|
150
|
+
accepted only when its probability is at least `min_confidence:` (default `0.8`). Otherwise, as with `neither`,
|
|
151
|
+
`DisagreementError` is raised with the choice and its probability in its `reason`. A judgment judge sends only its own
|
|
152
|
+
`params:` (as RubyLLM `provider_options:`). `attempts:` retries a failed request. Configure the provider's
|
|
153
|
+
credentials in RubyLLM as usual.
|
|
154
|
+
|
|
155
|
+
### The default judge prompt
|
|
156
|
+
|
|
157
|
+
```text
|
|
158
|
+
You are an impartial judge. Two independent attempts at the same operation returned different outputs,
|
|
159
|
+
candidate "a" and candidate "b". Both already match the required output format, so judge only their
|
|
160
|
+
content. You are given the operation's purpose and its input. Decide which candidate correctly and
|
|
161
|
+
faithfully carries out the purpose for this input. If both do (for example, they differ only in
|
|
162
|
+
wording), choose either one, preferring "a". Choose "neither" only when both are wrong or you can't tell
|
|
163
|
+
whether either is correct. Never combine them or invent a third answer.
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
It is available as `Squishling::Harness::DEFAULT_JUDGE_INSTRUCTIONS`. Replace it with `judge_instructions:`:
|
|
167
|
+
|
|
168
|
+
```ruby
|
|
169
|
+
squishling harness: { type: :judged_squishsum,
|
|
170
|
+
judge_instructions: "Pick the candidate whose priority follows our SLA rules: ..." }
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
## When the harness fails
|
|
174
|
+
|
|
175
|
+
A disagreement raises `Squishling::DisagreementError`, a subclass of `InvalidOutputError`. It's raised when the
|
|
176
|
+
samples differ with no judge, or when the judge rejects both. Like any failure on the LLM path, it goes to
|
|
177
|
+
[`squish_fallback`](failures.md#fallbacks), which can still return one of the candidates:
|
|
178
|
+
|
|
179
|
+
```ruby
|
|
180
|
+
squish_fallback do |error, **|
|
|
181
|
+
raise error unless error.is_a?(Squishling::DisagreementError)
|
|
182
|
+
|
|
183
|
+
error.candidates.first # already typed and validated; squished? is true
|
|
184
|
+
end
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
| Attribute | Value |
|
|
188
|
+
|---|---|
|
|
189
|
+
| `candidates` | The two typed results, `[a, b]` |
|
|
190
|
+
| `verdict` | `nil` when there was no judge; `:neither` when the judge rejected both |
|
|
191
|
+
| `reason` | The judge's reason (or, for a judgment model, the choice and its probability). A chat judge's reason is model-written, so it is not part of `message` or the logs; a judgment model's choice and probability are |
|
|
192
|
+
| `raw` | Both candidates as Hashes, `[a.to_h, b.to_h]` |
|
|
193
|
+
| `attempts` | `nil` |
|
|
194
|
+
| `models` | The model behind each role: sample `a`, sample `b`, then the judge if one ran (one entry each, not one per attempt) |
|
|
195
|
+
|
|
196
|
+
Other failures are unchanged. A sample or judge that never produces valid output raises `InvalidOutputError`, a
|
|
197
|
+
failed request on the last attempt raises `LLMError`, and setup mistakes (including a provider 400 on either
|
|
198
|
+
sample) raise `ConfigurationError`, which is never sent to the fallback.
|
|
@@ -0,0 +1,148 @@
|
|
|
1
|
+
# Measuring token spend
|
|
2
|
+
|
|
3
|
+
A squished path costs tokens on every call. A Ruby path costs engineer time once, plus upkeep. Measuring the
|
|
4
|
+
first tells you when the second is worth paying.
|
|
5
|
+
|
|
6
|
+
## What Squishling gives you
|
|
7
|
+
|
|
8
|
+
- `result.squished?` says which path served a call, so you can count LLM calls per class.
|
|
9
|
+
- The [`squawk` hook](failures.md#observing-every-attempt-squawk) runs after every LLM attempt and receives the
|
|
10
|
+
method (`label`, e.g. `"ShipmentUpdate#call"`), the model and provider, which attempt it was, and the token usage
|
|
11
|
+
RubyLLM reported. That is enough to attribute spend to a squishling method without any other tooling.
|
|
12
|
+
- Cost in dollars comes from RubyLLM's model pricing. Multiply the reported tokens by your model's price, or read it
|
|
13
|
+
from RubyLLM directly (see [the instrumenter](#lower-level-rubyllms-instrumenter)).
|
|
14
|
+
|
|
15
|
+
## Deciding when to write the code
|
|
16
|
+
|
|
17
|
+
For each squished path, over a month:
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
token cost = calls × (input tokens × input price + output tokens × output price)
|
|
21
|
+
code cost = time to write it + upkeep (new cases, format changes, on-call)
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Write the Ruby when the token cost keeps exceeding the code cost, and narrow `squish_when` so only the inputs
|
|
25
|
+
Ruby can't handle still go to the LLM. See [Hardening a path](routing.md#hardening-a-path). If the token cost
|
|
26
|
+
stays small, the debt you haven't taken on is cheaper than the code.
|
|
27
|
+
|
|
28
|
+
## Tools
|
|
29
|
+
|
|
30
|
+
| Option | Good for | Setup |
|
|
31
|
+
|---|---|---|
|
|
32
|
+
| [Coolhand Labs](#coolhand-labs) | Picking up your squishling calls automatically and measuring their cost and accuracy | The `coolhand` gem and an initializer |
|
|
33
|
+
| [Custom: the `squawk` hook](#custom-the-squawk-hook) | Logging or metrics in your own stack, with no new dependency | A lambda and one config line |
|
|
34
|
+
| [OpenTelemetry](#opentelemetry) | Sending traces to Datadog, Langfuse, Arize, Braintrust, LangSmith, or any OTel backend | The OpenTelemetry SDK, an exporter, and one line |
|
|
35
|
+
| [LangSmith](#langsmith) | Browsing traces and token usage in LangSmith | Through OpenTelemetry |
|
|
36
|
+
|
|
37
|
+
### Coolhand Labs
|
|
38
|
+
|
|
39
|
+
Add and initialize the [`coolhand`](https://github.com/Coolhand-Labs/coolhand-ruby) gem. It intercepts the LLM
|
|
40
|
+
requests your app makes, so it picks up your squishling calls automatically and measures their cost, and their
|
|
41
|
+
accuracy through Coolhand's feedback tools. No changes to your squishling classes.
|
|
42
|
+
|
|
43
|
+
```ruby
|
|
44
|
+
# Gemfile
|
|
45
|
+
gem "coolhand"
|
|
46
|
+
|
|
47
|
+
# config/initializers/coolhand.rb
|
|
48
|
+
Coolhand.configure do |config|
|
|
49
|
+
config.api_key = ENV.fetch("COOLHAND_API_KEY")
|
|
50
|
+
end
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Get an API key at [coolhandlabs.com](https://coolhandlabs.com/). To keep the data in your own infrastructure, point
|
|
54
|
+
`config.base_url` at a Coolhand-compatible endpoint of your own.
|
|
55
|
+
|
|
56
|
+
### Custom: the `squawk` hook
|
|
57
|
+
|
|
58
|
+
`config.squawk` runs after every LLM attempt, accepted or rejected, and is handed the method's label, the model,
|
|
59
|
+
and the token usage RubyLLM reported. A lambda can take just the keywords it needs:
|
|
60
|
+
|
|
61
|
+
```ruby
|
|
62
|
+
Squishling.configure do |config|
|
|
63
|
+
config.squawk = lambda do |metadata:, **|
|
|
64
|
+
usage = metadata[:usage] || {} # nil when the provider call failed; counts it didn't report are absent
|
|
65
|
+
Rails.logger.info("squishling #{metadata[:label]} model=#{metadata[:model]} " \
|
|
66
|
+
"attempt=#{metadata[:attempt]}/#{metadata[:attempts]} " \
|
|
67
|
+
"in=#{usage[:input_tokens]} out=#{usage[:output_tokens]}")
|
|
68
|
+
end
|
|
69
|
+
end
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
Each line carries the `"Class#method"` that made the call, so grouping by label gives you tokens per squishling. A call
|
|
73
|
+
that stays in Ruby makes no LLM request and never reaches the hook, so a label that stops appearing is a path that has
|
|
74
|
+
been hardened. You can also set the hook per class (`squishling squawk: ...`) or per method. See
|
|
75
|
+
[Observing every attempt](failures.md#observing-every-attempt-squawk) for every keyword and the rules for exceptions.
|
|
76
|
+
|
|
77
|
+
### Lower level: RubyLLM's instrumenter
|
|
78
|
+
|
|
79
|
+
To meter every RubyLLM request in your app, not only Squishling's, use RubyLLM's own instrumentation. It reports a
|
|
80
|
+
`usage.ruby_llm` event for each provider request, with `model`, `provider`, `status`, `tokens` (`input`, `output`,
|
|
81
|
+
and more) and `cost`. Give it an object that responds to `instrument(name, payload)` and optionally takes a block. In
|
|
82
|
+
Rails, `ActiveSupport::Notifications` is used automatically.
|
|
83
|
+
|
|
84
|
+
```ruby
|
|
85
|
+
class TokenMeter
|
|
86
|
+
def instrument(name, payload = {})
|
|
87
|
+
if name == "usage.ruby_llm"
|
|
88
|
+
Rails.logger.info("llm model=#{payload[:model]} workflow=#{payload[:workflow_name]} " \
|
|
89
|
+
"in=#{payload[:tokens].input} out=#{payload[:tokens].output}")
|
|
90
|
+
end
|
|
91
|
+
block_given? ? yield : nil
|
|
92
|
+
end
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
RubyLLM.configure { |config| config.instrumenter = TokenMeter.new }
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
These events don't know which squishling made the request. To attribute the spend, wrap the call in a RubyLLM
|
|
99
|
+
workflow, and every event emitted inside the block carries its name, id and any `metadata:` you pass. (The
|
|
100
|
+
`squawk` hook above needs no wrapping.)
|
|
101
|
+
|
|
102
|
+
```ruby
|
|
103
|
+
RubyLLM.workflow("ShipmentUpdate", metadata: { carrier: "acme-freight" }) do
|
|
104
|
+
ShipmentUpdate.call(carrier: "acme-freight", payload: body)
|
|
105
|
+
end
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
### OpenTelemetry
|
|
109
|
+
|
|
110
|
+
RubyLLM 2.1+ ships its own OpenTelemetry tracer. Each request becomes a span with the model, request settings and
|
|
111
|
+
token usage (`gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`), following the OpenTelemetry GenAI
|
|
112
|
+
conventions. It does not record message text. Your app owns the SDK and exporters.
|
|
113
|
+
|
|
114
|
+
```ruby
|
|
115
|
+
# Gemfile
|
|
116
|
+
gem "opentelemetry-sdk"
|
|
117
|
+
gem "opentelemetry-exporter-otlp" # or your backend's exporter
|
|
118
|
+
|
|
119
|
+
# config/initializers/opentelemetry.rb
|
|
120
|
+
OpenTelemetry::SDK.configure # reads the standard OTEL_* environment variables
|
|
121
|
+
RubyLLM::OpenTelemetry.enable # starts tracing RubyLLM requests
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
Spans don't say which squishling made the request, so pair this with a `RubyLLM.workflow` (see
|
|
125
|
+
[above](#lower-level-rubyllms-instrumenter)) or group by model and prompt in your backend.
|
|
126
|
+
|
|
127
|
+
### LangSmith
|
|
128
|
+
|
|
129
|
+
LangSmith accepts OpenTelemetry traces, so the OpenTelemetry setup above covers it: point the OTLP exporter at
|
|
130
|
+
LangSmith. See LangSmith's
|
|
131
|
+
[OpenTelemetry docs](https://docs.smith.langchain.com/observability/how_to_guides/tracing/trace_with_opentelemetry)
|
|
132
|
+
for the endpoint and headers. A community Ruby gem, [`langsmith-sdk`](https://github.com/felipekb/langsmith-ruby-sdk),
|
|
133
|
+
also exists; it isn't published by LangChain.
|
|
134
|
+
|
|
135
|
+
## Privacy
|
|
136
|
+
|
|
137
|
+
Squishling only sends the model what you name: the method arguments, any `squish_context` values, and any source you
|
|
138
|
+
append. Tools that sit on the request path are a separate route out of your process, and they see both sides:
|
|
139
|
+
|
|
140
|
+
- **Request interceptors, such as `coolhand`,** record the prompts (your purpose text and those inputs) and the
|
|
141
|
+
model's responses. Coolhand sends them to coolhandlabs.com unless you point `config.base_url` at an endpoint of your
|
|
142
|
+
own.
|
|
143
|
+
- **A `squawk` hook that forwards `output:` or `metadata[:input]`** sends model output and your inputs wherever you
|
|
144
|
+
send them. That is the sanctioned route for model output; see
|
|
145
|
+
[Observing every attempt](failures.md#observing-every-attempt-squawk). The examples above log only metadata.
|
|
146
|
+
- **OpenTelemetry** with RubyLLM's tracer records usage and request settings, not message text.
|
|
147
|
+
|
|
148
|
+
Check what your tool stores before pointing it at inputs that contain personal data.
|
data/docs/naming.md
ADDED
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# Naming and collisions
|
|
2
|
+
|
|
3
|
+
`include Squishling` adds a small set of methods to your class. This page lists them and explains what happens
|
|
4
|
+
when one of those names is already taken.
|
|
5
|
+
|
|
6
|
+
## What Squishling adds
|
|
7
|
+
|
|
8
|
+
| Where | Names |
|
|
9
|
+
|---|---|
|
|
10
|
+
| Class-level DSL | `squishling`, `purpose`, `append_to_purpose`, `output_schema`, `squish`, `squish_when`, `squish_fallback`, `squish_validate`, `squish_context`, `squished_methods`, `call`, and the `squishling_*` readers |
|
|
11
|
+
| Instance | `squishling_result`, `squish!`, and `result` (an alias for `squishling_result`) |
|
|
12
|
+
|
|
13
|
+
Methods your class defines itself always win over these, whether you define them before or after the `include`.
|
|
14
|
+
Only methods your class *inherits* can be shadowed.
|
|
15
|
+
|
|
16
|
+
## Inherited collisions raise
|
|
17
|
+
|
|
18
|
+
If a parent class or an earlier-included module already defines a class-level DSL method, `squish!` or
|
|
19
|
+
`squishling_result`, `include Squishling` raises `Squishling::ConfigurationError` listing every clash and where
|
|
20
|
+
it comes from:
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
ConfigurationError: MyApp inherits methods that Squishling would override: MyApp.call (from Sinatra::Base's
|
|
24
|
+
class methods). Include Squishling in a plain Ruby class instead ...
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
This is deliberate. Squishling's modules sit ahead of inherited ones, so without the check
|
|
28
|
+
`Sinatra::Base.call(env)` (the Rack entry point), or any `call` or `purpose` your parent class defines, would be
|
|
29
|
+
replaced silently. Put Squishling in a plain Ruby class and have your framework object call it instead.
|
|
30
|
+
|
|
31
|
+
## `result` is skipped when taken
|
|
32
|
+
|
|
33
|
+
`result` is a convenience for `squishling_result`. If the class already has a `result` (its own or inherited), it is
|
|
34
|
+
left alone and Squishling does not add the alias. Use `squishling_result(...)` instead; it is always available.
|
|
35
|
+
|
|
36
|
+
**ActiveRecord caveat.** Column readers are generated lazily, so Squishling can't see a `result` column when you
|
|
37
|
+
include it. Its `result` would then shadow the column reader. On a model with a `result` column, call
|
|
38
|
+
`squishling_result` (or keep Squishling out of the model and use a plain class).
|
|
39
|
+
|
|
40
|
+
## Not a collision
|
|
41
|
+
|
|
42
|
+
ActiveSupport adds `String#squish` and `String#squish!`. Those live on `String`; Squishling's `squish!` lives on
|
|
43
|
+
your class, so they never meet. They only look alike when reading code in a Rails app.
|