squishling 0.2.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +83 -0
- data/README.md +85 -30
- data/docs/configuration.md +40 -16
- data/docs/failures.md +68 -8
- data/docs/harnesses.md +198 -0
- data/docs/measuring-tokens.md +148 -0
- data/docs/naming.md +43 -0
- data/docs/routing.md +36 -27
- data/docs/schemas.md +4 -3
- data/lib/squishling/appendices.rb +3 -3
- data/lib/squishling/class_methods.rb +60 -29
- data/lib/squishling/collisions.rb +54 -0
- data/lib/squishling/configuration.rb +19 -0
- data/lib/squishling/definition.rb +48 -23
- data/lib/squishling/errors.rb +22 -1
- data/lib/squishling/harness.rb +153 -0
- data/lib/squishling/invoker.rb +188 -133
- data/lib/squishling/judge.rb +142 -0
- data/lib/squishling/llm_client.rb +150 -0
- data/lib/squishling/model_path.rb +21 -7
- data/lib/squishling/output_check.rb +109 -0
- data/lib/squishling/params.rb +37 -1
- data/lib/squishling/router.rb +8 -2
- data/lib/squishling/source.rb +5 -5
- data/lib/squishling/squawk.rb +52 -0
- data/lib/squishling/version.rb +1 -1
- data/lib/squishling.rb +24 -7
- metadata +12 -3
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: '0748f552c985521890d365ebc0818ec59c4336f93fb183d1769f49ed8cd32349'
|
|
4
|
+
data.tar.gz: '05128ecf8d7156e0c8ca084bf8a1860b754596017df2cdfdcc35e2242138fa7b'
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 2be443b51d2dc1460a410fbf0850bed09e33fff5f7ad894d4bd92e3f47d428d22df7bfc0371036c090a6b806f466b1d930de3a9caf8d1b044185fd4000b765bf
|
|
7
|
+
data.tar.gz: 3a18d6e041f5807c966677fea5499de8af5d0f28146cb98bbe205141598ee6a01a2d598b1326df3b78227455dcc124c69cf95afacabf9a1689a1cc382a652414
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,89 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.3.0] - 2026-10-09
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- Harnesses: a `harness:` option (on `squishling`, `squish`, `squish!`, and `default_harness` in
|
|
15
|
+
`Squishling.configure`) chooses how a call uses its escalation. `:escalation` is the existing behavior and stays
|
|
16
|
+
the default. `:squishsum` asks the first escalation step twice, concurrently, and accepts the result only when the
|
|
17
|
+
two samples agree (exact match, or a custom `compare:`). `:judged_squishsum` sends a disagreement to a judge: the
|
|
18
|
+
next escalation step or a dedicated `judge:` step, either a chat model or a Jev-style decision model through
|
|
19
|
+
`RubyLLM.judge`, with a default prompt you can override through `judge_instructions:`. Samples that still disagree
|
|
20
|
+
raise the new `Squishling::DisagreementError < InvalidOutputError`, which carries both typed results, and
|
|
21
|
+
`squish_fallback` remains the only fallback. See [Harnesses](docs/harnesses.md). (#16)
|
|
22
|
+
- `:ensemble` and `:judged_ensemble` harnesses: like the squishsum ones, but the two concurrent samples come from the
|
|
23
|
+
escalation's first and second steps (a cross-model check) instead of the first step twice, each retrying within its
|
|
24
|
+
own step's `attempts:`. A judged ensemble's judge defaults to the third step, or `judge:`, and
|
|
25
|
+
`DisagreementError#models` names the model behind each sample. An ensemble needs at least two distinct steps (adjacent
|
|
26
|
+
identical steps count as one step's attempts) or it raises `ConfigurationError` before any request. (#20)
|
|
27
|
+
- `squawk`, an opt-in hook for sending what the model returned to an error tracker or tracing tool. It is called after
|
|
28
|
+
every LLM attempt (every sample and judge attempt under a harness) with `output:`, `metadata:`, and `error:`, and
|
|
29
|
+
can be set on `Squishling.configure`, `squishling`, and `squish`. It does nothing by default, and exceptions it
|
|
30
|
+
raises propagate unchanged. (#15)
|
|
31
|
+
- `forward_rejected:` on escalation steps. When an attempt is rejected and the next step starts, the previous
|
|
32
|
+
output is forwarded to that step's provider by default (unchanged behavior); set `forward_rejected: false` to start
|
|
33
|
+
that step from the original input alone. (#14)
|
|
34
|
+
- `include Squishling` raises `ConfigurationError` when the class inherits a method Squishling would override (for
|
|
35
|
+
example `Sinatra::Base.call`) instead of silently shadowing it. See [Naming and collisions](docs/naming.md). (#17)
|
|
36
|
+
|
|
37
|
+
### Documentation
|
|
38
|
+
|
|
39
|
+
- A new guide, [Measuring token spend](docs/measuring-tokens.md), covers working out what each squished path costs
|
|
40
|
+
(with Coolhand Labs, the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter), and the README is
|
|
41
|
+
reworked around using Squishling to harden code paths as the economics justify it. (#21)
|
|
42
|
+
|
|
43
|
+
### Changed
|
|
44
|
+
|
|
45
|
+
- **Breaking:** the DSL method `instructions` is now `purpose`, and `append_instructions` is now `append_to_purpose`,
|
|
46
|
+
including the `instructions:`/`append_instructions:` keywords on `squishling`, `squish`, and `squish!`. There are no
|
|
47
|
+
aliases. Rename them in your classes: `instructions "..."` becomes `purpose "..."`, and
|
|
48
|
+
`append_instructions: [...]` becomes `append_to_purpose: [...]`. The error for a missing prompt now says
|
|
49
|
+
"has no purpose". (#17)
|
|
50
|
+
- **Breaking:** `ruby_llm` must now be `~> 2.1` (was `~> 2.0`); run `bundle update ruby_llm`. (#16)
|
|
51
|
+
- **Breaking:** every value a method with an `output_schema` returns is now validated and typed, not only Hashes. On
|
|
52
|
+
the Ruby path and in `squish_fallback`, a `nil`, String, Array, or other non-Hash return that used to pass through
|
|
53
|
+
now raises `InvalidOutputError`; return a Hash (or `result(...)`) that matches the schema. A result of the schema's
|
|
54
|
+
own class passes through, another schema's result is re-validated, and values JSON can't represent (`NaN`,
|
|
55
|
+
`Infinity`, cycles) raise `InvalidOutputError` instead of a raw JSON error. Inner calls (`super`, or the method
|
|
56
|
+
called from `squish_when` or a fallback) still only type Hashes, so overrides can reshape values. (#13)
|
|
57
|
+
- **Behavior change:** a provider now travels with the model declared at the same level on a class too. A subclass
|
|
58
|
+
that declares its own `model:` or `escalation:` no longer inherits its parent's `provider:`, which could send the
|
|
59
|
+
whole payload to the parent's provider under another provider's model name. If you relied on that, set `provider:`
|
|
60
|
+
next to the subclass's model.
|
|
61
|
+
- **Behavior change:** a `provider:` with no model beside it now raises `ConfigurationError` instead of being
|
|
62
|
+
silently ignored (the default provider was used). `squish ..., provider: ...` needs a `model:` or `escalation:` in
|
|
63
|
+
the same call, as `squish!` already did; `squishling provider: ...` needs one in the call or already declared on that
|
|
64
|
+
class. Move the `provider:` next to a model, or use `config.default_provider` with the default model.
|
|
65
|
+
- The `result` alias for `squishling_result` is no longer added when the class already has a `result` (its own or
|
|
66
|
+
inherited); use `squishling_result` there. (#17)
|
|
67
|
+
- `InvalidOutputError#message` and `config.logger` warnings no longer quote the model's response. For unparseable JSON
|
|
68
|
+
they report only the line and column; the raw output stays in `InvalidOutputError#raw`. (#15)
|
|
69
|
+
|
|
70
|
+
### Fixed
|
|
71
|
+
|
|
72
|
+
- A number a model writes outside the range JSON can represent (such as `1e400`, which parses to `Infinity`), or a
|
|
73
|
+
string with invalid UTF-8, was accepted as valid, while the same value from a Ruby method was rejected. It is now
|
|
74
|
+
invalid output on the LLM path too, and no longer raises a bare `JSON::GeneratorError` when `:squishsum` compares samples.
|
|
75
|
+
|
|
76
|
+
### Security
|
|
77
|
+
|
|
78
|
+
- Params can no longer replace the strict output format or tools through the containers that also hold ordinary
|
|
79
|
+
settings: `generationConfig`/`generation_config` (Google Gemini) and `completion_args` (Mistral Conversations) now
|
|
80
|
+
reject their format and tool keys, with symbol or string keys, and a non-Hash value for them is rejected.
|
|
81
|
+
`tool_config` is reserved at the top level. Settings such as `topK` still work. (#12)
|
|
82
|
+
- The previous model's rejected output forwarded to the next escalation step is capped at 4,000 characters
|
|
83
|
+
(`Invoker::MAX_FORWARDED_CHARS`), with a truncation marker. (#14)
|
|
84
|
+
- Model output is kept out of error messages and logs (see the `InvalidOutputError#message` change above); `squawk`
|
|
85
|
+
is the one sanctioned way for it to leave the process. (#15) A chat judge's `reason` is model-written, so
|
|
86
|
+
`DisagreementError#message` and the logs no longer include it; read it from `DisagreementError#reason`. A
|
|
87
|
+
`:judgment` judge's message still carries the choice and probability, and a choice other than `a`, `b`, or `neither`
|
|
88
|
+
is no longer repeated.
|
|
89
|
+
- Only the first 20 validation errors (plus a count of the rest) are fed back to the model, logged, and put in
|
|
90
|
+
`InvalidOutputError#message` for one invalid output, so a very large malformed response can no longer produce a
|
|
91
|
+
megabyte-sized retry message or log line.
|
|
92
|
+
|
|
10
93
|
## [0.2.0] - 2026-10-09
|
|
11
94
|
|
|
12
95
|
### Added
|
data/README.md
CHANGED
|
@@ -1,46 +1,54 @@
|
|
|
1
|
-
|
|
1
|
+
<h1 align="center">
|
|
2
|
+
<img src="assets/squishling-logo.png" alt="Squishling" width="420">
|
|
3
|
+
</h1>
|
|
2
4
|
|
|
3
5
|
[](https://github.com/Coolhand-Labs/squishling/actions/workflows/ci.yml)
|
|
6
|
+
[](https://badge.fury.io/rb/squishling)
|
|
4
7
|
|
|
5
|
-
|
|
6
|
-
(via [RubyLLM](https://rubyllm.com)), and either way returns the same strict-schema-validated, typed result.
|
|
7
|
-
Callers can't tell the difference.
|
|
8
|
+
**Don't just use AI to write code. Use AI to not need code.** Write code only where it's worth maintaining.
|
|
8
9
|
|
|
9
|
-
|
|
10
|
-
|
|
10
|
+
A squished method either runs its Ruby implementation, with the LLM as a failover, or routes some callers
|
|
11
|
+
through an LLM as a way of eliminating (or discovering) the cost of maintaining them as deterministic code.
|
|
12
|
+
Either way it returns the same strict-schema-validated, typed result, so callers can't tell the difference.
|
|
13
|
+
|
|
14
|
+
Think of code as a cache for AI judgment. Every line you write is something to maintain, so the LLM is the
|
|
15
|
+
default and code has to earn its place: you build it for the inputs that carry the volume, and the LLM keeps
|
|
16
|
+
handling the rest. New to the idea? Read
|
|
17
|
+
[Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
|
|
11
18
|
|
|
12
19
|
## Use cases
|
|
13
20
|
|
|
14
|
-
###
|
|
21
|
+
### Zero-code integrations
|
|
15
22
|
|
|
16
23
|
Accept a new data source today, before anyone writes a parser. Declare what you want back and leave the
|
|
17
24
|
method unimplemented: every call goes to the LLM, and its output is validated against your schema.
|
|
18
25
|
|
|
19
26
|
```ruby
|
|
20
|
-
class
|
|
27
|
+
class ShipmentUpdate
|
|
21
28
|
include Squishling
|
|
22
29
|
|
|
23
|
-
|
|
30
|
+
purpose "Normalize this carrier's tracking webhook into our shipment update."
|
|
24
31
|
output_schema do
|
|
25
|
-
string
|
|
26
|
-
|
|
27
|
-
string
|
|
28
|
-
string :external_id
|
|
32
|
+
string :status, enum: %w[label_created in_transit out_for_delivery delivered exception]
|
|
33
|
+
string :tracking_number
|
|
34
|
+
string :location
|
|
29
35
|
end
|
|
30
36
|
# No `def call` yet, so every webhook goes to the LLM.
|
|
31
37
|
end
|
|
32
38
|
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
event.squished? # => true
|
|
39
|
+
update = ShipmentUpdate.call(carrier: "acme-freight", payload: request.raw_post)
|
|
40
|
+
update.status # => "out_for_delivery"
|
|
41
|
+
update.squished? # => true
|
|
37
42
|
```
|
|
38
43
|
|
|
39
|
-
When one
|
|
40
|
-
`squish_when { |
|
|
44
|
+
When one carrier carries the volume, write `def call` for it and add
|
|
45
|
+
`squish_when { |carrier:, **| carrier != "ups" }`. UPS then runs in Ruby, everything else stays on the
|
|
41
46
|
LLM, and callers don't change. See [Hardening a path](docs/routing.md#hardening-a-path).
|
|
42
47
|
|
|
43
|
-
|
|
48
|
+
Validation guarantees the *shape* of the result, not that it's true. Where a wrong value is costly (money,
|
|
49
|
+
identity), cross-check it against another source or keep that path in Ruby.
|
|
50
|
+
|
|
51
|
+
### Rescue errors your code can't handle yet
|
|
44
52
|
|
|
45
53
|
Keep the Ruby you have for the inputs it understands, and hand the rest to the LLM instead of failing. When the
|
|
46
54
|
parser raises, `squish!` sends this call to the LLM with the error and the parser's own source as context.
|
|
@@ -49,8 +57,8 @@ parser raises, `squish!` sends this call to the LLM with the error and the parse
|
|
|
49
57
|
class InvoiceParser
|
|
50
58
|
include Squishling
|
|
51
59
|
|
|
52
|
-
|
|
53
|
-
|
|
60
|
+
purpose "Extract the invoice fields from the vendor's document."
|
|
61
|
+
append_to_purpose "The Ruby parser that handles well-formed invoices:", self # this class's source
|
|
54
62
|
output_schema do
|
|
55
63
|
string :invoice_number
|
|
56
64
|
number :total
|
|
@@ -60,7 +68,7 @@ class InvoiceParser
|
|
|
60
68
|
invoice = VendorFormats.fetch(vendor).parse(document)
|
|
61
69
|
result(invoice_number: invoice.number, total: invoice.total)
|
|
62
70
|
rescue VendorFormats::ParseError => e
|
|
63
|
-
squish!(
|
|
71
|
+
squish!(append_to_purpose: "The parser failed on this document; the error is in the context.",
|
|
64
72
|
context: { parse_error: e })
|
|
65
73
|
end
|
|
66
74
|
end
|
|
@@ -73,13 +81,36 @@ Both calls return the same result class. If the LLM can't deliver either, `squis
|
|
|
73
81
|
return, or the error is raised with the original `ParseError` as its cause. See
|
|
74
82
|
[Handing off to the LLM](docs/routing.md#handing-off-to-the-llm-with-squish).
|
|
75
83
|
|
|
84
|
+
### Measure your tech debt in tokens
|
|
85
|
+
|
|
86
|
+
Code you haven't written is debt you haven't taken on. A squished path has a running cost you can read off a
|
|
87
|
+
meter, so "should we write a parser for this?" becomes arithmetic: what the path costs in tokens each month,
|
|
88
|
+
against what it costs to write and maintain the code. Write it when the first number is bigger.
|
|
89
|
+
|
|
90
|
+
Add and initialize the [`coolhand`](https://github.com/Coolhand-Labs/coolhand-ruby) gem, and it picks up your
|
|
91
|
+
squishling calls automatically and measures their accuracy and cost:
|
|
92
|
+
|
|
93
|
+
```ruby
|
|
94
|
+
# Gemfile
|
|
95
|
+
gem "coolhand"
|
|
96
|
+
|
|
97
|
+
# config/initializers/coolhand.rb
|
|
98
|
+
Coolhand.configure do |config|
|
|
99
|
+
config.api_key = ENV.fetch("COOLHAND_API_KEY")
|
|
100
|
+
end
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
`squished?` tells you which path served a call. Coolhand records your LLM requests and responses; see
|
|
104
|
+
[what each tool sees](docs/measuring-tokens.md#privacy). Want to roll your own metrics? Check out
|
|
105
|
+
[our guide](docs/measuring-tokens.md) for doing it with other tools.
|
|
106
|
+
|
|
76
107
|
## Installation
|
|
77
108
|
|
|
78
109
|
```ruby
|
|
79
110
|
gem "squishling"
|
|
80
111
|
```
|
|
81
112
|
|
|
82
|
-
Requires Ruby 3.3+ and RubyLLM 2.
|
|
113
|
+
Requires Ruby 3.3+ and RubyLLM 2.1+. Configure your provider API keys in RubyLLM as usual, then optionally set a
|
|
83
114
|
universal model, or an escalation of models to try in order:
|
|
84
115
|
|
|
85
116
|
```ruby
|
|
@@ -91,7 +122,7 @@ end
|
|
|
91
122
|
|
|
92
123
|
## How it works
|
|
93
124
|
|
|
94
|
-
A squishling class needs **
|
|
125
|
+
A squishling class needs a **purpose** (the system prompt) and an **output schema** (the shape of the result,
|
|
95
126
|
validated on both paths). A squished call goes to the LLM when:
|
|
96
127
|
|
|
97
128
|
- its `squish_when` predicate is truthy for these inputs,
|
|
@@ -104,13 +135,13 @@ Otherwise the Ruby runs, and whatever it returns is validated and typed like LLM
|
|
|
104
135
|
|
|
105
136
|
- **One contract, two paths**: Ruby returns and LLM output are validated against the same strict schema and
|
|
106
137
|
returned as the same typed `Data` objects. `squished?` tells you which path served a call.
|
|
107
|
-
- **Your code as context**: `
|
|
138
|
+
- **Your code as context**: `append_to_purpose` adds sections to the prompt, including a class's or method's
|
|
108
139
|
own Ruby source.
|
|
109
140
|
- **Any RubyLLM provider and model**: OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, OpenRouter, and more.
|
|
110
141
|
Set a universal, per-class, per-method, or per-call model, plus layered generation params (temperature,
|
|
111
142
|
reasoning effort, top_p, …).
|
|
112
143
|
- **Model escalation**: declare an `escalation:` instead of a `model:`, with per-step `attempts:`, `order:`, provider,
|
|
113
|
-
and
|
|
144
|
+
params, and `forward_rejected:`. Invalid output (with its errors fed back) or a provider failure moves on to the next attempt. Steps can
|
|
114
145
|
cross providers, e.g. a local Qwen model on Ollama, then Anthropic Claude Haiku on AWS Bedrock, then Claude Opus:
|
|
115
146
|
|
|
116
147
|
```ruby
|
|
@@ -130,23 +161,36 @@ Otherwise the Ruby runs, and whatever it returns is validated and typed like LLM
|
|
|
130
161
|
{ model: "claude-haiku-4-5", provider: :bedrock }, # small hosted model
|
|
131
162
|
{ model: "claude-opus-5-5", provider: :anthropic } # last resort
|
|
132
163
|
]
|
|
133
|
-
#
|
|
164
|
+
# purpose, output_schema, ...
|
|
134
165
|
end
|
|
135
166
|
```
|
|
167
|
+
- **[Harnesses](docs/harnesses.md)**: choose how a call uses its escalation with `harness:`
|
|
168
|
+
- [`:escalation`](docs/configuration.md#models-and-escalation) (default): try each attempt until one passes
|
|
169
|
+
- [`:squishsum`](docs/harnesses.md#samples): two concurrent samples, accepted only if they agree
|
|
170
|
+
- [`:judged_squishsum`](docs/harnesses.md#the-judge): a chat or Jev judge picks between disagreeing samples
|
|
171
|
+
- [`:ensemble` and `:judged_ensemble`](docs/harnesses.md#ensembles): the same checks across two different models
|
|
136
172
|
- **Output contracts**: beyond the strict schema, conditional rules (`given`) and Ruby checks (`squish_validate`)
|
|
137
173
|
reject bad output and trigger the next attempt.
|
|
138
174
|
- **Defined failure behavior**: provider errors become `Squishling::LLMError`, and `squish_fallback` lets you decide
|
|
139
175
|
what to return when every attempt fails.
|
|
176
|
+
- **Observability**: the opt-in [`squawk`](docs/failures.md#observing-every-attempt-squawk) hook sends every attempt's
|
|
177
|
+
raw output to your error tracker. Model output stays out of error messages and logs otherwise.
|
|
140
178
|
- **Opt-in context**: only method arguments, the instance state you name with `squish_context`, the `context:` you
|
|
141
|
-
pass to `squish!`, and the source you choose to append are sent to the provider.
|
|
179
|
+
pass to `squish!`, and the source you choose to append are sent to the provider. The one addition: when
|
|
180
|
+
[escalation](docs/configuration.md#models-and-escalation) moves to the next step, that step also sees the previous
|
|
181
|
+
model's rejected output, unless the step sets `forward_rejected: false`.
|
|
142
182
|
|
|
143
183
|
## Documentation
|
|
144
184
|
|
|
145
185
|
- [Configuration](docs/configuration.md): options, models and escalation, providers, generation params, inheritance
|
|
146
|
-
- [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `
|
|
186
|
+
- [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `append_to_purpose`, hardening a path,
|
|
147
187
|
what the LLM sees
|
|
148
188
|
- [Output schemas](docs/schemas.md): schema forms, strict mode, typed results, optional vs. empty, contracts
|
|
149
189
|
- [Failure handling](docs/failures.md): escalation, `squish_validate`, error classes, fallbacks
|
|
190
|
+
- [Harnesses](docs/harnesses.md): escalation, squishsum, ensemble, and judges (chat or Jev)
|
|
191
|
+
- [Naming and collisions](docs/naming.md): the methods `include Squishling` adds and what happens when a name is taken
|
|
192
|
+
- [Measuring token spend](docs/measuring-tokens.md): see what each squished path costs, with Coolhand Labs,
|
|
193
|
+
the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter
|
|
150
194
|
- [Live examples](examples/README.md): end-to-end tests against Anthropic Claude Haiku and OpenAI GPT-6 Luna
|
|
151
195
|
|
|
152
196
|
## Development
|
|
@@ -158,6 +202,17 @@ bundle exec rake # RSpec (offline; RubyLLM is stubbed) + RuboCop
|
|
|
158
202
|
|
|
159
203
|
See [AGENTS.md](AGENTS.md) for repo conventions, and [SECURITY.md](SECURITY.md) for reporting vulnerabilities.
|
|
160
204
|
|
|
205
|
+
## Credits
|
|
206
|
+
|
|
207
|
+
Inspired by [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
|
|
208
|
+
|
|
161
209
|
## License
|
|
162
210
|
|
|
163
211
|
Apache-2.0
|
|
212
|
+
|
|
213
|
+
---
|
|
214
|
+
|
|
215
|
+
<p align="center">
|
|
216
|
+
This open source project is supported by <a href="https://coolhandlabs.com/">Coolhand Labs</a>.<br><br>
|
|
217
|
+
<a href="https://coolhandlabs.com/"><img src="assets/coolhand-labs.png" alt="Coolhand Labs" width="220"></a>
|
|
218
|
+
</p>
|
data/docs/configuration.md
CHANGED
|
@@ -17,6 +17,7 @@ Squishling.configure do |config|
|
|
|
17
17
|
config.default_model = "claude-sonnet-5-5" # or default_escalation, below
|
|
18
18
|
config.default_provider = nil
|
|
19
19
|
config.default_params = { temperature: 0 }
|
|
20
|
+
config.default_harness = :escalation
|
|
20
21
|
config.logger = Rails.logger
|
|
21
22
|
end
|
|
22
23
|
```
|
|
@@ -27,7 +28,9 @@ end
|
|
|
27
28
|
| `default_escalation` | `nil` | The universal [escalation](#models-and-escalation): models tried in order. Assigning it clears `default_model`, and vice versa. |
|
|
28
29
|
| `default_provider` | `nil` | The provider for `default_model`/`default_escalation` steps that don't name one. Only needed for models missing from RubyLLM's registry (see below). |
|
|
29
30
|
| `default_params` | `{}` | Generation params for every call (temperature, reasoning effort, top_p, …), overridable per class, per method, and per escalation step. See [Generation params](#generation-params). |
|
|
30
|
-
| `
|
|
31
|
+
| `squawk` | `nil` | A callable run after every LLM attempt with the raw output, for sending it to an error tracker or tracing tool. Also settable per class and per method. See [Observing every attempt](failures.md#observing-every-attempt-squawk). |
|
|
32
|
+
| `default_harness` | `nil` (`:escalation`) | How every squishling class that doesn't declare its own uses its escalation: `:escalation`, `:squishsum`, `:judged_squishsum`, `:ensemble`, `:judged_ensemble`, or a Hash with `type:` and options. See [Harnesses](harnesses.md). |
|
|
33
|
+
| `logger` | `nil` | Any `Logger`. Debug lines when a call routes to the LLM, warnings on each escalation and when a fallback is used. Never includes the raw model response; see [what ends up in errors and logs](failures.md#what-ends-up-in-errors-and-logs). |
|
|
31
34
|
|
|
32
35
|
Transport-level retries (rate limits, 5xx, timeouts) are configured on RubyLLM itself
|
|
33
36
|
(`RubyLLM.config.max_retries`, `request_timeout`).
|
|
@@ -61,6 +64,7 @@ Each step is a model name, or a Hash with:
|
|
|
61
64
|
| `attempts:` | `1` | How many attempts this step gets before escalating to the next one. |
|
|
62
65
|
| `order:` | position | An Integer; steps run lowest first. Give every step an `order:` or none (duplicates are rejected). Gaps are fine. |
|
|
63
66
|
| `provider:` | the level's `provider:` | See [Providers](#providers-and-newly-released-models). |
|
|
67
|
+
| `forward_rejected:` | `true` | Set `false` to start this step from the original input alone, without the previous step's rejected output and its errors (see below). |
|
|
64
68
|
| `params:` | `{}` | [Generation params](#generation-params) for this step, merged over the class and method params key by key. |
|
|
65
69
|
|
|
66
70
|
Without `order:`, steps run in the order written. With it, the order is explicit and doesn't depend on position:
|
|
@@ -105,10 +109,13 @@ end
|
|
|
105
109
|
Passing both `model:` and `escalation:` in one declaration raises `ConfigurationError`, and so does a list passed as
|
|
106
110
|
`model:`. Across separate declarations (a reopened class, or two config assignments), the latest one wins.
|
|
107
111
|
|
|
108
|
-
Consecutive attempts of the same step (same model, provider,
|
|
109
|
-
conversation. Moving to a different step starts a fresh chat with the original input plus the previous
|
|
110
|
-
it was rejected, so no provider-specific history crosses providers.
|
|
111
|
-
next
|
|
112
|
+
Consecutive attempts of the same step (same model, provider, params, and `forward_rejected:`, e.g. `attempts: 2`)
|
|
113
|
+
continue one conversation. Moving to a different step starts a fresh chat with the original input plus the previous
|
|
114
|
+
output and why it was rejected, so no provider-specific history crosses providers. That rejected output is
|
|
115
|
+
model-generated and is sent to the next step's provider, even when it is a different one, truncated to 4,000
|
|
116
|
+
characters. If you don't want a step to see it (say, because the next step is a different provider), set
|
|
117
|
+
`forward_rejected: false` on that step and it starts from the original input only. Configuration errors (bad
|
|
118
|
+
credentials, an unknown model, a request the provider rejects) never move to the next attempt. See [Failure handling](failures.md).
|
|
112
119
|
|
|
113
120
|
## Providers and newly released models
|
|
114
121
|
|
|
@@ -116,14 +123,21 @@ RubyLLM looks models up in its bundled registry. To use a model that isn't there
|
|
|
116
123
|
OpenAI or Anthropic model, name its provider next to it. Squishling then tells RubyLLM to assume the model exists:
|
|
117
124
|
|
|
118
125
|
```ruby
|
|
119
|
-
squishling model: "gpt-
|
|
126
|
+
squishling model: "gpt-7-preview", provider: :openai
|
|
120
127
|
squish :triage, model: "claude-haiku-4-5", provider: :anthropic
|
|
121
|
-
Squishling.configure { |c| c.default_model = "gpt-
|
|
128
|
+
Squishling.configure { |c| c.default_model = "gpt-7-preview"; c.default_provider = :openai }
|
|
122
129
|
```
|
|
123
130
|
|
|
124
131
|
A provider is paired with the model or escalation declared at the same level, and applies to that escalation's steps
|
|
125
|
-
that don't name their own. A
|
|
126
|
-
|
|
132
|
+
that don't name their own. A model never inherits a provider meant for a different model: a per-method model doesn't
|
|
133
|
+
take the class's provider, and a subclass that declares its own `model:` or `escalation:` doesn't take its parent's
|
|
134
|
+
(a subclass that only changes params or the purpose keeps the parent's model and provider together). For models that
|
|
135
|
+
are in the registry, a provider is optional and the normal registry lookup is kept.
|
|
136
|
+
|
|
137
|
+
A `provider:` with no model beside it would apply to nothing, so it raises `ConfigurationError`:
|
|
138
|
+
`squish ..., provider: ...` needs a `model:` or `escalation:` in the same call (as `squish!` always has), and
|
|
139
|
+
`squishling provider: ...` needs one in the call or already declared on that same class (not inherited). Set the
|
|
140
|
+
provider next to the model, or, for the configured default model, use `config.default_provider`.
|
|
127
141
|
|
|
128
142
|
## Generation params
|
|
129
143
|
|
|
@@ -164,22 +178,32 @@ end
|
|
|
164
178
|
|
|
165
179
|
- Keys that Squishling or RubyLLM control (`model`, `messages`, `input`, `instructions`, `system`, `stream`,
|
|
166
180
|
`response_format`, `text`, `output_config`, `tools`, `tool_choice`, `schema`, and the camelCase and plural
|
|
167
|
-
spellings other providers use, such as `systemInstruction`, `toolConfig`, `outputConfig`, `inputs`, …)
|
|
168
|
-
`ConfigurationError`, because they would override the model, the conversation, or the strict output format.
|
|
181
|
+
spellings other providers use, such as `systemInstruction`, `toolConfig`/`tool_config`, `outputConfig`, `inputs`, …)
|
|
182
|
+
raise `ConfigurationError`, because they would override the model, the conversation, or the strict output format.
|
|
183
|
+
- **Nested containers** such as Gemini's `generationConfig` (and `generation_config` for Gemini Interactions) and
|
|
184
|
+
Mistral Conversations' `completion_args` accept ordinary settings, but reject the keys inside them that carry the
|
|
185
|
+
output format or tools (`responseMimeType`, `responseSchema`, `responseJsonSchema`, `response_format`, `tools`,
|
|
186
|
+
`tool_choice`, `toolConfig`, and their snake_case spellings). Use `output_schema` instead:
|
|
187
|
+
|
|
188
|
+
```ruby
|
|
189
|
+
squishling params: { generationConfig: { topK: 5 } } # ok
|
|
190
|
+
squishling params: { generationConfig: { responseMimeType: "text/plain" } } # ConfigurationError
|
|
191
|
+
```
|
|
192
|
+
|
|
169
193
|
- If a provider rejects a param, the call raises `Squishling::ConfigurationError` naming the params. It isn't
|
|
170
194
|
retried or sent to `squish_fallback`. See [Failure handling](failures.md).
|
|
171
195
|
|
|
172
196
|
## Inheritance
|
|
173
197
|
|
|
174
|
-
Subclasses inherit the model or escalation
|
|
198
|
+
Subclasses inherit the model or escalation and its provider (as a pair), harness, generation params (merged key by key), purpose, output schema,
|
|
175
199
|
`squish_when` predicate, `squish_context` names, `squish_validate`, `squish_fallback`, and every `squish` declaration. Overrides in a subclass, including
|
|
176
200
|
overridden methods, are routed the same way.
|
|
177
201
|
|
|
178
|
-
`
|
|
179
|
-
`
|
|
202
|
+
`append_to_purpose` sections are added to, not replaced: a subclass's sections follow its parent's, and
|
|
203
|
+
`append_to_purpose false` drops the inherited ones. See [Appending to the purpose](routing.md#appending-to-the-purpose).
|
|
180
204
|
|
|
181
205
|
## Per-call overrides
|
|
182
206
|
|
|
183
|
-
Inside a squished method, `squish!` sends the call to the LLM with its own `
|
|
184
|
-
`
|
|
207
|
+
Inside a squished method, `squish!` sends the call to the LLM with its own `purpose:`,
|
|
208
|
+
`append_to_purpose:`, `context:`, `model:` or `escalation:` (with `provider:`), `params:`, and `harness:`. Each layers over the method and class
|
|
185
209
|
settings the same way they layer over each other. See [Handing off to the LLM](routing.md#handing-off-to-the-llm-with-squish).
|
data/docs/failures.md
CHANGED
|
@@ -15,14 +15,73 @@ A failed attempt moves on to the next one; once the last fails, the error is rai
|
|
|
15
15
|
| Malformed or truncated JSON | Moves to the next attempt, then raises `InvalidOutputError`. JSON wrapped in a markdown code fence is accepted. |
|
|
16
16
|
| JSON that doesn't match the schema (wrong types, missing or extra keys, `null` in a non-`optional` field, a broken [conditional rule](schemas.md#contracts-beyond-the-schema), root not an object) | Moves to the next attempt with the validation errors, then raises `InvalidOutputError` |
|
|
17
17
|
| Schema-valid output that [`squish_validate`](#output-checks-squish_validate) rejects | Moves to the next attempt with your messages, then raises `InvalidOutputError` |
|
|
18
|
+
| With a [squishsum or ensemble harness](harnesses.md): two valid samples that differ, with no judge or a judge that rejects both | Raises `Squishling::DisagreementError` (an `InvalidOutputError`), carrying both typed `candidates` |
|
|
18
19
|
|
|
19
20
|
Squishling validates output itself with [json_schemer](https://github.com/davishmcclurg/json_schemer), because
|
|
20
21
|
RubyLLM doesn't, and some providers don't enforce strict mode. When the next attempt is on the same step (same
|
|
21
|
-
model, provider, and
|
|
22
|
-
different step gets a fresh chat with the original input plus the rejected output
|
|
23
|
-
sent to the next step's provider, which may not be the one that
|
|
24
|
-
|
|
25
|
-
|
|
22
|
+
model, provider, params, and `forward_rejected:`), the re-ask happens in the same conversation, so the model sees
|
|
23
|
+
what it got wrong. A different step gets a fresh chat with the original input plus the rejected output (truncated to
|
|
24
|
+
4,000 characters) and its errors, so that output is sent to the next step's provider, which may not be the one that
|
|
25
|
+
produced it. Set `forward_rejected: false` on a step to leave both out. `InvalidOutputError` exposes `errors`, `raw`
|
|
26
|
+
(the last response), `attempts`, and `models` (the model tried on each attempt); for a
|
|
27
|
+
[`DisagreementError`](harnesses.md#when-the-harness-fails), `raw` holds both candidates, `attempts` is `nil`, and
|
|
28
|
+
`models` has one entry per role (sample a, sample b, then the judge). With a `logger` configured, every
|
|
29
|
+
escalation is logged as a warning.
|
|
30
|
+
|
|
31
|
+
### What ends up in errors and logs
|
|
32
|
+
|
|
33
|
+
Model output can echo sensitive input, so the raw response is kept in one place: `InvalidOutputError#raw` (plus the opt-in `squawk` hook below).
|
|
34
|
+
`InvalidOutputError#message`, `#errors`, and `config.logger` warnings are built from these:
|
|
35
|
+
|
|
36
|
+
- For unparseable JSON, only the position (`response was not valid JSON (at line 1 column 7)`), never the parser's
|
|
37
|
+
snippet of the response.
|
|
38
|
+
- JSON Schema errors, which name the schema path and the rule that failed. They never contain the rejected value, but
|
|
39
|
+
they do name an extra key the model added (`object property at /foo is a disallowed additional property`), and
|
|
40
|
+
that key name is chosen by the model.
|
|
41
|
+
- Your own `squish_validate` messages, verbatim. If you interpolate result values into them, those values reach
|
|
42
|
+
the message, the logger, and the next attempt's prompt.
|
|
43
|
+
- A fixed message when a response contains a value JSON can't represent (`1e400` parses to `Infinity`, or a string
|
|
44
|
+
with invalid UTF-8).
|
|
45
|
+
- At most the first 20 of these errors, plus a count of the rest, so a very large malformed response can't flood
|
|
46
|
+
the retry message or the log.
|
|
47
|
+
- For a [`DisagreementError`](harnesses.md#when-the-harness-fails), only what Squishling wrote: a `:judgment` judge's
|
|
48
|
+
choice and probability. A chat judge's `reason` is model-written, so it is available as `#reason` but never in the
|
|
49
|
+
message or logs.
|
|
50
|
+
|
|
51
|
+
The same applies to anything you log yourself, such as `error.message` in a [fallback](#fallbacks). Use `error.raw`
|
|
52
|
+
only where you are willing to store model output. To send the output of every attempt somewhere on purpose, use
|
|
53
|
+
[`squawk`](#observing-every-attempt-squawk).
|
|
54
|
+
|
|
55
|
+
## Observing every attempt (`squawk`)
|
|
56
|
+
|
|
57
|
+
`squawk` is an opt-in hook for sending what the model returned to an error tracker or tracing tool. It runs after
|
|
58
|
+
every LLM attempt, accepted or rejected, and does nothing unless you set it:
|
|
59
|
+
|
|
60
|
+
```ruby
|
|
61
|
+
Squishling.configure do |config|
|
|
62
|
+
config.squawk = lambda do |output:, metadata:, error:|
|
|
63
|
+
ObservabilitySolution.record(output, metadata, error) if error
|
|
64
|
+
end
|
|
65
|
+
end
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
| Keyword | Value |
|
|
69
|
+
|---|---|
|
|
70
|
+
| `output` | The raw response: a Hash or String, or `nil` when the provider call itself failed |
|
|
71
|
+
| `error` | `nil` for an accepted attempt. Otherwise the `InvalidOutputError` (with `errors` and `raw`) or `LLMError` that ended the attempt |
|
|
72
|
+
| `metadata` | `label` (`"Class#method"`), `attempt`, `attempts` (the escalation's length; under a harness, the sample's or judge's own attempts), `final` (the last attempt), `model`, `provider`, `params`, `input` (the JSON sent: arguments and named context only), `usage` (token counts, when RubyLLM reports them) |
|
|
73
|
+
|
|
74
|
+
- Set it globally with `config.squawk`, per class with `squishling squawk: ...`, or per method with
|
|
75
|
+
`squish :triage, squawk: ...`. The method's hook wins over the class's, which wins over the configured one;
|
|
76
|
+
`squawk: false` silences an inherited hook. Subclasses and `squish!` calls inherit it.
|
|
77
|
+
- Any object that responds to `call` works. The hook receives only the keywords it declares, unless it takes `**`,
|
|
78
|
+
so a lambda that wants just `error:` is fine and new metadata fields won't break it.
|
|
79
|
+
- It runs inline, so keep it quick. Exceptions it raises propagate unchanged and fail the call, like any of your own
|
|
80
|
+
code, so rescue inside the hook if an outage in your tracing tool shouldn't.
|
|
81
|
+
- `output` is the same object the result is built from, so treat it as read-only.
|
|
82
|
+
- Under a [harness](harnesses.md), it runs for every sample and judge attempt too, each with its own `input`.
|
|
83
|
+
- It isn't called for a `ConfigurationError`, for deterministic and fallback returns, or for an attempt where your own
|
|
84
|
+
`squish_validate` raises (that exception propagates first).
|
|
26
85
|
|
|
27
86
|
## Output checks (`squish_validate`)
|
|
28
87
|
|
|
@@ -57,16 +116,17 @@ end
|
|
|
57
116
|
| Error | Raised when |
|
|
58
117
|
|---|---|
|
|
59
118
|
| `Squishling::Error` | Base class for everything below. Raised directly for misuse at call time, e.g. `result` or `squish!` called outside a squished method |
|
|
60
|
-
| `Squishling::ConfigurationError` | Missing
|
|
119
|
+
| `Squishling::ConfigurationError` | Missing purpose or schema, a non-strict schema, invalid or reserved params, an invalid `model:`/`escalation:` declaration, an invalid `append_to_purpose` item or unavailable source, bad credentials, an unknown model, a request the provider rejects (400) |
|
|
61
120
|
| `Squishling::InvalidOutputError` | LLM output still invalid (schema or `squish_validate`) after every attempt in the escalation, or a deterministic/fallback return that doesn't match the schema |
|
|
121
|
+
| `Squishling::DisagreementError` | A subclass of `InvalidOutputError`: a [squishsum or ensemble harness](harnesses.md#when-the-harness-fails)'s samples disagreed and no judge accepted either. `candidates`, `verdict`, and `reason` describe what happened |
|
|
62
122
|
| `Squishling::LLMError` | The provider call failed on the last attempt in the escalation, after RubyLLM's own retries, including context-length errors (`cause` holds the original) |
|
|
63
123
|
|
|
64
124
|
Errors raised by your own Ruby code are not wrapped.
|
|
65
125
|
|
|
66
126
|
## Fallbacks
|
|
67
127
|
|
|
68
|
-
Use `squish_fallback` to decide what happens when
|
|
69
|
-
`LLMError`.
|
|
128
|
+
Use `squish_fallback` to decide what happens when the LLM path fails with `InvalidOutputError` (including a
|
|
129
|
+
`DisagreementError`) or `LLMError`. It's the one fallback for every [harness](harnesses.md).
|
|
70
130
|
It receives the error plus the method's inputs as keywords, and runs against the instance:
|
|
71
131
|
|
|
72
132
|
```ruby
|