squishling 0.1.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +123 -0
- data/README.md +117 -35
- data/docs/configuration.md +113 -32
- data/docs/failures.md +111 -14
- data/docs/harnesses.md +198 -0
- data/docs/measuring-tokens.md +148 -0
- data/docs/naming.md +43 -0
- data/docs/routing.md +46 -31
- data/docs/schemas.md +32 -2
- data/lib/squishling/appendices.rb +3 -3
- data/lib/squishling/class_methods.rb +81 -31
- data/lib/squishling/collisions.rb +54 -0
- data/lib/squishling/configuration.rb +58 -9
- data/lib/squishling/definition.rb +78 -39
- data/lib/squishling/errors.rb +29 -5
- data/lib/squishling/harness.rb +153 -0
- data/lib/squishling/invoker.rb +223 -74
- data/lib/squishling/judge.rb +142 -0
- data/lib/squishling/llm_client.rb +150 -0
- data/lib/squishling/model_path.rb +129 -0
- data/lib/squishling/output_check.rb +109 -0
- data/lib/squishling/params.rb +39 -2
- data/lib/squishling/router.rb +13 -7
- data/lib/squishling/schema.rb +45 -3
- data/lib/squishling/source.rb +5 -5
- data/lib/squishling/squawk.rb +52 -0
- data/lib/squishling/version.rb +1 -1
- data/lib/squishling.rb +25 -5
- metadata +13 -3
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: '0748f552c985521890d365ebc0818ec59c4336f93fb183d1769f49ed8cd32349'
|
|
4
|
+
data.tar.gz: '05128ecf8d7156e0c8ca084bf8a1860b754596017df2cdfdcc35e2242138fa7b'
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 2be443b51d2dc1460a410fbf0850bed09e33fff5f7ad894d4bd92e3f47d428d22df7bfc0371036c090a6b806f466b1d930de3a9caf8d1b044185fd4000b765bf
|
|
7
|
+
data.tar.gz: 3a18d6e041f5807c966677fea5499de8af5d0f28146cb98bbe205141598ee6a01a2d598b1326df3b78227455dcc124c69cf95afacabf9a1689a1cc382a652414
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,129 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.3.0] - 2026-10-09
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- Harnesses: a `harness:` option (on `squishling`, `squish`, `squish!`, and `default_harness` in
|
|
15
|
+
`Squishling.configure`) chooses how a call uses its escalation. `:escalation` is the existing behavior and stays
|
|
16
|
+
the default. `:squishsum` asks the first escalation step twice, concurrently, and accepts the result only when the
|
|
17
|
+
two samples agree (exact match, or a custom `compare:`). `:judged_squishsum` sends a disagreement to a judge: the
|
|
18
|
+
next escalation step or a dedicated `judge:` step, either a chat model or a Jev-style decision model through
|
|
19
|
+
`RubyLLM.judge`, with a default prompt you can override through `judge_instructions:`. Samples that still disagree
|
|
20
|
+
raise the new `Squishling::DisagreementError < InvalidOutputError`, which carries both typed results, and
|
|
21
|
+
`squish_fallback` remains the only fallback. See [Harnesses](docs/harnesses.md). (#16)
|
|
22
|
+
- `:ensemble` and `:judged_ensemble` harnesses: like the squishsum ones, but the two concurrent samples come from the
|
|
23
|
+
escalation's first and second steps (a cross-model check) instead of the first step twice, each retrying within its
|
|
24
|
+
own step's `attempts:`. A judged ensemble's judge defaults to the third step, or `judge:`, and
|
|
25
|
+
`DisagreementError#models` names the model behind each sample. An ensemble needs at least two distinct steps (adjacent
|
|
26
|
+
identical steps count as one step's attempts) or it raises `ConfigurationError` before any request. (#20)
|
|
27
|
+
- `squawk`, an opt-in hook for sending what the model returned to an error tracker or tracing tool. It is called after
|
|
28
|
+
every LLM attempt (every sample and judge attempt under a harness) with `output:`, `metadata:`, and `error:`, and
|
|
29
|
+
can be set on `Squishling.configure`, `squishling`, and `squish`. It does nothing by default, and exceptions it
|
|
30
|
+
raises propagate unchanged. (#15)
|
|
31
|
+
- `forward_rejected:` on escalation steps. When an attempt is rejected and the next step starts, the previous
|
|
32
|
+
output is forwarded to that step's provider by default (unchanged behavior); set `forward_rejected: false` to start
|
|
33
|
+
that step from the original input alone. (#14)
|
|
34
|
+
- `include Squishling` raises `ConfigurationError` when the class inherits a method Squishling would override (for
|
|
35
|
+
example `Sinatra::Base.call`) instead of silently shadowing it. See [Naming and collisions](docs/naming.md). (#17)
|
|
36
|
+
|
|
37
|
+
### Documentation
|
|
38
|
+
|
|
39
|
+
- A new guide, [Measuring token spend](docs/measuring-tokens.md), covers working out what each squished path costs
|
|
40
|
+
(with Coolhand Labs, the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter), and the README is
|
|
41
|
+
reworked around using Squishling to harden code paths as the economics justify it. (#21)
|
|
42
|
+
|
|
43
|
+
### Changed
|
|
44
|
+
|
|
45
|
+
- **Breaking:** the DSL method `instructions` is now `purpose`, and `append_instructions` is now `append_to_purpose`,
|
|
46
|
+
including the `instructions:`/`append_instructions:` keywords on `squishling`, `squish`, and `squish!`. There are no
|
|
47
|
+
aliases. Rename them in your classes: `instructions "..."` becomes `purpose "..."`, and
|
|
48
|
+
`append_instructions: [...]` becomes `append_to_purpose: [...]`. The error for a missing prompt now says
|
|
49
|
+
"has no purpose". (#17)
|
|
50
|
+
- **Breaking:** `ruby_llm` must now be `~> 2.1` (was `~> 2.0`); run `bundle update ruby_llm`. (#16)
|
|
51
|
+
- **Breaking:** every value a method with an `output_schema` returns is now validated and typed, not only Hashes. On
|
|
52
|
+
the Ruby path and in `squish_fallback`, a `nil`, String, Array, or other non-Hash return that used to pass through
|
|
53
|
+
now raises `InvalidOutputError`; return a Hash (or `result(...)`) that matches the schema. A result of the schema's
|
|
54
|
+
own class passes through, another schema's result is re-validated, and values JSON can't represent (`NaN`,
|
|
55
|
+
`Infinity`, cycles) raise `InvalidOutputError` instead of a raw JSON error. Inner calls (`super`, or the method
|
|
56
|
+
called from `squish_when` or a fallback) still only type Hashes, so overrides can reshape values. (#13)
|
|
57
|
+
- **Behavior change:** a provider now travels with the model declared at the same level on a class too. A subclass
|
|
58
|
+
that declares its own `model:` or `escalation:` no longer inherits its parent's `provider:`, which could send the
|
|
59
|
+
whole payload to the parent's provider under another provider's model name. If you relied on that, set `provider:`
|
|
60
|
+
next to the subclass's model.
|
|
61
|
+
- **Behavior change:** a `provider:` with no model beside it now raises `ConfigurationError` instead of being
|
|
62
|
+
silently ignored (the default provider was used). `squish ..., provider: ...` needs a `model:` or `escalation:` in
|
|
63
|
+
the same call, as `squish!` already did; `squishling provider: ...` needs one in the call or already declared on that
|
|
64
|
+
class. Move the `provider:` next to a model, or use `config.default_provider` with the default model.
|
|
65
|
+
- The `result` alias for `squishling_result` is no longer added when the class already has a `result` (its own or
|
|
66
|
+
inherited); use `squishling_result` there. (#17)
|
|
67
|
+
- `InvalidOutputError#message` and `config.logger` warnings no longer quote the model's response. For unparseable JSON
|
|
68
|
+
they report only the line and column; the raw output stays in `InvalidOutputError#raw`. (#15)
|
|
69
|
+
|
|
70
|
+
### Fixed
|
|
71
|
+
|
|
72
|
+
- A number a model writes outside the range JSON can represent (such as `1e400`, which parses to `Infinity`), or a
|
|
73
|
+
string with invalid UTF-8, was accepted as valid, while the same value from a Ruby method was rejected. It is now
|
|
74
|
+
invalid output on the LLM path too, and no longer raises a bare `JSON::GeneratorError` when `:squishsum` compares samples.
|
|
75
|
+
|
|
76
|
+
### Security
|
|
77
|
+
|
|
78
|
+
- Params can no longer replace the strict output format or tools through the containers that also hold ordinary
|
|
79
|
+
settings: `generationConfig`/`generation_config` (Google Gemini) and `completion_args` (Mistral Conversations) now
|
|
80
|
+
reject their format and tool keys, with symbol or string keys, and a non-Hash value for them is rejected.
|
|
81
|
+
`tool_config` is reserved at the top level. Settings such as `topK` still work. (#12)
|
|
82
|
+
- The previous model's rejected output forwarded to the next escalation step is capped at 4,000 characters
|
|
83
|
+
(`Invoker::MAX_FORWARDED_CHARS`), with a truncation marker. (#14)
|
|
84
|
+
- Model output is kept out of error messages and logs (see the `InvalidOutputError#message` change above); `squawk`
|
|
85
|
+
is the one sanctioned way for it to leave the process. (#15) A chat judge's `reason` is model-written, so
|
|
86
|
+
`DisagreementError#message` and the logs no longer include it; read it from `DisagreementError#reason`. A
|
|
87
|
+
`:judgment` judge's message still carries the choice and probability, and a choice other than `a`, `b`, or `neither`
|
|
88
|
+
is no longer repeated.
|
|
89
|
+
- Only the first 20 validation errors (plus a count of the rest) are fed back to the model, logged, and put in
|
|
90
|
+
`InvalidOutputError#message` for one invalid output, so a very large malformed response can no longer produce a
|
|
91
|
+
megabyte-sized retry message or log line.
|
|
92
|
+
|
|
93
|
+
## [0.2.0] - 2026-10-09
|
|
94
|
+
|
|
95
|
+
### Added
|
|
96
|
+
|
|
97
|
+
- Model escalation: `escalation:` (and `default_escalation` in `Squishling.configure`) declares an ordered list of
|
|
98
|
+
steps, each a model name or a Hash with `model:`, `attempts:`, an optional explicit `order:`, `provider:`, and
|
|
99
|
+
`params:`. Invalid output or an `LLMError` moves on to the next attempt. Attempts on the same step continue one
|
|
100
|
+
conversation; a new step starts a fresh chat that is told about the rejected output and its errors.
|
|
101
|
+
`ConfigurationError` never escalates. Works per method, per class, in config, and per `squish!` call; the first
|
|
102
|
+
level that declares a model or escalation wins and levels are never merged. `InvalidOutputError#models` lists the
|
|
103
|
+
model tried on each attempt. (#6)
|
|
104
|
+
- `squish_validate` (and `validate:` on `squish`): Ruby checks on schema-valid LLM output. Return `nil` or `true` to
|
|
105
|
+
accept, or `false`, a String, an Array of Strings, or a dry-validation style result to reject the output and
|
|
106
|
+
trigger the next attempt. Deterministic and fallback returns are not run through it. (#6)
|
|
107
|
+
- Conditional schema rules: Schematist's `given` and `dependent` (JSON Schema `if`/`then`/`else`,
|
|
108
|
+
`dependentRequired`, and `dependentSchemas`) are kept out of the schema sent to the provider, whose strict mode
|
|
109
|
+
doesn't support them, and are still enforced locally on every result. (#6)
|
|
110
|
+
|
|
111
|
+
### Changed
|
|
112
|
+
|
|
113
|
+
- **Breaking:** `config.max_retries` is removed; reading or setting it raises `ConfigurationError` with a migration
|
|
114
|
+
hint. The escalation now decides how many attempts run, and a plain `model:` makes a single attempt. To keep the
|
|
115
|
+
old behavior (two attempts on one model), set
|
|
116
|
+
`config.default_escalation = [{ model: "your-model", attempts: 2 }]`. (#6)
|
|
117
|
+
- `squish!`'s Ruby-to-LLM handoff is now described as "handing off", so "escalation" only means the model list.
|
|
118
|
+
`squish!` also accepts `escalation:` (instead of `model:`) for one call. (#6)
|
|
119
|
+
|
|
120
|
+
### Fixed
|
|
121
|
+
|
|
122
|
+
- A method with an unnamed positional parameter (a destructuring parameter such as `def call((a, b), second)`) sent
|
|
123
|
+
every later argument to the LLM under the wrong name. Each argument now keeps its own name, and unnamed ones are
|
|
124
|
+
sent as `arg0`, `arg1`, and so on.
|
|
125
|
+
|
|
126
|
+
### Security
|
|
127
|
+
|
|
128
|
+
- Params can no longer override the system prompt, tool config, or structured-output format through the camelCase
|
|
129
|
+
and plural request keys used by Google Gemini, Amazon Bedrock Converse, and Mistral Conversations
|
|
130
|
+
(`systemInstruction`, `cachedContent`, `toolConfig`, `outputConfig`, `inputs`); these now raise
|
|
131
|
+
`ConfigurationError` like the other reserved keys.
|
|
132
|
+
|
|
10
133
|
## [0.1.0] - 2026-10-08
|
|
11
134
|
|
|
12
135
|
Initial release.
|
data/README.md
CHANGED
|
@@ -1,46 +1,54 @@
|
|
|
1
|
-
|
|
1
|
+
<h1 align="center">
|
|
2
|
+
<img src="assets/squishling-logo.png" alt="Squishling" width="420">
|
|
3
|
+
</h1>
|
|
2
4
|
|
|
3
5
|
[](https://github.com/Coolhand-Labs/squishling/actions/workflows/ci.yml)
|
|
6
|
+
[](https://badge.fury.io/rb/squishling)
|
|
4
7
|
|
|
5
|
-
|
|
6
|
-
(via [RubyLLM](https://rubyllm.com)), and either way returns the same strict-schema-validated, typed result.
|
|
7
|
-
Callers can't tell the difference.
|
|
8
|
+
**Don't just use AI to write code. Use AI to not need code.** Write code only where it's worth maintaining.
|
|
8
9
|
|
|
9
|
-
|
|
10
|
-
|
|
10
|
+
A squished method either runs its Ruby implementation, with the LLM as a failover, or routes some callers
|
|
11
|
+
through an LLM as a way of eliminating (or discovering) the cost of maintaining them as deterministic code.
|
|
12
|
+
Either way it returns the same strict-schema-validated, typed result, so callers can't tell the difference.
|
|
13
|
+
|
|
14
|
+
Think of code as a cache for AI judgment. Every line you write is something to maintain, so the LLM is the
|
|
15
|
+
default and code has to earn its place: you build it for the inputs that carry the volume, and the LLM keeps
|
|
16
|
+
handling the rest. New to the idea? Read
|
|
17
|
+
[Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
|
|
11
18
|
|
|
12
19
|
## Use cases
|
|
13
20
|
|
|
14
|
-
###
|
|
21
|
+
### Zero-code integrations
|
|
15
22
|
|
|
16
23
|
Accept a new data source today, before anyone writes a parser. Declare what you want back and leave the
|
|
17
24
|
method unimplemented: every call goes to the LLM, and its output is validated against your schema.
|
|
18
25
|
|
|
19
26
|
```ruby
|
|
20
|
-
class
|
|
27
|
+
class ShipmentUpdate
|
|
21
28
|
include Squishling
|
|
22
29
|
|
|
23
|
-
|
|
30
|
+
purpose "Normalize this carrier's tracking webhook into our shipment update."
|
|
24
31
|
output_schema do
|
|
25
|
-
string
|
|
26
|
-
|
|
27
|
-
string
|
|
28
|
-
string :external_id
|
|
32
|
+
string :status, enum: %w[label_created in_transit out_for_delivery delivered exception]
|
|
33
|
+
string :tracking_number
|
|
34
|
+
string :location
|
|
29
35
|
end
|
|
30
36
|
# No `def call` yet, so every webhook goes to the LLM.
|
|
31
37
|
end
|
|
32
38
|
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
event.squished? # => true
|
|
39
|
+
update = ShipmentUpdate.call(carrier: "acme-freight", payload: request.raw_post)
|
|
40
|
+
update.status # => "out_for_delivery"
|
|
41
|
+
update.squished? # => true
|
|
37
42
|
```
|
|
38
43
|
|
|
39
|
-
When one
|
|
40
|
-
`squish_when { |
|
|
44
|
+
When one carrier carries the volume, write `def call` for it and add
|
|
45
|
+
`squish_when { |carrier:, **| carrier != "ups" }`. UPS then runs in Ruby, everything else stays on the
|
|
41
46
|
LLM, and callers don't change. See [Hardening a path](docs/routing.md#hardening-a-path).
|
|
42
47
|
|
|
43
|
-
|
|
48
|
+
Validation guarantees the *shape* of the result, not that it's true. Where a wrong value is costly (money,
|
|
49
|
+
identity), cross-check it against another source or keep that path in Ruby.
|
|
50
|
+
|
|
51
|
+
### Rescue errors your code can't handle yet
|
|
44
52
|
|
|
45
53
|
Keep the Ruby you have for the inputs it understands, and hand the rest to the LLM instead of failing. When the
|
|
46
54
|
parser raises, `squish!` sends this call to the LLM with the error and the parser's own source as context.
|
|
@@ -49,8 +57,8 @@ parser raises, `squish!` sends this call to the LLM with the error and the parse
|
|
|
49
57
|
class InvoiceParser
|
|
50
58
|
include Squishling
|
|
51
59
|
|
|
52
|
-
|
|
53
|
-
|
|
60
|
+
purpose "Extract the invoice fields from the vendor's document."
|
|
61
|
+
append_to_purpose "The Ruby parser that handles well-formed invoices:", self # this class's source
|
|
54
62
|
output_schema do
|
|
55
63
|
string :invoice_number
|
|
56
64
|
number :total
|
|
@@ -60,7 +68,7 @@ class InvoiceParser
|
|
|
60
68
|
invoice = VendorFormats.fetch(vendor).parse(document)
|
|
61
69
|
result(invoice_number: invoice.number, total: invoice.total)
|
|
62
70
|
rescue VendorFormats::ParseError => e
|
|
63
|
-
squish!(
|
|
71
|
+
squish!(append_to_purpose: "The parser failed on this document; the error is in the context.",
|
|
64
72
|
context: { parse_error: e })
|
|
65
73
|
end
|
|
66
74
|
end
|
|
@@ -71,7 +79,30 @@ InvoiceParser.call(vendor: "acme", document: scanned_text).squished? # => true
|
|
|
71
79
|
|
|
72
80
|
Both calls return the same result class. If the LLM can't deliver either, `squish_fallback` decides what to
|
|
73
81
|
return, or the error is raised with the original `ParseError` as its cause. See
|
|
74
|
-
[
|
|
82
|
+
[Handing off to the LLM](docs/routing.md#handing-off-to-the-llm-with-squish).
|
|
83
|
+
|
|
84
|
+
### Measure your tech debt in tokens
|
|
85
|
+
|
|
86
|
+
Code you haven't written is debt you haven't taken on. A squished path has a running cost you can read off a
|
|
87
|
+
meter, so "should we write a parser for this?" becomes arithmetic: what the path costs in tokens each month,
|
|
88
|
+
against what it costs to write and maintain the code. Write it when the first number is bigger.
|
|
89
|
+
|
|
90
|
+
Add and initialize the [`coolhand`](https://github.com/Coolhand-Labs/coolhand-ruby) gem, and it picks up your
|
|
91
|
+
squishling calls automatically and measures their accuracy and cost:
|
|
92
|
+
|
|
93
|
+
```ruby
|
|
94
|
+
# Gemfile
|
|
95
|
+
gem "coolhand"
|
|
96
|
+
|
|
97
|
+
# config/initializers/coolhand.rb
|
|
98
|
+
Coolhand.configure do |config|
|
|
99
|
+
config.api_key = ENV.fetch("COOLHAND_API_KEY")
|
|
100
|
+
end
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
`squished?` tells you which path served a call. Coolhand records your LLM requests and responses; see
|
|
104
|
+
[what each tool sees](docs/measuring-tokens.md#privacy). Want to roll your own metrics? Check out
|
|
105
|
+
[our guide](docs/measuring-tokens.md) for doing it with other tools.
|
|
75
106
|
|
|
76
107
|
## Installation
|
|
77
108
|
|
|
@@ -79,18 +110,19 @@ return, or the error is raised with the original `ParseError` as its cause. See
|
|
|
79
110
|
gem "squishling"
|
|
80
111
|
```
|
|
81
112
|
|
|
82
|
-
Requires Ruby 3.3+ and RubyLLM 2.
|
|
83
|
-
universal model:
|
|
113
|
+
Requires Ruby 3.3+ and RubyLLM 2.1+. Configure your provider API keys in RubyLLM as usual, then optionally set a
|
|
114
|
+
universal model, or an escalation of models to try in order:
|
|
84
115
|
|
|
85
116
|
```ruby
|
|
86
117
|
Squishling.configure do |config|
|
|
87
118
|
config.default_model = "claude-sonnet-5-5" # falls back to RubyLLM's default when nil
|
|
119
|
+
# or: config.default_escalation = [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5", "claude-opus-5-5"]
|
|
88
120
|
end
|
|
89
121
|
```
|
|
90
122
|
|
|
91
123
|
## How it works
|
|
92
124
|
|
|
93
|
-
A squishling class needs **
|
|
125
|
+
A squishling class needs a **purpose** (the system prompt) and an **output schema** (the shape of the result,
|
|
94
126
|
validated on both paths). A squished call goes to the LLM when:
|
|
95
127
|
|
|
96
128
|
- its `squish_when` predicate is truthy for these inputs,
|
|
@@ -103,23 +135,62 @@ Otherwise the Ruby runs, and whatever it returns is validated and typed like LLM
|
|
|
103
135
|
|
|
104
136
|
- **One contract, two paths**: Ruby returns and LLM output are validated against the same strict schema and
|
|
105
137
|
returned as the same typed `Data` objects. `squished?` tells you which path served a call.
|
|
106
|
-
- **Your code as context**: `
|
|
138
|
+
- **Your code as context**: `append_to_purpose` adds sections to the prompt, including a class's or method's
|
|
107
139
|
own Ruby source.
|
|
108
140
|
- **Any RubyLLM provider and model**: OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, OpenRouter, and more.
|
|
109
141
|
Set a universal, per-class, per-method, or per-call model, plus layered generation params (temperature,
|
|
110
142
|
reasoning effort, top_p, …).
|
|
111
|
-
- **
|
|
112
|
-
|
|
143
|
+
- **Model escalation**: declare an `escalation:` instead of a `model:`, with per-step `attempts:`, `order:`, provider,
|
|
144
|
+
params, and `forward_rejected:`. Invalid output (with its errors fed back) or a provider failure moves on to the next attempt. Steps can
|
|
145
|
+
cross providers, e.g. a local Qwen model on Ollama, then Anthropic Claude Haiku on AWS Bedrock, then Claude Opus:
|
|
146
|
+
|
|
147
|
+
```ruby
|
|
148
|
+
RubyLLM.configure do |config|
|
|
149
|
+
config.ollama_api_base = "http://localhost:11434/v1"
|
|
150
|
+
config.bedrock_api_key = ENV["AWS_ACCESS_KEY_ID"]
|
|
151
|
+
config.bedrock_secret_key = ENV["AWS_SECRET_ACCESS_KEY"]
|
|
152
|
+
config.bedrock_region = "us-east-1"
|
|
153
|
+
config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
|
|
154
|
+
end
|
|
155
|
+
|
|
156
|
+
class TicketTriager
|
|
157
|
+
include Squishling
|
|
158
|
+
|
|
159
|
+
squishling escalation: [
|
|
160
|
+
{ model: "qwen3:8b", provider: :ollama, attempts: 2 }, # local and free, tried twice
|
|
161
|
+
{ model: "claude-haiku-4-5", provider: :bedrock }, # small hosted model
|
|
162
|
+
{ model: "claude-opus-5-5", provider: :anthropic } # last resort
|
|
163
|
+
]
|
|
164
|
+
# purpose, output_schema, ...
|
|
165
|
+
end
|
|
166
|
+
```
|
|
167
|
+
- **[Harnesses](docs/harnesses.md)**: choose how a call uses its escalation with `harness:`
|
|
168
|
+
- [`:escalation`](docs/configuration.md#models-and-escalation) (default): try each attempt until one passes
|
|
169
|
+
- [`:squishsum`](docs/harnesses.md#samples): two concurrent samples, accepted only if they agree
|
|
170
|
+
- [`:judged_squishsum`](docs/harnesses.md#the-judge): a chat or Jev judge picks between disagreeing samples
|
|
171
|
+
- [`:ensemble` and `:judged_ensemble`](docs/harnesses.md#ensembles): the same checks across two different models
|
|
172
|
+
- **Output contracts**: beyond the strict schema, conditional rules (`given`) and Ruby checks (`squish_validate`)
|
|
173
|
+
reject bad output and trigger the next attempt.
|
|
174
|
+
- **Defined failure behavior**: provider errors become `Squishling::LLMError`, and `squish_fallback` lets you decide
|
|
175
|
+
what to return when every attempt fails.
|
|
176
|
+
- **Observability**: the opt-in [`squawk`](docs/failures.md#observing-every-attempt-squawk) hook sends every attempt's
|
|
177
|
+
raw output to your error tracker. Model output stays out of error messages and logs otherwise.
|
|
113
178
|
- **Opt-in context**: only method arguments, the instance state you name with `squish_context`, the `context:` you
|
|
114
|
-
pass to `squish!`, and the source you choose to append are sent to the provider.
|
|
179
|
+
pass to `squish!`, and the source you choose to append are sent to the provider. The one addition: when
|
|
180
|
+
[escalation](docs/configuration.md#models-and-escalation) moves to the next step, that step also sees the previous
|
|
181
|
+
model's rejected output, unless the step sets `forward_rejected: false`.
|
|
115
182
|
|
|
116
183
|
## Documentation
|
|
117
184
|
|
|
118
|
-
- [Configuration](docs/configuration.md): options,
|
|
119
|
-
- [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `
|
|
185
|
+
- [Configuration](docs/configuration.md): options, models and escalation, providers, generation params, inheritance
|
|
186
|
+
- [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `append_to_purpose`, hardening a path,
|
|
120
187
|
what the LLM sees
|
|
121
|
-
- [Output schemas](docs/schemas.md): schema forms, strict mode, typed results, optional vs. empty
|
|
122
|
-
- [Failure handling](docs/failures.md):
|
|
188
|
+
- [Output schemas](docs/schemas.md): schema forms, strict mode, typed results, optional vs. empty, contracts
|
|
189
|
+
- [Failure handling](docs/failures.md): escalation, `squish_validate`, error classes, fallbacks
|
|
190
|
+
- [Harnesses](docs/harnesses.md): escalation, squishsum, ensemble, and judges (chat or Jev)
|
|
191
|
+
- [Naming and collisions](docs/naming.md): the methods `include Squishling` adds and what happens when a name is taken
|
|
192
|
+
- [Measuring token spend](docs/measuring-tokens.md): see what each squished path costs, with Coolhand Labs,
|
|
193
|
+
the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter
|
|
123
194
|
- [Live examples](examples/README.md): end-to-end tests against Anthropic Claude Haiku and OpenAI GPT-6 Luna
|
|
124
195
|
|
|
125
196
|
## Development
|
|
@@ -131,6 +202,17 @@ bundle exec rake # RSpec (offline; RubyLLM is stubbed) + RuboCop
|
|
|
131
202
|
|
|
132
203
|
See [AGENTS.md](AGENTS.md) for repo conventions, and [SECURITY.md](SECURITY.md) for reporting vulnerabilities.
|
|
133
204
|
|
|
205
|
+
## Credits
|
|
206
|
+
|
|
207
|
+
Inspired by [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
|
|
208
|
+
|
|
134
209
|
## License
|
|
135
210
|
|
|
136
211
|
Apache-2.0
|
|
212
|
+
|
|
213
|
+
---
|
|
214
|
+
|
|
215
|
+
<p align="center">
|
|
216
|
+
This open source project is supported by <a href="https://coolhandlabs.com/">Coolhand Labs</a>.<br><br>
|
|
217
|
+
<a href="https://coolhandlabs.com/"><img src="assets/coolhand-labs.png" alt="Coolhand Labs" width="220"></a>
|
|
218
|
+
</p>
|
data/docs/configuration.md
CHANGED
|
@@ -14,65 +14,135 @@ Then configure Squishling:
|
|
|
14
14
|
|
|
15
15
|
```ruby
|
|
16
16
|
Squishling.configure do |config|
|
|
17
|
-
config.default_model = "claude-sonnet-5-5"
|
|
17
|
+
config.default_model = "claude-sonnet-5-5" # or default_escalation, below
|
|
18
18
|
config.default_provider = nil
|
|
19
19
|
config.default_params = { temperature: 0 }
|
|
20
|
-
config.
|
|
20
|
+
config.default_harness = :escalation
|
|
21
21
|
config.logger = Rails.logger
|
|
22
22
|
end
|
|
23
23
|
```
|
|
24
24
|
|
|
25
25
|
| Option | Default | Description |
|
|
26
26
|
|---|---|---|
|
|
27
|
-
| `default_model` | `nil` | The universal model for every squishling class that doesn't declare one. `nil` uses RubyLLM's `default_model`. |
|
|
28
|
-
| `
|
|
29
|
-
| `
|
|
30
|
-
| `
|
|
31
|
-
| `
|
|
27
|
+
| `default_model` | `nil` | The universal model (a single attempt) for every squishling class that doesn't declare one. `nil` uses RubyLLM's `default_model`. |
|
|
28
|
+
| `default_escalation` | `nil` | The universal [escalation](#models-and-escalation): models tried in order. Assigning it clears `default_model`, and vice versa. |
|
|
29
|
+
| `default_provider` | `nil` | The provider for `default_model`/`default_escalation` steps that don't name one. Only needed for models missing from RubyLLM's registry (see below). |
|
|
30
|
+
| `default_params` | `{}` | Generation params for every call (temperature, reasoning effort, top_p, …), overridable per class, per method, and per escalation step. See [Generation params](#generation-params). |
|
|
31
|
+
| `squawk` | `nil` | A callable run after every LLM attempt with the raw output, for sending it to an error tracker or tracing tool. Also settable per class and per method. See [Observing every attempt](failures.md#observing-every-attempt-squawk). |
|
|
32
|
+
| `default_harness` | `nil` (`:escalation`) | How every squishling class that doesn't declare its own uses its escalation: `:escalation`, `:squishsum`, `:judged_squishsum`, `:ensemble`, `:judged_ensemble`, or a Hash with `type:` and options. See [Harnesses](harnesses.md). |
|
|
33
|
+
| `logger` | `nil` | Any `Logger`. Debug lines when a call routes to the LLM, warnings on each escalation and when a fallback is used. Never includes the raw model response; see [what ends up in errors and logs](failures.md#what-ends-up-in-errors-and-logs). |
|
|
32
34
|
|
|
33
35
|
Transport-level retries (rate limits, 5xx, timeouts) are configured on RubyLLM itself
|
|
34
36
|
(`RubyLLM.config.max_retries`, `request_timeout`).
|
|
35
37
|
|
|
36
|
-
|
|
38
|
+
> **`max_retries` was removed.** The escalation now decides how many attempts run, and setting
|
|
39
|
+
> `config.max_retries` raises `ConfigurationError`. The old default (`max_retries = 1`, two attempts on one model) is
|
|
40
|
+
> `default_escalation = [{ model: "your-model", attempts: 2 }]`. A plain `model:` now makes one attempt.
|
|
37
41
|
|
|
38
|
-
|
|
42
|
+
## Models and escalation
|
|
39
43
|
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
44
|
+
Declare **either** a `model:`, one model with a single attempt, **or** an `escalation:`, the models to try in
|
|
45
|
+
order. When an attempt fails (invalid output, a [`squish_validate`](failures.md#output-checks-squish_validate)
|
|
46
|
+
rejection, or a provider failure), the call moves on to the next attempt. It stops at the first valid output, or
|
|
47
|
+
raises once the last attempt fails.
|
|
48
|
+
|
|
49
|
+
```ruby
|
|
50
|
+
Squishling.configure do |config|
|
|
51
|
+
config.default_escalation = [
|
|
52
|
+
{ model: "claude-haiku-4-5", attempts: 2 }, # attempts 1-2: same conversation, told what was wrong
|
|
53
|
+
"claude-sonnet-5-5", # attempt 3: fresh chat, shown Haiku's last output and errors
|
|
54
|
+
"claude-opus-5-5" # attempt 4
|
|
55
|
+
]
|
|
56
|
+
end
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Each step is a model name, or a Hash with:
|
|
60
|
+
|
|
61
|
+
| Key | Default | Description |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| `model:` | required | The model id. |
|
|
64
|
+
| `attempts:` | `1` | How many attempts this step gets before escalating to the next one. |
|
|
65
|
+
| `order:` | position | An Integer; steps run lowest first. Give every step an `order:` or none (duplicates are rejected). Gaps are fine. |
|
|
66
|
+
| `provider:` | the level's `provider:` | See [Providers](#providers-and-newly-released-models). |
|
|
67
|
+
| `forward_rejected:` | `true` | Set `false` to start this step from the original input alone, without the previous step's rejected output and its errors (see below). |
|
|
68
|
+
| `params:` | `{}` | [Generation params](#generation-params) for this step, merged over the class and method params key by key. |
|
|
69
|
+
|
|
70
|
+
Without `order:`, steps run in the order written. With it, the order is explicit and doesn't depend on position:
|
|
71
|
+
|
|
72
|
+
```ruby
|
|
73
|
+
squishling escalation: [
|
|
74
|
+
{ model: "claude-opus-5-5", order: 20, params: { thinking: { effort: :high } } },
|
|
75
|
+
{ model: "claude-haiku-4-5", order: 0, attempts: 2 },
|
|
76
|
+
{ model: "claude-sonnet-5-5", order: 10 }
|
|
77
|
+
]
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
`params:` on a step let one step switch to a reasoning model that rejects sampling params:
|
|
81
|
+
|
|
82
|
+
```ruby
|
|
83
|
+
squishling params: { temperature: 0 },
|
|
84
|
+
escalation: ["gpt-5-mini", { model: "gpt-6-luna", provider: :openai,
|
|
85
|
+
params: { temperature: nil, thinking: { effort: :high } } }]
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
### Which model or escalation applies
|
|
89
|
+
|
|
90
|
+
The first level that declares a `model:` or an `escalation:` supplies the whole thing (levels aren't merged):
|
|
91
|
+
|
|
92
|
+
1. per call: `squish!(model: ...)` or `squish!(escalation: ...)` (see [Per-call overrides](#per-call-overrides))
|
|
93
|
+
2. per method: `squish :name, model: ...` or `escalation: ...`
|
|
94
|
+
3. per class: `squishling model: ...` or `escalation: ...` (inherited by subclasses)
|
|
95
|
+
4. universal: `Squishling.config.default_model` or `default_escalation`
|
|
96
|
+
5. RubyLLM's `default_model` (a single attempt)
|
|
45
97
|
|
|
46
98
|
```ruby
|
|
47
99
|
class InvoiceParser
|
|
48
100
|
include Squishling
|
|
49
|
-
squishling
|
|
101
|
+
squishling escalation: %w[claude-sonnet-5-5 claude-opus-5-5] # class escalation
|
|
50
102
|
|
|
51
|
-
squish :classify, model: "claude-haiku-4-5" do
|
|
103
|
+
squish :classify, model: "claude-haiku-4-5" do # one cheap attempt for this method
|
|
52
104
|
string :category
|
|
53
105
|
end
|
|
54
106
|
end
|
|
55
107
|
```
|
|
56
108
|
|
|
109
|
+
Passing both `model:` and `escalation:` in one declaration raises `ConfigurationError`, and so does a list passed as
|
|
110
|
+
`model:`. Across separate declarations (a reopened class, or two config assignments), the latest one wins.
|
|
111
|
+
|
|
112
|
+
Consecutive attempts of the same step (same model, provider, params, and `forward_rejected:`, e.g. `attempts: 2`)
|
|
113
|
+
continue one conversation. Moving to a different step starts a fresh chat with the original input plus the previous
|
|
114
|
+
output and why it was rejected, so no provider-specific history crosses providers. That rejected output is
|
|
115
|
+
model-generated and is sent to the next step's provider, even when it is a different one, truncated to 4,000
|
|
116
|
+
characters. If you don't want a step to see it (say, because the next step is a different provider), set
|
|
117
|
+
`forward_rejected: false` on that step and it starts from the original input only. Configuration errors (bad
|
|
118
|
+
credentials, an unknown model, a request the provider rejects) never move to the next attempt. See [Failure handling](failures.md).
|
|
119
|
+
|
|
57
120
|
## Providers and newly released models
|
|
58
121
|
|
|
59
122
|
RubyLLM looks models up in its bundled registry. To use a model that isn't there yet, such as a newly released
|
|
60
123
|
OpenAI or Anthropic model, name its provider next to it. Squishling then tells RubyLLM to assume the model exists:
|
|
61
124
|
|
|
62
125
|
```ruby
|
|
63
|
-
squishling model: "gpt-
|
|
126
|
+
squishling model: "gpt-7-preview", provider: :openai
|
|
64
127
|
squish :triage, model: "claude-haiku-4-5", provider: :anthropic
|
|
65
|
-
Squishling.configure { |c| c.default_model = "gpt-
|
|
128
|
+
Squishling.configure { |c| c.default_model = "gpt-7-preview"; c.default_provider = :openai }
|
|
66
129
|
```
|
|
67
130
|
|
|
68
|
-
A provider is paired with the model declared at the same level,
|
|
69
|
-
|
|
70
|
-
|
|
131
|
+
A provider is paired with the model or escalation declared at the same level, and applies to that escalation's steps
|
|
132
|
+
that don't name their own. A model never inherits a provider meant for a different model: a per-method model doesn't
|
|
133
|
+
take the class's provider, and a subclass that declares its own `model:` or `escalation:` doesn't take its parent's
|
|
134
|
+
(a subclass that only changes params or the purpose keeps the parent's model and provider together). For models that
|
|
135
|
+
are in the registry, a provider is optional and the normal registry lookup is kept.
|
|
136
|
+
|
|
137
|
+
A `provider:` with no model beside it would apply to nothing, so it raises `ConfigurationError`:
|
|
138
|
+
`squish ..., provider: ...` needs a `model:` or `escalation:` in the same call (as `squish!` always has), and
|
|
139
|
+
`squishling provider: ...` needs one in the call or already declared on that same class (not inherited). Set the
|
|
140
|
+
provider next to the model, or, for the configured default model, use `config.default_provider`.
|
|
71
141
|
|
|
72
142
|
## Generation params
|
|
73
143
|
|
|
74
|
-
`params` holds generation settings. Set them at any of three levels
|
|
75
|
-
**key by key**:
|
|
144
|
+
`params` holds generation settings. Set them at any of three levels, plus per step in an
|
|
145
|
+
[escalation](#models-and-escalation); each level overrides the one above it **key by key**:
|
|
76
146
|
|
|
77
147
|
```ruby
|
|
78
148
|
Squishling.configure { |c| c.default_params = { temperature: 0 } } # every call
|
|
@@ -107,22 +177,33 @@ end
|
|
|
107
177
|
```
|
|
108
178
|
|
|
109
179
|
- Keys that Squishling or RubyLLM control (`model`, `messages`, `input`, `instructions`, `system`, `stream`,
|
|
110
|
-
`response_format`, `text`, `output_config`, `tools`, `tool_choice`, `schema`,
|
|
111
|
-
|
|
180
|
+
`response_format`, `text`, `output_config`, `tools`, `tool_choice`, `schema`, and the camelCase and plural
|
|
181
|
+
spellings other providers use, such as `systemInstruction`, `toolConfig`/`tool_config`, `outputConfig`, `inputs`, …)
|
|
182
|
+
raise `ConfigurationError`, because they would override the model, the conversation, or the strict output format.
|
|
183
|
+
- **Nested containers** such as Gemini's `generationConfig` (and `generation_config` for Gemini Interactions) and
|
|
184
|
+
Mistral Conversations' `completion_args` accept ordinary settings, but reject the keys inside them that carry the
|
|
185
|
+
output format or tools (`responseMimeType`, `responseSchema`, `responseJsonSchema`, `response_format`, `tools`,
|
|
186
|
+
`tool_choice`, `toolConfig`, and their snake_case spellings). Use `output_schema` instead:
|
|
187
|
+
|
|
188
|
+
```ruby
|
|
189
|
+
squishling params: { generationConfig: { topK: 5 } } # ok
|
|
190
|
+
squishling params: { generationConfig: { responseMimeType: "text/plain" } } # ConfigurationError
|
|
191
|
+
```
|
|
192
|
+
|
|
112
193
|
- If a provider rejects a param, the call raises `Squishling::ConfigurationError` naming the params. It isn't
|
|
113
194
|
retried or sent to `squish_fallback`. See [Failure handling](failures.md).
|
|
114
195
|
|
|
115
196
|
## Inheritance
|
|
116
197
|
|
|
117
|
-
Subclasses inherit the model
|
|
118
|
-
`squish_when` predicate, `squish_context` names, `squish_fallback`, and every `squish` declaration. Overrides in a subclass, including
|
|
198
|
+
Subclasses inherit the model or escalation and its provider (as a pair), harness, generation params (merged key by key), purpose, output schema,
|
|
199
|
+
`squish_when` predicate, `squish_context` names, `squish_validate`, `squish_fallback`, and every `squish` declaration. Overrides in a subclass, including
|
|
119
200
|
overridden methods, are routed the same way.
|
|
120
201
|
|
|
121
|
-
`
|
|
122
|
-
`
|
|
202
|
+
`append_to_purpose` sections are added to, not replaced: a subclass's sections follow its parent's, and
|
|
203
|
+
`append_to_purpose false` drops the inherited ones. See [Appending to the purpose](routing.md#appending-to-the-purpose).
|
|
123
204
|
|
|
124
205
|
## Per-call overrides
|
|
125
206
|
|
|
126
|
-
Inside a squished method, `squish!` sends the call to the LLM with its own `
|
|
127
|
-
`
|
|
128
|
-
settings the same way they layer over each other. See [
|
|
207
|
+
Inside a squished method, `squish!` sends the call to the LLM with its own `purpose:`,
|
|
208
|
+
`append_to_purpose:`, `context:`, `model:` or `escalation:` (with `provider:`), `params:`, and `harness:`. Each layers over the method and class
|
|
209
|
+
settings the same way they layer over each other. See [Handing off to the LLM](routing.md#handing-off-to-the-llm-with-squish).
|