squishling 0.1.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 76e867f7a93f81c06fb80016fcbeb1e1db39a064f1b8bb1e7cd98194bbf6ef6e
4
- data.tar.gz: 2c2ea6a4994f4fb764337c9f194c26eb264cea7a7ecae3e881fa3a4d75303e83
3
+ metadata.gz: '0748f552c985521890d365ebc0818ec59c4336f93fb183d1769f49ed8cd32349'
4
+ data.tar.gz: '05128ecf8d7156e0c8ca084bf8a1860b754596017df2cdfdcc35e2242138fa7b'
5
5
  SHA512:
6
- metadata.gz: ac0484e5a496beb3c3e27a51d13765c8b62c243d2253eacd12c1b532017e141780f723df5164bcaa7548e37b28f27fa901ff493fb343a34426ffafe120d96bde
7
- data.tar.gz: 483a5344615d1bc0b1cbb9f3608e4f48df8891eb05dfcb7ad9ac31f22fb3574392f3621d5227037f0a28b48e5389f6e375ef336181ccfd9b3c8836969856cae6
6
+ metadata.gz: 2be443b51d2dc1460a410fbf0850bed09e33fff5f7ad894d4bd92e3f47d428d22df7bfc0371036c090a6b806f466b1d930de3a9caf8d1b044185fd4000b765bf
7
+ data.tar.gz: 3a18d6e041f5807c966677fea5499de8af5d0f28146cb98bbe205141598ee6a01a2d598b1326df3b78227455dcc124c69cf95afacabf9a1689a1cc382a652414
data/CHANGELOG.md CHANGED
@@ -7,6 +7,129 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.3.0] - 2026-10-09
11
+
12
+ ### Added
13
+
14
+ - Harnesses: a `harness:` option (on `squishling`, `squish`, `squish!`, and `default_harness` in
15
+ `Squishling.configure`) chooses how a call uses its escalation. `:escalation` is the existing behavior and stays
16
+ the default. `:squishsum` asks the first escalation step twice, concurrently, and accepts the result only when the
17
+ two samples agree (exact match, or a custom `compare:`). `:judged_squishsum` sends a disagreement to a judge: the
18
+ next escalation step or a dedicated `judge:` step, either a chat model or a Jev-style decision model through
19
+ `RubyLLM.judge`, with a default prompt you can override through `judge_instructions:`. Samples that still disagree
20
+ raise the new `Squishling::DisagreementError < InvalidOutputError`, which carries both typed results, and
21
+ `squish_fallback` remains the only fallback. See [Harnesses](docs/harnesses.md). (#16)
22
+ - `:ensemble` and `:judged_ensemble` harnesses: like the squishsum ones, but the two concurrent samples come from the
23
+ escalation's first and second steps (a cross-model check) instead of the first step twice, each retrying within its
24
+ own step's `attempts:`. A judged ensemble's judge defaults to the third step, or `judge:`, and
25
+ `DisagreementError#models` names the model behind each sample. An ensemble needs at least two distinct steps (adjacent
26
+ identical steps count as one step's attempts) or it raises `ConfigurationError` before any request. (#20)
27
+ - `squawk`, an opt-in hook for sending what the model returned to an error tracker or tracing tool. It is called after
28
+ every LLM attempt (every sample and judge attempt under a harness) with `output:`, `metadata:`, and `error:`, and
29
+ can be set on `Squishling.configure`, `squishling`, and `squish`. It does nothing by default, and exceptions it
30
+ raises propagate unchanged. (#15)
31
+ - `forward_rejected:` on escalation steps. When an attempt is rejected and the next step starts, the previous
32
+ output is forwarded to that step's provider by default (unchanged behavior); set `forward_rejected: false` to start
33
+ that step from the original input alone. (#14)
34
+ - `include Squishling` raises `ConfigurationError` when the class inherits a method Squishling would override (for
35
+ example `Sinatra::Base.call`) instead of silently shadowing it. See [Naming and collisions](docs/naming.md). (#17)
36
+
37
+ ### Documentation
38
+
39
+ - A new guide, [Measuring token spend](docs/measuring-tokens.md), covers working out what each squished path costs
40
+ (with Coolhand Labs, the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter), and the README is
41
+ reworked around using Squishling to harden code paths as the economics justify it. (#21)
42
+
43
+ ### Changed
44
+
45
+ - **Breaking:** the DSL method `instructions` is now `purpose`, and `append_instructions` is now `append_to_purpose`,
46
+ including the `instructions:`/`append_instructions:` keywords on `squishling`, `squish`, and `squish!`. There are no
47
+ aliases. Rename them in your classes: `instructions "..."` becomes `purpose "..."`, and
48
+ `append_instructions: [...]` becomes `append_to_purpose: [...]`. The error for a missing prompt now says
49
+ "has no purpose". (#17)
50
+ - **Breaking:** `ruby_llm` must now be `~> 2.1` (was `~> 2.0`); run `bundle update ruby_llm`. (#16)
51
+ - **Breaking:** every value a method with an `output_schema` returns is now validated and typed, not only Hashes. On
52
+ the Ruby path and in `squish_fallback`, a `nil`, String, Array, or other non-Hash return that used to pass through
53
+ now raises `InvalidOutputError`; return a Hash (or `result(...)`) that matches the schema. A result of the schema's
54
+ own class passes through, another schema's result is re-validated, and values JSON can't represent (`NaN`,
55
+ `Infinity`, cycles) raise `InvalidOutputError` instead of a raw JSON error. Inner calls (`super`, or the method
56
+ called from `squish_when` or a fallback) still only type Hashes, so overrides can reshape values. (#13)
57
+ - **Behavior change:** a provider now travels with the model declared at the same level on a class too. A subclass
58
+ that declares its own `model:` or `escalation:` no longer inherits its parent's `provider:`, which could send the
59
+ whole payload to the parent's provider under another provider's model name. If you relied on that, set `provider:`
60
+ next to the subclass's model.
61
+ - **Behavior change:** a `provider:` with no model beside it now raises `ConfigurationError` instead of being
62
+ silently ignored (the default provider was used). `squish ..., provider: ...` needs a `model:` or `escalation:` in
63
+ the same call, as `squish!` already did; `squishling provider: ...` needs one in the call or already declared on that
64
+ class. Move the `provider:` next to a model, or use `config.default_provider` with the default model.
65
+ - The `result` alias for `squishling_result` is no longer added when the class already has a `result` (its own or
66
+ inherited); use `squishling_result` there. (#17)
67
+ - `InvalidOutputError#message` and `config.logger` warnings no longer quote the model's response. For unparseable JSON
68
+ they report only the line and column; the raw output stays in `InvalidOutputError#raw`. (#15)
69
+
70
+ ### Fixed
71
+
72
+ - A number a model writes outside the range JSON can represent (such as `1e400`, which parses to `Infinity`), or a
73
+ string with invalid UTF-8, was accepted as valid, while the same value from a Ruby method was rejected. It is now
74
+ invalid output on the LLM path too, and no longer raises a bare `JSON::GeneratorError` when `:squishsum` compares samples.
75
+
76
+ ### Security
77
+
78
+ - Params can no longer replace the strict output format or tools through the containers that also hold ordinary
79
+ settings: `generationConfig`/`generation_config` (Google Gemini) and `completion_args` (Mistral Conversations) now
80
+ reject their format and tool keys, with symbol or string keys, and a non-Hash value for them is rejected.
81
+ `tool_config` is reserved at the top level. Settings such as `topK` still work. (#12)
82
+ - The previous model's rejected output forwarded to the next escalation step is capped at 4,000 characters
83
+ (`Invoker::MAX_FORWARDED_CHARS`), with a truncation marker. (#14)
84
+ - Model output is kept out of error messages and logs (see the `InvalidOutputError#message` change above); `squawk`
85
+ is the one sanctioned way for it to leave the process. (#15) A chat judge's `reason` is model-written, so
86
+ `DisagreementError#message` and the logs no longer include it; read it from `DisagreementError#reason`. A
87
+ `:judgment` judge's message still carries the choice and probability, and a choice other than `a`, `b`, or `neither`
88
+ is no longer repeated.
89
+ - Only the first 20 validation errors (plus a count of the rest) are fed back to the model, logged, and put in
90
+ `InvalidOutputError#message` for one invalid output, so a very large malformed response can no longer produce a
91
+ megabyte-sized retry message or log line.
92
+
93
+ ## [0.2.0] - 2026-10-09
94
+
95
+ ### Added
96
+
97
+ - Model escalation: `escalation:` (and `default_escalation` in `Squishling.configure`) declares an ordered list of
98
+ steps, each a model name or a Hash with `model:`, `attempts:`, an optional explicit `order:`, `provider:`, and
99
+ `params:`. Invalid output or an `LLMError` moves on to the next attempt. Attempts on the same step continue one
100
+ conversation; a new step starts a fresh chat that is told about the rejected output and its errors.
101
+ `ConfigurationError` never escalates. Works per method, per class, in config, and per `squish!` call; the first
102
+ level that declares a model or escalation wins and levels are never merged. `InvalidOutputError#models` lists the
103
+ model tried on each attempt. (#6)
104
+ - `squish_validate` (and `validate:` on `squish`): Ruby checks on schema-valid LLM output. Return `nil` or `true` to
105
+ accept, or `false`, a String, an Array of Strings, or a dry-validation style result to reject the output and
106
+ trigger the next attempt. Deterministic and fallback returns are not run through it. (#6)
107
+ - Conditional schema rules: Schematist's `given` and `dependent` (JSON Schema `if`/`then`/`else`,
108
+ `dependentRequired`, and `dependentSchemas`) are kept out of the schema sent to the provider, whose strict mode
109
+ doesn't support them, and are still enforced locally on every result. (#6)
110
+
111
+ ### Changed
112
+
113
+ - **Breaking:** `config.max_retries` is removed; reading or setting it raises `ConfigurationError` with a migration
114
+ hint. The escalation now decides how many attempts run, and a plain `model:` makes a single attempt. To keep the
115
+ old behavior (two attempts on one model), set
116
+ `config.default_escalation = [{ model: "your-model", attempts: 2 }]`. (#6)
117
+ - `squish!`'s Ruby-to-LLM handoff is now described as "handing off", so "escalation" only means the model list.
118
+ `squish!` also accepts `escalation:` (instead of `model:`) for one call. (#6)
119
+
120
+ ### Fixed
121
+
122
+ - A method with an unnamed positional parameter (a destructuring parameter such as `def call((a, b), second)`) sent
123
+ every later argument to the LLM under the wrong name. Each argument now keeps its own name, and unnamed ones are
124
+ sent as `arg0`, `arg1`, and so on.
125
+
126
+ ### Security
127
+
128
+ - Params can no longer override the system prompt, tool config, or structured-output format through the camelCase
129
+ and plural request keys used by Google Gemini, Amazon Bedrock Converse, and Mistral Conversations
130
+ (`systemInstruction`, `cachedContent`, `toolConfig`, `outputConfig`, `inputs`); these now raise
131
+ `ConfigurationError` like the other reserved keys.
132
+
10
133
  ## [0.1.0] - 2026-10-08
11
134
 
12
135
  Initial release.
data/README.md CHANGED
@@ -1,46 +1,54 @@
1
- # Squishling
1
+ <h1 align="center">
2
+ <img src="assets/squishling-logo.png" alt="Squishling" width="420">
3
+ </h1>
2
4
 
3
5
  [![CI](https://github.com/Coolhand-Labs/squishling/actions/workflows/ci.yml/badge.svg)](https://github.com/Coolhand-Labs/squishling/actions/workflows/ci.yml)
6
+ [![Gem Version](https://badge.fury.io/rb/squishling.svg)](https://badge.fury.io/rb/squishling)
4
7
 
5
- Elastic Ruby classes. A squished method either runs its Ruby implementation or sends its inputs through an LLM
6
- (via [RubyLLM](https://rubyllm.com)), and either way returns the same strict-schema-validated, typed result.
7
- Callers can't tell the difference.
8
+ **Don't just use AI to write code. Use AI to not need code.** Write code only where it's worth maintaining.
8
9
 
9
- Inspired by [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai):
10
- start flexible with AI, then harden high-volume paths into code as the economics justify it.
10
+ A squished method either runs its Ruby implementation, with the LLM as a failover, or routes some callers
11
+ through an LLM as a way of eliminating (or discovering) the cost of maintaining them as deterministic code.
12
+ Either way it returns the same strict-schema-validated, typed result, so callers can't tell the difference.
13
+
14
+ Think of code as a cache for AI judgment. Every line you write is something to maintain, so the LLM is the
15
+ default and code has to earn its place: you build it for the inputs that carry the volume, and the LLM keeps
16
+ handling the rest. New to the idea? Read
17
+ [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
11
18
 
12
19
  ## Use cases
13
20
 
14
- ### Instant integration
21
+ ### Zero-code integrations
15
22
 
16
23
  Accept a new data source today, before anyone writes a parser. Declare what you want back and leave the
17
24
  method unimplemented: every call goes to the LLM, and its output is validated against your schema.
18
25
 
19
26
  ```ruby
20
- class PaymentWebhook
27
+ class ShipmentUpdate
21
28
  include Squishling
22
29
 
23
- instructions "Normalize this payment provider's webhook into our payment event."
30
+ purpose "Normalize this carrier's tracking webhook into our shipment update."
24
31
  output_schema do
25
- string :event, enum: %w[succeeded failed refunded disputed]
26
- integer :amount_cents
27
- string :currency
28
- string :external_id
32
+ string :status, enum: %w[label_created in_transit out_for_delivery delivered exception]
33
+ string :tracking_number
34
+ string :location
29
35
  end
30
36
  # No `def call` yet, so every webhook goes to the LLM.
31
37
  end
32
38
 
33
- event = PaymentWebhook.call(provider: "adyen", payload: request.raw_post)
34
- event.event # => "refunded"
35
- event.amount_cents # => 4200
36
- event.squished? # => true
39
+ update = ShipmentUpdate.call(carrier: "acme-freight", payload: request.raw_post)
40
+ update.status # => "out_for_delivery"
41
+ update.squished? # => true
37
42
  ```
38
43
 
39
- When one provider carries the volume, write `def call` for it and add
40
- `squish_when { |provider:, **| provider != "stripe" }`. Stripe then runs in Ruby, everything else stays on the
44
+ When one carrier carries the volume, write `def call` for it and add
45
+ `squish_when { |carrier:, **| carrier != "ups" }`. UPS then runs in Ruby, everything else stays on the
41
46
  LLM, and callers don't change. See [Hardening a path](docs/routing.md#hardening-a-path).
42
47
 
43
- ### Error recovery
48
+ Validation guarantees the *shape* of the result, not that it's true. Where a wrong value is costly (money,
49
+ identity), cross-check it against another source or keep that path in Ruby.
50
+
51
+ ### Rescue errors your code can't handle yet
44
52
 
45
53
  Keep the Ruby you have for the inputs it understands, and hand the rest to the LLM instead of failing. When the
46
54
  parser raises, `squish!` sends this call to the LLM with the error and the parser's own source as context.
@@ -49,8 +57,8 @@ parser raises, `squish!` sends this call to the LLM with the error and the parse
49
57
  class InvoiceParser
50
58
  include Squishling
51
59
 
52
- instructions "Extract the invoice fields from the vendor's document."
53
- append_instructions "The Ruby parser that handles well-formed invoices:", self # this class's source
60
+ purpose "Extract the invoice fields from the vendor's document."
61
+ append_to_purpose "The Ruby parser that handles well-formed invoices:", self # this class's source
54
62
  output_schema do
55
63
  string :invoice_number
56
64
  number :total
@@ -60,7 +68,7 @@ class InvoiceParser
60
68
  invoice = VendorFormats.fetch(vendor).parse(document)
61
69
  result(invoice_number: invoice.number, total: invoice.total)
62
70
  rescue VendorFormats::ParseError => e
63
- squish!(append_instructions: "The parser failed on this document; the error is in the context.",
71
+ squish!(append_to_purpose: "The parser failed on this document; the error is in the context.",
64
72
  context: { parse_error: e })
65
73
  end
66
74
  end
@@ -71,7 +79,30 @@ InvoiceParser.call(vendor: "acme", document: scanned_text).squished? # => true
71
79
 
72
80
  Both calls return the same result class. If the LLM can't deliver either, `squish_fallback` decides what to
73
81
  return, or the error is raised with the original `ParseError` as its cause. See
74
- [Escalating from Ruby](docs/routing.md#escalating-from-ruby-with-squish).
82
+ [Handing off to the LLM](docs/routing.md#handing-off-to-the-llm-with-squish).
83
+
84
+ ### Measure your tech debt in tokens
85
+
86
+ Code you haven't written is debt you haven't taken on. A squished path has a running cost you can read off a
87
+ meter, so "should we write a parser for this?" becomes arithmetic: what the path costs in tokens each month,
88
+ against what it costs to write and maintain the code. Write it when the first number is bigger.
89
+
90
+ Add and initialize the [`coolhand`](https://github.com/Coolhand-Labs/coolhand-ruby) gem, and it picks up your
91
+ squishling calls automatically and measures their accuracy and cost:
92
+
93
+ ```ruby
94
+ # Gemfile
95
+ gem "coolhand"
96
+
97
+ # config/initializers/coolhand.rb
98
+ Coolhand.configure do |config|
99
+ config.api_key = ENV.fetch("COOLHAND_API_KEY")
100
+ end
101
+ ```
102
+
103
+ `squished?` tells you which path served a call. Coolhand records your LLM requests and responses; see
104
+ [what each tool sees](docs/measuring-tokens.md#privacy). Want to roll your own metrics? Check out
105
+ [our guide](docs/measuring-tokens.md) for doing it with other tools.
75
106
 
76
107
  ## Installation
77
108
 
@@ -79,18 +110,19 @@ return, or the error is raised with the original `ParseError` as its cause. See
79
110
  gem "squishling"
80
111
  ```
81
112
 
82
- Requires Ruby 3.3+ and RubyLLM 2.x. Configure your provider API keys in RubyLLM as usual, then optionally set a
83
- universal model:
113
+ Requires Ruby 3.3+ and RubyLLM 2.1+. Configure your provider API keys in RubyLLM as usual, then optionally set a
114
+ universal model, or an escalation of models to try in order:
84
115
 
85
116
  ```ruby
86
117
  Squishling.configure do |config|
87
118
  config.default_model = "claude-sonnet-5-5" # falls back to RubyLLM's default when nil
119
+ # or: config.default_escalation = [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5", "claude-opus-5-5"]
88
120
  end
89
121
  ```
90
122
 
91
123
  ## How it works
92
124
 
93
- A squishling class needs **instructions** (the system prompt) and an **output schema** (the shape of the result,
125
+ A squishling class needs a **purpose** (the system prompt) and an **output schema** (the shape of the result,
94
126
  validated on both paths). A squished call goes to the LLM when:
95
127
 
96
128
  - its `squish_when` predicate is truthy for these inputs,
@@ -103,23 +135,62 @@ Otherwise the Ruby runs, and whatever it returns is validated and typed like LLM
103
135
 
104
136
  - **One contract, two paths**: Ruby returns and LLM output are validated against the same strict schema and
105
137
  returned as the same typed `Data` objects. `squished?` tells you which path served a call.
106
- - **Your code as context**: `append_instructions` adds sections to the prompt, including a class's or method's
138
+ - **Your code as context**: `append_to_purpose` adds sections to the prompt, including a class's or method's
107
139
  own Ruby source.
108
140
  - **Any RubyLLM provider and model**: OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, OpenRouter, and more.
109
141
  Set a universal, per-class, per-method, or per-call model, plus layered generation params (temperature,
110
142
  reasoning effort, top_p, …).
111
- - **Defined failure behavior**: invalid output is re-asked with the validation errors, provider errors become
112
- `Squishling::LLMError`, and `squish_fallback` decides what to return when the LLM can't deliver.
143
+ - **Model escalation**: declare an `escalation:` instead of a `model:`, with per-step `attempts:`, `order:`, provider,
144
+ params, and `forward_rejected:`. Invalid output (with its errors fed back) or a provider failure moves on to the next attempt. Steps can
145
+ cross providers, e.g. a local Qwen model on Ollama, then Anthropic Claude Haiku on AWS Bedrock, then Claude Opus:
146
+
147
+ ```ruby
148
+ RubyLLM.configure do |config|
149
+ config.ollama_api_base = "http://localhost:11434/v1"
150
+ config.bedrock_api_key = ENV["AWS_ACCESS_KEY_ID"]
151
+ config.bedrock_secret_key = ENV["AWS_SECRET_ACCESS_KEY"]
152
+ config.bedrock_region = "us-east-1"
153
+ config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
154
+ end
155
+
156
+ class TicketTriager
157
+ include Squishling
158
+
159
+ squishling escalation: [
160
+ { model: "qwen3:8b", provider: :ollama, attempts: 2 }, # local and free, tried twice
161
+ { model: "claude-haiku-4-5", provider: :bedrock }, # small hosted model
162
+ { model: "claude-opus-5-5", provider: :anthropic } # last resort
163
+ ]
164
+ # purpose, output_schema, ...
165
+ end
166
+ ```
167
+ - **[Harnesses](docs/harnesses.md)**: choose how a call uses its escalation with `harness:`
168
+ - [`:escalation`](docs/configuration.md#models-and-escalation) (default): try each attempt until one passes
169
+ - [`:squishsum`](docs/harnesses.md#samples): two concurrent samples, accepted only if they agree
170
+ - [`:judged_squishsum`](docs/harnesses.md#the-judge): a chat or Jev judge picks between disagreeing samples
171
+ - [`:ensemble` and `:judged_ensemble`](docs/harnesses.md#ensembles): the same checks across two different models
172
+ - **Output contracts**: beyond the strict schema, conditional rules (`given`) and Ruby checks (`squish_validate`)
173
+ reject bad output and trigger the next attempt.
174
+ - **Defined failure behavior**: provider errors become `Squishling::LLMError`, and `squish_fallback` lets you decide
175
+ what to return when every attempt fails.
176
+ - **Observability**: the opt-in [`squawk`](docs/failures.md#observing-every-attempt-squawk) hook sends every attempt's
177
+ raw output to your error tracker. Model output stays out of error messages and logs otherwise.
113
178
  - **Opt-in context**: only method arguments, the instance state you name with `squish_context`, the `context:` you
114
- pass to `squish!`, and the source you choose to append are sent to the provider.
179
+ pass to `squish!`, and the source you choose to append are sent to the provider. The one addition: when
180
+ [escalation](docs/configuration.md#models-and-escalation) moves to the next step, that step also sees the previous
181
+ model's rejected output, unless the step sets `forward_rejected: false`.
115
182
 
116
183
  ## Documentation
117
184
 
118
- - [Configuration](docs/configuration.md): options, model and provider resolution, generation params, inheritance
119
- - [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `append_instructions`, hardening a path,
185
+ - [Configuration](docs/configuration.md): options, models and escalation, providers, generation params, inheritance
186
+ - [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `append_to_purpose`, hardening a path,
120
187
  what the LLM sees
121
- - [Output schemas](docs/schemas.md): schema forms, strict mode, typed results, optional vs. empty
122
- - [Failure handling](docs/failures.md): retries, error classes, fallbacks
188
+ - [Output schemas](docs/schemas.md): schema forms, strict mode, typed results, optional vs. empty, contracts
189
+ - [Failure handling](docs/failures.md): escalation, `squish_validate`, error classes, fallbacks
190
+ - [Harnesses](docs/harnesses.md): escalation, squishsum, ensemble, and judges (chat or Jev)
191
+ - [Naming and collisions](docs/naming.md): the methods `include Squishling` adds and what happens when a name is taken
192
+ - [Measuring token spend](docs/measuring-tokens.md): see what each squished path costs, with Coolhand Labs,
193
+ the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter
123
194
  - [Live examples](examples/README.md): end-to-end tests against Anthropic Claude Haiku and OpenAI GPT-6 Luna
124
195
 
125
196
  ## Development
@@ -131,6 +202,17 @@ bundle exec rake # RSpec (offline; RubyLLM is stubbed) + RuboCop
131
202
 
132
203
  See [AGENTS.md](AGENTS.md) for repo conventions, and [SECURITY.md](SECURITY.md) for reporting vulnerabilities.
133
204
 
205
+ ## Credits
206
+
207
+ Inspired by [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
208
+
134
209
  ## License
135
210
 
136
211
  Apache-2.0
212
+
213
+ ---
214
+
215
+ <p align="center">
216
+ This open source project is supported by <a href="https://coolhandlabs.com/">Coolhand Labs</a>.<br><br>
217
+ <a href="https://coolhandlabs.com/"><img src="assets/coolhand-labs.png" alt="Coolhand Labs" width="220"></a>
218
+ </p>
@@ -14,65 +14,135 @@ Then configure Squishling:
14
14
 
15
15
  ```ruby
16
16
  Squishling.configure do |config|
17
- config.default_model = "claude-sonnet-5-5"
17
+ config.default_model = "claude-sonnet-5-5" # or default_escalation, below
18
18
  config.default_provider = nil
19
19
  config.default_params = { temperature: 0 }
20
- config.max_retries = 1
20
+ config.default_harness = :escalation
21
21
  config.logger = Rails.logger
22
22
  end
23
23
  ```
24
24
 
25
25
  | Option | Default | Description |
26
26
  |---|---|---|
27
- | `default_model` | `nil` | The universal model for every squishling class that doesn't declare one. `nil` uses RubyLLM's `default_model`. |
28
- | `default_provider` | `nil` | The provider for `default_model`. Only needed for models missing from RubyLLM's registry (see below). |
29
- | `default_params` | `{}` | Generation params for every call (temperature, reasoning effort, top_p, …), overridable per class and per method. See [Generation params](#generation-params). |
30
- | `max_retries` | `1` | How many times to re-ask the LLM after schema-invalid output before raising `InvalidOutputError`. `0` means a single attempt. See [Failure handling](failures.md). |
31
- | `logger` | `nil` | Any `Logger`. Debug lines when a call routes to the LLM, warnings when a fallback is used. |
27
+ | `default_model` | `nil` | The universal model (a single attempt) for every squishling class that doesn't declare one. `nil` uses RubyLLM's `default_model`. |
28
+ | `default_escalation` | `nil` | The universal [escalation](#models-and-escalation): models tried in order. Assigning it clears `default_model`, and vice versa. |
29
+ | `default_provider` | `nil` | The provider for `default_model`/`default_escalation` steps that don't name one. Only needed for models missing from RubyLLM's registry (see below). |
30
+ | `default_params` | `{}` | Generation params for every call (temperature, reasoning effort, top_p, …), overridable per class, per method, and per escalation step. See [Generation params](#generation-params). |
31
+ | `squawk` | `nil` | A callable run after every LLM attempt with the raw output, for sending it to an error tracker or tracing tool. Also settable per class and per method. See [Observing every attempt](failures.md#observing-every-attempt-squawk). |
32
+ | `default_harness` | `nil` (`:escalation`) | How every squishling class that doesn't declare its own uses its escalation: `:escalation`, `:squishsum`, `:judged_squishsum`, `:ensemble`, `:judged_ensemble`, or a Hash with `type:` and options. See [Harnesses](harnesses.md). |
33
+ | `logger` | `nil` | Any `Logger`. Debug lines when a call routes to the LLM, warnings on each escalation and when a fallback is used. Never includes the raw model response; see [what ends up in errors and logs](failures.md#what-ends-up-in-errors-and-logs). |
32
34
 
33
35
  Transport-level retries (rate limits, 5xx, timeouts) are configured on RubyLLM itself
34
36
  (`RubyLLM.config.max_retries`, `request_timeout`).
35
37
 
36
- ## Model resolution
38
+ > **`max_retries` was removed.** The escalation now decides how many attempts run, and setting
39
+ > `config.max_retries` raises `ConfigurationError`. The old default (`max_retries = 1`, two attempts on one model) is
40
+ > `default_escalation = [{ model: "your-model", attempts: 2 }]`. A plain `model:` now makes one attempt.
37
41
 
38
- The model for a squished call comes from the first level that declares one:
42
+ ## Models and escalation
39
43
 
40
- 1. per call: `squish!(model: "...")` (see [Per-call overrides](#per-call-overrides))
41
- 2. per method: `squish :name, model: "..."`
42
- 3. per class: `squishling model: "..."` (inherited by subclasses)
43
- 4. universal: `Squishling.config.default_model`
44
- 5. RubyLLM's `default_model`
44
+ Declare **either** a `model:`, one model with a single attempt, **or** an `escalation:`, the models to try in
45
+ order. When an attempt fails (invalid output, a [`squish_validate`](failures.md#output-checks-squish_validate)
46
+ rejection, or a provider failure), the call moves on to the next attempt. It stops at the first valid output, or
47
+ raises once the last attempt fails.
48
+
49
+ ```ruby
50
+ Squishling.configure do |config|
51
+ config.default_escalation = [
52
+ { model: "claude-haiku-4-5", attempts: 2 }, # attempts 1-2: same conversation, told what was wrong
53
+ "claude-sonnet-5-5", # attempt 3: fresh chat, shown Haiku's last output and errors
54
+ "claude-opus-5-5" # attempt 4
55
+ ]
56
+ end
57
+ ```
58
+
59
+ Each step is a model name, or a Hash with:
60
+
61
+ | Key | Default | Description |
62
+ |---|---|---|
63
+ | `model:` | required | The model id. |
64
+ | `attempts:` | `1` | How many attempts this step gets before escalating to the next one. |
65
+ | `order:` | position | An Integer; steps run lowest first. Give every step an `order:` or none (duplicates are rejected). Gaps are fine. |
66
+ | `provider:` | the level's `provider:` | See [Providers](#providers-and-newly-released-models). |
67
+ | `forward_rejected:` | `true` | Set `false` to start this step from the original input alone, without the previous step's rejected output and its errors (see below). |
68
+ | `params:` | `{}` | [Generation params](#generation-params) for this step, merged over the class and method params key by key. |
69
+
70
+ Without `order:`, steps run in the order written. With it, the order is explicit and doesn't depend on position:
71
+
72
+ ```ruby
73
+ squishling escalation: [
74
+ { model: "claude-opus-5-5", order: 20, params: { thinking: { effort: :high } } },
75
+ { model: "claude-haiku-4-5", order: 0, attempts: 2 },
76
+ { model: "claude-sonnet-5-5", order: 10 }
77
+ ]
78
+ ```
79
+
80
+ `params:` on a step let one step switch to a reasoning model that rejects sampling params:
81
+
82
+ ```ruby
83
+ squishling params: { temperature: 0 },
84
+ escalation: ["gpt-5-mini", { model: "gpt-6-luna", provider: :openai,
85
+ params: { temperature: nil, thinking: { effort: :high } } }]
86
+ ```
87
+
88
+ ### Which model or escalation applies
89
+
90
+ The first level that declares a `model:` or an `escalation:` supplies the whole thing (levels aren't merged):
91
+
92
+ 1. per call: `squish!(model: ...)` or `squish!(escalation: ...)` (see [Per-call overrides](#per-call-overrides))
93
+ 2. per method: `squish :name, model: ...` or `escalation: ...`
94
+ 3. per class: `squishling model: ...` or `escalation: ...` (inherited by subclasses)
95
+ 4. universal: `Squishling.config.default_model` or `default_escalation`
96
+ 5. RubyLLM's `default_model` (a single attempt)
45
97
 
46
98
  ```ruby
47
99
  class InvoiceParser
48
100
  include Squishling
49
- squishling model: "claude-sonnet-5-5" # class default
101
+ squishling escalation: %w[claude-sonnet-5-5 claude-opus-5-5] # class escalation
50
102
 
51
- squish :classify, model: "claude-haiku-4-5" do # cheaper model for one method
103
+ squish :classify, model: "claude-haiku-4-5" do # one cheap attempt for this method
52
104
  string :category
53
105
  end
54
106
  end
55
107
  ```
56
108
 
109
+ Passing both `model:` and `escalation:` in one declaration raises `ConfigurationError`, and so does a list passed as
110
+ `model:`. Across separate declarations (a reopened class, or two config assignments), the latest one wins.
111
+
112
+ Consecutive attempts of the same step (same model, provider, params, and `forward_rejected:`, e.g. `attempts: 2`)
113
+ continue one conversation. Moving to a different step starts a fresh chat with the original input plus the previous
114
+ output and why it was rejected, so no provider-specific history crosses providers. That rejected output is
115
+ model-generated and is sent to the next step's provider, even when it is a different one, truncated to 4,000
116
+ characters. If you don't want a step to see it (say, because the next step is a different provider), set
117
+ `forward_rejected: false` on that step and it starts from the original input only. Configuration errors (bad
118
+ credentials, an unknown model, a request the provider rejects) never move to the next attempt. See [Failure handling](failures.md).
119
+
57
120
  ## Providers and newly released models
58
121
 
59
122
  RubyLLM looks models up in its bundled registry. To use a model that isn't there yet, such as a newly released
60
123
  OpenAI or Anthropic model, name its provider next to it. Squishling then tells RubyLLM to assume the model exists:
61
124
 
62
125
  ```ruby
63
- squishling model: "gpt-6-luna", provider: :openai
126
+ squishling model: "gpt-7-preview", provider: :openai
64
127
  squish :triage, model: "claude-haiku-4-5", provider: :anthropic
65
- Squishling.configure { |c| c.default_model = "gpt-6-luna"; c.default_provider = :openai }
128
+ Squishling.configure { |c| c.default_model = "gpt-7-preview"; c.default_provider = :openai }
66
129
  ```
67
130
 
68
- A provider is paired with the model declared at the same level, so a per-method model never inherits a
69
- class-level provider meant for a different model. For models that are in the registry, a provider is optional
70
- and the normal registry lookup is kept.
131
+ A provider is paired with the model or escalation declared at the same level, and applies to that escalation's steps
132
+ that don't name their own. A model never inherits a provider meant for a different model: a per-method model doesn't
133
+ take the class's provider, and a subclass that declares its own `model:` or `escalation:` doesn't take its parent's
134
+ (a subclass that only changes params or the purpose keeps the parent's model and provider together). For models that
135
+ are in the registry, a provider is optional and the normal registry lookup is kept.
136
+
137
+ A `provider:` with no model beside it would apply to nothing, so it raises `ConfigurationError`:
138
+ `squish ..., provider: ...` needs a `model:` or `escalation:` in the same call (as `squish!` always has), and
139
+ `squishling provider: ...` needs one in the call or already declared on that same class (not inherited). Set the
140
+ provider next to the model, or, for the configured default model, use `config.default_provider`.
71
141
 
72
142
  ## Generation params
73
143
 
74
- `params` holds generation settings. Set them at any of three levels; each level overrides the one above it
75
- **key by key**:
144
+ `params` holds generation settings. Set them at any of three levels, plus per step in an
145
+ [escalation](#models-and-escalation); each level overrides the one above it **key by key**:
76
146
 
77
147
  ```ruby
78
148
  Squishling.configure { |c| c.default_params = { temperature: 0 } } # every call
@@ -107,22 +177,33 @@ end
107
177
  ```
108
178
 
109
179
  - Keys that Squishling or RubyLLM control (`model`, `messages`, `input`, `instructions`, `system`, `stream`,
110
- `response_format`, `text`, `output_config`, `tools`, `tool_choice`, `schema`, …) raise `ConfigurationError`,
111
- because they would override the model, the conversation, or the strict output format.
180
+ `response_format`, `text`, `output_config`, `tools`, `tool_choice`, `schema`, and the camelCase and plural
181
+ spellings other providers use, such as `systemInstruction`, `toolConfig`/`tool_config`, `outputConfig`, `inputs`, …)
182
+ raise `ConfigurationError`, because they would override the model, the conversation, or the strict output format.
183
+ - **Nested containers** such as Gemini's `generationConfig` (and `generation_config` for Gemini Interactions) and
184
+ Mistral Conversations' `completion_args` accept ordinary settings, but reject the keys inside them that carry the
185
+ output format or tools (`responseMimeType`, `responseSchema`, `responseJsonSchema`, `response_format`, `tools`,
186
+ `tool_choice`, `toolConfig`, and their snake_case spellings). Use `output_schema` instead:
187
+
188
+ ```ruby
189
+ squishling params: { generationConfig: { topK: 5 } } # ok
190
+ squishling params: { generationConfig: { responseMimeType: "text/plain" } } # ConfigurationError
191
+ ```
192
+
112
193
  - If a provider rejects a param, the call raises `Squishling::ConfigurationError` naming the params. It isn't
113
194
  retried or sent to `squish_fallback`. See [Failure handling](failures.md).
114
195
 
115
196
  ## Inheritance
116
197
 
117
- Subclasses inherit the model, provider, generation params (merged key by key), instructions, output schema,
118
- `squish_when` predicate, `squish_context` names, `squish_fallback`, and every `squish` declaration. Overrides in a subclass, including
198
+ Subclasses inherit the model or escalation and its provider (as a pair), harness, generation params (merged key by key), purpose, output schema,
199
+ `squish_when` predicate, `squish_context` names, `squish_validate`, `squish_fallback`, and every `squish` declaration. Overrides in a subclass, including
119
200
  overridden methods, are routed the same way.
120
201
 
121
- `append_instructions` sections are added to, not replaced: a subclass's sections follow its parent's, and
122
- `append_instructions false` drops the inherited ones. See [Appending to the instructions](routing.md#appending-to-the-instructions).
202
+ `append_to_purpose` sections are added to, not replaced: a subclass's sections follow its parent's, and
203
+ `append_to_purpose false` drops the inherited ones. See [Appending to the purpose](routing.md#appending-to-the-purpose).
123
204
 
124
205
  ## Per-call overrides
125
206
 
126
- Inside a squished method, `squish!` sends the call to the LLM with its own `instructions:`,
127
- `append_instructions:`, `context:`, `model:`/`provider:`, and `params:`. Each layers over the method and class
128
- settings the same way they layer over each other. See [Escalating from Ruby](routing.md#escalating-from-ruby-with-squish).
207
+ Inside a squished method, `squish!` sends the call to the LLM with its own `purpose:`,
208
+ `append_to_purpose:`, `context:`, `model:` or `escalation:` (with `provider:`), `params:`, and `harness:`. Each layers over the method and class
209
+ settings the same way they layer over each other. See [Handing off to the LLM](routing.md#handing-off-to-the-llm-with-squish).