squishling 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: ad18b43555b041d5c0ef6eabda578fc6a2111ecfbf962b2003a7b502584f0c91
4
- data.tar.gz: f4ce9bb54bd2c57abd73c5a2f0e50472c338e299f25fe7ed0dd2edb2f4854f06
3
+ metadata.gz: '0748f552c985521890d365ebc0818ec59c4336f93fb183d1769f49ed8cd32349'
4
+ data.tar.gz: '05128ecf8d7156e0c8ca084bf8a1860b754596017df2cdfdcc35e2242138fa7b'
5
5
  SHA512:
6
- metadata.gz: 726d27bcc6067039d3bba294a923c302f49fc756ce0882541461525d8421be48d5cbb9c02e1d201eb1c062c9519353f62b66f9cf43795fc46a15c9834d7e24fe
7
- data.tar.gz: 9e4093ad080e9742b38e6bfe3d59f2b5be21f0bf3ddfd5a8288b552bbe9d0abdffe3d47f6fea8d5d0c54afdb0172c69769451249e65f69cfc2665b96682b1b52
6
+ metadata.gz: 2be443b51d2dc1460a410fbf0850bed09e33fff5f7ad894d4bd92e3f47d428d22df7bfc0371036c090a6b806f466b1d930de3a9caf8d1b044185fd4000b765bf
7
+ data.tar.gz: 3a18d6e041f5807c966677fea5499de8af5d0f28146cb98bbe205141598ee6a01a2d598b1326df3b78227455dcc124c69cf95afacabf9a1689a1cc382a652414
data/CHANGELOG.md CHANGED
@@ -7,6 +7,89 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.3.0] - 2026-10-09
11
+
12
+ ### Added
13
+
14
+ - Harnesses: a `harness:` option (on `squishling`, `squish`, `squish!`, and `default_harness` in
15
+ `Squishling.configure`) chooses how a call uses its escalation. `:escalation` is the existing behavior and stays
16
+ the default. `:squishsum` asks the first escalation step twice, concurrently, and accepts the result only when the
17
+ two samples agree (exact match, or a custom `compare:`). `:judged_squishsum` sends a disagreement to a judge: the
18
+ next escalation step or a dedicated `judge:` step, either a chat model or a Jev-style decision model through
19
+ `RubyLLM.judge`, with a default prompt you can override through `judge_instructions:`. Samples that still disagree
20
+ raise the new `Squishling::DisagreementError < InvalidOutputError`, which carries both typed results, and
21
+ `squish_fallback` remains the only fallback. See [Harnesses](docs/harnesses.md). (#16)
22
+ - `:ensemble` and `:judged_ensemble` harnesses: like the squishsum ones, but the two concurrent samples come from the
23
+ escalation's first and second steps (a cross-model check) instead of the first step twice, each retrying within its
24
+ own step's `attempts:`. A judged ensemble's judge defaults to the third step, or `judge:`, and
25
+ `DisagreementError#models` names the model behind each sample. An ensemble needs at least two distinct steps (adjacent
26
+ identical steps count as one step's attempts) or it raises `ConfigurationError` before any request. (#20)
27
+ - `squawk`, an opt-in hook for sending what the model returned to an error tracker or tracing tool. It is called after
28
+ every LLM attempt (every sample and judge attempt under a harness) with `output:`, `metadata:`, and `error:`, and
29
+ can be set on `Squishling.configure`, `squishling`, and `squish`. It does nothing by default, and exceptions it
30
+ raises propagate unchanged. (#15)
31
+ - `forward_rejected:` on escalation steps. When an attempt is rejected and the next step starts, the previous
32
+ output is forwarded to that step's provider by default (unchanged behavior); set `forward_rejected: false` to start
33
+ that step from the original input alone. (#14)
34
+ - `include Squishling` raises `ConfigurationError` when the class inherits a method Squishling would override (for
35
+ example `Sinatra::Base.call`) instead of silently shadowing it. See [Naming and collisions](docs/naming.md). (#17)
36
+
37
+ ### Documentation
38
+
39
+ - A new guide, [Measuring token spend](docs/measuring-tokens.md), covers working out what each squished path costs
40
+ (with Coolhand Labs, the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter), and the README is
41
+ reworked around using Squishling to harden code paths as the economics justify it. (#21)
42
+
43
+ ### Changed
44
+
45
+ - **Breaking:** the DSL method `instructions` is now `purpose`, and `append_instructions` is now `append_to_purpose`,
46
+ including the `instructions:`/`append_instructions:` keywords on `squishling`, `squish`, and `squish!`. There are no
47
+ aliases. Rename them in your classes: `instructions "..."` becomes `purpose "..."`, and
48
+ `append_instructions: [...]` becomes `append_to_purpose: [...]`. The error for a missing prompt now says
49
+ "has no purpose". (#17)
50
+ - **Breaking:** `ruby_llm` must now be `~> 2.1` (was `~> 2.0`); run `bundle update ruby_llm`. (#16)
51
+ - **Breaking:** every value a method with an `output_schema` returns is now validated and typed, not only Hashes. On
52
+ the Ruby path and in `squish_fallback`, a `nil`, String, Array, or other non-Hash return that used to pass through
53
+ now raises `InvalidOutputError`; return a Hash (or `result(...)`) that matches the schema. A result of the schema's
54
+ own class passes through, another schema's result is re-validated, and values JSON can't represent (`NaN`,
55
+ `Infinity`, cycles) raise `InvalidOutputError` instead of a raw JSON error. Inner calls (`super`, or the method
56
+ called from `squish_when` or a fallback) still only type Hashes, so overrides can reshape values. (#13)
57
+ - **Behavior change:** a provider now travels with the model declared at the same level on a class too. A subclass
58
+ that declares its own `model:` or `escalation:` no longer inherits its parent's `provider:`, which could send the
59
+ whole payload to the parent's provider under another provider's model name. If you relied on that, set `provider:`
60
+ next to the subclass's model.
61
+ - **Behavior change:** a `provider:` with no model beside it now raises `ConfigurationError` instead of being
62
+ silently ignored (the default provider was used). `squish ..., provider: ...` needs a `model:` or `escalation:` in
63
+ the same call, as `squish!` already did; `squishling provider: ...` needs one in the call or already declared on that
64
+ class. Move the `provider:` next to a model, or use `config.default_provider` with the default model.
65
+ - The `result` alias for `squishling_result` is no longer added when the class already has a `result` (its own or
66
+ inherited); use `squishling_result` there. (#17)
67
+ - `InvalidOutputError#message` and `config.logger` warnings no longer quote the model's response. For unparseable JSON
68
+ they report only the line and column; the raw output stays in `InvalidOutputError#raw`. (#15)
69
+
70
+ ### Fixed
71
+
72
+ - A number a model writes outside the range JSON can represent (such as `1e400`, which parses to `Infinity`), or a
73
+ string with invalid UTF-8, was accepted as valid, while the same value from a Ruby method was rejected. It is now
74
+ invalid output on the LLM path too, and no longer raises a bare `JSON::GeneratorError` when `:squishsum` compares samples.
75
+
76
+ ### Security
77
+
78
+ - Params can no longer replace the strict output format or tools through the containers that also hold ordinary
79
+ settings: `generationConfig`/`generation_config` (Google Gemini) and `completion_args` (Mistral Conversations) now
80
+ reject their format and tool keys, with symbol or string keys, and a non-Hash value for them is rejected.
81
+ `tool_config` is reserved at the top level. Settings such as `topK` still work. (#12)
82
+ - The previous model's rejected output forwarded to the next escalation step is capped at 4,000 characters
83
+ (`Invoker::MAX_FORWARDED_CHARS`), with a truncation marker. (#14)
84
+ - Model output is kept out of error messages and logs (see the `InvalidOutputError#message` change above); `squawk`
85
+ is the one sanctioned way for it to leave the process. (#15) A chat judge's `reason` is model-written, so
86
+ `DisagreementError#message` and the logs no longer include it; read it from `DisagreementError#reason`. A
87
+ `:judgment` judge's message still carries the choice and probability, and a choice other than `a`, `b`, or `neither`
88
+ is no longer repeated.
89
+ - Only the first 20 validation errors (plus a count of the rest) are fed back to the model, logged, and put in
90
+ `InvalidOutputError#message` for one invalid output, so a very large malformed response can no longer produce a
91
+ megabyte-sized retry message or log line.
92
+
10
93
  ## [0.2.0] - 2026-10-09
11
94
 
12
95
  ### Added
data/README.md CHANGED
@@ -1,46 +1,54 @@
1
- # Squishling
1
+ <h1 align="center">
2
+ <img src="assets/squishling-logo.png" alt="Squishling" width="420">
3
+ </h1>
2
4
 
3
5
  [![CI](https://github.com/Coolhand-Labs/squishling/actions/workflows/ci.yml/badge.svg)](https://github.com/Coolhand-Labs/squishling/actions/workflows/ci.yml)
6
+ [![Gem Version](https://badge.fury.io/rb/squishling.svg)](https://badge.fury.io/rb/squishling)
4
7
 
5
- Elastic Ruby classes. A squished method either runs its Ruby implementation or sends its inputs through an LLM
6
- (via [RubyLLM](https://rubyllm.com)), and either way returns the same strict-schema-validated, typed result.
7
- Callers can't tell the difference.
8
+ **Don't just use AI to write code. Use AI to not need code.** Write code only where it's worth maintaining.
8
9
 
9
- Inspired by [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai):
10
- start flexible with AI, then harden high-volume paths into code as the economics justify it.
10
+ A squished method either runs its Ruby implementation, with the LLM as a failover, or routes some callers
11
+ through an LLM as a way of eliminating (or discovering) the cost of maintaining them as deterministic code.
12
+ Either way it returns the same strict-schema-validated, typed result, so callers can't tell the difference.
13
+
14
+ Think of code as a cache for AI judgment. Every line you write is something to maintain, so the LLM is the
15
+ default and code has to earn its place: you build it for the inputs that carry the volume, and the LLM keeps
16
+ handling the rest. New to the idea? Read
17
+ [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
11
18
 
12
19
  ## Use cases
13
20
 
14
- ### Instant integration
21
+ ### Zero-code integrations
15
22
 
16
23
  Accept a new data source today, before anyone writes a parser. Declare what you want back and leave the
17
24
  method unimplemented: every call goes to the LLM, and its output is validated against your schema.
18
25
 
19
26
  ```ruby
20
- class PaymentWebhook
27
+ class ShipmentUpdate
21
28
  include Squishling
22
29
 
23
- instructions "Normalize this payment provider's webhook into our payment event."
30
+ purpose "Normalize this carrier's tracking webhook into our shipment update."
24
31
  output_schema do
25
- string :event, enum: %w[succeeded failed refunded disputed]
26
- integer :amount_cents
27
- string :currency
28
- string :external_id
32
+ string :status, enum: %w[label_created in_transit out_for_delivery delivered exception]
33
+ string :tracking_number
34
+ string :location
29
35
  end
30
36
  # No `def call` yet, so every webhook goes to the LLM.
31
37
  end
32
38
 
33
- event = PaymentWebhook.call(provider: "adyen", payload: request.raw_post)
34
- event.event # => "refunded"
35
- event.amount_cents # => 4200
36
- event.squished? # => true
39
+ update = ShipmentUpdate.call(carrier: "acme-freight", payload: request.raw_post)
40
+ update.status # => "out_for_delivery"
41
+ update.squished? # => true
37
42
  ```
38
43
 
39
- When one provider carries the volume, write `def call` for it and add
40
- `squish_when { |provider:, **| provider != "stripe" }`. Stripe then runs in Ruby, everything else stays on the
44
+ When one carrier carries the volume, write `def call` for it and add
45
+ `squish_when { |carrier:, **| carrier != "ups" }`. UPS then runs in Ruby, everything else stays on the
41
46
  LLM, and callers don't change. See [Hardening a path](docs/routing.md#hardening-a-path).
42
47
 
43
- ### Error recovery
48
+ Validation guarantees the *shape* of the result, not that it's true. Where a wrong value is costly (money,
49
+ identity), cross-check it against another source or keep that path in Ruby.
50
+
51
+ ### Rescue errors your code can't handle yet
44
52
 
45
53
  Keep the Ruby you have for the inputs it understands, and hand the rest to the LLM instead of failing. When the
46
54
  parser raises, `squish!` sends this call to the LLM with the error and the parser's own source as context.
@@ -49,8 +57,8 @@ parser raises, `squish!` sends this call to the LLM with the error and the parse
49
57
  class InvoiceParser
50
58
  include Squishling
51
59
 
52
- instructions "Extract the invoice fields from the vendor's document."
53
- append_instructions "The Ruby parser that handles well-formed invoices:", self # this class's source
60
+ purpose "Extract the invoice fields from the vendor's document."
61
+ append_to_purpose "The Ruby parser that handles well-formed invoices:", self # this class's source
54
62
  output_schema do
55
63
  string :invoice_number
56
64
  number :total
@@ -60,7 +68,7 @@ class InvoiceParser
60
68
  invoice = VendorFormats.fetch(vendor).parse(document)
61
69
  result(invoice_number: invoice.number, total: invoice.total)
62
70
  rescue VendorFormats::ParseError => e
63
- squish!(append_instructions: "The parser failed on this document; the error is in the context.",
71
+ squish!(append_to_purpose: "The parser failed on this document; the error is in the context.",
64
72
  context: { parse_error: e })
65
73
  end
66
74
  end
@@ -73,13 +81,36 @@ Both calls return the same result class. If the LLM can't deliver either, `squis
73
81
  return, or the error is raised with the original `ParseError` as its cause. See
74
82
  [Handing off to the LLM](docs/routing.md#handing-off-to-the-llm-with-squish).
75
83
 
84
+ ### Measure your tech debt in tokens
85
+
86
+ Code you haven't written is debt you haven't taken on. A squished path has a running cost you can read off a
87
+ meter, so "should we write a parser for this?" becomes arithmetic: what the path costs in tokens each month,
88
+ against what it costs to write and maintain the code. Write it when the first number is bigger.
89
+
90
+ Add and initialize the [`coolhand`](https://github.com/Coolhand-Labs/coolhand-ruby) gem, and it picks up your
91
+ squishling calls automatically and measures their accuracy and cost:
92
+
93
+ ```ruby
94
+ # Gemfile
95
+ gem "coolhand"
96
+
97
+ # config/initializers/coolhand.rb
98
+ Coolhand.configure do |config|
99
+ config.api_key = ENV.fetch("COOLHAND_API_KEY")
100
+ end
101
+ ```
102
+
103
+ `squished?` tells you which path served a call. Coolhand records your LLM requests and responses; see
104
+ [what each tool sees](docs/measuring-tokens.md#privacy). Want to roll your own metrics? Check out
105
+ [our guide](docs/measuring-tokens.md) for doing it with other tools.
106
+
76
107
  ## Installation
77
108
 
78
109
  ```ruby
79
110
  gem "squishling"
80
111
  ```
81
112
 
82
- Requires Ruby 3.3+ and RubyLLM 2.x. Configure your provider API keys in RubyLLM as usual, then optionally set a
113
+ Requires Ruby 3.3+ and RubyLLM 2.1+. Configure your provider API keys in RubyLLM as usual, then optionally set a
83
114
  universal model, or an escalation of models to try in order:
84
115
 
85
116
  ```ruby
@@ -91,7 +122,7 @@ end
91
122
 
92
123
  ## How it works
93
124
 
94
- A squishling class needs **instructions** (the system prompt) and an **output schema** (the shape of the result,
125
+ A squishling class needs a **purpose** (the system prompt) and an **output schema** (the shape of the result,
95
126
  validated on both paths). A squished call goes to the LLM when:
96
127
 
97
128
  - its `squish_when` predicate is truthy for these inputs,
@@ -104,13 +135,13 @@ Otherwise the Ruby runs, and whatever it returns is validated and typed like LLM
104
135
 
105
136
  - **One contract, two paths**: Ruby returns and LLM output are validated against the same strict schema and
106
137
  returned as the same typed `Data` objects. `squished?` tells you which path served a call.
107
- - **Your code as context**: `append_instructions` adds sections to the prompt, including a class's or method's
138
+ - **Your code as context**: `append_to_purpose` adds sections to the prompt, including a class's or method's
108
139
  own Ruby source.
109
140
  - **Any RubyLLM provider and model**: OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, OpenRouter, and more.
110
141
  Set a universal, per-class, per-method, or per-call model, plus layered generation params (temperature,
111
142
  reasoning effort, top_p, …).
112
143
  - **Model escalation**: declare an `escalation:` instead of a `model:`, with per-step `attempts:`, `order:`, provider,
113
- and params. Invalid output (with its errors fed back) or a provider failure moves on to the next attempt. Steps can
144
+ params, and `forward_rejected:`. Invalid output (with its errors fed back) or a provider failure moves on to the next attempt. Steps can
114
145
  cross providers, e.g. a local Qwen model on Ollama, then Anthropic Claude Haiku on AWS Bedrock, then Claude Opus:
115
146
 
116
147
  ```ruby
@@ -130,23 +161,36 @@ Otherwise the Ruby runs, and whatever it returns is validated and typed like LLM
130
161
  { model: "claude-haiku-4-5", provider: :bedrock }, # small hosted model
131
162
  { model: "claude-opus-5-5", provider: :anthropic } # last resort
132
163
  ]
133
- # instructions, output_schema, ...
164
+ # purpose, output_schema, ...
134
165
  end
135
166
  ```
167
+ - **[Harnesses](docs/harnesses.md)**: choose how a call uses its escalation with `harness:`
168
+ - [`:escalation`](docs/configuration.md#models-and-escalation) (default): try each attempt until one passes
169
+ - [`:squishsum`](docs/harnesses.md#samples): two concurrent samples, accepted only if they agree
170
+ - [`:judged_squishsum`](docs/harnesses.md#the-judge): a chat or Jev judge picks between disagreeing samples
171
+ - [`:ensemble` and `:judged_ensemble`](docs/harnesses.md#ensembles): the same checks across two different models
136
172
  - **Output contracts**: beyond the strict schema, conditional rules (`given`) and Ruby checks (`squish_validate`)
137
173
  reject bad output and trigger the next attempt.
138
174
  - **Defined failure behavior**: provider errors become `Squishling::LLMError`, and `squish_fallback` lets you decide
139
175
  what to return when every attempt fails.
176
+ - **Observability**: the opt-in [`squawk`](docs/failures.md#observing-every-attempt-squawk) hook sends every attempt's
177
+ raw output to your error tracker. Model output stays out of error messages and logs otherwise.
140
178
  - **Opt-in context**: only method arguments, the instance state you name with `squish_context`, the `context:` you
141
- pass to `squish!`, and the source you choose to append are sent to the provider.
179
+ pass to `squish!`, and the source you choose to append are sent to the provider. The one addition: when
180
+ [escalation](docs/configuration.md#models-and-escalation) moves to the next step, that step also sees the previous
181
+ model's rejected output, unless the step sets `forward_rejected: false`.
142
182
 
143
183
  ## Documentation
144
184
 
145
185
  - [Configuration](docs/configuration.md): options, models and escalation, providers, generation params, inheritance
146
- - [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `append_instructions`, hardening a path,
186
+ - [Routing](docs/routing.md): when a call goes to the LLM, `squish!`, `append_to_purpose`, hardening a path,
147
187
  what the LLM sees
148
188
  - [Output schemas](docs/schemas.md): schema forms, strict mode, typed results, optional vs. empty, contracts
149
189
  - [Failure handling](docs/failures.md): escalation, `squish_validate`, error classes, fallbacks
190
+ - [Harnesses](docs/harnesses.md): escalation, squishsum, ensemble, and judges (chat or Jev)
191
+ - [Naming and collisions](docs/naming.md): the methods `include Squishling` adds and what happens when a name is taken
192
+ - [Measuring token spend](docs/measuring-tokens.md): see what each squished path costs, with Coolhand Labs,
193
+ the `squawk` hook, OpenTelemetry, LangSmith, or RubyLLM's instrumenter
150
194
  - [Live examples](examples/README.md): end-to-end tests against Anthropic Claude Haiku and OpenAI GPT-6 Luna
151
195
 
152
196
  ## Development
@@ -158,6 +202,17 @@ bundle exec rake # RSpec (offline; RubyLLM is stubbed) + RuboCop
158
202
 
159
203
  See [AGENTS.md](AGENTS.md) for repo conventions, and [SECURITY.md](SECURITY.md) for reporting vulnerabilities.
160
204
 
205
+ ## Credits
206
+
207
+ Inspired by [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai).
208
+
161
209
  ## License
162
210
 
163
211
  Apache-2.0
212
+
213
+ ---
214
+
215
+ <p align="center">
216
+ This open source project is supported by <a href="https://coolhandlabs.com/">Coolhand Labs</a>.<br><br>
217
+ <a href="https://coolhandlabs.com/"><img src="assets/coolhand-labs.png" alt="Coolhand Labs" width="220"></a>
218
+ </p>
@@ -17,6 +17,7 @@ Squishling.configure do |config|
17
17
  config.default_model = "claude-sonnet-5-5" # or default_escalation, below
18
18
  config.default_provider = nil
19
19
  config.default_params = { temperature: 0 }
20
+ config.default_harness = :escalation
20
21
  config.logger = Rails.logger
21
22
  end
22
23
  ```
@@ -27,7 +28,9 @@ end
27
28
  | `default_escalation` | `nil` | The universal [escalation](#models-and-escalation): models tried in order. Assigning it clears `default_model`, and vice versa. |
28
29
  | `default_provider` | `nil` | The provider for `default_model`/`default_escalation` steps that don't name one. Only needed for models missing from RubyLLM's registry (see below). |
29
30
  | `default_params` | `{}` | Generation params for every call (temperature, reasoning effort, top_p, …), overridable per class, per method, and per escalation step. See [Generation params](#generation-params). |
30
- | `logger` | `nil` | Any `Logger`. Debug lines when a call routes to the LLM, warnings on each escalation and when a fallback is used. |
31
+ | `squawk` | `nil` | A callable run after every LLM attempt with the raw output, for sending it to an error tracker or tracing tool. Also settable per class and per method. See [Observing every attempt](failures.md#observing-every-attempt-squawk). |
32
+ | `default_harness` | `nil` (`:escalation`) | How every squishling class that doesn't declare its own uses its escalation: `:escalation`, `:squishsum`, `:judged_squishsum`, `:ensemble`, `:judged_ensemble`, or a Hash with `type:` and options. See [Harnesses](harnesses.md). |
33
+ | `logger` | `nil` | Any `Logger`. Debug lines when a call routes to the LLM, warnings on each escalation and when a fallback is used. Never includes the raw model response; see [what ends up in errors and logs](failures.md#what-ends-up-in-errors-and-logs). |
31
34
 
32
35
  Transport-level retries (rate limits, 5xx, timeouts) are configured on RubyLLM itself
33
36
  (`RubyLLM.config.max_retries`, `request_timeout`).
@@ -61,6 +64,7 @@ Each step is a model name, or a Hash with:
61
64
  | `attempts:` | `1` | How many attempts this step gets before escalating to the next one. |
62
65
  | `order:` | position | An Integer; steps run lowest first. Give every step an `order:` or none (duplicates are rejected). Gaps are fine. |
63
66
  | `provider:` | the level's `provider:` | See [Providers](#providers-and-newly-released-models). |
67
+ | `forward_rejected:` | `true` | Set `false` to start this step from the original input alone, without the previous step's rejected output and its errors (see below). |
64
68
  | `params:` | `{}` | [Generation params](#generation-params) for this step, merged over the class and method params key by key. |
65
69
 
66
70
  Without `order:`, steps run in the order written. With it, the order is explicit and doesn't depend on position:
@@ -105,10 +109,13 @@ end
105
109
  Passing both `model:` and `escalation:` in one declaration raises `ConfigurationError`, and so does a list passed as
106
110
  `model:`. Across separate declarations (a reopened class, or two config assignments), the latest one wins.
107
111
 
108
- Consecutive attempts of the same step (same model, provider, and params, e.g. `attempts: 2`) continue one
109
- conversation. Moving to a different step starts a fresh chat with the original input plus the previous output and why
110
- it was rejected, so no provider-specific history crosses providers. Configuration errors (bad credentials, an unknown model, a request the provider rejects) never move to the
111
- next attempt. See [Failure handling](failures.md).
112
+ Consecutive attempts of the same step (same model, provider, params, and `forward_rejected:`, e.g. `attempts: 2`)
113
+ continue one conversation. Moving to a different step starts a fresh chat with the original input plus the previous
114
+ output and why it was rejected, so no provider-specific history crosses providers. That rejected output is
115
+ model-generated and is sent to the next step's provider, even when it is a different one, truncated to 4,000
116
+ characters. If you don't want a step to see it (say, because the next step is a different provider), set
117
+ `forward_rejected: false` on that step and it starts from the original input only. Configuration errors (bad
118
+ credentials, an unknown model, a request the provider rejects) never move to the next attempt. See [Failure handling](failures.md).
112
119
 
113
120
  ## Providers and newly released models
114
121
 
@@ -116,14 +123,21 @@ RubyLLM looks models up in its bundled registry. To use a model that isn't there
116
123
  OpenAI or Anthropic model, name its provider next to it. Squishling then tells RubyLLM to assume the model exists:
117
124
 
118
125
  ```ruby
119
- squishling model: "gpt-6-luna", provider: :openai
126
+ squishling model: "gpt-7-preview", provider: :openai
120
127
  squish :triage, model: "claude-haiku-4-5", provider: :anthropic
121
- Squishling.configure { |c| c.default_model = "gpt-6-luna"; c.default_provider = :openai }
128
+ Squishling.configure { |c| c.default_model = "gpt-7-preview"; c.default_provider = :openai }
122
129
  ```
123
130
 
124
131
  A provider is paired with the model or escalation declared at the same level, and applies to that escalation's steps
125
- that don't name their own. A per-method model never inherits a class-level provider meant for a different model. For models that are in the registry, a provider is optional
126
- and the normal registry lookup is kept.
132
+ that don't name their own. A model never inherits a provider meant for a different model: a per-method model doesn't
133
+ take the class's provider, and a subclass that declares its own `model:` or `escalation:` doesn't take its parent's
134
+ (a subclass that only changes params or the purpose keeps the parent's model and provider together). For models that
135
+ are in the registry, a provider is optional and the normal registry lookup is kept.
136
+
137
+ A `provider:` with no model beside it would apply to nothing, so it raises `ConfigurationError`:
138
+ `squish ..., provider: ...` needs a `model:` or `escalation:` in the same call (as `squish!` always has), and
139
+ `squishling provider: ...` needs one in the call or already declared on that same class (not inherited). Set the
140
+ provider next to the model, or, for the configured default model, use `config.default_provider`.
127
141
 
128
142
  ## Generation params
129
143
 
@@ -164,22 +178,32 @@ end
164
178
 
165
179
  - Keys that Squishling or RubyLLM control (`model`, `messages`, `input`, `instructions`, `system`, `stream`,
166
180
  `response_format`, `text`, `output_config`, `tools`, `tool_choice`, `schema`, and the camelCase and plural
167
- spellings other providers use, such as `systemInstruction`, `toolConfig`, `outputConfig`, `inputs`, …) raise
168
- `ConfigurationError`, because they would override the model, the conversation, or the strict output format.
181
+ spellings other providers use, such as `systemInstruction`, `toolConfig`/`tool_config`, `outputConfig`, `inputs`, …)
182
+ raise `ConfigurationError`, because they would override the model, the conversation, or the strict output format.
183
+ - **Nested containers** such as Gemini's `generationConfig` (and `generation_config` for Gemini Interactions) and
184
+ Mistral Conversations' `completion_args` accept ordinary settings, but reject the keys inside them that carry the
185
+ output format or tools (`responseMimeType`, `responseSchema`, `responseJsonSchema`, `response_format`, `tools`,
186
+ `tool_choice`, `toolConfig`, and their snake_case spellings). Use `output_schema` instead:
187
+
188
+ ```ruby
189
+ squishling params: { generationConfig: { topK: 5 } } # ok
190
+ squishling params: { generationConfig: { responseMimeType: "text/plain" } } # ConfigurationError
191
+ ```
192
+
169
193
  - If a provider rejects a param, the call raises `Squishling::ConfigurationError` naming the params. It isn't
170
194
  retried or sent to `squish_fallback`. See [Failure handling](failures.md).
171
195
 
172
196
  ## Inheritance
173
197
 
174
- Subclasses inherit the model or escalation, provider, generation params (merged key by key), instructions, output schema,
198
+ Subclasses inherit the model or escalation and its provider (as a pair), harness, generation params (merged key by key), purpose, output schema,
175
199
  `squish_when` predicate, `squish_context` names, `squish_validate`, `squish_fallback`, and every `squish` declaration. Overrides in a subclass, including
176
200
  overridden methods, are routed the same way.
177
201
 
178
- `append_instructions` sections are added to, not replaced: a subclass's sections follow its parent's, and
179
- `append_instructions false` drops the inherited ones. See [Appending to the instructions](routing.md#appending-to-the-instructions).
202
+ `append_to_purpose` sections are added to, not replaced: a subclass's sections follow its parent's, and
203
+ `append_to_purpose false` drops the inherited ones. See [Appending to the purpose](routing.md#appending-to-the-purpose).
180
204
 
181
205
  ## Per-call overrides
182
206
 
183
- Inside a squished method, `squish!` sends the call to the LLM with its own `instructions:`,
184
- `append_instructions:`, `context:`, `model:` or `escalation:` (with `provider:`), and `params:`. Each layers over the method and class
207
+ Inside a squished method, `squish!` sends the call to the LLM with its own `purpose:`,
208
+ `append_to_purpose:`, `context:`, `model:` or `escalation:` (with `provider:`), `params:`, and `harness:`. Each layers over the method and class
185
209
  settings the same way they layer over each other. See [Handing off to the LLM](routing.md#handing-off-to-the-llm-with-squish).
data/docs/failures.md CHANGED
@@ -15,14 +15,73 @@ A failed attempt moves on to the next one; once the last fails, the error is rai
15
15
  | Malformed or truncated JSON | Moves to the next attempt, then raises `InvalidOutputError`. JSON wrapped in a markdown code fence is accepted. |
16
16
  | JSON that doesn't match the schema (wrong types, missing or extra keys, `null` in a non-`optional` field, a broken [conditional rule](schemas.md#contracts-beyond-the-schema), root not an object) | Moves to the next attempt with the validation errors, then raises `InvalidOutputError` |
17
17
  | Schema-valid output that [`squish_validate`](#output-checks-squish_validate) rejects | Moves to the next attempt with your messages, then raises `InvalidOutputError` |
18
+ | With a [squishsum or ensemble harness](harnesses.md): two valid samples that differ, with no judge or a judge that rejects both | Raises `Squishling::DisagreementError` (an `InvalidOutputError`), carrying both typed `candidates` |
18
19
 
19
20
  Squishling validates output itself with [json_schemer](https://github.com/davishmcclurg/json_schemer), because
20
21
  RubyLLM doesn't, and some providers don't enforce strict mode. When the next attempt is on the same step (same
21
- model, provider, and params), the re-ask happens in the same conversation, so the model sees what it got wrong. A
22
- different step gets a fresh chat with the original input plus the rejected output and its errors, so that output is
23
- sent to the next step's provider, which may not be the one that produced it. `InvalidOutputError` exposes `errors`,
24
- `raw` (the last response), `attempts`, and `models` (the model tried on each attempt). With a `logger` configured,
25
- every escalation is logged as a warning.
22
+ model, provider, params, and `forward_rejected:`), the re-ask happens in the same conversation, so the model sees
23
+ what it got wrong. A different step gets a fresh chat with the original input plus the rejected output (truncated to
24
+ 4,000 characters) and its errors, so that output is sent to the next step's provider, which may not be the one that
25
+ produced it. Set `forward_rejected: false` on a step to leave both out. `InvalidOutputError` exposes `errors`, `raw`
26
+ (the last response), `attempts`, and `models` (the model tried on each attempt); for a
27
+ [`DisagreementError`](harnesses.md#when-the-harness-fails), `raw` holds both candidates, `attempts` is `nil`, and
28
+ `models` has one entry per role (sample a, sample b, then the judge). With a `logger` configured, every
29
+ escalation is logged as a warning.
30
+
31
+ ### What ends up in errors and logs
32
+
33
+ Model output can echo sensitive input, so the raw response is kept in one place: `InvalidOutputError#raw` (plus the opt-in `squawk` hook below).
34
+ `InvalidOutputError#message`, `#errors`, and `config.logger` warnings are built from these:
35
+
36
+ - For unparseable JSON, only the position (`response was not valid JSON (at line 1 column 7)`), never the parser's
37
+ snippet of the response.
38
+ - JSON Schema errors, which name the schema path and the rule that failed. They never contain the rejected value, but
39
+ they do name an extra key the model added (`object property at /foo is a disallowed additional property`), and
40
+ that key name is chosen by the model.
41
+ - Your own `squish_validate` messages, verbatim. If you interpolate result values into them, those values reach
42
+ the message, the logger, and the next attempt's prompt.
43
+ - A fixed message when a response contains a value JSON can't represent (`1e400` parses to `Infinity`, or a string
44
+ with invalid UTF-8).
45
+ - At most the first 20 of these errors, plus a count of the rest, so a very large malformed response can't flood
46
+ the retry message or the log.
47
+ - For a [`DisagreementError`](harnesses.md#when-the-harness-fails), only what Squishling wrote: a `:judgment` judge's
48
+ choice and probability. A chat judge's `reason` is model-written, so it is available as `#reason` but never in the
49
+ message or logs.
50
+
51
+ The same applies to anything you log yourself, such as `error.message` in a [fallback](#fallbacks). Use `error.raw`
52
+ only where you are willing to store model output. To send the output of every attempt somewhere on purpose, use
53
+ [`squawk`](#observing-every-attempt-squawk).
54
+
55
+ ## Observing every attempt (`squawk`)
56
+
57
+ `squawk` is an opt-in hook for sending what the model returned to an error tracker or tracing tool. It runs after
58
+ every LLM attempt, accepted or rejected, and does nothing unless you set it:
59
+
60
+ ```ruby
61
+ Squishling.configure do |config|
62
+ config.squawk = lambda do |output:, metadata:, error:|
63
+ ObservabilitySolution.record(output, metadata, error) if error
64
+ end
65
+ end
66
+ ```
67
+
68
+ | Keyword | Value |
69
+ |---|---|
70
+ | `output` | The raw response: a Hash or String, or `nil` when the provider call itself failed |
71
+ | `error` | `nil` for an accepted attempt. Otherwise the `InvalidOutputError` (with `errors` and `raw`) or `LLMError` that ended the attempt |
72
+ | `metadata` | `label` (`"Class#method"`), `attempt`, `attempts` (the escalation's length; under a harness, the sample's or judge's own attempts), `final` (the last attempt), `model`, `provider`, `params`, `input` (the JSON sent: arguments and named context only), `usage` (token counts, when RubyLLM reports them) |
73
+
74
+ - Set it globally with `config.squawk`, per class with `squishling squawk: ...`, or per method with
75
+ `squish :triage, squawk: ...`. The method's hook wins over the class's, which wins over the configured one;
76
+ `squawk: false` silences an inherited hook. Subclasses and `squish!` calls inherit it.
77
+ - Any object that responds to `call` works. The hook receives only the keywords it declares, unless it takes `**`,
78
+ so a lambda that wants just `error:` is fine and new metadata fields won't break it.
79
+ - It runs inline, so keep it quick. Exceptions it raises propagate unchanged and fail the call, like any of your own
80
+ code, so rescue inside the hook if an outage in your tracing tool shouldn't.
81
+ - `output` is the same object the result is built from, so treat it as read-only.
82
+ - Under a [harness](harnesses.md), it runs for every sample and judge attempt too, each with its own `input`.
83
+ - It isn't called for a `ConfigurationError`, for deterministic and fallback returns, or for an attempt where your own
84
+ `squish_validate` raises (that exception propagates first).
26
85
 
27
86
  ## Output checks (`squish_validate`)
28
87
 
@@ -57,16 +116,17 @@ end
57
116
  | Error | Raised when |
58
117
  |---|---|
59
118
  | `Squishling::Error` | Base class for everything below. Raised directly for misuse at call time, e.g. `result` or `squish!` called outside a squished method |
60
- | `Squishling::ConfigurationError` | Missing instructions or schema, a non-strict schema, invalid or reserved params, an invalid `model:`/`escalation:` declaration, an invalid `append_instructions` item or unavailable source, bad credentials, an unknown model, a request the provider rejects (400) |
119
+ | `Squishling::ConfigurationError` | Missing purpose or schema, a non-strict schema, invalid or reserved params, an invalid `model:`/`escalation:` declaration, an invalid `append_to_purpose` item or unavailable source, bad credentials, an unknown model, a request the provider rejects (400) |
61
120
  | `Squishling::InvalidOutputError` | LLM output still invalid (schema or `squish_validate`) after every attempt in the escalation, or a deterministic/fallback return that doesn't match the schema |
121
+ | `Squishling::DisagreementError` | A subclass of `InvalidOutputError`: a [squishsum or ensemble harness](harnesses.md#when-the-harness-fails)'s samples disagreed and no judge accepted either. `candidates`, `verdict`, and `reason` describe what happened |
62
122
  | `Squishling::LLMError` | The provider call failed on the last attempt in the escalation, after RubyLLM's own retries, including context-length errors (`cause` holds the original) |
63
123
 
64
124
  Errors raised by your own Ruby code are not wrapped.
65
125
 
66
126
  ## Fallbacks
67
127
 
68
- Use `squish_fallback` to decide what happens when every attempt fails, with `InvalidOutputError` or
69
- `LLMError`.
128
+ Use `squish_fallback` to decide what happens when the LLM path fails with `InvalidOutputError` (including a
129
+ `DisagreementError`) or `LLMError`. It's the one fallback for every [harness](harnesses.md).
70
130
  It receives the error plus the method's inputs as keywords, and runs against the instance:
71
131
 
72
132
  ```ruby