squishling 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/docs/harnesses.md ADDED
@@ -0,0 +1,198 @@
1
+ # Harnesses: escalation, squishsum, and ensemble
2
+
3
+ A squished call's **harness** decides how it uses its [escalation](configuration.md#models-and-escalation):
4
+
5
+ | Harness | What it does | Requests per call |
6
+ |---|---|---|
7
+ | `:escalation` (default) | Tries each attempt in order until an output passes the schema and `squish_validate`. | 1 or more |
8
+ | `:squishsum` | Asks the escalation's **first step** twice, concurrently. Identical outputs are accepted; different ones raise `Squishling::DisagreementError`. | 2 or more |
9
+ | `:judged_squishsum` | Like `:squishsum`, but when the two outputs differ, a **judge** picks one or rejects both. | 2 or more, plus 1 or more judge requests when they differ |
10
+ | `:ensemble` | Asks the escalation's **first and second steps** once each, concurrently, so two different models are compared. Identical outputs are accepted; different ones raise `DisagreementError`. | 2 or more |
11
+ | `:judged_ensemble` | Like `:ensemble`, but when the two outputs differ, a **judge** picks one or rejects both. | 2 or more, plus 1 or more judge requests when they differ |
12
+
13
+ Use a squishsum harness when a confidently wrong answer costs more than a second request: one sample can make a
14
+ mistake, but two independent samples rarely make the same one. Use an [ensemble](#ensembles) when the two samples
15
+ should come from different models: a model rarely disagrees with itself at low temperature, but a cheap and a
16
+ stronger model often disagree on exactly the cases the cheap one gets confidently wrong.
17
+
18
+ ```ruby
19
+ class TicketTriager
20
+ include Squishling
21
+
22
+ squishling escalation: [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5"],
23
+ harness: :judged_squishsum # two Haiku samples; Sonnet judges when they differ
24
+ purpose "Assign a priority and a team."
25
+ output_schema do
26
+ string :priority, enum: %w[low medium high]
27
+ string :team
28
+ end
29
+ end
30
+ ```
31
+
32
+ ## Declaring a harness
33
+
34
+ `harness:` takes a type, or a Hash with `type:` and its options:
35
+
36
+ ```ruby
37
+ Squishling.configure { |c| c.default_harness = :squishsum } # every class without its own
38
+
39
+ squishling harness: :judged_squishsum # this class (and subclasses)
40
+ squish :triage, harness: { type: :squishsum, # one method
41
+ compare: ->(a, b, **) { a.priority == b.priority } }
42
+ squish!(harness: :escalation) # one call (see routing.md)
43
+ ```
44
+
45
+ The first level that declares one wins: `squish!`, then the method, then the class, then
46
+ `Squishling.config.default_harness`, then `:escalation`. Options aren't merged across levels.
47
+
48
+ | Option | Harnesses | Description |
49
+ |---|---|---|
50
+ | `type:` | all | `:escalation`, `:squishsum`, `:judged_squishsum`, `:ensemble`, or `:judged_ensemble` (required in the Hash form). |
51
+ | `compare:` | squishsum and ensemble ones | `->(a, b, **inputs) { ... }`, evaluated against the instance with the two typed results and the method's inputs. Truthy means the samples agree. Default: the two outputs are exactly equal. |
52
+ | `judge:` | judged ones | The judge step (see [The judge](#the-judge)). Default: the escalation's next step after the sampled ones. |
53
+ | `judge_instructions:` | judged ones | A String, or a Proc evaluated against the instance, that replaces the [default judge prompt](#the-default-judge-prompt). |
54
+
55
+ ## Samples
56
+
57
+ Under `:squishsum`, both samples run on the escalation's first step, with its model, provider, and params
58
+ (an [ensemble](#ensembles) runs them on its first and second steps):
59
+
60
+ - **They're independent.** Each sample has its own chat, and the two requests are sent concurrently on separate
61
+ threads. Only the provider requests run on those threads. Parsing, schema validation, `squish_validate`,
62
+ `compare:`, and the judge all run on the calling thread, so your Squishling callbacks never run concurrently.
63
+ RubyLLM instrumentation subscribers (e.g. `ActiveSupport::Notifications` listeners for `chat.ruby_llm`) do see
64
+ each sample's request on its worker thread, outside the Rails executor, so keep them thread-safe.
65
+ - **Each one is retried like an escalation step.** A sample with invalid output is re-asked in its own chat, up to
66
+ their step's `attempts:`, while a sample that already passed keeps its result. A failed request gets a fresh
67
+ chat. A sample still failing after its attempts fails the call with `InvalidOutputError` or `LLMError`, as
68
+ usual. Under `:squishsum`, later escalation steps are never used for samples (see [Ensembles](#ensembles)).
69
+ - **Agreement is exact by default.** Free-text fields seldom match word for word, so pass `compare:` to decide
70
+ which fields must match:
71
+
72
+ ```ruby
73
+ squishling harness: { type: :squishsum, compare: ->(a, b, **) { a.priority == b.priority && a.team == b.team } }
74
+ ```
75
+
76
+ When the samples agree, the first one is returned.
77
+
78
+ ## Ensembles
79
+
80
+ `:ensemble` and `:judged_ensemble` sample two *different* steps instead of one step twice. Sample `a` is the
81
+ escalation's first step and sample `b` its second, each with its own model, provider, and params:
82
+
83
+ ```ruby
84
+ class TicketTriager
85
+ include Squishling
86
+
87
+ squishling escalation: [{ model: "claude-haiku-4-5", attempts: 2 }, "claude-sonnet-5-5", "claude-opus-5-5"],
88
+ harness: :judged_ensemble # a = Haiku, b = Sonnet, Opus judges when they differ
89
+ # purpose, output_schema, ...
90
+ end
91
+ ```
92
+
93
+ - A step's `attempts:` are that sample's retries (here Haiku gets two attempts, Sonnet one), exactly as in
94
+ [Samples](#samples). Steps after the second are never sampled.
95
+ - Everything else matches squishsum: independence and concurrency, `compare:`, `DisagreementError`, and the
96
+ judge. `DisagreementError#models` names the model behind each sample.
97
+ - The judge defaults to the escalation's **third** step. With only two steps, a `:judged_ensemble` needs `judge:`.
98
+ An ensemble with fewer than two steps, or a judged one with neither a third step nor `judge:`, raises
99
+ `ConfigurationError` before any request.
100
+ - A global `default_harness = :ensemble` therefore needs every class to declare at least two steps, and a
101
+ `squish!(model: ...)` override (one step) can't be combined with an ensemble harness.
102
+ - Steps are told apart by value, so two adjacent steps with the same model, provider, and params count as one step.
103
+
104
+ ## The judge
105
+
106
+ With a judged harness, two valid samples that disagree go to a judge. The judge picks candidate `a` or `b`,
107
+ whose typed result is returned (`squished?` is `true`), or it picks `neither`, which raises
108
+ `DisagreementError`.
109
+
110
+ The judge is:
111
+
112
+ 1. the `judge:` step, if declared: a model name, or a Hash with `model:`, `provider:`, `params:`, and `attempts:`
113
+ (as in an escalation step, without `order:`), plus `type:`; or
114
+ 2. the escalation's next step after the sampled ones (the second under `:judged_squishsum`, the third under
115
+ `:judged_ensemble`), with its attempts.
116
+
117
+ A judged harness with neither raises `ConfigurationError` before any sample is requested. A `judge:`
118
+ step uses only its own `provider:`; it doesn't inherit the class or method `provider:`, which belongs to their models
119
+ (a chat judge's params, though, do merge over the method's).
120
+
121
+ ### Chat judges
122
+
123
+ By default the judge is a chat model, asked for a strict `{ "verdict": "a" | "b" | "neither", "reason": "..." }`.
124
+ It gets the same context the samples had, plus the samples themselves:
125
+
126
+ - **System prompt:** the judge prompt, then the operation's own purpose (including any `append_to_purpose`
127
+ sections).
128
+ - **User message:** JSON with the samples' `input` (`arguments` and `context`), the `output_schema`, and the
129
+ `candidates` `a` and `b`.
130
+
131
+ A verdict that doesn't match its schema is retried per the judge step's `attempts:`; a judge that never returns a
132
+ valid verdict raises `InvalidOutputError`. A chat judge's params are the method's
133
+ [generation params](configuration.md#generation-params) with its own `params:` merged over them.
134
+
135
+ ### System One judgment models (Jev)
136
+
137
+ A judge can also be a System One decision model, such as TypeSafe's Jev, via
138
+ [RubyLLM judgments](https://rubyllm.com/judgments/) (RubyLLM 2.1+). Decision models return calibrated
139
+ probabilities rather than text, so they are much faster and cheaper than a chat judge.
140
+
141
+ ```ruby
142
+ squishling harness: {
143
+ type: :judged_squishsum,
144
+ judge: { model: "jev-latest", type: :judgment, min_confidence: 0.8 }
145
+ }
146
+ ```
147
+
148
+ Squishling asks one `choice` question (`a`, `b`, or `neither`), with the judge prompt as its instructions and the
149
+ operation's purpose, input, output schema, and candidates as the judgment input. A pick of `a` or `b` is
150
+ accepted only when its probability is at least `min_confidence:` (default `0.8`). Otherwise, as with `neither`,
151
+ `DisagreementError` is raised with the choice and its probability in its `reason`. A judgment judge sends only its own
152
+ `params:` (as RubyLLM `provider_options:`). `attempts:` retries a failed request. Configure the provider's
153
+ credentials in RubyLLM as usual.
154
+
155
+ ### The default judge prompt
156
+
157
+ ```text
158
+ You are an impartial judge. Two independent attempts at the same operation returned different outputs,
159
+ candidate "a" and candidate "b". Both already match the required output format, so judge only their
160
+ content. You are given the operation's purpose and its input. Decide which candidate correctly and
161
+ faithfully carries out the purpose for this input. If both do (for example, they differ only in
162
+ wording), choose either one, preferring "a". Choose "neither" only when both are wrong or you can't tell
163
+ whether either is correct. Never combine them or invent a third answer.
164
+ ```
165
+
166
+ It is available as `Squishling::Harness::DEFAULT_JUDGE_INSTRUCTIONS`. Replace it with `judge_instructions:`:
167
+
168
+ ```ruby
169
+ squishling harness: { type: :judged_squishsum,
170
+ judge_instructions: "Pick the candidate whose priority follows our SLA rules: ..." }
171
+ ```
172
+
173
+ ## When the harness fails
174
+
175
+ A disagreement raises `Squishling::DisagreementError`, a subclass of `InvalidOutputError`. It's raised when the
176
+ samples differ with no judge, or when the judge rejects both. Like any failure on the LLM path, it goes to
177
+ [`squish_fallback`](failures.md#fallbacks), which can still return one of the candidates:
178
+
179
+ ```ruby
180
+ squish_fallback do |error, **|
181
+ raise error unless error.is_a?(Squishling::DisagreementError)
182
+
183
+ error.candidates.first # already typed and validated; squished? is true
184
+ end
185
+ ```
186
+
187
+ | Attribute | Value |
188
+ |---|---|
189
+ | `candidates` | The two typed results, `[a, b]` |
190
+ | `verdict` | `nil` when there was no judge; `:neither` when the judge rejected both |
191
+ | `reason` | The judge's reason (or, for a judgment model, the choice and its probability). A chat judge's reason is model-written, so it is not part of `message` or the logs; a judgment model's choice and probability are |
192
+ | `raw` | Both candidates as Hashes, `[a.to_h, b.to_h]` |
193
+ | `attempts` | `nil` |
194
+ | `models` | The model behind each role: sample `a`, sample `b`, then the judge if one ran (one entry each, not one per attempt) |
195
+
196
+ Other failures are unchanged. A sample or judge that never produces valid output raises `InvalidOutputError`, a
197
+ failed request on the last attempt raises `LLMError`, and setup mistakes (including a provider 400 on either
198
+ sample) raise `ConfigurationError`, which is never sent to the fallback.
@@ -0,0 +1,148 @@
1
+ # Measuring token spend
2
+
3
+ A squished path costs tokens on every call. A Ruby path costs engineer time once, plus upkeep. Measuring the
4
+ first tells you when the second is worth paying.
5
+
6
+ ## What Squishling gives you
7
+
8
+ - `result.squished?` says which path served a call, so you can count LLM calls per class.
9
+ - The [`squawk` hook](failures.md#observing-every-attempt-squawk) runs after every LLM attempt and receives the
10
+ method (`label`, e.g. `"ShipmentUpdate#call"`), the model and provider, which attempt it was, and the token usage
11
+ RubyLLM reported. That is enough to attribute spend to a squishling method without any other tooling.
12
+ - Cost in dollars comes from RubyLLM's model pricing. Multiply the reported tokens by your model's price, or read it
13
+ from RubyLLM directly (see [the instrumenter](#lower-level-rubyllms-instrumenter)).
14
+
15
+ ## Deciding when to write the code
16
+
17
+ For each squished path, over a month:
18
+
19
+ ```
20
+ token cost = calls × (input tokens × input price + output tokens × output price)
21
+ code cost = time to write it + upkeep (new cases, format changes, on-call)
22
+ ```
23
+
24
+ Write the Ruby when the token cost keeps exceeding the code cost, and narrow `squish_when` so only the inputs
25
+ Ruby can't handle still go to the LLM. See [Hardening a path](routing.md#hardening-a-path). If the token cost
26
+ stays small, the debt you haven't taken on is cheaper than the code.
27
+
28
+ ## Tools
29
+
30
+ | Option | Good for | Setup |
31
+ |---|---|---|
32
+ | [Coolhand Labs](#coolhand-labs) | Picking up your squishling calls automatically and measuring their cost and accuracy | The `coolhand` gem and an initializer |
33
+ | [Custom: the `squawk` hook](#custom-the-squawk-hook) | Logging or metrics in your own stack, with no new dependency | A lambda and one config line |
34
+ | [OpenTelemetry](#opentelemetry) | Sending traces to Datadog, Langfuse, Arize, Braintrust, LangSmith, or any OTel backend | The OpenTelemetry SDK, an exporter, and one line |
35
+ | [LangSmith](#langsmith) | Browsing traces and token usage in LangSmith | Through OpenTelemetry |
36
+
37
+ ### Coolhand Labs
38
+
39
+ Add and initialize the [`coolhand`](https://github.com/Coolhand-Labs/coolhand-ruby) gem. It intercepts the LLM
40
+ requests your app makes, so it picks up your squishling calls automatically and measures their cost, and their
41
+ accuracy through Coolhand's feedback tools. No changes to your squishling classes.
42
+
43
+ ```ruby
44
+ # Gemfile
45
+ gem "coolhand"
46
+
47
+ # config/initializers/coolhand.rb
48
+ Coolhand.configure do |config|
49
+ config.api_key = ENV.fetch("COOLHAND_API_KEY")
50
+ end
51
+ ```
52
+
53
+ Get an API key at [coolhandlabs.com](https://coolhandlabs.com/). To keep the data in your own infrastructure, point
54
+ `config.base_url` at a Coolhand-compatible endpoint of your own.
55
+
56
+ ### Custom: the `squawk` hook
57
+
58
+ `config.squawk` runs after every LLM attempt, accepted or rejected, and is handed the method's label, the model,
59
+ and the token usage RubyLLM reported. A lambda can take just the keywords it needs:
60
+
61
+ ```ruby
62
+ Squishling.configure do |config|
63
+ config.squawk = lambda do |metadata:, **|
64
+ usage = metadata[:usage] || {} # nil when the provider call failed; counts it didn't report are absent
65
+ Rails.logger.info("squishling #{metadata[:label]} model=#{metadata[:model]} " \
66
+ "attempt=#{metadata[:attempt]}/#{metadata[:attempts]} " \
67
+ "in=#{usage[:input_tokens]} out=#{usage[:output_tokens]}")
68
+ end
69
+ end
70
+ ```
71
+
72
+ Each line carries the `"Class#method"` that made the call, so grouping by label gives you tokens per squishling. A call
73
+ that stays in Ruby makes no LLM request and never reaches the hook, so a label that stops appearing is a path that has
74
+ been hardened. You can also set the hook per class (`squishling squawk: ...`) or per method. See
75
+ [Observing every attempt](failures.md#observing-every-attempt-squawk) for every keyword and the rules for exceptions.
76
+
77
+ ### Lower level: RubyLLM's instrumenter
78
+
79
+ To meter every RubyLLM request in your app, not only Squishling's, use RubyLLM's own instrumentation. It reports a
80
+ `usage.ruby_llm` event for each provider request, with `model`, `provider`, `status`, `tokens` (`input`, `output`,
81
+ and more) and `cost`. Give it an object that responds to `instrument(name, payload)` and optionally takes a block. In
82
+ Rails, `ActiveSupport::Notifications` is used automatically.
83
+
84
+ ```ruby
85
+ class TokenMeter
86
+ def instrument(name, payload = {})
87
+ if name == "usage.ruby_llm"
88
+ Rails.logger.info("llm model=#{payload[:model]} workflow=#{payload[:workflow_name]} " \
89
+ "in=#{payload[:tokens].input} out=#{payload[:tokens].output}")
90
+ end
91
+ block_given? ? yield : nil
92
+ end
93
+ end
94
+
95
+ RubyLLM.configure { |config| config.instrumenter = TokenMeter.new }
96
+ ```
97
+
98
+ These events don't know which squishling made the request. To attribute the spend, wrap the call in a RubyLLM
99
+ workflow, and every event emitted inside the block carries its name, id and any `metadata:` you pass. (The
100
+ `squawk` hook above needs no wrapping.)
101
+
102
+ ```ruby
103
+ RubyLLM.workflow("ShipmentUpdate", metadata: { carrier: "acme-freight" }) do
104
+ ShipmentUpdate.call(carrier: "acme-freight", payload: body)
105
+ end
106
+ ```
107
+
108
+ ### OpenTelemetry
109
+
110
+ RubyLLM 2.1+ ships its own OpenTelemetry tracer. Each request becomes a span with the model, request settings and
111
+ token usage (`gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`), following the OpenTelemetry GenAI
112
+ conventions. It does not record message text. Your app owns the SDK and exporters.
113
+
114
+ ```ruby
115
+ # Gemfile
116
+ gem "opentelemetry-sdk"
117
+ gem "opentelemetry-exporter-otlp" # or your backend's exporter
118
+
119
+ # config/initializers/opentelemetry.rb
120
+ OpenTelemetry::SDK.configure # reads the standard OTEL_* environment variables
121
+ RubyLLM::OpenTelemetry.enable # starts tracing RubyLLM requests
122
+ ```
123
+
124
+ Spans don't say which squishling made the request, so pair this with a `RubyLLM.workflow` (see
125
+ [above](#lower-level-rubyllms-instrumenter)) or group by model and prompt in your backend.
126
+
127
+ ### LangSmith
128
+
129
+ LangSmith accepts OpenTelemetry traces, so the OpenTelemetry setup above covers it: point the OTLP exporter at
130
+ LangSmith. See LangSmith's
131
+ [OpenTelemetry docs](https://docs.smith.langchain.com/observability/how_to_guides/tracing/trace_with_opentelemetry)
132
+ for the endpoint and headers. A community Ruby gem, [`langsmith-sdk`](https://github.com/felipekb/langsmith-ruby-sdk),
133
+ also exists; it isn't published by LangChain.
134
+
135
+ ## Privacy
136
+
137
+ Squishling only sends the model what you name: the method arguments, any `squish_context` values, and any source you
138
+ append. Tools that sit on the request path are a separate route out of your process, and they see both sides:
139
+
140
+ - **Request interceptors, such as `coolhand`,** record the prompts (your purpose text and those inputs) and the
141
+ model's responses. Coolhand sends them to coolhandlabs.com unless you point `config.base_url` at an endpoint of your
142
+ own.
143
+ - **A `squawk` hook that forwards `output:` or `metadata[:input]`** sends model output and your inputs wherever you
144
+ send them. That is the sanctioned route for model output; see
145
+ [Observing every attempt](failures.md#observing-every-attempt-squawk). The examples above log only metadata.
146
+ - **OpenTelemetry** with RubyLLM's tracer records usage and request settings, not message text.
147
+
148
+ Check what your tool stores before pointing it at inputs that contain personal data.
data/docs/naming.md ADDED
@@ -0,0 +1,43 @@
1
+ # Naming and collisions
2
+
3
+ `include Squishling` adds a small set of methods to your class. This page lists them and explains what happens
4
+ when one of those names is already taken.
5
+
6
+ ## What Squishling adds
7
+
8
+ | Where | Names |
9
+ |---|---|
10
+ | Class-level DSL | `squishling`, `purpose`, `append_to_purpose`, `output_schema`, `squish`, `squish_when`, `squish_fallback`, `squish_validate`, `squish_context`, `squished_methods`, `call`, and the `squishling_*` readers |
11
+ | Instance | `squishling_result`, `squish!`, and `result` (an alias for `squishling_result`) |
12
+
13
+ Methods your class defines itself always win over these, whether you define them before or after the `include`.
14
+ Only methods your class *inherits* can be shadowed.
15
+
16
+ ## Inherited collisions raise
17
+
18
+ If a parent class or an earlier-included module already defines a class-level DSL method, `squish!` or
19
+ `squishling_result`, `include Squishling` raises `Squishling::ConfigurationError` listing every clash and where
20
+ it comes from:
21
+
22
+ ```
23
+ ConfigurationError: MyApp inherits methods that Squishling would override: MyApp.call (from Sinatra::Base's
24
+ class methods). Include Squishling in a plain Ruby class instead ...
25
+ ```
26
+
27
+ This is deliberate. Squishling's modules sit ahead of inherited ones, so without the check
28
+ `Sinatra::Base.call(env)` (the Rack entry point), or any `call` or `purpose` your parent class defines, would be
29
+ replaced silently. Put Squishling in a plain Ruby class and have your framework object call it instead.
30
+
31
+ ## `result` is skipped when taken
32
+
33
+ `result` is a convenience for `squishling_result`. If the class already has a `result` (its own or inherited), it is
34
+ left alone and Squishling does not add the alias. Use `squishling_result(...)` instead; it is always available.
35
+
36
+ **ActiveRecord caveat.** Column readers are generated lazily, so Squishling can't see a `result` column when you
37
+ include it. Its `result` would then shadow the column reader. On a model with a `result` column, call
38
+ `squishling_result` (or keep Squishling out of the model and use a plain class).
39
+
40
+ ## Not a collision
41
+
42
+ ActiveSupport adds `String#squish` and `String#squish!`. Those live on `String`; Squishling's `squish!` lives on
43
+ your class, so they never meet. They only look alike when reading code in a Rails app.
data/docs/routing.md CHANGED
@@ -35,24 +35,27 @@ end
35
35
  def summarize(text) = raise NotImplementedError # elastic until someone writes it
36
36
  ```
37
37
 
38
- When the Ruby implementation runs, a `Hash` it returns is validated against the schema and turned into the
39
- same typed result the LLM path produces. `result(...)` (alias `squishling_result`) does the same explicitly.
40
- Invalid deterministic output raises `Squishling::InvalidOutputError` too, so a hardened path can't silently
41
- drift from the contract.
38
+ When the Ruby implementation runs, whatever it returns (a `Hash`, `nil`, a string, another schema's result, …)
39
+ is validated against the schema and turned into the same typed result the LLM path produces.
40
+ `squishling_result(...)` does the same explicitly; `result(...)` is a convenience alias, skipped when the class
41
+ already has a `result` (see [Naming and collisions](naming.md)). Invalid deterministic output raises
42
+ `Squishling::InvalidOutputError` too, so a hardened path can't silently drift from the contract.
42
43
 
43
44
  A `NotImplementedError` raised anywhere inside the method, including from code it calls, also routes to the
44
45
  LLM.
45
46
 
46
- Each call is routed on its own, including a squished method that calls itself on smaller inputs. A subclass
47
+ Each call is routed on its own, including a squished method that calls itself on smaller inputs (those
48
+ recursive calls must return schema-valid values too). A subclass
47
49
  override that calls `super` is one call: it's routed once, at the subclass. A call to the method from its own
48
- `squish_when` or `squish_fallback` (or an instructions proc) isn't routed again: it runs the Ruby
49
- implementation, so a fallback can hand the input back to Ruby with `call(**inputs)`.
50
+ `squish_when` or `squish_fallback` (or a purpose proc) isn't routed again: it runs the Ruby
51
+ implementation, so a fallback can hand the input back to Ruby with `call(**inputs)`. These inner calls
52
+ return a `Hash` as the typed result and any other value unchanged; only the outermost return is validated in full.
50
53
 
51
54
  ## Hardening a path
52
55
 
53
56
  This is the workflow from [Elastic Software](https://everythingengineer.substack.com/p/beginners-write-software-with-ai):
54
57
 
55
- 1. Ship the class with instructions and a schema but no implementation. Every call goes to the LLM.
58
+ 1. Ship the class with purpose and a schema but no implementation. Every call goes to the LLM.
56
59
  2. Watch which inputs carry the volume. `squished?` on each result tells you which path served it.
57
60
  3. Write Ruby for the high-volume cases and narrow `squish_when` so only the rest go to the LLM.
58
61
 
@@ -68,7 +71,7 @@ class TicketTriager
68
71
 
69
72
  squish_context :customer_tier, :product # instance state sent alongside the arguments
70
73
 
71
- squish :triage, instructions: "Assign a priority and team.", model: "claude-haiku-4-5" do
74
+ squish :triage, purpose: "Assign a priority and team.", model: "claude-haiku-4-5" do
72
75
  string :priority, enum: %w[low med high]
73
76
  string :team
74
77
  end
@@ -83,10 +86,11 @@ end
83
86
 
84
87
  `squish` accepts:
85
88
 
86
- - `instructions:` (replaces the class's)
87
- - `append_instructions:` (added to the class's, see [Appending to the instructions](#appending-to-the-instructions))
89
+ - `purpose:` (replaces the class's)
90
+ - `append_to_purpose:` (added to the class's, see [Appending to the purpose](#appending-to-the-purpose))
88
91
  - `output_schema:` (or a schema block)
89
92
  - `model:` or `escalation:`, `provider:`, and `params:` (generation params; see [Configuration](configuration.md))
93
+ - `harness:` (see [Harnesses](harnesses.md))
90
94
  - `when:`, a predicate proc
91
95
  - `validate:` and `fallback:` (see [Failure handling](failures.md))
92
96
 
@@ -103,8 +107,8 @@ from the method.
103
107
  class InvoiceParser
104
108
  include Squishling
105
109
 
106
- instructions "Extract invoice fields from the client's raw data."
107
- append_instructions "Here is the Ruby that parses well-formed invoices, for context on the logic and goals:",
110
+ purpose "Extract invoice fields from the client's raw data."
111
+ append_to_purpose "Here is the Ruby that parses well-formed invoices, for context on the logic and goals:",
108
112
  self
109
113
  output_schema do
110
114
  string :invoice_number
@@ -115,7 +119,7 @@ class InvoiceParser
115
119
  parsed = AcmeParser.parse(data)
116
120
  result(invoice_number: parsed.id, total: parsed.sum)
117
121
  rescue AcmeParser::ParseError => e
118
- squish!(append_instructions: "The Ruby parser above failed on this input; the error is in the context.",
122
+ squish!(append_to_purpose: "The Ruby parser above failed on this input; the error is in the context.",
119
123
  context: { parse_error: e })
120
124
  end
121
125
  end
@@ -126,9 +130,10 @@ end
126
130
  | Option | Effect |
127
131
  |---|---|
128
132
  | `context:` | A Hash sent under `"context"` with any `squish_context` values (a same-named key wins). Exceptions are sent as `{ "class", "message" }`, never their backtrace. |
129
- | `append_instructions:` | Added to the declared sections; `false` (alone or first in an Array) drops them for this call |
130
- | `instructions:` | Replaces the instructions |
133
+ | `append_to_purpose:` | Added to the declared sections; `false` (alone or first in an Array) drops them for this call |
134
+ | `purpose:` | Replaces the purpose |
131
135
  | `model:` or `escalation:`, `provider:`, `params:` | E.g. send this call to a stronger model, or a whole [escalation](configuration.md#models-and-escalation), when Ruby fails. A `provider:` needs a `model:` or `escalation:`; `params:` merge key by key over the declared ones. |
136
+ | `harness:` | Replaces the declared [harness](harnesses.md), e.g. `harness: :judged_squishsum` to double-check a hand-off |
132
137
 
133
138
  - **The output schema can't be overridden.** The call still returns the method's result type.
134
139
  - **Failures** go through the normal LLM path: a declared `squish_fallback` is used, otherwise
@@ -148,9 +153,9 @@ end
148
153
  or in a parent implementation reached through `super`. Calling it anywhere else, or from a `squish_fallback`
149
154
  (which would loop), raises `Squishling::Error`.
150
155
 
151
- ## Appending to the instructions
156
+ ## Appending to the purpose
152
157
 
153
- `append_instructions` adds sections to the system prompt after the instructions. It takes items, an Array of
158
+ `append_to_purpose` adds sections to the system prompt after the purpose. It takes items, an Array of
154
159
  them, or a block (treated as a Proc item). Each item is one of:
155
160
 
156
161
  | Item | Sent as |
@@ -164,19 +169,19 @@ them, or a block (treated as a Proc item). Each item is one of:
164
169
  class InvoiceParser
165
170
  include Squishling
166
171
 
167
- append_instructions "The Ruby that parses well-formed invoices:", self, AcmeParser
168
- append_instructions -> { "This client's invoices are in #{currency}." }
172
+ append_to_purpose "The Ruby that parses well-formed invoices:", self, AcmeParser
173
+ append_to_purpose -> { "This client's invoices are in #{currency}." }
169
174
 
170
- squish :summarize, append_instructions: false do # no appendices for this method
175
+ squish :summarize, append_to_purpose: false do # no appendices for this method
171
176
  string :summary
172
177
  end
173
178
  end
174
179
  ```
175
180
 
176
- The same option goes in the class-wide call: `squishling append_instructions: ["...", self]`.
181
+ The same option goes in the class-wide call: `squishling append_to_purpose: ["...", self]`.
177
182
 
178
- Sections are added down the chain: class, then subclass, then `squish :name, append_instructions:`, then
179
- `squish!(append_instructions:)`. `false` drops everything declared above it, so `[false, "Only this."]`
183
+ Sections are added down the chain: class, then subclass, then `squish :name, append_to_purpose:`, then
184
+ `squish!(append_to_purpose:)`. `false` drops everything declared above it, so `[false, "Only this."]`
180
185
  replaces it.
181
186
 
182
187
  Source is read with Ruby's own parser (Prism) the first time it's needed and cached. Some limits:
@@ -195,8 +200,8 @@ Source is read with Ruby's own parser (Prism) the first time it's needed and cac
195
200
 
196
201
  ## What the LLM sees
197
202
 
198
- - **System prompt:** your instructions (a String, or a Proc evaluated against the instance), then any
199
- `append_instructions` sections, then a short note describing the input format.
203
+ - **System prompt:** your purpose (a String, or a Proc evaluated against the instance), then any
204
+ `append_to_purpose` sections, then a short note describing the input format.
200
205
  - **User message:** JSON with the method's arguments, mapped to their parameter names. Any
201
206
  `squish_context` values, and a `squish!` call's `context:`, go under `"context"`:
202
207
 
@@ -205,11 +210,15 @@ Source is read with Ruby's own parser (Prism) the first time it's needed and cac
205
210
  "context": { "customer_tier": "enterprise", "product": "API" } }
206
211
  ```
207
212
 
213
+ - **Squishsum and ensemble harnesses:** both samples get exactly this input, and a [judge](harnesses.md#the-judge) gets
214
+ the same purpose and input plus both samples' outputs.
208
215
  - **Retries and escalation:** another attempt of the same step gets the validation errors in the same conversation. A
209
216
  later [escalation](configuration.md#models-and-escalation) step, which may be a different provider (say, local
210
217
  Ollama, then hosted Anthropic Claude), gets the same JSON plus the previous model's rejected output and the
211
218
  errors, including any messages your [`squish_validate`](failures.md#output-checks-squish_validate) check
212
- returned. Don't put data in those messages that you wouldn't send as an argument.
219
+ returned. The rejected output is model-generated and is truncated to 4,000 characters; set
220
+ `forward_rejected: false` on a step to start it from the original input only, without the rejected output or
221
+ its errors. Don't put data in those messages that you wouldn't send as an argument.
213
222
 
214
223
  `squish_context` names are read from a method of that name if there is one, otherwise from the instance
215
224
  variable. Only context you name is sent. Instance variables are never dumped wholesale, so API clients,
data/docs/schemas.md CHANGED
@@ -59,7 +59,8 @@ r.squished? # => true when it came from the LLM
59
59
  ```
60
60
 
61
61
  On the deterministic path, return a `Hash` or build the result with `result(...)`. Either way it's validated
62
- against the schema.
62
+ against the schema, and so is anything else you return (`nil`, a string, another schema's result), which raises
63
+ `Squishling::InvalidOutputError` unless it matches.
63
64
 
64
65
  ## Optional vs. empty
65
66
 
@@ -84,7 +85,7 @@ end
84
85
  branch as usual: `vitals` is a `Data` object or `nil`, and a list of objects is a list of `Data` objects. In a raw
85
86
  JSON Schema, `type: ["array", "null"]` works too. A union with more than one non-null branch is ambiguous, so its
86
87
  values come back as plain hashes. If you need the model to tell "none" apart from "not mentioned", say so in your
87
- instructions.
88
+ purpose.
88
89
 
89
90
  ## Contracts beyond the schema
90
91
 
@@ -110,7 +111,7 @@ Providers' strict modes don't support conditional keywords (`if`/`then`/`else` f
110
111
  `dependentSchemas` from `dependent`), so Squishling keeps them out of the schema it sends and enforces them itself.
111
112
  Output that breaks a rule is rejected like any other invalid output: the call moves on to the next attempt in its
112
113
  [escalation](configuration.md#models-and-escalation) with the errors. The model doesn't see these rules
113
- up front, so state them in your instructions too.
114
+ up front, so state them in your purpose too.
114
115
 
115
116
  For anything a schema can't express (sums, lookups against your data), use
116
117
  [`squish_validate`](failures.md#output-checks-squish_validate).
@@ -1,7 +1,7 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Squishling
4
- # append_instructions: extra system-prompt sections placed after the instructions, layered
4
+ # append_to_purpose: extra system-prompt sections placed after the purpose, layered
5
5
  # class -> subclass -> method -> call. Each level adds to the levels above it; `false` drops them.
6
6
  module Appendices
7
7
  ITEM_TYPES = [String, Module, Method, UnboundMethod, Proc].freeze
@@ -29,7 +29,7 @@ module Squishling
29
29
 
30
30
  values = receiver.instance_exec(&item)
31
31
  (values.is_a?(Array) ? values : [values]).select(&:itself).map do |value|
32
- check!(value, "#{label} append_instructions proc", procs: false)
32
+ check!(value, "#{label} append_to_purpose proc", procs: false)
33
33
  render_item(value)
34
34
  end
35
35
  end
@@ -43,7 +43,7 @@ module Squishling
43
43
  allowed = procs ? ITEM_TYPES : ITEM_TYPES - [Proc]
44
44
  return if allowed.any? { |type| item.is_a?(type) }
45
45
 
46
- raise ConfigurationError, "#{label}: append_instructions items must be Strings, classes or modules, " \
46
+ raise ConfigurationError, "#{label}: append_to_purpose items must be Strings, classes or modules, " \
47
47
  "#{procs ? 'methods, or procs' : 'or methods'} (got #{describe(item)})"
48
48
  end
49
49