llm.rb 13.1.0 → 14.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +320 -0
- data/README.md +340 -31
- data/bin/llm.rb +36 -12
- data/data/anthropic.json +206 -263
- data/data/bedrock.json +2138 -1860
- data/data/deepinfra.json +1003 -624
- data/data/deepseek.json +38 -34
- data/data/google.json +1079 -371
- data/data/mistral.json +448 -368
- data/data/moonshot.json +384 -0
- data/data/openai.json +974 -1343
- data/data/xai.json +154 -126
- data/data/zai.json +191 -191
- data/lib/llm/agent.rb +47 -14
- data/lib/llm/context.rb +71 -88
- data/lib/llm/cost.rb +23 -17
- data/lib/llm/error.rb +0 -8
- data/lib/llm/function/async/task.rb +2 -0
- data/lib/llm/function/fiber/task.rb +2 -0
- data/lib/llm/function/fork/task.rb +2 -0
- data/lib/llm/function/ractor/task.rb +2 -0
- data/lib/llm/function/sequential/group.rb +4 -1
- data/lib/llm/function/sequential/task.rb +1 -1
- data/lib/llm/function/task.rb +4 -0
- data/lib/llm/function/thread/task.rb +2 -0
- data/lib/llm/function.rb +32 -4
- data/lib/llm/guard/loop.rb +89 -0
- data/lib/llm/guard/null.rb +19 -0
- data/lib/llm/guard.rb +61 -0
- data/lib/llm/provider.rb +36 -0
- data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
- data/lib/llm/providers/anthropic.rb +1 -8
- data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
- data/lib/llm/providers/bedrock.rb +1 -8
- data/lib/llm/providers/google/stream_parser.rb +1 -0
- data/lib/llm/providers/google.rb +1 -8
- data/lib/llm/providers/moonshot.rb +76 -0
- data/lib/llm/providers/ollama.rb +1 -8
- data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
- data/lib/llm/providers/openai/responses.rb +6 -8
- data/lib/llm/providers/openai/stream_parser.rb +1 -0
- data/lib/llm/providers/openai.rb +3 -10
- data/lib/llm/repl/bar.rb +4 -3
- data/lib/llm/repl/buffer.rb +42 -15
- data/lib/llm/repl/color.rb +78 -0
- data/lib/llm/repl/input/char.rb +46 -0
- data/lib/llm/repl/input/row.rb +39 -0
- data/lib/llm/repl/input.rb +251 -66
- data/lib/llm/repl/markdown/table.rb +6 -2
- data/lib/llm/repl/markdown.rb +31 -5
- data/lib/llm/repl/status.rb +38 -3
- data/lib/llm/repl/stream.rb +16 -4
- data/lib/llm/repl/walker.rb +3 -2
- data/lib/llm/repl/window.rb +25 -5
- data/lib/llm/repl.rb +29 -13
- data/lib/llm/stream.rb +8 -7
- data/lib/llm/tool.rb +29 -0
- data/lib/llm/transformer/null.rb +21 -0
- data/lib/llm/transformer.rb +55 -0
- data/lib/llm/version.rb +1 -1
- data/lib/llm.rb +12 -2
- data/llm.gemspec +1 -0
- data/resources/deepdive/advanced/cancellation.md +74 -0
- data/resources/deepdive/advanced/compaction.md +83 -0
- data/resources/deepdive/advanced/context.md +267 -0
- data/resources/deepdive/advanced/guard.md +371 -0
- data/resources/deepdive/advanced/tracer.md +180 -0
- data/resources/deepdive/advanced/transformer.md +67 -0
- data/resources/deepdive/advanced/transports.md +45 -0
- data/resources/deepdive/everything_else/audio.md +122 -0
- data/resources/deepdive/everything_else/cost.md +99 -0
- data/resources/deepdive/everything_else/images.md +89 -0
- data/resources/deepdive/everything_else/object.md +108 -0
- data/resources/deepdive/everything_else/ocr.md +48 -0
- data/resources/deepdive/fundamentals/agents.md +202 -0
- data/resources/deepdive/fundamentals/builtin_tools.md +191 -0
- data/resources/deepdive/fundamentals/concurrency.md +104 -0
- data/resources/deepdive/fundamentals/database.md +449 -0
- data/resources/deepdive/fundamentals/embeddings.md +157 -0
- data/resources/deepdive/fundamentals/repl.md +87 -0
- data/resources/deepdive/fundamentals/schema.md +61 -0
- data/resources/deepdive/fundamentals/skills.md +106 -0
- data/resources/deepdive/fundamentals/stream.md +110 -0
- data/resources/deepdive/fundamentals/tools.md +265 -0
- data/resources/deepdive/protocols/a2a.md +106 -0
- data/resources/deepdive/protocols/mcp.md +111 -0
- data/resources/deepdive.md +7 -1
- metadata +36 -3
- data/lib/llm/loop_guard.rb +0 -107
|
@@ -0,0 +1,371 @@
|
|
|
1
|
+
|
|
2
|
+
## Guard
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
|
|
9
|
+
is the superclass for context-level supervisors. A guard is bound
|
|
10
|
+
to a context and inspects each pending tool call before it runs.
|
|
11
|
+
It can let the call through, cancel it, block it with an error, or
|
|
12
|
+
answer for it with a synthesized result. Beyond loop detection,
|
|
13
|
+
guards handle policy, validation, quotas, cost control, caching,
|
|
14
|
+
and approval workflows.
|
|
15
|
+
|
|
16
|
+
#### How it works
|
|
17
|
+
|
|
18
|
+
A guard is a subclass of
|
|
19
|
+
[`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
|
|
20
|
+
that implements
|
|
21
|
+
[`LLM::Guard#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html#call-instance_method).
|
|
22
|
+
The guard is stamped onto the functions the context binds, and when a
|
|
23
|
+
function's task runs the guard is checked on the calling thread before
|
|
24
|
+
the tool is handed to the strategy. `call` receives the pending
|
|
25
|
+
`function:` and returns an
|
|
26
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
27
|
+
to close that single call, or `nil` to let it run. The guard inspects
|
|
28
|
+
the current conversation through the `messages` helper and reads the
|
|
29
|
+
pending call through the function's own accessors.
|
|
30
|
+
|
|
31
|
+
Configure a guard by passing a class through the `guard:` option;
|
|
32
|
+
options given in `guard_options:` are forwarded to `call` as keyword
|
|
33
|
+
arguments:
|
|
34
|
+
|
|
35
|
+
```ruby
|
|
36
|
+
class RateLimitGuard < LLM::Guard
|
|
37
|
+
def call(function:, limit: 10)
|
|
38
|
+
if messages.count(&:tool_return?) >= limit
|
|
39
|
+
function.return(error: true, type: "guard_error",
|
|
40
|
+
message: "too many tool calls")
|
|
41
|
+
end
|
|
42
|
+
end
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
agent = LLM::Agent.new(
|
|
46
|
+
llm,
|
|
47
|
+
guard: RateLimitGuard,
|
|
48
|
+
guard_options: {limit: 3}
|
|
49
|
+
)
|
|
50
|
+
agent.talk "Research the market", tools: [FetchNews, FetchStocks]
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
#### Why would I use it?
|
|
54
|
+
|
|
55
|
+
A guard is the one hook that sees every pending tool call and the
|
|
56
|
+
full conversation at once. It can observe what the model is about
|
|
57
|
+
to do, cancel a call that needs approval, block one that violates
|
|
58
|
+
policy, answer a cheap question without running a tool, or stop
|
|
59
|
+
work when a budget is spent. Because the guard runs before the
|
|
60
|
+
tool, anything it intercepts never executes.
|
|
61
|
+
|
|
62
|
+
#### Notes
|
|
63
|
+
|
|
64
|
+
[`LLM::Guard::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Null.html)
|
|
65
|
+
is the default guard for
|
|
66
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html);
|
|
67
|
+
it never blocks tool work.
|
|
68
|
+
[`LLM::Context#guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#guard)
|
|
69
|
+
returns the configured guard class. Because `call` returns an
|
|
70
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
71
|
+
rather than a warning string, a guard can block an individual tool
|
|
72
|
+
call while the rest of the batch still executes. The guard is stamped
|
|
73
|
+
onto every function the context binds, so it runs whenever a task is
|
|
74
|
+
spawned — including tool calls a stream queues itself during a
|
|
75
|
+
streaming turn. A blocked call yields its return without executing.
|
|
76
|
+
|
|
77
|
+
The runtime binds a guard instance to the context and stamps it onto
|
|
78
|
+
the functions it resolves, so a guard cannot carry state in instance
|
|
79
|
+
variables between calls. Anything a guard needs to remember, like
|
|
80
|
+
how many calls already ran, must come from the conversation
|
|
81
|
+
(`messages`) or from class-level state.
|
|
82
|
+
|
|
83
|
+
### Inspect
|
|
84
|
+
|
|
85
|
+
#### Overview
|
|
86
|
+
|
|
87
|
+
The guard receives the pending
|
|
88
|
+
[`LLM::Function`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
89
|
+
and can read what the model is about to do before anything runs.
|
|
90
|
+
The function exposes its `name`, `arguments`, and parameter schema,
|
|
91
|
+
so a guard can base its decision on the actual call rather than on
|
|
92
|
+
the whole conversation. The guard also holds the context, so the
|
|
93
|
+
message history, token usage, and cost are visible too.
|
|
94
|
+
|
|
95
|
+
#### How it works
|
|
96
|
+
|
|
97
|
+
When you want to inspect a pending call before it runs, read the
|
|
98
|
+
function's accessors inside `call`.
|
|
99
|
+
[`LLM::Function#name`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#name)
|
|
100
|
+
is the tool name, and
|
|
101
|
+
[`LLM::Function#arguments`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#arguments)
|
|
102
|
+
is an
|
|
103
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
104
|
+
with the parsed arguments. The full parameter schema is available
|
|
105
|
+
through
|
|
106
|
+
[`LLM::Function#params`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#params).
|
|
107
|
+
Returning `nil` lets the call run, so an inspection-only guard is a
|
|
108
|
+
pure observer:
|
|
109
|
+
|
|
110
|
+
```ruby
|
|
111
|
+
class AuditGuard < LLM::Guard
|
|
112
|
+
def call(function:)
|
|
113
|
+
warn "pending: #{function.name}(#{function.arguments.inspect})"
|
|
114
|
+
nil # let it run
|
|
115
|
+
end
|
|
116
|
+
end
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
#### Why would I use it?
|
|
120
|
+
|
|
121
|
+
Inspecting the pending call is the basis for policy, validation,
|
|
122
|
+
and audit. A guard can log every call into a telemetry stream, or
|
|
123
|
+
reject a single call because its arguments violate a rule while
|
|
124
|
+
letting every other call through. The same read underpins the
|
|
125
|
+
richer decisions in the rest of this document.
|
|
126
|
+
|
|
127
|
+
#### Notes
|
|
128
|
+
|
|
129
|
+
The guard sees the parsed arguments exactly as the model requested
|
|
130
|
+
them. Reading them costs nothing and never executes the tool.
|
|
131
|
+
Returning `nil` means the call proceeds normally.
|
|
132
|
+
|
|
133
|
+
### Cancel
|
|
134
|
+
|
|
135
|
+
#### Overview
|
|
136
|
+
|
|
137
|
+
Cancelling is distinct from blocking with an error.
|
|
138
|
+
[`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
|
|
139
|
+
produces a return the model reads as a declined call, not a failed
|
|
140
|
+
one. The conversation stays honest: the model asked for something,
|
|
141
|
+
and the runtime declined it with a reason.
|
|
142
|
+
|
|
143
|
+
#### How it works
|
|
144
|
+
|
|
145
|
+
When you want to cancel a single pending call, call the
|
|
146
|
+
[`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
|
|
147
|
+
method with a `reason:` and return the result from `call`. The
|
|
148
|
+
pending call never executes, and the rest of the batch still runs.
|
|
149
|
+
A guard can cancel one call in a two-call batch and the other call
|
|
150
|
+
still executes:
|
|
151
|
+
|
|
152
|
+
```ruby
|
|
153
|
+
class ApprovalGuard < LLM::Guard
|
|
154
|
+
def call(function:)
|
|
155
|
+
if function.name == "delete-file"
|
|
156
|
+
function.cancel(reason: "delete requires human approval")
|
|
157
|
+
end
|
|
158
|
+
end
|
|
159
|
+
end
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
#### Why would I use it?
|
|
163
|
+
|
|
164
|
+
Cancellation suits approval workflows and environment rules. A call
|
|
165
|
+
that needs a human in the loop, a tool that is disabled for a
|
|
166
|
+
particular user, or an operation that must not run in production
|
|
167
|
+
can all be declined without pretending the tool failed. Because
|
|
168
|
+
cancellation is per call, it declines only the offending tool and
|
|
169
|
+
leaves the rest of the batch intact.
|
|
170
|
+
|
|
171
|
+
#### Notes
|
|
172
|
+
|
|
173
|
+
The model receives
|
|
174
|
+
`{cancelled: true, reason: "delete requires human approval"}` as
|
|
175
|
+
that call's result and can react, for example by asking for
|
|
176
|
+
permission or skipping the operation. A cancelled call still counts
|
|
177
|
+
as resolved, like any tool return, so the model keeps moving. The
|
|
178
|
+
reason string is what the model sees, so write it as guidance
|
|
179
|
+
("ask the user first") rather than a raw error dump.
|
|
180
|
+
|
|
181
|
+
### Block
|
|
182
|
+
|
|
183
|
+
#### Overview
|
|
184
|
+
|
|
185
|
+
Blocking is how a guard enforces policy. A blocked call never runs,
|
|
186
|
+
and the model receives an in-band error explaining why. Unlike a
|
|
187
|
+
cancellation, an error tells the model the tool could not produce
|
|
188
|
+
a result at all.
|
|
189
|
+
|
|
190
|
+
#### How it works
|
|
191
|
+
|
|
192
|
+
When you want to block a call, return an
|
|
193
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
194
|
+
with `error: true`. The `type` and `message` are free-form, and the
|
|
195
|
+
model sees them inside the return value. Returning `nil` lets the
|
|
196
|
+
call through:
|
|
197
|
+
|
|
198
|
+
```ruby
|
|
199
|
+
class PolicyGuard < LLM::Guard
|
|
200
|
+
def call(function:)
|
|
201
|
+
if function.name == "shell"
|
|
202
|
+
function.return(error: true, type: "policy_error",
|
|
203
|
+
message: "shell is disabled")
|
|
204
|
+
end
|
|
205
|
+
end
|
|
206
|
+
end
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+
#### Why would I use it?
|
|
210
|
+
|
|
211
|
+
Blocking denies a dangerous tool, rejects an out-of-policy
|
|
212
|
+
argument, or stops a quota violation before it happens. The model
|
|
213
|
+
receives the return and can adapt, so the conversation stays valid.
|
|
214
|
+
|
|
215
|
+
#### Notes
|
|
216
|
+
|
|
217
|
+
A blocked call never executes. The return it produces is sent back
|
|
218
|
+
through the model like any tool result, so the model sees why the
|
|
219
|
+
call was blocked and can change course.
|
|
220
|
+
|
|
221
|
+
### Answer
|
|
222
|
+
|
|
223
|
+
#### Overview
|
|
224
|
+
|
|
225
|
+
A guard's return is injected into the conversation as if the tool
|
|
226
|
+
had executed, so a guard can answer for a tool that never runs.
|
|
227
|
+
From the model's side, a synthesized result is indistinguishable
|
|
228
|
+
from a real one.
|
|
229
|
+
|
|
230
|
+
#### How it works
|
|
231
|
+
|
|
232
|
+
When you want to answer a pending call yourself, return a value
|
|
233
|
+
through the
|
|
234
|
+
[`LLM::Function#return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
235
|
+
helper. The value you pass becomes the tool's result as-is, so it
|
|
236
|
+
must look like a plausible answer:
|
|
237
|
+
|
|
238
|
+
```ruby
|
|
239
|
+
class CacheGuard < LLM::Guard
|
|
240
|
+
CACHE = {"get-weather:tokyo" => {forecast: "sunny"}}
|
|
241
|
+
|
|
242
|
+
def call(function:)
|
|
243
|
+
key = "#{function.name}:#{function.arguments[:city]}"
|
|
244
|
+
function.return(CACHE[key]) if CACHE.key?(key)
|
|
245
|
+
end
|
|
246
|
+
end
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
#### Why would I use it?
|
|
250
|
+
|
|
251
|
+
Standing in for a tool is how you cache expensive calls, mock tools
|
|
252
|
+
in tests, or degrade gracefully when a service is down. The same
|
|
253
|
+
hook can return a fixed answer for tools that should not run in a
|
|
254
|
+
given environment.
|
|
255
|
+
|
|
256
|
+
#### Notes
|
|
257
|
+
|
|
258
|
+
[`LLM::Function#unavailable`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
259
|
+
marks a tool as not found, and
|
|
260
|
+
[`LLM::Function#budget_spent`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
261
|
+
reports that the tool budget is exhausted. Both are returns too,
|
|
262
|
+
so they flow back to the model like any tool result. A guard that
|
|
263
|
+
synthesizes a result runs instead of the tool, so side effects the
|
|
264
|
+
tool would have performed, like a database write, are skipped.
|
|
265
|
+
|
|
266
|
+
### Budget
|
|
267
|
+
|
|
268
|
+
#### Overview
|
|
269
|
+
|
|
270
|
+
A guard can read the accumulated usage and cost of the conversation
|
|
271
|
+
through the context. That makes it the natural place to enforce a
|
|
272
|
+
hard ceiling: stop calling tools once a cost or token budget is
|
|
273
|
+
spent, or once the context window is nearly full.
|
|
274
|
+
|
|
275
|
+
#### How it works
|
|
276
|
+
|
|
277
|
+
When you want to enforce a budget, compare
|
|
278
|
+
[`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage)
|
|
279
|
+
or
|
|
280
|
+
[`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
|
|
281
|
+
against a limit inside `call`. Pass the limit through
|
|
282
|
+
`guard_options:` so it can be tuned per context. Returning a
|
|
283
|
+
cancellation or an error closes the call before it runs:
|
|
284
|
+
|
|
285
|
+
```ruby
|
|
286
|
+
class BudgetGuard < LLM::Guard
|
|
287
|
+
def call(function:, limit: 0.05)
|
|
288
|
+
if ctx.cost.total >= limit
|
|
289
|
+
function.cancel(reason: "cost ceiling reached, ask before continuing")
|
|
290
|
+
end
|
|
291
|
+
end
|
|
292
|
+
end
|
|
293
|
+
|
|
294
|
+
ctx = LLM::Context.new(
|
|
295
|
+
llm,
|
|
296
|
+
guard: BudgetGuard,
|
|
297
|
+
guard_options: {limit: 0.01}
|
|
298
|
+
)
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
#### Why would I use it?
|
|
302
|
+
|
|
303
|
+
Budgets matter for long autonomous runs where the model decides how
|
|
304
|
+
many tools to call. A cost ceiling keeps a runaway agent from
|
|
305
|
+
spending money, and a quota derived from `messages` (as in the
|
|
306
|
+
rate-limit example above) keeps a single user from exhausting
|
|
307
|
+
shared resources. The same comparison works with
|
|
308
|
+
[`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage)
|
|
309
|
+
against
|
|
310
|
+
[`LLM::Context#context_window`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_window)
|
|
311
|
+
to keep the conversation inside the model's window.
|
|
312
|
+
|
|
313
|
+
#### Notes
|
|
314
|
+
|
|
315
|
+
The cost reported by
|
|
316
|
+
[`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
|
|
317
|
+
reflects the conversation so far, so a limit is enforced per call:
|
|
318
|
+
once the ceiling is crossed, the next pending call is declined. The
|
|
319
|
+
guard is not a replacement for
|
|
320
|
+
[`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method),
|
|
321
|
+
which caps the number of tool calls in a single turn. The two
|
|
322
|
+
compose: the budget caps call count, and a guard enforces cost.
|
|
323
|
+
|
|
324
|
+
### Loop
|
|
325
|
+
|
|
326
|
+
#### Overview
|
|
327
|
+
|
|
328
|
+
[`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
|
|
329
|
+
is the built-in loop-detection guard. It reduces each assistant
|
|
330
|
+
tool call to a `[tool name, arguments]` signature and checks whether
|
|
331
|
+
the tail of the sequence is repeating.
|
|
332
|
+
|
|
333
|
+
#### How it works
|
|
334
|
+
|
|
335
|
+
When you want to detect repeated tool-call patterns, enable
|
|
336
|
+
[`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
|
|
337
|
+
and tune the `threshold:` option, which is the number of repeated
|
|
338
|
+
patterns required before the guard intervenes (default `3`). When
|
|
339
|
+
the guard detects a repeat, it returns an in-band
|
|
340
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
341
|
+
with type `"guard_error"` and a message that tells the model it is
|
|
342
|
+
stuck and should change approach:
|
|
343
|
+
|
|
344
|
+
```ruby
|
|
345
|
+
ctx = LLM::Context.new(
|
|
346
|
+
llm,
|
|
347
|
+
guard: LLM::Guard::Loop,
|
|
348
|
+
guard_options: {threshold: 2}
|
|
349
|
+
)
|
|
350
|
+
ctx.talk "Research the market", tools: [FetchNews, FetchStocks]
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
#### Why would I use it?
|
|
354
|
+
|
|
355
|
+
Loop detection matters for long, autonomous agent runs. Without it,
|
|
356
|
+
a model that repeats a tool call with the same arguments can
|
|
357
|
+
bounce between calls forever. The guard turns that into a bounded
|
|
358
|
+
conversation: after the threshold, the model receives a message
|
|
359
|
+
telling it to stop and try a different strategy.
|
|
360
|
+
|
|
361
|
+
#### Notes
|
|
362
|
+
|
|
363
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
364
|
+
enables
|
|
365
|
+
[`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
|
|
366
|
+
by default, so agents get loop protection without configuration.
|
|
367
|
+
A custom guard can be passed through the `guard:` option to replace
|
|
368
|
+
the loop guard entirely. Guards and the agent's tool budget
|
|
369
|
+
complement each other: a guard blocks work that looks stuck, while
|
|
370
|
+
[`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
|
|
371
|
+
caps the total number of tool calls in a single turn.
|
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
|
|
2
|
+
## Tracer
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
Tracers let you observe what the runtime is doing. They hook into
|
|
9
|
+
requests, tool calls, compactions, and other events. Debug a
|
|
10
|
+
misbehaving agent, monitor request latency, or export spans to
|
|
11
|
+
an observability backend. A provider-wide tracer intercepts every
|
|
12
|
+
request through that provider. An agent-local tracer only covers
|
|
13
|
+
requests made by that agent.
|
|
14
|
+
|
|
15
|
+
#### How it works
|
|
16
|
+
|
|
17
|
+
When you want to observe runtime events, subclass
|
|
18
|
+
[`LLM::Tracer`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html)
|
|
19
|
+
and implement the hooks you need, or use one of the built-in
|
|
20
|
+
tracers. The built-in tracers include
|
|
21
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
22
|
+
(writes structured JSON to stdout or a file),
|
|
23
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
24
|
+
(writes human-readable single-line logs to stderr), and
|
|
25
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
26
|
+
(exports spans via OTLP for OpenTelemetry). Attach a tracer to a
|
|
27
|
+
provider or an agent:
|
|
28
|
+
|
|
29
|
+
```ruby
|
|
30
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
31
|
+
llm.tracer = LLM::Tracer::PrettyLogger.new(llm)
|
|
32
|
+
agent = LLM::Agent.new(llm)
|
|
33
|
+
agent.talk "Hello"
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
#### Why would I use it?
|
|
37
|
+
|
|
38
|
+
Tracers give you visibility into what the runtime is doing. Debug
|
|
39
|
+
a misbehaving agent by tracing every request it makes. Monitor
|
|
40
|
+
request latency and token usage across providers. Export spans to
|
|
41
|
+
OpenTelemetry for integration with existing observability pipelines.
|
|
42
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
43
|
+
is the best choice during development for compact, human-readable
|
|
44
|
+
output.
|
|
45
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
46
|
+
provides structured JSON, and
|
|
47
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
48
|
+
exports spans to OpenTelemetry for production observability.
|
|
49
|
+
|
|
50
|
+
#### Notes
|
|
51
|
+
|
|
52
|
+
The tracer is extensible. You can implement custom hooks for any
|
|
53
|
+
runtime event. The scope can be an individual agent or every
|
|
54
|
+
request a provider makes. Three built-in tracers are available:
|
|
55
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
56
|
+
(human-readable),
|
|
57
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
58
|
+
(structured JSON), and
|
|
59
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
60
|
+
(OpenTelemetry).
|
|
61
|
+
|
|
62
|
+
### Provider
|
|
63
|
+
|
|
64
|
+
#### Overview
|
|
65
|
+
|
|
66
|
+
A provider-wide tracer intercepts every request made through that
|
|
67
|
+
provider. All agents sharing the same provider share the same
|
|
68
|
+
tracer. Use this to trace at the infrastructure level without
|
|
69
|
+
configuring each agent individually.
|
|
70
|
+
|
|
71
|
+
#### How it works
|
|
72
|
+
|
|
73
|
+
When you want every request through a provider to be traced, set
|
|
74
|
+
the tracer on the provider directly. Every request made through
|
|
75
|
+
that provider, regardless of which agent initiates it, flows
|
|
76
|
+
through the same tracer hooks. The provider holds a reference to the
|
|
77
|
+
tracer and passes it to every new context it creates. This ensures
|
|
78
|
+
consistent observability without configuring each agent individually.
|
|
79
|
+
|
|
80
|
+
```ruby
|
|
81
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
82
|
+
llm.tracer = LLM::Tracer::Logger.new(llm, io: $stdout)
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
#### Why would I use it?
|
|
86
|
+
|
|
87
|
+
A provider-wide tracer captures every request at the infrastructure level.
|
|
88
|
+
All agents sharing the same provider share the same tracer.
|
|
89
|
+
|
|
90
|
+
#### Notes
|
|
91
|
+
|
|
92
|
+
The tracer can also write to a file with the `path:` option to
|
|
93
|
+
[`LLM::Tracer::Logger.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html#initialize-instance_method).
|
|
94
|
+
|
|
95
|
+
### Agent
|
|
96
|
+
|
|
97
|
+
#### Overview
|
|
98
|
+
|
|
99
|
+
An agent-local tracer only covers requests made by that agent.
|
|
100
|
+
Attach it via the `tracer:` keyword argument to
|
|
101
|
+
[`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#initialize-instance_method)
|
|
102
|
+
and it follows that agent wherever it goes. Different agents can
|
|
103
|
+
have different tracers.
|
|
104
|
+
|
|
105
|
+
#### How it works
|
|
106
|
+
|
|
107
|
+
When you want a tracer for a specific agent, pass it to the agent
|
|
108
|
+
on creation. Only requests made by
|
|
109
|
+
that agent flow through the tracer, leaving other agents on the
|
|
110
|
+
same provider unaffected.
|
|
111
|
+
|
|
112
|
+
```ruby
|
|
113
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
114
|
+
agent = LLM::Agent.new(llm, tracer: LLM::Tracer::Logger.new(llm, io: $stdout))
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
#### Why would I use it?
|
|
118
|
+
|
|
119
|
+
Agent-local tracers let each agent log differently.
|
|
120
|
+
One agent might log to stdout, another to a file, a third to
|
|
121
|
+
OpenTelemetry.
|
|
122
|
+
|
|
123
|
+
#### Notes
|
|
124
|
+
|
|
125
|
+
The tracer can also write to a file with the `path:` option to
|
|
126
|
+
[`LLM::Tracer::Logger.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html#initialize-instance_method).
|
|
127
|
+
|
|
128
|
+
### PrettyLogger
|
|
129
|
+
|
|
130
|
+
#### Overview
|
|
131
|
+
|
|
132
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
133
|
+
writes human-readable single-line logs to stderr. Unlike
|
|
134
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
135
|
+
(which emits structured JSON), the pretty logger is designed for
|
|
136
|
+
interactive development sessions where you want to see request and
|
|
137
|
+
tool-call activity at a glance.
|
|
138
|
+
|
|
139
|
+
#### How it works
|
|
140
|
+
|
|
141
|
+
Each request and tool call produces a single line on stderr with
|
|
142
|
+
the model, duration, and a summary of the activity. The logger
|
|
143
|
+
accepts an `io:` option to redirect output.
|
|
144
|
+
|
|
145
|
+
##### Provider-wide
|
|
146
|
+
|
|
147
|
+
```ruby
|
|
148
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
149
|
+
llm.tracer = LLM::Tracer::PrettyLogger.new(llm)
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
##### Agent-local
|
|
153
|
+
|
|
154
|
+
```ruby
|
|
155
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
156
|
+
agent = LLM::Agent.new(llm, tracer: LLM::Tracer::PrettyLogger.new(llm))
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
##### Custom output
|
|
160
|
+
|
|
161
|
+
```ruby
|
|
162
|
+
tracer = LLM::Tracer::PrettyLogger.new(llm, io: $stdout)
|
|
163
|
+
tracer = LLM::Tracer::PrettyLogger.new(llm, io: File.open("trace.log", "a"))
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
#### Why would I use it?
|
|
167
|
+
|
|
168
|
+
The pretty logger is the best choice for development. The output is
|
|
169
|
+
compact enough to follow in real time while still showing the model
|
|
170
|
+
name, duration, and tool calls. Switch to
|
|
171
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
172
|
+
when you need structured JSON for programmatic analysis, or to
|
|
173
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
174
|
+
when you need OpenTelemetry exports.
|
|
175
|
+
|
|
176
|
+
#### Notes
|
|
177
|
+
|
|
178
|
+
The pretty logger writes to `$stderr` by default. Set `io:` to
|
|
179
|
+
redirect output. All three built-in tracers share the same interface,
|
|
180
|
+
so switching between them requires changing only the class name.
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
|
|
2
|
+
## Transformer
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
|
|
9
|
+
is the superclass for message transformers. A transformer is bound
|
|
10
|
+
to a context and rewrites a single message before it is sent to the
|
|
11
|
+
provider. This lets you scrub sensitive data, inject context, or
|
|
12
|
+
otherwise modify outgoing messages without changing your prompt
|
|
13
|
+
code.
|
|
14
|
+
|
|
15
|
+
#### How it works
|
|
16
|
+
|
|
17
|
+
A transformer is a subclass of
|
|
18
|
+
[`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
|
|
19
|
+
that implements
|
|
20
|
+
[`LLM::Transformer#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html#call-instance_method).
|
|
21
|
+
The method receives the message to transform and returns a
|
|
22
|
+
message.
|
|
23
|
+
You can mutate the message in place or return a new one; either way,
|
|
24
|
+
the returned message is what gets sent.
|
|
25
|
+
|
|
26
|
+
Configure the transformer on a context with the `transformer:` option,
|
|
27
|
+
passing a class rather than an instance. The runtime instantiates it
|
|
28
|
+
once per turn. Options passed through `transformer_options:` are
|
|
29
|
+
forwarded to `call` as keyword arguments. The transformer runs on
|
|
30
|
+
the most recent message in both chat completions and Responses API
|
|
31
|
+
turns, before the request reaches the provider:
|
|
32
|
+
|
|
33
|
+
```ruby
|
|
34
|
+
class RedactEmails < LLM::Transformer
|
|
35
|
+
def call(message:)
|
|
36
|
+
content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
|
|
37
|
+
LLM::Message.new(message.role, content, message.extra)
|
|
38
|
+
end
|
|
39
|
+
end
|
|
40
|
+
|
|
41
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
42
|
+
ctx = LLM::Context.new(
|
|
43
|
+
llm,
|
|
44
|
+
transformer: RedactEmails
|
|
45
|
+
)
|
|
46
|
+
ctx.talk "Contact support@example.com for help"
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
#### Why would I use it?
|
|
50
|
+
|
|
51
|
+
Transformers give you a single hook point for all outgoing messages.
|
|
52
|
+
Common uses include redacting PII before it leaves your process,
|
|
53
|
+
injecting a timestamp or request ID, or normalizing content for a
|
|
54
|
+
particular provider. Because the transformer runs automatically on
|
|
55
|
+
every turn, you never need to remember to apply the transform in
|
|
56
|
+
your prompt code.
|
|
57
|
+
|
|
58
|
+
#### Notes
|
|
59
|
+
|
|
60
|
+
[`LLM::Transformer::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer/Null.html)
|
|
61
|
+
is the default transformer; it returns the message unchanged. The
|
|
62
|
+
`transformer_options:` hash is passed to `call` on every turn.
|
|
63
|
+
Streams can observe transformation through the
|
|
64
|
+
[`LLM::Stream#on_transform`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform-instance_method)
|
|
65
|
+
and
|
|
66
|
+
[`LLM::Stream#on_transform_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform_finish-instance_method)
|
|
67
|
+
callbacks, which receive the transformer instance.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
|
|
2
|
+
## Transports
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
The transport option selects which HTTP library is used for network
|
|
9
|
+
communication.
|
|
10
|
+
[`LLM::Provider`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html),
|
|
11
|
+
[`LLM::MCP`](https://r.uby.dev/api-docs/llm.rb/LLM/MCP.html), and
|
|
12
|
+
[`LLM::A2A`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html)
|
|
13
|
+
all accept this option. Three backends are available out of the box:
|
|
14
|
+
`net/http` is always there, `net/http/persistent` pools connections,
|
|
15
|
+
and `curb` wraps libcurl. Each implements the same internal interface
|
|
16
|
+
so switching between them is a one-word change.
|
|
17
|
+
|
|
18
|
+
#### How it works
|
|
19
|
+
|
|
20
|
+
**Net/HTTP** (`:net_http`) is the
|
|
21
|
+
default and is always available. **Net/HTTP/Persistent**
|
|
22
|
+
(`:net_http_persistent`) maintains a connection pool so the cost
|
|
23
|
+
of tearing down and setting up connections is kept low. **Curb**
|
|
24
|
+
(`:curb`) provides bindings for libcurl.
|
|
25
|
+
|
|
26
|
+
```ruby
|
|
27
|
+
llm = LLM.deepseek(key: "...", transport: :net_http)
|
|
28
|
+
mcp = LLM::MCP.http(url: "...", transport: :net_http_persistent)
|
|
29
|
+
a2a = LLM::A2A.rest(url: "...", transport: :curb)
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
#### Why would I use it?
|
|
33
|
+
|
|
34
|
+
Different environments have different HTTP needs.
|
|
35
|
+
|
|
36
|
+
The default `net/http` transport is always available and works
|
|
37
|
+
everywhere. The persistent transport reduces connection overhead
|
|
38
|
+
when your agent makes many requests to the same provider in quick
|
|
39
|
+
succession. Curb gives you libcurl bindings for environments where
|
|
40
|
+
that is already configured or preferred.
|
|
41
|
+
|
|
42
|
+
#### Notes
|
|
43
|
+
|
|
44
|
+
The persistent transport is built on top of `net/http`. Curb
|
|
45
|
+
requires the `curb` gem.
|