llm.rb 13.0.0 → 14.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +505 -14
- data/README.md +484 -50
- data/bin/llm.rb +148 -0
- data/data/anthropic.json +206 -263
- data/data/bedrock.json +2138 -1860
- data/data/deepinfra.json +1003 -624
- data/data/deepseek.json +38 -34
- data/data/google.json +1079 -371
- data/data/mistral.json +448 -368
- data/data/moonshot.json +384 -0
- data/data/openai.json +974 -1343
- data/data/xai.json +154 -126
- data/data/zai.json +191 -191
- data/lib/llm/agent.rb +123 -20
- data/lib/llm/context.rb +71 -88
- data/lib/llm/cost.rb +23 -17
- data/lib/llm/error.rb +0 -8
- data/lib/llm/function/array.rb +3 -3
- data/lib/llm/function/async/task.rb +2 -0
- data/lib/llm/function/fiber/task.rb +2 -0
- data/lib/llm/function/fork/task.rb +2 -0
- data/lib/llm/function/ractor/task.rb +2 -0
- data/lib/llm/function/sequential/group.rb +4 -1
- data/lib/llm/function/sequential/task.rb +1 -1
- data/lib/llm/function/task.rb +4 -0
- data/lib/llm/function/thread/task.rb +2 -0
- data/lib/llm/function.rb +33 -6
- data/lib/llm/guard/loop.rb +89 -0
- data/lib/llm/guard/null.rb +19 -0
- data/lib/llm/guard.rb +61 -0
- data/lib/llm/provider.rb +36 -0
- data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
- data/lib/llm/providers/anthropic.rb +2 -9
- data/lib/llm/providers/bedrock/request_adapter.rb +1 -1
- data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
- data/lib/llm/providers/bedrock.rb +1 -8
- data/lib/llm/providers/google/stream_parser.rb +1 -0
- data/lib/llm/providers/google.rb +1 -8
- data/lib/llm/providers/mistral.rb +1 -1
- data/lib/llm/providers/moonshot.rb +76 -0
- data/lib/llm/providers/ollama.rb +2 -9
- data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
- data/lib/llm/providers/openai/responses.rb +7 -9
- data/lib/llm/providers/openai/stream_parser.rb +1 -0
- data/lib/llm/providers/openai.rb +4 -11
- data/lib/llm/repl/bar.rb +4 -3
- data/lib/llm/repl/{transcript.rb → buffer.rb} +69 -29
- data/lib/llm/repl/color.rb +78 -0
- data/lib/llm/repl/command.rb +12 -5
- data/lib/llm/repl/commands/compact.rb +2 -2
- data/lib/llm/repl/commands/help.rb +3 -5
- data/lib/llm/repl/input/char.rb +46 -0
- data/lib/llm/repl/input/row.rb +39 -0
- data/lib/llm/repl/input.rb +251 -66
- data/lib/llm/repl/markdown/table.rb +11 -3
- data/lib/llm/repl/markdown.rb +34 -8
- data/lib/llm/repl/node.rb +37 -0
- data/lib/llm/repl/status.rb +42 -7
- data/lib/llm/repl/stream.rb +18 -6
- data/lib/llm/repl/walker.rb +3 -2
- data/lib/llm/repl/window.rb +54 -35
- data/lib/llm/repl.rb +74 -32
- data/lib/llm/skill.rb +20 -4
- data/lib/llm/stream.rb +8 -7
- data/lib/llm/tool.rb +29 -0
- data/lib/llm/tools/{swap_text.rb → edit-file.rb} +3 -3
- data/lib/llm/tools/git.rb +3 -0
- data/lib/llm/tools/mkdir.rb +3 -0
- data/lib/llm/tools/rg.rb +3 -0
- data/lib/llm/tools/ruby.rb +46 -0
- data/lib/llm/tools/shell.rb +3 -0
- data/lib/llm/tracer/pretty_logger.rb +127 -0
- data/lib/llm/tracer.rb +1 -0
- data/lib/llm/transformer/null.rb +21 -0
- data/lib/llm/transformer.rb +55 -0
- data/lib/llm/version.rb +1 -1
- data/lib/llm.rb +12 -2
- data/llm.gemspec +9 -2
- data/resources/deepdive/advanced/cancellation.md +74 -0
- data/resources/deepdive/advanced/compaction.md +83 -0
- data/resources/deepdive/advanced/context.md +267 -0
- data/resources/deepdive/advanced/guard.md +371 -0
- data/resources/deepdive/advanced/tracer.md +180 -0
- data/resources/deepdive/advanced/transformer.md +67 -0
- data/resources/deepdive/advanced/transports.md +45 -0
- data/resources/deepdive/everything_else/audio.md +122 -0
- data/resources/deepdive/everything_else/cost.md +99 -0
- data/resources/deepdive/everything_else/images.md +89 -0
- data/resources/deepdive/everything_else/object.md +108 -0
- data/resources/deepdive/everything_else/ocr.md +48 -0
- data/resources/deepdive/fundamentals/agents.md +202 -0
- data/resources/deepdive/fundamentals/builtin_tools.md +191 -0
- data/resources/deepdive/fundamentals/concurrency.md +104 -0
- data/resources/deepdive/fundamentals/database.md +449 -0
- data/resources/deepdive/fundamentals/embeddings.md +157 -0
- data/resources/deepdive/fundamentals/repl.md +87 -0
- data/resources/deepdive/fundamentals/schema.md +61 -0
- data/resources/deepdive/fundamentals/skills.md +106 -0
- data/resources/deepdive/fundamentals/stream.md +110 -0
- data/resources/deepdive/fundamentals/tools.md +265 -0
- data/resources/deepdive/protocols/a2a.md +106 -0
- data/resources/deepdive/protocols/mcp.md +111 -0
- data/resources/deepdive.md +58 -1792
- metadata +51 -7
- data/lib/llm/loop_guard.rb +0 -107
|
@@ -0,0 +1,267 @@
|
|
|
1
|
+
|
|
2
|
+
## Context
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
9
|
+
is the runtime that powers every agent. When you call
|
|
10
|
+
[`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk),
|
|
11
|
+
the agent delegates to its internal context. The context manages
|
|
12
|
+
the message history, sends requests to the provider, tracks pending
|
|
13
|
+
tool calls, and feeds results back to the model. Everything an agent
|
|
14
|
+
does, a context does too, but without the automatic tool loop.
|
|
15
|
+
|
|
16
|
+
Using a context directly gives you finer control over each step
|
|
17
|
+
of the conversation. You decide when to send messages, when to
|
|
18
|
+
execute tools, and when to stop. This is useful for custom
|
|
19
|
+
confirmation flows, mixed concurrency strategies per tool, or
|
|
20
|
+
any workflow where the agent's automatic loop gets in the way.
|
|
21
|
+
|
|
22
|
+
#### How it works
|
|
23
|
+
|
|
24
|
+
A context wraps a provider and maintains the conversation state.
|
|
25
|
+
Call
|
|
26
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
|
|
27
|
+
to send input to the model, check
|
|
28
|
+
[`LLM::Context#pending_functions?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions?)
|
|
29
|
+
to see if tools were requested, and use
|
|
30
|
+
[`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
|
|
31
|
+
to execute them. Each call to
|
|
32
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
|
|
33
|
+
appends
|
|
34
|
+
to the conversation and returns the model's response. The context
|
|
35
|
+
serializes its state with
|
|
36
|
+
[`LLM::Context#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#to_h)
|
|
37
|
+
and
|
|
38
|
+
[`LLM::Context#to_json`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#to_json),
|
|
39
|
+
and restores it
|
|
40
|
+
with
|
|
41
|
+
[`LLM::Context#restore`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#restore).
|
|
42
|
+
This is how the ORM integrations and filesystem
|
|
43
|
+
persistence work under the hood:
|
|
44
|
+
|
|
45
|
+
```ruby
|
|
46
|
+
require "llm"
|
|
47
|
+
|
|
48
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
49
|
+
ctx = LLM::Context.new(llm)
|
|
50
|
+
|
|
51
|
+
res = ctx.talk "What's the weather in Tokyo?"
|
|
52
|
+
puts res.content
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
#### Why would I use it?
|
|
56
|
+
|
|
57
|
+
A bare context gives you control that the agent
|
|
58
|
+
abstraction does not expose. Pre-flight checks on tool requests,
|
|
59
|
+
per-tool confirmation prompts, mixed concurrency strategies across
|
|
60
|
+
tools, or manual iteration until a condition is met are all easier
|
|
61
|
+
with a bare context.
|
|
62
|
+
|
|
63
|
+
#### Notes
|
|
64
|
+
|
|
65
|
+
The agent uses
|
|
66
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
67
|
+
internally. Anything you can do with
|
|
68
|
+
a context, you can also do through an agent. The trade-off is
|
|
69
|
+
convenience versus control. Contexts support the same concurrency
|
|
70
|
+
strategies, compaction, cancellation, and serialization as agents.
|
|
71
|
+
|
|
72
|
+
### Manual loop
|
|
73
|
+
|
|
74
|
+
#### Overview
|
|
75
|
+
|
|
76
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
77
|
+
manages the tool loop automatically. It calls
|
|
78
|
+
the model, checks for tool requests, runs the tools, feeds results
|
|
79
|
+
back, and repeats until the model produces text. You can bypass
|
|
80
|
+
this and drive
|
|
81
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
82
|
+
directly instead. This
|
|
83
|
+
gives you finer control over each step of the loop at the cost of
|
|
84
|
+
more code.
|
|
85
|
+
|
|
86
|
+
#### How it works
|
|
87
|
+
|
|
88
|
+
When you want to control the tool loop yourself, drive
|
|
89
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
90
|
+
directly instead of using an agent. Start a conversation, check
|
|
91
|
+
for tool requests, execute them, and feed results back. The full
|
|
92
|
+
loop is under your control. Each call to
|
|
93
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
|
|
94
|
+
appends to the conversation and returns the model's response, and
|
|
95
|
+
[`LLM::Context#pending_functions?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions?)
|
|
96
|
+
tells you whether tools were requested.
|
|
97
|
+
From that foundation you can inspect, iterate, or confirm
|
|
98
|
+
per-tool in a single flow:
|
|
99
|
+
|
|
100
|
+
```ruby
|
|
101
|
+
require "llm"
|
|
102
|
+
|
|
103
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
104
|
+
ctx = LLM::Context.new(llm)
|
|
105
|
+
|
|
106
|
+
loop do
|
|
107
|
+
res = ctx.talk("What's the weather in Tokyo?")
|
|
108
|
+
break unless ctx.pending_functions?
|
|
109
|
+
|
|
110
|
+
puts "Model requested #{ctx.pending_functions.size} tool(s)"
|
|
111
|
+
|
|
112
|
+
results = ctx.pending_functions.map do |fn|
|
|
113
|
+
print "Run #{fn.name} with #{fn.arguments}? [y/N] "
|
|
114
|
+
if $stdin.gets&.match?(/\Ay\z/i)
|
|
115
|
+
fn.task(:thread).wait
|
|
116
|
+
else
|
|
117
|
+
fn.cancel(reason: "user declined")
|
|
118
|
+
end
|
|
119
|
+
end
|
|
120
|
+
|
|
121
|
+
ctx.talk(results)
|
|
122
|
+
end
|
|
123
|
+
|
|
124
|
+
puts res.content
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
#### Why would I use it?
|
|
128
|
+
|
|
129
|
+
Manual control gives you pre-execution checks, custom confirmation
|
|
130
|
+
flows, different strategies per tool, and fine-grained error
|
|
131
|
+
recovery that the default tool loop does not expose.
|
|
132
|
+
|
|
133
|
+
#### Notes
|
|
134
|
+
|
|
135
|
+
[`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
|
|
136
|
+
picks up pending functions, spawns them using the chosen
|
|
137
|
+
strategy, waits for results, and records them back in the context.
|
|
138
|
+
Each strategy is supported: `:sequential`, `:thread`, `:fiber`,
|
|
139
|
+
`:async`, `:fork`, and `:ractor`. Functions are reset after each
|
|
140
|
+
[`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
|
|
141
|
+
or
|
|
142
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
|
|
143
|
+
call. Store the array if you need to
|
|
144
|
+
preserve them.
|
|
145
|
+
|
|
146
|
+
### Pending functions
|
|
147
|
+
|
|
148
|
+
#### Overview
|
|
149
|
+
|
|
150
|
+
Pending function calls represent the model's tool requests. After
|
|
151
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
|
|
152
|
+
returns, the context may have pending function
|
|
153
|
+
calls if the model requested tools. These are available through
|
|
154
|
+
[`LLM::Context#pending_functions`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions)
|
|
155
|
+
which returns an array of
|
|
156
|
+
[`LLM::Function`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
157
|
+
objects. Each function has a name, arguments, and methods for
|
|
158
|
+
execution or cancellation.
|
|
159
|
+
|
|
160
|
+
#### How it works
|
|
161
|
+
|
|
162
|
+
When you want to check whether the model requested tools, call
|
|
163
|
+
[`LLM::Context#pending_functions?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions?)
|
|
164
|
+
after each
|
|
165
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
|
|
166
|
+
call. Each pending
|
|
167
|
+
function has a name, arguments, and
|
|
168
|
+
methods for execution or cancellation. Call
|
|
169
|
+
[`LLM::Function#task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#task)
|
|
170
|
+
to execute it or
|
|
171
|
+
[`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
|
|
172
|
+
to skip it. Iterate over all
|
|
173
|
+
pending functions to inspect or handle them individually:
|
|
174
|
+
|
|
175
|
+
```ruby
|
|
176
|
+
res = ctx.talk "What's the weather in Tokyo?"
|
|
177
|
+
|
|
178
|
+
if ctx.pending_functions?
|
|
179
|
+
puts "Model requested #{ctx.pending_functions.size} tool(s)"
|
|
180
|
+
results = ctx.pending_functions.map do |fn|
|
|
181
|
+
print "Run #{fn.name} with #{fn.arguments}? [y/N] "
|
|
182
|
+
if $stdin.gets&.match?(/\Ay\z/i)
|
|
183
|
+
fn.task(:thread).wait
|
|
184
|
+
else
|
|
185
|
+
fn.cancel(reason: "user declined")
|
|
186
|
+
end
|
|
187
|
+
end
|
|
188
|
+
ctx.talk(results)
|
|
189
|
+
end
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
#### Why would I use it?
|
|
193
|
+
|
|
194
|
+
Inspecting pending functions lets you decide which tools to run,
|
|
195
|
+
in what order, and with what strategy. This is essential for
|
|
196
|
+
confirmation flows, selective execution, or logging which tools
|
|
197
|
+
the model requested.
|
|
198
|
+
|
|
199
|
+
#### Notes
|
|
200
|
+
|
|
201
|
+
Pending functions are reset after each
|
|
202
|
+
[`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
|
|
203
|
+
or
|
|
204
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
|
|
205
|
+
call. If you need to preserve them, store the array before
|
|
206
|
+
executing. Functions that are cancelled still count as completed
|
|
207
|
+
from the model's perspective; the model sees a cancellation
|
|
208
|
+
result, not a tool error.
|
|
209
|
+
|
|
210
|
+
### Tool responses
|
|
211
|
+
|
|
212
|
+
#### Overview
|
|
213
|
+
|
|
214
|
+
A tool interrupt gives you two choices. When a tool receives
|
|
215
|
+
[`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html),
|
|
216
|
+
it can either cancel the turn or return a result. The choice
|
|
217
|
+
depends on the situation.
|
|
218
|
+
A hard cancel aborts the request outright and is the default.
|
|
219
|
+
Returning a value lets the model adapt and continue the
|
|
220
|
+
conversation, which can be useful when the interrupt is
|
|
221
|
+
temporary, like a timeout or a user pause.
|
|
222
|
+
|
|
223
|
+
#### How it works
|
|
224
|
+
|
|
225
|
+
When a tool receives
|
|
226
|
+
[`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html),
|
|
227
|
+
re-raise to abort the turn or return a value to continue the loop.
|
|
228
|
+
The model receives the result and decides what to do next.
|
|
229
|
+
|
|
230
|
+
Re-raise to abort the turn entirely:
|
|
231
|
+
|
|
232
|
+
```ruby
|
|
233
|
+
class MyTool < LLM::Tool
|
|
234
|
+
def call
|
|
235
|
+
# do work
|
|
236
|
+
rescue LLM::Interrupt
|
|
237
|
+
cleanup
|
|
238
|
+
raise
|
|
239
|
+
end
|
|
240
|
+
end
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
Return a value to continue the loop:
|
|
244
|
+
|
|
245
|
+
```ruby
|
|
246
|
+
class MyTool < LLM::Tool
|
|
247
|
+
def call
|
|
248
|
+
# do work
|
|
249
|
+
rescue LLM::Interrupt
|
|
250
|
+
cleanup
|
|
251
|
+
{ok: false, reason: "interrupted"}
|
|
252
|
+
end
|
|
253
|
+
end
|
|
254
|
+
```
|
|
255
|
+
|
|
256
|
+
#### Why would I use it?
|
|
257
|
+
|
|
258
|
+
A hard cancel aborts the request outright. Useful when continuing
|
|
259
|
+
would produce garbage. Returning a value lets the model adapt,
|
|
260
|
+
which can be helpful when the interrupt is temporary.
|
|
261
|
+
|
|
262
|
+
#### Notes
|
|
263
|
+
|
|
264
|
+
The mechanism is the same across all six concurrency strategies.
|
|
265
|
+
The `:ractor` strategy delivers the interrupt through ractor
|
|
266
|
+
message passing. The `:fork` strategy delivers it via xchan.
|
|
267
|
+
|
|
@@ -0,0 +1,371 @@
|
|
|
1
|
+
|
|
2
|
+
## Guard
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
|
|
9
|
+
is the superclass for context-level supervisors. A guard is bound
|
|
10
|
+
to a context and inspects each pending tool call before it runs.
|
|
11
|
+
It can let the call through, cancel it, block it with an error, or
|
|
12
|
+
answer for it with a synthesized result. Beyond loop detection,
|
|
13
|
+
guards handle policy, validation, quotas, cost control, caching,
|
|
14
|
+
and approval workflows.
|
|
15
|
+
|
|
16
|
+
#### How it works
|
|
17
|
+
|
|
18
|
+
A guard is a subclass of
|
|
19
|
+
[`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
|
|
20
|
+
that implements
|
|
21
|
+
[`LLM::Guard#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html#call-instance_method).
|
|
22
|
+
The guard is stamped onto the functions the context binds, and when a
|
|
23
|
+
function's task runs the guard is checked on the calling thread before
|
|
24
|
+
the tool is handed to the strategy. `call` receives the pending
|
|
25
|
+
`function:` and returns an
|
|
26
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
27
|
+
to close that single call, or `nil` to let it run. The guard inspects
|
|
28
|
+
the current conversation through the `messages` helper and reads the
|
|
29
|
+
pending call through the function's own accessors.
|
|
30
|
+
|
|
31
|
+
Configure a guard by passing a class through the `guard:` option;
|
|
32
|
+
options given in `guard_options:` are forwarded to `call` as keyword
|
|
33
|
+
arguments:
|
|
34
|
+
|
|
35
|
+
```ruby
|
|
36
|
+
class RateLimitGuard < LLM::Guard
|
|
37
|
+
def call(function:, limit: 10)
|
|
38
|
+
if messages.count(&:tool_return?) >= limit
|
|
39
|
+
function.return(error: true, type: "guard_error",
|
|
40
|
+
message: "too many tool calls")
|
|
41
|
+
end
|
|
42
|
+
end
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
agent = LLM::Agent.new(
|
|
46
|
+
llm,
|
|
47
|
+
guard: RateLimitGuard,
|
|
48
|
+
guard_options: {limit: 3}
|
|
49
|
+
)
|
|
50
|
+
agent.talk "Research the market", tools: [FetchNews, FetchStocks]
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
#### Why would I use it?
|
|
54
|
+
|
|
55
|
+
A guard is the one hook that sees every pending tool call and the
|
|
56
|
+
full conversation at once. It can observe what the model is about
|
|
57
|
+
to do, cancel a call that needs approval, block one that violates
|
|
58
|
+
policy, answer a cheap question without running a tool, or stop
|
|
59
|
+
work when a budget is spent. Because the guard runs before the
|
|
60
|
+
tool, anything it intercepts never executes.
|
|
61
|
+
|
|
62
|
+
#### Notes
|
|
63
|
+
|
|
64
|
+
[`LLM::Guard::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Null.html)
|
|
65
|
+
is the default guard for
|
|
66
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html);
|
|
67
|
+
it never blocks tool work.
|
|
68
|
+
[`LLM::Context#guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#guard)
|
|
69
|
+
returns the configured guard class. Because `call` returns an
|
|
70
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
71
|
+
rather than a warning string, a guard can block an individual tool
|
|
72
|
+
call while the rest of the batch still executes. The guard is stamped
|
|
73
|
+
onto every function the context binds, so it runs whenever a task is
|
|
74
|
+
spawned — including tool calls a stream queues itself during a
|
|
75
|
+
streaming turn. A blocked call yields its return without executing.
|
|
76
|
+
|
|
77
|
+
The runtime binds a guard instance to the context and stamps it onto
|
|
78
|
+
the functions it resolves, so a guard cannot carry state in instance
|
|
79
|
+
variables between calls. Anything a guard needs to remember, like
|
|
80
|
+
how many calls already ran, must come from the conversation
|
|
81
|
+
(`messages`) or from class-level state.
|
|
82
|
+
|
|
83
|
+
### Inspect
|
|
84
|
+
|
|
85
|
+
#### Overview
|
|
86
|
+
|
|
87
|
+
The guard receives the pending
|
|
88
|
+
[`LLM::Function`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
89
|
+
and can read what the model is about to do before anything runs.
|
|
90
|
+
The function exposes its `name`, `arguments`, and parameter schema,
|
|
91
|
+
so a guard can base its decision on the actual call rather than on
|
|
92
|
+
the whole conversation. The guard also holds the context, so the
|
|
93
|
+
message history, token usage, and cost are visible too.
|
|
94
|
+
|
|
95
|
+
#### How it works
|
|
96
|
+
|
|
97
|
+
When you want to inspect a pending call before it runs, read the
|
|
98
|
+
function's accessors inside `call`.
|
|
99
|
+
[`LLM::Function#name`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#name)
|
|
100
|
+
is the tool name, and
|
|
101
|
+
[`LLM::Function#arguments`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#arguments)
|
|
102
|
+
is an
|
|
103
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
104
|
+
with the parsed arguments. The full parameter schema is available
|
|
105
|
+
through
|
|
106
|
+
[`LLM::Function#params`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#params).
|
|
107
|
+
Returning `nil` lets the call run, so an inspection-only guard is a
|
|
108
|
+
pure observer:
|
|
109
|
+
|
|
110
|
+
```ruby
|
|
111
|
+
class AuditGuard < LLM::Guard
|
|
112
|
+
def call(function:)
|
|
113
|
+
warn "pending: #{function.name}(#{function.arguments.inspect})"
|
|
114
|
+
nil # let it run
|
|
115
|
+
end
|
|
116
|
+
end
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
#### Why would I use it?
|
|
120
|
+
|
|
121
|
+
Inspecting the pending call is the basis for policy, validation,
|
|
122
|
+
and audit. A guard can log every call into a telemetry stream, or
|
|
123
|
+
reject a single call because its arguments violate a rule while
|
|
124
|
+
letting every other call through. The same read underpins the
|
|
125
|
+
richer decisions in the rest of this document.
|
|
126
|
+
|
|
127
|
+
#### Notes
|
|
128
|
+
|
|
129
|
+
The guard sees the parsed arguments exactly as the model requested
|
|
130
|
+
them. Reading them costs nothing and never executes the tool.
|
|
131
|
+
Returning `nil` means the call proceeds normally.
|
|
132
|
+
|
|
133
|
+
### Cancel
|
|
134
|
+
|
|
135
|
+
#### Overview
|
|
136
|
+
|
|
137
|
+
Cancelling is distinct from blocking with an error.
|
|
138
|
+
[`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
|
|
139
|
+
produces a return the model reads as a declined call, not a failed
|
|
140
|
+
one. The conversation stays honest: the model asked for something,
|
|
141
|
+
and the runtime declined it with a reason.
|
|
142
|
+
|
|
143
|
+
#### How it works
|
|
144
|
+
|
|
145
|
+
When you want to cancel a single pending call, call the
|
|
146
|
+
[`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
|
|
147
|
+
method with a `reason:` and return the result from `call`. The
|
|
148
|
+
pending call never executes, and the rest of the batch still runs.
|
|
149
|
+
A guard can cancel one call in a two-call batch and the other call
|
|
150
|
+
still executes:
|
|
151
|
+
|
|
152
|
+
```ruby
|
|
153
|
+
class ApprovalGuard < LLM::Guard
|
|
154
|
+
def call(function:)
|
|
155
|
+
if function.name == "delete-file"
|
|
156
|
+
function.cancel(reason: "delete requires human approval")
|
|
157
|
+
end
|
|
158
|
+
end
|
|
159
|
+
end
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
#### Why would I use it?
|
|
163
|
+
|
|
164
|
+
Cancellation suits approval workflows and environment rules. A call
|
|
165
|
+
that needs a human in the loop, a tool that is disabled for a
|
|
166
|
+
particular user, or an operation that must not run in production
|
|
167
|
+
can all be declined without pretending the tool failed. Because
|
|
168
|
+
cancellation is per call, it declines only the offending tool and
|
|
169
|
+
leaves the rest of the batch intact.
|
|
170
|
+
|
|
171
|
+
#### Notes
|
|
172
|
+
|
|
173
|
+
The model receives
|
|
174
|
+
`{cancelled: true, reason: "delete requires human approval"}` as
|
|
175
|
+
that call's result and can react, for example by asking for
|
|
176
|
+
permission or skipping the operation. A cancelled call still counts
|
|
177
|
+
as resolved, like any tool return, so the model keeps moving. The
|
|
178
|
+
reason string is what the model sees, so write it as guidance
|
|
179
|
+
("ask the user first") rather than a raw error dump.
|
|
180
|
+
|
|
181
|
+
### Block
|
|
182
|
+
|
|
183
|
+
#### Overview
|
|
184
|
+
|
|
185
|
+
Blocking is how a guard enforces policy. A blocked call never runs,
|
|
186
|
+
and the model receives an in-band error explaining why. Unlike a
|
|
187
|
+
cancellation, an error tells the model the tool could not produce
|
|
188
|
+
a result at all.
|
|
189
|
+
|
|
190
|
+
#### How it works
|
|
191
|
+
|
|
192
|
+
When you want to block a call, return an
|
|
193
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
194
|
+
with `error: true`. The `type` and `message` are free-form, and the
|
|
195
|
+
model sees them inside the return value. Returning `nil` lets the
|
|
196
|
+
call through:
|
|
197
|
+
|
|
198
|
+
```ruby
|
|
199
|
+
class PolicyGuard < LLM::Guard
|
|
200
|
+
def call(function:)
|
|
201
|
+
if function.name == "shell"
|
|
202
|
+
function.return(error: true, type: "policy_error",
|
|
203
|
+
message: "shell is disabled")
|
|
204
|
+
end
|
|
205
|
+
end
|
|
206
|
+
end
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+
#### Why would I use it?
|
|
210
|
+
|
|
211
|
+
Blocking denies a dangerous tool, rejects an out-of-policy
|
|
212
|
+
argument, or stops a quota violation before it happens. The model
|
|
213
|
+
receives the return and can adapt, so the conversation stays valid.
|
|
214
|
+
|
|
215
|
+
#### Notes
|
|
216
|
+
|
|
217
|
+
A blocked call never executes. The return it produces is sent back
|
|
218
|
+
through the model like any tool result, so the model sees why the
|
|
219
|
+
call was blocked and can change course.
|
|
220
|
+
|
|
221
|
+
### Answer
|
|
222
|
+
|
|
223
|
+
#### Overview
|
|
224
|
+
|
|
225
|
+
A guard's return is injected into the conversation as if the tool
|
|
226
|
+
had executed, so a guard can answer for a tool that never runs.
|
|
227
|
+
From the model's side, a synthesized result is indistinguishable
|
|
228
|
+
from a real one.
|
|
229
|
+
|
|
230
|
+
#### How it works
|
|
231
|
+
|
|
232
|
+
When you want to answer a pending call yourself, return a value
|
|
233
|
+
through the
|
|
234
|
+
[`LLM::Function#return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
235
|
+
helper. The value you pass becomes the tool's result as-is, so it
|
|
236
|
+
must look like a plausible answer:
|
|
237
|
+
|
|
238
|
+
```ruby
|
|
239
|
+
class CacheGuard < LLM::Guard
|
|
240
|
+
CACHE = {"get-weather:tokyo" => {forecast: "sunny"}}
|
|
241
|
+
|
|
242
|
+
def call(function:)
|
|
243
|
+
key = "#{function.name}:#{function.arguments[:city]}"
|
|
244
|
+
function.return(CACHE[key]) if CACHE.key?(key)
|
|
245
|
+
end
|
|
246
|
+
end
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
#### Why would I use it?
|
|
250
|
+
|
|
251
|
+
Standing in for a tool is how you cache expensive calls, mock tools
|
|
252
|
+
in tests, or degrade gracefully when a service is down. The same
|
|
253
|
+
hook can return a fixed answer for tools that should not run in a
|
|
254
|
+
given environment.
|
|
255
|
+
|
|
256
|
+
#### Notes
|
|
257
|
+
|
|
258
|
+
[`LLM::Function#unavailable`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
259
|
+
marks a tool as not found, and
|
|
260
|
+
[`LLM::Function#budget_spent`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
|
|
261
|
+
reports that the tool budget is exhausted. Both are returns too,
|
|
262
|
+
so they flow back to the model like any tool result. A guard that
|
|
263
|
+
synthesizes a result runs instead of the tool, so side effects the
|
|
264
|
+
tool would have performed, like a database write, are skipped.
|
|
265
|
+
|
|
266
|
+
### Budget
|
|
267
|
+
|
|
268
|
+
#### Overview
|
|
269
|
+
|
|
270
|
+
A guard can read the accumulated usage and cost of the conversation
|
|
271
|
+
through the context. That makes it the natural place to enforce a
|
|
272
|
+
hard ceiling: stop calling tools once a cost or token budget is
|
|
273
|
+
spent, or once the context window is nearly full.
|
|
274
|
+
|
|
275
|
+
#### How it works
|
|
276
|
+
|
|
277
|
+
When you want to enforce a budget, compare
|
|
278
|
+
[`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage)
|
|
279
|
+
or
|
|
280
|
+
[`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
|
|
281
|
+
against a limit inside `call`. Pass the limit through
|
|
282
|
+
`guard_options:` so it can be tuned per context. Returning a
|
|
283
|
+
cancellation or an error closes the call before it runs:
|
|
284
|
+
|
|
285
|
+
```ruby
|
|
286
|
+
class BudgetGuard < LLM::Guard
|
|
287
|
+
def call(function:, limit: 0.05)
|
|
288
|
+
if ctx.cost.total >= limit
|
|
289
|
+
function.cancel(reason: "cost ceiling reached, ask before continuing")
|
|
290
|
+
end
|
|
291
|
+
end
|
|
292
|
+
end
|
|
293
|
+
|
|
294
|
+
ctx = LLM::Context.new(
|
|
295
|
+
llm,
|
|
296
|
+
guard: BudgetGuard,
|
|
297
|
+
guard_options: {limit: 0.01}
|
|
298
|
+
)
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
#### Why would I use it?
|
|
302
|
+
|
|
303
|
+
Budgets matter for long autonomous runs where the model decides how
|
|
304
|
+
many tools to call. A cost ceiling keeps a runaway agent from
|
|
305
|
+
spending money, and a quota derived from `messages` (as in the
|
|
306
|
+
rate-limit example above) keeps a single user from exhausting
|
|
307
|
+
shared resources. The same comparison works with
|
|
308
|
+
[`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage)
|
|
309
|
+
against
|
|
310
|
+
[`LLM::Context#context_window`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_window)
|
|
311
|
+
to keep the conversation inside the model's window.
|
|
312
|
+
|
|
313
|
+
#### Notes
|
|
314
|
+
|
|
315
|
+
The cost reported by
|
|
316
|
+
[`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
|
|
317
|
+
reflects the conversation so far, so a limit is enforced per call:
|
|
318
|
+
once the ceiling is crossed, the next pending call is declined. The
|
|
319
|
+
guard is not a replacement for
|
|
320
|
+
[`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method),
|
|
321
|
+
which caps the number of tool calls in a single turn. The two
|
|
322
|
+
compose: the budget caps call count, and a guard enforces cost.
|
|
323
|
+
|
|
324
|
+
### Loop
|
|
325
|
+
|
|
326
|
+
#### Overview
|
|
327
|
+
|
|
328
|
+
[`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
|
|
329
|
+
is the built-in loop-detection guard. It reduces each assistant
|
|
330
|
+
tool call to a `[tool name, arguments]` signature and checks whether
|
|
331
|
+
the tail of the sequence is repeating.
|
|
332
|
+
|
|
333
|
+
#### How it works
|
|
334
|
+
|
|
335
|
+
When you want to detect repeated tool-call patterns, enable
|
|
336
|
+
[`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
|
|
337
|
+
and tune the `threshold:` option, which is the number of repeated
|
|
338
|
+
patterns required before the guard intervenes (default `3`). When
|
|
339
|
+
the guard detects a repeat, it returns an in-band
|
|
340
|
+
[`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
|
|
341
|
+
with type `"guard_error"` and a message that tells the model it is
|
|
342
|
+
stuck and should change approach:
|
|
343
|
+
|
|
344
|
+
```ruby
|
|
345
|
+
ctx = LLM::Context.new(
|
|
346
|
+
llm,
|
|
347
|
+
guard: LLM::Guard::Loop,
|
|
348
|
+
guard_options: {threshold: 2}
|
|
349
|
+
)
|
|
350
|
+
ctx.talk "Research the market", tools: [FetchNews, FetchStocks]
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
#### Why would I use it?
|
|
354
|
+
|
|
355
|
+
Loop detection matters for long, autonomous agent runs. Without it,
|
|
356
|
+
a model that repeats a tool call with the same arguments can
|
|
357
|
+
bounce between calls forever. The guard turns that into a bounded
|
|
358
|
+
conversation: after the threshold, the model receives a message
|
|
359
|
+
telling it to stop and try a different strategy.
|
|
360
|
+
|
|
361
|
+
#### Notes
|
|
362
|
+
|
|
363
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
364
|
+
enables
|
|
365
|
+
[`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
|
|
366
|
+
by default, so agents get loop protection without configuration.
|
|
367
|
+
A custom guard can be passed through the `guard:` option to replace
|
|
368
|
+
the loop guard entirely. Guards and the agent's tool budget
|
|
369
|
+
complement each other: a guard blocks work that looks stuck, while
|
|
370
|
+
[`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
|
|
371
|
+
caps the total number of tool calls in a single turn.
|