llm.rb 14.0.0 → 15.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +396 -1826
- data/README.md +596 -490
- data/bin/llm.rb +148 -72
- data/data/alibaba.json +1999 -0
- data/data/anthropic.json +205 -205
- data/data/bedrock.json +2171 -2114
- data/data/deepinfra.json +1143 -938
- data/data/deepseek.json +4 -5
- data/data/google.json +691 -779
- data/data/mistral.json +450 -450
- data/data/moonshot.json +100 -100
- data/data/openai.json +976 -976
- data/data/xai.json +193 -116
- data/data/zai.json +187 -187
- data/{resources → docs}/deepdive/advanced/cancellation.md +2 -2
- data/{resources → docs}/deepdive/advanced/compaction.md +5 -3
- data/{resources → docs}/deepdive/advanced/context.md +13 -0
- data/{resources → docs}/deepdive/advanced/guard.md +1 -1
- data/{resources/deepdive/fundamentals → docs/deepdive/features}/concurrency.md +6 -0
- data/{resources/deepdive/fundamentals → docs/deepdive/features}/repl.md +56 -2
- data/{resources → docs}/deepdive/fundamentals/agents.md +52 -1
- data/docs/deepdive/fundamentals/providers.md +159 -0
- data/{resources → docs}/deepdive/fundamentals/skills.md +5 -0
- data/{resources → docs}/deepdive/fundamentals/stream.md +36 -3
- data/{resources → docs}/deepdive/fundamentals/tools.md +87 -23
- data/{resources/deepdive/everything_else → docs/deepdive/reference}/cost.md +20 -10
- data/docs/deepdive/reference/model_registry.md +271 -0
- data/{resources/deepdive/advanced → docs/deepdive/reference}/tracer.md +7 -0
- data/{resources → docs}/deepdive.md +35 -27
- data/lib/llm/a2a/transport/http.rb +1 -1
- data/lib/llm/active_record/acts_as_llm.rb +19 -5
- data/lib/llm/agent.rb +62 -5
- data/lib/llm/context.rb +93 -46
- data/lib/llm/cost.rb +110 -51
- data/lib/llm/error.rb +7 -0
- data/lib/llm/function/array.rb +1 -1
- data/lib/llm/function/fork/task.rb +14 -1
- data/lib/llm/function/sequential/group.rb +20 -13
- data/lib/llm/function/sequential/task.rb +1 -8
- data/lib/llm/function.rb +5 -4
- data/lib/llm/message.rb +5 -4
- data/lib/llm/provider.rb +7 -0
- data/lib/llm/providers/alibaba/error_handler.rb +34 -0
- data/lib/llm/providers/alibaba/request_adapter.rb +13 -0
- data/lib/llm/providers/alibaba.rb +93 -0
- data/lib/llm/providers/anthropic.rb +0 -1
- data/lib/llm/providers/bedrock.rb +8 -1
- data/lib/llm/providers/deepseek/request_adapter.rb +2 -33
- data/lib/llm/providers/google.rb +0 -1
- data/lib/llm/providers/ollama.rb +0 -1
- data/lib/llm/providers/openai/responses.rb +0 -1
- data/lib/llm/providers/openai/schema.rb +37 -0
- data/lib/llm/providers/openai.rb +1 -2
- data/lib/llm/registry/model.rb +186 -0
- data/lib/llm/registry.rb +45 -14
- data/lib/llm/repl/bar.rb +11 -13
- data/lib/llm/repl/buffer.rb +1 -1
- data/lib/llm/repl/color.rb +8 -1
- data/lib/llm/repl/command.rb +12 -0
- data/lib/llm/repl/commands/model.rb +39 -0
- data/lib/llm/repl/input/cache.rb +45 -0
- data/lib/llm/repl/input/char.rb +2 -2
- data/lib/llm/repl/input.rb +86 -23
- data/lib/llm/repl/markdown.rb +25 -1
- data/lib/llm/repl/node.rb +7 -0
- data/lib/llm/repl/status.rb +17 -3
- data/lib/llm/repl/window.rb +91 -11
- data/lib/llm/repl.rb +18 -8
- data/lib/llm/sequel/plugin.rb +19 -5
- data/lib/llm/skill.rb +21 -8
- data/lib/llm/stream.rb +27 -0
- data/lib/llm/tool.rb +3 -5
- data/lib/llm/tools/rg.rb +2 -1
- data/lib/llm/transport/curb.rb +23 -3
- data/lib/llm/usage.rb +155 -9
- data/lib/llm/version.rb +1 -1
- data/lib/llm.rb +111 -29
- data/llm.gemspec +16 -11
- metadata +89 -35
- /data/{resources → docs}/deepdive/advanced/transformer.md +0 -0
- /data/{resources → docs}/deepdive/advanced/transports.md +0 -0
- /data/{resources/deepdive/fundamentals → docs/deepdive/features}/builtin_tools.md +0 -0
- /data/{resources/deepdive/fundamentals → docs/deepdive/features}/database.md +0 -0
- /data/{resources/deepdive/fundamentals → docs/deepdive/features}/embeddings.md +0 -0
- /data/{resources → docs}/deepdive/fundamentals/schema.md +0 -0
- /data/{resources/deepdive/everything_else → docs/deepdive/media}/audio.md +0 -0
- /data/{resources/deepdive/everything_else → docs/deepdive/media}/images.md +0 -0
- /data/{resources/deepdive/everything_else → docs/deepdive/media}/ocr.md +0 -0
- /data/{resources → docs}/deepdive/protocols/a2a.md +0 -0
- /data/{resources → docs}/deepdive/protocols/mcp.md +0 -0
- /data/{resources/deepdive/everything_else → docs/deepdive/reference}/object.md +0 -0
data/README.md
CHANGED
|
@@ -14,155 +14,17 @@
|
|
|
14
14
|
|
|
15
15
|
Welcome to the canonical llm.rb repository.
|
|
16
16
|
|
|
17
|
-
llm.rb is an advanced runtime for building
|
|
18
|
-
on CRuby. It has zero runtime dependencies by default,
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
contexts, and agents.
|
|
17
|
+
llm.rb is an advanced runtime for building agentic AI applications
|
|
18
|
+
on CRuby. It has zero runtime dependencies by default, it supports
|
|
19
|
+
concurrent and parallel tool execution and has a single coherent API
|
|
20
|
+
that spans 13+ providers. Streaming, tools, guards, compaction, the
|
|
21
|
+
REPL, builtin MCP/A2A support and the database integrations all build
|
|
22
|
+
on the same three concepts: providers, contexts, and agents.
|
|
23
23
|
|
|
24
24
|
Once you learn the fundamentals, everything else falls into place
|
|
25
25
|
naturally. Some features, such as ActiveRecord support, require
|
|
26
26
|
optional dependencies that are opt-in.
|
|
27
27
|
|
|
28
|
-
## Features
|
|
29
|
-
|
|
30
|
-
One runtime, 12+ providers. The same API drives OpenAI, Anthropic,
|
|
31
|
-
Google Gemini, Moonshot (kimi), Mistral, DeepSeek, DeepInfra, xAI, Z.ai,
|
|
32
|
-
AWS Bedrock, Ollama, and llama.cpp, so switching models or providers
|
|
33
|
-
can be done with minimal code change.
|
|
34
|
-
|
|
35
|
-
<details>
|
|
36
|
-
<summary><b>Agents</b></summary>
|
|
37
|
-
|
|
38
|
-
* **First-class support** <br>
|
|
39
|
-
llm.rb is designed to build agents. They can be attached to a
|
|
40
|
-
terminal-based read-eval-print loop (repl), persisted to disk
|
|
41
|
-
or a database column, run tools concurrently and be safely
|
|
42
|
-
interrupted.
|
|
43
|
-
|
|
44
|
-
* **Builtin REPL** <br>
|
|
45
|
-
A curses-based TUI for talking to an agent interactively. It
|
|
46
|
-
renders markdown, shows a live status line with context usage
|
|
47
|
-
and running cost, and recalls previous turns, so a
|
|
48
|
-
conversation survives a restart.
|
|
49
|
-
|
|
50
|
-
* **Persistence** <br>
|
|
51
|
-
Set `path:` and the agent saves its conversation to disk
|
|
52
|
-
automatically. ActiveRecord and Sequel support keep the same
|
|
53
|
-
state in a single database column, so you pick the storage
|
|
54
|
-
and the API stays identical.
|
|
55
|
-
|
|
56
|
-
</details>
|
|
57
|
-
|
|
58
|
-
<details>
|
|
59
|
-
<summary><b>MCP & A2A</b></summary>
|
|
60
|
-
|
|
61
|
-
* **MCP** <br>
|
|
62
|
-
The Model Context Protocol is first-class. Point an MCP client
|
|
63
|
-
at any tool server over stdio or HTTP, and its tools translate
|
|
64
|
-
into local `LLM::Tool` subclasses, with the same tracing and
|
|
65
|
-
error handling.
|
|
66
|
-
|
|
67
|
-
* **A2A** <br>
|
|
68
|
-
The Agent 2 Agent protocol is first-class. Point an A2A client
|
|
69
|
-
at another agent over HTTP or JSON-RPC, and call its skills
|
|
70
|
-
exactly like local tools.
|
|
71
|
-
|
|
72
|
-
</details>
|
|
73
|
-
|
|
74
|
-
<details>
|
|
75
|
-
<summary><b>ORM</b></summary>
|
|
76
|
-
|
|
77
|
-
* **ActiveRecord** <br>
|
|
78
|
-
Add `acts_as_agent` to a model and the agent state lives in a
|
|
79
|
-
single database column, saved after every turn and restored
|
|
80
|
-
on load. Works in Rack and Rails apps, with `jsonb` on
|
|
81
|
-
PostgreSQL.
|
|
82
|
-
|
|
83
|
-
* **Sequel** <br>
|
|
84
|
-
Add `plugin :agent` to a Sequel model for the same single-
|
|
85
|
-
column persistence, with the `pg_json` extension loaded
|
|
86
|
-
automatically on PostgreSQL.
|
|
87
|
-
|
|
88
|
-
</details>
|
|
89
|
-
|
|
90
|
-
<details>
|
|
91
|
-
<summary><b>RAG</b></summary>
|
|
92
|
-
|
|
93
|
-
* **RAG, out of the box** <br>
|
|
94
|
-
Embeddings, OCR, and OpenAI's vector stores API come first-
|
|
95
|
-
class. Ground answers in your own documents, with vectors in
|
|
96
|
-
a managed store or in your own database such as sqlite-vec
|
|
97
|
-
or pgvector.
|
|
98
|
-
|
|
99
|
-
</details>
|
|
100
|
-
|
|
101
|
-
<details>
|
|
102
|
-
<summary><b>Runtime</b></summary>
|
|
103
|
-
|
|
104
|
-
* **Streaming** <br>
|
|
105
|
-
Streaming is first-class, with structured callbacks for
|
|
106
|
-
content, reasoning, and tool calls. Tools can start while
|
|
107
|
-
the model is still talking, so the first result lands
|
|
108
|
-
before the response finishes.
|
|
109
|
-
|
|
110
|
-
* **Concurrency** <br>
|
|
111
|
-
Six ways to run tools: sequential, threads, async, fibers,
|
|
112
|
-
forks, and ractors. Plus three HTTP backends, so you pick
|
|
113
|
-
the concurrency model that fits the workload, not the other
|
|
114
|
-
way around.
|
|
115
|
-
|
|
116
|
-
* **Interruption** <br>
|
|
117
|
-
Cancel an in-flight request or a running tool at any moment,
|
|
118
|
-
on any transport or concurrency strategy. A stuck call never
|
|
119
|
-
leaves a thread running that you can't stop.
|
|
120
|
-
|
|
121
|
-
</details>
|
|
122
|
-
|
|
123
|
-
<details>
|
|
124
|
-
<summary><b>Provider extras</b></summary>
|
|
125
|
-
|
|
126
|
-
* **DeepSeek-optimized** <br>
|
|
127
|
-
DeepSeek is the most cost-effective option for API users, and the
|
|
128
|
-
runtime closes its gaps: [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
|
|
129
|
-
makes structured outputs work despite no official API, and
|
|
130
|
-
`images.create`/`edit` produce SVG vector graphics.
|
|
131
|
-
</details>
|
|
132
|
-
|
|
133
|
-
<details>
|
|
134
|
-
<summary><b>Portable</b></summary>
|
|
135
|
-
|
|
136
|
-
* **mruby-llm** <br>
|
|
137
|
-
The same runtime runs on mruby as
|
|
138
|
-
[mruby-llm](https://github.com/r-uby-dev/mruby-llm), with an
|
|
139
|
-
almost identical interface and the same set of capabilities.
|
|
140
|
-
|
|
141
|
-
</details>
|
|
142
|
-
|
|
143
|
-
<details>
|
|
144
|
-
<summary><b>Everything else</b></summary>
|
|
145
|
-
|
|
146
|
-
* **Skills** <br>
|
|
147
|
-
Write a SKILL.md, get a tool. The runtime spawns a
|
|
148
|
-
disposable subagent with the skill's instructions and tool
|
|
149
|
-
set for one turn, then discards it. Fresh and stateless
|
|
150
|
-
every call.
|
|
151
|
-
|
|
152
|
-
* **A unified plugin family** <br>
|
|
153
|
-
Compactors, transformers, and guards all share one
|
|
154
|
-
interface. Context management, message rewriting, and tool
|
|
155
|
-
supervision (policy, quotas, loop detection) plug in the
|
|
156
|
-
same way and compose freely.
|
|
157
|
-
|
|
158
|
-
* **Cost and usage tracking** <br>
|
|
159
|
-
Every context tracks its own cost and token usage, per turn.
|
|
160
|
-
Break the spend down by input, output, cache, and reasoning,
|
|
161
|
-
so the exact cost of any conversation is visible at a
|
|
162
|
-
glance.
|
|
163
|
-
|
|
164
|
-
</details>
|
|
165
|
-
|
|
166
28
|
## Install
|
|
167
29
|
|
|
168
30
|
```bash
|
|
@@ -171,7 +33,7 @@ gem install llm.rb
|
|
|
171
33
|
|
|
172
34
|
## Quick start
|
|
173
35
|
|
|
174
|
-
|
|
36
|
+
### Agents
|
|
175
37
|
|
|
176
38
|
The
|
|
177
39
|
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
@@ -186,82 +48,82 @@ require "llm"
|
|
|
186
48
|
|
|
187
49
|
llm = LLM.deepseek(key: ENV["KEY"])
|
|
188
50
|
agent = LLM::Agent.new(llm, stream: $stdout)
|
|
189
|
-
agent.talk "
|
|
51
|
+
agent.talk "hello world"
|
|
190
52
|
```
|
|
53
|
+
<details>
|
|
54
|
+
<summary>Stream</summary>
|
|
55
|
+
<br>
|
|
191
56
|
|
|
192
|
-
|
|
57
|
+
Streams can be simple IO objects or subclasses of
|
|
58
|
+
[`LLM::Stream`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
|
|
59
|
+
with structured callbacks for content,
|
|
60
|
+
reasoning, tool calls, tool returns, and compaction.
|
|
61
|
+
Streams can also observe message transformers, which rewrite
|
|
62
|
+
outgoing messages before they reach the provider.
|
|
193
63
|
|
|
194
|
-
[
|
|
195
|
-
|
|
196
|
-
corresponding class accessor: `name`, `description`, `model`, `tools`,
|
|
197
|
-
`instructions`, `schema`, `stream`, `tracer`, `concurrency`, `confirm`,
|
|
198
|
-
`path`, `skills`, and `tool_budget`. All options are optional; zero or
|
|
199
|
-
more can be set.
|
|
200
|
-
An error is raised for unknown keys so that typos are caught early.
|
|
64
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/stream/)
|
|
65
|
+
to learn more.
|
|
201
66
|
|
|
202
67
|
```ruby
|
|
203
|
-
class
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
end
|
|
68
|
+
class MyStream < LLM::Stream
|
|
69
|
+
# Visible assistant output.
|
|
70
|
+
def on_content(content)
|
|
71
|
+
print content
|
|
72
|
+
end
|
|
209
73
|
|
|
210
|
-
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
74
|
+
# Reasoning output streamed separately from visible content.
|
|
75
|
+
def on_reasoning_content(content)
|
|
76
|
+
warn content
|
|
77
|
+
end
|
|
214
78
|
|
|
215
|
-
|
|
79
|
+
# A streamed tool call has been fully parsed.
|
|
80
|
+
def on_tool_call(tool)
|
|
81
|
+
end
|
|
216
82
|
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
code. For database-backed persistence, ActiveRecord and Sequel
|
|
221
|
-
integrations are also available (see the
|
|
222
|
-
[database deepdive](https://r.uby.dev/llm/deepdive/advanced/database)
|
|
223
|
-
for details). All persistence options use the same underlying
|
|
224
|
-
serialization.
|
|
83
|
+
# Queued streamed tool work has returned.
|
|
84
|
+
def on_tool_return(tool, result)
|
|
85
|
+
end
|
|
225
86
|
|
|
226
|
-
|
|
227
|
-
|
|
87
|
+
# Before a transformer rewrites an outgoing message.
|
|
88
|
+
def on_transform(transformer)
|
|
89
|
+
end
|
|
228
90
|
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
|
|
91
|
+
# Aftter a transformer rewrites an outgoing message.
|
|
92
|
+
def on_transform_finish(transformer)
|
|
93
|
+
end
|
|
232
94
|
|
|
233
|
-
#
|
|
234
|
-
|
|
235
|
-
|
|
236
|
-
```
|
|
95
|
+
# Before a compactor trims the conversation.
|
|
96
|
+
def on_compaction(compactor)
|
|
97
|
+
end
|
|
237
98
|
|
|
238
|
-
|
|
99
|
+
# After a compactor trims the conversation.
|
|
100
|
+
def on_compaction_finish(compactor)
|
|
101
|
+
end
|
|
239
102
|
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
and it is what
|
|
244
|
-
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
245
|
-
uses under the hood.
|
|
246
|
-
It requires that the tool call loop be managed manually -
|
|
247
|
-
sometimes that can be useful, but usually for advanced use-cases.
|
|
248
|
-
If you're new to llm.rb, try
|
|
249
|
-
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html) first.
|
|
103
|
+
# Before a skill's subagent runs.
|
|
104
|
+
def on_skill_call(skill)
|
|
105
|
+
end
|
|
250
106
|
|
|
251
|
-
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
|
|
107
|
+
# After a skill's subagent runs.
|
|
108
|
+
# The subagent that ran it, the skill, and its response are passed
|
|
109
|
+
# through, so you can introspect the agent, tally skill usage, or
|
|
110
|
+
# track costs.
|
|
111
|
+
def on_skill_return(agent, skill, result)
|
|
112
|
+
end
|
|
255
113
|
|
|
256
|
-
|
|
257
|
-
|
|
114
|
+
# A request was rate limited and will be retried.
|
|
115
|
+
def on_rate_limit(error)
|
|
116
|
+
end
|
|
117
|
+
end
|
|
258
118
|
|
|
259
119
|
llm = LLM.deepseek(key: ENV["KEY"])
|
|
260
|
-
|
|
261
|
-
|
|
120
|
+
agent = LLM::Agent.new(llm, stream: MyStream.new)
|
|
121
|
+
agent.talk "Explain Ruby fibers."
|
|
262
122
|
```
|
|
123
|
+
</details>
|
|
263
124
|
|
|
264
|
-
|
|
125
|
+
<details><summary>Tools</summary>
|
|
126
|
+
<br>
|
|
265
127
|
|
|
266
128
|
Subclasses of
|
|
267
129
|
[`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
|
|
@@ -271,9 +133,7 @@ call them on your behalf, and they're one of the most powerful features
|
|
|
271
133
|
for extending the feature set or abilities of a model.
|
|
272
134
|
|
|
273
135
|
The runtime also ships with a catalog of built-in tools for
|
|
274
|
-
filesystem, search, and shell operations. See the
|
|
275
|
-
[deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/builtin_tools)
|
|
276
|
-
for details.
|
|
136
|
+
filesystem, search, and shell operations. <br> See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/tools) to learn more.
|
|
277
137
|
|
|
278
138
|
```ruby
|
|
279
139
|
class ReadFile < LLM::Tool
|
|
@@ -286,114 +146,145 @@ class ReadFile < LLM::Tool
|
|
|
286
146
|
{contents: File.read(path)}
|
|
287
147
|
end
|
|
288
148
|
end
|
|
149
|
+
|
|
150
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
151
|
+
agent = LLM::Agent.new(llm, tools: [ReadFile], stream: $stdout)
|
|
152
|
+
agent.talk "summarize README.md"
|
|
289
153
|
```
|
|
154
|
+
</details>
|
|
155
|
+
<details>
|
|
156
|
+
<summary>Skills</summary>
|
|
157
|
+
<br>
|
|
290
158
|
|
|
291
|
-
|
|
159
|
+
A skill turns a markdown file into a callable tool. When the model
|
|
160
|
+
calls it, the runtime spawns a subagent with the skill's instructions
|
|
161
|
+
as its system prompt and the skill's own tool set. The subagent runs
|
|
162
|
+
one turn and returns the result, then is discarded. Each call
|
|
163
|
+
is fresh and stateless.
|
|
292
164
|
|
|
293
|
-
[
|
|
294
|
-
|
|
295
|
-
the
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
`description`, `parameters`, `required`, and `defaults`:
|
|
165
|
+
A [LLM::Stream](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
|
|
166
|
+
can be notified as a skill starts and when it returns. The `on_skill_return`
|
|
167
|
+
callback hands back the subagent that ran the skill, so you can inspect
|
|
168
|
+
its conversation, measure its usage, track costs or add a verification
|
|
169
|
+
step (eg `subagent.talk("verify your work")`).
|
|
299
170
|
|
|
300
|
-
|
|
301
|
-
class MathTool < LLM::Tool
|
|
302
|
-
set name: "math",
|
|
303
|
-
description: "Performs arithmetic",
|
|
304
|
-
parameters: [
|
|
305
|
-
[:x, Integer, "first number" , {required: true}],
|
|
306
|
-
[:y, Integer, "second number", {default: 0}]
|
|
307
|
-
]
|
|
308
|
-
|
|
309
|
-
def call(x:, y: 0)
|
|
310
|
-
{result: x + y}
|
|
311
|
-
end
|
|
312
|
-
end
|
|
313
|
-
```
|
|
171
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/skills) to learn more.
|
|
314
172
|
|
|
315
|
-
|
|
173
|
+
##### summary.md
|
|
316
174
|
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
[deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transformer)).
|
|
175
|
+
```markdown
|
|
176
|
+
---
|
|
177
|
+
name: summary
|
|
178
|
+
description: Reads recent git history and writes a summary
|
|
179
|
+
tools: all
|
|
180
|
+
---
|
|
324
181
|
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
print content
|
|
329
|
-
end
|
|
182
|
+
Collect the recent git log, analyze each commit,
|
|
183
|
+
and write a summary to summary.txt.
|
|
184
|
+
```
|
|
330
185
|
|
|
331
|
-
|
|
332
|
-
warn content
|
|
333
|
-
end
|
|
334
|
-
end
|
|
186
|
+
##### agent.rb
|
|
335
187
|
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
|
|
188
|
+
```ruby
|
|
189
|
+
require "llm"
|
|
190
|
+
|
|
191
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
192
|
+
agent = LLM::Agent.new(llm, skills: ["summary.md"])
|
|
193
|
+
agent.talk "Summarize the last week of work"
|
|
339
194
|
```
|
|
195
|
+
</details>
|
|
340
196
|
|
|
341
|
-
|
|
197
|
+
<details>
|
|
198
|
+
<summary>Concurrency</summary>
|
|
199
|
+
<br>
|
|
342
200
|
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk-instance_method),
|
|
347
|
-
[`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk-instance_method),
|
|
348
|
-
or
|
|
349
|
-
[`LLM::Provider#complete`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#complete-instance_method)
|
|
350
|
-
to receive validated JSON instead of free text. Schemas work alongside tools and streams.
|
|
201
|
+
The runtime supports six different concurrency strategies that have
|
|
202
|
+
different attributes. The choice between all of them often depends
|
|
203
|
+
on the requirements of your application.
|
|
351
204
|
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
|
|
357
|
-
|
|
205
|
+
IO-bound tools are a good fit for the `:async`, `:thread`,
|
|
206
|
+
and `:fiber` strategies while true parallelism can be achieved
|
|
207
|
+
with the `:fork` and `:ractor` strategies. The
|
|
208
|
+
`:sequential` strategy runs tools one at a time and is the default.
|
|
209
|
+
The `:fork` strategy also provides a separate process that offers
|
|
210
|
+
isolation from its parent.
|
|
358
211
|
|
|
359
|
-
|
|
360
|
-
[`LLM::DeepSeek`](https://r.uby.dev/api-docs/llm.rb/LLM/DeepSeek.html)
|
|
361
|
-
provider includes runtime-level optimisations such as structured
|
|
362
|
-
output support (despite no official structured outputs API) and
|
|
363
|
-
SVG image generation. This example uses
|
|
364
|
-
[`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html) with
|
|
365
|
-
DeepSeek:
|
|
212
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/concurrency) to learn more.
|
|
366
213
|
|
|
367
214
|
```ruby
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
property :temperature, Float, "Current temperature"
|
|
371
|
-
property :conditions, String, "Weather conditions"
|
|
372
|
-
required %i[city temperature conditions]
|
|
373
|
-
end
|
|
215
|
+
require "llm"
|
|
216
|
+
require "llm/tools"
|
|
374
217
|
|
|
375
|
-
llm
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
218
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
219
|
+
tools = LLM::Tool.subclasses
|
|
220
|
+
agent = LLM::Agent.new(llm, tools:, concurrency: :fork)
|
|
221
|
+
agent.talk "Run the tools in parallel"
|
|
379
222
|
```
|
|
380
223
|
|
|
381
|
-
|
|
224
|
+
</details>
|
|
225
|
+
<details>
|
|
226
|
+
<summary>Cancellation</summary>
|
|
227
|
+
<br>
|
|
228
|
+
|
|
229
|
+
Abort a request mid-stream and interrupt any running tools with
|
|
230
|
+
[`LLM::Agent#interrupt!`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#interrupt!)
|
|
231
|
+
(or `cancel!`), from any thread. The runtime raises
|
|
232
|
+
[`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html)
|
|
233
|
+
on the caller and on every active tool. A forked tool gets interrupted over
|
|
234
|
+
the control channel, a ractor via message passing, and pending tools
|
|
235
|
+
are stopped before they run. The in-flight HTTP request is closed
|
|
236
|
+
too, so a turn you no longer want stops without burning tokens.
|
|
237
|
+
|
|
238
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/cancellation) to learn more.
|
|
239
|
+
|
|
240
|
+
```ruby
|
|
241
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
242
|
+
agent = LLM::Agent.new(llm)
|
|
243
|
+
Thread.new { sleep(1); agent.cancel! }
|
|
244
|
+
|
|
245
|
+
begin
|
|
246
|
+
agent.talk "write a very long poem", stream: $stdout
|
|
247
|
+
rescue LLM::Interrupt
|
|
248
|
+
puts "cancelled"
|
|
249
|
+
end
|
|
250
|
+
```
|
|
251
|
+
</details>
|
|
252
|
+
<details>
|
|
253
|
+
<summary>Console (<code>binding.pry</code> for agents)</summary>
|
|
254
|
+
<br>
|
|
382
255
|
|
|
383
256
|
The [LLM::Agent#repl](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#repl-instance_method)
|
|
384
|
-
method drops you into a
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
|
|
257
|
+
method drops you into a highly capable read-eval-print loop (REPL)
|
|
258
|
+
that is built on top of curses. It can help you debug agents,
|
|
259
|
+
test your tools, connect to MCP servers, and even A2A agents.
|
|
260
|
+
The REPL stands out because it connects to the surrounding
|
|
261
|
+
runtime and it can be extended by your code. Think of it as
|
|
389
262
|
`binding.pry` but for agents.
|
|
390
263
|
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
[
|
|
396
|
-
|
|
264
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/repl) to learn more.
|
|
265
|
+
|
|
266
|
+
##### Demo
|
|
267
|
+
|
|
268
|
+
[Watch in high quality on asciinema](https://asciinema.org/a/OsS8wwaasKasoDDz)
|
|
269
|
+
|
|
270
|
+

|
|
271
|
+
|
|
272
|
+
|
|
273
|
+
##### Installation
|
|
274
|
+
|
|
275
|
+
The REPL is distributed with llm.rb so you don't have to install
|
|
276
|
+
a separate gem but it requires a number of optional dependencies
|
|
277
|
+
to be installed separately. The following gems provide the full
|
|
278
|
+
experience:
|
|
279
|
+
|
|
280
|
+
gem install curses kramdown xchan.rb test-cmd.rb
|
|
281
|
+
|
|
282
|
+
##### Persistence
|
|
283
|
+
|
|
284
|
+
The `path:` option can be set on an agent for automatic persistence
|
|
285
|
+
across REPL sessions. The `tools:` option attaches extra tools
|
|
286
|
+
for the duration of the session. Recall previous turns with Ctrl+P and
|
|
287
|
+
Ctrl+N.
|
|
397
288
|
|
|
398
289
|
```ruby
|
|
399
290
|
require "llm"
|
|
@@ -407,20 +298,110 @@ agent.repl(tools: LLM::Tool.subclasses)
|
|
|
407
298
|
##### CLI
|
|
408
299
|
|
|
409
300
|
The `llm.rb` executable is available on your PATH after installation.
|
|
410
|
-
It starts a REPL session from any directory
|
|
301
|
+
It starts a REPL session from any directory.The CLI auto-detects your
|
|
302
|
+
provider from standard environment variables (`DEEPSEEK_API_KEY`,
|
|
303
|
+
`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, etc.). Persistent sessions are
|
|
304
|
+
stored under `~/.llm.rb/` and restored automatically on your next visit.
|
|
411
305
|
|
|
412
306
|
```bash
|
|
413
307
|
llm.rb # auto-detect from $DEEPSEEK_API_KEY
|
|
414
308
|
llm.rb -p openai # use OpenAI explicitly
|
|
415
309
|
llm.rb -t # temporary session, no persistence
|
|
416
310
|
```
|
|
311
|
+
</details>
|
|
312
|
+
<details>
|
|
313
|
+
<summary>Persistence</summary>
|
|
314
|
+
<br>
|
|
315
|
+
|
|
316
|
+
Set `path:` on an agent for automatic filesystem persistence:
|
|
317
|
+
the agent restores conversation history from the file on startup
|
|
318
|
+
and saves it back after every turn, with no manual serialization
|
|
319
|
+
code. For database-backed persistence, ActiveRecord and Sequel
|
|
320
|
+
integrations are also available. All persistence options use the same
|
|
321
|
+
underlying serialization.
|
|
322
|
+
|
|
323
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/database) to learn more.
|
|
324
|
+
|
|
325
|
+
```ruby
|
|
326
|
+
require "llm"
|
|
327
|
+
|
|
328
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
329
|
+
agent = LLM::Agent.new(llm, path: "session.json")
|
|
330
|
+
agent.talk "remember my name is robert"
|
|
331
|
+
|
|
332
|
+
# Next time, the conversation is restored automatically:
|
|
333
|
+
agent = LLM::Agent.new(llm, path: "session.json")
|
|
334
|
+
agent.talk "what's my name?"
|
|
335
|
+
```
|
|
336
|
+
</details>
|
|
337
|
+
<details><summary>ActiveRecord | Sequel</summary>
|
|
338
|
+
<br>
|
|
339
|
+
|
|
340
|
+
Because both
|
|
341
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) and
|
|
342
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
343
|
+
can be serialized to JSON and stored in a simple string, both ActiveRecord
|
|
344
|
+
and Sequel support can be implemented within a single column on a single row.
|
|
345
|
+
|
|
346
|
+
The runtime includes first-class support for both ActiveRecord / Sequel, and
|
|
347
|
+
for both Rack-based / Rails-based applications. On databases
|
|
348
|
+
where it is supported, such as PostgreSQL, the column can be optimized by using
|
|
349
|
+
the `jsonb` type.
|
|
350
|
+
|
|
351
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/database) to learn more.
|
|
352
|
+
|
|
353
|
+
```ruby
|
|
354
|
+
require "active_record"
|
|
355
|
+
require "llm"
|
|
356
|
+
require "llm/active_record"
|
|
357
|
+
|
|
358
|
+
class Email < ApplicationRecord
|
|
359
|
+
acts_as_agent do |agent|
|
|
360
|
+
agent.set name: "mail",
|
|
361
|
+
instructions: "Write concise, friendly replies to emails",
|
|
362
|
+
model: "deepseek-v4-pro"
|
|
363
|
+
end
|
|
364
|
+
|
|
365
|
+
def draft_reply!
|
|
366
|
+
talk("Draft a reply to:\n\n#{body}")
|
|
367
|
+
end
|
|
368
|
+
|
|
369
|
+
def summarize
|
|
370
|
+
talk("Summarize this email thread in a few sentences")
|
|
371
|
+
end
|
|
372
|
+
|
|
373
|
+
private
|
|
374
|
+
|
|
375
|
+
##
|
|
376
|
+
# By convention, this method defines the provider for a model.
|
|
377
|
+
# If necessary, it can be renamed with: provider: :your_method.
|
|
378
|
+
def set_provider
|
|
379
|
+
LLM.deepseek(key: ENV["KEY"])
|
|
380
|
+
end
|
|
381
|
+
|
|
382
|
+
##
|
|
383
|
+
# By convention, this method returns the context options given
|
|
384
|
+
# to LLM::Context or LLM::Agent. This method can be left undefined.
|
|
385
|
+
def set_context
|
|
386
|
+
{}
|
|
387
|
+
end
|
|
388
|
+
end
|
|
417
389
|
|
|
418
|
-
|
|
419
|
-
|
|
420
|
-
|
|
421
|
-
|
|
390
|
+
email = Email.create!(subject: "Streaming support", body: "How do I stream responses?")
|
|
391
|
+
email.draft_reply!
|
|
392
|
+
|
|
393
|
+
##
|
|
394
|
+
# The conversation (the email and the draft
|
|
395
|
+
# reply) is persisted to the email's column. A
|
|
396
|
+
# fresh instance restores it and continues the
|
|
397
|
+
# thread, so the summary below knows what was
|
|
398
|
+
# already drafted:
|
|
399
|
+
Email.find(email.id).summarize
|
|
400
|
+
```
|
|
401
|
+
</details>
|
|
422
402
|
|
|
423
|
-
|
|
403
|
+
<details><summary>MCP</summary>
|
|
404
|
+
<br>
|
|
424
405
|
|
|
425
406
|
The Model Context Protocol (MCP) has first-class support
|
|
426
407
|
in llm.rb. The stdio and http transports work out of the
|
|
@@ -430,6 +411,9 @@ used with
|
|
|
430
411
|
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) or
|
|
431
412
|
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
|
|
432
413
|
|
|
414
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/protocols/mcp/), and the
|
|
415
|
+
[deepdive.md on persistent connections](https://r.uby.dev/llm/deepdive/features/transports) to learn more.
|
|
416
|
+
|
|
433
417
|
```ruby
|
|
434
418
|
require "llm"
|
|
435
419
|
|
|
@@ -438,24 +422,9 @@ mcp = LLM::MCP.stdio(argv: ["ruby", "server.rb"])
|
|
|
438
422
|
agent = LLM::Agent.new(llm, stream: $stdout, tools: mcp.tools)
|
|
439
423
|
agent.talk "Run the tool"
|
|
440
424
|
```
|
|
441
|
-
|
|
442
|
-
|
|
443
|
-
|
|
444
|
-
Set `persistent: true` on HTTP transports to reuse connections
|
|
445
|
-
across requests. This uses
|
|
446
|
-
[`Net::HTTP::Persistent`](https://github.com/drbrain/net-http-persistent)
|
|
447
|
-
under the hood and avoids opening a new TCP connection for every
|
|
448
|
-
request:
|
|
449
|
-
|
|
450
|
-
```ruby
|
|
451
|
-
mcp = LLM::MCP.http(
|
|
452
|
-
url: "https://api.githubcopilot.com/mcp/",
|
|
453
|
-
headers: {"Authorization" => "Bearer #{ENV.fetch('GITHUB_PAT')}"},
|
|
454
|
-
persistent: true
|
|
455
|
-
)
|
|
456
|
-
```
|
|
457
|
-
|
|
458
|
-
#### LLM::A2A
|
|
425
|
+
</details>
|
|
426
|
+
<details><summary>A2A</summary>
|
|
427
|
+
<br>
|
|
459
428
|
|
|
460
429
|
The Agent 2 Agent (A2A) protocol has first-class support
|
|
461
430
|
in llm.rb. The http and jsonrpc transports work out of the
|
|
@@ -465,6 +434,9 @@ used with
|
|
|
465
434
|
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) or
|
|
466
435
|
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
|
|
467
436
|
|
|
437
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/protocols/a2a/), and the
|
|
438
|
+
[deepdive.md on persistent connections](https://r.uby.dev/llm/deepdive/features/transports) to learn more.
|
|
439
|
+
|
|
468
440
|
```ruby
|
|
469
441
|
require "llm"
|
|
470
442
|
|
|
@@ -473,21 +445,51 @@ a2a = LLM::A2A.rest(url: "https://remote-agent.example.com")
|
|
|
473
445
|
agent = LLM::Agent.new(llm, stream: $stdout, tools: a2a.skills)
|
|
474
446
|
agent.talk "Run the skill"
|
|
475
447
|
```
|
|
448
|
+
</details>
|
|
449
|
+
|
|
450
|
+
<details><summary>Structured outputs</summary>
|
|
451
|
+
<br>
|
|
452
|
+
|
|
453
|
+
[`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
|
|
454
|
+
subclasses produce typed, structured
|
|
455
|
+
output from any model call. Pass a schema to
|
|
456
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk-instance_method),
|
|
457
|
+
[`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk-instance_method),
|
|
458
|
+
or
|
|
459
|
+
[`LLM::Provider#complete`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#complete-instance_method)
|
|
460
|
+
to receive validated JSON instead of free text. Schemas work alongside tools and streams.
|
|
476
461
|
|
|
477
|
-
|
|
462
|
+
[`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
|
|
463
|
+
can define objects, arrays, enums, nested schemas,
|
|
464
|
+
and more. It is also used internally by
|
|
465
|
+
[`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html) for parameter
|
|
466
|
+
definitions, so you already benefit from it when you declare tool
|
|
467
|
+
parameters.
|
|
478
468
|
|
|
479
|
-
|
|
480
|
-
|
|
481
|
-
|
|
482
|
-
|
|
483
|
-
|
|
469
|
+
The
|
|
470
|
+
[`LLM::DeepSeek`](https://r.uby.dev/api-docs/llm.rb/LLM/DeepSeek.html)
|
|
471
|
+
provider includes runtime-level optimisations such as structured
|
|
472
|
+
output support (despite no official structured outputs API) and
|
|
473
|
+
SVG image generation. This example uses
|
|
474
|
+
[`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html) with
|
|
475
|
+
DeepSeek:
|
|
484
476
|
|
|
485
477
|
```ruby
|
|
486
|
-
|
|
487
|
-
|
|
488
|
-
|
|
478
|
+
class Weather < LLM::Schema
|
|
479
|
+
property :city, String, "The city name"
|
|
480
|
+
property :temperature, Number, "Current temperature"
|
|
481
|
+
property :conditions, String, "Weather conditions"
|
|
482
|
+
required %i[city temperature conditions]
|
|
483
|
+
end
|
|
489
484
|
|
|
490
|
-
|
|
485
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
486
|
+
agent = LLM::Agent.new(llm, schema: Weather)
|
|
487
|
+
res = agent.talk "Weather in Paris?"
|
|
488
|
+
res.content! # => {city: "Paris", temperature: 15.0, conditions: "Cloudy"}
|
|
489
|
+
```
|
|
490
|
+
</details>
|
|
491
|
+
<details><summary>Guards</summary>
|
|
492
|
+
<br>
|
|
491
493
|
|
|
492
494
|
[`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
|
|
493
495
|
is the hook that sees every tool call before it runs. A guard
|
|
@@ -507,6 +509,8 @@ and implement
|
|
|
507
509
|
The pending call arrives as `function:`. Return a value to close
|
|
508
510
|
the call, or `nil` to let it run:
|
|
509
511
|
|
|
512
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/guard) to learn more.
|
|
513
|
+
|
|
510
514
|
```ruby
|
|
511
515
|
class PolicyGuard < LLM::Guard
|
|
512
516
|
def call(function:)
|
|
@@ -517,141 +521,285 @@ class PolicyGuard < LLM::Guard
|
|
|
517
521
|
end
|
|
518
522
|
end
|
|
519
523
|
|
|
520
|
-
|
|
524
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
525
|
+
agent = LLM::Agent.new(llm, tools: [Shell, ReadFile], guard: PolicyGuard)
|
|
521
526
|
```
|
|
527
|
+
</details>
|
|
522
528
|
|
|
523
|
-
|
|
529
|
+
<details>
|
|
530
|
+
<summary>Transformers</summary>
|
|
531
|
+
<br>
|
|
524
532
|
|
|
525
|
-
|
|
526
|
-
|
|
527
|
-
|
|
528
|
-
|
|
529
|
-
|
|
530
|
-
[deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/skills).
|
|
533
|
+
It is possible to rewrite outgoing messages before they reach the provider with
|
|
534
|
+
[`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html).
|
|
535
|
+
Create a subclass and implement `call(message:)` to scrub sensitive data,
|
|
536
|
+
inject context, or normalize content. The transform runs automatically
|
|
537
|
+
on every turn, so you never have to change your prompt code.
|
|
531
538
|
|
|
532
|
-
|
|
539
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transformer) to learn more.
|
|
533
540
|
|
|
534
|
-
```
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
|
|
538
|
-
|
|
539
|
-
|
|
541
|
+
```ruby
|
|
542
|
+
class RedactEmails < LLM::Transformer
|
|
543
|
+
def call(message:)
|
|
544
|
+
content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
|
|
545
|
+
LLM::Message.new(message.role, content, message.extra)
|
|
546
|
+
end
|
|
547
|
+
end
|
|
540
548
|
|
|
541
|
-
|
|
542
|
-
|
|
549
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
550
|
+
agent = LLM::Agent.new(llm, transformer: RedactEmails)
|
|
551
|
+
agent.talk "Contact support@example.com for help"
|
|
543
552
|
```
|
|
553
|
+
</details>
|
|
544
554
|
|
|
545
|
-
|
|
555
|
+
<details>
|
|
556
|
+
<summary>Compactors</summary>
|
|
557
|
+
<br>
|
|
558
|
+
|
|
559
|
+
Every model has a context window: the finite number of tokens it can
|
|
560
|
+
consider in a single request. Generally a compactor will drop or
|
|
561
|
+
summarize older messages to keep the conversation within that window,
|
|
562
|
+
and it runs automatically before every turn. By default it is disabled
|
|
563
|
+
so it is a feature you must opt into.
|
|
564
|
+
|
|
565
|
+
[`LLM::Compactor::Truncate`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Truncate.html)
|
|
566
|
+
keeps the most recent messages via an integer count or a percentage like
|
|
567
|
+
`"80%"`. It preserves tool call and return pairs so the conversation
|
|
568
|
+
never contains an orphaned result. It is also possible to subclass
|
|
569
|
+
[`LLM::Compactor`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor.html)
|
|
570
|
+
to implement your own compactor with its own logic. Streams can observe the
|
|
571
|
+
process through the
|
|
572
|
+
[`LLM::Stream#on_compaction`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction)
|
|
573
|
+
and
|
|
574
|
+
[`LLM::Stream#on_compaction_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction_finish)
|
|
575
|
+
callbacks.
|
|
576
|
+
|
|
577
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/compaction) to learn more.
|
|
578
|
+
|
|
579
|
+
```ruby
|
|
580
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
581
|
+
agent = LLM::Agent.new(
|
|
582
|
+
llm,
|
|
583
|
+
compactor: LLM::Compactor::Truncate,
|
|
584
|
+
compactor_options: {keep: 64}
|
|
585
|
+
)
|
|
586
|
+
agent.talk "Hello"
|
|
587
|
+
```
|
|
588
|
+
</details>
|
|
589
|
+
|
|
590
|
+
<details>
|
|
591
|
+
<summary>Automatic retries</summary>
|
|
592
|
+
<br>
|
|
593
|
+
|
|
594
|
+
Rate-limited requests are retried automatically by default. Agents
|
|
595
|
+
retry a 429 up to five times with a growing backoff before giving
|
|
596
|
+
up, so most request failures resolve on their own. Set `retry_budget`
|
|
597
|
+
to change the number of retries, or `retry_budget: 0` to disable
|
|
598
|
+
them.
|
|
546
599
|
|
|
547
600
|
```ruby
|
|
548
601
|
require "llm"
|
|
549
602
|
|
|
550
|
-
llm
|
|
551
|
-
agent = LLM::Agent.new(llm,
|
|
552
|
-
agent.talk "
|
|
603
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
604
|
+
agent = LLM::Agent.new(llm, retry_budget: 0)
|
|
605
|
+
agent.talk "Hello"
|
|
553
606
|
```
|
|
554
607
|
|
|
555
|
-
|
|
608
|
+
</details>
|
|
609
|
+
|
|
556
610
|
|
|
557
|
-
|
|
558
|
-
|
|
559
|
-
|
|
560
|
-
be stored in a database that is optimized for storing
|
|
561
|
-
and querying vectors, such as SQLite's [sqlite-vec](https://github.com/asg017/sqlite-vec)
|
|
562
|
-
or PostgreSQL's [pg-vector](https://github.com/pgvector/pgvector).
|
|
611
|
+
<details>
|
|
612
|
+
<summary>Observability</summary>
|
|
613
|
+
<br>
|
|
563
614
|
|
|
564
|
-
|
|
565
|
-
|
|
566
|
-
|
|
567
|
-
|
|
615
|
+
Trace what an agent is doing by attaching a tracer. Hook into
|
|
616
|
+
requests, tool calls, and other runtime events to debug a
|
|
617
|
+
misbehaving agent, monitor latency, or export spans to an
|
|
618
|
+
observability backend. All built-in tracers share one interface,
|
|
619
|
+
so switching between them means changing a class name:
|
|
620
|
+
|
|
621
|
+
* [`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html): human-readable single-line logs to stderr, ideal during development.
|
|
622
|
+
* [`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html):
|
|
623
|
+
exports spans via OTLP for OpenTelemetry in production.
|
|
624
|
+
* [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html):
|
|
625
|
+
structured JSON to stdout or a file.
|
|
626
|
+
|
|
627
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/reference/tracer) to learn more.
|
|
628
|
+
|
|
629
|
+
```ruby
|
|
630
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
631
|
+
agent = LLM::Agent.new(llm, tracer: LLM::Tracer::PrettyLogger.new(llm))
|
|
632
|
+
agent.talk "Hello"
|
|
633
|
+
```
|
|
634
|
+
</details>
|
|
635
|
+
|
|
636
|
+
<details>
|
|
637
|
+
<summary>As a subclass</summary>
|
|
638
|
+
<br>
|
|
639
|
+
|
|
640
|
+
[`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
|
|
641
|
+
is a class-level DSL that accepts a Hash of properties. Each key resolves to a
|
|
642
|
+
corresponding class accessor: `name`, `description`, `model`, `tools`,
|
|
643
|
+
`instructions`, `schema`, `stream`, `tracer`, `concurrency`, `confirm`,
|
|
644
|
+
`path`, `skills`, `tool_budget`, and `retry_budget`. All options are
|
|
645
|
+
optional; zero or more can be set.
|
|
646
|
+
An error is raised for unknown keys so that typos are caught early.
|
|
568
647
|
|
|
569
648
|
```ruby
|
|
570
649
|
require "llm"
|
|
650
|
+
require "llm/tools"
|
|
571
651
|
|
|
572
|
-
|
|
573
|
-
|
|
574
|
-
|
|
652
|
+
class Agent < LLM::Agent
|
|
653
|
+
set name: "sysadmin",
|
|
654
|
+
description: "system administration agent",
|
|
655
|
+
model: "deepseek-v4-pro",
|
|
656
|
+
tools: [LLM::Tool::Shell]
|
|
657
|
+
end
|
|
575
658
|
|
|
576
|
-
|
|
577
|
-
|
|
578
|
-
|
|
579
|
-
title: "llm.rb",
|
|
580
|
-
body:,
|
|
581
|
-
embedding:,
|
|
582
|
-
)
|
|
659
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
660
|
+
agent = Agent.new(llm)
|
|
661
|
+
agent.talk "Run 'date'"
|
|
583
662
|
```
|
|
663
|
+
</details>
|
|
584
664
|
|
|
585
|
-
|
|
665
|
+
### Providers
|
|
586
666
|
|
|
587
|
-
|
|
588
|
-
|
|
589
|
-
|
|
667
|
+
Each provider is constructed with a class-level factory method on
|
|
668
|
+
`LLM`, and the resulting instance is passed to
|
|
669
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
670
|
+
or
|
|
671
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html). The
|
|
672
|
+
same API drives every one of them, so switching models is a one-line
|
|
673
|
+
change. See the [deepdive](https://r.uby.dev/llm/deepdive/fundamentals/providers)
|
|
674
|
+
for a full provider reference.
|
|
675
|
+
|
|
676
|
+
#### What providers does llm.rb support?
|
|
677
|
+
|
|
678
|
+
* **Anthropic** (`LLM.anthropic`)
|
|
679
|
+
* **Google** (`LLM.google`)
|
|
680
|
+
* **OpenAI** (`LLM.openai`)
|
|
681
|
+
* **DeepSeek** (`LLM.deepseek`)
|
|
682
|
+
* **DeepInfra** (`LLM.deepinfra`)
|
|
683
|
+
* **xAI** (`LLM.xai`)
|
|
684
|
+
* **Z.ai** (`LLM.zai`)
|
|
685
|
+
* **Moonshot (Kimi)** (`LLM.moonshot`)
|
|
686
|
+
* **Alibaba (Qwen3)** (`LLM.alibaba`, also `LLM.aliyun`)
|
|
687
|
+
* **Mistral** (`LLM.mistral`)
|
|
688
|
+
* **AWS Bedrock** (`LLM.bedrock`)
|
|
689
|
+
* **Ollama** (`LLM.ollama`)
|
|
690
|
+
* **llama.cpp** (`LLM.llamacpp`)
|
|
590
691
|
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
|
|
594
|
-
`:sequential` strategy runs tools one at a time and is the default.
|
|
595
|
-
The `:fork` strategy also provides a separate process that offers
|
|
596
|
-
isolation from its parent.
|
|
692
|
+
<details>
|
|
693
|
+
<summary>Implicit</summary>
|
|
694
|
+
<br>
|
|
597
695
|
|
|
598
|
-
|
|
599
|
-
|
|
696
|
+
Cloud providers can infer their API key automatically
|
|
697
|
+
from a set of common defaults that are defined by
|
|
698
|
+
the [models.dev](https://models.dev) registry that
|
|
699
|
+
is also distributed with llm.rb.
|
|
600
700
|
|
|
601
701
|
```ruby
|
|
602
|
-
|
|
702
|
+
llm = LLM.openai
|
|
703
|
+
llm = LLM.anthropic
|
|
704
|
+
llm = LLM.deepseek
|
|
705
|
+
llm = LLM.alibaba # also: LLM.aliyun
|
|
706
|
+
llm = LLM.moonshot
|
|
707
|
+
llm = LLM.mistral
|
|
708
|
+
```
|
|
709
|
+
</details>
|
|
710
|
+
<details>
|
|
711
|
+
<summary>Explicit</summary>
|
|
712
|
+
<br>
|
|
603
713
|
|
|
604
|
-
|
|
605
|
-
|
|
606
|
-
|
|
607
|
-
|
|
714
|
+
The `key` option can also be providied explicitly, and certain
|
|
715
|
+
providers (eg ollama, llamacpp) usually do not require an API
|
|
716
|
+
key at all.
|
|
717
|
+
|
|
718
|
+
```ruby
|
|
719
|
+
llm = LLM.openai(key: ENV["OPENAI_API_KEY"])
|
|
720
|
+
llm = LLM.anthropic(key: ENV["ANTHROPIC_API_KEY"])
|
|
721
|
+
llm = LLM.deepseek(key: ENV["DEEPSEEK_API_KEY"])
|
|
722
|
+
llm = LLM.alibaba(key: ENV["DASHSCOPE_API_KEY"]) # also: LLM.aliyun
|
|
723
|
+
llm = LLM.moonshot(key: ENV["MOONSHOT_API_KEY"])
|
|
724
|
+
llm = LLM.mistral(key: ENV["MISTRAL_API_KEY"])
|
|
608
725
|
```
|
|
726
|
+
</details>
|
|
609
727
|
|
|
610
|
-
|
|
728
|
+
<details>
|
|
729
|
+
<summary>Model Registry</summary>
|
|
730
|
+
<br>
|
|
611
731
|
|
|
612
|
-
|
|
613
|
-
[
|
|
614
|
-
|
|
615
|
-
|
|
616
|
-
and Sequel support can be implemented within a single column on a single row.
|
|
732
|
+
Each provider ships its model catalog, pricing, limits, and
|
|
733
|
+
modalities with the gem, sourced from [models.dev](https://models.dev).
|
|
734
|
+
Reach it from any provider, context, or agent, enumerate models, or
|
|
735
|
+
sort them by price.
|
|
617
736
|
|
|
618
|
-
|
|
619
|
-
for both Rack-based applications *and* Rails-based applications. On databases
|
|
620
|
-
where it is supported, such as PostgreSQL, the column can be optimized by using
|
|
621
|
-
the `jsonb` type.
|
|
737
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/reference/model_registry) to learn more.
|
|
622
738
|
|
|
623
739
|
```ruby
|
|
624
|
-
require "active_record"
|
|
625
740
|
require "llm"
|
|
626
|
-
require "llm/active_record"
|
|
627
741
|
|
|
628
|
-
|
|
629
|
-
|
|
630
|
-
|
|
631
|
-
|
|
632
|
-
|
|
633
|
-
|
|
742
|
+
llm = LLM.openai
|
|
743
|
+
registry = llm.registry # => LLM::Provider#registry
|
|
744
|
+
cheapest = registry.models.sort.first # => LLM::Model
|
|
745
|
+
cheapest.id # => "text-embedding-3-small"
|
|
746
|
+
cheapest.context_window # => 8191
|
|
747
|
+
cheapest.structured_output? # => false
|
|
748
|
+
```
|
|
749
|
+
</details>
|
|
634
750
|
|
|
635
|
-
|
|
751
|
+
<details>
|
|
752
|
+
<summary>Transports</summary>
|
|
753
|
+
<br>
|
|
636
754
|
|
|
637
|
-
|
|
638
|
-
|
|
639
|
-
|
|
640
|
-
|
|
641
|
-
|
|
755
|
+
The `transport:` option selects which HTTP library a provider uses for
|
|
756
|
+
network communication. Three backends ship out of the box: `net/http`
|
|
757
|
+
is always available and the default, `net/http/persistent` pools
|
|
758
|
+
connections for many requests to the same host, and `curb` wraps
|
|
759
|
+
libcurl. They share one interface, so switching is a one-word change.
|
|
642
760
|
|
|
643
|
-
|
|
644
|
-
# to LLM::Context or LLM::Agent.
|
|
645
|
-
def set_context
|
|
646
|
-
{}
|
|
647
|
-
end
|
|
648
|
-
end
|
|
761
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transports) to learn more.
|
|
649
762
|
|
|
650
|
-
|
|
651
|
-
|
|
763
|
+
```ruby
|
|
764
|
+
llm = LLM.deepseek(
|
|
765
|
+
key: ENV["KEY"],
|
|
766
|
+
transport: :net_http_persistent
|
|
767
|
+
)
|
|
768
|
+
```
|
|
769
|
+
</details>
|
|
770
|
+
|
|
771
|
+
### RAG
|
|
772
|
+
|
|
773
|
+
Most providers offer an embedding model that can be
|
|
774
|
+
used for semantic search, or similarity search. An
|
|
775
|
+
embedding model can generate embeddings that can then
|
|
776
|
+
be stored in a database that is optimized for storing
|
|
777
|
+
and querying vectors, such as SQLite's [sqlite-vec](https://github.com/asg017/sqlite-vec)
|
|
778
|
+
or PostgreSQL's [pg-vector](https://github.com/pgvector/pgvector).
|
|
779
|
+
|
|
780
|
+
llm.rb also includes support for OpenAI's vector store API. It
|
|
781
|
+
provides a vector database as a HTTP service but we won't cover
|
|
782
|
+
that here.
|
|
783
|
+
|
|
784
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/embeddings) to learn more.
|
|
785
|
+
|
|
786
|
+
```ruby
|
|
787
|
+
require "llm"
|
|
788
|
+
|
|
789
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
790
|
+
body = "llm.rb is Ruby's capable AI runtime."
|
|
791
|
+
embedding = llm.embed([body]).embeddings.first
|
|
792
|
+
|
|
793
|
+
# Document is your ActiveRecord or Sequel model
|
|
794
|
+
# with a vector column (e.g. sqlite-vec or pgvector)
|
|
795
|
+
Document.create!(
|
|
796
|
+
title: "llm.rb",
|
|
797
|
+
body:,
|
|
798
|
+
embedding:,
|
|
799
|
+
)
|
|
652
800
|
```
|
|
653
801
|
|
|
654
|
-
|
|
802
|
+
### Images
|
|
655
803
|
|
|
656
804
|
A handful of providers can generate images from a text prompt.
|
|
657
805
|
OpenAI, Google, xAI, and DeepInfra all support it. The API is
|
|
@@ -692,46 +840,16 @@ IO.copy_stream res.images[0], "rocket-with-dog.svg"
|
|
|
692
840
|
## FAQ
|
|
693
841
|
|
|
694
842
|
<details>
|
|
695
|
-
<summary>What
|
|
843
|
+
<summary>What about local LLM support?</summary>
|
|
696
844
|
<br>
|
|
697
845
|
<p>
|
|
698
|
-
|
|
699
|
-
|
|
700
|
-
|
|
701
|
-
|
|
702
|
-
In no particular order:
|
|
703
|
-
|
|
704
|
-
🇺🇸 OpenAI <br>
|
|
705
|
-
🇺🇸 DeepInfra <br>
|
|
706
|
-
🇺🇸 xAI <br>
|
|
707
|
-
🇺🇸 Google (Gemini) <br>
|
|
708
|
-
🇺🇸 AWS bedrock <br>
|
|
709
|
-
🇺🇸 Anthropic <br>
|
|
710
|
-
🇨🇳 DeepSeek <br>
|
|
711
|
-
🇨🇳 zAI <br>
|
|
712
|
-
🇨🇳 Moonshot AI (Kimi) <br>
|
|
713
|
-
🇪🇺 Mistral <br>
|
|
714
|
-
|
|
715
|
-
**Weights**
|
|
716
|
-
|
|
717
|
-
The following providers provide access to open-weight models. <br>
|
|
718
|
-
In no particular order:
|
|
719
|
-
|
|
720
|
-
🇺🇸 DeepInfra <br>
|
|
721
|
-
🇺🇸 AWS bedrock <br>
|
|
722
|
-
🇨🇳 DeepSeek <br>
|
|
723
|
-
🇨🇳 zAI <br>
|
|
724
|
-
🇨🇳 Moonshot AI (Kimi) <br>
|
|
725
|
-
🇪🇺 Mistral <br>
|
|
726
|
-
|
|
727
|
-
**Local**
|
|
728
|
-
|
|
729
|
-
The following providers can be run locally on your own hardware. <br>
|
|
730
|
-
In no particular order:
|
|
846
|
+
The following providers can be run used with models that
|
|
847
|
+
are running on your own hardware. They're reasonably well
|
|
848
|
+
tested but not my main driver:
|
|
849
|
+
</p>
|
|
731
850
|
|
|
732
851
|
* Ollama
|
|
733
852
|
* Llamacpp
|
|
734
|
-
</p>
|
|
735
853
|
</details>
|
|
736
854
|
|
|
737
855
|
<details>
|
|
@@ -763,11 +881,10 @@ type.
|
|
|
763
881
|
If you're on a budget, DeepSeek is hard to beat.
|
|
764
882
|
</details>
|
|
765
883
|
<details>
|
|
766
|
-
<summary>
|
|
767
|
-
<br>
|
|
768
|
-
Yes.
|
|
884
|
+
<summary>Sources other than GitHub?</summary>
|
|
769
885
|
<br>
|
|
770
|
-
|
|
886
|
+
<p>
|
|
887
|
+
We are on the <a href="https://radicle.network">radicle.network</a> as well.
|
|
771
888
|
<br>
|
|
772
889
|
Every commit that lands on GitHub also lands on Radicle.
|
|
773
890
|
<br>
|
|
@@ -776,6 +893,30 @@ Our repository ID is z2PtfQ6dYwyYaW2aGrztG1sMyDmCE.
|
|
|
776
893
|
Browse on <a
|
|
777
894
|
href="https://radicle.network/nodes/iris.radicle.network/z2PtfQ6dYwyYaW2aGrztG1sMyDmCE">the
|
|
778
895
|
web</a>.
|
|
896
|
+
</p>
|
|
897
|
+
</details>
|
|
898
|
+
|
|
899
|
+
<details>
|
|
900
|
+
<summary>Who maintains llm.rb?</summary>
|
|
901
|
+
<br>
|
|
902
|
+
|
|
903
|
+
The llm.rb project is maintained primarily by one
|
|
904
|
+
person. llm.rb has been in active development for more
|
|
905
|
+
than three years and over that time multiple other
|
|
906
|
+
contributors have contributed to llm.rb as well. New
|
|
907
|
+
contributors are always welcome.
|
|
908
|
+
|
|
909
|
+
I use the repl that is distributed with llm.rb to build
|
|
910
|
+
llm.rb itself so there is a healthy feedback loop and
|
|
911
|
+
llm.rb has also been battle tested in production
|
|
912
|
+
environments.
|
|
913
|
+
|
|
914
|
+
I have also also written llm.rb agents within the
|
|
915
|
+
repository that help me maintain the documentation,
|
|
916
|
+
and backport changes to the mruby-llm runtime as well.
|
|
917
|
+
|
|
918
|
+
I am constantly focused on improving llm.rb by using
|
|
919
|
+
it as my primary driver for development.
|
|
779
920
|
</details>
|
|
780
921
|
|
|
781
922
|
## Resources
|
|
@@ -786,41 +927,6 @@ wasn't possible to cover every feature without the README becoming a small book.
|
|
|
786
927
|
The [r.uby.dev](https://r.uby.dev) homepage also includes more learning material
|
|
787
928
|
and resources.
|
|
788
929
|
|
|
789
|
-
## Developers
|
|
790
|
-
|
|
791
|
-
The llm.rb project is quite large and maintained primarily by one
|
|
792
|
-
person. It would be near impossible for me to maintain both the codebase
|
|
793
|
-
and its documentation, especially the [deepdive.md](https://r.uby.dev/llm/deepdive/)
|
|
794
|
-
so I have written agents that maintain the documentation assets and that
|
|
795
|
-
allows me to put more focus on the code.
|
|
796
|
-
|
|
797
|
-
The following agents are available for those tasks, and all of them
|
|
798
|
-
use the most cost effective option: DeepSeek. Feel free to use them
|
|
799
|
-
in your own fork.
|
|
800
|
-
|
|
801
|
-
```sh
|
|
802
|
-
##
|
|
803
|
-
# Maintains the deepdive and API docs
|
|
804
|
-
rake agents:scribe:yardoc
|
|
805
|
-
rake agents:scribe:coverage
|
|
806
|
-
rake agents:scribe:regressions
|
|
807
|
-
rake agents:scribe:style
|
|
808
|
-
|
|
809
|
-
##
|
|
810
|
-
# Maintains the release
|
|
811
|
-
rake agents:dexter:changelog
|
|
812
|
-
rake agents:dexter:release
|
|
813
|
-
|
|
814
|
-
##
|
|
815
|
-
# Maintains mruby-llm backports
|
|
816
|
-
rake agents:mruby:research
|
|
817
|
-
rake agents:mruby:implement
|
|
818
|
-
|
|
819
|
-
##
|
|
820
|
-
# Refresh the data/ registry
|
|
821
|
-
rake models.dev:download
|
|
822
|
-
```
|
|
823
|
-
|
|
824
930
|
## License
|
|
825
931
|
|
|
826
932
|
This software is released under the terms of the MIT license. <br>
|