llm.rb 13.1.0 β 15.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +716 -1826
- data/README.md +668 -253
- data/bin/llm.rb +156 -56
- data/data/alibaba.json +1999 -0
- data/data/anthropic.json +195 -252
- data/data/bedrock.json +2189 -1854
- data/data/deepinfra.json +1312 -728
- data/data/deepseek.json +42 -39
- data/data/google.json +1024 -404
- data/data/mistral.json +481 -401
- data/data/moonshot.json +384 -0
- data/data/openai.json +985 -1354
- data/data/xai.json +220 -115
- data/data/zai.json +166 -166
- data/docs/deepdive/advanced/cancellation.md +74 -0
- data/docs/deepdive/advanced/compaction.md +85 -0
- data/docs/deepdive/advanced/context.md +280 -0
- data/docs/deepdive/advanced/guard.md +371 -0
- data/docs/deepdive/advanced/transformer.md +67 -0
- data/docs/deepdive/advanced/transports.md +45 -0
- data/docs/deepdive/features/builtin_tools.md +191 -0
- data/docs/deepdive/features/concurrency.md +110 -0
- data/docs/deepdive/features/database.md +449 -0
- data/docs/deepdive/features/embeddings.md +157 -0
- data/docs/deepdive/features/repl.md +141 -0
- data/docs/deepdive/fundamentals/agents.md +253 -0
- data/docs/deepdive/fundamentals/providers.md +159 -0
- data/docs/deepdive/fundamentals/schema.md +61 -0
- data/docs/deepdive/fundamentals/skills.md +111 -0
- data/docs/deepdive/fundamentals/stream.md +143 -0
- data/docs/deepdive/fundamentals/tools.md +329 -0
- data/docs/deepdive/media/audio.md +122 -0
- data/docs/deepdive/media/images.md +89 -0
- data/docs/deepdive/media/ocr.md +48 -0
- data/docs/deepdive/protocols/a2a.md +106 -0
- data/docs/deepdive/protocols/mcp.md +111 -0
- data/docs/deepdive/reference/cost.md +109 -0
- data/docs/deepdive/reference/model_registry.md +271 -0
- data/docs/deepdive/reference/object.md +108 -0
- data/docs/deepdive/reference/tracer.md +187 -0
- data/{resources β docs}/deepdive.md +37 -23
- data/lib/llm/a2a/transport/http.rb +1 -1
- data/lib/llm/active_record/acts_as_llm.rb +19 -5
- data/lib/llm/agent.rb +107 -17
- data/lib/llm/context.rb +164 -134
- data/lib/llm/cost.rb +114 -49
- data/lib/llm/error.rb +7 -8
- data/lib/llm/function/array.rb +1 -1
- data/lib/llm/function/async/task.rb +2 -0
- data/lib/llm/function/fiber/task.rb +2 -0
- data/lib/llm/function/fork/task.rb +16 -1
- data/lib/llm/function/ractor/task.rb +2 -0
- data/lib/llm/function/sequential/group.rb +20 -10
- data/lib/llm/function/sequential/task.rb +2 -9
- data/lib/llm/function/task.rb +4 -0
- data/lib/llm/function/thread/task.rb +2 -0
- data/lib/llm/function.rb +33 -4
- data/lib/llm/guard/loop.rb +89 -0
- data/lib/llm/guard/null.rb +19 -0
- data/lib/llm/guard.rb +61 -0
- data/lib/llm/message.rb +5 -4
- data/lib/llm/provider.rb +43 -0
- data/lib/llm/providers/alibaba/error_handler.rb +34 -0
- data/lib/llm/providers/alibaba/request_adapter.rb +13 -0
- data/lib/llm/providers/alibaba.rb +93 -0
- data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
- data/lib/llm/providers/anthropic.rb +1 -9
- data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
- data/lib/llm/providers/bedrock.rb +9 -9
- data/lib/llm/providers/deepseek/request_adapter.rb +2 -33
- data/lib/llm/providers/google/stream_parser.rb +1 -0
- data/lib/llm/providers/google.rb +1 -9
- data/lib/llm/providers/moonshot.rb +76 -0
- data/lib/llm/providers/ollama.rb +1 -9
- data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
- data/lib/llm/providers/openai/responses.rb +6 -9
- data/lib/llm/providers/openai/schema.rb +37 -0
- data/lib/llm/providers/openai/stream_parser.rb +1 -0
- data/lib/llm/providers/openai.rb +4 -12
- data/lib/llm/registry/model.rb +186 -0
- data/lib/llm/registry.rb +45 -14
- data/lib/llm/repl/bar.rb +11 -12
- data/lib/llm/repl/buffer.rb +43 -16
- data/lib/llm/repl/color.rb +85 -0
- data/lib/llm/repl/command.rb +12 -0
- data/lib/llm/repl/commands/model.rb +39 -0
- data/lib/llm/repl/input/cache.rb +45 -0
- data/lib/llm/repl/input/char.rb +46 -0
- data/lib/llm/repl/input/row.rb +39 -0
- data/lib/llm/repl/input.rb +327 -79
- data/lib/llm/repl/markdown/table.rb +6 -2
- data/lib/llm/repl/markdown.rb +56 -6
- data/lib/llm/repl/node.rb +7 -0
- data/lib/llm/repl/status.rb +54 -5
- data/lib/llm/repl/stream.rb +16 -4
- data/lib/llm/repl/walker.rb +3 -2
- data/lib/llm/repl/window.rb +111 -11
- data/lib/llm/repl.rb +47 -21
- data/lib/llm/sequel/plugin.rb +19 -5
- data/lib/llm/skill.rb +21 -8
- data/lib/llm/stream.rb +35 -7
- data/lib/llm/tool.rb +27 -0
- data/lib/llm/tools/rg.rb +2 -1
- data/lib/llm/transformer/null.rb +21 -0
- data/lib/llm/transformer.rb +55 -0
- data/lib/llm/transport/curb.rb +23 -3
- data/lib/llm/usage.rb +155 -9
- data/lib/llm/version.rb +1 -1
- data/lib/llm.rb +121 -29
- data/llm.gemspec +16 -10
- metadata +100 -13
- data/lib/llm/loop_guard.rb +0 -107
data/README.md
CHANGED
|
@@ -14,32 +14,17 @@
|
|
|
14
14
|
|
|
15
15
|
Welcome to the canonical llm.rb repository.
|
|
16
16
|
|
|
17
|
-
llm.rb is an advanced runtime for building
|
|
18
|
-
on CRuby.
|
|
19
|
-
|
|
17
|
+
llm.rb is an advanced runtime for building agentic AI applications
|
|
18
|
+
on CRuby. It has zero runtime dependencies by default, it supports
|
|
19
|
+
concurrent and parallel tool execution and has a single coherent API
|
|
20
|
+
that spans 13+ providers. Streaming, tools, guards, compaction, the
|
|
21
|
+
REPL, builtin MCP/A2A support and the database integrations all build
|
|
22
|
+
on the same three concepts: providers, contexts, and agents.
|
|
23
|
+
|
|
24
|
+
Once you learn the fundamentals, everything else falls into place
|
|
25
|
+
naturally. Some features, such as ActiveRecord support, require
|
|
20
26
|
optional dependencies that are opt-in.
|
|
21
27
|
|
|
22
|
-
When you want to learn more than what the README covers, checkout
|
|
23
|
-
the [deepdive.md](https://r.uby.dev/llm/deepdive/).
|
|
24
|
-
|
|
25
|
-
## Features
|
|
26
|
-
|
|
27
|
-
The runtime supports OpenAI, OpenAI-compatible endpoints, Anthropic, Google
|
|
28
|
-
Gemini, Mistral, DeepSeek, DeepInfra, xAI, Z.ai, AWS Bedrock, Ollama, and llama.cpp.
|
|
29
|
-
It has first-class support for streaming, tool calls, MCP
|
|
30
|
-
and A2A, embeddings, vector stores, OCR, context compaction,
|
|
31
|
-
and the RAG pattern.
|
|
32
|
-
|
|
33
|
-
There are multiple HTTP backends to choose from, tools can be run concurrently
|
|
34
|
-
or in parallel via threads, async tasks, fibers, ractors, and fork, and it is
|
|
35
|
-
also possible to make a tool call while the model is still streaming.
|
|
36
|
-
|
|
37
|
-
The runtime builds on top of three core concepts: providers, contexts, and agents,
|
|
38
|
-
so once you learn the fundamentals, everything else falls into place naturally. And once
|
|
39
|
-
you learn llm.rb, you will also be able to use
|
|
40
|
-
<a href="https://r.uby.dev/mruby-llm">mruby-llm</a> and
|
|
41
|
-
<a href="https://r.uby.dev/wasm-llm">wasm-llm</a> because the API is pretty much identical.
|
|
42
|
-
|
|
43
28
|
## Install
|
|
44
29
|
|
|
45
30
|
```bash
|
|
@@ -48,90 +33,97 @@ gem install llm.rb
|
|
|
48
33
|
|
|
49
34
|
## Quick start
|
|
50
35
|
|
|
51
|
-
|
|
36
|
+
### Agents
|
|
52
37
|
|
|
53
38
|
The
|
|
54
39
|
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
55
40
|
class is the default high-level interface,
|
|
56
41
|
and it is recommended for most use-cases. It manages tool execution
|
|
57
|
-
automatically
|
|
58
|
-
|
|
42
|
+
automatically and
|
|
43
|
+
[guards against infinite loops](https://r.uby.dev/llm/deepdive/advanced/guard),
|
|
44
|
+
manages conversation state, and much more.
|
|
59
45
|
|
|
60
46
|
```ruby
|
|
61
47
|
require "llm"
|
|
62
48
|
|
|
63
49
|
llm = LLM.deepseek(key: ENV["KEY"])
|
|
64
50
|
agent = LLM::Agent.new(llm, stream: $stdout)
|
|
65
|
-
agent.talk "
|
|
51
|
+
agent.talk "hello world"
|
|
66
52
|
```
|
|
53
|
+
<details>
|
|
54
|
+
<summary>Stream</summary>
|
|
55
|
+
<br>
|
|
67
56
|
|
|
68
|
-
|
|
57
|
+
Streams can be simple IO objects or subclasses of
|
|
58
|
+
[`LLM::Stream`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
|
|
59
|
+
with structured callbacks for content,
|
|
60
|
+
reasoning, tool calls, tool returns, and compaction.
|
|
61
|
+
Streams can also observe message transformers, which rewrite
|
|
62
|
+
outgoing messages before they reach the provider.
|
|
69
63
|
|
|
70
|
-
[
|
|
71
|
-
|
|
72
|
-
corresponding class accessor: `name`, `description`, `model`, `tools`,
|
|
73
|
-
`instructions`, `schema`, `stream`, `tracer`, `concurrency`, `confirm`,
|
|
74
|
-
`path`, and `skills`. All options are optional; zero or more can be set.
|
|
75
|
-
An error is raised for unknown keys so that typos are caught early.
|
|
64
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/stream/)
|
|
65
|
+
to learn more.
|
|
76
66
|
|
|
77
67
|
```ruby
|
|
78
|
-
class
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
end
|
|
68
|
+
class MyStream < LLM::Stream
|
|
69
|
+
# Visible assistant output.
|
|
70
|
+
def on_content(content)
|
|
71
|
+
print content
|
|
72
|
+
end
|
|
84
73
|
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
74
|
+
# Reasoning output streamed separately from visible content.
|
|
75
|
+
def on_reasoning_content(content)
|
|
76
|
+
warn content
|
|
77
|
+
end
|
|
89
78
|
|
|
90
|
-
|
|
79
|
+
# A streamed tool call has been fully parsed.
|
|
80
|
+
def on_tool_call(tool)
|
|
81
|
+
end
|
|
91
82
|
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
code. For database-backed persistence, ActiveRecord and Sequel
|
|
96
|
-
integrations are also available (see the
|
|
97
|
-
[database deepdive](https://r.uby.dev/llm/deepdive/advanced/database)
|
|
98
|
-
for details). All persistence options use the same underlying
|
|
99
|
-
serialization.
|
|
83
|
+
# Queued streamed tool work has returned.
|
|
84
|
+
def on_tool_return(tool, result)
|
|
85
|
+
end
|
|
100
86
|
|
|
101
|
-
|
|
102
|
-
|
|
87
|
+
# Before a transformer rewrites an outgoing message.
|
|
88
|
+
def on_transform(transformer)
|
|
89
|
+
end
|
|
103
90
|
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
91
|
+
# Aftter a transformer rewrites an outgoing message.
|
|
92
|
+
def on_transform_finish(transformer)
|
|
93
|
+
end
|
|
107
94
|
|
|
108
|
-
#
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
```
|
|
95
|
+
# Before a compactor trims the conversation.
|
|
96
|
+
def on_compaction(compactor)
|
|
97
|
+
end
|
|
112
98
|
|
|
113
|
-
|
|
99
|
+
# After a compactor trims the conversation.
|
|
100
|
+
def on_compaction_finish(compactor)
|
|
101
|
+
end
|
|
114
102
|
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
and it is what
|
|
119
|
-
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
120
|
-
uses under the hood.
|
|
121
|
-
It requires that the tool call loop be managed manually -
|
|
122
|
-
sometimes that can be useful, but usually for advanced use-cases.
|
|
123
|
-
If you're new to llm.rb, try
|
|
124
|
-
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html) first.
|
|
103
|
+
# Before a skill's subagent runs.
|
|
104
|
+
def on_skill_call(skill)
|
|
105
|
+
end
|
|
125
106
|
|
|
126
|
-
|
|
127
|
-
|
|
107
|
+
# After a skill's subagent runs.
|
|
108
|
+
# The subagent that ran it, the skill, and its response are passed
|
|
109
|
+
# through, so you can introspect the agent, tally skill usage, or
|
|
110
|
+
# track costs.
|
|
111
|
+
def on_skill_return(agent, skill, result)
|
|
112
|
+
end
|
|
113
|
+
|
|
114
|
+
# A request was rate limited and will be retried.
|
|
115
|
+
def on_rate_limit(error)
|
|
116
|
+
end
|
|
117
|
+
end
|
|
128
118
|
|
|
129
119
|
llm = LLM.deepseek(key: ENV["KEY"])
|
|
130
|
-
|
|
131
|
-
|
|
120
|
+
agent = LLM::Agent.new(llm, stream: MyStream.new)
|
|
121
|
+
agent.talk "Explain Ruby fibers."
|
|
132
122
|
```
|
|
123
|
+
</details>
|
|
133
124
|
|
|
134
|
-
|
|
125
|
+
<details><summary>Tools</summary>
|
|
126
|
+
<br>
|
|
135
127
|
|
|
136
128
|
Subclasses of
|
|
137
129
|
[`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
|
|
@@ -140,6 +132,9 @@ an optional set of typed parameters. <br> The model can choose to
|
|
|
140
132
|
call them on your behalf, and they're one of the most powerful features
|
|
141
133
|
for extending the feature set or abilities of a model.
|
|
142
134
|
|
|
135
|
+
The runtime also ships with a catalog of built-in tools for
|
|
136
|
+
filesystem, search, and shell operations. <br> See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/tools) to learn more.
|
|
137
|
+
|
|
143
138
|
```ruby
|
|
144
139
|
class ReadFile < LLM::Tool
|
|
145
140
|
name "read-file"
|
|
@@ -151,78 +146,145 @@ class ReadFile < LLM::Tool
|
|
|
151
146
|
{contents: File.read(path)}
|
|
152
147
|
end
|
|
153
148
|
end
|
|
149
|
+
|
|
150
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
151
|
+
agent = LLM::Agent.new(llm, tools: [ReadFile], stream: $stdout)
|
|
152
|
+
agent.talk "summarize README.md"
|
|
154
153
|
```
|
|
154
|
+
</details>
|
|
155
|
+
<details>
|
|
156
|
+
<summary>Skills</summary>
|
|
157
|
+
<br>
|
|
155
158
|
|
|
156
|
-
|
|
159
|
+
A skill turns a markdown file into a callable tool. When the model
|
|
160
|
+
calls it, the runtime spawns a subagent with the skill's instructions
|
|
161
|
+
as its system prompt and the skill's own tool set. The subagent runs
|
|
162
|
+
one turn and returns the result, then is discarded. Each call
|
|
163
|
+
is fresh and stateless.
|
|
157
164
|
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
165
|
+
A [LLM::Stream](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
|
|
166
|
+
can be notified as a skill starts and when it returns. The `on_skill_return`
|
|
167
|
+
callback hands back the subagent that ran the skill, so you can inspect
|
|
168
|
+
its conversation, measure its usage, track costs or add a verification
|
|
169
|
+
step (eg `subagent.talk("verify your work")`).
|
|
162
170
|
|
|
163
|
-
|
|
164
|
-
class MyStream < LLM::Stream
|
|
165
|
-
def on_content(content)
|
|
166
|
-
print content
|
|
167
|
-
end
|
|
171
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/skills) to learn more.
|
|
168
172
|
|
|
169
|
-
|
|
170
|
-
warn content
|
|
171
|
-
end
|
|
172
|
-
end
|
|
173
|
+
##### summary.md
|
|
173
174
|
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
175
|
+
```markdown
|
|
176
|
+
---
|
|
177
|
+
name: summary
|
|
178
|
+
description: Reads recent git history and writes a summary
|
|
179
|
+
tools: all
|
|
180
|
+
---
|
|
181
|
+
|
|
182
|
+
Collect the recent git log, analyze each commit,
|
|
183
|
+
and write a summary to summary.txt.
|
|
177
184
|
```
|
|
178
185
|
|
|
179
|
-
|
|
186
|
+
##### agent.rb
|
|
180
187
|
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
output from any model call. Pass a schema to `LLM::Context#talk`,
|
|
184
|
-
`LLM::Agent#talk`, or `LLM::Provider#complete` to receive validated
|
|
185
|
-
JSON instead of free text. Schemas work alongside tools and streams.
|
|
188
|
+
```ruby
|
|
189
|
+
require "llm"
|
|
186
190
|
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
parameters.
|
|
191
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
192
|
+
agent = LLM::Agent.new(llm, skills: ["summary.md"])
|
|
193
|
+
agent.talk "Summarize the last week of work"
|
|
194
|
+
```
|
|
195
|
+
</details>
|
|
193
196
|
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
197
|
+
<details>
|
|
198
|
+
<summary>Concurrency</summary>
|
|
199
|
+
<br>
|
|
200
|
+
|
|
201
|
+
The runtime supports six different concurrency strategies that have
|
|
202
|
+
different attributes. The choice between all of them often depends
|
|
203
|
+
on the requirements of your application.
|
|
204
|
+
|
|
205
|
+
IO-bound tools are a good fit for the `:async`, `:thread`,
|
|
206
|
+
and `:fiber` strategies while true parallelism can be achieved
|
|
207
|
+
with the `:fork` and `:ractor` strategies. The
|
|
208
|
+
`:sequential` strategy runs tools one at a time and is the default.
|
|
209
|
+
The `:fork` strategy also provides a separate process that offers
|
|
210
|
+
isolation from its parent.
|
|
211
|
+
|
|
212
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/concurrency) to learn more.
|
|
201
213
|
|
|
202
214
|
```ruby
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
property :temperature, Float, "Current temperature"
|
|
206
|
-
property :conditions, String, "Weather conditions"
|
|
207
|
-
required %i[city temperature conditions]
|
|
208
|
-
end
|
|
215
|
+
require "llm"
|
|
216
|
+
require "llm/tools"
|
|
209
217
|
|
|
210
|
-
llm
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
218
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
219
|
+
tools = LLM::Tool.subclasses
|
|
220
|
+
agent = LLM::Agent.new(llm, tools:, concurrency: :fork)
|
|
221
|
+
agent.talk "Run the tools in parallel"
|
|
214
222
|
```
|
|
215
223
|
|
|
216
|
-
|
|
224
|
+
</details>
|
|
225
|
+
<details>
|
|
226
|
+
<summary>Cancellation</summary>
|
|
227
|
+
<br>
|
|
228
|
+
|
|
229
|
+
Abort a request mid-stream and interrupt any running tools with
|
|
230
|
+
[`LLM::Agent#interrupt!`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#interrupt!)
|
|
231
|
+
(or `cancel!`), from any thread. The runtime raises
|
|
232
|
+
[`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html)
|
|
233
|
+
on the caller and on every active tool. A forked tool gets interrupted over
|
|
234
|
+
the control channel, a ractor via message passing, and pending tools
|
|
235
|
+
are stopped before they run. The in-flight HTTP request is closed
|
|
236
|
+
too, so a turn you no longer want stops without burning tokens.
|
|
237
|
+
|
|
238
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/cancellation) to learn more.
|
|
239
|
+
|
|
240
|
+
```ruby
|
|
241
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
242
|
+
agent = LLM::Agent.new(llm)
|
|
243
|
+
Thread.new { sleep(1); agent.cancel! }
|
|
244
|
+
|
|
245
|
+
begin
|
|
246
|
+
agent.talk "write a very long poem", stream: $stdout
|
|
247
|
+
rescue LLM::Interrupt
|
|
248
|
+
puts "cancelled"
|
|
249
|
+
end
|
|
250
|
+
```
|
|
251
|
+
</details>
|
|
252
|
+
<details>
|
|
253
|
+
<summary>Console (<code>binding.pry</code> for agents)</summary>
|
|
254
|
+
<br>
|
|
217
255
|
|
|
218
256
|
The [LLM::Agent#repl](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#repl-instance_method)
|
|
219
|
-
method drops you into a
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
257
|
+
method drops you into a highly capable read-eval-print loop (REPL)
|
|
258
|
+
that is built on top of curses. It can help you debug agents,
|
|
259
|
+
test your tools, connect to MCP servers, and even A2A agents.
|
|
260
|
+
The REPL stands out because it connects to the surrounding
|
|
261
|
+
runtime and it can be extended by your code. Think of it as
|
|
262
|
+
`binding.pry` but for agents.
|
|
263
|
+
|
|
264
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/repl) to learn more.
|
|
265
|
+
|
|
266
|
+
##### Demo
|
|
267
|
+
|
|
268
|
+
[Watch in high quality on asciinema](https://asciinema.org/a/OsS8wwaasKasoDDz)
|
|
269
|
+
|
|
270
|
+

|
|
271
|
+
|
|
272
|
+
|
|
273
|
+
##### Installation
|
|
274
|
+
|
|
275
|
+
The REPL is distributed with llm.rb so you don't have to install
|
|
276
|
+
a separate gem but it requires a number of optional dependencies
|
|
277
|
+
to be installed separately. The following gems provide the full
|
|
278
|
+
experience:
|
|
279
|
+
|
|
280
|
+
gem install curses kramdown xchan.rb test-cmd.rb
|
|
281
|
+
|
|
282
|
+
##### Persistence
|
|
283
|
+
|
|
284
|
+
The `path:` option can be set on an agent for automatic persistence
|
|
285
|
+
across REPL sessions. The `tools:` option attaches extra tools
|
|
286
|
+
for the duration of the session. Recall previous turns with Ctrl+P and
|
|
287
|
+
Ctrl+N.
|
|
226
288
|
|
|
227
289
|
```ruby
|
|
228
290
|
require "llm"
|
|
@@ -236,20 +298,110 @@ agent.repl(tools: LLM::Tool.subclasses)
|
|
|
236
298
|
##### CLI
|
|
237
299
|
|
|
238
300
|
The `llm.rb` executable is available on your PATH after installation.
|
|
239
|
-
It starts a REPL session from any directory
|
|
301
|
+
It starts a REPL session from any directory.The CLI auto-detects your
|
|
302
|
+
provider from standard environment variables (`DEEPSEEK_API_KEY`,
|
|
303
|
+
`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, etc.). Persistent sessions are
|
|
304
|
+
stored under `~/.llm.rb/` and restored automatically on your next visit.
|
|
240
305
|
|
|
241
306
|
```bash
|
|
242
307
|
llm.rb # auto-detect from $DEEPSEEK_API_KEY
|
|
243
308
|
llm.rb -p openai # use OpenAI explicitly
|
|
244
309
|
llm.rb -t # temporary session, no persistence
|
|
245
310
|
```
|
|
311
|
+
</details>
|
|
312
|
+
<details>
|
|
313
|
+
<summary>Persistence</summary>
|
|
314
|
+
<br>
|
|
315
|
+
|
|
316
|
+
Set `path:` on an agent for automatic filesystem persistence:
|
|
317
|
+
the agent restores conversation history from the file on startup
|
|
318
|
+
and saves it back after every turn, with no manual serialization
|
|
319
|
+
code. For database-backed persistence, ActiveRecord and Sequel
|
|
320
|
+
integrations are also available. All persistence options use the same
|
|
321
|
+
underlying serialization.
|
|
322
|
+
|
|
323
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/database) to learn more.
|
|
324
|
+
|
|
325
|
+
```ruby
|
|
326
|
+
require "llm"
|
|
327
|
+
|
|
328
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
329
|
+
agent = LLM::Agent.new(llm, path: "session.json")
|
|
330
|
+
agent.talk "remember my name is robert"
|
|
331
|
+
|
|
332
|
+
# Next time, the conversation is restored automatically:
|
|
333
|
+
agent = LLM::Agent.new(llm, path: "session.json")
|
|
334
|
+
agent.talk "what's my name?"
|
|
335
|
+
```
|
|
336
|
+
</details>
|
|
337
|
+
<details><summary>ActiveRecord | Sequel</summary>
|
|
338
|
+
<br>
|
|
339
|
+
|
|
340
|
+
Because both
|
|
341
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) and
|
|
342
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
343
|
+
can be serialized to JSON and stored in a simple string, both ActiveRecord
|
|
344
|
+
and Sequel support can be implemented within a single column on a single row.
|
|
345
|
+
|
|
346
|
+
The runtime includes first-class support for both ActiveRecord / Sequel, and
|
|
347
|
+
for both Rack-based / Rails-based applications. On databases
|
|
348
|
+
where it is supported, such as PostgreSQL, the column can be optimized by using
|
|
349
|
+
the `jsonb` type.
|
|
350
|
+
|
|
351
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/database) to learn more.
|
|
352
|
+
|
|
353
|
+
```ruby
|
|
354
|
+
require "active_record"
|
|
355
|
+
require "llm"
|
|
356
|
+
require "llm/active_record"
|
|
357
|
+
|
|
358
|
+
class Email < ApplicationRecord
|
|
359
|
+
acts_as_agent do |agent|
|
|
360
|
+
agent.set name: "mail",
|
|
361
|
+
instructions: "Write concise, friendly replies to emails",
|
|
362
|
+
model: "deepseek-v4-pro"
|
|
363
|
+
end
|
|
364
|
+
|
|
365
|
+
def draft_reply!
|
|
366
|
+
talk("Draft a reply to:\n\n#{body}")
|
|
367
|
+
end
|
|
368
|
+
|
|
369
|
+
def summarize
|
|
370
|
+
talk("Summarize this email thread in a few sentences")
|
|
371
|
+
end
|
|
372
|
+
|
|
373
|
+
private
|
|
374
|
+
|
|
375
|
+
##
|
|
376
|
+
# By convention, this method defines the provider for a model.
|
|
377
|
+
# If necessary, it can be renamed with: provider: :your_method.
|
|
378
|
+
def set_provider
|
|
379
|
+
LLM.deepseek(key: ENV["KEY"])
|
|
380
|
+
end
|
|
381
|
+
|
|
382
|
+
##
|
|
383
|
+
# By convention, this method returns the context options given
|
|
384
|
+
# to LLM::Context or LLM::Agent. This method can be left undefined.
|
|
385
|
+
def set_context
|
|
386
|
+
{}
|
|
387
|
+
end
|
|
388
|
+
end
|
|
246
389
|
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
Persistent sessions are stored under `~/.llm.rb/` and restored
|
|
250
|
-
automatically on your next visit.
|
|
390
|
+
email = Email.create!(subject: "Streaming support", body: "How do I stream responses?")
|
|
391
|
+
email.draft_reply!
|
|
251
392
|
|
|
252
|
-
|
|
393
|
+
##
|
|
394
|
+
# The conversation (the email and the draft
|
|
395
|
+
# reply) is persisted to the email's column. A
|
|
396
|
+
# fresh instance restores it and continues the
|
|
397
|
+
# thread, so the summary below knows what was
|
|
398
|
+
# already drafted:
|
|
399
|
+
Email.find(email.id).summarize
|
|
400
|
+
```
|
|
401
|
+
</details>
|
|
402
|
+
|
|
403
|
+
<details><summary>MCP</summary>
|
|
404
|
+
<br>
|
|
253
405
|
|
|
254
406
|
The Model Context Protocol (MCP) has first-class support
|
|
255
407
|
in llm.rb. The stdio and http transports work out of the
|
|
@@ -259,6 +411,9 @@ used with
|
|
|
259
411
|
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) or
|
|
260
412
|
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
|
|
261
413
|
|
|
414
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/protocols/mcp/), and the
|
|
415
|
+
[deepdive.md on persistent connections](https://r.uby.dev/llm/deepdive/features/transports) to learn more.
|
|
416
|
+
|
|
262
417
|
```ruby
|
|
263
418
|
require "llm"
|
|
264
419
|
|
|
@@ -267,8 +422,9 @@ mcp = LLM::MCP.stdio(argv: ["ruby", "server.rb"])
|
|
|
267
422
|
agent = LLM::Agent.new(llm, stream: $stdout, tools: mcp.tools)
|
|
268
423
|
agent.talk "Run the tool"
|
|
269
424
|
```
|
|
270
|
-
|
|
271
|
-
|
|
425
|
+
</details>
|
|
426
|
+
<details><summary>A2A</summary>
|
|
427
|
+
<br>
|
|
272
428
|
|
|
273
429
|
The Agent 2 Agent (A2A) protocol has first-class support
|
|
274
430
|
in llm.rb. The http and jsonrpc transports work out of the
|
|
@@ -278,6 +434,9 @@ used with
|
|
|
278
434
|
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) or
|
|
279
435
|
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
|
|
280
436
|
|
|
437
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/protocols/a2a/), and the
|
|
438
|
+
[deepdive.md on persistent connections](https://r.uby.dev/llm/deepdive/features/transports) to learn more.
|
|
439
|
+
|
|
281
440
|
```ruby
|
|
282
441
|
require "llm"
|
|
283
442
|
|
|
@@ -286,40 +445,330 @@ a2a = LLM::A2A.rest(url: "https://remote-agent.example.com")
|
|
|
286
445
|
agent = LLM::Agent.new(llm, stream: $stdout, tools: a2a.skills)
|
|
287
446
|
agent.talk "Run the skill"
|
|
288
447
|
```
|
|
448
|
+
</details>
|
|
289
449
|
|
|
290
|
-
|
|
450
|
+
<details><summary>Structured outputs</summary>
|
|
451
|
+
<br>
|
|
291
452
|
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
453
|
+
[`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
|
|
454
|
+
subclasses produce typed, structured
|
|
455
|
+
output from any model call. Pass a schema to
|
|
456
|
+
[`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk-instance_method),
|
|
457
|
+
[`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk-instance_method),
|
|
458
|
+
or
|
|
459
|
+
[`LLM::Provider#complete`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#complete-instance_method)
|
|
460
|
+
to receive validated JSON instead of free text. Schemas work alongside tools and streams.
|
|
298
461
|
|
|
299
|
-
|
|
462
|
+
[`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
|
|
463
|
+
can define objects, arrays, enums, nested schemas,
|
|
464
|
+
and more. It is also used internally by
|
|
465
|
+
[`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html) for parameter
|
|
466
|
+
definitions, so you already benefit from it when you declare tool
|
|
467
|
+
parameters.
|
|
300
468
|
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
|
|
469
|
+
The
|
|
470
|
+
[`LLM::DeepSeek`](https://r.uby.dev/api-docs/llm.rb/LLM/DeepSeek.html)
|
|
471
|
+
provider includes runtime-level optimisations such as structured
|
|
472
|
+
output support (despite no official structured outputs API) and
|
|
473
|
+
SVG image generation. This example uses
|
|
474
|
+
[`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html) with
|
|
475
|
+
DeepSeek:
|
|
307
476
|
|
|
308
|
-
|
|
309
|
-
|
|
477
|
+
```ruby
|
|
478
|
+
class Weather < LLM::Schema
|
|
479
|
+
property :city, String, "The city name"
|
|
480
|
+
property :temperature, Number, "Current temperature"
|
|
481
|
+
property :conditions, String, "Weather conditions"
|
|
482
|
+
required %i[city temperature conditions]
|
|
483
|
+
end
|
|
484
|
+
|
|
485
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
486
|
+
agent = LLM::Agent.new(llm, schema: Weather)
|
|
487
|
+
res = agent.talk "Weather in Paris?"
|
|
488
|
+
res.content! # => {city: "Paris", temperature: 15.0, conditions: "Cloudy"}
|
|
310
489
|
```
|
|
490
|
+
</details>
|
|
491
|
+
<details><summary>Guards</summary>
|
|
492
|
+
<br>
|
|
311
493
|
|
|
312
|
-
|
|
494
|
+
[`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
|
|
495
|
+
is the hook that sees every tool call before it runs. A guard
|
|
496
|
+
can let a call through, cancel it, block it with an error, or
|
|
497
|
+
even answer for it. Because it runs before the tool, anything
|
|
498
|
+
it intercepts never executes. Policy, validation, quotas, and
|
|
499
|
+
cost ceilings all live here.
|
|
500
|
+
|
|
501
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
502
|
+
enables
|
|
503
|
+
[`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
|
|
504
|
+
by default, so agents get loop protection out of the box. To
|
|
505
|
+
write your own guard, subclass
|
|
506
|
+
[`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
|
|
507
|
+
and implement
|
|
508
|
+
[`LLM::Guard#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html#call-instance_method).
|
|
509
|
+
The pending call arrives as `function:`. Return a value to close
|
|
510
|
+
the call, or `nil` to let it run:
|
|
511
|
+
|
|
512
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/guard) to learn more.
|
|
513
|
+
|
|
514
|
+
```ruby
|
|
515
|
+
class PolicyGuard < LLM::Guard
|
|
516
|
+
def call(function:)
|
|
517
|
+
if function.name == "shell"
|
|
518
|
+
function.return(error: true, type: "policy_error",
|
|
519
|
+
message: "shell is disabled")
|
|
520
|
+
end
|
|
521
|
+
end
|
|
522
|
+
end
|
|
523
|
+
|
|
524
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
525
|
+
agent = LLM::Agent.new(llm, tools: [Shell, ReadFile], guard: PolicyGuard)
|
|
526
|
+
```
|
|
527
|
+
</details>
|
|
528
|
+
|
|
529
|
+
<details>
|
|
530
|
+
<summary>Transformers</summary>
|
|
531
|
+
<br>
|
|
532
|
+
|
|
533
|
+
It is possible to rewrite outgoing messages before they reach the provider with
|
|
534
|
+
[`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html).
|
|
535
|
+
Create a subclass and implement `call(message:)` to scrub sensitive data,
|
|
536
|
+
inject context, or normalize content. The transform runs automatically
|
|
537
|
+
on every turn, so you never have to change your prompt code.
|
|
538
|
+
|
|
539
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transformer) to learn more.
|
|
540
|
+
|
|
541
|
+
```ruby
|
|
542
|
+
class RedactEmails < LLM::Transformer
|
|
543
|
+
def call(message:)
|
|
544
|
+
content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
|
|
545
|
+
LLM::Message.new(message.role, content, message.extra)
|
|
546
|
+
end
|
|
547
|
+
end
|
|
548
|
+
|
|
549
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
550
|
+
agent = LLM::Agent.new(llm, transformer: RedactEmails)
|
|
551
|
+
agent.talk "Contact support@example.com for help"
|
|
552
|
+
```
|
|
553
|
+
</details>
|
|
554
|
+
|
|
555
|
+
<details>
|
|
556
|
+
<summary>Compactors</summary>
|
|
557
|
+
<br>
|
|
558
|
+
|
|
559
|
+
Every model has a context window: the finite number of tokens it can
|
|
560
|
+
consider in a single request. Generally a compactor will drop or
|
|
561
|
+
summarize older messages to keep the conversation within that window,
|
|
562
|
+
and it runs automatically before every turn. By default it is disabled
|
|
563
|
+
so it is a feature you must opt into.
|
|
564
|
+
|
|
565
|
+
[`LLM::Compactor::Truncate`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Truncate.html)
|
|
566
|
+
keeps the most recent messages via an integer count or a percentage like
|
|
567
|
+
`"80%"`. It preserves tool call and return pairs so the conversation
|
|
568
|
+
never contains an orphaned result. It is also possible to subclass
|
|
569
|
+
[`LLM::Compactor`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor.html)
|
|
570
|
+
to implement your own compactor with its own logic. Streams can observe the
|
|
571
|
+
process through the
|
|
572
|
+
[`LLM::Stream#on_compaction`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction)
|
|
573
|
+
and
|
|
574
|
+
[`LLM::Stream#on_compaction_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction_finish)
|
|
575
|
+
callbacks.
|
|
576
|
+
|
|
577
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/compaction) to learn more.
|
|
578
|
+
|
|
579
|
+
```ruby
|
|
580
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
581
|
+
agent = LLM::Agent.new(
|
|
582
|
+
llm,
|
|
583
|
+
compactor: LLM::Compactor::Truncate,
|
|
584
|
+
compactor_options: {keep: 64}
|
|
585
|
+
)
|
|
586
|
+
agent.talk "Hello"
|
|
587
|
+
```
|
|
588
|
+
</details>
|
|
589
|
+
|
|
590
|
+
<details>
|
|
591
|
+
<summary>Automatic retries</summary>
|
|
592
|
+
<br>
|
|
593
|
+
|
|
594
|
+
Rate-limited requests are retried automatically by default. Agents
|
|
595
|
+
retry a 429 up to five times with a growing backoff before giving
|
|
596
|
+
up, so most request failures resolve on their own. Set `retry_budget`
|
|
597
|
+
to change the number of retries, or `retry_budget: 0` to disable
|
|
598
|
+
them.
|
|
313
599
|
|
|
314
600
|
```ruby
|
|
315
601
|
require "llm"
|
|
316
602
|
|
|
317
|
-
llm
|
|
318
|
-
agent = LLM::Agent.new(llm,
|
|
319
|
-
agent.talk "
|
|
603
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
604
|
+
agent = LLM::Agent.new(llm, retry_budget: 0)
|
|
605
|
+
agent.talk "Hello"
|
|
606
|
+
```
|
|
607
|
+
|
|
608
|
+
</details>
|
|
609
|
+
|
|
610
|
+
|
|
611
|
+
<details>
|
|
612
|
+
<summary>Observability</summary>
|
|
613
|
+
<br>
|
|
614
|
+
|
|
615
|
+
Trace what an agent is doing by attaching a tracer. Hook into
|
|
616
|
+
requests, tool calls, and other runtime events to debug a
|
|
617
|
+
misbehaving agent, monitor latency, or export spans to an
|
|
618
|
+
observability backend. All built-in tracers share one interface,
|
|
619
|
+
so switching between them means changing a class name:
|
|
620
|
+
|
|
621
|
+
* [`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html): human-readable single-line logs to stderr, ideal during development.
|
|
622
|
+
* [`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html):
|
|
623
|
+
exports spans via OTLP for OpenTelemetry in production.
|
|
624
|
+
* [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html):
|
|
625
|
+
structured JSON to stdout or a file.
|
|
626
|
+
|
|
627
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/reference/tracer) to learn more.
|
|
628
|
+
|
|
629
|
+
```ruby
|
|
630
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
631
|
+
agent = LLM::Agent.new(llm, tracer: LLM::Tracer::PrettyLogger.new(llm))
|
|
632
|
+
agent.talk "Hello"
|
|
320
633
|
```
|
|
634
|
+
</details>
|
|
635
|
+
|
|
636
|
+
<details>
|
|
637
|
+
<summary>As a subclass</summary>
|
|
638
|
+
<br>
|
|
321
639
|
|
|
322
|
-
|
|
640
|
+
[`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
|
|
641
|
+
is a class-level DSL that accepts a Hash of properties. Each key resolves to a
|
|
642
|
+
corresponding class accessor: `name`, `description`, `model`, `tools`,
|
|
643
|
+
`instructions`, `schema`, `stream`, `tracer`, `concurrency`, `confirm`,
|
|
644
|
+
`path`, `skills`, `tool_budget`, and `retry_budget`. All options are
|
|
645
|
+
optional; zero or more can be set.
|
|
646
|
+
An error is raised for unknown keys so that typos are caught early.
|
|
647
|
+
|
|
648
|
+
```ruby
|
|
649
|
+
require "llm"
|
|
650
|
+
require "llm/tools"
|
|
651
|
+
|
|
652
|
+
class Agent < LLM::Agent
|
|
653
|
+
set name: "sysadmin",
|
|
654
|
+
description: "system administration agent",
|
|
655
|
+
model: "deepseek-v4-pro",
|
|
656
|
+
tools: [LLM::Tool::Shell]
|
|
657
|
+
end
|
|
658
|
+
|
|
659
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
660
|
+
agent = Agent.new(llm)
|
|
661
|
+
agent.talk "Run 'date'"
|
|
662
|
+
```
|
|
663
|
+
</details>
|
|
664
|
+
|
|
665
|
+
### Providers
|
|
666
|
+
|
|
667
|
+
Each provider is constructed with a class-level factory method on
|
|
668
|
+
`LLM`, and the resulting instance is passed to
|
|
669
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
670
|
+
or
|
|
671
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html). The
|
|
672
|
+
same API drives every one of them, so switching models is a one-line
|
|
673
|
+
change. See the [deepdive](https://r.uby.dev/llm/deepdive/fundamentals/providers)
|
|
674
|
+
for a full provider reference.
|
|
675
|
+
|
|
676
|
+
#### What providers does llm.rb support?
|
|
677
|
+
|
|
678
|
+
* **Anthropic** (`LLM.anthropic`)
|
|
679
|
+
* **Google** (`LLM.google`)
|
|
680
|
+
* **OpenAI** (`LLM.openai`)
|
|
681
|
+
* **DeepSeek** (`LLM.deepseek`)
|
|
682
|
+
* **DeepInfra** (`LLM.deepinfra`)
|
|
683
|
+
* **xAI** (`LLM.xai`)
|
|
684
|
+
* **Z.ai** (`LLM.zai`)
|
|
685
|
+
* **Moonshot (Kimi)** (`LLM.moonshot`)
|
|
686
|
+
* **Alibaba (Qwen3)** (`LLM.alibaba`, also `LLM.aliyun`)
|
|
687
|
+
* **Mistral** (`LLM.mistral`)
|
|
688
|
+
* **AWS Bedrock** (`LLM.bedrock`)
|
|
689
|
+
* **Ollama** (`LLM.ollama`)
|
|
690
|
+
* **llama.cpp** (`LLM.llamacpp`)
|
|
691
|
+
|
|
692
|
+
<details>
|
|
693
|
+
<summary>Implicit</summary>
|
|
694
|
+
<br>
|
|
695
|
+
|
|
696
|
+
Cloud providers can infer their API key automatically
|
|
697
|
+
from a set of common defaults that are defined by
|
|
698
|
+
the [models.dev](https://models.dev) registry that
|
|
699
|
+
is also distributed with llm.rb.
|
|
700
|
+
|
|
701
|
+
```ruby
|
|
702
|
+
llm = LLM.openai
|
|
703
|
+
llm = LLM.anthropic
|
|
704
|
+
llm = LLM.deepseek
|
|
705
|
+
llm = LLM.alibaba # also: LLM.aliyun
|
|
706
|
+
llm = LLM.moonshot
|
|
707
|
+
llm = LLM.mistral
|
|
708
|
+
```
|
|
709
|
+
</details>
|
|
710
|
+
<details>
|
|
711
|
+
<summary>Explicit</summary>
|
|
712
|
+
<br>
|
|
713
|
+
|
|
714
|
+
The `key` option can also be providied explicitly, and certain
|
|
715
|
+
providers (eg ollama, llamacpp) usually do not require an API
|
|
716
|
+
key at all.
|
|
717
|
+
|
|
718
|
+
```ruby
|
|
719
|
+
llm = LLM.openai(key: ENV["OPENAI_API_KEY"])
|
|
720
|
+
llm = LLM.anthropic(key: ENV["ANTHROPIC_API_KEY"])
|
|
721
|
+
llm = LLM.deepseek(key: ENV["DEEPSEEK_API_KEY"])
|
|
722
|
+
llm = LLM.alibaba(key: ENV["DASHSCOPE_API_KEY"]) # also: LLM.aliyun
|
|
723
|
+
llm = LLM.moonshot(key: ENV["MOONSHOT_API_KEY"])
|
|
724
|
+
llm = LLM.mistral(key: ENV["MISTRAL_API_KEY"])
|
|
725
|
+
```
|
|
726
|
+
</details>
|
|
727
|
+
|
|
728
|
+
<details>
|
|
729
|
+
<summary>Model Registry</summary>
|
|
730
|
+
<br>
|
|
731
|
+
|
|
732
|
+
Each provider ships its model catalog, pricing, limits, and
|
|
733
|
+
modalities with the gem, sourced from [models.dev](https://models.dev).
|
|
734
|
+
Reach it from any provider, context, or agent, enumerate models, or
|
|
735
|
+
sort them by price.
|
|
736
|
+
|
|
737
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/reference/model_registry) to learn more.
|
|
738
|
+
|
|
739
|
+
```ruby
|
|
740
|
+
require "llm"
|
|
741
|
+
|
|
742
|
+
llm = LLM.openai
|
|
743
|
+
registry = llm.registry # => LLM::Provider#registry
|
|
744
|
+
cheapest = registry.models.sort.first # => LLM::Model
|
|
745
|
+
cheapest.id # => "text-embedding-3-small"
|
|
746
|
+
cheapest.context_window # => 8191
|
|
747
|
+
cheapest.structured_output? # => false
|
|
748
|
+
```
|
|
749
|
+
</details>
|
|
750
|
+
|
|
751
|
+
<details>
|
|
752
|
+
<summary>Transports</summary>
|
|
753
|
+
<br>
|
|
754
|
+
|
|
755
|
+
The `transport:` option selects which HTTP library a provider uses for
|
|
756
|
+
network communication. Three backends ship out of the box: `net/http`
|
|
757
|
+
is always available and the default, `net/http/persistent` pools
|
|
758
|
+
connections for many requests to the same host, and `curb` wraps
|
|
759
|
+
libcurl. They share one interface, so switching is a one-word change.
|
|
760
|
+
|
|
761
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transports) to learn more.
|
|
762
|
+
|
|
763
|
+
```ruby
|
|
764
|
+
llm = LLM.deepseek(
|
|
765
|
+
key: ENV["KEY"],
|
|
766
|
+
transport: :net_http_persistent
|
|
767
|
+
)
|
|
768
|
+
```
|
|
769
|
+
</details>
|
|
770
|
+
|
|
771
|
+
### RAG
|
|
323
772
|
|
|
324
773
|
Most providers offer an embedding model that can be
|
|
325
774
|
used for semantic search, or similarity search. An
|
|
@@ -332,6 +781,8 @@ llm.rb also includes support for OpenAI's vector store API. It
|
|
|
332
781
|
provides a vector database as a HTTP service but we won't cover
|
|
333
782
|
that here.
|
|
334
783
|
|
|
784
|
+
See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/embeddings) to learn more.
|
|
785
|
+
|
|
335
786
|
```ruby
|
|
336
787
|
require "llm"
|
|
337
788
|
|
|
@@ -348,116 +799,57 @@ Document.create!(
|
|
|
348
799
|
)
|
|
349
800
|
```
|
|
350
801
|
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
The runtime supports six different concurrency strategies that have
|
|
354
|
-
different attributes. The choice between all of them often depends
|
|
355
|
-
on the requirements of your application.
|
|
356
|
-
|
|
357
|
-
IO-bound tools are a good fit for the `:async`, `:thread`,
|
|
358
|
-
and `:fiber` strategies while true parallelism can be achieved
|
|
359
|
-
with the `:fork` and `:ractor` strategies. The
|
|
360
|
-
`:sequential` strategy runs tools one at a time and is the default.
|
|
361
|
-
The `:fork` strategy also provides a separate process that offers
|
|
362
|
-
isolation from its parent.
|
|
802
|
+
### Images
|
|
363
803
|
|
|
364
|
-
|
|
365
|
-
|
|
804
|
+
A handful of providers can generate images from a text prompt.
|
|
805
|
+
OpenAI, Google, xAI, and DeepInfra all support it. The API is
|
|
806
|
+
the same across providers:
|
|
366
807
|
|
|
367
808
|
```ruby
|
|
368
809
|
require "llm"
|
|
369
810
|
|
|
370
|
-
llm
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
agent.talk "Run the tools in parallel"
|
|
811
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
812
|
+
res = llm.images.create(prompt: "a dog on a rocket to the moon")
|
|
813
|
+
IO.copy_stream res.images[0], "rocket.png"
|
|
374
814
|
```
|
|
375
815
|
|
|
376
|
-
|
|
816
|
+
##### DeepSeek
|
|
377
817
|
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
The runtime includes first-class support for both ActiveRecord *and* Sequel, and
|
|
385
|
-
for both Rack-based applications *and* Rails-based applications. On databases
|
|
386
|
-
where it is supported, such as PostgreSQL, the column can be optimized by using
|
|
387
|
-
the `jsonb` type.
|
|
818
|
+
DeepSeek does not have a dedicated image model, but the runtime
|
|
819
|
+
generates SVG vector graphics through its text model. Each
|
|
820
|
+
generation produces a valid SVG document that can be converted
|
|
821
|
+
to PNG with tools like `rsvg-convert`. Pass an existing agent
|
|
822
|
+
to maintain a session across generations:
|
|
388
823
|
|
|
389
824
|
```ruby
|
|
390
|
-
require "active_record"
|
|
391
825
|
require "llm"
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
class Agent < ApplicationRecord
|
|
395
|
-
acts_as_agent
|
|
396
|
-
set name: "my-agent",
|
|
397
|
-
instructions: "solve the user's query",
|
|
398
|
-
model: "deepseek-v4-pro",
|
|
399
|
-
tools: [Research, FinalizeResearch, ActOnResearch]
|
|
400
|
-
|
|
401
|
-
private
|
|
402
|
-
|
|
403
|
-
# By convention, this method defines the provider for a model.
|
|
404
|
-
# If necessary, it can be renamed with: provider: :your_method.
|
|
405
|
-
def set_provider
|
|
406
|
-
LLM.deepseek(key: ENV["KEY"])
|
|
407
|
-
end
|
|
826
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
408
827
|
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
end
|
|
414
|
-
end
|
|
828
|
+
##
|
|
829
|
+
# First generation
|
|
830
|
+
res = llm.images.create(prompt: "a rocket on the moon")
|
|
831
|
+
IO.copy_stream res.images[0], "rocket.svg"
|
|
415
832
|
|
|
416
|
-
|
|
417
|
-
|
|
833
|
+
##
|
|
834
|
+
# Refine with follow-up prompts (shares context)
|
|
835
|
+
res = llm.images.create(prompt: "add a dog next to the rocket",
|
|
836
|
+
agent: res.agent)
|
|
837
|
+
IO.copy_stream res.images[0], "rocket-with-dog.svg"
|
|
418
838
|
```
|
|
419
839
|
|
|
420
840
|
## FAQ
|
|
421
841
|
|
|
422
842
|
<details>
|
|
423
|
-
<summary>What
|
|
843
|
+
<summary>What about local LLM support?</summary>
|
|
424
844
|
<br>
|
|
425
845
|
<p>
|
|
426
|
-
|
|
427
|
-
|
|
428
|
-
|
|
429
|
-
|
|
430
|
-
In no particular order:
|
|
431
|
-
|
|
432
|
-
πΊπΈ OpenAI <br>
|
|
433
|
-
πΊπΈ DeepInfra <br>
|
|
434
|
-
πΊπΈ xAI <br>
|
|
435
|
-
πΊπΈ Google (Gemini) <br>
|
|
436
|
-
πΊπΈ AWS bedrock <br>
|
|
437
|
-
πΊπΈ Anthropic <br>
|
|
438
|
-
π¨π³ DeepSeek <br>
|
|
439
|
-
π¨π³ zAI <br>
|
|
440
|
-
πͺπΊ Mistral <br>
|
|
441
|
-
|
|
442
|
-
**Weights**
|
|
443
|
-
|
|
444
|
-
The following providers provide access to open-weight models. <br>
|
|
445
|
-
In no particular order:
|
|
446
|
-
|
|
447
|
-
πΊπΈ DeepInfra <br>
|
|
448
|
-
πΊπΈ AWS bedrock <br>
|
|
449
|
-
π¨π³ DeepSeek <br>
|
|
450
|
-
π¨π³ zAI <br>
|
|
451
|
-
πͺπΊ Mistral <br>
|
|
452
|
-
|
|
453
|
-
**Local**
|
|
454
|
-
|
|
455
|
-
The following providers can be run locally on your own hardware. <br>
|
|
456
|
-
In no particular order:
|
|
846
|
+
The following providers can be run used with models that
|
|
847
|
+
are running on your own hardware. They're reasonably well
|
|
848
|
+
tested but not my main driver:
|
|
849
|
+
</p>
|
|
457
850
|
|
|
458
851
|
* Ollama
|
|
459
852
|
* Llamacpp
|
|
460
|
-
</p>
|
|
461
853
|
</details>
|
|
462
854
|
|
|
463
855
|
<details>
|
|
@@ -489,11 +881,10 @@ type.
|
|
|
489
881
|
If you're on a budget, DeepSeek is hard to beat.
|
|
490
882
|
</details>
|
|
491
883
|
<details>
|
|
492
|
-
<summary>
|
|
493
|
-
<br>
|
|
494
|
-
Yes.
|
|
884
|
+
<summary>Sources other than GitHub?</summary>
|
|
495
885
|
<br>
|
|
496
|
-
|
|
886
|
+
<p>
|
|
887
|
+
We are on the <a href="https://radicle.network">radicle.network</a> as well.
|
|
497
888
|
<br>
|
|
498
889
|
Every commit that lands on GitHub also lands on Radicle.
|
|
499
890
|
<br>
|
|
@@ -502,6 +893,30 @@ Our repository ID is z2PtfQ6dYwyYaW2aGrztG1sMyDmCE.
|
|
|
502
893
|
Browse on <a
|
|
503
894
|
href="https://radicle.network/nodes/iris.radicle.network/z2PtfQ6dYwyYaW2aGrztG1sMyDmCE">the
|
|
504
895
|
web</a>.
|
|
896
|
+
</p>
|
|
897
|
+
</details>
|
|
898
|
+
|
|
899
|
+
<details>
|
|
900
|
+
<summary>Who maintains llm.rb?</summary>
|
|
901
|
+
<br>
|
|
902
|
+
|
|
903
|
+
The llm.rb project is maintained primarily by one
|
|
904
|
+
person. llm.rb has been in active development for more
|
|
905
|
+
than three years and over that time multiple other
|
|
906
|
+
contributors have contributed to llm.rb as well. New
|
|
907
|
+
contributors are always welcome.
|
|
908
|
+
|
|
909
|
+
I use the repl that is distributed with llm.rb to build
|
|
910
|
+
llm.rb itself so there is a healthy feedback loop and
|
|
911
|
+
llm.rb has also been battle tested in production
|
|
912
|
+
environments.
|
|
913
|
+
|
|
914
|
+
I have also also written llm.rb agents within the
|
|
915
|
+
repository that help me maintain the documentation,
|
|
916
|
+
and backport changes to the mruby-llm runtime as well.
|
|
917
|
+
|
|
918
|
+
I am constantly focused on improving llm.rb by using
|
|
919
|
+
it as my primary driver for development.
|
|
505
920
|
</details>
|
|
506
921
|
|
|
507
922
|
## Resources
|