llm.rb 14.0.0 → 15.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (92) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +396 -1826
  3. data/README.md +596 -490
  4. data/bin/llm.rb +148 -72
  5. data/data/alibaba.json +1999 -0
  6. data/data/anthropic.json +205 -205
  7. data/data/bedrock.json +2171 -2114
  8. data/data/deepinfra.json +1143 -938
  9. data/data/deepseek.json +4 -5
  10. data/data/google.json +691 -779
  11. data/data/mistral.json +450 -450
  12. data/data/moonshot.json +100 -100
  13. data/data/openai.json +976 -976
  14. data/data/xai.json +193 -116
  15. data/data/zai.json +187 -187
  16. data/{resources → docs}/deepdive/advanced/cancellation.md +2 -2
  17. data/{resources → docs}/deepdive/advanced/compaction.md +5 -3
  18. data/{resources → docs}/deepdive/advanced/context.md +13 -0
  19. data/{resources → docs}/deepdive/advanced/guard.md +1 -1
  20. data/{resources/deepdive/fundamentals → docs/deepdive/features}/concurrency.md +6 -0
  21. data/{resources/deepdive/fundamentals → docs/deepdive/features}/repl.md +56 -2
  22. data/{resources → docs}/deepdive/fundamentals/agents.md +52 -1
  23. data/docs/deepdive/fundamentals/providers.md +159 -0
  24. data/{resources → docs}/deepdive/fundamentals/skills.md +5 -0
  25. data/{resources → docs}/deepdive/fundamentals/stream.md +36 -3
  26. data/{resources → docs}/deepdive/fundamentals/tools.md +87 -23
  27. data/{resources/deepdive/everything_else → docs/deepdive/reference}/cost.md +20 -10
  28. data/docs/deepdive/reference/model_registry.md +271 -0
  29. data/{resources/deepdive/advanced → docs/deepdive/reference}/tracer.md +7 -0
  30. data/{resources → docs}/deepdive.md +35 -27
  31. data/lib/llm/a2a/transport/http.rb +1 -1
  32. data/lib/llm/active_record/acts_as_llm.rb +19 -5
  33. data/lib/llm/agent.rb +62 -5
  34. data/lib/llm/context.rb +93 -46
  35. data/lib/llm/cost.rb +110 -51
  36. data/lib/llm/error.rb +7 -0
  37. data/lib/llm/function/array.rb +1 -1
  38. data/lib/llm/function/fork/task.rb +14 -1
  39. data/lib/llm/function/sequential/group.rb +20 -13
  40. data/lib/llm/function/sequential/task.rb +1 -8
  41. data/lib/llm/function.rb +5 -4
  42. data/lib/llm/message.rb +5 -4
  43. data/lib/llm/provider.rb +7 -0
  44. data/lib/llm/providers/alibaba/error_handler.rb +34 -0
  45. data/lib/llm/providers/alibaba/request_adapter.rb +13 -0
  46. data/lib/llm/providers/alibaba.rb +93 -0
  47. data/lib/llm/providers/anthropic.rb +0 -1
  48. data/lib/llm/providers/bedrock.rb +8 -1
  49. data/lib/llm/providers/deepseek/request_adapter.rb +2 -33
  50. data/lib/llm/providers/google.rb +0 -1
  51. data/lib/llm/providers/ollama.rb +0 -1
  52. data/lib/llm/providers/openai/responses.rb +0 -1
  53. data/lib/llm/providers/openai/schema.rb +37 -0
  54. data/lib/llm/providers/openai.rb +1 -2
  55. data/lib/llm/registry/model.rb +186 -0
  56. data/lib/llm/registry.rb +45 -14
  57. data/lib/llm/repl/bar.rb +11 -13
  58. data/lib/llm/repl/buffer.rb +1 -1
  59. data/lib/llm/repl/color.rb +8 -1
  60. data/lib/llm/repl/command.rb +12 -0
  61. data/lib/llm/repl/commands/model.rb +39 -0
  62. data/lib/llm/repl/input/cache.rb +45 -0
  63. data/lib/llm/repl/input/char.rb +2 -2
  64. data/lib/llm/repl/input.rb +86 -23
  65. data/lib/llm/repl/markdown.rb +25 -1
  66. data/lib/llm/repl/node.rb +7 -0
  67. data/lib/llm/repl/status.rb +17 -3
  68. data/lib/llm/repl/window.rb +91 -11
  69. data/lib/llm/repl.rb +18 -8
  70. data/lib/llm/sequel/plugin.rb +19 -5
  71. data/lib/llm/skill.rb +21 -8
  72. data/lib/llm/stream.rb +27 -0
  73. data/lib/llm/tool.rb +3 -5
  74. data/lib/llm/tools/rg.rb +2 -1
  75. data/lib/llm/transport/curb.rb +23 -3
  76. data/lib/llm/usage.rb +155 -9
  77. data/lib/llm/version.rb +1 -1
  78. data/lib/llm.rb +111 -29
  79. data/llm.gemspec +16 -11
  80. metadata +89 -35
  81. /data/{resources → docs}/deepdive/advanced/transformer.md +0 -0
  82. /data/{resources → docs}/deepdive/advanced/transports.md +0 -0
  83. /data/{resources/deepdive/fundamentals → docs/deepdive/features}/builtin_tools.md +0 -0
  84. /data/{resources/deepdive/fundamentals → docs/deepdive/features}/database.md +0 -0
  85. /data/{resources/deepdive/fundamentals → docs/deepdive/features}/embeddings.md +0 -0
  86. /data/{resources → docs}/deepdive/fundamentals/schema.md +0 -0
  87. /data/{resources/deepdive/everything_else → docs/deepdive/media}/audio.md +0 -0
  88. /data/{resources/deepdive/everything_else → docs/deepdive/media}/images.md +0 -0
  89. /data/{resources/deepdive/everything_else → docs/deepdive/media}/ocr.md +0 -0
  90. /data/{resources → docs}/deepdive/protocols/a2a.md +0 -0
  91. /data/{resources → docs}/deepdive/protocols/mcp.md +0 -0
  92. /data/{resources/deepdive/everything_else → docs/deepdive/reference}/object.md +0 -0
data/README.md CHANGED
@@ -14,155 +14,17 @@
14
14
 
15
15
  Welcome to the canonical llm.rb repository.
16
16
 
17
- llm.rb is an advanced runtime for building capable AI applications
18
- on CRuby. It has zero runtime dependencies by default, and a single
19
- coherent API that spans 12+ providers. Streaming, tools, guards,
20
- compaction, the REPL, builtin MCP/A2A support and the database
21
- integrations all build on the same three concepts: providers,
22
- contexts, and agents.
17
+ llm.rb is an advanced runtime for building agentic AI applications
18
+ on CRuby. It has zero runtime dependencies by default, it supports
19
+ concurrent and parallel tool execution and has a single coherent API
20
+ that spans 13+ providers. Streaming, tools, guards, compaction, the
21
+ REPL, builtin MCP/A2A support and the database integrations all build
22
+ on the same three concepts: providers, contexts, and agents.
23
23
 
24
24
  Once you learn the fundamentals, everything else falls into place
25
25
  naturally. Some features, such as ActiveRecord support, require
26
26
  optional dependencies that are opt-in.
27
27
 
28
- ## Features
29
-
30
- One runtime, 12+ providers. The same API drives OpenAI, Anthropic,
31
- Google Gemini, Moonshot (kimi), Mistral, DeepSeek, DeepInfra, xAI, Z.ai,
32
- AWS Bedrock, Ollama, and llama.cpp, so switching models or providers
33
- can be done with minimal code change.
34
-
35
- <details>
36
- <summary><b>Agents</b></summary>
37
-
38
- * **First-class support** <br>
39
- llm.rb is designed to build agents. They can be attached to a
40
- terminal-based read-eval-print loop (repl), persisted to disk
41
- or a database column, run tools concurrently and be safely
42
- interrupted.
43
-
44
- * **Builtin REPL** <br>
45
- A curses-based TUI for talking to an agent interactively. It
46
- renders markdown, shows a live status line with context usage
47
- and running cost, and recalls previous turns, so a
48
- conversation survives a restart.
49
-
50
- * **Persistence** <br>
51
- Set `path:` and the agent saves its conversation to disk
52
- automatically. ActiveRecord and Sequel support keep the same
53
- state in a single database column, so you pick the storage
54
- and the API stays identical.
55
-
56
- </details>
57
-
58
- <details>
59
- <summary><b>MCP &amp; A2A</b></summary>
60
-
61
- * **MCP** <br>
62
- The Model Context Protocol is first-class. Point an MCP client
63
- at any tool server over stdio or HTTP, and its tools translate
64
- into local `LLM::Tool` subclasses, with the same tracing and
65
- error handling.
66
-
67
- * **A2A** <br>
68
- The Agent 2 Agent protocol is first-class. Point an A2A client
69
- at another agent over HTTP or JSON-RPC, and call its skills
70
- exactly like local tools.
71
-
72
- </details>
73
-
74
- <details>
75
- <summary><b>ORM</b></summary>
76
-
77
- * **ActiveRecord** <br>
78
- Add `acts_as_agent` to a model and the agent state lives in a
79
- single database column, saved after every turn and restored
80
- on load. Works in Rack and Rails apps, with `jsonb` on
81
- PostgreSQL.
82
-
83
- * **Sequel** <br>
84
- Add `plugin :agent` to a Sequel model for the same single-
85
- column persistence, with the `pg_json` extension loaded
86
- automatically on PostgreSQL.
87
-
88
- </details>
89
-
90
- <details>
91
- <summary><b>RAG</b></summary>
92
-
93
- * **RAG, out of the box** <br>
94
- Embeddings, OCR, and OpenAI's vector stores API come first-
95
- class. Ground answers in your own documents, with vectors in
96
- a managed store or in your own database such as sqlite-vec
97
- or pgvector.
98
-
99
- </details>
100
-
101
- <details>
102
- <summary><b>Runtime</b></summary>
103
-
104
- * **Streaming** <br>
105
- Streaming is first-class, with structured callbacks for
106
- content, reasoning, and tool calls. Tools can start while
107
- the model is still talking, so the first result lands
108
- before the response finishes.
109
-
110
- * **Concurrency** <br>
111
- Six ways to run tools: sequential, threads, async, fibers,
112
- forks, and ractors. Plus three HTTP backends, so you pick
113
- the concurrency model that fits the workload, not the other
114
- way around.
115
-
116
- * **Interruption** <br>
117
- Cancel an in-flight request or a running tool at any moment,
118
- on any transport or concurrency strategy. A stuck call never
119
- leaves a thread running that you can't stop.
120
-
121
- </details>
122
-
123
- <details>
124
- <summary><b>Provider extras</b></summary>
125
-
126
- * **DeepSeek-optimized** <br>
127
- DeepSeek is the most cost-effective option for API users, and the
128
- runtime closes its gaps: [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
129
- makes structured outputs work despite no official API, and
130
- `images.create`/`edit` produce SVG vector graphics.
131
- </details>
132
-
133
- <details>
134
- <summary><b>Portable</b></summary>
135
-
136
- * **mruby-llm** <br>
137
- The same runtime runs on mruby as
138
- [mruby-llm](https://github.com/r-uby-dev/mruby-llm), with an
139
- almost identical interface and the same set of capabilities.
140
-
141
- </details>
142
-
143
- <details>
144
- <summary><b>Everything else</b></summary>
145
-
146
- * **Skills** <br>
147
- Write a SKILL.md, get a tool. The runtime spawns a
148
- disposable subagent with the skill's instructions and tool
149
- set for one turn, then discards it. Fresh and stateless
150
- every call.
151
-
152
- * **A unified plugin family** <br>
153
- Compactors, transformers, and guards all share one
154
- interface. Context management, message rewriting, and tool
155
- supervision (policy, quotas, loop detection) plug in the
156
- same way and compose freely.
157
-
158
- * **Cost and usage tracking** <br>
159
- Every context tracks its own cost and token usage, per turn.
160
- Break the spend down by input, output, cache, and reasoning,
161
- so the exact cost of any conversation is visible at a
162
- glance.
163
-
164
- </details>
165
-
166
28
  ## Install
167
29
 
168
30
  ```bash
@@ -171,7 +33,7 @@ gem install llm.rb
171
33
 
172
34
  ## Quick start
173
35
 
174
- #### LLM::Agent
36
+ ### Agents
175
37
 
176
38
  The
177
39
  [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
@@ -186,82 +48,82 @@ require "llm"
186
48
 
187
49
  llm = LLM.deepseek(key: ENV["KEY"])
188
50
  agent = LLM::Agent.new(llm, stream: $stdout)
189
- agent.talk "Hello world"
51
+ agent.talk "hello world"
190
52
  ```
53
+ <details>
54
+ <summary>Stream</summary>
55
+ <br>
191
56
 
192
- ##### set
57
+ Streams can be simple IO objects or subclasses of
58
+ [`LLM::Stream`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
59
+ with structured callbacks for content,
60
+ reasoning, tool calls, tool returns, and compaction.
61
+ Streams can also observe message transformers, which rewrite
62
+ outgoing messages before they reach the provider.
193
63
 
194
- [`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
195
- is a class-level DSL that accepts a Hash of properties. Each key resolves to a
196
- corresponding class accessor: `name`, `description`, `model`, `tools`,
197
- `instructions`, `schema`, `stream`, `tracer`, `concurrency`, `confirm`,
198
- `path`, `skills`, and `tool_budget`. All options are optional; zero or
199
- more can be set.
200
- An error is raised for unknown keys so that typos are caught early.
64
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/stream/)
65
+ to learn more.
201
66
 
202
67
  ```ruby
203
- class SystemAdmin < LLM::Agent
204
- set name: "sysadmin",
205
- description: "system administration agent",
206
- model: "deepseek-v4-pro",
207
- tools: [Shell]
208
- end
68
+ class MyStream < LLM::Stream
69
+ # Visible assistant output.
70
+ def on_content(content)
71
+ print content
72
+ end
209
73
 
210
- llm = LLM.deepseek(key: ENV["KEY"])
211
- agent = SystemAdmin.new(llm)
212
- agent.talk "Run 'date'"
213
- ```
74
+ # Reasoning output streamed separately from visible content.
75
+ def on_reasoning_content(content)
76
+ warn content
77
+ end
214
78
 
215
- ##### Persistence
79
+ # A streamed tool call has been fully parsed.
80
+ def on_tool_call(tool)
81
+ end
216
82
 
217
- Set `path:` on an agent for automatic filesystem persistence;
218
- the agent restores conversation history from the file on startup
219
- and saves it back after every turn, with no manual serialization
220
- code. For database-backed persistence, ActiveRecord and Sequel
221
- integrations are also available (see the
222
- [database deepdive](https://r.uby.dev/llm/deepdive/advanced/database)
223
- for details). All persistence options use the same underlying
224
- serialization.
83
+ # Queued streamed tool work has returned.
84
+ def on_tool_return(tool, result)
85
+ end
225
86
 
226
- ```ruby
227
- require "llm"
87
+ # Before a transformer rewrites an outgoing message.
88
+ def on_transform(transformer)
89
+ end
228
90
 
229
- llm = LLM.deepseek(key: ENV["KEY"])
230
- agent = LLM::Agent.new(llm, path: "session.json")
231
- agent.talk "remember my name is robert"
91
+ # Aftter a transformer rewrites an outgoing message.
92
+ def on_transform_finish(transformer)
93
+ end
232
94
 
233
- # Next time, the conversation is restored automatically:
234
- agent = LLM::Agent.new(llm, path: "session.json")
235
- agent.talk "what's my name?"
236
- ```
95
+ # Before a compactor trims the conversation.
96
+ def on_compaction(compactor)
97
+ end
237
98
 
238
- #### LLM::Context
99
+ # After a compactor trims the conversation.
100
+ def on_compaction_finish(compactor)
101
+ end
239
102
 
240
- The
241
- [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
242
- class is at the heart of the runtime
243
- and it is what
244
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
245
- uses under the hood.
246
- It requires that the tool call loop be managed manually -
247
- sometimes that can be useful, but usually for advanced use-cases.
248
- If you're new to llm.rb, try
249
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html) first.
103
+ # Before a skill's subagent runs.
104
+ def on_skill_call(skill)
105
+ end
250
106
 
251
- Every context tracks its own token usage and estimated cost. After any
252
- turn, you can read the cost breakdown through
253
- [`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method),
254
- and the REPL shows the running total in the status line.
107
+ # After a skill's subagent runs.
108
+ # The subagent that ran it, the skill, and its response are passed
109
+ # through, so you can introspect the agent, tally skill usage, or
110
+ # track costs.
111
+ def on_skill_return(agent, skill, result)
112
+ end
255
113
 
256
- ```ruby
257
- require "llm"
114
+ # A request was rate limited and will be retried.
115
+ def on_rate_limit(error)
116
+ end
117
+ end
258
118
 
259
119
  llm = LLM.deepseek(key: ENV["KEY"])
260
- ctx = LLM::Context.new(llm, stream: $stdout)
261
- ctx.talk "Hello world"
120
+ agent = LLM::Agent.new(llm, stream: MyStream.new)
121
+ agent.talk "Explain Ruby fibers."
262
122
  ```
123
+ </details>
263
124
 
264
- #### LLM::Tool
125
+ <details><summary>Tools</summary>
126
+ <br>
265
127
 
266
128
  Subclasses of
267
129
  [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
@@ -271,9 +133,7 @@ call them on your behalf, and they're one of the most powerful features
271
133
  for extending the feature set or abilities of a model.
272
134
 
273
135
  The runtime also ships with a catalog of built-in tools for
274
- filesystem, search, and shell operations. See the
275
- [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/builtin_tools)
276
- for details.
136
+ filesystem, search, and shell operations. <br> See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/tools) to learn more.
277
137
 
278
138
  ```ruby
279
139
  class ReadFile < LLM::Tool
@@ -286,114 +146,145 @@ class ReadFile < LLM::Tool
286
146
  {contents: File.read(path)}
287
147
  end
288
148
  end
149
+
150
+ llm = LLM.deepseek(key: ENV["KEY"])
151
+ agent = LLM::Agent.new(llm, tools: [ReadFile], stream: $stdout)
152
+ agent.talk "summarize README.md"
289
153
  ```
154
+ </details>
155
+ <details>
156
+ <summary>Skills</summary>
157
+ <br>
290
158
 
291
- ##### set
159
+ A skill turns a markdown file into a callable tool. When the model
160
+ calls it, the runtime spawns a subagent with the skill's instructions
161
+ as its system prompt and the skill's own tool set. The subagent runs
162
+ one turn and returns the result, then is discarded. Each call
163
+ is fresh and stateless.
292
164
 
293
- [`LLM::Tool.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html#set-class_method)
294
- is an alternative way to define tool properties using a Hash. It works
295
- the same way as
296
- [`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
297
- and accepts the same keys that the individual methods do: `name`,
298
- `description`, `parameters`, `required`, and `defaults`:
165
+ A [LLM::Stream](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
166
+ can be notified as a skill starts and when it returns. The `on_skill_return`
167
+ callback hands back the subagent that ran the skill, so you can inspect
168
+ its conversation, measure its usage, track costs or add a verification
169
+ step (eg `subagent.talk("verify your work")`).
299
170
 
300
- ```ruby
301
- class MathTool < LLM::Tool
302
- set name: "math",
303
- description: "Performs arithmetic",
304
- parameters: [
305
- [:x, Integer, "first number" , {required: true}],
306
- [:y, Integer, "second number", {default: 0}]
307
- ]
308
-
309
- def call(x:, y: 0)
310
- {result: x + y}
311
- end
312
- end
313
- ```
171
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/skills) to learn more.
314
172
 
315
- #### LLM::Stream
173
+ ##### summary.md
316
174
 
317
- Streams can be simple IO objects or subclasses of
318
- [`LLM::Stream`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
319
- with structured callbacks for content,
320
- reasoning, tool calls, tool returns, and compaction.
321
- Streams can also observe message transformers, which rewrite
322
- outgoing messages before they reach the provider (see the
323
- [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transformer)).
175
+ ```markdown
176
+ ---
177
+ name: summary
178
+ description: Reads recent git history and writes a summary
179
+ tools: all
180
+ ---
324
181
 
325
- ```ruby
326
- class MyStream < LLM::Stream
327
- def on_content(content)
328
- print content
329
- end
182
+ Collect the recent git log, analyze each commit,
183
+ and write a summary to summary.txt.
184
+ ```
330
185
 
331
- def on_reasoning_content(content)
332
- warn content
333
- end
334
- end
186
+ ##### agent.rb
335
187
 
336
- llm = LLM.deepseek(key: ENV["KEY"])
337
- agent = LLM::Agent.new(llm, stream: MyStream.new)
338
- agent.talk "Explain Ruby fibers."
188
+ ```ruby
189
+ require "llm"
190
+
191
+ llm = LLM.deepseek(key: ENV["KEY"])
192
+ agent = LLM::Agent.new(llm, skills: ["summary.md"])
193
+ agent.talk "Summarize the last week of work"
339
194
  ```
195
+ </details>
340
196
 
341
- #### LLM::Schema
197
+ <details>
198
+ <summary>Concurrency</summary>
199
+ <br>
342
200
 
343
- [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
344
- subclasses produce typed, structured
345
- output from any model call. Pass a schema to
346
- [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk-instance_method),
347
- [`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk-instance_method),
348
- or
349
- [`LLM::Provider#complete`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#complete-instance_method)
350
- to receive validated JSON instead of free text. Schemas work alongside tools and streams.
201
+ The runtime supports six different concurrency strategies that have
202
+ different attributes. The choice between all of them often depends
203
+ on the requirements of your application.
351
204
 
352
- [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
353
- can define objects, arrays, enums, nested schemas,
354
- and more. It is also used internally by
355
- [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html) for parameter
356
- definitions, so you already benefit from it when you declare tool
357
- parameters.
205
+ IO-bound tools are a good fit for the `:async`, `:thread`,
206
+ and `:fiber` strategies while true parallelism can be achieved
207
+ with the `:fork` and `:ractor` strategies. The
208
+ `:sequential` strategy runs tools one at a time and is the default.
209
+ The `:fork` strategy also provides a separate process that offers
210
+ isolation from its parent.
358
211
 
359
- The
360
- [`LLM::DeepSeek`](https://r.uby.dev/api-docs/llm.rb/LLM/DeepSeek.html)
361
- provider includes runtime-level optimisations such as structured
362
- output support (despite no official structured outputs API) and
363
- SVG image generation. This example uses
364
- [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html) with
365
- DeepSeek:
212
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/concurrency) to learn more.
366
213
 
367
214
  ```ruby
368
- class Weather < LLM::Schema
369
- property :city, String, "The city name"
370
- property :temperature, Float, "Current temperature"
371
- property :conditions, String, "Weather conditions"
372
- required %i[city temperature conditions]
373
- end
215
+ require "llm"
216
+ require "llm/tools"
374
217
 
375
- llm = LLM.deepseek(key: ENV["KEY"])
376
- agent = LLM::Agent.new(llm, schema: Weather)
377
- res = agent.talk "Weather in Paris?"
378
- res.content! # => {city: "Paris", temperature: 15.0, conditions: "Cloudy"}
218
+ llm = LLM.deepseek(key: ENV["KEY"])
219
+ tools = LLM::Tool.subclasses
220
+ agent = LLM::Agent.new(llm, tools:, concurrency: :fork)
221
+ agent.talk "Run the tools in parallel"
379
222
  ```
380
223
 
381
- #### LLM::REPL
224
+ </details>
225
+ <details>
226
+ <summary>Cancellation</summary>
227
+ <br>
228
+
229
+ Abort a request mid-stream and interrupt any running tools with
230
+ [`LLM::Agent#interrupt!`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#interrupt!)
231
+ (or `cancel!`), from any thread. The runtime raises
232
+ [`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html)
233
+ on the caller and on every active tool. A forked tool gets interrupted over
234
+ the control channel, a ractor via message passing, and pending tools
235
+ are stopped before they run. The in-flight HTTP request is closed
236
+ too, so a turn you no longer want stops without burning tokens.
237
+
238
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/cancellation) to learn more.
239
+
240
+ ```ruby
241
+ llm = LLM.deepseek(key: ENV["KEY"])
242
+ agent = LLM::Agent.new(llm)
243
+ Thread.new { sleep(1); agent.cancel! }
244
+
245
+ begin
246
+ agent.talk "write a very long poem", stream: $stdout
247
+ rescue LLM::Interrupt
248
+ puts "cancelled"
249
+ end
250
+ ```
251
+ </details>
252
+ <details>
253
+ <summary>Console (<code>binding.pry</code> for agents)</summary>
254
+ <br>
382
255
 
383
256
  The [LLM::Agent#repl](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#repl-instance_method)
384
- method drops you into a curses-based TUI for talking to an
385
- agent interactively. It renders markdown directly in the
386
- terminal and shows a live status line with context usage,
387
- running cost, and the current tool call. A second thread keeps
388
- the UI responsive while the model works. Think of it as
257
+ method drops you into a highly capable read-eval-print loop (REPL)
258
+ that is built on top of curses. It can help you debug agents,
259
+ test your tools, connect to MCP servers, and even A2A agents.
260
+ The REPL stands out because it connects to the surrounding
261
+ runtime and it can be extended by your code. Think of it as
389
262
  `binding.pry` but for agents.
390
263
 
391
- Set `path:` on the agent for automatic persistence across REPL
392
- sessions. The `tools:` option attaches extra tools for the
393
- duration of the session. Recall previous turns with Ctrl+P and
394
- Ctrl+N. For the full reference, see the
395
- [REPL section](https://r.uby.dev/llm/deepdive/fundamentals/repl) in the
396
- deepdive.
264
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/repl) to learn more.
265
+
266
+ ##### Demo
267
+
268
+ [Watch in high quality on asciinema](https://asciinema.org/a/OsS8wwaasKasoDDz)
269
+
270
+ ![llm.rb REPL demo](demo.gif)
271
+
272
+
273
+ ##### Installation
274
+
275
+ The REPL is distributed with llm.rb so you don't have to install
276
+ a separate gem but it requires a number of optional dependencies
277
+ to be installed separately. The following gems provide the full
278
+ experience:
279
+
280
+ gem install curses kramdown xchan.rb test-cmd.rb
281
+
282
+ ##### Persistence
283
+
284
+ The `path:` option can be set on an agent for automatic persistence
285
+ across REPL sessions. The `tools:` option attaches extra tools
286
+ for the duration of the session. Recall previous turns with Ctrl+P and
287
+ Ctrl+N.
397
288
 
398
289
  ```ruby
399
290
  require "llm"
@@ -407,20 +298,110 @@ agent.repl(tools: LLM::Tool.subclasses)
407
298
  ##### CLI
408
299
 
409
300
  The `llm.rb` executable is available on your PATH after installation.
410
- It starts a REPL session from any directory:
301
+ It starts a REPL session from any directory.The CLI auto-detects your
302
+ provider from standard environment variables (`DEEPSEEK_API_KEY`,
303
+ `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, etc.). Persistent sessions are
304
+ stored under `~/.llm.rb/` and restored automatically on your next visit.
411
305
 
412
306
  ```bash
413
307
  llm.rb # auto-detect from $DEEPSEEK_API_KEY
414
308
  llm.rb -p openai # use OpenAI explicitly
415
309
  llm.rb -t # temporary session, no persistence
416
310
  ```
311
+ </details>
312
+ <details>
313
+ <summary>Persistence</summary>
314
+ <br>
315
+
316
+ Set `path:` on an agent for automatic filesystem persistence:
317
+ the agent restores conversation history from the file on startup
318
+ and saves it back after every turn, with no manual serialization
319
+ code. For database-backed persistence, ActiveRecord and Sequel
320
+ integrations are also available. All persistence options use the same
321
+ underlying serialization.
322
+
323
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/database) to learn more.
324
+
325
+ ```ruby
326
+ require "llm"
327
+
328
+ llm = LLM.deepseek(key: ENV["KEY"])
329
+ agent = LLM::Agent.new(llm, path: "session.json")
330
+ agent.talk "remember my name is robert"
331
+
332
+ # Next time, the conversation is restored automatically:
333
+ agent = LLM::Agent.new(llm, path: "session.json")
334
+ agent.talk "what's my name?"
335
+ ```
336
+ </details>
337
+ <details><summary>ActiveRecord | Sequel</summary>
338
+ <br>
339
+
340
+ Because both
341
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) and
342
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
343
+ can be serialized to JSON and stored in a simple string, both ActiveRecord
344
+ and Sequel support can be implemented within a single column on a single row.
345
+
346
+ The runtime includes first-class support for both ActiveRecord / Sequel, and
347
+ for both Rack-based / Rails-based applications. On databases
348
+ where it is supported, such as PostgreSQL, the column can be optimized by using
349
+ the `jsonb` type.
350
+
351
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/database) to learn more.
352
+
353
+ ```ruby
354
+ require "active_record"
355
+ require "llm"
356
+ require "llm/active_record"
357
+
358
+ class Email < ApplicationRecord
359
+ acts_as_agent do |agent|
360
+ agent.set name: "mail",
361
+ instructions: "Write concise, friendly replies to emails",
362
+ model: "deepseek-v4-pro"
363
+ end
364
+
365
+ def draft_reply!
366
+ talk("Draft a reply to:\n\n#{body}")
367
+ end
368
+
369
+ def summarize
370
+ talk("Summarize this email thread in a few sentences")
371
+ end
372
+
373
+ private
374
+
375
+ ##
376
+ # By convention, this method defines the provider for a model.
377
+ # If necessary, it can be renamed with: provider: :your_method.
378
+ def set_provider
379
+ LLM.deepseek(key: ENV["KEY"])
380
+ end
381
+
382
+ ##
383
+ # By convention, this method returns the context options given
384
+ # to LLM::Context or LLM::Agent. This method can be left undefined.
385
+ def set_context
386
+ {}
387
+ end
388
+ end
417
389
 
418
- The CLI auto-detects your provider from standard environment variables
419
- (`DEEPSEEK_API_KEY`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, etc.).
420
- Persistent sessions are stored under `~/.llm.rb/` and restored
421
- automatically on your next visit.
390
+ email = Email.create!(subject: "Streaming support", body: "How do I stream responses?")
391
+ email.draft_reply!
392
+
393
+ ##
394
+ # The conversation (the email and the draft
395
+ # reply) is persisted to the email's column. A
396
+ # fresh instance restores it and continues the
397
+ # thread, so the summary below knows what was
398
+ # already drafted:
399
+ Email.find(email.id).summarize
400
+ ```
401
+ </details>
422
402
 
423
- #### LLM::MCP
403
+ <details><summary>MCP</summary>
404
+ <br>
424
405
 
425
406
  The Model Context Protocol (MCP) has first-class support
426
407
  in llm.rb. The stdio and http transports work out of the
@@ -430,6 +411,9 @@ used with
430
411
  [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) or
431
412
  [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
432
413
 
414
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/protocols/mcp/), and the
415
+ [deepdive.md on persistent connections](https://r.uby.dev/llm/deepdive/features/transports) to learn more.
416
+
433
417
  ```ruby
434
418
  require "llm"
435
419
 
@@ -438,24 +422,9 @@ mcp = LLM::MCP.stdio(argv: ["ruby", "server.rb"])
438
422
  agent = LLM::Agent.new(llm, stream: $stdout, tools: mcp.tools)
439
423
  agent.talk "Run the tool"
440
424
  ```
441
-
442
- ##### Persistent connections
443
-
444
- Set `persistent: true` on HTTP transports to reuse connections
445
- across requests. This uses
446
- [`Net::HTTP::Persistent`](https://github.com/drbrain/net-http-persistent)
447
- under the hood and avoids opening a new TCP connection for every
448
- request:
449
-
450
- ```ruby
451
- mcp = LLM::MCP.http(
452
- url: "https://api.githubcopilot.com/mcp/",
453
- headers: {"Authorization" => "Bearer #{ENV.fetch('GITHUB_PAT')}"},
454
- persistent: true
455
- )
456
- ```
457
-
458
- #### LLM::A2A
425
+ </details>
426
+ <details><summary>A2A</summary>
427
+ <br>
459
428
 
460
429
  The Agent 2 Agent (A2A) protocol has first-class support
461
430
  in llm.rb. The http and jsonrpc transports work out of the
@@ -465,6 +434,9 @@ used with
465
434
  [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) or
466
435
  [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
467
436
 
437
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/protocols/a2a/), and the
438
+ [deepdive.md on persistent connections](https://r.uby.dev/llm/deepdive/features/transports) to learn more.
439
+
468
440
  ```ruby
469
441
  require "llm"
470
442
 
@@ -473,21 +445,51 @@ a2a = LLM::A2A.rest(url: "https://remote-agent.example.com")
473
445
  agent = LLM::Agent.new(llm, stream: $stdout, tools: a2a.skills)
474
446
  agent.talk "Run the skill"
475
447
  ```
448
+ </details>
449
+
450
+ <details><summary>Structured outputs</summary>
451
+ <br>
452
+
453
+ [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
454
+ subclasses produce typed, structured
455
+ output from any model call. Pass a schema to
456
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk-instance_method),
457
+ [`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk-instance_method),
458
+ or
459
+ [`LLM::Provider#complete`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#complete-instance_method)
460
+ to receive validated JSON instead of free text. Schemas work alongside tools and streams.
476
461
 
477
- ##### Persistent connections
462
+ [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
463
+ can define objects, arrays, enums, nested schemas,
464
+ and more. It is also used internally by
465
+ [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html) for parameter
466
+ definitions, so you already benefit from it when you declare tool
467
+ parameters.
478
468
 
479
- Set `persistent: true` on HTTP transports to reuse connections
480
- across requests. This uses
481
- [`Net::HTTP::Persistent`](https://github.com/drbrain/net-http-persistent)
482
- under the hood and avoids opening a new TCP connection for every
483
- request:
469
+ The
470
+ [`LLM::DeepSeek`](https://r.uby.dev/api-docs/llm.rb/LLM/DeepSeek.html)
471
+ provider includes runtime-level optimisations such as structured
472
+ output support (despite no official structured outputs API) and
473
+ SVG image generation. This example uses
474
+ [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html) with
475
+ DeepSeek:
484
476
 
485
477
  ```ruby
486
- a2a = LLM::A2A.rest(url: "https://agent.example.com", persistent: true)
487
- a2a = LLM::A2A.jsonrpc(url: "https://agent.example.com", persistent: true)
488
- ```
478
+ class Weather < LLM::Schema
479
+ property :city, String, "The city name"
480
+ property :temperature, Number, "Current temperature"
481
+ property :conditions, String, "Weather conditions"
482
+ required %i[city temperature conditions]
483
+ end
489
484
 
490
- #### LLM::Guard
485
+ llm = LLM.deepseek(key: ENV["KEY"])
486
+ agent = LLM::Agent.new(llm, schema: Weather)
487
+ res = agent.talk "Weather in Paris?"
488
+ res.content! # => {city: "Paris", temperature: 15.0, conditions: "Cloudy"}
489
+ ```
490
+ </details>
491
+ <details><summary>Guards</summary>
492
+ <br>
491
493
 
492
494
  [`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
493
495
  is the hook that sees every tool call before it runs. A guard
@@ -507,6 +509,8 @@ and implement
507
509
  The pending call arrives as `function:`. Return a value to close
508
510
  the call, or `nil` to let it run:
509
511
 
512
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/guard) to learn more.
513
+
510
514
  ```ruby
511
515
  class PolicyGuard < LLM::Guard
512
516
  def call(function:)
@@ -517,141 +521,285 @@ class PolicyGuard < LLM::Guard
517
521
  end
518
522
  end
519
523
 
520
- agent = LLM::Agent.new(llm, guard: PolicyGuard)
524
+ llm = LLM.deepseek(key: ENV["KEY"])
525
+ agent = LLM::Agent.new(llm, tools: [Shell, ReadFile], guard: PolicyGuard)
521
526
  ```
527
+ </details>
522
528
 
523
- #### LLM::Skill
529
+ <details>
530
+ <summary>Transformers</summary>
531
+ <br>
524
532
 
525
- A skill turns a markdown file into a callable tool. When the model
526
- calls it, the runtime spawns a subagent with the skill's instructions
527
- as its system prompt and the skill's own tool set. The subagent runs
528
- one turn and returns the result, then is discarded. Each call
529
- is fresh and stateless. For a deeper explanation see the
530
- [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/skills).
533
+ It is possible to rewrite outgoing messages before they reach the provider with
534
+ [`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html).
535
+ Create a subclass and implement `call(message:)` to scrub sensitive data,
536
+ inject context, or normalize content. The transform runs automatically
537
+ on every turn, so you never have to change your prompt code.
531
538
 
532
- ##### SKILL.md
539
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transformer) to learn more.
533
540
 
534
- ```markdown
535
- ---
536
- name: summary
537
- description: Reads recent git history and writes a summary
538
- tools: all
539
- ---
541
+ ```ruby
542
+ class RedactEmails < LLM::Transformer
543
+ def call(message:)
544
+ content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
545
+ LLM::Message.new(message.role, content, message.extra)
546
+ end
547
+ end
540
548
 
541
- Collect the recent git log, analyze each commit,
542
- and write a summary to summary.txt.
549
+ llm = LLM.deepseek(key: ENV["KEY"])
550
+ agent = LLM::Agent.new(llm, transformer: RedactEmails)
551
+ agent.talk "Contact support@example.com for help"
543
552
  ```
553
+ </details>
544
554
 
545
- ##### agent.rb
555
+ <details>
556
+ <summary>Compactors</summary>
557
+ <br>
558
+
559
+ Every model has a context window: the finite number of tokens it can
560
+ consider in a single request. Generally a compactor will drop or
561
+ summarize older messages to keep the conversation within that window,
562
+ and it runs automatically before every turn. By default it is disabled
563
+ so it is a feature you must opt into.
564
+
565
+ [`LLM::Compactor::Truncate`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Truncate.html)
566
+ keeps the most recent messages via an integer count or a percentage like
567
+ `"80%"`. It preserves tool call and return pairs so the conversation
568
+ never contains an orphaned result. It is also possible to subclass
569
+ [`LLM::Compactor`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor.html)
570
+ to implement your own compactor with its own logic. Streams can observe the
571
+ process through the
572
+ [`LLM::Stream#on_compaction`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction)
573
+ and
574
+ [`LLM::Stream#on_compaction_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction_finish)
575
+ callbacks.
576
+
577
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/compaction) to learn more.
578
+
579
+ ```ruby
580
+ llm = LLM.deepseek(key: ENV["KEY"])
581
+ agent = LLM::Agent.new(
582
+ llm,
583
+ compactor: LLM::Compactor::Truncate,
584
+ compactor_options: {keep: 64}
585
+ )
586
+ agent.talk "Hello"
587
+ ```
588
+ </details>
589
+
590
+ <details>
591
+ <summary>Automatic retries</summary>
592
+ <br>
593
+
594
+ Rate-limited requests are retried automatically by default. Agents
595
+ retry a 429 up to five times with a growing backoff before giving
596
+ up, so most request failures resolve on their own. Set `retry_budget`
597
+ to change the number of retries, or `retry_budget: 0` to disable
598
+ them.
546
599
 
547
600
  ```ruby
548
601
  require "llm"
549
602
 
550
- llm = LLM.deepseek(key: ENV["KEY"])
551
- agent = LLM::Agent.new(llm, skills: ["./skills/summary"])
552
- agent.talk "Summarize the last week of work"
603
+ llm = LLM.deepseek(key: ENV["KEY"])
604
+ agent = LLM::Agent.new(llm, retry_budget: 0)
605
+ agent.talk "Hello"
553
606
  ```
554
607
 
555
- #### RAG
608
+ </details>
609
+
556
610
 
557
- Most providers offer an embedding model that can be
558
- used for semantic search, or similarity search. An
559
- embedding model can generate embeddings that can then
560
- be stored in a database that is optimized for storing
561
- and querying vectors, such as SQLite's [sqlite-vec](https://github.com/asg017/sqlite-vec)
562
- or PostgreSQL's [pg-vector](https://github.com/pgvector/pgvector).
611
+ <details>
612
+ <summary>Observability</summary>
613
+ <br>
563
614
 
564
- llm.rb also includes support for OpenAI's vector store API. It
565
- provides a vector database as a HTTP service but we won't cover
566
- that here. For a deeper explanation see the
567
- [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/embeddings).
615
+ Trace what an agent is doing by attaching a tracer. Hook into
616
+ requests, tool calls, and other runtime events to debug a
617
+ misbehaving agent, monitor latency, or export spans to an
618
+ observability backend. All built-in tracers share one interface,
619
+ so switching between them means changing a class name:
620
+
621
+ * [`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html): human-readable single-line logs to stderr, ideal during development.
622
+ * [`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html):
623
+ exports spans via OTLP for OpenTelemetry in production.
624
+ * [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html):
625
+ structured JSON to stdout or a file.
626
+
627
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/reference/tracer) to learn more.
628
+
629
+ ```ruby
630
+ llm = LLM.deepseek(key: ENV["KEY"])
631
+ agent = LLM::Agent.new(llm, tracer: LLM::Tracer::PrettyLogger.new(llm))
632
+ agent.talk "Hello"
633
+ ```
634
+ </details>
635
+
636
+ <details>
637
+ <summary>As a subclass</summary>
638
+ <br>
639
+
640
+ [`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
641
+ is a class-level DSL that accepts a Hash of properties. Each key resolves to a
642
+ corresponding class accessor: `name`, `description`, `model`, `tools`,
643
+ `instructions`, `schema`, `stream`, `tracer`, `concurrency`, `confirm`,
644
+ `path`, `skills`, `tool_budget`, and `retry_budget`. All options are
645
+ optional; zero or more can be set.
646
+ An error is raised for unknown keys so that typos are caught early.
568
647
 
569
648
  ```ruby
570
649
  require "llm"
650
+ require "llm/tools"
571
651
 
572
- llm = LLM.openai(key: ENV["KEY"])
573
- body = "llm.rb is Ruby's capable AI runtime."
574
- embedding = llm.embed([body]).embeddings.first
652
+ class Agent < LLM::Agent
653
+ set name: "sysadmin",
654
+ description: "system administration agent",
655
+ model: "deepseek-v4-pro",
656
+ tools: [LLM::Tool::Shell]
657
+ end
575
658
 
576
- # Document is your ActiveRecord or Sequel model
577
- # with a vector column (e.g. sqlite-vec or pgvector)
578
- Document.create!(
579
- title: "llm.rb",
580
- body:,
581
- embedding:,
582
- )
659
+ llm = LLM.deepseek(key: ENV["KEY"])
660
+ agent = Agent.new(llm)
661
+ agent.talk "Run 'date'"
583
662
  ```
663
+ </details>
584
664
 
585
- #### Concurrency
665
+ ### Providers
586
666
 
587
- The runtime supports six different concurrency strategies that have
588
- different attributes. The choice between all of them often depends
589
- on the requirements of your application.
667
+ Each provider is constructed with a class-level factory method on
668
+ `LLM`, and the resulting instance is passed to
669
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
670
+ or
671
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html). The
672
+ same API drives every one of them, so switching models is a one-line
673
+ change. See the [deepdive](https://r.uby.dev/llm/deepdive/fundamentals/providers)
674
+ for a full provider reference.
675
+
676
+ #### What providers does llm.rb support?
677
+
678
+ * **Anthropic** (`LLM.anthropic`)
679
+ * **Google** (`LLM.google`)
680
+ * **OpenAI** (`LLM.openai`)
681
+ * **DeepSeek** (`LLM.deepseek`)
682
+ * **DeepInfra** (`LLM.deepinfra`)
683
+ * **xAI** (`LLM.xai`)
684
+ * **Z.ai** (`LLM.zai`)
685
+ * **Moonshot (Kimi)** (`LLM.moonshot`)
686
+ * **Alibaba (Qwen3)** (`LLM.alibaba`, also `LLM.aliyun`)
687
+ * **Mistral** (`LLM.mistral`)
688
+ * **AWS Bedrock** (`LLM.bedrock`)
689
+ * **Ollama** (`LLM.ollama`)
690
+ * **llama.cpp** (`LLM.llamacpp`)
590
691
 
591
- IO-bound tools are a good fit for the `:async`, `:thread`,
592
- and `:fiber` strategies while true parallelism can be achieved
593
- with the `:fork` and `:ractor` strategies. The
594
- `:sequential` strategy runs tools one at a time and is the default.
595
- The `:fork` strategy also provides a separate process that offers
596
- isolation from its parent.
692
+ <details>
693
+ <summary>Implicit</summary>
694
+ <br>
597
695
 
598
- You can learn more about the llm.rb concurrency model in the
599
- [deepdive.md](https://r.uby.dev/llm/deepdive/fundamentals/concurrency).
696
+ Cloud providers can infer their API key automatically
697
+ from a set of common defaults that are defined by
698
+ the [models.dev](https://models.dev) registry that
699
+ is also distributed with llm.rb.
600
700
 
601
701
  ```ruby
602
- require "llm"
702
+ llm = LLM.openai
703
+ llm = LLM.anthropic
704
+ llm = LLM.deepseek
705
+ llm = LLM.alibaba # also: LLM.aliyun
706
+ llm = LLM.moonshot
707
+ llm = LLM.mistral
708
+ ```
709
+ </details>
710
+ <details>
711
+ <summary>Explicit</summary>
712
+ <br>
603
713
 
604
- llm = LLM.deepseek(key: ENV["KEY"])
605
- tools = [FetchNews, FetchStocks, FetchFeeds]
606
- agent = LLM::Agent.new(llm, tools:, concurrency: :fork)
607
- agent.talk "Run the tools in parallel"
714
+ The `key` option can also be providied explicitly, and certain
715
+ providers (eg ollama, llamacpp) usually do not require an API
716
+ key at all.
717
+
718
+ ```ruby
719
+ llm = LLM.openai(key: ENV["OPENAI_API_KEY"])
720
+ llm = LLM.anthropic(key: ENV["ANTHROPIC_API_KEY"])
721
+ llm = LLM.deepseek(key: ENV["DEEPSEEK_API_KEY"])
722
+ llm = LLM.alibaba(key: ENV["DASHSCOPE_API_KEY"]) # also: LLM.aliyun
723
+ llm = LLM.moonshot(key: ENV["MOONSHOT_API_KEY"])
724
+ llm = LLM.mistral(key: ENV["MISTRAL_API_KEY"])
608
725
  ```
726
+ </details>
609
727
 
610
- #### ORM
728
+ <details>
729
+ <summary>Model Registry</summary>
730
+ <br>
611
731
 
612
- Because both
613
- [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) and
614
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
615
- can be serialized to JSON and stored in a simple string, both ActiveRecord
616
- and Sequel support can be implemented within a single column on a single row.
732
+ Each provider ships its model catalog, pricing, limits, and
733
+ modalities with the gem, sourced from [models.dev](https://models.dev).
734
+ Reach it from any provider, context, or agent, enumerate models, or
735
+ sort them by price.
617
736
 
618
- The runtime includes first-class support for both ActiveRecord *and* Sequel, and
619
- for both Rack-based applications *and* Rails-based applications. On databases
620
- where it is supported, such as PostgreSQL, the column can be optimized by using
621
- the `jsonb` type.
737
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/reference/model_registry) to learn more.
622
738
 
623
739
  ```ruby
624
- require "active_record"
625
740
  require "llm"
626
- require "llm/active_record"
627
741
 
628
- class Agent < ApplicationRecord
629
- acts_as_agent
630
- set name: "my-agent",
631
- instructions: "solve the user's query",
632
- model: "deepseek-v4-pro",
633
- tools: [Research, FinalizeResearch, ActOnResearch]
742
+ llm = LLM.openai
743
+ registry = llm.registry # => LLM::Provider#registry
744
+ cheapest = registry.models.sort.first # => LLM::Model
745
+ cheapest.id # => "text-embedding-3-small"
746
+ cheapest.context_window # => 8191
747
+ cheapest.structured_output? # => false
748
+ ```
749
+ </details>
634
750
 
635
- private
751
+ <details>
752
+ <summary>Transports</summary>
753
+ <br>
636
754
 
637
- # By convention, this method defines the provider for a model.
638
- # If necessary, it can be renamed with: provider: :your_method.
639
- def set_provider
640
- LLM.deepseek(key: ENV["KEY"])
641
- end
755
+ The `transport:` option selects which HTTP library a provider uses for
756
+ network communication. Three backends ship out of the box: `net/http`
757
+ is always available and the default, `net/http/persistent` pools
758
+ connections for many requests to the same host, and `curb` wraps
759
+ libcurl. They share one interface, so switching is a one-word change.
642
760
 
643
- # By convention, this method returns the context options given
644
- # to LLM::Context or LLM::Agent.
645
- def set_context
646
- {}
647
- end
648
- end
761
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/advanced/transports) to learn more.
649
762
 
650
- agent = Agent.create!
651
- agent.talk "perform research"
763
+ ```ruby
764
+ llm = LLM.deepseek(
765
+ key: ENV["KEY"],
766
+ transport: :net_http_persistent
767
+ )
768
+ ```
769
+ </details>
770
+
771
+ ### RAG
772
+
773
+ Most providers offer an embedding model that can be
774
+ used for semantic search, or similarity search. An
775
+ embedding model can generate embeddings that can then
776
+ be stored in a database that is optimized for storing
777
+ and querying vectors, such as SQLite's [sqlite-vec](https://github.com/asg017/sqlite-vec)
778
+ or PostgreSQL's [pg-vector](https://github.com/pgvector/pgvector).
779
+
780
+ llm.rb also includes support for OpenAI's vector store API. It
781
+ provides a vector database as a HTTP service but we won't cover
782
+ that here.
783
+
784
+ See the [deepdive.md](https://r.uby.dev/llm/deepdive/features/embeddings) to learn more.
785
+
786
+ ```ruby
787
+ require "llm"
788
+
789
+ llm = LLM.openai(key: ENV["KEY"])
790
+ body = "llm.rb is Ruby's capable AI runtime."
791
+ embedding = llm.embed([body]).embeddings.first
792
+
793
+ # Document is your ActiveRecord or Sequel model
794
+ # with a vector column (e.g. sqlite-vec or pgvector)
795
+ Document.create!(
796
+ title: "llm.rb",
797
+ body:,
798
+ embedding:,
799
+ )
652
800
  ```
653
801
 
654
- #### Images
802
+ ### Images
655
803
 
656
804
  A handful of providers can generate images from a text prompt.
657
805
  OpenAI, Google, xAI, and DeepInfra all support it. The API is
@@ -692,46 +840,16 @@ IO.copy_stream res.images[0], "rocket-with-dog.svg"
692
840
  ## FAQ
693
841
 
694
842
  <details>
695
- <summary>What providers does llm.rb support?</summary>
843
+ <summary>What about local LLM support?</summary>
696
844
  <br>
697
845
  <p>
698
-
699
- **Cloud**
700
-
701
- The following cloud-based providers are available to choose from. <br>
702
- In no particular order:
703
-
704
- 🇺🇸 OpenAI <br>
705
- 🇺🇸 DeepInfra <br>
706
- 🇺🇸 xAI <br>
707
- 🇺🇸 Google (Gemini) <br>
708
- 🇺🇸 AWS bedrock <br>
709
- 🇺🇸 Anthropic <br>
710
- 🇨🇳 DeepSeek <br>
711
- 🇨🇳 zAI <br>
712
- 🇨🇳 Moonshot AI (Kimi) <br>
713
- 🇪🇺 Mistral <br>
714
-
715
- **Weights**
716
-
717
- The following providers provide access to open-weight models. <br>
718
- In no particular order:
719
-
720
- 🇺🇸 DeepInfra <br>
721
- 🇺🇸 AWS bedrock <br>
722
- 🇨🇳 DeepSeek <br>
723
- 🇨🇳 zAI <br>
724
- 🇨🇳 Moonshot AI (Kimi) <br>
725
- 🇪🇺 Mistral <br>
726
-
727
- **Local**
728
-
729
- The following providers can be run locally on your own hardware. <br>
730
- In no particular order:
846
+ The following providers can be run used with models that
847
+ are running on your own hardware. They're reasonably well
848
+ tested but not my main driver:
849
+ </p>
731
850
 
732
851
  * Ollama
733
852
  * Llamacpp
734
- </p>
735
853
  </details>
736
854
 
737
855
  <details>
@@ -763,11 +881,10 @@ type.
763
881
  If you're on a budget, DeepSeek is hard to beat.
764
882
  </details>
765
883
  <details>
766
- <summary>Can I download llm.rb via a decentralized network?</summary>
767
- <br>
768
- Yes.
884
+ <summary>Sources other than GitHub?</summary>
769
885
  <br>
770
- We are on the <a href="https://radicle.network">radicle.network</a>
886
+ <p>
887
+ We are on the <a href="https://radicle.network">radicle.network</a> as well.
771
888
  <br>
772
889
  Every commit that lands on GitHub also lands on Radicle.
773
890
  <br>
@@ -776,6 +893,30 @@ Our repository ID is z2PtfQ6dYwyYaW2aGrztG1sMyDmCE.
776
893
  Browse on <a
777
894
  href="https://radicle.network/nodes/iris.radicle.network/z2PtfQ6dYwyYaW2aGrztG1sMyDmCE">the
778
895
  web</a>.
896
+ </p>
897
+ </details>
898
+
899
+ <details>
900
+ <summary>Who maintains llm.rb?</summary>
901
+ <br>
902
+
903
+ The llm.rb project is maintained primarily by one
904
+ person. llm.rb has been in active development for more
905
+ than three years and over that time multiple other
906
+ contributors have contributed to llm.rb as well. New
907
+ contributors are always welcome.
908
+
909
+ I use the repl that is distributed with llm.rb to build
910
+ llm.rb itself so there is a healthy feedback loop and
911
+ llm.rb has also been battle tested in production
912
+ environments.
913
+
914
+ I have also also written llm.rb agents within the
915
+ repository that help me maintain the documentation,
916
+ and backport changes to the mruby-llm runtime as well.
917
+
918
+ I am constantly focused on improving llm.rb by using
919
+ it as my primary driver for development.
779
920
  </details>
780
921
 
781
922
  ## Resources
@@ -786,41 +927,6 @@ wasn't possible to cover every feature without the README becoming a small book.
786
927
  The [r.uby.dev](https://r.uby.dev) homepage also includes more learning material
787
928
  and resources.
788
929
 
789
- ## Developers
790
-
791
- The llm.rb project is quite large and maintained primarily by one
792
- person. It would be near impossible for me to maintain both the codebase
793
- and its documentation, especially the [deepdive.md](https://r.uby.dev/llm/deepdive/)
794
- so I have written agents that maintain the documentation assets and that
795
- allows me to put more focus on the code.
796
-
797
- The following agents are available for those tasks, and all of them
798
- use the most cost effective option: DeepSeek. Feel free to use them
799
- in your own fork.
800
-
801
- ```sh
802
- ##
803
- # Maintains the deepdive and API docs
804
- rake agents:scribe:yardoc
805
- rake agents:scribe:coverage
806
- rake agents:scribe:regressions
807
- rake agents:scribe:style
808
-
809
- ##
810
- # Maintains the release
811
- rake agents:dexter:changelog
812
- rake agents:dexter:release
813
-
814
- ##
815
- # Maintains mruby-llm backports
816
- rake agents:mruby:research
817
- rake agents:mruby:implement
818
-
819
- ##
820
- # Refresh the data/ registry
821
- rake models.dev:download
822
- ```
823
-
824
930
  ## License
825
931
 
826
932
  This software is released under the terms of the MIT license. <br>