llm.rb 13.0.0 → 13.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -14,1807 +14,67 @@
14
14
 
15
15
  ## Welcome
16
16
 
17
- Welcome to the llm.rb deepdive. You are reading this document
18
- in the markdown format. An optimized version exists
19
- at [https://r.uby.dev/llm/deepdive](https://r.uby.dev/llm/deepdive)
20
- and it is both easier to read and navigate.
21
-
22
- This document is a continuation of the [homepage documentation](https://r.uby.dev/llm).
23
- It assumes you are familiar with the basics already, and focuses on
24
- features that didn't make it into the homepage documentation.
25
-
26
- ## Table of contents
27
-
28
- **Overview**
29
-
30
- - [Welcome](#welcome)
31
-
32
- **Core**
33
-
34
- <details>
35
- <summary>Agents</summary>
36
-
37
- - [As a subclass](#as-a-subclass)
38
- - [As an object](#as-an-object)
39
- </details>
40
-
41
- <details>
42
- <summary>Tools</summary>
43
-
44
- - [LLM::Tool](#llmtool)
45
- - [Errors](#errors)
46
- - [Confirmation](#confirmation)
47
- - [Manual tool loop](#manual-tool-loop)
48
- - [Executing](#executing)
49
- - [Per-tool confirmation](#per-tool-confirmation)
50
- - [Full loop](#full-loop)
51
- - [Trade-offs](#trade-offs)
52
- </details>
53
-
54
- <details>
55
- <summary>Skills</summary>
56
-
57
- - [SKILL.md](#skillmd)
58
- - [Run it](#run-it)
59
- </details>
60
-
61
- <details>
62
- <summary>Schema</summary>
63
-
64
- - [Estimation](#estimation)
65
- </details>
66
-
67
- **Runtime**
68
-
69
- <details>
70
- <summary>Stream</summary>
71
-
72
- - [IO-like object](#io-like-object)
73
- - [LLM::Stream](#llmstream)
74
- </details>
75
-
76
- <details>
77
- <summary>Concurrency</summary>
78
-
79
- - [Overview](#overview)
80
- - [sequential](#sequential)
81
- - [thread](#thread)
82
- - [fiber](#fiber)
83
- - [async](#async)
84
- - [fork](#fork)
85
- - [ractor](#ractor)
86
- - [Quick reference](#quick-reference)
87
- </details>
88
-
89
- <details>
90
- <summary>Context Compaction</summary>
91
-
92
- - [Configuration](#configuration)
93
- - [Standalone usage](#standalone-usage)
94
- - [Strategies](#strategies)
95
- - [Manual compaction](#manual-compaction)
96
- - [Lifecycle callbacks](#lifecycle-callbacks)
97
- </details>
98
-
99
- <details>
100
- <summary>Cancellation</summary>
101
-
102
- - [Cancel a request](#cancel-a-request)
103
- - [Tool interrupts](#tool-interrupts)
104
- </details>
105
-
106
- <details>
107
- <summary>Transports</summary>
108
-
109
- - [net/http](#nethttp)
110
- - [net/http/persistent](#nethttppersistent)
111
- - [curb](#curb)
112
- </details>
113
-
114
- <details>
115
- <summary>Tracer</summary>
116
-
117
- - [Provider-wide tracer](#provider-wide-tracer)
118
- - [Agent-local tracer](#agent-local-tracer)
119
- </details>
120
-
121
- <details>
122
- <summary>REPL</summary>
123
-
124
- - [LLM::Agent](#llmagent)
125
- - [Persistence](#persistence)
126
- - [Tools](#tools)
127
- - [Skills](#skills-1)
128
- - [Tracer](#tracer-1)
129
- - [Input](#input)
130
- - [Commands](#commands)
131
- </details>
132
-
133
- **Persistence**
134
-
135
- <details>
136
- <summary>Serialization</summary>
137
-
138
- - [Save to disk](#save-to-disk)
139
- </details>
140
-
141
- <details>
142
- <summary>ORM</summary>
143
-
144
- - [ActiveRecord](#activerecord)
145
- - [Sequel](#sequel)
146
- </details>
147
-
148
- **Media**
149
-
150
- <details>
151
- <summary>Images</summary>
152
-
153
- - [Generation](#generation)
154
- - [Edits](#edits)
155
- - [DeepSeek](#deepseek)
156
- </details>
157
-
158
- <details>
159
- <summary>Audio</summary>
160
-
161
- - [text-to-speech](#text-to-speech)
162
- - [speech-to-text](#speech-to-text)
163
- - [translation](#translation)
164
- </details>
165
-
166
- <details>
167
- <summary>OCR</summary>
168
-
169
- - [Mistral](#mistral)
170
- </details>
171
-
172
- **Protocols**
173
-
174
- <details>
175
- <summary>MCP</summary>
176
-
177
- - [stdio](#stdio)
178
- - [http](#http)
179
- </details>
180
-
181
- <details>
182
- <summary>A2A</summary>
183
-
184
- - [rest](#rest)
185
- - [jsonrpc](#jsonrpc)
186
- </details>
187
-
188
- ## Agents
189
-
190
- An agent is represented by the
191
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
192
- class, and it is built on top of
193
- [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) -
194
- the heart of the runtime. An agent manages the tool loop automatically,
195
- implements a tool loop guard for misbehaving models, and
196
- it can use six different concurrency strategies to execute
197
- tools.
198
-
199
- An agent can be a subclass of
200
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html),
201
- or a direct
202
- instance of it. The subclass approach is useful when you
203
- want reusable agents that can attach behavior (as methods)
204
- to their own class.
205
-
206
- #### As a subclass
207
-
208
- A subclass of
209
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
210
- can define its model, tools,
211
- and other attributes at the class-level. All of these
212
- attributes are optional, and they act as defaults that
213
- can be overriden on the instance level.
214
-
215
- The example uses the `:fork` concurrency model. It has
216
- two primary benefits: tools are run in parallel, and in
217
- a separate process with a separate memory address space.
218
-
219
- The example purposefully demonstrates how the attributes
220
- can be lazily defined with a block, or a Symbol that is
221
- evaluated as an instance method on the subclass. It is
222
- not strictly neccessary, though, and the example would
223
- be simpler without it.
224
-
225
- ```ruby
226
- class Agent < LLM::Agent
227
- set model: "deepseek-v4-pro",
228
- tools: [DoResearch, FinalizeResearch, ActOnResearch],
229
- stream: -> { $stdout },
230
- tracer: :set_tracer,
231
- concurrency: :fork
232
-
233
- def research!
234
- talk "start the research"
235
- end
236
-
237
- private
238
-
239
- def set_tracer
240
- LLM::Tracer::Logger.new(llm, io: $stderr)
241
- end
242
- end
243
- llm = LLM.deepseek(key: ENV["KEY"])
244
- agent = Agent.new(llm).tap(&:research!)
245
- agent.talk "How did the research go?"
246
- ```
247
-
248
- #### As an object
249
-
250
- The more direct, and sometimes more convienent approach, is to
251
- create an instance of
252
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
253
- directly. The same attributes can be provided as the
254
- second argument given to
255
- [`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html),
256
- and the same lazy evaluation rules apply. This approach can be
257
- great for prototyping quickly, and you can always turn to a
258
- subclass later if that makes more sense.
259
-
260
- ```ruby
261
- llm = LLM.deepseek(key: ENV["KEY"])
262
- agent = LLM::Agent.new(llm, stream: $stdout)
263
- agent.talk "Hello, fellow agent"
264
- ```
265
-
266
- [Back to top](#table-of-contents)
267
-
268
- ## Tools
269
-
270
- A tool extends the capabilities of a model. <br>
271
- A tool is a subclass of
272
- [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
273
- that has a name,
274
- a description, and an optional set of typed parameters.
275
-
276
- A tool also has a method associated with it, and when the
277
- model calls a tool it will do so through this method &ndash;
278
- alongside any parameters the tool might have defined.
279
-
280
- In other words, a tool provides a way for a model to
281
- call a method you have written, and it returns a value
282
- to the model that is considered the tool's response.
283
- The model then proceeds to process the tool's response,
284
- and then might generate its own response, or perhaps call
285
- another tool.
286
-
287
- There is exactly one rule: a tool call must always produce
288
- a tool response. If a tool raises an exception, the runtime
289
- rescues it and returns a structured error to the model
290
- instead. The conversation never enters an invalid state
291
- because of a crashed tool &mdash; the model always has
292
- something to work with. This is by design. Keeping the
293
- tool loop alive is the highest priority.
294
-
295
- #### LLM::Tool
296
-
297
- A tool can be defined by subclassing
298
- [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
299
- with
300
- a name, description, and optional set of parameters. The
301
- tool name, and description should be informative so the
302
- model can understand what the tool does and how it can
303
- serve a user's query.
304
-
305
- ```ruby
306
- require "llm"
307
- require "shellwords"
308
-
309
- class Shell < LLM::Tool
310
- name "shell"
311
- description "execute a shell command"
312
- parameter :name, String, "the command's name"
313
- parameter :arguments, Array[String], "One or more arguments"
314
- required %i[name]
315
- defaults arguments: []
316
-
317
- def call(name:, arguments: [])
318
- out = `#{name.shellescape} #{arguments.map(&:shellescape).join(" ")}`
319
- {ok: $?.success?, out:}
320
- end
321
- end
322
-
323
- llm = LLM.deepseek(key: ENV["KEY"])
324
- agent = LLM::Agent.new(llm, tools: [Shell], stream: $stdout)
325
- agent.talk "What files are in the current working directory?"
326
- ```
327
-
328
- #### Errors
329
-
330
- Exceptions raised by a tool are automatically rescued and
331
- returned to the model as a structured error. The model sees
332
- something like this:
333
-
334
- ```ruby
335
- class Error < LLM::Tool
336
- name "error"
337
- description "demo how errors are handled"
338
-
339
- ##
340
- # Returns
341
- # {error: true, kind: "RuntimeError", message: "boom"}
342
- def call
343
- raise "boom"
344
- end
345
- end
346
- ```
347
-
348
- The runtime wraps the exception into `{error: true, kind: "RuntimeError",
349
- message: "boom"}` and returns it to the model as the tool response. From
350
- the model's perspective the tool completed &mdash; it just completed with
351
- an error. The model can read the error, decide what went wrong, and try
352
- something else. The conversation stays valid.
353
-
354
- You can also handle errors yourself inside `call`. Rescue the exception
355
- and return whatever shape makes sense for your tool:
356
-
357
- ```ruby
358
- class Shell < LLM::Tool
359
- name "shell"
360
- description "execute a shell command"
361
-
362
- def call(name:, arguments: [])
363
- out = `#{name} #{arguments.join(" ")}`
364
- {ok: $?.success?, out:}
365
- rescue Errno::ENOENT
366
- {ok: false, error: "command not found: #{name}"}
367
- end
368
- end
369
- ```
370
-
371
- The model receives `{ok: false, error: "command not found: ls"}` and can
372
- react accordingly &mdash; maybe it corrects the command name and tries
373
- again. This is often better than letting the runtime's generic error
374
- wrapper speak for you, because you can provide domain-specific detail
375
- that helps the model recover.
376
-
377
- The principle is the same either way: **return something**. A tool call
378
- must complete with a tool response. If you don't return a value, and you
379
- don't raise, the runtime has nothing to send back and the conversation
380
- is stuck. As long as you return a Hash (or anything the model can
381
- interpret), the tool loop continues.
382
-
383
- #### Confirmation
384
-
385
- Tools that perform destructive actions can be gated behind
386
- explicit confirmation. List their names in
387
- [`LLM::Agent.confirm`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#confirm-class_method)
388
- to block execution until you override `on_tool_confirmation`.
389
-
390
- The default handler cancels the tool. Override it per-agent to
391
- prompt the user, log the decision, or auto-approve certain tools.
392
-
393
- ```ruby
394
- class AdminAgent < LLM::Agent
395
- set confirm: %w[delete destroy shutdown]
396
-
397
- def on_tool_confirmation(fn, strategy)
398
- print "Run #{fn.name} with #{fn.arguments}? [y/N] "
399
- $stdin.gets&.match?(/\Ay\z/i) ? wait(strategy) : fn.cancel
400
- end
401
- end
402
-
403
- llm = LLM.deepseek(key: ENV["KEY"])
404
- agent = AdminAgent.new(llm)
405
- ```
406
-
407
- Confirmation also accepts a Symbol for lazy resolution, which
408
- allows the list of confirmed tools to change per-instance:
409
-
410
- ```ruby
411
- class AdaptiveAgent < LLM::Agent
412
- set confirm: :tools_that_need_confirmation
413
-
414
- def tools_that_need_confirmation
415
- some_condition ? %w[delete destroy] : %w[delete]
416
- end
417
- end
418
- ```
419
-
420
- ## Manual tool loop
421
-
422
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html) manages
423
- the tool loop automatically — it calls the model, checks for tool calls,
424
- executes them, sends results back, and repeats until the model responds
425
- with text. You can bypass this and drive the loop yourself for finer
426
- control.
427
-
428
- Start with a bare [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
429
- (no agent):
430
-
431
- ```ruby
432
- require "llm"
433
-
434
- llm = LLM.deepseek(key: ENV["KEY"])
435
- ctx = LLM::Context.new(llm)
436
- ```
437
-
438
- Send a message and check whether the model wants to call tools:
439
-
440
- ```ruby
441
- res = ctx.talk "What's the weather in Tokyo?"
442
-
443
- if ctx.pending_functions?
444
- puts "Model requested #{ctx.pending_functions.size} tool(s)"
445
- else
446
- puts res.text
447
- end
448
- ```
17
+ ### Introduction
449
18
 
450
- ### Executing
19
+ #### Overview
451
20
 
452
- When the model asks to call tools, call `ctx.wait(:strategy)` to run
453
- them. `ctx.wait` picks up pending functions, spawns them, waits for
454
- results, and records them back in the context. Every strategy is
455
- supported: `:sequential`, `:thread`, `:fiber`, `:async`, `:fork`,
456
- and `:ractor`.
21
+ Welcome to the llm.rb deepdive. This document is a continuation
22
+ of the [homepage documentation](https://r.uby.dev/llm). It assumes
23
+ you are familiar with the basics already, and focuses on features
24
+ that didn't make it into the homepage documentation.
457
25
 
458
- ```ruby
459
- require "llm"
460
-
461
- llm = LLM.deepseek(key: ENV["KEY"])
462
- ctx = LLM::Context.new(llm)
463
-
464
- ctx.talk("What's the weather in Tokyo?")
465
- ctx.talk ctx.wait(:thread)
466
- ```
467
-
468
- ### Per-tool confirmation
469
-
470
- Because `pending_functions` returns a regular array, you can inspect
471
- each function before execution. Call `ctx.wait(:thread)` to execute
472
- all pending tools — but you can also selectively exclude functions or
473
- run individual ones through `fn.task(:thread).wait` for ad-hoc execution
474
- that bypasses guards and streaming hooks:
475
-
476
- ```ruby
477
- results = ctx.pending_functions.map do |fn|
478
- print "Run #{fn.name} with #{fn.arguments}? [y/N] "
479
- if $stdin.gets&.match?(/\Ay\z/i)
480
- fn.task(:thread).wait
481
- else
482
- fn.cancel(reason: "user declined")
483
- end
484
- end
485
- ctx.talk(results)
486
- ```
487
-
488
- This pattern is what
489
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)'s
490
- built-in confirmation feature uses internally, but doing it manually
491
- gives you full control — route the decision through a web socket, a
492
- background job, or a multi-user approval flow.
493
-
494
- ### Full loop
495
-
496
- A complete manual tool loop looks like this:
497
-
498
- ```ruby
499
- require "llm"
500
-
501
- llm = LLM.deepseek(key: ENV["KEY"])
502
- ctx = LLM::Context.new(llm)
503
- res = nil
504
-
505
- loop do
506
- res = ctx.talk("What's the weather in Tokyo?")
507
- break unless ctx.pending_functions?
508
-
509
- results = ctx.wait(:thread)
510
- ctx.talk(results)
511
- end
512
-
513
- puts res.content
514
- ```
515
-
516
- The loop exits when the model responds with text rather than tool
517
- calls. You can extend it with timeouts, user confirmation gates, or
518
- custom error handling for each tool.
519
-
520
- ### Trade-offs
521
-
522
- Manual control is more code but gives you:
523
-
524
- - **Arbitrary pre-execution checks** — inspect, rewrite, or skip tool
525
- calls before they run.
526
- - **Custom confirmation flows** — async approval over HTTP, Slack,
527
- or email instead of the built-in terminal prompt.
528
- - **Different strategies per tool** — run one tool on a thread and
529
- another in a forked process within the same turn.
530
- - **Fine-grained error recovery** — rescue per-tool failures and
531
- decide which results to feed back.
26
+ An optimized version exists
27
+ at [https://r.uby.dev/llm/deepdive](https://r.uby.dev/llm/deepdive)
28
+ that is both easier to read and navigate.
532
29
 
533
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html) handles
534
- all of this automatically and is the right choice for most applications.
535
- Drop down to the manual loop when you need control that the agent
536
- abstraction doesn't expose.
30
+ #### How it works
537
31
 
538
- ## Skills
32
+ Each topic file follows a consistent four-part pattern:
33
+ `#### Overview` introduces the concept, `#### How it works` shows
34
+ code, `#### Why would I use it?` explains the use case, and
35
+ `#### Notes` covers caveats and edge cases.
539
36
 
540
- The skill concept is borrowed from tools like Claude and
541
- Codex, but llm.rb gives it a runtime of its own. A skill
542
- is a directory with a `SKILL.md` file. That file contains
543
- frontmatter where the skill's name, description, and tools
544
- can be declared.
37
+ #### Why would I use it?
545
38
 
546
- #### SKILL.md
39
+ The deepdive documents everything the homepage leaves out:
40
+ advanced patterns, configuration options, ORM integrations,
41
+ protocol support, and edge cases. Read it when you need to go
42
+ beyond the basics.
547
43
 
548
- The `SKILL.md` file can look like this. When a skill runs,
549
- the runtime spawns a subagent with its own context window
550
- and message history. Some context is inherited from the
551
- parent agent, though.
44
+ #### Notes
552
45
 
553
- By default the subagent can only access the tools declared
554
- by the skill. The `inherit` directive lets it inherit the
555
- parent agent's tools instead, including A2A and MCP tools.
46
+ The deepdive is a living document. Sections are added as new
47
+ features land. The [homepage](https://r.uby.dev/llm) is the best
48
+ place to start if you are new to llm.rb.
556
49
 
557
- ```markdown
558
- ---
559
- name: git-skill
560
- description: reads my git history and writes a summary
561
- tools: ['git-log', 'git-show', 'write-file']
562
50
  ---
563
51
 
564
- ## Task
565
-
566
- Collect a log of recent history.
567
- Analyze each commit.
568
- Write a summary to summary.txt
569
- ```
570
-
571
- #### Run it
572
-
573
- Given the skill above, llm.rb only needs the path to the
574
- directory that contains `SKILL.md`. Under the hood, a skill
575
- is represented as a tool the model can call. That means
576
- a skill can be called whenever it satisfies the user's
577
- request &ndash; in the same way that a regular tool can.
578
-
579
- This feature also works with both the ActiveRecord, and
580
- Sequel integrations.
581
-
582
- ```ruby
583
- require "llm"
584
-
585
- llm = LLM.deepseek(key: ENV["KEY"])
586
- agent = LLM::Agent.new(llm, skills: [__dir__])
587
- agent.talk "run the git skill"
588
- ```
589
-
590
- [Back to top](#table-of-contents)
591
-
592
- ## MCP
593
-
594
- #### stdio
595
-
596
- The stdio transport connects to an MCP server that is launched as a
597
- separate process, and both its standard input and standard output
598
- streams are used for communication. It is recommended but not
599
- required to execute commands for a stdio transport over a
600
- persistent session via the
601
- [`LLM::MCP#session`](https://r.uby.dev/api-docs/llm.rb/LLM/MCP.html#session-instance_method)
602
- method &ndash; otherwise
603
- you could end up launching the same process multiple times.
604
-
605
- ```ruby
606
- require "llm"
607
-
608
- llm = LLM.deepseek(key: ENV["KEY"])
609
- mcp = LLM::MCP.stdio(argv: ["npx", "-y", "@forgejo/mcp-server"])
610
- agent = LLM::Agent.new(llm)
611
-
612
- mcp.session do
613
- agent.talk "What's happening on forgejo?", tools: mcp.tools
614
- end
615
- ```
616
-
617
- #### http
618
-
619
- The http transport connects to an MCP server over HTTP, and unlike
620
- the stdio transport, the MCP server does not have to be running
621
- locally. Popular services like GitHub provide their own MCP server
622
- over HTTP, and it is one of the most capable MCP servers I have
623
- used.
624
-
625
- Unlike the stdio transport,
626
- [`LLM::MCP#session`](https://r.uby.dev/api-docs/llm.rb/LLM/MCP.html#session-instance_method)
627
- carries little benefit for the http transport and it can be
628
- omitted. It is recommended to consider the `net_http_persistent`
629
- transport for MCP interactions that run over HTTP, otherwise
630
- you could end up tearing down and setting up the same connection
631
- multiple times.
632
-
633
- ```ruby
634
- require "llm"
635
-
636
- llm = LLM.deepseek(key: ENV["KEY"])
637
- mcp = LLM::MCP.http(
638
- url: "https://api.githubcopilot.com/mcp/",
639
- headers: {
640
- "Authorization" => "Bearer #{ENV.fetch('GITHUB_PAT')}"
641
- },
642
- transport: :net_http_persistent
643
- )
644
- agent = LLM::Agent.new(llm)
645
- agent.talk "What's happening on GitHub?", tools: mcp.tools
646
- ```
647
-
648
- [Back to top](#table-of-contents)
649
-
650
- ## A2A
651
-
652
- #### rest
653
-
654
- The rest transport communicates with other agents via A2A
655
- endpoints that speak both HTTP and JSON. The skills advertised
656
- by an agent become subclasses of
657
- [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
658
- that can be used by both
659
- [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html),
660
- and [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
661
- &ndash; similar to how MCP tools become subclasses of
662
- [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html).
663
-
664
- ```ruby
665
- require "llm"
666
-
667
- llm = LLM.deepseek(key: ENV["KEY"])
668
- a2a = LLM::A2A.rest(url: "https://agent.example.com")
669
- agent = LLM::Agent.new(llm, tools: a2a.skills)
670
- agent.talk "What's happening, fellow agent?"
671
- ```
672
-
673
- #### jsonrpc
674
-
675
- The jsonrpc transport communicates with other agents via HTTP
676
- and a protocol known as jsonrpc. Sometimes an agent will
677
- implement both, or just one of each. An agent's card, which
678
- is represented by an instance of
679
- [`LLM::A2A::Card`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A/Card.html),
680
- can be
681
- used to discover available transports via the
682
- [`LLM::A2A::Card#interfaces`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A/Card.html#interfaces-instance_method)
683
- method.
684
-
685
- ```ruby
686
- require "llm"
687
- llm = LLM.deepseek(key: ENV["KEY"])
688
- a2a = LLM::A2A.jsonrpc(url: "https://agent.example.com")
689
- agent = LLM::Agent.new(llm, tools: a2a.skills)
690
- agent.talk "What's happening, fellow agent?"
691
- ```
692
-
693
- [Back to top](#table-of-contents)
694
-
695
- ## Transports
696
-
697
- The [`LLM::Provider`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html),
698
- [`LLM::MCP`](https://r.uby.dev/api-docs/llm.rb/LLM/MCP.html), and
699
- [`LLM::A2A`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html) classes
700
- all accept a `transport` option that decides which library
701
- will be used for HTTP communication. There are three options out
702
- of the box:
703
- [`net-http`](https://github.com/ruby/net-http),
704
- [`net-http-persistent`](https://github.com/drbrain/net-http-persistent),
705
- and [`curb`](https://github.com/taf2/curb).
706
-
707
- #### net/http
708
-
709
- The [`net/http`](https://github.com/ruby/net-http) transport is represented by the symbol `:net_http`. <br>
710
- It is the default transport.
711
-
712
- ```ruby
713
- require "llm"
714
-
715
- llm = LLM.deepseek(key: "...", transport: :net_http)
716
- mcp = LLM::MCP.http(url: "...", transport: :net_http)
717
- a2a = LLM::A2A.rest(url: "...", transport: :net_http)
718
- ```
719
-
720
- #### net/http/persistent
721
-
722
- The [`net/http/persistent`](https://github.com/drbrain/net-http-persistent) transport is represented by the symbol `:net_http_persistent`. <br>
723
- It maintains a connection pool so the cost of tearing down and
724
- setting up a connection repeatedly is kept low, and it is built
725
- on top of [`net/http`](https://github.com/ruby/net-http).
726
-
727
- ```ruby
728
- require "llm"
729
-
730
- llm = LLM.deepseek(key: "...", transport: :net_http_persistent)
731
- mcp = LLM::MCP.http(url: "...", transport: :net_http_persistent)
732
- a2a = LLM::A2A.rest(url: "...", transport: :net_http_persistent)
733
- ```
734
-
735
- #### curb
736
-
737
- The [`curb`](https://github.com/taf2/curb) transport is represented by the symbol `:curb`. <br>
738
- It provides bindings for libcurl &ndash; a widely used, highly portable
739
- and feature-rich HTTP library written in C.
740
-
741
- ```ruby
742
- require "llm"
743
-
744
- llm = LLM.deepseek(key: "...", transport: :curb)
745
- mcp = LLM::MCP.http(url: "...", transport: :curb)
746
- a2a = LLM::A2A.rest(url: "...", transport: :curb)
747
- ```
748
-
749
- [Back to top](#table-of-contents)
750
-
751
- ## Stream
752
-
753
- #### IO-like object
754
-
755
- Any object that implements the `#<<` method can receive
756
- chunks from a stream. That includes objects like `$stdout`.
757
- This form of streaming is simple and limited. It is the
758
- equivalent of
759
- [`LLM::Stream#on_content`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_content-instance_method),
760
- and doesn't include
761
- any of the other
762
- [`LLM::Stream`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
763
- hooks.
764
-
765
- ```ruby
766
- require "llm"
767
-
768
- llm = LLM.deepseek(key: ENV["KEY"])
769
- agent = LLM::Agent.new(llm, stream: $stdout)
770
- agent.talk "hello world"
771
- ```
772
-
773
- #### LLM::Stream
774
-
775
- The [`LLM::Stream`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html)
776
- class provides many hooks that a subclass
777
- can implement. They range from being notified when a tool call
778
- starts to when a tool call finishes, or when a conversation is
779
- due to be compacted because the context window exceeded a defined
780
- limit. All these callbacks support a responsive user interface
781
- where the user is always aware of what is happening behind the
782
- scenes.
783
-
784
- ```ruby
785
- class Stream < LLM::Stream
786
- def on_content(content)
787
- puts content
788
- end
789
-
790
- def on_reasoning_content(content)
791
- puts content
792
- end
793
-
794
- def on_tool_call(tool)
795
- # this callback can be used to either log a tool call,
796
- # or execute a tool call during a stream.
797
- end
798
-
799
- def on_tool_return(tool, result)
800
- end
801
-
802
- def on_compaction(compactor)
803
- # this callback is called *before* a compact happens
804
- end
805
-
806
- def on_compaction_finish(compactor)
807
- # this callback is called *after* a compact happens
808
- end
809
- end
810
- ```
811
-
812
- [Back to top](#table-of-contents)
813
-
814
- ## Concurrency
815
-
816
- llm.rb supports six concurrency strategies for tool execution &ndash;
817
- `:sequential`, `:thread`, `:fiber`, `:async`, `:fork`, and `:ractor`.
818
- Each one implements the same interface &mdash; `spawn`, `wait`, `alive?`,
819
- `interrupt!` &mdash; so the caller never has to care which strategy is
820
- behind a given task.
821
-
822
- Choose a strategy per-agent or per-call:
823
-
824
- ```ruby
825
- ## Per-agent — every tool loop uses :fork
826
- agent = LLM::Agent.new(llm, concurrency: :fork, tools: [...])
827
-
828
- ## Per-call — run a single tool on a thread
829
- fn = FetchStocks.function
830
- fn.task(:thread).wait
831
- ```
832
-
833
- Interruption is reliable across all six. No matter the backing &mdash;
834
- thread, fiber, process, ractor &mdash; `LLM::Interrupt` reaches the
835
- tool and it can rescue, clean up, and either re-raise to cancel the
836
- turn or return a value to continue.
837
-
838
- #### sequential
839
-
840
- The default. Tools run one at a time on the calling thread. No
841
- concurrency, no overhead. `spawn` is a no-op &mdash; execution happens
842
- in `wait`. `alive?` always returns `false`.
843
-
844
- Best for simple agents with a tool or two, debugging, or when tool
845
- order matters.
846
-
847
- #### thread
848
-
849
- Each tool runs in its own `Thread`. The thread is created lazily &mdash;
850
- you can build a task, pass it around, and decide when to run it.
851
- Threads have `report_on_exception` disabled so errors surface through
852
- `wait` rather than stderr.
853
-
854
- Interruption raises `LLM::Interrupt` directly on the tool's thread,
855
- which stops it mid-flight.
856
-
857
- Best for IO-bound tools &mdash; HTTP calls, database queries. CRuby
858
- releases the GVL during blocking IO, so you get real concurrency.
859
-
860
- #### fiber
861
-
862
- Each tool runs in a scheduler-backed `Fiber` via `Fiber.schedule`.
863
- Requires `Fiber.scheduler` &mdash; raises `ArgumentError` without one.
864
- Fibers yield cooperatively at IO boundaries, so this pairs well with
865
- async libraries that set a scheduler.
866
-
867
- Interruption raises `LLM::Interrupt` on the fiber, which stops at
868
- the next yield point.
869
-
870
- Best for IO-bound tools inside an async framework. Much lighter than
871
- threads.
872
-
873
- #### async
874
-
875
- Each tool runs as an `Async::Task` inside a managed background
876
- reactor. A dedicated thread runs an `Async::Reactor` event loop.
877
- Work is submitted through a thread-safe `Queue` inbox and consumed
878
- by the reactor. All fibers stay on one thread &mdash; no shared-memory
879
- contention between them.
880
-
881
- The reactor is created on demand and shared across all tasks in a
882
- group. When `Group#wait` is called, tasks are submitted, the reactor
883
- runs them concurrently, and results are bridged back to the caller
884
- through per-task queues. The reactor is torn down after `wait`
885
- completes.
886
-
887
- Interruption pushes an `LLM::Interrupt` sentinel into the task's
888
- result queue instead of using `Fiber#raise` &mdash; cleaner, and it
889
- avoids surprising the reactor's internal fibers.
890
-
891
- Best for IO-bound tools when you want Async's structured concurrency
892
- model without running your whole application inside a reactor. The
893
- reactor is self-contained &mdash; your main thread stays synchronous.
894
- Requires the `async` gem.
895
-
896
- #### fork
897
-
898
- Each tool runs in a forked child process. Communication uses
899
- [`xchan`](https://github.com/1robertrb/xchan.rb) (marshal-based
900
- channels): the parent sends control messages, the child sends
901
- results back. Each child is a separate OS process with its own
902
- memory space &mdash; a crash in the tool cannot touch the parent.
903
-
904
- `Fork::Task` checks liveness with `Process.waitpid(WNOHANG)` and
905
- delivers interrupts as messages over the control channel. The child
906
- raises `LLM::Interrupt` on `Thread.main` when it receives the
907
- interrupt message. Tracer callbacks fire in both parent and child.
908
-
909
- Best for process isolation &mdash; shell commands, native extensions,
910
- anything you don't want touching the parent's memory. True parallelism
911
- too, since there's no GVL in separate processes. Requires the
912
- `xchan` gem.
913
-
914
- #### ractor
915
-
916
- Each class-based tool runs in a Ruby `Ractor`. `Ractor::Task`
917
- coordinates through `Ractor::Mailbox`. Interruption sends a message
918
- through the mailbox; a listener thread inside the ractor raises
919
- `LLM::Interrupt` on `Thread.main`.
920
-
921
- Ractors have restrictions: only class-based tools are supported (no
922
- blocks, skills, or MCP tools), and arguments must be
923
- ractor-shareable. The runtime raises `LLM::RactorError` early if
924
- you try to run an unsupported tool type.
925
-
926
- Best for CPU-bound tools, true parallelism without the overhead of
927
- forking full processes. More restrictive than `:fork` but lighter.
928
-
929
- #### Quick reference
930
-
931
- | Strategy | Backing | Parallel? | Isolation? | Requires |
932
- |---|---|---|---|---|
933
- | `:sequential` | direct call | No | No | &mdash; |
934
- | `:thread` | `Thread` | IO only (GVL) | No | &mdash; |
935
- | `:fiber` | `Fiber.schedule` | Cooperative | No | `Fiber.scheduler` |
936
- | `:async` | `Async::Reactor` on bg thread | Cooperative | No | `async` gem |
937
- | `:fork` | `Kernel.fork` | Yes (process) | Yes (memory) | `xchan` gem |
938
- | `:ractor` | `Ractor` | Yes (CPU) | Limited | &mdash; |
939
-
940
- [Back to top](#table-of-contents)
941
-
942
- ## Context Compaction
943
-
944
- Long-running conversations consume tokens. Without intervention, every turn
945
- pushes toward the model's context window limit, at which point the provider
946
- rejects the request.
947
-
948
- llm.rb provides compaction through pluggable strategies. All strategies
949
- inherit from [`LLM::Compactor`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor.html)
950
- and are invoked automatically before each `ctx.talk(...)` call.
951
-
952
- ### Configuration
953
-
954
- Pass a compactor class and options when creating a context or agent:
955
-
956
- ```ruby
957
- ctx = LLM::Context.new(
958
- llm,
959
- compactor: LLM::Compactor::Truncate,
960
- compactor_options: {keep: 64}
961
- )
962
-
963
- # LLM::Agent accepts the same options
964
- agent = LLM::Agent.new(
965
- llm,
966
- compactor: LLM::Compactor::Truncate,
967
- compactor_options: {keep: 128}
968
- )
969
- ```
970
-
971
- The compactor runs automatically before every `talk` call. This keeps the
972
- conversation constantly alive — there is no chance of exhausting the context
973
- window because old messages are dropped before they accumulate. The trade-off
974
- is that dropped messages are gone, so information may be lost. Set `keep` to
975
- a higher number to retain more context at the cost of slower accumulation.
976
-
977
- The default is [`LLM::Compactor::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Null.html)
978
- — compaction is disabled unless you opt in.
979
-
980
- ### Standalone usage
981
-
982
- A compactor can be used independently of a context or agent:
983
-
984
- ```ruby
985
- compactor = LLM::Compactor::Truncate.new(agent)
986
- compactor.call(keep: 200) # or ctx, agent, etc.
987
- ```
988
-
989
- This is useful for one-off compaction outside the automatic per-turn cycle,
990
- or when you want to compact on a different schedule.
991
-
992
- ### Strategies
993
-
994
- **[`LLM::Compactor::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Null.html)**
995
- — the default. Does nothing.
996
-
997
- **[`LLM::Compactor::Truncate`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Truncate.html)**
998
- — drops the oldest messages, keeping only the N most recent.
999
-
1000
- - **Fast** — no network call, no LLM overhead. Operates entirely in memory
1001
- with a single pass over the message list.
1002
- - **No dependencies** — works offline, has no model or API requirements, and
1003
- introduces no additional cost.
1004
- - **Tool-loop safe** — when a tool result (return) falls at the truncation
1005
- boundary, the corresponding tool call is kept. Without this, the
1006
- conversation would contain an orphaned result with no matching call,
1007
- causing API-level errors on the next turn.
1008
-
1009
- The `keep:` parameter accepts either an integer count or a percentage
1010
- string like `"80%"`, which keeps approximately 80% of the most recent
1011
- messages. This is useful when you want to trim proportionally rather
1012
- than to an absolute number.
1013
-
1014
- ```ruby
1015
- ctx = LLM::Context.new(
1016
- llm,
1017
- compactor: LLM::Compactor::Truncate,
1018
- compactor_options: {keep: 128}
1019
- )
1020
- ```
1021
-
1022
- ### Manual compaction
1023
-
1024
- The REPL provides a `/compact` command that invokes Truncate on the current
1025
- agent's context:
1026
-
1027
- ```
1028
- /compact # keep last 128 messages
1029
- /compact 50 # keep last 50 messages
1030
- /compact 75% # keep approximately 75% of messages
1031
- ```
1032
-
1033
- ### Lifecycle callbacks
1034
-
1035
- Both strategies call stream hooks so the UI can show progress:
1036
-
1037
- ```ruby
1038
- def on_compaction(compactor)
1039
- # called before compaction begins
1040
- end
1041
-
1042
- def on_compaction_finish(compactor)
1043
- # called after compaction completes
1044
- end
1045
- ```
1046
-
1047
- The context's `compacted?` flag is `true` between compaction and the next
1048
- model response.
1049
-
1050
- [Back to top](#table-of-contents)
1051
-
1052
- ## Serialization
1053
-
1054
- The [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
1055
- class can be serialized to JSON and stored in a string or on disk.
1056
- That is powerful because a context contains runtime state that can
1057
- be restored later, in a different process or even on a different
1058
- machine. And because an agent is implemented on top of
1059
- [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
1060
- this feature works for [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html),
1061
- too.
1062
-
1063
- #### Save to disk
1064
-
1065
- The runtime can serialize its state to a string, a text file, or
1066
- a database column. The option that fits best depends on your application
1067
- and environment. Web applications might be more interested in the [ORM](#orm)
1068
- feature, which is built on top of the serialization feature.
1069
-
1070
- ```ruby
1071
- ##
1072
- # Create a provider
1073
- llm = LLM.deepseek(key: ENV["KEY"])
1074
-
1075
- ##
1076
- # Save agent
1077
- agent1 = LLM::Agent.new(llm)
1078
- agent1.talk "remember my name is robert"
1079
- agent1.save(path: "agent.json")
1080
-
1081
- ##
1082
- # Restore agent
1083
- agent2 = LLM::Agent.new(llm, stream: $stdout)
1084
- agent2.restore(path: "agent.json")
1085
- agent2.talk "what's my name?"
1086
- ```
1087
-
1088
- ## ORM
1089
-
1090
- Both ActiveRecord, and Sequel have first-class support on the
1091
- llm.rb runtime. In both cases an ActiveRecord or Sequel model
1092
- can be turned into a model that has the same capabilities as
1093
- [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html),
1094
- or [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
1095
-
1096
- The main difference is that the runtime persists directly into
1097
- the database with no requirements beyond a single column on a
1098
- single row. That means it is usually trivial to turn an existing
1099
- model into an AI-aware model.
1100
-
1101
- #### ActiveRecord
1102
-
1103
- The ActiveRecord interface for
1104
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
1105
- is
1106
- [`acts_as_agent`](https://r.uby.dev/api-docs/llm.rb/LLM/ActiveRecord/ActsAsAgent.html).
1107
- It yields an instance of
1108
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html),
1109
- and that can be used
1110
- to configure the agent (eg which model, instructions, skills,
1111
- tools, etc).
1112
-
1113
- An interesting option is the `format` option, by default it
1114
- defaults to `:string` but it can also be changed to `:json`
1115
- or `:jsonb` depending on the configuration and type of underlying
1116
- column. The JSONB column type is recommended.
1117
-
1118
- ```ruby
1119
- require "active_record"
1120
- require "llm"
1121
- require "llm/active_record"
1122
-
1123
- class Agent < ApplicationRecord
1124
- acts_as_agent(format: :jsonb) do |agent|
1125
- agent.model "deepseek-v4-pro"
1126
- agent.instructions "solve the user's query"
1127
- agent.tools [Research, FinalizeResearch, ActOnResearch]
1128
- end
1129
-
1130
- private
1131
-
1132
- ##
1133
- # By convention, this method defines the provider
1134
- # for a model. If neccessary, it can be renamed and
1135
- # configured via `provider: :your_method` instead.
1136
- def set_provider
1137
- LLM.deepseek(key: ENV["KEY"])
1138
- end
1139
-
1140
- ##
1141
- # By convention, this method should return what is
1142
- # given as the second argument to `LLM::Context` or
1143
- # `LLM::Agent`.
1144
- #
1145
- # Often, there is no need to set it, so it can be left
1146
- # undefined or it can be reassigned in the same way as
1147
- # `set_provider`. For example: `context: :your_method`
1148
- def set_context
1149
- {}
1150
- end
1151
- end
1152
-
1153
- agent = Agent.create!
1154
- agent.talk "perform research"
1155
- ```
1156
-
1157
- #### Sequel
1158
-
1159
- The following is a Sequel equivalent to the ActiveRecord example,
1160
- but to keep it interesting and informative, this example also
1161
- configures a per-model tracer that logs to `$stdout`. Works the
1162
- same for ActiveRecord.
1163
-
1164
- ```ruby
1165
- require "sequel"
1166
- require "llm"
1167
- require "llm/sequel/plugin"
1168
-
1169
- class Agent < Sequel::Model
1170
- plugin(:agent, format: :jsonb) do |agent|
1171
- agent.model "deepseek-v4-pro"
1172
- agent.instructions "solve the user's query"
1173
- agent.tools [Research, FinalizeResearch, ActOnResearch]
1174
- agent.tracer { LLM::Tracer::Logger.new(llm, io: $stdout) }
1175
- end
1176
-
1177
- private
1178
-
1179
- def set_provider
1180
- LLM.deepseek(key: ENV["KEY"])
1181
- end
1182
- end
1183
-
1184
- agent = Agent.create
1185
- agent.talk "perform research"
1186
- ```
1187
-
1188
- [Back to top](#table-of-contents)
1189
-
1190
- ## Schema
1191
-
1192
- The [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html)
1193
- class can be subclassed to describe
1194
- the shape of a JSON object or objects that you expect
1195
- the model to respond with.
1196
-
1197
- It can be useful for a wide range of use cases but the
1198
- most popular might be classification, data extraction,
1199
- and transferring structured data between different software
1200
- rather than blobs of text that a machine cannot easily parse
1201
- in a structured way.
1202
-
1203
- #### Estimation
1204
-
1205
- The following example asks the model to estimate the age
1206
- of a person in a photo. The model provides a structured response
1207
- that's represented by an instance of
1208
- [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html).
1209
-
1210
- The object returned by
1211
- [`LLM::Response#content!`](https://r.uby.dev/api-docs/llm.rb/LLM/Contract/Completion.html#content!-instance_method)
1212
- has methods that can access the age, confidence, and comments
1213
- properties.
1214
- This approach can also work for extracting data or an analysis
1215
- from a PDF, and other file types.
1216
-
1217
- ```ruby
1218
- require "llm"
1219
- require "pp"
1220
-
1221
- class Estimation < LLM::Schema
1222
- property :age, Integer, "The estimated age of the person"
1223
- property :confidence, Number, "Your confidence in the estimate"
1224
- property :applicable, Boolean, "True when the photo contains a person"
1225
- property :comments, String, "Any additional comments or input"
1226
- required %i[age confidence applicable comments]
1227
- end
1228
-
1229
- llm = LLM.openai(key: ENV["KEY"])
1230
- agent = LLM::Agent.new(llm, schema: Estimation)
1231
- res = agent.ask "Given this photo, provide an age estimate", with: "photo.jpg"
1232
-
1233
- ##
1234
- # Coerces the model's response from a JSON string
1235
- # to an instance of LLM::Object.
1236
- estimate = res.content!
1237
-
1238
- ##
1239
- # Let's print the estimate
1240
- if estimate.applicable
1241
- print "The person is approx ", estimate.age.to_s, " years old", "\n"
1242
- print "I have a confidence rating of ", estimate.confidence.to_s, "\n"
1243
- else
1244
- print "This photo is not applicable:", "\n"
1245
- print estimate.comments
1246
- end
1247
- ```
1248
-
1249
- [Back to top](#table-of-contents)
1250
-
1251
- ## Cancellation
1252
-
1253
- #### Cancel a request
1254
-
1255
- A common scenario when communicating with a model is to
1256
- want to cancel the request mid-stream. This could be done
1257
- for a number of different reasons, most often because the
1258
- user made a mistake, or the model is making a mistake and
1259
- the user wants to cancel the action.
1260
-
1261
- The runtime has built-in support for cancellation. Call
1262
- `agent.cancel!` or `ctx.cancel!` from any thread and two
1263
- things happen at once:
1264
-
1265
- [`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html)
1266
- is raised on the thread where `talk` is running, so the
1267
- caller can rescue it and know the request was cancelled.
1268
-
1269
- At the same time, `LLM::Interrupt` is raised on every tool
1270
- that is currently executing &mdash; regardless of which
1271
- concurrency strategy it's using. A tool running in a thread
1272
- gets it on that thread. A tool in a fiber gets it on that
1273
- fiber. A tool in a forked process gets it via a message
1274
- over the xchan control channel. The
1275
- delivery mechanism depends on the strategy, but the effect is
1276
- the same: the tool can rescue `LLM::Interrupt`, clean up
1277
- resources, close connections, flush buffers, and either
1278
- re-raise to abort or return a partial result.
1279
-
1280
- Pending tools &mdash; those the model requested but that haven't
1281
- started running yet &mdash; are cancelled through
1282
- [`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel-instance_method)
1283
- without ever being executed.
1284
-
1285
- The transport layer also cancels the in-flight HTTP request,
1286
- closing the connection to the provider.
1287
-
1288
- ```ruby
1289
- require "llm"
1290
-
1291
- llm = LLM.deepseek(key: ENV["DEEPSEEK_SECRET"])
1292
- agent = LLM::Agent.new(llm)
1293
- queue = Queue.new
1294
-
1295
- Thread.new do
1296
- queue.push(nil)
1297
- sleep(2)
1298
- agent.cancel!
1299
- end
1300
-
1301
- begin
1302
- queue.pop
1303
- agent.talk "write me a very long poem", stream: $stdout
1304
- rescue LLM::Interrupt
1305
- puts "request cancelled!"
1306
- end
1307
- ```
1308
-
1309
- #### Tool interrupts
1310
-
1311
- When a running tool is interrupted &ndash; for example the user presses
1312
- ESC in the [REPL](#repl) &ndash; the runtime raises
1313
- [`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html)
1314
- on the tool's execution context. This behavior is uniform across
1315
- all six concurrency strategies (`:sequential`, `:thread`, `:fiber`,
1316
- `:async`, `:fork`, and `:ractor`).
1317
-
1318
- A tool has two choices:
1319
-
1320
- **Re-raise** `LLM::Interrupt` to cancel the entire turn. The
1321
- exception propagates out of the tool loop and the request is
1322
- aborted. This is the default when you don't rescue the exception.
1323
-
1324
- ```ruby
1325
- def call
1326
- # ... do work ...
1327
- rescue LLM::Interrupt
1328
- cleanup
1329
- raise # cancel the turn
1330
- end
1331
- ```
1332
-
1333
- **Return a value** to continue the tool loop. The model receives
1334
- the result and decides what to do next, aware that the tool was
1335
- interrupted.
1336
-
1337
- ```ruby
1338
- def call
1339
- # ... do work ...
1340
- rescue LLM::Interrupt
1341
- cleanup
1342
- {ok: false, reason: "interrupted"} # continue the loop
1343
- end
1344
- ```
1345
-
1346
- The right choice depends on the situation. A hard cancel aborts
1347
- the request outright &ndash; useful when continuing would produce
1348
- garbage. Returning a value lets the model adapt, which can be
1349
- helpful when the interrupt is temporary (e.g. a timeout).
1350
-
1351
- The `:ractor` strategy delivers the interrupt through ractor
1352
- message passing &mdash; a listener thread inside the tool ractor
1353
- receives the interrupt message and raises `LLM::Interrupt` on the
1354
- ractor's main thread. The end result is the same as every other
1355
- strategy: the tool can rescue, clean up, and decide.
1356
- The `:fork` strategy delivers the interrupt via a message
1357
- over the xchan control channel, which a listener thread in the
1358
- child process picks up and raises on `Thread.main`. All other strategies
1359
- raise the exception directly on the executing thread or fiber.
1360
-
1361
- [Back to top](#table-of-contents)
1362
-
1363
- ## Tracer
1364
-
1365
- The runtime can be observed by subclasses of
1366
- [`LLM::Tracer`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html). <br>
1367
- The default tracers include a tracer that can write to standard
1368
- output
1369
- ([`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)),
1370
- and a generic OpenTelemetry tracer that can export spans via OTLP
1371
- ([`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)).
1372
-
1373
- llm.rb has numerous hooks implemented throughout the runtime that
1374
- [`LLM::Tracer`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html)
1375
- subclasses can hook into, and the tracer is
1376
- purposefully designed to be extensible. The scope of a trace
1377
- can vary from an individual agent (an instance of
1378
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)),
1379
- or for every request a provider makes (an indirect instance of
1380
- [`LLM::Provider`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html)).
1381
-
1382
- #### Provider-wide tracer
1383
-
1384
- The following two examples demonstrate provider-wide tracers that
1385
- cover every request made for a single provider.
1386
-
1387
- ```ruby
1388
- ##
1389
- # Provider-wide tracer
1390
- # Writes to $stdout
1391
- llm = LLM.deepseek(key: ENV["KEY"])
1392
- llm.tracer = LLM::Tracer::Logger.new(llm, io: $stdout)
1393
-
1394
- ##
1395
- # Provider-wide tracer
1396
- # Writes to deepseek.log
1397
- llm = LLM.deepseek(key: ENV["KEY"])
1398
- llm.tracer = LLM::Tracer::Logger.new(llm, path: "deepseek.log")
1399
- ```
1400
-
1401
- #### Agent-local tracer
1402
-
1403
- The next two examples demonstrate a tracer that is local
1404
- to an agent.
1405
-
1406
- ```ruby
1407
- ##
1408
- # Agent-local
1409
- # Writes to $stdout
1410
- llm = LLM.deepseek(key: ENV["KEY"])
1411
- agent = LLM::Agent.new(llm, tracer: LLM::Tracer::Logger.new(llm, io: $stdout))
1412
-
1413
- ##
1414
- # Agent-local
1415
- # Writes to deepseek-agent.log
1416
- llm = LLM.deepseek(key: ENV["KEY"])
1417
- agent = LLM::Agent.new(llm, tracer: LLM::Tracer::Logger.new(llm, path: "deepseek-agent.log"))
1418
- ```
1419
-
1420
- [Back to top](#table-of-contents)
1421
-
1422
- ## REPL
1423
-
1424
- During the development and operation of agents it can often
1425
- be helpful to drop into a read-eval-print loop. This gives
1426
- you a way to confirm the work was successful, inspect
1427
- anything that went wrong, and keep talking to the same
1428
- agent while its state is still intact.
1429
-
1430
- The REPL is a curses-based TUI with a status line showing
1431
- a context-usage bar and cost counter, a scrollable transcript
1432
- that renders markdown, and a multi-line input area. The UI
1433
- thread stays responsive while a second thread communicates
1434
- with the model.
1435
-
1436
- #### LLM::Agent
1437
-
1438
- The [LLM::Agent#repl](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#repl-instance_method)
1439
- method allows an agent to spawn a read-eval-print loop
1440
- that can be useful while developing or operating agents.
1441
- It can be used to debug tool calls, confirm an
1442
- agent has done what was expected, or improve an agent by
1443
- asking questions about what it has done up to that point.
1444
-
1445
- This feature requires that the [curses](https://github.com/ruby/curses)
1446
- and [kramdown](https://github.com/gettalong/kramdown) libraries are
1447
- installed and available to require.
1448
-
1449
- The `name:` option labels the agent throughout the TUI &mdash;
1450
- useful when working with multiple agents. The `path:` option
1451
- persists state across sessions. The `tools:` option attaches
1452
- extra tools for the duration of the session.
1453
-
1454
- ```ruby
1455
- require "llm"
1456
-
1457
- llm = LLM.deepseek(key: ENV["KEY"])
1458
- agent = LLM::Agent.new(llm, name: "my-agent")
1459
- agent.repl(path: "session.json", tools: LLM::Tool.subclasses)
1460
- ```
1461
-
1462
- #### Persistence
1463
-
1464
- The `path:` option accepts a file path where runtime state
1465
- is read from and written to. This lets you resume a
1466
- conversation across REPL sessions. When the file does not
1467
- exist the agent starts fresh; when it does, the agent
1468
- restores its previous state.
1469
-
1470
- ```ruby
1471
- llm = LLM.deepseek(key: ENV["KEY"])
1472
- agent = LLM::Agent.new(llm)
1473
- agent.repl(path: "session.json")
1474
- ```
1475
-
1476
- #### Tools
1477
-
1478
- The read-eval-print loop accepts a `tools` option that lets
1479
- you attach additional tools for the duration of the session.
1480
-
1481
- ```ruby
1482
- llm = LLM.deepseek(key: ENV["KEY"])
1483
- agent = LLM::Agent.new(llm)
1484
- agent.repl(tools: [Debugger])
1485
- ```
1486
-
1487
- Load every built-in tool with `LLM::Tool.subclasses`:
1488
-
1489
- ```ruby
1490
- require "llm/tools"
1491
-
1492
- llm = LLM.deepseek(key: ENV["KEY"])
1493
- agent = LLM::Agent.new(llm)
1494
- agent.repl(tools: LLM::Tool.subclasses)
1495
- ```
1496
-
1497
- #### Skills
1498
-
1499
- The read-eval-print loop also accepts a `skills` option.
1500
- This can be useful when you want to load extra skills
1501
- without attaching them to an agent permanently.
1502
-
1503
- ```ruby
1504
- llm = LLM.deepseek(key: ENV["KEY"])
1505
- agent = LLM::Agent.new(llm)
1506
- agent.repl(skills: [__dir__])
1507
- ```
1508
-
1509
- #### Tracer
1510
-
1511
- By default the tracer is disabled for the duration of the
1512
- session. This can be configured through the
1513
- `tracer` option. Setting it to `true` will configure
1514
- the REPL to use the tracer associated with an instance
1515
- of [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
1516
-
1517
- ```ruby
1518
- llm = LLM.deepseek(key: ENV["KEY"])
1519
- agent = LLM::Agent.new(llm, tracer: LLM.logger(llm, path: "agent.log"))
1520
- agent.repl(tracer: true, tools: [Debugger])
1521
- ```
1522
-
1523
- #### Input
1524
-
1525
- The input area supports several keyboard shortcuts.
1526
- When characters arrive faster than a threshold the REPL
1527
- detects that text is being pasted rather than typed. In
1528
- paste mode pressing `Enter` inserts a newline instead of
1529
- submitting, allowing multi-line prompts.
1530
-
1531
- | Key | Action |
1532
- |---|---|
1533
- | `Enter` | Submit the current prompt |
1534
- | `Ctrl+A` | Jump to the start of the line |
1535
- | `Ctrl+E` | Jump to the end of the line |
1536
- | `Ctrl+F` | Move the cursor forward |
1537
- | `Ctrl+K` | Erase from cursor to the end of the line |
1538
- | `Ctrl+Y` | Paste previously killed text |
1539
- | `Ctrl+D` | Delete the character at the cursor |
1540
- | `Ctrl+P` | Recall the previous user message |
1541
- | `Ctrl+N` | Recall the next user message |
1542
- | `Left / Right` | Move the cursor |
1543
- | `Up / Down` | Scroll the transcript one line |
1544
- | `PgUp` / `PgDn` | Scroll the transcript by one page |
1545
- | `Tab` | Complete `/command` names |
1546
- | `Esc` | Cancel the current request |
1547
-
1548
- #### Commands
1549
-
1550
- Commands are recognized by a `/` prefix on the input line.
1551
- Type `/compact` to free context window space by dropping the
1552
- oldest messages. Type `/exit` to leave the REPL.
1553
-
1554
- The [`LLM::Repl::Command`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl/Command.html)
1555
- class is intentionally similar to [`LLM::Tool`](#llmtool) and
1556
- [`LLM::Schema`](#schema) in its interface &ndash; you declare a
1557
- name, description, and parameters with the same vocabulary.
1558
- A subclass is automatically registered and available as `/name`.
1559
-
1560
- ##### Parameters
1561
-
1562
- Parameters are declared with `parameter :name, Type, "description"`
1563
- and marked required with `required %i[name]`. The `call` method
1564
- receives them as keyword arguments matching the parameter names.
1565
- Parameters without a user-supplied value fall back to the
1566
- method signature's default.
1567
-
1568
- ```ruby
1569
- class Greeter < LLM::Command
1570
- name "greet"
1571
- description "Greets the given name"
1572
- parameter :name, String, "The person's name"
1573
- required %i[name]
1574
-
1575
- def call(name:)
1576
- write "Welcome #{name}!\n"
1577
- end
1578
- end
1579
- ```
1580
-
1581
- ##### Output
1582
-
1583
- A command writes to the transcript with `write(str, who:)`.
1584
- The `who:` label is rendered in bold. It defaults to
1585
- `command(name): ` where `name` is the command's registered
1586
- name.
1587
-
1588
- ```ruby
1589
- def call(name:)
1590
- write("Greetings #{name}!\n")
1591
- end
1592
- ```
1593
-
1594
- ##### Help
1595
-
1596
- The built-in [`help`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl/Command.html#help-class_method)
1597
- class method formats the name, description, and parameter list
1598
- automatically. Use it from inside a command or via `/help`.
1599
-
1600
- ```ruby
1601
- class Greeter < LLM::Command
1602
- name "greet"
1603
- # ...
1604
- end
1605
-
1606
- # /help greet displays:
1607
- # Command: greet
1608
- # Description: Greets the given name
1609
- #
1610
- # Parameters:
1611
- # name [String] - The person's name (required)
1612
- ```
1613
-
1614
- ##### Aliases
1615
-
1616
- Subclassing an existing command inherits its name, description,
1617
- and parameters. This is how `/quit` is an alias of `/exit`:
1618
-
1619
- ```ruby
1620
- class Quit < LLM::Repl::Command::Exit
1621
- name "quit"
1622
- end
1623
- ```
1624
-
1625
- [Back to top](#table-of-contents)
1626
-
1627
- ## Images
1628
-
1629
- The OpenAI, Google, xAI, DeepInfra, and DeepSeek providers have
1630
- builtin image generation capabilities. OpenAI, xAI, and DeepInfra
1631
- also support image edits. Google only supports image generation.
1632
- DeepSeek supports generation and edits too, but only through SVG
1633
- output rather than raster image models.
1634
-
1635
- #### Generation
1636
-
1637
- The [`LLM::Provider#images`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#images-instance_method)
1638
- method returns an Image
1639
- object that a subset of providers implement. At the
1640
- moment Google, xAI, OpenAI, DeepInfra, and DeepSeek have image
1641
- generation capabilities. DeepSeek is the odd one out: it generates
1642
- SVG documents rather than raster images.
1643
-
1644
- ```ruby
1645
- require "llm"
1646
-
1647
- ##
1648
- # Store dogrocket.png
1649
- llm = LLM.openai(key: ENV["KEY"])
1650
- res = llm.images.create(prompt: "a dog on a rocket to the moon")
1651
- IO.copy_stream res.images[0], "dogrocket.png"
1652
- ```
1653
-
1654
- The API is the same across providers. <br>
1655
- For example &ndash; xAI:
1656
-
1657
- ```ruby
1658
- require "llm"
1659
-
1660
- ##
1661
- # Store dogrocket.png
1662
- # Same API as OpenAI
1663
- llm = LLM.xai(key: ENV["KEY"])
1664
- res = llm.images.create(prompt: "a dog on a rocket to the moon")
1665
- IO.copy_stream res.images[0], "dogrocket.png"
1666
- ```
1667
-
1668
- #### Edits
1669
-
1670
- OpenAI, xAI, and DeepInfra have the same interface for image edits. <br>
1671
- DeepSeek also supports edits, but only for SVG files. <br>
1672
- Google does not have edit image support. <br>
1673
-
1674
- ```ruby
1675
- require "llm"
1676
-
1677
- ##
1678
- # Edit self.jpg and add a mustache
1679
- # Save to mustache.png
1680
- llm = LLM.openai(key: ENV["KEY"])
1681
- res = llm.images.edit(prompt: "add a mustache", image: "self.jpg")
1682
- IO.copy_stream res.images[0], "mustache.png"
1683
- ```
1684
-
1685
- #### DeepSeek
1686
-
1687
- The DeepSeek provider does not provide an image generation model
1688
- but it is possible to ask a text-to-text model to produce
1689
- vector graphics (SVGs), and in that limited sense, it can become
1690
- a capable text-to-image model.
1691
-
1692
- ```ruby
1693
- require "llm"
1694
-
1695
- ##
1696
- # Edit rocket.svg and change its color
1697
- # Save to rocket-edited.svg
1698
- llm = LLM.deepseek(key: ENV["KEY"])
1699
- res = llm.images.edit(prompt: "make the rocket red", image: "rocket.svg")
1700
- IO.copy_stream res.images[0], "rocket-edited.svg"
1701
- ```
1702
-
1703
- An interesting property of the DeepSeek implementation is that
1704
- it can maintain a session that can perform multiple image generations
1705
- or edits rather than just one-shot generations.
1706
-
1707
- It's possible because under the hood
1708
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html),
1709
- is attached to the
1710
- [`LLM::Response`](https://r.uby.dev/api-docs/llm.rb/LLM/Response.html)
1711
- object that is returned to the caller. So the response includes an
1712
- `agent` method, and it can be carried across multiple generations.
1713
- It is specific to this endpoint though. It works like this:
1714
-
1715
- ```ruby
1716
- require "llm"
1717
-
1718
- llm = LLM.deepseek(key: ENV["DEEPSEEK_SECRET"])
1719
- agent = nil
1720
- loop do
1721
- print "> "
1722
- prompt = $stdin.gets
1723
- res = llm.images.create(prompt:, agent:)
1724
- agent = res.agent
1725
- IO.copy_stream res.images[0], "image.svg"
1726
- print "ok: saved image.svg", "\n"
1727
- end
1728
- ```
1729
-
1730
- [Back to top](#table-of-contents)
1731
-
1732
- ## Audio
1733
-
1734
- The audio interface defined by llm.rb describes three methods,
1735
- although not every provider implements all of them. Generally
1736
- speaking the audio interface is for text-to-speech, and
1737
- speech-to-text models.
1738
-
1739
- The following providers have audio support:
1740
-
1741
- * OpenAI - full support
1742
- * Google - partial support
1743
- * DeepInfra - partial support
1744
-
1745
- #### text-to-speech
1746
-
1747
- The `create_speech` method generates an audio clip based
1748
- on the given input. This method returns a
1749
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
1750
- object. OpenAI, and DeepInfra support this method.
1751
-
1752
- ```ruby
1753
- require "llm"
1754
-
1755
- llm = LLM.openai(key: ENV["KEY"])
1756
- res = llm.audio.create_speech(input: "Hello world")
1757
- IO.copy_stream res.audio.decoded, "helloworld.mp3"
1758
- ```
1759
-
1760
- #### speech-to-text
1761
-
1762
- The `create_transcription` method transcribes a given
1763
- audio clip as text. OpenAI, Google and DeepInfra support
1764
- this method.
1765
-
1766
- ```ruby
1767
- require "llm"
1768
-
1769
- llm = LLM.google(key: ENV["KEY"])
1770
- res = llm.audio.create_transcription(file: "helloworld.mp3")
1771
- res.text # => "Hello world"
1772
- ```
1773
-
1774
- #### translation
1775
-
1776
- The `create_translation` method translates a given audio
1777
- clip, then transcribes it as text. OpenAI, and Google
1778
- support this method.
1779
-
1780
- ```ruby
1781
- require "llm"
1782
-
1783
- llm = LLM.google(key: ENV["KEY"])
1784
- res = llm.audio.create_translation(file: "bomdia.mp3")
1785
- res.text # => "Good day"
1786
- ```
1787
-
1788
- [Back to top](#table-of-contents)
1789
-
1790
- ## OCR
1791
-
1792
- Optical Character Recognition extracts text from images and
1793
- documents.
52
+ ## Fundamentals
1794
53
 
1795
- #### Mistral
54
+ - [Agents](deepdive/fundamentals/agents.md)
55
+ - [Tools](deepdive/fundamentals/tools.md)
56
+ - [Skills](deepdive/fundamentals/skills.md)
57
+ - [Schema](deepdive/fundamentals/schema.md)
58
+ - [Stream](deepdive/fundamentals/stream.md)
59
+ - [Database](deepdive/fundamentals/database.md)
60
+ - [Concurrency](deepdive/fundamentals/concurrency.md)
61
+ - [REPL](deepdive/fundamentals/repl.md)
1796
62
 
1797
- Mistral is the only provider that currently supports OCR
1798
- through its dedicated API endpoint. The `ocr` method accepts
1799
- either an `image_url:` or a `document_url:` parameter.
1800
- Document URLs can point to PDFs. The response exposes pages
1801
- through `res.pages`, where each page has a `markdown` field
1802
- containing the extracted text.
63
+ ## Advanced
1803
64
 
1804
- ```ruby
1805
- require "llm"
65
+ - [Context](deepdive/advanced/context.md)
66
+ - [Compaction](deepdive/advanced/compaction.md)
67
+ - [Cancellation](deepdive/advanced/cancellation.md)
68
+ - [Transports](deepdive/advanced/transports.md)
69
+ - [Tracer](deepdive/advanced/tracer.md)
1806
70
 
1807
- llm = LLM.mistral(key: ENV["KEY"])
71
+ ## Protocols
1808
72
 
1809
- ##
1810
- # Extract text from an image
1811
- res = llm.ocr(image_url: "https://example.com/photo.png")
1812
- res.pages.each { |page| puts page.markdown }
73
+ - [MCP](deepdive/protocols/mcp.md)
74
+ - [A2A](deepdive/protocols/a2a.md)
1813
75
 
1814
- ##
1815
- # Extract text from a PDF
1816
- res = llm.ocr(document_url: "https://example.com/report.pdf")
1817
- res.pages.each { |page| puts page.markdown }
1818
- ```
76
+ ## Everything else
1819
77
 
1820
- [Back to top](#table-of-contents)
78
+ - [Images](deepdive/everything_else/images.md)
79
+ - [Audio](deepdive/everything_else/audio.md)
80
+ - [OCR](deepdive/everything_else/ocr.md)