llm.rb 13.1.0 → 14.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +320 -0
- data/README.md +340 -31
- data/bin/llm.rb +36 -12
- data/data/anthropic.json +206 -263
- data/data/bedrock.json +2138 -1860
- data/data/deepinfra.json +1003 -624
- data/data/deepseek.json +38 -34
- data/data/google.json +1079 -371
- data/data/mistral.json +448 -368
- data/data/moonshot.json +384 -0
- data/data/openai.json +974 -1343
- data/data/xai.json +154 -126
- data/data/zai.json +191 -191
- data/lib/llm/agent.rb +47 -14
- data/lib/llm/context.rb +71 -88
- data/lib/llm/cost.rb +23 -17
- data/lib/llm/error.rb +0 -8
- data/lib/llm/function/async/task.rb +2 -0
- data/lib/llm/function/fiber/task.rb +2 -0
- data/lib/llm/function/fork/task.rb +2 -0
- data/lib/llm/function/ractor/task.rb +2 -0
- data/lib/llm/function/sequential/group.rb +4 -1
- data/lib/llm/function/sequential/task.rb +1 -1
- data/lib/llm/function/task.rb +4 -0
- data/lib/llm/function/thread/task.rb +2 -0
- data/lib/llm/function.rb +32 -4
- data/lib/llm/guard/loop.rb +89 -0
- data/lib/llm/guard/null.rb +19 -0
- data/lib/llm/guard.rb +61 -0
- data/lib/llm/provider.rb +36 -0
- data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
- data/lib/llm/providers/anthropic.rb +1 -8
- data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
- data/lib/llm/providers/bedrock.rb +1 -8
- data/lib/llm/providers/google/stream_parser.rb +1 -0
- data/lib/llm/providers/google.rb +1 -8
- data/lib/llm/providers/moonshot.rb +76 -0
- data/lib/llm/providers/ollama.rb +1 -8
- data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
- data/lib/llm/providers/openai/responses.rb +6 -8
- data/lib/llm/providers/openai/stream_parser.rb +1 -0
- data/lib/llm/providers/openai.rb +3 -10
- data/lib/llm/repl/bar.rb +4 -3
- data/lib/llm/repl/buffer.rb +42 -15
- data/lib/llm/repl/color.rb +78 -0
- data/lib/llm/repl/input/char.rb +46 -0
- data/lib/llm/repl/input/row.rb +39 -0
- data/lib/llm/repl/input.rb +251 -66
- data/lib/llm/repl/markdown/table.rb +6 -2
- data/lib/llm/repl/markdown.rb +31 -5
- data/lib/llm/repl/status.rb +38 -3
- data/lib/llm/repl/stream.rb +16 -4
- data/lib/llm/repl/walker.rb +3 -2
- data/lib/llm/repl/window.rb +25 -5
- data/lib/llm/repl.rb +29 -13
- data/lib/llm/stream.rb +8 -7
- data/lib/llm/tool.rb +29 -0
- data/lib/llm/transformer/null.rb +21 -0
- data/lib/llm/transformer.rb +55 -0
- data/lib/llm/version.rb +1 -1
- data/lib/llm.rb +12 -2
- data/llm.gemspec +1 -0
- data/resources/deepdive/advanced/cancellation.md +74 -0
- data/resources/deepdive/advanced/compaction.md +83 -0
- data/resources/deepdive/advanced/context.md +267 -0
- data/resources/deepdive/advanced/guard.md +371 -0
- data/resources/deepdive/advanced/tracer.md +180 -0
- data/resources/deepdive/advanced/transformer.md +67 -0
- data/resources/deepdive/advanced/transports.md +45 -0
- data/resources/deepdive/everything_else/audio.md +122 -0
- data/resources/deepdive/everything_else/cost.md +99 -0
- data/resources/deepdive/everything_else/images.md +89 -0
- data/resources/deepdive/everything_else/object.md +108 -0
- data/resources/deepdive/everything_else/ocr.md +48 -0
- data/resources/deepdive/fundamentals/agents.md +202 -0
- data/resources/deepdive/fundamentals/builtin_tools.md +191 -0
- data/resources/deepdive/fundamentals/concurrency.md +104 -0
- data/resources/deepdive/fundamentals/database.md +449 -0
- data/resources/deepdive/fundamentals/embeddings.md +157 -0
- data/resources/deepdive/fundamentals/repl.md +87 -0
- data/resources/deepdive/fundamentals/schema.md +61 -0
- data/resources/deepdive/fundamentals/skills.md +106 -0
- data/resources/deepdive/fundamentals/stream.md +110 -0
- data/resources/deepdive/fundamentals/tools.md +265 -0
- data/resources/deepdive/protocols/a2a.md +106 -0
- data/resources/deepdive/protocols/mcp.md +111 -0
- data/resources/deepdive.md +7 -1
- metadata +36 -3
- data/lib/llm/loop_guard.rb +0 -107
|
@@ -0,0 +1,122 @@
|
|
|
1
|
+
|
|
2
|
+
## Audio
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
The audio interface covers three things: turning text into speech,
|
|
9
|
+
transcribing audio into text, and translating spoken language.
|
|
10
|
+
OpenAI supports all three. Google and DeepInfra support subsets.
|
|
11
|
+
Each method follows the same pattern: pass input, get output,
|
|
12
|
+
copy the result somewhere useful.
|
|
13
|
+
|
|
14
|
+
#### How it works
|
|
15
|
+
|
|
16
|
+
When you want to convert text to speech, call
|
|
17
|
+
[`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech).
|
|
18
|
+
The provider returns an audio clip as a
|
|
19
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
20
|
+
object. The generated audio can be copied
|
|
21
|
+
to a file or streamed directly. Each provider supports different
|
|
22
|
+
output formats and voice options:
|
|
23
|
+
|
|
24
|
+
```ruby
|
|
25
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
26
|
+
res = llm.audio.create_speech(input: "Hello world")
|
|
27
|
+
IO.copy_stream res.audio.decoded, "helloworld.mp3"
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
#### Why would I use it?
|
|
31
|
+
|
|
32
|
+
Audio support lets you build voice interfaces, add accessibility
|
|
33
|
+
features, and work across languages without wiring up a separate
|
|
34
|
+
speech service. Generate audio for notifications, narrate written
|
|
35
|
+
content, or add speech output to an existing application with a
|
|
36
|
+
single method call.
|
|
37
|
+
|
|
38
|
+
#### Notes
|
|
39
|
+
|
|
40
|
+
OpenAI has full audio support. Google and DeepInfra have partial
|
|
41
|
+
support. The
|
|
42
|
+
[`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech)
|
|
43
|
+
method returns a
|
|
44
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
45
|
+
object.
|
|
46
|
+
|
|
47
|
+
### Transcription
|
|
48
|
+
|
|
49
|
+
#### Overview
|
|
50
|
+
|
|
51
|
+
Transcription turns an audio file into text. Pass a file path
|
|
52
|
+
to
|
|
53
|
+
[`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription)
|
|
54
|
+
and get back the spoken content as
|
|
55
|
+
a string. OpenAI, Google, and DeepInfra support it. Transcribe
|
|
56
|
+
meeting notes, voice memos, or podcast episodes for search and
|
|
57
|
+
processing. The response is plain text you can feed into any
|
|
58
|
+
downstream pipeline.
|
|
59
|
+
|
|
60
|
+
#### How it works
|
|
61
|
+
|
|
62
|
+
When you want to transcribe audio into text, call
|
|
63
|
+
[`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription).
|
|
64
|
+
The provider processes the audio and returns the transcribed text.
|
|
65
|
+
The file can be a local path or a URL depending on provider support:
|
|
66
|
+
|
|
67
|
+
```ruby
|
|
68
|
+
llm = LLM.google(key: ENV["KEY"])
|
|
69
|
+
res = llm.audio.create_transcription(file: "helloworld.mp3")
|
|
70
|
+
res.text # => "Hello world"
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
#### Why would I use it?
|
|
74
|
+
|
|
75
|
+
Transcribe recorded meetings, voice memos, or podcast episodes for
|
|
76
|
+
search and processing. The returned text feeds directly into search
|
|
77
|
+
indexes, summarization pipelines, or downstream extraction. Spoken
|
|
78
|
+
content becomes as queryable as written text.
|
|
79
|
+
|
|
80
|
+
#### Notes
|
|
81
|
+
|
|
82
|
+
OpenAI has full audio support. Google and DeepInfra have partial
|
|
83
|
+
support.
|
|
84
|
+
|
|
85
|
+
### Translation
|
|
86
|
+
|
|
87
|
+
#### Overview
|
|
88
|
+
|
|
89
|
+
Translation transcribes audio and translates it into English in
|
|
90
|
+
one step. Pass a file to
|
|
91
|
+
[`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation)
|
|
92
|
+
and get back the
|
|
93
|
+
translated text. OpenAI and Google support it. Translate
|
|
94
|
+
multilingual podcasts, interviews, or any audio where you need
|
|
95
|
+
the content in English without running a separate translation
|
|
96
|
+
pipeline.
|
|
97
|
+
|
|
98
|
+
#### How it works
|
|
99
|
+
|
|
100
|
+
When you want to translate spoken audio into English, call
|
|
101
|
+
[`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation).
|
|
102
|
+
The provider transcribes the spoken language and translates the
|
|
103
|
+
result into English in a single operation. The returned text is the
|
|
104
|
+
English translation:
|
|
105
|
+
|
|
106
|
+
```ruby
|
|
107
|
+
llm = LLM.google(key: ENV["KEY"])
|
|
108
|
+
res = llm.audio.create_translation(file: "bomdia.mp3")
|
|
109
|
+
res.text # => "Good day"
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
#### Why would I use it?
|
|
113
|
+
|
|
114
|
+
Translate a podcast or interview into another language without a
|
|
115
|
+
separate speech service. The provider handles both transcription and
|
|
116
|
+
translation in a single call, so you get English text from any
|
|
117
|
+
supported source language without chaining two operations together.
|
|
118
|
+
|
|
119
|
+
#### Notes
|
|
120
|
+
|
|
121
|
+
Each method works independently and each provider supports a
|
|
122
|
+
different subset.
|
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
|
|
2
|
+
## LLM::Cost
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
|
|
9
|
+
represents the approximate cost of a conversation. It breaks the
|
|
10
|
+
total down by token type, so you can see how much was spent on input,
|
|
11
|
+
output, cached tokens, reasoning, audio, and images. Cost is computed
|
|
12
|
+
from token usage and the pricing data shipped in the model registry.
|
|
13
|
+
|
|
14
|
+
#### How it works
|
|
15
|
+
|
|
16
|
+
When you want to know what a conversation cost so far, call
|
|
17
|
+
[`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
|
|
18
|
+
(or
|
|
19
|
+
[`LLM::Agent#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#cost-instance_method))
|
|
20
|
+
and read the breakdown. The REPL shows this live in its status bar
|
|
21
|
+
after every turn:
|
|
22
|
+
|
|
23
|
+
```ruby
|
|
24
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
25
|
+
ctx = LLM::Context.new(llm)
|
|
26
|
+
ctx.talk "Hello"
|
|
27
|
+
|
|
28
|
+
cost = ctx.cost
|
|
29
|
+
cost.input # => 0.0000042
|
|
30
|
+
cost.output # => 0.0000084
|
|
31
|
+
cost.total # => 0.0000126
|
|
32
|
+
cost.to_s # => "0.0000126"
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
#### Why would I use it?
|
|
36
|
+
|
|
37
|
+
Cost tracking matters in production. Monitoring spend per
|
|
38
|
+
conversation, per agent, or per provider tells you which workflows
|
|
39
|
+
are expensive and when to switch models. Log a structured breakdown
|
|
40
|
+
with
|
|
41
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method),
|
|
42
|
+
or read
|
|
43
|
+
[`LLM::Cost#total`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#total-instance_method)
|
|
44
|
+
for a single number.
|
|
45
|
+
|
|
46
|
+
#### Notes
|
|
47
|
+
|
|
48
|
+
Cost is an approximation based on the pricing in the model registry.
|
|
49
|
+
[`LLM::Cost.from`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#from-class_method)
|
|
50
|
+
returns an empty cost when the model or registry cannot be found, so
|
|
51
|
+
a missing model never crashes your code.
|
|
52
|
+
|
|
53
|
+
### Reading the breakdown
|
|
54
|
+
|
|
55
|
+
#### Overview
|
|
56
|
+
|
|
57
|
+
[`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
|
|
58
|
+
exposes each cost component as a reader, plus
|
|
59
|
+
[`LLM::Cost#total`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#total-instance_method),
|
|
60
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method),
|
|
61
|
+
and
|
|
62
|
+
[`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_s-instance_method).
|
|
63
|
+
|
|
64
|
+
#### How it works
|
|
65
|
+
|
|
66
|
+
Each component is a Float, or `nil` when no tokens of that type were
|
|
67
|
+
used. The
|
|
68
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method)
|
|
69
|
+
method returns a Hash with only the non-nil components and the total:
|
|
70
|
+
|
|
71
|
+
```ruby
|
|
72
|
+
cost = ctx.cost
|
|
73
|
+
|
|
74
|
+
cost.input
|
|
75
|
+
cost.output
|
|
76
|
+
cost.cache_read
|
|
77
|
+
cost.cache_write
|
|
78
|
+
cost.reasoning
|
|
79
|
+
cost.input_audio
|
|
80
|
+
cost.output_audio
|
|
81
|
+
cost.input_image
|
|
82
|
+
|
|
83
|
+
cost.to_h # => {input: 4.2e-06, output: 8.4e-06, total: 1.26e-05}
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
#### Why would I use it?
|
|
87
|
+
|
|
88
|
+
The per-component breakdown shows where the money goes. High cache
|
|
89
|
+
read costs suggest a conversation benefits from prompt caching.
|
|
90
|
+
High reasoning costs point at a model that thinks a lot. Log
|
|
91
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method)
|
|
92
|
+
at the end of a session to keep a spend trail.
|
|
93
|
+
|
|
94
|
+
#### Notes
|
|
95
|
+
|
|
96
|
+
[`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_s-instance_method)
|
|
97
|
+
returns the total in a compact, human-friendly format
|
|
98
|
+
(`"0.0000126"`). Components that were not used are `nil`, so sum
|
|
99
|
+
them with `compact` if you aggregate across conversations.
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
|
|
2
|
+
## Images
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
A handful of providers can generate images from a text prompt.
|
|
9
|
+
OpenAI, Google, xAI, DeepInfra, and DeepSeek all support it.
|
|
10
|
+
OpenAI, xAI, and DeepInfra also let you edit existing images.
|
|
11
|
+
The API is the same across providers, so switching between them
|
|
12
|
+
requires no code changes.
|
|
13
|
+
|
|
14
|
+
#### How it works
|
|
15
|
+
|
|
16
|
+
The
|
|
17
|
+
[`LLM::Images#create`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#create)
|
|
18
|
+
method sends a prompt to the provider and
|
|
19
|
+
returns the result as a
|
|
20
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
21
|
+
object. The same API works
|
|
22
|
+
across providers: swap
|
|
23
|
+
[`LLM.openai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#openai-class_method) for
|
|
24
|
+
[`LLM.xai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#xai-class_method) and the rest
|
|
25
|
+
of the code is identical.
|
|
26
|
+
|
|
27
|
+
```ruby
|
|
28
|
+
require "llm"
|
|
29
|
+
|
|
30
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
31
|
+
res = llm.images.create(prompt: "a dog on a rocket to the moon")
|
|
32
|
+
IO.copy_stream res.images[0], "dogrocket.png"
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
#### Why would I use it?
|
|
36
|
+
|
|
37
|
+
Image generation and editing let the model produce visual output
|
|
38
|
+
directly from your prompts. The API is the same across providers,
|
|
39
|
+
so switching between OpenAI and xAI requires changing one line.
|
|
40
|
+
Prototype visual concepts, generate assets, or augment datasets
|
|
41
|
+
without wiring up a separate image API.
|
|
42
|
+
|
|
43
|
+
#### Notes
|
|
44
|
+
|
|
45
|
+
Google only supports image generation, not edits. DeepSeek
|
|
46
|
+
generates SVGs rather than raster images. DeepSeek can also
|
|
47
|
+
maintain a session across multiple generations through the
|
|
48
|
+
`agent` parameter on the response object.
|
|
49
|
+
|
|
50
|
+
### Editing
|
|
51
|
+
|
|
52
|
+
#### Overview
|
|
53
|
+
|
|
54
|
+
Image editing takes an existing image and a text prompt, then
|
|
55
|
+
produces a modified version. You can add objects, change colours,
|
|
56
|
+
or alter the scene while keeping the original composition.
|
|
57
|
+
OpenAI, xAI, and DeepInfra support raster image edits. DeepSeek
|
|
58
|
+
generates SVG documents which can be refined with follow-up
|
|
59
|
+
text-to-image prompts, giving you iterative vector editing.
|
|
60
|
+
|
|
61
|
+
#### How it works
|
|
62
|
+
|
|
63
|
+
The
|
|
64
|
+
[`LLM::Images#edit`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#edit)
|
|
65
|
+
method takes a prompt and an image path. It
|
|
66
|
+
returns a modified image as a
|
|
67
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
68
|
+
object. Copy the result
|
|
69
|
+
to a file the same way you would with generated images.
|
|
70
|
+
|
|
71
|
+
```ruby
|
|
72
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
73
|
+
res = llm.images.edit(prompt: "add a mustache", image: "self.jpg")
|
|
74
|
+
IO.copy_stream res.images[0], "mustache.png"
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
#### Why would I use it?
|
|
78
|
+
|
|
79
|
+
Editing lets the model modify existing images rather than starting
|
|
80
|
+
from scratch. DeepSeek's SVG output is particularly useful here
|
|
81
|
+
because vector graphics can be refined iteratively. Make targeted
|
|
82
|
+
adjustments: add objects, change colours, or alter the scene,
|
|
83
|
+
all while keeping the original composition intact.
|
|
84
|
+
|
|
85
|
+
#### Notes
|
|
86
|
+
|
|
87
|
+
Google does not support image edits. DeepSeek can maintain a
|
|
88
|
+
session across multiple generations through the `agent` parameter
|
|
89
|
+
on the response object.
|
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
|
|
2
|
+
## LLM::Object
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
9
|
+
is the hash-like object that llm.rb uses everywhere structured data
|
|
10
|
+
flows through the runtime. Response bodies, tool arguments, schema
|
|
11
|
+
results, usage and cost data, and function parameters all come back
|
|
12
|
+
as [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
13
|
+
instances. It is similar in spirit to OpenStruct,
|
|
14
|
+
and it was introduced after OpenStruct became a bundled gem rather
|
|
15
|
+
than a default gem in Ruby 3.5.
|
|
16
|
+
|
|
17
|
+
#### How it works
|
|
18
|
+
|
|
19
|
+
When you want to read a value from an
|
|
20
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html),
|
|
21
|
+
use either method-style or bracket access. Keys are indifferent, so
|
|
22
|
+
strings and symbols work interchangeably:
|
|
23
|
+
|
|
24
|
+
```ruby
|
|
25
|
+
obj = LLM::Object.from(city: "Paris", temperature: 15.0)
|
|
26
|
+
|
|
27
|
+
obj.city # => "Paris"
|
|
28
|
+
obj["city"] # => "Paris"
|
|
29
|
+
obj[:city] # => "Paris"
|
|
30
|
+
obj[:temperature] # => 15.0
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
Nested hashes and arrays are converted recursively, so deep chains
|
|
34
|
+
read naturally:
|
|
35
|
+
|
|
36
|
+
```ruby
|
|
37
|
+
obj = LLM::Object.from(person: {name: "John"})
|
|
38
|
+
obj.person.name # => "John"
|
|
39
|
+
obj.person.class # => LLM::Object
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
#### Why would I use it?
|
|
43
|
+
|
|
44
|
+
Most of the time you do not construct
|
|
45
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
46
|
+
instances yourself. They come back from `talk`, `ask`, `embed`, and
|
|
47
|
+
every other call that returns structured data. Knowing how they
|
|
48
|
+
behave lets you read response fields, pass tool arguments, and
|
|
49
|
+
inspect usage without reaching for `to_h` on every line.
|
|
50
|
+
|
|
51
|
+
#### Notes
|
|
52
|
+
|
|
53
|
+
An
|
|
54
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
55
|
+
is enumerable and supports the usual Hash operations: `keys`,
|
|
56
|
+
`values`, `key?`, `fetch`, `dig`, `slice`, `merge`, `merge!`,
|
|
57
|
+
`delete`, `size`, and `empty?`. Use `to_h` for a plain Hash and
|
|
58
|
+
`to_hash` for one with symbol keys. A missing key returns `nil`
|
|
59
|
+
rather than raising. Because it subclasses `BasicObject`, `to_json`
|
|
60
|
+
is defined explicitly and serializes through the configured JSON
|
|
61
|
+
adapter via
|
|
62
|
+
[`LLM.json.dump`](https://r.uby.dev/api-docs/llm.rb/LLM.html#json-class_method).
|
|
63
|
+
|
|
64
|
+
### Reading and writing
|
|
65
|
+
|
|
66
|
+
#### Overview
|
|
67
|
+
|
|
68
|
+
Beyond simple reads,
|
|
69
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
70
|
+
supports assignment, mutation, and iteration, so you can treat it
|
|
71
|
+
like a Hash in place.
|
|
72
|
+
|
|
73
|
+
#### How it works
|
|
74
|
+
|
|
75
|
+
Assign values with method or bracket syntax, mutate in place, and
|
|
76
|
+
iterate like a Hash:
|
|
77
|
+
|
|
78
|
+
```ruby
|
|
79
|
+
obj = LLM::Object.from({})
|
|
80
|
+
|
|
81
|
+
obj.city = "Paris" # method-style write
|
|
82
|
+
obj["country"] = "France"
|
|
83
|
+
|
|
84
|
+
obj.key?(:city) # => true
|
|
85
|
+
obj.keys # => ["city", "country"]
|
|
86
|
+
obj.merge!(population: 2_100_000)
|
|
87
|
+
|
|
88
|
+
obj.each { |key, value| puts "#{key}: #{value}" }
|
|
89
|
+
obj.transform_values!(&:upcase) if obj.any?
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
#### Why would I use it?
|
|
93
|
+
|
|
94
|
+
Tool implementations receive their arguments as an
|
|
95
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
|
|
96
|
+
and often build a result Hash from them. Merging defaults, deleting
|
|
97
|
+
optional keys, and transforming values in place keeps that code
|
|
98
|
+
concise without converting back and forth between Hash and
|
|
99
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html).
|
|
100
|
+
|
|
101
|
+
#### Notes
|
|
102
|
+
|
|
103
|
+
`merge` returns a new
|
|
104
|
+
[`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html);
|
|
105
|
+
`merge!` mutates in place.
|
|
106
|
+
Assignment always stores the key as a string internally, which is
|
|
107
|
+
why string and symbol lookups both work. Equality compares against
|
|
108
|
+
anything that responds to `to_h`.
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
|
|
2
|
+
## OCR
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
OCR pulls text out of images and PDFs so you can search, index,
|
|
9
|
+
or process scanned content. Mistral is the only provider with a
|
|
10
|
+
dedicated OCR endpoint, and it accepts both image URLs and
|
|
11
|
+
document URLs.
|
|
12
|
+
|
|
13
|
+
#### How it works
|
|
14
|
+
|
|
15
|
+
Mistral is the only provider with a dedicated OCR endpoint. The
|
|
16
|
+
[`LLM::Provider#ocr`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#ocr)
|
|
17
|
+
method accepts an `image_url:` or `document_url:` parameter.
|
|
18
|
+
Document URLs can point to PDFs.
|
|
19
|
+
|
|
20
|
+
```ruby
|
|
21
|
+
require "llm"
|
|
22
|
+
|
|
23
|
+
llm = LLM.mistral(key: ENV["KEY"])
|
|
24
|
+
|
|
25
|
+
# Extract text from an image
|
|
26
|
+
res = llm.ocr(image_url: "https://example.com/photo.png")
|
|
27
|
+
res.pages.each { |page| puts page.markdown }
|
|
28
|
+
|
|
29
|
+
# Extract text from a PDF
|
|
30
|
+
res = llm.ocr(document_url: "https://example.com/report.pdf")
|
|
31
|
+
res.pages.each { |page| puts page.markdown }
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
#### Why would I use it?
|
|
35
|
+
|
|
36
|
+
OCR extracts text from scanned documents and images. The result
|
|
37
|
+
feeds into search indexes, extraction pipelines that pull out dates
|
|
38
|
+
and amounts, or archives so scanned contracts become searchable
|
|
39
|
+
records. PDFs and images both work, and the response is structured
|
|
40
|
+
per page with markdown.
|
|
41
|
+
|
|
42
|
+
#### Notes
|
|
43
|
+
|
|
44
|
+
Only Mistral currently supports OCR through the llm.rb runtime.
|
|
45
|
+
The response exposes pages through
|
|
46
|
+
[`LLM::OCR::Response#pages`](https://r.uby.dev/api-docs/llm.rb/LLM/OCR/Response.html#pages),
|
|
47
|
+
where each page
|
|
48
|
+
has a `markdown` field containing the extracted text.
|
|
@@ -0,0 +1,202 @@
|
|
|
1
|
+
|
|
2
|
+
## Agents
|
|
3
|
+
### Introduction
|
|
4
|
+
|
|
5
|
+
#### Overview
|
|
6
|
+
|
|
7
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
8
|
+
is the recommended entry point for most use-cases. It provides a
|
|
9
|
+
class-level DSL for defining reusable, preconfigured assistants
|
|
10
|
+
with defaults for model, tools, schema, and instructions. Under
|
|
11
|
+
the hood it delegates to
|
|
12
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html),
|
|
13
|
+
so it has the same runtime surface: message history, streaming,
|
|
14
|
+
serialization, compaction, and concurrency.
|
|
15
|
+
|
|
16
|
+
#### How it works
|
|
17
|
+
|
|
18
|
+
An agent holds a conversation with a model. You send input with
|
|
19
|
+
[`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk),
|
|
20
|
+
the model responds, and if it requests tools the agent
|
|
21
|
+
executes them automatically and feeds the results back. It enables
|
|
22
|
+
[a loop guard by default](https://r.uby.dev/llm/deepdive/advanced/guard)
|
|
23
|
+
that detects repeated tool-call patterns
|
|
24
|
+
and blocks stuck execution. The tool loop can also be bounded with
|
|
25
|
+
[`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
|
|
26
|
+
(see the Tool budget section). Instructions are injected once
|
|
27
|
+
unless a system message is already present.
|
|
28
|
+
|
|
29
|
+
#### Why would I use it?
|
|
30
|
+
|
|
31
|
+
Agents manage the tool loop for you. They guard against infinite
|
|
32
|
+
loops, keep conversation state across turns, and let you define
|
|
33
|
+
reusable configurations at the class level. If you need manual
|
|
34
|
+
control over the tool loop, use
|
|
35
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
36
|
+
directly instead.
|
|
37
|
+
|
|
38
|
+
#### Notes
|
|
39
|
+
|
|
40
|
+
Agents support the same concurrency strategies, compaction,
|
|
41
|
+
cancellation, and serialization as contexts. The trade-off between
|
|
42
|
+
a subclass and a direct instance is only in how the agent is
|
|
43
|
+
organized, not in what it can do. Tool loop execution can be
|
|
44
|
+
configured with `concurrency: :sequential`, `:thread`, `:async`,
|
|
45
|
+
`:fiber`, `:fork`, or `:ractor`.
|
|
46
|
+
|
|
47
|
+
### Class-based
|
|
48
|
+
|
|
49
|
+
#### Overview
|
|
50
|
+
|
|
51
|
+
A subclass of
|
|
52
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
53
|
+
gives you a reusable agent with its own behavior. You define the model, tools, and other
|
|
54
|
+
attributes at the class level, and each instance picks them up
|
|
55
|
+
as defaults. Attributes can be overridden per-instance, and they
|
|
56
|
+
can be plain values, blocks, or Symbols that resolve to methods.
|
|
57
|
+
The class becomes a self-contained worker that you can instantiate
|
|
58
|
+
and talk to from anywhere.
|
|
59
|
+
|
|
60
|
+
#### How it works
|
|
61
|
+
|
|
62
|
+
A subclass declares its defaults with
|
|
63
|
+
[`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method).
|
|
64
|
+
Each key is a
|
|
65
|
+
class-level accessor: `name`, `description`, `model`, `tools`,
|
|
66
|
+
`skills`, `instructions`, `stream`, `tracer`, `concurrency`,
|
|
67
|
+
`schema`, `confirm`, `path`, `tool_budget`.
|
|
68
|
+
Keyword arguments in the constructor override these defaults.
|
|
69
|
+
|
|
70
|
+
```ruby
|
|
71
|
+
class Agent < LLM::Agent
|
|
72
|
+
set model: "deepseek-v4-pro",
|
|
73
|
+
description: "system administration agent",
|
|
74
|
+
tools: [Shell]
|
|
75
|
+
end
|
|
76
|
+
|
|
77
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
78
|
+
agent = Agent.new(llm)
|
|
79
|
+
agent.talk "Run 'date'"
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
#### Why would I use it?
|
|
83
|
+
|
|
84
|
+
A subclass is useful when multiple parts of an application need to
|
|
85
|
+
call the same agent. The configuration and any helper methods live in one place. Define a
|
|
86
|
+
`research!` method that kicks off the agent's work. The subclass
|
|
87
|
+
becomes a self-contained worker.
|
|
88
|
+
|
|
89
|
+
#### Notes
|
|
90
|
+
|
|
91
|
+
Attributes passed to
|
|
92
|
+
[`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
|
|
93
|
+
can be plain values, blocks, or
|
|
94
|
+
Symbols. A Symbol is evaluated as an instance method on the
|
|
95
|
+
subclass, so `tracer: :set_tracer` calls `set_tracer` on the
|
|
96
|
+
instance. A block like `stream: -> { $stdout }` is evaluated
|
|
97
|
+
when the attribute is first accessed.
|
|
98
|
+
|
|
99
|
+
Set `path:` on a subclass or instance for automatic filesystem
|
|
100
|
+
persistence; the agent restores conversation history from the
|
|
101
|
+
file on startup and saves it back after every turn with no
|
|
102
|
+
manual
|
|
103
|
+
[`LLM::Agent#save`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#save)/
|
|
104
|
+
[`LLM::Agent#restore`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#restore)
|
|
105
|
+
calls. See the
|
|
106
|
+
[database deepdive](../fundamentals/database.md) for details.
|
|
107
|
+
|
|
108
|
+
### Object-based
|
|
109
|
+
|
|
110
|
+
#### Overview
|
|
111
|
+
|
|
112
|
+
An
|
|
113
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
114
|
+
instance is the simplest way to get started. You pass a provider and any configuration as keyword
|
|
115
|
+
arguments, and the agent runs the tool loop and manages state
|
|
116
|
+
just like a subclass would. This is the right choice when you
|
|
117
|
+
are prototyping, running a one-off task, or when the agent's
|
|
118
|
+
configuration is determined at runtime.
|
|
119
|
+
|
|
120
|
+
#### How it works
|
|
121
|
+
|
|
122
|
+
A direct instance takes the same attributes as keyword arguments
|
|
123
|
+
to
|
|
124
|
+
[`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#initialize-instance_method).
|
|
125
|
+
The first argument is always the provider.
|
|
126
|
+
Everything else is optional. The agent runs the tool loop
|
|
127
|
+
and manages state under the hood through a
|
|
128
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html).
|
|
129
|
+
|
|
130
|
+
```ruby
|
|
131
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
132
|
+
agent = LLM::Agent.new(llm, stream: $stdout)
|
|
133
|
+
agent.talk "Hello world"
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
#### Why would I use it?
|
|
137
|
+
|
|
138
|
+
A direct instance is the right choice for quick experiments, one-shot
|
|
139
|
+
tasks, or when defining a class would be overkill. It is also
|
|
140
|
+
the right choice when the agent's configuration is determined at
|
|
141
|
+
runtime and a class hierarchy adds unnecessary complexity.
|
|
142
|
+
|
|
143
|
+
#### Notes
|
|
144
|
+
|
|
145
|
+
Direct instances accept all the same options as subclasses. The
|
|
146
|
+
difference is only in how the agent is organized, not in what it
|
|
147
|
+
can do. Under the hood,
|
|
148
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
|
|
149
|
+
creates a
|
|
150
|
+
[`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
|
|
151
|
+
that manages the message history.
|
|
152
|
+
|
|
153
|
+
### Tool budget
|
|
154
|
+
|
|
155
|
+
#### Overview
|
|
156
|
+
|
|
157
|
+
[`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
|
|
158
|
+
caps the number of tool calls allowed in a single turn. Once the
|
|
159
|
+
budget is spent, the agent sends an in-band advisory message back
|
|
160
|
+
through the model instead of running more tools. By default no
|
|
161
|
+
budget is set, so the feature is disabled.
|
|
162
|
+
|
|
163
|
+
#### How it works
|
|
164
|
+
|
|
165
|
+
When you want to bound how many tools an agent can call in one
|
|
166
|
+
turn, set the budget with
|
|
167
|
+
[`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method).
|
|
168
|
+
The budget can be a plain number or a block evaluated against the
|
|
169
|
+
agent instance. Once the agent has made the budgeted number of tool
|
|
170
|
+
calls, it stops and returns an in-band advisory message describing
|
|
171
|
+
the spent budget. A model will usually change course afterwards:
|
|
172
|
+
|
|
173
|
+
```ruby
|
|
174
|
+
class Researcher < LLM::Agent
|
|
175
|
+
set model: "deepseek-v4-pro",
|
|
176
|
+
tools: [FetchNews, FetchStocks],
|
|
177
|
+
tool_budget: 5
|
|
178
|
+
end
|
|
179
|
+
|
|
180
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
181
|
+
agent = Researcher.new(llm)
|
|
182
|
+
agent.talk "Research the market"
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
#### Why would I use it?
|
|
186
|
+
|
|
187
|
+
A tool budget prevents runaway tool loops. A misbehaving model can
|
|
188
|
+
otherwise keep calling tools, spending tokens on every round trip.
|
|
189
|
+
Capping the budget turns that into a bounded conversation: after
|
|
190
|
+
the cap, the model is told it has run out of tool calls and must
|
|
191
|
+
respond from what it has.
|
|
192
|
+
|
|
193
|
+
#### Notes
|
|
194
|
+
|
|
195
|
+
The budget is disabled by default (`nil`). Set it on a subclass
|
|
196
|
+
with the `tool_budget` DSL, through
|
|
197
|
+
[`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
|
|
198
|
+
with `tool_budget:`, or per-instance with the `tool_budget:` keyword
|
|
199
|
+
argument to
|
|
200
|
+
[`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#initialize-instance_method).
|
|
201
|
+
This replaces the previous `tool_attempts` parameter, which is no
|
|
202
|
+
longer used.
|