llm.rb 13.1.0 → 14.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (90) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +320 -0
  3. data/README.md +340 -31
  4. data/bin/llm.rb +36 -12
  5. data/data/anthropic.json +206 -263
  6. data/data/bedrock.json +2138 -1860
  7. data/data/deepinfra.json +1003 -624
  8. data/data/deepseek.json +38 -34
  9. data/data/google.json +1079 -371
  10. data/data/mistral.json +448 -368
  11. data/data/moonshot.json +384 -0
  12. data/data/openai.json +974 -1343
  13. data/data/xai.json +154 -126
  14. data/data/zai.json +191 -191
  15. data/lib/llm/agent.rb +47 -14
  16. data/lib/llm/context.rb +71 -88
  17. data/lib/llm/cost.rb +23 -17
  18. data/lib/llm/error.rb +0 -8
  19. data/lib/llm/function/async/task.rb +2 -0
  20. data/lib/llm/function/fiber/task.rb +2 -0
  21. data/lib/llm/function/fork/task.rb +2 -0
  22. data/lib/llm/function/ractor/task.rb +2 -0
  23. data/lib/llm/function/sequential/group.rb +4 -1
  24. data/lib/llm/function/sequential/task.rb +1 -1
  25. data/lib/llm/function/task.rb +4 -0
  26. data/lib/llm/function/thread/task.rb +2 -0
  27. data/lib/llm/function.rb +32 -4
  28. data/lib/llm/guard/loop.rb +89 -0
  29. data/lib/llm/guard/null.rb +19 -0
  30. data/lib/llm/guard.rb +61 -0
  31. data/lib/llm/provider.rb +36 -0
  32. data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
  33. data/lib/llm/providers/anthropic.rb +1 -8
  34. data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
  35. data/lib/llm/providers/bedrock.rb +1 -8
  36. data/lib/llm/providers/google/stream_parser.rb +1 -0
  37. data/lib/llm/providers/google.rb +1 -8
  38. data/lib/llm/providers/moonshot.rb +76 -0
  39. data/lib/llm/providers/ollama.rb +1 -8
  40. data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
  41. data/lib/llm/providers/openai/responses.rb +6 -8
  42. data/lib/llm/providers/openai/stream_parser.rb +1 -0
  43. data/lib/llm/providers/openai.rb +3 -10
  44. data/lib/llm/repl/bar.rb +4 -3
  45. data/lib/llm/repl/buffer.rb +42 -15
  46. data/lib/llm/repl/color.rb +78 -0
  47. data/lib/llm/repl/input/char.rb +46 -0
  48. data/lib/llm/repl/input/row.rb +39 -0
  49. data/lib/llm/repl/input.rb +251 -66
  50. data/lib/llm/repl/markdown/table.rb +6 -2
  51. data/lib/llm/repl/markdown.rb +31 -5
  52. data/lib/llm/repl/status.rb +38 -3
  53. data/lib/llm/repl/stream.rb +16 -4
  54. data/lib/llm/repl/walker.rb +3 -2
  55. data/lib/llm/repl/window.rb +25 -5
  56. data/lib/llm/repl.rb +29 -13
  57. data/lib/llm/stream.rb +8 -7
  58. data/lib/llm/tool.rb +29 -0
  59. data/lib/llm/transformer/null.rb +21 -0
  60. data/lib/llm/transformer.rb +55 -0
  61. data/lib/llm/version.rb +1 -1
  62. data/lib/llm.rb +12 -2
  63. data/llm.gemspec +1 -0
  64. data/resources/deepdive/advanced/cancellation.md +74 -0
  65. data/resources/deepdive/advanced/compaction.md +83 -0
  66. data/resources/deepdive/advanced/context.md +267 -0
  67. data/resources/deepdive/advanced/guard.md +371 -0
  68. data/resources/deepdive/advanced/tracer.md +180 -0
  69. data/resources/deepdive/advanced/transformer.md +67 -0
  70. data/resources/deepdive/advanced/transports.md +45 -0
  71. data/resources/deepdive/everything_else/audio.md +122 -0
  72. data/resources/deepdive/everything_else/cost.md +99 -0
  73. data/resources/deepdive/everything_else/images.md +89 -0
  74. data/resources/deepdive/everything_else/object.md +108 -0
  75. data/resources/deepdive/everything_else/ocr.md +48 -0
  76. data/resources/deepdive/fundamentals/agents.md +202 -0
  77. data/resources/deepdive/fundamentals/builtin_tools.md +191 -0
  78. data/resources/deepdive/fundamentals/concurrency.md +104 -0
  79. data/resources/deepdive/fundamentals/database.md +449 -0
  80. data/resources/deepdive/fundamentals/embeddings.md +157 -0
  81. data/resources/deepdive/fundamentals/repl.md +87 -0
  82. data/resources/deepdive/fundamentals/schema.md +61 -0
  83. data/resources/deepdive/fundamentals/skills.md +106 -0
  84. data/resources/deepdive/fundamentals/stream.md +110 -0
  85. data/resources/deepdive/fundamentals/tools.md +265 -0
  86. data/resources/deepdive/protocols/a2a.md +106 -0
  87. data/resources/deepdive/protocols/mcp.md +111 -0
  88. data/resources/deepdive.md +7 -1
  89. metadata +36 -3
  90. data/lib/llm/loop_guard.rb +0 -107
@@ -0,0 +1,122 @@
1
+
2
+ ## Audio
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ The audio interface covers three things: turning text into speech,
9
+ transcribing audio into text, and translating spoken language.
10
+ OpenAI supports all three. Google and DeepInfra support subsets.
11
+ Each method follows the same pattern: pass input, get output,
12
+ copy the result somewhere useful.
13
+
14
+ #### How it works
15
+
16
+ When you want to convert text to speech, call
17
+ [`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech).
18
+ The provider returns an audio clip as a
19
+ [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
20
+ object. The generated audio can be copied
21
+ to a file or streamed directly. Each provider supports different
22
+ output formats and voice options:
23
+
24
+ ```ruby
25
+ llm = LLM.openai(key: ENV["KEY"])
26
+ res = llm.audio.create_speech(input: "Hello world")
27
+ IO.copy_stream res.audio.decoded, "helloworld.mp3"
28
+ ```
29
+
30
+ #### Why would I use it?
31
+
32
+ Audio support lets you build voice interfaces, add accessibility
33
+ features, and work across languages without wiring up a separate
34
+ speech service. Generate audio for notifications, narrate written
35
+ content, or add speech output to an existing application with a
36
+ single method call.
37
+
38
+ #### Notes
39
+
40
+ OpenAI has full audio support. Google and DeepInfra have partial
41
+ support. The
42
+ [`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech)
43
+ method returns a
44
+ [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
45
+ object.
46
+
47
+ ### Transcription
48
+
49
+ #### Overview
50
+
51
+ Transcription turns an audio file into text. Pass a file path
52
+ to
53
+ [`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription)
54
+ and get back the spoken content as
55
+ a string. OpenAI, Google, and DeepInfra support it. Transcribe
56
+ meeting notes, voice memos, or podcast episodes for search and
57
+ processing. The response is plain text you can feed into any
58
+ downstream pipeline.
59
+
60
+ #### How it works
61
+
62
+ When you want to transcribe audio into text, call
63
+ [`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription).
64
+ The provider processes the audio and returns the transcribed text.
65
+ The file can be a local path or a URL depending on provider support:
66
+
67
+ ```ruby
68
+ llm = LLM.google(key: ENV["KEY"])
69
+ res = llm.audio.create_transcription(file: "helloworld.mp3")
70
+ res.text # => "Hello world"
71
+ ```
72
+
73
+ #### Why would I use it?
74
+
75
+ Transcribe recorded meetings, voice memos, or podcast episodes for
76
+ search and processing. The returned text feeds directly into search
77
+ indexes, summarization pipelines, or downstream extraction. Spoken
78
+ content becomes as queryable as written text.
79
+
80
+ #### Notes
81
+
82
+ OpenAI has full audio support. Google and DeepInfra have partial
83
+ support.
84
+
85
+ ### Translation
86
+
87
+ #### Overview
88
+
89
+ Translation transcribes audio and translates it into English in
90
+ one step. Pass a file to
91
+ [`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation)
92
+ and get back the
93
+ translated text. OpenAI and Google support it. Translate
94
+ multilingual podcasts, interviews, or any audio where you need
95
+ the content in English without running a separate translation
96
+ pipeline.
97
+
98
+ #### How it works
99
+
100
+ When you want to translate spoken audio into English, call
101
+ [`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation).
102
+ The provider transcribes the spoken language and translates the
103
+ result into English in a single operation. The returned text is the
104
+ English translation:
105
+
106
+ ```ruby
107
+ llm = LLM.google(key: ENV["KEY"])
108
+ res = llm.audio.create_translation(file: "bomdia.mp3")
109
+ res.text # => "Good day"
110
+ ```
111
+
112
+ #### Why would I use it?
113
+
114
+ Translate a podcast or interview into another language without a
115
+ separate speech service. The provider handles both transcription and
116
+ translation in a single call, so you get English text from any
117
+ supported source language without chaining two operations together.
118
+
119
+ #### Notes
120
+
121
+ Each method works independently and each provider supports a
122
+ different subset.
@@ -0,0 +1,99 @@
1
+
2
+ ## LLM::Cost
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ [`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
9
+ represents the approximate cost of a conversation. It breaks the
10
+ total down by token type, so you can see how much was spent on input,
11
+ output, cached tokens, reasoning, audio, and images. Cost is computed
12
+ from token usage and the pricing data shipped in the model registry.
13
+
14
+ #### How it works
15
+
16
+ When you want to know what a conversation cost so far, call
17
+ [`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
18
+ (or
19
+ [`LLM::Agent#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#cost-instance_method))
20
+ and read the breakdown. The REPL shows this live in its status bar
21
+ after every turn:
22
+
23
+ ```ruby
24
+ llm = LLM.deepseek(key: ENV["KEY"])
25
+ ctx = LLM::Context.new(llm)
26
+ ctx.talk "Hello"
27
+
28
+ cost = ctx.cost
29
+ cost.input # => 0.0000042
30
+ cost.output # => 0.0000084
31
+ cost.total # => 0.0000126
32
+ cost.to_s # => "0.0000126"
33
+ ```
34
+
35
+ #### Why would I use it?
36
+
37
+ Cost tracking matters in production. Monitoring spend per
38
+ conversation, per agent, or per provider tells you which workflows
39
+ are expensive and when to switch models. Log a structured breakdown
40
+ with
41
+ [`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method),
42
+ or read
43
+ [`LLM::Cost#total`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#total-instance_method)
44
+ for a single number.
45
+
46
+ #### Notes
47
+
48
+ Cost is an approximation based on the pricing in the model registry.
49
+ [`LLM::Cost.from`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#from-class_method)
50
+ returns an empty cost when the model or registry cannot be found, so
51
+ a missing model never crashes your code.
52
+
53
+ ### Reading the breakdown
54
+
55
+ #### Overview
56
+
57
+ [`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
58
+ exposes each cost component as a reader, plus
59
+ [`LLM::Cost#total`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#total-instance_method),
60
+ [`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method),
61
+ and
62
+ [`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_s-instance_method).
63
+
64
+ #### How it works
65
+
66
+ Each component is a Float, or `nil` when no tokens of that type were
67
+ used. The
68
+ [`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method)
69
+ method returns a Hash with only the non-nil components and the total:
70
+
71
+ ```ruby
72
+ cost = ctx.cost
73
+
74
+ cost.input
75
+ cost.output
76
+ cost.cache_read
77
+ cost.cache_write
78
+ cost.reasoning
79
+ cost.input_audio
80
+ cost.output_audio
81
+ cost.input_image
82
+
83
+ cost.to_h # => {input: 4.2e-06, output: 8.4e-06, total: 1.26e-05}
84
+ ```
85
+
86
+ #### Why would I use it?
87
+
88
+ The per-component breakdown shows where the money goes. High cache
89
+ read costs suggest a conversation benefits from prompt caching.
90
+ High reasoning costs point at a model that thinks a lot. Log
91
+ [`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method)
92
+ at the end of a session to keep a spend trail.
93
+
94
+ #### Notes
95
+
96
+ [`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_s-instance_method)
97
+ returns the total in a compact, human-friendly format
98
+ (`"0.0000126"`). Components that were not used are `nil`, so sum
99
+ them with `compact` if you aggregate across conversations.
@@ -0,0 +1,89 @@
1
+
2
+ ## Images
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ A handful of providers can generate images from a text prompt.
9
+ OpenAI, Google, xAI, DeepInfra, and DeepSeek all support it.
10
+ OpenAI, xAI, and DeepInfra also let you edit existing images.
11
+ The API is the same across providers, so switching between them
12
+ requires no code changes.
13
+
14
+ #### How it works
15
+
16
+ The
17
+ [`LLM::Images#create`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#create)
18
+ method sends a prompt to the provider and
19
+ returns the result as a
20
+ [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
21
+ object. The same API works
22
+ across providers: swap
23
+ [`LLM.openai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#openai-class_method) for
24
+ [`LLM.xai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#xai-class_method) and the rest
25
+ of the code is identical.
26
+
27
+ ```ruby
28
+ require "llm"
29
+
30
+ llm = LLM.openai(key: ENV["KEY"])
31
+ res = llm.images.create(prompt: "a dog on a rocket to the moon")
32
+ IO.copy_stream res.images[0], "dogrocket.png"
33
+ ```
34
+
35
+ #### Why would I use it?
36
+
37
+ Image generation and editing let the model produce visual output
38
+ directly from your prompts. The API is the same across providers,
39
+ so switching between OpenAI and xAI requires changing one line.
40
+ Prototype visual concepts, generate assets, or augment datasets
41
+ without wiring up a separate image API.
42
+
43
+ #### Notes
44
+
45
+ Google only supports image generation, not edits. DeepSeek
46
+ generates SVGs rather than raster images. DeepSeek can also
47
+ maintain a session across multiple generations through the
48
+ `agent` parameter on the response object.
49
+
50
+ ### Editing
51
+
52
+ #### Overview
53
+
54
+ Image editing takes an existing image and a text prompt, then
55
+ produces a modified version. You can add objects, change colours,
56
+ or alter the scene while keeping the original composition.
57
+ OpenAI, xAI, and DeepInfra support raster image edits. DeepSeek
58
+ generates SVG documents which can be refined with follow-up
59
+ text-to-image prompts, giving you iterative vector editing.
60
+
61
+ #### How it works
62
+
63
+ The
64
+ [`LLM::Images#edit`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#edit)
65
+ method takes a prompt and an image path. It
66
+ returns a modified image as a
67
+ [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
68
+ object. Copy the result
69
+ to a file the same way you would with generated images.
70
+
71
+ ```ruby
72
+ llm = LLM.openai(key: ENV["KEY"])
73
+ res = llm.images.edit(prompt: "add a mustache", image: "self.jpg")
74
+ IO.copy_stream res.images[0], "mustache.png"
75
+ ```
76
+
77
+ #### Why would I use it?
78
+
79
+ Editing lets the model modify existing images rather than starting
80
+ from scratch. DeepSeek's SVG output is particularly useful here
81
+ because vector graphics can be refined iteratively. Make targeted
82
+ adjustments: add objects, change colours, or alter the scene,
83
+ all while keeping the original composition intact.
84
+
85
+ #### Notes
86
+
87
+ Google does not support image edits. DeepSeek can maintain a
88
+ session across multiple generations through the `agent` parameter
89
+ on the response object.
@@ -0,0 +1,108 @@
1
+
2
+ ## LLM::Object
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
9
+ is the hash-like object that llm.rb uses everywhere structured data
10
+ flows through the runtime. Response bodies, tool arguments, schema
11
+ results, usage and cost data, and function parameters all come back
12
+ as [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
13
+ instances. It is similar in spirit to OpenStruct,
14
+ and it was introduced after OpenStruct became a bundled gem rather
15
+ than a default gem in Ruby 3.5.
16
+
17
+ #### How it works
18
+
19
+ When you want to read a value from an
20
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html),
21
+ use either method-style or bracket access. Keys are indifferent, so
22
+ strings and symbols work interchangeably:
23
+
24
+ ```ruby
25
+ obj = LLM::Object.from(city: "Paris", temperature: 15.0)
26
+
27
+ obj.city # => "Paris"
28
+ obj["city"] # => "Paris"
29
+ obj[:city] # => "Paris"
30
+ obj[:temperature] # => 15.0
31
+ ```
32
+
33
+ Nested hashes and arrays are converted recursively, so deep chains
34
+ read naturally:
35
+
36
+ ```ruby
37
+ obj = LLM::Object.from(person: {name: "John"})
38
+ obj.person.name # => "John"
39
+ obj.person.class # => LLM::Object
40
+ ```
41
+
42
+ #### Why would I use it?
43
+
44
+ Most of the time you do not construct
45
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
46
+ instances yourself. They come back from `talk`, `ask`, `embed`, and
47
+ every other call that returns structured data. Knowing how they
48
+ behave lets you read response fields, pass tool arguments, and
49
+ inspect usage without reaching for `to_h` on every line.
50
+
51
+ #### Notes
52
+
53
+ An
54
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
55
+ is enumerable and supports the usual Hash operations: `keys`,
56
+ `values`, `key?`, `fetch`, `dig`, `slice`, `merge`, `merge!`,
57
+ `delete`, `size`, and `empty?`. Use `to_h` for a plain Hash and
58
+ `to_hash` for one with symbol keys. A missing key returns `nil`
59
+ rather than raising. Because it subclasses `BasicObject`, `to_json`
60
+ is defined explicitly and serializes through the configured JSON
61
+ adapter via
62
+ [`LLM.json.dump`](https://r.uby.dev/api-docs/llm.rb/LLM.html#json-class_method).
63
+
64
+ ### Reading and writing
65
+
66
+ #### Overview
67
+
68
+ Beyond simple reads,
69
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
70
+ supports assignment, mutation, and iteration, so you can treat it
71
+ like a Hash in place.
72
+
73
+ #### How it works
74
+
75
+ Assign values with method or bracket syntax, mutate in place, and
76
+ iterate like a Hash:
77
+
78
+ ```ruby
79
+ obj = LLM::Object.from({})
80
+
81
+ obj.city = "Paris" # method-style write
82
+ obj["country"] = "France"
83
+
84
+ obj.key?(:city) # => true
85
+ obj.keys # => ["city", "country"]
86
+ obj.merge!(population: 2_100_000)
87
+
88
+ obj.each { |key, value| puts "#{key}: #{value}" }
89
+ obj.transform_values!(&:upcase) if obj.any?
90
+ ```
91
+
92
+ #### Why would I use it?
93
+
94
+ Tool implementations receive their arguments as an
95
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
96
+ and often build a result Hash from them. Merging defaults, deleting
97
+ optional keys, and transforming values in place keeps that code
98
+ concise without converting back and forth between Hash and
99
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html).
100
+
101
+ #### Notes
102
+
103
+ `merge` returns a new
104
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html);
105
+ `merge!` mutates in place.
106
+ Assignment always stores the key as a string internally, which is
107
+ why string and symbol lookups both work. Equality compares against
108
+ anything that responds to `to_h`.
@@ -0,0 +1,48 @@
1
+
2
+ ## OCR
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ OCR pulls text out of images and PDFs so you can search, index,
9
+ or process scanned content. Mistral is the only provider with a
10
+ dedicated OCR endpoint, and it accepts both image URLs and
11
+ document URLs.
12
+
13
+ #### How it works
14
+
15
+ Mistral is the only provider with a dedicated OCR endpoint. The
16
+ [`LLM::Provider#ocr`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#ocr)
17
+ method accepts an `image_url:` or `document_url:` parameter.
18
+ Document URLs can point to PDFs.
19
+
20
+ ```ruby
21
+ require "llm"
22
+
23
+ llm = LLM.mistral(key: ENV["KEY"])
24
+
25
+ # Extract text from an image
26
+ res = llm.ocr(image_url: "https://example.com/photo.png")
27
+ res.pages.each { |page| puts page.markdown }
28
+
29
+ # Extract text from a PDF
30
+ res = llm.ocr(document_url: "https://example.com/report.pdf")
31
+ res.pages.each { |page| puts page.markdown }
32
+ ```
33
+
34
+ #### Why would I use it?
35
+
36
+ OCR extracts text from scanned documents and images. The result
37
+ feeds into search indexes, extraction pipelines that pull out dates
38
+ and amounts, or archives so scanned contracts become searchable
39
+ records. PDFs and images both work, and the response is structured
40
+ per page with markdown.
41
+
42
+ #### Notes
43
+
44
+ Only Mistral currently supports OCR through the llm.rb runtime.
45
+ The response exposes pages through
46
+ [`LLM::OCR::Response#pages`](https://r.uby.dev/api-docs/llm.rb/LLM/OCR/Response.html#pages),
47
+ where each page
48
+ has a `markdown` field containing the extracted text.
@@ -0,0 +1,202 @@
1
+
2
+ ## Agents
3
+ ### Introduction
4
+
5
+ #### Overview
6
+
7
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
8
+ is the recommended entry point for most use-cases. It provides a
9
+ class-level DSL for defining reusable, preconfigured assistants
10
+ with defaults for model, tools, schema, and instructions. Under
11
+ the hood it delegates to
12
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html),
13
+ so it has the same runtime surface: message history, streaming,
14
+ serialization, compaction, and concurrency.
15
+
16
+ #### How it works
17
+
18
+ An agent holds a conversation with a model. You send input with
19
+ [`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk),
20
+ the model responds, and if it requests tools the agent
21
+ executes them automatically and feeds the results back. It enables
22
+ [a loop guard by default](https://r.uby.dev/llm/deepdive/advanced/guard)
23
+ that detects repeated tool-call patterns
24
+ and blocks stuck execution. The tool loop can also be bounded with
25
+ [`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
26
+ (see the Tool budget section). Instructions are injected once
27
+ unless a system message is already present.
28
+
29
+ #### Why would I use it?
30
+
31
+ Agents manage the tool loop for you. They guard against infinite
32
+ loops, keep conversation state across turns, and let you define
33
+ reusable configurations at the class level. If you need manual
34
+ control over the tool loop, use
35
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
36
+ directly instead.
37
+
38
+ #### Notes
39
+
40
+ Agents support the same concurrency strategies, compaction,
41
+ cancellation, and serialization as contexts. The trade-off between
42
+ a subclass and a direct instance is only in how the agent is
43
+ organized, not in what it can do. Tool loop execution can be
44
+ configured with `concurrency: :sequential`, `:thread`, `:async`,
45
+ `:fiber`, `:fork`, or `:ractor`.
46
+
47
+ ### Class-based
48
+
49
+ #### Overview
50
+
51
+ A subclass of
52
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
53
+ gives you a reusable agent with its own behavior. You define the model, tools, and other
54
+ attributes at the class level, and each instance picks them up
55
+ as defaults. Attributes can be overridden per-instance, and they
56
+ can be plain values, blocks, or Symbols that resolve to methods.
57
+ The class becomes a self-contained worker that you can instantiate
58
+ and talk to from anywhere.
59
+
60
+ #### How it works
61
+
62
+ A subclass declares its defaults with
63
+ [`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method).
64
+ Each key is a
65
+ class-level accessor: `name`, `description`, `model`, `tools`,
66
+ `skills`, `instructions`, `stream`, `tracer`, `concurrency`,
67
+ `schema`, `confirm`, `path`, `tool_budget`.
68
+ Keyword arguments in the constructor override these defaults.
69
+
70
+ ```ruby
71
+ class Agent < LLM::Agent
72
+ set model: "deepseek-v4-pro",
73
+ description: "system administration agent",
74
+ tools: [Shell]
75
+ end
76
+
77
+ llm = LLM.openai(key: ENV["KEY"])
78
+ agent = Agent.new(llm)
79
+ agent.talk "Run 'date'"
80
+ ```
81
+
82
+ #### Why would I use it?
83
+
84
+ A subclass is useful when multiple parts of an application need to
85
+ call the same agent. The configuration and any helper methods live in one place. Define a
86
+ `research!` method that kicks off the agent's work. The subclass
87
+ becomes a self-contained worker.
88
+
89
+ #### Notes
90
+
91
+ Attributes passed to
92
+ [`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
93
+ can be plain values, blocks, or
94
+ Symbols. A Symbol is evaluated as an instance method on the
95
+ subclass, so `tracer: :set_tracer` calls `set_tracer` on the
96
+ instance. A block like `stream: -> { $stdout }` is evaluated
97
+ when the attribute is first accessed.
98
+
99
+ Set `path:` on a subclass or instance for automatic filesystem
100
+ persistence; the agent restores conversation history from the
101
+ file on startup and saves it back after every turn with no
102
+ manual
103
+ [`LLM::Agent#save`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#save)/
104
+ [`LLM::Agent#restore`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#restore)
105
+ calls. See the
106
+ [database deepdive](../fundamentals/database.md) for details.
107
+
108
+ ### Object-based
109
+
110
+ #### Overview
111
+
112
+ An
113
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
114
+ instance is the simplest way to get started. You pass a provider and any configuration as keyword
115
+ arguments, and the agent runs the tool loop and manages state
116
+ just like a subclass would. This is the right choice when you
117
+ are prototyping, running a one-off task, or when the agent's
118
+ configuration is determined at runtime.
119
+
120
+ #### How it works
121
+
122
+ A direct instance takes the same attributes as keyword arguments
123
+ to
124
+ [`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#initialize-instance_method).
125
+ The first argument is always the provider.
126
+ Everything else is optional. The agent runs the tool loop
127
+ and manages state under the hood through a
128
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html).
129
+
130
+ ```ruby
131
+ llm = LLM.deepseek(key: ENV["KEY"])
132
+ agent = LLM::Agent.new(llm, stream: $stdout)
133
+ agent.talk "Hello world"
134
+ ```
135
+
136
+ #### Why would I use it?
137
+
138
+ A direct instance is the right choice for quick experiments, one-shot
139
+ tasks, or when defining a class would be overkill. It is also
140
+ the right choice when the agent's configuration is determined at
141
+ runtime and a class hierarchy adds unnecessary complexity.
142
+
143
+ #### Notes
144
+
145
+ Direct instances accept all the same options as subclasses. The
146
+ difference is only in how the agent is organized, not in what it
147
+ can do. Under the hood,
148
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
149
+ creates a
150
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
151
+ that manages the message history.
152
+
153
+ ### Tool budget
154
+
155
+ #### Overview
156
+
157
+ [`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
158
+ caps the number of tool calls allowed in a single turn. Once the
159
+ budget is spent, the agent sends an in-band advisory message back
160
+ through the model instead of running more tools. By default no
161
+ budget is set, so the feature is disabled.
162
+
163
+ #### How it works
164
+
165
+ When you want to bound how many tools an agent can call in one
166
+ turn, set the budget with
167
+ [`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method).
168
+ The budget can be a plain number or a block evaluated against the
169
+ agent instance. Once the agent has made the budgeted number of tool
170
+ calls, it stops and returns an in-band advisory message describing
171
+ the spent budget. A model will usually change course afterwards:
172
+
173
+ ```ruby
174
+ class Researcher < LLM::Agent
175
+ set model: "deepseek-v4-pro",
176
+ tools: [FetchNews, FetchStocks],
177
+ tool_budget: 5
178
+ end
179
+
180
+ llm = LLM.deepseek(key: ENV["KEY"])
181
+ agent = Researcher.new(llm)
182
+ agent.talk "Research the market"
183
+ ```
184
+
185
+ #### Why would I use it?
186
+
187
+ A tool budget prevents runaway tool loops. A misbehaving model can
188
+ otherwise keep calling tools, spending tokens on every round trip.
189
+ Capping the budget turns that into a bounded conversation: after
190
+ the cap, the model is told it has run out of tool calls and must
191
+ respond from what it has.
192
+
193
+ #### Notes
194
+
195
+ The budget is disabled by default (`nil`). Set it on a subclass
196
+ with the `tool_budget` DSL, through
197
+ [`LLM::Agent.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#set-class_method)
198
+ with `tool_budget:`, or per-instance with the `tool_budget:` keyword
199
+ argument to
200
+ [`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#initialize-instance_method).
201
+ This replaces the previous `tool_attempts` parameter, which is no
202
+ longer used.