llm.rb 15.2.2 → 15.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +197 -3
  3. data/README.md +180 -64
  4. data/bin/llm.rb +9 -2
  5. data/data/alibaba.json +45 -0
  6. data/data/anthropic.json +67 -0
  7. data/data/bedrock.json +1466 -397
  8. data/data/deepinfra.json +148 -16
  9. data/data/deepseek.json +3 -0
  10. data/data/mistral.json +42 -0
  11. data/data/openai.json +186 -0
  12. data/data/openrouter.json +1538 -378
  13. data/data/xai.json +53 -20
  14. data/data/zai.json +90 -4
  15. data/docs/deepdive/advanced/compaction.md +1 -2
  16. data/docs/deepdive/advanced/context.md +214 -1
  17. data/docs/deepdive/advanced/guard.md +9 -57
  18. data/docs/deepdive/features/builtin_tools.md +14 -16
  19. data/docs/deepdive/features/console.md +5 -0
  20. data/docs/deepdive/features/database.md +85 -10
  21. data/docs/deepdive/fundamentals/agents.md +7 -8
  22. data/docs/deepdive/fundamentals/providers.md +45 -5
  23. data/docs/deepdive/fundamentals/schema.md +73 -0
  24. data/docs/deepdive/fundamentals/tools.md +80 -27
  25. data/docs/deepdive/media/audio.md +8 -19
  26. data/docs/deepdive/media/images.md +8 -10
  27. data/docs/deepdive/media/ocr.md +1 -3
  28. data/docs/deepdive/reference/cost.md +48 -0
  29. data/docs/deepdive/reference/tracer.md +76 -0
  30. data/docs/deepdive.md +1 -1
  31. data/lib/llm/active_record/message.rb +113 -0
  32. data/lib/llm/active_record.rb +1 -0
  33. data/lib/llm/agent.rb +44 -19
  34. data/lib/llm/console/buffer.rb +9 -1
  35. data/lib/llm/console.rb +6 -1
  36. data/lib/llm/context/deserializer.rb +10 -3
  37. data/lib/llm/context.rb +62 -26
  38. data/lib/llm/guard.rb +2 -8
  39. data/lib/llm/message.rb +18 -7
  40. data/lib/llm/provider.rb +74 -16
  41. data/lib/llm/providers/alibaba.rb +15 -0
  42. data/lib/llm/providers/anthropic/error_handler.rb +5 -2
  43. data/lib/llm/providers/anthropic/files.rb +12 -12
  44. data/lib/llm/providers/anthropic/models.rb +2 -2
  45. data/lib/llm/providers/anthropic.rb +5 -3
  46. data/lib/llm/providers/bedrock/error_handler.rb +3 -2
  47. data/lib/llm/providers/bedrock/models.rb +5 -3
  48. data/lib/llm/providers/bedrock.rb +5 -3
  49. data/lib/llm/providers/deepinfra/audio.rb +4 -4
  50. data/lib/llm/providers/deepinfra/images.rb +4 -4
  51. data/lib/llm/providers/google/error_handler.rb +5 -2
  52. data/lib/llm/providers/google/files.rb +10 -10
  53. data/lib/llm/providers/google/images.rb +2 -2
  54. data/lib/llm/providers/google/models.rb +2 -2
  55. data/lib/llm/providers/google.rb +7 -7
  56. data/lib/llm/providers/mistral.rb +3 -1
  57. data/lib/llm/providers/ollama/error_handler.rb +5 -2
  58. data/lib/llm/providers/ollama/models.rb +2 -2
  59. data/lib/llm/providers/ollama.rb +7 -5
  60. data/lib/llm/providers/openai/audio.rb +6 -6
  61. data/lib/llm/providers/openai/error_handler.rb +5 -2
  62. data/lib/llm/providers/openai/files.rb +10 -10
  63. data/lib/llm/providers/openai/images.rb +4 -4
  64. data/lib/llm/providers/openai/models.rb +2 -2
  65. data/lib/llm/providers/openai/moderations.rb +2 -2
  66. data/lib/llm/providers/openai/request_adapter.rb +1 -1
  67. data/lib/llm/providers/openai/responses.rb +8 -8
  68. data/lib/llm/providers/openai/vector_stores.rb +22 -22
  69. data/lib/llm/providers/openai.rb +10 -8
  70. data/lib/llm/providers/xai/images.rb +4 -4
  71. data/lib/llm/schema.rb +24 -0
  72. data/lib/llm/tracer/telemetry.rb +4 -4
  73. data/lib/llm/tracer.rb +11 -3
  74. data/lib/llm/transport/execution.rb +8 -4
  75. data/lib/llm/utils.rb +13 -0
  76. data/lib/llm/version.rb +1 -1
  77. data/llm.gemspec +2 -2
  78. metadata +5 -5
  79. data/lib/llm/guard/loop.rb +0 -89
@@ -4,10 +4,11 @@
4
4
 
5
5
  #### Overview
6
6
 
7
- Persistence lets an agent outlive a single session. The conversation
8
- history, model name, and compaction status are serialized as JSON
9
- that can be stored in a file, a database column, or transmitted
10
- over a network. Four storage options are available:
7
+ Persistence lets an agent outlive a single session. The
8
+ conversation history, the context id, the model name, the compaction
9
+ status, and a snapshot of token usage are serialized as JSON that can
10
+ be stored in a file, a database column, or transmitted over a network.
11
+ Four storage options are available:
11
12
 
12
13
  - **Automatic filesystem persistence**: set `path:` on an agent
13
14
  for transparent auto-save after every turn (recommended for
@@ -15,9 +16,13 @@ over a network. Four storage options are available:
15
16
  - **Filesystem**: save and restore from a JSON file on disk
16
17
  - **ActiveRecord**: persist state in a database column using
17
18
  [`acts_as_agent`](https://r.uby.dev/api-docs/llm.rb/LLM/ActiveRecord.html#acts_as_agent-instance_method)
19
+ for an agent, or
20
+ [`acts_as_llm`](https://r.uby.dev/api-docs/llm.rb/LLM/ActiveRecord.html#acts_as_llm-instance_method)
21
+ for a context
18
22
  - **Sequel**: persist state in a database column using `plugin :agent`
23
+ for an agent, or `plugin :llm` for a context
19
24
 
20
- All three use the same serialization mechanism under the hood.
25
+ All four use the same serialization mechanism under the hood.
21
26
 
22
27
  #### How it works
23
28
 
@@ -161,8 +166,9 @@ and
161
166
  for serialization
162
167
  and
163
168
  [`LLM::Context#restore`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#restore)
164
- for deserialization. The serialized state includes
165
- the message history, model name, and compaction status. Save and
169
+ for deserialization. The serialized state includes the context id,
170
+ the message history, the model name, the compaction status, and a
171
+ snapshot of token usage. Save and
166
172
  restore work with file paths or in-memory strings. You can also
167
173
  serialize to a JSON string for database storage or network
168
174
  transmission:
@@ -230,9 +236,10 @@ When you want to add agent persistence to an ActiveRecord model,
230
236
  call
231
237
  [`LLM::ActiveRecord#acts_as_agent`](https://r.uby.dev/api-docs/llm.rb/LLM/ActiveRecord.html#acts_as_agent-instance_method)
232
238
  in the model class. The `data` column stores
233
- the full agent state (conversation history, model name, compaction
234
- status) as JSON. On first call, a fresh agent is created and the
235
- conversation starts from scratch.
239
+ the full agent state (the context id, conversation history, model
240
+ name, compaction status, and a token usage snapshot) as JSON. On
241
+ first call, a fresh agent is created and the conversation starts
242
+ from scratch.
236
243
 
237
244
  On subsequent calls, the stored state is restored and the
238
245
  conversation continues. Every
@@ -339,6 +346,67 @@ also work for backwards compatibility, but the block style
339
346
  (`agent.tracer ...`, `agent.stream ...`) is preferred for all
340
347
  agent-level configuration.
341
348
 
349
+ An agent built by `acts_as_agent` is bound to the record it was
350
+ loaded from, and
351
+ [`LLM::Agent#record`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#record)
352
+ returns it. To persist a context instead of an agent, use
353
+ `acts_as_llm`: the model gains `#llm` (the provider) and `#ctx` (the
354
+ context), and persists the same state without the automatic tool
355
+ loop.
356
+
357
+ ### SQL view
358
+
359
+ #### Overview
360
+
361
+ [`LLM::ActiveRecord::Message`](https://r.uby.dev/api-docs/llm.rb/LLM/ActiveRecord/Message.html)
362
+ is a virtual ActiveRecord model that never materializes as a table. It
363
+ exposes the messages stored inside an agent's `jsonb` column as a SQL
364
+ view, so a conversation can be filtered, ordered, and counted in the
365
+ database instead of in memory.
366
+
367
+ #### How it works
368
+
369
+ Call
370
+ [`LLM::ActiveRecord::Message.for`](https://r.uby.dev/api-docs/llm.rb/LLM/ActiveRecord/Message.html#for-class_method)
371
+ with an agent, and it returns an
372
+ [`ActiveRecord::Relation`](https://api.rubyonrails.org/classes/ActiveRecord/Relation.html)
373
+ scoped to that agent, with one row per message. The relation chains
374
+ like any other:
375
+
376
+ ```ruby
377
+ agent = Agent.find_by(id: 1)
378
+
379
+ messages = LLM::ActiveRecord::Message.for(agent:)
380
+ messages.where(role: "assistant").order(position: :desc).limit(10)
381
+ messages.count
382
+ ```
383
+
384
+ Each row carries a message flattened into columns: `id`, `role`,
385
+ `content`, `tools`, and `position` (its place in the conversation),
386
+ with the whole message kept as `data`.
387
+ [`LLM::ActiveRecord::Message#unwrap!`](https://r.uby.dev/api-docs/llm.rb/LLM/ActiveRecord/Message.html#unwrap!-instance_method)
388
+ rebuilds the message as the runtime would hand it back, so fields the
389
+ view does not name, such as usage and reasoning, survive the round
390
+ trip. `#tool_call?` and `#tool_return?` delegate to it.
391
+
392
+ #### Why would I use it?
393
+
394
+ Reading a conversation through the view keeps the work in the
395
+ database. "How many assistant messages has this agent produced" and
396
+ "show the last ten messages" become ordinary ActiveRecord queries
397
+ instead of loading and filtering the whole conversation in Ruby.
398
+
399
+ #### Notes
400
+
401
+ The view expects the agent and this class to share a connection, which
402
+ holds when both live on the same database. It requires
403
+ `format: :jsonb`, since a `:string` column stores the state as text and
404
+ cannot be expanded into rows. The queries the view runs are index
405
+ scans, because they expand one agent found by primary key, but a query
406
+ you write across every agent is not covered by default; index the state
407
+ column for those, for example with a GIN index on
408
+ `data jsonb_path_ops`.
409
+
342
410
  ### Sequel
343
411
 
344
412
  #### Overview
@@ -447,3 +515,10 @@ also work for backwards compatibility, but configuring tracer,
447
515
  stream, and other agent options in the block is the preferred
448
516
  approach.
449
517
 
518
+ To persist a context instead of an agent, use `plugin :llm`. The
519
+ model gains `#llm` (the provider) and `#ctx` (the context), and
520
+ persists the same state without the automatic tool loop. An agent
521
+ built by `plugin :agent` is bound to its record, which
522
+ [`LLM::Agent#record`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#record)
523
+ returns.
524
+
@@ -18,20 +18,19 @@ serialization, compaction, and concurrency.
18
18
  An agent holds a conversation with a model. You send input with
19
19
  [`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk),
20
20
  the model responds, and if it requests tools the agent
21
- executes them automatically and feeds the results back. It enables
22
- [a loop guard by default](https://r.uby.dev/llm/deepdive/advanced/guard)
23
- that detects repeated tool-call patterns
24
- and blocks stuck execution. The tool loop can also be bounded with
21
+ executes them automatically and feeds the results back. The tool
22
+ loop can be bounded with
25
23
  [`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
26
24
  (see the Tool budget section). Instructions are injected once
27
25
  unless a system message is already present.
28
26
 
29
27
  #### Why would I use it?
30
28
 
31
- Agents manage the tool loop for you. They guard against infinite
32
- loops, keep conversation state across turns, and let you define
33
- reusable configurations at the class level. If you need manual
34
- control over the tool loop, use
29
+ Agents manage the tool loop for you. They keep conversation state
30
+ across turns and let you define reusable configurations at the
31
+ class level. The loop runs until the model stops requesting tools
32
+ unless you bound it with `tool_budget` or a guard. If you need
33
+ manual control over the tool loop, use
35
34
  [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
36
35
  directly instead.
37
36
 
@@ -5,11 +5,13 @@
5
5
  #### Overview
6
6
 
7
7
  llm.rb talks to 14+ providers through one API. OpenAI-compatible
8
- providers (Anthropic, DeepSeek, DeepInfra, xAI, Z.ai, Moonshot,
9
- Alibaba, Ollama, and llama.cpp) share the same OpenAI code path, so
10
- switching models rarely means switching code. Each provider is
11
- constructed with a class-level factory method on `LLM`, and the
12
- result is passed to an `LLM::Context` or `LLM::Agent`.
8
+ providers (DeepSeek, DeepInfra, xAI, Z.ai, Moonshot, Alibaba,
9
+ Mistral, OpenRouter, and llama.cpp) share the same OpenAI code path,
10
+ while Anthropic, Google, Ollama, and Bedrock speak their own APIs
11
+ behind the same interface. Switching models therefore rarely means
12
+ switching code. Each provider is constructed with a class-level
13
+ factory method on `LLM`, and the result is passed to an
14
+ `LLM::Context` or `LLM::Agent`.
13
15
 
14
16
  #### How it works
15
17
 
@@ -90,6 +92,44 @@ missing model or registry raises `LLM::NoSuchModelError` or
90
92
  `LLM::NoSuchRegistryError`, which the runtime rescues to default
91
93
  gracefully (for example, an unknown context window reads as `nil`).
92
94
 
95
+ ### Request headers
96
+
97
+ #### Overview
98
+
99
+ Every request a provider sends can carry extra HTTP headers, either
100
+ for the lifetime of the provider or for a single call. Some headers
101
+ vary per request, so a provider needs a way to set one without
102
+ affecting the requests that come after it.
103
+
104
+ #### How it works
105
+
106
+ [`LLM::Provider#with`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#with-instance_method)
107
+ adds headers. Without a block the headers merge into the provider's
108
+ defaults and apply to every later request:
109
+
110
+ ```ruby
111
+ llm = LLM.openai(key: ENV["KEY"])
112
+ llm.with("OpenAI-Organization" => ENV["ORG"])
113
+ ```
114
+
115
+ With a block the headers apply only to the current fiber and are
116
+ restored when the block returns, so a header can be scoped to one
117
+ call:
118
+
119
+ ```ruby
120
+ llm.with("x-session-id" => "abc123") do
121
+ ctx.talk "Hello"
122
+ end
123
+ ```
124
+
125
+ #### Notes
126
+
127
+ The runtime uses the block form where a header varies per request.
128
+ An OpenRouter context sends its own id as `x-session-id`, so
129
+ consecutive requests share a session and OpenRouter can route them to
130
+ the same cached model, without pinning that header on the provider
131
+ for good.
132
+
93
133
  ### Moonshot
94
134
 
95
135
  #### Overview
@@ -59,3 +59,76 @@ They are also used internally by
59
59
  [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
60
60
  for parameter definitions, so you already benefit from them
61
61
  when you declare tool parameters.
62
+
63
+ ### Types
64
+
65
+ #### Overview
66
+
67
+ A schema is built from a small set of types: the JSON primitives, an
68
+ array of a type, an enum of allowed values, a nested schema, and the
69
+ combinators `any_of`, `all_of`, and `one_of`. Every type takes a
70
+ description, and any type can be marked required or given a default.
71
+
72
+ #### How it works
73
+
74
+ A property is declared with its type and description. `String`,
75
+ `Integer`, `Number`, and `Boolean` are the primitives, `Array[Type]`
76
+ wraps one, and `Enum[...]` constrains a value to a fixed set. A nested
77
+ schema is just another `LLM::Schema` subclass used as the type:
78
+
79
+ ```ruby
80
+ class Address < LLM::Schema
81
+ property :street, String, "Street address"
82
+ required %i[street]
83
+ end
84
+
85
+ class Person < LLM::Schema
86
+ property :name, String, "Person's name"
87
+ property :age, Integer, "Person's age"
88
+ property :hobbies, Array[String], "Person's hobbies"
89
+ property :address, Address, "Person's address"
90
+ required %i[name age hobbies address]
91
+ end
92
+ ```
93
+
94
+ `Array[String, Integer]` declares an array whose items may be either
95
+ type. Options on `property` set a leaf directly, so
96
+ `property :age, Integer, "Person's age", required: true` marks one
97
+ property required, and
98
+ [`LLM::Schema.defaults`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html#defaults-class_method)
99
+ sets defaults for several at once:
100
+
101
+ ```ruby
102
+ class Search < LLM::Schema
103
+ property :query, String, "The search query"
104
+ property :limit, Integer, "The number of results"
105
+ required %i[query]
106
+ defaults limit: 10
107
+ end
108
+ ```
109
+
110
+ The same schema can be built without a class, using the value methods:
111
+
112
+ ```ruby
113
+ schema = LLM::Schema.new
114
+ schema.object(
115
+ name: schema.string.required,
116
+ age: schema.integer.required,
117
+ colors: schema.array(schema.string.enum("red", "green")).required
118
+ )
119
+ ```
120
+
121
+ #### Why would I use it?
122
+
123
+ Types tell the model what shape to produce. A plain `String` is the
124
+ loosest, an `Enum` the tightest, and a nested schema lets a structured
125
+ response contain a structured value. Because the same machinery backs
126
+ tool parameters, anything you learn here applies to tools too.
127
+
128
+ #### Notes
129
+
130
+ `any_of`, `all_of`, and `one_of` combine types, and
131
+ [`LLM::Schema.to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html#to_s-class_method)
132
+ renders the schema as a prompt-friendly string. `required` and
133
+ `defaults` refer to properties that already exist, so declare the
134
+ property first and mark it afterwards.
@@ -51,17 +51,17 @@ class Exec < LLM::Tool
51
51
 
52
52
  name "exec"
53
53
  description "run a command without a shell"
54
- parameter :name, String, "the command's name"
55
- parameter :arguments, Array[String], "command args"
56
- required %i[name]
57
- defaults arguments: [], timeout: 60, max_bytes: :max_bytes
54
+ parameter :arguments, Array[String], "a command and its arguments"
55
+ required %i[arguments]
56
+ defaults timeout: 60, max_bytes: :max_bytes
58
57
 
59
58
  def self.max_bytes(bytes = nil)
60
59
  bytes ? (@max_bytes = bytes) : (@max_bytes || 75_000)
61
60
  end
62
61
 
63
- def call(name:, arguments: [], timeout: 60, max_bytes: self.class.max_bytes)
64
- command = spawn(name:, arguments:, max_bytes:)
62
+ def call(arguments: [], timeout: 60, max_bytes: self.class.max_bytes)
63
+ name = arguments[0]
64
+ command = spawn(name:, arguments: arguments[1..], max_bytes:)
65
65
  wait(command:, timeout:)
66
66
  {ok: command.success?, stdout: command.stdout, stderr: command.stderr}
67
67
  rescue LLM::Interrupt
@@ -172,33 +172,37 @@ matter what.
172
172
  If
173
173
  [`LLM::Tool#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html#call)
174
174
  raises, the runtime returns `{error: true, type: "RuntimeError",
175
- message: "boom"}` to the model. You can also rescue inside
176
- [`LLM::Tool#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html#call)
177
- and return your own error shape that gives the model more context.
175
+ message: "boom"}` to the model. A tool can also check for a known
176
+ failure and return its own error shape, which gives the model
177
+ more context than a generic error. The `exec` tool, for example,
178
+ reports a missing command through `command.not_found?` rather than
179
+ letting the spawn fail.
178
180
 
179
181
  ```ruby
180
182
  require "llm/tools/utils"
181
183
 
182
- class Exec < LLM::Tool
184
+ class SafeExec < LLM::Tool
183
185
  include Utils
184
186
 
185
- name "exec"
186
- description "run a command without a shell"
187
- parameter :name, String, "the command name"
188
- parameter :arguments, Array[String], "command args"
189
- required %i[name]
190
- defaults arguments: [], timeout: 60, max_bytes: :max_bytes
187
+ name "safe-exec"
188
+ description "run a command and report a missing one"
189
+ parameter :arguments, Array[String], "a command and its arguments"
190
+ required %i[arguments]
191
+ defaults timeout: 60, max_bytes: :max_bytes
191
192
 
192
193
  def self.max_bytes(bytes = nil)
193
194
  bytes ? (@max_bytes = bytes) : (@max_bytes || 75_000)
194
195
  end
195
196
 
196
- def call(name:, arguments: [], timeout: 60, max_bytes: self.class.max_bytes)
197
- command = spawn(name:, arguments:, max_bytes:)
197
+ def call(arguments: [], timeout: 60, max_bytes: self.class.max_bytes)
198
+ name = arguments[0]
199
+ command = spawn(name:, arguments: arguments[1..], max_bytes:)
198
200
  wait(command:, timeout:)
199
- {ok: command.success?, stdout: command.stdout, stderr: command.stderr}
200
- rescue Errno::ENOENT
201
- {ok: false, error: "command not found: #{name}"}
201
+ if command.not_found?
202
+ {ok: false, error: "command '#{name}' was not found on this system"}
203
+ else
204
+ {ok: command.success?, stdout: command.stdout, stderr: command.stderr}
205
+ end
202
206
  end
203
207
  end
204
208
  ```
@@ -207,8 +211,8 @@ end
207
211
 
208
212
  Custom error handling gives the model domain-specific detail that
209
213
  helps it recover. Instead of a generic "RuntimeError: boom", the
210
- model sees `{ok: false, error: "command not found: ls"}` and knows
211
- to correct the command name and try again.
214
+ model sees `{ok: false, error: "command 'ls' was not found on this
215
+ system"}` and knows to correct the command name and try again.
212
216
 
213
217
  #### Notes
214
218
 
@@ -244,8 +248,7 @@ class Exec < LLM::Tool
244
248
  set name: "exec",
245
249
  description: "run a command without a shell",
246
250
  parameters: [
247
- [:name, String, "the command's name", {required: true}],
248
- [:arguments, Array[String], "command args", {default: []}],
251
+ [:arguments, Array[String], "a command and its arguments", {required: true}],
249
252
  [:timeout, Integer, "timeout in seconds", {default: 60}],
250
253
  [:max_bytes, Integer, "max bytes to emit", {default: 75_000}]
251
254
  ]
@@ -254,8 +257,9 @@ class Exec < LLM::Tool
254
257
  bytes ? (@max_bytes = bytes) : (@max_bytes || 75_000)
255
258
  end
256
259
 
257
- def call(name:, arguments: [], timeout: 60, max_bytes: self.class.max_bytes)
258
- command = spawn(name:, arguments:, max_bytes:)
260
+ def call(arguments: [], timeout: 60, max_bytes: self.class.max_bytes)
261
+ name = arguments[0]
262
+ command = spawn(name:, arguments: arguments[1..], max_bytes:)
259
263
  wait(command:, timeout:)
260
264
  {ok: command.success?, stdout: command.stdout, stderr: command.stderr}
261
265
  end
@@ -326,6 +330,55 @@ loop, confirmation, and error handling behave identically. A single
326
330
  failing tool returns a structured error to the model, which can
327
331
  decide to retry or continue with the results it has.
328
332
 
333
+ ### Platform-native tools
334
+
335
+ #### Overview
336
+
337
+ Some capabilities live inside the provider rather than on your
338
+ machine. Web search, code execution, file search, and computer use are
339
+ examples: the provider runs them on its own infrastructure, and the
340
+ model can call them directly. The runtime represents these with
341
+ [`LLM::ServerTool`](https://r.uby.dev/api-docs/llm.rb/LLM/ServerTool.html).
342
+ A server tool is not an
343
+ [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html)
344
+ subclass, so it never appears in `LLM::Tool.subclasses`.
345
+
346
+ #### How it works
347
+
348
+ Build a server tool from a provider with
349
+ [`LLM::Provider#server_tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#server_tool-instance_method),
350
+ or read a ready-made one from the provider's catalog with
351
+ [`LLM::Provider#server_tools`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#server_tools-instance_method).
352
+ Pass server tools in the same `tools:` list as local tools:
353
+
354
+ ```ruby
355
+ require "llm"
356
+
357
+ llm = LLM.google(key: ENV["KEY"])
358
+ ctx = LLM::Context.new(llm, tools: [llm.server_tool(:google_search)])
359
+ ctx.talk "Summarize today's news"
360
+ ```
361
+
362
+ Each provider defines its own catalog. OpenAI offers `web_search`,
363
+ `file_search`, `image_generation`, `code_interpreter`, and
364
+ `computer_use`; Google offers `google_search`, `code_execution`, and
365
+ `url_context`; Anthropic offers `bash`, `web_search`, and
366
+ `text_editor`.
367
+
368
+ #### Why would I use it?
369
+
370
+ A server tool saves you from building and hosting the capability
371
+ yourself. Search, code execution, and file search run on the
372
+ provider's side, so the model can call them without a local tool
373
+ loop, a service of your own, or extra credentials.
374
+
375
+ #### Notes
376
+
377
+ OpenAI, Google, and Anthropic also expose a `web_search(query:)`
378
+ method that performs a search in one call and returns the results. A
379
+ server tool accepts whatever options the provider documents, for
380
+ example `llm.server_tool(:web_search, max_uses: 5)` on Anthropic.
381
+
329
382
  ### Built-in tools
330
383
 
331
384
  #### Overview
@@ -14,10 +14,7 @@ copy the result somewhere useful.
14
14
  #### How it works
15
15
 
16
16
  When you want to convert text to speech, call
17
- [`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech).
18
- The provider returns an audio clip as a
19
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
20
- object. The generated audio can be copied
17
+ `audio.create_speech`. The response holds the audio clip, which can be copied
21
18
  to a file or streamed directly. Each provider supports different
22
19
  output formats and voice options:
23
20
 
@@ -38,20 +35,14 @@ single method call.
38
35
  #### Notes
39
36
 
40
37
  OpenAI has full audio support. Google and DeepInfra have partial
41
- support. The
42
- [`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech)
43
- method returns a
44
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
45
- object.
38
+ support.
46
39
 
47
40
  ### Transcription
48
41
 
49
42
  #### Overview
50
43
 
51
- Transcription turns an audio file into text. Pass a file path
52
- to
53
- [`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription)
54
- and get back the spoken content as
44
+ Transcription turns an audio file into text. Pass a file path to
45
+ `audio.create_transcription` and get back the spoken content as
55
46
  a string. OpenAI, Google, and DeepInfra support it. Transcribe
56
47
  meeting notes, voice memos, or podcast episodes for search and
57
48
  processing. The response is plain text you can feed into any
@@ -60,8 +51,8 @@ downstream pipeline.
60
51
  #### How it works
61
52
 
62
53
  When you want to transcribe audio into text, call
63
- [`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription).
64
- The provider processes the audio and returns the transcribed text.
54
+ `audio.create_transcription`. The provider processes the audio and
55
+ returns the transcribed text.
65
56
  The file can be a local path or a URL depending on provider support:
66
57
 
67
58
  ```ruby
@@ -88,8 +79,7 @@ support.
88
79
 
89
80
  Translation transcribes audio and translates it into English in
90
81
  one step. Pass a file to
91
- [`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation)
92
- and get back the
82
+ `audio.create_translation` and get back the
93
83
  translated text. OpenAI and Google support it. Translate
94
84
  multilingual podcasts, interviews, or any audio where you need
95
85
  the content in English without running a separate translation
@@ -98,8 +88,7 @@ pipeline.
98
88
  #### How it works
99
89
 
100
90
  When you want to translate spoken audio into English, call
101
- [`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation).
102
- The provider transcribes the spoken language and translates the
91
+ `audio.create_translation`. The provider transcribes the spoken language and translates the
103
92
  result into English in a single operation. The returned text is the
104
93
  English translation:
105
94
 
@@ -14,11 +14,10 @@ requires no code changes.
14
14
  #### How it works
15
15
 
16
16
  The
17
- [`LLM::Images#create`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#create)
18
- method sends a prompt to the provider and
19
- returns the result as a
20
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
21
- object. The same API works
17
+ `images.create`
18
+ method sends a prompt to the provider and returns an
19
+ [`LLM::Response`](https://r.uby.dev/api-docs/llm.rb/LLM/Response.html)
20
+ whose `images` array holds the generated images. The same API works
22
21
  across providers: swap
23
22
  [`LLM.openai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#openai-class_method) for
24
23
  [`LLM.xai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#xai-class_method) and the rest
@@ -61,11 +60,10 @@ text-to-image prompts, giving you iterative vector editing.
61
60
  #### How it works
62
61
 
63
62
  The
64
- [`LLM::Images#edit`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#edit)
65
- method takes a prompt and an image path. It
66
- returns a modified image as a
67
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
68
- object. Copy the result
63
+ `images.edit`
64
+ method takes a prompt and an image path. It returns an
65
+ [`LLM::Response`](https://r.uby.dev/api-docs/llm.rb/LLM/Response.html)
66
+ whose `images` array holds the modified image. Copy the result
69
67
  to a file the same way you would with generated images.
70
68
 
71
69
  ```ruby
@@ -42,7 +42,5 @@ per page with markdown.
42
42
  #### Notes
43
43
 
44
44
  Only Mistral currently supports OCR through the llm.rb runtime.
45
- The response exposes pages through
46
- [`LLM::OCR::Response#pages`](https://r.uby.dev/api-docs/llm.rb/LLM/OCR/Response.html#pages),
47
- where each page
45
+ The response exposes pages through its `pages` reader, where each page
48
46
  has a `markdown` field containing the extracted text.
@@ -107,3 +107,51 @@ The console renders context usage as a proportion, not a cost.
107
107
  returns a `Rational` of the tokens used over the context window
108
108
  (for example `Rational(100, 10_000)`), or `nil` when the window is
109
109
  unknown or the conversation is too short.
110
+
111
+ ### Token usage
112
+
113
+ #### Overview
114
+
115
+ [`LLM::Usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Usage.html)
116
+ holds the token counts a provider reports for a request: input,
117
+ output, reasoning, cache read, cache write, and the audio and image
118
+ tokens where a model supports them. Cost is computed from it, and so
119
+ is the console's context meter.
120
+
121
+ #### How it works
122
+
123
+ Every context and agent exposes four readers, each answering a
124
+ different question:
125
+
126
+ ```ruby
127
+ ctx.token_usage # => LLM::Usage, summed over the whole conversation
128
+ ctx.context_used # => tokens in the most recent assistant message
129
+ ctx.context_usage # => Rational fraction of the context window in use
130
+ ctx.context_window # => the model's window, or nil when unknown
131
+ ```
132
+
133
+ [`LLM::Context#token_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#token_usage-instance_method)
134
+ accumulates across the conversation and returns `LLM::Usage.zero`
135
+ before any provider usage has been recorded.
136
+ [`LLM::Context#context_used`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_used-instance_method)
137
+ is the live size of a single turn, which is what the context window
138
+ is really being spent on.
139
+ [`LLM::Context#context_window`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_window-instance_method)
140
+ reads the limit from the model registry, and
141
+ [`LLM::Context#context_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_usage-instance_method)
142
+ divides one by the other.
143
+
144
+ #### Why would I use it?
145
+
146
+ `token_usage` answers "what has this conversation cost so far", while
147
+ `context_used` and `context_usage` answer "how much room is left".
148
+ Showing both lets a user see spend and headroom without either number
149
+ being mistaken for the other.
150
+
151
+ #### Notes
152
+
153
+ `LLM::Context#usage` is an alias of `token_usage`, kept for
154
+ compatibility. `context_used` and `context_usage` return `nil` when
155
+ the model is unknown to the registry, or before the conversation has
156
+ an assistant message to measure. An agent delegates all four readers
157
+ to the context it wraps.