llm.rb 13.1.0 → 14.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (90) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +320 -0
  3. data/README.md +340 -31
  4. data/bin/llm.rb +36 -12
  5. data/data/anthropic.json +206 -263
  6. data/data/bedrock.json +2138 -1860
  7. data/data/deepinfra.json +1003 -624
  8. data/data/deepseek.json +38 -34
  9. data/data/google.json +1079 -371
  10. data/data/mistral.json +448 -368
  11. data/data/moonshot.json +384 -0
  12. data/data/openai.json +974 -1343
  13. data/data/xai.json +154 -126
  14. data/data/zai.json +191 -191
  15. data/lib/llm/agent.rb +47 -14
  16. data/lib/llm/context.rb +71 -88
  17. data/lib/llm/cost.rb +23 -17
  18. data/lib/llm/error.rb +0 -8
  19. data/lib/llm/function/async/task.rb +2 -0
  20. data/lib/llm/function/fiber/task.rb +2 -0
  21. data/lib/llm/function/fork/task.rb +2 -0
  22. data/lib/llm/function/ractor/task.rb +2 -0
  23. data/lib/llm/function/sequential/group.rb +4 -1
  24. data/lib/llm/function/sequential/task.rb +1 -1
  25. data/lib/llm/function/task.rb +4 -0
  26. data/lib/llm/function/thread/task.rb +2 -0
  27. data/lib/llm/function.rb +32 -4
  28. data/lib/llm/guard/loop.rb +89 -0
  29. data/lib/llm/guard/null.rb +19 -0
  30. data/lib/llm/guard.rb +61 -0
  31. data/lib/llm/provider.rb +36 -0
  32. data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
  33. data/lib/llm/providers/anthropic.rb +1 -8
  34. data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
  35. data/lib/llm/providers/bedrock.rb +1 -8
  36. data/lib/llm/providers/google/stream_parser.rb +1 -0
  37. data/lib/llm/providers/google.rb +1 -8
  38. data/lib/llm/providers/moonshot.rb +76 -0
  39. data/lib/llm/providers/ollama.rb +1 -8
  40. data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
  41. data/lib/llm/providers/openai/responses.rb +6 -8
  42. data/lib/llm/providers/openai/stream_parser.rb +1 -0
  43. data/lib/llm/providers/openai.rb +3 -10
  44. data/lib/llm/repl/bar.rb +4 -3
  45. data/lib/llm/repl/buffer.rb +42 -15
  46. data/lib/llm/repl/color.rb +78 -0
  47. data/lib/llm/repl/input/char.rb +46 -0
  48. data/lib/llm/repl/input/row.rb +39 -0
  49. data/lib/llm/repl/input.rb +251 -66
  50. data/lib/llm/repl/markdown/table.rb +6 -2
  51. data/lib/llm/repl/markdown.rb +31 -5
  52. data/lib/llm/repl/status.rb +38 -3
  53. data/lib/llm/repl/stream.rb +16 -4
  54. data/lib/llm/repl/walker.rb +3 -2
  55. data/lib/llm/repl/window.rb +25 -5
  56. data/lib/llm/repl.rb +29 -13
  57. data/lib/llm/stream.rb +8 -7
  58. data/lib/llm/tool.rb +29 -0
  59. data/lib/llm/transformer/null.rb +21 -0
  60. data/lib/llm/transformer.rb +55 -0
  61. data/lib/llm/version.rb +1 -1
  62. data/lib/llm.rb +12 -2
  63. data/llm.gemspec +1 -0
  64. data/resources/deepdive/advanced/cancellation.md +74 -0
  65. data/resources/deepdive/advanced/compaction.md +83 -0
  66. data/resources/deepdive/advanced/context.md +267 -0
  67. data/resources/deepdive/advanced/guard.md +371 -0
  68. data/resources/deepdive/advanced/tracer.md +180 -0
  69. data/resources/deepdive/advanced/transformer.md +67 -0
  70. data/resources/deepdive/advanced/transports.md +45 -0
  71. data/resources/deepdive/everything_else/audio.md +122 -0
  72. data/resources/deepdive/everything_else/cost.md +99 -0
  73. data/resources/deepdive/everything_else/images.md +89 -0
  74. data/resources/deepdive/everything_else/object.md +108 -0
  75. data/resources/deepdive/everything_else/ocr.md +48 -0
  76. data/resources/deepdive/fundamentals/agents.md +202 -0
  77. data/resources/deepdive/fundamentals/builtin_tools.md +191 -0
  78. data/resources/deepdive/fundamentals/concurrency.md +104 -0
  79. data/resources/deepdive/fundamentals/database.md +449 -0
  80. data/resources/deepdive/fundamentals/embeddings.md +157 -0
  81. data/resources/deepdive/fundamentals/repl.md +87 -0
  82. data/resources/deepdive/fundamentals/schema.md +61 -0
  83. data/resources/deepdive/fundamentals/skills.md +106 -0
  84. data/resources/deepdive/fundamentals/stream.md +110 -0
  85. data/resources/deepdive/fundamentals/tools.md +265 -0
  86. data/resources/deepdive/protocols/a2a.md +106 -0
  87. data/resources/deepdive/protocols/mcp.md +111 -0
  88. data/resources/deepdive.md +7 -1
  89. metadata +36 -3
  90. data/lib/llm/loop_guard.rb +0 -107
@@ -0,0 +1,83 @@
1
+
2
+ ## Compaction
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ Long-running conversations consume tokens. Without intervention,
9
+ every turn pushes toward the model's context window limit. Compaction
10
+ drops old messages to keep the conversation alive. The runtime runs
11
+ a compactor automatically before each
12
+ [`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk)
13
+ call, trimming the oldest messages when the conversation exceeds a configured size.
14
+ This keeps the context window healthy without manual intervention.
15
+
16
+ #### How it works
17
+
18
+ Compactors run automatically before each
19
+ [`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk)
20
+ call.
21
+ [`LLM::Compactor::Truncate`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Truncate.html)
22
+ strategy drops the oldest messages, keeping only the N most recent.
23
+ It preserves tool call/return pairs so the conversation never
24
+ contains an orphaned result.
25
+
26
+ The `keep:` parameter accepts an integer count or a percentage
27
+ string like `"80%"`. The default compactor is
28
+ [`LLM::Compactor::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Compactor/Null.html),
29
+ which does nothing. A compactor can also be used standalone:
30
+
31
+ ```ruby
32
+ ctx = LLM::Context.new(
33
+ llm,
34
+ compactor: LLM::Compactor::Truncate,
35
+ compactor_options: {keep: 64}
36
+ )
37
+
38
+ agent = LLM::Agent.new(
39
+ llm,
40
+ compactor: LLM::Compactor::Truncate,
41
+ compactor_options: {keep: 128}
42
+ )
43
+
44
+ compactor = LLM::Compactor::Truncate.new(agent)
45
+ compactor.call(keep: 200)
46
+ ```
47
+
48
+ ##### The `/compact` command
49
+
50
+ The REPL provides a `/compact` command that accepts a count or
51
+ percentage:
52
+
53
+ ```ruby
54
+ # /compact # keep last 128 messages
55
+ # /compact 50 # keep last 50 messages
56
+ # /compact 75% # keep approximately 75% of messages
57
+ ```
58
+
59
+ #### Why would I use it?
60
+
61
+ Without compaction, the provider eventually rejects requests because
62
+ the context window is full. Compaction keeps the conversation alive
63
+ by discarding old messages before they cause a problem.
64
+
65
+ Long-running agents that span hundreds of turns need this. A bug
66
+ investigation that bounces back and forth between diagnosis and
67
+ fix cannot fit every exchange in memory. Compaction trims the
68
+ unimportant parts and keeps the conversation alive.
69
+
70
+ #### Notes
71
+
72
+ The Truncate strategy is fast and has no dependencies. It operates
73
+ entirely in memory with a single pass over the message list. The
74
+ trade-off is that dropped messages are gone, so information may be
75
+ lost. Set `keep` to a higher number to retain more context.
76
+
77
+ Both strategies emit
78
+ [`LLM::Stream#on_compaction`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction)
79
+ and
80
+ [`LLM::Stream#on_compaction_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_compaction_finish)
81
+ stream callbacks so the UI can show progress. The context's
82
+ [`LLM::Context#compacted?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#compacted?)
83
+ flag is `true` between compaction and the next model response.
@@ -0,0 +1,267 @@
1
+
2
+ ## Context
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
9
+ is the runtime that powers every agent. When you call
10
+ [`LLM::Agent#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#talk),
11
+ the agent delegates to its internal context. The context manages
12
+ the message history, sends requests to the provider, tracks pending
13
+ tool calls, and feeds results back to the model. Everything an agent
14
+ does, a context does too, but without the automatic tool loop.
15
+
16
+ Using a context directly gives you finer control over each step
17
+ of the conversation. You decide when to send messages, when to
18
+ execute tools, and when to stop. This is useful for custom
19
+ confirmation flows, mixed concurrency strategies per tool, or
20
+ any workflow where the agent's automatic loop gets in the way.
21
+
22
+ #### How it works
23
+
24
+ A context wraps a provider and maintains the conversation state.
25
+ Call
26
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
27
+ to send input to the model, check
28
+ [`LLM::Context#pending_functions?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions?)
29
+ to see if tools were requested, and use
30
+ [`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
31
+ to execute them. Each call to
32
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
33
+ appends
34
+ to the conversation and returns the model's response. The context
35
+ serializes its state with
36
+ [`LLM::Context#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#to_h)
37
+ and
38
+ [`LLM::Context#to_json`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#to_json),
39
+ and restores it
40
+ with
41
+ [`LLM::Context#restore`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#restore).
42
+ This is how the ORM integrations and filesystem
43
+ persistence work under the hood:
44
+
45
+ ```ruby
46
+ require "llm"
47
+
48
+ llm = LLM.deepseek(key: ENV["KEY"])
49
+ ctx = LLM::Context.new(llm)
50
+
51
+ res = ctx.talk "What's the weather in Tokyo?"
52
+ puts res.content
53
+ ```
54
+
55
+ #### Why would I use it?
56
+
57
+ A bare context gives you control that the agent
58
+ abstraction does not expose. Pre-flight checks on tool requests,
59
+ per-tool confirmation prompts, mixed concurrency strategies across
60
+ tools, or manual iteration until a condition is met are all easier
61
+ with a bare context.
62
+
63
+ #### Notes
64
+
65
+ The agent uses
66
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
67
+ internally. Anything you can do with
68
+ a context, you can also do through an agent. The trade-off is
69
+ convenience versus control. Contexts support the same concurrency
70
+ strategies, compaction, cancellation, and serialization as agents.
71
+
72
+ ### Manual loop
73
+
74
+ #### Overview
75
+
76
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
77
+ manages the tool loop automatically. It calls
78
+ the model, checks for tool requests, runs the tools, feeds results
79
+ back, and repeats until the model produces text. You can bypass
80
+ this and drive
81
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
82
+ directly instead. This
83
+ gives you finer control over each step of the loop at the cost of
84
+ more code.
85
+
86
+ #### How it works
87
+
88
+ When you want to control the tool loop yourself, drive
89
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
90
+ directly instead of using an agent. Start a conversation, check
91
+ for tool requests, execute them, and feed results back. The full
92
+ loop is under your control. Each call to
93
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
94
+ appends to the conversation and returns the model's response, and
95
+ [`LLM::Context#pending_functions?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions?)
96
+ tells you whether tools were requested.
97
+ From that foundation you can inspect, iterate, or confirm
98
+ per-tool in a single flow:
99
+
100
+ ```ruby
101
+ require "llm"
102
+
103
+ llm = LLM.deepseek(key: ENV["KEY"])
104
+ ctx = LLM::Context.new(llm)
105
+
106
+ loop do
107
+ res = ctx.talk("What's the weather in Tokyo?")
108
+ break unless ctx.pending_functions?
109
+
110
+ puts "Model requested #{ctx.pending_functions.size} tool(s)"
111
+
112
+ results = ctx.pending_functions.map do |fn|
113
+ print "Run #{fn.name} with #{fn.arguments}? [y/N] "
114
+ if $stdin.gets&.match?(/\Ay\z/i)
115
+ fn.task(:thread).wait
116
+ else
117
+ fn.cancel(reason: "user declined")
118
+ end
119
+ end
120
+
121
+ ctx.talk(results)
122
+ end
123
+
124
+ puts res.content
125
+ ```
126
+
127
+ #### Why would I use it?
128
+
129
+ Manual control gives you pre-execution checks, custom confirmation
130
+ flows, different strategies per tool, and fine-grained error
131
+ recovery that the default tool loop does not expose.
132
+
133
+ #### Notes
134
+
135
+ [`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
136
+ picks up pending functions, spawns them using the chosen
137
+ strategy, waits for results, and records them back in the context.
138
+ Each strategy is supported: `:sequential`, `:thread`, `:fiber`,
139
+ `:async`, `:fork`, and `:ractor`. Functions are reset after each
140
+ [`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
141
+ or
142
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
143
+ call. Store the array if you need to
144
+ preserve them.
145
+
146
+ ### Pending functions
147
+
148
+ #### Overview
149
+
150
+ Pending function calls represent the model's tool requests. After
151
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
152
+ returns, the context may have pending function
153
+ calls if the model requested tools. These are available through
154
+ [`LLM::Context#pending_functions`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions)
155
+ which returns an array of
156
+ [`LLM::Function`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
157
+ objects. Each function has a name, arguments, and methods for
158
+ execution or cancellation.
159
+
160
+ #### How it works
161
+
162
+ When you want to check whether the model requested tools, call
163
+ [`LLM::Context#pending_functions?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#pending_functions?)
164
+ after each
165
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
166
+ call. Each pending
167
+ function has a name, arguments, and
168
+ methods for execution or cancellation. Call
169
+ [`LLM::Function#task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#task)
170
+ to execute it or
171
+ [`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
172
+ to skip it. Iterate over all
173
+ pending functions to inspect or handle them individually:
174
+
175
+ ```ruby
176
+ res = ctx.talk "What's the weather in Tokyo?"
177
+
178
+ if ctx.pending_functions?
179
+ puts "Model requested #{ctx.pending_functions.size} tool(s)"
180
+ results = ctx.pending_functions.map do |fn|
181
+ print "Run #{fn.name} with #{fn.arguments}? [y/N] "
182
+ if $stdin.gets&.match?(/\Ay\z/i)
183
+ fn.task(:thread).wait
184
+ else
185
+ fn.cancel(reason: "user declined")
186
+ end
187
+ end
188
+ ctx.talk(results)
189
+ end
190
+ ```
191
+
192
+ #### Why would I use it?
193
+
194
+ Inspecting pending functions lets you decide which tools to run,
195
+ in what order, and with what strategy. This is essential for
196
+ confirmation flows, selective execution, or logging which tools
197
+ the model requested.
198
+
199
+ #### Notes
200
+
201
+ Pending functions are reset after each
202
+ [`LLM::Context#wait`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#wait)
203
+ or
204
+ [`LLM::Context#talk`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#talk)
205
+ call. If you need to preserve them, store the array before
206
+ executing. Functions that are cancelled still count as completed
207
+ from the model's perspective; the model sees a cancellation
208
+ result, not a tool error.
209
+
210
+ ### Tool responses
211
+
212
+ #### Overview
213
+
214
+ A tool interrupt gives you two choices. When a tool receives
215
+ [`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html),
216
+ it can either cancel the turn or return a result. The choice
217
+ depends on the situation.
218
+ A hard cancel aborts the request outright and is the default.
219
+ Returning a value lets the model adapt and continue the
220
+ conversation, which can be useful when the interrupt is
221
+ temporary, like a timeout or a user pause.
222
+
223
+ #### How it works
224
+
225
+ When a tool receives
226
+ [`LLM::Interrupt`](https://r.uby.dev/api-docs/llm.rb/LLM/Interrupt.html),
227
+ re-raise to abort the turn or return a value to continue the loop.
228
+ The model receives the result and decides what to do next.
229
+
230
+ Re-raise to abort the turn entirely:
231
+
232
+ ```ruby
233
+ class MyTool < LLM::Tool
234
+ def call
235
+ # do work
236
+ rescue LLM::Interrupt
237
+ cleanup
238
+ raise
239
+ end
240
+ end
241
+ ```
242
+
243
+ Return a value to continue the loop:
244
+
245
+ ```ruby
246
+ class MyTool < LLM::Tool
247
+ def call
248
+ # do work
249
+ rescue LLM::Interrupt
250
+ cleanup
251
+ {ok: false, reason: "interrupted"}
252
+ end
253
+ end
254
+ ```
255
+
256
+ #### Why would I use it?
257
+
258
+ A hard cancel aborts the request outright. Useful when continuing
259
+ would produce garbage. Returning a value lets the model adapt,
260
+ which can be helpful when the interrupt is temporary.
261
+
262
+ #### Notes
263
+
264
+ The mechanism is the same across all six concurrency strategies.
265
+ The `:ractor` strategy delivers the interrupt through ractor
266
+ message passing. The `:fork` strategy delivers it via xchan.
267
+