llm.rb 14.0.0 → 15.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (92) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +396 -1826
  3. data/README.md +596 -490
  4. data/bin/llm.rb +148 -72
  5. data/data/alibaba.json +1999 -0
  6. data/data/anthropic.json +205 -205
  7. data/data/bedrock.json +2171 -2114
  8. data/data/deepinfra.json +1143 -938
  9. data/data/deepseek.json +4 -5
  10. data/data/google.json +691 -779
  11. data/data/mistral.json +450 -450
  12. data/data/moonshot.json +100 -100
  13. data/data/openai.json +976 -976
  14. data/data/xai.json +193 -116
  15. data/data/zai.json +187 -187
  16. data/{resources → docs}/deepdive/advanced/cancellation.md +2 -2
  17. data/{resources → docs}/deepdive/advanced/compaction.md +5 -3
  18. data/{resources → docs}/deepdive/advanced/context.md +13 -0
  19. data/{resources → docs}/deepdive/advanced/guard.md +1 -1
  20. data/{resources/deepdive/fundamentals → docs/deepdive/features}/concurrency.md +6 -0
  21. data/{resources/deepdive/fundamentals → docs/deepdive/features}/repl.md +56 -2
  22. data/{resources → docs}/deepdive/fundamentals/agents.md +52 -1
  23. data/docs/deepdive/fundamentals/providers.md +159 -0
  24. data/{resources → docs}/deepdive/fundamentals/skills.md +5 -0
  25. data/{resources → docs}/deepdive/fundamentals/stream.md +36 -3
  26. data/{resources → docs}/deepdive/fundamentals/tools.md +87 -23
  27. data/{resources/deepdive/everything_else → docs/deepdive/reference}/cost.md +20 -10
  28. data/docs/deepdive/reference/model_registry.md +271 -0
  29. data/{resources/deepdive/advanced → docs/deepdive/reference}/tracer.md +7 -0
  30. data/{resources → docs}/deepdive.md +35 -27
  31. data/lib/llm/a2a/transport/http.rb +1 -1
  32. data/lib/llm/active_record/acts_as_llm.rb +19 -5
  33. data/lib/llm/agent.rb +62 -5
  34. data/lib/llm/context.rb +93 -46
  35. data/lib/llm/cost.rb +110 -51
  36. data/lib/llm/error.rb +7 -0
  37. data/lib/llm/function/array.rb +1 -1
  38. data/lib/llm/function/fork/task.rb +14 -1
  39. data/lib/llm/function/sequential/group.rb +20 -13
  40. data/lib/llm/function/sequential/task.rb +1 -8
  41. data/lib/llm/function.rb +5 -4
  42. data/lib/llm/message.rb +5 -4
  43. data/lib/llm/provider.rb +7 -0
  44. data/lib/llm/providers/alibaba/error_handler.rb +34 -0
  45. data/lib/llm/providers/alibaba/request_adapter.rb +13 -0
  46. data/lib/llm/providers/alibaba.rb +93 -0
  47. data/lib/llm/providers/anthropic.rb +0 -1
  48. data/lib/llm/providers/bedrock.rb +8 -1
  49. data/lib/llm/providers/deepseek/request_adapter.rb +2 -33
  50. data/lib/llm/providers/google.rb +0 -1
  51. data/lib/llm/providers/ollama.rb +0 -1
  52. data/lib/llm/providers/openai/responses.rb +0 -1
  53. data/lib/llm/providers/openai/schema.rb +37 -0
  54. data/lib/llm/providers/openai.rb +1 -2
  55. data/lib/llm/registry/model.rb +186 -0
  56. data/lib/llm/registry.rb +45 -14
  57. data/lib/llm/repl/bar.rb +11 -13
  58. data/lib/llm/repl/buffer.rb +1 -1
  59. data/lib/llm/repl/color.rb +8 -1
  60. data/lib/llm/repl/command.rb +12 -0
  61. data/lib/llm/repl/commands/model.rb +39 -0
  62. data/lib/llm/repl/input/cache.rb +45 -0
  63. data/lib/llm/repl/input/char.rb +2 -2
  64. data/lib/llm/repl/input.rb +86 -23
  65. data/lib/llm/repl/markdown.rb +25 -1
  66. data/lib/llm/repl/node.rb +7 -0
  67. data/lib/llm/repl/status.rb +17 -3
  68. data/lib/llm/repl/window.rb +91 -11
  69. data/lib/llm/repl.rb +18 -8
  70. data/lib/llm/sequel/plugin.rb +19 -5
  71. data/lib/llm/skill.rb +21 -8
  72. data/lib/llm/stream.rb +27 -0
  73. data/lib/llm/tool.rb +3 -5
  74. data/lib/llm/tools/rg.rb +2 -1
  75. data/lib/llm/transport/curb.rb +23 -3
  76. data/lib/llm/usage.rb +155 -9
  77. data/lib/llm/version.rb +1 -1
  78. data/lib/llm.rb +111 -29
  79. data/llm.gemspec +16 -11
  80. metadata +89 -35
  81. /data/{resources → docs}/deepdive/advanced/transformer.md +0 -0
  82. /data/{resources → docs}/deepdive/advanced/transports.md +0 -0
  83. /data/{resources/deepdive/fundamentals → docs/deepdive/features}/builtin_tools.md +0 -0
  84. /data/{resources/deepdive/fundamentals → docs/deepdive/features}/database.md +0 -0
  85. /data/{resources/deepdive/fundamentals → docs/deepdive/features}/embeddings.md +0 -0
  86. /data/{resources → docs}/deepdive/fundamentals/schema.md +0 -0
  87. /data/{resources/deepdive/everything_else → docs/deepdive/media}/audio.md +0 -0
  88. /data/{resources/deepdive/everything_else → docs/deepdive/media}/images.md +0 -0
  89. /data/{resources/deepdive/everything_else → docs/deepdive/media}/ocr.md +0 -0
  90. /data/{resources → docs}/deepdive/protocols/a2a.md +0 -0
  91. /data/{resources → docs}/deepdive/protocols/mcp.md +0 -0
  92. /data/{resources/deepdive/everything_else → docs/deepdive/reference}/object.md +0 -0
data/CHANGELOG.md CHANGED
@@ -15,7 +15,398 @@
15
15
 
16
16
  ## What's next
17
17
 
18
- *No unreleased changes yet. Check back after the next release.*
18
+ ## v15.0.0
19
+
20
+ Changes since `v14.0.0`.
21
+
22
+ This release renames `usage` to `token_usage` across contexts, agents,
23
+ and messages, makes `context_window` return `nil` when unknown, and
24
+ makes `LLM::Cost` accessors always return `Float`. It also adds the
25
+ Alibaba provider, automatic API key discovery from the environment,
26
+ new context-usage and context-used methods, a retry budget for
27
+ rate-limited requests, and a range of REPL and skills improvements.
28
+
29
+ ### Breaking
30
+
31
+ #### Migration
32
+
33
+ | Old | New |
34
+ |-----|-----|
35
+ | `ctx.usage` / `agent.usage` / `msg.usage` | `ctx.token_usage` / `agent.token_usage` / `msg.token_usage` (`usage` remains an alias) |
36
+ | `LLM::Message#usage` returns `LLM::Object` | `LLM::Message#token_usage` returns a copy of `LLM::Usage`, for assistant messages only |
37
+ | `ctx.context_window` returns `0` when the model isn't in the registry | returns `nil` when unknown |
38
+ | `LLM::Cost#input` (and other accessors) return `nil` when unused | return `0.0` |
39
+ | `ctx.usage` returns the most recent assistant message usage | sums token usage across all assistant messages |
40
+ | a skill exposes its tool as `weather` (the skill name) | the generated tool is now named `weather-skill` |
41
+
42
+ * **rename `#usage` to `#token_usage` across contexts, agents, and messages** <br>
43
+ [`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage-instance_method),
44
+ [`LLM::Agent#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#usage-instance_method),
45
+ and
46
+ [`LLM::Message#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Message.html#usage-instance_method)
47
+ are now aliases of `token_usage`.
48
+ [`LLM::Message#token_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Message.html#token_usage-instance_method)
49
+ now returns a copy of `LLM::Usage` instead of `LLM::Object`, and only
50
+ returns a value for assistant messages.
51
+
52
+ * **`LLM::Context#context_window` now returns `nil` when unknown** <br>
53
+ [`LLM::Context#context_window`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_window-instance_method)
54
+ now returns `nil` when the model's context window size is not known to the
55
+ runtime, instead of `0`. This makes the code check for a window instead
56
+ of a number, so an unknown window no longer reads as a real (zero) size.
57
+
58
+ * **`LLM::Cost` accessors always return `Float` objects** <br>
59
+ The cost accessors on
60
+ [`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
61
+ (`input`, `output`, `input_audio`, `output_audio`, `input_image`,
62
+ `cache_read`, `cache_write`, and `reasoning`) now always return a
63
+ `Float`, returning `0.0` when no tokens of that kind were used, instead
64
+ of `nil`. Callers can sum and compare cost values without guarding
65
+ against `nil`.
66
+
67
+ * **expose new context methods on the ActiveRecord and Sequel wrappers** <br>
68
+ The `acts_as_llm` (ActiveRecord) and `plugin :llm` (Sequel) wrappers
69
+ now expose `context_used` and `context_usage`, delegating to the
70
+ wrapped `LLM::Context`. `token_usage` replaces `usage` (which remains
71
+ as an alias), and `context_window` now returns `nil` when the model's
72
+ context window is unknown instead of `0`.
73
+
74
+ ### Core
75
+
76
+ * **discover API keys from the environment** <br>
77
+ [Cloud provider factories](https://r.uby.dev/api-docs/llm.rb/LLM.html)
78
+ (`LLM.anthropic`, `LLM.google`, `LLM.deepseek`, `LLM.openai`,
79
+ `LLM.xai`, `LLM.mistral`, `LLM.zai`, `LLM.moonshot`,
80
+ `LLM.alibaba`, and `LLM.aliyun`) now resolve the provider's API key
81
+ automatically when no `key:` is given, by walking the environment
82
+ variable names listed in the models.dev registry. So `LLM.openai`
83
+ works without an explicit key as long as `OPENAI_API_KEY` (or one of
84
+ the registry's alternative names) is set in the environment. A
85
+ missing key raises `ArgumentError`.
86
+
87
+ * **cli: auto-discover credentials and support Bedrock** <br>
88
+ `bin/llm.rb` now resolves the provider through the `LLM` factory
89
+ methods instead of mapping environment variable names directly, so it
90
+ picks up Bedrock (all three AWS credentials) and relies on the same
91
+ automatic key discovery as the library. The CLI also always starts
92
+ now: without arguments it falls back to `ollama` or `llamacpp`. A
93
+ provider whose credentials are not set exits with status 1.
94
+
95
+ * **cli: add `-c` and `-n` switches** <br>
96
+ `bin/llm.rb` now accepts a `-c STRATEGY` switch to choose the
97
+ concurrency strategy used for tool calls (`thread`, `async`, `fork`,
98
+ or any of the other strategies) and a `-n TRANSPORT` switch to choose
99
+ the HTTP transport (`net-http`, `net-http-persistent`, or `curb`),
100
+ both forwarded to the session's agent and provider.
101
+
102
+ * **context: keep runtime parameters from reaching the provider** <br>
103
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
104
+ now strips its runtime-only parameters (`guard`, `retry_budget`,
105
+ `concurrency`, `transformer`, and `compactor`) before merging params
106
+ into a provider request, so they can never cross the context-provider
107
+ boundary and risk an API-level error.
108
+
109
+ * **add `retry_budget` support for rate-limited requests** <br>
110
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
111
+ now accepts a `retry_budget:` that automatically sleeps and retries a
112
+ rate-limited request up to the given number of times before raising
113
+ `LLM::RateLimitError`. Each retry sleeps a growing interval (2s, 4s,
114
+ 6s, ...) and notifies the stream through
115
+ [`LLM::Stream#on_rate_limit`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_rate_limit-instance_method).
116
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
117
+ enables a budget of 5 by default, while a raw context disables it (0)
118
+ unless configured.
119
+
120
+ * **add `LLM::Usage.zero`** <br>
121
+ Add
122
+ [`LLM::Usage.zero`](https://r.uby.dev/api-docs/llm.rb/LLM/Usage.html#zero-class_method)
123
+ as a zero-valued usage object. `LLM::Context#usage`, `LLM::Agent#usage`,
124
+ and the ActiveRecord and Sequel wrappers now return `LLM::Usage` objects
125
+ instead of `LLM::Object` when no provider usage has been recorded yet.
126
+
127
+ * **add `LLM::Context#context_usage` and `LLM::Agent#context_usage`** <br>
128
+ Add
129
+ [`LLM::Context#context_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_usage-instance_method)
130
+ and
131
+ [`LLM::Agent#context_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#context_usage-instance_method),
132
+ which return the fraction of the model's context window currently used as
133
+ a `Rational` (for example `Rational(100, 10_000)`), or `nil` when the used
134
+ amount or the window size is unknown. The REPL status bar now renders this
135
+ fraction instead of computing the remainder from raw token counts.
136
+
137
+ * **add `LLM::Context#context_used` and `LLM::Agent#context_used`** <br>
138
+ Add
139
+ [`LLM::Context#context_used`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_used-instance_method)
140
+ and
141
+ [`LLM::Agent#context_used`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#context_used-instance_method),
142
+ which return the live context size (in tokens) of the most recent
143
+ assistant message, or `nil` when no assistant message has a recorded token
144
+ usage. This fills the gap left after `token_usage` became accumulative and
145
+ no longer represented a single turn, so callers can read how much of the
146
+ context window has been used without walking the messages themselves.
147
+
148
+ ### Provider
149
+
150
+ * **add `LLM::Provider#registry`** <br>
151
+ Add [`LLM::Provider#registry`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#registry-instance_method),
152
+ which returns the provider's model registry. `LLM::Context#registry`
153
+ and `LLM::Agent#registry` now delegate to their underlying provider
154
+ instead of looking it up on their own.
155
+
156
+ * **add `LLM::Alibaba` for Alibaba Cloud Model Studio** <br>
157
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
158
+ is a new provider that talks to
159
+ [Alibaba Cloud Model Studio](https://www.alibabacloud.com/help/en/model-studio/models)
160
+ through its OpenAI-compatible API, including the Qwen3 family of
161
+ models. Create an instance with
162
+ [`LLM.alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM.html#alibaba-class_method),
163
+ also aliased as `LLM.aliyun`, which accepts the same `key:`, `host:`,
164
+ and `base_path:` options as the OpenAI provider. The provider defaults
165
+ to the `deepseek-v4-flash-0731` model and supports chat completions,
166
+ streaming, tool calls, and structured output through the shared
167
+ OpenAI-compatible path; image, audio, moderation, responses, and
168
+ vector store endpoints raise `NotImplementedError`. Model metadata
169
+ ships in `data/alibaba.json` for the registry.
170
+
171
+ * **alibaba: support structured outputs via `json_object`** <br>
172
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
173
+ now supports structured output through a shared `json_object` fallback,
174
+ since Alibaba models do not support `json_schema` natively. The schema is
175
+ described in an injected system message that also satisfies the
176
+ "messages must contain the word json" requirement. The same shared
177
+ fallback now also backs DeepSeek.
178
+
179
+ * **alibaba: default to the pay-as-you-go host** <br>
180
+ The default
181
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
182
+ host is now `dashscope-intl.aliyuncs.com`. Override it globally with
183
+ the `DASHSCOPE_API_HOST` environment variable, or per instance with
184
+ `LLM.alibaba(host: ...)`, for example to point at a Token Plan
185
+ endpoint.
186
+
187
+ * **alibaba: use `DASHSCOPE_API_KEY` as the default key env var** <br>
188
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
189
+ now discovers its API key from `DASHSCOPE_API_KEY` instead of
190
+ `ALIBABA_API_KEY`, following the models.dev registry convention.
191
+
192
+ * **alibaba: raise `LLM::InsufficientQuotaError` for exhausted quota** <br>
193
+ Add
194
+ [`LLM::InsufficientQuotaError`](https://r.uby.dev/api-docs/llm.rb/LLM/InsufficientQuotaError.html),
195
+ a subclass of `LLM::RateLimitError`, for when a provider reports a
196
+ tokens-per-minute (TPM) quota limit.
197
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
198
+ now raises it when Alibaba responds with an `insufficient_quota`
199
+ error. Since it subclasses `RateLimitError`, quota errors are retried
200
+ like other rate limits.
201
+
202
+ * **bedrock: auto-discover AWS credentials from the environment** <br>
203
+ [`LLM.bedrock`](https://r.uby.dev/api-docs/llm.rb/LLM.html#bedrock-class_method)
204
+ now infers its credentials from the `AWS_ACCESS_KEY_ID`,
205
+ `AWS_SECRET_ACCESS_KEY`, and `AWS_REGION` environment variables when
206
+ they are not passed explicitly, matching the other cloud providers. A
207
+ missing key raises `ArgumentError`.
208
+
209
+ * **add `LLM::Bedrock#key?`** <br>
210
+ Add
211
+ [`LLM::Bedrock#key?`](https://r.uby.dev/api-docs/llm.rb/LLM/Bedrock.html#key%3F-instance_method),
212
+ which overrides the superclass method to check all three Bedrock
213
+ credentials (`access_key_id`, `secret_access_key`, and `region`)
214
+ instead of a single API key.
215
+
216
+ ### Function
217
+
218
+ * **make `Sequential::Group` abide by the `LLM::Function::Group` contract** <br>
219
+ [`LLM::Function::Sequential::Group`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Sequential/Group.html)
220
+ now receives an array of
221
+ [`LLM::Function::Sequential::Task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Sequential/Task.html)
222
+ objects instead of raw `LLM::Function` objects, matching the interface
223
+ shared by every other concurrency strategy.
224
+ [`LLM::Function::Array#task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Array.html#task-instance_method)
225
+ wraps each function as a `Sequential::Task` before constructing the
226
+ group, and the group delegates `spawn`, `alive?`, and `wait` to those
227
+ tasks. This fixes `Sequential::Group#alive?`, which always returned
228
+ `false`, and restores guard handling for sequential execution by
229
+ honoring the shared `guarded:` option on `Sequential::Task`.
230
+
231
+ * **function: redirect output streams in `:fork` tool processes** <br>
232
+ The `:fork` concurrency strategy (via
233
+ [`LLM::Function::Fork::Task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Fork/Task.html))
234
+ now redirects the child process's `$stdout` and `$stderr` to
235
+ `File::NULL`, so a forked tool can no longer clobber the parent
236
+ terminal, for example by blanking the curses REPL display. A tool that
237
+ genuinely needs the terminal can still reopen `/dev/tty`; the file
238
+ descriptor stays available to the child.
239
+
240
+ ### Fix
241
+
242
+ * **a2a: fix a typo in the HTTP transport** <br>
243
+ Fix a bug in
244
+ [`LLM::A2A::Transport::HTTP`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A/Transport/HTTP.html)
245
+ where the constructor read `uri.port` instead of `@uri.port`, which
246
+ crashed the program whenever
247
+ [`LLM::A2A.rest`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html#rest-class_method)
248
+ or
249
+ [`LLM::A2A.jsonrpc`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html#jsonrpc-class_method)
250
+ was used. The transport now reads the port from the parsed `@uri`.
251
+
252
+ * **openai: report usage for streamed completions requests** <br>
253
+ Fix a bug in the OpenAI completions path where `params[:stream]` was
254
+ checked after it had been deleted from the params hash, so the check
255
+ always evaluated to `false`. The fix checks the resolved stream's
256
+ `enabled?` instead, so `stream_options: {include_usage: true}` is
257
+ added to streamed requests and API usage is reported back to the
258
+ caller.
259
+
260
+ * **curb: read the stream body and resolve streaming requests** <br>
261
+ Fix two bugs in [`LLM::Transport::Curb`](https://r.uby.dev/api-docs/llm.rb/LLM/Transport/Curb.html)
262
+ that left the `curb` transport unusable. The request body setter now
263
+ reads a streaming request's body stream into a string (dropping the
264
+ chunked transfer header, which curb replaces with a content length),
265
+ and the result builder now accumulates the response body from the
266
+ `on_body` callback instead of leaving it empty.
267
+
268
+ * **cli: handle errors in `main`** <br>
269
+ Wrap all of `bin/llm.rb`'s `main` method in error handling: an
270
+ interrupted session exits gracefully with `Bye!`, an explicit provider
271
+ is passed the resolved transport, and any unexpected error prints a
272
+ formatted diagnostic with a link to issue tracking before exiting.
273
+
274
+ * **cli: persist the session mapping file** <br>
275
+ Fix a bug where `bin/llm.rb` saved the session file at
276
+ `~/.llm.rb/<provider>/<uuid>.json` but never wrote the updated
277
+ working-directory mapping back to `~/.llm.rb/<provider>.json`. The
278
+ mapping file is now written whenever a new session is registered.
279
+
280
+ * **context: aggregate usage across all assistant messages** <br>
281
+ [`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage-instance_method)
282
+ now sums token usage across every assistant message in the conversation
283
+ instead of returning only the first message's usage.
284
+ [`LLM::Cost.from`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
285
+ now subtracts reasoning tokens from the output total and cache-read
286
+ tokens from the input total before pricing, and prices reasoning tokens
287
+ with the model's reasoning rate when one is available.
288
+
289
+ ### Repl
290
+
291
+ * **draw a top chrome row with the cwd and active model** <br>
292
+ The curses-based REPL now draws a white-on-blue row at the very top of
293
+ the screen showing the current working directory on the left and the
294
+ active model on the right. The row is drawn above the transcript and
295
+ uses a new blue status-bar color pair.
296
+
297
+ * **redraw the window on resize** <br>
298
+ The curses-based REPL now handles the terminal resize signal
299
+ (`KEY_RESIZE`) while reading input, clearing and redrawing the entire
300
+ window so the layout stays aligned after the terminal is resized.
301
+
302
+ * **hide the cursor until the window is ready** <br>
303
+ Fix a visual glitch where the curses-based REPL showed the cursor at
304
+ position 0,0 at startup and then jumped it to the input area once the
305
+ window was drawn. The cursor is now hidden until the input field has
306
+ been drawn and the cursor can be placed directly into it.
307
+
308
+ * **collapse the cost to two decimal places** <br>
309
+ [`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
310
+ now renders the total cost with two decimal places (for example
311
+ `$0.01`), so the REPL status bar shows a compact cost estimate instead
312
+ of a long run of digits.
313
+
314
+ * **add auto-complete ability for commands** <br>
315
+ [`LLM::Repl::Command`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl/Command.html)
316
+ subclasses can now override a `complete` method to autocomplete their
317
+ arguments. The method receives the command's parameters as keyword
318
+ arguments, with the non-nil keyword being the active fragment, and
319
+ returns candidate completions. Repeated TAB presses cycle through the
320
+ candidate list.
321
+
322
+ * **add `LLM::Repl#model` and `LLM::Repl#model=`** <br>
323
+ [`LLM::Repl`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl.html#model-instance_method)
324
+ now tracks the active model in its own `model` attribute, seeded from
325
+ the wrapped agent's model. The status bar reads the model through the
326
+ repl instead of the agent, so the model can be switched within a
327
+ session.
328
+
329
+ * **add `/model` command** <br>
330
+ A new `/model <name>` command switches the active model within a
331
+ single REPL session. Its argument auto-completes through the
332
+ text-to-text models in the registry.
333
+
334
+ * **repl: restrict autocomplete to text-to-text models** <br>
335
+ The `/model` command's argument auto-complete now suggests only
336
+ text-to-text models, so embedding and other non-chat models are
337
+ left out of the completion list.
338
+
339
+ * **repl: highlight GitHub-flavored codeblocks** <br>
340
+ The curses-based REPL now parses the GitHub-style ``` fences that
341
+ models commonly emit as real code blocks. Kramdown's native fenced-code
342
+ syntax uses `~~~`, so the ``` fences were previously parsed as inline
343
+ code spans. The language name is now shown in bold white above the code,
344
+ which renders in green.
345
+
346
+ * **repl: fix a scroll render artifact** <br>
347
+ Fix a bug where scrolling upward could leave a piece of text just
348
+ above the status row as a render artifact. The row above the status
349
+ row is now cleared on every buffer render.
350
+
351
+ * **repl: add a buffer row below the blue status bar** <br>
352
+ The curses-based REPL buffer now starts with an empty row below the
353
+ blue status bar, improving the visual spacing of the first exchange
354
+ in the chat.
355
+
356
+ * **repl: pin `curses` and `kramdown` to tested versions** <br>
357
+ The REPL now pins `curses` to `~> 1.6` and `kramdown` to `~> 2.5`
358
+ through `LLM.require`, so it loads gem versions known to have been
359
+ tested instead of whatever happens to be installed.
360
+
361
+ ### Skills
362
+
363
+ * **skills: append `-skill` to the generated tool name** <br>
364
+ A skill is now exposed as a tool named `"<skill>-skill"` instead of
365
+ just the skill's name, so a skill like `weather` that also uses a tool
366
+ named `weather` no longer collides with it (or a same-named global
367
+ tool) in the tool registry.
368
+
369
+ * **skills: add `LLM::Stream` skill lifecycle callbacks** <br>
370
+ Add
371
+ [`LLM::Stream#on_skill_call`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_skill_call-instance_method)
372
+ and
373
+ [`LLM::Stream#on_skill_return`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_skill_return-instance_method),
374
+ which are called before a skill's sub-agent runs and after it finishes.
375
+ `on_skill_return` receives the `LLM::Agent` sub-agent that ran the skill
376
+ along with the resulting `LLM::Response`, so a stream can inspect the
377
+ sub-agent's conversation, tally its usage, or add a verification step. A
378
+ stream can use the two callbacks to know when a skill sub-agent is
379
+ running.
380
+
381
+ ### Registry
382
+
383
+ * **add `LLM::Registry::Model` as a comparable model wrapper** <br>
384
+ Add [`LLM::Registry::Model`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry/Model.html),
385
+ a wrapper around a model's registry metadata (pricing, limits,
386
+ capabilities, and modalities). Models are comparable by price, so
387
+ `models.sort` orders them from cheapest to most expensive. The class
388
+ exposes predicate helpers such as `tool_call?`, `reasoning?`,
389
+ `structured_output?`, `open_weights?`, `text?`, `image?`, `audio?`,
390
+ `pdf?`, and `video?`, plus `input_cost`, `output_cost`, and
391
+ `context_window` accessors.
392
+
393
+ * **gemspec: bundle the deepdive guide from `docs/`** <br>
394
+ The gemspec now packages the deepdive guide from `docs/deepdive.md`
395
+ and `docs/deepdive/*/*.md` after the deepdive sources moved from
396
+ `resources/` to `docs/`, so the full guide ships with the gem.
397
+
398
+ * **make `LLM::Registry#models` return model objects** <br>
399
+ [`LLM::Registry#models`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry.html#models-instance_method)
400
+ now returns a list of
401
+ [`LLM::Registry::Model`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry/Model.html)
402
+ objects instead of model name strings. Use the new
403
+ [`LLM::Registry#keys`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry.html#keys-instance_method)
404
+ method to get the model names.
405
+
406
+ * **refresh DeepInfra model metadata** <br>
407
+ Update `data/deepinfra.json` with current pricing for the DeepSeek
408
+ V4, DeepSeek-V3, DeepSeek-R1-0528, and Kimi-K3 models, and mark
409
+ `structured_output` support for one model.
19
410
 
20
411
  ## v14.0.0
21
412
 
@@ -510,10 +901,9 @@ several agent and tool bugs around persistence, interruption, and naming.
510
901
 
511
902
  ## v13.0.0
512
903
 
513
- v13.0.0 relicenses the project under the MIT license, replacing
514
- the Business Source License that was introduced in v12.0.0. No
515
- commercial license is needed. Commercial, personal, educational, and
516
- all other uses are now permitted under the standard MIT terms.
904
+ v13.0.0 is released under the MIT license. Commercial, personal,
905
+ educational, and all other uses are permitted under the standard
906
+ MIT terms.
517
907
 
518
908
  Seven breaking changes. Concurrency strategies have been renamed
519
909
  (`:call` → `:sequential`, `:task` → `:async`), `spawn` is now
@@ -538,7 +928,7 @@ reliable across all six concurrency backends. The `functions` and
538
928
  | `Compactor.new(model:, token_threshold:)` | `Compactor::Truncate.new(ctx)` |
539
929
  | `on_compaction(ctx, compactor)` | `on_compaction(compactor)` |
540
930
  | `ctx.functions` / `ctx.functions?` | `ctx.pending_functions` / `ctx.pending_functions?` |
541
- | `agent.functions` / `agent.functions?` | `agent.pending_functions` / `agent.pending_functions?` |
931
+ | `agent.functions` / `agent.functions?` | `agent.pending_functions` |
542
932
 
543
933
  ### Breaking
544
934
 
@@ -1681,10 +2071,6 @@ Multiple _opt-in_ tools have been added to the `llm/tools/*.rb`
1681
2071
  directory. They serve as examples and as general-purpose tools
1682
2072
  that happen to power the repository's agents.
1683
2073
 
1684
- The BSL license has been extended to grant additional free waivers
1685
- for non-profits, charities and for companies with 50 or less
1686
- employees.
1687
-
1688
2074
  Other changes include small-ish bug fixes. <br>
1689
2075
  As always, see the changelog details for a thorough overview.
1690
2076
 
@@ -1747,12 +2133,6 @@ As always, see the changelog details for a thorough overview.
1747
2133
 
1748
2134
  ### Change
1749
2135
 
1750
- * **Extend BSL additional use grant** <br>
1751
- The Business Source License additional use grant has been extended to
1752
- include non-profits, charities, and companies with 50 or fewer
1753
- employees, in addition to the existing personal, education, and
1754
- evaluation uses.
1755
-
1756
2136
  * **Change LlamaCpp default port (8080 => 8013)** <br>
1757
2137
  The default port for the LlamaCpp provider has changed from `8080` to
1758
2138
  `8013` since llamacpp itself defaults to that port.
@@ -1776,1813 +2156,3 @@ As always, see the changelog details for a thorough overview.
1776
2156
  beyond. The `LLM::JSONAdapter.dump` method now walks serialized data
1777
2157
  and encodes every string into UTF-8, using `String#scrub` to replace
1778
2158
  bytes that are not valid UTF-8.
1779
-
1780
- ## v12.0.0
1781
-
1782
- Changes since `v11.3.1`.
1783
-
1784
- This release relicenses the project under the Business Source License,
1785
- defaults OpenAI to the Responses API and gpt-image models, adds the
1786
- DeepInfra provider with audio and image support, introduces
1787
- DeepSeek vector-graphics generation and schema support, extends xAI
1788
- image editing, adds `LLM::Schema.defaults` and schema string rendering,
1789
- and makes ActiveRecord and Sequel agent wrappers yield `LLM::Agent`
1790
- instead of polluting the model namespace.
1791
-
1792
- ### Breaking
1793
-
1794
- * **License change** <br>
1795
- The llm.rb runtime has been developed primarily by one
1796
- person for 3 years. That was done on my own time, and
1797
- I haven't made a dime from that work.
1798
-
1799
- So when I saw a multi-million dollar company benefit from
1800
- the work and for it to become the backbone of their AI
1801
- infrastructure and then see them not contribute back or
1802
- offer any kind of support, I decided this is not sustainable,
1803
- or fair.
1804
-
1805
- I assumed good faith and for people to act in the spirit of
1806
- open source but sadly, that's just not the case. I
1807
- have to choose a license that respects my time and effort.
1808
-
1809
- For those reasons, llm.rb is being relicensed under the
1810
- [Business Source license](https://mariadb.com/bsl11/).
1811
- So what does that mean?
1812
-
1813
- In a nutshell:
1814
-
1815
- * Free for personal use.
1816
- * Free for education.
1817
- * Free for evaluation, development, and testing.
1818
- * Commercial production use requires a commercial license.
1819
- * Exemptions on a case-by-case basis
1820
-
1821
- After 4 years, the license expires and it will become
1822
- available under the 0BSDL as it was before v12.0.0.
1823
- These 4 years apply to a specific version, and not the
1824
- project overall.
1825
-
1826
- Going forward, v12.0.0 will be relicensed to respect
1827
- my time, energy, and effort. llm.rb took an incredible
1828
- amount of time and effort, and continues to do so, so
1829
- I want to protect myself from companies who benefit
1830
- from my work but don't respect the time or effort that
1831
- was put into it.
1832
-
1833
- * **OpenAI: default to the Responses API** <br>
1834
- The responses API has both models and features that are unavailable
1835
- on the chat completions API, and the responses API appears to be
1836
- the API of the future for OpenAI.
1837
-
1838
- Worth noting: the llm.rb implementation does **not** store state
1839
- server-side by default. This can be changed with the `store: true`
1840
- option. The legacy chat completions API can be accessed with the
1841
- `mode: :completions` option.
1842
-
1843
- llm.rb has had support for the responses API for quite
1844
- a while but it was not the default, and a number of bugs
1845
- were found and fixed during the process of making it the
1846
- default.
1847
-
1848
- * **OpenAI: use gpt-image for image generation** <br>
1849
- The `dalle` models are in the process of being deprecated, and support
1850
- has been dropped from llm.rb. The `gpt-image` models are the next-generation
1851
- image-generation models from OpenAI.
1852
-
1853
- * **xAI: provide images as base64-encoded data** <br>
1854
- Both xAI, and OpenAI had the option to generate images via a URL
1855
- you can fetch, or as a base64-encoded string embedded directly
1856
- in the response.
1857
-
1858
- OpenAI is moving away from the URL transport since deprecating dalle,
1859
- and with that in mind, llm.rb has dropped support for the URL transport
1860
- across all providers that supported it.
1861
-
1862
- Google, xAI, and OpenAI now consistently provide generated and modified
1863
- images as a base64-encoded string.
1864
-
1865
- * **ActiveRecord: yield `LLM::Agent` to `acts_as_agent`** <br>
1866
- With this change we yield an instance of `LLM::Agent` to the `acts_as_agent`
1867
- method, and drop the methods (such as `model`, `instructions`, etc) that
1868
- were previously defined directly on the model. This keeps the number of
1869
- methods that llm.rb adds to an ActiveRecord model at a minimum and retains
1870
- the same capabilities as before.
1871
-
1872
- * **Sequel: yield `LLM::Agent` to `plugin(:agent)`** <br>
1873
- Ditto as above but for Sequel.
1874
-
1875
- * **Remove the langsmith tracer** <br>
1876
- This code was contributed by a third party but contains
1877
- many anti-patterns that are against llm.rb conventions
1878
- and best practices. It was merged without oversight or
1879
- review, and basically against the ethos of open source.
1880
-
1881
- I also don't have a langsmith account to maintain the
1882
- code. The alternative is the `LLM::Tracer::Telemetry` class
1883
- that was originally written by me, and serves as a
1884
- general-purpose OTP tracer.
1885
-
1886
- ### Add
1887
-
1888
- * **Add a new provider: LLM::DeepInfra** <br>
1889
- [DeepInfra](https://deepinfra.com) provide OpenAI-compatible
1890
- endpoints for a large catalog of hosted open-source and
1891
- open-weight models. <br> Capabilities like tool calling, structured outputs, and
1892
- reasoning can depend on the model.
1893
-
1894
- * **Add new image provider: LLM::DeepInfra::Images** <br>
1895
- [DeepInfra](https://deepinfra.com) provide access to
1896
- diverse set of text-to-image models. <br> Learn more about the
1897
- available models on their [text-to-image models](https://deepinfra.com/models/text-to-image)
1898
- page.
1899
-
1900
- * **DeepSeek: add `LLM::DeepSeek::Images#create` and `#edit`** <br>
1901
- This new API can generate and edit vector graphics (SVGs). <br>
1902
- It is an experimental approach and API.
1903
-
1904
- DeepSeek does not provide an image generation model however
1905
- its text-to-text models can generate SVG documents, and
1906
- that's the approach this feature takes. It is limited
1907
- to vector graphics rather than raster images.
1908
-
1909
- * **DeepSeek: attach `LLM::Response#agent` to image responses** <br>
1910
- The DeepSeek image API is built on top of
1911
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
1912
- Image responses now expose that agent via `res.agent`, which makes
1913
- it possible to carry the same session across multiple generations
1914
- or edits.
1915
-
1916
- * **xAI: add `LLM::XAI::Images#edit`** <br>
1917
- With this change it is possible to both generate images
1918
- from a prompt, and edit an existing image with a prompt.
1919
- xAI now has the same edit and create capabilities that
1920
- OpenAI has.
1921
-
1922
- * **Add `LLM::Schema.defaults`** <br>
1923
- This method lets you map multiple property names to
1924
- different default values. It is similar to `LLM::Schema.required`
1925
- in the sense that it is called after the properties of
1926
- a schema have been defined.
1927
-
1928
- * **Add `LLM::Schema#to_s` and `LLM::Schema.to_s`** <br>
1929
- Schemas can now be rendered as a prompt-friendly string.
1930
- This is useful when the shape of a schema needs to be
1931
- described in natural-language instructions rather than
1932
- passed through a native structured output interface.
1933
-
1934
- * **DeepSeek: add `LLM::Schema` support** <br>
1935
- DeepSeek can now use `schema:` for structured output.
1936
- llm.rb handles this by setting `response_format: {type: "json_object"}`
1937
- and describing the schema in a system message.
1938
-
1939
- * **OpenAI: add local file support to the Responses API** <br>
1940
- Our responses API implementation lacked local file support. <br>
1941
- This change fixes that by supporting both image, document,
1942
- and other media types that OpenAI may support.
1943
-
1944
- * **Add `LLM::Response#id` across all providers** <br>
1945
- This method was previously implemented via `method_missing`,
1946
- and the field name could change depending on the provider.
1947
- The new method is a catch-all that provides a single method
1948
- that works across all providers.
1949
-
1950
- * **Add `LLM::DeepInfra::Audio`** <br>
1951
- DeepInfra implements most of the llm.rb audio interface
1952
- with both the `create_speech` and `create_transcription`
1953
- methods. The `create_translation` method is not implemented,
1954
- and the available text-to-speech and speech-to-text models
1955
- are more varied than other providers.
1956
-
1957
- * **OpenAI: normalize text-to-speech responses** <br>
1958
- The `res.audio` method now returns an
1959
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
1960
- object for OpenAI text-to-speech responses. The object provides
1961
- `encoded`, `decoded`, `content_type`, and `encoding_type`.
1962
-
1963
- * **DeepInfra: normalize text-to-speech responses** <br>
1964
- The `res.audio` method now returns an
1965
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
1966
- object for DeepInfra text-to-speech responses. The object provides
1967
- `encoded`, `decoded`, `content_type`, and `encoding_type`.
1968
-
1969
- ### Fix
1970
-
1971
- * **Fix Google `temperature` parameter fall-through** <br>
1972
- Ensure provider-level `temperature` and other `generationConfig`
1973
- parameters are forwarded to the API correctly instead of being
1974
- silently dropped.
1975
-
1976
- * **Fix Google `generationConfig` collisions** <br>
1977
- Prevent duplicate or conflicting `generationConfig` keys in the
1978
- Google request adapter.
1979
-
1980
- ### Change
1981
-
1982
- * **Change OpenAI defaults** <br>
1983
- The default chat model is now `gpt-5.4-mini`. <br>
1984
- The default image model is now `gpt-image`.
1985
-
1986
- * **Change google defaults** <br>
1987
- The default chat model is now `gemini-3.1-flash-lite` <br>
1988
- The default embeddings model is now `gemini-embedding-2`
1989
-
1990
- * **Change xAI defaults** <br>
1991
- The default chat model is now `grok-4.3`. <br>
1992
- The default image model is now `grok-imagine-image-quality`.
1993
-
1994
- * **Return an `LLM::Object` from `LLM::Response#content!`** <br>
1995
- The Hash-like, indifferent access data structure known as
1996
- `LLM::Object` provides a convenient interface around a Hash
1997
- object. It allows method access via `obj.key`, and decays
1998
- into a Hash in many cases.
1999
-
2000
- The `LLM::Response#content!` method now wraps its content
2001
- in an `LLM::Object` but only after it has parsed its
2002
- content (a JSON string) into a Ruby data structure.
2003
-
2004
- * **Refresh model metadata** <br>
2005
- Update `data/*.json` files with current provider model listings,
2006
- pricing, and capabilities.
2007
-
2008
- ## v11.3.1
2009
-
2010
- Changes since `v11.3.0`.
2011
-
2012
- This release rebrands the project under the r.uby.dev umbrella, removes
2013
- the Jekyll-based docs site in favor of a pure-markdown deepdive, and
2014
- cleans up YARD documentation across the codebase.
2015
-
2016
- ### Change
2017
-
2018
- * **Rebrand to r.uby.dev** <br>
2019
- Update README.md with the new logo, streamlined copy, and r.uby.dev
2020
- URLs. Rewrite `resources/deepdive.md` as a concise walkthrough and
2021
- bundle it with the gem. Remove the `docs/` directory (Jekyll site).
2022
- Update all references from `llmrb.github.io` to `r.uby.dev`.
2023
-
2024
- * **Update gemspec** <br>
2025
- Update homepage, metadata URLs, email, and author list. Switch the
2026
- YARD markdown processor from kramdown to redcarpet.
2027
-
2028
- ### Fix
2029
-
2030
- * **Fix YARD documentation** <br>
2031
- Fix unnamed, misnamed, and missing `@param` tags across provider
2032
- adapters, transport classes, stream, tool, schema, registry, agent,
2033
- and ActiveRecord integration files. Fix backtick-wrapped constant
2034
- references and other YARD formatting issues.
2035
-
2036
- ## v11.3.0
2037
-
2038
- Changes since `v11.2.0`.
2039
-
2040
- This release promotes `LLM::Agent` as the default high-level runtime,
2041
- raises `LLM::NotFoundError` for provider 404 responses, and adds
2042
- Symbol resolution to `LLM::Agent.confirm` and `LLM::Agent.skills` for
2043
- dynamic tool confirmation and skill lists.
2044
-
2045
- ### Add
2046
-
2047
- * **Raise `LLM::NotFoundError` for provider 404 responses** <br>
2048
- Raise `LLM::NotFoundError` when a provider returns HTTP 404. One
2049
- example is calling the embeddings API on DeepSeek
2050
- (`LLM.deepseek(...).embed(["foobar"])`), which returns 404 because
2051
- DeepSeek does not implement that endpoint.
2052
-
2053
- * **Add Symbol resolution to `LLM::Agent.confirm`** <br>
2054
- When `confirm` receives a single Symbol argument, it stores it
2055
- as-is instead of converting it to a string array. At initialization
2056
- time, `resolve_option` resolves the Symbol by calling the method
2057
- with that name on the agent instance, and the result is converted
2058
- to strings. This allows dynamic tool confirmation lists:
2059
-
2060
- class MyAgent < LLM::Agent
2061
- confirm :tools_that_need_confirmation
2062
-
2063
- def tools_that_need_confirmation
2064
- some_condition ? %w[delete destroy] : %w[delete]
2065
- end
2066
- end
2067
-
2068
- Ported from llmrb/mruby-llm@89a232e3 and @2dd04e2d.
2069
-
2070
- Extend the same pattern to `LLM::Agent.skills` so the skills DSL
2071
- accepts a Symbol that resolves through the agent instance at
2072
- initialization time.
2073
-
2074
- ### Change
2075
-
2076
- * **Clarify `LLM::Agent` as the default high-level runtime** <br>
2077
- Document that `LLM::Context` remains at the heart of llm.rb, but
2078
- `LLM::Agent` is the better default unless an application needs advanced
2079
- manual tool loops. `LLM::Agent` manages the tool loop for callers and
2080
- enables guards against runaway or repeated tool-call loops.
2081
-
2082
- ## v11.2.0
2083
-
2084
- Changes since `v11.1.0`.
2085
-
2086
- This release adds `LLM::Function#skill?` and `LLM::Tool#skill?` so
2087
- callers can inspect whether a function or tool is backed by a skill.
2088
-
2089
- It introduces `LLM::Transport::Request` as a transport-agnostic request
2090
- object so providers no longer depend directly on `Net::HTTP` request
2091
- classes, and adds an optional Curb (libcurl) backend alongside symbolic
2092
- transport shortcuts such as `transport: :curb`.
2093
-
2094
- MCP and A2A clients now accept `persistent: true` matching provider configuration.
2095
- Several fixes land for tool return callback emission, function comparison by
2096
- tool call ID, function array filtering, skill tool inheritance, and JSON generator
2097
- state compatibility on Ruby 4.
2098
-
2099
- ### Add
2100
-
2101
- * **Add `LLM::Function#skill?`** <br>
2102
- Add `skill?` to `LLM::Function` so callers can check whether a
2103
- function is backed by a skill tool.
2104
-
2105
- * **Add `LLM::Tool.skill?` and `LLM::Tool#skill?`** <br>
2106
- Add class-level `skill?` and instance-level `skill?` to
2107
- `LLM::Tool`, matching the existing `mcp?` and `a2a?` pattern.
2108
-
2109
- * **Add `LLM::Transport::Request`** <br>
2110
- Add `LLM::Transport::Request` as a transport-agnostic request object
2111
- and update providers to build requests without depending directly on
2112
- Net::HTTP request classes. The built-in Net::HTTP transports still
2113
- accept existing Net::HTTP request objects through a compatibility
2114
- bridge, while alternative transports can handle the generic request
2115
- shape directly.
2116
-
2117
- * **Add optional Curb transport support** <br>
2118
- Add `LLM::Transport::Curb`, an optional libcurl-backed transport
2119
- that can be selected with `transport: :curb`. Providers already
2120
- emit `LLM::Transport::Request` objects, so the Curb backend can
2121
- execute requests without routing through Net::HTTP.
2122
-
2123
- * **Add symbolic transport shortcuts** <br>
2124
- Allow providers, MCP HTTP clients, and A2A HTTP clients to accept
2125
- transport shortcuts such as `transport: :curb` and
2126
- `transport: :net_http_persistent`.
2127
-
2128
- * **Add persistent HTTP selection to MCP and A2A clients** <br>
2129
- Allow MCP and A2A HTTP clients to accept `persistent: true`, matching
2130
- provider configuration and selecting the persistent Net::HTTP
2131
- transport by default.
2132
-
2133
- ### Fix
2134
-
2135
- * **Support JSON generation state on Ruby 4** <br>
2136
- Handle JSON generator state objects in the standard JSON adapter so
2137
- schema objects serialize correctly when Ruby 4 calls custom `to_json`
2138
- methods during provider request generation.
2139
-
2140
- * **Emit tool return callbacks for direct context waits** <br>
2141
- Emit `LLM::Stream#on_tool_return` when `LLM::Context#wait` executes
2142
- pending tool work directly instead of draining `LLM::Stream::Queue`.
2143
-
2144
- * **Emit confirmed tool return callbacks once** <br>
2145
- Emit `LLM::Stream#on_tool_return` for confirmed and cancelled tool
2146
- calls, and exclude confirmed functions from later waits so mixed
2147
- confirmed and unconfirmed tool batches do not execute confirmed tools
2148
- twice.
2149
-
2150
- * **Compare functions by tool call ID** <br>
2151
- Add `LLM::Function#==`, `#eql?`, and `#hash` so pending function
2152
- collections can compare tool calls by provider-assigned ID instead of
2153
- object identity.
2154
-
2155
- * **Preserve function array behavior after filtering** <br>
2156
- Preserve `LLM::Function::Array` behavior when subtracting function
2157
- arrays so filtered tool batches can still spawn through the normal
2158
- function array API.
2159
-
2160
- * **Prevent skills from inheriting skill-backed tools** <br>
2161
- Exclude skill-backed tools when a skill sub-agent uses `tools:
2162
- inherit`, preventing skills loaded through a parent context from
2163
- being recursively exposed to nested skill agents.
2164
-
2165
- ## v11.1.0
2166
-
2167
- Changes since `v11.0.0`.
2168
-
2169
- This release adds the `inherit` directive for skill sub-agents so they can
2170
- inherit access to the local, MCP, and A2A tools available to their parent
2171
- agent. It introduces class-level `required %i[...]` declarations to
2172
- `LLM::Schema` and wraps `LLM::Function#arguments` in `LLM::Object` for
2173
- method-style argument access. The OpenTelemetry tracer now samples all spans
2174
- regardless of environment, and the tool-call loop repair step prevents stale
2175
- history from being sent on follow-up requests.
2176
-
2177
- ### Add
2178
-
2179
- * **Add support for the `inherit` directive in skills** <br>
2180
- Add support for the `inherit` directive so a skill sub-agent can
2181
- inherit access to the local, MCP, and A2A tools available to its
2182
- parent agent.
2183
-
2184
- * **Add class-level `required %i[...]` support to `LLM::Schema`** <br>
2185
- Add class-level `required %i[...]` declarations to `LLM::Schema`, so
2186
- schema classes can mark existing properties as required the same way
2187
- `LLM::Tool` params already can.
2188
-
2189
- * **Wrap function arguments in `LLM::Object`** <br>
2190
- Wrap `LLM::Function#arguments` in `LLM::Object`, so function
2191
- implementations can read arguments with method-style access while
2192
- still invoking runners with keyword arguments.
2193
-
2194
- ### Fix
2195
-
2196
- * **Ensure all traces are sampled regardless of environment** <br>
2197
- Explicitly pass `Samplers::ALWAYS_ON` when creating the OpenTelemetry
2198
- `TracerProvider` so the in-memory exporter always captures every span,
2199
- regardless of the `OTEL_TRACES_SAMPLER` environment variable.
2200
-
2201
- * **Always close the tool call loop before sending follow-up requests** <br>
2202
- Add a repair step in `Context#talk` that closes assistant tool-call
2203
- messages without matching tool responses before the next provider
2204
- request is sent. This prevents stale tool-call history from being sent
2205
- on follow-up requests, which some providers reject as invalid.
2206
-
2207
- ## v11.0.0
2208
-
2209
- Changes since `v10.0.0`.
2210
-
2211
- This release removes several deprecated or unused APIs, including the `#chat`
2212
- alias from contexts and agents, the `LLM::Function#register` alias, and the
2213
- unused positional `llm` argument from MCP constructors. Generated MCP and A2A
2214
- tools are no longer added to the global tool registry by default.
2215
-
2216
- On the additions side, it introduces the A2A (Agent2Agent) protocol client,
2217
- a new `#ask` convenience interface on contexts and agents, one-shot stdio MCP
2218
- requests outside `#session`, `LLM::Function#def` as a short alias for
2219
- `LLM::Function#define`, `LLM::File#exist?`, and `LLM::Tool.a2a?`.
2220
-
2221
- ### Breaking
2222
-
2223
- * **Remove the unused `llm` argument from MCP clients** <br>
2224
- Remove the unused positional `llm` argument from `LLM::MCP.new`,
2225
- `LLM::MCP.stdio`, `LLM::MCP.http`, and `LLM.mcp`.
2226
-
2227
- * **Stop globally registering generated MCP and A2A tools** <br>
2228
- Generated tools returned by `LLM::Tool.mcp(...)` and
2229
- `LLM::Tool.a2a(...)` are no longer added to the global
2230
- `LLM::Tool.registry` or `LLM::Function.registry`. They still work
2231
- when passed directly to a context or agent, but registry-based lookup
2232
- now only sees normal loaded `LLM::Tool` subclasses.
2233
-
2234
- * **Remove `LLM::Function#register`** <br>
2235
- Remove the `LLM::Function#register` alias and prefer
2236
- `LLM::Function#define` or `LLM::Function#def` when binding a
2237
- function to its implementation. The `register` alias was too easy to
2238
- confuse with the class-level `LLM::Tool.register` and
2239
- `LLM::Function.register` registry APIs.
2240
-
2241
- * **Remove the `#chat` alias from contexts and agents** <br>
2242
- Remove the `LLM::Context#chat` and `LLM::Agent#chat` aliases. Prefer
2243
- `#talk` for all context and agent turns.
2244
-
2245
- ### Add
2246
-
2247
- * **Add `LLM::Function#def`** <br>
2248
- Add `LLM::Function#def` as a short alias for
2249
- `LLM::Function#define` when binding a function instance to its
2250
- implementation.
2251
-
2252
- * **Add `LLM::MCP#session`** <br>
2253
- Add `LLM::MCP#session` as an alias for `LLM::MCP#run`, and prefer it
2254
- in examples for scoped stdio MCP sessions that should stay alive
2255
- across discovery and tool calls.
2256
-
2257
- * **Add `#ask` to contexts and agents** <br>
2258
- Add `LLM::Context#ask` and `LLM::Agent#ask` as a RubyLLM-compatible
2259
- convenience interface over `#talk`. `#ask` accepts a prompt, optional
2260
- `with:` attachments, an optional `stream:` target, and an optional
2261
- block for streamed chunks, and returns an `LLM::Response`.
2262
-
2263
- * **Add `LLM::File#exist?`** <br>
2264
- Add `LLM::File#exist?` as a small convenience wrapper for checking
2265
- whether a local file exists on disk.
2266
-
2267
- * **Allow one-shot stdio MCP requests outside `#session`** <br>
2268
- Allow `mcp.tools`, `mcp.prompts`, `mcp.find_prompt(...)`, and
2269
- `mcp.call_tool(...)` to work outside `mcp.session` by starting and
2270
- stopping a stdio transport on demand when needed. This makes stdio
2271
- MCP usable without an explicit session block, while keeping
2272
- `mcp.session` as the preferred pattern for efficient, stateful
2273
- stdio workflows.
2274
-
2275
- * **Add A2A client support** <br>
2276
- Add `LLM::A2A`, a client for the Agent2Agent (A2A) protocol with
2277
- REST and JSON-RPC bindings. Remote agent skills can be exposed as
2278
- `LLM::Tool` classes and used through `LLM::Context` or `LLM::Agent`,
2279
- and the client also supports direct messaging, streaming, task
2280
- operations, push notification configuration, extended agent cards,
2281
- persistent HTTP transport selection, and optional REST `base_path`
2282
- prefixing.
2283
-
2284
- Refactor shared MCP/A2A HTTP transport setup into
2285
- `LLM::Transport::Utils`, and extend
2286
- `LLM::Transport::StreamDecoder` to accept a callback block directly.
2287
-
2288
- * **Add `LLM::Tool.a2a?`** <br>
2289
- Add `LLM::Tool.a2a?` and mark generated A2A-backed tool classes so
2290
- callers can distinguish them from local or MCP tools.
2291
-
2292
- ### Fix
2293
-
2294
- * **Fix context and agent JSON serialization through `LLM.json`** <br>
2295
- Fix `LLM::Context#to_json` and `LLM::Agent#to_json` to serialize
2296
- through `LLM.json.dump(...)` instead of plain `to_json`.
2297
-
2298
- * **Fix block-form ORM agent DSL forwarding** <br>
2299
- Fix block-form `model { ... }`, `tools { ... }`, and
2300
- `schema { ... }` declarations in the ActiveRecord and Sequel agent
2301
- wrappers so persisted agent models configure the internal agent class
2302
- the same way `LLM::Agent` does.
2303
-
2304
- * **Fix missing `skills` in ORM agent wrappers** <br>
2305
- Fix the ActiveRecord and Sequel agent wrappers to expose `skills`, so
2306
- persisted agent models can declare skills the same way as
2307
- `LLM::Agent`.
2308
-
2309
- * **Fix `acts_as_agent#ctx` return type** <br>
2310
- Fix the ActiveRecord `acts_as_agent` wrapper so its `ctx` helper
2311
- returns the wrapped `LLM::Agent` instead of returning the underlying
2312
- `LLM::Context` directly.
2313
-
2314
- ## v10.0.0
2315
-
2316
- Changes since `v9.0.0`.
2317
-
2318
- This release removes the `LLM::Context#respond` method, and
2319
- also removes the deprecated `LLM::Bot` alias. **All** class-level
2320
- agent tunables can now be resolved lazily via a Symbol (method name),
2321
- or a Proc. The `LLM::Agent` class can now confirm a tool call
2322
- before it happens, and the `LLM::Schema` class has been extended
2323
- to support `Array[String,Integer]` as a shorthand for
2324
- `Array[AnyOf[String, Integer]]`. The `LLM::Stream` class has
2325
- had its public method surface reduced to help avoid accidental
2326
- collisions.
2327
-
2328
- ### Breaking
2329
-
2330
- * **Unify context turns under `#talk`** <br>
2331
- Remove `LLM::Context#respond` and route responses-mode turns through
2332
- `LLM::Context#talk` with `mode: :responses` instead.
2333
-
2334
- * **Remove the `LLM::Bot` alias** <br>
2335
- Remove the backward-compatible `LLM::Bot` alias for `LLM::Context`.
2336
- Use `LLM::Context` directly instead.
2337
-
2338
- ### Add
2339
-
2340
- * **Add shared option resolution through `LLM::Utils`** <br>
2341
- Add `LLM::Utils.resolve_option` for resolving configured values as
2342
- literals, procs, symbol-named methods, or duplicated hashes, and use
2343
- it in agent and ORM option resolution paths.
2344
-
2345
- * **Resolve all class-level agent tunables via Proc** <br>
2346
- Let `model`, `tools`, `skills`, `schema`, `stream`, and `tracer`
2347
- declared with a block be lazily evaluated against the agent instance
2348
- at initialization time, matching how `stream` and `tracer` already
2349
- worked.
2350
-
2351
- Add `LLM::Agent#params` for direct access to the underlying context
2352
- parameters.
2353
-
2354
- Ported from mruby-llm.
2355
-
2356
- * **Support `Array[...]` schema and tool param types** <br>
2357
- Let `LLM::Schema` properties and `LLM::Tool` params accept
2358
- `Array[...]` type declarations, including mixed item unions that are
2359
- serialized as `anyOf` array items.
2360
-
2361
- * **Add `LLM::Provider#key?`** <br>
2362
- Add `key?` to providers so callers can check whether a non-blank API
2363
- key has been configured.
2364
-
2365
- * **Add agent tool confirmation hooks** <br>
2366
- Add `LLM::Agent.confirm` and `LLM::Agent#on_tool_confirmation` so
2367
- selected tools can be approved or cancelled before execution. Pending
2368
- tool resolution now relies on `LLM::Context#functions` so confirmed
2369
- tools are not executed twice when mixed with unconfirmed tool calls.
2370
-
2371
- * **Add `LLM::Function#spawn(:call).wait`** <br>
2372
- Add task-shaped sequential execution support for direct
2373
- `LLM::Function#spawn(:call).wait`.
2374
-
2375
- ### Fix
2376
-
2377
- * **Reduce private internal methods on `LLM::Stream`** <br>
2378
- Remove `tool_not_found` and `__tools__` from `LLM::Stream`. The
2379
- `__tools__` logic is inlined directly into `__find__` since that
2380
- was its only caller. The `tool_not_found` utility method was unused
2381
- externally and added unnecessary surface to LLM::Stream.
2382
-
2383
- Ported from mruby-llm.
2384
-
2385
- ## v9.0.0
2386
-
2387
- Changes since `v8.1.0`.
2388
-
2389
- This release deepens llm.rb's transport and cost-tracking surface. It
2390
- replaces the old mutable `persist!` API with constructor-driven transport
2391
- selection, removes `#call` from contexts and agents in favor of explicit
2392
- `ctx.wait(:call)`, makes queued stream waits strategy-free, and deletes
2393
- the unused `LLM::Utils` module.
2394
-
2395
- It adds cache read/write token tracking
2396
- with corresponding cost components, audio and image token pricing,
2397
- `LLM::Context#functions?` for queue-aware tool loops,
2398
- `LLM::Agent.stream` DSL support, and exposes `#stream` readers on
2399
- contexts and agents.
2400
-
2401
- The HTTP transport layer has been refactored around shared backends so
2402
- providers, MCP, and custom transports all use the same normalized
2403
- response interface.
2404
-
2405
- ### Breaking
2406
-
2407
- * **Remove `#call` as a context and agent tool-loop API** <br>
2408
- Remove `LLM::Context#call(:functions)` and `LLM::Agent#call(:functions)`.
2409
- Tool loops should use `ctx.wait(:call)` or `agent.wait(:call)` instead.
2410
- The ActiveRecord and Sequel wrappers no longer expose `#call` passthroughs
2411
- for stored llm.rb contexts.
2412
-
2413
- * **Make HTTP transport selection constructor-driven** <br>
2414
- Remove public `persist!` and `.persistent` mutation APIs from
2415
- providers, transports, and MCP clients. Select persistent behavior at
2416
- construction time with `persistent: true`, `LLM::Transport.net_http`,
2417
- `LLM::Transport.net_http_persistent`, or an explicit `transport:`
2418
- override.
2419
-
2420
- * **Make queued stream waits strategy-free** <br>
2421
- Change `LLM::Stream::Queue#wait` to resolve queued work by the actual
2422
- task types already present in the queue instead of accepting an
2423
- external wait strategy. `LLM::Stream#wait(...)` remains compatible but
2424
- now ignores its arguments when delegating to the queue.
2425
-
2426
- * **Remove unused `LLM::Utils`** <br>
2427
- Delete the `LLM::Utils` module and remove its remaining unused
2428
- provider includes and top-level require.
2429
-
2430
- ### Add
2431
-
2432
- * **Expose `#stream` readers on contexts and agents** <br>
2433
- Add public `LLM::Context#stream` and `LLM::Agent#stream` accessors so
2434
- callers can inspect the active stream object directly.
2435
-
2436
- * **Track cache read and write tokens in usage** <br>
2437
- Add `cache_read_tokens` and `cache_write_tokens` to `LLM::Usage` and
2438
- preserve them through completion usage adaptation and context usage
2439
- aggregation.
2440
-
2441
- * **Add `LLM::Context#functions?` for queue-aware tool loops** <br>
2442
- Add `functions?` to `LLM::Context` and the ActiveRecord and Sequel
2443
- wrappers so callers can detect pending tool work through either the
2444
- bound stream queue or unresolved functions, and update the docs to
2445
- prefer `while ctx.functions?` over `ctx.functions.any?` in tool-loop
2446
- examples.
2447
-
2448
- * **Add `:call` as a first-class wait strategy** <br>
2449
- Add `:call` to pending-function wait paths so `ctx.wait(:call)` can
2450
- prefer queued streamed work when present and otherwise fall back to
2451
- direct sequential function execution through `spawn(:call).wait`.
2452
-
2453
- * **Read provider cache usage into completion responses** <br>
2454
- Read cache read tokens from provider usage metadata, including OpenAI
2455
- `usage.prompt_tokens_details` and Anthropic
2456
- `usage.cache_read_input_tokens`. Read Anthropic cache write tokens
2457
- from `usage.cache_creation_input_tokens`, and expose explicit
2458
- zero-valued `cache_write_tokens` methods on providers that do not
2459
- report cache creation usage.
2460
-
2461
- * **Extend cost tracking with cache write pricing** <br>
2462
- Extend `LLM::Cost` with `cache_read_costs`, `cache_write_costs`, and
2463
- `reasoning_costs` alongside the existing `input_costs` and
2464
- `output_costs`. Add `#to_h` for structured cost insight and update
2465
- `ctx.cost` to calculate all available components from registry
2466
- pricing data.
2467
-
2468
- * **Price input and output audio separately** <br>
2469
- Track `input_audio_tokens` and `output_audio_tokens` in usage and
2470
- include `input_audio_costs` and `output_audio_costs` in `LLM::Cost`
2471
- so multimodal requests report accurate audio spend.
2472
-
2473
- * **Track image tokens in input cost reporting** <br>
2474
- Add `input_image_tokens` to usage and include `input_image_costs` in
2475
- `LLM::Cost` using the model's generic input rate so image-bearing
2476
- prompts report their input spend.
2477
-
2478
- * **Add `LLM::Agent.stream` DSL support** <br>
2479
- Let agents define a default `stream` through the class DSL, including
2480
- block-based stream construction so each agent instance can resolve its
2481
- stream the same way `tracer` does.
2482
-
2483
- ### Change
2484
-
2485
- * **Refactor HTTP transports around shared backends** <br>
2486
- Split `Net::HTTP` and `Net::HTTP::Persistent` into separate
2487
- `LLM::Transport` implementations, move HTTP-specific request helpers
2488
- and response execution into the shared transport layer, and let MCP
2489
- HTTP wrap those transports instead of maintaining a separate
2490
- transient/persistent client split.
2491
-
2492
- * **Share transport overrides across providers and MCP** <br>
2493
- Let both provider construction and `LLM::MCP.http(...)` accept
2494
- `LLM::Transport` instances or classes as HTTP transport overrides, so
2495
- callers can reuse the same transport implementation across the
2496
- runtime.
2497
-
2498
- * **Let custom transports adapt their own response objects** <br>
2499
- Introduce a transport response interface so custom transports can
2500
- adapt backend-specific response objects to one normalized shape and
2501
- have them work with the existing provider execution and error-handling
2502
- code.
2503
-
2504
- ## v8.1.0
2505
-
2506
- Changes since `v8.0.0`.
2507
-
2508
- This release adds Amazon Bedrock provider support through the Converse
2509
- API, including AWS SigV4 request signing, event stream decoding,
2510
- structured output through `schema:`, and a models.dev-backed registry.
2511
- It exposes `llm.models.all` for Bedrock via the ListFoundationModels
2512
- API and adds `LLM::Object#transform_values!` for in-place value
2513
- transformation. Several Bedrock-specific fixes land as well, including
2514
- response id exposure, blank text block suppression in tool turns, and
2515
- DSML tool-marker filtering in streamed text.
2516
-
2517
- ### Add
2518
-
2519
- * **Add AWS Bedrock provider support** <br>
2520
- Add `LLM.bedrock(...)` with Bedrock Converse chat support, AWS SigV4
2521
- request signing, Bedrock event stream decoding, structured output
2522
- support through `schema:`, and models.dev-backed `bedrock.json`
2523
- registry generation.
2524
-
2525
- * **Add AWS Bedrock Models endpoint support** <br>
2526
- Add `llm.models.all` for Bedrock via the ListFoundationModels API,
2527
- including SigV4 signing for the control-plane endpoint and normalized
2528
- `LLM::Model` collection responses.
2529
-
2530
- * **Add `LLM::Object#transform_values!`** <br>
2531
- Let `LLM::Object` transform stored values in place through
2532
- `#transform_values!`.
2533
-
2534
- ### Fix
2535
-
2536
- * **Expose response ids on Bedrock completion responses** <br>
2537
- Read the Bedrock request id into `LLM::Response#id` for completion
2538
- responses adapted from the Converse API.
2539
-
2540
- * **Avoid blank assistant text blocks in Bedrock tool turns** <br>
2541
- Stop replaying assistant tool-call messages with empty text content
2542
- blocks that Bedrock rejects.
2543
-
2544
- * **Suppress Bedrock DSML tool markers in streamed text** <br>
2545
- Filter `"\u003c\u003cDSML\u003efunction_calls\u003e\u003e"` markers out of streamed Bedrock
2546
- assistant text so tool-call sentinels do not leak into user-visible
2547
- output.
2548
-
2549
- ## v8.0.0
2550
-
2551
- Changes since `v7.0.0`.
2552
-
2553
- This release adds Unix-fork concurrency for process-isolated tool
2554
- execution, extends `LLM::Object` with `#merge` and `#delete`, and drops
2555
- Ruby 3.2 support due to a segfault observed with the `:fork` path. It
2556
- promotes `LLM::Pipe` to the top-level namespace and adds
2557
- `persistent: true` on `LLM::MCP.http` for direct persistent transport
2558
- configuration. `LLM::Function#runner` is exposed as public API, agent
2559
- tracer overrides are supported, fiber execution now uses `Fiber.schedule`,
2560
- missing optional dependencies raise clearer `LLM::LoadError` guidance,
2561
- and ActiveRecord wrapper plumbing is deduplicated between `acts_as_llm`
2562
- and `acts_as_agent`.
2563
-
2564
- ### Breaking
2565
-
2566
- * **Drop Ruby 3.2 support** <br>
2567
- Stop supporting Ruby 3.2 due to a segfault observed with the `:fork`
2568
- tool concurrency strategy.
2569
-
2570
- ### Add
2571
-
2572
- * **Add `LLM::Object#merge`** <br>
2573
- Let `LLM::Object` return a new wrapped object when merging hash-like
2574
- data through `#merge`.
2575
-
2576
- * **Add `LLM::Object#delete`** <br>
2577
- Let `LLM::Object` delete keys directly through `#delete`.
2578
-
2579
- ### Change
2580
-
2581
- * **Add fork-based tool concurrency** <br>
2582
- Add `:fork` as a new concurrency strategy for `LLM::Function#spawn`,
2583
- `LLM::Function::Array#wait`, and `LLM::Agent.concurrency` that runs
2584
- class-based tools in isolated child processes. Fork-backed tools support
2585
- tracer callbacks, `on_interrupt`/`on_cancel` hooks, and `alive?` checks.
2586
- Requires the `xchan` gem for inter-process communication with `:fork`.
2587
- This is especially useful for tools that need process isolation, such as
2588
- running shell commands or handling unsafe data.
2589
-
2590
- * **Promote `LLM::Pipe` from MCP namespace to top-level** <br>
2591
- Move `LLM::MCP::Pipe` to `LLM::Pipe` so the pipe abstraction is available
2592
- outside MCP internals. The new class adds a `binmode:` option for binary
2593
- pipes. `LLM::MCP::Command` and related MCP transport code have been updated
2594
- to use `LLM::Pipe`.
2595
-
2596
- * **Allow `persistent: true` on `LLM::MCP.http`** <br>
2597
- Let `LLM::MCP.http(...)` enable persistent HTTP transport directly
2598
- through `persistent: true` at construction time.
2599
-
2600
- * **Expose `LLM::Function#runner` as public API** <br>
2601
- Promote the internal runner instantiation to a public `runner` method on
2602
- `LLM::Function`, so callers can inspect or reuse the resolved tool instance
2603
- that a function wraps.
2604
-
2605
- * **Allow agent instance tracer overrides** <br>
2606
- Let `LLM::Agent.new(..., tracer: ...)` override the class-level tracer
2607
- for that agent instance.
2608
-
2609
- * **Make `:fiber` use scheduler-backed fibers** <br>
2610
- Change `:fiber` tool execution to use `Fiber.schedule` and require
2611
- `Fiber.scheduler`, instead of wrapping direct calls in raw fibers. This
2612
- gives `:fiber` a real cooperative concurrency model instead of acting as
2613
- a thin wrapper around sequential execution.
2614
-
2615
- * **Read stored values from zero-argument `LLM::Object` method calls** <br>
2616
- Let calls like `obj.delete`, `obj.fetch`, `obj.merge`, `obj.key?`,
2617
- `obj.dig`, `obj.slice`, or `obj.keys` return a stored value when that
2618
- method name exists as a key and no arguments are given.
2619
-
2620
- * **Harden `LLM::Object` against arbitrary key names** <br>
2621
- Move internal lookup logic off `LLM::Object` instances and onto the
2622
- singleton class instead, making stored keys like `method_missing`
2623
- more resilient while preserving normal dynamic field access.
2624
-
2625
- * **Deduplicate ActiveRecord wrapper plumbing** <br>
2626
- Move shared ActiveRecord wrapper defaults and utility methods into
2627
- `LLM::ActiveRecord`, reducing duplication between `acts_as_llm` and
2628
- `acts_as_agent`.
2629
-
2630
- * **Raise clearer errors for missing optional runtime dependencies** <br>
2631
- Route optional `async`, `xchan`, and `net/http/persistent` loads
2632
- through `LLM.require` so missing runtime gems raise `LLM::LoadError`
2633
- with installation guidance instead of leaking raw `LoadError`
2634
- exceptions.
2635
-
2636
- ### Fix
2637
-
2638
- * **Avoid `RuntimeError` from `Async::Task.current` lookups** <br>
2639
- Check `Async::Task.current?` before reading the current Async task so
2640
- provider transports fall back to `Fiber.current` without raising when
2641
- no Async task is active.
2642
-
2643
- * **Serialize `LLM::Object` values correctly through `LLM.json`** <br>
2644
- Make `LLM::Object#to_json` call `LLM.json.dump(to_h, ...)` so
2645
- `LLM::Object` values serialize through the llm.rb JSON adapter.
2646
-
2647
- ## v7.0.0
2648
-
2649
- Changes since `v6.1.0`.
2650
-
2651
- This release turns agent tool-loop limit errors into in-band advisory
2652
- returns so the LLM can react to rate limits and continue the loop. It
2653
- adds `tool_attempts: nil` as a way to opt out of advisory tool-limit
2654
- returns entirely, and fixes the default provider HTTP path to keep
2655
- `net-http-persistent` optional when not explicitly enabled.
2656
-
2657
- ### Breaking
2658
-
2659
- * **Return in-band tool-loop limit errors from agents** <br>
2660
- Stop raising `LLM::ToolLoopError` when an agent exhausts its tool loop
2661
- attempt budget, and instead send advisory `LLM::Function::Return`
2662
- errors back through the model so the LLM can react to the rate limit
2663
- in-band and continue the loop.
2664
-
2665
- * **Allow `tool_attempts: nil` to disable advisory tool-limit returns** <br>
2666
- Keep the default `tool_attempts` budget at `25`, but treat an explicit
2667
- `tool_attempts: nil` as an opt-out that disables advisory tool-limit
2668
- returns entirely.
2669
-
2670
- ### Fix
2671
-
2672
- * **Keep `net-http-persistent` optional on normal HTTP requests** <br>
2673
- Stop the default provider HTTP path from loading `net/http/persistent`
2674
- unless persistent transport support is explicitly enabled.
2675
-
2676
- ## v6.1.0
2677
-
2678
- Changes since `v6.0.0`.
2679
-
2680
- This release tightens interrupt and compaction behavior for long-running
2681
- contexts. It adds `LLM::Buffer#rindex`, supports percentage-based token
2682
- thresholds in `LLM::Compactor`, tracks persisted compaction state through
2683
- context serialization, reliably interrupts Async-backed requests, preserves
2684
- valid tool-call history on cancellation, keeps concurrent skill tool loops
2685
- running on streamed agents, and returns zero-valued usage objects when no
2686
- provider usage has been recorded yet.
2687
-
2688
- ### Change
2689
-
2690
- * **Add `LLM::Buffer#rindex`** <br>
2691
- Add `LLM::Buffer#rindex` as a direct forward to the underlying message
2692
- array so callers can find the last matching message index through the
2693
- buffer API.
2694
-
2695
- * **Support percentage compaction token thresholds** <br>
2696
- Let `LLM::Compactor` accept `token_threshold:` values like `"90%"` so
2697
- compaction can trigger at a percentage of the active model context
2698
- window.
2699
-
2700
- ### Fix
2701
-
2702
- * **Interrupt Async-backed requests reliably** <br>
2703
- Track request ownership through the provider transport so contexts use
2704
- the active Async task when available, letting `ctx.interrupt!`
2705
- reliably cancel streamed requests under Async runtimes and surface
2706
- them as `LLM::Interrupt`.
2707
-
2708
- * **Preserve valid tool-call history on cancellation** <br>
2709
- Append cancelled tool-return messages for unresolved tool calls during
2710
- `ctx.interrupt!` so follow-up provider requests do not fail with
2711
- invalid tool-call history after pending tool work is cancelled.
2712
-
2713
- * **Preserve concurrent skill tool loops on streamed agents** <br>
2714
- Propagate the active agent concurrency through the effective request
2715
- stream so nested skill agents keep using queued `wait(...)` tool
2716
- execution instead of falling back to direct `:call` execution.
2717
-
2718
- * **Track persisted compaction state on contexts** <br>
2719
- Mark contexts as compacted after `LLM::Compactor#compact!`, persist and
2720
- restore that state through context serialization, and clear it after the
2721
- next successful model response.
2722
-
2723
- * **Return zero-valued usage objects from contexts** <br>
2724
- Make `LLM::Context#usage` consistently return an `LLM::Object`, using a
2725
- zero-valued usage object when no provider usage has been recorded yet.
2726
-
2727
- ## v6.0.0
2728
-
2729
- Changes since `v5.4.0`.
2730
-
2731
- This release simplifies the ORM persistence contract around serialized
2732
- `data` state, removing the assumption of reserved `provider`, `model`, and
2733
- usage columns. Provider selection must now come from `provider:` hooks,
2734
- model defaults come from `context:` or agent DSL, and usage is read from the
2735
- serialized runtime state. Alongside this breaking change, Sequel JSON and
2736
- JSONB persistence is fixed, ractor-backed tools now fire tracer callbacks,
2737
- and `LLM::RactorError` is raised for unsupported ractor tool work.
2738
-
2739
- ### Change
2740
-
2741
- * **Simplify ORM persistence to serialized `data` state** <br>
2742
- Change the built-in ActiveRecord and Sequel wrappers to treat serialized
2743
- `data` as the persistence contract, instead of assuming reserved
2744
- `provider`, `model`, and usage columns. Provider selection must now come
2745
- from `provider:` hooks that resolve a real `LLM::Provider` instance, model
2746
- defaults come from `context:` or agent DSL, and `usage` is read from the
2747
- serialized runtime state.
2748
-
2749
- ### Fix
2750
-
2751
- * **Fix Sequel JSON and JSONB persistence** <br>
2752
- Load Sequel PostgreSQL JSON support when `plugin :llm` is configured with
2753
- `format: :json` or `:jsonb`, and wrap structured payloads correctly so
2754
- persisted context state can be stored in PostgreSQL JSON columns.
2755
-
2756
- * **Trace ractor-backed tool callbacks** <br>
2757
- Make tool tracers fire `on_tool_start` and `on_tool_finish` for
2758
- class-based `:ractor` execution too, so ractor-backed tool calls show up
2759
- in tracer callbacks like the other concurrent tool paths.
2760
-
2761
- * **Raise `LLM::RactorError` for unsupported ractor tool work** <br>
2762
- Add `LLM::RactorError` and fail fast when `:ractor` execution is requested
2763
- for unsupported tool types such as skill-backed tools, instead of letting
2764
- deeper Ruby isolation errors leak out later in execution.
2765
-
2766
- * **Delegate interrupt to concurrent task implementations** <br>
2767
- Make `LLM::Function::Task#interrupt!` delegate to the underlying fork or
2768
- ractor task when it supports interruption, so `ctx.interrupt!` and
2769
- `task.interrupt!` work correctly for fork- and ractor-backed tool
2770
- execution.
2771
-
2772
- ## v5.4.0
2773
-
2774
- Changes since `v5.3.0`.
2775
-
2776
- This release expands tracer support around agentic execution. It lets
2777
- `LLM::Agent` define scoped tracers through the agent DSL and fixes concurrent
2778
- tool execution so those scoped tracers stay attached when work crosses
2779
- thread, task, fiber, and skill boundaries.
2780
-
2781
- ### Change
2782
-
2783
- * **Add agent-scoped tracers** <br>
2784
- Let `LLM::Agent` classes define `tracer ...` or `tracer { ... }` so an
2785
- agent can carry its own tracer without replacing the provider's default
2786
- tracer. The resolved tracer is scoped to that agent's turns, tool loops,
2787
- and pending tool access. Available through the `acts_as_agent` and Sequel
2788
- agent plugin `tracer` DSL too.
2789
-
2790
- ### Fix
2791
-
2792
- * **Preserve scoped tracers across concurrent tool work** <br>
2793
- Keep agent- and request-scoped tracers attached when tool execution
2794
- crosses `:thread`, `:task`, or `:fiber` boundaries, including skill
2795
- execution, so spawned work does not fall back to the provider default
2796
- tracer.
2797
-
2798
- ## v5.3.0
2799
-
2800
- Changes since `v5.2.1`.
2801
-
2802
- This release deepens llm.rb's request-rewriting and tool-definition surface.
2803
- It adds transformer lifecycle hooks to `LLM::Stream` so UIs can surface work
2804
- like PII scrubbing before a request is sent, and it adds a more explicit
2805
- OmniAI-style tool DSL form with `parameter` plus separate `required`
2806
- declarations while keeping the older `param ... required: true` style working.
2807
-
2808
- ### Change
2809
-
2810
- * **Add transformer stream lifecycle hooks** <br>
2811
- Add `on_transform` and `on_transform_finish` to
2812
- `LLM::Stream` so UIs can surface request rewriting work such as PII
2813
- scrubbing before a request is sent to the model.
2814
-
2815
- * **Add a separate `required` tool DSL form** <br>
2816
- Add `parameter` as an alias of `param` and support `required %i[...]`
2817
- as a separate declaration, inspired by OmniAI-style tools, while keeping
2818
- the existing `param ... required: true` form working too.
2819
-
2820
- ## v5.2.1
2821
-
2822
- Changes since `v5.2.0`.
2823
-
2824
- This release tightens the streamed queue fix from `v5.2.0` for concurrent
2825
- workloads. Request-local streams now stay bound long enough for `wait` to
2826
- drain queued work and then clear cleanly so later waits fall back to the
2827
- context's configured stream.
2828
-
2829
- ### Fix
2830
-
2831
- * **Reset request-local streams after `wait` drains queued work** <br>
2832
- Keep per-call `stream:` bindings alive through `LLM::Context#wait` so
2833
- queued streamed tool work still resolves correctly, then clear the
2834
- request-local stream after the wait completes to avoid leaking it into
2835
- later turns.
2836
-
2837
- ## v5.2.0
2838
-
2839
- Changes since `v5.1.0`.
2840
-
2841
- This release adds current DeepSeek V4 support through refreshed provider
2842
- metadata, including `deepseek-v4-flash` and `deepseek-v4-pro`, while fixing
2843
- request-local queue handling for concurrent streamed workloads so `wait` and
2844
- interruption use the active per-call stream correctly.
2845
-
2846
- ### Change
2847
-
2848
- * **Add `LLM::MCP#run` for scoped MCP client lifecycle** <br>
2849
- Add `LLM::MCP#run` so MCP clients can be started for the duration of a
2850
- block and then stopped automatically, which simplifies the usual
2851
- `start`/`stop` pattern in examples and application code.
2852
-
2853
- * **Refresh provider model metadata** <br>
2854
- Add current DeepSeek and OpenAI model metadata to `data/` and update the
2855
- Google Gemini model entry to match the current provider naming.
2856
-
2857
- ### Fix
2858
-
2859
- * **Reject unsupported DeepSeek multimodal prompt objects early** <br>
2860
- Raise `LLM::PromptError` for `image_url`, `local_file`, and
2861
- `remote_file` in DeepSeek chat requests instead of sending invalid
2862
- OpenAI-compatible payloads that the provider rejects at runtime.
2863
-
2864
- * **Preserve DeepSeek reasoning content across tool turns** <br>
2865
- Replay `reasoning_content` when serializing prior assistant messages for
2866
- DeepSeek chat completions, so thinking-mode tool calls can continue into
2867
- follow-up requests without triggering invalid request errors.
2868
-
2869
- * **Default DeepSeek to `deepseek-v4-flash`** <br>
2870
- Change `LLM::DeepSeek#default_model` to `deepseek-v4-flash` so new
2871
- contexts and default provider usage align with the current preferred chat
2872
- model.
2873
-
2874
- * **Use per-call streams when waiting on streamed tool work** <br>
2875
- Track request-local streams bound through `talk(..., stream:)` and
2876
- `respond(..., stream:)` so `LLM::Context#wait` and interruption-aware
2877
- queue handling use the active stream instead of falling back to pending
2878
- function spawning.
2879
-
2880
- ## v5.1.0
2881
-
2882
- Changes since `v5.0.0`.
2883
-
2884
- This release tightens streamed tool execution around the actual request-local
2885
- runtime state. It fixes streamed resolution of per-request tools and makes
2886
- that streamed path work cleanly with `LLM.function(...)`, MCP tools, bound
2887
- tool instances, and normal tool classes.
2888
-
2889
- ### Fix
2890
-
2891
- * **Resolve request-local tools during streaming** <br>
2892
- Resolve streamed tool calls through `LLM::Stream` request-local tools
2893
- before falling back to the global registry, so per-request tools and bound
2894
- tool instances work correctly during streaming.
2895
-
2896
- * **Support `LLM.function(...)` and MCP tools in streamed tool resolution** <br>
2897
- Let streamed tool resolution use the current request tool set, so
2898
- `LLM.function(...)`, MCP tools, bound tool instances, and normal
2899
- `LLM::Tool` classes all work through the same streamed tool path.
2900
-
2901
- ## v5.0.0
2902
-
2903
- Changes since `v4.23.0`.
2904
-
2905
- This release expands llm.rb from an execution runtime into a more explicit
2906
- supervision and transformation runtime. It adds context-level guards,
2907
- transformers, and loop supervision through `LLM::LoopGuard`, while deepening
2908
- long-lived context behavior through compaction, interruption hooks, and
2909
- streamed `ctx.spawn(...)` tool execution.
2910
-
2911
- ### Change
2912
-
2913
- * **Make compactor thresholds explicit** <br>
2914
- Require `message_threshold:` and `token_threshold:` to be opted into
2915
- explicitly, so `LLM::Compactor` only compacts automatically when one of
2916
- those thresholds is configured. Context-window-derived token limits can be
2917
- computed by the caller when needed.
2918
-
2919
- * **Allow assigning a compactor through `LLM::Context`** <br>
2920
- Let `LLM::Context` accept `ctx.compactor = ...` in addition to the
2921
- constructor `compactor:` option, so compactor config can be assigned or
2922
- replaced after context initialization.
2923
-
2924
- * **Mark compaction summaries in message metadata** <br>
2925
- Mark compaction summaries with `extra[:compaction]` and
2926
- `LLM::Message#compaction?`, so applications can detect or hide synthetic
2927
- summary messages in conversation history.
2928
-
2929
- * **Add cooperative tool interruption hooks** <br>
2930
- Let `ctx.interrupt!` notify queued tool work through `on_interrupt`, so
2931
- running tools can clean up cooperatively when a context is cancelled.
2932
-
2933
- * **Add `LLM::Context` guards** <br>
2934
- Add a new `guard` capability to `LLM::Context` so execution can be
2935
- supervised at the runtime level. The built-in `LLM::LoopGuard` detects
2936
- repeated tool-call patterns and stops stuck agentic loops through in-band
2937
- `LLM::GuardError` returns. `LLM::Agent` enables this guard by default.
2938
-
2939
- * **Add `LLM::Context` transformers** <br>
2940
- Add a new `transformer` capability to `LLM::Context` so prompts and params
2941
- can be rewritten before provider requests are sent. This makes it possible
2942
- to apply context-wide behaviors such as PII scrubbing or request-level
2943
- param injection without rewriting every `talk` and `respond` call site.
2944
-
2945
- ## v4.23.0
2946
-
2947
- Changes since `v4.22.0`.
2948
-
2949
- This release expands llm.rb's runtime surface for long-lived contexts and
2950
- stateful tools. It adds built-in context compaction through `LLM::Compactor`,
2951
- lets explicit `tools:` arrays accept bound `LLM::Tool` instances, and fixes
2952
- OpenAI-compatible no-arg tool schemas for stricter providers such as xAI.
2953
-
2954
- ### Change
2955
-
2956
- * **Add `LLM::Compactor` for long-lived contexts** <br>
2957
- Add built-in context compaction through `LLM::Compactor`, so older history
2958
- can be summarized, retained windows can stay bounded, compaction can run on
2959
- its own `model:`, thresholds can be configured explicitly, and
2960
- `LLM::Stream` can observe the lifecycle through `on_compaction` and
2961
- `on_compaction_finish`.
2962
-
2963
- * **Allow bound tool instances in explicit tool lists** <br>
2964
- Let explicit `tools:` arrays accept `LLM::Tool` instances such as
2965
- `MyTool.new(foo: 1)`, so tools can carry bound state without changing the
2966
- global tool registry model.
2967
-
2968
- ### Fix
2969
-
2970
- * **Fix xAI/OpenAI-compatible no-arg tool schemas** <br>
2971
- Send an empty object schema for tools without declared parameters instead
2972
- of `null`, so stricter providers such as xAI accept mixed tool sets that
2973
- include no-arg tools.
2974
-
2975
- ## v4.22.0
2976
-
2977
- Changes since `v4.21.0`.
2978
-
2979
- This release deepens the runtime shape of llm.rb. It reduces helper-method
2980
- surface on persisted ORM models, expands real ORM coverage, and makes skills
2981
- behave more like bounded sub-agents with inherited recent context and proper
2982
- instruction injection.
2983
-
2984
- ### Change
2985
-
2986
- * **Reduce ActiveRecord wrapper model surface** <br>
2987
- Move helper methods such as option resolution, column mapping,
2988
- serialization, and persistence into `Utils` for the ActiveRecord
2989
- wrappers so wrapped models include fewer internal helper methods.
2990
-
2991
- * **Reduce Sequel wrapper model surface** <br>
2992
- Move helper methods such as option resolution, column mapping,
2993
- serialization, and persistence into `Utils` for the Sequel wrappers
2994
- so wrapped models include fewer internal helper methods.
2995
-
2996
- * **Expand ORM integration coverage** <br>
2997
- Add broader ActiveRecord and Sequel coverage for persisted context and
2998
- agent wrappers, including real SQLite-backed records and cassette-backed
2999
- OpenAI persistence paths.
3000
-
3001
- * **Make skills inherit recent parent context** <br>
3002
- Run `LLM::Skill` with a curated slice of recent parent user and assistant
3003
- messages, prefixed with `Recent context:`, so skills behave more like
3004
- task-scoped sub-agents instead of instruction-only helpers.
3005
-
3006
- ### Fix
3007
-
3008
- * **Fix Sequel `plugin :agent` load order** <br>
3009
- Require the shared Sequel plugin support from `LLM::Sequel::Agent` so
3010
- `plugin :agent` can load independently without raising
3011
- `uninitialized constant LLM::Sequel::Plugin`.
3012
-
3013
- * **Make skill execution inherit parent context request settings** <br>
3014
- Run `LLM::Skill` through a parent `LLM::Context` instead of a bare
3015
- provider so nested skill agents inherit context-level settings such as
3016
- `mode: :responses`, `store: false`, streaming, and other request defaults,
3017
- while still keeping skill-local tools and avoiding parent schemas.
3018
-
3019
- * **Keep agent instructions when history is preseeded** <br>
3020
- Inject `LLM::Agent` instructions once unless a system message is already
3021
- present, so agents and nested skills still get their instructions when
3022
- they start with inherited non-system context.
3023
-
3024
- ## v4.21.0
3025
-
3026
- Changes since `v4.20.2`.
3027
-
3028
- This release expands higher-level composition in llm.rb. It adds Sequel agent
3029
- persistence through `plugin :agent` and introduces directory-backed skills
3030
- that load from `SKILL.md`, resolve named tools, and plug directly into
3031
- `LLM::Context` and `LLM::Agent`.
3032
-
3033
- ### Change
3034
-
3035
- * **Add `plugin :agent` for Sequel models** <br>
3036
- Add Sequel support for `plugin :agent`, similar to ActiveRecord's
3037
- `acts_as_agent`, so models can wrap `LLM::Agent` with built-in
3038
- persistence.
3039
-
3040
- * **Load directory-backed skills through `LLM::Context` and `LLM::Agent`** <br>
3041
- Add `skills:` to `LLM::Context` and `skills ...` to `LLM::Agent` so
3042
- directories with `SKILL.md` can be loaded, resolved into tools, and run
3043
- through the normal llm.rb tool path.
3044
-
3045
- ## v4.20.2
3046
-
3047
- Changes since `v4.20.1`.
3048
-
3049
- This patch release improves runtime behavior around interruption and mixed
3050
- concurrency waits. It also rounds out response API uniformity for Google
3051
- completion responses.
3052
-
3053
- ### Fix
3054
-
3055
- * **Expose Google completion response IDs through `.id`** <br>
3056
- Add `LLM::Response#id` support to Google completion responses so tracer
3057
- and caller code can rely on the same API used by other providers.
3058
-
3059
- * **Track interrupt ownership on the active request** <br>
3060
- Bind `LLM::Context` interruption to the fiber running `talk` or `respond`
3061
- so `interrupt!` works correctly when requests are started outside the
3062
- context's initialization fiber.
3063
-
3064
- ### Change
3065
-
3066
- * **Allow mixed concurrency strategies in `wait(...)`** <br>
3067
- Let `LLM::Context#wait`, `LLM::Stream#wait`, and `LLM::Agent.concurrency`
3068
- accept arrays such as `[:thread, :ractor]` so mixed tool sets can wait on
3069
- more than one concurrency strategy.
3070
-
3071
- ## v4.20.1
3072
-
3073
- Changes since `v4.20.0`.
3074
-
3075
- This patch release fixes ORM option resolution in the Sequel and
3076
- ActiveRecord wrappers. Symbol-based `provider:` and `context:` hooks now
3077
- resolve correctly, and internal default option constants are referenced
3078
- explicitly instead of relying on nested constant lookup.
3079
-
3080
- ### Fix
3081
-
3082
- * **Fix symbol-based ORM option hooks for provider and context hashes** <br>
3083
- Make `provider:` and `context:` resolve symbol hooks through the model in
3084
- the Sequel plugin and ActiveRecord wrappers instead of falling back to an
3085
- empty hash.
3086
-
3087
- * **Fix ORM wrapper constant lookup for option defaults** <br>
3088
- Qualify internal `EMPTY_HASH` / `DEFAULTS` references in the Sequel plugin
3089
- and ActiveRecord wrappers so option resolution does not depend on nested
3090
- constant lookup quirks.
3091
-
3092
- ## v4.20.0
3093
-
3094
- Changes since `v4.19.0`.
3095
-
3096
- This release adds better support for tagged prompt content. `LLM::Context`
3097
- can now serialize and restore `image_url`, `local_file`, and `remote_file`
3098
- content cleanly, and `LLM::Message` now exposes helpers for inspecting
3099
- tagged image and file attachments.
3100
-
3101
- ### Change
3102
-
3103
- * **Round-trip tagged prompt objects through `LLM::Context`** <br>
3104
- Teach `LLM::Context` serialization and restore to preserve
3105
- `image_url`, `local_file`, and `remote_file` content across
3106
- `to_json` / `restore`.
3107
-
3108
- * **Add attachment helpers to `LLM::Message`** <br>
3109
- Add `image_url?`, `image_urls`, `file?`, and `files` so callers can
3110
- inspect messages for tagged image and file content more directly.
3111
-
3112
- ## v4.19.0
3113
-
3114
- Changes since `v4.18.0`.
3115
-
3116
- This release tightens the ActiveRecord and ORM integration layer. It adds
3117
- inline agent DSL blocks to `acts_as_agent` so agent defaults can be defined
3118
- where the wrapper is declared, and it exposes the resolved provider through
3119
- public `llm` methods on the ActiveRecord and Sequel wrappers.
3120
-
3121
- ### Change
3122
-
3123
- * **Make ORM provider access public through `llm`** <br>
3124
- Expose the resolved provider on the Sequel plugin and the ActiveRecord
3125
- `acts_as_llm` / `acts_as_agent` wrappers through a public `llm` method.
3126
-
3127
- * **Allow inline agent DSL blocks in `acts_as_agent`** <br>
3128
- Let ActiveRecord models configure `model`, `tools`, `schema`,
3129
- `instructions`, and `concurrency` directly inside the `acts_as_agent`
3130
- declaration block.
3131
-
3132
- ## v4.18.0
3133
-
3134
- Changes since `v4.17.0`.
3135
-
3136
- This release improves tracing and tool execution behavior across llm.rb.
3137
- It makes provider tracers default to the provider instance, adds
3138
- `LLM::Provider#with_tracer` for scoped overrides, restores tool tracing for
3139
- concurrent and streamed tool execution, extends streamed tracing to MCP tools,
3140
- and adds symbol-based ORM option hooks alongside experimental ractor tool
3141
- concurrency.
3142
-
3143
- ### Change
3144
-
3145
- * **Make provider tracers default to the provider instance** <br>
3146
- Change `llm.tracer = ...` so it sets a provider default tracer instead of
3147
- relying on scoped fiber-local state alone. This makes tracer configuration
3148
- behave more predictably across normal tasks, threads, and fibers that share
3149
- the same provider instance.
3150
-
3151
- * **Add `LLM::Provider#with_tracer` for scoped overrides** <br>
3152
- Add `with_tracer` as the opt-in escape hatch for request- or turn-scoped
3153
- tracer overrides. Use it when you want temporary tracing on the current
3154
- fiber without replacing the provider's default tracer.
3155
-
3156
- * **Trace concurrent tool calls outside ractors** <br>
3157
- Make tool tracing fire correctly when functions run through `:thread`,
3158
- `:task`, or `:fiber` concurrency. Experimental `:ractor` execution still
3159
- does not emit tool tracer events.
3160
-
3161
- * **Trace streamed tool calls, including MCP tools** <br>
3162
- Bind stream metadata through `LLM::Stream#extra` so streamed tool calls
3163
- inherit tracer and model context before they are handed to `on_tool_call`.
3164
- This restores tool tracing for streamed MCP and local tool execution.
3165
-
3166
- * **Support symbol-based ORM option hooks** <br>
3167
- Let `provider:`, `context:`, and `tracer:` on the Sequel plugin and
3168
- the ActiveRecord `acts_as_llm` / `acts_as_agent` wrappers resolve through
3169
- model method names as well as procs.
3170
-
3171
- * **Add experimental ractor tool concurrency** <br>
3172
- Add `:ractor` support to `LLM::Function#spawn`, `LLM::Function::Array#wait`,
3173
- `LLM::Stream#wait`, and `LLM::Agent.concurrency` so class-based tools with
3174
- ractor-safe arguments and return values can run in Ruby ractors and report
3175
- their results back into the normal LLM tool-return path. MCP tools are not
3176
- supported by the current `:ractor` mode, but mixed workloads can still
3177
- branch on `tool.mcp?` and choose a supported strategy per tool. `:ractor`
3178
- is especially useful for CPU-bound tools, while `:task`, `:fiber`, or
3179
- `:thread` may be a better fit for I/O-bound work.
3180
-
3181
- ## v4.17.0
3182
-
3183
- Changes since `v4.16.1`.
3184
-
3185
- This release expands agent support across llm.rb. It brings `LLM::Agent`
3186
- closer to `LLM::Context`, adds configurable automatic tool concurrency
3187
- including experimental ractor support for class-based tools,
3188
- extends persisted ORM wrappers with more of the context runtime surface and
3189
- tracer hooks, and introduces built-in ActiveRecord agent persistence through
3190
- `acts_as_agent`.
3191
-
3192
- ### Change
3193
-
3194
- * **Add configurable tool concurrency to `LLM::Agent`** <br>
3195
- Add the class-level `concurrency` DSL to `LLM::Agent` so automatic
3196
- tool loops can run with `:call`, `:thread`, `:task`, `:fiber`, or
3197
- experimental `:ractor` support for class-based tools instead of
3198
- always executing sequentially.
3199
-
3200
- * **Bring `LLM::Agent` closer to `LLM::Context`** <br>
3201
- Expand `LLM::Agent` so it exposes more of the same runtime surface as
3202
- `LLM::Context`, including returns, interruption, mode, cost, context
3203
- window, structured serialization, and other context-backed helpers,
3204
- while still auto-managing tool loops.
3205
-
3206
- * **Refresh agent docs and coverage** <br>
3207
- Update the README and deep dive to explain the current role of
3208
- `LLM::Agent`, add examples that show automatic tool execution and
3209
- concurrency, and add focused specs for the expanded agent surface and
3210
- tool-loop behavior.
3211
-
3212
- * **Add ORM tracer hooks for persisted contexts** <br>
3213
- Add `tracer:` to both the Sequel plugin and `acts_as_llm` so models
3214
- can resolve and assign tracers onto the provider used by their persisted
3215
- `LLM::Context`.
3216
-
3217
- * **Bring persisted ORM wrappers closer to `LLM::Context`** <br>
3218
- Expand both the Sequel plugin and `acts_as_llm` so record-backed
3219
- contexts expose more of the same runtime surface as `LLM::Context`,
3220
- including mode, returns, interruption, prompt helpers, file helpers,
3221
- and tracer access.
3222
-
3223
- * **Add ActiveRecord agent persistence with `acts_as_agent`** <br>
3224
- Add `acts_as_agent` for ActiveRecord models that should wrap
3225
- `LLM::Agent`, reusing the same record-backed runtime shape as
3226
- `acts_as_llm` while letting tool execution be managed by the agent.
3227
-
3228
- ## v4.16.1
3229
-
3230
- Changes since `v4.16.0`.
3231
-
3232
- This release tightens ORM persistence by removing an unnecessary JSON
3233
- round-trip when restoring structured `:json` and `:jsonb` context
3234
- payloads.
3235
-
3236
- ### Change
3237
-
3238
- * **Restore structured ORM payloads directly** <br>
3239
- Teach `LLM::Context#restore` to accept parsed data payloads and use
3240
- that path from the ActiveRecord and Sequel persistence wrappers for
3241
- `format: :json` and `:jsonb`, avoiding a redundant
3242
- `Hash -> JSON string -> Hash` round-trip on restore.
3243
-
3244
- ## v4.16.0
3245
-
3246
- Changes since `v4.15.0`.
3247
-
3248
- This release expands ORM support with built-in ActiveRecord persistence
3249
- and improves compatibility with OpenAI-compatible gateways, proxies, and
3250
- self-hosted servers that use non-standard API root paths.
3251
-
3252
- ### Change
3253
-
3254
- * **Support OpenAI-compatible base paths** <br>
3255
- Add `base_path:` to provider configuration so OpenAI-compatible
3256
- endpoints can vary both host and API prefix. This supports providers,
3257
- proxies, and gateways that keep OpenAI request shapes but use
3258
- non-standard URL layouts such as DeepInfra's `/v1/openai/...`.
3259
-
3260
- * **Add ActiveRecord context persistence with `acts_as_llm`** <br>
3261
- Add a built-in ActiveRecord wrapper that mirrors the Sequel plugin
3262
- API so applications can persist `LLM::Context` state on records with
3263
- default columns, provider/context hooks, validation-backed writes,
3264
- and `format: :string`, `:json`, or `:jsonb` storage.
3265
-
3266
- ## v4.15.0
3267
-
3268
- Changes since `v4.14.0`.
3269
-
3270
- ### Change
3271
-
3272
- * **Reduce OpenAI stream parser merge overhead** <br>
3273
- Special-case the most common single-field deltas, streamline
3274
- incremental tool-call merging, and avoid repeated JSON parse attempts
3275
- until streamed tool arguments look complete.
3276
-
3277
- * **Cache streaming callback capabilities in parsers** <br>
3278
- Cache callback support checks once at parser initialization time in
3279
- the OpenAI, OpenAI Responses, Anthropic, Google, and Ollama stream
3280
- parsers instead of repeating `respond_to?` checks on hot streaming
3281
- paths.
3282
-
3283
- * **Reduce OpenAI Responses parser lookup overhead** <br>
3284
- Special-case the hot Responses API event paths and cache the current
3285
- output item and content part so streamed output text deltas do less
3286
- repeated nested lookup work.
3287
-
3288
- * **Add a Sequel context persistence plugin** <br>
3289
- Add `plugin :llm` for Sequel models so apps can persist
3290
- `LLM::Context` state with default columns and pass provider setup
3291
- through `provider:` when needed. The plugin now also supports
3292
- `format: :string`, `:json`, or `:jsonb` for text and native JSON
3293
- storage when Sequel JSON typecasting is enabled.
3294
-
3295
- * **Improve streaming parser performance** <br>
3296
- In the local replay-based `stream_parser` benchmark versus `v4.14.0`
3297
- (median of 20 samples, 5000 iterations), plain Ruby is a
3298
- small overall win: the generic eventstream path is about 0.4%
3299
- faster, the OpenAI stream parser is about 0.5% faster, and the
3300
- OpenAI Responses parser is about 1.6% faster, with unchanged
3301
- allocations. Under YJIT on the same benchmark harness, the generic
3302
- eventstream path is about 0.9% faster and the OpenAI stream parser
3303
- is about 0.4% faster, while the OpenAI Responses parser is about
3304
- 0.7% slower, also with unchanged allocations.
3305
-
3306
- Compared to `v4.13.0`, the larger `v4.14.0` streaming gains still
3307
- hold. The generic eventstream path remains dramatically faster than
3308
- `v4.13.0`, the OpenAI stream parser remains modestly faster, and the
3309
- OpenAI Responses parser is roughly flat to slightly better depending
3310
- on runtime. In other words, current keeps the large eventstream win
3311
- from `v4.14.0`, adds only small incremental changes beyond that, and
3312
- does not turn the post-`v4.14.0` parser work into another large
3313
- benchmark jump.
3314
-
3315
- ## v4.14.0
3316
-
3317
- Changes since `v4.13.0`.
3318
-
3319
- This release adds request interruption for contexts, reworks provider
3320
- HTTP internals for lower-overhead streaming, and fixes MCP clients so
3321
- parallel tool calls can safely share one connection.
3322
-
3323
- ### Add
3324
-
3325
- * **Add request interruption support** <br>
3326
- Add `LLM::Context#interrupt!`, `LLM::Context#cancel!`, and
3327
- `LLM::Interrupt` for interrupting in-flight provider requests,
3328
- inspired by Go's context cancellation.
3329
-
3330
- ### Change
3331
-
3332
- * **Rework provider HTTP transport internals** <br>
3333
- Rework provider HTTP around `LLM::Provider::Transport::HTTP` with
3334
- explicit transient and persistent transport handling.
3335
-
3336
- * **Reduce SSE parser overhead** <br>
3337
- Dispatch raw parsed values to registered visitors instead of building
3338
- an `Event` object for every streamed line.
3339
-
3340
- * **Reduce provider streaming allocations** <br>
3341
- Decode streamed provider payloads directly in
3342
- `LLM::Provider::Transport::HTTP` before handing them to provider
3343
- parsers, which cuts allocation churn and gives a small streaming
3344
- speed bump.
3345
-
3346
- * **Reduce generic SSE parser allocations** <br>
3347
- Keep unread event-stream buffer data in place until compaction is
3348
- worthwhile, which lowers allocation churn in the remaining generic
3349
- SSE path.
3350
-
3351
- * **Improve streaming parser performance** <br>
3352
- In the local replay-based `stream_parser` benchmark versus `v4.13.0`
3353
- (median of 20 samples, 5000 iterations):
3354
- Plain Ruby: the generic eventstream path is about 53% faster with
3355
- about 32% fewer allocations, the OpenAI stream parser is about 11%
3356
- faster with about 4% fewer allocations, and the OpenAI Responses
3357
- parser is about 3% faster with unchanged allocations.
3358
- YJIT on the current parser benchmark harness: the current tree is
3359
- about 26% faster than non-YJIT on the generic eventstream path,
3360
- about 18% faster on the OpenAI stream parser, and about 16% faster
3361
- on the OpenAI Responses parser, with allocations unchanged.
3362
-
3363
- ### Fix
3364
-
3365
- * **Support parallel MCP tool calls on one client** <br>
3366
- Route MCP responses by JSON-RPC id so concurrent tool calls can
3367
- share one client and transport without mismatching replies.
3368
-
3369
- * **Use explicit MCP non-blocking read errors** <br>
3370
- Use `IO::EAGAINWaitReadable` while continuing to retry on
3371
- `IO::WaitReadable`.
3372
-
3373
- ## v4.13.0
3374
-
3375
- Changes since `v4.12.0`.
3376
-
3377
- This release expands MCP prompt support, improves reasoning support in the
3378
- OpenAI Responses API, and refreshes the docs around llm.rb's runtime model,
3379
- contexts, and advanced workflows.
3380
-
3381
- ### Add
3382
-
3383
- - Add `LLM::MCP#prompts` and `LLM::MCP#find_prompt` for MCP prompt support.
3384
-
3385
- ### Change
3386
-
3387
- - Rework the README around llm.rb as a runtime for AI systems.
3388
- - Add a dedicated deep dive guide for providers, contexts, persistence,
3389
- tools, agents, MCP, tracing, multimodal prompts, and retrieval.
3390
-
3391
- ### Fix
3392
-
3393
- All of these fixes apply to MCP:
3394
-
3395
- - fix(mcp): raise `LLM::MCP::MismatchError` on mismatched response ids.
3396
- - fix(mcp): normalize prompt message content while preserving the original payload.
3397
-
3398
- All of these fixes apply to OpenAI's Responses API:
3399
-
3400
- - fix(openai): emit `on_reasoning_content` for streamed reasoning summaries.
3401
- - fix(openai): skip `previous_response_id` on `store: false` follow-up calls.
3402
- - fix(openai): fall back to an empty object schema for tools without params.
3403
- - fix(openai): preserve original tool-call payloads on re-sent assistant tool messages.
3404
- - fix(openai): emit `output_text` for assistant-authored response content.
3405
- - fix(openai): return `nil` for `system_fingerprint` on normalized response objects.
3406
-
3407
- ## v4.12.0
3408
-
3409
- Changes since `v4.11.1`.
3410
-
3411
- This release expands advanced streaming and MCP execution while reframing
3412
- llm.rb more clearly as a system integration layer for LLMs, tools, MCP
3413
- sources, and application APIs.
3414
-
3415
- ### Add
3416
-
3417
- - Add `persistent` as an alias for `persist!` on providers and MCP transports.
3418
- - Add `LLM::Stream#on_tool_return` for observing completed streamed tool work.
3419
- - Add `LLM::Function::Return#error?`.
3420
-
3421
- ### Change
3422
-
3423
- - Expect advanced streaming callbacks to use `LLM::Stream` subclasses
3424
- instead of duck-typing them onto arbitrary objects. Basic `#<<`
3425
- streaming remains supported.
3426
-
3427
- ### Fix
3428
-
3429
- - Fix Anthropic tools without params by always emitting `input_schema`.
3430
- - Fix Anthropic tool-only responses to still produce an assistant message.
3431
- - Fix Anthropic tool results to use the `user` role.
3432
- - Fix Anthropic tool input normalization.
3433
-
3434
- ## v4.11.1
3435
-
3436
- Changes since `v4.11.0`.
3437
-
3438
- ### Fix
3439
-
3440
- * Cast OpenTelemetry tool-related values to strings. <br>
3441
- Otherwise they're rejected by opentelemetry-sdk as invalid attributes.
3442
-
3443
- ## v4.11.0
3444
-
3445
- Changes since `v4.10.0`.
3446
-
3447
- ### Add
3448
-
3449
- - Add `LLM::Stream` for richer streaming callbacks, including `on_content`,
3450
- `on_reasoning_content`, and `on_tool_call` for concurrent tool execution.
3451
- - Add `LLM::Stream#wait` as a shortcut for `queue.wait`.
3452
- - Add `LLM::Context#wait` as a shortcut for the configured stream's `wait`.
3453
- - Add `LLM::Context#call(:functions)` as a shortcut for `functions.call`.
3454
- - Add `LLM::Function.registry` and enhanced support for MCP tools in
3455
- `LLM::Tool.registry` for tool resolution during streaming.
3456
- - Add normalized `LLM::Response` for OpenAI Responses, providing `content`,
3457
- `content!`, `messages` / `choices`, `usage`, and `reasoning_content`.
3458
- - Add `mode: :responses` to `LLM::Context` for routing `talk` through the
3459
- Responses API.
3460
- - Add `LLM::Context#returns` for collecting pending tool returns from the context.
3461
- - Add persistent HTTP connection pooling for repeated MCP tool calls via
3462
- `LLM.mcp(http: ...).persist!`.
3463
- - Add explicit MCP transport constructors via `LLM::MCP.stdio(...)` and
3464
- `LLM::MCP.http(...)`.
3465
-
3466
- ### Fix
3467
-
3468
- - Fix Google tool-call handling by synthesizing stable ids when Gemini does
3469
- not provide a direct tool-call id.
3470
-
3471
- ## v4.10.0
3472
-
3473
- Changes since `v4.9.0`.
3474
-
3475
- ### Add
3476
-
3477
- - Add HTTP transport for MCP with `LLM::MCP::Transport::HTTP` for remote servers
3478
- - Add JSON Schema union types (`any_of`, `all_of`, `one_of`) with parser integration
3479
- - Add JSON Schema type array union support (e.g., `"type": ["object", "null"]`)
3480
- - Add JSON Schema type inference from `const`, `enum`, or `default` fields
3481
-
3482
- ### Change
3483
-
3484
- - Update `LLM::MCP` constructor for exclusive `http:` or `stdio:` transport
3485
- - Update `LLM::MCP` documentation for HTTP transport support
3486
-
3487
- ## v4.9.0
3488
-
3489
- Changes since `v4.8.0`.
3490
-
3491
- ### Add
3492
-
3493
- - Add fiber-based concurrency with `LLM::Function::FiberGroup` and
3494
- `LLM::Function::TaskGroup` classes for lightweight async execution.
3495
- - Add `:thread`, `:task`, and `:fiber` strategy parameter to
3496
- `LLM::Function#spawn` for explicit concurrency control.
3497
- - Add stdio MCP client support, including remote tool discovery and
3498
- invocation through `LLM.mcp`, `LLM::Context`, and existing function/tool
3499
- APIs.
3500
- - Add model registry support via `LLM::Registry`, including model
3501
- metadata lookup, pricing, modalities, limits, and cost estimation.
3502
- - Add context access to a model context window via
3503
- `LLM::Context#context_window`.
3504
- - Add tracking of defined tools in the tool registry.
3505
- - Add `LLM::Schema::Enum`, enabling `Enum[...]` as a schema/tool
3506
- parameter type.
3507
- - Add top-level Anthropic system instruction support using Anthropic's
3508
- provider-specific request format.
3509
- - Add richer tracing hooks and extra metadata support for
3510
- LangSmith/OpenTelemetry-style traces.
3511
- - Add rack/websocket and Relay-related example work, including MCP-focused
3512
- examples.
3513
- - Add concurrent tool execution with `LLM::Function#spawn`,
3514
- `LLM::Function::Array` (`call`, `wait`, `spawn`), and
3515
- `LLM::Function::ThreadGroup`.
3516
- - Add `LLM::Function::ThreadGroup#alive?` method for non-blocking
3517
- monitoring of concurrent tool execution.
3518
- - Add `LLM::Function::ThreadGroup#value` alias for `ThreadGroup#wait` for
3519
- consistency with Ruby's `Thread#value`.
3520
-
3521
- ### Change
3522
-
3523
- - Rename `LLM::Session` to `LLM::Context` throughout the codebase to better
3524
- reflect the concept of a stateful interaction environment.
3525
- - Rename `LLM::Gemini` to `LLM::Google` to better reflect provider naming.
3526
- - Standardize model objects across providers around a smaller common
3527
- interface.
3528
- - Switch registry cost internals from `LLM::Estimate` to `LLM::Cost`.
3529
- - Update image generation defaults so OpenAI and xAI consistently return
3530
- base64-encoded image data by default.
3531
- - Update `LLM::Bot` deprecation warning from v5.0 to v6.0, giving users
3532
- more time to migrate to `LLM::Context`.
3533
- - Rework the README and screencast documentation to better cover MCP,
3534
- registry, contexts, prompts, concurrency, providers, and example flow.
3535
- - Expand the README with architecture, production, and provider guidance
3536
- while improving readability and example ordering.
3537
-
3538
- ### Fix
3539
-
3540
- - Fix local schema `$ref` resolution in `LLM::Schema::Parser`.
3541
- - Fix multiple MCP issues around stdio env handling, request IDs, registry
3542
- interaction, tool registration, and filtering of MCP tools from the
3543
- standard tool registry.
3544
- - Fix stream parsing issues, including chunk-splitting bugs and safer
3545
- handling of streamed error responses.
3546
- - Fix prompt handling across contexts, agents, and provider adapters so
3547
- prompt turns remain consistent in history and completions.
3548
- - Fix several tool/context issues, including function return wrapping,
3549
- tool lookup after deserialization, unnamed subclass filtering, and
3550
- thread-safety around tool registry mutations.
3551
- - Fix Google tool-call handling to preserve `thoughtSignature`.
3552
- - Fix `LLM::Tracer::Logger` argument handling.
3553
- - Fix packaging/docs issues such as registry files in the gemspec and
3554
- stale provider docs.
3555
- - Fix Google provider handling of `nil` function IDs during context
3556
- deserialization.
3557
- - Fix MCP stdio transport by increasing poll timeout for better
3558
- reliability.
3559
- - Fix Google provider to properly cast non-Hash tool results into Hash
3560
- format for API compatibility.
3561
- - Fix schema parser to support recursive normalization of `Array`,
3562
- `LLM::Object`, and nested structures.
3563
- - Fix DeepSeek provider to tolerate malformed tool arguments.
3564
- - Fix `LLM::Function::TaskGroup#alive?` to properly delegate to
3565
- `Async::Task#alive?`.
3566
- - Fix various RuboCop errors across the codebase.
3567
- - Fix DeepSeek provider to handle JSON that might be valid but unexpected.
3568
-
3569
- ### Notes
3570
-
3571
- Notable merged work in this range includes:
3572
-
3573
- - `feat(function): add fiber-based concurrency for async environments (#64)`
3574
- - `feat(mcp): add stdio MCP support (#134)`
3575
- - `Add LLM::Registry + cost support (#133)`
3576
- - `Consistent model objects across providers (#131)`
3577
- - `Add rack + websocket example (#130)`
3578
- - `feat(gemspec): add changelog URI (#136)`
3579
- - `feat(function): alias ThreadGroup#wait as ThreadGroup#value (#62)`
3580
- - `README and screencast refresh across `#66`, `#68`, `#71`, and
3581
- `#72`
3582
- - `chore(bot): update deprecation warning from v5.0 to v6.0`
3583
- - `fix(deepseek): tolerate malformed tool arguments`
3584
- - `refactor(context): Rename Session as Context (#70)`
3585
-
3586
- Comparison base:
3587
- - Latest tag: `v4.8.0` (`6468f2426ee125823b7ae43b4af507b125f96ffc`)
3588
- - HEAD used for this changelog: `915c48da6fda9bef1554ff613947a6ce26d382e3`