llm.rb 13.1.0 → 15.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (113) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +716 -1826
  3. data/README.md +668 -253
  4. data/bin/llm.rb +156 -56
  5. data/data/alibaba.json +1999 -0
  6. data/data/anthropic.json +195 -252
  7. data/data/bedrock.json +2189 -1854
  8. data/data/deepinfra.json +1312 -728
  9. data/data/deepseek.json +42 -39
  10. data/data/google.json +1024 -404
  11. data/data/mistral.json +481 -401
  12. data/data/moonshot.json +384 -0
  13. data/data/openai.json +985 -1354
  14. data/data/xai.json +220 -115
  15. data/data/zai.json +166 -166
  16. data/docs/deepdive/advanced/cancellation.md +74 -0
  17. data/docs/deepdive/advanced/compaction.md +85 -0
  18. data/docs/deepdive/advanced/context.md +280 -0
  19. data/docs/deepdive/advanced/guard.md +371 -0
  20. data/docs/deepdive/advanced/transformer.md +67 -0
  21. data/docs/deepdive/advanced/transports.md +45 -0
  22. data/docs/deepdive/features/builtin_tools.md +191 -0
  23. data/docs/deepdive/features/concurrency.md +110 -0
  24. data/docs/deepdive/features/database.md +449 -0
  25. data/docs/deepdive/features/embeddings.md +157 -0
  26. data/docs/deepdive/features/repl.md +141 -0
  27. data/docs/deepdive/fundamentals/agents.md +253 -0
  28. data/docs/deepdive/fundamentals/providers.md +159 -0
  29. data/docs/deepdive/fundamentals/schema.md +61 -0
  30. data/docs/deepdive/fundamentals/skills.md +111 -0
  31. data/docs/deepdive/fundamentals/stream.md +143 -0
  32. data/docs/deepdive/fundamentals/tools.md +329 -0
  33. data/docs/deepdive/media/audio.md +122 -0
  34. data/docs/deepdive/media/images.md +89 -0
  35. data/docs/deepdive/media/ocr.md +48 -0
  36. data/docs/deepdive/protocols/a2a.md +106 -0
  37. data/docs/deepdive/protocols/mcp.md +111 -0
  38. data/docs/deepdive/reference/cost.md +109 -0
  39. data/docs/deepdive/reference/model_registry.md +271 -0
  40. data/docs/deepdive/reference/object.md +108 -0
  41. data/docs/deepdive/reference/tracer.md +187 -0
  42. data/{resources → docs}/deepdive.md +37 -23
  43. data/lib/llm/a2a/transport/http.rb +1 -1
  44. data/lib/llm/active_record/acts_as_llm.rb +19 -5
  45. data/lib/llm/agent.rb +107 -17
  46. data/lib/llm/context.rb +164 -134
  47. data/lib/llm/cost.rb +114 -49
  48. data/lib/llm/error.rb +7 -8
  49. data/lib/llm/function/array.rb +1 -1
  50. data/lib/llm/function/async/task.rb +2 -0
  51. data/lib/llm/function/fiber/task.rb +2 -0
  52. data/lib/llm/function/fork/task.rb +16 -1
  53. data/lib/llm/function/ractor/task.rb +2 -0
  54. data/lib/llm/function/sequential/group.rb +20 -10
  55. data/lib/llm/function/sequential/task.rb +2 -9
  56. data/lib/llm/function/task.rb +4 -0
  57. data/lib/llm/function/thread/task.rb +2 -0
  58. data/lib/llm/function.rb +33 -4
  59. data/lib/llm/guard/loop.rb +89 -0
  60. data/lib/llm/guard/null.rb +19 -0
  61. data/lib/llm/guard.rb +61 -0
  62. data/lib/llm/message.rb +5 -4
  63. data/lib/llm/provider.rb +43 -0
  64. data/lib/llm/providers/alibaba/error_handler.rb +34 -0
  65. data/lib/llm/providers/alibaba/request_adapter.rb +13 -0
  66. data/lib/llm/providers/alibaba.rb +93 -0
  67. data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
  68. data/lib/llm/providers/anthropic.rb +1 -9
  69. data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
  70. data/lib/llm/providers/bedrock.rb +9 -9
  71. data/lib/llm/providers/deepseek/request_adapter.rb +2 -33
  72. data/lib/llm/providers/google/stream_parser.rb +1 -0
  73. data/lib/llm/providers/google.rb +1 -9
  74. data/lib/llm/providers/moonshot.rb +76 -0
  75. data/lib/llm/providers/ollama.rb +1 -9
  76. data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
  77. data/lib/llm/providers/openai/responses.rb +6 -9
  78. data/lib/llm/providers/openai/schema.rb +37 -0
  79. data/lib/llm/providers/openai/stream_parser.rb +1 -0
  80. data/lib/llm/providers/openai.rb +4 -12
  81. data/lib/llm/registry/model.rb +186 -0
  82. data/lib/llm/registry.rb +45 -14
  83. data/lib/llm/repl/bar.rb +11 -12
  84. data/lib/llm/repl/buffer.rb +43 -16
  85. data/lib/llm/repl/color.rb +85 -0
  86. data/lib/llm/repl/command.rb +12 -0
  87. data/lib/llm/repl/commands/model.rb +39 -0
  88. data/lib/llm/repl/input/cache.rb +45 -0
  89. data/lib/llm/repl/input/char.rb +46 -0
  90. data/lib/llm/repl/input/row.rb +39 -0
  91. data/lib/llm/repl/input.rb +327 -79
  92. data/lib/llm/repl/markdown/table.rb +6 -2
  93. data/lib/llm/repl/markdown.rb +56 -6
  94. data/lib/llm/repl/node.rb +7 -0
  95. data/lib/llm/repl/status.rb +54 -5
  96. data/lib/llm/repl/stream.rb +16 -4
  97. data/lib/llm/repl/walker.rb +3 -2
  98. data/lib/llm/repl/window.rb +111 -11
  99. data/lib/llm/repl.rb +47 -21
  100. data/lib/llm/sequel/plugin.rb +19 -5
  101. data/lib/llm/skill.rb +21 -8
  102. data/lib/llm/stream.rb +35 -7
  103. data/lib/llm/tool.rb +27 -0
  104. data/lib/llm/tools/rg.rb +2 -1
  105. data/lib/llm/transformer/null.rb +21 -0
  106. data/lib/llm/transformer.rb +55 -0
  107. data/lib/llm/transport/curb.rb +23 -3
  108. data/lib/llm/usage.rb +155 -9
  109. data/lib/llm/version.rb +1 -1
  110. data/lib/llm.rb +121 -29
  111. data/llm.gemspec +16 -10
  112. metadata +100 -13
  113. data/lib/llm/loop_guard.rb +0 -107
data/CHANGELOG.md CHANGED
@@ -15,7 +15,718 @@
15
15
 
16
16
  ## What's next
17
17
 
18
- *No unreleased changes yet. Check back after the next release.*
18
+ ## v15.0.0
19
+
20
+ Changes since `v14.0.0`.
21
+
22
+ This release renames `usage` to `token_usage` across contexts, agents,
23
+ and messages, makes `context_window` return `nil` when unknown, and
24
+ makes `LLM::Cost` accessors always return `Float`. It also adds the
25
+ Alibaba provider, automatic API key discovery from the environment,
26
+ new context-usage and context-used methods, a retry budget for
27
+ rate-limited requests, and a range of REPL and skills improvements.
28
+
29
+ ### Breaking
30
+
31
+ #### Migration
32
+
33
+ | Old | New |
34
+ |-----|-----|
35
+ | `ctx.usage` / `agent.usage` / `msg.usage` | `ctx.token_usage` / `agent.token_usage` / `msg.token_usage` (`usage` remains an alias) |
36
+ | `LLM::Message#usage` returns `LLM::Object` | `LLM::Message#token_usage` returns a copy of `LLM::Usage`, for assistant messages only |
37
+ | `ctx.context_window` returns `0` when the model isn't in the registry | returns `nil` when unknown |
38
+ | `LLM::Cost#input` (and other accessors) return `nil` when unused | return `0.0` |
39
+ | `ctx.usage` returns the most recent assistant message usage | sums token usage across all assistant messages |
40
+ | a skill exposes its tool as `weather` (the skill name) | the generated tool is now named `weather-skill` |
41
+
42
+ * **rename `#usage` to `#token_usage` across contexts, agents, and messages** <br>
43
+ [`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage-instance_method),
44
+ [`LLM::Agent#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#usage-instance_method),
45
+ and
46
+ [`LLM::Message#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Message.html#usage-instance_method)
47
+ are now aliases of `token_usage`.
48
+ [`LLM::Message#token_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Message.html#token_usage-instance_method)
49
+ now returns a copy of `LLM::Usage` instead of `LLM::Object`, and only
50
+ returns a value for assistant messages.
51
+
52
+ * **`LLM::Context#context_window` now returns `nil` when unknown** <br>
53
+ [`LLM::Context#context_window`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_window-instance_method)
54
+ now returns `nil` when the model's context window size is not known to the
55
+ runtime, instead of `0`. This makes the code check for a window instead
56
+ of a number, so an unknown window no longer reads as a real (zero) size.
57
+
58
+ * **`LLM::Cost` accessors always return `Float` objects** <br>
59
+ The cost accessors on
60
+ [`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
61
+ (`input`, `output`, `input_audio`, `output_audio`, `input_image`,
62
+ `cache_read`, `cache_write`, and `reasoning`) now always return a
63
+ `Float`, returning `0.0` when no tokens of that kind were used, instead
64
+ of `nil`. Callers can sum and compare cost values without guarding
65
+ against `nil`.
66
+
67
+ * **expose new context methods on the ActiveRecord and Sequel wrappers** <br>
68
+ The `acts_as_llm` (ActiveRecord) and `plugin :llm` (Sequel) wrappers
69
+ now expose `context_used` and `context_usage`, delegating to the
70
+ wrapped `LLM::Context`. `token_usage` replaces `usage` (which remains
71
+ as an alias), and `context_window` now returns `nil` when the model's
72
+ context window is unknown instead of `0`.
73
+
74
+ ### Core
75
+
76
+ * **discover API keys from the environment** <br>
77
+ [Cloud provider factories](https://r.uby.dev/api-docs/llm.rb/LLM.html)
78
+ (`LLM.anthropic`, `LLM.google`, `LLM.deepseek`, `LLM.openai`,
79
+ `LLM.xai`, `LLM.mistral`, `LLM.zai`, `LLM.moonshot`,
80
+ `LLM.alibaba`, and `LLM.aliyun`) now resolve the provider's API key
81
+ automatically when no `key:` is given, by walking the environment
82
+ variable names listed in the models.dev registry. So `LLM.openai`
83
+ works without an explicit key as long as `OPENAI_API_KEY` (or one of
84
+ the registry's alternative names) is set in the environment. A
85
+ missing key raises `ArgumentError`.
86
+
87
+ * **cli: auto-discover credentials and support Bedrock** <br>
88
+ `bin/llm.rb` now resolves the provider through the `LLM` factory
89
+ methods instead of mapping environment variable names directly, so it
90
+ picks up Bedrock (all three AWS credentials) and relies on the same
91
+ automatic key discovery as the library. The CLI also always starts
92
+ now: without arguments it falls back to `ollama` or `llamacpp`. A
93
+ provider whose credentials are not set exits with status 1.
94
+
95
+ * **cli: add `-c` and `-n` switches** <br>
96
+ `bin/llm.rb` now accepts a `-c STRATEGY` switch to choose the
97
+ concurrency strategy used for tool calls (`thread`, `async`, `fork`,
98
+ or any of the other strategies) and a `-n TRANSPORT` switch to choose
99
+ the HTTP transport (`net-http`, `net-http-persistent`, or `curb`),
100
+ both forwarded to the session's agent and provider.
101
+
102
+ * **context: keep runtime parameters from reaching the provider** <br>
103
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
104
+ now strips its runtime-only parameters (`guard`, `retry_budget`,
105
+ `concurrency`, `transformer`, and `compactor`) before merging params
106
+ into a provider request, so they can never cross the context-provider
107
+ boundary and risk an API-level error.
108
+
109
+ * **add `retry_budget` support for rate-limited requests** <br>
110
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
111
+ now accepts a `retry_budget:` that automatically sleeps and retries a
112
+ rate-limited request up to the given number of times before raising
113
+ `LLM::RateLimitError`. Each retry sleeps a growing interval (2s, 4s,
114
+ 6s, ...) and notifies the stream through
115
+ [`LLM::Stream#on_rate_limit`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_rate_limit-instance_method).
116
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
117
+ enables a budget of 5 by default, while a raw context disables it (0)
118
+ unless configured.
119
+
120
+ * **add `LLM::Usage.zero`** <br>
121
+ Add
122
+ [`LLM::Usage.zero`](https://r.uby.dev/api-docs/llm.rb/LLM/Usage.html#zero-class_method)
123
+ as a zero-valued usage object. `LLM::Context#usage`, `LLM::Agent#usage`,
124
+ and the ActiveRecord and Sequel wrappers now return `LLM::Usage` objects
125
+ instead of `LLM::Object` when no provider usage has been recorded yet.
126
+
127
+ * **add `LLM::Context#context_usage` and `LLM::Agent#context_usage`** <br>
128
+ Add
129
+ [`LLM::Context#context_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_usage-instance_method)
130
+ and
131
+ [`LLM::Agent#context_usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#context_usage-instance_method),
132
+ which return the fraction of the model's context window currently used as
133
+ a `Rational` (for example `Rational(100, 10_000)`), or `nil` when the used
134
+ amount or the window size is unknown. The REPL status bar now renders this
135
+ fraction instead of computing the remainder from raw token counts.
136
+
137
+ * **add `LLM::Context#context_used` and `LLM::Agent#context_used`** <br>
138
+ Add
139
+ [`LLM::Context#context_used`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_used-instance_method)
140
+ and
141
+ [`LLM::Agent#context_used`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#context_used-instance_method),
142
+ which return the live context size (in tokens) of the most recent
143
+ assistant message, or `nil` when no assistant message has a recorded token
144
+ usage. This fills the gap left after `token_usage` became accumulative and
145
+ no longer represented a single turn, so callers can read how much of the
146
+ context window has been used without walking the messages themselves.
147
+
148
+ ### Provider
149
+
150
+ * **add `LLM::Provider#registry`** <br>
151
+ Add [`LLM::Provider#registry`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#registry-instance_method),
152
+ which returns the provider's model registry. `LLM::Context#registry`
153
+ and `LLM::Agent#registry` now delegate to their underlying provider
154
+ instead of looking it up on their own.
155
+
156
+ * **add `LLM::Alibaba` for Alibaba Cloud Model Studio** <br>
157
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
158
+ is a new provider that talks to
159
+ [Alibaba Cloud Model Studio](https://www.alibabacloud.com/help/en/model-studio/models)
160
+ through its OpenAI-compatible API, including the Qwen3 family of
161
+ models. Create an instance with
162
+ [`LLM.alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM.html#alibaba-class_method),
163
+ also aliased as `LLM.aliyun`, which accepts the same `key:`, `host:`,
164
+ and `base_path:` options as the OpenAI provider. The provider defaults
165
+ to the `deepseek-v4-flash-0731` model and supports chat completions,
166
+ streaming, tool calls, and structured output through the shared
167
+ OpenAI-compatible path; image, audio, moderation, responses, and
168
+ vector store endpoints raise `NotImplementedError`. Model metadata
169
+ ships in `data/alibaba.json` for the registry.
170
+
171
+ * **alibaba: support structured outputs via `json_object`** <br>
172
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
173
+ now supports structured output through a shared `json_object` fallback,
174
+ since Alibaba models do not support `json_schema` natively. The schema is
175
+ described in an injected system message that also satisfies the
176
+ "messages must contain the word json" requirement. The same shared
177
+ fallback now also backs DeepSeek.
178
+
179
+ * **alibaba: default to the pay-as-you-go host** <br>
180
+ The default
181
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
182
+ host is now `dashscope-intl.aliyuncs.com`. Override it globally with
183
+ the `DASHSCOPE_API_HOST` environment variable, or per instance with
184
+ `LLM.alibaba(host: ...)`, for example to point at a Token Plan
185
+ endpoint.
186
+
187
+ * **alibaba: use `DASHSCOPE_API_KEY` as the default key env var** <br>
188
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
189
+ now discovers its API key from `DASHSCOPE_API_KEY` instead of
190
+ `ALIBABA_API_KEY`, following the models.dev registry convention.
191
+
192
+ * **alibaba: raise `LLM::InsufficientQuotaError` for exhausted quota** <br>
193
+ Add
194
+ [`LLM::InsufficientQuotaError`](https://r.uby.dev/api-docs/llm.rb/LLM/InsufficientQuotaError.html),
195
+ a subclass of `LLM::RateLimitError`, for when a provider reports a
196
+ tokens-per-minute (TPM) quota limit.
197
+ [`LLM::Alibaba`](https://r.uby.dev/api-docs/llm.rb/LLM/Alibaba.html)
198
+ now raises it when Alibaba responds with an `insufficient_quota`
199
+ error. Since it subclasses `RateLimitError`, quota errors are retried
200
+ like other rate limits.
201
+
202
+ * **bedrock: auto-discover AWS credentials from the environment** <br>
203
+ [`LLM.bedrock`](https://r.uby.dev/api-docs/llm.rb/LLM.html#bedrock-class_method)
204
+ now infers its credentials from the `AWS_ACCESS_KEY_ID`,
205
+ `AWS_SECRET_ACCESS_KEY`, and `AWS_REGION` environment variables when
206
+ they are not passed explicitly, matching the other cloud providers. A
207
+ missing key raises `ArgumentError`.
208
+
209
+ * **add `LLM::Bedrock#key?`** <br>
210
+ Add
211
+ [`LLM::Bedrock#key?`](https://r.uby.dev/api-docs/llm.rb/LLM/Bedrock.html#key%3F-instance_method),
212
+ which overrides the superclass method to check all three Bedrock
213
+ credentials (`access_key_id`, `secret_access_key`, and `region`)
214
+ instead of a single API key.
215
+
216
+ ### Function
217
+
218
+ * **make `Sequential::Group` abide by the `LLM::Function::Group` contract** <br>
219
+ [`LLM::Function::Sequential::Group`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Sequential/Group.html)
220
+ now receives an array of
221
+ [`LLM::Function::Sequential::Task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Sequential/Task.html)
222
+ objects instead of raw `LLM::Function` objects, matching the interface
223
+ shared by every other concurrency strategy.
224
+ [`LLM::Function::Array#task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Array.html#task-instance_method)
225
+ wraps each function as a `Sequential::Task` before constructing the
226
+ group, and the group delegates `spawn`, `alive?`, and `wait` to those
227
+ tasks. This fixes `Sequential::Group#alive?`, which always returned
228
+ `false`, and restores guard handling for sequential execution by
229
+ honoring the shared `guarded:` option on `Sequential::Task`.
230
+
231
+ * **function: redirect output streams in `:fork` tool processes** <br>
232
+ The `:fork` concurrency strategy (via
233
+ [`LLM::Function::Fork::Task`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Fork/Task.html))
234
+ now redirects the child process's `$stdout` and `$stderr` to
235
+ `File::NULL`, so a forked tool can no longer clobber the parent
236
+ terminal, for example by blanking the curses REPL display. A tool that
237
+ genuinely needs the terminal can still reopen `/dev/tty`; the file
238
+ descriptor stays available to the child.
239
+
240
+ ### Fix
241
+
242
+ * **a2a: fix a typo in the HTTP transport** <br>
243
+ Fix a bug in
244
+ [`LLM::A2A::Transport::HTTP`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A/Transport/HTTP.html)
245
+ where the constructor read `uri.port` instead of `@uri.port`, which
246
+ crashed the program whenever
247
+ [`LLM::A2A.rest`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html#rest-class_method)
248
+ or
249
+ [`LLM::A2A.jsonrpc`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html#jsonrpc-class_method)
250
+ was used. The transport now reads the port from the parsed `@uri`.
251
+
252
+ * **openai: report usage for streamed completions requests** <br>
253
+ Fix a bug in the OpenAI completions path where `params[:stream]` was
254
+ checked after it had been deleted from the params hash, so the check
255
+ always evaluated to `false`. The fix checks the resolved stream's
256
+ `enabled?` instead, so `stream_options: {include_usage: true}` is
257
+ added to streamed requests and API usage is reported back to the
258
+ caller.
259
+
260
+ * **curb: read the stream body and resolve streaming requests** <br>
261
+ Fix two bugs in [`LLM::Transport::Curb`](https://r.uby.dev/api-docs/llm.rb/LLM/Transport/Curb.html)
262
+ that left the `curb` transport unusable. The request body setter now
263
+ reads a streaming request's body stream into a string (dropping the
264
+ chunked transfer header, which curb replaces with a content length),
265
+ and the result builder now accumulates the response body from the
266
+ `on_body` callback instead of leaving it empty.
267
+
268
+ * **cli: handle errors in `main`** <br>
269
+ Wrap all of `bin/llm.rb`'s `main` method in error handling: an
270
+ interrupted session exits gracefully with `Bye!`, an explicit provider
271
+ is passed the resolved transport, and any unexpected error prints a
272
+ formatted diagnostic with a link to issue tracking before exiting.
273
+
274
+ * **cli: persist the session mapping file** <br>
275
+ Fix a bug where `bin/llm.rb` saved the session file at
276
+ `~/.llm.rb/<provider>/<uuid>.json` but never wrote the updated
277
+ working-directory mapping back to `~/.llm.rb/<provider>.json`. The
278
+ mapping file is now written whenever a new session is registered.
279
+
280
+ * **context: aggregate usage across all assistant messages** <br>
281
+ [`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage-instance_method)
282
+ now sums token usage across every assistant message in the conversation
283
+ instead of returning only the first message's usage.
284
+ [`LLM::Cost.from`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
285
+ now subtracts reasoning tokens from the output total and cache-read
286
+ tokens from the input total before pricing, and prices reasoning tokens
287
+ with the model's reasoning rate when one is available.
288
+
289
+ ### Repl
290
+
291
+ * **draw a top chrome row with the cwd and active model** <br>
292
+ The curses-based REPL now draws a white-on-blue row at the very top of
293
+ the screen showing the current working directory on the left and the
294
+ active model on the right. The row is drawn above the transcript and
295
+ uses a new blue status-bar color pair.
296
+
297
+ * **redraw the window on resize** <br>
298
+ The curses-based REPL now handles the terminal resize signal
299
+ (`KEY_RESIZE`) while reading input, clearing and redrawing the entire
300
+ window so the layout stays aligned after the terminal is resized.
301
+
302
+ * **hide the cursor until the window is ready** <br>
303
+ Fix a visual glitch where the curses-based REPL showed the cursor at
304
+ position 0,0 at startup and then jumped it to the input area once the
305
+ window was drawn. The cursor is now hidden until the input field has
306
+ been drawn and the cursor can be placed directly into it.
307
+
308
+ * **collapse the cost to two decimal places** <br>
309
+ [`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
310
+ now renders the total cost with two decimal places (for example
311
+ `$0.01`), so the REPL status bar shows a compact cost estimate instead
312
+ of a long run of digits.
313
+
314
+ * **add auto-complete ability for commands** <br>
315
+ [`LLM::Repl::Command`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl/Command.html)
316
+ subclasses can now override a `complete` method to autocomplete their
317
+ arguments. The method receives the command's parameters as keyword
318
+ arguments, with the non-nil keyword being the active fragment, and
319
+ returns candidate completions. Repeated TAB presses cycle through the
320
+ candidate list.
321
+
322
+ * **add `LLM::Repl#model` and `LLM::Repl#model=`** <br>
323
+ [`LLM::Repl`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl.html#model-instance_method)
324
+ now tracks the active model in its own `model` attribute, seeded from
325
+ the wrapped agent's model. The status bar reads the model through the
326
+ repl instead of the agent, so the model can be switched within a
327
+ session.
328
+
329
+ * **add `/model` command** <br>
330
+ A new `/model <name>` command switches the active model within a
331
+ single REPL session. Its argument auto-completes through the
332
+ text-to-text models in the registry.
333
+
334
+ * **repl: restrict autocomplete to text-to-text models** <br>
335
+ The `/model` command's argument auto-complete now suggests only
336
+ text-to-text models, so embedding and other non-chat models are
337
+ left out of the completion list.
338
+
339
+ * **repl: highlight GitHub-flavored codeblocks** <br>
340
+ The curses-based REPL now parses the GitHub-style ``` fences that
341
+ models commonly emit as real code blocks. Kramdown's native fenced-code
342
+ syntax uses `~~~`, so the ``` fences were previously parsed as inline
343
+ code spans. The language name is now shown in bold white above the code,
344
+ which renders in green.
345
+
346
+ * **repl: fix a scroll render artifact** <br>
347
+ Fix a bug where scrolling upward could leave a piece of text just
348
+ above the status row as a render artifact. The row above the status
349
+ row is now cleared on every buffer render.
350
+
351
+ * **repl: add a buffer row below the blue status bar** <br>
352
+ The curses-based REPL buffer now starts with an empty row below the
353
+ blue status bar, improving the visual spacing of the first exchange
354
+ in the chat.
355
+
356
+ * **repl: pin `curses` and `kramdown` to tested versions** <br>
357
+ The REPL now pins `curses` to `~> 1.6` and `kramdown` to `~> 2.5`
358
+ through `LLM.require`, so it loads gem versions known to have been
359
+ tested instead of whatever happens to be installed.
360
+
361
+ ### Skills
362
+
363
+ * **skills: append `-skill` to the generated tool name** <br>
364
+ A skill is now exposed as a tool named `"<skill>-skill"` instead of
365
+ just the skill's name, so a skill like `weather` that also uses a tool
366
+ named `weather` no longer collides with it (or a same-named global
367
+ tool) in the tool registry.
368
+
369
+ * **skills: add `LLM::Stream` skill lifecycle callbacks** <br>
370
+ Add
371
+ [`LLM::Stream#on_skill_call`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_skill_call-instance_method)
372
+ and
373
+ [`LLM::Stream#on_skill_return`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_skill_return-instance_method),
374
+ which are called before a skill's sub-agent runs and after it finishes.
375
+ `on_skill_return` receives the `LLM::Agent` sub-agent that ran the skill
376
+ along with the resulting `LLM::Response`, so a stream can inspect the
377
+ sub-agent's conversation, tally its usage, or add a verification step. A
378
+ stream can use the two callbacks to know when a skill sub-agent is
379
+ running.
380
+
381
+ ### Registry
382
+
383
+ * **add `LLM::Registry::Model` as a comparable model wrapper** <br>
384
+ Add [`LLM::Registry::Model`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry/Model.html),
385
+ a wrapper around a model's registry metadata (pricing, limits,
386
+ capabilities, and modalities). Models are comparable by price, so
387
+ `models.sort` orders them from cheapest to most expensive. The class
388
+ exposes predicate helpers such as `tool_call?`, `reasoning?`,
389
+ `structured_output?`, `open_weights?`, `text?`, `image?`, `audio?`,
390
+ `pdf?`, and `video?`, plus `input_cost`, `output_cost`, and
391
+ `context_window` accessors.
392
+
393
+ * **gemspec: bundle the deepdive guide from `docs/`** <br>
394
+ The gemspec now packages the deepdive guide from `docs/deepdive.md`
395
+ and `docs/deepdive/*/*.md` after the deepdive sources moved from
396
+ `resources/` to `docs/`, so the full guide ships with the gem.
397
+
398
+ * **make `LLM::Registry#models` return model objects** <br>
399
+ [`LLM::Registry#models`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry.html#models-instance_method)
400
+ now returns a list of
401
+ [`LLM::Registry::Model`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry/Model.html)
402
+ objects instead of model name strings. Use the new
403
+ [`LLM::Registry#keys`](https://r.uby.dev/api-docs/llm.rb/LLM/Registry.html#keys-instance_method)
404
+ method to get the model names.
405
+
406
+ * **refresh DeepInfra model metadata** <br>
407
+ Update `data/deepinfra.json` with current pricing for the DeepSeek
408
+ V4, DeepSeek-V3, DeepSeek-R1-0528, and Kimi-K3 models, and mark
409
+ `structured_output` support for one model.
410
+
411
+ ## v14.0.0
412
+
413
+ Changes since `v13.1.0`.
414
+
415
+ This release replaces the `transformer=` setter with the new
416
+ `LLM::Transformer` class hierarchy, refactors guards into the
417
+ `LLM::Guard` superclass with per-tool-call interception, and replaces
418
+ the agent `tool_attempts` parameter with the `tool_budget` class DSL.
419
+ It also adds the Moonshot (Kimi) provider, the `LLM::Tool.set`
420
+ bulk-assignment DSL, a `LLM::Function#return` shorthand, and a wide
421
+ range of REPL improvements.
422
+
423
+ ### Breaking
424
+
425
+ #### Migration
426
+
427
+ | Old | New |
428
+ |-----|-----|
429
+ | `ctx.transformer = MyTransformer` | `LLM::Context.new(transformer: MyTransformer)` |
430
+ | `transformer.call(ctx, prompt, params)` | `transformer.call(message:, **opts)` |
431
+ | `~/.llm.rb/session.json` (shared across providers) | `~/.llm.rb/<provider>/<uuid>.json` (scoped per provider and directory) |
432
+ | `agent.talk(tool_attempts: 25)` | `set :tool_budget => 50` (disabled by default) |
433
+ | `LLM::LoopGuard` | `LLM::Guard::Loop` |
434
+ | `guard: true` / `ctx.guard = MyGuard` | `guard: MyGuard, guard_options: {}` |
435
+ | `guard.call(ctx)` (warning string) | `guard.call(function:)` (`LLM::Function::Return` or nil) |
436
+ | `LLM::GuardError` | `"guard_error"` |
437
+
438
+ * **replace the transformer setter with `LLM::Transformer`** <br>
439
+ The previous `transformer=` setter and 3-argument
440
+ `call(ctx, prompt, params)` interface on `LLM::Context` have been
441
+ replaced by the new
442
+ [`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
443
+ class interface. Configure a transformer class through `transformer:`
444
+ and options through `transformer_options:` instead.
445
+
446
+ * **cli: scope session persistence per provider and directory** <br>
447
+ `bin/llm.rb` no longer shares a single session file between providers.
448
+ Each provider now has a `~/.llm.rb/<provider>.json` file that maps the
449
+ current working directory to a UUID-scoped session file under
450
+ `~/.llm.rb/<provider>/<uuid>.json`, so sessions are scoped to both the
451
+ provider and the directory they were started in.
452
+
453
+ * **cli: harden the executable against bad inputs** <br>
454
+ `bin/llm.rb` now prints an error message followed by the help menu and
455
+ exits with status 1 when the `-p` switch is given without an argument or
456
+ when an unknown option is passed. Previously unknown options produced a
457
+ warning but the run continued. The session-file lookup also no longer
458
+ rewrites `~/.llm.rb/<provider>.json` when it already exists.
459
+
460
+ * **agent: replace `tool_attempts` with the `tool_budget` class DSL** <br>
461
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
462
+ replaces the `tool_attempts` parameter with a `tool_budget` class DSL
463
+ (`tool_budget { 50 }`) that caps the number of tool calls allowed in a
464
+ single turn. Once the budget is spent, the agent sends an in-band
465
+ advisory message back through the model telling it to solve the problem
466
+ with fewer tool calls.
467
+ <br><br>
468
+ The feature is now disabled by default; the old `tool_attempts`
469
+ parameter defaulted to 25, which long-horizon agents could easily
470
+ exhaust in a single turn.
471
+
472
+ * **guard: replace `LLM::LoopGuard` with the `LLM::Guard` class hierarchy** <br>
473
+ [`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
474
+ is a new superclass for context-level supervisors, with
475
+ [`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
476
+ (replacing `LLM::LoopGuard`) and
477
+ [`LLM::Guard::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Null.html)
478
+ as the built-in implementations.
479
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html)
480
+ now accepts `guard:` (a guard class defaulting to `LLM::Guard::Null`)
481
+ and `guard_options:` (a hash forwarded to the guard's `call` method),
482
+ matching the transformer and compactor interfaces. The old boolean and
483
+ hash forms of `guard` and the `guard=` setter are removed. `LLM::Agent`
484
+ enables `LLM::Guard::Loop` by default.
485
+
486
+ * **guard: block individual tool calls instead of the whole batch** <br>
487
+ [`LLM::Guard#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html#call-instance_method)
488
+ now receives the pending `function:` and returns an
489
+ [`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
490
+ (or nil) instead of a warning string for the entire batch, so a guard
491
+ can block a single tool call while the rest of the batch still
492
+ executes. Custom guards that implemented the old `call(ctx)`
493
+ warning-string interface must be updated to return a
494
+ `LLM::Function::Return` instead.
495
+
496
+ * **errors: drop `LLM::GuardError`** <br>
497
+ Remove `LLM::GuardError`. The constant was never raised as an
498
+ exception; it only named the in-band error type for guarded tool
499
+ returns. Guarded tool returns now use the string `"guard_error"` as
500
+ their error type.
501
+
502
+ ### Core
503
+
504
+ * **add `LLM::Provider#build_messages` for assembling outgoing messages** <br>
505
+ [`LLM::Provider#build_messages`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#build_messages-instance_method)
506
+ normalizes a prompt into `LLM::Message` objects and prepends the existing
507
+ history, replacing the per-provider `build_complete_messages`
508
+ implementation. The method is idempotent: prompts that are already
509
+ [`LLM::Message`](https://r.uby.dev/api-docs/llm.rb/LLM/Message.html)
510
+ instances or arrays of messages are returned as-is.
511
+
512
+ * **copy the `params` hash in `LLM::Context` and `LLM::Agent`** <br>
513
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) and
514
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html) now copy
515
+ the `params` hash in their constructors before mutating it, leaving the
516
+ caller's hash untouched. Previously the constructors deleted keys from
517
+ the caller's hash in place.
518
+
519
+ * **gemspec: ship the deepdive sub-files in the gem** <br>
520
+ The gemspec now includes `resources/deepdive/*/*.md` in the gem
521
+ package, so the full deepdive guide (fundamentals, advanced,
522
+ protocols, and everything-else chapters) is available after
523
+ installation.
524
+
525
+ * **add short aliases to `LLM::Cost`** <br>
526
+ [`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html) now
527
+ offers short aliases for its cost accessors: `input`, `output`,
528
+ `input_audio`, `output_audio`, `input_image`, `cache_read`,
529
+ `cache_write`, and `reasoning`. Each alias matches the key used by
530
+ `#to_h`, so `cost.input` reads the same value as `cost.input_costs`.
531
+
532
+ ### Provider
533
+
534
+ * **add `LLM::Moonshot` for the Moonshot AI provider** <br>
535
+ [`LLM::Moonshot`](https://r.uby.dev/api-docs/llm.rb/LLM/Moonshot.html)
536
+ is a new provider that talks to
537
+ [Moonshot AI](https://platform.moonshot.ai) through its
538
+ OpenAI-compatible Kimi API. Create an instance with
539
+ [`LLM.moonshot`](https://r.uby.dev/api-docs/llm.rb/LLM.html#moonshot-class_method),
540
+ which accepts the same `key:`, `host:`, and `base_path:` options as the
541
+ OpenAI provider. The provider defaults to the `kimi-k3` model and
542
+ supports chat completions, streaming, tool calls, and structured output
543
+ through the shared OpenAI-compatible path; image, audio, moderation,
544
+ responses, and vector store endpoints raise `NotImplementedError`.
545
+ Model metadata ships in `data/moonshot.json` for the registry.
546
+
547
+ ### Transformer
548
+
549
+ * **add `LLM::Transformer` for rewriting messages before they reach the provider** <br>
550
+ [`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
551
+ is a new superclass for message transformers. A transformer is bound to a
552
+ context and rewrites a single message before it is sent to the provider,
553
+ which makes it possible to redact personal information or rewrite any
554
+ message before it goes out over the wire. Each subclass implements
555
+ `call(message:, **opts)` and returns the message to send, either by
556
+ mutating it in place or returning a new one.
557
+ [`LLM::Transformer::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer/Null.html)
558
+ is a no-op transformer used as the default.
559
+
560
+ * **hook the transformer API into `LLM::Context`** <br>
561
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html) now
562
+ accepts `transformer:` (a transformer class defaulting to
563
+ `LLM::Transformer::Null`) and `transformer_options:` (a hash forwarded to
564
+ the transformer's `call` method). The transformer runs on the most recent
565
+ message in both chat and responses turns.
566
+ [`LLM::Stream#on_transform`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform-instance_method)
567
+ and
568
+ [`LLM::Stream#on_transform_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform_finish-instance_method)
569
+ now receive the transformer instance as their single argument.
570
+
571
+ ### Tool
572
+
573
+ * **add `LLM::Tool.set` for bulk-assigning tool properties** <br>
574
+ [`LLM::Tool.set`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html#set-class_method)
575
+ accepts a hash of `name`, `description`, `parameters`, `required`, and
576
+ `defaults` to configure a tool subclass in a single call. Parameters are
577
+ defined as tuples of `[name, type, description, options]`, matching the
578
+ same interface as the existing `parameter` DSL. Unknown keys raise
579
+ `KeyError`.
580
+
581
+ ### Function
582
+
583
+ * **add `LLM::Function#return` for building tool returns** <br>
584
+ [`LLM::Function#return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#return-instance_method)
585
+ returns an
586
+ [`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
587
+ built from the function's own id and name, using the given hash as its
588
+ value. It is a shorthand mainly useful inside a
589
+ [`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
590
+ subclass and is defined via `define_method` because `return` is a Ruby
591
+ keyword.
592
+
593
+ ### Guard
594
+
595
+ * **run the guard for streamed tool calls** <br>
596
+ Fix a gap where the guard was not consulted when a tool call was queued
597
+ while a response was still streaming. The guard is now stamped onto the
598
+ functions a context binds, so it runs wherever a task is spawned,
599
+ including tool calls queued from a stream. A blocked call yields its
600
+ `guard_error` return without executing.
601
+
602
+ ### Agent
603
+
604
+ * **add `LLM::Agent#compacted?`** <br>
605
+ [`LLM::Agent#compacted?`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#compacted%3F-instance_method)
606
+ delegates to the wrapped
607
+ [`LLM::Context#compacted?`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#compacted%3F-instance_method)
608
+ and reports whether the conversation has been compacted, so callers
609
+ can detect when history was trimmed.
610
+
611
+ ### Change
612
+
613
+ * **openai: default to `gpt-5.6-luna`** <br>
614
+ The default OpenAI chat model has changed from `gpt-5.4-mini` to
615
+ `gpt-5.6-luna`. The new model is OpenAI's fastest and most affordable
616
+ option, matching the kind of default llm.rb aims for.
617
+
618
+ ### Repl
619
+
620
+ * **show an unknown context state after `/compact`** <br>
621
+ After running `/compact`, the REPL status line now renders `Context
622
+ compacted` and the context-usage bar shows `???` instead of a percentage,
623
+ because the used context is unknown until the next response.
624
+
625
+ * **land on a blank line after Ctrl+N at the end of history** <br>
626
+ When recalling history with Ctrl+P and Ctrl+N, Ctrl+N at the last item
627
+ now advances to a blank line so you can start typing new input, instead
628
+ of staying stuck on the last item in history (the previous behavior).
629
+ Recalling with Ctrl+P or Ctrl+N also no longer overwrites the input when
630
+ there is no history to show.
631
+
632
+ * **restore history wrap for Ctrl+P and Ctrl+N** <br>
633
+ Fix a regression where Ctrl+P and Ctrl+N recalled history text without
634
+ reflowing it into rows, so recalled lines wider than the terminal were
635
+ clipped. Recalled text now flows through the same word-wrap path as
636
+ typed input and wraps at the terminal width.
637
+
638
+ * **restore Ctrl+D deletion across rows** <br>
639
+ Fix a bug where Ctrl+D at the end of an input row was a no-op, so
640
+ multiline input could not be joined by deleting a row break. Deleting
641
+ at the end of a row now consumes the break and pulls the next row up,
642
+ restoring the split space so merged words do not run together.
643
+
644
+ * **center the buffer with 20% gutters** <br>
645
+ The curses-based REPL now centers
646
+ [`LLM::Repl::Buffer`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl/Buffer.html)
647
+ in a content area that is 60% of the terminal width, with an unused 20%
648
+ gutter on each side. The drawing area is based on the available rows and
649
+ columns instead of a fixed 80-column width, and `Buffer#wrap` now
650
+ hard-breaks words that overflow the width onto the next row, fixing a
651
+ bug where a word could be cut off between rows.
652
+
653
+ * **apply markdown to previous messages** <br>
654
+ The curses-based REPL now renders every message in the buffer with
655
+ markdown styling, including messages that were already present when
656
+ the session started or restored from disk. Previously only newly
657
+ streamed responses were styled; older messages fell back to plain
658
+ text.
659
+
660
+ * **add `LLM::Repl#sender` for the user label** <br>
661
+ [`LLM::Repl#sender`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl.html#sender-instance_method)
662
+ returns the label used for user messages in the curses-based REPL. It
663
+ defaults to `"You"` (previously `"user"`), and the buffer layout now
664
+ places each label on its own line followed by the message content and a
665
+ blank line.
666
+
667
+ * **add `LLM::Repl::Color` for coloring the curses UI** <br>
668
+ Add
669
+ [`LLM::Repl::Color`](https://r.uby.dev/api-docs/llm.rb/LLM/Repl/Color.html)
670
+ as a new module that returns Curses color bitmasks. `Color.enable`
671
+ initializes 8 color pairs, and methods like `Color.blue` return the
672
+ corresponding `Curses.color_pair(X)` bitmask, which can be bitwise
673
+ OR'ed with other attributes such as `Curses::A_BOLD`. User labels in
674
+ the REPL are now rendered in blue instead of plain bold text.
675
+
676
+ * **split on words rather than characters** <br>
677
+ `LLM::Repl::Buffer#wrap` now breaks text on word boundaries instead of
678
+ wrapping one character at a time. A word that does not fit on the
679
+ current row moves to the next, and only a single word longer than the
680
+ whole width is hard-broken, so text is never clipped by the window.
681
+
682
+ * **render kramdown typographic symbols and smart quotes** <br>
683
+ Fix a bug where certain character sequences such as `...` were not
684
+ rendered at all in the curses-based REPL. Kramdown parses them into
685
+ `:typographic_sym` and `:smart_quote` nodes, which previously fell
686
+ through to the children clause and were dropped. The markdown renderer
687
+ now maps them to their unicode equivalents: ellipsis, en and em
688
+ dashes, guillemets, and single and double quotation marks.
689
+
690
+ * **apply colors to the markdown renderer** <br>
691
+ The curses-based REPL now renders markdown with the `LLM::Repl::Color`
692
+ palette: headers and strong text in white, code spans and code blocks
693
+ in green, and links in underlined green, on the black background.
694
+ Previously markdown styling used bold, underline, and reverse video
695
+ attributes only.
696
+
697
+ * **wrap the input line at word boundaries** <br>
698
+ The curses-based REPL input line now wraps words whole onto the next
699
+ row at the terminal width instead of cutting them in half. A word that
700
+ does not fit on the current row moves to the next row, and only a
701
+ single word longer than the whole width is hard-broken, so typed text
702
+ is never clipped by the window.
703
+
704
+ * **distinguish the connecting and thinking status bar phases** <br>
705
+ The curses-based REPL status bar now shows `Connecting • Esc to
706
+ cancel` while the model is establishing a connection, then switches
707
+ to `Thinking • Esc to cancel` once a tool call or text fragment
708
+ arrives on the stream. Active tool calls appear in the status bar
709
+ with a lambda indicator.
710
+
711
+ * **add emoji to the status bar phases** <br>
712
+ The curses-based REPL status bar now uses emoji to identify each
713
+ phase at a glance: a globe (`🌐`) while the model is connecting, and
714
+ a brain (`🧠`) while it is thinking. The text after the emoji still
715
+ reads `Connecting • Esc to cancel` and `Thinking • Esc to cancel`
716
+ respectively.
717
+
718
+ * **render the status bar with color and attributes** <br>
719
+ The curses-based REPL status bar now supports colored and attributed
720
+ status text, so the lambda indicator for active tool calls is drawn
721
+ in bold red.
722
+
723
+ ### Registry
724
+
725
+ * **refresh model metadata across providers** <br>
726
+ Update `data/*.json` files with current provider model listings and
727
+ pricing. Remove the deprecated Claude Opus 4.1 entries from the
728
+ Anthropic registry, add `Qwen/Qwen3.8-Max` to DeepInfra, and add a
729
+ `low` reasoning-effort option to DeepSeek.
19
730
 
20
731
  ## v13.1.0
21
732
 
@@ -190,10 +901,9 @@ several agent and tool bugs around persistence, interruption, and naming.
190
901
 
191
902
  ## v13.0.0
192
903
 
193
- v13.0.0 relicenses the project under the MIT license, replacing
194
- the Business Source License that was introduced in v12.0.0. No
195
- commercial license is needed. Commercial, personal, educational, and
196
- all other uses are now permitted under the standard MIT terms.
904
+ v13.0.0 is released under the MIT license. Commercial, personal,
905
+ educational, and all other uses are permitted under the standard
906
+ MIT terms.
197
907
 
198
908
  Seven breaking changes. Concurrency strategies have been renamed
199
909
  (`:call` → `:sequential`, `:task` → `:async`), `spawn` is now
@@ -218,7 +928,7 @@ reliable across all six concurrency backends. The `functions` and
218
928
  | `Compactor.new(model:, token_threshold:)` | `Compactor::Truncate.new(ctx)` |
219
929
  | `on_compaction(ctx, compactor)` | `on_compaction(compactor)` |
220
930
  | `ctx.functions` / `ctx.functions?` | `ctx.pending_functions` / `ctx.pending_functions?` |
221
- | `agent.functions` / `agent.functions?` | `agent.pending_functions` / `agent.pending_functions?` |
931
+ | `agent.functions` / `agent.functions?` | `agent.pending_functions` |
222
932
 
223
933
  ### Breaking
224
934
 
@@ -1361,10 +2071,6 @@ Multiple _opt-in_ tools have been added to the `llm/tools/*.rb`
1361
2071
  directory. They serve as examples and as general-purpose tools
1362
2072
  that happen to power the repository's agents.
1363
2073
 
1364
- The BSL license has been extended to grant additional free waivers
1365
- for non-profits, charities and for companies with 50 or less
1366
- employees.
1367
-
1368
2074
  Other changes include small-ish bug fixes. <br>
1369
2075
  As always, see the changelog details for a thorough overview.
1370
2076
 
@@ -1427,12 +2133,6 @@ As always, see the changelog details for a thorough overview.
1427
2133
 
1428
2134
  ### Change
1429
2135
 
1430
- * **Extend BSL additional use grant** <br>
1431
- The Business Source License additional use grant has been extended to
1432
- include non-profits, charities, and companies with 50 or fewer
1433
- employees, in addition to the existing personal, education, and
1434
- evaluation uses.
1435
-
1436
2136
  * **Change LlamaCpp default port (8080 => 8013)** <br>
1437
2137
  The default port for the LlamaCpp provider has changed from `8080` to
1438
2138
  `8013` since llamacpp itself defaults to that port.
@@ -1456,1813 +2156,3 @@ As always, see the changelog details for a thorough overview.
1456
2156
  beyond. The `LLM::JSONAdapter.dump` method now walks serialized data
1457
2157
  and encodes every string into UTF-8, using `String#scrub` to replace
1458
2158
  bytes that are not valid UTF-8.
1459
-
1460
- ## v12.0.0
1461
-
1462
- Changes since `v11.3.1`.
1463
-
1464
- This release relicenses the project under the Business Source License,
1465
- defaults OpenAI to the Responses API and gpt-image models, adds the
1466
- DeepInfra provider with audio and image support, introduces
1467
- DeepSeek vector-graphics generation and schema support, extends xAI
1468
- image editing, adds `LLM::Schema.defaults` and schema string rendering,
1469
- and makes ActiveRecord and Sequel agent wrappers yield `LLM::Agent`
1470
- instead of polluting the model namespace.
1471
-
1472
- ### Breaking
1473
-
1474
- * **License change** <br>
1475
- The llm.rb runtime has been developed primarily by one
1476
- person for 3 years. That was done on my own time, and
1477
- I haven't made a dime from that work.
1478
-
1479
- So when I saw a multi-million dollar company benefit from
1480
- the work and for it to become the backbone of their AI
1481
- infrastructure and then see them not contribute back or
1482
- offer any kind of support, I decided this is not sustainable,
1483
- or fair.
1484
-
1485
- I assumed good faith and for people to act in the spirit of
1486
- open source but sadly, that's just not the case. I
1487
- have to choose a license that respects my time and effort.
1488
-
1489
- For those reasons, llm.rb is being relicensed under the
1490
- [Business Source license](https://mariadb.com/bsl11/).
1491
- So what does that mean?
1492
-
1493
- In a nutshell:
1494
-
1495
- * Free for personal use.
1496
- * Free for education.
1497
- * Free for evaluation, development, and testing.
1498
- * Commercial production use requires a commercial license.
1499
- * Exemptions on a case-by-case basis
1500
-
1501
- After 4 years, the license expires and it will become
1502
- available under the 0BSDL as it was before v12.0.0.
1503
- These 4 years apply to a specific version, and not the
1504
- project overall.
1505
-
1506
- Going forward, v12.0.0 will be relicensed to respect
1507
- my time, energy, and effort. llm.rb took an incredible
1508
- amount of time and effort, and continues to do so, so
1509
- I want to protect myself from companies who benefit
1510
- from my work but don't respect the time or effort that
1511
- was put into it.
1512
-
1513
- * **OpenAI: default to the Responses API** <br>
1514
- The responses API has both models and features that are unavailable
1515
- on the chat completions API, and the responses API appears to be
1516
- the API of the future for OpenAI.
1517
-
1518
- Worth noting: the llm.rb implementation does **not** store state
1519
- server-side by default. This can be changed with the `store: true`
1520
- option. The legacy chat completions API can be accessed with the
1521
- `mode: :completions` option.
1522
-
1523
- llm.rb has had support for the responses API for quite
1524
- a while but it was not the default, and a number of bugs
1525
- were found and fixed during the process of making it the
1526
- default.
1527
-
1528
- * **OpenAI: use gpt-image for image generation** <br>
1529
- The `dalle` models are in the process of being deprecated, and support
1530
- has been dropped from llm.rb. The `gpt-image` models are the next-generation
1531
- image-generation models from OpenAI.
1532
-
1533
- * **xAI: provide images as base64-encoded data** <br>
1534
- Both xAI, and OpenAI had the option to generate images via a URL
1535
- you can fetch, or as a base64-encoded string embedded directly
1536
- in the response.
1537
-
1538
- OpenAI is moving away from the URL transport since deprecating dalle,
1539
- and with that in mind, llm.rb has dropped support for the URL transport
1540
- across all providers that supported it.
1541
-
1542
- Google, xAI, and OpenAI now consistently provide generated and modified
1543
- images as a base64-encoded string.
1544
-
1545
- * **ActiveRecord: yield `LLM::Agent` to `acts_as_agent`** <br>
1546
- With this change we yield an instance of `LLM::Agent` to the `acts_as_agent`
1547
- method, and drop the methods (such as `model`, `instructions`, etc) that
1548
- were previously defined directly on the model. This keeps the number of
1549
- methods that llm.rb adds to an ActiveRecord model at a minimum and retains
1550
- the same capabilities as before.
1551
-
1552
- * **Sequel: yield `LLM::Agent` to `plugin(:agent)`** <br>
1553
- Ditto as above but for Sequel.
1554
-
1555
- * **Remove the langsmith tracer** <br>
1556
- This code was contributed by a third party but contains
1557
- many anti-patterns that are against llm.rb conventions
1558
- and best practices. It was merged without oversight or
1559
- review, and basically against the ethos of open source.
1560
-
1561
- I also don't have a langsmith account to maintain the
1562
- code. The alternative is the `LLM::Tracer::Telemetry` class
1563
- that was originally written by me, and serves as a
1564
- general-purpose OTP tracer.
1565
-
1566
- ### Add
1567
-
1568
- * **Add a new provider: LLM::DeepInfra** <br>
1569
- [DeepInfra](https://deepinfra.com) provide OpenAI-compatible
1570
- endpoints for a large catalog of hosted open-source and
1571
- open-weight models. <br> Capabilities like tool calling, structured outputs, and
1572
- reasoning can depend on the model.
1573
-
1574
- * **Add new image provider: LLM::DeepInfra::Images** <br>
1575
- [DeepInfra](https://deepinfra.com) provide access to
1576
- diverse set of text-to-image models. <br> Learn more about the
1577
- available models on their [text-to-image models](https://deepinfra.com/models/text-to-image)
1578
- page.
1579
-
1580
- * **DeepSeek: add `LLM::DeepSeek::Images#create` and `#edit`** <br>
1581
- This new API can generate and edit vector graphics (SVGs). <br>
1582
- It is an experimental approach and API.
1583
-
1584
- DeepSeek does not provide an image generation model however
1585
- its text-to-text models can generate SVG documents, and
1586
- that's the approach this feature takes. It is limited
1587
- to vector graphics rather than raster images.
1588
-
1589
- * **DeepSeek: attach `LLM::Response#agent` to image responses** <br>
1590
- The DeepSeek image API is built on top of
1591
- [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html).
1592
- Image responses now expose that agent via `res.agent`, which makes
1593
- it possible to carry the same session across multiple generations
1594
- or edits.
1595
-
1596
- * **xAI: add `LLM::XAI::Images#edit`** <br>
1597
- With this change it is possible to both generate images
1598
- from a prompt, and edit an existing image with a prompt.
1599
- xAI now has the same edit and create capabilities that
1600
- OpenAI has.
1601
-
1602
- * **Add `LLM::Schema.defaults`** <br>
1603
- This method lets you map multiple property names to
1604
- different default values. It is similar to `LLM::Schema.required`
1605
- in the sense that it is called after the properties of
1606
- a schema have been defined.
1607
-
1608
- * **Add `LLM::Schema#to_s` and `LLM::Schema.to_s`** <br>
1609
- Schemas can now be rendered as a prompt-friendly string.
1610
- This is useful when the shape of a schema needs to be
1611
- described in natural-language instructions rather than
1612
- passed through a native structured output interface.
1613
-
1614
- * **DeepSeek: add `LLM::Schema` support** <br>
1615
- DeepSeek can now use `schema:` for structured output.
1616
- llm.rb handles this by setting `response_format: {type: "json_object"}`
1617
- and describing the schema in a system message.
1618
-
1619
- * **OpenAI: add local file support to the Responses API** <br>
1620
- Our responses API implementation lacked local file support. <br>
1621
- This change fixes that by supporting both image, document,
1622
- and other media types that OpenAI may support.
1623
-
1624
- * **Add `LLM::Response#id` across all providers** <br>
1625
- This method was previously implemented via `method_missing`,
1626
- and the field name could change depending on the provider.
1627
- The new method is a catch-all that provides a single method
1628
- that works across all providers.
1629
-
1630
- * **Add `LLM::DeepInfra::Audio`** <br>
1631
- DeepInfra implements most of the llm.rb audio interface
1632
- with both the `create_speech` and `create_transcription`
1633
- methods. The `create_translation` method is not implemented,
1634
- and the available text-to-speech and speech-to-text models
1635
- are more varied than other providers.
1636
-
1637
- * **OpenAI: normalize text-to-speech responses** <br>
1638
- The `res.audio` method now returns an
1639
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
1640
- object for OpenAI text-to-speech responses. The object provides
1641
- `encoded`, `decoded`, `content_type`, and `encoding_type`.
1642
-
1643
- * **DeepInfra: normalize text-to-speech responses** <br>
1644
- The `res.audio` method now returns an
1645
- [`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
1646
- object for DeepInfra text-to-speech responses. The object provides
1647
- `encoded`, `decoded`, `content_type`, and `encoding_type`.
1648
-
1649
- ### Fix
1650
-
1651
- * **Fix Google `temperature` parameter fall-through** <br>
1652
- Ensure provider-level `temperature` and other `generationConfig`
1653
- parameters are forwarded to the API correctly instead of being
1654
- silently dropped.
1655
-
1656
- * **Fix Google `generationConfig` collisions** <br>
1657
- Prevent duplicate or conflicting `generationConfig` keys in the
1658
- Google request adapter.
1659
-
1660
- ### Change
1661
-
1662
- * **Change OpenAI defaults** <br>
1663
- The default chat model is now `gpt-5.4-mini`. <br>
1664
- The default image model is now `gpt-image`.
1665
-
1666
- * **Change google defaults** <br>
1667
- The default chat model is now `gemini-3.1-flash-lite` <br>
1668
- The default embeddings model is now `gemini-embedding-2`
1669
-
1670
- * **Change xAI defaults** <br>
1671
- The default chat model is now `grok-4.3`. <br>
1672
- The default image model is now `grok-imagine-image-quality`.
1673
-
1674
- * **Return an `LLM::Object` from `LLM::Response#content!`** <br>
1675
- The Hash-like, indifferent access data structure known as
1676
- `LLM::Object` provides a convenient interface around a Hash
1677
- object. It allows method access via `obj.key`, and decays
1678
- into a Hash in many cases.
1679
-
1680
- The `LLM::Response#content!` method now wraps its content
1681
- in an `LLM::Object` but only after it has parsed its
1682
- content (a JSON string) into a Ruby data structure.
1683
-
1684
- * **Refresh model metadata** <br>
1685
- Update `data/*.json` files with current provider model listings,
1686
- pricing, and capabilities.
1687
-
1688
- ## v11.3.1
1689
-
1690
- Changes since `v11.3.0`.
1691
-
1692
- This release rebrands the project under the r.uby.dev umbrella, removes
1693
- the Jekyll-based docs site in favor of a pure-markdown deepdive, and
1694
- cleans up YARD documentation across the codebase.
1695
-
1696
- ### Change
1697
-
1698
- * **Rebrand to r.uby.dev** <br>
1699
- Update README.md with the new logo, streamlined copy, and r.uby.dev
1700
- URLs. Rewrite `resources/deepdive.md` as a concise walkthrough and
1701
- bundle it with the gem. Remove the `docs/` directory (Jekyll site).
1702
- Update all references from `llmrb.github.io` to `r.uby.dev`.
1703
-
1704
- * **Update gemspec** <br>
1705
- Update homepage, metadata URLs, email, and author list. Switch the
1706
- YARD markdown processor from kramdown to redcarpet.
1707
-
1708
- ### Fix
1709
-
1710
- * **Fix YARD documentation** <br>
1711
- Fix unnamed, misnamed, and missing `@param` tags across provider
1712
- adapters, transport classes, stream, tool, schema, registry, agent,
1713
- and ActiveRecord integration files. Fix backtick-wrapped constant
1714
- references and other YARD formatting issues.
1715
-
1716
- ## v11.3.0
1717
-
1718
- Changes since `v11.2.0`.
1719
-
1720
- This release promotes `LLM::Agent` as the default high-level runtime,
1721
- raises `LLM::NotFoundError` for provider 404 responses, and adds
1722
- Symbol resolution to `LLM::Agent.confirm` and `LLM::Agent.skills` for
1723
- dynamic tool confirmation and skill lists.
1724
-
1725
- ### Add
1726
-
1727
- * **Raise `LLM::NotFoundError` for provider 404 responses** <br>
1728
- Raise `LLM::NotFoundError` when a provider returns HTTP 404. One
1729
- example is calling the embeddings API on DeepSeek
1730
- (`LLM.deepseek(...).embed(["foobar"])`), which returns 404 because
1731
- DeepSeek does not implement that endpoint.
1732
-
1733
- * **Add Symbol resolution to `LLM::Agent.confirm`** <br>
1734
- When `confirm` receives a single Symbol argument, it stores it
1735
- as-is instead of converting it to a string array. At initialization
1736
- time, `resolve_option` resolves the Symbol by calling the method
1737
- with that name on the agent instance, and the result is converted
1738
- to strings. This allows dynamic tool confirmation lists:
1739
-
1740
- class MyAgent < LLM::Agent
1741
- confirm :tools_that_need_confirmation
1742
-
1743
- def tools_that_need_confirmation
1744
- some_condition ? %w[delete destroy] : %w[delete]
1745
- end
1746
- end
1747
-
1748
- Ported from llmrb/mruby-llm@89a232e3 and @2dd04e2d.
1749
-
1750
- Extend the same pattern to `LLM::Agent.skills` so the skills DSL
1751
- accepts a Symbol that resolves through the agent instance at
1752
- initialization time.
1753
-
1754
- ### Change
1755
-
1756
- * **Clarify `LLM::Agent` as the default high-level runtime** <br>
1757
- Document that `LLM::Context` remains at the heart of llm.rb, but
1758
- `LLM::Agent` is the better default unless an application needs advanced
1759
- manual tool loops. `LLM::Agent` manages the tool loop for callers and
1760
- enables guards against runaway or repeated tool-call loops.
1761
-
1762
- ## v11.2.0
1763
-
1764
- Changes since `v11.1.0`.
1765
-
1766
- This release adds `LLM::Function#skill?` and `LLM::Tool#skill?` so
1767
- callers can inspect whether a function or tool is backed by a skill.
1768
-
1769
- It introduces `LLM::Transport::Request` as a transport-agnostic request
1770
- object so providers no longer depend directly on `Net::HTTP` request
1771
- classes, and adds an optional Curb (libcurl) backend alongside symbolic
1772
- transport shortcuts such as `transport: :curb`.
1773
-
1774
- MCP and A2A clients now accept `persistent: true` matching provider configuration.
1775
- Several fixes land for tool return callback emission, function comparison by
1776
- tool call ID, function array filtering, skill tool inheritance, and JSON generator
1777
- state compatibility on Ruby 4.
1778
-
1779
- ### Add
1780
-
1781
- * **Add `LLM::Function#skill?`** <br>
1782
- Add `skill?` to `LLM::Function` so callers can check whether a
1783
- function is backed by a skill tool.
1784
-
1785
- * **Add `LLM::Tool.skill?` and `LLM::Tool#skill?`** <br>
1786
- Add class-level `skill?` and instance-level `skill?` to
1787
- `LLM::Tool`, matching the existing `mcp?` and `a2a?` pattern.
1788
-
1789
- * **Add `LLM::Transport::Request`** <br>
1790
- Add `LLM::Transport::Request` as a transport-agnostic request object
1791
- and update providers to build requests without depending directly on
1792
- Net::HTTP request classes. The built-in Net::HTTP transports still
1793
- accept existing Net::HTTP request objects through a compatibility
1794
- bridge, while alternative transports can handle the generic request
1795
- shape directly.
1796
-
1797
- * **Add optional Curb transport support** <br>
1798
- Add `LLM::Transport::Curb`, an optional libcurl-backed transport
1799
- that can be selected with `transport: :curb`. Providers already
1800
- emit `LLM::Transport::Request` objects, so the Curb backend can
1801
- execute requests without routing through Net::HTTP.
1802
-
1803
- * **Add symbolic transport shortcuts** <br>
1804
- Allow providers, MCP HTTP clients, and A2A HTTP clients to accept
1805
- transport shortcuts such as `transport: :curb` and
1806
- `transport: :net_http_persistent`.
1807
-
1808
- * **Add persistent HTTP selection to MCP and A2A clients** <br>
1809
- Allow MCP and A2A HTTP clients to accept `persistent: true`, matching
1810
- provider configuration and selecting the persistent Net::HTTP
1811
- transport by default.
1812
-
1813
- ### Fix
1814
-
1815
- * **Support JSON generation state on Ruby 4** <br>
1816
- Handle JSON generator state objects in the standard JSON adapter so
1817
- schema objects serialize correctly when Ruby 4 calls custom `to_json`
1818
- methods during provider request generation.
1819
-
1820
- * **Emit tool return callbacks for direct context waits** <br>
1821
- Emit `LLM::Stream#on_tool_return` when `LLM::Context#wait` executes
1822
- pending tool work directly instead of draining `LLM::Stream::Queue`.
1823
-
1824
- * **Emit confirmed tool return callbacks once** <br>
1825
- Emit `LLM::Stream#on_tool_return` for confirmed and cancelled tool
1826
- calls, and exclude confirmed functions from later waits so mixed
1827
- confirmed and unconfirmed tool batches do not execute confirmed tools
1828
- twice.
1829
-
1830
- * **Compare functions by tool call ID** <br>
1831
- Add `LLM::Function#==`, `#eql?`, and `#hash` so pending function
1832
- collections can compare tool calls by provider-assigned ID instead of
1833
- object identity.
1834
-
1835
- * **Preserve function array behavior after filtering** <br>
1836
- Preserve `LLM::Function::Array` behavior when subtracting function
1837
- arrays so filtered tool batches can still spawn through the normal
1838
- function array API.
1839
-
1840
- * **Prevent skills from inheriting skill-backed tools** <br>
1841
- Exclude skill-backed tools when a skill sub-agent uses `tools:
1842
- inherit`, preventing skills loaded through a parent context from
1843
- being recursively exposed to nested skill agents.
1844
-
1845
- ## v11.1.0
1846
-
1847
- Changes since `v11.0.0`.
1848
-
1849
- This release adds the `inherit` directive for skill sub-agents so they can
1850
- inherit access to the local, MCP, and A2A tools available to their parent
1851
- agent. It introduces class-level `required %i[...]` declarations to
1852
- `LLM::Schema` and wraps `LLM::Function#arguments` in `LLM::Object` for
1853
- method-style argument access. The OpenTelemetry tracer now samples all spans
1854
- regardless of environment, and the tool-call loop repair step prevents stale
1855
- history from being sent on follow-up requests.
1856
-
1857
- ### Add
1858
-
1859
- * **Add support for the `inherit` directive in skills** <br>
1860
- Add support for the `inherit` directive so a skill sub-agent can
1861
- inherit access to the local, MCP, and A2A tools available to its
1862
- parent agent.
1863
-
1864
- * **Add class-level `required %i[...]` support to `LLM::Schema`** <br>
1865
- Add class-level `required %i[...]` declarations to `LLM::Schema`, so
1866
- schema classes can mark existing properties as required the same way
1867
- `LLM::Tool` params already can.
1868
-
1869
- * **Wrap function arguments in `LLM::Object`** <br>
1870
- Wrap `LLM::Function#arguments` in `LLM::Object`, so function
1871
- implementations can read arguments with method-style access while
1872
- still invoking runners with keyword arguments.
1873
-
1874
- ### Fix
1875
-
1876
- * **Ensure all traces are sampled regardless of environment** <br>
1877
- Explicitly pass `Samplers::ALWAYS_ON` when creating the OpenTelemetry
1878
- `TracerProvider` so the in-memory exporter always captures every span,
1879
- regardless of the `OTEL_TRACES_SAMPLER` environment variable.
1880
-
1881
- * **Always close the tool call loop before sending follow-up requests** <br>
1882
- Add a repair step in `Context#talk` that closes assistant tool-call
1883
- messages without matching tool responses before the next provider
1884
- request is sent. This prevents stale tool-call history from being sent
1885
- on follow-up requests, which some providers reject as invalid.
1886
-
1887
- ## v11.0.0
1888
-
1889
- Changes since `v10.0.0`.
1890
-
1891
- This release removes several deprecated or unused APIs, including the `#chat`
1892
- alias from contexts and agents, the `LLM::Function#register` alias, and the
1893
- unused positional `llm` argument from MCP constructors. Generated MCP and A2A
1894
- tools are no longer added to the global tool registry by default.
1895
-
1896
- On the additions side, it introduces the A2A (Agent2Agent) protocol client,
1897
- a new `#ask` convenience interface on contexts and agents, one-shot stdio MCP
1898
- requests outside `#session`, `LLM::Function#def` as a short alias for
1899
- `LLM::Function#define`, `LLM::File#exist?`, and `LLM::Tool.a2a?`.
1900
-
1901
- ### Breaking
1902
-
1903
- * **Remove the unused `llm` argument from MCP clients** <br>
1904
- Remove the unused positional `llm` argument from `LLM::MCP.new`,
1905
- `LLM::MCP.stdio`, `LLM::MCP.http`, and `LLM.mcp`.
1906
-
1907
- * **Stop globally registering generated MCP and A2A tools** <br>
1908
- Generated tools returned by `LLM::Tool.mcp(...)` and
1909
- `LLM::Tool.a2a(...)` are no longer added to the global
1910
- `LLM::Tool.registry` or `LLM::Function.registry`. They still work
1911
- when passed directly to a context or agent, but registry-based lookup
1912
- now only sees normal loaded `LLM::Tool` subclasses.
1913
-
1914
- * **Remove `LLM::Function#register`** <br>
1915
- Remove the `LLM::Function#register` alias and prefer
1916
- `LLM::Function#define` or `LLM::Function#def` when binding a
1917
- function to its implementation. The `register` alias was too easy to
1918
- confuse with the class-level `LLM::Tool.register` and
1919
- `LLM::Function.register` registry APIs.
1920
-
1921
- * **Remove the `#chat` alias from contexts and agents** <br>
1922
- Remove the `LLM::Context#chat` and `LLM::Agent#chat` aliases. Prefer
1923
- `#talk` for all context and agent turns.
1924
-
1925
- ### Add
1926
-
1927
- * **Add `LLM::Function#def`** <br>
1928
- Add `LLM::Function#def` as a short alias for
1929
- `LLM::Function#define` when binding a function instance to its
1930
- implementation.
1931
-
1932
- * **Add `LLM::MCP#session`** <br>
1933
- Add `LLM::MCP#session` as an alias for `LLM::MCP#run`, and prefer it
1934
- in examples for scoped stdio MCP sessions that should stay alive
1935
- across discovery and tool calls.
1936
-
1937
- * **Add `#ask` to contexts and agents** <br>
1938
- Add `LLM::Context#ask` and `LLM::Agent#ask` as a RubyLLM-compatible
1939
- convenience interface over `#talk`. `#ask` accepts a prompt, optional
1940
- `with:` attachments, an optional `stream:` target, and an optional
1941
- block for streamed chunks, and returns an `LLM::Response`.
1942
-
1943
- * **Add `LLM::File#exist?`** <br>
1944
- Add `LLM::File#exist?` as a small convenience wrapper for checking
1945
- whether a local file exists on disk.
1946
-
1947
- * **Allow one-shot stdio MCP requests outside `#session`** <br>
1948
- Allow `mcp.tools`, `mcp.prompts`, `mcp.find_prompt(...)`, and
1949
- `mcp.call_tool(...)` to work outside `mcp.session` by starting and
1950
- stopping a stdio transport on demand when needed. This makes stdio
1951
- MCP usable without an explicit session block, while keeping
1952
- `mcp.session` as the preferred pattern for efficient, stateful
1953
- stdio workflows.
1954
-
1955
- * **Add A2A client support** <br>
1956
- Add `LLM::A2A`, a client for the Agent2Agent (A2A) protocol with
1957
- REST and JSON-RPC bindings. Remote agent skills can be exposed as
1958
- `LLM::Tool` classes and used through `LLM::Context` or `LLM::Agent`,
1959
- and the client also supports direct messaging, streaming, task
1960
- operations, push notification configuration, extended agent cards,
1961
- persistent HTTP transport selection, and optional REST `base_path`
1962
- prefixing.
1963
-
1964
- Refactor shared MCP/A2A HTTP transport setup into
1965
- `LLM::Transport::Utils`, and extend
1966
- `LLM::Transport::StreamDecoder` to accept a callback block directly.
1967
-
1968
- * **Add `LLM::Tool.a2a?`** <br>
1969
- Add `LLM::Tool.a2a?` and mark generated A2A-backed tool classes so
1970
- callers can distinguish them from local or MCP tools.
1971
-
1972
- ### Fix
1973
-
1974
- * **Fix context and agent JSON serialization through `LLM.json`** <br>
1975
- Fix `LLM::Context#to_json` and `LLM::Agent#to_json` to serialize
1976
- through `LLM.json.dump(...)` instead of plain `to_json`.
1977
-
1978
- * **Fix block-form ORM agent DSL forwarding** <br>
1979
- Fix block-form `model { ... }`, `tools { ... }`, and
1980
- `schema { ... }` declarations in the ActiveRecord and Sequel agent
1981
- wrappers so persisted agent models configure the internal agent class
1982
- the same way `LLM::Agent` does.
1983
-
1984
- * **Fix missing `skills` in ORM agent wrappers** <br>
1985
- Fix the ActiveRecord and Sequel agent wrappers to expose `skills`, so
1986
- persisted agent models can declare skills the same way as
1987
- `LLM::Agent`.
1988
-
1989
- * **Fix `acts_as_agent#ctx` return type** <br>
1990
- Fix the ActiveRecord `acts_as_agent` wrapper so its `ctx` helper
1991
- returns the wrapped `LLM::Agent` instead of returning the underlying
1992
- `LLM::Context` directly.
1993
-
1994
- ## v10.0.0
1995
-
1996
- Changes since `v9.0.0`.
1997
-
1998
- This release removes the `LLM::Context#respond` method, and
1999
- also removes the deprecated `LLM::Bot` alias. **All** class-level
2000
- agent tunables can now be resolved lazily via a Symbol (method name),
2001
- or a Proc. The `LLM::Agent` class can now confirm a tool call
2002
- before it happens, and the `LLM::Schema` class has been extended
2003
- to support `Array[String,Integer]` as a shorthand for
2004
- `Array[AnyOf[String, Integer]]`. The `LLM::Stream` class has
2005
- had its public method surface reduced to help avoid accidental
2006
- collisions.
2007
-
2008
- ### Breaking
2009
-
2010
- * **Unify context turns under `#talk`** <br>
2011
- Remove `LLM::Context#respond` and route responses-mode turns through
2012
- `LLM::Context#talk` with `mode: :responses` instead.
2013
-
2014
- * **Remove the `LLM::Bot` alias** <br>
2015
- Remove the backward-compatible `LLM::Bot` alias for `LLM::Context`.
2016
- Use `LLM::Context` directly instead.
2017
-
2018
- ### Add
2019
-
2020
- * **Add shared option resolution through `LLM::Utils`** <br>
2021
- Add `LLM::Utils.resolve_option` for resolving configured values as
2022
- literals, procs, symbol-named methods, or duplicated hashes, and use
2023
- it in agent and ORM option resolution paths.
2024
-
2025
- * **Resolve all class-level agent tunables via Proc** <br>
2026
- Let `model`, `tools`, `skills`, `schema`, `stream`, and `tracer`
2027
- declared with a block be lazily evaluated against the agent instance
2028
- at initialization time, matching how `stream` and `tracer` already
2029
- worked.
2030
-
2031
- Add `LLM::Agent#params` for direct access to the underlying context
2032
- parameters.
2033
-
2034
- Ported from mruby-llm.
2035
-
2036
- * **Support `Array[...]` schema and tool param types** <br>
2037
- Let `LLM::Schema` properties and `LLM::Tool` params accept
2038
- `Array[...]` type declarations, including mixed item unions that are
2039
- serialized as `anyOf` array items.
2040
-
2041
- * **Add `LLM::Provider#key?`** <br>
2042
- Add `key?` to providers so callers can check whether a non-blank API
2043
- key has been configured.
2044
-
2045
- * **Add agent tool confirmation hooks** <br>
2046
- Add `LLM::Agent.confirm` and `LLM::Agent#on_tool_confirmation` so
2047
- selected tools can be approved or cancelled before execution. Pending
2048
- tool resolution now relies on `LLM::Context#functions` so confirmed
2049
- tools are not executed twice when mixed with unconfirmed tool calls.
2050
-
2051
- * **Add `LLM::Function#spawn(:call).wait`** <br>
2052
- Add task-shaped sequential execution support for direct
2053
- `LLM::Function#spawn(:call).wait`.
2054
-
2055
- ### Fix
2056
-
2057
- * **Reduce private internal methods on `LLM::Stream`** <br>
2058
- Remove `tool_not_found` and `__tools__` from `LLM::Stream`. The
2059
- `__tools__` logic is inlined directly into `__find__` since that
2060
- was its only caller. The `tool_not_found` utility method was unused
2061
- externally and added unnecessary surface to LLM::Stream.
2062
-
2063
- Ported from mruby-llm.
2064
-
2065
- ## v9.0.0
2066
-
2067
- Changes since `v8.1.0`.
2068
-
2069
- This release deepens llm.rb's transport and cost-tracking surface. It
2070
- replaces the old mutable `persist!` API with constructor-driven transport
2071
- selection, removes `#call` from contexts and agents in favor of explicit
2072
- `ctx.wait(:call)`, makes queued stream waits strategy-free, and deletes
2073
- the unused `LLM::Utils` module.
2074
-
2075
- It adds cache read/write token tracking
2076
- with corresponding cost components, audio and image token pricing,
2077
- `LLM::Context#functions?` for queue-aware tool loops,
2078
- `LLM::Agent.stream` DSL support, and exposes `#stream` readers on
2079
- contexts and agents.
2080
-
2081
- The HTTP transport layer has been refactored around shared backends so
2082
- providers, MCP, and custom transports all use the same normalized
2083
- response interface.
2084
-
2085
- ### Breaking
2086
-
2087
- * **Remove `#call` as a context and agent tool-loop API** <br>
2088
- Remove `LLM::Context#call(:functions)` and `LLM::Agent#call(:functions)`.
2089
- Tool loops should use `ctx.wait(:call)` or `agent.wait(:call)` instead.
2090
- The ActiveRecord and Sequel wrappers no longer expose `#call` passthroughs
2091
- for stored llm.rb contexts.
2092
-
2093
- * **Make HTTP transport selection constructor-driven** <br>
2094
- Remove public `persist!` and `.persistent` mutation APIs from
2095
- providers, transports, and MCP clients. Select persistent behavior at
2096
- construction time with `persistent: true`, `LLM::Transport.net_http`,
2097
- `LLM::Transport.net_http_persistent`, or an explicit `transport:`
2098
- override.
2099
-
2100
- * **Make queued stream waits strategy-free** <br>
2101
- Change `LLM::Stream::Queue#wait` to resolve queued work by the actual
2102
- task types already present in the queue instead of accepting an
2103
- external wait strategy. `LLM::Stream#wait(...)` remains compatible but
2104
- now ignores its arguments when delegating to the queue.
2105
-
2106
- * **Remove unused `LLM::Utils`** <br>
2107
- Delete the `LLM::Utils` module and remove its remaining unused
2108
- provider includes and top-level require.
2109
-
2110
- ### Add
2111
-
2112
- * **Expose `#stream` readers on contexts and agents** <br>
2113
- Add public `LLM::Context#stream` and `LLM::Agent#stream` accessors so
2114
- callers can inspect the active stream object directly.
2115
-
2116
- * **Track cache read and write tokens in usage** <br>
2117
- Add `cache_read_tokens` and `cache_write_tokens` to `LLM::Usage` and
2118
- preserve them through completion usage adaptation and context usage
2119
- aggregation.
2120
-
2121
- * **Add `LLM::Context#functions?` for queue-aware tool loops** <br>
2122
- Add `functions?` to `LLM::Context` and the ActiveRecord and Sequel
2123
- wrappers so callers can detect pending tool work through either the
2124
- bound stream queue or unresolved functions, and update the docs to
2125
- prefer `while ctx.functions?` over `ctx.functions.any?` in tool-loop
2126
- examples.
2127
-
2128
- * **Add `:call` as a first-class wait strategy** <br>
2129
- Add `:call` to pending-function wait paths so `ctx.wait(:call)` can
2130
- prefer queued streamed work when present and otherwise fall back to
2131
- direct sequential function execution through `spawn(:call).wait`.
2132
-
2133
- * **Read provider cache usage into completion responses** <br>
2134
- Read cache read tokens from provider usage metadata, including OpenAI
2135
- `usage.prompt_tokens_details` and Anthropic
2136
- `usage.cache_read_input_tokens`. Read Anthropic cache write tokens
2137
- from `usage.cache_creation_input_tokens`, and expose explicit
2138
- zero-valued `cache_write_tokens` methods on providers that do not
2139
- report cache creation usage.
2140
-
2141
- * **Extend cost tracking with cache write pricing** <br>
2142
- Extend `LLM::Cost` with `cache_read_costs`, `cache_write_costs`, and
2143
- `reasoning_costs` alongside the existing `input_costs` and
2144
- `output_costs`. Add `#to_h` for structured cost insight and update
2145
- `ctx.cost` to calculate all available components from registry
2146
- pricing data.
2147
-
2148
- * **Price input and output audio separately** <br>
2149
- Track `input_audio_tokens` and `output_audio_tokens` in usage and
2150
- include `input_audio_costs` and `output_audio_costs` in `LLM::Cost`
2151
- so multimodal requests report accurate audio spend.
2152
-
2153
- * **Track image tokens in input cost reporting** <br>
2154
- Add `input_image_tokens` to usage and include `input_image_costs` in
2155
- `LLM::Cost` using the model's generic input rate so image-bearing
2156
- prompts report their input spend.
2157
-
2158
- * **Add `LLM::Agent.stream` DSL support** <br>
2159
- Let agents define a default `stream` through the class DSL, including
2160
- block-based stream construction so each agent instance can resolve its
2161
- stream the same way `tracer` does.
2162
-
2163
- ### Change
2164
-
2165
- * **Refactor HTTP transports around shared backends** <br>
2166
- Split `Net::HTTP` and `Net::HTTP::Persistent` into separate
2167
- `LLM::Transport` implementations, move HTTP-specific request helpers
2168
- and response execution into the shared transport layer, and let MCP
2169
- HTTP wrap those transports instead of maintaining a separate
2170
- transient/persistent client split.
2171
-
2172
- * **Share transport overrides across providers and MCP** <br>
2173
- Let both provider construction and `LLM::MCP.http(...)` accept
2174
- `LLM::Transport` instances or classes as HTTP transport overrides, so
2175
- callers can reuse the same transport implementation across the
2176
- runtime.
2177
-
2178
- * **Let custom transports adapt their own response objects** <br>
2179
- Introduce a transport response interface so custom transports can
2180
- adapt backend-specific response objects to one normalized shape and
2181
- have them work with the existing provider execution and error-handling
2182
- code.
2183
-
2184
- ## v8.1.0
2185
-
2186
- Changes since `v8.0.0`.
2187
-
2188
- This release adds Amazon Bedrock provider support through the Converse
2189
- API, including AWS SigV4 request signing, event stream decoding,
2190
- structured output through `schema:`, and a models.dev-backed registry.
2191
- It exposes `llm.models.all` for Bedrock via the ListFoundationModels
2192
- API and adds `LLM::Object#transform_values!` for in-place value
2193
- transformation. Several Bedrock-specific fixes land as well, including
2194
- response id exposure, blank text block suppression in tool turns, and
2195
- DSML tool-marker filtering in streamed text.
2196
-
2197
- ### Add
2198
-
2199
- * **Add AWS Bedrock provider support** <br>
2200
- Add `LLM.bedrock(...)` with Bedrock Converse chat support, AWS SigV4
2201
- request signing, Bedrock event stream decoding, structured output
2202
- support through `schema:`, and models.dev-backed `bedrock.json`
2203
- registry generation.
2204
-
2205
- * **Add AWS Bedrock Models endpoint support** <br>
2206
- Add `llm.models.all` for Bedrock via the ListFoundationModels API,
2207
- including SigV4 signing for the control-plane endpoint and normalized
2208
- `LLM::Model` collection responses.
2209
-
2210
- * **Add `LLM::Object#transform_values!`** <br>
2211
- Let `LLM::Object` transform stored values in place through
2212
- `#transform_values!`.
2213
-
2214
- ### Fix
2215
-
2216
- * **Expose response ids on Bedrock completion responses** <br>
2217
- Read the Bedrock request id into `LLM::Response#id` for completion
2218
- responses adapted from the Converse API.
2219
-
2220
- * **Avoid blank assistant text blocks in Bedrock tool turns** <br>
2221
- Stop replaying assistant tool-call messages with empty text content
2222
- blocks that Bedrock rejects.
2223
-
2224
- * **Suppress Bedrock DSML tool markers in streamed text** <br>
2225
- Filter `"\u003c\u003cDSML\u003efunction_calls\u003e\u003e"` markers out of streamed Bedrock
2226
- assistant text so tool-call sentinels do not leak into user-visible
2227
- output.
2228
-
2229
- ## v8.0.0
2230
-
2231
- Changes since `v7.0.0`.
2232
-
2233
- This release adds Unix-fork concurrency for process-isolated tool
2234
- execution, extends `LLM::Object` with `#merge` and `#delete`, and drops
2235
- Ruby 3.2 support due to a segfault observed with the `:fork` path. It
2236
- promotes `LLM::Pipe` to the top-level namespace and adds
2237
- `persistent: true` on `LLM::MCP.http` for direct persistent transport
2238
- configuration. `LLM::Function#runner` is exposed as public API, agent
2239
- tracer overrides are supported, fiber execution now uses `Fiber.schedule`,
2240
- missing optional dependencies raise clearer `LLM::LoadError` guidance,
2241
- and ActiveRecord wrapper plumbing is deduplicated between `acts_as_llm`
2242
- and `acts_as_agent`.
2243
-
2244
- ### Breaking
2245
-
2246
- * **Drop Ruby 3.2 support** <br>
2247
- Stop supporting Ruby 3.2 due to a segfault observed with the `:fork`
2248
- tool concurrency strategy.
2249
-
2250
- ### Add
2251
-
2252
- * **Add `LLM::Object#merge`** <br>
2253
- Let `LLM::Object` return a new wrapped object when merging hash-like
2254
- data through `#merge`.
2255
-
2256
- * **Add `LLM::Object#delete`** <br>
2257
- Let `LLM::Object` delete keys directly through `#delete`.
2258
-
2259
- ### Change
2260
-
2261
- * **Add fork-based tool concurrency** <br>
2262
- Add `:fork` as a new concurrency strategy for `LLM::Function#spawn`,
2263
- `LLM::Function::Array#wait`, and `LLM::Agent.concurrency` that runs
2264
- class-based tools in isolated child processes. Fork-backed tools support
2265
- tracer callbacks, `on_interrupt`/`on_cancel` hooks, and `alive?` checks.
2266
- Requires the `xchan` gem for inter-process communication with `:fork`.
2267
- This is especially useful for tools that need process isolation, such as
2268
- running shell commands or handling unsafe data.
2269
-
2270
- * **Promote `LLM::Pipe` from MCP namespace to top-level** <br>
2271
- Move `LLM::MCP::Pipe` to `LLM::Pipe` so the pipe abstraction is available
2272
- outside MCP internals. The new class adds a `binmode:` option for binary
2273
- pipes. `LLM::MCP::Command` and related MCP transport code have been updated
2274
- to use `LLM::Pipe`.
2275
-
2276
- * **Allow `persistent: true` on `LLM::MCP.http`** <br>
2277
- Let `LLM::MCP.http(...)` enable persistent HTTP transport directly
2278
- through `persistent: true` at construction time.
2279
-
2280
- * **Expose `LLM::Function#runner` as public API** <br>
2281
- Promote the internal runner instantiation to a public `runner` method on
2282
- `LLM::Function`, so callers can inspect or reuse the resolved tool instance
2283
- that a function wraps.
2284
-
2285
- * **Allow agent instance tracer overrides** <br>
2286
- Let `LLM::Agent.new(..., tracer: ...)` override the class-level tracer
2287
- for that agent instance.
2288
-
2289
- * **Make `:fiber` use scheduler-backed fibers** <br>
2290
- Change `:fiber` tool execution to use `Fiber.schedule` and require
2291
- `Fiber.scheduler`, instead of wrapping direct calls in raw fibers. This
2292
- gives `:fiber` a real cooperative concurrency model instead of acting as
2293
- a thin wrapper around sequential execution.
2294
-
2295
- * **Read stored values from zero-argument `LLM::Object` method calls** <br>
2296
- Let calls like `obj.delete`, `obj.fetch`, `obj.merge`, `obj.key?`,
2297
- `obj.dig`, `obj.slice`, or `obj.keys` return a stored value when that
2298
- method name exists as a key and no arguments are given.
2299
-
2300
- * **Harden `LLM::Object` against arbitrary key names** <br>
2301
- Move internal lookup logic off `LLM::Object` instances and onto the
2302
- singleton class instead, making stored keys like `method_missing`
2303
- more resilient while preserving normal dynamic field access.
2304
-
2305
- * **Deduplicate ActiveRecord wrapper plumbing** <br>
2306
- Move shared ActiveRecord wrapper defaults and utility methods into
2307
- `LLM::ActiveRecord`, reducing duplication between `acts_as_llm` and
2308
- `acts_as_agent`.
2309
-
2310
- * **Raise clearer errors for missing optional runtime dependencies** <br>
2311
- Route optional `async`, `xchan`, and `net/http/persistent` loads
2312
- through `LLM.require` so missing runtime gems raise `LLM::LoadError`
2313
- with installation guidance instead of leaking raw `LoadError`
2314
- exceptions.
2315
-
2316
- ### Fix
2317
-
2318
- * **Avoid `RuntimeError` from `Async::Task.current` lookups** <br>
2319
- Check `Async::Task.current?` before reading the current Async task so
2320
- provider transports fall back to `Fiber.current` without raising when
2321
- no Async task is active.
2322
-
2323
- * **Serialize `LLM::Object` values correctly through `LLM.json`** <br>
2324
- Make `LLM::Object#to_json` call `LLM.json.dump(to_h, ...)` so
2325
- `LLM::Object` values serialize through the llm.rb JSON adapter.
2326
-
2327
- ## v7.0.0
2328
-
2329
- Changes since `v6.1.0`.
2330
-
2331
- This release turns agent tool-loop limit errors into in-band advisory
2332
- returns so the LLM can react to rate limits and continue the loop. It
2333
- adds `tool_attempts: nil` as a way to opt out of advisory tool-limit
2334
- returns entirely, and fixes the default provider HTTP path to keep
2335
- `net-http-persistent` optional when not explicitly enabled.
2336
-
2337
- ### Breaking
2338
-
2339
- * **Return in-band tool-loop limit errors from agents** <br>
2340
- Stop raising `LLM::ToolLoopError` when an agent exhausts its tool loop
2341
- attempt budget, and instead send advisory `LLM::Function::Return`
2342
- errors back through the model so the LLM can react to the rate limit
2343
- in-band and continue the loop.
2344
-
2345
- * **Allow `tool_attempts: nil` to disable advisory tool-limit returns** <br>
2346
- Keep the default `tool_attempts` budget at `25`, but treat an explicit
2347
- `tool_attempts: nil` as an opt-out that disables advisory tool-limit
2348
- returns entirely.
2349
-
2350
- ### Fix
2351
-
2352
- * **Keep `net-http-persistent` optional on normal HTTP requests** <br>
2353
- Stop the default provider HTTP path from loading `net/http/persistent`
2354
- unless persistent transport support is explicitly enabled.
2355
-
2356
- ## v6.1.0
2357
-
2358
- Changes since `v6.0.0`.
2359
-
2360
- This release tightens interrupt and compaction behavior for long-running
2361
- contexts. It adds `LLM::Buffer#rindex`, supports percentage-based token
2362
- thresholds in `LLM::Compactor`, tracks persisted compaction state through
2363
- context serialization, reliably interrupts Async-backed requests, preserves
2364
- valid tool-call history on cancellation, keeps concurrent skill tool loops
2365
- running on streamed agents, and returns zero-valued usage objects when no
2366
- provider usage has been recorded yet.
2367
-
2368
- ### Change
2369
-
2370
- * **Add `LLM::Buffer#rindex`** <br>
2371
- Add `LLM::Buffer#rindex` as a direct forward to the underlying message
2372
- array so callers can find the last matching message index through the
2373
- buffer API.
2374
-
2375
- * **Support percentage compaction token thresholds** <br>
2376
- Let `LLM::Compactor` accept `token_threshold:` values like `"90%"` so
2377
- compaction can trigger at a percentage of the active model context
2378
- window.
2379
-
2380
- ### Fix
2381
-
2382
- * **Interrupt Async-backed requests reliably** <br>
2383
- Track request ownership through the provider transport so contexts use
2384
- the active Async task when available, letting `ctx.interrupt!`
2385
- reliably cancel streamed requests under Async runtimes and surface
2386
- them as `LLM::Interrupt`.
2387
-
2388
- * **Preserve valid tool-call history on cancellation** <br>
2389
- Append cancelled tool-return messages for unresolved tool calls during
2390
- `ctx.interrupt!` so follow-up provider requests do not fail with
2391
- invalid tool-call history after pending tool work is cancelled.
2392
-
2393
- * **Preserve concurrent skill tool loops on streamed agents** <br>
2394
- Propagate the active agent concurrency through the effective request
2395
- stream so nested skill agents keep using queued `wait(...)` tool
2396
- execution instead of falling back to direct `:call` execution.
2397
-
2398
- * **Track persisted compaction state on contexts** <br>
2399
- Mark contexts as compacted after `LLM::Compactor#compact!`, persist and
2400
- restore that state through context serialization, and clear it after the
2401
- next successful model response.
2402
-
2403
- * **Return zero-valued usage objects from contexts** <br>
2404
- Make `LLM::Context#usage` consistently return an `LLM::Object`, using a
2405
- zero-valued usage object when no provider usage has been recorded yet.
2406
-
2407
- ## v6.0.0
2408
-
2409
- Changes since `v5.4.0`.
2410
-
2411
- This release simplifies the ORM persistence contract around serialized
2412
- `data` state, removing the assumption of reserved `provider`, `model`, and
2413
- usage columns. Provider selection must now come from `provider:` hooks,
2414
- model defaults come from `context:` or agent DSL, and usage is read from the
2415
- serialized runtime state. Alongside this breaking change, Sequel JSON and
2416
- JSONB persistence is fixed, ractor-backed tools now fire tracer callbacks,
2417
- and `LLM::RactorError` is raised for unsupported ractor tool work.
2418
-
2419
- ### Change
2420
-
2421
- * **Simplify ORM persistence to serialized `data` state** <br>
2422
- Change the built-in ActiveRecord and Sequel wrappers to treat serialized
2423
- `data` as the persistence contract, instead of assuming reserved
2424
- `provider`, `model`, and usage columns. Provider selection must now come
2425
- from `provider:` hooks that resolve a real `LLM::Provider` instance, model
2426
- defaults come from `context:` or agent DSL, and `usage` is read from the
2427
- serialized runtime state.
2428
-
2429
- ### Fix
2430
-
2431
- * **Fix Sequel JSON and JSONB persistence** <br>
2432
- Load Sequel PostgreSQL JSON support when `plugin :llm` is configured with
2433
- `format: :json` or `:jsonb`, and wrap structured payloads correctly so
2434
- persisted context state can be stored in PostgreSQL JSON columns.
2435
-
2436
- * **Trace ractor-backed tool callbacks** <br>
2437
- Make tool tracers fire `on_tool_start` and `on_tool_finish` for
2438
- class-based `:ractor` execution too, so ractor-backed tool calls show up
2439
- in tracer callbacks like the other concurrent tool paths.
2440
-
2441
- * **Raise `LLM::RactorError` for unsupported ractor tool work** <br>
2442
- Add `LLM::RactorError` and fail fast when `:ractor` execution is requested
2443
- for unsupported tool types such as skill-backed tools, instead of letting
2444
- deeper Ruby isolation errors leak out later in execution.
2445
-
2446
- * **Delegate interrupt to concurrent task implementations** <br>
2447
- Make `LLM::Function::Task#interrupt!` delegate to the underlying fork or
2448
- ractor task when it supports interruption, so `ctx.interrupt!` and
2449
- `task.interrupt!` work correctly for fork- and ractor-backed tool
2450
- execution.
2451
-
2452
- ## v5.4.0
2453
-
2454
- Changes since `v5.3.0`.
2455
-
2456
- This release expands tracer support around agentic execution. It lets
2457
- `LLM::Agent` define scoped tracers through the agent DSL and fixes concurrent
2458
- tool execution so those scoped tracers stay attached when work crosses
2459
- thread, task, fiber, and skill boundaries.
2460
-
2461
- ### Change
2462
-
2463
- * **Add agent-scoped tracers** <br>
2464
- Let `LLM::Agent` classes define `tracer ...` or `tracer { ... }` so an
2465
- agent can carry its own tracer without replacing the provider's default
2466
- tracer. The resolved tracer is scoped to that agent's turns, tool loops,
2467
- and pending tool access. Available through the `acts_as_agent` and Sequel
2468
- agent plugin `tracer` DSL too.
2469
-
2470
- ### Fix
2471
-
2472
- * **Preserve scoped tracers across concurrent tool work** <br>
2473
- Keep agent- and request-scoped tracers attached when tool execution
2474
- crosses `:thread`, `:task`, or `:fiber` boundaries, including skill
2475
- execution, so spawned work does not fall back to the provider default
2476
- tracer.
2477
-
2478
- ## v5.3.0
2479
-
2480
- Changes since `v5.2.1`.
2481
-
2482
- This release deepens llm.rb's request-rewriting and tool-definition surface.
2483
- It adds transformer lifecycle hooks to `LLM::Stream` so UIs can surface work
2484
- like PII scrubbing before a request is sent, and it adds a more explicit
2485
- OmniAI-style tool DSL form with `parameter` plus separate `required`
2486
- declarations while keeping the older `param ... required: true` style working.
2487
-
2488
- ### Change
2489
-
2490
- * **Add transformer stream lifecycle hooks** <br>
2491
- Add `on_transform` and `on_transform_finish` to
2492
- `LLM::Stream` so UIs can surface request rewriting work such as PII
2493
- scrubbing before a request is sent to the model.
2494
-
2495
- * **Add a separate `required` tool DSL form** <br>
2496
- Add `parameter` as an alias of `param` and support `required %i[...]`
2497
- as a separate declaration, inspired by OmniAI-style tools, while keeping
2498
- the existing `param ... required: true` form working too.
2499
-
2500
- ## v5.2.1
2501
-
2502
- Changes since `v5.2.0`.
2503
-
2504
- This release tightens the streamed queue fix from `v5.2.0` for concurrent
2505
- workloads. Request-local streams now stay bound long enough for `wait` to
2506
- drain queued work and then clear cleanly so later waits fall back to the
2507
- context's configured stream.
2508
-
2509
- ### Fix
2510
-
2511
- * **Reset request-local streams after `wait` drains queued work** <br>
2512
- Keep per-call `stream:` bindings alive through `LLM::Context#wait` so
2513
- queued streamed tool work still resolves correctly, then clear the
2514
- request-local stream after the wait completes to avoid leaking it into
2515
- later turns.
2516
-
2517
- ## v5.2.0
2518
-
2519
- Changes since `v5.1.0`.
2520
-
2521
- This release adds current DeepSeek V4 support through refreshed provider
2522
- metadata, including `deepseek-v4-flash` and `deepseek-v4-pro`, while fixing
2523
- request-local queue handling for concurrent streamed workloads so `wait` and
2524
- interruption use the active per-call stream correctly.
2525
-
2526
- ### Change
2527
-
2528
- * **Add `LLM::MCP#run` for scoped MCP client lifecycle** <br>
2529
- Add `LLM::MCP#run` so MCP clients can be started for the duration of a
2530
- block and then stopped automatically, which simplifies the usual
2531
- `start`/`stop` pattern in examples and application code.
2532
-
2533
- * **Refresh provider model metadata** <br>
2534
- Add current DeepSeek and OpenAI model metadata to `data/` and update the
2535
- Google Gemini model entry to match the current provider naming.
2536
-
2537
- ### Fix
2538
-
2539
- * **Reject unsupported DeepSeek multimodal prompt objects early** <br>
2540
- Raise `LLM::PromptError` for `image_url`, `local_file`, and
2541
- `remote_file` in DeepSeek chat requests instead of sending invalid
2542
- OpenAI-compatible payloads that the provider rejects at runtime.
2543
-
2544
- * **Preserve DeepSeek reasoning content across tool turns** <br>
2545
- Replay `reasoning_content` when serializing prior assistant messages for
2546
- DeepSeek chat completions, so thinking-mode tool calls can continue into
2547
- follow-up requests without triggering invalid request errors.
2548
-
2549
- * **Default DeepSeek to `deepseek-v4-flash`** <br>
2550
- Change `LLM::DeepSeek#default_model` to `deepseek-v4-flash` so new
2551
- contexts and default provider usage align with the current preferred chat
2552
- model.
2553
-
2554
- * **Use per-call streams when waiting on streamed tool work** <br>
2555
- Track request-local streams bound through `talk(..., stream:)` and
2556
- `respond(..., stream:)` so `LLM::Context#wait` and interruption-aware
2557
- queue handling use the active stream instead of falling back to pending
2558
- function spawning.
2559
-
2560
- ## v5.1.0
2561
-
2562
- Changes since `v5.0.0`.
2563
-
2564
- This release tightens streamed tool execution around the actual request-local
2565
- runtime state. It fixes streamed resolution of per-request tools and makes
2566
- that streamed path work cleanly with `LLM.function(...)`, MCP tools, bound
2567
- tool instances, and normal tool classes.
2568
-
2569
- ### Fix
2570
-
2571
- * **Resolve request-local tools during streaming** <br>
2572
- Resolve streamed tool calls through `LLM::Stream` request-local tools
2573
- before falling back to the global registry, so per-request tools and bound
2574
- tool instances work correctly during streaming.
2575
-
2576
- * **Support `LLM.function(...)` and MCP tools in streamed tool resolution** <br>
2577
- Let streamed tool resolution use the current request tool set, so
2578
- `LLM.function(...)`, MCP tools, bound tool instances, and normal
2579
- `LLM::Tool` classes all work through the same streamed tool path.
2580
-
2581
- ## v5.0.0
2582
-
2583
- Changes since `v4.23.0`.
2584
-
2585
- This release expands llm.rb from an execution runtime into a more explicit
2586
- supervision and transformation runtime. It adds context-level guards,
2587
- transformers, and loop supervision through `LLM::LoopGuard`, while deepening
2588
- long-lived context behavior through compaction, interruption hooks, and
2589
- streamed `ctx.spawn(...)` tool execution.
2590
-
2591
- ### Change
2592
-
2593
- * **Make compactor thresholds explicit** <br>
2594
- Require `message_threshold:` and `token_threshold:` to be opted into
2595
- explicitly, so `LLM::Compactor` only compacts automatically when one of
2596
- those thresholds is configured. Context-window-derived token limits can be
2597
- computed by the caller when needed.
2598
-
2599
- * **Allow assigning a compactor through `LLM::Context`** <br>
2600
- Let `LLM::Context` accept `ctx.compactor = ...` in addition to the
2601
- constructor `compactor:` option, so compactor config can be assigned or
2602
- replaced after context initialization.
2603
-
2604
- * **Mark compaction summaries in message metadata** <br>
2605
- Mark compaction summaries with `extra[:compaction]` and
2606
- `LLM::Message#compaction?`, so applications can detect or hide synthetic
2607
- summary messages in conversation history.
2608
-
2609
- * **Add cooperative tool interruption hooks** <br>
2610
- Let `ctx.interrupt!` notify queued tool work through `on_interrupt`, so
2611
- running tools can clean up cooperatively when a context is cancelled.
2612
-
2613
- * **Add `LLM::Context` guards** <br>
2614
- Add a new `guard` capability to `LLM::Context` so execution can be
2615
- supervised at the runtime level. The built-in `LLM::LoopGuard` detects
2616
- repeated tool-call patterns and stops stuck agentic loops through in-band
2617
- `LLM::GuardError` returns. `LLM::Agent` enables this guard by default.
2618
-
2619
- * **Add `LLM::Context` transformers** <br>
2620
- Add a new `transformer` capability to `LLM::Context` so prompts and params
2621
- can be rewritten before provider requests are sent. This makes it possible
2622
- to apply context-wide behaviors such as PII scrubbing or request-level
2623
- param injection without rewriting every `talk` and `respond` call site.
2624
-
2625
- ## v4.23.0
2626
-
2627
- Changes since `v4.22.0`.
2628
-
2629
- This release expands llm.rb's runtime surface for long-lived contexts and
2630
- stateful tools. It adds built-in context compaction through `LLM::Compactor`,
2631
- lets explicit `tools:` arrays accept bound `LLM::Tool` instances, and fixes
2632
- OpenAI-compatible no-arg tool schemas for stricter providers such as xAI.
2633
-
2634
- ### Change
2635
-
2636
- * **Add `LLM::Compactor` for long-lived contexts** <br>
2637
- Add built-in context compaction through `LLM::Compactor`, so older history
2638
- can be summarized, retained windows can stay bounded, compaction can run on
2639
- its own `model:`, thresholds can be configured explicitly, and
2640
- `LLM::Stream` can observe the lifecycle through `on_compaction` and
2641
- `on_compaction_finish`.
2642
-
2643
- * **Allow bound tool instances in explicit tool lists** <br>
2644
- Let explicit `tools:` arrays accept `LLM::Tool` instances such as
2645
- `MyTool.new(foo: 1)`, so tools can carry bound state without changing the
2646
- global tool registry model.
2647
-
2648
- ### Fix
2649
-
2650
- * **Fix xAI/OpenAI-compatible no-arg tool schemas** <br>
2651
- Send an empty object schema for tools without declared parameters instead
2652
- of `null`, so stricter providers such as xAI accept mixed tool sets that
2653
- include no-arg tools.
2654
-
2655
- ## v4.22.0
2656
-
2657
- Changes since `v4.21.0`.
2658
-
2659
- This release deepens the runtime shape of llm.rb. It reduces helper-method
2660
- surface on persisted ORM models, expands real ORM coverage, and makes skills
2661
- behave more like bounded sub-agents with inherited recent context and proper
2662
- instruction injection.
2663
-
2664
- ### Change
2665
-
2666
- * **Reduce ActiveRecord wrapper model surface** <br>
2667
- Move helper methods such as option resolution, column mapping,
2668
- serialization, and persistence into `Utils` for the ActiveRecord
2669
- wrappers so wrapped models include fewer internal helper methods.
2670
-
2671
- * **Reduce Sequel wrapper model surface** <br>
2672
- Move helper methods such as option resolution, column mapping,
2673
- serialization, and persistence into `Utils` for the Sequel wrappers
2674
- so wrapped models include fewer internal helper methods.
2675
-
2676
- * **Expand ORM integration coverage** <br>
2677
- Add broader ActiveRecord and Sequel coverage for persisted context and
2678
- agent wrappers, including real SQLite-backed records and cassette-backed
2679
- OpenAI persistence paths.
2680
-
2681
- * **Make skills inherit recent parent context** <br>
2682
- Run `LLM::Skill` with a curated slice of recent parent user and assistant
2683
- messages, prefixed with `Recent context:`, so skills behave more like
2684
- task-scoped sub-agents instead of instruction-only helpers.
2685
-
2686
- ### Fix
2687
-
2688
- * **Fix Sequel `plugin :agent` load order** <br>
2689
- Require the shared Sequel plugin support from `LLM::Sequel::Agent` so
2690
- `plugin :agent` can load independently without raising
2691
- `uninitialized constant LLM::Sequel::Plugin`.
2692
-
2693
- * **Make skill execution inherit parent context request settings** <br>
2694
- Run `LLM::Skill` through a parent `LLM::Context` instead of a bare
2695
- provider so nested skill agents inherit context-level settings such as
2696
- `mode: :responses`, `store: false`, streaming, and other request defaults,
2697
- while still keeping skill-local tools and avoiding parent schemas.
2698
-
2699
- * **Keep agent instructions when history is preseeded** <br>
2700
- Inject `LLM::Agent` instructions once unless a system message is already
2701
- present, so agents and nested skills still get their instructions when
2702
- they start with inherited non-system context.
2703
-
2704
- ## v4.21.0
2705
-
2706
- Changes since `v4.20.2`.
2707
-
2708
- This release expands higher-level composition in llm.rb. It adds Sequel agent
2709
- persistence through `plugin :agent` and introduces directory-backed skills
2710
- that load from `SKILL.md`, resolve named tools, and plug directly into
2711
- `LLM::Context` and `LLM::Agent`.
2712
-
2713
- ### Change
2714
-
2715
- * **Add `plugin :agent` for Sequel models** <br>
2716
- Add Sequel support for `plugin :agent`, similar to ActiveRecord's
2717
- `acts_as_agent`, so models can wrap `LLM::Agent` with built-in
2718
- persistence.
2719
-
2720
- * **Load directory-backed skills through `LLM::Context` and `LLM::Agent`** <br>
2721
- Add `skills:` to `LLM::Context` and `skills ...` to `LLM::Agent` so
2722
- directories with `SKILL.md` can be loaded, resolved into tools, and run
2723
- through the normal llm.rb tool path.
2724
-
2725
- ## v4.20.2
2726
-
2727
- Changes since `v4.20.1`.
2728
-
2729
- This patch release improves runtime behavior around interruption and mixed
2730
- concurrency waits. It also rounds out response API uniformity for Google
2731
- completion responses.
2732
-
2733
- ### Fix
2734
-
2735
- * **Expose Google completion response IDs through `.id`** <br>
2736
- Add `LLM::Response#id` support to Google completion responses so tracer
2737
- and caller code can rely on the same API used by other providers.
2738
-
2739
- * **Track interrupt ownership on the active request** <br>
2740
- Bind `LLM::Context` interruption to the fiber running `talk` or `respond`
2741
- so `interrupt!` works correctly when requests are started outside the
2742
- context's initialization fiber.
2743
-
2744
- ### Change
2745
-
2746
- * **Allow mixed concurrency strategies in `wait(...)`** <br>
2747
- Let `LLM::Context#wait`, `LLM::Stream#wait`, and `LLM::Agent.concurrency`
2748
- accept arrays such as `[:thread, :ractor]` so mixed tool sets can wait on
2749
- more than one concurrency strategy.
2750
-
2751
- ## v4.20.1
2752
-
2753
- Changes since `v4.20.0`.
2754
-
2755
- This patch release fixes ORM option resolution in the Sequel and
2756
- ActiveRecord wrappers. Symbol-based `provider:` and `context:` hooks now
2757
- resolve correctly, and internal default option constants are referenced
2758
- explicitly instead of relying on nested constant lookup.
2759
-
2760
- ### Fix
2761
-
2762
- * **Fix symbol-based ORM option hooks for provider and context hashes** <br>
2763
- Make `provider:` and `context:` resolve symbol hooks through the model in
2764
- the Sequel plugin and ActiveRecord wrappers instead of falling back to an
2765
- empty hash.
2766
-
2767
- * **Fix ORM wrapper constant lookup for option defaults** <br>
2768
- Qualify internal `EMPTY_HASH` / `DEFAULTS` references in the Sequel plugin
2769
- and ActiveRecord wrappers so option resolution does not depend on nested
2770
- constant lookup quirks.
2771
-
2772
- ## v4.20.0
2773
-
2774
- Changes since `v4.19.0`.
2775
-
2776
- This release adds better support for tagged prompt content. `LLM::Context`
2777
- can now serialize and restore `image_url`, `local_file`, and `remote_file`
2778
- content cleanly, and `LLM::Message` now exposes helpers for inspecting
2779
- tagged image and file attachments.
2780
-
2781
- ### Change
2782
-
2783
- * **Round-trip tagged prompt objects through `LLM::Context`** <br>
2784
- Teach `LLM::Context` serialization and restore to preserve
2785
- `image_url`, `local_file`, and `remote_file` content across
2786
- `to_json` / `restore`.
2787
-
2788
- * **Add attachment helpers to `LLM::Message`** <br>
2789
- Add `image_url?`, `image_urls`, `file?`, and `files` so callers can
2790
- inspect messages for tagged image and file content more directly.
2791
-
2792
- ## v4.19.0
2793
-
2794
- Changes since `v4.18.0`.
2795
-
2796
- This release tightens the ActiveRecord and ORM integration layer. It adds
2797
- inline agent DSL blocks to `acts_as_agent` so agent defaults can be defined
2798
- where the wrapper is declared, and it exposes the resolved provider through
2799
- public `llm` methods on the ActiveRecord and Sequel wrappers.
2800
-
2801
- ### Change
2802
-
2803
- * **Make ORM provider access public through `llm`** <br>
2804
- Expose the resolved provider on the Sequel plugin and the ActiveRecord
2805
- `acts_as_llm` / `acts_as_agent` wrappers through a public `llm` method.
2806
-
2807
- * **Allow inline agent DSL blocks in `acts_as_agent`** <br>
2808
- Let ActiveRecord models configure `model`, `tools`, `schema`,
2809
- `instructions`, and `concurrency` directly inside the `acts_as_agent`
2810
- declaration block.
2811
-
2812
- ## v4.18.0
2813
-
2814
- Changes since `v4.17.0`.
2815
-
2816
- This release improves tracing and tool execution behavior across llm.rb.
2817
- It makes provider tracers default to the provider instance, adds
2818
- `LLM::Provider#with_tracer` for scoped overrides, restores tool tracing for
2819
- concurrent and streamed tool execution, extends streamed tracing to MCP tools,
2820
- and adds symbol-based ORM option hooks alongside experimental ractor tool
2821
- concurrency.
2822
-
2823
- ### Change
2824
-
2825
- * **Make provider tracers default to the provider instance** <br>
2826
- Change `llm.tracer = ...` so it sets a provider default tracer instead of
2827
- relying on scoped fiber-local state alone. This makes tracer configuration
2828
- behave more predictably across normal tasks, threads, and fibers that share
2829
- the same provider instance.
2830
-
2831
- * **Add `LLM::Provider#with_tracer` for scoped overrides** <br>
2832
- Add `with_tracer` as the opt-in escape hatch for request- or turn-scoped
2833
- tracer overrides. Use it when you want temporary tracing on the current
2834
- fiber without replacing the provider's default tracer.
2835
-
2836
- * **Trace concurrent tool calls outside ractors** <br>
2837
- Make tool tracing fire correctly when functions run through `:thread`,
2838
- `:task`, or `:fiber` concurrency. Experimental `:ractor` execution still
2839
- does not emit tool tracer events.
2840
-
2841
- * **Trace streamed tool calls, including MCP tools** <br>
2842
- Bind stream metadata through `LLM::Stream#extra` so streamed tool calls
2843
- inherit tracer and model context before they are handed to `on_tool_call`.
2844
- This restores tool tracing for streamed MCP and local tool execution.
2845
-
2846
- * **Support symbol-based ORM option hooks** <br>
2847
- Let `provider:`, `context:`, and `tracer:` on the Sequel plugin and
2848
- the ActiveRecord `acts_as_llm` / `acts_as_agent` wrappers resolve through
2849
- model method names as well as procs.
2850
-
2851
- * **Add experimental ractor tool concurrency** <br>
2852
- Add `:ractor` support to `LLM::Function#spawn`, `LLM::Function::Array#wait`,
2853
- `LLM::Stream#wait`, and `LLM::Agent.concurrency` so class-based tools with
2854
- ractor-safe arguments and return values can run in Ruby ractors and report
2855
- their results back into the normal LLM tool-return path. MCP tools are not
2856
- supported by the current `:ractor` mode, but mixed workloads can still
2857
- branch on `tool.mcp?` and choose a supported strategy per tool. `:ractor`
2858
- is especially useful for CPU-bound tools, while `:task`, `:fiber`, or
2859
- `:thread` may be a better fit for I/O-bound work.
2860
-
2861
- ## v4.17.0
2862
-
2863
- Changes since `v4.16.1`.
2864
-
2865
- This release expands agent support across llm.rb. It brings `LLM::Agent`
2866
- closer to `LLM::Context`, adds configurable automatic tool concurrency
2867
- including experimental ractor support for class-based tools,
2868
- extends persisted ORM wrappers with more of the context runtime surface and
2869
- tracer hooks, and introduces built-in ActiveRecord agent persistence through
2870
- `acts_as_agent`.
2871
-
2872
- ### Change
2873
-
2874
- * **Add configurable tool concurrency to `LLM::Agent`** <br>
2875
- Add the class-level `concurrency` DSL to `LLM::Agent` so automatic
2876
- tool loops can run with `:call`, `:thread`, `:task`, `:fiber`, or
2877
- experimental `:ractor` support for class-based tools instead of
2878
- always executing sequentially.
2879
-
2880
- * **Bring `LLM::Agent` closer to `LLM::Context`** <br>
2881
- Expand `LLM::Agent` so it exposes more of the same runtime surface as
2882
- `LLM::Context`, including returns, interruption, mode, cost, context
2883
- window, structured serialization, and other context-backed helpers,
2884
- while still auto-managing tool loops.
2885
-
2886
- * **Refresh agent docs and coverage** <br>
2887
- Update the README and deep dive to explain the current role of
2888
- `LLM::Agent`, add examples that show automatic tool execution and
2889
- concurrency, and add focused specs for the expanded agent surface and
2890
- tool-loop behavior.
2891
-
2892
- * **Add ORM tracer hooks for persisted contexts** <br>
2893
- Add `tracer:` to both the Sequel plugin and `acts_as_llm` so models
2894
- can resolve and assign tracers onto the provider used by their persisted
2895
- `LLM::Context`.
2896
-
2897
- * **Bring persisted ORM wrappers closer to `LLM::Context`** <br>
2898
- Expand both the Sequel plugin and `acts_as_llm` so record-backed
2899
- contexts expose more of the same runtime surface as `LLM::Context`,
2900
- including mode, returns, interruption, prompt helpers, file helpers,
2901
- and tracer access.
2902
-
2903
- * **Add ActiveRecord agent persistence with `acts_as_agent`** <br>
2904
- Add `acts_as_agent` for ActiveRecord models that should wrap
2905
- `LLM::Agent`, reusing the same record-backed runtime shape as
2906
- `acts_as_llm` while letting tool execution be managed by the agent.
2907
-
2908
- ## v4.16.1
2909
-
2910
- Changes since `v4.16.0`.
2911
-
2912
- This release tightens ORM persistence by removing an unnecessary JSON
2913
- round-trip when restoring structured `:json` and `:jsonb` context
2914
- payloads.
2915
-
2916
- ### Change
2917
-
2918
- * **Restore structured ORM payloads directly** <br>
2919
- Teach `LLM::Context#restore` to accept parsed data payloads and use
2920
- that path from the ActiveRecord and Sequel persistence wrappers for
2921
- `format: :json` and `:jsonb`, avoiding a redundant
2922
- `Hash -> JSON string -> Hash` round-trip on restore.
2923
-
2924
- ## v4.16.0
2925
-
2926
- Changes since `v4.15.0`.
2927
-
2928
- This release expands ORM support with built-in ActiveRecord persistence
2929
- and improves compatibility with OpenAI-compatible gateways, proxies, and
2930
- self-hosted servers that use non-standard API root paths.
2931
-
2932
- ### Change
2933
-
2934
- * **Support OpenAI-compatible base paths** <br>
2935
- Add `base_path:` to provider configuration so OpenAI-compatible
2936
- endpoints can vary both host and API prefix. This supports providers,
2937
- proxies, and gateways that keep OpenAI request shapes but use
2938
- non-standard URL layouts such as DeepInfra's `/v1/openai/...`.
2939
-
2940
- * **Add ActiveRecord context persistence with `acts_as_llm`** <br>
2941
- Add a built-in ActiveRecord wrapper that mirrors the Sequel plugin
2942
- API so applications can persist `LLM::Context` state on records with
2943
- default columns, provider/context hooks, validation-backed writes,
2944
- and `format: :string`, `:json`, or `:jsonb` storage.
2945
-
2946
- ## v4.15.0
2947
-
2948
- Changes since `v4.14.0`.
2949
-
2950
- ### Change
2951
-
2952
- * **Reduce OpenAI stream parser merge overhead** <br>
2953
- Special-case the most common single-field deltas, streamline
2954
- incremental tool-call merging, and avoid repeated JSON parse attempts
2955
- until streamed tool arguments look complete.
2956
-
2957
- * **Cache streaming callback capabilities in parsers** <br>
2958
- Cache callback support checks once at parser initialization time in
2959
- the OpenAI, OpenAI Responses, Anthropic, Google, and Ollama stream
2960
- parsers instead of repeating `respond_to?` checks on hot streaming
2961
- paths.
2962
-
2963
- * **Reduce OpenAI Responses parser lookup overhead** <br>
2964
- Special-case the hot Responses API event paths and cache the current
2965
- output item and content part so streamed output text deltas do less
2966
- repeated nested lookup work.
2967
-
2968
- * **Add a Sequel context persistence plugin** <br>
2969
- Add `plugin :llm` for Sequel models so apps can persist
2970
- `LLM::Context` state with default columns and pass provider setup
2971
- through `provider:` when needed. The plugin now also supports
2972
- `format: :string`, `:json`, or `:jsonb` for text and native JSON
2973
- storage when Sequel JSON typecasting is enabled.
2974
-
2975
- * **Improve streaming parser performance** <br>
2976
- In the local replay-based `stream_parser` benchmark versus `v4.14.0`
2977
- (median of 20 samples, 5000 iterations), plain Ruby is a
2978
- small overall win: the generic eventstream path is about 0.4%
2979
- faster, the OpenAI stream parser is about 0.5% faster, and the
2980
- OpenAI Responses parser is about 1.6% faster, with unchanged
2981
- allocations. Under YJIT on the same benchmark harness, the generic
2982
- eventstream path is about 0.9% faster and the OpenAI stream parser
2983
- is about 0.4% faster, while the OpenAI Responses parser is about
2984
- 0.7% slower, also with unchanged allocations.
2985
-
2986
- Compared to `v4.13.0`, the larger `v4.14.0` streaming gains still
2987
- hold. The generic eventstream path remains dramatically faster than
2988
- `v4.13.0`, the OpenAI stream parser remains modestly faster, and the
2989
- OpenAI Responses parser is roughly flat to slightly better depending
2990
- on runtime. In other words, current keeps the large eventstream win
2991
- from `v4.14.0`, adds only small incremental changes beyond that, and
2992
- does not turn the post-`v4.14.0` parser work into another large
2993
- benchmark jump.
2994
-
2995
- ## v4.14.0
2996
-
2997
- Changes since `v4.13.0`.
2998
-
2999
- This release adds request interruption for contexts, reworks provider
3000
- HTTP internals for lower-overhead streaming, and fixes MCP clients so
3001
- parallel tool calls can safely share one connection.
3002
-
3003
- ### Add
3004
-
3005
- * **Add request interruption support** <br>
3006
- Add `LLM::Context#interrupt!`, `LLM::Context#cancel!`, and
3007
- `LLM::Interrupt` for interrupting in-flight provider requests,
3008
- inspired by Go's context cancellation.
3009
-
3010
- ### Change
3011
-
3012
- * **Rework provider HTTP transport internals** <br>
3013
- Rework provider HTTP around `LLM::Provider::Transport::HTTP` with
3014
- explicit transient and persistent transport handling.
3015
-
3016
- * **Reduce SSE parser overhead** <br>
3017
- Dispatch raw parsed values to registered visitors instead of building
3018
- an `Event` object for every streamed line.
3019
-
3020
- * **Reduce provider streaming allocations** <br>
3021
- Decode streamed provider payloads directly in
3022
- `LLM::Provider::Transport::HTTP` before handing them to provider
3023
- parsers, which cuts allocation churn and gives a small streaming
3024
- speed bump.
3025
-
3026
- * **Reduce generic SSE parser allocations** <br>
3027
- Keep unread event-stream buffer data in place until compaction is
3028
- worthwhile, which lowers allocation churn in the remaining generic
3029
- SSE path.
3030
-
3031
- * **Improve streaming parser performance** <br>
3032
- In the local replay-based `stream_parser` benchmark versus `v4.13.0`
3033
- (median of 20 samples, 5000 iterations):
3034
- Plain Ruby: the generic eventstream path is about 53% faster with
3035
- about 32% fewer allocations, the OpenAI stream parser is about 11%
3036
- faster with about 4% fewer allocations, and the OpenAI Responses
3037
- parser is about 3% faster with unchanged allocations.
3038
- YJIT on the current parser benchmark harness: the current tree is
3039
- about 26% faster than non-YJIT on the generic eventstream path,
3040
- about 18% faster on the OpenAI stream parser, and about 16% faster
3041
- on the OpenAI Responses parser, with allocations unchanged.
3042
-
3043
- ### Fix
3044
-
3045
- * **Support parallel MCP tool calls on one client** <br>
3046
- Route MCP responses by JSON-RPC id so concurrent tool calls can
3047
- share one client and transport without mismatching replies.
3048
-
3049
- * **Use explicit MCP non-blocking read errors** <br>
3050
- Use `IO::EAGAINWaitReadable` while continuing to retry on
3051
- `IO::WaitReadable`.
3052
-
3053
- ## v4.13.0
3054
-
3055
- Changes since `v4.12.0`.
3056
-
3057
- This release expands MCP prompt support, improves reasoning support in the
3058
- OpenAI Responses API, and refreshes the docs around llm.rb's runtime model,
3059
- contexts, and advanced workflows.
3060
-
3061
- ### Add
3062
-
3063
- - Add `LLM::MCP#prompts` and `LLM::MCP#find_prompt` for MCP prompt support.
3064
-
3065
- ### Change
3066
-
3067
- - Rework the README around llm.rb as a runtime for AI systems.
3068
- - Add a dedicated deep dive guide for providers, contexts, persistence,
3069
- tools, agents, MCP, tracing, multimodal prompts, and retrieval.
3070
-
3071
- ### Fix
3072
-
3073
- All of these fixes apply to MCP:
3074
-
3075
- - fix(mcp): raise `LLM::MCP::MismatchError` on mismatched response ids.
3076
- - fix(mcp): normalize prompt message content while preserving the original payload.
3077
-
3078
- All of these fixes apply to OpenAI's Responses API:
3079
-
3080
- - fix(openai): emit `on_reasoning_content` for streamed reasoning summaries.
3081
- - fix(openai): skip `previous_response_id` on `store: false` follow-up calls.
3082
- - fix(openai): fall back to an empty object schema for tools without params.
3083
- - fix(openai): preserve original tool-call payloads on re-sent assistant tool messages.
3084
- - fix(openai): emit `output_text` for assistant-authored response content.
3085
- - fix(openai): return `nil` for `system_fingerprint` on normalized response objects.
3086
-
3087
- ## v4.12.0
3088
-
3089
- Changes since `v4.11.1`.
3090
-
3091
- This release expands advanced streaming and MCP execution while reframing
3092
- llm.rb more clearly as a system integration layer for LLMs, tools, MCP
3093
- sources, and application APIs.
3094
-
3095
- ### Add
3096
-
3097
- - Add `persistent` as an alias for `persist!` on providers and MCP transports.
3098
- - Add `LLM::Stream#on_tool_return` for observing completed streamed tool work.
3099
- - Add `LLM::Function::Return#error?`.
3100
-
3101
- ### Change
3102
-
3103
- - Expect advanced streaming callbacks to use `LLM::Stream` subclasses
3104
- instead of duck-typing them onto arbitrary objects. Basic `#<<`
3105
- streaming remains supported.
3106
-
3107
- ### Fix
3108
-
3109
- - Fix Anthropic tools without params by always emitting `input_schema`.
3110
- - Fix Anthropic tool-only responses to still produce an assistant message.
3111
- - Fix Anthropic tool results to use the `user` role.
3112
- - Fix Anthropic tool input normalization.
3113
-
3114
- ## v4.11.1
3115
-
3116
- Changes since `v4.11.0`.
3117
-
3118
- ### Fix
3119
-
3120
- * Cast OpenTelemetry tool-related values to strings. <br>
3121
- Otherwise they're rejected by opentelemetry-sdk as invalid attributes.
3122
-
3123
- ## v4.11.0
3124
-
3125
- Changes since `v4.10.0`.
3126
-
3127
- ### Add
3128
-
3129
- - Add `LLM::Stream` for richer streaming callbacks, including `on_content`,
3130
- `on_reasoning_content`, and `on_tool_call` for concurrent tool execution.
3131
- - Add `LLM::Stream#wait` as a shortcut for `queue.wait`.
3132
- - Add `LLM::Context#wait` as a shortcut for the configured stream's `wait`.
3133
- - Add `LLM::Context#call(:functions)` as a shortcut for `functions.call`.
3134
- - Add `LLM::Function.registry` and enhanced support for MCP tools in
3135
- `LLM::Tool.registry` for tool resolution during streaming.
3136
- - Add normalized `LLM::Response` for OpenAI Responses, providing `content`,
3137
- `content!`, `messages` / `choices`, `usage`, and `reasoning_content`.
3138
- - Add `mode: :responses` to `LLM::Context` for routing `talk` through the
3139
- Responses API.
3140
- - Add `LLM::Context#returns` for collecting pending tool returns from the context.
3141
- - Add persistent HTTP connection pooling for repeated MCP tool calls via
3142
- `LLM.mcp(http: ...).persist!`.
3143
- - Add explicit MCP transport constructors via `LLM::MCP.stdio(...)` and
3144
- `LLM::MCP.http(...)`.
3145
-
3146
- ### Fix
3147
-
3148
- - Fix Google tool-call handling by synthesizing stable ids when Gemini does
3149
- not provide a direct tool-call id.
3150
-
3151
- ## v4.10.0
3152
-
3153
- Changes since `v4.9.0`.
3154
-
3155
- ### Add
3156
-
3157
- - Add HTTP transport for MCP with `LLM::MCP::Transport::HTTP` for remote servers
3158
- - Add JSON Schema union types (`any_of`, `all_of`, `one_of`) with parser integration
3159
- - Add JSON Schema type array union support (e.g., `"type": ["object", "null"]`)
3160
- - Add JSON Schema type inference from `const`, `enum`, or `default` fields
3161
-
3162
- ### Change
3163
-
3164
- - Update `LLM::MCP` constructor for exclusive `http:` or `stdio:` transport
3165
- - Update `LLM::MCP` documentation for HTTP transport support
3166
-
3167
- ## v4.9.0
3168
-
3169
- Changes since `v4.8.0`.
3170
-
3171
- ### Add
3172
-
3173
- - Add fiber-based concurrency with `LLM::Function::FiberGroup` and
3174
- `LLM::Function::TaskGroup` classes for lightweight async execution.
3175
- - Add `:thread`, `:task`, and `:fiber` strategy parameter to
3176
- `LLM::Function#spawn` for explicit concurrency control.
3177
- - Add stdio MCP client support, including remote tool discovery and
3178
- invocation through `LLM.mcp`, `LLM::Context`, and existing function/tool
3179
- APIs.
3180
- - Add model registry support via `LLM::Registry`, including model
3181
- metadata lookup, pricing, modalities, limits, and cost estimation.
3182
- - Add context access to a model context window via
3183
- `LLM::Context#context_window`.
3184
- - Add tracking of defined tools in the tool registry.
3185
- - Add `LLM::Schema::Enum`, enabling `Enum[...]` as a schema/tool
3186
- parameter type.
3187
- - Add top-level Anthropic system instruction support using Anthropic's
3188
- provider-specific request format.
3189
- - Add richer tracing hooks and extra metadata support for
3190
- LangSmith/OpenTelemetry-style traces.
3191
- - Add rack/websocket and Relay-related example work, including MCP-focused
3192
- examples.
3193
- - Add concurrent tool execution with `LLM::Function#spawn`,
3194
- `LLM::Function::Array` (`call`, `wait`, `spawn`), and
3195
- `LLM::Function::ThreadGroup`.
3196
- - Add `LLM::Function::ThreadGroup#alive?` method for non-blocking
3197
- monitoring of concurrent tool execution.
3198
- - Add `LLM::Function::ThreadGroup#value` alias for `ThreadGroup#wait` for
3199
- consistency with Ruby's `Thread#value`.
3200
-
3201
- ### Change
3202
-
3203
- - Rename `LLM::Session` to `LLM::Context` throughout the codebase to better
3204
- reflect the concept of a stateful interaction environment.
3205
- - Rename `LLM::Gemini` to `LLM::Google` to better reflect provider naming.
3206
- - Standardize model objects across providers around a smaller common
3207
- interface.
3208
- - Switch registry cost internals from `LLM::Estimate` to `LLM::Cost`.
3209
- - Update image generation defaults so OpenAI and xAI consistently return
3210
- base64-encoded image data by default.
3211
- - Update `LLM::Bot` deprecation warning from v5.0 to v6.0, giving users
3212
- more time to migrate to `LLM::Context`.
3213
- - Rework the README and screencast documentation to better cover MCP,
3214
- registry, contexts, prompts, concurrency, providers, and example flow.
3215
- - Expand the README with architecture, production, and provider guidance
3216
- while improving readability and example ordering.
3217
-
3218
- ### Fix
3219
-
3220
- - Fix local schema `$ref` resolution in `LLM::Schema::Parser`.
3221
- - Fix multiple MCP issues around stdio env handling, request IDs, registry
3222
- interaction, tool registration, and filtering of MCP tools from the
3223
- standard tool registry.
3224
- - Fix stream parsing issues, including chunk-splitting bugs and safer
3225
- handling of streamed error responses.
3226
- - Fix prompt handling across contexts, agents, and provider adapters so
3227
- prompt turns remain consistent in history and completions.
3228
- - Fix several tool/context issues, including function return wrapping,
3229
- tool lookup after deserialization, unnamed subclass filtering, and
3230
- thread-safety around tool registry mutations.
3231
- - Fix Google tool-call handling to preserve `thoughtSignature`.
3232
- - Fix `LLM::Tracer::Logger` argument handling.
3233
- - Fix packaging/docs issues such as registry files in the gemspec and
3234
- stale provider docs.
3235
- - Fix Google provider handling of `nil` function IDs during context
3236
- deserialization.
3237
- - Fix MCP stdio transport by increasing poll timeout for better
3238
- reliability.
3239
- - Fix Google provider to properly cast non-Hash tool results into Hash
3240
- format for API compatibility.
3241
- - Fix schema parser to support recursive normalization of `Array`,
3242
- `LLM::Object`, and nested structures.
3243
- - Fix DeepSeek provider to tolerate malformed tool arguments.
3244
- - Fix `LLM::Function::TaskGroup#alive?` to properly delegate to
3245
- `Async::Task#alive?`.
3246
- - Fix various RuboCop errors across the codebase.
3247
- - Fix DeepSeek provider to handle JSON that might be valid but unexpected.
3248
-
3249
- ### Notes
3250
-
3251
- Notable merged work in this range includes:
3252
-
3253
- - `feat(function): add fiber-based concurrency for async environments (#64)`
3254
- - `feat(mcp): add stdio MCP support (#134)`
3255
- - `Add LLM::Registry + cost support (#133)`
3256
- - `Consistent model objects across providers (#131)`
3257
- - `Add rack + websocket example (#130)`
3258
- - `feat(gemspec): add changelog URI (#136)`
3259
- - `feat(function): alias ThreadGroup#wait as ThreadGroup#value (#62)`
3260
- - `README and screencast refresh across `#66`, `#68`, `#71`, and
3261
- `#72`
3262
- - `chore(bot): update deprecation warning from v5.0 to v6.0`
3263
- - `fix(deepseek): tolerate malformed tool arguments`
3264
- - `refactor(context): Rename Session as Context (#70)`
3265
-
3266
- Comparison base:
3267
- - Latest tag: `v4.8.0` (`6468f2426ee125823b7ae43b4af507b125f96ffc`)
3268
- - HEAD used for this changelog: `915c48da6fda9bef1554ff613947a6ce26d382e3`