llm.rb 15.0.3 → 15.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +433 -3
- data/README.md +188 -71
- data/bin/llm.rb +50 -7
- data/data/alibaba.json +912 -823
- data/data/anthropic.json +234 -187
- data/data/bedrock.json +3702 -2058
- data/data/deepinfra.json +1288 -951
- data/data/deepseek.json +87 -53
- data/data/google.json +670 -670
- data/data/mistral.json +501 -460
- data/data/moonshot.json +43 -248
- data/data/openai.json +1008 -914
- data/data/openrouter.json +14417 -0
- data/data/xai.json +213 -201
- data/data/zai.json +242 -149
- data/docs/deepdive/advanced/compaction.md +5 -5
- data/docs/deepdive/advanced/context.md +8 -6
- data/docs/deepdive/advanced/guard.md +2 -2
- data/docs/deepdive/features/builtin_tools.md +93 -22
- data/docs/deepdive/features/{repl.md → console.md} +28 -28
- data/docs/deepdive/features/database.md +3 -3
- data/docs/deepdive/fundamentals/agents.md +13 -12
- data/docs/deepdive/fundamentals/providers.md +91 -6
- data/docs/deepdive/fundamentals/skills.md +14 -6
- data/docs/deepdive/fundamentals/stream.md +4 -4
- data/docs/deepdive/fundamentals/tools.md +63 -31
- data/docs/deepdive/reference/cost.md +2 -2
- data/docs/deepdive/reference/model_registry.md +2 -2
- data/docs/deepdive/reference/tracer.md +15 -13
- data/docs/deepdive.md +2 -2
- data/lib/llm/active_record/acts_as_agent.rb +9 -5
- data/lib/llm/agent.rb +40 -15
- data/lib/llm/{repl → console}/bar.rb +3 -3
- data/lib/llm/{repl → console}/buffer.rb +24 -9
- data/lib/llm/{repl → console}/color.rb +2 -2
- data/lib/llm/{repl → console}/command.rb +12 -12
- data/lib/llm/{repl → console}/commands/exit.rb +4 -4
- data/lib/llm/{repl → console}/commands/help.rb +1 -1
- data/lib/llm/{repl/commands/compact.rb → console/commands/keep.rb} +11 -9
- data/lib/llm/{repl → console}/commands/model.rb +2 -2
- data/lib/llm/{repl → console}/input/cache.rb +2 -2
- data/lib/llm/{repl → console}/input/char.rb +2 -2
- data/lib/llm/{repl → console}/input/row.rb +1 -1
- data/lib/llm/{repl → console}/input.rb +18 -10
- data/lib/llm/console/markdown/parser.rb +78 -0
- data/lib/llm/{repl → console}/markdown/table.rb +8 -5
- data/lib/llm/{repl → console}/markdown.rb +13 -30
- data/lib/llm/console/node.rb +69 -0
- data/lib/llm/{repl → console}/status.rb +11 -11
- data/lib/llm/{repl → console}/stream.rb +36 -9
- data/lib/llm/{repl → console}/walker.rb +1 -1
- data/lib/llm/{repl → console}/window.rb +17 -17
- data/lib/llm/{repl.rb → console.rb} +39 -19
- data/lib/llm/context/deserializer.rb +2 -1
- data/lib/llm/context.rb +29 -13
- data/lib/llm/cost.rb +13 -0
- data/lib/llm/function/async/reactor.rb +20 -1
- data/lib/llm/function/fork/task.rb +14 -10
- data/lib/llm/function.rb +1 -1
- data/lib/llm/json_adapter.rb +40 -28
- data/lib/llm/message.rb +7 -0
- data/lib/llm/provider.rb +31 -10
- data/lib/llm/providers/alibaba.rb +1 -1
- data/lib/llm/providers/anthropic.rb +1 -1
- data/lib/llm/providers/bedrock/models.rb +2 -2
- data/lib/llm/providers/bedrock.rb +1 -1
- data/lib/llm/providers/deepseek.rb +1 -1
- data/lib/llm/providers/google.rb +1 -1
- data/lib/llm/providers/ollama.rb +1 -1
- data/lib/llm/providers/openai/responses.rb +2 -1
- data/lib/llm/providers/openai.rb +2 -1
- data/lib/llm/providers/openrouter.rb +87 -0
- data/lib/llm/schema/leaf.rb +34 -2
- data/lib/llm/schema.rb +4 -2
- data/lib/llm/sequel/agent.rb +9 -5
- data/lib/llm/skill.rb +7 -1
- data/lib/llm/stream.rb +8 -3
- data/lib/llm/tool/param.rb +5 -1
- data/lib/llm/tool.rb +5 -0
- data/lib/llm/tools/bundle.rb +53 -0
- data/lib/llm/tools/edit-file.rb +7 -2
- data/lib/llm/tools/exec.rb +78 -0
- data/lib/llm/tools/git.rb +27 -26
- data/lib/llm/tools/mkdir.rb +12 -19
- data/lib/llm/tools/read_file.rb +69 -9
- data/lib/llm/tools/rg.rb +20 -24
- data/lib/llm/tools/ruby.rb +17 -25
- data/lib/llm/tools/utils.rb +75 -2
- data/lib/llm/tools/write_file.rb +4 -1
- data/lib/llm/tracer/logger.rb +2 -2
- data/lib/llm/tracer/pretty_logger.rb +4 -4
- data/lib/llm/tracer/telemetry.rb +2 -2
- data/lib/llm/tracer.rb +33 -0
- data/lib/llm/transport/curb.rb +5 -3
- data/lib/llm/transport/http.rb +5 -2
- data/lib/llm/transport/persistent_http.rb +6 -4
- data/lib/llm/transport/utils.rb +8 -6
- data/lib/llm/version.rb +1 -1
- data/lib/llm.rb +18 -12
- data/llm.gemspec +8 -8
- metadata +80 -37
- data/lib/llm/repl/node.rb +0 -44
- data/lib/llm/tools/shell.rb +0 -55
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 5c55a3fa1151afad5720fc33bbe60a71feed5577242e53b4e8d4e2f9962a18ed
|
|
4
|
+
data.tar.gz: 3099ef73435ed6d91b003a40e2e5c195a82a8f631418b21e7c8cbce9088a5c76
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 8e0d5dd6c98355412a31dcb638081ea9f0553725cbf30aa6598eed5d16cd6d4693b5edfae8c063625ee2f5198b342e9b11850897370404267d291ddae4d747fe
|
|
7
|
+
data.tar.gz: 25102554a840fee5046804d17ff91673793603b96ab9ae246a30b7d3192a792e9bd7e384d76ce20abf2d504910330202cb2d4aef0057c7875c1ce3866ccfe177
|
data/CHANGELOG.md
CHANGED
|
@@ -11,10 +11,440 @@
|
|
|
11
11
|
</p>
|
|
12
12
|
|
|
13
13
|
> Changelog <br>
|
|
14
|
-
>
|
|
14
|
+
> [r.uby.dev](https://r.uby.dev) project
|
|
15
15
|
|
|
16
16
|
## What's next
|
|
17
17
|
|
|
18
|
+
*No unreleased changes yet. Check back after the next release.*
|
|
19
|
+
|
|
20
|
+
## v15.2.0
|
|
21
|
+
|
|
22
|
+
Changes since `v15.1.0`.
|
|
23
|
+
|
|
24
|
+
This release renames the REPL to `LLM::Console` (with `/keep` replacing
|
|
25
|
+
`/compact`) and routes every shell-out tool through a shared, bounded
|
|
26
|
+
`exec` runner. It also adds `LLM::Message#created_at`, the `LLM::Tracer`
|
|
27
|
+
factory methods, a `bundle` tool, per-tool `max_bytes` output limits,
|
|
28
|
+
and a `-v` switch to the CLI, and refreshes the model registry.
|
|
29
|
+
|
|
30
|
+
### Core
|
|
31
|
+
|
|
32
|
+
* **message: add `LLM::Message#created_at`** <br>
|
|
33
|
+
[`LLM::Message#created_at`](https://r.uby.dev/api-docs/llm.rb/LLM/Message.html#created_at-instance_method)
|
|
34
|
+
returns the time the message was created, defaulting to the moment the
|
|
35
|
+
message is initialized. The timestamp is serialized into
|
|
36
|
+
[`LLM::Context#to_json`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#to_json-instance_method)
|
|
37
|
+
as an ISO-8601 string and restored on deserialization, so it can be
|
|
38
|
+
stored alongside the rest of the conversation.
|
|
39
|
+
|
|
40
|
+
### Agent
|
|
41
|
+
|
|
42
|
+
* **agent: inherit the ORM model's name** <br>
|
|
43
|
+
An `acts_as_agent` (ActiveRecord) or `plugin :agent` (Sequel) model now
|
|
44
|
+
names its generated
|
|
45
|
+
[`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html) after
|
|
46
|
+
the model class. Previously the agent was an anonymous subclass, so
|
|
47
|
+
without an explicit name it defaulted to a gibberish `#<Class:0x...>`
|
|
48
|
+
string. The wrapper now initializes the agent's name before `.agent`
|
|
49
|
+
returns, and
|
|
50
|
+
[`LLM::Agent.name`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#name-class_method)
|
|
51
|
+
kebab-cases a `Class` argument, so an `AdminUser` model yields an agent
|
|
52
|
+
named `admin-user`.
|
|
53
|
+
|
|
54
|
+
### Cli
|
|
55
|
+
|
|
56
|
+
* **cli: add a `-v` switch** <br>
|
|
57
|
+
`bin/llm.rb` gains a `-v` switch that prints the version (`llm.rb v#{LLM::VERSION}`) and exits.
|
|
58
|
+
|
|
59
|
+
### Console
|
|
60
|
+
|
|
61
|
+
* **console: raise `LLM::Interrupt` on the agent's thread** <br>
|
|
62
|
+
Pressing Esc to cancel now also raises `LLM::Interrupt` on the agent's
|
|
63
|
+
thread. `LLM::Agent#cancel!` alone can be a no-op at some stages of the
|
|
64
|
+
request lifecycle, so the console backs it up by interrupting the thread
|
|
65
|
+
that runs the agent.
|
|
66
|
+
|
|
67
|
+
* **console: add a `/keep` command and retire `/compact`** <br>
|
|
68
|
+
The console now offers `/keep` for freeing space in the context window;
|
|
69
|
+
the `/compact` command is removed. `/keep` takes the same argument, so
|
|
70
|
+
`/keep 20%` keeps 20% of the context window. Closes
|
|
71
|
+
[issue #161](https://github.com/r-uby-dev/llm.rb/issues/161).
|
|
72
|
+
|
|
73
|
+
* **console: keep the UI responsive during long streams** <br>
|
|
74
|
+
A model can emit many chunks in a single turn. The console now draws
|
|
75
|
+
at most four streamed chunks at a time, then checks for input, so the
|
|
76
|
+
UI stays responsive even when a turn produces a large amount of
|
|
77
|
+
output.
|
|
78
|
+
|
|
79
|
+
* **console: persist the conversation when a turn is done** <br>
|
|
80
|
+
The console now saves the agent's state after the turn finishes,
|
|
81
|
+
rather than while the response is still streaming. State is still
|
|
82
|
+
saved every turn, but not until the turn has completed.
|
|
83
|
+
|
|
84
|
+
* **console: render markdown text as typed** <br>
|
|
85
|
+
Fix a bug where [`LLM::Console::Markdown`](https://r.uby.dev/api-docs/llm.rb/LLM/Console/Markdown.html)
|
|
86
|
+
mangled the model's output: HTML could render invisible, and
|
|
87
|
+
sequences like `...` were converted to unicode glyphs. The renderer
|
|
88
|
+
now uses a custom kramdown parser that disables the HTML, smart-quote,
|
|
89
|
+
and typographic-symbol parsers, so tags and punctuation come through
|
|
90
|
+
exactly as written.
|
|
91
|
+
|
|
92
|
+
* **console: fix a crash in the markdown parser** <br>
|
|
93
|
+
Fix a bug where the markdown renderer raised an error on an unclosed
|
|
94
|
+
HTML tag or a partial tag taken out of context, such as `4 < 5`. The
|
|
95
|
+
parser now emits the `<...` run literally when there is no closing
|
|
96
|
+
`>`, so the text renders instead of crashing.
|
|
97
|
+
|
|
98
|
+
* **console: stop rendering bare pipes as tables** <br>
|
|
99
|
+
Fix a bug where the markdown renderer treated a lone `|foo|` in prose
|
|
100
|
+
as a table and mangled its output. A pipe line now parses as a table
|
|
101
|
+
only when a header row is followed by a delimiter row, so bare pipes
|
|
102
|
+
come through literally while real tables still render.
|
|
103
|
+
|
|
104
|
+
* **console: find the worker thread when cancelling** <br>
|
|
105
|
+
Fix a bug where pressing Esc to cancel raised `LLM::Interrupt` on an
|
|
106
|
+
instance variable that does not exist, so the interrupt was a no-op and
|
|
107
|
+
a cancel could leave the turn running. The console now resolves the
|
|
108
|
+
worker thread through its `#thread` reader and interrupts it.
|
|
109
|
+
|
|
110
|
+
* **console: protect the state write from cancellation** <br>
|
|
111
|
+
The console now defers `LLM::Interrupt` while it saves the agent's
|
|
112
|
+
state after a turn, so a cancel that arrives during the write cannot
|
|
113
|
+
interrupt `agent.save` mid-flight and risk a lost or corrupted session
|
|
114
|
+
file.
|
|
115
|
+
|
|
116
|
+
### Tools
|
|
117
|
+
|
|
118
|
+
* **tools: the command runner is now `exec`** <br>
|
|
119
|
+
The command tool that spawns a process without a shell is now
|
|
120
|
+
[`LLM::Tool::Exec`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool/Exec.html),
|
|
121
|
+
with the tool name `exec` instead of the previous `shell`. This is an
|
|
122
|
+
internal refactor of the shell-out tools: `git`, `rg`, `mkdir`,
|
|
123
|
+
`ruby`, and `bundle` all route through it and inherit
|
|
124
|
+
its bounded output.
|
|
125
|
+
|
|
126
|
+
* **tools: report when a command cannot be found** <br>
|
|
127
|
+
`LLM::Tool::Exec` now returns `{ok: false, error: "command 'NAME' was
|
|
128
|
+
not found on this system"}` when the requested command is missing,
|
|
129
|
+
instead of a bare `{ok: false}` result that did not tell the model why
|
|
130
|
+
the tool failed.
|
|
131
|
+
|
|
132
|
+
* **tools: drop the `name:` parameter from `LLM::Tool::Exec#call`** <br>
|
|
133
|
+
`LLM::Tool::Exec#call` now takes a single `arguments:` array instead of
|
|
134
|
+
separate `name:` and `arguments:` parameters, with the command name as
|
|
135
|
+
the first element (for example `arguments: ["rg", "-m", "10", "lib"]`).
|
|
136
|
+
The `Git`, `Mkdir`, `Rg`, `Ruby`, and `Bundle` tools build their calls
|
|
137
|
+
the same way. The change was made after models were observed confusing
|
|
138
|
+
the two parameters, so a single list is simpler and more reliable.
|
|
139
|
+
|
|
140
|
+
* **tools: rename `repl` as `console`** <br>
|
|
141
|
+
The interactive loop is renamed to
|
|
142
|
+
[`LLM::Console`](https://r.uby.dev/api-docs/llm.rb/LLM/Console.html),
|
|
143
|
+
which better reflects what it does. `agent.console` is the primary
|
|
144
|
+
entry point, and the require path moves from `llm/repl` to
|
|
145
|
+
`llm/console`. Backwards-compatible aliases remain: `LLM::Repl`,
|
|
146
|
+
`LLM::Agent#repl`, the ORM wrappers' `#repl`, and `LLM::Command =`
|
|
147
|
+
`LLM::Console::Command`.
|
|
148
|
+
|
|
149
|
+
* **tools: `LLM::Tool::Git#call` takes an `arguments:` array** <br>
|
|
150
|
+
[`LLM::Tool::Git#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool/Git.html)
|
|
151
|
+
now takes a single `arguments:` array in place of the previous
|
|
152
|
+
`subcommand:` parameter. The first element must be one of `log`,
|
|
153
|
+
`diff`, `commit`, `checkout`, `branch`, or `show`, validated before the
|
|
154
|
+
command is spawned; the remaining elements are forwarded to git.
|
|
155
|
+
|
|
156
|
+
* **tools: `LLM::Tool::Utils` now owns command spawning** <br>
|
|
157
|
+
The shared [`LLM::Tool::Utils`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool/Utils.html)
|
|
158
|
+
module now requires the `test-cmd.rb` gem (at `~> 2.7.1`) itself and
|
|
159
|
+
exposes the `spawn` and `wait` helpers, so any tool that includes
|
|
160
|
+
`Utils` gets command spawning without requiring `exec` directly. The
|
|
161
|
+
`Git`, `Mkdir`, `Rg`, `Ruby`, `Exec`, and `Bundle` tools all
|
|
162
|
+
inherit their bounded-output protections from this shared runner.
|
|
163
|
+
|
|
164
|
+
* **tools: route `git`, `rg`, `mkdir`, and `ruby` through `exec`** <br>
|
|
165
|
+
`LLM::Tool::Git`, `LLM::Tool::Rg`, `LLM::Tool::Mkdir`, and
|
|
166
|
+
`LLM::Tool::Ruby` now implement their calls through the `exec` tool,
|
|
167
|
+
completing the refactor so every tool that shells out flows through
|
|
168
|
+
the shared command runner with its bounded output.
|
|
169
|
+
|
|
170
|
+
* **tools: read-file returns structured lines** <br>
|
|
171
|
+
`LLM::Tool::ReadFile#call` now returns its content as structured
|
|
172
|
+
`{lineno:, content:}` lines under a `lines:` key instead of a single
|
|
173
|
+
`content:` string, and adds a `truncated:` flag. A reversed range
|
|
174
|
+
(`start: 20, stop: 2`) is swapped to read lines 2 through 20. The
|
|
175
|
+
truncation marker is kept out of the returned lines, so the model
|
|
176
|
+
does not mistake it for a real file line.
|
|
177
|
+
|
|
178
|
+
* **tools: write-file appends a trailing newline by default** <br>
|
|
179
|
+
`LLM::Tool::WriteFile` now ensures written content ends with a newline,
|
|
180
|
+
adding one when the content does not already end with `\n`. It previously
|
|
181
|
+
wrote the content exactly as given. A new `newline:` parameter (default
|
|
182
|
+
`true`) controls this, so `newline: false` writes the content exactly as
|
|
183
|
+
given.
|
|
184
|
+
|
|
185
|
+
* **tools: fix `edit-file` treating `before` as a regex** <br>
|
|
186
|
+
`LLM::Tool::EditFile` now escapes the `before` snippet with
|
|
187
|
+
`Regexp.escape`, so regex metacharacters are matched literally, and
|
|
188
|
+
switches to the block form of `sub` so the `after` replacement keeps
|
|
189
|
+
backslash sequences like `\1` and `\&` literal.
|
|
190
|
+
|
|
191
|
+
* **tools: bound tool output with a per-tool `max_bytes`** <br>
|
|
192
|
+
Each of the `Exec`, `ReadFile`, `Rg`, `Mkdir`, `Ruby`, and
|
|
193
|
+
`Bundle` tools gains a `max_bytes` limit (default 75,000) for the
|
|
194
|
+
maximum number of bytes a tool returns to the model. `Exec` and
|
|
195
|
+
`ReadFile` add the class-level `max_bytes` accessor, which the other
|
|
196
|
+
tools inherit through `Exec`, so each tool's cap can be configured
|
|
197
|
+
independently, for example `LLM::Tool::ReadFile.max_bytes(175_000)`.
|
|
198
|
+
It does not enforce the limit by itself;
|
|
199
|
+
[`LLM::Tool::Utils#truncate`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool/Utils.html#truncate-instance_method)
|
|
200
|
+
trims a string within the limit and marks the trailing content as
|
|
201
|
+
truncated, and `truncate!` returns a `[content, truncated]` tuple for
|
|
202
|
+
callers that structure truncated output themselves. `rg` also gains a
|
|
203
|
+
`max_count:` parameter that caps the number of results per file.
|
|
204
|
+
|
|
205
|
+
* **tools: add a `bundle` tool** <br>
|
|
206
|
+
A new [`LLM::Tool::Bundle`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool/Bundle.html)
|
|
207
|
+
tool runs a command through `bundle`. It uses the `BUNDLE_GEMFILE`
|
|
208
|
+
environment variable when set, or a `Gemfile` in the current working
|
|
209
|
+
directory otherwise. The tool takes an `arguments:` array, so the
|
|
210
|
+
model passes the bundle command and its arguments as a single list,
|
|
211
|
+
for example `arguments: ["exec", "rspec"]`.
|
|
212
|
+
|
|
213
|
+
* **tools: resolve defaults through `LLM::Utils.resolve_option`** <br>
|
|
214
|
+
A tool parameter default can now be an immediate value, a Symbol resolved
|
|
215
|
+
as a method on the tool, or a Proc evaluated lazily at runtime, matching
|
|
216
|
+
how `LLM::Agent` resolves its attributes. This lets a default track a
|
|
217
|
+
value that can change between boot and runtime, such as a tool's
|
|
218
|
+
`max_bytes`.
|
|
219
|
+
|
|
220
|
+
### Registry
|
|
221
|
+
|
|
222
|
+
* **refresh model metadata** <br>
|
|
223
|
+
Update `data/` with current pricing, limits, and capabilities for the
|
|
224
|
+
OpenRouter, OpenAI, Bedrock, DeepInfra, DeepSeek, Google, Mistral,
|
|
225
|
+
Moonshot, Z.ai, and Alibaba registries.
|
|
226
|
+
|
|
227
|
+
### Provider
|
|
228
|
+
|
|
229
|
+
* **provider: retry `Net::WriteTimeout`, too** <br>
|
|
230
|
+
Requests that raise `Net::WriteTimeout` are now retried alongside the
|
|
231
|
+
other timed-out and rate-limited requests, up to the `retry_budget`,
|
|
232
|
+
matching how `Net::OpenTimeout` and `Net::ReadTimeout` are handled. The
|
|
233
|
+
console status bar also reports a write timeout as `Timed out`.
|
|
234
|
+
|
|
235
|
+
* **alibaba: default to a retry budget of 8** <br>
|
|
236
|
+
An agent that runs on the Alibaba provider now defaults to a retry
|
|
237
|
+
budget of 8 instead of 5, because Alibaba (token plan) frequently rate
|
|
238
|
+
limits and times out requests that it later recovers from. An explicit
|
|
239
|
+
`retry_budget:` still overrides the default.
|
|
240
|
+
|
|
241
|
+
* **deepseek: default to the `deepseek-flash` model** <br>
|
|
242
|
+
The default DeepSeek chat model is now `deepseek-flash` instead of
|
|
243
|
+
`deepseek-v4-flash`. DeepSeek resolves `deepseek-flash` to
|
|
244
|
+
`deepseek-v4.1-flash` and recommends the name in its documentation and
|
|
245
|
+
API error messages, so the default follows the current model alias
|
|
246
|
+
instead of a pinned version.
|
|
247
|
+
|
|
248
|
+
### Tracer
|
|
249
|
+
|
|
250
|
+
* **tracer: add `LLM::Tracer` factory methods** <br>
|
|
251
|
+
Add
|
|
252
|
+
[`LLM::Tracer.logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html#logger-class_method),
|
|
253
|
+
[`LLM::Tracer.pretty_logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html#pretty_logger-class_method),
|
|
254
|
+
and
|
|
255
|
+
[`LLM::Tracer.telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html#telemetry-class_method)
|
|
256
|
+
as the preferred way to build a tracer for a provider, so switching
|
|
257
|
+
between tracers means changing a factory method instead of a class name.
|
|
258
|
+
The old `LLM.logger(llm, ...)` convenience method is removed in favor
|
|
259
|
+
of `LLM::Tracer.logger(llm, ...)`.
|
|
260
|
+
|
|
261
|
+
* **tracer: add `path:` support to `LLM::Tracer::PrettyLogger`** <br>
|
|
262
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
263
|
+
now accepts a `path:` option to write its human-readable entries to a
|
|
264
|
+
file, matching `LLM::Tracer::Logger`. It previously only accepted `io:`.
|
|
265
|
+
|
|
266
|
+
### Fix
|
|
267
|
+
|
|
268
|
+
* **json: scrub invalid UTF-8 on dump** <br>
|
|
269
|
+
Fix a bug where [`LLM::JSONAdapter`](https://r.uby.dev/api-docs/llm.rb/LLM/JSONAdapter.html)
|
|
270
|
+
raised a JSON generator error when dumping a string tagged as UTF-8 that
|
|
271
|
+
carried invalid bytes. The normalize step now transcodes every string to
|
|
272
|
+
valid UTF-8, replacing invalid sequences with the replacement character,
|
|
273
|
+
so dumping works on `json ~> 3.0`. The `oj` and `yajl` adapters now run
|
|
274
|
+
the same normalization, so every backend scrubs invalid bytes before
|
|
275
|
+
serializing.
|
|
276
|
+
|
|
277
|
+
* **fork: require xchan.rb `~> 0.23`** <br>
|
|
278
|
+
The `:fork` concurrency strategy now requires the `xchan.rb` gem at
|
|
279
|
+
`~> 0.23` instead of `~> 0.22`. xchan.rb 0.23.0 replaces the external
|
|
280
|
+
`lockf.rb` gem with a built-in, Fiddle-based `Chan::Lockf`, so fork
|
|
281
|
+
channels no longer carry that extra dependency. (The socket
|
|
282
|
+
length-header deadlock fix shipped earlier, in xchan.rb 0.22.0.)
|
|
283
|
+
|
|
284
|
+
* **async: fix a shutdown exception on the reactor thread** <br>
|
|
285
|
+
Fix a bug where the `:async` strategy's
|
|
286
|
+
[`LLM::Function::Async::Reactor`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Async/Reactor.html)
|
|
287
|
+
raised a `TypeError` on shutdown with recent `async` and `io-event`
|
|
288
|
+
versions, because their internals tried to raise an integer as an
|
|
289
|
+
exception. The scheduler is now detached from the reactor thread
|
|
290
|
+
before it exits, which avoids that code path entirely, and teardown
|
|
291
|
+
is managed by `reactor.stop`, so the thread exits promptly instead of
|
|
292
|
+
abruptly or hanging.
|
|
293
|
+
|
|
294
|
+
* **openai: prevent the loss of user messages in the completions path** <br>
|
|
295
|
+
Fix a bug where the request body was built from `params[:messages]`
|
|
296
|
+
alone when that key was present, discarding the messages built from the
|
|
297
|
+
prompt. The DeepSeek and Alibaba schema support injects the schema
|
|
298
|
+
system message into `params[:messages]`, so a request with a `schema:`
|
|
299
|
+
could be sent with the schema message only, dropping the user's
|
|
300
|
+
messages. The built messages now lead with `params[:messages]`, and the
|
|
301
|
+
key is removed before the body is assembled.
|
|
302
|
+
|
|
303
|
+
## v15.1.0
|
|
304
|
+
|
|
305
|
+
Changes since `v15.0.3`.
|
|
306
|
+
|
|
307
|
+
This release adds the OpenRouter provider, splits provider timeouts
|
|
308
|
+
into `connect_timeout` and `read_timeout`, retries timed-out requests,
|
|
309
|
+
and renames `on_rate_limit` to `on_retry`. Skills now gain a
|
|
310
|
+
frontmatter `model:` parameter and inherit the parent agent's model,
|
|
311
|
+
while the CLI gains `-m` and `-x` switches, and the console shows retry
|
|
312
|
+
progress and measures text by display width.
|
|
313
|
+
|
|
314
|
+
### Provider
|
|
315
|
+
|
|
316
|
+
* **add `LLM::OpenRouter` for the OpenRouter provider** <br>
|
|
317
|
+
[`LLM::OpenRouter`](https://r.uby.dev/api-docs/llm.rb/LLM/OpenRouter.html)
|
|
318
|
+
is a new provider that talks to [OpenRouter](https://openrouter.ai)
|
|
319
|
+
through its OpenAI-compatible API, contributed via
|
|
320
|
+
[PR #165](https://github.com/r-uby-dev/llm.rb/pull/165). Create an
|
|
321
|
+
instance with
|
|
322
|
+
[`LLM.openrouter`](https://r.uby.dev/api-docs/llm.rb/LLM.html#openrouter-class_method),
|
|
323
|
+
which accepts the same `key:`, `host:`, and `base_path:` options as the
|
|
324
|
+
OpenAI provider. It defaults to the `openrouter/auto` router model and
|
|
325
|
+
supports chat completions, streaming, tool calls, structured output,
|
|
326
|
+
and embeddings; image, audio, moderation, files, and vector store
|
|
327
|
+
endpoints raise `NotImplementedError`. Model metadata ships in
|
|
328
|
+
`data/openrouter.json` for the registry.
|
|
329
|
+
|
|
330
|
+
* **provider: `LLM::Provider#with` accepts headers without the `headers:` keyword** <br>
|
|
331
|
+
[`LLM::Provider#with`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html#with-instance_method)
|
|
332
|
+
now accepts headers directly as a Hash (`llm.with("User-Agent" => "llmrb/1.0")`)
|
|
333
|
+
in addition to the legacy `headers:` keyword form. Both are merged, and
|
|
334
|
+
the keyword form remains for backwards compatibility.
|
|
335
|
+
|
|
336
|
+
* **providers: `model: nil` falls back to `default_model`** <br>
|
|
337
|
+
[`LLM::Provider`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html)
|
|
338
|
+
now treats a `model: nil` payload as "use the default model" instead of
|
|
339
|
+
sending a model value that providers reject. Previously a `{model: nil}`
|
|
340
|
+
param overwrote the default and then was removed by `.compact`, leaving
|
|
341
|
+
undefined behavior where providers could reject the request.
|
|
342
|
+
|
|
343
|
+
* **openai: handle `model: nil` in the Responses API** <br>
|
|
344
|
+
The OpenAI Responses API path now also falls back to the default model
|
|
345
|
+
when `model: nil`, matching the completions path.
|
|
346
|
+
|
|
347
|
+
* **provider: separate `connect_timeout` and `read_timeout`** <br>
|
|
348
|
+
[`LLM::Provider`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html)
|
|
349
|
+
now accepts a `connect_timeout:` for opening the connection (default 5
|
|
350
|
+
seconds) and a `read_timeout:` for waiting on an idle connection (default
|
|
351
|
+
600 seconds), instead of a single `timeout`. The legacy `timeout:` option
|
|
352
|
+
remains as a shorthand for `read_timeout`. The provider exposes the new
|
|
353
|
+
`#read_timeout` and `#connect_timeout` accessors.
|
|
354
|
+
|
|
355
|
+
* **provider: retry timed-out requests** <br>
|
|
356
|
+
Requests that time out (`Net::OpenTimeout` or `Net::ReadTimeout`) are now
|
|
357
|
+
retried along with rate-limited requests, up to the `retry_budget`. When
|
|
358
|
+
a heavily loaded provider (for example DeepSeek) drops the connection
|
|
359
|
+
during the connect or read phase, the request is retried instead of
|
|
360
|
+
failing, and the stream is notified through `on_retry`.
|
|
361
|
+
|
|
362
|
+
* **context: retry on `LLM::InsufficientQuotaError`, too** <br>
|
|
363
|
+
A request that raises
|
|
364
|
+
[`LLM::InsufficientQuotaError`](https://r.uby.dev/api-docs/llm.rb/LLM/InsufficientQuotaError.html)
|
|
365
|
+
is now retried like other rate-limited requests. The error is a subclass
|
|
366
|
+
of `LLM::RateLimitError` and can be raised at regular intervals, in
|
|
367
|
+
particular by the Alibaba provider.
|
|
368
|
+
|
|
369
|
+
* **cost: return `LLM::Cost.zero` when the registry has no pricing** <br>
|
|
370
|
+
Add
|
|
371
|
+
[`LLM::Cost.zero`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#zero-class_method),
|
|
372
|
+
a factory for a zero-valued cost breakdown.
|
|
373
|
+
[`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
|
|
374
|
+
now returns it when the active model has no pricing in the registry (for
|
|
375
|
+
example OpenRouter's `openrouter/auto` router model), instead of crashing
|
|
376
|
+
on a nil pricing entry.
|
|
377
|
+
|
|
378
|
+
### Skills
|
|
379
|
+
|
|
380
|
+
* **skills: add a `model` frontmatter parameter** <br>
|
|
381
|
+
A skill's `SKILL.md` frontmatter can now declare a `model:` value, so a
|
|
382
|
+
skill's sub-agent runs on a specific model instead of the default. This
|
|
383
|
+
lets a parent agent run on one model (for example `deepseek-v4-flash`)
|
|
384
|
+
while the skill runs on another (`deepseek-v4-pro`). The `model` value is
|
|
385
|
+
not strictly portable between providers.
|
|
386
|
+
|
|
387
|
+
* **skills: inherit the model of the parent agent** <br>
|
|
388
|
+
A skill's sub-agent now inherits the active model of the agent that
|
|
389
|
+
spawns it instead of falling back to the provider default. The frontmatter
|
|
390
|
+
`model:` parameter overrides it explicitly when set.
|
|
391
|
+
[`LLM::Context#model`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#model-instance_method)
|
|
392
|
+
now falls back to the provider's default model when no model is set.
|
|
393
|
+
|
|
394
|
+
### Stream
|
|
395
|
+
|
|
396
|
+
* **stream: rename `on_rate_limit` to `on_retry`** <br>
|
|
397
|
+
[`LLM::Stream#on_rate_limit`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_rate_limit-instance_method)
|
|
398
|
+
is renamed to
|
|
399
|
+
[`LLM::Stream#on_retry`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_retry-instance_method),
|
|
400
|
+
which fires when a request is retried after a rate limit or a timeout,
|
|
401
|
+
and now receives the one-based retry attempt number as a second argument.
|
|
402
|
+
The old `on_rate_limit` name remains as an alias.
|
|
403
|
+
|
|
404
|
+
### CLI
|
|
405
|
+
|
|
406
|
+
* **cli: add a `-m` switch for choosing the model** <br>
|
|
407
|
+
`bin/llm.rb` now accepts `-m MODEL` to run the session on a model other
|
|
408
|
+
than the provider default, for example
|
|
409
|
+
`llm.rb -p deepseek -m deepseek-v4-pro`. A model unknown to the provider's
|
|
410
|
+
registry prints an error and exits.
|
|
411
|
+
|
|
412
|
+
* **cli: add a `-x` switch for the read timeout** <br>
|
|
413
|
+
`bin/llm.rb` now accepts `-x SECONDS` to set the provider read timeout.
|
|
414
|
+
|
|
415
|
+
* **cli: print a backtrace on fatal crashes** <br>
|
|
416
|
+
When `bin/llm.rb` hits an unexpected error, the crash message now includes
|
|
417
|
+
up to three stack lines from the backtrace, so the failure is easier to
|
|
418
|
+
locate and report than a bare diagnostic.
|
|
419
|
+
|
|
420
|
+
### Console
|
|
421
|
+
|
|
422
|
+
* **console: show retry progress in the status bar** <br>
|
|
423
|
+
When a request is rate limited or times out, the curses-based console status
|
|
424
|
+
bar shows a retry indicator with the error and the remaining attempts, for
|
|
425
|
+
example `🔁 Rate limited • attempt 2 of 5`.
|
|
426
|
+
|
|
427
|
+
* **console: measure text width with `unicode-display_width`** <br>
|
|
428
|
+
The curses-based console now counts and slices text by display column width
|
|
429
|
+
instead of character count, so wrapping, table columns, and clipping stay
|
|
430
|
+
aligned for wide characters such as emoji. It requires the optional
|
|
431
|
+
`unicode-display_width` gem.
|
|
432
|
+
|
|
433
|
+
* **console: treat `LLM::InsufficientQuotaError` as a rate limit in the status bar** <br>
|
|
434
|
+
The curses-based console status bar now shows `Rate limited` when a request
|
|
435
|
+
raises
|
|
436
|
+
[`LLM::InsufficientQuotaError`](https://r.uby.dev/api-docs/llm.rb/LLM/InsufficientQuotaError.html),
|
|
437
|
+
matching how ordinary `LLM::RateLimitError`s are shown, instead of falling
|
|
438
|
+
through to the raw class name.
|
|
439
|
+
|
|
440
|
+
### Registry
|
|
441
|
+
|
|
442
|
+
* **refresh model metadata** <br>
|
|
443
|
+
Update `data/*.json` with current model listings and pricing, adding
|
|
444
|
+
GPT-5.6 Sol, Terra, and Luna models to Bedrock, Grok 4.6 and Grok Imagine
|
|
445
|
+
Image 2.0 to xAI, DeepSeek V4 Flash Vision to DeepSeek, and DeepSeek V4
|
|
446
|
+
Pro 0813, Qwen3 VL, and Qwen3.8 models to DeepInfra.
|
|
447
|
+
|
|
18
448
|
## v15.0.3
|
|
19
449
|
|
|
20
450
|
Changes since `v15.0.2`.
|
|
@@ -1077,7 +1507,7 @@ reliable across all six concurrency backends. The `functions` and
|
|
|
1077
1507
|
`LLM.require` now accepts a second `version` parameter that is passed
|
|
1078
1508
|
to `Kernel#gem` before loading, enabling version constraints for
|
|
1079
1509
|
optional runtime dependencies. For example,
|
|
1080
|
-
`LLM.require "test-cmd.rb", "~>
|
|
1510
|
+
`LLM.require "test-cmd.rb", "~> 2.2"` ensures a minimum gem version
|
|
1081
1511
|
is available. This is used internally by the `Git`, `Rg`, `Mkdir`,
|
|
1082
1512
|
and `Shell` tools to enforce compatibility with the `test-cmd.rb` gem.
|
|
1083
1513
|
|
|
@@ -2141,7 +2571,7 @@ As always, see the changelog details for a thorough overview.
|
|
|
2141
2571
|
* **Add `LLM::Agent#repl`** <br>
|
|
2142
2572
|
Add a curses-based read-eval-print loop for `LLM::Agent` that lets
|
|
2143
2573
|
developers interact with an agent after it has been set up or has
|
|
2144
|
-
performed a task. It is similar to `binding.
|
|
2574
|
+
performed a task. It is similar to `binding.irb`: once you exit,
|
|
2145
2575
|
you can continue with the rest of your program. It requires the
|
|
2146
2576
|
`curses` gem.
|
|
2147
2577
|
|