axn-ruby_llm 0.3.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 1c013cfd72ea5be2662e93875a9f0f01b8d8761cfe008aabed920bb3ee4ab8d1
4
- data.tar.gz: 2977d6ab65433f53e59f019c796a9a10adc016efa290990ee3596d4b0680df4f
3
+ metadata.gz: 6976e9d15c05e921240e525fa9a421ea5af2f66bee88fe182d51e6dcf042e6a1
4
+ data.tar.gz: c979116f6bde6b20a5cf89664c9d9c04180f7165f48a2685e86db7db0385becc
5
5
  SHA512:
6
- metadata.gz: 532fa712897c472636dff0085d9df2b9696f0a355e88db3f18978da9128b6d4c9d0a7251170b8f84b9354e55c7fd83406b0d0ab7863b4291b1f255dfc9d7c5b1
7
- data.tar.gz: 8f77714a318a5bbf674d467bd2a0d5bb4b82c6f4d199d0faf84dc82a5d298fba9825aa083d2a71ca79c89e67d147cf150d02f2d044d6a573f88a936602782173
6
+ metadata.gz: 77698ef5d994d06e264b56845b507721058c2945357342182c9557582a9ea20572dc47164182fd541e2636115485fd46dcfb77ee63ff2646279d4b5805d6cd57
7
+ data.tar.gz: ebbae53a4026c405b5bb2636169bb916c1b77b085f1f06872ccc609c25f05b01050d59ea4d4e29871035e24d2ed4f572a8afb28dceb1b473683748460962f1ee
data/CHANGELOG.md CHANGED
@@ -1,5 +1,111 @@
1
1
  # Changelog
2
2
 
3
+ ## [Unreleased]
4
+
5
+ ## [0.4.0] - 2026-09-25
6
+
7
+ `Ask` now covers all of `RubyLLM::Chat`'s settings, plus attachments, seeded history, streaming, and
8
+ a tool-call cap, and the gem gains app-side remote MCP tools. Most of this is additive, but three
9
+ behaviour changes need checking on upgrade.
10
+
11
+ ### ⚠️ Breaking changes
12
+
13
+ - **Truncated or content-filtered responses now fail instead of succeeding.** A response cut off by
14
+ the output token limit (`finish_reason: :max_tokens`) now fails with "Response was cut off by the
15
+ output token limit before it finished". One blocked by a provider content filter fails with
16
+ "Response was blocked by the provider's content filter". Both used to come back as a success
17
+ carrying partial or empty text; with `schema:`, truncation surfaced as a misleading "Response was
18
+ not valid JSON". Usage, `cost`, `raw_message`, `transcript` and `finish_reason` are still exposed
19
+ on the failure.
20
+ **What to check:** callers that relied on partial text from a truncated response should handle
21
+ the failure, or raise `max_output_tokens:`.
22
+ - **The OpenTelemetry attribute `gen_ai.usage.input_tokens` now includes cached tokens.** It used to
23
+ carry RubyLLM's uncached count, which under-counts badly under prompt caching (e.g. 9 instead of
24
+ ~58k on a `gpt-5.6` tool loop). It now carries `total_input_tokens` (uncached + cache reads + cache
25
+ writes), matching the OTel GenAI definition. New span attributes:
26
+ `gen_ai.usage.cache_read.input_tokens`, `gen_ai.usage.cache_creation.input_tokens`, and
27
+ `axn.ruby_llm.version`.
28
+ **What to check:** dashboards or alerts summing `gen_ai.usage.input_tokens` will step up as each
29
+ service upgrades. Filter on `@axn.ruby_llm.version:*` to see only spans with the new meaning. Cost
30
+ attributes are unchanged.
31
+ - **Specs that stub `RubyLLM::Chat` or `RubyLLM::Message` directly** (rather than using
32
+ `stub_axn_ruby_llm`) may need updating. `Ask` now calls `chat.ask(prompt, with: nil)`, reads
33
+ `finish_reason`, `max_tokens?` and `content_filtered?` on the final message, and reads `thinking`
34
+ and `server_tool_use` from `chat.tokens`. A verifying double missing these fails with an
35
+ unexpected-message error.
36
+ **What to check:** switch those specs to `stub_axn_ruby_llm`, which handles all of this.
37
+
38
+ ### Deprecated
39
+
40
+ - **`Ask`'s `input_tokens` and `prompt_tokens` exposures**, replaced by `uncached_input_tokens` and
41
+ `total_input_tokens` respectively, with unchanged values. The bare `input_tokens` name meant
42
+ uncached input in RubyLLM's convention but all input in OpenTelemetry's, so it's retired rather
43
+ than redefined. Both still work; removal is scheduled for 1.0 (see `DEPRECATIONS.md`).
44
+
45
+ ### Added
46
+
47
+ **`Ask` options.** Every public `RubyLLM::Chat#with_*` method now has a matching input, forwarded
48
+ only when given:
49
+
50
+ - `provider:`, `protocol:`, `assume_model_exists:`, `context:` → `RubyLLM.chat(...)`, alongside
51
+ `model:`
52
+ - `fallbacks:` / `fallback_on:` → `with_fallbacks(*fallbacks, on:)`
53
+ - `max_output_tokens:` → `with_max_output_tokens`
54
+ - `thinking:` → `with_thinking`; `citations:` → `with_citations`; `caching:` → `with_caching`;
55
+ `compaction:` → `with_compaction`. These four also forward an explicit `false`.
56
+ - `end_user:` → `with_end_user`; `provider_options:` → `with_provider_options`; `headers:` →
57
+ `with_headers`
58
+ - `provider_tools:` → `with_provider_tools` (e.g. a provider-hosted remote MCP server:
59
+ `{ mcp: { name:, url:, headers:, allowed_tools:, require_approval: } }`); `tool_options:` →
60
+ `with_tool_options`
61
+ - `cache_system_prompt: true` → `with_instructions(..., cache_until_here: true)`
62
+
63
+ **Beyond the `with_*` methods:**
64
+
65
+ - `attachments:` is passed to `Chat#ask` as `with:` (paths, URLs, or IO).
66
+ - `history:` seeds earlier turns before the prompt; `transcript` leaves them out.
67
+ - `on_chunk:` receives each streamed `RubyLLM::Chunk`. `response` is still the complete message, so
68
+ `schema:` is unaffected.
69
+ - `max_tool_calls:` is a soft cap on app-executed tool calls per `ask`, shared across every tool.
70
+ Past the cap, each call returns an error result telling the model to wrap up.
71
+ - `on_remote_tool_approval:` approves or denies each call waiting on `Chat#awaiting_approval?`
72
+ (provider-hosted calls with `require_approval:`, or local tools declared with
73
+ `Tool.requires_approval`) and drives the chat to completion.
74
+
75
+ **New result fields:** `transcript` (every message the chat exchanged, including the provider-hosted
76
+ calls a remote MCP server ran), `total_input_tokens`, `uncached_input_tokens`, `thinking_tokens`,
77
+ `server_tool_use` (per-use counters for provider-hosted tools, such as web searches), and
78
+ `finish_reason`.
79
+
80
+ **`Axn::RubyLLM.remote_mcp_tools(url:, ...)`** connects to a remote MCP server from this app (using
81
+ the official `mcp` gem) and wraps its tools as `RubyLLM::Tool`s, so they sit next to Axn tools in
82
+ `tools:`.
83
+
84
+ - `allowed_tools:` limits which of the server's tools are exposed; pass it, since a server's full
85
+ list can include write/admin tools.
86
+ - `max_calls:` is a call budget shared across the toolset for its lifetime; connect one toolset per
87
+ request for a per-request cap. `timeout:` and `max_result_chars:` bound each call.
88
+ - Auth: static `headers:`; `bearer_token:` (a String, or a callable run on every request so the
89
+ caller can rotate it); or `oauth:` (an `mcp`-gem `OAuth::Provider` / `ClientCredentialsProvider`,
90
+ passed through). Token auth requires an https or loopback URL.
91
+ - If setup fails, the connection is closed before the error propagates.
92
+
93
+ **`stub_axn_ruby_llm`** accepts `finish_reason:`, `thinking_tokens:` and `server_tool_use:`.
94
+
95
+ **RubyLLM parity spec.** `spec/axn/ruby_llm/ruby_llm_parity_spec.rb` classifies every public method
96
+ on `RubyLLM::Chat`, `RubyLLM::Tokens` and `RubyLLM::Message` as covered or intentionally skipped,
97
+ and snapshots the signatures of the `Chat` methods `Ask` calls. A `ruby_llm` bump that adds or
98
+ re-signatures one fails CI by name.
99
+
100
+ ### Changed
101
+
102
+ - `stub_axn_ruby_llm` now stubs every public `RubyLLM::Chat#with_*` method (taken from the class
103
+ itself, so future ones are covered), plus `messages=` and `awaiting_approval?`, and matches
104
+ `model:` alongside the model-resolution keywords. It returns a real `RubyLLM::Tokens` rather than
105
+ a verifying double.
106
+ - `Ask`'s `provider_tools:`, `headers:` and `context:` inputs are marked `sensitive`, so axn keeps
107
+ them (and the API keys they usually carry) out of logs.
108
+
3
109
  ## [0.3.0] - 2026-09-22
4
110
 
5
111
  RubyLLM 2.0 is a breaking rewrite (renamed Tool DSL, restructured error hierarchy, a usage ledger
data/README.md CHANGED
@@ -2,23 +2,27 @@
2
2
 
3
3
  Call LLMs from [Axn](https://github.com/teamshares/axn) actions using [RubyLLM](https://github.com/crmne/ruby_llm), with declarative error handling, schema-based structured output, configurable defaults, and cost/token tracking — and wrap any Axn as a `RubyLLM::Tool` a chat can call.
4
4
 
5
- > **RubyLLM 2.0 required.** As of `0.3.0`, this gem requires `ruby_llm >= 2.0, < 3.0` and no longer supports RubyLLM 1.x — see [CHANGELOG.md](CHANGELOG.md) for the full breaking-change rundown if you're upgrading from an earlier `axn-ruby_llm` release.
5
+ > **RubyLLM 2.0 required.** As of `0.3.0`, this gem requires `ruby_llm >= 2.0, < 3.0` and no longer supports RubyLLM 1.x.
6
+ >
7
+ > **Upgrading to 0.4.0?** Two behaviour changes: a response cut off by the output token limit or blocked by a content filter now **fails** instead of succeeding, and the OpenTelemetry attribute `gen_ai.usage.input_tokens` now includes cached tokens. `input_tokens` / `prompt_tokens` are deprecated in favour of `uncached_input_tokens` / `total_input_tokens`. See [CHANGELOG.md](CHANGELOG.md) for the full list and what to check.
6
8
 
7
9
  Part of the `axn-*` extension ecosystem — see also [axn-mcp](https://github.com/teamshares/axn-mcp).
8
10
 
9
11
  ### Why use this over calling RubyLLM directly?
10
12
 
11
- Four things you'd otherwise hand-build:
13
+ Five things you'd otherwise hand-build:
12
14
 
13
- 1. **Structured error handling.** The Axn error DSL declaratively maps `RateLimitError`, `JSON::ParserError`, and generic `StandardError` to clean failure messages. Callers check `result.ok?` instead of wrapping every call in `begin/rescue`.
15
+ 1. **Structured error handling.** The Axn error DSL declaratively maps `RateLimitError`, `JSON::ParserError`, truncated or content-filtered responses, and generic `StandardError` to clean failure messages. Callers check `result.ok?` instead of wrapping every call in `begin/rescue`.
14
16
 
15
- 2. **Production gating.** A single `c.enabled = -> { Rails.env.production? }` in an initializer stubs every LLM call in non-prod environments — no per-callsite guards needed. The stub is typed (`stubbed: true`, `input_tokens: 0`, etc.) so downstream code doesn't need to branch on it either.
17
+ 2. **Production gating.** A single `c.enabled = -> { Rails.env.production? }` in an initializer stubs every LLM call in non-prod environments — no per-callsite guards needed. The stub is typed (`stubbed: true`, `total_input_tokens: 0`, etc.) so downstream code doesn't need to branch on it either.
16
18
 
17
- 3. **Cost/token tracking, exposed automatically.** Every call exposes `input_tokens`, `output_tokens`, `cache_read_tokens`, `cache_write_tokens`, `prompt_tokens` (the total), `cost`, and `cost_breakdown`, read straight off RubyLLM's own usage ledger (`Chat#tokens` / `Chat#cost`) — no manual model lookup, and a tool-call loop's retries and multiple round-trips are already aggregated for you. If your app uses OpenTelemetry, these values are also set as attributes on the existing `axn.call` span — no configuration required.
19
+ 3. **Cost/token tracking, exposed automatically.** Every call exposes `total_input_tokens`, `uncached_input_tokens`, `cache_read_tokens`, `cache_write_tokens`, `output_tokens`, `thinking_tokens`, `server_tool_use`, `cost`, and `cost_breakdown`, read straight off RubyLLM's own usage ledger (`Chat#tokens` / `Chat#cost`) — no manual model lookup, and a tool-call loop's retries and multiple round-trips are already aggregated for you. If your app uses OpenTelemetry, these values are also set as attributes on the existing `axn.call` span — no configuration required.
18
20
 
19
21
  4. **Author-once tools.** `Axn::RubyLLM.wrap` turns any Axn into a `RubyLLM::Tool` your chat can call — reuse the same Axn classes you already expose through [axn-mcp](https://github.com/teamshares/axn-mcp), or plain Axns, with no rewrite. The tool's name, JSON Schema, and argument validation all come from the Axn's own contract.
20
22
 
21
- > **Scope note:** This gem covers the subset of RubyLLM functionality that [Teamshares](https://github.com/teamshares) uses internally — single-turn chat, structured output, basic observability, and wrapping Axns as tools. It is intentionally minimal rather than a full-featured wrapper. Feedback and pull requests to extend it are very welcome.
23
+ 5. **Remote MCP tools.** `Axn::RubyLLM.remote_mcp_tools` wraps a remote MCP server's tools so they sit next to your Axn tools, with an allowlist, a call budget, timeouts, and bearer-token or OAuth auth. Or let the provider connect to the server itself with `provider_tools:`.
24
+
25
+ > **Scope note:** `ask` runs one prompt through to a final answer and covers every `RubyLLM::Chat#with_*` setting, plus attachments, seeded history, streaming, and a tool-call cap (a spec fails CI when a RubyLLM release adds a method that isn't covered or explicitly skipped). For anything beyond one prompt-to-answer run, such as keeping a chat alive across turns, stepping the loop by hand, or RubyLLM's lifecycle callbacks, drive `RubyLLM::Chat` directly; see "Using wrapped tools with RubyLLM directly" below. RubyLLM's non-chat capabilities (embeddings, images, transcription, and so on) aren't wrapped. Feedback and pull requests are very welcome.
22
26
 
23
27
  ---
24
28
 
@@ -63,6 +67,42 @@ result = Axn::RubyLLM.ask(
63
67
  )
64
68
  ```
65
69
 
70
+ ### Chat options
71
+
72
+ Every other `RubyLLM::Chat#with_*` setting has a matching `ask` input. Each is forwarded only when given, so omitting it leaves RubyLLM's default in place:
73
+
74
+ | `ask` input | Forwarded to | Notes |
75
+ |---|---|---|
76
+ | `provider:`, `protocol:`, `assume_model_exists:` | `RubyLLM.chat` (alongside `model:`) | Choose the provider or wire protocol yourself instead of relying on autodetection. `assume_model_exists: true` skips the registry lookup and requires `provider:`. |
77
+ | `context:` | `RubyLLM.chat(context:)` | A `RubyLLM.context { \|c\| ... }` whose API keys and base URLs replace the global config for this call. |
78
+ | `fallbacks:`, `fallback_on:` | `with_fallbacks(*fallbacks, on: fallback_on)` | Models to try in order when generation fails. `fallback_on:` defaults to RubyLLM's transient provider and network errors. |
79
+ | `cache_system_prompt:` | `with_instructions(system_prompt, cache_until_here: true)` | Marks the system prompt as an explicit prompt-cache boundary. Worth it for a long system prompt you reuse. |
80
+ | `max_output_tokens:` | `with_max_output_tokens` | |
81
+ | `thinking:` | `with_thinking` | `true`, `false`, or `{ effort:, budget:, display: }`. |
82
+ | `citations:` | `with_citations` | Use with `attachments:` to get `raw_message.citations` back. |
83
+ | `caching:` | `with_caching` | `true`, `false`, or `{ ttl:, id: }`. |
84
+ | `compaction:` | `with_compaction` | `true`, `false`, or `{ at:, instructions:, pause_after: }`. |
85
+ | `end_user:` | `with_end_user` | An opaque id. RubyLLM sends it as given, so never pass PII. |
86
+ | `provider_options:` | `with_provider_options` | Raw keys merged into the request payload, e.g. `{ service_tier: "flex" }`. |
87
+ | `headers:` | `with_headers` | Extra HTTP headers, e.g. a provider beta flag. |
88
+
89
+ `thinking:`, `citations:`, `caching:`, and `compaction:` also forward an explicit `false`, which you'd use, for example, to turn off thinking on a model that enables it by default. `context:`, `headers:`, and `provider_tools:` are marked `sensitive`, so axn keeps them out of logs.
90
+
91
+ ### Attachments, history, and streaming
92
+
93
+ ```ruby
94
+ Axn::RubyLLM.ask(
95
+ prompt: "What changed since last quarter?",
96
+ attachments: ["q3.pdf", "https://example.com/q2.pdf"], # Chat#ask's with:
97
+ history: prior_turns, # seeded before the prompt
98
+ on_chunk: ->(chunk) { print chunk.content }, # streamed as it arrives
99
+ )
100
+ ```
101
+
102
+ - **`attachments:`** is passed to `Chat#ask` as `with:`. RubyLLM reads a String from disk or fetches it over HTTP, so never pass an unvalidated, user-supplied path or URL.
103
+ - **`history:`** is anything `Chat#messages=` accepts: `RubyLLM::Message`s, `{ role:, content: }` Hashes, or records that respond to `#to_llm`. `system_prompt:` replaces any system message in it. `transcript` includes only the turns this call exchanged, not the seeded ones.
104
+ - **`on_chunk:`** receives each `RubyLLM::Chunk`. `response` and `raw_message` are still the complete message, so `schema:` works as usual. The stubbed (disabled) path never calls it.
105
+
66
106
 
67
107
  ### Structured output via schema
68
108
 
@@ -125,11 +165,13 @@ Every successful result exposes token usage and cost, read off RubyLLM's own usa
125
165
  ```ruby
126
166
  result = Axn::RubyLLM.ask(prompt: "...")
127
167
 
128
- result.input_tokens # => 412 (non-cached input tokens only)
129
- result.cache_read_tokens # => 80 (tokens served from cache; nil if provider didn't return them)
130
- result.cache_write_tokens # => 20 (tokens written to cache; nil if provider didn't return them)
131
- result.prompt_tokens # => 512 (input_tokens + cache_read_tokens + cache_write_tokens — total request-side tokens, OpenAI-style)
132
- result.output_tokens # => 78
168
+ result.total_input_tokens # => 512 (every input token: uncached + cache reads + cache writes — use this for prompt size)
169
+ result.uncached_input_tokens # => 412 (standard-rate input only — RubyLLM's Tokens#input)
170
+ result.cache_read_tokens # => 80 (served from the provider's prompt cache)
171
+ result.cache_write_tokens # => 20 (written to the provider's prompt cache)
172
+ result.output_tokens # => 78
173
+ result.thinking_tokens # => 40 (reasoning tokens, where the provider reports them separately)
174
+ result.server_tool_use # => { "web_search_requests" => 2 } (per-use counters for provider-hosted tools)
133
175
  result.cost # => 0.00056 (Float USD total; nil if RubyLLM has no pricing for the model)
134
176
 
135
177
  # Full breakdown — RubyLLM::Cost, RubyLLM's own aggregated-cost object
@@ -139,11 +181,17 @@ result.cost_breakdown # => #<Cost input: 0.0004, output: 0.00016, cache_read: 0
139
181
  result.raw_message # => #<RubyLLM::Message ...>
140
182
  ```
141
183
 
142
- `cost` is `nil` when RubyLLM lacks pricing for the model (e.g. unknown/custom endpoints); `cost_breakdown` itself is still a `Cost` object in that case (only its component readers are `nil`). Token counts are nil only if the provider did not return them. `prompt_tokens` is nil only if all three input token fields are nil.
184
+ Use `total_input_tokens` for prompt size. With prompt caching (explicit via `caching:`, or automatic on OpenAI's newer models), a tool loop's repeated prefix is billed as cache reads and new content as cache writes, so `uncached_input_tokens` can be single digits on a 50k-token prompt. `cost` prices each bucket at its own rate, so it's right either way.
185
+
186
+ There's deliberately no plain `input_tokens` here: RubyLLM's `Tokens#input` means *uncached* input (a billing bucket), while OpenTelemetry's `gen_ai.usage.input_tokens` means *all* input, and each field name says which one it is. The older `input_tokens` (= `uncached_input_tokens`) and `prompt_tokens` (= `total_input_tokens`) still work but are deprecated; see [DEPRECATIONS.md](DEPRECATIONS.md).
187
+
188
+ `cost` is `nil` when RubyLLM lacks pricing for the model (e.g. unknown/custom endpoints); `cost_breakdown` itself is still a `Cost` object in that case (only its component readers are `nil`). Token counts are nil only if the provider did not return them. `total_input_tokens` is nil only if all three input token fields are nil.
143
189
 
144
190
  ### Errors
145
191
 
146
192
  Errors are handled via Axn's declarative `error` DSL. Every failure shares a consistent `"LLM request failed: <reason>"` headline (the headline itself is configurable via `c.error_headline =`, e.g. to `"Something went wrong calling the LLM"`; the reasons below are unaffected):
193
+ - The response hit the output token limit (`finish_reason: :max_tokens`) → `"LLM request failed: Response was cut off by the output token limit before it finished"`. This is checked before `schema:` parsing, so a truncated structured response reports this rather than a JSON error.
194
+ - A provider content filter blocked the response (`finish_reason: :content_filter`) → `"LLM request failed: Response was blocked by the provider's content filter"`
147
195
  - `JSON::ParserError` → `"LLM request failed: Response was not valid JSON"`
148
196
  - `RubyLLM::RateLimitError` (HTTP 429, provider-agnostic) → `"LLM request failed: Rate limit reached: <message>"`
149
197
  - `RubyLLM::OverloadedError` / `ServiceUnavailableError` / `ServerError` (5xx, transient) → `"LLM request failed: Provider temporarily unavailable, try again later: <message>"`
@@ -152,6 +200,8 @@ Errors are handled via Axn's declarative `error` DSL. Every failure shares a con
152
200
  - Any other known RubyLLM error — `RubyLLM::Error` (auth, bad request, payment, etc.), `RubyLLM::ConfigurationError`, `ModelNotFoundError`, `ModelRegistryError`, `PromptNotFoundError`, `InvalidRoleError`, `InvalidToolChoiceError`, `PendingToolCallsError`, `CancelledError`, `UnsupportedAttachmentError` — or `Faraday::Error` (network/transport failure) → `"LLM request failed: <message>"`
153
201
  - Any other `StandardError` (i.e. not a recognized RubyLLM/network failure — most likely a bug) → `"LLM request failed"`, with no exception detail leaked into the message
154
202
 
203
+ The truncation and content-filter failures still expose the usage fields, `cost`, `raw_message`, `transcript` and `finish_reason`, since the call was paid for. `result.finish_reason` is also set on success (normally `:stop`).
204
+
155
205
  ## Tool adapter — wrap any Axn as a RubyLLM::Tool
156
206
 
157
207
  Any [Axn](https://github.com/teamshares/axn) can be exposed as a `::RubyLLM::Tool` — no adapter-specific mixin required, it's just a normal Axn:
@@ -234,7 +284,7 @@ Passing `ambient_context:` returns a tool **instance** (closing over that contex
234
284
 
235
285
  ### Using wrapped tools with RubyLLM directly
236
286
 
237
- `Axn::RubyLLM.ask(tools:)` covers the common single-call case. When you're driving `RubyLLM.chat` yourself — multi-turn conversations, streaming, or anything else beyond `ask` — register wrapped tools with RubyLLM's own `with_tools`, which accepts one or many `RubyLLM::Tool` classes/instances:
287
+ `Axn::RubyLLM.ask` runs one prompt through to a final answer. Drive `RubyLLM.chat` yourself when you need more control than that: keeping one chat alive across turns, stepping the loop by hand (`ask_later`/`step`), or RubyLLM's lifecycle callbacks (`before_tool_call`, `after_tool_result`, `after_message`, `before_request`, `before_fallback`). Register wrapped tools with RubyLLM's own `with_tools`, which accepts one or many `RubyLLM::Tool` classes/instances:
238
288
 
239
289
  ```ruby
240
290
  chat = RubyLLM.chat.with_tools(Axn::RubyLLM.wrap(CreateWidget))
@@ -244,6 +294,90 @@ chat.ask("Create a widget called Sprocket")
244
294
  chat = RubyLLM.chat.with_tools(*Axn::RubyLLM.tools)
245
295
  ```
246
296
 
297
+ ### Provider-hosted tools and remote MCP (`provider_tools:`)
298
+
299
+ `provider_tools:` forwards verbatim to RubyLLM's `Chat#with_provider_tools` — a Hash of alias => options. This is how the *provider* (not your app) runs a tool, including a **remote MCP server**: the provider connects to the server directly, with whatever `headers:` you pass, and its own tool calls happen inside the provider's request rather than yours.
300
+
301
+ ```ruby
302
+ Axn::RubyLLM.ask(
303
+ prompt: "Why did this company's margin drop?",
304
+ provider_tools: {
305
+ mcp: {
306
+ name: "metabase", url: ENV.fetch("METABASE_MCP_URL"),
307
+ headers: { "X-API-KEY" => ENV.fetch("METABASE_MCP_API_KEY") },
308
+ allowed_tools: %w[search execute_sql],
309
+ require_approval: "never",
310
+ },
311
+ },
312
+ )
313
+ ```
314
+
315
+ See RubyLLM's `Chat#with_provider_tools` for the other aliases (`:web_search`, `:code_execution`, ...) and which providers support each one.
316
+
317
+ ### App-side remote MCP (`Axn::RubyLLM.remote_mcp_tools`)
318
+
319
+ `remote_mcp_tools` works the other way round: *your app* connects to the MCP server, using the official `mcp` gem, and wraps each of the server's tools as an ordinary `RubyLLM::Tool`. They sit next to your Axn tools in `tools:`, and every call is one your app makes, logs, times out, and can cap.
320
+
321
+ ```ruby
322
+ toolset = Axn::RubyLLM.remote_mcp_tools(
323
+ url: ENV.fetch("METABASE_MCP_URL"),
324
+ headers: { "X-API-KEY" => ENV.fetch("METABASE_MCP_API_KEY") },
325
+ allowed_tools: %w[search execute_sql], # pass it: a server's full list can include write/admin tools
326
+ max_calls: 20, # shared across the whole toolset, for its lifetime (default 20)
327
+ )
328
+ Axn::RubyLLM.ask(prompt: "...", tools: [*Axn::RubyLLM.tools, *toolset.tools])
329
+ toolset.close
330
+ ```
331
+
332
+ The `max_calls` budget never resets, so connect a fresh toolset per request (and `close` it) when you want a per-request cap; a toolset reused across requests shares one budget between them. `timeout:` (default 60s) and `max_result_chars:` (default 20,000) bound each call. When a server wants more than a static header for auth:
333
+
334
+ - **`bearer_token:`** takes a String, or a callable that returns one, and sends it as `Authorization: Bearer <token>`. A callable runs on **every request**, so your code owns fetching, caching, and rotating the token, e.g. `bearer_token: -> { MyOAuthStore.current_token }`.
335
+ - **`oauth:`** takes an `MCP::Client::OAuth::ClientCredentialsProvider` (machine-to-machine) or `MCP::Client::OAuth::Provider` (interactive authorization code + PKCE). It's passed straight to `MCP::Client::HTTP`, which handles discovery, token exchange, refresh, and retrying after a 401.
336
+
337
+ Pass one or the other, not both. Both refuse a URL that is neither https nor loopback http.
338
+
339
+ ### Pausing on approval before a provider-hosted call runs (`on_remote_tool_approval:`)
340
+
341
+ Set `require_approval: "always"` (or per-tool) on a `provider_tools:` entry and the provider pauses before running the call, asking your app to approve or deny it first. Without a decision, `ask` would return whatever the chat produced at that pause — not a final answer. Pass `on_remote_tool_approval:`, a callable given each pending `RubyLLM::ToolCall` (`.name`, `.arguments`, `.remote?`), and `ask` drives the chat past every approval until it's genuinely done:
342
+
343
+ ```ruby
344
+ Axn::RubyLLM.ask(
345
+ prompt: "...",
346
+ provider_tools: { mcp: { name: "metabase", url: ..., require_approval: "always" } },
347
+ on_remote_tool_approval: ->(tool_call) {
348
+ Rails.logger.info("[probe] #{tool_call.name}: #{tool_call.arguments}")
349
+ tool_call.name == "execute_sql" # only allow the tools you actually want to approve
350
+ },
351
+ )
352
+ ```
353
+
354
+ The same callback also resolves a **local** tool declared with `Tool.requires_approval` — `pending_approvals` doesn't distinguish where the call runs, only whether a decision is still owed.
355
+
356
+ ### Tool concurrency and other chat options (`tool_options:`)
357
+
358
+ `tool_options:` forwards verbatim to `Chat#with_tool_options` (`choice:`, `calls:`, `concurrency:`). `concurrency: :threads` or `:fibers` runs one turn's **local** Axn tool calls in parallel — useful when a tool is pure I/O (an HTTP call, a remote MCP round-trip wrapped as a local tool) and each call doesn't share mutable state. It is not a good fit for a tool that checks out an ActiveRecord connection: parallel local calls check out one connection each, from a pool sized for the process's normal concurrency.
359
+
360
+ ```ruby
361
+ Axn::RubyLLM.ask(prompt: "...", tools: [...], tool_options: { concurrency: :threads, calls: :many })
362
+ ```
363
+
364
+ ### Capping tool calls (`max_tool_calls:`)
365
+
366
+ `ask` has no limit of its own on how long the tool loop runs, because RubyLLM 2.0 removed `halt_after`. `max_tool_calls:` caps how many tool calls your app runs in one `ask`, counted across every tool in `tools:`. Once the cap is reached, each further call returns `{ error: "Tool call budget exhausted ... write your final answer with what you have so far." }` and the model is left to finish. The cap is soft on purpose: raising would throw away everything the chat had gathered. Calls to `remote_mcp_tools` tools count here as well as against their own `max_calls`. Provider-hosted calls (`provider_tools:`) never reach your app, so they aren't counted. The caller's tool instances are copied, not modified, so reusing a `wrap(axn, ambient_context:)` instance across calls is safe.
367
+
368
+ ```ruby
369
+ Axn::RubyLLM.ask(prompt: "...", tools: Axn::RubyLLM.tools, max_tool_calls: 15)
370
+ ```
371
+
372
+ ### Reading what the chat actually did (`transcript`)
373
+
374
+ `ask` always exposes `transcript`: every message the chat exchanged (the system prompt excluded), each shaped as `{ role:, content:, tool_calls:, tool_call_id:, server_tool_calls: }`. `server_tool_calls` is where a provider-hosted MCP call's name, arguments, and result show up — your app never receives that request, so this is the only place to see what query actually ran.
375
+
376
+ ```ruby
377
+ result = Axn::RubyLLM.ask(prompt: "...", provider_tools: { mcp: { ... } })
378
+ result.transcript.flat_map { |m| m[:server_tool_calls]&.values || [] }.each { |c| puts "#{c.name}: #{c.arguments}" }
379
+ ```
380
+
247
381
  ### Tool naming
248
382
 
249
383
  The name is axn core's canonical, provider-safe `tool_name`: lowercased to `[a-z0-9_]`, leading configured prefixes stripped, snake_cased with single underscores, and never blank (`Admin::CreateWidget` → `admin_create_widget`; a truly anonymous Axn → `"tool"`). Declare `axn_name "..."` on the Axn to override the default. Because it's the same core derivation every adapter uses, a class wrapped by both `Axn::RubyLLM.wrap` and `Axn::MCP.wrap` advertises an identical name — the contract is declared once.
@@ -341,9 +475,11 @@ The response is the only required argument — pass it positionally (as above) o
341
475
  stub_axn_ruby_llm({ "company_id" => 42 }, schema: CompanyMatch)
342
476
  stub_axn_ruby_llm("...", model: "gpt-4o", input_tokens: 100, output_tokens: 50, cost: 0.0023)
343
477
  stub_axn_ruby_llm("...", cache_read_tokens: 500, cache_write_tokens: 200)
478
+ stub_axn_ruby_llm("...", thinking_tokens: 300, server_tool_use: { "web_search_requests" => 1 })
479
+ stub_axn_ruby_llm("partial", finish_reason: :max_tokens) # exercises Ask's truncation failure
344
480
  ```
345
481
 
346
- Remaining keywords: `model:`, `schema:`, `input_tokens:`, `output_tokens:`, `cache_read_tokens:`, `cache_write_tokens:`, `cost:`. Returns the stubbed chat instance double for further assertions if you need it.
482
+ Remaining keywords: `model:`, `schema:`, `input_tokens:` (RubyLLM's uncached count, so it becomes `uncached_input_tokens`), `output_tokens:`, `cache_read_tokens:`, `cache_write_tokens:`, `thinking_tokens:`, `server_tool_use:`, `cost:`, and `finish_reason:` (default `:stop`). It stubs every `RubyLLM::Chat#with_*` method, so specs can pass any `Ask` option. Returns the stubbed chat instance double for further assertions if you need it.
347
483
 
348
484
  ## OpenTelemetry
349
485
 
@@ -353,10 +489,13 @@ If your app uses OpenTelemetry, `axn` already wraps every action in an `axn.call
353
489
  |---|---|
354
490
  | `gen_ai.request.model` | The model requested |
355
491
  | `gen_ai.response.model` | The model that responded |
356
- | `gen_ai.usage.input_tokens` | Non-cached input token count |
492
+ | `gen_ai.usage.input_tokens` | `total_input_tokens` — all input, cached included, per the OTel GenAI conventions |
493
+ | `gen_ai.usage.cache_read.input_tokens` | `cache_read_tokens` (omitted when the provider doesn't report it) |
494
+ | `gen_ai.usage.cache_creation.input_tokens` | `cache_write_tokens` (omitted when the provider doesn't report it) |
357
495
  | `gen_ai.usage.output_tokens` | Completion token count |
358
496
  | `gen_ai.usage.cost` | USD total (non-standard; useful for spend filtering) |
359
497
  | `axn.ruby_llm.stubbed` | `true` when production gating returned a stub |
498
+ | `axn.ruby_llm.version` | This gem's version — separates spans from before and after a change in an attribute's meaning |
360
499
  | `axn.dimension.invoked_via` | `"ruby_llm"` — set by axn core on every tool call (including nested sub-Axns), not by this gem; lets you separate tool-driven traffic from ordinary direct `.call`s in the same span schema |
361
500
 
362
501
  For LLM-level tracing (individual `RubyLLM.chat` calls, tool calls, embeddings, prompt content), add [`opentelemetry-instrumentation-ruby_llm`](https://github.com/thoughtbot/opentelemetry-instrumentation-ruby_llm) to your own Gemfile and configure it per its README. It is not a dependency of this gem.
@@ -381,7 +520,9 @@ When disabled, `Axn::RubyLLM.ask` returns a **success** result with obvious stub
381
520
  |---|---|
382
521
  | `response` | `"stubbed response value"` (plain) / `{ "stubbed" => true }` (`schema:`) |
383
522
  | `raw_message` | Stub struct with `.content`, `.tokens` (a real `RubyLLM::Tokens`, all zero), `.model`, `.parsed` |
384
- | `input_tokens` / `output_tokens` / `cache_read_tokens` / `cache_write_tokens` / `prompt_tokens` | `0` |
523
+ | `total_input_tokens` / `uncached_input_tokens` / `cache_read_tokens` / `cache_write_tokens` / `output_tokens` | `0` (the deprecated `input_tokens` / `prompt_tokens` too) |
524
+ | `thinking_tokens` / `server_tool_use` / `finish_reason` | `nil` |
525
+ | `transcript` | `[]` |
385
526
  | `cost` | `0.0` |
386
527
  | `cost_breakdown` | `nil` |
387
528
  | `stubbed` | `true` |
@@ -6,20 +6,107 @@ module Axn
6
6
  include Axn
7
7
 
8
8
  expects :prompt
9
+ # Files/URLs attached to the prompt -- Chat#ask's `with:` (a path, URL, IO, or an Array of them).
10
+ # A String is read from disk or fetched over HTTP, so never pass an unvalidated user-supplied one.
11
+ expects :attachments, optional: true
12
+ # Prior turns to seed the chat with before `prompt` -- anything Chat#messages= accepts
13
+ # (RubyLLM::Message objects, `{ role:, content: }` Hashes, or records responding to #to_llm).
14
+ # Not echoed back in `transcript`, which covers only what this call exchanged.
15
+ expects :history, optional: true
9
16
  expects :schema, optional: true
10
17
  expects :model, optional: true
18
+ # Model resolution alongside `model:` (Chat.new's own keywords): `provider:` disambiguates a
19
+ # model several providers serve, `protocol:` overrides the wire protocol, and
20
+ # `assume_model_exists: true` skips the registry lookup (requires `provider:`).
21
+ expects :provider, optional: true
22
+ expects :protocol, optional: true
23
+ expects :assume_model_exists, optional: true
24
+ # A RubyLLM::Context (RubyLLM.context { |c| ... }) whose config -- API keys, base URLs --
25
+ # replaces the process-wide RubyLLM.config for this call.
26
+ expects :context, optional: true, sensitive: true
27
+ # Ordered models to retry on when generation fails (Chat#with_fallbacks); `fallback_on:` narrows
28
+ # the triggering error classes (default: RubyLLM's transient provider/network errors).
29
+ expects :fallbacks, optional: true
30
+ expects :fallback_on, optional: true
11
31
  expects :system_prompt, optional: true
32
+ # true marks the system prompt as an explicit prompt-cache boundary
33
+ # (with_instructions(cache_until_here: true)) -- worth it for a long, reused system prompt.
34
+ expects :cache_system_prompt, optional: true
12
35
  expects :temperature, optional: true
36
+ expects :max_output_tokens, optional: true
37
+ # true, false (disable for a model that thinks by default), or `{ effort:, budget:, display: }`.
38
+ expects :thinking, optional: true
39
+ # true to get `raw_message.citations` back from attached documents.
40
+ expects :citations, optional: true
41
+ # true, false, or `{ ttl:, id: }` -- provider prompt caching (Chat#with_caching).
42
+ expects :caching, optional: true
43
+ # true, false, or `{ at:, instructions:, pause_after: }` -- provider-side context compaction.
44
+ expects :compaction, optional: true
45
+ # An opaque end-user id for the provider's abuse monitoring -- sent as given, so never PII.
46
+ expects :end_user, optional: true
47
+ # Raw keys merged into the provider request payload (Chat#with_provider_options), e.g.
48
+ # `{ service_tier: "flex" }` -- distinct from a tool's own `provider_options`.
49
+ expects :provider_options, optional: true
50
+ # Extra HTTP headers on the completion request (e.g. a provider beta flag).
51
+ expects :headers, optional: true, sensitive: true
52
+ # A callable given each streamed RubyLLM::Chunk as it arrives. `response`/`raw_message` are still
53
+ # the complete message (RubyLLM assembles it), so `schema:` works unchanged. Never called on the
54
+ # disabled (stubbed) path.
55
+ expects :on_chunk, optional: true
13
56
  expects :tools, optional: true
57
+ # Caps how many tool calls this app executes in one ask, across every tool in `tools:`
58
+ # (including remote_mcp_tools ones, which also keep their own max_calls budget). Past the cap
59
+ # each call returns an error result telling the model to answer with what it has -- see
60
+ # ToolBudget. Provider-hosted calls (provider_tools:) never reach this app and aren't counted.
61
+ expects :max_tool_calls, optional: true
62
+ # Forwarded verbatim to Chat#with_provider_tools -- a Hash of alias => options, e.g.
63
+ # `provider_tools: { mcp: { name:, url:, headers:, allowed_tools:, require_approval: } }` for a
64
+ # provider-hosted remote MCP server, or `{ web_search: {} }`. See RubyLLM::Chat#with_provider_tools.
65
+ expects :provider_tools, optional: true, sensitive: true
66
+ # Forwarded verbatim to Chat#with_tool_options (choice:/calls:/concurrency:).
67
+ expects :tool_options, optional: true
68
+ # A callable given each pending ToolCall (from Chat#pending_approvals -- local tools declared
69
+ # with `requires_approval`, or provider-hosted remote calls awaiting an mcp_approval_request) and
70
+ # returning truthy to approve, falsy to deny. Required to drive a chat past #awaiting_approval? --
71
+ # without it, `llm_response` is whatever #ask returned when the loop first parked, and the
72
+ # response schema/tool loop never completes. Only meaningful alongside provider_tools using
73
+ # require_approval, or a local tool declared with `requires_approval`.
74
+ expects :on_remote_tool_approval, optional: true
14
75
 
15
76
  exposes :response
16
77
  exposes :raw_message
17
- exposes :input_tokens, allow_nil: true
18
- exposes :output_tokens, allow_nil: true
78
+ # A plain-Hash summary of every message the chat exchanged (system prompt excluded), in order:
79
+ # role, content, tool_calls (name/arguments/remote?), tool_call_id, and server_tool_calls (the
80
+ # provider-executed calls a hosted MCP server ran, e.g. Metabase queries). Built once per call
81
+ # from Chat#messages, which RubyLLM already retains -- this only shapes it for a caller that
82
+ # doesn't want to reach into raw RubyLLM::Message objects.
83
+ exposes :transcript, type: Array, allow_blank: true
84
+ # Input token fields deliberately avoid a bare `input_tokens` name: RubyLLM's Tokens#input is a
85
+ # billing bucket (standard-rate, non-cached input only), while OTel's gen_ai.usage.input_tokens
86
+ # is the whole prompt including cached tokens. Each name here says which one it is.
87
+ #
88
+ # All input tokens: uncached + cache_read + cache_write (the OTel meaning).
89
+ exposes :total_input_tokens, allow_nil: true
90
+ # Standard-rate input only -- RubyLLM's Tokens#input.
91
+ exposes :uncached_input_tokens, allow_nil: true
19
92
  exposes :cache_read_tokens, allow_nil: true
20
93
  exposes :cache_write_tokens, allow_nil: true
94
+ exposes :output_tokens, allow_nil: true
95
+ # Reasoning tokens, where the provider reports them separately (usually also counted in
96
+ # output_tokens when billed as output).
97
+ exposes :thinking_tokens, allow_nil: true
98
+ # Per-use counters for provider-hosted tools, e.g. { "web_search_requests" => 2 } -- billed per
99
+ # use, not per token.
100
+ exposes :server_tool_use, allow_nil: true
101
+ # Deprecated (see DEPRECATIONS.md): input_tokens == uncached_input_tokens,
102
+ # prompt_tokens == total_input_tokens.
103
+ exposes :input_tokens, allow_nil: true
21
104
  exposes :prompt_tokens, allow_nil: true
22
105
  exposes :cost, allow_nil: true
106
+ # The final message's normalized finish reason (:stop, :tool_calls, :max_tokens,
107
+ # :content_filter, ...). :max_tokens and :content_filter fail the call -- see
108
+ # fail_on_incomplete_response! -- with this still exposed.
109
+ exposes :finish_reason, allow_nil: true
23
110
  exposes :cost_breakdown, allow_nil: true
24
111
  exposes :stubbed, type: :boolean, default: false
25
112
 
@@ -76,38 +163,22 @@ module Axn
76
163
  before do
77
164
  if disabled?
78
165
  exposures = stubbed_exposures
79
- record_otel_attributes!(
80
- input_tokens: exposures[:input_tokens],
81
- output_tokens: exposures[:output_tokens],
82
- cost: exposures[:cost],
83
- response_model: nil,
84
- stubbed: true,
85
- )
166
+ record_otel_attributes!(exposures, response_model: nil)
86
167
  # Reason attaches to the "LLM request completed" base via the parenthetical join: above.
87
168
  done!("using stubbed values - actual LLM request disabled", **exposures)
88
169
  end
89
170
  end
90
171
 
91
172
  def call
92
- expose(
93
- response: parsed_response,
94
- raw_message: llm_response,
95
- input_tokens: token_usage.input,
96
- output_tokens: token_usage.output,
97
- cache_read_tokens: token_usage.cache_read,
98
- cache_write_tokens: token_usage.cache_write,
99
- prompt_tokens: total_input_tokens,
100
- cost_breakdown:,
101
- cost: cost_breakdown&.total,
102
- stubbed: false,
103
- )
104
- record_otel_attributes!(
105
- input_tokens: token_usage.input,
106
- output_tokens: token_usage.output,
107
- cost: cost_breakdown&.total,
108
- response_model: llm_response&.model,
109
- stubbed: false,
110
- )
173
+ # llm_response runs the chat; usage must be read after it -- token_usage/cost_breakdown
174
+ # memoize Chat#tokens/#cost, which are an empty ledger until #ask has run.
175
+ message = llm_response
176
+ usage = usage_exposures
177
+ record_otel_attributes!(usage, response_model: message&.model)
178
+ fail_on_incomplete_response!(message, usage)
179
+
180
+ expose(response: parsed_response, raw_message: message, transcript: transcript_entries,
181
+ finish_reason: message&.finish_reason, **usage)
111
182
  rescue ::RubyLLM::RateLimitError => e
112
183
  fail! "Rate limit reached: #{e.message}"
113
184
  end
@@ -123,11 +194,17 @@ module Axn
123
194
  {
124
195
  response: parsed_content || content,
125
196
  raw_message: StubMessage.new(content:, tokens: zero_tokens, model: "stubbed"),
126
- input_tokens: 0,
127
- output_tokens: 0,
197
+ transcript: [],
198
+ total_input_tokens: 0,
199
+ uncached_input_tokens: 0,
128
200
  cache_read_tokens: 0,
129
201
  cache_write_tokens: 0,
202
+ output_tokens: 0,
203
+ input_tokens: 0,
130
204
  prompt_tokens: 0,
205
+ thinking_tokens: nil,
206
+ server_tool_use: nil,
207
+ finish_reason: nil,
131
208
  cost: 0.0,
132
209
  cost_breakdown: nil,
133
210
  stubbed: true,
@@ -153,6 +230,40 @@ module Axn
153
230
  memo def token_usage = chat.tokens
154
231
  memo def cost_breakdown = chat.cost
155
232
 
233
+ # A truncated or filtered response is not an answer: returning it as a success hands callers
234
+ # partial text (or, with schema:, a misleading JSON parse error). Checked before
235
+ # parsed_response for that reason. Usage, cost and the raw message are still exposed, since
236
+ # the call was paid for.
237
+ def fail_on_incomplete_response!(message, usage)
238
+ reason =
239
+ if message&.max_tokens?
240
+ "Response was cut off by the output token limit before it finished"
241
+ elsif message&.content_filtered?
242
+ "Response was blocked by the provider's content filter"
243
+ end
244
+ return unless reason
245
+
246
+ fail!(reason, **usage, raw_message: message, transcript: transcript_entries, finish_reason: message.finish_reason)
247
+ end
248
+
249
+ def usage_exposures
250
+ total = total_input_tokens
251
+ {
252
+ total_input_tokens: total,
253
+ uncached_input_tokens: token_usage.input,
254
+ cache_read_tokens: token_usage.cache_read,
255
+ cache_write_tokens: token_usage.cache_write,
256
+ output_tokens: token_usage.output,
257
+ thinking_tokens: token_usage.thinking,
258
+ server_tool_use: token_usage.server_tool_use,
259
+ input_tokens: token_usage.input,
260
+ prompt_tokens: total,
261
+ cost_breakdown:,
262
+ cost: cost_breakdown&.total,
263
+ stubbed: false,
264
+ }
265
+ end
266
+
156
267
  # nil only when NO turn reported the field (preserving the "nil if the provider didn't return
157
268
  # it" contract); otherwise the summed count, treating a missing component as 0.
158
269
  def total_input_tokens
@@ -160,14 +271,72 @@ module Axn
160
271
  vals.all?(&:nil?) ? nil : vals.sum(&:to_i)
161
272
  end
162
273
 
163
- memo def llm_response = chat.ask(prompt)
274
+ # Without on_remote_tool_approval, this is exactly #ask -- one #complete run, parked (like
275
+ # before this feature existed) if the chat lands on an unresolved approval. With it, drive
276
+ # #complete past every approval Chat#pending_approvals reports, until the loop truly finishes.
277
+ # Chat#complete already loops through ordinary tool calls on its own; this only extends that
278
+ # loop across the approval pauses it otherwise stops at.
279
+ memo def llm_response
280
+ message = chat.ask(prompt, with: attachments, &on_chunk)
281
+ return message unless on_remote_tool_approval
164
282
 
165
- memo def chat
166
- ::RubyLLM.chat(model: resolved_model).tap do |c|
167
- c.with_instructions(system_prompt) if system_prompt
283
+ while chat.awaiting_approval?
284
+ chat.pending_approvals.each do |tool_call|
285
+ on_remote_tool_approval.call(tool_call) ? chat.approve(tool_call) : chat.deny(tool_call)
286
+ end
287
+ message = chat.complete(&on_chunk)
288
+ end
289
+ message
290
+ end
291
+
292
+ # provider:/protocol:/assume_model_exists:/context: go to Chat.new rather than a later
293
+ # #with_model/#with_context, so the model resolves once, against the right config.
294
+ #
295
+ # thinking:/citations:/caching:/compaction: forward `false` too (`unless nil?`, not `if`) --
296
+ # false is a real instruction (e.g. turn off thinking a model enables by default).
297
+ memo def chat # rubocop:disable Metrics/AbcSize
298
+ ::RubyLLM.chat(model: resolved_model, **{ provider:, protocol:, assume_model_exists:, context: }.compact).tap do |c|
299
+ seed_history(c) if history
300
+ c.with_instructions(system_prompt, **{ cache_until_here: cache_system_prompt }.compact) if system_prompt
168
301
  c.with_schema(resolved_schema) if schema
169
302
  c.with_temperature(temperature) if temperature
303
+ c.with_max_output_tokens(max_output_tokens) if max_output_tokens
304
+ c.with_thinking(thinking) unless thinking.nil?
305
+ c.with_citations(citations) unless citations.nil?
306
+ c.with_caching(caching) unless caching.nil?
307
+ c.with_compaction(compaction) unless compaction.nil?
308
+ c.with_fallbacks(*fallbacks, **{ on: fallback_on }.compact) if fallbacks
309
+ c.with_end_user(end_user) if end_user
310
+ c.with_provider_options(provider_options) if provider_options
311
+ c.with_headers(headers) if headers
170
312
  c.with_tools(*resolved_tools) if resolved_tools.any?
313
+ c.with_provider_tools(**provider_tools) if provider_tools
314
+ c.with_tool_options(**tool_options) if tool_options
315
+ end
316
+ end
317
+
318
+ # Runs before with_instructions, since Chat#messages= replaces the whole conversation -- the
319
+ # system prompt included -- and with_instructions then replaces any system message the history
320
+ # carried. The seeded count lets `transcript` skip what the caller already had.
321
+ def seed_history(chat)
322
+ chat.messages = history
323
+ @seeded_message_count = chat.messages.count { |m| m.role != :system }
324
+ end
325
+
326
+ # chat.messages already retains the whole exchange (RubyLLM's own Chat#tokens / Chat#cost read
327
+ # off it the same way) -- this only reshapes each Message into a plain Hash a caller can log or
328
+ # assert against without reaching into RubyLLM::Message/ToolCall objects. The system prompt is
329
+ # excluded: it's an Ask input the caller already has (system_prompt), not something the loop
330
+ # produced.
331
+ def transcript_entries
332
+ chat.messages.reject { |m| m.role == :system }.drop(@seeded_message_count || 0).map do |m|
333
+ {
334
+ role: m.role,
335
+ content: m.content.is_a?(String) ? m.content : nil,
336
+ tool_calls: m.tool_calls&.transform_values { |tc| { name: tc.name, arguments: tc.arguments, remote: tc.remote? } },
337
+ tool_call_id: m.tool_call_id,
338
+ server_tool_calls: m.server_tool_calls,
339
+ }
171
340
  end
172
341
  end
173
342
 
@@ -251,8 +420,14 @@ module Axn
251
420
  # straight in) and already-wrapped `::RubyLLM::Tool`s -- a class or an instance, the latter being
252
421
  # how you pass a tool that closed over explicit context via `Axn::RubyLLM.wrap(axn, ambient_context:)`.
253
422
  # RubyLLM's `with_tools` accepts either a class or an instance, so wrapped classes register as-is.
254
- def resolved_tools
255
- Array(tools).map { |tool| _as_ruby_llm_tool(tool) }
423
+ #
424
+ # Memoized: with max_tool_calls, every guarded tool must share the ONE budget built here.
425
+ memo def resolved_tools
426
+ wrapped = Array(tools).map { |tool| _as_ruby_llm_tool(tool) }
427
+ return wrapped unless max_tool_calls
428
+
429
+ budget = ToolBudget.new(max_tool_calls)
430
+ wrapped.map { |tool| budget.guard(tool) }
256
431
  end
257
432
 
258
433
  def _as_ruby_llm_tool(tool)
@@ -262,14 +437,21 @@ module Axn
262
437
  Axn::RubyLLM.wrap(tool)
263
438
  end
264
439
 
265
- def record_otel_attributes!(input_tokens:, output_tokens:, cost:, response_model:, stubbed:)
440
+ # OTel GenAI semconv: gen_ai.usage.input_tokens "SHOULD include all types of input tokens,
441
+ # including cached tokens", with the cache counts as sub-totals of it -- so it gets
442
+ # total_input_tokens, never RubyLLM's uncached Tokens#input. annotate_span skips nil values.
443
+ # axn.ruby_llm.version separates spans from before/after a change in what an attribute means.
444
+ def record_otel_attributes!(usage, response_model:)
266
445
  Axn::Extensions::Tracing.annotate_span(
267
446
  "gen_ai.request.model" => resolved_model,
268
447
  "gen_ai.response.model" => response_model,
269
- "gen_ai.usage.input_tokens" => input_tokens,
270
- "gen_ai.usage.output_tokens" => output_tokens,
271
- "gen_ai.usage.cost" => cost,
272
- "axn.ruby_llm.stubbed" => stubbed,
448
+ "gen_ai.usage.input_tokens" => usage[:total_input_tokens],
449
+ "gen_ai.usage.cache_read.input_tokens" => usage[:cache_read_tokens],
450
+ "gen_ai.usage.cache_creation.input_tokens" => usage[:cache_write_tokens],
451
+ "gen_ai.usage.output_tokens" => usage[:output_tokens],
452
+ "gen_ai.usage.cost" => usage[:cost],
453
+ "axn.ruby_llm.stubbed" => usage[:stubbed],
454
+ "axn.ruby_llm.version" => Axn::RubyLLM::VERSION,
273
455
  )
274
456
  end
275
457
  end
@@ -0,0 +1,169 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "mcp"
4
+
5
+ module Axn
6
+ module RubyLLM
7
+ # Wraps a remote MCP server's tools as ordinary ::RubyLLM::Tool subclasses, so they can sit
8
+ # next to Axn-wrapped local tools in the same `tools:` array (Ask, or `chat.with_tools`).
9
+ #
10
+ # This is the APP-SIDE alternative to `Chat#with_provider_tools(mcp: {...})`: with that, the
11
+ # *provider* connects to the server directly and its tool calls happen inside the provider's own
12
+ # request, invisible to this app. Here, THIS process connects, so every call is one this app makes,
13
+ # logs, times out, and can cap -- at the cost of an extra round-trip per remote tool call (the
14
+ # provider calls back into this app's chat loop, rather than looping against the MCP server itself).
15
+ module RemoteMcp
16
+ DEFAULT_MAX_CALLS = 20
17
+ DEFAULT_TIMEOUT = 60
18
+ DEFAULT_MAX_RESULT_CHARS = 20_000
19
+
20
+ # Holds the connected client alongside the wrapped tools so the caller can `close` it when
21
+ # done (an MCP::Client::HTTP session is a real HTTP connection, possibly with a live SSE
22
+ # listener thread -- see MCP::Client::HTTP#close). `tools` is what actually gets passed to
23
+ # `chat.with_tools` / `Axn::RubyLLM.ask(tools:)`.
24
+ Toolset = Struct.new(:tools, :client, keyword_init: true) do
25
+ # MCP::Client itself has no #close -- only its transport does (a real HTTP connection,
26
+ # possibly with a live SSE listener thread).
27
+ def close
28
+ client.transport.close
29
+ end
30
+ end
31
+
32
+ class << self
33
+ # Connects to a remote MCP server over Streamable HTTP and returns a Toolset: one
34
+ # ::RubyLLM::Tool subclass per allowed remote tool, plus the underlying client for `close`.
35
+ #
36
+ # toolset = Axn::RubyLLM.remote_mcp_tools(
37
+ # url: ENV.fetch("METABASE_MCP_URL"),
38
+ # headers: { "X-API-KEY" => ENV.fetch("METABASE_MCP_API_KEY") },
39
+ # allowed_tools: %w[search read_resource execute_sql],
40
+ # )
41
+ # Axn::RubyLLM.ask(prompt: "...", tools: [*Axn::RubyLLM.tools, *toolset.tools])
42
+ # toolset.close
43
+ #
44
+ # allowed_tools is not a convenience -- pass it. A server's full tool list can include
45
+ # write/admin tools (e.g. Metabase's create_dashboard, update_question) that have no
46
+ # business being reachable from an LLM's own tool choices; nil (the default) keeps
47
+ # everything the server advertises, so an omitted allowlist is a deliberate "trust every
48
+ # tool this server exposes," not a safe default.
49
+ #
50
+ # Auth beyond a static `headers:` value:
51
+ #
52
+ # - `bearer_token:` -- a String, or a callable returning one, sent as `Authorization: Bearer
53
+ # <token>`. A callable runs on EVERY request (Faraday's :authorization middleware), so the
54
+ # caller owns caching/refresh/rotation (e.g. a token fetched from its own OAuth store).
55
+ # - `oauth:` -- an `MCP::Client::OAuth::ClientCredentialsProvider` (machine-to-machine) or
56
+ # `MCP::Client::OAuth::Provider` (interactive authorization code + PKCE), passed straight to
57
+ # `MCP::Client::HTTP`, which runs discovery, token exchange, refresh, and the 401 retry itself.
58
+ #
59
+ # Either one refuses a URL that is neither https nor loopback http, so a token never crosses
60
+ # the wire in plaintext.
61
+ def remote_mcp_tools(url:, headers: {}, bearer_token: nil, oauth: nil, allowed_tools: nil,
62
+ max_calls: DEFAULT_MAX_CALLS, timeout: DEFAULT_TIMEOUT, max_result_chars: DEFAULT_MAX_RESULT_CHARS)
63
+ raise ArgumentError, "pass bearer_token: or oauth:, not both" if bearer_token && oauth
64
+
65
+ client = build_client(url:, headers:, bearer_token:, oauth:, timeout:)
66
+ begin
67
+ client.connect
68
+ remote_tools = client.tools
69
+ remote_tools = remote_tools.select { |t| allowed_tools.include?(t.name) } if allowed_tools
70
+
71
+ # One budget per toolset, not per tool: the limit bounds total round-trips to this server, not
72
+ # calls to any single tool. It lives as long as the toolset and never resets, so connect a
73
+ # toolset per request (as the example above does) for a per-request cap.
74
+ budget = ToolBudget.new(max_calls, noun: "remote calls")
75
+ wrapped = remote_tools.map { |remote_tool| build_tool_class(remote_tool, client:, budget:, max_result_chars:) }
76
+
77
+ Toolset.new(tools: wrapped, client:)
78
+ rescue StandardError
79
+ # Until a Toolset is returned the caller has nothing to #close, so a failed handshake or
80
+ # tools/list would otherwise leak the HTTP session (and any SSE listener thread).
81
+ close_quietly(client.transport)
82
+ raise
83
+ end
84
+ end
85
+
86
+ private
87
+
88
+ def build_client(url:, headers:, bearer_token:, oauth:, timeout:)
89
+ # MCP::Client::HTTP enforces this itself for oauth:, but knows nothing about a token set by
90
+ # the Faraday middleware below.
91
+ raise ArgumentError, "bearer_token: requires an https (or loopback http) MCP URL" if bearer_token && !::MCP::Client::OAuth::Discovery.secure_url?(url)
92
+
93
+ transport = ::MCP::Client::HTTP.new(url:, headers:, **{ oauth: }.compact) do |faraday|
94
+ faraday.options.timeout = timeout
95
+ faraday.options.open_timeout = timeout
96
+ faraday.request :authorization, "Bearer", bearer_token if bearer_token
97
+ end
98
+ ::MCP::Client.new(transport:)
99
+ end
100
+
101
+ # Cleanup on the failure path must never replace the error that got us there.
102
+ def close_quietly(transport)
103
+ transport.close
104
+ rescue StandardError => e
105
+ Axn::Extensions.best_effort("logging a failed remote MCP transport close") do
106
+ Axn.config.logger.warn { "[axn-ruby_llm] closing remote MCP transport after a failed setup also failed: #{e.class}: #{e.message}" }
107
+ end
108
+ end
109
+
110
+ def build_tool_class(remote_tool, client:, budget:, max_result_chars:)
111
+ tool_name = remote_tool.name
112
+ tool_description = remote_tool.description
113
+ input_schema = remote_tool.input_schema || { type: "object", properties: {} }
114
+
115
+ Class.new(::RubyLLM::Tool) do
116
+ description(tool_description) if tool_description
117
+ parameters(input_schema)
118
+
119
+ define_singleton_method(:tool_name) { tool_name }
120
+
121
+ define_method(:execute) do |**args|
122
+ Axn::RubyLLM::RemoteMcp.send(:call_remote_tool, client:, tool_name:, args:, budget:, max_result_chars:)
123
+ end
124
+ end
125
+ end
126
+
127
+ # Isolated from build_tool_class's closure (rather than inlined in the define_method block)
128
+ # so the rescue clauses -- the part most likely to need a new case as real servers are
129
+ # exercised -- are easy to find and extend on their own, not buried inside class-building
130
+ # metaprogramming.
131
+ def call_remote_tool(client:, tool_name:, args:, budget:, max_result_chars:)
132
+ return budget.exhausted_result unless budget.consume!
133
+
134
+ response = client.call_tool(name: tool_name, arguments: args)
135
+ render_tool_result(response, max_result_chars:)
136
+ rescue ::MCP::Client::ServerError => e
137
+ { error: "Remote tool error: #{e.message}" }
138
+ rescue ::Faraday::TimeoutError
139
+ { error: "Remote tool call timed out" }
140
+ rescue ::Faraday::Error, ::MCP::Client::RequestHandlerError => e
141
+ { error: "Remote tool call failed: #{e.message}" }
142
+ rescue StandardError => e
143
+ # RubyLLM has no rescue around a tool's #execute (axn-ruby_llm's own ToolAdapter guard
144
+ # exists for exactly this reason on the Axn side) -- an unanticipated MCP client error
145
+ # here (a malformed response, a client-side bug) must not escape and break the whole chat.
146
+ # best_effort: a broken configured logger must not defeat this boundary either.
147
+ Axn::Extensions.best_effort("logging a remote MCP tool failure") do
148
+ Axn.config.logger.error { "[axn-ruby_llm] remote MCP tool #{tool_name.inspect} failed: #{e.class}: #{e.message}" }
149
+ end
150
+ { error: "The remote tool could not produce a valid response" }
151
+ end
152
+
153
+ # The MCP result shape is `{"result" => {"content" => [{"type" => "text", "text" => "..."}, ...], "isError" => bool}}`
154
+ # (see MCP::Client#call_tool's own doc example: `response.dig("result", "content")`). Only
155
+ # text blocks are supported for now -- Metabase's tools return text/JSON, and a truncated
156
+ # binary block would be meaningless to the model anyway.
157
+ def render_tool_result(response, max_result_chars:)
158
+ result = response["result"] || {}
159
+ text = Array(result["content"]).filter_map { |block| block["text"] }.join("\n")
160
+ text = "#{text[0...max_result_chars]}\n... (truncated at #{max_result_chars} characters)" if text.length > max_result_chars
161
+
162
+ return { error: text.empty? ? "Tool call failed" : text } if result["isError"]
163
+
164
+ text
165
+ end
166
+ end
167
+ end
168
+ end
169
+ end
@@ -16,20 +16,26 @@ module Axn
16
16
  # stub_axn_ruby_llm("Here is a summary.")
17
17
  # stub_axn_ruby_llm({ "k" => "v" }, schema: MySchema) # Hash passed through as `parsed`
18
18
  # stub_axn_ruby_llm("...", input_tokens: 100, output_tokens: 50, cost: 0.0023)
19
+ # (input_tokens: here is RubyLLM's uncached Tokens#input -- Ask's uncached_input_tokens)
19
20
  # stub_axn_ruby_llm("...", cache_read_tokens: 500, cache_write_tokens: 200)
21
+ # stub_axn_ruby_llm("...", thinking_tokens: 300, server_tool_use: { "web_search_requests" => 1 })
22
+ # stub_axn_ruby_llm("partial", finish_reason: :max_tokens) # exercise Ask's truncation failure
20
23
  # stub_axn_ruby_llm(response: "...") # keyword form still works
21
24
  #
22
25
  # Returns the chat instance double for further assertions if needed.
26
+ # rubocop:disable-next Metrics/ParameterLists -- one optional keyword per stubbable usage field
23
27
  def stub_axn_ruby_llm(positional_response = UNSET, response: UNSET, model: nil, schema: nil,
24
28
  input_tokens: nil, output_tokens: nil, cache_read_tokens: nil,
25
- cache_write_tokens: nil, cost: nil)
29
+ cache_write_tokens: nil, thinking_tokens: nil, server_tool_use: nil,
30
+ cost: nil, finish_reason: :stop)
26
31
  response = positional_response unless positional_response.equal?(UNSET)
27
32
  raise ArgumentError, "stub_axn_ruby_llm requires a response (positionally or as `response:`)" if response.equal?(UNSET)
28
33
 
29
34
  resolved_model_id = model || Axn::RubyLLM.config.default_model
30
- llm_message = _stub_axn_ruby_llm_message(response, resolved_model_id, schema:)
31
- _stub_axn_ruby_llm_chat(model, llm_message, input_tokens:, output_tokens:,
32
- cache_read_tokens:, cache_write_tokens:, cost:)
35
+ llm_message = _stub_axn_ruby_llm_message(response, resolved_model_id, schema:, finish_reason:)
36
+ tokens = ::RubyLLM::Tokens.new(input: input_tokens, output: output_tokens, cache_read: cache_read_tokens,
37
+ cache_write: cache_write_tokens, thinking: thinking_tokens, server_tool_use:)
38
+ _stub_axn_ruby_llm_chat(model, llm_message, tokens:, cost:)
33
39
  end
34
40
 
35
41
  private
@@ -37,30 +43,40 @@ module Axn
37
43
  # `content` mirrors real ::RubyLLM::Message#content (a read-only String, JSON text when
38
44
  # `schema:` is set); `parsed` mirrors #parsed (the Hash `schema:` callers actually want back
39
45
  # via Ask's `parsed_response`, which reads `.parsed` -- not `.content` -- once schema is set).
40
- def _stub_axn_ruby_llm_message(response, model_id, schema:)
46
+ # finish_reason/max_tokens?/content_filtered? mirror the real predicates, so a helper-stubbed
47
+ # call goes through Ask's truncation/filter check like a real one.
48
+ def _stub_axn_ruby_llm_message(response, model_id, schema:, finish_reason:)
41
49
  content = schema ? response.to_json : response.to_s
42
50
  parsed = schema ? response : nil
43
- instance_double(::RubyLLM::Message, content:, parsed:, model: model_id)
51
+ instance_double(::RubyLLM::Message, content:, parsed:, model: model_id, finish_reason:,
52
+ max_tokens?: finish_reason == :max_tokens,
53
+ content_filtered?: finish_reason == :content_filter)
44
54
  end
45
55
 
46
- def _stub_axn_ruby_llm_chat(model, llm_message, input_tokens:, output_tokens:,
47
- cache_read_tokens:, cache_write_tokens:, cost:)
56
+ def _stub_axn_ruby_llm_chat(model, llm_message, tokens:, cost:)
48
57
  chat_instance = instance_double(::RubyLLM::Chat)
49
58
  if model
50
- allow(::RubyLLM).to receive(:chat).with(model:).and_return(chat_instance)
59
+ # hash_including: Ask also passes provider:/protocol:/assume_model_exists:/context: when set.
60
+ allow(::RubyLLM).to receive(:chat).with(hash_including(model:)).and_return(chat_instance)
51
61
  else
52
62
  allow(::RubyLLM).to receive(:chat).and_return(chat_instance)
53
63
  end
54
- %i[with_instructions with_schema with_temperature with_provider_options with_tools].each do |method|
64
+ # Every public with_* on the real Chat, not a hand-kept list -- Ask forwards whichever inputs
65
+ # the caller set, and the verifying double rejects any method left unstubbed.
66
+ ::RubyLLM::Chat.public_instance_methods(false).grep(/\Awith_/).each do |method|
55
67
  allow(chat_instance).to receive(method).and_return(chat_instance)
56
68
  end
69
+ allow(chat_instance).to receive(:messages=) # history:
70
+ allow(chat_instance).to receive(:awaiting_approval?).and_return(false) # on_remote_tool_approval:
57
71
  allow(chat_instance).to receive(:ask).and_return(llm_message)
72
+ # A stubbed call has no real conversation for transcript_entries to walk -- matches how the
73
+ # disabled/stubbed-config path (Ask#stubbed_exposures) also exposes an empty transcript.
74
+ allow(chat_instance).to receive(:messages).and_return([])
58
75
  # Ask reads usage off the chat's own ledger (Chat#tokens / Chat#cost), not per-message --
59
76
  # a stubbed call is single-turn, so the ledger is just these values directly.
60
- allow(chat_instance).to receive(:tokens).and_return(
61
- instance_double(::RubyLLM::Tokens, input: input_tokens, output: output_tokens,
62
- cache_read: cache_read_tokens, cache_write: cache_write_tokens),
63
- )
77
+ # A real ::RubyLLM::Tokens (a plain value object), not a double, so new ledger fields read as nil
78
+ # instead of tripping a verifying double.
79
+ allow(chat_instance).to receive(:tokens).and_return(tokens)
64
80
  # Default to zero cost so specs exercise the "cost computed" path.
65
81
  # Pass cost: explicitly to assert a specific value.
66
82
  allow(chat_instance).to receive(:cost).and_return(instance_double(::RubyLLM::Cost, total: cost || 0.0))
@@ -0,0 +1,57 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Axn
4
+ module RubyLLM
5
+ # A shared cap on how many tool calls one chat may make. Backs both `remote_mcp_tools(max_calls:)`
6
+ # (one budget per toolset) and `Ask`'s `max_tool_calls:` (one budget per `ask`, across every
7
+ # app-executed tool). RubyLLM 2.0 removed `Tool::Halt` / `halt_after:`, so `Chat#complete` has no
8
+ # iteration cap of its own.
9
+ #
10
+ # Soft by design: an exhausted budget answers each further call with an error result telling the
11
+ # model to wrap up, rather than raising -- running out is a normal end state for a long research
12
+ # loop, and raising would throw away everything the chat gathered so far. The Mutex keeps the
13
+ # count exact under `tool_options: { concurrency: :threads }`.
14
+ class ToolBudget
15
+ attr_reader :max_calls
16
+
17
+ def initialize(max_calls, noun: "tool calls")
18
+ @max_calls = max_calls
19
+ @noun = noun
20
+ @count = 0
21
+ @mutex = Mutex.new
22
+ end
23
+
24
+ # Returns true (and reserves a slot) when a call is still allowed, false once max_calls has
25
+ # been reached.
26
+ def consume!
27
+ @mutex.synchronize do
28
+ return false if @count >= @max_calls
29
+
30
+ @count += 1
31
+ true
32
+ end
33
+ end
34
+
35
+ def exhausted_result
36
+ { error: "Tool call budget exhausted (#{@max_calls} #{@noun} allowed) -- " \
37
+ "write your final answer with what you have so far." }
38
+ end
39
+
40
+ # Returns a tool instance whose every call draws from this budget. A class is instantiated; an
41
+ # instance is dup'd first, so the caller's own instance (e.g. one from
42
+ # `Axn::RubyLLM.wrap(axn, ambient_context:)`, possibly reused across asks) is never mutated.
43
+ # RubyLLM's Chat#with_tools registers an instance as-is, and runs it via Tool#call.
44
+ def guard(tool)
45
+ budget = self
46
+ instance = tool.is_a?(::Class) ? tool.new : tool.dup
47
+ instance.extend(Module.new do
48
+ define_method(:call) do |tool_call: nil, **arguments|
49
+ next budget.exhausted_result unless budget.consume!
50
+
51
+ super(tool_call:, **arguments)
52
+ end
53
+ end)
54
+ end
55
+ end
56
+ end
57
+ end
@@ -2,6 +2,6 @@
2
2
 
3
3
  module Axn
4
4
  module RubyLLM
5
- VERSION = "0.3.0"
5
+ VERSION = "0.4.0"
6
6
  end
7
7
  end
data/lib/axn/ruby_llm.rb CHANGED
@@ -4,7 +4,9 @@ require "ruby_llm"
4
4
  require "axn"
5
5
 
6
6
  require_relative "ruby_llm/version"
7
+ require_relative "ruby_llm/tool_budget"
7
8
  require_relative "ruby_llm/ask"
9
+ require_relative "ruby_llm/remote_mcp"
8
10
 
9
11
  module Axn
10
12
  module RubyLLM
@@ -43,6 +45,10 @@ module Axn
43
45
  value = config.enabled
44
46
  value.respond_to?(:call) ? !!value.call : !!value
45
47
  end
48
+
49
+ def remote_mcp_tools(...)
50
+ RemoteMcp.remote_mcp_tools(...)
51
+ end
46
52
  end
47
53
  end
48
54
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: axn-ruby_llm
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.3.0
4
+ version: 0.4.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Kali Donovan
@@ -63,6 +63,26 @@ dependencies:
63
63
  - - ">="
64
64
  - !ruby/object:Gem::Version
65
65
  version: 1.10.0
66
+ - !ruby/object:Gem::Dependency
67
+ name: mcp
68
+ requirement: !ruby/object:Gem::Requirement
69
+ requirements:
70
+ - - ">="
71
+ - !ruby/object:Gem::Version
72
+ version: '1.3'
73
+ - - "<"
74
+ - !ruby/object:Gem::Version
75
+ version: '2.0'
76
+ type: :runtime
77
+ prerelease: false
78
+ version_requirements: !ruby/object:Gem::Requirement
79
+ requirements:
80
+ - - ">="
81
+ - !ruby/object:Gem::Version
82
+ version: '1.3'
83
+ - - "<"
84
+ - !ruby/object:Gem::Version
85
+ version: '2.0'
66
86
  description: Call LLMs from Axn actions using RubyLLM, with structured error handling,
67
87
  schema-based structured output, and cost/token tracking.
68
88
  email:
@@ -77,8 +97,10 @@ files:
77
97
  - lib/axn-ruby_llm.rb
78
98
  - lib/axn/ruby_llm.rb
79
99
  - lib/axn/ruby_llm/ask.rb
100
+ - lib/axn/ruby_llm/remote_mcp.rb
80
101
  - lib/axn/ruby_llm/rspec.rb
81
102
  - lib/axn/ruby_llm/tool_adapter.rb
103
+ - lib/axn/ruby_llm/tool_budget.rb
82
104
  - lib/axn/ruby_llm/version.rb
83
105
  homepage: https://github.com/teamshares/axn-ruby_llm
84
106
  licenses: