axn-ruby_llm 0.3.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +106 -0
- data/README.md +157 -16
- data/lib/axn/ruby_llm/ask.rb +223 -41
- data/lib/axn/ruby_llm/remote_mcp.rb +169 -0
- data/lib/axn/ruby_llm/rspec.rb +30 -14
- data/lib/axn/ruby_llm/tool_budget.rb +57 -0
- data/lib/axn/ruby_llm/version.rb +1 -1
- data/lib/axn/ruby_llm.rb +6 -0
- metadata +23 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 6976e9d15c05e921240e525fa9a421ea5af2f66bee88fe182d51e6dcf042e6a1
|
|
4
|
+
data.tar.gz: c979116f6bde6b20a5cf89664c9d9c04180f7165f48a2685e86db7db0385becc
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 77698ef5d994d06e264b56845b507721058c2945357342182c9557582a9ea20572dc47164182fd541e2636115485fd46dcfb77ee63ff2646279d4b5805d6cd57
|
|
7
|
+
data.tar.gz: ebbae53a4026c405b5bb2636169bb916c1b77b085f1f06872ccc609c25f05b01050d59ea4d4e29871035e24d2ed4f572a8afb28dceb1b473683748460962f1ee
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,111 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## [Unreleased]
|
|
4
|
+
|
|
5
|
+
## [0.4.0] - 2026-09-25
|
|
6
|
+
|
|
7
|
+
`Ask` now covers all of `RubyLLM::Chat`'s settings, plus attachments, seeded history, streaming, and
|
|
8
|
+
a tool-call cap, and the gem gains app-side remote MCP tools. Most of this is additive, but three
|
|
9
|
+
behaviour changes need checking on upgrade.
|
|
10
|
+
|
|
11
|
+
### ⚠️ Breaking changes
|
|
12
|
+
|
|
13
|
+
- **Truncated or content-filtered responses now fail instead of succeeding.** A response cut off by
|
|
14
|
+
the output token limit (`finish_reason: :max_tokens`) now fails with "Response was cut off by the
|
|
15
|
+
output token limit before it finished". One blocked by a provider content filter fails with
|
|
16
|
+
"Response was blocked by the provider's content filter". Both used to come back as a success
|
|
17
|
+
carrying partial or empty text; with `schema:`, truncation surfaced as a misleading "Response was
|
|
18
|
+
not valid JSON". Usage, `cost`, `raw_message`, `transcript` and `finish_reason` are still exposed
|
|
19
|
+
on the failure.
|
|
20
|
+
**What to check:** callers that relied on partial text from a truncated response should handle
|
|
21
|
+
the failure, or raise `max_output_tokens:`.
|
|
22
|
+
- **The OpenTelemetry attribute `gen_ai.usage.input_tokens` now includes cached tokens.** It used to
|
|
23
|
+
carry RubyLLM's uncached count, which under-counts badly under prompt caching (e.g. 9 instead of
|
|
24
|
+
~58k on a `gpt-5.6` tool loop). It now carries `total_input_tokens` (uncached + cache reads + cache
|
|
25
|
+
writes), matching the OTel GenAI definition. New span attributes:
|
|
26
|
+
`gen_ai.usage.cache_read.input_tokens`, `gen_ai.usage.cache_creation.input_tokens`, and
|
|
27
|
+
`axn.ruby_llm.version`.
|
|
28
|
+
**What to check:** dashboards or alerts summing `gen_ai.usage.input_tokens` will step up as each
|
|
29
|
+
service upgrades. Filter on `@axn.ruby_llm.version:*` to see only spans with the new meaning. Cost
|
|
30
|
+
attributes are unchanged.
|
|
31
|
+
- **Specs that stub `RubyLLM::Chat` or `RubyLLM::Message` directly** (rather than using
|
|
32
|
+
`stub_axn_ruby_llm`) may need updating. `Ask` now calls `chat.ask(prompt, with: nil)`, reads
|
|
33
|
+
`finish_reason`, `max_tokens?` and `content_filtered?` on the final message, and reads `thinking`
|
|
34
|
+
and `server_tool_use` from `chat.tokens`. A verifying double missing these fails with an
|
|
35
|
+
unexpected-message error.
|
|
36
|
+
**What to check:** switch those specs to `stub_axn_ruby_llm`, which handles all of this.
|
|
37
|
+
|
|
38
|
+
### Deprecated
|
|
39
|
+
|
|
40
|
+
- **`Ask`'s `input_tokens` and `prompt_tokens` exposures**, replaced by `uncached_input_tokens` and
|
|
41
|
+
`total_input_tokens` respectively, with unchanged values. The bare `input_tokens` name meant
|
|
42
|
+
uncached input in RubyLLM's convention but all input in OpenTelemetry's, so it's retired rather
|
|
43
|
+
than redefined. Both still work; removal is scheduled for 1.0 (see `DEPRECATIONS.md`).
|
|
44
|
+
|
|
45
|
+
### Added
|
|
46
|
+
|
|
47
|
+
**`Ask` options.** Every public `RubyLLM::Chat#with_*` method now has a matching input, forwarded
|
|
48
|
+
only when given:
|
|
49
|
+
|
|
50
|
+
- `provider:`, `protocol:`, `assume_model_exists:`, `context:` → `RubyLLM.chat(...)`, alongside
|
|
51
|
+
`model:`
|
|
52
|
+
- `fallbacks:` / `fallback_on:` → `with_fallbacks(*fallbacks, on:)`
|
|
53
|
+
- `max_output_tokens:` → `with_max_output_tokens`
|
|
54
|
+
- `thinking:` → `with_thinking`; `citations:` → `with_citations`; `caching:` → `with_caching`;
|
|
55
|
+
`compaction:` → `with_compaction`. These four also forward an explicit `false`.
|
|
56
|
+
- `end_user:` → `with_end_user`; `provider_options:` → `with_provider_options`; `headers:` →
|
|
57
|
+
`with_headers`
|
|
58
|
+
- `provider_tools:` → `with_provider_tools` (e.g. a provider-hosted remote MCP server:
|
|
59
|
+
`{ mcp: { name:, url:, headers:, allowed_tools:, require_approval: } }`); `tool_options:` →
|
|
60
|
+
`with_tool_options`
|
|
61
|
+
- `cache_system_prompt: true` → `with_instructions(..., cache_until_here: true)`
|
|
62
|
+
|
|
63
|
+
**Beyond the `with_*` methods:**
|
|
64
|
+
|
|
65
|
+
- `attachments:` is passed to `Chat#ask` as `with:` (paths, URLs, or IO).
|
|
66
|
+
- `history:` seeds earlier turns before the prompt; `transcript` leaves them out.
|
|
67
|
+
- `on_chunk:` receives each streamed `RubyLLM::Chunk`. `response` is still the complete message, so
|
|
68
|
+
`schema:` is unaffected.
|
|
69
|
+
- `max_tool_calls:` is a soft cap on app-executed tool calls per `ask`, shared across every tool.
|
|
70
|
+
Past the cap, each call returns an error result telling the model to wrap up.
|
|
71
|
+
- `on_remote_tool_approval:` approves or denies each call waiting on `Chat#awaiting_approval?`
|
|
72
|
+
(provider-hosted calls with `require_approval:`, or local tools declared with
|
|
73
|
+
`Tool.requires_approval`) and drives the chat to completion.
|
|
74
|
+
|
|
75
|
+
**New result fields:** `transcript` (every message the chat exchanged, including the provider-hosted
|
|
76
|
+
calls a remote MCP server ran), `total_input_tokens`, `uncached_input_tokens`, `thinking_tokens`,
|
|
77
|
+
`server_tool_use` (per-use counters for provider-hosted tools, such as web searches), and
|
|
78
|
+
`finish_reason`.
|
|
79
|
+
|
|
80
|
+
**`Axn::RubyLLM.remote_mcp_tools(url:, ...)`** connects to a remote MCP server from this app (using
|
|
81
|
+
the official `mcp` gem) and wraps its tools as `RubyLLM::Tool`s, so they sit next to Axn tools in
|
|
82
|
+
`tools:`.
|
|
83
|
+
|
|
84
|
+
- `allowed_tools:` limits which of the server's tools are exposed; pass it, since a server's full
|
|
85
|
+
list can include write/admin tools.
|
|
86
|
+
- `max_calls:` is a call budget shared across the toolset for its lifetime; connect one toolset per
|
|
87
|
+
request for a per-request cap. `timeout:` and `max_result_chars:` bound each call.
|
|
88
|
+
- Auth: static `headers:`; `bearer_token:` (a String, or a callable run on every request so the
|
|
89
|
+
caller can rotate it); or `oauth:` (an `mcp`-gem `OAuth::Provider` / `ClientCredentialsProvider`,
|
|
90
|
+
passed through). Token auth requires an https or loopback URL.
|
|
91
|
+
- If setup fails, the connection is closed before the error propagates.
|
|
92
|
+
|
|
93
|
+
**`stub_axn_ruby_llm`** accepts `finish_reason:`, `thinking_tokens:` and `server_tool_use:`.
|
|
94
|
+
|
|
95
|
+
**RubyLLM parity spec.** `spec/axn/ruby_llm/ruby_llm_parity_spec.rb` classifies every public method
|
|
96
|
+
on `RubyLLM::Chat`, `RubyLLM::Tokens` and `RubyLLM::Message` as covered or intentionally skipped,
|
|
97
|
+
and snapshots the signatures of the `Chat` methods `Ask` calls. A `ruby_llm` bump that adds or
|
|
98
|
+
re-signatures one fails CI by name.
|
|
99
|
+
|
|
100
|
+
### Changed
|
|
101
|
+
|
|
102
|
+
- `stub_axn_ruby_llm` now stubs every public `RubyLLM::Chat#with_*` method (taken from the class
|
|
103
|
+
itself, so future ones are covered), plus `messages=` and `awaiting_approval?`, and matches
|
|
104
|
+
`model:` alongside the model-resolution keywords. It returns a real `RubyLLM::Tokens` rather than
|
|
105
|
+
a verifying double.
|
|
106
|
+
- `Ask`'s `provider_tools:`, `headers:` and `context:` inputs are marked `sensitive`, so axn keeps
|
|
107
|
+
them (and the API keys they usually carry) out of logs.
|
|
108
|
+
|
|
3
109
|
## [0.3.0] - 2026-09-22
|
|
4
110
|
|
|
5
111
|
RubyLLM 2.0 is a breaking rewrite (renamed Tool DSL, restructured error hierarchy, a usage ledger
|
data/README.md
CHANGED
|
@@ -2,23 +2,27 @@
|
|
|
2
2
|
|
|
3
3
|
Call LLMs from [Axn](https://github.com/teamshares/axn) actions using [RubyLLM](https://github.com/crmne/ruby_llm), with declarative error handling, schema-based structured output, configurable defaults, and cost/token tracking — and wrap any Axn as a `RubyLLM::Tool` a chat can call.
|
|
4
4
|
|
|
5
|
-
> **RubyLLM 2.0 required.** As of `0.3.0`, this gem requires `ruby_llm >= 2.0, < 3.0` and no longer supports RubyLLM 1.x
|
|
5
|
+
> **RubyLLM 2.0 required.** As of `0.3.0`, this gem requires `ruby_llm >= 2.0, < 3.0` and no longer supports RubyLLM 1.x.
|
|
6
|
+
>
|
|
7
|
+
> **Upgrading to 0.4.0?** Two behaviour changes: a response cut off by the output token limit or blocked by a content filter now **fails** instead of succeeding, and the OpenTelemetry attribute `gen_ai.usage.input_tokens` now includes cached tokens. `input_tokens` / `prompt_tokens` are deprecated in favour of `uncached_input_tokens` / `total_input_tokens`. See [CHANGELOG.md](CHANGELOG.md) for the full list and what to check.
|
|
6
8
|
|
|
7
9
|
Part of the `axn-*` extension ecosystem — see also [axn-mcp](https://github.com/teamshares/axn-mcp).
|
|
8
10
|
|
|
9
11
|
### Why use this over calling RubyLLM directly?
|
|
10
12
|
|
|
11
|
-
|
|
13
|
+
Five things you'd otherwise hand-build:
|
|
12
14
|
|
|
13
|
-
1. **Structured error handling.** The Axn error DSL declaratively maps `RateLimitError`, `JSON::ParserError`, and generic `StandardError` to clean failure messages. Callers check `result.ok?` instead of wrapping every call in `begin/rescue`.
|
|
15
|
+
1. **Structured error handling.** The Axn error DSL declaratively maps `RateLimitError`, `JSON::ParserError`, truncated or content-filtered responses, and generic `StandardError` to clean failure messages. Callers check `result.ok?` instead of wrapping every call in `begin/rescue`.
|
|
14
16
|
|
|
15
|
-
2. **Production gating.** A single `c.enabled = -> { Rails.env.production? }` in an initializer stubs every LLM call in non-prod environments — no per-callsite guards needed. The stub is typed (`stubbed: true`, `
|
|
17
|
+
2. **Production gating.** A single `c.enabled = -> { Rails.env.production? }` in an initializer stubs every LLM call in non-prod environments — no per-callsite guards needed. The stub is typed (`stubbed: true`, `total_input_tokens: 0`, etc.) so downstream code doesn't need to branch on it either.
|
|
16
18
|
|
|
17
|
-
3. **Cost/token tracking, exposed automatically.** Every call exposes `
|
|
19
|
+
3. **Cost/token tracking, exposed automatically.** Every call exposes `total_input_tokens`, `uncached_input_tokens`, `cache_read_tokens`, `cache_write_tokens`, `output_tokens`, `thinking_tokens`, `server_tool_use`, `cost`, and `cost_breakdown`, read straight off RubyLLM's own usage ledger (`Chat#tokens` / `Chat#cost`) — no manual model lookup, and a tool-call loop's retries and multiple round-trips are already aggregated for you. If your app uses OpenTelemetry, these values are also set as attributes on the existing `axn.call` span — no configuration required.
|
|
18
20
|
|
|
19
21
|
4. **Author-once tools.** `Axn::RubyLLM.wrap` turns any Axn into a `RubyLLM::Tool` your chat can call — reuse the same Axn classes you already expose through [axn-mcp](https://github.com/teamshares/axn-mcp), or plain Axns, with no rewrite. The tool's name, JSON Schema, and argument validation all come from the Axn's own contract.
|
|
20
22
|
|
|
21
|
-
|
|
23
|
+
5. **Remote MCP tools.** `Axn::RubyLLM.remote_mcp_tools` wraps a remote MCP server's tools so they sit next to your Axn tools, with an allowlist, a call budget, timeouts, and bearer-token or OAuth auth. Or let the provider connect to the server itself with `provider_tools:`.
|
|
24
|
+
|
|
25
|
+
> **Scope note:** `ask` runs one prompt through to a final answer and covers every `RubyLLM::Chat#with_*` setting, plus attachments, seeded history, streaming, and a tool-call cap (a spec fails CI when a RubyLLM release adds a method that isn't covered or explicitly skipped). For anything beyond one prompt-to-answer run, such as keeping a chat alive across turns, stepping the loop by hand, or RubyLLM's lifecycle callbacks, drive `RubyLLM::Chat` directly; see "Using wrapped tools with RubyLLM directly" below. RubyLLM's non-chat capabilities (embeddings, images, transcription, and so on) aren't wrapped. Feedback and pull requests are very welcome.
|
|
22
26
|
|
|
23
27
|
---
|
|
24
28
|
|
|
@@ -63,6 +67,42 @@ result = Axn::RubyLLM.ask(
|
|
|
63
67
|
)
|
|
64
68
|
```
|
|
65
69
|
|
|
70
|
+
### Chat options
|
|
71
|
+
|
|
72
|
+
Every other `RubyLLM::Chat#with_*` setting has a matching `ask` input. Each is forwarded only when given, so omitting it leaves RubyLLM's default in place:
|
|
73
|
+
|
|
74
|
+
| `ask` input | Forwarded to | Notes |
|
|
75
|
+
|---|---|---|
|
|
76
|
+
| `provider:`, `protocol:`, `assume_model_exists:` | `RubyLLM.chat` (alongside `model:`) | Choose the provider or wire protocol yourself instead of relying on autodetection. `assume_model_exists: true` skips the registry lookup and requires `provider:`. |
|
|
77
|
+
| `context:` | `RubyLLM.chat(context:)` | A `RubyLLM.context { \|c\| ... }` whose API keys and base URLs replace the global config for this call. |
|
|
78
|
+
| `fallbacks:`, `fallback_on:` | `with_fallbacks(*fallbacks, on: fallback_on)` | Models to try in order when generation fails. `fallback_on:` defaults to RubyLLM's transient provider and network errors. |
|
|
79
|
+
| `cache_system_prompt:` | `with_instructions(system_prompt, cache_until_here: true)` | Marks the system prompt as an explicit prompt-cache boundary. Worth it for a long system prompt you reuse. |
|
|
80
|
+
| `max_output_tokens:` | `with_max_output_tokens` | |
|
|
81
|
+
| `thinking:` | `with_thinking` | `true`, `false`, or `{ effort:, budget:, display: }`. |
|
|
82
|
+
| `citations:` | `with_citations` | Use with `attachments:` to get `raw_message.citations` back. |
|
|
83
|
+
| `caching:` | `with_caching` | `true`, `false`, or `{ ttl:, id: }`. |
|
|
84
|
+
| `compaction:` | `with_compaction` | `true`, `false`, or `{ at:, instructions:, pause_after: }`. |
|
|
85
|
+
| `end_user:` | `with_end_user` | An opaque id. RubyLLM sends it as given, so never pass PII. |
|
|
86
|
+
| `provider_options:` | `with_provider_options` | Raw keys merged into the request payload, e.g. `{ service_tier: "flex" }`. |
|
|
87
|
+
| `headers:` | `with_headers` | Extra HTTP headers, e.g. a provider beta flag. |
|
|
88
|
+
|
|
89
|
+
`thinking:`, `citations:`, `caching:`, and `compaction:` also forward an explicit `false`, which you'd use, for example, to turn off thinking on a model that enables it by default. `context:`, `headers:`, and `provider_tools:` are marked `sensitive`, so axn keeps them out of logs.
|
|
90
|
+
|
|
91
|
+
### Attachments, history, and streaming
|
|
92
|
+
|
|
93
|
+
```ruby
|
|
94
|
+
Axn::RubyLLM.ask(
|
|
95
|
+
prompt: "What changed since last quarter?",
|
|
96
|
+
attachments: ["q3.pdf", "https://example.com/q2.pdf"], # Chat#ask's with:
|
|
97
|
+
history: prior_turns, # seeded before the prompt
|
|
98
|
+
on_chunk: ->(chunk) { print chunk.content }, # streamed as it arrives
|
|
99
|
+
)
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
- **`attachments:`** is passed to `Chat#ask` as `with:`. RubyLLM reads a String from disk or fetches it over HTTP, so never pass an unvalidated, user-supplied path or URL.
|
|
103
|
+
- **`history:`** is anything `Chat#messages=` accepts: `RubyLLM::Message`s, `{ role:, content: }` Hashes, or records that respond to `#to_llm`. `system_prompt:` replaces any system message in it. `transcript` includes only the turns this call exchanged, not the seeded ones.
|
|
104
|
+
- **`on_chunk:`** receives each `RubyLLM::Chunk`. `response` and `raw_message` are still the complete message, so `schema:` works as usual. The stubbed (disabled) path never calls it.
|
|
105
|
+
|
|
66
106
|
|
|
67
107
|
### Structured output via schema
|
|
68
108
|
|
|
@@ -125,11 +165,13 @@ Every successful result exposes token usage and cost, read off RubyLLM's own usa
|
|
|
125
165
|
```ruby
|
|
126
166
|
result = Axn::RubyLLM.ask(prompt: "...")
|
|
127
167
|
|
|
128
|
-
result.
|
|
129
|
-
result.
|
|
130
|
-
result.
|
|
131
|
-
result.
|
|
132
|
-
result.output_tokens
|
|
168
|
+
result.total_input_tokens # => 512 (every input token: uncached + cache reads + cache writes — use this for prompt size)
|
|
169
|
+
result.uncached_input_tokens # => 412 (standard-rate input only — RubyLLM's Tokens#input)
|
|
170
|
+
result.cache_read_tokens # => 80 (served from the provider's prompt cache)
|
|
171
|
+
result.cache_write_tokens # => 20 (written to the provider's prompt cache)
|
|
172
|
+
result.output_tokens # => 78
|
|
173
|
+
result.thinking_tokens # => 40 (reasoning tokens, where the provider reports them separately)
|
|
174
|
+
result.server_tool_use # => { "web_search_requests" => 2 } (per-use counters for provider-hosted tools)
|
|
133
175
|
result.cost # => 0.00056 (Float USD total; nil if RubyLLM has no pricing for the model)
|
|
134
176
|
|
|
135
177
|
# Full breakdown — RubyLLM::Cost, RubyLLM's own aggregated-cost object
|
|
@@ -139,11 +181,17 @@ result.cost_breakdown # => #<Cost input: 0.0004, output: 0.00016, cache_read: 0
|
|
|
139
181
|
result.raw_message # => #<RubyLLM::Message ...>
|
|
140
182
|
```
|
|
141
183
|
|
|
142
|
-
`
|
|
184
|
+
Use `total_input_tokens` for prompt size. With prompt caching (explicit via `caching:`, or automatic on OpenAI's newer models), a tool loop's repeated prefix is billed as cache reads and new content as cache writes, so `uncached_input_tokens` can be single digits on a 50k-token prompt. `cost` prices each bucket at its own rate, so it's right either way.
|
|
185
|
+
|
|
186
|
+
There's deliberately no plain `input_tokens` here: RubyLLM's `Tokens#input` means *uncached* input (a billing bucket), while OpenTelemetry's `gen_ai.usage.input_tokens` means *all* input, and each field name says which one it is. The older `input_tokens` (= `uncached_input_tokens`) and `prompt_tokens` (= `total_input_tokens`) still work but are deprecated; see [DEPRECATIONS.md](DEPRECATIONS.md).
|
|
187
|
+
|
|
188
|
+
`cost` is `nil` when RubyLLM lacks pricing for the model (e.g. unknown/custom endpoints); `cost_breakdown` itself is still a `Cost` object in that case (only its component readers are `nil`). Token counts are nil only if the provider did not return them. `total_input_tokens` is nil only if all three input token fields are nil.
|
|
143
189
|
|
|
144
190
|
### Errors
|
|
145
191
|
|
|
146
192
|
Errors are handled via Axn's declarative `error` DSL. Every failure shares a consistent `"LLM request failed: <reason>"` headline (the headline itself is configurable via `c.error_headline =`, e.g. to `"Something went wrong calling the LLM"`; the reasons below are unaffected):
|
|
193
|
+
- The response hit the output token limit (`finish_reason: :max_tokens`) → `"LLM request failed: Response was cut off by the output token limit before it finished"`. This is checked before `schema:` parsing, so a truncated structured response reports this rather than a JSON error.
|
|
194
|
+
- A provider content filter blocked the response (`finish_reason: :content_filter`) → `"LLM request failed: Response was blocked by the provider's content filter"`
|
|
147
195
|
- `JSON::ParserError` → `"LLM request failed: Response was not valid JSON"`
|
|
148
196
|
- `RubyLLM::RateLimitError` (HTTP 429, provider-agnostic) → `"LLM request failed: Rate limit reached: <message>"`
|
|
149
197
|
- `RubyLLM::OverloadedError` / `ServiceUnavailableError` / `ServerError` (5xx, transient) → `"LLM request failed: Provider temporarily unavailable, try again later: <message>"`
|
|
@@ -152,6 +200,8 @@ Errors are handled via Axn's declarative `error` DSL. Every failure shares a con
|
|
|
152
200
|
- Any other known RubyLLM error — `RubyLLM::Error` (auth, bad request, payment, etc.), `RubyLLM::ConfigurationError`, `ModelNotFoundError`, `ModelRegistryError`, `PromptNotFoundError`, `InvalidRoleError`, `InvalidToolChoiceError`, `PendingToolCallsError`, `CancelledError`, `UnsupportedAttachmentError` — or `Faraday::Error` (network/transport failure) → `"LLM request failed: <message>"`
|
|
153
201
|
- Any other `StandardError` (i.e. not a recognized RubyLLM/network failure — most likely a bug) → `"LLM request failed"`, with no exception detail leaked into the message
|
|
154
202
|
|
|
203
|
+
The truncation and content-filter failures still expose the usage fields, `cost`, `raw_message`, `transcript` and `finish_reason`, since the call was paid for. `result.finish_reason` is also set on success (normally `:stop`).
|
|
204
|
+
|
|
155
205
|
## Tool adapter — wrap any Axn as a RubyLLM::Tool
|
|
156
206
|
|
|
157
207
|
Any [Axn](https://github.com/teamshares/axn) can be exposed as a `::RubyLLM::Tool` — no adapter-specific mixin required, it's just a normal Axn:
|
|
@@ -234,7 +284,7 @@ Passing `ambient_context:` returns a tool **instance** (closing over that contex
|
|
|
234
284
|
|
|
235
285
|
### Using wrapped tools with RubyLLM directly
|
|
236
286
|
|
|
237
|
-
`Axn::RubyLLM.ask
|
|
287
|
+
`Axn::RubyLLM.ask` runs one prompt through to a final answer. Drive `RubyLLM.chat` yourself when you need more control than that: keeping one chat alive across turns, stepping the loop by hand (`ask_later`/`step`), or RubyLLM's lifecycle callbacks (`before_tool_call`, `after_tool_result`, `after_message`, `before_request`, `before_fallback`). Register wrapped tools with RubyLLM's own `with_tools`, which accepts one or many `RubyLLM::Tool` classes/instances:
|
|
238
288
|
|
|
239
289
|
```ruby
|
|
240
290
|
chat = RubyLLM.chat.with_tools(Axn::RubyLLM.wrap(CreateWidget))
|
|
@@ -244,6 +294,90 @@ chat.ask("Create a widget called Sprocket")
|
|
|
244
294
|
chat = RubyLLM.chat.with_tools(*Axn::RubyLLM.tools)
|
|
245
295
|
```
|
|
246
296
|
|
|
297
|
+
### Provider-hosted tools and remote MCP (`provider_tools:`)
|
|
298
|
+
|
|
299
|
+
`provider_tools:` forwards verbatim to RubyLLM's `Chat#with_provider_tools` — a Hash of alias => options. This is how the *provider* (not your app) runs a tool, including a **remote MCP server**: the provider connects to the server directly, with whatever `headers:` you pass, and its own tool calls happen inside the provider's request rather than yours.
|
|
300
|
+
|
|
301
|
+
```ruby
|
|
302
|
+
Axn::RubyLLM.ask(
|
|
303
|
+
prompt: "Why did this company's margin drop?",
|
|
304
|
+
provider_tools: {
|
|
305
|
+
mcp: {
|
|
306
|
+
name: "metabase", url: ENV.fetch("METABASE_MCP_URL"),
|
|
307
|
+
headers: { "X-API-KEY" => ENV.fetch("METABASE_MCP_API_KEY") },
|
|
308
|
+
allowed_tools: %w[search execute_sql],
|
|
309
|
+
require_approval: "never",
|
|
310
|
+
},
|
|
311
|
+
},
|
|
312
|
+
)
|
|
313
|
+
```
|
|
314
|
+
|
|
315
|
+
See RubyLLM's `Chat#with_provider_tools` for the other aliases (`:web_search`, `:code_execution`, ...) and which providers support each one.
|
|
316
|
+
|
|
317
|
+
### App-side remote MCP (`Axn::RubyLLM.remote_mcp_tools`)
|
|
318
|
+
|
|
319
|
+
`remote_mcp_tools` works the other way round: *your app* connects to the MCP server, using the official `mcp` gem, and wraps each of the server's tools as an ordinary `RubyLLM::Tool`. They sit next to your Axn tools in `tools:`, and every call is one your app makes, logs, times out, and can cap.
|
|
320
|
+
|
|
321
|
+
```ruby
|
|
322
|
+
toolset = Axn::RubyLLM.remote_mcp_tools(
|
|
323
|
+
url: ENV.fetch("METABASE_MCP_URL"),
|
|
324
|
+
headers: { "X-API-KEY" => ENV.fetch("METABASE_MCP_API_KEY") },
|
|
325
|
+
allowed_tools: %w[search execute_sql], # pass it: a server's full list can include write/admin tools
|
|
326
|
+
max_calls: 20, # shared across the whole toolset, for its lifetime (default 20)
|
|
327
|
+
)
|
|
328
|
+
Axn::RubyLLM.ask(prompt: "...", tools: [*Axn::RubyLLM.tools, *toolset.tools])
|
|
329
|
+
toolset.close
|
|
330
|
+
```
|
|
331
|
+
|
|
332
|
+
The `max_calls` budget never resets, so connect a fresh toolset per request (and `close` it) when you want a per-request cap; a toolset reused across requests shares one budget between them. `timeout:` (default 60s) and `max_result_chars:` (default 20,000) bound each call. When a server wants more than a static header for auth:
|
|
333
|
+
|
|
334
|
+
- **`bearer_token:`** takes a String, or a callable that returns one, and sends it as `Authorization: Bearer <token>`. A callable runs on **every request**, so your code owns fetching, caching, and rotating the token, e.g. `bearer_token: -> { MyOAuthStore.current_token }`.
|
|
335
|
+
- **`oauth:`** takes an `MCP::Client::OAuth::ClientCredentialsProvider` (machine-to-machine) or `MCP::Client::OAuth::Provider` (interactive authorization code + PKCE). It's passed straight to `MCP::Client::HTTP`, which handles discovery, token exchange, refresh, and retrying after a 401.
|
|
336
|
+
|
|
337
|
+
Pass one or the other, not both. Both refuse a URL that is neither https nor loopback http.
|
|
338
|
+
|
|
339
|
+
### Pausing on approval before a provider-hosted call runs (`on_remote_tool_approval:`)
|
|
340
|
+
|
|
341
|
+
Set `require_approval: "always"` (or per-tool) on a `provider_tools:` entry and the provider pauses before running the call, asking your app to approve or deny it first. Without a decision, `ask` would return whatever the chat produced at that pause — not a final answer. Pass `on_remote_tool_approval:`, a callable given each pending `RubyLLM::ToolCall` (`.name`, `.arguments`, `.remote?`), and `ask` drives the chat past every approval until it's genuinely done:
|
|
342
|
+
|
|
343
|
+
```ruby
|
|
344
|
+
Axn::RubyLLM.ask(
|
|
345
|
+
prompt: "...",
|
|
346
|
+
provider_tools: { mcp: { name: "metabase", url: ..., require_approval: "always" } },
|
|
347
|
+
on_remote_tool_approval: ->(tool_call) {
|
|
348
|
+
Rails.logger.info("[probe] #{tool_call.name}: #{tool_call.arguments}")
|
|
349
|
+
tool_call.name == "execute_sql" # only allow the tools you actually want to approve
|
|
350
|
+
},
|
|
351
|
+
)
|
|
352
|
+
```
|
|
353
|
+
|
|
354
|
+
The same callback also resolves a **local** tool declared with `Tool.requires_approval` — `pending_approvals` doesn't distinguish where the call runs, only whether a decision is still owed.
|
|
355
|
+
|
|
356
|
+
### Tool concurrency and other chat options (`tool_options:`)
|
|
357
|
+
|
|
358
|
+
`tool_options:` forwards verbatim to `Chat#with_tool_options` (`choice:`, `calls:`, `concurrency:`). `concurrency: :threads` or `:fibers` runs one turn's **local** Axn tool calls in parallel — useful when a tool is pure I/O (an HTTP call, a remote MCP round-trip wrapped as a local tool) and each call doesn't share mutable state. It is not a good fit for a tool that checks out an ActiveRecord connection: parallel local calls check out one connection each, from a pool sized for the process's normal concurrency.
|
|
359
|
+
|
|
360
|
+
```ruby
|
|
361
|
+
Axn::RubyLLM.ask(prompt: "...", tools: [...], tool_options: { concurrency: :threads, calls: :many })
|
|
362
|
+
```
|
|
363
|
+
|
|
364
|
+
### Capping tool calls (`max_tool_calls:`)
|
|
365
|
+
|
|
366
|
+
`ask` has no limit of its own on how long the tool loop runs, because RubyLLM 2.0 removed `halt_after`. `max_tool_calls:` caps how many tool calls your app runs in one `ask`, counted across every tool in `tools:`. Once the cap is reached, each further call returns `{ error: "Tool call budget exhausted ... write your final answer with what you have so far." }` and the model is left to finish. The cap is soft on purpose: raising would throw away everything the chat had gathered. Calls to `remote_mcp_tools` tools count here as well as against their own `max_calls`. Provider-hosted calls (`provider_tools:`) never reach your app, so they aren't counted. The caller's tool instances are copied, not modified, so reusing a `wrap(axn, ambient_context:)` instance across calls is safe.
|
|
367
|
+
|
|
368
|
+
```ruby
|
|
369
|
+
Axn::RubyLLM.ask(prompt: "...", tools: Axn::RubyLLM.tools, max_tool_calls: 15)
|
|
370
|
+
```
|
|
371
|
+
|
|
372
|
+
### Reading what the chat actually did (`transcript`)
|
|
373
|
+
|
|
374
|
+
`ask` always exposes `transcript`: every message the chat exchanged (the system prompt excluded), each shaped as `{ role:, content:, tool_calls:, tool_call_id:, server_tool_calls: }`. `server_tool_calls` is where a provider-hosted MCP call's name, arguments, and result show up — your app never receives that request, so this is the only place to see what query actually ran.
|
|
375
|
+
|
|
376
|
+
```ruby
|
|
377
|
+
result = Axn::RubyLLM.ask(prompt: "...", provider_tools: { mcp: { ... } })
|
|
378
|
+
result.transcript.flat_map { |m| m[:server_tool_calls]&.values || [] }.each { |c| puts "#{c.name}: #{c.arguments}" }
|
|
379
|
+
```
|
|
380
|
+
|
|
247
381
|
### Tool naming
|
|
248
382
|
|
|
249
383
|
The name is axn core's canonical, provider-safe `tool_name`: lowercased to `[a-z0-9_]`, leading configured prefixes stripped, snake_cased with single underscores, and never blank (`Admin::CreateWidget` → `admin_create_widget`; a truly anonymous Axn → `"tool"`). Declare `axn_name "..."` on the Axn to override the default. Because it's the same core derivation every adapter uses, a class wrapped by both `Axn::RubyLLM.wrap` and `Axn::MCP.wrap` advertises an identical name — the contract is declared once.
|
|
@@ -341,9 +475,11 @@ The response is the only required argument — pass it positionally (as above) o
|
|
|
341
475
|
stub_axn_ruby_llm({ "company_id" => 42 }, schema: CompanyMatch)
|
|
342
476
|
stub_axn_ruby_llm("...", model: "gpt-4o", input_tokens: 100, output_tokens: 50, cost: 0.0023)
|
|
343
477
|
stub_axn_ruby_llm("...", cache_read_tokens: 500, cache_write_tokens: 200)
|
|
478
|
+
stub_axn_ruby_llm("...", thinking_tokens: 300, server_tool_use: { "web_search_requests" => 1 })
|
|
479
|
+
stub_axn_ruby_llm("partial", finish_reason: :max_tokens) # exercises Ask's truncation failure
|
|
344
480
|
```
|
|
345
481
|
|
|
346
|
-
Remaining keywords: `model:`, `schema:`, `input_tokens
|
|
482
|
+
Remaining keywords: `model:`, `schema:`, `input_tokens:` (RubyLLM's uncached count, so it becomes `uncached_input_tokens`), `output_tokens:`, `cache_read_tokens:`, `cache_write_tokens:`, `thinking_tokens:`, `server_tool_use:`, `cost:`, and `finish_reason:` (default `:stop`). It stubs every `RubyLLM::Chat#with_*` method, so specs can pass any `Ask` option. Returns the stubbed chat instance double for further assertions if you need it.
|
|
347
483
|
|
|
348
484
|
## OpenTelemetry
|
|
349
485
|
|
|
@@ -353,10 +489,13 @@ If your app uses OpenTelemetry, `axn` already wraps every action in an `axn.call
|
|
|
353
489
|
|---|---|
|
|
354
490
|
| `gen_ai.request.model` | The model requested |
|
|
355
491
|
| `gen_ai.response.model` | The model that responded |
|
|
356
|
-
| `gen_ai.usage.input_tokens` |
|
|
492
|
+
| `gen_ai.usage.input_tokens` | `total_input_tokens` — all input, cached included, per the OTel GenAI conventions |
|
|
493
|
+
| `gen_ai.usage.cache_read.input_tokens` | `cache_read_tokens` (omitted when the provider doesn't report it) |
|
|
494
|
+
| `gen_ai.usage.cache_creation.input_tokens` | `cache_write_tokens` (omitted when the provider doesn't report it) |
|
|
357
495
|
| `gen_ai.usage.output_tokens` | Completion token count |
|
|
358
496
|
| `gen_ai.usage.cost` | USD total (non-standard; useful for spend filtering) |
|
|
359
497
|
| `axn.ruby_llm.stubbed` | `true` when production gating returned a stub |
|
|
498
|
+
| `axn.ruby_llm.version` | This gem's version — separates spans from before and after a change in an attribute's meaning |
|
|
360
499
|
| `axn.dimension.invoked_via` | `"ruby_llm"` — set by axn core on every tool call (including nested sub-Axns), not by this gem; lets you separate tool-driven traffic from ordinary direct `.call`s in the same span schema |
|
|
361
500
|
|
|
362
501
|
For LLM-level tracing (individual `RubyLLM.chat` calls, tool calls, embeddings, prompt content), add [`opentelemetry-instrumentation-ruby_llm`](https://github.com/thoughtbot/opentelemetry-instrumentation-ruby_llm) to your own Gemfile and configure it per its README. It is not a dependency of this gem.
|
|
@@ -381,7 +520,9 @@ When disabled, `Axn::RubyLLM.ask` returns a **success** result with obvious stub
|
|
|
381
520
|
|---|---|
|
|
382
521
|
| `response` | `"stubbed response value"` (plain) / `{ "stubbed" => true }` (`schema:`) |
|
|
383
522
|
| `raw_message` | Stub struct with `.content`, `.tokens` (a real `RubyLLM::Tokens`, all zero), `.model`, `.parsed` |
|
|
384
|
-
| `
|
|
523
|
+
| `total_input_tokens` / `uncached_input_tokens` / `cache_read_tokens` / `cache_write_tokens` / `output_tokens` | `0` (the deprecated `input_tokens` / `prompt_tokens` too) |
|
|
524
|
+
| `thinking_tokens` / `server_tool_use` / `finish_reason` | `nil` |
|
|
525
|
+
| `transcript` | `[]` |
|
|
385
526
|
| `cost` | `0.0` |
|
|
386
527
|
| `cost_breakdown` | `nil` |
|
|
387
528
|
| `stubbed` | `true` |
|
data/lib/axn/ruby_llm/ask.rb
CHANGED
|
@@ -6,20 +6,107 @@ module Axn
|
|
|
6
6
|
include Axn
|
|
7
7
|
|
|
8
8
|
expects :prompt
|
|
9
|
+
# Files/URLs attached to the prompt -- Chat#ask's `with:` (a path, URL, IO, or an Array of them).
|
|
10
|
+
# A String is read from disk or fetched over HTTP, so never pass an unvalidated user-supplied one.
|
|
11
|
+
expects :attachments, optional: true
|
|
12
|
+
# Prior turns to seed the chat with before `prompt` -- anything Chat#messages= accepts
|
|
13
|
+
# (RubyLLM::Message objects, `{ role:, content: }` Hashes, or records responding to #to_llm).
|
|
14
|
+
# Not echoed back in `transcript`, which covers only what this call exchanged.
|
|
15
|
+
expects :history, optional: true
|
|
9
16
|
expects :schema, optional: true
|
|
10
17
|
expects :model, optional: true
|
|
18
|
+
# Model resolution alongside `model:` (Chat.new's own keywords): `provider:` disambiguates a
|
|
19
|
+
# model several providers serve, `protocol:` overrides the wire protocol, and
|
|
20
|
+
# `assume_model_exists: true` skips the registry lookup (requires `provider:`).
|
|
21
|
+
expects :provider, optional: true
|
|
22
|
+
expects :protocol, optional: true
|
|
23
|
+
expects :assume_model_exists, optional: true
|
|
24
|
+
# A RubyLLM::Context (RubyLLM.context { |c| ... }) whose config -- API keys, base URLs --
|
|
25
|
+
# replaces the process-wide RubyLLM.config for this call.
|
|
26
|
+
expects :context, optional: true, sensitive: true
|
|
27
|
+
# Ordered models to retry on when generation fails (Chat#with_fallbacks); `fallback_on:` narrows
|
|
28
|
+
# the triggering error classes (default: RubyLLM's transient provider/network errors).
|
|
29
|
+
expects :fallbacks, optional: true
|
|
30
|
+
expects :fallback_on, optional: true
|
|
11
31
|
expects :system_prompt, optional: true
|
|
32
|
+
# true marks the system prompt as an explicit prompt-cache boundary
|
|
33
|
+
# (with_instructions(cache_until_here: true)) -- worth it for a long, reused system prompt.
|
|
34
|
+
expects :cache_system_prompt, optional: true
|
|
12
35
|
expects :temperature, optional: true
|
|
36
|
+
expects :max_output_tokens, optional: true
|
|
37
|
+
# true, false (disable for a model that thinks by default), or `{ effort:, budget:, display: }`.
|
|
38
|
+
expects :thinking, optional: true
|
|
39
|
+
# true to get `raw_message.citations` back from attached documents.
|
|
40
|
+
expects :citations, optional: true
|
|
41
|
+
# true, false, or `{ ttl:, id: }` -- provider prompt caching (Chat#with_caching).
|
|
42
|
+
expects :caching, optional: true
|
|
43
|
+
# true, false, or `{ at:, instructions:, pause_after: }` -- provider-side context compaction.
|
|
44
|
+
expects :compaction, optional: true
|
|
45
|
+
# An opaque end-user id for the provider's abuse monitoring -- sent as given, so never PII.
|
|
46
|
+
expects :end_user, optional: true
|
|
47
|
+
# Raw keys merged into the provider request payload (Chat#with_provider_options), e.g.
|
|
48
|
+
# `{ service_tier: "flex" }` -- distinct from a tool's own `provider_options`.
|
|
49
|
+
expects :provider_options, optional: true
|
|
50
|
+
# Extra HTTP headers on the completion request (e.g. a provider beta flag).
|
|
51
|
+
expects :headers, optional: true, sensitive: true
|
|
52
|
+
# A callable given each streamed RubyLLM::Chunk as it arrives. `response`/`raw_message` are still
|
|
53
|
+
# the complete message (RubyLLM assembles it), so `schema:` works unchanged. Never called on the
|
|
54
|
+
# disabled (stubbed) path.
|
|
55
|
+
expects :on_chunk, optional: true
|
|
13
56
|
expects :tools, optional: true
|
|
57
|
+
# Caps how many tool calls this app executes in one ask, across every tool in `tools:`
|
|
58
|
+
# (including remote_mcp_tools ones, which also keep their own max_calls budget). Past the cap
|
|
59
|
+
# each call returns an error result telling the model to answer with what it has -- see
|
|
60
|
+
# ToolBudget. Provider-hosted calls (provider_tools:) never reach this app and aren't counted.
|
|
61
|
+
expects :max_tool_calls, optional: true
|
|
62
|
+
# Forwarded verbatim to Chat#with_provider_tools -- a Hash of alias => options, e.g.
|
|
63
|
+
# `provider_tools: { mcp: { name:, url:, headers:, allowed_tools:, require_approval: } }` for a
|
|
64
|
+
# provider-hosted remote MCP server, or `{ web_search: {} }`. See RubyLLM::Chat#with_provider_tools.
|
|
65
|
+
expects :provider_tools, optional: true, sensitive: true
|
|
66
|
+
# Forwarded verbatim to Chat#with_tool_options (choice:/calls:/concurrency:).
|
|
67
|
+
expects :tool_options, optional: true
|
|
68
|
+
# A callable given each pending ToolCall (from Chat#pending_approvals -- local tools declared
|
|
69
|
+
# with `requires_approval`, or provider-hosted remote calls awaiting an mcp_approval_request) and
|
|
70
|
+
# returning truthy to approve, falsy to deny. Required to drive a chat past #awaiting_approval? --
|
|
71
|
+
# without it, `llm_response` is whatever #ask returned when the loop first parked, and the
|
|
72
|
+
# response schema/tool loop never completes. Only meaningful alongside provider_tools using
|
|
73
|
+
# require_approval, or a local tool declared with `requires_approval`.
|
|
74
|
+
expects :on_remote_tool_approval, optional: true
|
|
14
75
|
|
|
15
76
|
exposes :response
|
|
16
77
|
exposes :raw_message
|
|
17
|
-
|
|
18
|
-
|
|
78
|
+
# A plain-Hash summary of every message the chat exchanged (system prompt excluded), in order:
|
|
79
|
+
# role, content, tool_calls (name/arguments/remote?), tool_call_id, and server_tool_calls (the
|
|
80
|
+
# provider-executed calls a hosted MCP server ran, e.g. Metabase queries). Built once per call
|
|
81
|
+
# from Chat#messages, which RubyLLM already retains -- this only shapes it for a caller that
|
|
82
|
+
# doesn't want to reach into raw RubyLLM::Message objects.
|
|
83
|
+
exposes :transcript, type: Array, allow_blank: true
|
|
84
|
+
# Input token fields deliberately avoid a bare `input_tokens` name: RubyLLM's Tokens#input is a
|
|
85
|
+
# billing bucket (standard-rate, non-cached input only), while OTel's gen_ai.usage.input_tokens
|
|
86
|
+
# is the whole prompt including cached tokens. Each name here says which one it is.
|
|
87
|
+
#
|
|
88
|
+
# All input tokens: uncached + cache_read + cache_write (the OTel meaning).
|
|
89
|
+
exposes :total_input_tokens, allow_nil: true
|
|
90
|
+
# Standard-rate input only -- RubyLLM's Tokens#input.
|
|
91
|
+
exposes :uncached_input_tokens, allow_nil: true
|
|
19
92
|
exposes :cache_read_tokens, allow_nil: true
|
|
20
93
|
exposes :cache_write_tokens, allow_nil: true
|
|
94
|
+
exposes :output_tokens, allow_nil: true
|
|
95
|
+
# Reasoning tokens, where the provider reports them separately (usually also counted in
|
|
96
|
+
# output_tokens when billed as output).
|
|
97
|
+
exposes :thinking_tokens, allow_nil: true
|
|
98
|
+
# Per-use counters for provider-hosted tools, e.g. { "web_search_requests" => 2 } -- billed per
|
|
99
|
+
# use, not per token.
|
|
100
|
+
exposes :server_tool_use, allow_nil: true
|
|
101
|
+
# Deprecated (see DEPRECATIONS.md): input_tokens == uncached_input_tokens,
|
|
102
|
+
# prompt_tokens == total_input_tokens.
|
|
103
|
+
exposes :input_tokens, allow_nil: true
|
|
21
104
|
exposes :prompt_tokens, allow_nil: true
|
|
22
105
|
exposes :cost, allow_nil: true
|
|
106
|
+
# The final message's normalized finish reason (:stop, :tool_calls, :max_tokens,
|
|
107
|
+
# :content_filter, ...). :max_tokens and :content_filter fail the call -- see
|
|
108
|
+
# fail_on_incomplete_response! -- with this still exposed.
|
|
109
|
+
exposes :finish_reason, allow_nil: true
|
|
23
110
|
exposes :cost_breakdown, allow_nil: true
|
|
24
111
|
exposes :stubbed, type: :boolean, default: false
|
|
25
112
|
|
|
@@ -76,38 +163,22 @@ module Axn
|
|
|
76
163
|
before do
|
|
77
164
|
if disabled?
|
|
78
165
|
exposures = stubbed_exposures
|
|
79
|
-
record_otel_attributes!(
|
|
80
|
-
input_tokens: exposures[:input_tokens],
|
|
81
|
-
output_tokens: exposures[:output_tokens],
|
|
82
|
-
cost: exposures[:cost],
|
|
83
|
-
response_model: nil,
|
|
84
|
-
stubbed: true,
|
|
85
|
-
)
|
|
166
|
+
record_otel_attributes!(exposures, response_model: nil)
|
|
86
167
|
# Reason attaches to the "LLM request completed" base via the parenthetical join: above.
|
|
87
168
|
done!("using stubbed values - actual LLM request disabled", **exposures)
|
|
88
169
|
end
|
|
89
170
|
end
|
|
90
171
|
|
|
91
172
|
def call
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
cost: cost_breakdown&.total,
|
|
102
|
-
stubbed: false,
|
|
103
|
-
)
|
|
104
|
-
record_otel_attributes!(
|
|
105
|
-
input_tokens: token_usage.input,
|
|
106
|
-
output_tokens: token_usage.output,
|
|
107
|
-
cost: cost_breakdown&.total,
|
|
108
|
-
response_model: llm_response&.model,
|
|
109
|
-
stubbed: false,
|
|
110
|
-
)
|
|
173
|
+
# llm_response runs the chat; usage must be read after it -- token_usage/cost_breakdown
|
|
174
|
+
# memoize Chat#tokens/#cost, which are an empty ledger until #ask has run.
|
|
175
|
+
message = llm_response
|
|
176
|
+
usage = usage_exposures
|
|
177
|
+
record_otel_attributes!(usage, response_model: message&.model)
|
|
178
|
+
fail_on_incomplete_response!(message, usage)
|
|
179
|
+
|
|
180
|
+
expose(response: parsed_response, raw_message: message, transcript: transcript_entries,
|
|
181
|
+
finish_reason: message&.finish_reason, **usage)
|
|
111
182
|
rescue ::RubyLLM::RateLimitError => e
|
|
112
183
|
fail! "Rate limit reached: #{e.message}"
|
|
113
184
|
end
|
|
@@ -123,11 +194,17 @@ module Axn
|
|
|
123
194
|
{
|
|
124
195
|
response: parsed_content || content,
|
|
125
196
|
raw_message: StubMessage.new(content:, tokens: zero_tokens, model: "stubbed"),
|
|
126
|
-
|
|
127
|
-
|
|
197
|
+
transcript: [],
|
|
198
|
+
total_input_tokens: 0,
|
|
199
|
+
uncached_input_tokens: 0,
|
|
128
200
|
cache_read_tokens: 0,
|
|
129
201
|
cache_write_tokens: 0,
|
|
202
|
+
output_tokens: 0,
|
|
203
|
+
input_tokens: 0,
|
|
130
204
|
prompt_tokens: 0,
|
|
205
|
+
thinking_tokens: nil,
|
|
206
|
+
server_tool_use: nil,
|
|
207
|
+
finish_reason: nil,
|
|
131
208
|
cost: 0.0,
|
|
132
209
|
cost_breakdown: nil,
|
|
133
210
|
stubbed: true,
|
|
@@ -153,6 +230,40 @@ module Axn
|
|
|
153
230
|
memo def token_usage = chat.tokens
|
|
154
231
|
memo def cost_breakdown = chat.cost
|
|
155
232
|
|
|
233
|
+
# A truncated or filtered response is not an answer: returning it as a success hands callers
|
|
234
|
+
# partial text (or, with schema:, a misleading JSON parse error). Checked before
|
|
235
|
+
# parsed_response for that reason. Usage, cost and the raw message are still exposed, since
|
|
236
|
+
# the call was paid for.
|
|
237
|
+
def fail_on_incomplete_response!(message, usage)
|
|
238
|
+
reason =
|
|
239
|
+
if message&.max_tokens?
|
|
240
|
+
"Response was cut off by the output token limit before it finished"
|
|
241
|
+
elsif message&.content_filtered?
|
|
242
|
+
"Response was blocked by the provider's content filter"
|
|
243
|
+
end
|
|
244
|
+
return unless reason
|
|
245
|
+
|
|
246
|
+
fail!(reason, **usage, raw_message: message, transcript: transcript_entries, finish_reason: message.finish_reason)
|
|
247
|
+
end
|
|
248
|
+
|
|
249
|
+
def usage_exposures
|
|
250
|
+
total = total_input_tokens
|
|
251
|
+
{
|
|
252
|
+
total_input_tokens: total,
|
|
253
|
+
uncached_input_tokens: token_usage.input,
|
|
254
|
+
cache_read_tokens: token_usage.cache_read,
|
|
255
|
+
cache_write_tokens: token_usage.cache_write,
|
|
256
|
+
output_tokens: token_usage.output,
|
|
257
|
+
thinking_tokens: token_usage.thinking,
|
|
258
|
+
server_tool_use: token_usage.server_tool_use,
|
|
259
|
+
input_tokens: token_usage.input,
|
|
260
|
+
prompt_tokens: total,
|
|
261
|
+
cost_breakdown:,
|
|
262
|
+
cost: cost_breakdown&.total,
|
|
263
|
+
stubbed: false,
|
|
264
|
+
}
|
|
265
|
+
end
|
|
266
|
+
|
|
156
267
|
# nil only when NO turn reported the field (preserving the "nil if the provider didn't return
|
|
157
268
|
# it" contract); otherwise the summed count, treating a missing component as 0.
|
|
158
269
|
def total_input_tokens
|
|
@@ -160,14 +271,72 @@ module Axn
|
|
|
160
271
|
vals.all?(&:nil?) ? nil : vals.sum(&:to_i)
|
|
161
272
|
end
|
|
162
273
|
|
|
163
|
-
|
|
274
|
+
# Without on_remote_tool_approval, this is exactly #ask -- one #complete run, parked (like
|
|
275
|
+
# before this feature existed) if the chat lands on an unresolved approval. With it, drive
|
|
276
|
+
# #complete past every approval Chat#pending_approvals reports, until the loop truly finishes.
|
|
277
|
+
# Chat#complete already loops through ordinary tool calls on its own; this only extends that
|
|
278
|
+
# loop across the approval pauses it otherwise stops at.
|
|
279
|
+
memo def llm_response
|
|
280
|
+
message = chat.ask(prompt, with: attachments, &on_chunk)
|
|
281
|
+
return message unless on_remote_tool_approval
|
|
164
282
|
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
283
|
+
while chat.awaiting_approval?
|
|
284
|
+
chat.pending_approvals.each do |tool_call|
|
|
285
|
+
on_remote_tool_approval.call(tool_call) ? chat.approve(tool_call) : chat.deny(tool_call)
|
|
286
|
+
end
|
|
287
|
+
message = chat.complete(&on_chunk)
|
|
288
|
+
end
|
|
289
|
+
message
|
|
290
|
+
end
|
|
291
|
+
|
|
292
|
+
# provider:/protocol:/assume_model_exists:/context: go to Chat.new rather than a later
|
|
293
|
+
# #with_model/#with_context, so the model resolves once, against the right config.
|
|
294
|
+
#
|
|
295
|
+
# thinking:/citations:/caching:/compaction: forward `false` too (`unless nil?`, not `if`) --
|
|
296
|
+
# false is a real instruction (e.g. turn off thinking a model enables by default).
|
|
297
|
+
memo def chat # rubocop:disable Metrics/AbcSize
|
|
298
|
+
::RubyLLM.chat(model: resolved_model, **{ provider:, protocol:, assume_model_exists:, context: }.compact).tap do |c|
|
|
299
|
+
seed_history(c) if history
|
|
300
|
+
c.with_instructions(system_prompt, **{ cache_until_here: cache_system_prompt }.compact) if system_prompt
|
|
168
301
|
c.with_schema(resolved_schema) if schema
|
|
169
302
|
c.with_temperature(temperature) if temperature
|
|
303
|
+
c.with_max_output_tokens(max_output_tokens) if max_output_tokens
|
|
304
|
+
c.with_thinking(thinking) unless thinking.nil?
|
|
305
|
+
c.with_citations(citations) unless citations.nil?
|
|
306
|
+
c.with_caching(caching) unless caching.nil?
|
|
307
|
+
c.with_compaction(compaction) unless compaction.nil?
|
|
308
|
+
c.with_fallbacks(*fallbacks, **{ on: fallback_on }.compact) if fallbacks
|
|
309
|
+
c.with_end_user(end_user) if end_user
|
|
310
|
+
c.with_provider_options(provider_options) if provider_options
|
|
311
|
+
c.with_headers(headers) if headers
|
|
170
312
|
c.with_tools(*resolved_tools) if resolved_tools.any?
|
|
313
|
+
c.with_provider_tools(**provider_tools) if provider_tools
|
|
314
|
+
c.with_tool_options(**tool_options) if tool_options
|
|
315
|
+
end
|
|
316
|
+
end
|
|
317
|
+
|
|
318
|
+
# Runs before with_instructions, since Chat#messages= replaces the whole conversation -- the
|
|
319
|
+
# system prompt included -- and with_instructions then replaces any system message the history
|
|
320
|
+
# carried. The seeded count lets `transcript` skip what the caller already had.
|
|
321
|
+
def seed_history(chat)
|
|
322
|
+
chat.messages = history
|
|
323
|
+
@seeded_message_count = chat.messages.count { |m| m.role != :system }
|
|
324
|
+
end
|
|
325
|
+
|
|
326
|
+
# chat.messages already retains the whole exchange (RubyLLM's own Chat#tokens / Chat#cost read
|
|
327
|
+
# off it the same way) -- this only reshapes each Message into a plain Hash a caller can log or
|
|
328
|
+
# assert against without reaching into RubyLLM::Message/ToolCall objects. The system prompt is
|
|
329
|
+
# excluded: it's an Ask input the caller already has (system_prompt), not something the loop
|
|
330
|
+
# produced.
|
|
331
|
+
def transcript_entries
|
|
332
|
+
chat.messages.reject { |m| m.role == :system }.drop(@seeded_message_count || 0).map do |m|
|
|
333
|
+
{
|
|
334
|
+
role: m.role,
|
|
335
|
+
content: m.content.is_a?(String) ? m.content : nil,
|
|
336
|
+
tool_calls: m.tool_calls&.transform_values { |tc| { name: tc.name, arguments: tc.arguments, remote: tc.remote? } },
|
|
337
|
+
tool_call_id: m.tool_call_id,
|
|
338
|
+
server_tool_calls: m.server_tool_calls,
|
|
339
|
+
}
|
|
171
340
|
end
|
|
172
341
|
end
|
|
173
342
|
|
|
@@ -251,8 +420,14 @@ module Axn
|
|
|
251
420
|
# straight in) and already-wrapped `::RubyLLM::Tool`s -- a class or an instance, the latter being
|
|
252
421
|
# how you pass a tool that closed over explicit context via `Axn::RubyLLM.wrap(axn, ambient_context:)`.
|
|
253
422
|
# RubyLLM's `with_tools` accepts either a class or an instance, so wrapped classes register as-is.
|
|
254
|
-
|
|
255
|
-
|
|
423
|
+
#
|
|
424
|
+
# Memoized: with max_tool_calls, every guarded tool must share the ONE budget built here.
|
|
425
|
+
memo def resolved_tools
|
|
426
|
+
wrapped = Array(tools).map { |tool| _as_ruby_llm_tool(tool) }
|
|
427
|
+
return wrapped unless max_tool_calls
|
|
428
|
+
|
|
429
|
+
budget = ToolBudget.new(max_tool_calls)
|
|
430
|
+
wrapped.map { |tool| budget.guard(tool) }
|
|
256
431
|
end
|
|
257
432
|
|
|
258
433
|
def _as_ruby_llm_tool(tool)
|
|
@@ -262,14 +437,21 @@ module Axn
|
|
|
262
437
|
Axn::RubyLLM.wrap(tool)
|
|
263
438
|
end
|
|
264
439
|
|
|
265
|
-
|
|
440
|
+
# OTel GenAI semconv: gen_ai.usage.input_tokens "SHOULD include all types of input tokens,
|
|
441
|
+
# including cached tokens", with the cache counts as sub-totals of it -- so it gets
|
|
442
|
+
# total_input_tokens, never RubyLLM's uncached Tokens#input. annotate_span skips nil values.
|
|
443
|
+
# axn.ruby_llm.version separates spans from before/after a change in what an attribute means.
|
|
444
|
+
def record_otel_attributes!(usage, response_model:)
|
|
266
445
|
Axn::Extensions::Tracing.annotate_span(
|
|
267
446
|
"gen_ai.request.model" => resolved_model,
|
|
268
447
|
"gen_ai.response.model" => response_model,
|
|
269
|
-
"gen_ai.usage.input_tokens" =>
|
|
270
|
-
"gen_ai.usage.
|
|
271
|
-
"gen_ai.usage.
|
|
272
|
-
"
|
|
448
|
+
"gen_ai.usage.input_tokens" => usage[:total_input_tokens],
|
|
449
|
+
"gen_ai.usage.cache_read.input_tokens" => usage[:cache_read_tokens],
|
|
450
|
+
"gen_ai.usage.cache_creation.input_tokens" => usage[:cache_write_tokens],
|
|
451
|
+
"gen_ai.usage.output_tokens" => usage[:output_tokens],
|
|
452
|
+
"gen_ai.usage.cost" => usage[:cost],
|
|
453
|
+
"axn.ruby_llm.stubbed" => usage[:stubbed],
|
|
454
|
+
"axn.ruby_llm.version" => Axn::RubyLLM::VERSION,
|
|
273
455
|
)
|
|
274
456
|
end
|
|
275
457
|
end
|
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "mcp"
|
|
4
|
+
|
|
5
|
+
module Axn
|
|
6
|
+
module RubyLLM
|
|
7
|
+
# Wraps a remote MCP server's tools as ordinary ::RubyLLM::Tool subclasses, so they can sit
|
|
8
|
+
# next to Axn-wrapped local tools in the same `tools:` array (Ask, or `chat.with_tools`).
|
|
9
|
+
#
|
|
10
|
+
# This is the APP-SIDE alternative to `Chat#with_provider_tools(mcp: {...})`: with that, the
|
|
11
|
+
# *provider* connects to the server directly and its tool calls happen inside the provider's own
|
|
12
|
+
# request, invisible to this app. Here, THIS process connects, so every call is one this app makes,
|
|
13
|
+
# logs, times out, and can cap -- at the cost of an extra round-trip per remote tool call (the
|
|
14
|
+
# provider calls back into this app's chat loop, rather than looping against the MCP server itself).
|
|
15
|
+
module RemoteMcp
|
|
16
|
+
DEFAULT_MAX_CALLS = 20
|
|
17
|
+
DEFAULT_TIMEOUT = 60
|
|
18
|
+
DEFAULT_MAX_RESULT_CHARS = 20_000
|
|
19
|
+
|
|
20
|
+
# Holds the connected client alongside the wrapped tools so the caller can `close` it when
|
|
21
|
+
# done (an MCP::Client::HTTP session is a real HTTP connection, possibly with a live SSE
|
|
22
|
+
# listener thread -- see MCP::Client::HTTP#close). `tools` is what actually gets passed to
|
|
23
|
+
# `chat.with_tools` / `Axn::RubyLLM.ask(tools:)`.
|
|
24
|
+
Toolset = Struct.new(:tools, :client, keyword_init: true) do
|
|
25
|
+
# MCP::Client itself has no #close -- only its transport does (a real HTTP connection,
|
|
26
|
+
# possibly with a live SSE listener thread).
|
|
27
|
+
def close
|
|
28
|
+
client.transport.close
|
|
29
|
+
end
|
|
30
|
+
end
|
|
31
|
+
|
|
32
|
+
class << self
|
|
33
|
+
# Connects to a remote MCP server over Streamable HTTP and returns a Toolset: one
|
|
34
|
+
# ::RubyLLM::Tool subclass per allowed remote tool, plus the underlying client for `close`.
|
|
35
|
+
#
|
|
36
|
+
# toolset = Axn::RubyLLM.remote_mcp_tools(
|
|
37
|
+
# url: ENV.fetch("METABASE_MCP_URL"),
|
|
38
|
+
# headers: { "X-API-KEY" => ENV.fetch("METABASE_MCP_API_KEY") },
|
|
39
|
+
# allowed_tools: %w[search read_resource execute_sql],
|
|
40
|
+
# )
|
|
41
|
+
# Axn::RubyLLM.ask(prompt: "...", tools: [*Axn::RubyLLM.tools, *toolset.tools])
|
|
42
|
+
# toolset.close
|
|
43
|
+
#
|
|
44
|
+
# allowed_tools is not a convenience -- pass it. A server's full tool list can include
|
|
45
|
+
# write/admin tools (e.g. Metabase's create_dashboard, update_question) that have no
|
|
46
|
+
# business being reachable from an LLM's own tool choices; nil (the default) keeps
|
|
47
|
+
# everything the server advertises, so an omitted allowlist is a deliberate "trust every
|
|
48
|
+
# tool this server exposes," not a safe default.
|
|
49
|
+
#
|
|
50
|
+
# Auth beyond a static `headers:` value:
|
|
51
|
+
#
|
|
52
|
+
# - `bearer_token:` -- a String, or a callable returning one, sent as `Authorization: Bearer
|
|
53
|
+
# <token>`. A callable runs on EVERY request (Faraday's :authorization middleware), so the
|
|
54
|
+
# caller owns caching/refresh/rotation (e.g. a token fetched from its own OAuth store).
|
|
55
|
+
# - `oauth:` -- an `MCP::Client::OAuth::ClientCredentialsProvider` (machine-to-machine) or
|
|
56
|
+
# `MCP::Client::OAuth::Provider` (interactive authorization code + PKCE), passed straight to
|
|
57
|
+
# `MCP::Client::HTTP`, which runs discovery, token exchange, refresh, and the 401 retry itself.
|
|
58
|
+
#
|
|
59
|
+
# Either one refuses a URL that is neither https nor loopback http, so a token never crosses
|
|
60
|
+
# the wire in plaintext.
|
|
61
|
+
def remote_mcp_tools(url:, headers: {}, bearer_token: nil, oauth: nil, allowed_tools: nil,
|
|
62
|
+
max_calls: DEFAULT_MAX_CALLS, timeout: DEFAULT_TIMEOUT, max_result_chars: DEFAULT_MAX_RESULT_CHARS)
|
|
63
|
+
raise ArgumentError, "pass bearer_token: or oauth:, not both" if bearer_token && oauth
|
|
64
|
+
|
|
65
|
+
client = build_client(url:, headers:, bearer_token:, oauth:, timeout:)
|
|
66
|
+
begin
|
|
67
|
+
client.connect
|
|
68
|
+
remote_tools = client.tools
|
|
69
|
+
remote_tools = remote_tools.select { |t| allowed_tools.include?(t.name) } if allowed_tools
|
|
70
|
+
|
|
71
|
+
# One budget per toolset, not per tool: the limit bounds total round-trips to this server, not
|
|
72
|
+
# calls to any single tool. It lives as long as the toolset and never resets, so connect a
|
|
73
|
+
# toolset per request (as the example above does) for a per-request cap.
|
|
74
|
+
budget = ToolBudget.new(max_calls, noun: "remote calls")
|
|
75
|
+
wrapped = remote_tools.map { |remote_tool| build_tool_class(remote_tool, client:, budget:, max_result_chars:) }
|
|
76
|
+
|
|
77
|
+
Toolset.new(tools: wrapped, client:)
|
|
78
|
+
rescue StandardError
|
|
79
|
+
# Until a Toolset is returned the caller has nothing to #close, so a failed handshake or
|
|
80
|
+
# tools/list would otherwise leak the HTTP session (and any SSE listener thread).
|
|
81
|
+
close_quietly(client.transport)
|
|
82
|
+
raise
|
|
83
|
+
end
|
|
84
|
+
end
|
|
85
|
+
|
|
86
|
+
private
|
|
87
|
+
|
|
88
|
+
def build_client(url:, headers:, bearer_token:, oauth:, timeout:)
|
|
89
|
+
# MCP::Client::HTTP enforces this itself for oauth:, but knows nothing about a token set by
|
|
90
|
+
# the Faraday middleware below.
|
|
91
|
+
raise ArgumentError, "bearer_token: requires an https (or loopback http) MCP URL" if bearer_token && !::MCP::Client::OAuth::Discovery.secure_url?(url)
|
|
92
|
+
|
|
93
|
+
transport = ::MCP::Client::HTTP.new(url:, headers:, **{ oauth: }.compact) do |faraday|
|
|
94
|
+
faraday.options.timeout = timeout
|
|
95
|
+
faraday.options.open_timeout = timeout
|
|
96
|
+
faraday.request :authorization, "Bearer", bearer_token if bearer_token
|
|
97
|
+
end
|
|
98
|
+
::MCP::Client.new(transport:)
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
# Cleanup on the failure path must never replace the error that got us there.
|
|
102
|
+
def close_quietly(transport)
|
|
103
|
+
transport.close
|
|
104
|
+
rescue StandardError => e
|
|
105
|
+
Axn::Extensions.best_effort("logging a failed remote MCP transport close") do
|
|
106
|
+
Axn.config.logger.warn { "[axn-ruby_llm] closing remote MCP transport after a failed setup also failed: #{e.class}: #{e.message}" }
|
|
107
|
+
end
|
|
108
|
+
end
|
|
109
|
+
|
|
110
|
+
def build_tool_class(remote_tool, client:, budget:, max_result_chars:)
|
|
111
|
+
tool_name = remote_tool.name
|
|
112
|
+
tool_description = remote_tool.description
|
|
113
|
+
input_schema = remote_tool.input_schema || { type: "object", properties: {} }
|
|
114
|
+
|
|
115
|
+
Class.new(::RubyLLM::Tool) do
|
|
116
|
+
description(tool_description) if tool_description
|
|
117
|
+
parameters(input_schema)
|
|
118
|
+
|
|
119
|
+
define_singleton_method(:tool_name) { tool_name }
|
|
120
|
+
|
|
121
|
+
define_method(:execute) do |**args|
|
|
122
|
+
Axn::RubyLLM::RemoteMcp.send(:call_remote_tool, client:, tool_name:, args:, budget:, max_result_chars:)
|
|
123
|
+
end
|
|
124
|
+
end
|
|
125
|
+
end
|
|
126
|
+
|
|
127
|
+
# Isolated from build_tool_class's closure (rather than inlined in the define_method block)
|
|
128
|
+
# so the rescue clauses -- the part most likely to need a new case as real servers are
|
|
129
|
+
# exercised -- are easy to find and extend on their own, not buried inside class-building
|
|
130
|
+
# metaprogramming.
|
|
131
|
+
def call_remote_tool(client:, tool_name:, args:, budget:, max_result_chars:)
|
|
132
|
+
return budget.exhausted_result unless budget.consume!
|
|
133
|
+
|
|
134
|
+
response = client.call_tool(name: tool_name, arguments: args)
|
|
135
|
+
render_tool_result(response, max_result_chars:)
|
|
136
|
+
rescue ::MCP::Client::ServerError => e
|
|
137
|
+
{ error: "Remote tool error: #{e.message}" }
|
|
138
|
+
rescue ::Faraday::TimeoutError
|
|
139
|
+
{ error: "Remote tool call timed out" }
|
|
140
|
+
rescue ::Faraday::Error, ::MCP::Client::RequestHandlerError => e
|
|
141
|
+
{ error: "Remote tool call failed: #{e.message}" }
|
|
142
|
+
rescue StandardError => e
|
|
143
|
+
# RubyLLM has no rescue around a tool's #execute (axn-ruby_llm's own ToolAdapter guard
|
|
144
|
+
# exists for exactly this reason on the Axn side) -- an unanticipated MCP client error
|
|
145
|
+
# here (a malformed response, a client-side bug) must not escape and break the whole chat.
|
|
146
|
+
# best_effort: a broken configured logger must not defeat this boundary either.
|
|
147
|
+
Axn::Extensions.best_effort("logging a remote MCP tool failure") do
|
|
148
|
+
Axn.config.logger.error { "[axn-ruby_llm] remote MCP tool #{tool_name.inspect} failed: #{e.class}: #{e.message}" }
|
|
149
|
+
end
|
|
150
|
+
{ error: "The remote tool could not produce a valid response" }
|
|
151
|
+
end
|
|
152
|
+
|
|
153
|
+
# The MCP result shape is `{"result" => {"content" => [{"type" => "text", "text" => "..."}, ...], "isError" => bool}}`
|
|
154
|
+
# (see MCP::Client#call_tool's own doc example: `response.dig("result", "content")`). Only
|
|
155
|
+
# text blocks are supported for now -- Metabase's tools return text/JSON, and a truncated
|
|
156
|
+
# binary block would be meaningless to the model anyway.
|
|
157
|
+
def render_tool_result(response, max_result_chars:)
|
|
158
|
+
result = response["result"] || {}
|
|
159
|
+
text = Array(result["content"]).filter_map { |block| block["text"] }.join("\n")
|
|
160
|
+
text = "#{text[0...max_result_chars]}\n... (truncated at #{max_result_chars} characters)" if text.length > max_result_chars
|
|
161
|
+
|
|
162
|
+
return { error: text.empty? ? "Tool call failed" : text } if result["isError"]
|
|
163
|
+
|
|
164
|
+
text
|
|
165
|
+
end
|
|
166
|
+
end
|
|
167
|
+
end
|
|
168
|
+
end
|
|
169
|
+
end
|
data/lib/axn/ruby_llm/rspec.rb
CHANGED
|
@@ -16,20 +16,26 @@ module Axn
|
|
|
16
16
|
# stub_axn_ruby_llm("Here is a summary.")
|
|
17
17
|
# stub_axn_ruby_llm({ "k" => "v" }, schema: MySchema) # Hash passed through as `parsed`
|
|
18
18
|
# stub_axn_ruby_llm("...", input_tokens: 100, output_tokens: 50, cost: 0.0023)
|
|
19
|
+
# (input_tokens: here is RubyLLM's uncached Tokens#input -- Ask's uncached_input_tokens)
|
|
19
20
|
# stub_axn_ruby_llm("...", cache_read_tokens: 500, cache_write_tokens: 200)
|
|
21
|
+
# stub_axn_ruby_llm("...", thinking_tokens: 300, server_tool_use: { "web_search_requests" => 1 })
|
|
22
|
+
# stub_axn_ruby_llm("partial", finish_reason: :max_tokens) # exercise Ask's truncation failure
|
|
20
23
|
# stub_axn_ruby_llm(response: "...") # keyword form still works
|
|
21
24
|
#
|
|
22
25
|
# Returns the chat instance double for further assertions if needed.
|
|
26
|
+
# rubocop:disable-next Metrics/ParameterLists -- one optional keyword per stubbable usage field
|
|
23
27
|
def stub_axn_ruby_llm(positional_response = UNSET, response: UNSET, model: nil, schema: nil,
|
|
24
28
|
input_tokens: nil, output_tokens: nil, cache_read_tokens: nil,
|
|
25
|
-
cache_write_tokens: nil,
|
|
29
|
+
cache_write_tokens: nil, thinking_tokens: nil, server_tool_use: nil,
|
|
30
|
+
cost: nil, finish_reason: :stop)
|
|
26
31
|
response = positional_response unless positional_response.equal?(UNSET)
|
|
27
32
|
raise ArgumentError, "stub_axn_ruby_llm requires a response (positionally or as `response:`)" if response.equal?(UNSET)
|
|
28
33
|
|
|
29
34
|
resolved_model_id = model || Axn::RubyLLM.config.default_model
|
|
30
|
-
llm_message = _stub_axn_ruby_llm_message(response, resolved_model_id, schema:)
|
|
31
|
-
|
|
32
|
-
|
|
35
|
+
llm_message = _stub_axn_ruby_llm_message(response, resolved_model_id, schema:, finish_reason:)
|
|
36
|
+
tokens = ::RubyLLM::Tokens.new(input: input_tokens, output: output_tokens, cache_read: cache_read_tokens,
|
|
37
|
+
cache_write: cache_write_tokens, thinking: thinking_tokens, server_tool_use:)
|
|
38
|
+
_stub_axn_ruby_llm_chat(model, llm_message, tokens:, cost:)
|
|
33
39
|
end
|
|
34
40
|
|
|
35
41
|
private
|
|
@@ -37,30 +43,40 @@ module Axn
|
|
|
37
43
|
# `content` mirrors real ::RubyLLM::Message#content (a read-only String, JSON text when
|
|
38
44
|
# `schema:` is set); `parsed` mirrors #parsed (the Hash `schema:` callers actually want back
|
|
39
45
|
# via Ask's `parsed_response`, which reads `.parsed` -- not `.content` -- once schema is set).
|
|
40
|
-
|
|
46
|
+
# finish_reason/max_tokens?/content_filtered? mirror the real predicates, so a helper-stubbed
|
|
47
|
+
# call goes through Ask's truncation/filter check like a real one.
|
|
48
|
+
def _stub_axn_ruby_llm_message(response, model_id, schema:, finish_reason:)
|
|
41
49
|
content = schema ? response.to_json : response.to_s
|
|
42
50
|
parsed = schema ? response : nil
|
|
43
|
-
instance_double(::RubyLLM::Message, content:, parsed:, model: model_id
|
|
51
|
+
instance_double(::RubyLLM::Message, content:, parsed:, model: model_id, finish_reason:,
|
|
52
|
+
max_tokens?: finish_reason == :max_tokens,
|
|
53
|
+
content_filtered?: finish_reason == :content_filter)
|
|
44
54
|
end
|
|
45
55
|
|
|
46
|
-
def _stub_axn_ruby_llm_chat(model, llm_message,
|
|
47
|
-
cache_read_tokens:, cache_write_tokens:, cost:)
|
|
56
|
+
def _stub_axn_ruby_llm_chat(model, llm_message, tokens:, cost:)
|
|
48
57
|
chat_instance = instance_double(::RubyLLM::Chat)
|
|
49
58
|
if model
|
|
50
|
-
|
|
59
|
+
# hash_including: Ask also passes provider:/protocol:/assume_model_exists:/context: when set.
|
|
60
|
+
allow(::RubyLLM).to receive(:chat).with(hash_including(model:)).and_return(chat_instance)
|
|
51
61
|
else
|
|
52
62
|
allow(::RubyLLM).to receive(:chat).and_return(chat_instance)
|
|
53
63
|
end
|
|
54
|
-
|
|
64
|
+
# Every public with_* on the real Chat, not a hand-kept list -- Ask forwards whichever inputs
|
|
65
|
+
# the caller set, and the verifying double rejects any method left unstubbed.
|
|
66
|
+
::RubyLLM::Chat.public_instance_methods(false).grep(/\Awith_/).each do |method|
|
|
55
67
|
allow(chat_instance).to receive(method).and_return(chat_instance)
|
|
56
68
|
end
|
|
69
|
+
allow(chat_instance).to receive(:messages=) # history:
|
|
70
|
+
allow(chat_instance).to receive(:awaiting_approval?).and_return(false) # on_remote_tool_approval:
|
|
57
71
|
allow(chat_instance).to receive(:ask).and_return(llm_message)
|
|
72
|
+
# A stubbed call has no real conversation for transcript_entries to walk -- matches how the
|
|
73
|
+
# disabled/stubbed-config path (Ask#stubbed_exposures) also exposes an empty transcript.
|
|
74
|
+
allow(chat_instance).to receive(:messages).and_return([])
|
|
58
75
|
# Ask reads usage off the chat's own ledger (Chat#tokens / Chat#cost), not per-message --
|
|
59
76
|
# a stubbed call is single-turn, so the ledger is just these values directly.
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
)
|
|
77
|
+
# A real ::RubyLLM::Tokens (a plain value object), not a double, so new ledger fields read as nil
|
|
78
|
+
# instead of tripping a verifying double.
|
|
79
|
+
allow(chat_instance).to receive(:tokens).and_return(tokens)
|
|
64
80
|
# Default to zero cost so specs exercise the "cost computed" path.
|
|
65
81
|
# Pass cost: explicitly to assert a specific value.
|
|
66
82
|
allow(chat_instance).to receive(:cost).and_return(instance_double(::RubyLLM::Cost, total: cost || 0.0))
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Axn
|
|
4
|
+
module RubyLLM
|
|
5
|
+
# A shared cap on how many tool calls one chat may make. Backs both `remote_mcp_tools(max_calls:)`
|
|
6
|
+
# (one budget per toolset) and `Ask`'s `max_tool_calls:` (one budget per `ask`, across every
|
|
7
|
+
# app-executed tool). RubyLLM 2.0 removed `Tool::Halt` / `halt_after:`, so `Chat#complete` has no
|
|
8
|
+
# iteration cap of its own.
|
|
9
|
+
#
|
|
10
|
+
# Soft by design: an exhausted budget answers each further call with an error result telling the
|
|
11
|
+
# model to wrap up, rather than raising -- running out is a normal end state for a long research
|
|
12
|
+
# loop, and raising would throw away everything the chat gathered so far. The Mutex keeps the
|
|
13
|
+
# count exact under `tool_options: { concurrency: :threads }`.
|
|
14
|
+
class ToolBudget
|
|
15
|
+
attr_reader :max_calls
|
|
16
|
+
|
|
17
|
+
def initialize(max_calls, noun: "tool calls")
|
|
18
|
+
@max_calls = max_calls
|
|
19
|
+
@noun = noun
|
|
20
|
+
@count = 0
|
|
21
|
+
@mutex = Mutex.new
|
|
22
|
+
end
|
|
23
|
+
|
|
24
|
+
# Returns true (and reserves a slot) when a call is still allowed, false once max_calls has
|
|
25
|
+
# been reached.
|
|
26
|
+
def consume!
|
|
27
|
+
@mutex.synchronize do
|
|
28
|
+
return false if @count >= @max_calls
|
|
29
|
+
|
|
30
|
+
@count += 1
|
|
31
|
+
true
|
|
32
|
+
end
|
|
33
|
+
end
|
|
34
|
+
|
|
35
|
+
def exhausted_result
|
|
36
|
+
{ error: "Tool call budget exhausted (#{@max_calls} #{@noun} allowed) -- " \
|
|
37
|
+
"write your final answer with what you have so far." }
|
|
38
|
+
end
|
|
39
|
+
|
|
40
|
+
# Returns a tool instance whose every call draws from this budget. A class is instantiated; an
|
|
41
|
+
# instance is dup'd first, so the caller's own instance (e.g. one from
|
|
42
|
+
# `Axn::RubyLLM.wrap(axn, ambient_context:)`, possibly reused across asks) is never mutated.
|
|
43
|
+
# RubyLLM's Chat#with_tools registers an instance as-is, and runs it via Tool#call.
|
|
44
|
+
def guard(tool)
|
|
45
|
+
budget = self
|
|
46
|
+
instance = tool.is_a?(::Class) ? tool.new : tool.dup
|
|
47
|
+
instance.extend(Module.new do
|
|
48
|
+
define_method(:call) do |tool_call: nil, **arguments|
|
|
49
|
+
next budget.exhausted_result unless budget.consume!
|
|
50
|
+
|
|
51
|
+
super(tool_call:, **arguments)
|
|
52
|
+
end
|
|
53
|
+
end)
|
|
54
|
+
end
|
|
55
|
+
end
|
|
56
|
+
end
|
|
57
|
+
end
|
data/lib/axn/ruby_llm/version.rb
CHANGED
data/lib/axn/ruby_llm.rb
CHANGED
|
@@ -4,7 +4,9 @@ require "ruby_llm"
|
|
|
4
4
|
require "axn"
|
|
5
5
|
|
|
6
6
|
require_relative "ruby_llm/version"
|
|
7
|
+
require_relative "ruby_llm/tool_budget"
|
|
7
8
|
require_relative "ruby_llm/ask"
|
|
9
|
+
require_relative "ruby_llm/remote_mcp"
|
|
8
10
|
|
|
9
11
|
module Axn
|
|
10
12
|
module RubyLLM
|
|
@@ -43,6 +45,10 @@ module Axn
|
|
|
43
45
|
value = config.enabled
|
|
44
46
|
value.respond_to?(:call) ? !!value.call : !!value
|
|
45
47
|
end
|
|
48
|
+
|
|
49
|
+
def remote_mcp_tools(...)
|
|
50
|
+
RemoteMcp.remote_mcp_tools(...)
|
|
51
|
+
end
|
|
46
52
|
end
|
|
47
53
|
end
|
|
48
54
|
end
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: axn-ruby_llm
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.4.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Kali Donovan
|
|
@@ -63,6 +63,26 @@ dependencies:
|
|
|
63
63
|
- - ">="
|
|
64
64
|
- !ruby/object:Gem::Version
|
|
65
65
|
version: 1.10.0
|
|
66
|
+
- !ruby/object:Gem::Dependency
|
|
67
|
+
name: mcp
|
|
68
|
+
requirement: !ruby/object:Gem::Requirement
|
|
69
|
+
requirements:
|
|
70
|
+
- - ">="
|
|
71
|
+
- !ruby/object:Gem::Version
|
|
72
|
+
version: '1.3'
|
|
73
|
+
- - "<"
|
|
74
|
+
- !ruby/object:Gem::Version
|
|
75
|
+
version: '2.0'
|
|
76
|
+
type: :runtime
|
|
77
|
+
prerelease: false
|
|
78
|
+
version_requirements: !ruby/object:Gem::Requirement
|
|
79
|
+
requirements:
|
|
80
|
+
- - ">="
|
|
81
|
+
- !ruby/object:Gem::Version
|
|
82
|
+
version: '1.3'
|
|
83
|
+
- - "<"
|
|
84
|
+
- !ruby/object:Gem::Version
|
|
85
|
+
version: '2.0'
|
|
66
86
|
description: Call LLMs from Axn actions using RubyLLM, with structured error handling,
|
|
67
87
|
schema-based structured output, and cost/token tracking.
|
|
68
88
|
email:
|
|
@@ -77,8 +97,10 @@ files:
|
|
|
77
97
|
- lib/axn-ruby_llm.rb
|
|
78
98
|
- lib/axn/ruby_llm.rb
|
|
79
99
|
- lib/axn/ruby_llm/ask.rb
|
|
100
|
+
- lib/axn/ruby_llm/remote_mcp.rb
|
|
80
101
|
- lib/axn/ruby_llm/rspec.rb
|
|
81
102
|
- lib/axn/ruby_llm/tool_adapter.rb
|
|
103
|
+
- lib/axn/ruby_llm/tool_budget.rb
|
|
82
104
|
- lib/axn/ruby_llm/version.rb
|
|
83
105
|
homepage: https://github.com/teamshares/axn-ruby_llm
|
|
84
106
|
licenses:
|