llm_logs 0.5.1 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: a687cddb64dc51d51cdf1bfa833833523b54194054ae32fe631354f187c9f256
4
- data.tar.gz: ec0375795a8d27afa1266c1a8ebacd0984395fd8d41753c6ebd4c2658cfef026
3
+ metadata.gz: c131b25cf1499cbebd755808d206f17e7c4dc4199fe18a7d282cabef947b349f
4
+ data.tar.gz: fadedd362588224fad8fc191dc6f9a3ee4d65746ed1d86f239b37e228b448606
5
5
  SHA512:
6
- metadata.gz: fe87b0ed824568ba3c1104e600fdc670a43919c5b061159a7e1e94bffe59257272b74914fd83256633236fd02da3b1ae99efd240ea8a6e645c93880e07d3a7af
7
- data.tar.gz: 81986848eec310e3c2360ceecfbe894a58fc5836adc1a5a6832c90a2c60abbc72898521c6a7b44cf9a1aa118c0837aea861d2e919aeb4fd33b5d83c251e6b22d
6
+ metadata.gz: d94d3962f820cce448da3171f364eeaea2995a707a8fd79e863fae6b9889db321e0ccf6bcead6f12ac48aa50709178e5c8dc2de858db6e111aa4be8ba7bc5cce
7
+ data.tar.gz: 8237d89556f129e2010ee04229d85f2ebe03d3bd11ad7799e1150a3d7c88134c1b4c63fe54e8555f9a3e53b1119e604332b184bb787c9586d87559db26453343
data/README.md CHANGED
@@ -209,16 +209,7 @@ Running the task creates missing prompts, updates metadata, and creates a new pr
209
209
 
210
210
  Send latency-insensitive requests through a provider's Batch API for roughly half the cost. LlmLogs persists each request, groups pending requests into a provider batch, reconciles results, and records a trace per request — so batched work shows up in the dashboard alongside synchronous calls.
211
211
 
212
- Two batch backends are supported, selected **per model**:
213
-
214
- - **[OpenAI Responses Batch API](https://platform.openai.com/docs/guides/batch)** via [`ruby_llm-responses_api`](https://rubygems.org/gems/ruby_llm-responses_api) — the default for OpenAI models.
215
- - **[AWS Bedrock Batch API](#aws-bedrock-batches)** (`CreateModelInvocationJob`) for Anthropic Claude models.
216
-
217
- Add the OpenAI provider to your app's Gemfile:
218
-
219
- ```ruby
220
- gem "ruby_llm-responses_api"
221
- ```
212
+ Batching runs on the **[AWS Bedrock Batch API](#aws-bedrock-batches)** (`CreateModelInvocationJob`) for Anthropic Claude models. Models Bedrock does not serve are not batchable; run them synchronously.
222
213
 
223
214
  ### Enqueue a Request
224
215
 
@@ -227,10 +218,10 @@ Requests are persisted immediately and grouped by `purpose` + `model` when submi
227
218
  ```ruby
228
219
  LlmLogs::Batch.enqueue(
229
220
  purpose: "chat_summary",
230
- model: "gpt-4.1-mini",
221
+ model: "anthropic.claude-haiku-4-5-20251001-v1:0",
231
222
  instructions: "Summarize the conversation in two sentences.",
232
223
  input: conversation_text,
233
- schema: SummarySchema, # optional RubyLLM::Schema for structured output
224
+ schema: SUMMARY_SCHEMA, # optional JSON schema Hash for structured output
234
225
  routing: { conversation_id: 42 }, # your keys, echoed into the trace metadata
235
226
  temperature: 0.2 # optional
236
227
  )
@@ -247,7 +238,8 @@ Register one handler per `purpose`. The gem owns the batch lifecycle; your app o
247
238
  LlmLogs.register_batch_handler("chat_summary", ChatSummaryHandler.new)
248
239
 
249
240
  class ChatSummaryHandler
250
- # Called once a request succeeds. `message` is the RubyLLM::Message.
241
+ # Called once a request succeeds. `message` is a `LlmLogs::Batch::Adapters::Bedrock::Result`
242
+ # (content, input_tokens, output_tokens, model_id).
251
243
  def call(request, message)
252
244
  Conversation.find(request.routing["conversation_id"])
253
245
  .update!(summary: message.content)
@@ -302,7 +294,7 @@ LlmLogs.configuration.bedrock_batch = LlmLogs::Configuration::BedrockBatch.new(
302
294
  LlmLogs.register_batch_adapter(:bedrock, LlmLogs::Batch::Adapters::Bedrock.new)
303
295
  ```
304
296
 
305
- Provider selection is per model: `LlmLogs::Batch.batch_provider_for(model)` returns `:bedrock` when the adapter is registered and `model_matcher` matches, otherwise the OpenAI backend when the model resolves there, otherwise `nil` (not batchable — run it synchronously). Bedrock enforces a **minimum records per job**, so check `LlmLogs::Batch.min_records_for(model)` and fall back to a synchronous call when a batch would be under the floor.
297
+ Provider selection is per model: `LlmLogs::Batch.batch_provider_for(model)` returns `:bedrock` when the adapter is registered and `model_matcher` matches, otherwise `nil` (not batchable — run it synchronously). Bedrock enforces a **minimum records per job**, so check `LlmLogs::Batch.min_records_for(model)` and fall back to a synchronous call when a batch would be under the floor.
306
298
 
307
299
  The adapter builds its AWS clients from the ambient credential chain by default; inject your own to authenticate explicitly:
308
300
 
@@ -330,16 +322,37 @@ Browse traces and manage prompts at `/llm_logs`.
330
322
  ```ruby
331
323
  LlmLogs.setup do |config|
332
324
  config.enabled = true # master switch for logging
333
- config.auto_instrument = true # auto-prepend on RubyLLM::Chat
325
+ config.auto_instrument = true # subscribe to ruby_llm notifications (needs RubyLLM.config.instrumenter = ActiveSupport::Notifications)
334
326
  config.retention_days = 30 # for future cleanup job
335
327
  config.prompts_source_path = Rails.root.join("db/data/prompts")
336
328
  config.prompt_subfolders = %w[skills fragments templates]
337
329
  config.batch_enabled = true # enable the batch API integration
338
- config.batch_provider = :openai_responses # default (OpenAI) backend; Bedrock is registered separately (see Batches)
339
330
  config.page_size = 50 # rows per page on all index pages
340
331
  end
341
332
  ```
342
333
 
334
+ ## ruby_llm 2.0 integration
335
+
336
+ With `auto_instrument` on, LlmLogs subscribes to ruby_llm's `chat.ruby_llm` and `tool_call.ruby_llm` notifications: one `llm` span per provider round and one tool span per tool call, all siblings under the trace. Tokens and cost come from each round's usage entries. This requires `RubyLLM.config.instrumenter` to be `ActiveSupport::Notifications`; the Railtie sets that default, and `LlmLogs::Instrumentation::RubyLlmChat.install!` sets it when unset.
337
+
338
+ ### Patches for ruby_llm 2.0.x
339
+
340
+ The engine prepends five fixes (`LlmLogs::RubyLLMPatches`) that ruby_llm 2.0.0 lacks. Each one backports an upstream pull request, so dropping it once a ruby_llm release includes that PR changes no behaviour (exceptions are noted). They install only on ruby_llm >= 2.0.0, and log a warning on 2.1 or later so you re-check them.
341
+
342
+ | Patch | Upstream PR |
343
+ |---|---|
344
+ | BedrockSigV4, RetrySSLError | [crmne/ruby_llm#1024](https://github.com/crmne/ruby_llm/pull/1024) |
345
+ | ConverseReasoningConfig, ConverseClaudeAdaptiveThinking | [crmne/ruby_llm#1025](https://github.com/crmne/ruby_llm/pull/1025) |
346
+ | ConverseForeignReasoning | [crmne/ruby_llm#1026](https://github.com/crmne/ruby_llm/pull/1026) |
347
+
348
+ - **BedrockSigV4** ([#1024](https://github.com/crmne/ruby_llm/pull/1024), "Sign each Bedrock attempt when it is sent") re-signs every Bedrock attempt at send time, over the body being sent, and removes an `X-Amz-Security-Token` the new signature no longer carries. Upstream 2.0.0 signs once before the retry middleware, so a retry after a long timeout replays an expired signature (403 "Signature expired"). Drop when a ruby_llm release includes #1024.
349
+ - **RetrySSLError** ([#1024](https://github.com/crmne/ruby_llm/pull/1024), "Retry requests that fail with a TLS error") adds `Faraday::SSLError` (for example "SSL_connect ... unexpected eof") to the retried exceptions. This is the same trade-off ruby_llm already accepts for read timeouts: a request that reached the server may be sent twice. ruby_llm's `retry_if` still refuses to retry requests marked non-idempotent and streams that already delivered content. Drop when a ruby_llm release includes #1024.
350
+ - **ConverseReasoningConfig** ([#1025](https://github.com/crmne/ruby_llm/pull/1025), "Send reasoning_config to models that publish it") sends a reasoning effort on Bedrock Converse as `additionalModelRequestFields: {reasoning_config: "<effort>"}` (including `none`, which is also what `with_thinking(false)` resolves to for these models) when the model's Converse `additionalRequestFieldsSchema`, or that of another registry entry for the same foundation model, publishes a `reasoning_config` enum (OpenAI GPT 5.6/6 Astra, xAI Grok 4.6, ...). Upstream 2.0.0 sends `reasoning_effort`, which Bedrock rejects for GPT (400 unknown_parameter). Thinking turned off, budgets, Claude, Nova, and gpt-oss are unchanged. Beyond #1025, OpenAI GPT ids with no published schema (ruby_llm 2.0.0's registry lacks GPT-6 Sol/Luna) get the same shape, under any region prefix. Drop when a ruby_llm release includes #1025; that fallback goes with it, so first check that the registry you run publishes the schema for every GPT id you use (or add `metadata.converse.additionalRequestFieldsSchema` to your catalog entry), or those ids go back to `reasoning_effort`.
351
+ - **ConverseClaudeAdaptiveThinking** ([#1025](https://github.com/crmne/ruby_llm/pull/1025), "Think adaptively on adaptive-only Claude via Converse") sends adaptive thinking to adaptive-only Claude models on Bedrock Converse (the registry advertises an `effort` option and no `budget_tokens` option: Sonnet 5, Opus 4.7/4.8/5, Fable 5) as `additionalModelRequestFields: {thinking: {type: "adaptive"}, output_config: {effort: ...}}` (every advertised tier, `xhigh` and `max` included; `{thinking: {type: "adaptive"}}` alone when no effort is set, as for `with_thinking(true)` or a `display` alone; nothing for `none`; no `display` is sent). Upstream 2.0.0 sends a fixed `reasoning_config` budget, which these models reject ("thinking.type.enabled" is not supported for this model). Any budget, thinking turned off, budget-style Claude (Haiku 4.5, Sonnet 4.6), and other vendors are unchanged. An unregistered id takes its reasoning options from a registry entry for the same foundation model. Drop when a ruby_llm release includes #1025.
352
+ - **ConverseForeignReasoning** ([#1026](https://github.com/crmne/ruby_llm/pull/1026), "Drop Bedrock reasoning another model family produced" and "Give application inference profiles no vendor") replays a message's reasoning only to a model of the same family on Bedrock Converse. Upstream 2.0.0 replays every assistant message's reasoning to whatever model is next, and Claude and GPT are both the `bedrock` provider: a chat Claude answered and GPT continues fails with "This model doesn't support the reasoningContent.reasoningText.text field", and a chat GPT answered and Claude continues hands Claude GPT's encrypted reasoning. A model's family is `anthropic`, `openai`, or the vendor segment of any other Bedrock id (`amazon`, `meta`, ...), whatever its region prefix; an application inference profile ARN names no family, even with a dot in its id (`.../application-inference-profile/team.prod`). The producer is `RubyLLM::Message#model` (for a persisted message, the model of its last successful usage entry). Reasoning from another family is dropped. Anthropic targets otherwise get upstream's replay unchanged, and so does a target that names no family. **This patch is stricter than #1026 on purpose:** a non-Anthropic target whose producer is unknown or of the same family gets only the `redactedContent` blocks stored in `raw_reasoning["converse"]` (GPT's own encrypted reasoning across tool-call rounds), never `reasoningText` or the thinking text/signature fallback, where #1026 replays them all. Drop when a ruby_llm release includes #1026, accepting that looser replay (or keep the stricter rule as a patch of its own).
353
+
354
+ Not backported: #1025's first commit ("Recognise the in. Bedrock inference profile prefix"), which adds `in` to `Converse::REGION_PREFIXES` and also changes region rewriting and Mantle routing. The patches here match ids under any region prefix instead.
355
+
343
356
  ## Requirements
344
357
 
345
358
  - Rails 8.0+
@@ -1,12 +1,13 @@
1
1
  module LlmLogs
2
2
  class Batch
3
- # Groups pending BatchRequests of one purpose+model into a single OpenAI batch via
4
- # ruby_llm-responses_api. To prevent two concurrent FlushJobs from double-submitting
5
- # the same requests, it first CLAIMS the pending rows in a `FOR UPDATE SKIP LOCKED`
6
- # transaction (assigning them to a placeholder Batch with no openai_batch_id, which
7
- # flips them out of the `pending` scope and which PollJob ignores). It then submits to
8
- # OpenAI and records the batch id. If submission fails, the claim is released (requests
9
- # return to `pending`) and the placeholder batch is dropped, so the work retries next flush.
3
+ # Groups pending BatchRequests of one purpose+model into a single provider batch (Bedrock).
4
+ # To keep two concurrent FlushJobs from double-submitting the same requests, it first
5
+ # CLAIMS the pending rows in a `FOR UPDATE SKIP LOCKED` transaction, assigning them to a
6
+ # placeholder Batch with no provider batch id (stored in the legacy `openai_batch_id`
7
+ # column), which flips them out of the `pending` scope and which PollJob ignores. It then
8
+ # submits the batch and records its id. If submission fails, the claim is released
9
+ # (requests return to `pending`) and the placeholder batch is dropped, so the work retries
10
+ # on the next flush.
10
11
  class Submitter
11
12
  def initialize(purpose:, model:, metadata: {})
12
13
  @purpose = purpose
@@ -27,23 +27,35 @@ module LlmLogs
27
27
  output: { "content" => span.serialize_content(message.content) },
28
28
  input_tokens: message.input_tokens,
29
29
  output_tokens: message.output_tokens,
30
- cost: compute_cost(message)
30
+ cost: compute_cost(message, request: request, provider: provider)
31
31
  )
32
32
  span.finish
33
33
  end
34
34
  trace
35
35
  end
36
36
 
37
- def compute_cost(message)
38
- model_info = RubyLLM.models.find(message.model_id)
39
- return nil unless model_info&.input_price_per_million && model_info&.output_price_per_million
37
+ # Prices the request at the submitted model's standard rates for this provider (the
38
+ # Bedrock id, e.g. us.anthropic.claude-haiku-4-5-...), falling back to the id the result
39
+ # reports (a native Anthropic id resolves to the anthropic provider's rates). nil when
40
+ # neither is in the registry or the registry has no price for the tokens used.
41
+ def compute_cost(message, request:, provider:)
42
+ model = pricing_model(request.model, provider) || pricing_model(message.model_id, nil)
43
+ return nil unless model
40
44
 
41
- raw = (message.input_tokens.to_f * model_info.input_price_per_million +
42
- message.output_tokens.to_f * model_info.output_price_per_million) / 1_000_000
43
- (raw * BATCH_COST_MULTIPLIER).round(6)
45
+ tokens = RubyLLM::Tokens.new(input: message.input_tokens.to_i, output: message.output_tokens.to_i)
46
+ total = model.cost_for(tokens).total
47
+ total && (total * BATCH_COST_MULTIPLIER).round(6)
44
48
  rescue StandardError
45
49
  nil
46
50
  end
51
+
52
+ def pricing_model(model_id, provider)
53
+ return nil if model_id.blank?
54
+
55
+ provider.present? ? RubyLLM.models.find(model_id, provider: provider) : RubyLLM.models.find(model_id)
56
+ rescue RubyLLM::ModelNotFoundError
57
+ nil
58
+ end
47
59
  end
48
60
  end
49
61
  end
@@ -66,14 +66,10 @@ module LlmLogs
66
66
  !batch_provider_for(model).nil?
67
67
  end
68
68
 
69
- # Which batch provider (if any) serves this model. Bedrock wins for Claude models when
70
- # the Bedrock adapter is configured; otherwise fall back to the OpenAI provider when the
71
- # model resolves there; otherwise nil (run synchronously).
69
+ # Which batch provider (if any) serves this model: :bedrock when the Bedrock adapter is
70
+ # registered and its model_matcher matches, otherwise nil (run synchronously).
72
71
  def self.batch_provider_for(model)
73
- return :bedrock if bedrock_serves?(model)
74
- return :openai_responses if openai_serves?(model)
75
-
76
- nil
72
+ bedrock_serves?(model) ? :bedrock : nil
77
73
  end
78
74
 
79
75
  # The Bedrock minimum records-per-job floor for this model (0 when Bedrock does not serve it).
@@ -89,25 +85,5 @@ module LlmLogs
89
85
  matcher = config.model_matcher
90
86
  matcher.respond_to?(:call) ? matcher.call(model.to_s) : matcher.match?(model.to_s)
91
87
  end
92
-
93
- def self.openai_serves?(model)
94
- return false unless defined?(RubyLLM::Providers::OpenAIResponses)
95
-
96
- servable_by_batch_provider?(model)
97
- end
98
-
99
- # The batch path submits via RubyLLM.batch(provider: batch_provider). A model is
100
- # only batchable if that provider can actually serve it -- i.e. the model resolves
101
- # under batch_provider. Models that belong to a different provider (e.g. Bedrock /
102
- # Anthropic) don't resolve there, so they return false and the caller runs them
103
- # synchronously instead of enqueueing work that would fail at submit time.
104
- def self.servable_by_batch_provider?(model)
105
- RubyLLM::Models.resolve(
106
- model, provider: LlmLogs.batch_provider, assume_exists: false, config: RubyLLM.config
107
- )
108
- true
109
- rescue RubyLLM::ModelNotFoundError
110
- false
111
- end
112
88
  end
113
89
  end
@@ -19,11 +19,23 @@ module LlmLogs
19
19
  LlmLogs::Tracer.current_span = parent_span
20
20
  end
21
21
 
22
+ # +message+ is a RubyLLM 2.x Message (or anything with +content+ and +tokens+).
22
23
  def record_response(message)
23
24
  self.output = { content: serialize_content(message.content) }
24
- self.input_tokens = message.input_tokens
25
- self.output_tokens = message.output_tokens
26
- self.cached_tokens = message.cached_tokens
25
+ record_tokens(message.tokens)
26
+ end
27
+
28
+ # Copies a RubyLLM::Tokens onto the span. +input+ is non-cached input (RubyLLM 2.x no
29
+ # longer subtracts cache buckets from it); cache writes and thinking have no column and
30
+ # go to metadata.
31
+ def record_tokens(tokens)
32
+ return unless tokens
33
+
34
+ self.input_tokens = tokens.input
35
+ self.output_tokens = tokens.output
36
+ self.cached_tokens = tokens.cache_read
37
+ set_attribute("cache_write_tokens", tokens.cache_write) unless tokens.cache_write.nil?
38
+ set_attribute("thinking_tokens", tokens.thinking) unless tokens.thinking.nil?
27
39
  end
28
40
 
29
41
  # Structured (schema) responses arrive as a Hash/Array; keep them as-is so the
@@ -3,7 +3,7 @@ module LlmLogs
3
3
  BedrockBatch = Struct.new(:role_arn, :s3_bucket, :s3_prefix, :min_records, :model_matcher, :region, keyword_init: true)
4
4
 
5
5
  attr_accessor :enabled, :auto_instrument, :retention_days, :prompts_source_path, :prompt_subfolders,
6
- :batch_enabled, :batch_provider, :page_size, :bedrock_batch, :reasoning_effort_options
6
+ :batch_enabled, :page_size, :bedrock_batch, :reasoning_effort_options
7
7
 
8
8
  def initialize
9
9
  @enabled = true
@@ -12,7 +12,6 @@ module LlmLogs
12
12
  @prompts_source_path = nil
13
13
  @prompt_subfolders = %w[skills fragments templates]
14
14
  @batch_enabled = true
15
- @batch_provider = :openai_responses
16
15
  @page_size = 50
17
16
  @bedrock_batch = nil
18
17
  # Effort tiers offered in the prompt form. Union of what current providers
@@ -6,11 +6,20 @@ module LlmLogs
6
6
  ActiveSupport.on_load(:active_record) do
7
7
  if LlmLogs.auto_instrument && defined?(RubyLLM::Chat)
8
8
  require "llm_logs/instrumentation/ruby_llm_chat"
9
- RubyLLM::Chat.prepend(LlmLogs::Instrumentation::RubyLlmChat)
9
+ LlmLogs::Instrumentation::RubyLlmChat.install!
10
10
  end
11
11
  end
12
12
  end
13
13
 
14
+ # Bedrock fixes missing from upstream ruby_llm 2.0.0. Installed whether or not
15
+ # auto-instrumentation is on: they change request behaviour, not logging.
16
+ initializer "llm_logs.ruby_llm_patches" do
17
+ if defined?(RubyLLM::VERSION)
18
+ require "llm_logs/ruby_llm_patches"
19
+ LlmLogs::RubyLLMPatches.install!
20
+ end
21
+ end
22
+
14
23
  rake_tasks do
15
24
  load File.expand_path("../tasks/llm_logs.rake", __dir__)
16
25
  end
@@ -1,105 +1,213 @@
1
1
  module LlmLogs
2
2
  module Instrumentation
3
+ # Records RubyLLM 2.x chat activity as llm_logs spans by subscribing to the
4
+ # events RubyLLM emits for observability adapters (docs/_advanced/instrumentation.md):
5
+ #
6
+ # chat.ruby_llm -- one per model request (RubyLLM::Chat#generate_once): an "llm" span
7
+ # per provider round. Wraps every transport retry of that round, so a
8
+ # round that fails after its retries surfaces as :exception_object.
9
+ # With fallbacks, each model tried gets its own span.
10
+ # tool_call.ruby_llm -- one per local tool execution (RubyLLM::Chat#execute_tool): a "tool" span.
11
+ #
12
+ # Rounds and tools are siblings under the current span/trace:
13
+ # trace -> llm (round 1, asks for a tool) -> tool.x -> llm (round 2, final answer)
14
+ #
15
+ # The subscription is global (every RubyLLM::Chat, including acts_as_chat's to_llm and
16
+ # RubyLLM::Agent), but only fires while RubyLLM's instrumenter is ActiveSupport::Notifications,
17
+ # which RubyLLM's Railtie sets by default.
3
18
  module RubyLlmChat
4
- def complete(&block)
5
- return super unless LlmLogs.enabled?
6
-
7
- span = LlmLogs::Tracer.start_span(
8
- name: "chat.complete",
9
- span_type: "llm",
10
- model: @model&.id,
11
- provider: @model&.provider,
12
- input: messages.map { |m| { role: m.role, content: llm_logs_serialize_content(m.content) } }
13
- )
14
-
15
- messages_before = messages.size
16
-
17
- begin
18
- result = super(&block)
19
- span.record_response(result)
20
- span.cost = llm_logs_compute_cost(result)
21
- result
22
- rescue => e
23
- span.record_error(e)
24
- llm_logs_capture_partial_tokens(span, messages_before)
25
- raise
26
- ensure
27
- span.finish
19
+ SPAN_KEY = :llm_logs_span
20
+ USAGE_START_KEY = :llm_logs_usage_start
21
+
22
+ class << self
23
+ def install!
24
+ return if installed?
25
+
26
+ RubyLLM.config.instrumenter ||= ActiveSupport::Notifications
27
+ @subscribers = [
28
+ ActiveSupport::Notifications.subscribe("chat.ruby_llm", ChatSubscriber.new),
29
+ ActiveSupport::Notifications.subscribe("tool_call.ruby_llm", ToolSubscriber.new)
30
+ ]
28
31
  end
29
- end
30
32
 
31
- private
33
+ def uninstall!
34
+ Array(@subscribers).each { |subscriber| ActiveSupport::Notifications.unsubscribe(subscriber) }
35
+ @subscribers = nil
36
+ end
32
37
 
33
- def llm_logs_serialize_content(content)
34
- content.is_a?(Hash) || content.is_a?(Array) ? content.to_json : content.to_s
35
- end
38
+ def installed?
39
+ !@subscribers.nil?
40
+ end
36
41
 
37
- def llm_logs_serialize_tool_result(result)
38
- if result.is_a?(Hash)
39
- result
40
- elsif defined?(RubyLLM::Tool::Halt) && result.is_a?(RubyLLM::Tool::Halt)
41
- { halted: result.message }
42
- else
43
- { result: result.to_s }
42
+ # Instrumentation must never break the LLM call it observes. ActiveSupport re-raises
43
+ # subscriber exceptions into the instrumented block, so every hook rescues.
44
+ def guard
45
+ yield
46
+ rescue StandardError => e
47
+ Rails.logger&.error("[llm_logs] instrumentation error: #{e.class}: #{e.message}")
48
+ nil
44
49
  end
45
50
  end
46
51
 
47
- def llm_logs_capture_partial_tokens(span, messages_before)
48
- assistant_msg = messages[messages_before..].find { |m| m.role == :assistant }
49
- return unless assistant_msg&.input_tokens
52
+ # Serializers shared by both subscribers.
53
+ module Serialize
54
+ module_function
50
55
 
51
- span.input_tokens = assistant_msg.input_tokens
52
- span.output_tokens = assistant_msg.output_tokens
53
- span.cached_tokens = assistant_msg.cached_tokens
54
- span.cost = llm_logs_compute_cost(assistant_msg)
55
- rescue StandardError
56
- nil
57
- end
56
+ def messages(list)
57
+ Array(list).map { |message| message_entry(message) }
58
+ end
59
+
60
+ def message_entry(message)
61
+ entry = {role: message.role, content: message.content}
62
+ calls = tool_calls(message.tool_calls)
63
+ entry[:tool_calls] = calls if calls
64
+ entry[:tool_call_id] = message.tool_call_id if message.tool_call_id
65
+ entry
66
+ end
67
+
68
+ def tool_calls(calls)
69
+ return nil if calls.nil? || calls.empty?
70
+
71
+ calls.values.map { |call| {id: call.id, name: call.name, arguments: call.arguments} }
72
+ end
58
73
 
59
- def llm_logs_compute_cost(message)
60
- model_info = llm_logs_resolve_pricing_model
61
- return unless model_info
74
+ # Structured (schema) answers are JSON text in 2.x; store the parsed Hash so the
75
+ # UI renders nested fields instead of an escaped string.
76
+ def response(message, schema:)
77
+ content = message.content
78
+ content = parse_json(content) if schema && content.is_a?(String) && !content.empty?
79
+ output = {content: content}
80
+ calls = tool_calls(message.tool_calls)
81
+ output[:tool_calls] = calls if calls
82
+ output[:thinking] = message.thinking.text if message.thinking&.text.present?
83
+ output
84
+ end
62
85
 
63
- input_price = model_info.input_price_per_million
64
- output_price = model_info.output_price_per_million
65
- return unless input_price && output_price
86
+ def parse_json(text)
87
+ JSON.parse(text)
88
+ rescue JSON::ParserError
89
+ text
90
+ end
66
91
 
67
- (message.input_tokens.to_f * input_price + message.output_tokens.to_f * output_price) / 1_000_000
92
+ def tool_result(result)
93
+ case result
94
+ when Hash then result
95
+ when Array then {result: result}
96
+ else {result: result.to_s}
97
+ end
98
+ end
68
99
  end
69
100
 
70
- def llm_logs_resolve_pricing_model
71
- return @model if @model&.input_price_per_million
101
+ class ChatSubscriber
102
+ def start(_name, _id, payload)
103
+ RubyLlmChat.guard do
104
+ next unless LlmLogs.enabled?
105
+
106
+ chat = payload[:chat]
107
+ payload[USAGE_START_KEY] = chat.usage_entries.length if chat.respond_to?(:usage_entries)
108
+ payload[SPAN_KEY] = LlmLogs::Tracer.start_span(
109
+ name: "chat.complete",
110
+ span_type: "llm",
111
+ model: payload[:model],
112
+ provider: payload[:provider],
113
+ input: Serialize.messages(payload[:input_messages]),
114
+ metadata: request_metadata(payload)
115
+ )
116
+ end
117
+ end
72
118
 
73
- RubyLLM.models.find(@model.id) if @model&.id
74
- rescue StandardError
75
- nil
119
+ def finish(_name, _id, payload)
120
+ span = payload.delete(SPAN_KEY)
121
+ return unless span
122
+
123
+ RubyLlmChat.guard do
124
+ if (error = payload[:exception_object])
125
+ span.record_error(error)
126
+ else
127
+ span.output = Serialize.response(payload[:response], schema: payload[:schema])
128
+ span.set_attribute("finish_reason", payload[:response].finish_reason&.to_s)
129
+ response_model = payload[:response_model]
130
+ span.set_attribute("response_model", response_model) if response_model && response_model != payload[:model]
131
+ end
132
+ record_usage(span, payload)
133
+ end
134
+ ensure
135
+ RubyLlmChat.guard { span&.finish }
136
+ end
137
+
138
+ private
139
+
140
+ def request_metadata(payload)
141
+ {
142
+ "streaming" => payload[:streaming],
143
+ "tools" => Array(payload[:tools]).map(&:to_s),
144
+ "temperature" => payload[:temperature],
145
+ "thinking" => thinking_metadata(payload[:thinking]),
146
+ "schema" => payload[:schema] && payload[:schema][:name]
147
+ }.compact
148
+ end
149
+
150
+ def thinking_metadata(thinking)
151
+ return nil unless thinking
152
+
153
+ {"effort" => thinking.effort&.to_s, "budget" => thinking.budget, "enabled" => thinking.enabled}.compact.presence
154
+ end
155
+
156
+ # Tokens and cost of every transport attempt this round made (retries included), read
157
+ # from the chat's usage ledger. Falls back to the event's tokens/cost when the ledger is
158
+ # unavailable. A failed attempt whose usage is unknown (e.g. a timeout that may have
159
+ # been billed) is left out of the cost and flagged with cost_complete: false.
160
+ def record_usage(span, payload)
161
+ entries = round_usage_entries(payload)
162
+ if entries
163
+ tokens = RubyLLM::Tokens.aggregate(entries.map(&:tokens))
164
+ cost = RubyLLM::Cost.aggregate(entries.map(&:cost))
165
+ span.set_attribute("attempts", entries.size) if entries.size > 1
166
+ span.set_attribute("cost_complete", false) unless entries.all?(&:cost_available?)
167
+ else
168
+ tokens = payload[:tokens]
169
+ cost = payload[:cost]
170
+ end
171
+ span.record_tokens(tokens)
172
+ span.cost = cost&.total
173
+ end
174
+
175
+ def round_usage_entries(payload)
176
+ start = payload[USAGE_START_KEY]
177
+ chat = payload[:chat]
178
+ return nil unless start && chat.respond_to?(:usage_entries)
179
+
180
+ chat.usage_entries.drop(start)
181
+ end
76
182
  end
77
183
 
78
- public
79
-
80
- def execute_tool(tool_call)
81
- return super unless LlmLogs.enabled?
82
-
83
- span = LlmLogs::Tracer.start_span(
84
- name: "tool.#{tool_call.name}",
85
- span_type: "tool",
86
- input: tool_call.arguments,
87
- metadata: { tool_name: tool_call.name }
88
- )
89
-
90
- begin
91
- result = super
92
- span.output = llm_logs_serialize_tool_result(result)
93
- span.set_attribute("tool.halted", true) if defined?(RubyLLM::Tool::Halt) && result.is_a?(RubyLLM::Tool::Halt)
94
- # Convert hash/array results to JSON so RubyLLM stores clean JSON
95
- # in message content, not Ruby hash syntax from Hash#to_s
96
- result = result.to_json if result.is_a?(Hash) || result.is_a?(Array)
97
- result
98
- rescue => e
99
- span.record_error(e)
100
- raise
184
+ class ToolSubscriber
185
+ def start(_name, _id, payload)
186
+ RubyLlmChat.guard do
187
+ next unless LlmLogs.enabled?
188
+
189
+ payload[SPAN_KEY] = LlmLogs::Tracer.start_span(
190
+ name: "tool.#{payload[:tool_name]}",
191
+ span_type: "tool",
192
+ input: payload[:tool_arguments],
193
+ metadata: {tool_name: payload[:tool_name].to_s, tool_call_id: payload[:tool_call_id]}.compact
194
+ )
195
+ end
196
+ end
197
+
198
+ def finish(_name, _id, payload)
199
+ span = payload.delete(SPAN_KEY)
200
+ return unless span
201
+
202
+ RubyLlmChat.guard do
203
+ if (error = payload[:exception_object])
204
+ span.record_error(error)
205
+ else
206
+ span.output = Serialize.tool_result(payload[:result])
207
+ end
208
+ end
101
209
  ensure
102
- span.finish
210
+ RubyLlmChat.guard { span&.finish }
103
211
  end
104
212
  end
105
213
  end
@@ -0,0 +1,75 @@
1
+ require "faraday"
2
+
3
+ module LlmLogs
4
+ module RubyLLMPatches
5
+ # Signs every Bedrock attempt when it is sent.
6
+ # Backport of crmne/ruby_llm#1024 ("Sign each Bedrock attempt when it is sent").
7
+ #
8
+ # ruby_llm 2.0.0 signs inside the request block (Protocols::Converse#signed_post,
9
+ # Converse::Streaming#stream_response, InvokeModel, Mantle, Guardrails, Rerank, ...),
10
+ # i.e. once, before Faraday's retry middleware runs. A retry after a long timeout
11
+ # (request_timeout 300s vs SigV4's 5-minute window) then replays an expired
12
+ # X-Amz-Date/Authorization and AWS answers 403 "Signature expired", which is not retried.
13
+ #
14
+ # This adds a middleware innermost in every Bedrock Transport::Connection stack (after
15
+ # :retry and after the :json encoder), so each attempt is re-signed over the exact body
16
+ # bytes being sent, at send time. It only touches requests that already carry a SigV4
17
+ # Authorization header and keeps the signing service named in it (bedrock,
18
+ # bedrock-mantle, ...). Auth#signed_get/#signed_post (model listing, batch control
19
+ # endpoints) use Connection.basic, which has no retry, so they need no change.
20
+ module BedrockSigV4
21
+ class Middleware < Faraday::Middleware
22
+ CREDENTIAL = %r{\AAWS4-HMAC-SHA256 Credential=[^/]+/\d{8}/[^/]+/(?<service>[^/]+)/aws4_request}
23
+
24
+ def initialize(app, provider:)
25
+ super(app)
26
+ @provider = provider
27
+ end
28
+
29
+ def call(env)
30
+ resign(env)
31
+ @app.call(env)
32
+ end
33
+
34
+ private
35
+
36
+ def resign(env)
37
+ match = CREDENTIAL.match(env.request_headers["Authorization"].to_s)
38
+ return unless match
39
+
40
+ body = env.body.nil? ? "" : env.body
41
+ return unless body.is_a?(String) # multipart/IO bodies keep the original signature
42
+
43
+ url = env.url
44
+ path = url.query ? "#{url.path}?#{url.query}" : url.path
45
+ headers = @provider.sign_headers(
46
+ env.method.to_s.upcase, path, body,
47
+ base_url: "#{url.scheme}://#{url.host}", service: match[:service]
48
+ )
49
+ env.request_headers.delete("X-Amz-Security-Token") # credentials may have rotated
50
+ env.request_headers.merge!(headers.except("Content-Type"))
51
+ end
52
+ end
53
+
54
+ module ConnectionPatch
55
+ private
56
+
57
+ def setup_middleware(faraday)
58
+ super
59
+ faraday.use(Middleware, provider: @provider) if BedrockSigV4.bedrock?(@provider)
60
+ end
61
+ end
62
+
63
+ module_function
64
+
65
+ def bedrock?(provider)
66
+ defined?(RubyLLM::Providers::Bedrock) && provider.is_a?(RubyLLM::Providers::Bedrock) &&
67
+ provider.respond_to?(:sign_headers)
68
+ end
69
+
70
+ def install!
71
+ RubyLLMPatches.prepend_checked(RubyLLM::Transport::Connection, ConnectionPatch, %i[setup_middleware])
72
+ end
73
+ end
74
+ end
75
+ end
@@ -0,0 +1,70 @@
1
+ module LlmLogs
2
+ module RubyLLMPatches
3
+ # Adaptive thinking for adaptive-only Claude models on Bedrock Converse.
4
+ # Backport of crmne/ruby_llm#1025 ("Think adaptively on adaptive-only Claude via Converse").
5
+ #
6
+ # ruby_llm 2.0.0 (Protocols::Converse::Chat#format_reasoning_fields) turns a Claude
7
+ # reasoning effort into a fixed budget, `additionalModelRequestFields: {reasoning_config:
8
+ # {type: "enabled", budget_tokens: N}}`, sized from the Converse budgetTokens schema.
9
+ # Claude models that only think adaptively (the registry advertises an effort option and
10
+ # no budget_tokens option: Sonnet 5, Opus 4.7/4.8/5, Fable 5) reject it with
11
+ # `"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive"
12
+ # and "output_config.effort"`, and with_thinking(true) sends no thinking at all. They take
13
+ # `{thinking: {type: "adaptive"}, output_config: {effort: "low"}}`, the rule upstream's own
14
+ # Anthropic protocol already applies (Protocols::Anthropic::Chat#thinking_mode).
15
+ #
16
+ # For an adaptive-only Claude target with no budget set:
17
+ # - effort other than "none": adaptive thinking, effort in output_config (every advertised
18
+ # tier, xhigh and max included);
19
+ # - no effort (with_thinking(true), or a display alone): adaptive thinking alone;
20
+ # - effort "none": nothing.
21
+ # No `display` is sent (Converse has no field for it). Everything else goes to super: any
22
+ # budget, thinking turned off (reasoning_config disabled), budget-style Claude (Haiku 4.5,
23
+ # Sonnet 4.6), GPT, Nova and other vendors.
24
+ #
25
+ # The target is Anthropic under any region prefix (ConverseForeignReasoning::ANTHROPIC).
26
+ # A model with no reasoning options of its own (an unregistered id, such as an "in."
27
+ # inference profile) takes them from a Bedrock registry entry for the same foundation model.
28
+ module ConverseClaudeAdaptiveThinking
29
+ ANTHROPIC = ConverseForeignReasoning::ANTHROPIC
30
+ REGION = /\A[a-z0-9-]+\.(?=anthropic\.)/
31
+
32
+ private
33
+
34
+ def format_reasoning_fields(thinking, model, max_output_tokens = nil)
35
+ return super unless thinking&.enabled? && thinking.enabled != false && thinking.budget.nil?
36
+ return super unless llm_logs_adaptive_only_claude?(model)
37
+
38
+ effort = thinking.effort.to_s
39
+ return nil if effort == "none"
40
+ return {thinking: {type: "adaptive"}} if effort.empty?
41
+
42
+ {thinking: {type: "adaptive"}, output_config: {effort: effort}}
43
+ end
44
+
45
+ def llm_logs_adaptive_only_claude?(model)
46
+ id = foundation_model_id(model&.id)
47
+ return false unless ANTHROPIC.match?(id)
48
+
49
+ options = llm_logs_reasoning_options(model, id.sub(REGION, ""))
50
+ options.any? { |o| o[:type] == "effort" } && options.none? { |o| o[:type] == "budget_tokens" }
51
+ end
52
+
53
+ def llm_logs_reasoning_options(model, bare_id)
54
+ return model.reasoning_options if model.reasoning_options.any?
55
+
56
+ sibling = RubyLLM.models.all.find do |candidate|
57
+ candidate.provider == "bedrock" && candidate.reasoning_options.any? &&
58
+ foundation_model_id(candidate.id).sub(REGION, "") == bare_id
59
+ end
60
+ sibling ? sibling.reasoning_options : []
61
+ end
62
+
63
+ def self.install!
64
+ RubyLLMPatches.prepend_checked(
65
+ RubyLLM::Protocols::Converse, self, %i[format_reasoning_fields foundation_model_id]
66
+ )
67
+ end
68
+ end
69
+ end
70
+ end
@@ -0,0 +1,80 @@
1
+ module LlmLogs
2
+ module RubyLLMPatches
3
+ # Bedrock Converse replays a message's reasoning only to a model of the same family.
4
+ # Backport of crmne/ruby_llm#1026, stricter than it on purpose (see below).
5
+ #
6
+ # ruby_llm 2.0.0 (Protocols::Converse::Chat#format_thinking_blocks) replays every assistant
7
+ # message's reasoning (raw_reasoning["converse"] blocks, else thinking text/signature) as
8
+ # reasoningContent whatever the target model. Upstream only strips thinking a *different
9
+ # provider* produced (Protocol#foreign_thinking?), and Claude and GPT are both "bedrock".
10
+ # So a chat Claude answered and GPT continues fails with 400 "This model doesn't support
11
+ # the reasoningContent.reasoningText.text field for assistant message", and a chat GPT
12
+ # answered and Claude continues hands Claude GPT's encrypted reasoning (redactedContent,
13
+ # or a reasoningText with an empty text and GPT's blob as signature for rows migrated from
14
+ # 1.16), which Claude cannot validate.
15
+ #
16
+ # The family of a model id is "anthropic", "openai", or the vendor segment of any other
17
+ # Bedrock id ("amazon", "meta", ...), whatever region prefix it carries. The producer is
18
+ # RubyLLM::Message#model (the reply's model, or for a restored row the model of its last
19
+ # successful ruby_llm_usages entry).
20
+ #
21
+ # - Producer known and of another family: no reasoning is replayed.
22
+ # - Anthropic target (same family or producer unknown): super, unchanged.
23
+ # - Other named-vendor target (same family or producer unknown): only the redactedContent
24
+ # blocks in raw_reasoning["converse"] (how GPT keeps its own encrypted reasoning across
25
+ # tool-call rounds); reasoningText and the thinking text/signature fallback are dropped.
26
+ # - Target whose id names no model (application-inference-profile ARN): super, unchanged.
27
+ #
28
+ # An application inference profile id is opaque, and AWS allows dots in it
29
+ # (".../application-inference-profile/team.prod"), so it names no family as producer or
30
+ # target (backport of crmne/ruby_llm#1026, "Give application inference profiles no vendor").
31
+ # The rest of this module is stricter than #1026 on purpose: for a non-Anthropic target
32
+ # whose producer is unknown or of the same family, upstream replays all the reasoning, while
33
+ # this keeps only redactedContent.
34
+ module ConverseForeignReasoning
35
+ ANTHROPIC = /(?:\A|\.)anthropic\./
36
+ OPENAI = /(?:\A|\.)openai\./
37
+ VENDOR = /\A[a-z0-9-]+(?=\.)/
38
+ APPLICATION_INFERENCE_PROFILE = ":application-inference-profile/"
39
+
40
+ private
41
+
42
+ def format_thinking_blocks(msg)
43
+ target = llm_logs_model_family(model&.id)
44
+ return super if target.nil?
45
+
46
+ producer = llm_logs_model_family(msg.model) if msg.respond_to?(:model)
47
+ return [] if producer && producer != target
48
+ return super if target == "anthropic"
49
+
50
+ blocks = msg.raw_reasoning["converse"] || msg.raw_reasoning[:converse] if msg.raw_reasoning.is_a?(Hash)
51
+ kept = Array(blocks).select { |block| llm_logs_redacted_only?(block) }
52
+ RubyLLM::Support::Utils.deep_dup(kept)
53
+ end
54
+
55
+ # "anthropic", "openai", another vendor segment, or nil when the id names no model.
56
+ def llm_logs_model_family(model_id)
57
+ return if model_id.to_s.include?(APPLICATION_INFERENCE_PROFILE)
58
+
59
+ id = foundation_model_id(model_id)
60
+ return "anthropic" if ANTHROPIC.match?(id)
61
+ return "openai" if OPENAI.match?(id)
62
+
63
+ id[VENDOR]
64
+ end
65
+
66
+ def llm_logs_redacted_only?(block)
67
+ content = block["reasoningContent"] || block[:reasoningContent] if block.is_a?(Hash)
68
+ return false unless content.is_a?(Hash)
69
+
70
+ content.keys.map(&:to_s) == ["redactedContent"]
71
+ end
72
+
73
+ def self.install!
74
+ RubyLLMPatches.prepend_checked(
75
+ RubyLLM::Protocols::Converse, self, %i[format_thinking_blocks foundation_model_id model]
76
+ )
77
+ end
78
+ end
79
+ end
80
+ end
@@ -0,0 +1,78 @@
1
+ require "json"
2
+
3
+ module LlmLogs
4
+ module RubyLLMPatches
5
+ # Reasoning effort as the reasoning_config value on Bedrock Converse.
6
+ # Backport of crmne/ruby_llm#1025 ("Send reasoning_config to models that publish it").
7
+ #
8
+ # ruby_llm 2.0.0 (Protocols::Converse::Chat#format_reasoning_fields) sends the effort of a
9
+ # model without a reasoning-budget schema as `additionalModelRequestFields:
10
+ # {reasoning_effort: "low"}`, which Bedrock rejects for OpenAI GPT models (400
11
+ # unknown_parameter). Their Converse additionalRequestFieldsSchema publishes
12
+ # `{"reasoning_config": {"type": "enum", "enum": ["none", "low", ...]}}`, and they take the
13
+ # effort as that value: `{reasoning_config: "low"}` ("none" turns reasoning off).
14
+ #
15
+ # Upstream rule, applied as #1025 does: when the model's schema, or the schema of another
16
+ # Bedrock registry entry for the same foundation model, publishes that enum, an effort that
17
+ # would have become reasoning_effort becomes the reasoning_config value, "none" included.
18
+ # Thinking off (reasoning_config disabled), budgets (explicit, or from a budget schema),
19
+ # Nova (reasoningConfig) and an empty effort are unchanged.
20
+ #
21
+ # Fallback (not in #1025): OpenAI GPT ids with no published schema get the same shape. ruby_llm
22
+ # 2.0.0's packaged registry lacks GPT-6 Sol/Luna, and an app catalog that registers them
23
+ # without Converse metadata leaves them schema-less. The match is prefix-agnostic (upstream
24
+ # REGION_PREFIXES lacks "in." in 2.0.0) and skips gpt-oss, which publishes no such enum.
25
+ module ConverseReasoningConfig
26
+ OPENAI_GPT = /(?:\A|\.)openai\.gpt-(?!oss)/
27
+
28
+ private
29
+
30
+ def format_reasoning_fields(thinking, model, max_output_tokens = nil)
31
+ return super unless thinking&.enabled? && thinking.enabled != false
32
+ return super unless llm_logs_reasoning_config?(model)
33
+
34
+ effort = thinking.effort.to_s
35
+ return super if effort.empty? || nova_model?(model)
36
+ return super if reasoning_budget(thinking, effort, model, max_output_tokens)
37
+
38
+ {reasoning_config: effort}
39
+ end
40
+
41
+ def llm_logs_reasoning_config?(model)
42
+ return false unless model
43
+
44
+ llm_logs_reasoning_config_schema?(model) || OPENAI_GPT.match?(foundation_model_id(model.id))
45
+ end
46
+
47
+ # Bedrock publishes Converse metadata for only some regional entries of a model.
48
+ def llm_logs_reasoning_config_schema?(model)
49
+ return true if llm_logs_publishes_reasoning_config?(model)
50
+
51
+ foundation_id = foundation_model_id(model.id)
52
+ RubyLLM.models.all.any? do |candidate|
53
+ candidate.provider == "bedrock" && candidate.id != model.id &&
54
+ foundation_model_id(candidate.id) == foundation_id && llm_logs_publishes_reasoning_config?(candidate)
55
+ end
56
+ end
57
+
58
+ def llm_logs_publishes_reasoning_config?(model)
59
+ metadata = RubyLLM::Support::Utils.deep_symbolize_keys(model.metadata || {})
60
+ raw_schema = metadata.dig(:converse, :additionalRequestFieldsSchema)
61
+ return false unless raw_schema.is_a?(String)
62
+
63
+ schema = JSON.parse(raw_schema, symbolize_names: true)
64
+ config = schema.is_a?(Hash) ? schema[:reasoning_config] : nil
65
+ config.is_a?(Hash) && config[:type] == "enum"
66
+ rescue JSON::ParserError
67
+ false
68
+ end
69
+
70
+ def self.install!
71
+ RubyLLMPatches.prepend_checked(
72
+ RubyLLM::Protocols::Converse, self,
73
+ %i[format_reasoning_fields foundation_model_id nova_model? reasoning_budget]
74
+ )
75
+ end
76
+ end
77
+ end
78
+ end
@@ -0,0 +1,25 @@
1
+ require "faraday"
2
+
3
+ module LlmLogs
4
+ module RubyLLMPatches
5
+ # Backport of crmne/ruby_llm#1024 ("Retry requests that fail with a TLS error").
6
+ # Retries TLS handshake failures (Faraday::SSLError, e.g. "SSL_connect ... unexpected
7
+ # eof"). faraday-net_http wraps every OpenSSL::SSL::SSLError as Faraday::SSLError, including
8
+ # read errors after the request was written, so this is the same trade-off ruby_llm
9
+ # already accepts for read timeouts. ruby_llm's retry_if (transport/connection.rb) still
10
+ # refuses non-idempotent requests and streams that already delivered content.
11
+ # Faraday::SSLError is not a Faraday::ConnectionFailed, so 2.0.0's retry list misses it.
12
+ module RetrySSLError
13
+ private
14
+
15
+ def retry_exceptions
16
+ exceptions = super
17
+ exceptions.include?(Faraday::SSLError) ? exceptions : exceptions + [Faraday::SSLError]
18
+ end
19
+
20
+ def self.install!
21
+ RubyLLMPatches.prepend_checked(RubyLLM::Transport::Connection, self, %i[retry_exceptions])
22
+ end
23
+ end
24
+ end
25
+ end
@@ -0,0 +1,54 @@
1
+ require "llm_logs/ruby_llm_patches/bedrock_sigv4"
2
+ require "llm_logs/ruby_llm_patches/converse_reasoning_config"
3
+ require "llm_logs/ruby_llm_patches/converse_foreign_reasoning"
4
+ require "llm_logs/ruby_llm_patches/converse_claude_adaptive_thinking"
5
+ require "llm_logs/ruby_llm_patches/retry_ssl_error"
6
+
7
+ module LlmLogs
8
+ # Bedrock fixes that upstream ruby_llm 2.0.0 lacks, applied with Module#prepend. Each
9
+ # backports an open upstream PR (crmne/ruby_llm#1024, #1025, #1026; see the README), so it
10
+ # can be dropped once a ruby_llm release includes that PR.
11
+ # Each was verified against 2.0.x only: on 1.x they are skipped (different internals),
12
+ # on a newer ruby_llm they still install but log a warning so we re-check whether
13
+ # upstream fixed the bug (then drop the patch) or moved the method (then the patch
14
+ # skips itself and logs why).
15
+ module RubyLLMPatches
16
+ PATCHES = [
17
+ BedrockSigV4, ConverseReasoningConfig, ConverseForeignReasoning, ConverseClaudeAdaptiveThinking, RetrySSLError
18
+ ].freeze
19
+ MINIMUM = Gem::Version.new("2.0.0")
20
+ VERIFIED = Gem::Requirement.new(">= 2.0.0", "< 2.1")
21
+
22
+ module_function
23
+
24
+ def install!
25
+ version = Gem::Version.new(RubyLLM::VERSION)
26
+ return [] if version < MINIMUM
27
+
28
+ unless VERIFIED.satisfied_by?(version)
29
+ warn_log("verified against ruby_llm #{VERIFIED}, running #{version}: check whether upstream " \
30
+ "now fixes #{PATCHES.map(&:name).join(', ')} and drop the patches it does")
31
+ end
32
+ PATCHES.select(&:install!)
33
+ end
34
+
35
+ def warn_log(message)
36
+ logger = defined?(Rails) && Rails.respond_to?(:logger) && Rails.logger
37
+ logger ? logger.warn("[llm_logs] #{message}") : Kernel.warn("[llm_logs] #{message}")
38
+ end
39
+
40
+ # Prepends +patch+ to +target+ when +target+ still defines every method in +methods+.
41
+ def prepend_checked(target, patch, methods)
42
+ return true if target <= patch
43
+
44
+ missing = methods.reject { |name| target.method_defined?(name) || target.private_method_defined?(name) }
45
+ if missing.any?
46
+ warn_log("#{patch.name} not installed: #{target} no longer defines #{missing.join(', ')}")
47
+ return false
48
+ end
49
+
50
+ target.prepend(patch)
51
+ true
52
+ end
53
+ end
54
+ end
@@ -1,3 +1,3 @@
1
1
  module LlmLogs
2
- VERSION = "0.5.1"
2
+ VERSION = "0.6.0"
3
3
  end
data/lib/llm_logs.rb CHANGED
@@ -54,10 +54,6 @@ module LlmLogs
54
54
  configuration.batch_enabled
55
55
  end
56
56
 
57
- def self.batch_provider
58
- configuration.batch_provider
59
- end
60
-
61
57
  def self.bedrock_batch
62
58
  configuration.bedrock_batch
63
59
  end
@@ -71,7 +67,7 @@ module LlmLogs
71
67
  end
72
68
 
73
69
  def self.batch_adapters
74
- @batch_adapters ||= {openai_responses: LlmLogs::Batch::Adapters::OpenaiResponses.new}
70
+ @batch_adapters ||= {}
75
71
  end
76
72
 
77
73
  def self.register_batch_adapter(provider, adapter)
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: llm_logs
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.5.1
4
+ version: 0.6.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Anton
@@ -97,30 +97,36 @@ dependencies:
97
97
  name: ruby_llm
98
98
  requirement: !ruby/object:Gem::Requirement
99
99
  requirements:
100
- - - "~>"
100
+ - - ">="
101
101
  - !ruby/object:Gem::Version
102
- version: '1.16'
103
- type: :development
102
+ version: '2.0'
103
+ - - "<"
104
+ - !ruby/object:Gem::Version
105
+ version: '3'
106
+ type: :runtime
104
107
  prerelease: false
105
108
  version_requirements: !ruby/object:Gem::Requirement
106
109
  requirements:
107
- - - "~>"
110
+ - - ">="
111
+ - !ruby/object:Gem::Version
112
+ version: '2.0'
113
+ - - "<"
108
114
  - !ruby/object:Gem::Version
109
- version: '1.16'
115
+ version: '3'
110
116
  - !ruby/object:Gem::Dependency
111
- name: ruby_llm-responses_api
117
+ name: aws-eventstream
112
118
  requirement: !ruby/object:Gem::Requirement
113
119
  requirements:
114
- - - "~>"
120
+ - - ">="
115
121
  - !ruby/object:Gem::Version
116
- version: '0.6'
122
+ version: '0'
117
123
  type: :development
118
124
  prerelease: false
119
125
  version_requirements: !ruby/object:Gem::Requirement
120
126
  requirements:
121
- - - "~>"
127
+ - - ">="
122
128
  - !ruby/object:Gem::Version
123
- version: '0.6'
129
+ version: '0'
124
130
  - !ruby/object:Gem::Dependency
125
131
  name: webmock
126
132
  requirement: !ruby/object:Gem::Requirement
@@ -186,10 +192,8 @@ files:
186
192
  - app/models/llm_logs/application_record.rb
187
193
  - app/models/llm_logs/batch.rb
188
194
  - app/models/llm_logs/batch/adapters/bedrock.rb
189
- - app/models/llm_logs/batch/adapters/openai_responses.rb
190
195
  - app/models/llm_logs/batch/handler_registry.rb
191
196
  - app/models/llm_logs/batch/reconciler.rb
192
- - app/models/llm_logs/batch/schema_format.rb
193
197
  - app/models/llm_logs/batch/submitter.rb
194
198
  - app/models/llm_logs/batch/trace_recorder.rb
195
199
  - app/models/llm_logs/batch_request.rb
@@ -237,6 +241,12 @@ files:
237
241
  - lib/llm_logs/engine.rb
238
242
  - lib/llm_logs/instrumentation/ruby_llm_chat.rb
239
243
  - lib/llm_logs/prompt_renderer.rb
244
+ - lib/llm_logs/ruby_llm_patches.rb
245
+ - lib/llm_logs/ruby_llm_patches/bedrock_sigv4.rb
246
+ - lib/llm_logs/ruby_llm_patches/converse_claude_adaptive_thinking.rb
247
+ - lib/llm_logs/ruby_llm_patches/converse_foreign_reasoning.rb
248
+ - lib/llm_logs/ruby_llm_patches/converse_reasoning_config.rb
249
+ - lib/llm_logs/ruby_llm_patches/retry_ssl_error.rb
240
250
  - lib/llm_logs/tracer.rb
241
251
  - lib/llm_logs/version.rb
242
252
  - lib/tasks/llm_logs.rake
@@ -261,7 +271,7 @@ required_rubygems_version: !ruby/object:Gem::Requirement
261
271
  - !ruby/object:Gem::Version
262
272
  version: '0'
263
273
  requirements: []
264
- rubygems_version: 4.0.20
274
+ rubygems_version: 3.7.2
265
275
  specification_version: 4
266
276
  summary: Rails engine for LLM logging and prompt management
267
277
  test_files: []
@@ -1,63 +0,0 @@
1
- module LlmLogs
2
- class Batch
3
- module Adapters
4
- # Wraps ruby_llm-responses_api's OpenAI Batch API. Contains the remote-API logic
5
- # previously inline in Submitter/Reconciler; the DB/claim lifecycle stays in those.
6
- class OpenaiResponses
7
- PROVIDER = :openai_responses
8
-
9
- def submit(_batch, requests)
10
- rubyllm_batch = RubyLLM.batch(model: requests.first.model, provider: PROVIDER)
11
- requests.each do |request|
12
- payload = request.payload
13
- rubyllm_batch.add(
14
- payload["input"],
15
- id: request.custom_id,
16
- instructions: payload["instructions"],
17
- temperature: payload["temperature"],
18
- **reasoning_extra(payload["reasoning_effort"]),
19
- **schema_extra(payload["schema"])
20
- )
21
- end
22
- rubyllm_batch.create!
23
- {provider_batch_id: rubyllm_batch.id, openai_batch_id: rubyllm_batch.id, provider_metadata: {}}
24
- end
25
-
26
- def terminal_status(batch)
27
- resume(batch).status
28
- end
29
-
30
- def results(batch)
31
- resume(batch).results
32
- end
33
-
34
- def error_ids(batch)
35
- resume(batch).errors.filter_map { |e| e["custom_id"] }
36
- end
37
-
38
- private
39
-
40
- # The Responses API nests effort under `reasoning:`; `reasoning_effort` is the
41
- # Chat Completions spelling and is silently ignored here.
42
- def reasoning_extra(effort)
43
- return {} if effort.blank?
44
-
45
- {reasoning: {effort: effort}}
46
- end
47
-
48
- # No memoization: LlmLogs.batch_adapters holds one shared instance of this adapter
49
- # for the life of the process, so caching the resumed handle here would go stale
50
- # for a long-running PollJob worker (in-progress batches would never re-fetch a
51
- # later terminal status) and would grow unbounded. Build a fresh handle every call.
52
- def resume(batch)
53
- RubyLLM.batch(id: batch.provider_batch_id, provider: PROVIDER)
54
- end
55
-
56
- def schema_extra(schema)
57
- format = SchemaFormat.call(schema)
58
- format ? {text: format} : {}
59
- end
60
- end
61
- end
62
- end
63
- end
@@ -1,25 +0,0 @@
1
- module LlmLogs
2
- class Batch
3
- # Translates a {name:, schema:, strict:} schema (the shape RubyLLM::Chat#with_schema
4
- # produces) into the OpenAI Responses API `text.format` block. The batch path builds
5
- # request bodies directly via RubyLLM.batch#add(**extra), bypassing with_schema, so
6
- # we must hand the json_schema block in ourselves.
7
- module SchemaFormat
8
- module_function
9
-
10
- def call(schema)
11
- return nil if schema.nil?
12
-
13
- schema = schema.symbolize_keys if schema.respond_to?(:symbolize_keys)
14
- {
15
- format: {
16
- type: "json_schema",
17
- name: schema[:name] || "response",
18
- schema: schema[:schema] || schema,
19
- strict: schema.key?(:strict) ? schema[:strict] : true
20
- }
21
- }
22
- end
23
- end
24
- end
25
- end