railwatch 0.1.2 → 0.1.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +47 -0
- data/docs/records.md +72 -0
- data/lib/railwatch/configuration.rb +8 -2
- data/lib/railwatch/execution.rb +1 -1
- data/lib/railwatch/record.rb +1 -1
- data/lib/railwatch/subscribers/llm.rb +335 -0
- data/lib/railwatch/subscribers.rb +2 -1
- data/lib/railwatch/version.rb +1 -1
- data/lib/railwatch.rb +1 -1
- metadata +3 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 0b78b98999447f64098f912cdfdcbb6202c9331ac4e07988fb350080d084f79c
|
|
4
|
+
data.tar.gz: e230f37a0a728c381fc4b84849bc0db762ad14e2b9c37aa5802850b97cba7271
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 27c7659440a7e1592b99d10ad2a409c711a2b5a44cafee9a4707ff8672e0952fcf6819ad5b98c96f4e9d25ae6e1c5b86193ebb70a74b52904f8aee808977c72d
|
|
7
|
+
data.tar.gz: c1101b7933f5703f4dfd7469b07bf9ac3f39cfaee7bd2a9d451e30126921c33a0326c5976d8605e7dce39f13b10b8d65fd8052007051775bc9e46f0b5cbb6368
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,52 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.1.4 (2026-09-15)
|
|
4
|
+
|
|
5
|
+
- `llm_call` records what the call carried and how it was configured, not
|
|
6
|
+
just what it cost. `finish_reason` shows when an answer was cut off
|
|
7
|
+
(`max_tokens`) or filtered, which previously read exactly like a complete
|
|
8
|
+
one. `attachments` and `attachment_types` show that a call carried two
|
|
9
|
+
images and a PDF, so a document read is no longer indistinguishable from
|
|
10
|
+
an expensive prompt -- on a document-reading call the attachments are most
|
|
11
|
+
of the input tokens. `params` holds the settings that produced the answer
|
|
12
|
+
(temperature, max_output_tokens, tool_choice, thinking, caching, whether a
|
|
13
|
+
schema was used, and the per-operation ones) so a surprising result can be
|
|
14
|
+
reproduced. `tools` names what the model could reach; `tool_call_id` joins
|
|
15
|
+
a tool call to the turn that asked for it.
|
|
16
|
+
- `provider_request_id` is read from the response headers (`request-id`,
|
|
17
|
+
`x-request-id`, `x-amzn-requestid`). It is the only key that joins a
|
|
18
|
+
Railwatch record to the provider's own record of the same call, and it is
|
|
19
|
+
what a provider support ticket asks for.
|
|
20
|
+
- `cost_reported` distinguishes a price the provider stated from one
|
|
21
|
+
estimated against the model registry.
|
|
22
|
+
- `provider_options` is filtered twice before it is stored: through the
|
|
23
|
+
app's own parameter filter, and again against the credential-name matcher
|
|
24
|
+
that catches `X-Api-Key` on a header. The default parameter filter is
|
|
25
|
+
password-shaped, so an `api_key` passed per call went through it
|
|
26
|
+
untouched. Attachment filenames stay behind `capture_llm_content`, since
|
|
27
|
+
a filename is business data rather than metadata.
|
|
28
|
+
|
|
29
|
+
## 0.1.3 (2026-09-15)
|
|
30
|
+
|
|
31
|
+
- LLM calls are recorded from RubyLLM's own instrumentation. Every model
|
|
32
|
+
call it emits -- `chat`, `compaction`, `embedding`, `image`, `speech`,
|
|
33
|
+
`transcription`, `moderation`, `rerank`, `ocr` -- plus each tool
|
|
34
|
+
invocation becomes one
|
|
35
|
+
`llm_call` child record on the request, job, or command that made it,
|
|
36
|
+
carrying provider, model, duration, token counts per bucket, and cost.
|
|
37
|
+
Nothing is patched: RubyLLM publishes ActiveSupport::Notifications events
|
|
38
|
+
and Railwatch subscribes to them like any Rails event.
|
|
39
|
+
- Both RubyLLM generations are read from the same subscriber. 1.16 puts
|
|
40
|
+
token counts on the event as scalars and reports no cost; 2.0 sends its
|
|
41
|
+
`Tokens` and `Cost` objects, and stamps `workflow_id` and step identity on
|
|
42
|
+
every event inside `RubyLLM.workflow`, which is recorded so an agent run
|
|
43
|
+
can be reassembled from its steps. A 1.16 app has no cost rather than a
|
|
44
|
+
cost of zero, and a model the registry cannot price is unpriced, not free.
|
|
45
|
+
- `capture_llm_content` (default off, `RAILWATCH_CAPTURE_LLM_CONTENT`)
|
|
46
|
+
records the last user turn and the reply, capped at 4 KiB of bytes each. Token
|
|
47
|
+
counts, model, and cost are always captured; prompts are not, because
|
|
48
|
+
they are whatever the app sent a provider.
|
|
49
|
+
|
|
3
50
|
## 0.1.2 (2026-09-14)
|
|
4
51
|
|
|
5
52
|
- A failed job's exception is reported once. Solid Queue re-raises it out
|
data/docs/records.md
CHANGED
|
@@ -502,6 +502,78 @@ never double-recorded.
|
|
|
502
502
|
| `response_body` | First 4 KiB of the response body, but only when `config.capture_response_body_on_error` is on (off by default) *and* the response was an error. A body that parses as a JSON object is run through the same parameter filter as request params and re-serialized; anything else is stored as it arrived. nil in every other case — including a connection failure, where there is no response (on the Net::HTTP path a body is read only if Net::HTTP already buffered it, so a response being streamed through `read_body` is never consumed; on the Faraday path the body is taken only once a status came back, so an outgoing request payload can never be filed as a response). |
|
|
503
503
|
| `source` | App-code call site (Net::HTTP path only). |
|
|
504
504
|
|
|
505
|
+
### `llm_call`
|
|
506
|
+
|
|
507
|
+
Every RubyLLM model call and tool invocation, from
|
|
508
|
+
`lib/railwatch/subscribers/llm.rb`. RubyLLM publishes its own
|
|
509
|
+
`ActiveSupport::Notifications` events, so nothing is patched and RubyLLM is
|
|
510
|
+
not a dependency — an app without it never emits these. Requires RubyLLM
|
|
511
|
+
1.16 or later, which is where its instrumentation landed.
|
|
512
|
+
|
|
513
|
+
The model call also appears as an `outgoing_request`, since it is an HTTP
|
|
514
|
+
call like any other. The two are different grains on purpose: the
|
|
515
|
+
`outgoing_request` is the HTTP truth, the `llm_call` is what it cost. That
|
|
516
|
+
difference is useful: RubyLLM retries through Faraday, so one `llm_call`
|
|
517
|
+
with several `outgoing_request` rows against it in the same execution is a
|
|
518
|
+
call that was retried. Over a window, `outgoing_requests - llm_calls` to the
|
|
519
|
+
same provider host is the number of *extra attempts*, not a rate -- the
|
|
520
|
+
share of calls that were retried needs counting the calls with more than one
|
|
521
|
+
request against them, which the execution id supports.
|
|
522
|
+
|
|
523
|
+
**Token counts and cost differ by RubyLLM version.** 1.16 reports token
|
|
524
|
+
counts and no cost at all. 2.0 reports both, from its usage ledger, and
|
|
525
|
+
adds the `workflow_*` fields. `cost_nanos` is null rather than zero
|
|
526
|
+
whenever RubyLLM reported no cost or the model registry could not price
|
|
527
|
+
it — an unpriced call is not a free one.
|
|
528
|
+
|
|
529
|
+
**Concurrent tool calls are not recorded.** RubyLLM's opt-in
|
|
530
|
+
`tool_concurrency` (`:threads` or `:fibers`) runs each tool in a fresh
|
|
531
|
+
thread or fiber. `Railwatch::Current` is backed by
|
|
532
|
+
`ActiveSupport::IsolatedExecutionState`, which a new thread does not
|
|
533
|
+
inherit, so the `tool_call.ruby_llm` event fires with no execution to
|
|
534
|
+
attach to and the record is dropped rather than misattributed. This
|
|
535
|
+
affects every Railwatch subscriber in an app-spawned thread, not just this
|
|
536
|
+
one. Tool concurrency is off by default; with it off, tool calls are
|
|
537
|
+
recorded normally. The model calls themselves are unaffected either way,
|
|
538
|
+
so cost is always complete.
|
|
539
|
+
|
|
540
|
+
| Field | Meaning |
|
|
541
|
+
|---|---|
|
|
542
|
+
| `group` | Hash of provider + model + operation, or of `"tool"` + tool name. |
|
|
543
|
+
| `operation` | `"chat"`, `"compaction"`, `"embedding"`, `"image"`, `"speech"`, `"transcription"`, `"moderation"`, `"rerank"`, `"ocr"`, or `"tool"`. |
|
|
544
|
+
| `provider` | Provider slug, e.g. `"anthropic"`. |
|
|
545
|
+
| `model` | Model the call was made with. Empty for a provider that selects its own (moderation). |
|
|
546
|
+
| `response_model` | Model the provider says answered, which can differ from the one asked for. |
|
|
547
|
+
| `tool_name` | Tool name, for `operation: "tool"`. |
|
|
548
|
+
| `duration` | Microseconds. |
|
|
549
|
+
| `status` | `"ok"`, or `"failed"` if the call raised. |
|
|
550
|
+
| `error` | `"Class: message"`, truncated to 255 chars, if the call raised. |
|
|
551
|
+
| `streaming` | Whether the call was streamed. |
|
|
552
|
+
| `message_count` | Conversation length at the time of the call. |
|
|
553
|
+
| `tool_count` | Number of tools the model was offered. |
|
|
554
|
+
| `input_tokens` | Standard (non-cached) input tokens. |
|
|
555
|
+
| `output_tokens` | Billable output tokens. |
|
|
556
|
+
| `cache_read_tokens` | Tokens served from the provider's prompt cache. |
|
|
557
|
+
| `cache_write_tokens` | Tokens written to the provider's prompt cache. |
|
|
558
|
+
| `thinking_tokens` | Reasoning tokens, where the provider reports them separately. |
|
|
559
|
+
| `cost_nanos` | Cost in billionths of a US dollar. Null when unpriced — see above. Nanodollars because a cheap call is well under a microdollar and floats do not sum to an invoice. |
|
|
560
|
+
| `workflow_id` | `RubyLLM.workflow` identifier (2.0+). Null outside a workflow. |
|
|
561
|
+
| `workflow_name` | Workflow name (2.0+). |
|
|
562
|
+
| `workflow_step_id` | Step identifier within the workflow (2.0+). |
|
|
563
|
+
| `workflow_step_name` | Step name (2.0+). |
|
|
564
|
+
| `workflow_step_parent_id` | Enclosing step, for nested steps — what reconstructs the tree (2.0+). |
|
|
565
|
+
| `finish_reason` | Why the model stopped: `stop`, `max_tokens`, `tool_calls`, `content_filter`, or whatever the provider spelled it. `max_tokens` means the answer was cut off -- without this a truncated extraction reads exactly like a complete one. |
|
|
566
|
+
| `provider_request_id` | The provider's own id for the request, read from the response headers (`request-id`, `x-request-id`, `x-amzn-requestid`). The only key that joins this record to the provider's side of it, and what a support ticket asks for. |
|
|
567
|
+
| `tools` | Comma-separated names of the tools the model could reach, first 50. `tool_count` says how many; retracing needs which. |
|
|
568
|
+
| `cost_reported` | Whether the provider priced the call itself, or the amount is an estimate from the model registry. Null on gems or operations that report no cost. |
|
|
569
|
+
| `attachments` | How many files the last user turn carried. Absent when it carried none. Only the last turn is measured: earlier turns were counted by the calls that sent them. |
|
|
570
|
+
| `attachment_types` | What they were, by category and count, e.g. `imagex2,pdf`. Categories are RubyLLM's: image, pdf, audio, video, text, document, unknown. On a document-reading call the attachments are most of the input tokens, so without this an expensive scan is indistinguishable from an expensive prompt. |
|
|
571
|
+
| `attachment_names` | Filenames, only when `config.capture_llm_content` is on. A filename like `ACME_invoice_88231.pdf` is business data, not metadata, so it follows the same switch as prompts. |
|
|
572
|
+
| `params` | JSON of the settings that produced the answer, so a surprising one can be reproduced: `temperature`, `max_output_tokens`, `tool_choice`, `tool_call_limit`, `thinking`, `caching`, `citations`, whether a `schema` was used, plus the per-operation ones (`dimensions`, `task_type`, `size`, `count`, `voice`, `format`, `language`, `pages`, `document_count`, `top_n`), `server_tools` and the provider's `server_tool_use` counters. `provider_options` is included, filtered twice: through the app's own parameter filter, and again against the credential-name matcher that catches `X-Api-Key` on a header -- an `api_key` passed per call sails straight through a password-shaped filter. For `operation: "tool"` this holds the tool result's class instead. |
|
|
573
|
+
| `tool_call_id` | The provider's id for a tool invocation, for joining a tool call to the assistant turn that asked for it. |
|
|
574
|
+
| `prompt` | Last user turn, only when `config.capture_llm_content` is on (off by default). Capped at 4 KiB of bytes. |
|
|
575
|
+
| `completion` | The reply, same condition and cap. For `operation: "tool"` these two hold the tool's arguments and result instead. |
|
|
576
|
+
|
|
505
577
|
### `storage_op`
|
|
506
578
|
|
|
507
579
|
Every Active Storage service operation. See
|
|
@@ -5,7 +5,8 @@ module Railwatch
|
|
|
5
5
|
# Laravel Nightwatch's config so the two products document the same knobs.
|
|
6
6
|
class Configuration
|
|
7
7
|
RECORD_TYPES = %i[queries cache_events mail broadcasts notifications outgoing_requests
|
|
8
|
-
storage_ops view_renders logs transactions deprecations sessions
|
|
8
|
+
storage_ops view_renders logs transactions deprecations sessions
|
|
9
|
+
llm_calls].freeze
|
|
9
10
|
|
|
10
11
|
# Framework/vendor noise excluded by default so a fresh install isn't
|
|
11
12
|
# dominated by Rails' own housekeeping. Both lists are opt-in to disable
|
|
@@ -78,7 +79,8 @@ module Railwatch
|
|
|
78
79
|
:profile_sample, :profile_slow_ms, :profile_interval_us, :profiler,
|
|
79
80
|
:capture_job_arguments, :capture_job_retry_errors, :capture_response_body_on_error, :max_attachment_bytes,
|
|
80
81
|
:track_sessions, :session_flush_interval, :session_timeout,
|
|
81
|
-
:capture_console, :interactive_runner_paths, :ignored_request_paths
|
|
82
|
+
:capture_console, :interactive_runner_paths, :ignored_request_paths,
|
|
83
|
+
:capture_llm_content
|
|
82
84
|
|
|
83
85
|
attr_reader :deploy, :deploy_source, :detect_deploy, :user_resolver, :beacon_user_resolver,
|
|
84
86
|
:fingerprint_resolver, :redactors, :rejectors, :before_ingest, :backpressure_high_water
|
|
@@ -175,6 +177,10 @@ module Railwatch
|
|
|
175
177
|
@capture_job_arguments = env_bool("RAILWATCH_CAPTURE_JOB_ARGUMENTS", false)
|
|
176
178
|
@capture_job_retry_errors = env_bool("RAILWATCH_CAPTURE_JOB_RETRY_ERRORS", false)
|
|
177
179
|
@capture_response_body_on_error = env_bool("RAILWATCH_CAPTURE_RESPONSE_BODY_ON_ERROR", false)
|
|
180
|
+
# Prompts and completions are whatever the app sent a provider, so
|
|
181
|
+
# they are off until an operator opts in. Token counts, model, and
|
|
182
|
+
# cost -- the reason the record exists -- are always captured.
|
|
183
|
+
@capture_llm_content = env_bool("RAILWATCH_CAPTURE_LLM_CONTENT", false)
|
|
178
184
|
@max_attachment_bytes = env_int("RAILWATCH_MAX_ATTACHMENT_BYTES", 1_048_576)
|
|
179
185
|
# Release health: one `session` record per browser tab (the beacon
|
|
180
186
|
# client) and per authenticated/cookied server session (Railwatch::Sessions).
|
data/lib/railwatch/execution.rb
CHANGED
|
@@ -12,7 +12,7 @@ module Railwatch
|
|
|
12
12
|
MAX_RECORDS = 10_000
|
|
13
13
|
COUNTERS = %i[queries cached_queries exceptions logs cache_events jobs_enqueued mail
|
|
14
14
|
broadcasts notifications outgoing_requests storage_ops view_renders
|
|
15
|
-
transactions hydrated_models lazy_loads deprecations spans].freeze
|
|
15
|
+
transactions hydrated_models lazy_loads deprecations spans llm_calls].freeze
|
|
16
16
|
# GC.stat with no key builds the whole stat hash; whether this Ruby
|
|
17
17
|
# reports GC time never changes, so ask once.
|
|
18
18
|
GC_TIME_SUPPORTED = GC.stat.key?(:time)
|
data/lib/railwatch/record.rb
CHANGED
|
@@ -10,7 +10,7 @@ module Railwatch
|
|
|
10
10
|
query: 1, n_plus_one: 1, transaction: 1, exception: 1, cache_event: 1, mail: 1,
|
|
11
11
|
broadcast: 1, notification: 1, outgoing_request: 1, storage_op: 1, view_render: 1,
|
|
12
12
|
log: 1, enqueued_job: 1, user: 1, deprecation: 1, visit: 1, process: 1, span: 1, health: 1,
|
|
13
|
-
profile: 1, attachment: 1, session: 1
|
|
13
|
+
profile: 1, attachment: 1, session: 1, llm_call: 1
|
|
14
14
|
}.freeze
|
|
15
15
|
|
|
16
16
|
# Used in place of an execution's envelope when there is no execution, so
|
|
@@ -0,0 +1,335 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Railwatch
|
|
4
|
+
module Subscribers
|
|
5
|
+
# RubyLLM's own instrumentation. It emits an ActiveSupport::Notifications
|
|
6
|
+
# event per model call -- chat, embedding, image, and the rest -- plus one
|
|
7
|
+
# per tool invocation, so nothing here patches RubyLLM; we subscribe the
|
|
8
|
+
# same way we subscribe to Rails.
|
|
9
|
+
#
|
|
10
|
+
# One `llm_call` record per event, with `operation` naming which kind it
|
|
11
|
+
# was. Tool calls are the same record with operation "tool": they sit in
|
|
12
|
+
# the same execution waterfall and carry no tokens or cost.
|
|
13
|
+
#
|
|
14
|
+
# Two RubyLLM generations are supported. 1.16 puts token counts on the
|
|
15
|
+
# event as scalars and reports no cost at all; 2.0 sends Tokens and Cost
|
|
16
|
+
# objects built from its usage ledger. Both normalise to the same wire
|
|
17
|
+
# record, so a 1.16 app simply has no cost. Subscribing to an event
|
|
18
|
+
# RubyLLM never emits costs nothing, so the operations 2.0 added are
|
|
19
|
+
# subscribed unconditionally rather than behind a version check.
|
|
20
|
+
module Llm
|
|
21
|
+
extend Base
|
|
22
|
+
|
|
23
|
+
module_function
|
|
24
|
+
|
|
25
|
+
# Every usage-bearing operation in RubyLLM 2.0. The ones 1.16 knows
|
|
26
|
+
# about (chat, embedding, image, transcription, moderation) emit the
|
|
27
|
+
# same event names, so this list needs no version branch. `compaction`
|
|
28
|
+
# is a chat call by another name -- same payload, same tokens, same
|
|
29
|
+
# cost -- and is billed, so it belongs here rather than being invisible
|
|
30
|
+
# spend.
|
|
31
|
+
OPERATIONS = %w[chat compaction embedding image speech transcription
|
|
32
|
+
moderation rerank ocr].freeze
|
|
33
|
+
|
|
34
|
+
# Costs are fractions of a cent: a cheap model's call is well under a
|
|
35
|
+
# microdollar, and floats summed across a month of rollups do not add
|
|
36
|
+
# up to an invoice. Nanodollars keep it exact in an integer column
|
|
37
|
+
# ($1,000 is 1e12, comfortably inside i64).
|
|
38
|
+
NANOS_PER_DOLLAR = 1_000_000_000
|
|
39
|
+
|
|
40
|
+
# Matches capture_response_body_on_error's cap. Prompts and completions
|
|
41
|
+
# are free text an app controls, so this is a size bound, not redaction.
|
|
42
|
+
CONTENT_MAX = 4096
|
|
43
|
+
|
|
44
|
+
def install!(_app)
|
|
45
|
+
OPERATIONS.each do |operation|
|
|
46
|
+
subscribe("#{operation}.ruby_llm") { |event| record_call(operation, event) }
|
|
47
|
+
end
|
|
48
|
+
subscribe("tool_call.ruby_llm") { |event| record_tool(event) }
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
def record_call(operation, event)
|
|
52
|
+
exe = execution
|
|
53
|
+
exe&.count(:llm_calls)
|
|
54
|
+
return unless recording?
|
|
55
|
+
|
|
56
|
+
p = event.payload
|
|
57
|
+
provider = p[:provider].to_s
|
|
58
|
+
model = p[:model].to_s
|
|
59
|
+
Railwatch.record(:llm_call,
|
|
60
|
+
group: Record.group_hash(provider, model, operation),
|
|
61
|
+
timestamp: started_at(event),
|
|
62
|
+
operation: operation,
|
|
63
|
+
provider: provider,
|
|
64
|
+
model: model,
|
|
65
|
+
response_model: p[:response_model]&.to_s&.slice(0, 255),
|
|
66
|
+
duration: micros(event),
|
|
67
|
+
streaming: p[:streaming] == true,
|
|
68
|
+
message_count: p[:message_count],
|
|
69
|
+
tool_count: Array(p[:tools]).size,
|
|
70
|
+
tools: tool_names(p),
|
|
71
|
+
cost_nanos: cost_nanos(p),
|
|
72
|
+
cost_reported: cost_reported(p),
|
|
73
|
+
finish_reason: finish_reason(p),
|
|
74
|
+
provider_request_id: provider_request_id(p),
|
|
75
|
+
params: params(operation, p),
|
|
76
|
+
**attachments(p),
|
|
77
|
+
**tokens(p),
|
|
78
|
+
**workflow(p),
|
|
79
|
+
**outcome(p),
|
|
80
|
+
prompt: content(prompt_text(p)),
|
|
81
|
+
completion: content(message_text(p[:response])))
|
|
82
|
+
end
|
|
83
|
+
|
|
84
|
+
# RubyLLM's opt-in tool_concurrency (:threads or :fibers) runs each
|
|
85
|
+
# tool in a fresh thread or fiber, and Current is backed by
|
|
86
|
+
# IsolatedExecutionState, which a new thread does not inherit. So this
|
|
87
|
+
# fires with no execution and the record is dropped.
|
|
88
|
+
#
|
|
89
|
+
# It cannot be fixed from here: by the time the event is delivered we
|
|
90
|
+
# are already inside the worker, with no reference to the execution
|
|
91
|
+
# that spawned it, and the thread is RubyLLM's to create
|
|
92
|
+
# (chat/tool_concurrency.rb propagates its own workflow context across
|
|
93
|
+
# that boundary, but knows nothing of ours). The same is true of every
|
|
94
|
+
# subscriber in an app-spawned thread. Dropping beats guessing: a
|
|
95
|
+
# process-wide fallback would file one request's tool call under
|
|
96
|
+
# another's execution. Concurrency is off by default, and the model
|
|
97
|
+
# calls are unaffected either way, so cost stays complete.
|
|
98
|
+
def record_tool(event)
|
|
99
|
+
exe = execution
|
|
100
|
+
exe&.count(:llm_calls)
|
|
101
|
+
return unless recording?
|
|
102
|
+
|
|
103
|
+
p = event.payload
|
|
104
|
+
tool_name = p[:tool_name].to_s
|
|
105
|
+
Railwatch.record(:llm_call,
|
|
106
|
+
group: Record.group_hash("tool", tool_name),
|
|
107
|
+
timestamp: started_at(event),
|
|
108
|
+
operation: "tool",
|
|
109
|
+
provider: p[:provider].to_s,
|
|
110
|
+
model: p[:model].to_s,
|
|
111
|
+
tool_name: tool_name[0, 255],
|
|
112
|
+
tool_call_id: p[:tool_call_id]&.to_s&.slice(0, 128),
|
|
113
|
+
params: ({ result_class: p[:result_class].to_s[0, 128] } if p[:result_class]),
|
|
114
|
+
duration: micros(event),
|
|
115
|
+
**workflow(p),
|
|
116
|
+
**outcome(p),
|
|
117
|
+
prompt: content(p[:tool_arguments]),
|
|
118
|
+
completion: content(p[:result_content]))
|
|
119
|
+
end
|
|
120
|
+
|
|
121
|
+
# Why the model stopped. :max_tokens means the answer was cut off --
|
|
122
|
+
# a truncated extraction reads exactly like a complete one without
|
|
123
|
+
# this, which is the failure most worth being able to see.
|
|
124
|
+
def finish_reason(payload)
|
|
125
|
+
response = payload[:response]
|
|
126
|
+
return nil unless response.respond_to?(:finish_reason)
|
|
127
|
+
|
|
128
|
+
response.finish_reason&.to_s&.slice(0, 32)
|
|
129
|
+
end
|
|
130
|
+
|
|
131
|
+
# Message#raw is the Faraday response (protocol.rb hands it in), so the
|
|
132
|
+
# provider's own request id is in its headers. It is what a provider
|
|
133
|
+
# support ticket asks for, and the only key that joins our record to
|
|
134
|
+
# theirs.
|
|
135
|
+
REQUEST_ID_HEADERS = %w[request-id x-request-id x-amzn-requestid].freeze
|
|
136
|
+
|
|
137
|
+
def provider_request_id(payload)
|
|
138
|
+
raw = payload[:response]
|
|
139
|
+
raw = raw.raw if raw.respond_to?(:raw)
|
|
140
|
+
headers = raw.respond_to?(:headers) ? raw.headers : nil
|
|
141
|
+
return nil unless headers.respond_to?(:[])
|
|
142
|
+
|
|
143
|
+
REQUEST_ID_HEADERS.each do |name|
|
|
144
|
+
value = headers[name]
|
|
145
|
+
return value.to_s[0, 128] if value.present?
|
|
146
|
+
end
|
|
147
|
+
nil
|
|
148
|
+
end
|
|
149
|
+
|
|
150
|
+
# Which tools the model could reach on this call. tool_count alone says
|
|
151
|
+
# how many; retracing needs which.
|
|
152
|
+
def tool_names(payload)
|
|
153
|
+
names = Array(payload[:tools]).map(&:to_s)
|
|
154
|
+
names.empty? ? nil : names.first(50).join(",")[0, 1024]
|
|
155
|
+
end
|
|
156
|
+
|
|
157
|
+
# Whether the provider priced the call itself, or we estimated it from
|
|
158
|
+
# the registry. The difference matters when a total is queried against
|
|
159
|
+
# an invoice.
|
|
160
|
+
def cost_reported(payload)
|
|
161
|
+
tokens = payload[:tokens]
|
|
162
|
+
return nil unless tokens.respond_to?(:reported_cost)
|
|
163
|
+
# No cost means no provenance to report. false would claim the
|
|
164
|
+
# registry priced it, which is the same false certainty cost_nanos
|
|
165
|
+
# avoids by being nil rather than zero.
|
|
166
|
+
return nil if cost_nanos(payload).nil?
|
|
167
|
+
|
|
168
|
+
!tokens.reported_cost.nil?
|
|
169
|
+
end
|
|
170
|
+
|
|
171
|
+
# What the call carried besides text. Images and PDFs are most of the
|
|
172
|
+
# input tokens on a document-reading call, and without this an
|
|
173
|
+
# expensive scan is indistinguishable from an expensive prompt.
|
|
174
|
+
# Only the last user turn is measured: earlier turns were counted by
|
|
175
|
+
# the calls that sent them, and walking the whole history would both
|
|
176
|
+
# double-count and cost O(messages) on every call.
|
|
177
|
+
def attachments(payload)
|
|
178
|
+
message = last_user_message(payload)
|
|
179
|
+
list = message.respond_to?(:attachments) ? Array(message.attachments) : []
|
|
180
|
+
return {} if list.empty?
|
|
181
|
+
|
|
182
|
+
types = list.filter_map { |a| a.type.to_s if a.respond_to?(:type) }.tally
|
|
183
|
+
.sort_by { |_, n| -n }.map { |type, n| n > 1 ? "#{type}x#{n}" : type }.join(",")
|
|
184
|
+
{ attachments: list.size, attachment_types: types[0, 128],
|
|
185
|
+
attachment_names: content(list.filter_map { |a| a.filename if a.respond_to?(:filename) }.join(", ")) }
|
|
186
|
+
end
|
|
187
|
+
|
|
188
|
+
# The knobs that change what a call costs and what it returns, so a
|
|
189
|
+
# surprising result can be reproduced with the settings that produced
|
|
190
|
+
# it. Provider options go through the app's own parameter filter: they
|
|
191
|
+
# are request configuration, but an app can put anything in them.
|
|
192
|
+
COMMON_PARAMS = %i[temperature max_output_tokens tool_choice tool_call_limit
|
|
193
|
+
thinking caching citations dimensions task_type size count
|
|
194
|
+
voice format language pages document_count top_n].freeze
|
|
195
|
+
|
|
196
|
+
def params(operation, payload)
|
|
197
|
+
out = {}
|
|
198
|
+
COMMON_PARAMS.each do |key|
|
|
199
|
+
value = payload[key]
|
|
200
|
+
next if value.nil?
|
|
201
|
+
# false is kept, not dropped: `caching` defaults to nil, so
|
|
202
|
+
# caching: false is a deliberate choice, and reproducing a call
|
|
203
|
+
# needs the settings it ran with. RubyLLM does not distinguish a
|
|
204
|
+
# boolean that was set from one that defaulted, so record both
|
|
205
|
+
# rather than guess which mattered.
|
|
206
|
+
out[key] = value.is_a?(Numeric) || [ true, false ].include?(value) ? value : value.to_s[0, 128]
|
|
207
|
+
end
|
|
208
|
+
out[:schema] = true if payload[:schema]
|
|
209
|
+
out[:server_tools] = Array(payload[:server_tools]).map(&:to_s).first(20) if payload[:server_tools].present?
|
|
210
|
+
if (usage = payload[:tokens]).respond_to?(:server_tool_use) && usage.server_tool_use.present?
|
|
211
|
+
out[:server_tool_use] = usage.server_tool_use
|
|
212
|
+
end
|
|
213
|
+
if (options = payload[:provider_options]).is_a?(Hash) && !options.empty?
|
|
214
|
+
out[:provider_options] = provider_options(options)
|
|
215
|
+
end
|
|
216
|
+
out[:operation] = operation unless out.empty?
|
|
217
|
+
out.empty? ? nil : out
|
|
218
|
+
end
|
|
219
|
+
|
|
220
|
+
# Two filters, because one is not enough here. The app's parameter
|
|
221
|
+
# filter defaults to password-shaped names only, and provider_options
|
|
222
|
+
# is the one place in this payload where a per-request credential
|
|
223
|
+
# plausibly lives -- an api_key passed per call sails straight through
|
|
224
|
+
# a password filter. The redactor's credential-name matcher (the same
|
|
225
|
+
# one that catches X-Api-Key on a header) closes that.
|
|
226
|
+
def provider_options(options)
|
|
227
|
+
redact_credentials(Railwatch.redactor.params(options.transform_keys(&:to_s)))
|
|
228
|
+
end
|
|
229
|
+
|
|
230
|
+
# Recursive, because provider options nest: extra_headers carrying an
|
|
231
|
+
# authorization value is a hash inside the hash, and a top-level scan
|
|
232
|
+
# walks straight past it.
|
|
233
|
+
def redact_credentials(value, depth = 0)
|
|
234
|
+
return value if depth > 4
|
|
235
|
+
|
|
236
|
+
case value
|
|
237
|
+
when Hash
|
|
238
|
+
value.each_with_object({}) do |(key, item), out|
|
|
239
|
+
# The matcher is written for header names, which are hyphenated;
|
|
240
|
+
# provider options are Ruby-ish and use underscores, so api_key
|
|
241
|
+
# would sail past a pattern expecting api-key.
|
|
242
|
+
out[key] = if Railwatch.redactor.redact_header?(key.to_s.tr("_", "-"))
|
|
243
|
+
Redactor::FILTERED
|
|
244
|
+
else
|
|
245
|
+
redact_credentials(item, depth + 1)
|
|
246
|
+
end
|
|
247
|
+
end
|
|
248
|
+
when Array then value.map { |item| redact_credentials(item, depth + 1) }
|
|
249
|
+
else value
|
|
250
|
+
end
|
|
251
|
+
end
|
|
252
|
+
|
|
253
|
+
def last_user_message(payload)
|
|
254
|
+
messages = payload[:input_messages]
|
|
255
|
+
return nil unless messages.respond_to?(:reverse_each)
|
|
256
|
+
|
|
257
|
+
messages.reverse_each.find { |m| m.respond_to?(:role) && m.role.to_s == "user" }
|
|
258
|
+
end
|
|
259
|
+
|
|
260
|
+
# 2.0 sends a RubyLLM::Tokens; 1.16 sends bare counts on the event.
|
|
261
|
+
def tokens(payload)
|
|
262
|
+
counts = payload[:tokens]
|
|
263
|
+
if counts.respond_to?(:input)
|
|
264
|
+
{ input_tokens: counts.input, output_tokens: counts.output,
|
|
265
|
+
cache_read_tokens: counts.cache_read, cache_write_tokens: counts.cache_write,
|
|
266
|
+
thinking_tokens: counts.thinking }
|
|
267
|
+
else
|
|
268
|
+
{ input_tokens: payload[:input_tokens], output_tokens: payload[:output_tokens],
|
|
269
|
+
cache_read_tokens: payload[:cached_tokens], cache_write_tokens: payload[:cache_creation_tokens],
|
|
270
|
+
thinking_tokens: payload[:thinking_tokens] }
|
|
271
|
+
end
|
|
272
|
+
end
|
|
273
|
+
|
|
274
|
+
# nil in three distinct cases that all mean "we do not know": RubyLLM
|
|
275
|
+
# 1.16 (which reports no cost), a model the registry has no pricing
|
|
276
|
+
# for, and an operation that used no tokens. Never zero -- an unpriced
|
|
277
|
+
# call is not a free call, and the difference matters on a bill.
|
|
278
|
+
def cost_nanos(payload)
|
|
279
|
+
cost = payload[:cost]
|
|
280
|
+
return nil unless cost.respond_to?(:total)
|
|
281
|
+
|
|
282
|
+
total = cost.total
|
|
283
|
+
total && (total * NANOS_PER_DOLLAR).round
|
|
284
|
+
end
|
|
285
|
+
|
|
286
|
+
# Present only on 2.0, and only inside RubyLLM.workflow. Stamped on
|
|
287
|
+
# every event the block emits, which is what lets an agent run be
|
|
288
|
+
# reassembled from its steps.
|
|
289
|
+
def workflow(payload)
|
|
290
|
+
return {} unless payload[:workflow_id]
|
|
291
|
+
|
|
292
|
+
{ workflow_id: payload[:workflow_id].to_s[0, 64],
|
|
293
|
+
workflow_name: payload[:workflow_name].to_s[0, 255],
|
|
294
|
+
workflow_step_id: payload[:workflow_step_id]&.to_s&.slice(0, 64),
|
|
295
|
+
workflow_step_name: payload[:workflow_step_name]&.to_s&.slice(0, 255),
|
|
296
|
+
workflow_step_parent_id: payload[:workflow_step_parent_id]&.to_s&.slice(0, 64) }
|
|
297
|
+
end
|
|
298
|
+
|
|
299
|
+
# Rails adds :exception to the payload when the instrumented block
|
|
300
|
+
# raised; RubyLLM leaves the rest of the event untouched in that case.
|
|
301
|
+
def outcome(payload)
|
|
302
|
+
error = payload[:exception]
|
|
303
|
+
return { status: "ok" } unless error
|
|
304
|
+
|
|
305
|
+
{ status: "failed", error: "#{Array(error).first}: #{Array(error).last}"[0, 255] }
|
|
306
|
+
end
|
|
307
|
+
|
|
308
|
+
# The last thing the app asked, which is the half of a conversation
|
|
309
|
+
# worth seeing next to a cost. Earlier turns are the app's own records.
|
|
310
|
+
def prompt_text(payload)
|
|
311
|
+
return payload[:input] || payload[:prompt] || payload[:query] unless payload[:input_messages].respond_to?(:reverse_each)
|
|
312
|
+
|
|
313
|
+
message_text(last_user_message(payload))
|
|
314
|
+
end
|
|
315
|
+
|
|
316
|
+
def message_text(message)
|
|
317
|
+
return nil if message.nil?
|
|
318
|
+
|
|
319
|
+
message.respond_to?(:content) ? message.content : message
|
|
320
|
+
end
|
|
321
|
+
|
|
322
|
+
def content(value)
|
|
323
|
+
return nil unless Railwatch.config.capture_llm_content
|
|
324
|
+
return nil if value.nil?
|
|
325
|
+
|
|
326
|
+
text = value.to_s
|
|
327
|
+
return nil if text.empty?
|
|
328
|
+
# byteslice, not [0, n]: the cap bounds what is buffered and shipped,
|
|
329
|
+
# and 4096 characters of CJK is three times that in bytes. scrub
|
|
330
|
+
# repairs the multibyte character the slice may have cut in half.
|
|
331
|
+
text.bytesize > CONTENT_MAX ? text.byteslice(0, CONTENT_MAX).scrub("") : text
|
|
332
|
+
end
|
|
333
|
+
end
|
|
334
|
+
end
|
|
335
|
+
end
|
|
@@ -12,6 +12,7 @@ require "railwatch/subscribers/storage"
|
|
|
12
12
|
require "railwatch/subscribers/views"
|
|
13
13
|
require "railwatch/subscribers/logs"
|
|
14
14
|
require "railwatch/subscribers/jobs"
|
|
15
|
+
require "railwatch/subscribers/llm"
|
|
15
16
|
require "railwatch/subscribers/deprecations"
|
|
16
17
|
require "railwatch/subscribers/users"
|
|
17
18
|
require "railwatch/subscribers/process_info"
|
|
@@ -19,7 +20,7 @@ require "railwatch/subscribers/process_info"
|
|
|
19
20
|
module Railwatch
|
|
20
21
|
module Subscribers
|
|
21
22
|
ALL = [ Requests, Queries, Exceptions, Cache, Mail, Broadcasts, Notifications,
|
|
22
|
-
Storage, Views, Logs, Jobs, Deprecations, Users, ProcessInfo ].freeze
|
|
23
|
+
Storage, Views, Logs, Jobs, Deprecations, Users, ProcessInfo, Llm ].freeze
|
|
23
24
|
|
|
24
25
|
module_function
|
|
25
26
|
|
data/lib/railwatch/version.rb
CHANGED
data/lib/railwatch.rb
CHANGED
|
@@ -496,7 +496,7 @@ module Railwatch
|
|
|
496
496
|
query: :queries, n_plus_one: :queries, transaction: :transactions, cache_event: :cache_events,
|
|
497
497
|
mail: :mail, broadcast: :broadcasts, notification: :notifications, outgoing_request: :outgoing_requests,
|
|
498
498
|
storage_op: :storage_ops, view_render: :view_renders, log: :logs, deprecation: :deprecations,
|
|
499
|
-
session: :sessions
|
|
499
|
+
session: :sessions, llm_call: :llm_calls
|
|
500
500
|
}.freeze
|
|
501
501
|
|
|
502
502
|
def type_plural(type)
|
metadata
CHANGED
|
@@ -1,13 +1,13 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: railwatch
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.1.
|
|
4
|
+
version: 0.1.4
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Cole Robertson
|
|
8
8
|
bindir: bin
|
|
9
9
|
cert_chain: []
|
|
10
|
-
date: 2026-09-
|
|
10
|
+
date: 2026-09-16 00:00:00.000000000 Z
|
|
11
11
|
dependencies:
|
|
12
12
|
- !ruby/object:Gem::Dependency
|
|
13
13
|
name: rails
|
|
@@ -115,6 +115,7 @@ files:
|
|
|
115
115
|
- lib/railwatch/subscribers/deprecations.rb
|
|
116
116
|
- lib/railwatch/subscribers/exceptions.rb
|
|
117
117
|
- lib/railwatch/subscribers/jobs.rb
|
|
118
|
+
- lib/railwatch/subscribers/llm.rb
|
|
118
119
|
- lib/railwatch/subscribers/logs.rb
|
|
119
120
|
- lib/railwatch/subscribers/mail.rb
|
|
120
121
|
- lib/railwatch/subscribers/notifications.rb
|