dead_bro 0.2.31 → 0.2.33

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: b6c0f13b23ff342d0ea4365872e2f380a949a8aaf3fd69d98606336854d14c9c
4
- data.tar.gz: a9b8a2fae51c26e7cd9b2f1d676c6742a5a5b4829f0c182417af7eabdf2581c6
3
+ metadata.gz: f1569a9a3be524d19ba616313570f01cb251647a76384347993655a8ffbaef9e
4
+ data.tar.gz: 3a091b0d844b12cfb9b61cf7b0d2eec04514bb5ef988fba7203d45b4ab237e67
5
5
  SHA512:
6
- metadata.gz: e70b0d0b091dfacf1db61ea287a47bb1944153968bc3f225c943baf6a195c099dfd88199e59a942230284f3cfc94dcbf207d942e03630ac14850f818c1dc0928
7
- data.tar.gz: 1e5f8605da3bf8c418d4c16025be5a24b7f89af1ca47d7bb6c767f790205c498e3023543b739b1568156547c21b8cfaf2bf327c573d8177aad88049e395352f2
6
+ metadata.gz: ebc44f95253369e7c8ab6f3873ed9143e0d72382c477bae39ca09491c22c93b236b2bc1dcc07599dd755797564f4a0f25f34104cf1f95246ea8d91fa9b4c1935
7
+ data.tar.gz: 6dd40034800398ceac4e82895d427e484ee7fd1b321a22882c1138c0e73514bfb4675a14437abd88cb6f4e5327c01ad85f2c7e75269956c4dd9d4e240a946714
data/CHANGELOG.md CHANGED
@@ -1,5 +1,18 @@
1
1
  ## [Unreleased]
2
2
 
3
+ ## [0.2.32] - 2026-09-22
4
+
5
+ ### Fixed
6
+ - **Errored requests now ship the SQL/cache/memory/etc data captured before the exception, not just the exception itself.** Previously, as soon as a request raised (e.g. an `ActiveRecord::LockWaitTimeout` mid-transaction), the subscriber discarded every already-captured `sql_queries`, `cache_events`, `ar_instantiation_count`, `memory_events`, `gc_stats`, etc. entry and sent only exception metadata — so error pages showed a backtrace with zero query context even when the gem had captured it. The error payload now carries the same detail fields a successful request would.
7
+ - **Background job failures were reported to the dashboard as successful completions, not dropped.** There is no `exception.active_job` event anywhere in Rails/ActiveJob — the handler that subscribed to it never fired. `perform.active_job` (the real event) fires whether or not the job raised, but its handler built `status: "completed"` unconditionally and never inspected `data[:exception_object]`, so a failing job wasn't invisible — it was recorded as if it had succeeded. This silently skewed job success-rate and duration numbers for every account on a version prior to this one; there's nothing to backfill, since no exception data was ever sent for those occurrences. The handler now branches on `data[:exception_object]`: a failing job reports `status: "failed"` with `error`, `fingerprint`, and `cause_chain`, and ships regardless of sampling.
8
+ - **Known limitation:** this only covers a job whose exception escapes `perform_now` uncaught. A job using `retry_on`, `discard_on`, or a custom `rescue_from` has the exception handled *inside* `perform_now`, before `perform.active_job`'s payload is ever built — so a retried or discarded job still reports `status: "completed"` for every handled attempt. Only the final attempt of a `retry_on` job that exhausts its retries with no block given re-raises far enough to be visible. Covering retried/discarded attempts would need separate instrumentation on `enqueue_retry.active_job` / `retry_stopped.active_job` / `discard.active_job` — not done here.
9
+ - **A request opening 5 or more transactions/savepoints could trip a false N+1 flag.** Every Rails adapter checked (MySQL/PostgreSQL/SQLite3, 7.1–8.1) logs transaction-control statements under the single name `"TRANSACTION"`, which the old skip list (`data[:name] == "BEGIN" || ...`) never matched — so `BEGIN`/`COMMIT`/etc. weren't ignored, they were tracked as ordinary queries: sanitized, counted toward N+1 detection, and folded into one aggregate per normalized statement. Five or more transactions in one request hit `N_PLUS_ONE_THRESHOLD` and flagged that aggregate `n_plus_one: true`, which the backend trusts directly for the aggregate payload format. Fixed as a consequence of the `TXN_CONTROL_NAMES` match below, not a separate change.
10
+
11
+ ### Added
12
+ - Transaction-control statements (`BEGIN`/`COMMIT`/`ROLLBACK`/`SAVEPOINT`/`RELEASE`) are now captured as lightweight breadcrumbs (`transaction_events`, with offsets) instead of being tracked as ordinary queries — so a transaction that rolled back before an error shows up explicitly in the request trace, instead of as an anonymous, sanitized `sql_queries` entry.
13
+
14
+ ## [0.2.31] - 2026-08-30
15
+
3
16
  ### Added
4
17
  - Monitor thread now sends a synchronous heartbeat on startup before the first collection tick. This ensures remote settings — including `monitor_enabled` — are applied from the very first reporting cycle, so Sidekiq workers and other non-web processes that have not yet sent any metrics still receive the correct configuration immediately on boot rather than waiting up to 60 seconds for the first scheduled tick.
5
18
 
data/FEATURES.md ADDED
@@ -0,0 +1,333 @@
1
+ # ApmBro Feature List
2
+
3
+ A comprehensive feature list for comparing ApmBro with other APM (Application Performance Monitoring) tools.
4
+
5
+ ## Core Architecture
6
+
7
+ - **Rails Integration**: Automatic subscription to Rails events via ActiveSupport::Notifications
8
+ - **Zero-Configuration Setup**: Works out of the box with minimal configuration
9
+ - **Asynchronous Metrics Posting**: Non-blocking HTTP requests using background threads
10
+ - **Thread-Local Storage**: Per-request metric collection using thread-local variables
11
+ - **Circuit Breaker Pattern**: Built-in circuit breaker to prevent cascading failures when APM endpoint is down
12
+ - **Deploy Tracking**: Automatic deploy ID resolution from multiple sources (Rails settings, ENV vars, Heroku, Git)
13
+
14
+ ## Request Tracking
15
+
16
+ ### Controller Action Monitoring
17
+ - **Automatic Tracking**: Tracks all controller actions automatically
18
+ - **Request Duration**: Measures total request processing time
19
+ - **HTTP Method & Path**: Captures HTTP method and request path
20
+ - **Status Codes**: Tracks HTTP response status codes
21
+ - **View Runtime**: Separate tracking of view rendering time
22
+ - **Database Runtime**: Separate tracking of database query time
23
+ - **Request Parameters**: Captures request parameters (with sensitive data filtering)
24
+ - **User Agent**: Tracks user agent strings
25
+ - **User ID Extraction**: Extracts authenticated user ID (supports Warden)
26
+ - **Environment Context**: Tracks Rails environment (development, staging, production)
27
+
28
+ ### Request Sampling
29
+ - **Configurable Sample Rate**: Percentage-based sampling (1-100%)
30
+ - **Random Sampling**: Each request has random chance of being tracked
31
+ - **Consistent Per-Request**: Sampling decision applies to all metrics for a request
32
+ - **Error Override**: Errors are always tracked regardless of sampling
33
+ - **Cost Optimization**: Reduces data volume and costs for high-traffic applications
34
+
35
+ ### Exclusion Rules
36
+ - **Controller Exclusion**: Exclude entire controllers from tracking
37
+ - **Action Exclusion**: Exclude specific controller#action combinations
38
+ - **Wildcard Support**: Pattern matching with `*` wildcards (e.g., `Admin::*`, `Admin::*#*`)
39
+ - **Job Exclusion**: Exclude specific background jobs from tracking
40
+ - **Flexible Configuration**: Configure via initializer, Rails settings, or environment variables
41
+
42
+ ## SQL Query Tracking
43
+
44
+ ### Query Details
45
+ - **Full SQL Tracking**: Captures all SQL queries executed during requests and jobs
46
+ - **Query Sanitization**: Automatically sanitizes SQL to remove sensitive data
47
+ - **Query Name**: Tracks query names (e.g., "User Load", "User Update")
48
+ - **Duration Measurement**: Precise query execution time in milliseconds
49
+ - **Cache Detection**: Identifies cached queries
50
+ - **Connection ID**: Tracks database connection ID
51
+ - **Call Stack Traces**: Full backtrace showing where queries were executed
52
+ - **Object Allocations**: Optional tracking of object allocations per query
53
+
54
+ ### Query Performance Analysis
55
+ - **Slow Query Detection**: Configurable threshold for identifying slow queries
56
+ - **EXPLAIN ANALYZE**: Automatic execution plan capture for slow queries
57
+ - **Background Execution**: EXPLAIN ANALYZE runs in separate thread (non-blocking)
58
+ - **Multi-Database Support**: Works with PostgreSQL, MySQL, SQLite, and others
59
+ - **Smart Filtering**: Automatically skips transaction queries (BEGIN, COMMIT, ROLLBACK)
60
+ - **Execution Plan Details**:
61
+ - PostgreSQL: Full EXPLAIN ANALYZE with buffer usage statistics
62
+ - MySQL: EXPLAIN ANALYZE with actual execution times
63
+ - SQLite: EXPLAIN QUERY PLAN output
64
+ - **Query Optimization Insights**: Helps identify missing indexes, full table scans, JOIN issues
65
+
66
+ ## View Rendering Tracking
67
+
68
+ ### View Performance
69
+ - **Template Rendering**: Tracks main template rendering
70
+ - **Partial Rendering**: Tracks partial template rendering with cache key information
71
+ - **Collection Rendering**: Tracks collection rendering (partials in loops)
72
+ - **Rendering Duration**: Precise timing for each view component
73
+ - **Virtual Path Tracking**: Tracks view virtual paths
74
+ - **Layout Information**: Captures layout usage
75
+
76
+ ### View Analysis
77
+ - **Slow View Detection**: Identifies the slowest rendering views
78
+ - **Frequency Analysis**: Tracks most frequently rendered views
79
+ - **Cache Hit Rate**: Calculates cache hit rates for partials
80
+ - **Collection Cache Analysis**: Tracks cache hit rates for collection rendering
81
+ - **Performance Metrics**:
82
+ - Total views rendered per request
83
+ - Total view rendering duration
84
+ - Average view rendering duration
85
+ - Breakdown by view type (template, partial, collection)
86
+
87
+ ## Memory Tracking & Leak Detection
88
+
89
+ ### Lightweight Memory Tracking (Default)
90
+ - **Memory Usage Monitoring**: Tracks memory consumption per request using GC stats
91
+ - **Memory Growth Tracking**: Measures memory growth during request processing
92
+ - **GC Statistics**: Tracks garbage collection count and heap pages
93
+ - **Minimal Performance Impact**: ~0.1ms overhead per request
94
+ - **Memory Before/After**: Captures memory state at request start and end
95
+
96
+ ### Detailed Allocation Tracking (Optional)
97
+ - **Object Allocation Tracking**: Detailed tracking of object allocations (disabled by default)
98
+ - **Allocation Sampling**: Configurable sampling rate for allocations
99
+ - **Large Object Detection**: Identifies objects larger than 1MB threshold
100
+ - **Memory Snapshots**: Periodic memory snapshots during request processing
101
+ - **Object Count Tracking**: Tracks object counts before and after requests
102
+ - **Performance Impact**: ~2-5ms overhead per request (only when enabled)
103
+
104
+ ### Memory Leak Detection
105
+ - **Pattern Detection**: Detects growing memory patterns over time
106
+ - **GC Efficiency Analysis**: Monitors garbage collection effectiveness
107
+ - **Heap Page Tracking**: Tracks heap page growth
108
+ - **Request Correlation**: Correlates memory growth with specific controllers/actions
109
+
110
+ ## Background Job Tracking
111
+
112
+ ### Job Execution Monitoring
113
+ - **ActiveJob Integration**: Automatic tracking when ActiveJob is available
114
+ - **Job Class Tracking**: Tracks job class names
115
+ - **Job ID**: Captures unique job identifiers
116
+ - **Queue Name**: Tracks which queue processed the job
117
+ - **Job Arguments**: Captures job arguments (with sensitive data filtering)
118
+ - **Duration Measurement**: Precise job execution time in milliseconds
119
+ - **Status Tracking**: Tracks job status (completed or failed)
120
+
121
+ ### Job Error Tracking
122
+ - **Exception Capture**: Captures exceptions from failed jobs
123
+ - **Exception Class**: Tracks exception class names
124
+ - **Exception Messages**: Captures exception messages (truncated to 1000 chars)
125
+ - **Backtraces**: Full exception backtraces (first 50 lines)
126
+ - **SQL Query Context**: Includes SQL queries executed during failed jobs
127
+ - **Memory Context**: Includes memory usage during job execution
128
+
129
+ ### Job SQL Tracking
130
+ - **SQL Query Tracking**: Tracks all SQL queries executed during job processing
131
+ - **Query Details**: Same detailed SQL tracking as request tracking
132
+ - **Query Context**: Full context of database operations in background jobs
133
+
134
+ ## Cache Tracking
135
+
136
+ ### Cache Operations
137
+ - **Read Operations**: Tracks cache read operations
138
+ - **Write Operations**: Tracks cache write operations
139
+ - **Delete Operations**: Tracks cache delete operations
140
+ - **Existence Checks**: Tracks cache existence checks
141
+ - **Fetch Operations**: Tracks cache fetch with hit/miss detection
142
+ - **Multi-Read Operations**: Tracks cache read_multi operations
143
+ - **Multi-Write Operations**: Tracks cache write_multi operations
144
+ - **Cache Generation**: Tracks cache generation events
145
+
146
+ ### Cache Analysis
147
+ - **Cache Hit Detection**: Identifies cache hits vs misses
148
+ - **Cache Key Tracking**: Tracks cache keys (truncated to 200 chars)
149
+ - **Store Information**: Identifies which cache store was used
150
+ - **Namespace Tracking**: Tracks cache namespaces
151
+ - **Duration Measurement**: Precise timing for each cache operation
152
+ - **Hit Rate Calculation**: Calculates cache hit rates per request
153
+
154
+ ## Redis Tracking
155
+
156
+ ### Redis Command Tracking
157
+ - **Command Monitoring**: Tracks all Redis commands executed
158
+ - **Command Name**: Captures Redis command names (GET, SET, etc.)
159
+ - **Key Tracking**: Tracks Redis keys (truncated to 200 chars)
160
+ - **Argument Count**: Tracks number of arguments per command
161
+ - **Database Selection**: Tracks which Redis database is used
162
+ - **Duration Measurement**: Precise timing for each Redis command
163
+ - **Error Tracking**: Captures Redis command errors
164
+
165
+ ### Advanced Redis Features
166
+ - **Pipeline Support**: Tracks Redis pipeline operations with command counts
167
+ - **Multi/Transaction Support**: Tracks Redis MULTI/EXEC transactions
168
+ - **ActiveSupport Integration**: Subscribes to ActiveSupport::Notifications for Redis events
169
+ - **Client Instrumentation**: Direct instrumentation of Redis::Client for comprehensive coverage
170
+
171
+ ## Error Tracking
172
+
173
+ ### Exception Handling
174
+ - **Automatic Exception Capture**: Captures exceptions from controller actions
175
+ - **Exception Class**: Tracks exception class names
176
+ - **Exception Messages**: Captures exception messages (truncated to 1000 chars)
177
+ - **Full Backtraces**: Captures complete exception backtraces (first 50 lines)
178
+ - **Request Context**: Includes full request context with exceptions
179
+ - **Error Flagging**: Errors are marked and always tracked (even with sampling)
180
+
181
+ ### Error Context
182
+ - **Controller/Action**: Identifies where the error occurred
183
+ - **Request Parameters**: Includes request parameters at time of error
184
+ - **User Information**: Includes user ID if available
185
+ - **SQL Queries**: Includes SQL queries executed before error
186
+ - **Memory State**: Includes memory usage at time of error
187
+ - **Log Messages**: Includes application logs captured during request
188
+
189
+ ## HTTP Instrumentation
190
+
191
+ ### Outgoing HTTP Tracking
192
+ - **HTTP Request Tracking**: Tracks outgoing HTTP requests (via middleware)
193
+ - **Request Context**: Captures HTTP request details
194
+ - **Response Context**: Captures HTTP response details
195
+ - **Duration Measurement**: Tracks HTTP request duration
196
+
197
+ ## Configuration & Flexibility
198
+
199
+ ### Configuration Options
200
+ - **API Key Management**: Multiple sources (config, Rails credentials, ENV)
201
+ - **Endpoint Configuration**: Configurable endpoint URL
202
+ - **Timeout Settings**: Configurable open and read timeouts
203
+ - **Enable/Disable Toggle**: Can be enabled/disabled via configuration
204
+ - **Environment Detection**: Automatic Rails environment detection
205
+
206
+ ### Circuit Breaker Configuration
207
+ - **Failure Threshold**: Configurable failure threshold (default: 3)
208
+ - **Recovery Timeout**: Configurable recovery timeout (default: 60 seconds)
209
+ - **Retry Timeout**: Configurable retry timeout (default: 300 seconds)
210
+ - **Enable/Disable**: Can enable/disable circuit breaker
211
+
212
+ ### Memory Tracking Configuration
213
+ - **Memory Tracking Toggle**: Enable/disable memory tracking
214
+ - **Allocation Tracking Toggle**: Enable/disable detailed allocation tracking
215
+ - **Sampling Configuration**: Configurable request sampling rate
216
+
217
+ ### Query Analysis Configuration
218
+ - **Slow Query Threshold**: Configurable threshold in milliseconds (default: 500ms)
219
+ - **EXPLAIN ANALYZE Toggle**: Enable/disable automatic EXPLAIN ANALYZE
220
+
221
+ ## Data Safety & Privacy
222
+
223
+ ### Data Sanitization
224
+ - **SQL Sanitization**: Automatically sanitizes SQL queries
225
+ - **Parameter Filtering**: Filters sensitive parameters (password, token, secret, key)
226
+ - **Argument Truncation**: Limits and truncates job arguments
227
+ - **Key Truncation**: Truncates cache and Redis keys to 200 characters
228
+ - **Value Truncation**: Recursively truncates nested values to prevent huge payloads
229
+ - **String Limits**: Limits string values (e.g., user agent to 200 chars, messages to 1000 chars)
230
+
231
+ ### Data Limits
232
+ - **Array Limits**: Limits array sizes (e.g., first 10 job arguments, first 5 array elements)
233
+ - **Hash Limits**: Limits hash key counts (e.g., first 20 hash keys, first 30 params)
234
+ - **Backtrace Limits**: Limits backtraces to first 50 lines
235
+ - **Allocation Limits**: Limits allocations tracked per request (max 1000)
236
+
237
+ ## Performance & Reliability
238
+
239
+ ### Performance Optimizations
240
+ - **Asynchronous Posting**: Non-blocking HTTP requests
241
+ - **Lightweight Default Mode**: Minimal overhead in default configuration
242
+ - **Sampling Support**: Reduces data volume for high-traffic applications
243
+ - **Thread-Local Storage**: Efficient per-request data collection
244
+ - **Background EXPLAIN**: EXPLAIN ANALYZE runs in background thread
245
+
246
+ ### Reliability Features
247
+ - **Circuit Breaker**: Prevents cascading failures
248
+ - **Error Handling**: Comprehensive error handling to prevent instrumentation failures
249
+ - **Graceful Degradation**: Continues working even if some features fail
250
+ - **Timeout Protection**: Configurable timeouts prevent hanging requests
251
+
252
+ ## Integration & Compatibility
253
+
254
+ ### Framework Support
255
+ - **Rails Integration**: Full Rails integration via Railtie
256
+ - **ActiveSupport Notifications**: Uses ActiveSupport::Notifications for event subscription
257
+ - **ActiveRecord Integration**: Tracks ActiveRecord SQL queries
258
+ - **ActiveJob Integration**: Tracks ActiveJob background jobs
259
+ - **ActionView Integration**: Tracks ActionView rendering
260
+
261
+ ### Database Support
262
+ - **PostgreSQL**: Full support with EXPLAIN ANALYZE
263
+ - **MySQL**: Full support with EXPLAIN ANALYZE
264
+ - **SQLite**: Full support with EXPLAIN QUERY PLAN
265
+ - **Other Databases**: Basic support with standard EXPLAIN
266
+
267
+ ### Cache Store Support
268
+ - **All Cache Stores**: Works with any Rails cache store
269
+ - **Multi-Store Support**: Tracks cache operations across different stores
270
+
271
+ ### Redis Support
272
+ - **Redis Gem**: Works with redis gem
273
+ - **Client Instrumentation**: Direct instrumentation of Redis::Client
274
+ - **Pipeline Support**: Tracks Redis pipelines
275
+ - **Transaction Support**: Tracks Redis MULTI/EXEC transactions
276
+
277
+ ## Logging & Debugging
278
+
279
+ ### Application Logging
280
+ - **Log Capture**: Captures application logs during request processing
281
+ - **Log Context**: Includes logs in metric payloads
282
+ - **Debug Logging**: Optional debug logging for skipped requests
283
+
284
+ ## Deployment & Environment
285
+
286
+ ### Deploy Tracking
287
+ - **Deploy ID Resolution**: Multiple sources for deploy identification (`Configuration#deploy_id=` wins when set, then ENV in `Configuration::DEPLOY_REVISION_ENV_KEYS` order—including `DEAD_BRO_DEPLOY_ID`, git/CI vars, `DD_VERSION`, etc.), otherwise a **per-process UUID** (fine for single dyno/process; unusable alone for fleets like ECS replicas)
288
+ - **Revision Tracking**: Includes deploy/revision ID in all metric payloads
289
+
290
+ ### Environment Support
291
+ - **Rails Environment**: Automatic Rails environment detection
292
+ - **Rack Environment**: Fallback to RACK_ENV or RAILS_ENV
293
+ - **Environment Context**: Includes environment in all metric payloads
294
+
295
+ ## Data Collection & Transmission
296
+
297
+ ### Metric Payload Structure
298
+ - **Structured Data**: Well-structured JSON payloads
299
+ - **Event Names**: Descriptive event names for different metric types
300
+ - **Timestamp Tracking**: ISO8601 timestamps for all metrics
301
+ - **Metadata**: Rich metadata including environment, host, deploy ID
302
+
303
+ ### HTTP Client
304
+ - **HTTPS Support**: Secure HTTPS communication
305
+ - **Bearer Token Auth**: API key authentication via Bearer tokens
306
+ - **JSON Encoding**: JSON-encoded payloads
307
+ - **Custom Headers**: Proper Content-Type and Authorization headers
308
+
309
+ ## Comparison-Ready Features
310
+
311
+ ### Unique Differentiators
312
+ 1. **Automatic EXPLAIN ANALYZE**: Background execution plan capture for slow queries
313
+ 2. **Lightweight Memory Tracking**: Low-overhead memory monitoring by default
314
+ 3. **Comprehensive Cache Tracking**: Detailed cache operation tracking
315
+ 4. **Redis Instrumentation**: Full Redis command tracking including pipelines
316
+ 5. **View Rendering Analysis**: Detailed view performance analysis with cache hit rates
317
+ 6. **Flexible Exclusion Rules**: Wildcard support for controller/job exclusion
318
+ 7. **Request Sampling**: Configurable percentage-based sampling
319
+ 8. **Circuit Breaker**: Built-in resilience for APM endpoint failures
320
+ 9. **Multi-Source Configuration**: Flexible configuration from multiple sources
321
+ 10. **Deploy Tracking**: Automatic deploy ID resolution from multiple sources
322
+
323
+ ### Standard APM Features
324
+ - Request/response tracking
325
+ - SQL query tracking
326
+ - Error tracking
327
+ - Background job tracking
328
+ - Memory tracking
329
+ - Performance metrics
330
+ - Exception handling
331
+ - User context
332
+ - Environment tracking
333
+
data/README.md CHANGED
@@ -94,11 +94,11 @@ DeadBro can automatically capture the query plan (`EXPLAIN`) of slow SELECT quer
94
94
 
95
95
  ### Configuration
96
96
 
97
- - **`explain_analyze_enabled`** (default: `false`) — this is the one setting that must be turned on in Ruby; the dashboard toggle can only turn capture *off*, never on, as an extra safety rail:
97
+ - **`explain_enabled`** (default: `false`) — this is the one setting that must be turned on in Ruby; the dashboard toggle can only turn capture *off*, never on, as an extra safety rail:
98
98
 
99
99
  ```ruby
100
100
  DeadBro.configure do |config|
101
- config.explain_analyze_enabled = true
101
+ config.explain_enabled = true
102
102
  end
103
103
  ```
104
104
 
@@ -12,8 +12,12 @@ module DeadBro
12
12
  # Local-only opt-in for EXPLAIN plan capture on slow queries. The gem runs
13
13
  # DB statements (plan-only EXPLAIN) when this is on, so it must be enabled
14
14
  # in the app's own config — remote settings can only turn it OFF, never on
15
- # (see apply_remote_settings). Effective state: #explain_analyze_active?
16
- attr_accessor :explain_analyze_enabled
15
+ # (see apply_remote_settings). Effective state: #explain_active?
16
+ attr_accessor :explain_enabled
17
+
18
+ # Compatibility for initializers written before the EXPLAIN option was renamed.
19
+ alias_method :explain_analyze_enabled, :explain_enabled
20
+ alias_method :explain_analyze_enabled=, :explain_enabled=
17
21
 
18
22
  # Remote-managed settings (overwritten by backend JSON `settings` on successful API responses)
19
23
  attr_accessor :memory_tracking_enabled, :allocation_tracking_enabled, :allocation_sample_rate,
@@ -66,7 +70,7 @@ module DeadBro
66
70
 
67
71
  REMOTE_SETTING_KEYS = %w[
68
72
  enabled sample_rate memory_tracking_enabled allocation_tracking_enabled allocation_sample_rate
69
- explain_analyze_enabled slow_query_threshold_ms max_sql_queries_to_send max_logs_to_send
73
+ explain_enabled slow_query_threshold_ms max_sql_queries_to_send max_logs_to_send
70
74
  watch_enabled excluded_controllers excluded_jobs exclusive_controllers exclusive_jobs
71
75
  monitor_enabled enable_db_stats enable_process_stats enable_system_stats
72
76
  sample_rates_by_type
@@ -94,11 +98,11 @@ module DeadBro
94
98
  # sampling, allocation-source tracing) runs on this % of requests so the
95
99
  # ~2-5ms overhead can be capped without turning the feature fully off.
96
100
  @allocation_sample_rate = 100
97
- @explain_analyze_enabled = false
101
+ @explain_enabled = false
98
102
  # Remote kill switch for EXPLAIN capture. Defaults to true so a local
99
103
  # opt-in works against older backends that never send the key; any
100
- # response that sends explain_analyze_enabled: false turns capture off.
101
- @remote_explain_analyze_enabled = true
104
+ # response that sends explain_enabled: false turns capture off.
105
+ @remote_explain_enabled = true
102
106
  @slow_query_threshold_ms = 500
103
107
  @max_sql_queries_to_send = 500
104
108
  @max_logs_to_send = 100
@@ -166,6 +170,12 @@ module DeadBro
166
170
  def apply_remote_settings(hash)
167
171
  return unless hash.is_a?(Hash)
168
172
 
173
+ # Accept older servers; the canonical key wins when both are present.
174
+ hash = hash.transform_keys(&:to_s)
175
+ if hash.key?("explain_analyze_enabled") && !hash.key?("explain_enabled")
176
+ hash["explain_enabled"] = hash["explain_analyze_enabled"]
177
+ end
178
+
169
179
  @settings_mutex.synchronize do
170
180
  hash.each do |key, value|
171
181
  k = key.to_s
@@ -174,11 +184,11 @@ module DeadBro
174
184
  case k
175
185
  when "sample_rate", "allocation_sample_rate", "slow_query_threshold_ms", "max_sql_queries_to_send", "max_logs_to_send"
176
186
  send(:"#{k}=", value.to_i)
177
- when "explain_analyze_enabled"
187
+ when "explain_enabled"
178
188
  # EXPLAIN runs statements against the customer DB, so the backend
179
189
  # must never be able to switch it on — only off. The local opt-in
180
- # (explain_analyze_enabled) stays untouched; see #explain_analyze_active?
181
- @remote_explain_analyze_enabled = !!value
190
+ # (explain_enabled) stays untouched; see #explain_active?
191
+ @remote_explain_enabled = !!value
182
192
  when "enabled", "memory_tracking_enabled", "allocation_tracking_enabled", "watch_enabled",
183
193
  "monitor_enabled", "enable_db_stats", "enable_process_stats", "enable_system_stats"
184
194
  send(:"#{k}=", !!value)
@@ -193,10 +203,12 @@ module DeadBro
193
203
 
194
204
  # EXPLAIN capture requires BOTH the local opt-in and the remote flag —
195
205
  # locally off means off no matter what the backend sends.
196
- def explain_analyze_active?
197
- !!(@explain_analyze_enabled && @remote_explain_analyze_enabled)
206
+ def explain_active?
207
+ !!(@explain_enabled && @remote_explain_enabled)
198
208
  end
199
209
 
210
+ alias_method :explain_analyze_active?, :explain_active?
211
+
200
212
  def heartbeat_due?
201
213
  return false if api_key.nil?
202
214
  last_heartbeat_attempt_at.nil? || (Time.now.utc - last_heartbeat_attempt_at) >= HEARTBEAT_INTERVAL
@@ -9,7 +9,6 @@ end
9
9
  module DeadBro
10
10
  class JobSubscriber
11
11
  JOB_EVENT_NAME = "perform.active_job"
12
- JOB_EXCEPTION_EVENT_NAME = "exception.active_job"
13
12
 
14
13
  def self.subscribe!(client: Client.new)
15
14
  # Snap GC state before the job runs so stop_request_tracking gets a valid diff
@@ -19,7 +18,28 @@ module DeadBro
19
18
  rescue
20
19
  end
21
20
 
22
- # Track job execution
21
+ # Track job execution — success AND failure both land here. ActiveJob wraps
22
+ # perform with `instrument(:perform) { super }` (see
23
+ # ActiveJob::Instrumentation#instrument), and AS::Notifications.instrument
24
+ # still runs its "finish" listeners (with :exception / :exception_object set
25
+ # in the payload) when the block raises, then re-raises. There is no separate
26
+ # "exception.active_job" event anywhere in Rails/ActiveJob — a prior version
27
+ # of this file subscribed to one, so it never fired; the job wasn't dropped,
28
+ # it fired *this* event and built status: "completed" unconditionally,
29
+ # silently reporting every failing job as a success. Branching on
30
+ # data[:exception_object] here is the only place a job failure can actually
31
+ # be detected.
32
+ #
33
+ # Known gap: this only sees an exception that escapes perform_now uncaught.
34
+ # ActiveJob::Base#perform_now (Execution) rescues internally and hands off to
35
+ # rescue_with_handler — which is exactly what retry_on/discard_on/rescue_from
36
+ # are built on (ActiveJob::Exceptions) — before Instrumentation's `super`
37
+ # even returns. A handled retry or discard returns normally with no
38
+ # exception attached, so perform.active_job still reports status:
39
+ # "completed" for every handled attempt; only a retry_on job's final,
40
+ # unhandled raise (attempts exhausted, no block given) is visible here.
41
+ # Covering handled attempts would need separate instrumentation on
42
+ # enqueue_retry.active_job / retry_stopped.active_job / discard.active_job.
23
43
  ActiveSupport::Notifications.subscribe(JOB_EVENT_NAME) do |name, started, finished, _unique_id, data|
24
44
  begin
25
45
  if DeadBro.configuration.skip_tracking?
@@ -40,12 +60,14 @@ module DeadBro
40
60
  rescue
41
61
  end
42
62
 
63
+ exception = data[:exception_object]
64
+ has_error = !exception.nil?
65
+
43
66
  # Skip out via sampling before we build any payload — jobs can be chatty
44
67
  # enough that even the "cheap" stop/analyze work matters under load.
45
- # Completions have no exception attached; the exception subscriber below
46
- # always sends errors with force: true.
68
+ # Errors always ship regardless of sampling, matching Subscriber's web path.
47
69
  job_type_key = "#{job_class_name}#perform"
48
- unless DeadBro.configuration.should_sample?(job_type_key)
70
+ unless has_error || DeadBro.configuration.should_sample?(job_type_key)
49
71
  drain_job_tracking
50
72
  next
51
73
  end
@@ -71,6 +93,7 @@ module DeadBro
71
93
 
72
94
  # Get SQL queries executed during this job
73
95
  sql_queries = DeadBro::SqlSubscriber.stop_request_tracking
96
+ transaction_events = DeadBro::SqlSubscriber.last_transaction_events
74
97
  dependency_events = job_dependency_payload
75
98
  db_connection_stats = defined?(DeadBro::DbConnectionSubscriber) ? DeadBro::DbConnectionSubscriber.stop_request_tracking : {}
76
99
  gc_pressure = defined?(DeadBro::GcTracker) ? DeadBro::GcTracker.stop_request_tracking : {}
@@ -117,8 +140,9 @@ module DeadBro
117
140
  db_connection_checkouts: db_connection_stats[:checkouts],
118
141
  gc_pressure: gc_pressure,
119
142
  ar_instantiation_count: ar_instantiation_count,
120
- status: "completed",
143
+ status: has_error ? "failed" : "completed",
121
144
  sql_queries: sql_queries,
145
+ transaction_events: transaction_events,
122
146
  rails_env: DeadBro.env,
123
147
  host: DeadBro.safe_hostname,
124
148
  process_kind: DeadBro.process_kind,
@@ -130,115 +154,19 @@ module DeadBro
130
154
  logs: DeadBro.logger.logs
131
155
  }.merge(dependency_events)
132
156
 
133
- # force: true — the sampling decision above already accounted for any
134
- # per-job-type override; client#post_metric must not re-roll it globally.
135
- client.post_metric(event_name: name, payload: payload, force: true)
136
- end
137
-
138
- # Track job exceptions
139
- ActiveSupport::Notifications.subscribe(JOB_EXCEPTION_EVENT_NAME) do |name, started, finished, _unique_id, data|
140
- begin
141
- if DeadBro.configuration.skip_tracking?
142
- drain_job_tracking
143
- next
144
- end
145
-
146
- job_class_name = data[:job].class.name
147
- if DeadBro.configuration.excluded_job?(job_class_name)
148
- next
149
- end
150
- # If exclusive_jobs is defined and not empty, only track matching jobs
151
- unless DeadBro.configuration.exclusive_job?(job_class_name)
152
- next
153
- end
154
- rescue
155
- end
156
-
157
- duration_ms = ((finished - started) * 1000.0).round(2)
158
- exception = data[:exception_object]
159
- queue_duration_ms = job_queue_duration_ms(data[:job], started)
160
-
161
- # Ensure tracking was started (fallback if perform_start.active_job didn't fire)
162
- unless DeadBro::SqlSubscriber.tracking_active?
163
- DeadBro.logger.clear
164
- Thread.current[DeadBro::TRACKING_START_TIME_KEY] = Time.now
165
- DeadBro::SqlSubscriber.start_request_tracking
166
- start_job_dependency_tracking
167
- DeadBro::DbConnectionSubscriber.start_request_tracking if defined?(DeadBro::DbConnectionSubscriber)
168
- DeadBro::WatchTracker.start_request_tracking if defined?(DeadBro::WatchTracker)
169
- if DeadBro.configuration.allocation_tracking_enabled && defined?(DeadBro::MemoryTrackingSubscriber)
170
- DeadBro::MemoryTrackingSubscriber.start_request_tracking
171
- else
172
- DeadBro::LightweightMemoryTracker.start_request_tracking if defined?(DeadBro::LightweightMemoryTracker)
173
- end
174
- end
175
-
176
- # Get SQL queries executed during this job
177
- sql_queries = DeadBro::SqlSubscriber.stop_request_tracking
178
- dependency_events = job_dependency_payload
179
- db_connection_stats = defined?(DeadBro::DbConnectionSubscriber) ? DeadBro::DbConnectionSubscriber.stop_request_tracking : {}
180
- gc_pressure = defined?(DeadBro::GcTracker) ? DeadBro::GcTracker.stop_request_tracking : {}
181
- ar_instantiation_count = defined?(DeadBro::ArObjectTracker) ? DeadBro::ArObjectTracker.stop_request_tracking : nil
182
- watch_events = defined?(DeadBro::WatchTracker) ? DeadBro::WatchTracker.stop_request_tracking : []
183
-
184
- # Stop memory tracking and get collected memory data
185
- if DeadBro.configuration.allocation_tracking_enabled && defined?(DeadBro::MemoryTrackingSubscriber)
186
- detailed_memory = DeadBro::MemoryTrackingSubscriber.stop_request_tracking
187
- memory_performance = DeadBro::MemoryTrackingSubscriber.analyze_memory_performance(detailed_memory)
188
- # Keep memory_events compact and user-friendly (no large raw arrays)
189
- memory_events = {
190
- memory_before: detailed_memory[:memory_before],
191
- memory_after: detailed_memory[:memory_after],
192
- duration_seconds: detailed_memory[:duration_seconds],
193
- allocations_count: (detailed_memory[:allocations] || []).length,
194
- memory_snapshots_count: (detailed_memory[:memory_snapshots] || []).length,
195
- large_objects_count: (detailed_memory[:large_objects] || []).length
196
- }
197
- else
198
- lightweight_memory = DeadBro::LightweightMemoryTracker.stop_request_tracking
199
- # Separate raw readings from derived performance metrics to avoid duplicating data
200
- memory_events = {
201
- memory_before: lightweight_memory[:memory_before],
202
- memory_after: lightweight_memory[:memory_after]
203
- }
204
- memory_performance = {
205
- memory_growth_mb: lightweight_memory[:memory_growth_mb],
206
- gc_count_increase: lightweight_memory[:gc_count_increase],
207
- heap_pages_increase: lightweight_memory[:heap_pages_increase],
208
- duration_seconds: lightweight_memory[:duration_seconds]
209
- }
157
+ if has_error
158
+ payload[:exception_class] = exception.class.name
159
+ payload[:message] = exception.message.to_s[0, 1000]
160
+ payload[:backtrace] = Array(exception.backtrace).first(50)
161
+ payload[:fingerprint] = DeadBro::Subscriber.compute_error_fingerprint(exception)
162
+ payload[:cause_chain] = DeadBro::Subscriber.build_cause_chain(exception)
163
+ payload[:error] = true
210
164
  end
211
165
 
212
- payload = {
213
- job_class: data[:job].class.name,
214
- job_id: data[:job].job_id,
215
- queue_name: data[:job].queue_name,
216
- arguments: safe_arguments(data[:job].arguments),
217
- started_at: started.utc.iso8601(3),
218
- duration_ms: duration_ms,
219
- queue_duration_ms: queue_duration_ms,
220
- db_connection_wait_ms: db_connection_stats[:wait_ms],
221
- db_connection_checkouts: db_connection_stats[:checkouts],
222
- gc_pressure: gc_pressure,
223
- ar_instantiation_count: ar_instantiation_count,
224
- status: "failed",
225
- sql_queries: sql_queries,
226
- exception_class: exception&.class&.name,
227
- message: exception&.message&.to_s&.[](0, 1000),
228
- backtrace: Array(exception&.backtrace).first(50),
229
- rails_env: DeadBro.env,
230
- host: DeadBro.safe_hostname,
231
- process_kind: DeadBro.process_kind,
232
- memory_usage: memory_usage_mb,
233
- gc_stats: gc_stats,
234
- memory_events: memory_events,
235
- memory_performance: memory_performance,
236
- watch_events: watch_events,
237
- logs: DeadBro.logger.logs
238
- }.merge(dependency_events)
239
-
240
- event_name = exception&.class&.name || "ActiveJob::Exception"
241
- client.post_metric(event_name: event_name, payload: payload, force: true)
166
+ # force: true — errors always ship, bypassing sampling by design; for
167
+ # completions the sampling decision above already accounted for any
168
+ # per-job-type override, so client#post_metric must not re-roll it globally.
169
+ client.post_metric(event_name: name, payload: payload, force: true)
242
170
  end
243
171
  rescue
244
172
  # Never raise from instrumentation install
@@ -248,6 +176,7 @@ module DeadBro
248
176
  # build a payload (excluded job / sampled out). Matches Subscriber.drain_request_tracking.
249
177
  def self.drain_job_tracking
250
178
  # wait_for_explains: false — result is discarded, don't block on pending plans.
179
+ # stop_request_tracking also pops the transaction-events stack (see its comment).
251
180
  DeadBro::SqlSubscriber.stop_request_tracking(wait_for_explains: false) if defined?(DeadBro::SqlSubscriber)
252
181
  Thread.current[:dead_bro_http_events] = nil
253
182
  DeadBro::CacheSubscriber.stop_request_tracking if defined?(DeadBro::CacheSubscriber)
@@ -16,7 +16,11 @@ module DeadBro
16
16
  THREAD_LOCAL_EXPLAIN_PENDING_KEY = :dead_bro_explain_pending
17
17
  THREAD_LOCAL_CALL_COUNTS_KEY = :dead_bro_sql_call_counts
18
18
  THREAD_LOCAL_AGGREGATES_KEY = :dead_bro_sql_aggregates
19
+ THREAD_LOCAL_TXN_EVENTS_KEY = :dead_bro_sql_txn_events
19
20
  MAX_TRACKED_QUERIES = 1000
21
+ # Transaction control breadcrumbs (BEGIN/COMMIT/ROLLBACK/...) are cheap and rare
22
+ # compared to queries, but a pathological retry loop could still spam them.
23
+ MAX_TRACKED_TXN_EVENTS = 200
20
24
 
21
25
  # Number of identical queries within one request that triggers N+1 detection.
22
26
  N_PLUS_ONE_THRESHOLD = 5
@@ -37,6 +41,43 @@ module DeadBro
37
41
  SANITIZE_SKIP_SENSITIVE_WHEN_NO_KEYWORDS = /password|token|secret|key|ssn|credit_card/i
38
42
  SANITIZE_SKIP_WHERE_WHEN_NO_KEYWORD = /WHERE/i
39
43
 
44
+ # Rails' own adapters (MySQL, PostgreSQL, SQLite3, at least 7.1 through 8.1 —
45
+ # the versions checked directly in this repo) log every transaction-control
46
+ # statement under the single generic name "TRANSACTION"
47
+ # (`internal_execute("BEGIN", "TRANSACTION", ...)`, `internal_execute("COMMIT",
48
+ # "TRANSACTION", ...)`, etc. — see AbstractMysqlAdapter/PostgreSQL::DatabaseStatements/
49
+ # SQLite3::DatabaseStatements). "BEGIN"/"COMMIT"/"ROLLBACK"/"SAVEPOINT"/"RELEASE"
50
+ # as literal `name` values are kept here only as a defensive fallback for
51
+ # adapters/older Rails versions that might still emit them directly — the actual
52
+ # operation for a "TRANSACTION"-named event is derived from the SQL text itself
53
+ # (see transaction_operation_for below), since the name alone can't distinguish
54
+ # BEGIN from COMMIT from a SAVEPOINT.
55
+ TXN_CONTROL_NAMES = %w[BEGIN COMMIT ROLLBACK SAVEPOINT RELEASE TRANSACTION].freeze
56
+
57
+ TXN_OPERATION_FROM_SQL = [
58
+ [/\A\s*ROLLBACK\s+TO\s+SAVEPOINT/i, "ROLLBACK TO SAVEPOINT"],
59
+ [/\A\s*RELEASE\s+SAVEPOINT/i, "RELEASE SAVEPOINT"],
60
+ [/\A\s*SAVEPOINT/i, "SAVEPOINT"],
61
+ [/\A\s*BEGIN/i, "BEGIN"],
62
+ [/\A\s*COMMIT/i, "COMMIT"],
63
+ [/\A\s*ROLLBACK/i, "ROLLBACK"]
64
+ ].freeze
65
+
66
+ # data[:name] is only ever "TRANSACTION" in current Rails — this derives the
67
+ # actual verb (BEGIN/COMMIT/ROLLBACK/SAVEPOINT/RELEASE SAVEPOINT/ROLLBACK TO
68
+ # SAVEPOINT) from the SQL text so breadcrumbs are still meaningful. A literal
69
+ # non-"TRANSACTION" name (older Rails/adapter) is passed through unchanged.
70
+ def self.transaction_operation_for(name, sql)
71
+ return name unless name == "TRANSACTION"
72
+ sql_str = sql.to_s
73
+ TXN_OPERATION_FROM_SQL.each do |re, op|
74
+ return op if sql_str.match?(re)
75
+ end
76
+ name
77
+ rescue
78
+ name
79
+ end
80
+
40
81
  # True when there is at least one active tracking context (e.g. for nested jobs).
41
82
  def self.tracking_active?
42
83
  stack = Thread.current[THREAD_LOCAL_KEY]
@@ -50,6 +91,13 @@ module DeadBro
50
91
  stack.last
51
92
  end
52
93
 
94
+ # Current transaction-events array (top of stack); nil if no active tracking.
95
+ def self.current_txn_events_array
96
+ stack = Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY]
97
+ return nil unless stack.is_a?(Array) && stack.any?
98
+ stack.last
99
+ end
100
+
53
101
  # Sum of SQL counts and durations recorded so far in the active tracking
54
102
  # context. Used by WatchTracker to attribute SQL to a DeadBro.watch block.
55
103
  def self.current_sql_metrics
@@ -96,6 +144,27 @@ module DeadBro
96
144
  sql.to_s.downcase
97
145
  end
98
146
 
147
+ # Records a lightweight breadcrumb for a transaction-control statement
148
+ # (BEGIN/COMMIT/ROLLBACK/SAVEPOINT/RELEASE) — enough to show, e.g., that the
149
+ # queries preceding an error ran inside a transaction that was rolled back,
150
+ # without the cost of treating it like a tracked query (no backtrace/EXPLAIN).
151
+ def self.record_transaction_event(name, started, finished)
152
+ current = current_txn_events_array
153
+ return unless current
154
+ return if current.length >= MAX_TRACKED_TXN_EVENTS
155
+
156
+ tracking_start = Thread.current[DeadBro::TRACKING_START_TIME_KEY]
157
+ start_offset_ms = tracking_start ? ((started - tracking_start) * 1000.0).round(2) : nil
158
+
159
+ current << {
160
+ event: name,
161
+ duration_ms: ((finished - started) * 1000.0).round(2),
162
+ start_offset_ms: start_offset_ms
163
+ }
164
+ rescue
165
+ nil
166
+ end
167
+
99
168
  def self.subscribe!
100
169
  # Subscribe with a start/finish listener to measure allocations per query
101
170
  if ActiveSupport::Notifications.notifier.respond_to?(:subscribe)
@@ -106,7 +175,13 @@ module DeadBro
106
175
  end
107
176
 
108
177
  ActiveSupport::Notifications.subscribe(SQL_EVENT_NAME) do |name, started, finished, _unique_id, data|
109
- next if data[:name] == "SCHEMA" || data[:name] == "CACHE" || data[:name] == "BEGIN" || data[:name] == "COMMIT" || data[:name] == "ROLLBACK" || data[:name] == "SAVEPOINT" || data[:name] == "RELEASE"
178
+ next if data[:name] == "SCHEMA" || data[:name] == "CACHE"
179
+
180
+ if TXN_CONTROL_NAMES.include?(data[:name])
181
+ record_transaction_event(transaction_operation_for(data[:name], data[:sql]), started, finished)
182
+ next
183
+ end
184
+
110
185
  # Only track queries that are part of the current request (top of stack for nested jobs)
111
186
  current = current_queries_array
112
187
  next unless current
@@ -208,6 +283,7 @@ module DeadBro
208
283
  Thread.current[THREAD_LOCAL_EXPLAIN_PENDING_KEY] = []
209
284
  (Thread.current[THREAD_LOCAL_CALL_COUNTS_KEY] ||= []) << {}
210
285
  (Thread.current[THREAD_LOCAL_AGGREGATES_KEY] ||= []) << {}
286
+ (Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY] ||= []) << []
211
287
  end
212
288
 
213
289
  # wait_for_explains: false on drain paths (excluded / sampled-out requests)
@@ -223,6 +299,21 @@ module DeadBro
223
299
  cc_stack = Thread.current[THREAD_LOCAL_CALL_COUNTS_KEY]
224
300
  cc_stack.pop if cc_stack.is_a?(Array) && cc_stack.any?
225
301
 
302
+ # Transaction-control breadcrumbs are popped here, as part of this method's
303
+ # existing lifecycle, rather than through a separate stop_transaction_tracking
304
+ # call. start_request_tracking has many callers across this gem (and its own
305
+ # specs) that only know about query tracking; requiring every one of them to
306
+ # also remember a second, separately-paired stop call is exactly how a frame
307
+ # gets pushed and never popped — leaking across every later request/job on
308
+ # the same thread. Piggybacking on stop_request_tracking, which is already
309
+ # reliably paired with every start_request_tracking call, means there is
310
+ # nothing new for any caller to remember. A caller that wants the events
311
+ # reads last_transaction_events immediately afterward.
312
+ txn_stack = Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY]
313
+ Thread.current[:dead_bro_last_transaction_events] =
314
+ (txn_stack.is_a?(Array) && txn_stack.any?) ? txn_stack.pop : []
315
+ Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY] = nil if txn_stack.nil? || txn_stack.empty?
316
+
226
317
  # Fold any completed EXPLAIN plans from raw queries into their aggregate entry
227
318
  raw_queries.each do |q|
228
319
  next unless q[:explain_plan]
@@ -244,6 +335,17 @@ module DeadBro
244
335
  aggregates_h.values.sort_by { |a| -a[:total_duration_ms] }
245
336
  end
246
337
 
338
+ # The transaction-control breadcrumbs captured during the tracking window that
339
+ # this thread's most recent stop_request_tracking call just ended. Consumes
340
+ # (clears) the thread-local on read, not just on write: otherwise it would sit
341
+ # pinned in Thread.current between requests (a Puma/Sidekiq thread is reused),
342
+ # and — worse — any caller that reads it without an immediately-preceding
343
+ # stop_request_tracking would silently get the *previous* request's
344
+ # breadcrumbs instead of an empty result.
345
+ def self.last_transaction_events
346
+ Thread.current[:dead_bro_last_transaction_events].tap { Thread.current[:dead_bro_last_transaction_events] = nil } || []
347
+ end
348
+
247
349
  # Upper bound on pending EXPLAIN threads per request. Each thread checks
248
350
  # out an AR pool connection, so this must stay well below common pool
249
351
  # sizes (Rails default is 5) to avoid starving the app under a storm.
@@ -304,7 +406,7 @@ module DeadBro
304
406
  EXPLAIN_DML_KEYWORD_RE = /\b(INSERT|UPDATE|DELETE|MERGE)\b/i.freeze
305
407
 
306
408
  def self.should_explain_query?(duration_ms, sql)
307
- return false unless DeadBro.configuration.explain_analyze_active?
409
+ return false unless DeadBro.configuration.explain_active?
308
410
  return false if duration_ms < DeadBro.configuration.slow_query_threshold_ms
309
411
  return false unless sql.is_a?(String)
310
412
 
@@ -94,6 +94,7 @@ module DeadBro
94
94
  if defined?(DeadBro::SqlSubscriber)
95
95
  Thread.current[:dead_bro_sql_queries]
96
96
  Thread.current[:dead_bro_sql_queries] = nil
97
+ Thread.current[DeadBro::SqlSubscriber::THREAD_LOCAL_TXN_EVENTS_KEY] = nil
97
98
  end
98
99
 
99
100
  if defined?(DeadBro::CacheSubscriber)
@@ -58,6 +58,7 @@ module DeadBro
58
58
 
59
59
  # Stop SQL tracking and get collected queries (this was started by the request)
60
60
  sql_queries = DeadBro::SqlSubscriber.stop_request_tracking
61
+ transaction_events = DeadBro::SqlSubscriber.last_transaction_events
61
62
 
62
63
  # Stop cache, redis, and elasticsearch tracking
63
64
  cache_events = defined?(DeadBro::CacheSubscriber) ? DeadBro::CacheSubscriber.stop_request_tracking : []
@@ -147,6 +148,42 @@ module DeadBro
147
148
  })
148
149
  end
149
150
 
151
+ # Everything captured during the request regardless of outcome — shared between
152
+ # the error and success payloads below so an errored request (e.g. a lock wait
153
+ # timeout raised mid-transaction) still ships the SQL/cache/memory/etc data that
154
+ # led up to it, not just the exception. This used to be built only for the
155
+ # success path, so every errored request shipped with sql_count: 0 and no trace
156
+ # even though the gem had already captured it.
157
+ detail_fields = {
158
+ view_runtime_ms: data[:view_runtime],
159
+ db_runtime_ms: data[:db_runtime],
160
+ memory_usage: memory_usage_mb,
161
+ gc_stats: gc_stats,
162
+ sql_count: sql_count(data),
163
+ sql_queries: sql_queries,
164
+ transaction_events: transaction_events,
165
+ http_outgoing: Thread.current[:dead_bro_http_events] || [],
166
+ cache_events: cache_events,
167
+ redis_events: redis_events,
168
+ elasticsearch_events: elasticsearch_events,
169
+ cache_hits: cache_hits(data),
170
+ cache_misses: cache_misses(data),
171
+ view_events: view_events,
172
+ view_performance: view_performance,
173
+ memory_events: memory_events,
174
+ memory_performance: memory_performance,
175
+ allocation_phases: allocation_phases,
176
+ rack_duration_ms: rack_duration_ms,
177
+ queue_duration_ms: Thread.current[:dead_bro_queue_duration_ms],
178
+ db_connection_wait_ms: db_connection_stats[:wait_ms],
179
+ db_connection_checkouts: db_connection_stats[:checkouts],
180
+ gc_pressure: gc_pressure,
181
+ ar_instantiation_count: ar_instantiation_count,
182
+ cpu_time_ms: cpu_time_ms,
183
+ watch_events: watch_events,
184
+ logs: DeadBro.logger.logs
185
+ }
186
+
150
187
  # Report exceptions attached to this action (e.g. controller/view errors)
151
188
  if data[:exception] || data[:exception_object]
152
189
  begin
@@ -154,7 +191,7 @@ module DeadBro
154
191
  exception_obj = data[:exception_object]
155
192
  backtrace = Array(exception_obj&.backtrace).first(50)
156
193
 
157
- error_payload = {
194
+ error_payload = detail_fields.merge(
158
195
  controller: data[:controller],
159
196
  action: data[:action],
160
197
  format: data[:format],
@@ -174,9 +211,8 @@ module DeadBro
174
211
  backtrace: backtrace,
175
212
  fingerprint: compute_error_fingerprint(exception_obj),
176
213
  cause_chain: build_cause_chain(exception_obj),
177
- error: true,
178
- logs: DeadBro.logger.logs
179
- }
214
+ error: true
215
+ )
180
216
 
181
217
  event_name = (exception_class || exception_obj&.class&.name || "exception").to_s
182
218
  client.post_metric(event_name: event_name, payload: error_payload, force: true)
@@ -186,7 +222,7 @@ module DeadBro
186
222
  end
187
223
  end
188
224
 
189
- payload = {
225
+ payload = detail_fields.merge(
190
226
  controller: data[:controller],
191
227
  action: data[:action],
192
228
  format: data[:format],
@@ -195,40 +231,14 @@ module DeadBro
195
231
  status: data[:status],
196
232
  started_at: started.utc.iso8601(3),
197
233
  duration_ms: duration_ms,
198
- view_runtime_ms: data[:view_runtime],
199
- db_runtime_ms: data[:db_runtime],
200
234
  host: DeadBro.safe_hostname,
201
235
  request_host: safe_request_host(data),
202
236
  rails_env: DeadBro.env,
203
237
  process_kind: DeadBro.process_kind,
204
238
  params: safe_params(data),
205
239
  user_agent: safe_user_agent(data),
206
- user_id: extract_user_id(data),
207
- memory_usage: memory_usage_mb,
208
- gc_stats: gc_stats,
209
- sql_count: sql_count(data),
210
- sql_queries: sql_queries,
211
- http_outgoing: Thread.current[:dead_bro_http_events] || [],
212
- cache_events: cache_events,
213
- redis_events: redis_events,
214
- elasticsearch_events: elasticsearch_events,
215
- cache_hits: cache_hits(data),
216
- cache_misses: cache_misses(data),
217
- view_events: view_events,
218
- view_performance: view_performance,
219
- memory_events: memory_events,
220
- memory_performance: memory_performance,
221
- allocation_phases: allocation_phases,
222
- rack_duration_ms: rack_duration_ms,
223
- queue_duration_ms: Thread.current[:dead_bro_queue_duration_ms],
224
- db_connection_wait_ms: db_connection_stats[:wait_ms],
225
- db_connection_checkouts: db_connection_stats[:checkouts],
226
- gc_pressure: gc_pressure,
227
- ar_instantiation_count: ar_instantiation_count,
228
- cpu_time_ms: cpu_time_ms,
229
- watch_events: watch_events,
230
- logs: DeadBro.logger.logs
231
- }
240
+ user_id: extract_user_id(data)
241
+ )
232
242
  # force: true — the sampling decision (global or per-request-type) was
233
243
  # already made above; client#post_metric must not re-roll it with the
234
244
  # global-only rate, which would silently override a per-type sample rate.
@@ -242,6 +252,7 @@ module DeadBro
242
252
  def self.drain_request_tracking
243
253
  # wait_for_explains: false — the result is discarded, so don't block this
244
254
  # thread waiting on pending EXPLAIN plans.
255
+ # stop_request_tracking also pops the transaction-events stack (see its comment).
245
256
  DeadBro::SqlSubscriber.stop_request_tracking(wait_for_explains: false) if defined?(DeadBro::SqlSubscriber)
246
257
  DeadBro::CacheSubscriber.stop_request_tracking if defined?(DeadBro::CacheSubscriber)
247
258
  DeadBro::RedisSubscriber.stop_request_tracking if defined?(DeadBro::RedisSubscriber)
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module DeadBro
4
- VERSION = "0.2.31"
4
+ VERSION = "0.2.33"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: dead_bro
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.2.31
4
+ version: 0.2.33
5
5
  platform: ruby
6
6
  authors:
7
7
  - Emanuel Comsa
@@ -18,6 +18,7 @@ extensions: []
18
18
  extra_rdoc_files: []
19
19
  files:
20
20
  - CHANGELOG.md
21
+ - FEATURES.md
21
22
  - README.md
22
23
  - lib/dead_bro.rb
23
24
  - lib/dead_bro/allocation_source_sampler.rb
@@ -80,7 +81,7 @@ required_rubygems_version: !ruby/object:Gem::Requirement
80
81
  - !ruby/object:Gem::Version
81
82
  version: '0'
82
83
  requirements: []
83
- rubygems_version: 4.0.10
84
+ rubygems_version: 4.0.16
84
85
  specification_version: 4
85
86
  summary: Minimal APM for Rails apps.
86
87
  test_files: []