fiber-profiler 0.7.0 → 0.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- checksums.yaml.gz.sig +3 -4
- data/context/capture-mode.md +80 -0
- data/context/getting-started.md +14 -119
- data/context/index.yaml +10 -1
- data/context/watchdog-mode.md +100 -0
- data/lib/fiber/profiler/version.rb +1 -1
- data/readme.md +5 -1
- data.tar.gz.sig +0 -0
- metadata +3 -1
- metadata.gz.sig +0 -0
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: e3d1446f213526360f5368bfb8a08884033484f4ab868160340709652349c25c
|
|
4
|
+
data.tar.gz: 844c5319c20f8b0929342978790efe2d1aeb4dd37ae260cd073ff36c82c730fd
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: '08e564d3dc9c9abfba57056f3ff5a7e0fce516a5969ee014820ff2dac72b60fd51603c258c21cbee80c46d686f9ca9bbd2376a11b1202635371d0201b8a3b372'
|
|
7
|
+
data.tar.gz: 88d96fba98c5b9578683ad7f13cb898d70e8da83ce11a3b756b15d29b4cbc29649a61c7ec2d589cb9a1185e9646bee4ed018bbcbb1946defdd2e7163ae951aa8
|
checksums.yaml.gz.sig
CHANGED
|
@@ -1,4 +1,3 @@
|
|
|
1
|
-
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
2�^���@I|���M�IiX����B�˹o]z�[�evs�K�f�m�@��*�X,���NL��$���b]4�
|
|
1
|
+
ɫq��}���ƃ�Gp�<�~O.�EgG}y������%ω�`h+���Ɉ[3�%z�K}U���2&�5J��T<.�`��
|
|
2
|
+
�����ZaD�,S��p���K��}���F�]ި���H@�e�b�&ִ� Q��X��.4EᱟT�Ş���-jxz���fq��w^���|�l�l��YF�-��2�*�H�"~4��K�/��z��-��]�w4�G���G��Y��0+���~&��gDta|�o�Zv�.jP��
|
|
3
|
+
>�@�q�C���a�pD�fֹ�nz
|
|
@@ -0,0 +1,80 @@
|
|
|
1
|
+
# Capture Mode
|
|
2
|
+
|
|
3
|
+
This guide explains how to trace fiber execution and analyze call timings with the capture profiler.
|
|
4
|
+
|
|
5
|
+
Use capture mode when you need to investigate the calls made during a stall. It records call timings for sampled fiber executions and reports after the fiber switches away. For sampled backtraces while a stall is still in progress, see [Watchdog Mode](../watchdog-mode/index).
|
|
6
|
+
|
|
7
|
+
## Usage
|
|
8
|
+
|
|
9
|
+
Add `fiber-profiler` to your application's bundle as described in [Getting Started](../getting-started/index), then select capture mode when starting Async or Falcon:
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
FIBER_PROFILER=capture bundle exec falcon serve
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Async starts and stops capture automatically. The legacy `FIBER_PROFILER_CAPTURE=true` setting also enables capture when `FIBER_PROFILER` is unset. An explicit `FIBER_PROFILER` value takes precedence.
|
|
16
|
+
|
|
17
|
+
### Manual Instrumentation
|
|
18
|
+
|
|
19
|
+
To reproduce a stall without a scheduler, save this as `capture.rb`. It starts a {ruby Fiber::Profiler::Capture} and simulates a blocking operation inside a fiber:
|
|
20
|
+
|
|
21
|
+
```ruby
|
|
22
|
+
require "fiber/profiler"
|
|
23
|
+
|
|
24
|
+
profiler = Fiber::Profiler::Capture.new
|
|
25
|
+
|
|
26
|
+
begin
|
|
27
|
+
profiler.start
|
|
28
|
+
|
|
29
|
+
# Simulate a blocking operation without a scheduler:
|
|
30
|
+
Fiber.new(blocking: false) do
|
|
31
|
+
sleep 0.1
|
|
32
|
+
end.resume
|
|
33
|
+
ensure
|
|
34
|
+
profiler.stop
|
|
35
|
+
end
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
bundle exec ruby capture.rb
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
This example starts the profiler explicitly, so it does not need `FIBER_PROFILER=capture`. The report should include `Kernel#sleep` and its elapsed duration. Start and stop profiling on the same thread.
|
|
43
|
+
|
|
44
|
+
## Configuration
|
|
45
|
+
|
|
46
|
+
Set these environment variables before loading `fiber-profiler`; the native extension reads them when it loads:
|
|
47
|
+
|
|
48
|
+
| Variable | Default | Meaning |
|
|
49
|
+
| --- | --- | --- |
|
|
50
|
+
| `FIBER_PROFILER_CAPTURE_STALL_THRESHOLD` | `0.01` | Minimum execution duration in seconds to exceed before reporting a stall. |
|
|
51
|
+
| `FIBER_PROFILER_CAPTURE_FILTER_THRESHOLD` | 10% of the stall threshold | Filter calls shorter than this duration in seconds. |
|
|
52
|
+
| `FIBER_PROFILER_CAPTURE_TRACK_CALLS` | `true` | Record call timings. Set to `false` to report stall durations without tracing calls. |
|
|
53
|
+
| `FIBER_PROFILER_CAPTURE_SAMPLE_RATE` | `1.0` | Fraction of eligible fiber executions to sample: `1.0` samples all, `0.1` samples approximately 10%. |
|
|
54
|
+
|
|
55
|
+
These settings apply to capture mode only. The sample rate controls selection at fiber switches; it is not a time interval between backtrace samples.
|
|
56
|
+
|
|
57
|
+
For explicit instrumentation, override the defaults with `Fiber::Profiler::Capture.new(stall_threshold: 0.05, filter_threshold: 0.005, track_calls: true, sample_rate: 0.1)`. Pass `output:` to select a writable IO; the default is standard error.
|
|
58
|
+
|
|
59
|
+
## Reading Reports
|
|
60
|
+
|
|
61
|
+
When standard error is a terminal, capture prints a readable call log. Redirected output uses one JSON object per stall, including the execution `duration` and a `calls` array. Each retained call includes its source location, class, method, duration, and nesting information.
|
|
62
|
+
|
|
63
|
+
To collect the example's reports and summarize call timings:
|
|
64
|
+
|
|
65
|
+
```bash
|
|
66
|
+
bundle exec ruby capture.rb 2> capture.ndjson
|
|
67
|
+
bundle exec bake input --file capture.ndjson fiber:profiler:analyze output
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
The analyzer aggregates call durations by source location and sorts the summary by total duration. Feed it capture reports; unrelated application output on standard error must be separated from the JSON first. With `track_calls: false`, reports contain no call timings to aggregate.
|
|
71
|
+
|
|
72
|
+
Call durations can include time spent in nested calls, so adding durations across different locations does not give total application runtime. Short or uninformative calls may be filtered from the report.
|
|
73
|
+
|
|
74
|
+
## Interpretation and Limits
|
|
75
|
+
|
|
76
|
+
Capture samples non-blocking fibers and excludes blocking fibers, including the event-loop fiber. It measures wall time between fiber switches, so an ordinary scheduler-aware wait that yields does not count its entire wait time as a stall. A blocking operation that keeps the same fiber executing can count toward a stall.
|
|
77
|
+
|
|
78
|
+
Reports are emitted when the fiber switches away. An operation that never yields or returns can therefore prevent its capture report from appearing. Use [Watchdog Mode](../watchdog-mode/index) to investigate an ongoing stall when Ruby thread scheduling is still possible.
|
|
79
|
+
|
|
80
|
+
Call tracing can substantially affect performance. Reduce the sample rate to trace fewer executions, or disable call tracking if you only need stall durations. Use the [Falcon overhead benchmark](https://github.com/socketry/fiber-profiler/tree/main/benchmark/falcon) to compare modes, and measure with your application's Ruby/JIT configuration and workload.
|
data/context/getting-started.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Getting Started
|
|
2
2
|
|
|
3
|
-
This guide explains how to
|
|
3
|
+
This guide explains how to install the fiber profiler and choose a mode for diagnosing event-loop stalls.
|
|
4
4
|
|
|
5
5
|
## Installation
|
|
6
6
|
|
|
@@ -10,7 +10,18 @@ Add the gem to your project:
|
|
|
10
10
|
$ bundle add fiber-profiler
|
|
11
11
|
```
|
|
12
12
|
|
|
13
|
-
##
|
|
13
|
+
## Choose a Mode
|
|
14
|
+
|
|
15
|
+
A fiber that runs for too long without yielding prevents other work on the same event loop from progressing. The profiler offers two ways to investigate it:
|
|
16
|
+
|
|
17
|
+
| Mode | Use it to | Reports |
|
|
18
|
+
| --- | --- | --- |
|
|
19
|
+
| [Watchdog](../watchdog-mode/index) | Find where a fiber is spending time during an ongoing stall. | Periodically sampled backtraces while the fiber is still executing. |
|
|
20
|
+
| [Capture](../capture-mode/index) | Investigate individual call timings during a fiber's execution. | Traced calls and their durations after the fiber switches away. |
|
|
21
|
+
|
|
22
|
+
Watchdog avoids method-call tracing and is a useful starting point for observing sustained stalls. Capture provides more detail at a higher instrumentation cost. Both affect the application being measured; see the [Falcon overhead benchmark](https://github.com/socketry/fiber-profiler/tree/main/benchmark/falcon) and measure with your own workload.
|
|
23
|
+
|
|
24
|
+
## Integration with Async and Falcon
|
|
14
25
|
|
|
15
26
|
Select the profiling mode when starting your application:
|
|
16
27
|
|
|
@@ -29,120 +40,4 @@ With Async, including Falcon, the scheduler starts and stops the selected profil
|
|
|
29
40
|
|
|
30
41
|
Unknown modes raise `ArgumentError`. Set environment variables before loading the gem: the native capture settings are read when the extension loads. The new mode selector and watchdog settings are read when `Fiber::Profiler.default` is called.
|
|
31
42
|
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
Instrument your code using the default profiler:
|
|
35
|
-
|
|
36
|
-
```ruby
|
|
37
|
-
#!/usr/bin/env ruby
|
|
38
|
-
|
|
39
|
-
require "fiber/profiler"
|
|
40
|
-
|
|
41
|
-
profiler = Fiber::Profiler.default
|
|
42
|
-
|
|
43
|
-
begin
|
|
44
|
-
profiler&.start
|
|
45
|
-
|
|
46
|
-
# Your application code:
|
|
47
|
-
Fiber.new do
|
|
48
|
-
sleep 0.1
|
|
49
|
-
end.resume
|
|
50
|
-
ensure
|
|
51
|
-
profiler&.stop
|
|
52
|
-
end
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
Running this program will output the following:
|
|
56
|
-
|
|
57
|
-
```bash
|
|
58
|
-
$ FIBER_PROFILER_CAPTURE=true bundle exec ./test.rb
|
|
59
|
-
Fiber stalled for 0.105 seconds
|
|
60
|
-
/Users/samuel/Developer/socketry/fiber-profiler/test.rb:11 in c-call 'Kernel#sleep' (0.105s)
|
|
61
|
-
```
|
|
62
|
-
|
|
63
|
-
## Integration with Async
|
|
64
|
-
|
|
65
|
-
The fiber profiler is optionally supported by `Async`. Enable either mode using `FIBER_PROFILER=watchdog` or `FIBER_PROFILER=capture`. The legacy `FIBER_PROFILER_CAPTURE=true` setting continues to enable capture mode when `FIBER_PROFILER` is unset.
|
|
66
|
-
|
|
67
|
-
Each scheduler obtains its own profiler. Watchdog binds to the thread when profiling starts, so schedulers in separate worker processes or threads are monitored independently. After `fork`, inherited profiling is stopped in the child; a new scheduler starts a new profiler normally.
|
|
68
|
-
|
|
69
|
-
## Watchdog Mode
|
|
70
|
-
|
|
71
|
-
Watchdog uses a thread-specific `:fiber_switch` tracepoint to track which fiber is executing. A separate Ruby thread periodically captures the monitored thread's backtrace. It does not install method-call tracing hooks or require a heartbeat task.
|
|
72
|
-
|
|
73
|
-
By default, it samples every 100 ms, retains at most five recent stacks, and reports when the same fiber execution lasts at least 500 ms. A continued stall produces at most one report per additional threshold interval. Returning to the event loop or switching to another fiber resets the sample window. This measures individual uninterrupted fiber executions, rather than starvation of a scheduled heartbeat.
|
|
74
|
-
|
|
75
|
-
To reproduce a stall, save this as `stall.rb`:
|
|
76
|
-
|
|
77
|
-
```ruby
|
|
78
|
-
require "async"
|
|
79
|
-
|
|
80
|
-
Sync do
|
|
81
|
-
deadline = Process.clock_gettime(Process::CLOCK_MONOTONIC) + 2
|
|
82
|
-
while Process.clock_gettime(Process::CLOCK_MONOTONIC) < deadline
|
|
83
|
-
end
|
|
84
|
-
end
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
```bash
|
|
88
|
-
FIBER_PROFILER=watchdog bundle exec ruby stall.rb
|
|
89
|
-
```
|
|
90
|
-
|
|
91
|
-
The reports should point to the busy loop. Replacing the loop with `sleep(2)` allows Async to keep scheduling fibers and should produce no reports.
|
|
92
|
-
|
|
93
|
-
Terminal output contains readable stacks. Redirected output uses one JSON object per report, with `mode: "watchdog"`, process/thread/fiber identifiers, elapsed `duration`, and a bounded `samples` array. Each sample contains its elapsed time and `backtrace`. Duration is elapsed wall time since the fiber resumed, and a report is emitted while the fiber is still executing.
|
|
94
|
-
|
|
95
|
-
### Configuration
|
|
96
|
-
|
|
97
|
-
| Variable | Default | Meaning |
|
|
98
|
-
| --- | --- | --- |
|
|
99
|
-
| `FIBER_PROFILER_WATCHDOG_STALL_THRESHOLD` | `0.5` | Minimum execution duration in seconds before reporting, and minimum interval between repeated reports. |
|
|
100
|
-
| `FIBER_PROFILER_WATCHDOG_SAMPLE_INTERVAL` | `0.1` | Delay in seconds between stack samples. |
|
|
101
|
-
|
|
102
|
-
Both values must be finite and positive. Actual sampling intervals depend on Ruby and operating-system scheduling. Capture-specific settings, including its sample rate, do not configure watchdog mode.
|
|
103
|
-
|
|
104
|
-
For explicit instrumentation, use `Fiber::Profiler::Watchdog.new(stall_threshold: 0.5, sample_interval: 0.1, max_samples: 5, output: $stderr)` and the same `start`/`stop` lifecycle as capture. The output must support writes from the watchdog thread. Stop the profiler before closing its output.
|
|
105
|
-
|
|
106
|
-
If sampling or reporting raises a `StandardError`, the watchdog disables its tracepoint and stops sampling. It retains the exception in `watchdog.error` and attempts one warning to standard error. Failure to write that warning is also contained. These errors do not escape through `stop` or replace application errors. Call `stop` normally to finish cleanup; a subsequent `start` clears the error and resumes monitoring.
|
|
107
|
-
|
|
108
|
-
### Interpretation and Limits
|
|
109
|
-
|
|
110
|
-
Repeated frames identify code worth investigating: optimize expensive work, add cooperative yield points, or offload suitable work to a bounded thread pool. Sampling still has overhead; measure it for your workload.
|
|
111
|
-
|
|
112
|
-
The repository includes a [Falcon overhead benchmark](https://github.com/socketry/fiber-profiler/tree/main/benchmark/falcon) comparing disabled, watchdog, and capture modes. It measures throughput, latency, server CPU time, and Ruby allocations across repeated runs. Use it as a starting point for measuring your own application before enabling continuous monitoring.
|
|
113
|
-
|
|
114
|
-
The watchdog requires Ruby thread scheduling. Native code that holds the GVL without allowing other threads to run can prevent sampling entirely. Short stalls can occur between samples, and normal scheduler/OS delays can extend measured execution time. A report is a diagnostic lead, not proof that every sampled frame is expensive.
|
|
115
|
-
|
|
116
|
-
When profiling starts on a blocking event-loop fiber, that fiber is excluded from monitoring, including its normal idle waits. Other fibers, including blocking application fibers, are monitored. When profiling starts inside a non-blocking application fiber, blocking fibers are excluded because the event-loop fiber is not known. Start and stop profiling on the monitored thread.
|
|
117
|
-
|
|
118
|
-
## Default Environment Variables
|
|
119
|
-
|
|
120
|
-
The following settings apply to capture mode only.
|
|
121
|
-
|
|
122
|
-
### `FIBER_PROFILER_CAPTURE`
|
|
123
|
-
|
|
124
|
-
Set to `true` to enable capturing of stalled fibers.
|
|
125
|
-
|
|
126
|
-
### `FIBER_PROFILER_CAPTURE_STALL_THRESHOLD`
|
|
127
|
-
|
|
128
|
-
Set the threshold in seconds for reporting a stalled fiber. Default is `0.01`.
|
|
129
|
-
|
|
130
|
-
### `FIBER_PROFILER_CAPTURE_TRACK_CALLS`
|
|
131
|
-
|
|
132
|
-
Set to `true` to track calls within the fiber. Default is `true`. This can be disabled to reduce overhead.
|
|
133
|
-
|
|
134
|
-
### `FIBER_PROFILER_CAPTURE_SAMPLE_RATE`
|
|
135
|
-
|
|
136
|
-
Set the sample rate of the profiler as a percentage of all context switches. The default is 1.0 (100%).
|
|
137
|
-
|
|
138
|
-
## Analyzing Logs
|
|
139
|
-
|
|
140
|
-
If you collect your logs in a file (e.g. as `ndjson`) you can analyze them using the included `bake` commands:
|
|
141
|
-
|
|
142
|
-
```bash
|
|
143
|
-
$ bundle exec bake input --file samples.ndjson fiber:profiler:analyze output
|
|
144
|
-
```
|
|
145
|
-
|
|
146
|
-
This will aggregate all the call logs and generate a short summary, ordered by duration.
|
|
147
|
-
|
|
148
|
-
The analyzer consumes capture-mode call timings. Watchdog reports contain sampled backtraces instead of call timings and are not aggregated by this command.
|
|
43
|
+
Each scheduler obtains its own profiler, so schedulers in separate worker processes or threads are monitored independently. After `fork`, inherited profiling is stopped in the child; a new scheduler starts a new profiler normally.
|
data/context/index.yaml
CHANGED
|
@@ -8,4 +8,13 @@ metadata:
|
|
|
8
8
|
files:
|
|
9
9
|
- path: getting-started.md
|
|
10
10
|
title: Getting Started
|
|
11
|
-
description: This guide explains how to
|
|
11
|
+
description: This guide explains how to install the fiber profiler and choose a
|
|
12
|
+
mode for diagnosing event-loop stalls.
|
|
13
|
+
- path: watchdog-mode.md
|
|
14
|
+
title: Watchdog Mode
|
|
15
|
+
description: This guide explains how to sample ongoing fiber stalls with the watchdog
|
|
16
|
+
profiler.
|
|
17
|
+
- path: capture-mode.md
|
|
18
|
+
title: Capture Mode
|
|
19
|
+
description: This guide explains how to trace fiber execution and analyze call timings
|
|
20
|
+
with the capture profiler.
|
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Watchdog Mode
|
|
2
|
+
|
|
3
|
+
This guide explains how to sample ongoing fiber stalls with the watchdog profiler.
|
|
4
|
+
|
|
5
|
+
Use watchdog mode to find where a fiber is spending time while it prevents other fibers from running. It reports during the stall, so you can investigate an operation that has not returned yet. For individual call timings, see [Capture Mode](../capture-mode/index).
|
|
6
|
+
|
|
7
|
+
## Usage
|
|
8
|
+
|
|
9
|
+
Add `fiber-profiler` to your application's bundle as described in [Getting Started](../getting-started/index), then select watchdog mode when starting Async or Falcon:
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
FIBER_PROFILER=watchdog bundle exec falcon serve
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Async starts and stops the watchdog automatically. Each scheduler gets its own profiler, which binds to the thread when profiling starts.
|
|
16
|
+
|
|
17
|
+
To reproduce a stall, save this as `stall.rb`:
|
|
18
|
+
|
|
19
|
+
```ruby
|
|
20
|
+
require "async"
|
|
21
|
+
|
|
22
|
+
Sync do
|
|
23
|
+
deadline = Process.clock_gettime(Process::CLOCK_MONOTONIC) + 2
|
|
24
|
+
while Process.clock_gettime(Process::CLOCK_MONOTONIC) < deadline
|
|
25
|
+
end
|
|
26
|
+
end
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
FIBER_PROFILER=watchdog bundle exec ruby stall.rb
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
The reports should point to the busy loop. Replacing the loop with `sleep(2)` allows Async to keep scheduling fibers and should produce no reports.
|
|
34
|
+
|
|
35
|
+
## How Sampling Works
|
|
36
|
+
|
|
37
|
+
Watchdog uses a thread-specific `:fiber_switch` tracepoint to track which fiber is executing. A separate Ruby thread periodically captures the monitored thread's backtrace. It does not install method-call tracing hooks or require a heartbeat task.
|
|
38
|
+
|
|
39
|
+
By default, it samples every 100 ms, retains at most five recent stacks, and reports when the same fiber execution lasts at least 500 ms. A continued stall produces at most one report per additional threshold interval. Returning to the event loop or switching to another fiber resets the sample window. The measured duration covers one uninterrupted fiber execution.
|
|
40
|
+
|
|
41
|
+
## Configuration
|
|
42
|
+
|
|
43
|
+
Set these environment variables before starting your application. They are read when the default watchdog is constructed:
|
|
44
|
+
|
|
45
|
+
| Variable | Default | Meaning |
|
|
46
|
+
| --- | --- | --- |
|
|
47
|
+
| `FIBER_PROFILER_WATCHDOG_STALL_THRESHOLD` | `0.5` | Minimum execution duration in seconds before reporting, and minimum interval between repeated reports. |
|
|
48
|
+
| `FIBER_PROFILER_WATCHDOG_SAMPLE_INTERVAL` | `0.1` | Delay in seconds between stack samples. |
|
|
49
|
+
|
|
50
|
+
Both values must be finite and positive. Actual sampling intervals depend on Ruby and operating-system scheduling. Capture-specific settings, including its sample rate, do not configure watchdog mode.
|
|
51
|
+
|
|
52
|
+
### Manual Instrumentation
|
|
53
|
+
|
|
54
|
+
When managing fibers directly, construct a {ruby Fiber::Profiler::Watchdog} and stop it in an `ensure` block:
|
|
55
|
+
|
|
56
|
+
```ruby
|
|
57
|
+
require "fiber/profiler/watchdog"
|
|
58
|
+
|
|
59
|
+
watchdog = Fiber::Profiler::Watchdog.new(
|
|
60
|
+
stall_threshold: 0.5,
|
|
61
|
+
sample_interval: 0.1,
|
|
62
|
+
max_samples: 5,
|
|
63
|
+
output: $stderr
|
|
64
|
+
)
|
|
65
|
+
|
|
66
|
+
begin
|
|
67
|
+
watchdog.start
|
|
68
|
+
|
|
69
|
+
# Simulate an application fiber blocked on an operation:
|
|
70
|
+
Fiber.new do
|
|
71
|
+
sleep 1
|
|
72
|
+
end.resume
|
|
73
|
+
ensure
|
|
74
|
+
watchdog.stop
|
|
75
|
+
end
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
This example starts the profiler explicitly, so it does not need `FIBER_PROFILER=watchdog`. Start and stop profiling on the monitored thread. `max_samples` must be a positive integer; there is no environment variable for it. The output must support writes from the watchdog thread. Stop the profiler before closing its output.
|
|
79
|
+
|
|
80
|
+
When profiling starts on a blocking event-loop fiber, that fiber is excluded from monitoring, including its normal idle waits. Other fibers, including blocking application fibers, are monitored. When profiling starts inside a non-blocking application fiber, blocking fibers are excluded because the event-loop fiber is not known.
|
|
81
|
+
|
|
82
|
+
If sampling or reporting raises a `StandardError`, the watchdog disables its tracepoint and stops sampling. It retains the exception in `watchdog.error` and attempts one warning to standard error. Failure to write that warning is also contained. These errors do not escape through `stop` or replace application errors. Call `stop` normally to finish cleanup; a subsequent `start` clears the error and resumes monitoring.
|
|
83
|
+
|
|
84
|
+
## Reading Reports
|
|
85
|
+
|
|
86
|
+
Terminal output contains readable stacks. Redirected output uses one JSON object per report:
|
|
87
|
+
|
|
88
|
+
```bash
|
|
89
|
+
FIBER_PROFILER=watchdog bundle exec ruby stall.rb 2> watchdog.ndjson
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
Each report contains `mode: "watchdog"`, process/thread/fiber identifiers, elapsed `duration`, and a bounded `samples` array. Each sample contains its elapsed time and `backtrace`. Duration is elapsed wall time since the fiber resumed, and a report is emitted while the fiber is still executing.
|
|
93
|
+
|
|
94
|
+
Repeated frames identify code worth investigating: optimize expensive work, add cooperative yield points, or offload suitable work to a bounded thread pool. Watchdog reports contain sampled backtraces, so the capture-mode `fiber:profiler:analyze` command does not aggregate them.
|
|
95
|
+
|
|
96
|
+
## Interpretation and Limits
|
|
97
|
+
|
|
98
|
+
The watchdog requires Ruby thread scheduling. Native code that holds the GVL without allowing other threads to run can prevent sampling entirely. Short stalls can occur between samples, and normal scheduler/OS delays can extend measured execution time. A report is a diagnostic lead, not proof that every sampled frame is expensive.
|
|
99
|
+
|
|
100
|
+
Sampling still has overhead. Use the [Falcon overhead benchmark](https://github.com/socketry/fiber-profiler/tree/main/benchmark/falcon) as a starting point for measuring your own application before enabling continuous monitoring.
|
data/readme.md
CHANGED
|
@@ -12,7 +12,11 @@ Migrating existing applications to the event loop can be tricky. One of the most
|
|
|
12
12
|
|
|
13
13
|
Please see the [project documentation](https://socketry.github.io/fiber-profiler/) for more details.
|
|
14
14
|
|
|
15
|
-
- [Getting Started](https://socketry.github.io/fiber-profiler/guides/getting-started/index) - This guide explains how to
|
|
15
|
+
- [Getting Started](https://socketry.github.io/fiber-profiler/guides/getting-started/index) - This guide explains how to install the fiber profiler and choose a mode for diagnosing event-loop stalls.
|
|
16
|
+
|
|
17
|
+
- [Watchdog Mode](https://socketry.github.io/fiber-profiler/guides/watchdog-mode/index) - This guide explains how to sample ongoing fiber stalls with the watchdog profiler.
|
|
18
|
+
|
|
19
|
+
- [Capture Mode](https://socketry.github.io/fiber-profiler/guides/capture-mode/index) - This guide explains how to trace fiber execution and analyze call timings with the capture profiler.
|
|
16
20
|
|
|
17
21
|
## Releases
|
|
18
22
|
|
data.tar.gz.sig
CHANGED
|
Binary file
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: fiber-profiler
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.7.
|
|
4
|
+
version: 0.7.1
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Samuel Williams
|
|
@@ -54,8 +54,10 @@ extensions:
|
|
|
54
54
|
extra_rdoc_files: []
|
|
55
55
|
files:
|
|
56
56
|
- bake/fiber/profiler/analyze.rb
|
|
57
|
+
- context/capture-mode.md
|
|
57
58
|
- context/getting-started.md
|
|
58
59
|
- context/index.yaml
|
|
60
|
+
- context/watchdog-mode.md
|
|
59
61
|
- ext/extconf.rb
|
|
60
62
|
- ext/fiber/profiler/capture.c
|
|
61
63
|
- ext/fiber/profiler/capture.h
|
metadata.gz.sig
CHANGED
|
Binary file
|