quonfig 1.4.1 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 8ecc1b6f2a6c327c36d7da4cf091623b8c44b81d6949b887356301961229a684
4
- data.tar.gz: bf357130f83dc29baf032c8dd0db652c8910797aac2fc6a499ec3461e5b56ff6
3
+ metadata.gz: 5be29a29cae00eb43e09f2aadef25191ec4dc9959ff0037cdfd7ba89dfbd5a10
4
+ data.tar.gz: b47703c3beb5a4f85f6b2237b3e4295e8e82055ab105168827e750dbd428d13e
5
5
  SHA512:
6
- metadata.gz: '08951e8e2261dce91f0308de7920ea46eef5077979efdd0bd8df55b403719d8fb1f2e7f6ed863a12d38108b30cdef8b92d08832059a68c45b96425a5ced249f6'
7
- data.tar.gz: 81738455b310e66693801fe2b2aa3ff027e676b28b032398271d1de17c9421212d5e0c9d3b20c37164a7299c74f2347a52f371abfabc36d8ebca3c75492d17d0
6
+ metadata.gz: 7616048fcd0a2a0727ac79a0a801f35d71be44cd54667dc4aaa8aa86009033194f20abbfb69f0b959dc5ba90cdf3e29069db1384e8a432caafe34898f69e3165
7
+ data.tar.gz: 31d412725f88752b8cbde55daa05c99b926212c2951519a177c24dae72b19d665b3452f4f723e5b2b7e78258cdb37e064733e18249316f7b055846408d2eda92
data/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.5.0 - 2026-09-25
4
+
5
+ - **Telemetry transport policy (qfg-y8je.8).** The telemetry POST had no timeout of its own (Faraday's defaults, 60s connect + 60s read); it now has a 15s overall deadline (`telemetry_timeout_ms`) and a 5s connect + TLS deadline (`telemetry_connect_timeout_ms`). A failed batch is kept byte-for-byte and resent (never merged with newer data, so the server dedups a resend of a batch that did land). Resends happen no sooner than 30s after a failure and honor `Retry-After` up to 10 min. The retained queue is capped at 5 batches / 2MB / 5 min (oldest dropped). At most one POST is in flight. 401/403/404 disable telemetry for the process with one ERROR; any other 4xx drops that batch with one ERROR. Before this, a failed batch was simply lost.
6
+ - **Oversize batches are dropped on failure.** A single batch larger than the 2MB byte cap (`telemetry_max_retained_bytes`) is POSTed once and, if that POST fails, dropped rather than kept; the drop counts toward the warning below. This fires for Ruby in practice: in a 24h production sample, sdk-ruby was the only SDK sending large batches (p95 874KB, p99 1.73MB, max 5.2MB, all example-context data) and 0.07% of its POSTs were over 2MB. Most of that size came from the old interval (next item), which let a window grow for up to 10 minutes; at 60s batches should be much smaller. To keep batches small regardless, use `context_upload_mode: :shapes_only` or a lower `context_max_size`.
7
+ - **Flush interval is a fixed 60s (`collect_sync_interval`, now explicitly defaulted).** The old default started at 8s and grew by 1.5x on every tick, including successful ones (a bug), reaching the 600s ceiling about 23 minutes after start; the exponential backoff is gone.
8
+ - **Logging (was a WARN on every failed POST):** a failed POST logs at debug; one WARN when data is actually dropped (then a summary at most every 10 min while drops continue); one INFO on recovery.
9
+ - **Shutdown:** `Client#stop` and the reporter's `at_exit` hook send the live window once with a 5s deadline and do not resend kept batches; an in-flight POST is abandoned. Exit is never blocked longer than that.
10
+ - **Memory caps:** the evaluation-summary, context-shape and example-context caps drop from 100,000 to 10,000 per window (`collect_max_evaluation_summaries`, `context_max_size`), the uniform server-SDK cap. The example-context once-per-hour rate-limit map is bounded at 100,000 keys. A summary key already seen keeps counting at the cap.
11
+ - **New options:** `telemetry_timeout_ms`, `telemetry_connect_timeout_ms`, `telemetry_max_retained_batches`, `telemetry_max_retained_bytes`, `telemetry_max_retained_age_ms`. Invalid values (non-numeric or <= 0) fall back to the default.
12
+ - Internal classes: `Quonfig::HttpConnection` gains `open_timeout_ms:` and sends a String body verbatim. `Quonfig::Telemetry::TelemetryReporter` gains `.build`, `tick`, `flush`, `close` and `debug_state`; `sync` and `stop` remain as aliases of `flush` and `close`. New `Quonfig::Telemetry::TransportQueue`.
13
+ - `context_upload_mode` default is unchanged (`:periodic_example`). No wire change, no removed public API, no new dependencies. Fork safety is unchanged: a forked child still builds its own reporter (and its own retained queue) on first use and never sends the parent's data.
14
+
3
15
  ## 1.4.1 - 2026-09-11
4
16
 
5
17
  - **Fix (fork/telemetry): a forked child now reports under its OWN `instanceHash` (qfg-xcym).** `@instance_hash` was minted once in `Client#initialize` and survived `fork(2)`, so every forked child POSTed its telemetry under the **parent's** hash. The Quonfig Debugger groups SDK last-seen by that hash, so an 8-worker Puma cluster (or a `parallel`-gem batch) collapsed into a single instance row with the parent's and the children's windows interleaved on top of each other. A child now mints a fresh hash as part of the same post-fork rebuild that gives it fresh aggregators and a fresh reporter — before the reporter is constructed, so the hash the child POSTs under is its own. This is parity with Reforge, where a forked process simply builds a whole new client. The parent's hash is never touched, and a child's row count is unchanged (it already reported its own window; it just did so under the wrong identity). Present since 0.0.16.
data/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  Ruby SDK for [Quonfig](https://quonfig.com) — Feature Flags, Live Config, and Dynamic Log Levels.
4
4
 
5
- > **Note:** This SDK is pre-1.0 and the API is not yet stable.
5
+ > **Note:** This SDK is stable (v1) and follows [Semantic Versioning](https://semver.org): breaking changes to the public API land only in a major (`x.0.0`) release.
6
6
 
7
7
  ## Installation
8
8
 
@@ -284,6 +284,16 @@ Quonfig::Client.new(
284
284
  | `data_dir_auto_reload` | `Boolean` | `false` | Datadir mode only. When `true`, the SDK watches the datadir and re-reads the envelope when files change. See [Datadir mode: auto-reload on file changes](#datadir-mode-auto-reload-on-file-changes). |
285
285
  | `data_dir_auto_reload_debounce_ms` | `Integer` (ms) | `200` | Debounce window for the auto-reload watcher — events arriving inside the window are coalesced into a single re-read. Ignored when `data_dir_auto_reload` is `false`. |
286
286
  | `logger` | Logger-like object | `nil` | Optional host-app logger (e.g. `Rails.logger`). Must respond to `debug`/`info`/`warn`/`error`. When set, all SDK warnings/errors flow through this logger instead of the default stderr / SemanticLogger backend. |
287
+ | `collect_evaluation_summaries` | `Boolean` | `true` | Send per-flag evaluation counts. See [Telemetry](#telemetry). |
288
+ | `collect_max_evaluation_summaries` | `Integer` | `10_000` | Distinct flags/configs counted per telemetry window; a key already seen keeps counting at the cap. |
289
+ | `context_upload_mode` | `Symbol` | `:periodic_example` | `:periodic_example` (context shapes + example contexts), `:shapes_only`, or `:none`. |
290
+ | `context_max_size` | `Integer` | `10_000` | Context-shape fields, and separately example contexts, kept per telemetry window. |
291
+ | `collect_sync_interval` | `Numeric` (s) | `60` | Seconds between telemetry POSTs. |
292
+ | `telemetry_timeout_ms` | `Integer` (ms) | `15_000` | Overall deadline for one telemetry POST. |
293
+ | `telemetry_connect_timeout_ms` | `Integer` (ms) | `5_000` | TCP connect + TLS deadline for one telemetry POST. |
294
+ | `telemetry_max_retained_batches` | `Integer` | `5` | Failed telemetry batches kept for resend. |
295
+ | `telemetry_max_retained_bytes` | `Integer` | `2_097_152` | Byte cap (2MB) on kept batches; a single batch larger than this is sent once and never kept. |
296
+ | `telemetry_max_retained_age_ms` | `Integer` (ms) | `300_000` | A kept batch older than this is discarded. |
287
297
 
288
298
  ## Failover & `QUONFIG_DOMAIN`
289
299
 
@@ -634,6 +644,58 @@ process.
634
644
 
635
645
  There is intentionally no `client.healthy?` primitive.
636
646
 
647
+ ## Telemetry
648
+
649
+ With an SDK key the client sends usage telemetry to `telemetry_url` so the
650
+ Quonfig dashboard can show which flags and configs are evaluated and with what
651
+ contexts. Telemetry never affects flag evaluation: every failure below is
652
+ contained in the background reporter thread.
653
+
654
+ **What is sent.** Evaluation summaries (per flag/config: counts per rule and
655
+ value), context shapes (context field names and types), example contexts (up to
656
+ one per context key per hour) and failover counters. Opt out with
657
+ `collect_evaluation_summaries: false` and `context_upload_mode: :shapes_only`
658
+ (no example contexts) or `:none` (no context data). With both off, no reporter
659
+ runs.
660
+
661
+ **How it is sent.**
662
+
663
+ - One POST every `collect_sync_interval` seconds (60), with at most one POST in
664
+ flight. A tick that fires while a POST is still out is skipped and its data
665
+ rolls into the next window.
666
+ - Each POST has an overall deadline of `telemetry_timeout_ms` (15s) and a
667
+ connect + TLS deadline of `telemetry_connect_timeout_ms` (5s).
668
+ - When a POST fails (timeout, network error, 408, 429 or 5xx), the serialized
669
+ batch is kept byte-for-byte and resent unchanged, never merged with newer
670
+ data, so the server can recognize a resend of a batch that did land. Up to 5
671
+ batches / 2MB are kept for up to 5 minutes; beyond that the oldest is
672
+ dropped. A single batch larger than 2MB is sent once and never kept: if that
673
+ one POST fails, the batch is dropped. Large batches come from example
674
+ contexts; `context_upload_mode: :shapes_only` or a lower `context_max_size`
675
+ keeps batches small.
676
+ - Resends happen no sooner than 30s after a failure and after any
677
+ `Retry-After` (honored up to 10 minutes), oldest first, then the current
678
+ window.
679
+ - A 401, 403 or 404 means the SDK key or `telemetry_url` is wrong: the SDK logs
680
+ one error and disables telemetry for the rest of the process. Any other 4xx
681
+ drops that one batch with an error (the server rejected the payload) and
682
+ telemetry continues.
683
+
684
+ **Logging.** A failed POST logs at debug only. The first batch actually dropped
685
+ logs one warning with the last POST result and queue depth; further drops log
686
+ at debug with a summary warning at most every 10 minutes; the first success
687
+ after failures logs one info line. The SDK's default logger prints warnings and
688
+ errors only; pass `logger:` to receive the debug and info lines.
689
+
690
+ **Shutdown.** `stop` (and the `at_exit` hook the reporter registers) sends the
691
+ current window once with a 5s deadline, does not resend kept batches, and never
692
+ blocks process exit longer than that.
693
+
694
+ **Memory.** Everything is bounded: at most 10,000 evaluation-summary keys,
695
+ 10,000 context-shape fields and 10,000 example contexts per window (keys
696
+ already seen keep counting at the cap), a 100,000-entry example-context
697
+ rate-limit map, and the 2MB retained queue.
698
+
637
699
  ## Documentation
638
700
 
639
701
  Full documentation, including SPEC, SDK reference, and operational guides, is
@@ -1123,40 +1123,12 @@ module Quonfig
1123
1123
  # The reporter runs on a background thread and periodically POSTs
1124
1124
  # context-shape and example-context batches to +telemetry_destination+.
1125
1125
  def initialize_telemetry(start: true)
1126
- shape_aggregator = nil
1127
- example_aggregator = nil
1128
- summaries_aggregator = nil
1129
-
1130
- if @options.collect_max_shapes.to_i.positive?
1131
- shape_aggregator = Quonfig::Telemetry::ContextShapeAggregator.new(
1132
- max_shapes: @options.collect_max_shapes
1133
- )
1134
- end
1135
-
1136
- if @options.collect_max_example_contexts.to_i.positive?
1137
- example_aggregator = Quonfig::Telemetry::ExampleContextsAggregator.new(
1138
- max_contexts: @options.collect_max_example_contexts
1139
- )
1140
- end
1141
-
1142
- if @options.collect_max_evaluation_summaries.to_i.positive?
1143
- summaries_aggregator = Quonfig::Telemetry::EvaluationSummariesAggregator.new(
1144
- max_keys: @options.collect_max_evaluation_summaries
1145
- )
1146
- end
1147
-
1148
- return if shape_aggregator.nil? && example_aggregator.nil? && summaries_aggregator.nil?
1149
-
1150
- @telemetry_reporter = Quonfig::Telemetry::TelemetryReporter.new(
1126
+ @telemetry_reporter = Quonfig::Telemetry::TelemetryReporter.build(
1151
1127
  options: @options,
1152
1128
  instance_hash: @instance_hash,
1153
- context_shape_aggregator: shape_aggregator,
1154
- example_contexts_aggregator: example_aggregator,
1155
- evaluation_summaries_aggregator: summaries_aggregator,
1156
- failover_aggregator: @failover_aggregator,
1157
- sync_interval: @options.collect_sync_interval
1129
+ failover_aggregator: @failover_aggregator
1158
1130
  )
1159
-
1131
+ return if @telemetry_reporter.nil?
1160
1132
  return unless @telemetry_reporter.enabled?
1161
1133
  return unless start
1162
1134
 
@@ -34,10 +34,16 @@ module Quonfig
34
34
  # Options#config_fetch_timeout_ms (sequential) or the hedge abort (hedged
35
35
  # legs) so a hung OR drip-feeding upstream aborts fast instead of blocking
36
36
  # the caller's whole init budget.
37
- def initialize(uri, sdk_key, timeout_ms: nil)
37
+ #
38
+ # +open_timeout_ms+ (qfg-y8je.8): a separate, usually shorter, bound on the
39
+ # connect (open) phase — TCP connect plus TLS handshake. nil uses
40
+ # +timeout_ms+ for it, as before. The telemetry reporter passes 5s here and
41
+ # 15s as +timeout_ms+.
42
+ def initialize(uri, sdk_key, timeout_ms: nil, open_timeout_ms: nil)
38
43
  @uri = uri
39
44
  @sdk_key = sdk_key
40
45
  @timeout_ms = timeout_ms
46
+ @open_timeout_ms = open_timeout_ms
41
47
  end
42
48
 
43
49
  attr_reader :uri
@@ -46,8 +52,11 @@ module Quonfig
46
52
  with_wall_clock_deadline { connection(headers).get(path) }
47
53
  end
48
54
 
55
+ # A String +body+ is sent verbatim (the telemetry reporter resends a
56
+ # retained batch byte-for-byte); anything else is serialized as JSON.
49
57
  def post(path, body)
50
- with_wall_clock_deadline { connection.post(path, body.to_json) }
58
+ payload = body.is_a?(String) ? body : body.to_json
59
+ with_wall_clock_deadline { connection.post(path, payload) }
51
60
  end
52
61
 
53
62
  def connection(headers = {})
@@ -63,6 +72,7 @@ module Quonfig
63
72
  conn.options.open_timeout = seconds
64
73
  conn.options.timeout = seconds
65
74
  end
75
+ conn.options.open_timeout = @open_timeout_ms / 1000.0 if @open_timeout_ms
66
76
  end
67
77
  end
68
78
 
@@ -8,7 +8,9 @@ module Quonfig
8
8
  attr_reader :sdk_key, :environment, :api_urls, :sse_api_urls, :telemetry_destination, :config_api_urls,
9
9
  :on_no_default, :init_timeout_ms, :on_init_failure, :collect_sync_interval, :datadir, :enable_sse, :fallback_poll_enabled, :fallback_poll_interval_ms, :global_context, :logger_key, :logger, :enable_quonfig_user_context,
10
10
  :data_dir_auto_reload, :data_dir_auto_reload_debounce_ms, :config_fetch_timeout_ms,
11
- :config_fetch_hedge_delay_ms, :config_fetch_hedge_abort_ms, :api_urls_explicit
11
+ :config_fetch_hedge_delay_ms, :config_fetch_hedge_abort_ms, :api_urls_explicit,
12
+ :telemetry_timeout_ms, :telemetry_connect_timeout_ms, :telemetry_max_retained_batches,
13
+ :telemetry_max_retained_bytes, :telemetry_max_retained_age_ms
12
14
  attr_accessor :is_fork
13
15
 
14
16
  # Default fallback poll interval, in milliseconds. The SDK polls api-delivery
@@ -85,9 +87,28 @@ module Quonfig
85
87
  end
86
88
 
87
89
  DEFAULT_MAX_PATHS = 1_000
88
- DEFAULT_MAX_KEYS = 100_000
89
- DEFAULT_MAX_EXAMPLE_CONTEXTS = 100_000
90
- DEFAULT_MAX_EVAL_SUMMARIES = 100_000
90
+ # Telemetry aggregator caps per flush window (P6 of the telemetry transport
91
+ # policy, qfg-y8je.8): the uniform server-SDK cap. Were 100,000 before 1.5.0.
92
+ DEFAULT_MAX_KEYS = 10_000
93
+ DEFAULT_MAX_EXAMPLE_CONTEXTS = 10_000
94
+ DEFAULT_MAX_EVAL_SUMMARIES = 10_000
95
+
96
+ # Telemetry transport defaults (qfg-y8je.8; policy P1-P5 in
97
+ # project/plans/2026-09-24-sdk-telemetry-transport-policy.md).
98
+ # Seconds between telemetry POSTs (a fixed cadence; was an 8s interval
99
+ # that grew to 600s).
100
+ DEFAULT_COLLECT_SYNC_INTERVAL = 60
101
+ # Overall deadline for one telemetry POST.
102
+ DEFAULT_TELEMETRY_TIMEOUT_MS = 15_000
103
+ # TCP connect + TLS deadline for one telemetry POST.
104
+ DEFAULT_TELEMETRY_CONNECT_TIMEOUT_MS = 5_000
105
+ # Failed batches kept for resend: at most this many...
106
+ DEFAULT_TELEMETRY_MAX_RETAINED_BATCHES = 5
107
+ # ...and at most this many serialized bytes (2MB); a single batch larger
108
+ # than this is POSTed once and never kept.
109
+ DEFAULT_TELEMETRY_MAX_RETAINED_BYTES = 2 * 1024 * 1024
110
+ # A kept batch older than this is discarded.
111
+ DEFAULT_TELEMETRY_MAX_RETAINED_AGE_MS = 300_000
91
112
 
92
113
  # Hardcoded fallback domain. Overridden by ENV['QUONFIG_DOMAIN'].
93
114
  DEFAULT_DOMAIN = 'quonfig.com'
@@ -213,6 +234,22 @@ module Quonfig
213
234
  # Debounce window in milliseconds. Filesystem events arriving
214
235
  # inside the window are coalesced into a single re-read. Ignored
215
236
  # when +:data_dir_auto_reload+ is +false+.
237
+ # @option options [Numeric] :collect_sync_interval (60)
238
+ # Seconds between telemetry POSTs, on a fixed cadence. At most one POST
239
+ # is in flight; a tick that fires while one is out is skipped and its
240
+ # data rolls into the next window.
241
+ # @option options [Integer] :telemetry_timeout_ms (15000)
242
+ # Overall deadline for one telemetry POST.
243
+ # @option options [Integer] :telemetry_connect_timeout_ms (5000)
244
+ # TCP connect + TLS deadline for one telemetry POST.
245
+ # @option options [Integer] :telemetry_max_retained_batches (5)
246
+ # Failed batches kept (byte-for-byte) for resend; the oldest is dropped
247
+ # beyond this.
248
+ # @option options [Integer] :telemetry_max_retained_bytes (2097152)
249
+ # Byte cap on kept batches. A single batch larger than this is POSTed
250
+ # once and dropped if that POST fails.
251
+ # @option options [Integer] :telemetry_max_retained_age_ms (300000)
252
+ # A kept batch older than this is discarded.
216
253
  # @option options [Boolean] :allow_telemetry_in_local_mode (false)
217
254
  # @deprecated No-op since 1.3.0 (qfg-5x9x). Telemetry is gated on SDK-key
218
255
  # presence alone, so datadir mode no longer suppresses it and this flag
@@ -240,6 +277,11 @@ module Quonfig
240
277
  config_fetch_hedge_abort_ms: nil,
241
278
  collect_max_paths: DEFAULT_MAX_PATHS,
242
279
  collect_sync_interval: nil,
280
+ telemetry_timeout_ms: nil,
281
+ telemetry_connect_timeout_ms: nil,
282
+ telemetry_max_retained_batches: nil,
283
+ telemetry_max_retained_bytes: nil,
284
+ telemetry_max_retained_age_ms: nil,
243
285
  context_upload_mode: :periodic_example, # :periodic_example, :shapes_only, :none
244
286
  context_max_size: DEFAULT_MAX_EVAL_SUMMARIES,
245
287
  collect_evaluation_summaries: true,
@@ -303,7 +345,15 @@ module Quonfig
303
345
  @config_fetch_hedge_abort_ms = config_fetch_hedge_abort_ms || DEFAULT_CONFIG_FETCH_HEDGE_ABORT_MS
304
346
 
305
347
  @collect_max_paths = collect_max_paths
306
- @collect_sync_interval = collect_sync_interval
348
+ @collect_sync_interval = collect_sync_interval.nil? ? DEFAULT_COLLECT_SYNC_INTERVAL : collect_sync_interval
349
+ # Telemetry transport (qfg-y8je.8). nil, non-numeric or <= 0 -> default.
350
+ @telemetry_timeout_ms = positive_or(telemetry_timeout_ms, DEFAULT_TELEMETRY_TIMEOUT_MS)
351
+ @telemetry_connect_timeout_ms = positive_or(telemetry_connect_timeout_ms, DEFAULT_TELEMETRY_CONNECT_TIMEOUT_MS)
352
+ @telemetry_max_retained_batches = positive_or(telemetry_max_retained_batches,
353
+ DEFAULT_TELEMETRY_MAX_RETAINED_BATCHES)
354
+ @telemetry_max_retained_bytes = positive_or(telemetry_max_retained_bytes, DEFAULT_TELEMETRY_MAX_RETAINED_BYTES)
355
+ @telemetry_max_retained_age_ms = positive_or(telemetry_max_retained_age_ms,
356
+ DEFAULT_TELEMETRY_MAX_RETAINED_AGE_MS)
307
357
  @collect_evaluation_summaries = collect_evaluation_summaries
308
358
  @collect_max_evaluation_summaries = collect_max_evaluation_summaries
309
359
  # Retained for back-compat only; nothing reads it (qfg-5x9x).
@@ -379,6 +429,10 @@ module Quonfig
379
429
  option && sdk_key?
380
430
  end
381
431
 
432
+ def positive_or(value, default)
433
+ value.is_a?(Numeric) && value.positive? && value.finite? ? value : default
434
+ end
435
+
382
436
  def remove_trailing_slash(url)
383
437
  url.end_with?('/') ? url[0..-2] : url
384
438
  end
@@ -10,11 +10,16 @@ module Quonfig
10
10
  # matching sdk-node and sdk-go. This is NOT the old Prefab protobuf.
11
11
  class ExampleContextsAggregator
12
12
  ONE_HOUR_SECONDS = 60 * 60
13
+ # Bound on the once-per-hour rate-limit map (P6 of the telemetry
14
+ # transport policy, qfg-y8je.8): a new key arriving when it is full
15
+ # prunes expired entries, and is not recorded if it is still full.
16
+ SEEN_CAP = 100_000
13
17
 
14
18
  attr_reader :data, :cache
15
19
 
16
- def initialize(max_contexts:, rate_limit_seconds: ONE_HOUR_SECONDS)
20
+ def initialize(max_contexts:, rate_limit_seconds: ONE_HOUR_SECONDS, seen_cap: SEEN_CAP)
17
21
  @max_contexts = max_contexts
22
+ @seen_cap = seen_cap
18
23
  @data = Concurrent::Array.new
19
24
  @cache = Quonfig::RateLimitCache.new(rate_limit_seconds)
20
25
  end
@@ -30,6 +35,7 @@ module Quonfig
30
35
  return if key.nil? || key.empty?
31
36
 
32
37
  return unless @data.size < @max_contexts && !@cache.fresh?(key)
38
+ return unless room_in_cache?
33
39
 
34
40
  @cache.set(key)
35
41
  @data.push([Quonfig::TimeHelpers.now_in_ms, context])
@@ -59,6 +65,13 @@ module Quonfig
59
65
 
60
66
  private
61
67
 
68
+ def room_in_cache?
69
+ return true if @cache.data.size < @seen_cap
70
+
71
+ @cache.prune
72
+ @cache.data.size < @seen_cap
73
+ end
74
+
62
75
  def grouped_key_for(context)
63
76
  return context.grouped_key if context.respond_to?(:grouped_key)
64
77
 
@@ -1,10 +1,16 @@
1
1
  # frozen_string_literal: true
2
2
 
3
+ require 'json'
4
+
3
5
  module Quonfig
4
6
  module Telemetry
5
- # Owns the background thread that periodically drains the context
6
- # aggregators and POSTs a JSON telemetry batch to
7
- # +<telemetry_destination>/api/v1/telemetry/+.
7
+ # Owns the background thread that drains the aggregators once per tick and
8
+ # hands the serialized window to a TransportQueue, which retains failed
9
+ # batches byte-for-byte and resends them under the telemetry transport
10
+ # policy (qfg-y8je.8: 60s ticks, one POST in flight, 15s timeout / 5s
11
+ # connect, 30s floor after a failure, Retry-After up to 10 min,
12
+ # 5 batches / 2MB / 5 min retention, disable on 401/403/404; see
13
+ # TransportQueue).
8
14
  #
9
15
  # Wire shape matches api-telemetry's TelemetryEventsSchema:
10
16
  #
@@ -24,16 +30,47 @@ module Quonfig
24
30
  class TelemetryReporter
25
31
  LOG = Quonfig::InternalLogger.new(self)
26
32
 
27
- DEFAULT_INITIAL_DELAY_SECONDS = 8
28
- DEFAULT_MAX_DELAY_SECONDS = 600
33
+ TELEMETRY_PATH = '/api/v1/telemetry/'
34
+ BODY_SNIPPET_BYTES = 1024
35
+ private_constant :BODY_SNIPPET_BYTES
36
+
37
+ # Build the aggregators the options enable and a reporter over them, or
38
+ # nil when every collector is off. Client#initialize_telemetry uses this;
39
+ # +clock+ is a test seam (defaults to the monotonic clock).
40
+ def self.build(options:, instance_hash:, failover_aggregator: nil, clock: nil)
41
+ shapes = (ContextShapeAggregator.new(max_shapes: options.collect_max_shapes) if options.collect_max_shapes.to_i.positive?)
42
+ examples = (ExampleContextsAggregator.new(max_contexts: options.collect_max_example_contexts) if options.collect_max_example_contexts.to_i.positive?)
43
+ summaries = (EvaluationSummariesAggregator.new(max_keys: options.collect_max_evaluation_summaries) if options.collect_max_evaluation_summaries.to_i.positive?)
44
+ return nil if shapes.nil? && examples.nil? && summaries.nil?
45
+
46
+ new(
47
+ options: options,
48
+ instance_hash: instance_hash,
49
+ context_shape_aggregator: shapes,
50
+ example_contexts_aggregator: examples,
51
+ evaluation_summaries_aggregator: summaries,
52
+ failover_aggregator: failover_aggregator,
53
+ clock: clock
54
+ )
55
+ end
56
+
57
+ # Resolved transport settings (the +telemetry_*+ options and
58
+ # +collect_sync_interval+). +flush_interval_ms+ is nil when the interval
59
+ # is a callable.
60
+ attr_reader :config
29
61
 
62
+ # +sync_interval+ (seconds, or a callable returning seconds) defaults to
63
+ # +options.collect_sync_interval+ (60). +http_connection+ is a test seam:
64
+ # anything answering +post(path, body_string)+ with a response that has
65
+ # +status+.
30
66
  def initialize(options:, instance_hash:,
31
67
  context_shape_aggregator: nil,
32
68
  example_contexts_aggregator: nil,
33
69
  evaluation_summaries_aggregator: nil,
34
70
  failover_aggregator: nil,
35
71
  sync_interval: nil,
36
- http_connection: nil)
72
+ http_connection: nil,
73
+ clock: nil)
37
74
  @options = options
38
75
  @instance_hash = instance_hash
39
76
  @sdk_key = options.sdk_key
@@ -47,8 +84,31 @@ module Quonfig
47
84
  # aggregator at the failover call sites; the reporter only drains it.
48
85
  @failover_aggregator = failover_aggregator
49
86
  @http_connection = http_connection
50
- @sync_interval = calculate_sync_interval(sync_interval)
51
- @stopped = Concurrent::AtomicBoolean.new(false)
87
+ @sync_interval = sync_interval.nil? ? options.collect_sync_interval : sync_interval
88
+ @config = {
89
+ flush_interval_ms: @sync_interval.is_a?(Numeric) ? (@sync_interval * 1000).to_i : nil,
90
+ timeout_ms: options.telemetry_timeout_ms,
91
+ connect_timeout_ms: options.telemetry_connect_timeout_ms,
92
+ max_retained_batches: options.telemetry_max_retained_batches,
93
+ max_retained_bytes: options.telemetry_max_retained_bytes,
94
+ max_retained_age_ms: options.telemetry_max_retained_age_ms
95
+ }.freeze
96
+ @queue = TransportQueue.new(
97
+ sender: method(:post_batch),
98
+ telemetry_url: "#{@telemetry_destination}#{TELEMETRY_PATH}",
99
+ timeout_ms: @config[:timeout_ms],
100
+ max_retained_batches: @config[:max_retained_batches],
101
+ max_retained_bytes: @config[:max_retained_bytes],
102
+ max_retained_age_ms: @config[:max_retained_age_ms],
103
+ clock: clock || MonotonicClock,
104
+ on_disabled: method(:on_disabled)
105
+ )
106
+ # Held for a whole tick (serialize + drain), so at most one POST is in
107
+ # flight (P2); a tick that finds it held is skipped.
108
+ @tick_mutex = Mutex.new
109
+ @state_mutex = Mutex.new
110
+ @wake = ConditionVariable.new
111
+ @closed = false
52
112
  @thread = nil
53
113
  @at_exit_registered = false
54
114
  # Set on #start. Everything that can EMIT is gated on it so a forked
@@ -69,97 +129,104 @@ module Quonfig
69
129
  # summaries are recorded separately via
70
130
  # +record_evaluation(...)+ since they require the evaluation result.
71
131
  def record(context)
72
- return if context.nil?
132
+ return if context.nil? || @queue.disabled?
73
133
 
74
134
  @context_shape_aggregator&.push(context)
75
135
  @example_contexts_aggregator&.record(context)
76
136
  end
77
137
 
78
138
  def record_evaluation(**kwargs)
139
+ return if @queue.disabled?
140
+
79
141
  @evaluation_summaries_aggregator&.record(**kwargs)
80
142
  end
81
143
 
82
144
  def start
83
145
  return if @thread&.alive?
84
146
  return unless enabled?
147
+ return if @closed || @queue.disabled?
85
148
 
86
149
  # Claim ownership for THIS process. fork(2) copies the reporter, its
87
150
  # aggregators, and the process-wide at_exit closure registered below;
88
151
  # the pid recorded here is what lets the copy know it is not the
89
152
  # owner and must stay silent (qfg-lv4n.1, dd-trace-rb's pattern).
90
153
  @owner_pid = Process.pid
91
- @stopped.make_false
92
154
  register_at_exit_handler
93
155
  @thread = Thread.new do
94
156
  Thread.current.name = 'quonfig-telemetry-reporter'
95
157
  LOG.debug "Telemetry reporter started instance_hash=#{@instance_hash} destination=#{@telemetry_destination}"
96
-
97
- until @stopped.true?
98
- begin
99
- sleep_duration = @sync_interval.call
100
- slept = 0.0
101
- step = 0.5
102
- while slept < sleep_duration && !@stopped.true?
103
- sleep([step, sleep_duration - slept].min)
104
- slept += step
105
- end
106
- break if @stopped.true?
107
-
108
- sync
109
- rescue StandardError => e
110
- LOG.warn "[quonfig] Telemetry reporter error: #{e.class}: #{e.message}"
111
- end
112
- end
158
+ run_loop
113
159
  end
114
160
  end
115
161
 
116
- def stop
117
- return if foreign_process?('stop')
162
+ # One tick of the contract's model: skip if closed, disabled or a POST is
163
+ # in flight (P2; the live window keeps aggregating); expire aged batches;
164
+ # skip if the 30s floor or Retry-After has not elapsed; serialize the
165
+ # live window once and append it; drain oldest-first.
166
+ def tick
167
+ return if foreign_process?('tick')
168
+ return unless @tick_mutex.try_lock
118
169
 
119
- @stopped.make_true
120
- thread = @thread
121
- @thread = nil
122
- thread&.wakeup if thread&.alive?
123
- # Final drain attempt on stop so tests / short-lived processes
124
- # don't silently drop pending telemetry.
125
170
  begin
126
- sync
127
- rescue StandardError => e
128
- LOG.debug "[quonfig] Final telemetry sync failed: #{e.class}: #{e.message}"
171
+ run_tick
172
+ ensure
173
+ @tick_mutex.unlock
129
174
  end
130
175
  end
131
176
 
132
- # Drain all aggregators and POST the batch. Public so tests can
133
- # trigger a sync without waiting for the background loop.
177
+ # Send the live window now. Waits for an in-flight POST first (bounded by
178
+ # the request timeout), then runs a tick, so after a failure it respects
179
+ # the 30s floor and Retry-After. Never raises.
134
180
  #
135
181
  # Silent in any process other than the one that started the reporter:
136
182
  # after a fork the child holds a full copy of the PARENT's un-flushed
137
183
  # window, and the parent is still going to flush it itself.
138
- def sync
184
+ def flush
139
185
  return if foreign_process?('sync')
140
186
 
141
- events = []
142
- if (summaries_event = @evaluation_summaries_aggregator&.drain_event)
143
- events << summaries_event
144
- end
145
- if (shape_event = @context_shape_aggregator&.drain_event)
146
- events << shape_event
147
- end
148
- if (example_event = @example_contexts_aggregator&.drain_event)
149
- events << example_event
150
- end
151
- if (failover_event = @failover_aggregator&.drain_event)
152
- events << failover_event
187
+ @tick_mutex.synchronize { run_tick }
188
+ rescue StandardError => e
189
+ LOG.debug "Telemetry flush failed: #{e.class}: #{e.message}"
190
+ end
191
+ alias sync flush
192
+
193
+ # Shutdown (P8): stop the reporter thread (aborting an in-flight POST),
194
+ # then give the live window one POST with a 5s deadline. The retained
195
+ # queue is not drained. Idempotent; never raises; never blocks exit for
196
+ # longer than that deadline.
197
+ def close
198
+ return if foreign_process?('close')
199
+
200
+ thread = @state_mutex.synchronize do
201
+ return if @closed
202
+
203
+ @closed = true
204
+ current = @thread
205
+ @thread = nil
206
+ current
153
207
  end
208
+ stop_thread(thread)
209
+ return if @queue.disabled?
154
210
 
155
- return if events.empty?
211
+ body = serialize_window
212
+ return if body.nil?
156
213
 
157
- payload = {
158
- 'instanceHash' => @instance_hash,
159
- 'events' => events
214
+ @queue.send_final(body, [TransportQueue::SHUTDOWN_FLUSH_DEADLINE_MS, @config[:timeout_ms]].min)
215
+ rescue StandardError => e
216
+ LOG.debug "Telemetry close failed: #{e.class}: #{e.message}"
217
+ end
218
+ alias stop close
219
+
220
+ # Test-visible state (the contract's retained_count / retained_bytes /
221
+ # telemetry_enabled).
222
+ def debug_state
223
+ {
224
+ retained_count: @queue.retained_count,
225
+ retained_bytes: @queue.retained_bytes,
226
+ enabled: !@queue.disabled?,
227
+ in_flight: @queue.in_flight?,
228
+ thread_alive: @thread&.alive? || false
160
229
  }
161
-
162
- post(payload)
163
230
  end
164
231
 
165
232
  # Visible for tests.
@@ -180,7 +247,7 @@ module Quonfig
180
247
  # not exist in the child, and the HTTP connection's fd is shared with
181
248
  # the parent.
182
249
  def discard_inherited!
183
- @stopped.make_true
250
+ @closed = true
184
251
  @thread = nil
185
252
  @context_shape_aggregator = nil
186
253
  @example_contexts_aggregator = nil
@@ -190,6 +257,106 @@ module Quonfig
190
257
 
191
258
  private
192
259
 
260
+ def run_tick
261
+ return if @closed || @queue.disabled?
262
+
263
+ @queue.expire
264
+ return unless @queue.send_allowed?
265
+
266
+ body = serialize_window
267
+ @queue.append(body) if body
268
+ @queue.drain
269
+ end
270
+
271
+ # Fixed cadence: tick k fires at k * interval regardless of how long a
272
+ # drain takes. The interval never grows (the old exponential backoff,
273
+ # which grew even on success, is gone: P4).
274
+ def run_loop
275
+ next_at = MonotonicClock.now_ms + next_interval_ms
276
+ until loop_done?
277
+ wait_ms = next_at - MonotonicClock.now_ms
278
+ if wait_ms.positive?
279
+ @state_mutex.synchronize { @wake.wait(@state_mutex, wait_ms / 1000.0) unless loop_done? }
280
+ next
281
+ end
282
+ next_at += next_interval_ms
283
+ next_at = MonotonicClock.now_ms + next_interval_ms if next_at <= MonotonicClock.now_ms
284
+ begin
285
+ tick
286
+ rescue StandardError => e
287
+ LOG.debug "Telemetry tick failed: #{e.class}: #{e.message}"
288
+ end
289
+ end
290
+ end
291
+
292
+ def loop_done?
293
+ @closed || @queue.disabled?
294
+ end
295
+
296
+ def next_interval_ms
297
+ seconds = @sync_interval.respond_to?(:call) ? @sync_interval.call : @sync_interval
298
+ [(seconds.to_f * 1000).to_i, 1].max
299
+ end
300
+
301
+ def wake_thread
302
+ @state_mutex.synchronize { @wake.broadcast }
303
+ end
304
+
305
+ # Stop the reporter thread. A thread parked in its wait exits on the
306
+ # broadcast; one still busy (a POST in flight) is killed, which aborts
307
+ # the POST like sdk-node's close() does. The aborted batch is never resent.
308
+ def stop_thread(thread)
309
+ return if thread.nil? || thread == Thread.current || !thread.alive?
310
+
311
+ wake_thread
312
+ thread.kill unless thread.join(0.1)
313
+ end
314
+
315
+ def on_disabled
316
+ # Nothing aggregates for a dead endpoint: drop what the window holds.
317
+ serialize_window
318
+ wake_thread
319
+ end
320
+
321
+ # Drain the aggregators into one serialized payload. This is the only
322
+ # serialization: the queue stores and resends these exact bytes (P5, P9).
323
+ def serialize_window
324
+ events = [
325
+ @evaluation_summaries_aggregator&.drain_event,
326
+ @context_shape_aggregator&.drain_event,
327
+ @example_contexts_aggregator&.drain_event,
328
+ # nil unless the window saw failover activity.
329
+ @failover_aggregator&.drain_event
330
+ ].compact
331
+ return nil if events.empty?
332
+
333
+ JSON.generate('instanceHash' => @instance_hash, 'events' => events)
334
+ end
335
+
336
+ # The TransportQueue sender: one POST of +body+ verbatim.
337
+ def post_batch(body, timeout_ms)
338
+ response = telemetry_connection(timeout_ms).post(TELEMETRY_PATH, body)
339
+ headers = response.respond_to?(:headers) ? response.headers : nil
340
+ snippet = response.respond_to?(:body) ? response.body.to_s.byteslice(0, BODY_SNIPPET_BYTES) : ''
341
+ TransportQueue::Result.new(
342
+ status: response.status.to_i,
343
+ retry_after: headers && headers['retry-after'],
344
+ body_snippet: snippet
345
+ )
346
+ end
347
+
348
+ # A connection bounded by +timeout_ms+ overall and the connect timeout
349
+ # (never longer than +timeout_ms+) for TCP connect + TLS (P1).
350
+ def telemetry_connection(timeout_ms)
351
+ return @http_connection if @http_connection
352
+
353
+ Quonfig::HttpConnection.new(
354
+ @telemetry_destination, @sdk_key,
355
+ timeout_ms: timeout_ms,
356
+ open_timeout_ms: [@config[:connect_timeout_ms], timeout_ms].min
357
+ )
358
+ end
359
+
193
360
  # True when this reporter belongs to a different process — i.e. we are
194
361
  # a fork(2) copy. Never true before #start (nothing has been claimed,
195
362
  # and nothing was registered at_exit either).
@@ -204,7 +371,7 @@ module Quonfig
204
371
 
205
372
  # Rails / Passenger / Puma workers often terminate via SIGTERM without
206
373
  # a chance to call Client#stop. Register a Kernel.at_exit hook on
207
- # first start so the in-flight batch still gets flushed.
374
+ # first start so the live window still gets its one final flush.
208
375
  def register_at_exit_handler
209
376
  return if @at_exit_registered
210
377
 
@@ -212,62 +379,11 @@ module Quonfig
212
379
  @at_exit_registered = true
213
380
  end
214
381
 
215
- # Wait this long for the background reporter thread to exit before
216
- # giving up. Bounded so a thread blocked on a dead telemetry endpoint
217
- # can't hang process exit.
218
- AT_EXIT_THREAD_JOIN_TIMEOUT_SECONDS = 1.0
219
- private_constant :AT_EXIT_THREAD_JOIN_TIMEOUT_SECONDS
220
-
221
- # Idempotent final drain. Safe to call after #stop has already
222
- # drained: aggregators return nil when empty and #sync becomes a
223
- # no-op. Bounded so a stuck reporter thread or dead telemetry
382
+ # Idempotent final flush (#close). Safe after #stop: a second close is a
383
+ # no-op. Bounded by the 5s shutdown deadline, so a dead telemetry
224
384
  # endpoint can't hang process exit.
225
385
  def final_drain_on_exit
226
- return if foreign_process?('at_exit drain')
227
-
228
- @stopped.make_true
229
- thread = @thread
230
- @thread = nil
231
- if thread&.alive?
232
- thread.wakeup
233
- thread.join(AT_EXIT_THREAD_JOIN_TIMEOUT_SECONDS)
234
- end
235
- sync
236
- rescue StandardError => e
237
- LOG.debug "[quonfig] at_exit telemetry drain failed: #{e.class}: #{e.message}"
238
- end
239
-
240
- def post(payload)
241
- conn = http_connection
242
- return if conn.nil?
243
-
244
- response = conn.post('/api/v1/telemetry/', payload)
245
- status = response.respond_to?(:status) ? response.status : nil
246
- if status && status >= 400
247
- LOG.warn "[quonfig] Telemetry POST failed: #{status}"
248
- else
249
- LOG.debug "[quonfig] Telemetry POST ok: events=#{payload['events'].size}"
250
- end
251
- response
252
- end
253
-
254
- def http_connection
255
- @http_connection ||= begin
256
- return nil if @sdk_key.nil? || @telemetry_destination.nil?
257
-
258
- Quonfig::HttpConnection.new(@telemetry_destination, @sdk_key)
259
- end
260
- end
261
-
262
- def calculate_sync_interval(sync_interval)
263
- return proc { sync_interval } if sync_interval.is_a?(Numeric)
264
- return sync_interval if sync_interval.respond_to?(:call)
265
-
266
- Quonfig::ExponentialBackoff.new(
267
- initial_delay: DEFAULT_INITIAL_DELAY_SECONDS,
268
- max_delay: DEFAULT_MAX_DELAY_SECONDS,
269
- multiplier: 1.5
270
- )
386
+ close
271
387
  end
272
388
  end
273
389
  end
@@ -0,0 +1,311 @@
1
+ # frozen_string_literal: true
2
+
3
+ require 'time'
4
+
5
+ module Quonfig
6
+ module Telemetry
7
+ # Monotonic milliseconds for the telemetry transport: the 30s floor,
8
+ # Retry-After, retained-batch age and the WARN cadence all read it. Tests
9
+ # inject a manual clock through TelemetryReporter.build(clock:).
10
+ module MonotonicClock
11
+ def self.now_ms
12
+ Process.clock_gettime(Process::CLOCK_MONOTONIC, :millisecond)
13
+ end
14
+ end
15
+
16
+ # Telemetry transport policy (qfg-y8je.8; policy P1-P10 in
17
+ # project/plans/2026-09-24-sdk-telemetry-transport-policy.md, contract tests
18
+ # in integration-test-data/chaos/telemetry-transport-contract.md; mirrors
19
+ # sdk-node's src/telemetry/transportQueue.ts).
20
+ #
21
+ # Owns the retained queue of serialized batches, the send gate (30s floor
22
+ # after a failure + Retry-After), the drain loop, disable-on-auth and the P7
23
+ # logging episodes. It knows nothing about aggregators or payload shape: it
24
+ # stores and resends opaque bytes, never re-serializing or merging them.
25
+ #
26
+ # Not thread-safe on its own: TelemetryReporter serializes every call under
27
+ # its tick mutex, which is also what keeps one POST in flight (P2).
28
+ class TransportQueue
29
+ LOG = Quonfig::InternalLogger.new(self)
30
+
31
+ # No send sooner than this after a failed POST (P4).
32
+ RESEND_FLOOR_MS = 30_000
33
+ # Retry-After is honored up to this (P4).
34
+ RETRY_AFTER_CAP_MS = 600_000
35
+ # At most one drop WARN per this interval while dropping continues (P7).
36
+ DROP_WARN_INTERVAL_MS = 600_000
37
+ # close() gives the live window one POST with this deadline (P8).
38
+ SHUTDOWN_FLUSH_DEADLINE_MS = 5_000
39
+
40
+ # Outcome of one POST that got an HTTP response.
41
+ Result = Struct.new(:status, :retry_after, :body_snippet, keyword_init: true)
42
+
43
+ Batch = Struct.new(:body, :bytes, :created_at, :oversize, keyword_init: true)
44
+
45
+ # 2xx -> :ok; 401, 403, 404 -> :auth; 408, 429, 5xx -> :retryable; every
46
+ # other status (other 4xx, 3xx, 1xx) -> :rejected (P3).
47
+ def self.classify_status(status)
48
+ return :ok if status >= 200 && status < 300
49
+ return :auth if [401, 403, 404].include?(status)
50
+ return :retryable if [408, 429].include?(status) || (status >= 500 && status < 600)
51
+
52
+ :rejected
53
+ end
54
+
55
+ # Parse a Retry-After header into a wait in ms: delta-seconds, or an
56
+ # HTTP-date relative to now (past dates -> 0). Unparseable -> nil.
57
+ # Clamped to RETRY_AFTER_CAP_MS.
58
+ def self.parse_retry_after_ms(header, wall_now: Time.now)
59
+ return nil if header.nil?
60
+
61
+ value = header.to_s.strip
62
+ return nil if value.empty?
63
+
64
+ ms =
65
+ if value.match?(/\A\d+\z/)
66
+ value.to_i * 1000
67
+ else
68
+ begin
69
+ [((Time.httpdate(value) - wall_now) * 1000).round, 0].max
70
+ rescue ArgumentError
71
+ return nil
72
+ end
73
+ end
74
+ [ms, RETRY_AFTER_CAP_MS].min
75
+ end
76
+
77
+ # +sender+ is called as +sender.call(body, timeout_ms)+ and returns a
78
+ # Result, or raises (a Faraday/Timeout error for a timeout, anything else
79
+ # for a network failure).
80
+ def initialize(sender:, telemetry_url:, timeout_ms:, max_retained_batches:,
81
+ max_retained_bytes:, max_retained_age_ms:, clock: MonotonicClock,
82
+ on_disabled: nil)
83
+ @sender = sender
84
+ @telemetry_url = telemetry_url
85
+ @timeout_ms = timeout_ms
86
+ @max_retained_batches = max_retained_batches
87
+ @max_retained_bytes = max_retained_bytes
88
+ @max_retained_age_ms = max_retained_age_ms
89
+ @clock = clock
90
+ @on_disabled = on_disabled
91
+
92
+ @queue = []
93
+ @in_flight = false
94
+ @last_failure_at = nil
95
+ @retry_after_until = nil
96
+ @disabled = false
97
+
98
+ # Outage episode (P7).
99
+ @failures_since_success = 0
100
+ @first_failure_at = nil
101
+ @last_result = nil
102
+ @last_drop_warn_at = nil
103
+ @drops_since_warn = 0
104
+ @drops_this_outage = 0
105
+
106
+ # Rejected-batch (other 4xx) cadence.
107
+ @last_reject_error_at = nil
108
+ @rejects_since_error = 0
109
+ end
110
+
111
+ def in_flight? = @in_flight
112
+ def disabled? = @disabled
113
+
114
+ # Every queued batch, including a not-yet-sent oversize one.
115
+ def retained_count = @queue.size
116
+
117
+ def retained_bytes = @queue.sum(&:bytes)
118
+
119
+ # Discard batches older than the max age (strictly greater). Tick step 2.
120
+ def expire
121
+ now = @clock.now_ms
122
+ while (head = @queue.first) && now - head.created_at > @max_retained_age_ms
123
+ @queue.shift
124
+ record_drop("batch older than #{(@max_retained_age_ms / 60_000.0).round} min")
125
+ end
126
+ end
127
+
128
+ # The 30s floor after a failure and any Retry-After have both elapsed.
129
+ # Tick step 3.
130
+ def send_allowed?
131
+ now = @clock.now_ms
132
+ return false if @last_failure_at && now < @last_failure_at + RESEND_FLOOR_MS
133
+ return false if @retry_after_until && now < @retry_after_until
134
+
135
+ true
136
+ end
137
+
138
+ # Append a serialized window and enforce the caps, dropping oldest. An
139
+ # oversize batch is never counted against the caps and never evicted by
140
+ # them: it is sent once and then dropped (see #drain). Tick step 4.
141
+ def append(body)
142
+ oversize = body.bytesize > @max_retained_bytes
143
+ @queue << Batch.new(body: body, bytes: body.bytesize, created_at: @clock.now_ms, oversize: oversize)
144
+
145
+ kept = @queue.reject(&:oversize)
146
+ count = kept.size
147
+ bytes = kept.sum(&:bytes)
148
+ while count > @max_retained_batches || bytes > @max_retained_bytes
149
+ index = @queue.index { |b| !b.oversize }
150
+ break if index.nil?
151
+
152
+ evicted = @queue.delete_at(index)
153
+ count -= 1
154
+ bytes -= evicted.bytes
155
+ record_drop('retained queue full')
156
+ end
157
+ end
158
+
159
+ # POST queued batches oldest-first, one at a time; stop at the first
160
+ # failure. Tick step 5.
161
+ def drain
162
+ until @queue.empty? || @disabled
163
+ batch = @queue.first
164
+ outcome = post(batch.body, @timeout_ms)
165
+
166
+ if outcome.is_a?(String)
167
+ on_retryable_failure(batch, outcome, nil)
168
+ break
169
+ end
170
+
171
+ case self.class.classify_status(outcome.status)
172
+ when :ok
173
+ @queue.shift
174
+ on_success
175
+ when :retryable
176
+ on_retryable_failure(batch, outcome.status.to_s, outcome.retry_after)
177
+ break
178
+ when :auth
179
+ disable(outcome.status)
180
+ return
181
+ else
182
+ # Rejected: drop this batch, report, carry on with the next one.
183
+ @queue.shift
184
+ on_rejected(outcome.status, batch.bytes, outcome.body_snippet)
185
+ end
186
+ end
187
+
188
+ # Oversize batches are never carried across ticks.
189
+ @queue.select(&:oversize).each do |batch|
190
+ remove(batch)
191
+ record_drop('batch larger than the byte cap')
192
+ end
193
+ end
194
+
195
+ # close(): one POST of the live window bounded by +deadline_ms+. Never
196
+ # retains, never touches the outage episode, never raises.
197
+ def send_final(body, deadline_ms)
198
+ outcome = post(body, deadline_ms)
199
+ result =
200
+ if outcome.is_a?(String)
201
+ outcome
202
+ elsif self.class.classify_status(outcome.status) != :ok
203
+ outcome.status.to_s
204
+ end
205
+ return if result.nil?
206
+
207
+ LOG.debug "Telemetry final flush at shutdown failed (#{result}); #{body.bytesize} bytes dropped, " \
208
+ "#{retained_count} retained batch(es) abandoned"
209
+ end
210
+
211
+ private
212
+
213
+ # Remove by identity (two batches may carry equal bytes).
214
+ def remove(batch)
215
+ @queue.reject! { |b| b.equal?(batch) }
216
+ end
217
+
218
+ # One POST. Returns a Result, or a String describing a request that got
219
+ # no HTTP response ("timeout", "network error: ...").
220
+ def post(body, timeout_ms)
221
+ @in_flight = true
222
+ @sender.call(body, timeout_ms)
223
+ rescue Faraday::TimeoutError, Timeout::Error
224
+ 'timeout'
225
+ rescue StandardError => e
226
+ "network error: #{e.class}: #{e.message}"
227
+ ensure
228
+ @in_flight = false
229
+ end
230
+
231
+ def on_success
232
+ return if @failures_since_success.zero?
233
+
234
+ seconds = ((@clock.now_ms - (@first_failure_at || @clock.now_ms)) / 1000.0).round
235
+ LOG.info "Telemetry recovered: POST succeeded after #{@failures_since_success} failed attempt(s) " \
236
+ "over #{seconds}s; #{@drops_this_outage} batch(es) were dropped."
237
+ @failures_since_success = 0
238
+ @first_failure_at = nil
239
+ @drops_this_outage = 0
240
+ @last_drop_warn_at = nil
241
+ @drops_since_warn = 0
242
+ end
243
+
244
+ def on_retryable_failure(batch, result, retry_after)
245
+ now = @clock.now_ms
246
+ @failures_since_success += 1
247
+ @first_failure_at ||= now
248
+ @last_failure_at = now
249
+ @last_result = result
250
+ wait = self.class.parse_retry_after_ms(retry_after)
251
+ @retry_after_until = now + wait if wait
252
+
253
+ next_ms = [@last_failure_at + RESEND_FLOOR_MS, @retry_after_until || 0].max - now
254
+ LOG.debug "Telemetry POST failed (#{result}); #{retained_count} batch(es) / #{retained_bytes} bytes " \
255
+ "retained, next send in >= #{(next_ms / 1000.0).ceil}s"
256
+
257
+ return unless batch.oversize
258
+
259
+ remove(batch)
260
+ record_drop('batch larger than the byte cap')
261
+ end
262
+
263
+ def disable(status)
264
+ hint = status == 404 ? 'wrong telemetry_url' : 'the SDK key was rejected'
265
+ LOG.error "Telemetry disabled for this process: #{@telemetry_url} answered #{status} (#{hint}). " \
266
+ 'Flag evaluation is unaffected.'
267
+ @queue.clear
268
+ @disabled = true
269
+ @on_disabled&.call
270
+ end
271
+
272
+ def on_rejected(status, bytes, body_snippet)
273
+ now = @clock.now_ms
274
+ if @last_reject_error_at.nil? || now - @last_reject_error_at >= DROP_WARN_INTERVAL_MS
275
+ more = @rejects_since_error.positive? ? ", #{@rejects_since_error} more since the last report" : ''
276
+ LOG.error "Telemetry batch rejected with #{status} and dropped (#{bytes} bytes#{more}): " \
277
+ "#{body_snippet}. This is likely an SDK bug; please report it."
278
+ @last_reject_error_at = now
279
+ @rejects_since_error = 0
280
+ else
281
+ @rejects_since_error += 1
282
+ LOG.debug "Telemetry batch rejected with #{status} and dropped (#{bytes} bytes)"
283
+ end
284
+ end
285
+
286
+ def record_drop(reason)
287
+ now = @clock.now_ms
288
+ @drops_since_warn += 1
289
+ @drops_this_outage += 1
290
+ last_result = @last_result || 'none'
291
+ if @last_drop_warn_at.nil?
292
+ LOG.warn "Telemetry is dropping data: #{reason} (last POST result: #{last_result}). " \
293
+ "#{@drops_this_outage} batch(es) dropped so far; retained queue " \
294
+ "#{retained_count}/#{@max_retained_batches} batches, #{retained_bytes} bytes. " \
295
+ 'Flag evaluation is unaffected; further drops log at debug with a summary every 10 min.'
296
+ @last_drop_warn_at = now
297
+ @drops_since_warn = 0
298
+ elsif now - @last_drop_warn_at >= DROP_WARN_INTERVAL_MS
299
+ minutes = ((now - @last_drop_warn_at) / 60_000.0).round
300
+ LOG.warn "Telemetry still dropping data: #{@drops_since_warn} batch(es) dropped in the last " \
301
+ "#{minutes} min (last POST result: #{last_result}); retained queue #{retained_count} " \
302
+ "batches, #{retained_bytes} bytes."
303
+ @last_drop_warn_at = now
304
+ @drops_since_warn = 0
305
+ else
306
+ LOG.debug "Telemetry dropped a batch: #{reason}; #{@drops_since_warn} since the last warning"
307
+ end
308
+ end
309
+ end
310
+ end
311
+ end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Quonfig
4
- VERSION = '1.4.1'
4
+ VERSION = '1.5.0'
5
5
  end
data/lib/quonfig.rb CHANGED
@@ -62,6 +62,7 @@ require 'quonfig/telemetry/context_shape_aggregator'
62
62
  require 'quonfig/telemetry/example_contexts_aggregator'
63
63
  require 'quonfig/telemetry/evaluation_summaries_aggregator'
64
64
  require 'quonfig/telemetry/failover_aggregator'
65
+ require 'quonfig/telemetry/transport_queue'
65
66
  require 'quonfig/telemetry/telemetry_reporter'
66
67
  require 'quonfig/client'
67
68
  require 'quonfig/bound_client'
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: quonfig
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.4.1
4
+ version: 1.5.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Jeff Dwyer
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-11 00:00:00.000000000 Z
11
+ date: 2026-09-25 00:00:00.000000000 Z
12
12
  dependencies:
13
13
  - !ruby/object:Gem::Dependency
14
14
  name: activesupport
@@ -132,6 +132,7 @@ files:
132
132
  - lib/quonfig/telemetry/example_contexts_aggregator.rb
133
133
  - lib/quonfig/telemetry/failover_aggregator.rb
134
134
  - lib/quonfig/telemetry/telemetry_reporter.rb
135
+ - lib/quonfig/telemetry/transport_queue.rb
135
136
  - lib/quonfig/time_helpers.rb
136
137
  - lib/quonfig/types.rb
137
138
  - lib/quonfig/version.rb