eco-helpers 3.3.3 → 3.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +43 -12
- data/README.md +62 -1
- data/lib/eco/api/session/concurrency/bounded_worker_pool.rb +150 -0
- data/lib/eco/api/session/concurrency/retry_policy.rb +139 -0
- data/lib/eco/language/klass/auto_loader.rb +12 -3
- data/lib/eco/version.rb +1 -1
- metadata +7 -5
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 0d846f9dd8c67e9ece8e82f76b48f57abb6f2281a827906fa0deea759793a2a4
|
|
4
|
+
data.tar.gz: 31010d6385deea25c65e005f2d056e07cc1c00e18cc3257c080863a9042c5393
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: ca7ad30915e09c003d99c0cf291b0ef2a1834179883de98402380d144bb231cf9f12f2fe802b5b2fed2387b8d63e6e3eee8e4363decdd3636863af0bc702edc8
|
|
7
|
+
data.tar.gz: ed532533a05a25f09b65118104a8627f7bf0fde3b212b2fcd562d499770fdd19de4588caefd388eebb5c91cd9689ab6cbfcf3456e47dafeb69acfab9231956a8
|
data/CHANGELOG.md
CHANGED
|
@@ -4,18 +4,49 @@ All notable changes to this project will be documented in this file.
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
-
## [3.
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
-
|
|
18
|
-
|
|
7
|
+
## [3.4.0] - 2026-09-29
|
|
8
|
+
|
|
9
|
+
**MERGE ONLY AFTER `ecoportal-api-graphql` 3.0.0 IS PUBLISHED** -- this release's gemspec
|
|
10
|
+
floor (`~> 3.0`) cannot resolve against rubygems.org until then; tested here against a
|
|
11
|
+
locally built 3.0.0 CANDIDATE gem (see `docs/worklog.md`'s 2026-09-29 entry for the exact
|
|
12
|
+
build/install steps), not the published artifact.
|
|
13
|
+
|
|
14
|
+
### Changed
|
|
15
|
+
|
|
16
|
+
- **BREAKING for consumers still on `ecoportal-api-graphql` 2.x:** requires
|
|
17
|
+
`ecoportal-api-graphql ~> 3.0` (was `~> 2.0`). That gem's own 3.0.0 removes
|
|
18
|
+
`Base::Page::DataField::ImageGallery#file_container_ids` / `#file_container_ids=` --
|
|
19
|
+
an Image Gallery image was never a `FileContainer`. Replaced by `#images` / `#image_ids` /
|
|
20
|
+
`#add_source_images([{source_id:, file_name:}, ...])` after
|
|
21
|
+
`FileUpload::Client#upload_image`. A repo-wide sweep of `eco-helpers` itself for
|
|
22
|
+
`file_container_ids` / `fileContainerIds` / `fileContainers` found **0 hits** -- no caller
|
|
23
|
+
in `lib/`, `spec/`, `docs/`, `bin/` used the removed accessors (the one look-alike,
|
|
24
|
+
`FileField#file_container_id` in `lib/eco/api/usecases/ooze_samples/helpers_migration/
|
|
25
|
+
copying.rb`, is the UNRELATED APIv2/REST `File` field type's own singular accessor, not
|
|
26
|
+
Image Gallery, not affected). No code change was needed in `eco-helpers` for this bump.
|
|
27
|
+
- `Eco::API::Session::Concurrency::BoundedWorkerPool` / `RetryPolicy`
|
|
28
|
+
(`lib/eco/api/session/concurrency/`) are now the CANONICAL, supported concurrency
|
|
29
|
+
primitives in this gem (dropped the "draft, nothing has switched to this copy yet"
|
|
30
|
+
header wording) -- see `README.md`'s new "Concurrency primitives" section. No
|
|
31
|
+
pre-existing inline `Thread.new`/`Queue.new`/subprocess-retry-loop exists anywhere under
|
|
32
|
+
`lib/eco/api/session` or `lib/eco/api/usecases` to migrate onto them (checked by grep at
|
|
33
|
+
promotion time); a future caller picks these up directly, as-is.
|
|
34
|
+
|
|
35
|
+
### Fixed
|
|
36
|
+
|
|
37
|
+
- `Eco::Language::Klass::AutoLoader#autoload_children!` only tolerated a `TypeError` from a
|
|
38
|
+
malformed pending child (its own comment: "must be the singleton class"); a pending child
|
|
39
|
+
raising any OTHER `StandardError` while being constructed (e.g. an anonymous, never fully
|
|
40
|
+
configured `Class.new(SomeAutoloadedBase)` test double left alive in `ObjectSpace` by an
|
|
41
|
+
entirely unrelated spec elsewhere in the same process -- `Parser`/`ErrorHandler`
|
|
42
|
+
characterization specs create these on purpose) propagated out of the NEXT, unrelated
|
|
43
|
+
caller's own `autoload_children!` cycle instead. Broadened the rescue to `TypeError,
|
|
44
|
+
StandardError`, matching the existing rescue's own documented intent ("can't create from
|
|
45
|
+
this class... just ignore"). Pre-existing, GC-timing-sensitive flake -- confirmed
|
|
46
|
+
reproducible (non-deterministically) against the CURRENT `ecoportal-api-graphql` 2.2.0
|
|
47
|
+
floor too, not something the 3.0 bump introduced; surfaced while testing this release
|
|
48
|
+
against the 3.0.0 candidate, where it reproduced deterministically across three
|
|
49
|
+
consecutive runs before the fix and cleared across three consecutive runs after.
|
|
19
50
|
|
|
20
51
|
## [3.3.2] - 2026-09-02
|
|
21
52
|
|
data/README.md
CHANGED
|
@@ -1,6 +1,34 @@
|
|
|
1
1
|
# API Helpers (eco-helpers)
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
## Why this repo exists
|
|
4
|
+
|
|
5
|
+
`eco-helpers` is the team's **central integrations library** -- one of the oldest repos in the
|
|
6
|
+
fleet. It sits between the raw EcoPortal API clients (`ecoportal-api`, `ecoportal-api-v2`,
|
|
7
|
+
`ecoportal-api-graphql`) and the end scripts, adding file management/access, SFTP access, and
|
|
8
|
+
more (`net-sftp`, `net-ssh`, `aws-sdk-s3` are runtime dependencies -- see `eco-helpers.gemspec`).
|
|
9
|
+
|
|
10
|
+
> "If there is a repo that should somehow be the reference of our integrations that is the
|
|
11
|
+
> eco-helpers repo... [ep-graphql is powerful but raw] ...but if we have an intermediate layer
|
|
12
|
+
> with helpers that allow to build integrations by taking into account not just raw api access,
|
|
13
|
+
> but also file managing/access, SFTP access, etc. that is the eco-helpers for sure."
|
|
14
|
+
> -- owner, 2026-09-23
|
|
15
|
+
|
|
16
|
+
```mermaid
|
|
17
|
+
flowchart LR
|
|
18
|
+
A["Raw API clients<br/>ecoportal-api / -v2 / -graphql"] --> B["eco-helpers<br/>integrations layer"]
|
|
19
|
+
B --> C["End scripts<br/>automation, CLI, use cases"]
|
|
20
|
+
```
|
|
21
|
+
*Layering per the owner's own description (source: memory
|
|
22
|
+
`project-imports-process-and-eco-helpers-owner-truth.md`, 2026-09-24).*
|
|
23
|
+
|
|
24
|
+
### What is inside
|
|
25
|
+
|
|
26
|
+
- `lib/eco/api/` -- API integration layer (session, use-cases, microcases, policies, org resources)
|
|
27
|
+
- `lib/eco/cli/` + `lib/eco/cli_default/` -- CLI framework and default options/filters/people workflows
|
|
28
|
+
- `lib/eco/csv/` -- CSV reading, streaming, splitting
|
|
29
|
+
- `lib/eco/data/` -- data utilities (fuzzy match, hashes, locations, strings, files)
|
|
30
|
+
- `lib/eco/language/` -- logging, curry, and auxiliary utilities
|
|
31
|
+
- `lib/eco/assets/` -- static assets (language files etc.)
|
|
4
32
|
|
|
5
33
|
## Installation
|
|
6
34
|
|
|
@@ -19,6 +47,39 @@ Or install it yourself as:
|
|
|
19
47
|
$ gem install eco-helpers
|
|
20
48
|
|
|
21
49
|
|
|
50
|
+
## Concurrency primitives
|
|
51
|
+
|
|
52
|
+
`Eco::API::Session::Concurrency` (`lib/eco/api/session/concurrency/`) is the supported home
|
|
53
|
+
for session-level concurrency primitives, as of 3.4.0:
|
|
54
|
+
|
|
55
|
+
- **`BoundedWorkerPool`** -- a generic, bounded pool of OS threads (pure `Thread`/`Queue`/
|
|
56
|
+
`Mutex`, no extra gem) for running the SAME block over a list of items with a hard ceiling
|
|
57
|
+
on how many run simultaneously. Results come back in the same order as the input,
|
|
58
|
+
regardless of finish order; one item's block raising is captured as that item's own
|
|
59
|
+
`Result#error`, never re-raised, never aborts sibling work already in flight or queued
|
|
60
|
+
(unless `stop_on_error: true`). See the class's own header for the full contract and
|
|
61
|
+
thread-safety notes.
|
|
62
|
+
- **`RetryPolicy`** -- a pure decision service (never sleeps itself) for whether/how long to
|
|
63
|
+
wait before retrying a whole failed subprocess call, given its captured output and exit
|
|
64
|
+
status: rate-limit/5xx-aware, bounded exponential backoff with full jitter, honours a
|
|
65
|
+
captured `Retry-After` header. See the module's own header for why an OUTER,
|
|
66
|
+
subprocess-level retry is worth having even though the underlying GraphQL gem already
|
|
67
|
+
retries transient errors INSIDE one call.
|
|
68
|
+
|
|
69
|
+
Both are generic (no case-specific or org-specific logic) and fully specced against
|
|
70
|
+
synthetic blocks only -- no subprocess, no network, nothing case-specific. See
|
|
71
|
+
`spec/eco/api/session/concurrency/`.
|
|
72
|
+
|
|
22
73
|
## Changelog
|
|
23
74
|
|
|
24
75
|
See {file:CHANGELOG.md} for a list of changes.
|
|
76
|
+
|
|
77
|
+
<!-- audit-status:start -->
|
|
78
|
+
### Fleet audit status
|
|
79
|
+
|
|
80
|
+
This repo is under the Fleet Audit & Amendment Programme (FAAP). No audit run yet -- this is the notice MR; the first `granular`-tier run populates this block with counts by severity and status.
|
|
81
|
+
|
|
82
|
+
Guidelines (read before any change): `docs/audits/guidelines.md`
|
|
83
|
+
Deviations register: `docs/audits/deviations.md`
|
|
84
|
+
Run index: `docs/audits/README.md`
|
|
85
|
+
<!-- audit-status:end -->
|
|
@@ -0,0 +1,150 @@
|
|
|
1
|
+
# PROMOTED FROM a downstream script repo (services/bounded_worker_pool.rb @ 41dedad)
|
|
2
|
+
# NAMESPACE: PROVISIONAL (owner ruling DSL-02) -- kept the original class/module name
|
|
3
|
+
# (`BoundedWorkerPool`) under `Eco::API::Session::Concurrency::*`, the session-level
|
|
4
|
+
# concurrency primitives location proposed by WP-6 (see
|
|
5
|
+
# the repo's internal docs, Option A).
|
|
6
|
+
#
|
|
7
|
+
# CANONICAL, 3.4.0 (release/3.4.0-prep) -- this is the supported concurrency primitive for
|
|
8
|
+
# a bounded pool of OS threads in this gem; see `README.md`'s own "Concurrency primitives"
|
|
9
|
+
# section. No PRE-EXISTING inline `Thread.new`/`Queue.new` concurrency exists anywhere else
|
|
10
|
+
# under `lib/eco/api/session` or `lib/eco/api/usecases` to migrate onto this (checked by
|
|
11
|
+
# grep, both directories, at promotion time) -- there was nothing to switch, so none of this
|
|
12
|
+
# gem's own call sites change as part of this promotion; a FUTURE caller (this gem's own
|
|
13
|
+
# batch/CLI code, or a downstream consumer such as a downstream script's wave runner, the
|
|
14
|
+
# original motivating caller) picks this up directly, as-is.
|
|
15
|
+
#
|
|
16
|
+
# the wave-concurrency work (2026-09-19) -- generic, reusable bounded-concurrency runner. Pure
|
|
17
|
+
# Ruby stdlib (`Thread`/`Queue`/`Mutex`) -- no new gem, matching this repo's own "PLAIN DATA IN/
|
|
18
|
+
# service-first" convention (see any other `Custom::Template::Services::*` file's header).
|
|
19
|
+
#
|
|
20
|
+
# WHY THIS EXISTS: `run_wave.rb` shells out to `ruby main.rb ...` once PER ITEM, per stage
|
|
21
|
+
# (build/force-install/publish/tooltips) -- ~47 items, each a live network round trip, run
|
|
22
|
+
# strictly sequentially. A pass is minutes of network WAIT, not CPU work, so a bounded pool of
|
|
23
|
+
# OS threads (each blocking on its own `Open3.capture3` call) is the right tool: Ruby (MRI)
|
|
24
|
+
# threads release the GVL during blocking I/O (`Open3.capture3` waits on the child process, a
|
|
25
|
+
# blocking syscall), so N threads genuinely overlap N in-flight subprocess calls rather than
|
|
26
|
+
# contending for CPU.
|
|
27
|
+
#
|
|
28
|
+
# CONTRACT (every point specced in isolation, synthetic blocks only -- no subprocess, no
|
|
29
|
+
# network, nothing case-specific here):
|
|
30
|
+
# - `concurrency:` is a HARD CEILING on simultaneously-RUNNING blocks -- never exceeded.
|
|
31
|
+
# - Results come back in the SAME order as `items`, regardless of which one finishes first
|
|
32
|
+
# (a slow item 1 and a fast item 2 must never swap positions in the returned Array).
|
|
33
|
+
# - One item's block RAISING is CAPTURED as that item's own `Result#error`, never re-raised
|
|
34
|
+
# to the caller, and never aborts sibling work already in flight or already queued --
|
|
35
|
+
# UNLESS `stop_on_error:` says otherwise (below).
|
|
36
|
+
# - Empty `items` returns `[]` immediately, no thread ever spawned.
|
|
37
|
+
#
|
|
38
|
+
# `stop_on_error:` (default `false`, OPT-IN, generic -- not `run_wave.rb`-specific): once ANY
|
|
39
|
+
# dispatched item's block raises, no item still WAITING in the queue is ever started (each
|
|
40
|
+
# such item's own `Result#skipped?` reads `true`, `#value`/`#error` both `nil`) -- but anything
|
|
41
|
+
# ALREADY running when the failure is observed is allowed to finish (there is no clean way to
|
|
42
|
+
# interrupt a live `Open3.capture3` mid-flight, and killing the child process would leave that
|
|
43
|
+
# item's own live state ambiguous, worse than letting it complete). At `concurrency: 1` this
|
|
44
|
+
# reproduces "stop at the first failure, never attempt the next item" EXACTLY -- with a single
|
|
45
|
+
# worker, "already running" and "already failed" can never overlap, so the second item is
|
|
46
|
+
# always still in the QUEUE (never started) the instant the first one fails. `run_wave.rb`'s
|
|
47
|
+
# own force-install/build wiring relies on this exact property for byte-for-byte parity with
|
|
48
|
+
# its pre-concurrency sequential behaviour at `--concurrency 1` (see
|
|
49
|
+
# `run_wave_concurrency_spec.rb`'s own "concurrency 1 equivalence" examples).
|
|
50
|
+
#
|
|
51
|
+
# THREAD SAFETY OF `results[index] = ...` FROM MULTIPLE THREADS: safe under MRI (this repo's
|
|
52
|
+
# only Ruby implementation, per rbenv) -- the GVL serializes bytecode execution, so concurrent
|
|
53
|
+
# writes to DISTINCT indices of one Array never race or corrupt its internal structure (a
|
|
54
|
+
# well-established MRI pattern; JRuby/TruffleRuby would need an explicit lock here too, but
|
|
55
|
+
# this repo does not target them). The ONE piece of state actually SHARED and mutated across
|
|
56
|
+
# threads (`aborted`, for `stop_on_error:`) is behind an explicit `Mutex` regardless, so this
|
|
57
|
+
# file makes no silent assumptions beyond that documented one.
|
|
58
|
+
|
|
59
|
+
module Eco
|
|
60
|
+
module API
|
|
61
|
+
class Session
|
|
62
|
+
module Concurrency
|
|
63
|
+
module BoundedWorkerPool
|
|
64
|
+
# One item's own outcome. Exactly one of `value`/`error` is non-nil unless `skipped?`
|
|
65
|
+
# is true, in which case both are `nil` -- the block was never even called for it.
|
|
66
|
+
Result = Struct.new(:item, :value, :error, :skipped, keyword_init: true) do
|
|
67
|
+
def skipped?
|
|
68
|
+
!!skipped
|
|
69
|
+
end
|
|
70
|
+
|
|
71
|
+
# @return [Boolean] true only for a NORMAL completion -- neither raised nor skipped.
|
|
72
|
+
def ok?
|
|
73
|
+
!skipped? && error.nil?
|
|
74
|
+
end
|
|
75
|
+
end
|
|
76
|
+
|
|
77
|
+
class << self
|
|
78
|
+
# @param items [Array] anything, plain data -- never mutated.
|
|
79
|
+
# @param concurrency [Integer] hard ceiling on simultaneously-RUNNING blocks. Values
|
|
80
|
+
# `<= 0` are treated as `1` (never zero threads, never negative) -- callers own
|
|
81
|
+
# their own upper-bound policy (`run_wave.rb`'s own `--concurrency` option caps at
|
|
82
|
+
# 8, enforced there, not here: this file is generic and takes no opinion on what a
|
|
83
|
+
# sane ceiling is for any particular caller).
|
|
84
|
+
# @param stop_on_error [Boolean] see the file header.
|
|
85
|
+
# @yieldparam item [Object] one element of `items`.
|
|
86
|
+
# @yieldreturn [Object] becomes that item's `Result#value`.
|
|
87
|
+
# @return [Array<Result>] one per `items` element, in `items`' own order.
|
|
88
|
+
def run(items, concurrency:, stop_on_error: false, &block)
|
|
89
|
+
list = Array(items)
|
|
90
|
+
return [] if list.empty?
|
|
91
|
+
|
|
92
|
+
worker_count = [concurrency.to_i, 1].max
|
|
93
|
+
run_state = RunState.new(
|
|
94
|
+
queue: build_queue(list, worker_count), results: Array.new(list.size),
|
|
95
|
+
abort_state: {aborted: false}, state_mutex: Mutex.new, stop_on_error: stop_on_error, block: block
|
|
96
|
+
)
|
|
97
|
+
|
|
98
|
+
threads = Array.new(worker_count) { Thread.new { worker_loop(run_state) } }
|
|
99
|
+
threads.each(&:join)
|
|
100
|
+
|
|
101
|
+
run_state.results
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
# Bundles the per-run shared state every worker thread reads/writes, so
|
|
105
|
+
# `#worker_loop`/`#run_one` stay under the 5-parameter style ceiling without losing
|
|
106
|
+
# any of it -- plain data, no behaviour of its own. Internal (declared above `private`
|
|
107
|
+
# only because Ruby constants ignore access modifiers -- rubocop Lint/UselessConstantScoping).
|
|
108
|
+
RunState = Struct.new(:queue, :results, :abort_state, :state_mutex, :stop_on_error, :block,
|
|
109
|
+
keyword_init: true)
|
|
110
|
+
|
|
111
|
+
private
|
|
112
|
+
|
|
113
|
+
# One `[index, item]` pair per real item, `worker_count` `:done` sentinels at the
|
|
114
|
+
# tail (one per worker, so every worker eventually sees its own stop signal and the
|
|
115
|
+
# pool's own `Thread#join`s all return -- a `Queue` blocks `#pop` forever otherwise).
|
|
116
|
+
def build_queue(list, worker_count)
|
|
117
|
+
queue = Queue.new
|
|
118
|
+
list.each_with_index {|item, index| queue << [index, item]}
|
|
119
|
+
worker_count.times { queue << :done }
|
|
120
|
+
queue
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
def worker_loop(run_state)
|
|
124
|
+
loop do
|
|
125
|
+
entry = run_state.queue.pop
|
|
126
|
+
break if entry == :done
|
|
127
|
+
|
|
128
|
+
index, item = entry
|
|
129
|
+
if run_state.stop_on_error && run_state.state_mutex.synchronize {run_state.abort_state[:aborted]}
|
|
130
|
+
run_state.results[index] = Result.new(item: item, value: nil, error: nil, skipped: true)
|
|
131
|
+
next
|
|
132
|
+
end
|
|
133
|
+
|
|
134
|
+
run_one(index, item, run_state)
|
|
135
|
+
end
|
|
136
|
+
end
|
|
137
|
+
|
|
138
|
+
def run_one(index, item, run_state)
|
|
139
|
+
value = run_state.block.call(item)
|
|
140
|
+
run_state.results[index] = Result.new(item: item, value: value, error: nil, skipped: false)
|
|
141
|
+
rescue StandardError => e
|
|
142
|
+
run_state.results[index] = Result.new(item: item, value: nil, error: e, skipped: false)
|
|
143
|
+
run_state.state_mutex.synchronize {run_state.abort_state[:aborted] = true} if run_state.stop_on_error
|
|
144
|
+
end
|
|
145
|
+
end
|
|
146
|
+
end
|
|
147
|
+
end
|
|
148
|
+
end
|
|
149
|
+
end
|
|
150
|
+
end
|
|
@@ -0,0 +1,139 @@
|
|
|
1
|
+
# PROMOTED FROM a downstream script repo (services/retry_policy.rb @ 41dedad)
|
|
2
|
+
# NAMESPACE: PROVISIONAL (owner ruling DSL-02) -- kept the original class/module name
|
|
3
|
+
# (`RetryPolicy`) under `Eco::API::Session::Concurrency::*`, the session-level concurrency
|
|
4
|
+
# primitives location proposed by WP-6 (see
|
|
5
|
+
# the repo's internal docs, Option A).
|
|
6
|
+
#
|
|
7
|
+
# CANONICAL, 3.4.0 (release/3.4.0-prep) -- this is the supported subprocess-level retry
|
|
8
|
+
# decision primitive in this gem; see `README.md`'s own "Concurrency primitives" section.
|
|
9
|
+
# `lib/eco/api/session/batch/launcher/retry.rb`'s own `offer_retry_on`/`batch_mode_on` are a
|
|
10
|
+
# DIFFERENT, PRE-EXISTING retry mechanism (interactive Y/n prompt + a fixed
|
|
11
|
+
# `ALLOWED_RETRIES` count around the people-batch HTTP endpoint's own error classes,
|
|
12
|
+
# `Ecoportal::API::Errors::TimeOut`/`StartTimeOut`) -- not a shelled-out-subprocess retry,
|
|
13
|
+
# nothing to switch onto this file. `lib/eco/api/usecases/graphql/samples/location/service/
|
|
14
|
+
# tree_diff.rb`'s `sleep(5)` is an unconditional cooldown pause between comparisons, not a
|
|
15
|
+
# retry loop either. Neither is this file's concern; checked (grep, both directories) at
|
|
16
|
+
# promotion time -- there is no PRE-EXISTING inline concurrency/retry loop that duplicates
|
|
17
|
+
# this file's own job to wire onto it. A FUTURE caller (this gem's own batch/CLI code, or a
|
|
18
|
+
# downstream consumer such as a downstream script's wave runner, the original motivating
|
|
19
|
+
# caller) picks this up directly, as-is.
|
|
20
|
+
#
|
|
21
|
+
# the wave-concurrency work (2026-09-19) -- decides WHETHER/HOW LONG to wait before retrying
|
|
22
|
+
# one shelled-out subprocess call, never sleeps itself (pure decision service, easily specced
|
|
23
|
+
# without a slow test) and never distinguishes idempotent-vs-not (that policy is the CALLER's
|
|
24
|
+
# own job -- `run_wave.rb` only ever calls this for stages it has independently classified as
|
|
25
|
+
# safe to retry -- see that file's own PER-STAGE RETRY TABLE comment).
|
|
26
|
+
#
|
|
27
|
+
# WHY A SUBPROCESS-LEVEL RETRY EXISTS AT ALL, GIVEN THE GEM ALREADY RETRIES: confirmed by
|
|
28
|
+
# reading `ecoportal-api-graphql` (2.2.0) `lib/ecoportal/api/common/graphql/http_client.rb`
|
|
29
|
+
# (`#execute`'s own comment, verbatim): every GraphQL call already routes through the
|
|
30
|
+
# inherited `Common::Client` pipeline ("instrument -> with_retry { rate_throttling }"), which
|
|
31
|
+
# already retries 429/Cloudflare-1015/5xx/connection errors INSIDE one call, transparently, long
|
|
32
|
+
# before either the case code or this file ever sees anything. So by the time a shelled-out
|
|
33
|
+
# `ruby main.rb ...` subprocess actually EXITS non-zero for an HTTP reason, that gem-level
|
|
34
|
+
# budget has ALREADY been exhausted (or the whole process crashed before ever making a call,
|
|
35
|
+
# e.g. a DNS blip at startup). Retrying the WHOLE subprocess is therefore a coarser, OUTER
|
|
36
|
+
# layer of defence -- a fresh process gets a FRESH inner retry budget too (useful for
|
|
37
|
+
# transient token/DNS issues an in-process retry cannot fix) -- not a duplicate of the gem's
|
|
38
|
+
# own retries.
|
|
39
|
+
#
|
|
40
|
+
# THE MARKER THIS FILE READS -- NO CASE-FILE CHANGE NEEDED: grepped every
|
|
41
|
+
# downstream template case file for an HTTP-status-distinguishing marker
|
|
42
|
+
# (429/Retry-After/5xx) -- none add one themselves; they let the gem's own exception surface
|
|
43
|
+
# verbatim on an uncaught crash. That exception already IS a stable, distinguishing marker:
|
|
44
|
+
# `Ecoportal::API::Common::Client::Error::UnexpectedServerError#initialize` (ecoportal-api
|
|
45
|
+
# 0.10.18, lib/ecoportal/api/common/client/error.rb) formats its message
|
|
46
|
+
# `"Code: #{code} -- Error: #{msg}"` -- printed verbatim to STDERR (captured into `execute`'s
|
|
47
|
+
# own combined stdout+stderr) whenever a subprocess crashes uncaught on a non-2xx HTTP
|
|
48
|
+
# response. `HTTP_CODE_PATTERN` below matches THAT exact, already-existing text -- nothing was
|
|
49
|
+
# added to any case file for this to work. Recorded here as an INFERRED-FROM-STATIC-READ
|
|
50
|
+
# finding, not live-verified against a real 429 -- if a live retry never actually fires when
|
|
51
|
+
# expected, re-check this pattern against a real crash's captured output first.
|
|
52
|
+
module Eco
|
|
53
|
+
module API
|
|
54
|
+
class Session
|
|
55
|
+
module Concurrency
|
|
56
|
+
module RetryPolicy
|
|
57
|
+
# HTTP statuses worth retrying the WHOLE subprocess for: rate limiting (429) and the
|
|
58
|
+
# standard 5xx server-error family. Never 4xx other than 429 (a 400/403/404/422 is a
|
|
59
|
+
# REQUEST defect -- retrying it verbatim would just fail the same way every time).
|
|
60
|
+
RETRYABLE_HTTP_CODES = [429, 500, 502, 503, 504].freeze
|
|
61
|
+
|
|
62
|
+
# See the file header -- `Ecoportal::API::Common::Client::Error::UnexpectedServerError`'s
|
|
63
|
+
# own message format, verbatim.
|
|
64
|
+
HTTP_CODE_PATTERN = /Code:\s*(\d+)\s*--\s*Error:/
|
|
65
|
+
|
|
66
|
+
# Case-insensitive: HTTP header names are conventionally capitalised in a raw dump, but
|
|
67
|
+
# nothing here guarantees a subprocess's captured text preserves that casing exactly.
|
|
68
|
+
RETRY_AFTER_PATTERN = /Retry-After:\s*(\d+)/i
|
|
69
|
+
|
|
70
|
+
BASE_DELAY_SECONDS = 1.0
|
|
71
|
+
MAX_DELAY_SECONDS = 30.0
|
|
72
|
+
|
|
73
|
+
# One decision, after ONE attempt has already run and been captured.
|
|
74
|
+
Decision = Struct.new(:should_retry, :delay_seconds, :reason, keyword_init: true) do
|
|
75
|
+
def retry?
|
|
76
|
+
!!should_retry
|
|
77
|
+
end
|
|
78
|
+
end
|
|
79
|
+
|
|
80
|
+
class << self
|
|
81
|
+
# @param output [String] combined stdout+stderr from the attempt just completed.
|
|
82
|
+
# @param status [Process::Status, nil] `nil` when the subprocess itself could not be
|
|
83
|
+
# spawned at all (see `run_wave.rb#execute`'s own `rescue StandardError` branch) --
|
|
84
|
+
# treated as a hard failure, never retryable (nothing about a spawn failure is an
|
|
85
|
+
# HTTP status).
|
|
86
|
+
# @param attempt [Integer] 1-based -- the attempt that JUST ran.
|
|
87
|
+
# @param max_attempts [Integer] total attempts allowed, INCLUDING the first -- so
|
|
88
|
+
# `max_attempts: 3` permits at most 2 retries after the initial attempt.
|
|
89
|
+
# @return [Decision]
|
|
90
|
+
def decide(output:, status:, attempt:, max_attempts: 3)
|
|
91
|
+
return Decision.new(should_retry: false, delay_seconds: 0, reason: :ok) if status&.success?
|
|
92
|
+
if attempt >= max_attempts
|
|
93
|
+
return Decision.new(should_retry: false, delay_seconds: 0,
|
|
94
|
+
reason: :max_attempts_reached)
|
|
95
|
+
end
|
|
96
|
+
|
|
97
|
+
code = retryable_http_code(output)
|
|
98
|
+
return Decision.new(should_retry: false, delay_seconds: 0, reason: :not_retryable) unless code
|
|
99
|
+
|
|
100
|
+
delay = retry_after_seconds(output) || backoff_delay_seconds(attempt)
|
|
101
|
+
Decision.new(should_retry: true, delay_seconds: delay, reason: :"http_#{code}")
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
# @return [Integer, nil] the HTTP status code found in `output`, only when it is one
|
|
105
|
+
# of {RETRYABLE_HTTP_CODES} -- `nil` for anything else (including a code that WAS
|
|
106
|
+
# found but is not on the retryable list, e.g. a 400).
|
|
107
|
+
def retryable_http_code(output)
|
|
108
|
+
match = HTTP_CODE_PATTERN.match(output.to_s)
|
|
109
|
+
return nil unless match
|
|
110
|
+
|
|
111
|
+
code = match[1].to_i
|
|
112
|
+
RETRYABLE_HTTP_CODES.include?(code) ? code : nil
|
|
113
|
+
end
|
|
114
|
+
|
|
115
|
+
# @return [Integer, nil] the server-requested wait, in seconds, when a `Retry-After`
|
|
116
|
+
# header value was captured in the output -- takes priority over the computed
|
|
117
|
+
# backoff below when present (the server knows better than a guess).
|
|
118
|
+
def retry_after_seconds(output)
|
|
119
|
+
match = RETRY_AFTER_PATTERN.match(output.to_s)
|
|
120
|
+
match && match[1].to_i
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
# Bounded exponential backoff with FULL jitter (AWS Architecture Blog's own
|
|
124
|
+
# "Exponential Backoff And Jitter", the well-established shape): a uniform random
|
|
125
|
+
# pick in `[0, capped)`, `capped = min(BASE * 2**(attempt-1), MAX)`. Full jitter
|
|
126
|
+
# (not merely +/- a percentage) is deliberate -- it is what actually de-correlates N
|
|
127
|
+
# concurrent workers that all hit a 429 in the SAME instant (`--concurrency N`'s own
|
|
128
|
+
# reason to exist) from retrying in lockstep and re-triggering the same rate limit.
|
|
129
|
+
# @return [Float] seconds, `0 <= delay < capped`.
|
|
130
|
+
def backoff_delay_seconds(attempt)
|
|
131
|
+
capped = [BASE_DELAY_SECONDS * (2**(attempt - 1)), MAX_DELAY_SECONDS].min
|
|
132
|
+
rand * capped
|
|
133
|
+
end
|
|
134
|
+
end
|
|
135
|
+
end
|
|
136
|
+
end
|
|
137
|
+
end
|
|
138
|
+
end
|
|
139
|
+
end
|
|
@@ -21,9 +21,18 @@ module Eco::Language::Klass
|
|
|
21
21
|
|
|
22
22
|
pending_children.each do |klass|
|
|
23
23
|
@child = klass.new(object)
|
|
24
|
-
rescue
|
|
25
|
-
# Can't create from this class (must be the singleton class)
|
|
26
|
-
#
|
|
24
|
+
rescue StandardError => _e
|
|
25
|
+
# Can't create from this class (must be the singleton class), OR the class was
|
|
26
|
+
# never fully configured for a real registration -- e.g. an anonymous
|
|
27
|
+
# `Class.new(SomeAutoloadedBase)` test double left behind, un-configured, by an
|
|
28
|
+
# ENTIRELY UNRELATED spec elsewhere in the same process (this scan walks
|
|
29
|
+
# `ObjectSpace`, a process-wide, cross-spec-file concern -- see
|
|
30
|
+
# `spec/eco/language/klass/auto_loader_spec.rb`'s own "malformed pending child"
|
|
31
|
+
# example, and `spec/support/fakes/c7_person_entry_fakes.rb`'s header, for two
|
|
32
|
+
# already-documented instances of this exact class of fragility). Best-effort
|
|
33
|
+
# discovery: one malformed/unrelated pending child must never abort loading the
|
|
34
|
+
# REST of the pending children, still less raise into a caller that has nothing
|
|
35
|
+
# to do with it. Just ignore, same as the original TypeError case above.
|
|
27
36
|
ensure
|
|
28
37
|
autoloaded_children.push(klass)
|
|
29
38
|
end
|
data/lib/eco/version.rb
CHANGED
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: eco-helpers
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 3.
|
|
4
|
+
version: 3.4.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Oscar Segura
|
|
@@ -259,20 +259,20 @@ dependencies:
|
|
|
259
259
|
requirements:
|
|
260
260
|
- - "~>"
|
|
261
261
|
- !ruby/object:Gem::Version
|
|
262
|
-
version: '
|
|
262
|
+
version: '3.0'
|
|
263
263
|
- - ">="
|
|
264
264
|
- !ruby/object:Gem::Version
|
|
265
|
-
version:
|
|
265
|
+
version: 3.0.0
|
|
266
266
|
type: :runtime
|
|
267
267
|
prerelease: false
|
|
268
268
|
version_requirements: !ruby/object:Gem::Requirement
|
|
269
269
|
requirements:
|
|
270
270
|
- - "~>"
|
|
271
271
|
- !ruby/object:Gem::Version
|
|
272
|
-
version: '
|
|
272
|
+
version: '3.0'
|
|
273
273
|
- - ">="
|
|
274
274
|
- !ruby/object:Gem::Version
|
|
275
|
-
version:
|
|
275
|
+
version: 3.0.0
|
|
276
276
|
- !ruby/object:Gem::Dependency
|
|
277
277
|
name: ecoportal-api-v2
|
|
278
278
|
requirement: !ruby/object:Gem::Requirement
|
|
@@ -713,6 +713,8 @@ files:
|
|
|
713
713
|
- lib/eco/api/session/batch/policies.rb
|
|
714
714
|
- lib/eco/api/session/batch/searcher.rb
|
|
715
715
|
- lib/eco/api/session/batch/status.rb
|
|
716
|
+
- lib/eco/api/session/concurrency/bounded_worker_pool.rb
|
|
717
|
+
- lib/eco/api/session/concurrency/retry_policy.rb
|
|
716
718
|
- lib/eco/api/session/config.rb
|
|
717
719
|
- lib/eco/api/session/config/api.rb
|
|
718
720
|
- lib/eco/api/session/config/apis.rb
|