scrapeunblocker 0.7.1 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 250fdc0480188972b4735ad8a308cd16df5702a07f98de0d1ce024db61907fe5
4
- data.tar.gz: 82f6efb47746aabab6045e47776709aa56de498e32265e635eb0e258c0514af1
3
+ metadata.gz: f20529c692d1de01f3f02acca1accaf677cd5df2ec9d075bfb949b4b230272d2
4
+ data.tar.gz: b0eff5160bbd56e6e45700609ab17966d061d74537909c7a1a06df9242cddb7a
5
5
  SHA512:
6
- metadata.gz: fb3abf429337da70c37605c7c37d13976da20d69c6b53226bd0eed943bf85a032c8f6c401cb0daa8a432e8b8394a239217b8f0c73d8beda15da648d1724f97e7
7
- data.tar.gz: 95717ff391f8fbf0a5297141edef98a17cea5fcc051dbf0f1a05ce463c8ac7a17d2f89134dfbf48b6f451703ee9ca97042e30053e0ba8b4c6085580aadd83596
6
+ metadata.gz: 2421c59b2bc7c32d769d01546b4f796b7f71237edf500af4ee21b56f0a1393532b6361cc4207264c6072c5510bd31dc840986adda94dfbc18e7d2fdbf03a1df8
7
+ data.tar.gz: 98532dd9a066ef4e52f1123a958e36f1f40df7f603a3055b78784337bedfb256a237343d011e086ba8d0feb94cd17100e2501d5317055b35bae15fc57bea27ea
data/CHANGELOG.md CHANGED
@@ -1,5 +1,18 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.9.0 (2026-10-07)
4
+
5
+ - New `ScrapeUnblocker::BudgetExceededError` (a subclass of `PaymentRequiredError`): raised when the API answers 402 with `User set budget exceeded`, meaning this billing period's spend reached the monthly budget limit (EUR, excluding VAT) set in the [profile](https://app.scrapeunblocker.com/dashboard/profile). It clears at the start of the next billing period, or within about a minute after the limit is raised or removed. Like the other 402s it is never retried and not billed. Before, this body surfaced as a plain `PaymentRequiredError`.
6
+
7
+ No breaking changes: `rescue ScrapeUnblocker::PaymentRequiredError` still catches it.
8
+
9
+ ## 0.8.0 (2026-10-06)
10
+
11
+ - `get_parsed` on a page that rendered but held no structured data now returns a normal `ParsedPage` instead of raising `NoDataExtractedError`. The API answers this case with a billed 200: `data` is empty, the new `data_extracted?` is `false`, the new `detail` explains it and the new `html` carries the rendered page, so no second call is needed. On a successful parse `data_extracted?` is `true` and `html` / `detail` are `nil`.
12
+ - `NoDataExtractedError` is deprecated: the API no longer sends the 422 `no_data_extracted`. The class stays so existing `rescue` clauses keep loading.
13
+
14
+ Behaviour change: code that rescued `NoDataExtractedError` should check `page.data_extracted?` instead.
15
+
3
16
  ## 0.7.1 (2026-09-30)
4
17
 
5
18
  - New `ScrapeUnblocker::NoDataExtractedError` (a subclass of `ValidationError`): raised by `get_parsed` when the page rendered but no structured data could be extracted from it. The API answers 422 with `{"error": "no_data_extracted", "detail": ...}`; the error carries `detail`. The call is not billed and is never retried - use `get_page_source` for the HTML. Before, this came back as a billed 200 with empty `data`.
data/README.md CHANGED
@@ -108,7 +108,6 @@ result["elements"].each { |el| p el }
108
108
  ```ruby
109
109
  result = su.get_parsed("https://www.walmart.com/ip/12345")
110
110
  puts result.page_type # e.g. "product"
111
- puts result.source # how it was extracted
112
111
  p result.data # the fields
113
112
 
114
113
  # If a parse ever comes back wrong, force a fresh set of rules:
@@ -259,7 +258,7 @@ begin
259
258
  rescue ScrapeUnblocker::BlockedError
260
259
  # 403: the target blocked every bypass path (not billed)
261
260
  rescue ScrapeUnblocker::PaymentRequiredError
262
- # 402: quota, credit limit, or a failed payment - fix billing
261
+ # 402: quota, credit limit, your budget limit, or a failed payment - fix billing
263
262
  rescue ScrapeUnblocker::RateLimitError
264
263
  # 429: slow down
265
264
  rescue ScrapeUnblocker::UpstreamOutageError
@@ -272,9 +271,10 @@ end
272
271
  | `InvalidRequestError` | 400 | Bad URL, unsupported scheme, or the API key header was not sent |
273
272
  | `AuthenticationError` | 401 | Key not recognised - typo, stray whitespace, or a rotated key |
274
273
  | `NoSubscriptionError` | 401 | Key is fine, but the account has no active plan |
275
- | `PaymentRequiredError` | 402 | Billing block - base class for the three below |
274
+ | `PaymentRequiredError` | 402 | Billing block - base class for the four below |
276
275
  | `QuotaExceededError` | 402 | The plan's requests for this period are used up |
277
276
  | `CreditLimitExceededError` | 402 | Unpaid balance is past the account's credit limit |
277
+ | `BudgetExceededError` | 402 | This billing period's spend reached the monthly budget limit you set |
278
278
  | `PaymentFailedError` | 402 | A card payment was declined three times |
279
279
  | `BlockedError` | 403 | Blocked by bot protection on every path |
280
280
  | `NotFoundError` | 404 | What you asked for does not exist - no image on the page (`get_image`), or a plugin lookup found nothing |
@@ -282,7 +282,6 @@ end
282
282
  | `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
283
283
  | `UnsupportedContentError` | 415 | The URL serves something other than HTML |
284
284
  | `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
285
- | `NoDataExtractedError` | 422 | `get_parsed`: the page rendered but held no structured data; carries `detail` (subclass of `ValidationError`, not billed) |
286
285
  | `RateLimitError` | 429 | Too many requests |
287
286
  | `UpstreamOutageError` | 503 | The target origin is down |
288
287
  | `ServerError` | 5xx | Unexpected server error, including a 504 upstream timeout |
@@ -306,26 +305,25 @@ end
306
305
 
307
306
  With `get_parsed` the body is the parsed-data JSON (`{"data": {"page_type": "not_found", ...}}`), so `html` is `nil` there; the raw JSON is on `#body`.
308
307
 
309
- ### No structured data on the page (422)
308
+ ### No structured data on the page
310
309
 
311
- When `get_parsed` renders the page but can extract no structured data from it, the API answers 422 and the client raises `ScrapeUnblocker::NoDataExtractedError`. The call is **not billed**, and retrying gives the same result - fetch the HTML with `get_page_source` instead:
310
+ When `get_parsed` renders the page but can extract no structured data from it, the API still answers 200 and the client returns a normal `ParsedPage`: `data` is empty, `data_extracted?` is `false`, `detail` says why, and `html` carries the rendered page. The call is billed like any delivered page, so you already have the HTML - no second call needed:
312
311
 
313
312
  ```ruby
314
- begin
315
- page = su.get_parsed("https://example.com/some-page")
316
- rescue ScrapeUnblocker::NoDataExtractedError => e
317
- puts e.detail # the API's explanation
318
- html = su.get_page_source("https://example.com/some-page")
313
+ page = su.get_parsed("https://example.com/some-page")
314
+ unless page.data_extracted?
315
+ puts page.detail # the API's explanation
316
+ html = page.html # the rendered page, parse it yourself
319
317
  end
320
318
  ```
321
319
 
322
- `NoDataExtractedError` subclasses `ValidationError`, so `rescue ScrapeUnblocker::ValidationError` catches it too.
320
+ On a successful parse `data_extracted?` is `true` and `html` is `nil`. Up to 0.7.1 this case raised `NoDataExtractedError` (a 422); the API no longer sends that, and the class stays only so existing `rescue` clauses keep loading.
323
321
 
324
322
  Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
325
323
 
326
324
  ### Billing errors (402)
327
325
 
328
- The three billing blocks share a status code and differ only in their message, so the client raises a dedicated error for each:
326
+ The four billing blocks share a status code and differ only in their message, so the client raises a dedicated error for each:
329
327
 
330
328
  ```ruby
331
329
  begin
@@ -334,12 +332,16 @@ rescue ScrapeUnblocker::QuotaExceededError
334
332
  # plan quota (plus any overage allowance) is used up for this period
335
333
  rescue ScrapeUnblocker::CreditLimitExceededError
336
334
  # unpaid balance passed the account credit limit
335
+ rescue ScrapeUnblocker::BudgetExceededError
336
+ # this period's spend reached the monthly budget limit set in your profile
337
337
  rescue ScrapeUnblocker::PaymentFailedError
338
338
  # card declined three times - update the payment method
339
339
  end
340
340
  ```
341
341
 
342
- When more than one applies, the most serious wins: failed payment outranks credit limit, which outranks quota. All three lift by themselves once the billing state changes - access returns within about a minute, and the API key stays the same. One catch worth knowing: subscribing to a new plan does **not** clear `PaymentFailedError`, because the old unpaid invoice stays open until it is paid.
342
+ When more than one applies, the most serious wins: failed payment outranks credit limit, which outranks quota, which outranks your own budget limit. All four lift by themselves once the billing state changes - access returns within about a minute, and the API key stays the same. One catch worth knowing: subscribing to a new plan does **not** clear `PaymentFailedError`, because the old unpaid invoice stays open until it is paid.
343
+
344
+ `BudgetExceededError` (body `User set budget exceeded`) means this billing period's spend reached the monthly budget limit you set in your [profile](https://app.scrapeunblocker.com/dashboard/profile?utm_source=rubygems&utm_medium=integration&utm_campaign=ruby-sdk) (EUR, excluding VAT). The key works again at the start of the next billing period, or within about a minute after you raise or remove the limit.
343
345
 
344
346
  Full details for every status code: [docs.scrapeunblocker.com/errors](https://docs.scrapeunblocker.com/errors).
345
347
 
@@ -74,10 +74,11 @@ module ScrapeUnblocker
74
74
 
75
75
  # Fetch a URL and return structured JSON instead of HTML.
76
76
  #
77
- # Raises ScrapeUnblocker::NoDataExtractedError (not billed) when the page
78
- # rendered but held no structured data - use #get_page_source for the HTML -
79
- # and ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the
80
- # target page itself answered 404 or 410.
77
+ # A page that rendered but held no structured data comes back as a
78
+ # ParsedPage with #data_extracted? false and the rendered page on #html
79
+ # (billed like any delivered page). Raises
80
+ # ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the target
81
+ # page itself answered 404 or 410.
81
82
  def get_parsed(url, proxy_country: nil, time_sleep: nil, refresh_rules: false, rules_hint: nil)
82
83
  body = request("/getPageSource",
83
84
  url: url, parsed_data: true, proxy_country: proxy_country,
@@ -36,15 +36,16 @@ module ScrapeUnblocker
36
36
  # The account has a billing problem (HTTP 402).
37
37
  #
38
38
  # Credentials are fine - the request was stopped for a billing reason. There
39
- # are three, each raised as a dedicated subclass: QuotaExceededError,
40
- # CreditLimitExceededError and PaymentFailedError. Rescue this base class to
41
- # handle all three.
39
+ # are four, each raised as a dedicated subclass: QuotaExceededError,
40
+ # CreditLimitExceededError, BudgetExceededError and PaymentFailedError.
41
+ # Rescue this base class to handle all four.
42
42
  #
43
43
  # When more than one applies, the most serious wins: failed payment outranks
44
- # credit limit, which outranks quota. All three lift by themselves once the
45
- # billing state changes - access returns within roughly a minute, with no key
46
- # change needed. Like a 401, a 402 is refused before anything is scraped, so
47
- # it is never billed. Retrying is pointless; fix the billing state first.
44
+ # credit limit, which outranks quota, which outranks your own budget limit.
45
+ # All four lift by themselves once the billing state changes - access returns
46
+ # within roughly a minute, with no key change needed. Like a 401, a 402 is
47
+ # refused before anything is scraped, so it is never billed. Retrying is
48
+ # pointless; fix the billing state first.
48
49
  class PaymentRequiredError < APIError; end
49
50
 
50
51
  # Every request the plan allows this period has been used (HTTP 402).
@@ -64,6 +65,17 @@ module ScrapeUnblocker
64
65
  # usually clears itself within about a minute.
65
66
  class CreditLimitExceededError < PaymentRequiredError; end
66
67
 
68
+ # This billing period's spend reached the monthly budget limit you set
69
+ # (HTTP 402).
70
+ #
71
+ # The limit is set in your profile (EUR, excluding VAT), and spend is counted
72
+ # the way the invoice is: the plan's fixed monthly fee, if any, plus the
73
+ # requests billed on top of it. Requests paid from coupon credit do not count.
74
+ # The key works again at the start of the next billing period, or within
75
+ # about a minute after you raise or remove the limit at
76
+ # https://app.scrapeunblocker.com/dashboard/profile.
77
+ class BudgetExceededError < PaymentRequiredError; end
78
+
67
79
  # A card payment has been declined three times (HTTP 402).
68
80
  #
69
81
  # Those attempts are the payment provider's automatic retries spread over
@@ -128,13 +140,13 @@ module ScrapeUnblocker
128
140
  # each problem field. Read it from #body.
129
141
  class ValidationError < APIError; end
130
142
 
131
- # The page rendered but no structured data came out of it (HTTP 422).
143
+ # Deprecated: the API no longer sends this 422.
132
144
  #
133
- # Raised by #get_parsed when the API loaded the page but could not extract any
134
- # structured fields from it. The API answers 422 with a JSON body of
135
- # {"error": "no_data_extracted", "detail": ...}; +detail+ holds the API's
136
- # explanation. The call is not billed and retrying returns the same answer;
137
- # call #get_page_source for the HTML.
145
+ # A page that rendered but held no structured data now comes back from
146
+ # #get_parsed as a normal ParsedPage with #data_extracted? false and the page
147
+ # on #html (a billed 200). The class stays so existing rescue clauses keep
148
+ # loading; it is raised only for a legacy 422 body of
149
+ # {"error": "no_data_extracted", "detail": ...}.
138
150
  class NoDataExtractedError < ValidationError
139
151
  attr_reader :detail
140
152
 
@@ -167,7 +179,7 @@ module ScrapeUnblocker
167
179
  BASE_MESSAGES = {
168
180
  400 => "Invalid request (bad URL, unsupported scheme, or missing API key header)",
169
181
  401 => "Authentication failed - key not recognised, or account has no active plan",
170
- 402 => "Billing block - quota exceeded, credit limit exceeded, or a failed payment",
182
+ 402 => "Billing block - quota exceeded, credit limit exceeded, budget limit reached, or a failed payment",
171
183
  403 => "Target blocked by bot protection on every bypass path",
172
184
  404 => "Requested element not found on the page",
173
185
  408 => "Browser run timed out before the page was ready",
@@ -187,12 +199,13 @@ module ScrapeUnblocker
187
199
  end
188
200
  private_class_method :auth_error_class
189
201
 
190
- # The three billing blocks share a status code and differ only in their
202
+ # The four billing blocks share a status code and differ only in their
191
203
  # plain-text body. An unrecognised body falls back to PaymentRequiredError.
192
204
  def self.billing_error_class(body)
193
205
  text = (body || "").downcase
194
206
  return QuotaExceededError if text.include?("quota exceeded")
195
207
  return CreditLimitExceededError if text.include?("credit limit exceeded")
208
+ return BudgetExceededError if text.include?("user set budget exceeded")
196
209
  return PaymentFailedError if text.include?("payment failed")
197
210
 
198
211
  PaymentRequiredError
@@ -222,9 +235,9 @@ module ScrapeUnblocker
222
235
  end
223
236
  private_class_method :target_not_found_error
224
237
 
225
- # parsed_data answers 422 with {"error": "no_data_extracted", "detail"} when
226
- # the page rendered but held no structured data. Anything else returns nil so
227
- # the general ValidationError applies.
238
+ # Maps a legacy 422 {"error": "no_data_extracted", "detail"} body; the API now
239
+ # answers an empty parse with a 200 (data_extracted: false). Anything else
240
+ # returns nil so the general ValidationError applies.
228
241
  def self.no_data_extracted_error(status, body)
229
242
  return nil unless status == 422
230
243
 
@@ -11,12 +11,25 @@ module ScrapeUnblocker
11
11
  attr_reader :data
12
12
  # @return [Hash] the full JSON payload as returned by the API
13
13
  attr_reader :raw
14
+ # @return [String, nil] the rendered page when #data_extracted? is false
15
+ attr_reader :html
16
+ # @return [String, nil] the API's explanation when #data_extracted? is false
17
+ attr_reader :detail
14
18
 
15
- def initialize(page_type:, source:, data:, raw:)
19
+ def initialize(page_type:, source:, data:, raw:, data_extracted: true, html: nil, detail: nil)
16
20
  @page_type = page_type
17
21
  @source = source
18
22
  @data = data
19
23
  @raw = raw
24
+ @data_extracted = data_extracted
25
+ @html = html
26
+ @detail = detail
27
+ end
28
+
29
+ # @return [Boolean] false when the page rendered but no structured data
30
+ # could be extracted from it (#data is then empty, #html holds the page)
31
+ def data_extracted?
32
+ @data_extracted
20
33
  end
21
34
 
22
35
  def self.from_hash(payload)
@@ -25,7 +38,10 @@ module ScrapeUnblocker
25
38
  page_type: inner["page_type"],
26
39
  source: inner["source"],
27
40
  data: inner["data"],
28
- raw: payload
41
+ raw: payload,
42
+ data_extracted: payload["data_extracted"] != false,
43
+ html: payload["html"],
44
+ detail: payload["detail"]
29
45
  )
30
46
  end
31
47
  end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module ScrapeUnblocker
4
- VERSION = "0.7.1"
4
+ VERSION = "0.9.0"
5
5
  end
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: scrapeunblocker
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.7.1
4
+ version: 0.9.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - ScrapeUnblocker
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-30 00:00:00.000000000 Z
11
+ date: 2026-10-07 00:00:00.000000000 Z
12
12
  dependencies: []
13
13
  description: JS-rendered pages that bypass Cloudflare, DataDome, PerimeterX and Akamai,
14
14
  plus Google SERP and Skyscanner flights/hotels/car-hire scraping as JSON.