scrapeunblocker 0.7.1 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +13 -0
- data/README.md +16 -14
- data/lib/scrapeunblocker/client.rb +5 -4
- data/lib/scrapeunblocker/errors.rb +31 -18
- data/lib/scrapeunblocker/parsed_page.rb +18 -2
- data/lib/scrapeunblocker/version.rb +1 -1
- metadata +2 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: f20529c692d1de01f3f02acca1accaf677cd5df2ec9d075bfb949b4b230272d2
|
|
4
|
+
data.tar.gz: b0eff5160bbd56e6e45700609ab17966d061d74537909c7a1a06df9242cddb7a
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 2421c59b2bc7c32d769d01546b4f796b7f71237edf500af4ee21b56f0a1393532b6361cc4207264c6072c5510bd31dc840986adda94dfbc18e7d2fdbf03a1df8
|
|
7
|
+
data.tar.gz: 98532dd9a066ef4e52f1123a958e36f1f40df7f603a3055b78784337bedfb256a237343d011e086ba8d0feb94cd17100e2501d5317055b35bae15fc57bea27ea
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.9.0 (2026-10-07)
|
|
4
|
+
|
|
5
|
+
- New `ScrapeUnblocker::BudgetExceededError` (a subclass of `PaymentRequiredError`): raised when the API answers 402 with `User set budget exceeded`, meaning this billing period's spend reached the monthly budget limit (EUR, excluding VAT) set in the [profile](https://app.scrapeunblocker.com/dashboard/profile). It clears at the start of the next billing period, or within about a minute after the limit is raised or removed. Like the other 402s it is never retried and not billed. Before, this body surfaced as a plain `PaymentRequiredError`.
|
|
6
|
+
|
|
7
|
+
No breaking changes: `rescue ScrapeUnblocker::PaymentRequiredError` still catches it.
|
|
8
|
+
|
|
9
|
+
## 0.8.0 (2026-10-06)
|
|
10
|
+
|
|
11
|
+
- `get_parsed` on a page that rendered but held no structured data now returns a normal `ParsedPage` instead of raising `NoDataExtractedError`. The API answers this case with a billed 200: `data` is empty, the new `data_extracted?` is `false`, the new `detail` explains it and the new `html` carries the rendered page, so no second call is needed. On a successful parse `data_extracted?` is `true` and `html` / `detail` are `nil`.
|
|
12
|
+
- `NoDataExtractedError` is deprecated: the API no longer sends the 422 `no_data_extracted`. The class stays so existing `rescue` clauses keep loading.
|
|
13
|
+
|
|
14
|
+
Behaviour change: code that rescued `NoDataExtractedError` should check `page.data_extracted?` instead.
|
|
15
|
+
|
|
3
16
|
## 0.7.1 (2026-09-30)
|
|
4
17
|
|
|
5
18
|
- New `ScrapeUnblocker::NoDataExtractedError` (a subclass of `ValidationError`): raised by `get_parsed` when the page rendered but no structured data could be extracted from it. The API answers 422 with `{"error": "no_data_extracted", "detail": ...}`; the error carries `detail`. The call is not billed and is never retried - use `get_page_source` for the HTML. Before, this came back as a billed 200 with empty `data`.
|
data/README.md
CHANGED
|
@@ -108,7 +108,6 @@ result["elements"].each { |el| p el }
|
|
|
108
108
|
```ruby
|
|
109
109
|
result = su.get_parsed("https://www.walmart.com/ip/12345")
|
|
110
110
|
puts result.page_type # e.g. "product"
|
|
111
|
-
puts result.source # how it was extracted
|
|
112
111
|
p result.data # the fields
|
|
113
112
|
|
|
114
113
|
# If a parse ever comes back wrong, force a fresh set of rules:
|
|
@@ -259,7 +258,7 @@ begin
|
|
|
259
258
|
rescue ScrapeUnblocker::BlockedError
|
|
260
259
|
# 403: the target blocked every bypass path (not billed)
|
|
261
260
|
rescue ScrapeUnblocker::PaymentRequiredError
|
|
262
|
-
# 402: quota, credit limit, or a failed payment - fix billing
|
|
261
|
+
# 402: quota, credit limit, your budget limit, or a failed payment - fix billing
|
|
263
262
|
rescue ScrapeUnblocker::RateLimitError
|
|
264
263
|
# 429: slow down
|
|
265
264
|
rescue ScrapeUnblocker::UpstreamOutageError
|
|
@@ -272,9 +271,10 @@ end
|
|
|
272
271
|
| `InvalidRequestError` | 400 | Bad URL, unsupported scheme, or the API key header was not sent |
|
|
273
272
|
| `AuthenticationError` | 401 | Key not recognised - typo, stray whitespace, or a rotated key |
|
|
274
273
|
| `NoSubscriptionError` | 401 | Key is fine, but the account has no active plan |
|
|
275
|
-
| `PaymentRequiredError` | 402 | Billing block - base class for the
|
|
274
|
+
| `PaymentRequiredError` | 402 | Billing block - base class for the four below |
|
|
276
275
|
| `QuotaExceededError` | 402 | The plan's requests for this period are used up |
|
|
277
276
|
| `CreditLimitExceededError` | 402 | Unpaid balance is past the account's credit limit |
|
|
277
|
+
| `BudgetExceededError` | 402 | This billing period's spend reached the monthly budget limit you set |
|
|
278
278
|
| `PaymentFailedError` | 402 | A card payment was declined three times |
|
|
279
279
|
| `BlockedError` | 403 | Blocked by bot protection on every path |
|
|
280
280
|
| `NotFoundError` | 404 | What you asked for does not exist - no image on the page (`get_image`), or a plugin lookup found nothing |
|
|
@@ -282,7 +282,6 @@ end
|
|
|
282
282
|
| `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
|
|
283
283
|
| `UnsupportedContentError` | 415 | The URL serves something other than HTML |
|
|
284
284
|
| `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
|
|
285
|
-
| `NoDataExtractedError` | 422 | `get_parsed`: the page rendered but held no structured data; carries `detail` (subclass of `ValidationError`, not billed) |
|
|
286
285
|
| `RateLimitError` | 429 | Too many requests |
|
|
287
286
|
| `UpstreamOutageError` | 503 | The target origin is down |
|
|
288
287
|
| `ServerError` | 5xx | Unexpected server error, including a 504 upstream timeout |
|
|
@@ -306,26 +305,25 @@ end
|
|
|
306
305
|
|
|
307
306
|
With `get_parsed` the body is the parsed-data JSON (`{"data": {"page_type": "not_found", ...}}`), so `html` is `nil` there; the raw JSON is on `#body`.
|
|
308
307
|
|
|
309
|
-
### No structured data on the page
|
|
308
|
+
### No structured data on the page
|
|
310
309
|
|
|
311
|
-
When `get_parsed` renders the page but can extract no structured data from it, the API answers
|
|
310
|
+
When `get_parsed` renders the page but can extract no structured data from it, the API still answers 200 and the client returns a normal `ParsedPage`: `data` is empty, `data_extracted?` is `false`, `detail` says why, and `html` carries the rendered page. The call is billed like any delivered page, so you already have the HTML - no second call needed:
|
|
312
311
|
|
|
313
312
|
```ruby
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
html = su.get_page_source("https://example.com/some-page")
|
|
313
|
+
page = su.get_parsed("https://example.com/some-page")
|
|
314
|
+
unless page.data_extracted?
|
|
315
|
+
puts page.detail # the API's explanation
|
|
316
|
+
html = page.html # the rendered page, parse it yourself
|
|
319
317
|
end
|
|
320
318
|
```
|
|
321
319
|
|
|
322
|
-
`
|
|
320
|
+
On a successful parse `data_extracted?` is `true` and `html` is `nil`. Up to 0.7.1 this case raised `NoDataExtractedError` (a 422); the API no longer sends that, and the class stays only so existing `rescue` clauses keep loading.
|
|
323
321
|
|
|
324
322
|
Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
|
|
325
323
|
|
|
326
324
|
### Billing errors (402)
|
|
327
325
|
|
|
328
|
-
The
|
|
326
|
+
The four billing blocks share a status code and differ only in their message, so the client raises a dedicated error for each:
|
|
329
327
|
|
|
330
328
|
```ruby
|
|
331
329
|
begin
|
|
@@ -334,12 +332,16 @@ rescue ScrapeUnblocker::QuotaExceededError
|
|
|
334
332
|
# plan quota (plus any overage allowance) is used up for this period
|
|
335
333
|
rescue ScrapeUnblocker::CreditLimitExceededError
|
|
336
334
|
# unpaid balance passed the account credit limit
|
|
335
|
+
rescue ScrapeUnblocker::BudgetExceededError
|
|
336
|
+
# this period's spend reached the monthly budget limit set in your profile
|
|
337
337
|
rescue ScrapeUnblocker::PaymentFailedError
|
|
338
338
|
# card declined three times - update the payment method
|
|
339
339
|
end
|
|
340
340
|
```
|
|
341
341
|
|
|
342
|
-
When more than one applies, the most serious wins: failed payment outranks credit limit, which outranks quota. All
|
|
342
|
+
When more than one applies, the most serious wins: failed payment outranks credit limit, which outranks quota, which outranks your own budget limit. All four lift by themselves once the billing state changes - access returns within about a minute, and the API key stays the same. One catch worth knowing: subscribing to a new plan does **not** clear `PaymentFailedError`, because the old unpaid invoice stays open until it is paid.
|
|
343
|
+
|
|
344
|
+
`BudgetExceededError` (body `User set budget exceeded`) means this billing period's spend reached the monthly budget limit you set in your [profile](https://app.scrapeunblocker.com/dashboard/profile?utm_source=rubygems&utm_medium=integration&utm_campaign=ruby-sdk) (EUR, excluding VAT). The key works again at the start of the next billing period, or within about a minute after you raise or remove the limit.
|
|
343
345
|
|
|
344
346
|
Full details for every status code: [docs.scrapeunblocker.com/errors](https://docs.scrapeunblocker.com/errors).
|
|
345
347
|
|
|
@@ -74,10 +74,11 @@ module ScrapeUnblocker
|
|
|
74
74
|
|
|
75
75
|
# Fetch a URL and return structured JSON instead of HTML.
|
|
76
76
|
#
|
|
77
|
-
#
|
|
78
|
-
#
|
|
79
|
-
#
|
|
80
|
-
#
|
|
77
|
+
# A page that rendered but held no structured data comes back as a
|
|
78
|
+
# ParsedPage with #data_extracted? false and the rendered page on #html
|
|
79
|
+
# (billed like any delivered page). Raises
|
|
80
|
+
# ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the target
|
|
81
|
+
# page itself answered 404 or 410.
|
|
81
82
|
def get_parsed(url, proxy_country: nil, time_sleep: nil, refresh_rules: false, rules_hint: nil)
|
|
82
83
|
body = request("/getPageSource",
|
|
83
84
|
url: url, parsed_data: true, proxy_country: proxy_country,
|
|
@@ -36,15 +36,16 @@ module ScrapeUnblocker
|
|
|
36
36
|
# The account has a billing problem (HTTP 402).
|
|
37
37
|
#
|
|
38
38
|
# Credentials are fine - the request was stopped for a billing reason. There
|
|
39
|
-
# are
|
|
40
|
-
# CreditLimitExceededError and PaymentFailedError.
|
|
41
|
-
# handle all
|
|
39
|
+
# are four, each raised as a dedicated subclass: QuotaExceededError,
|
|
40
|
+
# CreditLimitExceededError, BudgetExceededError and PaymentFailedError.
|
|
41
|
+
# Rescue this base class to handle all four.
|
|
42
42
|
#
|
|
43
43
|
# When more than one applies, the most serious wins: failed payment outranks
|
|
44
|
-
# credit limit, which outranks quota
|
|
45
|
-
#
|
|
46
|
-
# change needed. Like a 401, a 402 is
|
|
47
|
-
# it is never billed. Retrying is
|
|
44
|
+
# credit limit, which outranks quota, which outranks your own budget limit.
|
|
45
|
+
# All four lift by themselves once the billing state changes - access returns
|
|
46
|
+
# within roughly a minute, with no key change needed. Like a 401, a 402 is
|
|
47
|
+
# refused before anything is scraped, so it is never billed. Retrying is
|
|
48
|
+
# pointless; fix the billing state first.
|
|
48
49
|
class PaymentRequiredError < APIError; end
|
|
49
50
|
|
|
50
51
|
# Every request the plan allows this period has been used (HTTP 402).
|
|
@@ -64,6 +65,17 @@ module ScrapeUnblocker
|
|
|
64
65
|
# usually clears itself within about a minute.
|
|
65
66
|
class CreditLimitExceededError < PaymentRequiredError; end
|
|
66
67
|
|
|
68
|
+
# This billing period's spend reached the monthly budget limit you set
|
|
69
|
+
# (HTTP 402).
|
|
70
|
+
#
|
|
71
|
+
# The limit is set in your profile (EUR, excluding VAT), and spend is counted
|
|
72
|
+
# the way the invoice is: the plan's fixed monthly fee, if any, plus the
|
|
73
|
+
# requests billed on top of it. Requests paid from coupon credit do not count.
|
|
74
|
+
# The key works again at the start of the next billing period, or within
|
|
75
|
+
# about a minute after you raise or remove the limit at
|
|
76
|
+
# https://app.scrapeunblocker.com/dashboard/profile.
|
|
77
|
+
class BudgetExceededError < PaymentRequiredError; end
|
|
78
|
+
|
|
67
79
|
# A card payment has been declined three times (HTTP 402).
|
|
68
80
|
#
|
|
69
81
|
# Those attempts are the payment provider's automatic retries spread over
|
|
@@ -128,13 +140,13 @@ module ScrapeUnblocker
|
|
|
128
140
|
# each problem field. Read it from #body.
|
|
129
141
|
class ValidationError < APIError; end
|
|
130
142
|
|
|
131
|
-
#
|
|
143
|
+
# Deprecated: the API no longer sends this 422.
|
|
132
144
|
#
|
|
133
|
-
#
|
|
134
|
-
#
|
|
135
|
-
#
|
|
136
|
-
#
|
|
137
|
-
#
|
|
145
|
+
# A page that rendered but held no structured data now comes back from
|
|
146
|
+
# #get_parsed as a normal ParsedPage with #data_extracted? false and the page
|
|
147
|
+
# on #html (a billed 200). The class stays so existing rescue clauses keep
|
|
148
|
+
# loading; it is raised only for a legacy 422 body of
|
|
149
|
+
# {"error": "no_data_extracted", "detail": ...}.
|
|
138
150
|
class NoDataExtractedError < ValidationError
|
|
139
151
|
attr_reader :detail
|
|
140
152
|
|
|
@@ -167,7 +179,7 @@ module ScrapeUnblocker
|
|
|
167
179
|
BASE_MESSAGES = {
|
|
168
180
|
400 => "Invalid request (bad URL, unsupported scheme, or missing API key header)",
|
|
169
181
|
401 => "Authentication failed - key not recognised, or account has no active plan",
|
|
170
|
-
402 => "Billing block - quota exceeded, credit limit exceeded, or a failed payment",
|
|
182
|
+
402 => "Billing block - quota exceeded, credit limit exceeded, budget limit reached, or a failed payment",
|
|
171
183
|
403 => "Target blocked by bot protection on every bypass path",
|
|
172
184
|
404 => "Requested element not found on the page",
|
|
173
185
|
408 => "Browser run timed out before the page was ready",
|
|
@@ -187,12 +199,13 @@ module ScrapeUnblocker
|
|
|
187
199
|
end
|
|
188
200
|
private_class_method :auth_error_class
|
|
189
201
|
|
|
190
|
-
# The
|
|
202
|
+
# The four billing blocks share a status code and differ only in their
|
|
191
203
|
# plain-text body. An unrecognised body falls back to PaymentRequiredError.
|
|
192
204
|
def self.billing_error_class(body)
|
|
193
205
|
text = (body || "").downcase
|
|
194
206
|
return QuotaExceededError if text.include?("quota exceeded")
|
|
195
207
|
return CreditLimitExceededError if text.include?("credit limit exceeded")
|
|
208
|
+
return BudgetExceededError if text.include?("user set budget exceeded")
|
|
196
209
|
return PaymentFailedError if text.include?("payment failed")
|
|
197
210
|
|
|
198
211
|
PaymentRequiredError
|
|
@@ -222,9 +235,9 @@ module ScrapeUnblocker
|
|
|
222
235
|
end
|
|
223
236
|
private_class_method :target_not_found_error
|
|
224
237
|
|
|
225
|
-
#
|
|
226
|
-
#
|
|
227
|
-
# the general ValidationError applies.
|
|
238
|
+
# Maps a legacy 422 {"error": "no_data_extracted", "detail"} body; the API now
|
|
239
|
+
# answers an empty parse with a 200 (data_extracted: false). Anything else
|
|
240
|
+
# returns nil so the general ValidationError applies.
|
|
228
241
|
def self.no_data_extracted_error(status, body)
|
|
229
242
|
return nil unless status == 422
|
|
230
243
|
|
|
@@ -11,12 +11,25 @@ module ScrapeUnblocker
|
|
|
11
11
|
attr_reader :data
|
|
12
12
|
# @return [Hash] the full JSON payload as returned by the API
|
|
13
13
|
attr_reader :raw
|
|
14
|
+
# @return [String, nil] the rendered page when #data_extracted? is false
|
|
15
|
+
attr_reader :html
|
|
16
|
+
# @return [String, nil] the API's explanation when #data_extracted? is false
|
|
17
|
+
attr_reader :detail
|
|
14
18
|
|
|
15
|
-
def initialize(page_type:, source:, data:, raw:)
|
|
19
|
+
def initialize(page_type:, source:, data:, raw:, data_extracted: true, html: nil, detail: nil)
|
|
16
20
|
@page_type = page_type
|
|
17
21
|
@source = source
|
|
18
22
|
@data = data
|
|
19
23
|
@raw = raw
|
|
24
|
+
@data_extracted = data_extracted
|
|
25
|
+
@html = html
|
|
26
|
+
@detail = detail
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
# @return [Boolean] false when the page rendered but no structured data
|
|
30
|
+
# could be extracted from it (#data is then empty, #html holds the page)
|
|
31
|
+
def data_extracted?
|
|
32
|
+
@data_extracted
|
|
20
33
|
end
|
|
21
34
|
|
|
22
35
|
def self.from_hash(payload)
|
|
@@ -25,7 +38,10 @@ module ScrapeUnblocker
|
|
|
25
38
|
page_type: inner["page_type"],
|
|
26
39
|
source: inner["source"],
|
|
27
40
|
data: inner["data"],
|
|
28
|
-
raw: payload
|
|
41
|
+
raw: payload,
|
|
42
|
+
data_extracted: payload["data_extracted"] != false,
|
|
43
|
+
html: payload["html"],
|
|
44
|
+
detail: payload["detail"]
|
|
29
45
|
)
|
|
30
46
|
end
|
|
31
47
|
end
|
metadata
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: scrapeunblocker
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.9.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- ScrapeUnblocker
|
|
8
8
|
autorequire:
|
|
9
9
|
bindir: bin
|
|
10
10
|
cert_chain: []
|
|
11
|
-
date: 2026-
|
|
11
|
+
date: 2026-10-07 00:00:00.000000000 Z
|
|
12
12
|
dependencies: []
|
|
13
13
|
description: JS-rendered pages that bypass Cloudflare, DataDome, PerimeterX and Akamai,
|
|
14
14
|
plus Google SERP and Skyscanner flights/hotels/car-hire scraping as JSON.
|