scrapeunblocker 0.7.1 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +7 -0
- data/README.md +7 -9
- data/lib/scrapeunblocker/client.rb +5 -4
- data/lib/scrapeunblocker/errors.rb +9 -9
- data/lib/scrapeunblocker/parsed_page.rb +18 -2
- data/lib/scrapeunblocker/version.rb +1 -1
- metadata +2 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: e8096ca55e9f93c01077d947e60e36ec71c595e7ee92d1405a413a0746c303f0
|
|
4
|
+
data.tar.gz: 28c0667c78459edf026b26effd0209ea0635409eb078f152b8e4963be2312dd3
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: f1b51dee641104d7d77fd600c5cb2299f30d6dcdf3e155686befc2d3024c981fb0c904d2aa8d5d1f360797a9cdbd331356ec2c99f1a3832907704198689306c6
|
|
7
|
+
data.tar.gz: 8274485fa1644d6dfb1512bc33e93fc034f3c01a0248e450869315f3d0669244761ef146db9e11cac0ccfb3fe6361148d5b3ef64a4c706d967d825ef23fba9e6
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,12 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.8.0 (2026-10-06)
|
|
4
|
+
|
|
5
|
+
- `get_parsed` on a page that rendered but held no structured data now returns a normal `ParsedPage` instead of raising `NoDataExtractedError`. The API answers this case with a billed 200: `data` is empty, the new `data_extracted?` is `false`, the new `detail` explains it and the new `html` carries the rendered page, so no second call is needed. On a successful parse `data_extracted?` is `true` and `html` / `detail` are `nil`.
|
|
6
|
+
- `NoDataExtractedError` is deprecated: the API no longer sends the 422 `no_data_extracted`. The class stays so existing `rescue` clauses keep loading.
|
|
7
|
+
|
|
8
|
+
Behaviour change: code that rescued `NoDataExtractedError` should check `page.data_extracted?` instead.
|
|
9
|
+
|
|
3
10
|
## 0.7.1 (2026-09-30)
|
|
4
11
|
|
|
5
12
|
- New `ScrapeUnblocker::NoDataExtractedError` (a subclass of `ValidationError`): raised by `get_parsed` when the page rendered but no structured data could be extracted from it. The API answers 422 with `{"error": "no_data_extracted", "detail": ...}`; the error carries `detail`. The call is not billed and is never retried - use `get_page_source` for the HTML. Before, this came back as a billed 200 with empty `data`.
|
data/README.md
CHANGED
|
@@ -282,7 +282,6 @@ end
|
|
|
282
282
|
| `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
|
|
283
283
|
| `UnsupportedContentError` | 415 | The URL serves something other than HTML |
|
|
284
284
|
| `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
|
|
285
|
-
| `NoDataExtractedError` | 422 | `get_parsed`: the page rendered but held no structured data; carries `detail` (subclass of `ValidationError`, not billed) |
|
|
286
285
|
| `RateLimitError` | 429 | Too many requests |
|
|
287
286
|
| `UpstreamOutageError` | 503 | The target origin is down |
|
|
288
287
|
| `ServerError` | 5xx | Unexpected server error, including a 504 upstream timeout |
|
|
@@ -306,20 +305,19 @@ end
|
|
|
306
305
|
|
|
307
306
|
With `get_parsed` the body is the parsed-data JSON (`{"data": {"page_type": "not_found", ...}}`), so `html` is `nil` there; the raw JSON is on `#body`.
|
|
308
307
|
|
|
309
|
-
### No structured data on the page
|
|
308
|
+
### No structured data on the page
|
|
310
309
|
|
|
311
|
-
When `get_parsed` renders the page but can extract no structured data from it, the API answers
|
|
310
|
+
When `get_parsed` renders the page but can extract no structured data from it, the API still answers 200 and the client returns a normal `ParsedPage`: `data` is empty, `data_extracted?` is `false`, `detail` says why, and `html` carries the rendered page. The call is billed like any delivered page, so you already have the HTML - no second call needed:
|
|
312
311
|
|
|
313
312
|
```ruby
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
html = su.get_page_source("https://example.com/some-page")
|
|
313
|
+
page = su.get_parsed("https://example.com/some-page")
|
|
314
|
+
unless page.data_extracted?
|
|
315
|
+
puts page.detail # the API's explanation
|
|
316
|
+
html = page.html # the rendered page, parse it yourself
|
|
319
317
|
end
|
|
320
318
|
```
|
|
321
319
|
|
|
322
|
-
`
|
|
320
|
+
On a successful parse `data_extracted?` is `true` and `html` is `nil`. Up to 0.7.1 this case raised `NoDataExtractedError` (a 422); the API no longer sends that, and the class stays only so existing `rescue` clauses keep loading.
|
|
323
321
|
|
|
324
322
|
Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
|
|
325
323
|
|
|
@@ -74,10 +74,11 @@ module ScrapeUnblocker
|
|
|
74
74
|
|
|
75
75
|
# Fetch a URL and return structured JSON instead of HTML.
|
|
76
76
|
#
|
|
77
|
-
#
|
|
78
|
-
#
|
|
79
|
-
#
|
|
80
|
-
#
|
|
77
|
+
# A page that rendered but held no structured data comes back as a
|
|
78
|
+
# ParsedPage with #data_extracted? false and the rendered page on #html
|
|
79
|
+
# (billed like any delivered page). Raises
|
|
80
|
+
# ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the target
|
|
81
|
+
# page itself answered 404 or 410.
|
|
81
82
|
def get_parsed(url, proxy_country: nil, time_sleep: nil, refresh_rules: false, rules_hint: nil)
|
|
82
83
|
body = request("/getPageSource",
|
|
83
84
|
url: url, parsed_data: true, proxy_country: proxy_country,
|
|
@@ -128,13 +128,13 @@ module ScrapeUnblocker
|
|
|
128
128
|
# each problem field. Read it from #body.
|
|
129
129
|
class ValidationError < APIError; end
|
|
130
130
|
|
|
131
|
-
#
|
|
131
|
+
# Deprecated: the API no longer sends this 422.
|
|
132
132
|
#
|
|
133
|
-
#
|
|
134
|
-
#
|
|
135
|
-
#
|
|
136
|
-
#
|
|
137
|
-
#
|
|
133
|
+
# A page that rendered but held no structured data now comes back from
|
|
134
|
+
# #get_parsed as a normal ParsedPage with #data_extracted? false and the page
|
|
135
|
+
# on #html (a billed 200). The class stays so existing rescue clauses keep
|
|
136
|
+
# loading; it is raised only for a legacy 422 body of
|
|
137
|
+
# {"error": "no_data_extracted", "detail": ...}.
|
|
138
138
|
class NoDataExtractedError < ValidationError
|
|
139
139
|
attr_reader :detail
|
|
140
140
|
|
|
@@ -222,9 +222,9 @@ module ScrapeUnblocker
|
|
|
222
222
|
end
|
|
223
223
|
private_class_method :target_not_found_error
|
|
224
224
|
|
|
225
|
-
#
|
|
226
|
-
#
|
|
227
|
-
# the general ValidationError applies.
|
|
225
|
+
# Maps a legacy 422 {"error": "no_data_extracted", "detail"} body; the API now
|
|
226
|
+
# answers an empty parse with a 200 (data_extracted: false). Anything else
|
|
227
|
+
# returns nil so the general ValidationError applies.
|
|
228
228
|
def self.no_data_extracted_error(status, body)
|
|
229
229
|
return nil unless status == 422
|
|
230
230
|
|
|
@@ -11,12 +11,25 @@ module ScrapeUnblocker
|
|
|
11
11
|
attr_reader :data
|
|
12
12
|
# @return [Hash] the full JSON payload as returned by the API
|
|
13
13
|
attr_reader :raw
|
|
14
|
+
# @return [String, nil] the rendered page when #data_extracted? is false
|
|
15
|
+
attr_reader :html
|
|
16
|
+
# @return [String, nil] the API's explanation when #data_extracted? is false
|
|
17
|
+
attr_reader :detail
|
|
14
18
|
|
|
15
|
-
def initialize(page_type:, source:, data:, raw:)
|
|
19
|
+
def initialize(page_type:, source:, data:, raw:, data_extracted: true, html: nil, detail: nil)
|
|
16
20
|
@page_type = page_type
|
|
17
21
|
@source = source
|
|
18
22
|
@data = data
|
|
19
23
|
@raw = raw
|
|
24
|
+
@data_extracted = data_extracted
|
|
25
|
+
@html = html
|
|
26
|
+
@detail = detail
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
# @return [Boolean] false when the page rendered but no structured data
|
|
30
|
+
# could be extracted from it (#data is then empty, #html holds the page)
|
|
31
|
+
def data_extracted?
|
|
32
|
+
@data_extracted
|
|
20
33
|
end
|
|
21
34
|
|
|
22
35
|
def self.from_hash(payload)
|
|
@@ -25,7 +38,10 @@ module ScrapeUnblocker
|
|
|
25
38
|
page_type: inner["page_type"],
|
|
26
39
|
source: inner["source"],
|
|
27
40
|
data: inner["data"],
|
|
28
|
-
raw: payload
|
|
41
|
+
raw: payload,
|
|
42
|
+
data_extracted: payload["data_extracted"] != false,
|
|
43
|
+
html: payload["html"],
|
|
44
|
+
detail: payload["detail"]
|
|
29
45
|
)
|
|
30
46
|
end
|
|
31
47
|
end
|
metadata
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: scrapeunblocker
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.8.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- ScrapeUnblocker
|
|
8
8
|
autorequire:
|
|
9
9
|
bindir: bin
|
|
10
10
|
cert_chain: []
|
|
11
|
-
date: 2026-
|
|
11
|
+
date: 2026-10-06 00:00:00.000000000 Z
|
|
12
12
|
dependencies: []
|
|
13
13
|
description: JS-rendered pages that bypass Cloudflare, DataDome, PerimeterX and Akamai,
|
|
14
14
|
plus Google SERP and Skyscanner flights/hotels/car-hire scraping as JSON.
|