scrapeunblocker 0.7.0 → 0.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +5 -0
- data/README.md +18 -0
- data/lib/scrapeunblocker/client.rb +5 -0
- data/lib/scrapeunblocker/errors.rb +39 -0
- data/lib/scrapeunblocker/version.rb +1 -1
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 250fdc0480188972b4735ad8a308cd16df5702a07f98de0d1ce024db61907fe5
|
|
4
|
+
data.tar.gz: 82f6efb47746aabab6045e47776709aa56de498e32265e635eb0e258c0514af1
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: fb3abf429337da70c37605c7c37d13976da20d69c6b53226bd0eed943bf85a032c8f6c401cb0daa8a432e8b8394a239217b8f0c73d8beda15da648d1724f97e7
|
|
7
|
+
data.tar.gz: 95717ff391f8fbf0a5297141edef98a17cea5fcc051dbf0f1a05ce463c8ac7a17d2f89134dfbf48b6f451703ee9ca97042e30053e0ba8b4c6085580aadd83596
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,10 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.7.1 (2026-09-30)
|
|
4
|
+
|
|
5
|
+
- New `ScrapeUnblocker::NoDataExtractedError` (a subclass of `ValidationError`): raised by `get_parsed` when the page rendered but no structured data could be extracted from it. The API answers 422 with `{"error": "no_data_extracted", "detail": ...}`; the error carries `detail`. The call is not billed and is never retried - use `get_page_source` for the HTML. Before, this came back as a billed 200 with empty `data`.
|
|
6
|
+
- A dead URL fetched with `get_parsed` raises `TargetNotFoundError` with `html` nil (the body is the parsed-data JSON, on `body`).
|
|
7
|
+
|
|
3
8
|
## 0.7.0 (2026-09-30)
|
|
4
9
|
|
|
5
10
|
- New `ScrapeUnblocker::TargetNotFoundError` (a subclass of `NotFoundError`): raised by `get_page_source`, `get_parsed` and `get_page_with_cookies` when the target page itself answers 404 or 410. The API now passes the target's own status through instead of a 200, marked with the `X-Origin-Status` header. The error carries `origin_status`, the not-found page on `html` and `destination_url`. It is never retried, and the call is billed like any delivered page.
|
data/README.md
CHANGED
|
@@ -282,6 +282,7 @@ end
|
|
|
282
282
|
| `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
|
|
283
283
|
| `UnsupportedContentError` | 415 | The URL serves something other than HTML |
|
|
284
284
|
| `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
|
|
285
|
+
| `NoDataExtractedError` | 422 | `get_parsed`: the page rendered but held no structured data; carries `detail` (subclass of `ValidationError`, not billed) |
|
|
285
286
|
| `RateLimitError` | 429 | Too many requests |
|
|
286
287
|
| `UpstreamOutageError` | 503 | The target origin is down |
|
|
287
288
|
| `ServerError` | 5xx | Unexpected server error, including a 504 upstream timeout |
|
|
@@ -303,6 +304,23 @@ end
|
|
|
303
304
|
|
|
304
305
|
`TargetNotFoundError` subclasses `NotFoundError`, so `rescue ScrapeUnblocker::NotFoundError` catches it too. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
|
|
305
306
|
|
|
307
|
+
With `get_parsed` the body is the parsed-data JSON (`{"data": {"page_type": "not_found", ...}}`), so `html` is `nil` there; the raw JSON is on `#body`.
|
|
308
|
+
|
|
309
|
+
### No structured data on the page (422)
|
|
310
|
+
|
|
311
|
+
When `get_parsed` renders the page but can extract no structured data from it, the API answers 422 and the client raises `ScrapeUnblocker::NoDataExtractedError`. The call is **not billed**, and retrying gives the same result - fetch the HTML with `get_page_source` instead:
|
|
312
|
+
|
|
313
|
+
```ruby
|
|
314
|
+
begin
|
|
315
|
+
page = su.get_parsed("https://example.com/some-page")
|
|
316
|
+
rescue ScrapeUnblocker::NoDataExtractedError => e
|
|
317
|
+
puts e.detail # the API's explanation
|
|
318
|
+
html = su.get_page_source("https://example.com/some-page")
|
|
319
|
+
end
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
`NoDataExtractedError` subclasses `ValidationError`, so `rescue ScrapeUnblocker::ValidationError` catches it too.
|
|
323
|
+
|
|
306
324
|
Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
|
|
307
325
|
|
|
308
326
|
### Billing errors (402)
|
|
@@ -73,6 +73,11 @@ module ScrapeUnblocker
|
|
|
73
73
|
end
|
|
74
74
|
|
|
75
75
|
# Fetch a URL and return structured JSON instead of HTML.
|
|
76
|
+
#
|
|
77
|
+
# Raises ScrapeUnblocker::NoDataExtractedError (not billed) when the page
|
|
78
|
+
# rendered but held no structured data - use #get_page_source for the HTML -
|
|
79
|
+
# and ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the
|
|
80
|
+
# target page itself answered 404 or 410.
|
|
76
81
|
def get_parsed(url, proxy_country: nil, time_sleep: nil, refresh_rules: false, rules_hint: nil)
|
|
77
82
|
body = request("/getPageSource",
|
|
78
83
|
url: url, parsed_data: true, proxy_country: proxy_country,
|
|
@@ -128,6 +128,22 @@ module ScrapeUnblocker
|
|
|
128
128
|
# each problem field. Read it from #body.
|
|
129
129
|
class ValidationError < APIError; end
|
|
130
130
|
|
|
131
|
+
# The page rendered but no structured data came out of it (HTTP 422).
|
|
132
|
+
#
|
|
133
|
+
# Raised by #get_parsed when the API loaded the page but could not extract any
|
|
134
|
+
# structured fields from it. The API answers 422 with a JSON body of
|
|
135
|
+
# {"error": "no_data_extracted", "detail": ...}; +detail+ holds the API's
|
|
136
|
+
# explanation. The call is not billed and retrying returns the same answer;
|
|
137
|
+
# call #get_page_source for the HTML.
|
|
138
|
+
class NoDataExtractedError < ValidationError
|
|
139
|
+
attr_reader :detail
|
|
140
|
+
|
|
141
|
+
def initialize(message, status_code:, body: nil, detail: nil)
|
|
142
|
+
super(message, status_code: status_code, body: body)
|
|
143
|
+
@detail = detail
|
|
144
|
+
end
|
|
145
|
+
end
|
|
146
|
+
|
|
131
147
|
# The target site blocked every available bypass path (HTTP 403).
|
|
132
148
|
# Blocked calls are not billed.
|
|
133
149
|
class BlockedError < APIError; end
|
|
@@ -206,12 +222,35 @@ module ScrapeUnblocker
|
|
|
206
222
|
end
|
|
207
223
|
private_class_method :target_not_found_error
|
|
208
224
|
|
|
225
|
+
# parsed_data answers 422 with {"error": "no_data_extracted", "detail"} when
|
|
226
|
+
# the page rendered but held no structured data. Anything else returns nil so
|
|
227
|
+
# the general ValidationError applies.
|
|
228
|
+
def self.no_data_extracted_error(status, body)
|
|
229
|
+
return nil unless status == 422
|
|
230
|
+
|
|
231
|
+
data = begin
|
|
232
|
+
JSON.parse(body.to_s)
|
|
233
|
+
rescue JSON::ParserError
|
|
234
|
+
nil
|
|
235
|
+
end
|
|
236
|
+
return nil unless data.is_a?(Hash) && data["error"] == "no_data_extracted"
|
|
237
|
+
|
|
238
|
+
detail = data["detail"].is_a?(String) && !data["detail"].empty? ? data["detail"] : nil
|
|
239
|
+
message = detail || "The page was rendered, but no structured data could be extracted from it. Not billed."
|
|
240
|
+
message = "#{message} Not billed." unless message.downcase.include?("not billed")
|
|
241
|
+
NoDataExtractedError.new(message, status_code: status, body: body, detail: detail)
|
|
242
|
+
end
|
|
243
|
+
private_class_method :no_data_extracted_error
|
|
244
|
+
|
|
209
245
|
# Build a typed error from an HTTP status code, response body and headers
|
|
210
246
|
# (a Hash with lowercase names).
|
|
211
247
|
def self.error_for_status(status, body, headers = {})
|
|
212
248
|
target_error = target_not_found_error(status, body, headers || {})
|
|
213
249
|
return target_error if target_error
|
|
214
250
|
|
|
251
|
+
no_data_error = no_data_extracted_error(status, body)
|
|
252
|
+
return no_data_error if no_data_error
|
|
253
|
+
|
|
215
254
|
snippet = (body || "").strip.gsub(/\s+/, " ")
|
|
216
255
|
snippet = "#{snippet[0, 200]}..." if snippet.length > 200
|
|
217
256
|
base = BASE_MESSAGES.fetch(status, "API returned HTTP #{status}")
|