scrapeunblocker 0.7.0 → 0.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 941309503bf1a937b76c503a811aab8e7e738e4ec775360ad7f9f6cfc40b6259
4
- data.tar.gz: 68753623ec6ad6623632c520e688a5ba2f7697abf6588b8351cdd33e911f473c
3
+ metadata.gz: 250fdc0480188972b4735ad8a308cd16df5702a07f98de0d1ce024db61907fe5
4
+ data.tar.gz: 82f6efb47746aabab6045e47776709aa56de498e32265e635eb0e258c0514af1
5
5
  SHA512:
6
- metadata.gz: 14f6f077622d3669a8ef06944d4cb810b89d693d28eead00b1b5789242945478e0e83529ede25aed9a474ee79351ca5202a3c91c03a174fd0cc811a031a48e11
7
- data.tar.gz: 23314b94ae64e4ce3ade4bbef52278df0639681801981865e3faa8914fd71f6dec7f3047fbad96e4954dc70b8d603ef35487b5a3d5a9a65623be6e7de41071f7
6
+ metadata.gz: fb3abf429337da70c37605c7c37d13976da20d69c6b53226bd0eed943bf85a032c8f6c401cb0daa8a432e8b8394a239217b8f0c73d8beda15da648d1724f97e7
7
+ data.tar.gz: 95717ff391f8fbf0a5297141edef98a17cea5fcc051dbf0f1a05ce463c8ac7a17d2f89134dfbf48b6f451703ee9ca97042e30053e0ba8b4c6085580aadd83596
data/CHANGELOG.md CHANGED
@@ -1,5 +1,10 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.7.1 (2026-09-30)
4
+
5
+ - New `ScrapeUnblocker::NoDataExtractedError` (a subclass of `ValidationError`): raised by `get_parsed` when the page rendered but no structured data could be extracted from it. The API answers 422 with `{"error": "no_data_extracted", "detail": ...}`; the error carries `detail`. The call is not billed and is never retried - use `get_page_source` for the HTML. Before, this came back as a billed 200 with empty `data`.
6
+ - A dead URL fetched with `get_parsed` raises `TargetNotFoundError` with `html` nil (the body is the parsed-data JSON, on `body`).
7
+
3
8
  ## 0.7.0 (2026-09-30)
4
9
 
5
10
  - New `ScrapeUnblocker::TargetNotFoundError` (a subclass of `NotFoundError`): raised by `get_page_source`, `get_parsed` and `get_page_with_cookies` when the target page itself answers 404 or 410. The API now passes the target's own status through instead of a 200, marked with the `X-Origin-Status` header. The error carries `origin_status`, the not-found page on `html` and `destination_url`. It is never retried, and the call is billed like any delivered page.
data/README.md CHANGED
@@ -282,6 +282,7 @@ end
282
282
  | `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
283
283
  | `UnsupportedContentError` | 415 | The URL serves something other than HTML |
284
284
  | `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
285
+ | `NoDataExtractedError` | 422 | `get_parsed`: the page rendered but held no structured data; carries `detail` (subclass of `ValidationError`, not billed) |
285
286
  | `RateLimitError` | 429 | Too many requests |
286
287
  | `UpstreamOutageError` | 503 | The target origin is down |
287
288
  | `ServerError` | 5xx | Unexpected server error, including a 504 upstream timeout |
@@ -303,6 +304,23 @@ end
303
304
 
304
305
  `TargetNotFoundError` subclasses `NotFoundError`, so `rescue ScrapeUnblocker::NotFoundError` catches it too. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
305
306
 
307
+ With `get_parsed` the body is the parsed-data JSON (`{"data": {"page_type": "not_found", ...}}`), so `html` is `nil` there; the raw JSON is on `#body`.
308
+
309
+ ### No structured data on the page (422)
310
+
311
+ When `get_parsed` renders the page but can extract no structured data from it, the API answers 422 and the client raises `ScrapeUnblocker::NoDataExtractedError`. The call is **not billed**, and retrying gives the same result - fetch the HTML with `get_page_source` instead:
312
+
313
+ ```ruby
314
+ begin
315
+ page = su.get_parsed("https://example.com/some-page")
316
+ rescue ScrapeUnblocker::NoDataExtractedError => e
317
+ puts e.detail # the API's explanation
318
+ html = su.get_page_source("https://example.com/some-page")
319
+ end
320
+ ```
321
+
322
+ `NoDataExtractedError` subclasses `ValidationError`, so `rescue ScrapeUnblocker::ValidationError` catches it too.
323
+
306
324
  Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
307
325
 
308
326
  ### Billing errors (402)
@@ -73,6 +73,11 @@ module ScrapeUnblocker
73
73
  end
74
74
 
75
75
  # Fetch a URL and return structured JSON instead of HTML.
76
+ #
77
+ # Raises ScrapeUnblocker::NoDataExtractedError (not billed) when the page
78
+ # rendered but held no structured data - use #get_page_source for the HTML -
79
+ # and ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the
80
+ # target page itself answered 404 or 410.
76
81
  def get_parsed(url, proxy_country: nil, time_sleep: nil, refresh_rules: false, rules_hint: nil)
77
82
  body = request("/getPageSource",
78
83
  url: url, parsed_data: true, proxy_country: proxy_country,
@@ -128,6 +128,22 @@ module ScrapeUnblocker
128
128
  # each problem field. Read it from #body.
129
129
  class ValidationError < APIError; end
130
130
 
131
+ # The page rendered but no structured data came out of it (HTTP 422).
132
+ #
133
+ # Raised by #get_parsed when the API loaded the page but could not extract any
134
+ # structured fields from it. The API answers 422 with a JSON body of
135
+ # {"error": "no_data_extracted", "detail": ...}; +detail+ holds the API's
136
+ # explanation. The call is not billed and retrying returns the same answer;
137
+ # call #get_page_source for the HTML.
138
+ class NoDataExtractedError < ValidationError
139
+ attr_reader :detail
140
+
141
+ def initialize(message, status_code:, body: nil, detail: nil)
142
+ super(message, status_code: status_code, body: body)
143
+ @detail = detail
144
+ end
145
+ end
146
+
131
147
  # The target site blocked every available bypass path (HTTP 403).
132
148
  # Blocked calls are not billed.
133
149
  class BlockedError < APIError; end
@@ -206,12 +222,35 @@ module ScrapeUnblocker
206
222
  end
207
223
  private_class_method :target_not_found_error
208
224
 
225
+ # parsed_data answers 422 with {"error": "no_data_extracted", "detail"} when
226
+ # the page rendered but held no structured data. Anything else returns nil so
227
+ # the general ValidationError applies.
228
+ def self.no_data_extracted_error(status, body)
229
+ return nil unless status == 422
230
+
231
+ data = begin
232
+ JSON.parse(body.to_s)
233
+ rescue JSON::ParserError
234
+ nil
235
+ end
236
+ return nil unless data.is_a?(Hash) && data["error"] == "no_data_extracted"
237
+
238
+ detail = data["detail"].is_a?(String) && !data["detail"].empty? ? data["detail"] : nil
239
+ message = detail || "The page was rendered, but no structured data could be extracted from it. Not billed."
240
+ message = "#{message} Not billed." unless message.downcase.include?("not billed")
241
+ NoDataExtractedError.new(message, status_code: status, body: body, detail: detail)
242
+ end
243
+ private_class_method :no_data_extracted_error
244
+
209
245
  # Build a typed error from an HTTP status code, response body and headers
210
246
  # (a Hash with lowercase names).
211
247
  def self.error_for_status(status, body, headers = {})
212
248
  target_error = target_not_found_error(status, body, headers || {})
213
249
  return target_error if target_error
214
250
 
251
+ no_data_error = no_data_extracted_error(status, body)
252
+ return no_data_error if no_data_error
253
+
215
254
  snippet = (body || "").strip.gsub(/\s+/, " ")
216
255
  snippet = "#{snippet[0, 200]}..." if snippet.length > 200
217
256
  base = BASE_MESSAGES.fetch(status, "API returned HTTP #{status}")
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module ScrapeUnblocker
4
- VERSION = "0.7.0"
4
+ VERSION = "0.7.1"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: scrapeunblocker
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.7.0
4
+ version: 0.7.1
5
5
  platform: ruby
6
6
  authors:
7
7
  - ScrapeUnblocker