scrapeunblocker 0.6.0 → 0.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 3076debfbde8d8f4aac22c77b69c7d9117d2050743f7c5794e856b4c8f0618f4
4
- data.tar.gz: 19dd437fccd00ae77927c1863695bfd0599f20d6a424e61ead7aff0abffbfd04
3
+ metadata.gz: 250fdc0480188972b4735ad8a308cd16df5702a07f98de0d1ce024db61907fe5
4
+ data.tar.gz: 82f6efb47746aabab6045e47776709aa56de498e32265e635eb0e258c0514af1
5
5
  SHA512:
6
- metadata.gz: e0e42653b97107a96cd99ef480633c455b5833d38b57b1e18942232ff46684c48a7d616a3953c9ba887bf5eef848fa17cdafcba912b3014618d9dc45775570b2
7
- data.tar.gz: c7809f130d23a492dd1faec14b243a77b8d1618f796998fc2aad9c05fa6d3c681f175fc54444e59eed818c5fb0906596daa697260f3b4c29b4bde25b689f5257
6
+ metadata.gz: fb3abf429337da70c37605c7c37d13976da20d69c6b53226bd0eed943bf85a032c8f6c401cb0daa8a432e8b8394a239217b8f0c73d8beda15da648d1724f97e7
7
+ data.tar.gz: 95717ff391f8fbf0a5297141edef98a17cea5fcc051dbf0f1a05ce463c8ac7a17d2f89134dfbf48b6f451703ee9ca97042e30053e0ba8b4c6085580aadd83596
data/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.7.1 (2026-09-30)
4
+
5
+ - New `ScrapeUnblocker::NoDataExtractedError` (a subclass of `ValidationError`): raised by `get_parsed` when the page rendered but no structured data could be extracted from it. The API answers 422 with `{"error": "no_data_extracted", "detail": ...}`; the error carries `detail`. The call is not billed and is never retried - use `get_page_source` for the HTML. Before, this came back as a billed 200 with empty `data`.
6
+ - A dead URL fetched with `get_parsed` raises `TargetNotFoundError` with `html` nil (the body is the parsed-data JSON, on `body`).
7
+
8
+ ## 0.7.0 (2026-09-30)
9
+
10
+ - New `ScrapeUnblocker::TargetNotFoundError` (a subclass of `NotFoundError`): raised by `get_page_source`, `get_parsed` and `get_page_with_cookies` when the target page itself answers 404 or 410. The API now passes the target's own status through instead of a 200, marked with the `X-Origin-Status` header. The error carries `origin_status`, the not-found page on `html` and `destination_url`. It is never retried, and the call is billed like any delivered page.
11
+ - A custom `transport:` may return `headers:` alongside `status:` and `body:`; `ScrapeUnblocker.error_for_status` takes them as an optional third argument.
12
+
13
+ Behaviour change: a dead target URL used to return its not-found page as a normal String; it now raises `TargetNotFoundError`. `rescue ScrapeUnblocker::NotFoundError` still catches it. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
14
+
3
15
  ## 0.6.0 (2026-09-24)
4
16
 
5
17
  - `google_images` takes `pages:` (1-5): fetch up to five Google Images result pages of ~100 results each in one call. Each page fetched is billed as one request; the response's `pagesFetched` says how many. The Google market now follows `proxy_country` automatically, so `gl` is only an optional override, and `max_results` is an optional cap up to 500.
data/README.md CHANGED
@@ -277,16 +277,50 @@ end
277
277
  | `CreditLimitExceededError` | 402 | Unpaid balance is past the account's credit limit |
278
278
  | `PaymentFailedError` | 402 | A card payment was declined three times |
279
279
  | `BlockedError` | 403 | Blocked by bot protection on every path |
280
- | `NotFoundError` | 404 | Page loaded but held no image (`get_image` only) |
280
+ | `NotFoundError` | 404 | What you asked for does not exist - no image on the page (`get_image`), or a plugin lookup found nothing |
281
+ | `TargetNotFoundError` | 404 / 410 | The target page itself does not exist; carries `origin_status`, `html`, `destination_url` (subclass of `NotFoundError`, billed) |
281
282
  | `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
282
283
  | `UnsupportedContentError` | 415 | The URL serves something other than HTML |
283
284
  | `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
285
+ | `NoDataExtractedError` | 422 | `get_parsed`: the page rendered but held no structured data; carries `detail` (subclass of `ValidationError`, not billed) |
284
286
  | `RateLimitError` | 429 | Too many requests |
285
287
  | `UpstreamOutageError` | 503 | The target origin is down |
286
288
  | `ServerError` | 5xx | Unexpected server error, including a 504 upstream timeout |
287
289
  | `TimeoutError` | - | This client gave up locally before the API answered |
288
290
  | `ConnectionError` | - | Could not reach the API |
289
291
 
292
+ ### The target page does not exist (404 / 410)
293
+
294
+ When the site you scrape answers 404 or 410 itself, the API passes that status through with an `X-Origin-Status` header, and the client raises `ScrapeUnblocker::TargetNotFoundError`. It is the target's final answer, so it is never retried, and it is billed like any delivered page. The not-found page is on `#html`:
295
+
296
+ ```ruby
297
+ begin
298
+ html = su.get_page_source("https://example.com/removed-listing")
299
+ rescue ScrapeUnblocker::TargetNotFoundError => e
300
+ puts e.origin_status # 404 or 410
301
+ puts e.html # the target's own not-found page (can be empty)
302
+ end
303
+ ```
304
+
305
+ `TargetNotFoundError` subclasses `NotFoundError`, so `rescue ScrapeUnblocker::NotFoundError` catches it too. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
306
+
307
+ With `get_parsed` the body is the parsed-data JSON (`{"data": {"page_type": "not_found", ...}}`), so `html` is `nil` there; the raw JSON is on `#body`.
308
+
309
+ ### No structured data on the page (422)
310
+
311
+ When `get_parsed` renders the page but can extract no structured data from it, the API answers 422 and the client raises `ScrapeUnblocker::NoDataExtractedError`. The call is **not billed**, and retrying gives the same result - fetch the HTML with `get_page_source` instead:
312
+
313
+ ```ruby
314
+ begin
315
+ page = su.get_parsed("https://example.com/some-page")
316
+ rescue ScrapeUnblocker::NoDataExtractedError => e
317
+ puts e.detail # the API's explanation
318
+ html = su.get_page_source("https://example.com/some-page")
319
+ end
320
+ ```
321
+
322
+ `NoDataExtractedError` subclasses `ValidationError`, so `rescue ScrapeUnblocker::ValidationError` catches it too.
323
+
290
324
  Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
291
325
 
292
326
  ### Billing errors (402)
@@ -57,6 +57,9 @@ module ScrapeUnblocker
57
57
  # +list_elements+, when true, makes the API return a JSON summary of the
58
58
  # matched elements ({"url", "count", "elements"}) instead of HTML. This method
59
59
  # then returns that parsed Hash rather than an HTML String.
60
+ #
61
+ # When the target page itself answers 404 or 410 this raises
62
+ # ScrapeUnblocker::TargetNotFoundError (billed; the not-found page is on #html).
60
63
  def get_page_source(url, proxy_country: nil, time_sleep: nil, method: nil, value: nil, method_timeout: nil,
61
64
  steps: nil, list_elements: nil)
62
65
  body = request("/getPageSource",
@@ -70,6 +73,11 @@ module ScrapeUnblocker
70
73
  end
71
74
 
72
75
  # Fetch a URL and return structured JSON instead of HTML.
76
+ #
77
+ # Raises ScrapeUnblocker::NoDataExtractedError (not billed) when the page
78
+ # rendered but held no structured data - use #get_page_source for the HTML -
79
+ # and ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the
80
+ # target page itself answered 404 or 410.
73
81
  def get_parsed(url, proxy_country: nil, time_sleep: nil, refresh_rules: false, rules_hint: nil)
74
82
  body = request("/getPageSource",
75
83
  url: url, parsed_data: true, proxy_country: proxy_country,
@@ -284,7 +292,8 @@ module ScrapeUnblocker
284
292
 
285
293
  return { status: status, body: body } if status >= 200 && status < 300
286
294
 
287
- raise ScrapeUnblocker.error_for_status(status, body)
295
+ response_headers = (result[:headers] || {}).to_h { |k, v| [k.to_s.downcase, Array(v).join(", ")] }
296
+ raise ScrapeUnblocker.error_for_status(status, body, response_headers)
288
297
  end
289
298
  end
290
299
 
@@ -315,7 +324,7 @@ module ScrapeUnblocker
315
324
  raise ConnectionError, "Could not reach the API: #{e.message}"
316
325
  end
317
326
 
318
- { status: response.code.to_i, body: response.body }
327
+ { status: response.code.to_i, body: response.body, headers: response.each_header.to_h }
319
328
  end
320
329
  end
321
330
  end
@@ -1,5 +1,7 @@
1
1
  # frozen_string_literal: true
2
2
 
3
+ require "json"
4
+
3
5
  module ScrapeUnblocker
4
6
  # Base class for every error raised by this library.
5
7
  class Error < StandardError; end
@@ -78,10 +80,38 @@ module ScrapeUnblocker
78
80
  # use instead.
79
81
  class InvalidRequestError < APIError; end
80
82
 
81
- # The page loaded but the requested element was absent (HTTP 404).
82
- # Only #get_image raises this: the page rendered and held no <img> tag.
83
+ # Something the call asked for does not exist (HTTP 404).
84
+ #
85
+ # #get_image raises it when the page rendered but held no <img> tag, and
86
+ # plugin methods raise it when the item they look up does not exist. When the
87
+ # target page itself answered 404 or 410, the more specific
88
+ # TargetNotFoundError subclass is raised instead.
83
89
  class NotFoundError < APIError; end
84
90
 
91
+ # The target page itself does not exist (HTTP 404 or 410).
92
+ #
93
+ # Raised by #get_page_source, #get_parsed and #get_page_with_cookies when the
94
+ # site you asked for answered 404 or 410 on its own. The API passes that
95
+ # status through and marks it with the X-Origin-Status header, which is how
96
+ # this is told apart from an API-side 404. It is the target's final answer,
97
+ # not a block, so it is never retried - and the call is billed, because the
98
+ # page was fetched and delivered.
99
+ #
100
+ # +origin_status+ is the status the target answered with (404 or 410),
101
+ # +html+ the target's own not-found page as served (can be empty; nil when
102
+ # the body is a parsed-data JSON payload), and +destination_url+ the URL the
103
+ # target answered for, when the API sent X-Destination-URL.
104
+ class TargetNotFoundError < NotFoundError
105
+ attr_reader :origin_status, :html, :destination_url
106
+
107
+ def initialize(message, status_code:, origin_status:, body: nil, html: nil, destination_url: nil)
108
+ super(message, status_code: status_code, body: body)
109
+ @origin_status = origin_status
110
+ @html = html
111
+ @destination_url = destination_url
112
+ end
113
+ end
114
+
85
115
  # The browser run did not finish in time on our side (HTTP 408).
86
116
  #
87
117
  # Distinct from TimeoutError, which is this client giving up locally. Here
@@ -98,6 +128,22 @@ module ScrapeUnblocker
98
128
  # each problem field. Read it from #body.
99
129
  class ValidationError < APIError; end
100
130
 
131
+ # The page rendered but no structured data came out of it (HTTP 422).
132
+ #
133
+ # Raised by #get_parsed when the API loaded the page but could not extract any
134
+ # structured fields from it. The API answers 422 with a JSON body of
135
+ # {"error": "no_data_extracted", "detail": ...}; +detail+ holds the API's
136
+ # explanation. The call is not billed and retrying returns the same answer;
137
+ # call #get_page_source for the HTML.
138
+ class NoDataExtractedError < ValidationError
139
+ attr_reader :detail
140
+
141
+ def initialize(message, status_code:, body: nil, detail: nil)
142
+ super(message, status_code: status_code, body: body)
143
+ @detail = detail
144
+ end
145
+ end
146
+
101
147
  # The target site blocked every available bypass path (HTTP 403).
102
148
  # Blocked calls are not billed.
103
149
  class BlockedError < APIError; end
@@ -153,8 +199,58 @@ module ScrapeUnblocker
153
199
  end
154
200
  private_class_method :billing_error_class
155
201
 
156
- # Build a typed error from an HTTP status code and response body.
157
- def self.error_for_status(status, body)
202
+ # The API passes a target's "page does not exist" answer through with its
203
+ # status and an X-Origin-Status header. A 404 without that header is the
204
+ # API's own (a plugin lookup, a missing element) and returns nil so the
205
+ # general NotFoundError applies.
206
+ def self.target_not_found_error(status, body, headers)
207
+ origin = headers["x-origin-status"]
208
+ return nil unless [404, 410].include?(status) && origin && !origin.to_s.empty?
209
+
210
+ origin_status = Integer(origin.to_s, exception: false) || status
211
+ html = body
212
+ begin
213
+ data = JSON.parse(body.to_s)
214
+ html = data["html"].is_a?(String) ? data["html"] : nil if data.is_a?(Hash)
215
+ rescue JSON::ParserError
216
+ # Not JSON: the body is the target's own page.
217
+ end
218
+ message = "Target page does not exist (HTTP #{origin_status}). This is the " \
219
+ "target's own answer, not a block; the call is billed."
220
+ TargetNotFoundError.new(message, status_code: status, body: body, origin_status: origin_status,
221
+ html: html, destination_url: headers["x-destination-url"])
222
+ end
223
+ private_class_method :target_not_found_error
224
+
225
+ # parsed_data answers 422 with {"error": "no_data_extracted", "detail"} when
226
+ # the page rendered but held no structured data. Anything else returns nil so
227
+ # the general ValidationError applies.
228
+ def self.no_data_extracted_error(status, body)
229
+ return nil unless status == 422
230
+
231
+ data = begin
232
+ JSON.parse(body.to_s)
233
+ rescue JSON::ParserError
234
+ nil
235
+ end
236
+ return nil unless data.is_a?(Hash) && data["error"] == "no_data_extracted"
237
+
238
+ detail = data["detail"].is_a?(String) && !data["detail"].empty? ? data["detail"] : nil
239
+ message = detail || "The page was rendered, but no structured data could be extracted from it. Not billed."
240
+ message = "#{message} Not billed." unless message.downcase.include?("not billed")
241
+ NoDataExtractedError.new(message, status_code: status, body: body, detail: detail)
242
+ end
243
+ private_class_method :no_data_extracted_error
244
+
245
+ # Build a typed error from an HTTP status code, response body and headers
246
+ # (a Hash with lowercase names).
247
+ def self.error_for_status(status, body, headers = {})
248
+ target_error = target_not_found_error(status, body, headers || {})
249
+ return target_error if target_error
250
+
251
+ no_data_error = no_data_extracted_error(status, body)
252
+ return no_data_error if no_data_error
253
+
158
254
  snippet = (body || "").strip.gsub(/\s+/, " ")
159
255
  snippet = "#{snippet[0, 200]}..." if snippet.length > 200
160
256
  base = BASE_MESSAGES.fetch(status, "API returned HTTP #{status}")
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module ScrapeUnblocker
4
- VERSION = "0.6.0"
4
+ VERSION = "0.7.1"
5
5
  end
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: scrapeunblocker
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.6.0
4
+ version: 0.7.1
5
5
  platform: ruby
6
6
  authors:
7
7
  - ScrapeUnblocker
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-24 00:00:00.000000000 Z
11
+ date: 2026-09-30 00:00:00.000000000 Z
12
12
  dependencies: []
13
13
  description: JS-rendered pages that bypass Cloudflare, DataDome, PerimeterX and Akamai,
14
14
  plus Google SERP and Skyscanner flights/hotels/car-hire scraping as JSON.