scrapeunblocker 0.5.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 2668bf966d35f46fd8c65d4d78cdc0d93d989a54b9d3f257f57e7924cc4ad6f6
4
- data.tar.gz: 7ec5e5166b99712828dd771597ba0db89a09f398a1c5c0988af3256f077139f1
3
+ metadata.gz: 941309503bf1a937b76c503a811aab8e7e738e4ec775360ad7f9f6cfc40b6259
4
+ data.tar.gz: 68753623ec6ad6623632c520e688a5ba2f7697abf6588b8351cdd33e911f473c
5
5
  SHA512:
6
- metadata.gz: 98330917b1ddce4f2d24f3938a97ce719ef84e31a712a3e774ac5d5821b1df164d8defd3984d4732d736f374f357701b59bae7b3a3154539b06543b92b8ecc5a
7
- data.tar.gz: 0fb21ac2f6c04e979bb21415a9d393fd2853bdb1eb793e6ddf1d52cd9a38ab976d9ab403352f15bbe3b24ae64972c60df972942e231d85ff5e2ac805bf76bc34
6
+ metadata.gz: 14f6f077622d3669a8ef06944d4cb810b89d693d28eead00b1b5789242945478e0e83529ede25aed9a474ee79351ca5202a3c91c03a174fd0cc811a031a48e11
7
+ data.tar.gz: 23314b94ae64e4ce3ade4bbef52278df0639681801981865e3faa8914fd71f6dec7f3047fbad96e4954dc70b8d603ef35487b5a3d5a9a65623be6e7de41071f7
data/CHANGELOG.md CHANGED
@@ -1,5 +1,18 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.7.0 (2026-09-30)
4
+
5
+ - New `ScrapeUnblocker::TargetNotFoundError` (a subclass of `NotFoundError`): raised by `get_page_source`, `get_parsed` and `get_page_with_cookies` when the target page itself answers 404 or 410. The API now passes the target's own status through instead of a 200, marked with the `X-Origin-Status` header. The error carries `origin_status`, the not-found page on `html` and `destination_url`. It is never retried, and the call is billed like any delivered page.
6
+ - A custom `transport:` may return `headers:` alongside `status:` and `body:`; `ScrapeUnblocker.error_for_status` takes them as an optional third argument.
7
+
8
+ Behaviour change: a dead target URL used to return its not-found page as a normal String; it now raises `TargetNotFoundError`. `rescue ScrapeUnblocker::NotFoundError` still catches it. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
9
+
10
+ ## 0.6.0 (2026-09-24)
11
+
12
+ - `google_images` takes `pages:` (1-5): fetch up to five Google Images result pages of ~100 results each in one call. Each page fetched is billed as one request; the response's `pagesFetched` says how many. The Google market now follows `proxy_country` automatically, so `gl` is only an optional override, and `max_results` is an optional cap up to 500.
13
+
14
+ No breaking changes.
15
+
3
16
  ## 0.5.0 (2026-09-17)
4
17
 
5
18
  - Added `google_images(q, ...)` for the new Google Images plugin (`POST /images/google-search`). Given a keyword `q` it returns Google Images results as a Hash - each with the full-size `imageUrl` and its `sourceDomain`, plus the source page URL, title, source name, thumbnail URL, pixel dimensions and file size. Optional `gl` (ISO-2 lowercase market), `max_results` (1-100) and `proxy_country` (ISO-2) refine the search.
data/README.md CHANGED
@@ -131,7 +131,7 @@ local["results"].each { |biz| puts "#{biz['name']} #{biz['rating']} #{biz['addre
131
131
  ## Google Images
132
132
 
133
133
  ```ruby
134
- images = su.google_images("golden retriever puppy", proxy_country: "US", gl: "us")
134
+ images = su.google_images("golden retriever puppy", proxy_country: "US", pages: 3)
135
135
  images["results"].each { |img| puts "#{img['imageUrl']} #{img['sourceDomain']} #{img['title']}" }
136
136
  ```
137
137
 
@@ -277,7 +277,8 @@ end
277
277
  | `CreditLimitExceededError` | 402 | Unpaid balance is past the account's credit limit |
278
278
  | `PaymentFailedError` | 402 | A card payment was declined three times |
279
279
  | `BlockedError` | 403 | Blocked by bot protection on every path |
280
- | `NotFoundError` | 404 | Page loaded but held no image (`get_image` only) |
280
+ | `NotFoundError` | 404 | What you asked for does not exist - no image on the page (`get_image`), or a plugin lookup found nothing |
281
+ | `TargetNotFoundError` | 404 / 410 | The target page itself does not exist; carries `origin_status`, `html`, `destination_url` (subclass of `NotFoundError`, billed) |
281
282
  | `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
282
283
  | `UnsupportedContentError` | 415 | The URL serves something other than HTML |
283
284
  | `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
@@ -287,6 +288,21 @@ end
287
288
  | `TimeoutError` | - | This client gave up locally before the API answered |
288
289
  | `ConnectionError` | - | Could not reach the API |
289
290
 
291
+ ### The target page does not exist (404 / 410)
292
+
293
+ When the site you scrape answers 404 or 410 itself, the API passes that status through with an `X-Origin-Status` header, and the client raises `ScrapeUnblocker::TargetNotFoundError`. It is the target's final answer, so it is never retried, and it is billed like any delivered page. The not-found page is on `#html`:
294
+
295
+ ```ruby
296
+ begin
297
+ html = su.get_page_source("https://example.com/removed-listing")
298
+ rescue ScrapeUnblocker::TargetNotFoundError => e
299
+ puts e.origin_status # 404 or 410
300
+ puts e.html # the target's own not-found page (can be empty)
301
+ end
302
+ ```
303
+
304
+ `TargetNotFoundError` subclasses `NotFoundError`, so `rescue ScrapeUnblocker::NotFoundError` catches it too. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
305
+
290
306
  Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
291
307
 
292
308
  ### Billing errors (402)
@@ -57,6 +57,9 @@ module ScrapeUnblocker
57
57
  # +list_elements+, when true, makes the API return a JSON summary of the
58
58
  # matched elements ({"url", "count", "elements"}) instead of HTML. This method
59
59
  # then returns that parsed Hash rather than an HTML String.
60
+ #
61
+ # When the target page itself answers 404 or 410 this raises
62
+ # ScrapeUnblocker::TargetNotFoundError (billed; the not-found page is on #html).
60
63
  def get_page_source(url, proxy_country: nil, time_sleep: nil, method: nil, value: nil, method_timeout: nil,
61
64
  steps: nil, list_elements: nil)
62
65
  body = request("/getPageSource",
@@ -111,11 +114,13 @@ module ScrapeUnblocker
111
114
  # Returns image results, each with the full-size +imageUrl+ and its
112
115
  # +sourceDomain+, plus the source page URL, title, source name, thumbnail
113
116
  # URL, pixel dimensions and file size. +q+ is the search keyword. Set
114
- # +proxy_country+ (and optionally +gl+) to target a market, and
115
- # +max_results+ (1-100) to cap the count.
116
- def google_images(q, proxy_country: nil, gl: nil, max_results: nil)
117
+ # +proxy_country+ to target a market (the Google market follows it; +gl+
118
+ # is an optional override), +pages+ (1-5, ~100 results each; each page
119
+ # fetched is billed as one request, see +pagesFetched+) and +max_results+
120
+ # (1-500) to cap the count.
121
+ def google_images(q, proxy_country: nil, pages: nil, gl: nil, max_results: nil)
117
122
  post_json("/images/google-search",
118
- q: q, gl: gl, max_results: max_results, proxy_country: proxy_country)
123
+ q: q, pages: pages, gl: gl, max_results: max_results, proxy_country: proxy_country)
119
124
  end
120
125
 
121
126
  # Fetch an advertiser's Meta (Facebook) Ad Library ads and return them as a Hash.
@@ -282,7 +287,8 @@ module ScrapeUnblocker
282
287
 
283
288
  return { status: status, body: body } if status >= 200 && status < 300
284
289
 
285
- raise ScrapeUnblocker.error_for_status(status, body)
290
+ response_headers = (result[:headers] || {}).to_h { |k, v| [k.to_s.downcase, Array(v).join(", ")] }
291
+ raise ScrapeUnblocker.error_for_status(status, body, response_headers)
286
292
  end
287
293
  end
288
294
 
@@ -313,7 +319,7 @@ module ScrapeUnblocker
313
319
  raise ConnectionError, "Could not reach the API: #{e.message}"
314
320
  end
315
321
 
316
- { status: response.code.to_i, body: response.body }
322
+ { status: response.code.to_i, body: response.body, headers: response.each_header.to_h }
317
323
  end
318
324
  end
319
325
  end
@@ -1,5 +1,7 @@
1
1
  # frozen_string_literal: true
2
2
 
3
+ require "json"
4
+
3
5
  module ScrapeUnblocker
4
6
  # Base class for every error raised by this library.
5
7
  class Error < StandardError; end
@@ -78,10 +80,38 @@ module ScrapeUnblocker
78
80
  # use instead.
79
81
  class InvalidRequestError < APIError; end
80
82
 
81
- # The page loaded but the requested element was absent (HTTP 404).
82
- # Only #get_image raises this: the page rendered and held no <img> tag.
83
+ # Something the call asked for does not exist (HTTP 404).
84
+ #
85
+ # #get_image raises it when the page rendered but held no <img> tag, and
86
+ # plugin methods raise it when the item they look up does not exist. When the
87
+ # target page itself answered 404 or 410, the more specific
88
+ # TargetNotFoundError subclass is raised instead.
83
89
  class NotFoundError < APIError; end
84
90
 
91
+ # The target page itself does not exist (HTTP 404 or 410).
92
+ #
93
+ # Raised by #get_page_source, #get_parsed and #get_page_with_cookies when the
94
+ # site you asked for answered 404 or 410 on its own. The API passes that
95
+ # status through and marks it with the X-Origin-Status header, which is how
96
+ # this is told apart from an API-side 404. It is the target's final answer,
97
+ # not a block, so it is never retried - and the call is billed, because the
98
+ # page was fetched and delivered.
99
+ #
100
+ # +origin_status+ is the status the target answered with (404 or 410),
101
+ # +html+ the target's own not-found page as served (can be empty; nil when
102
+ # the body is a parsed-data JSON payload), and +destination_url+ the URL the
103
+ # target answered for, when the API sent X-Destination-URL.
104
+ class TargetNotFoundError < NotFoundError
105
+ attr_reader :origin_status, :html, :destination_url
106
+
107
+ def initialize(message, status_code:, origin_status:, body: nil, html: nil, destination_url: nil)
108
+ super(message, status_code: status_code, body: body)
109
+ @origin_status = origin_status
110
+ @html = html
111
+ @destination_url = destination_url
112
+ end
113
+ end
114
+
85
115
  # The browser run did not finish in time on our side (HTTP 408).
86
116
  #
87
117
  # Distinct from TimeoutError, which is this client giving up locally. Here
@@ -153,8 +183,35 @@ module ScrapeUnblocker
153
183
  end
154
184
  private_class_method :billing_error_class
155
185
 
156
- # Build a typed error from an HTTP status code and response body.
157
- def self.error_for_status(status, body)
186
+ # The API passes a target's "page does not exist" answer through with its
187
+ # status and an X-Origin-Status header. A 404 without that header is the
188
+ # API's own (a plugin lookup, a missing element) and returns nil so the
189
+ # general NotFoundError applies.
190
+ def self.target_not_found_error(status, body, headers)
191
+ origin = headers["x-origin-status"]
192
+ return nil unless [404, 410].include?(status) && origin && !origin.to_s.empty?
193
+
194
+ origin_status = Integer(origin.to_s, exception: false) || status
195
+ html = body
196
+ begin
197
+ data = JSON.parse(body.to_s)
198
+ html = data["html"].is_a?(String) ? data["html"] : nil if data.is_a?(Hash)
199
+ rescue JSON::ParserError
200
+ # Not JSON: the body is the target's own page.
201
+ end
202
+ message = "Target page does not exist (HTTP #{origin_status}). This is the " \
203
+ "target's own answer, not a block; the call is billed."
204
+ TargetNotFoundError.new(message, status_code: status, body: body, origin_status: origin_status,
205
+ html: html, destination_url: headers["x-destination-url"])
206
+ end
207
+ private_class_method :target_not_found_error
208
+
209
+ # Build a typed error from an HTTP status code, response body and headers
210
+ # (a Hash with lowercase names).
211
+ def self.error_for_status(status, body, headers = {})
212
+ target_error = target_not_found_error(status, body, headers || {})
213
+ return target_error if target_error
214
+
158
215
  snippet = (body || "").strip.gsub(/\s+/, " ")
159
216
  snippet = "#{snippet[0, 200]}..." if snippet.length > 200
160
217
  base = BASE_MESSAGES.fetch(status, "API returned HTTP #{status}")
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module ScrapeUnblocker
4
- VERSION = "0.5.0"
4
+ VERSION = "0.7.0"
5
5
  end
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: scrapeunblocker
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.5.0
4
+ version: 0.7.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - ScrapeUnblocker
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-17 00:00:00.000000000 Z
11
+ date: 2026-09-30 00:00:00.000000000 Z
12
12
  dependencies: []
13
13
  description: JS-rendered pages that bypass Cloudflare, DataDome, PerimeterX and Akamai,
14
14
  plus Google SERP and Skyscanner flights/hotels/car-hire scraping as JSON.