scrapeunblocker 0.5.0 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +13 -0
- data/README.md +18 -2
- data/lib/scrapeunblocker/client.rb +12 -6
- data/lib/scrapeunblocker/errors.rb +61 -4
- data/lib/scrapeunblocker/version.rb +1 -1
- metadata +2 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 941309503bf1a937b76c503a811aab8e7e738e4ec775360ad7f9f6cfc40b6259
|
|
4
|
+
data.tar.gz: 68753623ec6ad6623632c520e688a5ba2f7697abf6588b8351cdd33e911f473c
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 14f6f077622d3669a8ef06944d4cb810b89d693d28eead00b1b5789242945478e0e83529ede25aed9a474ee79351ca5202a3c91c03a174fd0cc811a031a48e11
|
|
7
|
+
data.tar.gz: 23314b94ae64e4ce3ade4bbef52278df0639681801981865e3faa8914fd71f6dec7f3047fbad96e4954dc70b8d603ef35487b5a3d5a9a65623be6e7de41071f7
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.7.0 (2026-09-30)
|
|
4
|
+
|
|
5
|
+
- New `ScrapeUnblocker::TargetNotFoundError` (a subclass of `NotFoundError`): raised by `get_page_source`, `get_parsed` and `get_page_with_cookies` when the target page itself answers 404 or 410. The API now passes the target's own status through instead of a 200, marked with the `X-Origin-Status` header. The error carries `origin_status`, the not-found page on `html` and `destination_url`. It is never retried, and the call is billed like any delivered page.
|
|
6
|
+
- A custom `transport:` may return `headers:` alongside `status:` and `body:`; `ScrapeUnblocker.error_for_status` takes them as an optional third argument.
|
|
7
|
+
|
|
8
|
+
Behaviour change: a dead target URL used to return its not-found page as a normal String; it now raises `TargetNotFoundError`. `rescue ScrapeUnblocker::NotFoundError` still catches it. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
|
|
9
|
+
|
|
10
|
+
## 0.6.0 (2026-09-24)
|
|
11
|
+
|
|
12
|
+
- `google_images` takes `pages:` (1-5): fetch up to five Google Images result pages of ~100 results each in one call. Each page fetched is billed as one request; the response's `pagesFetched` says how many. The Google market now follows `proxy_country` automatically, so `gl` is only an optional override, and `max_results` is an optional cap up to 500.
|
|
13
|
+
|
|
14
|
+
No breaking changes.
|
|
15
|
+
|
|
3
16
|
## 0.5.0 (2026-09-17)
|
|
4
17
|
|
|
5
18
|
- Added `google_images(q, ...)` for the new Google Images plugin (`POST /images/google-search`). Given a keyword `q` it returns Google Images results as a Hash - each with the full-size `imageUrl` and its `sourceDomain`, plus the source page URL, title, source name, thumbnail URL, pixel dimensions and file size. Optional `gl` (ISO-2 lowercase market), `max_results` (1-100) and `proxy_country` (ISO-2) refine the search.
|
data/README.md
CHANGED
|
@@ -131,7 +131,7 @@ local["results"].each { |biz| puts "#{biz['name']} #{biz['rating']} #{biz['addre
|
|
|
131
131
|
## Google Images
|
|
132
132
|
|
|
133
133
|
```ruby
|
|
134
|
-
images = su.google_images("golden retriever puppy", proxy_country: "US",
|
|
134
|
+
images = su.google_images("golden retriever puppy", proxy_country: "US", pages: 3)
|
|
135
135
|
images["results"].each { |img| puts "#{img['imageUrl']} #{img['sourceDomain']} #{img['title']}" }
|
|
136
136
|
```
|
|
137
137
|
|
|
@@ -277,7 +277,8 @@ end
|
|
|
277
277
|
| `CreditLimitExceededError` | 402 | Unpaid balance is past the account's credit limit |
|
|
278
278
|
| `PaymentFailedError` | 402 | A card payment was declined three times |
|
|
279
279
|
| `BlockedError` | 403 | Blocked by bot protection on every path |
|
|
280
|
-
| `NotFoundError` | 404 |
|
|
280
|
+
| `NotFoundError` | 404 | What you asked for does not exist - no image on the page (`get_image`), or a plugin lookup found nothing |
|
|
281
|
+
| `TargetNotFoundError` | 404 / 410 | The target page itself does not exist; carries `origin_status`, `html`, `destination_url` (subclass of `NotFoundError`, billed) |
|
|
281
282
|
| `BrowserTimeoutError` | 408 | Our browser run timed out before the page was ready |
|
|
282
283
|
| `UnsupportedContentError` | 415 | The URL serves something other than HTML |
|
|
283
284
|
| `ValidationError` | 422 | Missing or wrong-typed parameter; `body` holds the `detail` array |
|
|
@@ -287,6 +288,21 @@ end
|
|
|
287
288
|
| `TimeoutError` | - | This client gave up locally before the API answered |
|
|
288
289
|
| `ConnectionError` | - | Could not reach the API |
|
|
289
290
|
|
|
291
|
+
### The target page does not exist (404 / 410)
|
|
292
|
+
|
|
293
|
+
When the site you scrape answers 404 or 410 itself, the API passes that status through with an `X-Origin-Status` header, and the client raises `ScrapeUnblocker::TargetNotFoundError`. It is the target's final answer, so it is never retried, and it is billed like any delivered page. The not-found page is on `#html`:
|
|
294
|
+
|
|
295
|
+
```ruby
|
|
296
|
+
begin
|
|
297
|
+
html = su.get_page_source("https://example.com/removed-listing")
|
|
298
|
+
rescue ScrapeUnblocker::TargetNotFoundError => e
|
|
299
|
+
puts e.origin_status # 404 or 410
|
|
300
|
+
puts e.html # the target's own not-found page (can be empty)
|
|
301
|
+
end
|
|
302
|
+
```
|
|
303
|
+
|
|
304
|
+
`TargetNotFoundError` subclasses `NotFoundError`, so `rescue ScrapeUnblocker::NotFoundError` catches it too. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
|
|
305
|
+
|
|
290
306
|
Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
|
|
291
307
|
|
|
292
308
|
### Billing errors (402)
|
|
@@ -57,6 +57,9 @@ module ScrapeUnblocker
|
|
|
57
57
|
# +list_elements+, when true, makes the API return a JSON summary of the
|
|
58
58
|
# matched elements ({"url", "count", "elements"}) instead of HTML. This method
|
|
59
59
|
# then returns that parsed Hash rather than an HTML String.
|
|
60
|
+
#
|
|
61
|
+
# When the target page itself answers 404 or 410 this raises
|
|
62
|
+
# ScrapeUnblocker::TargetNotFoundError (billed; the not-found page is on #html).
|
|
60
63
|
def get_page_source(url, proxy_country: nil, time_sleep: nil, method: nil, value: nil, method_timeout: nil,
|
|
61
64
|
steps: nil, list_elements: nil)
|
|
62
65
|
body = request("/getPageSource",
|
|
@@ -111,11 +114,13 @@ module ScrapeUnblocker
|
|
|
111
114
|
# Returns image results, each with the full-size +imageUrl+ and its
|
|
112
115
|
# +sourceDomain+, plus the source page URL, title, source name, thumbnail
|
|
113
116
|
# URL, pixel dimensions and file size. +q+ is the search keyword. Set
|
|
114
|
-
# +proxy_country+
|
|
115
|
-
# +
|
|
116
|
-
|
|
117
|
+
# +proxy_country+ to target a market (the Google market follows it; +gl+
|
|
118
|
+
# is an optional override), +pages+ (1-5, ~100 results each; each page
|
|
119
|
+
# fetched is billed as one request, see +pagesFetched+) and +max_results+
|
|
120
|
+
# (1-500) to cap the count.
|
|
121
|
+
def google_images(q, proxy_country: nil, pages: nil, gl: nil, max_results: nil)
|
|
117
122
|
post_json("/images/google-search",
|
|
118
|
-
q: q, gl: gl, max_results: max_results, proxy_country: proxy_country)
|
|
123
|
+
q: q, pages: pages, gl: gl, max_results: max_results, proxy_country: proxy_country)
|
|
119
124
|
end
|
|
120
125
|
|
|
121
126
|
# Fetch an advertiser's Meta (Facebook) Ad Library ads and return them as a Hash.
|
|
@@ -282,7 +287,8 @@ module ScrapeUnblocker
|
|
|
282
287
|
|
|
283
288
|
return { status: status, body: body } if status >= 200 && status < 300
|
|
284
289
|
|
|
285
|
-
|
|
290
|
+
response_headers = (result[:headers] || {}).to_h { |k, v| [k.to_s.downcase, Array(v).join(", ")] }
|
|
291
|
+
raise ScrapeUnblocker.error_for_status(status, body, response_headers)
|
|
286
292
|
end
|
|
287
293
|
end
|
|
288
294
|
|
|
@@ -313,7 +319,7 @@ module ScrapeUnblocker
|
|
|
313
319
|
raise ConnectionError, "Could not reach the API: #{e.message}"
|
|
314
320
|
end
|
|
315
321
|
|
|
316
|
-
{ status: response.code.to_i, body: response.body }
|
|
322
|
+
{ status: response.code.to_i, body: response.body, headers: response.each_header.to_h }
|
|
317
323
|
end
|
|
318
324
|
end
|
|
319
325
|
end
|
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
+
require "json"
|
|
4
|
+
|
|
3
5
|
module ScrapeUnblocker
|
|
4
6
|
# Base class for every error raised by this library.
|
|
5
7
|
class Error < StandardError; end
|
|
@@ -78,10 +80,38 @@ module ScrapeUnblocker
|
|
|
78
80
|
# use instead.
|
|
79
81
|
class InvalidRequestError < APIError; end
|
|
80
82
|
|
|
81
|
-
#
|
|
82
|
-
#
|
|
83
|
+
# Something the call asked for does not exist (HTTP 404).
|
|
84
|
+
#
|
|
85
|
+
# #get_image raises it when the page rendered but held no <img> tag, and
|
|
86
|
+
# plugin methods raise it when the item they look up does not exist. When the
|
|
87
|
+
# target page itself answered 404 or 410, the more specific
|
|
88
|
+
# TargetNotFoundError subclass is raised instead.
|
|
83
89
|
class NotFoundError < APIError; end
|
|
84
90
|
|
|
91
|
+
# The target page itself does not exist (HTTP 404 or 410).
|
|
92
|
+
#
|
|
93
|
+
# Raised by #get_page_source, #get_parsed and #get_page_with_cookies when the
|
|
94
|
+
# site you asked for answered 404 or 410 on its own. The API passes that
|
|
95
|
+
# status through and marks it with the X-Origin-Status header, which is how
|
|
96
|
+
# this is told apart from an API-side 404. It is the target's final answer,
|
|
97
|
+
# not a block, so it is never retried - and the call is billed, because the
|
|
98
|
+
# page was fetched and delivered.
|
|
99
|
+
#
|
|
100
|
+
# +origin_status+ is the status the target answered with (404 or 410),
|
|
101
|
+
# +html+ the target's own not-found page as served (can be empty; nil when
|
|
102
|
+
# the body is a parsed-data JSON payload), and +destination_url+ the URL the
|
|
103
|
+
# target answered for, when the API sent X-Destination-URL.
|
|
104
|
+
class TargetNotFoundError < NotFoundError
|
|
105
|
+
attr_reader :origin_status, :html, :destination_url
|
|
106
|
+
|
|
107
|
+
def initialize(message, status_code:, origin_status:, body: nil, html: nil, destination_url: nil)
|
|
108
|
+
super(message, status_code: status_code, body: body)
|
|
109
|
+
@origin_status = origin_status
|
|
110
|
+
@html = html
|
|
111
|
+
@destination_url = destination_url
|
|
112
|
+
end
|
|
113
|
+
end
|
|
114
|
+
|
|
85
115
|
# The browser run did not finish in time on our side (HTTP 408).
|
|
86
116
|
#
|
|
87
117
|
# Distinct from TimeoutError, which is this client giving up locally. Here
|
|
@@ -153,8 +183,35 @@ module ScrapeUnblocker
|
|
|
153
183
|
end
|
|
154
184
|
private_class_method :billing_error_class
|
|
155
185
|
|
|
156
|
-
#
|
|
157
|
-
|
|
186
|
+
# The API passes a target's "page does not exist" answer through with its
|
|
187
|
+
# status and an X-Origin-Status header. A 404 without that header is the
|
|
188
|
+
# API's own (a plugin lookup, a missing element) and returns nil so the
|
|
189
|
+
# general NotFoundError applies.
|
|
190
|
+
def self.target_not_found_error(status, body, headers)
|
|
191
|
+
origin = headers["x-origin-status"]
|
|
192
|
+
return nil unless [404, 410].include?(status) && origin && !origin.to_s.empty?
|
|
193
|
+
|
|
194
|
+
origin_status = Integer(origin.to_s, exception: false) || status
|
|
195
|
+
html = body
|
|
196
|
+
begin
|
|
197
|
+
data = JSON.parse(body.to_s)
|
|
198
|
+
html = data["html"].is_a?(String) ? data["html"] : nil if data.is_a?(Hash)
|
|
199
|
+
rescue JSON::ParserError
|
|
200
|
+
# Not JSON: the body is the target's own page.
|
|
201
|
+
end
|
|
202
|
+
message = "Target page does not exist (HTTP #{origin_status}). This is the " \
|
|
203
|
+
"target's own answer, not a block; the call is billed."
|
|
204
|
+
TargetNotFoundError.new(message, status_code: status, body: body, origin_status: origin_status,
|
|
205
|
+
html: html, destination_url: headers["x-destination-url"])
|
|
206
|
+
end
|
|
207
|
+
private_class_method :target_not_found_error
|
|
208
|
+
|
|
209
|
+
# Build a typed error from an HTTP status code, response body and headers
|
|
210
|
+
# (a Hash with lowercase names).
|
|
211
|
+
def self.error_for_status(status, body, headers = {})
|
|
212
|
+
target_error = target_not_found_error(status, body, headers || {})
|
|
213
|
+
return target_error if target_error
|
|
214
|
+
|
|
158
215
|
snippet = (body || "").strip.gsub(/\s+/, " ")
|
|
159
216
|
snippet = "#{snippet[0, 200]}..." if snippet.length > 200
|
|
160
217
|
base = BASE_MESSAGES.fetch(status, "API returned HTTP #{status}")
|
metadata
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: scrapeunblocker
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.7.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- ScrapeUnblocker
|
|
8
8
|
autorequire:
|
|
9
9
|
bindir: bin
|
|
10
10
|
cert_chain: []
|
|
11
|
-
date: 2026-09-
|
|
11
|
+
date: 2026-09-30 00:00:00.000000000 Z
|
|
12
12
|
dependencies: []
|
|
13
13
|
description: JS-rendered pages that bypass Cloudflare, DataDome, PerimeterX and Akamai,
|
|
14
14
|
plus Google SERP and Skyscanner flights/hotels/car-hire scraping as JSON.
|