scrapeunblocker 0.7.0 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 941309503bf1a937b76c503a811aab8e7e738e4ec775360ad7f9f6cfc40b6259
4
- data.tar.gz: 68753623ec6ad6623632c520e688a5ba2f7697abf6588b8351cdd33e911f473c
3
+ metadata.gz: e8096ca55e9f93c01077d947e60e36ec71c595e7ee92d1405a413a0746c303f0
4
+ data.tar.gz: 28c0667c78459edf026b26effd0209ea0635409eb078f152b8e4963be2312dd3
5
5
  SHA512:
6
- metadata.gz: 14f6f077622d3669a8ef06944d4cb810b89d693d28eead00b1b5789242945478e0e83529ede25aed9a474ee79351ca5202a3c91c03a174fd0cc811a031a48e11
7
- data.tar.gz: 23314b94ae64e4ce3ade4bbef52278df0639681801981865e3faa8914fd71f6dec7f3047fbad96e4954dc70b8d603ef35487b5a3d5a9a65623be6e7de41071f7
6
+ metadata.gz: f1b51dee641104d7d77fd600c5cb2299f30d6dcdf3e155686befc2d3024c981fb0c904d2aa8d5d1f360797a9cdbd331356ec2c99f1a3832907704198689306c6
7
+ data.tar.gz: 8274485fa1644d6dfb1512bc33e93fc034f3c01a0248e450869315f3d0669244761ef146db9e11cac0ccfb3fe6361148d5b3ef64a4c706d967d825ef23fba9e6
data/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.8.0 (2026-10-06)
4
+
5
+ - `get_parsed` on a page that rendered but held no structured data now returns a normal `ParsedPage` instead of raising `NoDataExtractedError`. The API answers this case with a billed 200: `data` is empty, the new `data_extracted?` is `false`, the new `detail` explains it and the new `html` carries the rendered page, so no second call is needed. On a successful parse `data_extracted?` is `true` and `html` / `detail` are `nil`.
6
+ - `NoDataExtractedError` is deprecated: the API no longer sends the 422 `no_data_extracted`. The class stays so existing `rescue` clauses keep loading.
7
+
8
+ Behaviour change: code that rescued `NoDataExtractedError` should check `page.data_extracted?` instead.
9
+
10
+ ## 0.7.1 (2026-09-30)
11
+
12
+ - New `ScrapeUnblocker::NoDataExtractedError` (a subclass of `ValidationError`): raised by `get_parsed` when the page rendered but no structured data could be extracted from it. The API answers 422 with `{"error": "no_data_extracted", "detail": ...}`; the error carries `detail`. The call is not billed and is never retried - use `get_page_source` for the HTML. Before, this came back as a billed 200 with empty `data`.
13
+ - A dead URL fetched with `get_parsed` raises `TargetNotFoundError` with `html` nil (the body is the parsed-data JSON, on `body`).
14
+
3
15
  ## 0.7.0 (2026-09-30)
4
16
 
5
17
  - New `ScrapeUnblocker::TargetNotFoundError` (a subclass of `NotFoundError`): raised by `get_page_source`, `get_parsed` and `get_page_with_cookies` when the target page itself answers 404 or 410. The API now passes the target's own status through instead of a 200, marked with the `X-Origin-Status` header. The error carries `origin_status`, the not-found page on `html` and `destination_url`. It is never retried, and the call is billed like any delivered page.
data/README.md CHANGED
@@ -303,6 +303,22 @@ end
303
303
 
304
304
  `TargetNotFoundError` subclasses `NotFoundError`, so `rescue ScrapeUnblocker::NotFoundError` catches it too. A 404 without `X-Origin-Status` is the API's own and stays a plain `NotFoundError`.
305
305
 
306
+ With `get_parsed` the body is the parsed-data JSON (`{"data": {"page_type": "not_found", ...}}`), so `html` is `nil` there; the raw JSON is on `#body`.
307
+
308
+ ### No structured data on the page
309
+
310
+ When `get_parsed` renders the page but can extract no structured data from it, the API still answers 200 and the client returns a normal `ParsedPage`: `data` is empty, `data_extracted?` is `false`, `detail` says why, and `html` carries the rendered page. The call is billed like any delivered page, so you already have the HTML - no second call needed:
311
+
312
+ ```ruby
313
+ page = su.get_parsed("https://example.com/some-page")
314
+ unless page.data_extracted?
315
+ puts page.detail # the API's explanation
316
+ html = page.html # the rendered page, parse it yourself
317
+ end
318
+ ```
319
+
320
+ On a successful parse `data_extracted?` is `true` and `html` is `nil`. Up to 0.7.1 this case raised `NoDataExtractedError` (a 422); the API no longer sends that, and the class stays only so existing `rescue` clauses keep loading.
321
+
306
322
  Transient failures (429, 502, 503, 504 and network errors) are retried automatically with exponential backoff. A 401 or 402 is never retried - it clears when the key or the billing state changes, not on another attempt. Neither is billed or counted against your quota, because the request is refused before anything is scraped.
307
323
 
308
324
  ### Billing errors (402)
@@ -73,6 +73,12 @@ module ScrapeUnblocker
73
73
  end
74
74
 
75
75
  # Fetch a URL and return structured JSON instead of HTML.
76
+ #
77
+ # A page that rendered but held no structured data comes back as a
78
+ # ParsedPage with #data_extracted? false and the rendered page on #html
79
+ # (billed like any delivered page). Raises
80
+ # ScrapeUnblocker::TargetNotFoundError (billed, #html nil) when the target
81
+ # page itself answered 404 or 410.
76
82
  def get_parsed(url, proxy_country: nil, time_sleep: nil, refresh_rules: false, rules_hint: nil)
77
83
  body = request("/getPageSource",
78
84
  url: url, parsed_data: true, proxy_country: proxy_country,
@@ -128,6 +128,22 @@ module ScrapeUnblocker
128
128
  # each problem field. Read it from #body.
129
129
  class ValidationError < APIError; end
130
130
 
131
+ # Deprecated: the API no longer sends this 422.
132
+ #
133
+ # A page that rendered but held no structured data now comes back from
134
+ # #get_parsed as a normal ParsedPage with #data_extracted? false and the page
135
+ # on #html (a billed 200). The class stays so existing rescue clauses keep
136
+ # loading; it is raised only for a legacy 422 body of
137
+ # {"error": "no_data_extracted", "detail": ...}.
138
+ class NoDataExtractedError < ValidationError
139
+ attr_reader :detail
140
+
141
+ def initialize(message, status_code:, body: nil, detail: nil)
142
+ super(message, status_code: status_code, body: body)
143
+ @detail = detail
144
+ end
145
+ end
146
+
131
147
  # The target site blocked every available bypass path (HTTP 403).
132
148
  # Blocked calls are not billed.
133
149
  class BlockedError < APIError; end
@@ -206,12 +222,35 @@ module ScrapeUnblocker
206
222
  end
207
223
  private_class_method :target_not_found_error
208
224
 
225
+ # Maps a legacy 422 {"error": "no_data_extracted", "detail"} body; the API now
226
+ # answers an empty parse with a 200 (data_extracted: false). Anything else
227
+ # returns nil so the general ValidationError applies.
228
+ def self.no_data_extracted_error(status, body)
229
+ return nil unless status == 422
230
+
231
+ data = begin
232
+ JSON.parse(body.to_s)
233
+ rescue JSON::ParserError
234
+ nil
235
+ end
236
+ return nil unless data.is_a?(Hash) && data["error"] == "no_data_extracted"
237
+
238
+ detail = data["detail"].is_a?(String) && !data["detail"].empty? ? data["detail"] : nil
239
+ message = detail || "The page was rendered, but no structured data could be extracted from it. Not billed."
240
+ message = "#{message} Not billed." unless message.downcase.include?("not billed")
241
+ NoDataExtractedError.new(message, status_code: status, body: body, detail: detail)
242
+ end
243
+ private_class_method :no_data_extracted_error
244
+
209
245
  # Build a typed error from an HTTP status code, response body and headers
210
246
  # (a Hash with lowercase names).
211
247
  def self.error_for_status(status, body, headers = {})
212
248
  target_error = target_not_found_error(status, body, headers || {})
213
249
  return target_error if target_error
214
250
 
251
+ no_data_error = no_data_extracted_error(status, body)
252
+ return no_data_error if no_data_error
253
+
215
254
  snippet = (body || "").strip.gsub(/\s+/, " ")
216
255
  snippet = "#{snippet[0, 200]}..." if snippet.length > 200
217
256
  base = BASE_MESSAGES.fetch(status, "API returned HTTP #{status}")
@@ -11,12 +11,25 @@ module ScrapeUnblocker
11
11
  attr_reader :data
12
12
  # @return [Hash] the full JSON payload as returned by the API
13
13
  attr_reader :raw
14
+ # @return [String, nil] the rendered page when #data_extracted? is false
15
+ attr_reader :html
16
+ # @return [String, nil] the API's explanation when #data_extracted? is false
17
+ attr_reader :detail
14
18
 
15
- def initialize(page_type:, source:, data:, raw:)
19
+ def initialize(page_type:, source:, data:, raw:, data_extracted: true, html: nil, detail: nil)
16
20
  @page_type = page_type
17
21
  @source = source
18
22
  @data = data
19
23
  @raw = raw
24
+ @data_extracted = data_extracted
25
+ @html = html
26
+ @detail = detail
27
+ end
28
+
29
+ # @return [Boolean] false when the page rendered but no structured data
30
+ # could be extracted from it (#data is then empty, #html holds the page)
31
+ def data_extracted?
32
+ @data_extracted
20
33
  end
21
34
 
22
35
  def self.from_hash(payload)
@@ -25,7 +38,10 @@ module ScrapeUnblocker
25
38
  page_type: inner["page_type"],
26
39
  source: inner["source"],
27
40
  data: inner["data"],
28
- raw: payload
41
+ raw: payload,
42
+ data_extracted: payload["data_extracted"] != false,
43
+ html: payload["html"],
44
+ detail: payload["detail"]
29
45
  )
30
46
  end
31
47
  end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module ScrapeUnblocker
4
- VERSION = "0.7.0"
4
+ VERSION = "0.8.0"
5
5
  end
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: scrapeunblocker
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.7.0
4
+ version: 0.8.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - ScrapeUnblocker
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-30 00:00:00.000000000 Z
11
+ date: 2026-10-06 00:00:00.000000000 Z
12
12
  dependencies: []
13
13
  description: JS-rendered pages that bypass Cloudflare, DataDome, PerimeterX and Akamai,
14
14
  plus Google SERP and Skyscanner flights/hotels/car-hire scraping as JSON.