webscraping_ai 4.1.0 → 4.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: cd7230f48fdc600563c0e7fe84ba65411f6363007854a9bdaae6b2ee6bfbf17d
4
- data.tar.gz: a815bbc4645cd9eb23aba4e9cdafff0ac1c420f3a230580e002a59f7f19b28f6
3
+ metadata.gz: 12a9ee8045725e182d8c4600d17b6c2cdfb74264ae23b378ab59818d279f9e73
4
+ data.tar.gz: 8d26e529421132c3b615cf1b795e11546308f9e00c3e125c3e8d0f4d57b3c4af
5
5
  SHA512:
6
- metadata.gz: 8266c1f95d1432a37f817e3308e1c99c1795fba17fc105c4ef703561e89f84104f0d2734fda99dded525000baecd346d00c2ac3f0b838e251807d3651d20fae6
7
- data.tar.gz: 6d075d815694065fde5b830d577100c8c45c9943798689ecc50b9959120d0486b10ede96b77344fb0803b8094f2855cf60d72237161b07a480e31b4531c62b1c
6
+ metadata.gz: 28552acbb15d4b656f232a30e0476073ba55be3334b5752c87be65f900d950a7892491a907acc5f61cb90276cd4126db950a2555109d6df5138f948b5e7afb71
7
+ data.tar.gz: 3d8e917eb2ddba519649b89b3db73c2224a3185bf27946978120f5efe9778cc0685e7a43cce256434c3ae296ddc7418aa05d6ec1bebddde24fb7f292325c83e4
data/CHANGELOG.md CHANGED
@@ -2,6 +2,16 @@
2
2
 
3
3
  All notable changes to this project will be documented in this file. This project follows [Semantic Versioning](https://semver.org/).
4
4
 
5
+ ## 4.2.0 — 2026-09-25
6
+ ### Added
7
+
8
+ - `Client#data(url, country: nil, transcript: nil, transcript_language: nil, **params)` for the `GET /data` endpoint — structured JSON for a page on a supported site (e.g. YouTube, TikTok, X, LinkedIn, Instagram, Reddit) as a Hash with `request_parameters`, `parse_status` and `data`. Flat 15 credits per request. The client does not check the URL's site: new sites are added server-side, and an unsupported URL or page type returns a 400 that is not charged (`BadRequestError`); its message lists what is supported. Raises `ArgumentError` when `url` is blank or not a String. Extra keyword arguments are sent as-is as query params (String, Integer, Float or boolean; Floats are sent as plain decimals). `api_key`, `url`, `country`, `transcript` and `transcript_language` as extra params raise `ArgumentError` (use the named options).
9
+ - `bin/smoke.rb` checks `/data` on a YouTube video and that `https://example.com/` gets the server's 400 "Unsupported URL" error (~47 credits per sweep).
10
+
11
+ ### Fixed
12
+
13
+ - `TimeoutError` and `ConnectionError` redact `api_key=...` from the transport error's message and no longer carry the Faraday exception as `cause`, since it may include the request URL.
14
+
5
15
  ## 4.1.0 — 2026-09-25
6
16
 
7
17
  ### Added
data/README.md CHANGED
@@ -64,6 +64,10 @@ data = client.fields(
64
64
  results = client.serp(q: "coffee machines", gl: "us", hl: "en", page: 1)
65
65
  results["organic_results"].first["link"]
66
66
 
67
+ # Structured data for a page on a supported site (YouTube, TikTok, X, LinkedIn, Instagram, Reddit, ...)
68
+ video = client.data("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
69
+ video["data"]["title"]
70
+
67
71
  # Check your account quota
68
72
  info = client.account
69
73
  # => { "remaining_api_calls" => 200_000, "resets_at" => 1_617_073_667, "remaining_concurrency" => 100 }
@@ -153,6 +157,44 @@ Response keys: `search_parameters` (`engine`, `q`, `gl`, `hl`, `page`), `search_
153
157
  optional `snippet` and `date`), optional `related_searches` (`query`), and `pagination` (`current`, optional `next`).
154
158
  Optional keys are absent when Google does not show them.
155
159
 
160
+ ### Structured data (`#data`)
161
+
162
+ `#data(url, country: nil, transcript: nil, transcript_language: nil, **params)` returns structured JSON
163
+ for a public page on a supported site as a `Hash`. Pass the page's normal URL; the site (`provider`) and
164
+ page kind (`type`) are detected from it. Flat 15 credits per request, including pages that parse empty
165
+ (`parse_status` `"parse_failed"`) or no longer exist (`"not_found"`); failed fetches are not charged.
166
+ None of the page-fetch options above apply.
167
+
168
+ Supported sites today include, for example, YouTube (video/channel/playlist), TikTok (video/profile),
169
+ X/Twitter (tweet/profile), LinkedIn (company/job/profile), Instagram (post/reel/profile) and Reddit
170
+ (post/subreddit/user). **More sites are added server-side**, and they work with this gem without an
171
+ upgrade: the client never checks the URL's site. An unsupported URL or page type returns a 400 that is
172
+ not charged (`WebScrapingAI::BadRequestError`). Its message lists what is supported.
173
+
174
+ | Option | Type | Default | Description |
175
+ | --- | --- | --- | --- |
176
+ | `url` | `String` | — | Page URL (required, positional; blank or non-String raises `ArgumentError`) |
177
+ | `country` | `String` | `"us"` | Two-letter country code of the proxy used to fetch the page, `us` by default |
178
+ | `transcript` | `Boolean` | `false` | YouTube videos only. Also fetch the video's transcript into `data.transcript`. It's null when no matching captions are available. If the transcript fetch itself fails, the whole request fails with a 500 and is not charged |
179
+ | `transcript_language` | `String` | — | Caption language to pick, e.g. `en` or `de`. Without it, English is preferred, then the first available track. If the video has no captions in that language, `data.transcript` is null |
180
+ | `**params` | `String`, `Integer`, `Float`, boolean | — | Extra query params sent as-is, for provider-specific params added later (`nil` omits one). `api_key`, `url`, `country`, `transcript` and `transcript_language` raise `ArgumentError`; use the named options for the last three |
181
+
182
+ ```ruby
183
+ result = client.data("https://www.youtube.com/watch?v=dQw4w9WgXcQ", transcript: true)
184
+ result["request_parameters"] # => { "url" => "...", "provider" => "youtube", "type" => "video" }
185
+ result["parse_status"] # => "ok" (or "parse_failed" / "not_found")
186
+ result["data"]["title"] # shape depends on provider and type; may be nil
187
+
188
+ begin
189
+ client.data("https://example.com/")
190
+ rescue WebScrapingAI::BadRequestError => e
191
+ e.message # => "Unsupported URL for /data. Supported sites: youtube, tiktok, ..."
192
+ end
193
+ ```
194
+
195
+ `provider`, `type` and `parse_status` are open sets of strings, and `data` is the decoded JSON as-is
196
+ (no per-site classes), so new sites and fields show up without a gem release.
197
+
156
198
  ## Error handling
157
199
 
158
200
  All API errors inherit from `WebScrapingAI::ApiError` and expose `#status`, `#message`, `#status_code`, `#status_message`, `#body`, and `#response_body`.
@@ -191,7 +233,7 @@ bundle exec rubocop
191
233
 
192
234
  ## Smoke testing
193
235
 
194
- `bin/smoke.rb` hits every endpoint once against the live API, loading the gem from `lib/` so it tests the working tree. It is not part of the spec suite and costs ~32 credits per run: the four page calls run with `js: false` and `proxy: "datacenter"` (1 credit each), `question` and `fields` cost 6 each, and the SERP call is 15. Each case checks the result shape as well as exceptions (e.g. SERP must return organic results for the right query, `selected_multiple` must match something), and failure messages redact the API key.
236
+ `bin/smoke.rb` hits every endpoint once against the live API, loading the gem from `lib/` so it tests the working tree. It is not part of the spec suite and costs ~47 credits per run: the four page calls run with `js: false` and `proxy: "datacenter"` (1 credit each), `question` and `fields` cost 6 each, and the SERP and `/data` (YouTube video) calls are 15 each. A second `/data` call on `https://example.com/` must come back as the server's free 400, proving there is no client-side site filter. Each case checks the result shape as well as exceptions (e.g. SERP must return organic results for the right query, `/data` must parse a title, `selected_multiple` must match something), and failure messages redact the API key.
195
237
 
196
238
  ```bash
197
239
  WEBSCRAPING_AI_API_KEY=... bundle exec rake smoke
@@ -80,6 +80,34 @@ module WebScrapingAI
80
80
  get("/serp", q: q, engine: engine, gl: gl, hl: hl, page: page)
81
81
  end
82
82
 
83
+ # GET /data — returns structured JSON for a page on a supported site as a Hash
84
+ # (request_parameters: url/provider/type, parse_status, data). Flat 15 credits per request.
85
+ #
86
+ # The client never checks which site `url` belongs to: supported sites (e.g. YouTube, TikTok,
87
+ # X/Twitter, LinkedIn, Instagram, Reddit) are added server-side. An unsupported URL or page type
88
+ # returns a 400 that is not charged (BadRequestError); its message lists what is supported.
89
+ # `provider`, `type` and `parse_status` are open sets of strings; `data` may be nil.
90
+ #
91
+ # `country`: two-letter country code of the proxy used to fetch the page, `us` by default.
92
+ # `transcript`: YouTube videos only. Also fetch the video's transcript into `data.transcript`.
93
+ # It's null when no matching captions are available. If the transcript fetch itself fails, the
94
+ # whole request fails with a 500 and is not charged.
95
+ # `transcript_language`: caption language to pick, e.g. `en` or `de`. Without it, English is
96
+ # preferred, then the first available track. If the video has no captions in that language,
97
+ # `data.transcript` is null.
98
+ #
99
+ # Any other keyword arguments are sent as-is as extra query params (for provider-specific params
100
+ # added later). Values must be String, Integer, Float or boolean (nil omits the param); the names
101
+ # `api_key`, `url`, `country`, `transcript` and `transcript_language` raise ArgumentError.
102
+ # None of the page-fetch options apply.
103
+ def data(url, country: nil, transcript: nil, transcript_language: nil, **params)
104
+ raise ArgumentError, "url is required" if url.nil? || (url.is_a?(String) && url.strip.empty?)
105
+ raise ArgumentError, "url must be a String" unless url.is_a?(String)
106
+
107
+ named = { country: country, transcript: transcript, transcript_language: transcript_language }.compact
108
+ get("/data", url: url, **data_extra_params(params), **named)
109
+ end
110
+
83
111
  def inspect
84
112
  "#<#{self.class.name} base_url=#{configuration.base_url.inspect} api_key=\"[FILTERED]\">"
85
113
  end
@@ -91,6 +119,54 @@ module WebScrapingAI
91
119
 
92
120
  private
93
121
 
122
+ DATA_RESERVED_PARAMS = %w[api_key url].freeze
123
+ DATA_TYPED_PARAMS = %w[country transcript transcript_language].freeze
124
+ private_constant :DATA_RESERVED_PARAMS, :DATA_TYPED_PARAMS
125
+
126
+ # Extra /data query params are passed through untouched (same encoder, so `&`/`=` are escaped),
127
+ # except that they can't override the credentials or the target URL, and can't repeat a typed
128
+ # param (a string key like "country" would otherwise collide with `country:`).
129
+ def data_extra_params(params)
130
+ params.each_with_object({}) do |(key, value), extra|
131
+ name = key.to_s
132
+ raise ArgumentError, "#{name} can't be passed as an extra /data param" if DATA_RESERVED_PARAMS.include?(name)
133
+ if DATA_TYPED_PARAMS.include?(name)
134
+ raise ArgumentError, "#{name} can't be passed as an extra /data param; use the #{name}: option"
135
+ end
136
+
137
+ extra[name.to_sym] = data_extra_value(name, value)
138
+ end
139
+ end
140
+
141
+ def data_extra_value(name, value)
142
+ case value
143
+ when nil, String, Integer, true, false then value
144
+ when Float
145
+ raise ArgumentError, "extra /data param #{name} must be a finite number" unless value.finite?
146
+
147
+ plain_float(value)
148
+ else
149
+ raise ArgumentError, "extra /data param #{name} must be a String, Integer, Float or boolean"
150
+ end
151
+ end
152
+
153
+ # Float#to_s switches to exponent notation (1.0e+20, 1.5e-07); send plain decimals instead.
154
+ def plain_float(value)
155
+ text = value.to_s
156
+ return text unless text.include?("e")
157
+
158
+ mantissa, exponent = text.split("e")
159
+ fraction_digits = mantissa.split(".")[1].to_s.sub(/0+\z/, "").length
160
+ format("%.#{[fraction_digits - exponent.to_i, 0].max}f", value)
161
+ end
162
+
163
+ def redact(text)
164
+ key = configuration.api_key.to_s
165
+ text = text.to_s
166
+ text = text.gsub(key, "[FILTERED]") unless key.empty?
167
+ text.gsub(/api_key=[^&\s"']*/, "api_key=[FILTERED]")
168
+ end
169
+
94
170
  def connection
95
171
  @connection ||= Faraday.new(url: configuration.base_url) do |conn|
96
172
  conn.options.timeout = configuration.timeout
@@ -102,15 +178,17 @@ module WebScrapingAI
102
178
  end
103
179
  end
104
180
 
181
+ # Transport errors are re-raised with `cause: nil` and a redacted message: the Faraday error (and
182
+ # anything it wraps) may carry the request URL, which contains the API key.
105
183
  def get(path, **params)
106
184
  response = connection.get(path) do |req|
107
185
  req.params = params.merge(api_key: configuration.api_key)
108
186
  end
109
187
  handle_response(response)
110
188
  rescue Faraday::TimeoutError => e
111
- raise TimeoutError, e.message
189
+ raise TimeoutError.new(redact(e.message)), cause: nil
112
190
  rescue Faraday::ConnectionFailed => e
113
- raise ConnectionError, e.message
191
+ raise ConnectionError.new(redact(e.message)), cause: nil
114
192
  end
115
193
 
116
194
  def handle_response(response)
@@ -1,3 +1,3 @@
1
1
  module WebScrapingAI
2
- VERSION = "4.1.0".freeze
2
+ VERSION = "4.2.0".freeze
3
3
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: webscraping_ai
3
3
  version: !ruby/object:Gem::Version
4
- version: 4.1.0
4
+ version: 4.2.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - WebScraping.AI