webscraping_ai 4.0.1 → 4.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +21 -0
- data/README.md +86 -1
- data/lib/webscraping_ai/client.rb +98 -2
- data/lib/webscraping_ai/configuration.rb +6 -0
- data/lib/webscraping_ai/version.rb +1 -1
- metadata +2 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 12a9ee8045725e182d8c4600d17b6c2cdfb74264ae23b378ab59818d279f9e73
|
|
4
|
+
data.tar.gz: 8d26e529421132c3b615cf1b795e11546308f9e00c3e125c3e8d0f4d57b3c4af
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 28552acbb15d4b656f232a30e0476073ba55be3334b5752c87be65f900d950a7892491a907acc5f61cb90276cd4126db950a2555109d6df5138f948b5e7afb71
|
|
7
|
+
data.tar.gz: 3d8e917eb2ddba519649b89b3db73c2224a3185bf27946978120f5efe9778cc0685e7a43cce256434c3ae296ddc7418aa05d6ec1bebddde24fb7f292325c83e4
|
data/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,27 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to this project will be documented in this file. This project follows [Semantic Versioning](https://semver.org/).
|
|
4
4
|
|
|
5
|
+
## 4.2.0 — 2026-09-25
|
|
6
|
+
### Added
|
|
7
|
+
|
|
8
|
+
- `Client#data(url, country: nil, transcript: nil, transcript_language: nil, **params)` for the `GET /data` endpoint — structured JSON for a page on a supported site (e.g. YouTube, TikTok, X, LinkedIn, Instagram, Reddit) as a Hash with `request_parameters`, `parse_status` and `data`. Flat 15 credits per request. The client does not check the URL's site: new sites are added server-side, and an unsupported URL or page type returns a 400 that is not charged (`BadRequestError`); its message lists what is supported. Raises `ArgumentError` when `url` is blank or not a String. Extra keyword arguments are sent as-is as query params (String, Integer, Float or boolean; Floats are sent as plain decimals). `api_key`, `url`, `country`, `transcript` and `transcript_language` as extra params raise `ArgumentError` (use the named options).
|
|
9
|
+
- `bin/smoke.rb` checks `/data` on a YouTube video and that `https://example.com/` gets the server's 400 "Unsupported URL" error (~47 credits per sweep).
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- `TimeoutError` and `ConnectionError` redact `api_key=...` from the transport error's message and no longer carry the Faraday exception as `cause`, since it may include the request URL.
|
|
14
|
+
|
|
15
|
+
## 4.1.0 — 2026-09-25
|
|
16
|
+
|
|
17
|
+
### Added
|
|
18
|
+
|
|
19
|
+
- `Client#serp(q:, engine: nil, gl: nil, hl: nil, page: nil)` for the new `GET /serp` endpoint — parsed Google search results (organic results, related searches, pagination) as a Hash. Flat 15 credits per search. Raises `ArgumentError` when `q` is blank or not a String, or when `page` is not an Integer >= 1 (the server also rejects it with a 400, not billed; checking client-side saves the round trip). Pages are 1–100: the server rejects a `page` above 100 with a 400.
|
|
20
|
+
- `bin/smoke.rb` live smoke script: asserts result shapes (not just the absence of exceptions), runs page tools with `js: false` on datacenter proxies (~32 credits per sweep), and redacts the API key from failure output.
|
|
21
|
+
|
|
22
|
+
### Fixed
|
|
23
|
+
|
|
24
|
+
- `Client#inspect` and `Configuration#inspect` no longer print the API key; it is shown as `api_key="[FILTERED]"`.
|
|
25
|
+
|
|
5
26
|
## 4.0.1 — 2026-07-17
|
|
6
27
|
|
|
7
28
|
### Changed
|
data/README.md
CHANGED
|
@@ -13,7 +13,7 @@ structured field extraction on any page. See the
|
|
|
13
13
|
|
|
14
14
|
```ruby
|
|
15
15
|
# Gemfile
|
|
16
|
-
gem "webscraping_ai", "~> 4.
|
|
16
|
+
gem "webscraping_ai", "~> 4.1"
|
|
17
17
|
```
|
|
18
18
|
|
|
19
19
|
Or:
|
|
@@ -60,6 +60,14 @@ data = client.fields(
|
|
|
60
60
|
}
|
|
61
61
|
)
|
|
62
62
|
|
|
63
|
+
# Google search results (SERP) for a query
|
|
64
|
+
results = client.serp(q: "coffee machines", gl: "us", hl: "en", page: 1)
|
|
65
|
+
results["organic_results"].first["link"]
|
|
66
|
+
|
|
67
|
+
# Structured data for a page on a supported site (YouTube, TikTok, X, LinkedIn, Instagram, Reddit, ...)
|
|
68
|
+
video = client.data("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
|
|
69
|
+
video["data"]["title"]
|
|
70
|
+
|
|
63
71
|
# Check your account quota
|
|
64
72
|
info = client.account
|
|
65
73
|
# => { "remaining_api_calls" => 200_000, "resets_at" => 1_617_073_667, "remaining_concurrency" => 100 }
|
|
@@ -121,6 +129,72 @@ Endpoint-specific options:
|
|
|
121
129
|
|
|
122
130
|
Returns: `String` for HTML/text responses, `Hash`/`Array` for JSON responses.
|
|
123
131
|
|
|
132
|
+
### SERP (`#serp`)
|
|
133
|
+
|
|
134
|
+
`#serp` is query-shaped rather than URL-shaped, so none of the page-fetch options above apply.
|
|
135
|
+
It returns the parsed search results as a `Hash`. Flat 15 credits per search; failed searches are not charged.
|
|
136
|
+
|
|
137
|
+
| Option | Type | Default | Description |
|
|
138
|
+
| --- | --- | --- | --- |
|
|
139
|
+
| `q` | `String` | — | Search query (required; blank or non-String raises `ArgumentError`) |
|
|
140
|
+
| `engine` | `String` | `"google"` | Search engine; currently only `google` |
|
|
141
|
+
| `gl` | `String` | `"us"` | Two-letter country code for the search |
|
|
142
|
+
| `hl` | `String` | `"en"` | Two-letter language code for the results |
|
|
143
|
+
| `page` | `Integer` | `1` | Results page number (10 results per page). Must be an `Integer` >= 1, otherwise `ArgumentError`; the server rejects values above 100 with a 400 (not billed) |
|
|
144
|
+
|
|
145
|
+
```ruby
|
|
146
|
+
results = client.serp(q: "coffee machines", gl: "gb", page: 2)
|
|
147
|
+
results["search_information"]["organic_results_state"] # => "Results for exact spelling"
|
|
148
|
+
results["organic_results"].each do |r|
|
|
149
|
+
puts "#{r["position"]}. #{r["title"]} — #{r["link"]}"
|
|
150
|
+
end
|
|
151
|
+
results["pagination"] # => { "current" => 2, "next" => 3 }
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
Response keys: `search_parameters` (`engine`, `q`, `gl`, `hl`, `page`), `search_information`
|
|
155
|
+
(`query_displayed`, `organic_results_state`, optional `showing_results_for` and `total_results`),
|
|
156
|
+
`organic_results` (`position` — 1-based within the page — `title`, `link`, `domain`, `displayed_link`,
|
|
157
|
+
optional `snippet` and `date`), optional `related_searches` (`query`), and `pagination` (`current`, optional `next`).
|
|
158
|
+
Optional keys are absent when Google does not show them.
|
|
159
|
+
|
|
160
|
+
### Structured data (`#data`)
|
|
161
|
+
|
|
162
|
+
`#data(url, country: nil, transcript: nil, transcript_language: nil, **params)` returns structured JSON
|
|
163
|
+
for a public page on a supported site as a `Hash`. Pass the page's normal URL; the site (`provider`) and
|
|
164
|
+
page kind (`type`) are detected from it. Flat 15 credits per request, including pages that parse empty
|
|
165
|
+
(`parse_status` `"parse_failed"`) or no longer exist (`"not_found"`); failed fetches are not charged.
|
|
166
|
+
None of the page-fetch options above apply.
|
|
167
|
+
|
|
168
|
+
Supported sites today include, for example, YouTube (video/channel/playlist), TikTok (video/profile),
|
|
169
|
+
X/Twitter (tweet/profile), LinkedIn (company/job/profile), Instagram (post/reel/profile) and Reddit
|
|
170
|
+
(post/subreddit/user). **More sites are added server-side**, and they work with this gem without an
|
|
171
|
+
upgrade: the client never checks the URL's site. An unsupported URL or page type returns a 400 that is
|
|
172
|
+
not charged (`WebScrapingAI::BadRequestError`). Its message lists what is supported.
|
|
173
|
+
|
|
174
|
+
| Option | Type | Default | Description |
|
|
175
|
+
| --- | --- | --- | --- |
|
|
176
|
+
| `url` | `String` | — | Page URL (required, positional; blank or non-String raises `ArgumentError`) |
|
|
177
|
+
| `country` | `String` | `"us"` | Two-letter country code of the proxy used to fetch the page, `us` by default |
|
|
178
|
+
| `transcript` | `Boolean` | `false` | YouTube videos only. Also fetch the video's transcript into `data.transcript`. It's null when no matching captions are available. If the transcript fetch itself fails, the whole request fails with a 500 and is not charged |
|
|
179
|
+
| `transcript_language` | `String` | — | Caption language to pick, e.g. `en` or `de`. Without it, English is preferred, then the first available track. If the video has no captions in that language, `data.transcript` is null |
|
|
180
|
+
| `**params` | `String`, `Integer`, `Float`, boolean | — | Extra query params sent as-is, for provider-specific params added later (`nil` omits one). `api_key`, `url`, `country`, `transcript` and `transcript_language` raise `ArgumentError`; use the named options for the last three |
|
|
181
|
+
|
|
182
|
+
```ruby
|
|
183
|
+
result = client.data("https://www.youtube.com/watch?v=dQw4w9WgXcQ", transcript: true)
|
|
184
|
+
result["request_parameters"] # => { "url" => "...", "provider" => "youtube", "type" => "video" }
|
|
185
|
+
result["parse_status"] # => "ok" (or "parse_failed" / "not_found")
|
|
186
|
+
result["data"]["title"] # shape depends on provider and type; may be nil
|
|
187
|
+
|
|
188
|
+
begin
|
|
189
|
+
client.data("https://example.com/")
|
|
190
|
+
rescue WebScrapingAI::BadRequestError => e
|
|
191
|
+
e.message # => "Unsupported URL for /data. Supported sites: youtube, tiktok, ..."
|
|
192
|
+
end
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
`provider`, `type` and `parse_status` are open sets of strings, and `data` is the decoded JSON as-is
|
|
196
|
+
(no per-site classes), so new sites and fields show up without a gem release.
|
|
197
|
+
|
|
124
198
|
## Error handling
|
|
125
199
|
|
|
126
200
|
All API errors inherit from `WebScrapingAI::ApiError` and expose `#status`, `#message`, `#status_code`, `#status_message`, `#body`, and `#response_body`.
|
|
@@ -157,6 +231,17 @@ bundle exec rspec
|
|
|
157
231
|
bundle exec rubocop
|
|
158
232
|
```
|
|
159
233
|
|
|
234
|
+
## Smoke testing
|
|
235
|
+
|
|
236
|
+
`bin/smoke.rb` hits every endpoint once against the live API, loading the gem from `lib/` so it tests the working tree. It is not part of the spec suite and costs ~47 credits per run: the four page calls run with `js: false` and `proxy: "datacenter"` (1 credit each), `question` and `fields` cost 6 each, and the SERP and `/data` (YouTube video) calls are 15 each. A second `/data` call on `https://example.com/` must come back as the server's free 400, proving there is no client-side site filter. Each case checks the result shape as well as exceptions (e.g. SERP must return organic results for the right query, `/data` must parse a title, `selected_multiple` must match something), and failure messages redact the API key.
|
|
237
|
+
|
|
238
|
+
```bash
|
|
239
|
+
WEBSCRAPING_AI_API_KEY=... bundle exec rake smoke
|
|
240
|
+
# or: WEBSCRAPING_AI_API_KEY=... ruby bin/smoke.rb
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
Each endpoint prints an `ok` or `FAIL` line; the script exits non-zero if any call fails.
|
|
244
|
+
|
|
160
245
|
## Links
|
|
161
246
|
|
|
162
247
|
- [WebScraping.AI](https://webscraping.ai) — features, pricing, signup
|
|
@@ -8,6 +8,7 @@ module WebScrapingAI
|
|
|
8
8
|
DEVICES = %w[desktop mobile tablet].freeze
|
|
9
9
|
TEXT_FORMATS = %w[plain xml json].freeze
|
|
10
10
|
FORMATS = %w[json text].freeze
|
|
11
|
+
SERP_ENGINES = %w[google].freeze
|
|
11
12
|
|
|
12
13
|
PAGE_FETCH_OPTIONS = %i[
|
|
13
14
|
headers timeout js js_timeout wait_for proxy country
|
|
@@ -66,6 +67,51 @@ module WebScrapingAI
|
|
|
66
67
|
get("/selected-multiple", url: url, selectors: Array(selectors), **opts.slice(*PAGE_FETCH_OPTIONS))
|
|
67
68
|
end
|
|
68
69
|
|
|
70
|
+
# GET /serp — returns parsed search engine results for `q` as a Hash
|
|
71
|
+
# (search_parameters, search_information, organic_results, related_searches, pagination).
|
|
72
|
+
# Query-shaped: none of the page-fetch options apply. Flat 15 credits per search.
|
|
73
|
+
# `q` must be a non-blank String (sent as-is, untrimmed). `page`, when given, must be an
|
|
74
|
+
# Integer >= 1; the server rejects values above 100 with a 400 (not billed).
|
|
75
|
+
def serp(q:, engine: nil, gl: nil, hl: nil, page: nil)
|
|
76
|
+
raise ArgumentError, "q is required" if q.nil? || (q.is_a?(String) && q.strip.empty?)
|
|
77
|
+
raise ArgumentError, "q must be a String" unless q.is_a?(String)
|
|
78
|
+
raise ArgumentError, "page must be an Integer >= 1" unless page.nil? || (page.is_a?(Integer) && page >= 1)
|
|
79
|
+
|
|
80
|
+
get("/serp", q: q, engine: engine, gl: gl, hl: hl, page: page)
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
# GET /data — returns structured JSON for a page on a supported site as a Hash
|
|
84
|
+
# (request_parameters: url/provider/type, parse_status, data). Flat 15 credits per request.
|
|
85
|
+
#
|
|
86
|
+
# The client never checks which site `url` belongs to: supported sites (e.g. YouTube, TikTok,
|
|
87
|
+
# X/Twitter, LinkedIn, Instagram, Reddit) are added server-side. An unsupported URL or page type
|
|
88
|
+
# returns a 400 that is not charged (BadRequestError); its message lists what is supported.
|
|
89
|
+
# `provider`, `type` and `parse_status` are open sets of strings; `data` may be nil.
|
|
90
|
+
#
|
|
91
|
+
# `country`: two-letter country code of the proxy used to fetch the page, `us` by default.
|
|
92
|
+
# `transcript`: YouTube videos only. Also fetch the video's transcript into `data.transcript`.
|
|
93
|
+
# It's null when no matching captions are available. If the transcript fetch itself fails, the
|
|
94
|
+
# whole request fails with a 500 and is not charged.
|
|
95
|
+
# `transcript_language`: caption language to pick, e.g. `en` or `de`. Without it, English is
|
|
96
|
+
# preferred, then the first available track. If the video has no captions in that language,
|
|
97
|
+
# `data.transcript` is null.
|
|
98
|
+
#
|
|
99
|
+
# Any other keyword arguments are sent as-is as extra query params (for provider-specific params
|
|
100
|
+
# added later). Values must be String, Integer, Float or boolean (nil omits the param); the names
|
|
101
|
+
# `api_key`, `url`, `country`, `transcript` and `transcript_language` raise ArgumentError.
|
|
102
|
+
# None of the page-fetch options apply.
|
|
103
|
+
def data(url, country: nil, transcript: nil, transcript_language: nil, **params)
|
|
104
|
+
raise ArgumentError, "url is required" if url.nil? || (url.is_a?(String) && url.strip.empty?)
|
|
105
|
+
raise ArgumentError, "url must be a String" unless url.is_a?(String)
|
|
106
|
+
|
|
107
|
+
named = { country: country, transcript: transcript, transcript_language: transcript_language }.compact
|
|
108
|
+
get("/data", url: url, **data_extra_params(params), **named)
|
|
109
|
+
end
|
|
110
|
+
|
|
111
|
+
def inspect
|
|
112
|
+
"#<#{self.class.name} base_url=#{configuration.base_url.inspect} api_key=\"[FILTERED]\">"
|
|
113
|
+
end
|
|
114
|
+
|
|
69
115
|
# GET /account — returns Hash with remaining_api_calls, resets_at, remaining_concurrency, email.
|
|
70
116
|
def account
|
|
71
117
|
get("/account")
|
|
@@ -73,6 +119,54 @@ module WebScrapingAI
|
|
|
73
119
|
|
|
74
120
|
private
|
|
75
121
|
|
|
122
|
+
DATA_RESERVED_PARAMS = %w[api_key url].freeze
|
|
123
|
+
DATA_TYPED_PARAMS = %w[country transcript transcript_language].freeze
|
|
124
|
+
private_constant :DATA_RESERVED_PARAMS, :DATA_TYPED_PARAMS
|
|
125
|
+
|
|
126
|
+
# Extra /data query params are passed through untouched (same encoder, so `&`/`=` are escaped),
|
|
127
|
+
# except that they can't override the credentials or the target URL, and can't repeat a typed
|
|
128
|
+
# param (a string key like "country" would otherwise collide with `country:`).
|
|
129
|
+
def data_extra_params(params)
|
|
130
|
+
params.each_with_object({}) do |(key, value), extra|
|
|
131
|
+
name = key.to_s
|
|
132
|
+
raise ArgumentError, "#{name} can't be passed as an extra /data param" if DATA_RESERVED_PARAMS.include?(name)
|
|
133
|
+
if DATA_TYPED_PARAMS.include?(name)
|
|
134
|
+
raise ArgumentError, "#{name} can't be passed as an extra /data param; use the #{name}: option"
|
|
135
|
+
end
|
|
136
|
+
|
|
137
|
+
extra[name.to_sym] = data_extra_value(name, value)
|
|
138
|
+
end
|
|
139
|
+
end
|
|
140
|
+
|
|
141
|
+
def data_extra_value(name, value)
|
|
142
|
+
case value
|
|
143
|
+
when nil, String, Integer, true, false then value
|
|
144
|
+
when Float
|
|
145
|
+
raise ArgumentError, "extra /data param #{name} must be a finite number" unless value.finite?
|
|
146
|
+
|
|
147
|
+
plain_float(value)
|
|
148
|
+
else
|
|
149
|
+
raise ArgumentError, "extra /data param #{name} must be a String, Integer, Float or boolean"
|
|
150
|
+
end
|
|
151
|
+
end
|
|
152
|
+
|
|
153
|
+
# Float#to_s switches to exponent notation (1.0e+20, 1.5e-07); send plain decimals instead.
|
|
154
|
+
def plain_float(value)
|
|
155
|
+
text = value.to_s
|
|
156
|
+
return text unless text.include?("e")
|
|
157
|
+
|
|
158
|
+
mantissa, exponent = text.split("e")
|
|
159
|
+
fraction_digits = mantissa.split(".")[1].to_s.sub(/0+\z/, "").length
|
|
160
|
+
format("%.#{[fraction_digits - exponent.to_i, 0].max}f", value)
|
|
161
|
+
end
|
|
162
|
+
|
|
163
|
+
def redact(text)
|
|
164
|
+
key = configuration.api_key.to_s
|
|
165
|
+
text = text.to_s
|
|
166
|
+
text = text.gsub(key, "[FILTERED]") unless key.empty?
|
|
167
|
+
text.gsub(/api_key=[^&\s"']*/, "api_key=[FILTERED]")
|
|
168
|
+
end
|
|
169
|
+
|
|
76
170
|
def connection
|
|
77
171
|
@connection ||= Faraday.new(url: configuration.base_url) do |conn|
|
|
78
172
|
conn.options.timeout = configuration.timeout
|
|
@@ -84,15 +178,17 @@ module WebScrapingAI
|
|
|
84
178
|
end
|
|
85
179
|
end
|
|
86
180
|
|
|
181
|
+
# Transport errors are re-raised with `cause: nil` and a redacted message: the Faraday error (and
|
|
182
|
+
# anything it wraps) may carry the request URL, which contains the API key.
|
|
87
183
|
def get(path, **params)
|
|
88
184
|
response = connection.get(path) do |req|
|
|
89
185
|
req.params = params.merge(api_key: configuration.api_key)
|
|
90
186
|
end
|
|
91
187
|
handle_response(response)
|
|
92
188
|
rescue Faraday::TimeoutError => e
|
|
93
|
-
raise TimeoutError
|
|
189
|
+
raise TimeoutError.new(redact(e.message)), cause: nil
|
|
94
190
|
rescue Faraday::ConnectionFailed => e
|
|
95
|
-
raise ConnectionError
|
|
191
|
+
raise ConnectionError.new(redact(e.message)), cause: nil
|
|
96
192
|
end
|
|
97
193
|
|
|
98
194
|
def handle_response(response)
|
|
@@ -14,5 +14,11 @@ module WebScrapingAI
|
|
|
14
14
|
@adapter = nil
|
|
15
15
|
@user_agent = "webscraping_ai-ruby/#{WebScrapingAI::VERSION}"
|
|
16
16
|
end
|
|
17
|
+
|
|
18
|
+
def inspect
|
|
19
|
+
filtered = api_key.nil? ? "nil" : "\"[FILTERED]\""
|
|
20
|
+
"#<#{self.class.name} api_key=#{filtered} base_url=#{base_url.inspect} timeout=#{timeout.inspect} " \
|
|
21
|
+
"open_timeout=#{open_timeout.inspect} adapter=#{adapter.inspect} user_agent=#{user_agent.inspect}>"
|
|
22
|
+
end
|
|
17
23
|
end
|
|
18
24
|
end
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: webscraping_ai
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 4.0
|
|
4
|
+
version: 4.2.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- WebScraping.AI
|
|
@@ -66,7 +66,7 @@ required_rubygems_version: !ruby/object:Gem::Requirement
|
|
|
66
66
|
- !ruby/object:Gem::Version
|
|
67
67
|
version: '0'
|
|
68
68
|
requirements: []
|
|
69
|
-
rubygems_version: 4.0.
|
|
69
|
+
rubygems_version: 4.0.20
|
|
70
70
|
specification_version: 4
|
|
71
71
|
summary: Ruby client for the WebScraping.AI API.
|
|
72
72
|
test_files: []
|