webscraping_ai 4.0.0 → 4.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +17 -0
- data/README.md +63 -3
- data/lib/webscraping_ai/client.rb +18 -0
- data/lib/webscraping_ai/configuration.rb +6 -0
- data/lib/webscraping_ai/version.rb +1 -1
- metadata +3 -6
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: cd7230f48fdc600563c0e7fe84ba65411f6363007854a9bdaae6b2ee6bfbf17d
|
|
4
|
+
data.tar.gz: a815bbc4645cd9eb23aba4e9cdafff0ac1c420f3a230580e002a59f7f19b28f6
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 8266c1f95d1432a37f817e3308e1c99c1795fba17fc105c4ef703561e89f84104f0d2734fda99dded525000baecd346d00c2ac3f0b838e251807d3651d20fae6
|
|
7
|
+
data.tar.gz: 6d075d815694065fde5b830d577100c8c45c9943798689ecc50b9959120d0486b10ede96b77344fb0803b8094f2855cf60d72237161b07a480e31b4531c62b1c
|
data/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,23 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to this project will be documented in this file. This project follows [Semantic Versioning](https://semver.org/).
|
|
4
4
|
|
|
5
|
+
## 4.1.0 — 2026-09-25
|
|
6
|
+
|
|
7
|
+
### Added
|
|
8
|
+
|
|
9
|
+
- `Client#serp(q:, engine: nil, gl: nil, hl: nil, page: nil)` for the new `GET /serp` endpoint — parsed Google search results (organic results, related searches, pagination) as a Hash. Flat 15 credits per search. Raises `ArgumentError` when `q` is blank or not a String, or when `page` is not an Integer >= 1 (the server also rejects it with a 400, not billed; checking client-side saves the round trip). Pages are 1–100: the server rejects a `page` above 100 with a 400.
|
|
10
|
+
- `bin/smoke.rb` live smoke script: asserts result shapes (not just the absence of exceptions), runs page tools with `js: false` on datacenter proxies (~32 credits per sweep), and redacts the API key from failure output.
|
|
11
|
+
|
|
12
|
+
### Fixed
|
|
13
|
+
|
|
14
|
+
- `Client#inspect` and `Configuration#inspect` no longer print the API key; it is shown as `api_key="[FILTERED]"`.
|
|
15
|
+
|
|
16
|
+
## 4.0.1 — 2026-07-17
|
|
17
|
+
|
|
18
|
+
### Changed
|
|
19
|
+
|
|
20
|
+
- Documentation: expanded README — API docs, signup/dashboard links, badges, and links to the other official clients.
|
|
21
|
+
|
|
5
22
|
## [4.0.0] - Unreleased
|
|
6
23
|
|
|
7
24
|
### Changed
|
data/README.md
CHANGED
|
@@ -1,14 +1,19 @@
|
|
|
1
1
|
# WebScraping.AI Ruby Client
|
|
2
2
|
|
|
3
|
-
Official Ruby client for the [WebScraping.AI](https://webscraping.ai) API. Provides LLM-powered web scraping with Chromium JavaScript rendering, rotating proxies, and built-in HTML parsing.
|
|
4
|
-
|
|
5
3
|
[](https://rubygems.org/gems/webscraping_ai)
|
|
4
|
+
[](https://github.com/webscraping-ai/webscraping-ai-ruby/actions/workflows/ci.yml)
|
|
5
|
+
|
|
6
|
+
Official Ruby client for the [WebScraping.AI](https://webscraping.ai) API —
|
|
7
|
+
web scraping with Chromium JavaScript rendering, rotating
|
|
8
|
+
datacenter/residential/stealth proxies, and AI-powered question answering and
|
|
9
|
+
structured field extraction on any page. See the
|
|
10
|
+
[API documentation](https://webscraping.ai/docs) for the full parameter reference.
|
|
6
11
|
|
|
7
12
|
## Installation
|
|
8
13
|
|
|
9
14
|
```ruby
|
|
10
15
|
# Gemfile
|
|
11
|
-
gem "webscraping_ai", "~> 4.
|
|
16
|
+
gem "webscraping_ai", "~> 4.1"
|
|
12
17
|
```
|
|
13
18
|
|
|
14
19
|
Or:
|
|
@@ -21,6 +26,10 @@ Requires Ruby 3.1+.
|
|
|
21
26
|
|
|
22
27
|
## Quick start
|
|
23
28
|
|
|
29
|
+
[Sign up](https://webscraping.ai/auth/sign_up) to get an API key — the free
|
|
30
|
+
trial includes 2,000 credits, no credit card required. Your key lives in the
|
|
31
|
+
[dashboard](https://webscraping.ai/dashboard).
|
|
32
|
+
|
|
24
33
|
```ruby
|
|
25
34
|
require "webscraping_ai"
|
|
26
35
|
|
|
@@ -51,6 +60,10 @@ data = client.fields(
|
|
|
51
60
|
}
|
|
52
61
|
)
|
|
53
62
|
|
|
63
|
+
# Google search results (SERP) for a query
|
|
64
|
+
results = client.serp(q: "coffee machines", gl: "us", hl: "en", page: 1)
|
|
65
|
+
results["organic_results"].first["link"]
|
|
66
|
+
|
|
54
67
|
# Check your account quota
|
|
55
68
|
info = client.account
|
|
56
69
|
# => { "remaining_api_calls" => 200_000, "resets_at" => 1_617_073_667, "remaining_concurrency" => 100 }
|
|
@@ -112,6 +125,34 @@ Endpoint-specific options:
|
|
|
112
125
|
|
|
113
126
|
Returns: `String` for HTML/text responses, `Hash`/`Array` for JSON responses.
|
|
114
127
|
|
|
128
|
+
### SERP (`#serp`)
|
|
129
|
+
|
|
130
|
+
`#serp` is query-shaped rather than URL-shaped, so none of the page-fetch options above apply.
|
|
131
|
+
It returns the parsed search results as a `Hash`. Flat 15 credits per search; failed searches are not charged.
|
|
132
|
+
|
|
133
|
+
| Option | Type | Default | Description |
|
|
134
|
+
| --- | --- | --- | --- |
|
|
135
|
+
| `q` | `String` | — | Search query (required; blank or non-String raises `ArgumentError`) |
|
|
136
|
+
| `engine` | `String` | `"google"` | Search engine; currently only `google` |
|
|
137
|
+
| `gl` | `String` | `"us"` | Two-letter country code for the search |
|
|
138
|
+
| `hl` | `String` | `"en"` | Two-letter language code for the results |
|
|
139
|
+
| `page` | `Integer` | `1` | Results page number (10 results per page). Must be an `Integer` >= 1, otherwise `ArgumentError`; the server rejects values above 100 with a 400 (not billed) |
|
|
140
|
+
|
|
141
|
+
```ruby
|
|
142
|
+
results = client.serp(q: "coffee machines", gl: "gb", page: 2)
|
|
143
|
+
results["search_information"]["organic_results_state"] # => "Results for exact spelling"
|
|
144
|
+
results["organic_results"].each do |r|
|
|
145
|
+
puts "#{r["position"]}. #{r["title"]} — #{r["link"]}"
|
|
146
|
+
end
|
|
147
|
+
results["pagination"] # => { "current" => 2, "next" => 3 }
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Response keys: `search_parameters` (`engine`, `q`, `gl`, `hl`, `page`), `search_information`
|
|
151
|
+
(`query_displayed`, `organic_results_state`, optional `showing_results_for` and `total_results`),
|
|
152
|
+
`organic_results` (`position` — 1-based within the page — `title`, `link`, `domain`, `displayed_link`,
|
|
153
|
+
optional `snippet` and `date`), optional `related_searches` (`query`), and `pagination` (`current`, optional `next`).
|
|
154
|
+
Optional keys are absent when Google does not show them.
|
|
155
|
+
|
|
115
156
|
## Error handling
|
|
116
157
|
|
|
117
158
|
All API errors inherit from `WebScrapingAI::ApiError` and expose `#status`, `#message`, `#status_code`, `#status_message`, `#body`, and `#response_body`.
|
|
@@ -148,6 +189,25 @@ bundle exec rspec
|
|
|
148
189
|
bundle exec rubocop
|
|
149
190
|
```
|
|
150
191
|
|
|
192
|
+
## Smoke testing
|
|
193
|
+
|
|
194
|
+
`bin/smoke.rb` hits every endpoint once against the live API, loading the gem from `lib/` so it tests the working tree. It is not part of the spec suite and costs ~32 credits per run: the four page calls run with `js: false` and `proxy: "datacenter"` (1 credit each), `question` and `fields` cost 6 each, and the SERP call is 15. Each case checks the result shape as well as exceptions (e.g. SERP must return organic results for the right query, `selected_multiple` must match something), and failure messages redact the API key.
|
|
195
|
+
|
|
196
|
+
```bash
|
|
197
|
+
WEBSCRAPING_AI_API_KEY=... bundle exec rake smoke
|
|
198
|
+
# or: WEBSCRAPING_AI_API_KEY=... ruby bin/smoke.rb
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
Each endpoint prints an `ok` or `FAIL` line; the script exits non-zero if any call fails.
|
|
202
|
+
|
|
203
|
+
## Links
|
|
204
|
+
|
|
205
|
+
- [WebScraping.AI](https://webscraping.ai) — features, pricing, signup
|
|
206
|
+
- [API documentation](https://webscraping.ai/docs)
|
|
207
|
+
- [Dashboard](https://webscraping.ai/dashboard) — API key, usage, request builder
|
|
208
|
+
- Other official clients: [Python](https://github.com/webscraping-ai/webscraping-ai-python) · [JavaScript](https://github.com/webscraping-ai/webscraping-ai-js) · [PHP](https://github.com/webscraping-ai/webscraping-ai-php) · [Go](https://github.com/webscraping-ai/webscraping-ai-go) · [Java](https://github.com/webscraping-ai/webscraping-ai-java) · [.NET](https://github.com/webscraping-ai/webscraping-ai-dotnet) · [CLI](https://github.com/webscraping-ai/webscraping-ai-cli) · [MCP server](https://github.com/webscraping-ai/webscraping-ai-mcp-server) · [n8n node](https://github.com/webscraping-ai/webscraping-ai-n8n)
|
|
209
|
+
- Support: [support@webscraping.ai](mailto:support@webscraping.ai)
|
|
210
|
+
|
|
151
211
|
## License
|
|
152
212
|
|
|
153
213
|
MIT — see [LICENSE](LICENSE).
|
|
@@ -8,6 +8,7 @@ module WebScrapingAI
|
|
|
8
8
|
DEVICES = %w[desktop mobile tablet].freeze
|
|
9
9
|
TEXT_FORMATS = %w[plain xml json].freeze
|
|
10
10
|
FORMATS = %w[json text].freeze
|
|
11
|
+
SERP_ENGINES = %w[google].freeze
|
|
11
12
|
|
|
12
13
|
PAGE_FETCH_OPTIONS = %i[
|
|
13
14
|
headers timeout js js_timeout wait_for proxy country
|
|
@@ -66,6 +67,23 @@ module WebScrapingAI
|
|
|
66
67
|
get("/selected-multiple", url: url, selectors: Array(selectors), **opts.slice(*PAGE_FETCH_OPTIONS))
|
|
67
68
|
end
|
|
68
69
|
|
|
70
|
+
# GET /serp — returns parsed search engine results for `q` as a Hash
|
|
71
|
+
# (search_parameters, search_information, organic_results, related_searches, pagination).
|
|
72
|
+
# Query-shaped: none of the page-fetch options apply. Flat 15 credits per search.
|
|
73
|
+
# `q` must be a non-blank String (sent as-is, untrimmed). `page`, when given, must be an
|
|
74
|
+
# Integer >= 1; the server rejects values above 100 with a 400 (not billed).
|
|
75
|
+
def serp(q:, engine: nil, gl: nil, hl: nil, page: nil)
|
|
76
|
+
raise ArgumentError, "q is required" if q.nil? || (q.is_a?(String) && q.strip.empty?)
|
|
77
|
+
raise ArgumentError, "q must be a String" unless q.is_a?(String)
|
|
78
|
+
raise ArgumentError, "page must be an Integer >= 1" unless page.nil? || (page.is_a?(Integer) && page >= 1)
|
|
79
|
+
|
|
80
|
+
get("/serp", q: q, engine: engine, gl: gl, hl: hl, page: page)
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
def inspect
|
|
84
|
+
"#<#{self.class.name} base_url=#{configuration.base_url.inspect} api_key=\"[FILTERED]\">"
|
|
85
|
+
end
|
|
86
|
+
|
|
69
87
|
# GET /account — returns Hash with remaining_api_calls, resets_at, remaining_concurrency, email.
|
|
70
88
|
def account
|
|
71
89
|
get("/account")
|
|
@@ -14,5 +14,11 @@ module WebScrapingAI
|
|
|
14
14
|
@adapter = nil
|
|
15
15
|
@user_agent = "webscraping_ai-ruby/#{WebScrapingAI::VERSION}"
|
|
16
16
|
end
|
|
17
|
+
|
|
18
|
+
def inspect
|
|
19
|
+
filtered = api_key.nil? ? "nil" : "\"[FILTERED]\""
|
|
20
|
+
"#<#{self.class.name} api_key=#{filtered} base_url=#{base_url.inspect} timeout=#{timeout.inspect} " \
|
|
21
|
+
"open_timeout=#{open_timeout.inspect} adapter=#{adapter.inspect} user_agent=#{user_agent.inspect}>"
|
|
22
|
+
end
|
|
17
23
|
end
|
|
18
24
|
end
|
metadata
CHANGED
|
@@ -1,14 +1,13 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: webscraping_ai
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 4.
|
|
4
|
+
version: 4.1.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- WebScraping.AI
|
|
8
|
-
autorequire:
|
|
9
8
|
bindir: bin
|
|
10
9
|
cert_chain: []
|
|
11
|
-
date:
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
12
11
|
dependencies:
|
|
13
12
|
- !ruby/object:Gem::Dependency
|
|
14
13
|
name: faraday
|
|
@@ -53,7 +52,6 @@ metadata:
|
|
|
53
52
|
changelog_uri: https://github.com/webscraping-ai/webscraping-ai-ruby/blob/master/CHANGELOG.md
|
|
54
53
|
documentation_uri: https://webscraping.ai/docs/api
|
|
55
54
|
rubygems_mfa_required: 'true'
|
|
56
|
-
post_install_message:
|
|
57
55
|
rdoc_options: []
|
|
58
56
|
require_paths:
|
|
59
57
|
- lib
|
|
@@ -68,8 +66,7 @@ required_rubygems_version: !ruby/object:Gem::Requirement
|
|
|
68
66
|
- !ruby/object:Gem::Version
|
|
69
67
|
version: '0'
|
|
70
68
|
requirements: []
|
|
71
|
-
rubygems_version:
|
|
72
|
-
signing_key:
|
|
69
|
+
rubygems_version: 4.0.20
|
|
73
70
|
specification_version: 4
|
|
74
71
|
summary: Ruby client for the WebScraping.AI API.
|
|
75
72
|
test_files: []
|