webscraping_ai 4.0.1 → 4.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: e312f29ac8666f33be3757599d08f062908be2a1a815973cc591d2cce44a4f4e
4
- data.tar.gz: 6cb641d229bdf754fba1ffb2f717bc40db8bcad50158d6fc960d3d6de7448480
3
+ metadata.gz: cd7230f48fdc600563c0e7fe84ba65411f6363007854a9bdaae6b2ee6bfbf17d
4
+ data.tar.gz: a815bbc4645cd9eb23aba4e9cdafff0ac1c420f3a230580e002a59f7f19b28f6
5
5
  SHA512:
6
- metadata.gz: 8d6838f0b67bf4426e2b1c6f48d5047fe4413c3d1e342e576aedc4a8df755f2fe4b591ccf747e8df955690d66ebe3ebf5b95747cc93794fa8fe0d90572050d19
7
- data.tar.gz: 4d3966a1b3e7d4bd2b4a51cb1011f991ba43fc21f760fcce3496427af68256a5d2e45587089fc95e726f400577078935f83a178abd18c864f06da9312d29908e
6
+ metadata.gz: 8266c1f95d1432a37f817e3308e1c99c1795fba17fc105c4ef703561e89f84104f0d2734fda99dded525000baecd346d00c2ac3f0b838e251807d3651d20fae6
7
+ data.tar.gz: 6d075d815694065fde5b830d577100c8c45c9943798689ecc50b9959120d0486b10ede96b77344fb0803b8094f2855cf60d72237161b07a480e31b4531c62b1c
data/CHANGELOG.md CHANGED
@@ -2,6 +2,17 @@
2
2
 
3
3
  All notable changes to this project will be documented in this file. This project follows [Semantic Versioning](https://semver.org/).
4
4
 
5
+ ## 4.1.0 — 2026-09-25
6
+
7
+ ### Added
8
+
9
+ - `Client#serp(q:, engine: nil, gl: nil, hl: nil, page: nil)` for the new `GET /serp` endpoint — parsed Google search results (organic results, related searches, pagination) as a Hash. Flat 15 credits per search. Raises `ArgumentError` when `q` is blank or not a String, or when `page` is not an Integer >= 1 (the server also rejects it with a 400, not billed; checking client-side saves the round trip). Pages are 1–100: the server rejects a `page` above 100 with a 400.
10
+ - `bin/smoke.rb` live smoke script: asserts result shapes (not just the absence of exceptions), runs page tools with `js: false` on datacenter proxies (~32 credits per sweep), and redacts the API key from failure output.
11
+
12
+ ### Fixed
13
+
14
+ - `Client#inspect` and `Configuration#inspect` no longer print the API key; it is shown as `api_key="[FILTERED]"`.
15
+
5
16
  ## 4.0.1 — 2026-07-17
6
17
 
7
18
  ### Changed
data/README.md CHANGED
@@ -13,7 +13,7 @@ structured field extraction on any page. See the
13
13
 
14
14
  ```ruby
15
15
  # Gemfile
16
- gem "webscraping_ai", "~> 4.0"
16
+ gem "webscraping_ai", "~> 4.1"
17
17
  ```
18
18
 
19
19
  Or:
@@ -60,6 +60,10 @@ data = client.fields(
60
60
  }
61
61
  )
62
62
 
63
+ # Google search results (SERP) for a query
64
+ results = client.serp(q: "coffee machines", gl: "us", hl: "en", page: 1)
65
+ results["organic_results"].first["link"]
66
+
63
67
  # Check your account quota
64
68
  info = client.account
65
69
  # => { "remaining_api_calls" => 200_000, "resets_at" => 1_617_073_667, "remaining_concurrency" => 100 }
@@ -121,6 +125,34 @@ Endpoint-specific options:
121
125
 
122
126
  Returns: `String` for HTML/text responses, `Hash`/`Array` for JSON responses.
123
127
 
128
+ ### SERP (`#serp`)
129
+
130
+ `#serp` is query-shaped rather than URL-shaped, so none of the page-fetch options above apply.
131
+ It returns the parsed search results as a `Hash`. Flat 15 credits per search; failed searches are not charged.
132
+
133
+ | Option | Type | Default | Description |
134
+ | --- | --- | --- | --- |
135
+ | `q` | `String` | — | Search query (required; blank or non-String raises `ArgumentError`) |
136
+ | `engine` | `String` | `"google"` | Search engine; currently only `google` |
137
+ | `gl` | `String` | `"us"` | Two-letter country code for the search |
138
+ | `hl` | `String` | `"en"` | Two-letter language code for the results |
139
+ | `page` | `Integer` | `1` | Results page number (10 results per page). Must be an `Integer` >= 1, otherwise `ArgumentError`; the server rejects values above 100 with a 400 (not billed) |
140
+
141
+ ```ruby
142
+ results = client.serp(q: "coffee machines", gl: "gb", page: 2)
143
+ results["search_information"]["organic_results_state"] # => "Results for exact spelling"
144
+ results["organic_results"].each do |r|
145
+ puts "#{r["position"]}. #{r["title"]} — #{r["link"]}"
146
+ end
147
+ results["pagination"] # => { "current" => 2, "next" => 3 }
148
+ ```
149
+
150
+ Response keys: `search_parameters` (`engine`, `q`, `gl`, `hl`, `page`), `search_information`
151
+ (`query_displayed`, `organic_results_state`, optional `showing_results_for` and `total_results`),
152
+ `organic_results` (`position` — 1-based within the page — `title`, `link`, `domain`, `displayed_link`,
153
+ optional `snippet` and `date`), optional `related_searches` (`query`), and `pagination` (`current`, optional `next`).
154
+ Optional keys are absent when Google does not show them.
155
+
124
156
  ## Error handling
125
157
 
126
158
  All API errors inherit from `WebScrapingAI::ApiError` and expose `#status`, `#message`, `#status_code`, `#status_message`, `#body`, and `#response_body`.
@@ -157,6 +189,17 @@ bundle exec rspec
157
189
  bundle exec rubocop
158
190
  ```
159
191
 
192
+ ## Smoke testing
193
+
194
+ `bin/smoke.rb` hits every endpoint once against the live API, loading the gem from `lib/` so it tests the working tree. It is not part of the spec suite and costs ~32 credits per run: the four page calls run with `js: false` and `proxy: "datacenter"` (1 credit each), `question` and `fields` cost 6 each, and the SERP call is 15. Each case checks the result shape as well as exceptions (e.g. SERP must return organic results for the right query, `selected_multiple` must match something), and failure messages redact the API key.
195
+
196
+ ```bash
197
+ WEBSCRAPING_AI_API_KEY=... bundle exec rake smoke
198
+ # or: WEBSCRAPING_AI_API_KEY=... ruby bin/smoke.rb
199
+ ```
200
+
201
+ Each endpoint prints an `ok` or `FAIL` line; the script exits non-zero if any call fails.
202
+
160
203
  ## Links
161
204
 
162
205
  - [WebScraping.AI](https://webscraping.ai) — features, pricing, signup
@@ -8,6 +8,7 @@ module WebScrapingAI
8
8
  DEVICES = %w[desktop mobile tablet].freeze
9
9
  TEXT_FORMATS = %w[plain xml json].freeze
10
10
  FORMATS = %w[json text].freeze
11
+ SERP_ENGINES = %w[google].freeze
11
12
 
12
13
  PAGE_FETCH_OPTIONS = %i[
13
14
  headers timeout js js_timeout wait_for proxy country
@@ -66,6 +67,23 @@ module WebScrapingAI
66
67
  get("/selected-multiple", url: url, selectors: Array(selectors), **opts.slice(*PAGE_FETCH_OPTIONS))
67
68
  end
68
69
 
70
+ # GET /serp — returns parsed search engine results for `q` as a Hash
71
+ # (search_parameters, search_information, organic_results, related_searches, pagination).
72
+ # Query-shaped: none of the page-fetch options apply. Flat 15 credits per search.
73
+ # `q` must be a non-blank String (sent as-is, untrimmed). `page`, when given, must be an
74
+ # Integer >= 1; the server rejects values above 100 with a 400 (not billed).
75
+ def serp(q:, engine: nil, gl: nil, hl: nil, page: nil)
76
+ raise ArgumentError, "q is required" if q.nil? || (q.is_a?(String) && q.strip.empty?)
77
+ raise ArgumentError, "q must be a String" unless q.is_a?(String)
78
+ raise ArgumentError, "page must be an Integer >= 1" unless page.nil? || (page.is_a?(Integer) && page >= 1)
79
+
80
+ get("/serp", q: q, engine: engine, gl: gl, hl: hl, page: page)
81
+ end
82
+
83
+ def inspect
84
+ "#<#{self.class.name} base_url=#{configuration.base_url.inspect} api_key=\"[FILTERED]\">"
85
+ end
86
+
69
87
  # GET /account — returns Hash with remaining_api_calls, resets_at, remaining_concurrency, email.
70
88
  def account
71
89
  get("/account")
@@ -14,5 +14,11 @@ module WebScrapingAI
14
14
  @adapter = nil
15
15
  @user_agent = "webscraping_ai-ruby/#{WebScrapingAI::VERSION}"
16
16
  end
17
+
18
+ def inspect
19
+ filtered = api_key.nil? ? "nil" : "\"[FILTERED]\""
20
+ "#<#{self.class.name} api_key=#{filtered} base_url=#{base_url.inspect} timeout=#{timeout.inspect} " \
21
+ "open_timeout=#{open_timeout.inspect} adapter=#{adapter.inspect} user_agent=#{user_agent.inspect}>"
22
+ end
17
23
  end
18
24
  end
@@ -1,3 +1,3 @@
1
1
  module WebScrapingAI
2
- VERSION = "4.0.1".freeze
2
+ VERSION = "4.1.0".freeze
3
3
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: webscraping_ai
3
3
  version: !ruby/object:Gem::Version
4
- version: 4.0.1
4
+ version: 4.1.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - WebScraping.AI
@@ -66,7 +66,7 @@ required_rubygems_version: !ruby/object:Gem::Requirement
66
66
  - !ruby/object:Gem::Version
67
67
  version: '0'
68
68
  requirements: []
69
- rubygems_version: 4.0.16
69
+ rubygems_version: 4.0.20
70
70
  specification_version: 4
71
71
  summary: Ruby client for the WebScraping.AI API.
72
72
  test_files: []