adscrawl 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: 53190cac137e9d06f3175969488c05c975dc10ce41bae70c0682110242666b05
4
+ data.tar.gz: dbe222b0ff615996d4dbcb6071d9f4f9289d83467d850632c8778ba308be5c3b
5
+ SHA512:
6
+ metadata.gz: edc122d2a0d30101806275101c8def52f424a62348428405de2a91973559952cc032f84bf4088b744e493d0add8a400d3e580638317b069fcd041bf67bbcecc2
7
+ data.tar.gz: ae6f562c5cc5249c2f9b31ed7a2a424ffc93e01b1ce79f38e7b4b7272b9d4b2f7ba412cc20a268cb532fe299c949df1491fb647c84cca7a9f7ccb0d55e98f132
data/CHANGELOG.md ADDED
@@ -0,0 +1,7 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0 - 2026-09-28
4
+
5
+ - Initial Ruby SDK for rendered HTML, Markdown, articles, and PNG screenshots.
6
+ - SPA extraction, CDP session, and persistent cloud browser services.
7
+ - Request deadlines, response validation, credential-redacted errors, and runnable examples.
data/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 AdsCrawl
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,208 @@
1
+ <p align="center">
2
+ <a href="https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby">
3
+ <img src="assets/adscrawl-logo.svg" alt="AdsCrawl" width="360" />
4
+ </a>
5
+ </p>
6
+
7
+ <p align="center"><strong>Real browsers. Structured extraction. Screenshots and automation.</strong></p>
8
+
9
+ <h1 align="center">Ruby SDK</h1>
10
+
11
+ Turn a URL into rendered HTML, readable Markdown, structured data, or a PNG screenshot. Open remote CDP sessions and manage persistent cloud browsers with the same API as the TypeScript, Python, Go, and PHP SDKs.
12
+
13
+ [Website](https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [API documentation](https://www.adscrawl.net/docs/) · [Get an API key](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [中文](./README.zh-CN.md)
14
+
15
+ - Ruby 3.1+ with no runtime dependencies outside the standard library.
16
+ - Explicit HTTP deadlines, checked response shapes, and credential-redacted errors.
17
+ - Hash and keyword arguments use the HTTP API's camelCase field names.
18
+ - No automatic retries of metered requests or browser creation.
19
+
20
+ ## Install
21
+
22
+ ```bash
23
+ gem install adscrawl
24
+ ```
25
+
26
+ Or add it to your bundle:
27
+
28
+ ```bash
29
+ bundle add adscrawl
30
+ ```
31
+
32
+ ## Your first request
33
+
34
+ [Create an account](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby), create an API key, and set `ADSCRAWL_API_KEY` in your server environment.
35
+
36
+ ```ruby
37
+ require "adscrawl"
38
+
39
+ client = AdsCrawl::Client.new
40
+ markdown = client.markdown(
41
+ url: "https://www.adscrawl.net",
42
+ waitUntil: "domcontentloaded"
43
+ )
44
+ puts markdown
45
+ ```
46
+
47
+ You can also use `AdsCrawl::Client.new(api_key: "...")`. Keep API keys in server-side code. See the complete runnable examples for [Markdown](./examples/markdown.rb), [screenshots](./examples/screenshot.rb), [structured extraction](./examples/extract.rb), [fingerprints](./examples/fingerprint.rb), [CDP](./examples/cdp.rb), and [cloud browsers](./examples/cloud_browser.rb).
48
+
49
+ ## Rendered content and screenshots
50
+
51
+ ```ruby
52
+ require "adscrawl"
53
+
54
+ client = AdsCrawl::Client.new
55
+ html = client.html(url: "https://www.adscrawl.net")
56
+ article = client.article(url: "https://www.adscrawl.net")
57
+ puts [article["title"], article["textContent"]]
58
+
59
+ png = client.screenshot(
60
+ url: "https://www.adscrawl.net",
61
+ viewport: {width: 1440, height: 900},
62
+ fullPage: true,
63
+ waitUntil: "load"
64
+ )
65
+ File.binwrite("page.png", png)
66
+ ```
67
+
68
+ `html` returns HTML, `markdown` returns Markdown, `article` returns a Hash, and `screenshot` returns binary PNG data. Page options include viewport, locale, cookies, custom `proxy`, managed `countryCode`, fingerprint, server-side `timeoutMs`, and navigation `waitUntil`. A custom proxy and managed country cannot be combined.
69
+
70
+ ## Proxy and fingerprint: BrowserScan screenshot
71
+
72
+ AdsCrawl uses real browsers with configurable routing and browser fingerprints. This complete example uses managed `GLOBAL` routing unless proxy environment variables are present, opens [BrowserScan](https://www.browserscan.net/), and saves the returned screenshot. The same browser workflow can access and render [Pixelscan](https://pixelscan.net/) and [IPhey](https://iphey.com/); inspect the returned page for its current result.
73
+
74
+ ```ruby
75
+ require "adscrawl"
76
+
77
+ client = AdsCrawl::Client.new
78
+ server = ENV["ADSCRAWL_PROXY_SERVER"]
79
+ username = ENV["ADSCRAWL_PROXY_USERNAME"]
80
+ password = ENV["ADSCRAWL_PROXY_PASSWORD"]
81
+ raise "Set both proxy username and password, or neither." if username.nil? != password.nil?
82
+ raise "Set ADSCRAWL_PROXY_SERVER with proxy credentials." if server.nil? && (username || password)
83
+
84
+ routing = {countryCode: "GLOBAL"}
85
+ if server
86
+ proxy = {server: server}
87
+ proxy.merge!(username: username, password: password) if username && password
88
+ routing = {proxy: proxy}
89
+ end
90
+
91
+ png = client.screenshot(
92
+ **routing,
93
+ url: "https://www.browserscan.net/",
94
+ viewport: {width: 1440, height: 900},
95
+ fullPage: true,
96
+ waitUntil: "networkidle",
97
+ timeoutMs: 60_000,
98
+ userAgentMode: "random",
99
+ userAgentOs: "windows",
100
+ fingerprint: {
101
+ webRtc: "forward", webGl: "random", webGpu: "random",
102
+ webGlImage: "random", canvas: "random", audioContext: "random",
103
+ clientRects: "random", speechVoices: "random", fonts: "random",
104
+ hardware: "random", doNotTrack: "random"
105
+ },
106
+ timeout_ms: 75_000
107
+ )
108
+ File.binwrite("browserscan.png", png)
109
+ ```
110
+
111
+ Run the included version with `ruby examples/fingerprint.rb`. Pass a detection URL and output path as optional arguments. The example never combines a custom proxy with managed routing.
112
+
113
+ ## Structured extraction
114
+
115
+ ```ruby
116
+ require "adscrawl"
117
+
118
+ client = AdsCrawl::Client.new
119
+ puts client.spa.templates["templates"]
120
+
121
+ result = client.spa.extract(
122
+ url: "https://www.adscrawl.net",
123
+ fields: {
124
+ title: {source: "dom", selector: "h1", value: "text", required: true}
125
+ }
126
+ )
127
+ puts result.dig("data", "title")
128
+
129
+ inspection = client.spa.inspect(url: "https://www.adscrawl.net")
130
+ puts inspection["candidates"]
131
+ ```
132
+
133
+ ## Remote CDP browsers
134
+
135
+ ```ruby
136
+ require "adscrawl"
137
+
138
+ client = AdsCrawl::Client.new
139
+ session = client.cdp.create(
140
+ idleTimeoutMs: 600_000,
141
+ maxSessionMs: 3_600_000,
142
+ browserSettings: {viewport: {width: 1440, height: 900}}
143
+ )
144
+ begin
145
+ version = client.cdp.get_version(session)
146
+ puts version["Browser"]
147
+ # Connect a CDP-compatible Ruby library to session["cdpBaseUrl"].
148
+ ensure
149
+ client.cdp.close(session["sessionId"])
150
+ end
151
+ ```
152
+
153
+ `client.cdp.list` returns active sessions. `get_version` validates the session token URL and deliberately omits the API key. CDP connection URLs contain secrets; do not log them.
154
+
155
+ ## Persistent cloud browsers
156
+
157
+ ```ruby
158
+ require "adscrawl"
159
+
160
+ client = AdsCrawl::Client.new
161
+ launched = client.cloud_browsers.launch(
162
+ proxy: {server: ENV.fetch("ADSCRAWL_PROXY_SERVER")},
163
+ tabs: ["https://www.adscrawl.net"]
164
+ )
165
+ browser_id = launched["id"]
166
+ begin
167
+ profile = client.cloud_browsers.get(browser_id)
168
+ puts profile.dig("runtime", "status")
169
+ ensure
170
+ client.cloud_browsers.stop(browser_id)
171
+ end
172
+ ```
173
+
174
+ API-key starts and launches require an explicit top-level custom proxy. A `stopping` response means shutdown is pending; poll `get` until `stopped`. A failed launch may expose a cleanup profile ID as `APIError#resource_id`.
175
+
176
+ ## Configuration and errors
177
+
178
+ ```ruby
179
+ require "adscrawl"
180
+
181
+ client = AdsCrawl::Client.new(base_url: "https://api.adscrawl.net", timeout_ms: 90_000)
182
+ begin
183
+ puts client.markdown(
184
+ {url: "https://www.adscrawl.net", timeoutMs: 60_000},
185
+ timeout_ms: 75_000
186
+ )
187
+ rescue AdsCrawl::APIError => error
188
+ puts [error.status, error.api_code, error.trace_id]
189
+ rescue AdsCrawl::TimeoutError
190
+ puts "Inspect remote sessions before retrying."
191
+ end
192
+ ```
193
+
194
+ The API key defaults to `ADSCRAWL_API_KEY`; the base URL defaults to `ADSCRAWL_BASE_URL`, then `ADSCRAWL_API_URL`, then `https://api.adscrawl.net`. Errors include `APIError`, `TimeoutError`, `ConnectionError`, and `ResponseError`. API response bodies are credential-redacted. The HTTP deadline includes response body reading. Requests are not automatically retried.
195
+
196
+ ## Develop and release
197
+
198
+ ```bash
199
+ bundle install
200
+ bundle exec rake test
201
+ gem build adscrawl.gemspec
202
+ ```
203
+
204
+ Tests use local and fake responses and consume no AdsCrawl credits. Releases use RubyGems Trusted Publishing with no long-lived registry token; see [RELEASING.md](./RELEASING.md).
205
+
206
+ ## License
207
+
208
+ MIT
data/README.zh-CN.md ADDED
@@ -0,0 +1,208 @@
1
+ <p align="center">
2
+ <a href="https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby">
3
+ <img src="assets/adscrawl-logo.svg" alt="AdsCrawl" width="360" />
4
+ </a>
5
+ </p>
6
+
7
+ <p align="center"><strong>真实浏览器、结构化提取、截图与自动化。</strong></p>
8
+
9
+ <h1 align="center">Ruby SDK</h1>
10
+
11
+ 把 URL 转换为浏览器渲染后的 HTML、Markdown、结构化数据或 PNG 截图;也可以创建远程 CDP 会话和管理持久云浏览器。
12
+
13
+ [官网](https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [API 文档](https://www.adscrawl.net/docs/) · [获取 API Key](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [English](./README.md)
14
+
15
+ - 支持 Ruby 3.1+,除标准库外无运行时依赖。
16
+ - 支持 HTTP 超时、响应结构校验和错误信息凭据脱敏。
17
+ - Hash 和关键字参数与 HTTP API 使用相同的 camelCase 字段。
18
+ - 对计费请求和浏览器创建请求不会自动重试。
19
+
20
+ ## 安装
21
+
22
+ ```bash
23
+ gem install adscrawl
24
+ ```
25
+
26
+ 也可以添加到 Bundle:
27
+
28
+ ```bash
29
+ bundle add adscrawl
30
+ ```
31
+
32
+ ## 第一个请求
33
+
34
+ [注册账号](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby),创建 API Key,并在服务端环境变量中设置 `ADSCRAWL_API_KEY`。
35
+
36
+ ```ruby
37
+ require "adscrawl"
38
+
39
+ client = AdsCrawl::Client.new
40
+ markdown = client.markdown(
41
+ url: "https://www.adscrawl.net",
42
+ waitUntil: "domcontentloaded"
43
+ )
44
+ puts markdown
45
+ ```
46
+
47
+ 也可以使用 `AdsCrawl::Client.new(api_key: "...")`。API Key 应只保存在服务端。仓库包含完整可运行的 [Markdown](./examples/markdown.rb)、[截图](./examples/screenshot.rb)、[结构化提取](./examples/extract.rb)、[指纹检测](./examples/fingerprint.rb)、[CDP](./examples/cdp.rb) 和 [云浏览器](./examples/cloud_browser.rb) 示例。
48
+
49
+ ## 渲染内容与截图
50
+
51
+ ```ruby
52
+ require "adscrawl"
53
+
54
+ client = AdsCrawl::Client.new
55
+ html = client.html(url: "https://www.adscrawl.net")
56
+ article = client.article(url: "https://www.adscrawl.net")
57
+ puts [article["title"], article["textContent"]]
58
+
59
+ png = client.screenshot(
60
+ url: "https://www.adscrawl.net",
61
+ viewport: {width: 1440, height: 900},
62
+ fullPage: true,
63
+ waitUntil: "load"
64
+ )
65
+ File.binwrite("page.png", png)
66
+ ```
67
+
68
+ `html` 返回 HTML,`markdown` 返回 Markdown,`article` 返回 Hash,`screenshot` 返回 PNG 二进制数据。页面参数支持视口、语言、Cookie、自定义 `proxy`、托管 `countryCode`、浏览器指纹、服务端 `timeoutMs` 与导航 `waitUntil`。自定义代理和托管国家线路不能同时使用。
69
+
70
+ ## 代理与指纹:BrowserScan 截图
71
+
72
+ AdsCrawl 使用真实浏览器,并支持配置线路和浏览器指纹。以下示例默认使用 `GLOBAL` 托管线路;设置代理环境变量后则使用自定义代理。示例访问 [BrowserScan](https://www.browserscan.net/) 并保存返回截图。同一套浏览器流程也可访问和渲染 [Pixelscan](https://pixelscan.net/) 与 [IPhey](https://iphey.com/),可从返回页面查看网站当前的检测结果。
73
+
74
+ ```ruby
75
+ require "adscrawl"
76
+
77
+ client = AdsCrawl::Client.new
78
+ server = ENV["ADSCRAWL_PROXY_SERVER"]
79
+ username = ENV["ADSCRAWL_PROXY_USERNAME"]
80
+ password = ENV["ADSCRAWL_PROXY_PASSWORD"]
81
+ raise "代理用户名和密码必须同时设置。" if username.nil? != password.nil?
82
+ raise "使用代理凭据时必须设置 ADSCRAWL_PROXY_SERVER。" if server.nil? && (username || password)
83
+
84
+ routing = {countryCode: "GLOBAL"}
85
+ if server
86
+ proxy = {server: server}
87
+ proxy.merge!(username: username, password: password) if username && password
88
+ routing = {proxy: proxy}
89
+ end
90
+
91
+ png = client.screenshot(
92
+ **routing,
93
+ url: "https://www.browserscan.net/",
94
+ viewport: {width: 1440, height: 900},
95
+ fullPage: true,
96
+ waitUntil: "networkidle",
97
+ timeoutMs: 60_000,
98
+ userAgentMode: "random",
99
+ userAgentOs: "windows",
100
+ fingerprint: {
101
+ webRtc: "forward", webGl: "random", webGpu: "random",
102
+ webGlImage: "random", canvas: "random", audioContext: "random",
103
+ clientRects: "random", speechVoices: "random", fonts: "random",
104
+ hardware: "random", doNotTrack: "random"
105
+ },
106
+ timeout_ms: 75_000
107
+ )
108
+ File.binwrite("browserscan.png", png)
109
+ ```
110
+
111
+ 可运行 `ruby examples/fingerprint.rb`,也可以把检测网站 URL 和输出路径作为参数传入。示例不会把自定义代理和托管线路组合使用。
112
+
113
+ ## 结构化提取
114
+
115
+ ```ruby
116
+ require "adscrawl"
117
+
118
+ client = AdsCrawl::Client.new
119
+ puts client.spa.templates["templates"]
120
+
121
+ result = client.spa.extract(
122
+ url: "https://www.adscrawl.net",
123
+ fields: {
124
+ title: {source: "dom", selector: "h1", value: "text", required: true}
125
+ }
126
+ )
127
+ puts result.dig("data", "title")
128
+
129
+ inspection = client.spa.inspect(url: "https://www.adscrawl.net")
130
+ puts inspection["candidates"]
131
+ ```
132
+
133
+ ## 远程 CDP 浏览器
134
+
135
+ ```ruby
136
+ require "adscrawl"
137
+
138
+ client = AdsCrawl::Client.new
139
+ session = client.cdp.create(
140
+ idleTimeoutMs: 600_000,
141
+ maxSessionMs: 3_600_000,
142
+ browserSettings: {viewport: {width: 1440, height: 900}}
143
+ )
144
+ begin
145
+ version = client.cdp.get_version(session)
146
+ puts version["Browser"]
147
+ # 使用兼容 CDP 的 Ruby 库连接 session["cdpBaseUrl"]。
148
+ ensure
149
+ client.cdp.close(session["sessionId"])
150
+ end
151
+ ```
152
+
153
+ `client.cdp.list` 返回活动会话。`get_version` 会校验会话 Token URL,并且不会把 API Key 转发到数据端点。CDP 连接 URL 含有密钥,请勿写入日志。
154
+
155
+ ## 持久云浏览器
156
+
157
+ ```ruby
158
+ require "adscrawl"
159
+
160
+ client = AdsCrawl::Client.new
161
+ launched = client.cloud_browsers.launch(
162
+ proxy: {server: ENV.fetch("ADSCRAWL_PROXY_SERVER")},
163
+ tabs: ["https://www.adscrawl.net"]
164
+ )
165
+ browser_id = launched["id"]
166
+ begin
167
+ profile = client.cloud_browsers.get(browser_id)
168
+ puts profile.dig("runtime", "status")
169
+ ensure
170
+ client.cloud_browsers.stop(browser_id)
171
+ end
172
+ ```
173
+
174
+ 使用 API Key 启动或快速创建云浏览器时,每次都必须传顶层自定义代理。返回 `stopping` 表示仍在关闭,需要轮询 `get` 直到状态为 `stopped`。快速创建失败时,`APIError#resource_id` 可能包含需要清理的浏览器 ID。
175
+
176
+ ## 配置与错误处理
177
+
178
+ ```ruby
179
+ require "adscrawl"
180
+
181
+ client = AdsCrawl::Client.new(base_url: "https://api.adscrawl.net", timeout_ms: 90_000)
182
+ begin
183
+ puts client.markdown(
184
+ {url: "https://www.adscrawl.net", timeoutMs: 60_000},
185
+ timeout_ms: 75_000
186
+ )
187
+ rescue AdsCrawl::APIError => error
188
+ puts [error.status, error.api_code, error.trace_id]
189
+ rescue AdsCrawl::TimeoutError
190
+ puts "重试前请先检查远程会话。"
191
+ end
192
+ ```
193
+
194
+ API Key 默认读取 `ADSCRAWL_API_KEY`。基础 URL 依次读取 `ADSCRAWL_BASE_URL`、`ADSCRAWL_API_URL`,最后使用 `https://api.adscrawl.net`。错误类型包括 `APIError`、`TimeoutError`、`ConnectionError` 和 `ResponseError`。API 错误响应会进行凭据脱敏。HTTP 超时包含读取响应体的时间。SDK 不会自动重试请求。
195
+
196
+ ## 开发与发布
197
+
198
+ ```bash
199
+ bundle install
200
+ bundle exec rake test
201
+ gem build adscrawl.gemspec
202
+ ```
203
+
204
+ 测试只使用本机和模拟响应,不消耗 AdsCrawl 额度。发布使用 RubyGems Trusted Publishing,不保存长期注册表 Token,详见 [RELEASING.md](./RELEASING.md)。
205
+
206
+ ## 许可证
207
+
208
+ MIT
data/RELEASING.md ADDED
@@ -0,0 +1,15 @@
1
+ # Releasing
2
+
3
+ The release workflow uses RubyGems Trusted Publishing, so the repository stores no long-lived RubyGems API key.
4
+
5
+ For the first release, create a pending trusted publisher in the RubyGems profile with:
6
+
7
+ - Gem name: `adscrawl`
8
+ - Repository owner: `AdsCrawl`
9
+ - Repository name: `adscrawl-ruby`
10
+ - Workflow filename: `release.yml`
11
+ - Environment: `release`
12
+
13
+ Then run `bundle exec rake test`, update `CHANGELOG.md`, and push a semantic version tag matching `AdsCrawl::VERSION`, such as `v0.1.0`. The workflow tests, builds, and publishes the Gem with a short-lived OIDC credential.
14
+
15
+ Published Gem versions cannot be reused. If a release needs a correction, increment the version and publish a new tag.
@@ -0,0 +1,23 @@
1
+ <?xml version="1.0" encoding="UTF-8"?>
2
+ <svg id="_图层_2" data-name="图层 2" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1477.22 269.56">
3
+ <g id="_图层_1-2" data-name="图层 1">
4
+ <g>
5
+ <g>
6
+ <polygon points="126.17 164.02 32.77 238.88 40.36 171.56 43.32 141.82 17.48 134.9 4.69 261.79 33.69 269.56 156.33 172.1 126.17 164.02"/>
7
+ <polygon points="143.39 105.54 236.79 30.67 229.2 98 226.24 127.74 252.08 134.66 264.87 7.77 235.88 0 113.24 97.46 143.39 105.54"/>
8
+ <polygon points="105.54 126.16 30.67 32.77 98 40.36 127.74 43.32 134.66 17.48 7.77 4.69 0 33.68 97.46 156.32 105.54 126.16"/>
9
+ <polygon points="172.1 113.23 164.02 143.39 238.89 236.78 171.56 229.19 141.82 226.24 134.9 252.07 261.79 264.86 269.56 235.87 172.1 113.23"/>
10
+ </g>
11
+ <g>
12
+ <path d="M664.16,112.42c-3.12-3-6.7-5.54-10.75-7.59-6.84-3.46-14.54-5.19-23.1-5.19-11.2,0-21.12,2.72-29.76,8.15-8.65,5.43-15.48,12.84-20.5,22.23-5.02,9.39-7.53,20.01-7.53,31.86s2.51,22.23,7.53,31.62c5.02,9.39,11.89,16.8,20.63,22.23,8.72,5.43,18.61,8.15,29.64,8.15,8.56,0,16.3-1.81,23.22-5.43,4.13-2.16,7.74-4.8,10.87-7.9v10.86h32.11V48.32h-32.36v64.1ZM651.68,189.93c-4.53,2.72-9.84,4.08-15.93,4.08-5.77,0-10.95-1.36-15.56-4.08-4.61-2.72-8.19-6.5-10.74-11.36-2.55-4.86-3.83-10.5-3.83-16.92s1.27-11.77,3.83-16.55c2.55-4.77,6.09-8.56,10.62-11.36,4.53-2.8,9.84-4.2,15.93-4.2s11.15,1.36,15.68,4.08c4.53,2.72,8.07,6.51,10.62,11.36,2.55,4.86,3.83,10.42,3.83,16.67s-1.28,12.07-3.83,16.92c-2.56,4.86-6.1,8.65-10.62,11.36Z"/>
13
+ <path d="M797.68,156.5c-4.59-2.7-9.41-4.78-14.46-6.24-5.06-1.46-9.9-2.76-14.54-3.9-4.64-1.14-8.41-2.5-11.33-4.09-2.92-1.58-4.4-3.86-4.47-6.82-.05-2.63,1.18-4.72,3.7-6.25,2.52-1.53,6.17-2.35,10.94-2.45,5.27-.11,10.19.78,14.76,2.66,4.57,1.88,8.71,5.01,12.43,9.38l19.1-19.92c-5.41-6.8-12.15-11.85-20.21-15.14-8.06-3.29-17.02-4.83-26.9-4.63-9.38.2-17.5,1.93-24.35,5.2-6.85,3.27-12.07,7.75-15.66,13.42-3.59,5.68-5.3,12.3-5.14,19.87.15,7.25,1.71,13.1,4.69,17.57,2.97,4.47,6.79,7.93,11.46,10.38,4.66,2.46,9.52,4.41,14.58,5.87,5.05,1.46,9.9,2.76,14.53,3.9,4.63,1.14,8.41,2.58,11.33,4.33,2.92,1.75,4.41,4.27,4.48,7.56.06,2.96-1.25,5.21-3.93,6.75-2.69,1.54-6.66,2.36-11.93,2.47-6.59.14-12.62-.89-18.1-3.08-5.48-2.19-10.33-5.55-14.54-10.07l-18.85,19.91c4.05,4.53,8.82,8.34,14.32,11.44,5.5,3.1,11.52,5.48,18.06,7.16,6.54,1.67,13.18,2.44,19.94,2.3,14.48-.3,25.86-3.96,34.11-10.97,8.25-7.01,12.26-16.35,12.01-28.05-.15-7.24-1.72-13.14-4.69-17.69-2.98-4.55-6.76-8.17-11.34-10.88Z"/>
14
+ <polygon points="562.48 221.37 534.34 48.47 494.4 48.47 366.33 221.37 406.26 221.37 506.91 84.28 519 164.51 484.87 164.51 463.99 194.1 523.63 194.1 528.11 221.37 562.48 221.37"/>
15
+ <path d="M965.98,169.68c-10.34,12.89-26.19,21.17-43.96,21.17-31.08,0-56.37-25.29-56.37-56.37s25.29-56.37,56.37-56.37c17.77,0,33.62,8.28,43.96,21.17l23.48-23.48c-16.39-18.82-40.52-30.74-67.44-30.74-49.38,0-89.42,40.03-89.42,89.42s40.03,89.42,89.42,89.42c26.92,0,51.04-11.91,67.44-30.73l-23.48-23.48Z"/>
16
+ <path d="M1084.09,102.49c-4.61-1.89-9.8-2.84-15.56-2.84-13.34,0-23.55,4.24-30.63,12.72-.09.1-.16.22-.25.33v-10.58h-32.36v119.3h32.36v-65.95c0-8.89,2.18-15.52,6.55-19.88,4.36-4.36,10-6.55,16.92-6.55,3.29,0,6.21.49,8.77,1.48,2.55.99,4.73,2.47,6.55,4.45l20.25-23.22c-3.79-4.28-7.99-7.37-12.6-9.26Z"/>
17
+ <path d="M1191.19,113.5c-3.28-3.48-7.14-6.38-11.61-8.67-6.75-3.46-14.41-5.19-22.97-5.19-10.87,0-20.67,2.72-29.39,8.15-8.73,5.43-15.56,12.84-20.5,22.23-4.94,9.39-7.41,20.01-7.41,31.86s2.47,22.23,7.41,31.62c4.94,9.39,11.77,16.8,20.5,22.23,8.73,5.43,18.53,8.15,29.39,8.15,8.56,0,16.22-1.77,22.97-5.31,4.47-2.34,8.33-5.26,11.61-8.72v11.56h32.11v-119.3h-32.11v11.39ZM1184.52,184.99c-5.6,6.01-12.93,9.02-21.98,9.02-5.93,0-11.16-1.36-15.68-4.08-4.53-2.72-8.07-6.5-10.62-11.36-2.55-4.86-3.83-10.5-3.83-16.92s1.27-11.81,3.83-16.67c2.55-4.86,6.09-8.65,10.62-11.36,4.53-2.72,9.76-4.08,15.68-4.08s11.4,1.36,15.93,4.08c4.53,2.72,8.07,6.51,10.62,11.36,2.55,4.86,3.83,10.42,3.83,16.67,0,9.55-2.8,17.33-8.4,23.34Z"/>
18
+ <polygon points="1370.67 174.41 1345.82 102.12 1322.35 102.12 1297.5 174.41 1273.2 102.12 1240.59 102.12 1285.3 221.42 1308.52 221.42 1334.08 151.43 1359.65 221.42 1383.11 221.42 1427.58 102.12 1394.97 102.12 1370.67 174.41"/>
19
+ <rect x="1444.87" y="48.62" width="32.36" height="172.81"/>
20
+ </g>
21
+ </g>
22
+ </g>
23
+ </svg>
data/examples/cdp.rb ADDED
@@ -0,0 +1,17 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "adscrawl"
4
+
5
+ client = AdsCrawl::Client.new
6
+ session = client.cdp.create(
7
+ idleTimeoutMs: 600_000,
8
+ maxSessionMs: 3_600_000,
9
+ browserSettings: {viewport: {width: 1440, height: 900}}
10
+ )
11
+ begin
12
+ version = client.cdp.get_version(session)
13
+ puts version["Browser"] || "Unknown browser"
14
+ puts "Connect a CDP-compatible Ruby library to session['cdpBaseUrl']."
15
+ ensure
16
+ client.cdp.close(session["sessionId"])
17
+ end
@@ -0,0 +1,17 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "adscrawl"
4
+
5
+ proxy_server = ENV.fetch("ADSCRAWL_PROXY_SERVER")
6
+ client = AdsCrawl::Client.new
7
+ launched = client.cloud_browsers.launch(
8
+ proxy: {server: proxy_server},
9
+ tabs: ["https://www.adscrawl.net"]
10
+ )
11
+ browser_id = launched["id"]
12
+ begin
13
+ profile = client.cloud_browsers.get(browser_id)
14
+ puts "#{profile['id']} #{profile.dig('runtime', 'status')}"
15
+ ensure
16
+ client.cloud_browsers.stop(browser_id)
17
+ end
@@ -0,0 +1,16 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "json"
4
+ require "adscrawl"
5
+
6
+ client = AdsCrawl::Client.new
7
+ result = client.spa.extract(
8
+ url: "https://www.adscrawl.net",
9
+ fields: {
10
+ title: {source: "dom", selector: "h1", value: "text", required: true}
11
+ }
12
+ )
13
+ puts JSON.pretty_generate(
14
+ title: result.dig("data", "title"),
15
+ missingFields: result["missingFields"]
16
+ )
@@ -0,0 +1,39 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "adscrawl"
4
+
5
+ client = AdsCrawl::Client.new
6
+ server = ENV["ADSCRAWL_PROXY_SERVER"]
7
+ username = ENV["ADSCRAWL_PROXY_USERNAME"]
8
+ password = ENV["ADSCRAWL_PROXY_PASSWORD"]
9
+ raise "Set both proxy username and password, or neither." if username.nil? != password.nil?
10
+ raise "Set ADSCRAWL_PROXY_SERVER with proxy credentials." if server.nil? && (username || password)
11
+
12
+ routing = {countryCode: "GLOBAL"}
13
+ if server
14
+ proxy = {server: server}
15
+ proxy.merge!(username: username, password: password) if username && password
16
+ routing = {proxy: proxy}
17
+ end
18
+
19
+ url = ARGV[0] || "https://www.browserscan.net/"
20
+ output = ARGV[1] || "browserscan.png"
21
+ png = client.screenshot(
22
+ **routing,
23
+ url: url,
24
+ viewport: {width: 1440, height: 900},
25
+ fullPage: true,
26
+ waitUntil: "networkidle",
27
+ timeoutMs: 60_000,
28
+ userAgentMode: "random",
29
+ userAgentOs: "windows",
30
+ fingerprint: {
31
+ webRtc: "forward", webGl: "random", webGpu: "random",
32
+ webGlImage: "random", canvas: "random", audioContext: "random",
33
+ clientRects: "random", speechVoices: "random", fonts: "random",
34
+ hardware: "random", doNotTrack: "random"
35
+ },
36
+ timeout_ms: 75_000
37
+ )
38
+ File.binwrite(output, png)
39
+ puts "Saved fingerprint-check screenshot to #{output}"
@@ -0,0 +1,7 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "adscrawl"
4
+
5
+ client = AdsCrawl::Client.new
6
+ url = ARGV[0] || "https://www.adscrawl.net"
7
+ puts client.markdown(url: url, waitUntil: "domcontentloaded")
@@ -0,0 +1,10 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "adscrawl"
4
+
5
+ client = AdsCrawl::Client.new
6
+ url = ARGV[0] || "https://www.adscrawl.net"
7
+ output = ARGV[1] || "page.png"
8
+ png = client.screenshot(url: url, fullPage: true)
9
+ File.binwrite(output, png)
10
+ puts "Saved screenshot to #{output}"
@@ -0,0 +1,342 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "json"
4
+ require "net/http"
5
+ require "openssl"
6
+ require "socket"
7
+ require "timeout"
8
+ require "uri"
9
+
10
+ require_relative "errors"
11
+ require_relative "http_response"
12
+ require_relative "service/spa"
13
+ require_relative "service/cdp"
14
+ require_relative "service/cloud_browsers"
15
+
16
+ module AdsCrawl
17
+ class Client
18
+ DEFAULT_BASE_URL = "https://api.adscrawl.net"
19
+ DEFAULT_TIMEOUT_MS = 90_000
20
+ LAUNCH_TIMEOUT_MS = 195_000
21
+ MAX_TIMEOUT_MS = 2_147_483_647
22
+ MAX_BODY_BYTES = 64 * 1024 * 1024
23
+ PNG_SIGNATURE = "\x89PNG\r\n\x1a\n".b
24
+ RUNTIME_STATUSES = %w[starting running stopping stopped].freeze
25
+
26
+ attr_reader :spa, :cdp, :cloud_browsers
27
+ alias cloudBrowsers cloud_browsers
28
+
29
+ # Construction never performs a network request.
30
+ # A custom transport receives keyword arguments and returns HttpResponse.
31
+ def initialize(api_key: nil, base_url: nil, timeout_ms: nil, transport: nil)
32
+ selected_key = (api_key || ENV["ADSCRAWL_API_KEY"]).to_s.strip
33
+ raise ArgumentError, "Set ADSCRAWL_API_KEY or pass api_key to Client." if selected_key.empty?
34
+ raise ArgumentError, "api_key must not contain line breaks." if selected_key.match?(/[\r\n]/)
35
+
36
+ selected_url = base_url || ENV["ADSCRAWL_BASE_URL"] || ENV["ADSCRAWL_API_URL"] || DEFAULT_BASE_URL
37
+ @base_uri = validate_base_url(selected_url)
38
+ @base_url = selected_url.sub(%r{/+\z}, "")
39
+ @api_key = selected_key
40
+ @timeout_ms = validate_timeout(timeout_ms || DEFAULT_TIMEOUT_MS)
41
+ @custom_timeout = !timeout_ms.nil?
42
+ @transport = transport
43
+
44
+ @spa = Service::SPA.new(self)
45
+ @cdp = Service::CDP.new(self)
46
+ @cloud_browsers = Service::CloudBrowsers.new(self)
47
+ end
48
+
49
+ def html(options = nil, timeout_ms: nil, **keywords)
50
+ payload = options_hash(options, keywords)
51
+ validate_routing(payload)
52
+ body = send_request("POST", "/html", payload.merge("contentMode" => "html"), timeout_ms: timeout_ms)
53
+ raise ResponseError, "AdsCrawl returned empty page content." if body.strip.empty?
54
+
55
+ body.dup.force_encoding(Encoding::UTF_8)
56
+ end
57
+
58
+ def markdown(options = nil, timeout_ms: nil, **keywords)
59
+ payload = options_hash(options, keywords)
60
+ validate_routing(payload)
61
+ body = send_request("POST", "/html", payload.merge("contentMode" => "markdown"), timeout_ms: timeout_ms)
62
+ raise ResponseError, "AdsCrawl returned empty page content." if body.strip.empty?
63
+
64
+ body.dup.force_encoding(Encoding::UTF_8)
65
+ end
66
+
67
+ def article(options = nil, timeout_ms: nil, **keywords)
68
+ payload = options_hash(options, keywords)
69
+ validate_routing(payload)
70
+ request_json(
71
+ "POST",
72
+ "/html",
73
+ payload.merge("contentMode" => "json"),
74
+ timeout_ms: timeout_ms,
75
+ validate: lambda { |value|
76
+ string?(value, "title") && string?(value, "content") &&
77
+ string?(value, "textContent") && value["length"].is_a?(Numeric)
78
+ }
79
+ )
80
+ end
81
+
82
+ def screenshot(options = nil, timeout_ms: nil, **keywords)
83
+ payload = options_hash(options, keywords)
84
+ validate_routing(payload)
85
+ body = send_request("POST", "/screenshot", payload, timeout_ms: timeout_ms)
86
+ raise ResponseError, "AdsCrawl returned an invalid PNG screenshot." unless body.b.start_with?(PNG_SIGNATURE)
87
+
88
+ body.b
89
+ end
90
+
91
+ # Internal API used by service objects.
92
+ def send_request(method, path, body = nil, timeout_ms: nil, authenticate: true, extra_secrets: [])
93
+ deadline = validate_timeout(timeout_ms || @timeout_ms)
94
+ payload = body.nil? ? nil : normalize_hash(body)
95
+ secrets = [@api_key, *collect_secrets(payload), *extra_secrets].compact
96
+ encoded = payload.nil? ? nil : JSON.generate(payload)
97
+ headers = {"Accept" => "*/*", "User-Agent" => "adscrawl-ruby/#{VERSION}"}
98
+ headers["X-API-Key"] = @api_key if authenticate
99
+ headers["Content-Type"] = "application/json" unless encoded.nil?
100
+
101
+ response = if @transport
102
+ @transport.call(method: method, url: "#{@base_url}#{path}", headers: headers, body: encoded, timeout_ms: deadline)
103
+ else
104
+ net_http_request(method, "#{@base_url}#{path}", headers, encoded, deadline)
105
+ end
106
+ raise ResponseError, "The custom transport must return an AdsCrawl::HttpResponse." unless response.is_a?(HttpResponse)
107
+ raise ResponseError, "AdsCrawl response exceeds the 64 MiB safety limit." if response.body.bytesize > MAX_BODY_BYTES
108
+
109
+ status = Integer(response.status)
110
+ return response.body if status.between?(200, 299)
111
+
112
+ raise_api_error(status, response.body, response.headers || {}, secrets)
113
+ end
114
+
115
+ # Internal API used by service objects.
116
+ def request_json(method, path, body = nil, timeout_ms: nil, authenticate: true, validate: nil, extra_secrets: [])
117
+ raw = send_request(
118
+ method,
119
+ path,
120
+ body,
121
+ timeout_ms: timeout_ms,
122
+ authenticate: authenticate,
123
+ extra_secrets: extra_secrets
124
+ )
125
+ value = JSON.parse(raw)
126
+ raise ResponseError, "AdsCrawl returned invalid or unexpected JSON." unless value.is_a?(Hash)
127
+ if validate && !validate.call(value)
128
+ raise ResponseError, "AdsCrawl returned an unexpected JSON response shape."
129
+ end
130
+
131
+ value
132
+ rescue JSON::ParserError
133
+ raise ResponseError, "AdsCrawl returned invalid or unexpected JSON."
134
+ end
135
+
136
+ def launch_timeout_ms
137
+ @custom_timeout ? nil : LAUNCH_TIMEOUT_MS
138
+ end
139
+
140
+ def base_uri
141
+ @base_uri.dup
142
+ end
143
+
144
+ def normalize_hash(value)
145
+ raise ArgumentError, "options must be a Hash." unless value.is_a?(Hash)
146
+
147
+ normalize_value(value)
148
+ end
149
+
150
+ def options_hash(options, keywords)
151
+ if options && !keywords.empty?
152
+ raise ArgumentError, "Pass options as a Hash or keywords, not both."
153
+ end
154
+
155
+ normalize_hash(options || keywords)
156
+ end
157
+
158
+ def validate_routing(value)
159
+ return unless value.is_a?(Hash)
160
+
161
+ if value.key?("proxy") && !value["proxy"].nil? && value.key?("countryCode") && !value["countryCode"].to_s.empty?
162
+ raise ArgumentError, "proxy and countryCode cannot be combined."
163
+ end
164
+ validate_routing(value["browserSettings"]) if value["browserSettings"].is_a?(Hash)
165
+ end
166
+
167
+ def encode_id(value)
168
+ raise ArgumentError, "A non-empty resource id is required." unless value.is_a?(String) && !value.strip.empty? && !%w[. ..].include?(value)
169
+
170
+ URI.encode_www_form_component(value).tr("+", "%20")
171
+ end
172
+
173
+ def ok?(value)
174
+ value.is_a?(Hash) && value["ok"] == true
175
+ end
176
+
177
+ def session?(value)
178
+ value.is_a?(Hash) && string?(value, "sessionId") && string?(value, "expiresAt") && string?(value, "cdpBaseUrl")
179
+ end
180
+
181
+ def runtime?(value)
182
+ value.is_a?(Hash) && RUNTIME_STATUSES.include?(value["status"])
183
+ end
184
+
185
+ def cloud_browser?(value)
186
+ value.is_a?(Hash) && string?(value, "id") && runtime?(value["runtime"]) && value["browserSettings"].is_a?(Hash)
187
+ end
188
+
189
+ def page?(value)
190
+ value.is_a?(Hash) && string?(value, "url") && string?(value, "title")
191
+ end
192
+
193
+ def string?(value, key)
194
+ value.is_a?(Hash) && value[key].is_a?(String)
195
+ end
196
+
197
+ private
198
+
199
+ def validate_base_url(value)
200
+ uri = URI.parse(value.to_s)
201
+ unless %w[http https].include?(uri.scheme) && uri.host && !uri.userinfo && !uri.query && !uri.fragment
202
+ raise ArgumentError, "base_url must be an HTTP(S) URL without credentials, query, or fragment."
203
+ end
204
+
205
+ uri
206
+ rescue URI::InvalidURIError
207
+ raise ArgumentError, "base_url must be an HTTP(S) URL without credentials, query, or fragment."
208
+ end
209
+
210
+ def validate_timeout(value)
211
+ unless value.is_a?(Integer) && value.positive? && value <= MAX_TIMEOUT_MS
212
+ raise ArgumentError, "timeout_ms must be a positive Integer no greater than #{MAX_TIMEOUT_MS}."
213
+ end
214
+
215
+ value
216
+ end
217
+
218
+ def normalize_value(value)
219
+ case value
220
+ when Hash
221
+ value.each_with_object({}) { |(key, item), result| result[key.to_s] = normalize_value(item) }
222
+ when Array
223
+ value.map { |item| normalize_value(item) }
224
+ else
225
+ value
226
+ end
227
+ end
228
+
229
+ def net_http_request(method, url, headers, body, timeout_ms)
230
+ uri = URI.parse(url)
231
+ request_class = {
232
+ "GET" => Net::HTTP::Get,
233
+ "POST" => Net::HTTP::Post,
234
+ "DELETE" => Net::HTTP::Delete
235
+ }.fetch(method)
236
+ request = request_class.new(uri.request_uri, headers)
237
+ request.body = body unless body.nil?
238
+ timeout_seconds = timeout_ms / 1000.0
239
+ response_headers = {}
240
+ response_body = String.new(encoding: Encoding::BINARY)
241
+ status = nil
242
+
243
+ http = Net::HTTP.new(uri.host, uri.port)
244
+ http.use_ssl = uri.scheme == "https"
245
+ http.open_timeout = [timeout_seconds, 10.0].min
246
+ http.read_timeout = timeout_seconds
247
+ http.write_timeout = timeout_seconds if http.respond_to?(:write_timeout=)
248
+ http.request(request) do |response|
249
+ status = response.code.to_i
250
+ response.each_header { |name, value| response_headers[name.downcase] = value }
251
+ response.read_body do |chunk|
252
+ if response_body.bytesize + chunk.bytesize > MAX_BODY_BYTES
253
+ raise ResponseError, "AdsCrawl response exceeds the 64 MiB safety limit."
254
+ end
255
+ response_body << chunk.b
256
+ end
257
+ end
258
+
259
+ HttpResponse.new(status: status, headers: response_headers, body: response_body)
260
+ rescue Net::OpenTimeout, Net::ReadTimeout, Net::WriteTimeout, Timeout::Error
261
+ raise TimeoutError, timeout_ms
262
+ rescue ResponseError
263
+ raise
264
+ rescue SocketError, OpenSSL::SSL::SSLError, EOFError, IOError, SystemCallError, URI::InvalidURIError
265
+ raise ConnectionError
266
+ end
267
+
268
+ def raise_api_error(status, raw, headers, secrets)
269
+ parsed = JSON.parse(raw)
270
+ rescue JSON::ParserError
271
+ parsed = raw.to_s
272
+ ensure
273
+ safe_body = redact_value(parsed, secrets)
274
+ message = "AdsCrawl API returned HTTP #{status}"
275
+ api_code = resource_id = trace_id = nil
276
+ if safe_body.is_a?(Hash)
277
+ message = safe_body["error"] if safe_body["error"].is_a?(String)
278
+ api_code = safe_body["code"] if safe_body["code"].is_a?(String)
279
+ resource_id = safe_body["id"] if safe_body["id"].is_a?(String)
280
+ trace_id = safe_body["traceId"] if safe_body["traceId"].is_a?(String)
281
+ end
282
+ normalized_headers = headers.each_with_object({}) { |(key, value), result| result[key.to_s.downcase] = value.to_s }
283
+ trace_id = redact_string(normalized_headers["x-trace-id"], secrets) if normalized_headers["x-trace-id"]
284
+ request_id = normalized_headers["x-request-id"] && redact_string(normalized_headers["x-request-id"], secrets)
285
+ raise APIError.new(
286
+ message.to_s.slice(0, 2048),
287
+ status: status,
288
+ api_code: api_code,
289
+ resource_id: resource_id,
290
+ trace_id: trace_id,
291
+ request_id: request_id,
292
+ body: safe_body
293
+ )
294
+ end
295
+
296
+ def collect_secrets(value, result = [])
297
+ case value
298
+ when Hash
299
+ value.each do |key, item|
300
+ if key == "cookies" && item.is_a?(Array)
301
+ item.each { |cookie| result << cookie["value"] if cookie.is_a?(Hash) && cookie["value"].is_a?(String) && !cookie["value"].empty? }
302
+ end
303
+ if %w[password username token controltoken].include?(key.downcase) && item.is_a?(String) && !item.empty?
304
+ result << item
305
+ end
306
+ collect_secrets(item, result)
307
+ end
308
+ when Array
309
+ value.each { |item| collect_secrets(item, result) }
310
+ end
311
+ result
312
+ end
313
+
314
+ def redact_string(value, secrets)
315
+ output = value.to_s.dup
316
+ secrets.each { |secret| output.gsub!(secret, "[REDACTED]") unless secret.to_s.empty? }
317
+ output.gsub!(/([?&](?:token|controlToken|apiKey|api_key)=)[^&#\s"<>]+/i, '\\1[REDACTED]')
318
+ output.gsub!(/(\b(?:https?|socks5):\/\/)[^\s\/@]+:[^\s\/@]+@/i, '\\1[REDACTED]@')
319
+ output
320
+ end
321
+
322
+ def redact_value(value, secrets)
323
+ case value
324
+ when String
325
+ redact_string(value, secrets)
326
+ when Array
327
+ value.map { |item| redact_value(item, secrets) }
328
+ when Hash
329
+ value.each_with_object({}) do |(key, item), result|
330
+ normalized = key.to_s.downcase.delete("-_")
331
+ result[key] = if %w[apikey xapikey authorization password username cookie cookies token controltoken websocketdebuggerurl cdpbaseurl controlurl connecturl].include?(normalized)
332
+ "[REDACTED]"
333
+ else
334
+ redact_value(item, secrets)
335
+ end
336
+ end
337
+ else
338
+ value
339
+ end
340
+ end
341
+ end
342
+ end
@@ -0,0 +1,36 @@
1
+ # frozen_string_literal: true
2
+
3
+ module AdsCrawl
4
+ class Error < StandardError; end
5
+
6
+ class APIError < Error
7
+ attr_reader :status, :api_code, :resource_id, :trace_id, :request_id, :body
8
+
9
+ def initialize(message, status:, api_code: nil, resource_id: nil, trace_id: nil, request_id: nil, body: nil)
10
+ @status = status
11
+ @api_code = api_code
12
+ @resource_id = resource_id
13
+ @trace_id = trace_id
14
+ @request_id = request_id
15
+ @body = body
16
+ super(message)
17
+ end
18
+ end
19
+
20
+ class TimeoutError < Error
21
+ attr_reader :timeout_ms
22
+
23
+ def initialize(timeout_ms)
24
+ @timeout_ms = timeout_ms
25
+ super("AdsCrawl request timed out after #{timeout_ms}ms. Remote work may still be running; inspect sessions before retrying.")
26
+ end
27
+ end
28
+
29
+ class ConnectionError < Error
30
+ def initialize
31
+ super("Unable to complete the AdsCrawl request. Check your network and service availability.")
32
+ end
33
+ end
34
+
35
+ class ResponseError < Error; end
36
+ end
@@ -0,0 +1,6 @@
1
+ # frozen_string_literal: true
2
+
3
+ module AdsCrawl
4
+ # Public so custom transports can return a response without depending on internals.
5
+ HttpResponse = Struct.new(:status, :headers, :body, keyword_init: true)
6
+ end
@@ -0,0 +1,92 @@
1
+ # frozen_string_literal: true
2
+
3
+ module AdsCrawl
4
+ module Service
5
+ class CDP
6
+ def initialize(client)
7
+ @client = client
8
+ end
9
+
10
+ def create(options = nil, timeout_ms: nil, **keywords)
11
+ payload = @client.options_hash(options, keywords)
12
+ @client.validate_routing(payload)
13
+ @client.request_json(
14
+ "POST",
15
+ "/cdp/sessions",
16
+ payload,
17
+ timeout_ms: timeout_ms,
18
+ validate: ->(value) { @client.session?(value) }
19
+ )
20
+ end
21
+
22
+ def list(timeout_ms: nil)
23
+ @client.request_json(
24
+ "GET",
25
+ "/cdp/sessions",
26
+ timeout_ms: timeout_ms,
27
+ validate: lambda { |value|
28
+ @client.ok?(value) && value["data"].is_a?(Array) && value["data"].all? { |session| @client.session?(session) }
29
+ }
30
+ )
31
+ end
32
+
33
+ def close(session_id, timeout_ms: nil)
34
+ id = @client.encode_id(session_id)
35
+ @client.request_json(
36
+ "DELETE",
37
+ "/cdp/sessions/#{id}",
38
+ timeout_ms: timeout_ms,
39
+ validate: ->(value) { @client.ok?(value) }
40
+ )
41
+ end
42
+
43
+ # Discovers the WebSocket endpoint without forwarding the API key.
44
+ def get_version(session, timeout_ms: nil)
45
+ payload = @client.normalize_hash(session)
46
+ session_id = payload["sessionId"]
47
+ cdp_base_url = payload["cdpBaseUrl"]
48
+ unless session_id.is_a?(String) && cdp_base_url.is_a?(String)
49
+ raise ArgumentError, "The session must contain sessionId and a valid cdpBaseUrl."
50
+ end
51
+
52
+ id = @client.encode_id(session_id)
53
+ parsed = URI.parse(cdp_base_url)
54
+ base = @client.base_uri
55
+ token = URI.decode_www_form(parsed.query.to_s).to_h["token"]
56
+ expected_path = "#{base.path.to_s.sub(%r{/+\z}, "")}/cdp/sessions/#{id}"
57
+ unless parsed.scheme == base.scheme && parsed.host == base.host && parsed.port == base.port &&
58
+ parsed.path.sub(%r{/+\z}, "") == expected_path && !parsed.userinfo && !parsed.fragment &&
59
+ token.is_a?(String) && !token.empty?
60
+ raise ArgumentError, "cdpBaseUrl must belong to this API origin and session and include a data token."
61
+ end
62
+
63
+ path = "/cdp/sessions/#{id}/json/version?#{URI.encode_www_form(token: token)}"
64
+ @client.request_json(
65
+ "GET",
66
+ path,
67
+ timeout_ms: timeout_ms,
68
+ authenticate: false,
69
+ extra_secrets: [token],
70
+ validate: ->(value) { @client.string?(value, "Browser") && @client.string?(value, "webSocketDebuggerUrl") }
71
+ )
72
+ rescue URI::InvalidURIError, ArgumentError => error
73
+ raise error if error.message.start_with?("cdpBaseUrl")
74
+
75
+ raise ArgumentError, "The session must contain sessionId and a valid cdpBaseUrl."
76
+ end
77
+
78
+ def live_token(session_id, timeout_ms: nil)
79
+ @client.encode_id(session_id)
80
+ @client.request_json(
81
+ "POST",
82
+ "/cdp/live-token",
83
+ {"sessionId" => session_id},
84
+ timeout_ms: timeout_ms,
85
+ validate: lambda { |value|
86
+ @client.ok?(value) && @client.string?(value, "controlUrl") && value["expiresAt"].is_a?(Numeric)
87
+ }
88
+ )
89
+ end
90
+ end
91
+ end
92
+ end
@@ -0,0 +1,98 @@
1
+ # frozen_string_literal: true
2
+
3
+ module AdsCrawl
4
+ module Service
5
+ class CloudBrowsers
6
+ def initialize(client)
7
+ @client = client
8
+ end
9
+
10
+ def list(options = nil, timeout_ms: nil, **keywords)
11
+ payload = @client.options_hash(options, keywords)
12
+ query = payload.slice("page", "pageSize")
13
+ path = query.empty? ? "/cloud-browsers" : "/cloud-browsers?#{URI.encode_www_form(query)}"
14
+ @client.request_json(
15
+ "GET",
16
+ path,
17
+ timeout_ms: timeout_ms,
18
+ validate: lambda { |value|
19
+ @client.ok?(value) && value["data"].is_a?(Array) &&
20
+ value["data"].all? { |browser| @client.cloud_browser?(browser) } &&
21
+ value["pagination"].is_a?(Hash) &&
22
+ %w[limit runningLimit runningCount].all? { |key| value[key].is_a?(Integer) }
23
+ }
24
+ )
25
+ end
26
+
27
+ def create(options = nil, timeout_ms: nil, **keywords)
28
+ payload = @client.options_hash(options, keywords)
29
+ @client.validate_routing(payload)
30
+ @client.request_json(
31
+ "POST",
32
+ "/cloud-browsers",
33
+ payload,
34
+ timeout_ms: timeout_ms,
35
+ validate: ->(value) { @client.ok?(value) && @client.string?(value, "id") }
36
+ )
37
+ end
38
+
39
+ def get(browser_id, timeout_ms: nil)
40
+ id = @client.encode_id(browser_id)
41
+ @client.request_json(
42
+ "GET",
43
+ "/cloud-browsers/#{id}",
44
+ timeout_ms: timeout_ms,
45
+ validate: ->(value) { @client.cloud_browser?(value) }
46
+ )
47
+ end
48
+
49
+ def start(browser_id, options = nil, timeout_ms: nil, **keywords)
50
+ id = @client.encode_id(browser_id)
51
+ payload = require_proxy(@client.options_hash(options, keywords), "starts")
52
+ @client.request_json(
53
+ "POST",
54
+ "/cloud-browsers/#{id}/start",
55
+ payload,
56
+ timeout_ms: timeout_ms,
57
+ validate: ->(value) { @client.ok?(value) && @client.runtime?(value["runtime"]) }
58
+ )
59
+ end
60
+
61
+ def stop(browser_id, timeout_ms: nil)
62
+ id = @client.encode_id(browser_id)
63
+ @client.request_json(
64
+ "POST",
65
+ "/cloud-browsers/#{id}/stop",
66
+ timeout_ms: timeout_ms,
67
+ validate: ->(value) { @client.ok?(value) && @client.runtime?(value["runtime"]) }
68
+ )
69
+ end
70
+
71
+ def launch(options = nil, timeout_ms: nil, **keywords)
72
+ payload = require_proxy(@client.options_hash(options, keywords), "launches")
73
+ deadline = timeout_ms || @client.launch_timeout_ms
74
+ @client.request_json(
75
+ "POST",
76
+ "/cloud-browsers/launch",
77
+ payload,
78
+ timeout_ms: deadline,
79
+ validate: lambda { |value|
80
+ @client.ok?(value) && @client.string?(value, "id") && value["source"] == "launch" &&
81
+ value["deleteOnStop"] == false && @client.runtime?(value["runtime"])
82
+ }
83
+ )
84
+ end
85
+
86
+ private
87
+
88
+ def require_proxy(options, action)
89
+ payload = @client.normalize_hash(options)
90
+ unless payload["proxy"].is_a?(Hash)
91
+ raise ArgumentError, "Cloud browser #{action} require an explicit top-level proxy."
92
+ end
93
+
94
+ payload
95
+ end
96
+ end
97
+ end
98
+ end
@@ -0,0 +1,54 @@
1
+ # frozen_string_literal: true
2
+
3
+ module AdsCrawl
4
+ module Service
5
+ class SPA
6
+ def initialize(client)
7
+ @client = client
8
+ end
9
+
10
+ def templates(timeout_ms: nil)
11
+ @client.request_json(
12
+ "GET",
13
+ "/spa-extract/templates",
14
+ timeout_ms: timeout_ms,
15
+ validate: lambda { |value|
16
+ value["templates"].is_a?(Array) && value["templates"].all? do |template|
17
+ @client.string?(template, "id") && @client.string?(template, "name")
18
+ end
19
+ }
20
+ )
21
+ end
22
+
23
+ def extract(options = nil, timeout_ms: nil, **keywords)
24
+ payload = @client.options_hash(options, keywords)
25
+ @client.validate_routing(payload)
26
+ @client.request_json(
27
+ "POST",
28
+ "/spa-extract",
29
+ payload.merge("mode" => "extract"),
30
+ timeout_ms: timeout_ms,
31
+ validate: lambda { |value|
32
+ value["mode"] == "extract" && @client.page?(value["page"]) &&
33
+ value["data"].is_a?(Hash) && value["missingFields"].is_a?(Array)
34
+ }
35
+ )
36
+ end
37
+
38
+ def inspect(options = nil, timeout_ms: nil, **keywords)
39
+ payload = @client.options_hash(options, keywords)
40
+ @client.validate_routing(payload)
41
+ @client.request_json(
42
+ "POST",
43
+ "/spa-extract",
44
+ payload.merge("mode" => "inspect"),
45
+ timeout_ms: timeout_ms,
46
+ validate: lambda { |value|
47
+ value["mode"] == "inspect" && @client.page?(value["page"]) &&
48
+ value["candidates"].is_a?(Hash) && value["suggestedPlan"].is_a?(Hash)
49
+ }
50
+ )
51
+ end
52
+ end
53
+ end
54
+ end
@@ -0,0 +1,5 @@
1
+ # frozen_string_literal: true
2
+
3
+ module AdsCrawl
4
+ VERSION = "0.1.0"
5
+ end
data/lib/adscrawl.rb ADDED
@@ -0,0 +1,4 @@
1
+ # frozen_string_literal: true
2
+
3
+ require_relative "adscrawl/version"
4
+ require_relative "adscrawl/client"
metadata ADDED
@@ -0,0 +1,68 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: adscrawl
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - AdsCrawl
8
+ bindir: bin
9
+ cert_chain: []
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
+ dependencies: []
12
+ description: Render pages in real browsers, extract structured data, capture screenshots,
13
+ open CDP sessions, and manage cloud browsers.
14
+ email:
15
+ - support@adscrawl.net
16
+ executables: []
17
+ extensions: []
18
+ extra_rdoc_files: []
19
+ files:
20
+ - CHANGELOG.md
21
+ - LICENSE
22
+ - README.md
23
+ - README.zh-CN.md
24
+ - RELEASING.md
25
+ - assets/adscrawl-logo.svg
26
+ - examples/cdp.rb
27
+ - examples/cloud_browser.rb
28
+ - examples/extract.rb
29
+ - examples/fingerprint.rb
30
+ - examples/markdown.rb
31
+ - examples/screenshot.rb
32
+ - lib/adscrawl.rb
33
+ - lib/adscrawl/client.rb
34
+ - lib/adscrawl/errors.rb
35
+ - lib/adscrawl/http_response.rb
36
+ - lib/adscrawl/service/cdp.rb
37
+ - lib/adscrawl/service/cloud_browsers.rb
38
+ - lib/adscrawl/service/spa.rb
39
+ - lib/adscrawl/version.rb
40
+ homepage: https://www.adscrawl.net
41
+ licenses:
42
+ - MIT
43
+ metadata:
44
+ allowed_push_host: https://rubygems.org
45
+ bug_tracker_uri: https://github.com/AdsCrawl/adscrawl-ruby/issues
46
+ changelog_uri: https://github.com/AdsCrawl/adscrawl-ruby/blob/main/CHANGELOG.md
47
+ documentation_uri: https://www.adscrawl.net/docs/
48
+ homepage_uri: https://www.adscrawl.net
49
+ rubygems_mfa_required: 'true'
50
+ source_code_uri: https://github.com/AdsCrawl/adscrawl-ruby
51
+ rdoc_options: []
52
+ require_paths:
53
+ - lib
54
+ required_ruby_version: !ruby/object:Gem::Requirement
55
+ requirements:
56
+ - - ">="
57
+ - !ruby/object:Gem::Version
58
+ version: '3.1'
59
+ required_rubygems_version: !ruby/object:Gem::Requirement
60
+ requirements:
61
+ - - ">="
62
+ - !ruby/object:Gem::Version
63
+ version: '0'
64
+ requirements: []
65
+ rubygems_version: 4.0.20
66
+ specification_version: 4
67
+ summary: Official Ruby SDK for the AdsCrawl browser and extraction API
68
+ test_files: []