adscrawl 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/CHANGELOG.md +7 -0
- data/LICENSE +21 -0
- data/README.md +208 -0
- data/README.zh-CN.md +208 -0
- data/RELEASING.md +15 -0
- data/assets/adscrawl-logo.svg +23 -0
- data/examples/cdp.rb +17 -0
- data/examples/cloud_browser.rb +17 -0
- data/examples/extract.rb +16 -0
- data/examples/fingerprint.rb +39 -0
- data/examples/markdown.rb +7 -0
- data/examples/screenshot.rb +10 -0
- data/lib/adscrawl/client.rb +342 -0
- data/lib/adscrawl/errors.rb +36 -0
- data/lib/adscrawl/http_response.rb +6 -0
- data/lib/adscrawl/service/cdp.rb +92 -0
- data/lib/adscrawl/service/cloud_browsers.rb +98 -0
- data/lib/adscrawl/service/spa.rb +54 -0
- data/lib/adscrawl/version.rb +5 -0
- data/lib/adscrawl.rb +4 -0
- metadata +68 -0
checksums.yaml
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
SHA256:
|
|
3
|
+
metadata.gz: 53190cac137e9d06f3175969488c05c975dc10ce41bae70c0682110242666b05
|
|
4
|
+
data.tar.gz: dbe222b0ff615996d4dbcb6071d9f4f9289d83467d850632c8778ba308be5c3b
|
|
5
|
+
SHA512:
|
|
6
|
+
metadata.gz: edc122d2a0d30101806275101c8def52f424a62348428405de2a91973559952cc032f84bf4088b744e493d0add8a400d3e580638317b069fcd041bf67bbcecc2
|
|
7
|
+
data.tar.gz: ae6f562c5cc5249c2f9b31ed7a2a424ffc93e01b1ce79f38e7b4b7272b9d4b2f7ba412cc20a268cb532fe299c949df1491fb647c84cca7a9f7ccb0d55e98f132
|
data/CHANGELOG.md
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0 - 2026-09-28
|
|
4
|
+
|
|
5
|
+
- Initial Ruby SDK for rendered HTML, Markdown, articles, and PNG screenshots.
|
|
6
|
+
- SPA extraction, CDP session, and persistent cloud browser services.
|
|
7
|
+
- Request deadlines, response validation, credential-redacted errors, and runnable examples.
|
data/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 AdsCrawl
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
data/README.md
ADDED
|
@@ -0,0 +1,208 @@
|
|
|
1
|
+
<p align="center">
|
|
2
|
+
<a href="https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby">
|
|
3
|
+
<img src="assets/adscrawl-logo.svg" alt="AdsCrawl" width="360" />
|
|
4
|
+
</a>
|
|
5
|
+
</p>
|
|
6
|
+
|
|
7
|
+
<p align="center"><strong>Real browsers. Structured extraction. Screenshots and automation.</strong></p>
|
|
8
|
+
|
|
9
|
+
<h1 align="center">Ruby SDK</h1>
|
|
10
|
+
|
|
11
|
+
Turn a URL into rendered HTML, readable Markdown, structured data, or a PNG screenshot. Open remote CDP sessions and manage persistent cloud browsers with the same API as the TypeScript, Python, Go, and PHP SDKs.
|
|
12
|
+
|
|
13
|
+
[Website](https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [API documentation](https://www.adscrawl.net/docs/) · [Get an API key](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [中文](./README.zh-CN.md)
|
|
14
|
+
|
|
15
|
+
- Ruby 3.1+ with no runtime dependencies outside the standard library.
|
|
16
|
+
- Explicit HTTP deadlines, checked response shapes, and credential-redacted errors.
|
|
17
|
+
- Hash and keyword arguments use the HTTP API's camelCase field names.
|
|
18
|
+
- No automatic retries of metered requests or browser creation.
|
|
19
|
+
|
|
20
|
+
## Install
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
gem install adscrawl
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
Or add it to your bundle:
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
bundle add adscrawl
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## Your first request
|
|
33
|
+
|
|
34
|
+
[Create an account](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby), create an API key, and set `ADSCRAWL_API_KEY` in your server environment.
|
|
35
|
+
|
|
36
|
+
```ruby
|
|
37
|
+
require "adscrawl"
|
|
38
|
+
|
|
39
|
+
client = AdsCrawl::Client.new
|
|
40
|
+
markdown = client.markdown(
|
|
41
|
+
url: "https://www.adscrawl.net",
|
|
42
|
+
waitUntil: "domcontentloaded"
|
|
43
|
+
)
|
|
44
|
+
puts markdown
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
You can also use `AdsCrawl::Client.new(api_key: "...")`. Keep API keys in server-side code. See the complete runnable examples for [Markdown](./examples/markdown.rb), [screenshots](./examples/screenshot.rb), [structured extraction](./examples/extract.rb), [fingerprints](./examples/fingerprint.rb), [CDP](./examples/cdp.rb), and [cloud browsers](./examples/cloud_browser.rb).
|
|
48
|
+
|
|
49
|
+
## Rendered content and screenshots
|
|
50
|
+
|
|
51
|
+
```ruby
|
|
52
|
+
require "adscrawl"
|
|
53
|
+
|
|
54
|
+
client = AdsCrawl::Client.new
|
|
55
|
+
html = client.html(url: "https://www.adscrawl.net")
|
|
56
|
+
article = client.article(url: "https://www.adscrawl.net")
|
|
57
|
+
puts [article["title"], article["textContent"]]
|
|
58
|
+
|
|
59
|
+
png = client.screenshot(
|
|
60
|
+
url: "https://www.adscrawl.net",
|
|
61
|
+
viewport: {width: 1440, height: 900},
|
|
62
|
+
fullPage: true,
|
|
63
|
+
waitUntil: "load"
|
|
64
|
+
)
|
|
65
|
+
File.binwrite("page.png", png)
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
`html` returns HTML, `markdown` returns Markdown, `article` returns a Hash, and `screenshot` returns binary PNG data. Page options include viewport, locale, cookies, custom `proxy`, managed `countryCode`, fingerprint, server-side `timeoutMs`, and navigation `waitUntil`. A custom proxy and managed country cannot be combined.
|
|
69
|
+
|
|
70
|
+
## Proxy and fingerprint: BrowserScan screenshot
|
|
71
|
+
|
|
72
|
+
AdsCrawl uses real browsers with configurable routing and browser fingerprints. This complete example uses managed `GLOBAL` routing unless proxy environment variables are present, opens [BrowserScan](https://www.browserscan.net/), and saves the returned screenshot. The same browser workflow can access and render [Pixelscan](https://pixelscan.net/) and [IPhey](https://iphey.com/); inspect the returned page for its current result.
|
|
73
|
+
|
|
74
|
+
```ruby
|
|
75
|
+
require "adscrawl"
|
|
76
|
+
|
|
77
|
+
client = AdsCrawl::Client.new
|
|
78
|
+
server = ENV["ADSCRAWL_PROXY_SERVER"]
|
|
79
|
+
username = ENV["ADSCRAWL_PROXY_USERNAME"]
|
|
80
|
+
password = ENV["ADSCRAWL_PROXY_PASSWORD"]
|
|
81
|
+
raise "Set both proxy username and password, or neither." if username.nil? != password.nil?
|
|
82
|
+
raise "Set ADSCRAWL_PROXY_SERVER with proxy credentials." if server.nil? && (username || password)
|
|
83
|
+
|
|
84
|
+
routing = {countryCode: "GLOBAL"}
|
|
85
|
+
if server
|
|
86
|
+
proxy = {server: server}
|
|
87
|
+
proxy.merge!(username: username, password: password) if username && password
|
|
88
|
+
routing = {proxy: proxy}
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
png = client.screenshot(
|
|
92
|
+
**routing,
|
|
93
|
+
url: "https://www.browserscan.net/",
|
|
94
|
+
viewport: {width: 1440, height: 900},
|
|
95
|
+
fullPage: true,
|
|
96
|
+
waitUntil: "networkidle",
|
|
97
|
+
timeoutMs: 60_000,
|
|
98
|
+
userAgentMode: "random",
|
|
99
|
+
userAgentOs: "windows",
|
|
100
|
+
fingerprint: {
|
|
101
|
+
webRtc: "forward", webGl: "random", webGpu: "random",
|
|
102
|
+
webGlImage: "random", canvas: "random", audioContext: "random",
|
|
103
|
+
clientRects: "random", speechVoices: "random", fonts: "random",
|
|
104
|
+
hardware: "random", doNotTrack: "random"
|
|
105
|
+
},
|
|
106
|
+
timeout_ms: 75_000
|
|
107
|
+
)
|
|
108
|
+
File.binwrite("browserscan.png", png)
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Run the included version with `ruby examples/fingerprint.rb`. Pass a detection URL and output path as optional arguments. The example never combines a custom proxy with managed routing.
|
|
112
|
+
|
|
113
|
+
## Structured extraction
|
|
114
|
+
|
|
115
|
+
```ruby
|
|
116
|
+
require "adscrawl"
|
|
117
|
+
|
|
118
|
+
client = AdsCrawl::Client.new
|
|
119
|
+
puts client.spa.templates["templates"]
|
|
120
|
+
|
|
121
|
+
result = client.spa.extract(
|
|
122
|
+
url: "https://www.adscrawl.net",
|
|
123
|
+
fields: {
|
|
124
|
+
title: {source: "dom", selector: "h1", value: "text", required: true}
|
|
125
|
+
}
|
|
126
|
+
)
|
|
127
|
+
puts result.dig("data", "title")
|
|
128
|
+
|
|
129
|
+
inspection = client.spa.inspect(url: "https://www.adscrawl.net")
|
|
130
|
+
puts inspection["candidates"]
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
## Remote CDP browsers
|
|
134
|
+
|
|
135
|
+
```ruby
|
|
136
|
+
require "adscrawl"
|
|
137
|
+
|
|
138
|
+
client = AdsCrawl::Client.new
|
|
139
|
+
session = client.cdp.create(
|
|
140
|
+
idleTimeoutMs: 600_000,
|
|
141
|
+
maxSessionMs: 3_600_000,
|
|
142
|
+
browserSettings: {viewport: {width: 1440, height: 900}}
|
|
143
|
+
)
|
|
144
|
+
begin
|
|
145
|
+
version = client.cdp.get_version(session)
|
|
146
|
+
puts version["Browser"]
|
|
147
|
+
# Connect a CDP-compatible Ruby library to session["cdpBaseUrl"].
|
|
148
|
+
ensure
|
|
149
|
+
client.cdp.close(session["sessionId"])
|
|
150
|
+
end
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
`client.cdp.list` returns active sessions. `get_version` validates the session token URL and deliberately omits the API key. CDP connection URLs contain secrets; do not log them.
|
|
154
|
+
|
|
155
|
+
## Persistent cloud browsers
|
|
156
|
+
|
|
157
|
+
```ruby
|
|
158
|
+
require "adscrawl"
|
|
159
|
+
|
|
160
|
+
client = AdsCrawl::Client.new
|
|
161
|
+
launched = client.cloud_browsers.launch(
|
|
162
|
+
proxy: {server: ENV.fetch("ADSCRAWL_PROXY_SERVER")},
|
|
163
|
+
tabs: ["https://www.adscrawl.net"]
|
|
164
|
+
)
|
|
165
|
+
browser_id = launched["id"]
|
|
166
|
+
begin
|
|
167
|
+
profile = client.cloud_browsers.get(browser_id)
|
|
168
|
+
puts profile.dig("runtime", "status")
|
|
169
|
+
ensure
|
|
170
|
+
client.cloud_browsers.stop(browser_id)
|
|
171
|
+
end
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
API-key starts and launches require an explicit top-level custom proxy. A `stopping` response means shutdown is pending; poll `get` until `stopped`. A failed launch may expose a cleanup profile ID as `APIError#resource_id`.
|
|
175
|
+
|
|
176
|
+
## Configuration and errors
|
|
177
|
+
|
|
178
|
+
```ruby
|
|
179
|
+
require "adscrawl"
|
|
180
|
+
|
|
181
|
+
client = AdsCrawl::Client.new(base_url: "https://api.adscrawl.net", timeout_ms: 90_000)
|
|
182
|
+
begin
|
|
183
|
+
puts client.markdown(
|
|
184
|
+
{url: "https://www.adscrawl.net", timeoutMs: 60_000},
|
|
185
|
+
timeout_ms: 75_000
|
|
186
|
+
)
|
|
187
|
+
rescue AdsCrawl::APIError => error
|
|
188
|
+
puts [error.status, error.api_code, error.trace_id]
|
|
189
|
+
rescue AdsCrawl::TimeoutError
|
|
190
|
+
puts "Inspect remote sessions before retrying."
|
|
191
|
+
end
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
The API key defaults to `ADSCRAWL_API_KEY`; the base URL defaults to `ADSCRAWL_BASE_URL`, then `ADSCRAWL_API_URL`, then `https://api.adscrawl.net`. Errors include `APIError`, `TimeoutError`, `ConnectionError`, and `ResponseError`. API response bodies are credential-redacted. The HTTP deadline includes response body reading. Requests are not automatically retried.
|
|
195
|
+
|
|
196
|
+
## Develop and release
|
|
197
|
+
|
|
198
|
+
```bash
|
|
199
|
+
bundle install
|
|
200
|
+
bundle exec rake test
|
|
201
|
+
gem build adscrawl.gemspec
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
Tests use local and fake responses and consume no AdsCrawl credits. Releases use RubyGems Trusted Publishing with no long-lived registry token; see [RELEASING.md](./RELEASING.md).
|
|
205
|
+
|
|
206
|
+
## License
|
|
207
|
+
|
|
208
|
+
MIT
|
data/README.zh-CN.md
ADDED
|
@@ -0,0 +1,208 @@
|
|
|
1
|
+
<p align="center">
|
|
2
|
+
<a href="https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby">
|
|
3
|
+
<img src="assets/adscrawl-logo.svg" alt="AdsCrawl" width="360" />
|
|
4
|
+
</a>
|
|
5
|
+
</p>
|
|
6
|
+
|
|
7
|
+
<p align="center"><strong>真实浏览器、结构化提取、截图与自动化。</strong></p>
|
|
8
|
+
|
|
9
|
+
<h1 align="center">Ruby SDK</h1>
|
|
10
|
+
|
|
11
|
+
把 URL 转换为浏览器渲染后的 HTML、Markdown、结构化数据或 PNG 截图;也可以创建远程 CDP 会话和管理持久云浏览器。
|
|
12
|
+
|
|
13
|
+
[官网](https://www.adscrawl.net/?utm_source=github&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [API 文档](https://www.adscrawl.net/docs/) · [获取 API Key](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby) · [English](./README.md)
|
|
14
|
+
|
|
15
|
+
- 支持 Ruby 3.1+,除标准库外无运行时依赖。
|
|
16
|
+
- 支持 HTTP 超时、响应结构校验和错误信息凭据脱敏。
|
|
17
|
+
- Hash 和关键字参数与 HTTP API 使用相同的 camelCase 字段。
|
|
18
|
+
- 对计费请求和浏览器创建请求不会自动重试。
|
|
19
|
+
|
|
20
|
+
## 安装
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
gem install adscrawl
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
也可以添加到 Bundle:
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
bundle add adscrawl
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## 第一个请求
|
|
33
|
+
|
|
34
|
+
[注册账号](https://app.adscrawl.net/register/?utm_source=rubygems&utm_medium=sdk&utm_campaign=adscrawl-ruby),创建 API Key,并在服务端环境变量中设置 `ADSCRAWL_API_KEY`。
|
|
35
|
+
|
|
36
|
+
```ruby
|
|
37
|
+
require "adscrawl"
|
|
38
|
+
|
|
39
|
+
client = AdsCrawl::Client.new
|
|
40
|
+
markdown = client.markdown(
|
|
41
|
+
url: "https://www.adscrawl.net",
|
|
42
|
+
waitUntil: "domcontentloaded"
|
|
43
|
+
)
|
|
44
|
+
puts markdown
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
也可以使用 `AdsCrawl::Client.new(api_key: "...")`。API Key 应只保存在服务端。仓库包含完整可运行的 [Markdown](./examples/markdown.rb)、[截图](./examples/screenshot.rb)、[结构化提取](./examples/extract.rb)、[指纹检测](./examples/fingerprint.rb)、[CDP](./examples/cdp.rb) 和 [云浏览器](./examples/cloud_browser.rb) 示例。
|
|
48
|
+
|
|
49
|
+
## 渲染内容与截图
|
|
50
|
+
|
|
51
|
+
```ruby
|
|
52
|
+
require "adscrawl"
|
|
53
|
+
|
|
54
|
+
client = AdsCrawl::Client.new
|
|
55
|
+
html = client.html(url: "https://www.adscrawl.net")
|
|
56
|
+
article = client.article(url: "https://www.adscrawl.net")
|
|
57
|
+
puts [article["title"], article["textContent"]]
|
|
58
|
+
|
|
59
|
+
png = client.screenshot(
|
|
60
|
+
url: "https://www.adscrawl.net",
|
|
61
|
+
viewport: {width: 1440, height: 900},
|
|
62
|
+
fullPage: true,
|
|
63
|
+
waitUntil: "load"
|
|
64
|
+
)
|
|
65
|
+
File.binwrite("page.png", png)
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
`html` 返回 HTML,`markdown` 返回 Markdown,`article` 返回 Hash,`screenshot` 返回 PNG 二进制数据。页面参数支持视口、语言、Cookie、自定义 `proxy`、托管 `countryCode`、浏览器指纹、服务端 `timeoutMs` 与导航 `waitUntil`。自定义代理和托管国家线路不能同时使用。
|
|
69
|
+
|
|
70
|
+
## 代理与指纹:BrowserScan 截图
|
|
71
|
+
|
|
72
|
+
AdsCrawl 使用真实浏览器,并支持配置线路和浏览器指纹。以下示例默认使用 `GLOBAL` 托管线路;设置代理环境变量后则使用自定义代理。示例访问 [BrowserScan](https://www.browserscan.net/) 并保存返回截图。同一套浏览器流程也可访问和渲染 [Pixelscan](https://pixelscan.net/) 与 [IPhey](https://iphey.com/),可从返回页面查看网站当前的检测结果。
|
|
73
|
+
|
|
74
|
+
```ruby
|
|
75
|
+
require "adscrawl"
|
|
76
|
+
|
|
77
|
+
client = AdsCrawl::Client.new
|
|
78
|
+
server = ENV["ADSCRAWL_PROXY_SERVER"]
|
|
79
|
+
username = ENV["ADSCRAWL_PROXY_USERNAME"]
|
|
80
|
+
password = ENV["ADSCRAWL_PROXY_PASSWORD"]
|
|
81
|
+
raise "代理用户名和密码必须同时设置。" if username.nil? != password.nil?
|
|
82
|
+
raise "使用代理凭据时必须设置 ADSCRAWL_PROXY_SERVER。" if server.nil? && (username || password)
|
|
83
|
+
|
|
84
|
+
routing = {countryCode: "GLOBAL"}
|
|
85
|
+
if server
|
|
86
|
+
proxy = {server: server}
|
|
87
|
+
proxy.merge!(username: username, password: password) if username && password
|
|
88
|
+
routing = {proxy: proxy}
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
png = client.screenshot(
|
|
92
|
+
**routing,
|
|
93
|
+
url: "https://www.browserscan.net/",
|
|
94
|
+
viewport: {width: 1440, height: 900},
|
|
95
|
+
fullPage: true,
|
|
96
|
+
waitUntil: "networkidle",
|
|
97
|
+
timeoutMs: 60_000,
|
|
98
|
+
userAgentMode: "random",
|
|
99
|
+
userAgentOs: "windows",
|
|
100
|
+
fingerprint: {
|
|
101
|
+
webRtc: "forward", webGl: "random", webGpu: "random",
|
|
102
|
+
webGlImage: "random", canvas: "random", audioContext: "random",
|
|
103
|
+
clientRects: "random", speechVoices: "random", fonts: "random",
|
|
104
|
+
hardware: "random", doNotTrack: "random"
|
|
105
|
+
},
|
|
106
|
+
timeout_ms: 75_000
|
|
107
|
+
)
|
|
108
|
+
File.binwrite("browserscan.png", png)
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
可运行 `ruby examples/fingerprint.rb`,也可以把检测网站 URL 和输出路径作为参数传入。示例不会把自定义代理和托管线路组合使用。
|
|
112
|
+
|
|
113
|
+
## 结构化提取
|
|
114
|
+
|
|
115
|
+
```ruby
|
|
116
|
+
require "adscrawl"
|
|
117
|
+
|
|
118
|
+
client = AdsCrawl::Client.new
|
|
119
|
+
puts client.spa.templates["templates"]
|
|
120
|
+
|
|
121
|
+
result = client.spa.extract(
|
|
122
|
+
url: "https://www.adscrawl.net",
|
|
123
|
+
fields: {
|
|
124
|
+
title: {source: "dom", selector: "h1", value: "text", required: true}
|
|
125
|
+
}
|
|
126
|
+
)
|
|
127
|
+
puts result.dig("data", "title")
|
|
128
|
+
|
|
129
|
+
inspection = client.spa.inspect(url: "https://www.adscrawl.net")
|
|
130
|
+
puts inspection["candidates"]
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
## 远程 CDP 浏览器
|
|
134
|
+
|
|
135
|
+
```ruby
|
|
136
|
+
require "adscrawl"
|
|
137
|
+
|
|
138
|
+
client = AdsCrawl::Client.new
|
|
139
|
+
session = client.cdp.create(
|
|
140
|
+
idleTimeoutMs: 600_000,
|
|
141
|
+
maxSessionMs: 3_600_000,
|
|
142
|
+
browserSettings: {viewport: {width: 1440, height: 900}}
|
|
143
|
+
)
|
|
144
|
+
begin
|
|
145
|
+
version = client.cdp.get_version(session)
|
|
146
|
+
puts version["Browser"]
|
|
147
|
+
# 使用兼容 CDP 的 Ruby 库连接 session["cdpBaseUrl"]。
|
|
148
|
+
ensure
|
|
149
|
+
client.cdp.close(session["sessionId"])
|
|
150
|
+
end
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
`client.cdp.list` 返回活动会话。`get_version` 会校验会话 Token URL,并且不会把 API Key 转发到数据端点。CDP 连接 URL 含有密钥,请勿写入日志。
|
|
154
|
+
|
|
155
|
+
## 持久云浏览器
|
|
156
|
+
|
|
157
|
+
```ruby
|
|
158
|
+
require "adscrawl"
|
|
159
|
+
|
|
160
|
+
client = AdsCrawl::Client.new
|
|
161
|
+
launched = client.cloud_browsers.launch(
|
|
162
|
+
proxy: {server: ENV.fetch("ADSCRAWL_PROXY_SERVER")},
|
|
163
|
+
tabs: ["https://www.adscrawl.net"]
|
|
164
|
+
)
|
|
165
|
+
browser_id = launched["id"]
|
|
166
|
+
begin
|
|
167
|
+
profile = client.cloud_browsers.get(browser_id)
|
|
168
|
+
puts profile.dig("runtime", "status")
|
|
169
|
+
ensure
|
|
170
|
+
client.cloud_browsers.stop(browser_id)
|
|
171
|
+
end
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
使用 API Key 启动或快速创建云浏览器时,每次都必须传顶层自定义代理。返回 `stopping` 表示仍在关闭,需要轮询 `get` 直到状态为 `stopped`。快速创建失败时,`APIError#resource_id` 可能包含需要清理的浏览器 ID。
|
|
175
|
+
|
|
176
|
+
## 配置与错误处理
|
|
177
|
+
|
|
178
|
+
```ruby
|
|
179
|
+
require "adscrawl"
|
|
180
|
+
|
|
181
|
+
client = AdsCrawl::Client.new(base_url: "https://api.adscrawl.net", timeout_ms: 90_000)
|
|
182
|
+
begin
|
|
183
|
+
puts client.markdown(
|
|
184
|
+
{url: "https://www.adscrawl.net", timeoutMs: 60_000},
|
|
185
|
+
timeout_ms: 75_000
|
|
186
|
+
)
|
|
187
|
+
rescue AdsCrawl::APIError => error
|
|
188
|
+
puts [error.status, error.api_code, error.trace_id]
|
|
189
|
+
rescue AdsCrawl::TimeoutError
|
|
190
|
+
puts "重试前请先检查远程会话。"
|
|
191
|
+
end
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
API Key 默认读取 `ADSCRAWL_API_KEY`。基础 URL 依次读取 `ADSCRAWL_BASE_URL`、`ADSCRAWL_API_URL`,最后使用 `https://api.adscrawl.net`。错误类型包括 `APIError`、`TimeoutError`、`ConnectionError` 和 `ResponseError`。API 错误响应会进行凭据脱敏。HTTP 超时包含读取响应体的时间。SDK 不会自动重试请求。
|
|
195
|
+
|
|
196
|
+
## 开发与发布
|
|
197
|
+
|
|
198
|
+
```bash
|
|
199
|
+
bundle install
|
|
200
|
+
bundle exec rake test
|
|
201
|
+
gem build adscrawl.gemspec
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
测试只使用本机和模拟响应,不消耗 AdsCrawl 额度。发布使用 RubyGems Trusted Publishing,不保存长期注册表 Token,详见 [RELEASING.md](./RELEASING.md)。
|
|
205
|
+
|
|
206
|
+
## 许可证
|
|
207
|
+
|
|
208
|
+
MIT
|
data/RELEASING.md
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
1
|
+
# Releasing
|
|
2
|
+
|
|
3
|
+
The release workflow uses RubyGems Trusted Publishing, so the repository stores no long-lived RubyGems API key.
|
|
4
|
+
|
|
5
|
+
For the first release, create a pending trusted publisher in the RubyGems profile with:
|
|
6
|
+
|
|
7
|
+
- Gem name: `adscrawl`
|
|
8
|
+
- Repository owner: `AdsCrawl`
|
|
9
|
+
- Repository name: `adscrawl-ruby`
|
|
10
|
+
- Workflow filename: `release.yml`
|
|
11
|
+
- Environment: `release`
|
|
12
|
+
|
|
13
|
+
Then run `bundle exec rake test`, update `CHANGELOG.md`, and push a semantic version tag matching `AdsCrawl::VERSION`, such as `v0.1.0`. The workflow tests, builds, and publishes the Gem with a short-lived OIDC credential.
|
|
14
|
+
|
|
15
|
+
Published Gem versions cannot be reused. If a release needs a correction, increment the version and publish a new tag.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
<?xml version="1.0" encoding="UTF-8"?>
|
|
2
|
+
<svg id="_图层_2" data-name="图层 2" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1477.22 269.56">
|
|
3
|
+
<g id="_图层_1-2" data-name="图层 1">
|
|
4
|
+
<g>
|
|
5
|
+
<g>
|
|
6
|
+
<polygon points="126.17 164.02 32.77 238.88 40.36 171.56 43.32 141.82 17.48 134.9 4.69 261.79 33.69 269.56 156.33 172.1 126.17 164.02"/>
|
|
7
|
+
<polygon points="143.39 105.54 236.79 30.67 229.2 98 226.24 127.74 252.08 134.66 264.87 7.77 235.88 0 113.24 97.46 143.39 105.54"/>
|
|
8
|
+
<polygon points="105.54 126.16 30.67 32.77 98 40.36 127.74 43.32 134.66 17.48 7.77 4.69 0 33.68 97.46 156.32 105.54 126.16"/>
|
|
9
|
+
<polygon points="172.1 113.23 164.02 143.39 238.89 236.78 171.56 229.19 141.82 226.24 134.9 252.07 261.79 264.86 269.56 235.87 172.1 113.23"/>
|
|
10
|
+
</g>
|
|
11
|
+
<g>
|
|
12
|
+
<path d="M664.16,112.42c-3.12-3-6.7-5.54-10.75-7.59-6.84-3.46-14.54-5.19-23.1-5.19-11.2,0-21.12,2.72-29.76,8.15-8.65,5.43-15.48,12.84-20.5,22.23-5.02,9.39-7.53,20.01-7.53,31.86s2.51,22.23,7.53,31.62c5.02,9.39,11.89,16.8,20.63,22.23,8.72,5.43,18.61,8.15,29.64,8.15,8.56,0,16.3-1.81,23.22-5.43,4.13-2.16,7.74-4.8,10.87-7.9v10.86h32.11V48.32h-32.36v64.1ZM651.68,189.93c-4.53,2.72-9.84,4.08-15.93,4.08-5.77,0-10.95-1.36-15.56-4.08-4.61-2.72-8.19-6.5-10.74-11.36-2.55-4.86-3.83-10.5-3.83-16.92s1.27-11.77,3.83-16.55c2.55-4.77,6.09-8.56,10.62-11.36,4.53-2.8,9.84-4.2,15.93-4.2s11.15,1.36,15.68,4.08c4.53,2.72,8.07,6.51,10.62,11.36,2.55,4.86,3.83,10.42,3.83,16.67s-1.28,12.07-3.83,16.92c-2.56,4.86-6.1,8.65-10.62,11.36Z"/>
|
|
13
|
+
<path d="M797.68,156.5c-4.59-2.7-9.41-4.78-14.46-6.24-5.06-1.46-9.9-2.76-14.54-3.9-4.64-1.14-8.41-2.5-11.33-4.09-2.92-1.58-4.4-3.86-4.47-6.82-.05-2.63,1.18-4.72,3.7-6.25,2.52-1.53,6.17-2.35,10.94-2.45,5.27-.11,10.19.78,14.76,2.66,4.57,1.88,8.71,5.01,12.43,9.38l19.1-19.92c-5.41-6.8-12.15-11.85-20.21-15.14-8.06-3.29-17.02-4.83-26.9-4.63-9.38.2-17.5,1.93-24.35,5.2-6.85,3.27-12.07,7.75-15.66,13.42-3.59,5.68-5.3,12.3-5.14,19.87.15,7.25,1.71,13.1,4.69,17.57,2.97,4.47,6.79,7.93,11.46,10.38,4.66,2.46,9.52,4.41,14.58,5.87,5.05,1.46,9.9,2.76,14.53,3.9,4.63,1.14,8.41,2.58,11.33,4.33,2.92,1.75,4.41,4.27,4.48,7.56.06,2.96-1.25,5.21-3.93,6.75-2.69,1.54-6.66,2.36-11.93,2.47-6.59.14-12.62-.89-18.1-3.08-5.48-2.19-10.33-5.55-14.54-10.07l-18.85,19.91c4.05,4.53,8.82,8.34,14.32,11.44,5.5,3.1,11.52,5.48,18.06,7.16,6.54,1.67,13.18,2.44,19.94,2.3,14.48-.3,25.86-3.96,34.11-10.97,8.25-7.01,12.26-16.35,12.01-28.05-.15-7.24-1.72-13.14-4.69-17.69-2.98-4.55-6.76-8.17-11.34-10.88Z"/>
|
|
14
|
+
<polygon points="562.48 221.37 534.34 48.47 494.4 48.47 366.33 221.37 406.26 221.37 506.91 84.28 519 164.51 484.87 164.51 463.99 194.1 523.63 194.1 528.11 221.37 562.48 221.37"/>
|
|
15
|
+
<path d="M965.98,169.68c-10.34,12.89-26.19,21.17-43.96,21.17-31.08,0-56.37-25.29-56.37-56.37s25.29-56.37,56.37-56.37c17.77,0,33.62,8.28,43.96,21.17l23.48-23.48c-16.39-18.82-40.52-30.74-67.44-30.74-49.38,0-89.42,40.03-89.42,89.42s40.03,89.42,89.42,89.42c26.92,0,51.04-11.91,67.44-30.73l-23.48-23.48Z"/>
|
|
16
|
+
<path d="M1084.09,102.49c-4.61-1.89-9.8-2.84-15.56-2.84-13.34,0-23.55,4.24-30.63,12.72-.09.1-.16.22-.25.33v-10.58h-32.36v119.3h32.36v-65.95c0-8.89,2.18-15.52,6.55-19.88,4.36-4.36,10-6.55,16.92-6.55,3.29,0,6.21.49,8.77,1.48,2.55.99,4.73,2.47,6.55,4.45l20.25-23.22c-3.79-4.28-7.99-7.37-12.6-9.26Z"/>
|
|
17
|
+
<path d="M1191.19,113.5c-3.28-3.48-7.14-6.38-11.61-8.67-6.75-3.46-14.41-5.19-22.97-5.19-10.87,0-20.67,2.72-29.39,8.15-8.73,5.43-15.56,12.84-20.5,22.23-4.94,9.39-7.41,20.01-7.41,31.86s2.47,22.23,7.41,31.62c4.94,9.39,11.77,16.8,20.5,22.23,8.73,5.43,18.53,8.15,29.39,8.15,8.56,0,16.22-1.77,22.97-5.31,4.47-2.34,8.33-5.26,11.61-8.72v11.56h32.11v-119.3h-32.11v11.39ZM1184.52,184.99c-5.6,6.01-12.93,9.02-21.98,9.02-5.93,0-11.16-1.36-15.68-4.08-4.53-2.72-8.07-6.5-10.62-11.36-2.55-4.86-3.83-10.5-3.83-16.92s1.27-11.81,3.83-16.67c2.55-4.86,6.09-8.65,10.62-11.36,4.53-2.72,9.76-4.08,15.68-4.08s11.4,1.36,15.93,4.08c4.53,2.72,8.07,6.51,10.62,11.36,2.55,4.86,3.83,10.42,3.83,16.67,0,9.55-2.8,17.33-8.4,23.34Z"/>
|
|
18
|
+
<polygon points="1370.67 174.41 1345.82 102.12 1322.35 102.12 1297.5 174.41 1273.2 102.12 1240.59 102.12 1285.3 221.42 1308.52 221.42 1334.08 151.43 1359.65 221.42 1383.11 221.42 1427.58 102.12 1394.97 102.12 1370.67 174.41"/>
|
|
19
|
+
<rect x="1444.87" y="48.62" width="32.36" height="172.81"/>
|
|
20
|
+
</g>
|
|
21
|
+
</g>
|
|
22
|
+
</g>
|
|
23
|
+
</svg>
|
data/examples/cdp.rb
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "adscrawl"
|
|
4
|
+
|
|
5
|
+
client = AdsCrawl::Client.new
|
|
6
|
+
session = client.cdp.create(
|
|
7
|
+
idleTimeoutMs: 600_000,
|
|
8
|
+
maxSessionMs: 3_600_000,
|
|
9
|
+
browserSettings: {viewport: {width: 1440, height: 900}}
|
|
10
|
+
)
|
|
11
|
+
begin
|
|
12
|
+
version = client.cdp.get_version(session)
|
|
13
|
+
puts version["Browser"] || "Unknown browser"
|
|
14
|
+
puts "Connect a CDP-compatible Ruby library to session['cdpBaseUrl']."
|
|
15
|
+
ensure
|
|
16
|
+
client.cdp.close(session["sessionId"])
|
|
17
|
+
end
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "adscrawl"
|
|
4
|
+
|
|
5
|
+
proxy_server = ENV.fetch("ADSCRAWL_PROXY_SERVER")
|
|
6
|
+
client = AdsCrawl::Client.new
|
|
7
|
+
launched = client.cloud_browsers.launch(
|
|
8
|
+
proxy: {server: proxy_server},
|
|
9
|
+
tabs: ["https://www.adscrawl.net"]
|
|
10
|
+
)
|
|
11
|
+
browser_id = launched["id"]
|
|
12
|
+
begin
|
|
13
|
+
profile = client.cloud_browsers.get(browser_id)
|
|
14
|
+
puts "#{profile['id']} #{profile.dig('runtime', 'status')}"
|
|
15
|
+
ensure
|
|
16
|
+
client.cloud_browsers.stop(browser_id)
|
|
17
|
+
end
|
data/examples/extract.rb
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "json"
|
|
4
|
+
require "adscrawl"
|
|
5
|
+
|
|
6
|
+
client = AdsCrawl::Client.new
|
|
7
|
+
result = client.spa.extract(
|
|
8
|
+
url: "https://www.adscrawl.net",
|
|
9
|
+
fields: {
|
|
10
|
+
title: {source: "dom", selector: "h1", value: "text", required: true}
|
|
11
|
+
}
|
|
12
|
+
)
|
|
13
|
+
puts JSON.pretty_generate(
|
|
14
|
+
title: result.dig("data", "title"),
|
|
15
|
+
missingFields: result["missingFields"]
|
|
16
|
+
)
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "adscrawl"
|
|
4
|
+
|
|
5
|
+
client = AdsCrawl::Client.new
|
|
6
|
+
server = ENV["ADSCRAWL_PROXY_SERVER"]
|
|
7
|
+
username = ENV["ADSCRAWL_PROXY_USERNAME"]
|
|
8
|
+
password = ENV["ADSCRAWL_PROXY_PASSWORD"]
|
|
9
|
+
raise "Set both proxy username and password, or neither." if username.nil? != password.nil?
|
|
10
|
+
raise "Set ADSCRAWL_PROXY_SERVER with proxy credentials." if server.nil? && (username || password)
|
|
11
|
+
|
|
12
|
+
routing = {countryCode: "GLOBAL"}
|
|
13
|
+
if server
|
|
14
|
+
proxy = {server: server}
|
|
15
|
+
proxy.merge!(username: username, password: password) if username && password
|
|
16
|
+
routing = {proxy: proxy}
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
url = ARGV[0] || "https://www.browserscan.net/"
|
|
20
|
+
output = ARGV[1] || "browserscan.png"
|
|
21
|
+
png = client.screenshot(
|
|
22
|
+
**routing,
|
|
23
|
+
url: url,
|
|
24
|
+
viewport: {width: 1440, height: 900},
|
|
25
|
+
fullPage: true,
|
|
26
|
+
waitUntil: "networkidle",
|
|
27
|
+
timeoutMs: 60_000,
|
|
28
|
+
userAgentMode: "random",
|
|
29
|
+
userAgentOs: "windows",
|
|
30
|
+
fingerprint: {
|
|
31
|
+
webRtc: "forward", webGl: "random", webGpu: "random",
|
|
32
|
+
webGlImage: "random", canvas: "random", audioContext: "random",
|
|
33
|
+
clientRects: "random", speechVoices: "random", fonts: "random",
|
|
34
|
+
hardware: "random", doNotTrack: "random"
|
|
35
|
+
},
|
|
36
|
+
timeout_ms: 75_000
|
|
37
|
+
)
|
|
38
|
+
File.binwrite(output, png)
|
|
39
|
+
puts "Saved fingerprint-check screenshot to #{output}"
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "adscrawl"
|
|
4
|
+
|
|
5
|
+
client = AdsCrawl::Client.new
|
|
6
|
+
url = ARGV[0] || "https://www.adscrawl.net"
|
|
7
|
+
output = ARGV[1] || "page.png"
|
|
8
|
+
png = client.screenshot(url: url, fullPage: true)
|
|
9
|
+
File.binwrite(output, png)
|
|
10
|
+
puts "Saved screenshot to #{output}"
|
|
@@ -0,0 +1,342 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "json"
|
|
4
|
+
require "net/http"
|
|
5
|
+
require "openssl"
|
|
6
|
+
require "socket"
|
|
7
|
+
require "timeout"
|
|
8
|
+
require "uri"
|
|
9
|
+
|
|
10
|
+
require_relative "errors"
|
|
11
|
+
require_relative "http_response"
|
|
12
|
+
require_relative "service/spa"
|
|
13
|
+
require_relative "service/cdp"
|
|
14
|
+
require_relative "service/cloud_browsers"
|
|
15
|
+
|
|
16
|
+
module AdsCrawl
|
|
17
|
+
class Client
|
|
18
|
+
DEFAULT_BASE_URL = "https://api.adscrawl.net"
|
|
19
|
+
DEFAULT_TIMEOUT_MS = 90_000
|
|
20
|
+
LAUNCH_TIMEOUT_MS = 195_000
|
|
21
|
+
MAX_TIMEOUT_MS = 2_147_483_647
|
|
22
|
+
MAX_BODY_BYTES = 64 * 1024 * 1024
|
|
23
|
+
PNG_SIGNATURE = "\x89PNG\r\n\x1a\n".b
|
|
24
|
+
RUNTIME_STATUSES = %w[starting running stopping stopped].freeze
|
|
25
|
+
|
|
26
|
+
attr_reader :spa, :cdp, :cloud_browsers
|
|
27
|
+
alias cloudBrowsers cloud_browsers
|
|
28
|
+
|
|
29
|
+
# Construction never performs a network request.
|
|
30
|
+
# A custom transport receives keyword arguments and returns HttpResponse.
|
|
31
|
+
def initialize(api_key: nil, base_url: nil, timeout_ms: nil, transport: nil)
|
|
32
|
+
selected_key = (api_key || ENV["ADSCRAWL_API_KEY"]).to_s.strip
|
|
33
|
+
raise ArgumentError, "Set ADSCRAWL_API_KEY or pass api_key to Client." if selected_key.empty?
|
|
34
|
+
raise ArgumentError, "api_key must not contain line breaks." if selected_key.match?(/[\r\n]/)
|
|
35
|
+
|
|
36
|
+
selected_url = base_url || ENV["ADSCRAWL_BASE_URL"] || ENV["ADSCRAWL_API_URL"] || DEFAULT_BASE_URL
|
|
37
|
+
@base_uri = validate_base_url(selected_url)
|
|
38
|
+
@base_url = selected_url.sub(%r{/+\z}, "")
|
|
39
|
+
@api_key = selected_key
|
|
40
|
+
@timeout_ms = validate_timeout(timeout_ms || DEFAULT_TIMEOUT_MS)
|
|
41
|
+
@custom_timeout = !timeout_ms.nil?
|
|
42
|
+
@transport = transport
|
|
43
|
+
|
|
44
|
+
@spa = Service::SPA.new(self)
|
|
45
|
+
@cdp = Service::CDP.new(self)
|
|
46
|
+
@cloud_browsers = Service::CloudBrowsers.new(self)
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
def html(options = nil, timeout_ms: nil, **keywords)
|
|
50
|
+
payload = options_hash(options, keywords)
|
|
51
|
+
validate_routing(payload)
|
|
52
|
+
body = send_request("POST", "/html", payload.merge("contentMode" => "html"), timeout_ms: timeout_ms)
|
|
53
|
+
raise ResponseError, "AdsCrawl returned empty page content." if body.strip.empty?
|
|
54
|
+
|
|
55
|
+
body.dup.force_encoding(Encoding::UTF_8)
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
def markdown(options = nil, timeout_ms: nil, **keywords)
|
|
59
|
+
payload = options_hash(options, keywords)
|
|
60
|
+
validate_routing(payload)
|
|
61
|
+
body = send_request("POST", "/html", payload.merge("contentMode" => "markdown"), timeout_ms: timeout_ms)
|
|
62
|
+
raise ResponseError, "AdsCrawl returned empty page content." if body.strip.empty?
|
|
63
|
+
|
|
64
|
+
body.dup.force_encoding(Encoding::UTF_8)
|
|
65
|
+
end
|
|
66
|
+
|
|
67
|
+
def article(options = nil, timeout_ms: nil, **keywords)
|
|
68
|
+
payload = options_hash(options, keywords)
|
|
69
|
+
validate_routing(payload)
|
|
70
|
+
request_json(
|
|
71
|
+
"POST",
|
|
72
|
+
"/html",
|
|
73
|
+
payload.merge("contentMode" => "json"),
|
|
74
|
+
timeout_ms: timeout_ms,
|
|
75
|
+
validate: lambda { |value|
|
|
76
|
+
string?(value, "title") && string?(value, "content") &&
|
|
77
|
+
string?(value, "textContent") && value["length"].is_a?(Numeric)
|
|
78
|
+
}
|
|
79
|
+
)
|
|
80
|
+
end
|
|
81
|
+
|
|
82
|
+
def screenshot(options = nil, timeout_ms: nil, **keywords)
|
|
83
|
+
payload = options_hash(options, keywords)
|
|
84
|
+
validate_routing(payload)
|
|
85
|
+
body = send_request("POST", "/screenshot", payload, timeout_ms: timeout_ms)
|
|
86
|
+
raise ResponseError, "AdsCrawl returned an invalid PNG screenshot." unless body.b.start_with?(PNG_SIGNATURE)
|
|
87
|
+
|
|
88
|
+
body.b
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
# Internal API used by service objects.
|
|
92
|
+
def send_request(method, path, body = nil, timeout_ms: nil, authenticate: true, extra_secrets: [])
|
|
93
|
+
deadline = validate_timeout(timeout_ms || @timeout_ms)
|
|
94
|
+
payload = body.nil? ? nil : normalize_hash(body)
|
|
95
|
+
secrets = [@api_key, *collect_secrets(payload), *extra_secrets].compact
|
|
96
|
+
encoded = payload.nil? ? nil : JSON.generate(payload)
|
|
97
|
+
headers = {"Accept" => "*/*", "User-Agent" => "adscrawl-ruby/#{VERSION}"}
|
|
98
|
+
headers["X-API-Key"] = @api_key if authenticate
|
|
99
|
+
headers["Content-Type"] = "application/json" unless encoded.nil?
|
|
100
|
+
|
|
101
|
+
response = if @transport
|
|
102
|
+
@transport.call(method: method, url: "#{@base_url}#{path}", headers: headers, body: encoded, timeout_ms: deadline)
|
|
103
|
+
else
|
|
104
|
+
net_http_request(method, "#{@base_url}#{path}", headers, encoded, deadline)
|
|
105
|
+
end
|
|
106
|
+
raise ResponseError, "The custom transport must return an AdsCrawl::HttpResponse." unless response.is_a?(HttpResponse)
|
|
107
|
+
raise ResponseError, "AdsCrawl response exceeds the 64 MiB safety limit." if response.body.bytesize > MAX_BODY_BYTES
|
|
108
|
+
|
|
109
|
+
status = Integer(response.status)
|
|
110
|
+
return response.body if status.between?(200, 299)
|
|
111
|
+
|
|
112
|
+
raise_api_error(status, response.body, response.headers || {}, secrets)
|
|
113
|
+
end
|
|
114
|
+
|
|
115
|
+
# Internal API used by service objects.
|
|
116
|
+
def request_json(method, path, body = nil, timeout_ms: nil, authenticate: true, validate: nil, extra_secrets: [])
|
|
117
|
+
raw = send_request(
|
|
118
|
+
method,
|
|
119
|
+
path,
|
|
120
|
+
body,
|
|
121
|
+
timeout_ms: timeout_ms,
|
|
122
|
+
authenticate: authenticate,
|
|
123
|
+
extra_secrets: extra_secrets
|
|
124
|
+
)
|
|
125
|
+
value = JSON.parse(raw)
|
|
126
|
+
raise ResponseError, "AdsCrawl returned invalid or unexpected JSON." unless value.is_a?(Hash)
|
|
127
|
+
if validate && !validate.call(value)
|
|
128
|
+
raise ResponseError, "AdsCrawl returned an unexpected JSON response shape."
|
|
129
|
+
end
|
|
130
|
+
|
|
131
|
+
value
|
|
132
|
+
rescue JSON::ParserError
|
|
133
|
+
raise ResponseError, "AdsCrawl returned invalid or unexpected JSON."
|
|
134
|
+
end
|
|
135
|
+
|
|
136
|
+
def launch_timeout_ms
|
|
137
|
+
@custom_timeout ? nil : LAUNCH_TIMEOUT_MS
|
|
138
|
+
end
|
|
139
|
+
|
|
140
|
+
def base_uri
|
|
141
|
+
@base_uri.dup
|
|
142
|
+
end
|
|
143
|
+
|
|
144
|
+
def normalize_hash(value)
|
|
145
|
+
raise ArgumentError, "options must be a Hash." unless value.is_a?(Hash)
|
|
146
|
+
|
|
147
|
+
normalize_value(value)
|
|
148
|
+
end
|
|
149
|
+
|
|
150
|
+
def options_hash(options, keywords)
|
|
151
|
+
if options && !keywords.empty?
|
|
152
|
+
raise ArgumentError, "Pass options as a Hash or keywords, not both."
|
|
153
|
+
end
|
|
154
|
+
|
|
155
|
+
normalize_hash(options || keywords)
|
|
156
|
+
end
|
|
157
|
+
|
|
158
|
+
def validate_routing(value)
|
|
159
|
+
return unless value.is_a?(Hash)
|
|
160
|
+
|
|
161
|
+
if value.key?("proxy") && !value["proxy"].nil? && value.key?("countryCode") && !value["countryCode"].to_s.empty?
|
|
162
|
+
raise ArgumentError, "proxy and countryCode cannot be combined."
|
|
163
|
+
end
|
|
164
|
+
validate_routing(value["browserSettings"]) if value["browserSettings"].is_a?(Hash)
|
|
165
|
+
end
|
|
166
|
+
|
|
167
|
+
def encode_id(value)
|
|
168
|
+
raise ArgumentError, "A non-empty resource id is required." unless value.is_a?(String) && !value.strip.empty? && !%w[. ..].include?(value)
|
|
169
|
+
|
|
170
|
+
URI.encode_www_form_component(value).tr("+", "%20")
|
|
171
|
+
end
|
|
172
|
+
|
|
173
|
+
def ok?(value)
|
|
174
|
+
value.is_a?(Hash) && value["ok"] == true
|
|
175
|
+
end
|
|
176
|
+
|
|
177
|
+
def session?(value)
|
|
178
|
+
value.is_a?(Hash) && string?(value, "sessionId") && string?(value, "expiresAt") && string?(value, "cdpBaseUrl")
|
|
179
|
+
end
|
|
180
|
+
|
|
181
|
+
def runtime?(value)
|
|
182
|
+
value.is_a?(Hash) && RUNTIME_STATUSES.include?(value["status"])
|
|
183
|
+
end
|
|
184
|
+
|
|
185
|
+
def cloud_browser?(value)
|
|
186
|
+
value.is_a?(Hash) && string?(value, "id") && runtime?(value["runtime"]) && value["browserSettings"].is_a?(Hash)
|
|
187
|
+
end
|
|
188
|
+
|
|
189
|
+
def page?(value)
|
|
190
|
+
value.is_a?(Hash) && string?(value, "url") && string?(value, "title")
|
|
191
|
+
end
|
|
192
|
+
|
|
193
|
+
def string?(value, key)
|
|
194
|
+
value.is_a?(Hash) && value[key].is_a?(String)
|
|
195
|
+
end
|
|
196
|
+
|
|
197
|
+
private
|
|
198
|
+
|
|
199
|
+
def validate_base_url(value)
|
|
200
|
+
uri = URI.parse(value.to_s)
|
|
201
|
+
unless %w[http https].include?(uri.scheme) && uri.host && !uri.userinfo && !uri.query && !uri.fragment
|
|
202
|
+
raise ArgumentError, "base_url must be an HTTP(S) URL without credentials, query, or fragment."
|
|
203
|
+
end
|
|
204
|
+
|
|
205
|
+
uri
|
|
206
|
+
rescue URI::InvalidURIError
|
|
207
|
+
raise ArgumentError, "base_url must be an HTTP(S) URL without credentials, query, or fragment."
|
|
208
|
+
end
|
|
209
|
+
|
|
210
|
+
def validate_timeout(value)
|
|
211
|
+
unless value.is_a?(Integer) && value.positive? && value <= MAX_TIMEOUT_MS
|
|
212
|
+
raise ArgumentError, "timeout_ms must be a positive Integer no greater than #{MAX_TIMEOUT_MS}."
|
|
213
|
+
end
|
|
214
|
+
|
|
215
|
+
value
|
|
216
|
+
end
|
|
217
|
+
|
|
218
|
+
def normalize_value(value)
|
|
219
|
+
case value
|
|
220
|
+
when Hash
|
|
221
|
+
value.each_with_object({}) { |(key, item), result| result[key.to_s] = normalize_value(item) }
|
|
222
|
+
when Array
|
|
223
|
+
value.map { |item| normalize_value(item) }
|
|
224
|
+
else
|
|
225
|
+
value
|
|
226
|
+
end
|
|
227
|
+
end
|
|
228
|
+
|
|
229
|
+
def net_http_request(method, url, headers, body, timeout_ms)
|
|
230
|
+
uri = URI.parse(url)
|
|
231
|
+
request_class = {
|
|
232
|
+
"GET" => Net::HTTP::Get,
|
|
233
|
+
"POST" => Net::HTTP::Post,
|
|
234
|
+
"DELETE" => Net::HTTP::Delete
|
|
235
|
+
}.fetch(method)
|
|
236
|
+
request = request_class.new(uri.request_uri, headers)
|
|
237
|
+
request.body = body unless body.nil?
|
|
238
|
+
timeout_seconds = timeout_ms / 1000.0
|
|
239
|
+
response_headers = {}
|
|
240
|
+
response_body = String.new(encoding: Encoding::BINARY)
|
|
241
|
+
status = nil
|
|
242
|
+
|
|
243
|
+
http = Net::HTTP.new(uri.host, uri.port)
|
|
244
|
+
http.use_ssl = uri.scheme == "https"
|
|
245
|
+
http.open_timeout = [timeout_seconds, 10.0].min
|
|
246
|
+
http.read_timeout = timeout_seconds
|
|
247
|
+
http.write_timeout = timeout_seconds if http.respond_to?(:write_timeout=)
|
|
248
|
+
http.request(request) do |response|
|
|
249
|
+
status = response.code.to_i
|
|
250
|
+
response.each_header { |name, value| response_headers[name.downcase] = value }
|
|
251
|
+
response.read_body do |chunk|
|
|
252
|
+
if response_body.bytesize + chunk.bytesize > MAX_BODY_BYTES
|
|
253
|
+
raise ResponseError, "AdsCrawl response exceeds the 64 MiB safety limit."
|
|
254
|
+
end
|
|
255
|
+
response_body << chunk.b
|
|
256
|
+
end
|
|
257
|
+
end
|
|
258
|
+
|
|
259
|
+
HttpResponse.new(status: status, headers: response_headers, body: response_body)
|
|
260
|
+
rescue Net::OpenTimeout, Net::ReadTimeout, Net::WriteTimeout, Timeout::Error
|
|
261
|
+
raise TimeoutError, timeout_ms
|
|
262
|
+
rescue ResponseError
|
|
263
|
+
raise
|
|
264
|
+
rescue SocketError, OpenSSL::SSL::SSLError, EOFError, IOError, SystemCallError, URI::InvalidURIError
|
|
265
|
+
raise ConnectionError
|
|
266
|
+
end
|
|
267
|
+
|
|
268
|
+
def raise_api_error(status, raw, headers, secrets)
|
|
269
|
+
parsed = JSON.parse(raw)
|
|
270
|
+
rescue JSON::ParserError
|
|
271
|
+
parsed = raw.to_s
|
|
272
|
+
ensure
|
|
273
|
+
safe_body = redact_value(parsed, secrets)
|
|
274
|
+
message = "AdsCrawl API returned HTTP #{status}"
|
|
275
|
+
api_code = resource_id = trace_id = nil
|
|
276
|
+
if safe_body.is_a?(Hash)
|
|
277
|
+
message = safe_body["error"] if safe_body["error"].is_a?(String)
|
|
278
|
+
api_code = safe_body["code"] if safe_body["code"].is_a?(String)
|
|
279
|
+
resource_id = safe_body["id"] if safe_body["id"].is_a?(String)
|
|
280
|
+
trace_id = safe_body["traceId"] if safe_body["traceId"].is_a?(String)
|
|
281
|
+
end
|
|
282
|
+
normalized_headers = headers.each_with_object({}) { |(key, value), result| result[key.to_s.downcase] = value.to_s }
|
|
283
|
+
trace_id = redact_string(normalized_headers["x-trace-id"], secrets) if normalized_headers["x-trace-id"]
|
|
284
|
+
request_id = normalized_headers["x-request-id"] && redact_string(normalized_headers["x-request-id"], secrets)
|
|
285
|
+
raise APIError.new(
|
|
286
|
+
message.to_s.slice(0, 2048),
|
|
287
|
+
status: status,
|
|
288
|
+
api_code: api_code,
|
|
289
|
+
resource_id: resource_id,
|
|
290
|
+
trace_id: trace_id,
|
|
291
|
+
request_id: request_id,
|
|
292
|
+
body: safe_body
|
|
293
|
+
)
|
|
294
|
+
end
|
|
295
|
+
|
|
296
|
+
def collect_secrets(value, result = [])
|
|
297
|
+
case value
|
|
298
|
+
when Hash
|
|
299
|
+
value.each do |key, item|
|
|
300
|
+
if key == "cookies" && item.is_a?(Array)
|
|
301
|
+
item.each { |cookie| result << cookie["value"] if cookie.is_a?(Hash) && cookie["value"].is_a?(String) && !cookie["value"].empty? }
|
|
302
|
+
end
|
|
303
|
+
if %w[password username token controltoken].include?(key.downcase) && item.is_a?(String) && !item.empty?
|
|
304
|
+
result << item
|
|
305
|
+
end
|
|
306
|
+
collect_secrets(item, result)
|
|
307
|
+
end
|
|
308
|
+
when Array
|
|
309
|
+
value.each { |item| collect_secrets(item, result) }
|
|
310
|
+
end
|
|
311
|
+
result
|
|
312
|
+
end
|
|
313
|
+
|
|
314
|
+
def redact_string(value, secrets)
|
|
315
|
+
output = value.to_s.dup
|
|
316
|
+
secrets.each { |secret| output.gsub!(secret, "[REDACTED]") unless secret.to_s.empty? }
|
|
317
|
+
output.gsub!(/([?&](?:token|controlToken|apiKey|api_key)=)[^&#\s"<>]+/i, '\\1[REDACTED]')
|
|
318
|
+
output.gsub!(/(\b(?:https?|socks5):\/\/)[^\s\/@]+:[^\s\/@]+@/i, '\\1[REDACTED]@')
|
|
319
|
+
output
|
|
320
|
+
end
|
|
321
|
+
|
|
322
|
+
def redact_value(value, secrets)
|
|
323
|
+
case value
|
|
324
|
+
when String
|
|
325
|
+
redact_string(value, secrets)
|
|
326
|
+
when Array
|
|
327
|
+
value.map { |item| redact_value(item, secrets) }
|
|
328
|
+
when Hash
|
|
329
|
+
value.each_with_object({}) do |(key, item), result|
|
|
330
|
+
normalized = key.to_s.downcase.delete("-_")
|
|
331
|
+
result[key] = if %w[apikey xapikey authorization password username cookie cookies token controltoken websocketdebuggerurl cdpbaseurl controlurl connecturl].include?(normalized)
|
|
332
|
+
"[REDACTED]"
|
|
333
|
+
else
|
|
334
|
+
redact_value(item, secrets)
|
|
335
|
+
end
|
|
336
|
+
end
|
|
337
|
+
else
|
|
338
|
+
value
|
|
339
|
+
end
|
|
340
|
+
end
|
|
341
|
+
end
|
|
342
|
+
end
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module AdsCrawl
|
|
4
|
+
class Error < StandardError; end
|
|
5
|
+
|
|
6
|
+
class APIError < Error
|
|
7
|
+
attr_reader :status, :api_code, :resource_id, :trace_id, :request_id, :body
|
|
8
|
+
|
|
9
|
+
def initialize(message, status:, api_code: nil, resource_id: nil, trace_id: nil, request_id: nil, body: nil)
|
|
10
|
+
@status = status
|
|
11
|
+
@api_code = api_code
|
|
12
|
+
@resource_id = resource_id
|
|
13
|
+
@trace_id = trace_id
|
|
14
|
+
@request_id = request_id
|
|
15
|
+
@body = body
|
|
16
|
+
super(message)
|
|
17
|
+
end
|
|
18
|
+
end
|
|
19
|
+
|
|
20
|
+
class TimeoutError < Error
|
|
21
|
+
attr_reader :timeout_ms
|
|
22
|
+
|
|
23
|
+
def initialize(timeout_ms)
|
|
24
|
+
@timeout_ms = timeout_ms
|
|
25
|
+
super("AdsCrawl request timed out after #{timeout_ms}ms. Remote work may still be running; inspect sessions before retrying.")
|
|
26
|
+
end
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
class ConnectionError < Error
|
|
30
|
+
def initialize
|
|
31
|
+
super("Unable to complete the AdsCrawl request. Check your network and service availability.")
|
|
32
|
+
end
|
|
33
|
+
end
|
|
34
|
+
|
|
35
|
+
class ResponseError < Error; end
|
|
36
|
+
end
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module AdsCrawl
|
|
4
|
+
module Service
|
|
5
|
+
class CDP
|
|
6
|
+
def initialize(client)
|
|
7
|
+
@client = client
|
|
8
|
+
end
|
|
9
|
+
|
|
10
|
+
def create(options = nil, timeout_ms: nil, **keywords)
|
|
11
|
+
payload = @client.options_hash(options, keywords)
|
|
12
|
+
@client.validate_routing(payload)
|
|
13
|
+
@client.request_json(
|
|
14
|
+
"POST",
|
|
15
|
+
"/cdp/sessions",
|
|
16
|
+
payload,
|
|
17
|
+
timeout_ms: timeout_ms,
|
|
18
|
+
validate: ->(value) { @client.session?(value) }
|
|
19
|
+
)
|
|
20
|
+
end
|
|
21
|
+
|
|
22
|
+
def list(timeout_ms: nil)
|
|
23
|
+
@client.request_json(
|
|
24
|
+
"GET",
|
|
25
|
+
"/cdp/sessions",
|
|
26
|
+
timeout_ms: timeout_ms,
|
|
27
|
+
validate: lambda { |value|
|
|
28
|
+
@client.ok?(value) && value["data"].is_a?(Array) && value["data"].all? { |session| @client.session?(session) }
|
|
29
|
+
}
|
|
30
|
+
)
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
def close(session_id, timeout_ms: nil)
|
|
34
|
+
id = @client.encode_id(session_id)
|
|
35
|
+
@client.request_json(
|
|
36
|
+
"DELETE",
|
|
37
|
+
"/cdp/sessions/#{id}",
|
|
38
|
+
timeout_ms: timeout_ms,
|
|
39
|
+
validate: ->(value) { @client.ok?(value) }
|
|
40
|
+
)
|
|
41
|
+
end
|
|
42
|
+
|
|
43
|
+
# Discovers the WebSocket endpoint without forwarding the API key.
|
|
44
|
+
def get_version(session, timeout_ms: nil)
|
|
45
|
+
payload = @client.normalize_hash(session)
|
|
46
|
+
session_id = payload["sessionId"]
|
|
47
|
+
cdp_base_url = payload["cdpBaseUrl"]
|
|
48
|
+
unless session_id.is_a?(String) && cdp_base_url.is_a?(String)
|
|
49
|
+
raise ArgumentError, "The session must contain sessionId and a valid cdpBaseUrl."
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
id = @client.encode_id(session_id)
|
|
53
|
+
parsed = URI.parse(cdp_base_url)
|
|
54
|
+
base = @client.base_uri
|
|
55
|
+
token = URI.decode_www_form(parsed.query.to_s).to_h["token"]
|
|
56
|
+
expected_path = "#{base.path.to_s.sub(%r{/+\z}, "")}/cdp/sessions/#{id}"
|
|
57
|
+
unless parsed.scheme == base.scheme && parsed.host == base.host && parsed.port == base.port &&
|
|
58
|
+
parsed.path.sub(%r{/+\z}, "") == expected_path && !parsed.userinfo && !parsed.fragment &&
|
|
59
|
+
token.is_a?(String) && !token.empty?
|
|
60
|
+
raise ArgumentError, "cdpBaseUrl must belong to this API origin and session and include a data token."
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
path = "/cdp/sessions/#{id}/json/version?#{URI.encode_www_form(token: token)}"
|
|
64
|
+
@client.request_json(
|
|
65
|
+
"GET",
|
|
66
|
+
path,
|
|
67
|
+
timeout_ms: timeout_ms,
|
|
68
|
+
authenticate: false,
|
|
69
|
+
extra_secrets: [token],
|
|
70
|
+
validate: ->(value) { @client.string?(value, "Browser") && @client.string?(value, "webSocketDebuggerUrl") }
|
|
71
|
+
)
|
|
72
|
+
rescue URI::InvalidURIError, ArgumentError => error
|
|
73
|
+
raise error if error.message.start_with?("cdpBaseUrl")
|
|
74
|
+
|
|
75
|
+
raise ArgumentError, "The session must contain sessionId and a valid cdpBaseUrl."
|
|
76
|
+
end
|
|
77
|
+
|
|
78
|
+
def live_token(session_id, timeout_ms: nil)
|
|
79
|
+
@client.encode_id(session_id)
|
|
80
|
+
@client.request_json(
|
|
81
|
+
"POST",
|
|
82
|
+
"/cdp/live-token",
|
|
83
|
+
{"sessionId" => session_id},
|
|
84
|
+
timeout_ms: timeout_ms,
|
|
85
|
+
validate: lambda { |value|
|
|
86
|
+
@client.ok?(value) && @client.string?(value, "controlUrl") && value["expiresAt"].is_a?(Numeric)
|
|
87
|
+
}
|
|
88
|
+
)
|
|
89
|
+
end
|
|
90
|
+
end
|
|
91
|
+
end
|
|
92
|
+
end
|
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module AdsCrawl
|
|
4
|
+
module Service
|
|
5
|
+
class CloudBrowsers
|
|
6
|
+
def initialize(client)
|
|
7
|
+
@client = client
|
|
8
|
+
end
|
|
9
|
+
|
|
10
|
+
def list(options = nil, timeout_ms: nil, **keywords)
|
|
11
|
+
payload = @client.options_hash(options, keywords)
|
|
12
|
+
query = payload.slice("page", "pageSize")
|
|
13
|
+
path = query.empty? ? "/cloud-browsers" : "/cloud-browsers?#{URI.encode_www_form(query)}"
|
|
14
|
+
@client.request_json(
|
|
15
|
+
"GET",
|
|
16
|
+
path,
|
|
17
|
+
timeout_ms: timeout_ms,
|
|
18
|
+
validate: lambda { |value|
|
|
19
|
+
@client.ok?(value) && value["data"].is_a?(Array) &&
|
|
20
|
+
value["data"].all? { |browser| @client.cloud_browser?(browser) } &&
|
|
21
|
+
value["pagination"].is_a?(Hash) &&
|
|
22
|
+
%w[limit runningLimit runningCount].all? { |key| value[key].is_a?(Integer) }
|
|
23
|
+
}
|
|
24
|
+
)
|
|
25
|
+
end
|
|
26
|
+
|
|
27
|
+
def create(options = nil, timeout_ms: nil, **keywords)
|
|
28
|
+
payload = @client.options_hash(options, keywords)
|
|
29
|
+
@client.validate_routing(payload)
|
|
30
|
+
@client.request_json(
|
|
31
|
+
"POST",
|
|
32
|
+
"/cloud-browsers",
|
|
33
|
+
payload,
|
|
34
|
+
timeout_ms: timeout_ms,
|
|
35
|
+
validate: ->(value) { @client.ok?(value) && @client.string?(value, "id") }
|
|
36
|
+
)
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
def get(browser_id, timeout_ms: nil)
|
|
40
|
+
id = @client.encode_id(browser_id)
|
|
41
|
+
@client.request_json(
|
|
42
|
+
"GET",
|
|
43
|
+
"/cloud-browsers/#{id}",
|
|
44
|
+
timeout_ms: timeout_ms,
|
|
45
|
+
validate: ->(value) { @client.cloud_browser?(value) }
|
|
46
|
+
)
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
def start(browser_id, options = nil, timeout_ms: nil, **keywords)
|
|
50
|
+
id = @client.encode_id(browser_id)
|
|
51
|
+
payload = require_proxy(@client.options_hash(options, keywords), "starts")
|
|
52
|
+
@client.request_json(
|
|
53
|
+
"POST",
|
|
54
|
+
"/cloud-browsers/#{id}/start",
|
|
55
|
+
payload,
|
|
56
|
+
timeout_ms: timeout_ms,
|
|
57
|
+
validate: ->(value) { @client.ok?(value) && @client.runtime?(value["runtime"]) }
|
|
58
|
+
)
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
def stop(browser_id, timeout_ms: nil)
|
|
62
|
+
id = @client.encode_id(browser_id)
|
|
63
|
+
@client.request_json(
|
|
64
|
+
"POST",
|
|
65
|
+
"/cloud-browsers/#{id}/stop",
|
|
66
|
+
timeout_ms: timeout_ms,
|
|
67
|
+
validate: ->(value) { @client.ok?(value) && @client.runtime?(value["runtime"]) }
|
|
68
|
+
)
|
|
69
|
+
end
|
|
70
|
+
|
|
71
|
+
def launch(options = nil, timeout_ms: nil, **keywords)
|
|
72
|
+
payload = require_proxy(@client.options_hash(options, keywords), "launches")
|
|
73
|
+
deadline = timeout_ms || @client.launch_timeout_ms
|
|
74
|
+
@client.request_json(
|
|
75
|
+
"POST",
|
|
76
|
+
"/cloud-browsers/launch",
|
|
77
|
+
payload,
|
|
78
|
+
timeout_ms: deadline,
|
|
79
|
+
validate: lambda { |value|
|
|
80
|
+
@client.ok?(value) && @client.string?(value, "id") && value["source"] == "launch" &&
|
|
81
|
+
value["deleteOnStop"] == false && @client.runtime?(value["runtime"])
|
|
82
|
+
}
|
|
83
|
+
)
|
|
84
|
+
end
|
|
85
|
+
|
|
86
|
+
private
|
|
87
|
+
|
|
88
|
+
def require_proxy(options, action)
|
|
89
|
+
payload = @client.normalize_hash(options)
|
|
90
|
+
unless payload["proxy"].is_a?(Hash)
|
|
91
|
+
raise ArgumentError, "Cloud browser #{action} require an explicit top-level proxy."
|
|
92
|
+
end
|
|
93
|
+
|
|
94
|
+
payload
|
|
95
|
+
end
|
|
96
|
+
end
|
|
97
|
+
end
|
|
98
|
+
end
|
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module AdsCrawl
|
|
4
|
+
module Service
|
|
5
|
+
class SPA
|
|
6
|
+
def initialize(client)
|
|
7
|
+
@client = client
|
|
8
|
+
end
|
|
9
|
+
|
|
10
|
+
def templates(timeout_ms: nil)
|
|
11
|
+
@client.request_json(
|
|
12
|
+
"GET",
|
|
13
|
+
"/spa-extract/templates",
|
|
14
|
+
timeout_ms: timeout_ms,
|
|
15
|
+
validate: lambda { |value|
|
|
16
|
+
value["templates"].is_a?(Array) && value["templates"].all? do |template|
|
|
17
|
+
@client.string?(template, "id") && @client.string?(template, "name")
|
|
18
|
+
end
|
|
19
|
+
}
|
|
20
|
+
)
|
|
21
|
+
end
|
|
22
|
+
|
|
23
|
+
def extract(options = nil, timeout_ms: nil, **keywords)
|
|
24
|
+
payload = @client.options_hash(options, keywords)
|
|
25
|
+
@client.validate_routing(payload)
|
|
26
|
+
@client.request_json(
|
|
27
|
+
"POST",
|
|
28
|
+
"/spa-extract",
|
|
29
|
+
payload.merge("mode" => "extract"),
|
|
30
|
+
timeout_ms: timeout_ms,
|
|
31
|
+
validate: lambda { |value|
|
|
32
|
+
value["mode"] == "extract" && @client.page?(value["page"]) &&
|
|
33
|
+
value["data"].is_a?(Hash) && value["missingFields"].is_a?(Array)
|
|
34
|
+
}
|
|
35
|
+
)
|
|
36
|
+
end
|
|
37
|
+
|
|
38
|
+
def inspect(options = nil, timeout_ms: nil, **keywords)
|
|
39
|
+
payload = @client.options_hash(options, keywords)
|
|
40
|
+
@client.validate_routing(payload)
|
|
41
|
+
@client.request_json(
|
|
42
|
+
"POST",
|
|
43
|
+
"/spa-extract",
|
|
44
|
+
payload.merge("mode" => "inspect"),
|
|
45
|
+
timeout_ms: timeout_ms,
|
|
46
|
+
validate: lambda { |value|
|
|
47
|
+
value["mode"] == "inspect" && @client.page?(value["page"]) &&
|
|
48
|
+
value["candidates"].is_a?(Hash) && value["suggestedPlan"].is_a?(Hash)
|
|
49
|
+
}
|
|
50
|
+
)
|
|
51
|
+
end
|
|
52
|
+
end
|
|
53
|
+
end
|
|
54
|
+
end
|
data/lib/adscrawl.rb
ADDED
metadata
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
--- !ruby/object:Gem::Specification
|
|
2
|
+
name: adscrawl
|
|
3
|
+
version: !ruby/object:Gem::Version
|
|
4
|
+
version: 0.1.0
|
|
5
|
+
platform: ruby
|
|
6
|
+
authors:
|
|
7
|
+
- AdsCrawl
|
|
8
|
+
bindir: bin
|
|
9
|
+
cert_chain: []
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
11
|
+
dependencies: []
|
|
12
|
+
description: Render pages in real browsers, extract structured data, capture screenshots,
|
|
13
|
+
open CDP sessions, and manage cloud browsers.
|
|
14
|
+
email:
|
|
15
|
+
- support@adscrawl.net
|
|
16
|
+
executables: []
|
|
17
|
+
extensions: []
|
|
18
|
+
extra_rdoc_files: []
|
|
19
|
+
files:
|
|
20
|
+
- CHANGELOG.md
|
|
21
|
+
- LICENSE
|
|
22
|
+
- README.md
|
|
23
|
+
- README.zh-CN.md
|
|
24
|
+
- RELEASING.md
|
|
25
|
+
- assets/adscrawl-logo.svg
|
|
26
|
+
- examples/cdp.rb
|
|
27
|
+
- examples/cloud_browser.rb
|
|
28
|
+
- examples/extract.rb
|
|
29
|
+
- examples/fingerprint.rb
|
|
30
|
+
- examples/markdown.rb
|
|
31
|
+
- examples/screenshot.rb
|
|
32
|
+
- lib/adscrawl.rb
|
|
33
|
+
- lib/adscrawl/client.rb
|
|
34
|
+
- lib/adscrawl/errors.rb
|
|
35
|
+
- lib/adscrawl/http_response.rb
|
|
36
|
+
- lib/adscrawl/service/cdp.rb
|
|
37
|
+
- lib/adscrawl/service/cloud_browsers.rb
|
|
38
|
+
- lib/adscrawl/service/spa.rb
|
|
39
|
+
- lib/adscrawl/version.rb
|
|
40
|
+
homepage: https://www.adscrawl.net
|
|
41
|
+
licenses:
|
|
42
|
+
- MIT
|
|
43
|
+
metadata:
|
|
44
|
+
allowed_push_host: https://rubygems.org
|
|
45
|
+
bug_tracker_uri: https://github.com/AdsCrawl/adscrawl-ruby/issues
|
|
46
|
+
changelog_uri: https://github.com/AdsCrawl/adscrawl-ruby/blob/main/CHANGELOG.md
|
|
47
|
+
documentation_uri: https://www.adscrawl.net/docs/
|
|
48
|
+
homepage_uri: https://www.adscrawl.net
|
|
49
|
+
rubygems_mfa_required: 'true'
|
|
50
|
+
source_code_uri: https://github.com/AdsCrawl/adscrawl-ruby
|
|
51
|
+
rdoc_options: []
|
|
52
|
+
require_paths:
|
|
53
|
+
- lib
|
|
54
|
+
required_ruby_version: !ruby/object:Gem::Requirement
|
|
55
|
+
requirements:
|
|
56
|
+
- - ">="
|
|
57
|
+
- !ruby/object:Gem::Version
|
|
58
|
+
version: '3.1'
|
|
59
|
+
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
60
|
+
requirements:
|
|
61
|
+
- - ">="
|
|
62
|
+
- !ruby/object:Gem::Version
|
|
63
|
+
version: '0'
|
|
64
|
+
requirements: []
|
|
65
|
+
rubygems_version: 4.0.20
|
|
66
|
+
specification_version: 4
|
|
67
|
+
summary: Official Ruby SDK for the AdsCrawl browser and extraction API
|
|
68
|
+
test_files: []
|