urlpipe 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/CHANGELOG.md +13 -0
- data/LICENSE +21 -0
- data/README.md +245 -0
- data/lib/urlpipe/client.rb +347 -0
- data/lib/urlpipe/errors.rb +205 -0
- data/lib/urlpipe/response.rb +68 -0
- data/lib/urlpipe/screenshot.rb +45 -0
- data/lib/urlpipe/version.rb +5 -0
- data/lib/urlpipe/webhook.rb +118 -0
- data/lib/urlpipe.rb +17 -0
- metadata +60 -0
checksums.yaml
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
SHA256:
|
|
3
|
+
metadata.gz: be7381c8419e7711e1e5506df18decd8adf8a9174678a4bf6b14a30343cffe47
|
|
4
|
+
data.tar.gz: 496d714f333fb7d6a983e943a531a330b8b185064eae7ab3d933c10a19098fdf
|
|
5
|
+
SHA512:
|
|
6
|
+
metadata.gz: 54769706ad4bea7fddb97c0190793713011e3dc3d71023ec90d24ee63ea2810f0430120b037d73ed33edab04a3ae7bb8809dca722120c17083e3246f39842128
|
|
7
|
+
data.tar.gz: 1e8522f8eb99c9ffc2e982ef34412a738d9c5f53dcb0a25ab1e0e906be135567c60cf4913fce9cbf831c36f8ae41fbf9487b59a07cc8978e60e3f85f6d25fb58
|
data/CHANGELOG.md
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0
|
|
4
|
+
|
|
5
|
+
The first release.
|
|
6
|
+
|
|
7
|
+
- `Urlpipe::Client` with one method per operation: `markdown`, `html`, `summarize`, `screenshot`, `meta`, `keywords`, `console`, `lighthouse` and `scrape`, plus `result` and `wait` for tokens.
|
|
8
|
+
- Calls are synchronous by default. An analysis that outlives the API's 60-second sync window is polled until it lands.
|
|
9
|
+
- `Urlpipe::Response` with the result typed per operation, the token, labels and the parsed metadata headers.
|
|
10
|
+
- Screenshots decoded to bytes, with the MIME type and the key-free `result_url`.
|
|
11
|
+
- Automatic retries with an `Idempotency-Key` per call, so a retry never runs or bills the work twice.
|
|
12
|
+
- A typed error for every documented API error.
|
|
13
|
+
- `Urlpipe::Webhook.verify` for signed webhook deliveries.
|
data/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Aliat Partner S.L.
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
data/README.md
ADDED
|
@@ -0,0 +1,245 @@
|
|
|
1
|
+
# urlpipe
|
|
2
|
+
|
|
3
|
+
Turn any URL into clean Markdown, rendered HTML, a full-page screenshot, metadata, a summary, keywords, console errors or a Lighthouse audit, from Ruby. Pages are rendered in real Chrome, so JavaScript-heavy sites come back complete.
|
|
4
|
+
|
|
5
|
+
This is the official Ruby client for [URLpipe](https://urlpipe.dev). It has no runtime dependencies and runs on Ruby 3.1 and later.
|
|
6
|
+
|
|
7
|
+
## Install
|
|
8
|
+
|
|
9
|
+
```sh
|
|
10
|
+
bundle add urlpipe
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
or `gem install urlpipe`.
|
|
14
|
+
|
|
15
|
+
## Quickstart
|
|
16
|
+
|
|
17
|
+
Get an API key from a project in the [dashboard](https://urlpipe.dev) (the Free plan includes 1,000 credits a month, no card) and put it in `URLPIPE_API_KEY`:
|
|
18
|
+
|
|
19
|
+
```ruby
|
|
20
|
+
require "urlpipe"
|
|
21
|
+
|
|
22
|
+
client = Urlpipe::Client.new # reads ENV["URLPIPE_API_KEY"]
|
|
23
|
+
|
|
24
|
+
response = client.markdown("https://example.com")
|
|
25
|
+
puts response.data # "# Example Domain\n\n..."
|
|
26
|
+
response.meta.quota.remaining # => 943
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Every call waits for the result and hands it back in `response.data`. Cached results and failed calls cost nothing.
|
|
30
|
+
|
|
31
|
+
## The client
|
|
32
|
+
|
|
33
|
+
```ruby
|
|
34
|
+
Urlpipe::Client.new(
|
|
35
|
+
api_key: nil, # falls back to ENV["URLPIPE_API_KEY"]; raises Urlpipe::ConfigurationError if neither is set
|
|
36
|
+
base_url: "https://urlpipe.dev",
|
|
37
|
+
timeout: 90, # seconds per HTTP request; the API holds a sync request up to 60 s
|
|
38
|
+
max_retries: 2, # see "Retries and idempotency"
|
|
39
|
+
wait_timeout: 300, # how long a long analysis is polled before Urlpipe::WaitTimeoutError
|
|
40
|
+
poll_interval: 2 # seconds between polls of GET /result/:token
|
|
41
|
+
)
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
A client holds no connection state, so one instance can be shared across threads.
|
|
45
|
+
|
|
46
|
+
## Operations
|
|
47
|
+
|
|
48
|
+
Every method takes the URL first. These options work on all of them and are sent only when you give them:
|
|
49
|
+
|
|
50
|
+
| Option | What it does |
|
|
51
|
+
| --- | --- |
|
|
52
|
+
| `sync:` | `true` by default: the call returns the result. `false` returns straight away with a token (see [Async](#async-and-wait)). |
|
|
53
|
+
| `max_age:` | How fresh a cached result must be: seconds (`3600`) or a duration (`"3 days"`). `0` skips the cache. |
|
|
54
|
+
| `labels:` | Your own ids for the request, e.g. `{ client: "acme" }`. Returned with the result and the webhook. |
|
|
55
|
+
| `residential:` | `true` fetches the page from a home broadband address. |
|
|
56
|
+
| `report_to:` | A webhook URL the result is POSTed to (async calls). |
|
|
57
|
+
| `page_options:` | `wait_for_selector`, `delay`, `block_ads`, `block_cookie_banners`, `remove_selectors`. Not on `lighthouse`. |
|
|
58
|
+
| `idempotency_key:` | Sent as the `Idempotency-Key` header. |
|
|
59
|
+
| `extra:` | A Hash merged into the request body, for API options this version doesn't name yet. |
|
|
60
|
+
|
|
61
|
+
Option Hashes go to the API as given; it validates them and tells you what to change.
|
|
62
|
+
|
|
63
|
+
### markdown, html, summarize
|
|
64
|
+
|
|
65
|
+
`data` is a String.
|
|
66
|
+
|
|
67
|
+
```ruby
|
|
68
|
+
client.markdown("https://example.com", page_options: { block_cookie_banners: true }).data
|
|
69
|
+
client.html("https://example.com").data # the rendered document
|
|
70
|
+
client.summarize("https://example.com").data # a Markdown summary
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
### screenshot
|
|
74
|
+
|
|
75
|
+
`data` is a `Urlpipe::Screenshot`: the decoded image bytes, its `mime_type` (`image/png`, `image/jpeg` or `image/webp`) and `result_url`, a link to the image that needs no API key.
|
|
76
|
+
|
|
77
|
+
```ruby
|
|
78
|
+
shot = client.screenshot("https://example.com",
|
|
79
|
+
screenshot_options: { viewport_width: 390, format: "webp" }).data
|
|
80
|
+
shot.save("example.webp")
|
|
81
|
+
shot.mime_type # => "image/webp"
|
|
82
|
+
shot.result_url # => "https://..." - put it straight in an <img> tag
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
### meta
|
|
86
|
+
|
|
87
|
+
`data` is a Hash with `title`, `description`, `language`, `main_image_url`, `favicon_url`, `author_name`, `feed_url`, `publication_date` and `additional_author_information`. Any of them can be `nil`.
|
|
88
|
+
|
|
89
|
+
```ruby
|
|
90
|
+
client.meta("https://example.com").data["title"] # => "Example Domain"
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
### keywords
|
|
94
|
+
|
|
95
|
+
`data` is an Array of Strings, most relevant first.
|
|
96
|
+
|
|
97
|
+
```ruby
|
|
98
|
+
client.keywords("https://example.com").data # => ["example domain", "documentation", ...]
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
### console
|
|
102
|
+
|
|
103
|
+
`data` is an Array of `{"type" => "error" | "warning" | "exception", "text" => "..."}`.
|
|
104
|
+
|
|
105
|
+
```ruby
|
|
106
|
+
client.console("https://example.com").data.select { |entry| entry["type"] == "exception" }
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
### lighthouse
|
|
110
|
+
|
|
111
|
+
`data` is the audit as a Hash. `device:` is `"mobile"` (the default) or `"desktop"`; `include_audits: true` adds the full audits object.
|
|
112
|
+
|
|
113
|
+
```ruby
|
|
114
|
+
report = client.lighthouse("https://example.com", device: "desktop").data
|
|
115
|
+
report.dig("categories", "performance", "score")
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
### scrape
|
|
119
|
+
|
|
120
|
+
Several operations off one page visit, so you get them all much sooner. `data` is `{"url" => ..., "operations" => {"markdown" => {"success", "result", "error", "cached"}, ...}}`. Each operation succeeds or fails on its own. A screenshot inside a scrape stays Base64, as the API returns it.
|
|
121
|
+
|
|
122
|
+
```ruby
|
|
123
|
+
page = client.scrape("https://example.com", %w[markdown meta screenshot],
|
|
124
|
+
screenshot_options: { format: "jpeg" }).data
|
|
125
|
+
page.dig("operations", "meta", "result", "title")
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
## The Response
|
|
129
|
+
|
|
130
|
+
Every method returns a `Urlpipe::Response`:
|
|
131
|
+
|
|
132
|
+
| Field | |
|
|
133
|
+
| --- | --- |
|
|
134
|
+
| `status` | `"completed"`, `"accepted"` (an async request was accepted) or `"processing"` (still running). Also `completed?`, `accepted?`, `processing?`. |
|
|
135
|
+
| `data` | The result when completed, `nil` otherwise. |
|
|
136
|
+
| `token` | The request's result token. Fetching the result again with it is free for 30 days. |
|
|
137
|
+
| `labels` | The labels the request was made with; `{}` when it had none. |
|
|
138
|
+
| `meta` | `cache` (`"hit"`, `"miss"`, `"partial"`), `cache_age`, `processing_time_ms`, `quota` (`cost`, `limit`, `remaining`, `overage`, `resets_at`), `concurrency_limit`, `result_url`, `idempotent_replayed`. |
|
|
139
|
+
|
|
140
|
+
A value the API didn't send is `nil`; `idempotent_replayed` is always `true` or `false`. On an unlimited plan `quota.limit`, `quota.remaining` and `concurrency_limit` are the String `"unlimited"`.
|
|
141
|
+
|
|
142
|
+
## Long analyses
|
|
143
|
+
|
|
144
|
+
A sync request is held for up to 60 seconds. When an analysis takes longer (a Lighthouse audit, a slow site), the client polls `GET /result/:token` every 2 seconds and returns the result when it lands, so your code sees one call. After `wait_timeout` it raises `Urlpipe::WaitTimeoutError`, whose `token` you can collect later: the analysis keeps running.
|
|
145
|
+
|
|
146
|
+
## Async and wait
|
|
147
|
+
|
|
148
|
+
With `sync: false` the call returns at once with a token, and the work continues in the background. Collect it with `wait`, or with a webhook:
|
|
149
|
+
|
|
150
|
+
```ruby
|
|
151
|
+
accepted = client.lighthouse("https://example.com", sync: false)
|
|
152
|
+
accepted.status # => "accepted"
|
|
153
|
+
|
|
154
|
+
report = client.wait(accepted.token, operation: "lighthouse").data
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
`wait(token, operation: nil, timeout: wait_timeout, interval: poll_interval)` polls until the result is ready. `result(token, operation: nil)` checks once and returns a `"processing"` Response while it is still running.
|
|
158
|
+
|
|
159
|
+
The API doesn't say which operation produced a token, so pass `operation:` to get `data` typed the way that method returns it. Without it JSON comes back parsed and text stays a String, which means a screenshot stays Base64: `client.wait(token, operation: "screenshot")` decodes it into a `Urlpipe::Screenshot`.
|
|
160
|
+
|
|
161
|
+
## Webhooks
|
|
162
|
+
|
|
163
|
+
Turn on webhook signing for the project (Settings → Webhook Signing) and verify each delivery before trusting it. `Urlpipe::Webhook.verify` needs no client:
|
|
164
|
+
|
|
165
|
+
```ruby
|
|
166
|
+
# app/controllers/urlpipe_webhooks_controller.rb
|
|
167
|
+
class UrlpipeWebhooksController < ActionController::API
|
|
168
|
+
def create
|
|
169
|
+
payload = Urlpipe::Webhook.verify(request.raw_post, request.headers,
|
|
170
|
+
ENV.fetch("URLPIPE_WEBHOOK_SECRET"))
|
|
171
|
+
ProcessResultJob.perform_later(payload["token"])
|
|
172
|
+
head :ok
|
|
173
|
+
rescue Urlpipe::WebhookVerificationError
|
|
174
|
+
head :unauthorized
|
|
175
|
+
end
|
|
176
|
+
end
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
`verify(raw_body, headers, secret, tolerance: 300)` returns the payload Hash (`token`, `operation`, `labels`, `success`, `result`, `result_url`, `error`, `meta`), or raises `Urlpipe::WebhookVerificationError` saying why. It accepts any signature in the header that matches, so deliveries keep verifying through a secret rotation, and it refuses a timestamp more than `tolerance` seconds from now.
|
|
180
|
+
|
|
181
|
+
Pass the **raw** request body, exactly the bytes received: `request.raw_post` in Rails, `request.body.read` in Rack. Parsed and re-serialized JSON has different bytes and will not verify. `headers` can be Rails' `request.headers`, a Rack env or a plain Hash.
|
|
182
|
+
|
|
183
|
+
Make the handler idempotent on `token`: a delivery can be retried.
|
|
184
|
+
|
|
185
|
+
## Errors
|
|
186
|
+
|
|
187
|
+
Everything raised is a `Urlpipe::Error`, with `status`, `code` (the API's error code, when there is one), `message`, `body` (parsed JSON, or the raw text) and `token`.
|
|
188
|
+
|
|
189
|
+
| Error | When |
|
|
190
|
+
| --- | --- |
|
|
191
|
+
| `Urlpipe::AuthenticationError` | 401: the API key is missing or not an active project key. |
|
|
192
|
+
| `Urlpipe::EmailUnverifiedError` | 403: confirm the email address on the account. |
|
|
193
|
+
| `Urlpipe::InvalidRequestError` | 422 with a code: `invalid_url`, `invalid_max_age`, `invalid_options`, `invalid_labels`, `invalid_idempotency_key`, `idempotency_key_reused`, or a `report_to` the API won't deliver to. |
|
|
194
|
+
| `Urlpipe::AnalysisFailedError` | 422: the page couldn't be analysed; the message says why. A scrape where every operation failed keeps the scrape on `body`. |
|
|
195
|
+
| `Urlpipe::QuotaExceededError` | 429 on the Free plan: `limit`, `used`, `needed`, `resets_at`. |
|
|
196
|
+
| `Urlpipe::ConcurrencyLimitError` | 429: `limit` requests already `running`. |
|
|
197
|
+
| `Urlpipe::RateLimitedError` | 429: sending too fast; `retry_after` seconds. |
|
|
198
|
+
| `Urlpipe::NotFoundError` | 404: no result for this token in this project. |
|
|
199
|
+
| `Urlpipe::StaleResultError` | 410: the result is past the 30-day window. |
|
|
200
|
+
| `Urlpipe::ServerError` | Any other 5xx. |
|
|
201
|
+
| `Urlpipe::ConnectionError` | The API couldn't be reached, or a request outlived `timeout`. |
|
|
202
|
+
| `Urlpipe::WaitTimeoutError` | `wait_timeout` passed; `token` is still collectable. |
|
|
203
|
+
| `Urlpipe::ConfigurationError` | No API key. |
|
|
204
|
+
| `Urlpipe::WebhookVerificationError` | A webhook delivery didn't verify. |
|
|
205
|
+
|
|
206
|
+
```ruby
|
|
207
|
+
begin
|
|
208
|
+
client.markdown(url)
|
|
209
|
+
rescue Urlpipe::AnalysisFailedError => e
|
|
210
|
+
logger.info("Skipped #{url}: #{e.message}") # "The requested page was not found."
|
|
211
|
+
rescue Urlpipe::QuotaExceededError => e
|
|
212
|
+
retry_at(e.resets_at)
|
|
213
|
+
end
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
## Retries and idempotency
|
|
217
|
+
|
|
218
|
+
The client retries, up to `max_retries` times, what a retry can fix: connection errors, 500/502/503, `rate_limited` (after `Retry-After`, capped at 60 s) and `concurrency_limit` (after 1 s, 2 s, ...). It never retries 401, 403, 404, 410, 422 or `quota_exceeded`.
|
|
219
|
+
|
|
220
|
+
Every analysis request that may be retried carries an `Idempotency-Key`: yours if you pass `idempotency_key:`, otherwise a fresh UUID for that call, reused on each of its retries. The API answers a repeated key with the first request's token and result, so a retry after a dropped connection never runs or bills the work twice. Pass your own key (a job id, say) to get the same guarantee across process restarts:
|
|
221
|
+
|
|
222
|
+
```ruby
|
|
223
|
+
client.markdown(url, idempotency_key: "import-#{job.id}")
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
Set `max_retries: 0` to turn retries off; no key is generated then.
|
|
227
|
+
|
|
228
|
+
## Links
|
|
229
|
+
|
|
230
|
+
- API docs: https://urlpipe.dev/docs
|
|
231
|
+
- Pricing: https://urlpipe.dev/pricing
|
|
232
|
+
- MCP server, to give an AI assistant the same tools: https://github.com/URLpipe/mcp
|
|
233
|
+
|
|
234
|
+
## Development
|
|
235
|
+
|
|
236
|
+
```sh
|
|
237
|
+
bundle install
|
|
238
|
+
bundle exec rake test
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
The tests run against a local stub server and never touch the network. Set `URLPIPE_API_KEY` to also run one live smoke test against the real API.
|
|
242
|
+
|
|
243
|
+
## License
|
|
244
|
+
|
|
245
|
+
MIT. Copyright (c) 2026 Aliat Partner S.L.
|
|
@@ -0,0 +1,347 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "json"
|
|
4
|
+
require "net/http"
|
|
5
|
+
require "openssl"
|
|
6
|
+
require "securerandom"
|
|
7
|
+
require "uri"
|
|
8
|
+
|
|
9
|
+
module Urlpipe
|
|
10
|
+
# A client for the URLpipe API. Every analysis method takes the URL first
|
|
11
|
+
# and returns a Response. Calls are synchronous by default: the method
|
|
12
|
+
# returns once the result is ready, however long the analysis takes.
|
|
13
|
+
#
|
|
14
|
+
# client = Urlpipe::Client.new(api_key: "...")
|
|
15
|
+
# client.markdown("https://example.com").data # => "# Example Domain..."
|
|
16
|
+
class Client
|
|
17
|
+
DEFAULT_BASE_URL = "https://urlpipe.dev"
|
|
18
|
+
USER_AGENT = "urlpipe-ruby/#{VERSION}".freeze
|
|
19
|
+
|
|
20
|
+
TEXT_OPERATIONS = %w[markdown html summarize].freeze
|
|
21
|
+
JSON_OPERATIONS = %w[meta keywords console lighthouse scrape].freeze
|
|
22
|
+
|
|
23
|
+
RETRYABLE_STATUSES = [500, 502, 503].freeze
|
|
24
|
+
MAX_RETRY_DELAY = 60
|
|
25
|
+
|
|
26
|
+
NETWORK_ERRORS = [
|
|
27
|
+
SocketError, SystemCallError, IOError, Timeout::Error,
|
|
28
|
+
Net::OpenTimeout, Net::ReadTimeout, Net::WriteTimeout,
|
|
29
|
+
Net::HTTPBadResponse, OpenSSL::SSL::SSLError
|
|
30
|
+
].freeze
|
|
31
|
+
|
|
32
|
+
# The raw HTTP exchange, before it becomes a Response or an error.
|
|
33
|
+
RawResponse = Struct.new(:status, :headers, :text, keyword_init: true) do
|
|
34
|
+
def json? = headers["content-type"].to_s.include?("json")
|
|
35
|
+
end
|
|
36
|
+
private_constant :RawResponse
|
|
37
|
+
|
|
38
|
+
attr_reader :base_url, :timeout, :max_retries, :wait_timeout, :poll_interval
|
|
39
|
+
|
|
40
|
+
# - +api_key+: a project API key; falls back to ENV["URLPIPE_API_KEY"]
|
|
41
|
+
# - +timeout+: seconds per HTTP request (the API holds a sync request up
|
|
42
|
+
# to 60 s, so keep it above that)
|
|
43
|
+
# - +max_retries+: automatic retries on connection errors, 500/502/503,
|
|
44
|
+
# rate_limited and concurrency_limit
|
|
45
|
+
# - +wait_timeout+: how long a sync call or +wait+ keeps polling a long
|
|
46
|
+
# analysis before raising WaitTimeoutError
|
|
47
|
+
# - +poll_interval+: seconds between polls of GET /result/:token
|
|
48
|
+
def initialize(api_key: nil, base_url: DEFAULT_BASE_URL, timeout: 90, max_retries: 2,
|
|
49
|
+
wait_timeout: 300, poll_interval: 2)
|
|
50
|
+
api_key = ENV.fetch("URLPIPE_API_KEY", nil) if api_key.nil? || api_key.to_s.strip.empty?
|
|
51
|
+
if api_key.nil? || api_key.strip.empty?
|
|
52
|
+
raise ConfigurationError,
|
|
53
|
+
"Pass api_key: to Urlpipe::Client.new, or set URLPIPE_API_KEY. " \
|
|
54
|
+
"Each project's API key is in the URLpipe dashboard."
|
|
55
|
+
end
|
|
56
|
+
|
|
57
|
+
@api_key = api_key.strip
|
|
58
|
+
@base_url = base_url.to_s.sub(%r{/+\z}, "")
|
|
59
|
+
@base_uri = URI(@base_url)
|
|
60
|
+
@timeout = timeout
|
|
61
|
+
@max_retries = Integer(max_retries)
|
|
62
|
+
@wait_timeout = wait_timeout
|
|
63
|
+
@poll_interval = poll_interval
|
|
64
|
+
end
|
|
65
|
+
|
|
66
|
+
def inspect
|
|
67
|
+
"#<#{self.class.name} base_url=#{@base_url.inspect} timeout=#{@timeout} " \
|
|
68
|
+
"max_retries=#{@max_retries} wait_timeout=#{@wait_timeout}>"
|
|
69
|
+
end
|
|
70
|
+
|
|
71
|
+
# The page as Markdown. +data+ is a String.
|
|
72
|
+
def markdown(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
73
|
+
page_options: nil, idempotency_key: nil, extra: {})
|
|
74
|
+
analyse("markdown", url, { sync:, max_age:, labels:, residential:, report_to:, page_options: },
|
|
75
|
+
idempotency_key, extra)
|
|
76
|
+
end
|
|
77
|
+
|
|
78
|
+
# The rendered HTML document. +data+ is a String.
|
|
79
|
+
def html(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
80
|
+
page_options: nil, idempotency_key: nil, extra: {})
|
|
81
|
+
analyse("html", url, { sync:, max_age:, labels:, residential:, report_to:, page_options: },
|
|
82
|
+
idempotency_key, extra)
|
|
83
|
+
end
|
|
84
|
+
|
|
85
|
+
# A summary of the page, as Markdown. +data+ is a String.
|
|
86
|
+
def summarize(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
87
|
+
page_options: nil, idempotency_key: nil, extra: {})
|
|
88
|
+
analyse("summarize", url, { sync:, max_age:, labels:, residential:, report_to:, page_options: },
|
|
89
|
+
idempotency_key, extra)
|
|
90
|
+
end
|
|
91
|
+
|
|
92
|
+
# A full-page screenshot. +data+ is a Screenshot (decoded bytes,
|
|
93
|
+
# +mime_type+, +result_url+).
|
|
94
|
+
def screenshot(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
95
|
+
page_options: nil, screenshot_options: nil, idempotency_key: nil, extra: {})
|
|
96
|
+
analyse("screenshot", url,
|
|
97
|
+
{ sync:, max_age:, labels:, residential:, report_to:, page_options:, screenshot_options: },
|
|
98
|
+
idempotency_key, extra)
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
# Page metadata. +data+ is a Hash: title, description, language,
|
|
102
|
+
# main_image_url, favicon_url, author_name, feed_url, publication_date,
|
|
103
|
+
# additional_author_information.
|
|
104
|
+
def meta(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
105
|
+
page_options: nil, idempotency_key: nil, extra: {})
|
|
106
|
+
analyse("meta", url, { sync:, max_age:, labels:, residential:, report_to:, page_options: },
|
|
107
|
+
idempotency_key, extra)
|
|
108
|
+
end
|
|
109
|
+
|
|
110
|
+
# The page's keywords, most relevant first. +data+ is an Array of Strings.
|
|
111
|
+
def keywords(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
112
|
+
page_options: nil, idempotency_key: nil, extra: {})
|
|
113
|
+
analyse("keywords", url, { sync:, max_age:, labels:, residential:, report_to:, page_options: },
|
|
114
|
+
idempotency_key, extra)
|
|
115
|
+
end
|
|
116
|
+
|
|
117
|
+
# Console errors, warnings and uncaught exceptions. +data+ is an Array of
|
|
118
|
+
# Hashes with "type" ("error", "warning" or "exception") and "text".
|
|
119
|
+
def console(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
120
|
+
page_options: nil, idempotency_key: nil, extra: {})
|
|
121
|
+
analyse("console", url, { sync:, max_age:, labels:, residential:, report_to:, page_options: },
|
|
122
|
+
idempotency_key, extra)
|
|
123
|
+
end
|
|
124
|
+
|
|
125
|
+
# A Lighthouse audit. +data+ is a Hash. +device+ is "mobile" (the API's
|
|
126
|
+
# default) or "desktop".
|
|
127
|
+
def lighthouse(url, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
128
|
+
device: nil, include_audits: nil, idempotency_key: nil, extra: {})
|
|
129
|
+
analyse("lighthouse", url, { sync:, max_age:, labels:, residential:, report_to:, device:, include_audits: },
|
|
130
|
+
idempotency_key, extra)
|
|
131
|
+
end
|
|
132
|
+
|
|
133
|
+
# Several operations off one page visit. +operations+ is an Array such as
|
|
134
|
+
# ["markdown", "meta"]. +data+ is a Hash:
|
|
135
|
+
# {"url" => ..., "operations" => {"markdown" => {"success", "result", "error", "cached"}}}.
|
|
136
|
+
# A screenshot inside a scrape stays Base64, as the API returns it.
|
|
137
|
+
def scrape(url, operations, sync: true, max_age: nil, labels: nil, residential: nil, report_to: nil,
|
|
138
|
+
page_options: nil, device: nil, include_audits: nil, screenshot_options: nil,
|
|
139
|
+
idempotency_key: nil, extra: {})
|
|
140
|
+
analyse("scrape", url,
|
|
141
|
+
{ operations:, sync:, max_age:, labels:, residential:, report_to:, page_options:,
|
|
142
|
+
device:, include_audits:, screenshot_options: },
|
|
143
|
+
idempotency_key, extra)
|
|
144
|
+
end
|
|
145
|
+
|
|
146
|
+
# One GET /result/:token. Returns a completed Response (200) or a
|
|
147
|
+
# processing one (202); raises for everything else.
|
|
148
|
+
#
|
|
149
|
+
# +operation+ names the operation that produced the token, so +data+ is
|
|
150
|
+
# typed as that operation returns it. Without it, JSON stays parsed JSON
|
|
151
|
+
# and text stays a String (a screenshot stays Base64: pass
|
|
152
|
+
# <tt>operation: "screenshot"</tt> to get a Screenshot).
|
|
153
|
+
def result(token, operation: nil)
|
|
154
|
+
raw = execute(:get, "/result/#{URI.encode_www_form_component(token)}")
|
|
155
|
+
|
|
156
|
+
case raw.status
|
|
157
|
+
when 200 then completed(raw, operation&.to_s)
|
|
158
|
+
when 202 then pending(raw, "processing", fallback_token: token)
|
|
159
|
+
else
|
|
160
|
+
if processing_timeout?(raw)
|
|
161
|
+
pending(raw, "processing", fallback_token: token)
|
|
162
|
+
else
|
|
163
|
+
raise error_for(raw)
|
|
164
|
+
end
|
|
165
|
+
end
|
|
166
|
+
end
|
|
167
|
+
|
|
168
|
+
# Polls GET /result/:token every +interval+ seconds until the result is
|
|
169
|
+
# ready, then returns the completed Response. Raises WaitTimeoutError
|
|
170
|
+
# after +timeout+ seconds, or the API's error if the analysis failed.
|
|
171
|
+
def wait(token, operation: nil, timeout: @wait_timeout, interval: @poll_interval)
|
|
172
|
+
deadline = monotonic_now + timeout
|
|
173
|
+
|
|
174
|
+
loop do
|
|
175
|
+
response = result(token, operation: operation)
|
|
176
|
+
return response if response.completed?
|
|
177
|
+
|
|
178
|
+
if monotonic_now + interval > deadline
|
|
179
|
+
raise WaitTimeoutError.new(
|
|
180
|
+
"The result for #{token} was not ready within #{timeout} seconds. " \
|
|
181
|
+
"It keeps running: collect it later with client.wait(#{token.inspect}).",
|
|
182
|
+
token: token
|
|
183
|
+
)
|
|
184
|
+
end
|
|
185
|
+
|
|
186
|
+
sleep(interval)
|
|
187
|
+
end
|
|
188
|
+
end
|
|
189
|
+
|
|
190
|
+
private
|
|
191
|
+
|
|
192
|
+
def analyse(operation, url, options, idempotency_key, extra)
|
|
193
|
+
body = { "url" => url }
|
|
194
|
+
options.each { |key, value| body[key.to_s] = value unless value.nil? }
|
|
195
|
+
(extra || {}).each { |key, value| body[key.to_s] = value }
|
|
196
|
+
|
|
197
|
+
idempotency_key ||= SecureRandom.uuid if @max_retries.positive?
|
|
198
|
+
raw = execute(:post, "/#{operation}", body: body, idempotency_key: idempotency_key)
|
|
199
|
+
|
|
200
|
+
if raw.status == 200
|
|
201
|
+
sync?(body["sync"]) ? completed(raw, operation) : pending(raw, "accepted")
|
|
202
|
+
elsif processing_timeout?(raw) && (token = json_body(raw)["token"])
|
|
203
|
+
wait(token, operation: operation)
|
|
204
|
+
else
|
|
205
|
+
raise error_for(raw)
|
|
206
|
+
end
|
|
207
|
+
end
|
|
208
|
+
|
|
209
|
+
def sync?(value) = [true, "true"].include?(value)
|
|
210
|
+
|
|
211
|
+
def completed(raw, operation)
|
|
212
|
+
Response.new(
|
|
213
|
+
status: "completed",
|
|
214
|
+
data: typed_data(raw, operation),
|
|
215
|
+
token: raw.headers["x-result-token"],
|
|
216
|
+
labels: header_labels(raw) || {},
|
|
217
|
+
meta: Meta.from_headers(raw.headers)
|
|
218
|
+
)
|
|
219
|
+
end
|
|
220
|
+
|
|
221
|
+
def pending(raw, status, fallback_token: nil)
|
|
222
|
+
body = json_body(raw)
|
|
223
|
+
Response.new(
|
|
224
|
+
status: status,
|
|
225
|
+
data: nil,
|
|
226
|
+
token: raw.headers["x-result-token"] || body["token"] || fallback_token,
|
|
227
|
+
labels: header_labels(raw) || (body["labels"].is_a?(Hash) ? body["labels"] : {}),
|
|
228
|
+
meta: Meta.from_headers(raw.headers)
|
|
229
|
+
)
|
|
230
|
+
end
|
|
231
|
+
|
|
232
|
+
def typed_data(raw, operation)
|
|
233
|
+
if TEXT_OPERATIONS.include?(operation)
|
|
234
|
+
raw.text
|
|
235
|
+
elsif operation == "screenshot"
|
|
236
|
+
Screenshot.from_base64(raw.text, result_url: raw.headers["x-result-url"])
|
|
237
|
+
elsif JSON_OPERATIONS.include?(operation) || (operation.nil? && raw.json?)
|
|
238
|
+
parse_result_json(raw)
|
|
239
|
+
else
|
|
240
|
+
raw.text
|
|
241
|
+
end
|
|
242
|
+
end
|
|
243
|
+
|
|
244
|
+
def parse_result_json(raw)
|
|
245
|
+
JSON.parse(raw.text)
|
|
246
|
+
rescue JSON::ParserError => e
|
|
247
|
+
raise Error.new("The API answered #{raw.status} with a body that is not valid JSON: #{e.message}",
|
|
248
|
+
status: raw.status, body: raw.text)
|
|
249
|
+
end
|
|
250
|
+
|
|
251
|
+
def header_labels(raw)
|
|
252
|
+
value = raw.headers["x-labels"]
|
|
253
|
+
return nil if value.nil? || value.strip.empty?
|
|
254
|
+
|
|
255
|
+
JSON.parse(value)
|
|
256
|
+
rescue JSON::ParserError
|
|
257
|
+
nil
|
|
258
|
+
end
|
|
259
|
+
|
|
260
|
+
def json_body(raw)
|
|
261
|
+
parsed = ErrorFactory.parse_json(raw.text)
|
|
262
|
+
parsed.is_a?(Hash) ? parsed : {}
|
|
263
|
+
end
|
|
264
|
+
|
|
265
|
+
def processing_timeout?(raw)
|
|
266
|
+
raw.status == 504 && json_body(raw)["error"] == "processing_timeout"
|
|
267
|
+
end
|
|
268
|
+
|
|
269
|
+
def error_for(raw)
|
|
270
|
+
ErrorFactory.build(status: raw.status, text: raw.text, headers: raw.headers)
|
|
271
|
+
end
|
|
272
|
+
|
|
273
|
+
# Sends the request, retrying what is safe to retry. The same
|
|
274
|
+
# Idempotency-Key goes out on every attempt, so a retry can never run or
|
|
275
|
+
# bill the work twice.
|
|
276
|
+
def execute(method, path, body: nil, idempotency_key: nil)
|
|
277
|
+
attempt = 0
|
|
278
|
+
|
|
279
|
+
loop do
|
|
280
|
+
begin
|
|
281
|
+
raw = perform(method, path, body, idempotency_key)
|
|
282
|
+
rescue ConnectionError
|
|
283
|
+
raise if attempt >= @max_retries
|
|
284
|
+
|
|
285
|
+
sleep(backoff(attempt))
|
|
286
|
+
attempt += 1
|
|
287
|
+
next
|
|
288
|
+
end
|
|
289
|
+
|
|
290
|
+
delay = retry_delay(raw, attempt)
|
|
291
|
+
return raw if delay.nil? || attempt >= @max_retries
|
|
292
|
+
|
|
293
|
+
sleep(delay)
|
|
294
|
+
attempt += 1
|
|
295
|
+
end
|
|
296
|
+
end
|
|
297
|
+
|
|
298
|
+
def retry_delay(raw, attempt)
|
|
299
|
+
return backoff(attempt) if RETRYABLE_STATUSES.include?(raw.status)
|
|
300
|
+
return nil unless raw.status == 429
|
|
301
|
+
|
|
302
|
+
case json_body(raw)["error"]
|
|
303
|
+
when "rate_limited"
|
|
304
|
+
retry_after = ErrorFactory.retry_after(raw.headers, json_body(raw))
|
|
305
|
+
[retry_after || 1, MAX_RETRY_DELAY].min
|
|
306
|
+
when "concurrency_limit"
|
|
307
|
+
backoff(attempt)
|
|
308
|
+
end
|
|
309
|
+
end
|
|
310
|
+
|
|
311
|
+
def backoff(attempt) = [2**attempt, MAX_RETRY_DELAY].min
|
|
312
|
+
|
|
313
|
+
def perform(method, path, body, idempotency_key)
|
|
314
|
+
uri = @base_uri.dup
|
|
315
|
+
uri.path = "#{@base_uri.path}#{path}"
|
|
316
|
+
|
|
317
|
+
request = method == :post ? Net::HTTP::Post.new(uri) : Net::HTTP::Get.new(uri)
|
|
318
|
+
request["Authorization"] = "Bearer #{@api_key}"
|
|
319
|
+
request["User-Agent"] = USER_AGENT
|
|
320
|
+
request["Accept"] = "*/*"
|
|
321
|
+
request["Idempotency-Key"] = idempotency_key if idempotency_key
|
|
322
|
+
if body
|
|
323
|
+
request["Content-Type"] = "application/json"
|
|
324
|
+
request.body = JSON.generate(body)
|
|
325
|
+
end
|
|
326
|
+
|
|
327
|
+
response = http_for(uri).start { |http| http.request(request) }
|
|
328
|
+
headers = response.each_header.to_h { |name, value| [name.downcase, value] }
|
|
329
|
+
text = (response.body || +"").dup.force_encoding(Encoding::UTF_8)
|
|
330
|
+
RawResponse.new(status: response.code.to_i, headers: headers, text: text)
|
|
331
|
+
rescue *NETWORK_ERRORS => e
|
|
332
|
+
raise ConnectionError, "Could not reach #{@base_url}: #{e.class}: #{e.message}"
|
|
333
|
+
end
|
|
334
|
+
|
|
335
|
+
def http_for(uri)
|
|
336
|
+
http = Net::HTTP.new(uri.host, uri.port)
|
|
337
|
+
http.use_ssl = uri.scheme == "https"
|
|
338
|
+
http.open_timeout = @timeout
|
|
339
|
+
http.read_timeout = @timeout
|
|
340
|
+
http.write_timeout = @timeout
|
|
341
|
+
http.max_retries = 0
|
|
342
|
+
http
|
|
343
|
+
end
|
|
344
|
+
|
|
345
|
+
def monotonic_now = Process.clock_gettime(Process::CLOCK_MONOTONIC)
|
|
346
|
+
end
|
|
347
|
+
end
|
|
@@ -0,0 +1,205 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "json"
|
|
4
|
+
|
|
5
|
+
module Urlpipe
|
|
6
|
+
# Base class of every error this library raises.
|
|
7
|
+
#
|
|
8
|
+
# - +status+: the HTTP status, when the error came from a response
|
|
9
|
+
# - +code+: the body's +error+ value when it is a machine-readable code
|
|
10
|
+
# (+"invalid_url"+, +"quota_exceeded"+, ...)
|
|
11
|
+
# - +body+: the parsed JSON body, or the raw text when it was not JSON
|
|
12
|
+
# - +token+: the result token, when the body carried one
|
|
13
|
+
class Error < StandardError
|
|
14
|
+
attr_reader :status, :code, :body, :token
|
|
15
|
+
|
|
16
|
+
def initialize(message = nil, status: nil, code: nil, body: nil, token: nil)
|
|
17
|
+
super(message)
|
|
18
|
+
@status = status
|
|
19
|
+
@code = code
|
|
20
|
+
@body = body
|
|
21
|
+
@token = token
|
|
22
|
+
end
|
|
23
|
+
end
|
|
24
|
+
|
|
25
|
+
# No API key was given and URLPIPE_API_KEY is not set.
|
|
26
|
+
class ConfigurationError < Error; end
|
|
27
|
+
|
|
28
|
+
# 401: the API key is missing or not an active project key.
|
|
29
|
+
class AuthenticationError < Error; end
|
|
30
|
+
|
|
31
|
+
# 403 email_unverified: the key is valid, but the account's email address
|
|
32
|
+
# has not been confirmed yet.
|
|
33
|
+
class EmailUnverifiedError < Error; end
|
|
34
|
+
|
|
35
|
+
# 422 with a parameter code: invalid_url, invalid_max_age, invalid_options,
|
|
36
|
+
# invalid_labels, invalid_idempotency_key, idempotency_key_reused, or a
|
|
37
|
+
# report_to the API will not deliver to.
|
|
38
|
+
class InvalidRequestError < Error; end
|
|
39
|
+
|
|
40
|
+
# 422 whose error is a sentence: the page could not be analysed. The
|
|
41
|
+
# sentence is the message. A /scrape where every operation failed keeps the
|
|
42
|
+
# whole scrape object on +body+.
|
|
43
|
+
class AnalysisFailedError < Error; end
|
|
44
|
+
|
|
45
|
+
# 429 quota_exceeded (Free plan only): the month's credits are spent.
|
|
46
|
+
class QuotaExceededError < Error
|
|
47
|
+
attr_reader :limit, :used, :needed, :resets_at
|
|
48
|
+
|
|
49
|
+
def initialize(message = nil, limit: nil, used: nil, needed: nil, resets_at: nil, **rest)
|
|
50
|
+
super(message, **rest)
|
|
51
|
+
@limit = limit
|
|
52
|
+
@used = used
|
|
53
|
+
@needed = needed
|
|
54
|
+
@resets_at = resets_at
|
|
55
|
+
end
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
# 429 concurrency_limit: too many of your requests are already running.
|
|
59
|
+
class ConcurrencyLimitError < Error
|
|
60
|
+
attr_reader :limit, :running
|
|
61
|
+
|
|
62
|
+
def initialize(message = nil, limit: nil, running: nil, **rest)
|
|
63
|
+
super(message, **rest)
|
|
64
|
+
@limit = limit
|
|
65
|
+
@running = running
|
|
66
|
+
end
|
|
67
|
+
end
|
|
68
|
+
|
|
69
|
+
# 429 rate_limited: sending too fast. +retry_after+ is in seconds.
|
|
70
|
+
class RateLimitedError < Error
|
|
71
|
+
attr_reader :retry_after
|
|
72
|
+
|
|
73
|
+
def initialize(message = nil, retry_after: nil, **rest)
|
|
74
|
+
super(message, **rest)
|
|
75
|
+
@retry_after = retry_after
|
|
76
|
+
end
|
|
77
|
+
end
|
|
78
|
+
|
|
79
|
+
# 404 not_found: no result for this token under your project.
|
|
80
|
+
class NotFoundError < Error; end
|
|
81
|
+
|
|
82
|
+
# 410 stale: the result is older than the 30-day retention window.
|
|
83
|
+
class StaleResultError < Error; end
|
|
84
|
+
|
|
85
|
+
# A 5xx response.
|
|
86
|
+
class ServerError < Error; end
|
|
87
|
+
|
|
88
|
+
# The API could not be reached: DNS, refused or dropped connection, TLS,
|
|
89
|
+
# or a request that outlived +timeout+.
|
|
90
|
+
class ConnectionError < Error; end
|
|
91
|
+
|
|
92
|
+
# +wait+ (or a sync call that timed out on the server) ran out of
|
|
93
|
+
# +wait_timeout+. The analysis may still finish: keep +token+.
|
|
94
|
+
class WaitTimeoutError < Error; end
|
|
95
|
+
|
|
96
|
+
# A webhook delivery failed signature verification. The message says why.
|
|
97
|
+
class WebhookVerificationError < Error; end
|
|
98
|
+
|
|
99
|
+
# Builds the right error from an HTTP error response.
|
|
100
|
+
module ErrorFactory
|
|
101
|
+
CODE_PATTERN = /\A[a-z][a-z0-9_]*\z/
|
|
102
|
+
INVALID_REQUEST_CODES = %w[
|
|
103
|
+
invalid_url invalid_max_age invalid_options invalid_labels
|
|
104
|
+
invalid_idempotency_key idempotency_key_reused
|
|
105
|
+
].freeze
|
|
106
|
+
|
|
107
|
+
DEFAULT_MESSAGES = {
|
|
108
|
+
AuthenticationError => "The API key is missing or not an active project key. " \
|
|
109
|
+
"Copy your project's key from the dashboard.",
|
|
110
|
+
EmailUnverifiedError => "Confirm the email address on this account before using the API. " \
|
|
111
|
+
"The link is in the email we sent when you signed up.",
|
|
112
|
+
NotFoundError => "No result for this token under this project.",
|
|
113
|
+
StaleResultError => "This result is older than the 30-day retention window. Run a new analysis."
|
|
114
|
+
}.freeze
|
|
115
|
+
|
|
116
|
+
module_function
|
|
117
|
+
|
|
118
|
+
def build(status:, text:, headers: {})
|
|
119
|
+
parsed = parse_json(text)
|
|
120
|
+
hash = parsed.is_a?(Hash) ? parsed : {}
|
|
121
|
+
error_value = hash["error"].is_a?(String) ? hash["error"] : nil
|
|
122
|
+
code = error_value if error_value&.match?(CODE_PATTERN)
|
|
123
|
+
|
|
124
|
+
klass, extra = classify(status, code, error_value, hash, headers)
|
|
125
|
+
message = message_for(klass, hash, error_value, code, status, text)
|
|
126
|
+
|
|
127
|
+
klass.new(
|
|
128
|
+
message,
|
|
129
|
+
status: status,
|
|
130
|
+
code: code,
|
|
131
|
+
body: parsed.nil? ? text : parsed,
|
|
132
|
+
token: hash["token"].is_a?(String) ? hash["token"] : headers["x-result-token"],
|
|
133
|
+
**extra
|
|
134
|
+
)
|
|
135
|
+
end
|
|
136
|
+
|
|
137
|
+
def parse_json(text)
|
|
138
|
+
return nil if text.nil? || text.strip.empty?
|
|
139
|
+
|
|
140
|
+
JSON.parse(text)
|
|
141
|
+
rescue JSON::ParserError
|
|
142
|
+
nil
|
|
143
|
+
end
|
|
144
|
+
|
|
145
|
+
def classify(status, code, error_value, hash, headers)
|
|
146
|
+
case status
|
|
147
|
+
when 401 then [AuthenticationError, {}]
|
|
148
|
+
when 403 then [code == "email_unverified" ? EmailUnverifiedError : Error, {}]
|
|
149
|
+
when 404 then [NotFoundError, {}]
|
|
150
|
+
when 410 then [StaleResultError, {}]
|
|
151
|
+
when 422
|
|
152
|
+
[invalid_request?(error_value) ? InvalidRequestError : AnalysisFailedError, {}]
|
|
153
|
+
when 429 then classify_429(code, hash, headers)
|
|
154
|
+
when 500..599 then [ServerError, {}]
|
|
155
|
+
else [Error, {}]
|
|
156
|
+
end
|
|
157
|
+
end
|
|
158
|
+
|
|
159
|
+
def classify_429(code, hash, headers)
|
|
160
|
+
case code
|
|
161
|
+
when "quota_exceeded"
|
|
162
|
+
[QuotaExceededError,
|
|
163
|
+
{ limit: hash["limit"], used: hash["used"], needed: hash["needed"], resets_at: hash["resets_at"] }]
|
|
164
|
+
when "concurrency_limit"
|
|
165
|
+
[ConcurrencyLimitError, { limit: hash["limit"], running: hash["running"] }]
|
|
166
|
+
when "rate_limited"
|
|
167
|
+
[RateLimitedError, { retry_after: retry_after(headers, hash) }]
|
|
168
|
+
else
|
|
169
|
+
[Error, {}]
|
|
170
|
+
end
|
|
171
|
+
end
|
|
172
|
+
|
|
173
|
+
def invalid_request?(error_value)
|
|
174
|
+
return false unless error_value
|
|
175
|
+
|
|
176
|
+
INVALID_REQUEST_CODES.include?(error_value) || error_value.start_with?("report_to")
|
|
177
|
+
end
|
|
178
|
+
|
|
179
|
+
def retry_after(headers, hash)
|
|
180
|
+
value = headers["retry-after"]
|
|
181
|
+
return Integer(value, 10) if value.is_a?(String) && value.strip.match?(/\A\d+\z/)
|
|
182
|
+
|
|
183
|
+
hash["retry_after"].is_a?(Integer) ? hash["retry_after"] : nil
|
|
184
|
+
end
|
|
185
|
+
|
|
186
|
+
def message_for(klass, hash, error_value, code, status, text)
|
|
187
|
+
return hash["message"] if hash["message"].is_a?(String) && !hash["message"].empty?
|
|
188
|
+
return error_value if error_value && !error_value.empty? && code.nil?
|
|
189
|
+
return scrape_failure_message(hash) if klass == AnalysisFailedError && hash["operations"].is_a?(Hash)
|
|
190
|
+
return DEFAULT_MESSAGES[klass] if DEFAULT_MESSAGES.key?(klass)
|
|
191
|
+
return "The API refused the request with HTTP #{status} (#{code})." if code
|
|
192
|
+
|
|
193
|
+
snippet = text.to_s.strip[0, 200]
|
|
194
|
+
snippet.empty? ? "HTTP #{status}" : "HTTP #{status}: #{snippet}"
|
|
195
|
+
end
|
|
196
|
+
|
|
197
|
+
def scrape_failure_message(hash)
|
|
198
|
+
details = hash["operations"].map do |name, entry|
|
|
199
|
+
error = entry.is_a?(Hash) ? entry["error"] : nil
|
|
200
|
+
error ? "#{name}: #{error}" : name
|
|
201
|
+
end
|
|
202
|
+
"Every operation failed: #{details.join("; ")}"
|
|
203
|
+
end
|
|
204
|
+
end
|
|
205
|
+
end
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "json"
|
|
4
|
+
|
|
5
|
+
module Urlpipe
|
|
6
|
+
# What an analysis cost and what is left of the month's allowance.
|
|
7
|
+
# +limit+ and +remaining+ are Integers, or the String "unlimited".
|
|
8
|
+
Quota = Struct.new(:cost, :limit, :remaining, :overage, :resets_at, keyword_init: true)
|
|
9
|
+
|
|
10
|
+
# The response's metadata headers, parsed. A header that was not sent is
|
|
11
|
+
# +nil+ (+idempotent_replayed+ is always true or false).
|
|
12
|
+
#
|
|
13
|
+
# - +cache+: "hit", "miss" or (for /scrape) "partial"
|
|
14
|
+
# - +cache_age+: seconds, when served from cache
|
|
15
|
+
# - +processing_time_ms+: once the work has finished
|
|
16
|
+
# - +quota+: a Quota
|
|
17
|
+
# - +concurrency_limit+: Integer or "unlimited"
|
|
18
|
+
# - +result_url+: screenshots only, a link to the image that needs no key
|
|
19
|
+
# - +idempotent_replayed+: +true+ when this answers an earlier request with
|
|
20
|
+
# the same Idempotency-Key, +false+ otherwise
|
|
21
|
+
Meta = Struct.new(
|
|
22
|
+
:cache, :cache_age, :processing_time_ms, :quota, :concurrency_limit,
|
|
23
|
+
:result_url, :idempotent_replayed,
|
|
24
|
+
keyword_init: true
|
|
25
|
+
) do
|
|
26
|
+
# +headers+ is a Hash of lowercase header names to values.
|
|
27
|
+
def self.from_headers(headers)
|
|
28
|
+
new(
|
|
29
|
+
cache: headers["x-cache"],
|
|
30
|
+
cache_age: integer_or_nil(headers["x-cache-age"]),
|
|
31
|
+
processing_time_ms: integer_or_nil(headers["x-processing-time-ms"]),
|
|
32
|
+
quota: Quota.new(
|
|
33
|
+
cost: integer_or_given(headers["x-quota-cost"]),
|
|
34
|
+
limit: integer_or_given(headers["x-quota-limit"]),
|
|
35
|
+
remaining: integer_or_given(headers["x-quota-remaining"]),
|
|
36
|
+
overage: integer_or_given(headers["x-quota-overage"]),
|
|
37
|
+
resets_at: headers["x-quota-reset"]
|
|
38
|
+
),
|
|
39
|
+
concurrency_limit: integer_or_given(headers["x-concurrency-limit"]),
|
|
40
|
+
result_url: headers["x-result-url"],
|
|
41
|
+
idempotent_replayed: headers["idempotent-replayed"].to_s.strip.casecmp?("true")
|
|
42
|
+
)
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
def self.integer_or_nil(value)
|
|
46
|
+
value = value&.strip
|
|
47
|
+
value&.match?(/\A-?\d+\z/) ? Integer(value, 10) : nil
|
|
48
|
+
end
|
|
49
|
+
|
|
50
|
+
def self.integer_or_given(value)
|
|
51
|
+
integer_or_nil(value) || value&.strip
|
|
52
|
+
end
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
# What every method returns.
|
|
56
|
+
#
|
|
57
|
+
# - +status+: "completed", "accepted" (an async request was accepted) or
|
|
58
|
+
# "processing" (GET /result/:token is still working on it)
|
|
59
|
+
# - +data+: the result when completed, +nil+ otherwise
|
|
60
|
+
# - +token+: the request's result token
|
|
61
|
+
# - +labels+: the labels the request was made with; {} when it had none
|
|
62
|
+
# - +meta+: a Meta
|
|
63
|
+
Response = Struct.new(:status, :data, :token, :labels, :meta, keyword_init: true) do
|
|
64
|
+
def completed? = status == "completed"
|
|
65
|
+
def accepted? = status == "accepted"
|
|
66
|
+
def processing? = status == "processing"
|
|
67
|
+
end
|
|
68
|
+
end
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Urlpipe
|
|
4
|
+
# A decoded screenshot: +data+ is the image as a binary String,
|
|
5
|
+
# +mime_type+ is "image/png", "image/jpeg" or "image/webp", and
|
|
6
|
+
# +result_url+ is a link to the same image that needs no API key.
|
|
7
|
+
class Screenshot
|
|
8
|
+
attr_reader :data, :mime_type, :result_url
|
|
9
|
+
|
|
10
|
+
# Builds a Screenshot from the API's Base64 body.
|
|
11
|
+
def self.from_base64(text, result_url: nil)
|
|
12
|
+
new(text.to_s.unpack1("m"), result_url: result_url)
|
|
13
|
+
end
|
|
14
|
+
|
|
15
|
+
def self.detect_mime_type(bytes)
|
|
16
|
+
if bytes.start_with?("\x89PNG\r\n\x1A\n".b)
|
|
17
|
+
"image/png"
|
|
18
|
+
elsif bytes.start_with?("\xFF\xD8\xFF".b)
|
|
19
|
+
"image/jpeg"
|
|
20
|
+
elsif bytes.byteslice(0, 4) == "RIFF".b && bytes.byteslice(8, 4) == "WEBP".b
|
|
21
|
+
"image/webp"
|
|
22
|
+
else
|
|
23
|
+
"image/png"
|
|
24
|
+
end
|
|
25
|
+
end
|
|
26
|
+
|
|
27
|
+
def initialize(data, result_url: nil, mime_type: nil)
|
|
28
|
+
@data = data.b
|
|
29
|
+
@mime_type = mime_type || self.class.detect_mime_type(@data)
|
|
30
|
+
@result_url = result_url
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
# Writes the image to +path+ and returns the path.
|
|
34
|
+
def save(path)
|
|
35
|
+
File.binwrite(path, data)
|
|
36
|
+
path
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
def bytesize = data.bytesize
|
|
40
|
+
|
|
41
|
+
def inspect
|
|
42
|
+
"#<#{self.class.name} #{mime_type} #{bytesize} bytes result_url=#{result_url.inspect}>"
|
|
43
|
+
end
|
|
44
|
+
end
|
|
45
|
+
end
|
|
@@ -0,0 +1,118 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "json"
|
|
4
|
+
require "openssl"
|
|
5
|
+
|
|
6
|
+
module Urlpipe
|
|
7
|
+
# Verifies signed webhook deliveries. No Client needed.
|
|
8
|
+
#
|
|
9
|
+
# # Rails
|
|
10
|
+
# payload = Urlpipe::Webhook.verify(request.raw_post, request.headers,
|
|
11
|
+
# ENV["URLPIPE_WEBHOOK_SECRET"])
|
|
12
|
+
# payload["token"] # => "0Zx3...9aQ"
|
|
13
|
+
#
|
|
14
|
+
# Give it the RAW request body, exactly the bytes received. Parsed and
|
|
15
|
+
# re-serialized JSON has different bytes, and the signature will not match.
|
|
16
|
+
module Webhook
|
|
17
|
+
TIMESTAMP_HEADER = "X-URLpipe-Timestamp"
|
|
18
|
+
SIGNATURE_HEADER = "X-URLpipe-Signature"
|
|
19
|
+
DEFAULT_TOLERANCE = 300
|
|
20
|
+
|
|
21
|
+
module_function
|
|
22
|
+
|
|
23
|
+
# Checks the delivery's signature and freshness and returns the parsed
|
|
24
|
+
# payload, a Hash with "token", "operation", "labels", "success",
|
|
25
|
+
# "result", "result_url", "error" and "meta".
|
|
26
|
+
#
|
|
27
|
+
# - +raw_body+: the request body as received (a String, or an IO)
|
|
28
|
+
# - +headers+: the request headers. Accepts a plain Hash, Rails'
|
|
29
|
+
# +request.headers+, a Rack env (+HTTP_X_URLPIPE_SIGNATURE+) or any
|
|
30
|
+
# object with +[]+; names are matched case-insensitively.
|
|
31
|
+
# - +secret+: the project's signing secret (+whsec_...+), used as-is
|
|
32
|
+
# - +tolerance+: seconds a timestamp may be off from now, either way
|
|
33
|
+
#
|
|
34
|
+
# Raises WebhookVerificationError with the reason when the delivery
|
|
35
|
+
# does not verify.
|
|
36
|
+
def verify(raw_body, headers, secret, tolerance: DEFAULT_TOLERANCE)
|
|
37
|
+
raise WebhookVerificationError, "No webhook signing secret was given." if secret.to_s.empty?
|
|
38
|
+
|
|
39
|
+
body = read_body(raw_body)
|
|
40
|
+
timestamp = header(headers, TIMESTAMP_HEADER).to_s.strip
|
|
41
|
+
signature = header(headers, SIGNATURE_HEADER).to_s.strip
|
|
42
|
+
|
|
43
|
+
raise WebhookVerificationError, "The #{SIGNATURE_HEADER} header is missing." if signature.empty?
|
|
44
|
+
raise WebhookVerificationError, "The #{TIMESTAMP_HEADER} header is missing." if timestamp.empty?
|
|
45
|
+
unless timestamp.match?(/\A\d+\z/)
|
|
46
|
+
raise WebhookVerificationError, "The #{TIMESTAMP_HEADER} header is not a Unix timestamp."
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
age = Time.now.to_i - Integer(timestamp, 10)
|
|
50
|
+
if age.abs > tolerance
|
|
51
|
+
raise WebhookVerificationError,
|
|
52
|
+
"The delivery's timestamp is #{age.abs} seconds #{age.negative? ? "ahead" : "old"}, " \
|
|
53
|
+
"outside the #{tolerance}-second tolerance."
|
|
54
|
+
end
|
|
55
|
+
|
|
56
|
+
expected = OpenSSL::HMAC.hexdigest("SHA256", secret.to_s, "#{timestamp}.".b + body)
|
|
57
|
+
unless candidates(signature).any? { |candidate| OpenSSL.secure_compare(candidate, expected) }
|
|
58
|
+
raise WebhookVerificationError, "No v1 signature in #{SIGNATURE_HEADER} matches this body and secret."
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
parse(body)
|
|
62
|
+
end
|
|
63
|
+
|
|
64
|
+
def read_body(raw_body)
|
|
65
|
+
body = raw_body.respond_to?(:read) ? raw_body.read.tap { raw_body.rewind if raw_body.respond_to?(:rewind) } : raw_body
|
|
66
|
+
body.to_s.b
|
|
67
|
+
end
|
|
68
|
+
|
|
69
|
+
# The v1 values of a "v1=<hex>,v1=<hex>" header; other schemes are ignored.
|
|
70
|
+
def candidates(signature)
|
|
71
|
+
signature.split(",").filter_map do |part|
|
|
72
|
+
scheme, value = part.strip.split("=", 2)
|
|
73
|
+
value.strip.downcase if scheme == "v1" && value
|
|
74
|
+
end
|
|
75
|
+
end
|
|
76
|
+
|
|
77
|
+
def parse(body)
|
|
78
|
+
payload = JSON.parse(body.dup.force_encoding(Encoding::UTF_8))
|
|
79
|
+
raise WebhookVerificationError, "The delivery body is not a JSON object." unless payload.is_a?(Hash)
|
|
80
|
+
|
|
81
|
+
payload
|
|
82
|
+
rescue JSON::ParserError => e
|
|
83
|
+
raise WebhookVerificationError, "The delivery body is not valid JSON: #{e.message}"
|
|
84
|
+
end
|
|
85
|
+
|
|
86
|
+
def header(headers, name)
|
|
87
|
+
normalized = normalize(name)
|
|
88
|
+
|
|
89
|
+
if headers.respond_to?(:[])
|
|
90
|
+
[name, name.downcase, "HTTP_#{normalized}"].each do |key|
|
|
91
|
+
value = safe_lookup(headers, key)
|
|
92
|
+
return header_value(value) unless value.nil?
|
|
93
|
+
end
|
|
94
|
+
end
|
|
95
|
+
|
|
96
|
+
return nil unless headers.respond_to?(:each)
|
|
97
|
+
|
|
98
|
+
headers.each do |key, value|
|
|
99
|
+
return header_value(value) if normalize(key.to_s) == normalized
|
|
100
|
+
end
|
|
101
|
+
nil
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
def normalize(key)
|
|
105
|
+
key.upcase.tr("-", "_").delete_prefix("HTTP_")
|
|
106
|
+
end
|
|
107
|
+
|
|
108
|
+
def safe_lookup(headers, key)
|
|
109
|
+
headers[key]
|
|
110
|
+
rescue StandardError
|
|
111
|
+
nil
|
|
112
|
+
end
|
|
113
|
+
|
|
114
|
+
def header_value(value)
|
|
115
|
+
value.is_a?(Array) ? value.join(",") : value.to_s
|
|
116
|
+
end
|
|
117
|
+
end
|
|
118
|
+
end
|
data/lib/urlpipe.rb
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require_relative "urlpipe/version"
|
|
4
|
+
require_relative "urlpipe/errors"
|
|
5
|
+
require_relative "urlpipe/response"
|
|
6
|
+
require_relative "urlpipe/screenshot"
|
|
7
|
+
require_relative "urlpipe/client"
|
|
8
|
+
require_relative "urlpipe/webhook"
|
|
9
|
+
|
|
10
|
+
# Ruby client for the URLpipe API (https://urlpipe.dev): turn a URL into
|
|
11
|
+
# Markdown, rendered HTML, a screenshot, metadata, a summary, keywords,
|
|
12
|
+
# console errors or a Lighthouse audit.
|
|
13
|
+
#
|
|
14
|
+
# client = Urlpipe::Client.new # reads URLPIPE_API_KEY
|
|
15
|
+
# client.markdown("https://example.com").data
|
|
16
|
+
module Urlpipe
|
|
17
|
+
end
|
metadata
ADDED
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
--- !ruby/object:Gem::Specification
|
|
2
|
+
name: urlpipe
|
|
3
|
+
version: !ruby/object:Gem::Version
|
|
4
|
+
version: 0.1.0
|
|
5
|
+
platform: ruby
|
|
6
|
+
authors:
|
|
7
|
+
- Aliat Partner S.L.
|
|
8
|
+
bindir: bin
|
|
9
|
+
cert_chain: []
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
11
|
+
dependencies: []
|
|
12
|
+
description: Official Ruby client for URLpipe. Get a page's Markdown, rendered HTML,
|
|
13
|
+
a full-page screenshot, metadata, a summary, keywords, console errors or a Lighthouse
|
|
14
|
+
audit, rendered in real Chrome. Handles long analyses, retries with idempotency
|
|
15
|
+
keys, and webhook signature verification. No runtime dependencies.
|
|
16
|
+
email:
|
|
17
|
+
- contact@urlpipe.dev
|
|
18
|
+
executables: []
|
|
19
|
+
extensions: []
|
|
20
|
+
extra_rdoc_files: []
|
|
21
|
+
files:
|
|
22
|
+
- CHANGELOG.md
|
|
23
|
+
- LICENSE
|
|
24
|
+
- README.md
|
|
25
|
+
- lib/urlpipe.rb
|
|
26
|
+
- lib/urlpipe/client.rb
|
|
27
|
+
- lib/urlpipe/errors.rb
|
|
28
|
+
- lib/urlpipe/response.rb
|
|
29
|
+
- lib/urlpipe/screenshot.rb
|
|
30
|
+
- lib/urlpipe/version.rb
|
|
31
|
+
- lib/urlpipe/webhook.rb
|
|
32
|
+
homepage: https://urlpipe.dev
|
|
33
|
+
licenses:
|
|
34
|
+
- MIT
|
|
35
|
+
metadata:
|
|
36
|
+
homepage_uri: https://urlpipe.dev
|
|
37
|
+
documentation_uri: https://urlpipe.dev/docs
|
|
38
|
+
source_code_uri: https://github.com/URLpipe/urlpipe-ruby
|
|
39
|
+
changelog_uri: https://github.com/URLpipe/urlpipe-ruby/blob/main/CHANGELOG.md
|
|
40
|
+
bug_tracker_uri: https://github.com/URLpipe/urlpipe-ruby/issues
|
|
41
|
+
rubygems_mfa_required: 'true'
|
|
42
|
+
rdoc_options: []
|
|
43
|
+
require_paths:
|
|
44
|
+
- lib
|
|
45
|
+
required_ruby_version: !ruby/object:Gem::Requirement
|
|
46
|
+
requirements:
|
|
47
|
+
- - ">="
|
|
48
|
+
- !ruby/object:Gem::Version
|
|
49
|
+
version: '3.1'
|
|
50
|
+
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
51
|
+
requirements:
|
|
52
|
+
- - ">="
|
|
53
|
+
- !ruby/object:Gem::Version
|
|
54
|
+
version: '0'
|
|
55
|
+
requirements: []
|
|
56
|
+
rubygems_version: 4.0.20
|
|
57
|
+
specification_version: 4
|
|
58
|
+
summary: 'Ruby client for the URLpipe API: turn a URL into Markdown, HTML, a screenshot
|
|
59
|
+
and more.'
|
|
60
|
+
test_files: []
|