gento 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Potentially problematic release.
This version of gento might be problematic. Click here for more details.
- checksums.yaml +7 -0
- data/CHANGELOG.md +53 -0
- data/LICENSE.txt +21 -0
- data/README.md +143 -0
- data/exe/gento +7 -0
- data/lib/gento/adapters/base.rb +121 -0
- data/lib/gento/adapters/docswell.rb +58 -0
- data/lib/gento/adapters/google_slides.rb +90 -0
- data/lib/gento/adapters/slide_share.rb +82 -0
- data/lib/gento/adapters/speaker_deck.rb +79 -0
- data/lib/gento/charset.rb +45 -0
- data/lib/gento/cli.rb +121 -0
- data/lib/gento/client.rb +55 -0
- data/lib/gento/deck.rb +56 -0
- data/lib/gento/entities.rb +29 -0
- data/lib/gento/errors.rb +46 -0
- data/lib/gento/fetcher.rb +188 -0
- data/lib/gento/html.rb +225 -0
- data/lib/gento/registry.rb +49 -0
- data/lib/gento/response.rb +34 -0
- data/lib/gento/robots.rb +195 -0
- data/lib/gento/robots_fetcher.rb +126 -0
- data/lib/gento/slide.rb +40 -0
- data/lib/gento/version.rb +5 -0
- data/lib/gento.rb +44 -0
- metadata +76 -0
checksums.yaml
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
SHA256:
|
|
3
|
+
metadata.gz: bb89ec6dceb5b98779d471fc330c3ac141fec9fe45bc870696f7b947496db6d2
|
|
4
|
+
data.tar.gz: dc3aee5c63e67925826b74f43d28144d96525f9bf6f7d3099ba83a7ef987cf99
|
|
5
|
+
SHA512:
|
|
6
|
+
metadata.gz: 97e1e10f67231d1cdf2ec41ddd061fd6e2d052daab3c9acd7a9c5fa451ab9ea3452af191e919259bf341b546ab5520c6b3ee52de16866845ad7158dfe58a82b9
|
|
7
|
+
data.tar.gz: 9ffbe5105b485bc84592630b9b206f6167365e4e8f30b0dd3536c269873794bb08b52eaf3b273b14d6b2d6273bd46e296fbafd7ed085925a9a44e71476714367
|
data/CHANGELOG.md
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
|
|
7
|
+
- `Gento::Client#fetch`, which turns a slide URL into a `Deck` of
|
|
8
|
+
ordered page images.
|
|
9
|
+
- Adapters for Speaker Deck, SlideShare, Docswell and public Google Slides.
|
|
10
|
+
- A pure-Ruby HTML scanner, so the gem carries no native extensions and no
|
|
11
|
+
runtime dependencies.
|
|
12
|
+
- A pluggable `Fetcher` seam, with a net/http implementation that guards
|
|
13
|
+
against redirects into private address ranges.
|
|
14
|
+
- Charset detection, so Shift_JIS and EUC-JP decks decode correctly.
|
|
15
|
+
- `gento` CLI, printing JSON or one image URL per line.
|
|
16
|
+
|
|
17
|
+
- robots.txt, obeyed by default (RFC 9309). A `RobotsFetcher` wraps whatever
|
|
18
|
+
fetcher is in use, so every request an adapter makes is checked, not just
|
|
19
|
+
the deck page. `robots: false`, `--no-robots` and `GENTO_ROBOTS=off`
|
|
20
|
+
turn it off. A robots.txt that cannot be read is treated as a refusal
|
|
21
|
+
rather than as permission.
|
|
22
|
+
|
|
23
|
+
- Proxy support in the net/http fetcher, honouring `HTTPS_PROXY` (which
|
|
24
|
+
net/http ignores on its own) and bypassing it for private addresses.
|
|
25
|
+
|
|
26
|
+
- The demo's viewer: a lightbox on clicking a page, with arrow-key paging,
|
|
27
|
+
and a switch between the thumbnail grid and a continuous vertical read.
|
|
28
|
+
The choice is remembered. Built on `<dialog>` with no dependencies, and
|
|
29
|
+
layered over markup that still works with JavaScript off.
|
|
30
|
+
- `GENTO_USER_AGENT`, so a deployed container can change how it
|
|
31
|
+
identifies itself without a rebuild.
|
|
32
|
+
|
|
33
|
+
### Fixed
|
|
34
|
+
|
|
35
|
+
- The tag scanner consumed element content, so a tag nested inside another of
|
|
36
|
+
the same name was invisible — which is most of the markup on a real page.
|
|
37
|
+
- Typographic entities such as `…` and `—` were left undecoded.
|
|
38
|
+
- Titles carrying the author's own line breaks are collapsed to single spaces.
|
|
39
|
+
|
|
40
|
+
- Google Slides decks published to the web return the signed `viewpage` URLs
|
|
41
|
+
embedded in the document. `/export/png` answers 404 for those decks, so the
|
|
42
|
+
URLs it would have built were unusable. Ordinary shared decks keep using
|
|
43
|
+
`/export/png`, which does not expire.
|
|
44
|
+
|
|
45
|
+
### Known gaps
|
|
46
|
+
|
|
47
|
+
- SlideShare's bot protection decides on the caller's IP reputation, so from
|
|
48
|
+
a datacenter address it can refuse every request. Detected and reported as
|
|
49
|
+
itself; getting past it needs a JavaScript-capable `Fetcher`.
|
|
50
|
+
- The signed `viewpage` URLs returned for published Google Slides decks
|
|
51
|
+
expire; how long they last has not been measured.
|
|
52
|
+
- The robots.txt check does not follow redirects: the wrapped fetcher handles
|
|
53
|
+
those internally, so a redirect into a disallowed path is not caught.
|
data/LICENSE.txt
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 ngram
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
data/README.md
ADDED
|
@@ -0,0 +1,143 @@
|
|
|
1
|
+
# gento
|
|
2
|
+
|
|
3
|
+
スライド共有サービスの URL を渡すと、各ページの画像 URL を順番に並べて返す Ruby gem です。
|
|
4
|
+
|
|
5
|
+
対応サービス: Speaker Deck ・ SlideShare ・ Docswell ・ Google スライド(公開設定のもの)
|
|
6
|
+
|
|
7
|
+
ランタイム依存ゼロ・ネイティブ拡張ゼロで動きます。
|
|
8
|
+
|
|
9
|
+
## インストール
|
|
10
|
+
|
|
11
|
+
```console
|
|
12
|
+
$ gem install gento
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Gemfile なら:
|
|
16
|
+
|
|
17
|
+
```ruby
|
|
18
|
+
gem "gento"
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
## 使いかた
|
|
22
|
+
|
|
23
|
+
```ruby
|
|
24
|
+
require "gento"
|
|
25
|
+
|
|
26
|
+
deck = Gento.fetch("https://speakerdeck.com/user/talk")
|
|
27
|
+
|
|
28
|
+
deck.title # => "快適なスライド閲覧生活を実現する Web サービスの開発"
|
|
29
|
+
deck.author # => "ngram"
|
|
30
|
+
deck.page_count # => 42
|
|
31
|
+
deck.slides.map(&:url)
|
|
32
|
+
# => ["https://files.speakerdeck.com/.../slide_0.jpg", ...]
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
`Deck` と `Slide` は不変オブジェクトで、`to_h` / `to_json` でそのまま出力できます。
|
|
36
|
+
|
|
37
|
+
対応している URL かどうかは、取得する前に確認できます。
|
|
38
|
+
|
|
39
|
+
```ruby
|
|
40
|
+
Gento.supports?("https://speakerdeck.com/user/talk") # => true
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
### コマンドライン
|
|
44
|
+
|
|
45
|
+
```console
|
|
46
|
+
$ gento https://speakerdeck.com/user/talk
|
|
47
|
+
{
|
|
48
|
+
"provider": "speaker_deck",
|
|
49
|
+
"source_url": "https://speakerdeck.com/user/talk",
|
|
50
|
+
"title": "...",
|
|
51
|
+
"page_count": 42,
|
|
52
|
+
"slides": [ ... ]
|
|
53
|
+
}
|
|
54
|
+
|
|
55
|
+
$ gento --format urls https://speakerdeck.com/user/talk
|
|
56
|
+
https://files.speakerdeck.com/presentations/.../slide_0.jpg
|
|
57
|
+
...
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
### robots.txt
|
|
61
|
+
|
|
62
|
+
**取得先の `robots.txt` を既定で参照し、拒否されているパスは取りに行きません。**
|
|
63
|
+
デッキページだけでなく、アダプタが出すすべてのリクエスト(oEmbed、embed ビューなど)が対象です。
|
|
64
|
+
|
|
65
|
+
```ruby
|
|
66
|
+
Gento.fetch(url) # robots.txt に従う(既定)
|
|
67
|
+
Gento.fetch(url, robots: false) # 従わない(判断は利用者の責任)
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
```console
|
|
71
|
+
$ gento --no-robots https://speakerdeck.com/user/talk
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
拒否された場合は `Gento::RobotsDisallowedError` になります。
|
|
75
|
+
`robots.txt` 自体を読めなかったとき(5xx・429・接続失敗)も、RFC 9309 に従って拒否扱いです。
|
|
76
|
+
|
|
77
|
+
詳細は [doc/robots.md](doc/robots.md) を参照してください。
|
|
78
|
+
|
|
79
|
+
### HTTP クライアントの差し替え
|
|
80
|
+
|
|
81
|
+
通信はすべて `Gento::Fetcher` を経由するので、独自の HTTP スタックを差し込めます。
|
|
82
|
+
|
|
83
|
+
```ruby
|
|
84
|
+
Gento.fetch(url, fetcher: MyFetcher.new)
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
詳細は [doc/design.md](doc/design.md) を参照してください。
|
|
88
|
+
|
|
89
|
+
### エラー
|
|
90
|
+
|
|
91
|
+
| 例外 | 起きるとき |
|
|
92
|
+
| --- | --- |
|
|
93
|
+
| `Gento::UnsupportedURLError` | 対応アダプタのない URL |
|
|
94
|
+
| `Gento::RobotsDisallowedError` | `robots.txt` が許可していない(`FetchError` の一種) |
|
|
95
|
+
| `Gento::FetchError` | 取得に失敗(`ResponseError` を含む) |
|
|
96
|
+
| `Gento::ExtractionError` | 取得はできたがページを取り出せない |
|
|
97
|
+
|
|
98
|
+
いずれも `Gento::Error` を継承しています。
|
|
99
|
+
|
|
100
|
+
**SlideShare は接続元 IP によってボット判定で弾かれることがあり**、その場合は
|
|
101
|
+
`ExtractionError` になります。データセンターの IP からは恒常的に弾かれます
|
|
102
|
+
([doc/verification.md](doc/verification.md#slideshare-のボット判定について))。
|
|
103
|
+
|
|
104
|
+
## 利用にあたっての注意
|
|
105
|
+
|
|
106
|
+
**対象サービスの利用規約を確認し、遵守する責任は利用者にあります。**
|
|
107
|
+
[doc/legal.md](doc/legal.md) を読んでから使ってください。
|
|
108
|
+
|
|
109
|
+
## ドキュメント
|
|
110
|
+
|
|
111
|
+
- [robots.txt の扱い](doc/robots.md) — 判定ルール、無効化、各サービスの実際の内容
|
|
112
|
+
- [設計方針](doc/design.md) — 依存ゼロにした理由、Fetcher の差し替え、アダプタの足し方
|
|
113
|
+
- [検証状況と既知の制限](doc/verification.md) — サービスごとの注意点、SlideShare のボット判定
|
|
114
|
+
- [手元で試す](doc/development.md) — `main` を取得して gem・CLI・デモアプリを動かす
|
|
115
|
+
- [デモ Web サービス](doc/demo.md) — `web/` の動かしかたと Cloudflare へのデプロイ
|
|
116
|
+
- [利用にあたっての注意](doc/legal.md)
|
|
117
|
+
- [RubyGems への公開手順](doc/publishing.md)
|
|
118
|
+
|
|
119
|
+
## 開発
|
|
120
|
+
|
|
121
|
+
```console
|
|
122
|
+
$ bundle install
|
|
123
|
+
$ bundle exec rspec # gem 151 examples
|
|
124
|
+
$ bundle exec rubocop
|
|
125
|
+
$ (cd web && bundle exec rspec) # デモアプリ 27 examples
|
|
126
|
+
$ (cd worker && npm test && npm run typecheck) # Worker 15 tests
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
GitHub Codespaces なら `.devcontainer/` が上記をすべて用意します
|
|
130
|
+
([doc/codespaces.md](doc/codespaces.md))。
|
|
131
|
+
|
|
132
|
+
## 由来
|
|
133
|
+
|
|
134
|
+
技術書典5 で頒布した『てっくやみなべ Vol.1』所収
|
|
135
|
+
「快適なスライド閲覧⽣活を実現する Web サービスの開発」のサポートリポジトリを、
|
|
136
|
+
現在も動く形で作り直したものです。当時の Lambda / PaaS 向けデプロイコードは、
|
|
137
|
+
対象サービスが終了したため Cloudflare 向けに置き換えています。
|
|
138
|
+
|
|
139
|
+
名前の gento は、スライドを映写して見せる「幻灯(げんとう)」から取っています。
|
|
140
|
+
|
|
141
|
+
## ライセンス
|
|
142
|
+
|
|
143
|
+
MIT License. [LICENSE.txt](LICENSE.txt) を参照してください。
|
data/exe/gento
ADDED
|
@@ -0,0 +1,121 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "uri"
|
|
4
|
+
|
|
5
|
+
module Gento
|
|
6
|
+
module Adapters
|
|
7
|
+
# Shared behaviour for the per-site adapters.
|
|
8
|
+
#
|
|
9
|
+
# An adapter answers two questions: "is this URL mine?" and "what are the
|
|
10
|
+
# pages of this deck?". Everything else — HTTP, value objects, JSON output
|
|
11
|
+
# — is the framework's job.
|
|
12
|
+
class Base
|
|
13
|
+
class << self
|
|
14
|
+
# Host suffixes this adapter claims, e.g. %w[speakerdeck.com].
|
|
15
|
+
def hosts
|
|
16
|
+
raise NotImplementedError, "#{self} must declare .hosts"
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
def provider
|
|
20
|
+
name.split("::").last
|
|
21
|
+
.gsub(/([a-z\d])([A-Z])/, '\1_\2')
|
|
22
|
+
.downcase
|
|
23
|
+
end
|
|
24
|
+
|
|
25
|
+
def handles?(url)
|
|
26
|
+
uri = coerce_uri(url)
|
|
27
|
+
return false unless uri&.host
|
|
28
|
+
|
|
29
|
+
host = uri.host.downcase.sub(/\Awww\./, "")
|
|
30
|
+
hosts.any? { |claimed| host == claimed || host.end_with?(".#{claimed}") }
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
def coerce_uri(url)
|
|
34
|
+
uri = url.is_a?(URI) ? url : URI.parse(url.to_s)
|
|
35
|
+
uri.is_a?(URI::HTTP) ? uri : nil
|
|
36
|
+
rescue URI::InvalidURIError
|
|
37
|
+
nil
|
|
38
|
+
end
|
|
39
|
+
end
|
|
40
|
+
|
|
41
|
+
attr_reader :fetcher
|
|
42
|
+
|
|
43
|
+
def initialize(fetcher:)
|
|
44
|
+
@fetcher = fetcher
|
|
45
|
+
end
|
|
46
|
+
|
|
47
|
+
# @return [Gento::Deck]
|
|
48
|
+
def fetch(url)
|
|
49
|
+
raise NotImplementedError, "#{self.class} must implement #fetch"
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
private
|
|
53
|
+
|
|
54
|
+
def provider
|
|
55
|
+
self.class.provider
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
def get(url, headers: {})
|
|
59
|
+
fetcher.get(url, headers: headers)
|
|
60
|
+
end
|
|
61
|
+
|
|
62
|
+
def absolute_url(candidate, base)
|
|
63
|
+
return nil if candidate.nil? || candidate.to_s.empty?
|
|
64
|
+
|
|
65
|
+
URI.join(base.to_s, candidate.to_s).to_s
|
|
66
|
+
rescue URI::Error
|
|
67
|
+
nil
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
# Title/author/description as advertised by OpenGraph, which every one of
|
|
71
|
+
# these sites emits and keeps stable for the sake of link previews.
|
|
72
|
+
def open_graph_metadata(html)
|
|
73
|
+
{
|
|
74
|
+
title: html.meta("og:title", "twitter:title") || html.title,
|
|
75
|
+
description: html.meta("og:description", "description"),
|
|
76
|
+
author: html.meta("og:author", "author", "article:author")
|
|
77
|
+
}
|
|
78
|
+
end
|
|
79
|
+
|
|
80
|
+
# The oEmbed endpoint a page advertises in its head, if any.
|
|
81
|
+
#
|
|
82
|
+
# Every site here publishes one, so discovering it beats hardcoding a
|
|
83
|
+
# URL per adapter: the site tells us where its own endpoint lives.
|
|
84
|
+
def oembed_endpoint(html)
|
|
85
|
+
html.tags("link")
|
|
86
|
+
.find { |attrs| attrs["type"].to_s.include?("json+oembed") }
|
|
87
|
+
&.fetch("href", nil)
|
|
88
|
+
end
|
|
89
|
+
|
|
90
|
+
# oEmbed metadata, or an empty hash.
|
|
91
|
+
#
|
|
92
|
+
# It is always a nice-to-have: cleaner titles and real author names than
|
|
93
|
+
# OpenGraph offers, but never something a fetch should fail over.
|
|
94
|
+
def fetch_oembed(html)
|
|
95
|
+
endpoint = oembed_endpoint(html)
|
|
96
|
+
return {} unless endpoint
|
|
97
|
+
|
|
98
|
+
get(endpoint, headers: { "accept" => "application/json" }).json
|
|
99
|
+
rescue FetchError, ExtractionError
|
|
100
|
+
{}
|
|
101
|
+
end
|
|
102
|
+
|
|
103
|
+
def build_slides(urls, width: nil, height: nil)
|
|
104
|
+
urls.each_with_index.map do |url, index|
|
|
105
|
+
Slide.new(number: index + 1, url: url, width: width, height: height)
|
|
106
|
+
end
|
|
107
|
+
end
|
|
108
|
+
|
|
109
|
+
# The same image is often referenced at several sizes through a query
|
|
110
|
+
# string (`?width=160`, a cache-busting timestamp). Dropping the query
|
|
111
|
+
# before deduplicating collapses those into one page.
|
|
112
|
+
def without_query(url)
|
|
113
|
+
url.to_s.split("?").first.to_s
|
|
114
|
+
end
|
|
115
|
+
|
|
116
|
+
def fail_extraction(message)
|
|
117
|
+
raise ExtractionError, "#{provider}: #{message}"
|
|
118
|
+
end
|
|
119
|
+
end
|
|
120
|
+
end
|
|
121
|
+
end
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Gento
|
|
4
|
+
module Adapters
|
|
5
|
+
# Docswell (https://www.docswell.com/s/<user>/<id>-<slug>).
|
|
6
|
+
#
|
|
7
|
+
# The deck page only renders the first handful of pages; the rest arrive
|
|
8
|
+
# by script. The embed view it links to carries the whole deck, so that is
|
|
9
|
+
# what this adapter reads. Each page appears there twice, full size and as
|
|
10
|
+
# a `?width=160` thumbnail, which dropping the query collapses.
|
|
11
|
+
class Docswell < Base
|
|
12
|
+
EMBED_URL = %r{\Ahttps://(?:www\.)?docswell\.com/slide/[A-Za-z0-9]+/embed}i
|
|
13
|
+
PAGE_IMAGE = %r{\Ahttps://[a-z0-9.-]*docswell\.com/page/[^/]+\.(?:jpe?g|png|webp)}i
|
|
14
|
+
|
|
15
|
+
def self.hosts
|
|
16
|
+
%w[docswell.com]
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
def fetch(url)
|
|
20
|
+
page = get(url).html
|
|
21
|
+
|
|
22
|
+
# The page links the embed both plainly and with a ?mode= variant.
|
|
23
|
+
embed_url = page.urls(EMBED_URL).map { |found| without_query(found) }.first
|
|
24
|
+
fail_extraction("could not find the embed view for #{url}") unless embed_url
|
|
25
|
+
|
|
26
|
+
images = slide_images(embed_url)
|
|
27
|
+
fail_extraction("found no slide images for #{url}") if images.empty?
|
|
28
|
+
|
|
29
|
+
build_deck(url, page, images)
|
|
30
|
+
end
|
|
31
|
+
|
|
32
|
+
private
|
|
33
|
+
|
|
34
|
+
# Page filenames are opaque ids, so document order is the only ordering
|
|
35
|
+
# available. The embed view lists the pages in slide order, which the
|
|
36
|
+
# deck's og:image corroborates: it is always the first page.
|
|
37
|
+
def slide_images(embed_url)
|
|
38
|
+
get(embed_url).html.urls(PAGE_IMAGE).map { |url| without_query(url) }.uniq
|
|
39
|
+
end
|
|
40
|
+
|
|
41
|
+
# The deck page carries no author metadata at all, so oEmbed is the only
|
|
42
|
+
# place a name comes from here.
|
|
43
|
+
def build_deck(url, page, images)
|
|
44
|
+
metadata = open_graph_metadata(page)
|
|
45
|
+
oembed = fetch_oembed(page)
|
|
46
|
+
|
|
47
|
+
Deck.new(
|
|
48
|
+
provider: provider,
|
|
49
|
+
source_url: url,
|
|
50
|
+
title: oembed["title"] || metadata[:title],
|
|
51
|
+
author: oembed["author_name"] || metadata[:author],
|
|
52
|
+
description: metadata[:description],
|
|
53
|
+
slides: build_slides(images)
|
|
54
|
+
)
|
|
55
|
+
end
|
|
56
|
+
end
|
|
57
|
+
end
|
|
58
|
+
end
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Gento
|
|
4
|
+
module Adapters
|
|
5
|
+
# Google Slides, for decks shared publicly or published to the web.
|
|
6
|
+
#
|
|
7
|
+
# Google does not render page images into the deck page, so the job here
|
|
8
|
+
# is to find the ordered list of pages and name an image for each.
|
|
9
|
+
#
|
|
10
|
+
# The viewer at /embed is the obvious place to look and the wrong one: its
|
|
11
|
+
# payload mixes page ids with the ids of elements drawn on those pages,
|
|
12
|
+
# with nothing in the id itself to tell them apart. /htmlpresent instead
|
|
13
|
+
# references each page exactly once, in order, so that is what we read.
|
|
14
|
+
class GoogleSlides < Base
|
|
15
|
+
BASE_URL = "https://docs.google.com/presentation/d/%s"
|
|
16
|
+
|
|
17
|
+
# /presentation/d/<id>/... and the published /presentation/d/e/<id>/...
|
|
18
|
+
# Published ids contain hyphens, ordinary ones do not.
|
|
19
|
+
PRESENTATION_ID = %r{/presentation/d/(e/)?([A-Za-z0-9_-]{16,})}
|
|
20
|
+
PAGE_ID = /[?&]pageid=([A-Za-z0-9_]+)/
|
|
21
|
+
VIEWPAGE = %r{/viewpage\?}
|
|
22
|
+
|
|
23
|
+
def self.hosts
|
|
24
|
+
%w[docs.google.com]
|
|
25
|
+
end
|
|
26
|
+
|
|
27
|
+
def fetch(url)
|
|
28
|
+
published, id = presentation_id(url)
|
|
29
|
+
fail_extraction("could not find a presentation id in #{url}") unless id
|
|
30
|
+
|
|
31
|
+
page = get("#{deck_base(published, id)}/htmlpresent").html
|
|
32
|
+
images = published ? viewpage_images(page, url) : export_images(page, url, published, id)
|
|
33
|
+
|
|
34
|
+
Deck.new(
|
|
35
|
+
provider: provider,
|
|
36
|
+
source_url: url,
|
|
37
|
+
title: title_from(page),
|
|
38
|
+
slides: build_slides(images)
|
|
39
|
+
)
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
private
|
|
43
|
+
|
|
44
|
+
def deck_base(published, id)
|
|
45
|
+
format(BASE_URL, published ? "e/#{id}" : id)
|
|
46
|
+
end
|
|
47
|
+
|
|
48
|
+
# Ordinary decks: build export URLs from the page ids. These are stable,
|
|
49
|
+
# so a caller can keep them, and they redirect to a signed image.
|
|
50
|
+
def export_images(page, url, published, id)
|
|
51
|
+
base = deck_base(published, id)
|
|
52
|
+
|
|
53
|
+
page_ids(page, url).map { |page_id| "#{base}/export/png?id=#{id}&pageid=#{page_id}" }
|
|
54
|
+
end
|
|
55
|
+
|
|
56
|
+
# Published decks: /export/png answers 404 for them, whatever page id it
|
|
57
|
+
# is given. The only thing that serves their pages is the viewpage URL
|
|
58
|
+
# embedded in the document, which carries its own signature — so take it
|
|
59
|
+
# verbatim rather than trying to construct one.
|
|
60
|
+
#
|
|
61
|
+
# Those signatures expire, so unlike the export URLs these are worth
|
|
62
|
+
# fetching promptly rather than storing.
|
|
63
|
+
def viewpage_images(page, url)
|
|
64
|
+
images = page.urls(VIEWPAGE)
|
|
65
|
+
fail_extraction("found no slide pages for #{url}; is the deck still published?") if images.empty?
|
|
66
|
+
|
|
67
|
+
images
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
# Each page is referenced exactly once here, in order.
|
|
71
|
+
def page_ids(page, url)
|
|
72
|
+
ids = page.source.scan(PAGE_ID).flatten.uniq
|
|
73
|
+
fail_extraction("found no slide pages for #{url}; is the deck shared publicly?") if ids.empty?
|
|
74
|
+
|
|
75
|
+
ids
|
|
76
|
+
end
|
|
77
|
+
|
|
78
|
+
def presentation_id(url)
|
|
79
|
+
match = url.to_s.match(PRESENTATION_ID)
|
|
80
|
+
return [false, nil] unless match
|
|
81
|
+
|
|
82
|
+
[!match[1].nil?, match[2]]
|
|
83
|
+
end
|
|
84
|
+
|
|
85
|
+
def title_from(page)
|
|
86
|
+
page.title&.sub(/\s*-\s*Google (?:Slides|スライド)\s*\z/, "")
|
|
87
|
+
end
|
|
88
|
+
end
|
|
89
|
+
end
|
|
90
|
+
end
|
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Gento
|
|
4
|
+
module Adapters
|
|
5
|
+
# SlideShare (https://www.slideshare.net/<user>/<slug>).
|
|
6
|
+
#
|
|
7
|
+
# SlideShare usually serves the deck to an ordinary HTTP client, but it
|
|
8
|
+
# sits behind bot protection that intermittently answers with a small
|
|
9
|
+
# JavaScript interstitial instead — most readily under bursty access.
|
|
10
|
+
#
|
|
11
|
+
# That interstitial has no slides in it, so without recognising it the
|
|
12
|
+
# adapter would report an empty deck and leave the caller guessing. Naming
|
|
13
|
+
# it costs one check and turns a mystery into something actionable: wait,
|
|
14
|
+
# or supply a JavaScript-capable Fetcher.
|
|
15
|
+
class SlideShare < Base
|
|
16
|
+
CHALLENGE = /Client Challenge|_fs-ch-/
|
|
17
|
+
# <slug>-<page>-<width>.jpg on the slide CDN.
|
|
18
|
+
SLIDE_IMAGE = %r{\Ahttps://[a-z0-9.-]*slidesharecdn\.com/\S+-(\d+)-(\d+)\.(?:jpe?g|png|webp)}i
|
|
19
|
+
|
|
20
|
+
def self.hosts
|
|
21
|
+
%w[slideshare.net]
|
|
22
|
+
end
|
|
23
|
+
|
|
24
|
+
def fetch(url)
|
|
25
|
+
page = get(url).html
|
|
26
|
+
fail_challenge(url) if challenge?(page)
|
|
27
|
+
|
|
28
|
+
images = slide_images(page)
|
|
29
|
+
fail_extraction("found no slide images at #{url}") if images.empty?
|
|
30
|
+
|
|
31
|
+
build_deck(url, page, images)
|
|
32
|
+
end
|
|
33
|
+
|
|
34
|
+
private
|
|
35
|
+
|
|
36
|
+
def challenge?(page)
|
|
37
|
+
page.title.to_s.match?(CHALLENGE) || page.source.match?(CHALLENGE)
|
|
38
|
+
end
|
|
39
|
+
|
|
40
|
+
def fail_challenge(url)
|
|
41
|
+
fail_extraction(
|
|
42
|
+
"#{url} returned SlideShare's JavaScript bot challenge instead of the deck. " \
|
|
43
|
+
"Fetching from SlideShare needs a JavaScript-capable Gento::Fetcher; " \
|
|
44
|
+
"the default net/http one cannot get past it."
|
|
45
|
+
)
|
|
46
|
+
end
|
|
47
|
+
|
|
48
|
+
# Pages are published at several widths, not always the same set for
|
|
49
|
+
# every page, so keep the widest copy of each.
|
|
50
|
+
def slide_images(page)
|
|
51
|
+
page.urls(SLIDE_IMAGE)
|
|
52
|
+
.group_by { |image| Integer(image[SLIDE_IMAGE, 1]) }
|
|
53
|
+
.sort_by(&:first)
|
|
54
|
+
.map { |_, group| group.max_by { |image| Integer(image[SLIDE_IMAGE, 2]) } }
|
|
55
|
+
end
|
|
56
|
+
|
|
57
|
+
def build_deck(url, page, images)
|
|
58
|
+
metadata = open_graph_metadata(page)
|
|
59
|
+
document = json_ld_document(page)
|
|
60
|
+
|
|
61
|
+
Deck.new(
|
|
62
|
+
provider: provider,
|
|
63
|
+
source_url: url,
|
|
64
|
+
title: metadata[:title],
|
|
65
|
+
author: author_from(document) || metadata[:author],
|
|
66
|
+
description: metadata[:description],
|
|
67
|
+
published_at: document["datePublished"],
|
|
68
|
+
slides: build_slides(images)
|
|
69
|
+
)
|
|
70
|
+
end
|
|
71
|
+
|
|
72
|
+
def json_ld_document(page)
|
|
73
|
+
page.json_ld.find { |node| node["@type"].to_s.match?(/Presentation|Article|CreativeWork/i) } || {}
|
|
74
|
+
end
|
|
75
|
+
|
|
76
|
+
def author_from(document)
|
|
77
|
+
author = document["author"]
|
|
78
|
+
author.is_a?(Hash) ? author["name"] : author
|
|
79
|
+
end
|
|
80
|
+
end
|
|
81
|
+
end
|
|
82
|
+
end
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Gento
|
|
4
|
+
module Adapters
|
|
5
|
+
# Speaker Deck (https://speakerdeck.com/<user>/<slug>).
|
|
6
|
+
#
|
|
7
|
+
# The deck page carries every page image at full size, in <a href>
|
|
8
|
+
# elements. It also carries cover images for roughly thirty recommended
|
|
9
|
+
# decks, so matching on the CDN hostname alone would silently splice other
|
|
10
|
+
# people's slides into the result. Keying on this deck's own presentation
|
|
11
|
+
# id is what keeps them out.
|
|
12
|
+
class SpeakerDeck < Base
|
|
13
|
+
PLAYER_ID = %r{/player/([0-9a-f-]+)}i
|
|
14
|
+
DECK_ID = /\A[0-9a-f-]{8,}\z/i
|
|
15
|
+
|
|
16
|
+
def self.hosts
|
|
17
|
+
%w[speakerdeck.com]
|
|
18
|
+
end
|
|
19
|
+
|
|
20
|
+
def fetch(url)
|
|
21
|
+
page = get(url).html
|
|
22
|
+
oembed = fetch_oembed(page)
|
|
23
|
+
|
|
24
|
+
id = deck_id(page, oembed)
|
|
25
|
+
fail_extraction("could not find the presentation id for #{url}") unless id
|
|
26
|
+
|
|
27
|
+
images = slide_images(page, id)
|
|
28
|
+
fail_extraction("found no slide images at #{url}") if images.empty?
|
|
29
|
+
|
|
30
|
+
build_deck(url, page, oembed, images)
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
private
|
|
34
|
+
|
|
35
|
+
# Full-size pages only: the same presentation also publishes
|
|
36
|
+
# `preview_slide_N.jpg` thumbnails, which this pattern excludes.
|
|
37
|
+
def slide_images(page, id)
|
|
38
|
+
pattern = %r{
|
|
39
|
+
\Ahttps://files\.speakerdeck\.com/presentations/
|
|
40
|
+
#{Regexp.escape(id)}/slide_(\d+)\.[a-z]+
|
|
41
|
+
}xi
|
|
42
|
+
|
|
43
|
+
page.urls(pattern)
|
|
44
|
+
.map { |url| without_query(url) }
|
|
45
|
+
.uniq
|
|
46
|
+
.sort_by { |url| Integer(url[pattern, 1]) }
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
# oEmbed names the id inside a player URL; the page itself carries it
|
|
50
|
+
# bare, as the data-id of the embed placeholder.
|
|
51
|
+
def deck_id(page, oembed)
|
|
52
|
+
player_id(oembed["html"]) || embed_data_id(page) || player_id(page.source)
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
def player_id(text)
|
|
56
|
+
text.to_s[PLAYER_ID, 1]
|
|
57
|
+
end
|
|
58
|
+
|
|
59
|
+
def embed_data_id(page)
|
|
60
|
+
page.tags("div")
|
|
61
|
+
.filter_map { |attrs| attrs["data-id"] }
|
|
62
|
+
.find { |id| id.match?(DECK_ID) }
|
|
63
|
+
end
|
|
64
|
+
|
|
65
|
+
def build_deck(url, page, oembed, images)
|
|
66
|
+
metadata = open_graph_metadata(page)
|
|
67
|
+
|
|
68
|
+
Deck.new(
|
|
69
|
+
provider: provider,
|
|
70
|
+
source_url: url,
|
|
71
|
+
title: oembed["title"] || metadata[:title],
|
|
72
|
+
author: oembed["author_name"] || metadata[:author],
|
|
73
|
+
description: metadata[:description],
|
|
74
|
+
slides: build_slides(images)
|
|
75
|
+
)
|
|
76
|
+
end
|
|
77
|
+
end
|
|
78
|
+
end
|
|
79
|
+
end
|