gento 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: 531f3b054d96ec2e003e9dce8a1c25e45864231ab11389b8c53a8550adc289ba
4
+ data.tar.gz: ba66af96ea1af153b5dd51026690c90ca32584975960551d589d2563422e35d0
5
+ SHA512:
6
+ metadata.gz: a35d8e3d3979b56d4b674fc222c80099696192bab6fc8431da01926436a9ce8fba61ec54b05b0e12b3b94f7af84cc0041c3709fe0884f595575bb3da4f2df890
7
+ data.tar.gz: 33cc2a667bed4a1ba8b1518695725b88321c8d256b24198bdbcf0b18808b5c43611bc29b6b508e7752914b29058bba0eeba560f0236143c8d0855f69dd75b4ae
data/CHANGELOG.md ADDED
@@ -0,0 +1,60 @@
1
+ # Changelog
2
+
3
+ ## 0.1.1
4
+
5
+ ### Changed
6
+
7
+ - The gemspec no longer lists a contact email. Questions and bug reports go
8
+ to the issue tracker linked from the gem's page.
9
+
10
+ ## 0.1.0
11
+
12
+ ### Added
13
+
14
+ - `Gento::Client#fetch`, which turns a slide URL into a `Deck` of
15
+ ordered page images.
16
+ - Adapters for Speaker Deck, SlideShare, Docswell and public Google Slides.
17
+ - A pure-Ruby HTML scanner, so the gem carries no native extensions and no
18
+ runtime dependencies.
19
+ - A pluggable `Fetcher` seam, with a net/http implementation that guards
20
+ against redirects into private address ranges.
21
+ - Charset detection, so Shift_JIS and EUC-JP decks decode correctly.
22
+ - `gento` CLI, printing JSON or one image URL per line.
23
+
24
+ - robots.txt, obeyed by default (RFC 9309). A `RobotsFetcher` wraps whatever
25
+ fetcher is in use, so every request an adapter makes is checked, not just
26
+ the deck page. `robots: false`, `--no-robots` and `GENTO_ROBOTS=off`
27
+ turn it off. A robots.txt that cannot be read is treated as a refusal
28
+ rather than as permission.
29
+
30
+ - Proxy support in the net/http fetcher, honouring `HTTPS_PROXY` (which
31
+ net/http ignores on its own) and bypassing it for private addresses.
32
+
33
+ - The demo's viewer: a lightbox on clicking a page, with arrow-key paging,
34
+ and a switch between the thumbnail grid and a continuous vertical read.
35
+ The choice is remembered. Built on `<dialog>` with no dependencies, and
36
+ layered over markup that still works with JavaScript off.
37
+ - `GENTO_USER_AGENT`, so a deployed container can change how it
38
+ identifies itself without a rebuild.
39
+
40
+ ### Fixed
41
+
42
+ - The tag scanner consumed element content, so a tag nested inside another of
43
+ the same name was invisible — which is most of the markup on a real page.
44
+ - Typographic entities such as `&hellip;` and `&mdash;` were left undecoded.
45
+ - Titles carrying the author's own line breaks are collapsed to single spaces.
46
+
47
+ - Google Slides decks published to the web return the signed `viewpage` URLs
48
+ embedded in the document. `/export/png` answers 404 for those decks, so the
49
+ URLs it would have built were unusable. Ordinary shared decks keep using
50
+ `/export/png`, which does not expire.
51
+
52
+ ### Known gaps
53
+
54
+ - SlideShare's bot protection decides on the caller's IP reputation, so from
55
+ a datacenter address it can refuse every request. Detected and reported as
56
+ itself; getting past it needs a JavaScript-capable `Fetcher`.
57
+ - The signed `viewpage` URLs returned for published Google Slides decks
58
+ expire; how long they last has not been measured.
59
+ - The robots.txt check does not follow redirects: the wrapped fetcher handles
60
+ those internally, so a redirect into a disallowed path is not caught.
data/LICENSE.txt ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 ngram
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,143 @@
1
+ # gento
2
+
3
+ スライド共有サービスの URL を渡すと、各ページの画像 URL を順番に並べて返す Ruby gem です。
4
+
5
+ 対応サービス: Speaker Deck ・ SlideShare ・ Docswell ・ Google スライド(公開設定のもの)
6
+
7
+ ランタイム依存ゼロ・ネイティブ拡張ゼロで動きます。
8
+
9
+ ## インストール
10
+
11
+ ```console
12
+ $ gem install gento
13
+ ```
14
+
15
+ Gemfile なら:
16
+
17
+ ```ruby
18
+ gem "gento"
19
+ ```
20
+
21
+ ## 使いかた
22
+
23
+ ```ruby
24
+ require "gento"
25
+
26
+ deck = Gento.fetch("https://speakerdeck.com/user/talk")
27
+
28
+ deck.title # => "快適なスライド閲覧生活を実現する Web サービスの開発"
29
+ deck.author # => "ngram"
30
+ deck.page_count # => 42
31
+ deck.slides.map(&:url)
32
+ # => ["https://files.speakerdeck.com/.../slide_0.jpg", ...]
33
+ ```
34
+
35
+ `Deck` と `Slide` は不変オブジェクトで、`to_h` / `to_json` でそのまま出力できます。
36
+
37
+ 対応している URL かどうかは、取得する前に確認できます。
38
+
39
+ ```ruby
40
+ Gento.supports?("https://speakerdeck.com/user/talk") # => true
41
+ ```
42
+
43
+ ### コマンドライン
44
+
45
+ ```console
46
+ $ gento https://speakerdeck.com/user/talk
47
+ {
48
+ "provider": "speaker_deck",
49
+ "source_url": "https://speakerdeck.com/user/talk",
50
+ "title": "...",
51
+ "page_count": 42,
52
+ "slides": [ ... ]
53
+ }
54
+
55
+ $ gento --format urls https://speakerdeck.com/user/talk
56
+ https://files.speakerdeck.com/presentations/.../slide_0.jpg
57
+ ...
58
+ ```
59
+
60
+ ### robots.txt
61
+
62
+ **取得先の `robots.txt` を既定で参照し、拒否されているパスは取りに行きません。**
63
+ デッキページだけでなく、アダプタが出すすべてのリクエスト(oEmbed、embed ビューなど)が対象です。
64
+
65
+ ```ruby
66
+ Gento.fetch(url) # robots.txt に従う(既定)
67
+ Gento.fetch(url, robots: false) # 従わない(判断は利用者の責任)
68
+ ```
69
+
70
+ ```console
71
+ $ gento --no-robots https://speakerdeck.com/user/talk
72
+ ```
73
+
74
+ 拒否された場合は `Gento::RobotsDisallowedError` になります。
75
+ `robots.txt` 自体を読めなかったとき(5xx・429・接続失敗)も、RFC 9309 に従って拒否扱いです。
76
+
77
+ 詳細は [doc/robots.md](doc/robots.md) を参照してください。
78
+
79
+ ### HTTP クライアントの差し替え
80
+
81
+ 通信はすべて `Gento::Fetcher` を経由するので、独自の HTTP スタックを差し込めます。
82
+
83
+ ```ruby
84
+ Gento.fetch(url, fetcher: MyFetcher.new)
85
+ ```
86
+
87
+ 詳細は [doc/design.md](doc/design.md) を参照してください。
88
+
89
+ ### エラー
90
+
91
+ | 例外 | 起きるとき |
92
+ | --- | --- |
93
+ | `Gento::UnsupportedURLError` | 対応アダプタのない URL |
94
+ | `Gento::RobotsDisallowedError` | `robots.txt` が許可していない(`FetchError` の一種) |
95
+ | `Gento::FetchError` | 取得に失敗(`ResponseError` を含む) |
96
+ | `Gento::ExtractionError` | 取得はできたがページを取り出せない |
97
+
98
+ いずれも `Gento::Error` を継承しています。
99
+
100
+ **SlideShare は接続元 IP によってボット判定で弾かれることがあり**、その場合は
101
+ `ExtractionError` になります。データセンターの IP からは恒常的に弾かれます
102
+ ([doc/verification.md](doc/verification.md#slideshare-のボット判定について))。
103
+
104
+ ## 利用にあたっての注意
105
+
106
+ **対象サービスの利用規約を確認し、遵守する責任は利用者にあります。**
107
+ [doc/legal.md](doc/legal.md) を読んでから使ってください。
108
+
109
+ ## ドキュメント
110
+
111
+ - [robots.txt の扱い](doc/robots.md) — 判定ルール、無効化、各サービスの実際の内容
112
+ - [設計方針](doc/design.md) — 依存ゼロにした理由、Fetcher の差し替え、アダプタの足し方
113
+ - [検証状況と既知の制限](doc/verification.md) — サービスごとの注意点、SlideShare のボット判定
114
+ - [手元で試す](doc/development.md) — `main` を取得して gem・CLI・デモアプリを動かす
115
+ - [デモ Web サービス](doc/demo.md) — `web/` の動かしかたと Cloudflare へのデプロイ
116
+ - [利用にあたっての注意](doc/legal.md)
117
+ - [RubyGems への公開手順](doc/publishing.md)
118
+
119
+ ## 開発
120
+
121
+ ```console
122
+ $ bundle install
123
+ $ bundle exec rspec # gem 151 examples
124
+ $ bundle exec rubocop
125
+ $ (cd web && bundle exec rspec) # デモアプリ 27 examples
126
+ $ (cd worker && npm test && npm run typecheck) # Worker 15 tests
127
+ ```
128
+
129
+ GitHub Codespaces なら `.devcontainer/` が上記をすべて用意します
130
+ ([doc/codespaces.md](doc/codespaces.md))。
131
+
132
+ ## 由来
133
+
134
+ 技術書典5 で頒布した『てっくやみなべ Vol.1』所収
135
+ 「快適なスライド閲覧⽣活を実現する Web サービスの開発」のサポートリポジトリを、
136
+ 現在も動く形で作り直したものです。当時の Lambda / PaaS 向けデプロイコードは、
137
+ 対象サービスが終了したため Cloudflare 向けに置き換えています。
138
+
139
+ 名前の gento は、スライドを映写して見せる「幻灯(げんとう)」から取っています。
140
+
141
+ ## ライセンス
142
+
143
+ MIT License. [LICENSE.txt](LICENSE.txt) を参照してください。
data/exe/gento ADDED
@@ -0,0 +1,7 @@
1
+ #!/usr/bin/env ruby
2
+ # frozen_string_literal: true
3
+
4
+ require "gento"
5
+ require "gento/cli"
6
+
7
+ exit Gento::CLI.start(ARGV)
@@ -0,0 +1,121 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "uri"
4
+
5
+ module Gento
6
+ module Adapters
7
+ # Shared behaviour for the per-site adapters.
8
+ #
9
+ # An adapter answers two questions: "is this URL mine?" and "what are the
10
+ # pages of this deck?". Everything else — HTTP, value objects, JSON output
11
+ # — is the framework's job.
12
+ class Base
13
+ class << self
14
+ # Host suffixes this adapter claims, e.g. %w[speakerdeck.com].
15
+ def hosts
16
+ raise NotImplementedError, "#{self} must declare .hosts"
17
+ end
18
+
19
+ def provider
20
+ name.split("::").last
21
+ .gsub(/([a-z\d])([A-Z])/, '\1_\2')
22
+ .downcase
23
+ end
24
+
25
+ def handles?(url)
26
+ uri = coerce_uri(url)
27
+ return false unless uri&.host
28
+
29
+ host = uri.host.downcase.sub(/\Awww\./, "")
30
+ hosts.any? { |claimed| host == claimed || host.end_with?(".#{claimed}") }
31
+ end
32
+
33
+ def coerce_uri(url)
34
+ uri = url.is_a?(URI) ? url : URI.parse(url.to_s)
35
+ uri.is_a?(URI::HTTP) ? uri : nil
36
+ rescue URI::InvalidURIError
37
+ nil
38
+ end
39
+ end
40
+
41
+ attr_reader :fetcher
42
+
43
+ def initialize(fetcher:)
44
+ @fetcher = fetcher
45
+ end
46
+
47
+ # @return [Gento::Deck]
48
+ def fetch(url)
49
+ raise NotImplementedError, "#{self.class} must implement #fetch"
50
+ end
51
+
52
+ private
53
+
54
+ def provider
55
+ self.class.provider
56
+ end
57
+
58
+ def get(url, headers: {})
59
+ fetcher.get(url, headers: headers)
60
+ end
61
+
62
+ def absolute_url(candidate, base)
63
+ return nil if candidate.nil? || candidate.to_s.empty?
64
+
65
+ URI.join(base.to_s, candidate.to_s).to_s
66
+ rescue URI::Error
67
+ nil
68
+ end
69
+
70
+ # Title/author/description as advertised by OpenGraph, which every one of
71
+ # these sites emits and keeps stable for the sake of link previews.
72
+ def open_graph_metadata(html)
73
+ {
74
+ title: html.meta("og:title", "twitter:title") || html.title,
75
+ description: html.meta("og:description", "description"),
76
+ author: html.meta("og:author", "author", "article:author")
77
+ }
78
+ end
79
+
80
+ # The oEmbed endpoint a page advertises in its head, if any.
81
+ #
82
+ # Every site here publishes one, so discovering it beats hardcoding a
83
+ # URL per adapter: the site tells us where its own endpoint lives.
84
+ def oembed_endpoint(html)
85
+ html.tags("link")
86
+ .find { |attrs| attrs["type"].to_s.include?("json+oembed") }
87
+ &.fetch("href", nil)
88
+ end
89
+
90
+ # oEmbed metadata, or an empty hash.
91
+ #
92
+ # It is always a nice-to-have: cleaner titles and real author names than
93
+ # OpenGraph offers, but never something a fetch should fail over.
94
+ def fetch_oembed(html)
95
+ endpoint = oembed_endpoint(html)
96
+ return {} unless endpoint
97
+
98
+ get(endpoint, headers: { "accept" => "application/json" }).json
99
+ rescue FetchError, ExtractionError
100
+ {}
101
+ end
102
+
103
+ def build_slides(urls, width: nil, height: nil)
104
+ urls.each_with_index.map do |url, index|
105
+ Slide.new(number: index + 1, url: url, width: width, height: height)
106
+ end
107
+ end
108
+
109
+ # The same image is often referenced at several sizes through a query
110
+ # string (`?width=160`, a cache-busting timestamp). Dropping the query
111
+ # before deduplicating collapses those into one page.
112
+ def without_query(url)
113
+ url.to_s.split("?").first.to_s
114
+ end
115
+
116
+ def fail_extraction(message)
117
+ raise ExtractionError, "#{provider}: #{message}"
118
+ end
119
+ end
120
+ end
121
+ end
@@ -0,0 +1,58 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Gento
4
+ module Adapters
5
+ # Docswell (https://www.docswell.com/s/<user>/<id>-<slug>).
6
+ #
7
+ # The deck page only renders the first handful of pages; the rest arrive
8
+ # by script. The embed view it links to carries the whole deck, so that is
9
+ # what this adapter reads. Each page appears there twice, full size and as
10
+ # a `?width=160` thumbnail, which dropping the query collapses.
11
+ class Docswell < Base
12
+ EMBED_URL = %r{\Ahttps://(?:www\.)?docswell\.com/slide/[A-Za-z0-9]+/embed}i
13
+ PAGE_IMAGE = %r{\Ahttps://[a-z0-9.-]*docswell\.com/page/[^/]+\.(?:jpe?g|png|webp)}i
14
+
15
+ def self.hosts
16
+ %w[docswell.com]
17
+ end
18
+
19
+ def fetch(url)
20
+ page = get(url).html
21
+
22
+ # The page links the embed both plainly and with a ?mode= variant.
23
+ embed_url = page.urls(EMBED_URL).map { |found| without_query(found) }.first
24
+ fail_extraction("could not find the embed view for #{url}") unless embed_url
25
+
26
+ images = slide_images(embed_url)
27
+ fail_extraction("found no slide images for #{url}") if images.empty?
28
+
29
+ build_deck(url, page, images)
30
+ end
31
+
32
+ private
33
+
34
+ # Page filenames are opaque ids, so document order is the only ordering
35
+ # available. The embed view lists the pages in slide order, which the
36
+ # deck's og:image corroborates: it is always the first page.
37
+ def slide_images(embed_url)
38
+ get(embed_url).html.urls(PAGE_IMAGE).map { |url| without_query(url) }.uniq
39
+ end
40
+
41
+ # The deck page carries no author metadata at all, so oEmbed is the only
42
+ # place a name comes from here.
43
+ def build_deck(url, page, images)
44
+ metadata = open_graph_metadata(page)
45
+ oembed = fetch_oembed(page)
46
+
47
+ Deck.new(
48
+ provider: provider,
49
+ source_url: url,
50
+ title: oembed["title"] || metadata[:title],
51
+ author: oembed["author_name"] || metadata[:author],
52
+ description: metadata[:description],
53
+ slides: build_slides(images)
54
+ )
55
+ end
56
+ end
57
+ end
58
+ end
@@ -0,0 +1,90 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Gento
4
+ module Adapters
5
+ # Google Slides, for decks shared publicly or published to the web.
6
+ #
7
+ # Google does not render page images into the deck page, so the job here
8
+ # is to find the ordered list of pages and name an image for each.
9
+ #
10
+ # The viewer at /embed is the obvious place to look and the wrong one: its
11
+ # payload mixes page ids with the ids of elements drawn on those pages,
12
+ # with nothing in the id itself to tell them apart. /htmlpresent instead
13
+ # references each page exactly once, in order, so that is what we read.
14
+ class GoogleSlides < Base
15
+ BASE_URL = "https://docs.google.com/presentation/d/%s"
16
+
17
+ # /presentation/d/<id>/... and the published /presentation/d/e/<id>/...
18
+ # Published ids contain hyphens, ordinary ones do not.
19
+ PRESENTATION_ID = %r{/presentation/d/(e/)?([A-Za-z0-9_-]{16,})}
20
+ PAGE_ID = /[?&]pageid=([A-Za-z0-9_]+)/
21
+ VIEWPAGE = %r{/viewpage\?}
22
+
23
+ def self.hosts
24
+ %w[docs.google.com]
25
+ end
26
+
27
+ def fetch(url)
28
+ published, id = presentation_id(url)
29
+ fail_extraction("could not find a presentation id in #{url}") unless id
30
+
31
+ page = get("#{deck_base(published, id)}/htmlpresent").html
32
+ images = published ? viewpage_images(page, url) : export_images(page, url, published, id)
33
+
34
+ Deck.new(
35
+ provider: provider,
36
+ source_url: url,
37
+ title: title_from(page),
38
+ slides: build_slides(images)
39
+ )
40
+ end
41
+
42
+ private
43
+
44
+ def deck_base(published, id)
45
+ format(BASE_URL, published ? "e/#{id}" : id)
46
+ end
47
+
48
+ # Ordinary decks: build export URLs from the page ids. These are stable,
49
+ # so a caller can keep them, and they redirect to a signed image.
50
+ def export_images(page, url, published, id)
51
+ base = deck_base(published, id)
52
+
53
+ page_ids(page, url).map { |page_id| "#{base}/export/png?id=#{id}&pageid=#{page_id}" }
54
+ end
55
+
56
+ # Published decks: /export/png answers 404 for them, whatever page id it
57
+ # is given. The only thing that serves their pages is the viewpage URL
58
+ # embedded in the document, which carries its own signature — so take it
59
+ # verbatim rather than trying to construct one.
60
+ #
61
+ # Those signatures expire, so unlike the export URLs these are worth
62
+ # fetching promptly rather than storing.
63
+ def viewpage_images(page, url)
64
+ images = page.urls(VIEWPAGE)
65
+ fail_extraction("found no slide pages for #{url}; is the deck still published?") if images.empty?
66
+
67
+ images
68
+ end
69
+
70
+ # Each page is referenced exactly once here, in order.
71
+ def page_ids(page, url)
72
+ ids = page.source.scan(PAGE_ID).flatten.uniq
73
+ fail_extraction("found no slide pages for #{url}; is the deck shared publicly?") if ids.empty?
74
+
75
+ ids
76
+ end
77
+
78
+ def presentation_id(url)
79
+ match = url.to_s.match(PRESENTATION_ID)
80
+ return [false, nil] unless match
81
+
82
+ [!match[1].nil?, match[2]]
83
+ end
84
+
85
+ def title_from(page)
86
+ page.title&.sub(/\s*-\s*Google (?:Slides|スライド)\s*\z/, "")
87
+ end
88
+ end
89
+ end
90
+ end
@@ -0,0 +1,82 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Gento
4
+ module Adapters
5
+ # SlideShare (https://www.slideshare.net/<user>/<slug>).
6
+ #
7
+ # SlideShare usually serves the deck to an ordinary HTTP client, but it
8
+ # sits behind bot protection that intermittently answers with a small
9
+ # JavaScript interstitial instead — most readily under bursty access.
10
+ #
11
+ # That interstitial has no slides in it, so without recognising it the
12
+ # adapter would report an empty deck and leave the caller guessing. Naming
13
+ # it costs one check and turns a mystery into something actionable: wait,
14
+ # or supply a JavaScript-capable Fetcher.
15
+ class SlideShare < Base
16
+ CHALLENGE = /Client Challenge|_fs-ch-/
17
+ # <slug>-<page>-<width>.jpg on the slide CDN.
18
+ SLIDE_IMAGE = %r{\Ahttps://[a-z0-9.-]*slidesharecdn\.com/\S+-(\d+)-(\d+)\.(?:jpe?g|png|webp)}i
19
+
20
+ def self.hosts
21
+ %w[slideshare.net]
22
+ end
23
+
24
+ def fetch(url)
25
+ page = get(url).html
26
+ fail_challenge(url) if challenge?(page)
27
+
28
+ images = slide_images(page)
29
+ fail_extraction("found no slide images at #{url}") if images.empty?
30
+
31
+ build_deck(url, page, images)
32
+ end
33
+
34
+ private
35
+
36
+ def challenge?(page)
37
+ page.title.to_s.match?(CHALLENGE) || page.source.match?(CHALLENGE)
38
+ end
39
+
40
+ def fail_challenge(url)
41
+ fail_extraction(
42
+ "#{url} returned SlideShare's JavaScript bot challenge instead of the deck. " \
43
+ "Fetching from SlideShare needs a JavaScript-capable Gento::Fetcher; " \
44
+ "the default net/http one cannot get past it."
45
+ )
46
+ end
47
+
48
+ # Pages are published at several widths, not always the same set for
49
+ # every page, so keep the widest copy of each.
50
+ def slide_images(page)
51
+ page.urls(SLIDE_IMAGE)
52
+ .group_by { |image| Integer(image[SLIDE_IMAGE, 1]) }
53
+ .sort_by(&:first)
54
+ .map { |_, group| group.max_by { |image| Integer(image[SLIDE_IMAGE, 2]) } }
55
+ end
56
+
57
+ def build_deck(url, page, images)
58
+ metadata = open_graph_metadata(page)
59
+ document = json_ld_document(page)
60
+
61
+ Deck.new(
62
+ provider: provider,
63
+ source_url: url,
64
+ title: metadata[:title],
65
+ author: author_from(document) || metadata[:author],
66
+ description: metadata[:description],
67
+ published_at: document["datePublished"],
68
+ slides: build_slides(images)
69
+ )
70
+ end
71
+
72
+ def json_ld_document(page)
73
+ page.json_ld.find { |node| node["@type"].to_s.match?(/Presentation|Article|CreativeWork/i) } || {}
74
+ end
75
+
76
+ def author_from(document)
77
+ author = document["author"]
78
+ author.is_a?(Hash) ? author["name"] : author
79
+ end
80
+ end
81
+ end
82
+ end
@@ -0,0 +1,79 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Gento
4
+ module Adapters
5
+ # Speaker Deck (https://speakerdeck.com/<user>/<slug>).
6
+ #
7
+ # The deck page carries every page image at full size, in <a href>
8
+ # elements. It also carries cover images for roughly thirty recommended
9
+ # decks, so matching on the CDN hostname alone would silently splice other
10
+ # people's slides into the result. Keying on this deck's own presentation
11
+ # id is what keeps them out.
12
+ class SpeakerDeck < Base
13
+ PLAYER_ID = %r{/player/([0-9a-f-]+)}i
14
+ DECK_ID = /\A[0-9a-f-]{8,}\z/i
15
+
16
+ def self.hosts
17
+ %w[speakerdeck.com]
18
+ end
19
+
20
+ def fetch(url)
21
+ page = get(url).html
22
+ oembed = fetch_oembed(page)
23
+
24
+ id = deck_id(page, oembed)
25
+ fail_extraction("could not find the presentation id for #{url}") unless id
26
+
27
+ images = slide_images(page, id)
28
+ fail_extraction("found no slide images at #{url}") if images.empty?
29
+
30
+ build_deck(url, page, oembed, images)
31
+ end
32
+
33
+ private
34
+
35
+ # Full-size pages only: the same presentation also publishes
36
+ # `preview_slide_N.jpg` thumbnails, which this pattern excludes.
37
+ def slide_images(page, id)
38
+ pattern = %r{
39
+ \Ahttps://files\.speakerdeck\.com/presentations/
40
+ #{Regexp.escape(id)}/slide_(\d+)\.[a-z]+
41
+ }xi
42
+
43
+ page.urls(pattern)
44
+ .map { |url| without_query(url) }
45
+ .uniq
46
+ .sort_by { |url| Integer(url[pattern, 1]) }
47
+ end
48
+
49
+ # oEmbed names the id inside a player URL; the page itself carries it
50
+ # bare, as the data-id of the embed placeholder.
51
+ def deck_id(page, oembed)
52
+ player_id(oembed["html"]) || embed_data_id(page) || player_id(page.source)
53
+ end
54
+
55
+ def player_id(text)
56
+ text.to_s[PLAYER_ID, 1]
57
+ end
58
+
59
+ def embed_data_id(page)
60
+ page.tags("div")
61
+ .filter_map { |attrs| attrs["data-id"] }
62
+ .find { |id| id.match?(DECK_ID) }
63
+ end
64
+
65
+ def build_deck(url, page, oembed, images)
66
+ metadata = open_graph_metadata(page)
67
+
68
+ Deck.new(
69
+ provider: provider,
70
+ source_url: url,
71
+ title: oembed["title"] || metadata[:title],
72
+ author: oembed["author_name"] || metadata[:author],
73
+ description: metadata[:description],
74
+ slides: build_slides(images)
75
+ )
76
+ end
77
+ end
78
+ end
79
+ end