jekyll-fingerprint-flow 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: 99741c6430da05235dd51944c3e977522e1b0e9c61a4be73c77598ebf9e15007
4
+ data.tar.gz: ea62ad37859d5412030b1498435f32131e6a694f0f4eb478860f06150f0a9928
5
+ SHA512:
6
+ metadata.gz: d1fe48c5fec220446853fe54aebbdabea227ae975043b8ce6ac228513d550c695dd2258b569e8492bf349e12ea208f6e6aecb1058af93bb6a9471bf9773fd35b
7
+ data.tar.gz: 370926dbd79d93edabe1a9e776526dda449a98ea0c83533ec322926e04ebe45670f6a3e1fea53ec1bea6ed65eb42c586f7262ed937e3f8f99c494003c21eda9d
data/CHANGELOG.md ADDED
@@ -0,0 +1,33 @@
1
+ # Changelog
2
+
3
+ ## [Unreleased]
4
+
5
+ ## [0.1.0] - 2026-10-06
6
+
7
+ - `Jekyll::FingerprintFlow::Interface.to_h` — the config surface
8
+ (`fingerprint_flow` keys plus the `priority:` range) derived from
9
+ `DEFAULTS` and the code's actual `site.config` reads, serialized to a
10
+ committed `interface.yml` by `rake interface`;
11
+ `spec/interface_spec.rb` pins manifest freshness and completeness.
12
+
13
+ - Per-site `:site, :post_write` dispatcher that content-tags supported local
14
+ asset URLs after normal-priority processing and before compression (default
15
+ priority 12; configurable from 11 through 19).
16
+ - Source-preserving HTML tag/attribute scanning for `.html` and `.htm` output;
17
+ rewrites exact `src`, `href`, `xlink:href`, `srcset`, and `poster` attributes,
18
+ not comments, script/style text, or unrelated `data-*` attributes.
19
+ - `srcset` parsing preserves data URLs, candidate descriptors, and unchanged
20
+ formatting; handles local candidates individually.
21
+ - `?v=<md5(content)[0,10]>` tags are derived from bytes and memoized only for
22
+ the current build; no persistent stat-keyed cache can return an old digest.
23
+ - Resolves local `<base href>` values, honors `baseurl`, rejects targets
24
+ outside `_site` (including traversal and symlinks), and skips symlinked HTML
25
+ outputs so builds cannot write through them.
26
+ - Skips external/scheme URLs, non-allowlisted paths, missing targets, and
27
+ pages with invalid base-href encoding; replaces existing `?v=` params and
28
+ preserves other query params and fragments.
29
+ - Config: `fingerprint_flow.enabled` (default true),
30
+ `fingerprint_flow.priority` (default 12, integer 11–19),
31
+ `fingerprint_flow.extensions` (extend the allowlist), and
32
+ `fingerprint_flow.exclude` (URL path prefixes never rewritten).
33
+ - Writes each HTML file only when an attribute URL actually changes.
data/LICENSE.txt ADDED
@@ -0,0 +1,22 @@
1
+ GNU AFFERO GENERAL PUBLIC LICENSE
2
+ Version 3, 19 November 2007
3
+
4
+ Copyright (C) 2026 Svend Gundestrup
5
+
6
+ This project is free software: you can redistribute it and/or modify it under
7
+ the terms of the GNU Affero General Public License as published by the Free
8
+ Software Foundation, either version 3 of the License, or (at your option) any
9
+ later version.
10
+
11
+ This project is distributed in the hope that it will be useful, but WITHOUT
12
+ ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS
13
+ FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more
14
+ details.
15
+
16
+ You should have received a copy of the GNU Affero General Public License along
17
+ with this project. If not, see <https://www.gnu.org/licenses/agpl-3.0.html>.
18
+
19
+ The complete license text is available from the Free Software Foundation at:
20
+ https://www.gnu.org/licenses/agpl-3.0.txt
21
+
22
+ SPDX-License-Identifier: AGPL-3.0-or-later
data/README.md ADDED
@@ -0,0 +1,159 @@
1
+ # jekyll-fingerprint-flow
2
+
3
+ [![Status: Active](https://img.shields.io/badge/status-active-success)](https://github.com/gundestrup/jekyll-fingerprint-flow)
4
+ [![CI](https://github.com/gundestrup/jekyll-fingerprint-flow/actions/workflows/ci.yml/badge.svg)](https://github.com/gundestrup/jekyll-fingerprint-flow/actions/workflows/ci.yml)
5
+ [![Gem Version](https://img.shields.io/gem/v/jekyll-fingerprint-flow)](https://rubygems.org/gems/jekyll-fingerprint-flow)
6
+ [![Codecov](https://codecov.io/gh/gundestrup/jekyll-fingerprint-flow/graph/badge.svg)](https://codecov.io/gh/gundestrup/jekyll-fingerprint-flow)
7
+ [![Ruby](https://img.shields.io/badge/ruby-%E2%89%A5%203.3-red.svg)](https://www.ruby-lang.org/)
8
+ [![Jekyll](https://img.shields.io/badge/jekyll-4.x-blue.svg)](https://jekyllrb.com/)
9
+ [![Release](https://img.shields.io/github/v/tag/gundestrup/jekyll-fingerprint-flow)](https://github.com/gundestrup/jekyll-fingerprint-flow/tags)
10
+ [![License: AGPL v3](https://img.shields.io/badge/license-AGPL--3.0--or--later-blue.svg)](LICENSE.txt)
11
+ [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/gundestrup/jekyll-fingerprint-flow)
12
+ [![CodeFactor](https://www.codefactor.io/repository/github/gundestrup/jekyll-fingerprint-flow/badge)](https://www.codefactor.io/repository/github/gundestrup/jekyll-fingerprint-flow)
13
+ [![Semgrep CE](https://img.shields.io/badge/Semgrep_CE-security-success)](https://github.com/gundestrup/jekyll-fingerprint-flow/security/code-scanning)
14
+ [![Quality Gate Status](https://sonarcloud.io/api/project_badges/measure?project=gundestrup_jekyll-fingerprint-flow&metric=alert_status)](https://sonarcloud.io/summary/new_code?id=gundestrup_jekyll-fingerprint-flow)
15
+
16
+ Drop-in content-hash cache busting for Jekyll. After the site is written,
17
+ matching local asset URLs in supported attributes of generated `.html`
18
+ and `.htm` files receive a content-derived `?v=<md5>` tag:
19
+
20
+ ```html
21
+ <link rel="stylesheet" href="/assets/site.css">
22
+ <!-- becomes -->
23
+ <link rel="stylesheet" href="/assets/site.css?v=3f8a1c2e94">
24
+ ```
25
+
26
+ No template changes or asset-pipeline tags are needed. The post-write pass
27
+ covers matching URLs emitted by layouts, themes, and other plugins, without
28
+ mistaking comments, script text, or unrelated `data-*` attributes for markup.
29
+
30
+ ## Why
31
+
32
+ Browsers cache static assets aggressively. When a site serves assets with a
33
+ long `Cache-Control: max-age`, a new deploy can update the HTML while visitors
34
+ keep stale CSS or JavaScript referenced directly by that HTML. Changing the
35
+ URL invalidates the browser cache entry, and a **content hash** changes the
36
+ URL exactly when the referenced file's bytes change:
37
+
38
+ - same bytes → same URL → browser cache entry remains valid
39
+ - different bytes → new URL → browser fetches the new resource
40
+
41
+ The digest is deterministic across machines and builds; there is no persistent
42
+ version counter or manifest state to synchronize.
43
+
44
+ Query-param tagging (`?v=`) was chosen over renamed files
45
+ (`site-<hash>.css`) deliberately: filenames stay stable, so precompressed
46
+ siblings (`.br`, `.zst`, `.gz`) and direct document links keep working. The
47
+ post-write order is normal processing (20), fingerprinting (12), then
48
+ compression (10), so compressors see the rewritten HTML.
49
+
50
+ ## Install
51
+
52
+ ```ruby
53
+ # Gemfile
54
+ group :jekyll_plugins do
55
+ gem "jekyll-fingerprint-flow"
56
+ end
57
+ ```
58
+
59
+ That's it — the next `jekyll build` fingerprints matching references.
60
+
61
+ ## What gets tagged
62
+
63
+ The plugin scans real start-tag attributes in generated `.html` and `.htm`
64
+ files. It edits only the attribute value spans; it does not serialize the
65
+ whole document.
66
+
67
+ | Attribute | Example |
68
+ | ----------- | --------- |
69
+ | `src` | `<script>`, `<img>`, `<video>`, `<source>`, `<iframe>` |
70
+ | `href` | `<link>` and direct links to downloadable files |
71
+ | `xlink:href` | legacy SVG links |
72
+ | `srcset` | responsive image candidates (descriptors preserved) |
73
+ | `poster` | `<video>` poster frames |
74
+
75
+ A URL is tagged only when **all** of these hold:
76
+
77
+ - it is local (no scheme, no `//`, not `data:`/`mailto:`/…)
78
+ - its extension is in the allowlist (`css js mjs map json xml webmanifest
79
+ png jpg jpeg gif webp avif svg ico woff woff2 ttf otf eot mp4 webm mp3
80
+ wav ogg pdf txt`)
81
+ - its target resolves to a real file within `_site` (root-relative paths honor
82
+ `baseurl`; relative paths honor a local `<base href>` when present, otherwise
83
+ resolve from the generated document's directory)
84
+
85
+ Comments, raw-text elements (`script`, `style`, `textarea`, and similar), and
86
+ non-target attributes such as `data-src` are not scanned. External `<base`
87
+ URLs are left untouched along with that page's local-looking references,
88
+ since the browser would resolve them against the external origin.
89
+
90
+ `srcset` candidates are parsed separately so commas inside data URLs are not
91
+ treated as candidate separators. Candidates that cannot be tagged stay byte-
92
+ for-byte as supplied; the attribute is rewritten only if at least one URL
93
+ changes.
94
+
95
+ Existing query strings are preserved — a `?v=` already present is replaced
96
+ (not stacked), other parameters are retained, and `#fragments` survive. HTML
97
+ entities in attribute values are decoded for URL processing and escaped again
98
+ when written.
99
+
100
+ ### Scope boundary
101
+
102
+ This is a post-write HTML attribute rewriter, not a general parser for every
103
+ resource-bearing language. It does **not** rewrite CSS `url()` / `@import`,
104
+ inline CSS, JavaScript `fetch()`/dynamic imports, references in JSON/XML, or
105
+ other non-HTML output. Such references need format-specific processing; the
106
+ plugin intentionally does not guess by rewriting arbitrary strings in source
107
+ text.
108
+
109
+ ## Config
110
+
111
+ All optional — defaults are zero-config:
112
+
113
+ ```yaml
114
+ # _config.yml
115
+ fingerprint_flow:
116
+ enabled: true # set false to disable the pass entirely
117
+ priority: 12 # integer 11..19; default 12
118
+ extensions: [.gpx, .bin] # extra extensions to fingerprint
119
+ exclude: # URL path prefixes never rewritten
120
+ - "/internal/"
121
+ ```
122
+
123
+ `priority` is a site-configurable integer from 11 through 19 (default 12). A
124
+ dispatcher registers at each allowed slot and only the one selected by this
125
+ site's config rewrites output. Normal-priority work (20) runs first, then
126
+ fingerprinting, then compression (0–10).
127
+
128
+ ## Behaviour notes
129
+
130
+ - **Content-derived tag.** The value is `md5(file content)[0,10]`, a pure
131
+ function of the output bytes.
132
+ - **Build-local memoization.** Each unique resolved file is hashed once per
133
+ build. The memo is not persisted, so preserved mtimes or same-size edits
134
+ cannot make a later build reuse a stale digest.
135
+ - **Destination confinement.** Resolved files and symlink targets must remain
136
+ under `_site`; path traversal and external symlinks are skipped. Symlinked
137
+ HTML pages are not rewritten.
138
+ - **No mtime churn.** HTML files are written back only when a URL changes, so
139
+ rsync-based deploys do not re-upload untouched pages.
140
+ - **HTTP policy is separate.** The plugin changes URLs; it does not set
141
+ `Cache-Control`. The web server still controls freshness. Long-lived browser
142
+ caching is safe when caches key entries by the full URL, including `?v=`.
143
+
144
+ ## Development
145
+
146
+ ```bash
147
+ bundle install
148
+ bundle exec rake quick # rubocop + rspec
149
+ bundle exec rake ci # + bundler-audit, semgrep, package check
150
+ ```
151
+
152
+ Release: `bundle exec rake "version:bump[patch]"`, add a dated CHANGELOG
153
+ entry, commit, `git tag vX.Y.Z && git push --tags` — the release workflow
154
+ verifies the tag, builds the gem, attaches it to a GitHub release, and
155
+ publishes to RubyGems via OIDC trusted publishing.
156
+
157
+ ## License
158
+
159
+ AGPL-3.0-or-later — see [LICENSE.txt](LICENSE.txt).
data/interface.yml ADDED
@@ -0,0 +1,14 @@
1
+ # Generated by `rake interface` — do not edit.
2
+ # Public tag/config surface for tooling (editor extensions, doc linters).
3
+ ---
4
+ gem: jekyll-fingerprint-flow
5
+ version: 0.1.0
6
+ tags: {}
7
+ filters: []
8
+ config:
9
+ fingerprint_flow:
10
+ - enabled
11
+ - exclude
12
+ - extensions
13
+ - priority
14
+ enums: {}
@@ -0,0 +1,55 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Jekyll
4
+ module FingerprintFlow
5
+ # Reads the `fingerprint_flow:` section of _config.yml.
6
+ #
7
+ # fingerprint_flow:
8
+ # enabled: true # default
9
+ # extensions: [foo] # extra extensions to fingerprint
10
+ # exclude: ["/internal/"] # URL path prefixes never rewritten
11
+ # priority: 12 # integer 11..19; between normal processing and compression
12
+ class Configuration
13
+ DEFAULT_POST_WRITE_PRIORITY = 12
14
+ POST_WRITE_PRIORITIES = (11..19)
15
+ DEFAULT_EXTENSIONS = %w[
16
+ css js mjs map json xml webmanifest
17
+ png jpg jpeg gif webp avif svg ico
18
+ woff woff2 ttf otf eot
19
+ mp4 webm mp3 wav ogg pdf txt
20
+ ].freeze
21
+
22
+ attr_reader :extensions, :exclude, :post_write_priority
23
+
24
+ def self.post_write_priority(site)
25
+ config = site.config["fingerprint_flow"] || {}
26
+ priority = config.fetch("priority", DEFAULT_POST_WRITE_PRIORITY)
27
+ return priority if priority.is_a?(Integer) && POST_WRITE_PRIORITIES.cover?(priority)
28
+
29
+ raise Jekyll::Errors::FatalException,
30
+ "FingerprintFlow: fingerprint_flow.priority must be an integer from 11 through 19"
31
+ end
32
+
33
+ def initialize(site)
34
+ cfg = site.config["fingerprint_flow"] || {}
35
+ @enabled = cfg.fetch("enabled", true)
36
+ @post_write_priority = self.class.post_write_priority(site)
37
+ configure_asset_options(cfg)
38
+ end
39
+
40
+ def enabled?
41
+ @enabled
42
+ end
43
+
44
+ private
45
+
46
+ def configure_asset_options(config)
47
+ extra = Array(config["extensions"]).map do |extension|
48
+ extension.to_s.downcase.delete_prefix(".")
49
+ end
50
+ @extensions = (DEFAULT_EXTENSIONS | extra).freeze
51
+ @exclude = Array(config["exclude"]).map(&:to_s)
52
+ end
53
+ end
54
+ end
55
+ end
@@ -0,0 +1,22 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "digest"
4
+
5
+ module Jekyll
6
+ module FingerprintFlow
7
+ # Content-hash tags for files under _site. The tag is a pure function of
8
+ # file content — same bytes, same tag, on any machine, in any
9
+ # environment. Each output file is hashed once per build.
10
+ class Hasher
11
+ DIGEST_LENGTH = 10
12
+
13
+ def initialize
14
+ @memo = {}
15
+ end
16
+
17
+ def tag(path)
18
+ @memo[path] ||= Digest::MD5.file(path).hexdigest[0, DIGEST_LENGTH]
19
+ end
20
+ end
21
+ end
22
+ end
@@ -0,0 +1,115 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "cgi"
4
+
5
+ module Jekyll
6
+ module FingerprintFlow
7
+ class HtmlDocumentRewriter
8
+ URL_ATTRIBUTES = %w[src href srcset poster xlink:href].freeze
9
+ INVALID_BASE_HREF = :invalid_base_href
10
+
11
+ attr_reader :tagged, :skipped
12
+
13
+ def initialize(tagger, scanner: HtmlTagScanner.new)
14
+ @tagger = tagger
15
+ @scanner = scanner
16
+ @tagged = 0
17
+ @skipped = 0
18
+ @srcset_rewriter = SrcsetRewriter.new(@tagger) { |value| counted(value) }
19
+ end
20
+
21
+ def rewrite(source, path)
22
+ base_href = first_base_href(source)
23
+ return invalid_base(path) if base_href == INVALID_BASE_HREF
24
+
25
+ apply_changes(source, attribute_changes(source, path, base_href))
26
+ end
27
+
28
+ private
29
+
30
+ def first_base_href(source)
31
+ @scanner.each_tag(source) do |name, attributes|
32
+ next unless name == "base"
33
+
34
+ attribute = attributes.find { |item| item.name == "href" }
35
+ next unless attribute
36
+
37
+ value = decode_attribute(attribute_value(source, attribute))
38
+ return value || INVALID_BASE_HREF
39
+ end
40
+ nil
41
+ end
42
+
43
+ def attribute_changes(source, path, base_href)
44
+ changes = []
45
+ @scanner.each_tag(source) do |_name, attributes|
46
+ changes.concat(changes_for_attributes(source, attributes, File.dirname(path), base_href))
47
+ end
48
+ changes
49
+ end
50
+
51
+ def changes_for_attributes(source, attributes, html_dir, base_href)
52
+ attributes.filter_map do |attribute|
53
+ change_for_attribute(source, attribute, html_dir, base_href)
54
+ end
55
+ end
56
+
57
+ def change_for_attribute(source, attribute, html_dir, base_href)
58
+ return unless URL_ATTRIBUTES.include?(attribute.name)
59
+
60
+ value = decode_attribute(attribute_value(source, attribute))
61
+ return count_skip unless value
62
+
63
+ rewritten = rewrite_attribute(attribute, value, html_dir, base_href)
64
+ return unless rewritten && rewritten != value
65
+
66
+ [attribute.value_start, attribute.value_end, CGI.escapeHTML(rewritten).b]
67
+ end
68
+
69
+ def rewrite_attribute(attribute, value, html_dir, base_href)
70
+ if attribute.name == "srcset"
71
+ return @srcset_rewriter.rewrite(value, html_dir: html_dir, base_href: base_href)
72
+ end
73
+
74
+ counted(@tagger.tag(value, html_dir: html_dir, base_href: base_href))
75
+ end
76
+
77
+ def attribute_value(source, attribute)
78
+ source.byteslice(attribute.value_start, attribute.value_end - attribute.value_start)
79
+ end
80
+
81
+ def decode_attribute(raw)
82
+ value = raw.dup.force_encoding(Encoding::UTF_8)
83
+ CGI.unescapeHTML(value) if value.valid_encoding?
84
+ end
85
+
86
+ def invalid_base(path)
87
+ Jekyll.logger.warn("fingerprint-flow:", "skipped #{path} (invalid base href encoding)")
88
+ nil
89
+ end
90
+
91
+ def apply_changes(source, changes)
92
+ changes.sort_by(&:first).reverse_each do |start, finish, replacement|
93
+ prefix = source.byteslice(0, start)
94
+ suffix = source.byteslice(finish, source.bytesize - finish)
95
+ source = prefix + replacement + suffix
96
+ end
97
+ source
98
+ end
99
+
100
+ def count_skip
101
+ @skipped += 1
102
+ nil
103
+ end
104
+
105
+ def counted(value)
106
+ if value
107
+ @tagged += 1
108
+ value
109
+ else
110
+ count_skip
111
+ end
112
+ end
113
+ end
114
+ end
115
+ end
@@ -0,0 +1,122 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Jekyll
4
+ module FingerprintFlow
5
+ class HtmlTagParser
6
+ Attribute = Struct.new(:name, :value_start, :value_end, keyword_init: true)
7
+ ATTRIBUTE_TERMINATORS = [34, 39, 47, 60, 61, 62].freeze
8
+ QUOTES = [34, 39].freeze
9
+ ASCII_SPACE = [9, 10, 12, 13, 32].freeze
10
+
11
+ def initialize(source, opening)
12
+ @source = source
13
+ @cursor = opening + 1
14
+ end
15
+
16
+ def parse
17
+ return unless ascii_alpha?(@source.getbyte(@cursor))
18
+
19
+ name = read_tag_name
20
+ attributes, ending = read_attributes
21
+ [name, attributes, ending]
22
+ end
23
+
24
+ private
25
+
26
+ def read_tag_name
27
+ start = @cursor
28
+ @cursor += 1 while tag_name_byte?(@source.getbyte(@cursor))
29
+ @source.byteslice(start, @cursor - start).downcase
30
+ end
31
+
32
+ def read_attributes
33
+ attributes = []
34
+ seen = {}
35
+ loop do
36
+ @cursor = skip_spaces(@cursor)
37
+ ending = tag_ending
38
+ return [attributes, ending] if ending
39
+
40
+ append_current_attribute(attributes, seen)
41
+ end
42
+ end
43
+
44
+ def tag_ending
45
+ byte = @source.getbyte(@cursor)
46
+ return @cursor + 1 if byte == 62
47
+ return @cursor + 2 if byte == 47 && @source.getbyte(@cursor + 1) == 62
48
+
49
+ @source.bytesize unless byte
50
+ end
51
+
52
+ def append_current_attribute(attributes, seen)
53
+ name = read_attribute_name
54
+ unless name
55
+ @cursor += 1
56
+ return
57
+ end
58
+
59
+ start, finish = read_attribute_value
60
+ return if seen[name]
61
+
62
+ attributes << Attribute.new(name: name, value_start: start, value_end: finish) if start
63
+ seen[name] = true
64
+ end
65
+
66
+ def read_attribute_name
67
+ start = @cursor
68
+ @cursor += 1 while attribute_name_byte?(@source.getbyte(@cursor))
69
+ return if @cursor == start
70
+
71
+ @source.byteslice(start, @cursor - start).downcase
72
+ end
73
+
74
+ def read_attribute_value
75
+ @cursor = skip_spaces(@cursor)
76
+ return [nil, nil] unless @source.getbyte(@cursor) == 61
77
+
78
+ @cursor = skip_spaces(@cursor + 1)
79
+ quote = @source.getbyte(@cursor)
80
+ return read_quoted_value(quote) if QUOTES.include?(quote)
81
+
82
+ read_unquoted_value
83
+ end
84
+
85
+ def read_quoted_value(quote)
86
+ start = @cursor + 1
87
+ ending = @source.index(quote.chr.b, start)
88
+ @cursor = ending ? ending + 1 : @source.bytesize
89
+ ending ? [start, ending] : [nil, nil]
90
+ end
91
+
92
+ def read_unquoted_value
93
+ start = @cursor
94
+ @cursor += 1 while (byte = @source.getbyte(@cursor)) && !space?(byte) && byte != 62
95
+ [start, @cursor]
96
+ end
97
+
98
+ def skip_spaces(position)
99
+ position += 1 while space?(@source.getbyte(position))
100
+ position
101
+ end
102
+
103
+ def space?(byte)
104
+ ASCII_SPACE.include?(byte)
105
+ end
106
+
107
+ def ascii_alpha?(byte)
108
+ return false unless byte
109
+
110
+ byte.between?(65, 90) || byte.between?(97, 122)
111
+ end
112
+
113
+ def tag_name_byte?(byte)
114
+ byte && !space?(byte) && ![47, 61, 62].include?(byte)
115
+ end
116
+
117
+ def attribute_name_byte?(byte)
118
+ byte && !space?(byte) && !ATTRIBUTE_TERMINATORS.include?(byte)
119
+ end
120
+ end
121
+ end
122
+ end
@@ -0,0 +1,108 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "jekyll/fingerprint_flow/html_tag_parser"
4
+
5
+ module Jekyll
6
+ module FingerprintFlow
7
+ class HtmlTagScanner
8
+ RAW_TEXT_ELEMENTS = %w[
9
+ script style textarea title xmp iframe noembed noframes plaintext
10
+ ].freeze
11
+ COMMENT_OPEN = "<!--".b.freeze
12
+ COMMENT_CLOSE = "-->".b.freeze
13
+ CDATA_OPEN = "<![CDATA[".b.freeze
14
+
15
+ def each_tag(source, &block)
16
+ return enum_for(:each_tag, source) unless block
17
+
18
+ data = source.b
19
+ cursor = 0
20
+ while (opening = data.index("<".b, cursor))
21
+ cursor = skip_special(data, opening) || scan_tag(data, opening, &block)
22
+ end
23
+ self
24
+ end
25
+
26
+ private
27
+
28
+ def scan_tag(source, opening, &block)
29
+ parsed = HtmlTagParser.new(source, opening).parse
30
+ return opening + 1 unless parsed
31
+
32
+ name, attributes, cursor = parsed
33
+ block.call(name, attributes)
34
+ after_tag(source, name, cursor)
35
+ end
36
+
37
+ def after_tag(source, name, cursor)
38
+ return source.bytesize if name == "plaintext"
39
+ return raw_text_end(source, name, cursor) if RAW_TEXT_ELEMENTS.include?(name)
40
+
41
+ cursor
42
+ end
43
+
44
+ def skip_special(source, opening)
45
+ if starts_with?(source, COMMENT_OPEN, opening)
46
+ skip_delimited(source, opening + COMMENT_OPEN.bytesize)
47
+ elsif starts_with?(source, CDATA_OPEN, opening)
48
+ skip_cdata(source, opening + CDATA_OPEN.bytesize)
49
+ elsif markup_start?(source, opening)
50
+ skip_markup(source, opening + 2)
51
+ end
52
+ end
53
+
54
+ def markup_start?(source, opening)
55
+ [33, 47, 63].include?(source.getbyte(opening + 1))
56
+ end
57
+
58
+ def starts_with?(source, token, position)
59
+ source.byteslice(position, token.bytesize) == token
60
+ end
61
+
62
+ def skip_delimited(source, position)
63
+ ending = source.index(COMMENT_CLOSE, position)
64
+ ending ? ending + COMMENT_CLOSE.bytesize : source.bytesize
65
+ end
66
+
67
+ def skip_cdata(source, position)
68
+ while position + 2 < source.bytesize
69
+ return position + 3 if cdata_end?(source, position)
70
+
71
+ position += 1
72
+ end
73
+ source.bytesize
74
+ end
75
+
76
+ def cdata_end?(source, position)
77
+ source.getbyte(position) == 93 &&
78
+ source.getbyte(position + 1) == 93 &&
79
+ source.getbyte(position + 2) == 62
80
+ end
81
+
82
+ def skip_markup(source, position)
83
+ quote = nil
84
+ while position < source.bytesize
85
+ byte = source.getbyte(position)
86
+ return position + 1 if byte == 62 && quote.nil?
87
+
88
+ quote = next_quote(byte, quote)
89
+ position += 1
90
+ end
91
+ position
92
+ end
93
+
94
+ def next_quote(byte, quote)
95
+ return nil if quote && byte == quote
96
+ return byte if quote.nil? && [34, 39].include?(byte)
97
+
98
+ quote
99
+ end
100
+
101
+ def raw_text_end(source, name, position)
102
+ pattern = Regexp.new("</#{Regexp.escape(name)}(?=[\\x09\\x0a\\x0c\\x0d\\x20/>])".b, Regexp::IGNORECASE)
103
+ match = pattern.match(source, position)
104
+ match ? match.begin(0) : source.bytesize
105
+ end
106
+ end
107
+ end
108
+ end
@@ -0,0 +1,24 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Jekyll
4
+ module FingerprintFlow
5
+ # Machine-readable description of the plugin's public interface:
6
+ # config keys (this plugin registers no Liquid tags or filters).
7
+ # `rake interface` writes this as interface.yml (shipped in the gem) so
8
+ # tooling like editor extensions can consume it without parsing Ruby.
9
+ module Interface
10
+ CONFIG_KEYS = %w[enabled exclude extensions priority].freeze
11
+
12
+ def self.to_h
13
+ {
14
+ "gem" => "jekyll-fingerprint-flow",
15
+ "version" => VERSION,
16
+ "tags" => {},
17
+ "filters" => [],
18
+ "config" => { "fingerprint_flow" => CONFIG_KEYS.sort },
19
+ "enums" => {}
20
+ }
21
+ end
22
+ end
23
+ end
24
+ end
@@ -0,0 +1,47 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Jekyll
4
+ module FingerprintFlow
5
+ class Rewriter
6
+ def initialize(site, hasher: Hasher.new)
7
+ @config = Configuration.new(site)
8
+ @dest = File.expand_path(site.dest)
9
+ tagger = UrlTagger.new(site: site, config: @config, hasher: hasher)
10
+ @html_rewriter = HtmlDocumentRewriter.new(tagger)
11
+ end
12
+
13
+ def run
14
+ return unless @config.enabled?
15
+
16
+ files = Dir.glob(File.join(@dest, "**", "*.{html,htm}"))
17
+ changed = files.count { |file| rewrite_file(file) }
18
+ summary = "#{files.size} HTML files scanned, #{changed} rewritten, " \
19
+ "#{@html_rewriter.tagged} refs tagged"
20
+ Jekyll.logger.info("fingerprint-flow:", summary)
21
+ end
22
+
23
+ def rewrite_file(path)
24
+ return unless safe_output_path?(path)
25
+
26
+ source = File.binread(path)
27
+ output = @html_rewriter.rewrite(source, path)
28
+ return unless output && output != source
29
+
30
+ File.binwrite(path, output)
31
+ end
32
+
33
+ private
34
+
35
+ def safe_output_path?(path)
36
+ return false if File.symlink?(path)
37
+
38
+ root = File.realpath(@dest)
39
+ real = File.realpath(path)
40
+ prefix = root.end_with?(File::SEPARATOR) ? root : "#{root}#{File::SEPARATOR}"
41
+ File.file?(real) && real.start_with?(prefix)
42
+ rescue SystemCallError
43
+ false
44
+ end
45
+ end
46
+ end
47
+ end
@@ -0,0 +1,96 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Jekyll
4
+ module FingerprintFlow
5
+ class SrcsetRewriter
6
+ ASCII_SPACE = [9, 10, 12, 13, 32].freeze
7
+
8
+ def initialize(tagger, &counter)
9
+ @tagger = tagger
10
+ @counter = counter
11
+ end
12
+
13
+ def rewrite(value, html_dir:, base_href:)
14
+ changes = []
15
+ position = 0
16
+ while (candidate = next_candidate(value, position))
17
+ start, finish, url, position = candidate
18
+ tagged = @counter.call(@tagger.tag(url, html_dir: html_dir, base_href: base_href))
19
+ changes << [start, finish, tagged.b] if tagged && tagged != url
20
+ end
21
+ return if changes.empty?
22
+
23
+ apply_changes(value.b, changes).force_encoding(Encoding::UTF_8)
24
+ end
25
+
26
+ private
27
+
28
+ def next_candidate(value, position)
29
+ position = skip_separators(value, position)
30
+ return if position >= value.bytesize
31
+
32
+ start = position
33
+ position = skip_url(value, position)
34
+ finish = position
35
+ url = value.byteslice(start, finish - start)
36
+ url, finish, trailing_comma = strip_trailing_commas(url, finish)
37
+ position = skip_descriptors(value, position) unless trailing_comma
38
+ [start, finish, url, position]
39
+ end
40
+
41
+ def strip_trailing_commas(url, finish)
42
+ commas = url[/,+\z/]
43
+ return [url, finish, false] unless commas
44
+
45
+ [url.byteslice(0, url.bytesize - commas.bytesize), finish - commas.bytesize, true]
46
+ end
47
+
48
+ def skip_url(value, position)
49
+ position += 1 while position < value.bytesize && !space?(value.getbyte(position))
50
+ position
51
+ end
52
+
53
+ def skip_descriptors(value, position)
54
+ position = skip_spaces(value, position)
55
+ depth = 0
56
+ while position < value.bytesize
57
+ byte = value.getbyte(position)
58
+ return position + 1 if byte == 44 && depth.zero?
59
+
60
+ depth += 1 if byte == 40
61
+ depth -= 1 if byte == 41 && depth.positive?
62
+ position += 1
63
+ end
64
+ position
65
+ end
66
+
67
+ def skip_separators(value, position)
68
+ while position < value.bytesize
69
+ byte = value.getbyte(position)
70
+ break unless space?(byte) || byte == 44
71
+
72
+ position += 1
73
+ end
74
+ position
75
+ end
76
+
77
+ def skip_spaces(value, position)
78
+ position += 1 while position < value.bytesize && space?(value.getbyte(position))
79
+ position
80
+ end
81
+
82
+ def space?(byte)
83
+ ASCII_SPACE.include?(byte)
84
+ end
85
+
86
+ def apply_changes(source, changes)
87
+ changes.sort_by(&:first).reverse_each do |start, finish, replacement|
88
+ prefix = source.byteslice(0, start)
89
+ suffix = source.byteslice(finish, source.bytesize - finish)
90
+ source = prefix + replacement + suffix
91
+ end
92
+ source
93
+ end
94
+ end
95
+ end
96
+ end
@@ -0,0 +1,144 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "uri"
4
+
5
+ module Jekyll
6
+ module FingerprintFlow
7
+ # Decides whether a single URL from generated HTML should carry a
8
+ # content-hash tag, and produces the tagged URL. Pure URL logic — file
9
+ # resolution and hashing are injected.
10
+ #
11
+ # Tagged: local URLs with a fingerprintable extension resolving to a
12
+ # real file under _site → ?v=<md5> (merged into any existing
13
+ # query, fragment preserved)
14
+ # Skipped: external URLs (any scheme, //), fragment-only, data:,
15
+ # non-allowlisted extensions, and references whose target file does
16
+ # not exist inside _site.
17
+ class UrlTagger
18
+ EXTERNAL = %r{\A(?:[a-z][a-z0-9+.-]*:|//)}i
19
+
20
+ def initialize(site:, config:, hasher:)
21
+ @dest = File.expand_path(site.dest)
22
+ @baseurl = normalize_baseurl(site.baseurl)
23
+ @config = config
24
+ @hasher = hasher
25
+ end
26
+
27
+ # Returns the rewritten URL, or nil when the URL must not be touched.
28
+ def tag(url, html_dir:, base_href: nil)
29
+ parts = split(url.strip) or return
30
+ path, query, frag = parts
31
+ decoded_path = unescape(path) or return
32
+ return unless fingerprintable?(decoded_path)
33
+
34
+ resolved = resolve(decoded_path, html_dir, base_href) or return
35
+ merge(path, query, frag, @hasher.tag(resolved))
36
+ end
37
+
38
+ private
39
+
40
+ def normalize_baseurl(value)
41
+ normalized = value.to_s
42
+ normalized = normalized.delete_prefix("/") while normalized.start_with?("/")
43
+ normalized = normalized.delete_suffix("/") while normalized.end_with?("/")
44
+ normalized
45
+ end
46
+
47
+ def split(url)
48
+ return if url.empty? || url.match?(EXTERNAL)
49
+
50
+ path, _, fragment = url.partition("#")
51
+ path, _, query = path.partition("?")
52
+ return if path.empty?
53
+
54
+ [path, query, fragment]
55
+ end
56
+
57
+ def unescape(path)
58
+ URI::DEFAULT_PARSER.unescape(path)
59
+ rescue ArgumentError
60
+ nil
61
+ end
62
+
63
+ def fingerprintable?(path)
64
+ ext = File.extname(path).delete_prefix(".").downcase
65
+ return false unless @config.extensions.include?(ext)
66
+
67
+ @config.exclude.none? { |prefix| path.start_with?(prefix) }
68
+ end
69
+
70
+ def resolve(path, html_dir, base_href)
71
+ return if external_base_href?(base_href)
72
+
73
+ candidate = candidate_path(path, html_dir, base_href)
74
+ candidate && confined_file(candidate)
75
+ rescue ArgumentError, SystemCallError
76
+ nil
77
+ end
78
+
79
+ def external_base_href?(base_href)
80
+ base_href&.strip&.match?(EXTERNAL)
81
+ end
82
+
83
+ def candidate_path(path, html_dir, base_href)
84
+ return root_relative_path(path) if path.start_with?("/")
85
+
86
+ base_dir = relative_base_dir(base_href, html_dir)
87
+ base_dir && File.expand_path(path, base_dir)
88
+ end
89
+
90
+ def root_relative_path(path)
91
+ File.expand_path(without_baseurl(path.delete_prefix("/")), @dest)
92
+ end
93
+
94
+ def without_baseurl(relative)
95
+ return relative if @baseurl.empty?
96
+ return "" if relative == @baseurl
97
+
98
+ relative.start_with?("#{@baseurl}/") ? relative.delete_prefix("#{@baseurl}/") : relative
99
+ end
100
+
101
+ def relative_base_dir(base_href, html_dir)
102
+ return html_dir unless base_href
103
+
104
+ base_path = base_href.strip.partition("#").first.partition("?").first
105
+ decoded_base = unescape(base_path) or return
106
+ return html_dir if decoded_base.empty?
107
+
108
+ base_target = base_target(decoded_base, html_dir)
109
+ decoded_base.end_with?("/") ? base_target : File.dirname(base_target)
110
+ end
111
+
112
+ def base_target(path, html_dir)
113
+ return root_relative_path(path) if path.start_with?("/")
114
+
115
+ File.expand_path(path, html_dir)
116
+ end
117
+
118
+ def confined_file(candidate)
119
+ root = File.realpath(@dest)
120
+ resolved = File.realpath(candidate)
121
+ return unless File.file?(resolved) && inside_destination?(resolved, root)
122
+
123
+ resolved
124
+ end
125
+
126
+ def inside_destination?(path, root)
127
+ prefix = root.end_with?(File::SEPARATOR) ? root : "#{root}#{File::SEPARATOR}"
128
+ path.start_with?(prefix)
129
+ end
130
+
131
+ def merge(path, query, fragment, tag)
132
+ merged = if query.match?(/(?:^|&)v=/)
133
+ query.sub(/(^|&)v=[^&]*/, "\\1v=#{tag}")
134
+ elsif query.empty?
135
+ "v=#{tag}"
136
+ else
137
+ "#{query}&v=#{tag}"
138
+ end
139
+ suffix = fragment.empty? ? "" : "##{fragment}"
140
+ "#{path}?#{merged}#{suffix}"
141
+ end
142
+ end
143
+ end
144
+ end
@@ -0,0 +1,7 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Jekyll
4
+ module FingerprintFlow
5
+ VERSION = "0.1.0"
6
+ end
7
+ end
@@ -0,0 +1,25 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "jekyll"
4
+ require "jekyll/fingerprint_flow/version"
5
+ require "jekyll/fingerprint_flow/configuration"
6
+ require "jekyll/fingerprint_flow/hasher"
7
+ require "jekyll/fingerprint_flow/url_tagger"
8
+ require "jekyll/fingerprint_flow/html_tag_scanner"
9
+ require "jekyll/fingerprint_flow/srcset_rewriter"
10
+ require "jekyll/fingerprint_flow/html_document_rewriter"
11
+ require "jekyll/fingerprint_flow/rewriter"
12
+ require "jekyll/fingerprint_flow/interface"
13
+
14
+ module Jekyll
15
+ module FingerprintFlow
16
+ Configuration::POST_WRITE_PRIORITIES.each do |hook_priority|
17
+ registered_priority = hook_priority
18
+ Jekyll::Hooks.register :site, :post_write, priority: registered_priority do |site|
19
+ next unless Configuration.post_write_priority(site) == registered_priority
20
+
21
+ Rewriter.new(site).run
22
+ end
23
+ end
24
+ end
25
+ end
@@ -0,0 +1,3 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "jekyll/fingerprint_flow"
metadata ADDED
@@ -0,0 +1,84 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: jekyll-fingerprint-flow
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - Svend Gundestrup
8
+ bindir: bin
9
+ cert_chain: []
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
+ dependencies:
12
+ - !ruby/object:Gem::Dependency
13
+ name: jekyll
14
+ requirement: !ruby/object:Gem::Requirement
15
+ requirements:
16
+ - - ">="
17
+ - !ruby/object:Gem::Version
18
+ version: '4.0'
19
+ - - "<"
20
+ - !ruby/object:Gem::Version
21
+ version: '5.0'
22
+ type: :runtime
23
+ prerelease: false
24
+ version_requirements: !ruby/object:Gem::Requirement
25
+ requirements:
26
+ - - ">="
27
+ - !ruby/object:Gem::Version
28
+ version: '4.0'
29
+ - - "<"
30
+ - !ruby/object:Gem::Version
31
+ version: '5.0'
32
+ description: A zero-config Jekyll plugin that content-tags supported local asset URLs
33
+ in generated .html/.htm attributes at post_write. Covers matching markup emitted
34
+ by themes and plugins without template changes. CSS/JS runtime references and HTTP
35
+ cache headers are outside its scope.
36
+ email:
37
+ - svend@gundestrup.dk
38
+ executables: []
39
+ extensions: []
40
+ extra_rdoc_files: []
41
+ files:
42
+ - CHANGELOG.md
43
+ - LICENSE.txt
44
+ - README.md
45
+ - interface.yml
46
+ - lib/jekyll-fingerprint-flow.rb
47
+ - lib/jekyll/fingerprint_flow.rb
48
+ - lib/jekyll/fingerprint_flow/configuration.rb
49
+ - lib/jekyll/fingerprint_flow/hasher.rb
50
+ - lib/jekyll/fingerprint_flow/html_document_rewriter.rb
51
+ - lib/jekyll/fingerprint_flow/html_tag_parser.rb
52
+ - lib/jekyll/fingerprint_flow/html_tag_scanner.rb
53
+ - lib/jekyll/fingerprint_flow/interface.rb
54
+ - lib/jekyll/fingerprint_flow/rewriter.rb
55
+ - lib/jekyll/fingerprint_flow/srcset_rewriter.rb
56
+ - lib/jekyll/fingerprint_flow/url_tagger.rb
57
+ - lib/jekyll/fingerprint_flow/version.rb
58
+ homepage: https://github.com/gundestrup/jekyll-fingerprint-flow
59
+ licenses:
60
+ - AGPL-3.0-or-later
61
+ metadata:
62
+ homepage_uri: https://github.com/gundestrup/jekyll-fingerprint-flow
63
+ source_code_uri: https://github.com/gundestrup/jekyll-fingerprint-flow/tree/main
64
+ changelog_uri: https://github.com/gundestrup/jekyll-fingerprint-flow/blob/main/CHANGELOG.md
65
+ bug_tracker_uri: https://github.com/gundestrup/jekyll-fingerprint-flow/issues
66
+ rubygems_mfa_required: 'true'
67
+ rdoc_options: []
68
+ require_paths:
69
+ - lib
70
+ required_ruby_version: !ruby/object:Gem::Requirement
71
+ requirements:
72
+ - - ">="
73
+ - !ruby/object:Gem::Version
74
+ version: 3.3.0
75
+ required_rubygems_version: !ruby/object:Gem::Requirement
76
+ requirements:
77
+ - - ">="
78
+ - !ruby/object:Gem::Version
79
+ version: '0'
80
+ requirements: []
81
+ rubygems_version: 3.6.9
82
+ specification_version: 4
83
+ summary: Content-hash URL cache busting for Jekyll HTML output
84
+ test_files: []