mikser-io 9.70.0 → 9.73.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -64,6 +64,46 @@ These options are part of `runtime.options` and apply to the engine itself.
64
64
  | `server.requestTimeout` | — | number | node default (`300000`) | Milliseconds a single request may take, set on the underlying `http.Server`. Node's 5-minute default is effectively an **upload size limit expressed in seconds** — a large file over a slow link is indistinguishable from a stalled request, so it is cut off and the caller sees a truncated write rather than a readable error. Raised to two hours automatically when any route registers as `streaming` (an upload surface such as `mikser-io-drive`), because Node's default mismeasures exactly those; set this explicitly to override. `0` disables the cap: reasonable on a trusted-network build server, bad facing the internet, where it removes the only bound on how long a client can hold a connection open doing nothing. `headersTimeout` is clamped to stay at or below it. The server is also exposed as `runtime.options.httpServer`. |
65
65
  | `url` | `-u, --url <url>` | string | — | Public URL where this mikser is reachable (e.g. `https://blog.me.com`). Validated, trailing slash stripped, stamped on `runtime.options.url`. Read by webhook-capable plugins for push-vs-poll gating (`url.startsWith('https://')`); used by anything that surfaces absolute URLs externally — MCP preview URLs returned to agents, forms share links, email tracking pixels. Plugins that just need internal URLs keep using `runtime.options.port`. |
66
66
 
67
+ ### `siteRoots`
68
+
69
+ Which subtrees of the output folder are deployed as their own domain root.
70
+
71
+ ```js
72
+ export default {
73
+ // out/bg, out/en and out/mk each deploy to their own domain
74
+ siteRoots: ['bg', 'en', 'mk'],
75
+ }
76
+ ```
77
+
78
+ Read by the url helpers and by the broken-reference check.
79
+
80
+ **It changes the output.** `asset`, `href` and `resource` build a path from the
81
+ page to the target, and without this they measure from the output folder. When
82
+ `out/bg` is what gets deployed, that is one directory too far — every url
83
+ carries an extra `..` for the site segment. A browser floors a climb above the
84
+ origin root rather than failing, so those urls load and nothing reports them.
85
+ Declaring the roots makes the helpers measure from the site instead, and the
86
+ urls become what they should have been.
87
+
88
+ The check reads it for the same reason, and resolves urls the way a browser
89
+ does. Where the site root is decides whether `../../x.svg` on a given page is
90
+ correct, merely over-deep, or broken.
91
+
92
+ Default is the output folder itself, which is right for the ordinary case of
93
+ one site per build — and with nothing declared every url is byte-identical to
94
+ what it was. Only a build that emits several sites moves, and it moves from
95
+ working-by-flooring to correct.
96
+
97
+ Declaring it also asserts something: **a site root is a deployable unit.** It
98
+ is served alone, so anything its pages reference has to exist beneath it —
99
+ share a common assets folder into each root rather than beside them. A url to
100
+ a target in a *different* root is left as it was, because on a per-domain
101
+ deploy no relative path reaches another origin; the check will report it, which
102
+ is the honest answer.
103
+
104
+ Nothing can infer any of this: it is a fact about where the bytes get deployed,
105
+ not about the bytes.
106
+
67
107
  ## Engine Substrate
68
108
 
69
109
  The catalog, inverse-ref graph, render snapshot manifest, and per-cycle
@@ -231,7 +231,7 @@ the whole question:
231
231
  ```json
232
232
  { "destination": "/index.html", "reason": "query-matched",
233
233
  "matched": { "filter": { "id": { "$regex": "^/documents/devices/" } },
234
- "by": "/documents/devices/hera.md" } }
234
+ "by": "/documents/devices/model-a.md" } }
235
235
  ```
236
236
 
237
237
  A `matched.filter` of `null` is a different statement: the page's predicate
@@ -873,10 +873,30 @@ surfaces that turn silence into a statement:
873
873
  `resource()` *build* a URL from a naming convention rather than looking
874
874
  an entity up, so they cannot fail: a preset that never ran, or a
875
875
  template naming an extension the preset no longer emits, yields a
876
- well-formed URL to nothing. Every such call is recorded on the render
877
- track and checked against the output folder at finalize; what is missing
878
- is warned under `asset-missing`, naming the path and the pages that
879
- linked it.
876
+ well-formed URL to nothing. Two checks run at finalize, from different
877
+ evidence:
878
+ - **The emitted output**, read back and resolved the way a browser
879
+ would — `src`, `href`, `poster`, `srcset` and CSS `url()` across
880
+ html and css. Anything resolving to no file warns under
881
+ `reference-broken`, naming the url and the pages carrying it. This
882
+ one sees paths written by hand, not just helper output.
883
+ - **The render track**, which records every `asset()` / `resource()`
884
+ call and tests the destination it built. This catches a url that
885
+ never reaches an html file at all — one emitted into a feed or a
886
+ sitemap — and warns under `asset-missing`. Where both can see the
887
+ same file, the output scan reports it and this one stays quiet.
888
+ - **A link that works only by accident** — a url with one `..` too many
889
+ still loads, because a browser discards a climb above the origin root
890
+ rather than failing. Reported under `reference-over-deep`, separately
891
+ from the outright failures, and grouped by how far each climbed:
892
+ - **One url, or several climbing different distances** — each is a
893
+ latent 404, working today and broken as soon as the same markup
894
+ renders one level deeper.
895
+ - **Every url climbing the same distance** — not N problems but one
896
+ base that is off by a constant, reported once and flagged
897
+ `structural` in `--json`. The urls work at every depth. Usually it
898
+ means `siteRoots` is undeclared for a build that emits several
899
+ sites; see [configuration](./configuration.md#siteroots).
880
900
 
881
901
  ## See also
882
902
 
package/docs/rendering.md CHANGED
@@ -392,7 +392,15 @@ naming the path and the pages that linked it, with a summary under
392
392
  deployed into the output as a link resolves normally.
393
393
 
394
394
  It is a warning, not an error: a build can legitimately link a file that
395
- some later step supplies. What it removes is the silence.
395
+ some later step supplies, and a missing asset must not stop a dev server.
396
+ What it removes is the silence.
397
+
398
+ A second check reads the **emitted output** rather than the render track —
399
+ every `src` / `href` / `poster` / `srcset` / CSS `url()` in the html and
400
+ css that shipped, resolved the way a browser resolves it. It covers paths
401
+ written by hand, which no helper ever saw, and it is the only one that can
402
+ see a url which loads solely because the browser floored a `..` run at the
403
+ site root. See [diagnostics](./diagnostics.md#when-mikser-is-silent).
396
404
 
397
405
  ---
398
406
 
package/package.json CHANGED
@@ -1,6 +1,17 @@
1
1
  {
2
2
  "name": "mikser-io",
3
- "version": "9.70.0",
3
+ "version": "9.73.0",
4
+ "files": [
5
+ "app.js",
6
+ "index.js",
7
+ "src/",
8
+ "testing/",
9
+ "docs/",
10
+ "favicon.ico",
11
+ "favicon.svg",
12
+ "mikser-mark.svg",
13
+ "mikser-lockup-stacked.svg"
14
+ ],
4
15
  "description": "A mixer for content: entities in, configurable render pipelines, outputs of any kind. Static sites are the canonical recipe, not the definition — the same engine renders PDFs, emails and whatever a renderer plugin produces. Files are the source of truth, every lifecycle phase is observable, and the build graph is queryable by an agent.",
5
16
  "main": "index.js",
6
17
  "exports": {
package/src/engine.js CHANGED
@@ -12,6 +12,7 @@ import { globby } from 'globby'
12
12
  import { OPERATION, TASKS } from './constants.js'
13
13
  import { changeExtension, formatErrorContext, projectMeta, lookupKeys } from './utils.js'
14
14
  import { reportRendered, reportSkipped, reportError, renderErrorCount, emitReport, finishCycle, reportAssetUse, assetUse } from './report.js'
15
+ import { checkReferences } from './references.js'
15
16
  import { toolSchemas, invokeTool, toolResultText, toolResultFailed } from './tools.js'
16
17
  import { registerBuiltinTools } from './builtin-tools.js'
17
18
  import { useDatabase } from './database/index.js'
@@ -89,6 +90,87 @@ function workerSafeOptions(opts) {
89
90
  return result
90
91
  }
91
92
 
93
+ // Warn for anything the EMITTED output points at that is not there.
94
+ //
95
+ // Complements the helper-call check below rather than repeating it: this reads
96
+ // what shipped, so it also sees paths written by hand, and it resolves them the
97
+ // way a browser does, which is the only way to see a url that works solely
98
+ // because a `..` run was floored at the site root.
99
+ //
100
+ // A floored url is not broken today. It is the same markup one level deeper
101
+ // away from being broken, and it means the emitted depth does not match the
102
+ // page — so it is reported separately rather than folded in with the failures.
103
+ //
104
+ // Warn, never fail: a missing asset must not stop a dev server. Both lists
105
+ // carry stable codes into `--json` so a deploy script can decide for itself.
106
+ // Returns the set of broken targets so the helper-call check can skip them.
107
+ async function reportBrokenReferences(logger) {
108
+ const outputFolder = runtime.options.outputFolder
109
+ if (!outputFolder || !existsSync(outputFolder)) return new Set()
110
+
111
+ const siteRoots = runtime.config?.siteRoots ?? []
112
+ const { broken, overDeep, checked } = await checkReferences(outputFolder, { siteRoots })
113
+ if (!checked) return new Set()
114
+
115
+ const SHOWN = 10
116
+ const named = (files) =>
117
+ files.slice(0, 3).join(', ') + (files.length > 3 ? ` and ${files.length - 3} more` : '')
118
+
119
+ for (const { url, target, files } of broken.slice(0, SHOWN)) {
120
+ logger.warn({ code: 'reference-broken', url, target, files },
121
+ 'Resolves to nothing: %s (from %s) — %s', url, named(files), target)
122
+ }
123
+ if (broken.length) {
124
+ logger.warn({ code: 'reference-broken-summary', broken: broken.length, checked },
125
+ '%d of %d reference(s) in the output resolve to nothing%s. A URL helper builds the '
126
+ + 'path rather than looking it up, so these are links to files nothing produced.',
127
+ broken.length, checked, broken.length > SHOWN ? `, ${SHOWN} shown` : '')
128
+ }
129
+
130
+ // Grouped by how FAR each climbed, because a site whose every over-deep url
131
+ // climbs the same distance does not have N problems — it has one base that
132
+ // is off by a constant. Printing it N times is precisely how a real signal
133
+ // gets filtered out, which is the failure this check exists to prevent.
134
+ const byClimb = new Map()
135
+ for (const entry of overDeep) {
136
+ if (!byClimb.has(entry.floored)) byClimb.set(entry.floored, [])
137
+ byClimb.get(entry.floored).push(entry)
138
+ }
139
+ // One distance, many urls: structural. The helper's base is wrong, the urls
140
+ // are not — they load at every depth, because the climb is always floored.
141
+ // That wants a different reaction than a hand-written `../..` that happens
142
+ // to be right on one page and is a 404 waiting on the next.
143
+ const structural = byClimb.size === 1 && overDeep.length > 1
144
+
145
+ for (const [climb, entries] of [...byClimb].sort(([a], [b]) => a - b)) {
146
+ const examples = entries.slice(0, 3).map(e => e.url)
147
+ logger.warn(
148
+ {
149
+ code: 'reference-over-deep', climbs: climb, count: entries.length,
150
+ structural, urls: examples,
151
+ files: [...new Set(entries.flatMap(e => e.files))].slice(0, 3),
152
+ },
153
+ structural
154
+ ? '%d references climb %d level(s) above the site root — every one of them, by the '
155
+ + 'same amount. They load: a browser discards the extra `..`. What is wrong is the '
156
+ + 'base they were built from, not the links. Examples: %s'
157
+ : '%d reference(s) climb %d level(s) above the site root and load only because a '
158
+ + 'browser discards the extra `..`. Each breaks if the same markup renders one '
159
+ + 'level deeper. Examples: %s',
160
+ entries.length, climb, examples.join(', '))
161
+ }
162
+
163
+ if (overDeep.length) {
164
+ logger.warn({ code: 'reference-over-deep-summary', overDeep: overDeep.length, checked, structural },
165
+ '%d of %d reference(s) resolve above the site root.%s',
166
+ overDeep.length, checked,
167
+ siteRoots.length ? '' : ' No siteRoots are declared, so this resolved against the output '
168
+ + 'root — declare siteRoots if a subtree is deployed as its own domain.')
169
+ }
170
+
171
+ return new Set(broken.map(b => b.target))
172
+ }
173
+
92
174
  // Warn for anything a render linked to that is not in the output.
93
175
  //
94
176
  // Deliberately phrased as what was OBSERVED. Only entities that rendered this
@@ -96,7 +178,7 @@ function workerSafeOptions(opts) {
96
178
  // and says nothing about the rest — the same reasoning the assets plugin
97
179
  // already applies to its preset warning, and for the same reason: a warning
98
180
  // that overclaims gets filtered, and the filtered-out line is the real one.
99
- async function reportMissingAssets(logger) {
181
+ async function reportMissingAssets(logger, alreadyReported = new Set()) {
100
182
  const used = assetUse()
101
183
  if (!used.length) return
102
184
  const outputFolder = runtime.options.outputFolder
@@ -105,6 +187,10 @@ async function reportMissingAssets(logger) {
105
187
  const missing = []
106
188
  for (const [destination, ids] of used) {
107
189
  const file = path.join(outputFolder, destination.replace(/^\//, ''))
190
+ // The output scan resolves the same file the way a browser does and
191
+ // names the pages that link it, which is strictly more useful. Where
192
+ // both would fire, one warning is enough.
193
+ if (alreadyReported.has(destination.replace(/^\//, ''))) continue
108
194
  if (!existsSync(file)) missing.push([destination, ids])
109
195
  }
110
196
  if (!missing.length) return
@@ -412,6 +498,15 @@ export async function setup(options) {
412
498
  // runtime.config) and before any plugin's onLoaded.
413
499
  onLoad(async () => {
414
500
  const logger = useLogger()
501
+
502
+ // Onto OPTIONS, not left on config: the url helpers read it, and they
503
+ // run in render workers, which receive the worker-safe options and
504
+ // never see runtime.config.
505
+ runtime.options.siteRoots = runtime.config?.siteRoots ?? []
506
+ if (runtime.options.siteRoots.length) {
507
+ logger.info('Site roots: %s', runtime.options.siteRoots.join(', '))
508
+ }
509
+
415
510
  const cli = runtime.options.url
416
511
  const cfg = runtime.config?.url
417
512
  const raw = cli ?? cfg
@@ -1138,7 +1233,8 @@ export async function setup(options) {
1138
1233
  // Checked at the end of the cycle because that is the first moment the
1139
1234
  // answer is stable: derivatives are produced during the cycle, so
1140
1235
  // asking any earlier would report files that were about to appear.
1141
- await reportMissingAssets(useLogger())
1236
+ const brokenTargets = await reportBrokenReferences(useLogger())
1237
+ await reportMissingAssets(useLogger(), brokenTargets)
1142
1238
 
1143
1239
  // After the cycle, and only under --json. stdout has been kept clear
1144
1240
  // for exactly this (the logger writes to stderr under --json), so the
@@ -1,6 +1,6 @@
1
1
  import path from 'node:path'
2
2
 
3
- import { changeExtension } from '../../utils.js'
3
+ import { changeExtension, siteRelativeUrl } from '../../utils.js'
4
4
 
5
5
  // `{{asset 'web' '/media/hero.jpg'}}` — the deployed URL of a preset
6
6
  // derivative, relative to the page asking for it.
@@ -71,8 +71,10 @@ export function load({ runtime, entity, state, options, logger, track }) {
71
71
  // itself — it takes a path, not an entity, so there is nothing to look
72
72
  // up and the URL is well-formed whether or not anything produced it.
73
73
  track?.asset?.(destination)
74
- const from = path.dirname(entity.destination || '/')
75
- return { url: path.relative(from, destination) }
74
+ // Resolved within the site this page belongs to, not against the
75
+ // output root — see siteRelativeUrl. With no siteRoots declared the two
76
+ // are the same folder and the url is byte-identical.
77
+ return { url: siteRelativeUrl(entity.destination, destination, options?.siteRoots) }
76
78
  }
77
79
  }
78
80
 
@@ -72,7 +72,7 @@ function warnIfUntrackable(options, resolved, logger) {
72
72
  //
73
73
  // The first version of this hardcoded five content folders, and a project
74
74
  // registering its own collections through sources() has more than five. On
75
- // lmed that meant 63 warnings per build, one for every stylesheet and
75
+ // one real site that meant 63 warnings per build, one for every stylesheet and
76
76
  // script, all of them tracked correctly and every one of them saying the
77
77
  // opposite. Which is worse than not warning: 63 spurious lines a build
78
78
  // teaches you to filter the channel, and the filtered-out line is the real
@@ -1,4 +1,5 @@
1
1
  import path from 'node:path'
2
+ import { siteRelativeUrl } from '../../utils.js'
2
3
 
3
4
  export function load({ entity, runtime, options }) {
4
5
  const { clear } = options
@@ -19,16 +20,14 @@ export function load({ entity, runtime, options }) {
19
20
 
20
21
  let found = runtime.hrefLang(href)
21
22
  if (!found) {
22
- const from = path.dirname(entity.destination || '/')
23
- return { url: path.relative(from, href) }
23
+ return { url: siteRelativeUrl(entity.destination, href, options?.siteRoots) }
24
24
  } else {
25
25
  if (!found.id) {
26
26
  found = found[lang]
27
27
  }
28
28
  if (found?.destination) {
29
29
  const destination = clear ? found.destination.replace('index.html', '') : found.destination
30
- const from = path.dirname(entity.destination || '/')
31
- found.url = path.relative(from, destination)
30
+ found.url = siteRelativeUrl(entity.destination, destination, options?.siteRoots)
32
31
  }
33
32
  return found
34
33
  }
@@ -1,10 +1,14 @@
1
1
  import path from 'node:path'
2
+ import { matchesLibrary, siteRelativeUrl } from '../../utils.js'
2
3
 
3
4
  export function load({ runtime, entity, state, options, track }) {
4
5
  runtime.resource = (url) => {
5
6
  const { resourceLib } = state.resources
6
7
  for (let library in resourceLib) {
7
- if (url.match(library)) {
8
+ // Same matcher the resources plugin uses to decide what to
9
+ // DOWNLOAD. When these disagreed, this built urls for files the
10
+ // plugin never fetched.
11
+ if (matchesLibrary(url, library)) {
8
12
  const { origin } = new URL(url)
9
13
  const name = url.replace(origin, `${resourceLib[library]}`)
10
14
  const relative = url.replace(origin, `${state.resources.resourcesFolder}/${resourceLib[library]}`)
@@ -13,8 +17,7 @@ export function load({ runtime, entity, state, options, track }) {
13
17
  // than resolving one, so a library that was never copied
14
18
  // yields a link to nothing on a green build.
15
19
  track?.asset?.(destination)
16
- const from = path.dirname(entity.destination || '/')
17
- return { url: path.relative(from, destination), name }
20
+ return { url: siteRelativeUrl(entity.destination, destination, options?.siteRoots), name }
18
21
  }
19
22
  }
20
23
  }
@@ -10,6 +10,7 @@ import * as stream from 'stream'
10
10
  import { promisify } from 'util'
11
11
  import isUrl from 'is-url'
12
12
  import map from 'p-map'
13
+ import { matchesLibrary } from '../utils.js'
13
14
 
14
15
  export function resources(options = {}) {
15
16
  return ({
@@ -24,7 +25,6 @@ export function resources(options = {}) {
24
25
  checksum,
25
26
  trackProgress,
26
27
  updateProgress,
27
- matchEntity,
28
28
  constants: { OPERATION },
29
29
  }) => {
30
30
  const collection = 'resources'
@@ -52,6 +52,11 @@ export function resources(options = {}) {
52
52
 
53
53
  for (let library in (options.libraries || [])) {
54
54
  let resource = options.libraries[library]
55
+ // The key is a REGULAR EXPRESSION source, which is what the
56
+ // escapeStringRegexp call says: you only escape a string you are
57
+ // about to compile. The render helper has always read it that way
58
+ // (`url.match(library)`), so a library declared by `url` is a
59
+ // prefix pattern matching anything under it.
55
60
  runtime.state.resources.resourceLib[resource.match || escapeStringRegexp(resource.url)] = library
56
61
  }
57
62
  })
@@ -66,7 +71,15 @@ export function resources(options = {}) {
66
71
  _.eachDeep(entity.meta, resource => {
67
72
  if (typeof resource == 'string') {
68
73
  for (let library in resourceLib) {
69
- if (matchEntity(resource, library)) {
74
+ // Regex, matching the render helper. This used
75
+ // matchEntity, which is a GLOB demanding a full
76
+ // match — so a key derived from `url` (a bare
77
+ // prefix, no trailing wildcard) matched nothing and
78
+ // NO url-declared library was ever downloaded. The
79
+ // helper still built urls for them, so pages linked
80
+ // files nothing fetched and the build stayed green:
81
+ // one string read with two incompatible matchers.
82
+ if (matchesLibrary(resource, library)) {
70
83
  resourceMap[entity.id].push({ library, resource, entity })
71
84
  }
72
85
  }
@@ -0,0 +1,177 @@
1
+ // What the build actually shipped, checked against what it actually wrote.
2
+ //
3
+ // The URL helpers BUILD paths from a naming convention rather than resolving
4
+ // an entity, so they cannot fail — `asset` composes
5
+ // `<assetsFolder>/<preset>/<path>` with whatever extension it was handed and
6
+ // never asks whether that file exists. A wrong preset name, a wrong extension,
7
+ // or a source whose derivative silently failed to render all produce a
8
+ // well-formed url pointing at nothing, and every existing surface stays green:
9
+ // nothing threw, --verify compares snapshots against what was rendered rather
10
+ // than against what those renders point at, and mikser_refs_broken tracks
11
+ // document-to-document refs, not urls.
12
+ //
13
+ // This reads the emitted bytes instead. Everything it needs is on disk at the
14
+ // end of a cycle and nothing has to be inferred.
15
+ //
16
+ // It is deliberately NOT the same check as `asset-missing` in engine.js. That
17
+ // one records helper CALLS on the render track and tests the output-root
18
+ // absolute destination each one built; it knows the referencing entity, and it
19
+ // sees urls that never reach an html file at all (a sitemap, a feed). This one
20
+ // sees everything that shipped, including paths written by hand, and resolves
21
+ // them the way a browser would — which is the only way to catch a url that
22
+ // resolves solely because the browser floored a `..` run at the site root.
23
+
24
+ import path from 'node:path'
25
+ import { existsSync } from 'node:fs'
26
+ import { readFile } from 'node:fs/promises'
27
+ import { globby } from 'globby'
28
+ import { siteRootFor } from './utils.js'
29
+
30
+ // Documents that can carry a reference. Anything else in the output is either
31
+ // an asset itself or something whose internal structure this has no business
32
+ // guessing at.
33
+ const SCANNED = ['**/*.html', '**/*.htm', '**/*.css']
34
+
35
+ // Attributes whose value is a single url.
36
+ const ATTR = /(?:src|href|poster|data-bg)\s*=\s*["']([^"']*)["']/gi
37
+ // srcset / imagesrcset: a comma-separated list of `url [descriptor]`.
38
+ const SRCSET = /(?:img|image)?srcset\s*=\s*["']([^"']*)["']/gi
39
+ // css url(), in a stylesheet and in an inline style attribute alike.
40
+ const CSS_URL = /url\(\s*(['"]?)([^'")]*)\1\s*\)/gi
41
+
42
+ // A url this check has nothing to say about: another origin, an inline
43
+ // payload, a fragment or an in-page action. `//host/path` is protocol-relative
44
+ // and therefore external too.
45
+ function isExternal(url) {
46
+ if (!url) return true
47
+ const u = url.trim()
48
+ if (!u) return true
49
+ if (u.startsWith('#') || u.startsWith('//')) return true
50
+ // Template syntax that reached the output unrendered — `{{link}}`,
51
+ // `${x}`, `<%= y %>`. It is not a path, so "resolves to nothing" says
52
+ // nothing useful about it; the real problem is that it did not render,
53
+ // which is a different question than this one is asking. Documentation
54
+ // pages showing escaped template syntax are the common source, and they
55
+ // are not broken at all.
56
+ if (/\{\{|\}\}|\$\{|<%/.test(u)) return true
57
+ // A scheme — http:, data:, mailto:, tel:, javascript:.
58
+ if (/^[a-z][a-z0-9+.-]*:/i.test(u)) return true
59
+ // The same thing percent-encoded, which is how an external url arrives
60
+ // when it was built as a query parameter — `https%3A%2F%2F...` in a maps
61
+ // link. It has no scheme until it is decoded, so the test above misses it
62
+ // and the whole encoded string gets resolved as a path segment.
63
+ try {
64
+ if (/^[a-z][a-z0-9+.-]*:/i.test(decodeURIComponent(u))) return true
65
+ } catch { /* malformed escape — treat as a path and let it resolve */ }
66
+ return false
67
+ }
68
+
69
+ // Quotes inside an attribute value arrive encoded, and a CSS custom property
70
+ // in an inline style is the common way that happens:
71
+ //
72
+ // style="--icon-src:url(&quot;../media/raw/icons/x.svg&quot;)"
73
+ //
74
+ // Without decoding, the captured url is the entity text itself, which resolves
75
+ // nowhere and reports as broken — a false positive that would have buried the
76
+ // real ones. Only the quote and ampersand forms are decoded; turning &lt; back
77
+ // into a bracket could invent markup that was deliberately escaped.
78
+ function decodeEntities(source) {
79
+ return source
80
+ .replace(/&quot;|&#34;/g, '"')
81
+ .replace(/&apos;|&#39;/g, "'")
82
+ .replace(/&amp;/g, '&')
83
+ }
84
+
85
+ // Everything a page points at, as raw url strings.
86
+ export { siteRootFor }
87
+
88
+ export function extractReferences(rawSource) {
89
+ const source = decodeEntities(rawSource)
90
+ const found = new Set()
91
+ for (const [, url] of source.matchAll(ATTR)) found.add(url)
92
+ for (const [, , url] of source.matchAll(CSS_URL)) found.add(url)
93
+ for (const [, list] of source.matchAll(SRCSET)) {
94
+ for (const candidate of list.split(',')) {
95
+ const url = candidate.trim().split(/\s+/)[0]
96
+ if (url) found.add(url)
97
+ }
98
+ }
99
+ return [...found].filter(u => !isExternal(u))
100
+ }
101
+
102
+ // Resolve the way a browser does, which is the whole point.
103
+ //
104
+ // A browser walks the page's directory segments, pops one per `..`, and
105
+ // DISCARDS a `..` that would climb above the origin root — it does not error
106
+ // and it does not escape. So a url with one `..` too many still loads, and the
107
+ // page looks correct while carrying a path that breaks the moment the same
108
+ // markup is used one level deeper. That flooring is what `overDeep` records.
109
+ //
110
+ // `pageDir` and the result are both relative to `root`.
111
+ export function resolveUrl(pageDir, url, { root = '' } = {}) {
112
+ const clean = url.split('#')[0].split('?')[0]
113
+ const absolute = clean.startsWith('/')
114
+ const segments = clean.split('/').filter(s => s !== '' && s !== '.')
115
+
116
+ const parts = absolute ? [] : pageDir.split('/').filter(Boolean)
117
+ // How FAR above the root it climbed, not merely that it did. When every
118
+ // over-deep url on a site climbs the same distance, that is one base
119
+ // mismatch reported once — not N findings, which is how a real signal gets
120
+ // filtered.
121
+ let floored = 0
122
+ for (const segment of segments) {
123
+ if (segment !== '..') { parts.push(segment); continue }
124
+ if (parts.length) parts.pop()
125
+ else floored++ // a climb above the root, discarded
126
+ }
127
+ return { target: path.join(root, ...parts), overDeep: floored > 0, floored }
128
+ }
129
+
130
+ // Everything the output points at that is not there.
131
+ //
132
+ // Returns { broken, overDeep, checked }, where each entry is
133
+ // { url, target, files } — the target with the pages that named it, because
134
+ // "this is missing" is only actionable next to "and these link it".
135
+ export async function checkReferences(outputFolder, { siteRoots = [] } = {}) {
136
+ const files = await globby(SCANNED, {
137
+ cwd: outputFolder,
138
+ followSymbolicLinks: false,
139
+ suppressErrors: true,
140
+ })
141
+
142
+ const broken = new Map()
143
+ const overDeepRefs = new Map()
144
+ let checked = 0
145
+ // Existence is the expensive part and the same target repeats across a
146
+ // site — one lookup each.
147
+ const exists = new Map()
148
+
149
+ for (const file of files) {
150
+ let source
151
+ try { source = await readFile(path.join(outputFolder, file), 'utf8') }
152
+ catch { continue }
153
+
154
+ const root = siteRootFor(file, siteRoots)
155
+ // The page's directory, relative to its own site root.
156
+ const pageDir = path.dirname(file).slice(root.length).replace(/^\/+/, '')
157
+
158
+ for (const url of extractReferences(source)) {
159
+ const { target, overDeep, floored } = resolveUrl(pageDir, url, { root })
160
+ checked++
161
+
162
+ if (!exists.has(target)) {
163
+ exists.set(target, existsSync(path.join(outputFolder, target)))
164
+ }
165
+ // Broken outranks over-deep: a url that resolves nowhere is the
166
+ // failure, and adding that it is also one level too deep is noise.
167
+ const bucket = !exists.get(target) ? broken : (overDeep ? overDeepRefs : null)
168
+ if (!bucket) continue
169
+
170
+ const key = `${target} ${url}`
171
+ if (!bucket.has(key)) bucket.set(key, { url, target, floored, files: [] })
172
+ bucket.get(key).files.push(file)
173
+ }
174
+ }
175
+
176
+ return { broken: [...broken.values()], overDeep: [...overDeepRefs.values()], checked }
177
+ }
package/src/utils.js CHANGED
@@ -1316,3 +1316,77 @@ export function junkFilter() {
1316
1316
  return registered.match.some(test => test(name))
1317
1317
  }
1318
1318
  }
1319
+
1320
+ // Does a value fall under a resources library?
1321
+ //
1322
+ // The library key is a REGULAR EXPRESSION source — `resources()` derives it
1323
+ // with escapeStringRegexp(url), and you only escape a string you are about to
1324
+ // compile. It has two consumers: the plugin's discovery walk, which decides
1325
+ // what to download, and the `resource` render helper, which builds the url.
1326
+ // They read the same string with two different matchers — discovery used a
1327
+ // GLOB, which demands a full match, so a key derived from `url` (a bare prefix
1328
+ // with no trailing wildcard) matched nothing. Nothing was ever downloaded for
1329
+ // a url-declared library, while the helper happily built links to the files
1330
+ // that were not fetched. Green build, missing images.
1331
+ //
1332
+ // One function, so the two cannot drift again.
1333
+ const libraryPatterns = new Map()
1334
+ export function matchesLibrary(value, pattern) {
1335
+ if (typeof value !== 'string' || !pattern) return false
1336
+ if (!libraryPatterns.has(pattern)) {
1337
+ let re
1338
+ try { re = new RegExp(pattern) }
1339
+ // A hand-written `match` that is not valid regex would otherwise throw
1340
+ // mid-walk and take the build down.
1341
+ catch { re = { test: () => false } }
1342
+ libraryPatterns.set(pattern, re)
1343
+ }
1344
+ return libraryPatterns.get(pattern).test(value)
1345
+ }
1346
+
1347
+ // Which declared site root a path belongs to.
1348
+ //
1349
+ // A build can emit one subtree per language and deploy each as its own domain
1350
+ // root, which puts the site root at out/<lang>/ rather than at out/. Nothing
1351
+ // can derive that — it is a fact about where the bytes get deployed, not about
1352
+ // the bytes — so it is declared as `siteRoots` and the default is the output
1353
+ // root itself. Accepts a path with or without a leading slash, because an
1354
+ // entity destination has one and an output-relative file path does not.
1355
+ export function siteRootFor(file, roots = []) {
1356
+ const relative = String(file ?? '').replace(/^\/+/, '')
1357
+ let best = ''
1358
+ for (const root of roots) {
1359
+ if (!root) continue
1360
+ if (relative.startsWith(`${root}/`) && root.length > best.length) best = root
1361
+ }
1362
+ return best
1363
+ }
1364
+
1365
+ // A page-relative url from one output destination to another, addressed within
1366
+ // the site the page belongs to.
1367
+ //
1368
+ // Both are output-root absolute (`/bg/aparati/index.html`, `/derived/x.webp`),
1369
+ // which is the only shape the engine has. With one site per build that is also
1370
+ // the deployed root and this is a plain path.relative. With several, it is not:
1371
+ // out/bg IS the domain root, so a url computed against out/ carries one extra
1372
+ // `..` for the language segment. The browser floors that rather than failing,
1373
+ // which is why it worked and why nothing said so.
1374
+ //
1375
+ // Three cases, and the middle one is the reason this is not a one-liner:
1376
+ //
1377
+ // target outside every root a shared asset. It has to be reachable from
1378
+ // inside this page's site, so it is addressed
1379
+ // there — this is the case that was wrong.
1380
+ // target in the same root already correct; a plain relative path.
1381
+ // target in a DIFFERENT root a cross-site link. On a per-domain deploy the
1382
+ // other site is another origin and no relative
1383
+ // path reaches it. Left as it was, so the
1384
+ // reference check reports it broken instead of
1385
+ // this silently inventing a path that is not.
1386
+ export function siteRelativeUrl(pageDestination, target, siteRoots = []) {
1387
+ const from = path.dirname(pageDestination || '/')
1388
+ const pageRoot = siteRootFor(pageDestination, siteRoots)
1389
+ if (!pageRoot) return path.relative(from, target)
1390
+ if (siteRootFor(target, siteRoots)) return path.relative(from, target)
1391
+ return path.relative(from, path.join('/', pageRoot, target))
1392
+ }