scavold 0.2.0-rc.5 → 0.2.0-rc.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +87 -4
- package/FRONTMATTER.md +52 -0
- package/lib/config.js +211 -11
- package/lib/href.d.ts +21 -0
- package/lib/legacyUrls.js +179 -0
- package/lib/pageKeys.js +111 -0
- package/lib/verifyBuild.js +324 -0
- package/package.json +4 -4
- package/scripts/check-fixture.js +32 -1
package/CHANGELOG.md
CHANGED
|
@@ -9,6 +9,85 @@ changes may occur in any release.
|
|
|
9
9
|
|
|
10
10
|
## [Unreleased]
|
|
11
11
|
|
|
12
|
+
## [0.2.0-rc.7] — 2026-09-29
|
|
13
|
+
|
|
14
|
+
### Fixed
|
|
15
|
+
|
|
16
|
+
- A link to a media file whose extension VitePress does not know, such as `.diff`,
|
|
17
|
+
pointed at the file name with `.html` appended and therefore at nothing. It now names
|
|
18
|
+
the file as published.
|
|
19
|
+
|
|
20
|
+
## [0.2.0-rc.6] — 2026-09-13
|
|
21
|
+
|
|
22
|
+
### Added
|
|
23
|
+
|
|
24
|
+
- `aliases` in a page's front matter names the former addresses it replaces, and
|
|
25
|
+
`retired_urls` in `.cratly.config.yaml` those that are not coming back. A site taking
|
|
26
|
+
over from an existing website declares them once and the build writes
|
|
27
|
+
`cratly-redirects.json` into its output: an address, where it went, and `301` or `410`.
|
|
28
|
+
The format names no server — only a server-side redirect passes a ranking on, and which
|
|
29
|
+
server that is cratly does not assume, so the deploy step translates the file for its
|
|
30
|
+
target. Keeping an address is still better than redirecting to it, which is what `url`
|
|
31
|
+
is for; the build refuses an alias that two pages claim, or that the site serves itself
|
|
32
|
+
and the redirect would therefore hide.
|
|
33
|
+
- Every build now ends by reading its own result. A page that the generator listed but
|
|
34
|
+
never wrote, or a script a built page names and the output does not contain, ends the
|
|
35
|
+
build with the file named — the two states a bundling fault leaves behind while the log
|
|
36
|
+
says "build complete". Links and images pointing at absent files are reported and left
|
|
37
|
+
to the site, since one may legitimately name a path the web server provides. Turn it
|
|
38
|
+
down with `verify: { content: false }` or off with `verify: false`. A reference is read
|
|
39
|
+
the way a browser reads it: a relative one against the address of the page holding it, so
|
|
40
|
+
a link that forgot its protocol — `[Website](./www.example.com)`, which the build turns
|
|
41
|
+
into a sibling page nobody wrote — is seen as well. Proven against the
|
|
42
|
+
case-collision that used to build green: the page missing from the output is the only
|
|
43
|
+
trace it leaves on a file system that ignores letter case, and that is now enough.
|
|
44
|
+
Everything found also goes to `.cratly/build-report.json`, written before the build is
|
|
45
|
+
ended rather than after, each entry naming the source file an author would open. A build
|
|
46
|
+
log is read by whoever has access to it; a file can be handed to whoever wrote the page.
|
|
47
|
+
- `ignoreDeadLinks` now defaults to `true`. VitePress ends a build on a dead link and does
|
|
48
|
+
so before writing anything, which leaves the reason in a job log that whoever wrote the
|
|
49
|
+
link often cannot read, and lets one typo stand between a finished page and its
|
|
50
|
+
publication. Scavold finds the same mistake in the finished output and reports it there
|
|
51
|
+
instead. Whether it should also stop a pipeline is a question whose answer differs per
|
|
52
|
+
branch, so it is left to CI: the site template fails a merge request over it and lets the
|
|
53
|
+
default branch deploy, since blocking there would answer a broken link with a stale site.
|
|
54
|
+
Ask for VitePress' abort back with `ignoreDeadLinks: false`.
|
|
55
|
+
- A site whose page paths collide in VitePress's eyes now learns that from Scavold, by
|
|
56
|
+
name, before the build starts. VitePress identifies a page by its path flattened into
|
|
57
|
+
a single name — slash becomes underscore, and the page-hash map lower-cases it — and
|
|
58
|
+
uses that name for the page's bundle entry, its server module and its client chunk.
|
|
59
|
+
Two pages meeting there overwrite each other: either the build dies at render time in
|
|
60
|
+
`pageChunk.imports`, naming no file, or it succeeds and ships two pages pointing at a
|
|
61
|
+
script that was never written. Scavold now reproduces that identity when it reads the
|
|
62
|
+
config and aborts with both file names, whether they collide through an underscore in
|
|
63
|
+
a file name (`de_kontakt.md` beside `de/kontakt.md`), through letter case alone, or
|
|
64
|
+
through a `url` alias landing on a path another page already owns as its file — the
|
|
65
|
+
case the previous alias check, which compared aliases only with each other, let pass.
|
|
66
|
+
- `date` and `datetime` join the type vocabulary of `frontmatter_fields` and of a
|
|
67
|
+
container's `props`. Scavold passes such a value through as the ISO 8601 text it is;
|
|
68
|
+
what changes is that the cratly editor now offers a date picker for it instead of a
|
|
69
|
+
plain text box. A `datetime` always carries an offset — a time of day without one is
|
|
70
|
+
a different moment on every machine that builds the site.
|
|
71
|
+
|
|
72
|
+
### Changed
|
|
73
|
+
|
|
74
|
+
- `@cepharum/vue3-i18n` moves to `2.0.0`, two majors on from the `0.5.4` a site would
|
|
75
|
+
have installed. Scavold uses `useL10n`, `setLocale` and `setLoader`, none of which
|
|
76
|
+
changed; what 2.0 breaks are the locale helpers now returning full BCP-47 tags, and
|
|
77
|
+
Scavold calls none of them. A site inherits the new version with its next install, and
|
|
78
|
+
gains what 2.0 adds along the way: regional locales overlaying a bare language, plural
|
|
79
|
+
and gender selection resolved through `Intl`, and `Intl`-backed formatters in
|
|
80
|
+
placeholders.
|
|
81
|
+
|
|
82
|
+
### Fixed
|
|
83
|
+
|
|
84
|
+
- `lib/href.js` ships the declarations it had been missing since it was split out, so
|
|
85
|
+
`typecheck` passes again — and the pipeline runs it now, which is why it could go
|
|
86
|
+
unnoticed at all. Nothing about the module changed; TypeScript simply had nothing to
|
|
87
|
+
read about it, and treated every href the hierarchy composes as `any`.
|
|
88
|
+
|
|
89
|
+
## [0.2.0-rc.5] — 2026-08-17
|
|
90
|
+
|
|
12
91
|
### Added
|
|
13
92
|
|
|
14
93
|
- `:::pagelist` dates its entries the way the site wants them: `time-style` adds the
|
|
@@ -179,7 +258,11 @@ First published release. Version `0.1.0` existed in-tree only.
|
|
|
179
258
|
- Only images are processed out of `media_folder`; other file types (video, documents)
|
|
180
259
|
are never copied into the build output and have to live in the static folder.
|
|
181
260
|
|
|
182
|
-
[
|
|
183
|
-
[0.2.0-rc.
|
|
184
|
-
[0.2.0-rc.
|
|
185
|
-
[0.2.0-rc.
|
|
261
|
+
[Unreleased]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.7...main
|
|
262
|
+
[0.2.0-rc.7]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.6...v0.2.0-rc.7
|
|
263
|
+
[0.2.0-rc.6]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.5...v0.2.0-rc.6
|
|
264
|
+
[0.2.0-rc.5]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.4...v0.2.0-rc.5
|
|
265
|
+
[0.2.0-rc.4]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.3...v0.2.0-rc.4
|
|
266
|
+
[0.2.0-rc.3]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.2...v0.2.0-rc.3
|
|
267
|
+
[0.2.0-rc.2]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.1...v0.2.0-rc.2
|
|
268
|
+
[0.2.0-rc.1]: https://gitlab.com/cratly/scavold/-/tags/v0.2.0-rc.1
|
package/FRONTMATTER.md
CHANGED
|
@@ -107,6 +107,20 @@ Menu links and breadcrumbs produced by Scavold automatically use the alias URL.
|
|
|
107
107
|
with an error naming both conflicting files. Each alias must be unique across the
|
|
108
108
|
entire site.
|
|
109
109
|
|
|
110
|
+
The same applies when an alias lands on a path another page already owns as its own
|
|
111
|
+
file — a `url: de/kontakt` next to a `pages/de/kontakt.md`. VitePress identifies a
|
|
112
|
+
page by its path flattened into one name, slash becoming underscore, and it uses that
|
|
113
|
+
name for the page's bundle entry, its server module and its client chunk. Two pages
|
|
114
|
+
meeting there overwrite each other's output; the alias does not win, it collides. So
|
|
115
|
+
does a page whose file name spells out a path another page reaches through folders
|
|
116
|
+
(`de_kontakt.md` beside `de/kontakt.md`), and a pair differing only in letter case.
|
|
117
|
+
Scavold checks all of these when it reads the config and names both files. Prefer `-`
|
|
118
|
+
over `_` as the word separator in page names and the class cannot arise.
|
|
119
|
+
|
|
120
|
+
To move a page and keep its URL, rename the file and declare the old URL as its
|
|
121
|
+
alias — the alias belongs to the page that carries the content, not to a second file
|
|
122
|
+
left behind at the old location.
|
|
123
|
+
|
|
110
124
|
---
|
|
111
125
|
|
|
112
126
|
## `locale` / `lang`
|
|
@@ -260,6 +274,44 @@ in the hierarchy tree — they are only excluded from rendered navigation.
|
|
|
260
274
|
|
|
261
275
|
---
|
|
262
276
|
|
|
277
|
+
## `aliases`
|
|
278
|
+
|
|
279
|
+
**Type:** `string | string[]`
|
|
280
|
+
|
|
281
|
+
Former addresses this page replaces. Declared when a cratly site takes over from an
|
|
282
|
+
existing website: every address a search engine already knows and the new site does not
|
|
283
|
+
answer is a result its owner loses.
|
|
284
|
+
|
|
285
|
+
```yaml
|
|
286
|
+
---
|
|
287
|
+
aliases:
|
|
288
|
+
- /kontakt.html
|
|
289
|
+
- kontakt.php
|
|
290
|
+
---
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
Values are written as they stood in the browser, with or without a leading slash. A
|
|
294
|
+
fragment is dropped — it never reaches a server. A query string is kept; whether it can
|
|
295
|
+
be matched is up to whoever serves the site.
|
|
296
|
+
|
|
297
|
+
Scavold collects them and writes `cratly-redirects.json` into the build output: an
|
|
298
|
+
address, where it went, and a status — `301` for an inherited address, `410` for one a
|
|
299
|
+
site declares as gone for good in
|
|
300
|
+
[`retired_urls`](./docs/reference/config.md#retired-urls). Turning that into server
|
|
301
|
+
configuration is the deploy step's job, because cratly does not assume who serves the
|
|
302
|
+
site. Nothing about the file is Scavold-specific: another adapter writes the same format.
|
|
303
|
+
|
|
304
|
+
**Keeping the address beats redirecting to it.** A permanent redirect passes a ranking on;
|
|
305
|
+
a page that simply answers at the old address never spends it. Where the legacy path can
|
|
306
|
+
be a file, [`url`](#url) is the better tool, and it works on every static server, whether
|
|
307
|
+
or not it can redirect at all.
|
|
308
|
+
|
|
309
|
+
**Conflict detection:** the build stops when two pages claim the same former address, and
|
|
310
|
+
when a former address is one the new site serves itself — the redirect would win over the
|
|
311
|
+
page, leaving it unreachable while every menu still links to it.
|
|
312
|
+
|
|
313
|
+
---
|
|
314
|
+
|
|
263
315
|
## `redirect`
|
|
264
316
|
|
|
265
317
|
**Type:** `string | object`
|
package/lib/config.js
CHANGED
|
@@ -1,7 +1,9 @@
|
|
|
1
1
|
import { dirname as dirnameOfPath, join, relative, resolve } from "node:path";
|
|
2
2
|
import { dirname, join as joinPosix } from "node:path/posix";
|
|
3
3
|
import { readdirSync, readFileSync } from "node:fs";
|
|
4
|
-
import {
|
|
4
|
+
import {
|
|
5
|
+
clearCache, collectFrontmatterOfAllPages, compileHierarchy, compileRedirects, resolveHierarchyImages, sourceFolder,
|
|
6
|
+
} from "./pages.js";
|
|
5
7
|
import { extractFrontmatterMediaSrcs } from "./excerpt.js";
|
|
6
8
|
import { DEFAULT_FEED_LIMIT, collectFeedPages, renderRss } from "./feed.js";
|
|
7
9
|
import { useMedia } from "./media.js";
|
|
@@ -9,6 +11,13 @@ import { patchRenderer } from "./markdown.js";
|
|
|
9
11
|
import { extractContainerMediaSrcs, mediaPropsOf, registerContainers, resolveContainerMap } from "./containers.js";
|
|
10
12
|
import { buildSectionManifest, writeSectionManifest } from "./sectionManifest.js";
|
|
11
13
|
import { isExternalUrl, servableRedirectTarget } from "./redirectTarget.js";
|
|
14
|
+
import { assertUniquePageKeys } from "./pageKeys.js";
|
|
15
|
+
import {
|
|
16
|
+
BUILD_REPORT_DIR, BUILD_REPORT_FILE, formatProblems, renderBuildReport, verifyBuild,
|
|
17
|
+
} from "./verifyBuild.js";
|
|
18
|
+
import {
|
|
19
|
+
REDIRECT_MANIFEST_FILE, assertLegacyRedirects, compileLegacyRedirects, renderRedirectManifest,
|
|
20
|
+
} from "./legacyUrls.js";
|
|
12
21
|
import { pageHref, servableUrl } from "./href.js";
|
|
13
22
|
|
|
14
23
|
// Scavold's own version, reported under `adapter.version` in the emitted
|
|
@@ -180,6 +189,7 @@ export async function augmentConfig( rawConfig, options = {} ) {
|
|
|
180
189
|
let declaredImageSizes = null;
|
|
181
190
|
let declaredImageWidths = null;
|
|
182
191
|
let declaredExcerptLength = null;
|
|
192
|
+
let declaredRetiredUrls = [];
|
|
183
193
|
|
|
184
194
|
try {
|
|
185
195
|
const { readFile } = await import( "node:fs/promises" );
|
|
@@ -193,6 +203,7 @@ export async function augmentConfig( rawConfig, options = {} ) {
|
|
|
193
203
|
declaredImageSizes = cratlyConfig.image_sizes ?? null;
|
|
194
204
|
declaredImageWidths = cratlyConfig.image_widths ?? null;
|
|
195
205
|
declaredExcerptLength = cratlyConfig.excerpt_length ?? null;
|
|
206
|
+
declaredRetiredUrls = cratlyConfig.retired_urls ?? [];
|
|
196
207
|
} catch {
|
|
197
208
|
// file absent or unparseable — proceed with defaults
|
|
198
209
|
}
|
|
@@ -376,6 +387,99 @@ export async function augmentConfig( rawConfig, options = {} ) {
|
|
|
376
387
|
await writeFile( target, xml, "utf-8" );
|
|
377
388
|
}
|
|
378
389
|
|
|
390
|
+
/**
|
|
391
|
+
* Writes what the check found next to the section manifest, for readers that are
|
|
392
|
+
* not a terminal — the cratly editor shows it to whoever wrote the page.
|
|
393
|
+
*
|
|
394
|
+
* A report that cannot be written must not replace the finding it describes, so
|
|
395
|
+
* failure here is a warning and the build carries on to its own verdict.
|
|
396
|
+
*
|
|
397
|
+
* @param {{pages: object[], assets: object[], content: object[]}} problems what verifyBuild found
|
|
398
|
+
* @param {Object<string,string>} sources built page path → source file an author edits
|
|
399
|
+
* @returns {Promise<void>}
|
|
400
|
+
*/
|
|
401
|
+
async function writeBuildReport( problems, sources ) {
|
|
402
|
+
try {
|
|
403
|
+
const { writeFile, mkdir } = await import( "node:fs/promises" );
|
|
404
|
+
const dir = join( process.cwd(), BUILD_REPORT_DIR );
|
|
405
|
+
|
|
406
|
+
await mkdir( dir, { recursive: true } );
|
|
407
|
+
await writeFile(
|
|
408
|
+
join( dir, BUILD_REPORT_FILE ),
|
|
409
|
+
renderBuildReport( problems, {
|
|
410
|
+
adapter: { name: "scavold", version: SCAVOLD_VERSION },
|
|
411
|
+
sources,
|
|
412
|
+
} ),
|
|
413
|
+
"utf-8",
|
|
414
|
+
);
|
|
415
|
+
} catch ( error ) {
|
|
416
|
+
console.warn( `[scavold] could not write ${BUILD_REPORT_DIR}/${BUILD_REPORT_FILE} —`, error.message );
|
|
417
|
+
}
|
|
418
|
+
}
|
|
419
|
+
|
|
420
|
+
/**
|
|
421
|
+
* Reads the finished build and refuses one that references files it does not
|
|
422
|
+
* contain — the shape a framework bug takes when it does not announce itself.
|
|
423
|
+
*
|
|
424
|
+
* A site can turn the content half off (`verify: { content: false }`) or the whole
|
|
425
|
+
* check (`verify: false`) when it links to paths its web server provides rather
|
|
426
|
+
* than its build.
|
|
427
|
+
*
|
|
428
|
+
* @param {object} siteConfig VitePress site configuration, as handed to buildEnd
|
|
429
|
+
* @returns {Promise<void>}
|
|
430
|
+
*/
|
|
431
|
+
async function verifyBuiltSite( siteConfig ) {
|
|
432
|
+
if ( options.verify === false ) {
|
|
433
|
+
return;
|
|
434
|
+
}
|
|
435
|
+
|
|
436
|
+
const { content = true } = options.verify === true || options.verify == null ? {} : options.verify;
|
|
437
|
+
|
|
438
|
+
// What VitePress said it would write, in the spelling it serves it under —
|
|
439
|
+
// and, the other way round, which file an author would open to fix a page.
|
|
440
|
+
const sources = {};
|
|
441
|
+
|
|
442
|
+
for ( const page of siteConfig.pages ?? [] ) {
|
|
443
|
+
sources[( siteConfig.rewrites?.map?.[page] || page ).replace( /\.md$/, ".html" )] = page;
|
|
444
|
+
}
|
|
445
|
+
|
|
446
|
+
const problems = verifyBuild( siteConfig.outDir, {
|
|
447
|
+
base: siteConfig.site?.base ?? "/",
|
|
448
|
+
content,
|
|
449
|
+
pages: Object.keys( sources ),
|
|
450
|
+
} );
|
|
451
|
+
|
|
452
|
+
// Written before anything is thrown: the failure this check exists to cause
|
|
453
|
+
// would otherwise take the reason with it, leaving a red pipeline and a log
|
|
454
|
+
// only some people may read.
|
|
455
|
+
await writeBuildReport( problems, sources );
|
|
456
|
+
|
|
457
|
+
if ( problems.content.length ) {
|
|
458
|
+
console.warn(
|
|
459
|
+
`[scavold] ${problems.content.length} link(s) or image(s) point at files this build does not contain:\n` +
|
|
460
|
+
`${formatProblems( problems.content )}`
|
|
461
|
+
);
|
|
462
|
+
}
|
|
463
|
+
|
|
464
|
+
if ( problems.pages.length ) {
|
|
465
|
+
throw new Error(
|
|
466
|
+
`[scavold] ${problems.pages.length} page(s) the build listed were never written:\n` +
|
|
467
|
+
`${formatProblems( problems.pages )}\n` +
|
|
468
|
+
"Two pages sharing one build identity are the known cause — on a file system that ignores letter " +
|
|
469
|
+
"case, this is the only trace they leave."
|
|
470
|
+
);
|
|
471
|
+
}
|
|
472
|
+
|
|
473
|
+
if ( problems.assets.length ) {
|
|
474
|
+
throw new Error(
|
|
475
|
+
`[scavold] ${problems.assets.length} script(s) or stylesheet(s) named by a built page are missing from the ` +
|
|
476
|
+
`output:\n${formatProblems( problems.assets )}\n` +
|
|
477
|
+
"Pages in that state render but stay dead in the browser. This is a bundling fault, not a content " +
|
|
478
|
+
"mistake — two pages sharing one build identity are the known cause."
|
|
479
|
+
);
|
|
480
|
+
}
|
|
481
|
+
}
|
|
482
|
+
|
|
379
483
|
const pagesDir = sourceFolder( resolvedConfig );
|
|
380
484
|
|
|
381
485
|
// Emit the section-type manifest (.cratly/sections.json) so the cratly
|
|
@@ -405,9 +509,87 @@ export async function augmentConfig( rawConfig, options = {} ) {
|
|
|
405
509
|
},
|
|
406
510
|
};
|
|
407
511
|
|
|
512
|
+
// ─── Page identity guard ─────────────────────────────────────────────────
|
|
513
|
+
// VitePress identifies a page by its path flattened into a single name, and
|
|
514
|
+
// two pages colliding there quietly overwrite each other — see pageKeys.js.
|
|
515
|
+
// Checked here, on every config load, so dev server and build both report the
|
|
516
|
+
// pair by name instead of dying later on a chunk that was never emitted.
|
|
517
|
+
|
|
518
|
+
/**
|
|
519
|
+
* Lists the page files VitePress will pick up, relative to srcDir.
|
|
520
|
+
*
|
|
521
|
+
* `srcExclude` is honoured for plain paths and for prefix patterns ending in
|
|
522
|
+
* `/*` or `/**` — the shapes a site actually writes. A more exotic pattern
|
|
523
|
+
* just leaves its file in the list, which at worst reports a collision the
|
|
524
|
+
* site had excluded anyway.
|
|
525
|
+
*
|
|
526
|
+
* @returns {Promise<string[]>} page paths, posix-style, relative to srcDir
|
|
527
|
+
*/
|
|
528
|
+
async function listPageFiles() {
|
|
529
|
+
const { glob } = await import( "node:fs/promises" );
|
|
530
|
+
const excluded = ( resolvedConfig.srcExclude ?? [] ).map( pattern => pattern.replace( /\/\*{1,2}$/, "/" ) );
|
|
531
|
+
|
|
532
|
+
return ( await Array.fromAsync( glob( "**/*.md", {
|
|
533
|
+
cwd: sourceFolder( resolvedConfig ),
|
|
534
|
+
exclude: name => name === "node_modules" || name === "dist",
|
|
535
|
+
} ) ) )
|
|
536
|
+
.map( name => name.replace( /\\/g, "/" ) )
|
|
537
|
+
// Dynamic routes stand for pages named by their paths file, not by
|
|
538
|
+
// their own name — their identity is not this file's path.
|
|
539
|
+
.filter( name => !name.includes( "[" ) )
|
|
540
|
+
.filter( name => !excluded.some( pattern => ( pattern.endsWith( "/" ) ? name.startsWith( pattern ) : name === pattern ) ) );
|
|
541
|
+
}
|
|
542
|
+
|
|
543
|
+
const rewrites = {
|
|
544
|
+
...rawConfig?.rewrites,
|
|
545
|
+
...await compileRedirects( resolvedConfig ),
|
|
546
|
+
};
|
|
547
|
+
|
|
548
|
+
assertUniquePageKeys( await listPageFiles(), rewrites );
|
|
549
|
+
|
|
550
|
+
// ─── Former addresses ────────────────────────────────────────────────────
|
|
551
|
+
// Compiled here so a declaration a server could not honour is refused while the
|
|
552
|
+
// author is still looking at it, and written out at the end of the build for
|
|
553
|
+
// whoever configures the server. What Scavold produces is data, not server
|
|
554
|
+
// configuration: cratly does not assume who serves the site.
|
|
555
|
+
const legacyRedirects = compileLegacyRedirects(
|
|
556
|
+
await collectFrontmatterOfAllPages( resolvedConfig ),
|
|
557
|
+
{ rewrites, retired: declaredRetiredUrls },
|
|
558
|
+
);
|
|
559
|
+
|
|
560
|
+
assertLegacyRedirects( legacyRedirects.problems );
|
|
561
|
+
|
|
562
|
+
/**
|
|
563
|
+
* Writes the redirect manifest, unless the site declares no former addresses.
|
|
564
|
+
*
|
|
565
|
+
* @param {object} siteConfig VitePress site configuration, as handed to buildEnd
|
|
566
|
+
* @returns {Promise<void>}
|
|
567
|
+
*/
|
|
568
|
+
async function writeRedirectManifest( siteConfig ) {
|
|
569
|
+
if ( !legacyRedirects.rules.length ) {
|
|
570
|
+
return;
|
|
571
|
+
}
|
|
572
|
+
|
|
573
|
+
const { writeFile } = await import( "node:fs/promises" );
|
|
574
|
+
|
|
575
|
+
await writeFile(
|
|
576
|
+
join( siteConfig.outDir, REDIRECT_MANIFEST_FILE ),
|
|
577
|
+
renderRedirectManifest( legacyRedirects.rules, { name: "scavold", version: SCAVOLD_VERSION } ),
|
|
578
|
+
"utf-8",
|
|
579
|
+
);
|
|
580
|
+
}
|
|
581
|
+
|
|
408
582
|
return {
|
|
409
583
|
...resolvedConfig,
|
|
410
584
|
|
|
585
|
+
// VitePress ends a build on a dead link, and ends it before writing anything —
|
|
586
|
+
// so the reason lives in a job log, which whoever wrote the link often cannot
|
|
587
|
+
// read, and a single typo stands between a finished page and its publication.
|
|
588
|
+
// Scavold finds the same mistake in the finished output (see verifyBuiltSite)
|
|
589
|
+
// and reports it there, where it can be handed to the author. A site that
|
|
590
|
+
// wants the abort back asks for it: ignoreDeadLinks: false.
|
|
591
|
+
ignoreDeadLinks: rawConfig.ignoreDeadLinks ?? true,
|
|
592
|
+
|
|
411
593
|
// Without this, VitePress writes its page-to-hash map into an inline script whose
|
|
412
594
|
// content changes with every build. Under a content security policy that names
|
|
413
595
|
// script hashes, that one hash has to be updated on the server for every deploy —
|
|
@@ -440,13 +622,14 @@ export async function augmentConfig( rawConfig, options = {} ) {
|
|
|
440
622
|
createExternalRedirectMapPlugin( pagesDir ),
|
|
441
623
|
],
|
|
442
624
|
},
|
|
443
|
-
rewrites
|
|
444
|
-
...rawConfig?.rewrites,
|
|
445
|
-
...await compileRedirects( resolvedConfig ),
|
|
446
|
-
},
|
|
625
|
+
rewrites,
|
|
447
626
|
async buildEnd( siteConfig ) {
|
|
448
627
|
await writeFeed( siteConfig );
|
|
628
|
+
await writeRedirectManifest( siteConfig );
|
|
449
629
|
await rawConfig.buildEnd?.( siteConfig );
|
|
630
|
+
|
|
631
|
+
// Last, so it sees what every other hand has written.
|
|
632
|
+
await verifyBuiltSite( siteConfig );
|
|
450
633
|
},
|
|
451
634
|
async transformPageData( pageData, context ) {
|
|
452
635
|
clearCache();
|
|
@@ -470,7 +653,7 @@ export async function augmentConfig( rawConfig, options = {} ) {
|
|
|
470
653
|
// during in-app use.
|
|
471
654
|
//
|
|
472
655
|
// The redirect value can be either:
|
|
473
|
-
// - a site-relative URL ("/de/jobs") — written by the
|
|
656
|
+
// - a site-relative URL ("/de/jobs") — written by the cratly editor
|
|
474
657
|
// - a relative .md path ("jobs.md") — written by hand in frontmatter
|
|
475
658
|
// - an absolute external URL — e.g. "https://example.com"
|
|
476
659
|
const redirect = pageData.frontmatter?.redirect;
|
|
@@ -531,14 +714,31 @@ export async function augmentConfig( rawConfig, options = {} ) {
|
|
|
531
714
|
// offers, which is relative to that folder. Rewrite it to the URL the
|
|
532
715
|
// built site serves. Targets that are not media files (ordinary page
|
|
533
716
|
// links, anchors, remote URLs) resolve to null and stay untouched.
|
|
534
|
-
|
|
535
|
-
|
|
717
|
+
//
|
|
718
|
+
// VitePress then takes any extension missing from its own list of file
|
|
719
|
+
// types (`.diff`, say) for part of a page name and appends `.html`. The
|
|
720
|
+
// published URL is final, so that suffix is taken back — after VitePress
|
|
721
|
+
// ran, which leaves prefixing the site's base to it.
|
|
722
|
+
patchRenderer( md, "link_open", ( defaultHandler, tokens, idx, mdOptions, env, slf ) => {
|
|
723
|
+
const token = tokens[idx];
|
|
724
|
+
const published = media.collectAsset( token.attrGet( "href" ) );
|
|
725
|
+
|
|
726
|
+
if ( !published ) {
|
|
727
|
+
return defaultHandler( tokens, idx, mdOptions, env, slf );
|
|
728
|
+
}
|
|
729
|
+
|
|
730
|
+
token.attrSet( "href", published );
|
|
731
|
+
|
|
732
|
+
const rendered = defaultHandler( tokens, idx, mdOptions, env, slf );
|
|
733
|
+
const href = token.attrGet( "href" );
|
|
734
|
+
|
|
735
|
+
if ( !published.endsWith( ".html" ) && href?.endsWith( published + ".html" ) ) {
|
|
736
|
+
token.attrSet( "href", href.slice( 0, -".html".length ) );
|
|
536
737
|
|
|
537
|
-
|
|
538
|
-
tokens[idx].attrSet( "href", published );
|
|
738
|
+
return slf.renderToken( tokens, idx, mdOptions );
|
|
539
739
|
}
|
|
540
740
|
|
|
541
|
-
return
|
|
741
|
+
return rendered;
|
|
542
742
|
} );
|
|
543
743
|
|
|
544
744
|
registerContainers( md, containerMap, {
|
package/lib/href.d.ts
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
import type { Scavold } from "../index";
|
|
2
|
+
|
|
3
|
+
/** Extension every page is built into and therefore linked with. */
|
|
4
|
+
export const PAGE_EXTENSION: string;
|
|
5
|
+
|
|
6
|
+
/**
|
|
7
|
+
* Converts a page's source path into the root-relative URL the built site serves.
|
|
8
|
+
*/
|
|
9
|
+
export function pageHref( path: string ): string;
|
|
10
|
+
|
|
11
|
+
/**
|
|
12
|
+
* Converts a hierarchy node into the URL the built site serves it under, preferring
|
|
13
|
+
* the alias URL when the page declares one.
|
|
14
|
+
*/
|
|
15
|
+
export function nodePath( node: Scavold.HierarchyNode | undefined ): string;
|
|
16
|
+
|
|
17
|
+
/**
|
|
18
|
+
* Rewrites a site-relative URL to the file that answers it: a folder URL to the index
|
|
19
|
+
* page inside it, an extension-less path to its page.
|
|
20
|
+
*/
|
|
21
|
+
export function servableUrl( url: string ): string;
|
|
@@ -0,0 +1,179 @@
|
|
|
1
|
+
// ─── Legacy URLs ─────────────────────────────────────────────────────────────
|
|
2
|
+
// A site that replaces an existing one inherits its addresses. Every one of them
|
|
3
|
+
// that a search engine knows and the new site does not answer is a result the
|
|
4
|
+
// customer loses, and no build notices — the old URL simply is not there any more.
|
|
5
|
+
//
|
|
6
|
+
// Only the web server can answer an address that is not a file, and only a
|
|
7
|
+
// permanent redirect passes the ranking on; a redirect run by the browser is
|
|
8
|
+
// followed but counts for less. So the pages declare which old addresses they
|
|
9
|
+
// replace, and the build writes them out as data. What turns that data into
|
|
10
|
+
// server configuration is the deploy step of whoever serves the site — cratly
|
|
11
|
+
// does not assume any particular server, so this file knows about none.
|
|
12
|
+
//
|
|
13
|
+
// Where an old address can simply be kept, `url` in front matter is the better
|
|
14
|
+
// answer: a page served at the legacy path needs no redirect at all, and works on
|
|
15
|
+
// every static server there is.
|
|
16
|
+
|
|
17
|
+
import { pageKey } from "./pageKeys.js";
|
|
18
|
+
|
|
19
|
+
// Version of the cratly redirect-manifest grammar this writes, and its canonical
|
|
20
|
+
// URL — cratly owns the format, Scavold is one adapter producing it.
|
|
21
|
+
export const REDIRECT_MANIFEST_VERSION = 0;
|
|
22
|
+
export const REDIRECT_MANIFEST_SCHEMA = "https://cratly.io/schema/redirects/v0.json";
|
|
23
|
+
|
|
24
|
+
// Name of the file the build writes into its output folder.
|
|
25
|
+
export const REDIRECT_MANIFEST_FILE = "cratly-redirects.json";
|
|
26
|
+
|
|
27
|
+
/**
|
|
28
|
+
* Turns a page's source path into the URL its build is served under.
|
|
29
|
+
*
|
|
30
|
+
* @param {string} page path of the page's markdown file, relative to srcDir
|
|
31
|
+
* @param {Object<string,string>} [rewrites] VitePress `rewrites` map, source path → served path
|
|
32
|
+
* @returns {string} site-absolute URL of the built page
|
|
33
|
+
*/
|
|
34
|
+
export function servedPathOf( page, rewrites = {} ) {
|
|
35
|
+
const served = rewrites?.[page] || page;
|
|
36
|
+
|
|
37
|
+
return `/${served.replace( /\.md$/, ".html" )}`;
|
|
38
|
+
}
|
|
39
|
+
|
|
40
|
+
/**
|
|
41
|
+
* Reads a declared legacy address into the shape a redirect rule uses.
|
|
42
|
+
*
|
|
43
|
+
* Authors write these as they appeared in the browser, so a leading slash may be
|
|
44
|
+
* missing and a fragment may be attached — a fragment never reaches the server and
|
|
45
|
+
* is dropped. A query string is kept: whether it can be matched is the consuming
|
|
46
|
+
* server's business, not this format's.
|
|
47
|
+
*
|
|
48
|
+
* @param {string} value legacy address as declared
|
|
49
|
+
* @returns {string|null} site-absolute address, or null when the value says nothing
|
|
50
|
+
*/
|
|
51
|
+
export function normaliseLegacyPath( value ) {
|
|
52
|
+
if ( typeof value !== "string" ) {
|
|
53
|
+
return null;
|
|
54
|
+
}
|
|
55
|
+
|
|
56
|
+
const [withoutFragment] = value.trim().split( "#" );
|
|
57
|
+
|
|
58
|
+
if ( !withoutFragment ) {
|
|
59
|
+
return null;
|
|
60
|
+
}
|
|
61
|
+
|
|
62
|
+
return withoutFragment.startsWith( "/" ) ? withoutFragment : `/${withoutFragment}`;
|
|
63
|
+
}
|
|
64
|
+
|
|
65
|
+
/**
|
|
66
|
+
* Compiles the redirect rules a site's own declarations amount to.
|
|
67
|
+
*
|
|
68
|
+
* Two mistakes are worth stopping a build over, and both are reported rather than
|
|
69
|
+
* resolved: an old address that two pages claim, and one that collides with an
|
|
70
|
+
* address the new site itself serves. The latter is the interesting one — a
|
|
71
|
+
* redirect would shadow a real page, so the page would be unreachable while every
|
|
72
|
+
* link to it still points there.
|
|
73
|
+
*
|
|
74
|
+
* @param {Object<string,object>} frontmatter map of page path → front matter
|
|
75
|
+
* @param {object} [options] what the site declares elsewhere
|
|
76
|
+
* @param {Object<string,string>} [options.rewrites] VitePress `rewrites` map
|
|
77
|
+
* @param {string[]} [options.retired] addresses declared as gone for good
|
|
78
|
+
* @returns {{rules: {from: string, to?: string, status: number}[], problems: string[]}}
|
|
79
|
+
* rules in address order, and what a site has to fix first
|
|
80
|
+
*/
|
|
81
|
+
export function compileLegacyRedirects( frontmatter, { rewrites = {}, retired = [] } = {} ) {
|
|
82
|
+
const problems = [];
|
|
83
|
+
const claimedBy = new Map(); // legacy address → page claiming it
|
|
84
|
+
const rules = new Map(); // legacy address → rule
|
|
85
|
+
|
|
86
|
+
// Addresses the new site answers itself. A redirect must never cover one.
|
|
87
|
+
const served = new Map();
|
|
88
|
+
|
|
89
|
+
for ( const page of Object.keys( frontmatter ) ) {
|
|
90
|
+
served.set( pageKey( servedPathOf( page, rewrites ).slice( 1 ) ), page );
|
|
91
|
+
}
|
|
92
|
+
|
|
93
|
+
for ( const [ page, meta ] of Object.entries( frontmatter ) ) {
|
|
94
|
+
const declared = meta?.aliases == null ? [] : [meta.aliases].flat();
|
|
95
|
+
|
|
96
|
+
for ( const value of declared ) {
|
|
97
|
+
const from = normaliseLegacyPath( value );
|
|
98
|
+
|
|
99
|
+
if ( !from ) {
|
|
100
|
+
continue;
|
|
101
|
+
}
|
|
102
|
+
|
|
103
|
+
const owner = claimedBy.get( from );
|
|
104
|
+
|
|
105
|
+
if ( owner ) {
|
|
106
|
+
problems.push( `both "${owner}" and "${page}" declare the former address ${from}` );
|
|
107
|
+
continue;
|
|
108
|
+
}
|
|
109
|
+
|
|
110
|
+
const shadowed = served.get( pageKey( from.slice( 1 ) ) );
|
|
111
|
+
|
|
112
|
+
if ( shadowed ) {
|
|
113
|
+
problems.push(
|
|
114
|
+
`"${page}" declares the former address ${from}, which "${shadowed}" is served at — ` +
|
|
115
|
+
"a redirect there would hide that page"
|
|
116
|
+
);
|
|
117
|
+
continue;
|
|
118
|
+
}
|
|
119
|
+
|
|
120
|
+
claimedBy.set( from, page );
|
|
121
|
+
rules.set( from, { from, to: servedPathOf( page, rewrites ), status: 301 } );
|
|
122
|
+
}
|
|
123
|
+
}
|
|
124
|
+
|
|
125
|
+
for ( const value of retired ) {
|
|
126
|
+
const from = normaliseLegacyPath( value );
|
|
127
|
+
|
|
128
|
+
if ( !from ) {
|
|
129
|
+
continue;
|
|
130
|
+
}
|
|
131
|
+
|
|
132
|
+
if ( rules.has( from ) ) {
|
|
133
|
+
problems.push( `${from} is declared as retired and at the same time as a former address of "${claimedBy.get( from )}"` );
|
|
134
|
+
continue;
|
|
135
|
+
}
|
|
136
|
+
|
|
137
|
+
rules.set( from, { from, status: 410 } );
|
|
138
|
+
}
|
|
139
|
+
|
|
140
|
+
return {
|
|
141
|
+
rules: [...rules.values()].sort( ( a, b ) => ( a.from < b.from ? -1 : 1 ) ),
|
|
142
|
+
problems,
|
|
143
|
+
};
|
|
144
|
+
}
|
|
145
|
+
|
|
146
|
+
/**
|
|
147
|
+
* Refuses a set of declarations a server could not serve as meant.
|
|
148
|
+
*
|
|
149
|
+
* @param {string[]} problems what compileLegacyRedirects found
|
|
150
|
+
* @returns {void}
|
|
151
|
+
* @throws {Error} naming every problem
|
|
152
|
+
*/
|
|
153
|
+
export function assertLegacyRedirects( problems ) {
|
|
154
|
+
if ( problems.length ) {
|
|
155
|
+
throw new Error( `[scavold] former addresses cannot be served as declared: ${problems.join( "; " )}` );
|
|
156
|
+
}
|
|
157
|
+
}
|
|
158
|
+
|
|
159
|
+
/**
|
|
160
|
+
* Renders the redirect manifest, the artifact a deploy step reads.
|
|
161
|
+
*
|
|
162
|
+
* Deliberately free of any server's vocabulary: an address, where it went, and the
|
|
163
|
+
* status that says so. Translating that into a Caddyfile, an nginx map, an
|
|
164
|
+
* `.s3-http.config.yaml` or a bucket's routing rules is the deploy's job.
|
|
165
|
+
*
|
|
166
|
+
* @param {{from: string, to?: string, status: number}[]} rules compiled rules
|
|
167
|
+
* @param {object} [adapter] what produced the file
|
|
168
|
+
* @param {string} [adapter.name] name of the adapter
|
|
169
|
+
* @param {string} [adapter.version] its version
|
|
170
|
+
* @returns {string} the manifest as JSON text, newline-terminated
|
|
171
|
+
*/
|
|
172
|
+
export function renderRedirectManifest( rules, adapter = {} ) {
|
|
173
|
+
return `${JSON.stringify( {
|
|
174
|
+
$schema: REDIRECT_MANIFEST_SCHEMA,
|
|
175
|
+
specVersion: REDIRECT_MANIFEST_VERSION,
|
|
176
|
+
adapter,
|
|
177
|
+
rules,
|
|
178
|
+
}, null, "\t" )}\n`;
|
|
179
|
+
}
|
package/lib/pageKeys.js
ADDED
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
// ─── Page identity guard ─────────────────────────────────────────────────────
|
|
2
|
+
// VitePress does not carry a page's path through its build. It flattens the path
|
|
3
|
+
// into a single name and uses that name as the page's identity in four places:
|
|
4
|
+
//
|
|
5
|
+
// * the key of the page's entry in `rollupOptions.input` (the rewritten path
|
|
6
|
+
// wins here — `config.rewrites.map[file] || file`),
|
|
7
|
+
// * the file name of the page's SSR module in the temp folder,
|
|
8
|
+
// * the key of the page's entry in `pageToHashMap` — lower-cased,
|
|
9
|
+
// * the name of the page's client chunk, hence its `.lean.js` file name.
|
|
10
|
+
//
|
|
11
|
+
// Two pages whose paths flatten to the same name therefore share one identity
|
|
12
|
+
// and one set of output files. VitePress neither detects nor reports this:
|
|
13
|
+
// depending on which of the four collides, the build either dies at render time
|
|
14
|
+
// with `undefined is not an object (evaluating 'pageChunk.imports')`, naming no
|
|
15
|
+
// file, or it succeeds and ships two pages pointing at one script file that was
|
|
16
|
+
// never written.
|
|
17
|
+
//
|
|
18
|
+
// The guard below reproduces that identity and demands it stays unique, so the
|
|
19
|
+
// site learns which two files disagree instead of what VitePress makes of them.
|
|
20
|
+
|
|
21
|
+
// Copied from VitePress (which copies it from Vite) — the characters its
|
|
22
|
+
// `sanitizeFileName()` replaces with an underscore.
|
|
23
|
+
|
|
24
|
+
const INVALID_CHAR_RE = /[\u0000-\u001F"#$&*+,:;<=>?[\]^`{|}\u007F]/g;
|
|
25
|
+
|
|
26
|
+
/**
|
|
27
|
+
* Reproduces the identity VitePress derives from a page path.
|
|
28
|
+
*
|
|
29
|
+
* Lower-casing goes beyond what the bundle entry key does — it is what
|
|
30
|
+
* `pageToHashMap` keys on, and a pair colliding only there is the worse of the
|
|
31
|
+
* two failures, so the strictest of the four rules is the one to enforce.
|
|
32
|
+
*
|
|
33
|
+
* @param {string} page path of a page's markdown file, relative to srcDir
|
|
34
|
+
* @returns {string} the name VitePress identifies that page by
|
|
35
|
+
*/
|
|
36
|
+
export function pageKey( page ) {
|
|
37
|
+
return String( page )
|
|
38
|
+
.replace( /\\/g, "/" )
|
|
39
|
+
.replace( /\//g, "_" )
|
|
40
|
+
.replace( INVALID_CHAR_RE, "_" )
|
|
41
|
+
// VitePress strips underscores that a leading invalid character produced.
|
|
42
|
+
.replace( /^_+/, "" )
|
|
43
|
+
.toLowerCase();
|
|
44
|
+
}
|
|
45
|
+
|
|
46
|
+
/**
|
|
47
|
+
* Groups pages that VitePress cannot tell apart.
|
|
48
|
+
*
|
|
49
|
+
* @param {string[]} pages paths of all page files, relative to srcDir
|
|
50
|
+
* @param {Object<string,string>} [rewrites] VitePress `rewrites` map, source path → served path
|
|
51
|
+
* @returns {{key: string, sources: {page: string, alias: string}[]}[]} one entry per colliding identity, in path order
|
|
52
|
+
*/
|
|
53
|
+
export function findPageKeyCollisions( pages, rewrites = {} ) {
|
|
54
|
+
const byKey = new Map();
|
|
55
|
+
|
|
56
|
+
for ( const page of [...pages].sort() ) {
|
|
57
|
+
const alias = rewrites?.[page] || page;
|
|
58
|
+
const key = pageKey( alias );
|
|
59
|
+
const group = byKey.get( key );
|
|
60
|
+
|
|
61
|
+
if ( group ) {
|
|
62
|
+
group.push( { page, alias } );
|
|
63
|
+
} else {
|
|
64
|
+
byKey.set( key, [{ page, alias }] );
|
|
65
|
+
}
|
|
66
|
+
}
|
|
67
|
+
|
|
68
|
+
return [...byKey.entries()]
|
|
69
|
+
.filter( ( [ , sources ] ) => sources.length > 1 )
|
|
70
|
+
.map( ( [ key, sources ] ) => ( { key, sources } ) );
|
|
71
|
+
}
|
|
72
|
+
|
|
73
|
+
/**
|
|
74
|
+
* Renders one collision as a line naming both files and, where a rewrite or a
|
|
75
|
+
* `url` alias is what makes them collide, the path each is served at.
|
|
76
|
+
*
|
|
77
|
+
* @param {{key: string, sources: {page: string, alias: string}[]}} collision
|
|
78
|
+
* @returns {string} single-line description
|
|
79
|
+
*/
|
|
80
|
+
function describeCollision( { key, sources } ) {
|
|
81
|
+
const named = sources
|
|
82
|
+
.map( ( { page, alias } ) => ( alias === page ? `"${page}"` : `"${page}" (served as "${alias}")` ) )
|
|
83
|
+
.join( " and " );
|
|
84
|
+
|
|
85
|
+
return `${named} share the build identity "${key}"`;
|
|
86
|
+
}
|
|
87
|
+
|
|
88
|
+
/**
|
|
89
|
+
* Throws when two pages flatten to the same VitePress identity.
|
|
90
|
+
*
|
|
91
|
+
* @param {string[]} pages paths of all page files, relative to srcDir
|
|
92
|
+
* @param {Object<string,string>} [rewrites] VitePress `rewrites` map, source path → served path
|
|
93
|
+
* @returns {void}
|
|
94
|
+
* @throws {Error} naming every colliding pair and how to resolve it
|
|
95
|
+
*/
|
|
96
|
+
export function assertUniquePageKeys( pages, rewrites = {} ) {
|
|
97
|
+
const collisions = findPageKeyCollisions( pages, rewrites );
|
|
98
|
+
|
|
99
|
+
if ( collisions.length === 0 ) {
|
|
100
|
+
return;
|
|
101
|
+
}
|
|
102
|
+
|
|
103
|
+
throw new Error(
|
|
104
|
+
`[scavold] colliding page paths: ${collisions.map( describeCollision ).join( "; " )}. ` +
|
|
105
|
+
"VitePress flattens a page path into a single name — slash becomes underscore — and " +
|
|
106
|
+
"identifies the page by it, so pages colliding there overwrite each other's bundle " +
|
|
107
|
+
"entry and output files. Rename one of them; prefer \"-\" over \"_\" as the word " +
|
|
108
|
+
"separator in page names, and keep an existing URL alive with front matter `url` " +
|
|
109
|
+
"instead of the old file name."
|
|
110
|
+
);
|
|
111
|
+
}
|
|
@@ -0,0 +1,324 @@
|
|
|
1
|
+
// ─── Build verification ──────────────────────────────────────────────────────
|
|
2
|
+
// A static site can be wrong in ways its generator does not notice. VitePress
|
|
3
|
+
// identifies a page by its path flattened into one name and reuses that name for
|
|
4
|
+
// the page's script files; where two pages meet there, one of them ends up
|
|
5
|
+
// preloading a file that was never written — and the build reports success. The
|
|
6
|
+
// same silence covers a renamed page still linked from elsewhere and an image
|
|
7
|
+
// whose file went away.
|
|
8
|
+
//
|
|
9
|
+
// None of that is visible in the build log, and all of it is visible in the
|
|
10
|
+
// finished output: the pages name the files they need, and either those files are
|
|
11
|
+
// there or they are not. So the last thing Scavold does is read its own result.
|
|
12
|
+
//
|
|
13
|
+
// Three classes, treated differently:
|
|
14
|
+
//
|
|
15
|
+
// * A page the generator listed but did not write is never intentional either —
|
|
16
|
+
// and it is the only symptom left when two pages collide on a file system that
|
|
17
|
+
// ignores letter case, which is what the machine writing the page usually has.
|
|
18
|
+
// * A missing script or stylesheet is never intentional. It means the framework
|
|
19
|
+
// produced an inconsistent bundle, and the page is dead in the browser while
|
|
20
|
+
// the pre-rendered HTML still shows. That fails the build.
|
|
21
|
+
// * A missing link or image target is a content mistake, and one a site may make
|
|
22
|
+
// on purpose — a link to a path the web server provides rather than the build.
|
|
23
|
+
// That is reported and left to the site to decide on.
|
|
24
|
+
//
|
|
25
|
+
// All three go into .cratly/build-report.json as well, written before the build is
|
|
26
|
+
// ended rather than after, so what happened survives the failure. A build log is
|
|
27
|
+
// read by whoever has access to it; a file can be handed to whoever wrote the page.
|
|
28
|
+
|
|
29
|
+
import { readFileSync, readdirSync } from "node:fs";
|
|
30
|
+
import { join, posix, resolve } from "node:path";
|
|
31
|
+
|
|
32
|
+
// Where the report goes, and which shape a reader may expect. The version is
|
|
33
|
+
// bumped when a reader would have to be changed — the cratly editor is one.
|
|
34
|
+
export const BUILD_REPORT_DIR = ".cratly";
|
|
35
|
+
export const BUILD_REPORT_FILE = "build-report.json";
|
|
36
|
+
export const BUILD_REPORT_VERSION = 0;
|
|
37
|
+
|
|
38
|
+
// Attributes that name a file the browser must load for the page to work at all.
|
|
39
|
+
const ASSET_ATTR_RE = /<(?:script|link)\b[^>]*?\b(?:src|href)\s*=\s*["']([^"']+)["'][^>]*>/gi;
|
|
40
|
+
const ASSET_TAG_RE = /^<(?:script\b[^>]*\bsrc|link\b[^>]*\brel\s*=\s*["'](?:modulepreload|preload|stylesheet)["'])/i;
|
|
41
|
+
const LINK_RE = /<a\b[^>]*?\bhref\s*=\s*["']([^"']+)["']/gi;
|
|
42
|
+
const IMG_SRC_RE = /<(?:img|source|video)\b[^>]*?\b(?:src|poster)\s*=\s*["']([^"']+)["']/gi;
|
|
43
|
+
const SRCSET_RE = /<(?:img|source)\b[^>]*?\bsrcset\s*=\s*["']([^"']+)["']/gi;
|
|
44
|
+
|
|
45
|
+
// A reference naming its own protocol addresses something outside this build:
|
|
46
|
+
// http:, mailto:, tel: and every other scheme.
|
|
47
|
+
const SCHEME_RE = /^[a-z][a-z0-9+.-]*:/i;
|
|
48
|
+
|
|
49
|
+
/**
|
|
50
|
+
* Tells references that leave the site from those the build has to satisfy.
|
|
51
|
+
*
|
|
52
|
+
* Relative references count. A markdown link that forgot its protocol —
|
|
53
|
+
* `[Website](./www.example.com)` — leaves the build as a relative href pointing
|
|
54
|
+
* at a sibling page nobody wrote, and that is the mistake authors actually make.
|
|
55
|
+
* Where such a reference points is a question only the page holding it can
|
|
56
|
+
* answer, which is why resolving is left to {@link resolveReference}.
|
|
57
|
+
*
|
|
58
|
+
* @param {string} reference value of an href, src or srcset entry
|
|
59
|
+
* @returns {boolean} true when the reference names something this build produced
|
|
60
|
+
*/
|
|
61
|
+
function isLocalReference( reference ) {
|
|
62
|
+
return reference !== ""
|
|
63
|
+
&& !reference.startsWith( "//" )
|
|
64
|
+
&& !SCHEME_RE.test( reference )
|
|
65
|
+
&& !reference.includes( "?" );
|
|
66
|
+
}
|
|
67
|
+
|
|
68
|
+
/**
|
|
69
|
+
* Collects the scripts and stylesheets a page cannot work without.
|
|
70
|
+
*
|
|
71
|
+
* @param {string} html contents of one built page
|
|
72
|
+
* @returns {string[]} referenced paths, in document order and de-duplicated
|
|
73
|
+
*/
|
|
74
|
+
export function collectAssetRefs( html ) {
|
|
75
|
+
const found = new Set();
|
|
76
|
+
|
|
77
|
+
for ( const [ tag, reference ] of html.matchAll( ASSET_ATTR_RE ) ) {
|
|
78
|
+
if ( ASSET_TAG_RE.test( tag ) && isLocalReference( reference ) ) {
|
|
79
|
+
found.add( reference );
|
|
80
|
+
}
|
|
81
|
+
}
|
|
82
|
+
|
|
83
|
+
return [...found];
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
/**
|
|
87
|
+
* Collects the pages and files a page links to.
|
|
88
|
+
*
|
|
89
|
+
* @param {string} html contents of one built page
|
|
90
|
+
* @returns {string[]} referenced paths, de-duplicated
|
|
91
|
+
*/
|
|
92
|
+
export function collectLinkRefs( html ) {
|
|
93
|
+
const found = new Set();
|
|
94
|
+
|
|
95
|
+
for ( const [ , reference ] of html.matchAll( LINK_RE ) ) {
|
|
96
|
+
const [path] = reference.split( "#" );
|
|
97
|
+
|
|
98
|
+
if ( path && isLocalReference( path ) ) {
|
|
99
|
+
found.add( path );
|
|
100
|
+
}
|
|
101
|
+
}
|
|
102
|
+
|
|
103
|
+
return [...found];
|
|
104
|
+
}
|
|
105
|
+
|
|
106
|
+
/**
|
|
107
|
+
* Collects the images and media a page displays.
|
|
108
|
+
*
|
|
109
|
+
* Every candidate of a `srcset` counts: a browser on a wide screen asks for the
|
|
110
|
+
* widest one, so a variant missing from the set breaks the page for some readers
|
|
111
|
+
* and not for others — the worst kind to leave unreported.
|
|
112
|
+
*
|
|
113
|
+
* @param {string} html contents of one built page
|
|
114
|
+
* @returns {string[]} referenced paths, de-duplicated
|
|
115
|
+
*/
|
|
116
|
+
export function collectImageRefs( html ) {
|
|
117
|
+
const found = new Set();
|
|
118
|
+
|
|
119
|
+
for ( const [ , reference ] of html.matchAll( IMG_SRC_RE ) ) {
|
|
120
|
+
if ( isLocalReference( reference ) ) {
|
|
121
|
+
found.add( reference );
|
|
122
|
+
}
|
|
123
|
+
}
|
|
124
|
+
|
|
125
|
+
for ( const [ , srcset ] of html.matchAll( SRCSET_RE ) ) {
|
|
126
|
+
for ( const candidate of srcset.split( "," ) ) {
|
|
127
|
+
const [reference] = candidate.trim().split( /\s+/ );
|
|
128
|
+
|
|
129
|
+
if ( reference && isLocalReference( reference ) ) {
|
|
130
|
+
found.add( reference );
|
|
131
|
+
}
|
|
132
|
+
}
|
|
133
|
+
}
|
|
134
|
+
|
|
135
|
+
return [...found];
|
|
136
|
+
}
|
|
137
|
+
|
|
138
|
+
/**
|
|
139
|
+
* Turns a reference as written into the path below the output folder it addresses.
|
|
140
|
+
*
|
|
141
|
+
* A site-absolute reference is read the way the web server reads it, with the
|
|
142
|
+
* site's base path stripped. A relative one means nothing on its own: the browser
|
|
143
|
+
* resolves it against the address of the page it stands in, so that is what this
|
|
144
|
+
* resolves it against too. One that climbs above the site root stays as it is —
|
|
145
|
+
* no file of this build can satisfy it, and saying so is the point.
|
|
146
|
+
*
|
|
147
|
+
* @param {string} reference reference as written in the HTML
|
|
148
|
+
* @param {string} page output path of the page holding it, relative and posix-style
|
|
149
|
+
* @param {string} base site's base path, as configured
|
|
150
|
+
* @returns {string} path relative to the output folder, posix-style and without a leading slash
|
|
151
|
+
*/
|
|
152
|
+
function resolveReference( reference, page, base ) {
|
|
153
|
+
const decoded = decodeURI( reference );
|
|
154
|
+
|
|
155
|
+
if ( decoded.startsWith( "/" ) ) {
|
|
156
|
+
const withoutBase = base !== "/" && decoded.startsWith( base )
|
|
157
|
+
? decoded.slice( base.length - 1 )
|
|
158
|
+
: decoded;
|
|
159
|
+
|
|
160
|
+
return withoutBase.replace( /^\//, "" );
|
|
161
|
+
}
|
|
162
|
+
|
|
163
|
+
return posix.normalize( posix.join( posix.dirname( page ), decoded ) ).replace( /^\//, "" );
|
|
164
|
+
}
|
|
165
|
+
|
|
166
|
+
/**
|
|
167
|
+
* Tells whether a build wrote something a static server could answer a reference with.
|
|
168
|
+
*
|
|
169
|
+
* The three spellings are the ones every static server tries, cratly's own
|
|
170
|
+
* `s3-http` included: the path itself, an `index.html` inside it, and the path
|
|
171
|
+
* plus `.html`.
|
|
172
|
+
*
|
|
173
|
+
* The comparison runs against the names the build actually wrote, never against
|
|
174
|
+
* the file system — macOS and Windows answer `existsSync` case-insensitively, so a
|
|
175
|
+
* page preloading `de_Kontakt.md.js` while only `de_kontakt.md.js` exists would
|
|
176
|
+
* look sound on the machine of whoever wrote the page and break in production.
|
|
177
|
+
* That pair is one of the two ways two pages end up sharing a build identity.
|
|
178
|
+
*
|
|
179
|
+
* @param {string} reference reference as written in the HTML
|
|
180
|
+
* @param {string} page output path of the page holding it, relative and posix-style
|
|
181
|
+
* @param {Set<string>} files paths of every file in the output, relative and posix-style
|
|
182
|
+
* @param {string} base site's base path, as configured
|
|
183
|
+
* @returns {boolean} true when one of those files was written
|
|
184
|
+
*/
|
|
185
|
+
function resolvesToFile( reference, page, files, base ) {
|
|
186
|
+
const relativePath = resolveReference( reference, page, base );
|
|
187
|
+
|
|
188
|
+
if ( relativePath === "" || relativePath === "." ) {
|
|
189
|
+
return files.has( "index.html" );
|
|
190
|
+
}
|
|
191
|
+
|
|
192
|
+
return [ relativePath, posix.join( relativePath, "index.html" ), `${relativePath}.html` ]
|
|
193
|
+
.some( candidate => files.has( candidate ) );
|
|
194
|
+
}
|
|
195
|
+
|
|
196
|
+
/**
|
|
197
|
+
* Lists everything a build wrote.
|
|
198
|
+
*
|
|
199
|
+
* @param {string} outDir absolute path of the build output folder
|
|
200
|
+
* @returns {Set<string>} every file below outDir, relative and posix-style
|
|
201
|
+
*/
|
|
202
|
+
function collectFiles( outDir ) {
|
|
203
|
+
const files = new Set();
|
|
204
|
+
|
|
205
|
+
/**
|
|
206
|
+
* @param {string} dir folder to descend into
|
|
207
|
+
* @param {string} prefix path of that folder relative to outDir
|
|
208
|
+
* @returns {void}
|
|
209
|
+
*/
|
|
210
|
+
function walk( dir, prefix ) {
|
|
211
|
+
for ( const entry of readdirSync( dir, { withFileTypes: true } ) ) {
|
|
212
|
+
const relativePath = prefix ? `${prefix}/${entry.name}` : entry.name;
|
|
213
|
+
|
|
214
|
+
if ( entry.isDirectory() ) {
|
|
215
|
+
walk( join( dir, entry.name ), relativePath );
|
|
216
|
+
} else {
|
|
217
|
+
files.add( relativePath );
|
|
218
|
+
}
|
|
219
|
+
}
|
|
220
|
+
}
|
|
221
|
+
|
|
222
|
+
walk( outDir, "" );
|
|
223
|
+
|
|
224
|
+
return files;
|
|
225
|
+
}
|
|
226
|
+
|
|
227
|
+
/**
|
|
228
|
+
* Reads a finished build and reports what it references but does not contain.
|
|
229
|
+
*
|
|
230
|
+
* @param {string} outDir absolute path of the build output folder
|
|
231
|
+
* @param {object} [options] what to look at
|
|
232
|
+
* @param {string} [options.base] site's base path
|
|
233
|
+
* @param {boolean} [options.content] also check links and images, not just scripts
|
|
234
|
+
* @param {string[]} [options.pages] output files the generator said it would write, relative and posix-style
|
|
235
|
+
* @returns {{pages: object[], assets: object[], content: object[]}}
|
|
236
|
+
* what the output does not contain, split by how much it matters
|
|
237
|
+
*/
|
|
238
|
+
export function verifyBuild( outDir, { base = "/", content = true, pages = [] } = {} ) {
|
|
239
|
+
const root = resolve( outDir );
|
|
240
|
+
const files = collectFiles( root );
|
|
241
|
+
const problems = { pages: [], assets: [], content: [] };
|
|
242
|
+
|
|
243
|
+
for ( const page of pages ) {
|
|
244
|
+
if ( !files.has( page ) ) {
|
|
245
|
+
problems.pages.push( { page, reference: page } );
|
|
246
|
+
}
|
|
247
|
+
}
|
|
248
|
+
|
|
249
|
+
for ( const page of [...files].filter( file => file.endsWith( ".html" ) ).sort() ) {
|
|
250
|
+
const html = readFileSync( join( root, page ), "utf-8" );
|
|
251
|
+
|
|
252
|
+
for ( const reference of collectAssetRefs( html ) ) {
|
|
253
|
+
if ( !resolvesToFile( reference, page, files, base ) ) {
|
|
254
|
+
problems.assets.push( { page, reference } );
|
|
255
|
+
}
|
|
256
|
+
}
|
|
257
|
+
|
|
258
|
+
if ( content ) {
|
|
259
|
+
for ( const reference of [ ...collectLinkRefs( html ), ...collectImageRefs( html ) ] ) {
|
|
260
|
+
if ( !resolvesToFile( reference, page, files, base ) ) {
|
|
261
|
+
problems.content.push( { page, reference } );
|
|
262
|
+
}
|
|
263
|
+
}
|
|
264
|
+
}
|
|
265
|
+
}
|
|
266
|
+
|
|
267
|
+
return problems;
|
|
268
|
+
}
|
|
269
|
+
|
|
270
|
+
/**
|
|
271
|
+
* Renders a set of problems as one line per reference, capped so a systematic
|
|
272
|
+
* mistake does not bury the terminal.
|
|
273
|
+
*
|
|
274
|
+
* @param {{page: string, reference: string}[]} problems what verifyBuild found
|
|
275
|
+
* @param {number} [limit] how many to name
|
|
276
|
+
* @returns {string} indented, newline-separated list
|
|
277
|
+
*/
|
|
278
|
+
export function formatProblems( problems, limit = 10 ) {
|
|
279
|
+
const lines = problems.slice( 0, limit ).map( ( { page, reference } ) => ` ${page} → ${reference}` );
|
|
280
|
+
|
|
281
|
+
if ( problems.length > limit ) {
|
|
282
|
+
lines.push( ` … and ${problems.length - limit} more` );
|
|
283
|
+
}
|
|
284
|
+
|
|
285
|
+
return lines.join( "\n" );
|
|
286
|
+
}
|
|
287
|
+
|
|
288
|
+
/**
|
|
289
|
+
* Renders what a build found as the report file readers other than a terminal get.
|
|
290
|
+
*
|
|
291
|
+
* `fatal` says which of the two situations the file describes: a build that ended
|
|
292
|
+
* without writing a usable site, or one that finished and carries a content
|
|
293
|
+
* mistake. A reader showing this to whoever wrote the page needs that distinction
|
|
294
|
+
* and should not have to re-derive it from which array is empty.
|
|
295
|
+
*
|
|
296
|
+
* @param {{pages: object[], assets: object[], content: object[]}} problems what verifyBuild found
|
|
297
|
+
* @param {object} [options]
|
|
298
|
+
* @param {object} [options.adapter] name and version of what produced the build
|
|
299
|
+
* @param {Object<string,string>} [options.sources] built page path → source file an author edits
|
|
300
|
+
* @param {Date} [options.generated] when the build ran
|
|
301
|
+
* @returns {string} the file's content, newline-terminated
|
|
302
|
+
*/
|
|
303
|
+
export function renderBuildReport( problems, { adapter = {}, sources = {}, generated = new Date() } = {} ) {
|
|
304
|
+
/**
|
|
305
|
+
* @param {{page: string, reference: string}[]} entries one class of problem
|
|
306
|
+
* @returns {object[]} the same entries, each naming its source file where known
|
|
307
|
+
*/
|
|
308
|
+
const withSources = entries => entries.map( entry => ( {
|
|
309
|
+
...entry,
|
|
310
|
+
...sources[entry.page] != null && { source: sources[entry.page] },
|
|
311
|
+
} ) );
|
|
312
|
+
|
|
313
|
+
return `${JSON.stringify( {
|
|
314
|
+
specVersion: BUILD_REPORT_VERSION,
|
|
315
|
+
adapter,
|
|
316
|
+
generated: generated.toISOString(),
|
|
317
|
+
fatal: problems.pages.length > 0 || problems.assets.length > 0,
|
|
318
|
+
problems: {
|
|
319
|
+
pages: withSources( problems.pages ),
|
|
320
|
+
assets: withSources( problems.assets ),
|
|
321
|
+
content: withSources( problems.content ),
|
|
322
|
+
},
|
|
323
|
+
}, null, "\t" )}\n`;
|
|
324
|
+
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "scavold",
|
|
3
|
-
"version": "0.2.0-rc.
|
|
3
|
+
"version": "0.2.0-rc.7",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "VitePress theme framework — a scaffold for building custom VitePress themes with Vue at the core",
|
|
6
6
|
"keywords": [
|
|
@@ -18,10 +18,10 @@
|
|
|
18
18
|
"license": "MIT",
|
|
19
19
|
"repository": {
|
|
20
20
|
"type": "git",
|
|
21
|
-
"url": "https://gitlab.com/
|
|
21
|
+
"url": "https://gitlab.com/cratly/scavold.git"
|
|
22
22
|
},
|
|
23
23
|
"homepage": "https://scavold.io",
|
|
24
|
-
"bugs": "https://gitlab.com/
|
|
24
|
+
"bugs": "https://gitlab.com/cratly/scavold/-/issues",
|
|
25
25
|
"engines": {
|
|
26
26
|
"node": ">=20"
|
|
27
27
|
},
|
|
@@ -72,7 +72,7 @@
|
|
|
72
72
|
"vue": ">=3.0.0"
|
|
73
73
|
},
|
|
74
74
|
"dependencies": {
|
|
75
|
-
"@cepharum/vue3-i18n": "^0.
|
|
75
|
+
"@cepharum/vue3-i18n": "^2.0.0",
|
|
76
76
|
"markdown-it-container": "^4.0.0",
|
|
77
77
|
"sharp": "^0.34.5",
|
|
78
78
|
"yaml": "^2.8.3"
|
package/scripts/check-fixture.js
CHANGED
|
@@ -53,7 +53,7 @@ function linkSelf() {
|
|
|
53
53
|
|
|
54
54
|
/**
|
|
55
55
|
* Creates the fixture's media files. Images are real (the pipeline reads their
|
|
56
|
-
* dimensions); the
|
|
56
|
+
* dimensions); the others only ever get copied, so a few bytes carrying the right
|
|
57
57
|
* extension are enough.
|
|
58
58
|
*/
|
|
59
59
|
async function writeMedia() {
|
|
@@ -76,6 +76,8 @@ async function writeMedia() {
|
|
|
76
76
|
|
|
77
77
|
writeFileSync( join( media, "doc.pdf" ), "%PDF-1.4\n% fixture placeholder\n" );
|
|
78
78
|
writeFileSync( join( media, "clip.mp4" ), "fixture placeholder\n" );
|
|
79
|
+
// An extension VitePress does not know, so its link rules take the file for a page.
|
|
80
|
+
writeFileSync( join( media, "patch.diff" ), "--- a\n+++ b\n" );
|
|
79
81
|
}
|
|
80
82
|
|
|
81
83
|
/**
|
|
@@ -194,6 +196,8 @@ console.log( "\nmedia that is not an image" );
|
|
|
194
196
|
expect( published.includes( "doc.pdf" ), "copies a document into the site output" );
|
|
195
197
|
expect( published.includes( "clip.mp4" ), "copies a video into the site output" );
|
|
196
198
|
expect( media.includes( 'href="/media/doc.pdf"' ), "rewrites a link to a media file" );
|
|
199
|
+
expect( media.includes( 'href="/media/patch.diff"' ),
|
|
200
|
+
"links a media file of an extension VitePress does not know as that file, not as a page" );
|
|
197
201
|
// On the entry page, where the fixture's page links live. VitePress resolves them to
|
|
198
202
|
// .html; what matters is that the media pipeline did not touch them.
|
|
199
203
|
expect( built( "index.html" ).includes( 'href="/containers.html"' ), "leaves a link to a page alone" );
|
|
@@ -321,6 +325,33 @@ expect( [ "Sunday", "21", "June", "2026" ].every( part => ownDate.includes( part
|
|
|
321
325
|
"hands a site's own component the timestamp formatting as well",
|
|
322
326
|
`the entry reads ${JSON.stringify( ownDate )}` );
|
|
323
327
|
|
|
328
|
+
// ── build report ─────────────────────────────────────────────────────────────
|
|
329
|
+
|
|
330
|
+
console.log( "\nbuild report" );
|
|
331
|
+
|
|
332
|
+
// pages/dead-link.md links to two files this build does not contain, one written
|
|
333
|
+
// relative and one site-absolute. Neither ends the build — a content mistake must
|
|
334
|
+
// not stand between a finished page and its publication — and both have to be
|
|
335
|
+
// findable afterwards by someone who cannot read a job log.
|
|
336
|
+
expect( /1 dead link|dead link\(s\) found/.test( log ) === false,
|
|
337
|
+
"lets a dead link through instead of ending the build over it" );
|
|
338
|
+
expect( /2 link\(s\) or image\(s\) point at files this build does not contain/.test( log ),
|
|
339
|
+
"names both spellings of a dead link in the build log" );
|
|
340
|
+
|
|
341
|
+
const report = JSON.parse( readFileSync( join( SITE, ".cratly", "build-report.json" ), "utf-8" ) );
|
|
342
|
+
|
|
343
|
+
expect( report.adapter?.version === version, "reports this package's version as the adapter version",
|
|
344
|
+
`found ${report.adapter?.version}` );
|
|
345
|
+
expect( report.fatal === false, "says the build survived what it found" );
|
|
346
|
+
expect( report.problems?.content?.map( p => p.reference ).sort().join( " " ) === "./www.example.com.html /tour.html",
|
|
347
|
+
"carries both dead references",
|
|
348
|
+
`found ${report.problems?.content?.map( p => p.reference ).join( " " ) || "nothing"}` );
|
|
349
|
+
expect( report.problems.content.every( p => p.source === "dead-link.md" ),
|
|
350
|
+
"names the file an author would open, not just the file the build wrote",
|
|
351
|
+
`found ${report.problems.content.map( p => p.source ).join( " " )}` );
|
|
352
|
+
expect( report.problems?.pages?.length === 0 && report.problems?.assets?.length === 0,
|
|
353
|
+
"leaves the fatal classes empty on a build that finished" );
|
|
354
|
+
|
|
324
355
|
// ── content security policy ──────────────────────────────────────────────────
|
|
325
356
|
|
|
326
357
|
console.log( "\ncontent security policy" );
|