scavold 0.2.0-rc.5 → 0.2.0-rc.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -9,6 +9,85 @@ changes may occur in any release.
9
9
 
10
10
  ## [Unreleased]
11
11
 
12
+ ## [0.2.0-rc.7] — 2026-09-29
13
+
14
+ ### Fixed
15
+
16
+ - A link to a media file whose extension VitePress does not know, such as `.diff`,
17
+ pointed at the file name with `.html` appended and therefore at nothing. It now names
18
+ the file as published.
19
+
20
+ ## [0.2.0-rc.6] — 2026-09-13
21
+
22
+ ### Added
23
+
24
+ - `aliases` in a page's front matter names the former addresses it replaces, and
25
+ `retired_urls` in `.cratly.config.yaml` those that are not coming back. A site taking
26
+ over from an existing website declares them once and the build writes
27
+ `cratly-redirects.json` into its output: an address, where it went, and `301` or `410`.
28
+ The format names no server — only a server-side redirect passes a ranking on, and which
29
+ server that is cratly does not assume, so the deploy step translates the file for its
30
+ target. Keeping an address is still better than redirecting to it, which is what `url`
31
+ is for; the build refuses an alias that two pages claim, or that the site serves itself
32
+ and the redirect would therefore hide.
33
+ - Every build now ends by reading its own result. A page that the generator listed but
34
+ never wrote, or a script a built page names and the output does not contain, ends the
35
+ build with the file named — the two states a bundling fault leaves behind while the log
36
+ says "build complete". Links and images pointing at absent files are reported and left
37
+ to the site, since one may legitimately name a path the web server provides. Turn it
38
+ down with `verify: { content: false }` or off with `verify: false`. A reference is read
39
+ the way a browser reads it: a relative one against the address of the page holding it, so
40
+ a link that forgot its protocol — `[Website](./www.example.com)`, which the build turns
41
+ into a sibling page nobody wrote — is seen as well. Proven against the
42
+ case-collision that used to build green: the page missing from the output is the only
43
+ trace it leaves on a file system that ignores letter case, and that is now enough.
44
+ Everything found also goes to `.cratly/build-report.json`, written before the build is
45
+ ended rather than after, each entry naming the source file an author would open. A build
46
+ log is read by whoever has access to it; a file can be handed to whoever wrote the page.
47
+ - `ignoreDeadLinks` now defaults to `true`. VitePress ends a build on a dead link and does
48
+ so before writing anything, which leaves the reason in a job log that whoever wrote the
49
+ link often cannot read, and lets one typo stand between a finished page and its
50
+ publication. Scavold finds the same mistake in the finished output and reports it there
51
+ instead. Whether it should also stop a pipeline is a question whose answer differs per
52
+ branch, so it is left to CI: the site template fails a merge request over it and lets the
53
+ default branch deploy, since blocking there would answer a broken link with a stale site.
54
+ Ask for VitePress' abort back with `ignoreDeadLinks: false`.
55
+ - A site whose page paths collide in VitePress's eyes now learns that from Scavold, by
56
+ name, before the build starts. VitePress identifies a page by its path flattened into
57
+ a single name — slash becomes underscore, and the page-hash map lower-cases it — and
58
+ uses that name for the page's bundle entry, its server module and its client chunk.
59
+ Two pages meeting there overwrite each other: either the build dies at render time in
60
+ `pageChunk.imports`, naming no file, or it succeeds and ships two pages pointing at a
61
+ script that was never written. Scavold now reproduces that identity when it reads the
62
+ config and aborts with both file names, whether they collide through an underscore in
63
+ a file name (`de_kontakt.md` beside `de/kontakt.md`), through letter case alone, or
64
+ through a `url` alias landing on a path another page already owns as its file — the
65
+ case the previous alias check, which compared aliases only with each other, let pass.
66
+ - `date` and `datetime` join the type vocabulary of `frontmatter_fields` and of a
67
+ container's `props`. Scavold passes such a value through as the ISO 8601 text it is;
68
+ what changes is that the cratly editor now offers a date picker for it instead of a
69
+ plain text box. A `datetime` always carries an offset — a time of day without one is
70
+ a different moment on every machine that builds the site.
71
+
72
+ ### Changed
73
+
74
+ - `@cepharum/vue3-i18n` moves to `2.0.0`, two majors on from the `0.5.4` a site would
75
+ have installed. Scavold uses `useL10n`, `setLocale` and `setLoader`, none of which
76
+ changed; what 2.0 breaks are the locale helpers now returning full BCP-47 tags, and
77
+ Scavold calls none of them. A site inherits the new version with its next install, and
78
+ gains what 2.0 adds along the way: regional locales overlaying a bare language, plural
79
+ and gender selection resolved through `Intl`, and `Intl`-backed formatters in
80
+ placeholders.
81
+
82
+ ### Fixed
83
+
84
+ - `lib/href.js` ships the declarations it had been missing since it was split out, so
85
+ `typecheck` passes again — and the pipeline runs it now, which is why it could go
86
+ unnoticed at all. Nothing about the module changed; TypeScript simply had nothing to
87
+ read about it, and treated every href the hierarchy composes as `any`.
88
+
89
+ ## [0.2.0-rc.5] — 2026-08-17
90
+
12
91
  ### Added
13
92
 
14
93
  - `:::pagelist` dates its entries the way the site wants them: `time-style` adds the
@@ -179,7 +258,11 @@ First published release. Version `0.1.0` existed in-tree only.
179
258
  - Only images are processed out of `media_folder`; other file types (video, documents)
180
259
  are never copied into the build output and have to live in the static folder.
181
260
 
182
- [0.2.0-rc.4]: https://gitlab.com/cepharum-foss/cratly/scavold/-/compare/v0.2.0-rc.3...v0.2.0-rc.4
183
- [0.2.0-rc.3]: https://gitlab.com/cepharum-foss/cratly/scavold/-/compare/v0.2.0-rc.2...v0.2.0-rc.3
184
- [0.2.0-rc.2]: https://gitlab.com/cepharum-foss/cratly/scavold/-/compare/v0.2.0-rc.1...v0.2.0-rc.2
185
- [0.2.0-rc.1]: https://gitlab.com/cepharum-foss/cratly/scavold/-/tags/v0.2.0-rc.1
261
+ [Unreleased]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.7...main
262
+ [0.2.0-rc.7]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.6...v0.2.0-rc.7
263
+ [0.2.0-rc.6]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.5...v0.2.0-rc.6
264
+ [0.2.0-rc.5]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.4...v0.2.0-rc.5
265
+ [0.2.0-rc.4]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.3...v0.2.0-rc.4
266
+ [0.2.0-rc.3]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.2...v0.2.0-rc.3
267
+ [0.2.0-rc.2]: https://gitlab.com/cratly/scavold/-/compare/v0.2.0-rc.1...v0.2.0-rc.2
268
+ [0.2.0-rc.1]: https://gitlab.com/cratly/scavold/-/tags/v0.2.0-rc.1
package/FRONTMATTER.md CHANGED
@@ -107,6 +107,20 @@ Menu links and breadcrumbs produced by Scavold automatically use the alias URL.
107
107
  with an error naming both conflicting files. Each alias must be unique across the
108
108
  entire site.
109
109
 
110
+ The same applies when an alias lands on a path another page already owns as its own
111
+ file — a `url: de/kontakt` next to a `pages/de/kontakt.md`. VitePress identifies a
112
+ page by its path flattened into one name, slash becoming underscore, and it uses that
113
+ name for the page's bundle entry, its server module and its client chunk. Two pages
114
+ meeting there overwrite each other's output; the alias does not win, it collides. So
115
+ does a page whose file name spells out a path another page reaches through folders
116
+ (`de_kontakt.md` beside `de/kontakt.md`), and a pair differing only in letter case.
117
+ Scavold checks all of these when it reads the config and names both files. Prefer `-`
118
+ over `_` as the word separator in page names and the class cannot arise.
119
+
120
+ To move a page and keep its URL, rename the file and declare the old URL as its
121
+ alias — the alias belongs to the page that carries the content, not to a second file
122
+ left behind at the old location.
123
+
110
124
  ---
111
125
 
112
126
  ## `locale` / `lang`
@@ -260,6 +274,44 @@ in the hierarchy tree — they are only excluded from rendered navigation.
260
274
 
261
275
  ---
262
276
 
277
+ ## `aliases`
278
+
279
+ **Type:** `string | string[]`
280
+
281
+ Former addresses this page replaces. Declared when a cratly site takes over from an
282
+ existing website: every address a search engine already knows and the new site does not
283
+ answer is a result its owner loses.
284
+
285
+ ```yaml
286
+ ---
287
+ aliases:
288
+ - /kontakt.html
289
+ - kontakt.php
290
+ ---
291
+ ```
292
+
293
+ Values are written as they stood in the browser, with or without a leading slash. A
294
+ fragment is dropped — it never reaches a server. A query string is kept; whether it can
295
+ be matched is up to whoever serves the site.
296
+
297
+ Scavold collects them and writes `cratly-redirects.json` into the build output: an
298
+ address, where it went, and a status — `301` for an inherited address, `410` for one a
299
+ site declares as gone for good in
300
+ [`retired_urls`](./docs/reference/config.md#retired-urls). Turning that into server
301
+ configuration is the deploy step's job, because cratly does not assume who serves the
302
+ site. Nothing about the file is Scavold-specific: another adapter writes the same format.
303
+
304
+ **Keeping the address beats redirecting to it.** A permanent redirect passes a ranking on;
305
+ a page that simply answers at the old address never spends it. Where the legacy path can
306
+ be a file, [`url`](#url) is the better tool, and it works on every static server, whether
307
+ or not it can redirect at all.
308
+
309
+ **Conflict detection:** the build stops when two pages claim the same former address, and
310
+ when a former address is one the new site serves itself — the redirect would win over the
311
+ page, leaving it unreachable while every menu still links to it.
312
+
313
+ ---
314
+
263
315
  ## `redirect`
264
316
 
265
317
  **Type:** `string | object`
package/lib/config.js CHANGED
@@ -1,7 +1,9 @@
1
1
  import { dirname as dirnameOfPath, join, relative, resolve } from "node:path";
2
2
  import { dirname, join as joinPosix } from "node:path/posix";
3
3
  import { readdirSync, readFileSync } from "node:fs";
4
- import { clearCache, compileHierarchy, compileRedirects, resolveHierarchyImages, sourceFolder } from "./pages.js";
4
+ import {
5
+ clearCache, collectFrontmatterOfAllPages, compileHierarchy, compileRedirects, resolveHierarchyImages, sourceFolder,
6
+ } from "./pages.js";
5
7
  import { extractFrontmatterMediaSrcs } from "./excerpt.js";
6
8
  import { DEFAULT_FEED_LIMIT, collectFeedPages, renderRss } from "./feed.js";
7
9
  import { useMedia } from "./media.js";
@@ -9,6 +11,13 @@ import { patchRenderer } from "./markdown.js";
9
11
  import { extractContainerMediaSrcs, mediaPropsOf, registerContainers, resolveContainerMap } from "./containers.js";
10
12
  import { buildSectionManifest, writeSectionManifest } from "./sectionManifest.js";
11
13
  import { isExternalUrl, servableRedirectTarget } from "./redirectTarget.js";
14
+ import { assertUniquePageKeys } from "./pageKeys.js";
15
+ import {
16
+ BUILD_REPORT_DIR, BUILD_REPORT_FILE, formatProblems, renderBuildReport, verifyBuild,
17
+ } from "./verifyBuild.js";
18
+ import {
19
+ REDIRECT_MANIFEST_FILE, assertLegacyRedirects, compileLegacyRedirects, renderRedirectManifest,
20
+ } from "./legacyUrls.js";
12
21
  import { pageHref, servableUrl } from "./href.js";
13
22
 
14
23
  // Scavold's own version, reported under `adapter.version` in the emitted
@@ -180,6 +189,7 @@ export async function augmentConfig( rawConfig, options = {} ) {
180
189
  let declaredImageSizes = null;
181
190
  let declaredImageWidths = null;
182
191
  let declaredExcerptLength = null;
192
+ let declaredRetiredUrls = [];
183
193
 
184
194
  try {
185
195
  const { readFile } = await import( "node:fs/promises" );
@@ -193,6 +203,7 @@ export async function augmentConfig( rawConfig, options = {} ) {
193
203
  declaredImageSizes = cratlyConfig.image_sizes ?? null;
194
204
  declaredImageWidths = cratlyConfig.image_widths ?? null;
195
205
  declaredExcerptLength = cratlyConfig.excerpt_length ?? null;
206
+ declaredRetiredUrls = cratlyConfig.retired_urls ?? [];
196
207
  } catch {
197
208
  // file absent or unparseable — proceed with defaults
198
209
  }
@@ -376,6 +387,99 @@ export async function augmentConfig( rawConfig, options = {} ) {
376
387
  await writeFile( target, xml, "utf-8" );
377
388
  }
378
389
 
390
+ /**
391
+ * Writes what the check found next to the section manifest, for readers that are
392
+ * not a terminal — the cratly editor shows it to whoever wrote the page.
393
+ *
394
+ * A report that cannot be written must not replace the finding it describes, so
395
+ * failure here is a warning and the build carries on to its own verdict.
396
+ *
397
+ * @param {{pages: object[], assets: object[], content: object[]}} problems what verifyBuild found
398
+ * @param {Object<string,string>} sources built page path → source file an author edits
399
+ * @returns {Promise<void>}
400
+ */
401
+ async function writeBuildReport( problems, sources ) {
402
+ try {
403
+ const { writeFile, mkdir } = await import( "node:fs/promises" );
404
+ const dir = join( process.cwd(), BUILD_REPORT_DIR );
405
+
406
+ await mkdir( dir, { recursive: true } );
407
+ await writeFile(
408
+ join( dir, BUILD_REPORT_FILE ),
409
+ renderBuildReport( problems, {
410
+ adapter: { name: "scavold", version: SCAVOLD_VERSION },
411
+ sources,
412
+ } ),
413
+ "utf-8",
414
+ );
415
+ } catch ( error ) {
416
+ console.warn( `[scavold] could not write ${BUILD_REPORT_DIR}/${BUILD_REPORT_FILE} —`, error.message );
417
+ }
418
+ }
419
+
420
+ /**
421
+ * Reads the finished build and refuses one that references files it does not
422
+ * contain — the shape a framework bug takes when it does not announce itself.
423
+ *
424
+ * A site can turn the content half off (`verify: { content: false }`) or the whole
425
+ * check (`verify: false`) when it links to paths its web server provides rather
426
+ * than its build.
427
+ *
428
+ * @param {object} siteConfig VitePress site configuration, as handed to buildEnd
429
+ * @returns {Promise<void>}
430
+ */
431
+ async function verifyBuiltSite( siteConfig ) {
432
+ if ( options.verify === false ) {
433
+ return;
434
+ }
435
+
436
+ const { content = true } = options.verify === true || options.verify == null ? {} : options.verify;
437
+
438
+ // What VitePress said it would write, in the spelling it serves it under —
439
+ // and, the other way round, which file an author would open to fix a page.
440
+ const sources = {};
441
+
442
+ for ( const page of siteConfig.pages ?? [] ) {
443
+ sources[( siteConfig.rewrites?.map?.[page] || page ).replace( /\.md$/, ".html" )] = page;
444
+ }
445
+
446
+ const problems = verifyBuild( siteConfig.outDir, {
447
+ base: siteConfig.site?.base ?? "/",
448
+ content,
449
+ pages: Object.keys( sources ),
450
+ } );
451
+
452
+ // Written before anything is thrown: the failure this check exists to cause
453
+ // would otherwise take the reason with it, leaving a red pipeline and a log
454
+ // only some people may read.
455
+ await writeBuildReport( problems, sources );
456
+
457
+ if ( problems.content.length ) {
458
+ console.warn(
459
+ `[scavold] ${problems.content.length} link(s) or image(s) point at files this build does not contain:\n` +
460
+ `${formatProblems( problems.content )}`
461
+ );
462
+ }
463
+
464
+ if ( problems.pages.length ) {
465
+ throw new Error(
466
+ `[scavold] ${problems.pages.length} page(s) the build listed were never written:\n` +
467
+ `${formatProblems( problems.pages )}\n` +
468
+ "Two pages sharing one build identity are the known cause — on a file system that ignores letter " +
469
+ "case, this is the only trace they leave."
470
+ );
471
+ }
472
+
473
+ if ( problems.assets.length ) {
474
+ throw new Error(
475
+ `[scavold] ${problems.assets.length} script(s) or stylesheet(s) named by a built page are missing from the ` +
476
+ `output:\n${formatProblems( problems.assets )}\n` +
477
+ "Pages in that state render but stay dead in the browser. This is a bundling fault, not a content " +
478
+ "mistake — two pages sharing one build identity are the known cause."
479
+ );
480
+ }
481
+ }
482
+
379
483
  const pagesDir = sourceFolder( resolvedConfig );
380
484
 
381
485
  // Emit the section-type manifest (.cratly/sections.json) so the cratly
@@ -405,9 +509,87 @@ export async function augmentConfig( rawConfig, options = {} ) {
405
509
  },
406
510
  };
407
511
 
512
+ // ─── Page identity guard ─────────────────────────────────────────────────
513
+ // VitePress identifies a page by its path flattened into a single name, and
514
+ // two pages colliding there quietly overwrite each other — see pageKeys.js.
515
+ // Checked here, on every config load, so dev server and build both report the
516
+ // pair by name instead of dying later on a chunk that was never emitted.
517
+
518
+ /**
519
+ * Lists the page files VitePress will pick up, relative to srcDir.
520
+ *
521
+ * `srcExclude` is honoured for plain paths and for prefix patterns ending in
522
+ * `/*` or `/**` — the shapes a site actually writes. A more exotic pattern
523
+ * just leaves its file in the list, which at worst reports a collision the
524
+ * site had excluded anyway.
525
+ *
526
+ * @returns {Promise<string[]>} page paths, posix-style, relative to srcDir
527
+ */
528
+ async function listPageFiles() {
529
+ const { glob } = await import( "node:fs/promises" );
530
+ const excluded = ( resolvedConfig.srcExclude ?? [] ).map( pattern => pattern.replace( /\/\*{1,2}$/, "/" ) );
531
+
532
+ return ( await Array.fromAsync( glob( "**/*.md", {
533
+ cwd: sourceFolder( resolvedConfig ),
534
+ exclude: name => name === "node_modules" || name === "dist",
535
+ } ) ) )
536
+ .map( name => name.replace( /\\/g, "/" ) )
537
+ // Dynamic routes stand for pages named by their paths file, not by
538
+ // their own name — their identity is not this file's path.
539
+ .filter( name => !name.includes( "[" ) )
540
+ .filter( name => !excluded.some( pattern => ( pattern.endsWith( "/" ) ? name.startsWith( pattern ) : name === pattern ) ) );
541
+ }
542
+
543
+ const rewrites = {
544
+ ...rawConfig?.rewrites,
545
+ ...await compileRedirects( resolvedConfig ),
546
+ };
547
+
548
+ assertUniquePageKeys( await listPageFiles(), rewrites );
549
+
550
+ // ─── Former addresses ────────────────────────────────────────────────────
551
+ // Compiled here so a declaration a server could not honour is refused while the
552
+ // author is still looking at it, and written out at the end of the build for
553
+ // whoever configures the server. What Scavold produces is data, not server
554
+ // configuration: cratly does not assume who serves the site.
555
+ const legacyRedirects = compileLegacyRedirects(
556
+ await collectFrontmatterOfAllPages( resolvedConfig ),
557
+ { rewrites, retired: declaredRetiredUrls },
558
+ );
559
+
560
+ assertLegacyRedirects( legacyRedirects.problems );
561
+
562
+ /**
563
+ * Writes the redirect manifest, unless the site declares no former addresses.
564
+ *
565
+ * @param {object} siteConfig VitePress site configuration, as handed to buildEnd
566
+ * @returns {Promise<void>}
567
+ */
568
+ async function writeRedirectManifest( siteConfig ) {
569
+ if ( !legacyRedirects.rules.length ) {
570
+ return;
571
+ }
572
+
573
+ const { writeFile } = await import( "node:fs/promises" );
574
+
575
+ await writeFile(
576
+ join( siteConfig.outDir, REDIRECT_MANIFEST_FILE ),
577
+ renderRedirectManifest( legacyRedirects.rules, { name: "scavold", version: SCAVOLD_VERSION } ),
578
+ "utf-8",
579
+ );
580
+ }
581
+
408
582
  return {
409
583
  ...resolvedConfig,
410
584
 
585
+ // VitePress ends a build on a dead link, and ends it before writing anything —
586
+ // so the reason lives in a job log, which whoever wrote the link often cannot
587
+ // read, and a single typo stands between a finished page and its publication.
588
+ // Scavold finds the same mistake in the finished output (see verifyBuiltSite)
589
+ // and reports it there, where it can be handed to the author. A site that
590
+ // wants the abort back asks for it: ignoreDeadLinks: false.
591
+ ignoreDeadLinks: rawConfig.ignoreDeadLinks ?? true,
592
+
411
593
  // Without this, VitePress writes its page-to-hash map into an inline script whose
412
594
  // content changes with every build. Under a content security policy that names
413
595
  // script hashes, that one hash has to be updated on the server for every deploy —
@@ -440,13 +622,14 @@ export async function augmentConfig( rawConfig, options = {} ) {
440
622
  createExternalRedirectMapPlugin( pagesDir ),
441
623
  ],
442
624
  },
443
- rewrites: {
444
- ...rawConfig?.rewrites,
445
- ...await compileRedirects( resolvedConfig ),
446
- },
625
+ rewrites,
447
626
  async buildEnd( siteConfig ) {
448
627
  await writeFeed( siteConfig );
628
+ await writeRedirectManifest( siteConfig );
449
629
  await rawConfig.buildEnd?.( siteConfig );
630
+
631
+ // Last, so it sees what every other hand has written.
632
+ await verifyBuiltSite( siteConfig );
450
633
  },
451
634
  async transformPageData( pageData, context ) {
452
635
  clearCache();
@@ -470,7 +653,7 @@ export async function augmentConfig( rawConfig, options = {} ) {
470
653
  // during in-app use.
471
654
  //
472
655
  // The redirect value can be either:
473
- // - a site-relative URL ("/de/jobs") — written by the Cratly editor
656
+ // - a site-relative URL ("/de/jobs") — written by the cratly editor
474
657
  // - a relative .md path ("jobs.md") — written by hand in frontmatter
475
658
  // - an absolute external URL — e.g. "https://example.com"
476
659
  const redirect = pageData.frontmatter?.redirect;
@@ -531,14 +714,31 @@ export async function augmentConfig( rawConfig, options = {} ) {
531
714
  // offers, which is relative to that folder. Rewrite it to the URL the
532
715
  // built site serves. Targets that are not media files (ordinary page
533
716
  // links, anchors, remote URLs) resolve to null and stay untouched.
534
- patchRenderer( md, "link_open", ( defaultHandler, tokens, idx, ...args ) => {
535
- const published = media.collectAsset( tokens[idx].attrGet( "href" ) );
717
+ //
718
+ // VitePress then takes any extension missing from its own list of file
719
+ // types (`.diff`, say) for part of a page name and appends `.html`. The
720
+ // published URL is final, so that suffix is taken back — after VitePress
721
+ // ran, which leaves prefixing the site's base to it.
722
+ patchRenderer( md, "link_open", ( defaultHandler, tokens, idx, mdOptions, env, slf ) => {
723
+ const token = tokens[idx];
724
+ const published = media.collectAsset( token.attrGet( "href" ) );
725
+
726
+ if ( !published ) {
727
+ return defaultHandler( tokens, idx, mdOptions, env, slf );
728
+ }
729
+
730
+ token.attrSet( "href", published );
731
+
732
+ const rendered = defaultHandler( tokens, idx, mdOptions, env, slf );
733
+ const href = token.attrGet( "href" );
734
+
735
+ if ( !published.endsWith( ".html" ) && href?.endsWith( published + ".html" ) ) {
736
+ token.attrSet( "href", href.slice( 0, -".html".length ) );
536
737
 
537
- if ( published ) {
538
- tokens[idx].attrSet( "href", published );
738
+ return slf.renderToken( tokens, idx, mdOptions );
539
739
  }
540
740
 
541
- return defaultHandler( tokens, idx, ...args );
741
+ return rendered;
542
742
  } );
543
743
 
544
744
  registerContainers( md, containerMap, {
package/lib/href.d.ts ADDED
@@ -0,0 +1,21 @@
1
+ import type { Scavold } from "../index";
2
+
3
+ /** Extension every page is built into and therefore linked with. */
4
+ export const PAGE_EXTENSION: string;
5
+
6
+ /**
7
+ * Converts a page's source path into the root-relative URL the built site serves.
8
+ */
9
+ export function pageHref( path: string ): string;
10
+
11
+ /**
12
+ * Converts a hierarchy node into the URL the built site serves it under, preferring
13
+ * the alias URL when the page declares one.
14
+ */
15
+ export function nodePath( node: Scavold.HierarchyNode | undefined ): string;
16
+
17
+ /**
18
+ * Rewrites a site-relative URL to the file that answers it: a folder URL to the index
19
+ * page inside it, an extension-less path to its page.
20
+ */
21
+ export function servableUrl( url: string ): string;
@@ -0,0 +1,179 @@
1
+ // ─── Legacy URLs ─────────────────────────────────────────────────────────────
2
+ // A site that replaces an existing one inherits its addresses. Every one of them
3
+ // that a search engine knows and the new site does not answer is a result the
4
+ // customer loses, and no build notices — the old URL simply is not there any more.
5
+ //
6
+ // Only the web server can answer an address that is not a file, and only a
7
+ // permanent redirect passes the ranking on; a redirect run by the browser is
8
+ // followed but counts for less. So the pages declare which old addresses they
9
+ // replace, and the build writes them out as data. What turns that data into
10
+ // server configuration is the deploy step of whoever serves the site — cratly
11
+ // does not assume any particular server, so this file knows about none.
12
+ //
13
+ // Where an old address can simply be kept, `url` in front matter is the better
14
+ // answer: a page served at the legacy path needs no redirect at all, and works on
15
+ // every static server there is.
16
+
17
+ import { pageKey } from "./pageKeys.js";
18
+
19
+ // Version of the cratly redirect-manifest grammar this writes, and its canonical
20
+ // URL — cratly owns the format, Scavold is one adapter producing it.
21
+ export const REDIRECT_MANIFEST_VERSION = 0;
22
+ export const REDIRECT_MANIFEST_SCHEMA = "https://cratly.io/schema/redirects/v0.json";
23
+
24
+ // Name of the file the build writes into its output folder.
25
+ export const REDIRECT_MANIFEST_FILE = "cratly-redirects.json";
26
+
27
+ /**
28
+ * Turns a page's source path into the URL its build is served under.
29
+ *
30
+ * @param {string} page path of the page's markdown file, relative to srcDir
31
+ * @param {Object<string,string>} [rewrites] VitePress `rewrites` map, source path → served path
32
+ * @returns {string} site-absolute URL of the built page
33
+ */
34
+ export function servedPathOf( page, rewrites = {} ) {
35
+ const served = rewrites?.[page] || page;
36
+
37
+ return `/${served.replace( /\.md$/, ".html" )}`;
38
+ }
39
+
40
+ /**
41
+ * Reads a declared legacy address into the shape a redirect rule uses.
42
+ *
43
+ * Authors write these as they appeared in the browser, so a leading slash may be
44
+ * missing and a fragment may be attached — a fragment never reaches the server and
45
+ * is dropped. A query string is kept: whether it can be matched is the consuming
46
+ * server's business, not this format's.
47
+ *
48
+ * @param {string} value legacy address as declared
49
+ * @returns {string|null} site-absolute address, or null when the value says nothing
50
+ */
51
+ export function normaliseLegacyPath( value ) {
52
+ if ( typeof value !== "string" ) {
53
+ return null;
54
+ }
55
+
56
+ const [withoutFragment] = value.trim().split( "#" );
57
+
58
+ if ( !withoutFragment ) {
59
+ return null;
60
+ }
61
+
62
+ return withoutFragment.startsWith( "/" ) ? withoutFragment : `/${withoutFragment}`;
63
+ }
64
+
65
+ /**
66
+ * Compiles the redirect rules a site's own declarations amount to.
67
+ *
68
+ * Two mistakes are worth stopping a build over, and both are reported rather than
69
+ * resolved: an old address that two pages claim, and one that collides with an
70
+ * address the new site itself serves. The latter is the interesting one — a
71
+ * redirect would shadow a real page, so the page would be unreachable while every
72
+ * link to it still points there.
73
+ *
74
+ * @param {Object<string,object>} frontmatter map of page path → front matter
75
+ * @param {object} [options] what the site declares elsewhere
76
+ * @param {Object<string,string>} [options.rewrites] VitePress `rewrites` map
77
+ * @param {string[]} [options.retired] addresses declared as gone for good
78
+ * @returns {{rules: {from: string, to?: string, status: number}[], problems: string[]}}
79
+ * rules in address order, and what a site has to fix first
80
+ */
81
+ export function compileLegacyRedirects( frontmatter, { rewrites = {}, retired = [] } = {} ) {
82
+ const problems = [];
83
+ const claimedBy = new Map(); // legacy address → page claiming it
84
+ const rules = new Map(); // legacy address → rule
85
+
86
+ // Addresses the new site answers itself. A redirect must never cover one.
87
+ const served = new Map();
88
+
89
+ for ( const page of Object.keys( frontmatter ) ) {
90
+ served.set( pageKey( servedPathOf( page, rewrites ).slice( 1 ) ), page );
91
+ }
92
+
93
+ for ( const [ page, meta ] of Object.entries( frontmatter ) ) {
94
+ const declared = meta?.aliases == null ? [] : [meta.aliases].flat();
95
+
96
+ for ( const value of declared ) {
97
+ const from = normaliseLegacyPath( value );
98
+
99
+ if ( !from ) {
100
+ continue;
101
+ }
102
+
103
+ const owner = claimedBy.get( from );
104
+
105
+ if ( owner ) {
106
+ problems.push( `both "${owner}" and "${page}" declare the former address ${from}` );
107
+ continue;
108
+ }
109
+
110
+ const shadowed = served.get( pageKey( from.slice( 1 ) ) );
111
+
112
+ if ( shadowed ) {
113
+ problems.push(
114
+ `"${page}" declares the former address ${from}, which "${shadowed}" is served at — ` +
115
+ "a redirect there would hide that page"
116
+ );
117
+ continue;
118
+ }
119
+
120
+ claimedBy.set( from, page );
121
+ rules.set( from, { from, to: servedPathOf( page, rewrites ), status: 301 } );
122
+ }
123
+ }
124
+
125
+ for ( const value of retired ) {
126
+ const from = normaliseLegacyPath( value );
127
+
128
+ if ( !from ) {
129
+ continue;
130
+ }
131
+
132
+ if ( rules.has( from ) ) {
133
+ problems.push( `${from} is declared as retired and at the same time as a former address of "${claimedBy.get( from )}"` );
134
+ continue;
135
+ }
136
+
137
+ rules.set( from, { from, status: 410 } );
138
+ }
139
+
140
+ return {
141
+ rules: [...rules.values()].sort( ( a, b ) => ( a.from < b.from ? -1 : 1 ) ),
142
+ problems,
143
+ };
144
+ }
145
+
146
+ /**
147
+ * Refuses a set of declarations a server could not serve as meant.
148
+ *
149
+ * @param {string[]} problems what compileLegacyRedirects found
150
+ * @returns {void}
151
+ * @throws {Error} naming every problem
152
+ */
153
+ export function assertLegacyRedirects( problems ) {
154
+ if ( problems.length ) {
155
+ throw new Error( `[scavold] former addresses cannot be served as declared: ${problems.join( "; " )}` );
156
+ }
157
+ }
158
+
159
+ /**
160
+ * Renders the redirect manifest, the artifact a deploy step reads.
161
+ *
162
+ * Deliberately free of any server's vocabulary: an address, where it went, and the
163
+ * status that says so. Translating that into a Caddyfile, an nginx map, an
164
+ * `.s3-http.config.yaml` or a bucket's routing rules is the deploy's job.
165
+ *
166
+ * @param {{from: string, to?: string, status: number}[]} rules compiled rules
167
+ * @param {object} [adapter] what produced the file
168
+ * @param {string} [adapter.name] name of the adapter
169
+ * @param {string} [adapter.version] its version
170
+ * @returns {string} the manifest as JSON text, newline-terminated
171
+ */
172
+ export function renderRedirectManifest( rules, adapter = {} ) {
173
+ return `${JSON.stringify( {
174
+ $schema: REDIRECT_MANIFEST_SCHEMA,
175
+ specVersion: REDIRECT_MANIFEST_VERSION,
176
+ adapter,
177
+ rules,
178
+ }, null, "\t" )}\n`;
179
+ }
@@ -0,0 +1,111 @@
1
+ // ─── Page identity guard ─────────────────────────────────────────────────────
2
+ // VitePress does not carry a page's path through its build. It flattens the path
3
+ // into a single name and uses that name as the page's identity in four places:
4
+ //
5
+ // * the key of the page's entry in `rollupOptions.input` (the rewritten path
6
+ // wins here — `config.rewrites.map[file] || file`),
7
+ // * the file name of the page's SSR module in the temp folder,
8
+ // * the key of the page's entry in `pageToHashMap` — lower-cased,
9
+ // * the name of the page's client chunk, hence its `.lean.js` file name.
10
+ //
11
+ // Two pages whose paths flatten to the same name therefore share one identity
12
+ // and one set of output files. VitePress neither detects nor reports this:
13
+ // depending on which of the four collides, the build either dies at render time
14
+ // with `undefined is not an object (evaluating 'pageChunk.imports')`, naming no
15
+ // file, or it succeeds and ships two pages pointing at one script file that was
16
+ // never written.
17
+ //
18
+ // The guard below reproduces that identity and demands it stays unique, so the
19
+ // site learns which two files disagree instead of what VitePress makes of them.
20
+
21
+ // Copied from VitePress (which copies it from Vite) — the characters its
22
+ // `sanitizeFileName()` replaces with an underscore.
23
+
24
+ const INVALID_CHAR_RE = /[\u0000-\u001F"#$&*+,:;<=>?[\]^`{|}\u007F]/g;
25
+
26
+ /**
27
+ * Reproduces the identity VitePress derives from a page path.
28
+ *
29
+ * Lower-casing goes beyond what the bundle entry key does — it is what
30
+ * `pageToHashMap` keys on, and a pair colliding only there is the worse of the
31
+ * two failures, so the strictest of the four rules is the one to enforce.
32
+ *
33
+ * @param {string} page path of a page's markdown file, relative to srcDir
34
+ * @returns {string} the name VitePress identifies that page by
35
+ */
36
+ export function pageKey( page ) {
37
+ return String( page )
38
+ .replace( /\\/g, "/" )
39
+ .replace( /\//g, "_" )
40
+ .replace( INVALID_CHAR_RE, "_" )
41
+ // VitePress strips underscores that a leading invalid character produced.
42
+ .replace( /^_+/, "" )
43
+ .toLowerCase();
44
+ }
45
+
46
+ /**
47
+ * Groups pages that VitePress cannot tell apart.
48
+ *
49
+ * @param {string[]} pages paths of all page files, relative to srcDir
50
+ * @param {Object<string,string>} [rewrites] VitePress `rewrites` map, source path → served path
51
+ * @returns {{key: string, sources: {page: string, alias: string}[]}[]} one entry per colliding identity, in path order
52
+ */
53
+ export function findPageKeyCollisions( pages, rewrites = {} ) {
54
+ const byKey = new Map();
55
+
56
+ for ( const page of [...pages].sort() ) {
57
+ const alias = rewrites?.[page] || page;
58
+ const key = pageKey( alias );
59
+ const group = byKey.get( key );
60
+
61
+ if ( group ) {
62
+ group.push( { page, alias } );
63
+ } else {
64
+ byKey.set( key, [{ page, alias }] );
65
+ }
66
+ }
67
+
68
+ return [...byKey.entries()]
69
+ .filter( ( [ , sources ] ) => sources.length > 1 )
70
+ .map( ( [ key, sources ] ) => ( { key, sources } ) );
71
+ }
72
+
73
+ /**
74
+ * Renders one collision as a line naming both files and, where a rewrite or a
75
+ * `url` alias is what makes them collide, the path each is served at.
76
+ *
77
+ * @param {{key: string, sources: {page: string, alias: string}[]}} collision
78
+ * @returns {string} single-line description
79
+ */
80
+ function describeCollision( { key, sources } ) {
81
+ const named = sources
82
+ .map( ( { page, alias } ) => ( alias === page ? `"${page}"` : `"${page}" (served as "${alias}")` ) )
83
+ .join( " and " );
84
+
85
+ return `${named} share the build identity "${key}"`;
86
+ }
87
+
88
+ /**
89
+ * Throws when two pages flatten to the same VitePress identity.
90
+ *
91
+ * @param {string[]} pages paths of all page files, relative to srcDir
92
+ * @param {Object<string,string>} [rewrites] VitePress `rewrites` map, source path → served path
93
+ * @returns {void}
94
+ * @throws {Error} naming every colliding pair and how to resolve it
95
+ */
96
+ export function assertUniquePageKeys( pages, rewrites = {} ) {
97
+ const collisions = findPageKeyCollisions( pages, rewrites );
98
+
99
+ if ( collisions.length === 0 ) {
100
+ return;
101
+ }
102
+
103
+ throw new Error(
104
+ `[scavold] colliding page paths: ${collisions.map( describeCollision ).join( "; " )}. ` +
105
+ "VitePress flattens a page path into a single name — slash becomes underscore — and " +
106
+ "identifies the page by it, so pages colliding there overwrite each other's bundle " +
107
+ "entry and output files. Rename one of them; prefer \"-\" over \"_\" as the word " +
108
+ "separator in page names, and keep an existing URL alive with front matter `url` " +
109
+ "instead of the old file name."
110
+ );
111
+ }
@@ -0,0 +1,324 @@
1
+ // ─── Build verification ──────────────────────────────────────────────────────
2
+ // A static site can be wrong in ways its generator does not notice. VitePress
3
+ // identifies a page by its path flattened into one name and reuses that name for
4
+ // the page's script files; where two pages meet there, one of them ends up
5
+ // preloading a file that was never written — and the build reports success. The
6
+ // same silence covers a renamed page still linked from elsewhere and an image
7
+ // whose file went away.
8
+ //
9
+ // None of that is visible in the build log, and all of it is visible in the
10
+ // finished output: the pages name the files they need, and either those files are
11
+ // there or they are not. So the last thing Scavold does is read its own result.
12
+ //
13
+ // Three classes, treated differently:
14
+ //
15
+ // * A page the generator listed but did not write is never intentional either —
16
+ // and it is the only symptom left when two pages collide on a file system that
17
+ // ignores letter case, which is what the machine writing the page usually has.
18
+ // * A missing script or stylesheet is never intentional. It means the framework
19
+ // produced an inconsistent bundle, and the page is dead in the browser while
20
+ // the pre-rendered HTML still shows. That fails the build.
21
+ // * A missing link or image target is a content mistake, and one a site may make
22
+ // on purpose — a link to a path the web server provides rather than the build.
23
+ // That is reported and left to the site to decide on.
24
+ //
25
+ // All three go into .cratly/build-report.json as well, written before the build is
26
+ // ended rather than after, so what happened survives the failure. A build log is
27
+ // read by whoever has access to it; a file can be handed to whoever wrote the page.
28
+
29
+ import { readFileSync, readdirSync } from "node:fs";
30
+ import { join, posix, resolve } from "node:path";
31
+
32
+ // Where the report goes, and which shape a reader may expect. The version is
33
+ // bumped when a reader would have to be changed — the cratly editor is one.
34
+ export const BUILD_REPORT_DIR = ".cratly";
35
+ export const BUILD_REPORT_FILE = "build-report.json";
36
+ export const BUILD_REPORT_VERSION = 0;
37
+
38
+ // Attributes that name a file the browser must load for the page to work at all.
39
+ const ASSET_ATTR_RE = /<(?:script|link)\b[^>]*?\b(?:src|href)\s*=\s*["']([^"']+)["'][^>]*>/gi;
40
+ const ASSET_TAG_RE = /^<(?:script\b[^>]*\bsrc|link\b[^>]*\brel\s*=\s*["'](?:modulepreload|preload|stylesheet)["'])/i;
41
+ const LINK_RE = /<a\b[^>]*?\bhref\s*=\s*["']([^"']+)["']/gi;
42
+ const IMG_SRC_RE = /<(?:img|source|video)\b[^>]*?\b(?:src|poster)\s*=\s*["']([^"']+)["']/gi;
43
+ const SRCSET_RE = /<(?:img|source)\b[^>]*?\bsrcset\s*=\s*["']([^"']+)["']/gi;
44
+
45
+ // A reference naming its own protocol addresses something outside this build:
46
+ // http:, mailto:, tel: and every other scheme.
47
+ const SCHEME_RE = /^[a-z][a-z0-9+.-]*:/i;
48
+
49
+ /**
50
+ * Tells references that leave the site from those the build has to satisfy.
51
+ *
52
+ * Relative references count. A markdown link that forgot its protocol —
53
+ * `[Website](./www.example.com)` — leaves the build as a relative href pointing
54
+ * at a sibling page nobody wrote, and that is the mistake authors actually make.
55
+ * Where such a reference points is a question only the page holding it can
56
+ * answer, which is why resolving is left to {@link resolveReference}.
57
+ *
58
+ * @param {string} reference value of an href, src or srcset entry
59
+ * @returns {boolean} true when the reference names something this build produced
60
+ */
61
+ function isLocalReference( reference ) {
62
+ return reference !== ""
63
+ && !reference.startsWith( "//" )
64
+ && !SCHEME_RE.test( reference )
65
+ && !reference.includes( "?" );
66
+ }
67
+
68
+ /**
69
+ * Collects the scripts and stylesheets a page cannot work without.
70
+ *
71
+ * @param {string} html contents of one built page
72
+ * @returns {string[]} referenced paths, in document order and de-duplicated
73
+ */
74
+ export function collectAssetRefs( html ) {
75
+ const found = new Set();
76
+
77
+ for ( const [ tag, reference ] of html.matchAll( ASSET_ATTR_RE ) ) {
78
+ if ( ASSET_TAG_RE.test( tag ) && isLocalReference( reference ) ) {
79
+ found.add( reference );
80
+ }
81
+ }
82
+
83
+ return [...found];
84
+ }
85
+
86
+ /**
87
+ * Collects the pages and files a page links to.
88
+ *
89
+ * @param {string} html contents of one built page
90
+ * @returns {string[]} referenced paths, de-duplicated
91
+ */
92
+ export function collectLinkRefs( html ) {
93
+ const found = new Set();
94
+
95
+ for ( const [ , reference ] of html.matchAll( LINK_RE ) ) {
96
+ const [path] = reference.split( "#" );
97
+
98
+ if ( path && isLocalReference( path ) ) {
99
+ found.add( path );
100
+ }
101
+ }
102
+
103
+ return [...found];
104
+ }
105
+
106
+ /**
107
+ * Collects the images and media a page displays.
108
+ *
109
+ * Every candidate of a `srcset` counts: a browser on a wide screen asks for the
110
+ * widest one, so a variant missing from the set breaks the page for some readers
111
+ * and not for others — the worst kind to leave unreported.
112
+ *
113
+ * @param {string} html contents of one built page
114
+ * @returns {string[]} referenced paths, de-duplicated
115
+ */
116
+ export function collectImageRefs( html ) {
117
+ const found = new Set();
118
+
119
+ for ( const [ , reference ] of html.matchAll( IMG_SRC_RE ) ) {
120
+ if ( isLocalReference( reference ) ) {
121
+ found.add( reference );
122
+ }
123
+ }
124
+
125
+ for ( const [ , srcset ] of html.matchAll( SRCSET_RE ) ) {
126
+ for ( const candidate of srcset.split( "," ) ) {
127
+ const [reference] = candidate.trim().split( /\s+/ );
128
+
129
+ if ( reference && isLocalReference( reference ) ) {
130
+ found.add( reference );
131
+ }
132
+ }
133
+ }
134
+
135
+ return [...found];
136
+ }
137
+
138
+ /**
139
+ * Turns a reference as written into the path below the output folder it addresses.
140
+ *
141
+ * A site-absolute reference is read the way the web server reads it, with the
142
+ * site's base path stripped. A relative one means nothing on its own: the browser
143
+ * resolves it against the address of the page it stands in, so that is what this
144
+ * resolves it against too. One that climbs above the site root stays as it is —
145
+ * no file of this build can satisfy it, and saying so is the point.
146
+ *
147
+ * @param {string} reference reference as written in the HTML
148
+ * @param {string} page output path of the page holding it, relative and posix-style
149
+ * @param {string} base site's base path, as configured
150
+ * @returns {string} path relative to the output folder, posix-style and without a leading slash
151
+ */
152
+ function resolveReference( reference, page, base ) {
153
+ const decoded = decodeURI( reference );
154
+
155
+ if ( decoded.startsWith( "/" ) ) {
156
+ const withoutBase = base !== "/" && decoded.startsWith( base )
157
+ ? decoded.slice( base.length - 1 )
158
+ : decoded;
159
+
160
+ return withoutBase.replace( /^\//, "" );
161
+ }
162
+
163
+ return posix.normalize( posix.join( posix.dirname( page ), decoded ) ).replace( /^\//, "" );
164
+ }
165
+
166
+ /**
167
+ * Tells whether a build wrote something a static server could answer a reference with.
168
+ *
169
+ * The three spellings are the ones every static server tries, cratly's own
170
+ * `s3-http` included: the path itself, an `index.html` inside it, and the path
171
+ * plus `.html`.
172
+ *
173
+ * The comparison runs against the names the build actually wrote, never against
174
+ * the file system — macOS and Windows answer `existsSync` case-insensitively, so a
175
+ * page preloading `de_Kontakt.md.js` while only `de_kontakt.md.js` exists would
176
+ * look sound on the machine of whoever wrote the page and break in production.
177
+ * That pair is one of the two ways two pages end up sharing a build identity.
178
+ *
179
+ * @param {string} reference reference as written in the HTML
180
+ * @param {string} page output path of the page holding it, relative and posix-style
181
+ * @param {Set<string>} files paths of every file in the output, relative and posix-style
182
+ * @param {string} base site's base path, as configured
183
+ * @returns {boolean} true when one of those files was written
184
+ */
185
+ function resolvesToFile( reference, page, files, base ) {
186
+ const relativePath = resolveReference( reference, page, base );
187
+
188
+ if ( relativePath === "" || relativePath === "." ) {
189
+ return files.has( "index.html" );
190
+ }
191
+
192
+ return [ relativePath, posix.join( relativePath, "index.html" ), `${relativePath}.html` ]
193
+ .some( candidate => files.has( candidate ) );
194
+ }
195
+
196
+ /**
197
+ * Lists everything a build wrote.
198
+ *
199
+ * @param {string} outDir absolute path of the build output folder
200
+ * @returns {Set<string>} every file below outDir, relative and posix-style
201
+ */
202
+ function collectFiles( outDir ) {
203
+ const files = new Set();
204
+
205
+ /**
206
+ * @param {string} dir folder to descend into
207
+ * @param {string} prefix path of that folder relative to outDir
208
+ * @returns {void}
209
+ */
210
+ function walk( dir, prefix ) {
211
+ for ( const entry of readdirSync( dir, { withFileTypes: true } ) ) {
212
+ const relativePath = prefix ? `${prefix}/${entry.name}` : entry.name;
213
+
214
+ if ( entry.isDirectory() ) {
215
+ walk( join( dir, entry.name ), relativePath );
216
+ } else {
217
+ files.add( relativePath );
218
+ }
219
+ }
220
+ }
221
+
222
+ walk( outDir, "" );
223
+
224
+ return files;
225
+ }
226
+
227
+ /**
228
+ * Reads a finished build and reports what it references but does not contain.
229
+ *
230
+ * @param {string} outDir absolute path of the build output folder
231
+ * @param {object} [options] what to look at
232
+ * @param {string} [options.base] site's base path
233
+ * @param {boolean} [options.content] also check links and images, not just scripts
234
+ * @param {string[]} [options.pages] output files the generator said it would write, relative and posix-style
235
+ * @returns {{pages: object[], assets: object[], content: object[]}}
236
+ * what the output does not contain, split by how much it matters
237
+ */
238
+ export function verifyBuild( outDir, { base = "/", content = true, pages = [] } = {} ) {
239
+ const root = resolve( outDir );
240
+ const files = collectFiles( root );
241
+ const problems = { pages: [], assets: [], content: [] };
242
+
243
+ for ( const page of pages ) {
244
+ if ( !files.has( page ) ) {
245
+ problems.pages.push( { page, reference: page } );
246
+ }
247
+ }
248
+
249
+ for ( const page of [...files].filter( file => file.endsWith( ".html" ) ).sort() ) {
250
+ const html = readFileSync( join( root, page ), "utf-8" );
251
+
252
+ for ( const reference of collectAssetRefs( html ) ) {
253
+ if ( !resolvesToFile( reference, page, files, base ) ) {
254
+ problems.assets.push( { page, reference } );
255
+ }
256
+ }
257
+
258
+ if ( content ) {
259
+ for ( const reference of [ ...collectLinkRefs( html ), ...collectImageRefs( html ) ] ) {
260
+ if ( !resolvesToFile( reference, page, files, base ) ) {
261
+ problems.content.push( { page, reference } );
262
+ }
263
+ }
264
+ }
265
+ }
266
+
267
+ return problems;
268
+ }
269
+
270
+ /**
271
+ * Renders a set of problems as one line per reference, capped so a systematic
272
+ * mistake does not bury the terminal.
273
+ *
274
+ * @param {{page: string, reference: string}[]} problems what verifyBuild found
275
+ * @param {number} [limit] how many to name
276
+ * @returns {string} indented, newline-separated list
277
+ */
278
+ export function formatProblems( problems, limit = 10 ) {
279
+ const lines = problems.slice( 0, limit ).map( ( { page, reference } ) => ` ${page} → ${reference}` );
280
+
281
+ if ( problems.length > limit ) {
282
+ lines.push( ` … and ${problems.length - limit} more` );
283
+ }
284
+
285
+ return lines.join( "\n" );
286
+ }
287
+
288
+ /**
289
+ * Renders what a build found as the report file readers other than a terminal get.
290
+ *
291
+ * `fatal` says which of the two situations the file describes: a build that ended
292
+ * without writing a usable site, or one that finished and carries a content
293
+ * mistake. A reader showing this to whoever wrote the page needs that distinction
294
+ * and should not have to re-derive it from which array is empty.
295
+ *
296
+ * @param {{pages: object[], assets: object[], content: object[]}} problems what verifyBuild found
297
+ * @param {object} [options]
298
+ * @param {object} [options.adapter] name and version of what produced the build
299
+ * @param {Object<string,string>} [options.sources] built page path → source file an author edits
300
+ * @param {Date} [options.generated] when the build ran
301
+ * @returns {string} the file's content, newline-terminated
302
+ */
303
+ export function renderBuildReport( problems, { adapter = {}, sources = {}, generated = new Date() } = {} ) {
304
+ /**
305
+ * @param {{page: string, reference: string}[]} entries one class of problem
306
+ * @returns {object[]} the same entries, each naming its source file where known
307
+ */
308
+ const withSources = entries => entries.map( entry => ( {
309
+ ...entry,
310
+ ...sources[entry.page] != null && { source: sources[entry.page] },
311
+ } ) );
312
+
313
+ return `${JSON.stringify( {
314
+ specVersion: BUILD_REPORT_VERSION,
315
+ adapter,
316
+ generated: generated.toISOString(),
317
+ fatal: problems.pages.length > 0 || problems.assets.length > 0,
318
+ problems: {
319
+ pages: withSources( problems.pages ),
320
+ assets: withSources( problems.assets ),
321
+ content: withSources( problems.content ),
322
+ },
323
+ }, null, "\t" )}\n`;
324
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "scavold",
3
- "version": "0.2.0-rc.5",
3
+ "version": "0.2.0-rc.7",
4
4
  "type": "module",
5
5
  "description": "VitePress theme framework — a scaffold for building custom VitePress themes with Vue at the core",
6
6
  "keywords": [
@@ -18,10 +18,10 @@
18
18
  "license": "MIT",
19
19
  "repository": {
20
20
  "type": "git",
21
- "url": "https://gitlab.com/cepharum-foss/cratly/scavold.git"
21
+ "url": "https://gitlab.com/cratly/scavold.git"
22
22
  },
23
23
  "homepage": "https://scavold.io",
24
- "bugs": "https://gitlab.com/cepharum-foss/cratly/scavold/-/issues",
24
+ "bugs": "https://gitlab.com/cratly/scavold/-/issues",
25
25
  "engines": {
26
26
  "node": ">=20"
27
27
  },
@@ -72,7 +72,7 @@
72
72
  "vue": ">=3.0.0"
73
73
  },
74
74
  "dependencies": {
75
- "@cepharum/vue3-i18n": "^0.5.4",
75
+ "@cepharum/vue3-i18n": "^2.0.0",
76
76
  "markdown-it-container": "^4.0.0",
77
77
  "sharp": "^0.34.5",
78
78
  "yaml": "^2.8.3"
@@ -53,7 +53,7 @@ function linkSelf() {
53
53
 
54
54
  /**
55
55
  * Creates the fixture's media files. Images are real (the pipeline reads their
56
- * dimensions); the other two only ever get copied, so a few bytes carrying the right
56
+ * dimensions); the others only ever get copied, so a few bytes carrying the right
57
57
  * extension are enough.
58
58
  */
59
59
  async function writeMedia() {
@@ -76,6 +76,8 @@ async function writeMedia() {
76
76
 
77
77
  writeFileSync( join( media, "doc.pdf" ), "%PDF-1.4\n% fixture placeholder\n" );
78
78
  writeFileSync( join( media, "clip.mp4" ), "fixture placeholder\n" );
79
+ // An extension VitePress does not know, so its link rules take the file for a page.
80
+ writeFileSync( join( media, "patch.diff" ), "--- a\n+++ b\n" );
79
81
  }
80
82
 
81
83
  /**
@@ -194,6 +196,8 @@ console.log( "\nmedia that is not an image" );
194
196
  expect( published.includes( "doc.pdf" ), "copies a document into the site output" );
195
197
  expect( published.includes( "clip.mp4" ), "copies a video into the site output" );
196
198
  expect( media.includes( 'href="/media/doc.pdf"' ), "rewrites a link to a media file" );
199
+ expect( media.includes( 'href="/media/patch.diff"' ),
200
+ "links a media file of an extension VitePress does not know as that file, not as a page" );
197
201
  // On the entry page, where the fixture's page links live. VitePress resolves them to
198
202
  // .html; what matters is that the media pipeline did not touch them.
199
203
  expect( built( "index.html" ).includes( 'href="/containers.html"' ), "leaves a link to a page alone" );
@@ -321,6 +325,33 @@ expect( [ "Sunday", "21", "June", "2026" ].every( part => ownDate.includes( part
321
325
  "hands a site's own component the timestamp formatting as well",
322
326
  `the entry reads ${JSON.stringify( ownDate )}` );
323
327
 
328
+ // ── build report ─────────────────────────────────────────────────────────────
329
+
330
+ console.log( "\nbuild report" );
331
+
332
+ // pages/dead-link.md links to two files this build does not contain, one written
333
+ // relative and one site-absolute. Neither ends the build — a content mistake must
334
+ // not stand between a finished page and its publication — and both have to be
335
+ // findable afterwards by someone who cannot read a job log.
336
+ expect( /1 dead link|dead link\(s\) found/.test( log ) === false,
337
+ "lets a dead link through instead of ending the build over it" );
338
+ expect( /2 link\(s\) or image\(s\) point at files this build does not contain/.test( log ),
339
+ "names both spellings of a dead link in the build log" );
340
+
341
+ const report = JSON.parse( readFileSync( join( SITE, ".cratly", "build-report.json" ), "utf-8" ) );
342
+
343
+ expect( report.adapter?.version === version, "reports this package's version as the adapter version",
344
+ `found ${report.adapter?.version}` );
345
+ expect( report.fatal === false, "says the build survived what it found" );
346
+ expect( report.problems?.content?.map( p => p.reference ).sort().join( " " ) === "./www.example.com.html /tour.html",
347
+ "carries both dead references",
348
+ `found ${report.problems?.content?.map( p => p.reference ).join( " " ) || "nothing"}` );
349
+ expect( report.problems.content.every( p => p.source === "dead-link.md" ),
350
+ "names the file an author would open, not just the file the build wrote",
351
+ `found ${report.problems.content.map( p => p.source ).join( " " )}` );
352
+ expect( report.problems?.pages?.length === 0 && report.problems?.assets?.length === 0,
353
+ "leaves the fatal classes empty on a build that finished" );
354
+
324
355
  // ── content security policy ──────────────────────────────────────────────────
325
356
 
326
357
  console.log( "\ncontent security policy" );