pagetrace 0.13.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,21 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
  While the version is below 1.0.0, breaking changes ship in a minor release.
8
8
 
9
+ ## [0.14.0] - 2026-09-08
10
+
11
+ ### Added
12
+
13
+ - `--verify-all` also checks links to assets — PDFs, images, archives. They are
14
+ never crawled as pages, so they were skipped entirely: a link to a deleted
15
+ whitepaper reported nothing. A `--dir` run checks them against the filesystem;
16
+ a crawl spends a request each, which is why it is opt-in.
17
+
18
+ ### Changed
19
+
20
+ - Link verification now runs in parallel at `--concurrency`, and the cap rose
21
+ from 100 to 1000 unique targets. Hitting the cap now says so on stderr rather
22
+ than silently under-reporting, which read as a clean site.
23
+
9
24
  ## [0.13.0] - 2026-09-08
10
25
 
11
26
  ### Added
@@ -370,6 +385,7 @@ Initial release. `snapshot`, `check` and `audit` commands; filesystem and HTTP
370
385
  crawling; diff classified by transition; absolute, cross-page and hreflang audit
371
386
  rules; pretty, JSON, markdown, GitHub and HTML reporters.
372
387
 
388
+ [0.14.0]: https://github.com/shyamexe/pagetrace/compare/v0.13.0...v0.14.0
373
389
  [0.13.0]: https://github.com/shyamexe/pagetrace/compare/v0.12.0...v0.13.0
374
390
  [0.12.0]: https://github.com/shyamexe/pagetrace/compare/v0.11.0...v0.12.0
375
391
  [0.11.0]: https://github.com/shyamexe/pagetrace/compare/v0.10.0...v0.11.0
package/README.md CHANGED
@@ -19,6 +19,33 @@ npm install -D pagetrace
19
19
 
20
20
  Requires Node 20.19 or newer. No native modules, three small dependencies.
21
21
 
22
+ ## Commands
23
+
24
+ | Command | Answers | Crawls | Exit 1 when |
25
+ | --- | --- | --- | --- |
26
+ | `init` | "get me set up" | once, to write the first lockfile | never |
27
+ | `snapshot` | "record what the site looks like now" | whole site | never |
28
+ | `check` | "what did this deploy change?" | whole site | findings at or above `--fail-on` (default `error`) |
29
+ | `audit` | "what is wrong with this site?" | whole site | `--fail-on` (default `never`) |
30
+ | `links` | "are any links dead?" | whole site, links only in the report | any broken link |
31
+ | `page <url>` | "is this one page sound?" | that URL alone | `--fail-on` (default `error`) |
32
+ | `update` | "am I on the latest pagetrace?" | nothing | never |
33
+
34
+ Every crawling command takes `--dir <build>` or `--url <origin>`, plus:
35
+
36
+ | Flag | Default | Effect |
37
+ | --- | --- | --- |
38
+ | `--limit <n>` | 200 | Stop after this many pages |
39
+ | `--concurrency <n>` | 5 | Parallel requests |
40
+ | `--external` | off (on for `page`) | Also check links that leave the site |
41
+ | `--verify-all` | off | Also check links to assets — PDFs, images, archives |
42
+ | `--ignore-robots` | off | Crawl paths `robots.txt` disallows |
43
+ | `--fail-on <severity>` | varies | `error`, `warn`, `info` or `never` |
44
+ | `--format <format>` | `pretty` | `pretty`, `json`, `markdown`; `github` and `sarif` on `check` |
45
+ | `--config <file>` | `pagetrace.config.json` | Config file |
46
+
47
+ Exit codes are the same everywhere: `0` clean, `1` findings at or above `--fail-on`, `2` the run itself failed — bad flags, an unreadable build, an unreachable origin. CI can tell "the site regressed" from "the tool broke".
48
+
22
49
  ## Use
23
50
 
24
51
  Set up a config file and the first baseline in one step:
@@ -56,7 +83,7 @@ npx pagetrace check --dir ./out
56
83
  5 error, 9 warning, 2 info
57
84
  ```
58
85
 
59
- Exit codes are `0` for clean, `1` for findings at or above `--fail-on`, and `2` when the run itself failed — bad flags, an unreadable build directory, an unreachable origin. CI can tell "the site regressed" from "the tool broke". `--fail-on` takes `error` (the default), `warn` or `info`, and rejects anything else rather than quietly letting the build pass.
86
+ `--fail-on` rejects an unrecognised value rather than quietly letting the build pass.
60
87
 
61
88
  Note that `check` runs the absolute rules as well as the diff, so it can fail on a problem your build did not introduce. Use `--no-audit` for a pure regression gate.
62
89
 
@@ -193,7 +220,7 @@ Only `404` and `410` count as dead. A `403` from a bot wall, a `429`, a timeout
193
220
 
194
221
  Think twice before putting `--external` in `check`. A third party's bad afternoon becomes a diff in your repository and a red build you cannot fix.
195
222
 
196
- Internal links are checked too. Only the broken ones are stored, so a site's navigation never lands in the lockfile: `link.broken` for a link that is already dead, `link.broken.added` for one this build broke. A `--dir` crawl is authoritative — the build directory is the whole site — while a crawl confirms each candidate with a real request first, because a sitemap routinely omits pages that are live. External links are checked only with `--external`, and only a 404 or 410 counts.
223
+ Internal links are checked too. Only the broken ones are stored, so a site's navigation never lands in the lockfile: `link.broken` for a link that is already dead, `link.broken.added` for one this build broke. A `--dir` crawl is authoritative — the build directory is the whole site — while a crawl confirms each candidate with a real request first, because a sitemap routinely omits pages that are live. External links are checked only with `--external`, and only a 404 or 410 counts. Links to assets — PDFs, images, archives — are skipped unless `--verify-all`, since each one costs a request on a crawl (a `--dir` run checks them against the filesystem instead).
197
224
 
198
225
  Redirects are recorded from the response itself, so they cost no extra requests. A redirect that only adds or drops a trailing slash is server configuration rather than drift and is not reported. `canonical.redirects` is only raised when the canonical's target was actually crawled, so a `--limit` run cannot invent it.
199
226
 
package/dist/cli.cjs CHANGED
@@ -1698,12 +1698,20 @@ async function snapshotFromDir(dir, config = {}) {
1698
1698
  const llms = await (0, import_promises.readFile)((0, import_node_path.join)(dir, "llms.txt"), "utf8").catch(() => null);
1699
1699
  if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
1700
1700
  const origin = site.origin ?? "https://pagetrace.invalid";
1701
- await resolveBrokenLinks(pages, links, origin);
1701
+ await resolveBrokenLinks(pages, links, origin, {
1702
+ includeAssets: config.verifyAll,
1703
+ // The HTML route set is authoritative here, so a missing page is simply
1704
+ // broken. An asset is a file this crawl never walked, so it gets a stat.
1705
+ verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await (0, import_promises.access)((0, import_node_path.join)(dir, target)).then(
1706
+ () => true,
1707
+ () => false
1708
+ ) : true : void 0
1709
+ });
1702
1710
  if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
1703
1711
  return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
1704
1712
  }
1705
1713
  var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
1706
- function linkTarget(href, from, origin) {
1714
+ function linkTarget(href, from, origin, includeAssets = false) {
1707
1715
  let url;
1708
1716
  try {
1709
1717
  url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
@@ -1711,27 +1719,39 @@ function linkTarget(href, from, origin) {
1711
1719
  return null;
1712
1720
  }
1713
1721
  if (url.origin !== origin) return null;
1714
- if (ASSET_PATH.test(url.pathname)) return null;
1715
- return routeFromUrl(url.href);
1722
+ if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
1723
+ return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
1716
1724
  }
1717
- var MAX_LINK_CHECKS = 100;
1718
- async function resolveBrokenLinks(pages, links, origin, verify) {
1725
+ var MAX_LINK_CHECKS = 1e3;
1726
+ async function resolveBrokenLinks(pages, links, origin, options = {}) {
1727
+ const { includeAssets = false, verify, concurrency = 5 } = options;
1719
1728
  const known = new Set(Object.keys(pages));
1720
1729
  const candidates = /* @__PURE__ */ new Map();
1721
1730
  for (const [route, hrefs] of Object.entries(links)) {
1722
1731
  const missing = [
1723
1732
  ...new Set(
1724
- hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
1733
+ hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
1725
1734
  )
1726
1735
  ];
1727
1736
  if (missing.length > 0) candidates.set(route, missing);
1728
1737
  }
1729
1738
  const broken = /* @__PURE__ */ new Set();
1730
1739
  if (verify) {
1731
- const targets = [...new Set([...candidates.values()].flat())].slice(0, MAX_LINK_CHECKS);
1732
- for (const target of targets) {
1733
- if (await verify(target)) broken.add(target);
1740
+ const unique = [...new Set([...candidates.values()].flat())];
1741
+ const queue = unique.slice(0, MAX_LINK_CHECKS);
1742
+ if (unique.length > queue.length) {
1743
+ console.error(
1744
+ `pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
1745
+ );
1734
1746
  }
1747
+ await Promise.all(
1748
+ Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
1749
+ while (queue.length > 0) {
1750
+ const target = queue.shift();
1751
+ if (await verify(target)) broken.add(target);
1752
+ }
1753
+ })
1754
+ );
1735
1755
  }
1736
1756
  for (const [route, missing] of candidates) {
1737
1757
  const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
@@ -1812,11 +1832,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
1812
1832
  const pages = { [route]: page };
1813
1833
  const links = { [route]: extractLinks(doc.text) };
1814
1834
  const origin = site.origin ?? target.origin;
1815
- await resolveBrokenLinks(pages, links, origin, async (candidate) => {
1816
- try {
1817
- return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1818
- } catch {
1819
- return false;
1835
+ await resolveBrokenLinks(pages, links, origin, {
1836
+ includeAssets: options.verifyAll,
1837
+ concurrency: options.concurrency,
1838
+ verify: async (candidate) => {
1839
+ try {
1840
+ return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1841
+ } catch {
1842
+ return false;
1843
+ }
1820
1844
  }
1821
1845
  });
1822
1846
  if (options.checkExternal) {
@@ -1905,11 +1929,15 @@ async function snapshotFromOrigin(origin, options = {}) {
1905
1929
  }
1906
1930
  })
1907
1931
  );
1908
- await resolveBrokenLinks(pages, links, base.origin, async (route) => {
1909
- try {
1910
- return await fetchDoc(new URL(route, base).href, timeout) === null;
1911
- } catch {
1912
- return false;
1932
+ await resolveBrokenLinks(pages, links, base.origin, {
1933
+ includeAssets: options.verifyAll,
1934
+ concurrency,
1935
+ verify: async (route) => {
1936
+ try {
1937
+ return await fetchDoc(new URL(route, base).href, timeout) === null;
1938
+ } catch {
1939
+ return false;
1940
+ }
1913
1941
  }
1914
1942
  });
1915
1943
  if (options.checkExternal) {
@@ -1968,14 +1996,16 @@ async function loadConfig(path = DEFAULT_CONFIG) {
1968
1996
  }
1969
1997
  async function build(flags, config) {
1970
1998
  const checkExternal = flags.external ?? config.checkExternal;
1971
- if (flags.dir) return snapshotFromDir(flags.dir, { ...config, checkExternal });
1999
+ const verifyAll = flags.verifyAll ?? config.verifyAll;
2000
+ if (flags.dir) return snapshotFromDir(flags.dir, { ...config, checkExternal, verifyAll });
1972
2001
  if (flags.url)
1973
2002
  return snapshotFromOrigin(flags.url, {
1974
2003
  ...config,
1975
2004
  limit: flags.limit,
1976
2005
  concurrency: flags.concurrency,
1977
2006
  ignoreRobots: flags.ignoreRobots ?? config.ignoreRobots,
1978
- checkExternal
2007
+ checkExternal,
2008
+ verifyAll
1979
2009
  });
1980
2010
  throw new Error("Provide a source: --dir <build directory> or --url <origin>.");
1981
2011
  }
@@ -1994,7 +2024,7 @@ function render(findings, format) {
1994
2024
  }
1995
2025
  }
1996
2026
  var cli = (0, import_cac.cac)("pagetrace");
1997
- cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
2027
+ cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
1998
2028
  const config = await loadConfig(flags.config);
1999
2029
  const snapshot = await build(flags, config);
2000
2030
  const written = await writeLockfile(flags.out, snapshot);
@@ -2003,7 +2033,7 @@ cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").opti
2003
2033
  written ? import_picocolors2.default.green(`Wrote ${flags.out} \u2014 ${count} page${count === 1 ? "" : "s"}.`) : import_picocolors2.default.dim(`${flags.out} is already up to date \u2014 ${count} page${count === 1 ? "" : "s"}.`)
2004
2034
  );
2005
2035
  });
2006
- cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
2036
+ cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
2007
2037
  const failOn = parseFailOn(flags.failOn, false);
2008
2038
  const config = await loadConfig(flags.config);
2009
2039
  const next = await build(flags, config);
@@ -2033,7 +2063,7 @@ Failing: ${summary.error} error, ${summary.warn} warning (--fail-on ${failOn}).`
2033
2063
  process.exitCode = EXIT_FINDINGS;
2034
2064
  }
2035
2065
  });
2036
- cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
2066
+ cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
2037
2067
  const failOn = parseFailOn(flags.failOn, true);
2038
2068
  const config = await loadConfig(flags.config);
2039
2069
  const snapshot = await build(flags, config);
@@ -2073,13 +2103,14 @@ cli.command("audit", "Audit a site as it stands, with explanations and fixes").o
2073
2103
  process.exitCode = EXIT_FINDINGS;
2074
2104
  }
2075
2105
  });
2076
- cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
2106
+ cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
2077
2107
  const failOn = parseFailOn(flags.failOn, true);
2078
2108
  const config = await loadConfig(flags.config);
2079
2109
  const snapshot = await snapshotFromPage(url, {
2080
2110
  ...config,
2081
2111
  concurrency: flags.concurrency,
2082
- checkExternal: flags.external
2112
+ checkExternal: flags.external,
2113
+ verifyAll: flags.verifyAll ?? config.verifyAll
2083
2114
  });
2084
2115
  const [page] = Object.values(snapshot.pages);
2085
2116
  const platform = detectPlatform([page.generator], Object.values(page.og));
@@ -2109,7 +2140,7 @@ cli.command("page <url>", "Check one page: every link on it, and its own surface
2109
2140
  process.exitCode = EXIT_FINDINGS;
2110
2141
  }
2111
2142
  });
2112
- cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
2143
+ cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
2113
2144
  const failOn = parseFailOn(flags.failOn, true);
2114
2145
  const config = await loadConfig(flags.config);
2115
2146
  const snapshot = await build(flags, config);
@@ -2156,7 +2187,7 @@ cli.command("links", "Find internal links that point at no page").option("--url
2156
2187
  process.exitCode = EXIT_FINDINGS;
2157
2188
  }
2158
2189
  });
2159
- cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
2190
+ cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
2160
2191
  const existing = await (0, import_promises2.readFile)(flags.config, "utf8").catch(() => null);
2161
2192
  if (existing === null) {
2162
2193
  const config = flags.url ? { siteUrl: new URL(flags.url).origin } : {};
@@ -2179,11 +2210,11 @@ Commit both files, then run \`pagetrace check ${flags.dir ? `--dir ${flags.dir}`
2179
2210
  });
2180
2211
  cli.command("update", "Check npm for a newer pagetrace and install it").option("--check", "Only report whether an update exists").action(async (flags) => {
2181
2212
  const latest = await latestVersion("pagetrace");
2182
- if (!isNewer(latest, "0.13.0")) {
2183
- console.log(import_picocolors2.default.green(`pagetrace ${"0.13.0"} is the latest version.`));
2213
+ if (!isNewer(latest, "0.14.0")) {
2214
+ console.log(import_picocolors2.default.green(`pagetrace ${"0.14.0"} is the latest version.`));
2184
2215
  return;
2185
2216
  }
2186
- console.log(import_picocolors2.default.yellow(`Update available: ${"0.13.0"} \u2192 ${latest}`));
2217
+ console.log(import_picocolors2.default.yellow(`Update available: ${"0.14.0"} \u2192 ${latest}`));
2187
2218
  if (flags.check) return;
2188
2219
  if (process.argv[1]?.startsWith(process.cwd())) {
2189
2220
  console.log(
@@ -2201,7 +2232,7 @@ cli.command("update", "Check npm for a newer pagetrace and install it").option("
2201
2232
  console.log(import_picocolors2.default.green(`Updated to pagetrace ${latest}.`));
2202
2233
  });
2203
2234
  cli.help();
2204
- cli.version("0.13.0");
2235
+ cli.version("0.14.0");
2205
2236
  async function main() {
2206
2237
  try {
2207
2238
  cli.parse(process.argv, { run: false });
package/dist/cli.js CHANGED
@@ -1304,7 +1304,7 @@ ${cards || "<p>No issues found.</p>"}
1304
1304
 
1305
1305
  // src/snapshot.ts
1306
1306
  import { execFile } from "child_process";
1307
- import { readdir, readFile } from "fs/promises";
1307
+ import { access, readdir, readFile } from "fs/promises";
1308
1308
  import { join, relative, sep } from "path";
1309
1309
  import { promisify } from "util";
1310
1310
 
@@ -1675,12 +1675,20 @@ async function snapshotFromDir(dir, config = {}) {
1675
1675
  const llms = await readFile(join(dir, "llms.txt"), "utf8").catch(() => null);
1676
1676
  if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
1677
1677
  const origin = site.origin ?? "https://pagetrace.invalid";
1678
- await resolveBrokenLinks(pages, links, origin);
1678
+ await resolveBrokenLinks(pages, links, origin, {
1679
+ includeAssets: config.verifyAll,
1680
+ // The HTML route set is authoritative here, so a missing page is simply
1681
+ // broken. An asset is a file this crawl never walked, so it gets a stat.
1682
+ verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await access(join(dir, target)).then(
1683
+ () => true,
1684
+ () => false
1685
+ ) : true : void 0
1686
+ });
1679
1687
  if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
1680
1688
  return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
1681
1689
  }
1682
1690
  var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
1683
- function linkTarget(href, from, origin) {
1691
+ function linkTarget(href, from, origin, includeAssets = false) {
1684
1692
  let url;
1685
1693
  try {
1686
1694
  url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
@@ -1688,27 +1696,39 @@ function linkTarget(href, from, origin) {
1688
1696
  return null;
1689
1697
  }
1690
1698
  if (url.origin !== origin) return null;
1691
- if (ASSET_PATH.test(url.pathname)) return null;
1692
- return routeFromUrl(url.href);
1699
+ if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
1700
+ return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
1693
1701
  }
1694
- var MAX_LINK_CHECKS = 100;
1695
- async function resolveBrokenLinks(pages, links, origin, verify) {
1702
+ var MAX_LINK_CHECKS = 1e3;
1703
+ async function resolveBrokenLinks(pages, links, origin, options = {}) {
1704
+ const { includeAssets = false, verify, concurrency = 5 } = options;
1696
1705
  const known = new Set(Object.keys(pages));
1697
1706
  const candidates = /* @__PURE__ */ new Map();
1698
1707
  for (const [route, hrefs] of Object.entries(links)) {
1699
1708
  const missing = [
1700
1709
  ...new Set(
1701
- hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
1710
+ hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
1702
1711
  )
1703
1712
  ];
1704
1713
  if (missing.length > 0) candidates.set(route, missing);
1705
1714
  }
1706
1715
  const broken = /* @__PURE__ */ new Set();
1707
1716
  if (verify) {
1708
- const targets = [...new Set([...candidates.values()].flat())].slice(0, MAX_LINK_CHECKS);
1709
- for (const target of targets) {
1710
- if (await verify(target)) broken.add(target);
1717
+ const unique = [...new Set([...candidates.values()].flat())];
1718
+ const queue = unique.slice(0, MAX_LINK_CHECKS);
1719
+ if (unique.length > queue.length) {
1720
+ console.error(
1721
+ `pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
1722
+ );
1711
1723
  }
1724
+ await Promise.all(
1725
+ Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
1726
+ while (queue.length > 0) {
1727
+ const target = queue.shift();
1728
+ if (await verify(target)) broken.add(target);
1729
+ }
1730
+ })
1731
+ );
1712
1732
  }
1713
1733
  for (const [route, missing] of candidates) {
1714
1734
  const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
@@ -1789,11 +1809,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
1789
1809
  const pages = { [route]: page };
1790
1810
  const links = { [route]: extractLinks(doc.text) };
1791
1811
  const origin = site.origin ?? target.origin;
1792
- await resolveBrokenLinks(pages, links, origin, async (candidate) => {
1793
- try {
1794
- return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1795
- } catch {
1796
- return false;
1812
+ await resolveBrokenLinks(pages, links, origin, {
1813
+ includeAssets: options.verifyAll,
1814
+ concurrency: options.concurrency,
1815
+ verify: async (candidate) => {
1816
+ try {
1817
+ return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1818
+ } catch {
1819
+ return false;
1820
+ }
1797
1821
  }
1798
1822
  });
1799
1823
  if (options.checkExternal) {
@@ -1882,11 +1906,15 @@ async function snapshotFromOrigin(origin, options = {}) {
1882
1906
  }
1883
1907
  })
1884
1908
  );
1885
- await resolveBrokenLinks(pages, links, base.origin, async (route) => {
1886
- try {
1887
- return await fetchDoc(new URL(route, base).href, timeout) === null;
1888
- } catch {
1889
- return false;
1909
+ await resolveBrokenLinks(pages, links, base.origin, {
1910
+ includeAssets: options.verifyAll,
1911
+ concurrency,
1912
+ verify: async (route) => {
1913
+ try {
1914
+ return await fetchDoc(new URL(route, base).href, timeout) === null;
1915
+ } catch {
1916
+ return false;
1917
+ }
1890
1918
  }
1891
1919
  });
1892
1920
  if (options.checkExternal) {
@@ -1945,14 +1973,16 @@ async function loadConfig(path = DEFAULT_CONFIG) {
1945
1973
  }
1946
1974
  async function build(flags, config) {
1947
1975
  const checkExternal = flags.external ?? config.checkExternal;
1948
- if (flags.dir) return snapshotFromDir(flags.dir, { ...config, checkExternal });
1976
+ const verifyAll = flags.verifyAll ?? config.verifyAll;
1977
+ if (flags.dir) return snapshotFromDir(flags.dir, { ...config, checkExternal, verifyAll });
1949
1978
  if (flags.url)
1950
1979
  return snapshotFromOrigin(flags.url, {
1951
1980
  ...config,
1952
1981
  limit: flags.limit,
1953
1982
  concurrency: flags.concurrency,
1954
1983
  ignoreRobots: flags.ignoreRobots ?? config.ignoreRobots,
1955
- checkExternal
1984
+ checkExternal,
1985
+ verifyAll
1956
1986
  });
1957
1987
  throw new Error("Provide a source: --dir <build directory> or --url <origin>.");
1958
1988
  }
@@ -1971,7 +2001,7 @@ function render(findings, format) {
1971
2001
  }
1972
2002
  }
1973
2003
  var cli = cac("pagetrace");
1974
- cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
2004
+ cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
1975
2005
  const config = await loadConfig(flags.config);
1976
2006
  const snapshot = await build(flags, config);
1977
2007
  const written = await writeLockfile(flags.out, snapshot);
@@ -1980,7 +2010,7 @@ cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").opti
1980
2010
  written ? pc2.green(`Wrote ${flags.out} \u2014 ${count} page${count === 1 ? "" : "s"}.`) : pc2.dim(`${flags.out} is already up to date \u2014 ${count} page${count === 1 ? "" : "s"}.`)
1981
2011
  );
1982
2012
  });
1983
- cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
2013
+ cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
1984
2014
  const failOn = parseFailOn(flags.failOn, false);
1985
2015
  const config = await loadConfig(flags.config);
1986
2016
  const next = await build(flags, config);
@@ -2010,7 +2040,7 @@ Failing: ${summary.error} error, ${summary.warn} warning (--fail-on ${failOn}).`
2010
2040
  process.exitCode = EXIT_FINDINGS;
2011
2041
  }
2012
2042
  });
2013
- cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
2043
+ cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
2014
2044
  const failOn = parseFailOn(flags.failOn, true);
2015
2045
  const config = await loadConfig(flags.config);
2016
2046
  const snapshot = await build(flags, config);
@@ -2050,13 +2080,14 @@ cli.command("audit", "Audit a site as it stands, with explanations and fixes").o
2050
2080
  process.exitCode = EXIT_FINDINGS;
2051
2081
  }
2052
2082
  });
2053
- cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
2083
+ cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
2054
2084
  const failOn = parseFailOn(flags.failOn, true);
2055
2085
  const config = await loadConfig(flags.config);
2056
2086
  const snapshot = await snapshotFromPage(url, {
2057
2087
  ...config,
2058
2088
  concurrency: flags.concurrency,
2059
- checkExternal: flags.external
2089
+ checkExternal: flags.external,
2090
+ verifyAll: flags.verifyAll ?? config.verifyAll
2060
2091
  });
2061
2092
  const [page] = Object.values(snapshot.pages);
2062
2093
  const platform = detectPlatform([page.generator], Object.values(page.og));
@@ -2086,7 +2117,7 @@ cli.command("page <url>", "Check one page: every link on it, and its own surface
2086
2117
  process.exitCode = EXIT_FINDINGS;
2087
2118
  }
2088
2119
  });
2089
- cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
2120
+ cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
2090
2121
  const failOn = parseFailOn(flags.failOn, true);
2091
2122
  const config = await loadConfig(flags.config);
2092
2123
  const snapshot = await build(flags, config);
@@ -2133,7 +2164,7 @@ cli.command("links", "Find internal links that point at no page").option("--url
2133
2164
  process.exitCode = EXIT_FINDINGS;
2134
2165
  }
2135
2166
  });
2136
- cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
2167
+ cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
2137
2168
  const existing = await readFile2(flags.config, "utf8").catch(() => null);
2138
2169
  if (existing === null) {
2139
2170
  const config = flags.url ? { siteUrl: new URL(flags.url).origin } : {};
@@ -2156,11 +2187,11 @@ Commit both files, then run \`pagetrace check ${flags.dir ? `--dir ${flags.dir}`
2156
2187
  });
2157
2188
  cli.command("update", "Check npm for a newer pagetrace and install it").option("--check", "Only report whether an update exists").action(async (flags) => {
2158
2189
  const latest = await latestVersion("pagetrace");
2159
- if (!isNewer(latest, "0.13.0")) {
2160
- console.log(pc2.green(`pagetrace ${"0.13.0"} is the latest version.`));
2190
+ if (!isNewer(latest, "0.14.0")) {
2191
+ console.log(pc2.green(`pagetrace ${"0.14.0"} is the latest version.`));
2161
2192
  return;
2162
2193
  }
2163
- console.log(pc2.yellow(`Update available: ${"0.13.0"} \u2192 ${latest}`));
2194
+ console.log(pc2.yellow(`Update available: ${"0.14.0"} \u2192 ${latest}`));
2164
2195
  if (flags.check) return;
2165
2196
  if (process.argv[1]?.startsWith(process.cwd())) {
2166
2197
  console.log(
@@ -2178,7 +2209,7 @@ cli.command("update", "Check npm for a newer pagetrace and install it").option("
2178
2209
  console.log(pc2.green(`Updated to pagetrace ${latest}.`));
2179
2210
  });
2180
2211
  cli.help();
2181
- cli.version("0.13.0");
2212
+ cli.version("0.14.0");
2182
2213
  async function main() {
2183
2214
  try {
2184
2215
  cli.parse(process.argv, { run: false });
package/dist/index.cjs CHANGED
@@ -1740,12 +1740,20 @@ async function snapshotFromDir(dir, config = {}) {
1740
1740
  const llms = await (0, import_promises.readFile)((0, import_node_path.join)(dir, "llms.txt"), "utf8").catch(() => null);
1741
1741
  if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
1742
1742
  const origin = site.origin ?? "https://pagetrace.invalid";
1743
- await resolveBrokenLinks(pages, links, origin);
1743
+ await resolveBrokenLinks(pages, links, origin, {
1744
+ includeAssets: config.verifyAll,
1745
+ // The HTML route set is authoritative here, so a missing page is simply
1746
+ // broken. An asset is a file this crawl never walked, so it gets a stat.
1747
+ verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await (0, import_promises.access)((0, import_node_path.join)(dir, target)).then(
1748
+ () => true,
1749
+ () => false
1750
+ ) : true : void 0
1751
+ });
1744
1752
  if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
1745
1753
  return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
1746
1754
  }
1747
1755
  var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
1748
- function linkTarget(href, from, origin) {
1756
+ function linkTarget(href, from, origin, includeAssets = false) {
1749
1757
  let url;
1750
1758
  try {
1751
1759
  url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
@@ -1753,27 +1761,39 @@ function linkTarget(href, from, origin) {
1753
1761
  return null;
1754
1762
  }
1755
1763
  if (url.origin !== origin) return null;
1756
- if (ASSET_PATH.test(url.pathname)) return null;
1757
- return routeFromUrl(url.href);
1764
+ if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
1765
+ return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
1758
1766
  }
1759
- var MAX_LINK_CHECKS = 100;
1760
- async function resolveBrokenLinks(pages, links, origin, verify) {
1767
+ var MAX_LINK_CHECKS = 1e3;
1768
+ async function resolveBrokenLinks(pages, links, origin, options = {}) {
1769
+ const { includeAssets = false, verify, concurrency = 5 } = options;
1761
1770
  const known = new Set(Object.keys(pages));
1762
1771
  const candidates = /* @__PURE__ */ new Map();
1763
1772
  for (const [route, hrefs] of Object.entries(links)) {
1764
1773
  const missing = [
1765
1774
  ...new Set(
1766
- hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
1775
+ hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
1767
1776
  )
1768
1777
  ];
1769
1778
  if (missing.length > 0) candidates.set(route, missing);
1770
1779
  }
1771
1780
  const broken = /* @__PURE__ */ new Set();
1772
1781
  if (verify) {
1773
- const targets = [...new Set([...candidates.values()].flat())].slice(0, MAX_LINK_CHECKS);
1774
- for (const target of targets) {
1775
- if (await verify(target)) broken.add(target);
1782
+ const unique = [...new Set([...candidates.values()].flat())];
1783
+ const queue = unique.slice(0, MAX_LINK_CHECKS);
1784
+ if (unique.length > queue.length) {
1785
+ console.error(
1786
+ `pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
1787
+ );
1776
1788
  }
1789
+ await Promise.all(
1790
+ Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
1791
+ while (queue.length > 0) {
1792
+ const target = queue.shift();
1793
+ if (await verify(target)) broken.add(target);
1794
+ }
1795
+ })
1796
+ );
1777
1797
  }
1778
1798
  for (const [route, missing] of candidates) {
1779
1799
  const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
@@ -1854,11 +1874,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
1854
1874
  const pages = { [route]: page };
1855
1875
  const links = { [route]: extractLinks(doc.text) };
1856
1876
  const origin = site.origin ?? target.origin;
1857
- await resolveBrokenLinks(pages, links, origin, async (candidate) => {
1858
- try {
1859
- return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1860
- } catch {
1861
- return false;
1877
+ await resolveBrokenLinks(pages, links, origin, {
1878
+ includeAssets: options.verifyAll,
1879
+ concurrency: options.concurrency,
1880
+ verify: async (candidate) => {
1881
+ try {
1882
+ return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1883
+ } catch {
1884
+ return false;
1885
+ }
1862
1886
  }
1863
1887
  });
1864
1888
  if (options.checkExternal) {
@@ -1947,11 +1971,15 @@ async function snapshotFromOrigin(origin, options = {}) {
1947
1971
  }
1948
1972
  })
1949
1973
  );
1950
- await resolveBrokenLinks(pages, links, base.origin, async (route) => {
1951
- try {
1952
- return await fetchDoc(new URL(route, base).href, timeout) === null;
1953
- } catch {
1954
- return false;
1974
+ await resolveBrokenLinks(pages, links, base.origin, {
1975
+ includeAssets: options.verifyAll,
1976
+ concurrency,
1977
+ verify: async (route) => {
1978
+ try {
1979
+ return await fetchDoc(new URL(route, base).href, timeout) === null;
1980
+ } catch {
1981
+ return false;
1982
+ }
1955
1983
  }
1956
1984
  });
1957
1985
  if (options.checkExternal) {
package/dist/index.d.cts CHANGED
@@ -151,6 +151,11 @@ interface Config {
151
151
  * third party's bad afternoon into a diff in your repository.
152
152
  */
153
153
  checkExternal?: boolean;
154
+ /**
155
+ * Also check links to assets — PDFs, images, archives. They are never crawled
156
+ * as pages, so each one costs a request (or a filesystem check for --dir).
157
+ */
158
+ verifyAll?: boolean;
154
159
  }
155
160
 
156
161
  /**
package/dist/index.d.ts CHANGED
@@ -151,6 +151,11 @@ interface Config {
151
151
  * third party's bad afternoon into a diff in your repository.
152
152
  */
153
153
  checkExternal?: boolean;
154
+ /**
155
+ * Also check links to assets — PDFs, images, archives. They are never crawled
156
+ * as pages, so each one costs a request (or a filesystem check for --dir).
157
+ */
158
+ verifyAll?: boolean;
154
159
  }
155
160
 
156
161
  /**
package/dist/index.js CHANGED
@@ -1535,7 +1535,7 @@ ${cards || "<p>No issues found.</p>"}
1535
1535
 
1536
1536
  // src/snapshot.ts
1537
1537
  import { execFile } from "child_process";
1538
- import { readdir, readFile } from "fs/promises";
1538
+ import { access, readdir, readFile } from "fs/promises";
1539
1539
  import { join, relative, sep } from "path";
1540
1540
  import { promisify } from "util";
1541
1541
  function routeFromFilePath(root, filePath) {
@@ -1665,12 +1665,20 @@ async function snapshotFromDir(dir, config = {}) {
1665
1665
  const llms = await readFile(join(dir, "llms.txt"), "utf8").catch(() => null);
1666
1666
  if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
1667
1667
  const origin = site.origin ?? "https://pagetrace.invalid";
1668
- await resolveBrokenLinks(pages, links, origin);
1668
+ await resolveBrokenLinks(pages, links, origin, {
1669
+ includeAssets: config.verifyAll,
1670
+ // The HTML route set is authoritative here, so a missing page is simply
1671
+ // broken. An asset is a file this crawl never walked, so it gets a stat.
1672
+ verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await access(join(dir, target)).then(
1673
+ () => true,
1674
+ () => false
1675
+ ) : true : void 0
1676
+ });
1669
1677
  if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
1670
1678
  return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
1671
1679
  }
1672
1680
  var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
1673
- function linkTarget(href, from, origin) {
1681
+ function linkTarget(href, from, origin, includeAssets = false) {
1674
1682
  let url;
1675
1683
  try {
1676
1684
  url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
@@ -1678,27 +1686,39 @@ function linkTarget(href, from, origin) {
1678
1686
  return null;
1679
1687
  }
1680
1688
  if (url.origin !== origin) return null;
1681
- if (ASSET_PATH.test(url.pathname)) return null;
1682
- return routeFromUrl(url.href);
1689
+ if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
1690
+ return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
1683
1691
  }
1684
- var MAX_LINK_CHECKS = 100;
1685
- async function resolveBrokenLinks(pages, links, origin, verify) {
1692
+ var MAX_LINK_CHECKS = 1e3;
1693
+ async function resolveBrokenLinks(pages, links, origin, options = {}) {
1694
+ const { includeAssets = false, verify, concurrency = 5 } = options;
1686
1695
  const known = new Set(Object.keys(pages));
1687
1696
  const candidates = /* @__PURE__ */ new Map();
1688
1697
  for (const [route, hrefs] of Object.entries(links)) {
1689
1698
  const missing = [
1690
1699
  ...new Set(
1691
- hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
1700
+ hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
1692
1701
  )
1693
1702
  ];
1694
1703
  if (missing.length > 0) candidates.set(route, missing);
1695
1704
  }
1696
1705
  const broken = /* @__PURE__ */ new Set();
1697
1706
  if (verify) {
1698
- const targets = [...new Set([...candidates.values()].flat())].slice(0, MAX_LINK_CHECKS);
1699
- for (const target of targets) {
1700
- if (await verify(target)) broken.add(target);
1707
+ const unique = [...new Set([...candidates.values()].flat())];
1708
+ const queue = unique.slice(0, MAX_LINK_CHECKS);
1709
+ if (unique.length > queue.length) {
1710
+ console.error(
1711
+ `pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
1712
+ );
1701
1713
  }
1714
+ await Promise.all(
1715
+ Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
1716
+ while (queue.length > 0) {
1717
+ const target = queue.shift();
1718
+ if (await verify(target)) broken.add(target);
1719
+ }
1720
+ })
1721
+ );
1702
1722
  }
1703
1723
  for (const [route, missing] of candidates) {
1704
1724
  const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
@@ -1779,11 +1799,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
1779
1799
  const pages = { [route]: page };
1780
1800
  const links = { [route]: extractLinks(doc.text) };
1781
1801
  const origin = site.origin ?? target.origin;
1782
- await resolveBrokenLinks(pages, links, origin, async (candidate) => {
1783
- try {
1784
- return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1785
- } catch {
1786
- return false;
1802
+ await resolveBrokenLinks(pages, links, origin, {
1803
+ includeAssets: options.verifyAll,
1804
+ concurrency: options.concurrency,
1805
+ verify: async (candidate) => {
1806
+ try {
1807
+ return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
1808
+ } catch {
1809
+ return false;
1810
+ }
1787
1811
  }
1788
1812
  });
1789
1813
  if (options.checkExternal) {
@@ -1872,11 +1896,15 @@ async function snapshotFromOrigin(origin, options = {}) {
1872
1896
  }
1873
1897
  })
1874
1898
  );
1875
- await resolveBrokenLinks(pages, links, base.origin, async (route) => {
1876
- try {
1877
- return await fetchDoc(new URL(route, base).href, timeout) === null;
1878
- } catch {
1879
- return false;
1899
+ await resolveBrokenLinks(pages, links, base.origin, {
1900
+ includeAssets: options.verifyAll,
1901
+ concurrency,
1902
+ verify: async (route) => {
1903
+ try {
1904
+ return await fetchDoc(new URL(route, base).href, timeout) === null;
1905
+ } catch {
1906
+ return false;
1907
+ }
1880
1908
  }
1881
1909
  });
1882
1910
  if (options.checkExternal) {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pagetrace",
3
- "version": "0.13.0",
3
+ "version": "0.14.0",
4
4
  "description": "Baseline your site's SEO and AEO surface, diff every build against it, and fail CI on regressions.",
5
5
  "main": "./dist/index.cjs",
6
6
  "scripts": {