pagetrace 0.13.0 → 0.14.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +100 -0
- package/README.md +33 -2
- package/dist/cli.cjs +65 -34
- package/dist/cli.js +66 -35
- package/dist/index.cjs +49 -21
- package/dist/index.d.cts +5 -0
- package/dist/index.d.ts +5 -0
- package/dist/index.js +50 -22
- package/package.json +3 -2
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,100 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
|
6
6
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
7
|
While the version is below 1.0.0, breaking changes ship in a minor release.
|
|
8
8
|
|
|
9
|
+
## [0.14.5] - 2026-09-08
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- `extractLlmsTxt` used `Buffer.byteLength`, which does not exist in a browser,
|
|
14
|
+
so any page on a site that actually has an `/llms.txt` crashed the checker
|
|
15
|
+
with "Buffer is not defined". It counts UTF-8 bytes with `TextEncoder` now —
|
|
16
|
+
the same number, available everywhere. `extract.ts` is meant to be pure and
|
|
17
|
+
bundleable, and this was the one line that was not.
|
|
18
|
+
- The scan log blamed the fetch for failures that happened after it, because
|
|
19
|
+
the error handler settled a fixed step rather than the one still running.
|
|
20
|
+
|
|
21
|
+
## [0.14.4] - 2026-09-08
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- Whole-site mode in the checker. It crawls from the sitemap the way
|
|
26
|
+
`pagetrace audit --url` does and runs every rule over the result, so the
|
|
27
|
+
cross-page findings a single page cannot produce — duplicate titles, hreflang
|
|
28
|
+
reciprocity, dead internal links, sitemap health — finally appear on the site.
|
|
29
|
+
Findings are grouped by code and message and list the routes they touch, which
|
|
30
|
+
is how the CLI's own report aggregates.
|
|
31
|
+
|
|
32
|
+
### Fixed
|
|
33
|
+
|
|
34
|
+
- Every request now has an eight-second deadline. A proxy that is down hangs
|
|
35
|
+
rather than refusing, and the chain sat waiting on it.
|
|
36
|
+
- robots.txt and sitemap.xml were fetched through a reader that renders what it
|
|
37
|
+
fetches, so they came back as HTML documents and stopped looking like
|
|
38
|
+
themselves. Those requests now ask for text, and a sitemap whose XML did not
|
|
39
|
+
survive has its URLs recovered from the body rather than being given up on.
|
|
40
|
+
- Switching scope left the other scope's score standing under a heading that no
|
|
41
|
+
longer described it.
|
|
42
|
+
|
|
43
|
+
## [0.14.3] - 2026-09-08
|
|
44
|
+
|
|
45
|
+
### Fixed
|
|
46
|
+
|
|
47
|
+
- The checker could not fetch anything: it depended on one public CORS proxy,
|
|
48
|
+
and that proxy was returning Cloudflare 520s for every URL. It now tries
|
|
49
|
+
r.jina.ai, allorigins and codetabs in order and takes the first that answers
|
|
50
|
+
with HTML, so one service having a bad day no longer takes the feature down.
|
|
51
|
+
|
|
52
|
+
### Added
|
|
53
|
+
|
|
54
|
+
- A scan log. The check narrates itself as it runs — each request, its size and
|
|
55
|
+
which proxy served it, then the rule pass — instead of leaving a dead page
|
|
56
|
+
during the seconds it takes. A progress bar tracks the steps, the report dims
|
|
57
|
+
while it works, the score counts up to its value, and a category holding a
|
|
58
|
+
blocker gets a coloured edge.
|
|
59
|
+
- Example URLs to try, so the page can be used without owning a site.
|
|
60
|
+
|
|
61
|
+
## [0.14.2] - 2026-09-08
|
|
62
|
+
|
|
63
|
+
### Added
|
|
64
|
+
|
|
65
|
+
- An SEO checker page on the website: enter a URL and get all twenty-three
|
|
66
|
+
page-level checks scored and grouped, passes shown alongside failures. It
|
|
67
|
+
fetches robots.txt and llms.txt too, so the site-level checks answer from
|
|
68
|
+
evidence; anything not fetched reads "not checked" rather than passing or
|
|
69
|
+
failing.
|
|
70
|
+
|
|
71
|
+
### Fixed
|
|
72
|
+
|
|
73
|
+
- `npm run build:site` wrote into `docs/`, which tsup cleans, so building the
|
|
74
|
+
site deleted the site. Its output now has its own directory, `docs/assets/`,
|
|
75
|
+
which is the only thing the build can remove. The Pages deploy would have
|
|
76
|
+
published a lone JavaScript file.
|
|
77
|
+
|
|
78
|
+
## [0.14.1] - 2026-09-08
|
|
79
|
+
|
|
80
|
+
### Added
|
|
81
|
+
|
|
82
|
+
- A website at [shyamexe.github.io/pagetrace](https://shyamexe.github.io/pagetrace/), built from
|
|
83
|
+
`docs/` and deployed by GitHub Pages. Its playground compiles `extract`, `audit` and `diff` into
|
|
84
|
+
the page via a new `src/browser.ts` entry, so it runs the same rules as the CLI rather than a
|
|
85
|
+
reimplementation of them. `npm run build:site` produces the bundle, and the Pages workflow rebuilds
|
|
86
|
+
it on every deploy so it cannot drift from the rules it claims to run.
|
|
87
|
+
|
|
88
|
+
## [0.14.0] - 2026-09-08
|
|
89
|
+
|
|
90
|
+
### Added
|
|
91
|
+
|
|
92
|
+
- `--verify-all` also checks links to assets — PDFs, images, archives. They are
|
|
93
|
+
never crawled as pages, so they were skipped entirely: a link to a deleted
|
|
94
|
+
whitepaper reported nothing. A `--dir` run checks them against the filesystem;
|
|
95
|
+
a crawl spends a request each, which is why it is opt-in.
|
|
96
|
+
|
|
97
|
+
### Changed
|
|
98
|
+
|
|
99
|
+
- Link verification now runs in parallel at `--concurrency`, and the cap rose
|
|
100
|
+
from 100 to 1000 unique targets. Hitting the cap now says so on stderr rather
|
|
101
|
+
than silently under-reporting, which read as a clean site.
|
|
102
|
+
|
|
9
103
|
## [0.13.0] - 2026-09-08
|
|
10
104
|
|
|
11
105
|
### Added
|
|
@@ -370,6 +464,12 @@ Initial release. `snapshot`, `check` and `audit` commands; filesystem and HTTP
|
|
|
370
464
|
crawling; diff classified by transition; absolute, cross-page and hreflang audit
|
|
371
465
|
rules; pretty, JSON, markdown, GitHub and HTML reporters.
|
|
372
466
|
|
|
467
|
+
[0.14.5]: https://github.com/shyamexe/pagetrace/compare/v0.14.0...v0.14.5
|
|
468
|
+
[0.14.4]: https://github.com/shyamexe/pagetrace/commit/ec2a5e2cd590
|
|
469
|
+
[0.14.3]: https://github.com/shyamexe/pagetrace/commit/c45cfb49e265
|
|
470
|
+
[0.14.2]: https://github.com/shyamexe/pagetrace/commit/b3aa77388ee9
|
|
471
|
+
[0.14.1]: https://github.com/shyamexe/pagetrace/commit/7cc91383ed7a
|
|
472
|
+
[0.14.0]: https://github.com/shyamexe/pagetrace/compare/v0.13.0...v0.14.0
|
|
373
473
|
[0.13.0]: https://github.com/shyamexe/pagetrace/compare/v0.12.0...v0.13.0
|
|
374
474
|
[0.12.0]: https://github.com/shyamexe/pagetrace/compare/v0.11.0...v0.12.0
|
|
375
475
|
[0.11.0]: https://github.com/shyamexe/pagetrace/compare/v0.10.0...v0.11.0
|
package/README.md
CHANGED
|
@@ -19,6 +19,37 @@ npm install -D pagetrace
|
|
|
19
19
|
|
|
20
20
|
Requires Node 20.19 or newer. No native modules, three small dependencies.
|
|
21
21
|
|
|
22
|
+
## Playground
|
|
23
|
+
|
|
24
|
+
Try the rules without installing anything: **[shyamexe.github.io/pagetrace](https://shyamexe.github.io/pagetrace/)**. The site compiles `extract`, `audit` and `diff` — the pure half of this package — into the page, so it runs the same rules the CLI does, in your browser, with nothing uploaded.
|
|
25
|
+
|
|
26
|
+
## Commands
|
|
27
|
+
|
|
28
|
+
| Command | Answers | Crawls | Exit 1 when |
|
|
29
|
+
| --- | --- | --- | --- |
|
|
30
|
+
| `init` | "get me set up" | once, to write the first lockfile | never |
|
|
31
|
+
| `snapshot` | "record what the site looks like now" | whole site | never |
|
|
32
|
+
| `check` | "what did this deploy change?" | whole site | findings at or above `--fail-on` (default `error`) |
|
|
33
|
+
| `audit` | "what is wrong with this site?" | whole site | `--fail-on` (default `never`) |
|
|
34
|
+
| `links` | "are any links dead?" | whole site, links only in the report | any broken link |
|
|
35
|
+
| `page <url>` | "is this one page sound?" | that URL alone | `--fail-on` (default `error`) |
|
|
36
|
+
| `update` | "am I on the latest pagetrace?" | nothing | never |
|
|
37
|
+
|
|
38
|
+
Every crawling command takes `--dir <build>` or `--url <origin>`, plus:
|
|
39
|
+
|
|
40
|
+
| Flag | Default | Effect |
|
|
41
|
+
| --- | --- | --- |
|
|
42
|
+
| `--limit <n>` | 200 | Stop after this many pages |
|
|
43
|
+
| `--concurrency <n>` | 5 | Parallel requests |
|
|
44
|
+
| `--external` | off (on for `page`) | Also check links that leave the site |
|
|
45
|
+
| `--verify-all` | off | Also check links to assets — PDFs, images, archives |
|
|
46
|
+
| `--ignore-robots` | off | Crawl paths `robots.txt` disallows |
|
|
47
|
+
| `--fail-on <severity>` | varies | `error`, `warn`, `info` or `never` |
|
|
48
|
+
| `--format <format>` | `pretty` | `pretty`, `json`, `markdown`; `github` and `sarif` on `check` |
|
|
49
|
+
| `--config <file>` | `pagetrace.config.json` | Config file |
|
|
50
|
+
|
|
51
|
+
Exit codes are the same everywhere: `0` clean, `1` findings at or above `--fail-on`, `2` the run itself failed — bad flags, an unreadable build, an unreachable origin. CI can tell "the site regressed" from "the tool broke".
|
|
52
|
+
|
|
22
53
|
## Use
|
|
23
54
|
|
|
24
55
|
Set up a config file and the first baseline in one step:
|
|
@@ -56,7 +87,7 @@ npx pagetrace check --dir ./out
|
|
|
56
87
|
5 error, 9 warning, 2 info
|
|
57
88
|
```
|
|
58
89
|
|
|
59
|
-
|
|
90
|
+
`--fail-on` rejects an unrecognised value rather than quietly letting the build pass.
|
|
60
91
|
|
|
61
92
|
Note that `check` runs the absolute rules as well as the diff, so it can fail on a problem your build did not introduce. Use `--no-audit` for a pure regression gate.
|
|
62
93
|
|
|
@@ -193,7 +224,7 @@ Only `404` and `410` count as dead. A `403` from a bot wall, a `429`, a timeout
|
|
|
193
224
|
|
|
194
225
|
Think twice before putting `--external` in `check`. A third party's bad afternoon becomes a diff in your repository and a red build you cannot fix.
|
|
195
226
|
|
|
196
|
-
Internal links are checked too. Only the broken ones are stored, so a site's navigation never lands in the lockfile: `link.broken` for a link that is already dead, `link.broken.added` for one this build broke. A `--dir` crawl is authoritative — the build directory is the whole site — while a crawl confirms each candidate with a real request first, because a sitemap routinely omits pages that are live. External links are checked only with `--external`, and only a 404 or 410 counts.
|
|
227
|
+
Internal links are checked too. Only the broken ones are stored, so a site's navigation never lands in the lockfile: `link.broken` for a link that is already dead, `link.broken.added` for one this build broke. A `--dir` crawl is authoritative — the build directory is the whole site — while a crawl confirms each candidate with a real request first, because a sitemap routinely omits pages that are live. External links are checked only with `--external`, and only a 404 or 410 counts. Links to assets — PDFs, images, archives — are skipped unless `--verify-all`, since each one costs a request on a crawl (a `--dir` run checks them against the filesystem instead).
|
|
197
228
|
|
|
198
229
|
Redirects are recorded from the response itself, so they cost no extra requests. A redirect that only adds or drops a trailing slash is server configuration rather than drift and is not reported. `canonical.redirects` is only raised when the canonical's target was actually crawled, so a `--limit` run cannot invent it.
|
|
199
230
|
|
package/dist/cli.cjs
CHANGED
|
@@ -1549,7 +1549,7 @@ function extractLinks(html) {
|
|
|
1549
1549
|
}
|
|
1550
1550
|
function extractLlmsTxt(body) {
|
|
1551
1551
|
const sections = body.split(/\r?\n/).filter((line) => line.startsWith("## ")).map((line) => line.slice(3).trim());
|
|
1552
|
-
return { present: true, sections, bytes:
|
|
1552
|
+
return { present: true, sections, bytes: new TextEncoder().encode(body).length };
|
|
1553
1553
|
}
|
|
1554
1554
|
var XML_ENTITIES = {
|
|
1555
1555
|
amp: "&",
|
|
@@ -1698,12 +1698,20 @@ async function snapshotFromDir(dir, config = {}) {
|
|
|
1698
1698
|
const llms = await (0, import_promises.readFile)((0, import_node_path.join)(dir, "llms.txt"), "utf8").catch(() => null);
|
|
1699
1699
|
if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
|
|
1700
1700
|
const origin = site.origin ?? "https://pagetrace.invalid";
|
|
1701
|
-
await resolveBrokenLinks(pages, links, origin
|
|
1701
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1702
|
+
includeAssets: config.verifyAll,
|
|
1703
|
+
// The HTML route set is authoritative here, so a missing page is simply
|
|
1704
|
+
// broken. An asset is a file this crawl never walked, so it gets a stat.
|
|
1705
|
+
verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await (0, import_promises.access)((0, import_node_path.join)(dir, target)).then(
|
|
1706
|
+
() => true,
|
|
1707
|
+
() => false
|
|
1708
|
+
) : true : void 0
|
|
1709
|
+
});
|
|
1702
1710
|
if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
|
|
1703
1711
|
return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
|
|
1704
1712
|
}
|
|
1705
1713
|
var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
|
|
1706
|
-
function linkTarget(href, from, origin) {
|
|
1714
|
+
function linkTarget(href, from, origin, includeAssets = false) {
|
|
1707
1715
|
let url;
|
|
1708
1716
|
try {
|
|
1709
1717
|
url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
|
|
@@ -1711,27 +1719,39 @@ function linkTarget(href, from, origin) {
|
|
|
1711
1719
|
return null;
|
|
1712
1720
|
}
|
|
1713
1721
|
if (url.origin !== origin) return null;
|
|
1714
|
-
if (ASSET_PATH.test(url.pathname)) return null;
|
|
1715
|
-
return routeFromUrl(url.href);
|
|
1722
|
+
if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
|
|
1723
|
+
return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
|
|
1716
1724
|
}
|
|
1717
|
-
var MAX_LINK_CHECKS =
|
|
1718
|
-
async function resolveBrokenLinks(pages, links, origin,
|
|
1725
|
+
var MAX_LINK_CHECKS = 1e3;
|
|
1726
|
+
async function resolveBrokenLinks(pages, links, origin, options = {}) {
|
|
1727
|
+
const { includeAssets = false, verify, concurrency = 5 } = options;
|
|
1719
1728
|
const known = new Set(Object.keys(pages));
|
|
1720
1729
|
const candidates = /* @__PURE__ */ new Map();
|
|
1721
1730
|
for (const [route, hrefs] of Object.entries(links)) {
|
|
1722
1731
|
const missing = [
|
|
1723
1732
|
...new Set(
|
|
1724
|
-
hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
|
|
1733
|
+
hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
|
|
1725
1734
|
)
|
|
1726
1735
|
];
|
|
1727
1736
|
if (missing.length > 0) candidates.set(route, missing);
|
|
1728
1737
|
}
|
|
1729
1738
|
const broken = /* @__PURE__ */ new Set();
|
|
1730
1739
|
if (verify) {
|
|
1731
|
-
const
|
|
1732
|
-
|
|
1733
|
-
|
|
1740
|
+
const unique = [...new Set([...candidates.values()].flat())];
|
|
1741
|
+
const queue = unique.slice(0, MAX_LINK_CHECKS);
|
|
1742
|
+
if (unique.length > queue.length) {
|
|
1743
|
+
console.error(
|
|
1744
|
+
`pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
|
|
1745
|
+
);
|
|
1734
1746
|
}
|
|
1747
|
+
await Promise.all(
|
|
1748
|
+
Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
|
|
1749
|
+
while (queue.length > 0) {
|
|
1750
|
+
const target = queue.shift();
|
|
1751
|
+
if (await verify(target)) broken.add(target);
|
|
1752
|
+
}
|
|
1753
|
+
})
|
|
1754
|
+
);
|
|
1735
1755
|
}
|
|
1736
1756
|
for (const [route, missing] of candidates) {
|
|
1737
1757
|
const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
|
|
@@ -1812,11 +1832,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
|
|
|
1812
1832
|
const pages = { [route]: page };
|
|
1813
1833
|
const links = { [route]: extractLinks(doc.text) };
|
|
1814
1834
|
const origin = site.origin ?? target.origin;
|
|
1815
|
-
await resolveBrokenLinks(pages, links, origin,
|
|
1816
|
-
|
|
1817
|
-
|
|
1818
|
-
|
|
1819
|
-
|
|
1835
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1836
|
+
includeAssets: options.verifyAll,
|
|
1837
|
+
concurrency: options.concurrency,
|
|
1838
|
+
verify: async (candidate) => {
|
|
1839
|
+
try {
|
|
1840
|
+
return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
|
|
1841
|
+
} catch {
|
|
1842
|
+
return false;
|
|
1843
|
+
}
|
|
1820
1844
|
}
|
|
1821
1845
|
});
|
|
1822
1846
|
if (options.checkExternal) {
|
|
@@ -1905,11 +1929,15 @@ async function snapshotFromOrigin(origin, options = {}) {
|
|
|
1905
1929
|
}
|
|
1906
1930
|
})
|
|
1907
1931
|
);
|
|
1908
|
-
await resolveBrokenLinks(pages, links, base.origin,
|
|
1909
|
-
|
|
1910
|
-
|
|
1911
|
-
|
|
1912
|
-
|
|
1932
|
+
await resolveBrokenLinks(pages, links, base.origin, {
|
|
1933
|
+
includeAssets: options.verifyAll,
|
|
1934
|
+
concurrency,
|
|
1935
|
+
verify: async (route) => {
|
|
1936
|
+
try {
|
|
1937
|
+
return await fetchDoc(new URL(route, base).href, timeout) === null;
|
|
1938
|
+
} catch {
|
|
1939
|
+
return false;
|
|
1940
|
+
}
|
|
1913
1941
|
}
|
|
1914
1942
|
});
|
|
1915
1943
|
if (options.checkExternal) {
|
|
@@ -1968,14 +1996,16 @@ async function loadConfig(path = DEFAULT_CONFIG) {
|
|
|
1968
1996
|
}
|
|
1969
1997
|
async function build(flags, config) {
|
|
1970
1998
|
const checkExternal = flags.external ?? config.checkExternal;
|
|
1971
|
-
|
|
1999
|
+
const verifyAll = flags.verifyAll ?? config.verifyAll;
|
|
2000
|
+
if (flags.dir) return snapshotFromDir(flags.dir, { ...config, checkExternal, verifyAll });
|
|
1972
2001
|
if (flags.url)
|
|
1973
2002
|
return snapshotFromOrigin(flags.url, {
|
|
1974
2003
|
...config,
|
|
1975
2004
|
limit: flags.limit,
|
|
1976
2005
|
concurrency: flags.concurrency,
|
|
1977
2006
|
ignoreRobots: flags.ignoreRobots ?? config.ignoreRobots,
|
|
1978
|
-
checkExternal
|
|
2007
|
+
checkExternal,
|
|
2008
|
+
verifyAll
|
|
1979
2009
|
});
|
|
1980
2010
|
throw new Error("Provide a source: --dir <build directory> or --url <origin>.");
|
|
1981
2011
|
}
|
|
@@ -1994,7 +2024,7 @@ function render(findings, format) {
|
|
|
1994
2024
|
}
|
|
1995
2025
|
}
|
|
1996
2026
|
var cli = (0, import_cac.cac)("pagetrace");
|
|
1997
|
-
cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
2027
|
+
cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
1998
2028
|
const config = await loadConfig(flags.config);
|
|
1999
2029
|
const snapshot = await build(flags, config);
|
|
2000
2030
|
const written = await writeLockfile(flags.out, snapshot);
|
|
@@ -2003,7 +2033,7 @@ cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").opti
|
|
|
2003
2033
|
written ? import_picocolors2.default.green(`Wrote ${flags.out} \u2014 ${count} page${count === 1 ? "" : "s"}.`) : import_picocolors2.default.dim(`${flags.out} is already up to date \u2014 ${count} page${count === 1 ? "" : "s"}.`)
|
|
2004
2034
|
);
|
|
2005
2035
|
});
|
|
2006
|
-
cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
|
|
2036
|
+
cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
|
|
2007
2037
|
const failOn = parseFailOn(flags.failOn, false);
|
|
2008
2038
|
const config = await loadConfig(flags.config);
|
|
2009
2039
|
const next = await build(flags, config);
|
|
@@ -2033,7 +2063,7 @@ Failing: ${summary.error} error, ${summary.warn} warning (--fail-on ${failOn}).`
|
|
|
2033
2063
|
process.exitCode = EXIT_FINDINGS;
|
|
2034
2064
|
}
|
|
2035
2065
|
});
|
|
2036
|
-
cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
|
|
2066
|
+
cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
|
|
2037
2067
|
const failOn = parseFailOn(flags.failOn, true);
|
|
2038
2068
|
const config = await loadConfig(flags.config);
|
|
2039
2069
|
const snapshot = await build(flags, config);
|
|
@@ -2073,13 +2103,14 @@ cli.command("audit", "Audit a site as it stands, with explanations and fixes").o
|
|
|
2073
2103
|
process.exitCode = EXIT_FINDINGS;
|
|
2074
2104
|
}
|
|
2075
2105
|
});
|
|
2076
|
-
cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
|
|
2106
|
+
cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
|
|
2077
2107
|
const failOn = parseFailOn(flags.failOn, true);
|
|
2078
2108
|
const config = await loadConfig(flags.config);
|
|
2079
2109
|
const snapshot = await snapshotFromPage(url, {
|
|
2080
2110
|
...config,
|
|
2081
2111
|
concurrency: flags.concurrency,
|
|
2082
|
-
checkExternal: flags.external
|
|
2112
|
+
checkExternal: flags.external,
|
|
2113
|
+
verifyAll: flags.verifyAll ?? config.verifyAll
|
|
2083
2114
|
});
|
|
2084
2115
|
const [page] = Object.values(snapshot.pages);
|
|
2085
2116
|
const platform = detectPlatform([page.generator], Object.values(page.og));
|
|
@@ -2109,7 +2140,7 @@ cli.command("page <url>", "Check one page: every link on it, and its own surface
|
|
|
2109
2140
|
process.exitCode = EXIT_FINDINGS;
|
|
2110
2141
|
}
|
|
2111
2142
|
});
|
|
2112
|
-
cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
|
|
2143
|
+
cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
|
|
2113
2144
|
const failOn = parseFailOn(flags.failOn, true);
|
|
2114
2145
|
const config = await loadConfig(flags.config);
|
|
2115
2146
|
const snapshot = await build(flags, config);
|
|
@@ -2156,7 +2187,7 @@ cli.command("links", "Find internal links that point at no page").option("--url
|
|
|
2156
2187
|
process.exitCode = EXIT_FINDINGS;
|
|
2157
2188
|
}
|
|
2158
2189
|
});
|
|
2159
|
-
cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
2190
|
+
cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
2160
2191
|
const existing = await (0, import_promises2.readFile)(flags.config, "utf8").catch(() => null);
|
|
2161
2192
|
if (existing === null) {
|
|
2162
2193
|
const config = flags.url ? { siteUrl: new URL(flags.url).origin } : {};
|
|
@@ -2179,11 +2210,11 @@ Commit both files, then run \`pagetrace check ${flags.dir ? `--dir ${flags.dir}`
|
|
|
2179
2210
|
});
|
|
2180
2211
|
cli.command("update", "Check npm for a newer pagetrace and install it").option("--check", "Only report whether an update exists").action(async (flags) => {
|
|
2181
2212
|
const latest = await latestVersion("pagetrace");
|
|
2182
|
-
if (!isNewer(latest, "0.
|
|
2183
|
-
console.log(import_picocolors2.default.green(`pagetrace ${"0.
|
|
2213
|
+
if (!isNewer(latest, "0.14.5")) {
|
|
2214
|
+
console.log(import_picocolors2.default.green(`pagetrace ${"0.14.5"} is the latest version.`));
|
|
2184
2215
|
return;
|
|
2185
2216
|
}
|
|
2186
|
-
console.log(import_picocolors2.default.yellow(`Update available: ${"0.
|
|
2217
|
+
console.log(import_picocolors2.default.yellow(`Update available: ${"0.14.5"} \u2192 ${latest}`));
|
|
2187
2218
|
if (flags.check) return;
|
|
2188
2219
|
if (process.argv[1]?.startsWith(process.cwd())) {
|
|
2189
2220
|
console.log(
|
|
@@ -2201,7 +2232,7 @@ cli.command("update", "Check npm for a newer pagetrace and install it").option("
|
|
|
2201
2232
|
console.log(import_picocolors2.default.green(`Updated to pagetrace ${latest}.`));
|
|
2202
2233
|
});
|
|
2203
2234
|
cli.help();
|
|
2204
|
-
cli.version("0.
|
|
2235
|
+
cli.version("0.14.5");
|
|
2205
2236
|
async function main() {
|
|
2206
2237
|
try {
|
|
2207
2238
|
cli.parse(process.argv, { run: false });
|
package/dist/cli.js
CHANGED
|
@@ -1304,7 +1304,7 @@ ${cards || "<p>No issues found.</p>"}
|
|
|
1304
1304
|
|
|
1305
1305
|
// src/snapshot.ts
|
|
1306
1306
|
import { execFile } from "child_process";
|
|
1307
|
-
import { readdir, readFile } from "fs/promises";
|
|
1307
|
+
import { access, readdir, readFile } from "fs/promises";
|
|
1308
1308
|
import { join, relative, sep } from "path";
|
|
1309
1309
|
import { promisify } from "util";
|
|
1310
1310
|
|
|
@@ -1526,7 +1526,7 @@ function extractLinks(html) {
|
|
|
1526
1526
|
}
|
|
1527
1527
|
function extractLlmsTxt(body) {
|
|
1528
1528
|
const sections = body.split(/\r?\n/).filter((line) => line.startsWith("## ")).map((line) => line.slice(3).trim());
|
|
1529
|
-
return { present: true, sections, bytes:
|
|
1529
|
+
return { present: true, sections, bytes: new TextEncoder().encode(body).length };
|
|
1530
1530
|
}
|
|
1531
1531
|
var XML_ENTITIES = {
|
|
1532
1532
|
amp: "&",
|
|
@@ -1675,12 +1675,20 @@ async function snapshotFromDir(dir, config = {}) {
|
|
|
1675
1675
|
const llms = await readFile(join(dir, "llms.txt"), "utf8").catch(() => null);
|
|
1676
1676
|
if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
|
|
1677
1677
|
const origin = site.origin ?? "https://pagetrace.invalid";
|
|
1678
|
-
await resolveBrokenLinks(pages, links, origin
|
|
1678
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1679
|
+
includeAssets: config.verifyAll,
|
|
1680
|
+
// The HTML route set is authoritative here, so a missing page is simply
|
|
1681
|
+
// broken. An asset is a file this crawl never walked, so it gets a stat.
|
|
1682
|
+
verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await access(join(dir, target)).then(
|
|
1683
|
+
() => true,
|
|
1684
|
+
() => false
|
|
1685
|
+
) : true : void 0
|
|
1686
|
+
});
|
|
1679
1687
|
if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
|
|
1680
1688
|
return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
|
|
1681
1689
|
}
|
|
1682
1690
|
var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
|
|
1683
|
-
function linkTarget(href, from, origin) {
|
|
1691
|
+
function linkTarget(href, from, origin, includeAssets = false) {
|
|
1684
1692
|
let url;
|
|
1685
1693
|
try {
|
|
1686
1694
|
url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
|
|
@@ -1688,27 +1696,39 @@ function linkTarget(href, from, origin) {
|
|
|
1688
1696
|
return null;
|
|
1689
1697
|
}
|
|
1690
1698
|
if (url.origin !== origin) return null;
|
|
1691
|
-
if (ASSET_PATH.test(url.pathname)) return null;
|
|
1692
|
-
return routeFromUrl(url.href);
|
|
1699
|
+
if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
|
|
1700
|
+
return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
|
|
1693
1701
|
}
|
|
1694
|
-
var MAX_LINK_CHECKS =
|
|
1695
|
-
async function resolveBrokenLinks(pages, links, origin,
|
|
1702
|
+
var MAX_LINK_CHECKS = 1e3;
|
|
1703
|
+
async function resolveBrokenLinks(pages, links, origin, options = {}) {
|
|
1704
|
+
const { includeAssets = false, verify, concurrency = 5 } = options;
|
|
1696
1705
|
const known = new Set(Object.keys(pages));
|
|
1697
1706
|
const candidates = /* @__PURE__ */ new Map();
|
|
1698
1707
|
for (const [route, hrefs] of Object.entries(links)) {
|
|
1699
1708
|
const missing = [
|
|
1700
1709
|
...new Set(
|
|
1701
|
-
hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
|
|
1710
|
+
hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
|
|
1702
1711
|
)
|
|
1703
1712
|
];
|
|
1704
1713
|
if (missing.length > 0) candidates.set(route, missing);
|
|
1705
1714
|
}
|
|
1706
1715
|
const broken = /* @__PURE__ */ new Set();
|
|
1707
1716
|
if (verify) {
|
|
1708
|
-
const
|
|
1709
|
-
|
|
1710
|
-
|
|
1717
|
+
const unique = [...new Set([...candidates.values()].flat())];
|
|
1718
|
+
const queue = unique.slice(0, MAX_LINK_CHECKS);
|
|
1719
|
+
if (unique.length > queue.length) {
|
|
1720
|
+
console.error(
|
|
1721
|
+
`pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
|
|
1722
|
+
);
|
|
1711
1723
|
}
|
|
1724
|
+
await Promise.all(
|
|
1725
|
+
Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
|
|
1726
|
+
while (queue.length > 0) {
|
|
1727
|
+
const target = queue.shift();
|
|
1728
|
+
if (await verify(target)) broken.add(target);
|
|
1729
|
+
}
|
|
1730
|
+
})
|
|
1731
|
+
);
|
|
1712
1732
|
}
|
|
1713
1733
|
for (const [route, missing] of candidates) {
|
|
1714
1734
|
const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
|
|
@@ -1789,11 +1809,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
|
|
|
1789
1809
|
const pages = { [route]: page };
|
|
1790
1810
|
const links = { [route]: extractLinks(doc.text) };
|
|
1791
1811
|
const origin = site.origin ?? target.origin;
|
|
1792
|
-
await resolveBrokenLinks(pages, links, origin,
|
|
1793
|
-
|
|
1794
|
-
|
|
1795
|
-
|
|
1796
|
-
|
|
1812
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1813
|
+
includeAssets: options.verifyAll,
|
|
1814
|
+
concurrency: options.concurrency,
|
|
1815
|
+
verify: async (candidate) => {
|
|
1816
|
+
try {
|
|
1817
|
+
return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
|
|
1818
|
+
} catch {
|
|
1819
|
+
return false;
|
|
1820
|
+
}
|
|
1797
1821
|
}
|
|
1798
1822
|
});
|
|
1799
1823
|
if (options.checkExternal) {
|
|
@@ -1882,11 +1906,15 @@ async function snapshotFromOrigin(origin, options = {}) {
|
|
|
1882
1906
|
}
|
|
1883
1907
|
})
|
|
1884
1908
|
);
|
|
1885
|
-
await resolveBrokenLinks(pages, links, base.origin,
|
|
1886
|
-
|
|
1887
|
-
|
|
1888
|
-
|
|
1889
|
-
|
|
1909
|
+
await resolveBrokenLinks(pages, links, base.origin, {
|
|
1910
|
+
includeAssets: options.verifyAll,
|
|
1911
|
+
concurrency,
|
|
1912
|
+
verify: async (route) => {
|
|
1913
|
+
try {
|
|
1914
|
+
return await fetchDoc(new URL(route, base).href, timeout) === null;
|
|
1915
|
+
} catch {
|
|
1916
|
+
return false;
|
|
1917
|
+
}
|
|
1890
1918
|
}
|
|
1891
1919
|
});
|
|
1892
1920
|
if (options.checkExternal) {
|
|
@@ -1945,14 +1973,16 @@ async function loadConfig(path = DEFAULT_CONFIG) {
|
|
|
1945
1973
|
}
|
|
1946
1974
|
async function build(flags, config) {
|
|
1947
1975
|
const checkExternal = flags.external ?? config.checkExternal;
|
|
1948
|
-
|
|
1976
|
+
const verifyAll = flags.verifyAll ?? config.verifyAll;
|
|
1977
|
+
if (flags.dir) return snapshotFromDir(flags.dir, { ...config, checkExternal, verifyAll });
|
|
1949
1978
|
if (flags.url)
|
|
1950
1979
|
return snapshotFromOrigin(flags.url, {
|
|
1951
1980
|
...config,
|
|
1952
1981
|
limit: flags.limit,
|
|
1953
1982
|
concurrency: flags.concurrency,
|
|
1954
1983
|
ignoreRobots: flags.ignoreRobots ?? config.ignoreRobots,
|
|
1955
|
-
checkExternal
|
|
1984
|
+
checkExternal,
|
|
1985
|
+
verifyAll
|
|
1956
1986
|
});
|
|
1957
1987
|
throw new Error("Provide a source: --dir <build directory> or --url <origin>.");
|
|
1958
1988
|
}
|
|
@@ -1971,7 +2001,7 @@ function render(findings, format) {
|
|
|
1971
2001
|
}
|
|
1972
2002
|
}
|
|
1973
2003
|
var cli = cac("pagetrace");
|
|
1974
|
-
cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
2004
|
+
cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
1975
2005
|
const config = await loadConfig(flags.config);
|
|
1976
2006
|
const snapshot = await build(flags, config);
|
|
1977
2007
|
const written = await writeLockfile(flags.out, snapshot);
|
|
@@ -1980,7 +2010,7 @@ cli.command("snapshot", "Record the current SEO/AEO surface to a lockfile").opti
|
|
|
1980
2010
|
written ? pc2.green(`Wrote ${flags.out} \u2014 ${count} page${count === 1 ? "" : "s"}.`) : pc2.dim(`${flags.out} is already up to date \u2014 ${count} page${count === 1 ? "" : "s"}.`)
|
|
1981
2011
|
);
|
|
1982
2012
|
});
|
|
1983
|
-
cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
|
|
2013
|
+
cli.command("check", "Compare the current surface against the lockfile").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--lockfile <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | github | sarif", { default: "pretty" }).option("--fail-on <severity>", "error | warn | info", { default: "error" }).option("--audit", "Also run absolute rules, not just the diff", { default: true }).option("--update", "Write the new state to the lockfile after reporting").option("--baseline-branch <ref>", "Read the baseline lockfile from a git ref instead of disk").action(async (flags) => {
|
|
1984
2014
|
const failOn = parseFailOn(flags.failOn, false);
|
|
1985
2015
|
const config = await loadConfig(flags.config);
|
|
1986
2016
|
const next = await build(flags, config);
|
|
@@ -2010,7 +2040,7 @@ Failing: ${summary.error} error, ${summary.warn} warning (--fail-on ${failOn}).`
|
|
|
2010
2040
|
process.exitCode = EXIT_FINDINGS;
|
|
2011
2041
|
}
|
|
2012
2042
|
});
|
|
2013
|
-
cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
|
|
2043
|
+
cli.command("audit", "Audit a site as it stands, with explanations and fixes").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown | html", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "never" }).action(async (flags) => {
|
|
2014
2044
|
const failOn = parseFailOn(flags.failOn, true);
|
|
2015
2045
|
const config = await loadConfig(flags.config);
|
|
2016
2046
|
const snapshot = await build(flags, config);
|
|
@@ -2050,13 +2080,14 @@ cli.command("audit", "Audit a site as it stands, with explanations and fixes").o
|
|
|
2050
2080
|
process.exitCode = EXIT_FINDINGS;
|
|
2051
2081
|
}
|
|
2052
2082
|
});
|
|
2053
|
-
cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
|
|
2083
|
+
cli.command("page <url>", "Check one page: every link on it, and its own surface").option("--external", "Check links that leave the site", { default: true }).option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (url, flags) => {
|
|
2054
2084
|
const failOn = parseFailOn(flags.failOn, true);
|
|
2055
2085
|
const config = await loadConfig(flags.config);
|
|
2056
2086
|
const snapshot = await snapshotFromPage(url, {
|
|
2057
2087
|
...config,
|
|
2058
2088
|
concurrency: flags.concurrency,
|
|
2059
|
-
checkExternal: flags.external
|
|
2089
|
+
checkExternal: flags.external,
|
|
2090
|
+
verifyAll: flags.verifyAll ?? config.verifyAll
|
|
2060
2091
|
});
|
|
2061
2092
|
const [page] = Object.values(snapshot.pages);
|
|
2062
2093
|
const platform = detectPlatform([page.generator], Object.values(page.og));
|
|
@@ -2086,7 +2117,7 @@ cli.command("page <url>", "Check one page: every link on it, and its own surface
|
|
|
2086
2117
|
process.exitCode = EXIT_FINDINGS;
|
|
2087
2118
|
}
|
|
2088
2119
|
});
|
|
2089
|
-
cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
|
|
2120
|
+
cli.command("links", "Find internal links that point at no page").option("--url <origin>", "Live origin to crawl").option("--dir <dir>", "Directory of built HTML").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).option("--format <format>", "pretty | json | markdown", { default: "pretty" }).option("--out <file>", "Write the report to a file instead of stdout").option("--fail-on <severity>", "error | warn | info | never", { default: "error" }).action(async (flags) => {
|
|
2090
2121
|
const failOn = parseFailOn(flags.failOn, true);
|
|
2091
2122
|
const config = await loadConfig(flags.config);
|
|
2092
2123
|
const snapshot = await build(flags, config);
|
|
@@ -2133,7 +2164,7 @@ cli.command("links", "Find internal links that point at no page").option("--url
|
|
|
2133
2164
|
process.exitCode = EXIT_FINDINGS;
|
|
2134
2165
|
}
|
|
2135
2166
|
});
|
|
2136
|
-
cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
2167
|
+
cli.command("init", "Write a config file and take the first snapshot").option("--dir <dir>", "Directory of built HTML").option("--url <origin>", "Live origin to crawl").option("--limit <n>", "Max pages to crawl", { default: 200 }).option("--concurrency <n>", "Parallel requests", { default: 5 }).option("--ignore-robots", "Crawl paths that robots.txt disallows").option("--external", "Also check links that leave the site").option("--verify-all", "Also check links to assets (PDFs, images, archives)").option("--out <file>", "Lockfile path", { default: DEFAULT_LOCKFILE }).option("--config <file>", "Config file", { default: DEFAULT_CONFIG }).action(async (flags) => {
|
|
2137
2168
|
const existing = await readFile2(flags.config, "utf8").catch(() => null);
|
|
2138
2169
|
if (existing === null) {
|
|
2139
2170
|
const config = flags.url ? { siteUrl: new URL(flags.url).origin } : {};
|
|
@@ -2156,11 +2187,11 @@ Commit both files, then run \`pagetrace check ${flags.dir ? `--dir ${flags.dir}`
|
|
|
2156
2187
|
});
|
|
2157
2188
|
cli.command("update", "Check npm for a newer pagetrace and install it").option("--check", "Only report whether an update exists").action(async (flags) => {
|
|
2158
2189
|
const latest = await latestVersion("pagetrace");
|
|
2159
|
-
if (!isNewer(latest, "0.
|
|
2160
|
-
console.log(pc2.green(`pagetrace ${"0.
|
|
2190
|
+
if (!isNewer(latest, "0.14.5")) {
|
|
2191
|
+
console.log(pc2.green(`pagetrace ${"0.14.5"} is the latest version.`));
|
|
2161
2192
|
return;
|
|
2162
2193
|
}
|
|
2163
|
-
console.log(pc2.yellow(`Update available: ${"0.
|
|
2194
|
+
console.log(pc2.yellow(`Update available: ${"0.14.5"} \u2192 ${latest}`));
|
|
2164
2195
|
if (flags.check) return;
|
|
2165
2196
|
if (process.argv[1]?.startsWith(process.cwd())) {
|
|
2166
2197
|
console.log(
|
|
@@ -2178,7 +2209,7 @@ cli.command("update", "Check npm for a newer pagetrace and install it").option("
|
|
|
2178
2209
|
console.log(pc2.green(`Updated to pagetrace ${latest}.`));
|
|
2179
2210
|
});
|
|
2180
2211
|
cli.help();
|
|
2181
|
-
cli.version("0.
|
|
2212
|
+
cli.version("0.14.5");
|
|
2182
2213
|
async function main() {
|
|
2183
2214
|
try {
|
|
2184
2215
|
cli.parse(process.argv, { run: false });
|
package/dist/index.cjs
CHANGED
|
@@ -988,7 +988,7 @@ function extractLinks(html) {
|
|
|
988
988
|
}
|
|
989
989
|
function extractLlmsTxt(body) {
|
|
990
990
|
const sections = body.split(/\r?\n/).filter((line) => line.startsWith("## ")).map((line) => line.slice(3).trim());
|
|
991
|
-
return { present: true, sections, bytes:
|
|
991
|
+
return { present: true, sections, bytes: new TextEncoder().encode(body).length };
|
|
992
992
|
}
|
|
993
993
|
var XML_ENTITIES = {
|
|
994
994
|
amp: "&",
|
|
@@ -1740,12 +1740,20 @@ async function snapshotFromDir(dir, config = {}) {
|
|
|
1740
1740
|
const llms = await (0, import_promises.readFile)((0, import_node_path.join)(dir, "llms.txt"), "utf8").catch(() => null);
|
|
1741
1741
|
if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
|
|
1742
1742
|
const origin = site.origin ?? "https://pagetrace.invalid";
|
|
1743
|
-
await resolveBrokenLinks(pages, links, origin
|
|
1743
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1744
|
+
includeAssets: config.verifyAll,
|
|
1745
|
+
// The HTML route set is authoritative here, so a missing page is simply
|
|
1746
|
+
// broken. An asset is a file this crawl never walked, so it gets a stat.
|
|
1747
|
+
verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await (0, import_promises.access)((0, import_node_path.join)(dir, target)).then(
|
|
1748
|
+
() => true,
|
|
1749
|
+
() => false
|
|
1750
|
+
) : true : void 0
|
|
1751
|
+
});
|
|
1744
1752
|
if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
|
|
1745
1753
|
return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
|
|
1746
1754
|
}
|
|
1747
1755
|
var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
|
|
1748
|
-
function linkTarget(href, from, origin) {
|
|
1756
|
+
function linkTarget(href, from, origin, includeAssets = false) {
|
|
1749
1757
|
let url;
|
|
1750
1758
|
try {
|
|
1751
1759
|
url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
|
|
@@ -1753,27 +1761,39 @@ function linkTarget(href, from, origin) {
|
|
|
1753
1761
|
return null;
|
|
1754
1762
|
}
|
|
1755
1763
|
if (url.origin !== origin) return null;
|
|
1756
|
-
if (ASSET_PATH.test(url.pathname)) return null;
|
|
1757
|
-
return routeFromUrl(url.href);
|
|
1764
|
+
if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
|
|
1765
|
+
return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
|
|
1758
1766
|
}
|
|
1759
|
-
var MAX_LINK_CHECKS =
|
|
1760
|
-
async function resolveBrokenLinks(pages, links, origin,
|
|
1767
|
+
var MAX_LINK_CHECKS = 1e3;
|
|
1768
|
+
async function resolveBrokenLinks(pages, links, origin, options = {}) {
|
|
1769
|
+
const { includeAssets = false, verify, concurrency = 5 } = options;
|
|
1761
1770
|
const known = new Set(Object.keys(pages));
|
|
1762
1771
|
const candidates = /* @__PURE__ */ new Map();
|
|
1763
1772
|
for (const [route, hrefs] of Object.entries(links)) {
|
|
1764
1773
|
const missing = [
|
|
1765
1774
|
...new Set(
|
|
1766
|
-
hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
|
|
1775
|
+
hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
|
|
1767
1776
|
)
|
|
1768
1777
|
];
|
|
1769
1778
|
if (missing.length > 0) candidates.set(route, missing);
|
|
1770
1779
|
}
|
|
1771
1780
|
const broken = /* @__PURE__ */ new Set();
|
|
1772
1781
|
if (verify) {
|
|
1773
|
-
const
|
|
1774
|
-
|
|
1775
|
-
|
|
1782
|
+
const unique = [...new Set([...candidates.values()].flat())];
|
|
1783
|
+
const queue = unique.slice(0, MAX_LINK_CHECKS);
|
|
1784
|
+
if (unique.length > queue.length) {
|
|
1785
|
+
console.error(
|
|
1786
|
+
`pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
|
|
1787
|
+
);
|
|
1776
1788
|
}
|
|
1789
|
+
await Promise.all(
|
|
1790
|
+
Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
|
|
1791
|
+
while (queue.length > 0) {
|
|
1792
|
+
const target = queue.shift();
|
|
1793
|
+
if (await verify(target)) broken.add(target);
|
|
1794
|
+
}
|
|
1795
|
+
})
|
|
1796
|
+
);
|
|
1777
1797
|
}
|
|
1778
1798
|
for (const [route, missing] of candidates) {
|
|
1779
1799
|
const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
|
|
@@ -1854,11 +1874,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
|
|
|
1854
1874
|
const pages = { [route]: page };
|
|
1855
1875
|
const links = { [route]: extractLinks(doc.text) };
|
|
1856
1876
|
const origin = site.origin ?? target.origin;
|
|
1857
|
-
await resolveBrokenLinks(pages, links, origin,
|
|
1858
|
-
|
|
1859
|
-
|
|
1860
|
-
|
|
1861
|
-
|
|
1877
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1878
|
+
includeAssets: options.verifyAll,
|
|
1879
|
+
concurrency: options.concurrency,
|
|
1880
|
+
verify: async (candidate) => {
|
|
1881
|
+
try {
|
|
1882
|
+
return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
|
|
1883
|
+
} catch {
|
|
1884
|
+
return false;
|
|
1885
|
+
}
|
|
1862
1886
|
}
|
|
1863
1887
|
});
|
|
1864
1888
|
if (options.checkExternal) {
|
|
@@ -1947,11 +1971,15 @@ async function snapshotFromOrigin(origin, options = {}) {
|
|
|
1947
1971
|
}
|
|
1948
1972
|
})
|
|
1949
1973
|
);
|
|
1950
|
-
await resolveBrokenLinks(pages, links, base.origin,
|
|
1951
|
-
|
|
1952
|
-
|
|
1953
|
-
|
|
1954
|
-
|
|
1974
|
+
await resolveBrokenLinks(pages, links, base.origin, {
|
|
1975
|
+
includeAssets: options.verifyAll,
|
|
1976
|
+
concurrency,
|
|
1977
|
+
verify: async (route) => {
|
|
1978
|
+
try {
|
|
1979
|
+
return await fetchDoc(new URL(route, base).href, timeout) === null;
|
|
1980
|
+
} catch {
|
|
1981
|
+
return false;
|
|
1982
|
+
}
|
|
1955
1983
|
}
|
|
1956
1984
|
});
|
|
1957
1985
|
if (options.checkExternal) {
|
package/dist/index.d.cts
CHANGED
|
@@ -151,6 +151,11 @@ interface Config {
|
|
|
151
151
|
* third party's bad afternoon into a diff in your repository.
|
|
152
152
|
*/
|
|
153
153
|
checkExternal?: boolean;
|
|
154
|
+
/**
|
|
155
|
+
* Also check links to assets — PDFs, images, archives. They are never crawled
|
|
156
|
+
* as pages, so each one costs a request (or a filesystem check for --dir).
|
|
157
|
+
*/
|
|
158
|
+
verifyAll?: boolean;
|
|
154
159
|
}
|
|
155
160
|
|
|
156
161
|
/**
|
package/dist/index.d.ts
CHANGED
|
@@ -151,6 +151,11 @@ interface Config {
|
|
|
151
151
|
* third party's bad afternoon into a diff in your repository.
|
|
152
152
|
*/
|
|
153
153
|
checkExternal?: boolean;
|
|
154
|
+
/**
|
|
155
|
+
* Also check links to assets — PDFs, images, archives. They are never crawled
|
|
156
|
+
* as pages, so each one costs a request (or a filesystem check for --dir).
|
|
157
|
+
*/
|
|
158
|
+
verifyAll?: boolean;
|
|
154
159
|
}
|
|
155
160
|
|
|
156
161
|
/**
|
package/dist/index.js
CHANGED
|
@@ -913,7 +913,7 @@ function extractLinks(html) {
|
|
|
913
913
|
}
|
|
914
914
|
function extractLlmsTxt(body) {
|
|
915
915
|
const sections = body.split(/\r?\n/).filter((line) => line.startsWith("## ")).map((line) => line.slice(3).trim());
|
|
916
|
-
return { present: true, sections, bytes:
|
|
916
|
+
return { present: true, sections, bytes: new TextEncoder().encode(body).length };
|
|
917
917
|
}
|
|
918
918
|
var XML_ENTITIES = {
|
|
919
919
|
amp: "&",
|
|
@@ -1535,7 +1535,7 @@ ${cards || "<p>No issues found.</p>"}
|
|
|
1535
1535
|
|
|
1536
1536
|
// src/snapshot.ts
|
|
1537
1537
|
import { execFile } from "child_process";
|
|
1538
|
-
import { readdir, readFile } from "fs/promises";
|
|
1538
|
+
import { access, readdir, readFile } from "fs/promises";
|
|
1539
1539
|
import { join, relative, sep } from "path";
|
|
1540
1540
|
import { promisify } from "util";
|
|
1541
1541
|
function routeFromFilePath(root, filePath) {
|
|
@@ -1665,12 +1665,20 @@ async function snapshotFromDir(dir, config = {}) {
|
|
|
1665
1665
|
const llms = await readFile(join(dir, "llms.txt"), "utf8").catch(() => null);
|
|
1666
1666
|
if (llms !== null) site.llmsTxt = extractLlmsTxt(llms);
|
|
1667
1667
|
const origin = site.origin ?? "https://pagetrace.invalid";
|
|
1668
|
-
await resolveBrokenLinks(pages, links, origin
|
|
1668
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1669
|
+
includeAssets: config.verifyAll,
|
|
1670
|
+
// The HTML route set is authoritative here, so a missing page is simply
|
|
1671
|
+
// broken. An asset is a file this crawl never walked, so it gets a stat.
|
|
1672
|
+
verify: config.verifyAll ? async (target) => ASSET_PATH.test(target) ? !await access(join(dir, target)).then(
|
|
1673
|
+
() => true,
|
|
1674
|
+
() => false
|
|
1675
|
+
) : true : void 0
|
|
1676
|
+
});
|
|
1669
1677
|
if (config.checkExternal) await resolveDeadExternal(pages, links, origin, 5);
|
|
1670
1678
|
return { schemaVersion: 1, createdAt: (/* @__PURE__ */ new Date()).toISOString(), site, pages };
|
|
1671
1679
|
}
|
|
1672
1680
|
var ASSET_PATH = /\.(?!html?$)[a-z0-9]+$/i;
|
|
1673
|
-
function linkTarget(href, from, origin) {
|
|
1681
|
+
function linkTarget(href, from, origin, includeAssets = false) {
|
|
1674
1682
|
let url;
|
|
1675
1683
|
try {
|
|
1676
1684
|
url = new URL(href, `${origin}${from === "/" ? "" : from}/`);
|
|
@@ -1678,27 +1686,39 @@ function linkTarget(href, from, origin) {
|
|
|
1678
1686
|
return null;
|
|
1679
1687
|
}
|
|
1680
1688
|
if (url.origin !== origin) return null;
|
|
1681
|
-
if (ASSET_PATH.test(url.pathname)) return null;
|
|
1682
|
-
return routeFromUrl(url.href);
|
|
1689
|
+
if (!includeAssets && ASSET_PATH.test(url.pathname)) return null;
|
|
1690
|
+
return ASSET_PATH.test(url.pathname) ? url.pathname : routeFromUrl(url.href);
|
|
1683
1691
|
}
|
|
1684
|
-
var MAX_LINK_CHECKS =
|
|
1685
|
-
async function resolveBrokenLinks(pages, links, origin,
|
|
1692
|
+
var MAX_LINK_CHECKS = 1e3;
|
|
1693
|
+
async function resolveBrokenLinks(pages, links, origin, options = {}) {
|
|
1694
|
+
const { includeAssets = false, verify, concurrency = 5 } = options;
|
|
1686
1695
|
const known = new Set(Object.keys(pages));
|
|
1687
1696
|
const candidates = /* @__PURE__ */ new Map();
|
|
1688
1697
|
for (const [route, hrefs] of Object.entries(links)) {
|
|
1689
1698
|
const missing = [
|
|
1690
1699
|
...new Set(
|
|
1691
|
-
hrefs.map((href) => linkTarget(href, route, origin)).filter((target) => target !== null && !known.has(target))
|
|
1700
|
+
hrefs.map((href) => linkTarget(href, route, origin, includeAssets)).filter((target) => target !== null && !known.has(target))
|
|
1692
1701
|
)
|
|
1693
1702
|
];
|
|
1694
1703
|
if (missing.length > 0) candidates.set(route, missing);
|
|
1695
1704
|
}
|
|
1696
1705
|
const broken = /* @__PURE__ */ new Set();
|
|
1697
1706
|
if (verify) {
|
|
1698
|
-
const
|
|
1699
|
-
|
|
1700
|
-
|
|
1707
|
+
const unique = [...new Set([...candidates.values()].flat())];
|
|
1708
|
+
const queue = unique.slice(0, MAX_LINK_CHECKS);
|
|
1709
|
+
if (unique.length > queue.length) {
|
|
1710
|
+
console.error(
|
|
1711
|
+
`pagetrace: ${unique.length} link targets to confirm, checking the first ${MAX_LINK_CHECKS}.`
|
|
1712
|
+
);
|
|
1701
1713
|
}
|
|
1714
|
+
await Promise.all(
|
|
1715
|
+
Array.from({ length: Math.min(Math.max(1, concurrency), queue.length) }, async () => {
|
|
1716
|
+
while (queue.length > 0) {
|
|
1717
|
+
const target = queue.shift();
|
|
1718
|
+
if (await verify(target)) broken.add(target);
|
|
1719
|
+
}
|
|
1720
|
+
})
|
|
1721
|
+
);
|
|
1702
1722
|
}
|
|
1703
1723
|
for (const [route, missing] of candidates) {
|
|
1704
1724
|
const confirmed = verify ? missing.filter((target) => broken.has(target)) : missing;
|
|
@@ -1779,11 +1799,15 @@ async function snapshotFromPage(pageUrl, options = {}) {
|
|
|
1779
1799
|
const pages = { [route]: page };
|
|
1780
1800
|
const links = { [route]: extractLinks(doc.text) };
|
|
1781
1801
|
const origin = site.origin ?? target.origin;
|
|
1782
|
-
await resolveBrokenLinks(pages, links, origin,
|
|
1783
|
-
|
|
1784
|
-
|
|
1785
|
-
|
|
1786
|
-
|
|
1802
|
+
await resolveBrokenLinks(pages, links, origin, {
|
|
1803
|
+
includeAssets: options.verifyAll,
|
|
1804
|
+
concurrency: options.concurrency,
|
|
1805
|
+
verify: async (candidate) => {
|
|
1806
|
+
try {
|
|
1807
|
+
return await fetchDoc(new URL(candidate, target.origin).href, options.timeout) === null;
|
|
1808
|
+
} catch {
|
|
1809
|
+
return false;
|
|
1810
|
+
}
|
|
1787
1811
|
}
|
|
1788
1812
|
});
|
|
1789
1813
|
if (options.checkExternal) {
|
|
@@ -1872,11 +1896,15 @@ async function snapshotFromOrigin(origin, options = {}) {
|
|
|
1872
1896
|
}
|
|
1873
1897
|
})
|
|
1874
1898
|
);
|
|
1875
|
-
await resolveBrokenLinks(pages, links, base.origin,
|
|
1876
|
-
|
|
1877
|
-
|
|
1878
|
-
|
|
1879
|
-
|
|
1899
|
+
await resolveBrokenLinks(pages, links, base.origin, {
|
|
1900
|
+
includeAssets: options.verifyAll,
|
|
1901
|
+
concurrency,
|
|
1902
|
+
verify: async (route) => {
|
|
1903
|
+
try {
|
|
1904
|
+
return await fetchDoc(new URL(route, base).href, timeout) === null;
|
|
1905
|
+
} catch {
|
|
1906
|
+
return false;
|
|
1907
|
+
}
|
|
1880
1908
|
}
|
|
1881
1909
|
});
|
|
1882
1910
|
if (options.checkExternal) {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pagetrace",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.14.5",
|
|
4
4
|
"description": "Baseline your site's SEO and AEO surface, diff every build against it, and fail CI on regressions.",
|
|
5
5
|
"main": "./dist/index.cjs",
|
|
6
6
|
"scripts": {
|
|
@@ -10,7 +10,8 @@
|
|
|
10
10
|
"test:watch": "vitest",
|
|
11
11
|
"coverage": "vitest run --coverage",
|
|
12
12
|
"typecheck": "tsc --noEmit",
|
|
13
|
-
"prepublishOnly": "npm run typecheck && npm run test && npm run build"
|
|
13
|
+
"prepublishOnly": "npm run typecheck && npm run test && npm run build",
|
|
14
|
+
"build:site": "tsup src/browser.ts --format iife --global-name pagetrace --platform browser --minify --no-dts --out-dir docs/assets"
|
|
14
15
|
},
|
|
15
16
|
"keywords": [
|
|
16
17
|
"seo",
|