pagetrace 0.9.1 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +59 -0
- package/README.md +43 -3
- package/dist/cli.cjs +384 -25
- package/dist/cli.js +384 -25
- package/dist/index.cjs +311 -17
- package/dist/index.d.cts +67 -4
- package/dist/index.d.ts +67 -4
- package/dist/index.js +308 -17
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,63 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
|
6
6
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
7
|
While the version is below 1.0.0, breaking changes ship in a minor release.
|
|
8
8
|
|
|
9
|
+
## [0.11.0] - 2026-09-08
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- `pagetrace links`, a command for the one question. It crawls once and reports
|
|
14
|
+
only broken internal links, exiting 1 if there are any. The rules are the ones
|
|
15
|
+
`audit` runs, filtered rather than reimplemented, so the two cannot disagree.
|
|
16
|
+
|
|
17
|
+
## [0.10.0] - 2026-09-08
|
|
18
|
+
|
|
19
|
+
### Added
|
|
20
|
+
|
|
21
|
+
- Internal link integrity. `link.broken` (error) for a link pointing at no page,
|
|
22
|
+
and `link.broken.added` for one this build broke — the diff stays quiet about
|
|
23
|
+
breakage it did not introduce, so CI fails on the regression rather than on
|
|
24
|
+
the backlog. Only broken links are stored, keeping a site's navigation out of
|
|
25
|
+
the lockfile and its diff readable. A `--dir` crawl is authoritative; an
|
|
26
|
+
origin crawl confirms each candidate with a real request, because a sitemap
|
|
27
|
+
routinely omits pages that are live. External links are not checked.
|
|
28
|
+
|
|
29
|
+
- Redirect drift. A crawl now records where each route actually landed, taken
|
|
30
|
+
from the response itself at no extra request. `redirect.added` and
|
|
31
|
+
`redirect.changed` (warn) and `redirect.removed` (info) report a route that
|
|
32
|
+
starts, moves or stops redirecting — a whole class of change that was
|
|
33
|
+
previously invisible, since a redirected route still answers. A redirect that
|
|
34
|
+
only adds or drops a trailing slash is server configuration, not drift, and is
|
|
35
|
+
not reported.
|
|
36
|
+
- `canonical.redirects` (warn): a canonical naming a URL that redirects is one
|
|
37
|
+
the engine may ignore. Only raised when the target was actually crawled.
|
|
38
|
+
- Sitemap health. `sitemap.dead` (error) for a sitemap entry answering 404, and
|
|
39
|
+
`sitemap.redirect` (warn) for one that redirects. Both read data the crawl
|
|
40
|
+
already gathered.
|
|
41
|
+
- `--format sarif` on `check`, for `github/codeql-action/upload-sarif`. Findings
|
|
42
|
+
become pull request annotations and Security tab entries.
|
|
43
|
+
- `pagetrace init`: writes a config file and takes the first snapshot. An
|
|
44
|
+
existing config is never overwritten.
|
|
45
|
+
- Thirteen more Schema.org types in the rich-results table, taking it from 20 to
|
|
46
|
+
33: Book, Dataset, QAPage, Question, Answer, ClaimReview, Movie, ProfilePage,
|
|
47
|
+
DiscussionForumPosting, ImageObject, SpecialAnnouncement,
|
|
48
|
+
EmployerAggregateRating and PodcastSeries.
|
|
49
|
+
|
|
50
|
+
### Changed
|
|
51
|
+
|
|
52
|
+
- An origin crawl now honours `robots.txt` `Disallow`, with longest-match-wins
|
|
53
|
+
and Allow winning ties, so `Disallow: /` plus `Allow: /blog` crawls the blog.
|
|
54
|
+
**A staging origin that serves `Disallow: /` now yields no pages** — pass
|
|
55
|
+
`--ignore-robots`, or set `"ignoreRobots": true`, to crawl a site you own.
|
|
56
|
+
- **`sitemap.dead` is an error**, so a sitemap that already lists a dead URL
|
|
57
|
+
fails `check` on the first run after upgrading. The problem predates the rule.
|
|
58
|
+
|
|
59
|
+
### Fixed
|
|
60
|
+
|
|
61
|
+
- A field absent from an older lockfile read as a transition rather than as
|
|
62
|
+
missing information, because the scalar diff compared `undefined` against
|
|
63
|
+
`null`. Every optional field added from here on would have reported a phantom
|
|
64
|
+
change on every page of the first run.
|
|
65
|
+
|
|
9
66
|
## [0.9.1] - 2026-09-07
|
|
10
67
|
|
|
11
68
|
### Changed
|
|
@@ -281,6 +338,8 @@ Initial release. `snapshot`, `check` and `audit` commands; filesystem and HTTP
|
|
|
281
338
|
crawling; diff classified by transition; absolute, cross-page and hreflang audit
|
|
282
339
|
rules; pretty, JSON, markdown, GitHub and HTML reporters.
|
|
283
340
|
|
|
341
|
+
[0.11.0]: https://github.com/shyamexe/pagetrace/compare/v0.10.0...v0.11.0
|
|
342
|
+
[0.10.0]: https://github.com/shyamexe/pagetrace/compare/v0.9.1...v0.10.0
|
|
284
343
|
[0.9.1]: https://github.com/shyamexe/pagetrace/compare/v0.9.0...v0.9.1
|
|
285
344
|
[0.9.0]: https://github.com/shyamexe/pagetrace/compare/v0.8.1...v0.9.0
|
|
286
345
|
[0.8.1]: https://github.com/shyamexe/pagetrace/compare/v0.8.0...v0.8.1
|
package/README.md
CHANGED
|
@@ -21,13 +21,21 @@ Requires Node 20.19 or newer. No native modules, three small dependencies.
|
|
|
21
21
|
|
|
22
22
|
## Use
|
|
23
23
|
|
|
24
|
-
|
|
24
|
+
Set up a config file and the first baseline in one step:
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
npx pagetrace init --dir ./out
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
Or record the baseline on its own:
|
|
25
31
|
|
|
26
32
|
```bash
|
|
27
33
|
npx pagetrace snapshot --dir ./out
|
|
28
34
|
git add pagetrace.lock.json
|
|
29
35
|
```
|
|
30
36
|
|
|
37
|
+
`init` never overwrites an existing config; it refreshes the lockfile and leaves your edits alone.
|
|
38
|
+
|
|
31
39
|
Check every build against it:
|
|
32
40
|
|
|
33
41
|
```bash
|
|
@@ -113,6 +121,8 @@ npx pagetrace snapshot --url https://example.com --limit 200
|
|
|
113
121
|
|
|
114
122
|
Routes are discovered from `robots.txt` sitemap declarations, falling back to `/sitemap.xml`. Sitemap indexes are followed one level, up to 50 children, and expansion stops once `--limit` is satisfied. Gzipped children are recognised but not read.
|
|
115
123
|
|
|
124
|
+
Paths that `robots.txt` disallows are skipped, with the longest matching rule winning so an `Allow` exception still gets crawled. A staging origin that serves `Disallow: /` would therefore yield nothing — pass `--ignore-robots` (or set `"ignoreRobots": true` in the config) to crawl a site you own anyway.
|
|
125
|
+
|
|
116
126
|
URLs pointing at another host are skipped. A page that cannot be fetched stops the run with an error rather than being dropped from the snapshot — a page silently missing from a crawl is indistinguishable from a page you deleted, and reporting a transient outage as a site-wide deletion is worse than failing.
|
|
117
127
|
|
|
118
128
|
### Version and updates
|
|
@@ -146,10 +156,29 @@ pagetrace update # install it globally
|
|
|
146
156
|
| `aeo.crawler.newly_blocked` | error | `robots.txt` started blocking an AI crawler |
|
|
147
157
|
| `aeo.llmstxt.removed` | error | `/llms.txt` disappeared |
|
|
148
158
|
| `page.removed` | warn | A route in the lockfile is no longer there |
|
|
159
|
+
| `redirect.added` | warn | A route that used to answer directly now redirects |
|
|
160
|
+
| `redirect.changed` | warn | A route redirects somewhere new |
|
|
161
|
+
| `canonical.redirects` | warn | A canonical points at a URL that redirects |
|
|
162
|
+
| `link.broken.added` | error | A page started linking to a URL that does not exist |
|
|
163
|
+
| `sitemap.dead` | error | The sitemap lists a URL that answers 404 |
|
|
164
|
+
| `sitemap.redirect` | warn | The sitemap lists a URL that redirects |
|
|
149
165
|
| `og.removed` / `hreflang.removed` | warn | Social or i18n tags dropped |
|
|
150
166
|
| `title.changed` | info | Ordinary copy edit |
|
|
151
167
|
|
|
152
|
-
|
|
168
|
+
Broken links have a command of their own, when that is the only question you have:
|
|
169
|
+
|
|
170
|
+
```bash
|
|
171
|
+
npx pagetrace links --url https://example.com
|
|
172
|
+
npx pagetrace links --dir ./out --format json
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
It crawls once, runs the same rules as `audit`, and prints only the link findings — exit 1 if any, or `No broken links found — 42 pages checked.`
|
|
176
|
+
|
|
177
|
+
Internal links are checked too. Only the broken ones are stored, so a site's navigation never lands in the lockfile: `link.broken` for a link that is already dead, `link.broken.added` for one this build broke. A `--dir` crawl is authoritative — the build directory is the whole site — while a crawl confirms each candidate with a real request first, because a sitemap routinely omits pages that are live. External links are deliberately not checked: a Cloudflare 403 and a rate limit both look like a dead page, and that is where link checkers earn their reputation for false positives.
|
|
178
|
+
|
|
179
|
+
Redirects are recorded from the response itself, so they cost no extra requests. A redirect that only adds or drops a trailing slash is server configuration rather than drift and is not reported. `canonical.redirects` is only raised when the canonical's target was actually crawled, so a `--limit` run cannot invent it.
|
|
180
|
+
|
|
181
|
+
Alongside the diff, `check` runs absolute rules: missing title, canonical, `h1`, or description; JSON-LD required and recommended properties for the thirty-three Schema.org types Google supports as rich results; thin content; images without `alt`; and AEO signals like whether the page opens with something an answer engine can quote. Disable with `--no-audit`.
|
|
153
182
|
|
|
154
183
|
## Config
|
|
155
184
|
|
|
@@ -200,7 +229,18 @@ Without the Action:
|
|
|
200
229
|
- run: npx pagetrace check --dir ./out --baseline-branch origin/main --format github
|
|
201
230
|
```
|
|
202
231
|
|
|
203
|
-
`--format` accepts `pretty`, `json`, `markdown` (sized for a PR comment),
|
|
232
|
+
`--format` accepts `pretty`, `json`, `markdown` (sized for a PR comment), `github` (workflow annotations) and `sarif`.
|
|
233
|
+
|
|
234
|
+
SARIF puts the findings in the Security tab and on the pull request itself, which survives longer than a comment:
|
|
235
|
+
|
|
236
|
+
```yaml
|
|
237
|
+
- run: npx pagetrace check --dir ./out --baseline-branch origin/main --format sarif > pagetrace.sarif
|
|
238
|
+
- uses: github/codeql-action/upload-sarif@v3
|
|
239
|
+
with:
|
|
240
|
+
sarif_file: pagetrace.sarif
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
Needs `security-events: write`. Routes are not source files, so GitHub lists each finding without anchoring it to a line in the diff.
|
|
204
244
|
|
|
205
245
|
## Programmatic API
|
|
206
246
|
|