pagetrace 0.9.1 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,63 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
  While the version is below 1.0.0, breaking changes ship in a minor release.
8
8
 
9
+ ## [0.11.0] - 2026-09-08
10
+
11
+ ### Added
12
+
13
+ - `pagetrace links`, a command for the one question. It crawls once and reports
14
+ only broken internal links, exiting 1 if there are any. The rules are the ones
15
+ `audit` runs, filtered rather than reimplemented, so the two cannot disagree.
16
+
17
+ ## [0.10.0] - 2026-09-08
18
+
19
+ ### Added
20
+
21
+ - Internal link integrity. `link.broken` (error) for a link pointing at no page,
22
+ and `link.broken.added` for one this build broke — the diff stays quiet about
23
+ breakage it did not introduce, so CI fails on the regression rather than on
24
+ the backlog. Only broken links are stored, keeping a site's navigation out of
25
+ the lockfile and its diff readable. A `--dir` crawl is authoritative; an
26
+ origin crawl confirms each candidate with a real request, because a sitemap
27
+ routinely omits pages that are live. External links are not checked.
28
+
29
+ - Redirect drift. A crawl now records where each route actually landed, taken
30
+ from the response itself at no extra request. `redirect.added` and
31
+ `redirect.changed` (warn) and `redirect.removed` (info) report a route that
32
+ starts, moves or stops redirecting — a whole class of change that was
33
+ previously invisible, since a redirected route still answers. A redirect that
34
+ only adds or drops a trailing slash is server configuration, not drift, and is
35
+ not reported.
36
+ - `canonical.redirects` (warn): a canonical naming a URL that redirects is one
37
+ the engine may ignore. Only raised when the target was actually crawled.
38
+ - Sitemap health. `sitemap.dead` (error) for a sitemap entry answering 404, and
39
+ `sitemap.redirect` (warn) for one that redirects. Both read data the crawl
40
+ already gathered.
41
+ - `--format sarif` on `check`, for `github/codeql-action/upload-sarif`. Findings
42
+ become pull request annotations and Security tab entries.
43
+ - `pagetrace init`: writes a config file and takes the first snapshot. An
44
+ existing config is never overwritten.
45
+ - Thirteen more Schema.org types in the rich-results table, taking it from 20 to
46
+ 33: Book, Dataset, QAPage, Question, Answer, ClaimReview, Movie, ProfilePage,
47
+ DiscussionForumPosting, ImageObject, SpecialAnnouncement,
48
+ EmployerAggregateRating and PodcastSeries.
49
+
50
+ ### Changed
51
+
52
+ - An origin crawl now honours `robots.txt` `Disallow`, with longest-match-wins
53
+ and Allow winning ties, so `Disallow: /` plus `Allow: /blog` crawls the blog.
54
+ **A staging origin that serves `Disallow: /` now yields no pages** — pass
55
+ `--ignore-robots`, or set `"ignoreRobots": true`, to crawl a site you own.
56
+ - **`sitemap.dead` is an error**, so a sitemap that already lists a dead URL
57
+ fails `check` on the first run after upgrading. The problem predates the rule.
58
+
59
+ ### Fixed
60
+
61
+ - A field absent from an older lockfile read as a transition rather than as
62
+ missing information, because the scalar diff compared `undefined` against
63
+ `null`. Every optional field added from here on would have reported a phantom
64
+ change on every page of the first run.
65
+
9
66
  ## [0.9.1] - 2026-09-07
10
67
 
11
68
  ### Changed
@@ -281,6 +338,8 @@ Initial release. `snapshot`, `check` and `audit` commands; filesystem and HTTP
281
338
  crawling; diff classified by transition; absolute, cross-page and hreflang audit
282
339
  rules; pretty, JSON, markdown, GitHub and HTML reporters.
283
340
 
341
+ [0.11.0]: https://github.com/shyamexe/pagetrace/compare/v0.10.0...v0.11.0
342
+ [0.10.0]: https://github.com/shyamexe/pagetrace/compare/v0.9.1...v0.10.0
284
343
  [0.9.1]: https://github.com/shyamexe/pagetrace/compare/v0.9.0...v0.9.1
285
344
  [0.9.0]: https://github.com/shyamexe/pagetrace/compare/v0.8.1...v0.9.0
286
345
  [0.8.1]: https://github.com/shyamexe/pagetrace/compare/v0.8.0...v0.8.1
package/README.md CHANGED
@@ -21,13 +21,21 @@ Requires Node 20.19 or newer. No native modules, three small dependencies.
21
21
 
22
22
  ## Use
23
23
 
24
- Record a baseline from your build output:
24
+ Set up a config file and the first baseline in one step:
25
+
26
+ ```bash
27
+ npx pagetrace init --dir ./out
28
+ ```
29
+
30
+ Or record the baseline on its own:
25
31
 
26
32
  ```bash
27
33
  npx pagetrace snapshot --dir ./out
28
34
  git add pagetrace.lock.json
29
35
  ```
30
36
 
37
+ `init` never overwrites an existing config; it refreshes the lockfile and leaves your edits alone.
38
+
31
39
  Check every build against it:
32
40
 
33
41
  ```bash
@@ -113,6 +121,8 @@ npx pagetrace snapshot --url https://example.com --limit 200
113
121
 
114
122
  Routes are discovered from `robots.txt` sitemap declarations, falling back to `/sitemap.xml`. Sitemap indexes are followed one level, up to 50 children, and expansion stops once `--limit` is satisfied. Gzipped children are recognised but not read.
115
123
 
124
+ Paths that `robots.txt` disallows are skipped, with the longest matching rule winning so an `Allow` exception still gets crawled. A staging origin that serves `Disallow: /` would therefore yield nothing — pass `--ignore-robots` (or set `"ignoreRobots": true` in the config) to crawl a site you own anyway.
125
+
116
126
  URLs pointing at another host are skipped. A page that cannot be fetched stops the run with an error rather than being dropped from the snapshot — a page silently missing from a crawl is indistinguishable from a page you deleted, and reporting a transient outage as a site-wide deletion is worse than failing.
117
127
 
118
128
  ### Version and updates
@@ -146,10 +156,29 @@ pagetrace update # install it globally
146
156
  | `aeo.crawler.newly_blocked` | error | `robots.txt` started blocking an AI crawler |
147
157
  | `aeo.llmstxt.removed` | error | `/llms.txt` disappeared |
148
158
  | `page.removed` | warn | A route in the lockfile is no longer there |
159
+ | `redirect.added` | warn | A route that used to answer directly now redirects |
160
+ | `redirect.changed` | warn | A route redirects somewhere new |
161
+ | `canonical.redirects` | warn | A canonical points at a URL that redirects |
162
+ | `link.broken.added` | error | A page started linking to a URL that does not exist |
163
+ | `sitemap.dead` | error | The sitemap lists a URL that answers 404 |
164
+ | `sitemap.redirect` | warn | The sitemap lists a URL that redirects |
149
165
  | `og.removed` / `hreflang.removed` | warn | Social or i18n tags dropped |
150
166
  | `title.changed` | info | Ordinary copy edit |
151
167
 
152
- Alongside the diff, `check` runs absolute rules: missing title, canonical, `h1`, or description; JSON-LD required and recommended properties for the twenty Schema.org types Google supports as rich results; thin content; images without `alt`; and AEO signals like whether the page opens with something an answer engine can quote. Disable with `--no-audit`.
168
+ Broken links have a command of their own, when that is the only question you have:
169
+
170
+ ```bash
171
+ npx pagetrace links --url https://example.com
172
+ npx pagetrace links --dir ./out --format json
173
+ ```
174
+
175
+ It crawls once, runs the same rules as `audit`, and prints only the link findings — exit 1 if any, or `No broken links found — 42 pages checked.`
176
+
177
+ Internal links are checked too. Only the broken ones are stored, so a site's navigation never lands in the lockfile: `link.broken` for a link that is already dead, `link.broken.added` for one this build broke. A `--dir` crawl is authoritative — the build directory is the whole site — while a crawl confirms each candidate with a real request first, because a sitemap routinely omits pages that are live. External links are deliberately not checked: a Cloudflare 403 and a rate limit both look like a dead page, and that is where link checkers earn their reputation for false positives.
178
+
179
+ Redirects are recorded from the response itself, so they cost no extra requests. A redirect that only adds or drops a trailing slash is server configuration rather than drift and is not reported. `canonical.redirects` is only raised when the canonical's target was actually crawled, so a `--limit` run cannot invent it.
180
+
181
+ Alongside the diff, `check` runs absolute rules: missing title, canonical, `h1`, or description; JSON-LD required and recommended properties for the thirty-three Schema.org types Google supports as rich results; thin content; images without `alt`; and AEO signals like whether the page opens with something an answer engine can quote. Disable with `--no-audit`.
153
182
 
154
183
  ## Config
155
184
 
@@ -200,7 +229,18 @@ Without the Action:
200
229
  - run: npx pagetrace check --dir ./out --baseline-branch origin/main --format github
201
230
  ```
202
231
 
203
- `--format` accepts `pretty`, `json`, `markdown` (sized for a PR comment), and `github` (workflow annotations).
232
+ `--format` accepts `pretty`, `json`, `markdown` (sized for a PR comment), `github` (workflow annotations) and `sarif`.
233
+
234
+ SARIF puts the findings in the Security tab and on the pull request itself, which survives longer than a comment:
235
+
236
+ ```yaml
237
+ - run: npx pagetrace check --dir ./out --baseline-branch origin/main --format sarif > pagetrace.sarif
238
+ - uses: github/codeql-action/upload-sarif@v3
239
+ with:
240
+ sarif_file: pagetrace.sarif
241
+ ```
242
+
243
+ Needs `security-events: write`. Routes are not source files, so GitHub lists each finding without anchoring it to a line in the diff.
204
244
 
205
245
  ## Programmatic API
206
246