pagequiet 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 User0856
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,134 @@
1
+ Metadata-Version: 2.4
2
+ Name: pagequiet
3
+ Version: 0.1.0
4
+ Summary: Website change detection that stays quiet until something real changes.
5
+ Author: User0856
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/User0856/pagequiet
8
+ Project-URL: Issues, https://github.com/User0856/pagequiet/issues
9
+ Keywords: website monitoring,change detection,web page diff,playwright,screenshot,visualping alternative
10
+ Classifier: Programming Language :: Python :: 3
11
+ Classifier: Environment :: Console
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Topic :: Internet :: WWW/HTTP :: Site Management :: Link Checking
14
+ Classifier: Topic :: Software Development :: Testing
15
+ Requires-Python: >=3.11
16
+ Description-Content-Type: text/markdown
17
+ License-File: LICENSE
18
+ Provides-Extra: playwright
19
+ Requires-Dist: playwright>=1.40; extra == "playwright"
20
+ Provides-Extra: dev
21
+ Requires-Dist: pytest>=8; extra == "dev"
22
+ Dynamic: license-file
23
+
24
+ # pagequiet
25
+
26
+ Website change detection that stays quiet until something real changes.
27
+
28
+ Most change monitors compare HTML, so they alert on every rotating ad slot, CSRF token and "updated 3 minutes ago". After a week nobody reads the alerts. `pagequiet` compares what a reader sees instead: the rendered text of one element on the page, with relative times, counters and whitespace stripped. When that changes, you get the changed lines and a full-page screenshot of the new version.
29
+
30
+ ```
31
+ $ pagequiet check
32
+ CHANGED example-com-pricing https://example.com/pricing
33
+ - Pro $49 / month
34
+ + Pro $59 / month
35
+ screenshot: .pagequiet/captures/example-com-pricing-20261005T080012Z.png
36
+ 1 changed, 11 unchanged, 0 new, 0 errors
37
+ ```
38
+
39
+ It is a single command with a TOML file, built for cron and CI: exit code 1 means something changed.
40
+
41
+ ## Install
42
+
43
+ ```bash
44
+ pip install "pagequiet[playwright]"
45
+ playwright install chromium
46
+ ```
47
+
48
+ Python 3.11 or later. The Playwright extra renders pages with a local Chromium and costs nothing. If you would rather not run a browser (CI runners, small servers), see [Renderers](#renderers).
49
+
50
+ ## Quick start
51
+
52
+ ```bash
53
+ pagequiet init # writes pagequiet.toml
54
+ # edit the [[page]] entries
55
+ pagequiet check # first run records a baseline
56
+ pagequiet check # later runs report changes
57
+ ```
58
+
59
+ ## Configuration
60
+
61
+ ```toml
62
+ [settings]
63
+ renderer = "playwright" # or "snaprender"
64
+ state = ".pagequiet/state.json" # last text and fingerprint per page
65
+ captures = ".pagequiet/captures" # full-page screenshot of each new version
66
+ notify = ["stdout", "slack"] # stdout, slack, webhook
67
+ timeout = 45 # seconds per page
68
+ use_default_noise = true # strip relative times, timestamps, view counters
69
+
70
+ [[page]]
71
+ url = "https://example.com/pricing"
72
+ selector = "main" # watch one element, not the whole page
73
+ ignore = ['Offer ends in \d+ days'] # extra regex patterns removed before comparing
74
+ hide = [".testimonial-carousel"] # hidden before text and screenshot
75
+
76
+ [[page]]
77
+ url = "https://example.com/legal/terms"
78
+ selector = "article"
79
+ screenshot = false # text only
80
+ ```
81
+
82
+ Notifiers read their targets from the environment:
83
+
84
+ | Notifier | Variable | Sends |
85
+ |---|---|---|
86
+ | `stdout` | none | the changed lines and screenshot path |
87
+ | `slack` | `PAGEQUIET_SLACK_WEBHOOK` | a message to a Slack incoming webhook |
88
+ | `webhook` | `PAGEQUIET_WEBHOOK_URL` | JSON: `{"event": "page.changed", "name", "url", "changes", "screenshot"}` |
89
+
90
+ Exit codes: `0` nothing changed, `1` at least one page changed, `2` errors and no changes.
91
+
92
+ ## What counts as a change
93
+
94
+ 1. Render the page and take the visible text of `selector` (default `body`).
95
+ 2. Remove noise: relative times ("5 minutes ago", "just now"), ISO timestamps and clock times, view and visitor counters, extra whitespace, plus your `ignore` patterns.
96
+ 3. Hash the result. Same hash as last run: quiet. Different: report the changed lines and take a screenshot.
97
+
98
+ The defaults are deliberately few. Every pattern you strip is a kind of change you will never be told about, so add `ignore` patterns only after you have seen one cause a false alarm. If timestamps on a page are meaningful to you, set `use_default_noise = false`.
99
+
100
+ The biggest single improvement is `selector`. Watching `main` or the pricing table instead of the whole page ignores navigation, footers and "latest posts" widgets that change for reasons you do not care about.
101
+
102
+ A failed render or an empty selector is reported as an error and never as a change, and the previous text is kept, so a flaky site does not produce a false alarm the next time it loads.
103
+
104
+ ## Renderers
105
+
106
+ **`playwright`** (default) runs Chromium locally through Playwright. Free, private, and fine for most pages. Cookie consent banners are part of the page, so if one varies between visits, scope `selector` to the content or add the banner to `hide`.
107
+
108
+ **`snaprender`** uses the hosted [SnapRender](https://snap-render.com) API to render pages, so no browser runs on your machine. Cookie banners and ads are removed before the text and screenshot are taken. It needs `SNAPRENDER_API_KEY`; the free plan covers 200 renders a month. Each check is one render, plus one when a page changed. Disclosure: I build SnapRender. `pagequiet` works fully without it and always will.
109
+
110
+ ## Run it on a schedule
111
+
112
+ Cron, every six hours:
113
+
114
+ ```bash
115
+ 0 */6 * * * cd /opt/watch && pagequiet check -q >> pagequiet.log 2>&1
116
+ ```
117
+
118
+ GitHub Actions, keeping state in the repository: see [`examples/github-actions.yml`](examples/github-actions.yml). Scheduled workflows run at most every five minutes, can start late at busy times, and are paused in public repositories after 60 days without activity; committing the state file each run counts as activity.
119
+
120
+ ## Limits
121
+
122
+ - Pages behind a login are not supported.
123
+ - Text comparison misses purely visual changes (a new image, a broken layout). Compare the screenshots in `captures/` for those, or keep `screenshot = true` and review them when text changes.
124
+ - `hide` applies to the SnapRender renderer's screenshots only; scope its text with `selector`.
125
+
126
+ ## Development
127
+
128
+ ```bash
129
+ python -m venv .venv && . .venv/bin/activate
130
+ pip install -e ".[dev,playwright]"
131
+ pytest
132
+ ```
133
+
134
+ MIT licensed.
@@ -0,0 +1,111 @@
1
+ # pagequiet
2
+
3
+ Website change detection that stays quiet until something real changes.
4
+
5
+ Most change monitors compare HTML, so they alert on every rotating ad slot, CSRF token and "updated 3 minutes ago". After a week nobody reads the alerts. `pagequiet` compares what a reader sees instead: the rendered text of one element on the page, with relative times, counters and whitespace stripped. When that changes, you get the changed lines and a full-page screenshot of the new version.
6
+
7
+ ```
8
+ $ pagequiet check
9
+ CHANGED example-com-pricing https://example.com/pricing
10
+ - Pro $49 / month
11
+ + Pro $59 / month
12
+ screenshot: .pagequiet/captures/example-com-pricing-20261005T080012Z.png
13
+ 1 changed, 11 unchanged, 0 new, 0 errors
14
+ ```
15
+
16
+ It is a single command with a TOML file, built for cron and CI: exit code 1 means something changed.
17
+
18
+ ## Install
19
+
20
+ ```bash
21
+ pip install "pagequiet[playwright]"
22
+ playwright install chromium
23
+ ```
24
+
25
+ Python 3.11 or later. The Playwright extra renders pages with a local Chromium and costs nothing. If you would rather not run a browser (CI runners, small servers), see [Renderers](#renderers).
26
+
27
+ ## Quick start
28
+
29
+ ```bash
30
+ pagequiet init # writes pagequiet.toml
31
+ # edit the [[page]] entries
32
+ pagequiet check # first run records a baseline
33
+ pagequiet check # later runs report changes
34
+ ```
35
+
36
+ ## Configuration
37
+
38
+ ```toml
39
+ [settings]
40
+ renderer = "playwright" # or "snaprender"
41
+ state = ".pagequiet/state.json" # last text and fingerprint per page
42
+ captures = ".pagequiet/captures" # full-page screenshot of each new version
43
+ notify = ["stdout", "slack"] # stdout, slack, webhook
44
+ timeout = 45 # seconds per page
45
+ use_default_noise = true # strip relative times, timestamps, view counters
46
+
47
+ [[page]]
48
+ url = "https://example.com/pricing"
49
+ selector = "main" # watch one element, not the whole page
50
+ ignore = ['Offer ends in \d+ days'] # extra regex patterns removed before comparing
51
+ hide = [".testimonial-carousel"] # hidden before text and screenshot
52
+
53
+ [[page]]
54
+ url = "https://example.com/legal/terms"
55
+ selector = "article"
56
+ screenshot = false # text only
57
+ ```
58
+
59
+ Notifiers read their targets from the environment:
60
+
61
+ | Notifier | Variable | Sends |
62
+ |---|---|---|
63
+ | `stdout` | none | the changed lines and screenshot path |
64
+ | `slack` | `PAGEQUIET_SLACK_WEBHOOK` | a message to a Slack incoming webhook |
65
+ | `webhook` | `PAGEQUIET_WEBHOOK_URL` | JSON: `{"event": "page.changed", "name", "url", "changes", "screenshot"}` |
66
+
67
+ Exit codes: `0` nothing changed, `1` at least one page changed, `2` errors and no changes.
68
+
69
+ ## What counts as a change
70
+
71
+ 1. Render the page and take the visible text of `selector` (default `body`).
72
+ 2. Remove noise: relative times ("5 minutes ago", "just now"), ISO timestamps and clock times, view and visitor counters, extra whitespace, plus your `ignore` patterns.
73
+ 3. Hash the result. Same hash as last run: quiet. Different: report the changed lines and take a screenshot.
74
+
75
+ The defaults are deliberately few. Every pattern you strip is a kind of change you will never be told about, so add `ignore` patterns only after you have seen one cause a false alarm. If timestamps on a page are meaningful to you, set `use_default_noise = false`.
76
+
77
+ The biggest single improvement is `selector`. Watching `main` or the pricing table instead of the whole page ignores navigation, footers and "latest posts" widgets that change for reasons you do not care about.
78
+
79
+ A failed render or an empty selector is reported as an error and never as a change, and the previous text is kept, so a flaky site does not produce a false alarm the next time it loads.
80
+
81
+ ## Renderers
82
+
83
+ **`playwright`** (default) runs Chromium locally through Playwright. Free, private, and fine for most pages. Cookie consent banners are part of the page, so if one varies between visits, scope `selector` to the content or add the banner to `hide`.
84
+
85
+ **`snaprender`** uses the hosted [SnapRender](https://snap-render.com) API to render pages, so no browser runs on your machine. Cookie banners and ads are removed before the text and screenshot are taken. It needs `SNAPRENDER_API_KEY`; the free plan covers 200 renders a month. Each check is one render, plus one when a page changed. Disclosure: I build SnapRender. `pagequiet` works fully without it and always will.
86
+
87
+ ## Run it on a schedule
88
+
89
+ Cron, every six hours:
90
+
91
+ ```bash
92
+ 0 */6 * * * cd /opt/watch && pagequiet check -q >> pagequiet.log 2>&1
93
+ ```
94
+
95
+ GitHub Actions, keeping state in the repository: see [`examples/github-actions.yml`](examples/github-actions.yml). Scheduled workflows run at most every five minutes, can start late at busy times, and are paused in public repositories after 60 days without activity; committing the state file each run counts as activity.
96
+
97
+ ## Limits
98
+
99
+ - Pages behind a login are not supported.
100
+ - Text comparison misses purely visual changes (a new image, a broken layout). Compare the screenshots in `captures/` for those, or keep `screenshot = true` and review them when text changes.
101
+ - `hide` applies to the SnapRender renderer's screenshots only; scope its text with `selector`.
102
+
103
+ ## Development
104
+
105
+ ```bash
106
+ python -m venv .venv && . .venv/bin/activate
107
+ pip install -e ".[dev,playwright]"
108
+ pytest
109
+ ```
110
+
111
+ MIT licensed.
@@ -0,0 +1,3 @@
1
+ """pagequiet: website change detection that stays quiet until something real changes."""
2
+
3
+ __version__ = "0.1.0"
@@ -0,0 +1,5 @@
1
+ import sys
2
+
3
+ from .cli import main
4
+
5
+ sys.exit(main())
@@ -0,0 +1,86 @@
1
+ """One monitoring run: render each page, compare with the stored state, report changes."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import datetime as dt
6
+ import json
7
+ from dataclasses import dataclass, field
8
+ from pathlib import Path
9
+
10
+ from .config import Config
11
+ from .normalize import fingerprint, normalize, summarize
12
+ from .notify import Change
13
+ from .renderers import Renderer
14
+
15
+
16
+ @dataclass
17
+ class Result:
18
+ changed: list[Change] = field(default_factory=list)
19
+ baselined: list[str] = field(default_factory=list)
20
+ unchanged: list[str] = field(default_factory=list)
21
+ errors: list[str] = field(default_factory=list)
22
+
23
+
24
+ def _now() -> str:
25
+ return dt.datetime.now(dt.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
26
+
27
+
28
+ def load_state(path: Path) -> dict:
29
+ if not path.exists():
30
+ return {}
31
+ try:
32
+ return json.loads(path.read_text(encoding="utf-8"))
33
+ except ValueError:
34
+ return {}
35
+
36
+
37
+ def save_state(path: Path, state: dict) -> None:
38
+ path.parent.mkdir(parents=True, exist_ok=True)
39
+ tmp = path.with_suffix(path.suffix + ".tmp")
40
+ tmp.write_text(json.dumps(state, indent=2, sort_keys=True), encoding="utf-8")
41
+ tmp.replace(path) # atomic: a crash never leaves half-written state
42
+
43
+
44
+ def run(config: Config, renderer: Renderer) -> Result:
45
+ s = config.settings
46
+ state = load_state(s.state)
47
+ result = Result()
48
+
49
+ for page in config.pages:
50
+ try:
51
+ raw = renderer.text(page)
52
+ except Exception as e:
53
+ result.errors.append(f"{page.name}: {e}")
54
+ continue # keep the old fingerprint; an error is not a change
55
+
56
+ text = normalize(raw, page.ignore, s.use_default_noise)
57
+ if not text:
58
+ result.errors.append(f"{page.name}: rendered text is empty (check the selector {page.selector!r})")
59
+ continue
60
+ fp = fingerprint(text)
61
+ prev = state.get(page.url)
62
+
63
+ if prev and prev.get("fingerprint") == fp:
64
+ prev["checked_at"] = _now()
65
+ result.unchanged.append(page.name)
66
+ continue
67
+
68
+ shot_path = None
69
+ if page.screenshot:
70
+ s.captures.mkdir(parents=True, exist_ok=True)
71
+ stamp = dt.datetime.now(dt.timezone.utc).strftime("%Y%m%dT%H%M%SZ")
72
+ dest = s.captures / f"{page.name}-{stamp}.png"
73
+ try:
74
+ renderer.screenshot(page, dest)
75
+ shot_path = str(dest)
76
+ except Exception as e:
77
+ result.errors.append(f"{page.name}: screenshot failed: {e}")
78
+
79
+ if prev:
80
+ result.changed.append(Change(page.name, page.url, summarize(prev.get("text", ""), text), shot_path))
81
+ else:
82
+ result.baselined.append(page.name)
83
+ state[page.url] = {"fingerprint": fp, "text": text, "checked_at": _now(), "changed_at": _now()}
84
+
85
+ save_state(s.state, state)
86
+ return result
@@ -0,0 +1,71 @@
1
+ """Command line: `pagequiet init` and `pagequiet check`.
2
+
3
+ Exit codes for `check`: 0 nothing changed, 1 at least one page changed, 2 configuration
4
+ or rendering errors and no changes. Changes win over errors so CI jobs still alert.
5
+ """
6
+
7
+ from __future__ import annotations
8
+
9
+ import argparse
10
+ import sys
11
+ from pathlib import Path
12
+
13
+ from . import __version__
14
+ from .check import run
15
+ from .config import SAMPLE, ConfigError, load
16
+ from .notify import send
17
+ from .renderers import RenderError, make_renderer
18
+
19
+
20
+ def _init(path: Path) -> int:
21
+ if path.exists():
22
+ print(f"{path} already exists; leaving it alone.", file=sys.stderr)
23
+ return 2
24
+ path.write_text(SAMPLE, encoding="utf-8")
25
+ print(f"wrote {path}. Edit the [[page]] entries, then run: pagequiet check")
26
+ return 0
27
+
28
+
29
+ def _check(path: Path, quiet: bool) -> int:
30
+ try:
31
+ config = load(path)
32
+ renderer = make_renderer(config.settings)
33
+ except (ConfigError, RenderError) as e:
34
+ print(f"error: {e}", file=sys.stderr)
35
+ return 2
36
+ try:
37
+ result = run(config, renderer)
38
+ finally:
39
+ renderer.close()
40
+
41
+ errors = list(result.errors)
42
+ errors += send(config.settings.notify, result.changed)
43
+ if not quiet:
44
+ for name in result.baselined:
45
+ print(f"baseline recorded: {name}")
46
+ print(f"{len(result.changed)} changed, {len(result.unchanged)} unchanged, "
47
+ f"{len(result.baselined)} new, {len(result.errors)} errors")
48
+ for e in errors:
49
+ print(f"error: {e}", file=sys.stderr)
50
+ if result.changed:
51
+ return 1
52
+ return 2 if result.errors else 0
53
+
54
+
55
+ def main(argv: list[str] | None = None) -> int:
56
+ parser = argparse.ArgumentParser(prog="pagequiet", description="Website change detection that stays quiet until something real changes.")
57
+ parser.add_argument("--version", action="version", version=f"pagequiet {__version__}")
58
+ sub = parser.add_subparsers(dest="command", required=True)
59
+ p_init = sub.add_parser("init", help="write a sample pagequiet.toml")
60
+ p_init.add_argument("-c", "--config", default="pagequiet.toml")
61
+ p_check = sub.add_parser("check", help="render every page and report changes")
62
+ p_check.add_argument("-c", "--config", default="pagequiet.toml")
63
+ p_check.add_argument("-q", "--quiet", action="store_true", help="print only changes and errors")
64
+ args = parser.parse_args(argv)
65
+ if args.command == "init":
66
+ return _init(Path(args.config))
67
+ return _check(Path(args.config), args.quiet)
68
+
69
+
70
+ if __name__ == "__main__": # pragma: no cover
71
+ sys.exit(main())
@@ -0,0 +1,127 @@
1
+ """Load and validate pagequiet.toml."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import os
6
+ import re
7
+ import tomllib
8
+ from dataclasses import dataclass, field
9
+ from pathlib import Path
10
+
11
+
12
+ class ConfigError(ValueError):
13
+ pass
14
+
15
+
16
+ @dataclass
17
+ class Page:
18
+ url: str
19
+ name: str
20
+ selector: str = "body"
21
+ ignore: list[str] = field(default_factory=list)
22
+ hide: list[str] = field(default_factory=list)
23
+ screenshot: bool = True
24
+
25
+
26
+ @dataclass
27
+ class Settings:
28
+ renderer: str = "playwright"
29
+ state: Path = Path(".pagequiet/state.json")
30
+ captures: Path = Path(".pagequiet/captures")
31
+ notify: list[str] = field(default_factory=lambda: ["stdout"])
32
+ timeout: int = 45
33
+ use_default_noise: bool = True
34
+ width: int = 1280
35
+ height: int = 800
36
+
37
+
38
+ @dataclass
39
+ class Config:
40
+ settings: Settings
41
+ pages: list[Page]
42
+
43
+
44
+ RENDERERS = {"playwright", "snaprender"}
45
+ NOTIFIERS = {"stdout", "slack", "webhook"}
46
+
47
+
48
+ def _slug(url: str) -> str:
49
+ s = re.sub(r"^https?://", "", url)
50
+ s = re.sub(r"[^A-Za-z0-9]+", "-", s).strip("-").lower()
51
+ return s[:80] or "page"
52
+
53
+
54
+ def load(path: str | os.PathLike[str]) -> Config:
55
+ p = Path(path)
56
+ if not p.exists():
57
+ raise ConfigError(f"config not found: {p} (run `pagequiet init` to create one)")
58
+ try:
59
+ raw = tomllib.loads(p.read_text(encoding="utf-8"))
60
+ except tomllib.TOMLDecodeError as e:
61
+ raise ConfigError(f"{p}: {e}") from e
62
+
63
+ s = raw.get("settings", {})
64
+ settings = Settings(
65
+ renderer=s.get("renderer", "playwright"),
66
+ state=Path(s.get("state", ".pagequiet/state.json")),
67
+ captures=Path(s.get("captures", ".pagequiet/captures")),
68
+ notify=list(s.get("notify", ["stdout"])),
69
+ timeout=int(s.get("timeout", 45)),
70
+ use_default_noise=bool(s.get("use_default_noise", True)),
71
+ width=int(s.get("width", 1280)),
72
+ height=int(s.get("height", 800)),
73
+ )
74
+ if settings.renderer not in RENDERERS:
75
+ raise ConfigError(f"settings.renderer must be one of {sorted(RENDERERS)}, got {settings.renderer!r}")
76
+ unknown = [n for n in settings.notify if n not in NOTIFIERS]
77
+ if unknown:
78
+ raise ConfigError(f"settings.notify: unknown notifier(s) {unknown}; use {sorted(NOTIFIERS)}")
79
+
80
+ pages_raw = raw.get("page", [])
81
+ if not pages_raw:
82
+ raise ConfigError("no [[page]] entries in config")
83
+ pages: list[Page] = []
84
+ seen: set[str] = set()
85
+ for i, item in enumerate(pages_raw):
86
+ url = item.get("url")
87
+ if not url or not re.match(r"^https?://", url):
88
+ raise ConfigError(f"page #{i + 1}: url must start with http:// or https://")
89
+ for pattern in item.get("ignore", []):
90
+ try:
91
+ re.compile(pattern)
92
+ except re.error as e:
93
+ raise ConfigError(f"page {url}: bad ignore pattern {pattern!r}: {e}") from e
94
+ name = item.get("name") or _slug(url)
95
+ if name in seen:
96
+ name = f"{name}-{i + 1}"
97
+ seen.add(name)
98
+ pages.append(Page(
99
+ url=url,
100
+ name=name,
101
+ selector=item.get("selector", "body"),
102
+ ignore=list(item.get("ignore", [])),
103
+ hide=list(item.get("hide", [])),
104
+ screenshot=bool(item.get("screenshot", True)),
105
+ ))
106
+ return Config(settings=settings, pages=pages)
107
+
108
+
109
+ SAMPLE = """\
110
+ # pagequiet: website change detection that stays quiet until something real changes.
111
+
112
+ [settings]
113
+ renderer = "playwright" # free, runs Chromium locally. Or "snaprender" (needs SNAPRENDER_API_KEY)
114
+ state = ".pagequiet/state.json" # last fingerprint and text per page
115
+ captures = ".pagequiet/captures" # full-page screenshots of each changed version
116
+ notify = ["stdout"] # add "slack" (PAGEQUIET_SLACK_WEBHOOK) or "webhook" (PAGEQUIET_WEBHOOK_URL)
117
+
118
+ [[page]]
119
+ url = "https://example.com/"
120
+ selector = "body" # watch only the element that holds the content you care about
121
+
122
+ # [[page]]
123
+ # url = "https://example.com/pricing"
124
+ # selector = "main"
125
+ # ignore = ['Offer ends in \\d+ days'] # extra regex patterns to strip before comparing
126
+ # hide = [".testimonial-carousel"] # CSS selectors hidden before text and screenshot
127
+ """
@@ -0,0 +1,54 @@
1
+ """Turn rendered page text into a stable fingerprint and a readable diff.
2
+
3
+ Most false alarms in change monitoring come from content that changes on every
4
+ visit: relative times, view counters, dates, whitespace. These are stripped
5
+ before hashing. User patterns from the config are applied on top.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import difflib
11
+ import hashlib
12
+ import re
13
+ from collections.abc import Iterable
14
+
15
+ # Noise that changes on every visit without the page meaningfully changing.
16
+ # Kept deliberately short: every pattern is a kind of change you will never hear about.
17
+ DEFAULT_NOISE: tuple[str, ...] = (
18
+ r"\b\d+\s+(?:seconds?|secs?|minutes?|mins?|hours?|hrs?|days?|weeks?)\s+ago\b",
19
+ r"\b(?:just now|a moment ago|yesterday|today)\b",
20
+ r"\b(?:19|20)\d{2}-\d{2}-\d{2}(?:[ T]\d{2}:\d{2}(?::\d{2})?(?:\.\d+)?(?:Z|[+-]\d{2}:?\d{2})?)?\b",
21
+ r"\b\d{1,2}:\d{2}(?::\d{2})?\s?(?:AM|PM|am|pm)?\b",
22
+ r"\b\d[\d,.]*\s+(?:views|visitors|people (?:are )?(?:viewing|watching)|online now)\b",
23
+ )
24
+
25
+
26
+ def normalize(text: str, ignore: Iterable[str] = (), use_defaults: bool = True) -> str:
27
+ """Remove noise patterns and collapse whitespace, line by line."""
28
+ patterns = [*(DEFAULT_NOISE if use_defaults else ()), *ignore]
29
+ compiled = [re.compile(p, re.IGNORECASE) for p in patterns]
30
+ lines = []
31
+ for line in text.splitlines():
32
+ for rx in compiled:
33
+ line = rx.sub("", line)
34
+ line = re.sub(r"\s+", " ", line).strip()
35
+ if line:
36
+ lines.append(line)
37
+ return "\n".join(lines)
38
+
39
+
40
+ def fingerprint(normalized: str) -> str:
41
+ return hashlib.sha256(normalized.encode("utf-8")).hexdigest()
42
+
43
+
44
+ def summarize(old: str, new: str, limit: int = 8) -> list[str]:
45
+ """Changed lines as '- removed' / '+ added', at most `limit` of them."""
46
+ out: list[str] = []
47
+ for line in difflib.unified_diff(old.splitlines(), new.splitlines(), lineterm="", n=0):
48
+ if line.startswith(("+++", "---", "@@")):
49
+ continue
50
+ if line.startswith(("+", "-")):
51
+ out.append(f"{line[0]} {line[1:].strip()}")
52
+ if len(out) >= limit:
53
+ break
54
+ return out
@@ -0,0 +1,58 @@
1
+ """Send change notices: stdout, a Slack incoming webhook, or any JSON webhook."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import json
6
+ import os
7
+ import sys
8
+ import urllib.request
9
+ from dataclasses import dataclass
10
+
11
+
12
+ @dataclass
13
+ class Change:
14
+ name: str
15
+ url: str
16
+ lines: list[str]
17
+ screenshot: str | None
18
+
19
+
20
+ def _post(url: str, payload: dict) -> None:
21
+ req = urllib.request.Request(
22
+ url,
23
+ data=json.dumps(payload).encode("utf-8"),
24
+ headers={"Content-Type": "application/json", "User-Agent": "pagequiet"},
25
+ )
26
+ with urllib.request.urlopen(req, timeout=15):
27
+ pass
28
+
29
+
30
+ def to_text(c: Change) -> str:
31
+ body = "\n".join(f" {line}" for line in c.lines) or " (layout or wording changed; see screenshot)"
32
+ shot = f"\n screenshot: {c.screenshot}" if c.screenshot else ""
33
+ return f"CHANGED {c.name} {c.url}\n{body}{shot}"
34
+
35
+
36
+ def send(channels: list[str], changes: list[Change]) -> list[str]:
37
+ """Deliver each change to each channel. Returns error messages instead of raising."""
38
+ errors: list[str] = []
39
+ for c in changes:
40
+ for channel in channels:
41
+ try:
42
+ if channel == "stdout":
43
+ print(to_text(c), file=sys.stdout)
44
+ elif channel == "slack":
45
+ hook = os.environ.get("PAGEQUIET_SLACK_WEBHOOK")
46
+ if not hook:
47
+ errors.append("slack: PAGEQUIET_SLACK_WEBHOOK is not set")
48
+ continue
49
+ _post(hook, {"text": f"*Changed:* <{c.url}|{c.name}>\n```" + ("\n".join(c.lines) or "layout or wording changed") + "```"})
50
+ elif channel == "webhook":
51
+ hook = os.environ.get("PAGEQUIET_WEBHOOK_URL")
52
+ if not hook:
53
+ errors.append("webhook: PAGEQUIET_WEBHOOK_URL is not set")
54
+ continue
55
+ _post(hook, {"event": "page.changed", "name": c.name, "url": c.url, "changes": c.lines, "screenshot": c.screenshot})
56
+ except Exception as e: # a failed notice must never stop the run
57
+ errors.append(f"{channel}: {e}")
58
+ return errors
@@ -0,0 +1,128 @@
1
+ """Renderers turn a URL into visible text and, on change, a full-page screenshot.
2
+
3
+ PlaywrightRenderer runs Chromium locally and costs nothing. SnapRenderRenderer calls
4
+ the hosted SnapRender API, which removes cookie banners and ads before rendering
5
+ and needs no browser on the machine. Both implement the same two methods.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import json
11
+ import os
12
+ import urllib.error
13
+ import urllib.parse
14
+ import urllib.request
15
+ from pathlib import Path
16
+ from typing import Protocol
17
+
18
+ from .config import Page, Settings
19
+
20
+
21
+ class RenderError(RuntimeError):
22
+ pass
23
+
24
+
25
+ class Renderer(Protocol):
26
+ def text(self, page: Page) -> str: ...
27
+ def screenshot(self, page: Page, dest: Path) -> None: ...
28
+ def close(self) -> None: ...
29
+
30
+
31
+ class PlaywrightRenderer:
32
+ """Local Chromium through Playwright. Install with: pip install "pagequiet[playwright]"."""
33
+
34
+ def __init__(self, settings: Settings):
35
+ try:
36
+ from playwright.sync_api import sync_playwright
37
+ except ImportError as e: # pragma: no cover - depends on optional install
38
+ raise RenderError(
39
+ 'Playwright is not installed. Run: pip install "pagequiet[playwright]" && playwright install chromium'
40
+ ) from e
41
+ self._settings = settings
42
+ self._pw = sync_playwright().start()
43
+ self._browser = self._pw.chromium.launch()
44
+
45
+ def _open(self, page: Page):
46
+ ctx = self._browser.new_context(viewport={"width": self._settings.width, "height": self._settings.height})
47
+ tab = ctx.new_page()
48
+ try:
49
+ tab.goto(page.url, wait_until="networkidle", timeout=self._settings.timeout * 1000)
50
+ except Exception:
51
+ # Some pages never reach network idle (long polling, analytics); the DOM is usually ready anyway.
52
+ tab.wait_for_load_state("domcontentloaded", timeout=self._settings.timeout * 1000)
53
+ if page.hide:
54
+ css = ",".join(page.hide) + "{display:none!important}"
55
+ tab.add_style_tag(content=css)
56
+ return ctx, tab
57
+
58
+ def text(self, page: Page) -> str:
59
+ ctx, tab = self._open(page)
60
+ try:
61
+ loc = tab.locator(page.selector).first
62
+ if loc.count() == 0:
63
+ raise RenderError(f"selector {page.selector!r} not found on {page.url}")
64
+ return loc.inner_text(timeout=self._settings.timeout * 1000)
65
+ finally:
66
+ ctx.close()
67
+
68
+ def screenshot(self, page: Page, dest: Path) -> None:
69
+ ctx, tab = self._open(page)
70
+ try:
71
+ tab.screenshot(path=str(dest), full_page=True)
72
+ finally:
73
+ ctx.close()
74
+
75
+ def close(self) -> None:
76
+ self._browser.close()
77
+ self._pw.stop()
78
+
79
+
80
+ class SnapRenderRenderer:
81
+ """Hosted rendering through the SnapRender API (https://snap-render.com).
82
+
83
+ Cookie banners and ads are removed before text and screenshots are taken. Text comes
84
+ from GET /v1/extract?type=text scoped by selector; screenshots from GET /v1/screenshot.
85
+ Note: `hide` applies to screenshots only; scope the text with `selector` instead.
86
+ """
87
+
88
+ API = "https://app.snap-render.com"
89
+
90
+ def __init__(self, settings: Settings, api_key: str | None = None):
91
+ self._settings = settings
92
+ self._key = api_key or os.environ.get("SNAPRENDER_API_KEY", "")
93
+ if not self._key:
94
+ raise RenderError("SNAPRENDER_API_KEY is not set (get a free key at https://snap-render.com)")
95
+
96
+ def _get(self, path: str, params: dict[str, str]) -> bytes:
97
+ url = f"{self.API}{path}?{urllib.parse.urlencode(params)}"
98
+ req = urllib.request.Request(url, headers={"X-API-Key": self._key, "User-Agent": "pagequiet"})
99
+ try:
100
+ with urllib.request.urlopen(req, timeout=self._settings.timeout + 30) as resp:
101
+ return resp.read()
102
+ except urllib.error.HTTPError as e:
103
+ detail = e.read().decode("utf-8", "replace")[:300]
104
+ raise RenderError(f"SnapRender {path} returned {e.code}: {detail}") from e
105
+ except urllib.error.URLError as e:
106
+ raise RenderError(f"SnapRender {path} unreachable: {e.reason}") from e
107
+
108
+ def text(self, page: Page) -> str:
109
+ body = self._get("/v1/extract", {"url": page.url, "type": "text", "selector": page.selector})
110
+ try:
111
+ return json.loads(body)["content"]
112
+ except (ValueError, KeyError) as e:
113
+ raise RenderError(f"unexpected extract response for {page.url}") from e
114
+
115
+ def screenshot(self, page: Page, dest: Path) -> None:
116
+ params = {"url": page.url, "format": "png", "full_page": "true", "width": str(self._settings.width)}
117
+ if page.hide:
118
+ params["hide_selectors"] = ",".join(page.hide)
119
+ dest.write_bytes(self._get("/v1/screenshot", params))
120
+
121
+ def close(self) -> None:
122
+ pass
123
+
124
+
125
+ def make_renderer(settings: Settings) -> Renderer:
126
+ if settings.renderer == "snaprender":
127
+ return SnapRenderRenderer(settings)
128
+ return PlaywrightRenderer(settings)
@@ -0,0 +1,134 @@
1
+ Metadata-Version: 2.4
2
+ Name: pagequiet
3
+ Version: 0.1.0
4
+ Summary: Website change detection that stays quiet until something real changes.
5
+ Author: User0856
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/User0856/pagequiet
8
+ Project-URL: Issues, https://github.com/User0856/pagequiet/issues
9
+ Keywords: website monitoring,change detection,web page diff,playwright,screenshot,visualping alternative
10
+ Classifier: Programming Language :: Python :: 3
11
+ Classifier: Environment :: Console
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Topic :: Internet :: WWW/HTTP :: Site Management :: Link Checking
14
+ Classifier: Topic :: Software Development :: Testing
15
+ Requires-Python: >=3.11
16
+ Description-Content-Type: text/markdown
17
+ License-File: LICENSE
18
+ Provides-Extra: playwright
19
+ Requires-Dist: playwright>=1.40; extra == "playwright"
20
+ Provides-Extra: dev
21
+ Requires-Dist: pytest>=8; extra == "dev"
22
+ Dynamic: license-file
23
+
24
+ # pagequiet
25
+
26
+ Website change detection that stays quiet until something real changes.
27
+
28
+ Most change monitors compare HTML, so they alert on every rotating ad slot, CSRF token and "updated 3 minutes ago". After a week nobody reads the alerts. `pagequiet` compares what a reader sees instead: the rendered text of one element on the page, with relative times, counters and whitespace stripped. When that changes, you get the changed lines and a full-page screenshot of the new version.
29
+
30
+ ```
31
+ $ pagequiet check
32
+ CHANGED example-com-pricing https://example.com/pricing
33
+ - Pro $49 / month
34
+ + Pro $59 / month
35
+ screenshot: .pagequiet/captures/example-com-pricing-20261005T080012Z.png
36
+ 1 changed, 11 unchanged, 0 new, 0 errors
37
+ ```
38
+
39
+ It is a single command with a TOML file, built for cron and CI: exit code 1 means something changed.
40
+
41
+ ## Install
42
+
43
+ ```bash
44
+ pip install "pagequiet[playwright]"
45
+ playwright install chromium
46
+ ```
47
+
48
+ Python 3.11 or later. The Playwright extra renders pages with a local Chromium and costs nothing. If you would rather not run a browser (CI runners, small servers), see [Renderers](#renderers).
49
+
50
+ ## Quick start
51
+
52
+ ```bash
53
+ pagequiet init # writes pagequiet.toml
54
+ # edit the [[page]] entries
55
+ pagequiet check # first run records a baseline
56
+ pagequiet check # later runs report changes
57
+ ```
58
+
59
+ ## Configuration
60
+
61
+ ```toml
62
+ [settings]
63
+ renderer = "playwright" # or "snaprender"
64
+ state = ".pagequiet/state.json" # last text and fingerprint per page
65
+ captures = ".pagequiet/captures" # full-page screenshot of each new version
66
+ notify = ["stdout", "slack"] # stdout, slack, webhook
67
+ timeout = 45 # seconds per page
68
+ use_default_noise = true # strip relative times, timestamps, view counters
69
+
70
+ [[page]]
71
+ url = "https://example.com/pricing"
72
+ selector = "main" # watch one element, not the whole page
73
+ ignore = ['Offer ends in \d+ days'] # extra regex patterns removed before comparing
74
+ hide = [".testimonial-carousel"] # hidden before text and screenshot
75
+
76
+ [[page]]
77
+ url = "https://example.com/legal/terms"
78
+ selector = "article"
79
+ screenshot = false # text only
80
+ ```
81
+
82
+ Notifiers read their targets from the environment:
83
+
84
+ | Notifier | Variable | Sends |
85
+ |---|---|---|
86
+ | `stdout` | none | the changed lines and screenshot path |
87
+ | `slack` | `PAGEQUIET_SLACK_WEBHOOK` | a message to a Slack incoming webhook |
88
+ | `webhook` | `PAGEQUIET_WEBHOOK_URL` | JSON: `{"event": "page.changed", "name", "url", "changes", "screenshot"}` |
89
+
90
+ Exit codes: `0` nothing changed, `1` at least one page changed, `2` errors and no changes.
91
+
92
+ ## What counts as a change
93
+
94
+ 1. Render the page and take the visible text of `selector` (default `body`).
95
+ 2. Remove noise: relative times ("5 minutes ago", "just now"), ISO timestamps and clock times, view and visitor counters, extra whitespace, plus your `ignore` patterns.
96
+ 3. Hash the result. Same hash as last run: quiet. Different: report the changed lines and take a screenshot.
97
+
98
+ The defaults are deliberately few. Every pattern you strip is a kind of change you will never be told about, so add `ignore` patterns only after you have seen one cause a false alarm. If timestamps on a page are meaningful to you, set `use_default_noise = false`.
99
+
100
+ The biggest single improvement is `selector`. Watching `main` or the pricing table instead of the whole page ignores navigation, footers and "latest posts" widgets that change for reasons you do not care about.
101
+
102
+ A failed render or an empty selector is reported as an error and never as a change, and the previous text is kept, so a flaky site does not produce a false alarm the next time it loads.
103
+
104
+ ## Renderers
105
+
106
+ **`playwright`** (default) runs Chromium locally through Playwright. Free, private, and fine for most pages. Cookie consent banners are part of the page, so if one varies between visits, scope `selector` to the content or add the banner to `hide`.
107
+
108
+ **`snaprender`** uses the hosted [SnapRender](https://snap-render.com) API to render pages, so no browser runs on your machine. Cookie banners and ads are removed before the text and screenshot are taken. It needs `SNAPRENDER_API_KEY`; the free plan covers 200 renders a month. Each check is one render, plus one when a page changed. Disclosure: I build SnapRender. `pagequiet` works fully without it and always will.
109
+
110
+ ## Run it on a schedule
111
+
112
+ Cron, every six hours:
113
+
114
+ ```bash
115
+ 0 */6 * * * cd /opt/watch && pagequiet check -q >> pagequiet.log 2>&1
116
+ ```
117
+
118
+ GitHub Actions, keeping state in the repository: see [`examples/github-actions.yml`](examples/github-actions.yml). Scheduled workflows run at most every five minutes, can start late at busy times, and are paused in public repositories after 60 days without activity; committing the state file each run counts as activity.
119
+
120
+ ## Limits
121
+
122
+ - Pages behind a login are not supported.
123
+ - Text comparison misses purely visual changes (a new image, a broken layout). Compare the screenshots in `captures/` for those, or keep `screenshot = true` and review them when text changes.
124
+ - `hide` applies to the SnapRender renderer's screenshots only; scope its text with `selector`.
125
+
126
+ ## Development
127
+
128
+ ```bash
129
+ python -m venv .venv && . .venv/bin/activate
130
+ pip install -e ".[dev,playwright]"
131
+ pytest
132
+ ```
133
+
134
+ MIT licensed.
@@ -0,0 +1,18 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ pagequiet/__init__.py
5
+ pagequiet/__main__.py
6
+ pagequiet/check.py
7
+ pagequiet/cli.py
8
+ pagequiet/config.py
9
+ pagequiet/normalize.py
10
+ pagequiet/notify.py
11
+ pagequiet/renderers.py
12
+ pagequiet.egg-info/PKG-INFO
13
+ pagequiet.egg-info/SOURCES.txt
14
+ pagequiet.egg-info/dependency_links.txt
15
+ pagequiet.egg-info/entry_points.txt
16
+ pagequiet.egg-info/requires.txt
17
+ pagequiet.egg-info/top_level.txt
18
+ tests/test_pagequiet.py
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ pagequiet = pagequiet.cli:main
@@ -0,0 +1,6 @@
1
+
2
+ [dev]
3
+ pytest>=8
4
+
5
+ [playwright]
6
+ playwright>=1.40
@@ -0,0 +1 @@
1
+ pagequiet
@@ -0,0 +1,35 @@
1
+ [build-system]
2
+ requires = ["setuptools>=69"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "pagequiet"
7
+ version = "0.1.0"
8
+ description = "Website change detection that stays quiet until something real changes."
9
+ readme = "README.md"
10
+ license = "MIT"
11
+ requires-python = ">=3.11"
12
+ authors = [{ name = "User0856" }]
13
+ keywords = ["website monitoring", "change detection", "web page diff", "playwright", "screenshot", "visualping alternative"]
14
+ classifiers = [
15
+ "Programming Language :: Python :: 3",
16
+ "Environment :: Console",
17
+ "Intended Audience :: Developers",
18
+ "Topic :: Internet :: WWW/HTTP :: Site Management :: Link Checking",
19
+ "Topic :: Software Development :: Testing",
20
+ ]
21
+ dependencies = []
22
+
23
+ [project.optional-dependencies]
24
+ playwright = ["playwright>=1.40"]
25
+ dev = ["pytest>=8"]
26
+
27
+ [project.scripts]
28
+ pagequiet = "pagequiet.cli:main"
29
+
30
+ [project.urls]
31
+ Homepage = "https://github.com/User0856/pagequiet"
32
+ Issues = "https://github.com/User0856/pagequiet/issues"
33
+
34
+ [tool.setuptools]
35
+ packages = ["pagequiet"]
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,157 @@
1
+ import json
2
+ from pathlib import Path
3
+
4
+ import pytest
5
+
6
+ from pagequiet.check import run
7
+ from pagequiet.cli import main
8
+ from pagequiet.config import ConfigError, Page, load
9
+ from pagequiet.normalize import fingerprint, normalize, summarize
10
+
11
+
12
+ # ---------- normalize ----------
13
+
14
+ def test_relative_times_counters_and_whitespace_do_not_change_the_fingerprint():
15
+ a = "Pricing\nPro $49 / month\nUpdated 3 minutes ago\n1,204 people viewing"
16
+ b = "Pricing\n Pro $49 / month \nUpdated 2 hours ago\n987 people viewing\n\n"
17
+ assert fingerprint(normalize(a)) == fingerprint(normalize(b))
18
+
19
+
20
+ def test_a_real_change_changes_the_fingerprint():
21
+ a = "Pro $49 / month"
22
+ b = "Pro $59 / month"
23
+ assert fingerprint(normalize(a)) != fingerprint(normalize(b))
24
+
25
+
26
+ def test_iso_timestamps_are_ignored_by_default_but_can_be_kept():
27
+ a, b = "Build 2026-10-05T08:00:00Z ok", "Build 2026-10-06T09:30:00Z ok"
28
+ assert normalize(a) == normalize(b)
29
+ assert normalize(a, use_defaults=False) != normalize(b, use_defaults=False)
30
+
31
+
32
+ def test_user_ignore_patterns_apply():
33
+ a, b = "Offer ends in 3 days\nPlan A", "Offer ends in 2 days\nPlan A"
34
+ assert normalize(a) != normalize(b)
35
+ assert normalize(a, [r"Offer ends in \d+ days"]) == normalize(b, [r"Offer ends in \d+ days"])
36
+
37
+
38
+ def test_summary_lists_removed_and_added_lines():
39
+ lines = summarize("Plan A\nPro $49\nFooter", "Plan A\nPro $59\nFooter")
40
+ assert lines == ["- Pro $49", "+ Pro $59"]
41
+
42
+
43
+ def test_summary_is_capped():
44
+ old = "\n".join(f"line {i}" for i in range(50))
45
+ new = "\n".join(f"LINE {i}" for i in range(50))
46
+ assert len(summarize(old, new, limit=5)) == 5
47
+
48
+
49
+ # ---------- config ----------
50
+
51
+ def write(tmp_path: Path, body: str) -> Path:
52
+ p = tmp_path / "pagequiet.toml"
53
+ p.write_text(body, encoding="utf-8")
54
+ return p
55
+
56
+
57
+ def test_config_defaults_and_names(tmp_path):
58
+ cfg = load(write(tmp_path, '[[page]]\nurl = "https://example.com/pricing"\n'))
59
+ assert cfg.settings.renderer == "playwright"
60
+ assert cfg.pages[0].name == "example-com-pricing"
61
+ assert cfg.pages[0].selector == "body"
62
+
63
+
64
+ @pytest.mark.parametrize("body,needle", [
65
+ ('[settings]\nrenderer = "selenium"\n[[page]]\nurl = "https://x.y"\n', "renderer"),
66
+ ('[[page]]\nurl = "ftp://x.y"\n', "http"),
67
+ ('[settings]\nnotify = ["email"]\n[[page]]\nurl = "https://x.y"\n', "notifier"),
68
+ ('[[page]]\nurl = "https://x.y"\nignore = ["(unclosed"]\n', "ignore pattern"),
69
+ ('[settings]\nrenderer = "playwright"\n', r"no \[\[page\]\] entries"),
70
+ ])
71
+ def test_config_errors_are_explicit(tmp_path, body, needle):
72
+ with pytest.raises(ConfigError, match=needle):
73
+ load(write(tmp_path, body))
74
+
75
+
76
+ def test_missing_config_suggests_init(tmp_path):
77
+ with pytest.raises(ConfigError, match="pagequiet init"):
78
+ load(tmp_path / "nope.toml")
79
+
80
+
81
+ # ---------- run ----------
82
+
83
+ class FakeRenderer:
84
+ def __init__(self, texts):
85
+ self.texts = texts
86
+ self.shots = []
87
+
88
+ def text(self, page: Page) -> str:
89
+ value = self.texts[page.url]
90
+ if isinstance(value, Exception):
91
+ raise value
92
+ return value
93
+
94
+ def screenshot(self, page: Page, dest: Path) -> None:
95
+ dest.write_bytes(b"\x89PNG fake")
96
+ self.shots.append(dest)
97
+
98
+ def close(self):
99
+ pass
100
+
101
+
102
+ def cfg_for(tmp_path, urls):
103
+ pages = "\n".join(f'[[page]]\nurl = "{u}"\n' for u in urls)
104
+ return load(write(tmp_path, f'[settings]\nstate = "{tmp_path}/state.json"\ncaptures = "{tmp_path}/shots"\n{pages}'))
105
+
106
+
107
+ def test_first_run_baselines_then_quiet_then_alerts(tmp_path):
108
+ cfg = cfg_for(tmp_path, ["https://a.test/"])
109
+ r1 = run(cfg, FakeRenderer({"https://a.test/": "Hello\nPrice $10\nUpdated 1 minute ago"}))
110
+ assert r1.baselined == ["a-test"] and not r1.changed
111
+
112
+ r2 = run(cfg, FakeRenderer({"https://a.test/": "Hello\nPrice $10\nUpdated 5 minutes ago"}))
113
+ assert r2.unchanged == ["a-test"] and not r2.changed
114
+
115
+ fake = FakeRenderer({"https://a.test/": "Hello\nPrice $12\nUpdated just now"})
116
+ r3 = run(cfg, fake)
117
+ assert len(r3.changed) == 1
118
+ assert r3.changed[0].lines == ["- Price $10", "+ Price $12"]
119
+ assert r3.changed[0].screenshot and Path(r3.changed[0].screenshot).exists()
120
+
121
+
122
+ def test_render_errors_keep_the_previous_fingerprint(tmp_path):
123
+ cfg = cfg_for(tmp_path, ["https://a.test/"])
124
+ run(cfg, FakeRenderer({"https://a.test/": "v1"}))
125
+ r = run(cfg, FakeRenderer({"https://a.test/": RuntimeError("timeout")}))
126
+ assert r.errors and not r.changed
127
+ state = json.loads((tmp_path / "state.json").read_text())
128
+ assert state["https://a.test/"]["text"] == "v1"
129
+
130
+
131
+ def test_empty_text_is_an_error_not_a_change(tmp_path):
132
+ cfg = cfg_for(tmp_path, ["https://a.test/"])
133
+ run(cfg, FakeRenderer({"https://a.test/": "v1"}))
134
+ r = run(cfg, FakeRenderer({"https://a.test/": " \n "}))
135
+ assert r.errors and not r.changed
136
+
137
+
138
+ # ---------- cli ----------
139
+
140
+ def test_init_writes_a_loadable_sample(tmp_path, capsys):
141
+ target = tmp_path / "pagequiet.toml"
142
+ assert main(["init", "-c", str(target)]) == 0
143
+ assert load(target).pages[0].url == "https://example.com/"
144
+ assert main(["init", "-c", str(target)]) == 2 # never overwrites
145
+
146
+
147
+ def test_check_exit_codes(tmp_path, monkeypatch, capsys):
148
+ cfg_path = write(tmp_path, f'[settings]\nstate = "{tmp_path}/s.json"\ncaptures = "{tmp_path}/c"\n[[page]]\nurl = "https://a.test/"\n')
149
+ texts = {"https://a.test/": "v1"}
150
+ monkeypatch.setattr("pagequiet.cli.make_renderer", lambda settings: FakeRenderer(texts))
151
+ assert main(["check", "-c", str(cfg_path)]) == 0 # baseline
152
+ assert main(["check", "-c", str(cfg_path)]) == 0 # unchanged
153
+ texts["https://a.test/"] = "v2"
154
+ assert main(["check", "-c", str(cfg_path)]) == 1 # changed
155
+ out = capsys.readouterr().out
156
+ assert "CHANGED a-test https://a.test/" in out and "+ v2" in out
157
+ assert main(["check", "-c", str(tmp_path / "missing.toml")]) == 2