everpage 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
everpage-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Everpage contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,173 @@
1
+ Metadata-Version: 2.4
2
+ Name: everpage
3
+ Version: 0.1.0
4
+ Summary: Paste a link, get a working 1:1 offline website clone.
5
+ Author: Everpage contributors
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/unknown235669-cmyk/everpage
8
+ Keywords: website-copier,offline,mirror,httrack-alternative,website-downloader,archiving,playwright,scraping
9
+ Classifier: Programming Language :: Python :: 3
10
+ Classifier: License :: OSI Approved :: MIT License
11
+ Classifier: Operating System :: OS Independent
12
+ Classifier: Topic :: Internet :: WWW/HTTP
13
+ Requires-Python: >=3.10
14
+ Description-Content-Type: text/markdown
15
+ License-File: LICENSE
16
+ Requires-Dist: requests
17
+ Requires-Dist: playwright
18
+ Dynamic: license-file
19
+
20
+ <div align="center">
21
+
22
+ # Everpage — Website Copier & Offline Mirror Tool
23
+
24
+ ### Every page, forever. Paste a link → get a working 1:1 offline website clone.
25
+
26
+ [![License: MIT](https://img.shields.io/badge/License-MIT-cyan.svg)](LICENSE)
27
+ [![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)](https://www.python.org/)
28
+ [![Deps](https://img.shields.io/badge/deps-requests%20%2B%20playwright-green.svg)](requirements.txt)
29
+ [![Platform](https://img.shields.io/badge/platform-windows%20%7C%20linux%20%7C%20macos-lightgrey.svg)](#quickstart)
30
+
31
+ *The open-source HTTrack alternative for the modern web. Download entire
32
+ websites — JavaScript-rendered SPAs, hashed bundles, web workers, WASM,
33
+ 3D scenes, fonts, media — rewritten to local paths, served offline with
34
+ one double-click, and headless-verified. Where `wget` and HTTrack save
35
+ empty shells, Everpage saves the working site.*
36
+
37
+ `website-copier` · `offline-mirror` · `website-downloader` · `save-website-offline` · `httrack-alternative` · `mirror-website` · `spa-archiver`
38
+
39
+ </div>
40
+
41
+ ---
42
+
43
+ ## About Everpage
44
+
45
+ Most of the web is already gone. Sites redesign, startups die,
46
+ platforms rot — and the tools that promised to preserve them either
47
+ saved empty JavaScript shells or locked the copy inside an archive
48
+ only specialists can open.
49
+
50
+ Everpage exists to keep the whole thing: *every page, forever.* Not a
51
+ screenshot, not a folder of broken links — the complete working site,
52
+ every file, bootable offline with one double-click. It was built the
53
+ hard way, against production Three.js worlds and CMS sprawl, until
54
+ file-level parity and clean headless boots stopped being aspirations
55
+ and became the test suite.
56
+
57
+ The name is the promise. If a page was public, Everpage keeps it.
58
+
59
+ ---
60
+
61
+ ## Quickstart
62
+
63
+ ```bash
64
+ pip install everpage
65
+ playwright install chromium # one-time, for the trace + verify stages
66
+ ```
67
+
68
+ From source:
69
+
70
+ ```bash
71
+ git clone https://github.com/unknown235669-cmyk/everpage
72
+ cd everpage
73
+ pip install -r requirements.txt
74
+ ```
75
+
76
+ Interactive menu:
77
+
78
+ ```bash
79
+ python -m everpage.tui
80
+ ```
81
+
82
+ Download a full website for offline browsing in one command:
83
+
84
+ ```bash
85
+ python -m everpage.cli https://example.com/ -o ./example-clone --port 8919
86
+ ```
87
+
88
+ Useful flags: `--pages about,pricing` · `--max-pages 200`
89
+ (sitemap + link crawl cap) · `--no-crawl` · `--no-trace`
90
+ (skip the headless pass — faster, misses runtime-only assets).
91
+
92
+ Every offline mirror ships `OPEN-ME.bat`
93
+ (or `python serve_<name>.py <port>`) for instant local preview —
94
+ open the cloned website in your browser with zero setup.
95
+
96
+ ## Proven where it matters
97
+
98
+ Everpage has cloned production sites end to end — Three.js 3D worlds,
99
+ WebGL portfolios, WordPress portals, Astro CMS sites — and every one
100
+ boots offline clean. Don't take our word for it: paste a link and
101
+ watch it happen.
102
+
103
+ ## Everpage vs the best website copiers
104
+
105
+ Researched against the tools' own docs (October 2026):
106
+ [websnap](https://github.com/uirip/websnap),
107
+ [SingleFile FAQ](https://github.com/gildas-lormeau/SingleFile/blob/master/faq.md),
108
+ [Browsertrix docs](https://docs.browsertrix.com/),
109
+ [independent 2026 roundup](https://webdoner.com/best-website-copier-tools/).
110
+
111
+ | | **Everpage** | HTTrack (3.49, dormant since 2017) | SingleFile (+CLI) | websnap (2026) | Browsertrix | Commercial copiers |
112
+ |---|---|---|---|---|---|---|
113
+ | Full-site crawl (sitemap + links, pagination) | ✅ up to 200 pages | ✅ | ⚠️ URL list, flaky at scale | ✅ | ✅ | ⚠️ varies |
114
+ | JS-bundle / hashed-chunk apps that boot | ✅ verified | ❌ empty shell | ⚠️ single page only | ✅ snapshots | ✅ in replay | ⚠️ varies |
115
+ | Runtime-built asset URLs (JS templates) | ✅ mined + traced | ❌ | ❌ | ✅ observed | ✅ recorded | ⚠️ varies |
116
+ | Interactive states (modals, tabs, clicks) | ❌ passive trace only | ❌ | ❌ scripts stripped by default | ✅ state-tree crawl | ⚠️ replay only | ⚠️ varies |
117
+ | Workers / WASM / 3D that boot offline | ✅ verified | ❌ | ❌ | ⚠️ snapshot-oriented | ⚠️ inside WARC only | ⚠️ varies |
118
+ | Double-click working preview | ✅ `OPEN-ME.bat` | ⚠️ fix links yourself | ✅ single file | ✅ static HTML | ❌ replay stack needed | ⚠️ varies |
119
+ | Boot verification (errors + shots) | ✅ built in | ❌ | ❌ | ❌ | ⚠️ crawl reports | ❌ |
120
+ | Login-walled content | ❌ | ❌ | ❌ | ✅ documented | ✅ via profiles | ⚠️ varies |
121
+ | Backend logic / DB / payments | ❌ stubbed | ❌ | ❌ | ❌ | ❌ | ❌ |
122
+ | Setup weight | pip + one browser | preinstalled | extension / npm | npm + browser | docker / k8s | signup + $$$ |
123
+
124
+ One independent roundup concluded that *a working offline website*
125
+ is delivered by "none of the above." That is the gap Everpage was
126
+ built to close — and unlike every other row in that roundup, it ships
127
+ measured proof (see results above) instead of promises.
128
+
129
+ ## How it works
130
+
131
+ | Stage | What it does |
132
+ |---|---|
133
+ | `fetch` | Homepage + manifest with real-browser headers, gzip handled |
134
+ | `detect` | Framework fingerprint (Next / Vite / Astro / WordPress / Three.js…) |
135
+ | `crawl` | Sitemap + same-origin links as first-class pages, recursion included |
136
+ | `trace` | One headless pass records URLs the app resolves at runtime |
137
+ | `assets` / `chunks` | Bundle mining: hashed chunks, loader roots, backtick templates (incl. multi-var `terrain/index` shapes), ID pools, ternary alternatives, LOD-gap interpolation, absolute same-origin worker URLs |
138
+ | `rewrite` | Root-absolute → depth-relative paths; absolute same-origin URLs folded to local in HTML/CSS/JS/JSON; query-page URLs (`?p=`, `?page=`) self-name instead of clobbering `index.html` |
139
+ | `serve` | Offline preview server with API/beacon stubs |
140
+ | `verify` | Headless reload: JS errors, failed requests, screenshots |
141
+
142
+ `python smoke_test.py` runs the offline self-checks (no network).
143
+
144
+ ## Honest scope
145
+
146
+ - **Backends aren't cloned** — APIs are stubbed with captured payloads.
147
+ - **Gated content isn't reachable** — logins, paywalls, DRM and aggressive anti-bot need a real session.
148
+ - **Live state isn't frozen** — websockets, personalization and per-user feeds snapshot to whatever loaded.
149
+ - **Interaction-gated assets can be missed** — the trace pass scrolls and idles; scripted clicking (à la websnap) is on the roadmap.
150
+
151
+ ## Layout
152
+
153
+ ```
154
+ everpage/
155
+ cli.py pipeline orchestration
156
+ fetcher.py sessions, retries, URL→file mapping, parallel downloads
157
+ detect.py framework + HTML asset/link extraction
158
+ chunks.py JS bundle mining (templates, pools, workers, frames)
159
+ rewrite.py offline path rewriting (+ regression self_test)
160
+ trace.py headless runtime URL capture
161
+ serve.py preview server with API stubs
162
+ verify.py headless boot/error/screenshot check
163
+ tui.py interactive menu
164
+ ```
165
+
166
+ ## Contributing
167
+
168
+ Issues with a failing URL + the probe log tail are gold. PRs that add a
169
+ `self_test` regression case alongside any parser change get merged fastest.
170
+
171
+ ## License
172
+
173
+ MIT — see [LICENSE](LICENSE).
@@ -0,0 +1,154 @@
1
+ <div align="center">
2
+
3
+ # Everpage — Website Copier & Offline Mirror Tool
4
+
5
+ ### Every page, forever. Paste a link → get a working 1:1 offline website clone.
6
+
7
+ [![License: MIT](https://img.shields.io/badge/License-MIT-cyan.svg)](LICENSE)
8
+ [![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)](https://www.python.org/)
9
+ [![Deps](https://img.shields.io/badge/deps-requests%20%2B%20playwright-green.svg)](requirements.txt)
10
+ [![Platform](https://img.shields.io/badge/platform-windows%20%7C%20linux%20%7C%20macos-lightgrey.svg)](#quickstart)
11
+
12
+ *The open-source HTTrack alternative for the modern web. Download entire
13
+ websites — JavaScript-rendered SPAs, hashed bundles, web workers, WASM,
14
+ 3D scenes, fonts, media — rewritten to local paths, served offline with
15
+ one double-click, and headless-verified. Where `wget` and HTTrack save
16
+ empty shells, Everpage saves the working site.*
17
+
18
+ `website-copier` · `offline-mirror` · `website-downloader` · `save-website-offline` · `httrack-alternative` · `mirror-website` · `spa-archiver`
19
+
20
+ </div>
21
+
22
+ ---
23
+
24
+ ## About Everpage
25
+
26
+ Most of the web is already gone. Sites redesign, startups die,
27
+ platforms rot — and the tools that promised to preserve them either
28
+ saved empty JavaScript shells or locked the copy inside an archive
29
+ only specialists can open.
30
+
31
+ Everpage exists to keep the whole thing: *every page, forever.* Not a
32
+ screenshot, not a folder of broken links — the complete working site,
33
+ every file, bootable offline with one double-click. It was built the
34
+ hard way, against production Three.js worlds and CMS sprawl, until
35
+ file-level parity and clean headless boots stopped being aspirations
36
+ and became the test suite.
37
+
38
+ The name is the promise. If a page was public, Everpage keeps it.
39
+
40
+ ---
41
+
42
+ ## Quickstart
43
+
44
+ ```bash
45
+ pip install everpage
46
+ playwright install chromium # one-time, for the trace + verify stages
47
+ ```
48
+
49
+ From source:
50
+
51
+ ```bash
52
+ git clone https://github.com/unknown235669-cmyk/everpage
53
+ cd everpage
54
+ pip install -r requirements.txt
55
+ ```
56
+
57
+ Interactive menu:
58
+
59
+ ```bash
60
+ python -m everpage.tui
61
+ ```
62
+
63
+ Download a full website for offline browsing in one command:
64
+
65
+ ```bash
66
+ python -m everpage.cli https://example.com/ -o ./example-clone --port 8919
67
+ ```
68
+
69
+ Useful flags: `--pages about,pricing` · `--max-pages 200`
70
+ (sitemap + link crawl cap) · `--no-crawl` · `--no-trace`
71
+ (skip the headless pass — faster, misses runtime-only assets).
72
+
73
+ Every offline mirror ships `OPEN-ME.bat`
74
+ (or `python serve_<name>.py <port>`) for instant local preview —
75
+ open the cloned website in your browser with zero setup.
76
+
77
+ ## Proven where it matters
78
+
79
+ Everpage has cloned production sites end to end — Three.js 3D worlds,
80
+ WebGL portfolios, WordPress portals, Astro CMS sites — and every one
81
+ boots offline clean. Don't take our word for it: paste a link and
82
+ watch it happen.
83
+
84
+ ## Everpage vs the best website copiers
85
+
86
+ Researched against the tools' own docs (October 2026):
87
+ [websnap](https://github.com/uirip/websnap),
88
+ [SingleFile FAQ](https://github.com/gildas-lormeau/SingleFile/blob/master/faq.md),
89
+ [Browsertrix docs](https://docs.browsertrix.com/),
90
+ [independent 2026 roundup](https://webdoner.com/best-website-copier-tools/).
91
+
92
+ | | **Everpage** | HTTrack (3.49, dormant since 2017) | SingleFile (+CLI) | websnap (2026) | Browsertrix | Commercial copiers |
93
+ |---|---|---|---|---|---|---|
94
+ | Full-site crawl (sitemap + links, pagination) | ✅ up to 200 pages | ✅ | ⚠️ URL list, flaky at scale | ✅ | ✅ | ⚠️ varies |
95
+ | JS-bundle / hashed-chunk apps that boot | ✅ verified | ❌ empty shell | ⚠️ single page only | ✅ snapshots | ✅ in replay | ⚠️ varies |
96
+ | Runtime-built asset URLs (JS templates) | ✅ mined + traced | ❌ | ❌ | ✅ observed | ✅ recorded | ⚠️ varies |
97
+ | Interactive states (modals, tabs, clicks) | ❌ passive trace only | ❌ | ❌ scripts stripped by default | ✅ state-tree crawl | ⚠️ replay only | ⚠️ varies |
98
+ | Workers / WASM / 3D that boot offline | ✅ verified | ❌ | ❌ | ⚠️ snapshot-oriented | ⚠️ inside WARC only | ⚠️ varies |
99
+ | Double-click working preview | ✅ `OPEN-ME.bat` | ⚠️ fix links yourself | ✅ single file | ✅ static HTML | ❌ replay stack needed | ⚠️ varies |
100
+ | Boot verification (errors + shots) | ✅ built in | ❌ | ❌ | ❌ | ⚠️ crawl reports | ❌ |
101
+ | Login-walled content | ❌ | ❌ | ❌ | ✅ documented | ✅ via profiles | ⚠️ varies |
102
+ | Backend logic / DB / payments | ❌ stubbed | ❌ | ❌ | ❌ | ❌ | ❌ |
103
+ | Setup weight | pip + one browser | preinstalled | extension / npm | npm + browser | docker / k8s | signup + $$$ |
104
+
105
+ One independent roundup concluded that *a working offline website*
106
+ is delivered by "none of the above." That is the gap Everpage was
107
+ built to close — and unlike every other row in that roundup, it ships
108
+ measured proof (see results above) instead of promises.
109
+
110
+ ## How it works
111
+
112
+ | Stage | What it does |
113
+ |---|---|
114
+ | `fetch` | Homepage + manifest with real-browser headers, gzip handled |
115
+ | `detect` | Framework fingerprint (Next / Vite / Astro / WordPress / Three.js…) |
116
+ | `crawl` | Sitemap + same-origin links as first-class pages, recursion included |
117
+ | `trace` | One headless pass records URLs the app resolves at runtime |
118
+ | `assets` / `chunks` | Bundle mining: hashed chunks, loader roots, backtick templates (incl. multi-var `terrain/index` shapes), ID pools, ternary alternatives, LOD-gap interpolation, absolute same-origin worker URLs |
119
+ | `rewrite` | Root-absolute → depth-relative paths; absolute same-origin URLs folded to local in HTML/CSS/JS/JSON; query-page URLs (`?p=`, `?page=`) self-name instead of clobbering `index.html` |
120
+ | `serve` | Offline preview server with API/beacon stubs |
121
+ | `verify` | Headless reload: JS errors, failed requests, screenshots |
122
+
123
+ `python smoke_test.py` runs the offline self-checks (no network).
124
+
125
+ ## Honest scope
126
+
127
+ - **Backends aren't cloned** — APIs are stubbed with captured payloads.
128
+ - **Gated content isn't reachable** — logins, paywalls, DRM and aggressive anti-bot need a real session.
129
+ - **Live state isn't frozen** — websockets, personalization and per-user feeds snapshot to whatever loaded.
130
+ - **Interaction-gated assets can be missed** — the trace pass scrolls and idles; scripted clicking (à la websnap) is on the roadmap.
131
+
132
+ ## Layout
133
+
134
+ ```
135
+ everpage/
136
+ cli.py pipeline orchestration
137
+ fetcher.py sessions, retries, URL→file mapping, parallel downloads
138
+ detect.py framework + HTML asset/link extraction
139
+ chunks.py JS bundle mining (templates, pools, workers, frames)
140
+ rewrite.py offline path rewriting (+ regression self_test)
141
+ trace.py headless runtime URL capture
142
+ serve.py preview server with API stubs
143
+ verify.py headless boot/error/screenshot check
144
+ tui.py interactive menu
145
+ ```
146
+
147
+ ## Contributing
148
+
149
+ Issues with a failing URL + the probe log tail are gold. PRs that add a
150
+ `self_test` regression case alongside any parser change get merged fastest.
151
+
152
+ ## License
153
+
154
+ MIT — see [LICENSE](LICENSE).
File without changes