gnosis-markdown 1.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- gnosis_markdown-1.1.0/CHANGELOG.md +58 -0
- gnosis_markdown-1.1.0/LICENSE +21 -0
- gnosis_markdown-1.1.0/MANIFEST.in +9 -0
- gnosis_markdown-1.1.0/PKG-INFO +379 -0
- gnosis_markdown-1.1.0/README.md +339 -0
- gnosis_markdown-1.1.0/config/default.yaml +207 -0
- gnosis_markdown-1.1.0/gnosis/__init__.py +8 -0
- gnosis_markdown-1.1.0/gnosis/__main__.py +11 -0
- gnosis_markdown-1.1.0/gnosis/cli/__init__.py +5 -0
- gnosis_markdown-1.1.0/gnosis/cli/main.py +665 -0
- gnosis_markdown-1.1.0/gnosis/config/__init__.py +5 -0
- gnosis_markdown-1.1.0/gnosis/config/settings.py +353 -0
- gnosis_markdown-1.1.0/gnosis/core/__init__.py +7 -0
- gnosis_markdown-1.1.0/gnosis/core/converter.py +698 -0
- gnosis_markdown-1.1.0/gnosis/core/crawler.py +234 -0
- gnosis_markdown-1.1.0/gnosis/core/downloader.py +176 -0
- gnosis_markdown-1.1.0/gnosis/core/provenance.py +112 -0
- gnosis_markdown-1.1.0/gnosis/integrations/__init__.py +12 -0
- gnosis_markdown-1.1.0/gnosis/integrations/llm.py +255 -0
- gnosis_markdown-1.1.0/gnosis/integrations/qmd.py +264 -0
- gnosis_markdown-1.1.0/gnosis_markdown.egg-info/PKG-INFO +379 -0
- gnosis_markdown-1.1.0/gnosis_markdown.egg-info/SOURCES.txt +35 -0
- gnosis_markdown-1.1.0/gnosis_markdown.egg-info/dependency_links.txt +1 -0
- gnosis_markdown-1.1.0/gnosis_markdown.egg-info/entry_points.txt +2 -0
- gnosis_markdown-1.1.0/gnosis_markdown.egg-info/not-zip-safe +1 -0
- gnosis_markdown-1.1.0/gnosis_markdown.egg-info/requires.txt +10 -0
- gnosis_markdown-1.1.0/gnosis_markdown.egg-info/top_level.txt +2 -0
- gnosis_markdown-1.1.0/pyproject.toml +56 -0
- gnosis_markdown-1.1.0/requirements-qmd.txt +3 -0
- gnosis_markdown-1.1.0/requirements.txt +10 -0
- gnosis_markdown-1.1.0/setup.cfg +4 -0
- gnosis_markdown-1.1.0/setup.py +69 -0
- gnosis_markdown-1.1.0/tests/test_auth.py +119 -0
- gnosis_markdown-1.1.0/tests/test_cli.py +247 -0
- gnosis_markdown-1.1.0/tests/test_converter.py +222 -0
- gnosis_markdown-1.1.0/tests/test_crawler.py +84 -0
- gnosis_markdown-1.1.0/tests/test_provenance.py +79 -0
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## [1.1.0] - 2026-08-04
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
- **Provenance frontmatter on every output file** (on by default, `--no-frontmatter`
|
|
7
|
+
to opt out): `title`, `url`, `fetched_at` (UTC ISO 8601), `content_hash`
|
|
8
|
+
(SHA-256 of body), `status_code`, `language`, `author`, `description`,
|
|
9
|
+
`site_name`, `published_time`/`modified_time`, `etag`/`last_modified`,
|
|
10
|
+
`generator`. Standard YAML, parseable by python-frontmatter/Jekyll/Hugo.
|
|
11
|
+
- **Authentication**: Bearer, HTTP Basic (Confluence Cloud PAT pattern:
|
|
12
|
+
email + API token), and arbitrary header auth. Secrets are read from
|
|
13
|
+
environment variables only — via `--bearer-token-env`,
|
|
14
|
+
`--basic-user`/`--basic-token-env`, or `${ENV_VAR}` expansion in config files
|
|
15
|
+
and `--header` values.
|
|
16
|
+
- **User frontmatter extras**: `--frontmatter 'key: value'` (repeatable) and
|
|
17
|
+
`output.frontmatter_extra` in config; merged without overriding core fields.
|
|
18
|
+
- **Crawl manifest**: `--all` runs write `_manifest.json` with per-page URL,
|
|
19
|
+
file, content hash, timestamp, status, and title.
|
|
20
|
+
- **Metadata extraction**: `HTMLToMarkdownConverter.extract_metadata()`
|
|
21
|
+
(title with entity unescaping, author, language, OG fields).
|
|
22
|
+
- **Downloader `fetch_result()`**: returns `FetchResult` with final URL,
|
|
23
|
+
status code, fetch timestamp, and response headers. The crawler now carries
|
|
24
|
+
this provenance through crawl mode.
|
|
25
|
+
- **Boilerplate word stripping**: `converter.strip_class_words` — word-level
|
|
26
|
+
class matching catches framework-namespaced boilerplate
|
|
27
|
+
(`bd-sidebar-primary`) without false positives (`research-content`).
|
|
28
|
+
- **Content selector precedence**: `content_selectors` are now tried in order
|
|
29
|
+
and the first substantial match wins (matching the documented behavior);
|
|
30
|
+
platform containers (`.markdown-body`, `.ak-renderer-document`,
|
|
31
|
+
`.wiki-content`) precede chrome-wrapping landmarks (`main`, `#content`).
|
|
32
|
+
- **Test suite**: 47 pytest tests covering converter quality, provenance,
|
|
33
|
+
auth, crawler resolution, and CLI end-to-end behavior.
|
|
34
|
+
|
|
35
|
+
### Fixed
|
|
36
|
+
- HTML comments no longer leak into output as text (Confluence
|
|
37
|
+
`<!-- data-loadable-* -->` SSR markers polluted converted pages).
|
|
38
|
+
- Heading permalink anchors no longer glue `#` onto heading text
|
|
39
|
+
(`# Quickstart#` → `# Quickstart`).
|
|
40
|
+
- Table cells with multiple paragraphs no longer break markdown rows
|
|
41
|
+
(joined with `<br>`); pipes in cell text are escaped.
|
|
42
|
+
- Duplicate table header rows (Confluence thead+tbody repetition) and
|
|
43
|
+
single-row sticky-header clone tables are removed.
|
|
44
|
+
- Relative links below extensionless "directory" URLs now resolve correctly
|
|
45
|
+
(`/en/latest` + `quickstart.html` no longer escapes the crawl scope) —
|
|
46
|
+
this previously broke `--all` crawls of most docs sites.
|
|
47
|
+
- `data:`-URI images (spacers/tracking pixels) are skipped.
|
|
48
|
+
|
|
49
|
+
### Changed
|
|
50
|
+
- torch/transformers are no longer core dependencies; install the QMD
|
|
51
|
+
integration via `pip install gnosis[qmd]` (or `requirements-qmd.txt`).
|
|
52
|
+
- Blank-line collapsing in output is stricter (max one blank line).
|
|
53
|
+
- Default User-Agent updated to `Gnosis/1.1`.
|
|
54
|
+
|
|
55
|
+
## [1.0.0] - 2026-03-27
|
|
56
|
+
|
|
57
|
+
Initial public release: single-page download, `--all` crawling, configurable
|
|
58
|
+
extraction, QMD integration.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Steffen Hoehne, SHCV.IT
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,379 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: gnosis-markdown
|
|
3
|
+
Version: 1.1.0
|
|
4
|
+
Summary: Website to Markdown converter with provenance frontmatter for LLM knowledge bases
|
|
5
|
+
Home-page: https://github.com/shcv-it/gnosis
|
|
6
|
+
Author: Steffen Hoehne
|
|
7
|
+
Author-email: Steffen Hoehne <steffen.hoehne@shcv.it>
|
|
8
|
+
License-Expression: MIT
|
|
9
|
+
Project-URL: Homepage, https://github.com/SHCV-it/gnosis
|
|
10
|
+
Project-URL: Repository, https://github.com/SHCV-it/gnosis
|
|
11
|
+
Project-URL: Issues, https://github.com/SHCV-it/gnosis/issues
|
|
12
|
+
Project-URL: Changelog, https://github.com/SHCV-it/gnosis/blob/main/CHANGELOG.md
|
|
13
|
+
Keywords: markdown,web-scraping,crawler,html-to-markdown,knowledge-base,llm
|
|
14
|
+
Classifier: Development Status :: 5 - Production/Stable
|
|
15
|
+
Classifier: Intended Audience :: Developers
|
|
16
|
+
Classifier: Intended Audience :: Information Technology
|
|
17
|
+
Classifier: Operating System :: OS Independent
|
|
18
|
+
Classifier: Programming Language :: Python :: 3
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
21
|
+
Classifier: Topic :: Text Processing :: Markup
|
|
22
|
+
Classifier: Topic :: Internet :: WWW/HTTP
|
|
23
|
+
Classifier: Topic :: Documentation
|
|
24
|
+
Requires-Python: >=3.12
|
|
25
|
+
Description-Content-Type: text/markdown
|
|
26
|
+
License-File: LICENSE
|
|
27
|
+
Requires-Dist: click>=8.0.0
|
|
28
|
+
Requires-Dist: rich>=13.0.0
|
|
29
|
+
Requires-Dist: httpx>=0.25.0
|
|
30
|
+
Requires-Dist: beautifulsoup4>=4.12.0
|
|
31
|
+
Requires-Dist: lxml>=5.0.0
|
|
32
|
+
Requires-Dist: pyyaml>=6.0.0
|
|
33
|
+
Provides-Extra: qmd
|
|
34
|
+
Requires-Dist: torch>=2.0.0; extra == "qmd"
|
|
35
|
+
Requires-Dist: transformers>=4.30.0; extra == "qmd"
|
|
36
|
+
Dynamic: author
|
|
37
|
+
Dynamic: home-page
|
|
38
|
+
Dynamic: license-file
|
|
39
|
+
Dynamic: requires-python
|
|
40
|
+
|
|
41
|
+
# Gnosis
|
|
42
|
+
|
|
43
|
+
<p align="center">
|
|
44
|
+
<strong>Website → clean, provenance-stamped Markdown.</strong><br>
|
|
45
|
+
<em>Built for LLM knowledge bases, documentation pipelines, and audit-ready content archives.</em>
|
|
46
|
+
</p>
|
|
47
|
+
|
|
48
|
+
<p align="center">
|
|
49
|
+
<a href="https://pypi.org/project/gnosis/"><img alt="PyPI" src="https://img.shields.io/badge/python-3.12+-blue"></a>
|
|
50
|
+
<a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/license-MIT-green"></a>
|
|
51
|
+
<a href="https://github.com/SHCV-it/gnosis"><img alt="Version" src="https://img.shields.io/badge/version-1.1.0-blue"></a>
|
|
52
|
+
</p>
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
Point gnosis at a URL and get clean, LLM-friendly Markdown files with a **YAML
|
|
57
|
+
provenance frontmatter block** on every file — recording exactly where the
|
|
58
|
+
content came from, when it was fetched, and how to verify it. No external
|
|
59
|
+
bookkeeping, no hidden state. Every file is self-describing.
|
|
60
|
+
|
|
61
|
+
## Table of contents
|
|
62
|
+
|
|
63
|
+
- [Features](#features)
|
|
64
|
+
- [Installation](#installation)
|
|
65
|
+
- [Quick start](#quick-start)
|
|
66
|
+
- [Provenance: the contract](#provenance-the-contract)
|
|
67
|
+
- [Authenticated fetching](#authenticated-fetching)
|
|
68
|
+
- [CLI reference](#cli-reference)
|
|
69
|
+
- [Configuration reference](#configuration-reference)
|
|
70
|
+
- [Exit codes](#exit-codes)
|
|
71
|
+
- [How clean is the output?](#how-clean-is-the-output)
|
|
72
|
+
- [Development & testing](#development--testing)
|
|
73
|
+
- [License](#license)
|
|
74
|
+
|
|
75
|
+
## Features
|
|
76
|
+
|
|
77
|
+
| Feature | Description |
|
|
78
|
+
|---|---|
|
|
79
|
+
| **Single page or full site** | One URL, or crawl every page under a path with `--all` |
|
|
80
|
+
| **Provenance frontmatter** | Every file records `url`, `fetched_at` (UTC), SHA-256 `content_hash`, `status_code`, page metadata (title/author/language), and caching headers when the server sends them |
|
|
81
|
+
| **Authentication built-in** | Bearer tokens, HTTP Basic (Confluence Cloud API tokens), or arbitrary headers — secrets **always** from environment variables |
|
|
82
|
+
| **Clean extraction** | Main-content detection with ordered selectors, class-word boilerplate stripping (sidebars, breadcrumbs, cookie banners, permalink anchors), and framework-aware fixes for Confluence, Sphinx/RTD, and GitHub-style pages |
|
|
83
|
+
| **Valid GFM tables** | Multi-paragraph cells joined with `<br>`, pipes escaped, sticky-header clone tables removed — tables survive ingestion |
|
|
84
|
+
| **Scalable crawling** | Configurable concurrent fetches, politeness delay, retries with backoff, robots-aware |
|
|
85
|
+
| **Scheduler-friendly** | Headless CLI, meaningful exit codes, JSON crawl manifest — drops into cron, n8n, Airflow, CI |
|
|
86
|
+
| **Metadata extraction** | Title (entity-unescaped), author, language, description, Open Graph fields |
|
|
87
|
+
| **Optional QMD integration** | Index output into a QMD knowledge base with local-LLM context generation (`pip install gnosis[qmd]`) |
|
|
88
|
+
|
|
89
|
+
## Installation
|
|
90
|
+
|
|
91
|
+
Requires **Python 3.12+**.
|
|
92
|
+
|
|
93
|
+
```bash
|
|
94
|
+
git clone https://github.com/SHCV-it/gnosis.git
|
|
95
|
+
cd gnosis
|
|
96
|
+
python -m venv venv && source venv/bin/activate
|
|
97
|
+
pip install -e .
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
Optional, only if you use `--qmd-index` (pulls torch + transformers):
|
|
101
|
+
|
|
102
|
+
```bash
|
|
103
|
+
pip install -e .[qmd]
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
## Quick start
|
|
107
|
+
|
|
108
|
+
```bash
|
|
109
|
+
# One page → one markdown file with provenance
|
|
110
|
+
gnosis https://docs.python.org/3/tutorial/
|
|
111
|
+
|
|
112
|
+
# Crawl an entire section of a docs site
|
|
113
|
+
gnosis https://docs.python.org/3/tutorial/ --all -o ./python-docs/
|
|
114
|
+
|
|
115
|
+
# Preview how many pages would be crawled (no downloads)
|
|
116
|
+
gnosis https://docs.python.org/3/tutorial/ --all --dry-run
|
|
117
|
+
|
|
118
|
+
# Faster crawl with parallel fetches
|
|
119
|
+
gnosis https://docs.example.com/ --all --config myconfig.yaml
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
## Provenance: the contract
|
|
123
|
+
|
|
124
|
+
Every file gnosis writes is self-describing. Default frontmatter:
|
|
125
|
+
|
|
126
|
+
```yaml
|
|
127
|
+
---
|
|
128
|
+
title: Quickstart — Trafilatura 2.2.0 documentation
|
|
129
|
+
url: https://trafilatura.readthedocs.io/en/latest/quickstart.html
|
|
130
|
+
fetched_at: '2026-08-04T10:24:17Z'
|
|
131
|
+
content_hash: 1549512c...16fd
|
|
132
|
+
status_code: 200
|
|
133
|
+
generator: gnosis/1.1.0
|
|
134
|
+
language: en
|
|
135
|
+
etag: '"61e917f4cd107c3bce6182b633819fcf"'
|
|
136
|
+
last_modified: Fri, 31 Jul 2026 16:07:37 GMT
|
|
137
|
+
---
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
| Field | Required | Description |
|
|
141
|
+
|---|---|---|
|
|
142
|
+
| `title` | Always | Page title (from `og:title` or `<title>`, HTML entities unescaped) |
|
|
143
|
+
| `url` | Always | Final URL after redirects (`requested_url` added if it differs) |
|
|
144
|
+
| `fetched_at` | Always | UTC fetch timestamp, ISO 8601 |
|
|
145
|
+
| `content_hash` | Always | SHA-256 of the markdown body — use for dedup/change detection |
|
|
146
|
+
| `status_code` | Always | HTTP status of the final response |
|
|
147
|
+
| `generator` | Always | Gnosis version that produced this file |
|
|
148
|
+
| `language` | If present | From `<html lang>` or `og:locale` |
|
|
149
|
+
| `author` | If present | From `<meta name=author>`, `article:author`, or `dc.creator` |
|
|
150
|
+
| `description` | If present | From `<meta name=description>` or `og:description` |
|
|
151
|
+
| `site_name` | If present | From `og:site_name` |
|
|
152
|
+
| `published_time` | If present | From `article:published_time` |
|
|
153
|
+
| `modified_time` | If present | From `article:modified_time` |
|
|
154
|
+
| `etag` | If sent | Response `ETag` header |
|
|
155
|
+
| `last_modified` | If sent | Response `Last-Modified` header |
|
|
156
|
+
| `requested_url` | If redirected | Original URL before redirects |
|
|
157
|
+
|
|
158
|
+
Add your own constant fields per run (`--frontmatter`) or per config
|
|
159
|
+
(`output.frontmatter_extra`) — custom keys never override the core provenance
|
|
160
|
+
fields above:
|
|
161
|
+
|
|
162
|
+
```bash
|
|
163
|
+
gnosis https://example.com/docs --frontmatter 'tags: [customs, passar]' --frontmatter 'owner: kb-team'
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
The frontmatter is standard YAML between `---` fences: parseable by
|
|
167
|
+
python-frontmatter, Jekyll, Hugo, Obsidian, and any downstream knowledge
|
|
168
|
+
pipeline.
|
|
169
|
+
|
|
170
|
+
Opt out per run with `--no-frontmatter` or globally in config:
|
|
171
|
+
```yaml
|
|
172
|
+
output:
|
|
173
|
+
frontmatter: false
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
## Authenticated fetching
|
|
177
|
+
|
|
178
|
+
Secrets are read from **environment variables only**. They are never passed as
|
|
179
|
+
plain CLI arguments (which leak into shell history and process tables) and
|
|
180
|
+
never committed in config files.
|
|
181
|
+
|
|
182
|
+
### Confluence Cloud with a Personal Access Token
|
|
183
|
+
|
|
184
|
+
```bash
|
|
185
|
+
# Set up a PAT at https://id.atlassian.com/manage/api-tokens
|
|
186
|
+
export CONFLUENCE_PAT="your-api-token"
|
|
187
|
+
|
|
188
|
+
gnosis "https://your-domain.atlassian.net/wiki/spaces/SPACE/pages/PAGE_ID" \
|
|
189
|
+
--basic-user you@example.com \
|
|
190
|
+
--basic-token-env CONFLUENCE_PAT
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
### Bearer token (authenticated API docs, internal tools)
|
|
194
|
+
|
|
195
|
+
```bash
|
|
196
|
+
export MY_API_TOKEN="..."
|
|
197
|
+
gnosis https://internal.example.com/docs --bearer-token-env MY_API_TOKEN
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
### Custom headers
|
|
201
|
+
|
|
202
|
+
```bash
|
|
203
|
+
gnosis https://example.com --header "X-API-Key: ${MY_KEY}" --header "X-Team: docs"
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
### Via config file (multi-run / CI)
|
|
207
|
+
|
|
208
|
+
```yaml
|
|
209
|
+
downloader:
|
|
210
|
+
auth:
|
|
211
|
+
type: basic # bearer | basic | header
|
|
212
|
+
username: "you@example.com"
|
|
213
|
+
password: "${CONFLUENCE_PAT}" # ${ENV_VAR} expanded at load time
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
## CLI reference
|
|
217
|
+
|
|
218
|
+
```
|
|
219
|
+
gnosis URL [OPTIONS]
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
| Flag | Description |
|
|
223
|
+
|---|---|
|
|
224
|
+
| `-a, --all` | Crawl all child pages under the URL path |
|
|
225
|
+
| `-n, --dry-run` | Discover and count pages only (requires `--all`) |
|
|
226
|
+
| `-o, --output DIR` | Output directory (default: `./`) |
|
|
227
|
+
| `-c, --config FILE` | Path to YAML configuration file |
|
|
228
|
+
| `-f, --overwrite` | Overwrite existing output files |
|
|
229
|
+
| `-q, --quiet` | Suppress progress output |
|
|
230
|
+
| `-v, --verbose` | Show detailed conversion diagnostics |
|
|
231
|
+
| `--no-frontmatter` | Write bare markdown without provenance block |
|
|
232
|
+
| `--frontmatter KEY: VALUE` | Extra constant frontmatter field (repeatable) |
|
|
233
|
+
| `--header NAME: VALUE` | Extra request header, `${ENV_VAR}` expanded (repeatable) |
|
|
234
|
+
| `--bearer-token-env VAR` | Bearer token from environment variable |
|
|
235
|
+
| `--basic-user USER` | HTTP Basic username (requires `--basic-token-env`) |
|
|
236
|
+
| `--basic-token-env VAR` | HTTP Basic password/token from environment variable |
|
|
237
|
+
| `--qmd-index` | Index output into QMD (requires `[qmd]` extra) |
|
|
238
|
+
|
|
239
|
+
## Configuration reference
|
|
240
|
+
|
|
241
|
+
Copy [`config/default.yaml`](config/default.yaml) and pass it with `-c`.
|
|
242
|
+
Full reference:
|
|
243
|
+
|
|
244
|
+
```yaml
|
|
245
|
+
# ── Downloader ────────────────────────────
|
|
246
|
+
downloader:
|
|
247
|
+
timeout: 30 # Request timeout (seconds)
|
|
248
|
+
retries: 3 # Retries on 5xx / network errors
|
|
249
|
+
user_agent: "Gnosis/1.1" # User-Agent header
|
|
250
|
+
rate_limit_ms: 500 # Minimum delay between requests (0 = no limit)
|
|
251
|
+
respect_robots: true # Obey robots.txt (future)
|
|
252
|
+
headers: {} # Extra HTTP headers (${ENV_VAR} expanded)
|
|
253
|
+
auth: # Optional: bearer | basic | header
|
|
254
|
+
type: bearer
|
|
255
|
+
token: "${MY_API_TOKEN}"
|
|
256
|
+
|
|
257
|
+
# ── Crawler ───────────────────────────────
|
|
258
|
+
crawler:
|
|
259
|
+
max_depth: 10 # Crawl depth from seed URL
|
|
260
|
+
max_pages: 500 # Stop after this many pages
|
|
261
|
+
concurrent_requests: 5 # Parallel fetch batch size (1 = sequential)
|
|
262
|
+
|
|
263
|
+
# ── Converter ─────────────────────────────
|
|
264
|
+
converter:
|
|
265
|
+
excluded_tags: [...] # HTML tags stripped before conversion
|
|
266
|
+
content_selectors: [...] # Tried in order; first match ≥ 200 chars wins
|
|
267
|
+
strip_classes: [...] # Exact class-token matches to remove
|
|
268
|
+
strip_class_words: [...] # Word-level matches inside class names
|
|
269
|
+
include_images: true # Emit <img> as 
|
|
270
|
+
absolute_urls: true # Resolve relative links to absolute URLs
|
|
271
|
+
|
|
272
|
+
# ── Output ─────────────────────────────────
|
|
273
|
+
output:
|
|
274
|
+
directory: "./" # Where .md files go
|
|
275
|
+
overwrite: false # Skip existing files unless true
|
|
276
|
+
extension: ".md" # Output file extension
|
|
277
|
+
frontmatter: true # Write YAML provenance block
|
|
278
|
+
frontmatter_extra: {} # Constant fields added to every file
|
|
279
|
+
|
|
280
|
+
# ── QMD (optional) ──────────────────────────
|
|
281
|
+
qmd:
|
|
282
|
+
enabled: false # Enable QMD knowledge base indexing
|
|
283
|
+
llm_model: "Qwen/Qwen3-0.6B" # HuggingFace model for context generation
|
|
284
|
+
llm_device: "cpu" # cpu | cuda | auto
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
## Exit codes
|
|
288
|
+
|
|
289
|
+
| Code | Meaning |
|
|
290
|
+
|---|---|
|
|
291
|
+
| `0` | Success (crawl mode: at least one page saved) |
|
|
292
|
+
| `1` | Failure — download error, file exists without `-f`, nothing saved, bad flags |
|
|
293
|
+
| `130` | Interrupted (Ctrl-C) |
|
|
294
|
+
|
|
295
|
+
In `--all` mode a `_manifest.json` is written to the output directory listing
|
|
296
|
+
every page with `url`, `file`, `content_hash`, `fetched_at`, `status_code`, and
|
|
297
|
+
`title` — ready for schedulers and downstream audit.
|
|
298
|
+
|
|
299
|
+
## How clean is the output?
|
|
300
|
+
|
|
301
|
+
Gnosis is opinionated about boilerplate. By default it:
|
|
302
|
+
|
|
303
|
+
- **strips** script/style/nav/footer/aside/form/template tags and HTML comments
|
|
304
|
+
(including Confluence's `<!-- data-loadable-begin=... -->` SSR markers)
|
|
305
|
+
- **removes** elements matching exact `strip_classes` tokens AND boilerplate
|
|
306
|
+
*words* (`sidebar`, `toc`, `breadcrumb`, `cookie`, `headerlink`, `sourcelink`, …)
|
|
307
|
+
inside namespaced class names — `bd-sidebar-primary` is gone, but `research-content`
|
|
308
|
+
stays
|
|
309
|
+
- **cleans** permalink anchors from headings (`# Quickstart#` → `# Quickstart`)
|
|
310
|
+
- **picks** the main content by precedence-ordered selectors
|
|
311
|
+
(`.markdown-body`, `.ak-renderer-document`, `.wiki-content`, … before `main`/`#content`)
|
|
312
|
+
- **converts** tables to valid GFM: multi-line cells joined with `<br>`, `|`
|
|
313
|
+
escaped, duplicate/sticky-header rows removed
|
|
314
|
+
- **resolves** relative links to absolute URLs; skips `data:`-URI images
|
|
315
|
+
(spacers/tracking pixels)
|
|
316
|
+
- **unescapes** HTML entities in titles (`—` → `—`)
|
|
317
|
+
|
|
318
|
+
Everything is configurable — see [`config/default.yaml`](config/default.yaml).
|
|
319
|
+
|
|
320
|
+
## Development & testing
|
|
321
|
+
|
|
322
|
+
```bash
|
|
323
|
+
git clone https://github.com/SHCV-it/gnosis.git
|
|
324
|
+
cd gnosis
|
|
325
|
+
python -m venv venv && source venv/bin/activate
|
|
326
|
+
pip install -e . pytest python-frontmatter
|
|
327
|
+
|
|
328
|
+
# Run the test suite (offline — only localhost fixtures)
|
|
329
|
+
python -m pytest tests/ -q -v
|
|
330
|
+
```
|
|
331
|
+
|
|
332
|
+
### Project structure
|
|
333
|
+
|
|
334
|
+
```
|
|
335
|
+
gnosis/
|
|
336
|
+
cli/ Click CLI (single page, crawl, dry-run, manifest)
|
|
337
|
+
config/ YAML loading + typed settings dataclass
|
|
338
|
+
core/
|
|
339
|
+
downloader Async HTTP client, auth, retries, FetchResult
|
|
340
|
+
converter HTML → Markdown, boilerplate stripping, metadata extraction
|
|
341
|
+
crawler BFS crawler with concurrent batch fetching
|
|
342
|
+
provenance Frontmatter generation, content_hash, render_document
|
|
343
|
+
integrations/ QMD pipeline (optional, heavy deps)
|
|
344
|
+
```
|
|
345
|
+
|
|
346
|
+
The test suite covers converter quality (comment/anchor/boilerplate/table
|
|
347
|
+
handling, shadow-table dedup, metadata), provenance generation (fields,
|
|
348
|
+
round-trip parsing, extras merging), auth header injection (3 schemes),
|
|
349
|
+
crawler link resolution, and CLI end-to-end behavior. Runs entirely offline.
|
|
350
|
+
|
|
351
|
+
## Contributing
|
|
352
|
+
|
|
353
|
+
Contributions are welcome. Please open an issue first to discuss what you'd
|
|
354
|
+
like to change.
|
|
355
|
+
|
|
356
|
+
1. Fork the repository
|
|
357
|
+
2. Create a feature branch (`git checkout -b feature/amazing`)
|
|
358
|
+
3. Run the tests (`python -m pytest tests/ -q`)
|
|
359
|
+
4. Commit your changes with clear messages
|
|
360
|
+
5. Push and open a pull request against `main`
|
|
361
|
+
|
|
362
|
+
## Related projects
|
|
363
|
+
|
|
364
|
+
Gnosis is designed to feed documentation pipelines and LLM knowledge bases.
|
|
365
|
+
Pair it with:
|
|
366
|
+
|
|
367
|
+
- **n8n / cron / Airflow** — schedule gnosis runs and pipe results into your
|
|
368
|
+
downstream pipeline
|
|
369
|
+
- **QMD** — local vector search via the `--qmd-index` flag
|
|
370
|
+
- **Any Markdown-to-anything pipeline** — the YAML frontmatter is parseable by
|
|
371
|
+
python-frontmatter, Jekyll, Hugo, Obsidian, and standard static-site generators
|
|
372
|
+
|
|
373
|
+
## License
|
|
374
|
+
|
|
375
|
+
MIT — see [LICENSE](LICENSE).
|
|
376
|
+
|
|
377
|
+
---
|
|
378
|
+
|
|
379
|
+
**Author:** Steffen Hoehne, [SHCV.IT](https://shcv.it)
|