gnosis-markdown 1.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. gnosis_markdown-1.1.0/CHANGELOG.md +58 -0
  2. gnosis_markdown-1.1.0/LICENSE +21 -0
  3. gnosis_markdown-1.1.0/MANIFEST.in +9 -0
  4. gnosis_markdown-1.1.0/PKG-INFO +379 -0
  5. gnosis_markdown-1.1.0/README.md +339 -0
  6. gnosis_markdown-1.1.0/config/default.yaml +207 -0
  7. gnosis_markdown-1.1.0/gnosis/__init__.py +8 -0
  8. gnosis_markdown-1.1.0/gnosis/__main__.py +11 -0
  9. gnosis_markdown-1.1.0/gnosis/cli/__init__.py +5 -0
  10. gnosis_markdown-1.1.0/gnosis/cli/main.py +665 -0
  11. gnosis_markdown-1.1.0/gnosis/config/__init__.py +5 -0
  12. gnosis_markdown-1.1.0/gnosis/config/settings.py +353 -0
  13. gnosis_markdown-1.1.0/gnosis/core/__init__.py +7 -0
  14. gnosis_markdown-1.1.0/gnosis/core/converter.py +698 -0
  15. gnosis_markdown-1.1.0/gnosis/core/crawler.py +234 -0
  16. gnosis_markdown-1.1.0/gnosis/core/downloader.py +176 -0
  17. gnosis_markdown-1.1.0/gnosis/core/provenance.py +112 -0
  18. gnosis_markdown-1.1.0/gnosis/integrations/__init__.py +12 -0
  19. gnosis_markdown-1.1.0/gnosis/integrations/llm.py +255 -0
  20. gnosis_markdown-1.1.0/gnosis/integrations/qmd.py +264 -0
  21. gnosis_markdown-1.1.0/gnosis_markdown.egg-info/PKG-INFO +379 -0
  22. gnosis_markdown-1.1.0/gnosis_markdown.egg-info/SOURCES.txt +35 -0
  23. gnosis_markdown-1.1.0/gnosis_markdown.egg-info/dependency_links.txt +1 -0
  24. gnosis_markdown-1.1.0/gnosis_markdown.egg-info/entry_points.txt +2 -0
  25. gnosis_markdown-1.1.0/gnosis_markdown.egg-info/not-zip-safe +1 -0
  26. gnosis_markdown-1.1.0/gnosis_markdown.egg-info/requires.txt +10 -0
  27. gnosis_markdown-1.1.0/gnosis_markdown.egg-info/top_level.txt +2 -0
  28. gnosis_markdown-1.1.0/pyproject.toml +56 -0
  29. gnosis_markdown-1.1.0/requirements-qmd.txt +3 -0
  30. gnosis_markdown-1.1.0/requirements.txt +10 -0
  31. gnosis_markdown-1.1.0/setup.cfg +4 -0
  32. gnosis_markdown-1.1.0/setup.py +69 -0
  33. gnosis_markdown-1.1.0/tests/test_auth.py +119 -0
  34. gnosis_markdown-1.1.0/tests/test_cli.py +247 -0
  35. gnosis_markdown-1.1.0/tests/test_converter.py +222 -0
  36. gnosis_markdown-1.1.0/tests/test_crawler.py +84 -0
  37. gnosis_markdown-1.1.0/tests/test_provenance.py +79 -0
@@ -0,0 +1,58 @@
1
+ # Changelog
2
+
3
+ ## [1.1.0] - 2026-08-04
4
+
5
+ ### Added
6
+ - **Provenance frontmatter on every output file** (on by default, `--no-frontmatter`
7
+ to opt out): `title`, `url`, `fetched_at` (UTC ISO 8601), `content_hash`
8
+ (SHA-256 of body), `status_code`, `language`, `author`, `description`,
9
+ `site_name`, `published_time`/`modified_time`, `etag`/`last_modified`,
10
+ `generator`. Standard YAML, parseable by python-frontmatter/Jekyll/Hugo.
11
+ - **Authentication**: Bearer, HTTP Basic (Confluence Cloud PAT pattern:
12
+ email + API token), and arbitrary header auth. Secrets are read from
13
+ environment variables only — via `--bearer-token-env`,
14
+ `--basic-user`/`--basic-token-env`, or `${ENV_VAR}` expansion in config files
15
+ and `--header` values.
16
+ - **User frontmatter extras**: `--frontmatter 'key: value'` (repeatable) and
17
+ `output.frontmatter_extra` in config; merged without overriding core fields.
18
+ - **Crawl manifest**: `--all` runs write `_manifest.json` with per-page URL,
19
+ file, content hash, timestamp, status, and title.
20
+ - **Metadata extraction**: `HTMLToMarkdownConverter.extract_metadata()`
21
+ (title with entity unescaping, author, language, OG fields).
22
+ - **Downloader `fetch_result()`**: returns `FetchResult` with final URL,
23
+ status code, fetch timestamp, and response headers. The crawler now carries
24
+ this provenance through crawl mode.
25
+ - **Boilerplate word stripping**: `converter.strip_class_words` — word-level
26
+ class matching catches framework-namespaced boilerplate
27
+ (`bd-sidebar-primary`) without false positives (`research-content`).
28
+ - **Content selector precedence**: `content_selectors` are now tried in order
29
+ and the first substantial match wins (matching the documented behavior);
30
+ platform containers (`.markdown-body`, `.ak-renderer-document`,
31
+ `.wiki-content`) precede chrome-wrapping landmarks (`main`, `#content`).
32
+ - **Test suite**: 47 pytest tests covering converter quality, provenance,
33
+ auth, crawler resolution, and CLI end-to-end behavior.
34
+
35
+ ### Fixed
36
+ - HTML comments no longer leak into output as text (Confluence
37
+ `<!-- data-loadable-* -->` SSR markers polluted converted pages).
38
+ - Heading permalink anchors no longer glue `#` onto heading text
39
+ (`# Quickstart#` → `# Quickstart`).
40
+ - Table cells with multiple paragraphs no longer break markdown rows
41
+ (joined with `<br>`); pipes in cell text are escaped.
42
+ - Duplicate table header rows (Confluence thead+tbody repetition) and
43
+ single-row sticky-header clone tables are removed.
44
+ - Relative links below extensionless "directory" URLs now resolve correctly
45
+ (`/en/latest` + `quickstart.html` no longer escapes the crawl scope) —
46
+ this previously broke `--all` crawls of most docs sites.
47
+ - `data:`-URI images (spacers/tracking pixels) are skipped.
48
+
49
+ ### Changed
50
+ - torch/transformers are no longer core dependencies; install the QMD
51
+ integration via `pip install gnosis[qmd]` (or `requirements-qmd.txt`).
52
+ - Blank-line collapsing in output is stricter (max one blank line).
53
+ - Default User-Agent updated to `Gnosis/1.1`.
54
+
55
+ ## [1.0.0] - 2026-03-27
56
+
57
+ Initial public release: single-page download, `--all` crawling, configurable
58
+ extraction, QMD integration.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Steffen Hoehne, SHCV.IT
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,9 @@
1
+ include LICENSE
2
+ include README.md
3
+ include CHANGELOG.md
4
+ include requirements.txt
5
+ include requirements-qmd.txt
6
+ include config/default.yaml
7
+ recursive-include gnosis *.py
8
+ recursive-include tests *.py
9
+ graft config
@@ -0,0 +1,379 @@
1
+ Metadata-Version: 2.4
2
+ Name: gnosis-markdown
3
+ Version: 1.1.0
4
+ Summary: Website to Markdown converter with provenance frontmatter for LLM knowledge bases
5
+ Home-page: https://github.com/shcv-it/gnosis
6
+ Author: Steffen Hoehne
7
+ Author-email: Steffen Hoehne <steffen.hoehne@shcv.it>
8
+ License-Expression: MIT
9
+ Project-URL: Homepage, https://github.com/SHCV-it/gnosis
10
+ Project-URL: Repository, https://github.com/SHCV-it/gnosis
11
+ Project-URL: Issues, https://github.com/SHCV-it/gnosis/issues
12
+ Project-URL: Changelog, https://github.com/SHCV-it/gnosis/blob/main/CHANGELOG.md
13
+ Keywords: markdown,web-scraping,crawler,html-to-markdown,knowledge-base,llm
14
+ Classifier: Development Status :: 5 - Production/Stable
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Intended Audience :: Information Technology
17
+ Classifier: Operating System :: OS Independent
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Topic :: Text Processing :: Markup
22
+ Classifier: Topic :: Internet :: WWW/HTTP
23
+ Classifier: Topic :: Documentation
24
+ Requires-Python: >=3.12
25
+ Description-Content-Type: text/markdown
26
+ License-File: LICENSE
27
+ Requires-Dist: click>=8.0.0
28
+ Requires-Dist: rich>=13.0.0
29
+ Requires-Dist: httpx>=0.25.0
30
+ Requires-Dist: beautifulsoup4>=4.12.0
31
+ Requires-Dist: lxml>=5.0.0
32
+ Requires-Dist: pyyaml>=6.0.0
33
+ Provides-Extra: qmd
34
+ Requires-Dist: torch>=2.0.0; extra == "qmd"
35
+ Requires-Dist: transformers>=4.30.0; extra == "qmd"
36
+ Dynamic: author
37
+ Dynamic: home-page
38
+ Dynamic: license-file
39
+ Dynamic: requires-python
40
+
41
+ # Gnosis
42
+
43
+ <p align="center">
44
+ <strong>Website → clean, provenance-stamped Markdown.</strong><br>
45
+ <em>Built for LLM knowledge bases, documentation pipelines, and audit-ready content archives.</em>
46
+ </p>
47
+
48
+ <p align="center">
49
+ <a href="https://pypi.org/project/gnosis/"><img alt="PyPI" src="https://img.shields.io/badge/python-3.12+-blue"></a>
50
+ <a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/license-MIT-green"></a>
51
+ <a href="https://github.com/SHCV-it/gnosis"><img alt="Version" src="https://img.shields.io/badge/version-1.1.0-blue"></a>
52
+ </p>
53
+
54
+ ---
55
+
56
+ Point gnosis at a URL and get clean, LLM-friendly Markdown files with a **YAML
57
+ provenance frontmatter block** on every file — recording exactly where the
58
+ content came from, when it was fetched, and how to verify it. No external
59
+ bookkeeping, no hidden state. Every file is self-describing.
60
+
61
+ ## Table of contents
62
+
63
+ - [Features](#features)
64
+ - [Installation](#installation)
65
+ - [Quick start](#quick-start)
66
+ - [Provenance: the contract](#provenance-the-contract)
67
+ - [Authenticated fetching](#authenticated-fetching)
68
+ - [CLI reference](#cli-reference)
69
+ - [Configuration reference](#configuration-reference)
70
+ - [Exit codes](#exit-codes)
71
+ - [How clean is the output?](#how-clean-is-the-output)
72
+ - [Development & testing](#development--testing)
73
+ - [License](#license)
74
+
75
+ ## Features
76
+
77
+ | Feature | Description |
78
+ |---|---|
79
+ | **Single page or full site** | One URL, or crawl every page under a path with `--all` |
80
+ | **Provenance frontmatter** | Every file records `url`, `fetched_at` (UTC), SHA-256 `content_hash`, `status_code`, page metadata (title/author/language), and caching headers when the server sends them |
81
+ | **Authentication built-in** | Bearer tokens, HTTP Basic (Confluence Cloud API tokens), or arbitrary headers — secrets **always** from environment variables |
82
+ | **Clean extraction** | Main-content detection with ordered selectors, class-word boilerplate stripping (sidebars, breadcrumbs, cookie banners, permalink anchors), and framework-aware fixes for Confluence, Sphinx/RTD, and GitHub-style pages |
83
+ | **Valid GFM tables** | Multi-paragraph cells joined with `<br>`, pipes escaped, sticky-header clone tables removed — tables survive ingestion |
84
+ | **Scalable crawling** | Configurable concurrent fetches, politeness delay, retries with backoff, robots-aware |
85
+ | **Scheduler-friendly** | Headless CLI, meaningful exit codes, JSON crawl manifest — drops into cron, n8n, Airflow, CI |
86
+ | **Metadata extraction** | Title (entity-unescaped), author, language, description, Open Graph fields |
87
+ | **Optional QMD integration** | Index output into a QMD knowledge base with local-LLM context generation (`pip install gnosis[qmd]`) |
88
+
89
+ ## Installation
90
+
91
+ Requires **Python 3.12+**.
92
+
93
+ ```bash
94
+ git clone https://github.com/SHCV-it/gnosis.git
95
+ cd gnosis
96
+ python -m venv venv && source venv/bin/activate
97
+ pip install -e .
98
+ ```
99
+
100
+ Optional, only if you use `--qmd-index` (pulls torch + transformers):
101
+
102
+ ```bash
103
+ pip install -e .[qmd]
104
+ ```
105
+
106
+ ## Quick start
107
+
108
+ ```bash
109
+ # One page → one markdown file with provenance
110
+ gnosis https://docs.python.org/3/tutorial/
111
+
112
+ # Crawl an entire section of a docs site
113
+ gnosis https://docs.python.org/3/tutorial/ --all -o ./python-docs/
114
+
115
+ # Preview how many pages would be crawled (no downloads)
116
+ gnosis https://docs.python.org/3/tutorial/ --all --dry-run
117
+
118
+ # Faster crawl with parallel fetches
119
+ gnosis https://docs.example.com/ --all --config myconfig.yaml
120
+ ```
121
+
122
+ ## Provenance: the contract
123
+
124
+ Every file gnosis writes is self-describing. Default frontmatter:
125
+
126
+ ```yaml
127
+ ---
128
+ title: Quickstart — Trafilatura 2.2.0 documentation
129
+ url: https://trafilatura.readthedocs.io/en/latest/quickstart.html
130
+ fetched_at: '2026-08-04T10:24:17Z'
131
+ content_hash: 1549512c...16fd
132
+ status_code: 200
133
+ generator: gnosis/1.1.0
134
+ language: en
135
+ etag: '"61e917f4cd107c3bce6182b633819fcf"'
136
+ last_modified: Fri, 31 Jul 2026 16:07:37 GMT
137
+ ---
138
+ ```
139
+
140
+ | Field | Required | Description |
141
+ |---|---|---|
142
+ | `title` | Always | Page title (from `og:title` or `<title>`, HTML entities unescaped) |
143
+ | `url` | Always | Final URL after redirects (`requested_url` added if it differs) |
144
+ | `fetched_at` | Always | UTC fetch timestamp, ISO 8601 |
145
+ | `content_hash` | Always | SHA-256 of the markdown body — use for dedup/change detection |
146
+ | `status_code` | Always | HTTP status of the final response |
147
+ | `generator` | Always | Gnosis version that produced this file |
148
+ | `language` | If present | From `<html lang>` or `og:locale` |
149
+ | `author` | If present | From `<meta name=author>`, `article:author`, or `dc.creator` |
150
+ | `description` | If present | From `<meta name=description>` or `og:description` |
151
+ | `site_name` | If present | From `og:site_name` |
152
+ | `published_time` | If present | From `article:published_time` |
153
+ | `modified_time` | If present | From `article:modified_time` |
154
+ | `etag` | If sent | Response `ETag` header |
155
+ | `last_modified` | If sent | Response `Last-Modified` header |
156
+ | `requested_url` | If redirected | Original URL before redirects |
157
+
158
+ Add your own constant fields per run (`--frontmatter`) or per config
159
+ (`output.frontmatter_extra`) — custom keys never override the core provenance
160
+ fields above:
161
+
162
+ ```bash
163
+ gnosis https://example.com/docs --frontmatter 'tags: [customs, passar]' --frontmatter 'owner: kb-team'
164
+ ```
165
+
166
+ The frontmatter is standard YAML between `---` fences: parseable by
167
+ python-frontmatter, Jekyll, Hugo, Obsidian, and any downstream knowledge
168
+ pipeline.
169
+
170
+ Opt out per run with `--no-frontmatter` or globally in config:
171
+ ```yaml
172
+ output:
173
+ frontmatter: false
174
+ ```
175
+
176
+ ## Authenticated fetching
177
+
178
+ Secrets are read from **environment variables only**. They are never passed as
179
+ plain CLI arguments (which leak into shell history and process tables) and
180
+ never committed in config files.
181
+
182
+ ### Confluence Cloud with a Personal Access Token
183
+
184
+ ```bash
185
+ # Set up a PAT at https://id.atlassian.com/manage/api-tokens
186
+ export CONFLUENCE_PAT="your-api-token"
187
+
188
+ gnosis "https://your-domain.atlassian.net/wiki/spaces/SPACE/pages/PAGE_ID" \
189
+ --basic-user you@example.com \
190
+ --basic-token-env CONFLUENCE_PAT
191
+ ```
192
+
193
+ ### Bearer token (authenticated API docs, internal tools)
194
+
195
+ ```bash
196
+ export MY_API_TOKEN="..."
197
+ gnosis https://internal.example.com/docs --bearer-token-env MY_API_TOKEN
198
+ ```
199
+
200
+ ### Custom headers
201
+
202
+ ```bash
203
+ gnosis https://example.com --header "X-API-Key: ${MY_KEY}" --header "X-Team: docs"
204
+ ```
205
+
206
+ ### Via config file (multi-run / CI)
207
+
208
+ ```yaml
209
+ downloader:
210
+ auth:
211
+ type: basic # bearer | basic | header
212
+ username: "you@example.com"
213
+ password: "${CONFLUENCE_PAT}" # ${ENV_VAR} expanded at load time
214
+ ```
215
+
216
+ ## CLI reference
217
+
218
+ ```
219
+ gnosis URL [OPTIONS]
220
+ ```
221
+
222
+ | Flag | Description |
223
+ |---|---|
224
+ | `-a, --all` | Crawl all child pages under the URL path |
225
+ | `-n, --dry-run` | Discover and count pages only (requires `--all`) |
226
+ | `-o, --output DIR` | Output directory (default: `./`) |
227
+ | `-c, --config FILE` | Path to YAML configuration file |
228
+ | `-f, --overwrite` | Overwrite existing output files |
229
+ | `-q, --quiet` | Suppress progress output |
230
+ | `-v, --verbose` | Show detailed conversion diagnostics |
231
+ | `--no-frontmatter` | Write bare markdown without provenance block |
232
+ | `--frontmatter KEY: VALUE` | Extra constant frontmatter field (repeatable) |
233
+ | `--header NAME: VALUE` | Extra request header, `${ENV_VAR}` expanded (repeatable) |
234
+ | `--bearer-token-env VAR` | Bearer token from environment variable |
235
+ | `--basic-user USER` | HTTP Basic username (requires `--basic-token-env`) |
236
+ | `--basic-token-env VAR` | HTTP Basic password/token from environment variable |
237
+ | `--qmd-index` | Index output into QMD (requires `[qmd]` extra) |
238
+
239
+ ## Configuration reference
240
+
241
+ Copy [`config/default.yaml`](config/default.yaml) and pass it with `-c`.
242
+ Full reference:
243
+
244
+ ```yaml
245
+ # ── Downloader ────────────────────────────
246
+ downloader:
247
+ timeout: 30 # Request timeout (seconds)
248
+ retries: 3 # Retries on 5xx / network errors
249
+ user_agent: "Gnosis/1.1" # User-Agent header
250
+ rate_limit_ms: 500 # Minimum delay between requests (0 = no limit)
251
+ respect_robots: true # Obey robots.txt (future)
252
+ headers: {} # Extra HTTP headers (${ENV_VAR} expanded)
253
+ auth: # Optional: bearer | basic | header
254
+ type: bearer
255
+ token: "${MY_API_TOKEN}"
256
+
257
+ # ── Crawler ───────────────────────────────
258
+ crawler:
259
+ max_depth: 10 # Crawl depth from seed URL
260
+ max_pages: 500 # Stop after this many pages
261
+ concurrent_requests: 5 # Parallel fetch batch size (1 = sequential)
262
+
263
+ # ── Converter ─────────────────────────────
264
+ converter:
265
+ excluded_tags: [...] # HTML tags stripped before conversion
266
+ content_selectors: [...] # Tried in order; first match ≥ 200 chars wins
267
+ strip_classes: [...] # Exact class-token matches to remove
268
+ strip_class_words: [...] # Word-level matches inside class names
269
+ include_images: true # Emit <img> as ![alt](src)
270
+ absolute_urls: true # Resolve relative links to absolute URLs
271
+
272
+ # ── Output ─────────────────────────────────
273
+ output:
274
+ directory: "./" # Where .md files go
275
+ overwrite: false # Skip existing files unless true
276
+ extension: ".md" # Output file extension
277
+ frontmatter: true # Write YAML provenance block
278
+ frontmatter_extra: {} # Constant fields added to every file
279
+
280
+ # ── QMD (optional) ──────────────────────────
281
+ qmd:
282
+ enabled: false # Enable QMD knowledge base indexing
283
+ llm_model: "Qwen/Qwen3-0.6B" # HuggingFace model for context generation
284
+ llm_device: "cpu" # cpu | cuda | auto
285
+ ```
286
+
287
+ ## Exit codes
288
+
289
+ | Code | Meaning |
290
+ |---|---|
291
+ | `0` | Success (crawl mode: at least one page saved) |
292
+ | `1` | Failure — download error, file exists without `-f`, nothing saved, bad flags |
293
+ | `130` | Interrupted (Ctrl-C) |
294
+
295
+ In `--all` mode a `_manifest.json` is written to the output directory listing
296
+ every page with `url`, `file`, `content_hash`, `fetched_at`, `status_code`, and
297
+ `title` — ready for schedulers and downstream audit.
298
+
299
+ ## How clean is the output?
300
+
301
+ Gnosis is opinionated about boilerplate. By default it:
302
+
303
+ - **strips** script/style/nav/footer/aside/form/template tags and HTML comments
304
+ (including Confluence's `<!-- data-loadable-begin=... -->` SSR markers)
305
+ - **removes** elements matching exact `strip_classes` tokens AND boilerplate
306
+ *words* (`sidebar`, `toc`, `breadcrumb`, `cookie`, `headerlink`, `sourcelink`, …)
307
+ inside namespaced class names — `bd-sidebar-primary` is gone, but `research-content`
308
+ stays
309
+ - **cleans** permalink anchors from headings (`# Quickstart#` → `# Quickstart`)
310
+ - **picks** the main content by precedence-ordered selectors
311
+ (`.markdown-body`, `.ak-renderer-document`, `.wiki-content`, … before `main`/`#content`)
312
+ - **converts** tables to valid GFM: multi-line cells joined with `<br>`, `|`
313
+ escaped, duplicate/sticky-header rows removed
314
+ - **resolves** relative links to absolute URLs; skips `data:`-URI images
315
+ (spacers/tracking pixels)
316
+ - **unescapes** HTML entities in titles (`&#8212;` → `—`)
317
+
318
+ Everything is configurable — see [`config/default.yaml`](config/default.yaml).
319
+
320
+ ## Development & testing
321
+
322
+ ```bash
323
+ git clone https://github.com/SHCV-it/gnosis.git
324
+ cd gnosis
325
+ python -m venv venv && source venv/bin/activate
326
+ pip install -e . pytest python-frontmatter
327
+
328
+ # Run the test suite (offline — only localhost fixtures)
329
+ python -m pytest tests/ -q -v
330
+ ```
331
+
332
+ ### Project structure
333
+
334
+ ```
335
+ gnosis/
336
+ cli/ Click CLI (single page, crawl, dry-run, manifest)
337
+ config/ YAML loading + typed settings dataclass
338
+ core/
339
+ downloader Async HTTP client, auth, retries, FetchResult
340
+ converter HTML → Markdown, boilerplate stripping, metadata extraction
341
+ crawler BFS crawler with concurrent batch fetching
342
+ provenance Frontmatter generation, content_hash, render_document
343
+ integrations/ QMD pipeline (optional, heavy deps)
344
+ ```
345
+
346
+ The test suite covers converter quality (comment/anchor/boilerplate/table
347
+ handling, shadow-table dedup, metadata), provenance generation (fields,
348
+ round-trip parsing, extras merging), auth header injection (3 schemes),
349
+ crawler link resolution, and CLI end-to-end behavior. Runs entirely offline.
350
+
351
+ ## Contributing
352
+
353
+ Contributions are welcome. Please open an issue first to discuss what you'd
354
+ like to change.
355
+
356
+ 1. Fork the repository
357
+ 2. Create a feature branch (`git checkout -b feature/amazing`)
358
+ 3. Run the tests (`python -m pytest tests/ -q`)
359
+ 4. Commit your changes with clear messages
360
+ 5. Push and open a pull request against `main`
361
+
362
+ ## Related projects
363
+
364
+ Gnosis is designed to feed documentation pipelines and LLM knowledge bases.
365
+ Pair it with:
366
+
367
+ - **n8n / cron / Airflow** — schedule gnosis runs and pipe results into your
368
+ downstream pipeline
369
+ - **QMD** — local vector search via the `--qmd-index` flag
370
+ - **Any Markdown-to-anything pipeline** — the YAML frontmatter is parseable by
371
+ python-frontmatter, Jekyll, Hugo, Obsidian, and standard static-site generators
372
+
373
+ ## License
374
+
375
+ MIT — see [LICENSE](LICENSE).
376
+
377
+ ---
378
+
379
+ **Author:** Steffen Hoehne, [SHCV.IT](https://shcv.it)