docshelf-mcp 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. docshelf_mcp-0.2.0/.gitignore +38 -0
  2. docshelf_mcp-0.2.0/CHANGELOG.md +63 -0
  3. docshelf_mcp-0.2.0/LICENSE +21 -0
  4. docshelf_mcp-0.2.0/PKG-INFO +331 -0
  5. docshelf_mcp-0.2.0/README.md +295 -0
  6. docshelf_mcp-0.2.0/docs/ARCHITECTURE.md +93 -0
  7. docshelf_mcp-0.2.0/docs/PROJECT_PROMPT.md +163 -0
  8. docshelf_mcp-0.2.0/docs/USAGE.md +154 -0
  9. docshelf_mcp-0.2.0/docs/community/awesome-mcp-pr.md +70 -0
  10. docshelf_mcp-0.2.0/docs/community/reddit-show-hn.md +152 -0
  11. docshelf_mcp-0.2.0/docs/community/smithery-listing.md +137 -0
  12. docshelf_mcp-0.2.0/examples/homelab/README.md +84 -0
  13. docshelf_mcp-0.2.0/examples/recipes/README.md +65 -0
  14. docshelf_mcp-0.2.0/examples/research-papers/README.md +73 -0
  15. docshelf_mcp-0.2.0/pyproject.toml +95 -0
  16. docshelf_mcp-0.2.0/src/docshelf_mcp/__init__.py +20 -0
  17. docshelf_mcp-0.2.0/src/docshelf_mcp/__main__.py +6 -0
  18. docshelf_mcp-0.2.0/src/docshelf_mcp/config.py +28 -0
  19. docshelf_mcp-0.2.0/src/docshelf_mcp/core/__init__.py +6 -0
  20. docshelf_mcp-0.2.0/src/docshelf_mcp/core/converter.py +88 -0
  21. docshelf_mcp-0.2.0/src/docshelf_mcp/core/indexer.py +267 -0
  22. docshelf_mcp-0.2.0/src/docshelf_mcp/core/shelf.py +330 -0
  23. docshelf_mcp-0.2.0/src/docshelf_mcp/core/slugify.py +48 -0
  24. docshelf_mcp-0.2.0/src/docshelf_mcp/core/splitter.py +160 -0
  25. docshelf_mcp-0.2.0/src/docshelf_mcp/server.py +195 -0
  26. docshelf_mcp-0.2.0/src/docshelf_mcp/tools.py +345 -0
  27. docshelf_mcp-0.2.0/tests/__init__.py +0 -0
  28. docshelf_mcp-0.2.0/tests/fixtures/sample.md +22 -0
  29. docshelf_mcp-0.2.0/tests/test_indexer.py +125 -0
  30. docshelf_mcp-0.2.0/tests/test_server.py +144 -0
  31. docshelf_mcp-0.2.0/tests/test_shelf.py +141 -0
  32. docshelf_mcp-0.2.0/tests/test_slugify.py +46 -0
  33. docshelf_mcp-0.2.0/tests/test_splitter.py +97 -0
@@ -0,0 +1,38 @@
1
+ # Python
2
+ __pycache__/
3
+ *.py[cod]
4
+ *$py.class
5
+ *.so
6
+ .Python
7
+ *.egg-info/
8
+ *.egg
9
+ .eggs/
10
+ build/
11
+ dist/
12
+ .venv/
13
+ venv/
14
+ env/
15
+ ENV/
16
+ .python-version
17
+
18
+ # Tooling caches
19
+ .pytest_cache/
20
+ .ruff_cache/
21
+ .mypy_cache/
22
+ .coverage
23
+ .coverage.*
24
+ htmlcov/
25
+ pytest-cache-files-*/
26
+ pytest-of-*/
27
+
28
+ # Editor / OS
29
+ .DS_Store
30
+ *.swp
31
+ *.swo
32
+ .vscode/
33
+ .idea/
34
+
35
+ # macOS iCloud duplicates
36
+ * 2.md
37
+ * 2.py
38
+ * 2.json
@@ -0,0 +1,63 @@
1
+ # Changelog
2
+
3
+ All notable changes to docshelf-mcp will be documented in this file.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
+ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
+
8
+ ## [0.2.0] — 2026-05-14
9
+
10
+ Documentation, distribution, and community-onboarding pass. No runtime
11
+ changes — the library and MCP tools are byte-identical to v0.1.0.
12
+
13
+ ### Added
14
+ - `docs/PROJECT_PROMPT.md` — ready-to-use AI prompts (short / medium / full)
15
+ for projects that consume a docshelf via `INDEX.md`. Includes how-to
16
+ snippets for Claude Project, Claude Code, Claude Desktop, and the
17
+ Anthropic API.
18
+ - `.github/workflows/release.yml` — tag-triggered release workflow:
19
+ builds sdist + wheel, publishes to PyPI via **trusted publishing**
20
+ (OIDC, no API tokens), and creates a GitHub Release with the
21
+ matching CHANGELOG entry attached.
22
+ - `docs/community/` — submission materials for OSS distribution:
23
+ - `awesome-mcp-pr.md` — one-liner for the `awesome-mcp-servers` registry.
24
+ - `smithery-listing.md` — Smithery (smithery.ai) submission draft.
25
+ - `reddit-show-hn.md` — Reddit (`/r/ClaudeAI`, `/r/mcp`) and Show HN
26
+ announcement drafts.
27
+ - README: new **📋 Project Prompt** section linking to `docs/PROJECT_PROMPT.md`.
28
+ - README: PyPI install badge placeholder.
29
+
30
+ ### Changed
31
+ - `pyproject.toml` — version bumped to `0.2.0`.
32
+ - README install section expanded with the PyPI command + the git-source
33
+ fallback for users on `main`.
34
+
35
+ ### Notes
36
+ - First PyPI release will be cut by pushing the `v0.2.0` tag. The trusted
37
+ publisher needs to be configured once on the PyPI side
38
+ (`https://pypi.org/manage/account/publishing/` → add this repo, env name
39
+ `pypi`, workflow file `release.yml`).
40
+
41
+ ## [0.1.0] — 2026-05-13
42
+
43
+ Initial public release.
44
+
45
+ ### Added
46
+ - `init_shelf` tool — bootstrap a new shelf directory.
47
+ - `add_document` tool — PDF/Markdown ingestion with auto-split and INDEX
48
+ regeneration.
49
+ - `rebuild_index` tool — regenerate `INDEX.md` from on-disk state.
50
+ - `search` tool — plain-text grep with raw-URL enrichment.
51
+ - `list_documents` tool — categorized catalogue with size & section counts.
52
+ - `convert_pdf` tool — standalone PDF → Markdown utility.
53
+ - `pymupdf4llm` as the default conversion engine; optional `marker-pdf`
54
+ via `pip install docshelf-mcp[high-quality]`.
55
+ - `Shelf` class for direct library use (no MCP needed).
56
+ - Test suite (slugify, splitter, indexer, shelf, tool smoke).
57
+ - GitHub Actions CI (ruff + pytest on Python 3.10–3.12).
58
+ - Three example shelves: homelab, recipes, research-papers.
59
+
60
+ ### Notes
61
+ - Origin: this tool started life as a private script
62
+ (`outputs/homelab-encyclopedia.py`) that managed a homelab manuals repo.
63
+ v0.1.0 is the first generalised, course-agnostic release.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Filipp Ignatenko
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,331 @@
1
+ Metadata-Version: 2.4
2
+ Name: docshelf-mcp
3
+ Version: 0.2.0
4
+ Summary: MCP server for managing AI-friendly document collections — convert PDFs, split by chapter, index for chat projects.
5
+ Project-URL: Homepage, https://github.com/ignatenkofi/docshelf-mcp
6
+ Project-URL: Repository, https://github.com/ignatenkofi/docshelf-mcp
7
+ Project-URL: Issues, https://github.com/ignatenkofi/docshelf-mcp/issues
8
+ Project-URL: Changelog, https://github.com/ignatenkofi/docshelf-mcp/blob/main/CHANGELOG.md
9
+ Author-email: Filipp Ignatenko <ignatenkofi@gmail.com>
10
+ License: MIT
11
+ License-File: LICENSE
12
+ Keywords: claude,documentation,llm,markdown,mcp,model-context-protocol,pdf,rag
13
+ Classifier: Development Status :: 4 - Beta
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: License :: OSI Approved :: MIT License
16
+ Classifier: Operating System :: OS Independent
17
+ Classifier: Programming Language :: Python :: 3
18
+ Classifier: Programming Language :: Python :: 3.10
19
+ Classifier: Programming Language :: Python :: 3.11
20
+ Classifier: Programming Language :: Python :: 3.12
21
+ Classifier: Programming Language :: Python :: 3.13
22
+ Classifier: Topic :: Documentation
23
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
24
+ Classifier: Topic :: Text Processing :: Markup :: Markdown
25
+ Requires-Python: >=3.10
26
+ Requires-Dist: mcp>=1.2.0
27
+ Requires-Dist: pydantic>=2.6
28
+ Requires-Dist: pymupdf4llm>=0.0.17
29
+ Provides-Extra: dev
30
+ Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
31
+ Requires-Dist: pytest>=8.0; extra == 'dev'
32
+ Requires-Dist: ruff>=0.5.0; extra == 'dev'
33
+ Provides-Extra: high-quality
34
+ Requires-Dist: marker-pdf>=1.0.0; extra == 'high-quality'
35
+ Description-Content-Type: text/markdown
36
+
37
+ # docshelf-mcp
38
+
39
+ > Put your manuals on a shelf, hand the AI the index.
40
+
41
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
42
+ [![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)](https://www.python.org/downloads/)
43
+ [![MCP](https://img.shields.io/badge/MCP-compatible-purple.svg)](https://modelcontextprotocol.io/)
44
+ [![CI](https://github.com/ignatenkofi/docshelf-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/ignatenkofi/docshelf-mcp/actions/workflows/ci.yml)
45
+ [![PyPI](https://img.shields.io/pypi/v/docshelf-mcp.svg)](https://pypi.org/project/docshelf-mcp/)
46
+
47
+ ```text
48
+ ___ __ ____ ____ _ _ ____ __ ____
49
+ / __)/ \(_ _)/ ___)/ )( \( __)( ) ( __)
50
+ ( (_ \( O ) )( \___ \) __ ( ) _) / (_/\ ) _)
51
+ \___/ \__/ (__) (____/\_)(_/(____)\____/(__)
52
+ MCP server for AI-friendly doc shelves
53
+ ```
54
+
55
+ An [MCP](https://modelcontextprotocol.io/) server that turns a folder of PDFs and Markdown into a **chat-project-friendly document collection**: AI agents see a single `INDEX.md` and pull individual sections by raw GitHub URL on demand — instead of choking on a 4 MB datasheet.
56
+
57
+ ---
58
+
59
+ ## Why?
60
+
61
+ You have 30 hardware manuals, or 200 cooking recipes, or a stack of research PDFs.
62
+
63
+ You want Claude / ChatGPT / whatever to be able to answer questions across them — but:
64
+
65
+ - ❌ You can't dump 80 MB of PDFs into a chat project. It won't fit, and you'd burn the context window even if it did.
66
+ - ❌ You can manually copy-paste the relevant pages, but only after you remember which manual mentioned the thing you need.
67
+ - ❌ Long files mean retrieval is wasteful — the model loads the whole RouterOS guide just to answer a question about VLANs.
68
+
69
+ **docshelf-mcp** solves it like this:
70
+
71
+ 1. You drop a PDF onto the shelf.
72
+ 2. The shelf converts it to Markdown, splits big files chapter-by-chapter, and regenerates a navigation `INDEX.md`.
73
+ 3. You commit and push to a **public GitHub repo**.
74
+ 4. Add **only `INDEX.md`** to your Claude project. When the model needs a section, it fetches it via `raw.githubusercontent.com`.
75
+
76
+ Result: a 5 KB index pointing at a 50 MB collection. The model reads exactly the chapter it needs.
77
+
78
+ ---
79
+
80
+ ## 📦 Install
81
+
82
+ From PyPI (once the first tagged release is published):
83
+
84
+ ```bash
85
+ # uv (recommended)
86
+ uv pip install docshelf-mcp
87
+
88
+ # or plain pip
89
+ pip install docshelf-mcp
90
+ ```
91
+
92
+ Or straight from `main` (always-latest, no PyPI required):
93
+
94
+ ```bash
95
+ pip install "git+https://github.com/ignatenkofi/docshelf-mcp"
96
+ ```
97
+
98
+ Optional high-quality PDF engine (pulls ~2 GB of PyTorch — only if you need it):
99
+
100
+ ```bash
101
+ pip install "docshelf-mcp[high-quality]"
102
+ ```
103
+
104
+ ---
105
+
106
+ ## 📋 Project Prompt
107
+
108
+ Drop this into the **Custom Instructions** of any Claude project that consumes
109
+ a docshelf-style `INDEX.md`:
110
+
111
+ > This project uses the docshelf pattern. `INDEX.md` is the entry point.
112
+ > When answering: read INDEX → fetch ONLY the needed section file via its
113
+ > GitHub raw URL (use WebFetch / fetch / curl). Don't load full source files
114
+ > into context. For large manuals split into chapters, follow INDEX → chapter
115
+ > SUBINDEX → section file.
116
+
117
+ Medium (~150 words) and full (~400 words) versions, plus how-to snippets for
118
+ Claude Code, Claude Desktop, and the Anthropic API, live in
119
+ [`docs/PROJECT_PROMPT.md`](docs/PROJECT_PROMPT.md).
120
+
121
+ ---
122
+
123
+ ## Quickstart (Python library)
124
+
125
+ ```python
126
+ from docshelf_mcp import Shelf
127
+
128
+ shelf = Shelf("~/Documents/my-homelab-docs").init(
129
+ name="My HomeLab Docs",
130
+ remote="https://github.com/me/my-homelab-docs",
131
+ default_categories=["routers", "switches", "psu", "motherboards"],
132
+ )
133
+
134
+ shelf.add_document(
135
+ "~/Downloads/MIKROTIK_RouterOS.pdf",
136
+ category="routers",
137
+ title="Mikrotik RouterOS — full manual",
138
+ description="Official RouterOS reference, split by chapter.",
139
+ )
140
+ # → docs/routers/mikrotik-routeros-full-manual.md + docs/routers/.../001-..md, 002-..md, ...
141
+ # → INDEX.md is regenerated automatically.
142
+ ```
143
+
144
+ Then in the shelf directory: `git add . && git commit -m "docs: add RouterOS" && git push`.
145
+
146
+ In your Claude project, attach **only `INDEX.md`**. Done.
147
+
148
+ ---
149
+
150
+ ## Quickstart (MCP server)
151
+
152
+ ### 1. Add to Claude Desktop
153
+
154
+ Edit `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or `%APPDATA%/Claude/claude_desktop_config.json` (Windows):
155
+
156
+ ```json
157
+ {
158
+ "mcpServers": {
159
+ "docshelf": {
160
+ "command": "docshelf-mcp",
161
+ "env": {
162
+ "DOCSHELF_ROOT": "/Users/me/Documents/my-homelab-docs"
163
+ }
164
+ }
165
+ }
166
+ }
167
+ ```
168
+
169
+ Restart Claude Desktop. You now have six new tools available:
170
+
171
+ | Tool | What it does |
172
+ |---|---|
173
+ | `docshelf_init_shelf` | Bootstrap a new shelf directory. |
174
+ | `docshelf_add_document` | Add a PDF/MD file. Converts, splits, re-indexes. |
175
+ | `docshelf_rebuild_index` | Regenerate `INDEX.md` from disk. |
176
+ | `docshelf_search` | Plain-text search across the shelf, with raw URLs. |
177
+ | `docshelf_list_documents` | List documents by category. |
178
+ | `docshelf_convert_pdf` | Standalone PDF → Markdown (no shelf). |
179
+
180
+ ### 2. Add to Claude Code
181
+
182
+ ```bash
183
+ claude mcp add docshelf -- docshelf-mcp
184
+ # Optional: set the default shelf
185
+ claude mcp add docshelf --env DOCSHELF_ROOT=/path/to/shelf -- docshelf-mcp
186
+ ```
187
+
188
+ ### 3. Test from the command line
189
+
190
+ ```bash
191
+ # Sanity check — should print the server version then wait on stdin
192
+ docshelf-mcp
193
+ ```
194
+
195
+ ---
196
+
197
+ ## The shelf layout
198
+
199
+ ```text
200
+ my-shelf/
201
+ ├── .docshelf.json ← shelf metadata: name, remote, category order
202
+ ├── INDEX.md ← auto-generated navigation (your chat-project file)
203
+ ├── .gitignore
204
+ └── docs/
205
+ ├── routers/
206
+ │ ├── .meta.json ← per-document title/description overrides
207
+ │ ├── mikrotik-routeros.md (full document, lightly cleaned)
208
+ │ └── mikrotik-routeros/ (auto-split sections)
209
+ │ ├── 001-overview.md
210
+ │ ├── 002-bridging.md
211
+ │ └── 003-firewall.md
212
+ └── switches/
213
+ └── cudy-gs1010pe.md
214
+ ```
215
+
216
+ Everything in `docs/` is committed; everything is fetchable via raw URL once you push to GitHub.
217
+
218
+ ---
219
+
220
+ ## How splitting works
221
+
222
+ A document is split when **both** conditions hold:
223
+
224
+ 1. UTF-8 size > 50 KB (configurable via `.docshelf.json:split_threshold_bytes`).
225
+ 2. The document has at least two `## ` (H2) headings.
226
+
227
+ The splitter:
228
+
229
+ - Cleans PDF-extraction noise (collapses runaway blank lines, demotes CLI dumps mistaken for H1s).
230
+ - Slices on H2 boundaries.
231
+ - Names files `NNN-<slug>.md` so they sort naturally and survive title changes.
232
+ - Wipes the previous split directory before regenerating — fully idempotent.
233
+
234
+ If you want to keep a document whole, pass `split=False`.
235
+
236
+ ---
237
+
238
+ ## Examples
239
+
240
+ See the [`examples/`](examples) directory for three concrete use cases:
241
+
242
+ - **`examples/homelab/`** — original use case, hardware manuals for a home lab.
243
+ - **`examples/recipes/`** — a cookbook with one recipe per file.
244
+ - **`examples/research-papers/`** — academic PDFs with abstracts in `.meta.json`.
245
+
246
+ Each example shows the directory layout and the `INDEX.md` you'd end up with.
247
+
248
+ ---
249
+
250
+ ## Optional: high-quality PDF conversion
251
+
252
+ The default engine (`pymupdf4llm`) is fast and good enough for ~95% of technical documents. For papers with complex tables, math, or scanned content, install the `marker-pdf` backend:
253
+
254
+ ```bash
255
+ pip install "docshelf-mcp[high-quality]"
256
+ ```
257
+
258
+ Then pass `quality="high"`:
259
+
260
+ ```python
261
+ shelf.add_document("paper.pdf", category="research", title="...", quality="high")
262
+ ```
263
+
264
+ ⚠️ `marker-pdf` pulls in PyTorch (~2 GB) and is significantly slower (10–60 s per document on CPU). The library import is **deferred** — if you don't use `quality="high"`, the dependency is never loaded.
265
+
266
+ ---
267
+
268
+ ## FAQ
269
+
270
+ **Why GitHub raw URLs and not embeddings / RAG?**
271
+ Because it's dead simple, costs nothing to host, and the AI is already good at chasing links. You can layer embedding search on top later if you want — the on-disk shape is a normal git repo.
272
+
273
+ **Does this work with private repos?**
274
+ Not for the raw-URL trick — `raw.githubusercontent.com` won't serve them without auth. The local search tool works fine on private shelves; you just lose the "AI fetches sections directly" benefit. Make the doc repo public (separate from your code repo).
275
+
276
+ **Do I have to use GitHub?**
277
+ No. The shelf is just a directory. If you don't set a `github_remote`, INDEX.md still gets generated — entries just won't have URLs. You can host the static files anywhere that serves raw text (S3, Cloudflare R2, GitLab raw, Gitea, …) and post-process URLs yourself.
278
+
279
+ **Does it edit the source PDFs?**
280
+ No. PDFs are converted on `add_document` and the source is left in place. The shelf only writes inside its own directory.
281
+
282
+ **What about non-English documents?**
283
+ Slugify is Unicode-aware (NFKD-normalized, with `\w` under `re.UNICODE`). Cyrillic / CJK titles slug down to ASCII-ish forms; the body Markdown is preserved as-is.
284
+
285
+ **Can I use it without MCP?**
286
+ Yes — `from docshelf_mcp import Shelf` and use the class directly. See [`docs/USAGE.md`](docs/USAGE.md).
287
+
288
+ ---
289
+
290
+ ## Limitations
291
+
292
+ - **Public GitHub only** for the raw-URL trick (or whatever public static host you wire up).
293
+ - **Single repo per shelf.** If you outgrow one repo, run multiple shelves and attach multiple `INDEX.md`s.
294
+ - **Heuristic splitting.** The PDF→Markdown extract isn't always clean enough to split cleanly. For pathological cases (some 4+ MB datasheets), keep the file whole and rely on `docshelf_search`.
295
+ - **No automatic git commit.** Tools regenerate `INDEX.md` on disk, but the caller (you, or an agent) is responsible for `git add / commit / push`. This is intentional — staying out of git's way keeps the tool safe to call from agents.
296
+
297
+ ---
298
+
299
+ ## Demo
300
+
301
+ A short walkthrough video / GIF is planned: <https://github.com/ignatenkofi/docshelf-mcp/blob/main/docs/demo.md> *(coming soon)*
302
+
303
+ ---
304
+
305
+ ## Architecture
306
+
307
+ For a deeper dive, see [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md) — module layout, data flow, design rationale.
308
+
309
+ ---
310
+
311
+ ## Contributing
312
+
313
+ Bug reports and PRs welcome. To set up a dev env:
314
+
315
+ ```bash
316
+ git clone https://github.com/ignatenkofi/docshelf-mcp
317
+ cd docshelf-mcp
318
+ uv pip install -e ".[dev]"
319
+ ruff check src tests
320
+ pytest -v
321
+ ```
322
+
323
+ ---
324
+
325
+ ## License
326
+
327
+ MIT — see [`LICENSE`](LICENSE).
328
+
329
+ ## Origin
330
+
331
+ `docshelf-mcp` started life as a 350-line Python script (`homelab-encyclopedia.py`) that managed a single homelab manuals repo. The split / index / clean logic is the same code, generalised to work for any category-organised document collection.