dirag 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
dirag-0.1.0/.gitignore ADDED
@@ -0,0 +1,6 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ .venv/
4
+ dist/
5
+ build/
6
+ *.sqlite3
dirag-0.1.0/CLAUDE.md ADDED
@@ -0,0 +1,94 @@
1
+ # CLAUDE.md
2
+
3
+ dirag is a public, local-first tool: index a folder of PDF books, search it, and
4
+ read the page each result comes from. `README.md` is usage for people; this file
5
+ is how to work on the code.
6
+
7
+ ## Writing rules
8
+
9
+ - **Public project.** No personal names, machines, paths, libraries or usage
10
+ anywhere: code, comments, docs, examples, test data.
11
+ - **Current state, procedural.** Comments and docs say what the code does and
12
+ how to use it now. No history, no "used to", no reasons that only made sense
13
+ for an earlier version. Git keeps the story.
14
+ - **ASCII only** in code comments and Markdown. Check with
15
+ `LC_ALL=C grep -rn '[^ -~]' --include='*.py' --include='*.md' --include='*.js' --include='*.css' --include='*.html' .`
16
+ - **Minimal UI text.** Labels are one or two words; no instructions on screen.
17
+ Icons are Bootstrap Icons 1.11.3, paths copied verbatim from the upstream file
18
+ into `static/icons.js` with the icon's name beside each; never draw one.
19
+
20
+ ## Layout
21
+
22
+ ```text
23
+ src/dirag/
24
+ cli.py dirag [serve] | index | find | where | toc
25
+ config.py DIRAG_HOME, DIRAG_STATE, the library folder, per-folder index path
26
+ jobs.py the indexing job: job.json, start/stop from the app, stop checks
27
+ toc.py a PDF's table of contents -> page-partitioned sections
28
+ index.py update | reindex | rechunk: PDFs -> sections -> passages -> sqlite
29
+ embed.py the local models' device (GPU when it works, else CPU), embeddings
30
+ search.py modes (lexical, semantic, hybrid), rerank hook, per-book cap, chapters
31
+ rerank.py neural (fastembed cross-encoder) and llm rerankers
32
+ llm.py Ollama client: live availability, model pull, chat with an 8k context
33
+ answer.py quotes chosen by the model, verified, located on their pages
34
+ state.py positions, bookmarks, cards; per-key writes
35
+ server.py stdlib HTTP server and the JSON API
36
+ static/ index.html, app.js, app.css, icons.js, favicon.svg
37
+ vendor/ pdf.js (pdf.min.mjs, pdf.worker.min.mjs, LICENSE.pdfjs)
38
+ ```
39
+
40
+ ## Invariants
41
+
42
+ - **Books are keyed by path relative to the library folder**, in the index and
43
+ in every state file. Index ids are for requests only; they change on rebuild.
44
+ - **A request never names a file.** Books by id, passages by chunk id, static
45
+ files from the list built at startup.
46
+ - **State writes are per key**, merged into the file on disk, written through a
47
+ temp file and rename. Never accept a whole document from a client.
48
+ - **The stored passage text is verbatim.** The chapter header is embedded with
49
+ the passage but never stored on it or shown.
50
+ - **Two BM25 indexes**: `chunks_fts` (passage text) and `sections_fts` (section
51
+ titles). Do not fold titles into the passage index.
52
+ - **Pipeline order**: retrieve -> dedupe -> rerank -> per-book cap. The cap
53
+ comes last so it applies to the final order.
54
+ - **A language model never writes text shown as a source.** The llm reranker
55
+ returns passage numbers only; numbers out of range or repeated are ignored,
56
+ and the passages it leaves out stay in the list marked `dropped`. An answer
57
+ quote is shown only as the passage's own text for the span it matched (on
58
+ letters and digits); the model's response (one to three sentences) loses
59
+ citations to anything not quoted and is dropped when none remain.
60
+ - **Marks are rectangles in PDF points** (`rects`): a result's passage bbox, or
61
+ an answer quote's line boxes found through the page's words. The page image
62
+ and the reader draw the same list.
63
+ - **One install, no setup.** The dependencies are the GPU build of onnxruntime
64
+ (CPU build on macOS); `embed.build` uses the GPU only when nvidia-smi names
65
+ one and a first inference succeeds without onnxruntime falling back, and
66
+ otherwise runs on the CPU with the CPU provider named explicitly. The AI
67
+ features exist only while Ollama answers with the model; nothing else
68
+ depends on them.
69
+ - **No build step.** The browser loads `static/` as is; pdf.js is vendored. To
70
+ update pdf.js, copy `build/pdf.min.mjs`, `build/pdf.worker.min.mjs` and
71
+ `LICENSE` from the `pdfjs-dist` npm package into `static/vendor/`.
72
+ - **Indexing is resumable**: one commit per book; a book is skipped when its
73
+ size and mtime match, else when its content hash matches; cached page blocks
74
+ in `pages` mean `--rechunk` never re-parses.
75
+ - **One indexing job at a time**, recorded in `job.json` by the indexer itself,
76
+ so the app shows and stops a run started from either place. Stop is a request
77
+ checked before each book and after each embedding batch; the book in progress
78
+ rolls back. Liveness is checked through /proc (not a zombie, still a dirag
79
+ indexer), so a crashed run reads as failed.
80
+ - **One index per library folder** (`indexes/<name>-<hash>.sqlite3`); switching
81
+ folders never touches another folder's index.
82
+ - **The folder picker is confined** to the browse root, and `DIRAG_LIBRARY`
83
+ disables changing the folder from the app.
84
+
85
+ ## Running from a checkout
86
+
87
+ ```sh
88
+ uv venv --python 3.12 .venv
89
+ uv pip install --python .venv/bin/python -e .
90
+ .venv/bin/dirag
91
+ ```
92
+
93
+ Release: `uv build`, then `uv publish` (version in `pyproject.toml` and
94
+ `src/dirag/__init__.py`). Check the wheel carries `static/` and `static/vendor/`.
dirag-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Baris Arat
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
dirag-0.1.0/PKG-INFO ADDED
@@ -0,0 +1,196 @@
1
+ Metadata-Version: 2.5
2
+ Name: dirag
3
+ Version: 0.1.0
4
+ Summary: Search a folder of PDF books, read the page each result comes from, and get answers quoted from the books.
5
+ Project-URL: Homepage, https://github.com/barisarat/dirag
6
+ License-Expression: MIT
7
+ License-File: LICENSE
8
+ Keywords: bm25,books,embeddings,ollama,pdf,rag,search
9
+ Classifier: Environment :: Web Environment
10
+ Classifier: License :: OSI Approved :: MIT License
11
+ Classifier: Programming Language :: Python :: 3
12
+ Classifier: Topic :: Text Processing :: Indexing
13
+ Requires-Python: >=3.11
14
+ Requires-Dist: fastembed-gpu>=0.8; sys_platform == 'linux' or sys_platform == 'win32'
15
+ Requires-Dist: fastembed>=0.8; sys_platform == 'darwin'
16
+ Requires-Dist: onnxruntime-gpu[cuda,cudnn]>=1.22; sys_platform == 'linux' or sys_platform == 'win32'
17
+ Requires-Dist: pymupdf
18
+ Requires-Dist: sqlite-vec
19
+ Description-Content-Type: text/markdown
20
+
21
+ # dirag
22
+
23
+ Search a folder of PDF books, read the page each result comes from, and get
24
+ answers quoted from the books.
25
+
26
+ dirag indexes a library of PDFs once, then answers a query with ranked passages:
27
+ book, chapter, page and a snippet. Each result opens the rendered page with the
28
+ passage marked on it, or the whole book in the built-in reader at that page.
29
+
30
+ - Passages carry their chapter, taken from each PDF's table of contents.
31
+ - Three retrieval modes: lexical (BM25), semantic (embeddings), and hybrid
32
+ (both, fused by rank). An optional rerank reads query and passage together:
33
+ neural (a local cross-encoder) or an LLM.
34
+ - Results are capped per book, so one long series does not fill the list.
35
+ - The AI answer quotes the books: a language model picks the passages and the
36
+ words, dirag checks every quote against the text and shows it in the book's
37
+ own words, marked on its page.
38
+ - The reader remembers the page you stopped on in every book.
39
+ - Everything runs locally. No account, no network after the first downloads.
40
+
41
+ ## Start
42
+
43
+ ```sh
44
+ uvx dirag
45
+ ```
46
+
47
+ That needs only [uv](https://docs.astral.sh/uv/). The browser opens on the app:
48
+ choose the library folder (the folder button at the top of the shelf) and press
49
+ **Index new**. Subfolders are included.
50
+
51
+ dirag uses two things when they are there, with nothing to set:
52
+
53
+ | When present | Gives |
54
+ |---|---|
55
+ | An NVIDIA GPU | indexing in minutes per hundred books instead of hours; a faster neural rerank |
56
+ | [Ollama](https://ollama.com), running | the LLM rerank and the AI answer; dirag pulls `qwen2.5:7b-instruct` through it on the first start |
57
+
58
+ Without them, search, the neural rerank and the reader work as usual; the AI
59
+ button and the LLM rerank are hidden. The terminal shows what was found:
60
+
61
+ ```text
62
+ dirag on http://127.0.0.1:8008
63
+ GPU NVIDIA GeForce RTX 3060
64
+ AI qwen2.5:7b-instruct via Ollama
65
+ ```
66
+
67
+ The first start downloads the packages (about 2.5 GB on Linux and Windows,
68
+ GPU libraries included), the embedding and rerank models (about 250 MB) and,
69
+ with Ollama, the language model (about 4.7 GB). To keep dirag installed rather
70
+ than run it through uvx: `uv tool install dirag`, then `dirag`.
71
+
72
+ ## Indexing
73
+
74
+ Indexing is long by design: every page is parsed and every passage embedded.
75
+ The library view shows progress, a time estimate and a stop button. Stopping
76
+ keeps every finished book, and the next run continues from there. Search works
77
+ on the finished books while a run is going. **Index new** adds new and changed
78
+ PDFs and drops deleted ones; **Reindex all** parses and embeds everything again.
79
+
80
+ ## Search
81
+
82
+ | Key | Does |
83
+ |---|---|
84
+ | typing | lists quick exact-word matches under the box; pick one (click, or arrows and Enter) to open its page |
85
+ | Enter, or the search button | searches with the chosen mode and rerank |
86
+ | Ctrl+Enter, or the AI button | answers from the books (see below) |
87
+
88
+ The mode and rerank sit beside the search box:
89
+
90
+ | Mode | Finds |
91
+ |---|---|
92
+ | Hybrid | both of the below, fused by rank (the default) |
93
+ | Lexical | the exact words |
94
+ | Semantic | passages close in meaning, without the words |
95
+
96
+ | Rerank | Does |
97
+ |---|---|
98
+ | No rerank | keeps the retrieval order |
99
+ | Neural | scores the top 30 with a local cross-encoder; about a second on a CPU |
100
+ | LLM | asks the language model to order the top 20 and leave out the ones that do not help; those stay in the list, faded |
101
+
102
+ The line under the search box names what produced the results, with the count
103
+ and the time. While a search runs, the current results fade under a moving bar.
104
+
105
+ ## Answer
106
+
107
+ Ctrl+Enter or the AI button runs the search, then asks the model which passages
108
+ answer the question and which words to quote from each. Only what the model
109
+ chose is shown, as one card marked with the AI icon:
110
+
111
+ - its response, one to three sentences citing the quotes by number; citations
112
+ to anything not quoted are removed, and a response left with none is dropped;
113
+ - the quotes, as quotations in the books' own words, each with its source: the
114
+ source opens the reader scrolled to the quote, marked on the page, and the
115
+ page icon shows the page image in place.
116
+
117
+ Every quote is matched against its passage before it is shown; a quote that is
118
+ not in the text is dropped. When no passage answers the question, the answer
119
+ says so. The response is the model's reading, not the books; the quotes and
120
+ their pages are what to check.
121
+
122
+ ## Terminal
123
+
124
+ ```sh
125
+ dirag --port 9000 --no-browser # serve on another port, without opening the browser
126
+ dirag index --library ~/Books # choose the folder (remembered) and index it
127
+ dirag index # index new and changed PDFs, drop deleted ones
128
+ dirag index --reindex # parse and embed every PDF again
129
+ dirag index --rechunk # rebuild passages and vectors from cached pages
130
+ dirag find "martingale" --mode semantic --rerank neural # passages
131
+ dirag where "instrumental variables" # chapters
132
+ dirag toc book.pdf # the chapter map dirag reads from one PDF
133
+ ```
134
+
135
+ Ctrl-C stops an index run the same way the stop button does; press it twice to
136
+ stop at once.
137
+
138
+ ## Settings
139
+
140
+ All optional.
141
+
142
+ | Variable | Default | Sets |
143
+ |---|---|---|
144
+ | `DIRAG_HOME` | `~/.local/share/dirag` | indexes (one per library folder), models, the job record |
145
+ | `DIRAG_STATE` | `DIRAG_HOME` | `config.json`, `positions.json`, `bookmarks.json`, `cards.json` |
146
+ | `DIRAG_LIBRARY` | chosen in the app | fixes the library folder; the app then cannot change it |
147
+ | `DIRAG_BROWSE_ROOT` | the home directory | where the folder picker may browse |
148
+ | `DIRAG_LLM_MODEL` | `qwen2.5:7b-instruct` | the Ollama model |
149
+ | `DIRAG_LLM_URL` | `http://127.0.0.1:11434` | the Ollama server, which can be another machine |
150
+ | `DIRAG_RERANK_MODEL` | `Xenova/ms-marco-MiniLM-L-12-v2` | the cross-encoder; `BAAI/bge-reranker-base` is larger and slower |
151
+ | `DIRAG_DEVICE` | detected | `cpu` keeps the local models off the GPU |
152
+
153
+ Books are stored by their path inside the library, so the folder can move as
154
+ long as its contents keep their relative paths; a moved folder gets a new index
155
+ unless its index file is renamed to match.
156
+
157
+ `cards.json` overrides how a book is shown on the shelf. Keys are paths relative
158
+ to the library:
159
+
160
+ ```json
161
+ {"statistics/all-of-statistics.pdf": {"title": "All of Statistics", "subtitle": "A Concise Course",
162
+ "authors": ["Larry Wasserman"], "year": 2004, "pages": 442}}
163
+ ```
164
+
165
+ ## Running as a service
166
+
167
+ Install it with `uv tool install dirag`, then a systemd unit:
168
+
169
+ ```ini
170
+ [Unit]
171
+ Description=dirag
172
+ After=network-online.target
173
+
174
+ [Service]
175
+ ExecStart=%h/.local/bin/dirag --no-browser
176
+ Restart=on-failure
177
+
178
+ [Install]
179
+ WantedBy=default.target
180
+ ```
181
+
182
+ As a user unit (`~/.config/systemd/user/dirag.service`, then
183
+ `systemctl --user enable --now dirag`). Stopping the service stops a running
184
+ index job the same way the stop button does.
185
+
186
+ ## Security
187
+
188
+ The server binds 127.0.0.1 and has no login. Anyone who can reach the port can
189
+ read every book in the library, choose another folder under the browse root and
190
+ start indexing. Bind another address (`--host`) only on a network where every
191
+ device is trusted, and set `DIRAG_LIBRARY` to fix the folder.
192
+
193
+ ## License
194
+
195
+ MIT. pdf.js (Apache-2.0) and Bootstrap Icons (MIT) are included; see
196
+ `src/dirag/static/vendor/LICENSE.pdfjs`.
dirag-0.1.0/README.md ADDED
@@ -0,0 +1,176 @@
1
+ # dirag
2
+
3
+ Search a folder of PDF books, read the page each result comes from, and get
4
+ answers quoted from the books.
5
+
6
+ dirag indexes a library of PDFs once, then answers a query with ranked passages:
7
+ book, chapter, page and a snippet. Each result opens the rendered page with the
8
+ passage marked on it, or the whole book in the built-in reader at that page.
9
+
10
+ - Passages carry their chapter, taken from each PDF's table of contents.
11
+ - Three retrieval modes: lexical (BM25), semantic (embeddings), and hybrid
12
+ (both, fused by rank). An optional rerank reads query and passage together:
13
+ neural (a local cross-encoder) or an LLM.
14
+ - Results are capped per book, so one long series does not fill the list.
15
+ - The AI answer quotes the books: a language model picks the passages and the
16
+ words, dirag checks every quote against the text and shows it in the book's
17
+ own words, marked on its page.
18
+ - The reader remembers the page you stopped on in every book.
19
+ - Everything runs locally. No account, no network after the first downloads.
20
+
21
+ ## Start
22
+
23
+ ```sh
24
+ uvx dirag
25
+ ```
26
+
27
+ That needs only [uv](https://docs.astral.sh/uv/). The browser opens on the app:
28
+ choose the library folder (the folder button at the top of the shelf) and press
29
+ **Index new**. Subfolders are included.
30
+
31
+ dirag uses two things when they are there, with nothing to set:
32
+
33
+ | When present | Gives |
34
+ |---|---|
35
+ | An NVIDIA GPU | indexing in minutes per hundred books instead of hours; a faster neural rerank |
36
+ | [Ollama](https://ollama.com), running | the LLM rerank and the AI answer; dirag pulls `qwen2.5:7b-instruct` through it on the first start |
37
+
38
+ Without them, search, the neural rerank and the reader work as usual; the AI
39
+ button and the LLM rerank are hidden. The terminal shows what was found:
40
+
41
+ ```text
42
+ dirag on http://127.0.0.1:8008
43
+ GPU NVIDIA GeForce RTX 3060
44
+ AI qwen2.5:7b-instruct via Ollama
45
+ ```
46
+
47
+ The first start downloads the packages (about 2.5 GB on Linux and Windows,
48
+ GPU libraries included), the embedding and rerank models (about 250 MB) and,
49
+ with Ollama, the language model (about 4.7 GB). To keep dirag installed rather
50
+ than run it through uvx: `uv tool install dirag`, then `dirag`.
51
+
52
+ ## Indexing
53
+
54
+ Indexing is long by design: every page is parsed and every passage embedded.
55
+ The library view shows progress, a time estimate and a stop button. Stopping
56
+ keeps every finished book, and the next run continues from there. Search works
57
+ on the finished books while a run is going. **Index new** adds new and changed
58
+ PDFs and drops deleted ones; **Reindex all** parses and embeds everything again.
59
+
60
+ ## Search
61
+
62
+ | Key | Does |
63
+ |---|---|
64
+ | typing | lists quick exact-word matches under the box; pick one (click, or arrows and Enter) to open its page |
65
+ | Enter, or the search button | searches with the chosen mode and rerank |
66
+ | Ctrl+Enter, or the AI button | answers from the books (see below) |
67
+
68
+ The mode and rerank sit beside the search box:
69
+
70
+ | Mode | Finds |
71
+ |---|---|
72
+ | Hybrid | both of the below, fused by rank (the default) |
73
+ | Lexical | the exact words |
74
+ | Semantic | passages close in meaning, without the words |
75
+
76
+ | Rerank | Does |
77
+ |---|---|
78
+ | No rerank | keeps the retrieval order |
79
+ | Neural | scores the top 30 with a local cross-encoder; about a second on a CPU |
80
+ | LLM | asks the language model to order the top 20 and leave out the ones that do not help; those stay in the list, faded |
81
+
82
+ The line under the search box names what produced the results, with the count
83
+ and the time. While a search runs, the current results fade under a moving bar.
84
+
85
+ ## Answer
86
+
87
+ Ctrl+Enter or the AI button runs the search, then asks the model which passages
88
+ answer the question and which words to quote from each. Only what the model
89
+ chose is shown, as one card marked with the AI icon:
90
+
91
+ - its response, one to three sentences citing the quotes by number; citations
92
+ to anything not quoted are removed, and a response left with none is dropped;
93
+ - the quotes, as quotations in the books' own words, each with its source: the
94
+ source opens the reader scrolled to the quote, marked on the page, and the
95
+ page icon shows the page image in place.
96
+
97
+ Every quote is matched against its passage before it is shown; a quote that is
98
+ not in the text is dropped. When no passage answers the question, the answer
99
+ says so. The response is the model's reading, not the books; the quotes and
100
+ their pages are what to check.
101
+
102
+ ## Terminal
103
+
104
+ ```sh
105
+ dirag --port 9000 --no-browser # serve on another port, without opening the browser
106
+ dirag index --library ~/Books # choose the folder (remembered) and index it
107
+ dirag index # index new and changed PDFs, drop deleted ones
108
+ dirag index --reindex # parse and embed every PDF again
109
+ dirag index --rechunk # rebuild passages and vectors from cached pages
110
+ dirag find "martingale" --mode semantic --rerank neural # passages
111
+ dirag where "instrumental variables" # chapters
112
+ dirag toc book.pdf # the chapter map dirag reads from one PDF
113
+ ```
114
+
115
+ Ctrl-C stops an index run the same way the stop button does; press it twice to
116
+ stop at once.
117
+
118
+ ## Settings
119
+
120
+ All optional.
121
+
122
+ | Variable | Default | Sets |
123
+ |---|---|---|
124
+ | `DIRAG_HOME` | `~/.local/share/dirag` | indexes (one per library folder), models, the job record |
125
+ | `DIRAG_STATE` | `DIRAG_HOME` | `config.json`, `positions.json`, `bookmarks.json`, `cards.json` |
126
+ | `DIRAG_LIBRARY` | chosen in the app | fixes the library folder; the app then cannot change it |
127
+ | `DIRAG_BROWSE_ROOT` | the home directory | where the folder picker may browse |
128
+ | `DIRAG_LLM_MODEL` | `qwen2.5:7b-instruct` | the Ollama model |
129
+ | `DIRAG_LLM_URL` | `http://127.0.0.1:11434` | the Ollama server, which can be another machine |
130
+ | `DIRAG_RERANK_MODEL` | `Xenova/ms-marco-MiniLM-L-12-v2` | the cross-encoder; `BAAI/bge-reranker-base` is larger and slower |
131
+ | `DIRAG_DEVICE` | detected | `cpu` keeps the local models off the GPU |
132
+
133
+ Books are stored by their path inside the library, so the folder can move as
134
+ long as its contents keep their relative paths; a moved folder gets a new index
135
+ unless its index file is renamed to match.
136
+
137
+ `cards.json` overrides how a book is shown on the shelf. Keys are paths relative
138
+ to the library:
139
+
140
+ ```json
141
+ {"statistics/all-of-statistics.pdf": {"title": "All of Statistics", "subtitle": "A Concise Course",
142
+ "authors": ["Larry Wasserman"], "year": 2004, "pages": 442}}
143
+ ```
144
+
145
+ ## Running as a service
146
+
147
+ Install it with `uv tool install dirag`, then a systemd unit:
148
+
149
+ ```ini
150
+ [Unit]
151
+ Description=dirag
152
+ After=network-online.target
153
+
154
+ [Service]
155
+ ExecStart=%h/.local/bin/dirag --no-browser
156
+ Restart=on-failure
157
+
158
+ [Install]
159
+ WantedBy=default.target
160
+ ```
161
+
162
+ As a user unit (`~/.config/systemd/user/dirag.service`, then
163
+ `systemctl --user enable --now dirag`). Stopping the service stops a running
164
+ index job the same way the stop button does.
165
+
166
+ ## Security
167
+
168
+ The server binds 127.0.0.1 and has no login. Anyone who can reach the port can
169
+ read every book in the library, choose another folder under the browse root and
170
+ start indexing. Bind another address (`--host`) only on a network where every
171
+ device is trusted, and set `DIRAG_LIBRARY` to fix the folder.
172
+
173
+ ## License
174
+
175
+ MIT. pdf.js (Apache-2.0) and Bootstrap Icons (MIT) are included; see
176
+ `src/dirag/static/vendor/LICENSE.pdfjs`.
@@ -0,0 +1,36 @@
1
+ [build-system]
2
+ requires = ["hatchling"]
3
+ build-backend = "hatchling.build"
4
+
5
+ [project]
6
+ name = "dirag"
7
+ version = "0.1.0"
8
+ description = "Search a folder of PDF books, read the page each result comes from, and get answers quoted from the books."
9
+ readme = "README.md"
10
+ license = "MIT"
11
+ requires-python = ">=3.11"
12
+ # Linux and Windows get the GPU build of onnxruntime with CUDA and cuDNN as pip
13
+ # packages; it runs on the CPU when there is no NVIDIA GPU. macOS gets the CPU build.
14
+ dependencies = [
15
+ "pymupdf",
16
+ "sqlite-vec",
17
+ "fastembed-gpu>=0.8; sys_platform == 'linux' or sys_platform == 'win32'",
18
+ "onnxruntime-gpu[cuda,cudnn]>=1.22; sys_platform == 'linux' or sys_platform == 'win32'",
19
+ "fastembed>=0.8; sys_platform == 'darwin'",
20
+ ]
21
+ keywords = ["pdf", "books", "search", "rag", "bm25", "embeddings", "ollama"]
22
+ classifiers = [
23
+ "Programming Language :: Python :: 3",
24
+ "License :: OSI Approved :: MIT License",
25
+ "Environment :: Web Environment",
26
+ "Topic :: Text Processing :: Indexing",
27
+ ]
28
+
29
+ [project.urls]
30
+ Homepage = "https://github.com/barisarat/dirag"
31
+
32
+ [project.scripts]
33
+ dirag = "dirag.cli:main"
34
+
35
+ [tool.hatch.build.targets.wheel]
36
+ packages = ["src/dirag"]
@@ -0,0 +1,3 @@
1
+ """dirag: search a folder of PDF books and read the page each result comes from."""
2
+
3
+ __version__ = "0.1.0"
@@ -0,0 +1,5 @@
1
+ import sys
2
+
3
+ from .cli import main
4
+
5
+ sys.exit(main())
@@ -0,0 +1,129 @@
1
+ """The answer: quotes chosen by a language model, verified against the passages and located on the page.
2
+
3
+ The model sees the top passages of a search and returns passage numbers with
4
+ the words it quotes from each, plus a response of one to three sentences citing
5
+ them by number. Nothing
6
+ it writes is shown as a quote:
7
+
8
+ - Each quote is matched against its passage on letters and digits only, so
9
+ case, punctuation, spacing and line-break hyphens do not matter. A quote
10
+ with "..." is matched part by part, in order. A quote that does not match
11
+ is dropped.
12
+ - The text shown for a quote is the passage's own text for the matched span.
13
+ - The quote is located on its page through the page's words, giving one
14
+ rectangle per line for the reader and the page image to mark.
15
+ - Citations in the response to passages with no kept quote are removed, the
16
+ rest renumbered to the quotes. The response is kept only when at least one
17
+ citation remains; otherwise only the quotes are shown. With no kept quote the answer
18
+ is empty.
19
+ """
20
+
21
+ import json
22
+ import re
23
+
24
+ import pymupdf
25
+
26
+ from . import llm
27
+
28
+ DEPTH = 10
29
+ MAX_QUOTES = 5
30
+ MIN_ALNUM = 15
31
+
32
+ SYSTEM = ("You answer a question using only the numbered book passages given. Reply with JSON only:\n"
33
+ '{"quotes": [{"n": <passage number>, "text": "<words copied exactly from that passage>"}], '
34
+ '"summary": "<one to three sentences answering the question, each citing passages like [2]>"}\n'
35
+ "Copy every quote word for word from its passage; use ... only to skip words inside a quote. "
36
+ "Cite only passages you quoted. "
37
+ f"Each quote is one to three sentences. Order quotes from most to least useful; at most {MAX_QUOTES}. "
38
+ 'If no passage answers the question, reply {"quotes": [], "summary": ""}.')
39
+
40
+
41
+ def _alnum(text):
42
+ """Letters and digits of `text`, lowercased, with the index in `text` of each one."""
43
+ chars, where = [], []
44
+ for i, c in enumerate(text):
45
+ if c.isalnum():
46
+ chars.append(c.lower())
47
+ where.append(i)
48
+ return "".join(chars), where
49
+
50
+
51
+ def locate(quote, text):
52
+ """(start, end) of `quote` within `text`, matched on letters and digits, or None."""
53
+ haystack, where = _alnum(text)
54
+ parts = [p for p in (_alnum(part)[0] for part in re.split(r"\.\.\.|\u2026", quote)) if p]
55
+ if not parts or sum(len(p) for p in parts) < MIN_ALNUM:
56
+ return None
57
+ start, at = None, 0
58
+ for part in parts:
59
+ found = haystack.find(part, at)
60
+ if found == -1:
61
+ return None
62
+ start = found if start is None else start
63
+ at = found + len(part)
64
+ return where[start], where[at - 1] + 1
65
+
66
+
67
+ def page_rects(path, page_number, span_text):
68
+ """Line rectangles of `span_text` on a page, in PDF points, found through the page's words."""
69
+ doc = pymupdf.open(path)
70
+ try:
71
+ words = doc.load_page(page_number - 1).get_text("words")
72
+ finally:
73
+ doc.close()
74
+ joined, owner = [], []
75
+ for index, word in enumerate(words):
76
+ letters = _alnum(word[4])[0]
77
+ joined.append(letters)
78
+ owner.extend([index] * len(letters))
79
+ needle = _alnum(span_text)[0]
80
+ found = "".join(joined).find(needle)
81
+ if found == -1 or not needle:
82
+ return []
83
+ lines = {}
84
+ for index in sorted(set(owner[found:found + len(needle)])):
85
+ x0, y0, x1, y1, _, block, line, _ = words[index]
86
+ box = lines.setdefault((block, line), [x0, y0, x1, y1])
87
+ box[:] = [min(box[0], x0), min(box[1], y0), max(box[2], x1), max(box[3], y1)]
88
+ return [[round(v, 2) for v in box] for box in lines.values()]
89
+
90
+
91
+ def compose(question, rows, file_of):
92
+ """{"summary", "quotes"} for `rows` (search rows, best first). `file_of(rel)` gives a book's PDF path.
93
+
94
+ Raises llm.LLMError when the model call fails.
95
+ """
96
+ top = rows[:DEPTH]
97
+ listing = "\n\n".join(f"[{n}] {row['book']} > {row['section'] or ''}, p. {row['page']}\n{row['text']}"
98
+ for n, row in enumerate(top, 1))
99
+ reply = llm.chat_json(SYSTEM, f"Question: {question}\n\nPassages:\n\n{listing}")
100
+ quotes, numbers = [], {}
101
+ for item in reply.get("quotes", []) if isinstance(reply, dict) else []:
102
+ try:
103
+ n, said = int(item["n"]), str(item["text"])
104
+ except (KeyError, TypeError, ValueError):
105
+ continue
106
+ if not 1 <= n <= len(top) or len(quotes) >= MAX_QUOTES:
107
+ continue
108
+ row = top[n - 1]
109
+ span = locate(said, row["text"])
110
+ if span is None:
111
+ continue
112
+ text = row["text"][span[0]:span[1]]
113
+ if any(q["chunk_id"] == row["id"] and q["text"] == text for q in quotes):
114
+ continue
115
+ path = file_of(row["path"])
116
+ quotes.append({"n": len(quotes) + 1, "chunk_id": row["id"], "book_id": row["book_id"], "book": row["book"],
117
+ "year": row["year"], "section": row["section"], "page": row["page"], "text": text,
118
+ "rects": (page_rects(path, row["page"], text) if path else []) or _bbox(row)})
119
+ numbers.setdefault(n, quotes[-1]["n"])
120
+ summary = str(reply.get("summary") or "").strip() if quotes else ""
121
+ summary = re.sub(r"\s*\[(\d+)\]", lambda m: f" [{numbers[int(m.group(1))]}]" if int(m.group(1)) in numbers else "",
122
+ summary).strip()
123
+ if not re.search(r"\[\d+\]", summary):
124
+ summary = ""
125
+ return {"summary": summary, "quotes": quotes}
126
+
127
+
128
+ def _bbox(row):
129
+ return [json.loads(row["bbox_json"])] if row["bbox_json"] else []