dirag 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- dirag-0.1.0/.gitignore +6 -0
- dirag-0.1.0/CLAUDE.md +94 -0
- dirag-0.1.0/LICENSE +21 -0
- dirag-0.1.0/PKG-INFO +196 -0
- dirag-0.1.0/README.md +176 -0
- dirag-0.1.0/pyproject.toml +36 -0
- dirag-0.1.0/src/dirag/__init__.py +3 -0
- dirag-0.1.0/src/dirag/__main__.py +5 -0
- dirag-0.1.0/src/dirag/answer.py +129 -0
- dirag-0.1.0/src/dirag/cli.py +126 -0
- dirag-0.1.0/src/dirag/config.py +85 -0
- dirag-0.1.0/src/dirag/embed.py +130 -0
- dirag-0.1.0/src/dirag/index.py +405 -0
- dirag-0.1.0/src/dirag/jobs.py +129 -0
- dirag-0.1.0/src/dirag/llm.py +121 -0
- dirag-0.1.0/src/dirag/rerank.py +69 -0
- dirag-0.1.0/src/dirag/search.py +177 -0
- dirag-0.1.0/src/dirag/server.py +414 -0
- dirag-0.1.0/src/dirag/state.py +45 -0
- dirag-0.1.0/src/dirag/static/app.css +206 -0
- dirag-0.1.0/src/dirag/static/app.js +696 -0
- dirag-0.1.0/src/dirag/static/favicon.svg +1 -0
- dirag-0.1.0/src/dirag/static/icons.js +27 -0
- dirag-0.1.0/src/dirag/static/index.html +85 -0
- dirag-0.1.0/src/dirag/static/vendor/LICENSE.pdfjs +177 -0
- dirag-0.1.0/src/dirag/static/vendor/pdf.min.mjs +21 -0
- dirag-0.1.0/src/dirag/static/vendor/pdf.worker.min.mjs +21 -0
- dirag-0.1.0/src/dirag/toc.py +141 -0
dirag-0.1.0/.gitignore
ADDED
dirag-0.1.0/CLAUDE.md
ADDED
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
# CLAUDE.md
|
|
2
|
+
|
|
3
|
+
dirag is a public, local-first tool: index a folder of PDF books, search it, and
|
|
4
|
+
read the page each result comes from. `README.md` is usage for people; this file
|
|
5
|
+
is how to work on the code.
|
|
6
|
+
|
|
7
|
+
## Writing rules
|
|
8
|
+
|
|
9
|
+
- **Public project.** No personal names, machines, paths, libraries or usage
|
|
10
|
+
anywhere: code, comments, docs, examples, test data.
|
|
11
|
+
- **Current state, procedural.** Comments and docs say what the code does and
|
|
12
|
+
how to use it now. No history, no "used to", no reasons that only made sense
|
|
13
|
+
for an earlier version. Git keeps the story.
|
|
14
|
+
- **ASCII only** in code comments and Markdown. Check with
|
|
15
|
+
`LC_ALL=C grep -rn '[^ -~]' --include='*.py' --include='*.md' --include='*.js' --include='*.css' --include='*.html' .`
|
|
16
|
+
- **Minimal UI text.** Labels are one or two words; no instructions on screen.
|
|
17
|
+
Icons are Bootstrap Icons 1.11.3, paths copied verbatim from the upstream file
|
|
18
|
+
into `static/icons.js` with the icon's name beside each; never draw one.
|
|
19
|
+
|
|
20
|
+
## Layout
|
|
21
|
+
|
|
22
|
+
```text
|
|
23
|
+
src/dirag/
|
|
24
|
+
cli.py dirag [serve] | index | find | where | toc
|
|
25
|
+
config.py DIRAG_HOME, DIRAG_STATE, the library folder, per-folder index path
|
|
26
|
+
jobs.py the indexing job: job.json, start/stop from the app, stop checks
|
|
27
|
+
toc.py a PDF's table of contents -> page-partitioned sections
|
|
28
|
+
index.py update | reindex | rechunk: PDFs -> sections -> passages -> sqlite
|
|
29
|
+
embed.py the local models' device (GPU when it works, else CPU), embeddings
|
|
30
|
+
search.py modes (lexical, semantic, hybrid), rerank hook, per-book cap, chapters
|
|
31
|
+
rerank.py neural (fastembed cross-encoder) and llm rerankers
|
|
32
|
+
llm.py Ollama client: live availability, model pull, chat with an 8k context
|
|
33
|
+
answer.py quotes chosen by the model, verified, located on their pages
|
|
34
|
+
state.py positions, bookmarks, cards; per-key writes
|
|
35
|
+
server.py stdlib HTTP server and the JSON API
|
|
36
|
+
static/ index.html, app.js, app.css, icons.js, favicon.svg
|
|
37
|
+
vendor/ pdf.js (pdf.min.mjs, pdf.worker.min.mjs, LICENSE.pdfjs)
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
## Invariants
|
|
41
|
+
|
|
42
|
+
- **Books are keyed by path relative to the library folder**, in the index and
|
|
43
|
+
in every state file. Index ids are for requests only; they change on rebuild.
|
|
44
|
+
- **A request never names a file.** Books by id, passages by chunk id, static
|
|
45
|
+
files from the list built at startup.
|
|
46
|
+
- **State writes are per key**, merged into the file on disk, written through a
|
|
47
|
+
temp file and rename. Never accept a whole document from a client.
|
|
48
|
+
- **The stored passage text is verbatim.** The chapter header is embedded with
|
|
49
|
+
the passage but never stored on it or shown.
|
|
50
|
+
- **Two BM25 indexes**: `chunks_fts` (passage text) and `sections_fts` (section
|
|
51
|
+
titles). Do not fold titles into the passage index.
|
|
52
|
+
- **Pipeline order**: retrieve -> dedupe -> rerank -> per-book cap. The cap
|
|
53
|
+
comes last so it applies to the final order.
|
|
54
|
+
- **A language model never writes text shown as a source.** The llm reranker
|
|
55
|
+
returns passage numbers only; numbers out of range or repeated are ignored,
|
|
56
|
+
and the passages it leaves out stay in the list marked `dropped`. An answer
|
|
57
|
+
quote is shown only as the passage's own text for the span it matched (on
|
|
58
|
+
letters and digits); the model's response (one to three sentences) loses
|
|
59
|
+
citations to anything not quoted and is dropped when none remain.
|
|
60
|
+
- **Marks are rectangles in PDF points** (`rects`): a result's passage bbox, or
|
|
61
|
+
an answer quote's line boxes found through the page's words. The page image
|
|
62
|
+
and the reader draw the same list.
|
|
63
|
+
- **One install, no setup.** The dependencies are the GPU build of onnxruntime
|
|
64
|
+
(CPU build on macOS); `embed.build` uses the GPU only when nvidia-smi names
|
|
65
|
+
one and a first inference succeeds without onnxruntime falling back, and
|
|
66
|
+
otherwise runs on the CPU with the CPU provider named explicitly. The AI
|
|
67
|
+
features exist only while Ollama answers with the model; nothing else
|
|
68
|
+
depends on them.
|
|
69
|
+
- **No build step.** The browser loads `static/` as is; pdf.js is vendored. To
|
|
70
|
+
update pdf.js, copy `build/pdf.min.mjs`, `build/pdf.worker.min.mjs` and
|
|
71
|
+
`LICENSE` from the `pdfjs-dist` npm package into `static/vendor/`.
|
|
72
|
+
- **Indexing is resumable**: one commit per book; a book is skipped when its
|
|
73
|
+
size and mtime match, else when its content hash matches; cached page blocks
|
|
74
|
+
in `pages` mean `--rechunk` never re-parses.
|
|
75
|
+
- **One indexing job at a time**, recorded in `job.json` by the indexer itself,
|
|
76
|
+
so the app shows and stops a run started from either place. Stop is a request
|
|
77
|
+
checked before each book and after each embedding batch; the book in progress
|
|
78
|
+
rolls back. Liveness is checked through /proc (not a zombie, still a dirag
|
|
79
|
+
indexer), so a crashed run reads as failed.
|
|
80
|
+
- **One index per library folder** (`indexes/<name>-<hash>.sqlite3`); switching
|
|
81
|
+
folders never touches another folder's index.
|
|
82
|
+
- **The folder picker is confined** to the browse root, and `DIRAG_LIBRARY`
|
|
83
|
+
disables changing the folder from the app.
|
|
84
|
+
|
|
85
|
+
## Running from a checkout
|
|
86
|
+
|
|
87
|
+
```sh
|
|
88
|
+
uv venv --python 3.12 .venv
|
|
89
|
+
uv pip install --python .venv/bin/python -e .
|
|
90
|
+
.venv/bin/dirag
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Release: `uv build`, then `uv publish` (version in `pyproject.toml` and
|
|
94
|
+
`src/dirag/__init__.py`). Check the wheel carries `static/` and `static/vendor/`.
|
dirag-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Baris Arat
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
dirag-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,196 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: dirag
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Search a folder of PDF books, read the page each result comes from, and get answers quoted from the books.
|
|
5
|
+
Project-URL: Homepage, https://github.com/barisarat/dirag
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
License-File: LICENSE
|
|
8
|
+
Keywords: bm25,books,embeddings,ollama,pdf,rag,search
|
|
9
|
+
Classifier: Environment :: Web Environment
|
|
10
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
11
|
+
Classifier: Programming Language :: Python :: 3
|
|
12
|
+
Classifier: Topic :: Text Processing :: Indexing
|
|
13
|
+
Requires-Python: >=3.11
|
|
14
|
+
Requires-Dist: fastembed-gpu>=0.8; sys_platform == 'linux' or sys_platform == 'win32'
|
|
15
|
+
Requires-Dist: fastembed>=0.8; sys_platform == 'darwin'
|
|
16
|
+
Requires-Dist: onnxruntime-gpu[cuda,cudnn]>=1.22; sys_platform == 'linux' or sys_platform == 'win32'
|
|
17
|
+
Requires-Dist: pymupdf
|
|
18
|
+
Requires-Dist: sqlite-vec
|
|
19
|
+
Description-Content-Type: text/markdown
|
|
20
|
+
|
|
21
|
+
# dirag
|
|
22
|
+
|
|
23
|
+
Search a folder of PDF books, read the page each result comes from, and get
|
|
24
|
+
answers quoted from the books.
|
|
25
|
+
|
|
26
|
+
dirag indexes a library of PDFs once, then answers a query with ranked passages:
|
|
27
|
+
book, chapter, page and a snippet. Each result opens the rendered page with the
|
|
28
|
+
passage marked on it, or the whole book in the built-in reader at that page.
|
|
29
|
+
|
|
30
|
+
- Passages carry their chapter, taken from each PDF's table of contents.
|
|
31
|
+
- Three retrieval modes: lexical (BM25), semantic (embeddings), and hybrid
|
|
32
|
+
(both, fused by rank). An optional rerank reads query and passage together:
|
|
33
|
+
neural (a local cross-encoder) or an LLM.
|
|
34
|
+
- Results are capped per book, so one long series does not fill the list.
|
|
35
|
+
- The AI answer quotes the books: a language model picks the passages and the
|
|
36
|
+
words, dirag checks every quote against the text and shows it in the book's
|
|
37
|
+
own words, marked on its page.
|
|
38
|
+
- The reader remembers the page you stopped on in every book.
|
|
39
|
+
- Everything runs locally. No account, no network after the first downloads.
|
|
40
|
+
|
|
41
|
+
## Start
|
|
42
|
+
|
|
43
|
+
```sh
|
|
44
|
+
uvx dirag
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
That needs only [uv](https://docs.astral.sh/uv/). The browser opens on the app:
|
|
48
|
+
choose the library folder (the folder button at the top of the shelf) and press
|
|
49
|
+
**Index new**. Subfolders are included.
|
|
50
|
+
|
|
51
|
+
dirag uses two things when they are there, with nothing to set:
|
|
52
|
+
|
|
53
|
+
| When present | Gives |
|
|
54
|
+
|---|---|
|
|
55
|
+
| An NVIDIA GPU | indexing in minutes per hundred books instead of hours; a faster neural rerank |
|
|
56
|
+
| [Ollama](https://ollama.com), running | the LLM rerank and the AI answer; dirag pulls `qwen2.5:7b-instruct` through it on the first start |
|
|
57
|
+
|
|
58
|
+
Without them, search, the neural rerank and the reader work as usual; the AI
|
|
59
|
+
button and the LLM rerank are hidden. The terminal shows what was found:
|
|
60
|
+
|
|
61
|
+
```text
|
|
62
|
+
dirag on http://127.0.0.1:8008
|
|
63
|
+
GPU NVIDIA GeForce RTX 3060
|
|
64
|
+
AI qwen2.5:7b-instruct via Ollama
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
The first start downloads the packages (about 2.5 GB on Linux and Windows,
|
|
68
|
+
GPU libraries included), the embedding and rerank models (about 250 MB) and,
|
|
69
|
+
with Ollama, the language model (about 4.7 GB). To keep dirag installed rather
|
|
70
|
+
than run it through uvx: `uv tool install dirag`, then `dirag`.
|
|
71
|
+
|
|
72
|
+
## Indexing
|
|
73
|
+
|
|
74
|
+
Indexing is long by design: every page is parsed and every passage embedded.
|
|
75
|
+
The library view shows progress, a time estimate and a stop button. Stopping
|
|
76
|
+
keeps every finished book, and the next run continues from there. Search works
|
|
77
|
+
on the finished books while a run is going. **Index new** adds new and changed
|
|
78
|
+
PDFs and drops deleted ones; **Reindex all** parses and embeds everything again.
|
|
79
|
+
|
|
80
|
+
## Search
|
|
81
|
+
|
|
82
|
+
| Key | Does |
|
|
83
|
+
|---|---|
|
|
84
|
+
| typing | lists quick exact-word matches under the box; pick one (click, or arrows and Enter) to open its page |
|
|
85
|
+
| Enter, or the search button | searches with the chosen mode and rerank |
|
|
86
|
+
| Ctrl+Enter, or the AI button | answers from the books (see below) |
|
|
87
|
+
|
|
88
|
+
The mode and rerank sit beside the search box:
|
|
89
|
+
|
|
90
|
+
| Mode | Finds |
|
|
91
|
+
|---|---|
|
|
92
|
+
| Hybrid | both of the below, fused by rank (the default) |
|
|
93
|
+
| Lexical | the exact words |
|
|
94
|
+
| Semantic | passages close in meaning, without the words |
|
|
95
|
+
|
|
96
|
+
| Rerank | Does |
|
|
97
|
+
|---|---|
|
|
98
|
+
| No rerank | keeps the retrieval order |
|
|
99
|
+
| Neural | scores the top 30 with a local cross-encoder; about a second on a CPU |
|
|
100
|
+
| LLM | asks the language model to order the top 20 and leave out the ones that do not help; those stay in the list, faded |
|
|
101
|
+
|
|
102
|
+
The line under the search box names what produced the results, with the count
|
|
103
|
+
and the time. While a search runs, the current results fade under a moving bar.
|
|
104
|
+
|
|
105
|
+
## Answer
|
|
106
|
+
|
|
107
|
+
Ctrl+Enter or the AI button runs the search, then asks the model which passages
|
|
108
|
+
answer the question and which words to quote from each. Only what the model
|
|
109
|
+
chose is shown, as one card marked with the AI icon:
|
|
110
|
+
|
|
111
|
+
- its response, one to three sentences citing the quotes by number; citations
|
|
112
|
+
to anything not quoted are removed, and a response left with none is dropped;
|
|
113
|
+
- the quotes, as quotations in the books' own words, each with its source: the
|
|
114
|
+
source opens the reader scrolled to the quote, marked on the page, and the
|
|
115
|
+
page icon shows the page image in place.
|
|
116
|
+
|
|
117
|
+
Every quote is matched against its passage before it is shown; a quote that is
|
|
118
|
+
not in the text is dropped. When no passage answers the question, the answer
|
|
119
|
+
says so. The response is the model's reading, not the books; the quotes and
|
|
120
|
+
their pages are what to check.
|
|
121
|
+
|
|
122
|
+
## Terminal
|
|
123
|
+
|
|
124
|
+
```sh
|
|
125
|
+
dirag --port 9000 --no-browser # serve on another port, without opening the browser
|
|
126
|
+
dirag index --library ~/Books # choose the folder (remembered) and index it
|
|
127
|
+
dirag index # index new and changed PDFs, drop deleted ones
|
|
128
|
+
dirag index --reindex # parse and embed every PDF again
|
|
129
|
+
dirag index --rechunk # rebuild passages and vectors from cached pages
|
|
130
|
+
dirag find "martingale" --mode semantic --rerank neural # passages
|
|
131
|
+
dirag where "instrumental variables" # chapters
|
|
132
|
+
dirag toc book.pdf # the chapter map dirag reads from one PDF
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
Ctrl-C stops an index run the same way the stop button does; press it twice to
|
|
136
|
+
stop at once.
|
|
137
|
+
|
|
138
|
+
## Settings
|
|
139
|
+
|
|
140
|
+
All optional.
|
|
141
|
+
|
|
142
|
+
| Variable | Default | Sets |
|
|
143
|
+
|---|---|---|
|
|
144
|
+
| `DIRAG_HOME` | `~/.local/share/dirag` | indexes (one per library folder), models, the job record |
|
|
145
|
+
| `DIRAG_STATE` | `DIRAG_HOME` | `config.json`, `positions.json`, `bookmarks.json`, `cards.json` |
|
|
146
|
+
| `DIRAG_LIBRARY` | chosen in the app | fixes the library folder; the app then cannot change it |
|
|
147
|
+
| `DIRAG_BROWSE_ROOT` | the home directory | where the folder picker may browse |
|
|
148
|
+
| `DIRAG_LLM_MODEL` | `qwen2.5:7b-instruct` | the Ollama model |
|
|
149
|
+
| `DIRAG_LLM_URL` | `http://127.0.0.1:11434` | the Ollama server, which can be another machine |
|
|
150
|
+
| `DIRAG_RERANK_MODEL` | `Xenova/ms-marco-MiniLM-L-12-v2` | the cross-encoder; `BAAI/bge-reranker-base` is larger and slower |
|
|
151
|
+
| `DIRAG_DEVICE` | detected | `cpu` keeps the local models off the GPU |
|
|
152
|
+
|
|
153
|
+
Books are stored by their path inside the library, so the folder can move as
|
|
154
|
+
long as its contents keep their relative paths; a moved folder gets a new index
|
|
155
|
+
unless its index file is renamed to match.
|
|
156
|
+
|
|
157
|
+
`cards.json` overrides how a book is shown on the shelf. Keys are paths relative
|
|
158
|
+
to the library:
|
|
159
|
+
|
|
160
|
+
```json
|
|
161
|
+
{"statistics/all-of-statistics.pdf": {"title": "All of Statistics", "subtitle": "A Concise Course",
|
|
162
|
+
"authors": ["Larry Wasserman"], "year": 2004, "pages": 442}}
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
## Running as a service
|
|
166
|
+
|
|
167
|
+
Install it with `uv tool install dirag`, then a systemd unit:
|
|
168
|
+
|
|
169
|
+
```ini
|
|
170
|
+
[Unit]
|
|
171
|
+
Description=dirag
|
|
172
|
+
After=network-online.target
|
|
173
|
+
|
|
174
|
+
[Service]
|
|
175
|
+
ExecStart=%h/.local/bin/dirag --no-browser
|
|
176
|
+
Restart=on-failure
|
|
177
|
+
|
|
178
|
+
[Install]
|
|
179
|
+
WantedBy=default.target
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
As a user unit (`~/.config/systemd/user/dirag.service`, then
|
|
183
|
+
`systemctl --user enable --now dirag`). Stopping the service stops a running
|
|
184
|
+
index job the same way the stop button does.
|
|
185
|
+
|
|
186
|
+
## Security
|
|
187
|
+
|
|
188
|
+
The server binds 127.0.0.1 and has no login. Anyone who can reach the port can
|
|
189
|
+
read every book in the library, choose another folder under the browse root and
|
|
190
|
+
start indexing. Bind another address (`--host`) only on a network where every
|
|
191
|
+
device is trusted, and set `DIRAG_LIBRARY` to fix the folder.
|
|
192
|
+
|
|
193
|
+
## License
|
|
194
|
+
|
|
195
|
+
MIT. pdf.js (Apache-2.0) and Bootstrap Icons (MIT) are included; see
|
|
196
|
+
`src/dirag/static/vendor/LICENSE.pdfjs`.
|
dirag-0.1.0/README.md
ADDED
|
@@ -0,0 +1,176 @@
|
|
|
1
|
+
# dirag
|
|
2
|
+
|
|
3
|
+
Search a folder of PDF books, read the page each result comes from, and get
|
|
4
|
+
answers quoted from the books.
|
|
5
|
+
|
|
6
|
+
dirag indexes a library of PDFs once, then answers a query with ranked passages:
|
|
7
|
+
book, chapter, page and a snippet. Each result opens the rendered page with the
|
|
8
|
+
passage marked on it, or the whole book in the built-in reader at that page.
|
|
9
|
+
|
|
10
|
+
- Passages carry their chapter, taken from each PDF's table of contents.
|
|
11
|
+
- Three retrieval modes: lexical (BM25), semantic (embeddings), and hybrid
|
|
12
|
+
(both, fused by rank). An optional rerank reads query and passage together:
|
|
13
|
+
neural (a local cross-encoder) or an LLM.
|
|
14
|
+
- Results are capped per book, so one long series does not fill the list.
|
|
15
|
+
- The AI answer quotes the books: a language model picks the passages and the
|
|
16
|
+
words, dirag checks every quote against the text and shows it in the book's
|
|
17
|
+
own words, marked on its page.
|
|
18
|
+
- The reader remembers the page you stopped on in every book.
|
|
19
|
+
- Everything runs locally. No account, no network after the first downloads.
|
|
20
|
+
|
|
21
|
+
## Start
|
|
22
|
+
|
|
23
|
+
```sh
|
|
24
|
+
uvx dirag
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
That needs only [uv](https://docs.astral.sh/uv/). The browser opens on the app:
|
|
28
|
+
choose the library folder (the folder button at the top of the shelf) and press
|
|
29
|
+
**Index new**. Subfolders are included.
|
|
30
|
+
|
|
31
|
+
dirag uses two things when they are there, with nothing to set:
|
|
32
|
+
|
|
33
|
+
| When present | Gives |
|
|
34
|
+
|---|---|
|
|
35
|
+
| An NVIDIA GPU | indexing in minutes per hundred books instead of hours; a faster neural rerank |
|
|
36
|
+
| [Ollama](https://ollama.com), running | the LLM rerank and the AI answer; dirag pulls `qwen2.5:7b-instruct` through it on the first start |
|
|
37
|
+
|
|
38
|
+
Without them, search, the neural rerank and the reader work as usual; the AI
|
|
39
|
+
button and the LLM rerank are hidden. The terminal shows what was found:
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
dirag on http://127.0.0.1:8008
|
|
43
|
+
GPU NVIDIA GeForce RTX 3060
|
|
44
|
+
AI qwen2.5:7b-instruct via Ollama
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
The first start downloads the packages (about 2.5 GB on Linux and Windows,
|
|
48
|
+
GPU libraries included), the embedding and rerank models (about 250 MB) and,
|
|
49
|
+
with Ollama, the language model (about 4.7 GB). To keep dirag installed rather
|
|
50
|
+
than run it through uvx: `uv tool install dirag`, then `dirag`.
|
|
51
|
+
|
|
52
|
+
## Indexing
|
|
53
|
+
|
|
54
|
+
Indexing is long by design: every page is parsed and every passage embedded.
|
|
55
|
+
The library view shows progress, a time estimate and a stop button. Stopping
|
|
56
|
+
keeps every finished book, and the next run continues from there. Search works
|
|
57
|
+
on the finished books while a run is going. **Index new** adds new and changed
|
|
58
|
+
PDFs and drops deleted ones; **Reindex all** parses and embeds everything again.
|
|
59
|
+
|
|
60
|
+
## Search
|
|
61
|
+
|
|
62
|
+
| Key | Does |
|
|
63
|
+
|---|---|
|
|
64
|
+
| typing | lists quick exact-word matches under the box; pick one (click, or arrows and Enter) to open its page |
|
|
65
|
+
| Enter, or the search button | searches with the chosen mode and rerank |
|
|
66
|
+
| Ctrl+Enter, or the AI button | answers from the books (see below) |
|
|
67
|
+
|
|
68
|
+
The mode and rerank sit beside the search box:
|
|
69
|
+
|
|
70
|
+
| Mode | Finds |
|
|
71
|
+
|---|---|
|
|
72
|
+
| Hybrid | both of the below, fused by rank (the default) |
|
|
73
|
+
| Lexical | the exact words |
|
|
74
|
+
| Semantic | passages close in meaning, without the words |
|
|
75
|
+
|
|
76
|
+
| Rerank | Does |
|
|
77
|
+
|---|---|
|
|
78
|
+
| No rerank | keeps the retrieval order |
|
|
79
|
+
| Neural | scores the top 30 with a local cross-encoder; about a second on a CPU |
|
|
80
|
+
| LLM | asks the language model to order the top 20 and leave out the ones that do not help; those stay in the list, faded |
|
|
81
|
+
|
|
82
|
+
The line under the search box names what produced the results, with the count
|
|
83
|
+
and the time. While a search runs, the current results fade under a moving bar.
|
|
84
|
+
|
|
85
|
+
## Answer
|
|
86
|
+
|
|
87
|
+
Ctrl+Enter or the AI button runs the search, then asks the model which passages
|
|
88
|
+
answer the question and which words to quote from each. Only what the model
|
|
89
|
+
chose is shown, as one card marked with the AI icon:
|
|
90
|
+
|
|
91
|
+
- its response, one to three sentences citing the quotes by number; citations
|
|
92
|
+
to anything not quoted are removed, and a response left with none is dropped;
|
|
93
|
+
- the quotes, as quotations in the books' own words, each with its source: the
|
|
94
|
+
source opens the reader scrolled to the quote, marked on the page, and the
|
|
95
|
+
page icon shows the page image in place.
|
|
96
|
+
|
|
97
|
+
Every quote is matched against its passage before it is shown; a quote that is
|
|
98
|
+
not in the text is dropped. When no passage answers the question, the answer
|
|
99
|
+
says so. The response is the model's reading, not the books; the quotes and
|
|
100
|
+
their pages are what to check.
|
|
101
|
+
|
|
102
|
+
## Terminal
|
|
103
|
+
|
|
104
|
+
```sh
|
|
105
|
+
dirag --port 9000 --no-browser # serve on another port, without opening the browser
|
|
106
|
+
dirag index --library ~/Books # choose the folder (remembered) and index it
|
|
107
|
+
dirag index # index new and changed PDFs, drop deleted ones
|
|
108
|
+
dirag index --reindex # parse and embed every PDF again
|
|
109
|
+
dirag index --rechunk # rebuild passages and vectors from cached pages
|
|
110
|
+
dirag find "martingale" --mode semantic --rerank neural # passages
|
|
111
|
+
dirag where "instrumental variables" # chapters
|
|
112
|
+
dirag toc book.pdf # the chapter map dirag reads from one PDF
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
Ctrl-C stops an index run the same way the stop button does; press it twice to
|
|
116
|
+
stop at once.
|
|
117
|
+
|
|
118
|
+
## Settings
|
|
119
|
+
|
|
120
|
+
All optional.
|
|
121
|
+
|
|
122
|
+
| Variable | Default | Sets |
|
|
123
|
+
|---|---|---|
|
|
124
|
+
| `DIRAG_HOME` | `~/.local/share/dirag` | indexes (one per library folder), models, the job record |
|
|
125
|
+
| `DIRAG_STATE` | `DIRAG_HOME` | `config.json`, `positions.json`, `bookmarks.json`, `cards.json` |
|
|
126
|
+
| `DIRAG_LIBRARY` | chosen in the app | fixes the library folder; the app then cannot change it |
|
|
127
|
+
| `DIRAG_BROWSE_ROOT` | the home directory | where the folder picker may browse |
|
|
128
|
+
| `DIRAG_LLM_MODEL` | `qwen2.5:7b-instruct` | the Ollama model |
|
|
129
|
+
| `DIRAG_LLM_URL` | `http://127.0.0.1:11434` | the Ollama server, which can be another machine |
|
|
130
|
+
| `DIRAG_RERANK_MODEL` | `Xenova/ms-marco-MiniLM-L-12-v2` | the cross-encoder; `BAAI/bge-reranker-base` is larger and slower |
|
|
131
|
+
| `DIRAG_DEVICE` | detected | `cpu` keeps the local models off the GPU |
|
|
132
|
+
|
|
133
|
+
Books are stored by their path inside the library, so the folder can move as
|
|
134
|
+
long as its contents keep their relative paths; a moved folder gets a new index
|
|
135
|
+
unless its index file is renamed to match.
|
|
136
|
+
|
|
137
|
+
`cards.json` overrides how a book is shown on the shelf. Keys are paths relative
|
|
138
|
+
to the library:
|
|
139
|
+
|
|
140
|
+
```json
|
|
141
|
+
{"statistics/all-of-statistics.pdf": {"title": "All of Statistics", "subtitle": "A Concise Course",
|
|
142
|
+
"authors": ["Larry Wasserman"], "year": 2004, "pages": 442}}
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
## Running as a service
|
|
146
|
+
|
|
147
|
+
Install it with `uv tool install dirag`, then a systemd unit:
|
|
148
|
+
|
|
149
|
+
```ini
|
|
150
|
+
[Unit]
|
|
151
|
+
Description=dirag
|
|
152
|
+
After=network-online.target
|
|
153
|
+
|
|
154
|
+
[Service]
|
|
155
|
+
ExecStart=%h/.local/bin/dirag --no-browser
|
|
156
|
+
Restart=on-failure
|
|
157
|
+
|
|
158
|
+
[Install]
|
|
159
|
+
WantedBy=default.target
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
As a user unit (`~/.config/systemd/user/dirag.service`, then
|
|
163
|
+
`systemctl --user enable --now dirag`). Stopping the service stops a running
|
|
164
|
+
index job the same way the stop button does.
|
|
165
|
+
|
|
166
|
+
## Security
|
|
167
|
+
|
|
168
|
+
The server binds 127.0.0.1 and has no login. Anyone who can reach the port can
|
|
169
|
+
read every book in the library, choose another folder under the browse root and
|
|
170
|
+
start indexing. Bind another address (`--host`) only on a network where every
|
|
171
|
+
device is trusted, and set `DIRAG_LIBRARY` to fix the folder.
|
|
172
|
+
|
|
173
|
+
## License
|
|
174
|
+
|
|
175
|
+
MIT. pdf.js (Apache-2.0) and Bootstrap Icons (MIT) are included; see
|
|
176
|
+
`src/dirag/static/vendor/LICENSE.pdfjs`.
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
[build-system]
|
|
2
|
+
requires = ["hatchling"]
|
|
3
|
+
build-backend = "hatchling.build"
|
|
4
|
+
|
|
5
|
+
[project]
|
|
6
|
+
name = "dirag"
|
|
7
|
+
version = "0.1.0"
|
|
8
|
+
description = "Search a folder of PDF books, read the page each result comes from, and get answers quoted from the books."
|
|
9
|
+
readme = "README.md"
|
|
10
|
+
license = "MIT"
|
|
11
|
+
requires-python = ">=3.11"
|
|
12
|
+
# Linux and Windows get the GPU build of onnxruntime with CUDA and cuDNN as pip
|
|
13
|
+
# packages; it runs on the CPU when there is no NVIDIA GPU. macOS gets the CPU build.
|
|
14
|
+
dependencies = [
|
|
15
|
+
"pymupdf",
|
|
16
|
+
"sqlite-vec",
|
|
17
|
+
"fastembed-gpu>=0.8; sys_platform == 'linux' or sys_platform == 'win32'",
|
|
18
|
+
"onnxruntime-gpu[cuda,cudnn]>=1.22; sys_platform == 'linux' or sys_platform == 'win32'",
|
|
19
|
+
"fastembed>=0.8; sys_platform == 'darwin'",
|
|
20
|
+
]
|
|
21
|
+
keywords = ["pdf", "books", "search", "rag", "bm25", "embeddings", "ollama"]
|
|
22
|
+
classifiers = [
|
|
23
|
+
"Programming Language :: Python :: 3",
|
|
24
|
+
"License :: OSI Approved :: MIT License",
|
|
25
|
+
"Environment :: Web Environment",
|
|
26
|
+
"Topic :: Text Processing :: Indexing",
|
|
27
|
+
]
|
|
28
|
+
|
|
29
|
+
[project.urls]
|
|
30
|
+
Homepage = "https://github.com/barisarat/dirag"
|
|
31
|
+
|
|
32
|
+
[project.scripts]
|
|
33
|
+
dirag = "dirag.cli:main"
|
|
34
|
+
|
|
35
|
+
[tool.hatch.build.targets.wheel]
|
|
36
|
+
packages = ["src/dirag"]
|
|
@@ -0,0 +1,129 @@
|
|
|
1
|
+
"""The answer: quotes chosen by a language model, verified against the passages and located on the page.
|
|
2
|
+
|
|
3
|
+
The model sees the top passages of a search and returns passage numbers with
|
|
4
|
+
the words it quotes from each, plus a response of one to three sentences citing
|
|
5
|
+
them by number. Nothing
|
|
6
|
+
it writes is shown as a quote:
|
|
7
|
+
|
|
8
|
+
- Each quote is matched against its passage on letters and digits only, so
|
|
9
|
+
case, punctuation, spacing and line-break hyphens do not matter. A quote
|
|
10
|
+
with "..." is matched part by part, in order. A quote that does not match
|
|
11
|
+
is dropped.
|
|
12
|
+
- The text shown for a quote is the passage's own text for the matched span.
|
|
13
|
+
- The quote is located on its page through the page's words, giving one
|
|
14
|
+
rectangle per line for the reader and the page image to mark.
|
|
15
|
+
- Citations in the response to passages with no kept quote are removed, the
|
|
16
|
+
rest renumbered to the quotes. The response is kept only when at least one
|
|
17
|
+
citation remains; otherwise only the quotes are shown. With no kept quote the answer
|
|
18
|
+
is empty.
|
|
19
|
+
"""
|
|
20
|
+
|
|
21
|
+
import json
|
|
22
|
+
import re
|
|
23
|
+
|
|
24
|
+
import pymupdf
|
|
25
|
+
|
|
26
|
+
from . import llm
|
|
27
|
+
|
|
28
|
+
DEPTH = 10
|
|
29
|
+
MAX_QUOTES = 5
|
|
30
|
+
MIN_ALNUM = 15
|
|
31
|
+
|
|
32
|
+
SYSTEM = ("You answer a question using only the numbered book passages given. Reply with JSON only:\n"
|
|
33
|
+
'{"quotes": [{"n": <passage number>, "text": "<words copied exactly from that passage>"}], '
|
|
34
|
+
'"summary": "<one to three sentences answering the question, each citing passages like [2]>"}\n'
|
|
35
|
+
"Copy every quote word for word from its passage; use ... only to skip words inside a quote. "
|
|
36
|
+
"Cite only passages you quoted. "
|
|
37
|
+
f"Each quote is one to three sentences. Order quotes from most to least useful; at most {MAX_QUOTES}. "
|
|
38
|
+
'If no passage answers the question, reply {"quotes": [], "summary": ""}.')
|
|
39
|
+
|
|
40
|
+
|
|
41
|
+
def _alnum(text):
|
|
42
|
+
"""Letters and digits of `text`, lowercased, with the index in `text` of each one."""
|
|
43
|
+
chars, where = [], []
|
|
44
|
+
for i, c in enumerate(text):
|
|
45
|
+
if c.isalnum():
|
|
46
|
+
chars.append(c.lower())
|
|
47
|
+
where.append(i)
|
|
48
|
+
return "".join(chars), where
|
|
49
|
+
|
|
50
|
+
|
|
51
|
+
def locate(quote, text):
|
|
52
|
+
"""(start, end) of `quote` within `text`, matched on letters and digits, or None."""
|
|
53
|
+
haystack, where = _alnum(text)
|
|
54
|
+
parts = [p for p in (_alnum(part)[0] for part in re.split(r"\.\.\.|\u2026", quote)) if p]
|
|
55
|
+
if not parts or sum(len(p) for p in parts) < MIN_ALNUM:
|
|
56
|
+
return None
|
|
57
|
+
start, at = None, 0
|
|
58
|
+
for part in parts:
|
|
59
|
+
found = haystack.find(part, at)
|
|
60
|
+
if found == -1:
|
|
61
|
+
return None
|
|
62
|
+
start = found if start is None else start
|
|
63
|
+
at = found + len(part)
|
|
64
|
+
return where[start], where[at - 1] + 1
|
|
65
|
+
|
|
66
|
+
|
|
67
|
+
def page_rects(path, page_number, span_text):
|
|
68
|
+
"""Line rectangles of `span_text` on a page, in PDF points, found through the page's words."""
|
|
69
|
+
doc = pymupdf.open(path)
|
|
70
|
+
try:
|
|
71
|
+
words = doc.load_page(page_number - 1).get_text("words")
|
|
72
|
+
finally:
|
|
73
|
+
doc.close()
|
|
74
|
+
joined, owner = [], []
|
|
75
|
+
for index, word in enumerate(words):
|
|
76
|
+
letters = _alnum(word[4])[0]
|
|
77
|
+
joined.append(letters)
|
|
78
|
+
owner.extend([index] * len(letters))
|
|
79
|
+
needle = _alnum(span_text)[0]
|
|
80
|
+
found = "".join(joined).find(needle)
|
|
81
|
+
if found == -1 or not needle:
|
|
82
|
+
return []
|
|
83
|
+
lines = {}
|
|
84
|
+
for index in sorted(set(owner[found:found + len(needle)])):
|
|
85
|
+
x0, y0, x1, y1, _, block, line, _ = words[index]
|
|
86
|
+
box = lines.setdefault((block, line), [x0, y0, x1, y1])
|
|
87
|
+
box[:] = [min(box[0], x0), min(box[1], y0), max(box[2], x1), max(box[3], y1)]
|
|
88
|
+
return [[round(v, 2) for v in box] for box in lines.values()]
|
|
89
|
+
|
|
90
|
+
|
|
91
|
+
def compose(question, rows, file_of):
|
|
92
|
+
"""{"summary", "quotes"} for `rows` (search rows, best first). `file_of(rel)` gives a book's PDF path.
|
|
93
|
+
|
|
94
|
+
Raises llm.LLMError when the model call fails.
|
|
95
|
+
"""
|
|
96
|
+
top = rows[:DEPTH]
|
|
97
|
+
listing = "\n\n".join(f"[{n}] {row['book']} > {row['section'] or ''}, p. {row['page']}\n{row['text']}"
|
|
98
|
+
for n, row in enumerate(top, 1))
|
|
99
|
+
reply = llm.chat_json(SYSTEM, f"Question: {question}\n\nPassages:\n\n{listing}")
|
|
100
|
+
quotes, numbers = [], {}
|
|
101
|
+
for item in reply.get("quotes", []) if isinstance(reply, dict) else []:
|
|
102
|
+
try:
|
|
103
|
+
n, said = int(item["n"]), str(item["text"])
|
|
104
|
+
except (KeyError, TypeError, ValueError):
|
|
105
|
+
continue
|
|
106
|
+
if not 1 <= n <= len(top) or len(quotes) >= MAX_QUOTES:
|
|
107
|
+
continue
|
|
108
|
+
row = top[n - 1]
|
|
109
|
+
span = locate(said, row["text"])
|
|
110
|
+
if span is None:
|
|
111
|
+
continue
|
|
112
|
+
text = row["text"][span[0]:span[1]]
|
|
113
|
+
if any(q["chunk_id"] == row["id"] and q["text"] == text for q in quotes):
|
|
114
|
+
continue
|
|
115
|
+
path = file_of(row["path"])
|
|
116
|
+
quotes.append({"n": len(quotes) + 1, "chunk_id": row["id"], "book_id": row["book_id"], "book": row["book"],
|
|
117
|
+
"year": row["year"], "section": row["section"], "page": row["page"], "text": text,
|
|
118
|
+
"rects": (page_rects(path, row["page"], text) if path else []) or _bbox(row)})
|
|
119
|
+
numbers.setdefault(n, quotes[-1]["n"])
|
|
120
|
+
summary = str(reply.get("summary") or "").strip() if quotes else ""
|
|
121
|
+
summary = re.sub(r"\s*\[(\d+)\]", lambda m: f" [{numbers[int(m.group(1))]}]" if int(m.group(1)) in numbers else "",
|
|
122
|
+
summary).strip()
|
|
123
|
+
if not re.search(r"\[\d+\]", summary):
|
|
124
|
+
summary = ""
|
|
125
|
+
return {"summary": summary, "quotes": quotes}
|
|
126
|
+
|
|
127
|
+
|
|
128
|
+
def _bbox(row):
|
|
129
|
+
return [json.loads(row["bbox_json"])] if row["bbox_json"] else []
|