dirsql 0.3.69 → 0.3.71
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/docs/explanation.md +5 -0
- package/docs/howto/columns-from-paths.md +63 -0
- package/docs/howto/define-tables.md +70 -0
- package/docs/howto/embed.md +196 -0
- package/docs/howto/extract-from-contents.md +90 -0
- package/docs/howto/load-extension.md +91 -0
- package/docs/howto/persist.md +52 -0
- package/docs/howto/react-to-changes.md +67 -0
- package/docs/howto/search-by-meaning.md +148 -0
- package/docs/howto/skip-files.md +60 -0
- package/docs/reference/cli.md +132 -0
- package/docs/reference/columns.md +68 -0
- package/docs/reference/config.md +166 -0
- package/docs/reference/hooks.md +141 -0
- package/docs/reference/http-api.md +120 -0
- package/docs/reference/sdk.md +357 -0
- package/package.json +11 -11
- package/dist/cli/interpret/index.d.ts +0 -2
- package/dist/cli/interpret/index.d.ts.map +0 -1
- package/dist/cli/interpret/index.js +0 -15
- package/dist/cli/interpret/index.js.map +0 -1
package/README.md
CHANGED
|
@@ -89,7 +89,7 @@ Each event has `.action` (`'insert'` | `'update'` | `'delete'` | `'error'`), `.t
|
|
|
89
89
|
|
|
90
90
|
## CLI
|
|
91
91
|
|
|
92
|
-
`npx dirsql` runs an HTTP server exposing the SDK over HTTP: `POST /query` for SQL and `GET /events` for a Server-Sent Events change stream. Requires **Node >= 20.11**. See the [CLI
|
|
92
|
+
`npx dirsql` runs an HTTP server exposing the SDK over HTTP: `POST /query` for SQL and `GET /events` for a Server-Sent Events change stream. Requires **Node >= 20.11**. See the [CLI reference](https://thekevinscott.github.io/dirsql/reference/cli).
|
|
93
93
|
|
|
94
94
|
## License
|
|
95
95
|
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# Derive columns from file paths
|
|
2
|
+
|
|
3
|
+
Directory layouts often encode real data — an author, a year, a thread ID —
|
|
4
|
+
as path segments. A `{name}` capture in a table's glob turns such a segment
|
|
5
|
+
into a queryable column, no extraction code required.
|
|
6
|
+
|
|
7
|
+
## 1. Name the segment in the glob
|
|
8
|
+
|
|
9
|
+
Suppose photos are filed by year and month:
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
photos/2024/05/beach.jpg
|
|
13
|
+
photos/2024/11/hike.jpg
|
|
14
|
+
photos/2025/01/snow.jpg
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Capture both directory levels in `.dirsql.toml`:
|
|
18
|
+
|
|
19
|
+
```toml
|
|
20
|
+
[[table]]
|
|
21
|
+
ddl = "CREATE TABLE photos (year TEXT, month TEXT, _basename TEXT)"
|
|
22
|
+
glob = "photos/{year}/{month}/*.jpg"
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
A capture only populates a column when the DDL declares one with the same
|
|
26
|
+
name — here `year` and `month`. The capture rules (valid names, matching
|
|
27
|
+
within one path segment) are in
|
|
28
|
+
[glob captures](../reference/columns.md#glob-captures).
|
|
29
|
+
|
|
30
|
+
## 2. Query the captured columns
|
|
31
|
+
|
|
32
|
+
Start the server (`npx dirsql` / `uvx dirsql`) and query:
|
|
33
|
+
|
|
34
|
+
```bash
|
|
35
|
+
curl -s http://localhost:7117/query \
|
|
36
|
+
-H 'content-type: application/json' \
|
|
37
|
+
-d '{"sql":"SELECT year, month, _basename FROM photos ORDER BY year, month"}'
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
```json
|
|
41
|
+
[{"_basename":"beach.jpg","month":"05","year":"2024"},{"_basename":"hike.jpg","month":"11","year":"2024"},{"_basename":"snow.jpg","month":"01","year":"2025"}]
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Captures are real SQL columns, so aggregation works:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
curl -s http://localhost:7117/query \
|
|
48
|
+
-H 'content-type: application/json' \
|
|
49
|
+
-d '{"sql":"SELECT year, COUNT(*) AS photos FROM photos GROUP BY year"}'
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
```json
|
|
53
|
+
[{"photos":2,"year":"2024"},{"photos":1,"year":"2025"}]
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
## Going further
|
|
57
|
+
|
|
58
|
+
- Captures combine freely with [virtual columns](../reference/columns.md#virtual-columns)
|
|
59
|
+
(`_basename` above) — both are filesystem facts merged onto every row.
|
|
60
|
+
- The [tutorial](../getting-started.md) walks the same idea with an
|
|
61
|
+
`{author}` capture, starting from zero.
|
|
62
|
+
- When the value you need lives inside the file rather than in its path,
|
|
63
|
+
see [Extract rows from file contents](./extract-from-contents.md).
|
|
@@ -0,0 +1,70 @@
|
|
|
1
|
+
# Define tables for your files
|
|
2
|
+
|
|
3
|
+
Map a glob of files to a named SQL table so you query exactly the files you
|
|
4
|
+
care about, with exactly the columns you care about — instead of the
|
|
5
|
+
catch-all `files` table that [zero-config mode](../reference/cli.md#zero-config-mode)
|
|
6
|
+
serves.
|
|
7
|
+
|
|
8
|
+
## 1. Create a config next to your files
|
|
9
|
+
|
|
10
|
+
Suppose your blog posts live under `posts/`, one markdown file each. In the
|
|
11
|
+
directory you want to index, create a `.dirsql.toml` with one
|
|
12
|
+
[`[[table]]`](../reference/config.md#table) entry:
|
|
13
|
+
|
|
14
|
+
```toml
|
|
15
|
+
[[table]]
|
|
16
|
+
ddl = "CREATE TABLE posts (_path TEXT, _size INTEGER, _mtime INTEGER)"
|
|
17
|
+
glob = "posts/**/*.md"
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
- `glob` selects the files: every `.md` under `posts/`, at any depth,
|
|
21
|
+
relative to the directory containing the config.
|
|
22
|
+
- `ddl` is a plain SQLite `CREATE TABLE` naming the columns you want. Here
|
|
23
|
+
all three are [virtual columns](../reference/columns.md#virtual-columns) —
|
|
24
|
+
filesystem facts `dirsql` computes for every file. Facts are opt-in by
|
|
25
|
+
DDL: only the ones you declare become columns.
|
|
26
|
+
|
|
27
|
+
## 2. Start the server and query
|
|
28
|
+
|
|
29
|
+
::: code-group
|
|
30
|
+
|
|
31
|
+
```bash [npm]
|
|
32
|
+
npx dirsql
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
```bash [PyPI]
|
|
36
|
+
uvx dirsql
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
:::
|
|
40
|
+
|
|
41
|
+
Each matched file is one row:
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
curl -s http://localhost:7117/query \
|
|
45
|
+
-H 'content-type: application/json' \
|
|
46
|
+
-d '{"sql":"SELECT _path, _size FROM posts ORDER BY _path"}'
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
```json
|
|
50
|
+
[{"_path":"posts/2024/hello.md","_size":21},{"_path":"posts/2025/again.md","_size":55}]
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Files that don't match the glob (a `README.txt` next to `posts/`, say) are
|
|
54
|
+
simply not in the table. Once a config file exists, it fully replaces the
|
|
55
|
+
zero-config default — only the tables you define are served.
|
|
56
|
+
|
|
57
|
+
## Multiple tables
|
|
58
|
+
|
|
59
|
+
Add one `[[table]]` entry per table. When a file matches several globs, the
|
|
60
|
+
first matching table wins — see [`[[table]]`](../reference/config.md#table)
|
|
61
|
+
for that and the remaining keys (`strict`, `on-file`).
|
|
62
|
+
|
|
63
|
+
## Going further
|
|
64
|
+
|
|
65
|
+
- Your directory layout encodes data (authors, dates, IDs)? Capture path
|
|
66
|
+
segments as columns — [Derive columns from file paths](./columns-from-paths.md).
|
|
67
|
+
- Need columns from *inside* the files? A plain table never reads file
|
|
68
|
+
contents — [Extract rows from file contents](./extract-from-contents.md).
|
|
69
|
+
- Why one row per file, rebuilt from disk? See
|
|
70
|
+
[how `dirsql` thinks](../explanation.md).
|
|
@@ -0,0 +1,196 @@
|
|
|
1
|
+
# Embed `dirsql` in your application
|
|
2
|
+
|
|
3
|
+
Run the same engine in-process instead of as a server: construct a `DirSQL`
|
|
4
|
+
over a directory, hand it your tables, and query it like any library — no
|
|
5
|
+
HTTP, no separate process. The SDKs for Python, TypeScript, and Rust are
|
|
6
|
+
thin bindings over one shared core, so behavior matches the CLI exactly.
|
|
7
|
+
|
|
8
|
+
## 1. Install the SDK
|
|
9
|
+
|
|
10
|
+
::: code-group
|
|
11
|
+
|
|
12
|
+
```bash [Python]
|
|
13
|
+
uv add dirsql
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
```bash [TypeScript]
|
|
17
|
+
pnpm add dirsql
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
```bash [Rust]
|
|
21
|
+
cargo add dirsql
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
:::
|
|
25
|
+
|
|
26
|
+
## 2. Index a directory and query it
|
|
27
|
+
|
|
28
|
+
Suppose comments live as JSON files filed by thread:
|
|
29
|
+
|
|
30
|
+
```
|
|
31
|
+
comments/t1/c1.json # {"body": "first!", "author": "alice"}
|
|
32
|
+
comments/t1/c2.json # {"body": "nice post", "author": "bob"}
|
|
33
|
+
comments/t2/c1.json # {"body": "following up", "author": "alice"}
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
Unlike a config-file table, a programmatic table takes an `extract`
|
|
37
|
+
callback — your code reads each matched file and returns its rows, with
|
|
38
|
+
[glob captures and virtual columns](../reference/columns.md) merged on
|
|
39
|
+
automatically (here, `{thread}` from the path):
|
|
40
|
+
|
|
41
|
+
::: code-group
|
|
42
|
+
|
|
43
|
+
```python [Python]
|
|
44
|
+
import asyncio
|
|
45
|
+
import json
|
|
46
|
+
|
|
47
|
+
from dirsql import DirSQL, Table
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
def extract(path: str) -> list[dict]:
|
|
51
|
+
with open(path, encoding="utf-8") as f:
|
|
52
|
+
return [json.load(f)]
|
|
53
|
+
|
|
54
|
+
|
|
55
|
+
async def main() -> None:
|
|
56
|
+
db = DirSQL(
|
|
57
|
+
"./comments-root",
|
|
58
|
+
tables=[
|
|
59
|
+
Table(
|
|
60
|
+
ddl="CREATE TABLE comments (thread TEXT, author TEXT, body TEXT)",
|
|
61
|
+
glob="comments/{thread}/*.json",
|
|
62
|
+
extract=extract,
|
|
63
|
+
)
|
|
64
|
+
],
|
|
65
|
+
)
|
|
66
|
+
rows = await db.query("SELECT thread, author, body FROM comments ORDER BY thread, author")
|
|
67
|
+
for row in rows:
|
|
68
|
+
print(f"[{row['thread']}] {row['author']}: {row['body']}")
|
|
69
|
+
|
|
70
|
+
|
|
71
|
+
asyncio.run(main())
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
```typescript [TypeScript]
|
|
75
|
+
import { readFileSync } from "node:fs";
|
|
76
|
+
import { DirSQL } from "dirsql";
|
|
77
|
+
|
|
78
|
+
const db = new DirSQL({
|
|
79
|
+
root: "./comments-root",
|
|
80
|
+
tables: [
|
|
81
|
+
{
|
|
82
|
+
ddl: "CREATE TABLE comments (thread TEXT, author TEXT, body TEXT)",
|
|
83
|
+
glob: "comments/{thread}/*.json",
|
|
84
|
+
extract: (path) => [JSON.parse(readFileSync(path, "utf8"))],
|
|
85
|
+
},
|
|
86
|
+
],
|
|
87
|
+
});
|
|
88
|
+
|
|
89
|
+
const rows = await db.query(
|
|
90
|
+
"SELECT thread, author, body FROM comments ORDER BY thread, author",
|
|
91
|
+
);
|
|
92
|
+
for (const row of rows) {
|
|
93
|
+
console.log(`[${row.thread}] ${row.author}: ${row.body}`);
|
|
94
|
+
}
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
```rust [Rust]
|
|
98
|
+
use std::collections::HashMap;
|
|
99
|
+
|
|
100
|
+
use dirsql::{DirSQL, Table, Value};
|
|
101
|
+
|
|
102
|
+
fn extract(path: &str) -> Vec<HashMap<String, Value>> {
|
|
103
|
+
let raw = std::fs::read_to_string(path).unwrap_or_default();
|
|
104
|
+
let Ok(serde_json::Value::Object(obj)) = serde_json::from_str(&raw) else {
|
|
105
|
+
return vec![];
|
|
106
|
+
};
|
|
107
|
+
let mut row = HashMap::new();
|
|
108
|
+
for (key, value) in obj {
|
|
109
|
+
if let serde_json::Value::String(s) = value {
|
|
110
|
+
row.insert(key, Value::Text(s));
|
|
111
|
+
}
|
|
112
|
+
}
|
|
113
|
+
vec![row]
|
|
114
|
+
}
|
|
115
|
+
|
|
116
|
+
fn main() -> Result<(), Box<dyn std::error::Error>> {
|
|
117
|
+
let db = DirSQL::builder()
|
|
118
|
+
.root("./comments-root")
|
|
119
|
+
.table(Table::new(
|
|
120
|
+
"CREATE TABLE comments (thread TEXT, author TEXT, body TEXT)",
|
|
121
|
+
"comments/{thread}/*.json",
|
|
122
|
+
extract,
|
|
123
|
+
))
|
|
124
|
+
.build()?;
|
|
125
|
+
|
|
126
|
+
let rows = db.query("SELECT thread, author, body FROM comments ORDER BY thread, author")?;
|
|
127
|
+
for row in rows {
|
|
128
|
+
if let (Value::Text(thread), Value::Text(author), Value::Text(body)) =
|
|
129
|
+
(&row["thread"], &row["author"], &row["body"])
|
|
130
|
+
{
|
|
131
|
+
println!("[{thread}] {author}: {body}");
|
|
132
|
+
}
|
|
133
|
+
}
|
|
134
|
+
Ok(())
|
|
135
|
+
}
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
:::
|
|
139
|
+
|
|
140
|
+
Rows come back keyed by column name (a `dict` / object /
|
|
141
|
+
`HashMap<String, Value>` per row); all three programs print:
|
|
142
|
+
|
|
143
|
+
```
|
|
144
|
+
[t1] alice: first!
|
|
145
|
+
[t1] bob: nice post
|
|
146
|
+
[t2] alice: following up
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
A few things happened implicitly. In Python and TypeScript the constructor
|
|
150
|
+
returned immediately and the scan ran in the background — `query` awaited
|
|
151
|
+
readiness for you. In Rust, `.build()` scanned synchronously (there is a
|
|
152
|
+
[`build_async()`](../reference/sdk.md#asyncdirsql-rust) for the
|
|
153
|
+
non-blocking equivalent). Queries are read-only in every SDK. The complete
|
|
154
|
+
contract — constructor parameters, `ready`, value type mappings, error
|
|
155
|
+
shapes — is the [SDK reference](../reference/sdk.md).
|
|
156
|
+
|
|
157
|
+
## 3. React to changes in-process
|
|
158
|
+
|
|
159
|
+
`watch()` yields the same row events the server streams over
|
|
160
|
+
[`/events`](./react-to-changes.md), as your language's native async
|
|
161
|
+
iteration:
|
|
162
|
+
|
|
163
|
+
::: code-group
|
|
164
|
+
|
|
165
|
+
```python [Python]
|
|
166
|
+
async for event in db.watch():
|
|
167
|
+
print(event.action, event.table, event.row)
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
```typescript [TypeScript]
|
|
171
|
+
for await (const event of db.watch()) {
|
|
172
|
+
console.log(event.action, event.table, event.row);
|
|
173
|
+
}
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
```rust [Rust]
|
|
177
|
+
use futures::StreamExt;
|
|
178
|
+
|
|
179
|
+
let mut stream = db.watch()?;
|
|
180
|
+
while let Some(event) = stream.next().await {
|
|
181
|
+
println!("{event:?}");
|
|
182
|
+
}
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
:::
|
|
186
|
+
|
|
187
|
+
Event fields and actions are under
|
|
188
|
+
[`RowEvent`](../reference/sdk.md#rowevent).
|
|
189
|
+
|
|
190
|
+
## Reusing your CLI setup
|
|
191
|
+
|
|
192
|
+
Everything the config file expresses is available programmatically —
|
|
193
|
+
`ignore` patterns, `persist`, extensions — and the constructor's `config`
|
|
194
|
+
parameter loads a `.dirsql.toml` directly, so an application can share the
|
|
195
|
+
exact table definitions the CLI serves
|
|
196
|
+
([constructor reference](../reference/sdk.md#constructor)).
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
# Extract rows from file contents
|
|
2
|
+
|
|
3
|
+
Paths and stat metadata only get you so far — when the columns you want live
|
|
4
|
+
*inside* the files (JSON fields, frontmatter, log lines), add an
|
|
5
|
+
[`on-file`](../reference/config.md#table) command: it runs once per matched
|
|
6
|
+
file and its stdout becomes the file's rows.
|
|
7
|
+
|
|
8
|
+
## 1. Point a command at the files
|
|
9
|
+
|
|
10
|
+
Suppose each book is a JSON file:
|
|
11
|
+
|
|
12
|
+
```json
|
|
13
|
+
{"title": "Middlemarch", "author": "George Eliot", "year": 1871}
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
Any program that reads a file and prints a **JSON array of row objects** on
|
|
17
|
+
stdout works. With [`jq`](https://jqlang.org/):
|
|
18
|
+
|
|
19
|
+
```toml
|
|
20
|
+
[[table]]
|
|
21
|
+
ddl = "CREATE TABLE books (title TEXT, author TEXT, year INTEGER, _path TEXT)"
|
|
22
|
+
glob = "books/*.json"
|
|
23
|
+
on-file = "jq -c '[{title, author, year}]' {path}"
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
`{path}` is the matched file, relative to the index root — one of the
|
|
27
|
+
placeholders defined by the
|
|
28
|
+
[command hook contract](../reference/hooks.md#on-file), which also covers
|
|
29
|
+
the argv splitting, working directory, stdout protocol, and timeout shared
|
|
30
|
+
by every hook.
|
|
31
|
+
|
|
32
|
+
## 2. Query the extracted columns
|
|
33
|
+
|
|
34
|
+
Start the server (`npx dirsql` / `uvx dirsql`) and query:
|
|
35
|
+
|
|
36
|
+
```bash
|
|
37
|
+
curl -s http://localhost:7117/query \
|
|
38
|
+
-H 'content-type: application/json' \
|
|
39
|
+
-d '{"sql":"SELECT title, author, year, _path FROM books ORDER BY year"}'
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
```json
|
|
43
|
+
[{"_path":"books/bleak-house.json","author":"Charles Dickens","title":"Bleak House","year":1852},{"_path":"books/middlemarch.json","author":"George Eliot","title":"Middlemarch","year":1871}]
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
Filesystem facts are still merged onto every row — `_path` above comes from
|
|
47
|
+
`dirsql`, not from `jq`. When the command emits a key that collides with a
|
|
48
|
+
fact, the command wins
|
|
49
|
+
([precedence](../reference/columns.md#precedence)).
|
|
50
|
+
|
|
51
|
+
## Multiple rows per file
|
|
52
|
+
|
|
53
|
+
Each object in the printed array is one row. To turn a JSONL file into one
|
|
54
|
+
row per line, slurp it:
|
|
55
|
+
|
|
56
|
+
```toml
|
|
57
|
+
[[table]]
|
|
58
|
+
ddl = "CREATE TABLE events (event TEXT, user TEXT, _path TEXT)"
|
|
59
|
+
glob = "logs/*.jsonl"
|
|
60
|
+
on-file = "jq -c -s '.' {path}"
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
```bash
|
|
64
|
+
curl -s http://localhost:7117/query \
|
|
65
|
+
-H 'content-type: application/json' \
|
|
66
|
+
-d '{"sql":"SELECT event, user FROM events"}'
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
```json
|
|
70
|
+
[{"event":"login","user":"alice"},{"event":"logout","user":"alice"},{"event":"login","user":"bob"}]
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
## When a file fails
|
|
74
|
+
|
|
75
|
+
A file whose command errors (or prints something that isn't a JSON array of
|
|
76
|
+
objects) contributes no rows: `dirsql` warns on stderr and the scan
|
|
77
|
+
continues — one bad file never takes down the index. Details in
|
|
78
|
+
[failure semantics](../reference/hooks.md#failure-semantics); the JSON to
|
|
79
|
+
SQLite value mapping is under
|
|
80
|
+
[`on-file` row mapping](../reference/config.md#on-file-row-mapping).
|
|
81
|
+
|
|
82
|
+
## Going further
|
|
83
|
+
|
|
84
|
+
- The command re-runs on every startup and on every change to a matched
|
|
85
|
+
file. If it is expensive, [keep the index across restarts](./persist.md).
|
|
86
|
+
- The flagship use of `on-file` — computing embeddings — is
|
|
87
|
+
[Search documents by meaning](./search-by-meaning.md).
|
|
88
|
+
- Embedding `dirsql` in a program instead? The SDK's `extract` callback
|
|
89
|
+
fills the same role in-process — see
|
|
90
|
+
[Embed `dirsql` in your application](./embed.md).
|
|
@@ -0,0 +1,91 @@
|
|
|
1
|
+
# Load a SQLite extension
|
|
2
|
+
|
|
3
|
+
SQLite extensions add functions to your query surface — vector distances,
|
|
4
|
+
crypto digests, statistics. `dirsql` loads them at startup from a
|
|
5
|
+
[`[[dirsql.extension]]`](../reference/config.md#dirsql-extension) entry,
|
|
6
|
+
named either by file path or by installed package.
|
|
7
|
+
|
|
8
|
+
## Option A: a literal file path
|
|
9
|
+
|
|
10
|
+
Works everywhere — the `uvx`/`npx` launchers, the standalone
|
|
11
|
+
`cargo install` binary, and every SDK. Point `path` at the extension's
|
|
12
|
+
shared library (here [`sqlite-vec`](https://github.com/asg017/sqlite-vec)'s
|
|
13
|
+
`vec0.so`, copied into the project):
|
|
14
|
+
|
|
15
|
+
```toml
|
|
16
|
+
[[dirsql.extension]]
|
|
17
|
+
path = "./ext/vec0.so"
|
|
18
|
+
entrypoint = "sqlite3_vec_init"
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
Relative paths resolve against the config file's directory. `entrypoint`
|
|
22
|
+
overrides the init symbol when it doesn't match the filename-derived
|
|
23
|
+
default — `sqlite-vec` is exactly such a case
|
|
24
|
+
([reference](../reference/config.md#dirsql-extension)).
|
|
25
|
+
|
|
26
|
+
Start the server (`npx dirsql` / `uvx dirsql`) and the extension's
|
|
27
|
+
functions are callable:
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
curl -s http://localhost:7117/query \
|
|
31
|
+
-H 'content-type: application/json' \
|
|
32
|
+
-d '{"sql":"SELECT vec_version() AS vec_version"}'
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
```json
|
|
36
|
+
[{"vec_version":"v0.1.9"}]
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
## Option B: a package name
|
|
40
|
+
|
|
41
|
+
When the extension ships as a pip or npm package, name the package and let
|
|
42
|
+
the launcher find the loadable inside it — no path to copy or keep in sync.
|
|
43
|
+
How a bare name resolves (file-first probing, one-loadable rule) is defined
|
|
44
|
+
under [`[[dirsql.extension]]`](../reference/config.md#dirsql-extension);
|
|
45
|
+
what to actually write differs per runtime:
|
|
46
|
+
|
|
47
|
+
**Python (`uvx dirsql`, Python SDK).** Use the *importable module name* —
|
|
48
|
+
for the `sqlite-vec` pip package that is `sqlite_vec` (underscore), which
|
|
49
|
+
is looked up via `importlib`:
|
|
50
|
+
|
|
51
|
+
```toml
|
|
52
|
+
[[dirsql.extension]]
|
|
53
|
+
path = "sqlite_vec"
|
|
54
|
+
entrypoint = "sqlite3_vec_init"
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
uvx --with sqlite-vec dirsql
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
**Node (`npx dirsql`, TypeScript SDK).** Use the *npm package name* — but
|
|
62
|
+
note the named package must itself contain the loadable file. `sqlite-vec`
|
|
63
|
+
on npm is a meta-package whose loadable lives in a platform-specific
|
|
64
|
+
sub-package, so name that one (and install it) instead:
|
|
65
|
+
|
|
66
|
+
```toml
|
|
67
|
+
[[dirsql.extension]]
|
|
68
|
+
path = "sqlite-vec-linux-x64" # or -darwin-arm64, etc.
|
|
69
|
+
entrypoint = "sqlite3_vec_init"
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
If a platform-specific config is awkward, fall back to a literal path
|
|
73
|
+
(Option A).
|
|
74
|
+
|
|
75
|
+
**Rust (standalone binary, Rust SDK).** File paths only — there is no
|
|
76
|
+
interpreter to resolve package names with
|
|
77
|
+
([reference](../reference/config.md#dirsql-extension)).
|
|
78
|
+
|
|
79
|
+
## Notes
|
|
80
|
+
|
|
81
|
+
- Extensions add **functions**. An extension-backed virtual table cannot be
|
|
82
|
+
declared as a `[[table]]` — `dirsql` tables are per-file row tables
|
|
83
|
+
([reference](../reference/config.md#dirsql-extension)).
|
|
84
|
+
- Loading happens before any table DDL runs, and the SQL
|
|
85
|
+
`load_extension()` function is never exposed to queries
|
|
86
|
+
([reference](../reference/config.md#dirsql-extension)).
|
|
87
|
+
- Embedding `dirsql` in a program? The SDK constructor takes the same
|
|
88
|
+
specs via its `extensions` parameter
|
|
89
|
+
([SDK reference](../reference/sdk.md#constructor)).
|
|
90
|
+
- The payoff use case — `sqlite-vec` powering semantic search — is
|
|
91
|
+
[Search documents by meaning](./search-by-meaning.md).
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# Keep the index across restarts
|
|
2
|
+
|
|
3
|
+
By default the database is ephemeral: rebuilt from your files on every
|
|
4
|
+
startup and discarded on exit. [`persist`](../reference/config.md#dirsql-keys) keeps the
|
|
5
|
+
SQLite index on disk instead, so a restart only re-parses files that
|
|
6
|
+
actually changed — the difference between seconds and milliseconds on large
|
|
7
|
+
trees, and between re-running and skipping expensive
|
|
8
|
+
[`on-file`](./extract-from-contents.md) commands.
|
|
9
|
+
|
|
10
|
+
## 1. Turn it on
|
|
11
|
+
|
|
12
|
+
```toml
|
|
13
|
+
[dirsql]
|
|
14
|
+
persist = true
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
That's the whole change. On the next run the cache is written to
|
|
18
|
+
`.dirsql/cache.db` under the root; runs after that start from it. To put
|
|
19
|
+
the cache elsewhere (a CI cache dir, a tmpfs), set
|
|
20
|
+
[`persist_path`](../reference/config.md#dirsql-keys).
|
|
21
|
+
|
|
22
|
+
## 2. Keep the cache out of git
|
|
23
|
+
|
|
24
|
+
The cache is derived data — reproducible from the tree and frequently
|
|
25
|
+
large. Add it to `.gitignore`:
|
|
26
|
+
|
|
27
|
+
```
|
|
28
|
+
.dirsql/
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
The top-level `.dirsql/` directory is reserved for `dirsql`'s metadata and
|
|
32
|
+
is never scanned as data, so the cache can't index itself
|
|
33
|
+
([config reference](../reference/config.md#dirsql-keys)).
|
|
34
|
+
|
|
35
|
+
## What survives, what rebuilds
|
|
36
|
+
|
|
37
|
+
On startup `dirsql` validates the cache rather than trusting it blindly:
|
|
38
|
+
files whose stat metadata is unchanged keep their rows without being
|
|
39
|
+
re-read; changed, added, and deleted files are reconciled. When the cache
|
|
40
|
+
can't be trusted at all — the table/ignore configuration changed, or the
|
|
41
|
+
`dirsql` version did — it is discarded and rebuilt from scratch
|
|
42
|
+
automatically. You never need to delete it by hand; a full rebuild costs
|
|
43
|
+
exactly what a non-persistent startup does.
|
|
44
|
+
|
|
45
|
+
Persistence is a startup-time optimization, not a change in meaning: the
|
|
46
|
+
database remains a derived view of your files, and queries return the same
|
|
47
|
+
rows either way ([how `dirsql` thinks](../explanation.md)).
|
|
48
|
+
|
|
49
|
+
## Embedding `dirsql`?
|
|
50
|
+
|
|
51
|
+
The SDK constructors expose the same switch as `persist` / `persistPath`
|
|
52
|
+
parameters — see the [SDK reference](../reference/sdk.md#constructor).
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# React to file changes
|
|
2
|
+
|
|
3
|
+
`dirsql` watches the directory it indexes, so your tables update as files
|
|
4
|
+
change — and [`GET /events`](../reference/http-api.md#get-events) pushes
|
|
5
|
+
every row-level change to you as it happens. No polling, no diffing on your
|
|
6
|
+
side.
|
|
7
|
+
|
|
8
|
+
## 1. Open the stream
|
|
9
|
+
|
|
10
|
+
With the server running (`npx dirsql` / `uvx dirsql`), subscribe from
|
|
11
|
+
another terminal:
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
curl -N http://localhost:7117/events
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
The server immediately confirms the subscription is attached:
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
event: ready
|
|
21
|
+
data: {}
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
## 2. Change a file, receive an event
|
|
25
|
+
|
|
26
|
+
Create a file under the indexed directory:
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
echo 'second' > inbox/two.txt
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
The stream delivers the resulting row change:
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
event: row
|
|
36
|
+
data: {"action":"insert","file_path":"inbox/two.txt","old_row":null,"row":{"_basename":"two.txt","_ctime":1783170226,"_dir":"inbox","_ext":"txt","_mtime":1783170226,"_path":"inbox/two.txt","_size":7},"table":"files"}
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Edits arrive as `update` events carrying both the old and new row;
|
|
40
|
+
deletions as `delete`. The full payload schema — including `error` events —
|
|
41
|
+
is in [event payloads](../reference/http-api.md#event-payloads).
|
|
42
|
+
|
|
43
|
+
## Consuming it
|
|
44
|
+
|
|
45
|
+
Anything that speaks
|
|
46
|
+
[Server-Sent Events](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events)
|
|
47
|
+
works — `EventSource` in a browser, an SSE client library, or plain
|
|
48
|
+
`curl -N` in a shell pipeline. Two semantics worth designing around
|
|
49
|
+
([stream semantics](../reference/http-api.md#stream-semantics)):
|
|
50
|
+
|
|
51
|
+
- **Errors don't end the stream.** A malformed file produces an `error`
|
|
52
|
+
event and the stream continues — handle the event, don't reconnect.
|
|
53
|
+
- **Slow consumers skip.** A subscriber that falls behind misses the
|
|
54
|
+
overflowed events rather than stalling the server. If you must not miss
|
|
55
|
+
anything, re-query after catching up.
|
|
56
|
+
|
|
57
|
+
## Notes
|
|
58
|
+
|
|
59
|
+
- Events reflect *row* changes, not raw filesystem events: a file edit
|
|
60
|
+
that leaves a table's rows identical emits nothing, and one file change
|
|
61
|
+
can emit several events. The diffing model is part of
|
|
62
|
+
[how `dirsql` thinks](../explanation.md).
|
|
63
|
+
- Ignored files ([Skip files you don't want indexed](./skip-files.md))
|
|
64
|
+
never generate events.
|
|
65
|
+
- Embedding `dirsql` in a program? The SDK's `watch()` yields the same
|
|
66
|
+
events in-process — see the [SDK reference](../reference/sdk.md#watch)
|
|
67
|
+
and [Embed `dirsql` in your application](./embed.md).
|