dirsql 0.4.26 → 0.4.28

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -107,7 +107,9 @@ maintain nothing. `name` still has to be the row table (`posts`); the virtual
107
107
  table lives beside it under its own name. The batch runs once, when the table
108
108
  is created, and editing any part of it rebuilds a
109
109
  [persistent cache](./persist.md) from scratch. Full rules:
110
- [Batch `ddl`](../reference/config.md#batch-ddl).
110
+ [Batch `ddl`](../reference/config.md#batch-ddl); pasteable FTS5 and vector
111
+ templates, with the trigger mistakes that fail silently, are in
112
+ [Add a search index to a table](./search-indexes.md).
111
113
 
112
114
  ## Going further
113
115
 
@@ -55,7 +55,9 @@ question over a few hundred files, wasteful for a large tree you query
55
55
  repeatedly. When that day comes, [declare a table](./define-tables.md): it is
56
56
  indexed once, kept fresh by the watcher, and can persist across restarts. The
57
57
  zero-config query is the floor; a named table is the escalation path — nothing
58
- you write here has to be thrown away to get there.
58
+ you write here has to be thrown away to get there. The two sides in full, and
59
+ why indexes belong to only one of them, are in
60
+ [how `dirsql` thinks](../explanation.md#two-table-kinds-opposite-trade-offs).
59
61
 
60
62
  ## Notes
61
63
 
@@ -0,0 +1,220 @@
1
+ # Add a search index to a table
2
+
3
+ Make a declared table answer search queries instead of scanning: a B-tree for
4
+ lookups, an FTS5 index for keywords, a `vec0` index for meaning. All three are
5
+ statements in the table's [`ddl` batch](../reference/config.md#batch-ddl), so
6
+ SQLite does the work and `dirsql` maintains nothing.
7
+
8
+ The templates below are meant to be pasted and edited. Both are shaped around
9
+ one worked example — notes in `notes/*.md`, parsed into a `notes` table by an
10
+ `extract.py` that prints `{"slug": …, "title": …, "body": …}` per file, exactly
11
+ as in [Define tables for your files](./define-tables.md).
12
+
13
+ ## Why triggers, and why only two
14
+
15
+ `dirsql` writes file rows with plain `INSERT` and `DELETE`; an update is a
16
+ delete and an insert in one transaction, and there is no `UPDATE` path on user
17
+ rows. So an insert trigger and a delete trigger cover every way a row can
18
+ change — during the initial build, on each [watcher](./react-to-changes.md)
19
+ event, and when a [persistent cache](./persist.md) reconciles a tree that
20
+ moved on. There is no third trigger to write and no maintenance command to run.
21
+
22
+ The triggers are also the part that goes wrong, which is why the templates
23
+ spell them out rather than hiding them behind an abstraction.
24
+
25
+ ## Keyword search (FTS5)
26
+
27
+ FTS5 ships inside SQLite — no extension, no install. Declare an
28
+ external-content index beside the row table and let the two triggers feed it:
29
+
30
+ ```toml
31
+ [[table]]
32
+ name = "notes"
33
+ glob = "notes/**/*.md"
34
+ on-file = "python3 extract.py {path}"
35
+ ddl = '''
36
+ CREATE TABLE notes (slug TEXT, title TEXT, body TEXT);
37
+
38
+ CREATE VIRTUAL TABLE notes_fts USING fts5(
39
+ body,
40
+ content='notes',
41
+ content_rowid='rowid',
42
+ tokenize='porter unicode61'
43
+ );
44
+ CREATE TRIGGER notes_ai AFTER INSERT ON notes BEGIN
45
+ INSERT INTO notes_fts(rowid, body) VALUES (new.rowid, new.body);
46
+ END;
47
+ CREATE TRIGGER notes_ad AFTER DELETE ON notes BEGIN
48
+ INSERT INTO notes_fts(notes_fts, rowid, body)
49
+ VALUES ('delete', old.rowid, old.body);
50
+ END;
51
+ '''
52
+ ```
53
+
54
+ `content='notes'` makes the index store only its own search structures and
55
+ read column values back from `notes`, so the text is not duplicated.
56
+ `tokenize='porter unicode61'` stems English on top of the default
57
+ case-folding, so *deploying* matches *deploy*; drop the `porter` half to match
58
+ whole words only.
59
+
60
+ Query the index and join back to the row table on `rowid`:
61
+
62
+ ```bash
63
+ dirsql query "
64
+ SELECT n.slug,
65
+ bm25(notes_fts) AS score,
66
+ snippet(notes_fts, 0, '[', ']', '…', 8) AS excerpt
67
+ FROM notes_fts JOIN notes AS n ON n.rowid = notes_fts.rowid
68
+ WHERE notes_fts MATCH 'deploy'
69
+ ORDER BY score" -c ./.dirsql.toml
70
+ ```
71
+
72
+ ```json
73
+ [{"excerpt":"Rebase before you [deploy]. Keep pull requests small.","score":-1.0476190476190478e-6,"slug":"branches"},
74
+ {"excerpt":"…the spaghetti goes in. [Deploy] the garlic late.","score":-8.461538461538463e-7,"slug":"pasta"}]
75
+ ```
76
+
77
+ **`bm25()` returns negative scores, and better matches are more negative** —
78
+ so `ORDER BY score` ascending is best-first, with no `DESC`. `snippet()`'s
79
+ arguments are the table, the column index, the open and close markers, the
80
+ ellipsis, and the token budget.
81
+
82
+ ### Check the delete trigger
83
+
84
+ A wrong or missing delete trigger fails **silently**: the index keeps rows for
85
+ files that are gone, and they keep matching. Nothing errors.
86
+
87
+ The symptom is a hit that has no row behind it, which a `LEFT JOIN` exposes:
88
+
89
+ ```bash
90
+ dirsql query "
91
+ SELECT n.slug
92
+ FROM notes_fts LEFT JOIN notes AS n ON n.rowid = notes_fts.rowid
93
+ WHERE notes_fts MATCH 'frost'" -c ./.dirsql.toml --persist
94
+ ```
95
+
96
+ ```json
97
+ [{"slug":null}]
98
+ ```
99
+
100
+ A `null` slug is a stale index entry — the base row is gone and the index did
101
+ not hear about it. With the `notes_ad` trigger above, the deleted file's entry
102
+ is gone too and the query returns nothing.
103
+
104
+ Note the `--persist` flag: without it every run rebuilds the index from
105
+ scratch, so a broken delete trigger cannot show itself. Deletes are visible
106
+ against a [warm cache](./persist.md) and through the watcher, which is exactly
107
+ where the staleness would have bitten.
108
+
109
+ ## Vector search (`vec0`)
110
+
111
+ Two pieces beyond SQLite: the
112
+ [`sqlite-vec`](https://github.com/asg017/sqlite-vec) extension supplies the
113
+ `vec0` virtual table, and
114
+ [`dirsql-plugin-embeddings`](../plugins.md#dirsql-plugin-embeddings) supplies
115
+ `embed()`. Installing the plugin brings both — its config fragment declares
116
+ the extension too, and the launcher
117
+ [discovers it](../reference/cli.md#plugins):
118
+
119
+ ```bash
120
+ uvx --with dirsql-plugin-embeddings dirsql query "…" -c ./.dirsql.toml
121
+ ```
122
+
123
+ Loading `sqlite-vec` by hand instead (another runtime, a pinned build) is
124
+ [Load a SQLite extension](./load-extension.md).
125
+
126
+ ### Find your model's dimension
127
+
128
+ A `vec0` column declares a fixed width, so the template needs the number of
129
+ components your model emits. `embed()` returns the vector as JSON text, so ask
130
+ it:
131
+
132
+ ```bash
133
+ uvx --with dirsql-plugin-embeddings dirsql query \
134
+ "SELECT json_array_length(embed('probe', 'minishlab/potion-retrieval-32M')) AS dims"
135
+ ```
136
+
137
+ Run it once per model id you intend to use, and substitute the answer for the
138
+ `512` in the template below. (The first call for a model downloads it — see
139
+ [model](../plugins.md#model).) Getting it wrong is at least loud — the build fails naming both
140
+ numbers:
141
+
142
+ ```
143
+ dirsql query: failed to load config: SQLite error: Dimension mismatch for
144
+ inserted vector for the "embedding" column. Expected 8 dimensions but received 4.
145
+ ```
146
+
147
+ ### The template
148
+
149
+ ```toml
150
+ [[table]]
151
+ name = "notes"
152
+ glob = "notes/**/*.md"
153
+ on-file = "python3 extract.py {path}"
154
+ ddl = '''
155
+ CREATE TABLE notes (slug TEXT, title TEXT, body TEXT);
156
+
157
+ -- Width must equal what the probe printed for the model id named below.
158
+ CREATE VIRTUAL TABLE notes_vec USING vec0(embedding float[512]);
159
+ CREATE TRIGGER notes_vi AFTER INSERT ON notes BEGIN
160
+ INSERT INTO notes_vec(rowid, embedding)
161
+ VALUES (new.rowid, embed(new.body, 'minishlab/potion-retrieval-32M'));
162
+ END;
163
+ CREATE TRIGGER notes_vd AFTER DELETE ON notes BEGIN
164
+ DELETE FROM notes_vec WHERE rowid = old.rowid;
165
+ END;
166
+ '''
167
+ ```
168
+
169
+ The delete side is an ordinary `DELETE`, not FTS5's `'delete'` command row —
170
+ `vec0` is a normal-looking table that way.
171
+
172
+ Query it with `sqlite-vec`'s KNN form, joined back on `rowid`:
173
+
174
+ ```bash
175
+ uvx --with dirsql-plugin-embeddings dirsql query "
176
+ SELECT n.slug, v.distance
177
+ FROM notes_vec AS v JOIN notes AS n ON n.rowid = v.rowid
178
+ WHERE v.embedding MATCH embed('how do I cook pasta?', 'minishlab/potion-retrieval-32M')
179
+ AND k = 3
180
+ ORDER BY v.distance" -c ./.dirsql.toml
181
+ ```
182
+
183
+ `MATCH` plus `k = 3` is what makes this a top-k scan rather than a full one —
184
+ `vec0` uses `k`, not `LIMIT`. `distance` is supplied by the virtual table,
185
+ closest first.
186
+
187
+ **Name the same model id in the trigger and the query.** Nothing checks that
188
+ the two agree, and vectors from different models are not comparable. If the
189
+ models happen to share a width the mismatch is entirely silent — rankings just
190
+ get worse. Spelling the id out in both places, rather than leaning on the
191
+ default in either, is the cheap defence; it also puts the model in the
192
+ [config hash](../reference/config.md#batch-ddl), so changing it rebuilds the
193
+ index instead of mixing old vectors with new.
194
+
195
+ Rows are embedded once, at ingest, and cached on disk by content and model
196
+ ([vector cache](../plugins.md#vector-cache)). A search then costs one
197
+ `embed()` call for the query text plus an in-process scan.
198
+
199
+ ::: warning Editing `ddl` wedges a persisted `vec0` cache
200
+ Under `--persist`, editing any part of a `ddl` batch whose cache holds a
201
+ `vec0` table currently fails with `SQLite error: no such module: vec0`, and
202
+ keeps failing until the cache is deleted (`rm -rf <root>/.dirsql`). Tracked in
203
+ [#1008](https://github.com/thekevinscott/dirsql/issues/1008); FTS5 is
204
+ unaffected.
205
+ :::
206
+
207
+ ## What a rebuild does
208
+
209
+ `ddl` runs once, when the table is created. Editing any character of it
210
+ changes the config hash, which drops a persisted cache and re-ingests every
211
+ file — so a new index, a different tokenizer or a changed model id all rebuild
212
+ from scratch rather than leaving a half-migrated index behind. The full rules
213
+ are in [Batch `ddl`](../reference/config.md#batch-ddl).
214
+
215
+ ## Going further
216
+
217
+ - Ranked semantic search over files with no config at all —
218
+ [Search documents by meaning](./search-by-meaning.md).
219
+ - Why indexes belong to declared tables and not to path-tables —
220
+ [how `dirsql` thinks](../explanation.md).
@@ -9,9 +9,9 @@ The `dirsql` binary has these modes:
9
9
  | `dirsql query "<sql>"` | Explicit synonym for the default one-shot query. |
10
10
  | `dirsql server` | Start a long-lived HTTP server exposing a SQL view of a directory. See [HTTP API](./http-api.md). |
11
11
  | `dirsql init` | Generate a `.dirsql.toml`. |
12
+ | `dirsql` (bare) | Open a [REPL](#the-repl) over the current directory, reading statements until EOF. |
12
13
 
13
- Bare `dirsql` with no SQL is a usage error pointing at `dirsql server` — it
14
- does **not** start the server.
14
+ Bare `dirsql` does **not** start the server — that is `dirsql server`.
15
15
 
16
16
  ## Installation
17
17
 
@@ -47,6 +47,186 @@ dirsql "SELECT basename, size FROM './' ORDER BY size DESC LIMIT 5"
47
47
  pipeline, same flags, same output. See that section for config discovery,
48
48
  `--persist`, `--on-file`, hooks, and exit codes.
49
49
 
50
+ ## The REPL
51
+
52
+ `dirsql` with no subcommand and no SQL reads statements until EOF:
53
+
54
+ ```bash
55
+ dirsql
56
+ # dirsql 0.2.7 — this directory is a database.
57
+ #
58
+ # SELECT basename, size FROM './' ORDER BY size DESC LIMIT 5;
59
+ # SELECT path FROM './**/*.md' WHERE content LIKE '%TODO%';
60
+ #
61
+ # `exit`, `quit`, or Ctrl-D to leave.
62
+ #
63
+ # dirsql> SELECT count(*) AS files FROM './';
64
+ # files
65
+ # -----
66
+ # 128
67
+ #
68
+ # 1 row
69
+ # dirsql>
70
+ ```
71
+
72
+ Statements go through the same pipeline as [`dirsql query`](#dirsql-query) and
73
+ `POST /query`, so a statement typed at the prompt and one passed on the command
74
+ line return identical rows. Every config flag the default mode takes — `-c`,
75
+ `--persist`, `--no-ignore`, `--on-file` — applies unchanged:
76
+
77
+ ```bash
78
+ dirsql -c .dirsql.toml --persist
79
+ ```
80
+
81
+ The index is built **once**, before the first prompt: statements share one scan
82
+ rather than re-walking the directory each time, and the live watcher keeps it
83
+ fresh between them. Files the scan had to skip are named on stderr once, up
84
+ front.
85
+
86
+ ### Output format
87
+
88
+ Rows go where they are useful: a **table** when stdout is a terminal, the
89
+ **JSON array** when it is piped or redirected. `SELECT * FROM './'` in a
90
+ 5000-file tree should not put a 5000-element JSON array in front of a person,
91
+ and `dirsql "…" | jq` should not have to parse a table.
92
+
93
+ `--format` overrides that, in both directions, and is valid in the REPL and in
94
+ [`dirsql query`](#dirsql-query) alike:
95
+
96
+ | Value | Renders |
97
+ |---|---|
98
+ | `auto` (default) | Table if stdout is a terminal, JSON otherwise. |
99
+ | `table` | Always a table — including into a pipe or a file. |
100
+ | `json` | Always the JSON array — including at a terminal. |
101
+
102
+ ```bash
103
+ dirsql "SELECT basename, size FROM './' ORDER BY basename" --format table
104
+ # basename size
105
+ # -------- ----
106
+ # a.md 6
107
+ # bb.md 10
108
+ #
109
+ # 2 rows
110
+ ```
111
+
112
+ There is no `.mode`: dirsql has no dot-commands to extend (see
113
+ [Leaving](#leaving)), and a flag serves the one-shot query too. `dirsql server`
114
+ does not take `--format` — its transport is JSON over HTTP.
115
+
116
+ **`auto` keys on stdout, not stdin.** `dirsql > rows.json` typed at a terminal
117
+ is still headed for a file, and the file gets JSON.
118
+
119
+ Table rendering is deliberately plain: aligned columns, a rule under the
120
+ header, a row count, and `NULL` spelled out so it cannot be confused with an
121
+ empty string. Two things happen to a value on its way into a cell, both
122
+ because a `content` column holds a whole file: **newlines, tabs and other
123
+ control characters are escaped** (`\n`, `\t`, `\u{…}`) so one row cannot span
124
+ several lines, and **anything longer than 60 characters is truncated with
125
+ `…`** so one column cannot set the width of every row. `--format json`
126
+ returns the values unaltered.
127
+
128
+ Laying the table out to the terminal's width, and paging a long result, are
129
+ both out of scope; pipe to `less` for the latter.
130
+
131
+ ### Where a statement ends
132
+
133
+ At its semicolon — the same rule `sqlite3` uses, and **SQLite's own tokenizer**
134
+ decides where that semicolon is. So a statement can be laid out over as many
135
+ lines as it needs, and a `;` inside a string literal, a comment, or a
136
+ `BEGIN … END` body is not mistaken for the end of one:
137
+
138
+ ```
139
+ dirsql> SELECT basename, size
140
+ ...> FROM './'
141
+ ...> ORDER BY size DESC
142
+ ...> LIMIT 5;
143
+ ```
144
+
145
+ The `...>` prompt says the statement is still open. `exit`, `quit`, and a blank
146
+ line are not SQL, so they are taken as typed rather than waiting for a
147
+ terminator.
148
+
149
+ ### Editing and history
150
+
151
+ The prompt is a full line editor ([reedline](https://github.com/nushell/reedline)),
152
+ with the emacs bindings a shell prompt has:
153
+
154
+ | Key | Does |
155
+ |---|---|
156
+ | ↑ / ↓ | Walk back and forth through history. |
157
+ | Ctrl-R | Reverse-search history; type to narrow, Enter to accept. |
158
+ | Ctrl-A / Ctrl-E | Jump to the start / end of the line. |
159
+ | Alt-B / Alt-F | Move back / forward a word. |
160
+ | Ctrl-W, Ctrl-K, Ctrl-Y | Kill the previous word, kill to end of line, yank it back. |
161
+ | Ctrl-C | Abandon the line and return to a fresh prompt. **Does not exit.** |
162
+ | Ctrl-D | Leave. |
163
+
164
+ History is kept in one file for every directory — a query worked out in one
165
+ project is worth recalling in the next, the same way `sqlite3` keeps a single
166
+ `~/.sqlite_history`. It holds the last 1000 statements, at
167
+ `$XDG_DATA_HOME/dirsql/history`, falling back to
168
+ `~/.local/share/dirsql/history` (`%APPDATA%\dirsql\history` on Windows). If
169
+ none of those resolve, history is kept in memory for the session only.
170
+
171
+ ### Terminal vs. pipe
172
+
173
+ The prompt, banner, editor, and history exist only when **stdin is a terminal**.
174
+ From a pipe or a redirect there is none of that, and the terminator rule does
175
+ not apply either: a redirected script is not being typed, so there is no
176
+ continuation prompt to hang it off. **One statement per line, no `;` needed:**
177
+
178
+ ```bash
179
+ printf "SELECT 1 AS n\nSELECT 2 AS n\n" | dirsql
180
+ # [{"n":1}]
181
+ # [{"n":2}]
182
+
183
+ dirsql < queries.sql > rows.jsonl
184
+ ```
185
+
186
+ Blank lines do nothing in either mode.
187
+
188
+ ### Leaving
189
+
190
+ `exit`, `quit` (either case), or Ctrl-D. There are no dot-commands: the `.`
191
+ prefix exists in `sqlite3` to namespace meta-commands against SQL, and with no
192
+ meta-commands there is nothing to namespace. Schema questions are ordinary SQL:
193
+
194
+ ```sql
195
+ SELECT name FROM sqlite_master WHERE type = 'table';
196
+ ```
197
+
198
+ ### Errors
199
+
200
+ A statement that fails prints its diagnostic — the same string the HTTP
201
+ `{"error": …}` body carries — to stderr, and **the session continues**. This is
202
+ the one behavioral difference from `dirsql query`, which exits `1` on the first
203
+ failure:
204
+
205
+ ```
206
+ dirsql> SELECT nope FROM missing;
207
+ dirsql: SQLite error: no such table: missing
208
+ dirsql> SELECT 1 AS n;
209
+ n
210
+ -
211
+ 1
212
+
213
+ 1 row
214
+ ```
215
+
216
+ A config that cannot be loaded is different in kind: it fails identically for
217
+ every statement, so it is reported once and exits `1` before the first prompt.
218
+
219
+ ### Exit codes
220
+
221
+ | Code | Meaning |
222
+ |---|---|
223
+ | `0` | Clean EOF (Ctrl-D, `exit`, `quit`, or the end of a piped script) — **including when statements failed**. Matches interactive `sqlite3`; use [`dirsql query`](#dirsql-query) when a script needs a statement's exit status. |
224
+ | `1` | The index could not be built (a bad `-c`, an unresolvable `--on-file`), or stdin could not be read. Nothing was executed. |
225
+
226
+ `23` (partial scan) is not produced here: skipped files are reported before the
227
+ first prompt, and a session's exit code describes the session rather than one
228
+ scan.
229
+
50
230
  ## `dirsql server`
51
231
 
52
232
  ```bash
@@ -204,6 +384,12 @@ Errors print the same diagnostic the HTTP `{"error": …}` body carries —
204
384
  config failures, SQL errors, rejected reads, hook failures, timeouts — to
205
385
  stderr, with exit code `1`.
206
386
 
387
+ #### `--format {auto,table,json}`
388
+
389
+ How to render the result rows — the same flag [the REPL](#output-format)
390
+ takes, with the same `auto` default. A one-shot query is usually piped, so
391
+ `auto` usually means JSON; `--format table` is there for the times it is not.
392
+
207
393
  ### Exit codes
208
394
 
209
395
  | Code | Meaning |
@@ -190,7 +190,13 @@ That is the right trade for a hundreds-of-files, run-it-once question. When the
190
190
  same tree is queried repeatedly, or is large, declare a
191
191
  [table](/reference/config) for it instead — a declared table is indexed on
192
192
  build, kept fresh by the watcher, and (with `--persist`) survives restarts, so
193
- its rows are read from SQLite rather than re-walked each time.
193
+ its rows are read from SQLite rather than re-walked each time. It is also the
194
+ only place a keyword or vector index can live
195
+ ([templates](/howto/search-indexes)).
196
+
197
+ Neither kind is the better one: each buys what the other gives up, and the
198
+ choice is stated as one trade in
199
+ [how `dirsql` thinks](/explanation#two-table-kinds-opposite-trade-offs).
194
200
 
195
201
  ## Skip rules
196
202
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dirsql",
3
- "version": "0.4.26",
3
+ "version": "0.4.28",
4
4
  "description": "Ephemeral SQL index over a local directory",
5
5
  "license": "MIT",
6
6
  "repository": "https://github.com/thekevinscott/dirsql",
@@ -213,10 +213,10 @@
213
213
  ]
214
214
  },
215
215
  "optionalDependencies": {
216
- "@dirsql/lib-linux-x64-gnu": "0.4.26",
217
- "@dirsql/lib-linux-arm64-gnu": "0.4.26",
218
- "@dirsql/lib-darwin-x64": "0.4.26",
219
- "@dirsql/lib-darwin-arm64": "0.4.26",
220
- "@dirsql/lib-win32-x64-msvc": "0.4.26"
216
+ "@dirsql/lib-linux-x64-gnu": "0.4.28",
217
+ "@dirsql/lib-linux-arm64-gnu": "0.4.28",
218
+ "@dirsql/lib-darwin-x64": "0.4.28",
219
+ "@dirsql/lib-darwin-arm64": "0.4.28",
220
+ "@dirsql/lib-win32-x64-msvc": "0.4.28"
221
221
  }
222
222
  }