dirsql 0.3.124 → 0.3.126

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -89,7 +89,7 @@ Each event has `.action` (`'insert'` | `'update'` | `'delete'` | `'error'`), `.t
89
89
 
90
90
  ## CLI
91
91
 
92
- `npx dirsql` runs an HTTP server exposing the SDK over HTTP: `POST /query` for SQL and `GET /events` for a Server-Sent Events change stream. Requires **Node >= 20.11**. See the [CLI reference](https://thekevinscott.github.io/dirsql/reference/cli).
92
+ `npx dirsql "<sql>"` runs one query and prints the rows as JSON — the default. `npx dirsql server` starts an HTTP server exposing the SDK over HTTP: `POST /query` for SQL and `GET /events` for a Server-Sent Events change stream. Requires **Node >= 20.11**. See the [CLI reference](https://thekevinscott.github.io/dirsql/reference/cli).
93
93
 
94
94
  ## License
95
95
 
@@ -1,11 +1,12 @@
1
1
  # Your first dirsql database
2
2
 
3
3
  In this tutorial you will turn a directory of three tiny markdown files into
4
- a SQL database you can query over HTTP — without writing any code. You will:
4
+ a SQL database — with a single command, and without writing any code. You
5
+ will:
5
6
 
6
7
  1. Create the directory and files.
7
- 2. Start `dirsql` with zero configuration and query it with `curl`.
8
- 3. Define your own table in a `.dirsql.toml` and query the new shape.
8
+ 2. Query them straight away with zero configuration.
9
+ 3. Declare your own table to name and reuse a shape, and query it.
9
10
 
10
11
  It takes about five minutes.
11
12
 
@@ -13,9 +14,9 @@ It takes about five minutes.
13
14
  them — so it is safe to point at a real directory of your own once you are
14
15
  done here. See [Read-only by design](./explanation#read-only-by-design).
15
16
 
16
- **You need:** a terminal with `curl` and [`jq`](https://jqlang.org/), and
17
- Node ≥ 20.11 (for `npx`). Every `npx dirsql` step below also has a `uvx`
18
- tab that behaves identically, if you prefer Python tooling
17
+ **You need:** a terminal with [`jq`](https://jqlang.org/), and Node ≥ 20.11
18
+ (for `npx`). Every `npx dirsql` step below also has a `uvx` tab that behaves
19
+ identically, if you prefer Python tooling
19
20
  ([`uv`](https://docs.astral.sh/uv/)).
20
21
 
21
22
  ## 1. Create three files
@@ -59,61 +60,55 @@ notes/alice/welcome.md
59
60
  notes/bob/reading-list.md
60
61
  ```
61
62
 
62
- ## 2. Start the server
63
+ ## 2. Query your files
63
64
 
64
- From inside `my-notes`, start `dirsql`:
65
+ You wrote no configuration and no schema. Ask `dirsql` how many files are in
66
+ this directory anyway — from inside `my-notes`, run one command:
65
67
 
66
68
  ::: code-group
67
69
 
68
70
  ```bash [npm]
69
- npx dirsql
71
+ npx dirsql "SELECT COUNT(*) AS files FROM './'"
70
72
  ```
71
73
 
72
74
  ```bash [PyPI]
73
- uvx dirsql
75
+ uvx dirsql "SELECT COUNT(*) AS files FROM './'"
74
76
  ```
75
77
 
76
78
  :::
77
79
 
78
80
  The first run downloads the package (`npx` asks for confirmation — answer
79
- `y`; `uvx` prints download progress), then the server starts:
81
+ `y`; `uvx` prints download progress), then prints the result:
80
82
 
81
83
  ```
82
- Running at localhost:7117
84
+ [{"files":3}]
83
85
  ```
84
86
 
85
- That one command scanned the directory, built an ephemeral SQLite database
86
- with one row per file, and started an HTTP server. Leave it running and
87
- open a **second terminal** for the next step.
87
+ Three files, three rows. That one command scanned the directory, handed
88
+ SQLite one row per file, ran your SQL, and printed the answer as JSON.
88
89
 
89
- ## 3. Query your files
90
+ There is no named table here — you never declared one. `'./'` is a
91
+ [path-table](./reference/path-tables.md): a quoted path written where a table
92
+ name goes. `'./'` means everything under the directory you ran the command
93
+ in. The path *is* the query.
90
94
 
91
- You gave `dirsql` no configuration, so no named tables exist. Query the
92
- filesystem directly with a [path-table](./reference/path-tables.md) — a
93
- quoted path where a table name goes. `'./'` means everything under the
94
- index root. Ask it how many files there are:
95
+ ## 3. Select some columns
95
96
 
96
- ```bash
97
- curl -s http://localhost:7117/query \
98
- -H 'content-type: application/json' \
99
- -d '{"sql":"SELECT COUNT(*) AS files FROM \'./\'"}'
100
- ```
97
+ The response is always a JSON array of row objects, so from here on we pipe
98
+ it through `jq` to pretty-print. Ask for two columns instead of a count:
101
99
 
102
- ```
103
- [{"files":3}]
104
- ```
100
+ ::: code-group
105
101
 
106
- Three files, three rows. The response is always a JSON array of row
107
- objects ([HTTP API](./reference/http-api.md)) — from here on we pipe it
108
- through `jq` to pretty-print. Now select some columns:
102
+ ```bash [npm]
103
+ npx dirsql query "SELECT path, size FROM './' ORDER BY path" | jq
104
+ ```
109
105
 
110
- ```bash
111
- curl -s http://localhost:7117/query \
112
- -H 'content-type: application/json' \
113
- -d '{"sql":"SELECT path, size FROM \'./\' ORDER BY path"}' \
114
- | jq
106
+ ```bash [PyPI]
107
+ uvx dirsql query "SELECT path, size FROM './' ORDER BY path" | jq
115
108
  ```
116
109
 
110
+ :::
111
+
117
112
  ```json
118
113
  [
119
114
  {
@@ -131,126 +126,123 @@ curl -s http://localhost:7117/query \
131
126
  ]
132
127
  ```
133
128
 
134
- `path` and `size` are two of the built-in file columns `dirsql` collects
135
- for every file — see [stat columns](./reference/columns.md#stat-columns)
136
- for the full list. (The `size` values are byte counts; they match the
137
- output above because you pasted the files exactly.)
129
+ `path` and `size` are two of the built-in file columns `dirsql` collects for
130
+ every file — see [stat columns](./reference/columns.md#stat-columns) for the
131
+ full list. (The `size` values are byte counts; they match the output above
132
+ because you pasted the files exactly.)
138
133
 
139
- You have a working SQL database over your files. Next, teach it the
140
- structure your folders already encode.
134
+ You have a working SQL database over your files, and you never left the
135
+ command line.
141
136
 
142
- ## 4. Define a table
137
+ ## 4. Declare a table
143
138
 
144
- Look at the paths again: `notes/alice/ideas.md`, `notes/bob/reading-list.md`
145
- — the author's name is a directory segment. A config file can capture it as
146
- a real column.
139
+ A declared table fixes a shape once: you give it a name, scope it to exactly
140
+ the files you care about, and then query it by name instead of repeating a
141
+ path in every question. It is also the on-ramp to everything a path-table
142
+ can't do — a named table can be kept live by the watcher, persisted across
143
+ restarts, and given a parser that reads inside your files.
147
144
 
148
- In your second terminal, still inside `my-notes`, create a `.dirsql.toml`:
145
+ Still inside `my-notes`, create a `.dirsql.toml`:
149
146
 
150
147
  ```bash
151
148
  cat > .dirsql.toml <<'EOF'
152
149
  [[table]]
153
- ddl = "CREATE TABLE notes (author TEXT, basename TEXT, size INTEGER)"
154
- glob = "notes/{author}/*.md"
150
+ ddl = "CREATE TABLE notes (dir TEXT, basename TEXT, size INTEGER)"
151
+ glob = "notes/**/*.md"
155
152
  EOF
156
153
  ```
157
154
 
158
155
  Two keys define the table:
159
156
 
160
- - `glob` selects which files feed the table, and `{author}` is a
161
- [glob capture](./reference/columns.md#glob-captures): whatever directory
162
- name matches that segment becomes the row's `author` value.
157
+ - `glob` selects which files feed the table — every `.md` at any depth under
158
+ `notes/`, relative to the directory the config sits in.
163
159
  - `ddl` is ordinary `CREATE TABLE` SQL naming the columns you want to keep.
160
+ Each is a [stat column](./reference/columns.md#stat-columns) `dirsql`
161
+ computes for every file.
164
162
 
165
- ## 5. Restart and query the new shape
163
+ ## 5. Query the table
166
164
 
167
- Config is read at startup, so go back to the **first terminal**, stop the
168
- server with `Ctrl-C`, and start it again — this time pointing `dirsql` at
169
- your config with `-c` (`dirsql` does not auto-load a `.dirsql.toml` from the
170
- current directory; you always pass it explicitly):
165
+ `dirsql` does not auto-load a `.dirsql.toml` from the current directory, so
166
+ pass it explicitly with `-c`, **after** the SQL:
171
167
 
172
168
  ::: code-group
173
169
 
174
170
  ```bash [npm]
175
- npx dirsql -c .dirsql.toml
171
+ npx dirsql "SELECT dir, basename, size FROM notes ORDER BY dir, basename" -c .dirsql.toml | jq
176
172
  ```
177
173
 
178
174
  ```bash [PyPI]
179
- uvx dirsql -c .dirsql.toml
175
+ uvx dirsql "SELECT dir, basename, size FROM notes ORDER BY dir, basename" -c .dirsql.toml | jq
180
176
  ```
181
177
 
182
178
  :::
183
179
 
184
- ```
185
- Running at localhost:7117
186
- ```
187
-
188
- This time `dirsql` loaded your `.dirsql.toml` and served the `notes` table
189
- you defined. Query it from the second terminal:
190
-
191
- ```bash
192
- curl -s http://localhost:7117/query \
193
- -H 'content-type: application/json' \
194
- -d '{"sql":"SELECT author, basename, size FROM notes ORDER BY author, basename"}' \
195
- | jq
196
- ```
197
-
198
180
  ```json
199
181
  [
200
182
  {
201
183
  "basename": "ideas.md",
202
- "size": 52,
203
- "author": "alice"
184
+ "dir": "notes/alice",
185
+ "size": 52
204
186
  },
205
187
  {
206
188
  "basename": "welcome.md",
207
- "size": 66,
208
- "author": "alice"
189
+ "dir": "notes/alice",
190
+ "size": 66
209
191
  },
210
192
  {
211
193
  "basename": "reading-list.md",
212
- "size": 41,
213
- "author": "bob"
194
+ "dir": "notes/bob",
195
+ "size": 41
214
196
  }
215
197
  ]
216
198
  ```
217
199
 
218
- Every row now carries an `author` column extracted from its path — no
219
- extraction code, just a glob. And it is a real SQL column, so you can
220
- aggregate on it:
200
+ You queried `FROM notes` by name — no path, no glob to repeat. And because
201
+ `dir` is a real SQL column, you can aggregate on it. Count each author's
202
+ notes by their folder:
221
203
 
222
- ```bash
223
- curl -s http://localhost:7117/query \
224
- -H 'content-type: application/json' \
225
- -d '{"sql":"SELECT author, COUNT(*) AS notes FROM notes GROUP BY author"}' \
226
- | jq
204
+ ::: code-group
205
+
206
+ ```bash [npm]
207
+ npx dirsql query "SELECT dir, COUNT(*) AS notes FROM notes GROUP BY dir ORDER BY dir" -c .dirsql.toml | jq
208
+ ```
209
+
210
+ ```bash [PyPI]
211
+ uvx dirsql query "SELECT dir, COUNT(*) AS notes FROM notes GROUP BY dir ORDER BY dir" -c .dirsql.toml | jq
227
212
  ```
228
213
 
214
+ :::
215
+
229
216
  ```json
230
217
  [
231
218
  {
232
- "author": "alice",
219
+ "dir": "notes/alice",
233
220
  "notes": 2
234
221
  },
235
222
  {
236
- "author": "bob",
223
+ "dir": "notes/bob",
237
224
  "notes": 1
238
225
  }
239
226
  ]
240
227
  ```
241
228
 
242
- That's the whole loop: files in a directory, a declarative table on top,
243
- SQL over HTTP.
229
+ That's the whole loop: files in a directory, an instant query with no
230
+ configuration, and a declared table when you want a named shape to reuse.
244
231
 
245
232
  ## Where to go next
246
233
 
247
- - [Configuration file](./reference/config.md) — the complete `.dirsql.toml`
248
- reference: more tables, ignore patterns, persistence, hooks.
249
- - [CLI](./reference/cli.md) — flags like `--port` and `--config`, plus
250
- `dirsql init`.
251
- - [HTTP API](./reference/http-api.md) — `POST /query` in full, plus
252
- `GET /events`, a live stream of row changes as files change.
234
+ - [Query files without a config](./howto/query-without-config.md) — more
235
+ path-table questions you can ask with no setup at all.
236
+ - [Define tables for your files](./howto/define-tables.md) — the full
237
+ `[[table]]` recipe: multiple tables, ignore patterns.
238
+ - [Extract rows from file contents](./howto/extract-from-contents.md) —
239
+ pull columns out of *inside* your files with an `on-file` parser.
240
+ - [CLI](./reference/cli.md) — every flag, plus running `dirsql` as a
241
+ long-lived HTTP server instead of one-shot queries.
242
+ - [HTTP API](./reference/http-api.md) — `POST /query`, plus `GET /events`,
243
+ a live stream of row changes as files change.
253
244
  - [SDK](./reference/sdk.md) — embed `dirsql` in a Python, Rust, or
254
- TypeScript program instead of running the server.
255
- - Why is the database rebuilt from your files on every startup? See
245
+ TypeScript program instead of running the CLI.
246
+ - Why is the database rebuilt from your files on every query? See
256
247
  [how `dirsql` thinks](./explanation.md).
248
+ ```
@@ -53,7 +53,7 @@ entrypoint = "sqlite3_vec_init"
53
53
  ```
54
54
 
55
55
  ```bash
56
- uvx --with sqlite-vec dirsql -c ./.dirsql.toml
56
+ uvx --with sqlite-vec dirsql server -c ./.dirsql.toml
57
57
  ```
58
58
 
59
59
  **Node (`npx dirsql`, TypeScript SDK).** Use the *npm package name* — but
@@ -0,0 +1,122 @@
1
+ # Parse your files into columns
2
+
3
+ A [path-table](../reference/path-tables.md) gives you one row per file, with the
4
+ stat columns (`path`, `size`, `mtime`, …). When the columns you actually want
5
+ live *inside* each file, attach a parser with
6
+ [`--on-file`](../reference/cli.md#on-file-command): prototype it inline on the
7
+ path-table, then paste the **same** command into a config file the day you want
8
+ a watcher, persistence, or more than one table. The parser command never
9
+ changes across that move — that is the whole point.
10
+
11
+ ## 1. Start from a bare path-table
12
+
13
+ You have a directory of Markdown posts with YAML frontmatter:
14
+
15
+ ```
16
+ posts/hello-world.md
17
+ posts/on-recursion.md
18
+ ```
19
+
20
+ A path-table already answers questions about the files themselves:
21
+
22
+ ```bash
23
+ dirsql query "SELECT basename, size FROM './posts/*.md' ORDER BY basename"
24
+ ```
25
+
26
+ ```json
27
+ [{"basename":"hello-world.md","size":65},{"basename":"on-recursion.md","size":102}]
28
+ ```
29
+
30
+ But `title` and `author` are inside the frontmatter, not in the stat columns.
31
+ To get them, you need a parser.
32
+
33
+ ## 2. Attach a parser with `--on-file`
34
+
35
+ Any program that reads one file and prints a **JSON array of row objects** on
36
+ stdout is a parser. Here is a small one, `extract.py`, that reads a post's
37
+ frontmatter:
38
+
39
+ ```python
40
+ #!/usr/bin/env python3
41
+ import json, re, sys
42
+
43
+ text = open(sys.argv[1], encoding="utf-8").read()
44
+ m = re.match(r"^---\n(.*?)\n---", text, re.DOTALL)
45
+ fields = dict(
46
+ (k.strip(), v.strip())
47
+ for k, _, v in (line.partition(":") for line in (m.group(1).splitlines() if m else []))
48
+ )
49
+ print(json.dumps([{"title": fields.get("title"), "author": fields.get("author")}]))
50
+ ```
51
+
52
+ Point the path-table at it with `--on-file`:
53
+
54
+ ```bash
55
+ dirsql query "SELECT title, author FROM './posts/*.md' ORDER BY title" \
56
+ --on-file 'python3 extract.py {path}'
57
+ ```
58
+
59
+ ```json
60
+ [{"author":"Ada Lovelace","title":"Hello World"},{"author":"Alan Turing","title":"On Recursion"}]
61
+ ```
62
+
63
+ Now the parser's output *is* the table. `--on-file` runs the command once per
64
+ matched file; `{path}` is the file's absolute path, one of the placeholders in
65
+ the shared [`on-file` hook contract](../reference/hooks.md#on-file) (argv
66
+ splitting, timeout, and per-file failure isolation all come from there). The
67
+ stat columns are no longer reachable — a parser that wants the path emits it,
68
+ since it already has `{path}`. See
69
+ [Parsing rows with `--on-file`](../reference/path-tables.md#parsing-rows-with-on-file)
70
+ for the full behavior.
71
+
72
+ `--on-file` applies to every path-table in the query and may be given at most
73
+ once. It is a `query`-only flag: there is no config file involved yet, so it is
74
+ the fastest way to see whether a parser produces the rows you expect.
75
+
76
+ ## 3. Graduate to a config file
77
+
78
+ The inline flag re-scans and re-parses every file on every query, defines one
79
+ parser for the whole query, and forgets everything when the process exits. When
80
+ you want a **watcher** that keeps the rows fresh, **persistence** across
81
+ restarts, or **different parsers for different file sets**, move the parser into
82
+ a `.dirsql.toml` [`[[table]]`](../reference/config.md#table) — and paste the
83
+ command in verbatim:
84
+
85
+ ```toml
86
+ [[table]]
87
+ ddl = "CREATE TABLE posts (title TEXT, author TEXT)"
88
+ glob = "posts/*.md"
89
+ on-file = "python3 extract.py {path}"
90
+ ```
91
+
92
+ The `on-file` value is byte-for-byte the string you passed to `--on-file`. The
93
+ [`on-file` hook contract](../reference/hooks.md#on-file) is the same contract
94
+ the flag used — the flag and the config key are two spellings of the same
95
+ attachment. What you gain by moving to config is everything around the parser:
96
+
97
+ ```bash
98
+ dirsql query "SELECT title, author FROM posts ORDER BY title" -c ./.dirsql.toml
99
+ ```
100
+
101
+ ```json
102
+ [{"author":"Ada Lovelace","title":"Hello World"},{"author":"Alan Turing","title":"On Recursion"}]
103
+ ```
104
+
105
+ Same command, same rows. The difference is what a declared table brings: it is
106
+ indexed on build, kept fresh by the watcher, survives restarts with
107
+ [`--persist`](./persist.md), and a config can declare
108
+ [many tables](./define-tables.md) — each with its own `on-file` — where the flag
109
+ gives every path-table one parser. A declared table also merges the filesystem
110
+ facts back onto each row, so a `path TEXT` column in the DDL is populated by
111
+ `dirsql` even though the parser did not emit it
112
+ ([precedence](../reference/columns.md#precedence)).
113
+
114
+ ## Going further
115
+
116
+ - The full stat-vs-parsed behavior, failure isolation, and skip rules:
117
+ [Parsing rows with `--on-file`](../reference/path-tables.md#parsing-rows-with-on-file).
118
+ - Starting from a declared table instead of a path-table?
119
+ [Extract rows from file contents](./extract-from-contents.md) covers the
120
+ config-first path.
121
+ - One row per record *within* a file (JSONL, multiple frontmatter blocks): the
122
+ parser prints one object per row — the array length is the row count.
@@ -1,7 +1,7 @@
1
1
  # Keep the index across restarts
2
2
 
3
3
  By default the database is ephemeral: rebuilt from your files on every
4
- startup and discarded on exit. The [`--persist [PATH]`](../reference/cli.md#server-mode)
4
+ startup and discarded on exit. The [`--persist [PATH]`](../reference/cli.md#dirsql-server)
5
5
  flag keeps the SQLite index on disk instead, so a restart only re-parses
6
6
  files that actually changed — the difference between seconds and
7
7
  milliseconds on large trees, and between re-running and skipping expensive
@@ -18,8 +18,8 @@ glob = "**/*"
18
18
 
19
19
  ## 1. Open the stream
20
20
 
21
- With the server running (`npx dirsql -c ./.dirsql.toml` /
22
- `uvx dirsql -c ./.dirsql.toml`), subscribe from another terminal:
21
+ With the server running (`npx dirsql server -c ./.dirsql.toml` /
22
+ `uvx dirsql server -c ./.dirsql.toml`), subscribe from another terminal:
23
23
 
24
24
  ```bash
25
25
  curl -N http://localhost:7117/events
@@ -10,7 +10,7 @@ embed each question at query time.
10
10
  ::: tip Just want it working?
11
11
  [`dirsql-plugin-embeddings`](https://pypi.org/project/dirsql-plugin-embeddings/)
12
12
  packages exactly what this guide builds, ready to install:
13
- `uvx --with dirsql-plugin-embeddings dirsql`. Keep reading to see how it's
13
+ `uvx --with dirsql-plugin-embeddings dirsql server`. Keep reading to see how it's
14
14
  built — the same three pieces, from scratch.
15
15
  :::
16
16
 
package/docs/index.md CHANGED
@@ -34,7 +34,7 @@ db = DirSQL(
34
34
  "./my-project",
35
35
  tables=[
36
36
  Table(
37
- ddl="CREATE TABLE files (name TEXT, size INTEGER, type TEXT)",
37
+ ddl="CREATE TABLE records (name TEXT, size INTEGER, type TEXT)",
38
38
  glob="data/*.json",
39
39
  on_file=lambda path: [json.loads(open(path, encoding="utf-8").read())],
40
40
  ),
@@ -42,7 +42,7 @@ db = DirSQL(
42
42
  )
43
43
 
44
44
  # SQL queries over your filesystem
45
- large = db.query("SELECT * FROM files WHERE size > 1000")
45
+ large = db.query("SELECT * FROM records WHERE size > 1000")
46
46
  ```
47
47
 
48
48
  ```rust [Rust]
@@ -52,14 +52,14 @@ let db = DirSQL::new(
52
52
  "./my-project",
53
53
  vec![
54
54
  Table::new(
55
- "CREATE TABLE files (name TEXT, size INTEGER, type TEXT)",
55
+ "CREATE TABLE records (name TEXT, size INTEGER, type TEXT)",
56
56
  "data/*.json",
57
57
  |path| vec![serde_json::from_str(&std::fs::read_to_string(path).unwrap()).unwrap()],
58
58
  ),
59
59
  ],
60
60
  )?;
61
61
 
62
- let large = db.query("SELECT * FROM files WHERE size > 1000")?;
62
+ let large = db.query("SELECT * FROM records WHERE size > 1000")?;
63
63
  ```
64
64
 
65
65
  ```typescript [TypeScript]
@@ -70,14 +70,14 @@ const db = new DirSQL({
70
70
  root: './my-project',
71
71
  tables: [
72
72
  new Table({
73
- ddl: 'CREATE TABLE files (name TEXT, size INTEGER, type TEXT)',
73
+ ddl: 'CREATE TABLE records (name TEXT, size INTEGER, type TEXT)',
74
74
  glob: 'data/*.json',
75
75
  onFile: (path) => [JSON.parse(readFileSync(path, 'utf8'))],
76
76
  }),
77
77
  ],
78
78
  });
79
79
 
80
- const large = await db.query('SELECT * FROM files WHERE size > 1000');
80
+ const large = await db.query('SELECT * FROM records WHERE size > 1000');
81
81
  ```
82
82
 
83
83
  :::
@@ -1,39 +1,56 @@
1
1
  # CLI
2
2
 
3
- The `dirsql` binary has three modes:
3
+ Query is the default: `dirsql "<sql>"` runs one query and prints JSON rows.
4
+ The `dirsql` binary has these modes:
4
5
 
5
6
  | Invocation | Behavior |
6
7
  |---|---|
7
- | `dirsql` (no subcommand) | Start a long-lived HTTP server exposing a SQL view of a directory. See [HTTP API](./http-api.md). |
8
- | `dirsql query "<sql>"` | Query the file system as JSON |
9
- | `dirsql init` | Generate a `.dirsql.toml` |
8
+ | `dirsql "<sql>"` | Run one query over the directory and print the rows as JSON. The default; identical to `dirsql query "<sql>"`. |
9
+ | `dirsql query "<sql>"` | Explicit synonym for the default one-shot query. |
10
+ | `dirsql server` | Start a long-lived HTTP server exposing a SQL view of a directory. See [HTTP API](./http-api.md). |
11
+ | `dirsql init` | Generate a `.dirsql.toml`. |
12
+
13
+ Bare `dirsql` with no SQL is a usage error pointing at `dirsql server` — it
14
+ does **not** start the server.
10
15
 
11
16
  ## Installation
12
17
 
13
18
  ::: code-group
14
19
 
15
20
  ```bash [npm]
16
- npx dirsql
21
+ npx dirsql "SELECT * FROM './'"
17
22
  ```
18
23
 
19
24
  ```bash [PyPI]
20
- uvx dirsql
25
+ uvx dirsql "SELECT * FROM './'"
21
26
  ```
22
27
 
23
28
  ```bash [Cargo]
24
29
  # The `cli` feature is opt-in; this installs the binary only.
25
30
  cargo install dirsql --features cli
26
- dirsql
31
+ dirsql "SELECT * FROM './'"
27
32
  ```
28
33
 
29
34
  :::
30
35
 
31
36
  The npm launcher requires **Node ≥ 20.11**.
32
37
 
33
- ## Server mode
38
+ ## Default query mode
34
39
 
35
40
  ```bash
36
- dirsql
41
+ # No -c: query the filesystem with a path-table.
42
+ dirsql "SELECT basename, size FROM './' ORDER BY size DESC LIMIT 5"
43
+ # [{"basename":"model.bin","size":104857600}, …]
44
+ ```
45
+
46
+ `dirsql "<sql>"` is exactly [`dirsql query "<sql>"`](#dirsql-query) — same
47
+ pipeline, same flags, same output. See that section for config discovery,
48
+ `--persist`, `--on-file`, hooks, and exit codes.
49
+
50
+ ## `dirsql server`
51
+
52
+ ```bash
53
+ dirsql server
37
54
  # Running at localhost:7117
38
55
  ```
39
56
 
@@ -43,12 +60,15 @@ requests, closes open `/events` streams, and exits.
43
60
 
44
61
  ### Flags
45
62
 
63
+ Config flags are subcommand-local: pass them after `server`
64
+ (`dirsql server -c <cfg>`).
65
+
46
66
  | Flag | Default | Description |
47
67
  |---|---|---|
48
- | `-c, --config <path>` | none | Path to a [config file](./config.md). **Repeatable** (`-c a -c b`): the configs load and merge in argv order — see [Composing multiple configs](./config.md#composing-multiple-configs). The index is always rooted at the **invocation directory** (the current working directory), regardless of where a config lives — so `--config /elsewhere/.dirsql.toml` still indexes the directory you ran `dirsql` from. With none given, **no named tables are defined** — query the filesystem with a [path-table](./path-tables.md) (`FROM './'`). A `./.dirsql.toml` on disk is **not** auto-loaded; pass it explicitly. A `-c` naming a file that does not exist is an [error](#degraded-mode). |
68
+ | `-c, --config <path>` | none | Path to a [config file](./config.md). **Repeatable** (`-c a -c b`): the configs load and merge in argv order — see [Composing multiple configs](./config.md#composing-multiple-configs). The index is always rooted at the **invocation directory** (the current working directory), regardless of where a config lives — so `--config /elsewhere/.dirsql.toml` still indexes the directory you ran `dirsql server` from. With none given, **no named tables are defined** — query the filesystem with a [path-table](./path-tables.md) (`FROM './'`). A `./.dirsql.toml` on disk is **not** auto-loaded; pass it explicitly. A `-c` naming a file that does not exist is an [error](#degraded-mode). |
49
69
  | `--host <addr>` | `localhost` | Bind address. |
50
70
  | `--port <n>` | `7117` | TCP port to bind. |
51
- | `--persist [<path>]` | off | Keep the SQLite index on disk between runs so a restart only re-parses files that actually changed. Bare `--persist` caches at `<root>/.dirsql/cache.db`; `--persist <path>` caches at `<path>`. Off by default (the index is ephemeral). Also available on [`dirsql query`](#dirsql-query), passed after the subcommand. See [Keep the index across restarts](../howto/persist.md). |
71
+ | `--persist [<path>]` | off | Keep the SQLite index on disk between runs so a restart only re-parses files that actually changed. Bare `--persist` caches at `<root>/.dirsql/cache.db`; `--persist <path>` caches at `<path>`. Off by default (the index is ephemeral). Also available on [`dirsql query`](#dirsql-query). See [Keep the index across restarts](../howto/persist.md). |
52
72
  | `--extension <path>` | none | Load a SQLite extension by literal path, overriding the config's `[[dirsql.extension]]` entries. Repeatable. Format: `<path>` or `<path>::<entrypoint>`. Internal plumbing for the pip/npm launchers, which resolve package-name extensions and pass the resolved paths here — not intended for direct use. When any `--extension` is present, the config file's own extension entries are not loaded. |
53
73
  | `--version` | | Print the version and exit. |
54
74
  | `--help` | | Print usage and exit. |
@@ -124,8 +144,9 @@ dirsql query "SELECT COUNT(*) AS n FROM posts" -c ./.dirsql.toml | jq '.[0].n'
124
144
  Pass `-c`/`--config`, `--persist`, and `--extension` **after** `query`
125
145
  (`dirsql query "<sql>" -c <cfg>`). A config flag placed *before* the subcommand
126
146
  is a hard error — `error: the subcommand 'query' cannot be used with
127
- '--config <CONFIG>'` — never silently dropped. (In server mode, with no
128
- subcommand, the same flags are passed directly: `dirsql -c <cfg>`.)
147
+ '--config <CONFIG>'` — never silently dropped. (The default mode without the
148
+ `query` keyword takes the same flags after the SQL: `dirsql "<sql>" -c <cfg>`;
149
+ for the server they follow the subcommand: `dirsql server -c <cfg>`.)
129
150
  :::
130
151
 
131
152
  The subcommand builds the index, runs the SQL, prints the result rows as a
@@ -168,7 +189,9 @@ timeout. The parser's output is the whole schema; the stat columns are not
168
189
  reachable on a parsed path-table. `--on-file` may be given **at most once** (a
169
190
  repeat is an error pointing at config files) and never touches config-declared
170
191
  tables. It is a `query`-only flag — server mode rejects it as an unknown
171
- argument.
192
+ argument. The command string is copy-paste identical to a `[[table]]`
193
+ `on-file` key, so an inline parser graduates to a config file unchanged — see
194
+ [Parse your files into columns](../howto/parse-files-into-columns.md).
172
195
 
173
196
  Errors print the same diagnostic the HTTP `{"error": …}` body carries —
174
197
  config failures, SQL errors, rejected reads, hook failures, timeouts — to
@@ -197,7 +220,7 @@ The output does **not** auto-load. Once you've tweaked it, pass it explicitly
197
220
  to run against it:
198
221
 
199
222
  ```bash
200
- dirsql -c ./.dirsql.toml
223
+ dirsql "SELECT * FROM files" -c ./.dirsql.toml
201
224
  ```
202
225
 
203
226
  ### Flags
@@ -40,7 +40,7 @@ hook-timeout = 300
40
40
  ```
41
41
 
42
42
  Persistence is not a config key. Keep the SQLite index on disk between runs
43
- with the [`--persist [PATH]` CLI flag](./cli.md#server-mode) — a machine-local
43
+ with the [`--persist [PATH]` CLI flag](./cli.md#dirsql-server) — a machine-local
44
44
  operational choice that belongs to the runner, not to shareable config.
45
45
 
46
46
  ## `[[dirsql.extension]]`
@@ -99,6 +99,17 @@ every watched change. The command reads the file itself and prints a JSON
99
99
  array of row objects; see [`[[table]]`](./config.md#table) for the
100
100
  row-mapping rules.
101
101
 
102
+ The same command is attachable two ways, over the same contract: the
103
+ `on-file` config key on a `[[table]]`, and the
104
+ [`--on-file <command>`](./cli.md#on-file-command) flag on `dirsql query`,
105
+ which attaches it to every [path-table](./path-tables.md#parsing-rows-with-on-file)
106
+ in the query. The command string is identical between the two spellings — the
107
+ flag is the inline form, the config key the declared form (see
108
+ [Parse your files into columns](../howto/parse-files-into-columns.md)). The one
109
+ behavioral difference is not in this contract but in the surrounding table: a
110
+ declared `[[table]]` merges the filesystem facts onto each parsed row, while a
111
+ parsed path-table exposes only the parser's output.
112
+
102
113
  | Placeholder | Value |
103
114
  |---|---|
104
115
  | `{path}` | The matched file's **absolute** path. `on-file = "extract.py {path}"` — self-sufficient from any working directory, so the command resolves it even when the config lives outside the index. |
@@ -1,6 +1,6 @@
1
1
  # HTTP API
2
2
 
3
- The [`dirsql` server](./cli.md#server-mode) (default `localhost:7117`)
3
+ The [`dirsql` server](./cli.md#dirsql-server) (default `localhost:7117`)
4
4
  exposes two endpoints: `POST /query` and `GET /events`.
5
5
 
6
6
  ## `POST /query`
@@ -27,7 +27,7 @@ The `200` response is a JSON array of row objects keyed by column name:
27
27
  ```bash
28
28
  curl -s http://localhost:7117/query \
29
29
  -H 'content-type: application/json' \
30
- -d '{"sql":"SELECT COUNT(*) AS n FROM files"}' \
30
+ -d '{"sql":"SELECT COUNT(*) AS n FROM \'./\'"}' \
31
31
  | jq
32
32
  ```
33
33
 
@@ -198,9 +198,9 @@ pattern if you would rather not see them.
198
198
  Path-tables are ordinary SQLite tables once resolved, so they join freely:
199
199
 
200
200
  ```sql
201
- SELECT p.basename, f.size
201
+ SELECT p.basename, d.size
202
202
  FROM './docs/*.md' AS p
203
- JOIN files AS f ON f.path = p.path;
203
+ JOIN pages AS d ON d.path = p.path;
204
204
  ```
205
205
 
206
206
  A zero-match path-table is not an error — it is an empty table, and the query
@@ -130,7 +130,7 @@ shortcut was removed in #603 — use
130
130
  ephemeral, rebuilt every startup). The cache lives at
131
131
  `<root>/.dirsql/cache.db` by default; on restart, only files whose stat
132
132
  changed are re-parsed. (The CLI exposes the same switch as the
133
- [`--persist [PATH]`](./cli.md#server-mode) flag; it is not a config key.)
133
+ [`--persist [PATH]`](./cli.md#dirsql-server) flag; it is not a config key.)
134
134
  - `persist_path` / `persistPath` — Override the cache location. Ignored when
135
135
  persistence is off. Constructor values are used as given. In Rust these two
136
136
  parameters collapse into a single builder method: `.persist(None)` enables
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dirsql",
3
- "version": "0.3.124",
3
+ "version": "0.3.126",
4
4
  "description": "Ephemeral SQL index over a local directory",
5
5
  "license": "MIT",
6
6
  "repository": "https://github.com/thekevinscott/dirsql",
@@ -212,15 +212,15 @@
212
212
  ]
213
213
  },
214
214
  "optionalDependencies": {
215
- "@dirsql/lib-linux-x64-gnu": "0.3.124",
216
- "@dirsql/lib-linux-arm64-gnu": "0.3.124",
217
- "@dirsql/lib-darwin-x64": "0.3.124",
218
- "@dirsql/lib-darwin-arm64": "0.3.124",
219
- "@dirsql/lib-win32-x64-msvc": "0.3.124",
220
- "@dirsql/cli-linux-x64-gnu": "0.3.124",
221
- "@dirsql/cli-linux-arm64-gnu": "0.3.124",
222
- "@dirsql/cli-darwin-x64": "0.3.124",
223
- "@dirsql/cli-darwin-arm64": "0.3.124",
224
- "@dirsql/cli-win32-x64-msvc": "0.3.124"
215
+ "@dirsql/lib-linux-x64-gnu": "0.3.126",
216
+ "@dirsql/lib-linux-arm64-gnu": "0.3.126",
217
+ "@dirsql/lib-darwin-x64": "0.3.126",
218
+ "@dirsql/lib-darwin-arm64": "0.3.126",
219
+ "@dirsql/lib-win32-x64-msvc": "0.3.126",
220
+ "@dirsql/cli-linux-x64-gnu": "0.3.126",
221
+ "@dirsql/cli-linux-arm64-gnu": "0.3.126",
222
+ "@dirsql/cli-darwin-x64": "0.3.126",
223
+ "@dirsql/cli-darwin-arm64": "0.3.126",
224
+ "@dirsql/cli-win32-x64-msvc": "0.3.126"
225
225
  }
226
226
  }