dirsql 0.4.65 → 0.4.67

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  Ephemeral SQL index over a local directory. `dirsql` watches a filesystem, ingests structured files into an in-memory SQLite database, and exposes a SQL query interface -- the filesystem is always the source of truth. Built on the Rust core via napi-rs bindings.
4
4
 
5
- [Documentation](https://thekevinscott.github.io/dirsql/?lang=typescript)
5
+ [Documentation](https://dirsql.dev/?lang=typescript)
6
6
 
7
7
  Also available as [`dirsql` on crates.io](https://crates.io/crates/dirsql) and [`dirsql` on PyPI](https://pypi.org/project/dirsql/).
8
8
 
@@ -89,7 +89,7 @@ Each event has `.action` (`'insert'` | `'update'` | `'delete'` | `'error'`), `.t
89
89
 
90
90
  ## CLI
91
91
 
92
- `npx dirsql "<sql>"` runs one query and prints the rows as JSON — the default. `npx dirsql server` starts an HTTP server exposing the SDK over HTTP: `POST /query` for SQL and `GET /events` for a Server-Sent Events change stream. Requires **Node >= 20.11**. See the [CLI reference](https://thekevinscott.github.io/dirsql/reference/cli).
92
+ `npx dirsql "<sql>"` runs one query and prints the rows as JSON — the default. `npx dirsql server` starts an HTTP server exposing the SDK over HTTP: `POST /query` for SQL and `GET /events` for a Server-Sent Events change stream. Requires **Node >= 20.11**. See the [CLI reference](https://dirsql.dev/reference/cli).
93
93
 
94
94
  ## License
95
95
 
@@ -156,11 +156,17 @@ folder name), and prints them as a JSON row:
156
156
  ```bash
157
157
  cat > note.sh <<'EOF'
158
158
  #!/usr/bin/env sh
159
- title=$(sed -n 's/^# //p' "$1" | head -n1)
160
- author=$(basename "$(dirname "$1")")
161
- printf '[{"title":%s,"author":%s}]' \
162
- "$(jq -Rn --arg t "$title" '$t')" \
163
- "$(jq -Rn --arg a "$author" '$a')"
159
+ printf '['
160
+ sep=''
161
+ for f; do
162
+ title=$(sed -n 's/^# //p' "$f" | head -n1)
163
+ author=$(basename "$(dirname "$f")")
164
+ printf '%s{"title":%s,"author":%s}' "$sep" \
165
+ "$(jq -Rn --arg t "$title" '$t')" \
166
+ "$(jq -Rn --arg a "$author" '$a')"
167
+ sep=','
168
+ done
169
+ printf ']'
164
170
  EOF
165
171
  ```
166
172
 
@@ -172,7 +178,7 @@ cat > .dirsql.toml <<'EOF'
172
178
  name = "notes"
173
179
  ddl = "CREATE TABLE notes (title TEXT, author TEXT)"
174
180
  glob = "notes/**/*.md"
175
- on-file = "sh note.sh {path}"
181
+ on-file = "sh note.sh"
176
182
  EOF
177
183
  ```
178
184
 
@@ -181,9 +187,10 @@ Three keys define the table:
181
187
  - `glob` selects which files feed the table — every `.md` at any depth under
182
188
  `notes/`, relative to the directory the config sits in.
183
189
  - `ddl` is ordinary `CREATE TABLE` SQL naming the columns you want to keep.
184
- - `on-file` is the command run once per matched file; `{path}` is the file's
185
- path, and its printed JSON row becomes the file's row. The columns are
186
- exactly what it emits — `title` from the heading, `author` from the folder.
190
+ - `on-file` is the command run once for the table, with every matched file's
191
+ path appended as an argument; the JSON array it prints is the table's rows.
192
+ The columns are exactly what it emits — `title` from the heading, `author`
193
+ from the folder.
187
194
 
188
195
  ## 5. Query the table
189
196
 
@@ -15,18 +15,21 @@ photos/2024/11/hike.jpg
15
15
  photos/2025/01/snow.jpg
16
16
  ```
17
17
 
18
- Write a small parser, `pathcols.py`, that turns the path into a row. It
19
- receives the file's absolute path as its argument and prints a JSON array of
20
- row objects:
18
+ Write a small parser, `pathcols.py`, that turns each path into a row. It
19
+ receives the matched files' absolute paths as its arguments and prints one
20
+ JSON array of row objects:
21
21
 
22
22
  ```python
23
23
  #!/usr/bin/env python3
24
24
  import json, os, sys
25
25
 
26
- parts = sys.argv[1].split(os.sep)
27
- # .../photos/<year>/<month>/<file>
28
- print(json.dumps([{"year": parts[-3], "month": parts[-2],
29
- "basename": os.path.basename(sys.argv[1])}]))
26
+ rows = []
27
+ for path in sys.argv[1:]:
28
+ parts = path.split(os.sep)
29
+ # .../photos/<year>/<month>/<file>
30
+ rows.append({"year": parts[-3], "month": parts[-2],
31
+ "basename": os.path.basename(path)})
32
+ print(json.dumps(rows))
30
33
  ```
31
34
 
32
35
  Point a table at it in `.dirsql.toml`:
@@ -36,12 +39,12 @@ Point a table at it in `.dirsql.toml`:
36
39
  name = "photos"
37
40
  ddl = "CREATE TABLE photos (year TEXT, month TEXT, basename TEXT)"
38
41
  glob = "photos/*/*/*.jpg"
39
- on-file = "python3 pathcols.py {path}"
42
+ on-file = "python3 pathcols.py"
40
43
  ```
41
44
 
42
- The hook emits every column the table has — dirsql injects nothing. `{path}`
43
- is the matched file's absolute path, one of the placeholders in the
44
- [command hook contract](../reference/hooks.md#on-file). A column appears only
45
+ The hook emits every column the table has — dirsql injects nothing. The
46
+ matched files' absolute paths arrive as the command's trailing arguments, per
47
+ the [command hook contract](../reference/hooks.md#on-file). A column appears only
45
48
  because the DDL declares it *and* the hook emits it.
46
49
 
47
50
  ## 2. Query the derived columns
@@ -10,16 +10,19 @@ persist across restarts, instead of repeating an ad-hoc
10
10
  Suppose your blog posts live under `posts/`, one markdown file each. A named
11
11
  table needs three keys: a `glob` that selects the files, a `ddl` that names
12
12
  the columns, and an [`on-file`](../reference/config.md#table) hook that emits
13
- each file's rows. Put a small parser next to the config — `extract.py`, which
14
- reads a post's title line and prints a JSON array of row objects:
13
+ the table's rows. Put a small parser next to the config — `extract.py`, which
14
+ reads each post's title line and prints one JSON array of row objects:
15
15
 
16
16
  ```python
17
17
  #!/usr/bin/env python3
18
18
  import json, os, sys
19
19
 
20
- text = open(sys.argv[1], encoding="utf-8").read()
21
- title = next((l[2:].strip() for l in text.splitlines() if l.startswith("# ")), None)
22
- print(json.dumps([{"title": title, "slug": os.path.basename(sys.argv[1])[:-3]}]))
20
+ rows = []
21
+ for path in sys.argv[1:]:
22
+ text = open(path, encoding="utf-8").read()
23
+ title = next((l[2:].strip() for l in text.splitlines() if l.startswith("# ")), None)
24
+ rows.append({"title": title, "slug": os.path.basename(path)[:-3]})
25
+ print(json.dumps(rows))
23
26
  ```
24
27
 
25
28
  Then declare the table in `.dirsql.toml`:
@@ -29,15 +32,15 @@ Then declare the table in `.dirsql.toml`:
29
32
  name = "posts"
30
33
  ddl = "CREATE TABLE posts (title TEXT, slug TEXT)"
31
34
  glob = "posts/**/*.md"
32
- on-file = "python3 extract.py {path}"
35
+ on-file = "python3 extract.py"
33
36
  ```
34
37
 
35
38
  - `glob` selects the files: every `.md` under `posts/`, at any depth, relative
36
39
  to the directory containing the config.
37
40
  - `ddl` is a plain SQLite `CREATE TABLE` naming the columns you want to keep.
38
41
  - `on-file` is **required** — it is where the table's rows come from. dirsql
39
- injects nothing; the hook emits every column, reading the file (it has
40
- `{path}`) and deriving whatever it needs. A `[[table]]` with no `on-file` is
42
+ injects nothing; the hook emits every column, reading the files (their
43
+ paths are its arguments) and deriving whatever it needs. A `[[table]]` with no `on-file` is
41
44
  a [config error](../reference/config.md#parse-errors). For plain stat
42
45
  columns with no code, query the path directly with a path-table instead.
43
46
 
@@ -75,7 +78,7 @@ two triggers:
75
78
  [[table]]
76
79
  name = "posts"
77
80
  glob = "posts/**/*.md"
78
- on-file = "python3 extract.py {path}"
81
+ on-file = "python3 extract.py"
79
82
  ddl = '''
80
83
  CREATE TABLE posts (title TEXT, slug TEXT, body TEXT);
81
84
  CREATE INDEX posts_slug ON posts(slug);
@@ -21,11 +21,11 @@ stdout works. With [`jq`](https://jqlang.org/):
21
21
  name = "books"
22
22
  ddl = "CREATE TABLE books (title TEXT, author TEXT, year INTEGER)"
23
23
  glob = "books/*.json"
24
- on-file = "jq -c '[{title, author, year}]' {path}"
24
+ on-file = "jq -c -n '[inputs | {title, author, year}]'"
25
25
  ```
26
26
 
27
- `{path}` is the matched file's absolute path — one of the
28
- placeholders defined by the
27
+ The command runs once for the table, with every matched file's absolute path
28
+ appended as an argument, and prints one array for all of them — the
29
29
  [command hook contract](../reference/hooks.md#on-file), which also covers
30
30
  the argv splitting, working directory, stdout protocol, and timeout shared
31
31
  by every hook.
@@ -45,7 +45,7 @@ dirsql query "SELECT title, author, year FROM books ORDER BY year" -c ./.dirsql.
45
45
 
46
46
  The table's columns are exactly what the command emits, narrowed to the DDL —
47
47
  `dirsql` adds nothing. To include the file's `path`, have the command emit it
48
- (it has `{path}`); dirsql will not merge it in for you.
48
+ (it has the path); dirsql will not merge it in for you.
49
49
 
50
50
  ## Multiple rows per file
51
51
 
@@ -57,7 +57,7 @@ row per line, slurp it:
57
57
  name = "events"
58
58
  ddl = "CREATE TABLE events (event TEXT, user TEXT)"
59
59
  glob = "logs/*.jsonl"
60
- on-file = "jq -c -s '.' {path}"
60
+ on-file = "jq -c -s '.'"
61
61
  ```
62
62
 
63
63
  ```bash
@@ -32,40 +32,43 @@ To get them, you need a parser.
32
32
 
33
33
  ## 2. Attach a parser with `--on-file`
34
34
 
35
- Any program that reads one file and prints a **JSON array of row objects** on
36
- stdout is a parser. Here is a small one, `extract.py`, that reads a post's
37
- frontmatter:
35
+ Any program that reads the files named by its arguments and prints a **JSON
36
+ array of row objects** on stdout is a parser. Here is a small one,
37
+ `extract.py`, that reads each post's frontmatter:
38
38
 
39
39
  ```python
40
40
  #!/usr/bin/env python3
41
41
  import json, re, sys
42
42
 
43
- text = open(sys.argv[1], encoding="utf-8").read()
44
- m = re.match(r"^---\n(.*?)\n---", text, re.DOTALL)
45
- fields = dict(
46
- (k.strip(), v.strip())
47
- for k, _, v in (line.partition(":") for line in (m.group(1).splitlines() if m else []))
48
- )
49
- print(json.dumps([{"title": fields.get("title"), "author": fields.get("author")}]))
43
+ rows = []
44
+ for path in sys.argv[1:]:
45
+ text = open(path, encoding="utf-8").read()
46
+ m = re.match(r"^---\n(.*?)\n---", text, re.DOTALL)
47
+ fields = dict(
48
+ (k.strip(), v.strip())
49
+ for k, _, v in (line.partition(":") for line in (m.group(1).splitlines() if m else []))
50
+ )
51
+ rows.append({"title": fields.get("title"), "author": fields.get("author")})
52
+ print(json.dumps(rows))
50
53
  ```
51
54
 
52
55
  Point the path-table at it with `--on-file`:
53
56
 
54
57
  ```bash
55
58
  dirsql query "SELECT title, author FROM './posts/*.md' ORDER BY title" \
56
- --on-file 'python3 extract.py {path}'
59
+ --on-file 'python3 extract.py'
57
60
  ```
58
61
 
59
62
  ```json
60
63
  [{"author":"Ada Lovelace","title":"Hello World"},{"author":"Alan Turing","title":"On Recursion"}]
61
64
  ```
62
65
 
63
- Now the parser's output *is* the table. `--on-file` runs the command once per
64
- matched file; `{path}` is the file's absolute path, one of the placeholders in
65
- the shared [`on-file` hook contract](../reference/hooks.md#on-file) (argv
66
- splitting, timeout, and per-file failure isolation all come from there). The
67
- stat columns are no longer reachable — a parser that wants the path emits it,
68
- since it already has `{path}`. See
66
+ Now the parser's output *is* the table. `--on-file` runs the command once,
67
+ with every matched file's absolute path appended as an argument, under the
68
+ shared [`on-file` hook contract](../reference/hooks.md#on-file) (argv
69
+ splitting, timeout, and failure semantics all come from there). The stat
70
+ columns are no longer reachable — a parser that wants the path emits it, since
71
+ it already has the path. See
69
72
  [Parsing rows with `--on-file`](../reference/path-tables.md#parsing-rows-with-on-file)
70
73
  for the full behavior.
71
74
 
@@ -87,7 +90,7 @@ command in verbatim:
87
90
  name = "posts"
88
91
  ddl = "CREATE TABLE posts (title TEXT, author TEXT)"
89
92
  glob = "posts/*.md"
90
- on-file = "python3 extract.py {path}"
93
+ on-file = "python3 extract.py"
91
94
  ```
92
95
 
93
96
  The `on-file` value is byte-for-byte the string you passed to `--on-file`. The
@@ -109,7 +112,7 @@ indexed on build, kept fresh by the watcher, survives restarts with
109
112
  [many tables](./define-tables.md) — each with its own `on-file` — where the flag
110
113
  gives every path-table one parser. In both spellings the table's columns are
111
114
  exactly what the parser emits: `dirsql` merges no filesystem facts back on. A
112
- row that needs the file's `path` emits it (the parser has `{path}`).
115
+ row that needs the file's `path` emits it (the parser has the path).
113
116
 
114
117
  ## Going further
115
118
 
@@ -17,7 +17,7 @@ prints each file's basename:
17
17
  name = "files"
18
18
  ddl = "CREATE TABLE files (basename TEXT)"
19
19
  glob = "**/*"
20
- on-file = '''sh -c 'printf "[{\"basename\":\"%s\"}]" "${1##*/}"' sh {path}'''
20
+ on-file = '''sh -c 'printf "["; sep=""; for p; do printf "%s{\"basename\":\"%s\"}" "$sep" "${p##*/}"; sep=","; done; printf "]"' sh'''
21
21
  ```
22
22
 
23
23
  ## 1. Open the stream
@@ -31,7 +31,7 @@ external-content index beside the row table and let the two triggers feed it:
31
31
  [[table]]
32
32
  name = "notes"
33
33
  glob = "notes/**/*.md"
34
- on-file = "python3 extract.py {path}"
34
+ on-file = "python3 extract.py"
35
35
  ddl = '''
36
36
  CREATE TABLE notes (slug TEXT, title TEXT, body TEXT);
37
37
 
@@ -150,7 +150,7 @@ inserted vector for the "embedding" column. Expected 8 dimensions but received 4
150
150
  [[table]]
151
151
  name = "notes"
152
152
  glob = "notes/**/*.md"
153
- on-file = "python3 extract.py {path}"
153
+ on-file = "python3 extract.py"
154
154
  ddl = '''
155
155
  CREATE TABLE notes (slug TEXT, title TEXT, body TEXT);
156
156
 
@@ -25,7 +25,7 @@ ignore = ["notes/drafts/**", "**/*.tmp"]
25
25
  name = "notes"
26
26
  ddl = "CREATE TABLE notes (basename TEXT)"
27
27
  glob = "notes/**/*"
28
- on-file = '''sh -c 'printf "[{\"basename\":\"%s\"}]" "${1##*/}"' sh {path}'''
28
+ on-file = '''sh -c 'printf "["; sep=""; for p; do printf "%s{\"basename\":\"%s\"}" "$sep" "${p##*/}"; sep=","; done; printf "]"' sh'''
29
29
  ```
30
30
 
31
31
  Patterns match against root-relative paths, the same way table globs do. An
@@ -73,11 +73,11 @@ Both hook command styles from the [hook contract](../reference/hooks.md) work
73
73
  in a plugin fragment:
74
74
 
75
75
  - **Console scripts** — a `bin`-style entry point your package installs on
76
- `PATH` (`embed-file {path} {root}`). Recommended for published plugins: the command
76
+ `PATH` (`embed-file {root}`). Recommended for published plugins: the command
77
77
  is bound to your package's interpreter and dependencies, and it is
78
78
  language-neutral (the fragment names a command, not a Python file).
79
79
  - **Relative scripts** — a path resolved against the fragment's own directory
80
- (`uv run python embed.py {path}`). Convenient while developing the plugin
80
+ (`uv run python embed.py {root}`). Convenient while developing the plugin
81
81
  in-tree.
82
82
 
83
83
  Two facts from the [execution contract](../reference/hooks.md#execution-contract)
@@ -85,12 +85,12 @@ matter most for a published plugin:
85
85
 
86
86
  - **A hook runs in its declaring config's directory.** For a plugin that is
87
87
  the installed fragment's directory — inside **site-packages**. That is a
88
- read-only, shared location: **run from it, never write to it.** Use the
89
- absolute [`{path}`](../reference/hooks.md#on-file) placeholder to read the
90
- matched file, and [`{root}`](../reference/hooks.md#on-file) to reach the
91
- user's project directory. Write any cache to `{root}` or a real cache dir,
88
+ read-only, shared location: **run from it, never write to it.** Read the
89
+ matched files from the absolute paths the hook appends as arguments, and use
90
+ [`{root}`](../reference/hooks.md#on-file) to reach the user's project
91
+ directory. Write any cache to `{root}` or a real cache dir,
92
92
  never next to the fragment.
93
- - **`{path}` is absolute and `{root}` is the index root**, so a command is
93
+ - **The paths are absolute and `{root}` is the index root**, so a command is
94
94
  self-sufficient from any working directory — it works whether the plugin
95
95
  lives in the project or in site-packages.
96
96
 
@@ -112,28 +112,30 @@ entrypoint = "sqlite3_vec_init"
112
112
  name = "notes"
113
113
  ddl = "CREATE TABLE notes (path TEXT, text TEXT, embedding TEXT)"
114
114
  glob = "notes/*.md"
115
- on-file = "uv run --with model2vec python embed.py {path} {root}"
115
+ on-file = "uv run --with model2vec python embed.py {root}"
116
116
  ```
117
117
 
118
- `embed.py` turns one file into one row carrying its path, text, and embedding.
119
- dirsql injects no columns, so the script emits the path itself, from the
120
- `{path}`/`{root}` the hook passes in:
118
+ `embed.py` turns each file into one row carrying its path, text, and
119
+ embedding. dirsql injects no columns, so the script emits the path itself,
120
+ from the `{root}` the hook passes in and the paths it appends:
121
121
 
122
122
  ```python
123
- """Embed one file's text; print a dirsql row array on stdout."""
123
+ """Embed each file's text; print one dirsql row array on stdout."""
124
124
  import json
125
125
  import os
126
126
  import sys
127
127
 
128
128
  from model2vec import StaticModel
129
129
 
130
- path, root = sys.argv[1], sys.argv[2]
131
- text = open(path, encoding="utf-8").read()
130
+ root, paths = sys.argv[1], sys.argv[2:]
131
+ texts = [open(path, encoding="utf-8").read() for path in paths]
132
132
  model = StaticModel.from_pretrained("minishlab/potion-base-8M")
133
- vector = model.encode([text])[0]
134
- row = {"path": os.path.relpath(path, root), "text": text,
135
- "embedding": json.dumps([round(float(x), 6) for x in vector])}
136
- print(json.dumps([row]))
133
+ rows = [
134
+ {"path": os.path.relpath(path, root), "text": text,
135
+ "embedding": json.dumps([round(float(x), 6) for x in vector])}
136
+ for path, text, vector in zip(paths, texts, model.encode(texts))
137
+ ]
138
+ print(json.dumps(rows))
137
139
  ```
138
140
 
139
141
  The relative `embed.py` above resolves against the fragment directory, which
@@ -370,12 +370,12 @@ array of row objects) instead of the stat columns:
370
370
 
371
371
  ```sh
372
372
  dirsql query "SELECT title, author FROM './posts/*.md'" \
373
- --on-file 'extract.py {path}'
373
+ --on-file 'extract.py'
374
374
  ```
375
375
 
376
376
  The command follows the [`on-file` hook contract](./hooks.md#on-file) — argv
377
- splitting, `{path}`/`{root}` placeholders, per-file failure isolation, and the
378
- timeout. The parser's output is the whole schema; the stat columns are not
377
+ splitting, the `{root}` placeholder, every matched path as a trailing
378
+ argument, the failure semantics, and the timeout. The parser's output is the whole schema; the stat columns are not
379
379
  reachable on a parsed path-table. `--on-file` may be given **at most once** (a
380
380
  repeat is an error pointing at config files) and never touches config-declared
381
381
  tables. It is a `query`-only flag — server mode rejects it as an unknown
@@ -12,8 +12,8 @@ A named [`[[table]]`](./config.md#table) — or an SDK
12
12
  [`on-file` hook](./hooks.md#on-file) emits, narrowed to the columns the DDL
13
13
  declares. dirsql adds nothing on top: no `path`, no `size`, no value derived
14
14
  from the filename. A hook that wants any of those computes them itself. The
15
- hook receives the file's `{path}` (an SDK `on_file` callback receives the same
16
- path as its argument) and may stat or read the file however it likes.
15
+ hook receives every matched file's path as an argument (an SDK `on_file`
16
+ callback receives one path) and may stat or read the files however it likes.
17
17
 
18
18
  The hook is **required**. A `[[table]]` with no `on-file` is a
19
19
  [config-load error](./config.md#parse-errors): with nothing supplying columns,
@@ -58,12 +58,12 @@ ORDER BY mtime DESC;
58
58
  [Attaching a parser](./path-tables.md#parsing-rows-with-on-file) to a
59
59
  path-table with `--on-file` **replaces** these stat columns with the parser's
60
60
  own output — the two modes stay cleanly separate, exactly as for a named
61
- table. A parser that wants the path emits it; it has `{path}`.
61
+ table. A parser that wants the path emits it; it has the paths.
62
62
 
63
63
  ## Deriving columns from the path
64
64
 
65
65
  To turn path segments (an author, a year, a thread id) into columns, a hook
66
- splits `{path}` and emits the pieces — the same as any other column it
66
+ splits the path and emits the pieces — the same as any other column it
67
67
  produces. dirsql does not do this for you: a `{name}` segment in a glob is
68
68
  rewritten to `*` and matches a single path segment, but captures no value.
69
69
  [Derive columns from file paths](../howto/columns-from-paths.md) walks a
@@ -24,7 +24,7 @@ the config file's location. See [`--config`](./cli.md#flags).
24
24
 
25
25
  | Key | Type | Default | Description |
26
26
  |---|---|---|---|
27
- | `ignore` | array of strings | `[]` | Glob patterns matched against root-relative paths. Matched files are skipped entirely — excluded from the initial scan and from watch events. |
27
+ | `ignore` | array of strings | `[]` | Glob patterns matched against root-relative paths, under the [one glob rule](#glob-rule): `*` matches one level, `**` any depth. Matched files are skipped entirely — excluded from the initial scan and from watch events. |
28
28
 
29
29
  There is no timeout key. `on-file` hook runs are unbounded; to bound one, wrap
30
30
  its command in `timeout(1)` (see [Command hooks](./hooks.md#bounding-a-hook)).
@@ -43,6 +43,16 @@ directory.
43
43
  ignore = ["node_modules/**", ".git/**"]
44
44
  ```
45
45
 
46
+ ### Glob rule
47
+
48
+ Every glob in dirsql — `ignore`, a `[[table]]`'s `glob`, and a
49
+ [path-table](./path-tables.md#writing-the-path) — reads as the shell does:
50
+ `*` and `?` match within one path segment and never cross `/`; `**` matches
51
+ any depth. So `ignore = ["*"]` hides only the files directly inside the root,
52
+ `ignore = ["build/*"]` hides `build/a.o` but not `build/sub/b.o`, and
53
+ `glob = "*.json"` selects only the top-level `.json` files; write
54
+ `**/*.json` to select them at every depth.
55
+
46
56
  Persistence is not a config key. Keep the SQLite index on disk between runs
47
57
  with the [`--persist [PATH]` CLI flag](./cli.md#dirsql-server) — a machine-local
48
58
  operational choice that belongs to the runner, not to shareable config.
@@ -220,13 +230,13 @@ what its required `on-file` command emits — dirsql injects nothing (see
220
230
  |---|---|---|
221
231
  | `name` | yes | The table's SQL name — the name you query it by. Declared, never derived from `ddl`: dirsql does not read the DDL text. The `ddl` must create a table by this name; if it doesn't, loading fails. |
222
232
  | `ddl` | yes | A SQL batch, run verbatim — any number of statements. It must create a table called `name`; that table holds the file rows, and only the columns it declares are kept (keys the `on-file` command emits that are not declared are dropped). The rest of the batch is yours: indexes, virtual tables, triggers. See [Batch `ddl`](#batch-ddl). |
223
- | `glob` | yes | Glob pattern matched against root-relative paths. Every table whose glob matches a file receives that file's rows — a file can populate multiple tables. A `{name}` segment is rewritten to `*` (it matches one path segment but captures nothing). |
224
- | `on-file` | **yes** | A command run once per matched file; its stdout (a JSON array of row objects) becomes the file's rows. Must be non-empty. A `[[table]]` with no `on-file` is a load error (see [parse errors](#parse-errors)). See [Command hooks](./hooks.md#on-file). |
233
+ | `glob` | yes | Glob pattern matched against root-relative paths, under the [one glob rule](#glob-rule): `*` matches one level, `**` any depth. Every table whose glob matches a file receives that file's rows — a file can populate multiple tables. A `{name}` segment is rewritten to `*` (it matches one path segment but captures nothing). |
234
+ | `on-file` | **yes** | A command run once per table, with every matched file's absolute path appended as a trailing argument; its stdout (one JSON array of row objects) is the table's rows. Must be non-empty. A `[[table]]` with no `on-file` is a load error (see [parse errors](#parse-errors)). See [Command hooks](./hooks.md#on-file). |
225
235
  | `strict` | no (default `false`) | When `true`, rows whose keys do not exactly match the declared columns are rejected with an error: extra keys error, and every declared column must be supplied by the `on-file` output. When `false`, extra keys are dropped and missing columns become `NULL`. |
226
236
 
227
237
  `on-file` is required because a table's rows come from nowhere else. dirsql
228
238
  does not read file contents or merge filesystem facts on your behalf: the
229
- command reads the file (it receives `{path}`) and prints the rows, and those
239
+ command reads the files (it receives their paths as arguments) and prints the rows, and those
230
240
  rows — filtered to the DDL — are the table. For plain stat columns with no
231
241
  command, query the path directly with a [path-table](./path-tables.md)
232
242
  instead of declaring a table.
@@ -236,13 +246,13 @@ instead of declaring a table.
236
246
  name = "comments"
237
247
  ddl = "CREATE TABLE comments (path TEXT, author TEXT, body TEXT)"
238
248
  glob = "_comments/*/*.jsonl"
239
- on-file = "jq -c -s '.' {path}"
249
+ on-file = "jq -c -s '.'"
240
250
 
241
251
  [[table]]
242
252
  name = "papers"
243
253
  ddl = "CREATE TABLE papers (paper_id TEXT, title TEXT)"
244
254
  glob = "**/meta.json"
245
- on-file = "uv run python extract_papers.py {path}"
255
+ on-file = "uv run python extract_papers.py"
246
256
  strict = true
247
257
  ```
248
258
 
@@ -255,7 +265,7 @@ statement:
255
265
  [[table]]
256
266
  name = "messages"
257
267
  glob = "sessions/*/messages/*.json"
258
- on-file = "jq -c '.' {path}"
268
+ on-file = "jq -c -s add"
259
269
  ddl = '''
260
270
  CREATE TABLE messages (session TEXT, idx INT, role TEXT, text TEXT);
261
271
  CREATE INDEX messages_session ON messages(session);
@@ -419,11 +429,11 @@ timeout = "600s"
419
429
  name = "comments"
420
430
  ddl = "CREATE TABLE comments (author TEXT, body TEXT)"
421
431
  glob = "_comments/*/*.jsonl"
422
- on-file = "jq -c -s '.' {path}"
432
+ on-file = "jq -c -s '.'"
423
433
 
424
434
  [[table]]
425
435
  name = "documents"
426
436
  ddl = "CREATE TABLE documents (title TEXT, summary TEXT)"
427
437
  glob = "**/index.md"
428
- on-file = "uv run python extract_doc.py {path}"
438
+ on-file = "uv run python extract_doc.py"
429
439
  ```
@@ -10,7 +10,7 @@ external command under the execution contract below.
10
10
 
11
11
  The command string is split into an argv with shell-like quoting: whitespace
12
12
  separates arguments, and single or double quotes group them (so
13
- `sh -c 'grep foo {path} | sort'` keeps the quoted script as a single
13
+ `sh -c 'grep foo "$@" | sort' sh` keeps the quoted script as a single
14
14
  argument). **No shell is invoked** — there is no globbing, piping, `$VAR`
15
15
  expansion, or `&&`/`;` chaining. To get shell features, ask for a shell
16
16
  explicitly with `sh -c '…'`.
@@ -65,7 +65,7 @@ bound a hook, make the bound part of the command by wrapping it in
65
65
  `timeout(1)`:
66
66
 
67
67
  ```toml
68
- on-file = "timeout 30 my-extractor {path}"
68
+ on-file = "timeout 30 my-extractor"
69
69
  ```
70
70
 
71
71
  When the wrapper kills an overrunning command, the run exits non-zero and
@@ -117,9 +117,14 @@ flag is the inline form, the config key the declared form (see
117
117
  [Parse your files into columns](../howto/parse-files-into-columns.md)). In both
118
118
  spellings the table's columns are exactly what the command emits, narrowed to
119
119
  the DDL — `dirsql` injects no filesystem facts either way. A command that wants
120
- the path or stat metadata emits it (it has `{path}`).
120
+ the path or stat metadata emits it (it has the paths).
121
121
 
122
122
  | Placeholder | Value |
123
123
  |---|---|
124
- | `{path}` | The matched file's **absolute** path. `on-file = "extract.py {path}"` — self-sufficient from any working directory, so the command resolves it even when the config lives outside the index. |
125
- | `{root}` | The index root directory. Derive a root-relative path with `relpath({path}, {root})`. |
124
+ | `{root}` | The index root directory. Derive a root-relative path with `relpath(path, {root})`. |
125
+
126
+ The matched files' **absolute** paths are not placeholders: they are appended
127
+ to the command as trailing arguments, after everything written in the
128
+ command, so `on-file = "extract.py"` receives them as `sys.argv[1:]` — one run
129
+ per table, self-sufficient from any working directory. A command that still
130
+ spells `{path}` is rejected at startup.
@@ -47,8 +47,13 @@ beneath it; the non-recursive form is spelled explicitly with `*`.
47
47
  | `'./docs/**/*.md'` | markdown files at any depth under `docs/` |
48
48
  | `'./notes/today.md'` | exactly that one file — one file is one row |
49
49
 
50
- A path containing `*`, `?` or `[` is a glob and is used exactly as written: `*`
51
- matches within a single directory, `**` crosses directories.
50
+ A path containing `*`, `?`, `[` or `{` is a glob and is used exactly as
51
+ written: `*` matches within a single directory, `**` crosses directories.
52
+
53
+ The scan starts at the last directory named outright before the first glob
54
+ component -- `'./small/*.md'` walks `small/` and nothing else -- so a query
55
+ over one directory costs what `find ./small` costs, however large the
56
+ directories beside it.
52
57
 
53
58
  A path naming a single file yields exactly one row. dirsql never splits a file
54
59
  into rows on its own — that is what a table's `on_file` hook is for.
@@ -139,18 +144,18 @@ the `dirsql query` flag [`--on-file`](/reference/cli):
139
144
 
140
145
  ```sh
141
146
  dirsql query "SELECT title, author FROM './posts/*.md'" \
142
- --on-file 'extract.py {path}'
147
+ --on-file 'extract.py'
143
148
  ```
144
149
 
145
- The command runs once per matched file and prints a JSON array of row objects,
146
- exactly like a declared table's [`on-file` hook](/reference/hooks) — same argv
147
- splitting, same `{path}`/`{root}` placeholders, same timeout. Its output *is*
148
- the table:
150
+ The command runs once, with every matched file's absolute path appended as an
151
+ argument, and prints one JSON array of row objects, exactly like a declared
152
+ table's [`on-file` hook](/reference/hooks) — same argv splitting, same `{root}`
153
+ placeholder, same timeout. Its output *is* the table:
149
154
 
150
155
  - **The parser supplies the whole schema.** Columns are inferred from the keys
151
156
  across the emitted rows. The stat columns (`path`, `size`, …) are **not**
152
157
  reachable on a parsed path-table — a parser that wants the path emits it (it
153
- has `{path}`). The two modes stay cleanly separate.
158
+ has the paths). The two modes stay cleanly separate.
154
159
  - **Failures are isolated per file.** A file whose parser fails (spawn, non-zero
155
160
  exit, timeout, or no output) or whose output is not a JSON array of rows
156
161
  contributes no rows; a one-line warning naming the file goes to stderr and the
@@ -186,8 +191,9 @@ bug to design around.
186
191
  The table itself is per-connection: it lives in `temp`, so it cannot leak into
187
192
  `sqlite_master` or survive a restart. Under `--persist` a *parsed* table's rows
188
193
  outlive the connection in the cache (above), but the table is still minted
189
- fresh each run and the scan still decides what exists. The reserved top-level
190
- `.dirsql/` directory is excluded from the scan, as everywhere else.
194
+ fresh each run and the scan still decides what exists. A `.dirsql/` directory
195
+ at the top of the directory the scan starts in is reserved and excluded, as
196
+ everywhere else.
191
197
 
192
198
  ### When to promote to a declared table
193
199
 
@@ -207,7 +213,9 @@ choice is stated as one trade in
207
213
  ## Skip rules
208
214
 
209
215
  A path-table scan applies the same [`ignore`](/reference/config) patterns your
210
- declared tables use, plus two built-in defaults so a zero-config
216
+ declared tables use — matched against root-relative paths under the same
217
+ [glob rule](/reference/config#glob-rule) as the path itself, `*` one level and
218
+ `**` any depth — plus two built-in defaults so a zero-config
211
219
  `SELECT * FROM './'` does not drown in machinery:
212
220
 
213
221
  - `**/node_modules/**`
@@ -234,8 +242,8 @@ defaults and configured `ignore` patterns still apply.
234
242
 
235
243
  ### Naming a skipped directory
236
244
 
237
- Skip rules are judged on the part of the path *below* what you named outright,
238
- so pointing at a skipped directory — built-in or gitignored — still scans it:
245
+ Skip rules are judged from the directory the scan starts in, so pointing at a
246
+ skipped directory — built-in or gitignored — still scans it:
239
247
 
240
248
  ```sql
241
249
  SELECT path FROM './'; -- no node_modules rows
@@ -243,8 +251,8 @@ SELECT path FROM './node_modules/*/package.json'; -- scans it anyway
243
251
  SELECT path FROM './dist'; -- scans dist/ even when gitignored
244
252
  ```
245
253
 
246
- A `.gitignore` at or below the directory you named still filters beneath it;
247
- only rules inherited from above it are set aside.
254
+ A `.gitignore` at or below the directory the scan starts in still filters
255
+ beneath it; one above it is never read.
248
256
 
249
257
  ### Hidden files
250
258
 
@@ -116,8 +116,9 @@ shortcut was removed in #603 — use
116
116
  - `root` — Directory to index. When omitted, the index roots at the process
117
117
  cwd (even when `config` is supplied).
118
118
  - `tables` — Programmatic [`Table`](#table) definitions.
119
- - `ignore` — Glob patterns matched against root-relative paths; matched
120
- files are skipped entirely (scan and watch).
119
+ - `ignore` — Glob patterns matched against root-relative paths under the
120
+ [glob rule](./config.md#glob-rule) (`*` one level, `**` any depth);
121
+ matched files are skipped entirely (scan and watch).
121
122
  - `no_ignore` / `noIgnore` — Scan files a `.gitignore` would hide.
122
123
  [Path-tables](./path-tables.md#skip-rules) respect `.gitignore` files by
123
124
  default; the built-in `node_modules`/`.git` skips and any `ignore` patterns
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dirsql",
3
- "version": "0.4.65",
3
+ "version": "0.4.67",
4
4
  "description": "Ephemeral SQL index over a local directory",
5
5
  "license": "MIT",
6
6
  "repository": "https://github.com/thekevinscott/dirsql",
@@ -221,10 +221,10 @@
221
221
  ]
222
222
  },
223
223
  "optionalDependencies": {
224
- "@dirsql/lib-linux-x64-gnu": "0.4.65",
225
- "@dirsql/lib-linux-arm64-gnu": "0.4.65",
226
- "@dirsql/lib-darwin-x64": "0.4.65",
227
- "@dirsql/lib-darwin-arm64": "0.4.65",
228
- "@dirsql/lib-win32-x64-msvc": "0.4.65"
224
+ "@dirsql/lib-linux-x64-gnu": "0.4.67",
225
+ "@dirsql/lib-linux-arm64-gnu": "0.4.67",
226
+ "@dirsql/lib-darwin-x64": "0.4.67",
227
+ "@dirsql/lib-darwin-arm64": "0.4.67",
228
+ "@dirsql/lib-win32-x64-msvc": "0.4.67"
229
229
  }
230
230
  }