dirsql 0.3.128 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,12 +1,12 @@
1
1
  # Your first dirsql database
2
2
 
3
3
  In this tutorial you will turn a directory of three tiny markdown files into
4
- a SQL database — with a single command, and without writing any code. You
5
- will:
4
+ a SQL database. You will:
6
5
 
7
6
  1. Create the directory and files.
8
- 2. Query them straight away with zero configuration.
9
- 3. Declare your own table to name and reuse a shape, and query it.
7
+ 2. Query them straight away with zero configuration and no code.
8
+ 3. Declare your own named table — with a tiny parser that pulls a column out
9
+ of each file — and query it.
10
10
 
11
11
  It takes about five minutes.
12
12
 
@@ -140,25 +140,44 @@ A declared table fixes a shape once: you give it a name, scope it to exactly
140
140
  the files you care about, and then query it by name instead of repeating a
141
141
  path in every question. It is also the on-ramp to everything a path-table
142
142
  can't do — a named table can be kept live by the watcher, persisted across
143
- restarts, and given a parser that reads inside your files.
143
+ restarts, and given a parser that reads *inside* your files.
144
144
 
145
- Still inside `my-notes`, create a `.dirsql.toml`:
145
+ That parser is the point: a named table's columns are exactly what its
146
+ `on-file` command emits. `dirsql` adds nothing on its own — so this is where
147
+ you pull a value out of each file. Still inside `my-notes`, write a tiny
148
+ parser that reads a note's title line (the `# Heading`) and its author (the
149
+ folder name), and prints them as a JSON row:
150
+
151
+ ```bash
152
+ cat > note.sh <<'EOF'
153
+ #!/usr/bin/env sh
154
+ title=$(sed -n 's/^# //p' "$1" | head -n1)
155
+ author=$(basename "$(dirname "$1")")
156
+ printf '[{"title":%s,"author":%s}]' \
157
+ "$(jq -Rn --arg t "$title" '$t')" \
158
+ "$(jq -Rn --arg a "$author" '$a')"
159
+ EOF
160
+ ```
161
+
162
+ Now create a `.dirsql.toml` that points a table at it:
146
163
 
147
164
  ```bash
148
165
  cat > .dirsql.toml <<'EOF'
149
166
  [[table]]
150
- ddl = "CREATE TABLE notes (dir TEXT, basename TEXT, size INTEGER)"
151
- glob = "notes/**/*.md"
167
+ ddl = "CREATE TABLE notes (title TEXT, author TEXT)"
168
+ glob = "notes/**/*.md"
169
+ on-file = "sh note.sh {path}"
152
170
  EOF
153
171
  ```
154
172
 
155
- Two keys define the table:
173
+ Three keys define the table:
156
174
 
157
175
  - `glob` selects which files feed the table — every `.md` at any depth under
158
176
  `notes/`, relative to the directory the config sits in.
159
177
  - `ddl` is ordinary `CREATE TABLE` SQL naming the columns you want to keep.
160
- Each is a [stat column](./reference/columns.md#stat-columns) `dirsql`
161
- computes for every file.
178
+ - `on-file` is the command run once per matched file; `{path}` is the file's
179
+ path, and its printed JSON row becomes the file's row. The columns are
180
+ exactly what it emits — `title` from the heading, `author` from the folder.
162
181
 
163
182
  ## 5. Query the table
164
183
 
@@ -168,11 +187,11 @@ pass it explicitly with `-c`, **after** the SQL:
168
187
  ::: code-group
169
188
 
170
189
  ```bash [npm]
171
- npx dirsql "SELECT dir, basename, size FROM notes ORDER BY dir, basename" -c .dirsql.toml | jq
190
+ npx dirsql "SELECT title, author FROM notes ORDER BY author, title" -c .dirsql.toml | jq
172
191
  ```
173
192
 
174
193
  ```bash [PyPI]
175
- uvx dirsql "SELECT dir, basename, size FROM notes ORDER BY dir, basename" -c .dirsql.toml | jq
194
+ uvx dirsql "SELECT title, author FROM notes ORDER BY author, title" -c .dirsql.toml | jq
176
195
  ```
177
196
 
178
197
  :::
@@ -180,35 +199,33 @@ uvx dirsql "SELECT dir, basename, size FROM notes ORDER BY dir, basename" -c .di
180
199
  ```json
181
200
  [
182
201
  {
183
- "basename": "ideas.md",
184
- "dir": "notes/alice",
185
- "size": 52
202
+ "author": "alice",
203
+ "title": "Ideas"
186
204
  },
187
205
  {
188
- "basename": "welcome.md",
189
- "dir": "notes/alice",
190
- "size": 66
206
+ "author": "alice",
207
+ "title": "Welcome"
191
208
  },
192
209
  {
193
- "basename": "reading-list.md",
194
- "dir": "notes/bob",
195
- "size": 41
210
+ "author": "bob",
211
+ "title": "Reading list"
196
212
  }
197
213
  ]
198
214
  ```
199
215
 
200
- You queried `FROM notes` by name — no path, no glob to repeat. And because
201
- `dir` is a real SQL column, you can aggregate on it. Count each author's
202
- notes by their folder:
216
+ You queried `FROM notes` by name — no path, no glob to repeat — and `title`
217
+ came from *inside* each file, something a path-table can't reach. Because
218
+ `author` is a real SQL column, you can aggregate on it. Count each author's
219
+ notes:
203
220
 
204
221
  ::: code-group
205
222
 
206
223
  ```bash [npm]
207
- npx dirsql query "SELECT dir, COUNT(*) AS notes FROM notes GROUP BY dir ORDER BY dir" -c .dirsql.toml | jq
224
+ npx dirsql query "SELECT author, COUNT(*) AS notes FROM notes GROUP BY author ORDER BY author" -c .dirsql.toml | jq
208
225
  ```
209
226
 
210
227
  ```bash [PyPI]
211
- uvx dirsql query "SELECT dir, COUNT(*) AS notes FROM notes GROUP BY dir ORDER BY dir" -c .dirsql.toml | jq
228
+ uvx dirsql query "SELECT author, COUNT(*) AS notes FROM notes GROUP BY author ORDER BY author" -c .dirsql.toml | jq
212
229
  ```
213
230
 
214
231
  :::
@@ -216,11 +233,11 @@ uvx dirsql query "SELECT dir, COUNT(*) AS notes FROM notes GROUP BY dir ORDER BY
216
233
  ```json
217
234
  [
218
235
  {
219
- "dir": "notes/alice",
236
+ "author": "alice",
220
237
  "notes": 2
221
238
  },
222
239
  {
223
- "dir": "notes/bob",
240
+ "author": "bob",
224
241
  "notes": 1
225
242
  }
226
243
  ]
@@ -234,7 +251,7 @@ configuration, and a declared table when you want a named shape to reuse.
234
251
  - [Query files without a config](./howto/query-without-config.md) — more
235
252
  path-table questions you can ask with no setup at all.
236
253
  - [Define tables for your files](./howto/define-tables.md) — the full
237
- `[[table]]` recipe: multiple tables, ignore patterns.
254
+ `[[table]]` recipe: multiple tables, each with its own `on-file` parser.
238
255
  - [Extract rows from file contents](./howto/extract-from-contents.md) —
239
256
  pull columns out of *inside* your files with an `on-file` parser.
240
257
  - [CLI](./reference/cli.md) — every flag, plus running `dirsql` as a
@@ -1,10 +1,11 @@
1
1
  # Derive columns from file paths
2
2
 
3
3
  Directory layouts often encode real data — an author, a year, a thread ID —
4
- as path segments. A `{name}` capture in a table's glob turns such a segment
5
- into a queryable column, no extraction code required.
4
+ as path segments. An [`on-file`](../reference/config.md#table) hook receives
5
+ each file's path and can split it into columns, so a query can group and
6
+ filter on those segments.
6
7
 
7
- ## 1. Name the segment in the glob
8
+ ## 1. Split the path in a hook
8
9
 
9
10
  Suppose photos are filed by year and month:
10
11
 
@@ -14,20 +15,35 @@ photos/2024/11/hike.jpg
14
15
  photos/2025/01/snow.jpg
15
16
  ```
16
17
 
17
- Capture both directory levels in `.dirsql.toml`:
18
+ Write a small parser, `pathcols.py`, that turns the path into a row. It
19
+ receives the file's absolute path as its argument and prints a JSON array of
20
+ row objects:
21
+
22
+ ```python
23
+ #!/usr/bin/env python3
24
+ import json, os, sys
25
+
26
+ parts = sys.argv[1].split(os.sep)
27
+ # .../photos/<year>/<month>/<file>
28
+ print(json.dumps([{"year": parts[-3], "month": parts[-2],
29
+ "basename": os.path.basename(sys.argv[1])}]))
30
+ ```
31
+
32
+ Point a table at it in `.dirsql.toml`:
18
33
 
19
34
  ```toml
20
35
  [[table]]
21
- ddl = "CREATE TABLE photos (year TEXT, month TEXT, basename TEXT)"
22
- glob = "photos/{year}/{month}/*.jpg"
36
+ ddl = "CREATE TABLE photos (year TEXT, month TEXT, basename TEXT)"
37
+ glob = "photos/*/*/*.jpg"
38
+ on-file = "python3 pathcols.py {path}"
23
39
  ```
24
40
 
25
- A capture only populates a column when the DDL declares one with the same
26
- name — here `year` and `month`. The capture rules (valid names, matching
27
- within one path segment) are in
28
- [glob captures](../reference/columns.md#glob-captures).
41
+ The hook emits every column the table has — dirsql injects nothing. `{path}`
42
+ is the matched file's absolute path, one of the placeholders in the
43
+ [command hook contract](../reference/hooks.md#on-file). A column appears only
44
+ because the DDL declares it *and* the hook emits it.
29
45
 
30
- ## 2. Query the captured columns
46
+ ## 2. Query the derived columns
31
47
 
32
48
  Pass the config with [`-c`](../reference/cli.md#flags) (`dirsql` does not
33
49
  auto-load a `.dirsql.toml` from the current directory):
@@ -40,7 +56,7 @@ dirsql query "SELECT year, month, basename FROM photos ORDER BY year, month" -c
40
56
  [{"basename":"beach.jpg","month":"05","year":"2024"},{"basename":"hike.jpg","month":"11","year":"2024"},{"basename":"snow.jpg","month":"01","year":"2025"}]
41
57
  ```
42
58
 
43
- Captures are real SQL columns, so aggregation works:
59
+ They are real SQL columns, so aggregation works:
44
60
 
45
61
  ```bash
46
62
  dirsql query "SELECT year, COUNT(*) AS photos FROM photos GROUP BY year" -c ./.dirsql.toml
@@ -52,9 +68,10 @@ dirsql query "SELECT year, COUNT(*) AS photos FROM photos GROUP BY year" -c ./.d
52
68
 
53
69
  ## Going further
54
70
 
55
- - Captures combine freely with [stat columns](../reference/columns.md#stat-columns)
56
- (`basename` above) — both are filesystem facts merged onto every row.
57
- - The [tutorial](../getting-started.md) walks the same idea with an
58
- `{author}` capture, starting from zero.
71
+ - Only need the plain filesystem stat columns (`path`, `basename`, `dir`,
72
+ `ext`, `size`, …) with no code? Query the path directly — a
73
+ [path-table](../reference/path-tables.md) gives them for free, no config.
74
+ - The [tutorial](../getting-started.md) walks the same idea, deriving an
75
+ author from the folder name, starting from zero.
59
76
  - When the value you need lives inside the file rather than in its path,
60
77
  see [Extract rows from file contents](./extract-from-contents.md).
@@ -1,28 +1,44 @@
1
1
  # Define tables for your files
2
2
 
3
3
  Map a glob of files to a named SQL table so you query exactly the files you
4
- care about, with exactly the columns you care about — instead of the
5
- ad-hoc [path-tables](../reference/path-tables.md)
6
- [configless mode](../reference/cli.md#configless-mode) leaves you with.
4
+ care about, by name — a shape you can reuse, keep live with the watcher, and
5
+ persist across restarts, instead of repeating an ad-hoc
6
+ [path-table](../reference/path-tables.md) path in every query.
7
7
 
8
- ## 1. Create a config next to your files
8
+ ## 1. Create a config with a table
9
9
 
10
- Suppose your blog posts live under `posts/`, one markdown file each. In the
11
- directory you want to index, create a `.dirsql.toml` with one
12
- [`[[table]]`](../reference/config.md#table) entry:
10
+ Suppose your blog posts live under `posts/`, one markdown file each. A named
11
+ table needs three keys: a `glob` that selects the files, a `ddl` that names
12
+ the columns, and an [`on-file`](../reference/config.md#table) hook that emits
13
+ each file's rows. Put a small parser next to the config — `extract.py`, which
14
+ reads a post's title line and prints a JSON array of row objects:
15
+
16
+ ```python
17
+ #!/usr/bin/env python3
18
+ import json, os, sys
19
+
20
+ text = open(sys.argv[1], encoding="utf-8").read()
21
+ title = next((l[2:].strip() for l in text.splitlines() if l.startswith("# ")), None)
22
+ print(json.dumps([{"title": title, "slug": os.path.basename(sys.argv[1])[:-3]}]))
23
+ ```
24
+
25
+ Then declare the table in `.dirsql.toml`:
13
26
 
14
27
  ```toml
15
28
  [[table]]
16
- ddl = "CREATE TABLE posts (path TEXT, size INTEGER, mtime INTEGER)"
17
- glob = "posts/**/*.md"
29
+ ddl = "CREATE TABLE posts (title TEXT, slug TEXT)"
30
+ glob = "posts/**/*.md"
31
+ on-file = "python3 extract.py {path}"
18
32
  ```
19
33
 
20
- - `glob` selects the files: every `.md` under `posts/`, at any depth,
21
- relative to the directory containing the config.
22
- - `ddl` is a plain SQLite `CREATE TABLE` naming the columns you want. Here
23
- all three are [stat columns](../reference/columns.md#stat-columns) —
24
- filesystem facts `dirsql` computes for every file. Facts are opt-in by
25
- DDL: only the ones you declare become columns.
34
+ - `glob` selects the files: every `.md` under `posts/`, at any depth, relative
35
+ to the directory containing the config.
36
+ - `ddl` is a plain SQLite `CREATE TABLE` naming the columns you want to keep.
37
+ - `on-file` is **required** — it is where the table's rows come from. dirsql
38
+ injects nothing; the hook emits every column, reading the file (it has
39
+ `{path}`) and deriving whatever it needs. A `[[table]]` with no `on-file` is
40
+ a [config error](../reference/config.md#parse-errors). For plain stat
41
+ columns with no code, query the path directly with a path-table instead.
26
42
 
27
43
  ## 2. Query the table
28
44
 
@@ -31,11 +47,11 @@ auto-load a `.dirsql.toml` from the current directory. Each matched file is
31
47
  one row:
32
48
 
33
49
  ```bash
34
- dirsql query "SELECT path, size FROM posts ORDER BY path" -c ./.dirsql.toml
50
+ dirsql query "SELECT title, slug FROM posts ORDER BY slug" -c ./.dirsql.toml
35
51
  ```
36
52
 
37
53
  ```json
38
- [{"path":"posts/2024/hello.md","size":21},{"path":"posts/2025/again.md","size":55}]
54
+ [{"slug":"again","title":"On Recursion"},{"slug":"hello","title":"Hello World"}]
39
55
  ```
40
56
 
41
57
  Files that don't match the glob (a `README.txt` next to `posts/`, say) are
@@ -43,16 +59,18 @@ simply not in the table. Only the tables you define are served.
43
59
 
44
60
  ## Multiple tables
45
61
 
46
- Add one `[[table]]` entry per table. When a file matches several globs, it
47
- populates every matching table — each table is an independent view. See
48
- [`[[table]]`](../reference/config.md#table) for that and the remaining keys
49
- (`strict`, `on-file`).
62
+ Add one `[[table]]` entry per table — each with its own `glob`, `ddl`, and
63
+ `on-file`. When a file matches several globs, it populates every matching
64
+ table — each table is an independent view. See
65
+ [`[[table]]`](../reference/config.md#table) for the remaining key, `strict`.
50
66
 
51
67
  ## Going further
52
68
 
53
- - Your directory layout encodes data (authors, dates, IDs)? Capture path
54
- segments as columns — [Derive columns from file paths](./columns-from-paths.md).
55
- - Need columns from *inside* the files? A plain table never reads file
56
- contents — [Extract rows from file contents](./extract-from-contents.md).
69
+ - The parser mechanics — placeholders, stdout protocol, per-file failure
70
+ isolation — are the [`on-file` hook contract](../reference/hooks.md#on-file);
71
+ [Extract rows from file contents](./extract-from-contents.md) is the fuller
72
+ recipe.
73
+ - Your directory layout encodes data (authors, dates, IDs)? Split the path in
74
+ the hook — [Derive columns from file paths](./columns-from-paths.md).
57
75
  - Why one row per file, rebuilt from disk? See
58
76
  [how `dirsql` thinks](../explanation.md).
@@ -34,22 +34,25 @@ comments/t2/c1.json # {"body": "following up", "author": "alice"}
34
34
  ```
35
35
 
36
36
  Unlike a config-file table, a programmatic table takes an `on_file`
37
- callback — your code reads each matched file and returns its rows, with
38
- [glob captures and stat columns](../reference/columns.md) merged on
39
- automatically (here, `{thread}` from the path):
37
+ callback — your code reads each matched file and returns its rows. The row's
38
+ columns are exactly what the callback returns; `dirsql` merges nothing on top
39
+ (see [Columns](../reference/columns.md)), so the callback derives the `thread`
40
+ from the file's path itself:
40
41
 
41
42
  ::: code-group
42
43
 
43
44
  ```python [Python]
44
45
  import asyncio
45
46
  import json
47
+ import os
46
48
 
47
49
  from dirsql import DirSQL, Table
48
50
 
49
51
 
50
52
  def on_file(path: str) -> list[dict]:
53
+ thread = os.path.basename(os.path.dirname(path))
51
54
  with open(path, encoding="utf-8") as f:
52
- return [json.load(f)]
55
+ return [{**json.load(f), "thread": thread}]
53
56
 
54
57
 
55
58
  async def main() -> None:
@@ -58,7 +61,7 @@ async def main() -> None:
58
61
  tables=[
59
62
  Table(
60
63
  ddl="CREATE TABLE comments (thread TEXT, author TEXT, body TEXT)",
61
- glob="comments/{thread}/*.json",
64
+ glob="comments/*/*.json",
62
65
  on_file=on_file,
63
66
  )
64
67
  ],
@@ -73,6 +76,7 @@ asyncio.run(main())
73
76
 
74
77
  ```typescript [TypeScript]
75
78
  import { readFileSync } from "node:fs";
79
+ import { basename, dirname } from "node:path";
76
80
  import { DirSQL } from "dirsql";
77
81
 
78
82
  const db = new DirSQL({
@@ -80,8 +84,10 @@ const db = new DirSQL({
80
84
  tables: [
81
85
  {
82
86
  ddl: "CREATE TABLE comments (thread TEXT, author TEXT, body TEXT)",
83
- glob: "comments/{thread}/*.json",
84
- onFile: (path) => [JSON.parse(readFileSync(path, "utf8"))],
87
+ glob: "comments/*/*.json",
88
+ onFile: (path) => [
89
+ { ...JSON.parse(readFileSync(path, "utf8")), thread: basename(dirname(path)) },
90
+ ],
85
91
  },
86
92
  ],
87
93
  });
@@ -110,6 +116,12 @@ fn on_file(path: &str) -> Vec<HashMap<String, Value>> {
110
116
  row.insert(key, Value::Text(s));
111
117
  }
112
118
  }
119
+ let thread = std::path::Path::new(path)
120
+ .parent()
121
+ .and_then(|p| p.file_name())
122
+ .and_then(|n| n.to_str())
123
+ .unwrap_or_default();
124
+ row.insert("thread".into(), Value::Text(thread.into()));
113
125
  vec![row]
114
126
  }
115
127
 
@@ -118,7 +130,7 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
118
130
  .root("./comments-root")
119
131
  .table(Table::new(
120
132
  "CREATE TABLE comments (thread TEXT, author TEXT, body TEXT)",
121
- "comments/{thread}/*.json",
133
+ "comments/*/*.json",
122
134
  on_file,
123
135
  ))
124
136
  .build()?;
@@ -18,7 +18,7 @@ stdout works. With [`jq`](https://jqlang.org/):
18
18
 
19
19
  ```toml
20
20
  [[table]]
21
- ddl = "CREATE TABLE books (title TEXT, author TEXT, year INTEGER, path TEXT)"
21
+ ddl = "CREATE TABLE books (title TEXT, author TEXT, year INTEGER)"
22
22
  glob = "books/*.json"
23
23
  on-file = "jq -c '[{title, author, year}]' {path}"
24
24
  ```
@@ -35,17 +35,16 @@ Pass the config with [`-c`](../reference/cli.md#flags) (`dirsql` does not
35
35
  auto-load a `.dirsql.toml` from the current directory):
36
36
 
37
37
  ```bash
38
- dirsql query "SELECT title, author, year, path FROM books ORDER BY year" -c ./.dirsql.toml
38
+ dirsql query "SELECT title, author, year FROM books ORDER BY year" -c ./.dirsql.toml
39
39
  ```
40
40
 
41
41
  ```json
42
- [{"path":"books/bleak-house.json","author":"Charles Dickens","title":"Bleak House","year":1852},{"path":"books/middlemarch.json","author":"George Eliot","title":"Middlemarch","year":1871}]
42
+ [{"author":"Charles Dickens","title":"Bleak House","year":1852},{"author":"George Eliot","title":"Middlemarch","year":1871}]
43
43
  ```
44
44
 
45
- Filesystem facts are still merged onto every row — `path` above comes from
46
- `dirsql`, not from `jq`. When the command emits a key that collides with a
47
- fact, the command wins
48
- ([precedence](../reference/columns.md#precedence)).
45
+ The table's columns are exactly what the command emits, narrowed to the DDL —
46
+ `dirsql` adds nothing. To include the file's `path`, have the command emit it
47
+ (it has `{path}`); dirsql will not merge it in for you.
49
48
 
50
49
  ## Multiple rows per file
51
50
 
@@ -54,7 +53,7 @@ row per line, slurp it:
54
53
 
55
54
  ```toml
56
55
  [[table]]
57
- ddl = "CREATE TABLE events (event TEXT, user TEXT, path TEXT)"
56
+ ddl = "CREATE TABLE events (event TEXT, user TEXT)"
58
57
  glob = "logs/*.jsonl"
59
58
  on-file = "jq -c -s '.' {path}"
60
59
  ```
@@ -106,10 +106,9 @@ Same command, same rows. The difference is what a declared table brings: it is
106
106
  indexed on build, kept fresh by the watcher, survives restarts with
107
107
  [`--persist`](./persist.md), and a config can declare
108
108
  [many tables](./define-tables.md) — each with its own `on-file` — where the flag
109
- gives every path-table one parser. A declared table also merges the filesystem
110
- facts back onto each row, so a `path TEXT` column in the DDL is populated by
111
- `dirsql` even though the parser did not emit it
112
- ([precedence](../reference/columns.md#precedence)).
109
+ gives every path-table one parser. In both spellings the table's columns are
110
+ exactly what the parser emits: `dirsql` merges no filesystem facts back on. A
111
+ row that needs the file's `path` emits it (the parser has `{path}`).
113
112
 
114
113
  ## Going further
115
114
 
@@ -7,13 +7,16 @@ side.
7
7
 
8
8
  Row events are emitted for **named tables**, so this flow needs a config —
9
9
  [path-tables](../reference/path-tables.md) are scanned per query and are not
10
- watched. Define one next to your files:
10
+ watched. Define one next to your files. A named table's columns are whatever
11
+ its [`on-file`](../reference/config.md#table) hook emits; here a minimal hook
12
+ prints each file's basename:
11
13
 
12
14
  ```toml
13
15
  # .dirsql.toml
14
16
  [[table]]
15
- ddl = "CREATE TABLE files (path TEXT, basename TEXT, dir TEXT, ext TEXT, size INTEGER, mtime INTEGER, ctime INTEGER)"
16
- glob = "**/*"
17
+ ddl = "CREATE TABLE files (basename TEXT)"
18
+ glob = "**/*"
19
+ on-file = '''sh -c 'printf "[{\"basename\":\"%s\"}]" "${1##*/}"' sh {path}'''
17
20
  ```
18
21
 
19
22
  ## 1. Open the stream
@@ -44,7 +47,7 @@ The stream delivers the resulting row change:
44
47
 
45
48
  ```
46
49
  event: row
47
- data: {"action":"insert","file_path":"inbox/two.txt","old_row":null,"row":{"basename":"two.txt","ctime":1783170226,"dir":"inbox","ext":"txt","mtime":1783170226,"path":"inbox/two.txt","size":7},"table":"files"}
50
+ data: {"action":"insert","file_path":"inbox/two.txt","old_row":null,"row":{"basename":"two.txt"},"table":"files"}
48
51
  ```
49
52
 
50
53
  Edits arrive as `update` events carrying both the old and new row;
@@ -41,21 +41,25 @@ notes/tomatoes.md # planting tomato seedlings after the last frost
41
41
  ```
42
42
 
43
43
  Next to them, `embed.py` turns one file into one row carrying its text and
44
- its embedding (a JSON array, stored as TEXT — `sqlite-vec` accepts JSON
45
- vectors directly):
44
+ its `path`, text, and embedding (a JSON array, stored as TEXT — `sqlite-vec`
45
+ accepts JSON vectors directly). dirsql injects no columns, so the script emits
46
+ the path itself, deriving it from the `{path}`/`{root}` the hook passes in:
46
47
 
47
48
  ```python
48
49
  """Embed one file's text; print a dirsql row array on stdout."""
49
50
  import json
51
+ import os
50
52
  import sys
51
53
 
52
54
  from model2vec import StaticModel
53
55
 
54
- path = sys.argv[1]
56
+ path, root = sys.argv[1], sys.argv[2]
55
57
  text = open(path, encoding="utf-8").read()
56
58
  model = StaticModel.from_pretrained("minishlab/potion-base-8M")
57
59
  vector = model.encode([text])[0]
58
- print(json.dumps([{"text": text, "embedding": json.dumps([round(float(x), 6) for x in vector])}]))
60
+ row = {"path": os.path.relpath(path, root), "text": text,
61
+ "embedding": json.dumps([round(float(x), 6) for x in vector])}
62
+ print(json.dumps([row]))
59
63
  ```
60
64
 
61
65
  And `search.py` turns a `{"q": "..."}` request body into SQL, embedding the
@@ -98,7 +102,7 @@ entrypoint = "sqlite3_vec_init"
98
102
  [[table]]
99
103
  ddl = "CREATE TABLE notes (path TEXT, text TEXT, embedding TEXT)"
100
104
  glob = "notes/*.md"
101
- on-file = "uv run --with model2vec python embed.py {path}"
105
+ on-file = "uv run --with model2vec python embed.py {path} {root}"
102
106
  ```
103
107
 
104
108
  The extension is named by package: the Python launcher resolves the
@@ -22,8 +22,9 @@ Exclude the noise in `.dirsql.toml`:
22
22
  ignore = ["notes/drafts/**", "**/*.tmp"]
23
23
 
24
24
  [[table]]
25
- ddl = "CREATE TABLE notes (path TEXT)"
26
- glob = "notes/**/*"
25
+ ddl = "CREATE TABLE notes (basename TEXT)"
26
+ glob = "notes/**/*"
27
+ on-file = '''sh -c 'printf "[{\"basename\":\"%s\"}]" "${1##*/}"' sh {path}'''
27
28
  ```
28
29
 
29
30
  Patterns match against root-relative paths, the same way table globs do. An
@@ -35,11 +36,11 @@ Pass the config with [`-c`](../reference/cli.md#flags) (`dirsql` does not
35
36
  auto-load a `.dirsql.toml` from the current directory):
36
37
 
37
38
  ```bash
38
- dirsql query "SELECT path FROM notes ORDER BY path" -c ./.dirsql.toml
39
+ dirsql query "SELECT basename FROM notes ORDER BY basename" -c ./.dirsql.toml
39
40
  ```
40
41
 
41
42
  ```json
42
- [{"path":"notes/final.md"}]
43
+ [{"basename":"final.md"}]
43
44
  ```
44
45
 
45
46
  ## Notes
@@ -72,7 +72,7 @@ Both hook command styles from the [hook contract](../reference/hooks.md) work
72
72
  in a plugin fragment:
73
73
 
74
74
  - **Console scripts** — a `bin`-style entry point your package installs on
75
- `PATH` (`embed-file {path}`). Recommended for published plugins: the command
75
+ `PATH` (`embed-file {path} {root}`). Recommended for published plugins: the command
76
76
  is bound to your package's interpreter and dependencies, and it is
77
77
  language-neutral (the fragment names a command, not a Python file).
78
78
  - **Relative scripts** — a path resolved against the fragment's own directory
@@ -116,23 +116,28 @@ entrypoint = "sqlite3_vec_init"
116
116
  [[table]]
117
117
  ddl = "CREATE TABLE notes (path TEXT, text TEXT, embedding TEXT)"
118
118
  glob = "notes/*.md"
119
- on-file = "uv run --with model2vec python embed.py {path}"
119
+ on-file = "uv run --with model2vec python embed.py {path} {root}"
120
120
  ```
121
121
 
122
- `embed.py` turns one file into one row carrying its text and its embedding:
122
+ `embed.py` turns one file into one row carrying its path, text, and embedding.
123
+ dirsql injects no columns, so the script emits the path itself, from the
124
+ `{path}`/`{root}` the hook passes in:
123
125
 
124
126
  ```python
125
127
  """Embed one file's text; print a dirsql row array on stdout."""
126
128
  import json
129
+ import os
127
130
  import sys
128
131
 
129
132
  from model2vec import StaticModel
130
133
 
131
- path = sys.argv[1]
134
+ path, root = sys.argv[1], sys.argv[2]
132
135
  text = open(path, encoding="utf-8").read()
133
136
  model = StaticModel.from_pretrained("minishlab/potion-base-8M")
134
137
  vector = model.encode([text])[0]
135
- print(json.dumps([{"text": text, "embedding": json.dumps([round(float(x), 6) for x in vector])}]))
138
+ row = {"path": os.path.relpath(path, root), "text": text,
139
+ "embedding": json.dumps([round(float(x), 6) for x in vector])}
140
+ print(json.dumps([row]))
136
141
  ```
137
142
 
138
143
  `search.py` turns a `{"q": "..."}` request body into nearest-neighbor SQL:
@@ -1,31 +1,37 @@
1
- # Stat columns and glob captures
1
+ # Columns
2
2
 
3
- Every table — config-defined or programmatic — gets filesystem facts merged
4
- onto its rows automatically: seven **stat columns** derived from the file's
5
- path and stat metadata, plus one column per **`{name}` capture** in the
6
- table's glob.
3
+ Where a dirsql table's columns come from depends on which kind of table it is.
4
+ The rule that governs both is the same: **dirsql never injects a column your
5
+ table did not produce.** There is no automatic path/size/mtime merge, and glob
6
+ `{name}` segments capture nothing.
7
7
 
8
- These are ordinary, physically stored `TEXT`/`INTEGER` columns, computed
9
- once per file at scan time and written like any other column value — not
10
- SQLite's `GENERATED ... VIRTUAL` columns (computed on the fly, never stored)
11
- and not part of a `CREATE VIRTUAL TABLE` (dirsql tables are always real
12
- tables). "Stat" describes where the value comes from — the file's path and
13
- `stat` metadata, as opposed to its content — not how it's stored.
8
+ ## Named tables: exactly the hook's output
14
9
 
15
- Facts are **opt-in by DDL**: only facts whose name appears as a column in
16
- the table's `CREATE TABLE` are populated; the rest are silently dropped.
17
- Declaring them requires nothing else. The names below aren't a protected or
18
- enforced namespace — nothing stops you from declaring a column with one of
19
- these names for an unrelated purpose, in which case dirsql's computed value
20
- lands there like any other fact (unless your own row source — an `on-file`
21
- command or SDK `on_file` callback — supplies its own value for that name;
22
- see [Precedence](#precedence)).
10
+ A named [`[[table]]`](./config.md#table) — or an SDK
11
+ [`Table`](./sdk.md#table) — has exactly the columns its
12
+ [`on-file` hook](./hooks.md#on-file) emits, narrowed to the columns the DDL
13
+ declares. dirsql adds nothing on top: no `path`, no `size`, no value derived
14
+ from the filename. A hook that wants any of those computes them itself. The
15
+ hook receives the file's `{path}` (an SDK `on_file` callback receives the same
16
+ path as its argument) and may stat or read the file however it likes.
23
17
 
24
- ## Stat columns
18
+ The hook is **required**. A `[[table]]` with no `on-file` is a
19
+ [config-load error](./config.md#parse-errors): with nothing supplying columns,
20
+ every row would be all-NULL. The error points at the fix — add a hook, or, for
21
+ plain stat columns with no code, query the path directly with a path-table.
22
+
23
+ ## Path-tables: stat columns
24
+
25
+ A [path-table](./path-tables.md) (`FROM './'`) is the one place dirsql supplies
26
+ columns for you. Each matched file becomes one row carrying seven **stat
27
+ columns** — derived from the file's path and its `stat` metadata — plus a
28
+ lazily-read hidden [`content`](./path-tables.md#columns) column.
29
+
30
+ ### Stat columns
25
31
 
26
32
  | Column | Type | Value |
27
33
  |---|---|---|
28
- | `path` | TEXT | The file's path relative to the scan root (e.g. `posts/hello.md`). |
34
+ | `path` | TEXT | The file's path relative to the scan root (e.g. `posts/hello.md`); absolute for a `/`, `../` or `~/` path-table. |
29
35
  | `basename` | TEXT | The filename, including extension (`hello.md`). |
30
36
  | `dir` | TEXT | The parent directory relative to the root (`posts`); the empty string for files directly under the root. |
31
37
  | `ext` | TEXT | The file extension without the leading dot (`md`). Original case is preserved — `Photo.JPG` yields `JPG`; use `LOWER(ext)` for case-insensitive matching. `NULL` when the file has no extension. |
@@ -33,48 +39,32 @@ see [Precedence](#precedence)).
33
39
  | `mtime` | INTEGER | Last-modified time, Unix seconds. |
34
40
  | `ctime` | INTEGER | Creation (birth) time, Unix seconds. `NULL` when the platform or filesystem cannot supply it. |
35
41
 
36
- A fact that cannot be computed (an unreadable file's `size`/`mtime`/
37
- `ctime`, a missing extension's `ext`) is absent from the row: `NULL` in
38
- the default relaxed mode, a missing-column error for a
39
- [`strict`](./config.md#table) table that declares it.
42
+ These are ordinary stored `TEXT`/`INTEGER` values, computed once per file at
43
+ scan time — not SQLite `GENERATED ... VIRTUAL` columns and not part of a
44
+ `CREATE VIRTUAL TABLE`. "Stat" describes where the value comes from — the
45
+ file's path and `stat` metadata, as opposed to its content — not how it is
46
+ stored.
47
+
48
+ A stat value that cannot be computed (an unreadable file's `size`, a missing
49
+ extension's `ext`) is `NULL`.
40
50
 
41
51
  ```sql
42
52
  SELECT basename, size
43
- FROM posts
53
+ FROM './posts'
44
54
  WHERE mtime > strftime('%s', '2024-01-01')
45
55
  ORDER BY mtime DESC;
46
56
  ```
47
57
 
48
- ## Glob captures
49
-
50
- A `{name}` segment in a table's glob captures part of each matched path as
51
- a TEXT column named `name`:
52
-
53
- ```toml
54
- [[table]]
55
- ddl = "CREATE TABLE comments (thread_id TEXT, basename TEXT, mtime INTEGER)"
56
- glob = "_comments/{thread_id}/*.jsonl"
57
- ```
58
-
59
- A file at `_comments/abc123/2024-05-05.jsonl` produces a row with
60
- `thread_id = "abc123"`.
61
-
62
- - A capture name must be a valid identifier: a letter or underscore
63
- followed by letters, digits, or underscores (`[a-zA-Z_][a-zA-Z0-9_]*`).
64
- - A capture matches **within a single path segment** — one or more
65
- characters, never a `/`. For matching purposes, `{name}` behaves like
66
- `*`.
67
- - A glob may contain multiple captures (`{year}/{month}/*.jpg`).
68
- - Like stat columns, a capture populates a column only when the DDL
69
- declares a column of the same name.
70
-
71
- ## Precedence
58
+ [Attaching a parser](./path-tables.md#parsing-rows-with-on-file) to a
59
+ path-table with `--on-file` **replaces** these stat columns with the parser's
60
+ own output — the two modes stay cleanly separate, exactly as for a named
61
+ table. A parser that wants the path emits it; it has `{path}`.
72
62
 
73
- Values produced by a table's own row source — an `on-file` command's JSON
74
- output or an SDK `on_file` callback's return value — **win** over
75
- auto-injected facts of the same name. An `on_file` callback that explicitly emits
76
- `path` is honored.
63
+ ## Deriving columns from the path
77
64
 
78
- Injection order per row: stat columns first, then glob captures, then
79
- the row source's own values, each layer overwriting the previous, all
80
- filtered to the DDL's declared columns.
65
+ To turn path segments (an author, a year, a thread id) into columns, a hook
66
+ splits `{path}` and emits the pieces — the same as any other column it
67
+ produces. dirsql does not do this for you: a `{name}` segment in a glob is
68
+ rewritten to `*` and matches a single path segment, but captures no value.
69
+ [Derive columns from file paths](../howto/columns-from-paths.md) walks a
70
+ worked example.
@@ -90,28 +90,29 @@ rejected; call the extension's functions in queries instead.
90
90
 
91
91
  ## `[[table]]`
92
92
 
93
- Each entry maps a glob pattern to a SQL table. Every matched file produces
94
- rows whose columns come from filesystem facts — [glob captures and virtual
95
- columns](./columns.md) — plus, when `on-file` is set, the output of a
96
- per-file command.
93
+ Each entry maps a glob pattern to a SQL table. A table's columns are exactly
94
+ what its required `on-file` command emits — dirsql injects nothing (see
95
+ [Columns](./columns.md)).
97
96
 
98
97
  | Key | Required | Description |
99
98
  |---|---|---|
100
- | `ddl` | yes | A SQLite `CREATE TABLE` statement. The table name is parsed from it. Only columns declared here are populated; auto-injected facts not in the DDL are dropped. |
101
- | `glob` | yes | Glob pattern matched against root-relative paths. May contain `{name}` [capture segments](./columns.md#glob-captures). Every table whose glob matches a file receives that file's rows — a file can populate multiple tables. |
102
- | `strict` | no (default `false`) | When `true`, rows whose keys do not exactly match the declared columns are rejected with an error: extra keys error, and every declared column must be supplied (by the command/on-file output, a glob capture, or a stat column). When `false`, extra keys are dropped and missing columns become `NULL`. |
103
- | `on-file` | no | A command run once per matched file; its stdout (a JSON array of row objects) becomes the file's rows. Must be non-empty. See [Command hooks](./hooks.md#on-file). |
104
-
105
- Without `on-file`, a table produces exactly one row per matched file, built
106
- entirely from filesystem facts. Content interpretation (frontmatter, JSON
107
- fields, CSV parsing) is out of scope for plain config tables — use
108
- `on-file`, or a programmatic [SDK table](./sdk.md#table) with an `on_file`
109
- callback.
99
+ | `ddl` | yes | A SQLite `CREATE TABLE` statement. The table name is parsed from it. Only the columns declared here are kept; keys the `on-file` command emits that are not declared are dropped. |
100
+ | `glob` | yes | Glob pattern matched against root-relative paths. Every table whose glob matches a file receives that file's rows — a file can populate multiple tables. A `{name}` segment is rewritten to `*` (it matches one path segment but captures nothing). |
101
+ | `on-file` | **yes** | A command run once per matched file; its stdout (a JSON array of row objects) becomes the file's rows. Must be non-empty. A `[[table]]` with no `on-file` is a load error (see [parse errors](#parse-errors)). See [Command hooks](./hooks.md#on-file). |
102
+ | `strict` | no (default `false`) | When `true`, rows whose keys do not exactly match the declared columns are rejected with an error: extra keys error, and every declared column must be supplied by the `on-file` output. When `false`, extra keys are dropped and missing columns become `NULL`. |
103
+
104
+ `on-file` is required because a table's rows come from nowhere else. dirsql
105
+ does not read file contents or merge filesystem facts on your behalf: the
106
+ command reads the file (it receives `{path}`) and prints the rows, and those
107
+ rows — filtered to the DDL — are the table. For plain stat columns with no
108
+ command, query the path directly with a [path-table](./path-tables.md)
109
+ instead of declaring a table.
110
110
 
111
111
  ```toml
112
112
  [[table]]
113
- ddl = "CREATE TABLE comments (path TEXT, basename TEXT, mtime INTEGER)"
114
- glob = "_comments/*/*.jsonl"
113
+ ddl = "CREATE TABLE comments (path TEXT, author TEXT, body TEXT)"
114
+ glob = "_comments/*/*.jsonl"
115
+ on-file = "jq -c -s '.' {path}"
115
116
 
116
117
  [[table]]
117
118
  ddl = "CREATE TABLE papers (paper_id TEXT, title TEXT)"
@@ -127,10 +128,10 @@ JSON values map to SQLite as: `null` → `NULL`; `true`/`false` → `1`/`0`; an
127
128
  integral number → `INTEGER`, any other number → `REAL`; a string → `TEXT`; a
128
129
  nested array or object → its JSON text as `TEXT`.
129
130
 
130
- Filesystem facts are still merged onto every `on-file` row; a column emitted
131
- by the command wins over a same-named fact. Output that is not a JSON array
132
- of objects is a per-file failure: the file is skipped with a stderr warning
133
- and the scan continues (see [failure semantics](./hooks.md#failure-semantics)).
131
+ A row's columns are exactly the keys the command emits, narrowed to the DDL;
132
+ dirsql merges nothing else in. Output that is not a JSON array of objects is a
133
+ per-file failure: the file is skipped with a stderr warning and the scan
134
+ continues (see [failure semantics](./hooks.md#failure-semantics)).
134
135
 
135
136
  ## Composing multiple configs
136
137
 
@@ -174,8 +175,13 @@ SDKs raise/reject) when:
174
175
  - Any table contains an unknown key (top level, `[dirsql]`, `[[table]]`, or
175
176
  `[[dirsql.extension]]`). The error names the offending key.
176
177
  - A `[[table]]` entry omits `ddl` or `glob`.
178
+ - A `[[table]]` entry omits `on-file` (or it is empty/whitespace). The error
179
+ names the offending glob and points at the fix:
180
+
181
+ > `[[table]] '**/*.md' has no on-file hook, so every row would be all-NULL. Add an `on-file` hook that emits the columns, or, for stat columns with no code, query the path directly: `FROM './'``
182
+
177
183
  - A `[[dirsql.extension]]` entry omits `path`, or `path` is empty.
178
- - `on-file`, `pre-query`, or `post-query` is present but empty/whitespace.
184
+ - `pre-query` or `post-query` is present but empty/whitespace.
179
185
  - `hook-timeout` is zero or negative.
180
186
 
181
187
  ## Full example
@@ -193,10 +199,12 @@ path = "sqlite_vec" # Python module name; on Node use the
193
199
  entrypoint = "sqlite3_vec_init"
194
200
 
195
201
  [[table]]
196
- ddl = "CREATE TABLE comments (path TEXT, basename TEXT, mtime INTEGER)"
197
- glob = "_comments/*/*.jsonl"
202
+ ddl = "CREATE TABLE comments (author TEXT, body TEXT)"
203
+ glob = "_comments/*/*.jsonl"
204
+ on-file = "jq -c -s '.' {path}"
198
205
 
199
206
  [[table]]
200
- ddl = "CREATE TABLE documents (path TEXT, basename TEXT, size INTEGER)"
201
- glob = "**/index.md"
207
+ ddl = "CREATE TABLE documents (title TEXT, summary TEXT)"
208
+ glob = "**/index.md"
209
+ on-file = "uv run python extract_doc.py {path}"
202
210
  ```
@@ -105,10 +105,10 @@ The same command is attachable two ways, over the same contract: the
105
105
  which attaches it to every [path-table](./path-tables.md#parsing-rows-with-on-file)
106
106
  in the query. The command string is identical between the two spellings — the
107
107
  flag is the inline form, the config key the declared form (see
108
- [Parse your files into columns](../howto/parse-files-into-columns.md)). The one
109
- behavioral difference is not in this contract but in the surrounding table: a
110
- declared `[[table]]` merges the filesystem facts onto each parsed row, while a
111
- parsed path-table exposes only the parser's output.
108
+ [Parse your files into columns](../howto/parse-files-into-columns.md)). In both
109
+ spellings the table's columns are exactly what the command emits, narrowed to
110
+ the DDL — `dirsql` injects no filesystem facts either way. A command that wants
111
+ the path or stat metadata emits it (it has `{path}`).
112
112
 
113
113
  | Placeholder | Value |
114
114
  |---|---|
@@ -338,17 +338,17 @@ Maps files to table rows.
338
338
 
339
339
  - `ddl` — A SQLite `CREATE TABLE` statement; the table name is parsed from
340
340
  it. Table names must be unique across all tables.
341
- - `glob` — Glob pattern matched against root-relative paths. May contain
342
- `{name}` [captures](./columns.md#glob-captures). Every table whose glob
343
- matches a file receives that file's rows — a file can populate multiple
344
- tables.
341
+ - `glob` — Glob pattern matched against root-relative paths. Every table whose
342
+ glob matches a file receives that file's rows — a file can populate multiple
343
+ tables. A `{name}` segment is rewritten to `*` (it matches one path segment
344
+ but captures nothing).
345
345
  - `on_file` (`on_file` / `onFile`) — Callback receiving the matched file's full path (the root
346
346
  joined with the file's relative path — absolute when `root` is absolute)
347
- and returning the rows that file contributes. `dirsql` never reads file
348
- contents itself; a callback that needs the body reads the path. Return
349
- an empty list to skip a file. [Stat columns and glob
350
- captures](./columns.md) are merged onto each returned row; values the
351
- callback emits win over same-named facts.
347
+ and returning the rows that file contributes. Required. A row's columns are
348
+ exactly what the callback returns, narrowed to the DDL — `dirsql` injects
349
+ nothing (see [Columns](./columns.md)). A callback that wants the path, stat
350
+ metadata, or file body computes it from the path it receives; `dirsql` never
351
+ reads file contents itself. Return an empty list to skip a file.
352
352
  - `strict` — Default off: extra row keys are dropped and missing declared
353
353
  columns become `NULL`. When on, any extra or missing key is an error
354
354
  (see [`strict`](./config.md#table)).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dirsql",
3
- "version": "0.3.128",
3
+ "version": "0.4.0",
4
4
  "description": "Ephemeral SQL index over a local directory",
5
5
  "license": "MIT",
6
6
  "repository": "https://github.com/thekevinscott/dirsql",
@@ -212,15 +212,15 @@
212
212
  ]
213
213
  },
214
214
  "optionalDependencies": {
215
- "@dirsql/lib-linux-x64-gnu": "0.3.128",
216
- "@dirsql/lib-linux-arm64-gnu": "0.3.128",
217
- "@dirsql/lib-darwin-x64": "0.3.128",
218
- "@dirsql/lib-darwin-arm64": "0.3.128",
219
- "@dirsql/lib-win32-x64-msvc": "0.3.128",
220
- "@dirsql/cli-linux-x64-gnu": "0.3.128",
221
- "@dirsql/cli-linux-arm64-gnu": "0.3.128",
222
- "@dirsql/cli-darwin-x64": "0.3.128",
223
- "@dirsql/cli-darwin-arm64": "0.3.128",
224
- "@dirsql/cli-win32-x64-msvc": "0.3.128"
215
+ "@dirsql/lib-linux-x64-gnu": "0.4.0",
216
+ "@dirsql/lib-linux-arm64-gnu": "0.4.0",
217
+ "@dirsql/lib-darwin-x64": "0.4.0",
218
+ "@dirsql/lib-darwin-arm64": "0.4.0",
219
+ "@dirsql/lib-win32-x64-msvc": "0.4.0",
220
+ "@dirsql/cli-linux-x64-gnu": "0.4.0",
221
+ "@dirsql/cli-linux-arm64-gnu": "0.4.0",
222
+ "@dirsql/cli-darwin-x64": "0.4.0",
223
+ "@dirsql/cli-darwin-arm64": "0.4.0",
224
+ "@dirsql/cli-win32-x64-msvc": "0.4.0"
225
225
  }
226
226
  }