dirsql 0.4.64 → 0.4.66
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/getting-started.md +16 -9
- package/docs/howto/columns-from-paths.md +14 -11
- package/docs/howto/define-tables.md +12 -9
- package/docs/howto/extract-from-contents.md +5 -5
- package/docs/howto/parse-files-into-columns.md +22 -19
- package/docs/howto/react-to-changes.md +1 -1
- package/docs/howto/search-indexes.md +2 -2
- package/docs/howto/skip-files.md +1 -1
- package/docs/howto/write-a-plugin.md +20 -18
- package/docs/reference/cli.md +3 -3
- package/docs/reference/columns.md +4 -4
- package/docs/reference/config.md +7 -7
- package/docs/reference/hooks.md +10 -5
- package/docs/reference/path-tables.md +6 -6
- package/package.json +6 -6
package/docs/getting-started.md
CHANGED
|
@@ -156,11 +156,17 @@ folder name), and prints them as a JSON row:
|
|
|
156
156
|
```bash
|
|
157
157
|
cat > note.sh <<'EOF'
|
|
158
158
|
#!/usr/bin/env sh
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
"$(
|
|
159
|
+
printf '['
|
|
160
|
+
sep=''
|
|
161
|
+
for f; do
|
|
162
|
+
title=$(sed -n 's/^# //p' "$f" | head -n1)
|
|
163
|
+
author=$(basename "$(dirname "$f")")
|
|
164
|
+
printf '%s{"title":%s,"author":%s}' "$sep" \
|
|
165
|
+
"$(jq -Rn --arg t "$title" '$t')" \
|
|
166
|
+
"$(jq -Rn --arg a "$author" '$a')"
|
|
167
|
+
sep=','
|
|
168
|
+
done
|
|
169
|
+
printf ']'
|
|
164
170
|
EOF
|
|
165
171
|
```
|
|
166
172
|
|
|
@@ -172,7 +178,7 @@ cat > .dirsql.toml <<'EOF'
|
|
|
172
178
|
name = "notes"
|
|
173
179
|
ddl = "CREATE TABLE notes (title TEXT, author TEXT)"
|
|
174
180
|
glob = "notes/**/*.md"
|
|
175
|
-
on-file = "sh note.sh
|
|
181
|
+
on-file = "sh note.sh"
|
|
176
182
|
EOF
|
|
177
183
|
```
|
|
178
184
|
|
|
@@ -181,9 +187,10 @@ Three keys define the table:
|
|
|
181
187
|
- `glob` selects which files feed the table — every `.md` at any depth under
|
|
182
188
|
`notes/`, relative to the directory the config sits in.
|
|
183
189
|
- `ddl` is ordinary `CREATE TABLE` SQL naming the columns you want to keep.
|
|
184
|
-
- `on-file` is the command run once
|
|
185
|
-
path
|
|
186
|
-
exactly what it emits — `title` from the heading, `author`
|
|
190
|
+
- `on-file` is the command run once for the table, with every matched file's
|
|
191
|
+
path appended as an argument; the JSON array it prints is the table's rows.
|
|
192
|
+
The columns are exactly what it emits — `title` from the heading, `author`
|
|
193
|
+
from the folder.
|
|
187
194
|
|
|
188
195
|
## 5. Query the table
|
|
189
196
|
|
|
@@ -15,18 +15,21 @@ photos/2024/11/hike.jpg
|
|
|
15
15
|
photos/2025/01/snow.jpg
|
|
16
16
|
```
|
|
17
17
|
|
|
18
|
-
Write a small parser, `pathcols.py`, that turns
|
|
19
|
-
receives the
|
|
20
|
-
row objects:
|
|
18
|
+
Write a small parser, `pathcols.py`, that turns each path into a row. It
|
|
19
|
+
receives the matched files' absolute paths as its arguments and prints one
|
|
20
|
+
JSON array of row objects:
|
|
21
21
|
|
|
22
22
|
```python
|
|
23
23
|
#!/usr/bin/env python3
|
|
24
24
|
import json, os, sys
|
|
25
25
|
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
26
|
+
rows = []
|
|
27
|
+
for path in sys.argv[1:]:
|
|
28
|
+
parts = path.split(os.sep)
|
|
29
|
+
# .../photos/<year>/<month>/<file>
|
|
30
|
+
rows.append({"year": parts[-3], "month": parts[-2],
|
|
31
|
+
"basename": os.path.basename(path)})
|
|
32
|
+
print(json.dumps(rows))
|
|
30
33
|
```
|
|
31
34
|
|
|
32
35
|
Point a table at it in `.dirsql.toml`:
|
|
@@ -36,12 +39,12 @@ Point a table at it in `.dirsql.toml`:
|
|
|
36
39
|
name = "photos"
|
|
37
40
|
ddl = "CREATE TABLE photos (year TEXT, month TEXT, basename TEXT)"
|
|
38
41
|
glob = "photos/*/*/*.jpg"
|
|
39
|
-
on-file = "python3 pathcols.py
|
|
42
|
+
on-file = "python3 pathcols.py"
|
|
40
43
|
```
|
|
41
44
|
|
|
42
|
-
The hook emits every column the table has — dirsql injects nothing.
|
|
43
|
-
|
|
44
|
-
[command hook contract](../reference/hooks.md#on-file). A column appears only
|
|
45
|
+
The hook emits every column the table has — dirsql injects nothing. The
|
|
46
|
+
matched files' absolute paths arrive as the command's trailing arguments, per
|
|
47
|
+
the [command hook contract](../reference/hooks.md#on-file). A column appears only
|
|
45
48
|
because the DDL declares it *and* the hook emits it.
|
|
46
49
|
|
|
47
50
|
## 2. Query the derived columns
|
|
@@ -10,16 +10,19 @@ persist across restarts, instead of repeating an ad-hoc
|
|
|
10
10
|
Suppose your blog posts live under `posts/`, one markdown file each. A named
|
|
11
11
|
table needs three keys: a `glob` that selects the files, a `ddl` that names
|
|
12
12
|
the columns, and an [`on-file`](../reference/config.md#table) hook that emits
|
|
13
|
-
|
|
14
|
-
reads
|
|
13
|
+
the table's rows. Put a small parser next to the config — `extract.py`, which
|
|
14
|
+
reads each post's title line and prints one JSON array of row objects:
|
|
15
15
|
|
|
16
16
|
```python
|
|
17
17
|
#!/usr/bin/env python3
|
|
18
18
|
import json, os, sys
|
|
19
19
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
20
|
+
rows = []
|
|
21
|
+
for path in sys.argv[1:]:
|
|
22
|
+
text = open(path, encoding="utf-8").read()
|
|
23
|
+
title = next((l[2:].strip() for l in text.splitlines() if l.startswith("# ")), None)
|
|
24
|
+
rows.append({"title": title, "slug": os.path.basename(path)[:-3]})
|
|
25
|
+
print(json.dumps(rows))
|
|
23
26
|
```
|
|
24
27
|
|
|
25
28
|
Then declare the table in `.dirsql.toml`:
|
|
@@ -29,15 +32,15 @@ Then declare the table in `.dirsql.toml`:
|
|
|
29
32
|
name = "posts"
|
|
30
33
|
ddl = "CREATE TABLE posts (title TEXT, slug TEXT)"
|
|
31
34
|
glob = "posts/**/*.md"
|
|
32
|
-
on-file = "python3 extract.py
|
|
35
|
+
on-file = "python3 extract.py"
|
|
33
36
|
```
|
|
34
37
|
|
|
35
38
|
- `glob` selects the files: every `.md` under `posts/`, at any depth, relative
|
|
36
39
|
to the directory containing the config.
|
|
37
40
|
- `ddl` is a plain SQLite `CREATE TABLE` naming the columns you want to keep.
|
|
38
41
|
- `on-file` is **required** — it is where the table's rows come from. dirsql
|
|
39
|
-
injects nothing; the hook emits every column, reading the
|
|
40
|
-
|
|
42
|
+
injects nothing; the hook emits every column, reading the files (their
|
|
43
|
+
paths are its arguments) and deriving whatever it needs. A `[[table]]` with no `on-file` is
|
|
41
44
|
a [config error](../reference/config.md#parse-errors). For plain stat
|
|
42
45
|
columns with no code, query the path directly with a path-table instead.
|
|
43
46
|
|
|
@@ -75,7 +78,7 @@ two triggers:
|
|
|
75
78
|
[[table]]
|
|
76
79
|
name = "posts"
|
|
77
80
|
glob = "posts/**/*.md"
|
|
78
|
-
on-file = "python3 extract.py
|
|
81
|
+
on-file = "python3 extract.py"
|
|
79
82
|
ddl = '''
|
|
80
83
|
CREATE TABLE posts (title TEXT, slug TEXT, body TEXT);
|
|
81
84
|
CREATE INDEX posts_slug ON posts(slug);
|
|
@@ -21,11 +21,11 @@ stdout works. With [`jq`](https://jqlang.org/):
|
|
|
21
21
|
name = "books"
|
|
22
22
|
ddl = "CREATE TABLE books (title TEXT, author TEXT, year INTEGER)"
|
|
23
23
|
glob = "books/*.json"
|
|
24
|
-
on-file = "jq -c '[{title, author, year}]'
|
|
24
|
+
on-file = "jq -c -n '[inputs | {title, author, year}]'"
|
|
25
25
|
```
|
|
26
26
|
|
|
27
|
-
|
|
28
|
-
|
|
27
|
+
The command runs once for the table, with every matched file's absolute path
|
|
28
|
+
appended as an argument, and prints one array for all of them — the
|
|
29
29
|
[command hook contract](../reference/hooks.md#on-file), which also covers
|
|
30
30
|
the argv splitting, working directory, stdout protocol, and timeout shared
|
|
31
31
|
by every hook.
|
|
@@ -45,7 +45,7 @@ dirsql query "SELECT title, author, year FROM books ORDER BY year" -c ./.dirsql.
|
|
|
45
45
|
|
|
46
46
|
The table's columns are exactly what the command emits, narrowed to the DDL —
|
|
47
47
|
`dirsql` adds nothing. To include the file's `path`, have the command emit it
|
|
48
|
-
(it has
|
|
48
|
+
(it has the path); dirsql will not merge it in for you.
|
|
49
49
|
|
|
50
50
|
## Multiple rows per file
|
|
51
51
|
|
|
@@ -57,7 +57,7 @@ row per line, slurp it:
|
|
|
57
57
|
name = "events"
|
|
58
58
|
ddl = "CREATE TABLE events (event TEXT, user TEXT)"
|
|
59
59
|
glob = "logs/*.jsonl"
|
|
60
|
-
on-file = "jq -c -s '.'
|
|
60
|
+
on-file = "jq -c -s '.'"
|
|
61
61
|
```
|
|
62
62
|
|
|
63
63
|
```bash
|
|
@@ -32,40 +32,43 @@ To get them, you need a parser.
|
|
|
32
32
|
|
|
33
33
|
## 2. Attach a parser with `--on-file`
|
|
34
34
|
|
|
35
|
-
Any program that reads
|
|
36
|
-
stdout is a parser. Here is a small one,
|
|
37
|
-
frontmatter:
|
|
35
|
+
Any program that reads the files named by its arguments and prints a **JSON
|
|
36
|
+
array of row objects** on stdout is a parser. Here is a small one,
|
|
37
|
+
`extract.py`, that reads each post's frontmatter:
|
|
38
38
|
|
|
39
39
|
```python
|
|
40
40
|
#!/usr/bin/env python3
|
|
41
41
|
import json, re, sys
|
|
42
42
|
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
)
|
|
49
|
-
|
|
43
|
+
rows = []
|
|
44
|
+
for path in sys.argv[1:]:
|
|
45
|
+
text = open(path, encoding="utf-8").read()
|
|
46
|
+
m = re.match(r"^---\n(.*?)\n---", text, re.DOTALL)
|
|
47
|
+
fields = dict(
|
|
48
|
+
(k.strip(), v.strip())
|
|
49
|
+
for k, _, v in (line.partition(":") for line in (m.group(1).splitlines() if m else []))
|
|
50
|
+
)
|
|
51
|
+
rows.append({"title": fields.get("title"), "author": fields.get("author")})
|
|
52
|
+
print(json.dumps(rows))
|
|
50
53
|
```
|
|
51
54
|
|
|
52
55
|
Point the path-table at it with `--on-file`:
|
|
53
56
|
|
|
54
57
|
```bash
|
|
55
58
|
dirsql query "SELECT title, author FROM './posts/*.md' ORDER BY title" \
|
|
56
|
-
--on-file 'python3 extract.py
|
|
59
|
+
--on-file 'python3 extract.py'
|
|
57
60
|
```
|
|
58
61
|
|
|
59
62
|
```json
|
|
60
63
|
[{"author":"Ada Lovelace","title":"Hello World"},{"author":"Alan Turing","title":"On Recursion"}]
|
|
61
64
|
```
|
|
62
65
|
|
|
63
|
-
Now the parser's output *is* the table. `--on-file` runs the command once
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
splitting, timeout, and
|
|
67
|
-
|
|
68
|
-
|
|
66
|
+
Now the parser's output *is* the table. `--on-file` runs the command once,
|
|
67
|
+
with every matched file's absolute path appended as an argument, under the
|
|
68
|
+
shared [`on-file` hook contract](../reference/hooks.md#on-file) (argv
|
|
69
|
+
splitting, timeout, and failure semantics all come from there). The stat
|
|
70
|
+
columns are no longer reachable — a parser that wants the path emits it, since
|
|
71
|
+
it already has the path. See
|
|
69
72
|
[Parsing rows with `--on-file`](../reference/path-tables.md#parsing-rows-with-on-file)
|
|
70
73
|
for the full behavior.
|
|
71
74
|
|
|
@@ -87,7 +90,7 @@ command in verbatim:
|
|
|
87
90
|
name = "posts"
|
|
88
91
|
ddl = "CREATE TABLE posts (title TEXT, author TEXT)"
|
|
89
92
|
glob = "posts/*.md"
|
|
90
|
-
on-file = "python3 extract.py
|
|
93
|
+
on-file = "python3 extract.py"
|
|
91
94
|
```
|
|
92
95
|
|
|
93
96
|
The `on-file` value is byte-for-byte the string you passed to `--on-file`. The
|
|
@@ -109,7 +112,7 @@ indexed on build, kept fresh by the watcher, survives restarts with
|
|
|
109
112
|
[many tables](./define-tables.md) — each with its own `on-file` — where the flag
|
|
110
113
|
gives every path-table one parser. In both spellings the table's columns are
|
|
111
114
|
exactly what the parser emits: `dirsql` merges no filesystem facts back on. A
|
|
112
|
-
row that needs the file's `path` emits it (the parser has
|
|
115
|
+
row that needs the file's `path` emits it (the parser has the path).
|
|
113
116
|
|
|
114
117
|
## Going further
|
|
115
118
|
|
|
@@ -17,7 +17,7 @@ prints each file's basename:
|
|
|
17
17
|
name = "files"
|
|
18
18
|
ddl = "CREATE TABLE files (basename TEXT)"
|
|
19
19
|
glob = "**/*"
|
|
20
|
-
on-file = '''sh -c 'printf "[{\"basename\":\"%s\"}
|
|
20
|
+
on-file = '''sh -c 'printf "["; sep=""; for p; do printf "%s{\"basename\":\"%s\"}" "$sep" "${p##*/}"; sep=","; done; printf "]"' sh'''
|
|
21
21
|
```
|
|
22
22
|
|
|
23
23
|
## 1. Open the stream
|
|
@@ -31,7 +31,7 @@ external-content index beside the row table and let the two triggers feed it:
|
|
|
31
31
|
[[table]]
|
|
32
32
|
name = "notes"
|
|
33
33
|
glob = "notes/**/*.md"
|
|
34
|
-
on-file = "python3 extract.py
|
|
34
|
+
on-file = "python3 extract.py"
|
|
35
35
|
ddl = '''
|
|
36
36
|
CREATE TABLE notes (slug TEXT, title TEXT, body TEXT);
|
|
37
37
|
|
|
@@ -150,7 +150,7 @@ inserted vector for the "embedding" column. Expected 8 dimensions but received 4
|
|
|
150
150
|
[[table]]
|
|
151
151
|
name = "notes"
|
|
152
152
|
glob = "notes/**/*.md"
|
|
153
|
-
on-file = "python3 extract.py
|
|
153
|
+
on-file = "python3 extract.py"
|
|
154
154
|
ddl = '''
|
|
155
155
|
CREATE TABLE notes (slug TEXT, title TEXT, body TEXT);
|
|
156
156
|
|
package/docs/howto/skip-files.md
CHANGED
|
@@ -25,7 +25,7 @@ ignore = ["notes/drafts/**", "**/*.tmp"]
|
|
|
25
25
|
name = "notes"
|
|
26
26
|
ddl = "CREATE TABLE notes (basename TEXT)"
|
|
27
27
|
glob = "notes/**/*"
|
|
28
|
-
on-file = '''sh -c 'printf "[{\"basename\":\"%s\"}
|
|
28
|
+
on-file = '''sh -c 'printf "["; sep=""; for p; do printf "%s{\"basename\":\"%s\"}" "$sep" "${p##*/}"; sep=","; done; printf "]"' sh'''
|
|
29
29
|
```
|
|
30
30
|
|
|
31
31
|
Patterns match against root-relative paths, the same way table globs do. An
|
|
@@ -73,11 +73,11 @@ Both hook command styles from the [hook contract](../reference/hooks.md) work
|
|
|
73
73
|
in a plugin fragment:
|
|
74
74
|
|
|
75
75
|
- **Console scripts** — a `bin`-style entry point your package installs on
|
|
76
|
-
`PATH` (`embed-file {
|
|
76
|
+
`PATH` (`embed-file {root}`). Recommended for published plugins: the command
|
|
77
77
|
is bound to your package's interpreter and dependencies, and it is
|
|
78
78
|
language-neutral (the fragment names a command, not a Python file).
|
|
79
79
|
- **Relative scripts** — a path resolved against the fragment's own directory
|
|
80
|
-
(`uv run python embed.py {
|
|
80
|
+
(`uv run python embed.py {root}`). Convenient while developing the plugin
|
|
81
81
|
in-tree.
|
|
82
82
|
|
|
83
83
|
Two facts from the [execution contract](../reference/hooks.md#execution-contract)
|
|
@@ -85,12 +85,12 @@ matter most for a published plugin:
|
|
|
85
85
|
|
|
86
86
|
- **A hook runs in its declaring config's directory.** For a plugin that is
|
|
87
87
|
the installed fragment's directory — inside **site-packages**. That is a
|
|
88
|
-
read-only, shared location: **run from it, never write to it.**
|
|
89
|
-
absolute
|
|
90
|
-
|
|
91
|
-
|
|
88
|
+
read-only, shared location: **run from it, never write to it.** Read the
|
|
89
|
+
matched files from the absolute paths the hook appends as arguments, and use
|
|
90
|
+
[`{root}`](../reference/hooks.md#on-file) to reach the user's project
|
|
91
|
+
directory. Write any cache to `{root}` or a real cache dir,
|
|
92
92
|
never next to the fragment.
|
|
93
|
-
-
|
|
93
|
+
- **The paths are absolute and `{root}` is the index root**, so a command is
|
|
94
94
|
self-sufficient from any working directory — it works whether the plugin
|
|
95
95
|
lives in the project or in site-packages.
|
|
96
96
|
|
|
@@ -112,28 +112,30 @@ entrypoint = "sqlite3_vec_init"
|
|
|
112
112
|
name = "notes"
|
|
113
113
|
ddl = "CREATE TABLE notes (path TEXT, text TEXT, embedding TEXT)"
|
|
114
114
|
glob = "notes/*.md"
|
|
115
|
-
on-file = "uv run --with model2vec python embed.py {
|
|
115
|
+
on-file = "uv run --with model2vec python embed.py {root}"
|
|
116
116
|
```
|
|
117
117
|
|
|
118
|
-
`embed.py` turns
|
|
119
|
-
dirsql injects no columns, so the script emits the path itself,
|
|
120
|
-
`{
|
|
118
|
+
`embed.py` turns each file into one row carrying its path, text, and
|
|
119
|
+
embedding. dirsql injects no columns, so the script emits the path itself,
|
|
120
|
+
from the `{root}` the hook passes in and the paths it appends:
|
|
121
121
|
|
|
122
122
|
```python
|
|
123
|
-
"""Embed
|
|
123
|
+
"""Embed each file's text; print one dirsql row array on stdout."""
|
|
124
124
|
import json
|
|
125
125
|
import os
|
|
126
126
|
import sys
|
|
127
127
|
|
|
128
128
|
from model2vec import StaticModel
|
|
129
129
|
|
|
130
|
-
|
|
131
|
-
|
|
130
|
+
root, paths = sys.argv[1], sys.argv[2:]
|
|
131
|
+
texts = [open(path, encoding="utf-8").read() for path in paths]
|
|
132
132
|
model = StaticModel.from_pretrained("minishlab/potion-base-8M")
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
133
|
+
rows = [
|
|
134
|
+
{"path": os.path.relpath(path, root), "text": text,
|
|
135
|
+
"embedding": json.dumps([round(float(x), 6) for x in vector])}
|
|
136
|
+
for path, text, vector in zip(paths, texts, model.encode(texts))
|
|
137
|
+
]
|
|
138
|
+
print(json.dumps(rows))
|
|
137
139
|
```
|
|
138
140
|
|
|
139
141
|
The relative `embed.py` above resolves against the fragment directory, which
|
package/docs/reference/cli.md
CHANGED
|
@@ -370,12 +370,12 @@ array of row objects) instead of the stat columns:
|
|
|
370
370
|
|
|
371
371
|
```sh
|
|
372
372
|
dirsql query "SELECT title, author FROM './posts/*.md'" \
|
|
373
|
-
--on-file 'extract.py
|
|
373
|
+
--on-file 'extract.py'
|
|
374
374
|
```
|
|
375
375
|
|
|
376
376
|
The command follows the [`on-file` hook contract](./hooks.md#on-file) — argv
|
|
377
|
-
splitting, `{
|
|
378
|
-
timeout. The parser's output is the whole schema; the stat columns are not
|
|
377
|
+
splitting, the `{root}` placeholder, every matched path as a trailing
|
|
378
|
+
argument, the failure semantics, and the timeout. The parser's output is the whole schema; the stat columns are not
|
|
379
379
|
reachable on a parsed path-table. `--on-file` may be given **at most once** (a
|
|
380
380
|
repeat is an error pointing at config files) and never touches config-declared
|
|
381
381
|
tables. It is a `query`-only flag — server mode rejects it as an unknown
|
|
@@ -12,8 +12,8 @@ A named [`[[table]]`](./config.md#table) — or an SDK
|
|
|
12
12
|
[`on-file` hook](./hooks.md#on-file) emits, narrowed to the columns the DDL
|
|
13
13
|
declares. dirsql adds nothing on top: no `path`, no `size`, no value derived
|
|
14
14
|
from the filename. A hook that wants any of those computes them itself. The
|
|
15
|
-
hook receives
|
|
16
|
-
|
|
15
|
+
hook receives every matched file's path as an argument (an SDK `on_file`
|
|
16
|
+
callback receives one path) and may stat or read the files however it likes.
|
|
17
17
|
|
|
18
18
|
The hook is **required**. A `[[table]]` with no `on-file` is a
|
|
19
19
|
[config-load error](./config.md#parse-errors): with nothing supplying columns,
|
|
@@ -58,12 +58,12 @@ ORDER BY mtime DESC;
|
|
|
58
58
|
[Attaching a parser](./path-tables.md#parsing-rows-with-on-file) to a
|
|
59
59
|
path-table with `--on-file` **replaces** these stat columns with the parser's
|
|
60
60
|
own output — the two modes stay cleanly separate, exactly as for a named
|
|
61
|
-
table. A parser that wants the path emits it; it has
|
|
61
|
+
table. A parser that wants the path emits it; it has the paths.
|
|
62
62
|
|
|
63
63
|
## Deriving columns from the path
|
|
64
64
|
|
|
65
65
|
To turn path segments (an author, a year, a thread id) into columns, a hook
|
|
66
|
-
splits
|
|
66
|
+
splits the path and emits the pieces — the same as any other column it
|
|
67
67
|
produces. dirsql does not do this for you: a `{name}` segment in a glob is
|
|
68
68
|
rewritten to `*` and matches a single path segment, but captures no value.
|
|
69
69
|
[Derive columns from file paths](../howto/columns-from-paths.md) walks a
|
package/docs/reference/config.md
CHANGED
|
@@ -221,12 +221,12 @@ what its required `on-file` command emits — dirsql injects nothing (see
|
|
|
221
221
|
| `name` | yes | The table's SQL name — the name you query it by. Declared, never derived from `ddl`: dirsql does not read the DDL text. The `ddl` must create a table by this name; if it doesn't, loading fails. |
|
|
222
222
|
| `ddl` | yes | A SQL batch, run verbatim — any number of statements. It must create a table called `name`; that table holds the file rows, and only the columns it declares are kept (keys the `on-file` command emits that are not declared are dropped). The rest of the batch is yours: indexes, virtual tables, triggers. See [Batch `ddl`](#batch-ddl). |
|
|
223
223
|
| `glob` | yes | Glob pattern matched against root-relative paths. Every table whose glob matches a file receives that file's rows — a file can populate multiple tables. A `{name}` segment is rewritten to `*` (it matches one path segment but captures nothing). |
|
|
224
|
-
| `on-file` | **yes** | A command run once per matched file; its stdout (
|
|
224
|
+
| `on-file` | **yes** | A command run once per table, with every matched file's absolute path appended as a trailing argument; its stdout (one JSON array of row objects) is the table's rows. Must be non-empty. A `[[table]]` with no `on-file` is a load error (see [parse errors](#parse-errors)). See [Command hooks](./hooks.md#on-file). |
|
|
225
225
|
| `strict` | no (default `false`) | When `true`, rows whose keys do not exactly match the declared columns are rejected with an error: extra keys error, and every declared column must be supplied by the `on-file` output. When `false`, extra keys are dropped and missing columns become `NULL`. |
|
|
226
226
|
|
|
227
227
|
`on-file` is required because a table's rows come from nowhere else. dirsql
|
|
228
228
|
does not read file contents or merge filesystem facts on your behalf: the
|
|
229
|
-
command reads the
|
|
229
|
+
command reads the files (it receives their paths as arguments) and prints the rows, and those
|
|
230
230
|
rows — filtered to the DDL — are the table. For plain stat columns with no
|
|
231
231
|
command, query the path directly with a [path-table](./path-tables.md)
|
|
232
232
|
instead of declaring a table.
|
|
@@ -236,13 +236,13 @@ instead of declaring a table.
|
|
|
236
236
|
name = "comments"
|
|
237
237
|
ddl = "CREATE TABLE comments (path TEXT, author TEXT, body TEXT)"
|
|
238
238
|
glob = "_comments/*/*.jsonl"
|
|
239
|
-
on-file = "jq -c -s '.'
|
|
239
|
+
on-file = "jq -c -s '.'"
|
|
240
240
|
|
|
241
241
|
[[table]]
|
|
242
242
|
name = "papers"
|
|
243
243
|
ddl = "CREATE TABLE papers (paper_id TEXT, title TEXT)"
|
|
244
244
|
glob = "**/meta.json"
|
|
245
|
-
on-file = "uv run python extract_papers.py
|
|
245
|
+
on-file = "uv run python extract_papers.py"
|
|
246
246
|
strict = true
|
|
247
247
|
```
|
|
248
248
|
|
|
@@ -255,7 +255,7 @@ statement:
|
|
|
255
255
|
[[table]]
|
|
256
256
|
name = "messages"
|
|
257
257
|
glob = "sessions/*/messages/*.json"
|
|
258
|
-
on-file = "jq -c
|
|
258
|
+
on-file = "jq -c -s add"
|
|
259
259
|
ddl = '''
|
|
260
260
|
CREATE TABLE messages (session TEXT, idx INT, role TEXT, text TEXT);
|
|
261
261
|
CREATE INDEX messages_session ON messages(session);
|
|
@@ -419,11 +419,11 @@ timeout = "600s"
|
|
|
419
419
|
name = "comments"
|
|
420
420
|
ddl = "CREATE TABLE comments (author TEXT, body TEXT)"
|
|
421
421
|
glob = "_comments/*/*.jsonl"
|
|
422
|
-
on-file = "jq -c -s '.'
|
|
422
|
+
on-file = "jq -c -s '.'"
|
|
423
423
|
|
|
424
424
|
[[table]]
|
|
425
425
|
name = "documents"
|
|
426
426
|
ddl = "CREATE TABLE documents (title TEXT, summary TEXT)"
|
|
427
427
|
glob = "**/index.md"
|
|
428
|
-
on-file = "uv run python extract_doc.py
|
|
428
|
+
on-file = "uv run python extract_doc.py"
|
|
429
429
|
```
|
package/docs/reference/hooks.md
CHANGED
|
@@ -10,7 +10,7 @@ external command under the execution contract below.
|
|
|
10
10
|
|
|
11
11
|
The command string is split into an argv with shell-like quoting: whitespace
|
|
12
12
|
separates arguments, and single or double quotes group them (so
|
|
13
|
-
`sh -c 'grep foo
|
|
13
|
+
`sh -c 'grep foo "$@" | sort' sh` keeps the quoted script as a single
|
|
14
14
|
argument). **No shell is invoked** — there is no globbing, piping, `$VAR`
|
|
15
15
|
expansion, or `&&`/`;` chaining. To get shell features, ask for a shell
|
|
16
16
|
explicitly with `sh -c '…'`.
|
|
@@ -65,7 +65,7 @@ bound a hook, make the bound part of the command by wrapping it in
|
|
|
65
65
|
`timeout(1)`:
|
|
66
66
|
|
|
67
67
|
```toml
|
|
68
|
-
on-file = "timeout 30 my-extractor
|
|
68
|
+
on-file = "timeout 30 my-extractor"
|
|
69
69
|
```
|
|
70
70
|
|
|
71
71
|
When the wrapper kills an overrunning command, the run exits non-zero and
|
|
@@ -117,9 +117,14 @@ flag is the inline form, the config key the declared form (see
|
|
|
117
117
|
[Parse your files into columns](../howto/parse-files-into-columns.md)). In both
|
|
118
118
|
spellings the table's columns are exactly what the command emits, narrowed to
|
|
119
119
|
the DDL — `dirsql` injects no filesystem facts either way. A command that wants
|
|
120
|
-
the path or stat metadata emits it (it has
|
|
120
|
+
the path or stat metadata emits it (it has the paths).
|
|
121
121
|
|
|
122
122
|
| Placeholder | Value |
|
|
123
123
|
|---|---|
|
|
124
|
-
| `{
|
|
125
|
-
|
|
124
|
+
| `{root}` | The index root directory. Derive a root-relative path with `relpath(path, {root})`. |
|
|
125
|
+
|
|
126
|
+
The matched files' **absolute** paths are not placeholders: they are appended
|
|
127
|
+
to the command as trailing arguments, after everything written in the
|
|
128
|
+
command, so `on-file = "extract.py"` receives them as `sys.argv[1:]` — one run
|
|
129
|
+
per table, self-sufficient from any working directory. A command that still
|
|
130
|
+
spells `{path}` is rejected at startup.
|
|
@@ -139,18 +139,18 @@ the `dirsql query` flag [`--on-file`](/reference/cli):
|
|
|
139
139
|
|
|
140
140
|
```sh
|
|
141
141
|
dirsql query "SELECT title, author FROM './posts/*.md'" \
|
|
142
|
-
--on-file 'extract.py
|
|
142
|
+
--on-file 'extract.py'
|
|
143
143
|
```
|
|
144
144
|
|
|
145
|
-
The command runs once
|
|
146
|
-
|
|
147
|
-
splitting, same `{
|
|
148
|
-
the table:
|
|
145
|
+
The command runs once, with every matched file's absolute path appended as an
|
|
146
|
+
argument, and prints one JSON array of row objects, exactly like a declared
|
|
147
|
+
table's [`on-file` hook](/reference/hooks) — same argv splitting, same `{root}`
|
|
148
|
+
placeholder, same timeout. Its output *is* the table:
|
|
149
149
|
|
|
150
150
|
- **The parser supplies the whole schema.** Columns are inferred from the keys
|
|
151
151
|
across the emitted rows. The stat columns (`path`, `size`, …) are **not**
|
|
152
152
|
reachable on a parsed path-table — a parser that wants the path emits it (it
|
|
153
|
-
has
|
|
153
|
+
has the paths). The two modes stay cleanly separate.
|
|
154
154
|
- **Failures are isolated per file.** A file whose parser fails (spawn, non-zero
|
|
155
155
|
exit, timeout, or no output) or whose output is not a JSON array of rows
|
|
156
156
|
contributes no rows; a one-line warning naming the file goes to stderr and the
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "dirsql",
|
|
3
|
-
"version": "0.4.
|
|
3
|
+
"version": "0.4.66",
|
|
4
4
|
"description": "Ephemeral SQL index over a local directory",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"repository": "https://github.com/thekevinscott/dirsql",
|
|
@@ -221,10 +221,10 @@
|
|
|
221
221
|
]
|
|
222
222
|
},
|
|
223
223
|
"optionalDependencies": {
|
|
224
|
-
"@dirsql/lib-linux-x64-gnu": "0.4.
|
|
225
|
-
"@dirsql/lib-linux-arm64-gnu": "0.4.
|
|
226
|
-
"@dirsql/lib-darwin-x64": "0.4.
|
|
227
|
-
"@dirsql/lib-darwin-arm64": "0.4.
|
|
228
|
-
"@dirsql/lib-win32-x64-msvc": "0.4.
|
|
224
|
+
"@dirsql/lib-linux-x64-gnu": "0.4.66",
|
|
225
|
+
"@dirsql/lib-linux-arm64-gnu": "0.4.66",
|
|
226
|
+
"@dirsql/lib-darwin-x64": "0.4.66",
|
|
227
|
+
"@dirsql/lib-darwin-arm64": "0.4.66",
|
|
228
|
+
"@dirsql/lib-win32-x64-msvc": "0.4.66"
|
|
229
229
|
}
|
|
230
230
|
}
|