dirsql 0.3.64 → 0.3.66

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,221 +0,0 @@
1
- ---
2
- canonical: https://thekevinscott.github.io/dirsql/guide/querying
3
- ---
4
-
5
- # Querying
6
-
7
- > Online: <https://thekevinscott.github.io/dirsql/guide/querying>
8
-
9
- Once a `DirSQL` instance is created, the initial directory scan is complete and you can run SQL queries against the indexed data.
10
-
11
- ::: tip From the CLI
12
- The `dirsql` HTTP server exposes the same query interface over [`POST /query`](../cli/http-api.md#post-query). Send `{"sql": "..."}` and get back the same JSON array of row objects. See the [CLI section](../cli/) for the full server setup.
13
- :::
14
-
15
- ## Basic queries
16
-
17
- ::: code-group
18
-
19
- ```python [Python]
20
- # All rows from a table
21
- results = db.query("SELECT * FROM comments")
22
-
23
- # Filter with WHERE
24
- results = db.query("SELECT * FROM comments WHERE author = 'alice'")
25
-
26
- # Aggregations
27
- results = db.query("SELECT author, COUNT(*) as n FROM comments GROUP BY author")
28
-
29
- # JOINs across tables
30
- results = db.query("""
31
- SELECT posts.title, authors.name
32
- FROM posts
33
- JOIN authors ON posts.author_id = authors.id
34
- """)
35
- ```
36
-
37
- ```rust [Rust]
38
- // All rows from a table
39
- let results = db.query("SELECT * FROM comments")?;
40
-
41
- // Filter with WHERE
42
- let results = db.query("SELECT * FROM comments WHERE author = 'alice'")?;
43
-
44
- // Aggregations
45
- let results = db.query("SELECT author, COUNT(*) as n FROM comments GROUP BY author")?;
46
-
47
- // JOINs across tables
48
- let results = db.query(
49
- "SELECT posts.title, authors.name \
50
- FROM posts JOIN authors ON posts.author_id = authors.id"
51
- )?;
52
- ```
53
-
54
- ```typescript [TypeScript]
55
- // All rows from a table
56
- const results = await db.query('SELECT * FROM comments');
57
-
58
- // Filter with WHERE
59
- const filtered = await db.query("SELECT * FROM comments WHERE author = 'alice'");
60
-
61
- // Aggregations
62
- const counts = await db.query('SELECT author, COUNT(*) as n FROM comments GROUP BY author');
63
-
64
- // JOINs across tables
65
- const joined = await db.query(`
66
- SELECT posts.title, authors.name
67
- FROM posts
68
- JOIN authors ON posts.author_id = authors.id
69
- `);
70
- ```
71
-
72
- :::
73
-
74
- Any valid SQLite **SELECT** works. The in-memory database supports the full SQLite dialect including subqueries, CTEs, window functions, and aggregate functions. See [Read-only queries](#read-only-queries) below for why write statements (`INSERT`, `UPDATE`, `DELETE`, `DROP`, etc.) are rejected.
75
-
76
- ## Return format
77
-
78
- `query()` returns a list of dicts (Python), a `Vec<HashMap>` (Rust), or an array of objects (TypeScript). Each entry maps column names to values.
79
-
80
- ::: code-group
81
-
82
- ```python [Python]
83
- results = db.query("SELECT title, author FROM posts")
84
- # [
85
- # {"title": "Hello World", "author": "alice"},
86
- # {"title": "Second Post", "author": "bob"},
87
- # ]
88
- ```
89
-
90
- ```rust [Rust]
91
- let results = db.query("SELECT title, author FROM posts")?;
92
- // Vec<HashMap<String, Value>>
93
- // [{"title": "Hello World", "author": "alice"}, ...]
94
- ```
95
-
96
- ```typescript [TypeScript]
97
- const results = await db.query('SELECT title, author FROM posts');
98
- // [
99
- // { title: 'Hello World', author: 'alice' },
100
- // { title: 'Second Post', author: 'bob' },
101
- // ]
102
- ```
103
-
104
- :::
105
-
106
- SQLite types map back to language types:
107
-
108
- | SQLite type | Python | TypeScript | Rust |
109
- |-------------|---------|------------|------------------|
110
- | TEXT | `str` | `string` | `Value::Text` |
111
- | INTEGER | `int` | `number` | `Value::Integer` |
112
- | REAL | `float` | `number` | `Value::Real` |
113
- | BLOB | `bytes` | `Buffer` | `Value::Blob` |
114
- | NULL | `None` | `null` | `Value::Null` |
115
-
116
- ## Internal columns
117
-
118
- `dirsql` adds internal tracking columns (`_dirsql_file_path`, `_dirsql_row_index`) to each table for file-change diffing. These columns are automatically excluded from `SELECT *` results, so day-to-day queries don't need to account for them.
119
-
120
- If you want to know which file a row came from, you can name the tracking columns explicitly in the projection:
121
-
122
- ::: code-group
123
-
124
- ```python [Python]
125
- rows = db.query("SELECT title, _dirsql_file_path FROM posts")
126
- # [{"title": "Hello World", "_dirsql_file_path": "posts/hello.json"}, ...]
127
- ```
128
-
129
- ```rust [Rust]
130
- let rows = db.query("SELECT title, _dirsql_file_path FROM posts")?;
131
- // [{"title": "Hello World", "_dirsql_file_path": "posts/hello.json"}, ...]
132
- ```
133
-
134
- ```typescript [TypeScript]
135
- const rows = await db.query('SELECT title, _dirsql_file_path FROM posts');
136
- // [{ title: 'Hello World', _dirsql_file_path: 'posts/hello.json' }, ...]
137
- ```
138
-
139
- :::
140
-
141
- Tracking columns are only returned when named explicitly — `SELECT *` continues to exclude them.
142
-
143
- ## Read-only queries
144
-
145
- `query()` accepts only read-only statements. Each statement is prepared on SQLite and then classified via `sqlite3_stmt_readonly`; anything SQLite itself flags as a write — `INSERT`, `UPDATE`, `DELETE`, `DROP`, `CREATE`, `ALTER`, `REPLACE`, `VACUUM`, `ANALYZE`, etc. — is rejected before any rows are produced.
146
-
147
- This keeps the in-memory index consistent with the on-disk files that back it. Mutations only happen through the watcher/indexer pipeline: to change data, edit the underlying file and let the watcher re-extract rows.
148
-
149
- ::: code-group
150
-
151
- ```python [Python]
152
- # Raises a RuntimeError; the index is unchanged.
153
- db.query("DELETE FROM posts")
154
- ```
155
-
156
- ```rust [Rust]
157
- // Returns DirSqlError::WriteForbidden; the index is unchanged.
158
- let err = db.query("DELETE FROM posts").unwrap_err();
159
- assert!(matches!(err, dirsql::DirSqlError::WriteForbidden));
160
- ```
161
-
162
- ```typescript [TypeScript]
163
- // Rejects with an Error whose message explains writes are not accepted.
164
- // `db.query` is async, so assert on the rejected promise.
165
- await expect(db.query('DELETE FROM posts')).rejects.toThrow(/read-only/i);
166
- ```
167
-
168
- :::
169
-
170
- ## Error handling
171
-
172
- Invalid SQL raises an exception:
173
-
174
- ::: code-group
175
-
176
- ```python [Python]
177
- try:
178
- db.query("NOT VALID SQL")
179
- except Exception as e:
180
- print(f"Query error: {e}")
181
- ```
182
-
183
- ```rust [Rust]
184
- match db.query("NOT VALID SQL") {
185
- Ok(results) => println!("{:?}", results),
186
- Err(e) => eprintln!("Query error: {}", e),
187
- }
188
- ```
189
-
190
- ```typescript [TypeScript]
191
- try {
192
- await db.query('NOT VALID SQL');
193
- } catch (e) {
194
- console.error(`Query error: ${e}`);
195
- }
196
- ```
197
-
198
- :::
199
-
200
- ## Empty results
201
-
202
- Queries that match no rows return an empty collection:
203
-
204
- ::: code-group
205
-
206
- ```python [Python]
207
- results = db.query("SELECT * FROM posts WHERE author = 'nobody'")
208
- assert results == []
209
- ```
210
-
211
- ```rust [Rust]
212
- let results = db.query("SELECT * FROM posts WHERE author = 'nobody'")?;
213
- assert!(results.is_empty());
214
- ```
215
-
216
- ```typescript [TypeScript]
217
- const results = await db.query("SELECT * FROM posts WHERE author = 'nobody'");
218
- console.assert(results.length === 0);
219
- ```
220
-
221
- :::
@@ -1,269 +0,0 @@
1
- ---
2
- canonical: https://thekevinscott.github.io/dirsql/guide/tables
3
- ---
4
-
5
- # Defining Tables
6
-
7
- > Online: <https://thekevinscott.github.io/dirsql/guide/tables>
8
-
9
- Each table in `dirsql` maps a set of files to rows in an in-memory SQLite table. A table definition has three parts: DDL, a glob pattern, and an extract function.
10
-
11
- ## Table constructor
12
-
13
- ::: code-group
14
-
15
- ```python [Python]
16
- from dirsql import Table
17
-
18
- table = Table(
19
- ddl="CREATE TABLE comments (id TEXT, body TEXT, author TEXT)",
20
- glob="comments/**/index.jsonl",
21
- extract=lambda path: [
22
- {"id": "...", "body": "...", "author": "..."}
23
- ],
24
- )
25
- ```
26
-
27
- ```rust [Rust]
28
- use dirsql::{Table, Value};
29
- use std::collections::HashMap;
30
-
31
- let table = Table::new(
32
- "CREATE TABLE comments (id TEXT, body TEXT, author TEXT)",
33
- "comments/**/index.jsonl",
34
- |_path| {
35
- let mut row: HashMap<String, Value> = HashMap::new();
36
- row.insert("id".into(), Value::Text("...".into()));
37
- row.insert("body".into(), Value::Text("...".into()));
38
- row.insert("author".into(), Value::Text("...".into()));
39
- vec![row]
40
- },
41
- );
42
- ```
43
-
44
- ```typescript [TypeScript]
45
- import type { TableDef } from 'dirsql';
46
-
47
- const table: TableDef = {
48
- ddl: 'CREATE TABLE comments (id TEXT, body TEXT, author TEXT)',
49
- glob: 'comments/**/index.jsonl',
50
- extract: (_path) => [
51
- { id: '...', body: '...', author: '...' },
52
- ],
53
- };
54
- ```
55
-
56
- :::
57
-
58
- All three arguments are keyword-only (in Python). In Rust they are positional to `Table::new`. In TypeScript a table is a plain `TableDef` object literal — the TS SDK exports the `TableDef` type (not a class).
59
-
60
- ### `ddl`
61
-
62
- A SQLite `CREATE TABLE` statement. This defines the schema of the table. `dirsql` executes this DDL directly against the in-memory database, so any valid SQLite column types and constraints work.
63
-
64
- ```python
65
- # Simple text columns
66
- ddl="CREATE TABLE notes (title TEXT, body TEXT)"
67
-
68
- # Typed columns
69
- ddl="CREATE TABLE metrics (name TEXT, value REAL, count INTEGER)"
70
-
71
- # With constraints
72
- ddl="CREATE TABLE items (id TEXT PRIMARY KEY, name TEXT NOT NULL)"
73
- ```
74
-
75
- The table name is parsed from the DDL and must resolve to a valid SQLite identifier.
76
-
77
- ### `glob`
78
-
79
- A glob pattern that determines which files feed into this table. Matched relative to the root directory passed to `DirSQL`.
80
-
81
- ```python
82
- glob="*.json" # JSON files in root only
83
- glob="**/*.json" # JSON files at any depth
84
- glob="comments/**/index.jsonl" # JSONL files in comment subdirectories
85
- glob="data/*.csv" # CSV files in data/
86
- ```
87
-
88
- Glob syntax follows standard Unix globbing rules. `**` matches any number of directory levels.
89
-
90
- ### `extract`
91
-
92
- A callable `(path: str) -> list[dict]` that converts a file into rows.
93
-
94
- - `path` is the path of the matched file, **relative to the scan root** (or absolute when the `root` passed to `DirSQL` is absolute)
95
- - Return a list of dicts, where each dict maps column names to values
96
- - Return an empty list to skip a file
97
-
98
- `dirsql` does not read file contents for you. If your extract needs the file
99
- body, read it inside the callback using `path`. Callbacks that derive columns
100
- only from the path (or that rely solely on the auto-injected filesystem-fact
101
- columns) never touch the file at all.
102
-
103
- ```python
104
- import json
105
-
106
- # Single-object JSON files: one row per file
107
- def extract(path):
108
- with open(path, encoding="utf-8") as f:
109
- return [json.loads(f.read())]
110
-
111
- # JSONL files: one row per line
112
- def extract(path):
113
- with open(path, encoding="utf-8") as f:
114
- return [json.loads(line) for line in f]
115
-
116
- # Derive a value from the file path alone -- no file read
117
- import os
118
- extract = lambda path: [{"id": os.path.basename(os.path.dirname(path))}]
119
-
120
- # Conditionally skip files
121
- def extract(path):
122
- with open(path, encoding="utf-8") as f:
123
- data = json.loads(f.read())
124
- if data.get("draft"):
125
- return []
126
- return [data]
127
- ```
128
-
129
- ## Multiple tables
130
-
131
- Pass multiple `Table` definitions to index different file types into separate tables:
132
-
133
- ::: code-group
134
-
135
- ```python [Python]
136
- from dirsql import DirSQL, Table
137
- import json
138
-
139
- db = DirSQL(
140
- "./workspace",
141
- tables=[
142
- Table(
143
- ddl="CREATE TABLE posts (title TEXT, author_id TEXT)",
144
- glob="posts/*.json",
145
- extract=lambda path: [json.loads(open(path, encoding="utf-8").read())],
146
- ),
147
- Table(
148
- ddl="CREATE TABLE authors (id TEXT, name TEXT)",
149
- glob="authors/*.json",
150
- extract=lambda path: [json.loads(open(path, encoding="utf-8").read())],
151
- ),
152
- ],
153
- )
154
- ```
155
-
156
- ```rust [Rust]
157
- use dirsql::{DirSQL, Table, Value};
158
- use std::collections::HashMap;
159
-
160
- // See `row_from_json` in getting-started.md for a reusable helper.
161
- fn row_from_json(raw: &str) -> HashMap<String, Value> {
162
- let v: serde_json::Value = serde_json::from_str(raw).unwrap();
163
- let serde_json::Value::Object(obj) = v else { return HashMap::new() };
164
- obj.into_iter()
165
- .map(|(k, val)| {
166
- let v = match val {
167
- serde_json::Value::String(s) => Value::Text(s),
168
- serde_json::Value::Number(n) => n
169
- .as_i64()
170
- .map(Value::Integer)
171
- .unwrap_or_else(|| Value::Real(n.as_f64().unwrap_or(0.0))),
172
- serde_json::Value::Bool(b) => Value::Integer(b as i64),
173
- serde_json::Value::Null => Value::Null,
174
- other => Value::Text(other.to_string()),
175
- };
176
- (k, v)
177
- })
178
- .collect()
179
- }
180
-
181
- let db = DirSQL::new(
182
- "./workspace",
183
- vec![
184
- Table::new(
185
- "CREATE TABLE posts (title TEXT, author_id TEXT)",
186
- "posts/*.json",
187
- |path| vec![row_from_json(&std::fs::read_to_string(path).unwrap())],
188
- ),
189
- Table::new(
190
- "CREATE TABLE authors (id TEXT, name TEXT)",
191
- "authors/*.json",
192
- |path| vec![row_from_json(&std::fs::read_to_string(path).unwrap())],
193
- ),
194
- ],
195
- )?;
196
- ```
197
-
198
- ```typescript [TypeScript]
199
- import { DirSQL, type TableDef } from 'dirsql';
200
- import { readFileSync } from 'node:fs';
201
-
202
- const tables: TableDef[] = [
203
- {
204
- ddl: 'CREATE TABLE posts (title TEXT, author_id TEXT)',
205
- glob: 'posts/*.json',
206
- extract: (path) => [JSON.parse(readFileSync(path, 'utf8'))],
207
- },
208
- {
209
- ddl: 'CREATE TABLE authors (id TEXT, name TEXT)',
210
- glob: 'authors/*.json',
211
- extract: (path) => [JSON.parse(readFileSync(path, 'utf8'))],
212
- },
213
- ];
214
-
215
- const db = new DirSQL({ root: './workspace', tables });
216
- ```
217
-
218
- :::
219
-
220
- Each table has its own glob and extract function. A file can only match one table (the first matching glob wins).
221
-
222
- ## Ignore patterns
223
-
224
- Use the `ignore` parameter to exclude paths from all tables:
225
-
226
- ::: code-group
227
-
228
- ```python [Python]
229
- db = DirSQL(
230
- "./workspace",
231
- ignore=["**/node_modules/**", "**/.git/**"],
232
- tables=[...],
233
- )
234
- ```
235
-
236
- ```rust [Rust]
237
- let db = DirSQL::with_ignore(
238
- "./workspace",
239
- vec![/* tables */],
240
- vec!["**/node_modules/**", "**/.git/**"],
241
- )?;
242
- ```
243
-
244
- ```typescript [TypeScript]
245
- const db = new DirSQL({
246
- root: './workspace',
247
- tables: [/* tables */],
248
- ignore: ['**/node_modules/**', '**/.git/**'],
249
- });
250
- ```
251
-
252
- :::
253
-
254
- Ignore patterns are applied before glob matching. Any file matching an ignore pattern is skipped regardless of table globs.
255
-
256
- ## Supported value types
257
-
258
- The extract function can return these types, which map to SQLite types:
259
-
260
- | SQLite type | Python | TypeScript | Rust |
261
- |---------------|---------|-------------------------|------------------|
262
- | TEXT | `str` | `string` | `Value::Text` |
263
- | INTEGER | `int` | integral `number` | `Value::Integer` |
264
- | REAL | `float` | fractional `number` | `Value::Real` |
265
- | INTEGER (0/1) | `bool` | `boolean` | `Value::Integer` |
266
- | BLOB | `bytes` | `Buffer` / `Uint8Array` | `Value::Blob` |
267
- | NULL | `None` | `null` / `undefined` | `Value::Null` |
268
-
269
- Any other type is converted to its string representation (`str()` in Python, string coercion in TypeScript).