kbforge-okfquery 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- kbforge_okfquery-0.1.0.dist-info/METADATA +244 -0
- kbforge_okfquery-0.1.0.dist-info/RECORD +9 -0
- kbforge_okfquery-0.1.0.dist-info/WHEEL +4 -0
- kbforge_okfquery-0.1.0.dist-info/entry_points.txt +2 -0
- okfquery/__init__.py +14 -0
- okfquery/cli.py +144 -0
- okfquery/load.py +132 -0
- okfquery/parse.py +252 -0
- okfquery/schema.py +51 -0
|
@@ -0,0 +1,244 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: kbforge-okfquery
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: SQL over an OKF v0.2 bundle, via DuckDB
|
|
5
|
+
Author-email: Qing <qingye779@gmail.com>
|
|
6
|
+
License: MIT
|
|
7
|
+
Keywords: duckdb,knowledge-base,okf,open-knowledge-format,sql
|
|
8
|
+
Requires-Python: >=3.12
|
|
9
|
+
Requires-Dist: duckdb>=1.1
|
|
10
|
+
Requires-Dist: pytz>=2024.1
|
|
11
|
+
Requires-Dist: pyyaml>=6
|
|
12
|
+
Description-Content-Type: text/markdown
|
|
13
|
+
|
|
14
|
+
# okfquery
|
|
15
|
+
|
|
16
|
+
SQL over an [OKF v0.2](https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md)
|
|
17
|
+
bundle, via DuckDB. `okfquery` loads a published bundle into an in-memory DuckDB
|
|
18
|
+
connection and hands the connection back — no daemon, no cache, no committed
|
|
19
|
+
artifact. It depends on no kbforge code at runtime, so it reads any OKF v0.2
|
|
20
|
+
bundle, kbforge-built or not.
|
|
21
|
+
|
|
22
|
+
## Install and use
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
uv pip install kbforge-okfquery # or: pip install kbforge-okfquery
|
|
26
|
+
okfquery query "select count(*) from concepts" --bundle path/to/bundle
|
|
27
|
+
okfquery shell --bundle path/to/bundle
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
`query` runs one SQL statement and prints the result (`--format table|json|csv`).
|
|
31
|
+
`shell` writes a temporary `.duckdb` file, opens the `duckdb` CLI on it, and
|
|
32
|
+
deletes the file when the shell exits.
|
|
33
|
+
|
|
34
|
+
## Schema
|
|
35
|
+
|
|
36
|
+
```
|
|
37
|
+
$ okfquery schema
|
|
38
|
+
CREATE TABLE concepts (
|
|
39
|
+
path VARCHAR NOT NULL,
|
|
40
|
+
type VARCHAR,
|
|
41
|
+
title VARCHAR,
|
|
42
|
+
description VARCHAR,
|
|
43
|
+
generated_by VARCHAR,
|
|
44
|
+
-- TIMESTAMPTZ, never TIMESTAMP. Not because a naive column would corrupt
|
|
45
|
+
-- the instant on this package's load path -- `load` binds an aware Python
|
|
46
|
+
-- datetime through executemany, so the instant survives a naive column
|
|
47
|
+
-- fine. What a naive column loses is the type: values come back with no
|
|
48
|
+
-- tzinfo, so every comparison against now() (itself TIMESTAMPTZ) needs a
|
|
49
|
+
-- cast, and the column asserts UTC by convention with nothing recording
|
|
50
|
+
-- that it does. §4.4 law 4 exists to force an *aware* stamp; a column that
|
|
51
|
+
-- cannot hold one throws away exactly the property the law buys.
|
|
52
|
+
generated_at TIMESTAMPTZ,
|
|
53
|
+
facets JSON,
|
|
54
|
+
body VARCHAR
|
|
55
|
+
);
|
|
56
|
+
|
|
57
|
+
CREATE TABLE sources (
|
|
58
|
+
path VARCHAR NOT NULL,
|
|
59
|
+
-- A real column. kbforge's synthesize.assemble puts the OWNING anchor first
|
|
60
|
+
-- and grounding anchors after it, and that ordering is the only thing
|
|
61
|
+
-- telling owner from ground -- OKF has no field for it. Discarding ordinal
|
|
62
|
+
-- during unnest would destroy the signal.
|
|
63
|
+
ordinal INTEGER NOT NULL,
|
|
64
|
+
id VARCHAR,
|
|
65
|
+
resource VARCHAR,
|
|
66
|
+
content_hash VARCHAR
|
|
67
|
+
);
|
|
68
|
+
|
|
69
|
+
CREATE TABLE links (
|
|
70
|
+
path VARCHAR NOT NULL,
|
|
71
|
+
target VARCHAR NOT NULL
|
|
72
|
+
);
|
|
73
|
+
|
|
74
|
+
CREATE TABLE problems (
|
|
75
|
+
path VARCHAR NOT NULL,
|
|
76
|
+
kind VARCHAR NOT NULL,
|
|
77
|
+
detail VARCHAR NOT NULL
|
|
78
|
+
);
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
`path` is the bundle-relative concept path and the join key throughout.
|
|
82
|
+
`sources.ordinal = 0` is the owning source for a kbforge-produced bundle (the
|
|
83
|
+
convention `synthesize.assemble` writes) — on a foreign bundle it is merely
|
|
84
|
+
source order, with no error to tell you that. `facets` is one JSON column, open
|
|
85
|
+
by construction, so `facets->>'owner'` reaches whatever a connector attached.
|
|
86
|
+
`problems` is additive: a file that failed to parse still gets a `concepts` row
|
|
87
|
+
with NULLs for what could not be read, plus one `problems` row per distinct
|
|
88
|
+
failure — nothing a broken concept says gets silently dropped from an audit.
|
|
89
|
+
|
|
90
|
+
## Queries, run against the messy test fixture
|
|
91
|
+
|
|
92
|
+
All four queries below were run as shown against
|
|
93
|
+
`packages/okfquery/tests/fixtures/messy` — 9 `.md` files under `concepts/`, 7 of
|
|
94
|
+
them concepts (`concepts/index.md` and `concepts/log.md` are reserved and
|
|
95
|
+
fenceless, so they are skipped; `concepts/claims/index.md` is reserved-*named*
|
|
96
|
+
but carries frontmatter, so it counts).
|
|
97
|
+
|
|
98
|
+
**What is built on a given upstream document:**
|
|
99
|
+
|
|
100
|
+
```bash
|
|
101
|
+
$ okfquery query "
|
|
102
|
+
SELECT c.path, s.ordinal
|
|
103
|
+
FROM concepts c JOIN sources s USING (path)
|
|
104
|
+
WHERE s.resource = 'https://wiki/ok';
|
|
105
|
+
" --bundle packages/okfquery/tests/fixtures/messy
|
|
106
|
+
┌─────────────────────────┬─────────┐
|
|
107
|
+
│ path │ ordinal │
|
|
108
|
+
│ varchar │ int32 │
|
|
109
|
+
├─────────────────────────┼─────────┤
|
|
110
|
+
│ concepts/ok/overview.md │ 0 │
|
|
111
|
+
└─────────────────────────┴─────────┘
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
**Staleness by owning system** (`ordinal = 0` picks the owning source, not a
|
|
115
|
+
grounding one — see the ordinal note above):
|
|
116
|
+
|
|
117
|
+
```bash
|
|
118
|
+
$ okfquery query "
|
|
119
|
+
SELECT split_part(s.id, ':', 1) AS system,
|
|
120
|
+
count(*) AS concepts,
|
|
121
|
+
min(c.generated_at) AS oldest
|
|
122
|
+
FROM concepts c JOIN sources s USING (path)
|
|
123
|
+
WHERE s.ordinal = 0
|
|
124
|
+
GROUP BY 1 ORDER BY oldest;
|
|
125
|
+
" --bundle packages/okfquery/tests/fixtures/messy
|
|
126
|
+
┌─────────┬──────────┬──────────────────────────┐
|
|
127
|
+
│ system │ concepts │ oldest │
|
|
128
|
+
│ varchar │ int64 │ timestamp with time zone │
|
|
129
|
+
├─────────┼──────────┼──────────────────────────┤
|
|
130
|
+
│ wiki │ 3 │ 2026-08-23 00:00:00+00 │
|
|
131
|
+
└─────────┴──────────┴──────────────────────────┘
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
Only one system shows up on this fixture: most of the messy bundle's files
|
|
135
|
+
exist to exercise a `problems` kind and never got far enough to have a
|
|
136
|
+
`sources` entry at all, so they don't join into this query — which is correct,
|
|
137
|
+
not a bug in the query.
|
|
138
|
+
|
|
139
|
+
**Orphans — nothing links to them:**
|
|
140
|
+
|
|
141
|
+
```bash
|
|
142
|
+
$ okfquery query "
|
|
143
|
+
SELECT c.path FROM concepts c
|
|
144
|
+
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.target = c.path);
|
|
145
|
+
" --bundle packages/okfquery/tests/fixtures/messy
|
|
146
|
+
┌───────────────────────────────────┐
|
|
147
|
+
│ path │
|
|
148
|
+
│ varchar │
|
|
149
|
+
├───────────────────────────────────┤
|
|
150
|
+
│ concepts/badyaml/overview.md │
|
|
151
|
+
│ concepts/claims/index.md │
|
|
152
|
+
│ concepts/incomplete/overview.md │
|
|
153
|
+
│ concepts/nofence/overview.md │
|
|
154
|
+
│ concepts/notmapping/overview.md │
|
|
155
|
+
│ concepts/unterminated/overview.md │
|
|
156
|
+
└───────────────────────────────────┘
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
**Everything that failed to parse** (`--format csv` here, since the default
|
|
160
|
+
table renderer truncates long `detail` strings):
|
|
161
|
+
|
|
162
|
+
```bash
|
|
163
|
+
$ okfquery query "SELECT path, kind, detail FROM problems ORDER BY path, kind;" \
|
|
164
|
+
--bundle packages/okfquery/tests/fixtures/messy --format csv
|
|
165
|
+
path,kind,detail
|
|
166
|
+
concepts/badyaml/overview.md,invalid-yaml,frontmatter is not valid YAML: ParserError
|
|
167
|
+
concepts/incomplete/overview.md,bad-links,'links' entry 3 is not a string
|
|
168
|
+
concepts/incomplete/overview.md,bad-sources,'sources' entry 0 has no 'resource' (§5.1 requires one)
|
|
169
|
+
concepts/incomplete/overview.md,bad-timestamp,'generated.at' 'not-a-date' is not an ISO-8601 datetime
|
|
170
|
+
concepts/incomplete/overview.md,missing-required,"missing required OKF keys: title, description"
|
|
171
|
+
concepts/nofence/overview.md,no-frontmatter,file does not open a '---' frontmatter fence
|
|
172
|
+
concepts/notmapping/overview.md,frontmatter-not-mapping,"frontmatter parses to list, not a mapping"
|
|
173
|
+
concepts/unterminated/overview.md,unterminated-frontmatter,frontmatter fence is opened but never closed by a closing '---'
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
## The mirror, and why both obvious joins are wrong
|
|
177
|
+
|
|
178
|
+
`--mirror <path>` attaches a `mirror` view over a kbforge mirror directory
|
|
179
|
+
(`read_json_auto('<mirror>/*.json')`) — the same `CanonicalDocument` slots
|
|
180
|
+
`mirror.slot_key` writes, `anchor` struct included. It answers a different
|
|
181
|
+
question than the bundle alone can: not "what does this concept say" but "does
|
|
182
|
+
it still say what its source currently does."
|
|
183
|
+
|
|
184
|
+
**Both of the joins you'd reach for first are lossy, in different ways:**
|
|
185
|
+
|
|
186
|
+
- `sources.id = mirror.doc_id` holds for every shipped connector, but
|
|
187
|
+
`synthesize._source_entry` builds `id` from `anchor.system`/`anchor.native_id`
|
|
188
|
+
while `doc_id` is the mirror's own field — nothing enforces the two agree for
|
|
189
|
+
a third-party connector. Where they diverge, this join produces a silent
|
|
190
|
+
false miss, not an error.
|
|
191
|
+
- `sources.content_hash = mirror.anchor.content_hash` as an **equi-join** is not
|
|
192
|
+
the sound alternative either. `kbforge.pipeline.run` calls
|
|
193
|
+
`commit(mirror_path, docs)` right after `kbforge_publish`, and publishing only
|
|
194
|
+
opens a review request — kbforge never merges. Between publish and merge the
|
|
195
|
+
mirror already holds the new hash while the bundle still carries the previous
|
|
196
|
+
run's, so an equi-join on hash returns **zero rows for every concept with a
|
|
197
|
+
review request open** — exactly the ones an audit most wants to see.
|
|
198
|
+
|
|
199
|
+
The correct form: join on `id` with a **`LEFT JOIN`**, not an inner join, then
|
|
200
|
+
**compare** hashes instead of joining on them — an inner join silently drops
|
|
201
|
+
every source with no mirror row at all, which is exactly the "never
|
|
202
|
+
published" case the query needs to report. A match means current; a mismatch
|
|
203
|
+
means an update is sitting in review; a source with no mirror row was never
|
|
204
|
+
published. Run for real against the messy fixture (`concepts/ok/overview.md`'s
|
|
205
|
+
two sources — an owning `wiki:ok` and a grounding `notes:ok` — plus two other
|
|
206
|
+
concepts whose sources have no mirror row at all) with a synthetic
|
|
207
|
+
two-document mirror covering only `wiki:ok` and `notes:ok`, one hash left
|
|
208
|
+
matching and one changed to simulate an update stuck in review:
|
|
209
|
+
|
|
210
|
+
```bash
|
|
211
|
+
$ okfquery query "
|
|
212
|
+
SELECT s.path,
|
|
213
|
+
s.id,
|
|
214
|
+
s.content_hash AS bundle_hash,
|
|
215
|
+
m.anchor.content_hash AS mirror_hash,
|
|
216
|
+
CASE WHEN m.doc_id IS NULL THEN 'never published'
|
|
217
|
+
WHEN s.content_hash = m.anchor.content_hash THEN 'current'
|
|
218
|
+
ELSE 'update in review' END AS status
|
|
219
|
+
FROM sources s LEFT JOIN mirror m ON s.id = m.doc_id
|
|
220
|
+
ORDER BY s.path, s.id;
|
|
221
|
+
" --bundle packages/okfquery/tests/fixtures/messy --mirror /path/to/mirror
|
|
222
|
+
┌─────────────────────────────────┬─────────────────┬─────────────┬─────────────┬──────────────────┐
|
|
223
|
+
│ path │ id │ bundle_hash │ mirror_hash │ status │
|
|
224
|
+
│ varchar │ varchar │ varchar │ varchar │ varchar │
|
|
225
|
+
├─────────────────────────────────┼─────────────────┼─────────────┼─────────────┼──────────────────┤
|
|
226
|
+
│ concepts/claims/index.md │ wiki:claims │ h-claims │ NULL │ never published │
|
|
227
|
+
│ concepts/incomplete/overview.md │ wiki:incomplete │ NULL │ NULL │ never published │
|
|
228
|
+
│ concepts/ok/overview.md │ notes:ok │ h-ground │ h-ground-v2 │ update in review │
|
|
229
|
+
│ concepts/ok/overview.md │ wiki:ok │ h-own │ h-own │ current │
|
|
230
|
+
└─────────────────────────────────┴─────────────────┴─────────────┴─────────────┴──────────────────┘
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
That does not fix the `id` gap above: a connector whose anchor disagrees with
|
|
234
|
+
its `doc_id` still misses this join silently. There is no workaround for that
|
|
235
|
+
here — only the honest statement of it.
|
|
236
|
+
|
|
237
|
+
## Not built
|
|
238
|
+
|
|
239
|
+
- **No full-text search out of the box.** `body` is a plain column; `INSTALL
|
|
240
|
+
fts` and an index over it is a documented recipe, not a feature.
|
|
241
|
+
- **No remote bundles.** `--bundle` is a local checkout path; reading one over
|
|
242
|
+
HTTP wants a caching story this package deliberately avoids.
|
|
243
|
+
- **No MCP server.** `load()` is already small enough to wrap in one `query`
|
|
244
|
+
tool and a schema resource; nothing here does that yet.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
okfquery/__init__.py,sha256=kbuky3_HNMhxsf1pdOo49QiScaRc3UBa20q0Gdf_qeA,446
|
|
2
|
+
okfquery/cli.py,sha256=93mqOdVi7L2VWQTOJztMmpUVs-9ke7b_ZWdcGkAFryI,4986
|
|
3
|
+
okfquery/load.py,sha256=DVeDithpDRiZZxRglakY7NYWxPgwKcyoeo7XcX71-io,5058
|
|
4
|
+
okfquery/parse.py,sha256=DOnEJyLTrDX2qMsUHpe4Tt775r57KkrQ5yw7IqcyAn4,8752
|
|
5
|
+
okfquery/schema.py,sha256=z8E0vQLB5SNNlh3QAJWw9jxt1WAaQYY15vvPwR4QWJk,1856
|
|
6
|
+
kbforge_okfquery-0.1.0.dist-info/METADATA,sha256=SpivqOVFTmZBNdPClLYY52htoKZ6GRzDGbpIrF---6M,12493
|
|
7
|
+
kbforge_okfquery-0.1.0.dist-info/WHEEL,sha256=zOwg4jB6zX2kU910N-cMawjivD6tO8NEWvE12je1bVk,87
|
|
8
|
+
kbforge_okfquery-0.1.0.dist-info/entry_points.txt,sha256=r6vWPNoNezHdt_JlRk-ZTquezvP7lrsKcsrAZK1H89Y,47
|
|
9
|
+
kbforge_okfquery-0.1.0.dist-info/RECORD,,
|
okfquery/__init__.py
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
"""SQL over an OKF v0.2 bundle.
|
|
2
|
+
|
|
3
|
+
`load()` returns a DuckDB connection, deliberately unwrapped: an OKF bundle is
|
|
4
|
+
just a DuckDB database, and anything this package wrapped would be a worse
|
|
5
|
+
version of what DuckDB already offers."""
|
|
6
|
+
|
|
7
|
+
from __future__ import annotations
|
|
8
|
+
|
|
9
|
+
__version__ = "0.1.0"
|
|
10
|
+
|
|
11
|
+
from okfquery.load import EmptyMirrorError, load
|
|
12
|
+
from okfquery.schema import SCHEMA_SQL
|
|
13
|
+
|
|
14
|
+
__all__ = ["EmptyMirrorError", "SCHEMA_SQL", "__version__", "load"]
|
okfquery/cli.py
ADDED
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
"""argparse over `load()`. Four verbs, no query logic of its own.
|
|
2
|
+
|
|
3
|
+
There is deliberately no --format parquet and no --output: DuckDB writes parquet
|
|
4
|
+
from inside the SQL (`COPY (...) TO 'out.parquet'`), and a second export path
|
|
5
|
+
here would be a worse version of a feature already shipping."""
|
|
6
|
+
|
|
7
|
+
from __future__ import annotations
|
|
8
|
+
|
|
9
|
+
import argparse
|
|
10
|
+
import csv
|
|
11
|
+
import json
|
|
12
|
+
import subprocess
|
|
13
|
+
import sys
|
|
14
|
+
import tempfile
|
|
15
|
+
from pathlib import Path
|
|
16
|
+
|
|
17
|
+
import duckdb
|
|
18
|
+
|
|
19
|
+
from okfquery.load import EmptyMirrorError, load
|
|
20
|
+
from okfquery.schema import SCHEMA_SQL
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
def _bundle_missing(bundle: str) -> bool:
|
|
24
|
+
"""True when `bundle` has no `concepts/` directory to scan.
|
|
25
|
+
|
|
26
|
+
`load()` must keep answering an empty bundle (no .md files under an
|
|
27
|
+
existing `concepts/`) rather than erroring -- that is a real, if boring,
|
|
28
|
+
OKF bundle. A *missing* `concepts/` is different: it means the path is
|
|
29
|
+
wrong, and `scan()` silently returning `[]` for it would make `check`
|
|
30
|
+
report "no problems" and exit 0 against a bundle that was never loaded at
|
|
31
|
+
all. That is a CLI-level concern, not `load()`'s, so it is caught here.
|
|
32
|
+
"""
|
|
33
|
+
return not (Path(bundle) / "concepts").is_dir()
|
|
34
|
+
|
|
35
|
+
|
|
36
|
+
def _connect(args: argparse.Namespace) -> duckdb.DuckDBPyConnection:
|
|
37
|
+
mirror = Path(args.mirror) if args.mirror else None
|
|
38
|
+
return load(Path(args.bundle), mirror=mirror)
|
|
39
|
+
|
|
40
|
+
|
|
41
|
+
def _emit(con: duckdb.DuckDBPyConnection, sql: str, fmt: str) -> None:
|
|
42
|
+
if fmt == "table":
|
|
43
|
+
print(con.sql(sql))
|
|
44
|
+
return
|
|
45
|
+
cursor = con.execute(sql)
|
|
46
|
+
columns = [d[0] for d in cursor.description]
|
|
47
|
+
rows = cursor.fetchall()
|
|
48
|
+
if fmt == "json":
|
|
49
|
+
print(json.dumps([dict(zip(columns, r)) for r in rows], default=str))
|
|
50
|
+
return
|
|
51
|
+
writer = csv.writer(sys.stdout, lineterminator="\n")
|
|
52
|
+
writer.writerow(columns)
|
|
53
|
+
writer.writerows(rows)
|
|
54
|
+
|
|
55
|
+
|
|
56
|
+
def _shell(args: argparse.Namespace) -> int:
|
|
57
|
+
if _bundle_missing(args.bundle):
|
|
58
|
+
print(f"no concepts/ directory under {args.bundle}", file=sys.stderr)
|
|
59
|
+
return 2
|
|
60
|
+
# subprocess, NEVER exec. exec replaces this process, so the TemporaryDirectory
|
|
61
|
+
# finalizer would never run and a database holding every concept body would
|
|
62
|
+
# outlive the session -- quietly breaking the ephemeral guarantee the whole
|
|
63
|
+
# design rests on.
|
|
64
|
+
with tempfile.TemporaryDirectory(prefix="okfquery-") as tmp:
|
|
65
|
+
database = Path(tmp) / "bundle.duckdb"
|
|
66
|
+
mirror = Path(args.mirror) if args.mirror else None
|
|
67
|
+
try:
|
|
68
|
+
load(Path(args.bundle), mirror=mirror, database=str(database)).close()
|
|
69
|
+
except EmptyMirrorError as exc:
|
|
70
|
+
print(exc, file=sys.stderr)
|
|
71
|
+
return 2
|
|
72
|
+
try:
|
|
73
|
+
return subprocess.run(["duckdb", str(database)], check=False).returncode
|
|
74
|
+
except FileNotFoundError:
|
|
75
|
+
print(
|
|
76
|
+
"the `duckdb` CLI is not on PATH. Install it "
|
|
77
|
+
"(https://duckdb.org/docs/installation/) and run:\n"
|
|
78
|
+
f" duckdb {database}\n"
|
|
79
|
+
"note: that file is deleted when this command exits.",
|
|
80
|
+
file=sys.stderr,
|
|
81
|
+
)
|
|
82
|
+
return 2
|
|
83
|
+
|
|
84
|
+
|
|
85
|
+
def main(argv: list[str] | None = None) -> int:
|
|
86
|
+
parser = argparse.ArgumentParser(prog="okfquery", description=__doc__)
|
|
87
|
+
sub = parser.add_subparsers(dest="command", required=True)
|
|
88
|
+
|
|
89
|
+
def bundle_args(p: argparse.ArgumentParser) -> None:
|
|
90
|
+
p.add_argument("--bundle", default=".", help="bundle root (default: .)")
|
|
91
|
+
p.add_argument("--mirror", default=None, help="kbforge mirror to attach")
|
|
92
|
+
|
|
93
|
+
query = sub.add_parser("query", help="run one SQL statement")
|
|
94
|
+
query.add_argument("sql")
|
|
95
|
+
query.add_argument("--format", choices=("table", "json", "csv"), default="table")
|
|
96
|
+
bundle_args(query)
|
|
97
|
+
|
|
98
|
+
shell = sub.add_parser("shell", help="open the duckdb CLI on this bundle")
|
|
99
|
+
bundle_args(shell)
|
|
100
|
+
|
|
101
|
+
sub.add_parser("schema", help="print the DDL")
|
|
102
|
+
|
|
103
|
+
check = sub.add_parser("check", help="exit 1 if any file failed to parse")
|
|
104
|
+
bundle_args(check)
|
|
105
|
+
|
|
106
|
+
args = parser.parse_args(argv)
|
|
107
|
+
|
|
108
|
+
if args.command == "schema":
|
|
109
|
+
print(SCHEMA_SQL.strip())
|
|
110
|
+
return 0
|
|
111
|
+
if args.command == "shell":
|
|
112
|
+
return _shell(args)
|
|
113
|
+
|
|
114
|
+
if _bundle_missing(args.bundle):
|
|
115
|
+
print(f"no concepts/ directory under {args.bundle}", file=sys.stderr)
|
|
116
|
+
return 2
|
|
117
|
+
|
|
118
|
+
try:
|
|
119
|
+
con = _connect(args)
|
|
120
|
+
except EmptyMirrorError as exc:
|
|
121
|
+
print(exc, file=sys.stderr)
|
|
122
|
+
return 2
|
|
123
|
+
|
|
124
|
+
if args.command == "check":
|
|
125
|
+
rows = con.execute(
|
|
126
|
+
"select path, kind, detail from problems order by path, kind"
|
|
127
|
+
).fetchall()
|
|
128
|
+
if not rows:
|
|
129
|
+
print("no problems")
|
|
130
|
+
return 0
|
|
131
|
+
for path, kind, detail in rows:
|
|
132
|
+
print(f"{path}: {kind}: {detail}", file=sys.stderr)
|
|
133
|
+
return 1
|
|
134
|
+
|
|
135
|
+
try:
|
|
136
|
+
_emit(con, args.sql, args.format)
|
|
137
|
+
except duckdb.Error as exc:
|
|
138
|
+
print(exc, file=sys.stderr)
|
|
139
|
+
return 2
|
|
140
|
+
return 0
|
|
141
|
+
|
|
142
|
+
|
|
143
|
+
if __name__ == "__main__": # pragma: no cover
|
|
144
|
+
raise SystemExit(main())
|
okfquery/load.py
ADDED
|
@@ -0,0 +1,132 @@
|
|
|
1
|
+
"""Bundle on disk -> a DuckDB connection with the schema filled.
|
|
2
|
+
|
|
3
|
+
Ephemeral by design: nothing is cached and nothing is written, so the tables
|
|
4
|
+
cannot disagree with the files they came from. That is also why no index is
|
|
5
|
+
committed into a bundle -- an index derived from `ProposedChange.concepts` would
|
|
6
|
+
be a THIRD carrier of the same concept, and kbforge already has two."""
|
|
7
|
+
|
|
8
|
+
from __future__ import annotations
|
|
9
|
+
|
|
10
|
+
import json
|
|
11
|
+
from pathlib import Path
|
|
12
|
+
|
|
13
|
+
import duckdb
|
|
14
|
+
|
|
15
|
+
from okfquery.parse import is_reserved, parse
|
|
16
|
+
from okfquery.schema import SCHEMA_SQL
|
|
17
|
+
|
|
18
|
+
|
|
19
|
+
class EmptyMirrorError(Exception):
|
|
20
|
+
"""A --mirror path holding no document slots. Raised rather than passed
|
|
21
|
+
through, because read_json_auto's own error on a zero-match glob names a
|
|
22
|
+
pattern, not the directory the user typed."""
|
|
23
|
+
|
|
24
|
+
|
|
25
|
+
def scan(bundle: Path) -> list[Path]:
|
|
26
|
+
"""Concept files, sorted. Reserved OKF artifacts are excluded here and
|
|
27
|
+
nowhere else, so the count invariant has exactly one definition.
|
|
28
|
+
|
|
29
|
+
The glob is `concepts/**/*.md` because that is where kbforge's
|
|
30
|
+
`synthesize.concept_path` puts everything -- which keeps a bundle repo's own
|
|
31
|
+
README.md and docs/ out of the tables."""
|
|
32
|
+
root = bundle / "concepts"
|
|
33
|
+
if not root.is_dir():
|
|
34
|
+
return []
|
|
35
|
+
# errors="replace", not the bare default: a mis-encoded file is exactly the
|
|
36
|
+
# pathology this tool exists to surface. Raising here would abort the whole
|
|
37
|
+
# load over one file; U+FFFD replacement characters let it through to
|
|
38
|
+
# `parse()`, which turns it into a `concepts` row (NULLs where the decode
|
|
39
|
+
# broke the frontmatter) plus a `problems` row, same as any other bad file.
|
|
40
|
+
return [
|
|
41
|
+
p
|
|
42
|
+
for p in sorted(root.rglob("*.md"))
|
|
43
|
+
if not is_reserved(p.name, p.read_text("utf-8", errors="replace"))
|
|
44
|
+
]
|
|
45
|
+
|
|
46
|
+
|
|
47
|
+
def _mirror_view(con: duckdb.DuckDBPyConnection, mirror: Path) -> None:
|
|
48
|
+
# Shallow, never `**/*.json`: grounding.SIDECAR_DIR keeps flat
|
|
49
|
+
# {doc_id: content_hash} maps in <mirror>/_grounding/, and unioning those
|
|
50
|
+
# into a CanonicalDocument view would wreck read_json_auto's inferred schema.
|
|
51
|
+
slots = sorted(mirror.glob("*.json"))
|
|
52
|
+
if not slots:
|
|
53
|
+
raise EmptyMirrorError(
|
|
54
|
+
f"mirror {mirror} holds no document slots (*.json); "
|
|
55
|
+
"nothing to attach as the `mirror` view"
|
|
56
|
+
)
|
|
57
|
+
pattern = str(mirror / "*.json").replace("'", "''")
|
|
58
|
+
con.execute(f"CREATE VIEW mirror AS SELECT * FROM read_json_auto('{pattern}')")
|
|
59
|
+
|
|
60
|
+
|
|
61
|
+
def load(
|
|
62
|
+
bundle: Path,
|
|
63
|
+
mirror: Path | None = None,
|
|
64
|
+
*,
|
|
65
|
+
database: str = ":memory:",
|
|
66
|
+
) -> duckdb.DuckDBPyConnection:
|
|
67
|
+
"""An OKF v0.2 bundle as a DuckDB connection.
|
|
68
|
+
|
|
69
|
+
Returns the connection itself, unwrapped. An OKF bundle is just a DuckDB
|
|
70
|
+
database: COPY ... TO 'x.parquet', ATTACH, joining against your own CSV, and
|
|
71
|
+
INSTALL fts over `body` all work because the real object comes back.
|
|
72
|
+
|
|
73
|
+
`database` exists for `okfquery shell`, which needs the same tables in a file
|
|
74
|
+
the duckdb CLI can open. Everything else takes the default.
|
|
75
|
+
"""
|
|
76
|
+
bundle = Path(bundle)
|
|
77
|
+
con = duckdb.connect(database)
|
|
78
|
+
# Set before anything is inserted: TIMESTAMPTZ values render in the session
|
|
79
|
+
# timezone, and a bundle should read the same on every machine.
|
|
80
|
+
con.execute("SET timezone = 'UTC'")
|
|
81
|
+
con.execute(SCHEMA_SQL)
|
|
82
|
+
|
|
83
|
+
concepts: list[tuple] = []
|
|
84
|
+
sources: list[tuple] = []
|
|
85
|
+
links: list[tuple] = []
|
|
86
|
+
problems: list[tuple] = []
|
|
87
|
+
|
|
88
|
+
for file in scan(bundle):
|
|
89
|
+
path = file.relative_to(bundle).as_posix()
|
|
90
|
+
# Same errors="replace" as scan(), and for the same reason: one
|
|
91
|
+
# mis-encoded file must produce a row, not a traceback that aborts
|
|
92
|
+
# every other concept in the bundle along with it.
|
|
93
|
+
concept = parse(file.read_text("utf-8", errors="replace"))
|
|
94
|
+
concepts.append(
|
|
95
|
+
(
|
|
96
|
+
path,
|
|
97
|
+
concept.type,
|
|
98
|
+
concept.title,
|
|
99
|
+
concept.description,
|
|
100
|
+
concept.generated_by,
|
|
101
|
+
concept.generated_at,
|
|
102
|
+
json.dumps(concept.facets) if concept.facets else None,
|
|
103
|
+
concept.body,
|
|
104
|
+
)
|
|
105
|
+
)
|
|
106
|
+
for ordinal, source in enumerate(concept.sources):
|
|
107
|
+
sources.append(
|
|
108
|
+
(
|
|
109
|
+
path,
|
|
110
|
+
ordinal,
|
|
111
|
+
source["id"],
|
|
112
|
+
source["resource"],
|
|
113
|
+
source["content_hash"],
|
|
114
|
+
)
|
|
115
|
+
)
|
|
116
|
+
links.extend((path, target) for target in concept.links)
|
|
117
|
+
problems.extend((path, p.kind, p.detail) for p in concept.problems)
|
|
118
|
+
|
|
119
|
+
if concepts:
|
|
120
|
+
con.executemany(
|
|
121
|
+
"INSERT INTO concepts VALUES (?, ?, ?, ?, ?, ?, ?, ?)", concepts
|
|
122
|
+
)
|
|
123
|
+
if sources:
|
|
124
|
+
con.executemany("INSERT INTO sources VALUES (?, ?, ?, ?, ?)", sources)
|
|
125
|
+
if links:
|
|
126
|
+
con.executemany("INSERT INTO links VALUES (?, ?)", links)
|
|
127
|
+
if problems:
|
|
128
|
+
con.executemany("INSERT INTO problems VALUES (?, ?, ?)", problems)
|
|
129
|
+
|
|
130
|
+
if mirror is not None:
|
|
131
|
+
_mirror_view(con, Path(mirror))
|
|
132
|
+
return con
|
okfquery/parse.py
ADDED
|
@@ -0,0 +1,252 @@
|
|
|
1
|
+
"""Pure frontmatter parsing for OKF v0.2 concepts.
|
|
2
|
+
|
|
3
|
+
No filesystem, no clock, no DuckDB: text in, a `Concept` and its `Problem`s out.
|
|
4
|
+
That purity is why every row of the problem ladder is testable from a string
|
|
5
|
+
literal.
|
|
6
|
+
|
|
7
|
+
This module deliberately does NOT reuse kbforge's `validate._parse_frontmatter`,
|
|
8
|
+
and not only because okfquery imports no kbforge. That function collapses
|
|
9
|
+
no-fence, unterminated-fence, and broken-YAML alike to `{}` -- correct for a gate,
|
|
10
|
+
which needs only "no usable frontmatter" and fails either way. Here, *which* of
|
|
11
|
+
the three occurred is the entire finding: a missing fence is a renderer bug,
|
|
12
|
+
broken YAML is a hand-edit, an unterminated fence is a truncated write. Do not
|
|
13
|
+
"fix" this duplication by importing kbforge's version."""
|
|
14
|
+
|
|
15
|
+
from __future__ import annotations
|
|
16
|
+
|
|
17
|
+
from dataclasses import dataclass, field
|
|
18
|
+
from datetime import UTC, datetime
|
|
19
|
+
|
|
20
|
+
import yaml
|
|
21
|
+
|
|
22
|
+
RESERVED = frozenset({"index.md", "log.md"})
|
|
23
|
+
"""OKF §8 directory listings and change logs. See `is_reserved` for the rule."""
|
|
24
|
+
|
|
25
|
+
# The keys OKF owns at the head of a concept. Everything else in the frontmatter
|
|
26
|
+
# is a facet. Mirrors kbforge's synthesize.OKF_OWNED, which is what keeps a
|
|
27
|
+
# source field named `type` out of the facet map on the emit side.
|
|
28
|
+
OKF_OWNED = frozenset({"type", "title", "description", "generated", "sources", "links"})
|
|
29
|
+
|
|
30
|
+
_REQUIRED = ("type", "title", "description", "generated", "sources")
|
|
31
|
+
|
|
32
|
+
|
|
33
|
+
@dataclass(frozen=True)
|
|
34
|
+
class Problem:
|
|
35
|
+
kind: str
|
|
36
|
+
detail: str
|
|
37
|
+
|
|
38
|
+
|
|
39
|
+
@dataclass
|
|
40
|
+
class Concept:
|
|
41
|
+
type: str | None = None
|
|
42
|
+
title: str | None = None
|
|
43
|
+
description: str | None = None
|
|
44
|
+
generated_by: str | None = None
|
|
45
|
+
generated_at: datetime | None = None
|
|
46
|
+
facets: dict = field(default_factory=dict)
|
|
47
|
+
body: str = ""
|
|
48
|
+
sources: list[dict] = field(default_factory=list)
|
|
49
|
+
links: list[str] = field(default_factory=list)
|
|
50
|
+
problems: list[Problem] = field(default_factory=list)
|
|
51
|
+
|
|
52
|
+
|
|
53
|
+
def is_reserved(basename: str, content: str) -> bool:
|
|
54
|
+
"""Whether a file under `concepts/` is a non-concept OKF artifact.
|
|
55
|
+
|
|
56
|
+
Copied from kbforge's `validate._check_strict_okf`, deliberately and with
|
|
57
|
+
this note: a file is exempt only if its basename is reserved AND it does not
|
|
58
|
+
open a frontmatter fence, because an `index.md` bearing frontmatter is
|
|
59
|
+
claiming to be a concept and exempting it would skip every check. Keyed on
|
|
60
|
+
the raw fence, not on a parse result -- "no frontmatter" and "broken
|
|
61
|
+
frontmatter" must not look alike."""
|
|
62
|
+
return basename in RESERVED and not content.lstrip().startswith("---")
|
|
63
|
+
|
|
64
|
+
|
|
65
|
+
def _split(content: str) -> tuple[str | None, str, Problem | None]:
|
|
66
|
+
"""(raw frontmatter, body, problem). Raw is None when there is none to parse."""
|
|
67
|
+
text = content.lstrip()
|
|
68
|
+
if not text.startswith("---"):
|
|
69
|
+
return (
|
|
70
|
+
None,
|
|
71
|
+
content,
|
|
72
|
+
Problem("no-frontmatter", "file does not open a '---' frontmatter fence"),
|
|
73
|
+
)
|
|
74
|
+
_, _, rest = text.partition("---")
|
|
75
|
+
raw, sep, body = rest.partition("\n---")
|
|
76
|
+
if not sep:
|
|
77
|
+
return (
|
|
78
|
+
None,
|
|
79
|
+
content,
|
|
80
|
+
Problem(
|
|
81
|
+
"unterminated-frontmatter",
|
|
82
|
+
"frontmatter fence is opened but never closed by a closing '---'",
|
|
83
|
+
),
|
|
84
|
+
)
|
|
85
|
+
return raw, body.lstrip("\n"), None
|
|
86
|
+
|
|
87
|
+
|
|
88
|
+
def _generated(front: dict, out: Concept) -> None:
|
|
89
|
+
block = front.get("generated")
|
|
90
|
+
if not isinstance(block, dict):
|
|
91
|
+
out.problems.append(
|
|
92
|
+
Problem(
|
|
93
|
+
"bad-generated", f"'generated' is {type(block).__name__}, not a mapping"
|
|
94
|
+
)
|
|
95
|
+
)
|
|
96
|
+
return
|
|
97
|
+
by = block.get("by")
|
|
98
|
+
out.generated_by = by if isinstance(by, str) and by else None
|
|
99
|
+
if out.generated_by is None:
|
|
100
|
+
out.problems.append(
|
|
101
|
+
Problem("bad-generated", "'generated.by' is missing or empty")
|
|
102
|
+
)
|
|
103
|
+
at = block.get("at")
|
|
104
|
+
if at is None:
|
|
105
|
+
out.problems.append(Problem("bad-generated", "'generated.at' is missing"))
|
|
106
|
+
return
|
|
107
|
+
parsed = _timestamp(at, out)
|
|
108
|
+
if parsed is not None and parsed.utcoffset() is None:
|
|
109
|
+
out.problems.append(
|
|
110
|
+
Problem(
|
|
111
|
+
"naive-timestamp",
|
|
112
|
+
f"'generated.at' {at!r} carries no UTC offset (§4.4 law 4 requires "
|
|
113
|
+
"an aware stamp); read as UTC",
|
|
114
|
+
)
|
|
115
|
+
)
|
|
116
|
+
parsed = parsed.replace(tzinfo=UTC)
|
|
117
|
+
out.generated_at = parsed
|
|
118
|
+
|
|
119
|
+
|
|
120
|
+
def _timestamp(at: object, out: Concept) -> datetime | None:
|
|
121
|
+
if isinstance(at, datetime):
|
|
122
|
+
return at
|
|
123
|
+
try:
|
|
124
|
+
return datetime.fromisoformat(str(at))
|
|
125
|
+
except (TypeError, ValueError):
|
|
126
|
+
out.problems.append(
|
|
127
|
+
Problem(
|
|
128
|
+
"bad-timestamp", f"'generated.at' {at!r} is not an ISO-8601 datetime"
|
|
129
|
+
)
|
|
130
|
+
)
|
|
131
|
+
return None
|
|
132
|
+
|
|
133
|
+
|
|
134
|
+
def _sources(front: dict, out: Concept) -> None:
|
|
135
|
+
raw = front.get("sources")
|
|
136
|
+
if not isinstance(raw, list):
|
|
137
|
+
out.problems.append(
|
|
138
|
+
Problem("bad-sources", f"'sources' is {type(raw).__name__}, not a list")
|
|
139
|
+
)
|
|
140
|
+
return
|
|
141
|
+
for i, entry in enumerate(raw):
|
|
142
|
+
if not isinstance(entry, dict):
|
|
143
|
+
out.problems.append(
|
|
144
|
+
Problem(
|
|
145
|
+
"bad-sources",
|
|
146
|
+
f"'sources' entry {i} is {type(entry).__name__}, not a mapping",
|
|
147
|
+
)
|
|
148
|
+
)
|
|
149
|
+
continue
|
|
150
|
+
resource = entry.get("resource")
|
|
151
|
+
if not isinstance(resource, str) or not resource:
|
|
152
|
+
out.problems.append(
|
|
153
|
+
Problem(
|
|
154
|
+
"bad-sources",
|
|
155
|
+
f"'sources' entry {i} has no 'resource' (§5.1 requires one)",
|
|
156
|
+
)
|
|
157
|
+
)
|
|
158
|
+
resource = None
|
|
159
|
+
# Carried regardless: problems are additive, and dropping the entry
|
|
160
|
+
# would silently renumber every ordinal after it.
|
|
161
|
+
out.sources.append(
|
|
162
|
+
{
|
|
163
|
+
"id": entry.get("id") if isinstance(entry.get("id"), str) else None,
|
|
164
|
+
"resource": resource,
|
|
165
|
+
"content_hash": entry.get("content_hash")
|
|
166
|
+
if isinstance(entry.get("content_hash"), str)
|
|
167
|
+
else None,
|
|
168
|
+
}
|
|
169
|
+
)
|
|
170
|
+
|
|
171
|
+
|
|
172
|
+
def _links(front: dict, out: Concept) -> None:
|
|
173
|
+
raw = front.get("links")
|
|
174
|
+
if raw is None:
|
|
175
|
+
return
|
|
176
|
+
if not isinstance(raw, list):
|
|
177
|
+
out.problems.append(
|
|
178
|
+
Problem("bad-links", f"'links' is {type(raw).__name__}, not a list")
|
|
179
|
+
)
|
|
180
|
+
return
|
|
181
|
+
for entry in raw:
|
|
182
|
+
if isinstance(entry, str):
|
|
183
|
+
out.links.append(entry)
|
|
184
|
+
else:
|
|
185
|
+
out.problems.append(
|
|
186
|
+
Problem("bad-links", f"'links' entry {entry!r} is not a string")
|
|
187
|
+
)
|
|
188
|
+
|
|
189
|
+
|
|
190
|
+
def _facets(front: dict) -> dict:
|
|
191
|
+
"""Frontmatter keys OKF does not own. Mirrors synthesize._facets on the emit
|
|
192
|
+
side: scalars and scalar lists only, so a nested mapping never lands in a
|
|
193
|
+
column typed for values."""
|
|
194
|
+
scalar = (str, int, float, bool)
|
|
195
|
+
|
|
196
|
+
def ok(v: object) -> bool:
|
|
197
|
+
if isinstance(v, scalar):
|
|
198
|
+
return True
|
|
199
|
+
return isinstance(v, list) and all(isinstance(i, scalar) for i in v)
|
|
200
|
+
|
|
201
|
+
return {k: v for k, v in front.items() if k not in OKF_OWNED and ok(v)}
|
|
202
|
+
|
|
203
|
+
|
|
204
|
+
def parse(content: str) -> Concept:
|
|
205
|
+
"""One concept file's text into a `Concept` plus the problems it carries.
|
|
206
|
+
|
|
207
|
+
Never raises. A file this cannot read still yields a Concept -- with NULLs
|
|
208
|
+
for what was unreadable -- because dropping it would make an audit tool hide
|
|
209
|
+
exactly the files most likely to be wrong."""
|
|
210
|
+
out = Concept()
|
|
211
|
+
raw, body, problem = _split(content)
|
|
212
|
+
out.body = body
|
|
213
|
+
if problem is not None:
|
|
214
|
+
out.problems.append(problem)
|
|
215
|
+
return out
|
|
216
|
+
try:
|
|
217
|
+
front = yaml.safe_load(raw)
|
|
218
|
+
except yaml.YAMLError as exc:
|
|
219
|
+
out.problems.append(
|
|
220
|
+
Problem(
|
|
221
|
+
"invalid-yaml",
|
|
222
|
+
f"frontmatter is not valid YAML: {exc.__class__.__name__}",
|
|
223
|
+
)
|
|
224
|
+
)
|
|
225
|
+
return out
|
|
226
|
+
if not isinstance(front, dict):
|
|
227
|
+
out.problems.append(
|
|
228
|
+
Problem(
|
|
229
|
+
"frontmatter-not-mapping",
|
|
230
|
+
f"frontmatter parses to {type(front).__name__}, not a mapping",
|
|
231
|
+
)
|
|
232
|
+
)
|
|
233
|
+
return out
|
|
234
|
+
|
|
235
|
+
missing = [k for k in _REQUIRED if front.get(k) is None]
|
|
236
|
+
if missing:
|
|
237
|
+
out.problems.append(
|
|
238
|
+
Problem(
|
|
239
|
+
"missing-required", f"missing required OKF keys: {', '.join(missing)}"
|
|
240
|
+
)
|
|
241
|
+
)
|
|
242
|
+
|
|
243
|
+
for key in ("type", "title", "description"):
|
|
244
|
+
value = front.get(key)
|
|
245
|
+
setattr(out, key, value if isinstance(value, str) else None)
|
|
246
|
+
out.facets = _facets(front)
|
|
247
|
+
if front.get("generated") is not None:
|
|
248
|
+
_generated(front, out)
|
|
249
|
+
if front.get("sources") is not None:
|
|
250
|
+
_sources(front, out)
|
|
251
|
+
_links(front, out)
|
|
252
|
+
return out
|
okfquery/schema.py
ADDED
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
"""The DDL, defined once.
|
|
2
|
+
|
|
3
|
+
Explicit CREATE TABLE rather than a dataframe round-trip, for two reasons: an
|
|
4
|
+
empty bundle still answers queries instead of erroring on a missing table, and
|
|
5
|
+
the column types are pinned rather than inferred."""
|
|
6
|
+
|
|
7
|
+
from __future__ import annotations
|
|
8
|
+
|
|
9
|
+
SCHEMA_SQL = """
|
|
10
|
+
CREATE TABLE concepts (
|
|
11
|
+
path VARCHAR NOT NULL,
|
|
12
|
+
type VARCHAR,
|
|
13
|
+
title VARCHAR,
|
|
14
|
+
description VARCHAR,
|
|
15
|
+
generated_by VARCHAR,
|
|
16
|
+
-- TIMESTAMPTZ, never TIMESTAMP. Not because a naive column would corrupt
|
|
17
|
+
-- the instant on this package's load path -- `load` binds an aware Python
|
|
18
|
+
-- datetime through executemany, so the instant survives a naive column
|
|
19
|
+
-- fine. What a naive column loses is the type: values come back with no
|
|
20
|
+
-- tzinfo, so every comparison against now() (itself TIMESTAMPTZ) needs a
|
|
21
|
+
-- cast, and the column asserts UTC by convention with nothing recording
|
|
22
|
+
-- that it does. §4.4 law 4 exists to force an *aware* stamp; a column that
|
|
23
|
+
-- cannot hold one throws away exactly the property the law buys.
|
|
24
|
+
generated_at TIMESTAMPTZ,
|
|
25
|
+
facets JSON,
|
|
26
|
+
body VARCHAR
|
|
27
|
+
);
|
|
28
|
+
|
|
29
|
+
CREATE TABLE sources (
|
|
30
|
+
path VARCHAR NOT NULL,
|
|
31
|
+
-- A real column. kbforge's synthesize.assemble puts the OWNING anchor first
|
|
32
|
+
-- and grounding anchors after it, and that ordering is the only thing
|
|
33
|
+
-- telling owner from ground -- OKF has no field for it. Discarding ordinal
|
|
34
|
+
-- during unnest would destroy the signal.
|
|
35
|
+
ordinal INTEGER NOT NULL,
|
|
36
|
+
id VARCHAR,
|
|
37
|
+
resource VARCHAR,
|
|
38
|
+
content_hash VARCHAR
|
|
39
|
+
);
|
|
40
|
+
|
|
41
|
+
CREATE TABLE links (
|
|
42
|
+
path VARCHAR NOT NULL,
|
|
43
|
+
target VARCHAR NOT NULL
|
|
44
|
+
);
|
|
45
|
+
|
|
46
|
+
CREATE TABLE problems (
|
|
47
|
+
path VARCHAR NOT NULL,
|
|
48
|
+
kind VARCHAR NOT NULL,
|
|
49
|
+
detail VARCHAR NOT NULL
|
|
50
|
+
);
|
|
51
|
+
"""
|