sensemaking 0.24.4 → 0.24.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +16 -12
- package/dist/cjs/cli/status.js +14 -24
- package/dist/cjs/cli/status.js.map +1 -1
- package/dist/cjs/commands/search.js +22 -0
- package/dist/cjs/commands/search.js.map +1 -1
- package/dist/cjs/config/types.js.map +1 -1
- package/dist/cjs/errors.d.cts +1 -1
- package/dist/cjs/errors.d.ts +1 -1
- package/dist/cjs/errors.js.map +1 -1
- package/dist/cjs/lib/worker-file.d.cts +1 -0
- package/dist/cjs/lib/worker-file.d.ts +1 -0
- package/dist/cjs/lib/worker-file.js +34 -0
- package/dist/cjs/lib/worker-file.js.map +1 -0
- package/dist/cjs/output/search-error.js +38 -0
- package/dist/cjs/output/search-error.js.map +1 -1
- package/dist/cjs/scan/pool.js +2 -23
- package/dist/cjs/scan/pool.js.map +1 -1
- package/dist/cjs/store/builder.d.cts +2 -1
- package/dist/cjs/store/builder.d.ts +2 -1
- package/dist/cjs/store/builder.js +4 -4
- package/dist/cjs/store/builder.js.map +1 -1
- package/dist/cjs/store/duckdb/lexical.d.cts +1 -0
- package/dist/cjs/store/duckdb/lexical.d.ts +1 -0
- package/dist/cjs/store/duckdb/lexical.js +28 -9
- package/dist/cjs/store/duckdb/lexical.js.map +1 -1
- package/dist/cjs/store/duckdb/native.d.cts +2 -0
- package/dist/cjs/store/duckdb/native.d.ts +2 -0
- package/dist/cjs/store/duckdb/native.js +11 -1
- package/dist/cjs/store/duckdb/native.js.map +1 -1
- package/dist/cjs/store/duckdb/ordered-bm25.d.cts +17 -0
- package/dist/cjs/store/duckdb/ordered-bm25.d.ts +17 -0
- package/dist/cjs/store/duckdb/ordered-bm25.js +238 -0
- package/dist/cjs/store/duckdb/ordered-bm25.js.map +1 -0
- package/dist/cjs/store/lock-wait.d.cts +2 -2
- package/dist/cjs/store/lock-wait.d.ts +2 -2
- package/dist/cjs/store/lock-wait.js +14 -11
- package/dist/cjs/store/lock-wait.js.map +1 -1
- package/dist/cjs/store/native.d.cts +4 -0
- package/dist/cjs/store/native.d.ts +4 -0
- package/dist/cjs/store/native.js +64 -9
- package/dist/cjs/store/native.js.map +1 -1
- package/dist/cjs/store/open.js +11 -7
- package/dist/cjs/store/open.js.map +1 -1
- package/dist/cjs/store/reconcile.d.cts +5 -1
- package/dist/cjs/store/reconcile.d.ts +5 -1
- package/dist/cjs/store/reconcile.js +5 -4
- package/dist/cjs/store/reconcile.js.map +1 -1
- package/dist/cjs/store/turso/lexical.js +5 -1
- package/dist/cjs/store/turso/lexical.js.map +1 -1
- package/dist/cjs/text/segment.js +168 -3
- package/dist/cjs/text/segment.js.map +1 -1
- package/dist/cjs/watch-claim.d.cts +39 -0
- package/dist/cjs/watch-claim.d.ts +39 -0
- package/dist/cjs/watch-claim.js +446 -0
- package/dist/cjs/watch-claim.js.map +1 -0
- package/dist/cjs/watch.js +240 -141
- package/dist/cjs/watch.js.map +1 -1
- package/dist/cjs/workers/watch-heartbeat.d.cts +1 -0
- package/dist/cjs/workers/watch-heartbeat.d.ts +1 -0
- package/dist/cjs/workers/watch-heartbeat.js +308 -0
- package/dist/cjs/workers/watch-heartbeat.js.map +1 -0
- package/dist/esm/cli/status.js +6 -5
- package/dist/esm/cli/status.js.map +1 -1
- package/dist/esm/commands/search.js +21 -0
- package/dist/esm/commands/search.js.map +1 -1
- package/dist/esm/config/types.js.map +1 -1
- package/dist/esm/errors.d.ts +1 -1
- package/dist/esm/errors.js.map +1 -1
- package/dist/esm/lib/worker-file.d.ts +1 -0
- package/dist/esm/lib/worker-file.js +24 -0
- package/dist/esm/lib/worker-file.js.map +1 -0
- package/dist/esm/output/search-error.js +21 -0
- package/dist/esm/output/search-error.js.map +1 -1
- package/dist/esm/scan/pool.js +2 -22
- package/dist/esm/scan/pool.js.map +1 -1
- package/dist/esm/store/builder.d.ts +2 -1
- package/dist/esm/store/builder.js +4 -4
- package/dist/esm/store/builder.js.map +1 -1
- package/dist/esm/store/duckdb/lexical.d.ts +1 -0
- package/dist/esm/store/duckdb/lexical.js +16 -7
- package/dist/esm/store/duckdb/lexical.js.map +1 -1
- package/dist/esm/store/duckdb/native.d.ts +2 -0
- package/dist/esm/store/duckdb/native.js +5 -1
- package/dist/esm/store/duckdb/native.js.map +1 -1
- package/dist/esm/store/duckdb/ordered-bm25.d.ts +17 -0
- package/dist/esm/store/duckdb/ordered-bm25.js +56 -0
- package/dist/esm/store/duckdb/ordered-bm25.js.map +1 -0
- package/dist/esm/store/lock-wait.d.ts +2 -2
- package/dist/esm/store/lock-wait.js +17 -14
- package/dist/esm/store/lock-wait.js.map +1 -1
- package/dist/esm/store/native.d.ts +4 -0
- package/dist/esm/store/native.js +60 -6
- package/dist/esm/store/native.js.map +1 -1
- package/dist/esm/store/open.js +9 -6
- package/dist/esm/store/open.js.map +1 -1
- package/dist/esm/store/reconcile.d.ts +5 -1
- package/dist/esm/store/reconcile.js +4 -3
- package/dist/esm/store/reconcile.js.map +1 -1
- package/dist/esm/store/turso/lexical.js +5 -1
- package/dist/esm/store/turso/lexical.js.map +1 -1
- package/dist/esm/text/segment.js +130 -4
- package/dist/esm/text/segment.js.map +1 -1
- package/dist/esm/watch-claim.d.ts +39 -0
- package/dist/esm/watch-claim.js +176 -0
- package/dist/esm/watch-claim.js.map +1 -0
- package/dist/esm/watch.js +168 -117
- package/dist/esm/watch.js.map +1 -1
- package/dist/esm/workers/watch-heartbeat.d.ts +1 -0
- package/dist/esm/workers/watch-heartbeat.js +106 -0
- package/dist/esm/workers/watch-heartbeat.js.map +1 -0
- package/package.json +2 -2
- package/skills/sense/SKILL.md +105 -104
- package/skills/sense/references/search.md +47 -0
- package/skills/sense/references/sql.md +120 -0
- package/skills/sense/references/stores/duckdb.md +35 -0
- package/skills/sense/references/stores/sqlite.md +66 -0
- package/skills/sense/references/stores/turso.md +36 -0
- package/skills/sense-setup/SKILL.md +93 -34
- package/skills/sense-setup/references/embeddings.md +51 -0
- package/skills/sense-setup/references/store-benchmarks.md +16 -0
- package/skills/sense-setup/references/stores/duckdb.md +28 -0
- package/skills/sense-setup/references/stores/sqlite.md +21 -0
- package/skills/sense-setup/references/stores/turso.md +25 -0
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
import { parentPort, workerData } from 'node:worker_threads';
|
|
2
|
+
import { SenseError } from '../errors.js';
|
|
3
|
+
import { serializeError } from '../scan/worker-error.js';
|
|
4
|
+
import { WatchClaimDatabase } from '../watch-claim.js';
|
|
5
|
+
function requireParentPort() {
|
|
6
|
+
const port = parentPort;
|
|
7
|
+
if (!port) throw new Error('watch heartbeat worker requires a parent port');
|
|
8
|
+
return port;
|
|
9
|
+
}
|
|
10
|
+
const port = requireParentPort();
|
|
11
|
+
const data = workerData;
|
|
12
|
+
let database;
|
|
13
|
+
let timer;
|
|
14
|
+
let stopping = false;
|
|
15
|
+
let renewing = Promise.resolve();
|
|
16
|
+
function send(message) {
|
|
17
|
+
port.postMessage(message);
|
|
18
|
+
}
|
|
19
|
+
function asError(value) {
|
|
20
|
+
return value instanceof Error ? value : new Error(String(value));
|
|
21
|
+
}
|
|
22
|
+
function closeDatabase() {
|
|
23
|
+
if (timer) clearInterval(timer);
|
|
24
|
+
timer = undefined;
|
|
25
|
+
try {
|
|
26
|
+
database === null || database === void 0 ? void 0 : database.close();
|
|
27
|
+
database = undefined;
|
|
28
|
+
return undefined;
|
|
29
|
+
} catch (err) {
|
|
30
|
+
return asError(err);
|
|
31
|
+
}
|
|
32
|
+
}
|
|
33
|
+
function addFailure(primary, secondary, message) {
|
|
34
|
+
return secondary ? new AggregateError([
|
|
35
|
+
primary,
|
|
36
|
+
secondary
|
|
37
|
+
], message) : primary;
|
|
38
|
+
}
|
|
39
|
+
function fail(err) {
|
|
40
|
+
if (stopping) return;
|
|
41
|
+
stopping = true;
|
|
42
|
+
let failure = asError(err);
|
|
43
|
+
try {
|
|
44
|
+
database === null || database === void 0 ? void 0 : database.release(data.token);
|
|
45
|
+
} catch (releaseError) {
|
|
46
|
+
failure = addFailure(failure, asError(releaseError), 'watch heartbeat and claim release both failed');
|
|
47
|
+
}
|
|
48
|
+
const closeError = closeDatabase();
|
|
49
|
+
failure = addFailure(failure, closeError, 'watch heartbeat and database cleanup both failed');
|
|
50
|
+
try {
|
|
51
|
+
send({
|
|
52
|
+
type: 'failure',
|
|
53
|
+
error: serializeError(failure)
|
|
54
|
+
});
|
|
55
|
+
} finally{
|
|
56
|
+
port.close();
|
|
57
|
+
}
|
|
58
|
+
}
|
|
59
|
+
async function stop() {
|
|
60
|
+
if (stopping) return;
|
|
61
|
+
stopping = true;
|
|
62
|
+
if (timer) clearInterval(timer);
|
|
63
|
+
let failure;
|
|
64
|
+
try {
|
|
65
|
+
await renewing;
|
|
66
|
+
if (!(database === null || database === void 0 ? void 0 : database.release(data.token))) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');
|
|
67
|
+
} catch (err) {
|
|
68
|
+
failure = asError(err);
|
|
69
|
+
} finally{
|
|
70
|
+
const closeError = closeDatabase();
|
|
71
|
+
if (closeError) failure = failure ? addFailure(failure, closeError, 'watch heartbeat shutdown and database cleanup both failed') : closeError;
|
|
72
|
+
try {
|
|
73
|
+
if (failure) send({
|
|
74
|
+
type: 'failure',
|
|
75
|
+
error: serializeError(failure)
|
|
76
|
+
});
|
|
77
|
+
else send({
|
|
78
|
+
type: 'stopped'
|
|
79
|
+
});
|
|
80
|
+
} finally{
|
|
81
|
+
port.close();
|
|
82
|
+
}
|
|
83
|
+
}
|
|
84
|
+
}
|
|
85
|
+
async function main() {
|
|
86
|
+
try {
|
|
87
|
+
database = new WatchClaimDatabase(data.configDir);
|
|
88
|
+
await database.acquire(data.token, data.pid, data.force);
|
|
89
|
+
send({
|
|
90
|
+
type: 'acquired'
|
|
91
|
+
});
|
|
92
|
+
timer = setInterval(()=>{
|
|
93
|
+
if (stopping) return;
|
|
94
|
+
renewing = renewing.then(()=>{
|
|
95
|
+
if (!(database === null || database === void 0 ? void 0 : database.renew(data.token))) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');
|
|
96
|
+
});
|
|
97
|
+
renewing.catch(fail);
|
|
98
|
+
}, data.heartbeatIntervalMs);
|
|
99
|
+
port.on('message', (message)=>{
|
|
100
|
+
if (message.type === 'stop') void stop();
|
|
101
|
+
});
|
|
102
|
+
} catch (err) {
|
|
103
|
+
fail(err);
|
|
104
|
+
}
|
|
105
|
+
}
|
|
106
|
+
void main();
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/workers/watch-heartbeat.ts"],"sourcesContent":["import { parentPort, workerData } from 'node:worker_threads';\nimport { SenseError } from '../errors.ts';\nimport { serializeError } from '../scan/worker-error.ts';\nimport { WatchClaimDatabase, type WatchClaimWorkerData, type WatchClaimWorkerMessage } from '../watch-claim.ts';\n\nfunction requireParentPort() {\n const port = parentPort;\n if (!port) throw new Error('watch heartbeat worker requires a parent port');\n return port;\n}\n\nconst port = requireParentPort();\nconst data = workerData as WatchClaimWorkerData;\nlet database: WatchClaimDatabase | undefined;\nlet timer: NodeJS.Timeout | undefined;\nlet stopping = false;\nlet renewing = Promise.resolve();\n\nfunction send(message: WatchClaimWorkerMessage): void {\n port.postMessage(message);\n}\n\nfunction asError(value: unknown): Error {\n return value instanceof Error ? value : new Error(String(value));\n}\n\nfunction closeDatabase(): Error | undefined {\n if (timer) clearInterval(timer);\n timer = undefined;\n try {\n database?.close();\n database = undefined;\n return undefined;\n } catch (err) {\n return asError(err);\n }\n}\n\nfunction addFailure(primary: Error, secondary: Error | undefined, message: string): Error {\n return secondary ? new AggregateError([primary, secondary], message) : primary;\n}\n\nfunction fail(err: unknown): void {\n if (stopping) return;\n stopping = true;\n let failure = asError(err);\n try {\n database?.release(data.token);\n } catch (releaseError) {\n failure = addFailure(failure, asError(releaseError), 'watch heartbeat and claim release both failed');\n }\n const closeError = closeDatabase();\n failure = addFailure(failure, closeError, 'watch heartbeat and database cleanup both failed');\n try {\n send({ type: 'failure', error: serializeError(failure) });\n } finally {\n port.close();\n }\n}\n\nasync function stop(): Promise<void> {\n if (stopping) return;\n stopping = true;\n if (timer) clearInterval(timer);\n let failure: Error | undefined;\n try {\n await renewing;\n if (!database?.release(data.token)) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');\n } catch (err) {\n failure = asError(err);\n } finally {\n const closeError = closeDatabase();\n if (closeError) failure = failure ? addFailure(failure, closeError, 'watch heartbeat shutdown and database cleanup both failed') : closeError;\n try {\n if (failure) send({ type: 'failure', error: serializeError(failure) });\n else send({ type: 'stopped' });\n } finally {\n port.close();\n }\n }\n}\n\nasync function main(): Promise<void> {\n try {\n database = new WatchClaimDatabase(data.configDir);\n await database.acquire(data.token, data.pid, data.force);\n send({ type: 'acquired' });\n timer = setInterval(() => {\n if (stopping) return;\n renewing = renewing.then(() => {\n if (!database?.renew(data.token)) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');\n });\n renewing.catch(fail);\n }, data.heartbeatIntervalMs);\n port.on('message', (message: { type?: string }) => {\n if (message.type === 'stop') void stop();\n });\n } catch (err) {\n fail(err);\n }\n}\n\nvoid main();\n"],"names":["parentPort","workerData","SenseError","serializeError","WatchClaimDatabase","requireParentPort","port","Error","data","database","timer","stopping","renewing","Promise","resolve","send","message","postMessage","asError","value","String","closeDatabase","clearInterval","undefined","close","err","addFailure","primary","secondary","AggregateError","fail","failure","release","token","releaseError","closeError","type","error","stop","main","configDir","acquire","pid","force","setInterval","then","renew","catch","heartbeatIntervalMs","on"],"mappings":"AAAA,SAASA,UAAU,EAAEC,UAAU,QAAQ,sBAAsB;AAC7D,SAASC,UAAU,QAAQ,eAAe;AAC1C,SAASC,cAAc,QAAQ,0BAA0B;AACzD,SAASC,kBAAkB,QAAiE,oBAAoB;AAEhH,SAASC;IACP,MAAMC,OAAON;IACb,IAAI,CAACM,MAAM,MAAM,IAAIC,MAAM;IAC3B,OAAOD;AACT;AAEA,MAAMA,OAAOD;AACb,MAAMG,OAAOP;AACb,IAAIQ;AACJ,IAAIC;AACJ,IAAIC,WAAW;AACf,IAAIC,WAAWC,QAAQC,OAAO;AAE9B,SAASC,KAAKC,OAAgC;IAC5CV,KAAKW,WAAW,CAACD;AACnB;AAEA,SAASE,QAAQC,KAAc;IAC7B,OAAOA,iBAAiBZ,QAAQY,QAAQ,IAAIZ,MAAMa,OAAOD;AAC3D;AAEA,SAASE;IACP,IAAIX,OAAOY,cAAcZ;IACzBA,QAAQa;IACR,IAAI;QACFd,qBAAAA,+BAAAA,SAAUe,KAAK;QACff,WAAWc;QACX,OAAOA;IACT,EAAE,OAAOE,KAAK;QACZ,OAAOP,QAAQO;IACjB;AACF;AAEA,SAASC,WAAWC,OAAc,EAAEC,SAA4B,EAAEZ,OAAe;IAC/E,OAAOY,YAAY,IAAIC,eAAe;QAACF;QAASC;KAAU,EAAEZ,WAAWW;AACzE;AAEA,SAASG,KAAKL,GAAY;IACxB,IAAId,UAAU;IACdA,WAAW;IACX,IAAIoB,UAAUb,QAAQO;IACtB,IAAI;QACFhB,qBAAAA,+BAAAA,SAAUuB,OAAO,CAACxB,KAAKyB,KAAK;IAC9B,EAAE,OAAOC,cAAc;QACrBH,UAAUL,WAAWK,SAASb,QAAQgB,eAAe;IACvD;IACA,MAAMC,aAAad;IACnBU,UAAUL,WAAWK,SAASI,YAAY;IAC1C,IAAI;QACFpB,KAAK;YAAEqB,MAAM;YAAWC,OAAOlC,eAAe4B;QAAS;IACzD,SAAU;QACRzB,KAAKkB,KAAK;IACZ;AACF;AAEA,eAAec;IACb,IAAI3B,UAAU;IACdA,WAAW;IACX,IAAID,OAAOY,cAAcZ;IACzB,IAAIqB;IACJ,IAAI;QACF,MAAMnB;QACN,IAAI,EAACH,qBAAAA,+BAAAA,SAAUuB,OAAO,CAACxB,KAAKyB,KAAK,IAAG,MAAM,IAAI/B,WAAW,gBAAgB;IAC3E,EAAE,OAAOuB,KAAK;QACZM,UAAUb,QAAQO;IACpB,SAAU;QACR,MAAMU,aAAad;QACnB,IAAIc,YAAYJ,UAAUA,UAAUL,WAAWK,SAASI,YAAY,+DAA+DA;QACnI,IAAI;YACF,IAAIJ,SAAShB,KAAK;gBAAEqB,MAAM;gBAAWC,OAAOlC,eAAe4B;YAAS;iBAC/DhB,KAAK;gBAAEqB,MAAM;YAAU;QAC9B,SAAU;YACR9B,KAAKkB,KAAK;QACZ;IACF;AACF;AAEA,eAAee;IACb,IAAI;QACF9B,WAAW,IAAIL,mBAAmBI,KAAKgC,SAAS;QAChD,MAAM/B,SAASgC,OAAO,CAACjC,KAAKyB,KAAK,EAAEzB,KAAKkC,GAAG,EAAElC,KAAKmC,KAAK;QACvD5B,KAAK;YAAEqB,MAAM;QAAW;QACxB1B,QAAQkC,YAAY;YAClB,IAAIjC,UAAU;YACdC,WAAWA,SAASiC,IAAI,CAAC;gBACvB,IAAI,EAACpC,qBAAAA,+BAAAA,SAAUqC,KAAK,CAACtC,KAAKyB,KAAK,IAAG,MAAM,IAAI/B,WAAW,gBAAgB;YACzE;YACAU,SAASmC,KAAK,CAACjB;QACjB,GAAGtB,KAAKwC,mBAAmB;QAC3B1C,KAAK2C,EAAE,CAAC,WAAW,CAACjC;YAClB,IAAIA,QAAQoB,IAAI,KAAK,QAAQ,KAAKE;QACpC;IACF,EAAE,OAAOb,KAAK;QACZK,KAAKL;IACP;AACF;AAEA,KAAKc"}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "sensemaking",
|
|
3
|
-
"version": "0.24.
|
|
3
|
+
"version": "0.24.5",
|
|
4
4
|
"description": "Query and search your markdown notes with context-aware progressive disclosure: SQL over frontmatter, links, and text, plus semantic search and link-graph ranking. No server, no build step.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"markdown",
|
|
@@ -93,7 +93,7 @@
|
|
|
93
93
|
"yaml": "^2.9.0"
|
|
94
94
|
},
|
|
95
95
|
"devDependencies": {
|
|
96
|
-
"@duckdb/node-api": "
|
|
96
|
+
"@duckdb/node-api": "1.5.5-r.4",
|
|
97
97
|
"@tursodatabase/database": "*",
|
|
98
98
|
"@types/markdown-it-footnote": "^3.0.4",
|
|
99
99
|
"@types/mocha": "*",
|
package/skills/sense/SKILL.md
CHANGED
|
@@ -1,136 +1,137 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: sense
|
|
3
|
-
description: "Query a markdown tree with the sense CLI: filter notes by frontmatter,
|
|
3
|
+
description: "Query a markdown tree with the sense CLI: filter notes by frontmatter, search prose, follow wikilinks and backlinks, trace link paths, find similar but unlinked notes, and inspect note outlines. Use when a task needs to query, filter, count, search, or report on a markdown directory; when a directory has sense.config.json; or when adding a saved sense query."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# sense
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Use sense to locate evidence in a markdown tree before reading files. Every command reconciles changed files first. Results contain paths, metadata, excerpts, and line ranges. Read the returned files or ranges when the task needs their prose.
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
Setup, store selection, presets, and note design belong to the `sense-setup` skill. Translating an Obsidian Bases file belongs to `sense-bases`.
|
|
11
11
|
|
|
12
|
-
##
|
|
12
|
+
## Start with the tree
|
|
13
13
|
|
|
14
|
-
|
|
14
|
+
Run these when the tree or its configuration is unfamiliar:
|
|
15
15
|
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
16
|
+
```sh
|
|
17
|
+
sense status # config, store, cache, document count, preset coverage
|
|
18
|
+
sense map # fields, hubs, recent notes
|
|
19
|
+
sense --list # saved queries
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
If a command reports a missing config, run `sense init` only when the user wants the tree configured. A one-off query does not authorize changing the tree.
|
|
23
|
+
|
|
24
|
+
## Choose the command
|
|
23
25
|
|
|
24
|
-
|
|
26
|
+
| Need | Command |
|
|
27
|
+
|---|---|
|
|
28
|
+
| Exact count, filter, grouping, or known-field report | `sense sql` or a saved SQL query |
|
|
29
|
+
| Notes about a subject | `sense search` |
|
|
30
|
+
| One note's frontmatter, outline, links, and backlinks | `sense peek` |
|
|
31
|
+
| Shortest link chain between two notes | `sense path` |
|
|
32
|
+
| Similar notes that a note does not link to | `sense related` |
|
|
33
|
+
| Tree shape and available fields | `sense map` |
|
|
25
34
|
|
|
26
|
-
|
|
35
|
+
Start with a bounded result. Raise `--k`, widen the preset, or broaden the words only when the first result does not answer the question.
|
|
27
36
|
|
|
28
|
-
|
|
37
|
+
```sh
|
|
38
|
+
sense search "pricing" --k 10
|
|
39
|
+
sense search "sourcing quotes" --preset raw
|
|
40
|
+
sense peek notes/pricing-model.md
|
|
41
|
+
sense path onboarding.md pricing-model.md
|
|
42
|
+
sense related notes/pricing-model.md --k 10
|
|
43
|
+
sense sql "SELECT path FROM frontmatter WHERE status = ? LIMIT 50" active
|
|
44
|
+
```
|
|
29
45
|
|
|
30
|
-
|
|
31
|
-
|---|---|---|---|
|
|
32
|
-
| `words` | the text itself (BM25) | every literal occurrence: identifiers, error strings, names, exact phrases; deterministic | anything phrased differently, e.g. a note saying "compensation floor" for the query "minimum pay" |
|
|
33
|
-
| `links` | connections authors wrote | what the tree's own structure treats as related: context the matching text never restates | notes nobody linked |
|
|
34
|
-
| `vectors` | per-chunk meaning-vectors | notes sharing no words with the query: paraphrase, the category for an instance, a concept restated | proving presence. A vector row cannot show the words occur anywhere; `snippets` can |
|
|
46
|
+
## Read the store guide before composing syntax
|
|
35
47
|
|
|
36
|
-
|
|
48
|
+
The config's `store` key selects the SQL dialect and text-search grammar. An omitted key means `sqlite`. Bare words and quoted phrases work in `sense search` on every store. Advanced operators and raw text-index SQL differ.
|
|
37
49
|
|
|
38
|
-
|
|
50
|
+
Read the matching guide before writing raw SQL that uses engine functions, dates, JSON, text matching, or native types. Also read it before using advanced search operators.
|
|
39
51
|
|
|
40
|
-
|
|
52
|
+
- [SQLite query guide](references/stores/sqlite.md)
|
|
53
|
+
- [DuckDB query guide](references/stores/duckdb.md)
|
|
54
|
+
- [Turso query guide](references/stores/turso.md)
|
|
41
55
|
|
|
42
|
-
|
|
56
|
+
The tables, `?` placeholders, quoted identifiers, `has()`, `basename()`, and preset `scope` binding are shared. [Portable SQL](references/sql.md) documents the schema and queries that work without depending on one engine.
|
|
43
57
|
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
58
|
+
## Search and evidence
|
|
59
|
+
|
|
60
|
+
`search` combines the signals enabled by the selected preset:
|
|
61
|
+
|
|
62
|
+
| Signal | Evidence |
|
|
63
|
+
|---|---|
|
|
64
|
+
| `match` | The note contains the search words. `snippets` shows the matching passages. |
|
|
65
|
+
| `link` | A note that matched links to this note. |
|
|
66
|
+
| `vector` | The embedding model placed this note near the query. The search words may be absent. |
|
|
67
|
+
|
|
68
|
+
Combinations such as `match+link` mean that more than one signal produced the row. `score` ranks rows within that result only. Do not compare it across searches. `similarity` ranks vector evidence within the current result and model. Do not carry a fixed similarity cutoff between trees.
|
|
69
|
+
|
|
70
|
+
The `lines` value points at the section that earned the row. Read that range when it is present. A null range means the whole note is the reference. A vector-only row has no lexical snippet and is a lead, not proof that the note contains the query terms.
|
|
71
|
+
|
|
72
|
+
Read [search evidence and troubleshooting](references/search.md) when a search will support a factual claim, when absence matters, when results look noisy, or when tuning signal weights.
|
|
73
|
+
|
|
74
|
+
## Scope and output
|
|
75
|
+
|
|
76
|
+
Bare commands use the `default` preset. `--preset <name>` chooses another. For `search`, `--include`, `--exclude`, and `--no-exclude` change the query scope for one invocation, but cannot reach files that no preset indexes. `sense status` shows actual coverage.
|
|
77
|
+
|
|
78
|
+
`--where` filters search and graph commands against frontmatter alias `f`:
|
|
79
|
+
|
|
80
|
+
```sh
|
|
81
|
+
sense search "pricing" --where "f.status = 'active' AND has(f.tags, 'sales')"
|
|
54
82
|
```
|
|
55
83
|
|
|
56
|
-
|
|
57
|
-
- When a search misses, both sides have levers. Lexical: OR-in synonyms and concrete instances (the index only knows the words in the files; a note about a specific tool rarely names its category), raise `--k` (a row costs tens of tokens), widen the scope (`--preset`, or `--include` for an ad-hoc glob). Vector: restate the concept in different words. Vectors rank the meaning of the whole query text, so a rephrase moves them even when every literal word still misses. Then pivot through the nearest hit with `related <path>` to walk its meaning-neighbourhood. Vector-only rows in a miss are leads rather than lexical evidence: use `similarity` and the `lines` range to decide what to read; their `snippets` list is empty because the query words did not match. Each widening adds candidates and dilutes ranking, so the noise trade-off runs both ways.
|
|
58
|
-
- A frontmatter query enumerates its matches deterministically; search ranks by term overlap, so results shift as phrasing shifts. Trade-off: a query needs a known field, search doesn't.
|
|
59
|
-
- The `via` column says what produced each row: `match` (words hit), `link` (connected to notes that hit), `vector` (near in meaning), and combinations. The `lines` column, when set, points at the section that earned the row (the best-matching chunk on vector rows, the term cluster's section on large lexical notes) and is a direct `Read` range; null means the whole note is the reference. A row's `snippets` holds the marked passages that earned the lexical match (`[]` on a row that never contained the terms); `--snippet-char-limit` (default 80, the passage's rendered length, marks and ellipses counted) and `--snippet-count-limit` (default 1, passages per note) bound them, on `search` and a saved search alike.
|
|
60
|
-
- Scope is one vocabulary shared by `search`, `map`, `peek`, `path`, and `related`: bare command uses the config's `default` preset; `--preset <name>` picks another (unknown names error, listing what's declared); `--include <glob>` and `--exclude <glob>` are ad-hoc globs for one command, each overriding its own side of the preset, so one does not clear the other; `--no-exclude` drops the preset's `exclude` for one command, the only way to widen past it without editing config (it widens the query scope, not the index: a file no preset covers is never indexed). `--where` takes any SQL condition against frontmatter alias `f`, not only field equality: `"f.status = 'active' AND has(f.tags, 'x')"`, `"datetime(f.created) >= datetime(?)"`. There is no whole-index flag: a broad `default` preset, or a declared `all` preset (`include ["**/*"]`), is the whole tree. `sense status` shows every preset with its coverage. `sql` scopes differently: it runs over the whole index by default, and `--preset <name>` *binds* the scope as a temporary `scope(path)` table your statement joins, rather than filtering behind the query's back (`JOIN scope ON scope."path" = f.path`). Naming a preset without joining `scope` is a usage error, since it would return everything while reading as scoped. Without the flag, join `preset_files` directly, which is the same coverage under a preset name you write into the SQL.
|
|
61
|
-
- `score` is a rank-fusion value: it ranks rows within one result set and is not comparable across queries, not a relevance magnitude. It encodes how many signals fired and at what rank, so a perfect lexical hit and a weak vector-only hit can read the same number. With vectors active, rows carry `similarity`: the cosine (-1 to 1) of the query against that file's best-matching chunk (the same chunk the `lines` range points at). It orders vector evidence within a result set; the range it spans depends on the corpus and the embedding model, and compresses on small trees, where even a nonsense query has a moderately near neighbour somewhere. Compare similarities within a result set rather than against a fixed cutoff carried between trees.
|
|
62
|
-
- Vector rankings read the whole chunk, boilerplate included, so on a tree whose notes are mostly one template with a line of unique text (a directory of plugins, people, or assets), the shared scaffolding dominates every vector and `related` returns near-ties at the top of the range for any seed. Two signs, both visible in the output already: the top `similarity` values sit within a hair of each other, and the same few notes come back for unrelated seeds. Compare neighbour lists across two unlike seeds when a tree looks like this; matching lists mean the ranking is reading the template, not the content, and `search` over the distinguishing words is the answer instead.
|
|
63
|
-
- Absence evidence lives in the labels: a preset whose `signals` exclude `vectors` (or a tree with no `embed` block) returns 0 rows when the words are nowhere in it. Default `search` always returns up to `k` rows (nearest-neighbour search has a nearest neighbour for any input), so a result of only `via: vector` rows is the absence signal for the words themselves. Judge whether a vector row is a useful conceptual lead from `similarity`, then read its `lines` range; it has no lexical snippet.
|
|
64
|
-
- A `queries` entry names the verb it runs, mirroring the two commands: `"dead-links": { "sql": "SELECT src, target FROM links WHERE dst IS NULL AND lower(target) NOT GLOB '*.[a-z0-9]*'" }` runs as `sense dead-links`, and `"hot": { "search": "pricing OR billing", "preset": "raw", "k": 20 }` runs as `sense hot` with its settings baked in, so repeat runs need no flags. An invocation-level `--preset`, `--k`, or `--where` overrides a saved search's value; `--list` labels each entry `(sql)` or `(search)`.
|
|
65
|
-
- Running an entry is how it is validated: a typo'd column, stale SQL, bad FTS5 syntax or an unknown preset errors and exits nonzero. A parameterised entry validates with any argument, since SQL is prepared before parameters bind (`sense by-tag zzz` reports `no such column` if the column is wrong, `(0 rows)` if it is right). To sweep a whole config after editing it, read the exit code: 0 ran, 2 means it needs parameters (re-run it with any argument to validate the SQL), anything else is broken.
|
|
84
|
+
`sense sql` is index-wide by default. With `--preset`, the command binds a temporary `scope(path)` table. The SQL must join it:
|
|
66
85
|
|
|
67
86
|
```sh
|
|
68
|
-
|
|
69
|
-
sense "$q" --format json >/dev/null 2>&1
|
|
70
|
-
case $? in 0) ;; 2) echo "needs an argument: $q" ;; *) echo "broken: $q" ;; esac
|
|
71
|
-
done
|
|
87
|
+
sense sql "SELECT f.path FROM frontmatter f JOIN scope ON scope.path = f.path" --preset default
|
|
72
88
|
```
|
|
73
89
|
|
|
74
|
-
|
|
90
|
+
Table output is for people. Use `--format json` when code or an agent will parse rows. Use `--format csv` when redirecting a large row set to a file. `sql`, `search`, `related`, and saved queries support all three formats. `map`, `peek`, `status`, and `path` support table and JSON.
|
|
75
91
|
|
|
76
|
-
##
|
|
92
|
+
## Saved queries
|
|
77
93
|
|
|
78
|
-
|
|
94
|
+
Save a query in `sense.config.json` only when it will be reused:
|
|
79
95
|
|
|
80
|
-
```
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
sense sql "SELECT path FROM frontmatter WHERE path NOT IN (SELECT dst FROM links WHERE dst IS NOT NULL) AND path NOT IN (SELECT src FROM links)" # linked neither way (fine if intentional; linking is optional)
|
|
88
|
-
sense sql "SELECT src, target FROM links WHERE dst IS NULL AND lower(target) NOT GLOB '*.[a-z0-9]*'" # broken wikilinks, attachments excluded
|
|
89
|
-
sense sql "SELECT heading, start_line, tokens FROM sections WHERE path = ?" a.md # budget a read
|
|
90
|
-
sense sql "SELECT j.value, COUNT(*) n FROM frontmatter, json_each(frontmatter.tags) j GROUP BY j.value ORDER BY n DESC" # count per array member
|
|
96
|
+
```json
|
|
97
|
+
{
|
|
98
|
+
"queries": {
|
|
99
|
+
"by-tag": { "sql": "SELECT path, title FROM frontmatter WHERE has(tags, ?) ORDER BY path" },
|
|
100
|
+
"hot": { "search": "pricing", "preset": "raw", "k": 20 }
|
|
101
|
+
}
|
|
102
|
+
}
|
|
91
103
|
```
|
|
92
104
|
|
|
93
|
-
|
|
105
|
+
Run these as `sense by-tag urgent` and `sense hot`. Invocation flags override a saved search's `preset`, `k`, or `where` value.
|
|
94
106
|
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
107
|
+
Running a saved entry validates it. Exit code 0 means it ran, 2 means the invocation needs different arguments, and 1 means the query or store failed. An empty result can be valid data, so interpret it from the query's purpose.
|
|
108
|
+
|
|
109
|
+
## Tables
|
|
110
|
+
|
|
111
|
+
| Table | Holds |
|
|
112
|
+
|---|---|
|
|
113
|
+
| `frontmatter` | One discovered column per frontmatter key, plus `path`, `_mtime`, `_ctime`, `_size`, `_rank`, and `_parse_error` |
|
|
114
|
+
| `content` | `path`, `title`, `summary`, and authored text, plus store-owned search columns |
|
|
115
|
+
| `links` | `src`, written `target`, resolved `dst`, and `embed` |
|
|
116
|
+
| `tags` | Merged and deduplicated frontmatter and inline tags |
|
|
117
|
+
| `sections` | Heading, level, line range, and token estimate |
|
|
118
|
+
| `preset_files` | Paths covered by each preset |
|
|
119
|
+
|
|
120
|
+
Features can add tables. `sense map` and `sense status` show which features are active.
|
|
121
|
+
|
|
122
|
+
A non-null `_parse_error` means the file has no recovered frontmatter values. To distinguish a missing field from invalid frontmatter, include `_parse_error IS NULL` in the filter. Fixes appear on the next command because reconciliation runs first.
|
|
123
|
+
|
|
124
|
+
## Reading discipline
|
|
125
|
+
|
|
126
|
+
Select only the columns needed for the answer. Use `LIMIT` for row-returning exploration. Prefer `path`, `title`, `summary`, and bounded snippets over `content.text`. Aggregates such as `COUNT` and `GROUP BY` are already bounded by their result shape.
|
|
127
|
+
|
|
128
|
+
When a result identifies a large note, use `peek` and then read the relevant line range. Small files are often cheaper to read whole.
|
|
129
|
+
|
|
130
|
+
Worked command traces are in [EXAMPLES.md](EXAMPLES.md).
|
|
131
|
+
|
|
132
|
+
## Upkeep
|
|
110
133
|
|
|
111
|
-
-
|
|
112
|
-
- `
|
|
113
|
-
-
|
|
114
|
-
-
|
|
115
|
-
- Rank with `ORDER BY bm25(content, 10.0, 5.0, 1.0)` (title > summary > body); the full form `bm25(content, 10.0, 5.0, 1.0, 0, 10.0, 5.0, 1.0)` mirrors the same weights onto the `_seg` sidecars, so a title hit found through `title_seg` ranks like one found through `title` (the three-weight form still runs, FTS5 defaults unnamed columns to 1.0, but ranks a sidecar match at body weight). In hand-written SQLite SQL, `snippet(content, 2, '«', '»', '…', 10)` names the authored `text` column explicitly; `-1` can surface a machine-spaced sidecar instead. `search` does not call SQLite's `snippet()`: it computes bounded passages in shared code for every store. A raw SQL `snippet()` re-tokenizes its matched document, so guard large text with `CASE WHEN length(text) <= 16384 THEN snippet(...) END`, or select `title`/`summary`. Treat historical timing as diagnostic until a current identical-work sitting replaces it.
|
|
116
|
-
- Select `content.title`/`content.summary` (always exist, empty when absent) rather than `f.title`/`f.summary` (discovered columns; error on trees that never declare them).
|
|
117
|
-
- Frontmatter values keep their YAML type: strings are TEXT, whole numbers and booleans are INTEGER (`true` stores as 1, so `WHERE flag = 1` matches and `WHERE flag = 'true'` matches nothing), fractions are REAL, lists and maps are JSON text. On the `duckdb` store the discovered columns are VARIANT: homogeneous keys, which are nearly all of them, compare identically, but a numeric predicate against a key that holds numbers in some notes and text in others raises a comparison error where sqlite orders by storage class silently. `map` prints the observed type per field, and a field showing two types (`integer,text`) has drifted across notes. A list key written with no items (`tags:` above a bare `-`) is a list holding one null, stored as the JSON text `[null]`: `IS NULL` does not find it (the column holds a string), `json_each` yields one empty member per such row, and `map` counts it as covered because the key is present. `has(tags, 'x')` reads it correctly as no match. To separate written-but-empty from absent, compare against the text: `WHERE tags = '[null]'`.
|
|
118
|
-
- **Dead links need the attachment filter.** `dst IS NULL` alone is not "broken link": a wikilink to anything that is not markdown (`[[Board.base]]`, `![[Pasted image.png]]`, `[[spec.pdf]]`) can never resolve, because sense indexes markdown and resolution only tries the exact path or `+.md`. Those are out of the index's universe, not broken. On a 1,400-note Obsidian vault the unfiltered query returns 143 rows where 14 are real. Exclude anything carrying a file extension, as in the recipe above, and widen the exclusion if your notes have dotted titles (`[[Node.js]]` carries one too, so a stricter list, `'*.png'`, `'*.pdf'`, `'*.base'`, and whatever else your vault attaches, is safer on a tree whose titles use dots). Scope it with `preset_files` as well: template and skill files are full of `[[Note Name]]` examples that are deliberately unresolved.
|
|
119
|
-
- `has(field, value)`: array membership on JSON-array fields, substring on strings, false on NULL. This is the `includes()` convention. Substring means `has(f.status, 'active')` also matches `inactive`; exact scalar match is `f.status = ?`, deliberate substring is `LIKE`, exact array membership is `EXISTS (SELECT 1 FROM json_each(f.tags) WHERE value = ?)`. To aggregate per member instead, use `json_each(frontmatter.<field>)` (above). GROUP BY on the raw column splits `["a","b"]` and `["b","a"]` into separate buckets.
|
|
120
|
-
- Compare dates through `datetime()`, which resolves ISO 8601 offsets to UTC: `WHERE datetime(created) >= datetime(?)`. Bare string comparison is only safe when every note uses the same offset.
|
|
121
|
-
- Date spellings SQLite rejects (`-0800`, `-08`, a space separator) are normalized at index time, offset preserved. One it cannot fix is left as written and warned about by path: `datetime()` returns NULL there, so the row is invisible to a date comparison rather than excluded by it. List them with `WHERE d IS NOT NULL AND datetime(d) IS NULL`.
|
|
122
|
-
- **SQLite's `now` is UTC, so any query about "today" needs `'localtime'`.** `date('now')` reads as tomorrow from mid-afternoon onward in the Americas, which silently flips "scheduled today" into "overdue" every evening: write `date('now','localtime')` and `datetime('now','start of day','localtime')`. This only matters where the boundary carries the meaning; a `'-90 day'` window is unaffected by a few hours of skew.
|
|
123
|
-
- To bound what a query puts into context: `snippet()` excerpts just the matching text, `LIMIT` caps row counts, and selecting `path`/`title`/`summary` keeps rows small. A large result can also stay out of context entirely: `--format csv > file` writes it whole, and `grep`/`awk` over that file returns only the rows wanted. `SELECT text FROM content` returns the tree's entire prose (sense warns past 50 KB, after the rows have already printed, so the warning records the cost rather than preventing it). Aggregates (`COUNT`, `GROUP BY`) are already bounded. `SELECT * FROM frontmatter` is always safe: prose is not a frontmatter column.
|
|
124
|
-
|
|
125
|
-
Worked traces: [EXAMPLES.md](EXAMPLES.md).
|
|
126
|
-
|
|
127
|
-
## Setup and upkeep
|
|
128
|
-
|
|
129
|
-
- Missing CLI: `npm install -g sensemaking`. Missing config: `sense init` at the tree root. Discovery walks up from cwd; `--config <path>` overrides. Setting up or restructuring a tree (presets, frontmatter conventions, note design) is the `sense-setup` skill. Translating an Obsidian Bases `.base` file into equivalent queries is the `sense-bases` skill.
|
|
130
|
-
- `map` and `status` report each preset's coverage (files matched, embedded count). Indexing derives from presets, so the coverage numbers are how you see what a config actually indexes and embeds. A scope with fewer signals just uses fewer (a preset without the vectors signal searches lexically); a saved search naming an unknown preset errors when run, listing the declared ones.
|
|
131
|
-
- Save a query into `sense.config.json` only when it will be reused; run ad-hoc otherwise.
|
|
132
|
-
- A one-line `summary:` per note is optional and pays twice: it appears in result rows and is a weighted search field. Date comparisons work for dates written as ISO 8601 (`2026-08-12`, or with time and offset); other formats do not compare. Field names in examples (`status`, `tags`, `created`) are illustrative; your tree defines its own.
|
|
133
|
-
- Reserved frontmatter keys (dropped with a warning): `path`, `_mtime`, `_ctime`, `_size`, `_rank`, `_parse_error`, `content`, `links`, `sections`. The `tags` frontmatter column and the `tags` table coexist, mirroring Obsidian's own split: the column is the raw YAML list one note's frontmatter declares (Obsidian's `tags` property), the table is the merged, deduplicated frontmatter+inline set per note (what Obsidian's tag pane and Bases' `file.tags` read). "What is tagged X" is a table query; the column answers only what a note's frontmatter literally says. Inline tags inside `%%...%%` comments are indexed, and some trees run their whole maintenance-tag system in comments.
|
|
134
|
-
- A note whose frontmatter does not parse is indexed with **no** frontmatter columns and `_parse_error` set to the YAML message, which carries the line. Nothing is half-recovered: a non-NULL value is a value the author wrote. So a NULL column means the key was absent *or* the note did not parse, and `_parse_error` is how you tell: `WHERE status IS NULL AND _parse_error IS NULL` is "genuinely missing status". List what needs fixing with `sense sql "SELECT path, _parse_error FROM frontmatter WHERE _parse_error IS NOT NULL"`; fixing a file clears it on the next command. `sense status` reports the count.
|
|
135
|
-
- Exit codes: `0` ok, `1` error (store message verbatim), `2` usage (unknown query, wrong param count).
|
|
136
|
-
- Doubted cache: delete the directory `sense status` prints on its `cache:` line. Rarely needed; every query reconciles first.
|
|
134
|
+
- Install a missing CLI with `npm install -g sensemaking`.
|
|
135
|
+
- `sense status` prints the cache path and watcher state.
|
|
136
|
+
- Delete the cache directory printed by `sense status` only when the derived index is in doubt. The next command rebuilds it.
|
|
137
|
+
- Use `sense watch` when another process should keep the index warm during frequent edits. Queries remain responsible for their own freshness check.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Search evidence and troubleshooting
|
|
2
|
+
|
|
3
|
+
Read this guide when search results will support a factual claim, when absence matters, when results look noisy, or when changing signal weights.
|
|
4
|
+
|
|
5
|
+
## What each signal establishes
|
|
6
|
+
|
|
7
|
+
`words` ranks literal occurrences after stemming. It is the right signal for identifiers, error strings, names, and quoted phrases. A `match` row is lexical evidence because its `snippets` show the occurrence.
|
|
8
|
+
|
|
9
|
+
`links` expands from matching notes through links written by the authors. A `link` row establishes that relationship, but does not establish that the row contains the query words.
|
|
10
|
+
|
|
11
|
+
`vectors` ranks the meaning of chunks. It can find paraphrases and related concepts with no shared vocabulary. A `vector` row is a lead to read. It cannot prove that a word or claim appears in the note.
|
|
12
|
+
|
|
13
|
+
Search uses every signal named by the preset. If `signals` is absent, every signal whose prerequisites hold has weight 1. A number changes that signal's contribution to reciprocal-rank fusion:
|
|
14
|
+
|
|
15
|
+
```json
|
|
16
|
+
"signals": { "words": 1, "links": 1, "vectors": 4 }
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Weights are corpus and model choices. Compare representative queries before changing them. One weight does not transfer reliably between unrelated trees or embedding models.
|
|
20
|
+
|
|
21
|
+
## Reading the rows
|
|
22
|
+
|
|
23
|
+
- `via` names the evidence that produced the row.
|
|
24
|
+
- `snippets` contains marked lexical passages. It is empty when no word matched.
|
|
25
|
+
- `lines` names the best section to read. Null means the whole note is the reference.
|
|
26
|
+
- `score` orders the fused result. Its scale changes with the participating signals and ranks, so compare rows only inside one result.
|
|
27
|
+
- `similarity` is cosine similarity against the best chunk. Compare it inside one result and one model. Small trees can give unrelated text a moderately close nearest neighbor.
|
|
28
|
+
|
|
29
|
+
Relay the evidence label when reporting a result. "The note contains these words" and "the note is semantically related" support different claims.
|
|
30
|
+
|
|
31
|
+
## When a search misses
|
|
32
|
+
|
|
33
|
+
For word search, try concrete terms that the notes may use, widen the selected preset, or raise `--k`. Search grammar depends on the selected store. Read its guide before adding operators.
|
|
34
|
+
|
|
35
|
+
For vector search, restate the concept in different words. Then use `related <path>` on the nearest useful hit to inspect its semantic neighborhood.
|
|
36
|
+
|
|
37
|
+
Each widening step adds candidates and can dilute the ranking. Inspect the new rows before widening again.
|
|
38
|
+
|
|
39
|
+
## Absence
|
|
40
|
+
|
|
41
|
+
A words-only preset returns no rows when the terms do not occur in its indexed scope. With vectors enabled, nearest-neighbor search can still return vector-only rows for any input. Those rows show conceptual proximity, not lexical presence.
|
|
42
|
+
|
|
43
|
+
When absence matters, use a preset whose `signals` excludes `vectors`, or query the selected store's text index as described in its guide. Confirm the preset's coverage with `sense status` before concluding that the terms are absent from the tree.
|
|
44
|
+
|
|
45
|
+
## Template-heavy trees
|
|
46
|
+
|
|
47
|
+
Vectors rank whole chunks, including repeated boilerplate. In a tree whose notes share a large template and contain little unique prose, unrelated seeds can return the same neighbors with almost identical similarities. Compare two unlike seed notes. If their neighbor lists barely change, search the fields or words that distinguish the notes instead.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# Portable SQL
|
|
2
|
+
|
|
3
|
+
Read this guide for raw SQL over sense tables. These patterns use the shared schema and avoid native text-index syntax. Read the selected store's query guide as well when the statement uses dates, JSON table functions, native types, or text matching.
|
|
4
|
+
|
|
5
|
+
## Discover the tree before assuming fields
|
|
6
|
+
|
|
7
|
+
Frontmatter columns come from the indexed notes. Inspect them before writing a field query:
|
|
8
|
+
|
|
9
|
+
```sh
|
|
10
|
+
sense map
|
|
11
|
+
sense sql "SELECT name FROM pragma_table_info('frontmatter')"
|
|
12
|
+
sense sql "SELECT DISTINCT status FROM frontmatter ORDER BY status"
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Punctuated field names need double quotes. Values should use `?` parameters:
|
|
16
|
+
|
|
17
|
+
```sh
|
|
18
|
+
sense sql 'SELECT path FROM frontmatter WHERE "plugin-id" = ? LIMIT 50' example
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
`content.title` and `content.summary` always exist. A frontmatter column with either name exists only when at least one indexed note declares it.
|
|
22
|
+
|
|
23
|
+
## Shared tables
|
|
24
|
+
|
|
25
|
+
| Table | Main columns |
|
|
26
|
+
|---|---|
|
|
27
|
+
| `frontmatter` | `path`, discovered fields, `_mtime`, `_ctime`, `_size`, `_rank`, `_parse_error` |
|
|
28
|
+
| `content` | `path`, `title`, `summary`, `text` |
|
|
29
|
+
| `links` | `src`, `target`, `dst`, `embed` |
|
|
30
|
+
| `tags` | `path`, `tag` |
|
|
31
|
+
| `sections` | `path`, `heading`, `level`, `start_line`, `end_line`, `tokens` |
|
|
32
|
+
| `preset_files` | `path`, `preset` |
|
|
33
|
+
|
|
34
|
+
`links.dst` is null when the written target does not resolve to indexed markdown. `links.embed` is 1 for an embed and 0 for an ordinary wikilink. The `tags` table merges frontmatter and inline tags and stores a nested tag such as `book/scifi` in full.
|
|
35
|
+
|
|
36
|
+
## Shared query patterns
|
|
37
|
+
|
|
38
|
+
```sql
|
|
39
|
+
SELECT COUNT(*) AS notes FROM frontmatter;
|
|
40
|
+
|
|
41
|
+
SELECT path, title
|
|
42
|
+
FROM frontmatter
|
|
43
|
+
WHERE status = ?
|
|
44
|
+
ORDER BY path
|
|
45
|
+
LIMIT 50;
|
|
46
|
+
|
|
47
|
+
SELECT src
|
|
48
|
+
FROM links
|
|
49
|
+
WHERE dst = ?
|
|
50
|
+
ORDER BY src;
|
|
51
|
+
|
|
52
|
+
SELECT heading, start_line, end_line, tokens
|
|
53
|
+
FROM sections
|
|
54
|
+
WHERE path = ?
|
|
55
|
+
ORDER BY start_line;
|
|
56
|
+
|
|
57
|
+
SELECT path, tag
|
|
58
|
+
FROM tags
|
|
59
|
+
WHERE tag = ? OR tag LIKE ?
|
|
60
|
+
ORDER BY path;
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
For a nested tag family, bind the same value twice as `book` and `book/%`.
|
|
64
|
+
|
|
65
|
+
`has(field, value)` works on every store. It tests membership for a JSON array and substring presence for a scalar string. Use `field = ?` for exact scalar equality because `has(status, 'active')` also matches `inactive`.
|
|
66
|
+
|
|
67
|
+
`basename(path[, suffix])` works on every store. `segment(terms)` is available on SQLite and DuckDB only; the store guides explain when it is needed.
|
|
68
|
+
|
|
69
|
+
## Preset scope
|
|
70
|
+
|
|
71
|
+
`sense sql` covers the whole index unless it receives `--preset`. With that flag, sense binds a temporary `scope(path)` table and requires the statement to join it:
|
|
72
|
+
|
|
73
|
+
```sh
|
|
74
|
+
sense sql "SELECT f.path FROM frontmatter f JOIN scope ON scope.path = f.path ORDER BY f.path" --preset default
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
A saved SQL query written against `scope` can run under different presets without duplicating the statement. Without `--preset`, join `preset_files` directly and name the preset in SQL.
|
|
78
|
+
|
|
79
|
+
## Graph queries
|
|
80
|
+
|
|
81
|
+
Bound recursive walks. An unrestricted walk on a dense graph can enumerate paths faster than it eliminates them.
|
|
82
|
+
|
|
83
|
+
```sql
|
|
84
|
+
WITH RECURSIVE hop(path, d) AS (
|
|
85
|
+
SELECT ?, 0
|
|
86
|
+
UNION
|
|
87
|
+
SELECT CASE WHEN l.src = hop.path THEN l.dst ELSE l.src END, hop.d + 1
|
|
88
|
+
FROM hop
|
|
89
|
+
JOIN links l ON (l.src = hop.path OR l.dst = hop.path) AND l.dst IS NOT NULL
|
|
90
|
+
WHERE hop.d < 2
|
|
91
|
+
)
|
|
92
|
+
SELECT DISTINCT path FROM hop WHERE d > 0;
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
Use `sense path` for the shortest chain between two known notes. Use raw recursion when the task needs a set of neighbors to filter or join.
|
|
96
|
+
|
|
97
|
+
## Dead links
|
|
98
|
+
|
|
99
|
+
`dst IS NULL` includes links to attachments that sense never indexes, such as images, PDFs, and `.base` files. Exclude the attachment extensions used by the tree before treating the remaining rows as broken links. Trees with dotted markdown titles need an explicit extension list instead of a blanket "contains a dot" filter.
|
|
100
|
+
|
|
101
|
+
Templates and examples can also contain deliberately unresolved wikilinks. Scope the query to authored content when those files are indexed.
|
|
102
|
+
|
|
103
|
+
## Types and parse errors
|
|
104
|
+
|
|
105
|
+
YAML strings remain text, whole numbers and booleans remain integer-like values, fractions remain real-like values, and lists and maps remain structured or JSON-backed values according to the store. Use `sense map` to see the types observed for each field. Read the store guide before comparing a field that contains more than one type.
|
|
106
|
+
|
|
107
|
+
A malformed frontmatter block produces `_parse_error` and no recovered field values. This query lists the files to fix:
|
|
108
|
+
|
|
109
|
+
```sql
|
|
110
|
+
SELECT path, _parse_error
|
|
111
|
+
FROM frontmatter
|
|
112
|
+
WHERE _parse_error IS NOT NULL
|
|
113
|
+
ORDER BY path;
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
To find notes that genuinely omit `status`, use `status IS NULL AND _parse_error IS NULL`.
|
|
117
|
+
|
|
118
|
+
## Bound the output
|
|
119
|
+
|
|
120
|
+
Select only the columns required by the task and add `LIMIT` while exploring. Avoid `SELECT text FROM content` because it returns the tree's prose. Use `search` for bounded passages, or select `path`, `title`, and `summary` and read the relevant files afterward.
|