sensemaking 0.24.4 → 0.24.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (123) hide show
  1. package/README.md +16 -12
  2. package/dist/cjs/cli/status.js +14 -24
  3. package/dist/cjs/cli/status.js.map +1 -1
  4. package/dist/cjs/commands/search.js +22 -0
  5. package/dist/cjs/commands/search.js.map +1 -1
  6. package/dist/cjs/config/types.js.map +1 -1
  7. package/dist/cjs/errors.d.cts +1 -1
  8. package/dist/cjs/errors.d.ts +1 -1
  9. package/dist/cjs/errors.js.map +1 -1
  10. package/dist/cjs/lib/worker-file.d.cts +1 -0
  11. package/dist/cjs/lib/worker-file.d.ts +1 -0
  12. package/dist/cjs/lib/worker-file.js +34 -0
  13. package/dist/cjs/lib/worker-file.js.map +1 -0
  14. package/dist/cjs/output/search-error.js +38 -0
  15. package/dist/cjs/output/search-error.js.map +1 -1
  16. package/dist/cjs/scan/pool.js +2 -23
  17. package/dist/cjs/scan/pool.js.map +1 -1
  18. package/dist/cjs/store/builder.d.cts +2 -1
  19. package/dist/cjs/store/builder.d.ts +2 -1
  20. package/dist/cjs/store/builder.js +4 -4
  21. package/dist/cjs/store/builder.js.map +1 -1
  22. package/dist/cjs/store/duckdb/lexical.d.cts +1 -0
  23. package/dist/cjs/store/duckdb/lexical.d.ts +1 -0
  24. package/dist/cjs/store/duckdb/lexical.js +28 -9
  25. package/dist/cjs/store/duckdb/lexical.js.map +1 -1
  26. package/dist/cjs/store/duckdb/native.d.cts +2 -0
  27. package/dist/cjs/store/duckdb/native.d.ts +2 -0
  28. package/dist/cjs/store/duckdb/native.js +11 -1
  29. package/dist/cjs/store/duckdb/native.js.map +1 -1
  30. package/dist/cjs/store/duckdb/ordered-bm25.d.cts +17 -0
  31. package/dist/cjs/store/duckdb/ordered-bm25.d.ts +17 -0
  32. package/dist/cjs/store/duckdb/ordered-bm25.js +238 -0
  33. package/dist/cjs/store/duckdb/ordered-bm25.js.map +1 -0
  34. package/dist/cjs/store/lock-wait.d.cts +2 -2
  35. package/dist/cjs/store/lock-wait.d.ts +2 -2
  36. package/dist/cjs/store/lock-wait.js +14 -11
  37. package/dist/cjs/store/lock-wait.js.map +1 -1
  38. package/dist/cjs/store/native.d.cts +4 -0
  39. package/dist/cjs/store/native.d.ts +4 -0
  40. package/dist/cjs/store/native.js +64 -9
  41. package/dist/cjs/store/native.js.map +1 -1
  42. package/dist/cjs/store/open.js +11 -7
  43. package/dist/cjs/store/open.js.map +1 -1
  44. package/dist/cjs/store/reconcile.d.cts +5 -1
  45. package/dist/cjs/store/reconcile.d.ts +5 -1
  46. package/dist/cjs/store/reconcile.js +5 -4
  47. package/dist/cjs/store/reconcile.js.map +1 -1
  48. package/dist/cjs/store/turso/lexical.js +5 -1
  49. package/dist/cjs/store/turso/lexical.js.map +1 -1
  50. package/dist/cjs/text/segment.js +168 -3
  51. package/dist/cjs/text/segment.js.map +1 -1
  52. package/dist/cjs/watch-claim.d.cts +39 -0
  53. package/dist/cjs/watch-claim.d.ts +39 -0
  54. package/dist/cjs/watch-claim.js +446 -0
  55. package/dist/cjs/watch-claim.js.map +1 -0
  56. package/dist/cjs/watch.js +240 -141
  57. package/dist/cjs/watch.js.map +1 -1
  58. package/dist/cjs/workers/watch-heartbeat.d.cts +1 -0
  59. package/dist/cjs/workers/watch-heartbeat.d.ts +1 -0
  60. package/dist/cjs/workers/watch-heartbeat.js +308 -0
  61. package/dist/cjs/workers/watch-heartbeat.js.map +1 -0
  62. package/dist/esm/cli/status.js +6 -5
  63. package/dist/esm/cli/status.js.map +1 -1
  64. package/dist/esm/commands/search.js +21 -0
  65. package/dist/esm/commands/search.js.map +1 -1
  66. package/dist/esm/config/types.js.map +1 -1
  67. package/dist/esm/errors.d.ts +1 -1
  68. package/dist/esm/errors.js.map +1 -1
  69. package/dist/esm/lib/worker-file.d.ts +1 -0
  70. package/dist/esm/lib/worker-file.js +24 -0
  71. package/dist/esm/lib/worker-file.js.map +1 -0
  72. package/dist/esm/output/search-error.js +21 -0
  73. package/dist/esm/output/search-error.js.map +1 -1
  74. package/dist/esm/scan/pool.js +2 -22
  75. package/dist/esm/scan/pool.js.map +1 -1
  76. package/dist/esm/store/builder.d.ts +2 -1
  77. package/dist/esm/store/builder.js +4 -4
  78. package/dist/esm/store/builder.js.map +1 -1
  79. package/dist/esm/store/duckdb/lexical.d.ts +1 -0
  80. package/dist/esm/store/duckdb/lexical.js +16 -7
  81. package/dist/esm/store/duckdb/lexical.js.map +1 -1
  82. package/dist/esm/store/duckdb/native.d.ts +2 -0
  83. package/dist/esm/store/duckdb/native.js +5 -1
  84. package/dist/esm/store/duckdb/native.js.map +1 -1
  85. package/dist/esm/store/duckdb/ordered-bm25.d.ts +17 -0
  86. package/dist/esm/store/duckdb/ordered-bm25.js +56 -0
  87. package/dist/esm/store/duckdb/ordered-bm25.js.map +1 -0
  88. package/dist/esm/store/lock-wait.d.ts +2 -2
  89. package/dist/esm/store/lock-wait.js +17 -14
  90. package/dist/esm/store/lock-wait.js.map +1 -1
  91. package/dist/esm/store/native.d.ts +4 -0
  92. package/dist/esm/store/native.js +60 -6
  93. package/dist/esm/store/native.js.map +1 -1
  94. package/dist/esm/store/open.js +9 -6
  95. package/dist/esm/store/open.js.map +1 -1
  96. package/dist/esm/store/reconcile.d.ts +5 -1
  97. package/dist/esm/store/reconcile.js +4 -3
  98. package/dist/esm/store/reconcile.js.map +1 -1
  99. package/dist/esm/store/turso/lexical.js +5 -1
  100. package/dist/esm/store/turso/lexical.js.map +1 -1
  101. package/dist/esm/text/segment.js +130 -4
  102. package/dist/esm/text/segment.js.map +1 -1
  103. package/dist/esm/watch-claim.d.ts +39 -0
  104. package/dist/esm/watch-claim.js +176 -0
  105. package/dist/esm/watch-claim.js.map +1 -0
  106. package/dist/esm/watch.js +168 -117
  107. package/dist/esm/watch.js.map +1 -1
  108. package/dist/esm/workers/watch-heartbeat.d.ts +1 -0
  109. package/dist/esm/workers/watch-heartbeat.js +106 -0
  110. package/dist/esm/workers/watch-heartbeat.js.map +1 -0
  111. package/package.json +2 -2
  112. package/skills/sense/SKILL.md +105 -104
  113. package/skills/sense/references/search.md +47 -0
  114. package/skills/sense/references/sql.md +120 -0
  115. package/skills/sense/references/stores/duckdb.md +35 -0
  116. package/skills/sense/references/stores/sqlite.md +66 -0
  117. package/skills/sense/references/stores/turso.md +36 -0
  118. package/skills/sense-setup/SKILL.md +93 -34
  119. package/skills/sense-setup/references/embeddings.md +51 -0
  120. package/skills/sense-setup/references/store-benchmarks.md +16 -0
  121. package/skills/sense-setup/references/stores/duckdb.md +28 -0
  122. package/skills/sense-setup/references/stores/sqlite.md +21 -0
  123. package/skills/sense-setup/references/stores/turso.md +25 -0
@@ -0,0 +1,106 @@
1
+ import { parentPort, workerData } from 'node:worker_threads';
2
+ import { SenseError } from '../errors.js';
3
+ import { serializeError } from '../scan/worker-error.js';
4
+ import { WatchClaimDatabase } from '../watch-claim.js';
5
+ function requireParentPort() {
6
+ const port = parentPort;
7
+ if (!port) throw new Error('watch heartbeat worker requires a parent port');
8
+ return port;
9
+ }
10
+ const port = requireParentPort();
11
+ const data = workerData;
12
+ let database;
13
+ let timer;
14
+ let stopping = false;
15
+ let renewing = Promise.resolve();
16
+ function send(message) {
17
+ port.postMessage(message);
18
+ }
19
+ function asError(value) {
20
+ return value instanceof Error ? value : new Error(String(value));
21
+ }
22
+ function closeDatabase() {
23
+ if (timer) clearInterval(timer);
24
+ timer = undefined;
25
+ try {
26
+ database === null || database === void 0 ? void 0 : database.close();
27
+ database = undefined;
28
+ return undefined;
29
+ } catch (err) {
30
+ return asError(err);
31
+ }
32
+ }
33
+ function addFailure(primary, secondary, message) {
34
+ return secondary ? new AggregateError([
35
+ primary,
36
+ secondary
37
+ ], message) : primary;
38
+ }
39
+ function fail(err) {
40
+ if (stopping) return;
41
+ stopping = true;
42
+ let failure = asError(err);
43
+ try {
44
+ database === null || database === void 0 ? void 0 : database.release(data.token);
45
+ } catch (releaseError) {
46
+ failure = addFailure(failure, asError(releaseError), 'watch heartbeat and claim release both failed');
47
+ }
48
+ const closeError = closeDatabase();
49
+ failure = addFailure(failure, closeError, 'watch heartbeat and database cleanup both failed');
50
+ try {
51
+ send({
52
+ type: 'failure',
53
+ error: serializeError(failure)
54
+ });
55
+ } finally{
56
+ port.close();
57
+ }
58
+ }
59
+ async function stop() {
60
+ if (stopping) return;
61
+ stopping = true;
62
+ if (timer) clearInterval(timer);
63
+ let failure;
64
+ try {
65
+ await renewing;
66
+ if (!(database === null || database === void 0 ? void 0 : database.release(data.token))) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');
67
+ } catch (err) {
68
+ failure = asError(err);
69
+ } finally{
70
+ const closeError = closeDatabase();
71
+ if (closeError) failure = failure ? addFailure(failure, closeError, 'watch heartbeat shutdown and database cleanup both failed') : closeError;
72
+ try {
73
+ if (failure) send({
74
+ type: 'failure',
75
+ error: serializeError(failure)
76
+ });
77
+ else send({
78
+ type: 'stopped'
79
+ });
80
+ } finally{
81
+ port.close();
82
+ }
83
+ }
84
+ }
85
+ async function main() {
86
+ try {
87
+ database = new WatchClaimDatabase(data.configDir);
88
+ await database.acquire(data.token, data.pid, data.force);
89
+ send({
90
+ type: 'acquired'
91
+ });
92
+ timer = setInterval(()=>{
93
+ if (stopping) return;
94
+ renewing = renewing.then(()=>{
95
+ if (!(database === null || database === void 0 ? void 0 : database.renew(data.token))) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');
96
+ });
97
+ renewing.catch(fail);
98
+ }, data.heartbeatIntervalMs);
99
+ port.on('message', (message)=>{
100
+ if (message.type === 'stop') void stop();
101
+ });
102
+ } catch (err) {
103
+ fail(err);
104
+ }
105
+ }
106
+ void main();
@@ -0,0 +1 @@
1
+ {"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/workers/watch-heartbeat.ts"],"sourcesContent":["import { parentPort, workerData } from 'node:worker_threads';\nimport { SenseError } from '../errors.ts';\nimport { serializeError } from '../scan/worker-error.ts';\nimport { WatchClaimDatabase, type WatchClaimWorkerData, type WatchClaimWorkerMessage } from '../watch-claim.ts';\n\nfunction requireParentPort() {\n const port = parentPort;\n if (!port) throw new Error('watch heartbeat worker requires a parent port');\n return port;\n}\n\nconst port = requireParentPort();\nconst data = workerData as WatchClaimWorkerData;\nlet database: WatchClaimDatabase | undefined;\nlet timer: NodeJS.Timeout | undefined;\nlet stopping = false;\nlet renewing = Promise.resolve();\n\nfunction send(message: WatchClaimWorkerMessage): void {\n port.postMessage(message);\n}\n\nfunction asError(value: unknown): Error {\n return value instanceof Error ? value : new Error(String(value));\n}\n\nfunction closeDatabase(): Error | undefined {\n if (timer) clearInterval(timer);\n timer = undefined;\n try {\n database?.close();\n database = undefined;\n return undefined;\n } catch (err) {\n return asError(err);\n }\n}\n\nfunction addFailure(primary: Error, secondary: Error | undefined, message: string): Error {\n return secondary ? new AggregateError([primary, secondary], message) : primary;\n}\n\nfunction fail(err: unknown): void {\n if (stopping) return;\n stopping = true;\n let failure = asError(err);\n try {\n database?.release(data.token);\n } catch (releaseError) {\n failure = addFailure(failure, asError(releaseError), 'watch heartbeat and claim release both failed');\n }\n const closeError = closeDatabase();\n failure = addFailure(failure, closeError, 'watch heartbeat and database cleanup both failed');\n try {\n send({ type: 'failure', error: serializeError(failure) });\n } finally {\n port.close();\n }\n}\n\nasync function stop(): Promise<void> {\n if (stopping) return;\n stopping = true;\n if (timer) clearInterval(timer);\n let failure: Error | undefined;\n try {\n await renewing;\n if (!database?.release(data.token)) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');\n } catch (err) {\n failure = asError(err);\n } finally {\n const closeError = closeDatabase();\n if (closeError) failure = failure ? addFailure(failure, closeError, 'watch heartbeat shutdown and database cleanup both failed') : closeError;\n try {\n if (failure) send({ type: 'failure', error: serializeError(failure) });\n else send({ type: 'stopped' });\n } finally {\n port.close();\n }\n }\n}\n\nasync function main(): Promise<void> {\n try {\n database = new WatchClaimDatabase(data.configDir);\n await database.acquire(data.token, data.pid, data.force);\n send({ type: 'acquired' });\n timer = setInterval(() => {\n if (stopping) return;\n renewing = renewing.then(() => {\n if (!database?.renew(data.token)) throw new SenseError('WATCH_ACTIVE', 'watch ownership was replaced by another watcher');\n });\n renewing.catch(fail);\n }, data.heartbeatIntervalMs);\n port.on('message', (message: { type?: string }) => {\n if (message.type === 'stop') void stop();\n });\n } catch (err) {\n fail(err);\n }\n}\n\nvoid main();\n"],"names":["parentPort","workerData","SenseError","serializeError","WatchClaimDatabase","requireParentPort","port","Error","data","database","timer","stopping","renewing","Promise","resolve","send","message","postMessage","asError","value","String","closeDatabase","clearInterval","undefined","close","err","addFailure","primary","secondary","AggregateError","fail","failure","release","token","releaseError","closeError","type","error","stop","main","configDir","acquire","pid","force","setInterval","then","renew","catch","heartbeatIntervalMs","on"],"mappings":"AAAA,SAASA,UAAU,EAAEC,UAAU,QAAQ,sBAAsB;AAC7D,SAASC,UAAU,QAAQ,eAAe;AAC1C,SAASC,cAAc,QAAQ,0BAA0B;AACzD,SAASC,kBAAkB,QAAiE,oBAAoB;AAEhH,SAASC;IACP,MAAMC,OAAON;IACb,IAAI,CAACM,MAAM,MAAM,IAAIC,MAAM;IAC3B,OAAOD;AACT;AAEA,MAAMA,OAAOD;AACb,MAAMG,OAAOP;AACb,IAAIQ;AACJ,IAAIC;AACJ,IAAIC,WAAW;AACf,IAAIC,WAAWC,QAAQC,OAAO;AAE9B,SAASC,KAAKC,OAAgC;IAC5CV,KAAKW,WAAW,CAACD;AACnB;AAEA,SAASE,QAAQC,KAAc;IAC7B,OAAOA,iBAAiBZ,QAAQY,QAAQ,IAAIZ,MAAMa,OAAOD;AAC3D;AAEA,SAASE;IACP,IAAIX,OAAOY,cAAcZ;IACzBA,QAAQa;IACR,IAAI;QACFd,qBAAAA,+BAAAA,SAAUe,KAAK;QACff,WAAWc;QACX,OAAOA;IACT,EAAE,OAAOE,KAAK;QACZ,OAAOP,QAAQO;IACjB;AACF;AAEA,SAASC,WAAWC,OAAc,EAAEC,SAA4B,EAAEZ,OAAe;IAC/E,OAAOY,YAAY,IAAIC,eAAe;QAACF;QAASC;KAAU,EAAEZ,WAAWW;AACzE;AAEA,SAASG,KAAKL,GAAY;IACxB,IAAId,UAAU;IACdA,WAAW;IACX,IAAIoB,UAAUb,QAAQO;IACtB,IAAI;QACFhB,qBAAAA,+BAAAA,SAAUuB,OAAO,CAACxB,KAAKyB,KAAK;IAC9B,EAAE,OAAOC,cAAc;QACrBH,UAAUL,WAAWK,SAASb,QAAQgB,eAAe;IACvD;IACA,MAAMC,aAAad;IACnBU,UAAUL,WAAWK,SAASI,YAAY;IAC1C,IAAI;QACFpB,KAAK;YAAEqB,MAAM;YAAWC,OAAOlC,eAAe4B;QAAS;IACzD,SAAU;QACRzB,KAAKkB,KAAK;IACZ;AACF;AAEA,eAAec;IACb,IAAI3B,UAAU;IACdA,WAAW;IACX,IAAID,OAAOY,cAAcZ;IACzB,IAAIqB;IACJ,IAAI;QACF,MAAMnB;QACN,IAAI,EAACH,qBAAAA,+BAAAA,SAAUuB,OAAO,CAACxB,KAAKyB,KAAK,IAAG,MAAM,IAAI/B,WAAW,gBAAgB;IAC3E,EAAE,OAAOuB,KAAK;QACZM,UAAUb,QAAQO;IACpB,SAAU;QACR,MAAMU,aAAad;QACnB,IAAIc,YAAYJ,UAAUA,UAAUL,WAAWK,SAASI,YAAY,+DAA+DA;QACnI,IAAI;YACF,IAAIJ,SAAShB,KAAK;gBAAEqB,MAAM;gBAAWC,OAAOlC,eAAe4B;YAAS;iBAC/DhB,KAAK;gBAAEqB,MAAM;YAAU;QAC9B,SAAU;YACR9B,KAAKkB,KAAK;QACZ;IACF;AACF;AAEA,eAAee;IACb,IAAI;QACF9B,WAAW,IAAIL,mBAAmBI,KAAKgC,SAAS;QAChD,MAAM/B,SAASgC,OAAO,CAACjC,KAAKyB,KAAK,EAAEzB,KAAKkC,GAAG,EAAElC,KAAKmC,KAAK;QACvD5B,KAAK;YAAEqB,MAAM;QAAW;QACxB1B,QAAQkC,YAAY;YAClB,IAAIjC,UAAU;YACdC,WAAWA,SAASiC,IAAI,CAAC;gBACvB,IAAI,EAACpC,qBAAAA,+BAAAA,SAAUqC,KAAK,CAACtC,KAAKyB,KAAK,IAAG,MAAM,IAAI/B,WAAW,gBAAgB;YACzE;YACAU,SAASmC,KAAK,CAACjB;QACjB,GAAGtB,KAAKwC,mBAAmB;QAC3B1C,KAAK2C,EAAE,CAAC,WAAW,CAACjC;YAClB,IAAIA,QAAQoB,IAAI,KAAK,QAAQ,KAAKE;QACpC;IACF,EAAE,OAAOb,KAAK;QACZK,KAAKL;IACP;AACF;AAEA,KAAKc"}
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "sensemaking",
3
- "version": "0.24.4",
3
+ "version": "0.24.5",
4
4
  "description": "Query and search your markdown notes with context-aware progressive disclosure: SQL over frontmatter, links, and text, plus semantic search and link-graph ranking. No server, no build step.",
5
5
  "keywords": [
6
6
  "markdown",
@@ -93,7 +93,7 @@
93
93
  "yaml": "^2.9.0"
94
94
  },
95
95
  "devDependencies": {
96
- "@duckdb/node-api": "*",
96
+ "@duckdb/node-api": "1.5.5-r.4",
97
97
  "@tursodatabase/database": "*",
98
98
  "@types/markdown-it-footnote": "^3.0.4",
99
99
  "@types/mocha": "*",
@@ -1,136 +1,137 @@
1
1
  ---
2
2
  name: sense
3
- description: "Query a markdown tree with the sense CLI: filter notes by frontmatter, full-text search the prose, follow wikilinks/backlinks, trace how notes connect (link path, similar-but-unlinked), and read note outlines. Use when the user wants to query, filter, count, search, or report on a folder of markdown notes, when you need to find which notes discuss a topic before reading them, when you want a note's backlinks or structure, when you want to know how two notes connect, when a directory has a sense.config.json, or when asked to add a saved entry to one."
3
+ description: "Query a markdown tree with the sense CLI: filter notes by frontmatter, search prose, follow wikilinks and backlinks, trace link paths, find similar but unlinked notes, and inspect note outlines. Use when a task needs to query, filter, count, search, or report on a markdown directory; when a directory has sense.config.json; or when adding a saved sense query."
4
4
  ---
5
5
 
6
6
  # sense
7
7
 
8
- SQL over a markdown tree, kept fresh by a filesystem check on every query. Every file becomes rows in `frontmatter` (one column per key, plus `path`/`_mtime`/`_ctime`/`_size`/`_rank`/`_parse_error`; `_ctime` is filesystem birthtime, which a clone or copy resets just like `_mtime`), `content` (`title`, `summary`, `text`, `path`; an FTS5 index on the default store, with machine-written `title_seg`/`summary_seg`/`text_seg` sidecars used for matching Chinese, Japanese, Thai, Khmer, Lao, and Burmese text, not for reading), `links` (`src`, `target`, `dst`, `embed`; `NULL` dst = dead link; `embed` 1 for `![[...]]` embeds, 0 for links; one row per distinct written target and kind with alias and anchor stripped, so `[[Foo]]` and `[[Foo|alias]]` are one row while `[[Foo]]` and `[[notes/Foo]]` are two rows that can share a `dst`, and a target both linked and embedded is a row of each kind; extraction matches Obsidian's own graph: comments, code, and link-syntax text yield no rows, `[[#Anchor]]` is a self-edge, a frontmatter value that is exactly `[[X]]` is a link, and a basename collision resolves to the linking note itself, else the shortest path), `tags` (`path`, `tag`: frontmatter and inline `#tags` merged and deduplicated; nested tags stored full, so `book/scifi` matches `tag = 'book' OR tag LIKE 'book/%'`), `sections` (heading outline with line ranges and token estimates), and `preset_files` (`path`, `preset`: which presets cover which files). Features add their own storage; `map` and `status` report which are on.
8
+ Use sense to locate evidence in a markdown tree before reading files. Every command reconciles changed files first. Results contain paths, metadata, excerpts, and line ranges. Read the returned files or ranges when the task needs their prose.
9
9
 
10
- **Stores.** The config's `store` key picks the backing store: `sqlite` (default, zero-dependency, Node's built-in SQLite), or the experimental `duckdb` and `turso` (the first command that opens such a tree installs that engine's package on its own: `@duckdb/node-api`, a one-time native download of about 110 MB, or the much smaller `@tursodatabase/database`). The tables, `?` placeholders, quoted identifiers, and the `scope` binding are the same on all three, so ordinary frontmatter SQL ports as written. Two things do not port. First, FTS5: `content` is an FTS5 table on sqlite and a plain table on the other two, so hand-written `MATCH`, `snippet()`, `bm25()`, and sqlite's date-function forms run only on sqlite (duckdb has its own fts functions and date syntax; turso has Tantivy's `fts_match`/`fts_score`), and `search` text under duckdb and turso rejects FTS5's prefix (`foo*`), boolean (`AND`/`OR`/`NOT`), `NEAR`, initial-token (`^`), and column-filter (`title:foo`) operators with a named error (STORE_CAPABILITY_MISSING) that says how to rephrase or set `store` to `sqlite`; bare words and quoted phrases work on all three. Second, of the `has`/`basename`/`segment` functions, `has` and `basename` run on all three (turso rewrites them into portable SQL rather than registering them), while `segment` runs on sqlite and duckdb only, so a query calling `segment` under turso fails with a named error (STORE_CAPABILITY_MISSING) saying to rephrase or set `store` to sqlite. `sense watch` runs on all three; under duckdb and turso, which lock the cache file per connection, a concurrent command waits out the watcher's current cycle instead of failing.
10
+ Setup, store selection, presets, and note design belong to the `sense-setup` skill. Translating an Obsidian Bases file belongs to `sense-bases`.
11
11
 
12
- ## What each tool is for
12
+ ## Start with the tree
13
13
 
14
- Every result is a reference (path, metadata, excerpt), never file contents; prose enters context only when you Read it. Costs: `map` is fixed-size, a `search` row is tens of tokens, and a `peek` stays flat however large the note is. Which tool fits is a property of the question:
14
+ Run these when the tree or its configuration is unfamiliar:
15
15
 
16
- - A deterministic, factual answer over known fields (counts, filters, "which notes have X") is SQL: `sense sql`, a saved `{ sql }` entry, or `search --where`. Enumerates every match; same result regardless of phrasing.
17
- - Locating notes about something is `search`, one text through every engine the scope has: word match (bare words AND-join, one absent word = zero lexical rows; write `a OR b OR c` for any-word), link-graph expansion, and vector similarity, fused into one ranked list. Read `via` per row: `match` rows contained your words; `vector`-only rows did not. A `vector`-only row means the search words don't appear in that note; it showed up because the model judged it semantically related. Vector rows are conceptual similarity, not typo-tolerance; false positives are expected, labeled, and bounded by `--k`, and they are the only rows a search can produce when note and query share no vocabulary at all, the paraphrase and category-for-instance cases words cannot reach. A scope searches with vectors when its preset's `signals` include `vectors` (on by default whenever the tree names an `embed` model); a preset that declares `"signals": {"words": 1, "links": 1}` searches on words and links only. A preset that asks for vectors when a local model path is missing its files is an error naming the fix, not a quieter result that would make the same search answer differently before and after.
18
- - `map` answers "what is this tree" (fields, hub notes, recent changes) when the tree is unfamiliar.
19
- - `peek <path>` prices a file before you pay for it: outline with `[L143-162, ~380t]` ranges and links both ways. Every list shows its first 20 with the true total; the `sections` and `links` tables hold the rest, so a peek costs a few hundred tokens on any note. Its link totals count distinct notes: "links out" dedupes written targets by resolved note and lists unresolved targets separately, while `COUNT(*) FROM links WHERE src = ?` counts every written target, resolved or not, so the raw count can read higher without either number being wrong.
20
- - `path <a> <b>` walks the link graph for a chain connecting two notes, or reports none within the depth bound: it answers how they connect, not just that both exist.
21
- - `related <note>` ranks notes near in meaning to one note that it does not already link to: the links it is missing. It reads the meaning-vectors, so it needs the `vectors` signal on for the scope and scans them, costing about what a vector `search` does, not what a `peek` does. The model named in the config fetches once per machine at the first vector search (progress on stderr; `sense download` prefetches where that timing matters). A `model` pointing at a local directory with missing files is an error naming it, for `search` and `related` alike, not a quieter result.
22
- - When you know the file and need its contents, `Read` it. sense adds nothing there. On large files peek's ranges let you read just one section; small files are often cheaper whole.
16
+ ```sh
17
+ sense status # config, store, cache, document count, preset coverage
18
+ sense map # fields, hubs, recent notes
19
+ sense --list # saved queries
20
+ ```
21
+
22
+ If a command reports a missing config, run `sense init` only when the user wants the tree configured. A one-off query does not authorize changing the tree.
23
+
24
+ ## Choose the command
23
25
 
24
- Output defaults to a table, built for humans; `--format json` returns the same rows machine-parseable, and `--format csv` returns one row per line for `grep` and `awk`. csv keeps every character a value holds, embedded newlines included, but cannot express NULL versus an empty string or a value's type, which json can. The commands that emit rows (`sql`, `search`, `related`, saved queries) take all three; `map`, `peek`, `status`, and `path` render a structure rather than a row set and take table or json. That also makes a saved query usable as a CI/hook gate with zero added mechanism: `[ "$(sense <name> --format json)" = "[]" ]` is true exactly when it returned no rows.
26
+ | Need | Command |
27
+ |---|---|
28
+ | Exact count, filter, grouping, or known-field report | `sense sql` or a saved SQL query |
29
+ | Notes about a subject | `sense search` |
30
+ | One note's frontmatter, outline, links, and backlinks | `sense peek` |
31
+ | Shortest link chain between two notes | `sense path` |
32
+ | Similar notes that a note does not link to | `sense related` |
33
+ | Tree shape and available fields | `sense map` |
25
34
 
26
- ## Three signals
35
+ Start with a bounded result. Raise `--k`, widen the preset, or broaden the words only when the first result does not answer the question.
27
36
 
28
- `search` composes up to three engines; which ones run is the preset's `signals` (every signal whose prerequisites hold, by default). Each exists because it reads evidence the others cannot:
37
+ ```sh
38
+ sense search "pricing" --k 10
39
+ sense search "sourcing quotes" --preset raw
40
+ sense peek notes/pricing-model.md
41
+ sense path onboarding.md pricing-model.md
42
+ sense related notes/pricing-model.md --k 10
43
+ sense sql "SELECT path FROM frontmatter WHERE status = ? LIMIT 50" active
44
+ ```
29
45
 
30
- | signal | reads | uniquely finds | blind to |
31
- |---|---|---|---|
32
- | `words` | the text itself (BM25) | every literal occurrence: identifiers, error strings, names, exact phrases; deterministic | anything phrased differently, e.g. a note saying "compensation floor" for the query "minimum pay" |
33
- | `links` | connections authors wrote | what the tree's own structure treats as related: context the matching text never restates | notes nobody linked |
34
- | `vectors` | per-chunk meaning-vectors | notes sharing no words with the query: paraphrase, the category for an instance, a concept restated | proving presence. A vector row cannot show the words occur anywhere; `snippets` can |
46
+ ## Read the store guide before composing syntax
35
47
 
36
- Each signal in the map carries a weight (`{"words": 1, "links": 1, "vectors": 1}` is the default when the key is omitted entirely); presence turns a signal on, and the number scales its share of the fused ranking. Equal weight (1 for everything) is what every number elsewhere in this doc describes. A weight above 1 pulls the fused ranking toward that signal's own ordering: measured on nfcorpus with the default static model and on MIRACL zh with an HTTP encoder, equal-weight fusion helps English nfcorpus but ranks the encoder's MIRACL-zh results below its own cosine-only ranking (benchmark/reports/2026-08-27-embedding-model-selection.md, weight-sweep table). There is no single weight that is right for both, so a preset that leans on a strong encoder is a candidate for a higher `vectors` weight, checked against that table rather than assumed.
48
+ The config's `store` key selects the SQL dialect and text-search grammar. An omitted key means `sqlite`. Bare words and quoted phrases work in `sense search` on every store. Advanced operators and raw text-index SQL differ.
37
49
 
38
- Which signal carries a search is a property of the query, and the ends of the range are measured (BENCHMARKING.md, "Retrieval quality"): on the vocabulary-gap corpus, 31% of queries have no relevant lexical row in their top 10 and vector rows are the only recall; on a corpus whose queries quote their documents, words alone hit 99.7% and vectors add nothing. Real trees sit between. An exact identifier is words territory; "notes about X" where X could be phrased many ways is where vector rows carry; "how do these connect" is the link graph (`path`, `peek`).
50
+ Read the matching guide before writing raw SQL that uses engine functions, dates, JSON, text matching, or native types. Also read it before using advanced search operators.
39
51
 
40
- Attribution is already per-row: one fused search shows which engines earned each hit (`via`), so reading the labels across a few queries is the cheapest way to learn which signals this tree rewards. To isolate a signal harder, declare presets differing only in `signals` and run the same text through both; the row diff is the excluded signal's contribution. The labels also matter downstream: relayed to a human, they distinguish "contains these words" from "related by meaning", which is the difference between a citation and a lead.
52
+ - [SQLite query guide](references/stores/sqlite.md)
53
+ - [DuckDB query guide](references/stores/duckdb.md)
54
+ - [Turso query guide](references/stores/turso.md)
41
55
 
42
- ## Commands
56
+ The tables, `?` placeholders, quoted identifiers, `has()`, `basename()`, and preset `scope` binding are shared. [Portable SQL](references/sql.md) documents the schema and queries that work without depending on one engine.
43
57
 
44
- ```
45
- sense search "pricing OR billing OR invoicing" --where "f.status = 'active'" --k 10
46
- sense search "sourcing quotes" --preset raw # a named settings bundle from the config
47
- sense peek notes/pricing-model.md # a unique basename also works
48
- sense path onboarding.md pricing-model.md # link chain between two notes, or none within the bound
49
- sense related notes/pricing-model.md # notes similar by meaning it does not yet link to
50
- sense map
51
- sense sql "<statement>" [params...] # ad-hoc SQL; ? binds positional args, count-checked
52
- sense <name> [params...] # a saved query from sense.config.json
53
- sense --list | status | download
58
+ ## Search and evidence
59
+
60
+ `search` combines the signals enabled by the selected preset:
61
+
62
+ | Signal | Evidence |
63
+ |---|---|
64
+ | `match` | The note contains the search words. `snippets` shows the matching passages. |
65
+ | `link` | A note that matched links to this note. |
66
+ | `vector` | The embedding model placed this note near the query. The search words may be absent. |
67
+
68
+ Combinations such as `match+link` mean that more than one signal produced the row. `score` ranks rows within that result only. Do not compare it across searches. `similarity` ranks vector evidence within the current result and model. Do not carry a fixed similarity cutoff between trees.
69
+
70
+ The `lines` value points at the section that earned the row. Read that range when it is present. A null range means the whole note is the reference. A vector-only row has no lexical snippet and is a lead, not proof that the note contains the query terms.
71
+
72
+ Read [search evidence and troubleshooting](references/search.md) when a search will support a factual claim, when absence matters, when results look noisy, or when tuning signal weights.
73
+
74
+ ## Scope and output
75
+
76
+ Bare commands use the `default` preset. `--preset <name>` chooses another. For `search`, `--include`, `--exclude`, and `--no-exclude` change the query scope for one invocation, but cannot reach files that no preset indexes. `sense status` shows actual coverage.
77
+
78
+ `--where` filters search and graph commands against frontmatter alias `f`:
79
+
80
+ ```sh
81
+ sense search "pricing" --where "f.status = 'active' AND has(f.tags, 'sales')"
54
82
  ```
55
83
 
56
- - Terms pass verbatim to FTS5 MATCH. Bare words AND-join (one absent word means zero rows), so write `OR` yourself when you want any-word matching; double-quote punctuated terms (`"customer-facing"`, `"founder's"`); invalid syntax is an error, not a rewrite. The same rules apply to search commands you write into subagent briefs.
57
- - When a search misses, both sides have levers. Lexical: OR-in synonyms and concrete instances (the index only knows the words in the files; a note about a specific tool rarely names its category), raise `--k` (a row costs tens of tokens), widen the scope (`--preset`, or `--include` for an ad-hoc glob). Vector: restate the concept in different words. Vectors rank the meaning of the whole query text, so a rephrase moves them even when every literal word still misses. Then pivot through the nearest hit with `related <path>` to walk its meaning-neighbourhood. Vector-only rows in a miss are leads rather than lexical evidence: use `similarity` and the `lines` range to decide what to read; their `snippets` list is empty because the query words did not match. Each widening adds candidates and dilutes ranking, so the noise trade-off runs both ways.
58
- - A frontmatter query enumerates its matches deterministically; search ranks by term overlap, so results shift as phrasing shifts. Trade-off: a query needs a known field, search doesn't.
59
- - The `via` column says what produced each row: `match` (words hit), `link` (connected to notes that hit), `vector` (near in meaning), and combinations. The `lines` column, when set, points at the section that earned the row (the best-matching chunk on vector rows, the term cluster's section on large lexical notes) and is a direct `Read` range; null means the whole note is the reference. A row's `snippets` holds the marked passages that earned the lexical match (`[]` on a row that never contained the terms); `--snippet-char-limit` (default 80, the passage's rendered length, marks and ellipses counted) and `--snippet-count-limit` (default 1, passages per note) bound them, on `search` and a saved search alike.
60
- - Scope is one vocabulary shared by `search`, `map`, `peek`, `path`, and `related`: bare command uses the config's `default` preset; `--preset <name>` picks another (unknown names error, listing what's declared); `--include <glob>` and `--exclude <glob>` are ad-hoc globs for one command, each overriding its own side of the preset, so one does not clear the other; `--no-exclude` drops the preset's `exclude` for one command, the only way to widen past it without editing config (it widens the query scope, not the index: a file no preset covers is never indexed). `--where` takes any SQL condition against frontmatter alias `f`, not only field equality: `"f.status = 'active' AND has(f.tags, 'x')"`, `"datetime(f.created) >= datetime(?)"`. There is no whole-index flag: a broad `default` preset, or a declared `all` preset (`include ["**/*"]`), is the whole tree. `sense status` shows every preset with its coverage. `sql` scopes differently: it runs over the whole index by default, and `--preset <name>` *binds* the scope as a temporary `scope(path)` table your statement joins, rather than filtering behind the query's back (`JOIN scope ON scope."path" = f.path`). Naming a preset without joining `scope` is a usage error, since it would return everything while reading as scoped. Without the flag, join `preset_files` directly, which is the same coverage under a preset name you write into the SQL.
61
- - `score` is a rank-fusion value: it ranks rows within one result set and is not comparable across queries, not a relevance magnitude. It encodes how many signals fired and at what rank, so a perfect lexical hit and a weak vector-only hit can read the same number. With vectors active, rows carry `similarity`: the cosine (-1 to 1) of the query against that file's best-matching chunk (the same chunk the `lines` range points at). It orders vector evidence within a result set; the range it spans depends on the corpus and the embedding model, and compresses on small trees, where even a nonsense query has a moderately near neighbour somewhere. Compare similarities within a result set rather than against a fixed cutoff carried between trees.
62
- - Vector rankings read the whole chunk, boilerplate included, so on a tree whose notes are mostly one template with a line of unique text (a directory of plugins, people, or assets), the shared scaffolding dominates every vector and `related` returns near-ties at the top of the range for any seed. Two signs, both visible in the output already: the top `similarity` values sit within a hair of each other, and the same few notes come back for unrelated seeds. Compare neighbour lists across two unlike seeds when a tree looks like this; matching lists mean the ranking is reading the template, not the content, and `search` over the distinguishing words is the answer instead.
63
- - Absence evidence lives in the labels: a preset whose `signals` exclude `vectors` (or a tree with no `embed` block) returns 0 rows when the words are nowhere in it. Default `search` always returns up to `k` rows (nearest-neighbour search has a nearest neighbour for any input), so a result of only `via: vector` rows is the absence signal for the words themselves. Judge whether a vector row is a useful conceptual lead from `similarity`, then read its `lines` range; it has no lexical snippet.
64
- - A `queries` entry names the verb it runs, mirroring the two commands: `"dead-links": { "sql": "SELECT src, target FROM links WHERE dst IS NULL AND lower(target) NOT GLOB '*.[a-z0-9]*'" }` runs as `sense dead-links`, and `"hot": { "search": "pricing OR billing", "preset": "raw", "k": 20 }` runs as `sense hot` with its settings baked in, so repeat runs need no flags. An invocation-level `--preset`, `--k`, or `--where` overrides a saved search's value; `--list` labels each entry `(sql)` or `(search)`.
65
- - Running an entry is how it is validated: a typo'd column, stale SQL, bad FTS5 syntax or an unknown preset errors and exits nonzero. A parameterised entry validates with any argument, since SQL is prepared before parameters bind (`sense by-tag zzz` reports `no such column` if the column is wrong, `(0 rows)` if it is right). To sweep a whole config after editing it, read the exit code: 0 ran, 2 means it needs parameters (re-run it with any argument to validate the SQL), anything else is broken.
84
+ `sense sql` is index-wide by default. With `--preset`, the command binds a temporary `scope(path)` table. The SQL must join it:
66
85
 
67
86
  ```sh
68
- for q in $(sense --list | awk '{print $1}'); do
69
- sense "$q" --format json >/dev/null 2>&1
70
- case $? in 0) ;; 2) echo "needs an argument: $q" ;; *) echo "broken: $q" ;; esac
71
- done
87
+ sense sql "SELECT f.path FROM frontmatter f JOIN scope ON scope.path = f.path" --preset default
72
88
  ```
73
89
 
74
- Whether an empty result is good or bad is the reader's judgment: a dead-link query returning rows means broken citations to fix.
90
+ Table output is for people. Use `--format json` when code or an agent will parse rows. Use `--format csv` when redirecting a large row set to a file. `sql`, `search`, `related`, and saved queries support all three formats. `map`, `peek`, `status`, and `path` support table and JSON.
75
91
 
76
- ## SQL
92
+ ## Saved queries
77
93
 
78
- The commands are shorthands over those tables; anything they don't express, SQL does.
94
+ Save a query in `sense.config.json` only when it will be reused:
79
95
 
80
- ```
81
- sense sql "SELECT name FROM pragma_table_info('frontmatter')" # what fields exist
82
- sense sql 'SELECT path FROM frontmatter WHERE "plugin-id" IS NOT NULL' # punctuated keys need double quotes: unquoted, plugin-id reads as subtraction
83
- sense sql "SELECT DISTINCT status FROM frontmatter" # what values a field takes
84
- sense sql "SELECT src FROM links WHERE dst = ?" notes/pricing-model.md # backlinks
85
- sense sql "SELECT f.path FROM frontmatter f JOIN scope ON scope.path = f.path" --preset default # scope SQL to a preset
86
- sense sql "SELECT f.path FROM frontmatter f JOIN preset_files p ON p.path = f.path AND p.preset = 'default'" # the same, preset named in the SQL
87
- sense sql "SELECT path FROM frontmatter WHERE path NOT IN (SELECT dst FROM links WHERE dst IS NOT NULL) AND path NOT IN (SELECT src FROM links)" # linked neither way (fine if intentional; linking is optional)
88
- sense sql "SELECT src, target FROM links WHERE dst IS NULL AND lower(target) NOT GLOB '*.[a-z0-9]*'" # broken wikilinks, attachments excluded
89
- sense sql "SELECT heading, start_line, tokens FROM sections WHERE path = ?" a.md # budget a read
90
- sense sql "SELECT j.value, COUNT(*) n FROM frontmatter, json_each(frontmatter.tags) j GROUP BY j.value ORDER BY n DESC" # count per array member
96
+ ```json
97
+ {
98
+ "queries": {
99
+ "by-tag": { "sql": "SELECT path, title FROM frontmatter WHERE has(tags, ?) ORDER BY path" },
100
+ "hot": { "search": "pricing", "preset": "raw", "k": 20 }
101
+ }
102
+ }
91
103
  ```
92
104
 
93
- `path` covers the route between two notes, and `search`'s `via: link` rows are the ranked neighborhood around a query; a structural k-hop walk is a bounded `WITH RECURSIVE` over `links`. Bound the depth: an unbounded walk on a densely linked tree enumerates paths exponentially. Pass these through `sense sql "<sql>" <seed>` or save as `{ "sql": "..." }`.
105
+ Run these as `sense by-tag urgent` and `sense hot`. Invocation flags override a saved search's `preset`, `k`, or `where` value.
94
106
 
95
- ```
96
- -- notes within 2 hops of a seed, links both ways (UNION dedups, so it terminates)
97
- WITH RECURSIVE hop(path, d) AS (
98
- SELECT ?, 0
99
- UNION
100
- SELECT CASE WHEN l.src = hop.path THEN l.dst ELSE l.src END, hop.d + 1
101
- FROM hop JOIN links l ON (l.src = hop.path OR l.dst = hop.path) AND l.dst IS NOT NULL
102
- WHERE hop.d < 2
103
- )
104
- SELECT DISTINCT path FROM hop WHERE d > 0;
105
-
106
- -- notes cited alongside a seed: they share a note that links to both (co-citation)
107
- SELECT DISTINCT b.dst FROM links a JOIN links b ON a.src = b.src
108
- WHERE a.dst = ? AND b.dst IS NOT NULL AND b.dst <> a.dst;
109
- ```
107
+ Running a saved entry validates it. Exit code 0 means it ran, 2 means the invocation needs different arguments, and 1 means the query or store failed. An empty result can be valid data, so interpret it from the query's purpose.
108
+
109
+ ## Tables
110
+
111
+ | Table | Holds |
112
+ |---|---|
113
+ | `frontmatter` | One discovered column per frontmatter key, plus `path`, `_mtime`, `_ctime`, `_size`, `_rank`, and `_parse_error` |
114
+ | `content` | `path`, `title`, `summary`, and authored text, plus store-owned search columns |
115
+ | `links` | `src`, written `target`, resolved `dst`, and `embed` |
116
+ | `tags` | Merged and deduplicated frontmatter and inline tags |
117
+ | `sections` | Heading, level, line range, and token estimate |
118
+ | `preset_files` | Paths covered by each preset |
119
+
120
+ Features can add tables. `sense map` and `sense status` show which features are active.
121
+
122
+ A non-null `_parse_error` means the file has no recovered frontmatter values. To distinguish a missing field from invalid frontmatter, include `_parse_error IS NULL` in the filter. Fixes appear on the next command because reconciliation runs first.
123
+
124
+ ## Reading discipline
125
+
126
+ Select only the columns needed for the answer. Use `LIMIT` for row-returning exploration. Prefer `path`, `title`, `summary`, and bounded snippets over `content.text`. Aggregates such as `COUNT` and `GROUP BY` are already bounded by their result shape.
127
+
128
+ When a result identifies a large note, use `peek` and then read the relevant line range. Small files are often cheaper to read whole.
129
+
130
+ Worked command traces are in [EXAMPLES.md](EXAMPLES.md).
131
+
132
+ ## Upkeep
110
133
 
111
- - A saved `{ sql }` written against `scope` is preset-agnostic: `sense <name> --preset raw` re-points the same statement at another layer, so one entry serves every preset instead of one copy each.
112
- - `content MATCH` only works against the fts5 table by its own name, never through an alias or a view: `FROM content c ... WHERE c MATCH 'x'` fails with `no such column: c`. This is why `--preset` binds a table to join rather than shadowing the tables.
113
- - `content MATCH` takes FTS5 syntax: `a OR b`, `"phrase"`, `pref*`, `NEAR(a b, 5)`, `summary: term`. Stemmed; markdown stripped at index time. Double-quote any term with punctuation. Bare `customer-facing` errors (`-` reads as a column filter), bare apostrophes are syntax errors: write `"customer-facing"`, `"founder's"`. This is the default sqlite store's grammar: under `duckdb` and `turso` the operator forms (`a OR b`, `pref*`, `NEAR`, `^`, column filters) are a named error naming the rephrase, and `MATCH` itself does not run (see Stores).
114
- - A language written without word spaces (Chinese, Japanese, Thai, Khmer, Lao, Burmese) is indexed per grapheme and searched as an ordered grapheme phrase against the `_seg` sidecar columns: substring semantics, what `grep` gives, a query matches wherever its exact text occurs, including inside a longer run (`京都` matches `东京都政府`, correctly, because it's there at position 2), and needs no minimum length. No decision is needed for these languages. Hand-written SQL is not rewritten for you, so a raw `content MATCH '数据库'` finds nothing: write `content MATCH segment(?)` and bind the terms. `segment()` returns text with no such run unchanged, so it is safe to leave in a query whatever the tree's language.
115
- - Rank with `ORDER BY bm25(content, 10.0, 5.0, 1.0)` (title > summary > body); the full form `bm25(content, 10.0, 5.0, 1.0, 0, 10.0, 5.0, 1.0)` mirrors the same weights onto the `_seg` sidecars, so a title hit found through `title_seg` ranks like one found through `title` (the three-weight form still runs, FTS5 defaults unnamed columns to 1.0, but ranks a sidecar match at body weight). In hand-written SQLite SQL, `snippet(content, 2, '«', '»', '…', 10)` names the authored `text` column explicitly; `-1` can surface a machine-spaced sidecar instead. `search` does not call SQLite's `snippet()`: it computes bounded passages in shared code for every store. A raw SQL `snippet()` re-tokenizes its matched document, so guard large text with `CASE WHEN length(text) <= 16384 THEN snippet(...) END`, or select `title`/`summary`. Treat historical timing as diagnostic until a current identical-work sitting replaces it.
116
- - Select `content.title`/`content.summary` (always exist, empty when absent) rather than `f.title`/`f.summary` (discovered columns; error on trees that never declare them).
117
- - Frontmatter values keep their YAML type: strings are TEXT, whole numbers and booleans are INTEGER (`true` stores as 1, so `WHERE flag = 1` matches and `WHERE flag = 'true'` matches nothing), fractions are REAL, lists and maps are JSON text. On the `duckdb` store the discovered columns are VARIANT: homogeneous keys, which are nearly all of them, compare identically, but a numeric predicate against a key that holds numbers in some notes and text in others raises a comparison error where sqlite orders by storage class silently. `map` prints the observed type per field, and a field showing two types (`integer,text`) has drifted across notes. A list key written with no items (`tags:` above a bare `-`) is a list holding one null, stored as the JSON text `[null]`: `IS NULL` does not find it (the column holds a string), `json_each` yields one empty member per such row, and `map` counts it as covered because the key is present. `has(tags, 'x')` reads it correctly as no match. To separate written-but-empty from absent, compare against the text: `WHERE tags = '[null]'`.
118
- - **Dead links need the attachment filter.** `dst IS NULL` alone is not "broken link": a wikilink to anything that is not markdown (`[[Board.base]]`, `![[Pasted image.png]]`, `[[spec.pdf]]`) can never resolve, because sense indexes markdown and resolution only tries the exact path or `+.md`. Those are out of the index's universe, not broken. On a 1,400-note Obsidian vault the unfiltered query returns 143 rows where 14 are real. Exclude anything carrying a file extension, as in the recipe above, and widen the exclusion if your notes have dotted titles (`[[Node.js]]` carries one too, so a stricter list, `'*.png'`, `'*.pdf'`, `'*.base'`, and whatever else your vault attaches, is safer on a tree whose titles use dots). Scope it with `preset_files` as well: template and skill files are full of `[[Note Name]]` examples that are deliberately unresolved.
119
- - `has(field, value)`: array membership on JSON-array fields, substring on strings, false on NULL. This is the `includes()` convention. Substring means `has(f.status, 'active')` also matches `inactive`; exact scalar match is `f.status = ?`, deliberate substring is `LIKE`, exact array membership is `EXISTS (SELECT 1 FROM json_each(f.tags) WHERE value = ?)`. To aggregate per member instead, use `json_each(frontmatter.<field>)` (above). GROUP BY on the raw column splits `["a","b"]` and `["b","a"]` into separate buckets.
120
- - Compare dates through `datetime()`, which resolves ISO 8601 offsets to UTC: `WHERE datetime(created) >= datetime(?)`. Bare string comparison is only safe when every note uses the same offset.
121
- - Date spellings SQLite rejects (`-0800`, `-08`, a space separator) are normalized at index time, offset preserved. One it cannot fix is left as written and warned about by path: `datetime()` returns NULL there, so the row is invisible to a date comparison rather than excluded by it. List them with `WHERE d IS NOT NULL AND datetime(d) IS NULL`.
122
- - **SQLite's `now` is UTC, so any query about "today" needs `'localtime'`.** `date('now')` reads as tomorrow from mid-afternoon onward in the Americas, which silently flips "scheduled today" into "overdue" every evening: write `date('now','localtime')` and `datetime('now','start of day','localtime')`. This only matters where the boundary carries the meaning; a `'-90 day'` window is unaffected by a few hours of skew.
123
- - To bound what a query puts into context: `snippet()` excerpts just the matching text, `LIMIT` caps row counts, and selecting `path`/`title`/`summary` keeps rows small. A large result can also stay out of context entirely: `--format csv > file` writes it whole, and `grep`/`awk` over that file returns only the rows wanted. `SELECT text FROM content` returns the tree's entire prose (sense warns past 50 KB, after the rows have already printed, so the warning records the cost rather than preventing it). Aggregates (`COUNT`, `GROUP BY`) are already bounded. `SELECT * FROM frontmatter` is always safe: prose is not a frontmatter column.
124
-
125
- Worked traces: [EXAMPLES.md](EXAMPLES.md).
126
-
127
- ## Setup and upkeep
128
-
129
- - Missing CLI: `npm install -g sensemaking`. Missing config: `sense init` at the tree root. Discovery walks up from cwd; `--config <path>` overrides. Setting up or restructuring a tree (presets, frontmatter conventions, note design) is the `sense-setup` skill. Translating an Obsidian Bases `.base` file into equivalent queries is the `sense-bases` skill.
130
- - `map` and `status` report each preset's coverage (files matched, embedded count). Indexing derives from presets, so the coverage numbers are how you see what a config actually indexes and embeds. A scope with fewer signals just uses fewer (a preset without the vectors signal searches lexically); a saved search naming an unknown preset errors when run, listing the declared ones.
131
- - Save a query into `sense.config.json` only when it will be reused; run ad-hoc otherwise.
132
- - A one-line `summary:` per note is optional and pays twice: it appears in result rows and is a weighted search field. Date comparisons work for dates written as ISO 8601 (`2026-08-12`, or with time and offset); other formats do not compare. Field names in examples (`status`, `tags`, `created`) are illustrative; your tree defines its own.
133
- - Reserved frontmatter keys (dropped with a warning): `path`, `_mtime`, `_ctime`, `_size`, `_rank`, `_parse_error`, `content`, `links`, `sections`. The `tags` frontmatter column and the `tags` table coexist, mirroring Obsidian's own split: the column is the raw YAML list one note's frontmatter declares (Obsidian's `tags` property), the table is the merged, deduplicated frontmatter+inline set per note (what Obsidian's tag pane and Bases' `file.tags` read). "What is tagged X" is a table query; the column answers only what a note's frontmatter literally says. Inline tags inside `%%...%%` comments are indexed, and some trees run their whole maintenance-tag system in comments.
134
- - A note whose frontmatter does not parse is indexed with **no** frontmatter columns and `_parse_error` set to the YAML message, which carries the line. Nothing is half-recovered: a non-NULL value is a value the author wrote. So a NULL column means the key was absent *or* the note did not parse, and `_parse_error` is how you tell: `WHERE status IS NULL AND _parse_error IS NULL` is "genuinely missing status". List what needs fixing with `sense sql "SELECT path, _parse_error FROM frontmatter WHERE _parse_error IS NOT NULL"`; fixing a file clears it on the next command. `sense status` reports the count.
135
- - Exit codes: `0` ok, `1` error (store message verbatim), `2` usage (unknown query, wrong param count).
136
- - Doubted cache: delete the directory `sense status` prints on its `cache:` line. Rarely needed; every query reconciles first.
134
+ - Install a missing CLI with `npm install -g sensemaking`.
135
+ - `sense status` prints the cache path and watcher state.
136
+ - Delete the cache directory printed by `sense status` only when the derived index is in doubt. The next command rebuilds it.
137
+ - Use `sense watch` when another process should keep the index warm during frequent edits. Queries remain responsible for their own freshness check.
@@ -0,0 +1,47 @@
1
+ # Search evidence and troubleshooting
2
+
3
+ Read this guide when search results will support a factual claim, when absence matters, when results look noisy, or when changing signal weights.
4
+
5
+ ## What each signal establishes
6
+
7
+ `words` ranks literal occurrences after stemming. It is the right signal for identifiers, error strings, names, and quoted phrases. A `match` row is lexical evidence because its `snippets` show the occurrence.
8
+
9
+ `links` expands from matching notes through links written by the authors. A `link` row establishes that relationship, but does not establish that the row contains the query words.
10
+
11
+ `vectors` ranks the meaning of chunks. It can find paraphrases and related concepts with no shared vocabulary. A `vector` row is a lead to read. It cannot prove that a word or claim appears in the note.
12
+
13
+ Search uses every signal named by the preset. If `signals` is absent, every signal whose prerequisites hold has weight 1. A number changes that signal's contribution to reciprocal-rank fusion:
14
+
15
+ ```json
16
+ "signals": { "words": 1, "links": 1, "vectors": 4 }
17
+ ```
18
+
19
+ Weights are corpus and model choices. Compare representative queries before changing them. One weight does not transfer reliably between unrelated trees or embedding models.
20
+
21
+ ## Reading the rows
22
+
23
+ - `via` names the evidence that produced the row.
24
+ - `snippets` contains marked lexical passages. It is empty when no word matched.
25
+ - `lines` names the best section to read. Null means the whole note is the reference.
26
+ - `score` orders the fused result. Its scale changes with the participating signals and ranks, so compare rows only inside one result.
27
+ - `similarity` is cosine similarity against the best chunk. Compare it inside one result and one model. Small trees can give unrelated text a moderately close nearest neighbor.
28
+
29
+ Relay the evidence label when reporting a result. "The note contains these words" and "the note is semantically related" support different claims.
30
+
31
+ ## When a search misses
32
+
33
+ For word search, try concrete terms that the notes may use, widen the selected preset, or raise `--k`. Search grammar depends on the selected store. Read its guide before adding operators.
34
+
35
+ For vector search, restate the concept in different words. Then use `related <path>` on the nearest useful hit to inspect its semantic neighborhood.
36
+
37
+ Each widening step adds candidates and can dilute the ranking. Inspect the new rows before widening again.
38
+
39
+ ## Absence
40
+
41
+ A words-only preset returns no rows when the terms do not occur in its indexed scope. With vectors enabled, nearest-neighbor search can still return vector-only rows for any input. Those rows show conceptual proximity, not lexical presence.
42
+
43
+ When absence matters, use a preset whose `signals` excludes `vectors`, or query the selected store's text index as described in its guide. Confirm the preset's coverage with `sense status` before concluding that the terms are absent from the tree.
44
+
45
+ ## Template-heavy trees
46
+
47
+ Vectors rank whole chunks, including repeated boilerplate. In a tree whose notes share a large template and contain little unique prose, unrelated seeds can return the same neighbors with almost identical similarities. Compare two unlike seed notes. If their neighbor lists barely change, search the fields or words that distinguish the notes instead.
@@ -0,0 +1,120 @@
1
+ # Portable SQL
2
+
3
+ Read this guide for raw SQL over sense tables. These patterns use the shared schema and avoid native text-index syntax. Read the selected store's query guide as well when the statement uses dates, JSON table functions, native types, or text matching.
4
+
5
+ ## Discover the tree before assuming fields
6
+
7
+ Frontmatter columns come from the indexed notes. Inspect them before writing a field query:
8
+
9
+ ```sh
10
+ sense map
11
+ sense sql "SELECT name FROM pragma_table_info('frontmatter')"
12
+ sense sql "SELECT DISTINCT status FROM frontmatter ORDER BY status"
13
+ ```
14
+
15
+ Punctuated field names need double quotes. Values should use `?` parameters:
16
+
17
+ ```sh
18
+ sense sql 'SELECT path FROM frontmatter WHERE "plugin-id" = ? LIMIT 50' example
19
+ ```
20
+
21
+ `content.title` and `content.summary` always exist. A frontmatter column with either name exists only when at least one indexed note declares it.
22
+
23
+ ## Shared tables
24
+
25
+ | Table | Main columns |
26
+ |---|---|
27
+ | `frontmatter` | `path`, discovered fields, `_mtime`, `_ctime`, `_size`, `_rank`, `_parse_error` |
28
+ | `content` | `path`, `title`, `summary`, `text` |
29
+ | `links` | `src`, `target`, `dst`, `embed` |
30
+ | `tags` | `path`, `tag` |
31
+ | `sections` | `path`, `heading`, `level`, `start_line`, `end_line`, `tokens` |
32
+ | `preset_files` | `path`, `preset` |
33
+
34
+ `links.dst` is null when the written target does not resolve to indexed markdown. `links.embed` is 1 for an embed and 0 for an ordinary wikilink. The `tags` table merges frontmatter and inline tags and stores a nested tag such as `book/scifi` in full.
35
+
36
+ ## Shared query patterns
37
+
38
+ ```sql
39
+ SELECT COUNT(*) AS notes FROM frontmatter;
40
+
41
+ SELECT path, title
42
+ FROM frontmatter
43
+ WHERE status = ?
44
+ ORDER BY path
45
+ LIMIT 50;
46
+
47
+ SELECT src
48
+ FROM links
49
+ WHERE dst = ?
50
+ ORDER BY src;
51
+
52
+ SELECT heading, start_line, end_line, tokens
53
+ FROM sections
54
+ WHERE path = ?
55
+ ORDER BY start_line;
56
+
57
+ SELECT path, tag
58
+ FROM tags
59
+ WHERE tag = ? OR tag LIKE ?
60
+ ORDER BY path;
61
+ ```
62
+
63
+ For a nested tag family, bind the same value twice as `book` and `book/%`.
64
+
65
+ `has(field, value)` works on every store. It tests membership for a JSON array and substring presence for a scalar string. Use `field = ?` for exact scalar equality because `has(status, 'active')` also matches `inactive`.
66
+
67
+ `basename(path[, suffix])` works on every store. `segment(terms)` is available on SQLite and DuckDB only; the store guides explain when it is needed.
68
+
69
+ ## Preset scope
70
+
71
+ `sense sql` covers the whole index unless it receives `--preset`. With that flag, sense binds a temporary `scope(path)` table and requires the statement to join it:
72
+
73
+ ```sh
74
+ sense sql "SELECT f.path FROM frontmatter f JOIN scope ON scope.path = f.path ORDER BY f.path" --preset default
75
+ ```
76
+
77
+ A saved SQL query written against `scope` can run under different presets without duplicating the statement. Without `--preset`, join `preset_files` directly and name the preset in SQL.
78
+
79
+ ## Graph queries
80
+
81
+ Bound recursive walks. An unrestricted walk on a dense graph can enumerate paths faster than it eliminates them.
82
+
83
+ ```sql
84
+ WITH RECURSIVE hop(path, d) AS (
85
+ SELECT ?, 0
86
+ UNION
87
+ SELECT CASE WHEN l.src = hop.path THEN l.dst ELSE l.src END, hop.d + 1
88
+ FROM hop
89
+ JOIN links l ON (l.src = hop.path OR l.dst = hop.path) AND l.dst IS NOT NULL
90
+ WHERE hop.d < 2
91
+ )
92
+ SELECT DISTINCT path FROM hop WHERE d > 0;
93
+ ```
94
+
95
+ Use `sense path` for the shortest chain between two known notes. Use raw recursion when the task needs a set of neighbors to filter or join.
96
+
97
+ ## Dead links
98
+
99
+ `dst IS NULL` includes links to attachments that sense never indexes, such as images, PDFs, and `.base` files. Exclude the attachment extensions used by the tree before treating the remaining rows as broken links. Trees with dotted markdown titles need an explicit extension list instead of a blanket "contains a dot" filter.
100
+
101
+ Templates and examples can also contain deliberately unresolved wikilinks. Scope the query to authored content when those files are indexed.
102
+
103
+ ## Types and parse errors
104
+
105
+ YAML strings remain text, whole numbers and booleans remain integer-like values, fractions remain real-like values, and lists and maps remain structured or JSON-backed values according to the store. Use `sense map` to see the types observed for each field. Read the store guide before comparing a field that contains more than one type.
106
+
107
+ A malformed frontmatter block produces `_parse_error` and no recovered field values. This query lists the files to fix:
108
+
109
+ ```sql
110
+ SELECT path, _parse_error
111
+ FROM frontmatter
112
+ WHERE _parse_error IS NOT NULL
113
+ ORDER BY path;
114
+ ```
115
+
116
+ To find notes that genuinely omit `status`, use `status IS NULL AND _parse_error IS NULL`.
117
+
118
+ ## Bound the output
119
+
120
+ Select only the columns required by the task and add `LIMIT` while exploring. Avoid `SELECT text FROM content` because it returns the tree's prose. Use `search` for bounded passages, or select `path`, `title`, and `summary` and read the relevant files afterward.