sensemaking 0.24.4 → 0.24.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (130) hide show
  1. package/README.md +16 -12
  2. package/dist/cjs/cli/status.js +14 -24
  3. package/dist/cjs/cli/status.js.map +1 -1
  4. package/dist/cjs/commands/search.js +22 -0
  5. package/dist/cjs/commands/search.js.map +1 -1
  6. package/dist/cjs/config/types.js.map +1 -1
  7. package/dist/cjs/errors.d.cts +1 -1
  8. package/dist/cjs/errors.d.ts +1 -1
  9. package/dist/cjs/errors.js.map +1 -1
  10. package/dist/cjs/lib/watch-notification.d.cts +1 -0
  11. package/dist/cjs/lib/watch-notification.d.ts +1 -0
  12. package/dist/cjs/lib/watch-notification.js +15 -0
  13. package/dist/cjs/lib/watch-notification.js.map +1 -0
  14. package/dist/cjs/lib/worker-file.d.cts +1 -0
  15. package/dist/cjs/lib/worker-file.d.ts +1 -0
  16. package/dist/cjs/lib/worker-file.js +34 -0
  17. package/dist/cjs/lib/worker-file.js.map +1 -0
  18. package/dist/cjs/output/search-error.js +38 -0
  19. package/dist/cjs/output/search-error.js.map +1 -1
  20. package/dist/cjs/scan/pool.js +2 -23
  21. package/dist/cjs/scan/pool.js.map +1 -1
  22. package/dist/cjs/store/builder.d.cts +2 -1
  23. package/dist/cjs/store/builder.d.ts +2 -1
  24. package/dist/cjs/store/builder.js +4 -4
  25. package/dist/cjs/store/builder.js.map +1 -1
  26. package/dist/cjs/store/duckdb/lexical.d.cts +1 -0
  27. package/dist/cjs/store/duckdb/lexical.d.ts +1 -0
  28. package/dist/cjs/store/duckdb/lexical.js +28 -9
  29. package/dist/cjs/store/duckdb/lexical.js.map +1 -1
  30. package/dist/cjs/store/duckdb/native.d.cts +2 -0
  31. package/dist/cjs/store/duckdb/native.d.ts +2 -0
  32. package/dist/cjs/store/duckdb/native.js +11 -1
  33. package/dist/cjs/store/duckdb/native.js.map +1 -1
  34. package/dist/cjs/store/duckdb/ordered-bm25.d.cts +17 -0
  35. package/dist/cjs/store/duckdb/ordered-bm25.d.ts +17 -0
  36. package/dist/cjs/store/duckdb/ordered-bm25.js +238 -0
  37. package/dist/cjs/store/duckdb/ordered-bm25.js.map +1 -0
  38. package/dist/cjs/store/lock-wait.d.cts +2 -2
  39. package/dist/cjs/store/lock-wait.d.ts +2 -2
  40. package/dist/cjs/store/lock-wait.js +14 -11
  41. package/dist/cjs/store/lock-wait.js.map +1 -1
  42. package/dist/cjs/store/native.d.cts +4 -0
  43. package/dist/cjs/store/native.d.ts +4 -0
  44. package/dist/cjs/store/native.js +64 -9
  45. package/dist/cjs/store/native.js.map +1 -1
  46. package/dist/cjs/store/open.js +11 -7
  47. package/dist/cjs/store/open.js.map +1 -1
  48. package/dist/cjs/store/reconcile.d.cts +5 -1
  49. package/dist/cjs/store/reconcile.d.ts +5 -1
  50. package/dist/cjs/store/reconcile.js +5 -4
  51. package/dist/cjs/store/reconcile.js.map +1 -1
  52. package/dist/cjs/store/turso/lexical.js +5 -1
  53. package/dist/cjs/store/turso/lexical.js.map +1 -1
  54. package/dist/cjs/text/segment.js +168 -3
  55. package/dist/cjs/text/segment.js.map +1 -1
  56. package/dist/cjs/watch-claim.d.cts +39 -0
  57. package/dist/cjs/watch-claim.d.ts +39 -0
  58. package/dist/cjs/watch-claim.js +446 -0
  59. package/dist/cjs/watch-claim.js.map +1 -0
  60. package/dist/cjs/watch.js +246 -147
  61. package/dist/cjs/watch.js.map +1 -1
  62. package/dist/cjs/workers/watch-heartbeat.d.cts +1 -0
  63. package/dist/cjs/workers/watch-heartbeat.d.ts +1 -0
  64. package/dist/cjs/workers/watch-heartbeat.js +308 -0
  65. package/dist/cjs/workers/watch-heartbeat.js.map +1 -0
  66. package/dist/esm/cli/status.js +6 -5
  67. package/dist/esm/cli/status.js.map +1 -1
  68. package/dist/esm/commands/search.js +21 -0
  69. package/dist/esm/commands/search.js.map +1 -1
  70. package/dist/esm/config/types.js.map +1 -1
  71. package/dist/esm/errors.d.ts +1 -1
  72. package/dist/esm/errors.js.map +1 -1
  73. package/dist/esm/lib/watch-notification.d.ts +1 -0
  74. package/dist/esm/lib/watch-notification.js +4 -0
  75. package/dist/esm/lib/watch-notification.js.map +1 -0
  76. package/dist/esm/lib/worker-file.d.ts +1 -0
  77. package/dist/esm/lib/worker-file.js +24 -0
  78. package/dist/esm/lib/worker-file.js.map +1 -0
  79. package/dist/esm/output/search-error.js +21 -0
  80. package/dist/esm/output/search-error.js.map +1 -1
  81. package/dist/esm/scan/pool.js +2 -22
  82. package/dist/esm/scan/pool.js.map +1 -1
  83. package/dist/esm/store/builder.d.ts +2 -1
  84. package/dist/esm/store/builder.js +4 -4
  85. package/dist/esm/store/builder.js.map +1 -1
  86. package/dist/esm/store/duckdb/lexical.d.ts +1 -0
  87. package/dist/esm/store/duckdb/lexical.js +16 -7
  88. package/dist/esm/store/duckdb/lexical.js.map +1 -1
  89. package/dist/esm/store/duckdb/native.d.ts +2 -0
  90. package/dist/esm/store/duckdb/native.js +5 -1
  91. package/dist/esm/store/duckdb/native.js.map +1 -1
  92. package/dist/esm/store/duckdb/ordered-bm25.d.ts +17 -0
  93. package/dist/esm/store/duckdb/ordered-bm25.js +56 -0
  94. package/dist/esm/store/duckdb/ordered-bm25.js.map +1 -0
  95. package/dist/esm/store/lock-wait.d.ts +2 -2
  96. package/dist/esm/store/lock-wait.js +17 -14
  97. package/dist/esm/store/lock-wait.js.map +1 -1
  98. package/dist/esm/store/native.d.ts +4 -0
  99. package/dist/esm/store/native.js +60 -6
  100. package/dist/esm/store/native.js.map +1 -1
  101. package/dist/esm/store/open.js +9 -6
  102. package/dist/esm/store/open.js.map +1 -1
  103. package/dist/esm/store/reconcile.d.ts +5 -1
  104. package/dist/esm/store/reconcile.js +4 -3
  105. package/dist/esm/store/reconcile.js.map +1 -1
  106. package/dist/esm/store/turso/lexical.js +5 -1
  107. package/dist/esm/store/turso/lexical.js.map +1 -1
  108. package/dist/esm/text/segment.js +130 -4
  109. package/dist/esm/text/segment.js.map +1 -1
  110. package/dist/esm/watch-claim.d.ts +39 -0
  111. package/dist/esm/watch-claim.js +176 -0
  112. package/dist/esm/watch-claim.js.map +1 -0
  113. package/dist/esm/watch.js +169 -118
  114. package/dist/esm/watch.js.map +1 -1
  115. package/dist/esm/workers/watch-heartbeat.d.ts +1 -0
  116. package/dist/esm/workers/watch-heartbeat.js +106 -0
  117. package/dist/esm/workers/watch-heartbeat.js.map +1 -0
  118. package/package.json +2 -2
  119. package/skills/sense/SKILL.md +105 -104
  120. package/skills/sense/references/search.md +47 -0
  121. package/skills/sense/references/sql.md +120 -0
  122. package/skills/sense/references/stores/duckdb.md +35 -0
  123. package/skills/sense/references/stores/sqlite.md +66 -0
  124. package/skills/sense/references/stores/turso.md +36 -0
  125. package/skills/sense-setup/SKILL.md +93 -34
  126. package/skills/sense-setup/references/embeddings.md +51 -0
  127. package/skills/sense-setup/references/store-benchmarks.md +16 -0
  128. package/skills/sense-setup/references/stores/duckdb.md +28 -0
  129. package/skills/sense-setup/references/stores/sqlite.md +21 -0
  130. package/skills/sense-setup/references/stores/turso.md +25 -0
@@ -0,0 +1,35 @@
1
+ # DuckDB query guide
2
+
3
+ Read this guide when `store` is `duckdb`.
4
+
5
+ Sense stores the cache at `.sense/cache.duckdb`. It is a DuckDB database file with frontmatter values in `VARIANT` columns and embeddings in native fixed-width float arrays. Use this store when another DuckDB tool needs to inspect or analyze the cache.
6
+
7
+ The first open installs `@duckdb/node-api` if needed. Its native download is about 110 MB. DuckDB holds the cache file for the connection's lifetime, so another sense command waits for the current command or watcher cycle to close it.
8
+
9
+ ## `sense search` grammar
10
+
11
+ Bare words and quoted phrases work. Bare words are joined with AND, and phrases are verified against authored text.
12
+
13
+ The portable search command does not accept FTS5 prefix, boolean, `NEAR`, initial-token, column-filter, or grouping operators. Examples include `foo*`, `a OR b`, `NEAR(a b)`, `^term`, `title:term`, and `(a b)`. Sense raises `STORE_CAPABILITY_MISSING` instead of interpreting them as literals.
14
+
15
+ To express alternatives, run separate bounded searches and combine the paths, or use a frontmatter filter when the distinction is structured.
16
+
17
+ ## Raw text search
18
+
19
+ `content` is a plain table. SQLite's `content MATCH`, `snippet()`, and `bm25()` do not run here. Use `sense search` for ranked text retrieval and combine it with `--where` for frontmatter filters.
20
+
21
+ DuckDB maintains its own native FTS index for the search command. Its generated index name and scoring adapter are implementation details. Saved queries should not call them.
22
+
23
+ ## Functions and types
24
+
25
+ `has()`, `basename()`, and `segment()` are registered SQL functions. `segment()` is available for portable saved queries, although `sense search` already handles unspaced scripts.
26
+
27
+ Frontmatter columns use DuckDB `VARIANT`. Homogeneous values compare normally. A numeric comparison against a field that contains numbers in some notes and text in others can raise a type error. Check `sense map` for mixed observed types before writing the predicate.
28
+
29
+ DuckDB has its own date, JSON, list, and casting syntax. Use DuckDB SQL rather than copying SQLite functions into a saved query. For ISO 8601 timestamps with offsets, cast to `TIMESTAMPTZ` before comparing:
30
+
31
+ ```sql
32
+ WHERE created::TIMESTAMPTZ >= ?::TIMESTAMPTZ
33
+ ```
34
+
35
+ DuckDB's dialect closely follows PostgreSQL but has documented differences and its own extensions. Read the official [DuckDB SQL introduction](https://duckdb.org/docs/current/sql/introduction) for engine SQL. Queries written with DuckDB-only syntax stay tied to this store.
@@ -0,0 +1,66 @@
1
+ # SQLite query guide
2
+
3
+ Read this guide when `store` is `sqlite` or omitted.
4
+
5
+ SQLite is the default store. It uses Node's built-in SQLite, keeps the cache at `.sense/cache.db`, supports concurrent sense commands through WAL, and implements the full search grammar documented here.
6
+
7
+ ## `sense search` grammar
8
+
9
+ Search text uses FTS5 syntax:
10
+
11
+ - Bare words are joined with AND. If one word is absent, the lexical signal returns no row.
12
+ - `a OR b` matches either word.
13
+ - `pref*` matches a token prefix.
14
+ - `"exact phrase"` requires adjacent tokens.
15
+ - `NEAR(a b, 5)` limits token distance.
16
+ - `^term` requires the token at the start of a column.
17
+ - `title:term` and `summary:term` filter by indexed column.
18
+ - Parentheses group expressions.
19
+
20
+ Double-quote terms containing punctuation. Bare `customer-facing` treats the hyphen as syntax, and a bare apostrophe is invalid.
21
+
22
+ ## Raw text search
23
+
24
+ `content` is an FTS5 virtual table. Refer to it by its table name in `MATCH`, even when the `FROM` clause has an alias:
25
+
26
+ ```sql
27
+ SELECT content.path, content.title
28
+ FROM content
29
+ WHERE content MATCH ?
30
+ ORDER BY bm25(content, 10.0, 5.0, 1.0)
31
+ LIMIT 20;
32
+ ```
33
+
34
+ The three BM25 weights rank title above summary above body. For queries that can match the machine-written unspaced-script sidecars, use all eight weights so the sidecars receive the same field weighting:
35
+
36
+ ```sql
37
+ ORDER BY bm25(content, 10.0, 5.0, 1.0, 0, 10.0, 5.0, 1.0)
38
+ ```
39
+
40
+ Raw `snippet(content, 2, '«', '»', '…', 10)` targets the authored body column. Avoid `-1` because it may choose a machine-written sidecar. SQLite's `snippet()` tokenizes the matched text again, so prefer `sense search` for large notes. Sense builds its own bounded passages on every store.
41
+
42
+ ## Functions and unspaced scripts
43
+
44
+ `has()`, `basename()`, and `segment()` are registered SQL functions. `sense search` handles Chinese, Japanese, Thai, Khmer, Lao, and Burmese automatically. Hand-written `MATCH` cannot be rewritten, so pass such terms through `segment()`:
45
+
46
+ ```sql
47
+ SELECT path FROM content WHERE content MATCH segment(?);
48
+ ```
49
+
50
+ `segment()` leaves other text unchanged.
51
+
52
+ ## Dates and values
53
+
54
+ Use SQLite's `datetime()` to compare ISO 8601 values with different offsets:
55
+
56
+ ```sql
57
+ WHERE datetime(created) >= datetime(?)
58
+ ```
59
+
60
+ SQLite's `now` is UTC. A calendar-day query in the machine's local timezone needs `localtime`:
61
+
62
+ ```sql
63
+ WHERE date(created) = date('now', 'localtime')
64
+ ```
65
+
66
+ Frontmatter values use SQLite storage classes. Booleans are integers, so compare true with 1. Lists and maps are JSON text and can be expanded with `json_each()`.
@@ -0,0 +1,36 @@
1
+ # Turso query guide
2
+
3
+ Read this guide when `store` is `turso`.
4
+
5
+ Sense stores the cache at `.sense/cache.turso.db` and uses `@tursodatabase/database` with Tantivy full-text indexes. Use this experimental adapter when the tree benefits from the embedded Turso engine or its text-index behavior.
6
+
7
+ The first open installs the native package if needed. Turso holds the cache file for the connection's lifetime, so another Sense command waits for the current command or watcher cycle to close it. The upstream engine supports more deployment and concurrency options than this adapter currently exposes. Sense opens a local file and does not configure sync, remote access, concurrent writers, or encryption.
8
+
9
+ ## `sense search` grammar
10
+
11
+ Bare words and quoted phrases work. Sense applies shared stemming and joins bare words with AND. It verifies phrases against authored text and uses n-gram sidecars for unspaced scripts.
12
+
13
+ The portable search command does not accept FTS5 prefix, boolean, `NEAR`, initial-token or boost, column-filter, or grouping operators. Examples include `foo*`, `a OR b`, `NEAR(a b)`, `^term`, `title:term`, and `(a b)`. Sense raises `STORE_CAPABILITY_MISSING` instead of passing those forms to Tantivy.
14
+
15
+ To express alternatives, run separate bounded searches and combine the paths, or use a frontmatter filter when the distinction is structured.
16
+
17
+ ## Raw text search
18
+
19
+ `content` is a regular table with store-owned stem and n-gram columns. SQLite's `content MATCH`, `snippet()`, and `bm25()` do not run here. Use `sense search` for ranked text retrieval.
20
+
21
+ Turso exposes native `fts_match()` and `fts_score()`, but their valid scoring forms depend on the exact index projection. Sense owns those calls so that scores do not silently flatten. Saved queries should not call the store-owned text indexes.
22
+
23
+ ## Functions and types
24
+
25
+ `has()` and `basename()` work in raw SQL. Sense rewrites them into portable SQL because the client cannot register user-defined functions. Nested calls and placeholders are supported.
26
+
27
+ `segment()` is unavailable. A query that calls it raises `STORE_CAPABILITY_MISSING`. `sense search` still handles unspaced scripts without it.
28
+
29
+ Turso follows SQLite syntax for common date and JSON operations. Use `datetime()` for ISO 8601 comparisons and add `localtime` when a calendar-day query means the machine's local day:
30
+
31
+ ```sql
32
+ WHERE datetime(created) >= datetime(?)
33
+ WHERE date(created) = date('now', 'localtime')
34
+ ```
35
+
36
+ Frontmatter lists and maps are stored as JSON text and can be expanded with `json_each()`.
@@ -1,53 +1,112 @@
1
1
  ---
2
2
  name: sense-setup
3
- description: "Set up the sense CLI on a markdown tree and make the tree-design decisions that shape it: sense init, presets (which files, which settings, whether the scope searches by meaning), the embed block that names the model, and the trade-offs of frontmatter conventions, summaries, folder layout, and note size. Use when creating or restructuring a markdown knowledge base, running sense init, editing sense.config.json, choosing the backing store (sqlite, duckdb or turso), configuring search scope or vectors, or deciding how notes should be written for an agent to query later."
3
+ description: "Set up the sense CLI on a markdown tree and make the decisions that shape its index: run sense init, choose sqlite, duckdb, or turso, define presets and search signals, configure embeddings, and design queryable note conventions. Use when creating or restructuring a markdown knowledge base, editing sense.config.json, choosing a store or embedding model, or deciding how agents should write notes for later retrieval."
4
4
  ---
5
5
 
6
- # sense: setup and tree design
6
+ # sense setup and tree design
7
7
 
8
- Querying an existing tree is the `sense` skill. This one covers making a tree: installing, writing presets, and the design decisions a tree owner faces. Worked configurations for common tree shapes: [EXAMPLES.md](EXAMPLES.md).
8
+ Use this skill to create or change a sense configuration. Querying an existing tree belongs to the `sense` skill. Worked configurations for common tree shapes are in [EXAMPLES.md](EXAMPLES.md).
9
9
 
10
- ## Setup
10
+ ## Set up the tree
11
11
 
12
- - `npm install -g sensemaking`, then `sense init` at the tree root writes `sense.config.json`: two presets (`default`, and `large` showing what a big tree tunes) and an `embed` block naming the model. The model fetches once per machine at the first vector search (progress on stderr); `sense download` prefetches it instead where that timing matters (CI, air-gapped setup). Config discovery walks up from cwd; `--config <path>` overrides. A top-level `root` can point the config at another markdown tree: it resolves relative to the config directory, makes all globs and stored paths tree-relative, while `.sense/` remains beside the config. Use this when one vault needs independent consumer-specific presets, queries, and caches.
13
- - **Backing store.** The config's `store` key: `sqlite` (default, zero-dependency, Node's built-in SQLite), or the experimental `duckdb` and `turso`, whose engine package the first command that opens such a tree installs on its own (`@duckdb/node-api`, a one-time native download of about 110 MB; `@tursodatabase/database`, much smaller). The same commands and tables run on all three. Two things do not port, and each one decides a tree. **FTS5 syntax:** under `duckdb` and `turso`, `search` text and raw `MATCH` reject FTS5's prefix, boolean, `NEAR`, initial-token and column-filter operators with a named error, and sqlite's FTS5 SQL (`MATCH`, `snippet()`, `bm25()`) does not run, so saved queries written in that syntax are sqlite dialect; a tree whose saved queries or search vocabulary depend on FTS5 operators stays on `sqlite`. **SQL functions:** `has`/`basename` run on all three (`turso` rewrites them into portable SQL rather than registering them); `segment` runs on `sqlite` and `duckdb` only, so a tree whose queries call `segment` stays off `turso`. `sense watch` runs on all three; `duckdb` and `turso` lock their cache file per connection, so a concurrent command waits out the watcher's current cycle instead of failing. Each store keeps its own cache file (`.sense/cache.db`, `.sense/cache.duckdb`, `.sense/cache.turso.db`); switching stores is a rebuild, not a migration. Speed is a separate axis and must come from a current identical-work measurement. Historical per-store rows can include different ranked candidates and are diagnostics, not an overall store leaderboard. The cache remains an ordinary database file, so a tree indexed under duckdb is readable by DuckDB's ecosystem, while turso brings concurrent access, non-blocking I/O and encryption. Pick for the capabilities you need, then review the current evidence in BENCHMARKING.md.
14
- - Globs resolve relative to `root` when present, otherwise the config file directory; never the cwd.
15
- - `sense status` and `sense map` show each preset's coverage (files matched, embedded count), so what a config actually indexes is always visible in output. A config edit that changes coverage rebuilds the cache and names the preset that caused it on stderr.
12
+ Install the CLI and initialize it at the markdown root:
16
13
 
17
- ## Presets
14
+ ```sh
15
+ npm install -g sensemaking
16
+ cd path/to/notes
17
+ sense init
18
+ sense status
19
+ ```
18
20
 
19
- A preset is a named, self-contained bundle of settings. `default` (required) is what bare commands use; every other preset is addressed by name (`sense search "..." --preset raw`, or `"preset": "raw"` in a saved search). No inheritance: what a preset states is all it does.
21
+ `sense init` writes `sense.config.json`. Config discovery walks upward from the current directory. `--config <path>` selects another file.
20
22
 
21
- | field | means | default |
23
+ By default, the config indexes its own directory. A top-level `root` can point at another markdown tree. Relative roots resolve from the config directory, and `.sense/` remains beside the config. Use this when separate consumers need their own presets, saved queries, or caches over one tree.
24
+
25
+ Globs and indexed paths are relative to the selected root. Run `sense status` and `sense map` after changing the config to confirm preset coverage and the selected store. A change that affects indexed content rebuilds the relevant cache and reports the reason.
26
+
27
+ ## Choose the store
28
+
29
+ Choose from the intended use of the cache. Start with `sqlite` when no surrounding workflow favors
30
+ another engine, then consider interoperability, SQL, connection behavior, and representative Sense
31
+ measurements. A current timing result does not define an engine's long-term suitability.
32
+
33
+ | Store | Choose it when | Main trade-off |
34
+ |---|---|---|
35
+ | `sqlite` | The tree needs the smallest setup, advanced FTS5 search syntax, or concurrent Sense commands | Included with Node; raw SQL is SQLite |
36
+ | `duckdb` | The cache belongs in DuckDB analytical work over large datasets or in a local/cloud workflow, or needs DuckDB types and SQL | Experimental Sense adapter, large native install, one open connection at a time |
37
+ | `turso` | The tree benefits from the embedded Rust engine or Tantivy index, or is evaluating the engine as its adapter evolves | Experimental Sense adapter, restricted search grammar, one open connection at a time |
38
+
39
+ Read the matching selection guide before recommending or configuring a non-default store:
40
+
41
+ - [SQLite selection guide](references/stores/sqlite.md)
42
+ - [DuckDB selection guide](references/stores/duckdb.md)
43
+ - [Turso selection guide](references/stores/turso.md)
44
+
45
+ Read the [current store benchmark summary](references/store-benchmarks.md) only when performance could change the choice. The release assessment generates that shipped summary from the latest accepted all-store run.
46
+
47
+ Each store uses a separate cache file. Changing `store` rebuilds the index instead of migrating the old cache. Commands and public tables are shared, while raw SQL and advanced word-search syntax follow the selected engine.
48
+
49
+ ## Design presets
50
+
51
+ A preset is a named, self-contained scope. The required `default` preset serves bare commands. Other presets are selected by name.
52
+
53
+ | Field | Meaning | Default |
22
54
  |---|---|---|
23
- | `include` / `exclude` | which files this preset covers (globs) | required |
24
- | `k` | how many results a search returns | 10 |
25
- | `signals` | which engines this preset's searches compose, each mapped to its RRF weight (`{"words": 1, "links": 1, "vectors": 1}`) | every signal whose prerequisites hold, each at weight 1 |
26
- | `where` | a standing SQL filter on frontmatter | none |
55
+ | `include` and `exclude` | Files covered by the preset | `include` is required |
56
+ | `k` | Search result count | 10 |
57
+ | `signals` | Enabled search signals and their reciprocal-rank weights | Every available signal at weight 1 |
58
+ | `where` | Standing frontmatter filter | None |
59
+
60
+ A file is indexed when any preset includes it. Presets can overlap. Files outside every preset do not enter the index.
61
+
62
+ Use paths for stable layers such as `raw/`, `notes/`, and `archive/`. Use frontmatter filters for changing state such as `status`, `project`, and dates. A preset controls indexing, so its coverage must be computable from paths before frontmatter queries run.
63
+
64
+ Use separate presets when parts of the tree need different search signals or result counts. A single-purpose tree does not need extra preset vocabulary.
65
+
66
+ ## Configure vectors
67
+
68
+ Vectors require two choices. The top-level `embed` block names the model and provider. Each preset's `signals` decides whether that scope uses vectors.
69
+
70
+ ```json
71
+ {
72
+ "embed": {
73
+ "model": "minishlab/potion-retrieval-32M",
74
+ "provider": "static"
75
+ },
76
+ "presets": {
77
+ "default": {
78
+ "include": ["notes/**/*.md"],
79
+ "signals": { "words": 1, "links": 1, "vectors": 1 }
80
+ },
81
+ "raw": {
82
+ "include": ["raw/**/*.md"],
83
+ "signals": { "words": 1, "links": 1 }
84
+ }
85
+ }
86
+ }
87
+ ```
88
+
89
+ The first vector search downloads a named static model and embeds the covered notes. `sense download` fetches the model earlier when CI, offline work, or timing makes that useful. A config change that alters the model, vector coverage, or chunking can rebuild vectors.
90
+
91
+ Read [embedding setup](references/embeddings.md) when choosing a provider or model, supporting a non-English tree, changing chunk size, or tuning signal weights.
27
92
 
28
- **Indexing derives from presets.** A file is indexed if any preset includes it. Consequences worth designing around:
93
+ ## Design queryable notes
29
94
 
30
- - Files no preset includes are not indexed at all.
31
- - Presets may overlap; they are views, not partitions.
32
- - Global `features` (`links`, `sections`, `rank`) still toggle tree-wide; most trees never touch them.
95
+ Sense accepts heterogeneous markdown. These choices decide which queries will be reliable:
33
96
 
34
- **Vectors take two decisions, in two places.** The top-level `"embed": { "model", "provider": "static"|"openai"|"cohere", "url", "key", "chunkTokens" }` block names the model and says whether the tree has vectors at all. `chunkTokens` is a chunk size ceiling in estimated tokens for small-context models; default 500. A preset's `signals` says which engines that scope uses: a layer searched for exact wording (ingested sources, archives, generated output) declares `"signals": {"words": 1, "links": 1}`, costs no embedding, and its searches run on words and links. That is the main scale lever, and it is the llm-wiki split: compiled pages searched by meaning, raw sources searched for the phrasing you are citing. Each named signal's number is its RRF weight, not a toggle. 1 is the default and reproduces equal-weight fusion; a preset can instead raise one signal's number, e.g. `{"words": 1, "vectors": 4}`, to shift the fused ranking toward that signal without dropping the others. Whether that helps is corpus- and model-contingent, not a fixed rule: benchmark/reports/2026-08-27-embedding-model-selection.md's weight-sweep table measures equal weight against {0.5, 1, 2, 4} and vectors-only on nfcorpus and on MIRACL zh with an HTTP encoder, and the two corpora do not agree on which weight wins. `static` is the built-in pure-JS Model2Vec loader and handles paraphrase and reworded concepts; tight domain jargon ("heart attack" for "myocardial infarction") is where an `openai`- or `cohere`-shaped encoder model tends to do better, measured in the same report's encoder-tier tables. Naming the model in the config is the consent to fetch it: the first vector search downloads it once per machine into `~/.sense/models` (or `sense download` prefetches) (huggingface_hub's cache layout, one snapshot directory per resolved revision), so several models coexist and switching between them rebuilds the index rather than mixing vector spaces. A `model` holding a path instead of a Hugging Face id points at a local directory, which `sense download` reports as nothing to fetch. The first search after that embeds the tree (progress on stderr; minutes on tens of thousands of notes, seconds on small trees). A config edit that changes coverage or features rebuilds the cache, vectors included, so settle presets before the first vector search on a large tree or that embedding run is paid twice.
97
+ - Consistent frontmatter fields make exact filters and saved reports possible. Inspect actual coverage and types with `sense map`.
98
+ - ISO 8601 dates can be compared by the selected store's date functions. Mixed date formats can still be stored, but do not make a dependable time range.
99
+ - A one-line `summary` appears in results and receives more lexical weight than body text. Its cost is keeping it current.
100
+ - Folders are the natural unit for preset coverage. Frontmatter is the natural unit for status and ownership.
101
+ - Small notes give precise hits and cheap whole-file reads. Large notes still work because `peek`, `sections`, and search `lines` point at ranges, and vector chunking splits long sections.
102
+ - Save recurring questions under `queries`. Use `{ "sql": "..." }` for deterministic filters and reports, and `{ "search": "..." }` for ranked retrieval.
35
103
 
36
- **Finding a model.** The default, `potion-retrieval-32M`, is English only. A tree whose text is mostly a language the model does not declare fails loudly at embed time (`EMBED_MODEL_MISMATCH`, naming the fix), and `sense status` shows the detected language mix beside the model's declared languages. A model whose card declares no languages leaves that check off, and a mismatched pairing then degrades silently, nearest-neighbour search always returns a neighbour regardless of fit, so check the card's languages yourself when the tree is not English. Picking a model is a lookup against the source, not a name to memorize: for a static model, filter Hugging Face's model2vec library for the target language and read the card for its declared languages, its safetensors shape, and F32 weights (an int8 republish fails the loader's dtype check by design). For an encoder reached over HTTP, Ollama's embedding-model library and LM Studio's catalog list each model's languages and context length; both serve an OpenAI-shaped endpoint, so `provider: "openai"` with `url` pointing at the local port reaches either with no other config. Cohere's hosted models are reached with `provider: "cohere"` instead. One dated measurement stands in for a name table that would go stale: `potion-retrieval-32M` was the best-measured static English retrieval model as of 2026-08, per benchmark/reports/2026-08-27-embedding-model-selection.md's static-ladder table.
104
+ Field names in examples are illustrative. The tree defines its own schema. Reserved frontmatter keys are `path`, `_mtime`, `_ctime`, `_size`, `_rank`, `_parse_error`, `content`, `links`, and `sections`; sense drops them with a warning.
37
105
 
38
- **Large vaults**: everything except the vector build is measured linear to 100k notes with no tuning (BENCHMARKING.md). The knobs that matter are `k` (more, smaller results; rows carry `lines` section ranges, so agents read sections, not files) and `"signals": {"words": 1, "links": 1}` on the layers that do not earn vectors.
106
+ ## Validate the result
39
107
 
40
- ## Tree design decisions
108
+ Run `sense status` to confirm the store, cache, document count, embedding state, watcher state, and preset coverage. Run `sense map` to confirm the discovered fields and their observed types.
41
109
 
42
- These belong to the tree's owner. sense works with any of them and reads no instruction files of its own; each choice only changes what queries can do.
110
+ Run every saved query after editing it. A parameterized SQL query can use any value because preparing the statement validates its columns and syntax before the parameter changes the result.
43
111
 
44
- - **Frontmatter fields.** Columns are discovered per tree: whatever keys notes declare become queryable. Consistent fields across notes make SQL filters and saved queries possible (`WHERE status = 'active'`). The store's column limit bounds distinct keys per tree: sqlite's compiled 2,000 (sqlite.org/limits.html), turso's 2,000 result-set column limit, duckdb's 10,000 sanity fence (it has no compile-time cap). The crawl stops with an error naming the count and the levers. Reserved keys (dropped with a warning): `path`, `_mtime`, `_ctime`, `_size`, `_rank`, `content`, `links`, `sections`. Values keep their YAML type: strings TEXT, whole numbers and booleans INTEGER (`true` is 1), fractions REAL, lists and maps JSON text; `map` prints the observed type per field.
45
- - **Presets are path-shaped; frontmatter is state-shaped.** A preset's coverage must be computable from the path alone (it decides indexing, baked into the cache). Volatile state (`status`, `project`, dates) lives in frontmatter and filters at query time (`where`, `has()`, `datetime()`). A state worth different *indexing* (retired memory, superseded sources) is a state worth moving the file: the archive-folder pattern in EXAMPLES.md.
46
- - **What a note omits is also a filter.** A layer that deliberately carries none of the fields the saved views filter on is excluded from all of them without any view naming the layer. Sparse fields cut both ways: less of the tree filters when you want breadth, and exactly this separation when layers differ in authority.
47
- - **Dates.** `datetime()` comparisons work for dates written as ISO 8601, the only format it parses. A tree that mixes date formats can store them, but can't compare them in SQL.
48
- - **Language.** No decision needed. A language written without word spaces (Chinese, Japanese, Thai, Khmer, Lao, Burmese) is indexed per grapheme and searched as an ordered grapheme phrase, so `sense search "全文"` finds what it should. This is substring semantics, what `grep` gives: a query matches wherever its exact text occurs, including inside a longer run, and needs no minimum length. A language written with spaces is left exactly as it was, storing nothing extra, and its stemming is unaffected. The one place this does not reach is a hand-written `content MATCH '...'`, which cannot be rewritten for its author: pass the terms through `segment()` there. `content.tokenize` is a separate lever with a different purpose, substring matching inside a Latin word (`trigram`) or keeping hyphenated terms whole (`unicode61 tokenchars '-_'`); naming one turns the grapheme-phrase scheme off, since the tree has then chosen its own scheme. Changing it rebuilds the text index only; vectors, links, and sections are kept.
49
- - **Summaries.** A one-line `summary:` is optional and pays twice: it shows in every result row (often answering a question with no file read) and is a weighted search field ranked above body text. The cost is writing and maintaining the line as notes change.
50
- - **Folder shape.** Globs find the files, paths are queryable text, links resolve by basename at any depth, but presets make folders meaningful: a folder is the natural unit that gets its own coverage and settings.
51
- - **Note size.** Many small notes: precise search hits, whole-file reads stay cheap, more links to maintain. Fewer large notes: `sections`, `peek`, and the `lines` column carry the cost down to line-range reads. Both work. A long section no longer becomes one oversized vector either: chunking splits at headings and caps each chunk around 500 estimated tokens (CJK counted 1:1, spaced scripts ~4:1), so an oversized note degrades to more, smaller chunks rather than one truncated one; `embed.chunkTokens` lowers that cap further for a small-context model.
52
- - **Recurring questions.** Save a scenario an agent will repeat under `queries`, naming the verb it runs: `{ "sql": "..." }` for filters and reports, or `{ "search": "...", "preset": "raw", "k": 5 }` for a ranked search. Either runs as `sense <name>`, and running one is how it is validated: a typo'd column or an unknown preset errors and exits nonzero. A parameterised entry validates with any argument, since SQL is prepared before parameters bind.
53
- - **Where decisions live.** Choices that should outlive one conversation can be recorded in the agent's own instruction or skill files, or in a note in the tree itself; a one-off search over an existing corpus needs none of that.
112
+ Record choices that should outlive the setup conversation in the tree's maintained agent guidance or in an authored note. The sense config should contain executable settings and reusable queries, not prose policy.
@@ -0,0 +1,51 @@
1
+ # Embedding setup
2
+
3
+ Read this guide when choosing an embedding provider or model, supporting a non-English tree, changing chunk size, or tuning signal weights.
4
+
5
+ ## Provider choice
6
+
7
+ `static` loads a local Model2Vec model in JavaScript. The default model is `minishlab/potion-retrieval-32M`. A Hugging Face model id downloads into `~/.sense/models` on first vector use. A filesystem path points at a directory that already contains `model.safetensors` and `tokenizer.json`.
8
+
9
+ `openai` calls an OpenAI-compatible `/embeddings` endpoint. Local Ollama and LM Studio servers use this protocol. A hosted endpoint sends note text to that service.
10
+
11
+ `cohere` calls Cohere's native embedding endpoint. The `key` setting names the environment variable that contains the credential. Note text leaves the machine.
12
+
13
+ ## Model choice
14
+
15
+ The default static model is English-only. For another language, choose a model whose published card declares that language. A static Model2Vec model must have compatible safetensors and tokenizer files with supported floating-point weights.
16
+
17
+ Sense checks declared model languages against the indexed tree. A clear mismatch raises `EMBED_MODEL_MISMATCH`. A model card with no language declaration cannot be checked, so confirm it before indexing.
18
+
19
+ For a static model, inspect current Model2Vec models and their cards rather than relying on a fixed list in this skill. For a local HTTP provider, inspect the model catalog exposed by the chosen runtime. Model availability and model cards change independently of sense releases.
20
+
21
+ ## Download and rebuild behavior
22
+
23
+ Naming a remote static model in the config authorizes its download. `sense download` performs that fetch before the first query. The first vector-participating search then embeds the indexed notes.
24
+
25
+ Changing the model changes the vector space and rebuilds embeddings. Changing preset coverage can also add, remove, or rebuild vector rows. Settle the broad scope before embedding a large tree.
26
+
27
+ ## Chunk size
28
+
29
+ `embed.chunkTokens` sets the estimated token ceiling for chunks. The default is 500. Sense splits at headings and then applies the ceiling to long sections. Lower it for an embedding model with a smaller useful context window. Raising it trades fewer vectors for less precise line ranges and more text per vector.
30
+
31
+ ## Signal weights
32
+
33
+ A preset includes a signal by naming it:
34
+
35
+ ```json
36
+ "signals": { "words": 1, "links": 1, "vectors": 1 }
37
+ ```
38
+
39
+ The number is the signal's reciprocal-rank weight. It is not a boolean. Weight 1 gives equal contribution. A higher vector weight can help a strong encoder on some corpora and hurt another tree whose exact vocabulary carries the answer.
40
+
41
+ Test representative questions from the actual tree before changing weights. Compare returned paths and evidence labels, not only the fused score.
42
+
43
+ ## Layers without vectors
44
+
45
+ An `embed` block makes vectors available to the tree. A preset can still omit the vector signal:
46
+
47
+ ```json
48
+ "signals": { "words": 1, "links": 1 }
49
+ ```
50
+
51
+ Use that for layers where exact wording matters more than paraphrase, such as raw sources, archives, generated logs, or citation corpora. Those files can remain indexed without paying the embedding cost for that preset's scope.
@@ -0,0 +1,16 @@
1
+ <!-- sense-store-benchmark release=0.24.6 -->
2
+ # Current store benchmark summary
3
+
4
+ The release assessment generated this file for store selection. Release `0.24.6` passed on 2026-09-16, measured on Apple M4 Pro with Node v26.7.0.
5
+
6
+ The timing rows ran on the same 6,566-note tree. They include CLI startup and each store's complete selected path. Ranked candidates and downstream work can differ by store, so these are current operating measurements rather than an isolated database-engine contest.
7
+
8
+ | Store | Cold index | Warm count | Lexical search | Semantic search | Portable semantic nDCG@10 |
9
+ |---|---|---|---|---|---|
10
+ | sqlite | 1,471 ms | 143 ms | 192 ms | 318 ms | 0.3306 |
11
+ | duckdb | 2,149 ms | 162 ms | 279 ms | 399 ms | 0.3284 |
12
+ | turso | 2,727 ms | 138 ms | 237 ms | 584 ms | 0.3307 |
13
+
14
+ Cold index is the first `status` that builds the cache. Warm count is a no-change `COUNT(*)` query. Lexical and semantic search are steady-state `sense search` commands. Lower timing is faster. Higher nDCG@10 is better; that quality column uses the same NFCorpus queries, judgments, result count, and model on every store.
15
+
16
+ Choose from the intended workflow, capabilities, and SQL compatibility. These numbers describe the current Sense implementations, not a permanent ranking of the engines. Treat small timing or relevance differences as diagnostic unless a representative workload for the target tree reproduces them.
@@ -0,0 +1,28 @@
1
+ # Choose DuckDB
2
+
3
+ Choose DuckDB when the cache file is an analytical artifact that another DuckDB tool will query, when saved SQL needs DuckDB's types and functions, or when the cache belongs beside large local or cloud datasets.
4
+
5
+ DuckDB is designed for analytical SQL and supports larger-than-memory work by spilling to disk. Its SQL closely follows PostgreSQL while keeping documented differences and extensions. See the official [DuckDB performance guide](https://duckdb.org/docs/current/guides/performance/how_to_tune_workloads) and [SQL introduction](https://duckdb.org/docs/current/sql/introduction).
6
+
7
+ Sense's [store benchmark summary](../store-benchmarks.md) measures the current adapter paths on a named markdown tree. It does not measure DuckDB's general analytical ceiling. Use it when current Sense latency matters to the choice.
8
+
9
+ ## What the choice gives you
10
+
11
+ - A standard `.sense/cache.duckdb` database that DuckDB tools can open.
12
+ - Frontmatter in `VARIANT` columns, so lists, maps, booleans, numbers, and strings retain useful native shapes.
13
+ - Embeddings in native fixed-width float arrays.
14
+ - DuckDB SQL for downstream analysis.
15
+ - Compatibility with external DuckDB workflows. MotherDuck supports hybrid queries over local and cloud data, although Sense does not connect, upload, or replicate the cache for you. See [MotherDuck's hybrid execution overview](https://motherduck.com/research/motherduck-duckdb-in-the-cloud-and-in-the-client/).
16
+
17
+ ## Current Sense adapter boundaries
18
+
19
+ - DuckDB support is experimental.
20
+ - The first open installs `@duckdb/node-api`, with a native download of about 110 MB.
21
+ - DuckDB holds the cache file for the life of a connection. Concurrent sense commands wait for the current command or watcher cycle to close it.
22
+ - `sense search` accepts bare words and quoted phrases. It rejects FTS5 boolean, prefix, `NEAR`, initial-token, grouping, and column-filter syntax.
23
+ - SQLite's raw `MATCH`, `snippet()`, and `bm25()` queries do not port.
24
+ - Mixed-type `VARIANT` fields can require explicit casts in predicates.
25
+
26
+ Commands, public tables, `has()`, `basename()`, `segment()`, vectors, and `sense watch` remain available. Read the [DuckDB query guide](../../../sense/references/stores/duckdb.md) before translating saved SQL or advanced searches.
27
+
28
+ Changing to DuckDB builds its separate cache. It does not migrate the SQLite or Turso cache.
@@ -0,0 +1,21 @@
1
+ # Choose SQLite
2
+
3
+ SQLite is the default recommendation for a sense tree.
4
+
5
+ Choose it when the tree has no engine-specific requirement, when saved searches use FTS5 operators, or when people and agents may run sense commands at the same time.
6
+
7
+ ## What the choice gives you
8
+
9
+ - No optional store package. Sense uses Node's built-in SQLite.
10
+ - The full documented word-search grammar, including boolean, prefix, `NEAR`, initial-token, grouping, and column-filter expressions.
11
+ - Raw FTS5 SQL through `MATCH`, `bm25()`, and `snippet()`.
12
+ - The sense SQL functions `has()`, `basename()`, and `segment()`.
13
+ - WAL-backed concurrent sense commands.
14
+
15
+ The cache is `.sense/cache.db`. It is derived data and can be rebuilt from the markdown files.
16
+
17
+ ## When another store fits the workflow
18
+
19
+ Use DuckDB when downstream DuckDB tools need to query the cache, the tree needs DuckDB's analytical SQL and native types, or the cache belongs beside large datasets or in a local/cloud workflow. Use Turso when its embedded Rust engine or Tantivy index fits the intended integration. Compare current measurements in [the store benchmark summary](../store-benchmarks.md) when Sense performance affects the decision.
20
+
21
+ If existing queries use FTS5 syntax, moving away from SQLite requires rewriting those queries. Read the selected engine's query guide under the `sense` skill before changing the config.
@@ -0,0 +1,25 @@
1
+ # Choose Turso
2
+
3
+ Choose Turso when the tree benefits from the embedded Turso Database engine or its Tantivy text index, or when evaluating the engine as the Sense adapter evolves. The Sense adapter is experimental; the engine is a fast-moving Rust rewrite of SQLite whose trade-offs can change between releases.
4
+
5
+ Upstream Turso offers async I/O and concurrent writes in its local engine, plus sync through a separate SDK. Sense currently opens one local cache through `@tursodatabase/database` and serializes Sense commands against that file. It does not configure sync, remote access, or encryption. Read [Turso's current SDK guide](https://docs.turso.tech/sdk/introduction) to distinguish engine capabilities from the adapter features Sense exposes.
6
+
7
+ ## What the choice gives you
8
+
9
+ - A `.sense/cache.turso.db` file using the embedded Turso engine.
10
+ - Tantivy-backed lexical ranking with shared stemming and n-gram support for unspaced scripts.
11
+ - The asynchronous JavaScript database client used by sense's store adapter.
12
+ - Native vector storage and search.
13
+
14
+ ## Current Sense adapter boundaries
15
+
16
+ - Turso support is experimental.
17
+ - The first open installs the optional native package.
18
+ - Turso holds the cache file for the life of a connection. Concurrent sense commands wait for the current command or watcher cycle to close it.
19
+ - `sense search` accepts bare words and quoted phrases. It rejects FTS5 boolean, prefix, `NEAR`, initial-token, grouping, and column-filter syntax.
20
+ - SQLite's raw `MATCH`, `snippet()`, and `bm25()` queries do not port.
21
+ - `has()` and `basename()` work through SQL rewriting. `segment()` is unavailable.
22
+
23
+ Commands, public tables, vectors, and `sense watch` remain available. Read the [Turso query guide](../../../sense/references/stores/turso.md) before translating saved SQL or advanced searches.
24
+
25
+ Changing to Turso builds its separate cache. It does not migrate the SQLite or DuckDB cache. Review the shipped [store benchmark summary](../store-benchmarks.md) if performance affects the choice.