sensemaking 0.18.4 → 0.19.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +22 -4
- package/dist/cjs/chunk/group.js +10 -23
- package/dist/cjs/chunk/group.js.map +1 -1
- package/dist/cjs/chunk/parse.js +2 -5
- package/dist/cjs/chunk/parse.js.map +1 -1
- package/dist/cjs/cli/index.js.map +1 -1
- package/dist/cjs/cli/shared.js +6 -10
- package/dist/cjs/cli/shared.js.map +1 -1
- package/dist/cjs/cli/status.js +44 -10
- package/dist/cjs/cli/status.js.map +1 -1
- package/dist/cjs/cli.js +2 -8
- package/dist/cjs/cli.js.map +1 -1
- package/dist/cjs/commands/map.js +1 -3
- package/dist/cjs/commands/map.js.map +1 -1
- package/dist/cjs/commands/peek.js.map +1 -1
- package/dist/cjs/commands/related.js +3 -5
- package/dist/cjs/commands/related.js.map +1 -1
- package/dist/cjs/commands/scope.js.map +1 -1
- package/dist/cjs/commands/search.js +6 -8
- package/dist/cjs/commands/search.js.map +1 -1
- package/dist/cjs/commands/signals.js.map +1 -1
- package/dist/cjs/config/access.js.map +1 -1
- package/dist/cjs/config/index.d.cts +1 -1
- package/dist/cjs/config/index.d.ts +1 -1
- package/dist/cjs/config/index.js +3 -0
- package/dist/cjs/config/index.js.map +1 -1
- package/dist/cjs/config/load.js +10 -24
- package/dist/cjs/config/load.js.map +1 -1
- package/dist/cjs/config/resolve.js.map +1 -1
- package/dist/cjs/config/signals.js +2 -3
- package/dist/cjs/config/signals.js.map +1 -1
- package/dist/cjs/config/types.d.cts +2 -1
- package/dist/cjs/config/types.d.ts +2 -1
- package/dist/cjs/config/types.js +8 -0
- package/dist/cjs/config/types.js.map +1 -1
- package/dist/cjs/config/validate.d.cts +1 -1
- package/dist/cjs/config/validate.d.ts +1 -1
- package/dist/cjs/config/validate.js +9 -10
- package/dist/cjs/config/validate.js.map +1 -1
- package/dist/cjs/embed/cohere.js +2 -2
- package/dist/cjs/embed/cohere.js.map +1 -1
- package/dist/cjs/embed/distribution.js +3 -3
- package/dist/cjs/embed/distribution.js.map +1 -1
- package/dist/cjs/embed/handoff.d.cts +3 -0
- package/dist/cjs/embed/handoff.d.ts +3 -0
- package/dist/cjs/embed/handoff.js +39 -0
- package/dist/cjs/embed/handoff.js.map +1 -0
- package/dist/cjs/embed/identity.d.cts +11 -10
- package/dist/cjs/embed/identity.d.ts +11 -10
- package/dist/cjs/embed/identity.js +30 -15
- package/dist/cjs/embed/identity.js.map +1 -1
- package/dist/cjs/embed/langfit.js +2 -3
- package/dist/cjs/embed/langfit.js.map +1 -1
- package/dist/cjs/embed/languages.js +2 -4
- package/dist/cjs/embed/languages.js.map +1 -1
- package/dist/cjs/embed/openai.js +2 -2
- package/dist/cjs/embed/openai.js.map +1 -1
- package/dist/cjs/embed/query.js +58 -20
- package/dist/cjs/embed/query.js.map +1 -1
- package/dist/cjs/embed/registry.js +1 -2
- package/dist/cjs/embed/registry.js.map +1 -1
- package/dist/cjs/embed/static.d.cts +1 -1
- package/dist/cjs/embed/static.d.ts +1 -1
- package/dist/cjs/embed/static.js +8 -10
- package/dist/cjs/embed/static.js.map +1 -1
- package/dist/cjs/embed/store.d.cts +2 -2
- package/dist/cjs/embed/store.d.ts +2 -2
- package/dist/cjs/embed/store.js +48 -34
- package/dist/cjs/embed/store.js.map +1 -1
- package/dist/cjs/embed/types.js +1 -3
- package/dist/cjs/embed/types.js.map +1 -1
- package/dist/cjs/errors.d.cts +1 -1
- package/dist/cjs/errors.d.ts +1 -1
- package/dist/cjs/errors.js.map +1 -1
- package/dist/cjs/features/embed.js +19 -17
- package/dist/cjs/features/embed.js.map +1 -1
- package/dist/cjs/features/fences.js +5 -10
- package/dist/cjs/features/fences.js.map +1 -1
- package/dist/cjs/features/links.js +20 -37
- package/dist/cjs/features/links.js.map +1 -1
- package/dist/cjs/features/sections.js +21 -19
- package/dist/cjs/features/sections.js.map +1 -1
- package/dist/cjs/features/tags.js +5 -8
- package/dist/cjs/features/tags.js.map +1 -1
- package/dist/cjs/graph/traverse.js +1 -2
- package/dist/cjs/graph/traverse.js.map +1 -1
- package/dist/cjs/index.d.cts +1 -1
- package/dist/cjs/index.d.ts +1 -1
- package/dist/cjs/index.js +3 -0
- package/dist/cjs/index.js.map +1 -1
- package/dist/cjs/output/column-hint.js +3 -6
- package/dist/cjs/output/column-hint.js.map +1 -1
- package/dist/cjs/output/output.js +7 -14
- package/dist/cjs/output/output.js.map +1 -1
- package/dist/cjs/output/search-error.js +2 -5
- package/dist/cjs/output/search-error.js.map +1 -1
- package/dist/cjs/scan/frontmatter.d.cts +1 -0
- package/dist/cjs/scan/frontmatter.d.ts +1 -0
- package/dist/cjs/scan/frontmatter.js +14 -18
- package/dist/cjs/scan/frontmatter.js.map +1 -1
- package/dist/cjs/scan/list.js +2 -3
- package/dist/cjs/scan/list.js.map +1 -1
- package/dist/cjs/scan/reparse.js +12 -26
- package/dist/cjs/scan/reparse.js.map +1 -1
- package/dist/cjs/scan/worker-error.js.map +1 -1
- package/dist/cjs/store/cache.d.cts +2 -0
- package/dist/cjs/store/cache.d.ts +2 -0
- package/dist/cjs/store/cache.js +20 -0
- package/dist/cjs/store/cache.js.map +1 -0
- package/dist/cjs/store/duckdb/batch.js +4 -16
- package/dist/cjs/store/duckdb/batch.js.map +1 -1
- package/dist/cjs/store/duckdb/connection.js +8 -20
- package/dist/cjs/store/duckdb/connection.js.map +1 -1
- package/dist/cjs/store/duckdb/fieldStats.js +2 -8
- package/dist/cjs/store/duckdb/fieldStats.js.map +1 -1
- package/dist/cjs/store/duckdb/lexical.js +14 -28
- package/dist/cjs/store/duckdb/lexical.js.map +1 -1
- package/dist/cjs/store/duckdb/native.js.map +1 -1
- package/dist/cjs/store/duckdb/open.js +12 -13
- package/dist/cjs/store/duckdb/open.js.map +1 -1
- package/dist/cjs/store/duckdb/reconcile.js +4 -11
- package/dist/cjs/store/duckdb/reconcile.js.map +1 -1
- package/dist/cjs/store/duckdb/sql-functions.js +2 -4
- package/dist/cjs/store/duckdb/sql-functions.js.map +1 -1
- package/dist/cjs/store/duckdb/store.js +2 -3
- package/dist/cjs/store/duckdb/store.js.map +1 -1
- package/dist/cjs/store/duckdb/vectors.js +8 -18
- package/dist/cjs/store/duckdb/vectors.js.map +1 -1
- package/dist/cjs/store/index.d.cts +3 -2
- package/dist/cjs/store/index.d.ts +3 -2
- package/dist/cjs/store/index.js +7 -8
- package/dist/cjs/store/index.js.map +1 -1
- package/dist/cjs/store/native.js +6 -16
- package/dist/cjs/store/native.js.map +1 -1
- package/dist/cjs/store/sql-functions.js.map +1 -1
- package/dist/cjs/store/sqlite/connection.js +5 -9
- package/dist/cjs/store/sqlite/connection.js.map +1 -1
- package/dist/cjs/store/sqlite/fieldStats.js.map +1 -1
- package/dist/cjs/store/sqlite/lexical.js +3 -5
- package/dist/cjs/store/sqlite/lexical.js.map +1 -1
- package/dist/cjs/store/sqlite/open.d.cts +0 -1
- package/dist/cjs/store/sqlite/open.d.ts +0 -1
- package/dist/cjs/store/sqlite/open.js +47 -45
- package/dist/cjs/store/sqlite/open.js.map +1 -1
- package/dist/cjs/store/sqlite/reconcile.js +113 -104
- package/dist/cjs/store/sqlite/reconcile.js.map +1 -1
- package/dist/cjs/store/sqlite/sql-functions.js +2 -4
- package/dist/cjs/store/sqlite/sql-functions.js.map +1 -1
- package/dist/cjs/store/sqlite/store.js.map +1 -1
- package/dist/cjs/store/transaction.d.cts +2 -1
- package/dist/cjs/store/transaction.d.ts +2 -1
- package/dist/cjs/store/transaction.js +17 -9
- package/dist/cjs/store/transaction.js.map +1 -1
- package/dist/cjs/store/turso/connection.js +3 -7
- package/dist/cjs/store/turso/connection.js.map +1 -1
- package/dist/cjs/store/turso/fieldStats.js.map +1 -1
- package/dist/cjs/store/turso/lexical.d.cts +2 -0
- package/dist/cjs/store/turso/lexical.d.ts +2 -0
- package/dist/cjs/store/turso/lexical.js +347 -0
- package/dist/cjs/store/turso/lexical.js.map +1 -0
- package/dist/cjs/store/turso/native.js.map +1 -1
- package/dist/cjs/store/turso/open.d.cts +3 -1
- package/dist/cjs/store/turso/open.d.ts +3 -1
- package/dist/cjs/store/turso/open.js +150 -66
- package/dist/cjs/store/turso/open.js.map +1 -1
- package/dist/cjs/store/turso/reconcile.js +286 -152
- package/dist/cjs/store/turso/reconcile.js.map +1 -1
- package/dist/cjs/store/turso/store.js +18 -41
- package/dist/cjs/store/turso/store.js.map +1 -1
- package/dist/cjs/store/turso/vectors.d.cts +8 -0
- package/dist/cjs/store/turso/vectors.d.ts +8 -0
- package/dist/cjs/store/turso/vectors.js +435 -0
- package/dist/cjs/store/turso/vectors.js.map +1 -0
- package/dist/cjs/store/types.js +2 -3
- package/dist/cjs/store/types.js.map +1 -1
- package/dist/cjs/store/vectors.js.map +1 -1
- package/dist/cjs/text/segment.d.cts +1 -0
- package/dist/cjs/text/segment.d.ts +1 -0
- package/dist/cjs/text/segment.js +10 -14
- package/dist/cjs/text/segment.js.map +1 -1
- package/dist/cjs/watch.js +2 -5
- package/dist/cjs/watch.js.map +1 -1
- package/dist/cjs/workers/parse.js.map +1 -1
- package/dist/esm/chunk/group.js +10 -23
- package/dist/esm/chunk/group.js.map +1 -1
- package/dist/esm/chunk/parse.js +4 -8
- package/dist/esm/chunk/parse.js.map +1 -1
- package/dist/esm/cli/index.js +3 -5
- package/dist/esm/cli/index.js.map +1 -1
- package/dist/esm/cli/shared.js +12 -22
- package/dist/esm/cli/shared.js.map +1 -1
- package/dist/esm/cli/status.js +43 -9
- package/dist/esm/cli/status.js.map +1 -1
- package/dist/esm/cli.js +2 -8
- package/dist/esm/cli.js.map +1 -1
- package/dist/esm/commands/map.js +1 -3
- package/dist/esm/commands/map.js.map +1 -1
- package/dist/esm/commands/peek.js.map +1 -1
- package/dist/esm/commands/related.js +3 -5
- package/dist/esm/commands/related.js.map +1 -1
- package/dist/esm/commands/scope.js +3 -5
- package/dist/esm/commands/scope.js.map +1 -1
- package/dist/esm/commands/search.js +10 -16
- package/dist/esm/commands/search.js.map +1 -1
- package/dist/esm/commands/signals.js +3 -7
- package/dist/esm/commands/signals.js.map +1 -1
- package/dist/esm/config/access.js +8 -13
- package/dist/esm/config/access.js.map +1 -1
- package/dist/esm/config/index.d.ts +1 -1
- package/dist/esm/config/index.js +1 -1
- package/dist/esm/config/index.js.map +1 -1
- package/dist/esm/config/load.js +10 -24
- package/dist/esm/config/load.js.map +1 -1
- package/dist/esm/config/resolve.js +2 -3
- package/dist/esm/config/resolve.js.map +1 -1
- package/dist/esm/config/signals.js +3 -5
- package/dist/esm/config/signals.js.map +1 -1
- package/dist/esm/config/types.d.ts +2 -1
- package/dist/esm/config/types.js +8 -2
- package/dist/esm/config/types.js.map +1 -1
- package/dist/esm/config/validate.d.ts +1 -1
- package/dist/esm/config/validate.js +7 -10
- package/dist/esm/config/validate.js.map +1 -1
- package/dist/esm/embed/cohere.js +2 -2
- package/dist/esm/embed/cohere.js.map +1 -1
- package/dist/esm/embed/distribution.js +1 -1
- package/dist/esm/embed/distribution.js.map +1 -1
- package/dist/esm/embed/handoff.d.ts +3 -0
- package/dist/esm/embed/handoff.js +21 -0
- package/dist/esm/embed/handoff.js.map +1 -0
- package/dist/esm/embed/identity.d.ts +11 -10
- package/dist/esm/embed/identity.js +29 -26
- package/dist/esm/embed/identity.js.map +1 -1
- package/dist/esm/embed/langfit.js +2 -3
- package/dist/esm/embed/langfit.js.map +1 -1
- package/dist/esm/embed/languages.js +2 -4
- package/dist/esm/embed/languages.js.map +1 -1
- package/dist/esm/embed/openai.js +2 -2
- package/dist/esm/embed/openai.js.map +1 -1
- package/dist/esm/embed/query.js +21 -6
- package/dist/esm/embed/query.js.map +1 -1
- package/dist/esm/embed/registry.js +1 -2
- package/dist/esm/embed/registry.js.map +1 -1
- package/dist/esm/embed/static.d.ts +1 -1
- package/dist/esm/embed/static.js +8 -10
- package/dist/esm/embed/static.js.map +1 -1
- package/dist/esm/embed/store.d.ts +2 -2
- package/dist/esm/embed/store.js +28 -23
- package/dist/esm/embed/store.js.map +1 -1
- package/dist/esm/embed/types.js +1 -3
- package/dist/esm/embed/types.js.map +1 -1
- package/dist/esm/errors.d.ts +1 -1
- package/dist/esm/errors.js.map +1 -1
- package/dist/esm/features/embed.js +16 -16
- package/dist/esm/features/embed.js.map +1 -1
- package/dist/esm/features/fences.js +9 -19
- package/dist/esm/features/fences.js.map +1 -1
- package/dist/esm/features/links.js +24 -49
- package/dist/esm/features/links.js.map +1 -1
- package/dist/esm/features/sections.js +20 -20
- package/dist/esm/features/sections.js.map +1 -1
- package/dist/esm/features/tags.js +5 -8
- package/dist/esm/features/tags.js.map +1 -1
- package/dist/esm/features/types.js.map +1 -1
- package/dist/esm/graph/traverse.js +1 -2
- package/dist/esm/graph/traverse.js.map +1 -1
- package/dist/esm/index.d.ts +1 -1
- package/dist/esm/index.js +1 -1
- package/dist/esm/index.js.map +1 -1
- package/dist/esm/output/column-hint.js +3 -6
- package/dist/esm/output/column-hint.js.map +1 -1
- package/dist/esm/output/output.js +9 -20
- package/dist/esm/output/output.js.map +1 -1
- package/dist/esm/output/search-error.js +2 -5
- package/dist/esm/output/search-error.js.map +1 -1
- package/dist/esm/scan/frontmatter.d.ts +1 -0
- package/dist/esm/scan/frontmatter.js +13 -18
- package/dist/esm/scan/frontmatter.js.map +1 -1
- package/dist/esm/scan/list.js +2 -3
- package/dist/esm/scan/list.js.map +1 -1
- package/dist/esm/scan/reparse.js +16 -34
- package/dist/esm/scan/reparse.js.map +1 -1
- package/dist/esm/scan/worker-error.js.map +1 -1
- package/dist/esm/store/cache.d.ts +2 -0
- package/dist/esm/store/cache.js +11 -0
- package/dist/esm/store/cache.js.map +1 -0
- package/dist/esm/store/duckdb/batch.js +4 -16
- package/dist/esm/store/duckdb/batch.js.map +1 -1
- package/dist/esm/store/duckdb/connection.js +8 -20
- package/dist/esm/store/duckdb/connection.js.map +1 -1
- package/dist/esm/store/duckdb/fieldStats.js +4 -12
- package/dist/esm/store/duckdb/fieldStats.js.map +1 -1
- package/dist/esm/store/duckdb/lexical.js +19 -35
- package/dist/esm/store/duckdb/lexical.js.map +1 -1
- package/dist/esm/store/duckdb/native.js +2 -7
- package/dist/esm/store/duckdb/native.js.map +1 -1
- package/dist/esm/store/duckdb/open.js +16 -23
- package/dist/esm/store/duckdb/open.js.map +1 -1
- package/dist/esm/store/duckdb/reconcile.js +4 -11
- package/dist/esm/store/duckdb/reconcile.js.map +1 -1
- package/dist/esm/store/duckdb/sql-functions.js +4 -9
- package/dist/esm/store/duckdb/sql-functions.js.map +1 -1
- package/dist/esm/store/duckdb/store.js +6 -14
- package/dist/esm/store/duckdb/store.js.map +1 -1
- package/dist/esm/store/duckdb/vectors.js +14 -31
- package/dist/esm/store/duckdb/vectors.js.map +1 -1
- package/dist/esm/store/index.d.ts +3 -2
- package/dist/esm/store/index.js +5 -6
- package/dist/esm/store/index.js.map +1 -1
- package/dist/esm/store/native.js +10 -25
- package/dist/esm/store/native.js.map +1 -1
- package/dist/esm/store/sql-functions.js +3 -5
- package/dist/esm/store/sql-functions.js.map +1 -1
- package/dist/esm/store/sqlite/connection.js +6 -10
- package/dist/esm/store/sqlite/connection.js.map +1 -1
- package/dist/esm/store/sqlite/fieldStats.js +2 -6
- package/dist/esm/store/sqlite/fieldStats.js.map +1 -1
- package/dist/esm/store/sqlite/lexical.js +5 -9
- package/dist/esm/store/sqlite/lexical.js.map +1 -1
- package/dist/esm/store/sqlite/open.d.ts +0 -1
- package/dist/esm/store/sqlite/open.js +45 -47
- package/dist/esm/store/sqlite/open.js.map +1 -1
- package/dist/esm/store/sqlite/reconcile.js +13 -17
- package/dist/esm/store/sqlite/reconcile.js.map +1 -1
- package/dist/esm/store/sqlite/sql-functions.js +2 -4
- package/dist/esm/store/sqlite/sql-functions.js.map +1 -1
- package/dist/esm/store/sqlite/store.js +1 -2
- package/dist/esm/store/sqlite/store.js.map +1 -1
- package/dist/esm/store/transaction.d.ts +2 -1
- package/dist/esm/store/transaction.js +12 -10
- package/dist/esm/store/transaction.js.map +1 -1
- package/dist/esm/store/turso/connection.js +4 -8
- package/dist/esm/store/turso/connection.js.map +1 -1
- package/dist/esm/store/turso/fieldStats.js +2 -5
- package/dist/esm/store/turso/fieldStats.js.map +1 -1
- package/dist/esm/store/turso/lexical.d.ts +2 -0
- package/dist/esm/store/turso/lexical.js +128 -0
- package/dist/esm/store/turso/lexical.js.map +1 -0
- package/dist/esm/store/turso/native.js +2 -7
- package/dist/esm/store/turso/native.js.map +1 -1
- package/dist/esm/store/turso/open.d.ts +3 -1
- package/dist/esm/store/turso/open.js +45 -29
- package/dist/esm/store/turso/open.js.map +1 -1
- package/dist/esm/store/turso/reconcile.js +38 -24
- package/dist/esm/store/turso/reconcile.js.map +1 -1
- package/dist/esm/store/turso/store.js +17 -28
- package/dist/esm/store/turso/store.js.map +1 -1
- package/dist/esm/store/turso/vectors.d.ts +8 -0
- package/dist/esm/store/turso/vectors.js +98 -0
- package/dist/esm/store/turso/vectors.js.map +1 -0
- package/dist/esm/store/types.js +2 -3
- package/dist/esm/store/types.js.map +1 -1
- package/dist/esm/store/vectors.js +2 -3
- package/dist/esm/store/vectors.js.map +1 -1
- package/dist/esm/text/segment.d.ts +1 -0
- package/dist/esm/text/segment.js +13 -21
- package/dist/esm/text/segment.js.map +1 -1
- package/dist/esm/watch.js +2 -5
- package/dist/esm/watch.js.map +1 -1
- package/dist/esm/workers/parse.js.map +1 -1
- package/package.json +3 -1
- package/schema.json +2 -2
- package/skills/sense/SKILL.md +2 -2
- package/skills/sense-setup/SKILL.md +3 -3
- package/dist/cjs/store/meta.d.cts +0 -3
- package/dist/cjs/store/meta.d.ts +0 -3
- package/dist/cjs/store/meta.js +0 -215
- package/dist/cjs/store/meta.js.map +0 -1
- package/dist/esm/store/meta.d.ts +0 -3
- package/dist/esm/store/meta.js +0 -17
- package/dist/esm/store/meta.js.map +0 -1
|
@@ -1,21 +1,18 @@
|
|
|
1
|
-
import {
|
|
1
|
+
import { STORE_DIMS } from '../../embed/types.js';
|
|
2
2
|
import { getColumns } from '../shared.js';
|
|
3
3
|
import { withTransaction } from '../transaction.js';
|
|
4
4
|
import { hasVectorRow, pendingRows } from '../vectors.js';
|
|
5
5
|
import { fieldStats } from './fieldStats.js';
|
|
6
|
+
import { queryLexical } from './lexical.js';
|
|
6
7
|
import { reconcile } from './reconcile.js';
|
|
7
|
-
|
|
8
|
-
//
|
|
9
|
-
//
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
export const CAPABILITIES = new Set();
|
|
16
|
-
function missing(name) {
|
|
17
|
-
throw new SenseError('STORE_CAPABILITY_MISSING', `store "turso" does not implement "${name}" in this build; set "store" to "sqlite" or "duckdb" in this tree's config`);
|
|
18
|
-
}
|
|
8
|
+
import { scanCandidates, scanSimilar, writeVectorBatch } from './vectors.js';
|
|
9
|
+
// No 'snippets': fts_highlight returns the whole column, not a bounded window, so hits use
|
|
10
|
+
// the caller's JS excerpt. No 'watch-concurrency': single-process by default, as on duckdb.
|
|
11
|
+
export const CAPABILITIES = new Set([
|
|
12
|
+
'lexical',
|
|
13
|
+
'phrases',
|
|
14
|
+
'vectors'
|
|
15
|
+
]);
|
|
19
16
|
// Shares one Connection instance (conn) with open()'s own reconcile call so transaction depth
|
|
20
17
|
// (see transaction.ts) is tracked against the same object everywhere.
|
|
21
18
|
export function createStore(db, conn, cfg, baseDir) {
|
|
@@ -46,28 +43,20 @@ export function createStore(db, conn, cfg, baseDir) {
|
|
|
46
43
|
fieldStats: (columns, scopeWhere)=>fieldStats(conn, columns, scopeWhere)
|
|
47
44
|
},
|
|
48
45
|
lexical: {
|
|
49
|
-
|
|
50
|
-
return missing('lexical');
|
|
51
|
-
}
|
|
46
|
+
query: (terms, opts)=>queryLexical(conn, terms, opts)
|
|
52
47
|
},
|
|
48
|
+
// The column's fixed DDL width (STORE_DIMS) is what every scan binds against, not the
|
|
49
|
+
// interface's per-call storeDims -- see vectors.ts's padded() for why a shorter vector is still correct against a wider column.
|
|
53
50
|
vectors: {
|
|
54
51
|
pending: ()=>pendingRows(conn),
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
async candidates () {
|
|
59
|
-
return missing('vectors');
|
|
60
|
-
},
|
|
61
|
-
async similar () {
|
|
62
|
-
return missing('vectors');
|
|
63
|
-
},
|
|
52
|
+
writeVectors: (rows)=>writeVectorBatch(conn, STORE_DIMS, rows),
|
|
53
|
+
candidates: (qv, _storeDims, fetch, allowed)=>scanCandidates(conn, qv, STORE_DIMS, fetch, allowed),
|
|
54
|
+
similar: (path, opts)=>scanSimilar(conn, STORE_DIMS, path, opts),
|
|
64
55
|
hasVector: (path)=>hasVectorRow(conn, path)
|
|
65
56
|
},
|
|
66
57
|
async engineStatus () {
|
|
67
58
|
// Read back rather than recomputed: this is what open() actually set (3x the largest
|
|
68
|
-
// recorded reconcile, floored at 30s, capped at 10min), not
|
|
69
|
-
// Turso's raw PRAGMA busy_timeout names its result column "busy_timeout" (spike-verified),
|
|
70
|
-
// not "timeout" like real SQLite's own quirky column name for this one pragma.
|
|
59
|
+
// recorded reconcile, floored at 30s, capped at 10min). Turso's PRAGMA busy_timeout names its column "busy_timeout" (spike-verified), not "timeout" like real SQLite.
|
|
71
60
|
const row = await (await db.prepare('PRAGMA busy_timeout')).get();
|
|
72
61
|
return {
|
|
73
62
|
busy_timeout: `${row.busy_timeout}ms (derived: 3x the largest reconcile this cache has recorded, floored at 30000ms)`
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/store/turso/store.ts"],"sourcesContent":["import type { Database } from '@tursodatabase/database';\nimport type { Config } from '../../config/index.ts';\nimport {
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/store/turso/store.ts"],"sourcesContent":["import type { Database } from '@tursodatabase/database';\nimport type { Config } from '../../config/index.ts';\nimport { STORE_DIMS } from '../../embed/types.ts';\nimport { getColumns } from '../shared.ts';\nimport { withTransaction } from '../transaction.ts';\nimport type { Capability, Connection, Statement, Store } from '../types.ts';\nimport { hasVectorRow, pendingRows } from '../vectors.ts';\nimport { fieldStats } from './fieldStats.ts';\nimport { queryLexical } from './lexical.ts';\nimport { reconcile } from './reconcile.ts';\nimport { scanCandidates, scanSimilar, writeVectorBatch } from './vectors.ts';\n\n// No 'snippets': fts_highlight returns the whole column, not a bounded window, so hits use\n// the caller's JS excerpt. No 'watch-concurrency': single-process by default, as on duckdb.\nexport const CAPABILITIES: ReadonlySet<Capability> = new Set(['lexical', 'phrases', 'vectors']);\n\n// Shares one Connection instance (conn) with open()'s own reconcile call so transaction depth\n// (see transaction.ts) is tracked against the same object everywhere.\nexport function createStore(db: Database, conn: Connection, cfg: Config, baseDir: string): Store {\n return {\n name: 'turso',\n capabilities: CAPABILITIES,\n async exec(sql: string): Promise<void> {\n await conn.exec(sql);\n },\n async prepare(sql: string): Promise<Statement> {\n return conn.prepare(sql);\n },\n async runBatch(sql: string, paramRows: unknown[][]): Promise<void> {\n await conn.runBatch(sql, paramRows);\n },\n async transaction<T>(fn: () => Promise<T>): Promise<T> {\n return withTransaction(conn, fn);\n },\n async reconcile() {\n return reconcile(conn, cfg, baseDir);\n },\n docs: {\n async columns() {\n return [...(await getColumns(conn))];\n },\n fieldStats: (columns, scopeWhere) => fieldStats(conn, columns, scopeWhere),\n },\n lexical: {\n query: (terms, opts) => queryLexical(conn, terms, opts),\n },\n // The column's fixed DDL width (STORE_DIMS) is what every scan binds against, not the\n // interface's per-call storeDims -- see vectors.ts's padded() for why a shorter vector is still correct against a wider column.\n vectors: {\n pending: () => pendingRows(conn),\n writeVectors: (rows) => writeVectorBatch(conn, STORE_DIMS, rows),\n candidates: (qv, _storeDims, fetch, allowed) => scanCandidates(conn, qv, STORE_DIMS, fetch, allowed),\n similar: (path, opts) => scanSimilar(conn, STORE_DIMS, path, opts),\n hasVector: (path) => hasVectorRow(conn, path),\n },\n async engineStatus() {\n // Read back rather than recomputed: this is what open() actually set (3x the largest\n // recorded reconcile, floored at 30s, capped at 10min). Turso's PRAGMA busy_timeout names its column \"busy_timeout\" (spike-verified), not \"timeout\" like real SQLite.\n const row = (await (await db.prepare('PRAGMA busy_timeout')).get()) as { busy_timeout: number };\n return { busy_timeout: `${row.busy_timeout}ms (derived: 3x the largest reconcile this cache has recorded, floored at 30000ms)` };\n },\n raw: {\n async prepare(sql: string) {\n const stmt = await db.prepare(sql);\n stmt.safeIntegers(true); // int64 past 2^53 arrives as BigInt instead of losing precision\n return {\n columns: () => stmt.columns(),\n // sense sql streams through the client's own async generator.\n iterate: async function* (...params: unknown[]) {\n yield* stmt.iterate(...params);\n },\n };\n },\n },\n async close() {\n await db.close();\n },\n };\n}\n"],"names":["STORE_DIMS","getColumns","withTransaction","hasVectorRow","pendingRows","fieldStats","queryLexical","reconcile","scanCandidates","scanSimilar","writeVectorBatch","CAPABILITIES","Set","createStore","db","conn","cfg","baseDir","name","capabilities","exec","sql","prepare","runBatch","paramRows","transaction","fn","docs","columns","scopeWhere","lexical","query","terms","opts","vectors","pending","writeVectors","rows","candidates","qv","_storeDims","fetch","allowed","similar","path","hasVector","engineStatus","row","get","busy_timeout","raw","stmt","safeIntegers","iterate","params","close"],"mappings":"AAEA,SAASA,UAAU,QAAQ,uBAAuB;AAClD,SAASC,UAAU,QAAQ,eAAe;AAC1C,SAASC,eAAe,QAAQ,oBAAoB;AAEpD,SAASC,YAAY,EAAEC,WAAW,QAAQ,gBAAgB;AAC1D,SAASC,UAAU,QAAQ,kBAAkB;AAC7C,SAASC,YAAY,QAAQ,eAAe;AAC5C,SAASC,SAAS,QAAQ,iBAAiB;AAC3C,SAASC,cAAc,EAAEC,WAAW,EAAEC,gBAAgB,QAAQ,eAAe;AAE7E,2FAA2F;AAC3F,4FAA4F;AAC5F,OAAO,MAAMC,eAAwC,IAAIC,IAAI;IAAC;IAAW;IAAW;CAAU,EAAE;AAEhG,8FAA8F;AAC9F,sEAAsE;AACtE,OAAO,SAASC,YAAYC,EAAY,EAAEC,IAAgB,EAAEC,GAAW,EAAEC,OAAe;IACtF,OAAO;QACLC,MAAM;QACNC,cAAcR;QACd,MAAMS,MAAKC,GAAW;YACpB,MAAMN,KAAKK,IAAI,CAACC;QAClB;QACA,MAAMC,SAAQD,GAAW;YACvB,OAAON,KAAKO,OAAO,CAACD;QACtB;QACA,MAAME,UAASF,GAAW,EAAEG,SAAsB;YAChD,MAAMT,KAAKQ,QAAQ,CAACF,KAAKG;QAC3B;QACA,MAAMC,aAAeC,EAAoB;YACvC,OAAOxB,gBAAgBa,MAAMW;QAC/B;QACA,MAAMnB;YACJ,OAAOA,UAAUQ,MAAMC,KAAKC;QAC9B;QACAU,MAAM;YACJ,MAAMC;gBACJ,OAAO;uBAAK,MAAM3B,WAAWc;iBAAO;YACtC;YACAV,YAAY,CAACuB,SAASC,aAAexB,WAAWU,MAAMa,SAASC;QACjE;QACAC,SAAS;YACPC,OAAO,CAACC,OAAOC,OAAS3B,aAAaS,MAAMiB,OAAOC;QACpD;QACA,sFAAsF;QACtF,gIAAgI;QAChIC,SAAS;YACPC,SAAS,IAAM/B,YAAYW;YAC3BqB,cAAc,CAACC,OAAS3B,iBAAiBK,MAAMf,YAAYqC;YAC3DC,YAAY,CAACC,IAAIC,YAAYC,OAAOC,UAAYlC,eAAeO,MAAMwB,IAAIvC,YAAYyC,OAAOC;YAC5FC,SAAS,CAACC,MAAMX,OAASxB,YAAYM,MAAMf,YAAY4C,MAAMX;YAC7DY,WAAW,CAACD,OAASzC,aAAaY,MAAM6B;QAC1C;QACA,MAAME;YACJ,qFAAqF;YACrF,sKAAsK;YACtK,MAAMC,MAAO,MAAM,AAAC,CAAA,MAAMjC,GAAGQ,OAAO,CAAC,sBAAqB,EAAG0B,GAAG;YAChE,OAAO;gBAAEC,cAAc,GAAGF,IAAIE,YAAY,CAAC,kFAAkF,CAAC;YAAC;QACjI;QACAC,KAAK;YACH,MAAM5B,SAAQD,GAAW;gBACvB,MAAM8B,OAAO,MAAMrC,GAAGQ,OAAO,CAACD;gBAC9B8B,KAAKC,YAAY,CAAC,OAAO,gEAAgE;gBACzF,OAAO;oBACLxB,SAAS,IAAMuB,KAAKvB,OAAO;oBAC3B,8DAA8D;oBAC9DyB,SAAS,gBAAiB,GAAGC,MAAiB;wBAC5C,OAAOH,KAAKE,OAAO,IAAIC;oBACzB;gBACF;YACF;QACF;QACA,MAAMC;YACJ,MAAMzC,GAAGyC,KAAK;QAChB;IACF;AACF"}
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
import type { Connection, VectorCandidate, VectorSimilar, VectorWriteRow } from '../types.js';
|
|
2
|
+
export declare function writeVectorBatch(conn: Connection, dims: number, rows: VectorWriteRow[]): Promise<void>;
|
|
3
|
+
export declare function scanCandidates(conn: Connection, qv: Float32Array, dims: number, fetch: number, allowed?: Set<string>): Promise<VectorCandidate[]>;
|
|
4
|
+
export declare function scanSimilar(conn: Connection, dims: number, path: string, opts: {
|
|
5
|
+
exclude: Set<string>;
|
|
6
|
+
allowed?: Set<string>;
|
|
7
|
+
k: number;
|
|
8
|
+
}): Promise<VectorSimilar[]>;
|
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
import { asCosine, sampleEvenly } from '../vectors.js';
|
|
2
|
+
// Native F32_BLOB(dims) columns, dims fixed at DDL time (open.ts, STORE_DIMS). Both scans
|
|
3
|
+
// score in JS: a per-row vector_distance_cos measured 3.5-8x slower on this row engine.
|
|
4
|
+
// A model narrower than the column zero-pads: the added dimensions contribute nothing to
|
|
5
|
+
// either dot product or norm, so cosine is unchanged.
|
|
6
|
+
function padded(values, dims) {
|
|
7
|
+
const out = new Float32Array(dims);
|
|
8
|
+
for(let d = 0; d < Math.min(values.length, dims); d++)out[d] = values[d];
|
|
9
|
+
return out;
|
|
10
|
+
}
|
|
11
|
+
// A bound Float32Array stores truncated bytes with no error; a Buffer view over the same
|
|
12
|
+
// bytes round-trips exact (measured 0.7.2).
|
|
13
|
+
function toBlob(v) {
|
|
14
|
+
return Buffer.from(v.buffer, v.byteOffset, v.byteLength);
|
|
15
|
+
}
|
|
16
|
+
function decode(vector, dims) {
|
|
17
|
+
return new Float32Array(vector.buffer, vector.byteOffset, dims);
|
|
18
|
+
}
|
|
19
|
+
// Rows arrive int8-quantized with a per-vector scale (embed/query.ts's toStore); this store
|
|
20
|
+
// dequantizes to floats, as duckdb does.
|
|
21
|
+
function dequantize(row, dims) {
|
|
22
|
+
const q = new Int8Array(row.vector.buffer, row.vector.byteOffset, row.vector.byteLength);
|
|
23
|
+
const out = new Float32Array(dims);
|
|
24
|
+
for(let d = 0; d < Math.min(q.length, dims); d++)out[d] = q[d] * row.scale;
|
|
25
|
+
return out;
|
|
26
|
+
}
|
|
27
|
+
function cosineSimilarity(a, b) {
|
|
28
|
+
let dot = 0;
|
|
29
|
+
let na = 0;
|
|
30
|
+
let nb = 0;
|
|
31
|
+
for(let d = 0; d < a.length; d++){
|
|
32
|
+
dot += a[d] * b[d];
|
|
33
|
+
na += a[d] * a[d];
|
|
34
|
+
nb += b[d] * b[d];
|
|
35
|
+
}
|
|
36
|
+
return dot / (Math.sqrt(na) * Math.sqrt(nb) + 1e-32);
|
|
37
|
+
}
|
|
38
|
+
// `scale` stays NULL: the column exists so the row shape matches the other stores, and
|
|
39
|
+
// cosine is scale-invariant once the value is dequantized here.
|
|
40
|
+
export async function writeVectorBatch(conn, dims, rows) {
|
|
41
|
+
if (rows.length === 0) return;
|
|
42
|
+
await conn.runBatch('UPDATE embeddings SET vector = ? WHERE "path" = ? AND chunk = ?', rows.map((row)=>[
|
|
43
|
+
toBlob(dequantize(row, dims)),
|
|
44
|
+
row.path,
|
|
45
|
+
row.chunk
|
|
46
|
+
]));
|
|
47
|
+
}
|
|
48
|
+
// Best chunk per file by cosine. GROUP BY with MIN() cannot do this on 0.7.2: the aggregate
|
|
49
|
+
// is right but the bare start_line/end_line come from an arbitrary row in the group.
|
|
50
|
+
export async function scanCandidates(conn, qv, dims, fetch, allowed) {
|
|
51
|
+
const q = padded(qv, dims);
|
|
52
|
+
const stmt = await conn.prepare('SELECT "path", start_line, end_line, vector FROM embeddings WHERE vector IS NOT NULL');
|
|
53
|
+
const rows = await stmt.all();
|
|
54
|
+
const best = new Map();
|
|
55
|
+
for (const row of rows){
|
|
56
|
+
if (allowed && !allowed.has(row.path)) continue;
|
|
57
|
+
const score = cosineSimilarity(q, decode(row.vector, dims));
|
|
58
|
+
const existing = best.get(row.path);
|
|
59
|
+
if (!existing || score > existing.score) best.set(row.path, {
|
|
60
|
+
score,
|
|
61
|
+
lines: `L${row.start_line}-${row.end_line}`
|
|
62
|
+
});
|
|
63
|
+
}
|
|
64
|
+
return [
|
|
65
|
+
...best.entries()
|
|
66
|
+
].sort((a, b)=>b[1].score - a[1].score).slice(0, fetch).map(([path, b])=>({
|
|
67
|
+
path,
|
|
68
|
+
lines: b.lines,
|
|
69
|
+
similarity: asCosine(b.score)
|
|
70
|
+
}));
|
|
71
|
+
}
|
|
72
|
+
// Max cosine over (target chunk, other chunk) pairs. The SQL cross-join duckdb uses measured
|
|
73
|
+
// 6-8x slower here at 26k chunks (2026-08-30).
|
|
74
|
+
export async function scanSimilar(conn, dims, path, opts) {
|
|
75
|
+
if (opts.allowed && opts.allowed.size === 0) return [];
|
|
76
|
+
const targetStmt = await conn.prepare('SELECT vector FROM embeddings WHERE "path" = ? AND vector IS NOT NULL ORDER BY chunk');
|
|
77
|
+
const targetRows = await targetStmt.all(path);
|
|
78
|
+
if (targetRows.length === 0) return [];
|
|
79
|
+
const targets = sampleEvenly(targetRows).map((row)=>decode(row.vector, dims));
|
|
80
|
+
const stmt = await conn.prepare('SELECT "path", vector FROM embeddings WHERE vector IS NOT NULL');
|
|
81
|
+
const rows = await stmt.all();
|
|
82
|
+
const best = new Map();
|
|
83
|
+
for (const row of rows){
|
|
84
|
+
if (row.path === path || opts.exclude.has(row.path) || opts.allowed && !opts.allowed.has(row.path)) continue;
|
|
85
|
+
const other = decode(row.vector, dims);
|
|
86
|
+
for (const t of targets){
|
|
87
|
+
const score = cosineSimilarity(t, other);
|
|
88
|
+
const existing = best.get(row.path);
|
|
89
|
+
if (existing === undefined || score > existing) best.set(row.path, score);
|
|
90
|
+
}
|
|
91
|
+
}
|
|
92
|
+
return [
|
|
93
|
+
...best.entries()
|
|
94
|
+
].sort((a, b)=>b[1] - a[1]).slice(0, opts.k).map(([p, score])=>({
|
|
95
|
+
path: p,
|
|
96
|
+
similarity: asCosine(score)
|
|
97
|
+
}));
|
|
98
|
+
}
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/store/turso/vectors.ts"],"sourcesContent":["import type { Connection, VectorCandidate, VectorSimilar, VectorWriteRow } from '../types.ts';\nimport { asCosine, sampleEvenly } from '../vectors.ts';\n\n// Native F32_BLOB(dims) columns, dims fixed at DDL time (open.ts, STORE_DIMS). Both scans\n// score in JS: a per-row vector_distance_cos measured 3.5-8x slower on this row engine.\n\n// A model narrower than the column zero-pads: the added dimensions contribute nothing to\n// either dot product or norm, so cosine is unchanged.\nfunction padded(values: ArrayLike<number>, dims: number): Float32Array {\n const out = new Float32Array(dims);\n for (let d = 0; d < Math.min(values.length, dims); d++) out[d] = values[d];\n return out;\n}\n\n// A bound Float32Array stores truncated bytes with no error; a Buffer view over the same\n// bytes round-trips exact (measured 0.7.2).\nfunction toBlob(v: Float32Array): Buffer {\n return Buffer.from(v.buffer, v.byteOffset, v.byteLength);\n}\n\nfunction decode(vector: Buffer, dims: number): Float32Array {\n return new Float32Array(vector.buffer, vector.byteOffset, dims);\n}\n\n// Rows arrive int8-quantized with a per-vector scale (embed/query.ts's toStore); this store\n// dequantizes to floats, as duckdb does.\nfunction dequantize(row: VectorWriteRow, dims: number): Float32Array {\n const q = new Int8Array(row.vector.buffer, row.vector.byteOffset, row.vector.byteLength);\n const out = new Float32Array(dims);\n for (let d = 0; d < Math.min(q.length, dims); d++) out[d] = q[d] * row.scale;\n return out;\n}\n\nfunction cosineSimilarity(a: Float32Array, b: Float32Array): number {\n let dot = 0;\n let na = 0;\n let nb = 0;\n for (let d = 0; d < a.length; d++) {\n dot += a[d] * b[d];\n na += a[d] * a[d];\n nb += b[d] * b[d];\n }\n return dot / (Math.sqrt(na) * Math.sqrt(nb) + 1e-32);\n}\n\n// `scale` stays NULL: the column exists so the row shape matches the other stores, and\n// cosine is scale-invariant once the value is dequantized here.\nexport async function writeVectorBatch(conn: Connection, dims: number, rows: VectorWriteRow[]): Promise<void> {\n if (rows.length === 0) return;\n await conn.runBatch(\n 'UPDATE embeddings SET vector = ? WHERE \"path\" = ? AND chunk = ?',\n rows.map((row) => [toBlob(dequantize(row, dims)), row.path, row.chunk])\n );\n}\n\n// Best chunk per file by cosine. GROUP BY with MIN() cannot do this on 0.7.2: the aggregate\n// is right but the bare start_line/end_line come from an arbitrary row in the group.\nexport async function scanCandidates(conn: Connection, qv: Float32Array, dims: number, fetch: number, allowed?: Set<string>): Promise<VectorCandidate[]> {\n const q = padded(qv, dims);\n const stmt = await conn.prepare('SELECT \"path\", start_line, end_line, vector FROM embeddings WHERE vector IS NOT NULL');\n const rows = (await stmt.all()) as Array<{ path: string; start_line: number; end_line: number; vector: Buffer }>;\n\n const best = new Map<string, { score: number; lines: string }>();\n for (const row of rows) {\n if (allowed && !allowed.has(row.path)) continue;\n const score = cosineSimilarity(q, decode(row.vector, dims));\n const existing = best.get(row.path);\n if (!existing || score > existing.score) best.set(row.path, { score, lines: `L${row.start_line}-${row.end_line}` });\n }\n return [...best.entries()]\n .sort((a, b) => b[1].score - a[1].score)\n .slice(0, fetch)\n .map(([path, b]) => ({ path, lines: b.lines, similarity: asCosine(b.score) }));\n}\n\n// Max cosine over (target chunk, other chunk) pairs. The SQL cross-join duckdb uses measured\n// 6-8x slower here at 26k chunks (2026-08-30).\nexport async function scanSimilar(conn: Connection, dims: number, path: string, opts: { exclude: Set<string>; allowed?: Set<string>; k: number }): Promise<VectorSimilar[]> {\n if (opts.allowed && opts.allowed.size === 0) return [];\n const targetStmt = await conn.prepare('SELECT vector FROM embeddings WHERE \"path\" = ? AND vector IS NOT NULL ORDER BY chunk');\n const targetRows = (await targetStmt.all(path)) as Array<{ vector: Buffer }>;\n if (targetRows.length === 0) return [];\n const targets = sampleEvenly(targetRows).map((row) => decode(row.vector, dims));\n\n const stmt = await conn.prepare('SELECT \"path\", vector FROM embeddings WHERE vector IS NOT NULL');\n const rows = (await stmt.all()) as Array<{ path: string; vector: Buffer }>;\n const best = new Map<string, number>();\n for (const row of rows) {\n if (row.path === path || opts.exclude.has(row.path) || (opts.allowed && !opts.allowed.has(row.path))) continue;\n const other = decode(row.vector, dims);\n for (const t of targets) {\n const score = cosineSimilarity(t, other);\n const existing = best.get(row.path);\n if (existing === undefined || score > existing) best.set(row.path, score);\n }\n }\n\n return [...best.entries()]\n .sort((a, b) => b[1] - a[1])\n .slice(0, opts.k)\n .map(([p, score]) => ({ path: p, similarity: asCosine(score) }));\n}\n"],"names":["asCosine","sampleEvenly","padded","values","dims","out","Float32Array","d","Math","min","length","toBlob","v","Buffer","from","buffer","byteOffset","byteLength","decode","vector","dequantize","row","q","Int8Array","scale","cosineSimilarity","a","b","dot","na","nb","sqrt","writeVectorBatch","conn","rows","runBatch","map","path","chunk","scanCandidates","qv","fetch","allowed","stmt","prepare","all","best","Map","has","score","existing","get","set","lines","start_line","end_line","entries","sort","slice","similarity","scanSimilar","opts","size","targetStmt","targetRows","targets","exclude","other","t","undefined","k","p"],"mappings":"AACA,SAASA,QAAQ,EAAEC,YAAY,QAAQ,gBAAgB;AAEvD,0FAA0F;AAC1F,wFAAwF;AAExF,yFAAyF;AACzF,sDAAsD;AACtD,SAASC,OAAOC,MAAyB,EAAEC,IAAY;IACrD,MAAMC,MAAM,IAAIC,aAAaF;IAC7B,IAAK,IAAIG,IAAI,GAAGA,IAAIC,KAAKC,GAAG,CAACN,OAAOO,MAAM,EAAEN,OAAOG,IAAKF,GAAG,CAACE,EAAE,GAAGJ,MAAM,CAACI,EAAE;IAC1E,OAAOF;AACT;AAEA,yFAAyF;AACzF,4CAA4C;AAC5C,SAASM,OAAOC,CAAe;IAC7B,OAAOC,OAAOC,IAAI,CAACF,EAAEG,MAAM,EAAEH,EAAEI,UAAU,EAAEJ,EAAEK,UAAU;AACzD;AAEA,SAASC,OAAOC,MAAc,EAAEf,IAAY;IAC1C,OAAO,IAAIE,aAAaa,OAAOJ,MAAM,EAAEI,OAAOH,UAAU,EAAEZ;AAC5D;AAEA,4FAA4F;AAC5F,yCAAyC;AACzC,SAASgB,WAAWC,GAAmB,EAAEjB,IAAY;IACnD,MAAMkB,IAAI,IAAIC,UAAUF,IAAIF,MAAM,CAACJ,MAAM,EAAEM,IAAIF,MAAM,CAACH,UAAU,EAAEK,IAAIF,MAAM,CAACF,UAAU;IACvF,MAAMZ,MAAM,IAAIC,aAAaF;IAC7B,IAAK,IAAIG,IAAI,GAAGA,IAAIC,KAAKC,GAAG,CAACa,EAAEZ,MAAM,EAAEN,OAAOG,IAAKF,GAAG,CAACE,EAAE,GAAGe,CAAC,CAACf,EAAE,GAAGc,IAAIG,KAAK;IAC5E,OAAOnB;AACT;AAEA,SAASoB,iBAAiBC,CAAe,EAAEC,CAAe;IACxD,IAAIC,MAAM;IACV,IAAIC,KAAK;IACT,IAAIC,KAAK;IACT,IAAK,IAAIvB,IAAI,GAAGA,IAAImB,EAAEhB,MAAM,EAAEH,IAAK;QACjCqB,OAAOF,CAAC,CAACnB,EAAE,GAAGoB,CAAC,CAACpB,EAAE;QAClBsB,MAAMH,CAAC,CAACnB,EAAE,GAAGmB,CAAC,CAACnB,EAAE;QACjBuB,MAAMH,CAAC,CAACpB,EAAE,GAAGoB,CAAC,CAACpB,EAAE;IACnB;IACA,OAAOqB,MAAOpB,CAAAA,KAAKuB,IAAI,CAACF,MAAMrB,KAAKuB,IAAI,CAACD,MAAM,KAAI;AACpD;AAEA,uFAAuF;AACvF,gEAAgE;AAChE,OAAO,eAAeE,iBAAiBC,IAAgB,EAAE7B,IAAY,EAAE8B,IAAsB;IAC3F,IAAIA,KAAKxB,MAAM,KAAK,GAAG;IACvB,MAAMuB,KAAKE,QAAQ,CACjB,mEACAD,KAAKE,GAAG,CAAC,CAACf,MAAQ;YAACV,OAAOS,WAAWC,KAAKjB;YAAQiB,IAAIgB,IAAI;YAAEhB,IAAIiB,KAAK;SAAC;AAE1E;AAEA,4FAA4F;AAC5F,qFAAqF;AACrF,OAAO,eAAeC,eAAeN,IAAgB,EAAEO,EAAgB,EAAEpC,IAAY,EAAEqC,KAAa,EAAEC,OAAqB;IACzH,MAAMpB,IAAIpB,OAAOsC,IAAIpC;IACrB,MAAMuC,OAAO,MAAMV,KAAKW,OAAO,CAAC;IAChC,MAAMV,OAAQ,MAAMS,KAAKE,GAAG;IAE5B,MAAMC,OAAO,IAAIC;IACjB,KAAK,MAAM1B,OAAOa,KAAM;QACtB,IAAIQ,WAAW,CAACA,QAAQM,GAAG,CAAC3B,IAAIgB,IAAI,GAAG;QACvC,MAAMY,QAAQxB,iBAAiBH,GAAGJ,OAAOG,IAAIF,MAAM,EAAEf;QACrD,MAAM8C,WAAWJ,KAAKK,GAAG,CAAC9B,IAAIgB,IAAI;QAClC,IAAI,CAACa,YAAYD,QAAQC,SAASD,KAAK,EAAEH,KAAKM,GAAG,CAAC/B,IAAIgB,IAAI,EAAE;YAAEY;YAAOI,OAAO,CAAC,CAAC,EAAEhC,IAAIiC,UAAU,CAAC,CAAC,EAAEjC,IAAIkC,QAAQ,EAAE;QAAC;IACnH;IACA,OAAO;WAAIT,KAAKU,OAAO;KAAG,CACvBC,IAAI,CAAC,CAAC/B,GAAGC,IAAMA,CAAC,CAAC,EAAE,CAACsB,KAAK,GAAGvB,CAAC,CAAC,EAAE,CAACuB,KAAK,EACtCS,KAAK,CAAC,GAAGjB,OACTL,GAAG,CAAC,CAAC,CAACC,MAAMV,EAAE,GAAM,CAAA;YAAEU;YAAMgB,OAAO1B,EAAE0B,KAAK;YAAEM,YAAY3D,SAAS2B,EAAEsB,KAAK;QAAE,CAAA;AAC/E;AAEA,6FAA6F;AAC7F,+CAA+C;AAC/C,OAAO,eAAeW,YAAY3B,IAAgB,EAAE7B,IAAY,EAAEiC,IAAY,EAAEwB,IAAgE;IAC9I,IAAIA,KAAKnB,OAAO,IAAImB,KAAKnB,OAAO,CAACoB,IAAI,KAAK,GAAG,OAAO,EAAE;IACtD,MAAMC,aAAa,MAAM9B,KAAKW,OAAO,CAAC;IACtC,MAAMoB,aAAc,MAAMD,WAAWlB,GAAG,CAACR;IACzC,IAAI2B,WAAWtD,MAAM,KAAK,GAAG,OAAO,EAAE;IACtC,MAAMuD,UAAUhE,aAAa+D,YAAY5B,GAAG,CAAC,CAACf,MAAQH,OAAOG,IAAIF,MAAM,EAAEf;IAEzE,MAAMuC,OAAO,MAAMV,KAAKW,OAAO,CAAC;IAChC,MAAMV,OAAQ,MAAMS,KAAKE,GAAG;IAC5B,MAAMC,OAAO,IAAIC;IACjB,KAAK,MAAM1B,OAAOa,KAAM;QACtB,IAAIb,IAAIgB,IAAI,KAAKA,QAAQwB,KAAKK,OAAO,CAAClB,GAAG,CAAC3B,IAAIgB,IAAI,KAAMwB,KAAKnB,OAAO,IAAI,CAACmB,KAAKnB,OAAO,CAACM,GAAG,CAAC3B,IAAIgB,IAAI,GAAI;QACtG,MAAM8B,QAAQjD,OAAOG,IAAIF,MAAM,EAAEf;QACjC,KAAK,MAAMgE,KAAKH,QAAS;YACvB,MAAMhB,QAAQxB,iBAAiB2C,GAAGD;YAClC,MAAMjB,WAAWJ,KAAKK,GAAG,CAAC9B,IAAIgB,IAAI;YAClC,IAAIa,aAAamB,aAAapB,QAAQC,UAAUJ,KAAKM,GAAG,CAAC/B,IAAIgB,IAAI,EAAEY;QACrE;IACF;IAEA,OAAO;WAAIH,KAAKU,OAAO;KAAG,CACvBC,IAAI,CAAC,CAAC/B,GAAGC,IAAMA,CAAC,CAAC,EAAE,GAAGD,CAAC,CAAC,EAAE,EAC1BgC,KAAK,CAAC,GAAGG,KAAKS,CAAC,EACflC,GAAG,CAAC,CAAC,CAACmC,GAAGtB,MAAM,GAAM,CAAA;YAAEZ,MAAMkC;YAAGZ,YAAY3D,SAASiD;QAAO,CAAA;AACjE"}
|
package/dist/esm/store/types.js
CHANGED
|
@@ -1,4 +1,3 @@
|
|
|
1
|
-
// The backing-store interface: a minimal portable statement surface (exec/prepare,
|
|
2
|
-
//
|
|
3
|
-
// diverge (lexical index, vector scan, raw sql passthrough).
|
|
1
|
+
// The backing-store interface: a minimal portable statement surface (exec/prepare), plus
|
|
2
|
+
// dedicated interfaces exactly where engines diverge (lexical index, vector scan, raw sql).
|
|
4
3
|
export { };
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/store/types.ts"],"sourcesContent":["// The backing-store interface: a minimal portable statement surface (exec/prepare,
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/store/types.ts"],"sourcesContent":["// The backing-store interface: a minimal portable statement surface (exec/prepare), plus\n// dedicated interfaces exactly where engines diverge (lexical index, vector scan, raw sql).\n\nimport type { StoreName } from '../config/index.ts';\n\n// 'lexical'/'vectors': the store's LexicalIndex/VectorStore is functionally implemented, not\n// present-but-inert. The rest are finer FTS5-only behaviors; a missing one fails at open or first use.\nexport type Capability = 'phrases' | 'snippets' | 'watch-concurrency' | 'lexical' | 'vectors';\n\nexport interface RunResult {\n changes: number | bigint;\n lastInsertRowid: number | bigint;\n}\n\n// A prepared statement's async surface: every supported engine crosses a real async boundary\n// for a query (DuckDB has no synchronous client), so run/get/all return Promises.\nexport interface Statement {\n run(...params: unknown[]): Promise<RunResult>;\n get(...params: unknown[]): Promise<unknown>;\n all(...params: unknown[]): Promise<unknown[]>;\n iterate(...params: unknown[]): AsyncIterable<unknown>;\n columns(): Array<{ name: string }>;\n setReadBigInts(enabled: boolean): void;\n}\n\n// Connection surface feature-owned SQL runs against inside a store's hot loops (schema,\n// reconcile), portable so a feature never imports an engine-specific client type.\nexport interface Connection {\n exec(sql: string): Promise<void>;\n prepare(sql: string): Promise<Statement>;\n runBatch(sql: string, paramRows: unknown[][]): Promise<void>;\n}\n\nexport interface FieldStat {\n field: string;\n coverage: number;\n type: string;\n}\n\nexport interface DocumentStore {\n // Frontmatter column names (including internal ones; callers filter).\n columns(): Promise<string[]>;\n // Per-column coverage (non-null count) and observed type set, aggregated in one SQL query,\n // never one row per note. `scopeWhere` is a caller-built WHERE fragment, per LexicalQueryOptions.\n fieldStats(columns: string[], scopeWhere: string): Promise<FieldStat[]>;\n}\n\nexport interface LexicalHit {\n path: string;\n hit: string | null;\n}\n\nexport interface LexicalQueryOptions {\n whereJoin: string;\n whereCond: string;\n scopeCond: string;\n limit: number;\n}\n\nexport interface LexicalIndex {\n // Ranked word-match query with excerpt, scoped by the caller-built SQL fragments (the same\n // fragments narrowByWhere/materializeScope produce elsewhere).\n query(terms: string, opts: LexicalQueryOptions): Promise<LexicalHit[]>;\n}\n\nexport interface VectorCandidate {\n path: string;\n lines: string;\n similarity: number;\n}\n\nexport interface VectorSimilar {\n path: string;\n similarity: number;\n}\n\nexport interface VectorWriteRow {\n path: string;\n chunk: number;\n scale: number;\n vector: Buffer;\n}\n\nexport interface VectorStore {\n // Rows whose vector is NULL: never embedded, or added since.\n pending(): Promise<Array<{ path: string; chunk: number }>>;\n // One batch write per call (never per row) so a provider's embedding batch stays inside a\n // single store method.\n writeVectors(rows: VectorWriteRow[]): Promise<void>;\n candidates(queryVector: Float32Array, storeDims: number, fetch: number, allowed?: Set<string>): Promise<VectorCandidate[]>;\n similar(path: string, opts: { exclude: Set<string>; allowed?: Set<string>; k: number }): Promise<VectorSimilar[]>;\n hasVector(path: string): Promise<boolean>;\n}\n\n// The `sense sql` passthrough: string in, streamed rows out. Each store registers the same\n// sense-supplied functions and applies its own read-bigints/error-translation behavior.\nexport interface RawStatement {\n columns(): Array<{ name: string }>;\n iterate(...params: unknown[]): AsyncIterable<unknown>;\n}\n\nexport interface SqlSession {\n prepare(sql: string): Promise<RawStatement>;\n}\n\nexport interface Store {\n readonly name: StoreName;\n readonly capabilities: ReadonlySet<Capability>;\n exec(sql: string): Promise<void>;\n prepare(sql: string): Promise<Statement>;\n // One crossing, N rows: every external write loop (search's candidate insert, graph ring's\n // temp-table writes) goes through this instead of looping over `run()`. Same contract as Connection.runBatch.\n runBatch(sql: string, paramRows: unknown[][]): Promise<void>;\n // Pins one snapshot across multi-statement reads; `map` and `peek` use it. A network-bound\n // write (search's embedding top-up) must stay outside it. Nesting joins the enclosing transaction.\n transaction<T>(fn: () => Promise<T>): Promise<T>;\n // The whole file-sync pass (parse changed files, run every feature hook, resolve links,\n // recompute rank) behind one method: per-file iteration never crosses the async boundary.\n reconcile(): Promise<{ parsed: number; warnings: string[] }>;\n docs: DocumentStore;\n lexical: LexicalIndex;\n vectors: VectorStore;\n raw: SqlSession;\n // Engine-level facts for `sense status` (e.g. a derived busy_timeout PRAGMA reading); each\n // store owns what it reports and how it is worded. Empty when there is nothing to report.\n engineStatus(): Promise<Record<string, string>>;\n close(): Promise<void>;\n}\n"],"names":[],"mappings":"AAAA,yFAAyF;AACzF,4FAA4F;AAwG5F,WAsBC"}
|
|
@@ -7,9 +7,8 @@ export function sampleEvenly(rows, cap = TARGET_CHUNK_CAP) {
|
|
|
7
7
|
const step = Math.max(1, Math.ceil(rows.length / cap));
|
|
8
8
|
return rows.filter((_, i)=>i % step === 0);
|
|
9
9
|
}
|
|
10
|
-
// Dequantised int8 dot products (sqlite) land a little either side of a true cosine
|
|
11
|
-
//
|
|
12
|
-
// rounded the same way, so both stores print the same bounded number.
|
|
10
|
+
// Dequantised int8 dot products (sqlite) land a little either side of a true cosine (an identical
|
|
11
|
+
// pair can print 1.001); array_cosine_similarity (duckdb) has no such error but is rounded the same way, so both stores print the same bounded number.
|
|
13
12
|
export function asCosine(score) {
|
|
14
13
|
return Math.round(Math.min(1, Math.max(-1, score)) * 1000) / 1000;
|
|
15
14
|
}
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/store/vectors.ts"],"sourcesContent":["import type { Connection } from './types.ts';\n\n// Seed chunks that participate in a `related` scan. Cost is target_chunks x stored_chunks, so a\n// heading-dense seed multiplies a full-corpus scan (12.7s at 201 chunks/note unsampled).\nexport const TARGET_CHUNK_CAP = 16;\n\n// Evenly samples down to at most `cap` rows, so late sections of a long note still get a vote\n// instead of being cut off by a fixed prefix.\nexport function sampleEvenly<T>(rows: T[], cap: number = TARGET_CHUNK_CAP): T[] {\n const step = Math.max(1, Math.ceil(rows.length / cap));\n return rows.filter((_, i) => i % step === 0);\n}\n\n// Dequantised int8 dot products (sqlite) land a little either side of a true cosine
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/store/vectors.ts"],"sourcesContent":["import type { Connection } from './types.ts';\n\n// Seed chunks that participate in a `related` scan. Cost is target_chunks x stored_chunks, so a\n// heading-dense seed multiplies a full-corpus scan (12.7s at 201 chunks/note unsampled).\nexport const TARGET_CHUNK_CAP = 16;\n\n// Evenly samples down to at most `cap` rows, so late sections of a long note still get a vote\n// instead of being cut off by a fixed prefix.\nexport function sampleEvenly<T>(rows: T[], cap: number = TARGET_CHUNK_CAP): T[] {\n const step = Math.max(1, Math.ceil(rows.length / cap));\n return rows.filter((_, i) => i % step === 0);\n}\n\n// Dequantised int8 dot products (sqlite) land a little either side of a true cosine (an identical\n// pair can print 1.001); array_cosine_similarity (duckdb) has no such error but is rounded the same way, so both stores print the same bounded number.\nexport function asCosine(score: number): number {\n return Math.round(Math.min(1, Math.max(-1, score)) * 1000) / 1000;\n}\n\nexport async function pendingRows(conn: Connection): Promise<Array<{ path: string; chunk: number }>> {\n const stmt = await conn.prepare('SELECT \"path\", chunk FROM embeddings WHERE vector IS NULL ORDER BY \"path\", chunk');\n return (await stmt.all()) as Array<{ path: string; chunk: number }>;\n}\n\n// Whether a note has any chunk with a vector. Distinguishes \"nothing is near this note\" from\n// \"this note has no text\", which look the same in an empty result.\nexport async function hasVectorRow(conn: Connection, path: string): Promise<boolean> {\n const stmt = await conn.prepare('SELECT 1 AS ok FROM embeddings WHERE \"path\" = ? AND vector IS NOT NULL LIMIT 1');\n const row = (await stmt.get(path)) as { ok: number } | undefined;\n return row !== undefined;\n}\n"],"names":["TARGET_CHUNK_CAP","sampleEvenly","rows","cap","step","Math","max","ceil","length","filter","_","i","asCosine","score","round","min","pendingRows","conn","stmt","prepare","all","hasVectorRow","path","row","get","undefined"],"mappings":"AAEA,gGAAgG;AAChG,yFAAyF;AACzF,OAAO,MAAMA,mBAAmB,GAAG;AAEnC,8FAA8F;AAC9F,8CAA8C;AAC9C,OAAO,SAASC,aAAgBC,IAAS,EAAEC,MAAcH,gBAAgB;IACvE,MAAMI,OAAOC,KAAKC,GAAG,CAAC,GAAGD,KAAKE,IAAI,CAACL,KAAKM,MAAM,GAAGL;IACjD,OAAOD,KAAKO,MAAM,CAAC,CAACC,GAAGC,IAAMA,IAAIP,SAAS;AAC5C;AAEA,kGAAkG;AAClG,uJAAuJ;AACvJ,OAAO,SAASQ,SAASC,KAAa;IACpC,OAAOR,KAAKS,KAAK,CAACT,KAAKU,GAAG,CAAC,GAAGV,KAAKC,GAAG,CAAC,CAAC,GAAGO,UAAU,QAAQ;AAC/D;AAEA,OAAO,eAAeG,YAAYC,IAAgB;IAChD,MAAMC,OAAO,MAAMD,KAAKE,OAAO,CAAC;IAChC,OAAQ,MAAMD,KAAKE,GAAG;AACxB;AAEA,6FAA6F;AAC7F,mEAAmE;AACnE,OAAO,eAAeC,aAAaJ,IAAgB,EAAEK,IAAY;IAC/D,MAAMJ,OAAO,MAAMD,KAAKE,OAAO,CAAC;IAChC,MAAMI,MAAO,MAAML,KAAKM,GAAG,CAACF;IAC5B,OAAOC,QAAQE;AACjB"}
|
|
@@ -1,3 +1,4 @@
|
|
|
1
1
|
export declare const UNSPACED_SCRIPTS = "\\p{scx=Han}\\p{scx=Hiragana}\\p{scx=Katakana}\\p{scx=Thai}\\p{scx=Khmer}\\p{scx=Lao}\\p{scx=Myanmar}";
|
|
2
|
+
export declare function hasUnspacedRun(text: string): boolean;
|
|
2
3
|
export declare function segmentField(text: string): string;
|
|
3
4
|
export declare function segmentMatch(terms: string): string;
|
package/dist/esm/text/segment.js
CHANGED
|
@@ -1,12 +1,5 @@
|
|
|
1
|
-
// Word boundaries FTS5's tokenizers cannot find on their own.
|
|
2
|
-
//
|
|
3
|
-
// Segmentation is context-free: seg(query) appears inside seg(document) whenever the query is
|
|
4
|
-
// a substring of the document. That is the contract -- substring semantics, what `grep` and
|
|
5
|
-
// `LIKE '%..%'` give -- and it is what makes index and query agree on every input, always.
|
|
6
|
-
// Intl.Segmenter's word mode is context-dependent (`东京都政府` splits as `东 | 京都 | 政府`
|
|
7
|
-
// while the query `东京` splits as one word) and is rejected for that reason. Its grapheme
|
|
8
|
-
// mode has no such dependency -- UAX #29 cluster boundaries never look past adjacent
|
|
9
|
-
// characters -- so grapheme mode is used below; no dictionary, no word mode.
|
|
1
|
+
// Word boundaries FTS5's tokenizers cannot find on their own. Contract: seg(query) must appear in
|
|
2
|
+
// seg(document) for any substring query; word mode is context-dependent and breaks that, so grapheme mode (UAX #29) is used.
|
|
10
3
|
// Scripts written without word spaces: a closed set of writing systems.
|
|
11
4
|
// Script_Extensions, not Script: the katakana long-vowel mark ー is Script=Common.
|
|
12
5
|
export const UNSPACED_SCRIPTS = '\\p{scx=Han}\\p{scx=Hiragana}\\p{scx=Katakana}\\p{scx=Thai}\\p{scx=Khmer}\\p{scx=Lao}\\p{scx=Myanmar}';
|
|
@@ -14,6 +7,11 @@ export const UNSPACED_SCRIPTS = '\\p{scx=Han}\\p{scx=Hiragana}\\p{scx=Katakana}\
|
|
|
14
7
|
// Latin letter (decomposed é) never starts one.
|
|
15
8
|
const RUN = new RegExp(`((?:[${UNSPACED_SCRIPTS}]\\p{M}*)+)`, 'gu');
|
|
16
9
|
const HAS_RUN = new RegExp(`[${UNSPACED_SCRIPTS}]`, 'u');
|
|
10
|
+
// Whether text holds a script that marks no word boundaries, the predicate every store's
|
|
11
|
+
// sidecar population and query split turns on.
|
|
12
|
+
export function hasUnspacedRun(text) {
|
|
13
|
+
return HAS_RUN.test(text);
|
|
14
|
+
}
|
|
17
15
|
// Grapheme clusters, ECMA-402/UAX #29: base char plus its marks, ZWJ sequences, Hangul jamo.
|
|
18
16
|
// Built once (construction cost amortizes) and used for both index and query splitting.
|
|
19
17
|
const GRAPHEME_SEGMENTER = new Intl.Segmenter(undefined, {
|
|
@@ -40,9 +38,8 @@ function splitOnPunctuation(run) {
|
|
|
40
38
|
}
|
|
41
39
|
return groups.filter((g)=>g.length > 0);
|
|
42
40
|
}
|
|
43
|
-
// Index side: '' when the field has no unspaced-script run
|
|
44
|
-
//
|
|
45
|
-
// neighbors and from a punctuation split within itself.
|
|
41
|
+
// Index side: '' when the field has no unspaced-script run. Otherwise each run explodes into
|
|
42
|
+
// its graphemes, barrier-delimited from its neighbors and from a punctuation split within itself.
|
|
46
43
|
export function segmentField(text) {
|
|
47
44
|
if (!HAS_RUN.test(text)) return '';
|
|
48
45
|
const out = text.replace(RUN, (run)=>{
|
|
@@ -54,21 +51,16 @@ export function segmentField(text) {
|
|
|
54
51
|
// A `title:`/`summary:`/`text:` qualifier directly before a run that is about to become a
|
|
55
52
|
// quoted grapheme phrase, so the rewrite can retarget it at the matching `_seg` column.
|
|
56
53
|
const QUALIFIER = /(^|[\s(])(-?)(title|summary|text)\s*:\s*$/;
|
|
57
|
-
//
|
|
58
|
-
//
|
|
59
|
-
// as across a real gap (`数数` vs `数。数`) -- a false positive only the barriered `_seg` columns
|
|
60
|
-
// are safe from. FTS5's column-set filter, not parens+OR, so the group still composes under
|
|
61
|
-
// AND/OR/NOT/juxtaposition exactly like the single-column qualifier form below.
|
|
54
|
+
// Raw title/summary/text drop punctuation as unicode61's token separator, so `数数` and `数。数`
|
|
55
|
+
// falsely match there; only the barriered `_seg` columns are safe from it, hence this fallback target.
|
|
62
56
|
const SIDECAR_COLUMNS = '{title_seg summary_seg text_seg}:';
|
|
63
57
|
// A run's punctuation-free groups, each its own quoted phrase (bare token if one grapheme),
|
|
64
58
|
// space-joined -- the same split points segmentField barriers, so query and index agree.
|
|
65
59
|
function runQuery(run, columnPrefix) {
|
|
66
60
|
return splitOnPunctuation(run).map((g)=>`${columnPrefix}${g.length > 1 ? `"${g.join(' ')}"` : g[0]}`).join(' ');
|
|
67
61
|
}
|
|
68
|
-
// Query side: each unspaced run becomes phrases of its graphemes, matching
|
|
69
|
-
//
|
|
70
|
-
// three (SIDECAR_COLUMNS), since raw columns cannot express the contract. An author's own quoted
|
|
71
|
-
// phrase is their explicit escape hatch to FTS5's native syntax, and passes through byte-identical.
|
|
62
|
+
// Query side: each unspaced run becomes phrases of its graphemes, matching segmentField. A
|
|
63
|
+
// qualifier maps to its `_seg` column; unqualified maps to all three (SIDECAR_COLUMNS).
|
|
72
64
|
export function segmentMatch(terms) {
|
|
73
65
|
if (!HAS_RUN.test(terms)) return terms;
|
|
74
66
|
let out = '';
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/text/segment.ts"],"sourcesContent":["// Word boundaries FTS5's tokenizers cannot find on their own
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/text/segment.ts"],"sourcesContent":["// Word boundaries FTS5's tokenizers cannot find on their own. Contract: seg(query) must appear in\n// seg(document) for any substring query; word mode is context-dependent and breaks that, so grapheme mode (UAX #29) is used.\n\n// Scripts written without word spaces: a closed set of writing systems.\n// Script_Extensions, not Script: the katakana long-vowel mark ー is Script=Common.\nexport const UNSPACED_SCRIPTS = '\\\\p{scx=Han}\\\\p{scx=Hiragana}\\\\p{scx=Katakana}\\\\p{scx=Thai}\\\\p{scx=Khmer}\\\\p{scx=Lao}\\\\p{scx=Myanmar}';\n// A run is script BASE characters with their combining marks attached; a bare mark after a\n// Latin letter (decomposed é) never starts one.\nconst RUN = new RegExp(`((?:[${UNSPACED_SCRIPTS}]\\\\p{M}*)+)`, 'gu');\nconst HAS_RUN = new RegExp(`[${UNSPACED_SCRIPTS}]`, 'u');\n\n// Whether text holds a script that marks no word boundaries, the predicate every store's\n// sidecar population and query split turns on.\nexport function hasUnspacedRun(text: string): boolean {\n return HAS_RUN.test(text);\n}\n// Grapheme clusters, ECMA-402/UAX #29: base char plus its marks, ZWJ sequences, Hangul jamo.\n// Built once (construction cost amortizes) and used for both index and query splitting.\nconst GRAPHEME_SEGMENTER = new Intl.Segmenter(undefined, { granularity: 'grapheme' });\n// unicode61 drops Unicode punctuation as a separator. Some of it (。、「」) has Script_Extensions\n// into an unspaced script, so RUN keeps it -- a grapheme matching this becomes a split point.\nconst PUNCTUATION = /\\p{P}/u;\n// Token barrier between separate runs (or a punctuation split within one) so their graphemes\n// are never phrase-adjacent. U+A7F7: a letter (so FTS5 keeps it as a token) no one types.\nconst BARRIER = 'ꟷ';\n\nfunction graphemes(run: string): string[] {\n return Array.from(GRAPHEME_SEGMENTER.segment(run), (s) => s.segment);\n}\n\n// A run's graphemes, cut into punctuation-free groups at every punctuation grapheme (dropped,\n// matching unicode61). The one place punctuation is classified, so index and query agree.\nfunction splitOnPunctuation(run: string): string[][] {\n const groups: string[][] = [[]];\n for (const g of graphemes(run)) {\n if (PUNCTUATION.test(g)) groups.push([]);\n else groups[groups.length - 1].push(g);\n }\n return groups.filter((g) => g.length > 0);\n}\n\n// Index side: '' when the field has no unspaced-script run. Otherwise each run explodes into\n// its graphemes, barrier-delimited from its neighbors and from a punctuation split within itself.\nexport function segmentField(text: string): string {\n if (!HAS_RUN.test(text)) return '';\n const out = text.replace(RUN, (run) => {\n const body = splitOnPunctuation(run)\n .map((g) => g.join(' '))\n .join(` ${BARRIER} `);\n return ` ${BARRIER} ${body} ${BARRIER} `;\n });\n return out.replace(/\\s+/g, ' ').trim();\n}\n\n// A `title:`/`summary:`/`text:` qualifier directly before a run that is about to become a\n// quoted grapheme phrase, so the rewrite can retarget it at the matching `_seg` column.\nconst QUALIFIER = /(^|[\\s(])(-?)(title|summary|text)\\s*:\\s*$/;\n\n// Raw title/summary/text drop punctuation as unicode61's token separator, so `数数` and `数。数`\n// falsely match there; only the barriered `_seg` columns are safe from it, hence this fallback target.\nconst SIDECAR_COLUMNS = '{title_seg summary_seg text_seg}:';\n\n// A run's punctuation-free groups, each its own quoted phrase (bare token if one grapheme),\n// space-joined -- the same split points segmentField barriers, so query and index agree.\nfunction runQuery(run: string, columnPrefix: string): string {\n return splitOnPunctuation(run)\n .map((g) => `${columnPrefix}${g.length > 1 ? `\"${g.join(' ')}\"` : g[0]}`)\n .join(' ');\n}\n\n// Query side: each unspaced run becomes phrases of its graphemes, matching segmentField. A\n// qualifier maps to its `_seg` column; unqualified maps to all three (SIDECAR_COLUMNS).\nexport function segmentMatch(terms: string): string {\n if (!HAS_RUN.test(terms)) return terms;\n let out = '';\n let quoted = false;\n const pieces = terms.split(RUN); // split keeps captured runs at odd indices\n for (let i = 0; i < pieces.length; i++) {\n if (i % 2 === 0) {\n for (const ch of pieces[i]) if (ch === '\"') quoted = !quoted;\n out += pieces[i];\n continue;\n }\n if (quoted) {\n out += pieces[i]; // an author's phrase is matched as written\n continue;\n }\n const m = out.match(QUALIFIER);\n if (m) {\n out = `${out.slice(0, m.index)}${m[1]}${runQuery(pieces[i], `${m[2]}${m[3]}_seg:`)}`;\n } else {\n out += runQuery(pieces[i], SIDECAR_COLUMNS);\n }\n }\n return out;\n}\n"],"names":["UNSPACED_SCRIPTS","RUN","RegExp","HAS_RUN","hasUnspacedRun","text","test","GRAPHEME_SEGMENTER","Intl","Segmenter","undefined","granularity","PUNCTUATION","BARRIER","graphemes","run","Array","from","segment","s","splitOnPunctuation","groups","g","push","length","filter","segmentField","out","replace","body","map","join","trim","QUALIFIER","SIDECAR_COLUMNS","runQuery","columnPrefix","segmentMatch","terms","quoted","pieces","split","i","ch","m","match","slice","index"],"mappings":"AAAA,kGAAkG;AAClG,6HAA6H;AAE7H,wEAAwE;AACxE,kFAAkF;AAClF,OAAO,MAAMA,mBAAmB,wGAAwG;AACxI,2FAA2F;AAC3F,gDAAgD;AAChD,MAAMC,MAAM,IAAIC,OAAO,CAAC,KAAK,EAAEF,iBAAiB,WAAW,CAAC,EAAE;AAC9D,MAAMG,UAAU,IAAID,OAAO,CAAC,CAAC,EAAEF,iBAAiB,CAAC,CAAC,EAAE;AAEpD,yFAAyF;AACzF,+CAA+C;AAC/C,OAAO,SAASI,eAAeC,IAAY;IACzC,OAAOF,QAAQG,IAAI,CAACD;AACtB;AACA,6FAA6F;AAC7F,wFAAwF;AACxF,MAAME,qBAAqB,IAAIC,KAAKC,SAAS,CAACC,WAAW;IAAEC,aAAa;AAAW;AACnF,8FAA8F;AAC9F,8FAA8F;AAC9F,MAAMC,cAAc;AACpB,6FAA6F;AAC7F,0FAA0F;AAC1F,MAAMC,UAAU;AAEhB,SAASC,UAAUC,GAAW;IAC5B,OAAOC,MAAMC,IAAI,CAACV,mBAAmBW,OAAO,CAACH,MAAM,CAACI,IAAMA,EAAED,OAAO;AACrE;AAEA,8FAA8F;AAC9F,0FAA0F;AAC1F,SAASE,mBAAmBL,GAAW;IACrC,MAAMM,SAAqB;QAAC,EAAE;KAAC;IAC/B,KAAK,MAAMC,KAAKR,UAAUC,KAAM;QAC9B,IAAIH,YAAYN,IAAI,CAACgB,IAAID,OAAOE,IAAI,CAAC,EAAE;aAClCF,MAAM,CAACA,OAAOG,MAAM,GAAG,EAAE,CAACD,IAAI,CAACD;IACtC;IACA,OAAOD,OAAOI,MAAM,CAAC,CAACH,IAAMA,EAAEE,MAAM,GAAG;AACzC;AAEA,6FAA6F;AAC7F,kGAAkG;AAClG,OAAO,SAASE,aAAarB,IAAY;IACvC,IAAI,CAACF,QAAQG,IAAI,CAACD,OAAO,OAAO;IAChC,MAAMsB,MAAMtB,KAAKuB,OAAO,CAAC3B,KAAK,CAACc;QAC7B,MAAMc,OAAOT,mBAAmBL,KAC7Be,GAAG,CAAC,CAACR,IAAMA,EAAES,IAAI,CAAC,MAClBA,IAAI,CAAC,CAAC,CAAC,EAAElB,QAAQ,CAAC,CAAC;QACtB,OAAO,CAAC,CAAC,EAAEA,QAAQ,CAAC,EAAEgB,KAAK,CAAC,EAAEhB,QAAQ,CAAC,CAAC;IAC1C;IACA,OAAOc,IAAIC,OAAO,CAAC,QAAQ,KAAKI,IAAI;AACtC;AAEA,0FAA0F;AAC1F,wFAAwF;AACxF,MAAMC,YAAY;AAElB,4FAA4F;AAC5F,uGAAuG;AACvG,MAAMC,kBAAkB;AAExB,4FAA4F;AAC5F,yFAAyF;AACzF,SAASC,SAASpB,GAAW,EAAEqB,YAAoB;IACjD,OAAOhB,mBAAmBL,KACvBe,GAAG,CAAC,CAACR,IAAM,GAAGc,eAAed,EAAEE,MAAM,GAAG,IAAI,CAAC,CAAC,EAAEF,EAAES,IAAI,CAAC,KAAK,CAAC,CAAC,GAAGT,CAAC,CAAC,EAAE,EAAE,EACvES,IAAI,CAAC;AACV;AAEA,2FAA2F;AAC3F,wFAAwF;AACxF,OAAO,SAASM,aAAaC,KAAa;IACxC,IAAI,CAACnC,QAAQG,IAAI,CAACgC,QAAQ,OAAOA;IACjC,IAAIX,MAAM;IACV,IAAIY,SAAS;IACb,MAAMC,SAASF,MAAMG,KAAK,CAACxC,MAAM,2CAA2C;IAC5E,IAAK,IAAIyC,IAAI,GAAGA,IAAIF,OAAOhB,MAAM,EAAEkB,IAAK;QACtC,IAAIA,IAAI,MAAM,GAAG;YACf,KAAK,MAAMC,MAAMH,MAAM,CAACE,EAAE,CAAE,IAAIC,OAAO,KAAKJ,SAAS,CAACA;YACtDZ,OAAOa,MAAM,CAACE,EAAE;YAChB;QACF;QACA,IAAIH,QAAQ;YACVZ,OAAOa,MAAM,CAACE,EAAE,EAAE,2CAA2C;YAC7D;QACF;QACA,MAAME,IAAIjB,IAAIkB,KAAK,CAACZ;QACpB,IAAIW,GAAG;YACLjB,MAAM,GAAGA,IAAImB,KAAK,CAAC,GAAGF,EAAEG,KAAK,IAAIH,CAAC,CAAC,EAAE,GAAGT,SAASK,MAAM,CAACE,EAAE,EAAE,GAAGE,CAAC,CAAC,EAAE,GAAGA,CAAC,CAAC,EAAE,CAAC,KAAK,CAAC,GAAG;QACtF,OAAO;YACLjB,OAAOQ,SAASK,MAAM,CAACE,EAAE,EAAER;QAC7B;IACF;IACA,OAAOP;AACT"}
|
package/dist/esm/watch.js
CHANGED
|
@@ -71,11 +71,8 @@ export async function runWatch(cfg, opts = {}) {
|
|
|
71
71
|
})();
|
|
72
72
|
}, debounceMs);
|
|
73
73
|
};
|
|
74
|
-
// Ignore our own state dir, or the heartbeat write would retrigger itself forever. An event
|
|
75
|
-
//
|
|
76
|
-
// load) is attributed to nothing and so reconciles: one reconcile that parses nothing costs
|
|
77
|
-
// less than missing a real edit. That is why the guard cannot promise zero reconciles, only
|
|
78
|
-
// that an identified state-dir write is never one of them.
|
|
74
|
+
// Ignore our own state dir, or the heartbeat write would retrigger itself forever. An event with
|
|
75
|
+
// an unresolvable filename (null, which fs.watch delivers under load) reconciles: parsing nothing costs less than missing a real edit.
|
|
79
76
|
const watcher = fsWatch(baseDir, {
|
|
80
77
|
recursive: true
|
|
81
78
|
}, (_event, filename)=>{
|
package/dist/esm/watch.js.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/watch.ts"],"sourcesContent":["import { watch as fsWatch } from 'node:fs';\nimport type { ResolvedConfig } from './config/index.ts';\nimport { STATE_DIR } from './config/index.ts';\nimport { SenseError } from './errors.ts';\nimport { guardedTick } from './lib/guarded-tick.ts';\nimport { docCount, getMeta, openStore, requireWatchConcurrency, setMeta } from './store/index.ts';\n\n// Watch is a cache pre-warmer, not a correctness mechanism: open() always reconciles anyway, so any fs event just triggers a debounced full reconcile.\nconst DEBOUNCE_MS = 200;\nconst HEARTBEAT_INTERVAL_MS = 5000;\nconst STALE_HEARTBEAT_MS = 15000;\n\nexport type WatchEvent = { type: 'started'; baseDir: string; dbPath: string } | { type: 'reconciled'; parsed: number; total: number; warnings: string[] } | { type: 'reconcile-error'; message: string };\n\nexport interface WatchOptions {\n force?: boolean;\n onEvent?: (event: WatchEvent) => void;\n // Aborting runs the same shutdown path as SIGINT/SIGTERM.\n signal?: AbortSignal;\n debounceMs?: number;\n heartbeatIntervalMs?: number;\n}\n\n// Runs in the foreground until SIGINT/SIGTERM/signal abort. Throws WATCH_ACTIVE if another watcher's heartbeat is still fresh and --force wasn't given.\nexport async function runWatch(cfg: ResolvedConfig, opts: WatchOptions = {}): Promise<void> {\n requireWatchConcurrency(cfg);\n const onEvent = opts.onEvent ?? (() => {});\n const debounceMs = opts.debounceMs ?? DEBOUNCE_MS;\n const heartbeatIntervalMs = opts.heartbeatIntervalMs ?? HEARTBEAT_INTERVAL_MS;\n const { store, dbPath, warnings: initialWarnings, parsed: initialParsed } = await openStore(cfg);\n const baseDir = cfg.baseDir;\n\n const existingHeartbeat = await getMeta(store, 'watch_heartbeat');\n if (existingHeartbeat && !opts.force) {\n const age = Date.now() - Date.parse(existingHeartbeat);\n if (age >= 0 && age < STALE_HEARTBEAT_MS) {\n await store.close();\n throw new SenseError('WATCH_ACTIVE', `another watcher appears active (heartbeat ${Math.round(age / 1000)}s ago); use --force to override`);\n }\n }\n\n onEvent({ type: 'started', baseDir, dbPath });\n if (initialWarnings.length > 0 || initialParsed > 0) {\n onEvent({ type: 'reconciled', parsed: initialParsed, total: await docCount(store), warnings: initialWarnings });\n }\n\n let stopping = false;\n const touchHeartbeat = async () => {\n await setMeta(store, 'watch_heartbeat', new Date().toISOString());\n await setMeta(store, 'watch_pid', String(process.pid));\n };\n await touchHeartbeat();\n\n let debounceTimer: NodeJS.Timeout | null = null;\n // A reconcile that has already started owns the store, and on a bulk reparse a live worker\n // pool as well. Shutdown waits on this rather than closing the connection underneath it.\n let inFlight: Promise<void> | null = null;\n const scheduleReconcile = () => {\n if (debounceTimer) clearTimeout(debounceTimer);\n debounceTimer = setTimeout(() => {\n debounceTimer = null;\n inFlight = (async () => {\n try {\n const { parsed, warnings } = await store.reconcile();\n onEvent({ type: 'reconciled', parsed, total: await docCount(store), warnings });\n } catch (err) {\n onEvent({ type: 'reconcile-error', message: (err as Error).message });\n } finally {\n inFlight = null;\n }\n })();\n }, debounceMs);\n };\n\n // Ignore our own state dir, or the heartbeat write would retrigger itself forever. An event\n //
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/watch.ts"],"sourcesContent":["import { watch as fsWatch } from 'node:fs';\nimport type { ResolvedConfig } from './config/index.ts';\nimport { STATE_DIR } from './config/index.ts';\nimport { SenseError } from './errors.ts';\nimport { guardedTick } from './lib/guarded-tick.ts';\nimport { docCount, getMeta, openStore, requireWatchConcurrency, setMeta } from './store/index.ts';\n\n// Watch is a cache pre-warmer, not a correctness mechanism: open() always reconciles anyway, so any fs event just triggers a debounced full reconcile.\nconst DEBOUNCE_MS = 200;\nconst HEARTBEAT_INTERVAL_MS = 5000;\nconst STALE_HEARTBEAT_MS = 15000;\n\nexport type WatchEvent = { type: 'started'; baseDir: string; dbPath: string } | { type: 'reconciled'; parsed: number; total: number; warnings: string[] } | { type: 'reconcile-error'; message: string };\n\nexport interface WatchOptions {\n force?: boolean;\n onEvent?: (event: WatchEvent) => void;\n // Aborting runs the same shutdown path as SIGINT/SIGTERM.\n signal?: AbortSignal;\n debounceMs?: number;\n heartbeatIntervalMs?: number;\n}\n\n// Runs in the foreground until SIGINT/SIGTERM/signal abort. Throws WATCH_ACTIVE if another watcher's heartbeat is still fresh and --force wasn't given.\nexport async function runWatch(cfg: ResolvedConfig, opts: WatchOptions = {}): Promise<void> {\n requireWatchConcurrency(cfg);\n const onEvent = opts.onEvent ?? (() => {});\n const debounceMs = opts.debounceMs ?? DEBOUNCE_MS;\n const heartbeatIntervalMs = opts.heartbeatIntervalMs ?? HEARTBEAT_INTERVAL_MS;\n const { store, dbPath, warnings: initialWarnings, parsed: initialParsed } = await openStore(cfg);\n const baseDir = cfg.baseDir;\n\n const existingHeartbeat = await getMeta(store, 'watch_heartbeat');\n if (existingHeartbeat && !opts.force) {\n const age = Date.now() - Date.parse(existingHeartbeat);\n if (age >= 0 && age < STALE_HEARTBEAT_MS) {\n await store.close();\n throw new SenseError('WATCH_ACTIVE', `another watcher appears active (heartbeat ${Math.round(age / 1000)}s ago); use --force to override`);\n }\n }\n\n onEvent({ type: 'started', baseDir, dbPath });\n if (initialWarnings.length > 0 || initialParsed > 0) {\n onEvent({ type: 'reconciled', parsed: initialParsed, total: await docCount(store), warnings: initialWarnings });\n }\n\n let stopping = false;\n const touchHeartbeat = async () => {\n await setMeta(store, 'watch_heartbeat', new Date().toISOString());\n await setMeta(store, 'watch_pid', String(process.pid));\n };\n await touchHeartbeat();\n\n let debounceTimer: NodeJS.Timeout | null = null;\n // A reconcile that has already started owns the store, and on a bulk reparse a live worker\n // pool as well. Shutdown waits on this rather than closing the connection underneath it.\n let inFlight: Promise<void> | null = null;\n const scheduleReconcile = () => {\n if (debounceTimer) clearTimeout(debounceTimer);\n debounceTimer = setTimeout(() => {\n debounceTimer = null;\n inFlight = (async () => {\n try {\n const { parsed, warnings } = await store.reconcile();\n onEvent({ type: 'reconciled', parsed, total: await docCount(store), warnings });\n } catch (err) {\n onEvent({ type: 'reconcile-error', message: (err as Error).message });\n } finally {\n inFlight = null;\n }\n })();\n }, debounceMs);\n };\n\n // Ignore our own state dir, or the heartbeat write would retrigger itself forever. An event with\n // an unresolvable filename (null, which fs.watch delivers under load) reconciles: parsing nothing costs less than missing a real edit.\n const watcher = fsWatch(baseDir, { recursive: true }, (_event, filename) => {\n if (typeof filename === 'string' && filename.startsWith(STATE_DIR)) return;\n scheduleReconcile();\n });\n const heartbeatTimer = setInterval(\n guardedTick(touchHeartbeat, () => stopping),\n heartbeatIntervalMs\n );\n\n return new Promise<void>((resolveShutdown) => {\n // SIGINT/SIGTERM and an aborted signal all run this same path exactly once; each is\n // unregistered here too so a second runWatch call in the same process starts clean.\n const shutdown = async () => {\n if (stopping) return;\n stopping = true;\n process.off('SIGINT', shutdown);\n process.off('SIGTERM', shutdown);\n opts.signal?.removeEventListener('abort', shutdown);\n clearInterval(heartbeatTimer);\n if (debounceTimer) clearTimeout(debounceTimer);\n watcher.close();\n await inFlight;\n await setMeta(store, 'watch_heartbeat', null);\n await setMeta(store, 'watch_pid', null);\n await store.close();\n resolveShutdown();\n };\n process.once('SIGINT', shutdown);\n process.once('SIGTERM', shutdown);\n if (opts.signal?.aborted) shutdown();\n else opts.signal?.addEventListener('abort', shutdown, { once: true });\n });\n}\n"],"names":["watch","fsWatch","STATE_DIR","SenseError","guardedTick","docCount","getMeta","openStore","requireWatchConcurrency","setMeta","DEBOUNCE_MS","HEARTBEAT_INTERVAL_MS","STALE_HEARTBEAT_MS","runWatch","cfg","opts","onEvent","debounceMs","heartbeatIntervalMs","store","dbPath","warnings","initialWarnings","parsed","initialParsed","baseDir","existingHeartbeat","force","age","Date","now","parse","close","Math","round","type","length","total","stopping","touchHeartbeat","toISOString","String","process","pid","debounceTimer","inFlight","scheduleReconcile","clearTimeout","setTimeout","reconcile","err","message","watcher","recursive","_event","filename","startsWith","heartbeatTimer","setInterval","Promise","resolveShutdown","shutdown","off","signal","removeEventListener","clearInterval","once","aborted","addEventListener"],"mappings":"AAAA,SAASA,SAASC,OAAO,QAAQ,UAAU;AAE3C,SAASC,SAAS,QAAQ,oBAAoB;AAC9C,SAASC,UAAU,QAAQ,cAAc;AACzC,SAASC,WAAW,QAAQ,wBAAwB;AACpD,SAASC,QAAQ,EAAEC,OAAO,EAAEC,SAAS,EAAEC,uBAAuB,EAAEC,OAAO,QAAQ,mBAAmB;AAElG,uJAAuJ;AACvJ,MAAMC,cAAc;AACpB,MAAMC,wBAAwB;AAC9B,MAAMC,qBAAqB;AAa3B,wJAAwJ;AACxJ,OAAO,eAAeC,SAASC,GAAmB,EAAEC,OAAqB,CAAC,CAAC;QAEzDA,eACGA,kBACSA;IAH5BP,wBAAwBM;IACxB,MAAME,WAAUD,gBAAAA,KAAKC,OAAO,cAAZD,2BAAAA,gBAAiB,KAAO;IACxC,MAAME,cAAaF,mBAAAA,KAAKE,UAAU,cAAfF,8BAAAA,mBAAmBL;IACtC,MAAMQ,uBAAsBH,4BAAAA,KAAKG,mBAAmB,cAAxBH,uCAAAA,4BAA4BJ;IACxD,MAAM,EAAEQ,KAAK,EAAEC,MAAM,EAAEC,UAAUC,eAAe,EAAEC,QAAQC,aAAa,EAAE,GAAG,MAAMjB,UAAUO;IAC5F,MAAMW,UAAUX,IAAIW,OAAO;IAE3B,MAAMC,oBAAoB,MAAMpB,QAAQa,OAAO;IAC/C,IAAIO,qBAAqB,CAACX,KAAKY,KAAK,EAAE;QACpC,MAAMC,MAAMC,KAAKC,GAAG,KAAKD,KAAKE,KAAK,CAACL;QACpC,IAAIE,OAAO,KAAKA,MAAMhB,oBAAoB;YACxC,MAAMO,MAAMa,KAAK;YACjB,MAAM,IAAI7B,WAAW,gBAAgB,CAAC,0CAA0C,EAAE8B,KAAKC,KAAK,CAACN,MAAM,MAAM,+BAA+B,CAAC;QAC3I;IACF;IAEAZ,QAAQ;QAAEmB,MAAM;QAAWV;QAASL;IAAO;IAC3C,IAAIE,gBAAgBc,MAAM,GAAG,KAAKZ,gBAAgB,GAAG;QACnDR,QAAQ;YAAEmB,MAAM;YAAcZ,QAAQC;YAAea,OAAO,MAAMhC,SAASc;YAAQE,UAAUC;QAAgB;IAC/G;IAEA,IAAIgB,WAAW;IACf,MAAMC,iBAAiB;QACrB,MAAM9B,QAAQU,OAAO,mBAAmB,IAAIU,OAAOW,WAAW;QAC9D,MAAM/B,QAAQU,OAAO,aAAasB,OAAOC,QAAQC,GAAG;IACtD;IACA,MAAMJ;IAEN,IAAIK,gBAAuC;IAC3C,2FAA2F;IAC3F,yFAAyF;IACzF,IAAIC,WAAiC;IACrC,MAAMC,oBAAoB;QACxB,IAAIF,eAAeG,aAAaH;QAChCA,gBAAgBI,WAAW;YACzBJ,gBAAgB;YAChBC,WAAW,AAAC,CAAA;gBACV,IAAI;oBACF,MAAM,EAAEtB,MAAM,EAAEF,QAAQ,EAAE,GAAG,MAAMF,MAAM8B,SAAS;oBAClDjC,QAAQ;wBAAEmB,MAAM;wBAAcZ;wBAAQc,OAAO,MAAMhC,SAASc;wBAAQE;oBAAS;gBAC/E,EAAE,OAAO6B,KAAK;oBACZlC,QAAQ;wBAAEmB,MAAM;wBAAmBgB,SAAS,AAACD,IAAcC,OAAO;oBAAC;gBACrE,SAAU;oBACRN,WAAW;gBACb;YACF,CAAA;QACF,GAAG5B;IACL;IAEA,iGAAiG;IACjG,uIAAuI;IACvI,MAAMmC,UAAUnD,QAAQwB,SAAS;QAAE4B,WAAW;IAAK,GAAG,CAACC,QAAQC;QAC7D,IAAI,OAAOA,aAAa,YAAYA,SAASC,UAAU,CAACtD,YAAY;QACpE4C;IACF;IACA,MAAMW,iBAAiBC,YACrBtD,YAAYmC,gBAAgB,IAAMD,WAClCpB;IAGF,OAAO,IAAIyC,QAAc,CAACC;YAoBpB7C,cACCA;QApBL,oFAAoF;QACpF,oFAAoF;QACpF,MAAM8C,WAAW;gBAKf9C;YAJA,IAAIuB,UAAU;YACdA,WAAW;YACXI,QAAQoB,GAAG,CAAC,UAAUD;YACtBnB,QAAQoB,GAAG,CAAC,WAAWD;aACvB9C,eAAAA,KAAKgD,MAAM,cAAXhD,mCAAAA,aAAaiD,mBAAmB,CAAC,SAASH;YAC1CI,cAAcR;YACd,IAAIb,eAAeG,aAAaH;YAChCQ,QAAQpB,KAAK;YACb,MAAMa;YACN,MAAMpC,QAAQU,OAAO,mBAAmB;YACxC,MAAMV,QAAQU,OAAO,aAAa;YAClC,MAAMA,MAAMa,KAAK;YACjB4B;QACF;QACAlB,QAAQwB,IAAI,CAAC,UAAUL;QACvBnB,QAAQwB,IAAI,CAAC,WAAWL;QACxB,KAAI9C,eAAAA,KAAKgD,MAAM,cAAXhD,mCAAAA,aAAaoD,OAAO,EAAEN;cACrB9C,gBAAAA,KAAKgD,MAAM,cAAXhD,oCAAAA,cAAaqD,gBAAgB,CAAC,SAASP,UAAU;YAAEK,MAAM;QAAK;IACrE;AACF"}
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/workers/parse.ts"],"sourcesContent":["import Tinypool from 'tinypool';\nimport type { Config, FeatureName } from '../config/index.ts';\nimport { FEATURES } from '../features/index.ts';\nimport type { ParsedDoc } from '../scan/index.ts';\nimport { parseFile } from '../scan/index.ts';\nimport type { FileStat } from '../scan/list.ts';\nimport { featuresForFile } from '../scan/reparse.ts';\nimport type { WorkerErrorPayload } from '../scan/worker-error.ts';\nimport { serializeError } from '../scan/worker-error.ts';\n\n// Constant for the whole dispatch, so it crosses once per worker instead of once per task
|
|
1
|
+
{"version":3,"sources":["/Users/kevin/Dev/OpenSource/ai/sensemaking/src/workers/parse.ts"],"sourcesContent":["import Tinypool from 'tinypool';\nimport type { Config, FeatureName } from '../config/index.ts';\nimport { FEATURES } from '../features/index.ts';\nimport type { ParsedDoc } from '../scan/index.ts';\nimport { parseFile } from '../scan/index.ts';\nimport type { FileStat } from '../scan/list.ts';\nimport { featuresForFile } from '../scan/reparse.ts';\nimport type { WorkerErrorPayload } from '../scan/worker-error.ts';\nimport { serializeError } from '../scan/worker-error.ts';\n\n// Constant for the whole dispatch, so it crosses once per worker instead of once per task. A\n// Feature carries closures and cannot cross the thread boundary; its name can, and the registry here resolves it back.\nexport interface ParseWorkerData {\n cfg: Config;\n featureNames: FeatureName[];\n}\n\n// tinypool's task, in and out. The task itself is one FileStat. Result carries only what\n// parseFile already returns -- extracted text and per-feature values, never the mdast tree.\nexport type ParseTask = FileStat;\n\nexport type ParseTaskResult = { ok: true; doc: ParsedDoc; warnings: string[] } | { ok: false; error: WorkerErrorPayload };\n\n// Read once per worker, not per task.\nconst { cfg, featureNames } = Tinypool.workerData as ParseWorkerData;\n// Filtering the registry (rather than mapping the names) keeps registry order, which is the\n// order `extracted` keys land in on the serial path.\nconst selected = FEATURES.filter((feature) => featureNames.includes(feature.name));\n\nexport default function parseTask(file: ParseTask): ParseTaskResult {\n try {\n const { doc, warnings } = parseFile(file, featuresForFile(selected, cfg, file), cfg);\n return { ok: true, doc, warnings };\n } catch (err) {\n return { ok: false, error: serializeError(err) };\n }\n}\n"],"names":["Tinypool","FEATURES","parseFile","featuresForFile","serializeError","cfg","featureNames","workerData","selected","filter","feature","includes","name","parseTask","file","doc","warnings","ok","err","error"],"mappings":"AAAA,OAAOA,cAAc,WAAW;AAEhC,SAASC,QAAQ,QAAQ,uBAAuB;AAEhD,SAASC,SAAS,QAAQ,mBAAmB;AAE7C,SAASC,eAAe,QAAQ,qBAAqB;AAErD,SAASC,cAAc,QAAQ,0BAA0B;AAezD,sCAAsC;AACtC,MAAM,EAAEC,GAAG,EAAEC,YAAY,EAAE,GAAGN,SAASO,UAAU;AACjD,4FAA4F;AAC5F,qDAAqD;AACrD,MAAMC,WAAWP,SAASQ,MAAM,CAAC,CAACC,UAAYJ,aAAaK,QAAQ,CAACD,QAAQE,IAAI;AAEhF,eAAe,SAASC,UAAUC,IAAe;IAC/C,IAAI;QACF,MAAM,EAAEC,GAAG,EAAEC,QAAQ,EAAE,GAAGd,UAAUY,MAAMX,gBAAgBK,UAAUH,KAAKS,OAAOT;QAChF,OAAO;YAAEY,IAAI;YAAMF;YAAKC;QAAS;IACnC,EAAE,OAAOE,KAAK;QACZ,OAAO;YAAED,IAAI;YAAOE,OAAOf,eAAec;QAAK;IACjD;AACF"}
|
package/package.json
CHANGED
|
@@ -1,12 +1,14 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "sensemaking",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.19.1",
|
|
4
4
|
"description": "Query and search your markdown notes with context-aware progressive disclosure: SQL over frontmatter, links, and text, plus semantic search and link-graph ranking. No server, no build step",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"markdown",
|
|
7
7
|
"frontmatter",
|
|
8
8
|
"sql",
|
|
9
9
|
"sqlite",
|
|
10
|
+
"duckdb",
|
|
11
|
+
"turso",
|
|
10
12
|
"fts5",
|
|
11
13
|
"full-text-search",
|
|
12
14
|
"query",
|
package/schema.json
CHANGED
|
@@ -114,8 +114,8 @@
|
|
|
114
114
|
},
|
|
115
115
|
"store": {
|
|
116
116
|
"type": "string",
|
|
117
|
-
"enum": ["sqlite", "duckdb"],
|
|
118
|
-
"description": "Backing store engine, defaulting to \"sqlite\" (zero-dependency, Node's built-in SQLite). \"duckdb\"
|
|
117
|
+
"enum": ["sqlite", "duckdb", "turso"],
|
|
118
|
+
"description": "Backing store engine, defaulting to \"sqlite\" (zero-dependency, Node's built-in SQLite). \"duckdb\" and \"turso\" are experimental: the first command that opens such a tree installs its engine package on its own (@duckdb/node-api, a one-time native download of about 110 MB; @tursodatabase/database, much smaller). The same commands and table names run on all three; what does not port is FTS5 syntax. Under duckdb and turso, `search` text and raw MATCH reject FTS5's prefix (foo*), boolean (AND/OR/NOT), NEAR, initial-token (^), and column-filter (title:foo) operators with a named error (STORE_CAPABILITY_MISSING) that says how to rephrase or set \"store\" to \"sqlite\"; bare words and quoted phrases work on all three. Raw SQL is a per-store dialect: the tables are portable everywhere, but sqlite's FTS5 MATCH, snippet(), and bm25() do not run under duckdb (its fts extension has its own functions) or turso (Tantivy fts_match/fts_score), so saved queries written in FTS5 syntax are sqlite dialect. The has/basename/segment functions are registered on sqlite and duckdb but not turso, whose client cannot register SQL functions at all. `sense watch` requires sqlite; on duckdb and turso it errors at start naming the fix. Each store keeps its own cache file (.sense/cache.db, .sense/cache.duckdb, .sense/cache.turso.db); switching stores rebuilds the index rather than migrating it."
|
|
119
119
|
},
|
|
120
120
|
"queries": {
|
|
121
121
|
"type": "object",
|
package/skills/sense/SKILL.md
CHANGED
|
@@ -7,7 +7,7 @@ description: "Query a markdown tree with the sense CLI: filter notes by frontmat
|
|
|
7
7
|
|
|
8
8
|
SQL over a markdown tree, kept fresh by a filesystem check on every query. Every file becomes rows in `frontmatter` (one column per key, plus `path`/`_mtime`/`_ctime`/`_size`/`_rank`/`_parse_error`; `_ctime` is filesystem birthtime, which a clone or copy resets just like `_mtime`), `content` (`title`, `summary`, `text`, `path`; an FTS5 index on the default store, with machine-written `title_seg`/`summary_seg`/`text_seg` sidecars used for matching Chinese, Japanese, Thai, Khmer, Lao, and Burmese text, not for reading), `links` (`src`, `target`, `dst`, `embed`; `NULL` dst = dead link; `embed` 1 for `![[...]]` embeds, 0 for links; one row per distinct written target and kind with alias and anchor stripped, so `[[Foo]]` and `[[Foo|alias]]` are one row while `[[Foo]]` and `[[notes/Foo]]` are two rows that can share a `dst`, and a target both linked and embedded is a row of each kind; extraction matches Obsidian's own graph: comments, code, and link-syntax text yield no rows, `[[#Anchor]]` is a self-edge, a frontmatter value that is exactly `[[X]]` is a link, and a basename collision resolves to the linking note itself, else the shortest path), `tags` (`path`, `tag`: frontmatter and inline `#tags` merged and deduplicated; nested tags stored full, so `book/scifi` matches `tag = 'book' OR tag LIKE 'book/%'`), `sections` (heading outline with line ranges and token estimates), and `preset_files` (`path`, `preset`: which presets cover which files). Features add their own storage; `map` and `status` report which are on.
|
|
9
9
|
|
|
10
|
-
**Stores.** The config's `store` key picks the backing store: `sqlite` (default, zero-dependency, Node's built-in SQLite) or `duckdb` (
|
|
10
|
+
**Stores.** The config's `store` key picks the backing store: `sqlite` (default, zero-dependency, Node's built-in SQLite), or the experimental `duckdb` and `turso` (the first command that opens such a tree installs that engine's package on its own: `@duckdb/node-api`, a one-time native download of about 110 MB, or the much smaller `@tursodatabase/database`). The tables, `?` placeholders, quoted identifiers, and the `scope` binding are the same on all three, so ordinary frontmatter SQL ports as written. Two things do not port. First, FTS5: `content` is an FTS5 table on sqlite and a plain table on the other two, so hand-written `MATCH`, `snippet()`, `bm25()`, and sqlite's date-function forms run only on sqlite (duckdb has its own fts functions and date syntax; turso has Tantivy's `fts_match`/`fts_score`), and `search` text under duckdb and turso rejects FTS5's prefix (`foo*`), boolean (`AND`/`OR`/`NOT`), `NEAR`, initial-token (`^`), and column-filter (`title:foo`) operators with a named error (STORE_CAPABILITY_MISSING) that says how to rephrase or set `store` to `sqlite`; bare words and quoted phrases work on all three. Second, the `has`/`basename`/`segment` functions are registered on sqlite and duckdb but not turso, whose client cannot register SQL functions at all, so a query calling them under turso fails with `no such function`. `sense watch` requires sqlite; under duckdb and turso it errors at start naming the fix.
|
|
11
11
|
|
|
12
12
|
## What each tool is for
|
|
13
13
|
|
|
@@ -110,7 +110,7 @@ WHERE a.dst = ? AND b.dst IS NOT NULL AND b.dst <> a.dst;
|
|
|
110
110
|
|
|
111
111
|
- A saved `{ sql }` written against `scope` is preset-agnostic: `sense <name> --preset raw` re-points the same statement at another layer, so one entry serves every preset instead of one copy each.
|
|
112
112
|
- `content MATCH` only works against the fts5 table by its own name, never through an alias or a view: `FROM content c ... WHERE c MATCH 'x'` fails with `no such column: c`. This is why `--preset` binds a table to join rather than shadowing the tables.
|
|
113
|
-
- `content MATCH` takes FTS5 syntax: `a OR b`, `"phrase"`, `pref*`, `NEAR(a b, 5)`, `summary: term`. Stemmed; markdown stripped at index time. Double-quote any term with punctuation. Bare `customer-facing` errors (`-` reads as a column filter), bare apostrophes are syntax errors: write `"customer-facing"`, `"founder's"`. This is the default sqlite store's grammar: under `duckdb` the operator forms (`a OR b`, `pref*`, `NEAR`, `^`, column filters) are a named error naming the rephrase, and `MATCH` itself does not run (see Stores).
|
|
113
|
+
- `content MATCH` takes FTS5 syntax: `a OR b`, `"phrase"`, `pref*`, `NEAR(a b, 5)`, `summary: term`. Stemmed; markdown stripped at index time. Double-quote any term with punctuation. Bare `customer-facing` errors (`-` reads as a column filter), bare apostrophes are syntax errors: write `"customer-facing"`, `"founder's"`. This is the default sqlite store's grammar: under `duckdb` and `turso` the operator forms (`a OR b`, `pref*`, `NEAR`, `^`, column filters) are a named error naming the rephrase, and `MATCH` itself does not run (see Stores).
|
|
114
114
|
- A language written without word spaces (Chinese, Japanese, Thai, Khmer, Lao, Burmese) is indexed per grapheme and searched as an ordered grapheme phrase against the `_seg` sidecar columns: substring semantics, what `grep` gives, a query matches wherever its exact text occurs, including inside a longer run (`京都` matches `东京都政府`, correctly, because it's there at position 2), and needs no minimum length. No decision is needed for these languages. Hand-written SQL is not rewritten for you, so a raw `content MATCH '数据库'` finds nothing: write `content MATCH segment(?)` and bind the terms. `segment()` returns text with no such run unchanged, so it is safe to leave in a query whatever the tree's language.
|
|
115
115
|
- Rank with `ORDER BY bm25(content, 10.0, 5.0, 1.0)` (title > summary > body); the full form `bm25(content, 10.0, 5.0, 1.0, 0, 10.0, 5.0, 1.0)` mirrors the same weights onto the `_seg` sidecars, so a title hit found through `title_seg` ranks like one found through `title` (the three-weight form still runs, FTS5 defaults unnamed columns to 1.0, but ranks a sidecar match at body weight). Excerpt with `snippet(content, 2, '«', '»', '…', 10)`, naming the `text` column explicitly: `-1` means best column, which can surface a machine-spaced sidecar as the excerpt (`search` itself always names column 2). snippet() re-tokenizes each matched doc and its cost grows superlinearly with doc size, measured ~10 s per query on a tree holding one 1 MB note. `search` bounds this itself (docs past 16 KB get an equivalent excerpt another way); in hand-written SQL, guard it: `CASE WHEN length(text) <= 16384 THEN snippet(...) END`, or select `title`/`summary` instead of an excerpt.
|
|
116
116
|
- Select `content.title`/`content.summary` (always exist, empty when absent) rather than `f.title`/`f.summary` (discovered columns; error on trees that never declare them).
|