@tekmidian/pai 0.15.0 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +11 -3
- package/dist/{auto-route-CLYcToZ5.mjs → auto-route-CdvjEWs5.mjs} +4 -4
- package/dist/{auto-route-CLYcToZ5.mjs.map → auto-route-CdvjEWs5.mjs.map} +1 -1
- package/dist/auto-route-DkIRampF.mjs +86 -0
- package/dist/auto-route-DkIRampF.mjs.map +1 -0
- package/dist/{kg-extraction-C8DEUHTS.mjs → checkpoint-block-Bl0HD3U2.mjs} +442 -13
- package/dist/checkpoint-block-Bl0HD3U2.mjs.map +1 -0
- package/dist/checkpoint-block-C05sK_HK.mjs +1056 -0
- package/dist/checkpoint-block-C05sK_HK.mjs.map +1 -0
- package/dist/checkpoint-block-CUCG10Uh.mjs +1022 -0
- package/dist/checkpoint-block-CUCG10Uh.mjs.map +1 -0
- package/dist/checkpoint-block-D5IyLLTr.mjs +1062 -0
- package/dist/checkpoint-block-D5IyLLTr.mjs.map +1 -0
- package/dist/cli/index.mjs +14 -12
- package/dist/cli/index.mjs.map +1 -1
- package/dist/cli/program.d.mts.map +1 -1
- package/dist/cli/program.mjs +31 -16
- package/dist/cli/program.mjs.map +1 -1
- package/dist/{clusters-BYPw7vfW.mjs → clusters-CKWDwcMX.mjs} +2 -2
- package/dist/{clusters-BYPw7vfW.mjs.map → clusters-CKWDwcMX.mjs.map} +1 -1
- package/dist/{config-C8m-tPhP.mjs → config-DGOfoGfm.mjs} +1 -1
- package/dist/{config-C8m-tPhP.mjs.map → config-DGOfoGfm.mjs.map} +1 -1
- package/dist/daemon/index.mjs +18 -16
- package/dist/daemon/index.mjs.map +1 -1
- package/dist/{daemon-COSvHTNu.mjs → daemon-1rKbN-fd.mjs} +31 -30
- package/dist/{daemon-COSvHTNu.mjs.map → daemon-1rKbN-fd.mjs.map} +1 -1
- package/dist/daemon-9AL6Pwh-.mjs +1581 -0
- package/dist/daemon-9AL6Pwh-.mjs.map +1 -0
- package/dist/daemon-Bbk-_rV6.mjs +1581 -0
- package/dist/daemon-Bbk-_rV6.mjs.map +1 -0
- package/dist/daemon-BvcaMk4g.mjs +1581 -0
- package/dist/daemon-BvcaMk4g.mjs.map +1 -0
- package/dist/daemon-CTDEfwcq.mjs +1576 -0
- package/dist/daemon-CTDEfwcq.mjs.map +1 -0
- package/dist/daemon-ClbfaF60.mjs +1576 -0
- package/dist/daemon-ClbfaF60.mjs.map +1 -0
- package/dist/daemon-DDQNDgYF.mjs +1576 -0
- package/dist/daemon-DDQNDgYF.mjs.map +1 -0
- package/dist/daemon-DU8yDWlQ.mjs +1576 -0
- package/dist/daemon-DU8yDWlQ.mjs.map +1 -0
- package/dist/daemon-NJr4iSW3.mjs +1576 -0
- package/dist/daemon-NJr4iSW3.mjs.map +1 -0
- package/dist/daemon-mcp/index.mjs +2 -2
- package/dist/db-BtuN768f.mjs +206 -0
- package/dist/db-BtuN768f.mjs.map +1 -0
- package/dist/{db-CmYbAVCD.mjs → db-CYmBWcjh.mjs} +1 -1
- package/dist/{db-CmYbAVCD.mjs.map → db-CYmBWcjh.mjs.map} +1 -1
- package/dist/{detect-KjycLtXM.mjs → detect-Bf2z-oKB.mjs} +1 -1
- package/dist/{detect-KjycLtXM.mjs.map → detect-Bf2z-oKB.mjs.map} +1 -1
- package/dist/{detector-ZiOhHszd.mjs → detector-Bwzm3E_x.mjs} +2 -2
- package/dist/{detector-ZiOhHszd.mjs.map → detector-Bwzm3E_x.mjs.map} +1 -1
- package/dist/detector-rq9JGTmS.mjs +74 -0
- package/dist/detector-rq9JGTmS.mjs.map +1 -0
- package/dist/embeddings-Bn86ssxR.mjs +119 -0
- package/dist/embeddings-Bn86ssxR.mjs.map +1 -0
- package/dist/{embeddings-BJPOcbik.mjs → embeddings-DGRAPAYb.mjs} +1 -1
- package/dist/{embeddings-BJPOcbik.mjs.map → embeddings-DGRAPAYb.mjs.map} +1 -1
- package/dist/{factory-Q88X1bAN.mjs → factory-0-57Ac-7.mjs} +8 -5
- package/dist/{factory-Q88X1bAN.mjs.map → factory-0-57Ac-7.mjs.map} +1 -1
- package/dist/factory-HcQsQqh0.mjs +72 -0
- package/dist/factory-HcQsQqh0.mjs.map +1 -0
- package/dist/factory-PDXQdkLO.mjs +72 -0
- package/dist/factory-PDXQdkLO.mjs.map +1 -0
- package/dist/{helpers-IjZkXBhj.mjs → helpers-crDEr6S2.mjs} +1 -1
- package/dist/{helpers-IjZkXBhj.mjs.map → helpers-crDEr6S2.mjs.map} +1 -1
- package/dist/hooks/capture-all-events.mjs.map +1 -1
- package/dist/hooks/capture-session-summary.mjs.map +1 -1
- package/dist/hooks/capture-tool-output.mjs.map +1 -1
- package/dist/hooks/cleanup-session-files.mjs.map +1 -1
- package/dist/hooks/context-compression-hook.mjs +336 -48
- package/dist/hooks/context-compression-hook.mjs.map +4 -4
- package/dist/hooks/initialize-session.mjs.map +1 -1
- package/dist/hooks/inject-observations.mjs.map +1 -1
- package/dist/hooks/load-core-context.mjs.map +1 -1
- package/dist/hooks/load-project-context.mjs +188 -6
- package/dist/hooks/load-project-context.mjs.map +4 -4
- package/dist/hooks/observe.mjs.map +1 -1
- package/dist/hooks/post-compact-inject.mjs.map +1 -1
- package/dist/hooks/security-validator.mjs.map +1 -1
- package/dist/hooks/session-commands.mjs.map +1 -1
- package/dist/hooks/stop-hook.mjs +331 -43
- package/dist/hooks/stop-hook.mjs.map +4 -4
- package/dist/hooks/subagent-stop-hook.mjs.map +1 -1
- package/dist/hooks/sync-todo-to-md.mjs.map +1 -1
- package/dist/hooks/update-tab-on-action.mjs.map +1 -1
- package/dist/hooks/update-tab-titles.mjs.map +1 -1
- package/dist/hooks/whisper-rules.mjs.map +1 -1
- package/dist/index.d.mts +0 -9
- package/dist/index.d.mts.map +1 -1
- package/dist/index.mjs +10 -9
- package/dist/indexer-AEcT8wHf.mjs +1 -0
- package/dist/{indexer-backend-Cm5RS7y0.mjs → indexer-backend-BQZXoarl.mjs} +3 -3
- package/dist/{indexer-backend-Cm5RS7y0.mjs.map → indexer-backend-BQZXoarl.mjs.map} +1 -1
- package/dist/indexer-backend-BmaS2VcD.mjs +299 -0
- package/dist/indexer-backend-BmaS2VcD.mjs.map +1 -0
- package/dist/indexer-backend-COrJ5rjS.mjs +299 -0
- package/dist/indexer-backend-COrJ5rjS.mjs.map +1 -0
- package/dist/indexer-backend-NwD7TBGh.mjs +301 -0
- package/dist/indexer-backend-NwD7TBGh.mjs.map +1 -0
- package/dist/{ipc-client-CoyUHPod.mjs → ipc-client-vO2GV325.mjs} +1 -1
- package/dist/{ipc-client-CoyUHPod.mjs.map → ipc-client-vO2GV325.mjs.map} +1 -1
- package/dist/{kg-entity-LXblD7LZ.mjs → kg-entity-DbOMPdF9.mjs} +1 -1
- package/dist/{kg-entity-LXblD7LZ.mjs.map → kg-entity-DbOMPdF9.mjs.map} +1 -1
- package/dist/{latent-ideas-RyC6eyqI.mjs → latent-ideas-BwM3WdBp.mjs} +4 -4
- package/dist/{latent-ideas-RyC6eyqI.mjs.map → latent-ideas-BwM3WdBp.mjs.map} +1 -1
- package/dist/latent-ideas-C6IFFT0-.mjs +191 -0
- package/dist/latent-ideas-C6IFFT0-.mjs.map +1 -0
- package/dist/{main-resolver-uFxNiDy7.mjs → main-resolver-DimnS2qP.mjs} +22 -22
- package/dist/{main-resolver-uFxNiDy7.mjs.map → main-resolver-DimnS2qP.mjs.map} +1 -1
- package/dist/{migrate-B02wDDgS.mjs → migrate-fLD6rAdO.mjs} +2 -2
- package/dist/{migrate-B02wDDgS.mjs.map → migrate-fLD6rAdO.mjs.map} +1 -1
- package/dist/{neighborhood-CklGIB8r.mjs → neighborhood-CzRNG2Oy.mjs} +2 -2
- package/dist/{neighborhood-CklGIB8r.mjs.map → neighborhood-CzRNG2Oy.mjs.map} +1 -1
- package/dist/neighborhood-DZY0wkpE.mjs +135 -0
- package/dist/neighborhood-DZY0wkpE.mjs.map +1 -0
- package/dist/{note-context-qvZlVXVO.mjs → note-context-CrfMbr4R.mjs} +1 -1
- package/dist/{note-context-qvZlVXVO.mjs.map → note-context-CrfMbr4R.mjs.map} +1 -1
- package/dist/{pai-marker-HVBwwBW-.mjs → pai-marker-B20KqhA8.mjs} +44 -2
- package/dist/pai-marker-B20KqhA8.mjs.map +1 -0
- package/dist/pick-B0q5cy79.mjs +11785 -0
- package/dist/pick-B0q5cy79.mjs.map +1 -0
- package/dist/{pick-r95zvydr.mjs → pick-B3zEYjin.mjs} +2288 -497
- package/dist/pick-B3zEYjin.mjs.map +1 -0
- package/dist/pick-BHE9_iA0.mjs +11947 -0
- package/dist/pick-BHE9_iA0.mjs.map +1 -0
- package/dist/pick-BaTjoNqM.mjs +11544 -0
- package/dist/pick-BaTjoNqM.mjs.map +1 -0
- package/dist/pick-C6JbZfN4.mjs +11893 -0
- package/dist/pick-C6JbZfN4.mjs.map +1 -0
- package/dist/pick-CK-bJyjp.mjs +11883 -0
- package/dist/pick-CK-bJyjp.mjs.map +1 -0
- package/dist/pick-CLBkGHyI.mjs +11826 -0
- package/dist/pick-CLBkGHyI.mjs.map +1 -0
- package/dist/pick-CTOif21h.mjs +11544 -0
- package/dist/pick-CTOif21h.mjs.map +1 -0
- package/dist/pick-Ck2cdQgD.mjs +11893 -0
- package/dist/pick-Ck2cdQgD.mjs.map +1 -0
- package/dist/pick-Cr4nJzHA.mjs +11914 -0
- package/dist/pick-Cr4nJzHA.mjs.map +1 -0
- package/dist/pick-D4dFfKAu.mjs +11893 -0
- package/dist/pick-D4dFfKAu.mjs.map +1 -0
- package/dist/pick-D7-lYiGX.mjs +11543 -0
- package/dist/pick-D7-lYiGX.mjs.map +1 -0
- package/dist/pick-DMjCY0np.mjs +11875 -0
- package/dist/pick-DMjCY0np.mjs.map +1 -0
- package/dist/pick-DbX4OrCP.mjs +11543 -0
- package/dist/pick-DbX4OrCP.mjs.map +1 -0
- package/dist/pick-iCiPjlnC.mjs +11893 -0
- package/dist/pick-iCiPjlnC.mjs.map +1 -0
- package/dist/pick-sF9GnQUS.mjs +11893 -0
- package/dist/pick-sF9GnQUS.mjs.map +1 -0
- package/dist/pick-yHKUrnfG.mjs +11793 -0
- package/dist/pick-yHKUrnfG.mjs.map +1 -0
- package/dist/postgres-BJxiqtqg.mjs +891 -0
- package/dist/postgres-BJxiqtqg.mjs.map +1 -0
- package/dist/{postgres-B451Hnkm.mjs → postgres-zjKqPl4X.mjs} +2 -2
- package/dist/{postgres-B451Hnkm.mjs.map → postgres-zjKqPl4X.mjs.map} +1 -1
- package/dist/{query-feedback-fapbqRgh.mjs → query-feedback-BoY8_Dbb.mjs} +1 -1
- package/dist/{query-feedback-fapbqRgh.mjs.map → query-feedback-BoY8_Dbb.mjs.map} +1 -1
- package/dist/{reranker-C08R99zn.mjs → reranker-CMNZcfVx.mjs} +1 -1
- package/dist/{reranker-C08R99zn.mjs.map → reranker-CMNZcfVx.mjs.map} +1 -1
- package/dist/{search-HcdKtMla.mjs → search-CpTv1I24.mjs} +3 -3
- package/dist/{search-HcdKtMla.mjs.map → search-CpTv1I24.mjs.map} +1 -1
- package/dist/search-i2nlQ-JM.mjs +298 -0
- package/dist/search-i2nlQ-JM.mjs.map +1 -0
- package/dist/skills/End/SKILL.md +78 -0
- package/dist/skills/Pause/SKILL.md +84 -0
- package/dist/sqlite-CZ0LFili.mjs +271 -0
- package/dist/sqlite-CZ0LFili.mjs.map +1 -0
- package/dist/{sqlite-DQpY1Esi.mjs → sqlite-ChSCYwiC.mjs} +3 -3
- package/dist/{sqlite-DQpY1Esi.mjs.map → sqlite-ChSCYwiC.mjs.map} +1 -1
- package/dist/sqlite-Nw8ZW5Rj.mjs +271 -0
- package/dist/sqlite-Nw8ZW5Rj.mjs.map +1 -0
- package/dist/{state-BIlxNRUn.mjs → state-BXIdbxDs.mjs} +1 -1
- package/dist/{state-BIlxNRUn.mjs.map → state-BXIdbxDs.mjs.map} +1 -1
- package/dist/{stop-words-BwplsQ3z.mjs → stop-words-BaMEGVeY.mjs} +1 -1
- package/dist/{stop-words-BwplsQ3z.mjs.map → stop-words-BaMEGVeY.mjs.map} +1 -1
- package/dist/{sync-uPR4g438.mjs → sync--BoxBBok.mjs} +6 -205
- package/dist/sync--BoxBBok.mjs.map +1 -0
- package/dist/sync-CmBKOL3K.mjs +310 -0
- package/dist/sync-CmBKOL3K.mjs.map +1 -0
- package/dist/{themes-Cg8oG9Ra.mjs → themes-BBOlGXAg.mjs} +3 -3
- package/dist/{themes-Cg8oG9Ra.mjs.map → themes-BBOlGXAg.mjs.map} +1 -1
- package/dist/themes-nvRM8v84.mjs +148 -0
- package/dist/themes-nvRM8v84.mjs.map +1 -0
- package/dist/tools-B6BmaAvz.mjs +1939 -0
- package/dist/tools-B6BmaAvz.mjs.map +1 -0
- package/dist/tools-CAIY8TAb.mjs +1939 -0
- package/dist/tools-CAIY8TAb.mjs.map +1 -0
- package/dist/{tools-Op3C2Nm6.mjs → tools-CznjzfYz.mjs} +26 -26
- package/dist/{tools-Op3C2Nm6.mjs.map → tools-CznjzfYz.mjs.map} +1 -1
- package/dist/{trace-CLK-NPkb.mjs → trace-B8oz1ok0.mjs} +1 -1
- package/dist/{trace-CLK-NPkb.mjs.map → trace-B8oz1ok0.mjs.map} +1 -1
- package/dist/{utils-CqqgB0dH.mjs → utils-BAxjW3j8.mjs} +1 -1
- package/dist/{utils-CqqgB0dH.mjs.map → utils-BAxjW3j8.mjs.map} +1 -1
- package/dist/{vault-indexer-Dn_BccVK.mjs → vault-indexer-DrttVeEk.mjs} +2 -2
- package/dist/{vault-indexer-Dn_BccVK.mjs.map → vault-indexer-DrttVeEk.mjs.map} +1 -1
- package/dist/{work-queue-worker-BwouhVkb.mjs → work-queue-worker-BPXLkEik.mjs} +47 -30
- package/dist/work-queue-worker-BPXLkEik.mjs.map +1 -0
- package/dist/work-queue-worker-BoHIH70K.mjs +1856 -0
- package/dist/work-queue-worker-BoHIH70K.mjs.map +1 -0
- package/dist/work-queue-worker-CZGcQoLw.mjs +1856 -0
- package/dist/work-queue-worker-CZGcQoLw.mjs.map +1 -0
- package/dist/work-queue-worker-D05Ze9e9.mjs +1856 -0
- package/dist/work-queue-worker-D05Ze9e9.mjs.map +1 -0
- package/dist/work-queue-worker-DdXFINfH.mjs +1856 -0
- package/dist/work-queue-worker-DdXFINfH.mjs.map +1 -0
- package/dist/{zettelkasten-NwBD1NsT.mjs → zettelkasten-CTlt1dGv.mjs} +4 -4
- package/dist/{zettelkasten-NwBD1NsT.mjs.map → zettelkasten-CTlt1dGv.mjs.map} +1 -1
- package/dist/zettelkasten-D0N_A_OE.mjs +1063 -0
- package/dist/zettelkasten-D0N_A_OE.mjs.map +1 -0
- package/docs/commands/README.md +18 -0
- package/docs/commands/backup.md +1 -1
- package/docs/commands/clear-names.md +1 -1
- package/docs/commands/daemon.md +1 -1
- package/docs/commands/db.md +1 -1
- package/docs/commands/end.md +4 -1
- package/docs/commands/help.md +1 -1
- package/docs/commands/kg.md +1 -1
- package/docs/commands/mcp.md +1 -1
- package/docs/commands/memory.md +1 -1
- package/docs/commands/notify.md +1 -1
- package/docs/commands/observation.md +1 -1
- package/docs/commands/obsidian.md +1 -1
- package/docs/commands/pause.md +5 -2
- package/docs/commands/project.md +1 -1
- package/docs/commands/projects.md +1 -1
- package/docs/commands/registry.md +18 -1
- package/docs/commands/restore.md +1 -1
- package/docs/commands/session.md +295 -0
- package/docs/commands/sessions.md +1 -1
- package/docs/commands/setup.md +1 -1
- package/docs/commands/shell-init.md +1 -1
- package/docs/commands/skill.md +1 -1
- package/docs/commands/task.md +1 -1
- package/docs/commands/topic.md +1 -1
- package/docs/commands/update.md +1 -1
- package/docs/commands/zettel.md +1 -1
- package/package.json +1 -1
- package/plugins/context-preservation/hooks/hooks.json +12 -1
- package/scripts/build-hooks.mjs +33 -3
- package/src/hooks/session-autosave.sh +69 -0
- package/src/hooks/session-stop.sh +32 -1
- package/src/hooks/ts/lib/project-utils/todo.test.ts +132 -0
- package/src/hooks/ts/lib/project-utils/todo.ts +60 -34
- package/src/hooks/ts/session-start/load-project-context.ts +44 -0
- package/dist/kg-extraction-C8DEUHTS.mjs.map +0 -1
- package/dist/pai-marker-HVBwwBW-.mjs.map +0 -1
- package/dist/pick-r95zvydr.mjs.map +0 -1
- package/dist/sync-uPR4g438.mjs.map +0 -1
- package/dist/work-queue-worker-BwouhVkb.mjs.map +0 -1
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"search-HcdKtMla.mjs","names":[],"sources":["../src/memory/search.ts"],"sourcesContent":["/**\n * Search over the PAI federation memory index.\n *\n * Provides three search modes:\n * - keyword — BM25 full-text search (default, fast, no ML required)\n * - semantic — Brute-force cosine similarity over pre-computed embeddings\n * - hybrid — Normalized combination of BM25 + cosine scores\n *\n * BM25 uses SQLite's FTS5 extension. Semantic search requires embeddings to\n * have been generated first via `embedChunks()` in the indexer.\n */\n\nimport type { Database } from \"better-sqlite3\";\nimport { deserializeEmbedding, cosineSimilarity } from \"./embeddings.js\";\nimport { STOP_WORDS } from \"../utils/stop-words.js\";\n\n// ---------------------------------------------------------------------------\n// Types\n// ---------------------------------------------------------------------------\n\nexport interface SearchResult {\n projectId: number;\n projectSlug?: string; // populated from registry after search when available\n path: string;\n startLine: number;\n endLine: number;\n snippet: string;\n score: number; // raw BM25 score (lower = more relevant in FTS5)\n tier: string;\n source: string;\n updatedAt?: number; // Unix ms from memory_chunks.updated_at\n lastAccessedAt?: number; // Unix ms from memory_chunks.last_accessed_at (QW2)\n chunkId?: string; // chunk ID for last_accessed_at update (QW2)\n}\n\nexport interface SearchOptions {\n /** Restrict search to these project IDs. */\n projectIds?: number[];\n /** Restrict to 'memory' or 'notes' sources. */\n sources?: string[];\n /** Restrict to specific tier(s): 'evergreen' | 'daily' | 'topic' | 'session' */\n tiers?: string[];\n /** Maximum number of results to return. Default 10. */\n maxResults?: number;\n /** Minimum BM25 score threshold (FTS5 scores are negative; 0.0 means no filter). */\n minScore?: number;\n}\n\n// STOP_WORDS imported from utils/stop-words.ts\n\n// ---------------------------------------------------------------------------\n// Query builder\n// ---------------------------------------------------------------------------\n\n/**\n * Convert a free-text query into an FTS5 query string.\n *\n * Strategy:\n * 1. Tokenise by whitespace and punctuation\n * 2. Remove stop words and tokens shorter than 2 characters\n * 3. Double-quote each remaining token (exact word form)\n * 4. Join with OR so that any matching token returns a result\n *\n * Using OR instead of AND is critical for multi-word queries: the words rarely\n * all appear in the same chunk, so AND would return zero results. FTS5 BM25\n * scoring naturally ranks chunks where more terms match higher, so the most\n * relevant chunks still surface at the top.\n *\n * Example: \"Synchrotech interview follow-up Gilles\"\n * → `\"synchrotech\" OR \"interview\" OR \"follow\" OR \"gilles\"`\n * → chunks matching any term, ranked by how many terms match\n */\nexport function buildFtsQuery(query: string): string {\n const tokens = query\n .toLowerCase()\n .split(/[\\s\\p{P}]+/u)\n .filter(Boolean)\n .filter((t) => t.length >= 2)\n .filter((t) => !STOP_WORDS.has(t))\n // Escape any double-quotes inside the token (FTS5 uses them as delimiters)\n .map((t) => `\"${t.replace(/\"/g, '\"\"')}\"`)\n\n if (tokens.length === 0) {\n // Fallback: use original query as a raw string (may produce no results)\n return `\"${query.replace(/\"/g, '\"\"')}\"`;\n }\n\n return tokens.join(\" OR \");\n}\n\n// ---------------------------------------------------------------------------\n// Search\n// ---------------------------------------------------------------------------\n\n/**\n * Search across all indexed memory using FTS5 BM25 ranking.\n *\n * Results are ordered by BM25 score (most relevant first).\n * FTS5 bm25() returns negative values; closer to 0 = more relevant.\n * We negate the score so callers get positive values where higher = better.\n *\n * Multilingual note: SQLite FTS5 uses the `unicode61` tokenizer by default,\n * which handles Unicode correctly (German umlauts, French accents, etc.) without\n * language-specific stemming. No changes needed here — it is already\n * multilingual-safe.\n */\nexport function searchMemory(\n db: Database,\n query: string,\n opts?: SearchOptions,\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n const ftsQuery = buildFtsQuery(query);\n\n // Build the SQL with optional filters\n const conditions: string[] = [];\n const params: (string | number)[] = [ftsQuery];\n\n if (opts?.projectIds && opts.projectIds.length > 0) {\n const placeholders = opts.projectIds.map(() => \"?\").join(\", \");\n conditions.push(`c.project_id IN (${placeholders})`);\n params.push(...opts.projectIds);\n }\n\n if (opts?.sources && opts.sources.length > 0) {\n const placeholders = opts.sources.map(() => \"?\").join(\", \");\n conditions.push(`c.source IN (${placeholders})`);\n params.push(...opts.sources);\n }\n\n if (opts?.tiers && opts.tiers.length > 0) {\n const placeholders = opts.tiers.map(() => \"?\").join(\", \");\n conditions.push(`c.tier IN (${placeholders})`);\n params.push(...opts.tiers);\n }\n\n const whereClause = conditions.length > 0\n ? \"AND \" + conditions.join(\" AND \")\n : \"\";\n\n params.push(maxResults);\n\n // FTS5: join memory_fts with memory_chunks to get metadata\n // bm25(memory_fts) returns negative values (lower = better match)\n const sql = `\n SELECT\n c.id,\n c.project_id,\n c.path,\n c.start_line,\n c.end_line,\n c.text AS snippet,\n c.tier,\n c.source,\n c.updated_at,\n c.last_accessed_at,\n c.relevance_score,\n bm25(memory_fts) AS bm25_score\n FROM memory_fts\n JOIN memory_chunks c ON memory_fts.id = c.id\n WHERE memory_fts MATCH ?\n ${whereClause}\n ORDER BY bm25_score\n LIMIT ?\n `;\n\n let rows: Array<{\n id: string;\n project_id: number;\n path: string;\n start_line: number;\n end_line: number;\n snippet: string;\n tier: string;\n source: string;\n updated_at: number;\n last_accessed_at: number | null;\n relevance_score: number | null;\n bm25_score: number;\n }>;\n\n try {\n rows = db.prepare(sql).all(...params) as typeof rows;\n } catch {\n // FTS5 MATCH throws when the query is invalid — return empty results\n return [];\n }\n\n const minScore = opts?.minScore ?? 0.0;\n\n return rows\n .map((row) => {\n // Negate so higher = better match for callers\n const baseScore = -row.bm25_score;\n // MR2: scale by feedback relevance_score: multiplier in [0.5, 1.5]\n const relevanceScore = row.relevance_score ?? 0.5;\n const score = baseScore * (0.5 + relevanceScore);\n return {\n chunkId: row.id,\n projectId: row.project_id,\n path: row.path,\n startLine: row.start_line,\n endLine: row.end_line,\n snippet: row.snippet,\n score,\n tier: row.tier,\n source: row.source,\n updatedAt: row.updated_at,\n lastAccessedAt: row.last_accessed_at ?? undefined,\n };\n })\n .filter((r) => r.score >= minScore);\n}\n\n// ---------------------------------------------------------------------------\n// Semantic search\n// ---------------------------------------------------------------------------\n\n/**\n * Search chunks using brute-force cosine similarity over stored embeddings.\n *\n * Only chunks that have a non-null embedding BLOB are considered. Chunks\n * without embeddings are silently skipped (they can be embedded later via\n * `embedChunks()`).\n *\n * @param queryEmbedding Pre-computed Float32Array for the search query.\n */\nexport function searchMemorySemantic(\n db: Database,\n queryEmbedding: Float32Array,\n opts?: SearchOptions,\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n\n // Build the SQL filter conditions\n const conditions: string[] = [\"embedding IS NOT NULL\"];\n const params: (string | number)[] = [];\n\n if (opts?.projectIds && opts.projectIds.length > 0) {\n const placeholders = opts.projectIds.map(() => \"?\").join(\", \");\n conditions.push(`project_id IN (${placeholders})`);\n params.push(...opts.projectIds);\n }\n\n if (opts?.sources && opts.sources.length > 0) {\n const placeholders = opts.sources.map(() => \"?\").join(\", \");\n conditions.push(`source IN (${placeholders})`);\n params.push(...opts.sources);\n }\n\n if (opts?.tiers && opts.tiers.length > 0) {\n const placeholders = opts.tiers.map(() => \"?\").join(\", \");\n conditions.push(`tier IN (${placeholders})`);\n params.push(...opts.tiers);\n }\n\n const where = \"WHERE \" + conditions.join(\" AND \");\n\n // Hard cap for SQLite semantic path — prevents OOM on large corpora.\n // Use Postgres for production semantic search.\n const sql = `\n SELECT id, project_id, path, start_line, end_line, text, tier, source, embedding, updated_at, last_accessed_at, relevance_score\n FROM memory_chunks\n ${where}\n LIMIT 5000\n `;\n\n const rows = db.prepare(sql).all(...params) as Array<{\n id: string;\n project_id: number;\n path: string;\n start_line: number;\n end_line: number;\n text: string;\n tier: string;\n source: string;\n embedding: Buffer;\n updated_at: number;\n last_accessed_at: number | null;\n relevance_score: number | null;\n }>;\n\n if (rows.length === 0) return [];\n\n // Compute cosine similarity for every chunk\n const scored = rows.map((row) => {\n const vec = deserializeEmbedding(row.embedding);\n const baseScore = cosineSimilarity(queryEmbedding, vec);\n // MR2: scale by feedback relevance_score: multiplier in [0.5, 1.5]\n const relevanceScore = row.relevance_score ?? 0.5;\n const score = baseScore * (0.5 + relevanceScore);\n return {\n chunkId: row.id,\n projectId: row.project_id,\n path: row.path,\n startLine: row.start_line,\n endLine: row.end_line,\n snippet: row.text,\n score,\n tier: row.tier,\n source: row.source,\n updatedAt: row.updated_at,\n lastAccessedAt: row.last_accessed_at ?? undefined,\n };\n });\n\n // Sort by descending similarity, apply optional min score filter, limit\n const minScore = opts?.minScore ?? -Infinity;\n\n return scored\n .filter((r) => r.score >= minScore)\n .sort((a, b) => b.score - a.score)\n .slice(0, maxResults);\n}\n\n// ---------------------------------------------------------------------------\n// Hybrid search\n// ---------------------------------------------------------------------------\n\n/**\n * Combine BM25 keyword search and semantic search using normalized scores.\n *\n * Both score sets are min-max normalized to [0,1] before combining, so neither\n * dominates the other regardless of their raw scales.\n *\n * @param queryEmbedding Pre-computed embedding for the query.\n * @param keywordWeight Weight for BM25 score (default 0.5).\n * @param semanticWeight Weight for cosine similarity score (default 0.5).\n */\nexport function searchMemoryHybrid(\n db: Database,\n query: string,\n queryEmbedding: Float32Array,\n opts?: SearchOptions & { keywordWeight?: number; semanticWeight?: number },\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n const kw = opts?.keywordWeight ?? 0.5;\n const sw = opts?.semanticWeight ?? 0.5;\n\n // Fetch keyword results — 50 candidates is sufficient for min-max normalization\n const keywordResults = searchMemory(db, query, {\n ...opts,\n maxResults: 50,\n });\n\n // Fetch semantic results — 50 candidates is sufficient for min-max normalization\n const semanticResults = searchMemorySemantic(db, queryEmbedding, {\n ...opts,\n maxResults: 50,\n });\n\n if (keywordResults.length === 0 && semanticResults.length === 0) return [];\n\n // Build a map of chunk ID → combined result\n // Use \"projectId:path:startLine:endLine\" as a stable key (same as chunk IDs)\n const keyFor = (r: SearchResult) =>\n `${r.projectId}:${r.path}:${r.startLine}:${r.endLine}`;\n\n // Min-max normalize helper\n function minMaxNormalize(items: SearchResult[]): Map<string, number> {\n if (items.length === 0) return new Map();\n const min = Math.min(...items.map((r) => r.score));\n const max = Math.max(...items.map((r) => r.score));\n const range = max - min;\n const m = new Map<string, number>();\n for (const r of items) {\n m.set(keyFor(r), range === 0 ? 1 : (r.score - min) / range);\n }\n return m;\n }\n\n const kwNorm = minMaxNormalize(keywordResults);\n const semNorm = minMaxNormalize(semanticResults);\n\n // Union of all chunk keys\n const allKeys = new Set<string>([\n ...keywordResults.map(keyFor),\n ...semanticResults.map(keyFor),\n ]);\n\n // Build a lookup from key → result metadata\n const metaMap = new Map<string, SearchResult>();\n for (const r of [...keywordResults, ...semanticResults]) {\n metaMap.set(keyFor(r), r);\n }\n\n // Combine scores\n const combined: Array<SearchResult & { combinedScore: number }> = [];\n for (const key of allKeys) {\n const meta = metaMap.get(key)!;\n const kwScore = kwNorm.get(key) ?? 0;\n const semScore = semNorm.get(key) ?? 0;\n const combinedScore = kw * kwScore + sw * semScore;\n combined.push({ ...meta, score: combinedScore, combinedScore });\n }\n\n // Sort by combined score descending\n return combined\n .sort((a, b) => b.score - a.score)\n .slice(0, maxResults)\n .map(({ combinedScore: _unused, ...r }) => r);\n}\n\n// ---------------------------------------------------------------------------\n// Access timestamp tracking (QW2)\n// ---------------------------------------------------------------------------\n\n/**\n * Update last_accessed_at for a set of chunk IDs to the current timestamp.\n *\n * Called after a successful search to record that these chunks were retrieved.\n * This enables the recency boost to account for access patterns, not just\n * modification time.\n *\n * Best-effort: errors are silently ignored so search is never blocked.\n */\nexport function touchChunksLastAccessed(db: Database, chunkIds: string[]): void {\n if (chunkIds.length === 0) return;\n try {\n const now = Date.now();\n const placeholders = chunkIds.map(() => \"?\").join(\", \");\n db.prepare(\n `UPDATE memory_chunks SET last_accessed_at = ? WHERE id IN (${placeholders})`\n ).run(now, ...chunkIds);\n } catch {\n // non-critical — do not block search results\n }\n}\n\n// ---------------------------------------------------------------------------\n// Slug lookup helper\n// ---------------------------------------------------------------------------\n\n/**\n * Populate the projectSlug field on search results by looking up project IDs\n * in the registry database.\n */\nexport function populateSlugs(\n results: SearchResult[],\n registryDb: Database,\n): SearchResult[] {\n if (results.length === 0) return results;\n\n const ids = [...new Set(results.map((r) => r.projectId))];\n const placeholders = ids.map(() => \"?\").join(\", \");\n const rows = registryDb\n .prepare(`SELECT id, slug FROM projects WHERE id IN (${placeholders})`)\n .all(...ids) as Array<{ id: number; slug: string }>;\n\n const slugMap = new Map(rows.map((r) => [r.id, r.slug]));\n\n return results.map((r) => ({\n ...r,\n projectSlug: slugMap.get(r.projectId),\n }));\n}\n\n// ---------------------------------------------------------------------------\n// Recency boost\n// ---------------------------------------------------------------------------\n\n/**\n * Apply exponential recency boost to search scores.\n *\n * Scores are first min-max normalized to [0,1], then multiplied by an\n * exponential decay factor based on chunk age. Normalization is required\n * because the cross-encoder reranker produces negative logit scores — naive\n * multiplication of a negative score by a decay factor (0 < d ≤ 1) would\n * make the score *less* negative, effectively boosting old results instead\n * of penalizing them.\n *\n * Formula: score_final = normalized * exp(-lambda * age_days)\n * where lambda = ln(2) / halfLifeDays, normalized ∈ [0,1]\n *\n * With default halfLifeDays=90, a 3-month-old chunk retains 50% of its\n * normalized score, a 6-month-old retains 25%, and a 1-year-old ~6%.\n *\n * Results without an updatedAt timestamp receive no decay penalty.\n * Results are re-sorted by the boosted score after application.\n *\n * @param results Search results with optional updatedAt timestamps.\n * @param halfLifeDays Score halves every N days. Default 90 (~3 months).\n * @returns New array sorted by decayed normalized score (descending).\n */\nexport function applyRecencyBoost(\n results: SearchResult[],\n halfLifeDays = 90,\n): SearchResult[] {\n if (halfLifeDays <= 0 || results.length === 0) return results;\n\n const lambda = Math.LN2 / halfLifeDays;\n const now = Date.now();\n\n // Min-max normalize scores to [0,1] so multiplicative decay works\n // correctly regardless of the raw score sign/scale.\n const scores = results.map((r) => r.score);\n const minScore = Math.min(...scores);\n const maxScore = Math.max(...scores);\n const range = maxScore - minScore;\n\n return results\n .map((r) => {\n const normalized = range === 0 ? 1 : (r.score - minScore) / range;\n // QW2: use the more recent of updated_at and last_accessed_at for recency decay\n const effectiveTs = r.updatedAt != null && r.lastAccessedAt != null\n ? Math.max(r.updatedAt, r.lastAccessedAt)\n : (r.lastAccessedAt ?? r.updatedAt);\n const decay = effectiveTs\n ? Math.exp(-lambda * Math.max(0, (now - effectiveTs) / 86_400_000))\n : 1; // no timestamp → no penalty\n return { ...r, score: normalized * decay };\n })\n .sort((a, b) => b.score - a.score);\n}\n"],"mappings":";;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;AAwEA,SAAgB,cAAc,OAAuB;CACnD,MAAM,SAAS,MACZ,aAAa,CACb,MAAM,cAAc,CACpB,OAAO,QAAQ,CACf,QAAQ,MAAM,EAAE,UAAU,EAAE,CAC5B,QAAQ,MAAM,CAAC,WAAW,IAAI,EAAE,CAAC,CAEjC,KAAK,MAAM,IAAI,EAAE,QAAQ,MAAM,OAAK,CAAC,GAAG;AAE3C,KAAI,OAAO,WAAW,EAEpB,QAAO,IAAI,MAAM,QAAQ,MAAM,OAAK,CAAC;AAGvC,QAAO,OAAO,KAAK,OAAO;;;;;;;;;;;;;;AAmB5B,SAAgB,aACd,IACA,OACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CACvC,MAAM,WAAW,cAAc,MAAM;CAGrC,MAAM,aAAuB,EAAE;CAC/B,MAAM,SAA8B,CAAC,SAAS;AAE9C,KAAI,MAAM,cAAc,KAAK,WAAW,SAAS,GAAG;EAClD,MAAM,eAAe,KAAK,WAAW,UAAU,IAAI,CAAC,KAAK,KAAK;AAC9D,aAAW,KAAK,oBAAoB,aAAa,GAAG;AACpD,SAAO,KAAK,GAAG,KAAK,WAAW;;AAGjC,KAAI,MAAM,WAAW,KAAK,QAAQ,SAAS,GAAG;EAC5C,MAAM,eAAe,KAAK,QAAQ,UAAU,IAAI,CAAC,KAAK,KAAK;AAC3D,aAAW,KAAK,gBAAgB,aAAa,GAAG;AAChD,SAAO,KAAK,GAAG,KAAK,QAAQ;;AAG9B,KAAI,MAAM,SAAS,KAAK,MAAM,SAAS,GAAG;EACxC,MAAM,eAAe,KAAK,MAAM,UAAU,IAAI,CAAC,KAAK,KAAK;AACzD,aAAW,KAAK,cAAc,aAAa,GAAG;AAC9C,SAAO,KAAK,GAAG,KAAK,MAAM;;CAG5B,MAAM,cAAc,WAAW,SAAS,IACpC,SAAS,WAAW,KAAK,QAAQ,GACjC;AAEJ,QAAO,KAAK,WAAW;CAIvB,MAAM,MAAM;;;;;;;;;;;;;;;;;QAiBN,YAAY;;;;CAKlB,IAAI;AAeJ,KAAI;AACF,SAAO,GAAG,QAAQ,IAAI,CAAC,IAAI,GAAG,OAAO;SAC/B;AAEN,SAAO,EAAE;;CAGX,MAAM,WAAW,MAAM,YAAY;AAEnC,QAAO,KACJ,KAAK,QAAQ;EAKZ,MAAM,QAHY,CAAC,IAAI,cAGI,MADJ,IAAI,mBAAmB;AAE9C,SAAO;GACL,SAAS,IAAI;GACb,WAAW,IAAI;GACf,MAAM,IAAI;GACV,WAAW,IAAI;GACf,SAAS,IAAI;GACb,SAAS,IAAI;GACb;GACA,MAAM,IAAI;GACV,QAAQ,IAAI;GACZ,WAAW,IAAI;GACf,gBAAgB,IAAI,oBAAoB;GACzC;GACD,CACD,QAAQ,MAAM,EAAE,SAAS,SAAS;;;;;;;;;;;AAgBvC,SAAgB,qBACd,IACA,gBACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CAGvC,MAAM,aAAuB,CAAC,wBAAwB;CACtD,MAAM,SAA8B,EAAE;AAEtC,KAAI,MAAM,cAAc,KAAK,WAAW,SAAS,GAAG;EAClD,MAAM,eAAe,KAAK,WAAW,UAAU,IAAI,CAAC,KAAK,KAAK;AAC9D,aAAW,KAAK,kBAAkB,aAAa,GAAG;AAClD,SAAO,KAAK,GAAG,KAAK,WAAW;;AAGjC,KAAI,MAAM,WAAW,KAAK,QAAQ,SAAS,GAAG;EAC5C,MAAM,eAAe,KAAK,QAAQ,UAAU,IAAI,CAAC,KAAK,KAAK;AAC3D,aAAW,KAAK,cAAc,aAAa,GAAG;AAC9C,SAAO,KAAK,GAAG,KAAK,QAAQ;;AAG9B,KAAI,MAAM,SAAS,KAAK,MAAM,SAAS,GAAG;EACxC,MAAM,eAAe,KAAK,MAAM,UAAU,IAAI,CAAC,KAAK,KAAK;AACzD,aAAW,KAAK,YAAY,aAAa,GAAG;AAC5C,SAAO,KAAK,GAAG,KAAK,MAAM;;CAO5B,MAAM,MAAM;;;MAJE,WAAW,WAAW,KAAK,QAAQ,CAOvC;;;CAIV,MAAM,OAAO,GAAG,QAAQ,IAAI,CAAC,IAAI,GAAG,OAAO;AAe3C,KAAI,KAAK,WAAW,EAAG,QAAO,EAAE;CAGhC,MAAM,SAAS,KAAK,KAAK,QAAQ;EAK/B,MAAM,QAHY,iBAAiB,gBADvB,qBAAqB,IAAI,UAAU,CACQ,IAG5B,MADJ,IAAI,mBAAmB;AAE9C,SAAO;GACL,SAAS,IAAI;GACb,WAAW,IAAI;GACf,MAAM,IAAI;GACV,WAAW,IAAI;GACf,SAAS,IAAI;GACb,SAAS,IAAI;GACb;GACA,MAAM,IAAI;GACV,QAAQ,IAAI;GACZ,WAAW,IAAI;GACf,gBAAgB,IAAI,oBAAoB;GACzC;GACD;CAGF,MAAM,WAAW,MAAM,YAAY;AAEnC,QAAO,OACJ,QAAQ,MAAM,EAAE,SAAS,SAAS,CAClC,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM,CACjC,MAAM,GAAG,WAAW;;;;;;;;;;;;AAiBzB,SAAgB,mBACd,IACA,OACA,gBACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CACvC,MAAM,KAAK,MAAM,iBAAiB;CAClC,MAAM,KAAK,MAAM,kBAAkB;CAGnC,MAAM,iBAAiB,aAAa,IAAI,OAAO;EAC7C,GAAG;EACH,YAAY;EACb,CAAC;CAGF,MAAM,kBAAkB,qBAAqB,IAAI,gBAAgB;EAC/D,GAAG;EACH,YAAY;EACb,CAAC;AAEF,KAAI,eAAe,WAAW,KAAK,gBAAgB,WAAW,EAAG,QAAO,EAAE;CAI1E,MAAM,UAAU,MACd,GAAG,EAAE,UAAU,GAAG,EAAE,KAAK,GAAG,EAAE,UAAU,GAAG,EAAE;CAG/C,SAAS,gBAAgB,OAA4C;AACnE,MAAI,MAAM,WAAW,EAAG,wBAAO,IAAI,KAAK;EACxC,MAAM,MAAM,KAAK,IAAI,GAAG,MAAM,KAAK,MAAM,EAAE,MAAM,CAAC;EAElD,MAAM,QADM,KAAK,IAAI,GAAG,MAAM,KAAK,MAAM,EAAE,MAAM,CAAC,GAC9B;EACpB,MAAM,oBAAI,IAAI,KAAqB;AACnC,OAAK,MAAM,KAAK,MACd,GAAE,IAAI,OAAO,EAAE,EAAE,UAAU,IAAI,KAAK,EAAE,QAAQ,OAAO,MAAM;AAE7D,SAAO;;CAGT,MAAM,SAAS,gBAAgB,eAAe;CAC9C,MAAM,UAAU,gBAAgB,gBAAgB;CAGhD,MAAM,UAAU,IAAI,IAAY,CAC9B,GAAG,eAAe,IAAI,OAAO,EAC7B,GAAG,gBAAgB,IAAI,OAAO,CAC/B,CAAC;CAGF,MAAM,0BAAU,IAAI,KAA2B;AAC/C,MAAK,MAAM,KAAK,CAAC,GAAG,gBAAgB,GAAG,gBAAgB,CACrD,SAAQ,IAAI,OAAO,EAAE,EAAE,EAAE;CAI3B,MAAM,WAA4D,EAAE;AACpE,MAAK,MAAM,OAAO,SAAS;EACzB,MAAM,OAAO,QAAQ,IAAI,IAAI;EAC7B,MAAM,UAAU,OAAO,IAAI,IAAI,IAAI;EACnC,MAAM,WAAW,QAAQ,IAAI,IAAI,IAAI;EACrC,MAAM,gBAAgB,KAAK,UAAU,KAAK;AAC1C,WAAS,KAAK;GAAE,GAAG;GAAM,OAAO;GAAe;GAAe,CAAC;;AAIjE,QAAO,SACJ,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM,CACjC,MAAM,GAAG,WAAW,CACpB,KAAK,EAAE,eAAe,SAAS,GAAG,QAAQ,EAAE;;;;;;;;;;;AAgBjD,SAAgB,wBAAwB,IAAc,UAA0B;AAC9E,KAAI,SAAS,WAAW,EAAG;AAC3B,KAAI;EACF,MAAM,MAAM,KAAK,KAAK;EACtB,MAAM,eAAe,SAAS,UAAU,IAAI,CAAC,KAAK,KAAK;AACvD,KAAG,QACD,8DAA8D,aAAa,GAC5E,CAAC,IAAI,KAAK,GAAG,SAAS;SACjB;;;;;;AAaV,SAAgB,cACd,SACA,YACgB;AAChB,KAAI,QAAQ,WAAW,EAAG,QAAO;CAEjC,MAAM,MAAM,CAAC,GAAG,IAAI,IAAI,QAAQ,KAAK,MAAM,EAAE,UAAU,CAAC,CAAC;CACzD,MAAM,eAAe,IAAI,UAAU,IAAI,CAAC,KAAK,KAAK;CAClD,MAAM,OAAO,WACV,QAAQ,8CAA8C,aAAa,GAAG,CACtE,IAAI,GAAG,IAAI;CAEd,MAAM,UAAU,IAAI,IAAI,KAAK,KAAK,MAAM,CAAC,EAAE,IAAI,EAAE,KAAK,CAAC,CAAC;AAExD,QAAO,QAAQ,KAAK,OAAO;EACzB,GAAG;EACH,aAAa,QAAQ,IAAI,EAAE,UAAU;EACtC,EAAE;;;;;;;;;;;;;;;;;;;;;;;;;AA8BL,SAAgB,kBACd,SACA,eAAe,IACC;AAChB,KAAI,gBAAgB,KAAK,QAAQ,WAAW,EAAG,QAAO;CAEtD,MAAM,SAAS,KAAK,MAAM;CAC1B,MAAM,MAAM,KAAK,KAAK;CAItB,MAAM,SAAS,QAAQ,KAAK,MAAM,EAAE,MAAM;CAC1C,MAAM,WAAW,KAAK,IAAI,GAAG,OAAO;CAEpC,MAAM,QADW,KAAK,IAAI,GAAG,OAAO,GACX;AAEzB,QAAO,QACJ,KAAK,MAAM;EACV,MAAM,aAAa,UAAU,IAAI,KAAK,EAAE,QAAQ,YAAY;EAE5D,MAAM,cAAc,EAAE,aAAa,QAAQ,EAAE,kBAAkB,OAC3D,KAAK,IAAI,EAAE,WAAW,EAAE,eAAe,GACtC,EAAE,kBAAkB,EAAE;EAC3B,MAAM,QAAQ,cACV,KAAK,IAAI,CAAC,SAAS,KAAK,IAAI,IAAI,MAAM,eAAe,MAAW,CAAC,GACjE;AACJ,SAAO;GAAE,GAAG;GAAG,OAAO,aAAa;GAAO;GAC1C,CACD,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM"}
|
|
1
|
+
{"version":3,"file":"search-CpTv1I24.mjs","names":[],"sources":["../src/memory/search.ts"],"sourcesContent":["/**\n * Search over the PAI federation memory index.\n *\n * Provides three search modes:\n * - keyword — BM25 full-text search (default, fast, no ML required)\n * - semantic — Brute-force cosine similarity over pre-computed embeddings\n * - hybrid — Normalized combination of BM25 + cosine scores\n *\n * BM25 uses SQLite's FTS5 extension. Semantic search requires embeddings to\n * have been generated first via `embedChunks()` in the indexer.\n */\n\nimport type { Database } from \"better-sqlite3\";\nimport { deserializeEmbedding, cosineSimilarity } from \"./embeddings.js\";\nimport { STOP_WORDS } from \"../utils/stop-words.js\";\n\n// ---------------------------------------------------------------------------\n// Types\n// ---------------------------------------------------------------------------\n\nexport interface SearchResult {\n projectId: number;\n projectSlug?: string; // populated from registry after search when available\n path: string;\n startLine: number;\n endLine: number;\n snippet: string;\n score: number; // raw BM25 score (lower = more relevant in FTS5)\n tier: string;\n source: string;\n updatedAt?: number; // Unix ms from memory_chunks.updated_at\n lastAccessedAt?: number; // Unix ms from memory_chunks.last_accessed_at (QW2)\n chunkId?: string; // chunk ID for last_accessed_at update (QW2)\n}\n\nexport interface SearchOptions {\n /** Restrict search to these project IDs. */\n projectIds?: number[];\n /** Restrict to 'memory' or 'notes' sources. */\n sources?: string[];\n /** Restrict to specific tier(s): 'evergreen' | 'daily' | 'topic' | 'session' */\n tiers?: string[];\n /** Maximum number of results to return. Default 10. */\n maxResults?: number;\n /** Minimum BM25 score threshold (FTS5 scores are negative; 0.0 means no filter). */\n minScore?: number;\n}\n\n// STOP_WORDS imported from utils/stop-words.ts\n\n// ---------------------------------------------------------------------------\n// Query builder\n// ---------------------------------------------------------------------------\n\n/**\n * Convert a free-text query into an FTS5 query string.\n *\n * Strategy:\n * 1. Tokenise by whitespace and punctuation\n * 2. Remove stop words and tokens shorter than 2 characters\n * 3. Double-quote each remaining token (exact word form)\n * 4. Join with OR so that any matching token returns a result\n *\n * Using OR instead of AND is critical for multi-word queries: the words rarely\n * all appear in the same chunk, so AND would return zero results. FTS5 BM25\n * scoring naturally ranks chunks where more terms match higher, so the most\n * relevant chunks still surface at the top.\n *\n * Example: \"Synchrotech interview follow-up Gilles\"\n * → `\"synchrotech\" OR \"interview\" OR \"follow\" OR \"gilles\"`\n * → chunks matching any term, ranked by how many terms match\n */\nexport function buildFtsQuery(query: string): string {\n const tokens = query\n .toLowerCase()\n .split(/[\\s\\p{P}]+/u)\n .filter(Boolean)\n .filter((t) => t.length >= 2)\n .filter((t) => !STOP_WORDS.has(t))\n // Escape any double-quotes inside the token (FTS5 uses them as delimiters)\n .map((t) => `\"${t.replace(/\"/g, '\"\"')}\"`)\n\n if (tokens.length === 0) {\n // Fallback: use original query as a raw string (may produce no results)\n return `\"${query.replace(/\"/g, '\"\"')}\"`;\n }\n\n return tokens.join(\" OR \");\n}\n\n// ---------------------------------------------------------------------------\n// Search\n// ---------------------------------------------------------------------------\n\n/**\n * Search across all indexed memory using FTS5 BM25 ranking.\n *\n * Results are ordered by BM25 score (most relevant first).\n * FTS5 bm25() returns negative values; closer to 0 = more relevant.\n * We negate the score so callers get positive values where higher = better.\n *\n * Multilingual note: SQLite FTS5 uses the `unicode61` tokenizer by default,\n * which handles Unicode correctly (German umlauts, French accents, etc.) without\n * language-specific stemming. No changes needed here — it is already\n * multilingual-safe.\n */\nexport function searchMemory(\n db: Database,\n query: string,\n opts?: SearchOptions,\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n const ftsQuery = buildFtsQuery(query);\n\n // Build the SQL with optional filters\n const conditions: string[] = [];\n const params: (string | number)[] = [ftsQuery];\n\n if (opts?.projectIds && opts.projectIds.length > 0) {\n const placeholders = opts.projectIds.map(() => \"?\").join(\", \");\n conditions.push(`c.project_id IN (${placeholders})`);\n params.push(...opts.projectIds);\n }\n\n if (opts?.sources && opts.sources.length > 0) {\n const placeholders = opts.sources.map(() => \"?\").join(\", \");\n conditions.push(`c.source IN (${placeholders})`);\n params.push(...opts.sources);\n }\n\n if (opts?.tiers && opts.tiers.length > 0) {\n const placeholders = opts.tiers.map(() => \"?\").join(\", \");\n conditions.push(`c.tier IN (${placeholders})`);\n params.push(...opts.tiers);\n }\n\n const whereClause = conditions.length > 0\n ? \"AND \" + conditions.join(\" AND \")\n : \"\";\n\n params.push(maxResults);\n\n // FTS5: join memory_fts with memory_chunks to get metadata\n // bm25(memory_fts) returns negative values (lower = better match)\n const sql = `\n SELECT\n c.id,\n c.project_id,\n c.path,\n c.start_line,\n c.end_line,\n c.text AS snippet,\n c.tier,\n c.source,\n c.updated_at,\n c.last_accessed_at,\n c.relevance_score,\n bm25(memory_fts) AS bm25_score\n FROM memory_fts\n JOIN memory_chunks c ON memory_fts.id = c.id\n WHERE memory_fts MATCH ?\n ${whereClause}\n ORDER BY bm25_score\n LIMIT ?\n `;\n\n let rows: Array<{\n id: string;\n project_id: number;\n path: string;\n start_line: number;\n end_line: number;\n snippet: string;\n tier: string;\n source: string;\n updated_at: number;\n last_accessed_at: number | null;\n relevance_score: number | null;\n bm25_score: number;\n }>;\n\n try {\n rows = db.prepare(sql).all(...params) as typeof rows;\n } catch {\n // FTS5 MATCH throws when the query is invalid — return empty results\n return [];\n }\n\n const minScore = opts?.minScore ?? 0.0;\n\n return rows\n .map((row) => {\n // Negate so higher = better match for callers\n const baseScore = -row.bm25_score;\n // MR2: scale by feedback relevance_score: multiplier in [0.5, 1.5]\n const relevanceScore = row.relevance_score ?? 0.5;\n const score = baseScore * (0.5 + relevanceScore);\n return {\n chunkId: row.id,\n projectId: row.project_id,\n path: row.path,\n startLine: row.start_line,\n endLine: row.end_line,\n snippet: row.snippet,\n score,\n tier: row.tier,\n source: row.source,\n updatedAt: row.updated_at,\n lastAccessedAt: row.last_accessed_at ?? undefined,\n };\n })\n .filter((r) => r.score >= minScore);\n}\n\n// ---------------------------------------------------------------------------\n// Semantic search\n// ---------------------------------------------------------------------------\n\n/**\n * Search chunks using brute-force cosine similarity over stored embeddings.\n *\n * Only chunks that have a non-null embedding BLOB are considered. Chunks\n * without embeddings are silently skipped (they can be embedded later via\n * `embedChunks()`).\n *\n * @param queryEmbedding Pre-computed Float32Array for the search query.\n */\nexport function searchMemorySemantic(\n db: Database,\n queryEmbedding: Float32Array,\n opts?: SearchOptions,\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n\n // Build the SQL filter conditions\n const conditions: string[] = [\"embedding IS NOT NULL\"];\n const params: (string | number)[] = [];\n\n if (opts?.projectIds && opts.projectIds.length > 0) {\n const placeholders = opts.projectIds.map(() => \"?\").join(\", \");\n conditions.push(`project_id IN (${placeholders})`);\n params.push(...opts.projectIds);\n }\n\n if (opts?.sources && opts.sources.length > 0) {\n const placeholders = opts.sources.map(() => \"?\").join(\", \");\n conditions.push(`source IN (${placeholders})`);\n params.push(...opts.sources);\n }\n\n if (opts?.tiers && opts.tiers.length > 0) {\n const placeholders = opts.tiers.map(() => \"?\").join(\", \");\n conditions.push(`tier IN (${placeholders})`);\n params.push(...opts.tiers);\n }\n\n const where = \"WHERE \" + conditions.join(\" AND \");\n\n // Hard cap for SQLite semantic path — prevents OOM on large corpora.\n // Use Postgres for production semantic search.\n const sql = `\n SELECT id, project_id, path, start_line, end_line, text, tier, source, embedding, updated_at, last_accessed_at, relevance_score\n FROM memory_chunks\n ${where}\n LIMIT 5000\n `;\n\n const rows = db.prepare(sql).all(...params) as Array<{\n id: string;\n project_id: number;\n path: string;\n start_line: number;\n end_line: number;\n text: string;\n tier: string;\n source: string;\n embedding: Buffer;\n updated_at: number;\n last_accessed_at: number | null;\n relevance_score: number | null;\n }>;\n\n if (rows.length === 0) return [];\n\n // Compute cosine similarity for every chunk\n const scored = rows.map((row) => {\n const vec = deserializeEmbedding(row.embedding);\n const baseScore = cosineSimilarity(queryEmbedding, vec);\n // MR2: scale by feedback relevance_score: multiplier in [0.5, 1.5]\n const relevanceScore = row.relevance_score ?? 0.5;\n const score = baseScore * (0.5 + relevanceScore);\n return {\n chunkId: row.id,\n projectId: row.project_id,\n path: row.path,\n startLine: row.start_line,\n endLine: row.end_line,\n snippet: row.text,\n score,\n tier: row.tier,\n source: row.source,\n updatedAt: row.updated_at,\n lastAccessedAt: row.last_accessed_at ?? undefined,\n };\n });\n\n // Sort by descending similarity, apply optional min score filter, limit\n const minScore = opts?.minScore ?? -Infinity;\n\n return scored\n .filter((r) => r.score >= minScore)\n .sort((a, b) => b.score - a.score)\n .slice(0, maxResults);\n}\n\n// ---------------------------------------------------------------------------\n// Hybrid search\n// ---------------------------------------------------------------------------\n\n/**\n * Combine BM25 keyword search and semantic search using normalized scores.\n *\n * Both score sets are min-max normalized to [0,1] before combining, so neither\n * dominates the other regardless of their raw scales.\n *\n * @param queryEmbedding Pre-computed embedding for the query.\n * @param keywordWeight Weight for BM25 score (default 0.5).\n * @param semanticWeight Weight for cosine similarity score (default 0.5).\n */\nexport function searchMemoryHybrid(\n db: Database,\n query: string,\n queryEmbedding: Float32Array,\n opts?: SearchOptions & { keywordWeight?: number; semanticWeight?: number },\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n const kw = opts?.keywordWeight ?? 0.5;\n const sw = opts?.semanticWeight ?? 0.5;\n\n // Fetch keyword results — 50 candidates is sufficient for min-max normalization\n const keywordResults = searchMemory(db, query, {\n ...opts,\n maxResults: 50,\n });\n\n // Fetch semantic results — 50 candidates is sufficient for min-max normalization\n const semanticResults = searchMemorySemantic(db, queryEmbedding, {\n ...opts,\n maxResults: 50,\n });\n\n if (keywordResults.length === 0 && semanticResults.length === 0) return [];\n\n // Build a map of chunk ID → combined result\n // Use \"projectId:path:startLine:endLine\" as a stable key (same as chunk IDs)\n const keyFor = (r: SearchResult) =>\n `${r.projectId}:${r.path}:${r.startLine}:${r.endLine}`;\n\n // Min-max normalize helper\n function minMaxNormalize(items: SearchResult[]): Map<string, number> {\n if (items.length === 0) return new Map();\n const min = Math.min(...items.map((r) => r.score));\n const max = Math.max(...items.map((r) => r.score));\n const range = max - min;\n const m = new Map<string, number>();\n for (const r of items) {\n m.set(keyFor(r), range === 0 ? 1 : (r.score - min) / range);\n }\n return m;\n }\n\n const kwNorm = minMaxNormalize(keywordResults);\n const semNorm = minMaxNormalize(semanticResults);\n\n // Union of all chunk keys\n const allKeys = new Set<string>([\n ...keywordResults.map(keyFor),\n ...semanticResults.map(keyFor),\n ]);\n\n // Build a lookup from key → result metadata\n const metaMap = new Map<string, SearchResult>();\n for (const r of [...keywordResults, ...semanticResults]) {\n metaMap.set(keyFor(r), r);\n }\n\n // Combine scores\n const combined: Array<SearchResult & { combinedScore: number }> = [];\n for (const key of allKeys) {\n const meta = metaMap.get(key)!;\n const kwScore = kwNorm.get(key) ?? 0;\n const semScore = semNorm.get(key) ?? 0;\n const combinedScore = kw * kwScore + sw * semScore;\n combined.push({ ...meta, score: combinedScore, combinedScore });\n }\n\n // Sort by combined score descending\n return combined\n .sort((a, b) => b.score - a.score)\n .slice(0, maxResults)\n .map(({ combinedScore: _unused, ...r }) => r);\n}\n\n// ---------------------------------------------------------------------------\n// Access timestamp tracking (QW2)\n// ---------------------------------------------------------------------------\n\n/**\n * Update last_accessed_at for a set of chunk IDs to the current timestamp.\n *\n * Called after a successful search to record that these chunks were retrieved.\n * This enables the recency boost to account for access patterns, not just\n * modification time.\n *\n * Best-effort: errors are silently ignored so search is never blocked.\n */\nexport function touchChunksLastAccessed(db: Database, chunkIds: string[]): void {\n if (chunkIds.length === 0) return;\n try {\n const now = Date.now();\n const placeholders = chunkIds.map(() => \"?\").join(\", \");\n db.prepare(\n `UPDATE memory_chunks SET last_accessed_at = ? WHERE id IN (${placeholders})`\n ).run(now, ...chunkIds);\n } catch {\n // non-critical — do not block search results\n }\n}\n\n// ---------------------------------------------------------------------------\n// Slug lookup helper\n// ---------------------------------------------------------------------------\n\n/**\n * Populate the projectSlug field on search results by looking up project IDs\n * in the registry database.\n */\nexport function populateSlugs(\n results: SearchResult[],\n registryDb: Database,\n): SearchResult[] {\n if (results.length === 0) return results;\n\n const ids = [...new Set(results.map((r) => r.projectId))];\n const placeholders = ids.map(() => \"?\").join(\", \");\n const rows = registryDb\n .prepare(`SELECT id, slug FROM projects WHERE id IN (${placeholders})`)\n .all(...ids) as Array<{ id: number; slug: string }>;\n\n const slugMap = new Map(rows.map((r) => [r.id, r.slug]));\n\n return results.map((r) => ({\n ...r,\n projectSlug: slugMap.get(r.projectId),\n }));\n}\n\n// ---------------------------------------------------------------------------\n// Recency boost\n// ---------------------------------------------------------------------------\n\n/**\n * Apply exponential recency boost to search scores.\n *\n * Scores are first min-max normalized to [0,1], then multiplied by an\n * exponential decay factor based on chunk age. Normalization is required\n * because the cross-encoder reranker produces negative logit scores — naive\n * multiplication of a negative score by a decay factor (0 < d ≤ 1) would\n * make the score *less* negative, effectively boosting old results instead\n * of penalizing them.\n *\n * Formula: score_final = normalized * exp(-lambda * age_days)\n * where lambda = ln(2) / halfLifeDays, normalized ∈ [0,1]\n *\n * With default halfLifeDays=90, a 3-month-old chunk retains 50% of its\n * normalized score, a 6-month-old retains 25%, and a 1-year-old ~6%.\n *\n * Results without an updatedAt timestamp receive no decay penalty.\n * Results are re-sorted by the boosted score after application.\n *\n * @param results Search results with optional updatedAt timestamps.\n * @param halfLifeDays Score halves every N days. Default 90 (~3 months).\n * @returns New array sorted by decayed normalized score (descending).\n */\nexport function applyRecencyBoost(\n results: SearchResult[],\n halfLifeDays = 90,\n): SearchResult[] {\n if (halfLifeDays <= 0 || results.length === 0) return results;\n\n const lambda = Math.LN2 / halfLifeDays;\n const now = Date.now();\n\n // Min-max normalize scores to [0,1] so multiplicative decay works\n // correctly regardless of the raw score sign/scale.\n const scores = results.map((r) => r.score);\n const minScore = Math.min(...scores);\n const maxScore = Math.max(...scores);\n const range = maxScore - minScore;\n\n return results\n .map((r) => {\n const normalized = range === 0 ? 1 : (r.score - minScore) / range;\n // QW2: use the more recent of updated_at and last_accessed_at for recency decay\n const effectiveTs = r.updatedAt != null && r.lastAccessedAt != null\n ? Math.max(r.updatedAt, r.lastAccessedAt)\n : (r.lastAccessedAt ?? r.updatedAt);\n const decay = effectiveTs\n ? Math.exp(-lambda * Math.max(0, (now - effectiveTs) / 86_400_000))\n : 1; // no timestamp → no penalty\n return { ...r, score: normalized * decay };\n })\n .sort((a, b) => b.score - a.score);\n}\n"],"mappings":";;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;AAwEA,SAAgB,cAAc,OAAuB;CACnD,MAAM,SAAS,MACZ,aAAa,CACb,MAAM,cAAc,CACpB,OAAO,QAAQ,CACf,QAAQ,MAAM,EAAE,UAAU,EAAE,CAC5B,QAAQ,MAAM,CAAC,WAAW,IAAI,EAAE,CAAC,CAEjC,KAAK,MAAM,IAAI,EAAE,QAAQ,MAAM,OAAK,CAAC,GAAG;AAE3C,KAAI,OAAO,WAAW,EAEpB,QAAO,IAAI,MAAM,QAAQ,MAAM,OAAK,CAAC;AAGvC,QAAO,OAAO,KAAK,OAAO;;;;;;;;;;;;;;AAmB5B,SAAgB,aACd,IACA,OACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CACvC,MAAM,WAAW,cAAc,MAAM;CAGrC,MAAM,aAAuB,EAAE;CAC/B,MAAM,SAA8B,CAAC,SAAS;AAE9C,KAAI,MAAM,cAAc,KAAK,WAAW,SAAS,GAAG;EAClD,MAAM,eAAe,KAAK,WAAW,UAAU,IAAI,CAAC,KAAK,KAAK;AAC9D,aAAW,KAAK,oBAAoB,aAAa,GAAG;AACpD,SAAO,KAAK,GAAG,KAAK,WAAW;;AAGjC,KAAI,MAAM,WAAW,KAAK,QAAQ,SAAS,GAAG;EAC5C,MAAM,eAAe,KAAK,QAAQ,UAAU,IAAI,CAAC,KAAK,KAAK;AAC3D,aAAW,KAAK,gBAAgB,aAAa,GAAG;AAChD,SAAO,KAAK,GAAG,KAAK,QAAQ;;AAG9B,KAAI,MAAM,SAAS,KAAK,MAAM,SAAS,GAAG;EACxC,MAAM,eAAe,KAAK,MAAM,UAAU,IAAI,CAAC,KAAK,KAAK;AACzD,aAAW,KAAK,cAAc,aAAa,GAAG;AAC9C,SAAO,KAAK,GAAG,KAAK,MAAM;;CAG5B,MAAM,cAAc,WAAW,SAAS,IACpC,SAAS,WAAW,KAAK,QAAQ,GACjC;AAEJ,QAAO,KAAK,WAAW;CAIvB,MAAM,MAAM;;;;;;;;;;;;;;;;;QAiBN,YAAY;;;;CAKlB,IAAI;AAeJ,KAAI;AACF,SAAO,GAAG,QAAQ,IAAI,CAAC,IAAI,GAAG,OAAO;SAC/B;AAEN,SAAO,EAAE;;CAGX,MAAM,WAAW,MAAM,YAAY;AAEnC,QAAO,KACJ,KAAK,QAAQ;EAKZ,MAAM,QAHY,CAAC,IAAI,cAGI,MADJ,IAAI,mBAAmB;AAE9C,SAAO;GACL,SAAS,IAAI;GACb,WAAW,IAAI;GACf,MAAM,IAAI;GACV,WAAW,IAAI;GACf,SAAS,IAAI;GACb,SAAS,IAAI;GACb;GACA,MAAM,IAAI;GACV,QAAQ,IAAI;GACZ,WAAW,IAAI;GACf,gBAAgB,IAAI,oBAAoB;GACzC;GACD,CACD,QAAQ,MAAM,EAAE,SAAS,SAAS;;;;;;;;;;;AAgBvC,SAAgB,qBACd,IACA,gBACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CAGvC,MAAM,aAAuB,CAAC,wBAAwB;CACtD,MAAM,SAA8B,EAAE;AAEtC,KAAI,MAAM,cAAc,KAAK,WAAW,SAAS,GAAG;EAClD,MAAM,eAAe,KAAK,WAAW,UAAU,IAAI,CAAC,KAAK,KAAK;AAC9D,aAAW,KAAK,kBAAkB,aAAa,GAAG;AAClD,SAAO,KAAK,GAAG,KAAK,WAAW;;AAGjC,KAAI,MAAM,WAAW,KAAK,QAAQ,SAAS,GAAG;EAC5C,MAAM,eAAe,KAAK,QAAQ,UAAU,IAAI,CAAC,KAAK,KAAK;AAC3D,aAAW,KAAK,cAAc,aAAa,GAAG;AAC9C,SAAO,KAAK,GAAG,KAAK,QAAQ;;AAG9B,KAAI,MAAM,SAAS,KAAK,MAAM,SAAS,GAAG;EACxC,MAAM,eAAe,KAAK,MAAM,UAAU,IAAI,CAAC,KAAK,KAAK;AACzD,aAAW,KAAK,YAAY,aAAa,GAAG;AAC5C,SAAO,KAAK,GAAG,KAAK,MAAM;;CAO5B,MAAM,MAAM;;;MAJE,WAAW,WAAW,KAAK,QAAQ,CAOvC;;;CAIV,MAAM,OAAO,GAAG,QAAQ,IAAI,CAAC,IAAI,GAAG,OAAO;AAe3C,KAAI,KAAK,WAAW,EAAG,QAAO,EAAE;CAGhC,MAAM,SAAS,KAAK,KAAK,QAAQ;EAK/B,MAAM,QAHY,iBAAiB,gBADvB,qBAAqB,IAAI,UAAU,CACQ,IAG5B,MADJ,IAAI,mBAAmB;AAE9C,SAAO;GACL,SAAS,IAAI;GACb,WAAW,IAAI;GACf,MAAM,IAAI;GACV,WAAW,IAAI;GACf,SAAS,IAAI;GACb,SAAS,IAAI;GACb;GACA,MAAM,IAAI;GACV,QAAQ,IAAI;GACZ,WAAW,IAAI;GACf,gBAAgB,IAAI,oBAAoB;GACzC;GACD;CAGF,MAAM,WAAW,MAAM,YAAY;AAEnC,QAAO,OACJ,QAAQ,MAAM,EAAE,SAAS,SAAS,CAClC,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM,CACjC,MAAM,GAAG,WAAW;;;;;;;;;;;;AAiBzB,SAAgB,mBACd,IACA,OACA,gBACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CACvC,MAAM,KAAK,MAAM,iBAAiB;CAClC,MAAM,KAAK,MAAM,kBAAkB;CAGnC,MAAM,iBAAiB,aAAa,IAAI,OAAO;EAC7C,GAAG;EACH,YAAY;EACb,CAAC;CAGF,MAAM,kBAAkB,qBAAqB,IAAI,gBAAgB;EAC/D,GAAG;EACH,YAAY;EACb,CAAC;AAEF,KAAI,eAAe,WAAW,KAAK,gBAAgB,WAAW,EAAG,QAAO,EAAE;CAI1E,MAAM,UAAU,MACd,GAAG,EAAE,UAAU,GAAG,EAAE,KAAK,GAAG,EAAE,UAAU,GAAG,EAAE;CAG/C,SAAS,gBAAgB,OAA4C;AACnE,MAAI,MAAM,WAAW,EAAG,wBAAO,IAAI,KAAK;EACxC,MAAM,MAAM,KAAK,IAAI,GAAG,MAAM,KAAK,MAAM,EAAE,MAAM,CAAC;EAElD,MAAM,QADM,KAAK,IAAI,GAAG,MAAM,KAAK,MAAM,EAAE,MAAM,CAAC,GAC9B;EACpB,MAAM,oBAAI,IAAI,KAAqB;AACnC,OAAK,MAAM,KAAK,MACd,GAAE,IAAI,OAAO,EAAE,EAAE,UAAU,IAAI,KAAK,EAAE,QAAQ,OAAO,MAAM;AAE7D,SAAO;;CAGT,MAAM,SAAS,gBAAgB,eAAe;CAC9C,MAAM,UAAU,gBAAgB,gBAAgB;CAGhD,MAAM,UAAU,IAAI,IAAY,CAC9B,GAAG,eAAe,IAAI,OAAO,EAC7B,GAAG,gBAAgB,IAAI,OAAO,CAC/B,CAAC;CAGF,MAAM,0BAAU,IAAI,KAA2B;AAC/C,MAAK,MAAM,KAAK,CAAC,GAAG,gBAAgB,GAAG,gBAAgB,CACrD,SAAQ,IAAI,OAAO,EAAE,EAAE,EAAE;CAI3B,MAAM,WAA4D,EAAE;AACpE,MAAK,MAAM,OAAO,SAAS;EACzB,MAAM,OAAO,QAAQ,IAAI,IAAI;EAC7B,MAAM,UAAU,OAAO,IAAI,IAAI,IAAI;EACnC,MAAM,WAAW,QAAQ,IAAI,IAAI,IAAI;EACrC,MAAM,gBAAgB,KAAK,UAAU,KAAK;AAC1C,WAAS,KAAK;GAAE,GAAG;GAAM,OAAO;GAAe;GAAe,CAAC;;AAIjE,QAAO,SACJ,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM,CACjC,MAAM,GAAG,WAAW,CACpB,KAAK,EAAE,eAAe,SAAS,GAAG,QAAQ,EAAE;;;;;;;;;;;AAgBjD,SAAgB,wBAAwB,IAAc,UAA0B;AAC9E,KAAI,SAAS,WAAW,EAAG;AAC3B,KAAI;EACF,MAAM,MAAM,KAAK,KAAK;EACtB,MAAM,eAAe,SAAS,UAAU,IAAI,CAAC,KAAK,KAAK;AACvD,KAAG,QACD,8DAA8D,aAAa,GAC5E,CAAC,IAAI,KAAK,GAAG,SAAS;SACjB;;;;;;AAaV,SAAgB,cACd,SACA,YACgB;AAChB,KAAI,QAAQ,WAAW,EAAG,QAAO;CAEjC,MAAM,MAAM,CAAC,GAAG,IAAI,IAAI,QAAQ,KAAK,MAAM,EAAE,UAAU,CAAC,CAAC;CACzD,MAAM,eAAe,IAAI,UAAU,IAAI,CAAC,KAAK,KAAK;CAClD,MAAM,OAAO,WACV,QAAQ,8CAA8C,aAAa,GAAG,CACtE,IAAI,GAAG,IAAI;CAEd,MAAM,UAAU,IAAI,IAAI,KAAK,KAAK,MAAM,CAAC,EAAE,IAAI,EAAE,KAAK,CAAC,CAAC;AAExD,QAAO,QAAQ,KAAK,OAAO;EACzB,GAAG;EACH,aAAa,QAAQ,IAAI,EAAE,UAAU;EACtC,EAAE;;;;;;;;;;;;;;;;;;;;;;;;;AA8BL,SAAgB,kBACd,SACA,eAAe,IACC;AAChB,KAAI,gBAAgB,KAAK,QAAQ,WAAW,EAAG,QAAO;CAEtD,MAAM,SAAS,KAAK,MAAM;CAC1B,MAAM,MAAM,KAAK,KAAK;CAItB,MAAM,SAAS,QAAQ,KAAK,MAAM,EAAE,MAAM;CAC1C,MAAM,WAAW,KAAK,IAAI,GAAG,OAAO;CAEpC,MAAM,QADW,KAAK,IAAI,GAAG,OAAO,GACX;AAEzB,QAAO,QACJ,KAAK,MAAM;EACV,MAAM,aAAa,UAAU,IAAI,KAAK,EAAE,QAAQ,YAAY;EAE5D,MAAM,cAAc,EAAE,aAAa,QAAQ,EAAE,kBAAkB,OAC3D,KAAK,IAAI,EAAE,WAAW,EAAE,eAAe,GACtC,EAAE,kBAAkB,EAAE;EAC3B,MAAM,QAAQ,cACV,KAAK,IAAI,CAAC,SAAS,KAAK,IAAI,IAAI,MAAM,eAAe,MAAW,CAAC,GACjE;AACJ,SAAO;GAAE,GAAG;GAAG,OAAO,aAAa;GAAO;GAC1C,CACD,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM"}
|
|
@@ -0,0 +1,298 @@
|
|
|
1
|
+
import { t as __exportAll } from "./rolldown-runtime-95iHPtFO.mjs";
|
|
2
|
+
import { n as cosineSimilarity, r as deserializeEmbedding } from "./embeddings-DGRAPAYb.mjs";
|
|
3
|
+
import { t as STOP_WORDS } from "./stop-words-BaMEGVeY.mjs";
|
|
4
|
+
|
|
5
|
+
//#region src/memory/search.ts
|
|
6
|
+
var search_exports = /* @__PURE__ */ __exportAll({
|
|
7
|
+
applyRecencyBoost: () => applyRecencyBoost,
|
|
8
|
+
buildFtsQuery: () => buildFtsQuery,
|
|
9
|
+
populateSlugs: () => populateSlugs,
|
|
10
|
+
searchMemory: () => searchMemory,
|
|
11
|
+
searchMemoryHybrid: () => searchMemoryHybrid,
|
|
12
|
+
searchMemorySemantic: () => searchMemorySemantic,
|
|
13
|
+
touchChunksLastAccessed: () => touchChunksLastAccessed
|
|
14
|
+
});
|
|
15
|
+
/**
|
|
16
|
+
* Convert a free-text query into an FTS5 query string.
|
|
17
|
+
*
|
|
18
|
+
* Strategy:
|
|
19
|
+
* 1. Tokenise by whitespace and punctuation
|
|
20
|
+
* 2. Remove stop words and tokens shorter than 2 characters
|
|
21
|
+
* 3. Double-quote each remaining token (exact word form)
|
|
22
|
+
* 4. Join with OR so that any matching token returns a result
|
|
23
|
+
*
|
|
24
|
+
* Using OR instead of AND is critical for multi-word queries: the words rarely
|
|
25
|
+
* all appear in the same chunk, so AND would return zero results. FTS5 BM25
|
|
26
|
+
* scoring naturally ranks chunks where more terms match higher, so the most
|
|
27
|
+
* relevant chunks still surface at the top.
|
|
28
|
+
*
|
|
29
|
+
* Example: "Synchrotech interview follow-up Gilles"
|
|
30
|
+
* → `"synchrotech" OR "interview" OR "follow" OR "gilles"`
|
|
31
|
+
* → chunks matching any term, ranked by how many terms match
|
|
32
|
+
*/
|
|
33
|
+
function buildFtsQuery(query) {
|
|
34
|
+
const tokens = query.toLowerCase().split(/[\s\p{P}]+/u).filter(Boolean).filter((t) => t.length >= 2).filter((t) => !STOP_WORDS.has(t)).map((t) => `"${t.replace(/"/g, "\"\"")}"`);
|
|
35
|
+
if (tokens.length === 0) return `"${query.replace(/"/g, "\"\"")}"`;
|
|
36
|
+
return tokens.join(" OR ");
|
|
37
|
+
}
|
|
38
|
+
/**
|
|
39
|
+
* Search across all indexed memory using FTS5 BM25 ranking.
|
|
40
|
+
*
|
|
41
|
+
* Results are ordered by BM25 score (most relevant first).
|
|
42
|
+
* FTS5 bm25() returns negative values; closer to 0 = more relevant.
|
|
43
|
+
* We negate the score so callers get positive values where higher = better.
|
|
44
|
+
*
|
|
45
|
+
* Multilingual note: SQLite FTS5 uses the `unicode61` tokenizer by default,
|
|
46
|
+
* which handles Unicode correctly (German umlauts, French accents, etc.) without
|
|
47
|
+
* language-specific stemming. No changes needed here — it is already
|
|
48
|
+
* multilingual-safe.
|
|
49
|
+
*/
|
|
50
|
+
function searchMemory(db, query, opts) {
|
|
51
|
+
const maxResults = opts?.maxResults ?? 10;
|
|
52
|
+
const ftsQuery = buildFtsQuery(query);
|
|
53
|
+
const conditions = [];
|
|
54
|
+
const params = [ftsQuery];
|
|
55
|
+
if (opts?.projectIds && opts.projectIds.length > 0) {
|
|
56
|
+
const placeholders = opts.projectIds.map(() => "?").join(", ");
|
|
57
|
+
conditions.push(`c.project_id IN (${placeholders})`);
|
|
58
|
+
params.push(...opts.projectIds);
|
|
59
|
+
}
|
|
60
|
+
if (opts?.sources && opts.sources.length > 0) {
|
|
61
|
+
const placeholders = opts.sources.map(() => "?").join(", ");
|
|
62
|
+
conditions.push(`c.source IN (${placeholders})`);
|
|
63
|
+
params.push(...opts.sources);
|
|
64
|
+
}
|
|
65
|
+
if (opts?.tiers && opts.tiers.length > 0) {
|
|
66
|
+
const placeholders = opts.tiers.map(() => "?").join(", ");
|
|
67
|
+
conditions.push(`c.tier IN (${placeholders})`);
|
|
68
|
+
params.push(...opts.tiers);
|
|
69
|
+
}
|
|
70
|
+
const whereClause = conditions.length > 0 ? "AND " + conditions.join(" AND ") : "";
|
|
71
|
+
params.push(maxResults);
|
|
72
|
+
const sql = `
|
|
73
|
+
SELECT
|
|
74
|
+
c.id,
|
|
75
|
+
c.project_id,
|
|
76
|
+
c.path,
|
|
77
|
+
c.start_line,
|
|
78
|
+
c.end_line,
|
|
79
|
+
c.text AS snippet,
|
|
80
|
+
c.tier,
|
|
81
|
+
c.source,
|
|
82
|
+
c.updated_at,
|
|
83
|
+
c.last_accessed_at,
|
|
84
|
+
c.relevance_score,
|
|
85
|
+
bm25(memory_fts) AS bm25_score
|
|
86
|
+
FROM memory_fts
|
|
87
|
+
JOIN memory_chunks c ON memory_fts.id = c.id
|
|
88
|
+
WHERE memory_fts MATCH ?
|
|
89
|
+
${whereClause}
|
|
90
|
+
ORDER BY bm25_score
|
|
91
|
+
LIMIT ?
|
|
92
|
+
`;
|
|
93
|
+
let rows;
|
|
94
|
+
try {
|
|
95
|
+
rows = db.prepare(sql).all(...params);
|
|
96
|
+
} catch {
|
|
97
|
+
return [];
|
|
98
|
+
}
|
|
99
|
+
const minScore = opts?.minScore ?? 0;
|
|
100
|
+
return rows.map((row) => {
|
|
101
|
+
const score = -row.bm25_score * (.5 + (row.relevance_score ?? .5));
|
|
102
|
+
return {
|
|
103
|
+
chunkId: row.id,
|
|
104
|
+
projectId: row.project_id,
|
|
105
|
+
path: row.path,
|
|
106
|
+
startLine: row.start_line,
|
|
107
|
+
endLine: row.end_line,
|
|
108
|
+
snippet: row.snippet,
|
|
109
|
+
score,
|
|
110
|
+
tier: row.tier,
|
|
111
|
+
source: row.source,
|
|
112
|
+
updatedAt: row.updated_at,
|
|
113
|
+
lastAccessedAt: row.last_accessed_at ?? void 0
|
|
114
|
+
};
|
|
115
|
+
}).filter((r) => r.score >= minScore);
|
|
116
|
+
}
|
|
117
|
+
/**
|
|
118
|
+
* Search chunks using brute-force cosine similarity over stored embeddings.
|
|
119
|
+
*
|
|
120
|
+
* Only chunks that have a non-null embedding BLOB are considered. Chunks
|
|
121
|
+
* without embeddings are silently skipped (they can be embedded later via
|
|
122
|
+
* `embedChunks()`).
|
|
123
|
+
*
|
|
124
|
+
* @param queryEmbedding Pre-computed Float32Array for the search query.
|
|
125
|
+
*/
|
|
126
|
+
function searchMemorySemantic(db, queryEmbedding, opts) {
|
|
127
|
+
const maxResults = opts?.maxResults ?? 10;
|
|
128
|
+
const conditions = ["embedding IS NOT NULL"];
|
|
129
|
+
const params = [];
|
|
130
|
+
if (opts?.projectIds && opts.projectIds.length > 0) {
|
|
131
|
+
const placeholders = opts.projectIds.map(() => "?").join(", ");
|
|
132
|
+
conditions.push(`project_id IN (${placeholders})`);
|
|
133
|
+
params.push(...opts.projectIds);
|
|
134
|
+
}
|
|
135
|
+
if (opts?.sources && opts.sources.length > 0) {
|
|
136
|
+
const placeholders = opts.sources.map(() => "?").join(", ");
|
|
137
|
+
conditions.push(`source IN (${placeholders})`);
|
|
138
|
+
params.push(...opts.sources);
|
|
139
|
+
}
|
|
140
|
+
if (opts?.tiers && opts.tiers.length > 0) {
|
|
141
|
+
const placeholders = opts.tiers.map(() => "?").join(", ");
|
|
142
|
+
conditions.push(`tier IN (${placeholders})`);
|
|
143
|
+
params.push(...opts.tiers);
|
|
144
|
+
}
|
|
145
|
+
const sql = `
|
|
146
|
+
SELECT id, project_id, path, start_line, end_line, text, tier, source, embedding, updated_at, last_accessed_at, relevance_score
|
|
147
|
+
FROM memory_chunks
|
|
148
|
+
${"WHERE " + conditions.join(" AND ")}
|
|
149
|
+
LIMIT 5000
|
|
150
|
+
`;
|
|
151
|
+
const rows = db.prepare(sql).all(...params);
|
|
152
|
+
if (rows.length === 0) return [];
|
|
153
|
+
const scored = rows.map((row) => {
|
|
154
|
+
const score = cosineSimilarity(queryEmbedding, deserializeEmbedding(row.embedding)) * (.5 + (row.relevance_score ?? .5));
|
|
155
|
+
return {
|
|
156
|
+
chunkId: row.id,
|
|
157
|
+
projectId: row.project_id,
|
|
158
|
+
path: row.path,
|
|
159
|
+
startLine: row.start_line,
|
|
160
|
+
endLine: row.end_line,
|
|
161
|
+
snippet: row.text,
|
|
162
|
+
score,
|
|
163
|
+
tier: row.tier,
|
|
164
|
+
source: row.source,
|
|
165
|
+
updatedAt: row.updated_at,
|
|
166
|
+
lastAccessedAt: row.last_accessed_at ?? void 0
|
|
167
|
+
};
|
|
168
|
+
});
|
|
169
|
+
const minScore = opts?.minScore ?? -Infinity;
|
|
170
|
+
return scored.filter((r) => r.score >= minScore).sort((a, b) => b.score - a.score).slice(0, maxResults);
|
|
171
|
+
}
|
|
172
|
+
/**
|
|
173
|
+
* Combine BM25 keyword search and semantic search using normalized scores.
|
|
174
|
+
*
|
|
175
|
+
* Both score sets are min-max normalized to [0,1] before combining, so neither
|
|
176
|
+
* dominates the other regardless of their raw scales.
|
|
177
|
+
*
|
|
178
|
+
* @param queryEmbedding Pre-computed embedding for the query.
|
|
179
|
+
* @param keywordWeight Weight for BM25 score (default 0.5).
|
|
180
|
+
* @param semanticWeight Weight for cosine similarity score (default 0.5).
|
|
181
|
+
*/
|
|
182
|
+
function searchMemoryHybrid(db, query, queryEmbedding, opts) {
|
|
183
|
+
const maxResults = opts?.maxResults ?? 10;
|
|
184
|
+
const kw = opts?.keywordWeight ?? .5;
|
|
185
|
+
const sw = opts?.semanticWeight ?? .5;
|
|
186
|
+
const keywordResults = searchMemory(db, query, {
|
|
187
|
+
...opts,
|
|
188
|
+
maxResults: 50
|
|
189
|
+
});
|
|
190
|
+
const semanticResults = searchMemorySemantic(db, queryEmbedding, {
|
|
191
|
+
...opts,
|
|
192
|
+
maxResults: 50
|
|
193
|
+
});
|
|
194
|
+
if (keywordResults.length === 0 && semanticResults.length === 0) return [];
|
|
195
|
+
const keyFor = (r) => `${r.projectId}:${r.path}:${r.startLine}:${r.endLine}`;
|
|
196
|
+
function minMaxNormalize(items) {
|
|
197
|
+
if (items.length === 0) return /* @__PURE__ */ new Map();
|
|
198
|
+
const min = Math.min(...items.map((r) => r.score));
|
|
199
|
+
const range = Math.max(...items.map((r) => r.score)) - min;
|
|
200
|
+
const m = /* @__PURE__ */ new Map();
|
|
201
|
+
for (const r of items) m.set(keyFor(r), range === 0 ? 1 : (r.score - min) / range);
|
|
202
|
+
return m;
|
|
203
|
+
}
|
|
204
|
+
const kwNorm = minMaxNormalize(keywordResults);
|
|
205
|
+
const semNorm = minMaxNormalize(semanticResults);
|
|
206
|
+
const allKeys = new Set([...keywordResults.map(keyFor), ...semanticResults.map(keyFor)]);
|
|
207
|
+
const metaMap = /* @__PURE__ */ new Map();
|
|
208
|
+
for (const r of [...keywordResults, ...semanticResults]) metaMap.set(keyFor(r), r);
|
|
209
|
+
const combined = [];
|
|
210
|
+
for (const key of allKeys) {
|
|
211
|
+
const meta = metaMap.get(key);
|
|
212
|
+
const kwScore = kwNorm.get(key) ?? 0;
|
|
213
|
+
const semScore = semNorm.get(key) ?? 0;
|
|
214
|
+
const combinedScore = kw * kwScore + sw * semScore;
|
|
215
|
+
combined.push({
|
|
216
|
+
...meta,
|
|
217
|
+
score: combinedScore,
|
|
218
|
+
combinedScore
|
|
219
|
+
});
|
|
220
|
+
}
|
|
221
|
+
return combined.sort((a, b) => b.score - a.score).slice(0, maxResults).map(({ combinedScore: _unused, ...r }) => r);
|
|
222
|
+
}
|
|
223
|
+
/**
|
|
224
|
+
* Update last_accessed_at for a set of chunk IDs to the current timestamp.
|
|
225
|
+
*
|
|
226
|
+
* Called after a successful search to record that these chunks were retrieved.
|
|
227
|
+
* This enables the recency boost to account for access patterns, not just
|
|
228
|
+
* modification time.
|
|
229
|
+
*
|
|
230
|
+
* Best-effort: errors are silently ignored so search is never blocked.
|
|
231
|
+
*/
|
|
232
|
+
function touchChunksLastAccessed(db, chunkIds) {
|
|
233
|
+
if (chunkIds.length === 0) return;
|
|
234
|
+
try {
|
|
235
|
+
const now = Date.now();
|
|
236
|
+
const placeholders = chunkIds.map(() => "?").join(", ");
|
|
237
|
+
db.prepare(`UPDATE memory_chunks SET last_accessed_at = ? WHERE id IN (${placeholders})`).run(now, ...chunkIds);
|
|
238
|
+
} catch {}
|
|
239
|
+
}
|
|
240
|
+
/**
|
|
241
|
+
* Populate the projectSlug field on search results by looking up project IDs
|
|
242
|
+
* in the registry database.
|
|
243
|
+
*/
|
|
244
|
+
function populateSlugs(results, registryDb) {
|
|
245
|
+
if (results.length === 0) return results;
|
|
246
|
+
const ids = [...new Set(results.map((r) => r.projectId))];
|
|
247
|
+
const placeholders = ids.map(() => "?").join(", ");
|
|
248
|
+
const rows = registryDb.prepare(`SELECT id, slug FROM projects WHERE id IN (${placeholders})`).all(...ids);
|
|
249
|
+
const slugMap = new Map(rows.map((r) => [r.id, r.slug]));
|
|
250
|
+
return results.map((r) => ({
|
|
251
|
+
...r,
|
|
252
|
+
projectSlug: slugMap.get(r.projectId)
|
|
253
|
+
}));
|
|
254
|
+
}
|
|
255
|
+
/**
|
|
256
|
+
* Apply exponential recency boost to search scores.
|
|
257
|
+
*
|
|
258
|
+
* Scores are first min-max normalized to [0,1], then multiplied by an
|
|
259
|
+
* exponential decay factor based on chunk age. Normalization is required
|
|
260
|
+
* because the cross-encoder reranker produces negative logit scores — naive
|
|
261
|
+
* multiplication of a negative score by a decay factor (0 < d ≤ 1) would
|
|
262
|
+
* make the score *less* negative, effectively boosting old results instead
|
|
263
|
+
* of penalizing them.
|
|
264
|
+
*
|
|
265
|
+
* Formula: score_final = normalized * exp(-lambda * age_days)
|
|
266
|
+
* where lambda = ln(2) / halfLifeDays, normalized ∈ [0,1]
|
|
267
|
+
*
|
|
268
|
+
* With default halfLifeDays=90, a 3-month-old chunk retains 50% of its
|
|
269
|
+
* normalized score, a 6-month-old retains 25%, and a 1-year-old ~6%.
|
|
270
|
+
*
|
|
271
|
+
* Results without an updatedAt timestamp receive no decay penalty.
|
|
272
|
+
* Results are re-sorted by the boosted score after application.
|
|
273
|
+
*
|
|
274
|
+
* @param results Search results with optional updatedAt timestamps.
|
|
275
|
+
* @param halfLifeDays Score halves every N days. Default 90 (~3 months).
|
|
276
|
+
* @returns New array sorted by decayed normalized score (descending).
|
|
277
|
+
*/
|
|
278
|
+
function applyRecencyBoost(results, halfLifeDays = 90) {
|
|
279
|
+
if (halfLifeDays <= 0 || results.length === 0) return results;
|
|
280
|
+
const lambda = Math.LN2 / halfLifeDays;
|
|
281
|
+
const now = Date.now();
|
|
282
|
+
const scores = results.map((r) => r.score);
|
|
283
|
+
const minScore = Math.min(...scores);
|
|
284
|
+
const range = Math.max(...scores) - minScore;
|
|
285
|
+
return results.map((r) => {
|
|
286
|
+
const normalized = range === 0 ? 1 : (r.score - minScore) / range;
|
|
287
|
+
const effectiveTs = r.updatedAt != null && r.lastAccessedAt != null ? Math.max(r.updatedAt, r.lastAccessedAt) : r.lastAccessedAt ?? r.updatedAt;
|
|
288
|
+
const decay = effectiveTs ? Math.exp(-lambda * Math.max(0, (now - effectiveTs) / 864e5)) : 1;
|
|
289
|
+
return {
|
|
290
|
+
...r,
|
|
291
|
+
score: normalized * decay
|
|
292
|
+
};
|
|
293
|
+
}).sort((a, b) => b.score - a.score);
|
|
294
|
+
}
|
|
295
|
+
|
|
296
|
+
//#endregion
|
|
297
|
+
export { searchMemorySemantic as a, searchMemoryHybrid as i, populateSlugs as n, search_exports as o, searchMemory as r, touchChunksLastAccessed as s, buildFtsQuery as t };
|
|
298
|
+
//# sourceMappingURL=search-i2nlQ-JM.mjs.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"search-i2nlQ-JM.mjs","names":[],"sources":["../src/memory/search.ts"],"sourcesContent":["/**\n * Search over the PAI federation memory index.\n *\n * Provides three search modes:\n * - keyword — BM25 full-text search (default, fast, no ML required)\n * - semantic — Brute-force cosine similarity over pre-computed embeddings\n * - hybrid — Normalized combination of BM25 + cosine scores\n *\n * BM25 uses SQLite's FTS5 extension. Semantic search requires embeddings to\n * have been generated first via `embedChunks()` in the indexer.\n */\n\nimport type { Database } from \"better-sqlite3\";\nimport { deserializeEmbedding, cosineSimilarity } from \"./embeddings.js\";\nimport { STOP_WORDS } from \"../utils/stop-words.js\";\n\n// ---------------------------------------------------------------------------\n// Types\n// ---------------------------------------------------------------------------\n\nexport interface SearchResult {\n projectId: number;\n projectSlug?: string; // populated from registry after search when available\n path: string;\n startLine: number;\n endLine: number;\n snippet: string;\n score: number; // raw BM25 score (lower = more relevant in FTS5)\n tier: string;\n source: string;\n updatedAt?: number; // Unix ms from memory_chunks.updated_at\n lastAccessedAt?: number; // Unix ms from memory_chunks.last_accessed_at (QW2)\n chunkId?: string; // chunk ID for last_accessed_at update (QW2)\n}\n\nexport interface SearchOptions {\n /** Restrict search to these project IDs. */\n projectIds?: number[];\n /** Restrict to 'memory' or 'notes' sources. */\n sources?: string[];\n /** Restrict to specific tier(s): 'evergreen' | 'daily' | 'topic' | 'session' */\n tiers?: string[];\n /** Maximum number of results to return. Default 10. */\n maxResults?: number;\n /** Minimum BM25 score threshold (FTS5 scores are negative; 0.0 means no filter). */\n minScore?: number;\n}\n\n// STOP_WORDS imported from utils/stop-words.ts\n\n// ---------------------------------------------------------------------------\n// Query builder\n// ---------------------------------------------------------------------------\n\n/**\n * Convert a free-text query into an FTS5 query string.\n *\n * Strategy:\n * 1. Tokenise by whitespace and punctuation\n * 2. Remove stop words and tokens shorter than 2 characters\n * 3. Double-quote each remaining token (exact word form)\n * 4. Join with OR so that any matching token returns a result\n *\n * Using OR instead of AND is critical for multi-word queries: the words rarely\n * all appear in the same chunk, so AND would return zero results. FTS5 BM25\n * scoring naturally ranks chunks where more terms match higher, so the most\n * relevant chunks still surface at the top.\n *\n * Example: \"Synchrotech interview follow-up Gilles\"\n * → `\"synchrotech\" OR \"interview\" OR \"follow\" OR \"gilles\"`\n * → chunks matching any term, ranked by how many terms match\n */\nexport function buildFtsQuery(query: string): string {\n const tokens = query\n .toLowerCase()\n .split(/[\\s\\p{P}]+/u)\n .filter(Boolean)\n .filter((t) => t.length >= 2)\n .filter((t) => !STOP_WORDS.has(t))\n // Escape any double-quotes inside the token (FTS5 uses them as delimiters)\n .map((t) => `\"${t.replace(/\"/g, '\"\"')}\"`)\n\n if (tokens.length === 0) {\n // Fallback: use original query as a raw string (may produce no results)\n return `\"${query.replace(/\"/g, '\"\"')}\"`;\n }\n\n return tokens.join(\" OR \");\n}\n\n// ---------------------------------------------------------------------------\n// Search\n// ---------------------------------------------------------------------------\n\n/**\n * Search across all indexed memory using FTS5 BM25 ranking.\n *\n * Results are ordered by BM25 score (most relevant first).\n * FTS5 bm25() returns negative values; closer to 0 = more relevant.\n * We negate the score so callers get positive values where higher = better.\n *\n * Multilingual note: SQLite FTS5 uses the `unicode61` tokenizer by default,\n * which handles Unicode correctly (German umlauts, French accents, etc.) without\n * language-specific stemming. No changes needed here — it is already\n * multilingual-safe.\n */\nexport function searchMemory(\n db: Database,\n query: string,\n opts?: SearchOptions,\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n const ftsQuery = buildFtsQuery(query);\n\n // Build the SQL with optional filters\n const conditions: string[] = [];\n const params: (string | number)[] = [ftsQuery];\n\n if (opts?.projectIds && opts.projectIds.length > 0) {\n const placeholders = opts.projectIds.map(() => \"?\").join(\", \");\n conditions.push(`c.project_id IN (${placeholders})`);\n params.push(...opts.projectIds);\n }\n\n if (opts?.sources && opts.sources.length > 0) {\n const placeholders = opts.sources.map(() => \"?\").join(\", \");\n conditions.push(`c.source IN (${placeholders})`);\n params.push(...opts.sources);\n }\n\n if (opts?.tiers && opts.tiers.length > 0) {\n const placeholders = opts.tiers.map(() => \"?\").join(\", \");\n conditions.push(`c.tier IN (${placeholders})`);\n params.push(...opts.tiers);\n }\n\n const whereClause = conditions.length > 0\n ? \"AND \" + conditions.join(\" AND \")\n : \"\";\n\n params.push(maxResults);\n\n // FTS5: join memory_fts with memory_chunks to get metadata\n // bm25(memory_fts) returns negative values (lower = better match)\n const sql = `\n SELECT\n c.id,\n c.project_id,\n c.path,\n c.start_line,\n c.end_line,\n c.text AS snippet,\n c.tier,\n c.source,\n c.updated_at,\n c.last_accessed_at,\n c.relevance_score,\n bm25(memory_fts) AS bm25_score\n FROM memory_fts\n JOIN memory_chunks c ON memory_fts.id = c.id\n WHERE memory_fts MATCH ?\n ${whereClause}\n ORDER BY bm25_score\n LIMIT ?\n `;\n\n let rows: Array<{\n id: string;\n project_id: number;\n path: string;\n start_line: number;\n end_line: number;\n snippet: string;\n tier: string;\n source: string;\n updated_at: number;\n last_accessed_at: number | null;\n relevance_score: number | null;\n bm25_score: number;\n }>;\n\n try {\n rows = db.prepare(sql).all(...params) as typeof rows;\n } catch {\n // FTS5 MATCH throws when the query is invalid — return empty results\n return [];\n }\n\n const minScore = opts?.minScore ?? 0.0;\n\n return rows\n .map((row) => {\n // Negate so higher = better match for callers\n const baseScore = -row.bm25_score;\n // MR2: scale by feedback relevance_score: multiplier in [0.5, 1.5]\n const relevanceScore = row.relevance_score ?? 0.5;\n const score = baseScore * (0.5 + relevanceScore);\n return {\n chunkId: row.id,\n projectId: row.project_id,\n path: row.path,\n startLine: row.start_line,\n endLine: row.end_line,\n snippet: row.snippet,\n score,\n tier: row.tier,\n source: row.source,\n updatedAt: row.updated_at,\n lastAccessedAt: row.last_accessed_at ?? undefined,\n };\n })\n .filter((r) => r.score >= minScore);\n}\n\n// ---------------------------------------------------------------------------\n// Semantic search\n// ---------------------------------------------------------------------------\n\n/**\n * Search chunks using brute-force cosine similarity over stored embeddings.\n *\n * Only chunks that have a non-null embedding BLOB are considered. Chunks\n * without embeddings are silently skipped (they can be embedded later via\n * `embedChunks()`).\n *\n * @param queryEmbedding Pre-computed Float32Array for the search query.\n */\nexport function searchMemorySemantic(\n db: Database,\n queryEmbedding: Float32Array,\n opts?: SearchOptions,\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n\n // Build the SQL filter conditions\n const conditions: string[] = [\"embedding IS NOT NULL\"];\n const params: (string | number)[] = [];\n\n if (opts?.projectIds && opts.projectIds.length > 0) {\n const placeholders = opts.projectIds.map(() => \"?\").join(\", \");\n conditions.push(`project_id IN (${placeholders})`);\n params.push(...opts.projectIds);\n }\n\n if (opts?.sources && opts.sources.length > 0) {\n const placeholders = opts.sources.map(() => \"?\").join(\", \");\n conditions.push(`source IN (${placeholders})`);\n params.push(...opts.sources);\n }\n\n if (opts?.tiers && opts.tiers.length > 0) {\n const placeholders = opts.tiers.map(() => \"?\").join(\", \");\n conditions.push(`tier IN (${placeholders})`);\n params.push(...opts.tiers);\n }\n\n const where = \"WHERE \" + conditions.join(\" AND \");\n\n // Hard cap for SQLite semantic path — prevents OOM on large corpora.\n // Use Postgres for production semantic search.\n const sql = `\n SELECT id, project_id, path, start_line, end_line, text, tier, source, embedding, updated_at, last_accessed_at, relevance_score\n FROM memory_chunks\n ${where}\n LIMIT 5000\n `;\n\n const rows = db.prepare(sql).all(...params) as Array<{\n id: string;\n project_id: number;\n path: string;\n start_line: number;\n end_line: number;\n text: string;\n tier: string;\n source: string;\n embedding: Buffer;\n updated_at: number;\n last_accessed_at: number | null;\n relevance_score: number | null;\n }>;\n\n if (rows.length === 0) return [];\n\n // Compute cosine similarity for every chunk\n const scored = rows.map((row) => {\n const vec = deserializeEmbedding(row.embedding);\n const baseScore = cosineSimilarity(queryEmbedding, vec);\n // MR2: scale by feedback relevance_score: multiplier in [0.5, 1.5]\n const relevanceScore = row.relevance_score ?? 0.5;\n const score = baseScore * (0.5 + relevanceScore);\n return {\n chunkId: row.id,\n projectId: row.project_id,\n path: row.path,\n startLine: row.start_line,\n endLine: row.end_line,\n snippet: row.text,\n score,\n tier: row.tier,\n source: row.source,\n updatedAt: row.updated_at,\n lastAccessedAt: row.last_accessed_at ?? undefined,\n };\n });\n\n // Sort by descending similarity, apply optional min score filter, limit\n const minScore = opts?.minScore ?? -Infinity;\n\n return scored\n .filter((r) => r.score >= minScore)\n .sort((a, b) => b.score - a.score)\n .slice(0, maxResults);\n}\n\n// ---------------------------------------------------------------------------\n// Hybrid search\n// ---------------------------------------------------------------------------\n\n/**\n * Combine BM25 keyword search and semantic search using normalized scores.\n *\n * Both score sets are min-max normalized to [0,1] before combining, so neither\n * dominates the other regardless of their raw scales.\n *\n * @param queryEmbedding Pre-computed embedding for the query.\n * @param keywordWeight Weight for BM25 score (default 0.5).\n * @param semanticWeight Weight for cosine similarity score (default 0.5).\n */\nexport function searchMemoryHybrid(\n db: Database,\n query: string,\n queryEmbedding: Float32Array,\n opts?: SearchOptions & { keywordWeight?: number; semanticWeight?: number },\n): SearchResult[] {\n const maxResults = opts?.maxResults ?? 10;\n const kw = opts?.keywordWeight ?? 0.5;\n const sw = opts?.semanticWeight ?? 0.5;\n\n // Fetch keyword results — 50 candidates is sufficient for min-max normalization\n const keywordResults = searchMemory(db, query, {\n ...opts,\n maxResults: 50,\n });\n\n // Fetch semantic results — 50 candidates is sufficient for min-max normalization\n const semanticResults = searchMemorySemantic(db, queryEmbedding, {\n ...opts,\n maxResults: 50,\n });\n\n if (keywordResults.length === 0 && semanticResults.length === 0) return [];\n\n // Build a map of chunk ID → combined result\n // Use \"projectId:path:startLine:endLine\" as a stable key (same as chunk IDs)\n const keyFor = (r: SearchResult) =>\n `${r.projectId}:${r.path}:${r.startLine}:${r.endLine}`;\n\n // Min-max normalize helper\n function minMaxNormalize(items: SearchResult[]): Map<string, number> {\n if (items.length === 0) return new Map();\n const min = Math.min(...items.map((r) => r.score));\n const max = Math.max(...items.map((r) => r.score));\n const range = max - min;\n const m = new Map<string, number>();\n for (const r of items) {\n m.set(keyFor(r), range === 0 ? 1 : (r.score - min) / range);\n }\n return m;\n }\n\n const kwNorm = minMaxNormalize(keywordResults);\n const semNorm = minMaxNormalize(semanticResults);\n\n // Union of all chunk keys\n const allKeys = new Set<string>([\n ...keywordResults.map(keyFor),\n ...semanticResults.map(keyFor),\n ]);\n\n // Build a lookup from key → result metadata\n const metaMap = new Map<string, SearchResult>();\n for (const r of [...keywordResults, ...semanticResults]) {\n metaMap.set(keyFor(r), r);\n }\n\n // Combine scores\n const combined: Array<SearchResult & { combinedScore: number }> = [];\n for (const key of allKeys) {\n const meta = metaMap.get(key)!;\n const kwScore = kwNorm.get(key) ?? 0;\n const semScore = semNorm.get(key) ?? 0;\n const combinedScore = kw * kwScore + sw * semScore;\n combined.push({ ...meta, score: combinedScore, combinedScore });\n }\n\n // Sort by combined score descending\n return combined\n .sort((a, b) => b.score - a.score)\n .slice(0, maxResults)\n .map(({ combinedScore: _unused, ...r }) => r);\n}\n\n// ---------------------------------------------------------------------------\n// Access timestamp tracking (QW2)\n// ---------------------------------------------------------------------------\n\n/**\n * Update last_accessed_at for a set of chunk IDs to the current timestamp.\n *\n * Called after a successful search to record that these chunks were retrieved.\n * This enables the recency boost to account for access patterns, not just\n * modification time.\n *\n * Best-effort: errors are silently ignored so search is never blocked.\n */\nexport function touchChunksLastAccessed(db: Database, chunkIds: string[]): void {\n if (chunkIds.length === 0) return;\n try {\n const now = Date.now();\n const placeholders = chunkIds.map(() => \"?\").join(\", \");\n db.prepare(\n `UPDATE memory_chunks SET last_accessed_at = ? WHERE id IN (${placeholders})`\n ).run(now, ...chunkIds);\n } catch {\n // non-critical — do not block search results\n }\n}\n\n// ---------------------------------------------------------------------------\n// Slug lookup helper\n// ---------------------------------------------------------------------------\n\n/**\n * Populate the projectSlug field on search results by looking up project IDs\n * in the registry database.\n */\nexport function populateSlugs(\n results: SearchResult[],\n registryDb: Database,\n): SearchResult[] {\n if (results.length === 0) return results;\n\n const ids = [...new Set(results.map((r) => r.projectId))];\n const placeholders = ids.map(() => \"?\").join(\", \");\n const rows = registryDb\n .prepare(`SELECT id, slug FROM projects WHERE id IN (${placeholders})`)\n .all(...ids) as Array<{ id: number; slug: string }>;\n\n const slugMap = new Map(rows.map((r) => [r.id, r.slug]));\n\n return results.map((r) => ({\n ...r,\n projectSlug: slugMap.get(r.projectId),\n }));\n}\n\n// ---------------------------------------------------------------------------\n// Recency boost\n// ---------------------------------------------------------------------------\n\n/**\n * Apply exponential recency boost to search scores.\n *\n * Scores are first min-max normalized to [0,1], then multiplied by an\n * exponential decay factor based on chunk age. Normalization is required\n * because the cross-encoder reranker produces negative logit scores — naive\n * multiplication of a negative score by a decay factor (0 < d ≤ 1) would\n * make the score *less* negative, effectively boosting old results instead\n * of penalizing them.\n *\n * Formula: score_final = normalized * exp(-lambda * age_days)\n * where lambda = ln(2) / halfLifeDays, normalized ∈ [0,1]\n *\n * With default halfLifeDays=90, a 3-month-old chunk retains 50% of its\n * normalized score, a 6-month-old retains 25%, and a 1-year-old ~6%.\n *\n * Results without an updatedAt timestamp receive no decay penalty.\n * Results are re-sorted by the boosted score after application.\n *\n * @param results Search results with optional updatedAt timestamps.\n * @param halfLifeDays Score halves every N days. Default 90 (~3 months).\n * @returns New array sorted by decayed normalized score (descending).\n */\nexport function applyRecencyBoost(\n results: SearchResult[],\n halfLifeDays = 90,\n): SearchResult[] {\n if (halfLifeDays <= 0 || results.length === 0) return results;\n\n const lambda = Math.LN2 / halfLifeDays;\n const now = Date.now();\n\n // Min-max normalize scores to [0,1] so multiplicative decay works\n // correctly regardless of the raw score sign/scale.\n const scores = results.map((r) => r.score);\n const minScore = Math.min(...scores);\n const maxScore = Math.max(...scores);\n const range = maxScore - minScore;\n\n return results\n .map((r) => {\n const normalized = range === 0 ? 1 : (r.score - minScore) / range;\n // QW2: use the more recent of updated_at and last_accessed_at for recency decay\n const effectiveTs = r.updatedAt != null && r.lastAccessedAt != null\n ? Math.max(r.updatedAt, r.lastAccessedAt)\n : (r.lastAccessedAt ?? r.updatedAt);\n const decay = effectiveTs\n ? Math.exp(-lambda * Math.max(0, (now - effectiveTs) / 86_400_000))\n : 1; // no timestamp → no penalty\n return { ...r, score: normalized * decay };\n })\n .sort((a, b) => b.score - a.score);\n}\n"],"mappings":";;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;AAwEA,SAAgB,cAAc,OAAuB;CACnD,MAAM,SAAS,MACZ,aAAa,CACb,MAAM,cAAc,CACpB,OAAO,QAAQ,CACf,QAAQ,MAAM,EAAE,UAAU,EAAE,CAC5B,QAAQ,MAAM,CAAC,WAAW,IAAI,EAAE,CAAC,CAEjC,KAAK,MAAM,IAAI,EAAE,QAAQ,MAAM,OAAK,CAAC,GAAG;AAE3C,KAAI,OAAO,WAAW,EAEpB,QAAO,IAAI,MAAM,QAAQ,MAAM,OAAK,CAAC;AAGvC,QAAO,OAAO,KAAK,OAAO;;;;;;;;;;;;;;AAmB5B,SAAgB,aACd,IACA,OACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CACvC,MAAM,WAAW,cAAc,MAAM;CAGrC,MAAM,aAAuB,EAAE;CAC/B,MAAM,SAA8B,CAAC,SAAS;AAE9C,KAAI,MAAM,cAAc,KAAK,WAAW,SAAS,GAAG;EAClD,MAAM,eAAe,KAAK,WAAW,UAAU,IAAI,CAAC,KAAK,KAAK;AAC9D,aAAW,KAAK,oBAAoB,aAAa,GAAG;AACpD,SAAO,KAAK,GAAG,KAAK,WAAW;;AAGjC,KAAI,MAAM,WAAW,KAAK,QAAQ,SAAS,GAAG;EAC5C,MAAM,eAAe,KAAK,QAAQ,UAAU,IAAI,CAAC,KAAK,KAAK;AAC3D,aAAW,KAAK,gBAAgB,aAAa,GAAG;AAChD,SAAO,KAAK,GAAG,KAAK,QAAQ;;AAG9B,KAAI,MAAM,SAAS,KAAK,MAAM,SAAS,GAAG;EACxC,MAAM,eAAe,KAAK,MAAM,UAAU,IAAI,CAAC,KAAK,KAAK;AACzD,aAAW,KAAK,cAAc,aAAa,GAAG;AAC9C,SAAO,KAAK,GAAG,KAAK,MAAM;;CAG5B,MAAM,cAAc,WAAW,SAAS,IACpC,SAAS,WAAW,KAAK,QAAQ,GACjC;AAEJ,QAAO,KAAK,WAAW;CAIvB,MAAM,MAAM;;;;;;;;;;;;;;;;;QAiBN,YAAY;;;;CAKlB,IAAI;AAeJ,KAAI;AACF,SAAO,GAAG,QAAQ,IAAI,CAAC,IAAI,GAAG,OAAO;SAC/B;AAEN,SAAO,EAAE;;CAGX,MAAM,WAAW,MAAM,YAAY;AAEnC,QAAO,KACJ,KAAK,QAAQ;EAKZ,MAAM,QAHY,CAAC,IAAI,cAGI,MADJ,IAAI,mBAAmB;AAE9C,SAAO;GACL,SAAS,IAAI;GACb,WAAW,IAAI;GACf,MAAM,IAAI;GACV,WAAW,IAAI;GACf,SAAS,IAAI;GACb,SAAS,IAAI;GACb;GACA,MAAM,IAAI;GACV,QAAQ,IAAI;GACZ,WAAW,IAAI;GACf,gBAAgB,IAAI,oBAAoB;GACzC;GACD,CACD,QAAQ,MAAM,EAAE,SAAS,SAAS;;;;;;;;;;;AAgBvC,SAAgB,qBACd,IACA,gBACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CAGvC,MAAM,aAAuB,CAAC,wBAAwB;CACtD,MAAM,SAA8B,EAAE;AAEtC,KAAI,MAAM,cAAc,KAAK,WAAW,SAAS,GAAG;EAClD,MAAM,eAAe,KAAK,WAAW,UAAU,IAAI,CAAC,KAAK,KAAK;AAC9D,aAAW,KAAK,kBAAkB,aAAa,GAAG;AAClD,SAAO,KAAK,GAAG,KAAK,WAAW;;AAGjC,KAAI,MAAM,WAAW,KAAK,QAAQ,SAAS,GAAG;EAC5C,MAAM,eAAe,KAAK,QAAQ,UAAU,IAAI,CAAC,KAAK,KAAK;AAC3D,aAAW,KAAK,cAAc,aAAa,GAAG;AAC9C,SAAO,KAAK,GAAG,KAAK,QAAQ;;AAG9B,KAAI,MAAM,SAAS,KAAK,MAAM,SAAS,GAAG;EACxC,MAAM,eAAe,KAAK,MAAM,UAAU,IAAI,CAAC,KAAK,KAAK;AACzD,aAAW,KAAK,YAAY,aAAa,GAAG;AAC5C,SAAO,KAAK,GAAG,KAAK,MAAM;;CAO5B,MAAM,MAAM;;;MAJE,WAAW,WAAW,KAAK,QAAQ,CAOvC;;;CAIV,MAAM,OAAO,GAAG,QAAQ,IAAI,CAAC,IAAI,GAAG,OAAO;AAe3C,KAAI,KAAK,WAAW,EAAG,QAAO,EAAE;CAGhC,MAAM,SAAS,KAAK,KAAK,QAAQ;EAK/B,MAAM,QAHY,iBAAiB,gBADvB,qBAAqB,IAAI,UAAU,CACQ,IAG5B,MADJ,IAAI,mBAAmB;AAE9C,SAAO;GACL,SAAS,IAAI;GACb,WAAW,IAAI;GACf,MAAM,IAAI;GACV,WAAW,IAAI;GACf,SAAS,IAAI;GACb,SAAS,IAAI;GACb;GACA,MAAM,IAAI;GACV,QAAQ,IAAI;GACZ,WAAW,IAAI;GACf,gBAAgB,IAAI,oBAAoB;GACzC;GACD;CAGF,MAAM,WAAW,MAAM,YAAY;AAEnC,QAAO,OACJ,QAAQ,MAAM,EAAE,SAAS,SAAS,CAClC,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM,CACjC,MAAM,GAAG,WAAW;;;;;;;;;;;;AAiBzB,SAAgB,mBACd,IACA,OACA,gBACA,MACgB;CAChB,MAAM,aAAa,MAAM,cAAc;CACvC,MAAM,KAAK,MAAM,iBAAiB;CAClC,MAAM,KAAK,MAAM,kBAAkB;CAGnC,MAAM,iBAAiB,aAAa,IAAI,OAAO;EAC7C,GAAG;EACH,YAAY;EACb,CAAC;CAGF,MAAM,kBAAkB,qBAAqB,IAAI,gBAAgB;EAC/D,GAAG;EACH,YAAY;EACb,CAAC;AAEF,KAAI,eAAe,WAAW,KAAK,gBAAgB,WAAW,EAAG,QAAO,EAAE;CAI1E,MAAM,UAAU,MACd,GAAG,EAAE,UAAU,GAAG,EAAE,KAAK,GAAG,EAAE,UAAU,GAAG,EAAE;CAG/C,SAAS,gBAAgB,OAA4C;AACnE,MAAI,MAAM,WAAW,EAAG,wBAAO,IAAI,KAAK;EACxC,MAAM,MAAM,KAAK,IAAI,GAAG,MAAM,KAAK,MAAM,EAAE,MAAM,CAAC;EAElD,MAAM,QADM,KAAK,IAAI,GAAG,MAAM,KAAK,MAAM,EAAE,MAAM,CAAC,GAC9B;EACpB,MAAM,oBAAI,IAAI,KAAqB;AACnC,OAAK,MAAM,KAAK,MACd,GAAE,IAAI,OAAO,EAAE,EAAE,UAAU,IAAI,KAAK,EAAE,QAAQ,OAAO,MAAM;AAE7D,SAAO;;CAGT,MAAM,SAAS,gBAAgB,eAAe;CAC9C,MAAM,UAAU,gBAAgB,gBAAgB;CAGhD,MAAM,UAAU,IAAI,IAAY,CAC9B,GAAG,eAAe,IAAI,OAAO,EAC7B,GAAG,gBAAgB,IAAI,OAAO,CAC/B,CAAC;CAGF,MAAM,0BAAU,IAAI,KAA2B;AAC/C,MAAK,MAAM,KAAK,CAAC,GAAG,gBAAgB,GAAG,gBAAgB,CACrD,SAAQ,IAAI,OAAO,EAAE,EAAE,EAAE;CAI3B,MAAM,WAA4D,EAAE;AACpE,MAAK,MAAM,OAAO,SAAS;EACzB,MAAM,OAAO,QAAQ,IAAI,IAAI;EAC7B,MAAM,UAAU,OAAO,IAAI,IAAI,IAAI;EACnC,MAAM,WAAW,QAAQ,IAAI,IAAI,IAAI;EACrC,MAAM,gBAAgB,KAAK,UAAU,KAAK;AAC1C,WAAS,KAAK;GAAE,GAAG;GAAM,OAAO;GAAe;GAAe,CAAC;;AAIjE,QAAO,SACJ,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM,CACjC,MAAM,GAAG,WAAW,CACpB,KAAK,EAAE,eAAe,SAAS,GAAG,QAAQ,EAAE;;;;;;;;;;;AAgBjD,SAAgB,wBAAwB,IAAc,UAA0B;AAC9E,KAAI,SAAS,WAAW,EAAG;AAC3B,KAAI;EACF,MAAM,MAAM,KAAK,KAAK;EACtB,MAAM,eAAe,SAAS,UAAU,IAAI,CAAC,KAAK,KAAK;AACvD,KAAG,QACD,8DAA8D,aAAa,GAC5E,CAAC,IAAI,KAAK,GAAG,SAAS;SACjB;;;;;;AAaV,SAAgB,cACd,SACA,YACgB;AAChB,KAAI,QAAQ,WAAW,EAAG,QAAO;CAEjC,MAAM,MAAM,CAAC,GAAG,IAAI,IAAI,QAAQ,KAAK,MAAM,EAAE,UAAU,CAAC,CAAC;CACzD,MAAM,eAAe,IAAI,UAAU,IAAI,CAAC,KAAK,KAAK;CAClD,MAAM,OAAO,WACV,QAAQ,8CAA8C,aAAa,GAAG,CACtE,IAAI,GAAG,IAAI;CAEd,MAAM,UAAU,IAAI,IAAI,KAAK,KAAK,MAAM,CAAC,EAAE,IAAI,EAAE,KAAK,CAAC,CAAC;AAExD,QAAO,QAAQ,KAAK,OAAO;EACzB,GAAG;EACH,aAAa,QAAQ,IAAI,EAAE,UAAU;EACtC,EAAE;;;;;;;;;;;;;;;;;;;;;;;;;AA8BL,SAAgB,kBACd,SACA,eAAe,IACC;AAChB,KAAI,gBAAgB,KAAK,QAAQ,WAAW,EAAG,QAAO;CAEtD,MAAM,SAAS,KAAK,MAAM;CAC1B,MAAM,MAAM,KAAK,KAAK;CAItB,MAAM,SAAS,QAAQ,KAAK,MAAM,EAAE,MAAM;CAC1C,MAAM,WAAW,KAAK,IAAI,GAAG,OAAO;CAEpC,MAAM,QADW,KAAK,IAAI,GAAG,OAAO,GACX;AAEzB,QAAO,QACJ,KAAK,MAAM;EACV,MAAM,aAAa,UAAU,IAAI,KAAK,EAAE,QAAQ,YAAY;EAE5D,MAAM,cAAc,EAAE,aAAa,QAAQ,EAAE,kBAAkB,OAC3D,KAAK,IAAI,EAAE,WAAW,EAAE,eAAe,GACtC,EAAE,kBAAkB,EAAE;EAC3B,MAAM,QAAQ,cACV,KAAK,IAAI,CAAC,SAAS,KAAK,IAAI,IAAI,MAAM,eAAe,MAAW,CAAC,GACjE;AACJ,SAAO;GAAE,GAAG;GAAG,OAAO,aAAa;GAAO;GAC1C,CACD,MAAM,GAAG,MAAM,EAAE,QAAQ,EAAE,MAAM"}
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: End
|
|
3
|
+
description: "Finalize a session: checkpoint, mark note completed, commit pending work if asked, then exit safely. USE WHEN user says /end, end session, finish session, OR is done with this conversation entirely."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## End Skill
|
|
7
|
+
|
|
8
|
+
USE WHEN user says /end, end session, finish session, OR is done with this conversation entirely.
|
|
9
|
+
|
|
10
|
+
### The rule that matters
|
|
11
|
+
|
|
12
|
+
**The checkpoint must be written to a FILE before it is printed.** Printing it to the
|
|
13
|
+
terminal does not persist it — if the session ends, it is gone and the user has to copy it
|
|
14
|
+
off the screen by hand. Order is **compose → write file → persist via CLI → print**, and it
|
|
15
|
+
never varies.
|
|
16
|
+
|
|
17
|
+
### Procedure (extends Pause)
|
|
18
|
+
|
|
19
|
+
1. **Compose the checkpoint** — what was accomplished, what was left incomplete, open
|
|
20
|
+
decisions, anything built but not installed, and any follow-up actions. Write it for a
|
|
21
|
+
session that has none of your context.
|
|
22
|
+
|
|
23
|
+
1b. **Write it to a file** with the Write tool:
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
/tmp/pai-checkpoint-<session-id>.md
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Use the `sessionId` from the startup system reminders. Body markdown only — no
|
|
30
|
+
`## Continue` or `## Session Complete` heading; PAI adds those.
|
|
31
|
+
|
|
32
|
+
2. **Run `pai end`** via Bash:
|
|
33
|
+
|
|
34
|
+
```bash
|
|
35
|
+
pai end --body-file /tmp/pai-checkpoint-<session-id>.md --session-id <session-id>
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
This does everything `pai pause` does (writes the checkpoint into `## Continue` in
|
|
39
|
+
TODO.md and appends it to the session note) PLUS marks the current session note
|
|
40
|
+
**Status: Completed** with a timestamp.
|
|
41
|
+
|
|
42
|
+
**If this errors, stop and fix it.** It refuses to write a metadata-only checkpoint, so
|
|
43
|
+
an error here is the difference between a real checkpoint and a lost one.
|
|
44
|
+
|
|
45
|
+
2b. **File open items onto the task bus.**
|
|
46
|
+
|
|
47
|
+
`pai end` writes the session note and TODO.md. It does **not** touch Todoist — it is a
|
|
48
|
+
mechanical command, and deciding what counts as an open item needs judgement. So do it
|
|
49
|
+
here, explicitly:
|
|
50
|
+
|
|
51
|
+
```bash
|
|
52
|
+
pai task add "<open item>" --into <ProjectName> --owner <project> --body "<reasoning>"
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
**Always use `--into <ProjectName>`** — one sub-project per PAI project, created
|
|
56
|
+
automatically if missing. Doing so is *following the convention, not inventing
|
|
57
|
+
structure*: filing flat into the root out of caution buries findings across projects,
|
|
58
|
+
which is the mess the convention prevents. Omit `--into` only for something genuinely
|
|
59
|
+
cross-cutting.
|
|
60
|
+
|
|
61
|
+
Skip this only if nothing is actually open — say so rather than filing filler. Otherwise
|
|
62
|
+
the open items exist solely in a session note nobody re-reads, which is exactly the loss
|
|
63
|
+
the bus was built to stop.
|
|
64
|
+
|
|
65
|
+
3. **Check for uncommitted changes** with `git status`. If there are any, ask whether to
|
|
66
|
+
commit them. If yes, use clean conventional-commit format — no AI signatures, no
|
|
67
|
+
`--no-verify`.
|
|
68
|
+
|
|
69
|
+
4. **Print the handoff block** — the same content you wrote in step 1b, plus the session ID
|
|
70
|
+
from the `sessionId` field in startup system reminders, and:
|
|
71
|
+
|
|
72
|
+
```
|
|
73
|
+
To resume (if needed): claude --resume <uuid>
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
5. **Tell the user:**
|
|
77
|
+
"Now type `/exit` to safely close the session. DO NOT press Ctrl+C — that bypasses PAI
|
|
78
|
+
stop-hook and means the stop-hook finalization never runs."
|
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: Pause
|
|
3
|
+
description: "Save a checkpoint and prepare to exit safely, knowing you will come back to the same conversation. USE WHEN user says /pause, pause session, pause this, OR wants to step away knowing they will come back to the same conversation."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Pause Skill
|
|
7
|
+
|
|
8
|
+
USE WHEN user says /pause, pause session, pause this, OR wants to step away knowing they will come back to the same conversation.
|
|
9
|
+
|
|
10
|
+
### The rule that matters
|
|
11
|
+
|
|
12
|
+
**The checkpoint must be written to a FILE before it is printed.** Printing it to the
|
|
13
|
+
terminal does not persist it. If you print first and the session ends, the checkpoint is
|
|
14
|
+
gone and the user has to copy it off the screen by hand — which is the exact failure this
|
|
15
|
+
procedure exists to prevent.
|
|
16
|
+
|
|
17
|
+
Order is: **compose → write file → persist via CLI → print**. Never reorder.
|
|
18
|
+
|
|
19
|
+
### Procedure
|
|
20
|
+
|
|
21
|
+
1. **Compose the checkpoint.** It must be genuinely useful to a session that has none of
|
|
22
|
+
your context. Include:
|
|
23
|
+
|
|
24
|
+
- **Shipped/completed** — what actually landed, with versions or commit refs
|
|
25
|
+
- **Open decisions** — anything blocked on a judgement call, stated as the question
|
|
26
|
+
- **In flight** — built but not installed, written but not tested, etc.
|
|
27
|
+
- **Watch items** — changes whose effects need observing, and why
|
|
28
|
+
- **Cross-session work** — what other sessions are owed, or owe you
|
|
29
|
+
|
|
30
|
+
Be specific. "Continue the refactor" is worthless; "the poller is built and tested but
|
|
31
|
+
`pai task schedule install` has not been run" is a checkpoint.
|
|
32
|
+
|
|
33
|
+
2. **Write it to a file** with the Write tool:
|
|
34
|
+
|
|
35
|
+
```
|
|
36
|
+
/tmp/pai-checkpoint-<session-id>.md
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Use the `sessionId` from the startup system reminders for `<session-id>`. Write the
|
|
40
|
+
markdown body only — no `## Continue` heading and no `## Pause Checkpoint` heading.
|
|
41
|
+
PAI adds those.
|
|
42
|
+
|
|
43
|
+
3. **Persist it** via Bash:
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
pai pause --body-file /tmp/pai-checkpoint-<session-id>.md --session-id <session-id>
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
This writes the body into `## Continue` in the project TODO.md **and** appends it to
|
|
50
|
+
the current session note, then prints the safe-exit reminder.
|
|
51
|
+
|
|
52
|
+
**If this command errors, do not proceed to step 4.** It refuses to write a
|
|
53
|
+
metadata-only checkpoint, so an error here is the difference between a real checkpoint
|
|
54
|
+
and a lost one. Fix the cause and re-run.
|
|
55
|
+
|
|
56
|
+
4. **Print the checkpoint to the user** — the same content you wrote in step 2, plus:
|
|
57
|
+
|
|
58
|
+
```
|
|
59
|
+
To resume: claude --resume <session-id>
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
5. **Tell the user:**
|
|
63
|
+
"Now type `/exit` to safely close the session. DO NOT press Ctrl+C — that bypasses PAI
|
|
64
|
+
stop-hook and orphans the session so it cannot be resumed."
|
|
65
|
+
|
|
66
|
+
### Notes
|
|
67
|
+
|
|
68
|
+
- `--session-id` is what makes `claude --resume` recoverable from TODO.md alone, and it is
|
|
69
|
+
also the key that protects your checkpoint from being overwritten. Always pass it.
|
|
70
|
+
- The next session receives this checkpoint automatically: the SessionStart hook reads
|
|
71
|
+
`## Continue` and injects it. The user does not have to say "go" for it to arrive.
|
|
72
|
+
- The session-stop hook runs `pai session handover` on exit. It will **not** overwrite the
|
|
73
|
+
checkpoint you just wrote — preservation is keyed on the session UUID, which survives the
|
|
74
|
+
note rename and renumber that the same hook performs a few steps earlier. Unattributed
|
|
75
|
+
content in older blocks is carried forward rather than dropped.
|
|
76
|
+
- A rolling autosave (`pai session autosave`, wired to UserPromptSubmit and PostToolUse)
|
|
77
|
+
keeps a mechanical checkpoint fresh throughout the session, so an interrupted session
|
|
78
|
+
still leaves something behind. It never replaces an authored checkpoint for the same
|
|
79
|
+
session — yours always wins.
|
|
80
|
+
- `--no-body` exists for deliberate metadata-only checkpoints. Do not reach for it to work
|
|
81
|
+
around a failure in step 2 or 3.
|
|
82
|
+
- Do **not** hand-place state below the generated header lines as a workaround. That was
|
|
83
|
+
necessary when the block got clobbered; it no longer is, and `--body-file` puts the
|
|
84
|
+
content somewhere that survives.
|