@ninjaxtools/slopdex 0.11.0 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -57,7 +57,7 @@ Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable
57
57
  - `callable_provenance`: first-seen committed source identity.
58
58
  - `indexing_errors`: diagnostics associated with files.
59
59
 
60
- Current schema version is `6`. Earlier schemas are intentionally incompatible and require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. For a schema-6 index, `--force-reindex` clears logical index state while retaining content-addressed caches; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
60
+ Current schema version is `7`. Schema 6 is migrated in place by adding file-description state; earlier schemas require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. Forced rebuilds clear logical index state while retaining content-addressed caches when possible; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
61
61
 
62
62
  ### Diagnostics
63
63
 
@@ -74,21 +74,21 @@ Embedding profiles consist of provider, model, dimensions, and strategy version.
74
74
 
75
75
  Embedding inputs identify language, callable kind, qualified symbol, signature, documentation, and source. For Python, a function's first-statement docstring is included in a separate `documentation` section in addition to remaining part of the callable source.
76
76
 
77
- Purpose generation uses OpenAI's Responses API with `gpt-5.6-sol` and strategy `callable-purpose-v1`. The prompt asks for one to three sentences describing responsibility and visible relationships, using repository name, path, callable source, and full file context. Requests use `store: false`. Generated text is embedded with the configured embedding provider.
77
+ Purpose generation uses the AI SDK and strategy `callable-purpose-v2`. The default is OpenAI's Responses API with `gpt-5.6-sol`; OpenCode Zen and Go are also selectable, defaulting to `gpt-5.6-sol` and `gpt-5.6-luna`. OpenCode catalog models are routed through their published protocol using the OpenAI Responses, Anthropic Messages, Google Generative AI, or OpenAI-compatible adapter. Generation opens one conversation per file: stable instructions and complete file context form the prefix, the first request describes the file overall, then callable prompts and generated answers are appended sequentially in source order. This avoids repeating the file within a request and gives provider prompt caches an increasingly large reusable prefix. Responses requests use `store: false`. File and callable text are embedded with the configured embedding provider.
78
78
 
79
- Description inputs include contextual and profile information so file-context, path, or model changes invalidate relevant cached results. Description text identity is independent of the embedding profile, allowing an embedding-model change to reuse generation output while producing the required new vector. Every validated description and vector is cached before indexing continues. `useDescriptions` persists the profile and enabled state; `disableDescriptions` turns automatic updates and description search/scoring off without deleting cached artifacts. Function references are attached only by the final logical transaction, so provider failure cannot expose partially updated callable records. Deleting a function removes it from description search but retains reusable cache rows.
79
+ Description inputs include contextual and profile information so file-context, path, model, or generation-strategy changes invalidate relevant cached results. Description text identity is independent of the embedding profile, allowing an embedding-model change to reuse generation output while producing the required new vector. Every validated description and vector is cached before indexing continues. If generation resumes partway through a file, completed file/callable prompts and cached answers are replayed locally before the next request so the conversation prefix remains equivalent. Ordinary source updates retain the previous file description and its source hash, making staleness explicit without incurring automatic regeneration; `reindex-files` replaces stale file descriptions, optionally continuing through callable regeneration. `useDescriptions` persists the profile and enabled state; `disableDescriptions` turns automatic updates and description search/scoring off without deleting cached artifacts. Function references are attached only by the final logical transaction, so provider failure cannot expose partially updated callable records. Deleting a function removes it from description search but retains reusable cache rows.
80
80
 
81
81
  ## Similarity and analysis
82
82
 
83
- `search` embeds a query and searches code vectors; `search-description` searches description vectors. Analysis uses code-only cosine similarity unless all indexed callables have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
83
+ `search` embeds a query and, when descriptions are complete, scores code, callable-description, and containing-file-description vectors. `search-description` scores callable and file descriptions. Analysis uses code-only cosine similarity unless all indexed callables and files have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
84
84
 
85
85
  When descriptions are complete:
86
86
 
87
87
  ```text
88
- similarity = 0.5 * codeSimilarity + 0.5 * descriptionSimilarity
88
+ similarity = (codeSimilarity + descriptionSimilarity + fileDescriptionSimilarity) / 3
89
89
  ```
90
90
 
91
- Both component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and description profiles.
91
+ All component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and description profiles. Cohesion JSON uses schema version 3 for the three-component scoring contract.
92
92
 
93
93
  Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
94
94
 
@@ -133,7 +133,7 @@ Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation
133
133
  - Use `updateFromWorkingTree()` when Git is unavailable. Unlike the CLI, the library does not automatically refresh before queries or fall back from Git.
134
134
  - Set `sourceFilter.nameRegex` for source-only cross-search/cohesion filtering; combine it with `path` and a filter type (`all`, `changed-since`, or `uncommitted`). `changed-since` also accepts `uncommitted: true`.
135
135
  - Top-level analysis `nameRegex` filters both sources and candidates. Query `SimilaritySearchOptions.nameRegex` filters result names before limiting.
136
- - Call `await index.useDescriptions()`, then `await index.searchDescription({ query: "maintain the repository index" })`. Select a model via `descriptionProvider: new OpenAIDescriptionProvider({ model: "gpt-5.6-sol" })` in index options. Custom providers implement `EmbeddingProvider` or `DescriptionProvider`.
136
+ - Call `await index.useDescriptions()`, then `await index.searchDescription({ query: "maintain the repository index" })`. Select a model via `descriptionProvider: new OpenAIDescriptionProvider({ model: "gpt-5.6-sol" })` in index options. Custom description providers implement both stateless file/callable methods and may add `startFile()` for contextual sessions.
137
137
  - Inspect failures through `index.indexErrors()` or exported `readIndexErrors(indexPath)` without a provider. Records use `IndexingError`.
138
138
  - Public cohesion reports retain indexed-function data; the CLI presents compact function references and includes source bodies only with `--include-source`.
139
139
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ninjaxtools/slopdex",
3
- "version": "0.11.0",
3
+ "version": "0.13.0",
4
4
  "description": "Tree-sitter callable-level semantic indexing, search, and duplicate discovery for Python, JavaScript, TypeScript, Rust, Go, Java, and C",
5
5
  "license": "MIT",
6
6
  "repository": {
@@ -43,6 +43,11 @@
43
43
  "check:parser-parity": "node --import tsx scripts/check-parser-parity.mjs"
44
44
  },
45
45
  "dependencies": {
46
+ "@ai-sdk/anthropic": "^4.0.53",
47
+ "@ai-sdk/google": "^4.0.69",
48
+ "@ai-sdk/openai": "^4.0.66",
49
+ "@ai-sdk/openai-compatible": "^3.0.48",
50
+ "ai": "^7.0.99",
46
51
  "ignore": "^7.0.9",
47
52
  "install": "^0.13.0",
48
53
  "js-tiktoken": "^1.0.21",