@ninjaxtools/slopdex 0.10.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -7,12 +7,12 @@ For installation, command examples, configuration, and result interpretation, se
7
7
  | Module | Responsibility |
8
8
  | --- | --- |
9
9
  | `src/cli.ts` | Argument parsing, configuration, automatic refresh/recovery, diagnostics, and output selection. |
10
- | `src/code-index.ts` | Index lifecycle, file preparation, Git/working-tree reconciliation, embedding/summary caching, and search facade. |
10
+ | `src/code-index.ts` | Index lifecycle, file preparation, Git/working-tree reconciliation, embedding/description caching, and search facade. |
11
11
  | `src/parser/` | Language dispatch, native Tree-sitter extraction, callable identity, and recoverable diagnostics. |
12
12
  | `src/source-policy.ts`, `src/gitignore.ts` | Supported paths, built-in/config exclusions, nested ignore rules. |
13
13
  | `src/git/repository.ts` | Git commits, trees, blobs, diffs, ancestry, and working-tree changes. |
14
- | `src/embeddings/`, `src/summaries/` | Provider requests and provider profiles. |
15
- | `src/storage/database.ts` | SQLite schema, migrations, transactions, metadata, and vector queries. |
14
+ | `src/embeddings/`, `src/descriptions/` | Provider requests and provider profiles. |
15
+ | `src/storage/database.ts` | SQLite schema, durable artifact caches, transactions, metadata, and vector queries. |
16
16
  | `src/search/` | Analysis scoring selection and cross-index neighbor discovery. |
17
17
  | `src/analysis/cohesion.ts` | Physical distance, gap scores, aggregate affinity, and groups. |
18
18
  | `src/format.ts` | Human-readable search, cluster, and cohesion output. |
@@ -44,18 +44,20 @@ The CLI reconciles an index before most commands. Library callers choose when to
44
44
 
45
45
  Git updates resolve the target commit, validate ancestry against the checkpoint, reconcile committed blobs, and optionally overlay current working-tree files when the target equals HEAD. Overlays account for staged, unstaged, untracked, renamed, and deleted paths. The saved checkpoint remains the committed base. Historical/other-branch targets are committed-only. Explicit `updateFiles` operations do not advance the checkpoint.
46
46
 
47
- File preparation and provider calls precede database updates. Generation checks detect concurrent index changes; working-tree updates also verify source and ignore-rule stability. SQLite transactions apply related file, function, vector, summary, diagnostic, and checkpoint changes together. Failed initialization removes its incomplete index artifacts.
47
+ Tree-sitter results and each valid provider result are committed to content-addressed cache tables immediately. Generation checks detect concurrent logical index changes; working-tree updates also verify source and ignore-rule stability. A separate SQLite transaction atomically applies related file, function, diagnostic, description-reference, and checkpoint changes. Failed or aborted initialization retains the valid database and its cache rows so the next invocation resumes without repeating completed work.
48
48
 
49
- Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable WAL, foreign keys, and normal synchronization. The schema contains:
49
+ Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable WAL, foreign keys, and full synchronization. The schema contains:
50
50
 
51
- - `metadata`: repository root, embedding and summary profiles, generation, checkpoint, schema version, and feature/scan state.
51
+ - `metadata`: repository root, embedding and description profiles, generation, checkpoint, schema version, and feature/scan state.
52
52
  - `files`: paths, content hashes, blob IDs, source mode, previous paths, language, and size.
53
- - `functions`: identity, names, signatures, locations, source, provenance, and embedding/summary references.
54
- - `embeddings` and `summary_embeddings`: cached vectors, with summary text in the latter.
53
+ - `functions`: identity, names, signatures, locations, source, provenance, and embedding/description references.
54
+ - `embeddings`: code, description, and query vectors keyed by embedding profile, operation, and exact input.
55
+ - `description_cache`: generated description text keyed independently by description profile and complete source context.
56
+ - `parse_cache`: successful Tree-sitter extraction results keyed by parser strategy, path, and file-content hash.
55
57
  - `callable_provenance`: first-seen committed source identity.
56
58
  - `indexing_errors`: diagnostics associated with files.
57
59
 
58
- Current schema version is `4`. Supported older versions migrate automatically; the diagnostics migration marks a full rescan pending. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. CLI `--force-reindex` recreates incompatible indexes, preserving enabled OpenAI summary settings for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
60
+ Current schema version is `6`. Earlier schemas are intentionally incompatible and require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. For a schema-6 index, `--force-reindex` clears logical index state while retaining content-addressed caches; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
59
61
 
60
62
  ### Diagnostics
61
63
 
@@ -63,7 +65,7 @@ Diagnostics cover parse errors, parser exceptions, extraction failures, read fai
63
65
 
64
66
  Diagnostics commit with their corresponding file update. Updates retry failed files even when their Git blobs are unchanged; successful replacement, deletion, and exclusion clear failures. Standalone readers inspect saved diagnostics without constructing an embedding provider. The CLI's exit handler reports remaining failures for source and target indexes; version exits before registering that handler.
65
67
 
66
- ## Embeddings and summaries
68
+ ## Embeddings and descriptions
67
69
 
68
70
  Embedding profiles consist of provider, model, dimensions, and strategy version. Cross-index analysis requires matching profiles. Vectors are validated and normalized before storage/search.
69
71
 
@@ -74,19 +76,19 @@ Embedding inputs identify language, callable kind, qualified symbol, signature,
74
76
 
75
77
  Purpose generation uses OpenAI's Responses API with `gpt-5.6-sol` and strategy `callable-purpose-v1`. The prompt asks for one to three sentences describing responsibility and visible relationships, using repository name, path, callable source, and full file context. Requests use `store: false`. Generated text is embedded with the configured embedding provider.
76
78
 
77
- Summary inputs include contextual and profile information so file-context, path, or model changes invalidate relevant cached results. Unchanged inputs reuse summaries and vectors. `useSummaries` persists the profile and enabled state; `disableSummaries` turns automatic updates and summary search/scoring off without deleting cached summaries. Summary generation and embedding preparation finish before the corresponding database transaction, preventing partially updated callable records on provider failure. Deleting a function removes it from summary search.
79
+ Description inputs include contextual and profile information so file-context, path, or model changes invalidate relevant cached results. Description text identity is independent of the embedding profile, allowing an embedding-model change to reuse generation output while producing the required new vector. Every validated description and vector is cached before indexing continues. `useDescriptions` persists the profile and enabled state; `disableDescriptions` turns automatic updates and description search/scoring off without deleting cached artifacts. Function references are attached only by the final logical transaction, so provider failure cannot expose partially updated callable records. Deleting a function removes it from description search but retains reusable cache rows.
78
80
 
79
81
  ## Similarity and analysis
80
82
 
81
- `search` embeds a query and searches code vectors; `search-summary` searches summary vectors. Analysis uses code-only cosine similarity unless all indexed callables have enabled summary embeddings. Cross-index analysis requires completeness on both sides. Summary-generator models may differ across indexes even though embedding profiles must match.
83
+ `search` embeds a query and searches code vectors; `search-description` searches description vectors. Analysis uses code-only cosine similarity unless all indexed callables have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
82
84
 
83
- When summaries are complete:
85
+ When descriptions are complete:
84
86
 
85
87
  ```text
86
- similarity = 0.5 * codeSimilarity + 0.5 * summarySimilarity
88
+ similarity = 0.5 * codeSimilarity + 0.5 * descriptionSimilarity
87
89
  ```
88
90
 
89
- Both component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and summary profiles.
91
+ Both component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and description profiles.
90
92
 
91
93
  Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
92
94
 
@@ -102,7 +104,7 @@ separationWeight = 1 - exp(-physicalDistance / 2)
102
104
  cohesionGap = semanticWeight * separationWeight
103
105
  ```
104
106
 
105
- Affinity ratios and mean distance are weighted by `semanticWeight`. Summary metrics use all qualifying edges before output limiting. File reports aggregate internal, same-folder, and external affinity for selected-source files. Pairs rank by gap and tie-breakers; groups are connected components of reported pairs. Source/test classification is a path heuristic, not a dependency or call-graph analysis.
107
+ Affinity ratios and mean distance are weighted by `semanticWeight`. Cohesion metrics use all qualifying edges before output limiting. File reports aggregate internal, same-folder, and external affinity for selected-source files. Pairs rank by gap and tie-breakers; groups are connected components of reported pairs. Source/test classification is a path heuristic, not a dependency or call-graph analysis.
106
108
 
107
109
  ## Library API
108
110
 
@@ -126,12 +128,12 @@ try {
126
128
  }
127
129
  ```
128
130
 
129
- Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `JinaEmbeddingProvider`, `OpenAISummaryProvider`, error types, and the contracts in `src/types.ts`. Standalone update/search helpers wrap the corresponding index methods.
131
+ Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `JinaEmbeddingProvider`, `OpenAIDescriptionProvider`, error types, and the contracts in `src/types.ts`. Standalone update/search helpers wrap the corresponding index methods.
130
132
 
131
133
  - Use `updateFromWorkingTree()` when Git is unavailable. Unlike the CLI, the library does not automatically refresh before queries or fall back from Git.
132
134
  - Set `sourceFilter.nameRegex` for source-only cross-search/cohesion filtering; combine it with `path` and a filter type (`all`, `changed-since`, or `uncommitted`). `changed-since` also accepts `uncommitted: true`.
133
135
  - Top-level analysis `nameRegex` filters both sources and candidates. Query `SimilaritySearchOptions.nameRegex` filters result names before limiting.
134
- - Call `await index.useSummaries()`, then `await index.searchSummary({ query: "maintain the repository index" })`. Select a model via `summaryProvider: new OpenAISummaryProvider({ model: "gpt-5.6-sol" })` in index options. Custom providers implement `EmbeddingProvider` or `SummaryProvider`.
136
+ - Call `await index.useDescriptions()`, then `await index.searchDescription({ query: "maintain the repository index" })`. Select a model via `descriptionProvider: new OpenAIDescriptionProvider({ model: "gpt-5.6-sol" })` in index options. Custom providers implement `EmbeddingProvider` or `DescriptionProvider`.
135
137
  - Inspect failures through `index.indexErrors()` or exported `readIndexErrors(indexPath)` without a provider. Records use `IndexingError`.
136
138
  - Public cohesion reports retain indexed-function data; the CLI presents compact function references and includes source bodies only with `--include-source`.
137
139
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ninjaxtools/slopdex",
3
- "version": "0.10.0",
3
+ "version": "0.11.0",
4
4
  "description": "Tree-sitter callable-level semantic indexing, search, and duplicate discovery for Python, JavaScript, TypeScript, Rust, Go, Java, and C",
5
5
  "license": "MIT",
6
6
  "repository": {